sist newsletter vol.1, no. 4 publications the us national technical information service (ntis) announces the publication of a directory of computer software applications-energy . it is 80 page reference source which presents over 300 abstracts of available energy-related reports and computer programs developed from federally sponsored research. the abstracts detail what software is available in fossil, solar, nuclear, geothermal ocean thermal and other energy areas. the programs range in subject matter from comparative costs to social impact and are offered in several computer languages. you receive complete software program instructions (or a program source) and/or a report for a cost of a paper copy (average cost is $6). the ntis brochure suggests that the directory is an ideal source for reviewing available energy software programs for guidance before undertaking expensive program development of your own. the directory costs $20 and is available in paper copy or microfiche. ntis order no. is pb-264 200. write to ntis, us department of commerce, springfield, va 22161. the analysis of textual data has recently benefitted from the methods and tools which are the result of computer analysis. editions du cnrs announces a new text on the analysis of textual data: analyse et validation dans 1 'etude des donn^es textuelles . the brief description says that the objective of this work is to put forward a certain number of theoretical perspectives and concrete experiences that present the analysis of texts in the light of data-handling techniques and be a reference to the most valued acquisitions of contemporary linguistics". the book groups 14 papers from french and foreign specialists working in diverse fields of study but all of whom are concerned with the same fundamental questions. the more strictly methodological or theoretical parts presented in each of these contributions are illustrated by reports on particular experiences related to specific branches of knowledge and objectives (archaelogy and history, folklore, literature, anthropology). for further information, write editions du cnrs, 15 quai anatole france, 75700 paris, france; in canada: presses de i'universite de montreal, case postale 6128, montreal 101; usa: smpf, 14 east 60th street, new york, ny 10022. the social science data archive of the social science library at yale university announces the revision and updating of its directory of data holdings and services . cost is $5.00 and prepayment is requested. checks should be made payable to yale university, but sent to the social science data archive, social science library, box 1958, yale station, new haven, connecticut 06520. iassist quarterly fall winter 2012 5 iassist quarterly editor’s notes special issue: the organizational dimension of digital preservation welcome to the special double issue 3 & 4 of the iassist quarterly (iq) volume 36 (2012). this special issue addresses the organizational dimension of digital preservation as it was presented and discussed at the iassist conference in may 2013 in cologne, germany. the two guest editors astrid recker and natascha schumann from the gesis leibniz institute for the social sciences in cologne have earned special thanks. if you find their names familiar it is because they co-wrote a paper in the iq 36-2. they are concerned with data preservation and curation at the data archive for the social sciences, and as a member of the archive and data management training center, astrid also trains others in these areas. furthermore, they co-chaired the panel on ‘beyond bits and bytes: the organizational dimension of digital preservation’ at iassist 2013, both also participating as panelists in the session. they have now persuaded the other panelists to contribute to this combined special issue. thanks also therefore to michelle lindlar, stefan strathmann and achim oßwald, and yvonne friese. articles for the iassist quarterly are always very welcome. they can be papers from iassist conferences or other conferences and workshops, from local presentations or papers especially written for the iq. authors are permitted “deep links” where you link directly to your paper published in the iq. chairing a conference session with the purpose of aggregating and integrating papers for a special issue iq is also much appreciated as the information reaches many more people than the session participants, and will be readily available on the iassist website at http://www.iassistdata.org. authors are very welcome to take a look at the instructions and layout: http://iassistdata.org/iq/instructions-authors. authors can also contact me via e-mail: kbr@sam.sdu.dk. should you be interested in compiling a special issue for the iq as guest editor(s) i will also be delighted to hear from you. karsten boye rasmussen january 2014 editor fjl^sist newsletter vol. 1, no. 3 chairperson's report the second north american meeting was held in toronto, canada, may 11 -12th, 1977. sharon henry, the canadian secretariat was able to secure funds from canada council to support the travel of participants delivering formal presentations at the opening session. sixty-five individuals attended the conference. abstracts of the presentations were submitted prior to the conference and translated to french. the english and french versions of the abstracts, along with the complete english text of the presentations were available at registration. participants of the conference will automatically receive a copy of the proceedings and members of lassist will be able to purchase copies at a reduced rate. the registration fee for the conference will be used to defray reproduction costs of the proceedings . any additional monies from fees will be forwarded to the treasurer, ed hanis. the remainder of the meeting was devoted to the action group workshops. abstracts of the presentations and the progress reports resulting from the action group workshops are published in the body of this newsletter . per nielsen, the west european secretariat, has organized a european lassist action group workshop to be held at the danish data archives in copenhagen, june 26-29, 1977. the european ags will build on the products of the us and canadian efforts as well as previous work carried out by european archives, particularly in the area of classification and documentation. proceedings of this meeting will be published. [see newsletter section on the danish data archives sponsored lassist west european meeting.] the february 1978 lassist meeting has been scheduled for the 8th-12th, at nordic hills, itasca, illinois. the program is detailed in the newsletter section describing upcoming meetings. the chicago area was selected as the site of the meetings to permit maximum participation, given its central location in north america. topics for presentation at the symposia are invited. please submit the topics to tony falsetto, public readable archives, public archives of canada, 2850 cedarwood dr., apt. 909, ottawa, ontario, canada, with a copy to carolyn geda. the lassist program for the meetings to be held in conjunction with the international sociological association in august 1978, has been approved. the schedule of activities is presented in the newsletter section describing upcoming meetings. initial monitoring of lassist receipts and expenditures verifies that there are adequate funds to cover the cost of production of the newsletter , mailing costs, and subscriptions to s s data . registration fees for the conferences are sufficient to support the costs of the meeting and publication of the proceedings . a full treasurer's report will be included in the next issue of the newsletter . as has been noted in previous newsletters , the working manual for cataloguing mrdf, compiled by sue dodd, university of north carolina, will be the first lassistrelated publication. in recognition of this work and forthcoming pieces, the steering committee has proposed a set of publishing policies and has appointed an editorial board to work with authors and publishing houses in whatever capacity is facilitative. the editorial board is composed of elina almasy, carolyn geda, edward hanis, cees middendorp, and david nasitir. vol252 4 iassist quarterly fall 2001 editor's notes the data documentation initiative has had workshops, special sessions and general good publicity in iassist connections since the 1995 iassist conference in québec. this issue of iassist quarterly brings you a paper from the iassist/ifdo conference last year in amsterdam. the ddi has evolved and in a session on “ddi tables” this paper from the university of california (berkeley) and from the california digital library at the same university appeared. the paper is titled “using ddi extensions as an intermediary for data storage and data display” and is written and prepared by patricia cruse, marsha fanshier, fredric gey, and margaret low. in the paper they are addressing the problem of providing metadata for multidimensional files (or aggregate files or tables). the solution based upon use of xml is shown in the paper. at the iassist/ifdo conference last year a session was devoted to “digital archiving”. i believe that the subject is much too important to leave to archivists alone and it also has the attention of governments and the eu. the paper from concha fernández de la puente at the european commission, dg information society is titled “the european union initiatives in the digital archive area: achieving the information society” and the point of departure of the paper is that initiatives from the european union has had the goal of ensuring that all european citizens have easy access to information. the many initiatives in the areas of electronic document management and exchange are reviewed in the paper. among the initiatives are dlm-forum which have had widely participation from archivists. the paper surely points further in its conclusion where the citizen-friendliness of on-line government is mentioned as well as “the ideas and values which have shaped the european union”. another session at the same conference was on “qualitative data”. i could mention that this session also had its paper on ddi, but here we are looking at a very different content than the values of variables or columns that the ddi was invented for. culture often implies a notion of being analog, personal, and difficult to disseminate. tiina mahlamäki from the university of turku describes in the paper “from the field to the net: cataloguing and digitising cultural research material” how tape-recorded interviews was catalogued and transferred for use on the internet. tiina mahlamäki explains that the oral tradition became an important aspect of the folkloristic research in the nordic countries in the mid-1960ies and an important part of the fieldwork. the individualistic trait of culture also applied to the culture researchers as every project followed its own principles of archiving. the paper describes the project of standardization and digitisation of this cultural research material. the last paper in this issue of iassist quarterly is from the opening plenary at the amsterdam conference. the title of the plenary was “interand intranational archives in the new millennium: issues, strategies and models” where yvette hackett from the national archives of canada presented “a national research data management strategy for canada: the work of the national data archive consultation working group”. yvette hackett gives an account of how recommendations for preservation of data for the social sciences and the humanities lead to a working group and project that are investigating questions about how to meet the needs of the research community and to take advantage of new technologies. when this issue of the iassist quarterly reaches you the 2002 conference in storrs, connecticut will be over. i am certain it has been a successful conference and that you can look forward to receive papers from that conference in the coming issues of the iassist quarterly. may 2002. karsten boye rasmussen iassist quarterly fall 2011 5 iassist quarterly editor’s notes sharing data and building information with this issue (volume 35-3, 2011) of the iassist quarterly (iq) we return to the regular format of a collection of articles not within the same specialist subject area as we have seen in recent special issues of iq. naturally the three articles presented here are related to the iq subject area in general, as in: assisting research with data, acquiring data from research, and making good use of the user community. this last topic could also be spelled “involvement”. the hope is that these articles will carry involvement to the iassist community, so that the gained knowledge can be shared and practised widely. “mind the gap” is a caveat to passengers on the london underground. the authors of this article are susan noble, celia russell and richard wiseman, all affiliated with esds-international hosted by mimas at the university of manchester in the uk. the esds, standing for “economic and social data service”, are extending their reach beyond the uk. in the article “mind the gap: global data sharing” they are looking into how today’s research on the important topics of climate change, economic crises, migration and health requires cross-national data sharing. clearly these topics are international (e.g. the weather or air pollution does not stop at national borders), but the article discusses how existing barriers prevent global data sharing. the paper is based on a presentation in a session on “sharing data: high rewards, formidable barriers” at the iassist 2009 conference. it is demonstrated how even international data produced by intergovernmental organizations like the international monetary fund, the international energy agency, oecd, the united nations and the world bank are often only available with an expensive subscription, presented in complex incomprehensible tables, through special interfaces; such barriers are making the international use of the data difficult. because of missing metadata standards it is difficult to evaluate the quality of the dataset and to search for and locate the data resources required. the paper highlights the development of e-learning materials that can raise awareness and ease access to international data. in this case the example is e-learning for the “united nations millennium development goals”. the second paper is also related to the sharing of data with an introduction to the international level. “the research-data-centre in research-data-centre approach: a first step towards decentralised international data sharing” is written by stefan bender and jörg heining from the institute for employment research (iab) in nuremberg, germany. in order to preserve the confidentiality of single entities, access to complete datasets is often restricted to monitored on-site analysis. although off-site access is facilitated in other countries, germany has relied on on-site security. however, an opportunity has been presented where research data centre sites are placed at statistical offices around germany, and also at a michigan centre for demography. the article contains historical information on approaches and developments in other countries and has a special focus on the german solution. the project will gain experience in the complex balance between confidentiality and analysis, and the differences between national laws. the paper by stuart macdonald from edina in scotland originated as a poster session at the iassist 2010 conference. the name of the paper is “addressinghistory: a web2.0 community engagement tool and api”. the community consists of members within and outside academia, as local history groups and genealogists are using the software to enhance and combine data from historical scottish post office directories with largescale historical maps. the background and technical issues are presented in the paper, which also looks into issues and perspectives of user generated content. the “crowdsourcing” tool did successfully generate engagement and there are plans for further development, such as upload and attachment of photos of people, buildings, and landmarks to enrich the collection. articles for the iq are always very welcome. they can be papers from iassist conferences or other conferences and workshops, from local presentations or papers especially written for the iq. if you don’t have anything to offer right now, then please prepare yourself for the next iassist conference and start planning for participation in a session there. chairing a conference session with the purpose of aggregating and integrating papers for a special issue iq is much appreciated as the information in the form of an iq issue reaches many more people than the session participants and will be readily available on the iassist website at http://www.iassistdata.org. authors are very welcome to take a look at the instructions and layout: http://iassistdata.org/iq/instructions-authors authors can also contact me via e-mail: kbr@sam.sdu.dk. should you be interested in compiling a special issue for the iq as guest editor(s) i will also be delighted to hear from you. karsten boye rasmussen december 2011 editor 4 iassist quarterly fall 2009 editor’s notes torture, numbers, and digital tape welcome to the iassist quarterly (iq) volume 33 number 3. one third of 100 volumes! we have reached the autumn of 2009 in our iq chronology. this issue is very much about quantitative approaches, and without getting into a discussion on precision and accuracy here, we can report from our own world that most other activities presently experience themselves as being in autumn 2010. by now, you have probably learned that it is unwise to set your clock by the iq. with this issue concentrating on quantitative investigations and the use of statistics, i came to think of the mark twain citation "there are three kinds of lies: lies, damned lies, and statistics." mark twain did not take credit for the remark; he attributed it to disraeli but that has been questioned and that is another story. however, mark twain showed his goodwill by making the reference and demonstrating free dissemination of knowledge. with the violent title of the first article (see below), the article on numbers and statistics, and the report from a national statistics agency, the "torture, numbers, and digital tape" title of this editorial surfaced. a fourth article bends and uses a famous film title, so i used and bent another film title for the heading for this introduction. (a remark to the non-cineasts: steven soderbergh directed in 1989 a motion picture called "sex, lies, and videotape"). with that i welcome you to an issue of the iq that is filled with tales about data and their uses. "torturing nurses with data". now, that's a title to remember! maybe it could be a song too! kristi thompson, data librarian at the leddy library of the university of windsor, presented this paper at the iassist 2008 conference at stanford in the session "numeracy, quantitative reasoning and teaching about data". she describes and discusses two iterations in the creation of a short module in quantitative research in the programme for a masters in nursing. the module included a lecture as well as hands-on practice. in the second version of the module the hands-on part was more extensive and more structured including an analysis assignment using a preselected data set. the paper explains what worked and what didn’t work. we can reveal that the article includes an analysis of feedback from an anonymous questionnaire filled out by the students following the second unit. as a data provider for eager students kristi thompson speaks of the "impossible dataset" the dream dataset searched for by students that is very unlikely to exist as it would break nearly all confidentiality codes. this was remedied by spending more time on the realities of data collection. the second iteration had learned from the first but was also faced with its own quantitative problems as the number of students in the class was now much higher. for both of these reasons the student task was changed into choosing from a selection of prepared datasets and accompanying research questions. however, the module then was experienced more as a class in statistics than in quantitative research. this wisdom is what triggered the "torturing nurses with data" title. there needs to be more direct interest from the students. this can probably be achieved by again letting the students formulate their own research questions and thus fulfilling the subtitle of the article: "building a successful quantitative research module". the second article is also within the quantitative area. with the short title of "numbers", flavio bonifacio takes us on a tour of "numbers" from the representation of reality, through processing as a system of guarantee and numbers as model parts, to what is called a path between numbers and society. flavio bonifacio works at metis ricerche in torino, a company which does data collection and processing, forecasting, and analysis, and has used sas software in several of its projects. he starts by addressing the issue of objectivity and gives what looks as the concluding insight: "how numbers are not more true than other representations of reality, that every statement about reality must be responsibly supported, that it is necessary to support this responsibility with an explicit agreement between the producers of the data, that this agreement must be institutionally granted, that only this agreement can make it possible not to surrender the objectivity of the measurements to the tastes of the moment." the article is on "the misuse of numbers" and shows with some citations how the same event in a public bus takes many different forms in queneau's "exercises of style". the obvious differences in texts can lead to the assumption – and a jump to a false conclusion that numbers are "more true". in the session on "building on data: resources, tools and applications" at the 2009 iassist conference in tampere chiu-chuang (lu) chou, a senior special librarian at the data and information service at the university of wisconsin-madison presented her "the good, the bad and the ugly of playing a data custodian". the article concerns the national survey of families and households (nsfh) that is a longitudinal study on family life in the us. the survey has been carried out in three waves in 1987-1988, 1992-1994, and 2001-2003. it is a very expensive data collection that is heavily used; an icpsr database shows that so far the data has been used in more than a thousand publications. the nsfh project has ended but researchers continue to use the data. so data librarians face the job of documentation though they were not involved in the actual project; the people who were involved are now retired or unavailable. the paper describes the content of the three nsfh waves that are concentrated on family living arrangement, marriage, cohabitation, fertility, parenting relations, kin contact and economic and psychological well iassist quarterly fall 2009 5 being. it also includes some measure of the requirement for user support, as the design and structure of the nsfh waves are complex. the article shows that some of this support is quickly done but i noted that more than a quarter of the users' questions took two hours or "substantially" more (my statistics). so there is a great need for support even though most of the contact support takes place through email. the article also includes examples showing that acting as a data custodian for such complicated studies also involves making corrections in the data based upon users' reports. the last article is a presentation from the session "building data archives and user communities: greece, estonia and ethiopia" also at the 2009 iassist conference in tampere. yacob mudesir seid from the central statistical agency of ethiopia (csa) describes the agency as being responsible for providing accurate and timely statistical information for development planning and monitoring purposes. the paper describes the history of the use of information communication technology in the agency's data processing, archiving and dissemination efforts. the csa is considered as one of the leading institutions in ethiopia in utilizing ict and has developed through generations of hardware and software. on the archiving and dissemination side the csa has taken advantage of ddi (data documentation initiative) for improvement of the metadata documentation to an international standard. it is interesting how younger agencies as "late movers" very quickly move to the highest levels of standards. the article shows many graphic examples of ict use at the csa, among them geographical information systems (gis) that are providing easy access for decision makers. articles for the iassist quarterly are very welcome. articles can be papers from iassist conferences, from other conferences, from local presentations, discussion input, etcetera. authors are very welcome to contact me. if you don't have anything to offer right now, then please prepare yourselves for the coming iassist conference. you can start planning for participation in a session there. should you be interested in compiling special issues for the iq as guest editor(s) please also contact me. chairing a conference session with the purpose of aggregating and integrating papers for a special issue iq is much appreciated as the information reaches many more people than the session participants and will be readily available on the iassist website. by the way, if you have not experienced the new website then take a look at http:// www.iassistdata.org. also the iassist blog the iassist communiqué – is found at http://www.iassistblog.org. contact the editor via e-mail: kbr@sam.sdu.dk. karsten boye rasmussen october 2010 . vol273.indd iassist quarterly fall 2003 15 by nancy lemay * abstract as data users become technically knowledgeable, “data specialists”, with no technical support available, should endeavour to offer sophisticated means for acquiring and/or requesting data. providing users with the means to send data requests via the world wide web (web) is a realistic goal for any data specialist, even those without programming capabilities. data requests can either be sent via e-mail or stored in a database. this article will cover the latter option, i.e. storing data requests in a database. by empowering the data user, you are in a way empowering yourself. introduction this article provides guidance for the creation of webbased forms, using the example of the registration form created for the 2003 iassist conference. web-based forms can range from an online survey to a data request form. the underlying concepts presented in this article are pertinent for any database-linked web form. the tools necessary for the creation of such a form should be accessible to most data specialists, who are regularly required to manage information gathered using the web. context in the spring of 2003, the university of ottawa hosted the iassist conference. in light of the high participation rates in the past, the local arrangements committee (lac) anticipated a high participation rate for 2003. the lacʼs priority was therefore to ensure that the webbased registration procedure be simple and efficient. more importantly, a web-based registration procedure would minimize human errors since the number of people involved in processing the registrations and payments would be reduced. capturing the registration information electronically also allowed us to automate the creation of receipts, name tags and banquet tickets, thereby significantly reducing the lacʼs workload. considerations the suggestions provided in this article are meant to limit problems and complications which may be encountered when setting up a web form. it should be kept in mind that there are other ways to successfully produce the same end results. the very first thing required to create a web form is access to a microsoft internet server having msfrontpage extensions and msaccess software installed. the use of a test server is recommended to verify that the web form works properly and to avoid many pitfalls. keep in mind that once the form is live on the web, it is very difficult and inconvenient to make changes to it. it is also difficult to know how the database will react when changes are made while the form is being accessed simultaneously by people via the web. summary of steps to be taken1 once access to a microsoft internet server has been set-up and tested, we can move on to the conceptual information necessary to produce a web application. the first step to link a web form to a database is to create a simple form using msfrontpage (called “yourwebform. htm”). the creation of a simple form in msfrontpage is relatively easy since the whole process is wizard driven. any msfrontpage manual should also explain how to create a simple form. it is highly recommended to create a few textboxes in order to test the success of the operation before creating a complex web form. it is very important to ensure that the textboxes are named properly, preferably without spaces in the textbox names.2 the names given to every textbox should be noted, since this information will be needed for the database. it is also wise to decide upon a naming scheme ahead of time, especially if more than one person will be working on the form or database. for instance, we decided that all textbox names would be capitalized and any spaces replaced with an underscore. this can be particularly useful if you need to go back and make changes or remove textboxes. after creating the form, an msaccess database using the same field names given to the textboxes in the first step should be created. an example is the creation of three fields called first_name, last_name and email. then save the database as “yourdatabase.mdb”. the process is linking a frontpage 2000 web form to an access 2000 database 16 iassist quarterly fall 2003 simple as it is entirely wizard driven. after the creation of the form and database, import the database into the msfrontpage web page. to do so click on file � import option � add file and locate the database “yourdatabase.mdb”. at this point, the wizard will ask whether or not to create a database connection. it is highly recommended to do so. the connection may be created later, but it is always easier to do it at this point, especially using the wizardʼs instructions. a new folder named /fpdb, which will be used to store the database, should then be created. the form created, “yourwebform.htm”, is the applicationʼs data-entry form. a new asp page needs to be created with an sql insert statement. this can be done using the form properties dialog box to create a link to “enter_database_ insert.asp”. this asp page will link the form to the database. to create the “enter_database_insert.asp” page use the insert database results wizard to add the sql insert statement for copying records from the form to the msaccess database. the following is an example of an sql insert statement: insert into names (first_name, last_name, email) values (ʻ%%first_name%%ʼ, ʻ%%last_name%%ʼ, ʻ%%email%%ʼ) verify the query by using the verify query button in the dialog box. save the new asp page as “enter_database_ insert.asp”. at this point, it is necessary to rename “yourwebform.htm” form with an .asp extension. at this point, it should be possible to start testing the form on the local machine, or preferably, on a test server. access to the database the last major concern is attributing access rights to other project team members. unless a second person was actively involved in the setup of the database, only the database administrator should be given access to the msaccess database on the microsoft internet server. this will help ensure database integrity. when setting up our database, access to the information stored in the msaccess database was provided to our registration coordinator by linking an msexcel spreadsheet to the database. we wanted to limit the use of the database as much as possible. therefore, because the registration coordinator needed just the fields pertinent to the registration desk, only selected fields (such as registration information and id numbers) from the msaccess spreadsheet were linked to the excel spreadsheet. this gave the registration coordinator full access to the registration data and the ability to add columns or comments in excel without changing the msaccess database. furthermore, since we had a limited number of spaces available for each workshop, the registration coordinator could track workshop registrations and inform the technical support coordinator when a workshop was full. this allowed us to update the registration form on the web and remove the registration option for that particular workshop. in addition to the tracking capabilities, the registration coordinator could provide the lac with weekly updates about registration numbers and membership details. using msaccess reports for receipts, name tags and banquet tickets the use of a database and web form to capture our registration information allowed us to take advantage of many of msaccessʼs advanced features. for example, the report option in msaccess which enables the retrieval of data from a database while formatting it to meet the reportʼs purpose, or, the addition of graphics, tables and charts to reports. all of our 2003 iassist receipts were automatically created as reports in msaccess. we then simply arranged the registration summary in a formatted table along with the 2003 iassist logo. this meant that the registration coordinator only had to print and send the receipts to conference participants. name tags were also generated using the report options in msaccess. this was done by selecting the pertinent fields, formatting our report to follow the dimensions of precut inkjet name tags paper, and adding the conference logo. conclusion the linking of a web form to a database can be used for numerous tasks not only for conference registrations, but also for other data requests sent via an online form or for online surveys about your service area. the goal of this article was to encourage non-technical data specialists to explore this option in order to provide a sophisticated service. * contact: nancy lemay, data support specialist, geographic information and data centre, university of ottawa. nlemay@uottawa.ca. footnotes 1 the following article was used as a guide through specific steps in front page: slater, william f. iii. database wizardry with frontpage 2000. sept. 1999. < http://www. webtechniques.com/archives/1999/09/data/>. 2 an example is: first_name, last_name and email 4 iassist quarterly 2016 / vol 40 no 3 iassist quarterly editor’s notes being international and proud of it! iassist is proud of being international. these days some us of find it important to emphasize how international collaboration has improved and made our lives more efficient. in the small but around-the-globe-reaching world of iassist, many national data archives have come into existence as well as continuing their development, through friendly international support and spreading of knowledge and good practices among iassisters. so let us cherish the 'international' in iassist. we are proud of the lead 'i' for 'international' in the iassist acronym and have no intention of changing that to 'n' for 'national'. it is also my impression that data archives all over the world simply don't have the facilities for storing 'alternative facts' as they are shy of all kinds of documentation. welcome to the third issue of volume 40 of the iassist quarterly (iq 40:3, 2016). four papers with authors from three continents are presented in this issue. the paper 'demonstrating repository trustworthiness through the data seal of approval' is a summary of a panel session at the iassist 2015 conference in minneapolis with panel members stuart macdonald, ingrid dillo, sophia lafferty-hess, lynn woolfrey, and mary vardigan. the paper has an introduction from dans in the netherlands where the data seal of approval (dsa) originated. cases from the us and south africa are presented and the future of the dsa including possible harmonization with other systems is discussed. dsa certifications are basically consumer guidance, clearly assisting all the involved parties. depositors and funding bodies will be assured that data are reliably stored, researchers can reliably access the data repositories, and repositories are supported in their work of archiving and distribution of data. the second article brings us to the actual use of data. from the uk data service, rebecca parsons and scott summers in 'the role of case studies in effective data sharing, reuse and impact' take us into positive narratives around secondary data. the background is that although the publishing of data is now recognised by funders, the authors find that ‘showcasing’ brings motivation for data sharing and reuse as well as improving the quality of data and documentation. the impact of case studies is all-sided and research, depositing data, and the brand recognition of the uk data service are among the areas investigated. the future is likely to include new case studies developed for use in teaching in schools, with easy linking to datasets, as well as for researchers being assisted to build their own portfolios. the appendix presents case studies on research and impact. in the third article, we are situated in data creation. muhammad f. bhuiyan and paula lackie from carleton college in minnesota write on 'mitigating survey fraud and human error: lessons learned from a low budget village census in bangladesh'. as the 'fraud' term implies, they are looking into the problem of data creators being too creative, but more importantly they are investigating the essential area of data quality. the authors explain how selected technological assets like the use of geographic information systems (gis) and audio-capturing smart pens improved data quality. the use of these tools is exemplified through many scenarios described in the paper. furthermore, a procedure of daily monitoring and fast transcription lead to quick surveyor re-training and dismissal of others, thus minimising data errors. for those interested in false data and its detection, the introduction in particular has valuable references to literature. in the last paper the difficult task of handling images is addressed in 'image management as a data service' by berenica vejvoda, k. jane burpee, and paula lackie. vejvoda and burpee work at mcgill university in montreal. you have already met lackie from carleton college in relation to the third paper above. the 'images' in the article are digital images, and the authors suggest that the knowledge of digital data services across the 'research data lifecycle' also benefits the management of digital images. digital images are numerical data, and the article compares the data, metadata, and paradata of a survey respondent to the information on a digital image. considerations from normal data concerning system formats and storage space also apply to management of images. in the last section the paper introduces copyright issues that are complicated, to say the least. just as reuse of normal data can have ethical angles, it is even more apparent that images can have complicated issues of privacy and confidentiality. papers for the iassist quarterly are always very welcome. we welcome input from iassist conferences or other conferences and workshops, from local presentations or papers especially written for the iq. when you are preparing a presentation, give a thought to turning your one-time presentation into a lasting contribution. we permit authors 'deep links' into the iq as well as deposition of the paper in your local repository. chairing a conference session with the purpose of aggregating and integrating papers for a special issue iq is also much appreciated as the information reaches many more people than the session participants, and will be readily available on the iassist website at http://www.iassistd ata.org. authors are very welcome to take a look at the instructions and layout: http://iassistdata.org/iq/instructions-authors authors can also contact me via e-mail: kbr@sam.sdu.dk. should you be interested in compiling a special issue for the iq as guest editor(s) i will also be delighted to hear from you. karsten boye rasmussen january 2017 editor http://www.iassistdata.org http://www.iassistdata.org http://iassistdata.org/iq/instructions-authors http://www.sam.sdu.dk 34 iassisi quarterly access to unpublished social science data in government departments in the united kingdom by d. a. clarke ' both for the purposes of day-to-day administration and for the formulation of future policy, government departments gather data on a wide range of topics which are potentially valuable for social science research. (social science is here interpreted very broadly, to include economics and government as well as social studies in the more restricted sense.) a few random examples are agricullural economics and farm profits; consumption of food and drink; public service pay and personnel management; social planning; regional and town planning; population trends; the supply and training of teachers; employment statistics; labour mobility; industrial relations; immigration and race relations; safety, health and welfare; manpower planning; local prisons and penal policy; the police. the records of some other quasi-official bodies, e.g. the research councils and the nationalized industries, must similarly contain data of major value for research. 'presented at lassist/ifix) international conference may 1985, amsterdam though some types of material clearly need to remain closed for a considerable period (e.g. census material and other records of individuals, and data supplied in confidence by business) most of it is not inherently confidential and could be made accessible to bona fide researchers long before the expiry of the 30 years after which government records normally become accessible in the public record office (pro). the problem for the researcher is first to discover what data exists and then to obtain access to iu in the 1940's and early 1950's the inter-departmenial committee on social and economic research (the north committee) put in hand extensive listings of government material of social science interest (including some unpublished material) held by various departments, which resulted in the publication of a valuable series of six "guides to official sources". the social science research council, soon after it came into being in 1965, set up a social science and government committee, which served in a sense as a successor to the north committee. this latter committee speedily recommended that the north committee initiative be repealed and the guides brought up to date. the council therefore awarded a contract to the british library of political and economic science (at the lxdndon school of economics) to produce a new, single guide to all the unpublished government data likely to be of interest to research workers in the social sciences, and potentially accessible to them, that a small team could locale and list in a limited period. the resulting "guide to government data", published in 1974, reported the material revealed by surveys of 17 departments, carried out over a period of three years from september 1969. it is clearly of the greatest importance that these listings should now be updated, and the remaining departments covered. although sheer lack of lime necessitated the exclusion of some departments, the absence on fall/winter 1985 iassist quarterly 35 other grounds of certain key departments is greatly to be regretted; and it is important that their records should now be surveyed. in particular, the department of trade and industry did not find it possible to cooperate in the enterprise; and, though a great deal of material is published by the department, a guide 10 unpublished material there remains a most important desideratum for economists. the foreign and commonwealth office was excluded in view of the confidentiality of much of its unpublished material and the practical difficulty of sorting it; those researching contemporary history will nevertheless find it particularly unfortunate that such materials as press releases could not readily be assembled and made available to serious scholars. the records of the former department of scientific and industrial research contain a vast fund of valuable data relevant to the social sciences in many fields; but preliminary approaches made it clear that a worthwhile survey was impracticable at that time. though problems of personal privacy and commercial confidentiality will probably always prevent access to recent unpublished data in the office of population surveys and h.m. customs and excise, the remaining departments should now be included in a new survey. it seems very possible that in some cases the decision that a department could not cooperate in the earlier project may have been aftecied by traditional department, or even by personal, attitudes towards public access to official documents per se\ or by the fact that its library was particulariy hard pressed at the lime of the enquiry: so that a renewed approach to the department might now be favorably received and a proper survey carries out it is also to be hoped that the official attitude towards public access is becoming more liberal, and that the policy expressed at the time of the former enquiry by the civil service department, namely that it should "make available as much of the material as it reasonably can within the official limits most generously interpreted" is more widel>' accepted, so that it will be possible to fund the compilation of a third guide which would not only update the surveys already published but also reveal important sources of unpublished data held by departments not previously surveyed. for access to material restricted by the 30-year rule the researcher is normally dependent on the libraries of departments. the amount, type and quality of unpublished material actually made available to a researcher in a library varies considerably from department to department, according to the extent to which it is departmental policy to transfer papers from administrative divisions or branches to its central library. in turn the libraries have varying policies on the admission of researchers, and the extent to which readers, once admitted, are given access to papers, as well as on publication some departments publish a substantial proportion of the available material. it may be useful to explain that working papers in a government department are usually held in individual registries attached to the separate divisions or branches which they serve. although files are reviewed by the department after five years in order to arrange for the destruction of ephemeral material, files containing correspondence and minutes are retained by the department and are not accessible to the public as of right imtil they are transferred to the pro after 30 years; and some may be retained by the department after that date on administrative groimds. so if certain information, although not "classified" in the official sense of security classification, is neither readily available in a departmental librar\ open to the public, nor published, nor (as is usual in the case of statistics) provided on payment of a fee to the statistical department concerned, then it generally seems unlikely that under present arrangements members of the public will be able to consult the papers until they have been transfened to the pro. fall/winter 1985 36 iassist quarterly it is clear thai a substantial proportion of the statistical and other data gathered by government departments is being, or could be, recorded and analysed by automated data processing methods. both the actual progress made in this area and the scope for further development will no doubt be discussed in the course of this conference. meanwhile, in 1981 the committee on modem public records (the wilson committee) recommended (para. 423) that "the nucleus of a data archiving centre" should be established as soon as possible; and the government agreed (cmnd 8531 para 55) that a feasibility study should be carried out for the establishment of such a centre at the pro. it is to be hoped that any such centre will work in close cooperation with the social science data bank of the economic & social research council. the advent of data prcicessing techniques should also enable scholars to exploit records more fully and make new connections through the adoption of machine indexing. this should also assist with the hitherto intractable problem of "particular instance papers" (pips). these consist of groups of papers, often very numerous indeed, the subject of which is the same in all cases, though each individual paper relates to a different person, body or place. while each individual paper may in isolation be of little importance, when a set of pips is taken as a whole and analysed it may form a source from which conclusions as to historical, economic or social trends may be drawn. in 1979 the appointment of the wilson committee provided an opportunity to propose changes which would facilitate the use of government archives by researchers in the social sciences. some such opportunities which were taken, and others which were missed, are set out below. in view of the potential importance of data contained in the records of such bodies as the research councils and the nationalized industries. it was recommended that the pro should make proposals for consistent arrangements for the selection, preservation and public inspection of the records of the research councils and nationalized industries. substantial concern had been expressed to the committee about the "weeding" of records before they are transfened to the pro. it was felt that coherent, clearly applicable objectives should be established; and that those carrying out this task needed to have an appreciation of the needs of present and future scholars. it was also suggested that some types of "weeded" material might be transferred for preservation outside government departments (e.g. to universities); and that special note should be taken of the importance for research of the working, papers of royal commissions and select committees and the evidence presented to them. the committee recommended the adoption of a system of "sector panels", including both departmental interests and researchers with experience of the field in question, to advise departments on the official and wider purposes to which departmental records might be put, so that this advice could be taken into account by the "weeders"; but the government has rejected this recommendation. the committee made no reference to the suggestion that some "weeded" material might be preserved outside the relevant department: it was perhaps thought that the organization and exploitation of such deposits might require rather substantial expenditures by the recipient organizations. the committee made no reference to the suggestion that the press releases issued by government departments should be computerized; nor to the concern that has been expressed over the rigid application of the "30-year rule" when material on the same topic is already accessible outside the pro and when material of a routine nature, but of some historical interest, is automatically restricted under the rule.n fall/winter 1985 iassist quarterly 37 new products data user services division data developments bureau of the census u^. department of commerce wathington, d.c. 20233 census ofl^ers international data base on diskette the u.s. bureau of the census is now offering its international data base (idb) on disicette for 203 countries of the world. organized as a series of 94 statistical tables, the idb contains demographic, economic, and social data for all countries of the world. each table is fully annotated with information on sources, methods of collection, definitions of terms, methods used in computations, and qualifications of the data. each diskette contains tables and notes for a single country. a computer program, written in basic, is supplied to users to extract individual tables into formats acceptable to software packages such as lotus 1-2-3, ddbase, supercalc, and spss[sic]. file documentation describing the contents of the tables and notes is also included. the cost for the first country purchased on diskette, including file documentation, is $15. the price for each additional country is $7.50. the statistical tables from the idb are also available in printed copy. when ordering printed copies of statistical tables, the first 10 tables are free and additional tables are $0.25 each. on-line access to the idb is available to users within the federal govemmenl currently, there are 56 organizations within the federal government that have on-line access to the idb through computer terminals or microcomputers with dial-up capability. the data contained in the idb are collected from many sources, including national statistical offices throughout the world, international organizations such as the united nations, research centres and universities around the world, and federal agencies such as the u.s. agency for international developmenu data that originate in host countries usually come from census, vital registration systems, surveys, and administrative records. fall/winter 1985 vol223 fall 1998 9 the recent establishment of the california digital library creates an unprecedented opportunity to bring social science data to users in a more user-friendly format as well as making it available to a much wider audience. the california digital library was established at the university of california in the middle of last fall, in effect creating a tenth campus library, this one entirely virtual. in practice, it is based at the uc office of the president headquarters in oakland, california, it’s main manifestation in its initial year appears to have been acquiring licensing agreements with computerized databases of academic journals in the fields of science and technology. but with an anticipated infusion of $3 million from state coffers this coming year, plus another $1 million from the university of california itself, the cdl, as it is called, will soon be looking for new disciplines to cover, and new territory to conquer. according to press reports, california has had a budget surplus of $4.4 billion which can hopefully be partially allotted to the libraries. social sciences are a likely future field of attention for the cdl, and uc’s data archivists and librarians recently met with the cdl’s newly installed collections officer to discuss matters of mutual interest, including ways of collaborating in possible future joint endeavors. this paper, then, is part of a continuing, fairly new process of rethinking how the uc system collects and provides services to social science material in digitized formats. since the cdl bills itself as the 10th campus library collecting everything in digitized form, there might have been initial fears that collecting of digitized materials would end on the individual campuses that make up the uc system. while legally and constitutionally the university of california is one state-wide system, in actual practice, there is almost no collection development (in terms of library resources) at the system level, except for licenses to databases negotiated system-wide and made available via melvyl, the uc catalog and database gateway; and large purchases made in the past through “shared purchases” system-wide, largely made up of big ticket items and large microform sets. each campus is strongly independent, each headed by its own chancellor and strong faculty senates; campus libraries are no different, each with its own collecting focus, depending in a large part on the research and instructional needs of the faculty on the particular campus. despite being three years in the planning stage, the actual creation of the cdl may have caught many campus librarians by surprise. any initial fears that collecting of digitized materials on campus would be displaced by the cdl collecting the same stuff for all campuses soon gave way to the realization that there was no way individual campus digitized collections would disappear, nor would collection development cease. in fact, since data archives and data collection development vary across campuses, meeting individual needs not necessarily duplicated elsewhere, the likely model that will prevail is one more of collaboration with the cdl than one of the cdl usurping and displacing campus collection patterns. while it is relatively easy (and familiar) for the cdl to negotiate licenses over journals in the social sciences, it is a different kettle of fish for the cdl to enter the field of collecting (and making available) raw data that are the products of social science research. it is not just a matter of licensing, but also one of making it available to users — and which users need to be defined — in an appropriate format. how the cdl is structuring its own collection development may be instructive here. currently, within its science and technology collection focus, the cdl has selected biotechnology and computer science as the two areas where acquisition of digitized information in formats other than electronic journals or monographs would take place. this is described in a march 8, 1998 announcement on its listserv (cdlinfo-l) as providing a “laboratory for learning and planning for future cdl collections,” and for “gradually [adding] other types of content” beyond electronic journals or monographs. this is an ongoing experiment, and it is too early to report any results. but what is clear is that the cdl intends to acquire datasets in biotechnology and computer science initially, although what it would acquire in particular is not apparent. in addition, the cdl plans to structure its collection in the california digital library: implications for social science data files collections by daniel c. tsang* 10 iassist quarterly three “tiers”, namely, tier 1, material funded, in whole or in part, by the cdl; tier 2, material funded by two or more campuses; and tier 3, material that is paid for by only one campus. for the last tier, this implies the cdl would provide a gateway, on its web site apparently, to material on that campus, even if, as is likely, only users affiliated with that campus can use the material. for tier 2, then, material would likely not be accessible to those outside the campuses that funded it; and for tier 1, since the cdl funded the material, it would presumably be made available to most or all campuses (a campus can opt out of a particular acquisition). its current collection principles are that priority should be given to digital format acquisition of those resources which offer economies of scale by benefiting the most faculty and/or students both locally and system wide. also, electronic materials should be selected based on increase of access to the installed base of uc library collections and build on the investments already made by the university in digital resources. if and when the cdl enters the field of social sciences collecting and social science data in particular, a major issue appears to me to be the one of defining who would be the potential user population. on a traditional campus archive or data library, uc data archivists and librarians have generally acquired material only for the use of faculty and students on that campus. in addition, as i argued in an earlier iassist conference paper1, collection development, in uc’s as elsewhere, has generally, been reactive rather than proactive; i.e., datasets are acquired in response to specific requests, rather than, as with books, collected in advance of specific request, and often without anticipation of actual immediate use (e.g., through book approval plans). since that paper, however, with the advent of the main source of social science data now distributing, in effect, its entire newly acquired collection every quarter or so in the from of periodic release cd-roms, libraries are acquiring data without any selectivity, in actual practice, and certainly not in response to any specific request. in addition, with the icpsr’s entire archive much easier to retrieve (with the demise of distribution on round tapes), data archivists can potentially replicate much of the collection on their own campuses, although as of yet, there is no mirror site to the icpsr archive yet established. such a mirror site might be one area the cdl might fruitfully consider, especially since its mission appears not only to serve the campus-based academic community but also the community at large, although it is still unclear how that would work in practice. the cdl, after all, is called the california digital library, not the university of california digital library, strongly implying it is collecting digital material to serve the entire state, i.e., all the people of the state of california. that was the quid pro quo that apparently was necessary for the state legislature to fund the cdl. in announcing the creation of the cdl last fall, uc president richard c. atkinson spoke of creating “uc’s library without walls.” it would be a library allowing “scholars of all ages and interests to range worldwide in their quest for knowledge, using the internet, the world wide web and a computer.” if the target population is the scholarly community within uc and beyond, that would be one thing. libraries are used to collecting for scholars; the cdl could very well mirror icpsr’s archive (after becoming a fullfledged member like the other uc’s that are members) by paying them enough money so that they would not think that they are losing money. but since there is no physical campus associated with the cdl, what does this mean? or is the cdl membership to take the place of the individual campus memberships? that does not seem likely, given that existing campus archives and data collections have constituencies they have nurtured and served for years; they are not likely to disappear, at least not without a fight. but if the cdl promised access to all icpsr data, and provided front-end interfaces that facilitated the extraction of variable-level data (it would need to write the software etc.) of selected datasets, would that not make local service points less necessary, or even impractical? the countervailing argument, of course, is that with all secondary analysis of data, it is not sufficient merely to make the data available; there has to be a variety of services (metadata access; interpretation of metadata; statistical consulting; etc.) that, because computer setups vary across campuses, and even within, are best handled at the local level. in addition, the cdl can perhaps better negotiate licenses at the system-level with data vendors that provide data, such as economic time series, of interest across the uc system. but given its mandate to serve the state, the cdl and its pioneering vision, the cdl could very well do more than just provide access, even improved access, to what currently exists. the cdl could well get involved in one or more large-scale digitization projects, funded by industry and government, to archive and make accessible datasets previously not readily available, such as government data at both the state, county and municipal levels. the challenge here is for joint partnerships with communities local and state; what is sorely needed is a collaborative effort to make sure that government data does not go the way of the main frame and that they are archived and preserved, and eventually made accessible to users. as the cdl develops, one potentially controversial area involves intellectual property rights. who owns the research that faculty have invested their time in? the university is now arguing that it does, and that faculty who sign away rights to articles, for example, are just making commercial journal publishers much more rich. administrators are now wondering why a university should fall 1998 11 have to pay exorbitant fees to access the journal output of their own faculty, just because a commercial journal published it. richard lucier, the cdl’s head, argues that scholarly publishing must change, and that universities must on their own, compete with the industry and ”publish” on line the scholarly output. with data however, is it likely that individual faculty, even at uc, will be willing to deposit their research data at the cdl, if the university insists that it must? the university recently revised its policy on research; now, even if faculty use a dataset gathered elsewhere for secondary analysis, it must register with local institutional review boards. local boards could well insist that faculty deposit data thus gathered (or originally collected data for that matter) with the cdl, or the campus data archive. right now there is no such provision or mandate, not least because researchers are unlikely to be willing to part with their data, and there is no common understanding of who owns what. if a faculty member leaves, he or she is allowed to take research data he or she has gathered; thus far, i am not aware that the university has insisted on ownership. but a case could well be made, given that every employee is, upon hire, made to sign away most of his or her patent rights to the university. involvement in large-scale projects is especially likely, and necessary, if indeed, the user community stretches beyond the scholarly community. if indeed the vision is to let anyone access the information (and since the cdl is called the california digital library, not the university of california digital library, as originally envisioned), that suggests that anyone, “of any age”, would have access to the library. that would well be a mammoth task, devising a dataset (or more) that would be useful to such a mythical user. there are many more issues one could raise, not least that of digital archive maintenance, authentication, and dataset updating, as well as version control. at a minimum, the cdl could provide a union list of what exists in existing campus archives and collections, providing some bibliographic control to an existing situation that is anarchistic at best. but i see it as doing more than that; it can best make the use of data more appealing, by providing the necessary tools to access data from the hard-to-find to those most popularly requested. as such the cdl can help those of us in the data archive community by making, and educating, more people to be sophisticated data users and consumers. in conclusion, the cdl is unlikely to be the sole digital repository for california of social science data, but its creation and expansion will likely spur collaborative efforts with existing collections and archives as well as create new ways to provide improved access to these collections. 1. daniel c. tsang, “academic libraries and collection development of nonbibliographic machine readable data files,” iassist quarterly, volume 12, number 3, fall 1988, pp. 47-55. * paper presented at the iassist conference, may 21, 1998, at yale university, new haven, connecticut. http://www.yorku.ca/org/iassist/ confidentiality and access: legislative initiatives in the united states government. patricia aronsson ' director, documentation standards division national archives and records service introduction information specialists recognize the implications of our increased reliance upon automation. the conflicting demands for privacy protection ' the opinions expressed in this article are those of the author alone and do not reflect the official position of the national archives and records service, nor of any other federal agency. and access clash repeatedly. but, this conflict will not be readily resolved. recent initiatives in the united states government reflect the competing demands of enhancing privacy protection and broadening access to information. at the same time that the united states senate has proposed restricting access under the freedom of information act, we are seeing more computer matching programs than ever before. while the government implemented a strong privacy act, we failed to recognize the international implications of that effort. privacy legislation, freedom of information act revisions, and hearings about transborder data flow and computer matching represent the key areas where privacy and access considerations are evident. congressional concern about personal privacy has existed for several years; in 1974, the united states congress passed the privacy act. the house of representatives, through hearings in 1983, demonstrated its concern about the oversight of the act. and during this past congress (the 98th, 1983-1984), representative glenn english proposed the creation of a privacy protection commission to strengthen the implementation of the act. the freedom of information act continues to generate interest and activity. in 1974, congress substantially revised the act. during the 98th congress, the senate passed the freedom of information reform act which significantly modifies several provisions of the 1974 act. the house of representatives held -11hearings on the proposed reforms but took no final action. in recent years, the united states congress has demonstrated its concern about the impact of automation on access and confidentiality through several hearings it has held. there have been a few different congressional hearings concerning transborder data flow. and, in 1982, the senate held hearings on the subject of computer matching. without question, states congress i in privacy relate it has yet to foe to and confidenti automated informa highlighting and congressional act areas of privacy confidentiality, author's hope tha strengths and wea recent united sta initiatives will apparent . the united s interested d issues, but us on access ality of tion. by explaining ivity in the and it is the t the knesses of tes become privacy act the privacy act of 1974 (5 usc 552a) represented the first attempt by the united states congress to legislate general government-wide standards for the protection of individual privacy = but, the passage of this act was not motivated by c concern about the implications of automation for privacy. it is more likely that congress passed the privacy act in reaction to the watergate episode which had heightened national concern about executive secrecy. congress recognized the need for access to personal information about individuals, and specified exceptions to the privacy act requirement that an individual must authorize access to personally identifiable records. these several exceptions indicate that congress was aware of the need to provide access to information. it is less clear that they foresaw the need to take steps to ensure full privacy protection for individuals. the legislation did not designate a central authority for the adminstration and oversight of privacy act implementation. this has proven to be a significant shortcoming and has been the focus of recent congressional inquiries. in 1982, governme house of hearings entitled privacy? privacy office o and by t hearings the subc from sev about th implicat privacy a subcommi nt informat representa and issued "who care oversight act of 1974 f managemen he congress focused on ommittee al eral people e internati ions of the act. ttee on ion in the tives held a report s about of the by the t and budget ". while the oversight , so heard concerned onal u.s. as an outgrowth of these hearings. representative glenn english, the chairman of the subcommittee, and an advocate of privacy protection, proposed the creation of a privacy protection commission which would be a permanent and independent commission responsible for both domestic and international privacy issues. representative english described the proposed commission as followss domestically, the commission would be assigned an oversight role under the privacy act of 1974. the commission would develop guidelines and model regulations, investigate compliance with the act, and generally oversee agency private act activities. for international privacy issues, the commission would assist u.s. companies doing business abroad to comply with foreign data protection laws, assist in the coordination of u.s. privacy policies with those of foreign nations, accept complaints and otherwise consult with foreign data protection agencies. ^ while this legislation certainly would centralize privacy act oversight, it does not seem likely that any action on its creation will be taken in the near future. in the meantime, the courts offer the only recourse for private citizens who feel their privacy has been violated by the federal government. as is probably apparent, the federal government offers only limited privacy protection. a few categories of non-governmental records are the subject of federal privacy laws; personally identifiable records gathered by credit bureaus and by colleges and universities are protected. state and local government records are, for the most part, excluded from coverage. it is important to note, though, that in the united states, each state develops its own privacy legislation for state and local records. transborder flow ^congressional record, august 2, 1983, h6344 daily edition. the proposal for the privacy protection commission is the most recent legislative initiative which addresses the issue of transborder data flow. as is apparent from representative english's statement, though, his legislation grows out of a concern for privacy. many people in the united states contend that the major issue surrounding transborder data flow is one of economics, not privacy. both the house of representatives and the senate addressed economic concerns during the past congress. each house received legislation to create an entity to oversee international telecommunications and information; the senate bill proposed a white house office and the house bill proposed an interagency committee. the recent initiatives in the area of transborder data flow stem, in large measure, from a recognition that the united states government does not have an agency which focuses on the international exchange of information. several reasons make it unlikely that, in the near future, the united states congress will pass legislation -i3to increase the government's role in transborder data flow. in the united states, there is a clear separation between the public and private sectors and, in addition, the first amendment to our constitution explicitly limits the government's intrusion into the flow of information. additionally, the united states government views international regulation of data as restrictive. a senate report written in 1980 states: while other believe tha regulatory may be nece ensure equi opportuniti protect all the united been reluct advocate cr legal struc may prove r in light of technologic development changing ma dynamics. ^ nations t a new framework ssary to table es and to pa r t i e s , states has ant to eation of tures which estrictive rapid al s and rket until united states businesses are significantly hampered by transborder data flow. congress probably will not act. computer matching computer matching programs present significant threats to personal privacy and yet they have proliferated in recent ^"international telecommunications and information policy: selected issues for the i980's", senate report, 98-94, p. 24 years, even though the privacy act seems to prohibit them. computer matching is defined as the use of a computer to compare data in a privacy act system of records with other data for purposes of identifying individuals whose records appear in more than one set of records. computer matching programs are intended to detect and curtail fraud or abuse in federal assistance, loan, or benefit programs. the inspector general act of 1978 authorized the inspectors general to request and obtain information from other federal agencies and state and local governments * . as a result, in 1979, the office of management and budget issued guidelines for agencies acquiring computerized data files for use in computer matching programs. these guidelines, which were revised in 1982, require that ...the source agency should require the matching entity to agree in writing to certain conditions governing the use of the matching file.^ some of these conditions may state that the matching file will remain the property of the source agency, that the file will be destroyed or returned at the end of the project, that the file will be used only for the purposes stated in the agreement, and the file will " p.l. 97-252, sections 6 (a) (3) and (b) (l). ^david a. stockman, memorandum for the executive departments and establishments. may 11, 1982; published in the federal register, may 19, 1982. 1 4not be duplicated either within or outside the receiving agency. in addition, matching agencies are required to publish in the federal register a notice describing the matching programs and to provide a copy of this notice to the congress and to the office of management and budget. these procedures are designed to alleviate the potential abuse of personal privacy that is magnified by the increasing use of computers for the collection and maintenance of personal information. computer matching programs are becoming more widespread in government. in 1982, the president's council on integrity and efficiency initiated the long term computer matching project to encourage and facilitate computer matching. in july 1982, the project, which involves all inspectors general, published its first newsletter and reported that the inspectors general of ten departments and agencies were engaged in computer matching efforts of sufficient significance to merit reporting. it is likely that the newsletter did not mention all the computer matching programs underway, and it is just as likely that new matching programs have been started since then. while computer matching remains a controversial practice, it seems to be gaining wider acceptance. interestingly, while computer matching has become more acceptable, statistical research has been severely curtailed since the passage in 1976 of the tax reform act which places strict restrictions on the use of individual tax returns. freedom of information act the freedom of information act (foia) ^ affects privacy and access issues in a number of ways; some of them only tangentially. recent congressional initiatives to reform the foia reflect a growing concern about access to government information. one revision prohibits foreign nationals from acquiring information under the freedom of information act. another proposal suggests charging a fair value fee for commercially valuable information. yet another modification tightens access to informant information.' as mentioned earlier, the senate passed these reforms but since the house of representatives did not act, the bill expired and will have to be re-introduced during the next congress. one change to the freedom of information act did, however, become law this year; central intelligence agency files on intelligence sources and methods have been exempted from the provisions of the foia. the impact of various definitions on privacy and access ^ 5 use 552 's. 774, 98th congress. -15certainly a major problem we face in the united states government is determining which materials are covered by which legislation. the multiple definitions of "record" complicate the protection of privacy in the u.s. government. both the federal records act and the privacy act contain different definitions of what constitutes record material. while the freedom of. information act does not contain an explicit definition of records, it uses phrases defined in each of the other two laws. these multiple definitions of record present problems to those people concerned with privacy protection and access. it is difficult to properly protect information or to monitor compliance with privacy protection requirements if it is not clear what information is covered. than ever. the united states congress is beginning to recognize the need to address the problems surrounding the confidentiality of automated information. and it is likely that, in the coming years, we will see a great deal of congressional activity in the area of privacy and access. but, we can also expect that such action will come only in reaction to particular problems. we are likely to see a series of discrete pieces of legislation, each designed to address a specific problem. it would be unduly optimistic to expect the united states congress to develop a comprehensive and coherent policy concerning the increased use of automation for the collection of information. an additiona that these d only to the each state i establishing of what cons therefore, t privacy prot need to know state in whi located. 1 complication is efinitions apply federal government, s responsible for its own definition titutes a record, o ensure full ection, one would the laws in every ch records may be conclusion federal law plays a significant, albeit limited, role in the protection of privacy in the united states. our increased reliance on automated information makes such protection more important -16iassist quarterly 2013 71 iassist quarterly abstract the special interest group on data citation (sigdc) carries on the work of iassist established in the early years of the organization by sue dodd by advocating for the standardized citation of datasets. a major accomplishment of the sigdc is the development of the quick guide to data citation. this educational document simplifies the proliferation of data citation guidelines by presenting the five common elements of citation and offering examples in apa, mla, and chicago style formats. it can be used as a tool in education and advocacy to foster growth in data sharing behavior and research data management practice. keywords data citation, bibliographic references for data, data sharing, history of iassist introduction although many students find mastering citation formatting to be a tedious venture, as information professionals we recognize the significance of the citation. librarians appreciate the predictable structure of bibliographic references for the ability to easily look up known items. research faculty rely on citation counts for merit review. scholars seek out the complex connections that make up the scholarly conversation. data sharing advocates bemoan the failure to directly integrate datasets into the conversation and look to the citation as savior. if only we could get researchers to treat the contribution of re-usable data itself as equally valuable to the description and analysis of that data, then the cause of research data management and the curation of data would be furthered. iassist is an organization founded on the value of curation, preservation, and access to data. so it comes as no surprise that advancing the cause of data citation has been embraced by the membership since the beginning. as margaret o’neill adams reports, the classification action group was a strong component of the early iassist organization and sue dodd’s working manual for cataloging machine-readable data files and later book, cataloging machine-readable data files: an interpretive manual, were a crowning achievement for the organization (adams, 2006). a clear intended a practical approach to data citation: the special interest group on data citation and development of the quick guide to data citation by hailey mooney 1 sigdc authored the quick guide to data citation 72 iassist quarterly 2013 iassist quarterly outcome and extension of cataloging and classification for data files, providing for identification and access via bibliographic references was an essential component of the classification action group’s activities. dodd published a concise set of guidelines for the citation of data in the journal of the american society for information science in her capacity as chairperson of the iassist u.s. classification action group (dodd, 1979). in specifying the need for data citation, dodd explains that data files are cited irregularly and that in many cases, “the information provided is not sufficient to indicate the proper source of the data and consequently an interested party has to spend a considerable amount of time determining additional information on the availability of a particular data file that could easily be provided by a full and proper bibliographic reference” (dodd, 1979, p. 78). over the years, the topic of data citation has continued to receive intermittent attention as part of the conversation within iassist, with a focus on the need to further the practice through the development of norms and guidelines (e.g., altman & king, 2006; drolet, 2005; hankinson, 1988). unfortunately, despite the championing of dodd and others the status quo of lackluster data citation behavior has yet to significantly change (mooney & newton, 2012). the recent attention within the profession on data management planning, including priming researchers to participate more broadly in data sharing, has placed renewed emphasis on the importance of citing data as first-class scholarly objects within the literature. especially with the emergence of datacite () and the participation of iassisters in the organization, iassist experienced a rekindling of attention to data citation. this coalesced into the formation of the special interest group on data citation (sigdc). sigdc & the quick guide to data citation the sigdc was started in late 2010 with a goal to promote awareness of data-related research and scholarship through data citation. to that end, sigdc has sponsored several iassist conference sessions and posters. the group has sent advocacy letters to the major style guides imploring them to include instructions on best practices for the citation of data. most notably, sigdc authored the quick guide to data citation (international association for social science information services and technology, special interest group on data citation, 2012), an instructional pamphlet that distills the main elements of a data citation from the key existing and proposed standards and provides an example citation in the three major style formats. the core team behind the development of the quick guide to data citation included robert downs, michele hayslett, hailey mooney, and michael witt. elizabeth moss, mary vardigan, and other sigdc members also contributed to its final form the impetus for the quick guide came from the reality that many style guides do not provide adequate guidance for the citation of data (newton, mooney, & witt, 2010). much of the literature we reviewed in the process of creating the quick guide is captured on sigdc data citation resources bibliography (). the key examples which informed the guidelines were drawn from examples and standards from the scholarly literature (altman & king, 2007; dodd, 1979; mooney & newton, 2012), data producing organizations (e.g., gesis, 2013; green, 2009; international polar year data and information service, 2008; socioeconomic data and applications center, n.d.; statistics canada, 2009; uk data archive, n.d.), and other stakeholder groups (ball & duke, 2011; datacite metadata working group, 2013). so, although style guides and normative practices do not universally embrace data citation, a relative proliferation of various data citation suggestions exists across disciplinary areas. the goal of the quick guide then was to distill these guidelines to the essential elements and provide practical examples in common referencing styles. distribution of the quick guide in printed brochure format at the annual conference provided iasssters with a means to easily provide data citation advocacy materials in orientation packets or other venues; the online version can also easily be printed out. librarians are encouraged to use the quick guide as a basis for libguides that promote data citation best practices to their communities. it is a practice-orientated tool that simplifies what can be a complex endeavor for ease of use by researchers and scholars. the final set of the minimal baseline data citation elements recommended by the quick guide includes: • author: name(s) of each individual or organizational entity responsible for the creation of the dataset. • date of publication: year the dataset was published or disseminated. • title: complete title of the dataset, including the edition or version number, if applicable. • publisher and/or distributor: organizational entity that makes the dataset available by archiving, producing, publishing, and/or distributing the dataset. • electronic location or identifier: web address or unique, persistent, global identifier used to locate the dataset (such as a doi). append the date retrieved if the title and locator are not specific to the exact instance of the data you used. areas of debate during the development of these citation elements included the issue of whether edition/version and date retrieved should be elevated to a separate element. this information, especially the date retrieved, could be critical given the need to be precise about the exact dataset used in an analysis and the potential for some datasets to have multiple updates over time that may or may not be reflected in the assignment of a new edition or version number. ultimately, the desire for simplicity won out in keeping to a set of just five citation elements. edition and version are discussed within the title element and data retrieved is part of the definition for electronic location or identifier. these were seen to be closely tied with the other elements in aiding for precision of recall. it should be noted that dodd’s conception of the “imprint” (dodd, 1979), or producer and distributor statement, was a major influence in the conception of this citation component. this accounts for the somewhat unique characteristics of research datasets in that they may originally be produced and published at one institution, but trusted to another institution (the data archive) for curation and dissemination. while it is important to acknowledge the original producer, for purposes of retrieval the dataset distributor is a crucial piece of information. in crafting the example citations within the quick guide, the general social survey was chosen as an emblematic social science dataset. examples were provided using the general guidelines for iassist quarterly 2013 73 iassist quarterly formatting from the “big three” style guides in regular academic use: apa, mla, and chicago. • apa (6th edition) smith, t.w., marsden, p.v., & hout, m. (2011). general social survey, 1972-2010 cumulative file (icpsr31521-v1) [data file and codebook]. chicago, il: national opinion research center [producer]. ann arbor, mi: inter-university consortium for political and social research [distributor]. doi: 10.3886/ icpsr31521.v1 • mla (7th edition) smith, tom w., peter v. marsden, and michael hout. general social survey, 1972-2010 cumulative file. icpsr31521-v1. chicago, il: national opinion research center [producer]. ann arbor, mi: inter-university consortium for political and social research [distributor], 2011. web. 23 jan 2012. doi:10.3886/ icpsr31521.v1 • chicago (16th edition) (author-date) smith, tom w., peter v. marsden, and michael hout. 2011. general social survey, 1972-2010 cumulative file. icpsr31521-v1. chicago, il: national opinion research center. distributed by ann arbor, mi: inter-university consortium for political and social research. doi:10.3886/icpsr31521.v1 since the time the quick guide was published in 2012, the data citation landscape has continued to evolve as the conversation endures around required citation elements, best practices, and author behavior (e.g., codata-icsti task group on data citation standards and practices, 2013; force11, data citation synthesis group, 2013). educating scholars on the benefits of data citation is crucial in the move to support funder-mandated data sharing policies. developing and advocating for the establishment of best practices within the profession is a core part of the iassist mission. as a practical synthesis of many voices, the quick guide continues to stand as a valuable resource to information professionals, scholars, and researchers looking for a simple and direct way to cite data. references adams, m. o. (2006). the origins and early years of iassist. iassist quarterly, (fall), 5–14. retrieved from altman, m., & king, g. (2006). overview of a proposed standard for the scholarly citation of quantitative data. iassist quarterly, (summer), 18-19.. retrieved from altman, m., & king, g. (2007). a proposed standard for the scholarly citation of quantitative data. d-lib magazine, 13(3/4). doi:10.1045/ march2007-altman ball, a., & duke, m. (2011). how to cite datasets and link to publications (dcc how-to guides). edinburgh: digital curation centre. retrieved from codata-icsti task group on data citation standards and practices. (2013). out of cite, out of mind: the current state of practice, policy, and technology for the citation of data.. data science journal, 12, cidcr1–cidcr75. doi:10.2481/dsj.osom13-043 datacite metadata working group. (2013, june). datacite metadata schema for the publication and citation of research data. retrieved from dodd, s. a. (1979). bibliographic references for numeric social science data files: suggested guidelines. journal of the american society for information science, 30(2), 77–82. doi:10.1002/asi.4630300203 drolet, g. (2005). citing statistics and data: where are we today? presented at the iassist annual conference, edinburgh, scotland. retrieved from force11, data citation synthesis group. (2013). declaration of data citation principles draft. retrieved november 26, 2013, from gesis. (2013, october 10). bibliographic citation of research data and study related documents. retrieved december 18, 2013, from green, t. (2009). we need publishing standards for datasets and data tables (oecd publishing white papers). oecd publishing. retrieved from hankinson, r. (1988). issues concerning the bibliographic citation of machine-readable data files. iassist quarterly, 1988 (spring), 13-18. retrieved from . international association for social science information services and technology, special interest group on data citation. (2012). quick guide to data citation. retrieved december 18, 2013, from international polar year data and information service. (2008, june 19). how to cite a data set. international polar year data and information service. retrieved december 18, 2013, from mooney, h., & newton, m. (2012). the anatomy of a data citation: discovery, reuse, and credit. journal of librarianship and scholarly communication, 1(1). doi: newton, m. p., mooney, h., & witt, m. (2010). a description of data citation instructions in style guides. retrieved from socioeconomic data and applications center. (n.d.). citing our data. retrieved december 18, 2013, from statistics canada. (2009, may). how to cite statistics canada products. statistics canada. retrieved december 18, 2013, from uk data archive. (n.d.). terms & conditions: citing data. uk data archive. retrieved december 18, 2013, from appendix quick guide to data citation ( pages 74-77) notes 1 hailey mooney is the data services coordinator and social sciences librarian at michigan state university libraries. she can be reached by email: mooneyh@msu.edu. 74 iassist quarterly 2013 iassist quarterly appendix quick guide to data citation iassist quarterly 2013 75 iassist quarterly 76 iassist quarterly 2013 iassist quarterly iassist quarterly 2013 77 iassist quarterly modified for the iassist quarterly. original is found at : http://www.iassistdata.org/sites/default/files/quick_guide_to_data_citation_high-res_printer-ready.pdf microsoft word 5 tripepibookreview.docx 1/2 tripepi, chubing (2017) book review: databrarianship: the academic data librarian in theory and practice, iassist quarterly 41 (1-4), pp. 1-2. doi: https://doi.org/10.29173/iq13 book review: databrarianship: the academic data librarian in theory and practice lynda kellam and kristi thompson, eds. (2016) chicago: association of college and research libraries. 378pp. $68. isbn 978-0838987995 chubing tripepi lynda kellam and kristi thompson’s book, databrarianship: the academic data librarian in theory and practice, presents cutting-edge topics and pragmatic common practices in various aspects of the relatively new data librarian field. by inviting collaborations from data librarians across the world, the editors successfully deliver a masterpiece to showcase the front-end of data services in academic libraries. in the first part of book, data support services for researchers and learners, various case studies related to data service support are discussed. a diversity of data services models is offered for readers to adapt to their own situations, such as those presented in chapters 1, 2, 4, and 5. chapter 3 sheds light on how to get started with microdata services as a subject librarian, and chapter 6 offers handson lesson plans on teaching data literacy to undergraduate students who are new to quantitative methods and data analysis. students will learn how to discover, evaluate, and manipulate data through the lessons. in the second part of the book, data in the disciplines, readers get first-hand information about supporting geospatial data services in academic library settings. this is important, as gis data “has its own peculiarities, tools, and approaches”, and it is definitely gaining popularity in the academic research environment. to engage the reader further, evolving technologies related to gis data support, such as crowdsourcing and cloud computing, are elaborated upon in chapter 10. in chapter 11, the author explains different stages of supporting qualitative research and data, and presents exploratory findings both from an analysis of iassist listserv job postings and from a survey of qualitative data support practices, which used carefully designed questions to inquire about key aspects of qualitative data support in academic libraries. lastly, chapter 12 introduces scientific data and its relevance to librarians and libraries. this material can be read together with that in chapter 21 if one is interested in supporting science-focused data services. in the third part of the book, data preservation and access, key issues, such as data sharing policies, selection and appraisal of datasets, metadata, and scholarly communication are examined. a special 2/2 tripepi, chubing (2017) book review: databrarianship: the academic data librarian in theory and practice, iassist quarterly 41 (1-4), pp. 1-2. doi: https://doi.org/10.29173/iq13 case of collaboration with the city of calgary, canada is shared in chapter 16 for those who are interested in acquiring local data for their collection. while these aspects of data support are behind the scenes, they are just as important as the public-facing support that librarians provide to the research community. these technical components are the building blocks of the growing data collection and data services fields. in the last part of the book, data: past, present, and future, readers can learn about the data services models used in the uk and canada, including their opportunities and challenges. chapter 21 explains interesting results from a survey of science data librarians. it is fascinating and inspiring to learn from their experience and advice. observation from the social science perspective was a nice touch. in chapter 22, authors discuss their valuable experience in teaching data librarianship to lis students, which was rather forward-thinking in 2011. the course plans and outcomes are shared, along with lessons learned. one example is to leave enough time to teach students necessary tools in their case, spss. personally, i feel their course plan was a little too ambitious; my suggestion would be to pair this course with another standalone class teaching data analysis tools, such as (these days) r or python. another class for data visualization or gis would also be useful. it can be a struggle to gain proficiency in these concepts within one or two weeks. i have two main suggestions to editors of this book. first, some chapters seem to fit better in a different part than they originally appeared, such as chapter 8 for part iii, or chapter 21 for part ii. second, since data services is a rapidly changing field, it would be useful to include the online presence (twitter, blogs, etc.) of the authors in their biographies, so that people who are interested in their chapter can follow up on their later works. overall, i believe that databrarianship is the go-to book for any new librarians or librarians offering or interested in data support services. the data librarianship era is coming, and we are all part of it. this book provides priceless guidance to librarians and leads them in the right direction to start exploring with a community of awesome data cohorts! chubing tripepi cht2114@columbia.edu business research & data services librarian thomas j. watson library of business and economics columbia university ia3sist newsletter, vcj.. 2, lie. 2 (sprin.j 197d) the rcper centeh: cevelcprieni of k new organizational model for daia acci/ss by william j. gairinell the university or concecticut what is the roper cinteh: esta a jo, t search ter , in largest data in and su over 9, ried ou have be between ciai themsel source ices of dupiica or spec ti.an uo researc of the associa members tlished he roper center (n c.) is to archive the worl pporting d do invidi t in more en deposit two and scientists v€s of th ty request the cente tion, inf ific data colleges, h organiza internatio tion (isla hip ara of thir publ ow t aay of d. ocu m vual than ed thre an is ing r, s orma aaa uni tion nal ) , t the ty-t ic he the samp the eiita stu 7c in t e tn nual rich var uch ticn lysi vers £ a £urv he c hop wo pini t?ope olde j.e ra tion dies coun he c ousa dat lous as d ret s. itie re m ey l oope er c years en rer censt and survey v. data rrom , cartries , tnter . nd 30d vail a reservataset rieval more s, and embers ibrary rative enter. the roper center in transition in 1975, the trustees and staff of the boper center began to consider seriously a major revision of the center's organizational structure and institutional base. the first action following from this was to establish tne center in july 1975 as a non-profit corporation rormally governed by a board of trustees. at about the same time, discussions were initiated between center personnel and the officials of a nuirber of american universities about the possibility of a new hosting arrangement. the university or connecticut and yale university entered into such discussions with the roper center in the fall of 1975, prompte tion 01 foremos the iud limits do to the cen tial. accompl its thr the wi and des sociati tees ca priate the res vane d the new t amo qment to wh devel ter • s mind ishme ee de lliam irous on wi me t base earch ty o rope host n g th that at a op an to ful o nts o cades s col or tn hi o fee for t univ i co r boar ing po em, h ther small arcn its f u f the f the / app lege m a 1 n t a iliacs 1 that he rep ersity nsider d' s ex ssibii owever e are colle ive su llest signi cente reciat contri iling , tiie tne er cen ations f loraities. , was severe ge can ch as potenricant r over ive of jtution an astr usapproter is for the connect w the p rangemen uld host cialiy a bot:. ew that culd be iaborat sser deg scholar shion . clined t r m a t i o n survey ovide. e rooer ant " res ntly evo at couii pact, a uld like ir pa icut rospe t t tne ttrac insti soci based ory ree o s wo they o tap and rese thas c e nt e ource iving pro nd h to unive ivers disti scie reput „ st t ive t the ce re am pus thei . and the ersit t ste rt, an. ct hro pop tiv tut al to str n t rki wa ne new arc so via tne yale of a c ugh wh er cen e. ma ions s science a gr ea uctur es he priv n g in nt ea a w sour techno n and acult y n airea n this cial sc e much e one assocsee rsities. ity had nguished nee 3 wi ation fo ill, its leaders advanc search require r social tnat fo roper y would p i n t h uni ver uni ver oopera ter as ny fac hared rese ter ex and ate au a sing model ces of logy compu consia dy sig more ience , value which iated sit/ si ty tl vi taey esulty tne arch tent to i sin j uiar more insucii ters ered nifr eone and thev witr. yale un aether a the social national excellence administra vinced tha cial scien on their c ration of structures ation with host univ signif ican tion. for its part, the u connecticut in the started making a majo this sector of soci through the development data facility. the so data center (ssdc), es 1968, had assumed a c in both social science research at the instit 1976^ it had acguired stafr of twenty profe and women comprising skills and experience to a partnersaip with center. the ssdc would "in-place" structure offer immediate aid in ment of the roper f which would itself ben from a hosting arrange archive of tne scope tional associations or center . bro fac th a r s fac we emen and d th la rmax cent cons is ught toulty in n intercholar lyulty and re cont of soteaching e elaboboratory associer as a titute a elaboraniversity or late 19fa0's r effort in al science of a social cial science tabiished in entral place teaching and ution. by a competent ssional men many of the appropriate the roper provide, an which coula the developaciiity and ef it great ly ment with an and internathe roper yale and connecticut approached the possibility of a partnership 33 wit., the hoper ccntt;r, tner,, as a pitted (see "news and notes" m natural txtendion or jxai.s ai.o co.iitms issue of the newsletter., liiitaents wl.ici. haa taker. shape over a perioi; or time. ii. additioi. to this, faculty at the two scnools shared strongly the judgments of advantages of thz rcper tne roper trustees taat the recenter search university setting was apt^iut^ln 12 atiqn and propriate to zurtner center develel^nsk^il?! opuent, and that the two schools were well placed to provide the the roper center will be able to kind or assistance that the center function more effectively in its required. t.iey would de able to new setting than it has been able mare a contribution to social scito in the old one. we identify the ence nationally and inter nat icnally following gains as the most importhroaga provision of expanded factant. uity expertise ana tecnnicai facilities. it appeared, then, that a hosting arrangement with the roper center would mean a happy marriage access to impro ved techn ical of legitimate institutional inter"facilit ies ests and the reguirements of an important eociaj. science resource. the poper center requires sophisticated computational taciliey late fall, 1975. connecticut ties. as a result of the move, it and yale had agreed tnat their inhas full access to the yale and vitation to the roger center would university of connecticut computer be a joint one. it was felt that centers. to all of the software in such a cooperative venture would place at those facilities, and to represent a sensible utilization of the technical staffs associated the resources of two neighboring with them. (both centers, it might schools. yale and connecticut facbe noted, have fully compatible, ulty consiaered their institutions same-generation ib;] computers.) in well matchea in resources, in ataddition, the technical services tainments. in interests, and constaff of tne social science data cludea tnat their closeness geocenter, is available to assist the grapiiically woula make for an easy roper center ii. a number of ways, collaboration. the further develfor one thing, it will be possible opmeiit of the roper center would to achieve "economies of scale" require all that both institutions wnich are so important. it has acwould be able to contribute and cess to staff help enabling it to would be advanced especially by the respond to the directions in which fact that the contributions or the tne user community is likely to two schools would be, to such a proceed in the future. striking extent, complementary rather than overlapping ana redundant. ill^ ds^ aiililli§l£3tive arrangements in late 1976, the administration of williams college expressed its we are confident that the new support for the new hosting araaministrati ve arrangement will rangement and its desire to re aswork effectively, and that it repsociatea with it, ana all cf tne resents the best possible arrangenecessary elements were in piace. ment for attaining a variety of meeting on february 2, 1977, with quite disparate objectives. first, representatives of yale, connectiit was considered essential that cut and williams, the roper center the place of the roper center as a hoard of trustees forirally approved repository for commercial survey the move and reorganization. the data -foreign and domestic -be trustees of tne tnree institutions maintained. tne various commerical separately endorsed the new -jartsurvey firms have looked upon roper nersi.ip. " as "their own," and this encouraged them to contribute to the center. the rccer center it now an indeon the roper board of trustees have penaent corporation in formal partsat burns w. roper of the roper ornership witn the university of conganizaticn, george gallup or the necticut, yale university, ana american institute of public opinwiiliams college. divisions cf reion, wiiliam j. wilson of sponsibility have been agree.^ uponst arch/inr a/hooper , and other imtae develcom^nt or a new staff portant leaders of the survey structure "is proceeding. the world, transiticn or the archivax oevolopment section to storrs lo we^x unsecond, we wanted to add to this derway, ana t.ie transfer or user historic element of roper center sfcivices to t\ew haven i.as been comorganization the skills ana facili39 lassist mewslatter, vci. 2, nc. 2 (spring 1978) ties of the research university. the continued i.'iportance the absence of this ccmpouent has of tritl^opl!? c'et^te'r been a aeciaej debit ic past roper td the soctil erforts to service the social scil^iie^^? ence coniirunity. yalt, as tne setting for user services, is easij.y impleraent ation or the new or^anaccessicle to scholars from around izational a rrangemer.ts was prilithe world. cated upon the assumption tnat tne, would help an aiready valuable soft ccirpiete copy cr the entire cial science instrumentality becom: hoiuings of the center will be even more important, more successmaintaiced at the university of ful in meeting research and teachconnecticut, with another conpiete ing objectives of faculty in th ; copy at yale university. taere is social sciences. the move to ne no unnecessary duplication here, of facilities at connecticut and yali course, because one copy serves as becomes desirable, of course, onl' a backup for the ether. williams oecause the roper center is a valucollege will continue tc house varable resource, ious portions of the archive it chooses to maintain. the user tne following is a partial lisi services division will be located of some of the archival resource: in new haven. requests for the duavailable at the roper center: plication of data sets, for searches of the archive, ana the 1. it maintains the raw data like, will be the responsibility of {lu reformattea, numeric this staffthe arcuival developfiles) from polls conment ccmponent will be housed in ducted regularly by the storrs. it will bear responsiblity american institute of for bringing new data sets into tne public opinion /the gaiarchive, for interactions with data iup polij . this series suppliers. for reformatting data of nearxv 1,000 studies sets and bringing the data into one dating from 1936 is of various levels cr accesiblity clearly one of the most for most effective utilization. impressive collections of this division of respcnsiblit les is attitude and behavioral an entirely natural one.. the fact data available anywhere that the two schools are separated for the investigation of by a distance of about 50 miles social cnange. should net pose any serious obstacles, linkage or the two computer 2. it is the repository for centers will be achieved, and arivstudies from many other mg time between the two institumajor u.s. polling organtions is just over an hour. an opizations, including the erating committee, with two roper organization, narepresenta ti ves from each of the tional opinion research three host scaools, has been estabcenter (ncrc) . bureau of lished, and it is charged witn coapplied social research, otdinatmg institutional involvebelden associates ment. (texas) , the minnesota poll, and dozens more. the williams college continued these thousands of surto bear the major responsiblity for veys are keyed guestionservicing user requests during the by-question into a matransiticn year (july 1, 1977 chine-retrievable june 30, 1978), although user servinstrument which makes ices and archival development acsubject searches effitivities now reside ic new haven cient. and storrs respectively. but we look to a continued prominent wil3. it annually acquires beliams role, through such areas as tween 300 and 400 surveys the publications program or the conducted by major non-aroper center, the development of merican research organiteaching materials and packages zations and the overseas where the standing of williams as affiliates of u. s. surone of the outstanding undergraduvey groups. today, apate institutions ir. tr.e country proximately 7,000 non-awill be a major asset, in tae hostmerican survey data sets, ing of special seminars, conferfrom 72 different counences, and training programs. tries and dating from 1938, are available througn the roper center. these include.' for examgle, 275 studies from enmark, 415 from france, 450 from germany, 900 i,"ijsist newsietttir, vcj.. 1, \ 51; ' 8i5' l i s3»8 j? iti jiir ij jm i i ! pilijl i 3 c ' i ii s' ii i! i t i i ^ ? 3 nj 5 fall/winter j987 iassist quarterly 23 si ii ^ff in' n si to 5 •ml + 0) .§, jj • 0) -o s]udj5jluuji jo joqixinu |d:io| :: m :• o • 0) :: m u :: 00 s "i oo ' ••id o t 00 «" fall/winter 1987 iassist quarterly 27 (d o c "> \_ q_ n d) -m o c 0) e a q) c d ro 6 n 10 * »r^ n h0) n o> cm ^ .t^ ^ ^ 00 n oj co (d ^ cvi 't (0 (0 d c n y •0 c t3 c 3 1. •d lij v n i n c \. go i] 'c n c v r u t e o r n < q < $ c > * ^-i l n t> 4^ zi. d c u u < z q. z z o o i u) < m u cl --1 ^ t x ih v .? ifi ni *s m ^) <^ icl 01 ni li h m fn in 1 u n n 3 ^ 17> k o i a n (> o n , « o i^ r • u o o p. c fall/winter 1987 28 lassisl quarterly id 00 { e / 1 \ o " '^ m 0) m 0) n ro (n o o (£) x; o) o u u m lo -u en '^ 1 oo ' o in' 0) o 1 • '"" e! • ^i ' 1 0) in ~l 1 to it] 0) d 1 i uopd|ridod ooo'ooi j^d sopy fall/winter 1987 iassisl quarterly 29 population distribution by regions millions 32 r12.5% 11.4''/o 7.7% 35.0% 24.0% 9.0% 1951 1961 1971 1981 1991 2001 l-all w !uut l^s' 30 iassist quarterly 1 co . o x) 1 0) 1 o t; r1 .« o i 1 o „° / • o o 10 f • o q / • cm ^ / • • d / ^" •i-i • 0) ^ / • 1 6 / o 0) jj / • 0) .-1 u / / n / c ," d .»^ ^ 4j (d i ^ 00 u « *" c^ : 1 :t< j a " •h id ' 1 0) n o « (d 0) ^0) u 1^ ^ d h• • r"^ ti f ^§ ?, 1 0° " <^ • (0 «|" ti ^^______j^ •0) e^ ? .-h ~ " ;3 '^"^'^^^^ 8| 5 ii 8 &< \ . 1 eh 53 »-• '— -^-^^^^^ ' i c id 1 1 1 1 1 1 f 1 1 • a . x. . 1 1 1 d 1 1 1 1 1 1 1 1 1a({oaonoaq 1 1 1 1 1 1 1 uj 3 to^ % fall/ winter 1987 iassisl quarterly 31 m td u (d in o) ^ o f () ^ 0) r— cr c u o u fi u 01 o 'u r u m o ^ cm j o) in fall/winter 1987 32 iassist quarterly ^ co en en id 1 1 0) co (d in (d cd cj) cd cd (n cm 0^ cd in tn cn tsi n co 0) co id 1 0) 00 (0 (d 0) (0 'icd (d to to (d 0) to cd to o in "1 ^ ^ '^ ^ * 0) (u i. 0) 3 u z u. z z o c o 2 w m >f-all/ winter /w7 lassisl quarterly 33 fall/winter 1987 34 iassist quarterly (0 oo 05 in ^ 0) n (0 *i d i— i ft 6 a cm r in men 15-24 /' / c / / / / / 0) > c in / i 1 _ / \ / / ( c e \ • \_ -\ ! / / /' / \ \ \\ \ \——1 1 — —+\\ —1—1—1 — — h h h 1 f 0) 8 00 » 0) p si ro si 4a ti 01 d'^^ c * fall/winter j 987 iassisl quarterly 35 o n q) wl q 4-1 •r-l w > •l-l t 't oo 1 0) 1 ^ (0 oo 0) 1 1 i "^ 1 0) ' y-' vr t f t>i . ^q ^ ^ -• ^•.t o cn 3 in .'.£ sss859s5? sooo fall/winter 1987 36 lassist quarterly w m u (d > .i-h c , \ c i) \ \ <1j f h-o o 1 n si 00 t:-ul fn *^ qo: 5 e :^u fall/winier 1987 lassist quarterly 37 bdttdhlfiiibhhhhmbaha i^imgmsggiiiiiiiyia^gii contents features families: diversity the new norm new family l>pes are transfurming scxial and legal iradiliuns as well as market ihinking by maryi anjie burke expanding the choices cable leievision, saielliies and vcr's are creating a new home entertainment and information environment by ted \1'tiniiell nuj cruig mckie changing health risks the primar>' threat to the health of canadians has shifted from infectious diseases to socially or environmentally related conditions. fo)' mury antte burke the law — a changing profession profiling canada's lawyers by crciig mckie trend reports canada in the 21st century foreign students child care births to lnmanied women private education canadian social trends editor associate editor managing editor assistant editors consulting editor -^-— > art direction design marketing and proinution composition 13 3 11 19 20 27 david bruscgard craig mckie colin lindsay mar>' anne burke, jo anne harliameni —tisa kwong '" • " " greg moore kobeno guido, baicejamieson, jill rcid judith buehler, kaihr^n bonner monique legare, rachel mondou acknowledgements c>iiihia bieerb. gordon priest, boriss mazikins, owen adams, doug angus, jim nucdonald, .michel durand, u'endy hansen, howard clifford, anatole romaniuk, sandra ramstxjitom.john silins, david bray, ian macredle, andre libelle, sylvie mercler. beryl gorman, lucie laniadeleine, ricarda windthorst, georgciie gaulin, daniel scoll and sylvie blais tuvclrlsretulhijj'im dr)ninc ) ill cjfudd, t.>uv^-j u'/i x 2i)'/. in . ihyb. njlidiul issn 0tl31 5698 fall/winter 1987 38 iassist quarterly i : -\ contents low income in cunada by suzanne melbul the changing indumrial mix of employment, i95ii985 by w gilnut i haul ihe decline in employment among men aged 55-64, 1975-1985 by colin linjiuy increases in long-term unemployment byjoanne i'arluiinent lifestyle risks: smoking and drinking in canada by cruig mckie low educational attainment in canada, 1975-1985 by bi i^itla anidli migration between atlantic canada and ontario, 19511985 by mtiiy anne hiiike a profile of employed migrants between atlantic canada and ontario by hubert hiscutt social hiuicaturs art direction and composition design photos promotion 35 39 canadian social trends editor craig mckie managing editor colin lindsay assistant editors mary anne burke. jo-annc parliament hublications division, staiisiics canada griffe design plioio centre, ssc cheryllynn irelandttotry donaiucci ' ^ review committee jw coombs, j hagey, d b peine, ge priest, l t pryor, m rochon acknowledgements mariin dials, catherine bronson. litryl gorman, lucie lamadtlcinc, myriam laportc, isabelle lavoie, louise pavelcy, cheryl sarazin, daniel scon, caihy shea, tim stringer cin^dun social trend) i cjiil.i^ur 1 i uuhi i ii published luunimri a >c,irh> slaiiiiii.» canada. puhlilaiion sjic» uiij»j. unlit. o, (jiijdj kia 016, itlcphu^cl6l)|^l';^^u7a copyrinhi 1';h6 b) siiniik^ cinidi. illtmhiirtjcrvcd fic>nb» po>li8c pjid ii oiijwi oncjno, cjnjdj slbbt-hiition kates hiiititm cjnjdj huclicahccc sinijlcixuc • i 2 sucich.ik jnidi i is cl>c»hcfi: srnj >ub>ic.pihin u.dcrwiid jd dlciii.h-j»grwuslil.mki(_jnjdi, hublii ju.in siki oiii» i. oiujfiu cjnjdi ma 0tb pic j^c supply b,ith n.x.ctkslonliinbr i ..rtc.p.jndcnic mi) be ldd.c-»><:d li.lhc tdil.ir. ciiia rends, i hhhooi jtin 1 il.m bu.ld.nu uiiiwi, oniinu, k 1 a ofb cinadlin socul trcndj iny iriklc herein, piui idcd trcdii is given lo siiiiiiii> cinidi ind canadian social t: pcciil perniissiuni ui hulk nrders should be iddicssed in the ncircsl regiuuil ollue ^ iddi i.ihci t panuge. hunt tttand, b ( r idvi 1| llubhc: fall/winter 1987 iassist quarterly 39 yj^iiasii^gmiauimmaamitiumm tgmtttktuuuuii uuit tma contents annual review uf labour force trends 2 by culiu liniliuy and cruig mikie the growth of part-time worli 9 6)' midry anne burke education in canada: selected highlight!) 13 by jo anne parliament french immersion 22 by jo-anne farlianieni immlgraiion 23 by mary anne burke the labour force participation of immigrants 28 (adapted jroni an article by nancy mclaughlin) trends in the crime rate in canada, 1970 1985 33 by culin lindsay common-law: living together as husband and \^ ife 39 without marriage by craig mckie the value of household work in canada 42 (adapted from an article byj l swinanier) social indicators 43 canadian social trends editor craig mckic managing editor colin lindsay assistant editors mary aiinc burke. jo anne i'jrliamcni art direction and composition design photos review committee acknowledgements publilaiiuns division, siuiisiics canada griffc design regional and indusirial expansion. plioio ccnirc, bsc j w coombs. j hagcy. d b pctric. g e prics^e t pryor, m rochon sylvie blais, lucie lamadeleine, elizabcili marcclla, kale mcgregor, suzanne mclhot. louise quiiin, sandra ramsbollom, daniel scon. cailiy shea, tim sinngcr canaduo sotlal trciidsk aiilugut 1 1 008e) l^ published (our linio i ycir b) suiiiliti canjdj, publicj nun biles, oniwi, onlanu, oiuda, kia 0t6. ickphonc i6h| vyi-sirs lopinghl isiso b) slalislics cjrudi, ill rikhls tck-r\cd firsl cbss posiigc pjid il olljwa, onuno. tinjdj subscription rates in 1 )cir in cjnjdi. iso ckcwhclt single isiuc »i2 so cjih in cjnjdj. «i5 elsewhere send subslnpliun urders ind address lhanges (u siaiislks canada, hublilaiiiin sales. oilawa. unlario. canada kia oto please supply both uld and new addresses and alluw sii weeks tor change corre>pondencc may be addressed lu the editor, caaaduo social treads, 1 lib flour, jean talun building. ottawa. ontario, kia oto caoadlan social treads is not responsible lor unsolicited materials permission is granted b) ihe cupyriglii owner lor libraries and others lu pliotocopy any article herein, provided credii is gisen lu statistics canada and caaadlaa social trends. reciuests lur special permissions or bulk orders should be addressed to the editor cover; manlluba parly b) william kuteiek oil and pencil on board -ih" « 00' i'xh '(mrs ttilliain kurelekl the isaacs cillery , toronto collection national gilltry ul canadi, ottawa issn 0831 5698 fall/winter 1987 iassisl quarterly 3 downloading for pc users; part i: the u.s. government experience by donald f. harrison w. jon heddesheimer national archives and records administration' various ways of writing data on diskettes, formatting considerations (dbms or ascii) are discussed. past and fixture changes in the cost, speed and capacity of micros are stressed. program details such as pricing and contractor versus inhouse production of floppies are mentioned. the point is emphasized that federal downloading is restricted almost entirely to small, simple files formatted in a dbms because agencies do not consider the flexible diskette suitable for storing large amounts of data. major statistical producers expect mass storage devices like the cdrom to become generally available to micro users. they fiirlher anticipate thai soon the average pc user will be able to download and manipulate large ascii files. the authors discuss various electronic data communication systems, such as the nav/s dif and iso 821 1 as alternatives to media transfer. this leads finally to a discussion of digital commumcalion methods such as bitnet. netnorth and earn. the authors conclude with predictions of the effect on the future operations of data archives. jon (introduction): in this session we intend to describe the efforts of four federal agencies to serve pc users by downloading onto flexible diskettes information traditionally offered and sold on tape. we will use a dialog format to encourage you to interject your own comments. what follows is the edited manuscript of a dialog presented at the marina del rey meeting of the international association for social science information service and technology, on may 22. 1986. it focuses on the programs of four federal agencies which download onto flexible diskettes information traditionally offered on tape. after describing ' the opinions expressed in this article arc solely those of the authors and in no way reflect the official position of the u. s. national archives and records administrauon. don: there are many different ways to define the term downloading. my ibm pc users manual defines it in terms of simply "printing out" material in hardcopy. my pc talk users guide defines it as receiving data from a remote terminal. for the purposes of this session, the "down" part of the word refers to moving data "down" from a mainframe (or mini) to a smaller computer (mini or micro, in this case to a micro), and the "load" part of the word will refer to the writing of that data on a 5 1/4" fiexible diskette — commonly referred to as a "fioppy" — so that the data can be spring 1987 4 iassist quarterly manipulated on a micrckomputcr, in this session we will concenuaie on downloading as a reference service lo the user public. jon: the following four possibilities by which automated records can be downloaded from mainframe files onto floppy diskettes for users of microcomputers have been suggested in our research: a. downloading files in bulk. using a package, the 'downloader' transfers an entire file in straight ascii. this requires careful calculation to make certain the file does not overwhelm the pc's capacity. also, because the data are relatively unformatted, the pc user will need lo be well acquainted with the mainframe's procedures of access in order to retrieve and use the data. b. downloading files in a data base management system 8dbms) formal, such as lotus or dbase. these programs pre-format the information to allow the user to begin work at once. ttie vast majority of floppies offered by government agencies are in a dbms formal c. selective data access. this allows the user to request selected portions of mainframe data either from a single file or across a range of files tied together in some fashion. ttiis can be accomplished using generalized menus or customized packages, such as, in a corporation where the command "marketing department budget" might trigger the downloading of abstracted relevant information. d. cooperative processing. this method involves expensive custom packages which allow the user to establish an intelligent connection between mainframe and microcomputer applications. don. agencies and institutions which make machine-readable data available to researchers find that there is a growing interest in storing and manipulating data by microcomputer. some researchers are reluctant to use traditional computer service centers with mainframes and programmers. they would rather perform the research themselves in the privacy and convenience of the office or home. moreover, today's microcomputers have the storage capacity and the sophistication of yesterday's mainframes. jon: yes, don. this is part of a general trend away from large mainframes and central processing facilities, accompanied by a great increase in the capacity of pc's. offering data on floppies is one way for an archives to harness this trend and increase visibility and clientele. don: in preparing for this session, jon and 1 located several very small federal offices which download for their users on a "swap" basis: users send in formatted flexible diskettes which the agency vnites on and returns to the researcher. the national register of historic places, maintained on-line by the national park service, is one example of this practice. don: this presentation, however, will concentrate exclusively on the experiences wiui downloading of four federal agencies. these agencies have had considerable experience distributing data on tape and, just in the past few years, have begun to write data onto diskettes. they are: the national technical information service (ntis), the bureau of economic analysis (bea), the bureau of labor statistics (bls), and the bureau of the census (census). don: ntis was created by an act of congress to make available, to the public, material created by a variety of federal agencies because these agencies do not have revolving funds (such as the national archives trust fund) with which they can sell copies of records to the public. over the years, ntis has become a broker for both textual and non-textual records. spring 1987 [assist quarterly 5 ntls's revolving fund allows it to set prices, receive monies, and publish catalogs. ntis sells its technical information products and services under the provisions of title 15 of the united slates code. in the early 1970's, it staned to accept machine-readable data files (mrdf's) from federal agencies and. since then, has been selling copies of tape files to the public. in the last year or so, ntis began a program whereby any of its mrdfs could be ordered on fiexible diskettes written in straight ascii or in one of several commercial, packaged data base management systems. ntis sees absolutely no problem with very large orders, and it is up to the customer to determine if her/his microcomputer can accept the file volume; one popular offering is written on 87 hoppies. ntis is the largest seller of downloaded data; it far exceeds all the others combined. jon: bls was established in the early twentieth century and has been in the forefront of statistical analysis from the beginning. bls, as well as the remaining two agencies to be discussed, believes in making its files available to the public directly from the analysts who created them. consequently, bls has for 15 years been publishing a catalog of data files available on tape. two years ago it launched a program of making them available on flexible diskette as well. each division within bls does its own analysis, its own downloading, and sets its own prices. jon: bls is uniquely qualified to download data due to the existence of a labstat "umbrella" system. created for in-housc research, over 25,000 search data elements can be tapped across an enormous range of statistical data. because this agency thus has the capacity to allow researchers to browse through the data (on-line) and pick and choose from all files, it has the capability to pioneer a similar downloading service to the public. jon: the bea, like the bls, believes that its analysts alone are capable of dealing intelligently with the user. therefore, it makes data available to the public directly, not through any broker or archives. but its files are so voluminous and complicated that the downloading program encompasses only one data set, which was previously oftered on microfiche. therefore, this modest program is operated "in-house" by one person, merely as an extension of an existing service. the agency has no plans to expand the program. don: census has been gathering statistical data since 1790, and began working with machine manipulated data with the 1890 decennial census. not unexpectedly. census has also been in the "downloading business" longer than any other federal agency. the data user services division, which has been making public use sample data files available on tape for many years, began in 1984 to make these same files available on floppies. jon: yes, this program was the firsl all agencies contemplating using downloading as a part of their reference service began by visiting census. ttie agency spent a great deal of time and money aeating custom packages which subdivide extensive databases into "pc-sized" chunks. jon: thus we are discussing four programs, the oldest of which first offered downloaded files in march, 1984. this is a new service for the federal government, which has thus far been slow to provide data on floppies. this is due partly to lack of resources for new services. but even more important, it is our contention that the low level of downloading activity is due to beliefs about the immediate future of the pc and the data medium it uses, which we will discuss in this dialog. don: one must calculate carefully before writing mainframe data onto floppies. in the hrst place, not all files can be downloaded to the pc because of manufacturers' limitauons and spccincauons of the diskette. for example. spring 1987 6 lasstsl quarterly one must consider the size of ihe file as il exists on ihe mainframe and the complexity of the language in which it is written. one flexible diskette, filled to capacity with a fiat ascii file, contains 362.496 bytes, (ix)s 2.1 formatted, double sided, double density). this figure is insignificant when compared to a standard reel of magnetic tape (11" reel, wound with 1/2" tape. 2400' long, and filled to capacity with 9 track. 6250 bytes-per-inch). such a tape could contain approximately 120 million bytes, or the equivalent of more than 150 floppies. furthermore, when the data are formatted on the diskette to accommodate a microcomputer data base management system such as lotus or dbase, the diskette will probably hold far less than 362,496 bytes. of course, the obvious solution is to add more diskettes. ntis offers one cartographic data file from the central intelligence agency called "world data base 11" written on move than 80 diskettes. in order to manipulate the entire file, one would have to load all 80 diskettes into the pc first none of the other three agencies offers files of this magnitude on floppies, because they believe that users cannot or will not use them. jon: to illustrate don's point, one can subdivide downloading programs into "active" and "passive." one agency makes everything available from its extensive holdings through a contractor and in most commercial dbms formats. if the researcher is comfortable with 200 floppies, so be it this is "passive" and a growing trend. another agency tries to be somewhat active by limiting what it will offer to 2 diskettes per file, but does little beyond that two agencies offer their holdings only in lotus 1-2-3 format one agency, by far the most "active", spends a great deal of time and effort subdividing complex files into compact units convenient to the pc user. jon: outside the federal experience, but certainly worth mentioning, is one data archives with limited resources which offers workshops on how to download and formal data for use on a pc. this approach encourages users to master a fortran program which enables them to download, and thus make greater use of the tapes held by this archives. they are saying, in effect "we will give you the tools and turn you loose." of course as a pc owner, you have to be very serious to use this as a research strategy. don: software formatting is a second consideration. many persons we interviewed agreed that all micro users are divided into two types: the mainframe expert "data junkie" who can work with unformatted "flat" data, and a newer, less energetic user on the scene. this second user prefers to insert a fioppy, flip a switch, and let the machine do the rest in many cases, this new user just purchased her/his pc last week and expects instant results. at least two of the four agencies we spoke with accommodate this latter type by formatting the data in a data base management system which allows the user to begin data manipulationimmediately. (lotus and dbase are two of the several dbmss available). formatting involves considerable reworking by the data producer. the payoff, however, is in attracting many more users, in fact an entirely new market of users, far different and more numerous than those who have traditionally ordered magnetic tape for use on a mainframe. we suggest that these clients are best served with data distributed according to the specifications of the individual order. this avoids consumer complaints and unpleasant scenes with commercial software manufacturers. don: hardware formatting is another matter completely. there seems to be a universal preference (among data producers, data brokers and data users) for ibm-compatible diskettes. most prefer dos formatting over cp/m as well, though this doesn't seem to be an issue. jon: 1 agree that using hardware designed to be ibm-compatible is far more important than spring 19s7 iassisl quarterly 7 deciding on one software package. i would like to add that in the next few years, the ordinary user should be able to deal with software independent information and do easily what today only a "data junkie" can accomplish. first will come graphics packages, a process already underway. next will come the tools for the inexperienced to handle large quantities of raw data. along these lines, i expect the storage capacity of micros to increase twcv-fold in two years and ten-fold in five, with little increase in cosl software designers expect this to happen and already are hard at work. don: the strategies of pricing suggest as many solutions as there are agencies. they depend on what the individual agency perceives its internal costs to be, whether or not it is dedicated to the concept of service to users and, in the case of agencies using a contractor, a markup from the contractor's costs. these strategies may be compared with the production of microfilm publications in traditional archives and libraries. data producers (tape or floppy), like microfilm producers, may produce the product on demand, and charge the first customer the full amount required to recover costs; the second and third customers, in turn, pay only a marginal fee. a second solution is to spread the charge over several users. yet another solution is to absorb the cost of preparation and simply charge a fiat fee. in addition to production of microfiche publications, this range of solutions seems to occur with producuon of files written on either tape or diskette. in the case of diskettes, there is great diversity in charges: one agency charges $35 per diskette; another charges $75 for the first fioppy, and $15 for each additional diskette in the same file; yei another agency now charges $60 for the first diskette and $12 for successive ones. our fourth respondent, by producing only one file for the public and updating it each month, charges a fiat $240 per annua! subscription of 12 monthly installments; this amounts to $20 per diskette. jon: circular ado, issued in 1986 by the office of management and budget, instructs agencies to charge incremental costs incurred in serving researchers. information is defined as a marketable resource, thus tacitly refuting the notion of public service. pricing policies vary. for normal orders, the charges expressed are "so much per diskette," not "so much per file." in general, large user service organizations charge handsomely for special considerations, special formats, special tabulations, etc. one agency representative stated that "by-the-book" processing charges are so extreme that the final output (tape, floppies, hardcopy) costs about the same regardless of medium. this is partly due to the cost of getting the data ready for the user. processing data so that the user can work with them frequently constitutes the major portion of the entire cost of a service order. and, of course, a part of this problem is also that to write data onto a floppy sometimes requires reworking and reformatting. don: since floppies are fragile, hold a comparatively small amount of data and can be accidentally erased, each of the four agencies we spoke to have considered other media for use on the pc. almost all otncontacts discussed the use of the compact disk for digital data storage. since the data are written with a laser, there can be no problem with accidental erasure or overwriting. it is a read-only mode. unfortunately, it requires a separate and extremely expensive disk drive. all persons we spoke to predicted that compact disks will soon be both plentiful and economical. what makes the "compact disk-read only memory" (cd-rom) so attractive is that it will store the equivalent of 1500 flexible diskettes — or the equivalent of 4 high density mainframe tapes. jon: data professionals are also giving consideration to an interactive storage device with the characteristics of hard disks. this, if it becomes a reality, is further down the road than cd-rom. (at present, "read only" is considered a virtue.) the bernoulli box, while spring 1987 lassist quarterly loo expensive for individual use is nonetheless indicative of what might become commonplace once pc's are bener able to accommodate and manipulate large data bases. thus one can confidently predict that a standardized, economical, mass storage replacement for floppies will soon be available. the average user will not rely exclusively on floppies and indeed, may not use them at all. jon: however, considerations might be reversed in the future when pondering whether or not to store data on the cd-rom or the bernoulli box. there might be a case where too little data is requested, even for the economical use of a diskette. such cases are tailor-made for the electronic bulletin board, which is designed to transfer small bits and pieces of data rather than huge data bases. one agency routinely gives small bits of data to users for the cost of a phone call instead of charging the price of a floppy. electronic bulletin boards are becoming increasingly popular, for the presentation of finding aids, lists and other advertisements. don: electronic bulletin boards depend on electronic data transfer. instead of writing the dau onto a diskette and shipping the diskette to the researcher, it involves sending the data across a telephone or wireless circuit from the archives in which the data are stored directly to the researcher's microcomputer. it is a method that is gaining in popularity and could easily be the prefened data transfer method by the year 2000. don: however, electronic data transfer also introduces problems of interchange formats. these involve the use of data filtering devices which are important when machines of differing specifications from several manufacturers and using different software are communicating with each other on-line. in march, 1983 the u.s. department of the navy initiated a cooperative effort among government and leading office systems to define and test a document interchange format (dip) which vendors could support today dif permits the interchange of textual data between word processors, providing about 95% of their document formatting needs. twelve of the fourteen manufacturers who cooperated with the dif test were datapoint, data general, dec, hewlett packard, national cash register. sperry, motorola, at&t, four-phase, xerox, wang and ibm. the dif generally would filter textual data files between microcomputers. but what about statistical data files'' and suppose one of the computers is a mainframe communicating with a microcomputer? don: at about the same time that the navy developed its dif, the international organization for standardization developed a set of standards which incorporates a mechanism allowing statistical as well as textual data structures to be easily moved from one computer system to another, independent of the manufacturers. the resulting system is called, "iso 8211." it is much more flexible than the dif and will accommodate magnetic tape, disk packs, flexible diskettes and data interchange over communication lines —in any combination, either as a source or a target it will accommodate files with variable length records as well as those with fixed length records. user file structures such as sequential, hierarchical, relational or indexed, could be connected with the interchange structure. therefore, it can be said that iso 8211 is both content and media independent jon: once an operation involves more then a few files, or even perhaps from the very beginnmg, employing a good contractor is probably the best solution to the problems of disseminating downloaded data. the risks of floppies becoming obsolete, or of incorrectly anticipating researchers' needs are passed on to the contractor. contractors with extensive experience are numerous, and will allow the researcher (for a price) to define his specifications. contractors now offer agencies a hal price of under $25 per floppy, even if lhe\ spring 1987 iasstst quarterly 9 have to copy and reformat the original tape, and for this price will keep a copy in a contractor maintained library. thus it is now possible to make money on an initial order and still charge reasonable prices, a situation only true in the last few months. before that, even the largest organizations foimd it necessary to sell several sets of a file before recovering their costs. don: an undeniable advantage of an in-house operation, especially in the early stages of building a floppy program, is that you can work with a researcher in developing files. you can also help her/him select a simple, established file to get the "fee!" of things. individual program managers still tend to offer this flexibility in decentralized agencies (in two of the four agencies we contacted). also, an agency can send new files to sophisticated, established users for comment prior to public release. there will come a time, however, when economics coupled with instructions from the office of management and budget, will make such operations impossible for all four agencies we contacted. the present reality is that these four agencies will probably have the choice of using a contractor or having no program at all. jon: the justification for modest or non-existent downloading programs in major statistical agencies is that floppies will soon cease to be the medium of choice for downloading to pc's. don: also, given the consensus of opinion that micros will vastly increase in capacity and that the ordinary data user will be able to deal successfully with unformatted "fiat" ascii data, it pays to look down the road rather than be frozen in the present jon: 1 recommend that you use electronic bulleiin boards, first as a catalog to advertise holdings and later to hold simple updates and smaller data bases. using an electronic bulletin board as a catalog and as a vehicle to answer routine researcher inquiries should greatly increase the visibilty of your collection while eliminating a great deal of your reference load. ideally, reference persoimel should deal only with special requests. to ask reference personnel to answer the same questions over and over is both expensive for administrators and demeaning to professionals. gte sprjnts "pc pursuit" represents the wave of the future and is an example of how the general public can tap into electronic bulletin boards, or even on-line data bases. this relatively new service ofters unlimited nighttime and weekend access to 14 cities from over 200 telenet areas for $30 per month. jon: 1 would suggest also that you consider designing "umbrella" systems which tie your holdings together via common data elements, making the researcher her/his own boss regarding data access. i have already mentioned labstat. another example is to be found in the networks which have developed between ibm installations throughout the world. started originally as independent data collections to service specific needs, they have gradually become integrated and can be accessed by the casual user via standard menus, (given, of course, the constraints of access levels.) unlike labstat, which was developed from the beginning to be somewhat like it is now, these systems were cobbled together artificially to fend off competition from other sources and to satisfy internal corporate requirements. they serve as an excellent example of how data archivists can develop menus to link together their own collections and eventually develop the means whereby researchers may consult numerous data archives utilizing standard search commands. don: this technique is already in use in libraries and manuscript collections in the united states. consortia of library collections are accessible by computer networks in at least two very large systems; the research libraries information network (rlin) and the on-line computer library center (oclc). however, spring 1987 10 lassisl quarterly for the mosi pari, whal one reaches via tjiesc systems are descriptions of the collections, not the collections themselves. this is very different from our thesis. what we are advocating is that the data, as well as descriptions of the data, be made available to remote locations, in a library research area, for example, or in a researcher's work/terminal area, in short, anywhere a researcher has access to a microcomputer and a modem. don: variations of this are at work in digital communications networks. collaboration between academics at widely separated locations is becoming more and more common as technology allows them to exchange ideas and information more easily. at research locations throughout the united states, a network called bitnet is playing a major role in this information interchange. bitnet is a cooperative digital communications network connecting over 1,200 computers in universities and other educational and research institutions. by connecting to bitnet one can gain access to computers on the majlnet, earn (europe) and netnorth (canada) international networks. using the facilities at any one of these four systems, data archivists could easily transmit pan or all of their holdings to researchers located within the geographic confines of the others. a combination of the filtering systems laid between hardware of diverse manufacture together with the use of the new, world-wide, inter-computer communication networks will revolutionize the transfer of electronic data. conclusions as data archivists we must observe this research trend of exchanging mainframes for micros in order to perform statistical and other kinds of analysis on machine-readable records. a reference program offering data on floppies through a contractor is a good way to begin tapping this market which is already far larger than that for tape. further advances in the capacity of micros coupled with new software, an inexpensive mass storage medium to replace floppies, and continually declining costs will allow individual pc users to function more and more like data centers. all this should only serve to increase the market for downloaded machine-readable data files. information transfer via floppy diskettes or other media is but one way to move data from mainframe to micro. an electronic bulletin board could be started by a catalog listing, and later enlarged using electronic data interchange over one of many communications networks such as bitnet. finally, archivists should plan to develop "umbrella" systems which will tie holdings together in a given repository and ultimately establish access to other collections.n jon: data archivists, therefore must ponder the integration of their own holdings using umbrella systems, while learning to relate these to the contents of other repositories through digital communication networks. while accomplishing this may initially require additional resources, long run costs should decline or stabilize while user access increases. spring 1987 6 iassist quarterly 2016 / vol 40 no 3 iassist quarterly iassist quarterlyiassist quarterly demonstrating repository trustworthiness through the data seal of approval by stuart macdonald1, ingrid dillo2, sophia lafferty-hess3, lynn woolfrey4, mary vardigan5 data seal of approval offers a basic, lightweight certification standard. abstract this paper is a summary of a panel session which consisted of five presentations given on trusted digital repository certification through the data seal of approval (dsa) at iassist 2015 in minneapolis. the paper begins with an overview of the dsa complemented by case studies illustrating how archives undertake the process of certification and concludes with future plans. keywords dsa, data seal of approval, trusted digital repository, certification, data stewardship, digital preservation introduction the data seal of approval: a fitting label for trustworthy data repositories (ingrid dillo, dans) the data seal of approval is a basic, transparent process for digital repositories to certify that they are sustainable and trustworthy. assessments are conducted first internally by a repository and then reviewed by community peers. assessments help data communities – producers, repositories, and consumers – increase compliance with an awareness of established standards. national and international funders are increasingly likely to mandate open data and data management policies that call for the long-term storage and accessibility of data. if we want to share data, the long-term storage of those data in a trustworthy digital archive is a sine qua non. data created and used by scientists should be managed, curated and archived in order to preserve the initial investment in collecting them. researchers must be certain that research data provided by the archives for secondary use remain useful and meaningful, even in the long term. vol 40 no 3 / iassist quarterly 2016 7 iassist quarterly the concept of sustainability is challenging and crosses several dimensions: organizational, technical, financial, legal, etc. certification can be an important contribution for ensuring the reliability and durability of digital archives and, hence the possibilities for sharing data, over a long period of time. the data seal of approval offers a basic, lightweight certification standard. the dsa enables any organization, regardless of size or staff, to quickly self-assess how they are performing compared to data community standards. the dsa, developed by dans (data archiving and networked services) in the netherlands, was first presented at the first african digital curation conference in 2008. the dsa assessment criteria were initially developed for use in the netherlands, but were soon found to be very useful in an international context as well. thus, in 2009 the dsa was transferred to an international body, the dsa board, which has since managed and further developed the guidelines and the peer review process. the dsa aims to safeguard data, to ensure high quality and to guide reliable management of data for the future without requiring the implementation of new standards, regulations or heavy investments. the data seal of approval: • gives researchers the assurance that their data will be stored in a reliable manner and can be reused; • provides funding bodies with the confidence that research data will remain available for reuse; • enables researchers to reliably assess the repositories holding the data they want to reuse; and • supports data repositories in the efficient archiving and distribution of data there are 16 guidelines in the dsa: three focusing on the data producer, three on the data consumer, and ten on the data repository. 1. the data producer deposits the data in a data repository with sufficient information for others to assess the quality of the data and compliance with disciplinary and ethical norms. 2. the data producer provides the data in formats recommended by the data repository. 3. the data producer provides the data together with the metadata requested by the data repository. 4. the data repository has an explicit mission in the area of digital archiving and promulgates it. 5. the data repository uses due diligence to ensure compliance with legal regulations and contracts including: when applicable, regulations governing the protection of human subjects. 6. the data repository applies documented processes and procedures for managing data storage. 7. the data repository has a plan for long-term preservation of its digital assets. 8. archiving takes place according to explicit work flows across the data life cycle. 9. the data repository assumes responsibility from the data producers for access and availability of the digital objects. 10. the data repository enables the users to discover and use the data and reference them in a persistent way. 11. the data repository ensures the integrity of the digital objects and the metadata. 12. the data repository ensures the authenticity of the digital objects and the metadata. 13. the technical infrastructure explicitly supports the tasks and functions described in internationally-accepted archival standards like oais. 14. the data consumer complies with access regulations set by the data repository. 15. the data consumer conforms to and agrees with any codes of conduct that are generally accepted in the relevant sector for the exchange and proper use of knowledge and information. 16. the data consumer respects the applicable licences of the data repository regarding the use of the data. the dsa guidelines the guidelines are based on five criteria: the data can be found on the internet; the data are accessible (clear rights and licenses); the data are in a usable format; the data are reliable; and, the data are identified in a unique and persistent way so they can be referred to. 8 iassist quarterly 2016 / vol 40 no 3 iassist quarterly obtaining the dsa involves two stages. first, the repository conducts a self-assessment, documenting and compiling evidence of compliance into an online tool. a community peer then evaluates this self-assessment by confirming and validating the evidence. the self-assessment, including all evidence, will only be published on the websites of the dsa and the applicant’s data repository after the dsa has been awarded. since approved applications, including any evidence and peer review comments, are publicly available on the dsa website, they can be used as references or samples. this openness fosters trust and accountability as the assessment is accessible by all stakeholders. today a total of 55 seals, including 8 renewals, have been awarded, and some 45 digital archives are working on their dsa selfassessments. this steady growth shows that there is a clear demand for a less resource-intensive approach to certification of trustworthiness of digital archives. the dsa is also a good springboard for repositories interested in completing the more rigorous and comprehensive assessments and certifications available such as the iso16363 audit and certification of trustworthy digital repositories. in addition, the dsa can be used as a roadmap and a planning tool for repositories that are just getting started. case study 1 odum institute for research in social science, university of north carolina sophia lafferty-hess introduction the h. w. odum institute for research in social science at the university of north carolina at chapel hill was founded in 1924 making it one of the oldest social science research institutes in the united states. the odum institute data archive was established in the late 1960s and today supports researchers’ data management needs throughout the research lifecycle, including providing a trustworthy repository for the long-term preservation and dissemination of data. because of its commitment to data stewardship, the odum institute data archive has been actively working to demonstrate its trustworthiness through transparent policies and procedures, self-assessments, and certifications. as part of this initiative, the odum institute data archive successfully applied for, and was awarded, the data seal of approval (dsa) in september 2013. the primary rationale for applying for the seal was that it provides a transparent and public method to display to the broader community the archive’s commitment to following data curation standards and best practices through an external peer review process. the data seal of approval also plays an important part in the archive’s overall strategic plan for self-assessment and ongoing improvement. data seal of approval application process successfully completing the application for the dsa involved a four-step process that included: education, assessment, documentation, and submission. the education step included building familiarity with the dsa guidelines and criteria, examining successful applications published on the dsa website, and reviewing pertinent standards referenced in the guidelines (i.e., the oais reference model). the second step involved a comprehensive self-assessment of all current policies, procedures, and technical systems. during this step, policies and procedures were mapped to the relevant guidelines allowing the identification of documentation that needed to be updated or expanded to address the criteria in the guidelines. after completing this assessment, policy, procedural, and system documents were revised as needed and published on the odum institute website or internal wiki. archive staff then drafted responses to the guidelines using this updated documentation. the final step was submitting the online application upon completion of a final comprehensive review by archive staff. reflections on demonstrating trustworthiness completing a successful application for the data seal of approval allowed the odum institute data archive to not only demonstrate compliance with data stewardship best practices, but also provided an opportunity to reflect on the role of self-assessments and certifications in the broader context of their organization and the data curation field. detailed below are some key reflections as well as the odum institute’s plans for demonstrating trustworthiness into the future. demonstrating trustworthiness is a continuous process: demonstrating a repository’s trustworthiness is not a simple task and requires commitment to continually striving towards transparency and improvement in procedures and systems. the dsa provides a useful first step for demonstrating trustworthiness through a relatively lightweight certification process. vol 40 no 3 / iassist quarterly 2016 9 iassist quarterly certifications facilitate structured periods for assessment: often a repository can become so busy “doing data curation” that it requires conscious effort to assess whether procedures align with current best practices and standards within the field, which can be a moving target as technological and workflow developments continue to emerge. certifications, such as the dsa, provide an ideal mechanism for repositories to slow down, take stock, modify procedures and systems, and update documentation in accordance with new developments. demonstrating trustworthiness benefits from a supportive community: the dsa community-driven structure benefits the entire data curation community by providing a low-cost method for data repositories to demonstrate their trustworthiness. community members contribute by participating in the peer review process and these relatively nominal contributions, taken as a collective whole, make a significant impact on the data curation field and repository users who reap the benefits of trustworthy repositories. plans for continuing to demonstrate trustworthiness: since demonstrating trustworthiness is an ongoing process, the odum institute data archive plans to continue to strive for transparency and improvement through assessments and certifications. this plan includes renewing our dsa when updated guidelines are released and completing a self-audit using iso 16363 in preparation for undergoing an external formal audit. the odum institute data archive will also support the broader dsa community through participation within the dsa general assembly. case study 2 cornell institute for social and economic research (ciser) stuart macdonald introduction the cornell institute for social and economic research (ciser) was founded in 1981 and is home to one of the oldest, university-based social science data archives in the united states. its mission is to anticipate and support the evolving computational and data needs of cornell researchers throughout the entire data life cycle. the data archive houses an extensive collection of public and restricted-use numeric data files to support quantitative research in the social sciences with particular emphasis on studies that match the interests of cornell researchers: demography, economics and labor, political and social behavior, family life, and health. ciser data archive was awarded the data seal of approval in july 2014. dsa application process and approach ciser have long been committed to long-term archiving and providing access to scholarly research data in a sustainable way and trustworthy manner. formalisation of this commitment through the data seal of approval self-assessment process commenced with a review of documentation and existing case studies (archaeology data service and finnish social science data archive ) in november 2013. this was followed by a scoping exercise primarily to gain familiarity with dsa regulations and compliance statements (detailing quality aspects with regard to creation, storage and reuse of data as it applies to data producer, consumer and archive or repository) as well as 16 guidelines underpinned by the following criteria that determine whether or not data may be qualified as being sustainably archived: • data can be found on the internet • data are accessible • data are available in a usable format • data are reliable • data can be referred to. a cross-section of successful dsa applications were consulted that were identified as giving sufficient breadth of data archival practice in addition to providing discipline-specific guidance. elements deemed relevant and pertinent to ciser data archive practices were collected and collated in a series of spreadsheets in association with relevant requirements from the applicant manual for each statement. statements were then assigned to members of staff with particular expertise (storage, security, formatting, restricted data, metadata) and discussed in weekly meetings (1-2 hours) with separate meetings held to discuss in more depth individual assignments as required. separate meetings were also held to update policies and craft new policy documents to underpin archival process and workflow where they didn’t already exist. the application process from scoping to submission took approximately 12 person weeks (principally that of the data services librarian plus colleagues). observations and lessons to kick-start the application process ‘quick wins’ were identified and used to seed the submissions document such as referencing of existing policies, agreements, terms of use, guideline 0. information gathering and evaluation was in part an iterative process with knowledge, workflow and procedure being ‘scattered’ across the organisation, existing inside people’s heads, in technical documentation and legacy printed material (including policies), and in internal and external online links. as such, it was easy to underestimate the time required to assemble and craft new policies (such as preservation and storage, security, versioning, data collection), mission statement, and other public facing documentation as evidence to support the application. proofreading, consistency of language, terminology and narrative also took time to rationalise, bearing in mind the ‘different voices’ of staff member experts. 10 iassist quarterly 2016 / vol 40 no 3 iassist quarterly organisational and community benefits a number of organisational and community benefits were gained from the data seal of approval application process, namely: • clarification and articulation of organisation’s archival practices; • promotion of trust and confidence between the three stakeholders in the data supply chain producer, repository/archive and consumer are working to a common set of standards and principles; • easier to conduct future systematic reviews of technical/human processes and procedures; • better equipped to respond to necessary changes in data stewardship workflows as/when new compliant tools, technologies and standards emerge; • identification of service gaps and areas for improvement or modernisation in archival process and procedure; • raise the profile of the archive and preservation with cornell senior managers; • provide a holistic overview and perspective on the mechanics of a mature data archive for new archive staff; • foundation for further institutional trusted digital repository accreditation such as din 31644 (34 metrics) and iso 16363 certificate or tdr checklist (107 metrics); • uncover areas of mutual interworking and interaction between archival colleagues for the purposes of streamlining operations; • contribution to the social science data archiving community and the data stewardship profession by openly sharing processes, workflows and practice. summary in summary, the data seal of approval application process is beneficial as a learning and knowledge sharing experience for archival staff. it also provides the opportunity for an organisation to audit and enhance its archival operations. more importantly however, the data seal of approval is a public pronouncement of an organisation’s archival intent, to demonstrate reliable and trusted access to managed research data for the academic community both now and into the future. case study 3 datafirst, university of cape town lynn woolfrey introduction datafirst’s repository was awarded the data seal of approval in 2014, and is the only african institution to achieve this certification to date. datafirst is based at the university of cape town in south africa, but our repository gives online access to african data for researchers around the world. datafirst’s data service model and the dsa guidelines getting to the point of dsa certification was enabled by our twin strategies of, first, adhering to standards and, second, communicating regularly with our various stakeholders. the aim is to build confidence in our service as a trusted digital repository. trust in our service leads to use of our resources for data-intensive research and quality research output, generating further usage and demand. it also encourages data deposits, as data producers gain confidence in our abilities to handle and share their data in a responsible manner. dsa guidelines concerning the data life cycle dsa guidelines require that data curation should be carried out with a clear mission and according to documented and well-understood procedures. the requirement for repositories to understand and advertise their mission is stated as: the data repository has an explicit mission in the area of digital archiving and promulgates it (guideline 4) datafirst’s mission is to support top quality research on south africa and other african countries by providing researchers with access to african survey and administrative microdata. this is clearly stated on our website as a mission statement and in other information on the work we do. over 14 years of service provision, we have modelled the work of our data service. the model is based on the open archival information system (oais) model for digital repositories. this model was originally designed for space data systems. the oais has since become the standard for digital archives and is registered with the international standards organisation (iso) as iso 14721:2012. the model in figure 1 shows how we comply with dsa guidelines related to data curation processes undertaken by repositories. these include: figure 1. the datafirst microdata service model vol 40 no 3 / iassist quarterly 2016 11 iassist quarterly • the data repository applies documented processes and procedures for managing data storage (guideline 6) • archiving takes place according to explicit work flows across the data life-cycle (guideline 8) • the technical infrastructure explicitly supports the tasks and functions described in internationally accepted archival standards like oais (guideline 13) below we detail our compliance with the rest of the dsa guidelines, with reference to this model. dsa guidelines related to data deposit the data deposit stage is depicted as stage 2 in the model. dsa guidelines related to this stage are: • the data producer deposits the data in a data repository with sufficient information for others to assess the quality of the data and compliance with disciplinary and ethical norms (guideline 1) • the data producer provides the data in formats recommended by the data repository (guideline 2) • the data producer provides the data together with the metadata requested by the data repository (guideline 3) these guidelines ensure that data users are informed of the quality of the data in the repository. datafirst provides feedback from data users to data depositors, to address anomalies in the data. the service provides data quality notes on each dataset to highlight issues. data quality benefits can accrue from this type of independent assessment by academics. datafirst supports this “virtuous cycle of data reuse” (as depicted in the model). the dsa guidelines require depositors to provide data in formats that can be used to prepare a usable research dataset. depositors can provide data in a number of formats. for example, administrative datasets deposited with the service may be in inappropriate formats. however, datafirst staff have developed the skills needed to convert these files into research-ready formats. another dimension of quality is interpretability, which is dependent on the availability of useful documentation to support sound data analysis (statistics canada quality guidelines 2014 ). guideline 3 requires depositors to provide adequate information to help repository managers create useful metadata. datafirst communicates with depositors on an ongoing basis, to ensure this is the case. this interaction is also designed to ensure compliance with ethical norms. for example, depositors are required to confirm they have ownership of and permission to share the data. dsa guidelines concerning data assurance data security needs to be assured by those sharing the data, to engender trust among depositors and users. dsa guidelines related to data assurance are: • the data repository uses due diligence to ensure compliance with legal regulations and contracts including, when applicable, regulations governing the protection of human subjects (guideline 5) • the data repository ensures the authenticity of the digital objects and the metadata (guideline 12) these guidelines are fulfilled at the data assurance stage depicted in the model (stage 3), and also during data preservation (stage 6). protection of confidential data is assured through disclosure control routines followed closely at the service. we also encourage deposits of restricted-access data with our secure research data centre, which is a controlled environment at the university accessed by approved researchers. dsa guidelines on data preservation dsa compliance requires long-term planning by data repositories. the guideline dealing with this is: the data repository has a plan for long-term preservation of its digital assets (guideline 7) while nothing is permanent, datafirst has operated a data service since 2001 and we have the support of our parent institution, the university of cape town, to ensure our continued existence and ability to preserve and disseminate data. the integrity of preservation and dissemination of datasets needs to be protected, as dsa guideline 11 states: the data repository ensures the integrity of the digital objects and the metadata (11) at datafirst, all iterations of each dataset are stored on a secure server with password access. checksums are used to ensure the preserved and shared datasets are not altered inadvertently. dsa guidelines on data discovery, access and citation dsa guidelines dealing with data discovery and access are: • the data repository assumes responsibility from the data producers for access and availability of the digital objects (guideline 9) 12 iassist quarterly 2016 / vol 40 no 3 iassist quarterly • the data repository enables the users to discover and use the data and refer to them in a persistent way (guideline 10) • the data consumer complies with access regulations set by the data repository (guideline 14). • the data consumer conforms to and agrees with any codes of conduct that are generally accepted in the relevant sector for the exchange and proper use of knowledge and information (guideline 15) • the data consumer respects the applicable licences of the data repository regarding the use of the data (guideline 16) data accessibility is an important component of data quality. this refers to how easy the data are to find and obtain (statistics canada quality guidelines, 2014). datafirst’s data portal enables researchers to discover and access data. discovery is aided by detailed metadata prepared for each dataset. metadata is created using nesstar publisher, which is free data markup software for the creation of xml compliant metadata. the publisher software uses the data documentation initiative (ddi) metadata standard. public and licensed data can be downloaded from datafirst’s online data portal. the portal has been created using the national data archive (nada) software, an open source data dissemination package developed by the world bank. enabling researchers to use the data involves providing good metadata. it also involves answering user queries about the data. users can contact us through our online support site or facebook page. the support site is used mainly by early-career academics and postgraduates. queries range from requests for help with data portal registration and downloads, to questions related to in-depth analysis of the data. dsa guidelines concerned with codes of conduct and licensing relate to data security standards, as data licenses are a means of protecting data confidentiality. researchers who download data from our site agree to a standard data usage license. this agreement commits them to preserving data confidentiality and citing data sources. citing data in a standard manner assists other researchers to find data sources and assess or extend research based on the data. datafirst provides a recommended citation for each of our datasets, based on the datacite international data citation standard. on our website we also provide researchers with information on how to cite data in their publications. conclusion this paper describes how we built the service at datafirst by complying with data curation standards and by working with depositors to ensure they follow data quality standards. feedback from stakeholders is also essential to offering a good service. all stages of our data curation process have stakeholder communication dimensions built into them. interactions with experts in the international data curation community also assists us with best practice. government, academia and other data services are represented on our board, to provide input to our work. we interact daily with researchers, and this enables us to understand their data needs. finally, reviews like the data seal of approval process enable us to judge the services we offer against international standards and community best practice. future directions (mary vardigan, icpsr) sustainability as mentioned above, the rapid uptake of the data seal of approval around the world attests to the need for a basic, low-threshold certification process for repositories. as increasing numbers of repositories are applying for and being awarded the seal, the dsa community is growing in other ways as well. while the initiative started in the social sciences and humanities, we are now seeing repositories in the natural and physical sciences applying for the dsa, and the geographic spread is expanding also. these are all positive developments, but they raise the issue of sustainability: how do we ensure the future of the dsa initiative so that we can continue this forward momentum and provide more and more repositories with this validation of their trustworthiness? the dsa regulations offer a clear path to a sustainable and community-driven organization through the mechanism of a general assembly (ga). any repository that has acquired the seal is eligible to join the general assembly. membership in the ga naturally carries both rights and responsibilities. each general assembly member commits to conduct a maximum of three peer reviews a year per dsa repository, thereby receiving voting rights in the ga, which elects the dsa board and provides advice to the board when needed. the idea is that any certified repository having earned the seal itself should have enough expertise to participate in reviewing other repositories. this will enlarge and refresh the pool of reviewers and ultimately strengthen the organization. as of this writing, the general assembly had been convened and had elected a new board. we are confident that this governance mechanism will go a long way toward assuring the stability of the dsa initiative. at the same time we are also evaluating different business models and approaching potential funders. while the dsa organization does not have a lot of overhead, it does need a minimal amount of funding to maintain the dsa assessment tool and website and manage and train the pool of peer reviewers. so far these activities have been undertaken by individuals contributing their time, but such in-kind contributions are not sustainable in the long term and must be bolstered by actual funding, even if minimal. common requirements for basic certification another interesting development relating to the future of the dsa is a project begun in 2014 to harmonize the guidelines of the dsa and the certification criteria of the world data system, an interdisciplinary body of the international council for science (icsu) created in 2008. the world data system has built its own set of guidelines for trustworthy repositories, many of which overlap with the dsa criteria. thus, vol 40 no 3 / iassist quarterly 2016 13 iassist quarterly under the auspices of the research data alliance repository certification interest group, an rda working group was created to bring the requirements of the two certification catalogues together. the case statement of the working group also calls for the group to develop common procedures and to create a shared testbed for assessment. the ultimate goal is to create a shared framework for certification that includes other standards as well, such as the nestorseal and iso-16363/trac. goals of this dsa-wds harmonization effort are to: • simplify the array of certification options • show the value of a certification procedure requiring relatively low investment • stimulate more certifications • foster greater trust in repositories • promote data sharing representatives from the dsa and the wds have worked diligently to create harmonized criteria and as of this writing were poised to publish them to the rda community. the harmonization work included constructing two-way mappings between the standards, analyzing the gaps and commonalities, and then crafting new language to bring the standards together. certification procedures were also harmonized. what will this project mean to the dsa going forward? the plan is that in the future both the dsa and the wds will be using the common criteria and there will be greater collaboration and synergies across the organizations. this is still a work in progress, so stay tuned for more information as the project bears fruit in 2016. with all of this ongoing activity, it is clear that many exciting changes lie ahead as the dsa expands into new territory. but even more exciting is the steady accumulation of repositories gaining the seal, coming on board one by one, and together demonstrating the impact and importance of a growing federation of trusted repositories that are protecting valuable digital assets around the world. notes 1. data services librarian, ciser, cornell university. (university of edinburgh) email: stuart.macdonald@ed.ac.uk 2. deputy director, data archiving and networked services (dans). email: ingrid.dillo@dans.knaw.nl 3. research data manager, odum institute for research in social science, university of north carolina. email: slaffer@email.unc.edu 4. data services manager, datafirst, university of cape town. email: lynn.woolfrey@uct.ac.za 5. assistant director, inter-university consortium for political and social research (icpsr). email: vardigan@umich.edu 6. http://stardata.nrf.ac.za/nadicc/presentations/harmsen_henk.ppt 7. http://datasealofapproval.org/en/information/all-documentation/ 8. ads and the data seal of approval – case study for the dcc http://www.dcc.ac.uk/resources/case-studies/ads-dsa 9. finnish social science data archive and the dsa: a case study http://datasealofapproval.org/en/assessment/fsd-dsa-case-study/ 10. http://www.statcan.gc.ca/pub/12-539-x/12-539-x2009001-eng.htm iassist quarterly vol 24 no. 3 iassist quarterly fall 2000 15 abstract the use of numeric and statistical data for macro and micro level decision-making, development planning, and socioeconomic research has always been critical for governments, international organisations and society at large. while the developments in information and communication technologies (icts) have paved way for timely access to validated numeric data on the one hand, these have also posed many challenges before library and information professionals to exploit the opportunities made available by the icts to manage and disseminate the numeric and statistical data efficiently and effectively. in india, numeric data have been published regularly, mainly by the central and state governments. numeric data published by the government ministries, departments, and other agencies is largely print-based, basically brought out in the form of reports, as well as ad-hoc and regular publications. although the technology for digital storage and dissemination of numeric data had been available for a long time, its importance seems to be realized only recently by the government. although the nicnet (a national network for dissemination of the government data and information) has been operational since the 1980s, a comprehensive national policy on dissemination of statistical data (npdsp) was announced only in may 1999 by the government of india (goi). this paper critically evaluates the provisions of this policy and also looks at the infrastructure being made available in the form of nicnet. though a few efforts have been made in india (e.g. by the reserve bank of india, and registrar general office) to digitise the existing numeric and statistical data, access to the digitised data is not adequate and reliable. in fact, the conduit is available, but the content is far from satisfactory. there is a strong need for assessing user needs, enhancing their awareness, and consolidating the efforts of various ministries, departments and other source agencies to make the collection, validation, organisation and dissemination of numeric and statistical data efficient and effective. a beginning has already been made in this direction by making the department of statistics, government of india responsible to serve as a nodal agency in this regard. an effort has been in this paper to raise a few issues and put forward a few suggestions to ameliorate the situation and also to enhance global community’s awareness regarding the state-of-the-art in india. introduction the importance of reliable and accurate numeric and statistical data has been duly recognised by scientists, social scientists, planners, and decision-makers, entrepreneurs, and governments. now-a-days, a lot of time money and manpower is invested in collecting, analysing, and disseminating information in quantitative form. as the number and variety of data sources increase, be it an academic, industrial or government setting, the process of providing access to data gets complicated. the use of traditional methods of data collection are tedious and cumbersome. all sorts of clarifications and explanations regarding the data need to be mentioned to all data source agencies/individuals time and again. dissemination of the data analysed (in some cases unanalysed also) to all concerned poses problems for the library and information professionals. the problems are compounded with the increasing demand of users for numeric data customised according to their needs, or in ready to use formats. users from academic organisations, bureaucracy, government, business, industry, and non-government organisations rely heavily on numeric and statistical data for their work. though the use of computers and storage media on the one hand has solved the problems of storing and analysing large quantities of such data, it has given rise to operational problems on the other. with the convergence of computer and communication technologies and the emergence of networks, information handling processes have undergone a profound change. developments in information and communication technologies (icts), particularly in the 1990s have changed very significantly the way we manage information, including statistical information, right from its generation to use. now it is possible to access numeric and statistical data via the networks, particularly the internet. in the networked environment, access to, validation, security, and updating of statistical data are some of the challenges facing the library and information professionals. accessing indian numeric and statistical data: a critical study of the suprastructure and infrastructure in india by dr. jagtar singh & h. p. s. kalra* 16 iassist quarterly fall 2000 state-of-the-art report in the context of developments in icts at the global level, the situation in india regarding the availability of and access to numeric and statistical data is not so encouraging. numeric data are collected, analysed, validated and disseminated by ministries, departments, and agencies of the central and the state governments. these data are published as reports, ad-hoc and regular publications mainly in the printed form (e.g., publications of central statistical organisation, human development reports for various years produced by the state governments of karnataka, and madhya pradesh, and the statistical abstract of punjab published by the economic and statistical organisation, punjab). the statistics wing (sw) of the ministry of statistics and programme implementation (mospi), government of india, (earlier the department of statistics) is the apex body for official statistical system in india. two organisations under it, namely the national sample survey organisation (nsso) and the central statistical organisation (cso) are responsible for carrying out socioeconomic surveys, field work for surveys, training, dissemination and publication, and coordination of statistical activities. since government agencies are largely responsible for collection and dissemination of numeric data, there is generally a big time lag in the publication of numeric data. appropriate technology was available in india for quite a long time, but has not been used optimally for quick analysis and timely dissemination of such data. other agencies, such as the reserve bank of india (rbi) and the registrar general’s office, are also engaged in providing financial and census information respectively in statistical form. statistical publications by various government departments and agencies are in the broad areas of national income, industry, banking and finance, census, trade, agriculture, labour, and education. in the last few years however, the computer centre of mospi has also been instrumental in making available some of the publications of nsso in magnetic tapes, e.g., the annual survey of industries 1995-96, and the report on energy statistics, 1998-99. as far as the bibliographic control of publications containing statistical information is concerned, there is no single source in printed form, though efforts have been made, e.g., the statistical system in india, 1989; the catalogue of government of india’s civil publications; and the announcements in newspapers entitled ‘list of new arrivals’ by the controller of publications, government of india. national policy on dissemination of statistical data the government of india (goi) seems to have realized only recently the importance of disseminating the numeric information in digital form. in may 1999, the government of india announced a comprehensive national policy on dissemination of statistical data (npdsd) and specific guidelines for release of data. the provisions of the policy are reproduced as annex 1 given at the end of this paper. the policy of the goi is a welcome step in the direction of dissemination of numeric data in digital form, but certain lacunae in the policy are worth discussion. provisions of npdsd have been examined critically below. under clause (i) of the policy, it is written that the data would be available to users in the form of hard copies and magnetic media. as the policy was announced in may 1999, developments in data storage technology at that time were ahead of magnetic media. infrastructure for data storage in optical media (cd-roms, dvds) is also available with government departments and agencies. a comprehensive term to include optical media would have been better. similarly, provisions for availability of validated data via networks, particularly the internet, could also have been incorporated in the clause. clause (v) of the policy will act as a hindrance in quick and timely release of data. under this clause, it is said the data users shall have to wait for three years after the completion of field work to get the data, in case the reports based on survey data work can not be released by the concerned government agencies earlier. moreover, the access mechanism for data collection in such a situation has not been specified. other clauses in the policy such as clause (vi) and clause (viii), deal with non-commercial use of data, and the department of statistics (dos), goi, acting as the nodal agency for dissemination of statistical data, respectively. statistics wing (sw) in the mosip (earlier the dos) has been entrusted with the responsibility of data collection from source agencies; the organisation of data and ensuring its quality; conducting studies regarding data collection and validation for each type of data source; and the dissemination of official statistics under clause (viii) of the policy, and clauses (i), (ii), (iii), and (iv) of the guidelines for release of data. regarding secondary publications, e.g., bibliographies, indexes, and directories in electronic form, provision has been made in the guidelines for release of data (clause viii), but not much information is available on the web site of sw in mospi (http://www.nic.in/stat) though dissemination of statistical data is the focus, mechanisms have been spelt out only for release of data, and not for its dissemination. sw can act as disseminator of statistical information, to one and all only if it has a network of branch offices. it does not have any such network, and under the present provisions of the policy and guidelines, therefore, either the dissemination activity, largely print-based, will be centralised or will have to rely on some other government department/agency. public libraries offer such a network in almost all the states and union territories in the country. traditionally, public iassist quarterly fall 2000 17 libraries have been playing the role of information disseminators. many states in india have working public library systems today. ten out of 28 states have public library legislation. the policy could have incorporated the role of public libraries in collaboration with sw for dissemination of statistical data. national statistical commission the national statistical commission (nsc) was setup by the goi in january 2000 to critically examine the shortcomings and deficiencies of the present statistical system with a view to recommending measures for a systematic revamping of the system. nsc has 12 members under the chairmanship of dr. c. rangarajan, the governor, andhra pradesh. terms of reference of nsc include timeliness, reliability, and adequacy of statistical system in india; dissemination of these statistics; and the coordinating mechanism using statistical information for policy making and planning. although the commission was expected to submit its report to the government within a period of twelve months from the date of its establishment, it has not submitted the report. (the text version of the information on the nsc downloaded from the mospi web site is enclosed as annex 2.) infrastructure for dissemination of digital data the infrastructure for storing and disseminating the data in digital form was set up in 1977 by the government of india under the department of electronics. later, it was entrusted to the direct control of planning commission. recently, recognising the growing importance of icts, a separate ministry of information technology (mit) was created by the goi, with nic and its infrastructure under the mit. nic has a large distributed network infrastructure, known as nicnet with its nodes in all parts of the country, including many remote areas. the conduit, i.e. nicnet, is available, but the availability of and access to the content, i.e. the statistical data, via nicnet is not easy. effort of the nic to provide statistical, numeric and other factual information through general information service terminals of nic (gistnic) is not successful as these are placed in the district commissioners’ offices. even the web site of gistnic does not provide users with much information (http://gist.ap.nic.in). although the npdsd has been announced only in 1999, and the nsc commissioned in 2000, efforts by the concerned government agencies to provide statistical and numeric data in the digital form and via networks started earlier. examples of these are given below. the economic and monetary data and information are available via the reserve bank of india (rbi) web site (http:// www.rbi.org.in). rbi is the central bank of india. in addition to its main web site, it has also created special urls for frequently accessed documents, a list of which appears as annex 3. the documents available via the rbi web sites provide textual as well as numerical and statistical information. the weekly statistical supplement, available via its web site (http://www.wss.rbi.org.in) provides economic information in numeric form under many headings (the text version of a downloaded document is enclosed as annex 4.) census data in the floppy discs is also available at select institutions, but its format is not user-friendly for search purposes. brief information on census is also available from registrar general’s office (rgo) web site (http:// www.censusindia.net). in spite of the excellent web site maintained by the rgo, not much data are available via it. web sites of nearly all ministries and departments of goi and their agencies exist today. a list of these sites has been compiled by varun. a cursory look at the urls of the web sites of various ministries, departments, and agencies reveals that nic has created web sites for many government ministries, departments, and agencies, but adequate information is not available via the government web sites, and is not updated in some cases. in some cases, there is no email address for contacting the concerned department. thus it becomes clear that while, with the help of the elite nic the conduit for information dissemination has become available to data source agencies, the content (data and relevant information) available via the conduit is far from satisfactory. looking towards the future reforms in the telecom sector in india have been very rapid in the last few years, and with the increasing competition, prices of computers, telecommunication equipment, and services have come down heavily. more and more bodies in india are now using the computers and communication facilities and services. these are likely to increase manifold in the coming years. with such a market scenario, need for information (including numeric information) available via the computer and communication facilities, both at workplace and at home would increase considerably. therefore, assessing user needs for statistical and numeric information and enhancing users’ awareness of the existing resources and services through which they can access such information is the need of the hour. the initiatives by the goi, such as the announcement of npdsd in may 1999 and setting up of nsc in january 2000 are in the right direction, but are in reverse order. the government should have set up nsc earlier, as one of the terms of its reference is with regard to collection and dissemination of timely, reliable and adequate statistics. while the focus of npdsd is on centralisation of statistical information system, one of the terms of reference of nsc deals with decentralisation of statistical information system. in the light of report and recommendations of nsc (whenever submitted), the 18 iassist quarterly fall 2000 npdsd will have to be amended or a new policy will have to be announced. consolidation of the efforts of various ministries, departments and other source agencies of the goi is also needed to make the collection, validation, organisation and dissemination of numeric and statistical data efficient and effective. the sw has been entrusted with the job of coordination with other government ministries, departments, and data source agencies, and to act as a nodal agency for dissemination of statistical data (clause v of the guidelines). but nothing concrete has come out even after one and a half years of the adoption of the policy. some guidelines and clauses of the policy act as hindrance in accessing data. in such a situation, the role of library and information professionals would be to convince the government to provide validated quality numeric information to society at large. the goi through the sw can provide quality information by strengthening the existing infrastructure of the nic and ensuring access to the information by making available the nicnet terminals available to people via public libraries. further reading chidambaram, s. siva. (1999) access and availability of statistical information. iaslic bulletin, 44(3), sep., 133141. goswami, p.r. (1998) access to socio-economic data with particular reference to india. desidoc bulletin of information technology, 18(4), jul, 29-38. goswami, p.r. (2000) official statistical information: indian scenario. information today and tomorrow, 19(1) jan-mar, 11-16, 21. india. ministry of statistics and programme implementation. (2000) annual report 1999-2000. [new delhi: mospi] india. national statistical commission. http://www.nic.in/ stat (visited 10-1-2001) kathuria, rajat (2000) telecom policy reforms in india. global business review, 1(2) jul-dec, 301-326. national policy on dissemination of statistical data. (1999) information today and tomorrow, 18(3), jul-sep, 16-17. saha, a. and thulasi, k. (1998) techno-commercial information on the internet. information today and tomorrow, 17(1) jan-mar, 6-18. satish chander (1998) access to legal information in india. desidoc bulletin of information technology, 18(4), jul, 21-28. seshagiri, n. and reddy, c.l.m. (1997) evolution of ethical aspects of digital information in india. international information and library review, 29(2), june, 227-235. varun, v.k. (1998) ru on internet? information today & tomorrow, 17(4) oct-dec, 19-20. annex 1 national policy on dissemination of statistical data i. dissemination of official statistics in the form of reports, ad-hoc and regular publication etc. by the central ministries / departments / agencies as at present shall continue. validated data, though published, including unit/household/establishment level data after deleting their identification particulars to maintain confidentiality should also be made available to the national and international data users in the form of hard copies and on magnetic media on payment basis; ii. no data, which are considered by the concerned official in data source agency to be of sensitive nature and the supply of, which may be prejudicial to the interest, integrity, and security of the nation, should be supplied. the central government, or a state government or the concerned government agency, as the case may be, shall exercise its overriding prerogative to decide the degree of sensitivity of the official statistics produced by it. the data source agency will reserve the right to withhold its release altogether or to release selectively. iii. price of data to be supplied under (i) above should include the cost of stationery, computer consumables and computer time for sorting information. however iv. price may be fixed in indian currency as well as in sterling pound and american dollar. foreign currency prices may be determined using relevant official multiplier fixed from time to time for printed government publications; v. survey results/data should be made available to the data users in india and abroad simultaneously after the expiry of three years from the completion of the field work or after the reports based on survey data are released, whichever is earlier; vi. data users will give an undertaking in the prescribed form to the effect, inter alia, that the official statistics obtained by him for his own declared use will not be passed on with or without profit to any other data user or disseminator of data with or without commercial purpose; iassist quarterly fall 2000 19 vii. data users will have to acknowledge the data sources in their research work based on official statistics. one copy of research study along with short summary of conclusions, if required by the concerned data source agency, should be supplied in the form of hard copy or on electronic media, free of cost; and viii. the department of statistics will be the nodal agency for dissemination of official statistics provided by central government ministries and departments. however the concerned subject matter ministries and departments of the central government will be the final authority on issues arising out of this policy with a view to resolving any dispute between a data user and a data source agency. the guidelines for the release of the data are: i. a data warehouse in the department of statistics will be created to enable the data users and general public to have easy access to the published as well as unpublished validated data from one source. ii. the data warehouse will collect data from various source agencies, integrate the data into logical subject areas, store the data in a manner that is accessible and understandable to non-technical decisionmakers and deliver data/information to decision makers through report writing and query tools. iii. as data source agencies are generating data at various levels, the responsibility of data supplied and receipt will be shared between the respective central ministries/departments/ agencies and the department of statistics by establishing and maintaining close collaboration. iv. for each data type and source, detailed studies will be undertaken by the department of statistics in cooperation with the concerned data source agency on (a) the concepts, definitions, classifications and methods used in data collection and processing including validation, (b) formats of data collection, (c) media on which data will be supplied, (d) frequency of supply of data and (e) procedures and modalities for preservation, updation and dissemination of data. v. the volume of data flowing from each source agency into the data warehouse will be assessed by the department of statistics in order to formulate the various parameters required for designing, establishing and maintaining a data warehouse. vi. each data source agency will be required to adopt for itself a calendar for preparation and release of data data which it will share with the department of statistics. as part of its nodal responsibility of dissemination of data from the source the department of statistics will keep track of the data release calendar of each source agency. vii. the data source agency will be required to supply on computer compatible media, validated data, published or unpublished free of cost to the data warehouse. viii. the department of statistics will prepare directories of all available data in the data warehouse and update the same at frequent intervals. a web site will be created for the data warehouse and directories will be available on the web site. ix. from the data warehouse, data/information will be made available free of cost to the data source agencies for official use and also the approved research institutes and universities for research purposes. x. the price of data to be supplied to other users will depend upon system hardware and software used for data storage, retrieval etc. and also on medium of supply of data. annex 2 national statistical commission the government of india has setup a national statistical commission to critically examine the deficiencies of the present statistical system with a view to recommending measures for a systematic revamping of the system (gazette of india, extraordinary part 1, no.10, resolution no.m-13011/3/99-admn.iv dated 19.01.2000). the commission consists of dr. c.rangarajan, governor, andhra pradesh, as its part-time chairman and the following 11 eminent experts as part-time members: 1. mr. v.r.rao, ex-director, central statistical organisation and un advisor 2. mr. s.m.vidhwans, ex-director (economics & statistics), govt. of maharashtra and un expert 3. prof. j.roy, professor emeritus, indian statistical institute 4. dr. prem narain, emeritus scientist, iari and exdirector, indian agricultural statistics research institute 20 iassist quarterly fall 2000 5. dr. rakesh mohan, director-general, national council of applied economic research (ncaer) 6. dr. v.r.panchmukhi, director-general, research and information system for the non-aligned and other developing countries 7. dr. y.venugopal reddy, deputy governor, reserve bank of india 8. dr. k.srinivasan, executive director, population foundation of india and ex-director of international institute of population studies 9. prof s.tendulkar, delhi school of economics and vice-chairman, n.a.b.s. 10. dr. a.b.l.srivastava, chief consultant, educational consultants india ltd. & exprofessor,national council for educational research and training (ncert) 11. dr. fredie ardeshir mehta, eminent private sector economist and director, m/s tata sons ltd. the terms of reference of the national statistical commission are as follows: 1. to examine critically the deficiencies of the present statistical system in terms of timeliness, reliability and adequacy 2. to recommend measures to correct the deficiencies and revamp the statistical system to generate timely and reliable statistics for the purpose of policy and planning in government at different levels of administrative structure 3. to recommend permanent and effective coordinating mechanism for ensuring integrated development of the decentralised statistical system in the country 4. to review the existing legislation for the collection of statistical information and to recommend amendments where necessary, to achieve the objective of collection and dissemination of timely, reliable and adequate statistics 5. to review the existing organisation of the ministry of statistics and programme implementation (statistics wing) and other statistical units of the government and to make recommendations on their staffing and training requirements to enable them to cope with the increase and development of statistical sources 6. to examine the need for instituting statistical audit of the range of services provided by the government and the local bodies and make suitable recommendations thereof and 7. to recommend any other measures for improving the statistical system in the country. the commission is expected to submit its report to the government within a period of twelve months from the date of its establishment. annex 3 list of special urls for frequently accessed documents of reserve bank of india (rbi). * currency museum http://www.museum.rbi.org.in * exchange control manual http://www.ecm.rbi.org.in * weekly statistical supplement http:// www.wss.rbi.org.in * rbi bulletin http://www.bulletin.rbi.org.in * monetary and credit policy http:// www.cpolicy.rbi.org.in * 9% government of india relief bonds http:// www.goirb.rbi.org.in * rbi notifications http://www.notifics.rbi.org.in * rbi press releases http://www.pr.rbi.org.in * rbi speeches http://www.speeches.rbi.org.in * rbi annual report http://www.annualreport.rbi.org.in * credit information review http://www.cir.rbi.org.in * report on trend and progress of banking in india http://www.bankreport.rbi.org.in * faqs http://www.faqs.rbi.org.in * committee reports http://www.reports.rbi.org.in * y2k http://www.y2k.rbi.org.in * fii list http://www.fiilist.rbi.org.in * electronics clearing service http://www.ecs.rbi.org.in * facilities for nris http://www.nri.rbi.org.in * sdds-national summary data page-india http:// www.nsdp.rbi.org.in * foreign exchange management act, 1999 http:// www.fema.rbi.org.in annex 4 reserve bank of india.weekly statistical supplement dec 02, 2000 1. reserve bank of india 2. foreign exchange reserves 3. scheduled commercial banks business in india 4. interest rates (per cent per annum) 5. accommodation provided by scheduled commercial banks to commercial sector in the form of bank credit and investments in shares/debentures/bonds/commercial paper etc. http://www.museum.rbi.org.in http://www.ecm.rbi.org.in http://www.wss.rbi.org.in http://www.wss.rbi.org.in http://www.bulletin.rbi.org.in http://www.cpolicy.rbi.org.in http://www.cpolicy.rbi.org.in http://www.goirb.rbi.org.in http://www.goirb.rbi.org.in http://www.notifics.rbi.org.in http://www.pr.rbi.org.in http://www.speeches.rbi.org.in http://www.annualreport.rbi.org.in http://www.cir.rbi.org.in http://www.bankreport.rbi.org.in http://www.faqs.rbi.org.in http://www.reports.rbi.org.in http://www.y2k.rbi.org.in http://www.fiilist.rbi.org.in http://www.ecs.rbi.org.in http://www.nri.rbi.org.in http://www.nsdp.rbi.org.in http://www.nsdp.rbi.org.in http://www.fema.rbi.org.in http://www.fema.rbi.org.in iassist quarterly fall 2000 21 6. foreign exchange rates spot and forward premia 7. money stock: components and sources 8. reserve money: components and sources 9. auctions of 14-day government of india treasury bills 10. auctions of 91-daygovernment of india treasury bills 11. auctions of 182-day government of india treasury bills 12. auctions of 364-day government of india treasury bills 13. certificates of deposit issued by scheduled commercial banks 14. commercial paper issued by companies (at face value) 15. index numbers of wholesale prices (base: 1993-94 = 100) 16. bse sensitive index and nse nifty index of ordinary share prices mumbai 17a. average daily turnover in call money market 17b. turnover in government securities market (face value) 17c. turnover in foreign exchange market 17d. weekly traded volume in corporate debt at nse 18. bullion prices (spot) 19. government of india: treasury bills outstanding (face value) 20. government of india: long and medium term borrowings 2000-2001 21. secondary market transactions in government securities (face value) dr. jagtar singh and h. p. s. kalra, jagtar@pbi.ernet.in harry@pbi.ernet.in, department of library & information science, punjabi university, patiala 147 002. (india) {^ sist newsletter vol.1, no. 2 chairperson's report on february 1977, a north american lassist meeting was held in cocoa beach, florida. the primary objective of the meeting was to provide an opportunity for the us and canadian participants to assemble in ags and debate the commonality and differences of the ag mandates. alice robbin of wisconsin presented a most comprehensive outline for a guide to providing social science data services which served as a catalytic influence for the ags as they reviewed their activities in terms of contributions to the guide . establishing technical standards or guidelines would also seem a legitimate contribution of lassist. the workshop format was extremely successful and productive. the results are reported in detail by the us secretariat, judith rowe, who handled the arrangements for the meeting, the canadian secretariat, sharon chappie henry, and the chairpersons of the ag in the body of this issue of the newsletter . in late april or early may 1977, the european ags will convene at a workshop to review their progress and activities. this workshop is being organized by per nielsen, the west european secretariat. during may 1977, the canadian secretariat is organizing a north american lassist meeting to be held in toronto. this meeting will build on the achievements of the ags at the february 1977 meeting. an international lassist meeting was originally planned for august 1977. due to the given schedule of workshops and conflicts in other conference dates during august, this meeting has been slated for february 1978, in chicago. the theme will be "state of the art: perspectives." it will be a three day meeting, with one day devoted to symposia focusing on the ag activities and objectives. there will be three symposia during the morning and four during the afternoon. each symposia will represent one ag and eight papers will be presented at each. topics for papers should be sent to either carolyn geda of the inter-university consortium for political and social research or tony falsetto of the canadian machine-readable archives, public archives of canada [see details in section, "upcoming meetings"]. the remainder of the conference will be devoted to ag workshops. the program committee is composed of tony falsetto, public archives of canada, richard roistacher of the center for advanced computation, university of illinois, sheldon laube of cm. leinwond and carolyn geda of icpsr. a north american program committee has been established for the lassist program to be held in conjunction with the international sociological association in august 1978, in uppsala, sweden. the members of this committee are elliot avedon of university of waterloo, don harrison of national archives, and john devries of carleton university. the proposed scheme is the international interchange of data. details of this conference will be reported in the news letter as they are confirmed. the treasurer of lassist is ed hanis of university of western ontario. lassist accounts have been opened in both us and canadian currencies. membership fees should be remitted in one of these currencies. if this is not possible, contact ed hanis. the treasurer will forv;ard the names and addresses of lassist members to john kolp of s s data and members should be receiving s s data in the very near future. sist newsletter vol.1, no. 2 the newsletter will again be distributed to the full mailing list. subsequent issues will be circulated only to members. a french translation is being done by the canadian secretariat, should anyone be interested in the french version of the newsletter . kenneth brewer of australian national universities has offered to act as the distributor of lassist information throughout australia. given the distances in australia, ag groups will probably not be formed at this time. the distribution of lassist information however is perceived as most desirable. the steering committee list has been reproduced again in this newsletter . please note the following revisions: sharon chappie now is sharon chappie henry; ed hanis has been appointed treasurer; the address of the east european secretariat has changed. secretariat reports canadian secretariat report sharon chappie henry data clearing house for the social sciences the data clearinghouse for the social sciences [dchss] is continuing to publicize lassist activities in the dchss bulletin and other canadian professional journals. the campaign for lassist members is continuing. ten canadians actively participated in the conference of canadian and american action groups in cocoa beach, florida. plans are underway for lassist [canada] to host another working conference to be held in toronto in may [see "upcoming meetings" in newsletter ]. all members of lassist including those in other countries are welcome to attend. further details are available from sharon c. henry, canadian secretary for lassist. west european secretariat report per nielsen danish data archives per nielsen has recommended that action groups organize a series of workshops in their respective areas of interest. paul muller, chairperson of the process-produced data ag, has suggested that a conference devoted to the data law issue would be relevant at the present stage of development within his ag area. a call for papers concerning the data law issue will be made in advance of the conference and workshops. 5. iassisl quarterly 11 the poll database: roper center's online source for public opinion research by linda langschied rutgers university college avenue campus new brunswick, new jersey 08903 communications and mass media research. not only an archival facility, the roper center provides its clients with services such as data analysis and interpretation, and searches of the archive. the thousands of scholars who have made use of the center's data archives and services have tappped a rich source of machine-readable data, which is stored on magnetic tape. in 1980, the roper center began planning a new service, an online retrieval system that would allow researchers, with the use of a computer and modem, to tap directly a database that would contain survey questions and responses. in 1983 this system, poll, was constructed and is now available by subscription to researchers across the u.s. and internationally. the content of poll differs from machine-readable data files in that the poll database contains no raw data. rather, poll approximates a bibliographic database, with each record giving all the information necessary to form a complete citation. searchers retrieve, upon entering a topic, three basic kinds of information: the texts of poll questions related to that topic, the responses to each question, and "study level" information: the dates the poll was conducted, by whom, the type of sample, and so on. the following is an example of a record, which illustrates the various fields of information available: the roper center and poll established in 1946, the roper center contains the largest archives of public opinion research data in the worid. this collection includes the basic data from over 10,000 public opinion surveys conducted since 1936, which have been gathered from over 40 major united states suppliers, and over 70 foreign counuies. the roper center classifies approximately 65% of its materials as public affairs studies. 20% as market and consumer surveys, with the remaining studies focusing primarily on question: r12 the reagan administration has proposed giving $100 million in military, medical and economic aid to the rebels fighting the sandinista government in nicaragua. do you favor or oppose this proposal? responses: favor 33% spring j 987 12 iasstsl quarterly oppose 54% not sure 13% survey organization: nbc/wall street journal population: national adult population siz,e: 1599 interview method telephone beginning date: apr 28, 1986 ending date: apr 29, 1986 source document: nbc news/wall street journal date of source document: may 12, 1986 subject: latin diplomacy presidency flill question id: usnbcwsj.051286.r12 the need for an online system like poll that would provide accurate and inexpensive data to their clientele was very apparent to the roper staff who conducted manual searches of the archives for clients. prior to the start of the online service, the staft created manual subject indexes of the reports, or used indexes that were sometimes supplied by the contributing survey organizations. to conduct a search for a client on a particular topic, in a specific time frame, was very labor intensive: many times this included having to sit down and read the entire report data were then compiled for patrons through a great deal of cutting, pasting, and photocopying. it was not uncommon for the roper staff to have to pull the same study ten times within a year and document the same things ten times. obviously, this degree of labor intensiveness made jobs very costly for the paying patron. for many people, including academics, with limited financial resources, the cost was out of reach. in the online system, the data are simply input once, in a standard format, and are readily available. through poll, roper is meeting its goal to provide easy access to data which are timely and accurate, while holding down costs, both internally and to the clients it serves. scope of the poll database researchers using poll should have a good understanding of the types of questions contained in the database. poll incorporates all surveys with national samples that come into the archives, from most of the major polling organizations, such as gallup, roper, yankelovich, harris, and major media like abc, cbs, nbc, los angeles times, and many more. poll includes onlv national surveys. although the roper center does archive data files from international surveys, such as the u.s. international agency surveys, and state-level studies such as the minnesota poll, these are not added to poll. roper center has, however, begun planning to create a separate database of state-level surveys, though initiation of this database is still several years down the road. there are two types of study which are undertaken by the polling organizations and entered into poll. most organizations which store their survey results at the roper center conduct omnibus surveys—ongoing surveys which measure changes in attitudes over a period of time. in addition, omnibus surveys cover many different types of political data or policy issues in one study. the second type of study is the special, one-time study. typically, these studies focus on a specific, major topic: "as a result of the weapons deal (with iran and the contras in nicaragua) that is now coming to light, do you think the u.s. congress should cut back on american military aid to israel, or do you think congress shouldn't do that?" a single search in poll will find both types of studv. spring 1987 iassisl quarterly 13 who uses poll? the main users of poll are from academia, the media, business, and even polling organizations, who may use poll to design a study, or to compare their results to those of other organizations. academic institutions may also be members of isla (international survey library organization), the cooperative, educational arm of the roper center. isla members enjoy reduced rates for many of the roper center's services, including the use of poll. annual subscription rales for poll vary according to member/non-member status and whether the institution receives the reduced isla rate. the roper center is trying to alert users, especially those in academia, of the value of making poll available to students for classroom use, particularly in disciplines such as sociology, political science, journalism, and any of the increasing number of academic departments that utilize public opinion research as part of student coursework. the center is encouraging institutions to take out a poll account, and give students direct access to the database. at present. roper center is in the process of negotiating a subscription to one institution for the purpose of direct access for its students. to open such a subscription requires a great deal of planning on the part of both roper and the institution, mainly because of the internal bookkeeping required by allowing general access. the plan for building the poll database the staft at the roper center have set entering the most current data as their highest priority, with omnibus surveys taking precedence over special, one-time studies. typically, when an omnibus study comes in, unless it is extraordinarily large, it is entered into the database within two weeks, so that the information is very cunent special studies are entered as quickly as possible, though these surveys are more time-consuming to enter than the omnibus studies. at present, the entry of omnibus studies is almost complete for all studies conducted from 1974 to the present at the same time that current data are being loaded, roper center is retrospectively entering data from older surveys. ultimately, the database will include information going back to 1937. an average of 500 questions a week are entered into the system: over 79,000 questions have been entered thus far. no specific target date for the completion of the retrospective entry project has been set predicting how long it will take to enter all national surveys is difficult, as the number of collections available in the eariier years decreases. still, it will be several years before the retrospective entry of national poll data is completed. how can poll aid research? a researcher may need to use poll only to satisfy a public opinion research question. for example, a researcher interested in the question of whether people think children should be taught sex education courses might find this kind of information through a poll search: spring 1987 14 lassisl quarterly qucsiion: q14g do you think that public elementary schools in this community should or should not teach sex education in grades 4 through 8? (if favor sex education in grade 4 through 8, ask.:) should this program include discussions about aids (acquired immune deficiency syndrome), or not? responses: favor sex education/include discussions of aids 67% favor sex education/don't include discussions of aids 4% oppose sex education 21% no opinion 8% survey organization: gallup organization population: national adult population size: 503 interview method: telephone beginning date: feb 9, 1987 ending date: feb 25. 1987 source document: gallup poll date of source document: mar 22, 1987 subject: sex education health full question id: usgallup.032287.r1 however, if a more detailed analysis than this is needed, e.g. a breakdown of responses by age, race, men vs. women, etc., the machine-readable data files must be tapped. the researcher, having isolated the correct studies through a poll search, must contact the roper center to establish whether the center has the corresponding data sets. these data sets contain the raw data of individual responses to survey questions, and are stored on magnetic tape. poll, in this example, has served as a reference tool, leading the researcher to the data sets which are rich in information on the his topic. roper generally does have the data sets referred to in poll, (though there arc some exceptions, e.g. the harris poll, which is archived at the university of north carolina.) the center receives the data directly from the contributing survey organizations, though the lag time for receipt of the sets varies greatly from one organization to another. some organizations send out six months worth of data sets at a time, while others send their data each time a stud\ is conducted. in addition, academic researchers whose institutions are members of the isla, may find that their institution has purchased the tapes from roper, and that they may access the files on their own campuses. how is poll searched? anyone familiar with library card catalogs or printed periodical indexes understands that they may hunt for books or articles using set "fields" of information—author, title, or subject researchers who have made the transition from print catalogs and indexes to online searching—whether searching a library's online catalog or the databases of a commercial information vendor, such as lockheed information systems' dialog. system development corporation's orbit, or the bibliographic retrieval services (brs)— find the possibilities for searching are greatly broadened through the use of the computer. the computer can enable them to seek out, for example, individual words imbedded in a title or abstract of a book or article, or to limit their findings to a specific year of publicauon. perhaps most important, an online search allows the combination of two or more distinct concepts in a single search (e.g. learning disabilities and college students), thus retrieving very specific search results. spring j i lassisl quarterly 15 poll is, of course, a unique database, and while ihe fields available for searching differ from the usual bibliographic format, the same search concepts and strategies still apply. following the instructions in the user's manual for poll, which is distributed to subscribers, the researcher can execute a search with the knowledge of some basic commands. the following is a very brief explanation of the methodology used in searching poll: 1. the basic command for searching the poll database is the "find" command, which may be abbreviated as "fin". 2. the searcher must indicate what kind of search is to be executed, by choosing from the fields, or "indices" available: subject: topic(s) assigned to question word: actual text of questions and responses organization: organization conducting the study date: beginning date of the study examples of the general form that a search might take are: hnd name of index entry in index hn word reagan hn subject ethics hn organization gallup hn date 12/18/85 the user's manual lists, in its appendices, all subject category codes and definitions, and organization codes. a cautionar}' note: subject categories are very broad, and searchers should not assume thai their topic constitutes an official subject for example, a search done on reagan as "subject" will net z,ero hits, while the same search done m the "word" index will retrieve over 6,000 items. 3. searchers may use the truncation symbol "#" to pick up all forms of a word. a search of the word "librar#" will find items containing the words "library," "librarian." "libraries," or 4. the searcher must wait for an arrow to appear on the screen before entering a search. the prompt, "->" means that the system is ready for searching. a simple secirch might look like this: -» fin word librar# -result: 93 items the above example constitutes the simplest kind of search that might be done in poll the system also supports more sophisticated search strategies, utilizing standard boolean protocols (using and, not, and and not) to combine two or more search terms into a single statement: -»fin word reagan and subject diplomacy -result: 62 items or, items may be "nested" within parentheses to indicate the order in which terms are to be searched. ordinarily, the system searches items in the order in which they were entered, from left to right, and from lop to bottom. nesting terms makes the system search the items within parentheses first, and then combine the result with items outside the parentheses. here is an example of a fairly complicated nesting strategy: -»fin word boycott* and ((south and africa) or aparthied) searches may also be entered step-by-step. the same search might be done like this: -»fin word south and africa -result: 17 items ->x)r apartheid -result: 38 items -»and boycott -result: 9 items spring 1987 16 lassist quarterly these examples just begin lo touch on ihe possible search techniques that may be used to search poll. poll is designed to be flexible, and accurate search results may be obtained through many different avenues. this flexibility makes poll a relatively easy database to access and search. with the aid of the user's manual , and a little practice, researchers using poll can achieve proficiency in a fairly short time. once a search is completed, the user may view the records on the screen by issuing a "type" command. in order obtain hardcopy results of the search, the searcher may want to print the results directly, as the search results scroll by on the screen. in this case, it may be advisable to first "sort" the search results, eg. by date, by using poll's "sequence" command, before entering "type." otherwise the records will be printed out in random order. another possibility is to download the results directly to disk. these results can be edited later with any word-processing software that can handle a standard ascii file. in the event that a search should result in an extremely large number of items, using these techniques may not be practical. the file can instead be printed at the roper center. the searcher issues the command "..pollprt", sending the search result directly to the roper center, where it is printed and then mailed to the searcher. roper center does charge for this service, and rates vary according to the size of the file. doing searches on it at patron request, and charging a fiat fee to cover some of the cost of connect time and telecommunication charges. thus far, we have had two requests from graduate students in political science for poll searches—one on south africa and apartheid, and the other on the attitudes of gays on political elections. the reactions of these two students were favorable to poll as a system, although the student researching the attitudes of gays found little relevant materia! in poll. as the coordinator of online reference services at alexander library, i do much of the database searching, and am the person most familiar with poll. overall, i find il extremely helpful in locating polling results— it is an easy-to-use resource, and it produces results far faster than paging manually through indexes. it is also a relatively inexpensive database, costing $15.00 per hour of connect time for isla members. (in comparison, we often search databases on the dialog system which cost upwards of $100.00 per hour of connect time, plus $1.00 or more for each citation—and there are many databases which cost much more than this.) one minor drawback of poll, financially speaking, is that it must be accessed directly in storrs, connecticut, rather than through a telecommunications network like tymnet or telenet, so telephone charges can be rather high. still, when one considers the alternative of manual searching, poll more than pays for itself. conclusion rutgers university, a participating isla institution, has been a subscriber to poll since september of 1986. rutgers handles its poll activities through the alexander library, the graduate library for social sciences and humanities research. the librarians at alexander ueat poll like anv other database. i would also like to add that one of the chief strengths of poll is the support that roper center gives to poll users. new users are bound to have problems initially, with determining their software communications parameters, with bad phone lines, or with signing on to the system. at alexander library, wc encountered all of these problems, and each time the technical staff at roper analyzed our situation over the telephone, and coached us through to a solution. anyone interested in sprmg 1987 lasstsl quarierly — 17 contaciing ihc roper center with quesuons aboul poll may do so by writing lo marilyn potior, assistant director for user services and administration, roper center for public opinion research, p.o. box 440, storrs, ct 06268, or may call at (203) 486-4440. *a11 background information on poll was obtained directly from marilyn potter and sterling green, staff members at the roper center.n spring 1987 vol252 iassist quarterly fall 2001 17 abstract over the past few years, the european union has fostered a number of initiatives to enhance and implement technologies related to electronic document management and exchange with a view to ensuring that all european citizens have easy access to information in the same way. these initiatives have been launched under different programmes or strategies that were implemented at european level, with the support of the member states. this paper reviews these initiatives and examine some of their achievements. introduction we are living through a historic period of technological change, brought about by development and application of information and communication technologies (icts). this process is both different from, and faster than, anything we have seen before. it has a huge potential for wealth creation, higher standards of living and better services. icts are already an integral part of our daily life, providing us with useful tools and services in our homes, at our workplaces, everywhere. the information society is not a society far away in the future, but a reality in daily life. it is adding a new dimension to society as we know it, a dimension of growing importance. the production of goods as well as services is becoming more and more knowledge based. however, the speed of introduction of icts varies between countries, regions, sectors, industries and enterprises. the benefits, in the form of prosperity, and the costs, in the form of burden of change, are unevenly distributed between different parts of the european union (eu) and between citizens. understandably, people are worried and demand answers to questions about the impact of icts. their concerns can be summarised in two main questions: • the first has to do with employment. will these technologies not destroy more jobs than they create? will people be able to adapt to the changes in the way we work? • the second question has to do with democracy and equality. will the complexity and the cost of the new technologies not widen the gaps between industrialised and less developed areas, between the young and the old, between those in the know and those who are not? political framework to meet these concerns we need public policies to help us reap the benefits of technological progress, and which can ensure equitable access to the information society and a fair distribution of the potential for prosperity. the european commission suggests that public policies should, among other things: • improve democracy and social justice by ensuring that the potential of icts to provide relevant, up-todate, information on matters of common interest and to enable citizens to participate in public decision making, are fully supported by governments, with the involvement of non-governmental organisations. • reduce bureaucracy and improve the quality and efficiency of public administration at national, regional and local level, and improve the overall benefits of welfare state services, such as health care and education, through efficiency improvements and through the better matching of provisions and individual needs. in order to meet these and the other objectives, a number of political initiatives have been launched by the european union. as early as july 1996, the commission adopted the green paper living and working in the information society: people first1 on the key social challenges raised by the transition to the information society. the green paper points out that although the adoption and widespread use of information and communication technologies offer a huge the european union initiatives in the digital archive area: achieving the information society by concha fernàndez de la puente* 18 iassist quarterly fall 2001 potential for wealth creation and higher standards of living, many people are also concerned about the impact of the information society on their lives. the green paper examines how information and communication technologies (icts) are reshaping production and work organisation and are transforming peoples lives. in january 1999, the commission published a green paper on public sector information in the information society2. the main message of the green paper was that the use of public sector information in europe for the sake of citizens and enterprises had to be improved. on-line access to public information can decrease the gap between citizens and administrations and support the democratic process. better exploitation possibilities of public sector information will increase the competitiveness of european firms active in the content industries. at the same time, access to high quality government information can help to improve the competitiveness of european companies in general. this green paper attracted a number of responses which have led to more focused priorities in the form of a draft communication which is currently under discussion. improving citizens access to information was also a priority for the lisbon summit that took place on 23 and 24 march 2000. during this summit, the european commission presented the initiative launched in december 1999 entitled “eeurope: an information society for all”3, which proposes ambitious targets to bring the benefits of the information society within reach of all europeans. the initiative focuses on ten priority areas, from education to transport and from healthcare to the disabled. the key objectives of the eeurope initiative are: • bringing every citizen, home and school, every business and administration, online and into the digital age. • creating a digitally literate europe, supported by an entrepreneurial culture ready to finance and develop new ideas. • ensuring that the whole process is socially inclusive, builds consumer trust and strengthens social cohesion. to achieve these objectives, the commission has proposed 10 priority areas for action with ambitious targets to be achieved through joint action by the commission, the member states, industry and the citizens of europe. these areas of action are: • european youth into the digital age: bring internet and multimedia tools to schools and adapt education to the digital age. • cheaper internet access: increase competition to reduce prices and boost consumer choice. • accelerating e-commerce: speed up implementation of the legal framework and expand use of eprocurement. • fast internet for researchers and students: ensure high speed access to internet thereby facilitating cooperative learning and working. • smart cards for electronic access: facilitate the establishment of european-wide infrastructure to maximise uptake. • risk capital for high-tech smes: develop innovative approaches to maximise the availability of risk capital for high-tech smes. • “eparticipation” for the disabled: ensure that the development of the information society takes full account of the needs of disabled people. • healthcare online: maximise the use of networking and smart technologies for health monitoring, information access and healthcare. • intelligent transport: safer, more efficient transport through the use of digital technologies. • government online: ensure that citizens have easy access to government information, services and decision-making procedures on-line. digitisation is an essential first step to generating digital content that would underpin a fully digital europe. it is a vital activity in preserving europe’s collective cultural heritage, providing improved access for the citizen to that heritage, to enhancing education and tourism, and to the development of econtent industries. the critical role that it plays was recognised in the eeurope 2002 action plan endorsed by the eu member sates at the feira european council in june 2000. representatives and experts from the eu member states gathered in lund (sweden) on the 4th of april 2001 to identify ways in which ‘a co-ordination mechanism for digitisation programmes across the member states’ could be put in place to stimulate european content in global networks (objective 3(d) of the eeurope action plan). the meeting, which was arranged by the cultural heritage applications unit of the european commission’s directorate general information society and hosted under the auspices of the swedish presidency, began by endorsing the findings of an earlier meeting of eu experts (luxembourg, november 2000). the lund meeting agreed that digitisation provided a key mechanism to exploit europe’s unique heritage and to iassist quarterly fall 2001 19 support cultural diversity, education and the generation of content industries. although the member states were investing in enabling access to their cultural heritage there were still many obstacles to the near and longer term success of these initiatives. these hurdles included the diversity of approaches to digitisation, the risks associated with the use of inappropriate technologies and inadequate standards, the challenges posed long term preservation and access to digital objects, lack of consistency in approaches to intellectual property rights (ipr), and the lack of synergy between cultural and new technology programmes. the lund meeting concluded that these obstacles and the objectives of the eeurope action plan could be enabled if the member states were to establish an ongoing forum for co-ordination, support the developing of a european view on policies and programmes, develop mechanisms to promote good practice and consistency of practice and skills development, and work in a collaborative manner to make visible and accessible the digitised cultural and scientific heritage of europe. at the same time the participants in the meeting agreed that the european commission could help achieve the eeurope objectives by supporting co-ordination activities, enabling the creation of centres of competence, fostering the development of benchmarking standards for digitisation practices, encouraging a framework that would enable a shared vision of european content, and assisting member states to improve access and awareness for citizens through enhancing the quality and usability of content and the development of models to enable eculture enterprises. by working together to bring these principles into action the member states and european commission aim to ensure that the richness of our heritage will be made visible and usable both by citizens for learning, understanding and enjoyment and by european enterprise to enable industries for the new generation. these political initiatives are supported and endorsed by other activities such as the research programmes and other specific actions. rtd programmes support for european r&d in the area of digital document related technology has moved forward in the several phases and has been done within different programmes. the first actions in this area date from early esprit4 initiatives and have been expanded through its successor programmes, and through the advanced networking programmes, race and acts. many activities have also been developed under the telematics applications programme5 involving a number of different sectors including libraries6, information engineering7 and administrations8. they reveal an emerging interest in the creation and management of new digitised content, with increasing multimedia component, and in a more commercial or economic context. there have been a number of fundamental shifts in the focus of r&d funding in this area of document technology through the successive framework programmes, namely: • from highly technical, industry-targeted developments to the emerging focus on service and on access; • the emergence of more complex distributed models, linking different types of information objects in order to deliver new types of service; • emerging models for networked access. in summary, the trend has moved from the development of basic technologies and tools, which found test-application in document management centres, through to a greater emphasis on downstream applications and on more service and market oriented developments. one of the actions that has had a greater impact in the area of document management technology was the telematics for libraries programme9 that run from 1990 to 1998 and funded over 100 projects and support actions. the work carried out by this programme developed interfaces, systems and services for new digital collections. during the programme, networks of library actors worked together and with publishers and ict industry players in value chains for service provision. the main technological clusters of the libraries programme built on developing interoperable access to distributed services and collections. the services include document access and delivery both of electronic documents and via interlibrary loan; large-scale interoperable catalogues, accessible across borders; directory services; integrating access to archive, library and museum resources. they are built mainly around implementation of the z39.50 protocol and more recently on xml encoding for documents and data. overall the projects provide experience in: handling multilingual access; metadata requirements from different sources and for different types of material; managing the distributed environment. there was also an important cluster of projects based on developing image collections (e.g. fine arts, early printed books) which tested both the technologies for digitising library materials and the systems for digital image storage, management and access. linked to this cluster we find groups of projects which concentrate mainly on: the creation of local document stores either from bought-in or 20 iassist quarterly fall 2001 locally created materials; testing the technologies for storage and access, including issues relating to long-term provision; and developing managed digital collections and services, including access to copyrighted materials. the european commission’s fifth framework programme10 (fp5) takes eu research into the new millennium. it recognises the importance of content and of citizens’ access to knowledge and culture. the challenges are to make content accessible in mixed and multiple media formats and in real and virtual forms, to maintain and preserve information resources, and to strengthen alliances for content creation and learning provision. the information society technologies (ist)11 programme is the largest of the fp5 thematic programmes. the main focus of ist is on enhancing the user-friendliness of the information society: improving the accessibility, relevance and quality of public-services especially for the disabled and elderly; empowering citizens as employees, entrepreneurs and customers; facilitating creativity and access to learning; helping to develop a multi-lingual and multicultural information society; ensuring universally available access and the intuitiveness of next-generation interfaces; and encouraging design for all. the ist programme brings together and extends the acts, esprit and telematics applications programmes to provide a single and integrated programme that reflects the convergence of information processing, communications and media technologies. digital heritage and cultural content12 is one of the five main areas under multimedia content and tools key action 3 of the ist programme. work is aimed at expanding the contribution of libraries, museums and archives to the emerging culture economy, including economic, scientific and technological development. in this area, there are three research priorities: • ensuring integrated access to collections and materials held in libraries, museums and archives. • improving the operational efficiency of large-scale content holdings by means of powerful interfacing and management techniques • preserving and accessing multimedia content of various types, including electronic materials and surrogates of physical objects the key participants in the cultural heritage projects are europe’s memory institutions, both public and private, with a particular focus on new alliances with technical and content-related partners. work is building on achievements under the fourth framework programme addressing libraries, archives, museums and related institutions and attempts to encourage convergence in technical approaches and applications for the various cultural institutions and networked services. the projects and support measures13 resulted from the six calls for proposals launched by the digital heritage area from 1999 to 2001 concentrate on the development of technologies that support the next generation digital library applications, on strategies for the long term development and maintenance of repositories and archives of valuable digital objects, on exploring and experimenting with novel ways of creating, manipulating, managing and presenting new classes of intelligent, dynamically adaptive and selfaware digital cultural objects, either held by memory institutions or directly involving digitally born objects or art forms, and on creating a living record of the information society. the projects selected after these calls (some 50 rtd projects and 20 support actions) address in particular collaborative e-print archive environments, heterogeneous multimedia digital libraries and cultural service infrastructures for trading cultural assets based upon different architectures and business models. through a co-operation agreement with the us national science foundation (nsf) on digital library research some of the selected projects will also have formal relationships with us partners14. with a view to obtaining comments and suggestions on the possibility of developing new topics for r&d in the area of cultural heritage applications, some brainstorming meetings have been organised. one of the meetings discussed the topic “creating a living on-line record of europe’s cultural diversity”15 (luxembourg, 14 march 2000) and brought together experts from across europe with experience in the provision of new services in and around the library, museum and archive institutions, particularly at local or regional level. the meeting concluded with the ideas that in building a strategic approach likely to have lasting impact, particular reference would have to be made to: • the need to base an expanding record of europe’s information society on the valuable building blocks which were now beginning to emerge; • concentrating on a grass-roots approach starting with the needs of local communities and their citizens which had not been sufficiently appreciated elsewhere in the programme; • the importance of scalability which could emerge from local initiatives designed in the interests of replicability, e.g. scalable digital archival technologies; • the pressing need to ensure the involvement of citizens of all walks of life in order to overcome the iassist quarterly fall 2001 21 dangers of social exclusion which often seemed to result from the introduction of new it services; • the difficulty in assessing user needs in an area which was just beginning to emerge as a topic for action; • the opportunities at local level for drawing on the assets of all the various cultural institutions in creating a better appreciation of local heritage and identity and in encouraging the involvement of a wide cross-section of citizens whether for purposes of leisure, education or personal expression. in general, what is emerging as a focus for the future is to help create a european cultural information landscape by encouraging cultural memory organisations to participate in r&d actions providing innovative prototype networked services for both professional users and citizens. this future information landscape should be easy to identify, easy to access, and easy to navigate and should be extended to also encompass europe’s scientific and industrial heritage. equally tomorrow’s cultural content will be produced by generating new forms of digital media. what this cultural content will be, and how it will be created, managed, distributed and preserved remains uncertain and a fertile ground for future research and experimentation. some of the ec archive projects: amicitia asset management integration of cultural heritage in the interexchange between archivesstephan.schneider@tecmath.com brava broadcast restoration of archives through video analysishttp://www.ina.fr/recherche/brava/ index.en.html collatecollaboratory for annotation, indexing and retrieval of digitized historical archive material http:// dbs.cordis.lu/fep-cgi/ srchidadb?action=d&caller=proj_ist&qm_ep_rcn_a=53037 covax contemporary culture virtual archive in xml http://www.covax.org/ cyclades an open collaborative virtual archive environment http://galileo.iei.pi.cnr.it/cyclades/ dhm digital historical maps www.dhm.lm.se echo european chronicles on-linehttp://pcerato2.iei.pi.cnr.it/echo/ euan the european archive networkhttp:// 158.169.50.95:10080/info2000/en/factsheets/euan.html eva european visual archivehttp://192.87.107.12/eva/ search.asp laurin libraries and archives collecting newspaper clippings unified for their integration into networkshttp://laurin.uibk.ac.at/ leaf linking and exploring authority filesjutta.weber@sbb.spk-berlin.de malvine manuscripts and letters via networks in europehttp://www.malvine.org/ nedlib networked european deposit libraryhttp://www.kb.nl/coop/nedlib/ presto preservation technology for european broadcast archiveshttp://presto.joanneum.ac.at/index.html ida programme apart from the research programmes, the commission has launched some other initiatives in the area of electronic record management. one of them is ida16, the european union programme for the interchange of data between administrations that was adopted in november 1995. its specialists offer advice and access to the results of existing telematics projects, to help the public administrations build information links with their counterparts across europe. ida is a facilitator not a regulator. it offers advice and information to public administrations in a number of areas: • implementing an initial set of trans-european telematics networks in a variety of sectors. • facilitating the establishment of a common european telematics service for administrations. • advancing the development of a legal framework and guidelines for the exchange of electronic data. • documenting successful project results that can be transferred for use in other administrations. • offering guidelines for migration from paper-based to electronic administrative procedures. as a result of a constructive consensus building under the austrian and german eu presidencies, the council and european parliament adopted on 12 july 1999 the second phase of the ida programme17, with the objective of improving interoperability of networks and developing trans-european telematics services in priority areas. lately, ida has produced a model requirements for the management of electronic records (moreq specification), which focuses on the functional requirements for the 22 iassist quarterly fall 2001 management of electronic records. ida organised a symposium on open source software (oss) in brussels on 22 february 2001 to meet the growing interest in the use of oss in eu public administrations. this event brought together representatives of the european commission, national and local governments, and the information technology (it) industry and provided a platform where europe’s administrations could share their experiences. in addition, it permitted dialogue with the private sector on the benefits and pitfalls of oss in the public sector. the interchange of data between administration is high on the priority list of the commission. the ida programme facilitates the exchange of information between member states at trans-european level, with administrations, citizens and enterprises as beneficiaries. in this sense, ida will support eeurope’s government online priority area with a series of actions at european level such as portals to european level information and benchmarking and spread of best practice. dlm forum the other important initiative is the dlm-forum18, that was created as “a multidisciplinary forum … in the framework of the community on the problems of the management, storage, conservation and retrieval of machine-readable data, inviting public administrations and national archives services, as well as representatives of industry and of research, to take part in the forum”19. the european commission established the forum acting on the conclusions of a report entitled archives in the european union20, compiled by a group of high level european experts and currently under revision. the archives were the starting point and the driving force behind this european dlm-initiative, which from the very first moment brought together people from different disciplines involved in electronic information handling. while the main goals of the first dlm-forum that took place in 1996 were to find out what was going on in europe and the rest of the world in terms of electronic document and records management and to seek wider co-operation in this area between the member states of the european union (eu) and the european commission, the second dlmforum held in 199921 was clearly targeted at the ict industry. the second dlm-forum issued the following conclusions on three main areas: 1. development of a reference model for managing electronic documents and records in public administration; special dlm-message to ict-industry; 2. realisation of a modular european training programme for administrators and archivists on electronic documents and records management (eterm); improvement of skills and recruitment facilities in europe; 3. implementation of a reinforced dlm action plan, 1999-2004: access for the european citizen and funding priority activities. one of the important result of the dlm-forum ’99 was the elaboration of a “dlm-message” to the ict-industry22 to promote best practices in public administration and provide easily applicable and cost-effective records management and digital archival solutions23. later, the european commission together with the member states updated and forwarded this dlmmessage to the ict industry in which industry across europe was encouraged to exploit the field of electronic documents, records management and digital archiving as a new and viable market. the following challenges for the industry were identified: • a quicker and more thoughtful reaction to constant user demands and new groups of users is required. • solutions must be developed which on the one hand are capable of adapting to changing it evolution but on the other hand can guarantee long-term accessibility and intelligent retrieval of knowledge stored in document management and archive systems. • providers must clearly declare their support for standards and interoperability to ensure general use and distribution of information. • products must become more economical and simpler to use, integrate and operate. at the same time, the following challenges for the european commission were identified: • defining the concrete demands on electronic archiving and document management systems, • establishing documents generated by data processing systems with digital signatures as equal to paper documents with original signatures, • uniform regulations on signatures which also take developments in the us software industry into account, • uniform regulations on the legal recognition of electronically archived documents, • harmonisation of the various european commission iassist quarterly fall 2001 23 initiatives which directly or indirectly concern document management or electronic archiving. long discussions emerged within the ict industry with regard to the dlm-message. its leaders recognised that the public sector is one of the largest and most important vertical market for document related technologies and that today many proven and practical solutions are already in place in european administrations with the effective management of a vast amount of information. the problem is that many of these solutions have been developed individually and there are no standard software packages available for the specific needs of government and administration authorities. further, the benefits of such proven applications now need to be disseminated to a broader user base within the eu as there are many similar requirements and applications in each of the member states. the industry agreed that solutions must be developed that are, on the one hand capable of adapting to rapid it technological advancements, and on the other can guarantee shortand long-term accessibility and intelligent retrieval of knowledge stored in document management and archival systems. this is recognised by the ict industry as a critical factor in preserving the “memory of the information society“ within the eu. the ict industry also agreed that solutions also need to be cost effective and easy to implement using standard, compatible and widely accepted software and hardware platforms and must address security and privacy issues. the ict industry accepted the challenge given to it by the dlm forum and has declared that is prepared and willing to support the efforts of the european union for the preservation and public access to archives and records in a variety of practical ways. the dlm monitoring committee and its special working parties are currently planing the 3rd dlm-ict forum 2002 to take place during the forthcoming spanish eu presidency (1st half of 2002). this will provide an interdisciplinary european platform for the dlm group and the ict industry to jointly present best practices and concrete solutions and to promote, with the support of the european commission’s dg information society, the european network on electronic archives. other actions the info2000 programme (1996-1999) aimed at stimulating the emerging multimedia content industry to recognise and exploit new business opportunities. the central theme was the development of a european information content industry capable of competing on a global scale, and able to satisfy the needs of europe’s enterprises and citizens for information content, leading, on the one hand, to economic growth, competitiveness and employment, and, on the other hand, to individual, professional, social, and cultural development. the european commission launched several preparatory actions for a joint multi-annual follow-on activity to the info2000 programme and to the mlis programme (multilingual information society). as a result, econtent, a “multiannual community programme to stimulate the development and use of european digital content on the global networks and to promote the linguistic diversity in the information society”, was adopted by the council on 22 december 2000 for a period covering 2001 to 2005. the programme is aimed at supporting the production, dissemination and use of european digital content and to promote linguistic diversity on the global networks and it is based on three main strands of action where eu added value can be maximised: 1. improving access to and expanding use of public sector information. 2. enhancing content production in a multilingual and multicultural environment. 3. increasing dynamism of the digital content market. in particular, the econtent programme, as part of the eeurope action plan, contributes to its third objective: “to stimulate the use of internet”. as for the digital archives area, the most relevant action line addressed by the programme is “improving access to and expanding use of public sector information”. to tackle these issues, econtent will stimulate different types of activities, such as experiments in concrete projects showcasing the use of public sector information (psi) to make added-value services and products and the establishment of european digital data collections. at the same time, the activities addressing the policy dimension that have led under the info2000 programme to the publication of the green paper on public sector information in the information society will be pursued. the other interesting programme is ten-telecom26 (trans-european telecommunications networks) which is part of the ten (trans-european networks) community action, to support the trans-european deployment of esociety applications and services based on global telecommunications networks in areas of high socio-economic value. down-stream of the research results, ten-telecom helps project consortia to bridge the gap between technical mature developments and global market operation. it encourages essentially the validation phase for market feasibility of new telecommunications-based applications and generics services, and their early investment phase. 24 iassist quarterly fall 2001 the sectors of common interest identified in the tentelecom work programme include access to europe’s cultural heritage, applications for smes and city and regional information. conclusions the european social model is built both on competition between enterprises and solidarity between citizens and member states. the european information society must draw strongly from this economic, social and cultural strength, linking technological, economic and social aspects together in the creation of new opportunities for all its citizens. the european union fully recognises the importance of ensuring that citizens have access to information resources, particularly in the context of the information society. in the library sphere, this priority was endorsed by the european parliament in october 1998 when it adopted a report from the culture committee on the role of libraries in the modern world27. since that time, the commission adopted a green paper on public sector information. moreover, the eeurope initiative stresses the need for making government more open and more citizen-friendly by introducing support for on-line government. in the specific area of cultural heritage applications, a number of projects have contributed in their own way to these goals. and we hope to enhance citizens’ participation more directly by the action line on heritage for all. the european-wide and interdisciplinary co-operation in the field of electronic document and records management, firmly supported by the dlm forum, is to be fostered and enhanced. the important requirement for standardisation concerning the short-term and long-term preservation and accessibility of electronic information can only be met by using the expertise of all related professional groups, including industry. the information society represents the most fundamental change in our time, with enormous opportunities for society as a whole, but with risks for individuals and regions. the way we develop it must reflect the ideas and values which have shaped the european union. these ideas and values should be transparent and coherent with social justice in order to win the support of citizens. to this end, all interested parties should reflect on the possibility of formulating a set of common community principles for the development of the european information society. footnotes 1 http://europa.eu.int/comm/employment_social/soc-dial/ info_soc/green/green_en.pdf 2 http://www.europa.eu.int/ispo/docs/policy/docs/ com(98)585/index.html 3 http://www.europa.eu.int/comm/information_society/ eeurope/index_en.htm 4 http://www.cordis.lu/esprit/home.html 5 http://158.169.50.95:10080/telematics/ 6 http://www.cordis.lu/libraries/ 7 http://158.169.50.95:10080/ie/ 8 http://158.169.50.95:10080/telematics/admin/ administrations.html 9 http://www.cordis.lu/libraries/ 10 http://www.cordis.lu/fp5/home.html 11 http://www.cordis.lu/ist/ 12 http://www.cordis.lu/ist/ka3/digicult/home.html 13 http://www.cordis.lu/ist/ka3/digicult/en/projects.html 14 ftp://ftp.cordis.lu/pub/ist/docs/digicult/eu-nsf-call.pdf 15 http://www.cordis.lu/ist/ka3/digicult/en/livingrecord.html 16 http://europa.eu.int/ispo/ida/ 17 http://www.europa.eu.int/ispo/ida/ida2/guidelines/en.pdf 18 http://europa.eu.int/ispo/dlm/ 19 htpp://europa.eu.int/eur-lex/en/lif/dat/1994/ en_394y0823_03.html 20 archives in the european union: report of the group of experts on the coordination of archives. luxembourg: office for official publications of the european communities, 1994. isbn 92-826-8233-1. 21 http://www.europa.eu.int/ispo/dlm/dlm99/index.htm 22 http://www.europa.eu.int/ispo/dlm/dlm99/index.htm 23 http://www.dlmforum.eu.org 24 http://158.169.50.95:10080/info2000/infohome.html 25 http://www.cordis.lu/econtent/ 26 http://www.ten-telecom.org 27 http://www.cordis.lu/libraries/en/reportrole.html * paper presented at the iassist/ifdo conference 2001 in amsterdam. concha fernàndez de la puente, european commission, dg information society, cultural heritage applications. rue alcide de gasperi, l-2920 luxembourg. tel: +352-4301.34071 fax: +352-4301.33530. email: concha.fpuente@cec.eu.int http://www.cordis.lu/ist/ka3/ digicult/home.html 20 iassist quarterly 2016 / vol 40 no 3 iassist quarterly abstract the paper suggests effective strategies for collecting high quality data in developing countries based on lessons learned from implementing a household level census of three villages in bangladesh. in particular, we focus on low cost but effective techniques for reducing survey fraud (e.g. curbstoning) and human error (e.g. transcription errors) in conducting face-toface questionnaire-based interviews by hired surveyors. we find the following strategies to greatly improve data quality: use of a geographic information system (gis) and audio-capturing smart pens; daily monitoring; surveyor retraining; and swift firing of those showing consistent errors in judgment. transcribing the data as soon as the surveys were completed helped locate and contain human errors, as well as fraudulent activities. the techniques suggested here are geared towards prevention of errors, rather than detecting fraud during post-survey validation. keywords curbstoning, survey fraud, low budget survey, developing country, data quality introduction “[i]n spite of early evidence of cheating in market research and other fields where professional surveyors are employed ..., the problem of surveyor cheating has largely been ignored in recent literature.” – harrison and krauss (2002) the process of mitigating data falsification is more difficult than it may seem at first. while the challenges of collecting reliable interview-based data are well known (biemer and stokes, 1989; crespi, 1945), the literature on dealing with these challenges, particularly with regard to falsified data (harrison and krauss, 2002), is sparse. this lack of attention began to change in 2003 with the publication of the best practices list by aapor (2003). it succinctly describes the reality that “[e]ffective control of falsification is not the result of any single method, but of the combined aspects of the study-specific environment in which surveyors conduct their work.” in this paper, we briefly outline lessons learned during mitigating survey fraud and human error: lessons learned from a low budget village census in bangladesh1 by muhammad f. bhuiyan2 and paula lackie the implementation of a census of three contiguous villages in bangladesh (bhuiyan and szulga, 2013), denoted the tangail survey (ts). we focus on the strategies employed during the data collection process to mitigate human errors and data falsification. the research project had minimal funding so budgetary restrictions dictated many of the methods employed. roughly a dozen local graduate students were hired for the ts, which used a geographic information system (gis) and smart pens to conduct face-to-face interviews (approximately 40 minutes each) at the household level. our top priority was to gather a highly reliable and robust dataset on the villagers’ subjective well-being (swb) and their perceptions of relative economic position, by identifying households with international migrants, and collecting the geographic (latitude/ longitude) coordinates of household locations with a global positioning system (gps). in agreement with the techniques suggested by koczela et al. (2015), we provide evidence that the use of smart pens, a gis mapping of the household location prior to conducting the interviews, and daily on-site monitoring of the hired surveyors, to be quite effective in catching survey fraud and reducing unintentional errors. the use of the mentioned technologies also made postsurvey validation and catching transcription errors a relatively easy task. curbstoning, falsified data, and cheating the most common terms used to describe survey fraud are curbstoning (where face-to-face interview data are faked), partial falsification (where only a portion of the survey data are faked), and cheating (when the convenience of the surveyor takes precedence over the protocols of the survey). blasius and friedrichs (2012) provide a concise summary of the literature and describe faked interviews. they conclude that it is remarkably easy for surveyors to fabricate interviews in face-to-face surveys which may remain undetected when basic monitoring protocols are not followed. basic protocols, while essential to achieving high quality data, do not necessarily guarantee this quality of data: vol 40 no 3 / iassist quarterly 2016 21 iassist quarterly [while] there has been an old and long discussion on the reasons why surveyors fake interviews (cf. crespi, 1945, 1946; bennet, 1948; nelson and kiecker, 1996), the most elementary reason has hardly been discussed. falsifiers save a lot of time/earn more money if they (partly) fake their interviews. ... [and] the risk of getting detected is relatively low since control mechanisms as those proposed by aapor (2003) and murphy et al. (2004) are relatively easy to bypass. ... furthermore, a detailed introduction, good payment, no time pressure, and an interesting study do not necessarily guarantee well-done interviews. harrison and krauss (2002), waller (2012) and koczela et al. (2015) focus on the motivations for surveyors to cheat. waller (2012) provides the most comprehensive study of the motivations for and methods of falsifying data by the surveyors. in general, the literature focuses mostly on methods of detecting falsified survey data in pre-existing data (bredl et al., 2011, 2012; kuriakose and robbins, 2015) and less so on the on-site prevention of the collection of falsified data. we contribute to the latter. the tangail survey process the accuracy of survey data may be compromised due to a variety of reasons, ranging from human errors or misaligned incentives to technical problems before, during, and after the data gathering process. these errors are further intensified when the survey site is in a developing economy and the project has tight budgetary constraints. both are true for the ts. to understand how the ts methodology was driven by the overarching goal of collecting robust accurate data, we provide a brief overview of the research agenda and the project workflow. research objectives the choice of topics covered in the ts is primarily determined by the principal investigators’ (pis) research interest relative income, subjective well-being (swb) and international migration. while the literature on relative consumption is vast, there is very little empirical work looking into the economic position of individuals relative to local reference groups. for instance, when talking about income relative to neighbors, most papers operationalize the definition of neighbors as all individuals who live within broad geographical regions such as villages (fafchamps and shilpi, 2008), areas within the same zip code (knies et al., 2007), primary census units (luttmer, 2005), states (blanchflower and oswald, 2004) or even countries (easterlin, 1995). the pis decided to collect data fine enough to be connected to more realistic definitions of local reference groups such as neighbors. the ts gathers data on both objective and perceptive measures of relative economic position, as the literature does not provide a strong preference for either measure (fafchamps and shilpi, 2008; ferrer-i-carbonell, 2005; luttmer, 2005; mayraz et al., 2009; mcbride, 2001; senik, 2009). although asking respondents about their perception of relative income compared to neighbors, siblings, colleagues, etc., is somewhat straightforward, data of this type are largely missing for developing countries. hence, we decided to ask respondents about their perceptions of relative economic position directly. in terms of objective measures, we recognized that having geographic coordinates for every household in each village along with objective measures of their income, would be the best possible type of objective data on relative economic position that we could hope for. one of the pis research interests is understanding how local networks affect the choice of destination when it comes to international migration. for instance, if a potential migrant’s neighbor is already an international migrant in the middle east, is it more likely that they will put more weight on that region when choosing between multiple destinations. no data are available that can adequately address this question. after preliminary visits to bangladesh and in consultation with local experts, villagers, government agencies, and development workers, the pis chose three contiguous villages in the tangail district of bangladesh. these villages have a significant number of households with at least one international migrant and they are going to different regions of the world. the census nature of the survey along with the geographic coordinates of the household location, makes it possible to more closely study local neighborhood effects on migration decisions and subjective well-being. from a policy standpoint, the ts offers valuable insights into the interaction of local neighborhood/community effects with international migration, conspicuous consumption and quality of life measures in rural bangladesh. the paucity of data of this type, especially in the context of developing countries, makes this a very useful dataset for those interested in the aforementioned topics. project workflow in the pre-fieldwork and preparation phases, the pis enhanced their local social network and improved their rapport with the hired surveyors (figure 1). once the training and inter-coder reliability development stage commenced, the personal ties strengthened and surveyor specific figure 1: overview of data and project workflow 22 iassist quarterly 2016 / vol 40 no 3 iassist quarterly weaknesses with the delivery of certain survey questions became more apparent. during data production and monitoring, it was easy to verify suspicions and fire the surveyors who exhibited serious errors in judgement. the surveyors met to discuss questions of the survey, improve their understanding of how to deliver questions and interpret answers on a daily basis. this was both time consuming and exhausting, but an effective impediment to the would-be cheaters. (see appendix a for the overall project time-line and a daily schedule of survey activities.) due to financial and infrastructural constraints, we chose to (a) create a geographic information system of the survey area, and (b) use the inexpensive audio-capturing livescribetm smart pen technology. it not only allowed for collecting the data efficiently but also helped with monitoring and catching survey fraud and errors. it is worth noting that we were able to borrow the necessary gps equipment from a local institution for free which kept our costs in check. the use of a gis figure 2 provides a schematic example of the type of map employed for this project. the use of a gis made sense on several levels. mapping out the spatial location of households offered a nuanced understanding of local neighborhood effects. additionally, the development of the gis gave the supervisors, who were also the gis mappers, experience inside each village. this familiarity came in handy during the survey process. the gis also offered a simplified solution to: • dividing the villages into specific areas that were then assigned to surveyors. • assigning unique identifiers to the households. • managing the paradata (which households were non-responders, unavailable, or chose to delay the survey), and tracking the overall survey progress. • following up with the households to screen for fraudulent activities. • randomly choosing households during post-survey validation efforts. developing the gis itself produced some limited errors. a concrete example of this showed up during the ts when a son claimed he and his family ate separately while his father claimed the opposite. it was later found that the father and son had recently split and the father did not want to think of his son as living separately. these gis errors became evident during the questionnaire survey phase and it was relatively simple to correct them. yet another example of errors occurred when the responders would get confused about the definition of “household” and provide inaccurate answers. when the surveyors returned and found such households they were instructed to report them. the daily reporting session held every evening was a chance to discover these households and address the situation. the use of smart pens the basic feature of the smart pen is to have anything written on a special dot-paper recorded into the pen and synchronized with all audio that accompanies this writing. the resulting audio can be “played back” using paper replaytm with the pen and paper itself. it can also be replayed as an encapsulated pencasttm using desktop software such as adobe flash®, pdf, png or m4a or within the livescribe desktop software. a feature of this pen is that it captures only what the pen itself writes on special dot-paper and not what has been pre-printed (e.g. the survey form), on the paper (lackie et al., 2014). once the survey process was underway, the benefits and problems of the smart pens became clear. at the end of each survey day the interviewing teams returned to the base (where there was more reliable access to electricity than in the villages). the supervisors immediately began processing the day’s data: logging the survey forms, transferring the recorded “data” and audio from each pen to a computer, and backing all of that up again. most importantly, the smart pen provided a rapid recap of every recorded survey. each day the supervisors transcribed the paper surveys and listened to specific areas of the surveys that were suspect (transition points, questions that had been difficult to read or were simply skipped on the paper form). the figure 2: example household location/gis vol 40 no 3 / iassist quarterly 2016 23 iassist quarterly pencasttm made it possible for the supervisors to immediately fast-forward the recording to a specific point in the survey where surveyors were facing challenges. with this immediate and convenient method of double-checking the data, the surveyors and supervisors met daily to follow up on errors, specifically focusing on survey questions that were not being asked or recorded consistently across the surveyors. this daily iterative process continuously improved the quality of the data. it became clear to the surveyors that they had to conduct the surveys as instructed or risk getting caught and fired. the echotm smart pens: • are significantly less expensive than any computer tablet or screen-based device. • are small, easy to use, and not distracting in an interview. • are robust enough to withstand intense heat and humidity. • run on a single battery charge for the entire day (necessary as there was no way to recharge mid-day.) • are capable of storing a full day of interview data with full audio within the pen. • provides backup for every survey (i.e. paper forms and a digital document with full audio). lessons learned and effective strategies there are multiple ways of classifying errors that compromise data quality. certain errors are intentional while others unintentional. an example of an intentional error is turning in fake data to avoid the effort of conducting a genuine survey. reframing the survey question by modifying the language and mistakenly assuming this causes no bias is another type of unintentional error. from the pis’ perspective, the tools used to mitigate and repair these errors caused by unintentional mistakes, negligence, imprecision or fraud tend to be very similar. where it is clear that the hired surveyor is intentionally conducting fraud, it is best to terminate their contract right away. however, in cases where it is not as obvious, the surveyor’s ability to incorporate feedback and the magnitude of damage caused by their potentially unintentional errors, is the appropriate metric for deciding on termination decisions. as will become clear from the discussion below, the following four aspects of the ts significantly improved our ability to catch and prevent survey fraud and human errors: • having at our disposal the pre-survey gis map of the study area. • the use of livescribetm smart pens for audio-capturing the full interview. • checking the surveys daily (e.g. listening to the audio recording) for deviations from the survey protocols and debriefing the surveyors instantly. • sending supervisors to verify if certain households were surveyed properly when survey fraud was suspected. issues mitigated by requiring a full audio transcript of the interview in this section, we provide scenarios of survey fraud and human errors along with techniques used to mitigate the errors during the implementation of ts. in particular, we focus on survey fraud and errors that were effectively mitigated using the smart pens. scenarios 13 are examples of curbstoning while scenarios 4-9 are examples of cheating or unintentional errors. scenario 1: the surveyor was unsure whether the respondent would cooperate once they were located, and so ventured into the market place and asked some random individual to complete the survey. as surveyors were able to start or pause the audio recording of the smart pen as they wanted, some thought they would be able to hide the fact that they were interviewing the wrong person. solution: the fact that the audio capturing was turned off strategically to avoid recording the name of the respondent raised red flags. subsequently, supervisors were sent to these households to verify whether they were properly surveyed, if at all. fraudulent surveyors were caught and fired. scenario 2: surveyors claimed that the background noise from a weaving machine was too loud. consequently, the interview could not be heard in the audio-capture. solution: the supervisors were aware that the audio-capturing features of these smart pens are robust to these types of noises. thus, such claims also raised red flags and resulted in subsequent verification. scenario 3: surveyors claimed that the pen was not working on a specific day. solution: two approaches were used to deal with this. first, the surveyors were told that they could keep the smart pen at the end of the project, but only if it worked throughout the survey process. if their pen did not work, they would not be allowed to keep it. they valued the pen and consequently had proper incentives to protect the device. second, when they reported a pen not working at the end of a particular day, a supervisor was sent out to verify if households were actually surveyed the day for which the audio could not be captured. scenario 4: surveyors rephrased the questions incorrectly. for instance, replacing the phrase “life satisfaction” with “happy” in the question “how satisfied are you with your life?” note that in the swb literature they have very different meanings. solution: from the training sessions the supervisors and pi knew which questions would be most challenging and could quickly check the audio directly on the paper replaytm when the surveyors returned each evening. the surveyors who were caught not having followed their training were retrained. repeat offenders faced the prospect of being fired. 24 iassist quarterly 2016 / vol 40 no 3 iassist quarterly scenario 5:: surveyors changed the language of the survey and assumed answers. for instance, when asking a question on the perception of relative economic position compared with neighbors, the surveyor might truncate the response from a five-point scale of “much worse”, “worse”, “same”, “better” and “much better” to a three-point scale of “worse”, “same” and “’better’. the surveyor then uses their own judgement to convert the answer to a five-point scale response. scenario 6: surveyor unintentionally reverted to a pre-training way of speaking. solution (both 5 and 6): as a part of the training, the surveyors discussed their language use and developed a feeling for which questions were most likely to cause these problems. the immediate paper replaytm made it a simple matter to focus on the audio of how specific questions were asked and review them as the surveys were completed. in many cases, this was an unintentional result of fatigue and the tendency to revert back to their usual way of speaking. this was confirmed by listening to many of their surveys and hearing that they framed the questions properly during most interviews but slipped up on a few. repeated listening made them more aware of their language use and they quickly learned not to do it. scenario 7: socioeconomic differences played a role. the surveyors were educated and from the city, while the population surveyed are poor and rural. in the bangladeshi context there is an implicit understanding of socioeconomic hierarchies which lead both sides to act in certain ways. for example: a surveyor harshly demanding answers to questions. “why does it take you so long to answer this? answer quickly!” or showing anger or impatience with the respondent in any way. this lead to the respondents not thinking about their answers and answering quickly to get out of the circumstance. solution: the surveyors had to be trained about the deleterious effects of such behavior on the survey process and then to overcome these tendencies. they were regularly reminded that for the survey to be taken seriously, socioeconomic biases needed to be addressed. as the audio was being checked as they turned in their forms each night, this behavior was quickly detected. scenario 8:: although trained not to, sometimes the surveyor would prompt the respondents with an answer in trying to explain the question. scenario 9:: when transcribing multiple fields of data, sometimes it does not show internal consistency. for instance, a family indicated as not having an international migrant, nonetheless, seems to be receiving a non-zero remittance from one of its household members. another example involves transcribing the gender of a son of the household head to be a female. while it is obvious that there is a mistake here, it is not clear which field between the two contradictory ones is incorrect. solution (both 8 and 9): the audio playback provided the supervisor or pi with the necessary tools to repair the data entry. issues mitigated by the gis and post-survey validation the pre-survey gis data developed for the study area included information on the geographic coordinates of the household locations and the name of the household head. these data played a very important role in monitoring coverage, avoiding duplication of surveys, and dealing with erroneous transcription of household identifiers. it also made post survey validation a very quick and cost-effective process. here are some problematic scenarios that the gis helped to resolve: scenario a: occasionally the surveyors miscommunicate which houses had already been surveyed and would duplicate efforts and skip other households altogether. solution: this was caught when the household identifier was matched between the gis records and the survey data. scenario b: transcription errors of household identification numbers were more difficult to catch. it resulted in certain households mistakenly connected with a different gis location, such that two sets of data were then suspect. solution: comparing the name of the household head in the gis survey, with the questionnaire survey usually made it clear which survey had the incorrect household identification number. scenario c: some cheating was caught by re-surveying 20% of the households. these were chosen at random and the households were asked three verifying questions about, (a) whether someone surveyed their household (b) the name of all individuals who lived in the household, and (c) whether the surveyor instructed the respondent to not cooperate or answer in a fraudulent manner. solution: while the basic questions are simple, the process of verification is still a weak link. especially in the case of a violation of (c), the respondents may choose not to answer truthfully in the verification stage, out of fear that they would have to make the time to do the survey again. scenario d: surveyors occasionally chose to skip some households, but claimed that the household said they did not want to take part in the survey. it did not happen very often legitimately, due to the good relationship they had developed with the villagers. solution: all households who refused to participate were followed up on. vol 40 no 3 / iassist quarterly 2016 25 iassist quarterly issues mitigated by the gis and post-survey validation the tips from the few articles about mitigating surveyor fraud seemed to hold true for the ts (e.g. recording the interviews, providing random checks to confirm that the interview took place as expected, following up on protocol violations, and rigorous oversight with little isolation of surveyors). in addition, some new issues arose with this group. these hired surveyors were hand-picked, well-paid graduate university students who were gaining excellent field experience. several had hopes of attending school in the us or elsewhere and needed letters of reference. still, they required a great deal of attention and persistent following. a few students did not like being monitored and in retrospect conducted much of the survey fraud. this group tried to unionize and extort a higher wage early in the process. they started rumors about the pi making money off their hard work and being involved in financial fraud. they realized that the pi was under a binding time constraint and so started to engage in extortionary behavior. the solution to these issues included persistence, openness, and being willing to fire and replace them very quickly. these surveyors were also the ones who worked to cheat whenever possible and exhibited a disdain that their “usual” methods of survey fulfillment (e.g. survey fraud tactics) would not work because of the rigorous and prompt data verification process. when they were fired, the rest of the survey team took notice and worked very well. a few proactive bad apples can seriously hamper the process and it is important to deal with them swiftly and transparently. references aapor (2003) interviewer falsification in survey research: current best methods for prevention, detection, and repair of its effects, survey research, 2004 newsletter from the survey research laboratory, college of urban planning and public affairs, 35, 1, university of illinois at chicago. bhuiyan, m. f. (2012) relative consumption: a model of peers, status, and labor supply, american journal of agricultural economics. bhuiyan, m. and szulga, r. (2013) the tangail survey: household level census of subjective wellbeing, perceptions of relative economic position, and international migration: 2013 [tangail, bangladesh]., http://doi.org/10.3886/e61146v1. biemer, p. p. and stokes, s. l. (1989) the optimal design of quality control samples to detect interviewer cheating, journal of official statistics, 5, 23–39. blanchflower, d. g. and oswald, a. j. (2004) well-being over time in britain and the usa, journal of public economics, 88, 1359–1386. blasius, j. and friedrichs, j. (2012) faked interviews, in methods, theories, and empirical applications in the social sciences, springer, pp. 49–56. bredl, s., storfinger, n. and menold, n. (2011) a literature review of methods to detect fabricated survey data, zentrum für internationale entwicklungs und umweltforschung 56, courant research centre peg. bredl, s., winker, p. and kötschau, k. (2012) a statistical approach to detect interviewer falsification of survey data, survey methodology, 38, 1–10. clark, a. e. and oswald, a. j. (1996) satisfaction and comparison income, journal of public economics, 61, 359–381. cole, h. l., mailath, g. j. and postlewaite, a. (1998) class systems and the enforcement of social norms, journal of public economics, 70, 5–35. crespi, l. p. (1945) the cheater problem in polling, public opinion quarterly, 9, 431–445. deaton, a. and stone, a. a. (2013) two happiness puzzles, american economic review, 103, 591–597. diener, e., suh, e. m. and smith, h. (1999) subjective well-being: three decades of progress, phsychological bulletin, 125, 276–303. duncan, o. d. (1975) does money buy satisfaction?, social indicators research, 2, 267–274. dynan, k. e. and ravina, e. (2007) increasing income inequality, external habits, and self-reported happiness, american economic review, 97, 226–231. easterlin, r. a. (1995) will raising the incomes of all increase the happiness of all?, journal of economic behavior and organization, 27, 35–47. easterlin, r. a. (2001) income and happiness: towards an unified theory, economic journal, 111, 465–484. fafchamps, m. and shilpi, f. (2008) subjective welfare, isolation, and relative consumption, journal of development economics, 86, 43–60. ferrer-i-carbonell, a. (2005) income and well-being: an empirical analysis of the comparison income effect, journal of public economics, 89, 997–1019. harrison, d. e. and krauss, s. i. (2002) interviewer cheating: implications for research on entrepreneurship in africa, journal of developmental entrepreneurship, 7, 319. knies, g., burgess, s. and propper, c. (2007) keeping up with the schmidts: an empirical test of relative deprivation theory in the neighbourhood context, discussion papers of diw berlin 697, diw berlin, german institute for economic research. koczela, s., furlong, c., mccarthy, j. and mushtaq, a. (2015) curbstoning and beyond: confronting data fabrication in survey research, statistical journal of the iaos, pp. 1–10. kuriakose, n. and robbins, m. (2015) don’t get duped: fraud through duplication in public opinion surveys, statistical journal of the iaos (forthcoming). lackie, p., ketama, m., loery, g., mansour, c. b. and strauss, b. (2014) smartpens as data capture devices: survey data from handwriting on paper to automatic csv file, unpublished manuscript, carleton college. luttmer, e. f. p. (2005) neighbors as negatives: relative earnings and well-being, quarterly journal of economics, 120, 963–1002. mayraz, g., schupp, j. and wagner, g. g. (2009) life satisfaction and relative income: perceptions and evidence, cep discussion papers, centre for economic performance, lse. mcbride, m. (2001) relative-income effects on subjective well-being in the cross-section, journal of economic behavior and organization, 45, 251–278. mcguire, m. t., fairbanks, l. a. and raleigh, m. j. (1995) life-history strategies, adaptation variations, and behavior-physiologic interventions: the sociophysiology of vervet monkeys, new york: oxford university press. senik, c. (2009) direct evidence on income comparisons and their welfare effects, journal of economic behavior and organization, 72, 408–424. 26 iassist quarterly 2016 / vol 40 no 3 iassist quarterly solnick, s. j. and hemenway, d. (1998) is more always better?: a survey on positional concerns, journal of economic behavior and organization, 37, 373–383. van de stadt, h. a. k. and de geer, s. v. (1985) the impact of changes in income and family composition on subjective well-being, review of economics and statistics, 67, 179–187. waller, l. g. (2012) interviewing the surveyors: factors which contribute to questionnaire falsification (curbstoning) among jamaican field surveyors, international journal of social research methodology, 16, 155–164. notes 1. acknowledgements: we are very grateful to the department of economics at carleton college for their valuable advice. the data gathering process has been a joint project of radek szulga (economics, lyon college) and bhuiyan, and was funded by the dean’s office and the economics department of carleton college. paula lackie is the technology adviser for the project. in bangladesh, we were helped by a dozen seniors of dhaka university (du) from various disciplines and prof. a q m mahbub (geography & environment, du). we are greatly indebted to them. the respondents in these villages took part in the surveys voluntarily and were very kind in donating their precious time to our cause. last but not the least thanks to yuping huang, iris wang and caroline greenberg for their fantastic research assistance. 2. contact author muhammad f. bhuiyan is an assistant professor in the department of economics, carleton college, northfield, mn 55044, usa. email: fbhuiyan@carleton.edu. corresponding author paula lackie is the academic technologist for data, information technology services, carleton college, northfield, mn 55044, usa. email: plackie@carleton.edu 3. the two pis for this project are muhammad f. bhuiyan and radek szulga. paula lackie was the technical advisor for the project. the survey was funded by the discretionary fund of the dean of the college office (carleton college, mn) and the department of economics (carleton college). bhuiyan is currently an assistant professor of economics at carleton college (mn, usa) while szulga is an assistant professor of economics at lyon college (ar, usa). 4. see mcguire et al. (1995); van de stadt and de geer (1985); blanchflower and oswald (2004); clark and oswald (1996); dynan and ravina (2007); easterlin (2001); duncan (1975); cole et al. (1998); diener et al. (1999); deaton and stone (2013); luttmer (2005); j. solnick and hemenway (1998); bhuiyan (2012); mayraz et al. (2009). 5. while the survey was conducted in bengali, the answers were interpreted onto the survey form in english during the interview. surveyors needed to learn to deliver the survey in a strict format and interpret the answers onto each form in a consistent way. 6. livescribe, paper replay and never miss a word are trademarks of livescribe inc. all rights reserved. ©2014 livescribe inc. http://www.livescribe.com/en-us/faq/online_help/maps/connect_desktop/r_ formats-for-sending-notes-and-audio.html 7. for a discussion of downloadable dot paper: https://support.livescribe.com/entries/ 22263341-60013-printing-free-livescribe-3-or-livescribe-wifi-smartpen-dot-paper 8. paper replaytm is the process that the livescribetm smart pen uses to play back any audio recorded when specific text was written with the pen on special livescribetm dot paper. there is both a microphone and speaker built into the pen. sist newsletter vol.1, no. 4 book n t i c e s / kathleen m. heim machine-readable social science data special issue of the drexel library quarterly . volume 13 (january 1977). edited by howard d. white. issue available for $5.00 by writing: drexel library quarterly , graduate school of library science, drexel university, philadelphia, pennsylvania 19104. contents: "numeric data files: an introduction." howard d. white. "the pre-acquisi tion process: a strategy for locating and acquiring machine-readable data." alice robbin. "stalking the wild data set: the acquisition of machine-readable social science data at home and abroad." david nasatir. "cataloging machine-readable data files a first step?" sue a. dodd. "social science data files, the research library and the computing center." douglas ferguson. "the information transfer process on the university campus: the case for public use files." lorraine borman and richard hay, jr. "training the professional data librarian." judith s. rowe and carolyn l. geda. "this is d contribution to a growing discussion between and within the computing and library worlds." douglas ferguson though ferguson is referring to his own article in this special number of the drexel library quarterly , his statement quite accurately describes the entire issue, for this marks the first time that the library press has devoted a substantial amount of space to treatment of the problems that confront the data archivist. the contributors are familiar names to those active in the data archive movement--all have published extensively and have been involved in the conceptual definitions of the principles of data archiving. because the intended audience of this issue is the broad range of library and information professionals, the treatment of matters familiar to the working data archivist may seem elementary. howard d. white's introduction to numeric data files is a careful building of concepts and definitions that provides the newcomer with vocabulary and a general framework on which to base an understanding of the technical articles that follow. for the novice, douglas ferguson's article follows white's most logically for he assesses the role of the library in data archive development and provides the rationale for cooperation. his use of the specific example of the stanford university experience provides a cogent model of library-archive interaction. lorraine borman and richard hay, jr. extend the concepts defined by ferguson by presenting the present-day situation in universities for the use of public files. they attack the problem of low use by focusing upon the kind of public use files in a social science file collection; illustrate how these files can be accessed and manipulated by both novice and sophisticated users; and show how this type of use is related to the utilization of public information files and information systems, and assists in the total information transfer process. 37 newsletter vol. 1, no. 4 from these articles the more specialized essays by alice robbin, david nasatir, and sue a. dodd may be approached. robbin's scrutinization of the pre-acquisition process is a detailed, logical analysis of the complex factors that enter into the decision to obtain a particular data set. after a general introduction which places the data archive squarely as a type of special library, robbin explores the five factors which influence the pre-acquisition process and the search strategy for locating data. a diagrammatic presentation of the search strategy is presented with each part carefully explained in an appendix. nasatir's brief piece, "stalking the wild data set," is by far the most dramatic contribution to the entire issue. it is a simple, yet vivid rendering of the sometimes machiavellian, sometimes mendicant contortions which a data archivist must endure to procure material. working data archivists will empathize, and those unfamiliar with the field will marvel, at nasatir's travails. dodd's article on cataloging machine-readable data files (mrdf) may have been the most difficult piece to compose. it has fallen to dodd to unravel the lengthy history of both the library and data archive profession's attempts to establish conformity in the cataloging of mrdf. the article is loaded with acronyms, committees and subcommittees of the ala, asis, lassist and aect. though difficult at the outset her essay provides a solid introduction to the monumental problems involved in hammering out a cataloging format agreeable to the various factions. she examines progress to the present focusing on the lassist classification group project designed to test the feasibility of cataloging social science data files according to the ala subcommittee's recommended procedures. an appendix of sample catalog cards from those participating in the project is included. finally, judith s. rowe and carolyn l. geda discuss the training of the professional data librarian. the profession is still in its formative stages and current practioners are drawn from a diverse group: programmers, social scientists or librarians. the only formal training available is via the inter-university consortium for political and social research summer program held at ann arbor, michigan, which is described in detail. as stated above, the intended audience of this special issue is the working librarian or information professional. this audience requires definitions that may be elementary to the data archivist. however, howard d. white's editing has produced a diverse publication that operates on several levels. his contributors have at once provided a sourcebook of articles comprehensible to one unfamiliar with data archiving and yet valuable to the experienced archivist. dodd's patient unraveling of the machinations of cataloging mrdf and robbin's precise articulation of pre-acquisition procedures are the only explanations of these processes in the literature. the essays devoted to the reference and information component of data archiving extend current ideas about service and begin to outline the emergent user-models. this issue will stand as a vital resource for data archiving until some more comprehensive publication supersedes it. the only glaring defect is the lack of a simple introduction that provides a framework for the whole. nevertheless, this issue of the drexel library quarterly is a necessary item for all data archivists. it will provide the profession with a stronger realization of its role in the information transfer process and it will provide traditional librarians with an understanding of the role of data archives in that process. 38 sist newsletter vol.1, no. 4 richard c. roistacher and barbara c. noble. computer network support of social research communities . university of illinois. graduate college. center for advanced computation. urbana, illinois 61801. in march of 1977 following the first lassist north american conference, richard roistacher provided participants with publications and technical reports which fell into three major classes: 1) leaa data archive material; 2) data archiving and analysis methods; 3) sociology of computing. all of the papers advance the conceptualization of the role of data archives in the research community. mr. roistacher has agreed to provide copies to interested data archivists who contact him at the center for advanced computation. in this notice we would like to focus upon one of the papers which falls into the category of the sociology of computing: computer network support of social research communities . this paper, a pre-print of an article to be published by human organization , sets forth the idea that present computer technology will allow geographically separated social researchers to share data and data analysis tools, collaborate on research and writing, and communicate with their colleagues much as if they were in a single department. a description of the law enforcement assistance administration (leaa) computer network based research support center which operates over a commercial communication network is described in detail. the leaa data archive and research facility for social scientists and program evaluation groups which has been developed by the center for advanced computation at the university of illinois to provide research support services to clients scattered around the country is viewed by roistacher as a prototype. the services offered by a computer network fall into three major categories: data management and archiving; computer and statistical consulting; and communications and documentation. roistacher uses the generic term. research support facility (rsf), to describe the multi-purpose service which this prototype provides. services fall into two modes: those which do not require a computer network connection to the client (data archives; consulting on analysis of archive files; consulting on client-collected data; agency for other archives; bibliographic information retrieval; and software distribution) and those which can only be provided over a direct network connection (shared archival data; shared software; on-line consulting; shared user data; remote technical services; network mail; network conferencing; facilitating the formation of working groups; and a network journal). all services are described in detail. potential clients for rsf services include academic researchers, independent research organizations, state planning agencies, and program evaluation groups. computer and network costs are compared favorably with the costs of using local facilities. though charges seem high roistacher demonstrates that they are much lower than travel, postage, typing, duplicating and time considerations under the present system. management considerations in remote services are treated. these include training users, administrative procedures, accomodation of clients in distant time zones, terminal acquisitions, and clear and self-sufficient documentation. 39 sist newsletter vol.1, no. 4 the social consequences of networking as perceived by roistacher's observations of the department of defense's arpanet are discussed and the implications for collaboration among colleagues regardless of geographical considerations explored. though data archives are only one component in the larger rsf, they are a vital component and data archivists will want to ponder roistacher's predictions for computer-based research support. his experiences with a highly developed facility provide a model for future developments in behavioral science communication and augur a new perhaps more democratic era for the invisible college. statistical policy division. office of management and budget. in cooperation with the federal statistical agencies with responsibility for the collection, processing, analysis and dissemination from major federal statistical programs. framework for planning u.s. federal statistics 1978-1989 . (reprinted from the statistical reporter , 1974-1976) . copies available from the statistical policy division, office of management and budget, 726 jackson place, n.w., washington, d.c. 20503. this collection of reprints from the statistical reporter details the efforts of the statistical policy division of the office of management and budget to provide for a more systematic review of the needs for improved statistical planning and coordination. a topical outline of the planning framework is presented along with articles on its nature. the goal is improved coordination of the highly decentralized u.s. federal statistical system and setting of statistical priorities. of particular interest is the article, "user access-data banks," which includes a bibliography of major agency guides and indexes to statistical data. much concern is expressed over the failure of government agencies to comply with government regulations which provide for improvement of the dissemination of statistical information on a timely basis. feedback is requested from users of government generated data. lassist members can obtain a copy of the framework for planning u.s. federal statistics 1978-1989 from the above address and reply to the same on matters concerning shortcomings in the statistical services of the u.s. 40 vol183&4 9fall/winter 1994 introduction while there is no one model for providing services for data in colleges and universities, it is increasingly common for various constituencies to cooperate, especially in lean fiscal years. there are both positive and negative aspects to pooling resources in such a “marriage of convenience.” although not the solution for everyone, this paper will take a look at a partnership among two academic departments, computing services, the libraries, and the provost’s office at binghamton university, state university of new york. ‘ it will suggest advantages and disadvantages for those considering cooperative ventures at their institutions. from this day forward, for better for worse, for richer for poorer, in sickness and in health, to love and to cherish, till death us do part ... data library and service operations in academic institutions in north america have in many instances seen a reduction in resources in the last five years. in some cases, this has threatened the existence of some or even all services. in others, it has caused data professionals and administrators to be creative and forge new arrangements to maintain or even enhance basic levels of service for their clientele. because data service organizations vary considerably from academic institution to institution, there is no single or simple way to diagram a preferred organizational structure for data service. what works in one academic setting, may not in another. services seen as basic at one university may be on a wish list at others. size and diversity of user groups also vary depending on programmatic and research agendas. in any case, optimal staffing and funding levels are directly related to the level of service needed by an institution’s primary clientele. unfortunately, even minimal resource levels may not be possible at some institutions. pooling resources among departments and units across a college or university can be an option where a separately funded “data center” or “data library” does not exist, or, when an existing service is faced with dissolution. while these ‘marriages of convenience not suitable for all organizations, there are significant advantages and disadvantages of academic partnerships. they are especially worth exploring if an institution faces “rightsizing” or consolidating services. these partnerships rely on the ability of various constituencies to work together, an agreed upon common purpose, mutual respect, and tolerance. from this day forward... an institution’s history of providing quantitative, social research support on a given campus will often set the stage for future service configurations. because of this it can be difficult to change support paradigms, although it is certainly possible and even necessary in some cases. at binghamton university, state university of new york, the political science department in conjunction with an organized research center, provided support for quantitative social data for two decades. in 1990, a time of considerable fiscal uncertainly in the university system, the impending closing of that research center necessitated rethinking the way in which we were organized to provide data services. for the most part this meant fulfilling our inter-university consortium for political and social research (icpsr) membership responsibilities and related data services. after a series of extensive consultations with administrators, faculty and staff, the libraries agreed to assume responsibility for ‘data services.’ this primarily entailed maintaining formal relationships with icpsr and later the u.s. bureau of the census’ state data center program. ultimately this meant that the libraries would: . maintain formal relations with icpsr .serve as liaison for the state data center program . assume fiscal responsibility for icpsr membership after an initial transfer of monies from the provost’s office . provide customer services, particularly identifying and ordering data for better or for worse: academic partnerships for data services by diane geraci1 binghamton university state university of new york 10 iassist quarterly .collect and maintain codebooks, related technical documentation and statistical manuals .provide user consultations, research assistance and referrals .cooperate with academic computing, to make data available and to provide complementary services .cooperate with the economics and political science departments, and the assistant provost for graduate studies and teaching to assign two icpsr/data services graduate assistants to the libraries. the formal change in service occurred in july 1 1991 to coincide with the new fiscal year. however, academic computing, the libraries, the political science department, and the organized research center had already begun the process of working together several years before. this early period effectively served as a ‘getting to know you” phase where each unit’s service orientation and working patterns became known. evolving service plans and position descriptions assisted in making clear who would be responsible for which aspect of the reconstituted service. for better for worse... commitment of each constituency is essential for a service that exists through the shared agreement of its partners. the best strategy for success is creating a winwin situation whereby each of the partners benefits from contributing to the service. a benefit may mean better meeting the mission of the unit, such as a library or computing service that serves the entire academic community. from an institutional point of view it may mean reducing duplicate purchases or services. it certainly will mean providing the kind of research support desired by relevant academic programs. it can also mean acknowledging that going it alone might not provide the depth and range of services needed. while good will and intentions may characterize a shared agreement to provide service, a written plan is well worth the effort. support staff and administrators do change. a written service plan cannot absolutely guarantee the continued cooperation of each unit but it does provide a framework and codification of responsibilities. after seven years of sharing responsibility for data services on the binghamton campus, several benefits are evident. they include: • icpsr membership benefits are more widely available to all constituencies on campus. there had been a perception that everyone knew about the icpsr and the extent of their data holdings. this turned out not to the case. new faculty and graduate students continually arrive on campus and existing campus instructors and researchers have new data needs. researchers in departments not traditionally employing quantitative research methodologies may begin doing so. there is a continual need for dissemination of information about new data and related data news. for example, only one department knew about the icpsr summer program in quantitative methods before the libraries coordinated the membership services. • duplication of data acquisitions was reduced. because data support originally resided in the school of arts and sciences, other schools and divisions often bought there own data directly from producers. we found that much of these data were available via our icpsr membership. this was especially true for health data and economic time series data. • integration of data collected in several media is a positive by-product of centering access to data in the libraries. print resources, cd-roms, diskettes, remote access via the internet, and commercial services already are available in or through the libraries. making the libraries the first stop to ascertain if data are available on mainframe cartridge tape has brought together conceptually if not physically, access to related resources. • existing expertise is utilized; that is, information management skills, computing skills and service orientation in the libraries; technical, computing and statistical skills from computer services; research skills of the departmental graduate assistants. • skills shared across units increase the skills of all contributors to the service. graduate students particularly gain solid experience working with data and valuable statistical programming skills. . cooperation with other units on campus increases awareness of research needs as well as understanding of different campus cultures. daily contact with colleagues in other campus units greatly fosters understanding and respect for each other’s work. several difficulties or less positive aspects of the partnership also became apparent. we also found: • icpsr resources became more widely used on campus making it difficult for part-time staff in the several units providing support to keep up with demand. statistics showed a substantial increase in data use on our campus as a result of the reconstituted 11fall/winter 1994 service. staff in the libraries and in academic computing found that an increased percentage of their work week supported data services. some reorganization of duties occurred in each unit with the pressure being born by existing staff members. similarly, the service began with one graduate assistant. it soon became clear that one was insufficient and we were able to negotiate for another student. • reliance on graduate student support entails constant training and rotation of staff. considerable fluctuations in the quality of service regularly occur. • additional permanent staff is desirable, but thus far, has been unattainable. research level support is very time-consuming. permanent staff and new lines are difficult to acquire. they would assist in providing consistent service and allow for performance of needed tasks, especially as the number of users increases and users’ request an increased level of service. • keeping current with data services developments requires additional space and equipment. changes in computing platforms and storage devices require new hardware and software. decisions made in one unit may affect another. for example, the decision by computing services to stop maintenance of 9-track tape drives has consequences for the way that the libraries order data. • new skills are required. for example, knowledge of database maintenance, cataloging, or statistical programming, and understanding research design may be necessary for data services staff to provide certain services. for already overextended staff, there is not adequate time for learning new processes or acquiring necessary skills. the aptitudes of existing staff for acquiring new skills will also vary. • cooperating with other units on campus is difficult in practice. conflicting priorities in a unit or between units may be difficult to resolve. politics internal to a unit are less easily negotiated by those outside the unit. service orientations or philosophies of the partners may differ. in sickness and in poor-health... in times of staff reduction, fiscal uncertainty, competing demands in a unit, or simply a reprioritization of needs or goals, a joint service can suffer the consequences. there can be real concerns for the integrity of the service as a whole if a key group withdraws its support. when individual units experience shifting priorities or staff reductions the danger exists that the shared service will fall to the bottom of the list of things to do, or worse, will no longer be supported. when there are administrative changes the partners in the service may need to renew their “vows.” while living with a small degree of uncertainty is admissible, a crisis can arise if one contributor to the service can no longer participate or even temporarily suspends participation. major disruption of service or stress on the other partners can occur if one unit is unable to meet their obligations. there is not a way to absolutely ensure that no change will occur in a partner’s commitment to the relationship. there are ways, though, to engender support for the service and keep it on the priority list of each partner. relying on a core group of researchers as an “advisory group” is one way to get feedback from users. measuring the amount of data ordered, number of users assisted, computer usage, and any other relevant factor at an institution can demonstrate the utility and necessity of the data service to administrators. to love and to cherish... when there is stability in the service and researchers’ needs are being met, all partners deserve congratulations for cooperating across units and effectively working together to create a viable service. this is the ‘feel good” outcome of a win-win situation and should be enjoyed. lest complacency cause problems, it is a good idea to reaffirm what works with the arrangement and what can be handled in a better way. assessment during the good times is much less threatening then when the sky is falling due to impending budget cuts or some other “natural” academic disaster. several methods work well to evaluate the service including meeting with the primary front line staff in each unit, consulting an advisory group of researchers, and surveying past and prospective users of the service. taking the time for assessment is a positive way to renew the agreement and service plan(s) of the units involved and make any necessary adjustments. till death us do part? binghamton’s “marriage of convenience” came at time when data support on the campus was in jeopardy. it has served the university community well in its time. it does not mean that this is the only way to provide data services or that another type of service will not evolve from it. there are several reasons a partnership such as binghamton’s might cease to continue: . the service is no longer necessary. there may be 12 iassist quarterly other ways to meet the need of data users. schools or departments might decide to provide some of their own services. national or international consortia and computer networks may provide more data services negating the need for some local services. it is difficult to imagine, though, that some measure of local support will not be necessary, even in a future of distributed services over “the net.” there certainly will be a time when the service needs to be reformulated or reconstituted. . one or more of the partners cannot afford the commitment of staff and/or resources. a worst case scenario is the service dies. another possibility is that the other partners are able to pick up the slack. in the case where the partners are unable to absorb additional responsibilities, providing a reduced level of service may be necessary. . cooperation is no longer possible between the partners. one of more of the partners may experience a change in their mission, unresolvable disagreements may occur between partners, or administrative prerogative may preclude further cooperation. providing data services through an academic partnership can be very rewarding. forging key relationships between disparate units and seeing positive results in support of research and teaching are successful outcomes. before embarking on a cooperative venture, careful consideration of a partnership model’s suitability for the needs and culture of an institution is necessary. 1. paper presented at iassist 1994 in san francisco. data infrastructure for the social sciences in the netherlands by gerton heyne ' iva, institutefor socila research. university of tilburg, the netherlands introduction computers have become commonplace during the past few years. there has been an obvious increase in, on the one hand, the computerization of data files, and on the other, the analysis and processing of machine-readable data files by computer. the pressing need, particularly by the state, for reliable and up to date policy information and the desire to raise the level of efficiency have been factors in this development data stored in computerized form offers greater possibilities for analysis. as a result, scientific research is also able to tap into a growing reservoir of available machine-readable fdes. technical innovations have moreover made it possible for researchers to collect and process research data more quickly and efficiently in order to obtain suitable material for analysis. in part because of innovations in research methods and techniques, researchers are able to carry out complex data analysis not possible before. the technical infrastructure is growing steadily because researchers are increasingly utilizing pcs and mainframes to set up and conduct research and to analyze machine-readable data fdes. in addition, they also have access to a growing number of networks which can be used to exchange data: for example, the various local university networks and the national surfnel until a few years ago, when researchers wished to conduct data analysis one of their most important sources was data published in written form. the most important data fdes were compiled for the state by the central bureau for statistics (cbs). the availability of machinereadable data files on the one hand and the development of infrastructural facilities such as pcs and networks on the other are raising the demand for this information in machine-readable form. the possibility of conducting data analysis grows when researchers can access machine-readable fdes instead of written publications. in addition such fdes serve increasingly as "background information." researchers have access to ever greater amounts of data which can serve as important sources of secondary information for designing and conducting research, because access to files already in use can also, in principle, be made simpler and easier. the importance of a well-designed infrastructure is evident. researchers must have easy access to statistical material. research in the social sciences makes it possible to gain systematic and detailed insights into society. such insights serve two purposes that are, in fact, directiy related. firstiy, research based on statistical information is important for both basic scientific research and for more policy-directed research. insights gained in this way are of scientific value in and of themselves, but they are also important for efficient management and confident policy-making. secondly, research serves a fundamental, democratic purpose. a balanced process of policy-making in a democratic society requires that all individuals and parties involved possess relevant information. such a balance can only be strengthened by a vigorously independent research capacity autonomous of the expertise of the state. to achieve this, however, requires sufficient access to statistical information such as that compiled officially by cbs and other institutions. in recent years a number of reports have put forward the idea that the data infrastructure used by the social sciences in the netherlands needs to be augmented (de bie, 1989; de guchteneire & timmermans, 1990). the authors conclude that researchers spend a great deal of time collecting and converting data files before being able to start analysis. they observe the following problems: the high cost of purchasing fdes; the restrictive conditions under which files are made available; the uneven quality of both the files and the accompanying documentation. the assumption is that these problems prevent researchers from making adequate use of the data present. delivery of data files is not optimal. the service provided by suppliers in general, and cbs in particular, is insufficiently geared to the growing demand by researchers for data files and statistics in the form of machinereadable numerical material. the technical innovations described above have transformed desires and wishes on both the supply and demand side of data files. in addition, priorities have shifted. while the transformation process was still underway, the various parties formulated new goals and fall/winter 1990 interests that did not always turn out to be complementary. researchers insist that data files be made available at minimal cost and under the least restrictive terms possible. the policy of cbs, the most important producer of data files, is not entirely self-determined but depends partly on political considerations. this supplier is obliged to comply with political agreements concerning the cost of field work and development and with statutory regulations concerning privacy. part of the criticism that is being heard, therefore, concerns political leaders rather than cbs. it looks as if technical and social innovations have set a process of change in motion that has caught the parties involved inadequately prepared. measures have been taken to secure the interests of both producers and consumers of data files as well as of the individuals whose personal information has been collected. however, these measures often seem to seriously impede the work of researchers. institutes and individuals), it was almost impossible to ask all of them to participate in the study. the goal was in any case to approach all relevant faculties. for this reason a number of relevant university faculties and research institutes were selected. an attempt was made to assign each research group or institute a contact person in charge of coordination. only in a small number of cases did this contact person have a full understanding of every aspect of external data file use in his or her research group or institute. the rest of the time, contact persons were requested to distribute the questionnaires to those researchers who might answer questions concerning the use of external data files. this paper will provide a short description of a. the most important information as provided by two institutions that supply data files to third parties (cbs and the steinmetz archive); b. the most important results of the survey distributed to a number of data file users. design of the present study data infrastructure is becoming increasingly important to researchers. improving this infrastructure is extremely important for the quality of social scientific research. the present study considers the current state of affairs with respect to the availability, accessibility and use of statistical data files 2 in the social sciences3 . which files are actually being used in research? where do the bottlenecks occur in use or loan? the study comprises an inventory of the use of external * statistical data files in 1989. the files had to meet a number of criteria formulated beforehand.5 in order to gain as great an insight as possible into the use of external data files, both the suppliers and consumers of such data files were approached. this also offers insight into the way in which data files are made available for research in the social sciences. this is the first time that the above-mentioned suppositions concerning the functioning of data infrastructure have been subjected to systematic and relatively large-scale testing by those actually involved in this issue in practice: the researchers. a number of suppliers of data files were requested to provide information about files either sold, lent or offered in some other fashion to researchers in 1989. a number of questions concerning the price, the size and the consumers of these files were added. a number of users of data files received a questionnaire concerning files they had used in 1989. this also included questions concerning the price and size of the files, time of delivery, problems, quality, etc. given the large number of research institutions, both commercial and non-profit (including faculties, research groups, use according to two suppliers cbs and the steinmetz archive, the dutch data archive for the social sciences, were asked to answer a number of questions concerning the files that they supplied to researchers in 1989. central bureau for statistics (cbs) cbs provided an outline of a two-year period, specifically the microfiles and publication files 6 delivered in 1988 and 1989. the reason given was that in this way chance fluctuations would be less likely to misrepresent the information. the following deliveries took place in 1988 and 1989: 91 deliveries of 20 different microfiles to research institutions; 130 deliveries of 10 different publication files, of which: * 25 went to research institutions; * 66 went to businesses; * 39 went to the state. two publication files were delivered 105 times. two publication files were by far the most popular, being delivered 66 and 45 times respectively. in addition, eight micro-files were made available more than five times in the period concerned; the socio-economic panel was the most frequent (16 times). cbs additionally supplied aggregated data7 concerning the prices paid for various files by showing proceeds per file. it was impossible to deduce from this information how much individual deliveries yielded, or rather, what consumers must have paid for them. for this reason an outline of average prices8 must suffice. averages should iassist quarterly be interpreted with the necessary degree of caution, but the average amounts can serve as an indication. in 1988 and 1989 for individual deliveries of eleven microfiles average amounts have paid above 10,000 dutch guilders (approx. $ 6,060). the data shows that most microfiles are quite expensive. the highest (average) amount for one of these microfiles was 58,800 guilders (approx. $ 35,600). publication files, on the other hand, cost significantly less. prices of the publication files supplied most often averaged 630 guilders (approx. $ 380) and 340 guilders (approx. $ 205) respectively. deliveries yielded 1.75 million guilders (approx. $ 1.05 million) in total; 1.59 million guilders (approx. $ 0.96 million) came from microfile deliveries. publication files yielded 0.16 million guilders (approx. $ 0.10 million). the amounts mentioned do not include compensation for additional field work or development costs. the steinmetz archive the steinmetz archive distinguishes between research files for the social sciences and weekly public opinion research (weekly questionnaires). the following will deal only with the first category. in 1989 the steinmetz archive sold9 or made a file available 286 times. a total of 1 1 1 different files were offered; 42 remained after various volumes or waves belonging to the same file were combined. the table below indicates how often dutch and foreign files are used in the netherlands and how often dutch files are used in foreign countries. 10 table 1. use in and outside the netherlands of the steinmetz archive use in the netherlands dutch files 110 foreign files 71 total 181 use outside the netherlands 105 215 105 71 286 there are 1 10 dutch files and 71 foreign files in all used by dutch researchers. almost half of the dutch files were sent outside the netherlands. three distinct groups of foreign users can be distinguished: universities, data archives and research institutes, comprising respectively 81%, 13% and 6% of the foreign consumption. in the netherlands users are universities and research institutes, which received 87% and 13% of the files respectively. use according to the users in the survey 70 institutions were requested to complete one or more questionnaires concerning the use of external data files in 1989. the response was low. table 2. response approached response university research groups social sciences economics/business 44 31 13 18 11 7 41% 35% 54% research institutions (non-profit) 23 8 35% research institutions (commercial) 3 0% total 70 26 37% of the 44 institutions that did not return a completed form, ten stated that they did not work with external data files and can thus be categorized as non-users. the remaining institutions failed to respond at all. the controlled response was therefore upwards of 51%. the 26 institutions, research groups and faculties returned 80 questionnaires in all with information concerning 121 files. because of the high level of non-response we do not consider the data ultimately received as statistically representative. the data does not provide an exhaustive view of the use of external data files. in our opinion, the results are nevertheless important for gaining insight into the use of files and the issues related to this use, in view of the number of completed questionnaires and their distribution across various institutions. intensity of use the table below provides numerical data concerning the intensity with which the data files supplied were used. a distinction is made between cbs files and files coming from other supplying agencies. fall/winter 1990 table 3 intensity of use cbs rest total 1. a few analyses shortly after receipt 2% 4% 3% 2. many analyses shortly after receipt; no others after this 10% 11% 10% 3. a few analyses over a long period of time 16% 37% 26% 4. many analyses over a long period of time 62% 45% 54% 5. unknown 10% 4% 7% total 100% 99% 100% intensity refers to use in time and the number of analyses performed on the file. the figures indicate that the files were generally used intensively. approximately (54+26=) 80% of the files were used over a long period of time, and many analyses were performed on about (10+54=) 64%. for well over half (54%) of the files, many analyses were performed for long periods of time. use of the cbs files is proportionately more intensive, in the sense that in general many analyses were performed. types of research the study asked what type of research the files were used for. the respondents were asked how often the file concerned was used for contract and how often for noncontract research. the results showed that files were chiefly used for contract research, in particular the cbs files (see figure 1). indirecdy this was actually a query as to who subsidizes the research (and the files). this information made it possible to indicate which commissioning agencies were direcdy or indirecdy involved in the delivery terms for data files. for non-contract research this is the state. the respondents were asked to indicate who commissioned contract research. the state was named as commissioning agency in 75% of the cases. the conclusion is that the vast majority of research in the social sciences is commissioned and subsidized direcdy by the state. all in all 33 percent of this was charged to the ministry responsible for scientific research, the ministry of education and science. delivery time for files questions were posed as to the amount of time that passed between the first formal contact with the supplier and the actual delivery of the file ordered. delivery of cbs files took a long time in comparison with delivery of the remaining files. delivery of a cbs file can take well over seven months on the average; delivery of the remaining files averaged little more than a month. twothirds of the "remaining" files were delivered within a month; delivery for the rest of these took longer than three months. only 21% of the cbs files were at their destination within four weeks. most of the cbs files were delivered within three to twelve months; 21% did not arrive for more than a year. figure 1: use of data files non-contract res. contract research other 20 m 32 24 76 u ^mm^mmsiiums^ | 63 4m 6 j 70 60 other i :•:•:•:•:: 1 total iassist quarterly problems after delivery problems involving costs, privacy, etc. come up before delivery. the questionnaire included a number of questions about the sorts of problems that came up after files were delivered. of the cbs files, 60% caused researchers problems after delivery (see figure 2). this applied to 40% of the "remaining" files. the cbs files caused more problems for users with respect to cleaning and completeness than the remaining files: 25% as opposed to 19% and 13% as opposed to 5% respectively. widely diverging scores with respect to these important problems in particular. those who use cbs files find almost every aspect more problematic than those who use other files. 4. on the other hand, cbs scores relatively high on familiarity. archiving a data infrastructure that functions well presupposes facilities not only for efficient delivery of files but also figure 2: problems with files after delivery no problems bad cleaning bad documentation incomplete file physical damage — [ other don't know ] 1 10 l'i 24 17 92 20 40 cbs 60 other 80 100 total quantifying general problems at the end of the questionnaire the respondents were asked to give their opinion on six different statements describing the same number of problems. this was an attempt to discover to what degree the respondents found specific hindrances problematic. the results are given in figure 3. the data led to the following observations: 1 . almost all scores are negative. this means that all of the aspects illustrated in the statements were seen as more or less problematic. 2. the most important problems concern the availability and delivery of the files; files are too expensive, considerations of privacy hinder or prevent delivery, and when delivery does take place, it takes too long. 3. cbs file users and those who used other files had for file storage and management. the study asked what happens to files after use. remarkably, only 80% of the files were archived in any way (see figure 4). the fact that 8% of the files are returned to the supplier and 6% are destroyed may be related to the terms under which the files were delivered; a supplier can, for example, require that files be destroyed or returned after a project has been completed. nothing was done with 4% of the files. as this study focuses on files delivered by others, one may assume that the original files remained in the possession of the supplier. nonetheless, if the same practice applies for files put together by the institutions themselves, then there is an irrevocable loss of material. however, the study gives no definitive answer as to whether this fall/winter 1990 actually does occur. only 1% of the work of archiving is contracted out to others, and this although a national archive exists for storing all files that in theory may prove relevant in future social scientific research. once again, this study concerns files whose original versions are in all probability still with the supplier; copies therefore did not necessarily have to be archived by others. however, by the same token it was by no means strictly necessary to archive files locally, which is what happened in 80% of the cases anyway. for this reason the balance between "archiving" (80%) and "archiving by others" (1%) remains remarkable. the next question concerned the way in which archiving, management and registration take place. no uniform archiving method appears to exist archiving and management take place at different levels; sometimes the individual researcher does it; then again it may be left to someone within the research group, the faculty or the computer center. the lack of guidelines or agreements for storing data systematically hinders access to files that might be suitable for secondary analyses in the future. there is no clear survey of the various locations where files are available. conclusions we must first mention that this study focuses largely on the use of external data files and the problems researchers encounter in this use. the most important points of this issue are taken up in this article. suppliers of data files, on the other hand, are not given an opportunity in this study to express their grievances concerning file users. although this article says nothing about their complaints, this by no means suggests that they do not have problems with users' unclear, unreasonable or badly formulated wishes. the study has led us to formulate the following conclusions. 1 . poor response is one of the most important reasons why this study did not result in an exhaustive inventory of the use of data files in the social sciences. it is remarkable that researchers who depend so much figure 3: quantification of problems supply bad cleaning costs privacy privacy quality -0.4 t-'-ifi-miiii hi o.i -1.4 -1.4 1 1 | cbs hfffjii other i j total the respondents have given their opinion on a number of statements included in the questionnaire. scoring ranges from 1 = "agree completely" 3 = "indifferent" to 5 = "disagree completely". for clarity the results have been converted into a figure whose minimum score is equal to "problematic" and whose maximum score of 2 is equal to "not problematic". iassist quarterly on the cooperation of others in their work could respond so poorly. 2. the various institutions and research groups within the social sciences have not kept systematic records concerning the use of external data files. perception of how files are used is diffuse. in this respect, there is a lack of order in the way social scientific research is organized and conducted. in no sense does it fit the image of a professional, well-oiled machine. the data highways see very little orderly traffic. 3. cbs is actually the most constant factor on the data file market of supply and demand. as the most important supplier of data files, cbs has, firstly, a wide range of files to offer and secondly a number of clearly stated delivery terms. this gives cbs a welldefined policy concerning the delivery of data files, making it the cornerstone of data facilities. the demand side, or rather the research world, has little to offer in return. as mentioned before, it is a diffuse and unorganized field in which mutual interests and viewpoints are difficult to formulate. this unequal situation hinders coordination between the parties concerning delivery terms for the files. 4. the information provided by cbs makes clear that a demand with great purchasing power exists for a select number of files in the social sciences. cbs's files are better known than files offered by other suppliers. the study also confirms what has been observed in the literature: that researchers find the limited availability of microfiles due to cost, considerations of privacy and delayed delivery problematic. these problems are particularly in evidence with respect to cbs files. 5. the demand for steinmetz archive files is quite specific. in 67% of the cases, the steinmetz archive delivers files containing voting research data. remarkably, almost half of the dutch files go to foreign countries, and an important portion of the files used by dutch researchers are from abroad. 6. the survey reveals that 70% of the research is on a contract basis, and that almost all of this is financed by the state. given that the state also finances the remaining research either directly or indirectly, we can conclude that the state subsidizes almost all social scientific research for which data files are purchased. 7. the fact that most of the data files are used in contract research, together with the observation that non-contract research is much less likely to use cbs figure 4: archiving 20 40 60 in percentages 80 100 cbs other i total fall/winter 1990 data than data from other files, indicates the difficulty, if not impossibility, of acquiring costly files given the limited financial resources of the various research groups. 8. because most of the files are ordered for contract research, it may seem acceptable to pass the costs of the files on to the users. however, it is also important to note that while certain costs may be included in the price of the file, other costs should not be. 9. local storage and management of data files is conducted at various levels. there are no guidelines or agreements concerning systematic data storage. this hinders access to files that might theoretically be suitable for secondary analysis. recommendations technical innovations have made the processing of machine-readable statistical files increasingly important for social scientific research, a development requiring reflection upon the supply of files and the use of new technical facilities. 1 . this reflection can take shape in the form of coordination between the suppliers and consumers of data files. segers (1990) has already observed that the research world is too unorganized to set up a structural dialogue with suppliers; it must first gather its forces and negotiate with suppliers over new delivery terms and a broader range of user options. segers recommended establishing an independent organization to represent the research world. as part of its task this institution would coordinate the desires and wishes of the users on one side and cbs on the other. on the basis of this coordination a number of standard agreements could be formulated. 2. clearly cbs must play a vital role in improving the data infrastructure. its cooperation is necessary in implementing a number of possible adjustments concerning: first of all, the delivery policy of cbs. too much time is lost bargaining over terms of delivery. delivery of microfiles is often laborious, partly due to the disclosure risk; as a result, delivery is expensive and irregular. prices increase becarse files have to be custom-made every time. both researchers and cbs would be better off if they switched to making standard deliveries of the most important data, perhaps in the interim stage in the form of publication files. these would be "stripped" microfiles. second, the cost of microfiles. besides the risk of disclosure, this is the most important barrier to delivery of files. the standard delivery of data mentioned above, in the form of inexpensive publication files, may help to solve this problem, but the cost of microfiles must be reconsidered further. the main idea would be to generate a buyers' demand within non-contract research; in other words, to take a number of measures that will facilitate the acquisition of files by university research groups. the recommendation made by de guchteneire & timmermans to establish a central budget for the purchase of data files might serve as a point of departure. third, the use of other media than what has been used to date. this involves creating facilities destined for long-term use, for example cd-rom. this initiative clearly supports the standard delivery of important data. cbs has assumed that for the time being such media are more suitable for the dissemination of aggregated data than for the distribution of individual level data, for example all statistical data concerning municipalities. mention must be made here that university libraries will no doubt be unavoidably drawn into this issue. when libraries purchase and make such information available, it becomes available to a much wider audience of users. 3. data infrastructure would profit by improved "signposting". users or potential users must have better information about which files are to be found where, and how and under what conditions they can be acquired. at the moment most researchers do not know which paths to follow to get this kind of information; often they may not even know that such paths exist. consequently, policy should aim at improving information on user options; in other words, documentation, etc. with respect to finding and acquiring the files. 4. researchers should have a service organization available to them for information and assistance in finding and acquiring data files. in the united states in particular an increasing number of universities are beginning to set up and develop local data libraries. in most cases these libraries still form part of a "regular" library. such libraries perform the following tasks: * systematically storing used data files; * entering bibliographical information on these files into a computerized catalogue; * providing services to researchers or students, including: searching for particular files using the files analyzing the files. besides offering advice, other administrative duties can be included as part of the service: performing the tasks mentioned above, iassist quarterly finding and retrieving data files present elsewhere " upon request. 5. de guchteneire & timmermans proposed establishing a service organization which would perform the following tasks: * gathering and developing expertise directed at making sensitive material anonymous12 * developing techniques through which sensitive material can still be made available (for example, by working with test files), * advising researchers. to these tasks may be added: * designing sound rules concerning liability and other organizational terms concerning the safety of the files. as in point 4, the issue here is service to researchers. we advise describing the possible functions and tasks of a service organization in a follow-up project, and exploring which of these tasks should or could be performed locally or nationally. 6. to summarize, in shaping a new or adapted data infra-structure attention must be given to the following points: * organization and representation from the field; * conditions for making data files available; * the way in which the files are made available; * guidelines for data file storage and management; * service for researchers and students. an important consideration in describing the infrastructure is the current and future position of the remaining supply agencies, in particular the steinmetz archive. inconclusion technical innovations have made it increasingly possible to produce computerized data files, make them available and analyze them. data can be analyzed more often and from various perspectives. for this reason, we recommend that cbs files be made available without many restrictions to as broad a sector of researchers as possible, and that different types of analysis be conducted by various institutions. because cbs produces statistics, it must process the collected data in a number of ways. there are so many different manipulations that can be performed on files that it would be impracticable for cbs to carry out them all. this is not to say that cbs would no longer have to perform manipulations, on the contrary. however, microfiles should be made more readily available to a larger group of researchers. wider delivery of cbs files would serve not only the interests of researchers, but of society at large. adequate construction and efficient use of the data infrastructure requires the involvement of builders, managers and users. technology on the one hand and sound agreements on the other must lead to the development of an infrastructure that allows researchers to make high-speed use of the available "data highways." at present, however, unexpected roadblocks and no trespassing signs have been thrown up. regulations concerning availability must be amended in order to give data files the right of way. the interests of all parties involved must be kept in mind. it is time to take measures and come to agreements: a well functioning infrastructure is a necessity in a modem and democratic society. the respondents have given their opinion on a number of statements included in the questionnaire. scoring ranges from 1= "agree completely," 3= "indifferent" to 5= "disagree completely." for clarity the results have been converted into a figure whose minimum score of -2 is equal to "problematic" and whose maximum score of 2 is equal to "not problematic." bibliography bie, s.e. de, wetenschap en statistiek. standpunt van de vereniging van onderzoek instituten inzake de levering van cbs microdata ten behoeve van sociaal wetenschaopeliik onderzoek . leiden, voi, 1988. guchteneire, p.f.a. and j.g.m. timmermans. wetenschappeliik gebruik van overheidsdatabestanden . amsterdam, swidoc, 1990. segers, j., de data-infrastruktuur in nederland: snelwee of zandpad? presented during the "statistics day" conference, 1990. steinmetz archive, data catalogue & guide . amsterdam, 1990. 1 presented at the iassist '90 conference held in poughkeepsie, ny, united states, may 30-june 2, 1990. 'these files were set up for statistical analysis and can theoretically be used for planning and policy. as opposed to data in an administrative file, data stored in a statistical file is not altered again except for aggregations and scale or index constructions. for example, data included in the "population statistics" file undergoes no further change once they have been collected, whereas data in an administrative file such as the register of births, marriages and deaths is altered whenever a resident whose personal information is included in the file moves or receives a new passport or driver's license. in statistical files, just as in administrative files, informafall/winter 1990 tion is often collected at the level of the individual. in the analysis of statistical files, however, the object is manipulations at the aggregate or case level. data from administrative files, in contrast, is used at the individual level (de guchteneire & timmermans, 1990), as the above example makes clear. 3 the social sciences encompass a wide range of fields, and the borders are not always sharply drawn. even the dutch term "gamma sciences" does not solve the border issue. in the context of this study the following categories were used: economics (microeconomics, macroeconomics, business economics and econometry) and business administration; sociology; psychology; education; political science; public and policy administration; environmental planning, geography and related studies. 4 these are files whose data was not collected at the initiative of the institution itself, but were acquired or purchased from others. not included are files whose data was collected by the institution's own field workers; was collected by others at the request of the institution. the files had to meet the following criteria: * the data was in a file; the study specifically did dflt focus on tables that were delivered; * the files were stored in such a way (tape, diskette, cd-rom) that they could be accessed by computer (that is, as a datamatrix); * the files were in the possession of the researcher, research group or institution; on-line external files were not included in the study; * microfiles (data at the level of persons or households) had to include at least 1000 research units; mesofiles or macrofiles (data at the level of companies, institutions, sectors or countries) needed at least 100 research units; * the files were used for research in 1989; * the files were used for research in the social sciences. 6 in publication files the research data, anonymous or not, is reworked in such a way that disclosure of individuals is almost entirely ruled out. microfiles are also anonymous data files at the individual (or household) level, but they include such detailed information on variables that by making intelligent combinations, disclosure may be possible. the ethical code maintained by researchers does not permit such disclosures. microfiles are only supplied under the terms of an agreement; publication files are offered on the open market. unlike publication files, microfiles are supplied neither to government agencies serving a public administrative function, nor to business, but only to research institutions, meaning universities and the planning bureaus. 7 for example, the response states that ten deliveries of a particular file yielded 375,000 guilders. it is not clear whether this is two large deliveries yielding 300,000 guilders and eight others yielding 75,000 in all, or another combination. 8 cbs was unable to meet our request for a survey of purchase prices of the most complete files. because negotiations are underway concerning policy adjustments with respect to data file availability, cbs found it an inopportune moment to provide price information which would become outdated in the near future. 'the weekly questionnaires were offered 226 times in all. these were 226 different questionnaires. "these are files delivered by the steinmetz archive. 11 various developments in the united states have taken place in close collaboration with the international consortium for political and social research (icpsr). 12cbs comments here that "anonymous" is hardly a viable term. it is not only a matter of sensitive data files but also of data that cannot be identified at the individual level, and of non-disclosure, according the cbs. the bureau states that a high level of expertise has been focused on this issue and that the bureau, affiliated institutions outside the netherlands and certain foreign research institutions are continuously developing greater expertise. iassist quarterly iassist quarterly 3 public data in use: a case study of ireland by john blackwell' resource and environmental policy centre. university college, dublin introduction this paper is concerned with the policy issues which arise in enabling the public data which are most useful for research purposes to be disseminated in a timely fashion. case study material from ireland is drawn on, outlining problems which have arisen in identifying and in meeting the needs of users, and the extent to which it is possible to meet the requirements at a lime when there is no prospect of an increase in real government expenditure on the provision of statistics. particular examples in the areas of earnings, employment and social 'presented at lassist/ifdo international conference may 1985, amsterdam welfare statistics will be given. the extent to which administrative records can be a fruitful source of data, and can substitute for purpose-built surveys and censuses, especially in view of the cost constraints, is assessed. at the end of the paper, some policy issues are posed. the background data provision in ireland in ireland there is a centralized central statistics oftice (cso), whose output ranges from population and vital statistics to labour force, employment and unemployment statistics, industrial production, prices, earnings and hours worked in industry, trade, transpon and distribution. many of these statistics are produced by means of censuses and surveys. the methods of dissemination used by the cso vary from published volumes to regular mimeographed series (issued usually monthly or quarteriy); data on ireland which appear only in publications of international organizations such as eurosiat and oecd; small area data such as from the census of population and from the census of agriculture which are disseminated directly to the users; and special analyses which the cso does in answer to specific requests e.g., from the labour force survey and the household budget survey, which are usually not charged for. however, in recent years there has been an increase in the amount of data available from public sector bodies outside the cso, which in the main complement the existing range of data from the cso. the monetary and public expenditure data, produced by the central bank and the department of finance respectively, can be put with the national accounts as the basic raw material for macroeconomic analysis. much of the information on education and manpower fiows and science and technology, on activities such as transport, energy, and tourism, on fdl/winter 1985 4 iassist quarterly education, health, housing: and social services, and on taxation arc produced by a variety of bodies in the public sector outside the cso. in a number of cases, these data are produced by government departments (e.g. education and social welfare); in other cases the data comes from government agencies or state bodies (e.g., in the case of science and technology statistics). problems of resource allocation there has been no increase in the volume of resources provided to cso in recent years, rather, a decline. while there is no coherent account of the amounts of resources which are put into provision of statistics outside cso, these are unlikely to have increased. the budgetary problems in ireland rule out any prospects of an increase in resources for statistics. indeed, one of the elements which worked to the advantage of data users in ireland in recent years was ireland's entry into the european economic community. this meant that eec regulations and directives on statistics became binding. this provided, for example, a much needed set of statistics on the distribution of earnings tlirough the 1979 structure of earnings survey. as a reflection of eec budgetary pressures, a proposed 1986 survey will not now take place, and there is no domestic survey which can fill the gap. especially at a time of pressure on the government budget some means of ranking the statistical output by reference to the intensity of demand, is required. this is very difficult to achieve, for the following reasons: there is a heterogeneous body of users, ranging from government to research institutes, individual researchers and private sector firms. the implied value of attributes such as timeliness, frequency, accuracy of estimate, level of disaggregation, will vary among these groups of users. many private firms, and those in the financial sector, put a premium on timeliness. hence, as a result of timeliness problems in the case of the irish census of industrial productions (problems which are currently being righted, slowly), there has been minimal use of industrial statistics by business firms. there are other reasons for this neglect of industrial statistics on the part of the private sector. firms will typically want highly disaggregated data in order to answer questions such as market share or productivity change for a particular product in a country such as ireland, where one or two firms may dominate some sub-sectors, there often there can be no possibility that cso would be able to provide the data for reasons of confidentiality. hence, the popularity among business users of the foreign uade statistics, which are quite disaggregated and are more timely than production data. related to the preceding point, it is difficult to establish the nature of the "trade-offs" which exist between improvements along these different dimensions of statistics. for example, timeliness could be improved by sacrificing some accuracy of estimate. for some firms, such as those in the financial sector, timeliness is crucial. for them, it is often more important to have an early estimate which is subject to later revision, than to have a purer number which arrives after a long interval. by contrast in the case of social policy, the time-scale fall/winter 1985 iassist quarterly 5 is different here, the analysis of large data sets, gathered at longer intervals, concerned more with details at household or family level, is often needed. with regard to frequency, there is a danger that in a perverse way, an over-conceniiation on frequency could slow up decision-making, especially if estimates are subsequently revised to a marked extent in the case of national accounts data where there is already a time-lag due in part to slowness in provision of local authorities' data there has in ireland been a good deal of data revision long after the event the price mechanism is not used to ration demand so that those uses of highest value (or which use up least scarce resources) are effected. that is partly because information has many of the attributes of a "pure public good". this is, information which is made available to one person (through, for example published sources, in the widest sense) can be made available to the community as a whole, and one person's consumption of information does not "take away" from the consumption of anoiiier person. (of course, this must be qualified: information can be withheld from certain consumers and can be bought and sold on the market) moreover, a number of the statistics are provided at close to zero prices, which gives rise to difficulties in assessing the intensity of preferences among actual and potential users. in summary, therefore, a key issue is how many more, or how many fewer, resources should be devoted, at the margin, to producing particular data series. this is, in a sense, a problem of providing a surrogate for the market system in guiding resource allocation to the provision of public statistics. difficulties abound in assessing the intensity of user preferences, partly because of the various preferences of the different users. furthermore, it is difficult to get people to reveal their preferences, a refiection of the 'public good' aspect of statistics. users will be tempted to overstate their preferences, in the knowledge that if information is provided by somebody else, they in turn can benefit in other words, people can obtain information at no cost to themselves, even though, if they were confronted with the stark alternative of "pay up or do without the information" they would be prepared to pay something. the "classic" perception of the way in which information is gathered and organized has often centred on the following sequence by which decisions are made: identification of a problem; collection of data to illuminate that problem; identification of feasible solutions; choice of the best solution from the alternatives. however, in practice, this logical sequence is not always followed. information may be gathered only after a proposal has been decided on, either to monitor the implementation of policy, or to co-ordinate activities across departments, or to serve as a back-up for particular cases which a public body may wish to make. why does the classic model not apply? in pan it reflects the fluctuating pressures for decisions, often made under time pressure. there is inevitably a time lag before information can be gathered, by which time decisions may have already been made. in pan it reflects the lack of a clear set of signals from the users of data to the producers. in ireland, the cso has found it difficult to get users to articulate their requirements an unambiguous way. while there has been a generalized feeling of dissatisfaction among users, there has been no formalized method whereby the reactions of users can be channelled to the producers. at mid-1984, for the first lime, a one-day statistical users' seminar was held in order to help communications between users and producers. even here, there was a disappointing turnout fall/winter 1985 6 iassisl quarterly from private sector users. cso does consult with major users, but this consultation gives little weight to non-institutional users. it is almost trite to observe that the demand for statistics is a derived demand. yet it should be noted that these derived demands are governed by the ecomic structure and by the way in which policy analysis and research occurs. thus, for example, most social science research in ireland occurs in large research institutes. in the absence of a funding body for social science research, and given a weak tradition in team research in the social sciences, the countervailing power of the lone researcher, not to speak of ability to use large data sets, may be weak. indeed, for some uses such as macro-economic models, the data supply/demand position is more or less one of one monopoly producer (cso) and one or two large users such as the economic and social research institute. another example is the use, in other countries with more decentralized of government, of sub-national statistics to govern the spatial allocation of public resources: with the degree of centralization in ireland, this use is virtually absent a final example: the more rapid the pace of social change (such as women's participation in the labour force), the greater the likelihood that the full censuses will (again, given current methods) lag behind what would be needed to monitor change. this has implications for the potential use of sample surveys or else of simpler censuses. there will tend to be a chronic excess of user demands over the supply of statistics (given existing resources and methods). the question arises whether it would be possible to decide on allocating resources to statistics by using formal criteria such as cost-benefit analysis. the answer is "no". while it is possible to estimate costs, the diffuse nature of the uses of statistics would make the estimation of benefits extremely difficult except in a few particular cases, the monetary value of statistics will be almost impossible to compute. moreover, given the links between information, policy and "background" knowledge of the worid, the potential returns from statistics may be greater than the actual returns if there is potential to improve the linkage between public policy-making and the provision of statistics. this can be illustrated by reference to the irish census of population. the cancellation of the 1976 census had some direct effects on other data provision it led to increasingly inaccurate estimates of population and of the labour force and consequently of employment and of national output at the same time, there seems to be almost no accounting for the individual decisions, including those in physical planning, which have been affected. but the measure of these effects would be difterent with a difterent type of public administration, one that was more responsive to changes in economic and social information. in this sense there are interaction effects which complicate any attempt to ascertain a pure return to statistics. it is difficult in fact to point to particular policies in ireland which were critically infiuenced by the provision or absence of certain data. there arc a number of reasons for this. first the provision of data per se is best seen as contributing to the understanding and explanation of economic change and social processes. second, even if a starkly empirical view were (mistakenly) taken that the matter begins and ends with gathering "the facts" the dissemination of data would work its effects in a diffused way, with time lags. third, there are usually time pressures on policy-makers research findings and associated data must be speedily available, and this is not always possible. fourth, it is difficult to predict data needs ahead of lime the focus of policy attention tends to change over time, while the provision of statistics lags behind. finally, there are insufficient contacts between researchers and policy-makers (including lack of effective dissemination of findings) and there is certainly too little interaction between the three groups of fall/winter 1985 iassist quarterly 7 policy-makers, researchers and producers of data. however, it is possible to get a feel for the value of statistics by asking in any particular case: what are these statistics for? who is using them? in anwering these questions, due account should be taken of the nature of the benefits, many of which are indirect, rather than relating to specific uses by those who decide on public policy. and a number of benefits relate to users outside government in particular, there is need to concentrate on decisions at the margin what additional information should be provided, and what should be dispensed with? this would enable the identification of those data which are most crucial to policy decisions. it would also enable the recognition of those pieces of information which are redundant, or where a lowering in frequency or a degeneration in timeliness could be tolerated, especially if this meant that resources would be freed to devote to the provision of information for which there would be a high return. in ireland it has, however, been extremely difficult to identify "redundant" data series. this may in part reflect a system which is producing the "minimum critical mass" of statistics. it may also reflect a "ratchet effect" whereby, once a series is put in place, any attempts to discard it are met with opposition. scope for co-ordination and rationalizaton while the collection of statistics by a wide variety of public bodies outside the cso has filled a number of gaps in knowledge, it has given rise to a number of problems. no explicit co-ordination of this output of statistics occurs, nor is there any indication of the amount of resources which go into this activity. there is no "statistics budget" for the public sector as a whole, with an associated set of outputs. cso itself has no "leverage" over the gathering of statistics by other public sector bodies. the activities of cso could be greatly influenced by the data gathering of other bodies, either because they complemented the work of the cso, or because the cso could use the data for purposes such as national accounts compilation, or because data could be used either as a means of verification, a sampling frame, or a register for coding purposes. there is no quality control with regard to non-cso statistical outpul this has in part reflected the way in which data collection began sometimes as a by-product of filling internal plaiming needs and the need for co-ordination, sometimes as a result of implementing administrative schemes. unpublished data are available from an employment survey conducted annual by the industrial development authority, from a quarterly consumer survey undertaken by the agricultural institute for the eec, and in an information system on public sector employment in the department of the public service. the first and last cases exemplify data sets which have been compiled primarily for internal fall/winter 1985 iassisl quarterly managemenl purposes. there are no principles of access lo non-cso data. while there has been some dissemination of the industrial development authority and department of the public service data, this has been as a by-product of secondary analysis by "outside" workers. the industrial development authority data have been used as a source for information which, in the absence of cso industrial employment data on a "component of change" (job gains, job losses) basis, cannot be obtained from other sources. they could, in part, be substituted by cso data if the basic cso data were analysed in a form which facilitated the "component of change" work. government departments hold a great deal of information which is of interest to policy analysts, some of it in the form of raw data. a certain furtiveness in even disclosing the presence of data seems to be endemic. the ultimate banier is when a department or agency does not release a body of data. for example, twice since 1973, the results of housing conditions surveys, undertaken on behalf of the department of the environment have been kept within the department this is a notable case, because the information would complement the census data on housing and, in the areas of unfitness and obsolescence, provide information for policy which is more relevant than could ever be obtained from the census. while there are many ad hoc surveys conducted by departments and other bodies, surveys conducted by the economic and social research institute for the institute work, and surveys conducted by other research institutes and by individual researchers, none of this body of material is pulled together in one place. nor is it even catalogued. there is no central user's guide to the body of data. there is danger that the current round of computerization which is occurring in public sector bodies will result in further fragmentation. there is, for example, a danger that computerization among local authorities will miss out on the possibility of providing the means of a data bank. to take another example: the department of the public service is to expand its personnel information system. ideally, in this process, the cso should be consulted and compatible codes for occupations and other entities should be used. this leads to the issue of whether there would be potential for rationalizing the data-gathering on irish by state-sponsored bodies. a number of these bodie^ national training agency, the irish export board, the industrial development authority, the institute for industrial research and standards, and the national board for science and technology compile data on irish firms, as does the cso itself. on the surface, rationalization could contribute to the problems of non-response and of slowness in response, which in part account for the lime lags in the issuing of cso industrial statistics. this may be the case up to a point, but it has to be qualified in view of the client relationship which exists between many of these firms and the stale body in quesuon, which is likely to facilitate response. a more compelling reason for seeking some rationalization is the richer potential for cross-tabulation which would ensue, as at the moment this is limited to the variables on each individual inquiry. moreover, there would arise a common set of standards with regard to methods of compilation and of analysis, which would ideally be in accord with cso categories as regards coding. there would, of course, remain a sub-set of data which would be confidential to individual agencies. fall/winter j 985 iassisl quarterly 9 some co-ordinating role is needed in relation to the output of statistics trom sources other than the central statistics office. there are a number of ways in which the desired co-ordination of non-cso sources could occur. one would be through expanding the responsibilities of cso to include the co-ordination of all public statistics. another would be a co-ordinator of non-cso. the co-ordinator could be in a government department, or outside the departmental system, or could work through an inter-departmental committee. the main objectives of the co-ordinating function would be as follows: this could occur: through an "umbrella" survey, or through the linking of records. the latter would be much more difficult to achieve. some move towards an "umbrella" survey would be the most fruitful route. to consider the use of these statistics as a means of verfication of sampling frames or registers for coding purposes. to take part in a revised and explicit system of resource allocation for statistics, infiuencing the allocation of resources and pointing up the areas where attention needs to be focused. to encourage public sector bodies to consider the statistical possibilities of administrative records and to consult on this with bodies such as the cso. this is especially important at a time when new computer facilities are being installed. to ensure that statistics which are of importance to policy-makers but which are of little or no importance to the collecting agency, are produced. this may at times require that such a subordinate objective be given explicitly to the agency. to ensure quality control, i.e., some minimum standards and compatibility between individual data sets. the more standard classifications and definitions are used, the lesser the likely burden on respondents as their records would be in conformity with all requirements. to examine the potential for rationalization in the data-gathering activities of state bodies. this could lead to richer possibilities of data analysis, as it might be possible across the headings which are contained in the existing surveys. there are two ways by which to fulfill some gatekeeping functions with regard to the increasing volume of surveys of irish firms by those in the public sector. among the cases which are candidates for action are manpower and energy. another instance is the road vehicle file of the department of the environment which is computerized but which is not being used to produce data on the vehicle fieet by age, make, type of fuel used. perhaps the prime candidates are the department of the public service data on public sector employment (mentioned above) and revenue commissioner's data derived as a by-product of tax collection. this leads to the next section, on the potential for using administrative records. using administrative records in part because of the budgetary problems which face us, the use of administrative records as sources of statistics is an attractive option. limited resources can be devoted to the fall/winter 1985 10 iassist quarterly gathering of stalisiics by surveys; siatislics based on administiaiive records can be collected at relatively low cost, in contrast with surveys. in addition, there has been increasing resistance on the part of respondents to provide information to the statistical office. (indeed, this was the main reason the pre-1968 survey on the distribution of industrial earnings by the cso was dropped). in addition, income data derived from administrative records have potential for being more accurate than income data derived from surveys, which are subject both to sampling enor and to understatement of income. this potential is particulariy evident in the case of the service sector where relatively little data are available on earnings and employment even for the public sector, where the only readily available information is on civil service employment the potential for mak-ing greater use of administrative records as a data source, or as a basis for sampling, will increase as existing records become more computerized, thereby enhancing the potential for cross-tabulations. a number of departments and agencies are now apprised of the potential for decentralized collection and dissemination of data via micro-computers. the idealized picture, thus, is one where the contraints on resources which cunently bind us are loosened through a combination of the use of administrative records allied with computer collection and analysis. there are a number of potential snags which could disturb this pleasant sequence of events. first, the definitions and categories used in administrative records refiect the administrative process (e.g., entitlement of benefits, being in the tax net). they may not be ideal for statistical analysis, and many change over time. ideally, statisticians would be involved at an early stage, either when new systems are being put in place or when manual records are computerized. one feels apprehensive that the current wave of computerization could lead to a rash of incompatible variables and codes. experience to date with some of non-cso sources-with their lack of compatibility with cso data and changes over time in categories which impede time series analysis does not inspire confidence that this can be avoided without some policy initiatives. the department of social welfare data, which have a good deal of potential for providing data on social welfare recipients, are tied to benefit recipiency. hence, it is difficult to explore problems of lack of take-up of benefit and in the case of unemployment data, in ireland (as in other eec countries) a substantial minority of the unemployed (from the labour force survey) receive neither unemployment benefit nor unemployment assistance. second, much of policy analysis requires linkages either to cso data or to departmental data. an example would be the relationship between social welfare payments and patterns of work. such linkages are either quite difficult to effect or can raise problems of confidentiality. third, the only aspects of administrative records which are reliable are those which are regularly used by the authority in question and are, therefore, kept accurate and up to date. considerable clerical resources may be required to maintain administrative records for statistical analysis purposes (i.e., keeping designatory details up to date, elimination of duplication, etc.) and authorities have other priorities in mind. if administrative records are to be fruitful, they need constant attention. important administrative details required for analytic purposes (e.g., designatory details such as business descriptions, occupation, age) are not usually complete nor up-to-date because they are not frequently used by the authority. this is the biggest handicap to the use of these records, and arises in particular in the case of the irish revenue commissioners data which have the potential to provide earnings and employment data and to fill notable gaps in coverage, especially on the service sector. fall/winter 1985 iassist quarterly 11 fourth, there is need for close coordination between siausticians and administrators in order that the needs of both administration and provision of statistics are mel decisions have to be made about classifying data in cases where there may well be conflicts between the desires of different users. at times when schemes are designed, changed or computerized, there is need to be sensitive to the possibilities of using these records to provide statistical tabulations. this would mean consulting with statisticians in the design stage. if consultation were to occur too late, a costly redesign of systems could be required. in the past in the case of the irish health sector, where there is a certain ceniralizaton of the functions of health boards, computer systems which were incompatible across health boards were acquired. while there has since been a rethink of information requirements within the department of health, this has involved too little contact with those with a planning function in government fifth, a major constraint to using the revenue commissioners' data relates to objectives. this is linked both to the issue of resources and the need to use clerical resources. understandably, the revenue commissioners regard their main purpose as getting in money from the tax system and would regard information provision as, at best a subsidiary objective, or at worst having no jusdficaiion at all. in the case of a tax with high collection costs per yield, such as the residential property tax, the minimum amount of information would be provided. from the point of view of the public weal, and from the point of view of cost-minimization for the state as a whole, there is need to use the revenue commissioners data. but there is no incentive for the revenue commissioners themselves to take account of this. rather the opposite: if people are switched from revenue-raising in order to provide more information, the revenue commissioners will be in danger of receiving criticism for failure to maximize tax revenue. if the basic data were "clean", it would take litue resources to produce the required data, by using the automatic data processing facilities which revenue has. the biggest obstacle to producing tables which would help policy-making outside the specific remit of revenue is the need to deploy clerical workers to "clean" data with regard to designatory codes. this would especially be the case as revenue is prepared to consider the provision of a subset of basic data on magnetic tape to outside agencies with the individual identifying characteristics removed. users could then analyse these data using either a software package or a custom-built program. this would be a particularly valuable facility. there will of course, always remain areas of inquiry where there is no substitute for survey or census data. an example arises in the case of labour force statistics. it is manifest that existing data are often inadequate to answer the analytical and policy questions which arise. this is partly because most of the labour force data are given in terms of stocks. many of the questions raised need an analysis of flows between states. another reason for difficulty is that existing data come, in part from administrative records which are tied to benefit criteria, which differ over time, and in part from self-description of states which are inherently ambiguous, such as "unable to work due to disability". an instance of this from the irish labour force survey is that in 1977 a category "unable to work due to permanent sickness or disability" was introduced, but led some people who described themselves as "unemployed" in 1975 to opt for this new category. there is an interplay between disability, unemployment and disability, and between retirement and disability. if one is to capture the impact of social security payments on the labour market there is need to complement the existing stock data with fiow data. fall/winter 1985 12 iassist quarterly ideally, flow data (shown in table 1) by sex would be required. in this table, the intersection of a row and a column gives the flow between states. from this table, a table showing transition rates (for, say, one or two years) between labour force stales could be constructed. this would show entry rates into a particular state from all types of states, together with exit rates, and continuation rates (the proportions of people who remain in a particular state between one period and another). in ireland there is already one survey which in part can be used to provide flow data: an annual survey of school-leavers undertaken for the department of labour. this is an example of a survey which would provide the answer to a number of policy questions if the basic data were accessible. policy issues in summary some of the policy issues which arise are: there is need for some form of coordination (with no coercion*^) of non-cso data with cso data, in order to ensure quality control, i.e.. some minimum standards and compatibility between individual data sets. there is also scope for some gatekeeping functions with regard to the increasing volume of surveys, including surveys of irish firms, by those in the public sector. and there could be some rationalization of the existing surveys of irish firms, leading to richer possibilities for data analysis. should some minimum standards be set in relation to data which are collected as a by-product of government contracts, including their deposit in a central location? what forms of surrogates for the market system to identify user demands ("needs"?) can be developed e.g.. some mechanism for feedback and complaints, for wide consultation and receipt of suggestions at an early stage of the design of inquiries? for example, could there be greater use of charges in order to get an indication of the implicit value which different users put on different types of statistics? what forms of users' guide are required for the existing mass of data which public bodies hold? and what ground rules on access to public sector data by outside users should be established? there is need for public sector bodies to consult with cso on the statistical possibilities of using administrative records. related to this, how can the danger that the increased diffusion of computers among public bodies, including the local authorities, will lead to the onset of incompatible data sets, which are in turn incompatible with cso data, be avoided? despite the above remarks on the potential for using administrative records, there are many issues of economic and social policy where there is no substitute for survey or census statistics at the level of the individual or of the household, especially where it comes to establishing relationships between variables. this is especially true in view of the increasing fall/winter j985 iassist quarterly 13 incidence of "non-standard" households such as lone-parent families and the many family policy issues which result to what extent could existing surveys be used with additional questions being asked at marginal cost? in ireland, there are many possible uses of a general household survey with a central base of questions and a varying sub-set of questions from year to year. this could provide information on education and on use of health services, for example. there is no possibility of adding to the household budget survey (i.e., family expenditure type), which is already overloaded. to what extent should a simpler census of population be conducted, with a shorter time lag for dissemination, leaving matters such as housing to sample surveys? to what extent should one try to force the development of intermediaries who would take primary industrial data and engage in "packaging" and marketing to industrial customers; are there economies of scale in this activity? underlying the paper as a whole, there are the difficulties in deciding how much public resources to put into sialislics. for example, non-cso data will reflect producer preferences as to internal management needs or inter-departmental and inter-agency coordination: this may not be optimal from the point of view of society as a whole, n schcmalit outline ol lypc of labour markcl how dala whicli arc required initial state final state (numbered as in rows) li) lii) 2 3a) 4 5 6 7 1. employed a. full time b. pan time 2. unemployed 3. in partial retirement a. in the same job as when last fully employed b. in a different job 4. fully retired 5. in labour force, sick 6. unable to work due to disability 7. not in labour force due to "home duues" 1 fall/winter 1985 vol29-3.indd iassist quarterly fall 2005 by by bobray bordelon* cross-national & intergovernmental data: paying for one-stop shopping introduction many intergovernmental and governmental organizations produce their own interfaces for statistical data. with the exception of the common database which combines select data from many of the united nations’ independent agencies, there is typically little cooperation for integrating data between agencies. some universities have chosen to produce their own interface to merge data into a single source. one example is the economic and social data service from the united kingdom. one may also rely on commercial sources. there are various questions to consider. is a researcher likely to want to combine sources from various organizations? is it logical to do so? is the native interface good enough to use on its own and does it is easily allow the combining of data from various sources? are the data from the original producer easily extractable? does the original producer allow alteration of the data? does your organization have the staff and expertise to design and maintain its own interface? does your organization have the expertise to match changing data elements over time? is there a commercial service, already in existence, that already does what you want? if so, how do the short-term and long-term costs compare? how would you treat documentation? how does the commercial vendor treat documentation? commercial aggregators since most organizations lack the resources to build its own interface and provide the ongoing maintenance, the choices for a commercial aggregator will be explored. there are many elements to investigate when considering a commercial solution: • variety of data (due to cost, we want one source to cover as many organizations as possible; ideally it should be international and cover all countries and territories) • source of data should always be specified • fixed beginning date or rolling period • start date (should be easy to determine when each series starts) • frequency of data (daily, weekly, monthly, quarterly, annual) • update frequency • downloading capabilities and format (continuous series; text, tab, or comma delimited; excel; sas, spss, stata) • documentation (online, direct links, paper) • interface/ease in finding data • medium of storage and delivery (internet, web, cd, tape, terminal based) • cost • customer help (toll free phone number, responsive e-mail, online chat) • lease or purchase analyzing the aggregators four major aggregators will be examined: datastream international, eiu world data, global financial database, and global insight. 6 iassist quarterly fall 2005 datastream eiu world data global financial database global insight individual country data yes yes yes limited sub-national very little limited no yes for usa imf direction of trade statistics yes limited no yes imf balance of payment statistics no no no yes imf international financial statistics yes limited no yes imf government finance statistics no no no no oecd historical statistics/ national accounts yes limited no yes forecasts yes yes no yes but for an extra fee exchange rates yes yes yes yes interest rates yes yes yes yes platform client based web web client based/web soon documentation limited/extranet incorporated incorporated plus data encyclopedia limited extranet for actual series plus more from source for select others (equivalent of a codebook) another issue to consider is what metadata are provided. datastream provides the start date, frequency, unit and scale of measurement, adjustment factor, conversion method, and original source. global insight and the eiu provide similar information with the eiu also providing the name of the analyst. global financial database provides the most detail and its de� history of an indicator. what is missing? the sources tend to omit actual definitions of the terms being used. the eiu is the only one of the four sources to provide definitions. for the novice trying to choose between several choices, a hyperlink or mouse-over could provide much assistance. it would be useful to provide the prefaces to the works being used or the methodological documentation. for example, including international financial statistics country notes, balance of payments manual, and guide to direction of trade statistics would assist the researcher in understanding inconsistencies and subtle distinctions in series. what do we call an item? even when the same source is used, the vendor may not use the original terminology. the international monetary fund tries to standardize terminology between nations by having a consistent vocabulary. since items are sometimes measured quite differently, the imf tries to bring together similar concepts by grouping items by line number. for example, line 14a of international financial statistics represents the amount of money in circulation outside of banks. the united kingdom, iassist quarterly fall 2005 7 japan, brazil, and canada refer to this as “reserve money of which currency outside dmbs (deposit money banks)”. the united states calls it “reserve money of which currency outside banks”. the euro nations, switzerland, and sweden refer to it as “currency issued” while denmark defines it as “reserve money of which outside bis”. datastream refers to it as “currency in circulation” or “currency outside banks”. global insight uses four terms: “currency in circulation outside dmb”, “currency outside banks”, “currency outside banking institutions”, and “currency in circulation”. if this was not confusing enough, one would have to turn to the country notes and not the international financial statistics yearbook to determine what constitutes a “dmb” in each nation. a few examples are qatar considers dmbs to be locally owned commercial banks (including islamic banks) and branches of foreign banks; canada includes only chartered banks; and georgia only commercial banks. in comoros, it is one specific bank - banque pour l’industrie et le commerce-comores. the confusion continues when a source is listed but the notes provide contrasting information. for example, the eiu provides a definition of m1 for the united states and cites its source as international financial statistics. however a note states the data uses national concepts rather than imf concepts. throw in similar series having different base dates and one begins to understand why researchers are confused and crave simplicity. what documentation should be provided? at minimum the following should always be listed: • start date • frequency • update schedule • formats available for downloading • unit of measure • scale • adjustment factors (if any) • status (active/discontinued) • conversion method (if applicable) • source • name of analyst for forecast • definition • links to old data if discontinued • comparability charts how can we help researchers? talk to vendors about what is needed. research institutions help pay the bills. use the collective power of groups such as iassist and the american library association to persuade vendors to improve their products. provide guides that point out common mistakes and fallacies. long ago, princeton university provided a guide for datastream when the company itself was providing no documentation. this guide not only provided search examples but also pointed the user to where to look for items and the many misnomers of the program. provide training to staff and researchers. have the actual documentation available at a common reference point. while many of the sources are standard monographs that you may not have room for in a reference collection, have the basic guides to documentation and try to keep major statistical works on site. while the user accessing this information off site will not have immediate access, the serious researcher would know where to turn. perhaps most importantly, we can know the sources’ advantages and limitations and be able to guide the researcher to their appropriate use. * bobray bordelon is the pliny fisk librarian of economics and finance/data services librarian at princeton university. bordelon@princeton.edu. the article is based on a presentation delivered at the 31st iassist conference in edinburgh scotland on may 25, 2005. mailto:bordelon@princeton.edu vol223 fall 1998 17 introduction the rapid expansion in electronic communications and commerce over the past several years has raised concerns in the united states over personal privacy in an online environment. these concerns have captured the attention of the public, the media, and policy-makers, and there is new interest in the united states in explicit policies protecting the privacy of electronic transactions and personal information. these efforts continue a pattern of policies directed at subject-specific information, such as the national education statistics act of 1994 that tightened access to personal data collected in the field of education [1]. this pattern is a sharp contrast to the privacy and data protection polices in europe. where the u.s. approach has been to provide specific and narrowly applicable legislation, in europe there are unified supra-national policies for the region. most countries have implemented these policies with omnibus legislation. the european legislation outlines a set of rights and principle for the treatment of personal data, without regard to whether the data is held in the public or private sector. in the united states, the legal tradition is much more concerned with regulating data collected by the federal government. this paper will review and contrast the development of data protection policies in the united states and europe. what is privacy? privacy is an important, but illusive concept in law. the right to privacy is acknowledged in several broad-based international agreements. article 12 of the universal declaration of human rights and article 17 of the united nations international covenant on civil and political rights both state that, “no one shall be subjected to arbitrary interference with his privacy, family, home or correspondence, nor to attacks upon his honour and reputation. everyone has the right to the protection of the law against such interference or attacks.” the international concept of “the right to privacy” traces its roots to the u.s. constitution and to common law [2]. a hallmark article in the harvard law review in 1890 is widely credited as establishing the right to privacy as a tradition of common law [3]. in that article, samuel warren and louis brandeis defined that right as “the right to be let alone” [4]. they argued that the right to privacy that afforded to intellectual and artistic property in common law is founded, not on principle of protection of private property, but on that of “inviolate personality” [5]. the term “privacy” does not appear in the u.s. constitution or the bill of rights. however, the u.s. supreme court has ruled in favor of various privacy interests-deriving the right to privacy from the first, third, fourth, fifth, ninth, and fourteenth amendments to the constitution. in 1977 in whalen v. roe, the supreme court first recognized the right to information privacy [6]. it noted that the constitution protected two kinds of individual interests: “one is the individual interest in avoiding disclosure of personal matters, and another is the interest in independence in making certain kinds of important decisions” [7]. several other decisions have balanced the right to privacy against other compelling interests. the supreme court upheld a new york law that required the state to maintain computerized records of prescriptions for certain drugs because the program did not pose “a sufficiently grievous threat” [8]. in nixon v. administrators of general services, the court upheld the federal statute that required national archivists to examine written and recorded information accumulated by the president [9]. the court ruled that while “the appellant has a legitimate expectation of privacy in his personal communications,” that right must be weighed against the important public interest in preservation of materials [10]. the court did not believe that the appellant’s privacy interest was a match for the competing public interest [11]. u.s. data protection laws there is no single law in the united states that provides a comprehensive treatment of data protection or privacy issues. in addition to the constitutional interpretations provided by the courts and the international agreements mentioned above, there have been a number of laws and executive orders dealing specifically with the concept of data protection. the most important and broad based of these laws are the privacy act of 1974 and the computer matching and privacy act. these laws deal exclusively with personal information held by the federal government data protection and privacy in the united states and europe by jean slemmons stratford & juri stratford * 18 iassist quarterly and do not have any authority over the collection and use of personal information held by other private and public sector entities. the privacy act (pl 93-579) is a companion to and extension of the freedom of information act (foia) of 1966. foia was primarily intended to provide access to government information. it did exempt the disclosure of personnel and medical files that would constitute “a clearly unwarranted invasion of personal privacy” [12]. this provision was initially used to deny access to people requesting their own records. so the privacy act was also adopted both to protect personal information in federal databases and to provide individuals with certain rights over information contained in those databases. the act has been characterized as “the centerpiece of u.s. privacy law affecting government record-keeping” [13]. the act was developed explicitly to address the problems posed by electronic technologies and personal records systems and covers the vast majority of personal records systems maintained by the federal government. the act set forth some basic principles of “fair information practice,” and provided individuals with the right of access to information about themselves and the right to challenge the contents of records. it requires that personal information may only be disclosed with the individual’s consent or for purposes announced in advance. the act also requires federal agencies to publish an annual list of systems maintained by the agency that contain personal information. the law had originally proposed the creation of a privacy protection commission; however, then president gerald ford was opposed to such a bureaucracy. he wrote i do not favor establishing a separate commission or board bureaucracy empowered to define privacy in its own terms and to second-guess citizens and agencies. i vastly prefer an approach which makes federal agencies fully and publicly accountable for legally-mandated privacy protections and which gives the individual adequate legal remedies to enforce what he deems to be his own best privacy interests [14]. as a compromise, central oversight was assigned to the office of management and budget, and omb has exercised relatively weak leadership in the implementation of the privacy act. the law also calls for the designation of privacy act officers within federal executive agencies to handle requests and insure compliance with the code of practice. ultimately enforcement rests with the courts (as individuals bring suit to redress perceived grievances). under the umbrella of the privacy act, congress has also enacted the computer matching and privacy protection act of 1988 (pl 100-503). this act amended the privacy act by adding new provisions regulating the use of computer matching. computer matching is the computerized comparison of information about an individual for the purpose of determining eligibility for federal benefit programs, or for the purpose of recouping payments or delinquent debts under such programs. in general, matching programs involving federal records must be conducted under an agreement between the source and recipient agencies. this agreement describes the purpose and procedures for the matching and establishes protections for the matched records. the agreement is subject to review by a data integrity board and each agency involved in matching activities must establish such a board. while the law provides no special access rights to individuals; agencies must notify individuals of any findings based upon a computer matching program before taking any adverse actions; and individuals must be given the opportunity to contest such findings. the computer security act of 1987 (pl 100-235) also deals with personal information in federal record systems. it protects the security of sensitive personal information in federal computer systems. the act establishes governmentwide standards for computer security and assigns responsibility for those standards to the national institute of standards. the law also requires federal agencies to identify systems containing sensitive personal information and to develop security plans for those systems. narrowly applicable laws there are also numerous narrowly applicable laws on privacy and data protection. these laws generally fall into two distinct categories. the first governs the status of information held by the federal government. in general, these laws provide declarations regarding the confidentiality of specific types of personal information, provide guidelines for their disclosure and penalties for infringement of the individual’s right to privacy. as examples, 13 u.s.c. 9 absolutely prohibits any use of personally identifiable data from the census except by sworn officers and employees of the census bureau. similarly 42 u.s.c. 242m protects against the disclosure of personal information gathered by the national centers for health services research and for health statistics for research purposes. the national education statistics act (pl 103-382) re-authorized and amended provisions for the national center for educational statistics and the national assessment of educational progress. the act dramatically revised the confidentiality and dissemination practices of the center. the tax reform act (pl 94-455) makes tax returns and return information confidential, permits only limited disclosure of returns and returns information for specific purposes, and specifies procedures for disclosure. the law also authorizes persons whose tax returns or return information is disclosed in violation of this act to bring a civil action for damages and costs of the action, and establishes criminal penalties for wrongful disclosures. fall 1998 19 the united states has largely avoided legislation governing the treatment of sensitive personal information in records systems held by sources other than the federal government. the few laws that deal with these systems tend to address the treatment of personal financial information. for example, the fair credit reporting act (90-321) regulates the use of individual personal and financial information by consumer credit reporting agencies. it assures that information is accurate and complete, relevant to the purpose for which it is used, and upholds the individual’s right to privacy. a limited number of laws have been passed to deal with issues outside the financial arena. these laws have generally been implemented in response to specific perceived abuses. as an example, the video privacy protection act of 1988 (pl 100-618) amends the federal criminal code to prohibit, with certain exceptions, the disclosure of video rental records containing personally identifiable information. it permits any person who is aggrieved by a violation of this act to bring a civil action for damages; and requires the destruction of personally identifiable records within a specified period of time. this law was passed in the wake of criticism following the release and publication of robert bork’s video rental records, during his consideration as a nominee to the supreme court [15]. several u.s. laws do restrict the federal government’s access to records held by other sources. until the rise of the internet, misuse of personal data held by entities other than the federal government did not command much attention from policymakers as a threat to privacy or personal liberty. however, government access to these records did seem to be a cause for concern as several laws have restricted federal access to information held in such systems. typically, agencies must obtain permission or a court order to get access to these records [16] data protection in europe there are two important supra-national policies in europe in relation to data protection. the first is the council of europe’s convention on data protection, and the second is the eu data directive. in contrast to u.s. privacy law, privacy protection in europe is addressed by omnibus legislation covering both public and private sectors the council of europe was set up after the second world war to help unite europe by fostering closer relations between the states belonging to the community, ensuring economic and social progress by common action to eliminate the barriers which divide europe … and promoting democracy on the basis of the fundamental rights recognized in the constitutions and laws of the member states and in the european convention for the protection of human rights and fundamental freedoms [17]. that convention recognizes the right to privacy as one of the fundamental human rights. the council’s concern with the processing of personal information grew slowly with advances in information technology and the increase in the use of such data. in the late 1960s, the council’s committee of experts on human rights conducted a survey with regard to human rights and modern scientific and technological developments. it concluded that existing laws did not provide adequate protection for individuals given the developments in these areas. several other committees examined various aspects of the problem and came to similar conclusions. in 1976, the council established a committee of experts on data protection that reported its findings in early 1979 and the result was the council of europe’s convention for the protection of individuals with regard to automatic processing of personal data. the council of europe convention sets forth the data subject’s right to privacy, enumerates a series of basic principals for data, provides for transborder data flows, and calls for mutual assistance between parties to the treaty including the establishment of a consultative committee and a procedure for future amendments to the convention [18]. the commission of the european community recommended that member states ratify the council of europe convention and warned that it might introduce its own directive on the subject. when it did so, the primary purpose of the directive was to further standardize the level of protection across the community. the eu data protection directive reaffirms the principals outlined in the council of europe convention [19]. major components of the directive acknowledge the individual’s right to privacy. the directive sets standards for the treatment of personal data collected from individuals and for individuals rights of access, notification, and correction. of particular interest to the united states is the directive’s treatment of data transfers to countries outside the eu. article 25 governs the “transfer of personal data to third countries.” eu member states may transfer personal data only after determining that “the third country in question ensures an adequate level of [data] protection.” the eu shall consider the “rules of law…in the third country” to make this determination. the directive was adopted in october 1995, and called for member states to bring their national privacy laws into compliance within three years. these national laws are now going into force across europe. the absence of generic privacy legislation in the u.s. is a major concern to the eu nations and this will make determination that the u.s. ensures an adequate level of 20 iassist quarterly protection unlikely. while there are concerted efforts in the administration calling for privacy legislation covering various types of data (e.g. secretary of health and human services shalala made recommendations to congress on the confidentiality of individually-identifiable health information on september 11, 1997,) the large number of bills in congress dealing with privacy issues suggests that the u.s. may continue to take a piece-meal approach to privacy legislation [20]. however, the eu is unlikely to issue an across-the-board finding that u.s. privacy protections are inadequate. the eu could demonstrate its seriousness about the directive by initially singling out one or more u.s. companies or sectors as not meeting the adequacy test; e.g. any company handling personal medical information [21]. given policy traditions in the u.s., it is likely that data protection in the private sector will be largely selfregulatory. the federal trade commission has been working with the private sector to develop voluntary codes of conduct, but it is unclear where these efforts will lead. it is difficult to say whether the eu will be able to recognize such an approach as adequate. if the eu decides that the largely self-regulatory approach followed by the u.s. is not sufficient to justify an adequacy finding, a much broader embargo is possible [22]. privacy and data protection are likely to continue to be big issues in u.s. domestic and international policy. it will be interesting to see how these issues will resolve themselves, or if there is to be a major clash between the u.s. and europe. footnotes [1] u.s. national education statistics act of 1994. p.l. 103-382 u.s.c., 9001-9012. [2] united nations, general assembly, 3rd session. “resolution 217a universal declaration of human rights,” 1948. the international covenant on civil and political rights was adopted by the general assembly of the united nations in resolution 2200 (xxi) of 16 december 1966. for the full text of the resolution and the covenant, see official records of the general assembly, twenty-first session, supplement no. 16 (a/6316), 49. [3] samuel d. warren and louis d. brandeis, “the right to privacy,” harvard law review 4 (1890):193-220. [4] warren and brandeis, 193. [5] warren and brandeis, 205. [6] whalen v. roe, 429 u.s. reports (february 22, 1977), 589-604. [7] whalen v. roe, 599-600. [8] whalen v. roe, 600. [9] nixon v. administrators of general services, 433 u.s. reports (28 june 1977), 425-484. [10] nixon v. administrators of general services, 465. [11] nixon v. administrators of general services, 465. [12] freedom of information, title 5 u.s.c. 552(b) (6). [13] robert aldrich, “privacy protection law in the united states,” (ntia report 82-98) in u.s. congress. house. committee on government operations. oversight of the privacy act of 1974: hearings. 98th congress, 1st session, 7-8 june 1983, 489 (y4.g74/7:p93/11/974). [14] u.s. congress. house. committee on house administration. legislative history of the privacy act of 1974, s.3418 (public law 93-579): source book on privacy. 94th congress, 2nd session, 1976, joint committee print (y4.g74/6:l52/3). [15] priscilla regan, legislating privacy: technology, social values and public policy. (chapel hill: university of north carolina press, 1995), 199. [16] aldrich, “privacy protection,” 505-507. [17] commission of the european community. communications on the protection of individuals in relation to the processing of personal data in the community and information security, com (90)314.syn 287, 44. [18] sarah ellis and charles oppenheim. “legal issues for information professionals, part iii: data protection and the media – background to the data protection act 1984 and the ec draft directive on data protection,” journal of information science 19 (1993):85. [19] “directive 95/46/ec of the european parliament and of the council of 24 october on the protection of individuals with regard to the processing of personal data and the free movement of such data,” official journal of the european community 23 november 1995, no.l281, 31. [20] rebecca vesely. “cop-friendly approach to handling medical data,” wired news 12 (september 1997) (url http://www.wired.com/news/news/politics/story/6824.html) [21] peter b. swire and robert e. litan, avoiding a showdown over eu privacy laws, brookings policy brief, no. 29 (february 1998) (url http://www.brook.edu/comm/ policybriefs/pb029/pb29.htm) [22] ibid. * paper presented at the iassist conference, may 21, 1998, at yale university, new haven, connecticut. jean slemmons stratford, university of california, davis, and juri stratford, university of california, davis. http://www.wired.com/news/news/politics/story/6824.html vol30-3.indd iassist quarterly fall 2006 by by margaret o’neill adams* the origins and early years of iassist a personal prologue this essay is based upon my presentation on a plenary panel commemorating the 25th anniversary of the founding of iassist at the iassist annual conference in toronto, canada, may, 1999. my panel colleagues were carolyn geda and ekkehard mochmann; laine ruus chaired the session. all three had participated in iassist’s 1974 founding meeting, also in toronto, and remained active in iassist thereafter. although i was part of the late 1960s social science data archives community, at the time of the 1974 meeting i was “retired” – at least temporarily, and did not participate in iassist’s founding nor formative years. carolyn’s and ekkehard’s 1999 presentations drew upon their personal experiences and memory. carolyn focussed on her experiences as the chair of the ad hoc committee that created iassist and then as its first association president. ekkehard provided a european view, and brought the discussion of the evolution of the international social science data community closer to the present. laine, too, drew upon her memory of events to add her own coloring and a canadian perspective. we all spoke from rough outlines or notes, which we promised to convert to text. in the time since i have tinkered with my material, planning to finalize it for the iassist quarterly (iq). the iq editor, karsten boye rasmussen, patiently inquired about an article from time to time. when he wrote that he was coming from denmark to washington, d.c. in the summer, 2005, to pick it up, i expected to be able to oblige. instead, completion became a 2006 new year’s resolution. this essay, submitted in anticipation of the 32nd anniversary of the founding of iassist, is neither as detailed nor as extensive a history as i once planned. that qualification aside, the iassist history is a tribute to all iassisters: to the pioneers no longer engaged in data services or no longer with us, to those who carry on their work and the iassist traditions today, and to those who will do so into the indefinite future. introduction preparing a historical interpretation of the founding and evolution of iassist offers a chance to relive its early days vicariously and to celebrate the founding and subsequent contributions of the iassist community. it also provides an opportunity to survey the milieu from which iassist emerged and to document efforts of international cooperation in the collection, processing, storage, retrieval, exchange, and use of machine-readable social science data. looking back more than three decades on experiences of the early international data services community can also, perhaps, contextualize contemporary digital library and archives challenges, issues, and initiatives. maybe it can contribute to defining the unique professional identity of its multi-disciplinary members. hindsight often is 20/20. it seems to be human nature to minimize past challenges met when considering those that loom and are still to be resolved. the accomplishments of the international data community over the last three decades are not necessarily well known. to those unaware of the history of the social science data community, expectations for identifying, preserving, and providing access to the valuable portion of the immense outpouring of digital materials produced by the tools of office automation, or by digitization, are “new” challenges. yet from the perspective of past as well as present achievements managing social science data, and from subsequent accomplishments exploiting contemporary technologies to describe, disseminate, and preserve them, today’s challenges may be seen as simply the latest variation at the intersection of computing technology and the social sciences. valuable lessons learned by the pioneers who preserved, described, and provided access to computer-readable structured data can be applied to the archival challenges of continuing and new forms of digital information. the story of how the pioneers came together to collaborate in an international organization may yield contemporarily useful insights. historical sources two sets of materials, the iassist archives and the full run of social science information (ssi), offered plentiful primary material for considering iassist’s past and some of its early accomplishments. from 1962-1987, the international social science council (issc), with support from the united nations educational, social, and cultural 6 iassist quarterly fall 2006 organization (unesco), published ssi. ssi became a vehicle for reporting on the work of the issc as well as for publishing articles from interdisciplinary social science research the iassist archives rest for now in a traditional environment. they are paper documents foldered in approximately 30 hollinger boxes, the acid-free containers that traditional archives use to store paper records. they include an almost complete set of the early years of the iassist newsletter/quarterly, a significant volume of correspondence between the founders, meeting minutes, and other administrative material. the author’s ready access to the archival material provided her the opportunity to attempt this history. all of this essay’s references to unpublished materials are from the iassist archives, unless otherwise cited. elina almasy, the long-time editor of ssi and for many years the executive director of the issc, supported the quest for additional background materials by kindly offering access to all the volumes of ssi when the author visited her in paris in the spring, 1999. numerous ssi articles document post-world war ii international and interdisciplinary collaboration in the social sciences, including efforts to coordinate resources and support services. some of the earliest explicitly discuss the need for data libraries or data archives to support international, interdisciplinary, and comparative social science research.[1] mme. almasy, who was secretary of the issc’s standing committee on social science data archives at the time of the founding of iassist, also generously shared with the author some of her recollections of the personalities and milieu that contributed to iassist. one of the oft-told stories of iassist’s origins revolves around “the meeting in the bar.” thus one of the first pieces of evidence sought in the archives was any documentation or reference to such a meeting. the author also hoped that the archives would offer evidence of the role that michael t. aiken played in facilitating the emergence of iassist. aiken, professor of sociology at the university of wisconsin, madison (uw) in 1974 and eight years earlier, the founding faculty director of uw’s [social science] data and program library service (dpls), had been active in the work of the issc’s standing committee on social science data archives. he appears to have been largely responsible for planning a conference on data and program library services that was held in conjunction with the world sociology congress meeting in toronto in late august 1974. iassist emerged from that conference.[2] earlier in august 1974, mike had shared a copy of the conference “call” and general program outline with the author and enthusiastically told her of his hopes that this conference could be the catalyst for the formation of an association for the newly emerging profession of social science data archivists and librarians. they were the women and men who were developing data support services in numerous institutions around the world and establishing standards for managing and sharing computerreadable social science data. an internationally recognized association might solidify the professional status of data services personnel, something that was becoming especially important in north america in an era of increasing professionalization that was in part a by-product of new affirmative action and equal opportunity initiatives. in addition, an international association of social science data services professionals could further collaboration in social science research. that long had been a goal of the issc’s standing committee on social science data archives and of the by-then defunct u.s. council of social science data archives (cssda).[3] with a cutback in u.s. government funding, cssda met an early demise in 1969-1970. its final activity was co-sponsorship with the university of wisconsin of a national workshop on the management of a data and program library in madison, wisconsin in june 1969, and the subsequent publication of the proceedings from the workshop.[4] the archives the iassist archives do document the “meeting in the bar” and they verify michael aiken’s role in the formation of iassist. they also document much more, well beyond what can be drawn upon in a single essay. it therefore focuses solely on the formative years. before considering the specifics of iassist’s history, some general impressions gleaned from skimming the archival material offer insight about the environment from which iassist emerged and suggest the degree to which its accomplishments are so noteworthy. for example, on a mundane level, the archives include few photocopies but many onion skin carbons. the latter offer evidence of the way correspondents circulated multiple copies of letters and document drafts across continents. somewhat surprisingly by contemporary standards, there is much traditional formality in these materials, despite the collegiality and familiarity of the writers. one frequently finds original typewriting whose quality caused self-consciousness on the part of the authors. on the other hand, imagine the effort that judith rowe was reflecting upon when she wrote to carolyn geda on september 12, 1974, barely three weeks after the toronto meeting: “enclosed are the two drafts typed with my own hands. what greater love has anyone?” those drafts contrasted with many of the others, and contained no strike-overs, no typos. one was a three-page general memo concerning the organization of assist [sic], the association for social science information services technology. it presumably was intended for widespread distribution, while the second was a memo to members of assist’s ad hoc organizing committee. these drafts iassist quarterly fall 2006 7 were the first of many that judith composed in the early months, and that were revised many times by alice robbin, per nielsen, carolyn geda, and sometimes others, as well as by judith herself. they formed the beginning of a hefty collection of paper documents that carolyn gathered for a mailing to the ad hoc organizing committee later in the fall, 1974. the authors of the early iassist documentation comment frequently about delays in circulating correspondence within north america, and between the u.s., canada, western europe, and other parts of the world. there is no suggestion that the authors considered using international telephone lines to speed communication. questions of currency equivalencies and how they might apply to any effort to set a standard membership fee, the challenges of hedging any organizational bank accounts against possible currency deflations, and the problems of the varied financial costs of seemingly comparable activities from one country or continent to another are among the persistent organizational themes. the archives offer a window on the vicissitudes of the academic employment scene of the time, and the vagaries of institutional support for both international and national social science data infrastructures. you don’t have to read between the lines to understand that dealing with all of this took a toll on the abilities of individuals to do the organizing work they had promised. the reader also observes a group of dedicated professionals trying to build a worldwide network of achievement and support, apparently having no question about the language by which they would communicate among themselves: english. many of the concerns of the early iassisters are issues that still preoccupy the international data community. yet in terms of basic communication, and especially the technologies that support it, the environment was wholly different than that of today. while we all know this and understand on many different levels what this means, there is something about all those carbon paper copies that serves to symbolize just how much has changed, societally and professionally, since 1974. any foray into the iassist archives leaves the reader exhausted by the volume of correspondence and other written material produced in those early iassist years. what is especially remarkable was that most of it was authored by only four individuals: judith rowe of princeton university’s data library, alice robbin of uw’s data and program library service, carolyn geda of the inter-university consortium for political research (icpr) at the university of michigan, and the late per nielsen of the danish data archives. while many others contributed throughout the early period, those four were the primary actors, and by their unique commitment, enthusiasm, and tenacity, assured the emergence of a formal international membership organization to address the common problems facing all who were organizing and staffing social science data archives or libraries. we’re getting ahead of the story. background for the toronto meeting, 1974 reading between the lines in the materials reviewed for this essay, and recalling the author’s conversation with mike aiken shortly before the toronto meeting, it seems that organizing a conference on data archives and program library services concurrent with the 1974 world sociology congress in toronto was the way aiken found to advance his long-term goals. he had connections with at least three professional groups that shared his interests in influencing the future of international social science research and support services for it. his networks included the international profession of sociologists, his colleagues in issc’s standing committee on social science data archives, and the staffs of the emerging informal network of social science data services providers. it is fairly clear that aiken, and likely others, sought to use the toronto sociologists’ meeting, issc connections, and personal involvement in the nascent “data community” to bring issues of common interests to a joint gathering of social science researchers and data services providers. the international social science council (issc), formed in 1952, was an interdisciplinary nongovernmental organization (ngo) supported by unesco. by the early 1970s, stein rokkan, a norwegian social scientist (and founder of the norwegian social science data services, the nsd), was president of the issc. rokkan had been working since the late 1950s, if not earlier, on international efforts to facilitate access to social science data for comparative cross-national and cross-cultural analysis. he and his colleagues had identified the tools essential for access to data for comparative analysis. they included data inventories, archives of raw survey data, a current file of information on progress in cross-national and cross-cultural research, guides for research workers in need of data for other countries, standardization and manuals of information on existing and proposed data classification standards for the social sciences, and regional working conferences and seminars of senior social scientists as well as some for graduate students.[5] during the 1960s the issc sponsored several conferences on social science data archives, most of which focussed on european issues, with one in the u.s. hosted by yale university. rokkan, among others, viewed unesco, through the efforts of the issc, as the logical agent for internationalizing social science research and for making truly comparative research a reality.[6] as a result of a recommendation of the third european social science data archives conference in 1965, the issc decided at a conference in london in april 1966 to constitute a standing 8 iassist quarterly fall 2006 committee on social science data archives (scssda). its first meeting was in the u.s., at ann arbor, michigan in june the same year. by fiscal year 1967-68, unesco allocated [u.s.] $6000 for the standing committee’s work. the original standing committee members included stein rokkan, erwin k. scheuch of germany, john madge of england, and three other europeans. ralph l. bisco and warren miller of icpr, william glaser of columbia university, philip hastings of the roper center, and ithiel de sola pool of the massachusetts institute of technology formed the u.s. contingent. glaser and bisco wore two hats on the standing committee as they were also, respectively, chairman and executive director of the u.s. national science foundation (nsf)-financed council of social science data archives (cssda). over time, membership on the standing committee expanded to include other european and north american social scientists, as well as members from south america, east asia, south asia, and the ussr.[7] the issc’s standing committee and the u.s. cssda emerged at approximately the same time and for the years of their co-existence, jointly focused upon developing data inventories. none was produced before the demise of the cssda in 1969. cssda did, however, sponsor several national meetings and in 1967 published a directory of social science data archives in the united states, covering the activities of twenty-three u.s. data organizations and one in canada, and included an entry for the cssda itself.[8] both issc’s standing committee and the cssda focused primarily on the interests of social science researchers. professionals from the data services community who were not themselves researchers played no role in issc’s standing committee, though some participated in the work of the cssda.[9] the evolution of the standing committee on social science data archives continued as the cssda faded. it became more structured, recasting itself into task forces as the vehicles through which its work would be done. by the time of its 1968 meeting, the issc’s standing committee had seven such groups: task forces on data inventories and data retrieval, archives management, program library services, archives of aggregate data, historical data archives, training in secondary analysis, and archives development. the standing committee also decided that year that resolving policies on the privacy and confidentiality of “archived records” should be the joint concern of all its task forces.[10 in 1970 michael aiken became chairman of the program library services task force. this likely reflected not only his interests, but also a new project at the university of wisconsin: a nsf-sponsored national program library service. the 1974 toronto conference the formal convenors of the 1974 conference on data archives and program library services, august 21-22, 1974, sponsored by the issc’s standing committee on social science data archives were its chair, erwin scheuch, director of the zentralarchiv, cologne; michael aiken; and hagen stegemann, also of the zentralarchiv. these three also were the coordinator and discussants, respectively, of the final summary session of the conference. an overview of the conference program is reprinted here as appendix a. it shows the conference considering some of the topics that were a responsibility of one or more of the standing committee’s task forces. in addition, the program introduced some themes that reflected issues of importance to professionals who were supporting social science research, rather than undertaking it. uw’s data and program library service sent more than 300 invitations to the conference. session coordinators, program speakers, and the approximately 60 registrants came from the social science data services community, from the existing membership of the standing committee’s task forces, and other interested social scientists. this mix of people differed from participants in previous issc activities, wherein social science researchers prevailed. the conference program addressed issues that both complemented and expanded upon challenges the issc and the former cssda had confronted. note the similarity to the themes and issues of iassist conferences over the years. alice robbin’s account of the final session, included in the reprinted program overview, summarized the action areas agreed to by conference participants. taken together, they led to recognition that meeting the challenges of facilitating the kind of international and interdisciplinary social science research collaboration that stein rokkan and others had long advocated required moving beyond the organizational status quo. the group identified professionalization and training of data archivists, the people on whose work social science research depended, as the first means of accomplishing their goals. more than 30 years later, clarifying the professional status and complementary relationship between data archivists or librarians and those engaged in social science research iassist quarterly fall 2006 9 can sometimes seem elusive. compounding this is the challenge of integrating data services into the mainstream of the traditional archives and libraries as they embrace a digital era. robbin recalls that at the summary session of the conference, david nasatir, then director of the international data library and reference service at the university of california, berkeley, recommended that the group adjourn and continue discussing the idea of a new organization “down in the bar.”[11] the idea of a “meeting in the bar” thus can be credited to nasatir, embellishing further the reputation he earned for his many contributions to the social science data community.[12] there are several documentary references for “the meeting in the bar” to corroborate the memories of this happening. the first is a letter of september 10, 1974, less than three weeks after the conference, from alice robbin to carolyn geda and judith rowe. alice sent news that stein rokkan, president of the issc, had already written to mike aiken to say he was delighted with what he heard about the toronto meeting, which he had been unable to attend. rokkan urged that “programme points,” i.e., an outline of mandate, function, goals, etc. for the new organization be formulated quickly, so he could submit them to unesco. he evidently emphasized that unesco funding might be available, given the right approach. alice adds a “p.s.” to her letter, in which she refers to the “notes that carolyn and judith took at the long bar on thursday night.” she ends by suggesting that the newly-forming group may want to think of a name change.[13] the second reference to “the bar” comes in a letter from david nasatir to carolyn geda dated a day later, september 11, 1974, in which he writes, “recognizing that mike aiken will have to fill in, etc., i thought i would set down what i understand of the results of our meeting in the bar, the meeting the next morning, and the standing committee on social science data archives meeting on friday night, and possible implications for iassist.” apparently some shift in emphasis or support developed after the conference summary session and the meeting in the bar and before or during the standing committee’s meeting a day or so later. at a planning meeting for iassist the morning after the gathering in the bar, there had been tentative proposals that iassist might associate in some manner with the issc’s standing committee. one suggestion was that a chair and vice-chair of iassist serve on issc’s standing committee, alongside its task force chairs. however, it seems not everyone on the standing committee supported this overture or the idea of a joining of iassist and the standing committee in this manner. nasatir’s letter recommends that “everyone. . .should keep her cool” and remember a few things. iassist is [to be] an independent, autonomous organization that can (and should) do as its members please. they may wish to have working groups, for example. . . .[secondly] the [issc] standing committee on social science data archives. . .has an established structure. . . . [thirdly] nothing prevents (in fact everything supports) some or all of the working groups of iassist becoming task forces of the standing committee, without [its members] ceasing to exist as iassist members. [he continues:] . . .on the morning after the meeting in the bar, a group met to form iassist and agreed to form a provisional organizing committee with you [carolyn] as the chair, per as the european secretary, . . .and we decided to work towards a goal of an organizational meeting to be held in 1975. we hoped to seek funds for this, but the mechanism for funding remains vague in my mind. aiken wrote a few weeks later to carolyn geda, confirming the information in nasatir’s letter (david had sent him a copy), and said he was writing with hopes of clarifying the situation. he, or he and nasatir, had evidently reported to the standing committee about the organizing plans of “international assist,” and had there expressed the desire of the new group to be an independent professional organization, while recognizing that it probably would gain greater “legitimacy in the international social science [research] community immediately [if it had] association with the [issc] standing committee.” he expands a bit: “international assist is a stand alone organization with autonomous status, unless its membership decides otherwise. it is an association of professionals in the data archives field who will define projects of mutual concern, set up task forces to carry out these objectives, and we hope, be able to obtain sufficient resources to have national and international meetings from time to time.” he continues: “regarding its relationship to the [issc] standing committee on social science data archives... one way to maintain both its autonomy and to participate in the activities of the standing committee is for the task force chairmen of international assist also to occupy the position of task force chairmen in the standing committee.” he also mentions that when it met in toronto after the conference, the standing committee formed one new task force, on a worldwide data archives register, that was “eminently fundable” by the issc. he implied that such an effort would be of interest to international assist. nonetheless, it is likely that even at this early point, the potential for competition between the scssda and iassist for issc and/or unesco funds added to the tension related to the issue of organizational autonomy. post-toronto the stage was thus set for the flurry of organizing activity that consumed carolyn, judith, alice, per and colleagues 10 iassist quarterly fall 2006 for the next many months and years. overall it is a story of many small and not so small dramas, well beyond the scope of this essay. the organizing tensions evident in the early correspondence needed to be resolved in a manner that the interested parties could support should iassist affiliate with the issc’s standing committee on social science data archives? doing so might mean eligibility for unesco support. should it affiliate with other professional organizations? should iassist be a federation of regional associations or some other structure? perhaps the key decision, from the perspective of establishing the professional identity of data archivists and librarians, was whether membership was to be of individual professionals or institution-based. ultimately, when the iassist constitution was adopted a couple of years later and made clear that its membership was of individuals, the international federation of data organizations (ifdo) emerged to facilitate the types of collaboration that its organizers felt required institutional support.[14] a few organizing efforts offer a glimpse of the range and variety of work undertaken in iassist’s first years. the first was assembling the package of materials that carolyn geda sent to the iassist ad hoc organizing committee. mailed in early december 1974, it included five parts: • a list of most of the major meetings related to social science data archives that had been held between 1962 and 1969, including reports or agendas from some of them; • two sample constitutions from other professional associations that might help iassist draft its constitution; • the summary of the meetings on iassist in toronto; • a questionnaire on the interests of the committee members, including their suggestions for persons for the mailing list that the ad hoc organizing committee was building; and, • a list of suggested newsletters and publications which might be asked to print a notice publicizing iassist. while carolyn was busy those first couple of months with this mailing and the numerous other details that fell to the chair of an organizing committee, the archives provide evidence that judith rowe was also busy doing what she seemed to do without ceasing throughout her career. she was working every angle, spreading the word about the new organization and seeking allies in and among numerous other associations. for example, there is a letter dated september 19, 1974 from the canadian secretary/ treasurer of wapor (world association for public opinion research), yvan corbeil, responding to a letter he had received from irving crespi of the gallup organization in princeton about “the idea” proposed by judith rowe. he calls “the idea” interesting, though he “isn’t sure of the exact nature of the association she is planning.” however, he raises a concern that the two associations potentially could be competing for the same members, so proposes an alternative. iassist could become a research group within wapor. clearly he missed the point because being subsumed in this manner was not what the iassist founders had in mind. it is, however, representative of some of the response they encountered as they began to organize. for her part, alice robbin was contributing to the narratives about the toronto meeting, including authoring the article in ssi excerpted here as appendix a. she also put the growing mailing list into machine-readable form, corresponded with some of the others who had been at the toronto meeting, and during 1976 and 1977, edited the first volume of the iassist newsletter.[15] as european secretariat, per nielsen shouldered responsibility for rounding up commitment from “interested parties,” as he called them, from throughout europe. to some extent his efforts evolved in parallel with an informal coming together in the spring, 1976, of european data archives in what became cessda: committee of european social science data archives. his iassist organizing challenges also related to the continuing tensions that emerged in toronto between issc’s standing committee and iassist and are reflected in a letter he received from erwin scheuch, dated june 7, 1976, on the subject of the future structure for the issc standing committee. it echoed, in part, a letter scheuch had sent geda in january 1975 following the mailing described above to iassist’s ad hoc organizing committee. after discussing his interpretation of the differences between iassist and the standing committee, scheuch proposed some possible areas for collaboration, such as on a world registry of data archives and related data sources, and in the development of procedures for data handling in survey archives. while negotiating all of the above, per also responded at length to the various draft documents and other mailings that came from the north americans, and urged what he called “the aspect of the third world.” in addition, given that most of the sentiment for an individual-membership organization, which he personally seemed to support, was coming from north america, while the sentiment for an association of institutional members tended to be preferred by some of the europeans, the archives suggest that per was the point person for ameliorating the differences in these perspectives. both carolyn and per drafted constitutions for iassist for discussion with the ad hoc organizing committee, and subsequently its steering committee, and the constitution which was eventually adopted later in 1976 significantly merged their drafts. iassist quarterly fall 2006 11 the evolution of iassist evidence of organizational progress for iassist is shown in appendix b, a reprint from the first issue of the iassist newsletter (november 1976). it lists the original seventeen members of the iassist steering committee. the latter replaced the ad hoc organizing committee following meetings in april 1975 in london in conjunction with the meeting of the european consortium for political research (ecpr), and in august 1976, in edinburgh during the international political science association meeting. the list provides evidence of commitment to iassist from a very heterogeneous international social science community. the identification of six regional secretariats suggests the ambitious manner in which the organizers sought to be truly international. members of the steering committee came from thirteen different countries! iassist as an organization, and its members individually, can be credited with numerous contributions in the fields of archives and library services and also in interdisciplinary social science research. they occurred as iassisters participated in the programs of numerous allied professional associations, collaborated in data-related standards-setting efforts at both national and international levels, and through the work of the iassist action groups. the closing glimpse in this essay of the work undertaken by iassist in the early years is thus of its initial action groups. the fourth issue of the iassist newsletter, dated fall, 1977, included a list of all the initial groups and their respective canadian, european, and u.s. chairpersons. there were seven action groups, giving further evidence of the scope and ambition of the young iassist. they were: classification [of data], data acquisition, data archive development, data archive registry, data documentation, data organization and management, and process-produced data. perhaps the best known and most influential product to emerge from the early iassist years was the working manual for cataloging machine-readable data files, prepared by sue a. dodd, the u.s. chair of the classification action group. members of the action group and other iassisters tested the manual’s guidelines, using them to prepare descriptions of their holdings. this was the necessary first step toward the long-sought goal of a union catalog for data files. in 1982 the american library association (ala) published dodd’s manual, cataloging machine-readable data files: an interpretive manual and subsequently chose it as their prize book of the year. the classification action group, the last operational action group from the formative iassist years, disbanded shortly thereafter. conclusion in its first decade iassist succeeded in becoming an autonomous, vibrant, and productive association for professionals of the international community of social science data services. from the outset, iassist has been a forum for “early-adopting” collaboration between social science researchers and data services and information technology professionals. iassist’s continuing embrace, over time, of ever-new technologies for its own communications and organizational efforts, together with its workshops, annual meetings, and publications, have facilitated widespread adoption of innovations and standards for the management of collections of social science data.[16] the first decade’s activities represent the rich legacy from which present-day iassist has evolved. iassist’s more than 30 years of accomplishments, rooted in international professional collaboration, are a firm foundation for confidently addressing contemporary challenges at the intersection of technology and the social sciences. footnotes 1. examples are henning friis, “institutional means of collaboration between the social sciences,” social science information 1:2 (1962), pp. 5-22; and, stein rokkan, “the development of cross-national comparative research: a review of current problems and possibilities,” social science information 1:3 (1962), pp. 21-38. 2. michael t. aiken was chancellor of the university of illinois at champaign-urbana in 1999, when the author contacted him with respect to the material she was preparing for iassist’s 25th anniversary panel. when we talked, mike was modest in his recollections of his role in the founding of iassist. those who know how he worked and his commitment to the development of data services for social science teaching and research likely know of his contributions to the data services community. in addition, the author has her own personal debt to michael aiken, for he hired her in 1966 to be dpls’s founding data archivist/ librarian. 3. for background on the cssda, see william a. glaser, “note on the work of the council of social science data archives 1965-1968,” social science information 8:2 (1969), pp. 159-176. 4. the proceedings of the workshop on the management of a data and program library [june 19-20, 1969]. margaret o’neill adams, david elesh, and alice robbin, eds. madison, wi: the university of wisconsin, 1970. 5. rokkan, 1962. 6. stein rokkan, “international efforts to develop networks of data archives,” social science information 4:3 (1965), pp. 9-13. one cannot help but imagine how excited rokkan would be by the nesstar project, one of 12 iassist quarterly fall 2006 whose sponsors is the nsd, and which embodies, in an evolutionary way, his early innovative ideas. for further information, see: http://www.nesstar.org/. 7. reports on the work of the standing committee on social science data archives appeared periodically in social science information including in 6:4 (1967), pp. 179-180; 6:6 (1967), pp. 71-76; 8:2 (1969), pp. 141-145; 9:6 (1970),pp. 179-186. 8. council of social science data archives. social science data archives. 1967. 9. one person who was an exception to this was ralph l. bisco, whose participation in the issc’s standing committee devolved from his role as executive director of the u.s. cssda. bisco died in 1970. 10. “report on the fifth administrative meeting [of the sscda], social science information 8:2 (1969), p. 144. 11. alice robbin, in telephone conversation with the author,april 1999. 12. see for example, david nasatir, data archives for the social sciences: purposes, operations, and problems. paris: unesco. 1973. 13. there are several variations to the organizational name in materials from the early days of iassist: assist, international assist, i-assist, and iassist. the author has attempted to be faithful to the usage in the context under discussion, but has also reverted to the dominant usage following adoption of its constitution, i.e., iassist. 14. ekkehard mochmann’s international social science data service: scope and accessibility (report for the international social science council (issc)), cologne, 2002 includes a brief description of the distinction between ifdo and iassist on p. 9. 15. the iassist newsletter became the iassist quarterly in 1982. 16. a wealth of information related to iassist’s accomplishments can be reviewed from its website www. iassistdata.org. appendix a the conference on data archives and program library services toronto,canada august 21-22, 1974 in conjunction with the 8th world congress of sociology convened by the chairman of the international social science council’s standing committee on social science data, erwin k. scheuch, institut fur vergleichende sozialforschung, cologne; and two of its task force chairmen, michael t. aiken, department of sociology, university of wisconsin; and hagen stegemann, zentralarchiv, cologne. invitations (over 300) were sent by the data and program library service, university of wisconsin, madison. sixty people, from australia, latin and north america, india, the middle east, and europe registered for the conference. panel discussions with audience participation. each topical session was prepared by one or two coordinators, and their invited discussants. session 1: quality of the data base: the issue of data generation coordinators: carolyn geda, inter-university consortium for political research [icpr], ann arbor, mi; and frank aarebrot, the christian michelsen institute, bergen discussants: per nielsen, danish data archives, dda, copenhagen; erwin rose and mark karhausen, zentralarchiv fur empirische sozialforschung, cologne session 2: problems of inventorying data: classification schemes coordinator: ekkehard mochmann, zentralarchiv, cologne discussants: sue dodd, institute for research in social science, university of north carolina; paul r. voss, roper public opinion research center, williamstown, ma; per nielsen, danish data archives, dda, copenhagen; carolyn geda, icpr, ann arbor, mi; and paul peters, social sciences information utilization laboratory, pittsburgh, pa session 3: increasing the utilization of data archives coordinator: lorraine borman, vogelbeck computing center, northwestern univ., evanston, il discussants: tom atkinson, institute for behavioral research, york university, ontario; richard a. hay, jr., northwestern univ., evanston, il; mark karhausen and erwin rose, zentralarchiv, cologne session 4: specific problems facing the data archive coordinator: ramkrishna mukherjee, indian statistical iassist quarterly fall 2006 13 institute, calcutta sub-session 4a: the library and information center and the emerging social science data network -paul peters, social sciences information utilization laboratory, pittsburgh, pa sub-session 4b: introducing the access and use of machine-readable data to the traditional library -judith rowe, social science user services, princeton univ., princeton, nj sub-session 4c: data availability and diffusion in the third world: latin america, a case in point -manuel carvajal, latin american data bank, university of florida, gainesville, fl session 5: the problem of ownership and diffusion of data coordinator: joseph bonmariage, belgian archives for the social sciences, louvain discussants: evert ladd, social science data center, university of connecticut, storrs, ct; warren miller, icpr, ann arbor, mi; hagen stegemann, zentralarchiv, cologne session 6: exchange of information coordinator: erwin k. scheuch, institut fur vergleichende sozialforschung, cologne discussants: hagen stegemann, zentralarchiv,cologne; michael aiken, department of sociology, university of wisconsin, madison. the final session summarized the problem areas identified during the earlier sessions, suggesting topics for further work: 1. need for professionalization and training of data archivists -both upgrading of current data archive professionals and training of professionals for emerging data archives, generally and especially in third world nations; 2. confidentiality and the role of the archive: development of a position on confidential data; 3. establishment of standards for ownership and diffusion of data; 4. establishment of standards for classification of data files and data descriptors; 5. establishment of standards for text documentation of data files; 6. cooperation and exchange between traditional libraries and data archives; 7. a mechanism or organization to permit consideration of these problems in greater depth. ______________ all of the above comes directly from: alice robbin, “the conference on data archives and program library services,” social science information xiv:2, 1975, pp. 197-201. the same volume also includes an article by james c. taylor, “session on program library services,” (pp. 202-205). it discusses the session concerning computer program abstracting and information clearinghouse services that was also part of the above-described conference. 14 iassist quarterly fall 2006 appendix b iassist quarterly fall 2009 by by chiu-chuang (lu) chou 1 the good, the bad and the ugly of playing a data custodian abstract: the national survey of families and households (nsfh) is a prominent longitudinal study on family life in the united states. nsfh has been funded by the national institute of child health and human development (nichd) and national institute on aging (nia). the total amount of federal grant money for nsfh is 14.5 million dollars. three waves of surveys were conducted in 1987-1988, 19921994 and 2001-2003. according to the icpsr related literature database, there are 1,051 publications based on the nsfh data. researchers continue to use nsfh to study family living arrangements, marriage, cohabitation, fertility, parenting relations, kin contact and economic and psychological well-being. when the nsfh funding ended in 2006, the center for demography of health and aging (cdha) took over providing user support for nsfh. the decision was made because it is important to continue the stewardship of this prominent study. without the expertise of the original nsfh project team, cdha staff needs to learn the rich content of nsfh quickly to help nsfh researchers. this paper describes the challenges and rewards we have in the last few years as we play the custodial role for the complex nsfh study. also, our effective online analysis tool for disseminating nsfh data is explained. and finally, a future plan for repurposing nsfh data is presented. keywords: user support, data preservation and dissemination, national survey of families and households (nsfh), center for demography of health and aging (cdha) and better access to data for global interdisciplinary research (badgir) introduction the center for demography of health and aging (cdha) at the university of wisconsin -madison is one of fourteen p30 demography centers on aging sponsored by the national institute on aging (grant number: p30 ag17266). the national survey of families and households (nsfh) is one of five major projects associated with our center. nsfh was funded by the national institute of child health and human development (nichd) and national institute on aging (nia). the total amount of federal grant money for nsfh is 14.5 million dollars. three waves of surveys were conducted in 1987-1988, 1992-1994 and 2001-2003. nsfh data has been widely used to study family formation, marriage, cohabitation, fertility, parenting relations, kin contact and economic and psychological wellbeing in the last twenty years. when the funding for nsfh ended in 2006, cdha assumed the responsibility to maintain the nsfh project website in the spirit of good data stewardship. even though nsfh project ended, the research community continues to conduct secondary analysis on nsfh data. we knew it was important to continue user support for nsfh. however, none of our staff was involved in the nsfh project when it was active. both principal investigators are retired and all the project team members left. to carry out user support for nsfh, cdha staff dove into the nsfh documents and tried to learn nsfh as much as we can. this paper first illustrates the rich and complex content of three waves of nsfh, then continues to describe the scope and work load of our user support service, and concludes with a plan to repurpose nsfh data in the future. national survey of families and households (nsfh) american families were going through some significant changes in the 1970s and early 1980s. divorce rates were going up, more couples chose to cohabitate, more women were joining in the labor force and the fertility rate declined considerably (bumpass, 1990). all of these changes were affecting family life and household structure. the federal government had a need to understand these social changes and formulate public policies to address them. in 1983, the center for population research of the national institute of child health and human development (nichd) issued a request for proposal (ref no. nichd-dbs-83-8) for a large scale data collection to study the causes and consequences of the changes happening in the families and households in the u.s. in june of 1983, a research team consisting of larry bumpass, james sweet, maurice macdonald, sara mclanahan, annemette sorensen, and elizabeth thomson at the university of wisconsin-madison sent in a proposal to design a national study covering many aspects of family experiences and related life course events. this study would have a very broad scope and a comprehensive coverage on family living arrangements, 20 iassist quarterly fall 2009 kin contact, life histories including marriage, cohabitation, education and fertility, employment, economic and psychological well-being. the research team wanted to provide a rich data resource for comprehensive analyses of family experience from various theoretical perspectives. in january of 1984, the wisconsin proposal was accepted and the research team spent 18 months developing a basic design for the survey and possible question sequences. in june of 1985, bumpass and sweet submitted a grant proposal to the national institute of child health and human development for the implementation of the national survey of families and households (nsfh). a three-year, $4.8 million grant (r01hd021009) was awarded to undertake this cross-section survey, which was also designed to be the first round of a longitudinal study. nsfh research team had specific plans to address the limits of other large cross-sectional surveys such as, the panel study of income dynamics (psid), the national longitudinal surveys of labor market experience (nls), the survey of income and program participation (sipp) and the current population survey (cps). nsfh sample would represent the population of the entire united states and its sample size would be large enough to permit analysis of family process and structure among subpopulations. nsfh would have an exclusive focus on family issues covering a broad range of family structure, process and relationships to allow examination of the relations among them. it would cover issues important to several disciplines and sub-disciplines and permit testing of different hypotheses related to various aspects of the american family. the survey would collect respondents’ early experiences with their family and in other facets of life. such information would be the baseline for a longitudinal study of the causes and effects of family transitions and experiences. the main nsfh sample was a national, multi-stage area probability sample containing about 17,000 housing units drawn from 100 sampling areas in the 48 contiguous states in the u.s. the sample included a main crosssection sample of 9,643 households. the oversample of blacks, puerto ricans, mexican americans, single-parent families and families with stepchildren, cohabiting couples and recently married was accomplished by doubling the number of households selected within the 100 sampling areas. one adult (19 and older) was randomly selected from each household as the primary respondent. face to face person interviews were conducted either in english or spanish with portions of the interview using selfadministered questionnaires. the interview lasted about one hour and forty minutes. a shorter self-administered questionnaire was given to the spouse or cohabiting partner of the primary respondent. field work was conducted from march of 1987 to may of 1988 by the institute for survey research at temple university. there were 13,017 respondents completed wave 1 interviews. a re-interview of nsfh wave 1 respondents was conducted in 1992-1994 with a budget of 5.2 million. the sample of nsfh2 consisted of wave 1 main respondents (n=10,007), their current spouse/partner (n=5,643), their original spouse/partner (n=789) (if the respondents parted with their spouse/partner in wave 1), their focal children (n=2,495), their parents (n=3,347) and their proxy (n=802) (when the wave1 respondent was deceased or too ill to be interviewed.) the main respondents, their current spouse/partner, and their original spouse/partner were interviewed with computer assisted personal interviewing (capi) technology using laptop computers. focal children, parent and proxy interviews were done over phones with computer assisted telephone interviewing (cati) technique. this five-year follow-up was designed to collect information on life events including marriages, divorces, births, work experiences and other transitions since the first interview in 1987-1988. the consequences of family experiences, characteristics, attitudes, health, wellbeing, kinship, parenting, spousal relationship, social support, labor force participation, income sources, assets and debt enrich nsfh data for further examination on various factors influencing family dynamics, marriage, cohabitation, childbearing and marital disruption. wave 2 data combined with wave 1 data allow researchers to study union formation and dissolution with the implications of cohabitation. attitudes and behaviors from both male and female respondents provide a detailed picture on the transition from cohabitation to marriage, factors and effects of nonmarital childbearing among cohabitated couples, and differences on union stability. the third wave of interviews was conducted in 2001-2003 by the university of wisconsin-madison survey center using computer assisted telephone interview technology. due to budget constraint (4.5 million), only a subset of the original sample was interviewed. the subset consists of main respondents 45 and older by january 1, 2001 with no focal children and respondents who have a focal child, wave 1 spouses or partners of main respondents and respondents’ eligible focal children (age 18-34 in wave 3). for main respondents who were deceased or too ill to be interviewed and who did not have a spouse/partner to be interviewed, proxy interviews were conducted. the instruments for respondents and their spouses/partners were the same. the focal child’s questionnaire was shorter and very similar to that used in the older focal child interview in wave 2. the proxy interview used the same instrument from the wave 2 proxy interview. when all the interviews were completed, information on the 2nd follow-up survey was collected from 4,600 respondents, 2,677 spouses/partners, 1,952 focal children and 924 proxy interviews. information collected in wave 3 included living arrangements, histories of marriage, cohabitation, education, fertility, employment, marital and parenting iassist quarterly fall 2009 21 relationships, kin contact, and economic and psychological well-being. dissemination of nsfh data and documentation files nsfh data and document files are available from several places: nsfh project website (http://www.ssc.wisc. edu/nsfh), inter-university consortium for political and social research (icpsr) (http://www.icprs.umich.edu/), sociometics (http://www.socio.com), and better access to data for global interdisciplinary research (badgir) (http://nesstar.ssc.wisc.edu/webview/). each venue has its merits and caveats. the nsfh project website (http://www.ssc.wisc.edu/nsfh/) was created by the principal investigators to disseminate data and provide relevant information for interested users. an outline of major topics is described for each wave. so users can browse the rich nsfh content by broad categories. in addition to data, documentation files such as, sample design, methodology, codebooks, layout/dictionary files, appendices, questionnaires and skip patterns, working papers, bibliography, and faq for data users can be accessed at this website. data files from waves 1 and 2 are distributed in raw ascii format. users need to use the layout/dictionary files to write their own setup files to extract data. the wave 3 data files are in spss sav format only. users can download them to spss without writing their own setup files. however, users who prefer data in other formats are responsible for data conversion. all three waves of the nsfh data and documentation files are freely available there. this site has the most comprehensive coverage of nsfh both in data and documents. the ascii format of waves 1 and 2 data and the lack of statistical package setup files mean that users need to consult the codebook and dictionary files to create their own subsets. nsfh waves 1 and 2 were deposited with the interuniversity consortium for political and social research (icpsr) in 1994 and 1997 respectively. icpsr assigned study number 6041 to wave 1, study number 6906 to waves 1 and 2. for wave 3, icpsr provides only the study description (study # 171) and directs users to the nsfh project website for data access. icpsr provides sas and spss setup files for wave 1 data, but no sas or spss setup files are available for the wave 2 data files. the coverage of nsfh at icpsr is limited. however, users who are interested in publications based on the nsfh data can search the icpsr related literature database at this url: http://www.icpsr.umich.edu/icpsrweb/icpsr/citations/ index.jsp. sociometics has nsfh waves 1 and 2 available from their american family data archive (afda) for a fee. users can obtain sas and spss setup files for both waves of data. the sociometrics multivariate interactive data analysis system (midas) is a subscription-based online analysis tool that provides access to nsfh data and documents. they don’t have wave 3 files. the limited coverage and fee-based service might not appeal to some nsfh users. cdha data archivist, janet eisenhauer smith created an online archive called better access to data for global interdisciplinary research (badgir) in 2004. badgir is powered by the nesstar suite which implements metadata developed by the data documentation initiative (ddi). that same year, nsfh waves 1 and 2 data and documentation were marked up and published to the badgir catalog. users can browse and search the rich content of nsfh via a friendly web interface. simple statistical analyses like cross-tabulation and regression can be run interactively. custom extracts in sas, spss, stata, excel and raw ascii format can be done efficiently using data download feature. this online data tool gives users the flexibility to output nsfh data in a format they can use effectively. the number of wave 2 data files in badgir is smaller than those available from the nsfh website and icpsr. this is because when cdha staff marked up wave 2 data files, variables from the self-enumerated data files and constructed variables were combined in the main files for respondent, spouse and ex-spouse. in fall of 2008 we added waves 3 files to badgir. because we had an access to the original capi questionnaires, the wording of a question pertinent to each variable is included in the metadata description in wave 3. study methodology, sample design and other important documents are embedded in the metadata and might not be obvious to users if they access the variables directly. user support for nsfh the nsfh project included a user support component while it was funded. by august of 2006, only one graduate student was employed to provide user support. because nsfh is such a rich data source for issues related to family, researchers continue to conduct their secondary analysis of nsfh. hence, there is an ongoing need to provide user support. since cdha has nsfh in its badgir online archive, a decision was made by the center to continue user support when the nsfh project funding ran out. cdha staff also assumed the responsibility of maintaining the nsfh website and continued the geomerge service for users who need to include geographical characteristics in their analysis of nsfh data. we started to log nsfh user support questions, when cdha merged with the data and program library services (dpls) and the center for demography and ecology (cde) in january of 2007. our new organization is called the data and information service center (disc) and we used libstats to keep track of our reference services. libstats is an online utility developed by eric larson and nathan vack at the general library system in the university of wisconsin-madison. libstats classifies the extent of reference questions by the time our staff spent 22 iassist quarterly fall 2009 answering them. there are four categories: 0-9 minutes, 30 minutes, a couple hours and substantial. these categories are very simple and yet can reflect the complexity level of each reference question disc staff handles. from january 1st of 2007 to june 30th of 2010, we have answered 313 questions from nsfh users. some questions can be answered fairly quickly but some took significant time from the cdha staff. these three tables were created to illustrate types of patron who seek nsfh user support, the amount of time our staff spent to answer their questions and the format of their questions. the majority of questions arrived in nsfhhelp email account. nsfh is a prominent longitudinal study of family life and dynamics. its broad scope and national sample size make it a unique data source for research. twenty years after its first wave of data collection, research papers are still being published based on nsfh data. according to the icpsr related literature database (http://www.icpsr.umich.edu/ icpsrweb/icpsr/citations/index.jsp) on july 20, 2010, there are 1,051 publications based on nsfh data: 190 theses, 616 journal articles, 189 reports, 34 book sections, 11 conference proceedings and 11 books. such high numbers of publications clearly demonstrate the strength and richness of nsfh. the most significant contribution of cdha is to freely disseminate public-use nsfh data and documentation online via our badgir catalog. cdha staff marked up nsfh study using nesstar publisher with the metadata elements established by the data documentation initiative (ddi) in 2004. these ddi compliant data files were published in the badgir catalog which a nesstar server powers. in fall of 2008, all three waves of nsfh data have been described and documented down to the variable level in badgir. table 4 shows the number of files, cases, and variables available from badgir. there users can search any of the 26,927 variables of their interest, run cross tabulations and regression analyses, and download customized subsets in formats like sas, spss, stata, comma delimited, and spreadsheet for further analysis. such value-added enhancements gives users with various levels of experience an exceptionally easy point of access to the complex and rich nsfh data. (there is one caveat regarding sas. badgir actually creates a sas statement file and an ascii data file not a sas system file. this is because sas did not release its proprietary codes to the nesstar developers.) restricted geomerge files for nsfh users in response to users who need to include geographical characteristics in their analyses of nsfh data, we have created geo files which can be matched with users’ contextual data. to obtain geomerge data, a nsfh user first fills out a confidentiality agreement form and submits it for review by the nsfh principal investigator, larry bumpass. currently, we only offer geomerge for waves 1 and 2. the nsfh geo file consists of the following geographic units: state, zip code, state fips, county fips, place fips, mcd fips, msa fips, 1990 census tract, 1990 census block, 1990 census labor market and area from mable/geocorr v2.5 geographic correspondence engine. users choose the geographic unit and locate the contextual variables for that unit from other sources, such as the county and city data book. they then send us a copy of their contextual file in question formats walk-up email phone counts 6 274 33 table 3 question formats the majority of questions arrived in nsfhhelp email account. we do not conduct reference services using instant message or online chat. patron types researchersnon uwmadison undergraduate s uwmadison graduates uw-madison cdha affiliates communities counts 285 2 9 3 14 table 1: questions by patron types (communities include mass media and nonresearch related inquiries.) most questions came from researchers outside of uw-madison. time spent 0-9 minutes 30 minutes a couple hours substantial/more than two hours counts 132 96 17 68 table 2: time spent by cdha staff to answer users’ questions 42% of them were answered in less than 10 minutes iassist quarterly fall 2009 23 raw ascii format, a dictionary file listing the variables and their column locations, and a file with descriptive statistics for each variable in the contextual data file. my colleagues, charlie fiss and janet eisenhauser smith review these files and merge the contextual data with individual records in the restricted geo nsfh files. at the end, a geomerge file with case id and the characteristics of places of residence by the geographic unit chosen by the user but not geographical locations is sent to the user. this way, nsfh respondents’ confidentiality is protected while researchers can add geographical aspects to their analyses of nsfh data. challenges of user support one of the downsides of having data and documentation accessible down to the variable level in badgir is that it has created high expectations from nsfh users. some users are very anxious and jump too soon to the data without spending adequate time learning about the survey design, sampling and study methodology. this means that cdha staff receives many nsfh questions which could be answered by the documentation available from the website. users may miss this information when they readily access nsfh data from badgir. because none of the cdha staff members were involved with the nsfh project when it was active, we have spent hours reading its documentation on methodology, survey design, codebook and many supplementary documents to get familiar with its rich content. it is not uncommon for staff to re-read the same document many times over when we try to answer users’ questions because of the study complexities. on some occasions, we have gone to the pi for clarification and additional information. we are very careful to only provide the facts gleaned from the nsfh documentation and avoid giving advice on research methods to data users. cdha staff does not have the expertise to discuss what variables to use in users’ analyses. staff makes this clear to users who attempt to involve staff in their research most of the nsfh questions we answered are from researchers like faculty, academic staff and graduate students. the experience levels and research skills of these data users vary. graduate students who are new to secondary data analysis usually seek guidance about locating relevant sections of the codebook, questionnaires or skip patterns. more experienced users tend to ask for clarification on some coding issues for variables over wave file description number of cases number of variables 1 2 2 2 2 2 2 2 3 3 3 3 3 3 3 3 3 3 3 3 3 total main respondent 13,007 4,355 main respondent 10,005 4,887 spouse 5,624 4,887 ex-spouse 789 4,883 proxy 802 84 parent 3,347 1,192 focal child (age 10-17 1,415 179 focal child (age 18-23 1,090 717 best measures 10,211 444 main respondent and spouse combined 7,277 2,317 roster 1 (household members) 7,277 308 roster 2 (sons and daughters living elsewhere) 5,224 512 roster 3 (spouse/partner's sons/daughters living elsewhere) 1,122 243 marriage history 4,418 133 union history 4,418 177 t1t2t3 status 13,007 54 focal child interview 1,952 1,208 focal child household roster 1,670 125 focal child marriage history 1,952 55 focal child union history 1,952 97 proxy interview 924 70 97,484 26,927 table 4: file specification for nsfh data in badgir 24 iassist quarterly fall 2009 time. sometimes, a few ambitious undergraduate students have emailed us about their plans to use nsfh for a class project. when such requests arrive, staff explains to them that nsfh is probably too complicated for a class paper and suggests alternative data instead. because of budget constraint, wave 3 data lacks some important attributes. a set of constructed income variables were created in waves 1 and 2. users who use wave 3 data are disappointed to learn that there is no plan to create constructed income variables for this last wave. weight variables were not created for wave 3 files because the respondents were selected arbitrarily by their age (respondents who are older than 45 by january 21 of 2001). without weight variables, the wave 3 files are not nationally representative. there is no solution to this problem that is being planned. these limitations have reduced the value of the wave 3 data. a few users have reported data and documentation discrepancies to staff and we have made corrections accordingly. for example, a user reported to us that there was invalid data in his sas data file for nsfh wave 3 respondent and spouse combined file which he downloaded from badgir. our investigation revealed that those invalid values are related to income variables. when household income is over one million dollars, the export engine in badgir does not allocate enough columns when the requested output format is sas. we increased the column width for income variables and republished the data file in badgir. occasionally, cdha staff receives questions from readers of mass media because nsfh is mentioned in a recent article in a popular magazine or news story. on 2008 father’s day, lisa belkin in her new york times magazine article, “when mom and dad share it all”, mentioned that, “the most recent figures from the university of wisconsin’s national survey of families and households show that the average wife does 31 hours of housework a week while the average husband does 14 — a ratio of slightly more than two to one.” we got many calls and emails from readers of ms. belkin’s article. they asked us to verify the hours spent on household works by husband and wife reported in that article. those numbers were likely produced by researches using nsfh data -but because we don’t know who did the secondary data analysis or what variables were used, staff cannot recreate the analysis. thus, we suggested those inquirers to contact ms. belkin directly for assistance. future plans for nsfh data the long-term plan is to create longitudinal nsfh files combining all three waves of nsfh data. these files will be very useful for research on the changes in family and household and the associated causes and effects over time. a match of nsfh data with the national death index is also being considered. cdha will archive the address files for wave 3 and create a geocode file for these addresses. all these plans will be carried out if cdha has sufficient funding in our third round of grant from the national institute of aging (grant number: p30 ag17266: 20092014). conclusion research based on the nsfh surveys has contributed significantly to the understanding of causes and consequences of changes in u.s. families and households in the last 20 years. multiple access points to nsfh such as, icpsr, sociometics, the nsfh website, and badgir give users different options for obtaining its data and documents. badgir, an enhanced discovery tool, has made access to nsfh easier. it is clearly an efficient webbased dissemination model for the nsfh study. its free, versatile and friendly user interface has drawn attentions from professors who use nsfh in their classes. one of them is rachel gordon, an associate professor in the department of sociology and institute of government and public affairs in the university of illinois at chicago. in her text book titled regression analysis for the social sciences published in 2010 by routledge, she includes examples from nsfh data analyses using badgir. we are pleased that nsfh and badgir utility are being used in this new textbook to advance the learning in secondary data research. it has been four years since our center became the custodian of the nsfh study. we will continue to provide user support to the research community in the spirit of good stewardship of a milestone social science study. it is rewarding to know that our assistance to nsfh researchers is appreciated. rarely our good intention is damped by aggressive users who want us to do their work for them. keeping good communication with nsfh users is important to us. we know our continuing support to the nsfh research community is relevant and important. with badgir we are confident to deliver reliable and informative user support for nsfh community. references bumpass, larry l. 1990. “what's happening to the family? interactions between demographic and institutional change” demography, 27(4): 483-498. sweet, james, larry bumpass, and vaughn call. 1988. the design and content of the national survey of families and households. http://www.ssc.wisc.edu/cde/nsfhwp/nsfh1. pdf. trull, elaine, lisa famularo. 1996 national survey of families and households wave 2 field report. ftp://elaine. ssc.wisc.edu/pub/nsfh/cmapp_n.001. iassist quarterly fall 2009 25 wright, debra. 2003. national survey of families and households wave 3 field report. http://www.ssc.wisc.edu/ nsfh/wave3/fieldreport.doc. introduction to the nsfh1 codebooks and other documentation. http://www.ssc.wisc.edu/nsfh/intro.htm. introduction to the nsfh2 codebooks and other documentation. http://www.ssc.wisc.edu/nsfh/intro.htm. introduction to the nsfh3 codebooks and other documentation. http://www.ssc.wisc.edu/nsfh/intro.htm. icpsr related literature database. http://www.icpsr. umich.edu/icpsrweb/icpsr/citations/index.jsp. note 1. chiu-chuang (lu) chou, senior special librarian, data and information services center and center for demography of health and aging, university of wisconsin madison, email: cchou2@wisc.edu, 3308 social science building, 1180 observatory drive, madison, wi 53706, usa. iassist quarterly 2010 / 2011 5 iassist quarterly editor’s notes barking up the right tree welcome to this very special iassist quarterly issue. we now present volume 34 (3 & 4) of 2010 and volume 35 (1 & 2) of 2011. normally we have about three papers in a single issue. in this super-mega-special issue we have fourteen papers from the countries: finland, ireland, united kingdom, austria, czech republic, denmark, germany, norway, slovenia, belarus, hungary, lithuania, poland and switzerland. this will be known in iassist as the “the book of the bremen workshop”. the workshop took place in april 2009 at the university of bremen. the workshop was hosted by the archive for life course research at bremen and funded by the timescapes initiative with support from cessda. the background and context of the workshop as well as short introductions to the many papers are found in the editorial introduction by the guest editors bren neale and libby bishop. the many papers are the result of the effort of numerous authors that were instrumental in the development and fulfillment of the many outcomes of the workshop. the introduction by the guest editors shows impressive lists of short-term activities, agreed goals, and also strategies for development. there are future initiatives and the future looks bright and interesting. the focus of the bremen workshop is on “qualitative (q) and qualitative longitudinal (ql) research and resources across europe”. i would have called that a qualitative workshop but you can see from the introduction and the papers that this subject is often referred to as “qualitative and ql data”. the “and ql” emphasizes that the longitudinal aspect is the special and important issue. in the beginning of iassist data was equivalent to quantitative data. however, digital archives found in the next wave that the qualitative data also with great value were made available for secondary research. the aspect of “longitudinal” further accentuates that value creation. this is a growing subject area. during the processing one of the authors wanted to update her paper and asked for us to replace the sentence “80 archived qualitative datasets and yearly around 30-40 datasets are ordered for re-use” with “115 archived qualitative datasets and yearly around 50-60 datasets are ordered for re-use”. yes, we do have a somewhat long processing time but this is still a very fast growth rate. i want to thank libby bishop for not being annoyed when i persistently reminded her of the iq special issues. i’m sure the guest editors with similar persistency contacted the authors. it was worth it. as in sherlock holmes we might look for what is not there as when curiosity is raised by the fact that “the dog did not bark”. iassist has had and continues to have a majority of its membership in north america so it is also remarkable that we here present the initiative on “qualitative (q) and qualitative longitudinal (ql) research” with a european angle. hopefully the rest of the world will enjoy these papers and there will probably be more papers both from europe but also from the others regions covered by the iassist members. articles for the iq are always very welcome. they can be papers from iassist conferences or other conferences and workshops, from local presentations or papers especially written for the iq. if you don’t have anything to offer right now, then please prepare yourself for the next iassist conference and start planning for participation in a session there. chairing a conference session with the purpose of aggregating and integrating papers for a special issue iq is much appreciated as the information in the form of an iq issue reaches many more people than the session participants and will be readily available on the iassist website at http://www.iassistdata.org. authors are very welcome to take a look at the description for layout and sending papers to the iq: http://iassistdata.org/iq/instructions-authors authors can also contact me via e-mail: kbr@sam.sdu.dk. should you be interested in compiling a special issue for the iq as guest editor (editors) i will also delighted to hear from you. karsten boye rasmussen editor august 2011 sist newsletter vol. 1, no. 3 data archivists, especially those working in isolation from traditional libraries, will want to consider the implications of white's work for their field. greater acceptance by libraries of access tools to data sets will generate a broader base of potential archive users and, since white argues cogently for the continuing housing of the actual data sets in archives, there is no reason to fear that the very specialized services of archives will be subsumed by libraries without staff expertise to promote their exploitation. we hope that dr. white will consolidate his groundbreaking findings into articles to be disseminated in the library and data archive press. the thesis itself is a mandatory purchase for library schools and a critical accumulation of data for archivists. we hope that dr. white's thesis is the first of many dissertations which will explore in greater depth the problems and characteristics of archives and their users. quantum completes survey quantum members in germany have completed a survey of completed, ongoing, and planned research projects in quantitative history. this survey has just been published as volume i in a new series of klett verlag, stuttgart, the hsf, which will be concerned with quantitative social scientific analysis of historical and process-produced data. the book entitled, the quantum documentation , is available from ernst klett verlag, rotebuhlstrasse 77, postfach 809, d-7000, stuttgart 1, at a cost of dm 39. the full bibliographic citation is: w. bick, p. j. miiller, h. reinke. quantum documentation: quantitative historische forschung 1977/quantitative history 1977. stuttgart: ernst klett verlag, 1977. new organizations/ reorganization bass hosts cessda meeting to create if-do belgian archives for the social sciences hosts committee of european social science data archives to create international federation of data organizations on the invitation of the committee of european social science data archives (cessda), representatives of data organizations from the united states, canada and europe gathered together at louvain-la-neuve, belgium on the 20th and 21st may, 1977 to discuss and formulate further means of mutual cooperation in the area of social sciences data archiving services. i^^sist newsletter vol. 1, no. 3 this meeting was also attended by observers from france, sweden and switzerland and a representative of the division for the international development of the social sciences of unesco. on sunday, 21st may, they decided on the establishment of an international federation of data organizations (if-do), designated mr. guido martinotti (adpss, milan) as its president, and entrusted the duties of secretary to mr. erwin scheuch (za, cologne). this federation is open to all organizations prepared to participate in the federation and cooperate in the continuing development of data archiving services in the social sciences. more information or a copy of the status may be available by contacting the secretary: dr. erwin scheuch zentralarchiv fur empirische sozialforschung universitat zu ktiln d 5 koln (deutschland) bachermerstrasse, 40 data organization and management the laboratory for computer graphics and spatial analysis within the graduate school of design at harvard university has just released a new edition of lab-log, its catalog of computer programs, data bases and publications. research at the laboratory is principally concerned with the analysis and graphic display of geographic data used in the planning process. lab-log describes various products which have resulted from this work and which are currently available for distribution to universities, government agencies and private organizations. lab-log includes a description of six different computer programs for use in the graphical display of spatial data via a line printer, line plotter and cathode ray tubes. a wide variety of cartographic (x-y coordinate) data bases are also described. publications are available on the subjects of automated cartography, theoretical cartography and theoretical geography. lab-log also contains a brief description of the laboratory's history, research directions and operating policies copies of lab-log are available at a cost of $1.00 each upon request to: the laboratory for computer graphics and spatial analysis 520 gund hall harvard university 48 quincy street cambridge, ma 02138 payment must accompany your order. 32. 36 iassist quarterly 2013 iassist quarterly building the ddi by ann green and chuck humphrey1 abstract this paper describes the context, motivation, and requirements behind the design and development of the first version of the data documentation initiative2 (ddi) metadata community specification, with an emphasis upon the process of creating the initial element set for the “study level” of ddi version 1. we also offer a framework for understanding the infrastructural changes that contributed to the establishment of the ddi. by taking a close look at the confluence of influences on the earliest efforts to design and build the ddi, we can better understand what essential elements of metadata are necessary to support independent use of social science data over time. keywords: metadata, data description, documentation, social science data, ddi introduction initially, the ddi was built to provide a structured framework for metadata describing social science surveys, but has since grown to include a broad range of data types. from the beginning, it was clear that over time the ddi would need to provide a metadata structure for longitudinal data, comparative data across geographies, panel studies, aggregate data, and administrative records. by taking a close look at the confluence of influences on the earliest efforts to design and build the ddi, we can better understand what essential elements of metadata are necessary to support independent use of social science data over time. by the 1980’s, the types of descriptive information necessary for researchers to use data created by someone else were well established. such documentation needs to describe the content, quality, and features of a dataset, which in turn provides an indication of its fitness for use. the social sciences benefit from a long history of large-scale databases that were accompanied by extensive documentation (some examples in the united states are the general social survey, the american national election study, the california poll, the health interview survey and in canada there are the public use microdata files for the census of population beginning in 1971). these databases were designed for sharing and wide use, so the documentation needed to support the independent use of the data without making it necessary to return to the producer of the data to make sense of things such as the methodology, coding of variable, question text, interviewer characteristics, sampling, etc. (ccsds, 2012). data producers, primarily in government agencies and large research centers, were expected to publish printed documentation about the data they collected and disseminated to the world. at the time, almost all of this documentation was in printed form. data archives across the us, uk, europe, and canada, which had already begun collecting and taking stewardship of large collections of social science data files and accompanying documentation, were also building catalogues and access systems of technical capacity far ahead of their time. data professionals began to build a community of expertise to support the many services related to curating and preserving voluminous collections of social science data. libraries began collecting and cataloguing these materials to support their designated research communities as the reuse of large data collections became a standard practice in the social sciences. it was at this time that a shift from paper documentation to machine readable alternatives presented new capabilities for expanding the role of descriptive information and hyperlinking the intellectual components of that information. what was needed was a way of structuring this information so that it was both human readable and machine actionable. challenges to make data independently usable continue today and as such, the research community needs to be encouraged to continue its commitment to produce and distribute information about the data they collect, share, reuse, analyze, replicate, and publish. as important now as in the 1960’s, metadata are used to locate and review data for fitness of use, to have transparent access to methods and sampling, to understand the capacity for linking data, and to create maps, visualizations, or mine large data collections. metadata are integral to all data manipulation functions. without structured, complete, and accessible metadata these challenges cannot be met. in this review, we make connections between metadata activities today and developments in the past where the rich history of documenting social science data has been overlooked. we also offer a framework or model for understanding the iassist quarterly 2013 37 iassist quarterly infrastructural changes that contributed to the establishment of the ddi. finally, we explore the importance of cataloguing and citation standards, study description guidelines, and how information in ‘codebooks’ could be best represented in the ddi. throughout this exploration, we point to the contributions of the many participants in the ddi’s early development, especially acknowledging sue dodd’s influence and engagement. from cataloging to metadata and shifts in research infrastructure the recent flurry of interest around widely promoting data citation and attribution to entice researchers into sharing their data has been largely unconnected to the history of cataloguing machinereadable data files. in the introductory chapter to the national research council’s report (2012), for attribution – developing data attribution and citation practices and standards, christine borgman acknowledged that the current debate around drivers behind data citation and attribution has failed to recognize long-standing cataloguing practices for research data. simply put, if research data files can be catalogued, they can be cited. we have had standards for cataloging data files since the 1970s. objects that can be cataloged also can be cited. similarly, data archives have been promoting data citation practices for several decades. however, over this same period, very few journal editors required data citations, disciplines did not instill data citation as a fundamental practice of good research, granting agencies did not reward the data citations of applicants, tenure and reward committees did not recognize data citations in annual performance reviews, and researchers did not take responsibility for citing data sources. (national research council, 2012, p. 1) while the potential for widespread adoption of data citation practices has been present for several decades, the uptake has been slow largely because the production of catalogue records has been primarily associated with print objects. it is significant that the rules for cataloguing machinereadable data files were openly embraced by the social science data archive community in the late 1970’s and early 1980’s. however, the wider library community was slower to adopt these practices. for example, an oclc project in 1990 generated catalogue records for the complete set of icpsr class i codebooks in print, which reflected the bias at that time for print over digital objects. nevertheless, an increasing volume of descriptive information about machine-readable data files and accompanying documentation catalyzed the change from cataloguing to metadata. over the past three decades, the production and use of descriptive information supporting the discovery, access, usability and preservation of research data has fundamentally changed. we find it helpful to think of this transformation in terms of a shift in research infrastructure. paul edwards, steven jackson, geoffrey bowker, and cory knobel (edwards, et. al., 2007) provide a model that characterizes such changes in research infrastructure. adoption of this model requires seeing metadata as a component of research infrastructure, which is itself a mental shift. figure 1 represents the edwards, et. al. distribution of cyberinfrastructure solutions that take shape across the dimensions of global-local and social-technological contexts. the authors insist that building cyberinfrastructure is not a case of selecting an end point along these dimensions but one of choosing from the distribution of solutions that are availed across these factors. the following example about the choice of infrastructure to support guest internet access on a local campus will illustrate the application of this model and help pave the way to discussing changes in infrastructure options for research metadata. a typical university service providing visitors with wireless internet access requires visitors to be issued a guest account and password. a visitor will need to go to the campus office where the person responsible for issuing guest accounts is located, to show some identification, and to sign an internet use agreement before receiving an account and password. this specific solution to provide guest internet access is defined completely by local organizational practices, social norms, and technology and is characterized as a local-social solution in the above model. an alternative solution can be found through institutional membership in eduroam3. wireless network access is available to anyone at an eduroam site as long as the person is from an institution that is part of the eduroam network. using the credentials from one’s home institution, the eduroam service figure 1: a distribution of cyberinfrastructure solutions 38 iassist quarterly 2013 iassist quarterly authenticates a guest by automatically verifying their account with their home institution. local policies can also be configured on eduroam servers to make additional resources available to guests, such as, printers or access to licensed databases. in terms of the model, eduroam is a global-technological infrastructure solution. this service is available in over 65 countries worldwide and is governed through of confederation of national organizations. the array of infrastructure solutions at any one time is in flux. for example, as norms around privacy and confidentiality in today’s digital world swing, the range of social solutions will expand or contract. as interconnectivity using trusted, standards-based protocols shrinks the world, new global solutions become available. new local possibilities emerge as individual institutions develop policies, procedures, and guidelines around digital asset management. finally, rapid changes in information technology are constantly resulting in new ways of doing things. the story of the movement from cataloguing to metadata is expressed both in the changing array of infrastructure solutions over time and the dominant solutions that have emerged. the solutions for cataloguing machine-readable data files (which aacr2 now designates as an electronic resource) began largely within the local-social context in the late 1970’s when a group of university libraries began producing their own marc records for research data. in the 1980’s, the icpsr began circulating studylevel marc records of their data holdings, pushing catalogue infrastructure toward a global-social solution. increased automation by icpsr in the production of marc records and the oclc becoming a distributor of marc records for icpsr holdings moved cataloguing support for this collection into the space of global-technological infrastructure. however, with increased use of ddi metadata since 2000, marc records for social science research data have tended to be produced through a crosswalk between the ddi and other record formats, such as dublin core or marc21 xml. catalogue records can now be derived through other forms of metadata, largely making the practice of cataloguing research data unnecessary. catalogue records for research data have been useful for study level discovery purposes but the quest has long been for variable level discovery. in the 1980’s the icspr introduced a variables database searchable through spires for a subset of studies. this database demonstrated the value of variable level metadata but the workflow to produce such metadata was dependent on special, extended processing, namely, the generation of icpsr class i studies and osiris dictionaries. the array of infrastructure solutions to facilitate variable level discovery changed throughout the 1980’s and into the 1990’s. coming out of the 1970’s, documenting research data was treated primarily as a publication process. this information was assembled into a printed report, which was commonly referred to as a codebook. these volumes often contained sections dedicated to a technical description of the study and method of data collection, a detailed listing of the variables, their codes, and their layout in a data file, a copy of the data collection instrument, and any other contextual information providing background to the data. the production infrastructure for such documentation tended to be in the form of solutions that were local, social, and technological. they entailed some automation with a substantial amount of human resources within a local operation. the mindset of this era was one of assembling as much information as possible and organizing it in a printed booklet. subsequent use of this metadata, however, often required reentering information for other automated functions. for example, record layouts of variables would be rekeyed for multiple statistical packages, which would often be repeated locally across the many universities receiving copies of the same study and its data. an important change in the practices around metadata production occurred with a growing acceptance of reusing digital information for many purposes. this practice of enter-once-and-reuse-formany-purposes pushed the design of data documentation into incorporating digital content that is structured consistently, well defined, and universally sharable. concurrent with this view toward reuse of digital information was the introduction of dynamic digital texts. project xanadu and subsequent hypertext initiatives demonstrated the power of automating connections between bodies of text. hypercard (produced by apple) in the late 1980’s popularized the mapping of relationships in digital text and in the 1990’s the world wide web became the most successful implementation of hypertext. the web also proved the utility of linking key descriptive elements within a document without necessarily linking to the whole document: targeted reuse of specific information elements. the technology around hypertext coincided with the development of structured conceptual data models that supported the identification of key information elements. the introduction of sgml digital texts employing mark-up tags allowed designating text to a specific structural element. furthermore, sgml supported the description of conceptual layers of digital information. when hypertext and mark-up languages converged with the web, the power of describing layers of information in ways that could be linked and reused substantially altered our understanding of metadata for research data. the application of entering information digitally once and then reusing it for multiple purposes through conceptual linkages is now a dominant technological solution in metadata infrastructure. all of the technological changes leading up to the uses of digital text on the web shifted the array of solutions for metadata infrastructure from social to technological. in addition, the array of solutions has been expanded through two factors driving solutions from local to global infrastructure. first, metadata standards for research data were needed to define the key information elements in data documentation. the widespread acceptance and use of a standard such as ddi pushed metadata toward global solutions. standards enable information to be appropriately compared with predictable outcomes. second, the ultimate reuse of metadata occurs when this information is turned into machine actions. the movement from the systematic use of metadata to describe elements within a conceptual model to invoking machine actionable workflows is one headed toward creating global data interoperability. our thinking about description, discovery, access, usability and preservation has been altered through changes in metadata infrastructure. we are now being challenged about what should be described, how to structure the description, the purpose for which the content can be used, and the workflow processes that this information can drive. iassist quarterly 2013 39 iassist quarterly ddi committee 1995 – 1997 the first meeting of what was to become the data documentation initiative committee was held in conjunction with the annual meeting of the international association of social science information service and technology (iassist) in quebec city, quebec, in may of 1995. (see the accompanying iassist quarterly article by karsten boye rasmussen for some data description activities prior to this meeting.) at that time, merrill shanks, professor of political science at the university of california at berkeley, was named chair of what was the first iteration of the ddi committee: the icpsr committee on survey documentation (later to be called the committee on data documentation). invitations to join the committee were sent in february 1995.4 the charge to the committee5 made it clear that this was to be an international effort with inclusive involvement across research, survey, and data professional organizations, representing the interests of data producers, archives, distributors, and end-users of social science data. “as you probably know, this will be hard work; it will require production of a dtd [document type definition, ed.] that meets the needs of icpsr and that can be implemented immediately in production. i am even more ambitious than that: i would like for the product of this committee to provide a standard that is acceptable to iassist, apdu, and our european partner archives as well. it should fully conform to sgml, and this should be tested with standard sgml software. a document should be produced (and made available on the world wide web) defining this dtd and making it available to the entire community.” 6 it is important to note that the committee was encouraged “to continue to consult with other interested parties concerning both the short-term and long-term goals (or content) for our evolving dtd, and we should keep each other informed of any new development or second thoughts concerning our initial agreement in quebec.” this was, from the start, seen as a community effort and the committee members were charged with the responsibility of consulting with interested parties. some of the larger goals surrounding the development of the ddi were to: come up with a non-proprietary format that was ‘preservation friendly;’ streamline the process from data collection to metadata production; develop metadata authoring tools for specific purposes; produce and distribute software converters to automate the transport of metadata to varying formats; develop cross-walks and supporting linkages; make it easier to integrate ddi metadata into various systems for resource discovery and statistical analysis; improve linkages between data and metadata; analyze more than one study at a time; offer cross-domain searching and integration; and better integrate geospatial analysis and statistical analysis. (green, 1999) committee members and others from the social science community were given the task of developing a draft list of elements for the first version of the ddi (which was initially rendered in sgml in april 1996). two subcommittees were established to make recommendations and clarify issues; one of these subgroups concentrated on the different kinds of study level information, while the other focused on the detailed specifications at the variable level. david barber, at the university of michigan library, was charged with combining the suggestions from both subcommittees and with developing the first ddi dtd. it is important to note that the ddi was built to contain descriptive information not only about the variables and coding in a data file, but also to include descriptive information about the study itself and to provide a tagged structure for potentially all of the elements that were deemed important to fully documenting data. defining the ddi study level information the study level subcommittee, led by ann green and mary vardigan, coordinated the development of the elements for the study level description with the help of sue dodd, karsten boye rasmussen, laine ruus, bill bradley, carolyn geda, pat vanderberg, bridget winstanley, atle alvheim, rolf uhrer, richard rockwell, merrill shanks, david barber, and john brandt. their goals were to develop a common core of elements that could be understood and applied across communities of data producers, survey collecting agencies, libraries and archives, researchers, and software developers. the ddi was constructed from existing standards and guidelines that had been in use, in some cases, for decades at data archives in the us, canada, the uk, and europe. the ddi also was intended to support new applications for study description information: to enhance the ability to search, display, and manipulate metadata; to provide a means of discovering that a data set exists and how it might be obtained or accessed; to document the content, quality, and features of a data set and so give an indication of its fitness for use; to supply information for statistical analysis software; and to provide information for citations, cataloguing records, and electronic headers. the major challenge in developing the structure and individual elements for the study level portion of the dtd was to incorporate the concepts and parts of traditional printed codebooks and also to build compatibility with computer-generated data collection and documentation processes. at the same time the subcommittee gathered elements for the intellectual description, they needed to examine the processes and output of computergenerated surveys to understand the relationship between the survey instrument and the production of study-level descriptive information. even though codebooks describing datasets did not at the time have strict standard structures or formats, there was a standard set of intellectual content outlined in guidelines and reviewed in the major social science documentation literature. it was critical that these intellectual components be the defining force behind the “distillation” process of producing the individual elements in the codebook dtd. the goal was to distill a set of elements that were comprehensive and flexible, and capable of producing pieces that are compatible with automated methods of producing codebooks, as well as feeding into systems that describe, cite, catalogue and locate datasets. the procedure for identifying and defining the study level elements for the dtd included reviewing the following five kinds of resources, each of which is described in detail below. 1. review cataloguing and citation guidelines 2. study the ways in which social science data have been described by data archives. 3. examine data archive and data producer guidelines 4. examine the pieces of existing print and machine-readable codebooks 40 iassist quarterly 2013 iassist quarterly 5. review other encoding guidelines in development at the same time, and electronic header standards 1 review cataloguing and citation guidelines the first step was to review the standards that establish rules for producing cataloguing records for the study in an online public catalogue or online retrieval system, with an emphasis upon intellectual ownership and identification. this information serves as the basis of a standard bibliographic citation. the cataloguing of machine-readable computer files was well established by the early 1980’s. the anglo-american cataloguing rules, second edition (aacr2) was published in 1978 and included a new chapter for machine-readable data files (chapter 9).7 the library of congress, the national library of canada, the british library, and the australian national library adopted aacr2 in january 1981. ala published an interpretive manual by sue dodd focusing on machine-readable data files in 1982. this was followed by a manual for cataloguing microcomputer files in 1985 by dodd and ann sandburg-fox. prior to these works, dodd had already published on cataloguing standards in the asis journal (dodd, 1979a) and the journal of library automation (dodd, 1979b). her manual on cataloging machine readable data files was one of the founding documents for the bibliographic components of the ddi standard and other citation standards to come. the integration of bibliographic citations within data documentation is not new. in march, 1996, as part of the review of the proposed ddi elements, sue dodd encouraged via email9 that there be “a better distinction between required bibliographic information denoting intellectual identification (aka citation) and data abstract information (aka study description).” her recommendation was to “disconnect them (conceptually).” this was an important distinction, which clarified the necessity for the ddi instance to carry within it a distinct reference to the intellectual identity of its creators, producers, and distributors. from the beginning it was clear that citations to data are an essential aggregation of descriptive elements best compiled into a standard format. an element was added to the dtd called “bibliographic citation” so that a complete citation could be carried within the documentation instance. it was also important that the ddi include elements that could be compiled for multiple purposes (for example, compiling a citation in an alternate format, forming a title page if the codebook was printed or rendered as a separate document, or mapping intellectual identification to other metadata standards, e.g. marc or dublin core). this flexibility was illuminated by dodd’s insight into the relationships among instances of documentation, and the necessity for the ability to carry the requisite information to create various output formats from a single ddi instance. dodd also noted “you might want to include information on the difference between citations for the documentation and citations for the computer file – provided they are separate (and not “packaged” together).” this was one of the key challenges of creating the dtd – to clearly define all of the objects and intellectual components being described by the ddi instance. that meant that there was to be a citation to the document itself (the ddi instance), a citation to the study being described (study description), and detailed identification of the particular file/s that made up the physical object being described (file description). this was part of the motivation to produce the ddi as a modular entity, with components that clearly articulated and integrated these separate conceptual pieces. dodd also addressed the need to verify authenticity. she wrote: “the study number supplied by the producer and the archival number supplied by the distributer and archive may be different. this difference should be noted. there can be an original study number (e.g. harris a019) and an archival study number (e.g. icpsr 7657). they represent the same data, but different distributors and archives.” it may seem obvious now, but at the time it helped articulate the importance of retaining all distributor and archive assigned study numbers, a key component of trust that content is what it purports to be. defining citation principles for data has become a popular topic (codata, 2013), but data archives have been promoting data citation practices for approximately forty years, and have for decades included citations to data within data documentation. since the 1980’s, libraries have been producing bibliographic records containing the basic information for how data should be cited in a publication (mooney & newton, 2012). the ddi was built upon this history of promoting data citation, of cataloging data, and of including data citations in documentation. the history of data attribution and citation has always been at the core of the ddi. cataloging and citation guidelines • marc-mrdf: the work of the american library association sub-committee on rules for machine-readable data files. local variations of marc format have been developed in canada, the united kingdom, sweden, etc. • isbd-computer files: the international standard bibliographic description for computer files. recommended by the working group on the isbd set up by the international federation of library associations (ifla) committee on cataloging. (ifla, 1990) • gils: government information locator system: these locators provide users with descriptive, location, and access information for a wide range of [u.s.] federal government information resources. compliant with z39.50, a standard way for two computers to communicate for the purpose of information retrieval and facilitates the use of large information databases by standardizing the procedures and features for searching and retrieving. • iso 690-2: draft standard for bibliographic references to electronic documents iso 690-2 is a standard in review for the content, form and structure of bibliographic references to electronic documents, being developed by iso technical committee 46, subcommittee 9. • dublin core: oclc/ncsa metadata workshop, online computer library center 1996: university of warwick/ uk office for library and information networking oclc/ncsa metadata workshop recommendation for core data elements for discovery and retrieval of internet resources by a diverse group on internet users. listed data elements with possible equivalents in usmarc.. 2.examine the ways in which social science data have been described by data archives the ddi subcommittee also examined descriptive metadata contained in study descriptions and data catalogs to identify key descriptive material that should be included in the ideal iassist quarterly 2013 41 iassist quarterly comprehensive codebook. the focus was on capturing the intellectual content of the pieces rather than their variant names/ labels. data description • standard study description: developed by and for data archives, adopted by several members of the council of european social science data archives in 1974 and endorsed by the council of european social science data archives. further details regarding the origins of the study description can be found in: nielsen, per: report on standardization of study description schemes and classification of indicators, copenhagen: dda, september 1974, 62 pp. nielsen, per: study description guide and scheme, copenhagen: dda, april, 1975, 55 pp. • icpsr study description “template” manual. “every new or revised icpsr study requires a study description which is written by the staff member who processes or evaluates the study. these descriptions follow a strict format, called a template, which insures that standard information is recorded for each study. the template consists of named fields into which the staff person enters appropriate information about the study. completed templates ultimately reside in an online spires database.” • essex esrc data archive study description outline, supplied by bridget winstanley • federal geographic data committee (fgdc) subcommittee on cultural and demographic data. (scdd) content standards for cultural and demographic data metadata. (c&dd metadata standard) specifically the crosswalk. • ruus, laine. university of toronto. “a comparison of major descriptive systems in use to describe computer-readable data files,” 2nd edition, september 1992. “the objective of this document is to track major systems currently in use for the description of computer-readable data files. identifying the comparable fields in the descriptive systems, as well as the difference, will serve to support recommendations…to satisfy national and international requirements for the formal description of computer-readable data files” (p. 2) each field (element) was described in numeric marc tag number order, indicating if the particular field was mandatory, optional, and repeatable. that list is followed by definitions of each field, from each source. the compilation represents a very time consuming and precise effort with international scope. the comparison included the following: the canadian union list of machine readable data files at the university of alberta; statistics canada catalog; cataloging computer files in the uk: a practical guide to standards, edited by peter burnhill and ray templeton; the treasury board of ottawa’s guide to structure data model data dictionary; nasa’s directory interchange format manual; isbd international standard bibliographic description for computer files; rlg’s machine-readable data files memory aid; health and welfare canada’s microdata set documentation: reference guide; danish data archives’ study description completion guide; icpsr’s template manual; and the us library of congress usmarc concise formats for bibliographic, authority, and holdings data. 3. examine data archive and data producer guidelines the best practices for how to prepare data for archiving have been around for over three decades. it was essential that the ddi support the best practices of preparing and documenting data as described in these guides. guidelines for preparing data • roistacher, richard: a style manual for machine-readable data files and their documentation with sue dodd, barbara noble and alice robbin. (roistacher, 1980). note that numerous data archives being established in the 1980s and 1990s used this manual. the style manual was influential in the development of standardized documentation of data files. • geda, carolyn: 1980, icpsr, data preparation manual. • collins, patrick, 1996, depositing data with the national data archive on child abuse and neglect: a handbook for investigators. • us bureau of the census, statistical research division, statistical design and methods extension to cultural and demographic data metadata: cddm draft standard 1995. an extremely detailed table of contents and definitions of proposed documentation of census or surveys. • essex survey research center data archive: documentation guidelines committee (presentation to iassist in 1994) 4. examine the pieces of existing codebooks in print and machinereadable formats. an important step in the process of building the ddi was to review the components of standard printed codebooks, which included full bibliographic identification with a standard citation; an abstract; descriptive and contextual materials, such questionnaires, statements of methodology, appendices and glossaries, and coding schemes for things like geographic entities, topical recoding, or occupation and industrial classifications. it was also important to examine how statistical packages and data archives at the time included metadata in the system files of these packages and programs. the most comprehensive statistical package, in terms of metadata, was osiris (rattenbury & pelletier 1974), a set of computer programs that also included descriptive information. the osiris codebook was part of that system and carried structured information about the survey instrument, file descriptions, and the elements making up a bibliographic reference. one of the goals of the ddi committee was to come up with a replacement of the osiris codebook / data dictionary format. of course the ddi became more than simply a replacement for osiris as it grew to support the entire lifecycle of data. two other information systems influenced the construction of the ddi: one is a suite of software developed at the university of california at berkeley, cases (computer assisted survey execution system) and csa (conversational survey analysis). codebooks were produced as a by-product of computer assisted interviewing software (cati/capi) and were integrated into accompanying analysis software. not only was it useful to examine this system to understand how the survey process intersected with the documentation process, but these tools were in use by researchers and government agencies who could benefit from incorporating the ddi into their survey process and the subsequent dissemination of data and documentation to their user communities. another system that informed the development of the ddi was health canada’s ddms system (data dictionary and documentation management system). ddms was a pc-based package for producing social science data dictionaries and documentation and for managing research outputs from health canada. this tool furthermore interoperated with metadata registries. 42 iassist quarterly 2013 iassist quarterly a close review of each of these systems provided insights into what the ddi needed to contain to be integrated into and support the processing of surveys and the management of documentation within analysis systems and metadata registries. 5. review other encoding guidelines in development at the same time. other developing standards at the time, in particular the tei9 for encoding digital texts and the ead10 for encoding archival descriptions, informed the development of the ddi. they were especially influential because they addressed similar objectives, were based upon community standards, and initially used the sgml framework. all of these initiatives were working on file header standards and the ddi incorporated those guidelines. encoding guidelines • tei: text encoding initiative dtd for sgml: markup for primary materials (note that the tei was used to some extent in producing the ddi xml version 1.6x dated december 27, 1997) • ead: encoded archive description dtd for sgml: markup for metadata describing primary archival materials • file header information, drawing upon other encoding standards. this set of elements contain information about the marked up ddi instance itself. file headers support resource discovery and establish bibliographic identity of the ddi file itself. the ddi document description component of the ddi is essentially the “header.” refining the elements and moving from sgml to xml: 1995 – 1997 the ddi committee met in october 1995 and again in april 1996 to examine a sample sgml dtd prepared by john brandt and his colleagues at the university of michigan library. at a meeting in october 1997, subcommittees were formed to conduct a review of the elements of the dtd and to address the issue of handling aggregate data in the dtd. in december 1997, the dtd was made compliant with xml (extensible markup language) by jan nielsen of the danish data archives. (nielsen, 1998) impact as we have shown by describing the beginnings of the ddi, the specification was developed at a time of extraordinary shifts in research infrastructure and information science. the creators of the ddi were aware of these shifts and committed to producing a solution to the metadata challenge that could build upon the strengths of data description from, as in the edwards et. al. model, a highly social and local context as well as meeting the technical and global demands and capabilities of the time. however, with the production of the ddi came new dependencies to find solutions to capture and produce ddi compliant metadata and to take advantage of the constantly evolving technical capabilities and a rapidly changing research environment. the vision of the ddi and the metadata produced through its application went beyond merely structuring information necessary for using data. metadata was also seen as the connection between data producers and data users and the technical solutions required to meet the challenges of transferring knowledge in structured formats. as jostein ryssevik wrote (ryssevik, nd): “whereas the creators and primary users of statistics might possess “undocumented” and informal knowledge, which will guide them in the analysis process, secondary users must rely on the amount of formal meta-data that travels along with the data in order to exploit their full potential. for this reason it might be said that social science data are only made accessible through their metadata. the metadata provides the bridges between the producers of data and their users and convey information that is essential for secondary analysts.” but building that metadata bridge is a difficult task. at the time the ddi began, there were, and continue to be, major challenges in collecting and distributing metadata. the most obvious and disconcerting fact is that information about data, its context, and its content, is not recorded or is inadequately stored – for example in unstructured and incomplete ‘read me’ files. the commitment, workflow tools, and production of what needs to accompany data for informed use have not been widely or enthusiastically embraced across research teams. the reasons are primarily due to time and resource constraints (tenopir et al., 2011), but also a lack of integrated tools that capture metadata throughout the research lifecycle and that package the information in ways to support the sharing and archiving of data. recent requirements to share and preserve data have created a new conversation about research data management, yet at the same time data sharing platforms accept data without verifying the quality of the documentation. the norm with incoming data is not to review or to check for adequate metadata that would support reuse and replication. in spite of the presence of metadata specifications across many disciplines and detailed guidelines in preparing data, the challenges of producing and distributing good quality structured documentation continue to impede the reuse, replication, and sharing of data. the promise of the ddi cannot be met as long as metadata are not being captured. another challenge is to explore how the ddi could be incorporated into tools that capture metadata throughout the full lifecycle of the research process. as alice robbin wrote ”(d)ata documentation, the descriptive text accompanying a file, is the key to understanding its quality“ and “should be prepared at the time of a file’s creation and may contribute significantly to future use of the data….”(robbin, 1981). lifecycle models developed soon after the ddi emerged made it clear that metadata production was not simply a process that happened at the end of a research project (green and kent, 2002; ddi, 2004; humphrey, 2006). the idea of the metadata lifecycle, and its intersections with the research lifecycle, has become a common element in publications, conference presentations, and metadata modeling efforts. mary vardigan’s article in this issue carries our story forward into the development of the lifecycle model11 of the next version of the ddi. the ddi specification is dependent on new technological developments to reach its potential capacity. we draw particular attention to: • interoperate with other metadata systems for resource discovery and cataloging systems; • establish software for parsing, validating, viewing, searching, manipulating, authoring and converting; • exploit the ability to link to with other digital objects; • take advantage of non-proprietary and platform independent metadata in preservation systems; iassist quarterly 2013 43 iassist quarterly • integrate descriptive metadata into analysis, visualization, mapping, data mining; • realize interoperability with other data and the promise of open data, especially the demands related to comparability, privacy, authenticity, and attribution; and • inculcate into the habits and workflow of research and data production systems tools and incentives for creating metadata. responses to such challenges are best met through a concerted effort to enrich the evolving array of solutions identified within the metadata infrastructure model describe above. as a community, we especially want to exploit technologies that are flexible and responsive to local requirements, to incorporate social drivers and habits, and to have a clear goal of meeting global requirements for shared and open data. this can be done, just as the ddi was created, by incorporating potential solutions, carefully articulated requirements and expertise from across the communities of data producers, researchers, data archives, and institutional repositories. references american library association. (1978). anglo-american cataloguing rules second edition: chapter 9: computer files: draft revision, volume 9, michael gorman, joint steering committee for revision of aacr, american library association. codata/itsci task force on data citation. (2013). “out of cite, out of mind: the current state of practice, policy and technology for data citation.” data science journal 12, 1-75. data documentation initiative (ddi). (2004). ddi version 3.0 conceptual model structural reform group. background of the conceptual model. working draft. retrieved from dodd, s. a. (1979a) “bibliographic reference for numeric social science data files: suggested guidelines,” american society for information science journal 30:2, p. 77-82. dodd, s. a. (1979b) “building an on-line bibliographic/marc resource data base for machine-readable data files,” journal of library automation 12:1, 6-21. dodd, s. a. (1982). cataloging machine-readable data files: an interpretive manual. chicago: american library association. dodd, s. a. & sandberg-fox, a. m. (1985). cataloging microcomputer files: a manual of interpretation for aacr2. chicago: american library association. edwards, p. jackson, s., bowker, g., & knoble, c. (2007). “understanding infrastructure: dynamics, tension, and design.” report of a workshop on history & theory of infrastructure: lessons for new scientific cyberinfrastructures. retrieved from green, a. (1999). “why the data documentation initiative is so important.” presentation to the association of public data users. october, 1999. green, a. & kent, j-p. (2002) “the metadata life cycle.” in: kent, j-p., ed. metanet: a network of excellence for harmonising and synthesizing the development of statistical metadata. metanet work package 1: methodology and tools. the metanet project. p. 29-34. retrieved from humphrey, c. (2006). “e-science and the life cycle of research.” retrieved from mooney, h, newton, m.p. (2012). “the anatomy of a data citation: discovery, reuse, and credit.” journal of librarianship and scholarly communication 1(1):ep1035. national research council. (2012). for attribution -developing data attribution and citation practices and standards: summary of an international workshop. washington, dc: the national academies press. nielsen, j. (1998). “from osiris to xml: markup and internet presentation of structured data documentation.” unpublished thesis. rasmussen, k. b. (1995). “documentation what we have and what we want: report of an “enquete” of data archives and their staff.” iassist quarterly 19(1). retrieved from rasmussen, k. b. (2013) ”social science metadata and the foundations of the ddi.” iassist quarterly 37. rattenbury, j. & pelletier, p. (1974). data processing in the social sciences with osiris. ann arbor : survey research center, institute for social research, university of michigan. retrieved from robbin, a. (1981) “strategies for improving utilization of computerized statistical data by the social scientific community.” social science information studies 1, p 89-109. roistacher, r. c. with contributions from dodd, s.a., noble, b. b., & robbin, a. (1980). a style manual for machine-readable data files and their documentation. washington, d.c. u.s. dept. of justice, bureau of justice statistics. grant no. 78-ss-ax-0028 awarded to the bureau of social science research. report no. sd-t-3. ncj-62766. ryssevik, j. (nd) “the data documentation initiative (ddi) metadata specification.” retrieved from tenopir, c. a., allard, s., douglass, k., aydinoglu, a.u., wu, l., read, e., manoff, m., & frame, m. (2011). “data sharing by scientists: practices and perceptions.” plos one 6(6). doi:10.1371/journal.pone.0021101 retrieved from vardigan, m. (2013) vardigan, mary (2013) ”the ddi matures: 1997 to the present. ” iassist quarterly 37. vardigan, m. (2013) ”timeline.” iassist quarterly 37. 44 iassist quarterly 2013 iassist quarterly notes 1. ann green, digital lifecycle research & consulting (dlifecycle@gmail. com), is an independent consultant working in the areas of research data management and digital preservation. chuck humphrey (chuck.humphrey@ualberta.ca) is the research data management services coordinator in the university of alberta libraries and the academic director of the alberta research data centre. 2. data documentation initiative. 3. see 4. memorandum from merrill shanks, chair, to fellow members of the icpsr committee on survey documentation, hand dated june 7, 1995 5. letter of invitation from richard rockwell 6. memorandum to ann green, dated february 14, 1995 from richard rockwell, cosigned by j. merrill shanks, peter granda, and mary vardigan. subject: invitation to serve on icpsr committee. 7. a draft revision of aacr2 chapter 9 (renamed: computer files) was published in 1987. 8. email from ann green to mary vardigan, march 12, 1996. subject: dtd changes. enclosed are communications between sue dodd and ann green in regard to the review of ‘study level’ elements. 9. for information about the history of the tei see: 10. for information about the history of the ead see: 11. see the ddi lifecycle illustration here: keywords from vol 37 n0. 1-4. courtesy tagxedo.com vol273.indd 4 iassist quarterly fall 2003 editorʼs notes after some delay we welcome you to the third issue of the iassist quarterly vol. 27. i offer my sincere apologies for this delay. it is my hope that in the coming months we will be able to catch up by sending out the missing issues at a faster pace. the three articles in this 27-3 iq issue are presentations from the iassist 2004 conference and a “user experience” in setting up the 2003 conference. both were successful conferences – and both were as always “better than ever!”. the session on “facilitating data access and analysis” included the presentation and the article on “delivering the world: the establishment of an international data service”. this was presented by the main author susan noble from mimas (manchester information & associated services) at the manchester computing, university of manchester. co-authors on “delivering the world” are keith cole, celia russell, jim schumm, nicholas syrotiuk, gindo tampubolon. the esds (economic and social data service) international is a service providing access to the major socio-economic databanks produced by inter-governmental organizations institutions as the world bank and the united nations are mentioned – as well as to international survey datasets. the service delivers the world over the web and the service is free – well not to the world but to the uk academic community. up to now access has been very expensive and individual institutions had to have arrangements/subscription to each source. the work has established a national academic access in uk. the article discusses and describes the data acquisition strategy and the establishment of licensing agreements with the data providers. the delivery software and the development of a user interface are described and there are reports on the challenges of converting large and complex datasets. you can take a closer look at international data at the web-site: (web-site: http://www.esds.ac.uk/international). another of the sessions had the heading of “privacy, security, and information today”. the development of information technology has made more surveillance possible and unfortunately terrorism has made these technologies in demand. from the us-side juri stratford from university of california (davis) reports on “internet surveillance: recent u.s. developments”. the article describes some techniques that in parts resemble wellknown telephone surveillance, but in others are much more powerful. the u.s. federal bureau of investigation uses the carnivore system that can capture e-mails to and from a specifi c account or similar for a specifi c ip address. the dilemma becomes obvious when on one hand the carnivore system in 2000 was declared “not legal … within the context of any current wiretap law” and on the other hand the usa patriot act following the terrorist attacks in september 2001 introduced “new authorities” for the surveillance. the article takes a closer look into this. from the more comfortable editor chair – outside the u.s. the dilemma is that when more surveillance is heavily performed it transforms the society into a more rigid and totalitarian society which is exactly the goal and claim of the terror. the last article is related to the iassist conferences as well. but the connection here is to the administration of the iassist 2003 conference in ottawa. the conference used a web-based registration procedure. the experiences from this have a wider appeal. as the abstract mentions the web can empower the data user and thus free yourself as a mediator between the data requestor and the data. the technique describes using a web-interface and the storage in a database. this is the subject and title of the article from nancy lemay: “linking a frontpage 2000 web form to an access 2000 database”. the article is ms-linked too as it uses “a microsoft internet server having msfrontpage extensions and msaccess software installed”. but the main point is “to encourage non-technical data specialists to explore this option in order to provide a sophisticated service”. the 1st international conference on e-social science is to be held in manchester (uk) between june 22nd 24th 2005. you can read a small advertisement for the conference in this issue. remember to pay a visit at the iassist website on www. iassistdata.org. the web-crew members are making many efforts in keeping this site attractive and useful for you. papers for the iassist quarterly are most welcome. papers can be from iassist conferences, from other conferences, from local presentation, etc. so please contact the editor (kbr@sam.sdu.dk). you should also be aware that a contest is running. submissions are due january 14, 2005 and announcement of winner takes place on march 1, 2005. the award is $250 us and one year membership in iassist and the competition is not limited to current iassist members. the winning paper as well as other submissions meeting the criteria will be published in the iassist quarterly. karsten boye rasmussen, october 2004. you should also be aware that a contest is running. submissions are due january 14, 2005 and announcement of winner takes place on march 1, 2005. the award is $250 us and one year membership in iassist and the competition is not limited to current iassist members. the winning paper as well as other submissions meeting the criteria will be published in the iassist quarterly. 12 iassist quarterly 2010 / 2011 iassist quarterly abstract the finnish social science data archive started archiving qualitative data in 2003. many researchers found this to be highly problematic. their main reason for opposing the archiving of qualitative data was research ethics. researchers who oppose the archiving of qualitative interviews mainly appeal to the confidential nature of the interview situation. this kind of argument against archiving is put under scrutiny in this article. it covers issues such as the presentation of research subjects and the understanding of research relationship. researchers tend to define the interview relationship as unpredictable and private, and interviewees as helpless participants in need of protection. in contrast, the interviewees themselves define the relationship as an institutional one aiming to foster science. keywords: research ethics, qualitative research, data archiving, interviews methodological and ethical dilemmas of archiving qualitative data the focus of my article is to study ethical and methodological assumptions related to archiving qualitative data in order to question some researchers’ presumptions that archiving infringes on the idea or nature of qualitative research. before discussing the main topic i will characterize a few differences in research culture between the humanities and the social sciences concerning research data archiving. after that i will describe briefly the actual measures that the finnish social science data archive (fsd) has taken in establishing the archiving of social science qualitative data. the different phases and difficulties fsd has had to go through reflect also the general research culture in the social sciences. qualitative data can consist of memoirs, letters, pictures, movies, webpages and audio-visual recordings of different kinds of situations. due to the identifying nature of images and audio recording, they are probably the most challenging material to archive. i will however, concentrate on interviews. researchers often define them as difficult type of data to archive for re-use. those opposing archiving on ethical and methodological grounds perceive a qualitative interview as very intimate, sensitive, un-predictable, emotional and thus infeasible to be archived for re-use by a researcher who has not been in the field doing the interviews. in this article i try to challenge this argument. in addition to reviewing the literature on this issue, i will examine the results i have obtained when contacting a great number of research participants. contacting the participants was done in order to ask their permission for archiving data about them which the researcher had promised to keep totally confidential and restricted to his or her use only. according to the views of research participants, the researchers’ argument against archiving starts to be revealed as a methodological myth: research participants believe they have control over the interview and they do not interpret qualitative interviews as secret engagements that would hinder the archiving of the data for further use. instead, they see open access to research data for further uses as self-evident and a way for them to engage in the advancement of science. differences between humanities and social sciences the research culture in finland has been much more favorable towards qualitative than quantitative research especially since the 1980’s. this trend has been more common in europe compared to northern america where survey methods retained their place within mainstream methodological and ethical dilemmas of archiving qualitative data by arja kuula 1 iassist quarterly 2010 / 2011 13 iassist quarterly methodology (alastalo 2008). one reason for setting up the fsd a decade ago was to foster quantitative and comparative research in finland in a situation where new researchers seemed to have less ability and willingness to use statistical methods than previous generations. fsd has succeeded in its task of fostering quantitative research. at the same time fsd has maintained the idea of fostering the re-use of qualitative data as well. that has not been an easy task. in spite of a wide-ranging collection of finnish qualitative method books and internationally famous methodologists – such as pertti alasuutari or anssi peräkylä – we do not have traditions of sharing, reusing or archiving qualitative data in social sciences. the situation is somewhat different in the humanities. in humanities much research data, like sound records or different kinds of folklore and interview datasets, are archived in small departmentbased university archives, such as the archives of the turku university school of cultural research (see mahlamäki, 2001). in addition to department archives there are larger archives in humanities that have material not only in paper but also an increasing volume of electronic data. for instance, the folklore archives – a “finnish cousin” of the british mass observation archive and the research institute for the languages of finland both have a respectable tradition of archiving qualitative research data and strategies in order to follow the developments in the digital era. comparing the research culture in humanities with social sciences is illuminating in the context of data archiving. in humanities, research data are considered to be testimonies that ought to be available in case someone wants to check the interpretation and results of a published research. in humanities the research data are also seen as valuable common resources that ought to be preserved if we are to understand and study our culture and history. by contrast, in the social sciences data are seen more often as private property. the legal aspects of research data are also assessed differently. social scientists more often emphasize privacy issues, while in the humanities it is more common to stress the significance of research participants’ copyright instead of data protection. recently the emphasis on privacy and identification as a risk has been challenged in social sciences as well. there is a growing number of examples where research participants have expressed a wish to be referred to by their real names in research publications. this sign of the cultural change in defining the boundaries between privacy and publicity is not peculiar to finland. the same phenomenon has been reported elsewhere (grinyer 2002; wiles et al. 2004; kobayashi 2001; kelder 2005). first steps towards promoting re-use the finnish social science data archive first started to promote the re-use of qualitative data by developing and maintaining a database of available qualitative data without archiving the data itself. it proved to be a difficult task. the data collected in social sciences were mainly in the hands – or at homes, in attics or summer cottages – of the original researchers. it was very difficult to get the basic documentation of datasets and even more difficult to persuade researchers to give information about their data to a public database. researchers realized it would have meant extra work for them if someone had been interested in their data. humanities archives proved to be the most cooperative in collecting information about available datasets. starting to co-operate with traditional archives in humanities was a reasonable solution. resources in those archives were very interesting, including large collections of ordinary peoples’ accounts, writings and memories. the archives were also happy to extract and give basic information about those collections that we identified as potentially valuable to social scientists. the documentation in traditional archives had been based on dataunits – for instance documenting and key-wording each life story of an immigrant instead of documenting the whole collection including the writing instructions that were given to immigrants. the archives saw the extraction of basic information of certain data collections according to the ddi documentation2 as an interesting way of promoting the use of their resources to broader audiences. the database consisting of 30 documented collections of traditional archives was for fsd a way to give social scientists a concrete idea of documenting qualitative data and promoting its re-use. the archiving of qualitative data commenced in fsd in 2003. after that decision the even more demanding work began of trying to get qualitative social science datasets archived. researchers we contacted were concerned about several issues: the actual usability of their old datasets (either depending on subject matter or it-problems) and the inadvertent misuse of data or unclear agreements on ownership. the most common reasons were concerning ethics, confidentiality and data protection. researchers considered those the foremost reasons not to be able to archive their old qualitative datasets. in addition researchers often appealed to a basic premise or philosophy of qualitative research: that data from such research would not be suitable to be archived for use by the broader scientific community. since it was difficult to persuade researchers to archive their data, we contacted finland’s widely distributed daily newspaper’s weekly supplement editor. the weekly supplement nyt had conducted several internet surveys which included many open-ended questions. the first qualitative data catalogue of archived datasets was made from those surveys. the data catalogue was not very large, 15 datasets, but it included various subjects, such as experiences of domestic violence, alcohol and drug use, sexual identity, living with depression, being a mother for grown-ups etc. those datasets were our qualitative seed corn with which we were able little by little to show that qualitative data can be re-used since, in fact, they were in demand for methods courses and research purposes. another – and still a continuing – problem has been to smooth out the methodological prejudices of researchers doing qualitative research. in order to resolve this, fsd has gained knowledge about data protection and research ethics. by now fsd is considered one of the main information services when it comes to data protection and ethics concerning collecting, processing and re-using data in social sciences. we have extensive web-resources for researchers in finnish and a few also in english. the most often used are guidelines for informing research participants3 and guidelines for anonymization of data4 . the administrative and technical infrastructure for archiving and re-using qualitative data is excellent since in finland qualitative data archiving was embedded into the systems built for quantitative data archiving in the finnish social science data archive. at the moment fsd has 115 archived qualitative datasets and yearly around 50-60 datasets are ordered for re-use. but the culture of archiving and re-using of qualitative research data is still only slowly emerging. researchers need to be further assured about the advantages of archiving and especially not to over-exaggerate the ethical concerns related to archiving. 14 iassist quarterly 2010 / 2011 iassist quarterly the most important opportunities to promote and discuss in depth the ethical and methodological dilemmas related to data archiving have been provided through ethics courses and national seminars targeted at finnish researchers. during the last few years representatives of fsd have been in many of those events – and we are invited as speakers and lecturers with increasing frequency. a later part of this article will concentrate on methodological and ethical issues that have been most often discussed with researchers in those events. instead of referring to these informal conversations, i draw upon the published articles from britain where discussion of the issue has been active since esds qualidata was founded and especially after the economic and social research council set up its data policy in 1995. methodological prejudices towards archiving methodological obstacles connected to archiving have been discussed extensively in britain (mauthner and doucet 1998; mauthner, parry and backett-milburn 1998; parry and mauthner 2004; richardson and godfrey 2003; bishop 2005). such an energetic discussion has not occurred in journals in finland, but informal discussions with researchers are reminiscent of the debate in britain. researchers seem to be concerned about whether potential re-users of datasets will be able to follow the basic ethical norms which advise researchers not to compromise anonymity, privacy or confidentiality of research participants. this ethical concern is complemented by the assumption that qualitative research – or at least the qualitative interview relationship – is very open, confessional, truth-telling, intimate and sometimes emotional. thus opponents of archiving cite the power of the method as unpredictably revealing and positions research participants as vulnerable and lacking power or at least lacking competence to control their speech in research situations. richardson and codfrey (2002) assume that the well-being of research participants might be compromised if transcribed interviews are archived. they claim that one ethical risk of archiving is the possibility of identification. the other presupposition is that the integrity of research participants is violated when a researcher they do not know beforehand analyses confidential data. another type of risk is raised by mauthner, parry and backett-milburn (1998) and parry and mauthner (2004), who discuss the methodological obstacles to rigorous and truly self-reflective research with archived qualitative data. in both articles, archiving is placed in the realm of positivism and realism. the risk they talk about is the possibility of forcing qualitative data into rational, logical and partial datasets which do not represent the personal, in-depth, messy, haphazard, intuitive and creative real nature of that data (mauthner et. al. 1998; mauthner and doucet 1998). the archiving of data may compromise the quality of future interviews since researchers know that the interviews will be subject to scrutiny by other researchers. that may lessen the rapport in interviews, as well (parry and mauthner 2004the risk formulated in these articles is the possibility of revealing researchers’ professional performance, with the implication that it will be found wanting. although the expressed worry is articulated as a need to protect research participants, it may be that the unspoken real risk researchers attach to archiving is the unforeseen or unpredictable criticism by competitors or malicious researchers. despite the problems caused by a competitive research culture, transparency of research process is acknowledged as an essential part of science. for example, social scientists researching health care think that public matters, including public documents and professional performance of doctors, should be accessible to debate and scrutiny (hoeyer et. al. 2005). according to this logic, it is questionable that a researcher doing qualitative interviews acts in his/her own right in a private and individual role while doing publicly funded research. mauthner, parry and backett-milburn (1998) claim that qualitative data are not suitable to be archived because using archived data is incompatible with the interpretative and reflexive nature of the research paradigm. discussions of the difficulties of getting enough context information for re-use support that opinion. bishop (2006 and 2009) and moore (2007) have written responsive articles about this subject, but many researchers still think only they themselves are capable of using their data correctly. it is true that an interviewer can perceive and partly interpret the emotions, expressions and exclamations of the interviewee. social interaction may contain elements that are difficult to express verbally. however, researchers often employ field or research staff to collect and process the data. at the analysis stage, even those researchers who have personally collected the raw data mainly work with material derived from it. according to the conventions of science, researchers must be able to verbally express and validate all interpretations of data – including those formed in authentic situations – in their research reports. the idea of “pure” or “original” data is simply not feasible. research data are always a construction, as bishop (2006) says. the perception behind the idea that the original researcher is the only one capable of analyzing the data correctly means that the original methodology is the orthodox way to understand research data. what this implies is that the original researcher has an exclusive right to define the characteristics and nature of the empirical world under investigation. that is an odd presupposition for a research paradigm that often accuses quantitative research of naïve realist epistemology. there are few empirical methods in social sciences that can be defined as neutral or unbiased. even the ethnographic gaze is always partial, not all-embracing. it is good to keep in mind that re-use of qualitative data is never a replication of qualitative research. researchers re-using ethnographic field notes and interview transcriptions cannot claim to be doing ethnography himor herself. re-use is always partial and most of all, it is usually asks quite different questions from the original research. even in the case of quantitative data, pure replication of research is very rare. independent of method or data, researchers may have theoretical or ideological standpoints that affect the analyses process so that it is impossible to replicate the original research. most re-use of archived data focuses on different kinds of research questions and methods of analyses than the original research did. for instance, original research may have concentrated on memoirs of women living in the countryside, using long in-depth interviews to study the impact of the environment on the identities of the women. a re-user of that dataset may use parts of the interviews as additional comparative data for a study that collects primary data as well, and focuses on the definitions of mother-daughter relationships. if the dataset is well-transcribed (or preferably with audioor audiovisual recordings as well) there are many possibilities for analyzing emotions between the researcher and participant, or to carry out interaction analyses (southall, 2009). according to this view re-using qualitative data is more of a practical issue than an epistemological one. to ensure that data are reusable for further research, there must be sufficient documentation on the context of the research and on how the data were collected (fielding 2000 and corti 2006). iassist quarterly 2010 / 2011 15 iassist quarterly interviewees’ perceptions of research interaction those opposing archiving on methodological grounds seem to imply that some kind of deception occurs in these methods of reusing data. if research participants talk in an emotionally uncontrolled way, researchers seem to feel the need to protect research participants, and one way to protect them is to prevent the archiving of data. but do research participants lose their ability of control their speech in a research context and will they be hurt by the analysing gaze of a researcher unknown to them? very little empirical research has been done on this kind of research experience, but luckily there is one study. it is a british report called “ethics in social research: studying the views of research participants”, published by the national centre for social research (graham, grewal and lewis 2007). the study sought to look at research ethics from the perspective of research participants and to identify their ethical requirements. it consisted of 50 in-depth interviews with adults who had recently participated in research. ten participants in each of five studies were interviewed. they had participated in either qualitative or quantitative studies. the results showed that the interviewees had ways of withholding information if they so wished even though they had not said explicitly “i do not want to answer or discuss this topic”. participants told how they had given misinformation and how they sometimes had held back or gave an outline of a reply but no details. in addition, they explained how behaving in certain ways, for instance, showing discomfort, affected the interaction and pushed the interviewer to move on so they did not need to reveal personal information concerning the issue at hand (graham, grewal and lewis 2007these results show that research participants are not vulnerable persons who can be exploited by qualitative interviewing. on the contrary, participants seem to be quite capable of using different strategies to control their privacy. the report also enquired if people thought that asking upsetting questions could be justified. the general view was that it is justified to ask upsetting questions provided certain conditions are met: the research is important and worthwhile; people know the topics beforehand; interviewers are skilled and alert to how participants might be feeling and able to respond sensitively (graham, grewal and lewis 2007). the results above remind me of several conversations that i have had with researchers on ethics courses about the problems they have faced in their fieldwork. it seems that researchers tend to think there are ethical problems with their research every time an interview rouses emotions and especially when they themselves are emotionally and feel unable to help participants who have experienced difficulties in their lives. suffering can sometimes be transmitted, or at the very least make the researcher empathetic and sad. still, emotions are normal in research interaction in the same way as they are normal in everyday interaction when dealing with different aspects of human life. corbin and morse (2003) have reviewed several research publications which have been based on qualitative interview data of sensitive issues – such as recalling traumatic experiences in life. they found no evidence of interviews having caused long-term harm or that participants required referral for follow up counseling: “in fact, even though participants experienced some degree of emotional distress during and immediately afterward, the anecdotal evidence suggests that interviews are more beneficial than harmful” (corbin and morse 2003: 346). thus the seeing and feeling of emotions does not pose an imminent threat of ethical problems or risks in the research. interviewees’ perceptions about archiving researchers collecting qualitative data often assume that research participants would not accept the idea of archiving. to check this assumption, we in fsd have asked a few researchers to let us recontact their research participants. the researcher and i wrote a letter together to participants reminding them of the research project and telling them about the possibility that their data would be archived if they consent. in our telephone calls to selected research participants, we have been able to talk about the research, archiving and the terms of the future use of the data. we have re-contacted participants of four datasets. three datasets were interview studies and one consisted of university students’ written life stories. one interview dataset consisted of discussions of equality and gender issues in working life, another concentrated on environmental conflicts, and the third focused on the life and experiences of women living in the finnish countryside. it is almost never possible to locate all research participants after a study has been completed. we were able to find the addresses and re-contact 169 research participants, 165 (98%) agreed to archive their data and only four did not accept the idea of archiving. one can always ask whether these particular datasets were for some reason regarded as non-sensitive by the research participants. however, all the datasets included unique and personal stories, and occasionally sensitive experiences about the issues at hand. the interviews of rural women had taken two to four hours and were very candid. the participants had spoken widely about the joys and miseries of their personal lives. despite initial concern that consent would not be granted for this dataset, every one of those women agreed to the idea of archiving the interviews for future research purposes. during my phone conservations with the research participants i learned that for them the main reason to give consent for archiving seems to be a wish to advance science. people had participated in the research because they had thought the subjects of the interviews were worth studying. giving consent to archiving meant continuing to fulfil this wish. one research participant also said that the original research results did not convince him, and he warmly welcomed re-analysis by different researchers representing different disciplines. in fact, a few were a bit irritated by my contact since they had already made the decision to advance research and did not think that archiving and re-use by other, as yet unknown, researchers would conflict in any way with the original participation decision, no matter that the original researcher had said that she or he would be the only one to use the data. one person interviewed about gender issues and discrimination in working life, laughingly asked “what kind of a risk or harm could a university researcher possibly pose by studying my ten-year-old words, thoughts and experiences?’ the idea that the wellbeing of the participant could be compromised by allowing a third person to study and analyse the interview material was not a consideration. through this exchange i started to realize how differently researchers and research participants define the research relationship. it is worthwhile to note that research participants perceive open access to research data for other researchers as self-evident. that kind of perception of research data implies a certain kind of understanding of the relationship between researcher and the interviewees. we can naturally speculate about the extent of the research participant’s knowledge of the imaginable risks and harms that archiving may lead to. another possibility is that they do not regard the interview relationship as 16 iassist quarterly 2010 / 2011 iassist quarterly private or secret. for them, the interview relationship is an institutional interaction. the perception of research interview as institutional interaction supports the idea of participants as conscious subjects, not as ignorant or vulnerable people in need of protection. corbin and morse (2003) also point out in their article that research participants are given control over the course of interview and participants know that they are telling their experiences to an audience, even if during the interview there is only an audience of one interviewer. towards a reasonable perception of confidentiality most qualitative researchers have told their research participants that the people who collect the data will be the only ones using it. one reason for doing so is the presupposition that this way they will get more authentic and candid data. the other reason is the implied nature of qualitative interviews: they are perceived as being sensitive, intimate and thus fully confidential. as the previous results of research participants’ attitudes show, the participants can control their communication and they do not perceive the research data as secret and limited to the use of the original researcher. defining the research interview as an institutional interaction does not mean that qualitative interviews could not be confidential and include personally sensitive information. neither does it rule out unpredictable emotional investments by interviewees. it only means that the interaction is predefined as a research encounter whereby a researcher represents the institution of science. the interview is not to be taken as a casual conversation between two or more individuals in a private situation. unless it is a research design involving deception, both parties define the interaction as belonging to the domain of research. participants are fully aware that they are talking to a researcher for research purposes. as natasha mauthner and odette parry (2004,) say, the joint construction of qualitative data between researcher and respondent has important implications for the ownership and control of research data. because of that we should also respect the perceptions of research participants. disrupting peoples’ ordinary life by doing a qualitative interview can be tiresome and exhausting, especially if the interview proves to be long and emotionally stressful. after having invested their time and emotions in order to promote scientific research, people rarely appreciate the view that the data can be used for one research project only and at worst, only partially even for that project. if the views of research participants referred to in this article reflect the attitudes of people participating in research in general, we have to define in a more exact manner what confidentiality actually means. instead of secrecy, confidentiality should consist of agreements between the researcher and participants on the future use and preservation of the data. confidentiality would then mean that when data are collected for research purposes the data could be archived and used for further research unless otherwise agreed with research participants. confidentiality does not mean an all-inclusive secrecy that would hinder the archiving and future research use of interviews. but confidentiality certainly does mean that identifiable personal information gathered during an interview cannot be delivered or presented as such to the media or, for example, or to administrative officials for decisions concerning individual interviewees. it is usually the researcher who defines what confidentiality means in each case. researchers who perceive qualitative interviewing as private and secret tell this in the beginning of the research to the interviewees as well. i recommend that the starting point in defining confidentiality ought to be the archiving of data for broader research use by setting reasonable conditions for the secondary use of data. that would be practical and useful. respecting the research participants’ self-determination in defining the value and usability of data would be ethical as well. list of references alastalo, m. (2008). ‘the history of social research methods’ in the sage handbook of social research methods. alasuutari, p, bickman, l and brannen, j.(eds) london: sage. pp. 26–41. bishop, l. (2005). ‘protecting respondents and enabling data sharing: reply to parry and mauther’. sociology. 39(2). pp333–336. bishop, l. (2006). ‘a proposal for archiving context for secondary analysis’. methodological innovations online 1(2). [online]. available at: http://sirius.soc.plymouth.ac.uk/~andyp/viewarticle.php?id=26 bishop, l. (2009). ‘ethical sharing and reuse of qualitative data’. australian journal of social issues. 44(3). pp255-272. corbin, j and morse, j m. (2003). ‘the unstructured interactive interview: issues of reciprocity and risks when dealing with sensitive topics’. qualitative inquiry. 9(3). pp335–354. corti, l. (2006). ‘editorial’. methodological innovations online. 1(2). [online] available at: http://sirius.soc.plymouth.ac.uk/~andyp/viewarticle.php?id=33&layout=html fielding, n. (2000). ‘the shared fate of two innovations in qualitative methodology: the relationship of qualitative software and secondary analysis of archived qualitative data’. forum qualitative sozialforschung / forum: qualitative social research, 1(3). available: http://www.qualitative-research.net/fqs-texte/3-00/3-00fielding-e. htm graham, j., grewal, i. and lewis, j. (2006). ‘ethics in social research: the views of research participants’. prepared for the government social research unit by national centre for social research: london. grinyer, anne (2002) ‘the anonymity of research participants: assumptions, ethics and practicalities’. social research update. (36). [on-line]. available at: http://www.soc.surrey.ac.uk/sru/sru36.html hoeyer, k, dahlager, l and lynöe, n. (2005) ‘conflicting notions of research ethics. the mutually challenging traditions off social scientists and medical researchers’. social science and medicine. 61(8). pp1741–1749. kelder, j. (2005). “using someone else’s data: problems, pragmatics and provisions” forum qualitative sozialforschung / forum: qualitative social research 6(1). available: http://www.qualitative-research.net/ fqs-texte/1-05/05-1-39-e.htm kobayashi, a. (2001). ‘negotiating the personal and the political in critical qualitative research’. in qualitative methodologies for geographers. limb, m and dwyer, c. (eds) new york: oxford university press. pp. 55–70. iassist quarterly 2010 / 2011 17 iassist quarterly mahlamäki, t. (2001). ‘from field to the net: cataloguing and digitising cultural research material’. iassist quarterly. 25(3). 25-28. available: http://iassistdata.org/publications/iq/iq25/iqvol253mahlamaki.pdf mauthner, n and doucet, a. (1998). ‘reflections on a voice-centred relational method of data analysis: analysing maternal and domestic voices’ in feminist dilemmas in qualitative research: private lives and public texts. ribbens, j and edwards, r. (eds) london: sage. pp119–144. mauthner, n, parry, o and backett-milburn, k. (1998). ‘the data are out there, or are they? implications for archiving and revisiting qualitative data’. sociology. 32(4). pp733–745. niamh m. (2007). ‘(re)using qualitative data?’ sociological research online. 12(3). available: http://www.socresonline.org.uk/12/3/1.html parry, o and mauthner, n. (2004). ‘whose data are they anyway? practical, legal and ethical issues in archiving qualitative research data’. sociology. 38(1). pp139–152. richardson, j, and godfrey, b. (2003). ‘towards ethical practice in the use of archived transcribed interviews’. international journal of social research methodology. 6(4). pp347–355. southall, j. (2009). ‘is this thing working? – the challenge of digital audio for qualitative research’. australian journal of social issues. 44(3). pp321-334. wiles, r, charles, v, crow, g and heath, s. (2004). ‘informed consent and the research process’. paper presented at the esrc research methods festival at the university of oxford, 2nd july 2004. notes 1. arja kuula has a phd in sociology and works as a development manager in the finnish social science data archive. she is responsible for the archiving processes of qualitative data and information service on research ethics, privacy protection and copyright issues relating to both quantitative and qualitative data. in 2006, she published a handbook on research ethics and legislation regulating data collection and re-use. she has been a member of the finnish national advisory board on research ethics 2/2007-1/2010. arja.kuula@uta.fi 2. the data documentation initiative (ddi) is an effort to create an international standard for describing social science data. expressed in xml, the ddi metadata specification supports the entire life cycle of social science datasets. even though it is most suitable for quantitative data, the standard can be used in describing qualitative datasets. (for more information see http://www.ddialliance.org/) 3. for further information see, http://www.fsd.uta.fi/english/informing_guidelines/index.html 4. for further information see, http://www.fsd.uta.fi/english/anonymisation/index.html vol30-3.indd 4 iassist quarterly fall 2006 editor’s notes welcome to the second issue of the iassist quarterly, we are all interested in the future, because that is where we intend to spend the rest of our lives. at the 1999 iassist conference, there was a session called “bridging the past with the future” which featured central and important iassist people assembled to celebrate the 25th anniversary of iassist. iassist began as you will learn in this issue in toronto at “the meeting in the bar”. iassist 1999 was appropriately held in toronto and this special celebration session was chaired by laine ruus from the university of toronto. the session involved looking backwards and forward through the following presentations: “iassist: origins and evolution as revealed in its archives and other materials” by margaret adams, “early iassist as recalled by its first president” by carolyn geda, and “social research infrastructure from a european perspective” by ekkehard mochmann. better late than never, we are now able to present the revised manuscripts from the american side. these pages summarize original material and present eyewitness observations, some of which had been in the personal information processor for longer than the 25 year span that was commemorated. in the early years, iassist had an archivist and i’d like to point out that the position is currently open. as the adams’ article demonstrates, there is a lot of material available on the formation and establishing of iassist for anyone who wants to continue documenting the organization’s history. the young readers might not fully grasp that when we use the words “mail” and “mailing”, we are in the area of physics, the area of moving around physical envelopes with physical paper often by foot for the last bit. communication was slow; not instant like we are used to now that we can’t imagine a world without e-mail. the technological revolution and the practical ease nowadays might make you feel that “i was so much older then, i’m younger than that now”. however, our kids will without being asked tell us differently! the first article is a contribution from margaret o’neill adams: “the origins and early years of iassist”. in it, adams compiles and researches the information found in the iassist archives and presents a view of how “assist, international assist, i-assist, and iassist” started. no single person is given the credit for inventing the brilliant acronym. the adams’ article also explores some of the central issues, conflicts, and relations to other organizations. from the beginning, iassist has been an organization with people as members. carolyn l. geda, the first president of iassist looks further into the issue of the acronym in her article “recollections of the formative years of iassist”. she credits no particular person or people, which, in my view, could be taken as an example of the solidarity within iassist: that iassist is a group consisting of individuals but often acting as a group or organization. in the article, carolyn geda writes: “the first task of the committee was to construct and agree upon an acronym. this acronym evolved at the infamous bar in toronto. once the acronym of iassist was agreed upon, we had the problem of finding appropriate words for it. although we were very pleased with the acronym, we frequently had and continue to have problems remembering the actual name of the organization!” geda’s piece demonstrates the “international” character of iassist when mentioning how regions where supposed to be involved in all action groups and how this evolved into regional secretariats for recruitment, fee handling and also regional meetings. the third article is also from our very own world. it is not about 25 years ago, although the data are a bit aged. this is a self-look or what could be called an organization’s narcissistic research. in 2001, repke de vries and i collected information through a questionnaire about the use and the future use of the iassist web-site. some of this information has been presented earlier, and a lot of good people have worked hard and improved the iassist website tremendously since then. i gave the presentation a new wrapping when i used the information at the 2007 iassist conference in the panel session “care and maintenance of a global knowledge community”. more content was added and we ended up with the second article on iassist as a virtual community with the title: “open virtuality or virtually open? openness on the web as viewed by the iassist membership.” among the issues the article addresses is the balancing of access for and openness towards non-members of the iassist. the iassist acronym is indeed very, very good. it’s so good that others are using it too. ever heard of “the internet server manager for mac os x”? or how about “sj asset management iassist”. or iassist.org, not to be confused with our iassistdata.org? or “ iassisthr a virtual assistance firm”? the “iassist.co.uk”? or “iassist. ca”? etc. most of these are new, small, or discontinued. our iassist is certainly still here and with a history to tell. remember to have a look at the website http://iassistdata. org articles for the iassist quarterly are most welcome. articles can be papers from iassist conferences, from other conferences, from local presentations, discussion input, etc. contact the editor via e-mail: kbr@sam.sdu.dk. karsten boye rasmussen, august 2007 catching reader responses on the fly by rosanne g. potter' department ofenglish iowa state university ames, iowa 50011 literary criticism, the art of making discriminating judgments or evaluations of literary works, depends in large measure on perspective. reader response criticism approaches literary works from the perspective not of the writer, or the period, or the style, but from the perspective of the reader. one looks at interpretations of literary works and questions the knowledge base, interests, and psychological defenses of the readers who devised the interpretation. for more than 10 years, i have been using computational and statistical means to study hnguistic and stylistic features in texts in an attempt to find and quantify textual controls in dramatic literature. i conducted my first pragmatic study of real readers in the early 1980s but, at that time, did not concern myself with gender. in 1990 when i agreed to chair the women's studies program at iowa state, 1 began noticing differences between male and female student readers in my classes; i wondered whether males and females reacted to different cues or responded to the same cues differently and, most important, how i could catch the responses as the readers were reading. this paper describes my investigation of gendered reading; it begins with a brief summary of other pragmatic studies of readers noting the presence, or absence, of sex and gender as factors in these studies. reader response criticism has been the least developed kind of criticism because, until quite recently, readers were thought to make only one contribution to the critical process: they were the sources of error, private associations, and misreadings. these negative judgments of readers fiow from the assumption that there really is a "right" or "most complete" reading, and that, if properly trained, all readers can achieve it this "right reading" assumption has undermined practical critics from i. a. richards to elizabeth flynn and is still very much in practice in most american classrooms. although theorists like wolfgang iser have written book after book about "the reader," the researchers who have attempted pragmatic studies of real readers have been few and their methods painfully unsystematic. the first, and most famous, empirical study was performed at cambridge university by 1. a. richards in the twenties. richards' results were so devastating to him, and to many other, that for fifty years after, no one attempted to study the reading skills being learned in literature classrooms. in his 1929 book practical criticism . richards reported the results of asking cambridge undergraduates to read thirteen poems (authorship not identified and ranging from john donne to minor poets), then to "comment freely" on their "readings." richards documents— in excruciating detail — the many ways of misconstruing meaning in poetry when it is presented "without any hint of provenance"(5). richards presents his findings not to indict the "products of the most expensive kind of education" (292), but to demonstrate that in all types of educational settings "we must cease to regard misinterpretation as an unlucky accident we must treat it as the normal and probable event" (315) he rightly ascribed their poor readings to "bewilderment" (296) caused by the lack of "clues [about] authorship, period, school, the sanction of an anthology, or the hint of a context." (296)^ in her 1978 book the reader, the text, and the poem: the transactional theory of the literary work .^ louise rosenblatt reported on over twenty years of collecting student responses to unidentified poems; as i see it she, like richards, created abnormal test situations. in everyday reading, readers know the name of the author and can easily discover her/his dates, nationality, and school, and may have had their responses shaped by earlier readings of other works by the same writer or by earlier teachings about the writer or the work. by forcing readers to respond to the words of the text only, both researchers deny a reality condition in trying to create an unbiased test situation. this experimental design inevitably sets up perfect conditions for "errors" and encourages the discovery of differences between readers' responses. fifty years after the publication of practical criticism . norman n. holland in 5 readers reading decided to conduct "more or less undirected interviews with a few readers who had taken standard personality tests" (x).* he taped extensive interviews with five undergraduate students on faulkner's "a rose for emily" and then read their responses to the story in light of their "identity themes" (56-62). instead of finding a great deal of overlap between the readers, holland found that the readers perceived very different stories. holland arrived fall 1992 at "four principles that describe the inner dynamics of the reading experience: "style seeks itself (1 13), "defenses must be matched" (1 15), "fantasy projects fantasy" (1 17), and "character trarisforms character" (121). these principles are psychological descriptions of how readers transform the characters and events in stories to defend their own identity themes while reading fictions. holland, like richards, used a free-response method, but holland's method was molded by three interventions richards had not allowed. the readers knew the author and the name of the story (some had even read it before in other classes). the readers' responses and personalities were both elicited by questions from holland and reported by him. this method of gathering reader responses produces, in holland's words, no "uniform core" of meaning "from the text" as opposed to "individual variations" contributed "from the people" (366). thus, while richards is distressed at the general decrease in ability "to make plain sense of poetry" (12), holland explains "the way readers respond to literary characters as if they were real people" (xii) by applying freudian terms ("transformations," "defenses," "fantasies") to the reading experience. although they perceive the outcomes of their experiments very differently, neither richards nor holland has a model that can be replicated, because neither has an experimental design with clear-cut categories for grouping responses. they inevitably emphasize differences among readers because each reader is treated as a separate case rather than has having features that can be clustered with other similar reader responses. in any study where readers respond "freely" and no attempt is made to identify features in the responses which correlate with features in the texts read, the results will inevitably emphasize difference. my 1982 study of reader responses to the characters created in the first acts of 21 modem english-language plays could focus on agreement among readers because it asked all readers to respond to the same questions about seven character traits (dominance, intellect, excitability, speculation, poetry, education, attitude), using the same scale (e.g., markedly dominant, moderately dominant, neither dominant nor dominated, moderately dominated, markedly dominated). the reader responses were correlated to features in the language assigned those characters gike high or low use of questions, imperatives, fragments, exclamations, and seven other syntactic and/or semantic features). since the research design asked specific questions and correlated the results with countable features in texts, there was no difficulty either in finding areas of statistically robust agreement among readers on character traits or in regressing the character trait data against the linguistic features to show which features figured at what levels in readers' judgments of character traits. literary scholars who do not know about dependent and independent variables or about objective methods of handling data, and who never attempt to gather qualitative information in quantifiable ways, are destined to discover, as holland did, responses that have "nothing in common" (366). up through holland, no particular interest in sex or gender as variables shows up. if women participate in the studies, that fact is either not noted or not considered a significant enough factor to require any balancing of the groups being tested. in the 1980s, the idea of genderbalanced samples begins to emerge, but since the general research methods continue to be highly subjective, the introduction of this variable hardly matters; the presence of sex as possible variant does not change the basic methods of analysis. in elizabeth flynn and pau'ocinio schweickart's important 1986 book gender and reading . david bleich reports on conducting an admittedly unscientific study of four females and four males (one of whom was himselq in an attempt to discover differences in male and female ways of reading male and female writers. bleich's general conclusions are that males and females respond in similar ways to lyric poetry by male and female writers, but very differently to fictions. men conduct a dialogue with the author about the characters and situations, while women enter the fictional world and allow themselves to see feelings more quickly than men do. this graduate-student pilot study established the theses to be tested in a larger, apparently scientific, contrast of the retellings of faulkner's short story "bam buming" by 00 males and 00 females. bleich, whose most famous book is entitled subjective criticism , docs not use objective methods for handling the responses he collected. instead he reads the responses subjectively and finds, not surprisingly, that they not only confirm his earlier findings, but also allow him to go on to even larger generalizations. the narrator's voice is the "mother tongue" and since separation from the mother is less significant for women than for men, women perceive men as "less-other" than men perceive women. these assertions may be true, but we should not be misled into thinking that the use of objective methods to collect data means that bleich's generalizations have any more truth value than if he had arrived at them without consulting any readers. actually the data he collected and the conclusions are structurally unrelated. starting from a gender-balanced sample does not necessarily prove anything about gender. skill in research design is not widely distributed, especially among literary critics who have no training in even recognizing a well designed project when they read it. 12 lassist quarterly in her essay on "gender and reading" in the book of the same title, elizabeth flynn describes collecting large numbers of responses from males and females (even though that was quite inconvenient at a mostly male school) and using random selection to achieve a balanced number of responses. unfortunately, she also apphes no objective tests to the data so carefully assembled. like the male critics who preceded her, flynn reads the essays and judges their adequacy as readings against her own critical standards supplemented by psychological terms describing human interactions. examples of "domination" by or "submission" to texts are quoted and dismissed in favor of a "third possibility" in which "the reader learns from the experience without losing critical distance; reader and text interact with a degree of mutuality" (270). flynn concludes by asserting that many males react to disturbing stories by "rejecting" them in an attempt to "dominate" the text, and females "more often arrive at meaningful interpretations of stories because they more frequently break free of submissive entanglement in texts" (285). like holland's, flynn's conclusions are interesting and both sets of insights may have, as oscar wilde quipped, "the minor merit of being true"' but neither researcher has begun to prove anything. they have simply used a new method of establishing "authority" or ethical appeal. at this point, 1 wish i could say and here, tah dah, is the reader response study that does what none of the others have done, but i come more to discuss the theoretical issues that are at stake when gender appears in empirical studies of reader's responses to literature, than to make final report on research. since my work on male and female responses to modem british literature texts is still under development and has only been described in print in a belgian journal, i will describe it here in some detail. the study grew out of two occurrences in a 1990 modem british literature course in which the students wrote reader responses to each assignment before class discussion. the first reader response, to chapters 1 through 3 of oscar wilde's the picture of dorian gray , surprised me because the readers' responses seemed to be sex-linked. both male and female students commented on the "howery" language and the ornate tone of the writing, but the males then asked, and i quote, "is this guy queer?" or asserted "this guy must be a homosexual," while the females said "he is very sensitive" or "very poetic." in class, we discussed what they were reacting to and why some males drew conclusions that no females did. i gave the students the "sexual facts" on wilde: that he was a married man and, as far as his biographers can discem, had not had any homosexual experiences at the time of the writing of this novel; that he did subsequently have such experiences and, five years later, was imprisoned for "gross acts of indecency," the 1890s' code term for homosexuality. a month or so later, the second incident occurred. the students had read their first selection from bemard shaw's an intelligent woman's guide to socialism. capitalism, etc. (chapters 7-12"). here the difference was much more pronounced and intense: the female students related very positively to the text: they felt that shaw understood the realities of women's lives and was arguing for improved economic conditions for women. a strong majority of the male students felt that shaw was condescending to women and treating them as little better than children.* this difference empted into a fullscale classroom confrontation and to the discovery that the males and females were in some cases using the same passages to prove their different positions. the strongest male and female speakers both wrote papers in support of their readings. the female, who was simultaneously attending a senior seminar "language and gender," designed a questionnaire using selected passages (ones that "proved" her point, ones that "proved" the male student's, and ones that both asserted "proved" their opposing positions) and administered it to male and female student friends. although the sample was small and neither stratified nor random, the results provided more anecdotal support for the proposition that male and female readers drew opposing conclusions from their readings of the selected passages. as a result of these two cases of striking gender differences in reader responses to literary texts, i designed an interactive reading exf)erimeni for use during the next offering of the same course. the project 1 designed in 1990 and ran in 1991 goes back conceptually to michael riffaterre's 1959 insight ("criteria for style analysis" word) about the existence of places in literary texts that are commented on by almost all average readers (or ars). riffaterre asserts that disagreements (among critics about what passages mean) proves that stylistic devices (or sds) have surprised readers. the unexpected use of language, according to riffaterre, elicits interpretation.' i wanted to catch ars responses to sds in their first reading by inducing them to read new texts on a computer screen and to respond to anything that they found "surprising or unexpected."* my assumption about the best way to get a response (without interrupting the reading process) was to ask readers to take a simple action while reading, e.g., to "double-click" on words that seemed surprising or unexpected. there were six readings: balanced by gender of writer (three females, three males) and balanced by genre within each gender (two works of fiction, one of nonfiction). the males (oscar wilde, bemard shaw, and james joyce) were all from dublin, though from differeni classes and social spheres; the females (katherine mansfield, enid bagnold, and virginia woolo, were from london and wellington, new zealand, were all upper-middle class.' the students self-selected into either mac-lab readers and control-group readers. the mac-lab readers came to the lab, entered a few demographic facts (their sex, age, major, and home state) in a logon procedure, and read for the first time an initial segment (chapter, or part of chapter) or a complete short work. as they read, they double-clicked on words, thus highlighting them and (although this was not spelled out to the students), simultaneously moving each highlighted word into a list tagged with the student's demographic facts. after completing the reading, the students were asked to reply to six forced-choice post-reading questions designed to estabush whether they (1) enjoyed the reading (i.e., did they want to read more by this author), (2) were experienced readers of texts like this one, and (3) felt competent (in terms of vocabulary and general comprehension) in the face of this text they were then asked four expansion questions ("what is this work about?" what do you think of the writer?" "what do you notice as repeated?" etc.) and finally, they were asked to write a short paragraph of reaction to the reading. the control group read the same assignment before class, discussed it in small groups, and then wrote about it again afterwards. a contrast of the discursive writings by the mac-lab and the control group students was anticipated but, unfortunately, not performed in the pilot stage. the mac-lab readers were overwhelmingly female (18 to 6). the ratio of females to males in this course is routinely 3 to ; the same ratio self-selected into the experimental group. the female/male imbalance was intensified when four males completed only five of the six readings, and one completed only four. i had hoped to find that males and females responded to many of the same words, and that some words were more surprising to females than males and vice versa. the results were so skewed as to be statistically unreportable; the best that could be said for the six lists of "surprising words selected" was that they showed as much variance within sex as between sexes. this study of reader' responses, especially when it is seen in the context of other pragmatic and theoretical approaches to reading comprehension, can be refined in a number of ways. first, it needs to be conducted on larger samples of males and females which, as the research reported on in this paper shows, means moving out of the small upper-division classroom and into the large, always available. freshman english pool. second, more needs to be known about the readers; knowing their sex is not sufficient if readers could be arranged along a scale of more or less "masculine" or "feminine," their responses could grouped to see if the social construct of gender is more useful than the biological differentiation into male and female. third, the whole question of post-reading questions needs thorough re-thinking. many students reported having very little memory of the texts they read on the mac screens.' the students probably experienced some test and time anxiety because they knew they would have to answer questions after the reading; these anxieties may have interfered with their responses, their comprehension and subsequent memory of the texts. possibly, the most important re-design would be to presegment the texts, so that passages, rather than words, would identified by clicks. word orientation tended to mean that "difficult" vocabulary items (including british spellings or usages) dominated the word lists. passage orientation would group responses so that students who respond more slowly (at the end of a striking passage rather than near the beginning) would still be counted as responding to the same stimulus as the quicker, more experienced readers.'" this first of these improvements fiows directly out of the comparison between the small and larger studies; it is prima facia clear that if one wants to study male versus female readers, the numbers of males and females need to be larger and more balanced. (bleich's report on four subjects is used merely to investigate gender differences; the second study of 00 students is the one that is supposed to convince.) according to mack shelley, the statistician i worked with on the pilot stage of this project, i would need a sample size of at least 260 responses (balanced between males and females) to have enough degrees of freedom to start getting statistically significant results. the second improvement has to do with triangulation, a research design achieved in my "reader responses and character syntax" project: the relation of two countable features through a third. in my 1982 essay, i counted occurrences of syntactic and semantic features and correlated them with characters who used these features, through readers' judgments of those characters' personality traits. here, i counted the words and the sex of the reader, and probably should correlate these through the reader's scores on a standard test, like the bem sex role inventory. the bem scores could be used to arrange the readers along a female/androgynous/male spectrum and might conffibute to a better account of within-gender variance. the third improvement, eliminating the post-reading questions, would keep the readers' attention focussed on the reading. the students were certainly distracted from selecting words as surprising, because they were concerned about whether they would have enough lime to lassist quarterly complete the assignment in the class period. my study of gendered reading has just begun. the futurefor studies of gender in reader responses when reader response criticism is fully articulated, it will try to understand how readers' experiences, experiences tied to their class, age, ethnicity, and gender, affect their interpretations of literary works; it will try to assess the impact of information (or the lack of it), inclinations and disinclinations in readers. unless testing techniques are defined that categorize kinds of responses (based on knowledge, personality traits, gender roles), unless reader demographics: sex, age, ethnicity, education are factored in, and until all factors are correlated to features in the texts, reports about the impact of gender on the reading of literature will be based on assertions, not upon research. 1 presented at the lassist 92 conference held in madison, wisconsin, u.s.a. may 26 29, 1992. 2 in her 1987 the return of the reader (london & new york: methuen), elizabeth freund summarizes the "vices" as: "carelessness, self-indulgence and sentimentality [with] arrogance and obtuseness...not far behind." (32) 3 carbondale and edwardsville: southern illinois press. 4 holland had enunciated this model in his 1968 work the dynamics of literary response (new york: oxford university press, 1968). 5 "reader responses and character syntax" in computing and the humanities , ed. richard w. bailey, (amsterdam, new york, and oxford: north-holland). 6 the writings of oscar wild ed. isobel murray, 'the critic as artist, part ii," 278. 7 these responses formed 25% of the student's grade, were collected daily and returned at the beginning of the next class with brief positive reinforcement comments, i.e., no comments on writing problems, or "wrong" interpretations, just positive notes on insights, thought processes, and/or expression. 8 in her best selling 1990 book you just don't understand . deborah tannen's descriptions of the differences in conversation styles may explain for the differing responses. tannen asserts that women's conversational style emphasizes establishing community, while men's conversation style emphasizes competing for authority. the women students may feel recognized and valued when shaw explains how the economic system takes advantage of the unpaid labor of wives and mothers. the men students may feel condescended to, put in the onedown position, when shaw assumes the role of the authority explaining economic relationships to readers who do not understand the subject. 9 when many commentators mention a passage, regardless of whether they say the same things about it, riffaterre asserts that they do so because the passage is surprising and calls for interpretation. 10 relying on common language interpretations of these terms, i chose not to create a stipulative definition. 1 1 the texts were the first chapter of enid bagnold's diary without dates . "bliss" by katherine mansfield, and the first chapter of virginia woolfs to the lighthouse . chapters 1 through 3 of wilde's dorian gray , the first fifteen pages of james joyce's the portrait of an artist as a young man , and chapters 7 through 10 aji, intelligent woman... . each file was approximately fifteen pages long in microsoft word. 12 this was the only required part of the post-reading responses (all others could be skipped); something had to be entered here to logout. 13 the texts were re-read when they came up in their normal places in the syllabus. 14 if 1 choose to pre-segment the texts, that would mean a complete re-design of the project. a taxonomy, like the one described by teresa snelgrove in her essay on george eliot ("a method for the analysis of the structure of narrative texts" in literary and linguistic computing 5 [1990]: 221-225,of narrative-mode tags, possibly supplemented by persuasive-more tags could be employed to segment and pre-tag the texts; then the collected responses to any word within a segment could be accumulated to see whether narrative and/or persuasive modes and gender differences correlate. differences in segmentation and/or tagging choices might also turn out to be gender-matched. if a number of male and female critics segmented and tagged the same texts, similarities/differences could be recognized and added into the variables checked in the responses of student readers. this augmented approach definitely merits consideration. 15 hands on the census: microdata from the 1991 census of population in britain by catherine marsh ' department of sociology and department ofeonometrics university ofmanchester the information dilemma a basic dilemma confronts the producers of data in the public sector. on the one hand, they face users demanding that information collected at public expense be made available to the research community in ever increasing detail. on the other hand they face those concerned with privacy and confidentiality worrying about the possible release of identifiable information about individuals. in short, they have to reconcile competing moral claims: they are caught in the middle between one group of citizens asserting a right to information and another group asserting a right to privacy. the audience at this lassist meeting, as professionals in the business of information acquisition and dissemination, will have rehearsed the arguments many times about the rights to information. it is important to remember that the arguments for a right to privacy and confidentiality are also strong, but have changed their character somewhat during the last decade, and have had a profound influence on the preparedness of state authorities and census agencies to release microdata to the research community. the stale is no longer the only, and perhaps no longer even the major, collector of systematic personal data on individuals. the expansion of information technology has led to a new private industry of information collection and management. many companies take as their base the publicly available state collected information such as electoral registration lists. some then may link to this other published records e.g. on bankruptcies and criminal records. others specialise in the collation of information from pull-out questionnaires in magazines and so-called guarantee registration cards, others bring together personal financial information from major credit card companies and chain stores. sometimes the aim of this data collection is to assess the credit-worthiness of individuals. sometimes it is to facilitate modem direct marketing techniques, targeting prime areas for particular mail-shots. combined with advances in telephone technology, these techniques have become big business. in britain, several companies have datafiles containing information on over 90% of households. most of this information is referenced by a combination of name, address and post-code. there is widespread public concern about the activities of these commercial companies. the british public is very hostile, for example, to the sale of election registration lists to outsiders (campbell 1987). in the lead up to the 1991 census, there were media programmes and newspaper articles which expressed worry about the potential of census information about small areas to be linked to other private information databases both of these private companies. when the recent census (confidentiality) bill was being debated in parliament, certain members suggested amendments which would have made the linkage of census information to fine-grain post-coded information illegal (computer weekly 31 jan 1991). the attempts did not succeed, although plans to release statistics for post-coded areas in england and wales were modified. both the political right and left have turned to the state to protect citizen liberties and rights to privacy. throughout the 1980s, most european countries, unlike the us and canada, have enacted legislation to give citizens rights with respect to databanks of information which may be held about them. the data protection act of 1984 gives individuals in the uk rights to find out what computerised personal information any organisation might hold about them, to challenge it if wrong, and claim compensation if they suffer as a result owners of machine-readable lists are obliged to register any file of identifiable personal information they hold, and to say for what purposes it is held; the register is open for public inspection. the data protection registrar attempts lo police the activities of the information industry; he has, for example, recently attempted to curb the unrestricted use of address-based information for credit-redlining. however, in some ways, the individualistic focus of the existing data protection legislation weakens it as a tool for those seeking to ensure that information released could not be linked to private databanks. rights to know are restricted to rights granted to individuals lo find out what is held about themselves; there is no right for someone wanting to establish how much census information it is safe to release to obtain answers lo such critical questions as how many individuals are covered in the databank and which census variables are held which lassist quarterly would be available for matching purposes. the format of the data protection register does not enable it to be used as a source for answering such questions. nonetheless, in the office of the data protection registrar there exists a team of individuals whose job it is to know precisely who has what computerised information about whom, and to police the workings of the act; in practice they know in broad outline the information gathering activities of all the major data collection companies. as well as worry that information may be passed on to the commercial sector, the other major public concern about census and survey data relates the use to which the government itself might put the information. worries [lave been expressed in britain at the time of previous :ensuses about the passing of identifiable census information to immigration officials or to social security officers. a similar set of concerns were expressed in 1991 over the possibility that census information might be passed to community charge officers, responsible for compiling lists of the adult population in order to collect a new and very unpopular flat-rate local tax based on a head count (the "poll lax"). some compaigners against the tax (e.g. in tottenham in london) exphcitly called for a boycou of the census on these grounds. (it is important to emphasise that the census authorities would never in fact pass on census data to any outside the census office, even to other government departments.) taking a comparative perspective, it does appear to be the case that the british public is more sensitive to confidentiality issues and less willing to trust the census authorities than in other countries. the british public is more concerned about privacy than other anglophone countries; goyder and leiper (1985), for example, did a content analysis of letters to the press about the censuses in 1980 and 1981, and found much more concern over issues of privacy and confidentiality in britain than in us or in canada. given the concern about confidentiality, it is worrying to learn that disbehef in the absolute confidentiality guarantees given by census authorities is widespread. this is illustrated by a recent gallup survey undertaken in both gb and usa. it documents that the level of trust in the census authorities is low in both countries, but lower in britain: the survey shows that the british public is much less likely to believe the confidentiality pledges given by the census offices than the american public, as table 1 shows. in short, the general public has a range of worries that census information, gathered ostensibly for assisting government in planning purjxjses of various kinds, will be circulated to others for purposes which were not declared at the lime when the information was collected. in lights of these worries, it is not surprising that census offices do not just hand out microdata on request. it is also not surprising that the census offices that made the decision to release microdata earlier on were more liberal about what they were prepared to release than those trying to make the same decisions more recently. the resolution of the dilemma with respect to census data faced with the dilemma between rights to information on the one hand and rights to privacy on the other, the census bureaus in different countries have made different responses. in general, the english-speaking countries have tended to give primacy to rights to information, and have made census microdata available in various anonymised and sanitised forms, whereas european ta±)le 1 : perception of census confidentiality in britain and us question: "how confident are you that the census office will not release an individual's census information to other government agencies: are you ... gb usa % % very confident 17 23 somewhat confident 27 44 or not at all confident?" 42 28 (don't know) 14 5 fieldwork dates: march 1991 for gb: march 1990for us: source: gallup political and economic index, no. 368, april 1991 (gb) and gallup reportfor usa reproduced with the kind permission ofuk gallup. summer 1991 countries have been much more exercised about rights to privacy and have generally not gone down this public road. britain, standing as it does with a foot in either camp, has taken a long time to decide which way to go. the united states of america was the first country to release such public use files in the 1960s. from the time when the population census was first computerised in the us, discussion was initiated about releasing forms of microdata to the research community. partly because the administrative culture was open to research dissemination, and partly because of the existence of energetic individuals pushing from within the census bureau, "public use files", as they are termed in the us, were released retrospectively for the 1961 census and have become a routine part of census output. they were followed by canada in the 1970s and australia in the 1980s. however, canada and australia never released as much information, either in terms of sample size, detail of file structure or fine-grain detail of coding schemes as the us public use files (see marsh et al. 1991a for more details). the australian census microdata in particular contains only the state as a geographic identifier. despite similar requests for microdata, many european countries have declined to release samples into the public domain unfettered by limitations placed on uses or users. academic researchers in denmark and sweden may only receive microdata for specified and delimited purposes. some countries allow local state authorities access to census data but deny similar access to academics: local government researchers at all levels in italy can have access to microdata, for example, as can regional governments in spain and some government departments in luxembourg. in germany, anonymised census records may only be released to the communities^ and some countries only release a small subset of the available information in the form of microdata; in france, for example, only a restricted subset of census variables is released publicly as microdata in sampling fractions varying from 0.1% to 25%. (redfem 1987 briefly outlines the rules in each european country.) i am delighted to tell you that for the 1991 census in britain, agreement has been reached for samples of anonymised records to be released. at present the plans only include england, wales and scotland, but representations are also being made to the census office in northern ireland to grant a similar request with respect to the 1991 northern irish census. requests from british academics for census microdata go back at least fifteen years. spurred by interest in the release of public use samples in north america, a committee of interested academics was convened in the mid-1970s to discuss the possibility of obtaining similiar microdata in britain for the 1971 census. one problem which emerged early on was the geographers' demand for microdata was for very large samples (10% or more of census records) with very fine grain geography. during the 1970s, it seems that the needs of geographers dominated the requests, and the gulf between what academics seemed to require and what the census offices felt able to release while retaining confidence in the confidentiality of the records was wide. furthermore, the early discussions about microdata never progressed very far since the legality of releasing microdata under the terms of the 1920 census act was never resolved. there were those inside opcs who argued that release of microdata was permissible under the terms of the act (redfem 1976), but the legality of this move was contested by others at the time. more concerted efforts were made to obtain microdata from the 1981 census. the white paper outlining plans for the 1981 census made it clear that the census authorities were prepared to consider reasonable requests for samples of anonymised records. supfwrt for release of microdata was also available from other sources, some of them somewhat unexpected: from the british computer society team reviewing security provisions for the 1981 census (hmg 1981a), and from these advocating cuts in the government statistical service who argued that microdata could provide a cheap and flexible substitute for tables (hmg 1981b). there was also interest was shown by census office research staff (denham 1986). and requests from academics (e.g. norris 1983) for microdata persisted. a further committee was therefore convened under the auspices of the environment and planning committee of the economic and social research council between 1984 and 1985, to try to co-ordinate a request to be put to the census offices to obtain retrospective samples from the 1981 census. despite receiving evidence from several academics about the value of such data, the committee never reached the stage of putting a formal proposal to the census offices. there appear to have been several reasons for this. first, the demands of the geographers who wanted fine grain areal information to the detriment of detail elsewhere and the demands of social and policy researchers who wanted full household information, if necessary sacrificing geographical detail were never reconciled. second, it proved very hard to get agreement from others in the commercial and public sectors to form a purchasing consortium to buy the proposed data. third, in 1985 it seemed likely that there might be a major 10% household survey undertaken in 1986 which might meet needs more effectively; (in fact this survey was never undertaken). as lime furthermore, as time went by, the value of data relating to 1981 38 assist quarteriy seemed to decline^ the request for data from 1981 therefore lapsed. however, user demand did not (marsh el at. 1988). the economic and social research council therefore renewed its efforts in good time for the 1991 census, and set up a working party to negotiate with the census offices in great britain and to present a formal request. this working parly undertook some systematic work on quantifying the risks of disclosure from releasing census data, and concluded that the risks were minimal. on the basis of that work, it proposed that the census offices release two different files of microdata, one to meet the needs of those who wanted the maximum geographical detail and one to service those whose prime interests were in household structure. the request was presented in a lengthy report in 1989 (published in marsh et al. 1991) which was favourably received by the census offices. the committee also seciu-ed the agreement of the esrc to shoulder the total costs of the purchase if necessary. the census offices sought advice from their solicitors about whether microdata could legally be released under the terms of the 1920 census act they were advised that anonymised microdata came under the general heading of a 'statistical abstract', and could be released without changing the law. the office of the data protection registrar and liberty (the national council for civil liberties as was) were consulted, and neither had any major objections. the proposal to release microdata was therefore mentioned in the white paper outlining plans for the 1991 census (cm 430, 1988), subject to the overriding need to preserve census confidentiahty. it was also commented upon by the british computer society team who reviewed security arrangements for the 1991 census (her majesty's government 1991); since they had supported the idea for the 1981 census, they gave the idea their blessing. in july 1990, the agreement in principle of the registrars general to the esrc request was announced in a written parliamentary answer. background to the british census in the united kingdom, censuses of population are the responsibility of three separate offices: the office of population censuses and surveys under the supervision of the registrar general for england and wales, the general register office for scotland under the supervision of the registrar general for scotland, and the census office in the northern irish department of health and social services under the supervision of the registrar general for northern ireland. the office of population censuses and surveys (opcs) plays a coordinating role in the work of the three offices. the legal firamework for censuses in great britain is the 1920 census act as amended by the census (confidentiality) act 1991. the 1920 act makes provision for population censuses to be taken at no greater frequency than every five years. it is enabling legislation which requires there to be a new census order outlining arrangements and plans for content each time. the registrargeneral is given the responsibility for organising the census, the power to take on the necessary staff, and the right to present the final accounts to parliament under the terms of the 1920 census act, filling a census return is compulsory for householders. many in census offices feel that the compulsory natiu'e of the census puts it in a different moral category when it comes to the release of microdata. the logic of this position is not entirely clear, however, since similar confidentiality guarantees are given to those who take part in voluntary government surveys*. the registrar general is given the duty to ensure that summary reports of census data re prepared.furthemiore, "the registrar-general may, if he so thinks fit, at the request and cost of any local authority or person, cause abstracts to be prepared containing any such statistical information, being information which is not contained in the reports made by him under this section and which, in his opinion it is reasonable for that authority or person to require, as can be derived from the census returns." [section 4(2)]. the interpretation of this section of the act was critical for the release of microdata; it turned on whether microdata could be deemed a "statistical abstract". both by international standards and by comparison with previous british censuses, the censuses of 1981 and 1991 were fairly slim. in 1991 there were 8 questions about housing and 19 questions about individuals. there has never been an income question nor, since 1851, any question on religion in great britain, although there is such a (voluntary) question in northern ireland. enumeration is done on the basis of presence on census night; information is also obtained about the usual place of residence of visitors and about absent usual residents (with a voluntary return if the whole household is absent). there are no long and short forms on the british census. sampling is only undertaken at the processing stage. the answers to those questions which are laborious to code (such as occupation, industry and qualifications) are only fully coded for a 10% sample. when it comes to census output, the principle of dissemisummer 1991 39 nation is best expressed in the motto of the london statistical society, first expressed over 150 years ago: aliis exterendum^. at the time it was first enunciated, it was an unattainable ideal, as the technology for taking censuses and surveys had not progressed very far, and the information was usually published in the form of verbal commentaries, often of a very opinionated form (cullen 1975). however, the trend in census output has continuously been towards making more and more of the detailed information available, aiming for the user to be free to rework it and interpret it in any way that he or she chooses. first this was achieved by providing more tables. there will be 20 volumes of special topic statistics from the 1991 census and monitors for each county and parliamentary constituency. more recently, increased volume of information have been supplied in machine-readable form; the census offices estimate that 98% of census output nowadays is in electronic form. in particular, the small area statistics are particularly detailed; information is provided down to aggregates of around 200 households in england and wales (70 households in scotland) and around 9,000 pieces of information are available at this level. there are also special workplace and migration statistics released in machine-readable format for small areas. thus, up to 1991, the output from the british census in forms other than microdata was exceptionally detailed, especially in the amount of machine-readable data made available for small areas. this may also be part of the explanation why the demand for microdata in the past never became overwhelming. the information to be released at the time of writing, agreement in principle has been given to a plan to release two different samples of data from the 1991 census. their broad structure has been agreed, but negotiations are still continuing about the precise details. any of the details mentioned here could therefore change before release of the data. the first file proposed is a two percent sample of individuals with full housing information and some limited information about household structure attached. the census offices have been guarded about releasing too much information about other household members for fear of effectively releasing what amounts to a hierarchical file under a different guise. this file will show a geographical scheme identifying large local authority districts or groupings of smaller ones. the second file proposed is a one percent hierarchical file of households, containing housing information and full information about all the individual household members. for confidentiality reasons this file will classify data only to the 10 standard regions of great britain. both files are to be drawn from the 10% of records which are fully coded. the sample will be drawn systematically from this 10% file which is ordered by county, then by enumeration district and street. the two samples will be drawn without replacement so that there will be no overlapping subsamples. measures to protect conndentiality of the information great effort is put into ensuring that the data is safe from identification and disclosure. five different devices are being used. (i) sampling of records sampling itself is one of the most effective ways of reducing the risk of disclosure, provided that users cannot identify which individuals are selected. the sampling fractions are small (.01 and .02), and, because the geographical identifiers are different on the two files, it will not be possible to combine them. (ii) supresslon of variables names and addresses are not entered onto the census computer, and therefore obviously not even available for suppression. precise data of birth is to be suppressed; age will only be available in yearly bands, and will be top-coded (see below). there will also be no information on the imputed missing household', since these are excluded from the 10% coding operation. (iii) limiting geographical detail there were competing claims about the basis of sar geography. the main choices were: local government administrative geography • health service administrative geography political boundaries local labour market areas postcode sector defined boundaries grid square geography it was not possible to have more than one system, as the small overlaps between many of these different schemes would have led to unacccptably small areas being identified. on general utilitarian grounds of maximum benefit to maximum numbers, local government administrative geography was chosen. the geographical scheme in the individual level file will identify only those local districts with populations of 40 lassist quarterly 120,000 or more; the rest will be grouped into areas of at least 120,000 population. however, all bar one of the metropolitan local authority districts will be identified. on the hierarchical file, the geography will be even coarser. only standard regions will be identified, with london subdivided into inner and outer areas; this amounts to twelve areas in all, the smallest of which east anglia with an estimated population of 1.9 million. other geographical information is also collected on the census the usual address of visitors to the household, workplace address, students' term-time addresses and address one year ago. these are to be very heavily restricted; at the time of writing, the proposal is to identify standard region, a same district/different district identifier and perhaps a distance measure as well. the order of records in the microdata file are to be scrambled to ensure that locality information cannot be obtained from the proximity of one records to another. (iv) grouping of categories a rule is being used to restrict the fineness of the coding scheme for each variable: each category of each variable to be identified on the sar must have an expected sample count of at least 1 in the smallest geographical area permitted on either file. expectations are formed by scaling down the distribution of that variable for great britain as a whole. to illustrate, in the individual file, the smallest geographical area identified will have a population of 120,000. the sampling fraction will be 1/50. thus the cut-off on this file is 50/120,000 times the population of great britain of 56 miuion, yielding 23,3000; this is then rounded up to 25,000. accordingly, any category of any variable which in great britain is expected to have less than 25,000 people at the census will be grouped in with another category in the sar. in the household file, the smallest region identified is east anglia, with a population of around 1.9 million people, and the sampling fraction will be 1/100, so the cut-off being used for this file is 2,700 people. thus only the univariate distributions are used to identify categories at risk. some have been worried that it is unusual combinations of categories that cause problems: female banisters, small householders in accommodation with a lot of rooms, and so on. however, some recent work suggests that the strategy of worrying only about those categories that are small in the univariate distribution will predict people who have unique combinations of variables extremely efficiently (marsh et al. 1991b). the variables most affected by the grouping rule are occupation (where the 350 categories in the full coding scheme will probably be reduced to around 250), industry (270 categories reduced to around 2(x)), detailed educational qualifications (103 reduced to 53) and country of birth (reduced to 47 groups). the variables to be top-coded are age (to be grouped into two year bands between 91 and 94 and top-coded 95 and above), hours worked last week (top-coded after 70), and number of rooms in the accommodation (to be top-coded after 14). top-coding does not solve the problems of households with very large numbers of individuals; the number of people in a household affects the very structure of a hierarchical file rather than the categories of one variable in it. several suggestions are currently being explored to solve this problem, including removing all geographic identifiers from households with more than 12 members. (v) perbwbation the small area statistics from the british census have always been subject to a degree of random perturbation ('bamardisation') whereby a random -i-l , or -1 is added all cells. the comprehensive application of this technique has always been unpopular with users (marsh et al. 1988) and the cumulative effects of the small errors introduced can pose quite serious analytic problems (senior and cole 1991). wholesale bamardisation is of course not possible with microdata, and the prospect of deliberately adding more noise to many different variables on top of the natural levels of error already existing in the data was resisted by esrc team negotiating the release of the data. instead, the census offices have suggested that a small number of individuals in each area be switched with others in a nearby area (griffin et al. 1989), a technique which amounts to adding a small degree of random noise to the geographical identifier, but preserving the rest of the household data intact. in the usa and canada, the census offices appoint some sort of panel to oversee the arrangements for protecting confidentiality (gales 1988). such a microdata review panel was considered for gb, but, on the recommendation of the royal statistical society, the census offices have appointed one technical adviser instead, to oversee the specification of the sar files and the general arrangements made for confidentiahty. this advisor, a senior academic statistician, will make independent recommendations to the government minister responsible for the census offices. contractual arrangements the funders of the project are the computer board', who are paying for the data, and the economic and social research council, who are putting up the money for a research and distribution centre to be summer 1991 41 established to house the data. data. the contract to buy the data from the census offices is regulated in part by legislation; under the terms of the 1920 census act, the census offices are obliged to recoup the cost of producing extra tables (which includes the sars), but are prohibited from making profits on the data which they supply. the esrc team negotiating the release of sars attempted to find co-purchasers to enter into a consortium to share the costs, but none came forward. thus now, the academic purchasers are bearing the full developmental costs of the sars. the census offices have done their costings on the basis of passing on the full marginal costs of producing sars; at present the cost of the data seems likely to be around £200,000, but this sum will only be finalised when the file specifications have finally been approved. the contract to must safeguard the interests of the academic purchasers in their product. although final contractual details are still being sorted out, we hope that, in return for our payment, we will get full exploitation rights of the dataset, which will be sole rights for a limited period. since the academic purchasers of the data are bearing the full cost of the production of the files, the census offices will not receive any further royalties when value-added products of the census are passed on to third parties. the contract to purchase from the census offices will specify that all end-users of this data must give various undertakings about respecting census confidentiality. people will be expressly forbidden to try to identify individuals in the sar, link them with other sources of data or to claim to have done so. any breach of this agreement would lead to withdrawal of the sars. while not undoing the harm caused by a breach, but its threat would probably act as a deterrent since academics would subsequently be unable to publish any information based on the sars. various methods are under discussion to register users and keep tract of copies. it is possible that heads of departments will be required to be the data holders rather than individual researchers. the contractual arrangements between the purchasers of the data and the end users with respect to have yet to be discussed in detail. academic users wanting the data for research purposes will have free access to the data, but means will be sought to regulate the other ways in which the data may be used, to safeguard the interest of the public fundcrs. a graduated scale of charges seem likely, charging commercial users what the market will bear, perhaps with lesser charges to the rest of the public sector and the voluntary sector. university and college faculty who contract their services as consultants outside the academic sector will be charged for their use of the disseminating the data to users the economic and social research council is funding a centre at a university location to house, disseminate and act as a research focus for this dataset invitations to house this centre were put out to several institutions, and three teams submitted tenders. the decision about which will be successful is expected in late may 1991. four types of usage are envisaged: (i) on-line access over the academic network the joint academic network (janet) connects all universities and many polytechnics and other institutions in britain. it is funded by the computer board, and free at the point of use to all universities. one important method of giving access to the data will be to mount versions of the data at cenffal locations on the network for easy access by any academic users. the decision about which software to use for the data has not yet been made. spssx seems the most popular candidate for the individual file. the hierarchical file of households could also be set up in spss, but this might prove cumbersome, especially since there are two different hierarchies within the dataset: individuals can be grouped into either households of families. sas, sir and oracle are other candidates which might be considered. it is likely that eventually the same information may be held in different formats at different locations; while some may view this as inefficient, from the user's point of view there are great gains in terms of familiarity and ease of use. (ii) tables service the census is a benchmark source of social data. it provides denominators for many different researchers' numerators. however, it is not the prime data source for more than a few researchers. it is therefore important for there to be a service to academics which would provide a service of this kind to other academics, although the demand for this may lead to the need for rationing. (hi) customised subsets another central task that the service centre will undertake is to extract subsets of variables and cases for different users. this will need to have regard to the media most likely to be demanded; a pc/workstation platform is the most likely here, although demand for data on cd rom and other media will need to be monitored carefully. there will doubtless be demand for teaching datasets for use in schools and colleges. (iv) passing on whole dataset for efficiency reasons, other academic users would be lassist quarterly encouraged to use the service provided over the academic network. but the whole philosophy of providing samples of anonymised records is that users will be free to port them into their own hardware and software environment and not be restricted by earlier decisions. no restrictions will therefore be placed on others wanting entire copies of the dataset. public and commercial users do not in general have access to the academic network. they also tend to be familiar with somewhat different software to that used in universities; in britain, for example, the most popular tabulation software used by market researchers is a product marketed by quantime called quanvert. one of the things that the academic purchasers of the data may want to explore is licensing a commercial agency to provide an on-line service for the commercial sector. the centre housing the microdata will have responsibility for documenting the datasets, and for computing the sampling errors and documenting these. it will also need to establish a user group, and disseminate information about the database to users, both in electronic media and by hard copy newsletters. conclusion the existence of microdata from the 1991 census opens up a valuable new resource to british social researchers. our social research community has grown used to making do with small area aggregates, and drawing inferences from these; the existence of microdata should lead to some interesting work on ecological fallacies. survey and market researchers will be in a position to design much more efficient samples to locate specific subgroups of the population. we will have a source of microdata on a badly neglected part of the population, namely those who do not live in private households. research into different means of classifying families and households should blossom given the richness of the hierarchical information available. denham, cj. (1986) 'census microdata in great britain: the possibilities', nutzung von anonymisierten einzekangaben aus daten der amtlichen stalistik: bedingungen und moglichkeiten, verlag w. kohlhammer. gallup (1991) gallup political & economic index, report no. 368, april. gates, g.w. (1988) census bureau microdata: providing useful research data while protecting the anonymity of respondents, proceedings of the social statistics section of the american statistical association annual meeting, new orleans, louisiana, august 22-25. goyder, j., and leiper, j. mck. (1985) 'the decline in survey response: a social values interpretation'. sociology, vol. 19, 1, pp. 55-71. griffin, r.a., navarro, a., and florez-baez, l. (1989) 'disclosure avoidance for the 19(x) census', u.s. bureau of the census, paper prepared for presentation at the 1989 joint statistical meetings, washington, d.c., august her majesty's government (1981a) 1981 census of population: confidentiality and computing, presented to parliament by the secretary of state for health and the secretary of state for scotland, cmmd 8201, london: hmso. her majesty's government (1981b) 'initial study of the office of population censuses and surveys', annex to the review of the government statistical services, cmmd 8236, london: hmso. her majesty's government (1991) 1991 census of population: confidentiality and computing, presented to parliament by the secretary of state for health and the secretary of state for scotland, february, cmmd 1447, london: hmso. perhaps also we may hope that effort will be made to construct internationally comparable files of census microdata relating to the 1990 and 1991 censuses in different countries. one spin-off of presenting this information to an international audience at this lassist meeting might be to stimulate discussion in this direction. references campbell, d. (1987) 'the databank dossier'. new statesman, april 24. cuuen, m. (1975) the early victorian statistical societies, brighton: wheatsheaf. marsh, c, arber, s., wrigley, n., rhind, d., and bulmer, m. (1988) "the view of academic social scientists on the 1991 uk census of population: a report of the economic and social research council working group", environment and planning a, vol. 20:851-889. marsh, c, skinner, c, arber, s., penhale, b., openshaw, s., hobcraft, j., lievesley, d., and walford, n. (1991a) 'the case for samples of anonymised records from the 1991 census'. journal of the royal statistical society (a), vol. 154(2), 1991. marsh, c, dale, a., and skinner, c. (1991b) "safe data or safe settings: disseminating information from the british census of population", paper to be presented to summer 1991 43 the isi meetings in cairo, september. norris, p. (1983) 'microdata from the british census', in david rhind ed. a census user's handbook, london and new york: methuen. redfem, p. (1976) releasing statistics as aggregates (tables) or tapes of anonymised individual data (id). opcs, london: mimeo, available from the author at 17 fulwith close, harrogate, hg2 8hp. senior, m., and cole, k. (1991) '1981 census costs department of health £190,(xx)!', manchester computing centre newsletter, no. 183, p. 2. ' presented at the 1assist 91 conference held in edmonton, alberta, canada. may 14 17, 1991. ^the federal republic of germany has a policy against the general release of microdata collected by the state. it will not allow its labour force survey to be released for secondary analysis, for example. since the statistical office of the european commission takes the most restrictive poucies of its member states as guidelines on release of microdata, this means that europe-wide release of the labour force survey is precluded. ' in fact, the value of census data to academic researchers does not seem to decline as the census data becomes less timely; if usage of small area statistics is a guide, the usage of these increased throughout the 1980s. * perhaps this is the reason why two other data sources collected compulsorily by law (the new earnings survey and census of employment) are only released in aggregates, albeit of very small units. ' (trans.) for others to thresh, or, more colloquially, let someone else work out what it all means! it is still the motto of the royal statistical society. ' households with all members absent on census night and which fail to return a voluntary return to the census offices once they get back home. ' the body which funds purchases and support of mainframe computing in british universities. it is soon to become a subcommittee of the main universities funding council. assist quarterly vol26no1 10 iassist quarterly spring 2002 iassist quarterly spring 2002 11 counting california background government data serve a variety of clientele, ranging from businesses and private citizens to some of the most prominent educational and research institutions in the world. with a population and economy larger than many european countries, the state of californiaʼs need for continuous and uniform access to government-produced data has never been more urgent. while digital technologies have revolutionized data distribution, they have also created new problems. what was once a stable system of print materials has given way to a diffuse, constantly changing array of electronic media, using different formats and access methods. the current climate leaves many would-be users frustrated or bewildered; each new upgrade of software and web browsers only exacerbates the problem. the preservation and consolidation of historical, or timeseries, data are similarly at risk. government agency web sites often mount new information, but may follow no systematic plan to preserve older, historical data as each update supersedes the last. the lack of cumulative timeseries data can effectively cripple any attempt to discern long-term trends and changes. the need for uninterrupted data access and the preservation of historical data are two of the biggest problems facing government data distribution today. counting california (http://countingcalifornia.cdlib.org) is a new initiative committed to enhancing california citizens ̓access to the growing range of social science and economic data produced by government agencies by enabling users access to data compiled by federal, state, and local agencies through a single interface. counting californiaʼs goals respond directly to these challenges: • to provide flexible, user-friendly access that meets the diverse needs of the california citizenry; • to insure uniform, continuous access to both current and historical government data; and by patricia cruse, ilona einowski, and juri stratford* applications in the real world: the counting california experience with the ddi • to foster the ability to share data and work collaboratively between government agencies and other members of the data community. development of counting california began in june 2000. the project team for counting california includes patricia cruse (program leader), ilona einowski, fred gey, andjuri stratford (data specialists), brian tingle (www manager), margaret low (sybase programmer), and marsha fanshier (sas programmer). the project teamʼs programmers initially had to rely on a variety of sources developed by the data specialists to describe the data. however, the project is now using xml exclusively to develop the data descriptions and interfaces to counting californiaʼs data. one of the project teamʼs objectives was to have the metadata drive all the functionality from data discovery to data display: including the creation of citations and related materials, labeling, and the thesaurus, as well as enabling searching. the data documentation initiative (ddi) appeared to be the appropriate set of guidelines to develop the xml and achieve this goal.¡ the data documentation initiative for the past several years the social science data community has been working on the ddi to produce a metadata standard for social science data resources. the ddi “is an effort to establish an international criterion and methodology for the content, presentation, transportation, and preservation of ʻmetadata ̓about the datasets in the social and behavioral sciences.” the ddi committee has produced a document type definition (dtd) as the foundation of its data discovery and display system. the dtd provides the rules for applying xml to the ddi metadata. xml, a dialect of the more general sgml markup language, is used for documents containing structured information. in recent years there has been explosive growth in the use of xml and web-based xml tools including databases, search engines, and editors. as we build additional functionality into counting california, we will support other xml tools. in keeping with the projectʼs goals, we intend to make the ddi metadata available to the scholarly community. links to actual xml will be available from the web site. we will encourage 10 iassist quarterly spring 2002 iassist quarterly spring 2002 11 others to build upon our work in creating alternate applications that can also be shared. the ddi is currently a set of guidelines. throughout the development of counting california, the project team has relied heavily on the expertise of the data community. ddi committee members ann green (yale university), wendy thomas (university of minnesota), and cavan capps (bureau of the census) provided guidance in the early stages of the project by answering our many questions about the dtd. many of our early questions focused on the problem of representing tables of aggregate data using the ddi. these questions lead to a meeting in september 2000 with these ddi committee members. as a result of this meeting, we incorporated wendy thomas ̓proposed extensions to the ddi, which allows for the description of aggregate/tabular data, into the development of the metadata. at present, there are few concrete examples demonstrating how the ddi is to be applied. mary vardigan (inter-university consortium for political and social research) furnished reference support as we struggled to collect examples. project implementation to set the stage, let us begin by giving a brief overview of the contents of the elements in the ddi dtd. the data documentation initiative dtd consists of 5 sections: • section 1.0: document description (codebook header) consisting of bibliographic information describing the ddi-compliant document being created. • section 2.0: study description consisting of information about the data collection, study, or compilation that the ddi-compliant documentation file describes. • section 3.0: data files description consisting of information about the particular data file(s) containing numeric and/or numeric plus textual information that the ddi-compliant file describes. • section 4.0: variable description consisting of information about each of the individual variables/ observations that the ddi-compliant documentation file describes. • section 5.0: other study-related material provides a section for the inclusion of other materials that are related to the study as identified and labeled by the dtd users (encoders). other study-related materials may include: questionnaires, coding notes, spss/sas/ stata setups (and others), user manuals, continuity guides, sample computer software programs, glossaries of terms, interviewer/project instructions, maps, database schema, data dictionaries, show cards, coding information, interview schedules, missing values information, frequency files, variable maps, etc. there are a total of 178 elements in the ddi including wendy thomas ̓extensions, which describe aggregate/ tabular data tables. in order to facilitate the creation of the xml, we carefully examined the full list of elements, and selected only those elements that directly impacted the functionality of the system. the project developed its own internal standards incorporating a subset of the dtd needed to access data and to present displays. as the project evolved, we saw the need for additional elements that we then incorporated. since we were working with a manageable number of data titles, and the entire system was designed as a set of interacting modules, it was not difficult to incorporate these additional elements. we are currently using only thirty-five of the original 178 elements. the project used the xmetal software to both develop the xml and to validate the integrity of the xml̓ s adherence to the dtd. as we mentioned before, it is the intent of the project to share the xml with data community. the project is using the ddi to document aggregate data sets at the variable level. this will allow users to go directly from the variable descriptions to all related tables in any data file. the backbone of counting california s̓ data delivery capacities is sas/intranet, which allows for the integration of sas and the world wide web. specifically, sas allows for data extraction via the metadata database and provides the ability to format the data in tables, charts, maps, and graphs. in the future we hope to take full advantage of sas/intranetʼs capabilities as we add functionality to the system. some of the files included in counting california were available from the producers only as excel spreadsheets. for these files, we used ddi sections 1 and 2 to drive the data discovery. the data specialists created these sections, and the sybase programmer used the metadata to retrieve the corresponding spreadsheet. we hope to offer pdf versions of the spreadsheets as a future enhancement. we also hope to receive the actual data files, used to create the spreadsheets, directly from the producers thereby allowing the project more flexibility in presenting and combining data. one of the many advantages of basing the system on a fully loaded ddi is that we eliminate many problems of inconsistency. by conforming to the ddi, all titles, bibliographic citations, concepts, etc. will always appear in the same format. as it currently exists, counting california is not fully populated with ddi-generated material. some of the “early” screens were generated by hand. this was done to expedite the creation of certain screens when we were still experimenting with the xml. given all the other tasks on our plates, we have not gone back and reworked those screens. itʼs difficult to bring yourself to dismantle something that is currently functioning adequately, but it will be done. 12 iassist quarterly spring 2002 iassist quarterly spring 2002 13 project functionality counting california integrates data from a variety of sources, and provides a standard display regardless of the original format of data. the system provides a number of entry points into the data including subject, geography, study title, or issuing agency. counting california allows users to search the table titles, studies, and geographic areas as free text. users can search all datasets and tables simultaneously or search tables within an individual dataset. the system is designed to ensure that there are no “data dead-ends”: the system never returns zero results. the metadata drives functionality for counting california including the development of the thesaurus, the creation of citations and references to related materials, searching and the labeling of tables and variables. we have found that using the ddi driven version allows a modular approach, is less labor intensive when incorporating changes, and encourages sharing. at present this data documentation process is labor intensive. the data documentation process must be simplified and to some extent automated to encourage california state agencies to contribute new data sets to counting california. at present, counting california has only a limited set of features. more features will be added as we learn more about what users want, which will allow us to prioritize development activities. however, our experience has provided a real world application of the ddi standards. our team members now feel qualified to engage in a productive exchange with the ddi committee on future enhancements of the ddi, including additional elements and expanded documentation. while we are pleased with the performance of the ddi in providing metadata to power counting california, we are aware that the standards for the content of the elements are not fully developed. we found that we could populate the metadata database with ddi elements that would provide the functionality we desired as long as we were consistent in the application of the ddi within the project. the continued development of both the counting california project and the ddi is an iterative process. acknowledgements: counting california is a collaborative project funded by the california digital library and the library of california. a portion of the funding for phase i of counting california comes via special arrangement with the library of california. this interagency agreement stipulates that the library of california provides funding, and the cdl implements the project. some of the funding for phase ii comes through a grant from the library services and technology act (lsta), administered by the california state library. current strategic partners in data acquisition include the u.s. bureau of the census and the california department of finance. we are currently negotiating similar arrangements with numerous other state agencies. * paper presented at the iassist conference, amsterdam, may 2001. patricia cruse, california digital library, academic initiatives, university of california, patricia.cruse@ucop.edu ilona m. einowsky, assistant director, uc data/src, university of california, berkeley, ilona@ucdata.berkeley.edu; juri stratford, government information and maps, shields library, university of california, davis, jtstratford@ucdavis.edu. footnotes * note: activities described in this article do not reflect the current status of counting california, but rather the status of the project in may 2001. mailto:patricia.cruse@ucop.edu book reviews kathleen he i m university of illinois at urbana-ciiampagne graduate school of library and information science boruch, robert f.; wortman^ paul m.j cordray, david s.j and associates reanalyzing program evaluations: po licies and practices for secondary analysis of social and educati onal programs. (san francisco: jossey-bass publishers, igst). i+orpp. isbn 0875 89-^+9 5-x , aimed primarily at researchers who must capitalize on existing data by applying their findings in the interest of advancing science and policy, this volume provides nuch material for social science information specialists seeking to gain greater understanding of the publics they serve. the book is divided into four parts plus an introduction. part one represents policies of the federal government and selected key agencies (national archives; national institute of justice; government accounting office) and methods for locating public and private data resources; part two is comprised of a discussion of the need for better documentation of data files and technical guidelines for preparing and documenting them; part three contains essays suggesting nethods to improve reanalyses of evaluations; and part four includes case studies. perhaps the most useful chapters for data specialists are those on documentation. joan linsenmeier, paul m. wortman and michael hendricks outline their problems with the riverside school study of desegregation data set. their perceptions support the leed for generally accepted standards for data documentation. alice robb i n proposes a set of guidelines for the documentation of machinereadable files which reflect the concerns of a variety of users. these extensive guidelines (59 pages) are divided into four sections: 1) "technical standards for documentat i on ,"--wh i ch addresses five major components (general study overview, listory of the project, summary of the data file's processing history, codebook , appendices); 2) "technical standards for quality data"--which recommends good practice for organizing a data file, identifying record types and data items, describing nissing data, employing standardized coding schemes, maintaining data at their lowest ievel of aggregation, and recording standards for the transfer of data files from one computer to another; 3) "providing access to information about machine-readable 3ata fi les ,"--wh i ch suggests standard rules for identifying files for inclusion in library, scientific, and technical information systems; and h) "roles of the researcher, ijponsor, and government agencies." robbins' "guidelines" are organized into clear sections and subsections and are correlated with figures to illustrate how each should be developed. a subject index provides cross-referenced access. these "guidelines," if applied by developers of nachinereadable data files, should begin to move files from the fringes of respecta>i1ity close to the mainstream of information resources. this is an extremely important book for the data professional. it provides theoretical, intellectual, and technical background for data provision. the thl rtyr the magazine public opinion; data set news, announcements of data sets of special iterest to the research community. remarks: since 1977, the roper center operates irough a formal partnership of the university of connecticut, yale university, and ' lliams college. the university of connecticut branch of the roper centei — which !s primary operating respons i bi 1 i t ies-i s housed administratively within the stitute for social inquiry. st of general abbreviations, geographic abbreviations, and network acronyms, are ovided for easy look-ups. appendixes include lists of libraries for the blind and physically handicapped; tent depository libraries; federal information centers; federal job information nters; and united nations depository libraries. a detailed subject index with good ossreferenci ng provides access by topic. the scope of this directory is enormous but in spite of many entries of only ripheral interest to the social science data community, there is much pertinent and -to-date information that indicates that this volume should be examined and considered >r purchase. while no subject access is provided under "data archive" or data library," ladings under "survey research," "public opinion," and "social science" do include ijor data libraries. access under specific subjects such as "population," "political :iience," and "census" also lead to data holdings. there are omissions of some major 'tta libraries (notably icpsr, which is referred to only under the university of l:chiganinst i tute for social research, without reference to holdings) and federal lies such as the criminal justice date held by icpsr or the bls data bank files are itt identifiable through this pi rectory . if the lassist membership is interested in litter coverage of data libraries, comments should be forwarded to gale research i mpany , book tower, detroit, michigan ^48226. in addition to this major reference tool which is volume 1 of a set. gale produces :o additional allied volumes: volume 2, geographic and personnel indexes , lists all rtries according to state or province and an alphabetical list of all personnel ilkpp. isbn 0-8103-o2620'4 for $200); and new special libraries (follows the same rrmat as the di rectory but provides up-to-date information between major editions .(d includes cumulative indexing. isbn 8102-0281-0 for $210) is volume three. vol29-3.indd 8 iassist quarterly fall 2005 by louise corti 1 qualitative archiving and data sharing: extending the reach and impact of qualitative data introduction archived qualitative data are a rich and unique, yet too often unexploited, source of research material. they offer information that can be reanalysed, reworked, and compared with contemporary data. in time, too, archived research materials can prove to be a significant part of our cultural heritage and become resources for historical as well as contemporary research. but while there is a well-established tradition in social science of reanalysing quantitative data, there is not yet a well developed paradigm, nor a pervasive research culture of sharing or secondary analysis of qualitative data. the lack of discussion in the current literature on the benefits and limitations of such approaches is evident. in the uk, finland and france some interesting debate has begun, but it is still early days. further internationalisation of some of the arguments would certainly be welcomed, and i would encourage iassist members who are considering archiving qualitative data to consider kick starting some written discussion in their own countries1. this contribution provides an overview of some of the perceived barriers to re-use and highlights some of the positive pragmatic measures that are being taken to enable both sharing and re-use of qualitative data. i draw on the recent experiences of esds qualidata and a new research council funded programme intended to investigate innovative ways of extending the reach and impact of qualitative data. a brief history readers will be well aware that research data archiving has been around for some years. the data archiving movement began in the 1960s within a number of key social science departments in the united states who stored original data of survey interviews. the movement spread across europe and in 1967 a uk data archive (ukda) was established by the uk social science research council (ssrc). but the emphasis was strictly quantitative. the ssrc’s successor, the esrc introduced a formalised datasets policy in 1996 that contracted all award holders to submit data for possible accession to the uk data archive. the word ‘data’; was typically, and perhaps conveniently, taken by researchers to refer purely to numeric data. pressure from a small insistent minority fought for qualitative data to be explicitly embedded into the esrc portfolio of data resources. the qualidata centre set up in 1994 in the sociology department at essex complemented the uk data archive with a joint mission to actively acquire, curate, disseminate and promote the raw data from social science research. from 2003, qualidata became an integral part of a larger joined up one-stop-shop for data sharing, archiving and dissemination, under the esrc/jisc supported economic and social data service (esds). over the past ten years esds qualidata has contributed to the elucidation of some of the key perceived barriers to re-using data through extensive contact with 2000 or more qualitative researchers and through the experiences of handing many disparate data collections. these arguments have been rehearsed in a number of publications by qualidata staff (e.g. corti and thompson, 2004; corti, l., witzel, a. and bishop, l. 2005; bishop 2005; corti 2000). while esds qualidata has conquered some of the ‘mainstream’ methods for archiving and sharing qualitative data, it has made significant efforts to spark more general academic debate. however, there is still a significant under-use of archived qualitative data when compared with survey data. moreover, there is still a noticeable imbalance in attitudes towards sharing and re-using data across disciplines and types of methodological approaches. the key research issues facing re-use of qualitative data there are some insistent voices who suggest there is a widespread reluctance to deposit qualitative data with a research archive. while this was partially true some ten years ago, today we see a new generation of qualitative researchers who are more inclined to either embrace or gracefully accept the esrc’s datasets policy and its efforts to promote the value of sharing data (esrc 2005). at the uk data archive, where some 150 qualitative datasets are catalogued, user figures have soared, particularly for use in research methods teaching. nevertheless, there are still barriers. the six key main perceived barriers that have been identified through contact iassist quarterly fall 2005 9 with researchers in the uk over past ten years can be summarised as follows: • the practice of secondary analysis of qualitative data is not yet a common place research activity. the literature is not forthcoming on methodological guidance on how to approach the revisiting of data. corti and thompson (2004) provide the first inclusive and state-of the-art chapter on the topic invited for a high profile methods reader by seale et al. progress is also hindered by preconceptions and sometimes less than innovative approaches to qualitative research. a cultural shift is required and we believe that this has been progressively happening since 1994. • problems of the implicit nature of qualitative data collection and analysis, of context and reflexivity, which are sometimes proclaimed to be indefinable. what are needed here then are practical strategies. indeed, for research conducted in teams, data and fieldwork experiences are commonly shared, and for principal investigators who remain one step away from the field, it is imperative that they rely on their research staff on the ground to capture, document and communicate the nuances of the research process. it is vital to capture better and more systematically the context and the interrelationships among data and between data and other academic products, like analyses and write ups. • lack of time to get fully acquainted with research materials created by someone else. social historians have been more forthcoming in revisiting data sources because of their willingness to embrace the slow and rigorous but commonly accepted practice of document analysis and the need to evaluate methodically the very sources they are revisiting. however it can be terribly time-consuming to locate suitable data sources, and to locate, for example, paper materials that may reside in traditional archival locations with limited access. new ways and tools that more efficiently expose the content and context of digital data sources need to be developed, in order to reduce such researcher burden. • constraints of informed consent. informed consent is an ethical and legal requirement of the research process. it must be thought through at the time of research proposal planning and writing and be tailored towards the specific research questions and the sample. often consent is not addressed until late in the research process by many researchers, and verbal consent alone is typically no sufficient for longer-term sharing and for effective use of research findings by the original researcher. failure to realise the need to gain informed consent means that research efforts and the opportunities for archiving and secondary analysis are jeopardised from the start. but researchers require more guidance on this area to better understand the nature and implications of consent and confidentiality. additionally, ppragmatic strategies are also required to aid the commonly accepted practice of anonymisation or pseudonymisation. bishop provides a succinct reply to some of the recent scepticism of the possibilities of re-use (bishop 2005) • insecurity about exposure of one’s research practice, ipr or threat of misinterpretation. this may be relevant in some specific cases (e.g. an anthropologist's life work), but for the sake of data quality or auditing as hammersley describes it (hammersley 1997), exposure of data and methods is no bad thing. capturing evidence, or specific reasons, as to why data cannot be shared is valuable. • finally, lack of a wide range of publicly available catalogued research data. while in the uk, the economic and social data service (esds) has done much to facilitate common resource discovery points of access through the use of standards at the study description level, ways to delve deeper into the qualitative data resource have not been as forthcoming as they have for survey data. the nesstar system is a good example of how data can be browsed online through the use of detailed data description down to the survey question (variable) level (nesstar 2005). as the pool of rich and diverse shareable data expands, the greater the need also for interoperable and standardised description that will allow searching and location of key data across distributed sources. the means of enabling this information stock and flow to reach fruition needs to be investigated and common community methods agreed. prerequisites for making data shareable there are two major issues that appear to be at the heart of making data fully shareable. the first is producing rich and full documentation about the data and the research processes used to conceptualise, collect, manage, process and analyse data. full documentation enables effective resource discovery (i.e., catalogues) of distributed data sources and enables more informed re-use. the second challenge for sharing data is that of exposing data in the most flexible way possible so as to enable multiple methods of accessibility and innovative uses, for example, combine and link: activities that are the very core of some of the initial considerations of e-social scientists. both challenges require that: • data are collected to a high standard using appropriate sampling strategies, rigorous data gathering methods and, where appropriate, systematic interview transcription 10 iassist quarterly fall 2005 • research methods and practices (including the consent process) are fully documented • the context of the data collection and analysis is captured • the richness of the structure and features of data and are made available (use of mark-up) • the interrelationships between data and analyses (intra-project) are made available (issues of representation) • data are disseminated in sensitive ways that satisfy the ethical and legal requirements to which they are bound. • data are represented in appealing and digestible ways, such presenting academic findings alongside evidence from the raw data (that is more than anecdotal quotes) enabling these requirements entails practical as well as conceptual challenges. and fundamentally, the underlying need is for the formulation, adoption and community subscription to commonly agreed methods, standards, and ontologies for data description and exposure. previous work in this area has been spearheaded by esds qualidata and the uk data archive in social science archival documentation and data processing for qualitative data (corti, 2002). creatively exploring the barriers and looking forward esds has found that data creation workshops have been very useful in helping unpack some of the specific issues and problems arising in the course of projects that are considering data sharing. the medical research council (mrc) has also taken up an interest in data sharing of primary data from population and clinical trials data, with qualitative data firmly on the agenda. (corti and wright 2003). more recently the natural environment research council (nerc) has also implemented a formalised data management policy for a joint council programme on rural economy and land use (relu) that specifically includes all the qualitative data from the programme. context and consent are the two words that crop up most frequently in the debates. for researchers, like myself, who have been seeking academic funding opportunities to confront such consent, context and technical issues, 2003 – 2004 saw a bumper harvest of such prospects. in the past it has been unusual for esrc to fund research and development or consider dedicated methodological initiatives. but, over the past three years we have seen a much welcomed move towards dedicated funding for methods. these strands of money have enabled some innovative investigations to be undertaken, particularly for qualitative data. five main pots of esrc funding appeared on the scene, thanks to a number of champions to the cause of methods and data analysis: e-social science; the research methods programme; the national centre for social research; the qualitative longitudinal study and the quads scheme. innovation: the quads scheme quads is the esrc qualitative archiving and data sharing scheme, running from april 2005 until october 2006. the aim of the scheme is to develop and promote innovative methodological approaches to the archiving, sharing, re-use and secondary analysis of qualitative research and data. a range of new models for increasing access to qualitative data resources, and for extending the reach and impact of qualitative studies will be explored. the scheme also aims to disseminate good practice in qualitative data sharing and research archiving. this is part of the esrc’s initiative to increase the uk resource of highly skilled researchers, and to fully exploit the distinctive potential offered by qualitative research and data. the quads is a small initiative (some £500,000 over 18 months) but is dedicated to the mission of learning more about sharing, representation and re-use of qualitative data, in all of its disparate shape and forms. five small exploratory projects have been funded together with a coordination role. the co-ordination team based at esds qualidata have been charged with the task of providing a pivotal role in fostering communication and understanding between the five demonstrator projects. communication of the scheme’s innovative efforts to the broader spectrum of qualitative researchers is much needed. but equally it is must be appreciated that there exist various communities of practice with different data needs and methodological approaches to sharing and secondary analysis of qualitative research and data. fruitful collaboration is required which ca be achieved through guided discourse to inform and help guide the progress of quads demonstrators, and to encourage the broader acceptance and take up of data sharing and re-use. iassist quarterly fall 2005 11 key areas for quads projects four key areas of needs and commonality identified across all the quads projects point to: defining and capturing data context, audio-visual archiving; consent, confidentiality and ipr; and web and metadata standards. the debate on capturing context has been around for some time now on the qualitative data archiving scene. quads aims to devise and recommend a minimum set of contextual constructs that would be necessary to document a collection of qualitative data to enable informed secondary use. regarding audio-visual data, they are being handled by many of the projects and the scheme is providing an opportunity to share expertise on presenting and re-using such sources. on the hot topic of consent, confidentiality and copyright, while esds qualidata maintain up-to-date detailed information many the quads projects do have specific consent and copyright issues, and it will be invaluable to see how these are confronted by the different projects during the demonstrator period. they will afford unique case studies that can be used in the future. quads coordination will hold an end of scheme handson demonstrator workshop, where projects will be able to talk about their investigations and developments and demonstrate any working quads products to an open invitation audience. furthermore, quads coordination through its web site will mount papers, tools and training materials arising during the course of the projects in an easy-to-navigate manner. a session at the ncrm summer school is currently being planned. quads is an exciting pilot initiative that is exploring experimental methods. the team is therefore very happy to hear from anyone who is already working in this area or who would lie to contribute to these exciting projects. but what about these standards? in order to approach primary data now and in the future in years, we need that data to be accurately, richly and contextually described. and in turn, re-presentation of original data, methods and analytic interpretation and their interweaving requires agreed and exemplary standards and procedures. fielding’s scoping study that examined issues for the role of qualitative data in e-social science (fielding 2003) aptly confirmed that that ‘it is timely to anticipate emerging innovations in qualitative methods, including new data forms, sources, possibilities for research archiving and data mining and the potential for increased participation and access’. representation can be viewed across a spectrum starting from the simple publishing of anonymised digital qualitative data sources or banks (which are typically not present) through to the ability to link qualitative data to other distributed data sources (e.g. audio-visual or geo12 iassist quarterly fall 2005 coded data sources) and to creative and exciting ways of visualizing data. however, it is important to take a step back and see what is currently exposed. the researcher will find very little qualitative data even exposed to the web in any meaningful way. while there are archives of qualitative data to be found across the world, the majority is not even in digital format and “digitizing” these collections is often seen as merely providing an online catalogue of digitised metadata. the issue of how to make these data resources accessible to users has hitherto been a central concern for esds qualidata who has continually been seeking ways to meet users’ requirements. standards of relevance are those for: building sustainable web sites; harmonious data descriptions to enable rich resource discovery (metadata); and marking-up data content. esds qualidata recognised the need for standards and tools back in 2000 – tools that allow data to be published to the web and support online interrogation of data via standard web browsers. in 2000, qualidata undertook pioneering work in this area through developing the qualidata online system and a methodology for sharing data (corti and barker, 2002, 2003). the need to keep pace with the development numeric data browsing systems, that are now quite far advanced, is important not only for the uk data archive but also for other groups who wish to publish and share qualitative data. community efforts quads co-ordination is very aware that there are an increasing number of projects in the world that are looking at sharing qualitative data, typically via the web, and particularly in the wake of the e-science rush. but they are not linked up in any formal way. quads coordination will be building an interactive map will be built to show the location of such initiatives and the key contacts. additionally training and advice is being given on best practice in metadata creation and web standards for qualitative data. it is hoped that many of the e-science and methods research groups will be amenable to agreeing on some basic sets of standards. *for more details about esds qualidata see www.esds. ac.uk/qualidata and quads see: http://quads.esds.ac.uk/ or contact louise corti, uk data archive, university of essex, colchester co4 3sq corti@essex.ac.uk. an earlier version of this article appeared in issue 1 of qualiti, the newsletter of the national research methods centre, university of cardiff. www.cardiff.ac.uk/socsi/ qualiti/newsletter.html references bishop, l. (2005). ‘protecting respondents and enabling data sharing: reply to parry and mauthner’, sociology, 39(2) pp. 333-336, london: sage publications corti, l. (2000, december). progress and problems of preserving and providing access to qualitative data for social research the international picture of an emerging culture. forum qualitative sozialforschung/forum: qualitative social research [online journal], 1(3). http:// qualitative-research.net/fqs-texte/3-00/3-00corti-e.htm corti, l. and thompson, p. (2004) ‘secondary analysis of archive data’ in c. seale et al (eds.), qualitative research practice, london: sage publications corti, l., witzel, a. and bishop, l. (2005, january). secondary analysis of qualitative data forum qualitative sozialforschung/forum: qualitative social research [online journal], 6(1). http://qualitative-research.net/fqs/ fqs-e/inhalt1-05-e.htm corti, l. and barker, e. (2002) ‘edwardians online’ iassist quarterly vol. 26 (2002)no. 4, pp. 5-8 corti, l. and barker, e. (2003) ‘edwardians online: an xml application for qualitative data’ assignation, vol 120, no 2, january 2003. corti, louise (2002). qualilitative data processing guidelines. qualidata, uk data archive, university of essex, colchester corti, l. (2000, december). progress and problems of preserving and providing access to qualitative data for social research the international picture of an emerging culture. forum qualitative sozialforschung/forum: qualitative social research [online journal], 1(3). http:// qualitative-research.net/fqs-texte/3-00/3-00corti-e.htm esds www.esds.ac.uk esds qualidata www.esds.ac.uk/qualidata esds qualidata online http://www.esds.ac.uk/qualidata/ online/ esrc datasets policy www.esrcsocietytoday.ac.uk/ esrcinfocentre/images/annex%20c_tcm6-9738.pdf) fielding, n.(2003), qualitative research and e-social science: appraising the potential, university of surrey, http://www.ncess.ac.uk/docs/qualitative_research_and_e_ soc_sci.pdf hammersley, m. (1997). qualitative data archiving: some reflections on its prospects and problems. sociology, 31(1), 131-42 http://qualitative-research.net/fqs-texte/3-00/3-00corti-e.htm http://qualitative-research.net/fqs-texte/3-00/3-00corti-e.htm http://qualitative-research.net/fqs-texte/3-00/3-00corti-e.htm http://qualitative-research.net/fqs-texte/3-00/3-00corti-e.htm http://www.aslib.co.uk/sigs/assig/assignation/2002.html http://qualitative-research.net/fqs-texte/3-00/3-00corti-e.htm http://qualitative-research.net/fqs-texte/3-00/3-00corti-e.htm http://qualitative-research.net/fqs-texte/3-00/3-00corti-e.htm http://qualitative-research.net/fqs-texte/3-00/3-00corti-e.htm iassist quarterly fall 2005 13 ncess hub and nodes www.ncess.ac.uk ncrm hub and nodes www.ncrm.ac.uk nesstar (2005) http://nesstar.esds.ac.uk/webview quads http://quads.esds.ac.uk footnotes 1 see www.esds.ac.uk/qualidata/access/internationaldata. asp for an overview of progress on national data archives acquiring qualitative data http://quads.esds.ac.uk sist newsletter vol.1, no. 4 '^select if and mult response techniques for filtered marginals. resource *select if mult response number of jobs 2 1 kbyt-sec 21,016 24,957 elap-kbs 180,493 73,752 reader excp** 93 31 printer excp 14,630 3,844 public disk excp 10,476 360 private disk excp 10 5 tape excp 72 36 tape mounts 2 1 cost $31.70 $13.39 elapsed time 13.48 min. 6.23 min. *comparisons were performed on the university of connecticut's research computer center ibm 360/65 ibm 370/155 systems running under os and shared spool hasp. local costs may vary from installation to installation. **each excp represents the "execution of a channel program" and indicates the movement of one block of data. the mult response procedure can further be extended to filtered bivariate tables by including a second by_ statement on the tables card. this technique is not illustrated due to space limitations. readers are referred to their local installations for further documentation of spss version 7.0. discussion paper/riohard c. roistacher the following article describes the work of richard c. roistacher and barbara noble at the center for advanced computation at the university of illinois. they are involved in the development of guidelines for the descriptive materials which accompany a data file. a source documentation style manual by richard roistacher center for advanced computation university of illinois urbana, illinois barbara noble and richard roistacher of the university of illinois' center for advanced computation are currently developing a style manual for the documentation of machine readable data. the manual, which is being developed as part of a project funded by the u.s. department of justice's law enforcement assistance administration, is presently available in draft form. the manual, conforming to the sist newsletter vol.1, no. 4 naming conventions of the u.s. government's federal information processing standards, is titled leaa research support center machine readable source documentation system: user's guide . the style manual is designed to serve the needs of data producers, archivists, and users. it was developed from several current series of documentation, as well as from the leaa project's experience in archiving a number of large criminal justice data files. the manual gives an annotated example of each major section of the source documentation for a machine readable data file. the manual's major sections describe the cover, abstract, introduction, codebook, and appendices of source documentation. the section on the cover gives advice on the titling, authorship, and citation of files and documentation. the manual's chapters on the abstract and introduction sections gives a form for abstracts, and gives advice on how to describe the file's collection, methodological, and processing history. the chapter on the codebook section gives formats for data item blocks for several types of data. the chapter on appendices gives formats for listings of known errors, definitions of terms, dictionary descriptions, description of data shipments, and extended code listings. at present, the examples in the manual cover only rectangular files of survey data. additions to the manual will cover the documentation of hierarchical and other complex files, time series, network data, and generalized arrays. the documentation manual is designed to be independent of any particular manual or computer based system for producing documentation. however, the leaa project at the center for advanced computation has developed a computer system for producing machine readable source documentation conforming to the manual. the system merges the information in an osiris iii data dictionary file into a file of documentation text. the documentation text is processed with the university of british columbia's fmt document processing program to yield a finished codebook, which can be printed or copied to tape. the system will automatically reformat and reorder the documentation to match a data file which has been subsetted, reordered, or merged. new page numbers, tables of contents, cross references, and indices are generated by the format program without any intervention by the user. a later version of the system, which runs on ibm hardware, will extract format information from spss system files. this winter. noble and roistacher will hold a conference on the documentation of machine readable data. one of the major purposes of the conference will be to expand and refine the current draft of the documentation style manual. the authors would welcome advice and comments from data producers, archivists, and users. copies of the current draft are available to interested people, and any comments would be welcome. copies are available from barbara b. noble or richard c. roistacher, center for advanced computation, university of illinois, urbana, il 61801. [editor's note: the meeting was held on october 27 and 28, 1977 in boston. revisions to the manual were recommended. a full report will appear in the next issues of lassist as roistacher and noble further refine the manual .] 33 i^$sist newsletter vol.1, no. 4 data organization registry form june1977 ?s s ess 17 34 i^$sist newsletter vol.1, no. 4 35 1)^$sist newsletter vol.1 no. 4 .go. li^^.^-, ^sj-s^ 36 tf^3 sara and faadifig ef nriagialins tapeis patricia a. reslook center for naval analyses alexandria, virginia there are many obstacles to successfully acquiring machine-readable data. difficulties usually occur during the physical transfer of the data from one site to another, the most common method being magnetic tape. yet, even with higher quality tape and more efficient tape drives, tape processing is still plagued with problems. parity errors, incompatible tape labels, multiple encoding schemes, and poor quality media have all contributed to tape users' frustration. in an attempt to educate the nontechnical professional, this paper will give an overview of the technical aspects of magnetic tapes and their use. topics include the physical aspects of magnetic tape, how data is stored on tape, recording modes, tape errors, and their prevention. a data archivist faces two main tasks: data acquisition and data maintenance. to acquire data you must know how data is stored on tape and what your installation can handle. the information necessary to successfully transfer data between sites will be covered below, as will housekeeping measures essential to keeping magnetic tapes readable. data acquisition the physical aspects of magnetic tape. a standard magnetic tape has a 10?^-inch hub filled with '^-inch wide tape. the tape should be 2400 feet +50 feet, -0 feet. it is important to bear this in mind when purchasing new tapes, as tapes of less than 2400 feet are indicative of poor quality control . a magnetic tape is composed of three layers. the first layer is made of oxide; this is where the magnetized data is stored. this layer must be smooth and of uniform thickness (0.00045 of an inch). the second layer (the binder) is a glue used to make the oxide adhere to the backing. the binder must be flexible enough to reduce oxide chipping, yet, not too sticky, or the tape will stick to itself. the last layer is the backing, and it is usually made of mylar. creating a magnetic tape is a complex task; hence it is important to have rigid quality control standards. there are three important points to be aware of when exchanging data on tape: 1) is the tape labeled or unlabeled? 2) how many data files are on the tape? and 3) how many tape reels does the data span? tapes can come with or without labels, and files can come as a single file on a single tape reel, multiple files on a reel, a single file on multiple reels, or multiple files on multiple reels (reference (1),(2)). you must know what types of tapes your installation can handle and what types of utilities are available at your site for processing these types of tapes. check to see if one form is easier to handle than another. if so, see if you can receive the data from the sending organization in that form. here is a list of the other items you need to know when receiving a tape: 1. is the tape 9-track or 7-track (7-track is fairly old technology, but does still exist)? 2. how large is the blocksize of the data? some systems cannot handle blocked data. 3. at what density was the data recorded (6250 bpi , 1600 bpi , 800 bpi)? 4. at what parity was the data recorded (even or odd)? 5. what is the character set in which the data was recorded? some character set possibilities include: ascii american national standard code for information interchange. most commonly found on digital equipment corporation systems; ebcdic extended binary coded decimal interchange code. ibm, amdahl , burroughs; bcd binary coded decimal character code. cdc6600, honeywell; fieldata standarized military data transmission code. univac. the main point to all this is: know what your system has and what it can handle. your computing services department should have a handout telling the easiest way to receive data. this information is imperative for a successful transfer of data. i have found the government form "transmittal form for describing computer magnetic tape file properties" to be an invaluable reference when receiving or sending data (reference (3)). you can use this form in one of two ways. when receiving data, fill it out the way you need to receive the tape, and send the form along with your data request. when sending data, send the form filled out with the specifications of the way the tape was created. tape maintenance to contamination of the tape, physical mishandling, or problems with tape drives and/or tape cleaning equipment. here are some items that you want to consider to keep your tapes readable (reference (4)). reduce contamination. most tape contamination comes from the tapes themselves. oxide chips off the surface of the tape each time the tape is used. one way to reduce this chipping is to buy high quality tapes made with good binder. other contamination comes from carelessness in handling tapes or a dirty computer room. some ways to reduce contamination are: 1. reguarly clean tape drives. (at cna we clean the drives at the start of every shift and before tape intensive jobs. ) 2. maintain the proper recommended temperature and humidity--this helps to reduce the oxide chipping. 3. when cleaning the floors around tapes, clean the entire floor with a damp mop--do not sweep, dry mop, or dust. 4. minimize floor waxing— if you must wax, machine buff to remove the excess wax, damp mop with cold water to harden the surface, and buff again when dry. never use steel wool or other metal abrasives for buffing. another way to reduce contamination is to regularly clean tapes (once every eight uses). be very careful if you decide to do this, as some tape cleaners do more harm than good. i would not recommend using a tape cleaner on a tape that is error free. tape drives can also produce tape errors. your organization should have a regular tape-drive maintenance program where a field engineer does routine cleaning and alignment of the drives. this preventive maintenance is invaluable. how to minimize, correct, and prevent errors. most tape errors are due ftf long-term storage. once a tape has been hanging on a rack for about six months, it starts to deteriorate. pieces of dirt and chipped oxide start to cause dents in the backing of the tape. the tape may become unreadable because the dent is causing the tape to be too far away from the read head of the tape drive. the best way to prevent this type of problem is to spin tapes every six months. the preferred method is to have a program that scans the data on the tape to be sure the tape can be read. if there are problems, you will know it early and will be able to copy the data to another reel. if you wait for a long period of time to spin a tape, the damage could be permanent. the concept is the same as walking with a rock in your shoe. if you walk a block and take it out, you probably will not suffer any long-term damage. if you walk a mile with a rock in your shoe, you will have a hole in your foot. tape handling. tips for tape handling include: 1. keep tapes in the computer room. 2. tapes should not be laid on top of the tape unit. 3. external tape labels should be sticky labels that peel off and leave no residue. 4. never allow the beginning of the tape to trail on the floor. 5. smoking should not be permitted around tapes. tape storage. tapes should be stored in an upright position in a cabinet or shelf elevated from the floor, and as far away from sources of paper and card dust (line printers and card reader/punch) as possible. backups of valuable data sets should be kept off site. to save money, look for a sister institution to swap tapes with instad of paying for vault storage. tape transmittal. very often tapes are damaged in the mail or while being hand carried. tapes sent through the mail should be clearly marked magnetic tape--keep away from electric motors, scanning devices and magnets-do not x-ray. be sure to give the above adivce to the courier for tapes being hand carried. conclusions and recommendations to ease data transfer, know the answers to all of the questions on a form like "transmittal form for describing computer magnetic tape file properties" (reference (3)). be sure to ask for some type of a dump or map of the tape being created. this makes the sending installation look at the tape after it is written--that is, it forces them to be sure something was written to the tape. finally, be sure to get documentation on the data. these items include record layout, descriptions of all the variables and how they were derived, and the count of the total number of records in each file. to insure in-house data reliability, scan data tapes every six months to a year to be sure they are readable. at first sign of trouble, recopy the data. recopy vauable data sets to new tapes every three or four years. try to maintain cleanliness at your sit. do not use a cheap tape cleaner. it will do more damage than good. most important of all: buy the best quality tapes you can afford. referenaes (1) vax/vms magnetic taoe user's guide, vax/vms version 3.0. digital equipment corp, maynard, ma; chapter 2. may, 1982. (2) u.s. department of commerce. national bureau of standards. magnetic tape labels and file structure for information interchange. federal information processing standards publication 79. 1980. (3) u.s. department of commerce. national bureau of standards. transmittal form for describing continued on page 22 13 vol 41-2 lead editor's notes rebuilding, preserving and reproducing welcome to the second issue of volume 41 of the iassist quarterly (iq 41:2, 2017). the iassist quarterly has a focus on curation, preservation and reproduction of research, and all three bases are covered in this issue. the reproduction of earlier results from archived data is a validation of the data and also of the earlier research. the mimicking reuse of data for reproduction of the original results is the normal first step before use of the data for new purposes. this iq starts with a paper on reproduction. before reproduction is possible, intensive work is required at the earliest stage to curate the data, and in the case of older data as presented in this issue a costly process of rebuilding the data from old formats and forms of storage. between the establishment of the data as a resource and the subsequent reproduction, the preservation process secures the data for future use. the middle paper brings special attention to preservation of 3d digital data. at the iassist 2017 conference the presentation 'reproducing and preserving research with reprozip' was given at the session 'e3: tools for reproducible workflows across the research lifecycle'. this is presented here as a paper with the title 'using reprozip for reproducibility and library services' by vicky steeves, rémi rampin, and fernando chirigati. the authors work at new york university as librarian for research data management and reproducibility, phd candidate, and research engineer. they present reprozip, an open source tool designed to help overcome the technical difficulties involved in preserving and replicating research, ranging from digital humanities to machine learning as well as library services. the paper addresses the concept of computational reproducibility leading to capture and preservation of digital environments, and the creation of a file that encapsulates metadata about the computational environment including the operating system, hardware architecture, and software library dependencies in order to achieve reproducibility. the authors state that reprozip can be used to reproduce a plethora of applications, including data analysis tools, scripts and software. at the same conference in the session 'e1: preservation matters' jennifer moore of washington university libraries in st. louis and hannah scates kettler of university of iowa libraries presented their paper 'who cares about 3d data preservation?'. well, the iq does! 3d digital data preservation is necessary when for example an anthropologist produces digital 3d data as a preservation and presentation mechanism for an artefact. the 3d digital data has like other data to be treated for preservation. the artefact could be a building, and the paper holds much technical information and literature that refers to various interesting 3d projects; for example the augmented asbury park app that projects lost and now virtual buildings and attractions upon their earlier physical space using augmented reality. the last paper in this issue is 'retirement in the 1950s: rebuilding a longitudinal research database' by amy m. pienta and jared lyle, respectively associate research scientist and director of curation at icpsr at the university of michigan. this tells the story of the successful recovery of the important data from gordon streib’s cornell study of occupational retirement (csor). the paper includes the caveat that the work involved in rescuing these old data was many times more expensive than curating newer data would be. the csor followed a large (over 4,000 person) national cohort of retirement-age men and women in the period 1952 to 1958. the study is of great value for research in such areas as the relationships between health and gender and retirement. the data was deemed unrecoverable, as the punched cards did not directly match the documentation. further work and additional materials were required to make it possible. the data is enriched by collections of several types of health records and examinations; some remaining in paper form that can be consulted for closer investigation on-site at icpsr. submissions of papers for the iassist quarterly are always very welcome. we welcome input from iassist conferences or other conferences and workshops, from local presentations or papers especially written for the iq. when you are preparing a presentation, give a thought to turning your one-time presentation into a lasting contribution. we permit authors 'deep links' into the iq as well as deposition of the paper in your local repository. chairing a conference session with the purpose of aggregating and integrating papers for a special issue iq is also much appreciated as the information reaches many more people than the session participants, and will be readily available on the iassist website at http://www.iassistdata.org. authors are very welcome to take a look at the instructions and layout: http://iassistdata.org/iq/instructions-authors authors can also contact me via e-mail: kbr@sam.sdu.dk. should you be interested in compiling a special issue for the iq as guest editor(s) i will also be delighted to hear from you. karsten boye rasmussen december, 2017 vol30-1.indd 26 iassist quarterly spring 2006 the theme for the 33rd annual conference of the international association for social science information service and technology (iassist) is building global knowledge communities with open data. we invite your participation on may 16-18, 2007 in montreal, quebec, canada. the conference will be preceded by a day of workshops on may 15 and followed by a weekend of optional activities in the montreal area. details about the conference and iassist are available at http://www.iassistdata.org the theme, building global knowledge communities with open data, focuses our attention on the ever increasing globalization of knowledge and the importance of the “open data” concept in the development of knowledge communities. the conference will explore the inter-relationship of knowledge communities with open data. what is required to make data more “open” and available; what are the outcomes from open data; andwhat is the role of the data community in helping this happen? with this announcement, we seek proposals for papers, sessions, panel discussions, poster/ demonstration sessions and workshops on topics that address all aspects of the conference’s theme, including: * open data and the development of knowledge communities * the principles of open data * open data and its implications for documentation, metadata dissemination, preservation, curation and data authentication * new data partnerships in knowledge communities * e-science, cyberinfrastructure and open data * open data and digital repositories * developing trusted data repositories in knowledge communities * open data and issues of confi dentiality and disclosure * the development of statistical literacy in knowledge communities * developing an agenda for open data literacy * open spatial data and gis * open data and the role of the data librarian * life cycle models for managing data in knowledge communities * empirical research results on any of these areas for other key topics see previous iassist conferences at http://www.iassistdata.org/conferences/index.html the deadline for these proposals is january 16, 2007. iassist 2007 call for papers iassist quarterly spring 2006 27 iassist 2007 procedure individual presentation proposals and session proposals are welcome. proposals for complete sessions, typically a panel of three to four presentations within a 90-minute session, should provide information on the focus of the session, the organizer or moderator, and possible participants. the session organizer or moderator will be responsible for securing session participants, some of whom may submit paper proposals independently. workshops are typically for half a day (3 hours) and may include a hands on component. proposals should provide an outline of the content the workshop seeks to cover and the names of the presenters proposals should include the proposed title, an abstract (limited to 150 words) and 3 to 5 keywords based on the focus of the session. all proposals (paper, session, poster/demonstration and workshops) can be submitted using the online submission form under call for papers on the 2007 conference site soon to be available via the iassist conference webpage, http://www.iassistdata.org/conferences/ alternatively proposals may be sent via email to . please use a subject heading of “paper proposal your name”, “session proposal your name”,”poster/demonstration your name” or “workshop your name” replacing “your name” with the name of the author/session etc. organizer. following the january 16, 2007 deadline, the conference program committee will send notifi cation of the acceptance of proposals on or before february 15, 2007. all presenters are required to register and pay the registration fee for the conference. registration for individual days will be available. further information on travel and accommodation will be available shortly from the iassist ‘07 conference website: http://www.iassistdata.org/conferences/ online conference registration is scheduled to open in early february, 2007. make plans to come to montreal for iassist 2007 on may 15-18, 2007! about iassist iassist is an international organization of professionals working in and with information technology and data services to support research and teaching in the social sciences. the organization also explores issues of access, stewardship and the interconnections among social science, behavioral, biological, and health data. typical workplaces include quantitative and qualitative data archives/libraries, statistical agencies, research centers, libraries, academic departments, government departments, and non-profi t organizations. for further information see the iassist website at http://www.iassistdata.org. iassist conferences bring together data professionals, data producers, and data analysts from around the world for presentations and workshops covering new and persistent issues relating to access to data, its documentation, and digital preservation, with special emphasis on the social sciences. the social sciences have a long history of data sharing activity which will make the conference of interest to colleagues in disciplines where improving data access practices is on the policy agenda, and where there are clear overlaps with digital curation, data publishing, e-science/ cyberinfrastructure initiatives, and new interdisciplinary collaborations. questions may be sent to the program planning co-chairs, suzette giles and louise corti at iassist07@gmail.com vol21.3 28 iassist quarterly this paper discusses internet sites which enable the user to see images of plants, animals, and human disease conditions on computer monitors. the organization of the electronic archives at these sites commonly mimics bioscience taxonomy. among the entities experimenting with biological images for the internet are international and government science agencies, academic and research institutions, businesses, computer labs, interest groups, and innovative individuals. digital visualization for the internet is driven by evolving technology and legacy materials. most images presented are raster bit-mapped files in the graphics interchange format (gif, extension: .gif) or use the joint photographic experts group compression (jpeg/.jpg). gif and jpeg standards are incorporated in hypertext mark-up language (html) and in software for viewing, editing, and printing images. platforms, operating systems, and graphics software can almost universally import and export files as gif and jpeg often though the standard for transporting images (tiff) between formats. software “plug-ins” allow additional image file formats to be viewed. the physical source media for the electronic images of biological subjects include specimen, film and digital photographs, slides, prints, drawings, x-rays, magnetic resonating images, electrocardiograms, and sonograms. images displayed were originally digitized in dozens of raster and vector image file formats native to proprietary equipment and software used to capture, scan, or edit electronic images. human factors figures in visualization. can people recognize the subject of the image? recognition involves cognition, culturally determined perception, cultural preferences for style, and the visual experience and training (technical or aesthetic) of the viewer. composition and clarity critical for human perception may be affected by the view, shape, and indication of scale, details, or color if critical for identification: what about the subject is displayed. focus, size, tones, background or context, the placement of the subject, and the method used to create a digital image affect human perception: how technically the image is made. images with high resolution can convey finely detailed content and color tone gradients, can be magnified (zoomed up), and require a higher density of pixels: the differentiated units of equal size that make up an electronic image. high pixel density implies more bytes of data and larger files. large files tax the capacity of internet packet transmission, and serving and receiving computers. “lossy” principle compressions (including jpeg) reduce file byte size by literally “losing” visual bits, degrading images irretrievably. sizes of raster files posted for viewing on the internet to depict or identify biological species and human disease conditions were observed to range from around five kilobytes for highly compressed photo “thumbnails” and black and white line drawings to hundreds of thousands of kilobytes. the large size of electronic image files is one reason why more image files are posted for downloading than for immediate viewing on the internet. file size influences the commonly made decision to locate image files in a terminal position, at the end of an archivein an auxiliary computer graphics interface (cgi) bin or cdrom that exclusively warehouses image files. on home pages or directories, links to images may be embedded in items of an inventory list of subject or file names, or brief descriptive text, or lossy previews. (1) the terminal or attached position of image files fortuitously permits linking them to and from documents within and among sites on a unique file address. if directory and file names are semantically meaningful in human language, then their content can be found more easily by internet search engines which automatically check directory and file names. using latin or common names to label the directories and files where information or images about a species are located sets up a smooth interface to browsers. files named “c|/gif/birch.gif “ and “flora/asteraceae/ launeas/l-aborenscens/ laborescense.1.jpeg” can be identified as images from their extensions, .gif and .jpeg. one file identifies the subject of the image using a common name, birch, while the other uses a scientific name, l. aborescense. each latin genus and species name is short hand metadata for a unique taxon, defined and characterized in biological data bases and monographs. bioscience ranks life forms— kingdom, phylum, class, order, family, genus, species (and their subs and branches). the second file is nested in a cascade of directory names which mark the subject’s botanical by leslie a. brownrigg * emerging internet image archives visualizing biological species and medical conditions fall 1997 29 classification. several internet sites have adapted the taxonomic levels of biological classifications as a structure for their archives of electronic materials about species. materials in hierarchies. taxonomic levels that group similar life forms gave a jump start to the logical organization of elaborate archives of biological species. several internet enterprises are vying to construct taxonomies or phylogenic lists to describe and compile data for all the living species on earth— or in a large region. some of these sites incorporate illustration. others are choosing technologies incompatible with the special requirements of image files. the organization of “the tree of life: a phylogenetic navigation systems” is instructive. tree of life is a federated site distributed on 18 servers in three countries and is linked to additional servers where specialized sites note their on-line “place” in the tree. tree’s “page” on frogs (salientia) resides on the university of texas server along with “herps of texas.” adding a refereed description for a higher order “clade” or “terminal taxon” to the tree is facilitated by the data entry program macclade which can read from and to the widely used phylogenetic analysis using parsimony (paup) taxonomic program. each basic “page” in the tree of life covering a taxon or more generalized life form is illustrated with scientific line drawings and compressed photo images. some images hide the html indices to clade latin names by which tree’s internal web crawler navigates among its “pages.” images can be attached or linked at any point to tree documents and generated from tree’s random and searchable internal browser facility. tree’s photographic images range from 11,000 to 328,000+ bytes in size .(2) tree’s is a collaborative effort among entomologists at the the universities that serve and host this federation. its content is best developed for insects. unesco’s man and the biosphere (mab), national government scientific agencies, and research or academic institutions also initiated on-line sites with comprehensive ambitions. in a complex of sites sponsored and affiliated with u.s. government agencies, attempts are underway to link sites and to standardize meta-data aspects for the inventory and description of biological species. agencies in three cabinet-level departments of the united states federal government and collaborators are forging distributed internet inventory data bases. nexi include the us information center for the environment (ice), national biological information infrastructure (nbii), and “itis”the inter-agency taxonomic information system.(3) standards, data base layouts, and “meta maker” data entry programs being tested emphasize location and text data and cannot accommodate images. (4) images on the sites affiliated with this complex are rare, and tend to adorn home pages, or lie in ad hoc galleries, or buried in menus. (5) bioscience requires at least one “voucher specimen” be held in a scientific repository such as a botanic garden, herbarium, zoo, museum, university, or research organization to establish and reference the existence and scientific name of a unique species. many internet sites experimenting with visualizing originate from these institutions that collect and classify. some sites cling to traditional presentations of specimen, showing animals confined in aquaria or cages or dried plants laid out on conventional herbarium sheets (6). others deploy new media to photograph living organisms in fields and forests (7) or through the microscope. many institutions post digital peeks into their formidable collections (8) and label images as copyright, hesitant to reveal their research patrimony absent “pay per view” or use fees. proceeds from the market for digital visuals can potentially reimburse the high costs of processing, archiving, and serving image files. on-line mechanisms include limiting access (on passwords) to subscribers, collecting fees to view or to download files, and receiving a payment for each “hit” to the site as use royalties from subscription services (9) or from advertisers. commercial tie-ins are possible. several sites display biological visuals to catalog items for sale and businesses sponsor sites and popular federations to attract customers. until recently, one herbarium’s internet presence was hosted and sponsored by a garden supply site. two popular sites posting images of felines were sponsored respectively by a pet food company and an airline. access to the bulk of images on-line that visualize human disease conditions is already structured primarily through subscription services. access to electronic archives of medical imaging tends to be strictly privatized to the archive’s donor/users. an offline market for digital images exists in cd-roms, software, and print ads. some of the most visually and technically sophisticated sites featuring numerous images of biological species prepared for immediately viewing on the internet were designed and initiated by experimenting individuals and computer teams. henriette kress issues directories full of image files depicting plants, parsimoniously linking each file name she lists directly to an image file. kress’ botanical site served from helsinki, finland (10) is relayed by “mirror” sites on three continents, a tribute to the site’s outstanding content and popularity. tim knight developed and maintains sites dedicated to images and other electronic media to depict the living primates and designed internet sites for several zoos and scientific associations.(11) knight’s sites are worth “visiting” for their design and the quality and economy of images. a computer technology group at a texas university is posting images of crops and wildlife. (12) the korean research and development information center’s bioinfo computer imaging project serves 2300+ optimized animal pictures in its on-line archive which can be searched using english vernacular 30 iassist quarterly names in the roman alphabet or korean names and script. (13) independently, the designers of these four sites pioneered similar technical solutions to problems that internet visualization projects confront. the exact technology (equipment, software, file format, and post digitalization editing) they used differs. in common, the four sites post attractive color images that can be magnified at least once, in files sized in the modest 20-250 kilobyte range. in common, the four sites created the digital images displayed by by scanning film photographs or slides. they select among their own and contributed photographs those that can be successfully scanned at quite low pixel densities (under 100 dots per inch), a technique that limits image file size at the moment of digital creation. photographs suitable for this technical approach must be infocus, high contrast, high content, and centered on the subject. it is striking how many internet images show the faces of animals and flowers of plants, suggesting face or flower are regarded as the the most recognizable or visually interesting feature. images posted to portray biological species are not standardized as to perspective or view whether the organism is large or microscopic. by contrast, medical imaging by x-ray, magnetic resonance (mris), and other devices to diagnose or monitor patients’ disease conditions have strictly standardized views. protocols exist for positioning subjects in relation to the image capturing device to achieve the correct views. resulting images require a specialized training to interpret. a principle of file compression can be applied to highly standardized views : to reduce file size by removing extraneous features located in predictable areas of the image rather than to “loss” randomized pixels throughout the image. fixed view medical images (14) can be stored without the extraneous features, and templates of what typically appears in the removed sections can be reinserted to “decompress” the image for viewing. human disease conditions are not registered in a single universal taxonomic hierarchy like that for biological species. there is less consistency in approaches to organizing medical images on the internet. each disease considered unique will be identified by its own diagnostic code and by vernacular or scientific names. diseases are grouped into areas of specialization recognized in the medical community, that is, grouped either as related to a particular organ or system of the body, or to a disease processes, or stage in the life cycle. thus, medical images cluster on sites for medical specialities. images of melanoma skin tumors (15) can be found in dermatology and cancer archives. additional organizing principles are applied. medical images on the internet (and private intranets) are sometimes organized in archives containing a single technological type of medical image — just cat-scans, just mri, although from different institutions and concerning different patients, perhaps because interpreting each type of image is still another medical specialization. medical imaging is sometimes grouped across types by source (hospital, clinic, practice) and an identifier like date of deposit. medical images are organized in the “case” of one patient, either walled off in a confidential archive or broadcast as a didactic history in one of the commercial medical teaching and reference internet subscription services. electronic images depicting human medical conditions produced daily due to their role in diagnosis and treatment are lost to retrieval and research when buried in incompatibly indexed and organized archives. general observations follow relevant to sites with extensive on-line archives that include image files. site designs typically redress the intrinsically large size of image files by applying compressions and by exercising selectivity. giving computer directories and files names that have meaning in natural language makes them easier to search. archives organized by establishing layered classes that generically includes every particular thing in its class, are easy to navigate. end notes (internet addresses of sites): (1) examples of image to image directories are missouri botanic garden’s directory at http://www.mobot.org/mobot/research/ gallery.lagallery2.html http://www.turtlebackzoo.org/tbz25.htm http://www.cbs.nl/temp/lmi/ff300rnr.htm (2) tree of life’s format is explained at http://phylogeny.arizona.edu/tree/home.pages/ intro.html and see http://phylogeny.arizona.edu /tree/eukaryotes/animals/ chordata/osteostraci/osteostraci.gif http://phylogeny.arizona.edu/tree/eukaryotes/animals/ chordata/ichthyostegalia.gif http: //phylogeny.arizona.edu/ tree/eukaryotes animals/ anthropoda/hexapoda/coleoptera/colerptera.html. (3) for more information about this u.s. government internet enterprise see http://www.nbs.gov/ nbii http://trident.ftc.nrcs.usda.gov/itis/ http://ice.ucdavis.edu/ http://www.mobot.org/mobot/research/gallery.lagallery2.html http://www.mobot.org/mobot/research/gallery.lagallery2.html http://www.turtlebackzoo.org/tbz25.htm http://www.cbs.nl/temp/lmi/ff300rnr.htm http://phylogeny.arizona.edu/tree/home.pages/intro.html http://phylogeny.arizona.edu /tree/eukaryotes/animals/ chordata/osteostraci/osteostraci.gif http://phylogeny.arizona.edu /tree/eukaryotes/animals/ chordata/osteostraci/osteostraci.gif http://phylogeny.arizona.edu/tree/eukaryotes/animals/chordata/ichthyostegalia.gif http://phylogeny.arizona.edu/tree/eukaryotes/animals/chordata/ichthyostegalia.gif http: //phylogeny.arizona.edu/ tree/eukaryotes animals/anthropoda/hexapoda/coleoptera/colerptera.html. http: //phylogeny.arizona.edu/ tree/eukaryotes animals/anthropoda/hexapoda/coleoptera/colerptera.html. http://www.nbs.gov/ nbii http://trident.ftc.nrcs.usda.gov/itis/ http://ice.ucdavis.edu/ fall 1997 31 http://ice.ucdavis.edu/us_national_park_service/ national_park_service.html http://ice/about_nps_databases.html http:// www.fs.fed.us/ (4) http:://www.itis.usda.gov.images/twbscr.gif http://www.nbs.gov/nbii/current.status.html http://www.fw.vt.edu/fishex/macsis.html (5) http://www.nbs.gov/features/photogallery http://www.itis.usda.gov/images/twbscr (6) http://www.inbio.ac/cr/tipos/ herbario/cgrayumii.jpg http://www.mpiz-koeln.mpg.de:80/~stueber/ross/potato/ sucrense_complete.jpg (7) http://bluehen.ags.udel.edu/udgarden.html (8) http://www.nfrcg.gov.htm http://www.nmhn.cgi-in/wdb/fish/catalog/query/ catgno==00334647 (9) http://members.aol.com/zooweb/pictures (10) http://sunsite.unc.edu/herbmed/pictures http://sunsite.sut.ac.jp/pub/acadmeic/medicinal/ alternative-healthcare/herbal-medicine/unc.edu (11) http://www.selu.com/~bio/primategallery/primates/ species.html/index.html http://www.selu.com/~bio/ (12) http://www.ics.tamu.edu/flora/ gallery.htm http://leviathan.tamu.edu: 70/1s/slides.ef/ http://straylight.tamu.edu/tamu/tracy/chklcon.html http://ranch.tamu.edu/rlem/ (13) http://bioinfo.kordic.re.kr/animal/ (14) http://www.laurie.umdnj.edu/database/ collectedapps.acgi$ imagelist_jpeg?11843 (15) see the dermatology online atlas from the erlanger image database http://ultaet/med/kli/derma/biblddb/diagnose/dg_m.htm and see md challenger sample photo: http://www.embbs.com/img/i0000012.jpg * paper presented at iassist/ifdo ‘97, odense, denmark, may 6-9,1997. leslie a. brownrigg, united states bureau of the census http://ice.ucdavis.edu/us_national_park_service/national_park_service.html http://ice.ucdavis.edu/us_national_park_service/national_park_service.html http://ice/about_nps_databases.html http:// www.fs.fed.us/ http:://www.itis.usda.gov.images/twbscr.gif http://www.nbs.gov/nbii/current.status.html http://www.fw.vt.edu/fishex/macsis.html http://www.nbs.gov/features/photogallery/ http://www.itis.usda.gov/images/twbscr http://www.inbio.ac/cr/tipos/ herbario/cgrayumii.jpg http://www.mpiz-koeln.mpg.de:80/~stueber/ross/potato/sucrense_complete.jpg http://www.mpiz-koeln.mpg.de:80/~stueber/ross/potato/sucrense_complete.jpg http://bluehen.ags.udel.edu/udgarden.html http://www.nfrcg.gov.htm http://www.nmhn.cgi-in/wdb/fish/catalog/query/catgno==0033464 http://www.nmhn.cgi-in/wdb/fish/catalog/query/catgno==0033464 http://members.aol.com/zooweb/pictures http://sunsite.sut.ac.jp/pub/acadmeic/medicinal/alternative-healthcare/herbal-medicine/unc.edu http://sunsite.sut.ac.jp/pub/acadmeic/medicinal/alternative-healthcare/herbal-medicine/unc.edu http://www.selu.com/~bio/primategallery/primates/species.html/index.html http://www.selu.com/~bio/primategallery/primates/species.html/index.html http://www.selu.com/~bio/ http://www.ics.tamu.edu/flora/ gallery.htm http://leviathan.tamu.edu: 70/1s/slides.ef/ http://straylight.tamu.edu/tamu/tracy/chklcon.html http://ranch.tamu.edu/rlem/ http://bioinfo.kordic.re.kr/animal/ http://www.laurie.umdnj.edu/database/collectedapps.acgi$ imagelist_jpeg?11843 http://www.laurie.umdnj.edu/database/collectedapps.acgi$ imagelist_jpeg?11843 http://ultaet/med/kli/derma/biblddb/diagnose/dg_m.htm http://www.embbs.com/img/i0000012.jpg attrition and the longitudinal surveys of labor force behavior: avoidance, control, and correction by dr. patricia rhoton data archivist national longitudinal surveys center for human resource research ohio state university since 1966 the center for human resource research has been analyzing the longitudinal surveys conducted by the census bureau for the department of labor. the main purpose of these surveys is to study the labor force activity of different population groups. the original groups included men who were 45-59 years old in 1966, women who were 30-44 years old in 1967, men who were 14-24 years old in 1966 and women who were 14-24 years old in 1968. in 1979 a new survey, conducted by the national opinion research center in chicago, was added for young men and women who were 14-21 in that year. each of the five surveys is designed to collect information on all phases of the respondent's labor force activity and on other characteristics such as educational attainment, health, family composition, and financial status that are known to be related to such activity. the original plan in 1965 was to interview the same respondents each year for a period of five years. because of the usefulness of the data and the relatively small sample attrition, the decision was made at the end of the first five year period to continue for another five years. the interview pattern was changed at this time from a face to face yearly interview to a 2-2-1 pattern. each respondent was contacted by phone e^ery two years, then again in person one year after the second phone interview. this pattern was used again in the third five-year extension obtained in 1976 and during the fourth five-year extension, obtained in december, 1982. at the time of this most recent extension a study was done looking specifically at attrition within the different cohorts. longitudinal studies in general have several advantages over the more frequent cross-sectional studies. while longitudinal studies are very expensive, the data are collected in great detail over time with the respondent reporting events and attitudes as they occur rather than retrospectively. gathering the data in this way also enables the researcher to go beyond issues of correlations to address the more urgent issues of causality. the main advantage of a longitudinal survey, following the same set of respondents year after year, creates its two major problems, however. the first is the difficulty of locating the respondents for the subsequent interview and the second is maintaining respondent cooperation over repeated interviews. attrition in the nls table 1 shows the number and percentage of respondents for all interviews up to and including the 1983 questionnaire. the base year shows only those respondents who were interviewed that first year. between the original screening and the first interview, part of the eligible respondents were lost: 9.0 percent for the older men, 5.5 percent for the older women, 8.3 percent for the young men, 5.8 percent for the young women, and 11.5 percent for the new youth. while table 1 shows the distribution of noninterviews between and among the five cohorts. tables 2-5 show interview/noninterview status for the four older cohorts by reason for noninterview. while there are shifts in the distribution of a particular noninterview reason during a particular year, a consistency appears in the rate of attrition within each of the four older panels. the method of interview, whether face to face or by telephone, does not seem to affect the attrition rate. some of these losses to the sample are unavoidable. for the mature men (table 2), for example, an increasing percentage of the sample losses have been due to respondent's death. the mature women's survey (table 3) has the second highest retention rate among the four older cohorts. this high rate is probably due to the fact that this group is very stable and has low geographic mobility. the young men's cohort has the lowest rate of retention and has been the test case for new attempts to stop the gradual decline in sample size. a variety of factors account for the difficulty in locating these respondents: completion of school, acquisition of new jobs, formation of families and movement in and out of the military services. the higher rates of attrition in the earlier years were attributed to the influx into military since the sample was drawn and the initial interviewing done during the vietnam war. however, rates remained high even as the respondents returned from the military. the young women's cohort, which is similar to the young men's with respect to completion of school, acquisition of new jobs and formations of families, had the added challenge of name changes accompanying changes in marital status, yet the overall response rate has remained high. the new youth cohort has benefited greatly from the lessons taught by experience with the older four cohorts. in 1983, the response rate for this group was 96.3 percent. comparing this cohort with the young women in the first five years, the cohort that had the best retention rates of the older cohorts, shows that a difference in procedures and techniques can decrease attrition on a substantial basis. not only does norc have a higher overall interview rate, the organization seems to be better at retrieving respondents. in 1982, of the original 1979 sample, 96.0 percent were interviewed. some of these had not been interviewed in previous years: 2.2 percent in 1980, 1.1 percent in 1981, and 0.5 percent in 1980 or 1981. only 165 respondents (one percent) of the original sample has had only one interview after four rounds of the survey. in 1983, the number of respondents who had had only one interview dropped to 115. over 11 thousand (90.7) of the respondents were interviewed every year, and 5.5 percent had completed four out of the five interviews. ^2 01 t) is 01 o> i t-l o 1-1 v 0) i cc 0> o *j w o 00 •— »h .-( .1 00 £•:: sj^ njf: 0^ 2^ 52y.^ aj # o c — t > o co m c c l. 01 ^ «o (gx s £^85 »8s as s 2^ i o 00 tn i c^ co tn es) cd i tn ^r i i i i i tos coo c^ to • . u^ i i i i i i i in i i i 1 i i i cd i cm 00 o i cd e^ i cd e^ m 11 i co i-" i ^ c^ i o^ i o ^^ c^cdcdcd*0__,__osososososososososososos 0) w i) ci v 01 ai t3 0) x: od c =/ ii 1^ a*' ffl-j^ •0 2r ?'^ « c s l-h ^ o ^ o w «-! ffi co ol o o .-i ^ ^ ;d o ro ^ cd i t-i 0> i 0» t> i ^ t00 cj co ^ *rt cd t— oo tbttttto^ o o^ o) 93 o) o^ ^ co co co 00 oo f 2 to &-z i cd ^ ^' tirt in oj ^ o* ift 0> 0> 0> 00 00 « 2j? oi i oi i i i os -^ tift <*5 c>) ec c^ i o) t« 00 o 1-« f-h i c4 x: t. 03 c sb 9* ?* a> «o a*' e.jp sz be — _ ._ t:|^ oo ^-> tn » o£ bo o> « ^ be c » «: o « c'l s 3 rh « # x: he ^ ?i* co w os qj a> o oj as c -^ — 0; o a> oo co e^ ift c^ ^ oi es) ,-, .-h >oioo ^<-^o^ oooico^t-r-tci' > (c ^ f^ c^ c^ tc>3 c^ rt 1-h i/? oi cc to r5 oi cd 00 i-i oi c-:> <* o"' cc ^ o> ^ ^-i ca co ^ o cd ^ v c> '"f c^iooo ococo e^olu^ocdoooo*-! cd^co r5t-«m oioocoiits^cocirt*rs«— • oic>ao5 »-hoot/>i/5pooooi'»'cd ^•oocd oc-oo c^^coo^cottoi cocorn t~-<-ic•-i,-ojin. •-i cj cv)m wi tc^ cd ^ co .-< »ft o tt to o^c^o »—'o^'—"gi^-'qooool^ 1 co n l> m o ^ o > r,qj o '^' — ^ 03 st >.— •>. = bt 01 god-oi-cc oo.£»-wi-c__ c5 o c^ irt »-' • be o c c — o — — ._ t. ffl »3 »i*oie-oi— e c3 c (0 cc c l. .o 1. i, — o) 3 eh (1w n. beo — v 01 j: tt. "^ tie c <£> 0) w en s = '-j(? »—•ooiftin^tint-»—ico «-<^0^01»-h i 0)01000) oe cm co in —1 co — -v co co o 'o tin c* 00 co o) co ' (o in o in rtn en cd e>) co ^ ^ 00 ^ o t» o t00 co be c (o :«t: o o>cooof-iai^< > (dooosfc o> o> *t o) ^ r— 2 2*5 co c~0> co 9:> i «-< «r 05 ^h in oo f-i ^ rh 0> oo rh co o to <-" 00 co o 00 ^ ctto v^ootftino i cocotooo mcj-ncniojo e^jo^oi uo 0> co irt c^ t-h o> in « t-« e.j^ »-it-co^ct>co i caot-^ i-ic^inifi'j'tint-*-i^ in in oi o c4 to co »-* t^ o 00 <-< ^h m t^ m coin^i^hoom »—ico «-h 00 —• oi 00 c«i -n to ^r ^ ^ ^^ in c^ ^r t> cm »-t in co cm to i o o) o> to tto o ^ •** ^ os 00 "c to a> o o) ^ o) in »-" ^ o to ^ 1-< *» ^ 05 to «-' 1 cm 0> t^ in tr-^ o o 0> 1-1 cm oi in •-• ca i ^ 0) a> a> j3 > q be 0) — c to 00 t^ o r«-h o 's' in cm oi co in ^ <-" t^ ^ o) oo o o> ,-. ^ co ^^ 1-1 "u to a> o o> oi co cm ^^ to^ ,— , ,_ 1—i in e-s c o; . 01 >— 1« ' ^ ffl cj 3 ' c/2 u o > — c _ »£ o t— oj . o 0) ffi o 03 x cri b b. ^ < c a> ct> i i (— oi ' o o ( cb 3 o o o oo cm 10 £=# 2= 3 ^ w *5 ^ a> « 05 *— ' s =>-• ^ s< ^ ^ — o — c — • 1c !? ^ *-• oo co cm in cd o »-l to cs ^ cm cm cm ift i-» o in 1-1 ^ o cd oi 00 c^ in «d o a> ^ o to co ino cm o cm rt 00o co tc^ cm i-« i f-h — cm *t ol «-i cm ca co 00 ttt cco • oi in co co co 00 00 ^ c* co cm cm co i-h in cm co in in 00 c^ r»-< co cm ^h cm m ^ c* cm co ^t •-k o o ^ co o •-«»-" co co in tc^ 0> 00 oo c^ 00 •— ' cm cm co 00 cm c^ 1-1 ^ ta. cm to cm ^jcm 1-1 cc s .— .— o "c *oi ol _ c 05 i o o o c — cri o o o * 03 o o o o cfi 3 cm o cm oj di * o m bcw r^ eo ^ ^ + 11 the impact of attrition in representativeness this gradual decline in sample size over time becomes very important if it results in a biased sample. while each cohort was checked at the end of the first five year series of interviews and smaller checks were made in the context of reports on occupational distribution, educational attainment, age distributions and marital status with nationally represented published data, no one looked at all the cohorts systematically until 1982. at this point the issue of representativeness had to be addressed as part of the proposal to extend the cohorts for another five years. such a study could be done in essentially two ways. first, the remaining sample could be compared against some outside group, such as one from the decennial census or the current population survey. comparison with an outside sample was difficult given time contstraints and the fact that the decennial census data were not yet ready for release. while the cps data were available, differences between the cps and each of the four older cohorts had already been documented in the first year. the second alternative was to compare the characteristics of all respondents interviewed in the initial year to see how much difference, if any, there actually was. each cohort was checked for differences in the age distributions, educational attainment levels, employment status, industry and occupation distributions, educational attainment levels, employment status, industry and occupation distributions, marital status, smsa status, annual income distribution and wages and salary distribution. the young men and young women were also checked on enrollment status. a separate evaluation was done by race for each of the four cohorts. table 6 is an example of the type of table constructed for each group. the ten year sample was weighted using two methods: the entry level weight and a ten year weight, which includes successive adjustments for each year's noninterview. for all the cohorts except the young men the relevant comparison was between the entry year weighted figures and the ten year sample using the ten year weight. in the young men's cohort, the 1966 sample using the 1966 weights was compared to the 1976 sample using the 1966 weights because the 1976 weight had been adjusted to include individuals formerly in the military. since young men already in the military had been deliberately excluded from the young men's sample, using the 1976 weight could create apparent differences where none existed. for this group alone, it was more appropriate to use the 1966 weight. table 7 summarizes the distribution of differences by cohort and shows that for most of the characteristics the differences between the two samples were less than two percentage points. after the differences were identified, statistical tests of significance were computed for each of the comparisons. table 8 shows the number of statistically significant differences at various levels for each cohort by race. while the number of differences were higher than would be expected by chance, several were based upon small sample cases in the initial year and characteristics with only two values. in the latter cases a statistically significant result in one category means the other category will also be statistically different. after reviewing the entire set of tables it was clear that the noninterviews had not seriously distorted the representativeness of the sample. given this finding and the ability to change the weights to eliminate any potential bias, the decision was made to continue all four surveys for another five years. 12 table 7 nimber and percentatge of differences by panel absolute differences (%) panel 0^2 2^3 3+ total mature men black 34 (73.9 8 (17.4) 4 (8.7) 46 (100.0) white 43 (95.6) 2 (4.4) 45 (100.0) mature wcmen black 42 (93.3) 3 (6.7) 45 (100.0) white 45 (100.0) 45 (100.0) young men black 30 (73.2) 5 (12.2) 6 (14.6) 41 (100.0) white 43 (97.7) 1 (2.3) 44 (100.0) young wcmen black 33 (82.5) 6 (15.0) 1 (2.5) 40 (100.0) white 40 (95.2) 2 (4.8) 42 (100.0) table 8 nunber and percentage of statistically significant differences by panel level of significance panel 1% 2% 3% mature men black 4 (9.1) 7 (15.9) 12 (27.3) white 4 (9.1) 7 (15.9) 14 (31.8) mature wcmen black 2 (4.5) 2 (4.5) 3 (6.8) white 1 (2.3) 4 (9.3) 5 (11.6) young men black 1 (2.6) 4 (10.3) 6 (15.4) white 2 (4.7) 4 (9.3) 6 (14.0) young wcmen black 1 (2.6) 3 (7.9) 4 (10.5) white 1 (2.6) 2 (5.1) 2 (s.l) 13 it is unclear, however, how further erosion of the samples will affect this representativeness. concern with this issue, together with the high noninterview rates that norc was having with the new youth sample, led to an evaluation of the rules that had been established in the original five year period and an attempt to see if it was possible to retrieve some of the noninterview cases. retrieving former noninterview cases since the young men panel had lost the most respondents, it was the target for the first attempt at retrieval. respondents from the 1975, 1976, 1978 and 1980 survey years who normally would not have been included in the workload (i.e., attempted to be contacted) because of their noninterview status for those years (refused, unable to contact, institutionalized, moved outside the u.s.) were sorted and a sample of 279 respondents selected. several changes occurred in procedures for contacting these special respondents. no restrictions were placed on the number of telephone calls, mileage or time spent in locating and retrieving these respondents. each interviewing packet included the respondent's most recently completed interview and household record card, as well as the most recent questionnaire and all record cards for any other household members participating in any of the other cohorts. in addition, an expanded list of methods of locating respondents was included. as a result of these additional steps, 104 (37.3 percent) respondents were interviewed. these have been identified and will be checked as soon as the data tapes are available from the census bureau to see if they differ in any way from the rest of the respondents. if these respondents remain in the sample for the next round of interviews in the last part of 1983, a concerted effort may be made to use these procedures during the regular interviews and in similar attempts to retrieve noninterviews in the other three cohorts. differences between census and norc one of the biggest differences between census and norc is the amount of locating information obtained from the respondent. norc gets more information and asks for individuals with specific relationships depending upon the respondent's circumstances. the interviewer starts by asking the name, relationship, address, and phone number of the person most likely to know where the respondent is. if the respondent is living in a dormitory, fraternity, sorority, hospital or other temporary situation, the interviewer is instructed to obtain the name and relationship of a householder at a permanent home address. if the respondent is married and living apart from a spouse, the spouse's address and telephone number are requested. if the respondent is not living with a parent and has not provided a parent's name, this information is obtained, including whether or not the parents live together. the name of another relative with whom the respondent is in contact and the names of friends and places where the respondent goes when not spending spare time at home are also obtained. respondents are also asked nicknames, maiden names if they are married women, and whether or not they expect to move in the next 12 months. this extensive list gives the norc interviewer a real advantage when contacting someone on the list, since the ability to mention the respondent's parents, relatives, friends, hangouts or nicknames demonstrates that the interviewer knows the respondent to some degree and may make the reference more willing to give out 14 information about the respondent. another major advantage that the norc interviewer has over the census interviewer is the existence of a centralized locating shop in chicago. the person working at the locating shop has access to all previous questionnaires, original copies of locator documents and information about the respondent's brothers and sisters. working with this additional data, the respondent can usually be located by phone and reassigned to the same or another interviewer. the census interviewer starts out with less information to locate the respondent. s/he has a questionnaire with a label indicating the respondent's name and most recent home address. in addition, there is a household record card for each respondent that contains the telephone numbers, all the addresses where the respondent has lived since the survey began, the names of all persons who have lived with the respondent, and the names, addresses and telephone numbers of all persons who have lived with the respondent, and the names, addresses and telephone numbers of only two persons who will always know where s/he can be reached. besides the more extensive locating supplement that norc builds in the interview, several other differences appear. for the new youth cohort, each respondent is paid $10.00 for a completed interview, since many researchers believe that even a small amount of money helps in obtaining cooperation, especially among younger respondents. the new youth respondents also had the opportunity to take a series of tests that the department of defense needed to evaluate tests given to individuals in the military. for these tests, which take several hours, the respondents were paid $50.00. when the four older cohorts were first interviewed, paying respondents was not as well accepted. now there are fears that starting this procedure with the older cohorts would cause concern on the part of the respondents. another procedural difference is that in the new youth cohort, the respondents are told up front that they will be interviewed each year for the next several years and are therefore aware that they will be contacted about the same time next year. the census interviewers are told only that they may be conducting additional surveys, and should not tell the respondents that this is the last time s/he will be interviewed. the lack of an answer to give the respondent, in addition to the 2-2-1 pattern, probably leaves the respondent without a sense of when or if s/he will be contacted again. while this ambiguity may not have an impact on their cooperation in the survey, the norc approach leaves the respondent with a greater feeling of certainty about the interviewing schedule. revising the rule for dropping respondents after the first year respondents from the four older cohorts who refused to participate or had died were dropped from the census sample. those who were not interviewed for any reason for two consecutive years were also dropped. the only exception was made for the young men's sample with the respondents who were in the armed forces. since the sample was to represent the national civilian, noninstitutionalized population, the young men were not interviewed while they were in the armed forces but they were retained in the sample and picked up the first interview after they left the services. however, norc's success in retrieving respondents even after they refused and the success in the young men retrieval effort resulted in a change in these rules. currently no respondent will be dropped except those who have died. norc goes back each year and attempts to interview all living respondents. 15 maintaining respondent cooperation while both census and norc send out advance letters about the entire survey stressing the importance of the respondents' cooperation, norc also sends out a newsletter that tells respondents in a very "chatty" format about some general results of the previous survey. the census bureau had a short, formal fact sheet that goes out with the cover letter, but the interviewers reported that the respondents did not feel it was very useful. for the 1982 young women's survey, a more extensive description of the surveys and a list of the research results from the survey were sent to any respondent who filled out a postcard requesting additional information. over one-third of the respondents interviewed in that wave mailed in the postcard. a variable will be created identifying these respondents and if reception of the handbook increases the response rate for the next round, the handbook will be offered to respondents in the other three cohorts. conclusions the new youth survey at this time has considerably better response rates than any of the four older cohorts. a great part of this success can be traced to solving problems that developed over time in the older four cohorts. while the necessity of keeping the same measures over time prevented change in the older four cohorts, these problems were corrected in the first wave of the new youth. questions that the respondents or the interviewer had difficulty with in the older four cohorts were altered so that there was no confusion from the very beginning. perhaps most important, given the highly mobile nature of this age group, much more detail was obtained on individuals who would always know where the respondent was. in addition, more information about the survey was given to the respondent before, during and after each interview. all of these factors combined have resulted in a response rate that is very good for any survey and exceptional for a longitudinal survey in its fifth year. 16 44 iassist quarterly the potential for computer communications among icpsr representatives which impede its application. a recent study' revealed that among the important factors determining the usage of computer conferences was the prominence of a terminal^ within a person's immediate work environment the most active members of the conference of study were those who regularly used a terminal during their daily routine and whose equipment permitted the use of packet-switching networks. on the basis of these findings and with the advent of the consortium data network (cdnet), a survey was conducted of the participants attending the 1985 biennial meeting of the ofticial representatives (or's) to the inter-imiversity consortium for political and social research (icpsr). by charles humphrey computing services university of alberta a survey of official representatives immediately prior to the icpsr business meeting, a questionnaire was distributed which focused on two topics.^ the first of these dealt with the availability and use of terminals in the introduction availability of computer-based communication, especially electronic mail and computer conferencing, has become commonplace on most north american campuses. through such technology, staff at many universities can now communicate with their colleagues at other institutions as easily as they do with those on their own campus. the promise of computer communications lay in facilitating scholarly and professional exchanges which are immediate, easy, inexpensive and widespread. however, even though this technology has become extensively accessible, obstacles do exist 1^ charles humphrey and wendy watkins, "datalink: a computer conference for canadian data libraries and archives," report to the social sciences and humanities research council of canada, p. 24ft. ^ terminal is used here to refer to the same range of i/o devices that the term workstation has come to denote, which covers everything frpm teletypes to visual display units to microcomputers. however, since items in the questionnaire refened to terminals, we will continue to use this term below. ' ninety-one questionnaires were collected from the 125 representatives at the meeting, providing a 73% response rate. two factors, however, must be considered when generalizing from this poll. first, 30% of the participants were substitutes for official representatives. thus, generalized comments from this sample encompass more than just or's. secondly, only 46% of the icpsr membership were in attendance. spring 1986 iassist quarterly 45 or's workplace; while the second sought some indication of the scope of experience that representatives have had with computer commimications. this paper reviews specific characteristics of the ors and examines how prepared this group is to avoid or overcome the impediments to active computer-based communications. 70% indicated that they make use of a terminal throughout their workday. important characteristics for computer communications the availability and use of computer terminals lack of convenient access to a terminal does not appear to be a major problem for the vast majority of icpsr representatives. nearly 80% of the respondents have immediate access to a terminal at work (see table 1), and 48% have terminals both at home and work. only 14% do not have any convenient access, and of these thirteen respondents, eleven have a terminal available to them either on the same floor as their office or on another floor in the same building. having a terminal at your fingertips does not necessarily ensure use of the device. however, as shown in table 2, a clear relationship between access and use does exist in this data. those with a terminal immediately available to them during the workday report the highest usage rates. examining the breakdown across the categories of access, the proportion of those using a terminal several times a day declines monotonically as one moves from those with the highest degree of immediate access to those with no terminal directly available. the obvious conclusion is that most respondents make use of the equipment that they have. however, this is not necessarily the most significant conclusion. more important is the summation that a computer terminal is an integral tool in the work routine of a large majority of the icpsr representatives. over a few special features are desirable for the effective use of computer mail or conferencing systems. one feature is the capability of placing a call and making a connection with a central computer system and its mail or conferencing software. this type of terminal connection usually is supported by a modem attached to a standard telephone outlel such a configuration permits a user to call either their local computer system or a packet-switching network through which a myriad of computer systems are available. some terminals, however, are directly wired to a central computer. in such instances, the use of packet-switching networks is dependent upon a call-out facility on the mainframe. regardless of whether the terminal connection is through a modem or a mainframe call-out facility, the most flexible sittiation for the user is to be able to logon to the computer system housing the mail or conferencing software. of tjie survey respondents, 44% have a terminal with dial-out capabilities at work (see table 3). when those who have a terminal at home only are included in the group with dial-out capability, the overall percentage increases to 54%. furthermore, a call-out facility was present on the central computer systems of over 70% of the respondents. these figures reveal that a majority of the respondents have available some form of call-out facility which would permit them to cormect to the consortium networl another characteristic which encourages the use of computer communications is the availability spring 1986 46 iassist quarterly of a full-saeen editor. the backbone of computer communications is the typed word, and the ease with which text can be entered and modified significantly influences the amount of text contributed. just as was the case with dial-out facilities, respondents seem to have ready access to ftill-screen editors whether at home or work (see table 4). eighty-eight percent of those with terminals at work have such an editor; 82% with terminals at home also have one available. experiences with computer conununications two-thirds of the respondents reported that they had made use of at least one of the three communication methods — electronic mail, computer conferencing and networks (see table 5). nearly half (47%) had experience with more than one of these methods. in fact, those saying that they had used both networks and electronic mail constituted the largest single group (27%).* considering the three electronic media separately, 58% of the respondents noted some experience with electronic mail; 57% had used a network; only 20% had tried a computer conference. in terms of overall exposure, one in five indicated experience with all three methods. experience with these communication methods clearly varied by type of terminal access. eighty-one percent of those who have a terminal botii at work and home have had ' respondents may have been confused about the difterence between a carrier network such as telenet and an application network such as bitnel the former is a service which allows one to dial a local telephone number and to connect as a remote terminal to a computer system, while the latter type of networx refers to special application software making use of packet-switching technology to transmit mformation between sites. the item in the questionnaire was suppose to identify those who had experience with a carrier network. experience with at least one of the three methods (see table 6). antithetically, 70% of those without immediate access to a terminal indicated that they had no experience with any of the three commimication methods. the difference between these two groups accenttiates the gap that exists between those who have a terminal at their fingertips and those who do not in comparing the remaining two groups, the percentage of those having worked with at least one of the communication methods was virtually the same, 68% for those with a terminal at work only and 67% for those with a terminal at home only. the experience levels of these two groups are much closer to the group with terminals at both work and home. one interesting difterence is that a higher proportion of those with only a terminal at home had tried two or more of the communication methods. an indication of the extent to which these three types of communication have been incorporated into the work routines of the respondents is shown in table 7. nearly one-third of those making use of electronic mail check it on a daily basis and over 60% access it twice a week or more. similarly, 22% of those belonging to a computer conference use that medium daily, while only 10% of network users use the network that frequently. electronic mail is cleariy leading the way among these three methods of communication; and with the introduction of electronic mail service between universities, daily use of electronic mail will undoubtedly increase. its popularity is exemplified by the fact that 63% of those who use electronic mail daily also reported using bitnet, which is an inter-university electronic mail service. spring 1986 [assist quarterly 47 conclusion given both the availability of terminals to icpsr representatives and their experiences with computer communications, what are the implications for the consortium data network (cdnet)? the profiles described above point to a couple of possibilities. first, slightly more than half the respondents possess the proper mix of both equipment and experience, thus making the likelihood that they will use cdnet very high. fifty-two percent of the respondents reported immediate access to a terminal and indicated experience using a network. this is a significant group, since access to cdnet depends upon a remote terminal connection through a packet-switching network such as telenel secondly, an additional 12% have both a terminal available and some experience with electronic mail or computer conferencing. assimiing that some experience with either of these media develops skills that are easily transferable to the use of networks, this group should also readily use cdnet thus, 64% of the respondents appear to possess essential equipment and skills to use cdnet without major obstacles. an additional 21% of the respondents have ready access to terminals but no experience with the three methods of computer commimication. consequently, this group faces the task of learning some new computing skills. an important factor in this regard will be motivation. motivating people to use any of the three commimication media, even when they already possess the necessary skills, is in itself a challenge. thus, initiating a service such as cdnet is further complicated by the need to motivate first time users to acquire the additional skills. no data was collected in this survey to indicate directly how significant a factor motivation will be. factors other than the computing skills and motivation levels of or's will also influence the future use of cdnet certain environmental factors, such as past demands for icpsr services and the vitality of the member imiversity's research conmiunity, will contribute to usage patterns. these factors have not been examined here. rather, attention has been focused on a few known obstacles to the use of computer communications. as cdnet swings into production, the importance of these and other factors should become evident [ editor's note: the following is reprinted from icpsr's guide to resources and services 19851986, p24. "testing of a new remote service. consortium data network (cdnet), is currently underway and should be available in the fall. this new service is aimed initially at icpsr official representatives. cdnet will provide access to an on-line searchable version of the holdings section of the guide, an on-line data and codebook ordering facility, an interactive message and conferencing facility as well as access to statistical software for analysis of icpsr holdings. a data base containing information about each item in a large subset of the studies available through the icpsr is also being produced for inclusion in cdnet connection to cdnet will be available through the autonel, telenet and tymnet public data networks."]" spring 1986 48 iassist quarterly table 1 table 2 access to a computing terminal' terminal at work home total terminal at home have access don't have n % n % n % have access don't have 43 48% 25 28 9 10% 13 14 52 58% 38 42 work total 68 76% 22 24% 90' 100% percentages are based on the total number of respondents in the overall table. one questionnaire was excluded from analysis since information was provided for only one of the items. frequency of terminal use by location i. access to terminal location of immediate terminal access frequency of use' access at work and home access at work only access at home only no immediate access totals n % n % n % n % n % [1] [2] [3] [4] [5] [6] 36 84% 3 6 2 5 2 5 18 78% 2 9 3 13 6 67% 3 33 1 8% 4 34 2 16 4 34 1 8 61 71% 8 9 9 10 2 2 6 7 1 1 total n.a. 43 100% 23 100% 2 9 100% 12 100% 1 87 100% 3 l=several times a day 2=0nce a day 3=couple of times a week 4=0nce a week 5=couple of times a month 6 = 1 nf requent ly spring 1986 iassist quarterly 49 table 3 table 4 type of access to dial-out facilities is dial-out available through ... both terminal at work & via mainframe computer answer a terminal at work a mainframe computer n % n % n % yes no n.a. 40 44% 25 28 25 28 64 71% 15 17 11 12 29 32% 29 32 32 36 availability of a full screen editor does your terminal allow full screen editing? answer terminal at wor k terminal at home n % n % yes no n.a. 58 8 24 88% 12 41 9 40 82% 18 table 5 use of electronic mail, computer conferences, and networks experience with ... n % only networks only e-mail' networks & cc ' networks & e-mail networks, cc & e-mail none 9 1 1 1 24 17 28 10% 12 1 27 19 31 total 90 100% ' e-mai l=electronic mail cc=computer conference spring 1986 50 iassist quarterly table 6 computer communication experience by access to a terminal location of immediate terminal access type of access ' access at work & home access at work only access at home only no immediate access total n % n % n % n % n % [1] (2) [3] [4] [5] [6] 6 14% 7 16 12 28 10 23 8 19 1 4% 4 16 1 4 6 24 5 20 8 32 4 45% 2 22 3 33 2 15% 2 15 9 70 9 10% 11 12 1 i 24 27 17 19 28 31 total 43 100% 25 100% 9 100% 13 100% 90 100% l=network only 2=e-mail only 3=network s, computer conference 4=network (. e-mail 5=network, e-mail & cc 6=none table 7 how often computer communications are used use rate electronic mail computer conferences networks n % n % n % daily twice weekly once a week twice monthly i nf requently 16 32% 15 30 6 12 6 12 7 14 4 22% 2 1 1 3 17 2 1 1 7 39 5 1 1 8 7 16 10% 23 16 14 37 total 50 100% 18 100% 49 100% spring 1986 14 iassist quarterly :,«s«sss9!es<*i«1»-«s»»!«««' prospects and problems in the use of government generated data in india by n.k. nijhawan' director, data archives, indian council of social science research (icssr), new delhi this paper atlempls to discuss the nature and extent of governmcni effort, in india, in generating various kinds of socio-economic data and also to examine the problems faced by researchers in having access to such data. india is one of the few countries in the world which has such an elabourate, comprehensive and multilevel data-base. but it must be admitted that very few of these data are put to effective use. the government departments and its other organizations have a limited purpose and accordingly analyse the data only marginally. the researchers in the institutions and university departments who have interest, necessary skills and competence to use data in greater depth and dimensions do not have access to these data. thus despite a 'data revolution' in the country, both policy makers and researchers are unable to make effective use of a valuable data base for various reasons. an attempt has been made to analyse some of the problems underlying this process and also to explore possibilities that exist for remedying the situation. historical background the present statistical system in the country has evolved through centuries. it dates back at least to the sixteenth century, when the mughal rulers had data collected and tabulated systematically for about four thousand geographical units of the mughal empire. a comparable system for that century may not be found in any other pan of the world. the nature of tabulated data also indicates that the mughals may have borrowed some of its important features from the system prevalent in an earlier period. it seems even the british rulers, in india, leaned heavily on the data system developed by the mughals during the early period of their rule. the british, however, transformed, and in many ways enriched, the system as is evident from the fact that the first gazetteer was published in 1866, the statistical abstract of british india in 1868, and the first all-india population census conducted in 1881. in 1895, a statistical bureau was established at the centre to coordinate agricultural and foreign trade statistics and to collect data on prices, wages and industry. a number of committees and commissions were set up from time to time to examine in detail the available statistical material with a view to improving its quality. 'presented at lasslst/ifdo international conference may 1985, amsterdam. the views expressed in this paper are those of the author and not those of the icssr. fall/ winter j985 iassisi quarterly 15 post independence developments a real breakthrough in the official sialistical system, however, came only after independence. a nucleic statistical unit called central statistical organizaton (cso) was set up in 1951. subsequently, state sutistical bureaux (ssb's) were set up in almost all the states. following this statistical offices were set up, each in the charge of a district statistical officer. like the cso at the centre, the ssb's are entrusted with the task of coordination and publication of statistics collected by different government departments and district statistical offices. at the instigation of, then. prime minister jawaharlal nehru, a national sample survey (nss) was started in 1950 by the government to expeditiously collect data relating to all aspects of the national economy, on a continuing basis, using sampling methods. beginning with the first survey (called a round) in 1950-51. the national sample survey organization (nsso) is about to complete data collection for the 40th round by july, 1985. currenuy, nsso has a large staff its field staff alone consists of about 5000 full time research investigators and field supervisors engaged in data collection. the need for collection of relevant data for the purposes of national five-year plans, as well as for meeting the demands of various international organizations, has been increasingly felt since independence. a large body of these data are collected by the government in addition to those that are usually collected for its routine and normal functions. it is appropriate to say that the country is witnessing data explosion since 1951. a brief mention of the nature and extent of socioeconomic data generated by the government during the past three decades or so is in order. nature of data four decennial all india population censuses have been completed. basic population characteristics for all the districts (approximately 365) are available in published form for all four censuses. another extremely important tabulation of the indian census which goes right up to the village and urban blocks called the primary census abstract (pca) are available for each village/town in the district census handbooks. the census data also serve as the major source of basic economic information about the populaton. besides data on population composition and distribution, data on fertility and mortality are collected under the vital regisuation system. considering the significance of agriculture in indian life, information on agricultural holdings, their characteristics such as: size, form of tenancy, land utilizaton, production of crops, irrigation, livestock, prices, wages, etc., are being collected systematically by various government agencies. data pertaining to land utilization, areas under various crops and area irrigated are almost continuously available in published form since 1884. apart from the data collected from village records, nsso has also collected data on land holdings, land use, agricultural production, cost of cultivation, etc.. in its various rounds. besides, the nsso carried out a special survey on land holdings, at the instigation of fao, in its 16lh and 17th rounds to meet the requirements for the 1960 world agricultural census. unlike agricultural statistics, data on indian indusu7 is of relatively recent origin. before independence, industrial statistics in the country, were extremely inadequate and of poor quality. after independence, legislation had to be passed to facilitate collection of statistical data relating to industries. a breakthrough came with the passing of the collection of statistics act, 1953 fall/winter 1985 16 iassist quarterly and the colleclion of slalislics (central) rules. 1959. an annual survey of industries (asi) has been carried out since 1959. covering all the registered and licensed factories in the country. a number of government agencies periodically collect data on employment, output, earnings, etc., related to small scale and cottage industries. at present, the data relating to the industrial sector are available from at least twenty different sources of the govemmenl the national sample surveys are a rich treasure of data encompassing a large spectrum of socio-economic and cultural life of indian society in the post-independence era. the nss are multipurpose surveys covering several socio-economic and other subjects simultaneously. comprehensive data have been collected, in various rounds, on different aspects of agriculture including land utilization, cost of cultivation, rural indebtedness, rural labour etc. likewise, a great deal of eftorl has gone into data collection in various surveys on consumer expenditure, demographic characteristics of sampled households, internal migration, employment and unemployment, housing conditions, trade, transport and rural electrification. the quantity of data being collected through these surveys is quite voluminous. just to give an idea, it may be mentioned that the sample size varies from round to round, but normally a survey covers about 5000 urban blocks including 200,000 sampled households. sometime back, at the initiative of the indian council of social science research, an estimate was made of the number of schedules/pages canvassed in various nss rounds. it was estimated that in the first nineteen rounds alone (1950-51 to 1964-65) over two million schedules, comprising over nine million pages, had been canvassed. the large quantity of information collected during the course of all these years can thus be imagined. lack of perspective the above account, which is purely illustrative in nature, clearly reveals the nature and extent of data being collected by the government. by any standards, india can claim, to be one of the richest countries in the worjd in terms of data resources. but at the same time it cannot be denied that both policy makers and researchers are unable to make eftective use of this resource, as a large part of the available data is put to a very limited use. the obvious question to ask is: why is that so? the single most important factor responsible for this sorry state of affairs is the government's lack of perspective in this regard. it may be worth noting that there is a disproportionate emphasis on data collection as opposed to data utilizaton. emphasis continues to be on data collection rather than on its utilization and improvement of its quality. the main problem, therefore, is that the effort in data collection has multiplied manifold in the past 35 years, but the techniques of handling data continue to be quite rudimentary. most of the data collected by the government is tabulated manually even today. this results in long delays, even in bringing out limited tabulated data. for instance, the results of the 31st nss round on rural electrification, conducted in 1976-77. were published as late as january, 1984. likewise, although the summary results have been released, the detailed results of the 1978-79 annual survey of industries are still being processed by the cso. the last detailed results of asi available for public use relate to the year 1973-74. nothwithstanding unusually long delays in the publication of government' produced data, scholars can make little use of them owing to high levels of aggregation in the published data. most government agencies collect data at a micro level, such as, village, household or industrial unit, but these are not made available to researchers. the analysis of micro data at fall/winter 1985 lassisl quarterly 17 the village level would have been immensely useful to the researchers, but such information is not available to them. only all-india and state-level results are normally published by the government lack of uniformity multiplicity of government agencies in the collection of similar data also adds to the diftlculties of the researchers. for instance, as pointed out above, data relating to indian industrial sector are available from at least twenty different sources. these alternate sources of data differ in their conceptualization, coverage and measurement procedures. for example, the armtial survey of industries (asi) gives the gross output of industry in value terms where as the the index of industrial production (iip) gives the index in terms of quantities. similarly, the asi covers all companies registered under the factories act, where as the iip covers units registered with institutions, such as the director general of technical development, coal commissioner and textile commissioner. consequently one may arrive at different answers to the same question. for instance, to the question how much did indian industry grow in a particular year, the estimates vary depending upon the source used for estimation. in one such exercise, the divergence among the estimates of growth rates using different sources of data for the year 1978-79 was found to vary between 7.2 percent and 14.1 percent, at the extremes. non accessibility of unpublished data it may be noted further that researchers usually do not have easy access to the unpublished data. procedural complexities and red tape more often discourage them from making use of such data. worst yet, potential users do not even have any means of knowing about the available data. if, through personal contacts or by sheer coincidence, they come to know about some unpublished data, it is extremely difficult for them to get access to such data. the dictum of confidentiality and privacy in respect of government held data is so rigidly adhered to by various government agencies, that even interdepartmental exchange of data within the government is difticull the problem becomes all the more serious when the agency collects data under the provisions of the collection of statistics act of 1953. in denying access to unpublished data they invariably invoke the sanctity of rule 7(1) of the act which prohibits giving information on individual units. the said rule reads: "no information, no individual return and no part of an individual return with respect to any particular industrial or commercial concern given for the purpose of this act shall, without previous consent in writing of the owner for the time being of the industrial or commercial concern in relation to which the information or return was given or made or his authorized agent, be published in such a manner as would enable any particulars to be identified as referring to a particular concern." it can be seen that the act does not explicitly restrict the release of raw data to researchers. government agencies, in fact, use the act in self defense and to conceal inefficiency in handling data. it may also be appropriate to mention here that while the act can be invoked by the competent authority to pimish or impose a fine if any person "willfully refuses or fall/ ivinter j 985 18 iassist quarterly without lawful excuse neglects to furnish such information... as may be required under the act", the act is silent about the obligation of the 'statistics authority' in respect of utilization of the collected information. consequently, large volumes of collected information continue to gather dust while genuine researchers are denied access to the required data on some pretext or other. prospects a large majority of researchers are working in the universities and research institutions in various disciplines. the universities and deemed universities (institutions which have been given the status of a university), numbering over 150, are financially supported by the government via the university grants commission (ugc). the research institutions are financed or are in some way supported by national level organizations in various fields like agriculture, medicine, social sciences, education, science and technology and culture, etc. these national level organizations are, to name just a few: the indian council of agricultural research, indian council of medical research, indian council of social science research, indian council for scientific and industrial research, national council of educational research and training, etc. all such national level organizations as well as the ugc are financed by the government as already pointed out, researchers based in these organizations, or institutions supported by them, have quite often little or no access to unpublished government generated data. the published data, on the other hand, serve only a very limited purpose for the scholar. consequently, researchers have no option but to generate their own data. for want of adequate funds, most of the scholars both in the universities and research institutions end up taking only micro level studies. it is ironic that at least some of the scholars in the universities and research institutions, keen, and qualified to handle large data for macro analysis, must content themselves with only micro studies for want of financial resources. on the other hand vast body of data generated by the government remain largely unutilized. thus, research activity is seriously hampered not for want of resources, initiative, willingness, skills or expertise in the country, but for lack of perspective and policy for proper storage and sharing of data generated by different wings of the government thus what is needed is not additional output of financial resources for the collection of new data, but a perspective and policy for utilizing data collected by various government organizations. it is necessary to improve communication channels between the producers and users of data so that interested scholars can know what type of data is available, where, and in what form. various national level organizations may help to bridge this communication gap. besides, these organizations may try to persuade government agencies to release raw data to genuine researchers. computerization of govenment data may also greatly facilitate utilization of the existing data. in the absence of a computerized data base, it is extremely difficult for researchers to utilize available data even when these are released by the government the volume of data is so large, and the structure so complex, that individual scholars with limited funds at their disposal, would find it difficult to put them in proper shape for meaningful analysis. it should be mentioned that the indian council of social science research is making efforts to persuade the csc, nsso and other major data producing agencies to, at least release 'non-sensitive' and unclassified data to interested scholars. the icssr is also preparing an inventory of data available in various government agencies and research institutions which will, hopefully, go a long way towards fall/winter 1985 iassisl quarterly 19 disseminating information to scholars about available data. the icssr has also recently financed a project in collabouration with the bureau of industrial costs and prices, the planning commission and the industrial development bank, to develop an information system on industrial data. the system would be developed by pooling industrial data collected by various government agencies and standardizing them as much as possible. currently, the government is following a more liberal policy in the use of computers in various government activities. a number of government agencies have already set up computers and transformed their databases in machine-readable form. the national informatics centre, set up by the department of electronics, is also helping various government agencies to computerize their databases. with the help of • the national informatics centre, separate information systems are being developed, containing data on education and manpower. energy, health and demography, industrial resources, agriculture, transport and industry. the icssr is part of the nic network and. therefore, will be in a position to help scholars get access to these information systems, when developed. thus the potential of utilizing government generated data in india is enormous, provided proper perspective is developed and policy formulated. hopefully, given recent developments in government, particularly, in the computerization of government data, it may be possible for researchers to access primary data. proper utilization of existing data may in turn help government agencies to cut the cost of couecton of unnecessary and inelevant data. this will, however, require greater competence on the part of the researchers in handling complex data. recent developments do seem to hold promise but time alone will bear testimony to this." fail/winter 1985 1/14 moore, jennifer and kettler, hannah scates (2018) who cares about 3d preservation?, iassist quarterly 42 (1), pp. 1-14. doi https://doi.org/10.29173/iq20 who cares about 3d preservation? jennifer moore1, hannah scates kettler2 abstract the preservation of 3d research data is a present and emerging need. an increasing number of researchers are generating, capturing, and/or analyzing 3d data, but they rarely focus on preservation or reuse of that data. this paper and presentation describe models of 3d data creation and use, outline the specific concerns for this data type, unpack the complexities and challenges of preserving it, and examine existing 3d data preservation resources while working through local case studies from the field of anthropology. directions on how to move digital 3d data preservation forward will be discussed. keywords 3d digital data preservation; anthropology; archaeology, digital humanities introduction anthropologists are concerned with both physical and digital 3d data. in some cases, they are integrating digital 3d capture and/or creation into workflows as a preservation mechanism for artifacts, faunal remains, casts, sites, etc. however, as with other digital data types, creating a digital record of an object or place does not provide stable preservation for said artifact unless the digital data are also treated for preservation. digital preservation of data is widely accepted as necessary to protect data against loss and obsolescence, particularly ‘where the data are non-reproducible or extremely valuable'. (dcc, n.d.) digital preservation ensures that data are in an appropriate state for long-term access and reuse. inserting preservation actions and behaviors into any type of research data workflow can be tedious, as the process requires forethought and is best accomplished when built into the methodology of a project. in actuality, preservation often occurs as an afterthought, at the end of a project, at which point the preservation of the data becomes more difficult if the data have not been appropriately prepared and administered throughout the project lifecycle. digital preservation principles can be applied more easily to some types of data than others. 3d data fall within the latter category and require specific actions that may not fit typical curation treatments. digital 3d data preservation is a burgeoning topic among practitioners and data curators alike. as with all data, digital 3d data preservation actions would be most successful if workflows were established using best practices and standards at the outset, but this has proven to be much easier said than done. creating and capturing digital 3d data requires intense research and skill development in order to accomplish immediate goals, so the added complication of long-term data stability is typically not at the forefront of methodology considerations. while there is no shortage of literature describing capture and creation methods, the lack of consensus on how to do digital 3d preservation for the longterm is a major barrier. the best practices that exist are not yet expansive enough for adoption. research shows that institutions that are implementing any kind of preservation actions are often doing so ad hoc. the authors of this paper recently conducted a survey3 of practitioners and curators, which indicated there is a need and a great deal of support for collaboratively developed best practices and standards. https://doi.org/10.29173/iq20 2/14 moore, jennifer and kettler, hannah scates (2018) who cares about 3d preservation?, iassist quarterly 42 (1), pp. 1-14. doi https://doi.org/10.29173/iq20 literature review as early as 2009, julie doyle wrote about the need for digital 3d preservation standards, asserting that establishing principles of authenticity and reuse is of the utmost importance. she estimated that the most important step to be taken to ensure long-term preservation is setting up workflows for emulation and metadata creation (doyle, 2009). more recently, the archaeological data service (ads)4 produced a case study documented in curating research data, a handbook of current practice (johnston, 2016), which briefly outlines some 3d data curation concerns. the case study suggests that preservation of reusable 3d objects is challenging because projects are often focused on an end product. the ads & digital antiquity, the london charter, and a group called 3d-icons5 have put forth some limited recommendations that offer a foundation but are also purposefully vague or incomplete, making the preservation landscape difficult to navigate. the london charter6 was developed in 2006 ‘as a means of ensuring the methodological rigor of computer-based visualization as a means of researching and communicating cultural heritage. also sought was a means of achieving widespread recognition for this method’ (london charter, n.d.). it outlines six basic objectives, which are to create standards that ensure methods are rigorous, intelligible, useful, sustainable, and extensible by the community of practice. the objectives are manifested in the practice of six basic principles. of the principles, 1) implementation 2) aims and methods 3) research sources 4) documentation 5) sustainability and 6) access, aims and methods as well as documentation and sustainability are particularly relevant to data preservation. while these principles provide a valuable framework for considering necessary workflows at a high level, no suggestion of specific recommendations are made. the ads and digital antiquity7 collaborated to create the guides to good practice to ‘ensure digital data access and long-term preservation (guides to good practice, n.d.). the guides cover a multitude of methods and data types in use by archaeologists, including close-range photogrammetry and 3d laser scanning. the guides also offer information related to formats, approaches to documentation, and suggest minimum, file-level metadata elements and workflows, but they remain short of standards and are far from exhaustive. the mission of 3d icons is to create ‘highly accurate’ 3d models of important heritage sites in europe. 3d-icons produced a report that focused on adapting the carare schema8 for 3d models. the aim is to assure quality through providing provenance, transformation, and paradata description. the carare schema was developed for describing digital items for cultural heritage organizations in europe (d’andrea & fernie, 2013). the work done by 3d-icons, though also narrow in its focus on provenance and paradata, could be useful in building more widely accepted digital 3d data metadata standards. methods of 3d data creation and use there are four ways in which digital 3d data are created: through free-form, physical, real-world measurements (i.e. with a tape measure); algorithmically, as with the 3d scanning methods discussed later; observationally, and by using treatises or historical research. many methodologies use a mixture of these creation methods. data created through the free-form method can be an expression or visualization aid for concepts, or a representation of real-world space. free-form 3d modeling refers to the lack of any material that https://doi.org/10.29173/iq20 3/14 moore, jennifer and kettler, hannah scates (2018) who cares about 3d preservation?, iassist quarterly 42 (1), pp. 1-14. doi https://doi.org/10.29173/iq20 represents the object or space being modeled. much of this type of work is being done in virtual reality spheres. this process allows for some artistic license in 3d representation and some fluidity in regard to the requirements for 3d preservation. augmented asbury park9 is a representation of lost historical space. the project represents the heyday of the asbury park boardwalk and places the missing buildings and attractions in their physical space using augmented reality. a guided tour is available in the form of a physical booklet one can carry around the park, using coded images to bring in 3d models of some of the boardwalk attractions by means of an augmented reality application. 3d models were created to represent the buildings, which in many instances were created free form and without digital guides. the models are not 3d scans but are representations and interpretations of space based on photography of various similar objects and the allowances of the physical space at asbury park. in this instance, there is much more flexibility in terms of what is necessary to preserve. it may be more of a prerogative for the data preservation specialist and researcher to preserve the raw data, the experience of the data, or the presentation of the data. continuing with the example of augmented asbury park, it may make more sense to preserve the physical booklet and virtual 3d object relationship which would require that the 3d data be preserved in a way that can allow for display on a new platform or emulation in conjunction with the physical booklet artifact. as with many other projects, and especially digital humanities projects, the physical and the digital artifacts are inseparable and should maintain their relationship to support the research being conveyed. 3d models based on the physical process of measurement are, for the purposes of this article, distinct from algorithmically generated 3d models because the act of measurement is not a proprietary process and can easily be reproduced by anyone with a ruler. 3d data creation based on physical measurements requires one to go out into the field or handle the specific object and take detailed measurements that can be visualized using 3d modeling software. the technique requires extreme attention to detail as well as the ability to maintain accurate notes that would represent the analog of the 3d model. in a lot of these cases, the 3d model is the scientific redundancy (which has heretofore been missing) in fields like archaeology. a 3d model based on physical measurements is a much more transparent form of 3d data creation. the use of 3d scanning and geographical information systems (gis) in archaeology has expanded the possibilities of 3d modeling over much larger areas but is much more dependent on computationally generated measurement. the automation of 3d data creation places the processes of measurement in a ‘black box’ and makes the processes proprietary and obscure. algorithmically generated 3d data, like the 3d scanning and gis methods mentioned above, are much harder to evaluate, reproduce, and decouple from the technology used to create them. 3d scans measure distance, angle, reflectance, and color, depending on the type of scanning method one uses, such as computerized tomography (ct) scanning, laser, structured light scanning, and photogrammetry, to name a few. the processes used to evaluate said measurements are not open sourced, so that a researcher hoping to reproduce the data creation process can evaluate the integrity of the data. at first, this may not seem to be an issue for preservation, but upon further reflection it is integral to any 3d preservation process. when data need to be reproduced or migrated for preservation purposes, there is no way to evaluate whether the reproduction is faithful or the migration successful without a clear understanding of how the data were created in the first place. since much of today’s 3d data streams from scanners is entirely proprietary, the likelihood of successful reproducibility of 3d scanning data is low to non-existent. take, for example, the laser-scanned 3d collection of artifacts10. these data represent many hours of labor and are specimens used for anthropological research. these scans are, in many ways, digital surrogates for the physical specimens. lasers accurately measured and reproduced the objects at a level of faithfulness researchers were satisfied with, but the process that occurs during scanning and computational interpretation is entirely proprietary. this process has its own uniquely coded https://doi.org/10.29173/iq20 4/14 moore, jennifer and kettler, hannah scates (2018) who cares about 3d preservation?, iassist quarterly 42 (1), pp. 1-14. doi https://doi.org/10.29173/iq20 methodology and software, which generates the end product that is available on the web. in several years, these data will need migration to new servers and a new platform and will need new standards of preservation. although 3d scanning methods are popular, they produce a singular representation of raw 3d data that makes articulating the importance of data preservation much more difficult because of its opaqueness and unverified processes and methodologies. to mediate this, one could argue for the usage of open-sourced 3d scanning technologies to facilitate the preservation of their outputs and bolster the integrity of the generated 3d models as reproducible, or at the very least, transparent research objects. not all 3d data are based entirely on the presence of physical artifacts. a 3d data creator could rely on traditional research methods to determine how to represent artifacts in virtual space. this method of 3d data creation has much in common with free-form modeling. 3d data may be based on historical research and interpretation like any other work one would publish. the 3d data are imbued with thoughts, decisions, assumptions, and technological or expertise limitations, and may represent a range of artifact types, including imaginary, lost, or degraded objects. the 3d model represents a scholarly argument in much the same way as a scholarly monograph. the resulting 3d data therefore requires the same level of preservation as an article or book, with the same attention paid to reusability and shelf life. the long-running digital hardian’s villa project11 created 3d data that represents not only the empirical archaeological data, but also contextual research about the villa. the researchers have used a combination of free-form, physical measuring and scanning to extrapolate hardian’s villa footprint into a fully three-dimensional space. a website has been built around the research involved in the creation of the 3d data and houses not only the models, but also the bibliography and citations necessary to its creation, the paradata describing the decision processes, the photos on which the 3d model is based, the videos of the space, and interviews with specialists discussing what they know about the individual spaces. all of this contextualized research is underpinned by the process of 3d data creation as much as the contextualizing research bolsters the resulting 3d model. the 3d data are a part of a feedback loop with all other data types in this digital monograph and should be rendered as stable as other scholarly outputs. to complicate 3d data creation, much 3d work is done using a combination of the methods mentioned thus far. 3d data represents the gambit of individualized, single processing work (as with 3d scanning) but may be as complex as a 3d thesis or publication that combines 3d scanning, free-form modeling, data visualization, and traditional research methodologies. unpacking digital 3d data capture data can be captured to create 3d models in various ways, including close-range photogrammetry, computed tomography (ct) scanning, structured light scanning, laser (time of flight, phase shift, triangulation) scanning, etc. to describe all of the methods of collection is beyond the scope of this paper, which focuses on some of the methods commonly employed by anthropologists. the methods include close-range photogrammetry, structured light scanning, and triangulated laser scanning. close-range photogrammetry is a process by which a camera is calibrated to capture multiple images of an object. many types of cameras can be used to accomplish this task, but higher-quality images enable more reliable 3d reproductions. calibration is specific to the lens and camera used; this calibration accounts for any distortion by the lens, sensor, or processor. an external control may be added to define the data and/or geometric constraints (guides to good practice, n.d.). the images are stitched together and processed by software that will then allow for a number of outputs, including a point cloud, 3d polylines, mesh, or raster graphics. https://doi.org/10.29173/iq20 5/14 moore, jennifer and kettler, hannah scates (2018) who cares about 3d preservation?, iassist quarterly 42 (1), pp. 1-14. doi https://doi.org/10.29173/iq20 a point cloud can be converted to rcs/rcp files and attached to a cad drawing to create a 3d cad model, or its output(s) can be manipulated in other 3d software such as meshlab12 (martorelli, 2014). an example of this method is illustrated by the work of martorelli, pensa, and speranza in their paper 'digital photogrammetry for documentation of maritime heritage'. they presented three case studies that use photogrammetry to analyze historical boats: tomahawk, refola, and nada. the case study of the tomahawk described the methods, which included calibration, marker placement, image capture, processing, point cloud generation, and 3d model analysis. the authors outlined the process for creating the 3d model as a cycle of 'correction, surface layout, sectioning, further correction….' (martorelli, 2014). if information on calibration and processing was recorded, it was not offered in this article. it should be noted that 3d digital data preservation was not the purpose of this research. by the authors’ own estimation, the method of photogrammetry did not result in the greatest accuracy, but they determined it was a good fit because it was a 'quick and inexpensive' way to acquire the data they wanted. this example inspires many questions. without detailed documentation of the methods, are these data reusable? with iterative processing of the model, at what stage should it be preserved? are these data suitable for preservation, given that they are probably not accurate? if so, in which format(s) should they be preserved? would they meet selection criteria for preservation, assuming that such a criteria existed? triangular laser scanning is ideal for close-range scanning. as the name suggests, this type of scanning uses triangulation to collect precise point locations by projecting a laser line onto an object, which bounces back to the sensor. triangular laser scanning is also dependent on accurate calibration; some scanners can collect highly accurate data, and some have color capture capabilities. molloy et. al (2016) describe methods of using triangular laser scanning to investigate the edge wear on prehistoric tools to understand their function. the authors describe establishing resolution by setting the step rate of the laser, which determines the point-to-point distance. denser point clouds produce higher resolution. the project is said to have used the highest resolution settings, but specifics are not listed. eight to 12 static positions were rotated to capture the entirety of each object, and the software stitched the various scans together. it’s noted that once the scan finished producing a point cloud, the superfluous data were removed, which the software documented. molloy et. al were satisfied with the results of the laser scanning and commented on the portability and efficiency of the process. this research project also employed photogrammetry by capturing raw images at high resolution, processing, and post-processing to remove noise. by the authors' assessment, the laser scanner exceeded in capturing small details but was weaker in capturing sharp changes of direction, reflective surfaces, and damage to the edges of objects. in one case, the scanner was unable to capture the edge data, and, while either side of the blade was captured, there was a data gap that could only be completed by large-scale hole filling, which leads to inaccurate data. each method the authors employed had weaknesses, but it's noted that they intended to combine methods for a more complete model. questions of what to preserve are again raised by this example, due to the iterative workflow, the problem of processing, corrective measures (like hole filling), and the combining of data from disparate methods to make a complete model. this project seems to have recorded some essential information for the creation of useful metadata, although it was not offered as a part of the paper. structured light scanning is accomplished by projecting white light onto a surface to display a series of organized patterns (wachowiak, 2009). cameras record the changes to display a surface, and the software algorithmically calculates the distances. structured light at close range is good for capturing a small point-to-point distance or resolution. calibration of the light and camera is essential for accurate 3d reproduction. because the scanner only captures what it can see, a series of scans must https://doi.org/10.29173/iq20 6/14 moore, jennifer and kettler, hannah scates (2018) who cares about 3d preservation?, iassist quarterly 42 (1), pp. 1-14. doi https://doi.org/10.29173/iq20 be completed, often using a rotary table (for small objects) and changing the angle of the object to cover all of the area. this means that overlapping data are collected, and that the camera often collects noise from background interference, etc. laura niven (2009) used structured light scanning to create a reference collection of faunal skeletons for zoo archaeologists. researchers in this example spent four years developing a workflow, which is well documented in the article, with details such as specimen size ranges and lens choice. the structured light scanner they used allowed for capture with or without color. their scanning took place in a photographic light tent on a black background, which reduced shadow and noise and made processing more efficient. the specimen was scanned in 3-4 rotation steps. the software stitched the scans together after two steps, and alignment was automatically added throughout the process after the first two steps. following the first set of rotations, the specimen was repositioned, and the new set was scanned and aligned. once all of the scans were completed, the data were merged into a final scan, with overlap automatically adjusted. the outputs of this project were .ply (polygon file format, stores multiple properties) files for 'archival purposes', as well as .stl (stereolithography, basic surface description), and .obj (simple geometry), which were used to create 3d pdfs as well as some 2d image formats. this example was very detailed in local documentation and took preservation into account, but the problem of processing and the subsequent questions of when and what needs to be preserved still exists (what are the ‘raw’ data?), even in this careful example. local case studies fluxus digital collection the fluxus digital collection was a digital humanities collaboration between a faculty member, technologists in information technology, and the digital humanities librarian at the university of iowa. the intention was to create an online contextualizing exhibit of the fluxus west art collection housed in the university of iowa special collections. grappling with the library’s charge to preserve, the scholar’s prerogative to study, and the artists’ intent to make art interactive and accessible, the collaborators decided to reconcile these issues and create 3d models of select objects as a form of preservation and as a mode of interactivity for patrons. in 2012, the university of iowa did not have 3d scanning equipment or the expertise to create 3d models from scans. as an example and a consult, the university of iowa libraries used the sousa archive 3d flute collection at the university of illinois at urbana-champaign, purchasing the same strata 3d (photogrammetry) scanning setup and hiring a 3d modeler to scan and post-process the materials. the fluxus project was, in part, a multi-departmental library endeavor to test the potential for 3d scanning as a preservation tool for fragile special collections objects. as part of this endeavor, the university of iowa libraries created a bespoke 3d workflow and digitization effort involving various units: special collections organized and pulled materials; preservation & conservation assessed the stability of the materials and made recommendations on handling; the digitization department oversaw the standard scanning digitization of the fluxus west collection; and the né digital research & publishing department was in charge of figuring out what a 3d scanning workflow would look like and how it would fit into more traditional digitization efforts. this was not an effort to test the feasibility of 3d preservation but to test whether 3d scanning could be used as a form of preservation of special collection objects. this project, like many, was focused on the outcome of 3d scans and not proportionally concerned with the preservation of said scans (the departure of the digital preservation librarian during the project added another hurdle to the initiative). https://doi.org/10.29173/iq20 7/14 moore, jennifer and kettler, hannah scates (2018) who cares about 3d preservation?, iassist quarterly 42 (1), pp. 1-14. doi https://doi.org/10.29173/iq20 the creation process was documented for those wishing to recreate the scanning workflow, as was information contextualizing the 3d scanning data. the workflow that resulted from this project was one that complemented the more traditional 2d scanning workflow, and the resulting 3d scans were treated similarly. the london charter was used as a guide for best practices for the visualization, but it was deemed insufficient for a library implementation focused on 3d preservation. all objects were barcoded and identified in the digital catalogue, renamed and described, and the 3d scanning folders reflected the established file naming conventions used for all other catalogued special collection objects, folders, and boxes. once the objects were photographed and named in accordance with the institutional naming conventions, the post-processing of the 3d scans occurred. most of the 3d data cleaning occurs in the post-processing phase of a 3d scan. the scan goes through various types of transformations that corrupt the raw data but also support end-user needs and reflect the researcher’s intended portrayal. the resulting augmented data represent a new digital object that is in some cases rather loosely derived from the raw data collected during the scanning process. for the fluxus project, there was an interesting tension between the project as a projection of an intentionally ephemeral art movement, the necessary steps to create digital representations of the artifacts, and the motivation of the library to preserve that which was meant to defy cataloging and preservation. for the researchers, the creation of the 3d scans was an interesting aspect of the collection and was worth documenting and keeping for the long term. unlike other projects, where keeping ‘everything’ is more of a reflexive process, with the fluxus project, the need to preserve both the final 3d dataset and the raw data -including the photography images, original 3d datafiles (in the iso standard of stereo lithography (.stl) file format), and the proprietary strata/3dsom software files – was clearly articulated. due to a lack of guidance on the subject, the library was unsure what 3d data to keep and what would be useful in the future if the university of iowa 3d scanning processes needed to be replicated. the preservation process did not include preserving the mode of display of the data (i.e. the now-antiquated adobe flash player), nor the interaction of the user and the 3d object. the space required to house all of these data was significant, especially compared to other digitization efforts. after a reevaluation of the project scope, it was decided to include only a handful of 3d scans rather than a comprehensive 3d library as was originally intended. the space needed to preserve the various products of the 3d scanning was as unanticipated as the lack of guidance regarding the preservation of 3d data. additionally, the fluxus project coincided with the decline of flash as a standard for presenting multimedia objects (including 3d) online (howard, 2012), and when webgl (web graphics library) was just emerging as a possible replacement. packaged with the 3d object is the readme file that captures the workflows taken, the contextualizing paradata about the files and objects within the folder, and a list of potential exports and uses of the data. the outcome of this project was an evaluation and set of recommendations that remain internal to the university of iowa libraries regarding the workflow of 3d scanning as part of a digitization effort for special collection materials, as well as a set of recommendations moving forward regarding 3d as a potential form of preservation. as with many such projects, the documentation and recommendation remains relatively project-specific and internal to the institution. this lack of transparency is not uncommon, as evidenced by a recent survey of 3d practitioners and information specialists, but has the unintended consequence of perpetuating the ad hoc and esoteric nature of workflow and standards development for 3d preservation. https://doi.org/10.29173/iq20 8/14 moore, jennifer and kettler, hannah scates (2018) who cares about 3d preservation?, iassist quarterly 42 (1), pp. 1-14. doi https://doi.org/10.29173/iq20 digital baboon digital baboon is an overarching title for the ongoing digital representation of analog data collected by anthropology professor jane phillips-conroy of washington university in st. louis and anthropologist cliff jolly for their awash national park baboon research project in ethiopia. these data, which span a 30-year period, are being digitized, managed, and preserved for the long-term. analog data types include field sheets, slides, palm prints, and tooth casts.1 items are being scanned in 2d apart from the tooth casts. plaster casts were taken in the field from captured baboons and correspond with their other data, all of which were collected to determine the hybridity of the olive and hamadryas baboons from the study area’s hybrid zone. the tooth casts themselves are important for showing tooth wear, which is enhanced by the ability to scale the 3d models. more than 2000 casts were collected in the field. due to the fragile nature of the casts, it is imperative that they are scanned and preserved. to that end, the securely wrapped plaster casts are transferred to the 3d scanning lab, and each cast is given a year and id number that is related to the capture id embedded in the file naming convention. each item is scanned with an hdi advance r3x13 structured light scanner, which is calibrated regularly with a calibration board. the calibration attempts to achieve 80% coverage (at minimum); to achieve this calibration, the scanner must often do around 100 captures of the calibration board. locally developed documentation is recorded as a part of the technical metadata in a spreadsheet along the following additional attributes: date scanned, by whom, cast number, cast filename, resolution, coverage, average reprojection error, disparity error, point distance, smoothing, hole filling, editing other, file types generated. the scans are completed using a rotary table with 12-18 rotation stops. the scan is repositioned when the rotation set completes. each set of rotations align automatically, and the technician manually puts the sets together and finalizes the scan, removes the noise and fills only very small holes. the files are exported to ascii, .obj., and .stl. to date, they have not been exported to .ply, but all of the original project data has been kept, and an export to .ply may be added to the workflow. to assure that the scans are accurate, a sampling of them is measured using the software against the physical object using calipers. in addition to machine-generated metadata, descriptive readme files are created for the overall collection and subcollections, which include calibration details. scans are moved to washington university libraries’ servers and regularly backed up to the university’s servers. currently, there is not a public access point for the information, and so the data are not accessible to anyone outside the project. while the methodology for this project has been carefully thought through, guidelines for resolution, processing, and formats and standards for metadata were sought to no end. left without definitive answers, localized practices were implemented. dcn review funded by the alfred p. sloan foundation, the data curation network (dcn) is an initiative that enhances data curation services for participating university libraries. part of the mission of dcn is to evaluate and improve skills, workflows, and best practices for data curators. to this end, dcn ran a pilot in the fall of 2016 asking data curation reviewers to evaluate submissions of supplementary data and metadata to the data repository for the university of minnesota. the submission of interest to 3d preservation was supplementary data for reconstructing past craft networks: a case study using 3d scans of late bronze age swords to reconstruct specialized craft networks (golubiewski, 2016). careful review of the supplementary data and metadata found that the methods by which the submitted 3d data were collected were not fully documented in the metadata. the data consisted of https://doi.org/10.29173/iq20 9/14 moore, jennifer and kettler, hannah scates (2018) who cares about 3d preservation?, iassist quarterly 42 (1), pp. 1-14. doi https://doi.org/10.29173/iq20 .bmp files, and the metadata included documentation on necessary software for viewing. presumably, using methods of photogrammetry, these images could be stitched together. however, because the methods of data collection were not clearly described in the metadata, it was necessary to review the dissertation to understand the provenance of these .bmp files. the dissertation included clearer documentation; the methodology included scanning the objects with a structured light scanner, creating screen captures from those, and analyzing the shapes. the gaps in the metadata make reproducibility and reusability unlikely, because curators cannot preserve data that lacks the technical metadata needed to inform future users of what they are working with and how it came to be. this case demonstrates that, without standards and best practices, there is no criteria for researchers to report methods or reliable rubric by which data curators can assess whether metadata are complete. iowa city archaeological data-dump in fall 2016, a group of researchers from the university of iowa’s department of geographical and sustainability sciences approached the library about housing their lidar (light detection and ranging) data for public dissemination. this large dataset consists of many million vertices point clouds, each constituting a gigabyte to several gigabytes each. these data are also in a proprietary file format that is consistent across lidar data creators but is not a standard that the library world considers sustainable (digital preservation handbook, 2017). this data, however, represents the de facto data of an archaeological site that is, at present, inaccessible to the public. the researchers intend to make the data available online and to preserve it “long-term” for public use. the conundrum, insofar as the library is concerned, is determining what file format the data should be moved into, so that it can be 1) disseminated to the public 2) preserved for posterity as archaeological data, and 3) shared with an appropriate level of degradation, so that the files are not large and unwieldy using current technologies. deciding which data are the most appropriate for preservation is a key issue. using the model set forth by the fluxus project, the university of iowa libraries could commit to preserving everything – original files and all derivatives. at this point, the library does not know the state of the data which, as an institutional repository, they are required to preserve. likely, equipment metadata are sufficient, as they are typically created via an automated process like the aforementioned scanning project case studies, but the archaeological and anthropological data surrounding the scan data are most likely recorded elsewhere. additionally, this project will require that the libraries expand their definition of 3d preservation to include that of the user interface to the data. at present, the project will include a web interface for the 3d data using a javascript library called 3dhop14 and build upon this interface to include data points not present in the raw 3d lidar data. pending any upcoming recommendations and standards set by the library and archives community in coordination and conjunction with subject specialists, this project has the potential to be as ad hoc and bespoke as the ones listed thus far. with varying degrees of contextualizing documentation (for example a readme.txt file packaged with all the 3d data outlining the purpose, file types, and decision processes), and the potential for recreation of bespoke metadata schemas, the utility of the 3d lidar data preservation practices is suspect. the idiosyncratic creation of preservation practices has the potential to compound the technological barriers of sharing proprietary 3d data, so that one cannot find nor interface with other 3d datasets, hampering (or making impossible) new discoveries and 3d research. https://doi.org/10.29173/iq20 10/14 moore, jennifer and kettler, hannah scates (2018) who cares about 3d preservation?, iassist quarterly 42 (1), pp. 1-14. doi https://doi.org/10.29173/iq20 questions for the community these case studies demonstrate that coordination is needed between stakeholders, including data collectors (researchers) and data curators (archivists and librarians), to create best practices and standards that reflect the needs of everyone involved and foster digital preservation within and between institutions. such a community could collectively answer questions such as: what should be preserved? this is not only a question of selecting valuable, important, rare, or interesting data. other basic questions include: at what stage is a digital object most stable, and what does raw data mean in the context of a 3d model? because of the often-iterative nature of the scanning process -one 3d model is comprised of many scans and/or other 3d models -levels of processing, correction, and combination should be considered. further, there is the question of format: what is the most stable, authentic, and reusable format in which data should be saved? the challenge of storage space for relatively large objects is real, especially in a world where curators are cautious of destroying any type of 3d data. for even a small item, a project folder consisting of raw data and exports to .stl, .obj, and .ply formats can be relatively large for each object (e.g., 10gb) compared to other digital objects, like scanned manuscript materials. what metadata is necessary? integral to the preservation of the generated 3d data is the description and cataloging of the data to encourage discoverability and provide context to projects like the ones mentioned thus far. many methodologies, as well as software and hardware, can be used to create 3d models. what level of description is necessary to effectively document the creation, augmentation, and use of 3d data for future research endeavors? how would metadata look different for ‘raw data’ verses the final product for end use? what metadata is necessary to help the utility of the data? what metadata would facilitate emulation or the reproducibility of the 3d data? what is the role of emulation in 3d data preservation? open-source options exist for opening and viewing 3d model exports. however, working with 3d data in raw format is often done with proprietary software. the outputs of 3d work, therefore, are proprietary and do not directly translate to data that have the potential to be reused outside of emulation of the software in which they were created. in addition to questions around what constitutes raw data, concerns exist about the software in which the data are created and the preservation of the workflows used to create the data output. this is especially important when one is attempting to recreate the 3d data or replicate the 3d scanning process. coupled with the preservation of workflows or output processes is whether the software (again, typically proprietary) can be maintained alongside the raw 3d data. the reason for this is the maintenance of the 3d creation process as well as the output’s presentation and dissemination mechanism, potentially reaching into the possibilities of preserving user experience as a component in need of preservation consideration. should proprietary software be emulated? is emulation of a proprietary software prohibited? if it is not possible to emulate it, how can the original raw project data be protected from obsolescence? what kinds of commitment and coordination are required between practitioners, curators, and commercial vendors? the projects mentioned here represent only a few of the possibilities and idiosyncrasies that occur while working with digital 3d data. whether the data are created by free-form modeling practices or algorithmically generated as part of a 3d scanning project, without best practices and recommendations for 3d data preservation, the 3d data being captured in this era will be lost to time. such preservation standards are essential for building capacity in libraries, other repositories, and archives to accommodate the data long-term. that includes considerations of what exactly constitutes raw data, the outputs that result from the 3d data use, and the software and equipment used to https://doi.org/10.29173/iq20 11/14 moore, jennifer and kettler, hannah scates (2018) who cares about 3d preservation?, iassist quarterly 42 (1), pp. 1-14. doi https://doi.org/10.29173/iq20 create and disseminate the data. lack of standards is a detriment to the advancement of 3d scholarship as a whole. this paper has mentioned several different types of raw 3d data, all distinct and individually decided upon by the project members and not necessarily by some consensus born out of practice. at this point, the aforementioned project members are attempting to predict the possible future use and reuse of their data. as 3d projects extend beyond the lifespan of the technology, concern about preservation and migration increases. questions are arising about use and reuse of raw 3d data, and about which types of raw data will be useful in future. does the original, unaltered data represent the raw data that will be most useful in future? or is the priority the preservation of the 3d data that represents the intended outcome, i.e. the published data? moving forward an essential element that was present directly or implicitly in every example in this paper is metadata. developing metadata standards specifically for 3d objects has great potential for growing 3d data preservation. documentation and metadata are key to creating data that are reusable well into the future. good documentation and metadata standards, which a set of best practices would address, represent the waterfall impact of detailing data and processes that can inform and bolster the efforts of processes, software emulation, data migration, and dissemination. creating a standard for metadata and a set of best practice recommendations would have immense impact on the overall preservation and interoperability of 3d research. conversations are occurring at conferences15 and on listservs16 about the problem of preserving digital 3d data. efforts are underway to set the foundation for coordinated efforts regarding 3d preservation between stakeholders. the authors of this paper are involved in ongoing conversations regarding 3d support focused on the practices of data creators. these conversations typically focus on outlining the different ways in which 3d data are created by scholars, but little attention is given to preservation, mainly because there are too few preservation specialists in the room. the authors have since collaborated in administering a survey17 regarding preservation practices, which revealed that many invested library and archive professionals are thinking about the problem of 3d data but have yet to come together with data creators to figure out what is needed to support their research. the next step is pulling the information from these previous conversations and working toward solutions together. 0 50 100 use best practices don't use best practices community members using best practices 0 50 100 interested not interested other community members interested in collaborative development https://doi.org/10.29173/iq20 12/14 moore, jennifer and kettler, hannah scates (2018) who cares about 3d preservation?, iassist quarterly 42 (1), pp. 1-14. doi https://doi.org/10.29173/iq20 references the data curation network: discussing and providing solutions for data issues in a collaborative way | agricultural information management standards (aims). (n.d.). available from: http://aims.fao.org/activity/blog/data-curation-network-discussing-andproviding-solutions-data-issues-collaborative-way doyle, j., viktor, h., & paquet, e. (2009). long-term digital preservation: preserving authenticity and usability of 3-d data. international journal on digital libraries, 10(1), 33–47. available from: https://doi.org/10.1007/s00799-009-0051-7 golubiewski-davis, k. (2016). reconstructing past craft networks: a case study using 3d scans of late bronze age swords to reconstruct specialized craft networks. accessed available from: http://conservancy.umn.edu/handle/11299/181730 grosman, l., smikt, o., & smilansky, u. (2008). on the application of 3-d scanning technology for the documentation and typology of lithic artifacts. journal of archaeological science, 35(12), 3101–3110. available from: https://doi.org/10.1016/j.jas.2008.06.011 guides to good practice: 3d_toc. (n.d.). available from: http://guides.archaeologydataservice.ac.uk/g2gp/contents. [accessed may 14, 2018] guides to good practice: photogram_1-1. (n.d.). available from: http://guides.archaeologydataservice.ac.uk/g2gp/photogram_1-1. [accessed march 24, 2017] howard, bill. (2012). flash – chrome for android beta. retrieved november 20, 2017, from adobe air and adobe flash player team blog: blogs.adobe.com/flashplayer/2012/02/flashchrome-for-android-beta.html digital preservation handbook (2017). available from: http://www.dpconline.org/handbook/technical-solutions-and-tools/file-formats-and-standards. [accessed november 20, 2017] digitale 3d rekonstuktionen in virtuellen forschungsumgebungen. (n.d.). available from: https://www.herder-institut.de/forschung-projekte/laufende-projekte/digitale-3drekonstruktionen-in-virtuellen-forschungsumgebungen.html. [accessed march 24, 2017] london charter. (n.d.). available from: http://www.londoncharter.org/. [accessed march 24, 2017] cidoc crm. (n.d.). available from: http://www.cidoc-crm.org/. [accessed march 24, 2017] d’andrea, a., & fernie, k. (2013). 3d digitisation of icons of european architectural and archaeological heritage (final no. d6.1: report on metadata and thesaurii) (p. 45). european commission’s ict policy support programme. retrieved from http://www.3dicons-project.eu/eng/resources/d6.1report-on-metadata-and-thesauriihumanities heritage 3d visualization: theory and practice. (n.d.). available from: https://www.neh.gov/divisions/odh/institutes/humanitiesheritage-3d-visualization-theory-and-practice. [accessed march 24, 2017] johnston, l. r., association of college and research libraries, & american library association. (2017). curating reserach data. a handbook of current practice with 30 case studies contributed by practitioners in the field volume two volume two. accessed available from: http://www.ala.org/acrl/sites/ala.org.acrl/files/content/publications/booksanddigitalresources/digit al/9780838988633_crd_v2_oa.pdf martorelli, m., pensa, c., & speranza, d. (2014). digital photogrammetry for documentation of maritime heritage. journal of maritime archaeology, 9(1), 81–93. available from: https://doi.org/10.1007/s11457-014-9124-x molloy, b., wisniewski, m., lynam, f., o’neill, b., o’sullivan, a., & peatfield, a. (2016). tracing edges: a consideration of the applications of 3d modelling for metalwork wear analysis on bronze age bladed artefacts. journal of archaeological science, 76, 79–87. available from: https://doi.org/10.1016/j.jas.2016.09.007 https://doi.org/10.29173/iq20 http://guides.archaeologydataservice.ac.uk/g2gp/photogram_1-1 http://www.londoncharter.org/ http://www.cidoc-crm.org/ 13/14 moore, jennifer and kettler, hannah scates (2018) who cares about 3d preservation?, iassist quarterly 42 (1), pp. 1-14. doi https://doi.org/10.29173/iq20 neh announces protecting our cultural heritage. (2016, march 7). available from: https://www.neh.gov/news/press-release/2016-03-09. [accessed march 24, 2017] niven, l., steele, t. e., finke, h., gernat, t., & hublin, j.-j. (2009). virtual skeletons: using a structured light scanner to create a 3d faunal comparative collection. journal of archaeological science, 36(9), 2018–2023. available from: https://doi.org/10.1016/j.jas.2009.05.021 project history | mayaarch3d. (n.d.). available from: http://www.mayaarch3d.org/language/en/about/project-history/ [accessed march 24, 2017] sousa archives music instrument digital image and 3d model collection. (n.d.). available from: http://imagesearchnew.library.illinois.edu/cdm/landingpage/collection/sousa. [accessed march 24, 2017] staff, i. b. r. (2015, september 14). isu professor earns national grant to preserve artifacts with 3d scans. available from: http://idahobusinessreview.com/2015/09/14/isu-professor-earns-nationalgrant-to-preserve-artifacts-with-3d-scans/. [accessed march 24, 2017] wachowiak, m. j., & karas, b. v. (2013). 3d scanning and replication for museum and cultural heritage applications. journal of the american institute for conservation journal of the american institute for conservation, 48(2), 141–158. what is digital curation? | digital curation centre. (n.d.). accessed march 24, 2017, available from: http://www.dcc.ac.uk/resources/briefing-papers/introduction-curation/what-digital-curation. [accessed march 24, 2017] yastikli, n. (2007). documentation of cultural heritage using digital photogrammetry and laser scanning. journal of cultural heritage, 8(4), 423–427. https://doi.org/10.1016/j.culher.2007.06.003 1 jennifer moore, washington university in st. louis, 1 brookings drive, campus box 1169, st. louis, mo 63130. mail: j.moore@wustl.edu. 2 hannah scates kettler, university of iowa, 100 main library, digital research and publishing, iowa city, ia 52242 3 community standards for 3d preservation survey conducted in 2017 underscores the continued lack of community and standard practice. https://mfr.osf.io/render?url=https://osf.io/tcn6h/?action=download%26mode=render 4 http://archaeologydataservice.ac.uk/ 5 http://3dicons-project.eu/ 6 http://www.londoncharter.org/ 7 https://www.digitalantiquity.org/ 8 http://pro.carare.eu/doku.php?id=support:metadata-schema 9 http://augmentedasburypark.com/ 10 http://humanorigins.si.edu/evidence/3d-collection/artifact 11 http://vwhl.soic.indiana.edu/villa/index.php 12 http://www.meshlab.net/ 13 https://gomeasure3d.com/hdi/advance/ 14 http://3dhop.net/index.php 15 such conferences and institutes like the digital library federation forum (https://forum2017.diglib.org/), the neh funded advanced challenges in theory and practice in 3d modeling of cultural heritage sites https://doi.org/10.29173/iq20 14/14 moore, jennifer and kettler, hannah scates (2018) who cares about 3d preservation?, iassist quarterly 42 (1), pp. 1-14. doi https://doi.org/10.29173/iq20 (https://www.neh.gov/divisions/odh/institutes/advanced-challenges-in-theory-andpractice-in-3d-modeling-cultural-heritage). 16 cs3dp google group https://groups.google.com/forum/#!forum/community-standardsfor-3d-data-preservation-cs3dp 17 community standards for 3d preservation survey conducted in 2017 underscores the continued lack of community and standard practice. https://mfr.osf.io/render?url=https://osf.io/tcn6h/?action=download%26mode=render https://doi.org/10.29173/iq20 28 iassist quarterly fall & winter 2007 luis martinez-uribe* digital repository services for managing research data: what do oxford researchers need? abstract uk researchers are facing the challenges of having to comply with funder requirements to submit data management plans and make their data available. academic institutions have the responsibility to support their researchers to fulfill their contractual requirements with funding agencies. increasingly, repository services are dealing with the management of research output. managing research data to ensure digital materials adhere to the right standards and are securely stored, shared and preserved can be complex and resource intensive. understanding how researchers work is the key to designing university repository services to manage research data. this article describes the university of oxford’s federation of digital repositories and introduces the scoping digital repository services for research data management project to present the findings from the requirements gathering exercise carried out to understand oxford researchers’ practices and needs. . introduction the proliferation of gadgets that deliver information ming in the current climate of large scale international projects around research data such as the australian national data service, the us national science foundation datanet and the uk research data service, the management of research data is a topic that attracts interest from policy makers worldwide because of the importance of data in the age of the knowledge economy (pmseic 2006). there are many efforts to foster standards for data description and sharing as well as to establish national and international federations of digital repositories to deal with the management and curation of these digital resources. uk universities and their researchers are facing the challenges themselves. researchers in all disciplines are increasingly being asked by funding bodies to not only make their data available but also to submit data management and data sharing plans with their funding applications. these plans should describe what datasets will be created that are worth keeping, what standards will be used, how they will be made available and who will be responsible for their long-term preservation (weaver 2007). managing and curating research data poses many challenges because of their heterogeneity and a lack of standards as well as many unresolved ethical issues (carusi & jirotka 2008). some of the benefits of the active management and curation of research data include the possibility to replicate research results, avoiding expensive data collection by promoting data reuse and protecting the contributor’s sensitive information (schroeder & axelsson 2007). when planning how best to manage and curate research data, efforts are sometimes mainly focused on understanding the data themselves, the different types, their volume and specific technical or legal issues surrounding them. nonetheless engaging with the producers and users of those datasets is at times forgotten. understanding researchers’ needs and workflows will help to comprehend how they work with data, why they create them in the first place and their reasons for managing these resources the way they do. this approach will also assist to identify services to support them throughout their research process so that they can fulfill the requirements from funders. this paper describes briefly oxford’s federation of digital repositories and then introduces the scoping digital repository services for research data management project and the findings of the requirement gathering exercise carried out between may and june 2008 as part of the project. background: a federation of digital repositories in a collegiate university the university of oxford, the oldest university in the english speaking world, has a complex structure with divisional departments, institutes, independent colleges and more than a hundred libraries. a highly devolved institution which has on many occasions been compared to a microcosm of the entire uk higher education system; this organizational arrangement is mirrored with a devolved computing structure (jeffreys 2008). the main central ict service provider is oxford university computing services (oucs) that operates the primary computing infrastructure such as core networks, back up servers and core support services. another central service provider providing ict for its users is the oxford university library services (ouls). these centralized services are complemented with local ones embedded in departments, institutes and colleges with their own ict infrastructure and support teams. iassist quarterly fall & winter 2007 29 acting as an overall umbrella is the office of the director of it (odit) providing strategic direction for it at the university. in terms of digital repositories, oxford can be seen as a federation with the oxford digital repositories steering group providing a coordinating role for the development of repository services, see figure 1. the oxford research archive (ora) is the digital repository infrastructure for ouls providing permanent and secure online archive of research output materials produced by members of oxford. its content includes journal articles, conference papers, working papers, theses and other grey literature. however, there are a number of other digital collections and repository activities like the data resources from the oxford e-research centre (oerc), the google materials, the oxford digital library collections and others. this wealth of digital resources, the ”increasing need to manage curation and access to research data” (fraser 2005) and the lack of interaction between activities led to the creation of the digital repositories research coordinator post and the start of the scoping digital repository services for research data management project. the scoping digital repository services for research data management project the scoping digital repository services for research data management project is a joint effort between odit, oucs, ouls, oerc and reports to the oxford digital repositories steering group. the project scopes the requirements, including infrastructure and interoperability, for repository services to manage and curate research data generated by oxford researchers. before describing the project any further it is important to clarify the scope of what is meant here by research data, data management and repository services. research data are a heterogeneous type of research output which can take many forms (text, numbers, audio, images, moving images, etc) and might be created for different purposes during the research process. the national science foundation provides a useful categorization based on the origins of the data: observational, computational or experimental (2005). observational data are historical records such as opinion polls or precipitation measurements. these data cannot be recollected, thus the importance of preserving it indefinitely. on the other hand, figure 1. oxford digital repositories structure (jeffreys & fraser 2007) 30 iassist quarterly fall & winter 2007 computational data resulting from computing simulations can be reproduced. preserving the input files that allow replicating the simulation is more important than preserving the raw data obtained through the simulation. data produced through experiments poses other challenges. in many cases although the experiment could be reproduced this may be too costly. the research information network (2008) adds two more categories in their data typology: derived and canonical data. the former refers to data resulting from some form of processing to primary data while the latter refers to those reference datasets such as the gene sequence. data management is a vast discipline and includes activities such as database design, data compliance and data modeling, and many others. it takes from disciplines like information and knowledge management that focus on looking after these assets from the moment of creation and dissemination through the organization. these activities include storage, retrieval, use, access, preservation or disposal (macevi & wilson 2005; alavi & leidner 2001). the us department of defense defines this as “understanding current and future data needs” (parker 2000) and as pointed out by a study from virginia commonwealth university (aiken et al 2007) the management of data needs to be seen as a means to an end. in this project one of the main drivers for improving the current infrastructure for data management is the pressing requirement from funding bodies to make data available and provide data management and data sharing plans with funding applications. the term “repository services”, in its widest sense, refers not only to a technical infrastructure that allows storage, access, description, dissemination and preservation of digital objects but also the support services to assist researchers with technical and legal issues and policies for the creation, deposition and sharing of digital research outputs. requirements capture: methodology and participation one of the main objectives of the project was to document data management practices of oxford researchers as well as to capture their requirements for services to help them manage their data more effectively. in order to do this, thirty-seven face-to-face interviews were conducted between may and june with researchers from twentysix departments and faculties from oxford, see table 1. this positive response to the interview calls provided a good cross-section: 58% of which were on the ground researchers; 28% heads of departments/faculties or research teams; and the remaining 14% a mixture of data managers, it officers and administrators embedded in departments/faculties and research teams. the interview framework was largely based on the methodology employed by other oxford projects (the e-infrastructure use cases, building a virtual research environment for humanities and the integrative biology iassist quarterly fall & winter 2007 31 virtual research environment ) with some adjustments to fit with the requirements of this study. the interviews were semi-structured to engage in conversations with researchers and delve into participant reasons for doing things they way they do. a generic research life cycle model was referenced to structure the conversation. the research life cycle started with the funding application, moved into data collection and data processing, and finished with data publication. several approaches were used to identify interview candidates: the first choice of interviewees was guided by suggestions from members of oucs, ouls and oerc and then a call for interviewees circulated amongst research facilitators. for every researcher interviewed, a snowball sampling approach was used to identify further candidates. in addition to this an event, the research data management workshop, was organized to complement interview findings. this event brought 46 attendees throughout the day from 24 departments, research centres and colleges. both the interviews and the workshop also formed the basis of the oxford case study for the uk research data service feasibility study, a joint project between research libraries uk and the russell group it directors group, aiming to assess feasibility and cost of developing a national shared data service for research data generated in uk higher education institutions. findings: from funding application to data publication findings from the interviews and the workshop revealed that the management of research data occurs with varying degrees of maturity across oxford university. there are some departments that have been dealing with very sensitive data for several years and they have policies and procedures in place that support a robust technical infrastructure. on the other hand, many other units in the university tend to work on a more ad-hoc basis and data management relies on the individual researcher’s skills. overall, oxford researchers felt there were potential services to help them manage their data more effectively. at the funding application stage, researchers tend not to plan the management of the data at the outset of their research project in detail. as one of the researchers interviewed stated “…you are not interested in this because you are interested in the science; the technical issues come up later.” with the wide variety of funding sources available, they find it confusing to understand what the different requirements are for making their data available and the retention period. the types of data collected were, as expected at the beginning of the project, many and very diverse. there exist a significant variety of forms and formats, some of them proprietary. in addition, there are a wide range of sizes, from few megabytes to several petabytes. some of the data were highly sensitive, and strict ethical and security protocols needed to be followed to collect these data. the long-term usefulness of the data also varied enormously: in some fast moving disciplines the data would be relevant up to five years before better data could be produced, while in other cases it is impossible to reproduce the results, and the life-span is indefinite. once the data are collected they are mostly stored on personal computers or departmental servers with a variety of security and back up procedures; although there are some horror stories about shelves full of highly valuable data stored on cds and dvds. metadata accompanying the data tends to be minimal and data are organized in hierarchical folder structures with file names that make sense to the researcher. these data are then shared in informal ways and this happens mostly by email or portable media. problems arise when the size of the data increases and then the storage and sharing becomes difficult. very few of the researchers interviewed had deposited any data in domain specific data archives such as the natural environment research council data centres or the uk data archive . nonetheless, many of them were publishing data on their departmental websites. data ownership is seen as a major issue. there have been cases where data have been generated from human subjects as part of collaborative research projects between many institutions in different countries and with several funding bodies. in cases like these and others, researchers struggle to understand who owns the data produced in their research projects. in terms of sharing their data, although researchers tend to feel very attached to their data, they believe that if their research is publicly funded, then their data should be made publicly available. the top three requirements expressed by oxford researchers for services to help them manage their data more effectively were: • a secure and user-friendly solution that allows storage of large volume of data and sharing of these in a controlled fashion, allowing fine grained access control mechanisms. \ • a sustainable infrastructure that allows publication and long-term preservation of research data for those disciplines not currently served by domain specific services such as the uk data archive, nerc data centres, european bioinformatics institute and others. • advice on practical issues related to managing data across its life cycle. this help would range from assistance in producing a data management/sharing plan; advice on best formats for data creation and options for storing and sharing data securely; to guidance on publishing and preserving these research data. 32 iassist quarterly fall & winter 2007 conclusion and next steps as pointed out by lyon (2007), in order to manage and curate research data it is crucial that the different communities of researchers, librarians and computing services work together. nonetheless, in order to deploy an effective and usable infrastructure for managing research data it is key to understand the producers of the data themselves, their workflows and needs. the interviews and workshop have provided enough evidence about current data management practices at oxford and researchers’ requirements for services to help with these. the findings will be complemented with another workshop and a consultation with service providers in oxford to assist in producing a set of recommendations to improve and coordinate the provision of digital repository services for research data at oxford16. * contact: luis martinez-uribe, digital repositories research co-ordinator, oxford e-research centre, university of oxford. e-mail: luis.martinez-uribe@oerc. ox.ac.uk references aiken, p., d. allen, b. parker & a. mattia (2007 april) measuring data management practice maturity: a community's self-assessment. computer. retrieved march 15, 2008 alavi, m. & d. e. leidner (2001) review: knowledge management and knowledge management systems: conceptual foundations and research issues. mis quarterly, 25, 107-136 carusi, annamaria & jirotka, marina.(2008). from data archive to ethical labyrinth. qualitative research, forthcoming jeffreys, paul. (2008). oxford’s computing model : delivering computing services to the university. retrieved july 15,2008 from: http://www.ict.ox.ac.uk/odit/ itcoordination/oditoxford%27scomputingmodel.pdf jeffreys, paul & fraser, mike (2007). oxford digital repositories steering group meeting agenda job description for digital repositories research co-ordinator – annex 1: oxford digital repositories structure. retrieved july 15,2008 from: http://www.ict.ox.ac.uk/ repositories/meetings/odrsg_agenda-papers_4may07_2. pdf fraser, mike. (2005). towards a research repository for oxford university.retrieved march 10,2008 from: http:// ora.ouls.ox.ac.uk/objects/uuid:5fa206c9-c400-403c-ab69174ce8604a7a lyon, liz. (2007). dealing with data: roles, rights, responsabilities and relationships – consultancy report. retrieved may 17,2008 from: http://www.ukoln.ac.uk/ ukoln/staff/e.j.lyon/reports/dealing_with_data_report-final. pdf macevi, e. & t. wilson. (2005). introducing information management: an information research reader. london, uk : facet publishing national science foundation. (2005). long-lived digital data collections; enabling research and education in the 21st century. retrieved march 15, 2008 from http://nsf. gov/pubs/2005/nsb0540/ parker, b. (2000). enterprise data management process maturity. in data management handbook. auerbach publications. pmseic (prime minister’s science, engineering and innovation council) working group on data for science. (2006). from data to wisdom: pathways to succesful data management for australian science. retrieved march 12, 2008 from: http://www.dest.gov.au/sectors/science_ innovation/publications_resources/profiles/presentation_ data_for_science.htm research information network.(2008). stewardship of digital research data : a framework of principles and guidelines. retrieved july 15,2008 from http://www.rin. ac.uk/data-principles [proceedings] schroeder, ralph & axelsson, ann-sofie. (2007). making it open and keeping it safe: e-enabled data sharing in sweden and related issues. e-social science 2007 ann arbor, michigan us. retrieved july 2008 from: http://ess.si.umich.edu/papers/paper139.pdf weaver,belinda. (2007). constructing a research project data management plan. creating a data management strategy for new research projects workshop, university of queensland, australia. retrieved on july 15, 2008 from: http://www.library.uq.edu.au/escholarship/orca.html footnotes 1. http://www.ands.org.au/ 2. sustainable digital data preservation and access network partners (datanet) http://www.nsf.gov/funding/ pgm_summ.jsp?pims_id=503141&org=oci&from=home 3. http://www.ukrds.ac.uk/ 4. http://ora.ouls.ox.ac.uk/ 5. http://www.oerc.ox.ac.uk/resources 6. http://www.bodley.ox.ac.uk/google/ iassist quarterly fall & winter 2007 33 7. http://www.odl.ox.ac.uk/ 8. an up to date list of activities can be found at: http:// www.ict.ox.ac.uk/repositories/index.xml.id=body.1_div.6 9. http://www.ict.ox.ac.uk/odit/projects/digitalrepository/ 10. www.eius.ac.uk/ 11. http://bvreh.humanities.ox.ac.uk/ 12. http://www.vre.ox.ac.uk/ibvre/ 13. digitalrepository/workshops.xml 14. http://www.nerc.ac.uk/research/sites/data/ 15. http://www.data-archive.ac.uk 16. a detailed report of the findings from the interviews and the workshop, next steps after the interviews and other outputs from the project can be found at: http://www.ict. ox.ac.uk/odit/projects/digitalrepository/ vol273.indd by 12 iassist quarterly fall 2003 by juri stratford * internet surveillance: recent u.s. developments the u.s. federal government has implemented both technologies and policies related to internet surveillance. while the recent discussion tends to focus on the usa patriot act following the september 11 terrorist attacks, the u.s. congress held hearings addressing internet surveillance and fourth amendment protections as early as april 2000. at this point, congress criticized the lack of oversight on the federal bureau of investigationʼs internet surveillance system, carnivore. congress revisited the issue of internet surveillance days after the september 11 attacks when the attorney general presented draft legislation addressing “new surveillance authorities;” the usa patriot act developed from this proposal. the executive branch of the federal government has since pursued a number of policies and strategies dealing with internet surveillance and data mining. many of the difficulties surrounding the question of internet surveillance center on the analogies between internet surveillance and telephone surveillance. are these analogies appropriate; and if we accept that there is a place for telephone surveillance in law enforcement and intelligence activities, does internet surveillance or data mining naturally follow? carnivore (dsc 1000) carnivore is a microsoft windows based system developed and used by the u.s. federal bureau of investigation that directly connects to an ispʼs server. the fbi draws analogies from telephone surveillance to describe the carnivore system. carnivore is used in two ways: as a “content wiretap” and a “trap and trace/pen-register.” a telephone “content wiretap” is where law enforcement eavesdrops on a suspectʼs telephone calls, recording the oral communications on tape. carnivore provides analogous capabilities for e-mail, capturing all e-mail messages to and from a specific account or all the network traffic to and from a specific ip address. “trap and trace” technology tracks all caller ids of inbound telephone calls, while “pen-register” tracks all outbound telephone numbers dialed. similar functionality for e-mail consists of capturing all e-mail headers (including e-mail addresses) going to or from an e-mail account, but not the actual contents. for other forms of internet activity similar functionality consists of listing all the servers (web servers, ftp servers, etc.) accessed but not capturing the content of this communication, tracking everyone who accesses a specific web page or ftp file, or tracking all web pages or ftp files that a suspect accesses (independent technical review of the carnivore system; final report, 2002). earthlink carnivore first came to public attention through a february 4, 2000 court decision. an internet service provider, later identified as earthlink, questioned the legal authority of the court to issue an order requiring the installation of a “device which captures the time, date, source, and destination of electronic mail (e-mail) messages sent to and from an e-mail address maintained by a customer at the isp.” the court found that it had the legal authority under the pen register statute (18 usc 3122) to issue such an order (court order authorizing carnivore installation at earthlink, 2000). the court decision and an article in the wall street journal prompted congressional hearings on carnivore in april and july 2000. in testimony before the house committee on the judiciary, july 24, 2000, tom perrine, on behalf of the san diego supercomputer, argued that “[t]he current debate… is really about the risks in naively attempting to simply translate the policies, law, and practices of telephone wiretaps into the digital realm of the internet” (perrine, 2000). while internet surveillance through carnivore employed strategies similar to those employed in telephone surveillance, congress questioned the potential to sift through large quantities of private communications without regard to source or destination. house majority leader richard k. armey (r-tx) stated that “nobody can dispute the fact that this [carnivore technology] is not legal… within the context of any current wiretap law” (poole, 2000). iassist quarterly fall 2003 13 patriot act following the september 11 terrorist attacks, attorney general john ashcroft submitted a draft of the legislation, the “mobilization against terrorism act” to congress on september 19, 2001. the patriot act developed from this proposal. president bush signed the bill into public law 107-56 on october 26, 2001. while much of the patriot act builds the infrastructure necessary to respond to terrorist activities, a significant section of the patriot act deals with “new authorities” that enhance the governmentʼs ability to conduct surveillance and share information. privacy concerns were addressed in part through a sunset provision; many of these new surveillance authorities expire at the end 2005. however, much of the controversy surrounding the patriot act continues to focus on these “new authorities.” the fbi highlights these new authorities in their document entitled “field guidance on new authorities (redacted) enacted in the 2001 anti-terrorism legislation.” these highlights include the nationwide effect of court orders for pen registers or trap and trace installations; nationwide search warrants for e-mail; and the use of carnivore installations. another interesting point of clarification under the patriot act is that computer system administrators, e.g. an isp, can obtain the assistance of law enforcement to monitor activity on their own computers (field guidance on new authorities [redacted] enacted in the 2001 terrorism legislation, 2001). in a congressional research service report on the patriot act, charles doyle describes federal communications privacy law as a three tiered system protecting the confidentiality of private telephone, face-to-face, and computer communications while enabling authorities to identify and intercept criminal communications. first, title iii of the omnibus crime control and safe streets act of 1968 prohibits electronic eavesdropping on telephone conversations, or computer or other forms of electronic communications in most instances. it also gives authorities a narrowly defined process for electronic surveillance to be used as a last resort in serious criminal cases. next, 18 usc 2701-2709 covers telephone records, e-mail held in third party storage. finally, 18 usc 31213127 governs court orders approving the governmentʼs use of trap and trace devices and pen registers which identify the source and destination of calls made to and from a particular telephone. the patriot act modifies the procedures at each of these three levels. • permits pen register and trap-and-trace orders for electronic communications (e.g. e-mail); • authorizes nationwide execution of court orders for pen registers, trap-and trace-devices, and access to stored e-mail or communication records (i.e. carnivore technology); • treats stored voice mail like stored e-mail (rather than telephone conversations); • permits authorities to intercept communications to and from a trespasser within a computer system (with permission of the systemʼs owner); • adds terrorist and computer crimes to title iiiʼs predicate offense list; • reinforces protection for those who help execute title iii, ch. 121 and ch. 206 orders; • encourages cooperation between law enforcement and foreign intelligence investigators; • establishes a claim against the u.s. for certain communications privacy violations by government personnel a sunset provision terminates the authority found in many of these provisions and several of the foreign intelligence amendments on december 31, 2005. however, section 216 addressing the use of carnivore is not subject to the sunset provision (doyle, 2002). total information awareness (terrorist information awareness) the u.s. department of defense committed resources to develop data mining capabilities through the total information awareness program. the defense advanced research projects agency (darpa) began work on tia in 2003. the objective of the program was to integrate information technologies into a prototype that could determine the feasibility of searching vast quantities of data as well as determine links or patterns in the data that are indicative of terrorist activities. the program sought to develop information technology in three areas including language translation, data search with pattern recognition and privacy protection, and advanced collaborative and decision support tools as darpa is a research and development agency, the intent was for darpa to turn over their prototype for adoption to the department of defense and other federal agencies. while the tia itself is now defunct, the federal government continues to use “data mining” techniques in other initiatives such as the multi-state anti-terrorism information exchange (matrix). in a review of the tia project, the department of defense inspector general reported that the federal government is likely to adopt other versions of “data mining” in the future. (information 14 iassist quarterly fall 2003 technology management: terrorism information awareness program, 2003). recent developments the federal government continues to implement new policies to incorporate internet surveillance and data mining into law enforcement and terrorist investigations. on may 30, 2002, attorney general john ashcroft issued new guidelines to permit the fbi to tap commercial databases, employ data mining and search the internet for evidence of terrorist activity. these new guidelines relax restrictions that were imposed on the fbi in 1976 to curb excesses of the 1950s and 1960s, when the agency actively spied on americans involved in the civil rights movement, political dissent, and war protests (the attorney generalʼs guidelines on general crimes, racketeering enterprise and terrorism enterprise investigations, 2002). in 2003 and 2004, both president bush and the attorney general have made public appeals for the extension of the patriot act. these extensions refer to sections of the title ii surveillance authorities set to expire next year under the actʼs sunset provisions (u.s. president, 2004). if we accept that there is any appropriate need for surveillance activities such as telephone wiretapping then we canʼt dismiss the question of internet surveillance out of hand. while the scope of telephone surveillance is limited by the means of communications, the scope of internet surveillance is not. congress needs to revisit the question of internet surveillance in an impartial setting that protects citizens ̓privacy while enabling law enforcement and terrorist investigations. references the attorney generalʼs guidelines on general crimes, racketeering enterprise and terrorism enterprise investigations (2002, may 30), (retrieved from http://www. usdoj.gov/olp/generalcrimes2.pdf) court order authorizing carnivore installation at earthlink (2000, february 4), (retrieved from http://www.epic.org/ privacy/carnivore/cd_cal_order.html) doyle, charles (2002, april 15). the usa patriot act: a legal analysis, (retrieved from http://www.epic.org/ privacy/terrorism/usapatriot/rl31377.pdf) field guidance on new authorities (redacted) enacted in the 2001 terrorism legislation (2001), (retrieved from http://www.epic.org/privacy/terrorism/doj_guidance.pdf) independent technical review of the carnivore system: final report (2002, december 8), (retrieved from http:// www.epic.org/privacy/carnivore/carniv_final.pdf) perrine, tom (2000, july 24). testimony before the u.s. congress house committee on the judiciary, subcommittee on the constitution, (retrieved from http:// www.house.gov/judiciary/perr0724.htm) poole, patrick (2000). ʻcarnivore ̓under siege, (retrieved from http://www.worldnetdaily.com/news/article. asp?article_id=20036) u.s. department of defense. office of the inspector general (2003, december 12). information technology management: terrorism information awareness program, (retrieved from http://www.dodig.osd.mil/audit/reports/ fy04/04-033.pdf) u.s. president (2004, january 20), state of the union address, (retrieved from http://www.whitehouse.gov/ news/releases/2004/01/20040120-7.html) * paper presented at the iassist conference, madison, may 2004. juri stratford, government information and maps, shields library, university of california, davis, california, 95616. contact: jtstratford@ucdavis.edu data bank on the u.s.a. and soviet-american relations by tatyana yudina ' senior researcher, spasibo institute of the usa and canada studies ussr, academy ofsciences moscow, ussr introduction our institute institute of usa and canada studies of the ussr academy of sciences was established in 1969. there are several departments in it; usa domestic policy, foreign policy, military policy, agriculture, management system, and others. some 200 researchers and post graduates work in our institute. two years ago several dozen personal computers (ibm at and xt clones) were purchased and a new section was established that of applied research and informatics. invited to join the section were researchers interested in new technologies and new sources of information. i was one of 12 memebers to join this new group. we also have two professional full-time programmers and several part-time programmers working with us; not enough to meet our needs, however. in this we are not alone as i learned from my attendance of the 1990 iassist conference. i discovered that this is a common problem among data libraries throughout the world understaffed with the personnel that are needed most of all. our section has several major functions: teaching researchers computer skills studying the software market for emerging applicationsthat may be rquired by our department an providing traing in software use designing the computer-based information-retrieval system for our institute and coordinating the efforts of different departments of our institute taking into account the strategic aim of our data base to have as much data as possible, studying and evaluating new sources of information; cd-roms, on-line data-banks, archives of machinereadable data etc. these are our main tasks to say nothing about educating our chiefs and encouraging them to give money for ongoing operations and development of our programs. we were among the first of the humanitarian institutes to use personal computers, with limited access to consultants, and with no prior computing experience, we have made do through trial and error. computer-based information retrieval system "usa and soviet-american relations our institute is a section of the usa studies in the ussr, and our mandate is to serve as a national data bank on the usa and soviet-american relations. we are fortunate in that there is a great quantity and variaty of high quality data that the usa produces about itself. our problem is one of finding funds for the purchase of the data we choose. happily we have other responsibilties in addition to building data collections. the system is multifunctional and multilingual, it contains data of various genres, original full texts or created especially for the system in russian, in english and some in french. we hope that in future it will be supported by the automatic translation subsystem. there are two blocks of information in our system one about the usa and the other about the soviet-american relations. both blocks contain several types of data sets : the "usa reference module"; an electronic version of the usa encyclopedia recently published by our institute, a statistical module consisting of long range timeseries data on economic indicators in the usa, full-text documents, such as soviet-american agreements from 1933 to present (updated on a fortnightly basis), a chronology of usa-ussr relations, articles and statistics about economic and trade relations between the two countries, mutual soviet-american projects and joint ventures, biographies of the american political leaders, with a special module: "speeches" short references of main ideas processed by our linguistic means (further they will be mentioned), archive collections, bibliography of publications about the usa and soviet-american relations done by our library, which has started to create an on-line catalogue of its holdings. in our data collection we have data sets on such subjects as: the congress of the usa, president bush's cabinet, white house personnel and the structure of the american administration, the constitution of usa, etc. these iassist quarterly subject data sets are being created by our researchers we encourage such projects, since we simply can't afford to purchase such collections from the usa. in collectiong information our researchers make useof american online services such as dialogue, lexis-nexis, among others, and which is funded by the academy of sciences. work organization in developing our information system we have introduced some new elements of work organization our researchers are invited to build data collections in their field of expertice. we did this for several reasons, the first one is that we hope that a researcher expert in some particular sphere knows the information market better, can evaluate different sources, and choose the most informative ones. he is also supposed to process the data, to administer and update the collection. we encourage our researchers to start such data collections, though there are some social and psychological problems to this in that researchers become possive of their data. we hold competitions amoung the researchers, with the winners going to attend the summer workshops at the university of michigan or essex university , where they learn about new methods of analysis and modern software for personal computers. the idea of data collections built by experts is that they use the data as a base for their analysis while others mostly use it for information. we especially encourage such data sets that may be used for modelling. new sources of information as for the on-line sources of data we have on-line communication with soviet press agency which widely reports international news, and speeches of american political leaders. all the speeches of american political leaders which touche upon the problems of sovietamerican relations are supposed to be processed and added to our political "portraits" databases. one of the most important source of data for our research is the data archive of the inter-university consortium for political and social research at univeristy of michigan. we believe that soon we'll be honoured with the opportunity to join it and to access its data collections which will help us to raise the level of our research. for a decade we were unable to join icpsr because of not fulfilling one of its major conditions the ussr had no machine-readable information about itself in the market. by now the situation has changed our state statistics agency (goscomstat) announced that a special section has been organized with the aim to produce machinereadable collections of data about the development of the ussr. our other interest in joining the inter-university consortium for political and social research is its summer school program this year we send two of our researchers to icpsr's summer program and hope that it will help us to move to new levels of research using modem software and methodology in analyzing social and political proccesses. software the software that is now being designed by our main programmer, valentine ponomarenko, is supposed to allow to use direct entry, scanners, on-line communication, and cd-roms to add data to our information systems. it will also provide us with an interface transition from data set to data set. having information in several languages we come accross the problem of using english and cyrrilic alphabets which we manage to solve by having software designed by our prgrammers which is specific to our needs. linguistic means of analyses most of the documents we are dealing with are full-text and to create an effective information system is impossible without the built in and complex linguistic means of processing and analysing the full-text databases; thus a sub-system of advanced automatic analyses of the full text documents becomes an mandatory component. in full-text documents special knowledge and information are accumulated in natural language. the idea of "restricted language", an approach adopted by most artificial intelligence systems, is irrelevant for the field of social and political sciences. moreover as for political text it's insignificant parts may be conveyed by so called key words. the more exact expression of the content calls for taking into account the relations between key words and the use of logical inference. so we've started research in a field of structural linguistics which is absolutely untraditional for our institute. we have had a group of professional linguists assigned to our section, without costs to us, and are enthusiastically assisting our researchers in their tasks. the description of this project may be best addressed in a seperate article by nina leontieva, who is in charge of this work. financing our work is financed by the academy of sciences of the ussr. for the past two years we applied and received additional funds to work on our linguistic means of analysis project, to invite professional linguists and designers, and additional salaries for researchers working on this project i realize how difficult it must have been for our academy chiefs to make the decision to finance fall/winter 1990 our idea of linguistic analysis and i am very pleased that they decided to finance the project. our results to date are gratifying, we have developed several functional modules of linguistic analysis and other modules are in process most notably the russian semantic dictionary where political lexics are being fully described and the creation of the thesaurus of political terms. both of them may become commercial products and that is very important for us since it helps us to earn additional money to support our project. another funding source is the "ussr congress of peoples deputies" database which was created by our institute. it is of interest to several universities in the usa and europe. version 1 contains the biographical information on deputies; education, profession, career; and information about the district the deputy represents, or public organization that elected him/her; the deputy's position in the supreme soviet and membership in committees and commissions. the database is in russian. names, geographic names, names of public organization have equivalents in english. we are planning to have version 2 available soon which will have some additional fields voting behavior of the deputies, voting results, etc. so our section of applied research and informatics after two years of looking for our own way of developing new technologies and new thinking in our institute has some results^ lot of problems, and many ideas. in conclusion i would like to express my great gratitude to all the iassist people who have made it possible for me to attend this conference it was very helpful and informative. i hope that one day soon an iassist conference may be held in moscow, ussr. iassist 90 has helped a lot to make this possible. 1 a presentation to the iassist 90 conference held in poughkeepsie, n.y. may 30 june 2, 1990. career achievement award on febuary 2 1 the second iassist career achlevment award was awarded to don harrison. judith rowe presented the award to don on behalf of our organization, at a retirement reception held in his office at the national archives building in washington, d.c. the award read, "iassist honors donald fisher harrison for his career of dedicated service to the archival preservation of computerized information. february 1991." the authorization of the award followed the procedures approved last may by the general assembly and finalized by the admin committee last october. tom brown. iassist president and a colleague of don, noted that "this marks a personal passing for me. for it was don who introduced me to bits and bytes and everything else about machine readable data files sixteen years ago when i joined the staff of the national archives. many thanks don!" iasslst members join tom brown and the staff of the national archives in extending the warmest of thanks to don for his many contributions to iassist and his profession. iassist iassist quarterly vol21.3 22 iassist quarterly outsourcing as a temporary solution the national archives of finland faced a severe problem at the end of the 1980’s. electronic records had replaced traditional paper records in many cases. though the national archives preferred and still prefers paper records to electronic ones, it was impossible to generate hard copy records from electronic records in all cases. the reason is simple. the capacity of a computer enables it to handle records which are unmanageable in paper or microfilm. it is obvious that appraisal can’t be based on the physical format of a record, but on the information the record contains. our first attempt was to outsource the preservation function to in the main computer centre of finland in 1989. information was stored in ’traditional’ way, by 1/2" magnetic tape in mainframe environment. it was, without a doubt, a safe solution. but outsourcing had its obvious disadvatanges, too. preservation was costly. the charge was 50 fim/reel each month, including backup. this meant that the cost of preservation was calculated in a different way for electronic than for paper records. this fact had an influence in appraisal practices whereby preserving data was considered as ’expensive’ while paper was regarded as ’cheap’. it meant that a great amount of electronic records was not preserved because the solution was considered too expensive. access to electronic records was difficult. it was costly to load a tape and run queries in a mainframe. mainframe serves well as a multi-user database server, but it is too robust a tool for archived material. the basic problem of archiving data is to provide access to a very large volume of very little used data. only in very rare cases do you have more than one simultaneous user for the same file. you don’t need to worry about the speed of transactions so you don’t need a mainframe. one main point was whether it was wise for an organisation to outsource its key functions. it was clear, that the use of a computer centre had to be a temporary solution. solutions in middle of an economic crisis in the beginning of the 1990’s the economy of finland suffered a major setback. in 1991, the gross domestic product diminished 7,1% and continued to diminish until 1994. government consumption expenditure sank, too. it was a crisis. for the national archives it wasn’t possible to invest in a tape repository, although the need for a preservation system of its own became evident. it was, also impossible to have more personnel. that made us think about what we actually did need. who should be able to preserver more.. if we don’t have data, then nobody asks for it, and we may think that nobody needs data archives. this circular argument was sometimes made, when we considered our task. safety is a major problem secondly, we should do it in a safe way. safety means safety in every respect, and it means both safety of material and safety of citizens, too. when we think of the long-term preservation of a record, we must take into account the ageing of the material and other such threats. also, we mustn’t forget that the society tends to change as well. let us take an example. in the year 1897 a protocol of the överstyrelse för pressärendena the supreme office of press affairs written and then preserved. it is there, in the repository even today. in one hundred years it has ’lived’ through a coup de état in 1899, a general strike and rioting in 1905, abolition of överstyrelse för pressärendena and rioting in 1917, a civil war in 1918 and series of bombardements in 1939 1940 and in 1941 1944. it has been preserved through economic crises in 1973 74 and 1992 94. if we think of the permanent preservation of electronic records, we must accept, that the future is uncertain. we hope, of course, that the next century is happier than current one, but we can’t be sure of it. in the preservation of electronic records you always need theoretical and technical solutions for preservation of electronic records in finland by matti pulkkinen * fall 1997 23 some manpower and supplies. you can’t just forget your disks or cassettes in a repository. there is a risk, because a severe economic crisis may make the performance of your routine duties impossible. if you run out of money, you may loose your data. the nature of an electronic record makes its preservation during a severe crisis a difficult task. if we think of threats like a hostile occupation or a totalitarian coup de état, the main benefits of electronic records -, easy access and rapid altering become threats. the horrors of rwanda were enabled partially by misuse of census records. . safety by replication we selected a file in logical sense as the object of preservation. a physical media can’t be preserved readable long term. file formats and physical formats become technically obsolete. a record is a semantically coherent set of information, a semantical entity, and not a physical one. it is, however, true that a record can’t exist without some physical container for the information. as a corollary we can say, that you can use different types of containers for a record. this means in practical level that we have freedom to use any media we want to preserver a record as long the semantical integrity of a record can be guaranteed. we can replicate electronic records. indeed replication is a key technique in our concept of security. we always use two different media in preservation. we take a master copy in dat and a backup in 8 mm, and if we want, a cd rcopy, too. dat copy is preserved in national archives, backup in high security deposit in another place. the choice of storage media was dictated by the lack of money we couldn’t possibly construct a traditional repository for 1/2" tape. by selecting a more compact media, we lost perhaps some of proven security of traditional 1/2" tape. we can compensate for it by a faster cycle of recopying of records once in five years. for data cassettes, it is sufficient to use a safe, secure storage vault as a repository. it is difficult to gain illegal access and it will keep the temperature and relative humidity on a stable level. inside the vault the data cassettes can be stored in fireproof cupboards to secure them from fire, water and magnetic interference. there has been, alas, a difficulty with our use of cupboards. relative humidity has been too high, probably because the insulation material contained wasn’t perfectly dry before we started to use them. however, these cupboards are easy to evacuate, if necessary. changes of temperature are very slow inside the cupboard, and we can transport material for example in an open lorry in wintertime. downsizing with unix as a computer environment we use ibm rs/6000 c20 machines. the operating system is aix 4. there are 1/2" tape-, qic-, 8-mm exabyte-, cd rand 2 dat devices in the same machine, and we are about to connect a 3480device in this machine, too. for security reasons, this machine works as stand alone, but we can temporarily connect it to our network, if we want. unix has been an optimal platform for archival preservation. it contains powerful script language, ccompiler (unfortunately an option in aix 4) and all the device drivers we need. aix is easy to manage. you can do a lot with unix programs like cut and grep; you don’t need to load your data for example in a relational database to query it. the cost of investment for the construction of the vault and the purchase of storage cupboards and computer equipment is the equal of the cost of outsourcing the electronic records preservation function over a three-year period. the value of records the values that inhere in basically any records are of two kinds: primary values for the originating agency itself and secondary values for other agencies and private users. records are, of course, originally created to fulfil the first of these: they are created to accomplish those purposes for which an agency has been created administrative, fiscal, legal and operating, if we are referring to public records. but in archival perspective, more important is the second meaning records are, after all, preserved in an archival institution because they have values that will exist long after they cease to be of current use, and because their values will be for others than the creators. these secondary values of records can be furthermore divided into two kinds: 1. the evidence they contain, i.e. evidential value, referring to the value that depends on the character and importance of the matter evidenced. 2. the information that they contain, i.e. informational value, which may relate, in a general way, either to persons, or things, or phenomena. in modern archival science the evidential value is regarded as the principal value of a record. there has been, however, little or no discussion about the nature of these two values. a closer study of both these aspects of preservation is important, and especially when we are discussing the preservation of electronic records. informational value the semantics of informational value is quite clear. we can here adopt the so called ’the naive theory of truth’ on the semantic issue. from an epistemological point of view this theory in itself is not sufficient, but altogether it is ’good 24 iassist quarterly enough’ to be used in logic. this theory of truth implies, that, for example, the sentence ’there is an a’ is true only and only if ’a’ is a really existing entity. if there is no entity called ’a’ in the world, the sentence is obviously untrue. according to what we said earlier, we can conclude, that when any record includes true and meaningful statements, it also contains informational value. so if a record states, that there exist or has existed in the time that record in question was created an ’a’, and if this statement is true it can be verified to be true -, the particular record in hand contains unquestionable informational value. evidential value and speech act theory we can’t, however, adopt this theory of truth used above to the matter of the evidential value of a record. while a receipt can be genuine or not, but semantically it can’t be said to be either true or untrue. in fact, the evidential value is always bound in the use of the particular record the record does not have evidential value per se. the semantics of evidential value has to be derived from so called ’speech act theory’. this theory states, that a performative act of speech is neither true nor untrue. a more coherent way to make the distinction is to use the terms successful and unsuccessful. a judge, for example, can impose a penalty on a criminal in the court. he has the authority to do so. in this context his sentence ’the court fires you the sum of 1000 pounds’ is successful, when and if the judge follows formalities stated by the law. the judge can’t use the power given by the law to him outside of the court. so his sentence of punishment will be unsuccessful when imposed outside of the court. some conclusions we can now argue that the very nature of evidential value is based on semantics like this. it implies that we can use binary logic as a tool of falsification of evidential value. if and when we accept the argument generally shaped in this paper, we can furthermore define some formal criteria for the evaluation of evidential value. some essential although maybe not all questions of evidential value can be reduced in the two following concepts of authenticity and integrity. the first of these, the problem of authenticity, is one of the key problems concerning the preservation of electronic records. this question can’t be solved with those techniques used in data transfer. for example public key encryption is highly software-bound and sometimes also hardware-dependent. in archival preservation we cannot and we should not rely on public key encryption techniques. the other key issue in preservation, i.e. integrity, can also be easily violated, for example by loosing referential integrity or transactional integrity. so we can state, that in the strict sense evidential value is lost when authenticity or integrity is lost. many preservation programs for electronic records seem to be unaware of these problems mentioned, and to which should be given more investigation. * paper presented at iassist/ifdo ‘97, odense, denmark, may 6-9,1997. matti pulkkinen, kansallisarkisto. vol273.indd iassist quarterly fall 2003 5 by by susan noble, keith cole, celia russell, jim schumm, nicholas syrotiuk, gindo tampubolon * delivering the world: the establishment of an international data service abstract in this paper we describe esds international, a new data service providing access to the major socioeconomic databanks produced by intergovernmental organisations such as the world bank and the united nations. through the new service, these important databanks are delivered over the web, free at the point of access to the uk academic community. the paper discusses the principles behind the service, the data acquisition strategy and the establishment of licensing agreements with the data providers. the delivery software and the development of a user interface are described and we report on the challenges of converting large and complex datasets from a range of sources into a single user-friendly format. in addition to the data delivery, a pilot web-based data exploration and visualisation interface has been developed to encourage the use of the data in learning and teaching. finally, our paper outlines the strategies and value added services employed to promote the use of these previously under-utilised databanks across a broad range of social science disciplines. keywords: international data, social sciences, macroeconomic databanks, economic and social data service, esds international. introduction june 2003 saw the launch of the economic and social data service (esds), a national uk data service which provides preservation, dissemination, user support and training for an extensive range of key economic and social data, both quantitative and qualitative, spanning many disciplines and themes. the service, funded by the economic and social research council (esrc) and the joint information systems committee (jisc) comprises four specialist data services: esds government; esds international; esds longitudinal and esds qualidata. it has been established as a distributed service, based on a collaboration between four uk centres of expertise – the uk data archive (ukda) and institute for social and economic research, both based at the university of essex, and manchester information and associated services (mimas) and the cathie marsh centre for census and survey research, both based at the university of manchester. this paper describes the development of esds international, the specialist data service which provides the uk academic community with free online access to the statistical databanks produced by international governmental organizations such as the world bank and united nations. the service also provides access to international survey datasets. background international governmental organisations (igoʼs) such as the international monetary fund (imf), the united nations, the world bank and the organisation for economic cooperation and development (oecd) have long produced high quality, regularly updated time series databanks for their own internal use. these databanks typically contain a huge range of macro-economic and social indicators aggregated to national or regional level and collectively cover virtually every country in the world. recent years have seen a growing requirement from the academic sector for access to these types of data. as the integration of national societies and economies accelerates, the accompanying issues of growth, power and inequity have attracted increasing interest from the research community. access to international databanks also allows researchers to make cross country comparisons and interpret their findings in a broader perspective. moreover, in social science fields such as crime, employment or economics, a research question that may have been a single country study a few years ago now requires examination in an international context. many of our biggest challenges such as climate change, infectious disease, global security or other ʻproblems without borders ̓can only be addressed at an international level and the academic sector requires an evidence base in order to contribute and comment on the trans-national policy responses to these collective global problems. esds international was established to address these needs through the provision of free web-based access to a portfolio of authoritative, high quality international databanks. the igoʼs have a presence in every country in the world, the authority to create international standards and the technical and financial capacity to support the 6 iassist quarterly fall 2003 development of national statistical infrastructures. the result is that the igoʼs produce databanks of tremendous quality, scope and potential value to the academic sector. however, until now these databanks have largely been unavailable or very expensive for people outside these organisations to use with only those academics belonging to an institution with an institutional subscription to each igo having access to the data. the establishment of esds international gave us an opportunity to remove some of these barriers to use by adopting a new and unique approach to the delivery of international data to the uk academic community. in addition, the creation of this service in the uk was possible due to a number of timely factors: − an increasing demand for, and desire of the igoʼs to become more transparent in their decision making; − a recent move by the igoʼs to the production of datasets in a format which can be converted for web delivery; − the desire of funding bodies, such as the esrc to expand research into international issues. data acquistion strategy the strategy developed by esds international for the acquisition of data aimed in the first instance to ensure that the uk academic community would have continued access to regularly updated versions of the international macro datasets previously available through mimas, rcade (a discontinued european data service) and the ukda. this initial data portfolio had a strong macro-economic theme in keeping with the data providers ̓primary concerns. it covered topics such as economic performance, trade and international flows of capital and included the imfʼs direction of trade statistics, international financial statistics, balance of payments, government finance statistics, the oecdʼs main economic indicators and unido industrial statistics and demand supply databases. additional funding from the esrc enabled the service to enhance the data portfolio to include datasets on the social effects of greater global interdependency, economic growth and changes in public policy priorities. these datasets cover topics such as human development, globalisation, migration, labour markets, social expenditure, demography, environment, education and science and technology. a variety of selection criteria were used to evaluate whether a dataset should be included in the portfolio. a literature survey using the john rylands university library of manchester (jrulm) e-journals service was carried out on leading journals covering a range of social science disciplines and where relevant, the empirical data section of each paper was reviewed to identify any international datasets used in the research. a data mapping exercise provided content mapping between the topics covered in each potential dataset, the esrc thematic priorities and the 19 subject categories used by the council of european social science data archives to classify socio-economic datasets (see appendix 1). an esds user consultation survey identified international datasets that potential users of the service would like to gain access to. priority was given to those ʻresearch quality ̓datasets, that is those that are regarded internationally as being key sources of high quality, authoritative, reliable and up-todate statistics with good temporal and spatial coverage. the databanks were to provide long term potential for research and teaching, with long and consistent time series, relatively stable data domains and strong opportunities for comparable research. final consideration was given to those datasets where prohibitive data license costs were previously an obstacle to use. the data acquisition strategy identified a data portfolio that collectively would chart over 50 years of global economic, industrial and social activity. the esds international macro data portfolio is shown in appendix 2. through a series of reciprocal agreements with data archives worldwide, esds international also provides access to international survey datasets such as eurobarometer, international social survey programme and the world and european values surveys. these micro level datasets cover a range of social science topics including household and demographic information, income, employment, education and housing. data re-distribution licensing one of the key achievements of esds international has been the successful negotiation of uk wide academic data redistribution agreements with a number of igos including the international monetary fund, the world bank, the organisation for economic co-operation and development (oecd), the united nations, and the international labour organisation (ilo). in many instances, this was the first time that an igo had agreed to country wide data redistribution agreements for academic access to their databases. previously, in order for a uk academic to obtain access to sourceoecd, un common database and the world bankʼs world development indicators and global development finance statistics, their institution would have to have taken out an institutional subscription. the requirement for institutional subscriptions, which have to be funded by the library or a department, has significantly reduced the number of uk academics who could have iassist quarterly fall 2003 7 access to key international datasets and it also made it difficult for academics from different institutions to undertake collaborative research. for example, although the un common database is a relatively new product, there were in may 2003 a total of 185 academic subscribers worldwide of which 81 were in the us with only 2 academic subscribers in the uk. the aim was to negotiate a uk wide redistribution agreement that would deliver significant savings to the community as a whole and remove one of the major barriers to use of ʻcommercial ̓datasets in research and teaching. in some instances, additional discount was provided by negotiating a five-year deal with a single up-front payment in year 1. indeed, the consortium purchase approach for the portfolio has produced significant savings and opened up access to international macro data to the entire uk academic community. for the bulk of the licence agreements, esds international employed the services of databeuro a professional data negotiation agent. databeuro (http://www.databeuro.com) were selected as they already had expertise in working with intergovernmental organisations (igoʼs) such as oecd, imf, world bank, and un for access to on-line data and ebooks to the academic, government and corporate sectors. where possible, the data redistribution agreement negotiated with each igo was based on a model licence. the model licence adopted was based on one developed between the jisc and the publishers association for the licensing of commercial datasets (http://www.jisc.ac.uk/ index.cfm?name=wg_standardlicensing_report). the model licence set out generic terms and conditions of use (e.g. use of data for educational purposes only) but there was also scope for adding additional clauses to reflect special conditions required by different suppliers. for example, the world bank requires a user to obtain written permission if they wish to use more than 10 indicators in teaching materials. in circumstances where an igo had its own model licence agreement (e.g. oecd) it was necessary to get them modified to ensure that they were in line with all the other agreements in terms of key definitions and conditions of use. as the licence agreement was between the university of manchester (licensee) and the igo (licensor), it was necessary to ensure that the university was not exposed to any financial risk by taking on the role of licensing the data on behalf of the entire uk academic community. as a result, it was necessary to ensure that appropriate clauses relating to limitation of liability and dispute resolution were inserted to protect the university of manchester. delivering the data complex, non web-based interfaces have long been associated with accessing international databanks. mimas for example had previously offered a limited international data service providing access to datasets such as the unido industrial statistics via a complicated x-windows interface which meant knowledge of the unix operating system and sas statistical software package were essential in order to access the data. each igo provides custom interfaces to their datasets so a user wanting to access data from a number of sources would have to navigate a number of different interfaces. a key element therefore, to our data dissemination strategy was the provision of a web-based, user-friendly interface which would be common to all databases within the portfolio. it was also decided to adopt a proprietary solution to minimize interface development and maintenance therefore giving us more time to concentrate on value added activities. beyond 20/20 web data server (wds), a web-based data dissemination tool, was chosen as the software solution to provide our common user interface. it requires only a standard web browser, is accessibility compliant and can be used to display, subset, visualise and download time series data. in addition it is recognized within the uk social science community as the office for national statistics use the same interface to deliver their neighborhood statistics. in order to deliver the data via beyond 20/20 wds it was first necessary to convert the large and complex files from a range of sources into beyond 20/20ʼs own table format. this process was problematical as data was provided in various formats with differing quantities and quality of explanatory documentation which, in some cases proved difficult to match with a generic product like wds. the first stage was to understand the contents and structure of the datasets by analysing the data providerʼs original data files, documentation and user interface. once the structure of the database was understood it was possible to interpret how this could be presented. data processing programs were written to re-format the raw data, to load the data and to publish the tables on the web server. each dataset consisted of a unique set of data files, requiring a unique set of processing programs which varied considerably from dataset to dataset. a further complication to the delivery of the international data was the requirement to restrict access to uk academic users only. users are required to register and to authenticate before they can access the data. registration is completed by accepting the esds end user license conditions and authentication is provided by eduserv athens using their access management system. a key element of esds internationalʼs data delivery strategy was to enable users to search across the full portfolio of international datasets. however, as the data comes from various data providers with no common way of describing their data, it was necessary to create collection 8 iassist quarterly fall 2003 level descriptions of the datasets. this was accomplished by creating metadata records containing subject headings and classification numbers. these were all assigned using international metadata standards, including dewey decimal classification, library of congress subject headings (lcsh), humanities and social science electronic thesaurus (hasset) and the unesco thesaurus. users can currently search in the ukda catalogue (to be replaced by the esds catalogue) across all esds internationalʼs datasets and within the mimas metadata database across all of mimasʼs services. building a user community another key activity in the establishment of the international data service has been the promotion of international data use to raise awareness of the datasets and their research potential. awareness days and workshops have been organised by esds and the service ensures it has a presence at relevant social science events. introductory and advanced training courses on data analysis are provided and an online jisc mailing list is used to disseminate information to users of international data, including details of new data releases, updates, courses, workshops and events. an esds user consultation exercise carried out in march june 2003 highlighted a potential unmet demand for access to high quality international data for use within the academic community and user statistics so far collected indicate that esds international has been successful in meeting this demand. dataset usage logs have revealed a new and growing community of international data users with the release of new datasets (i.e. unido data in november 2003, imf data in december 2003 and world bank data in february 2004) generating a substantial increase in users. over 1300 users from 124 uk further and higher educational institutions have accessed the international macro datasets in the period of june 2003 to april 2004. one of esds internationalʼs aims was to attract users from a broad range of disciplines both in fe and he and analysis of the esds international user-base allows us to see who is actually using the service. appendix 3 lists the number of users from each discipline as recorded in the esds user registration database. as expected economists are the largest single user group but users from other subject areas such as accounting and finance, environment, geography, sociology, politics, international studies, history, pure mathematics, health, asian studies and engineering are also represented. data librarians and data support staff are the second largest group highlighting the key role they have in promoting and supporting use of the service locally. the majority of our users come from higher education institutions, but a very small percentage of users (1.5%) are from further education institutions. the table below presents a breakdown of the types of users accessing the data showing that post-graduates are the largest single group with student users in total forming over 60% of the registered users. value added activities in order to further promote the use of international databanks and increase the serviceʼs effectiveness across a broad range of social science disciplines and user types, a number of ʻvalue added ̓services are offered by esds international. this includes the provision of specialist advice via a dedicated helpdesk where the teamʼs familiarity with the contents and structure of the various databases within the portfolio enables timely responses and has proved integral in building serviceuser relationships. online teaching and learning materials, a structured guide to freely available international data and frequently updated questions are provided from the serviceʼs website. a programme of introductory and advanced training courses esds registered users (jun 2003 – feb 2004): user type no. of users (836 in total) % of total staff – class tutor 110 13.2 staff – other 210 25.1 postgraduate student 334 39.9 undergraduate student 156 18.7 student – other 23 2.8 academic visitor 3 0.3 common gis cross classification map for europe. y axis = death rate, x axis = birth rate iassist quarterly fall 2003 9 is provided and can be carried out at individual institutions on request. in addition to the data delivery, the web-based geographical information system software commongis has been used to create an interactive exploration and visualisation interface to a set of freely available international data (cia world factbook 2002) to encourage the use of the data in learning and teaching. this has proved very popular at esds international workshops and our aim is to produce a number of themed gis interfaces to the datasets we host using world boundary data. conclusion and the future for esds iinternational this paper has presented the approach taken by esds international to provide the academic community in the uk with access to regularly updated, high quality international data. having only been in operation since june 2003, the service is in its infancy and there remains work to be done. the service will complete delivery of its current data portfolio during 2004 and monitor demand for international datasets working closely with esrc to identify any additional international datasets that might be required to support future research programmes. the service will look at published research based on the databanks to discover how the data provided by esds international is being used and the types of research questions being addressed through the use of international data. promotional activities and materials will continue and be enhanced, for example with the provision of themed guides on concepts implicit in the databanks such as economic stability and growth and the development of support materials on the possibilities of linking macro and micro international datasets. a review of requirements to support teaching use of the service will be carried out and used as the basis for developing a guide to creating teaching datasets and using international macro data in teaching. there will be a particular emphasis on raising awareness of the use of international data within teaching in order to encourage a new generation of expert data users with the knowledge and skills to handle international data resources with confidence. appendix 1: esrc thematic priorities and cessda subject categories used in the content mapping exercise economic and social research councilʼs thematic priorities: the thematic priorities enable the esrc to respond to the most pressing issues facing the uk. because they are based on the views of a wide range of people and sectors, they ensure that the council s̓ activities are relevant to today s̓ problems, and that there is increasing ʻknowledge transfer ̓ between social scientists and the users of their research. the thematic priorities were first developed by the esrc in response to the science white paper ʻrealising our potentialʼ. the white paper wanted to tackle the problem that while world-class research was undertaken in the uk, it was not fully exploited, either commercially or in the public interest. all the research councils were required to address this problem by working more closely with users of research and introducing the criterion of ʻrelevance ̓more clearly and strongly into their funding decisions.ʻ thematic research priorities 2000, economic and social research council http://www.esrc.ac.uk/esrccontent/publicationslist/thematicp/themefirst.html the seven thematic priorities are: 1. economic performance and development 2. environment and human behaviour 3. governance and citizenship 4. knowledge, communication and learning 5. lifecourse, lifestyles and health 6. social stability and exclusion 7. work and organisations cessda socio-economic subject categories the 19 subject categories used by the council of european social science data archives (cessda) to classify socioeconomic datasets: 1. economics 2. trade, industry and markets 3. labour and employment 4. politics 5. law, crime and legal systems 6. education 10 iassist quarterly fall 2003 7. information and communication 8. health 9. natural environment 10. housing and land use planning 11. transport, travel and mobility 12. social stratifications and groupings 13. society and culture 14. demography 15. social welfare policy and systems 16. science and technology 17. psychology 18. history 19. reference and instructional resources appendix 2: esds international macro data portfolio imf direction of trade statistics* imf balance of payments statistics imf government finance statistics imf international financial statistics* unido industrial statistics databases* unido demand supply databases* national statistics time series data** oecd main economic indicators* oecd international development* oecd international direct investment* oecd international migration statistics* oecd main science and technology indicators * oecd measuring globalisation statistics* oecd statistics in international trade in services oecd statistics on value added and employment oecd social expenditure statistics* oecd quarterly labour force statistics* un common database ilo key indicators of the labour market world bank world development indicators* world bank global development finance* eurostat new cronos *those datasets highlighted in bold are currently available via beyond 20/20 wds, (may 2004) **the national statistics time series data is delivered via searchns – a bespoke piece of software developed inhouse at manchester computing. appendix 3: esds international registration statistics – discipline 836 users registered with esds in period from june 2003 – february 2004. economics and econometrics (264) library or data/information (118) economics, labour and employment (104) business and management (70) business studies and accountancy (44) geography (26) accounting and finance (22) politics and international studies (22) economic and social history (16) sociology (14) environment, housing and planning (13) social administration and social policy (9) computing service (8) education (8) statistics, computing and methodology (7) european studies (6) political science and international relations (6) history (5) administration (4) communication, cultural (4) computer science (4) engineering (4) manufacturing engineering (4) psychology (4) asian studies (3) built environment (3) health and medicine (3) hotel, catering, travel (3) mathematics (3) statistics and operational research (3) agriculture (2) iassist quarterly fall 2003 11 chemistry (2) electrical and electronic engineering (2) environmental sciences (2) information science (2) law (2) librarianship, information (2) medieval and modern history (2) philosophy (2) anthropology (1) archaeology (1) art and design (1) biological sciences (1) biology and biochemistry (1) classics, ancient history (1) community based clinical (1) mathematics (1) mechanical, aeronautical and media studies, journalism (1) other studies and professions allied to medicine (1) research council (1) social anthropology (1) socio-legal studies (1) town and country planning (1) * paper presented at the iassist conference, madison, may 2004, by susan noble (mimas, manchester computing, university of manchester, kilburn building, oxford road, manchester, m13 9pl united kingdom). contact: susan.noble@man.ac.uk. (web-site: url: http:// www.esds.ac.uk/international). 6 iassist quarterly 2013 iassist quarterly guest editor’s notes the evolution of a special issue of the iq in honor of sue a. dodd: how remembering a pioneer data librarian became a tribute to the contributions of iassist over the decades in the months before the 2012 annual iassist meeting, libbie stephenson contacted me to toss around ideas for some type of memorial we might plan for the coming meeting to honor our friend and colleague, sue a. dodd. sue had passed away in october 20101. libbie spoke of the important role sue had played as a mentor in her data librarian career. she talked about how with the passage of time, newer members of our professional community might not know about the early days of iassist nor of the important contributions made then by sue and others. i agreed and we decided to meet during the conference to discuss all of this further, allowing time for our ideas to percolate. when we met, quite informally, our discussion included others from sue’s cohort: tom brown, carolyn geda, and judith rowe among them. we agreed on the value of gathering original essays that would consider dodd’s contributions as a conceptual starting point, especially regarding data description broadly conceived. we decided that in moving beyond this, the essays should address the history and development of related topics of special interest, reflecting the expertise of the authors. subsequently we discussed our ideas with the iq editor, karsten boye rasmussen, who was fully supportive. i volunteered to guest edit the collection and libbie stated her interest in preparing a bibliographic essay on sue’s publications. in subsequent months we contacted colleagues whom we thought might be interested in participating in the project and were pleased at the responses. as can be seen by a quick scan of the table of contents, they were so positive that our initial goal for “an iq issue” grew to constitute the whole of iq volume 37 (2013)2. among the numerous themes that run through these essays, the collegiality within iassist is ever present. perhaps the most interesting trait the essays have in common is the unique way each alludes to dodd’s contributions. in explaining our concept of the project to potential authors, we did not expect that they would explicitly tie their work to hers, yet each essay does so in one way or another. among the authors who knew sue, some of the commentary reflects the personal as well as the professional. two of the authors, ann gray and jonathan crabtree, share experiences of working at the institute for research in social science at the university of north carolina, sue’s professional home for more than 30 years. peter burnhill charms with memories of his first encounters with sue while giving us a glimpse of the early days of the edinburgh university data library and then edina, and his almost 30 years-later encounter with her writing as he embarked on new ventures related to digital preservation of what he terms “scholarly statement.” several authors note both implicitly and explicitly the manner in which ideas in the 1960s – 1980s were both prescient in terms of challenges still vexing the social science data community and foundational in terms of setting the course that has influenced the collaborative efforts of the international data community ever since. one of the more exciting results of the project is that it offered a venue for an exclusive: a thorough telling of the history of the data documentation initiative (ddi). the comprehensive story of the ddi emerges for the first time in the essays contributed by karsten boye rasmussen, ann green and chuck humphrey, and mary vardigan, and through the detailed and richly documented timeline of the ddi prepared by mary vardigan. taking us from the early study description work at some of the european data archives and the contemporaneous documentation standards efforts in north america, their essays give us an overarching perspective on the context, motivation and requirements behind the design and development of metadata standards, the ddi, and the infrastructural and technological changes that contributed to the contemporary maturation of the ddi. in addition to a glimpse of the legacy of dodd at the university of north carolina, jon crabtree’s essay explores some of the more contemporary collaborative ventures of the wider data community, especially the data-pass (data preservation alliance for the social sciences) project. in the following essay micah altman and mercè crosas take us fully into the present and look to the future, discussing a range of work related to data citation in the context of contemporary frameworks for data management and research. hailey mooney brings the essays full circle with her discussion of iassist’s special interest group on data citation and its development of the quick guide to data citation. sue dodd would be very pleased to see this guide, we think, as it seems to be exactly the kind of product that she expected of future generations of data professionals, as expressed in the quote by which libbie stephenson closes the introduction of her comprehensive bibliographic essay. editing this volume of the iq was a stimulating and enjoyable experience; working with each of the authors introduced me to new facets of our shared experiences and interests in unexpected ways. my thanks to each of the authors for their enthusiasm, commitment, and iassist quarterly 2013 7 iassist quarterly patience. individually and as a group they personify what iassist has stood for all these years: collegial collaboration of social science data professionals across the world. that’s what made this endeavor so interesting: i discovered that in honoring the memory of one of our early members, we actually pay tribute to the many contributions of all of iassist. and as a bonus, the essays make for a good documentary read. thank you, all. margaret o’neill adams (peggy adams) april 2014 guest editor peggyoadams@yahoo.com notes 1. sue anna dodd, b. 6/18/1937 in lexington, ky; d. 10/6/2010 in pittsboro, nc. 2. there is an additional area of sue dodd’s professional contributions that is not covered in these essays because the work involved was outside their topical scope of consideration. however, in keeping with one of the goals for this project -to inform readers of sue dodd’s contributions -we add this note. building upon the work that her publications represent and her outreach to the professional library community, dodd was among those who also contributed to the professional community of traditional archivists. she participated in training workshops for them, focusing especially on the description and documentation of archival machinereadable records. most importantly and indicative of her civic commitment, she completed a multi-month consultancy during 1986-87, at the [u.s.] national archives and records administration. she produced a substantial but unpublished report that covered her study of computer records and the federal records management program, and case studies of the computer records in two federal agencies: the bureau of the census and the department of state. in the report she identified challenges, prioritized problems, and presented a set of recommendations, with cost estimates, to improve the management of computer records by the national archives. (dodd, s.a. “computer records and the national archives: an assessment with new directions,” 1986-1987) iassist quarterly vol 24 no. 3 8 iassist quarterly fall 2000 the social science electronic data library: serving the needs of data librarians and users by michael carley &. josefina j. card * the last decade has witnessed enormous strides in two areas: first, the development of numerous social science data sets of high quality; and, second, the development of the computing hardware and software capability and infrastructure needed to locate and analyze these data sets for minimal cost and to communicate data findings in interesting and easy-tounderstand fashion. hand in hand with these advances in data development have come technological advances which allow social science research and teaching laboratories, with the hardware and software needed to analyze the best data in a given field, to be set up with ease by an academic department or even by an individual professor. additionally, sophisticated data analysis software packages formerly available only for mainframe computers have become available for microcomputers at a much reduced cost. taken together, these developments make it possible for academic departments, research institutes, and government offices of all sizes and levels of financial resources to access and analyze exemplary data sets for research, teaching, and programand policy-development purposes. data archives, in both the private and public sectors, allow easy and open access to many hundreds of the best health and social science data sets covering a broad range of topics, study populations, and making use of a variety of research designs. the data available from many of these archives are clean and the documentation user-friendly. for these and other reasons, researchers and instructors who are considering the use of secondary data welcome the functions served by a well designed data archive. this data is used to: • conduct secondary analyses of outstanding data sets to serve the needs of policy, practice, or basic research; • perform meta analyses based on access to multiple original raw data sets; • prepare research proposals on various issues; • write publications comparing and contrasting results from related data sets; • prepare masters theses and doctoral dissertations; • produce classroom materials for teaching substantive, methodological, and statistical concepts from real-world data. in this paper, we will discuss three areas of major concern for data librarians and data users wishing to obtain and use data for secondary analysis: data quality, format, and dissemination. we review and contrast the issues and concerns of data librarians and data users, outlining areas of similarity and difference. we will explore how one large data collection, the social science electronic data library (ssedl), compiled over the last 17 years by sociometrics corporation, has addressed each of these issues and the conflicts and problems that arose during that process. finally, we peer into the future, assessing how data providers can bridge knowledge gaps via recent technological advances. the social science electronic data library (ssedl) the sociometrics social science electronic data library is a premium health and social science resource that consists of seven topically focused data archives. with over 300 data sets from 200 different studies comprising seven topically-focused collections, it is a unique source of high quality health and social science data and documentation for researchers, educators, students, and policy analysts. the electronic data library was made available in 1999 on a set of cd-roms and includes an online membership with free access to datasets for downloading by members. the collections: the data archive on adolescent pregnancy and pregnancy prevention (daappp) was established by the us office of population affairs (opa) in 1982 as the repository for the best social science data on the incidence, prevalence, antecedents and consequences of teenage pregnancy and family planning. in 1994, the scope of daappp was expanded to include studies that focus more broadly on adolescent sexual health issues, thereby including studies examining behavioral factors related to sexually transmitted diseases (stds) in addition to pregnancy. daappp currently holds data from over 150 premiere studies (many of them longitudinal) on sexuality, health, and adolescence. state-of-the-art research data on the american family are iassist quarterly fall 2000 9 available through the american family data archive (afda). afda, funded by the national institute for child health and human development, contains data and documentation from 20 nationally recognized studies on important issues relating to american family life, demographics, and family patterns. among the topics covered are educational, economic, health, social, and psychological indicators, child welfare, family violence, marriage, divorce, child care and child custody. the aids/std data archive (aids) consists of original research data and instruments from 11 premier studies on aids/hiv and other sexually transmitted diseases (stds). the collection was established with funding from the national institute of child health and human development (nichd. included data sets address the following topics: the incidence and prevalence of specific sexual behaviors (including abstinence, vaginal and anal intercourse, oralgenital sexual activity, masturbation); contraceptive and std-preventive behavior; attitudes and beliefs regarding sexual behavior and methods of contraception and std prophylaxis; aids/hiv knowledge, attitudes, behavior, and serostatus; current and past episodes of gonorrhea, syphilis, chlamydia, and other stds; and high-risk behavior, including alcohol/drug use and prostitution. the maternal drug abuse archive (mda) brings together seven state-of-the-art research databases on maternal alcohol and drug abuse. funded by the national institute on drug abuse, the collection includes data on the following topics: the prevalence of drug use among pregnant women and women of childbearing age; demographic characteristics of pregnant drug users; types and patterns of illicit drug use; social, psychological and economic antecedents of preand perinatal drug abuse; the effects of preand perinatal substance use on pregnancy complications and neonatal status; and the effects of fetal alcohol and drug exposure on children’s physical, neurobehavioral, psychological and social development. the data archive of social research on aging (dasra) was assembled with the support of a grant from the national institute on aging. dasra contains data and documentation from three very large nationally recognized studies. these three studies covered a variety of topics including functional status and impairment, living arrangements, caregiving and social support, health attitudes, retirement income and plans, mortality, health, financial resources and assets, expenditures, cognitive ability, medical conditions, housing, health insurance, and personal characteristics the research archive on disability in the united states (radius) was funded by the national center for medical rehabilitation research (ncmrr) within the national institute for child health and human development (nichd). the purpose of the project is to facilitate access to the best data sets on the prevalence, incidence, correlates, and consequences of disability in the u.s. the heart of the archive is a collection of 19 studies that address the topic of disability. these data sets permit analyses on topics such as: the incidence and prevalence of specific diseases, disorders, and impairments, including deficits of cognition, emotion, physiology, and anatomical structure; functional limitations across a variety of specific organ systems; disabilities in relation to major life roles and activities, such as work, parenting, education, and recreation; societal limitations including physical, attitudinal, and economical barriers that restrict full participation in society; psychosocial and interpersonal factors such as coping with stress, sexuality, feelings of control and productivity, quality of life, and family relations and support; health care and rehabilitation issues such as medical costs, coverage, service utilization, use of orthotic, prosthetic, assistive devices, effectiveness of rehabilitation; as well as a variety of basic demographic factors on respondents such as age, race, sex, income, occupation, marital status, family size, and living arrangements. to facilitate access to the best contextual data, sociometrics has developed a contextual data archive. by contextual data we mean data that describe the population, social, and economic characteristics of geographic areas, from census tracts to states, in which people reside or work. the contextual data archive consists of a series of files, each organized around a different geographic unit of analysis (such as census tracts, school districts, counties, states, etc). each file contains variables drawn from various sources, but having one common geographic unit of analysis. support for this project was provided by the national institute of child health and human development. data quality both data librarians and data users have an abiding interest in the availability of high quality digital data. the librarian must put her/his limited resources to the most efficient and effective use possible. the resources we mean here are not only financial assets such as approved budgets, but also material and human capital such as shelf space and the person-hours of those who must purchase, assemble, and maintain various data collections. because all of these resources are limited and precious, data librarians must make wise choices as to how to prioritize their use in order to achieve the highest quality collection of data for the users at their institutions. data users are also concerned about having data of the highest possible quality. the users of digital data wish to make the best possible contribution to the body of knowledge in their field. that contribution is placed in jeopardy if the data used is of questionable quality. data gathered via poor research design, or via a good design 10 iassist quarterly fall 2000 poorly executed are of little use in advancing knowledge. researchers using such data risk not only making a tainted contribution, but also of generating criticism from colleagues and associates who recognize the problems or limitations of the data being used. the research staff at sociometrics recognized these needs when compiling the seven data archives in the social science electronic data library. it was determined that each archive would be a ‘best of the lot’ collection, accepting only the best data available in each of the seven topic areas. to accomplish this, we formed national advisory panels of research scientists who were experts in both the substantive content of the particular archive and the research methods commonly applied in that field. the panel, usually consisting of six members, was asked to evaluate candidate data sets on the following five criteria: • technical quality: among the factors to be considered are high response rates, low attrition rates, use of reliable and valid measures, and sound sampling and design elements. • substantive importance to the field: factors include the potential to address contemporary issues, to break new ground, and to replicate or confirm important findings. • program or policy relevance: the ability of the data set to answer applied questions on how to improve public policy or shape intervention programs; • potential for secondary analysis, including: scope of sample the broader or more diverse the scope of the sample, the greater the potential of the data for generalization. size of sample sample size is always an important consideration. this is even more true for data intended for secondary analysis: sample sizes adequate to support the originally intended analysis may be too small to support other analyses, especially if the new analyses focus on data cells that have a very low proportion of cases. breadth of variables and constructs covered the potential for secondary analysis is directly related to the breadth of variables measured in the data set. the more numerous and diverse the set of variables, the more possibilities there are for new or expanded analyses. • disciplinary balance: an archive should attempt to be representative of the entire field of research. variations in state of the art exist between different sub-areas within any discipline. thus a somewhat flexible standard (as measured by the other criteria above) should be used to ensure that all major areas of the discipline are represented in the archive as a whole. in order to perform these evaluations, we provided each panel member with briefing materials consisting of a 2-4 page description of the data source, which covered: the purpose of the study; methods (including sampling design, periodicity, unit of analysis, response rates, and attrition); content (description of variables covered, number of variables, and topics covered); limitations; sponsorship; and a bibliography. in addition, we provided copies of original peer-reviewed publications for each data set, which allowed the panel members to review issues we may have not addressed in our briefing documents. panel members were encouraged to suggest additional data sets for consideration, and have often done so. panel members did not vote yes or no on each data set, but rather rated each with a ‘priority score’ from 1-10. only those data sets receiving an average score of 7 or above were accepted, and higher priority for archiving was given to those receiving higher scores. in addition to pre-screening the data sets, archivists for the ssedl data sets perform several other tasks designed to ensure data quality. we review the data thoroughly, checking to make sure that all variable and value labels are included and are sufficiently descriptive. we check the data for internal consistency and completeness, scanning in particular for variables with an excessive number of missing or out-of-range values. we also perform random checks verifying that the skip logic in the original instrument was followed and that the variables are consistent in relation to one another (no variables describing a female as a father or brother or a male as a mother or sister, etc). finally, we produce a user’s guide to the machine-readable files and documentation which notes any remaining limitations or inconsistencies. those archiving digital data face many challenges, not the least of which are the limitations on their own, as well as the users’ time and resources. consequently, different data archivists take a variety of approaches to address these challenges. our ‘best of the lot’ approach emphasizes quality over quantity. this means that our data collections, while not as large as some of those from other sources, are of higher overall quality and are better documented than is the industry average. this approach limits the size of our collections, but contributes to their popularity among researchers for their high quality and ease of use. formats the needs of data users and librarians diverge somewhat when it comes to format preferences for digital data collections. users look for data in the most easily accessible form, while librarians must be concerned with the big picture, and look for collections that serve the needs of as many users as possible, both in the present and the future. the format(s) in which data are provided also have an impact on the role the librarian will take in the data distribution process, which could vary from that of a facilitator to an active gatekeeper. the challenge for data iassist quarterly fall 2000 11 providers is to address both the preservation needs of the librarian and the ease of use needs of data users. distribution media. librarians must conserve their many precious resources, including both financial resources and shelf space. however, they also desire that data collections be in a form that is not easily lost, damaged, or misused by careless users. therefore, a data collection should be durable, and in a format that is not likely to change rapidly with evolving technology. cd-rom technology meets these requirements. cds are more durable than diskettes and do not have the associated (at least perceived) transient qualities of internet sites. data made available on a cdrom are not likely to be easily lost as library staff can, should they choose to, maintain tight control over them or copy their contents to a central repository for safekeeping. diskettes are more likely to become corrupted and internet sites often are revised and require more constant updating on the part of both the data provider and the librarian to keep all links accurate and up to date. data users, on the other hand, prefer that the data are made available in the simplest, easiest to access format possible. however data are made available, it must be transferred to the computer where the user will actually be working. with desktop computer speed and hard drive space increasing exponentially, users often prefer that data be available for copying to their own system, rather than residing at a central repository. internet downloads may be preferred to cd-roms as the data can be copied to the user’s own computer, then manipulated and transformed as necessary. to best accommodate these divergent needs, we found it necessary to make our data sets available in both cd-rom format and via our internet web site. ssedl volume i is distributed via 17 cd-roms, along with accompanying support material. additionally, each purchasing institution is given free web access to all of the data sets, as well as to new data that have not yet been added to the cd collection. by allowing the users to download the data sets or use the cd-roms, we were able to provide both the user and the data librarian with flexibility in both data format and data access. analytic software. a user who wishes to use data on a particular topic would be best served by data that can be retrieved quickly and effortlessly with a variety of software. given the wide variety of statistical software available to users in different fields, this can be a challenge. users should have the capacity to use the software of their choice, and the ability to access the data quickly with that software. at the same time, data providers must understand that software currently popular may change or become outdated, making files created from these programs unusable or at the least cumbersome and inefficient to use. most data collections have taken one of two approaches to this problem. first, the data provider may distribute raw data with a codebook. the raw data is typically stored in an ascii file which is simply useless text (numbers) without the codebook. the codebook provides the user with the location of specific variables and cases within the raw data file. the advantage to this approach is that it addresses the issue of durability well. users can access the data by writing a program using the statistical software package of their choice, inserting the variable and case locations given in the codebook to access the data. changes in software applications do not affect data distributed in this method, as users write their own program with the language in which they have expertise. however, writing the programs to read the raw data can be a time consuming process, causing users to waste much of their resources on mundane tasks. in addition, such writing is prone to error; one misplaced character can cause much of the data to be written incorrectly. other data providers address these issues by distributing the data in a pre-packaged format using one of the most popular statistical software packages (typically spss or sas). by distributing these formatted files (usually either complete system files or portable files), the user can access the data directly simply by opening the files in the appropriate software. when portable files are used, the data can be used with different versions of the same software package (either earlier versus later versions or versions for different operating systems) or in some limited cases, in other popular statistical packages. the advantage to this method is clear: quick, easy access to data. the disadvantage is that this approach cannot possibly be flexible enough to address all user needs. some users wish to access the data with a software package that is not among the most popular. also, data distributed in this method can become unusable when software packages radically change their formats. data made available in the most recent format could be inaccessible within just a few years. to address the limitations in each of the above approaches, data sets in ssedl are distributed with raw data files and machine-readable set-up statements for use with both spss and sas statistical software. these set-up statements provide for the best of both worlds: ease of use, combined with flexibility. users use these syntax files to create the system or portable files in whichever software they are using. for those users who are use software other than spss or sas, the set up statements serve essentially the same purpose as the codebook described above. because the set-up files are machine-readable, they can often be converted for use with other software with a minimal investment of time. should the syntax requirements of spss or sas change radically, these files would serve also accomplish this goal. 12 iassist quarterly fall 2000 in addition to the above files, ssedl data sets also are distributed with a machine-readable spss data dictionary file and an spss frequency and statistics file. these aid the user in making sure that they have created their system or portable files correctly. users can compare the statistics in the frequency file to their own, thereby preventing mistakes in data analysis. both the frequency and dictionary files are also useful in reviewing the contents of the data set. each data set is also accompanied by a printed user’s guide (provided in machine-readable form, in addition to printed form, for the more recent archives) comprised of a standard set of sections and subsections. the provision of standard machine-readable and printed documentation assists users in familiarizing themselves with the sociometrics data sets. once a user has worked with one sociometrics-packaged data set, it is easy for him or her to work with any of the others. the original instrument and codebook are offered as optional, supplementary documentation for each data set, when available. for the more recent archives, the original instrument is distributed in machine-readable form along with the data, as a set of graphics files (page images). search and retrieval software. as data sets get larger, both in the number and scope of variables covered and in the number of cases, users are faced with an increasingly overwhelming task of reviewing which parts of a study are necessary and appropriate to their needs. often, users will begin work with a data set containing over 5,000 variables (and often several thousand cases) only to find that their interests only require 30-40 of those variables. it is important for users to be able to quickly sort though the variable list and find those of interest. given current (though perhaps temporary) limitations in speed and disk space, users also need to be able to reduce the large data set into one with only those variables needed for analysis. to address this need, sociometrics staff developed powerful search & retrieval software which now accompanies each data archive. this software allows a user to search an entire topically-focused collection, a customized group of data sets created explicitly for a given user, or a single data set; to identify variables of interest across this designated search space and to save located variables as a search set. users can conduct: (1) full-text keyword searches, including variable names, words in variable labels (question descriptors), and words in value labels (response descriptors); (2) searches by assigned topic and type codes; and (3) searches by study name or assigned data set number. standard boolean operators (i.e., “and,” “or,” “not”) can be used to combine search sets. alongside this software, we provide data extract software which allows users of cd-rom versions of archived data sets to create customized spss or sas program files containing only those variables of interest to them. this capability permits analyses of subsets of large data sets to be conducted quickly (with rapid turn-around) on most microcomputers. it also saves users significant program development time writing and re-writing spss and sas program statements to define variables used in a given analysis. technical support the role of the data librarian. the role of the data librarian in this process varies a great deal among institutions. in some cases, the librarian serves as an expert gatekeeper to the data, allowing access to users as s/he deems appropriate and answering a wide range of questions users may have. others may serve a minimal role, simply providing access to the data and support materials and little else. librarians also vary in their level of statistical knowledge, as well as their expertise in the various topics that may be covered by the data in their collections. our approach to this issue was again to provide the greatest amount of flexibility in the collection as possible. while some data librarians do take on something close to a gatekeeper role, we chose to allow for those who had neither the time nor the expertise to do so. librarians need to be fully informed, not on all of the topics included in their data collections, but on the process by which they can aid users in finding data of interest to them. given the wide variety of topics covered in ssedl, one could never expect librarians to provide users with all of the help they may require. to this end, the social science electronic data library includes a variety of tools to facilitate this process. a user’s manual and contents manual detail the data sets included in the collection, and a quick start guide offers advice on how to implement the software and the knowledge needed to use the data sets. most importantly, a ‘guide to the social science electronic data library’ cd is provided with each collection. this cd takes the librarian through the process of using the data library using a brief step by step tutorial. ssedl’s research support group. in addition to the support given to the data librarian, users have direct access to help from sociometrics’ archiving and scientific staff through our research support group (rsg). the research support group consists of ph.d. and masters level social scientists who provide free technical assistance for users who have questions about accessing or using our data sets. in addition, the rsg occasionally performs consultant work such as the creation of customized data set extracts; user-defined statistical tables and analyses; data archiving, management and analysis services; customized cd-roms; and training workshops. these services greatly aid users with limited expertise or resources with which to conduct their own analyses. dissemination the manner and methods by which digital data are disseminated by data providers and eventually by data librarians are crucial to the usability of such data. users iassist quarterly fall 2000 13 must be made aware of the availability of data that meets their needs. however, with today’s rapidly expanding technologies, the problem of ‘information overload’ is a crucial one. even within our own collections, users can become overwhelmed with the sheer amount of information available to them. if these issues are not handled properly, they can inhibit the users’ ability to locate and make use of the most appropriate data. data providers must do what they can to ensure that users have the capability to quickly and easily locate the data sets and even the particular variables their topic of interest requires. while the search and retrieval software made available for each individual data archive helps users find variables within a study they have already chosen, it does not help when users have not yet selected the study that meets their needs. to address this issue, we created a search mechanism for use on our internet site which is cross archive. this allows for users to search by keyword(s), or designated variable topic or type, for variables of interest in all of the ssedl data sets. users can search for words in variable labels (question descriptors) or in value labels (response descriptors). in addition, we include a brief abstract of each study, and users can search for keywords within those abstracts. we chose to make this software available to the general public as well as users on our internet site, in order to allow researchers at nonpurchasing institutions the opportunity to find data sets of interest and order them individually. data librarians must also make an effort to make help users become aware of available data collections. in order to facilitate this process, we provided not only the above mentioned guide to ssedl on cd-rom, we provided informational flyers and brochures to help the librarian make potential users aware of the availability of our data collections. in addition, the research support group provides both librarians and users ongoing advice as to how to find data sets of interest in our collections. looking to the future: new technologies, new audiences the value of any data collection is in part predicated upon its ability to address issues of the day. therefore, any collection will inherently be of greater value the newer the data are that are contained within it. the preservation of historic data is clearly important, but any collection that aspires to be useful must also be kept up to date with the addition of more recent data. we will continue enlarging the content and capabilities of our data set collections. we will be adding to our current data archives as funds permit, and expanding our efforts by adding new topic areas to the collection. we expect to begin the establishment of a data archive on child well-being shortly. a feasibility study on the formation os a complementary and alternative medicine data archive has just been successfully completed. in putting together the ssedl, sociometrics’ staff have learned a great deal about data collection methods and ways to improve efficiency and reduce costs to researchers. currently, we are developing a software product that will aid researchers on this aspect of the process. sociometrics’ automated dataset development software (adds) is an integrated software program that, when completed, will develop and document social science research studies. the program will perform the following functions: 1) instrument generation—generate a fully formatted research instrument in print, ascii, and other machine-readable formats. 2) codebook generation—generate the data set documentation in a printed codebook (also in ascii and other formats), flow chart (skip map), and data file map. 3) data entry—provide for data entry from completed questionnaires, with simultaneous error checking. 4) program file generation—produce a raw data file in ascii format, and build the program statement files needed to transform the raw data file into spss and/or sas system files. the software will automate tasks best done by computer, improve instrumentation and documentation by providing a complete, high-quality structure and format, and reduce the post data-collection effort of documenting a public-use data set. additionally, we are also building an item bank of high quality, commonly used questions, scales, and interviewing tools from the ssedl collection. this bank will be accessible within the adds program to permit users to select questions to develop their own research instruments. the item bank will be filled with several thousand questionnaire items drawn from some of the leading studies in research on the american family. using questions or scales that have been previously tested will not only improve the choice of questions, but will also lead to greater comparability between studies and over time. in addition to keeping the social science electronic data library current and helping researchers imrpove their methods, we hope to reach new audiences through new technologically innovative products. bridging the ‘knowledge gap’ is of prime importance to those who wish to make practical contributions through social science research. rather than limiting our efforts to trained researchers, we must reach out to other professionals, and, when possible, the lay public as well. we are beginning our efforts to reach the ‘paraprofessional’ audience with two new products related to the ssedl. the u.s. social surveys: a sampler of questions and responses, will contain searchable, edit-ready, and print-ready machinereadable versions of the demographic, behavioral, and health science instrumentsæquestionnaires, medical forms, interview protocolsæused to collect the data in ssedl. questionnaire items will be linked to crosstabulations with age, race/ethnicity, and gender, obtained from the linked ssedl data archives. 14 iassist quarterly fall 2000 secondly, the multivariate interactive data analysis system (midas), will allow online analysis of the data in ssedl. online data analytic procedures will include weighted and unweighted frequencies, percentiles, and measures of dispersion and central tendency, as well as two-way and n-way tables with measures of association, comparison of means (2-group and anova) and correlations, and the calculation of complex variance estimations. users will be able to define case subsets, recodes, or aggregations for analysis, and then produce output which can be downloaded or printed. custom dataset downloads will also be available. the goal of adds is to aid expert researchers in handling the ‘front end’ of the research process. through the use of the adds software, researchers will be able to reduce their costs and improve the accuracy and efficiency of instrument development, data collection, input, management, and analysis. the goal of social surveys and midas is to increase the accessibility of the data to those who are not competent in the sophisticated statistical software packages such as spss or sas. these products will help the ‘paraprofessional’--people with college degrees who are not necessarily trained in complex data analysis— avail themselves of exemplary social science data. they will provide a basic introduction to social science methodologies as well, and will be linked to the ssedl data sets for those who wish to progress to the next stage of data analysis (e.g., advanced undergraduate students). in sum, the historic progression of ssedl has been to expand the definition and purpose of a data archive. ssedl staff have worked to enhance archiving methods to make data easily accessible to researchers. our advances in this field have helped to make the research process more efficient, especially for those conducting secondary analysis. adds will improve the process for primary research as well. non-researchers will be introduced to data analysis through the new products, social surveys and midas. together, these products will allow us to extract the greatest possible value from our research dollars as data will be used in as many ways as are feasible and by a much wider audience. * michael carley and josefina j. card, sociometrics corporation, contact name and address: josefina j. card sociometrics corporation, 170 state st. suite 260, los altos ca 94022, (650) 949-3282, ext. 211, fax (650) 949-3299, jjcard@socio.com mailtoi:jjcard@socio.com vol233 summer 1999 15 context in canada, the data liberation initiative (dli) approved in 1996 by the treasury board of canada has removed a significant obstacle to obtaining canadian data in our universities. with dli, canadian universities and statistics canada have solved the problem of obtaining canadian data at an acceptable price. i remind you that in canada, access to data is not free. in the province of quebec in particular, a number of universities obtained access to data but without really improving methods of consulting these data. it is important that people realize that there are no full time data librarians in any of the quebec universities. most universities (except mcgill, montreal and laval) are small in size, with equally small resources. we do not have a long-standing tradition of data use as in other canadian universities. this is why, in order to render microdata files more usable, university libraries in quebec have pooled their resources and expertise for the development of a common infrastructure to facilitate access and use of data. what is sherlock? no, we aren’t talking about the world-famous detective, sherlock holmes. according to the conference theme, sherlock is a kind of regional bridge to data. using sherlock, the data user becomes a detective of sorts. sherlock is a bilingual tool, designed by the numerical data file subgroup1 of the crepuq (conference of rectors and principals of quebec universities). at the conference of rectors and principals of quebec, we are a small but active group of four data librarians who are been working together since the beginning of the 90’s. we organise data workshops for our colleagues. we share our experiences and expertise. all of us are here at iassist. this paper is being read on behalf of the four of us. we are the designers and the managers of sherlock in our different institutions. the crepuq provided the place where quebec university libraries were able to initiate and discuss this co-operative project. we have a 30-year tradition of co-operation between libraries. sherlock was developed mainly for members of the quebec academic community to enable them to access and utilise the survey microdata of the dli (data liberation initiative) and the icpsr (university consortium for political and social research) data. project origin and description the first document submitted by the subgroup on numerical data files was rapport de la consultation sur l’intérêt et la faisabilité d’une approche collective à la gestion des données numériques (report on consultations concerning the value and feasibility of a collective approach to the management of numerical data,) crepuq, november 1996. our colleague chuck humphrey of the university of alberta acted as a consultant for this stage. after approving this report, the heads of the quebec university libraries asked the subgroup to conduct a preliminary analysis on a top-priority basis. the timing seemed to be right. when the subgroup took stock of data extractors in operation at the time, the landru system, developed at the university of calgary, stood out as one of the best although it did not meet all the requirements of the system to be implemented in quebec. we wanted a bilingual interface; a decentralized and distributed approach to encourage the sharing of expertise and responsibilities in many institutions; management of all survey files available in the quebec university network; compliance with licences; etc. therefore the four data librarians, who are members of the crepuq subgroup, with the help of an analyst from the library of laval university, conducted a preliminary analysis and designed a pilot project. in march 1997, the subgroup submitted its report, titled infrastructure collective pour la gestion des données numériques dans les bibliothèques universitaires québécoises (a common infrastructure for the management of microdata files in quebec university libraries) crepuq, march 1997. this report was subsequently accepted by library directors from eleven universities, and they asked to my library (laval university) to undertake the task of implementing phase 1 of the sherlock project. sherlock: a web magnifying glass for microdata files by gaëtan drolet* 16 iassist quarterly the pilot project phase 1 of the pilot project started in september 1997 and was completed in october 1998. the phase focused on developing all of the system’s capabilities and setting up a first server centre. development team the responsibility for implementing phase i of the project was assigned to the library of laval university, which established a development team made up of a project leader, the data librarian, a librarian and a computer analyst. the team’s mandate was to develop all the system’s capabilities, with bilingual interfaces, set up an initial server for a limited number of surveys, and make corrections as needed during the trial period. project co-ordination to ensure that the project went smoothly, the crepuq data librarians subgroup on data files was assigned the role of advisory committee. funding the funding for the pilot project was provided through contributions from quebec’s university libraries. twelve institutions participated in the funding of phase 1 out of a total number of 14. according to a complex formula, small universities invested less money than big institutions. institutions as clients all users of quebec universities, called client institutions, have access to sherlock, but the use of the actual survey data requires that the institution’s library be a member of the dli or the inter-university consortium for political and social research (icpsr). in addition to being user institutions, a few libraries will become server institutions. institutions as servers the management of the surveys and their files is a responsibility shared by different server centres. each institution (server centre) that has taken on responsibility for managing surveys in sherlock has designated a local manager who is responsible for the management and follow-up of these surveys in sherlock. these managers will be the only persons authorised to complete, to modify or delete a survey. a survey management module has been developed to facilitate these operations. for the implementation of phase 1 of the pilot project, only the laval university library acted as a server centre. surveys included for phase 1, fourteen surveys from statistics canada and one from icpsr were installed on the system’s first server centre. five of these support all the system’s capabilities (including extraction by variable and statistical analysis), while ten others support the basic level of use (retrieval, consultation of documentation and block files transfers, with no extraction). under access licences, owing to the number, diversity and breadth of the surveys, some files can only be downloaded as a block (ftp), with no data extraction, while others have limited access, specifically to member institutions of the icpsr. some local surveys (e.g., a survey of quebec public service retirees) could be loaded into sherlock and be accessible only to certain universities. it is the case for a survey on political attitudes done by a graduate students’ class in my institution last semester. the system access to sherlock is based on a bilingual web interface (french and english) offering a single and universal gateway to all the surveys. sherlock is accessible in quebec university libraries at the following url address [http://sherlock.crepuq.qc.ca]. the general purpose of the system is to provide for the management and optimum use of all microdata files available in the quebec university network. sherlock is not a teaching tool with a set of exercises, but it is easy to use by professors in undergraduate classes. capabilities the main capabilities of the public module are: • to provide access to the inventory and description of surveys by means of a retrieval module; • to provide the user with documentation (survey metadata) on data files (guides, user manuals or codebooks, sas or spss statements, record layouts and description of variables) when available; • to enable users to extract subsets of data files in different formats for later processing at a local workstation. intermediate and advanced users who can handle large sets of variables can download the complete dataset; • to enable users to obtain simple statistical results such as a frequency distribution, cross-tabulation, mean, median or regression analysis on a variable in using the module of analysis. more specifically, sherlock can be used • to ensure the compatibility of and access to information systems in twelve member institutions; • to make the greatest number of surveys available; • to promote the sharing of resources for data preparation, storage and use; http://sherlock.crepuq.qc.ca fall 1999 17 • to promote the development and sharing of expertise in the use of data among both the clientele and the reference staff of our libraries. computer infrastructure sherlock is a decentralized system made up of two modules: a public module and a management module. the public module is used to access web pages, conduct searches and access the forms used for retrieval and analysis. searches are conducted on a unix main server (sherlock.crepuq.qc.ca) located at the laval university library. the programs needed by the user are netscape, an e-mail software, winzip, acrobat reader, excel/sas or spss. whereas the documentation is accessible to the general public, access to data (transfer of complete file, extraction, analysis) is controlled by ip numbers, ensuring observance of licences governing use. access to metadata is public but access to file transfers, extraction and analysis is controlled. the html pages, for searching the description, the list of surveys resides on the main server (unix). the survey metadata (codebook, record layouts, sas and spss files, etc.) and data files reside in the different server centres on nt servers. the data extraction and analysis is also done on the different nt servers. extraction and analysis operations use perl procedures. management module the management module is used for the capture of data (description of surveys, metadata, and data files) from surveys that can be retrieved using the public module. the management module can be used only by the institutions who are server centres. the management unix search main server (sherlock.crepuq.qc.ca) consult html pages (except for part of extraction and analysis) data extraction transfer of data files retrieval of files resulting from extraction and analysis server center (université laval) (sherlock.bibl.ulaval.ca) server center (mcgill university) server center (université de montréal) server center (université du québec à rimouski) 18 iassist quarterly module has different functions. using html forms, it is possible to work on the surveys, the files (metadata, data sets) and the variables. only the french version of this module is available at this time. at the survey level, the data librarian can add a survey, modify it or delete it. the data librarian also decides the treatment level (e/t), server address where the files will be loaded, which universities will have access to the survey. the data librarian enters the description and the abstract in both languages. once inside a survey at the files level, you enter the files (metadata and data), giving a title to each file. inside a survey at the variables level, you can also add, modify or delete variables. the module includes technical notes which are like an online manual. they are guidelines and procedures to facilitate the entry of metadata information. sherlock also collects statistics on usage (monthly/ annual) by surveys, and by universities. with these statistics we can determine whether the users consult only the description, whether they transfer the complete dataset or whether they perform an extraction or an analysis. promotion now that the development of sherlock is complete, institutions participating in the project are responsible for promoting this collective tool among data users in their respective universities. to facilitate the marketing of sherlock, the crepuq data subgroup organized two sherlock information and familiarisation workshops. the first one took place at mcgill university on october 15, 1998 and the second one, at laval university (québec city) in december 1998. these activities drew more than 50 participants (data librarians and staff serving the public). the introduction of sherlock was supported by a press release and a presentation to the heads of university libraries. in the quebec universities network, library heads voted unanimously to continue the sherlock project. accordingly, phase ii was developed from november 1998 to may 1999. this phase had a two-fold objective: to install sherlock in three server centres (université du québec à rimouski, université de montréal and mcgill university) and to increase the number of surveys in the sherlock collection, because we have gathered around 40 surveys in our collective tool. in addition to maintaining the system, the development team of the laval university library has assisted the institutions with installation procedures. for the year 1 starting next month, a board of management has been created. this group will establish an annual program and will report to library directors. the users will be represented on the group. more recently, the sherlock project won a second prize among fifty projects presented at the caubo (canadian association of university business officers) as an academic initiative and development increasing productivity and effectiveness in higher education. the development team is very pleased with this recognition. conclusion among quebec university libraries’ the collective approach to the management of microdata files is two-fold : first, to “liberate” access to data, and secondly, to liberate their use. sherlock is also an active participant in the data liberation initiative in canada, which concerns the development of a data culture in our universities. in jointly supporting the development of this research infrastructure, quebec university libraries are 1) encouraging the analysis of the statistical information available in the quebec university network, 2) promoting student learning, 3) supporting the work of professors and researchers, and 4) participating in the demystification of data among library staff. i would especially like to thank my three data friends (les trois amis des données). these data friends are not the same as the “los tres data amigos”, well known at the icpsr summer institute. i invite you to meet sherlock in person at the poster session. 1 consisting of richard boily (université du québec à rimouski), jerry bull (université de montréal), gaëtan drolet (université laval) and anastassia khouri (mcgill university). *paper presented at the iassist conference, may 19, 1999, ryerson polytechnic university, toronto, ontario. . gaëtan drolet université laval. vol26no1 4 iassist quarterly spring 2002 iassist quarterly spring 2002 5 editor’s notes this issue vol. 26-1 of the iassist quarterly brings you three papers: vernon leighton from princeton university delivered a paper “developing a new data archive in a time of maturing standards” at the session “a new topical data archive: building a research infrastructure for arts and cultural policy studies in the u.s.” at the iassist 2002 conference in storrs, ct. the paper describes how the cpanda (cultural policy and the arts national data archive) team has worked and chosen standards among the many customs, standards and standard applications. the old joke: “it is so nice with standards, because there are so many to choose from” does actually imply a lot of work. the paper discusses techniques and standards like xml, ddi codebook formats, archival management software, use of controlled vocabularies, full text indexing in database management systems and much more. the successful strategies, followed by cpanda, have been to adopt previous efforts and to minimize duplication. the paper from robert g. cromley and patrick mcglamery on “integrating spatial metadata and data dissemination over the internet” has not been part of a session in an iassist conference. however, the two authors did present some other work at the poster session at the 2002 conference in storrs, ct. the authors cromley and mcglamery are working at the department of geography and the homer babbidge library at the university of connecticut. storrs, ct. the paper describes the components of the metadata records for spatial data and describes the supporting system that will bring this information to the users. patricia cruse, california digital library, university of california; ilona m. einowsky, uc data/src, university of california, berkeley, and juri stratford, government information and maps, shields library, university of california, davis presented at the iassist 2001 conference in amsterdam the paper “applications in the real world: the counting california experience with the ddi”. the counting california initiative is committed to enhancing california citizens ̓access to the growing range of social science and economic data produced by government agencies. this is accomplished by enabling users access to data compiled by federal, state, and local agencies through a single interface. as the title shows the paper is about metadata and the use of ddi (data documentation initiative). but you should also visit the impressive website http://countingcalifornia.cdlib.org and have a look at the data. remember that conference information is available on the web. you can see the powerpoint presentation at the iassist web-site www.iassistdata.org where many “multimedia presentations” are available. if your presentation is not included you should contact the collector lisa neidert lisan@umich.edu. papers from the conference are scheduled to appear in this and coming issues of the iassist quarterly. please contact the editor kbr@sam.sdu.dk. plan ahead and make a note about the iassist conference in ottawa in 2003 (may 27-30). karsten boye rasmussen, oct. 2002 http://countingcalifornia.cdlib.org http://www.iassistdata.org/ mailto:lisan@umich.edu mailto:kbr@sam.sdu.dk sist newsletter vol.1, no. 2 action group reports data archive registry canadalisa lasko, canadian consortium for social research, institute for behavioral research, york university, 4700 keele street, downsview, ontario m3j 1p3 europejoseph bonmariage, belgian archives for the social sciences, university of louvain, sh-2, 1348 louvain-la-neuve, belgium united statesdavid nasatir, behavioral sciences graduate program, california state college, dominguez hills, california 907m-7 unesco directory of data services the unesco social sciences sector has passed a contract with the international committee for social sciences information and documentation (icssd, 27 rue saint-guillaume, 75007 paris) to establish a first directory of data services . the icssd has sent out a complex questionnaire which some of our readers may already have seen. philippe laurent and stein rokkan discussed the project with jean meyriat of the icssd in november and tried to work out an arrangement under which the lassist action group for data archive registry could be associated with this work, but this proved difficult under the contract established with unesco. it was agreed that the current project should take its course and be looked upon as a pilot phase. the further work of systematization, computerization and updating would be undertaken by bass in co-operation with the action group for data archive registry. a detailed plan for this follow-up project will be prepared during 1977 and submitted to unesco and other funding agencies. in north america, efforts continue to survey existing directories of data libraries, archives, and data information services in order to determine the type of directory which will best meet both data information and general reference community needs and to determine how effective are existing directories, who uses them and with what success. data acquisition canadapierre lacasse, centre de recherches en ame'nagement regional, universite' de sherbrooke, sherbrooke, quebec europemarcia taylor, social science research council survey archive, university of essex, wivenhoe park, p.o. box 23, colchester, essex, england c04 3s0 united statesisist newsletter vol.1, no. 2 as the amount of data available in machine-readable form increases, it becomes even more important for data services to address on local, national, and international levels questions of what should be preserved, by whom, and in what form. in north america, this group will study the needs for "data services' collection statements similar to those produced by paper archives and traditional archives, appraisal of machine-readable data files, and de-acnuisition. data documentation canadadave l. salley, management and central services group, standards division, statistics canada, tunney's pasture, ottawa, ontario kia 0t6 europecees middendorp, steinmetzarchief , kleine-gartmanplantsoen 10, amsterdam-c, netherlands united statesjohn grasso, office of research and development, center for appalachian studies and development, west virginia university, morgantown, west virginia 26506 in north america, this action group will begin to address minimum standards for the documentation of social science mrdf. lassist has received requests from both funding agencies and data information services and individuals to provide some formal means by which documentation may be evaluated. some work has been done in this area and it should be possible at least to produce a check list in the near future. the special documentation requirements for process-produced data will be a prime concern of that ag. classification canadamohan sharma, humanities & social science library, university of alberta, rutherford north, edmonton, alberta europeekkehard mochmann, zentralarchiv fiir empirische sozialforschung, bachemer strasse 40, 5 koln 41, federal republic of germany united statessue dodd, data library, insitute for research in social sciences, manning hall, university of north carolina, chapel hill , north carolina 27514 n^fesist newsletter vol.1, no. 2 report of the joint united states-canadian action groups on classification (c ag) submitted by sue dodd, university of north carolina mohan sharma, university of alberta members present united states sue dodd, ch, university of north carolina nancy carmichael, center for social indicators, social science research center, washington, d.c. debra powell, dualabs, washington, d.c. canada mohan sharma, ch, university of alberta elliot m. avedon, university of waterloo sharon chappie henry, data clearing house for the social sciences agenda topics for the c ag included (1) discussion of useful areas of concern and future coordination of tasks between us and canadian ag members; (2) review of cataloguing efforts to date and discussion of any or all problems concerned with cataloguing tasks, including a review and discussion of the working rianual for cataloguing machine-readable data files [mrdf] compiled by sue dodd; (3) discussion of the ramifications of mrdf catalogue records, such as the national union list of social science data; shared cataloguinn; cataloouina-in-production; the marc ii record as a standard format for storing automated bibliographic records of data files and its flexibility for an expanded record of information which would resemble a data abstract or study description; network interactive systems among data centers and libraries; and, an on-line search and retrieval system for the expanded marc il-type record; (4) practical exercise in applying subject headings and descriptors for several large and uniguely held data sets, with a view towards compilina the beginning of the thesaurus or authority list of social science terms for data files; (5) discussion of existina printed and other available thesauri or authority lists in the social sciences, and a review of them in terms of future applications to data files; (6) discussion of the lack of adequate subject headings and sub-headings currently provided by the library of congress for social science data files, with a view toward providing constructive recommendations. given these broad topics of discussion, the classification ag decided on the following projects and sets of recommendations: 1. there are mutual areas of concern and projects on which the us and canadian representatives could work together. there would continue to be an active exchange of information, with the next meetinqs scheduled for may 1977 in toronto. 2. the ag's cataloguing project can be viewed as a tremendous success and as an example of a cooperative effort to evaluate both the ala recommended rules and the working manual for cataloguing mrdf. (a total of 40 examples were submitted by 8 individuals from control data corporation, rutgers university, the national archives, dualabs, yale university, university of pittsburgh, national opinion research center, and drexel university.) sist newsletter vol.1, no. 2 each cataloguing effort will be individually identified and circulated among the classification group. representative samples will be selected for inclusion in the manual (with permission of the cataloguer). the c ag will outline procedures for shared cataloguing of mrdf to avoid duplication of work and to maintain a high quality of effort. the catalonuing project will be extended to those now cataloguing data files or those wishing to be included in the project. the final report will follow the completion of the initial cataloguing effort and the final revision of the manual . 3. the working manual for mrdf will be revised based on numerous suggestions by members of the c ag and by participants of the working conference. these recommendations along with the evaluation forms, will be circulated among the c ag members. the draft of the revised manual should be ready for review by the c ag meetings in toronto. [a further measure of success of this project is the possibility that the manual may be accepted for publication by one of the 1 i brary-aff i 1 iated associations . ] 4. the c ag recommended that procedures for describing a data file, such as study description, cataloguing, and descriptors, (i.e., "cataloguing in production") be built into the daily routine of classification of any data file, and indeed, that these procedures be implemented at the early stage of data file creation. in addition, the ag recommended procedures that could be directed toward the major funding organizations (e.g., nsf) which would be built into the contractual provisions: (a) data be documented according to some type of recommended standard; (b) information be provided for cataloguing; (c) data descriptors be applied; and, (d) data be designated to the appropriate archive within a reasonable amount of time. 5. the c ag will prepare two position papers reflecting the pros and cons of establishing a national union list of mrdf. 6. a bibliography of existing thesauri or authority lists for social science terms will be collected and reviewed in view of their adaptability to social science data files. in addition, representatives of the c ag will be in close contact with existing work in this area, such as the committee on conceptual and terminological analysis (cocta). any professional groups currently involved in compiling thesauri or conceptual lists of social science terms will be contacted to coordinate and share information. as a voluntary and non-funded group without expertise in this area, the c ag is not equipped to assume the task of creating a thesaurus of social science data files. however, it will attempt, in so far as possible, to provide guidance in this area and to share experiences gained from c ag associated projects to be carried out by members within their respective institutions. two such projects are the compiling of an authority list for census and census-related data files [dualabs] and an authority list of terms representing question/item level descriptors for harris public opinion polls [north carolina social science data library]. 1 0. ^^sist newsletter vol.1, no. 2 7. a standards and quality-control review board (soqcrb) will be established within the north american joint classification ag to evaluate mutual projects, recommend procedures, establish standards, and work with other established organizations committed to or involved with the bibliographic control of mrdf within the social sciences. one of the first tasks of this board will be to provide guidance and examples of bibliographic references for social science data files. at the present time, there are no rules or a standard format for citing data in the published literature. the quality or amount of information varies among authors, editors, and students; often, the information is insufficient to enable a researcher to replicate the data file for the type of secondary analysis so important to the discipline. (for example, it is often very difficult to identify a data file, or its source, or data elements on which the published analysis has relied.) the standardization of bibliographic references for mrdf would pave the way for their inclusion into such reference work as the social science citation index . upon completion of this task, it was decided that a letter would be sent to the standards committee z39 of the american national standards institute (ansi) recommending an amendment to their forthcoming publication. the american national standards for bibliographic references . the amendment would be applied to this committee's method of citing "data files," which now provides examples for only one type of data (bibliographic data files) and which was created without any coordination with representatives of the library community who are establishing standards for cataloguing mrdf. the c ag feels very strongly that any standard citing data files should be compatible with the rules set forth in the angloamerican cataloguing rules ii [forthcoming]. the second task of the review board will be to compile a list of recommended subject headings and sub-division headings for social science data files. these recommendations would be forwarded to the head of the subject cataloguing division of the library of congress. the work of the c ag and any resulting standards or reports should be given the widest possible coverage among the existing scholarly and professional journals. someone will be appointed to write brief abstracts and notices of the work and will be responsible for circulating it in the available press. data archive development canadalaine ruus, data library, computing centre, university of british columbia, 2075 wesbrook place, vancouver, british columbia vst 1w5 europenot activated united statesalice robbin, data and program library service, 4452 social science building, university of wisconsin-madison wisconsin 53706 11 sist newsletter vol.1, no. 2 report of the joint united states-canadian action groups on data archive development (dad "ag] submitted by alice robbin, university of wisconsin-madison laine ruus, university of british columbia members present united states alice robbin, ch, university of wisconsin-madison laverne d. knezek, texas christian university [also dom ag] judith rowe, princeton university peter tolousis, temple university canada laine ruus, university of british columbia edward hanis, university of western ontario objectives of the meeting were to (1) develop a conceptual framework for a publication, the working title of which is, a guide to providing social science data services , (2) revise, modify and elaborate a preliminary outline for this publication, (3) identify areas of responsibilities to be assumed by other action groups, (4) establish deadlines of completion for various sections of the guide by contributing ags. the dad ag defined the target audience for this publication as those individuals already providing or intending to provide data services for research, policy, and planning purposes. these individuals are typically in academic, research, and governmental organizations. the nature of these data is non-bibliographic, primarily quantitative, including microand macro-level of aggregation, but may also be textual. the following chapters of the guide were agreed upon: (1) preface describing the target audience, purpose for the book, etc., (2) introduction providing the historical context for the development of social science data services, (3) functional overview defining a social science data services system and its components, (4) detailed functional description of the system which includes management or technical processing and user services aspects, (5) organizational models and services provided, (6) organizational development including resource requirements and scenarios for the development of services, (7) glossary of terms, (8) extended bibliographic references tied to each of the sections, (9) detailed index, (10) appendix of procedural, management, and administrative forms, with introductory remarks on the limits and usefulness of recordkeeping. the preliminary outline prepared for the meeting was revised and will be circulated to all ag members attending the joint conference. comments will be forwarded to ag coordinators, in time for the toronto meetings. further modifications based on recommendations made during the toronto meetings will result in a detailed outline whose projected completion date is june 30, 1977. at that time, individuals in the various ags will assume responsibility for the substantive topics/sections within each chapter. alice robbin will act as general editor for chapters (1) through (6); laine ruus, for chapters (7) 1 2. i^ssist newsletter vol.1, no. 2 through (10). the dad ag plans to have a first draft completed by august, 1978, and a second draft by december, 1978. discussions are now underway with a potential publishing house who has indicated an enthusiastic interest in the project and has already offered useful advice. a second project [discussed in greater detail by k. heim in the "book notices" section of this f^ewslett^r "! is now fully underway. this is the production of an annotated bibliography, whose provisional title is perspectives in social science data services: an annotated bibliography . thus far the dad ag has gathered almost 200 articles relating to this topic. a letter was sent to all data archives last fall asking that they participate in this report. the dad ag noted the importance of a collaborative effort, and therefore requests assistance: we would like every lassist newsletter reader to search his/her files for all pieces of literature (articles, unpublished papers, conference addresses, parts of books) concerning data archiving (retrospective and current) and to send citations to the lassist newsletter editor. all respondents will be acknowledged in the bibliography. process-produced data canadajohn devries, social science data archives, department of sociology, carleton university, ottawa, ontario kis 5b6 europepaul muller, institute for applied social research, university of cologne, greinstrasse 2, 5000-koln 41, federal republic of germany united statesdonald harrison, national archives (nnr), washington, d.c. 20408 report of the joint united states-canadian action groups on process-produced data (ppd ag) submitted by john devries, carleton university donald f. harrison, national archives, washington, d.c. members present united states donald f. harrison, ch, national archives, washington, d.c. charlotte boschan, national bureau of economic research, new york harriet dhanak, michigan state university shirley gilbert, princeton university [also dom ag] elizabeth powell, leaa/ncjiss, department of justice, washington, d.c. canada john devries, ch, carleton university tony falsetto, public archives of canada, ottawa 1 3. ^sist newsletter vol.1, no. 2 ex officio carolyn geda, icpsr, university of michigan definitional problems were first addressed by the members. the following definition of process-produced data v/as agreed upon. processed-produced data are data which are products of routine, administrative activities of private or public institutions or persons. they can be explained with respect to the functions of their origin: either for administrative or policy uses; can be merged, linked, abstracted, or summarized; and, may be used for research and/or statistical analysis. the paul muller paper [see edited version in the newsletter issue] was reviewed in order to arrive at the mandate of the action group . generally, we have asked ourselves, "how can we facilitate the movement of such data from the originating institution or person to the secondary user?" to do this, the ag explored areas of concern, some of which were rejected, and some where it was felt that the ag might logically make a contribution. we rejected the following: 1. laws regarding confidentiality: this subject logically should be a matter of concern for this ag. but, because there are so many other institutions/professional groups working on the same subject, the ag prefers to await their reports before proceeding in order to avoid duplication of work. the ag members are particularly anxious to await the final report of the project of ed hanis and dave flaherty at the university of western ontario, which deals with this issue. [hanis and flaherty are examining privacy/confidentiality problems as they are dealt with by central statistics bureaus in west germany, sweden, great britain, canada, and the united states. a final report is due at the end of this year.] 2. linkage techniques: a subject as complex as this requires so much technical and concentrated study, that the ag considers this an area of concern for another or a new ag. 3. uses made of machine-readable files: the ag rejected this subject as having no useful purpose. the following projects were accepted as within the mandate of the ag: 1a directory of catalogues which list data bases (inventory of inventories) first priority. the ag anticipates presenting lassist with a first draft in toronto in may, 1977. by february, 1978, it is expected that the directory will be presentable for publication in the lassist newsletter or in some other suitable journal, after proper lassist backing. 2an inventory of existing guidelines in use by the originating institution or archival institution for: (a) documentation and fb) preservatinn (not tn include acquisition activities). second priority. the ag will begin to collect and disseminate the information within the jag immediately, and plans to present an outline of the method for compilation by may 1977 in toronto. a first draft should be available for presentation at uppsala in august, 1978. 1 4. fiassist newsletter vol.1, no. 2 3. an inventory of procedures in use by originating and archival institutions for (a) documentation on processes; and (b) management decisions. second priority. like project #2, the ag will begin to collect and disseminate almost immediately those procedures. plans are to present an outline by may, 1977 in toronto. if progress is made, a first draft should be completed in time for the uppsala meetings. 4. an lassist publication of desirable documentation components necessary to service, retrieve and otherwise manipulate a process-produced data base . third priority. although extremely important, it logically follows that such a publication cannot be produced before first surveying the field and consulting practitioners. plans are to present an outline for consideration at the uppsala meetings. these projects have been arranged in priority, not by their relative importance, but in consideration of their reasonable and logical sequential order for work. all projects except the last one will be prepared separately by national action groups, each consulting with institutions in their parent country (i-e., the usag will prepare lists of inventories published or originated in the united states, and so forth). project #/ priorities may 1977 february 1978 august 1978 1/1 first draft finished product 2/2 outline outline first draft 3/2 outline outline first draft 4/3 outline data organization and management canadagreg morrison, social science data archive, department of sociology, carleton university, ottawa, ontario kis 5e6 europeeric tannenbaum, social science research council survey archive, university of essex, wivenhoe park, p.o. box 23, colchester, essex, england c04 3s0 united stateswilliam gammell, social science data center, university of connecticut, stnrrs, connecticut 06268 15. ^sist newsletter vol.1, no. 2 report of the joint united states-canadian action groups on data organization and management (dom ag) submitted by bill gammell, university of connecticut greg morrison, carleton university members present united states william gammell, ch, university of connecticut peggy cahn, federal reserve bank of new york shirley gilbert, princeton university pnina grinberg, columbia university gary klass, state university of new york at binghamton laverne d. knezek, texas christian university [also dad ag] sheldon laube, cm. leinwond associates , newton, massachusetts barbara noble, center for advanced computation, university of illinois pat peterson, agency for international development (aid) elizabeth powell, leaa/ncjiss, department of justice, washington, d.c. [also pp ag] richard c. roistacher, center for advanced computation, university of illinois canada greg morrison, ch, carleton university, ottawa, ontario rachel des rosiers, canadian data clearing house for the social sciences ex officio carolyn geda, icpsr, university of michigan during the conference, the following activities were carried out: 1. began compiling an inventory of software relating to data organization and management . this effort is viewed as an on-going project. the coordinators will complete a first draft to be distributed and approved during the lassist may meetings in toronto. the inventory wil 1 be published in the newsletter . contributions to this list are encouraged. 2. discussed advanced developments in data management. richard roistacher reported on the development of the data interchange file concept and on the use of computerized document processors for the creation and management of machinereadable codebooks. sheldon laube reviewed the capabilities of existing data cleaning software and described the design and development work now being done on a generalized data cleaning system. the ag will be active in providing input to these projects. 3. reviewed with representatives of the data archive development ag the outline for the guide [see description under dad ag report] and discussed possible contributions to this project. copies of the revised dad ag outline will be circulated to group members in time for extensive consideration at the toronto meetings. 1 6. ^sist newsletter vol.1, no. 2 planned several projects: (a) development of a list of recommendations addressed to potential researchers on study desiqn as it relates to data management (the "do's and don'ts" list). group members are to send in their suggestions to the dom ag coordinators. a consolidated list will be produced in toronto for publication in the newsletter and elsewhere, (b) collation of information on relevant monographs, technical reports, program writeups, etc., which the ag members would like to share. the intent of the collection is to publish this information in a "what's new in data flanagement" section of the newsletter . 5. prepared an agenda for the toronto meetings; items for that agenda have been referred to above. the dom ag coordinators would like to suggest the following revised mandate for future discussion at international lassist meetings: "this group addresses the problems of data organization and management confronting those archiving or usinq social science data. the ag will investigate and evaluate software and procedures for data and documentation preparation and management; recommend nuidelines for preparation procedures and software development; and, sponsor workshops and seminars for professional training and the exchange of information in these areas." [the original mandate includes "hardware" as an area to be addressed by this ag; see newsletter volume 1, number 1 for the full text of the mandate.] an overview of problems associated with proc ess prod uc ed data / paul miitler for the august 1976 lassist meetings, paul muller preoared a report which provided a broad overview of problems associated with process-produced data. what follows is an edited version of this report. [future issues of the news-letter will contain additional action group reports. the membership is encouraged to begin a dialogue on this and other issues of concern to the data archive community.] "administrative bookkeeping as a social science data base" by,, paul muller institute for applied social research university of cologne 1.0 [...] i will try to give a rather broad overview of problems associated with the "production, acguisition, preservation, processing, distribution, and utilization of machine-readable" process-produced data. this report is necessarily biased by my own viewpoints and experiences; other experiences may well be fundamentally different. but,' the function of this paper is to initiate discussion and later, joint actions. iassist quarterly 3 problems in the use of business data on tape by e kay worrell, ehrector, survey research center the conference board, new york introduction this will be a description of some types of business data available on magnetic tape, the uses made of these data by companies and business researchers, and some problems that may be encoimtered. examples will be taken from experience with conference board applications and from discussions with business users of such data among conference board associate companies. the conference board publishes research reports, economic forecasts and newsletters, and organizes conferences for the business community. it is supported by subscription income from associate companies, conference fees, and publication sales. the survey research center provides assistance to research staff in the collection and processing of survey data, including maintenance of research-related mailing lists. we also provide statistical support, training in microcomputer software, and assistance in database design and interface between the mainframe and mictocomputers. types of business data there are three broad types of business data available on magnetic tape. 1 directory information, including names, titles and addresses of executives 2 financial and other descriptive information about individual companies 3 aggregate statistical information on the economy, the workforce and industrial production and services. i will focus on the first two types: directoryinformation on executives, and descriptive information about individual companies. two of the most commonly used sources of data on tape are dim & bradstreet corporation and standard & poor's corporation. both companies were formerly best known for their printed directories, and tape products were derived from computerization of these data. dun & bradstreet also offers extensive business data processing and mailing services. both companies are producers of proprietary on-line databases. standard & poor's and dun & bradstreet rely to some extent on self-reporting by companies, in addition to publicly available financial reports and other published sources. the titles of individual executives are coded in some detail. standard & poor's codes up to four separate titles for each individual, using a two-part code consisting of 35 discrete designations of rank (e.g. assl v.p.) and 25 functions (e.g. finance) a maximum of 20 unique standard industrial classification (sic) codes describing products or services produced can be provided, though some tape products may not contain all. fall 1986 iassist quarterly an administrative database containing information on all publicly-held companies is now available in machine-readable form from disclosure inc. all companies whose stock is publicly traded in the united states must file annual and quarterly reports with the securities and exchange commission (sec). the annual forms — the 10-k for u.s. based companies, 2(>-k for foreign companies — and the 8-k quarterly forms have been available on mictofilm from disclosure inc. for more than ten years. data are extracted from these, as well as from annual reports and proceedings of annual meetings, and are now available on-line from disclosure, as well as on magnetic tape and microdiskene. information on the chief executive officer and directors of a corporation is also included. the cash compensation awarded to the lop five officers is provided, along with their names and titles. little other information on individuals is available, however. a recent entry into the tape product market is the national register publishing company (nrpc), publisher of the directory of corporate affiliations, the directory of advertisers, and the corporate bluebook of financial executives. nrpc is an especially good source of detailed information on subsidiary units of corporations. tape files from nrpc are now available as parallel products to each of the printed directories currently produced. each directory product is tailored to one specific business market like standard & poor's and dun & bradstreei, public data are supplemented by data solicited directly from companies, and all data are submitted annually to these companies for verification. uses of business data on tape three types of information may be of interest to the purchaser of business databases on tape, and the type that is of primary importance may determine the choice of vendor. the types of information are: 1. classificatory information, such as sic code, or even sales or assets of corporations, 2. names, titles and addresses of individuals, most often required coded by title — both level and function, 3. detailed financial information on corporations. companies purchase data on tape for a variety of reasons, but primarily for purposes related to marketing. in addition, they may use these data for analysis of financial and other information for plannmg and comparisons relating to investment, mergers, acquisitions and divestitures. marketing activities involving use of such data fall into three categories:. 1 production of labels for direct mail, 2 update and improvement of in-house lists, 3 analysis of information for marketing comparisons. there is some overiap among these activities. information may be used for mailing purposes directly from the purchased tape, or may be run against in-house computer files — a "merge/purge" run — lo produce a non-redundant mailing. this merge/purge run ma> be done by the company's in-house edp department or may be subcontracted to a list broker or service bureau. dun & bradstreet fall 1986 iasast quarterly 5 offers this service to purchasers. data from purchased tape may be merged with other machinereadable data, such as the corporation's own list of clients, for update purposes or for the addition of analytical variables which the user company may not collect and maintain, sub as sales, assets, or sic classification. an example of the latter was a recent request from one of our associate companies seeking advice on matching sic codes from a purchased tape file with their in-house client hsl they wanted to be able to analyze their client list stratified by industry groups. two problems were encountered. an alphabetic match on company name resulted in only a 60% match with their own list. use of a unique identification number might have facilitated a more complete match. also, the company was specifically interested in sic codes in the retail trade series. the dataset they had purchased included 6 sic codes per corporation. only the 6 sic codes thai described their primaryactivities, reflected in the proportion of the annual revenue generated by those activities, were listed. there was a distinct possibility that the codes in which they were interested were not always included for companies involved in retail trade activities. our own most recent use of a purchased tape was in conjunction with a survey on corporate benefits. the tape was acquired from standard & poor's, and the sas system was used to produce mail labels (in a time-sharing system). selecting subjects by title within certain industry group and size parameters. we wanted to mail a questionnaire to senior human resource officers. although a great variety of funcuons were coded, that particular release of tape product had a disappointingly low number of high-ranking officers coded as this funcuon. only about 65% of the largest 1000 corporations had an identifiable senior human resources officer coded. a final example is an investigation we concluded recently on the feasibility of using another external tape product the conference board has recently begun to identify a second level of major corporate leadership, the chief executive officers of large subsidiary companies. in our exploratory work we used the printed version of the directory of corporate affiliations. a complimentary product is now available on tape. this is the most comprehensive source we have found of information on levels of corporate ownership. the vendor relies heavily on descriptions of corporate levels provided by individtial companies, which are polled annually. because of the variety of reporting procediu'es used by the companies, it is difficult to extract reliable lists of second-level executives which are consistent from company to company. we have resoned to checking the armual reports and organizational charts that the conference board collects annually from cooperating companies. many of the companies listed in the national register directory appear to be "paper" entities rather than corporate profit centers. several may be headed by the same individual. preliminary work suggests that these tapes will not be useful for our purposes. summar) of problems and solutions tv-pes of problems one encounters in using commercial!) available business databases on tape are: a. the quality of the data b. the appropriateness of the content of the fall 1986 6 iassist quarterly data for the intended purposes c. logistic problems related to implementation of the intended use, e.g. matching external data with internally produced data d. format and documentation of the tape. the industrial classification codes assigned vary from tape source to tape source, as do printed sources. often, these codes are assigned by clerical staff on the basis of reported descriptions of business activities provided by the responding company; these may be incomplete or enoneous. they may also be inconsistently or incorrectly interpreted by the coders. the composition of the corporation may change through merger, acquisition, divestiture, or change in the revenues generated in specific sub-units, so that the original "primary sic-code" no longer applies. to facilitate merging information from an externally purchased list with one's own in-house company list, it may be advisable to add a unique identification number to the companies in one's own lisl the cusip number, used by the sec, the d-u-n-s number assigned by dun & bradstreet, and the fortune number used in the fortune data bank, have all been incorporated into at least one other database besides that of the generator of the number. one or more ticker symbols (stock market codes) may also be included, but these are not consistent from exchange to exchange. sundard & poor's includes the cusip number; disclosure includes both the fortune and d-u-n-s numbers, as well as the cusip. quality-checking. computer-produced, customized print lists and labels are also available from these same suppliers, and are frequently the most cost-efficient way to purchase the information. the list brokers use a variety of trade journal subscription surveys and other resoiu'ces. in order to update certain specialized lists, we have purchased sets of labels or print-outs from specialized vendors. periodically, in order to update our research mailing list of banks, we have purchased a printed listing from r. l. polk & co., whose primary research focus is on banks, and compared it to our own in-house list we find the primed copy easier to work with, as we have limited database management capability in-house. our own database, maintained on a burroughs system, can not readily accommodate external files. we request from polk's print-out of the 3,000 largest banks, in alphabetic order by state and city, showing the total deposit income of each. we then select for our list those banks with total deposit income of e.g. over $500 million (or $100 million or $1 billion). these we check against our own in-house list, especially noting possible bank mergers or name-changes. these are easier to identify in listings ordered geographically. the most difticult problems facing those who would merge external information with internal information are the problems of subsidiary units, and, more recently, the even more complicated problem of joint ventures.n there are also list service houses which will merge the data and do some specified checking, based on text-matching of compan\ names. this is far less successful than matching on a unique identification number, which allows the end user substantial control over fall 1986 i rights of researchers and governments to national records who owns contract and grant data and who can use it? by ekkehard mochmann zent ral archi v fur empirische sozi al forschung university of cologne this article is reprinted with permission from the editors of european political :data newsletter. i. freedom of research, but no access to data? the legal context for social science data access in germany is significantly different from the american situation. while the federal and state data protection laws have been enacted in the late seventies (like in most european countries), a general data access regulation as the equally important component of information legislation is still lacking. in this situation article 5 of the german federal republic's constitution is referenced as the most authoritative written norm. it guarantees freedom of arts and science, research and teaching in general, but does not say anything specific about data access. the interpretation of this article by courts and experts, however, acknowledges in principle the right to information access. this position is contrasted by the actual behaviour of the german administration. an orientation to keep information under its control is prevailing (1). in practice it is the researcher who has to justify the information request and has to prove that he cannot achieve his results by other means. there is no regulation demanding from the administration to justify its refusal. on the contrary the researcher has to convince the data holding administration of the impor:tance and legitimacy of his research intention. in short: there is no equivalent to the freedom of information act. before privacy legislation was enacted, the legal situation in western germany vas reasonably well characterized by the statement "that the owner of data-including personal data--was more or less regarded as proprietor who could dispose df them as long as he did not violate the rights of other persons" (2). in terms jf data accessibility (not protection) privacy legislation changed the situation :o the worse: it was frequently misused as an argument to prevent access--even in :ases where privacy legislation did not apply. ! access to data from the federal statistical agency the mandate of the federal statistical agency is (among others) to provide juantitative information about social and economic development of the society, 'he same can be said for the state and commune agencies. more than sixty legal norms regulate the procedure for more than two-hundred separate statistical counts, which have to be conducted. whenever a special legal norm is enacted, it specifies all details for the survey (sample, variables) including access and disposition rights. these normally rest with the office, which initiated the data collection. there are cases, in v\;hich statistical offices of the communes collected the data for the state office, but were not entitled to analyze the data themselves, either before or after transfer to the state office. in other cases public access was explicitely guaranteed. whether access actually can be achieved is frequently a question of the fees which have to be paid for usage (e.g., up to dm 6000 for the copy of a tape from a \% sample of the micro census). apart from these specific regulations the general procedure for the federal agency was defined in the federal statistics act in 1953 which was revised in i98o. this new version responded to problems arising from recent privacy legislation. according to the new regulation the statistics can be used by scientific institutions and other interested bodies, only if the data is anonymized. the data flow is seriously hampered by the fact that appropriate anonymi zat ion procedures have to be developed and implemented (3). in fact, given the problems of anonymi zat i on , even public use files are not available. on the other hand the federal statistic agency offers the services of statis-bund. this is a network-oriented service for access to data and appropriate statistical procedures. you can analyze the data in the bank, you can bring in additional data which you had collected yourself, but you cannot transfer the data to your own installation (k) . 3. access to administrative (process-produced) data the transfer of anonymous data is permissible, transfer of identifiable is restricted. again data can be transferred only if the researcher can convince the agency and data protection commissioner that the research interests are considerably higher than contradicting interests and that the research goal cannot be achieved by other means. basically the data access is under complete control of the admi n i st rat i on . h. access to contract data government agencies have a high demand of data, but hardly any personnel resources for data collection or data analysis. these activities are usually contracted out. given this situation they are not interested to lay their hands on the data itself. their needs are satisfied when they receive the research report and tables. there are hardly any resources for data analysis in the offices of the aermi ni st rat i on . as a consequence no attention is being paid to clarify the rights regarding the data in the contract. only recently attempts to alert the responsible administrators to the fact that publicly financed data collections are an important resource for secondary analysis are gaining increasing attention. let me characterize the situation regarding access to cross-section surveys by two contrasting experiences. since the first story given an example of very questionable performance in a critical political situation, it may be particularly important for discussion. nevertheless i will protect the identity of this office, since i cannot give a fair account of the detailed arguments here. in 1979 one of the more prominent offices of the federal administration asked a commercial research institute to contract a study on the right-wing radical potential. this research team had contracted the field work for the cross-national survey to one of the leading opinion research institutes. the results which were finally reported in the media were substantially contradicting the findings of a prominent german sociologist, who had contributed to most important research findings in this field. of course he wanted to reanalyze this new data set. we asked the financee for access to the data. the response was positive, but conditional on the agreement of the contract institute. this agency was positive too, but conditional on the o.k. of the sub-contractor. this was very positive, but we did not receive the data. after several i terat ions--and even political i ntervent ions--we were informed by the government office that data transfer seemed not to be advisable in this given situation. almost parallel to that we were informed by the contractor that the interested researchers certainly could inspect the data in the contractor's office; apart from that the results would be published on the book market in short. the book is available, we are still waiting for the data. one positive experience stems from negotiations with the state of horthrhinewestfalia and the federal post minister's office. they currently are conducting implementation and evaluation studies for the two-way communication system, which is called b i 1 dsch i rmtext in germany (prestel in england, telidon in canada and antiope in france, just to name a few). the contracts with the research institutes clearly define that all rights regarding the data rest with the financees, and they are interested in an intensive usage of these data sets. to prove this: we have produced the codebook and a character data set for two of the major studies al ready (5) • likewise we received and distributed three big data sets from the federal labour minister's office with results from recent studies of unemployment to quote just another positive experience. i could continue with an amazing example from a postord i nated agency (nachgeordnete bundesbehbrde) , which holds numerous labour market data sets. the most positive declaration to start data transfer to the social science community via the zent ra larchi v was unfortunately restricted by including a privacy clause, which explicitly stated that data can only be used for the purposes of this office. it might have been phrased slightly different to open access to the social science community in general. but i prefer to follow the positive terms described before. 5. federal archives act a law for the federal archives (bundesarchi v) has been drafted. there was none in existence before. reactions to privacy legislations promoted the idea of a federal archives act. this draft explicitely offers access to research--that , with some exceptions, can wait 30 to 120 or even i50 years. i understand that these regulations are similar to those in many other countries. they can be interesting for historical research, they hardly will be useful for empirical social research. all materials produced or received by federal agencies, will be subject to this law unless considered not worth archiving. nothing is said about contract data. all materials have to be offered to the federal archive as soon as they are no longer needed in the public administration. a decentralized principle is followed in so far as materials from local or state agencies can be sorted in their respective archives. nothing is said about contract data. this point, however, is tapped in the swiss neighbour's "guidelines for processing personal data in the federal admin i strata! on". in case a state (kanton) or commune, a private person or organization is given a contract, data protection rules have to be specified by contract or order and have to be supervi sed-i f possible. nothing is said about access regulations, howeve r (6) . 6. access to grant data this situation can be characterized very quickly. besides research foundations. government agencies give a significant amount of funds in form of research grants. as a rule the data collected under these grants is under complete control of the researcher who collected it (of course subject to current privacy and other legislation). unfortunately the practice of data sharing in the research community is still less popular than arguing against data protection and restricted data access regul at ions . 7. result i there is no clear answer to the question of ownership and access to contract data. discussion of negative effects of privacy legislation and positive experiences with special access regulations as well as references to the international development gradually develop the feeling for the importance of this subject in our research community. the experts' discussion about the topics is reported in specialized journals and monographs. wh i le the state of hessen was first to pass a privacy protection law we have to catch up with respect to data access regulations (7). part of the game is to balance political, commercial and research interests. references j. scherer "datenzuqang des forschers zwischen i nformat ionsanspruch and gehe imha 1 tungsgrundsatz", in: m. kaase ; h.j. krupp; m. pflanz; e.k. scheuch; s. simitis (eds.): datenzugang und datenschutz, konigstein 1980, p. 96. u. dammann und r. brennecke: "country report federal republic of germany", in: e. mochmann, p.j. muller (eds.): data protection and social science research, frankfurt 1979, p. 129b. pohl: "datenschutz in de r amt 1 i chen statistik", in: datenschutz und datens i cherung , volume 2, 1981, p. 69(k) before access is permitted you have to sign a contract specifying access conditions and fees. descriptive information about statisbund, its holdings and program libraries are available from "stat i st i ches bundesamt, wiesbaden" (5) i nf ratest-nul lerhebung im rahmen der wi ssenschaft 1 i chen beglei t forschung zu bi 1 dschi rmtext and bi 1 dschi rmtext-voruntersuchung (langenbucher , scheuch, treinen, date collection by infratest munchen) (6) r.j. schweizer "die richtlinien des schwe i zeri schen buncesrates uber den datenschutz in der bundesverwal tung", in: datenschutz und datens i cherung. vol. ^, 1981, pp. 239-2'i2, and "richtlinien fur die bearbeitung von pe rsonal daten in der bundesverwa 1 tung" , paragraph 3^*. in: datenschutz und datensi cherung. vol. ^4, i98i, p. 2't3the german science foundation has implemented a committee on privacy protection and data access problems. see also. m. kasse; h.j. krupp; m. pflanz, e. k. scheuch; s. simitis (eds.): "datenzugang, datenschutz", konigstein i98o. iassvol201 21fall 1996 establishing data and documentation standards for investigators who are required to archive research data by patrick t. collins1 , project director, national data archive on child abuse and neglect introduction this paper is about the approach that the national data archive on child abuse and neglect has taken to improving the quality and consistency of our documentation. of the many problems we have encounytered, including uncooperative investigators, dirty data files, and unusual file formats, poorly prepared or non-existent documentation has been the most difficult to handle. since investigators were not required to archive their data with our archive, we had to actively solicit contributors. most researchers were unwilling to contribute their data and the ones who were willing had little or no resources to dedicate to the task. in short, we were in the position of having to accept whatever investigators were willing to provide. in many cases we received nothing more than an spss or raw data file and a copy of the instrument, leaving us with the daunting task of creating a user’s guide from scratch. the task of preparing comprehensive documentation for these studies was so time consuming that were only able to process 2-3 datasets per year. while we appreciated the efforts of the investigators who chose to contribute data to the archjive, it became clear that the only way the archive could expand its holdings with any speed would be to improve the nature of the materials contributed by investigators. since our archive was funded by a federal agency with an active research program, we chose to work through that agency in order to establish data documentation standards for their research grantees. but before i describe this process in detail, let me tell you a little more about the archive. the national data archive on child abuse and neglect the national data archive on child abuse and neglect has been in operation for approximately six years. during this time ndacan has received all of its funding from the national center on child abuse and neglect (nccan) which is a division of the administration for children youth and families which in turn is a unit of the us department of health and human sercies. nccan is the federal agency witht the primary mission of responding to child abuse and neglect in the usa. one of nccan’s many responsibilities is a field initiated research program which is funded at approximately $1.5 million per year. the archive, which is funded through this program, works primarily with nccan’s research grantees. we are in the final year of our second three-year award from nccan and will apply this spring for continued funding. over its six years of operation, the archive has been flat-funded at $150,000 per year, leaving us with approximately $100,000/year in direct funds. most of these funds are used to support our 2.2 ftes. the archive’s primary mission is to acquire, process, preserve, and disseminate high quality datasets relevant to the study of child maltreatment. we have the secondary mission of networking and training child maltreatment workers. toward this end, the archive publishes a biannual newsletter, hosts a listserv with approxixmately 400 subscribers, and maintains a gopher/ftp server. in many ways we have been more successful in achieving our secondary mission of networking and training researchers than our primary mission of acuirring datasets. creating networking and training opportunities is a fairly straightforward job and such services are eagerly consumed by researchers. acquiring, processing, and disseminating high quality data is a far more complex task. for these and other reasons, ndacan began to advocate for mandatory data archiving for nccan research grantees. simultaneously, we lobbied nccan to establish tecnical standards for their research grantees. toward this end, jane powers and i co-authored, the preparation of data sets for analysis and dissemination: technical guidelines for machinereadable data, a manual which set forth standards for the preparation of research datasets and their associated documentation. we disseminated hundreds of copies of this manual to cnnan’s research grantees and to other child maltreatment researchers and we offered technical assistance to researchers willing to follow our guidelines. unfortrunately very few researchers responded with interests. the situation changed however when, in their 1993 rfp, nccan announced that their research grantees would be expected to prepare datasets and documentation according to ncacan’s standrads. this generated quite a bit of interest and for the first time applicants began to contact us for technical assistance and copies of our manual. the archive’s lobbying efforts came to fruition when, in their 1994 rfp, nccan set forth the requirement that applicants include in their proposal plans to prepare their data and documentation according to ndacan’s guidelines and to archive 22 iassist quarterly their data with ndacan upon the completion of their grant. as a result of this dramatic policy change, ndacan will have the opportunity work with investigators from the beginning of their projects to ensure that data and documentation are prepared properly. all grantees will be provided with free technical assistance during the start-up phase of their studies and will receive a new publication entititled, depositing data with the national data archive on child abuse and neglect: a handbook for investigators. the purpose of the handbook is to outline the investigators responsibilities and to provide a clear set of deliverables that must be submitted to ndacan. in some sense nccan’s policy change took us by surprise. after years of lobbying, we were happily surprised to learn that nccandecided to require research grantees to archive their data. the way the requirement was implemented was that nccan reviews rated applicants’ plans to prepare and archive data and documentation as one of the many criteria use to evaluate grant proposals. while it is not clear that this arrangement provides any method of enforcement, grantees are working under the assumption that they will be required to archive their data. while nccan set forth the requirement, the archive is in the position of defining all of the nuts and bolts of the arrangement. instead of reinvention the wheel, we have studied the data archiving programs of other federal agencies, such as the national institute of justice (nij) and the national science foundation (nsf). our approach has been to build on the successes of these programs and make adjustments where necessary. our goal is to build a program that meets the needs of ndacan and nccan’s research grantees. in our experience of working with researchers, their greatest concern is having adequate time to publish their results of their study before the data are released to the public. in response, we have created a policy that will allow all investigators a twoyear “grace period” after the termination of their grant which will allow them to publish the results of their study before the data are made available to the public. this is an area where our approach differs from that of nij. nij grantees must submit their data and documentation along with their final report at the termination of the award; grantees who fail to do so are not eligible for new nij funding. this policy has created a great deal of animosity among some nij grantees and there have been cases where a secondary data user published a study’s findings before the principal investigator. in other areas, we have closely followed nij’s lead. for example, in determining the investigators’ responsibilities and required deliverables we have essentially mimicked nij’s requirements. broadly defined, we see the investigators are responsible for, submitting data and sporting materials, responding to requests by ndacan staff for additional or clarifying information, and reviewing and correcting draft materials prepared by the ndacan staff. ndacan staff is responsible for preparing the data files in ready-to-use statistical file formats, preparing a user’s guide that describes the project and data, reviewing the codebook for completeness and accuracy and augmenting the codebook as necessary, and making copies of the datasets archived available to the research community (for a small fee) and providing technical support to data users. this new arrangement has the potential to solve the problem of inadequate and inconsistent documentation because we can specify exactly what materials the investigator must submit. grantees will be provided with a clear set of deliverables as well as clear written guidelines for the preparation of those materials. working with grantees early on in their projects in order to determine potential problem areas and needs for technical assistance will be integral to our approach. while it will be several years before the first grantees are required to submit data under this arrangement we have established a tentative list of deliverables. these include: (1) data file(s). (2) description of data files (3) data collection instruments(s) (4) references for data collection instruments (5) codebook or data dictionary (6) explanation of derived (computed) variables (7) final project report, project summary, or other description of the project (8) bibliography of publications pertaining to the data (9) printout of the first and last data records the draft handbook that i have distributed contains some guidelines and specifications for the preparation of these materials. our plan is to distill the most important guidelines in our technical standards manual and include them in the handbook. we want to keep the handbook as simple and free from jargon as possible. once it is completed, we will stop distributing the technical standards manual because it is both out-dated and too technical for investigators. we are very interested in your 23fall 1996 feedback and suggestions for the handbook so please read it if you have time. and let us know how we can improve it. once we receive these materials from the investigators we will create a comprehensive user’s guide with the following format: project overview purpose of the study sampling/selection information data collection instruments and measures description of machine-readable files list of files notes regarding the data files references references to publications from the dataset references to publications related to the dataset appendices data collection instruments codebook sample programs so far reaction to the policy has been relatively positive. we presented our general plan to a large group of researchers at the annual nccan grantees meeting and most of the grantees were receptive. we are in the process of forming an advisory committee to handle cases where investigators have special needs relative to archiving (e.g., longitudinal studies). we are hopeful that nccan’s data archiving policy will go a long way toward solving the problems i have described however, it will be several years before we know for sure. clearly, the approach has limitations but we feel it is a step in the right direction. 1. submitted for the iassist conference held in quebec, canada. may 1995 iassist quarterly 2014 21 iassist quarterly mapping the general social survey to the generic statistical business process model: norc’s experience by scot ausborn, julia rotondo and tim mulcahy1 abstract as a part of the metadata portal project, with support from the national science foundation, norc mapped the general social survey workflow to the generic statistical business process model (gsbpm) to determine where in the survey cycle ddi-based metadata could be more effectively captured. lessons learned from the process include a better understanding of utilizing the flexibility of the gsbpm model and a recommendation to collect paradata in a collaborative, facilitated workshop rather than mapping responses from individual staff. information gained from the mapping has proven useful in identifying areas of metadata and paradata collection enhancements. keywords: metadata, paradata, data documentation initiative (ddi), generic statistical business process model (gsbpm), workflows background the metadata portal project, funded by the national science foundation as part of the metadata for long-standing social science surveys (meta-sss, ses-1229957) initiative, is a collaborative effort among the general social survey at norc at the university of chicago, the american national election study at the university of michigan, and the inter-university consortium for political and social research with technical assistance from metadata technologies north america (mtna). the project’s objectives are: • to develop rich, structured metadata compliant with the data documentation initiative (ddi) standard for two premier time series studies in the social sciences — the gss and the anes • to showcase tools that can be built upon the foundation of rich metadata • to analyze and improve the projects’ workflows to capture more metadata at the source the primary deliverable of the project is a web-based portal leveraging ddi-compliant metadata to provide a range of new tools for researchers working with gss and anes data. the portal incorporates an enhanced search engine for both datasets, comprehensive variable and concept banks, and a subsetting feature for generating custom datasets. one of the project tasks supporting the creation of the metadata portal – and intended to sustain it going forward – is an analysis of the business processes surrounding the production of survey data to determine where in the survey cycle ddi-based metadata may be captured to avoid having to generate it retroactively. by taking this initial step the gsbpm is a schema for parsing statistical production workflow that consists of nine high-level processes 22 iassist quarterly 2014 iassist quarterly toward a ddi-based workflow, the goal is to enhance the metadata available to researchers in the portal and to realize greater efficiencies in the survey cycle itself by identifying redundant processes, such as duplicate data transformation, that could be remediated with a metadata-based approach. the gsbpm in order to better understand the workflow processes associated with the production of the general social survey (gss), norc conducted a survey of internal gss staff asking them to explicate their respective roles on the survey in terms of the generic statistical business process model (gsbpm)4 . the gsbpm is a schema for parsing statistical production workflow that consists of nine high-level processes and several sub-processes under these.: correlating aspects of the gss workflow to elements of the gsbpm allowed norc to gain a comprehensive and integrative view of the individual efforts that together produce the survey. additionally, gathering the gss paradata in this manner also facilitated the identification of processes in the workflow where metadata relevant to dissemination and discovery of the survey data is potentially being lost or ineffectively captured. by identifying and remediating these points, it is intended that the survey be produced more efficiently while better meeting the needs of researchers analyzing the data questionnaire design the questionnaire distributed to internal gss staff (hosted online at http://dataenclave.org/gss) was adapted directly from the language used in gsbpm descriptions of individual processes and source: http://www1.unece.org/stat/platform/display/metis/the+generic+statistical+business+process+model sub-processes. as a result, the questionnaire followed a highly structured format. respondents were asked to provide detailed information regarding the facets (inputs, outputs, actions, tools, etc.) of each sub-process in addition to a brief overview of the subprocess itself. one of the challenges of using the language provided by the gsbpm is that it is highly abstract, requiring some deduction to understand how the process being described corresponded with internal gss processes. thus in creating the questionnaire, additional explanatory text was required to help tailor gsbpm language to gss specific processes. a gss staff member with experience in several different aspects of the survey workflow was essential in order to create the additional explanatory text. similarly, the volume of gsbpm sub-processes required the selection of appropriate respondents for a particular sub-process in order to prevent survey fatigue. on the technical side, it was determined that web-based dissemination of survey questions would best facilitate data collection and analysis. to implement this norc used a standard lamp-stack design, with the webpage coded in standard html/ css and data stored in a mysql database using php. the database schema used a separate table for each sub-process, with the respective columns storing the respondent’s id, an overview of the process from their perspective, and the different facets of that process. iassist quarterly 2014 23 iassist quarterly survey execution prior to the survey link being distributed to the respondents, an email from the senior vice president of the department producing the gss was sent to reinforce the value of the survey and the expectations of completing it. once the survey link was distributed, respondents were given one week to enter the information for the specific sub-processes that were assigned to them. respondents were allowed to quit the survey and return later; however, they were unable to retrieve previously saved answers at a later time. in retrospect, allowing respondents to login and retrieve saved answers would have been helpful for respondents, but needed to be balanced against time to develop and implement. another possibility for achieving this functionality would have been to set a cookie in the respondent’s browser. given the direct request from senior management to respondents, the survey garnered a high response rate. however, the responses indicated that some respondents may not have been targeted well, with a few stating that the sub-processes they had been assigned to provide information for were not part of their work with the gss . collecting, compiling, and cleaning internal survey responses once the survey collection period had ended, the survey responses were downloaded from the mysql database to a csv file. from there, the file was opened in excel and cleaned to ensure that within a sub-process each respondent’s answers were only captured once. some respondents had encountered technical difficulties with the form, mostly due to browser compatibility issues, and ended up submitting the same answers in excess of five times for the same sub-process. responses that did not add new information to the survey (e.g., a respondent entering “skip” or “n/a” into the comment box) were deleted from the file as well. finally, formatting was introduced to improve legibility of the responses for analysis. mapping to the gsbpm after the responses had been cleaned, the norc team examined the responses given for each of the gsbpm sub-processes and attempted to create a comprehensive overview of the process. challenges became immediately apparent in the process, including: • determining if responses truly belonged in the sub-process they were placed in by respondents • determining what happened in cases where no responses were given for sub-processes • within the gss, it was often the case that multiple subprocesses were happening simultaneously while in the model they occurred linearly when consolidating the survey responses into an overview, research analysts noted that respondents would often reply to one sub-process with information that might better fit in another. for example, some responses were submitted under the 2.5 design statistical processing methodology that upon review seemed to be a better fit for 2.2 design data collection methodology. in instances such as these, the responses were moved to the new section, but annotated so that the team could track how responses had moved. responses were moved because while the respondents were the experts in the gss, they were not as familiar with the gsbpm as the research analysts working on mapping the gss to the gsbpm. therefore, while the team worked to ensure that all information submitted to the gsbpm was included in the combined workflow overview, if the responses given to a particular section seemed a better fit to another section, the team decided to move the response. other challenges included having no responses for parts of a sub-process or entire processes. for example, norc staff did not submit any responses for many of the sub-processes within the 1.0 specify needs section of the gsbpm, including 1.3 establish output objectives and 1.4 identify concepts. because there 24 iassist quarterly 2014 iassist quarterly were no responses, the team did not include these sub-processes in the overview. finally, a difficulty was that one sub-process within the gsbpm was sometimes too broad to clearly show multiple processes working simultaneously. for example, within the gsbpm sub-process 4.2 set up collection, many distinct action items are performed by different gss teams to accomplish this task. the responses revealed that while certain common steps occurred within that sub-process, it actually contained three separate team processes – each with their own steps, paperwork, inputs, outcomes, and purposes. the team kept all the responses together within the sub-process narrative to maintain cohesion within the gsbpm, but as can be seen below by the diagram of the sub-process, it was not a natural fit. lessons learned overall, the challenges faced by the research team in mapping the gss workflow to the gsbpm can be traced to a rigid adherence to the gsbpm model. the research team began the process of collecting paradata with the model and then asked gss staff to discuss their processes within its framework, leading to poorly fitting sub-processes – where some sub-processes are empty while others are so full that they lose clarity. in retrospect, a better process might have been to start by asking gss staff to detail their process, map out the steps, and then see how that process compared to the model. in that respect, the norc team did not fully exploit the main benefit of the gsbpm – namely that the tool is meant to be a customizable starting point rather than a rigid endpoint. going forward, if norc were to conduct this study again, we would hold a workshop in which gss staff would be able to engage with one another and discuss the workflow processes rather than having each person provide his or her input in isolation. furthermore, it would be highly beneficial to have an expert in gsbpm (or perhaps the complementary generic statistical information model) to conduct the mapping of workflow rather than asking staff members to conceptualize their work in terms of the abstract language provided by the gsbpm model. by doing this, staff would simply describe what they do, rather than reacting to a question or sub-process that might be interpreted as having no relevance to their work. nevertheless, while the method of gathering gss paradata had its difficulties, the information gleaned from this study has proven useful in terms of identifying points in the workflow where metadata might be enhanced as well as how the collection of paradata using a gsbpm or similar model might be improved upon. it is the desire of the norc team that this experience should prove instructive for other institutions wishing to conduct a workflow study of their own statistical production processes. references andritsos, periklis and keilty, patrick (2014) level-wise exploration of link notes 1. scot ausborn is a systems engineer for the data enclave at norc. email: ausborn-scot@norc.org 2. julia rotondo is a senior research analyst at norc. email: rotondojulia@norc.org 3. tim mulcahy is the program area director of the norc data enclave. email: mulcahy-tim@norc.org 4. http://www1.unece.org/stat/platform/display/metis/the+generic+s tatistical+business+process+model sist newsletter vol.1, no. 4 chai r person' s report the european lassist action group workshop was held at the danish data archive in copenhagen, june 26-29, 1977, under the direction of lassist west european secretary, per nielsen. a major product from this meeting was a data organization registration form (dorf) which is reproduced in this newsletter . copies of the dorf can be obtained by writing to elliott avedon, department of recreation, university of waterloo, waterloo, canada. the completed form should be returned to him. a study description questionnaire form developed by the zentralarchiv at cologne was reviewed and discussed. several archives are now testing this instrument. individuals interested in examining or testing the instrument should contact elliott avedon for copies. the workshop also addressed levels of documentation for machine-readable files. the results of this discussion are reported in this newsletter . the program for the february 8-11, 1978, lassist conference at nordic hills, itasca, illinois has been revised on the basis of membership feedback. the following panels have been scheduled: documentation, alternative structures for data access, problems and potentials of networking, privacy versus freedom of information, acquisition and preservation, and software analysis of non-rectangular files. in addition, workshops will be offered and action group meetings will be held. although panelists have been designated for each session, the sessions have not been closed. individuals interested in presenting a paper at any of the sessions should send the proposed title with a 100 word abstract to carolyn geda. please note that individuals who were following the original program theme and who anticipated doing a formal presentation within an action group should submit their paper to the chairperson. the schedule for the panels, workshops and action group meetings will be sent to lassist members with conference registration forms during november. please note that the conference will terminate on saturday, february 11, 1978. additional information on the february meetings is contained in the lassist newsletter section on "upcoming meetings." the panels for the lassist sessions at the international sociological association world congress are privacy versus freedom of information (chairperson guido martinotti), issues in comparative data and research (chairperson elliott avedon), and research problems associated with complex data bases (chairperson john de vries). individuals interested in presenting a paper should contact the chairperson as soon as possible. information on the availability of charter flights will be sent to lassist members in the near future. this issue of the newsletter completes the first volume and also contains an index for the four issues in this volume. the first volume has been edited and produced by alice robbin, university of wisconsin-madison. the effort she has devoted to the newsletter has been widely recognized and appreciated. she will remain affiliated with the newsletter in an ex officio capacity. articles and notices to be included in the newsletter should continue to be sent to her. the lassist newsletter will move to western kentucky university where thomas madron, academic computing and research services, will assume the role of editor. a nominating committee will be appointed shortly to prepare a slate for the election of the steering committee members. if you are interested in serving on this committee or on the steering committee or wish to recommend other individuals or nominations, please forward the names to carolyn geda. erratum in an earlier version of the article "journals in economic sciences: paying lip service to reproducible research?" by vlaeminck and podkrajac, the following appearred uncited on page 2: "while in other sciences replicability is regarded as a fundamental principle of research and a prerequisite for the publication of results, in economic sciences it is not treated as a top priority. in 2006..." this should instead have been cited as follows: "according to höffler (2017), replicability does not take a high priority in economics. in his opinion, this is in sharp contrast to other sciences where replicability is "regarded as a fundamental principle of research and a prerequisite for the publication of results" (höffler, 2017, p.1). already in 2006..." 1/17 vlaeminck, sven and podkrajac, felix (2017) journals in economic sciences: paying lip service to reproducible research? iassist quarterly 41 (1-4), pp. 1-16. doi: https://doi.org/10.29173/iq6 journals in economic sciences: paying lip service to reproducible research? sven vlaeminck1 and felix podkrajac2 abstract the findings of numerous replication studies in economics have raised serious concerns regarding the credibility and reliability of published applied economic research. literature suggests that economic research often is not replicable because (i) only a small proportion of journals in the field have implemented functional policies on the disclosure of employed datasets and program code, (ii) authors frequently do not comply with these data policies and (iii) editorial offices do not ensure that these policies are enforced. in this paper, we focus on the aspect last mentioned. we empirically evaluate 599 articles published in 37 journals with a data availability policy. we present the share of articles that fall under a data policy, because replication data is needed to verify the published results. afterwards, we check the journal data archives and supplemental information section of each article for the availability of replication files. for a reduced sub-sample of 245 data-based articles, we check in depth whether the replication files we found are compliant with the requirements of the journal’s respective data policy. thereby, we are able to determine how much journals in economic sciences enforce their data policies. our findings suggest a mixed picture: while some journals achieve high compliance rates, a significant share of journals only sporadically provides replication files for databased research papers. keywords reproducible research, economics, journals, data policies, data archives, 1. introduction in economics, data-based3 research has become increasingly important. according to the us economist hamermesh (2013), the number of contributions to top journals in which authors utilised self-collected or borrowed datasets, employed experimental designs, or used real data for simulations of theoretical models has massively increased over the last decades. while in 1963 the share of publications in economic journals with a solely theoretical orientation was 50.7%, this percentage dropped to 19.1% in 2011. with the growing relevance of data-based publications, new questions and challenges for academic publishing emerge. one of the most pressing challenges is the validation of published data-based research, which is closely connected to the principle of replicable research, as a cornerstone of the scientific method4: ’replication ensures that the method used to produce the results is known. whether the results are correct or not is another matter, but unless everyone knows how the results were produced, their correctness cannot be assessed. replicable research is subject to the scientific principle of verification; non-replicable research cannot be verified. second, and more importantly, replicable research speeds scientific progress. […] third, researchers will have an incentive to avoid sloppiness. […] fourth, the incidence of fraud will decrease’ (mccullough, 2009, p.118f.). replication refers to the duplication of scientific findings, but literature shows that the common understanding of replication differs widely: clemens, for instance, lists several terms used by https://doi.org/10.29173/iq6 2/17 vlaeminck, sven and podkrajac, felix (2017) journals in economic sciences: paying lip service to reproducible research? iassist quarterly 41 (1-4), pp. 1-16. doi: https://doi.org/10.29173/iq6 researchers (e.g. reproduction, verification, extensions, robustness tests,…) which all refer to replication. he concluded that ‘there is no consensus meaning of the term replication’ (clemens, 2015, 7). for our paper, we employ the definition of hamermesh’s ‘pure replication’, which he defines as to ‘duplicate, repeat, as in a statistical experiment’ and ‘to make or do something again in exactly the same way’ (hamermesh, 2007). according to höffler (2017), replicability does not take a high priority in economics. in his opinion, this is in sharp contrast to other sciences where replicability is “regarded as a fundamental principle of research and a prerequisite for the publication of results” (höffler, 2017, p.1). already in 2006, mccullough et al. criticised the status quo in economics: ‘results published in economics journals are accepted at face value and rarely subjected to the independent verification that is the cornerstone of the scientific method. most results published in economics journals cannot be subjected to verification, even in principle, because authors typically are not required to make their data and code available for verification.’ (mccullough, mcgeary & harrison, 2006, p.1093 f.) mccullough’s negative appraisal is based on the findings of his own studies and those of other researchers in the field, who systematically tried to replicate published applied economic research. one of the first of these studies was released by the economists dewald, thursby and anderson in 1986. in their seminal paper, the researchers collected programs and data from authors of the journal of money, credit and banking (jmcb) with the aim to replicate the results of 54 published articles. they were able to replicate the key findings of only two papers (3.7%) – a result that has led to an ongoing debate on reproducible research in economics (dewald, thursby & anderson, 1986). the data policy of jmcb at that time was based on the willingness of researchers to cooperate in cases of requests for data and code – which we label an ‘author responsibility policy’ (arp). but this willingness does not conform to the incentive structures in science and research, as mirowski & sklivas (1991) and feigenbaum & lewi (1993) point out. one important reason for the replication crisis in economics is based on such missing incentives for researchers to share their data. as duvendack, reed and palmer-jones (2015) argue, authors of original studies are concerned about the costs of compiling data and program code into (re-)usable forms. original authors may expect that the benefits of providing well documented, easily usable code are small or even negative. little credit is given to the original author if the replicating authors confirm the original results, while the damage to reputation may be large, if the original results cannot be confirmed (hamermesh, 2007). this situation is aggravated by the fact that replication is still viewed as a ‘parasitic activity’ by some researchers (hamermesh 1997, c.f. longo & drazen, 2016). also the wish to retain exclusive rights to data that had taken a lot of time to produce is an important reason for researchers not to share their data (savage & vickers, 2009 and mueller-langer & andreoliversbach, 2014). other concerns comprise legal issues and the fear of a misinterpretation/misuse of data (tenopir et al, 2009). these missing incentives to share data have also been noticed by dewald, thursby and anderson: the researchers reported that one-third of the authors never replied to their repeated requests, and an additional one-third replied that they could not furnish their programs or data. dewald, thursby and anderson concluded: ’our findings suggest that the existence of a requirement that authors submit to the journal their programs and data along with each manuscript would significantly reduce the frequency and magnitude of errors’ (dewald, thursby & anderson, 1986, p.588). https://doi.org/10.29173/iq6 3/17 vlaeminck, sven and podkrajac, felix (2017) journals in economic sciences: paying lip service to reproducible research? iassist quarterly 41 (1-4), pp. 1-16. doi: https://doi.org/10.29173/iq6 the implementation of such mandatory data policies and corresponding data archives that require both data and program code from authors of data-based papers could have several benefits, as highlighted by anderson et al.: with mandatory data and program code archives, checking robustness of a data-based paper should be a comparatively simple matter. data and specifically program code should provide a record of how the published research was produced. researchers who wish to build on previous research therefore no longer have to program everything from scratch. consequently, not only can data policies and data archives increase the accuracy of the published results, but they could also be better examined for robustness and can be more easily extended, thus also increasing the quality of research (anderson, greene, mccullough & vinod, 2008).5 in fact, some journals altered their data policy after publication of dewald’s, thursby’s and anderson’s article towards a mandatory data availability policy (abbreviated in the following to ‘dap’). a dap normally asks authors of data-based papers to provide their replication files to the editorial office prior to an article’s publication. which data and files are exactly needed to facilitate replications is exemplarily outlined by king (1995a), but also depends on the methodology utilised.6 the federal reserve bank of st. louis review (1993), the journal of money, credit and banking (1996), studies in nonlinear dynamics & econometrics (1996) and macroeconomic dynamics (1996) were among those journals who implemented or altered their data policy after dewald’s, thursby’s and anderson’s findings till the turn of the millennium (mccullough, 2009). but despite changes in journals’ data policies and implementations of data archives, replication studies still reported that only a minority of applied research papers were replicable: mccullough, mcgeary and harrison (2006) again tried to replicate papers published in jmcb by using data from the journal’s data archive. 193 of 266 articles published between 1996 and 2003 should have had entries in the data archive, but only 69 did. of these, the researchers tried to replicate 62, but succeeded only 14 times (22.6%). as a consequence, mccullough (2007, 327ff.) published recommendations for managing a journal’s data archive and listed files and conditions required for successful replications. in a subsequent study, mccullough, mcgeary and harrison (2008) tried to replicate 117 articles published by the federal reserve bank of st. louis review. only 9 (7.7%) papers could be reproduced. also a more recent study reported ongoing difficulties with the replication of published economic research: chang and li (2015) could replicate qualitative key results from 33% of 67 papers published in 13 well-regarded economics journals. with support from the authors, the share increased to 49% – which is still below half. the authors found that their replication success rate was significantly higher when they attempted to replicate papers from periodicals that have a mandatory replication data and program code submission policy. but how can it be explained that so many replication attempts in economic sciences fail, despite more journals starting to implement data policies and corresponding data archives? to answer this question, mccullough, mcgeary and harrison (2008) examined the data archives of four economics journals (federal reserve bank of st. louis review, journal of business and economic statistics, journal of applied econometrics and economic journal) and compared the number of data-based articles with the entries in the journals’ data archives. first, the researchers examined all articles in these journals, classifying each as requiring an entry in the archive or not. an article was classified as requiring an archive entry if the article displayed or otherwise represented numbers. therefore they included classically ‘empirical’ articles, as well as computational economics articles. subsequently, the authors https://doi.org/10.29173/iq6 4/17 vlaeminck, sven and podkrajac, felix (2017) journals in economic sciences: paying lip service to reproducible research? iassist quarterly 41 (1-4), pp. 1-16. doi: https://doi.org/10.29173/iq6 checked the data archives for entries of the respective articles. for each issue investigated, they show the number of articles that utilised data and therefore should have an archive entry, the number of articles that actually had an archive entry, and the ‘compliance’ percentage of articles that have an entry. as a result, they found compliance rates reaching from 13% (economic journal) to 99% (journal of applied econometrics). with this paper we provide an updated evaluation of journals’ practises when it comes to the enforcement of their data policies. because the availability of data and program code can be regarded as a prerequisite to validate the findings of published research, it is important to monitor the current practises of journals in economic research. especially against the background of the replication crises in the social sciences, this aspect is of special interest. our approach follows the study of mccullough, mcgeary and harrison (2008) in several aspects: we regard the practise of 37 journals in economic sciences – all of them have a dap. we give the number of articles that are data-based and therefore should have an entry in journal’s data archive (or if no separate data archive exists, the replication files should be available in the supplemental information section of the article). afterwards, we investigate for how many of these articles replication files are available. we distinguish between papers which are using restricted and non-restricted data, because dap often contain varying rules for research based on proprietary or confidential data (vlaeminck & herrmann, 2015).7 furthermore, we compare the share of articles with supplemental replication data of journals with a voluntary data policy to those with a mandatory data policy. by doing so we determine empirically how well voluntary and mandatory data policies perform against the background of missing incentives for data sharing. with a reduced sample of journals we subsequently evaluate in depth whether journals follow and enforce their data policies. for this purpose, we compare the requirements of the journal’s respective data policy to the replication files we found for each published data-based article. our paper is organised as follows: section 2 describes the methodology to collect the data for our analysis. section 3 presents the results of our study, while section 4 summarises and discusses the findings of the analyses. 2. data and methodology in the following subsections, the approach of our study is described. section 2.1 describes previous research and a dataset we utilised as a starting point for our analyses. section 2.2 sets out how we determine the share of data-based articles and expounds how we ascertain the share of accompanying replication files. section 2.3 explicates the methodology for the analysis of journals’ compliance rates with their data policies. 2.1 datasets available for our study, we used a publicly available dataset of vlaeminck and herrmann (2015b) as a basis. the corresponding research paper (vlaeminck & herrmann, 2015a) presents the results of an analysis of journals’ data policies in economic research. the authors evaluated the data policies in a sample of 346 journals and found 49 journals with a dap (14.2%). to compile the sample for their analyses, the https://doi.org/10.29173/iq6 http://www.edawax.de/wp-content/uploads/2016/05/alle-untersuchten-journals.pdf http://www.edawax.de/wp-content/uploads/2016/02/data-policies_mit-links.pdf http://www.edawax.de/wp-content/uploads/2016/02/data-policies_mit-links.pdf 5/17 vlaeminck, sven and podkrajac, felix (2017) journals in economic sciences: paying lip service to reproducible research? iassist quarterly 41 (1-4), pp. 1-16. doi: https://doi.org/10.29173/iq6 authors employed two lists of academic journals assembled by economic associations. one list (jourqual 2.1) is maintained by the german academic association for business research (vhb). vlaeminck and herrmann chose all journals from the jourqual list ranked a+, a or b and added each 60 journals ranked c, d and e to their sample. besides, a list of journals analysed by bräuninger, haucap and muck (2011) has been added to the research sample. in a next step, vlaeminck and herrmann removed double entries and specified the primary subject category of these periodicals according to their classification in thomson reuters journal citation reports and zbw’s indexing guidelines. 262 out of 346 (75.7%) journals investigated were listed in thomson reuters journal citation reports (jcr) 2013 and therefore have an impact factor. in their dataset, vlaeminck and herrmann also set out which of the 49 dap are mandatory and which are voluntary for authors. also they stated for each journal in their dataset what the data policies require from journals’ authors and how these journals provide replication files. for our analysis, we include 37 (75.5%) of these 49 journals with a dap in our research sample. the remaining journals have been excluded because either the subject category is not primarily located in economics (e.g. science, nature, pnas) or because we were not able to access the articles of these journals due to licensing restrictions (e.g. journal of law and economics, journal of labor economics, european accounting review,...) 2.2 approach to determine the share of data based articles to determine the share of data-based articles in our sample, we investigate 599 articles published by 37 journals in the issues 1/2014 (323) and 1/2013 (276). we choose only regular issues. in the event that one of the issues is a special issue, we use the following issue. this approach is necessary, because special issues tend to focus on single research questions or topics which might result in a bias towards a certain methodology. for each journal, we examine the whole issue. this leads to disparate numbers of articles per periodical in our sample: while some journals only publish four articles per issue, others publish 20 and more. therefore, journals like the journal of the american statistical association (9.8% of all articles within our sample), the review of economics and statistics (6.3%) or the journal of economic dynamics and control (5.8%) have the highest numbers of articles in our sample, while journals like the journal of political economy (0.8%), the econometrics journal (0.8%) or the jahrbuecher fuer wirtschaftswissenschaften/review of economics (0.8%) only have a few articles in our sample.8 subsequently, we examine the content and methods used by each of the 599 articles. first of all, we study the title and the abstract of each to determine whether an article is data-based. if we are still unsure, we go through the article and manually search for specific terms or keywords like ’data’, ’simulat*’, ’significant*’, ’experiment*’, ’evidence’, ’empirical*’ and read the specific paragraph.9 we only consider original research articles and remove from the survey literary material like review articles, editorials, comments and rejoinders. we define an article to be ’data-based’ when the methods include creating, reusing, and analysing research data, other empirical methods (e.g. experiments) or self-compiled program code (e.g. simulations). for two reasons, our approach is comparatively strict: while a few data policies10 consider papers not to be empirical if they use data only for case studies (for instance, some simulations use real data to demonstrate the effects of a theoretical model), we consider these articles https://doi.org/10.29173/iq6 http://vhbonline.org/en/service/jourqual/vhb-jourqual-archiv/vhb-jourqual-21-2011/ http://ipscience.thomsonreuters.com/product/journal-citation-reports/ 6/17 vlaeminck, sven and podkrajac, felix (2017) journals in economic sciences: paying lip service to reproducible research? iassist quarterly 41 (1-4), pp. 1-16. doi: https://doi.org/10.29173/iq6 to be data-based. first, would-be replicators need the program code of the simulation to verify the computation. second, when real data applications have been used, we also expect these case studies to be replicable – even though the case study is not the main result of the paper. our approach is in line with the specifications of the most robust data policies, like the one of the american economic review (aer) (c.f. glandon, 2010). in addition, we determine whether an article uses proprietary or confidential data (for instance, purchased datasets, firm data or microdata) to estimate the share of articles that rely on restricted data. we examine these articles for references on datasets employed by the authors. this information is of interest for our approach described in section 2.3, because several data policies possess special rules for research based on restricted data. afterwards, we check whether the data-based articles are accompanied by replication files (e.g. datasets and/or program code). in case of restricted data, we look for program codes and further references to these datasets. for this purpose, we check both the publisher’s web page and the separate data archive, if applicable,11 for replication files. for each data-based article, we note which files we found. 2.3 approach on journals’ data policies compliance rate to examine journals’ compliance with their own data policies, we collect some additional information for all of the data-based articles. because we have to ensure that a journal’s data policy was already in effect at the time of article’s submission, we note the publication history for each and the implementation date of journal’s data policy, if available. this approach is necessary because normally a paper only falls under a data policy, when the policy is effective at the time of submission. also, some journals (e.g. the journal of the american statistical association) demand authors to submit their data already with the submission of their manuscripts. at that point, certain difficulties arise because not many journals provide the publication history of their articles and also the implementation date of a data policy often is not easy to confirm, because only few journals explicitly mention the date when the policy took effect. to solve these challenges, we first determine how long the publication process in economics usually lasts. literature shows that the publishing process in economic journals can easily exceed more than a year: ellison (2002) ascertained the average review time (from submission to acceptance) of 16.5 months for 25 top economics journals in 1999. björk and solomon (2013) found that the whole publication period from submission to publication lasts 18 months, on average. besides, by reusing a list containing the full text of journals’ data policies compiled by the edawax project in the spring of 2012, we made certain that there are at least 29 journals with a dap at that time. for some other journals, we could determine the date the policy took effect, either because it was mentioned in the literature or because the policies specified the point in time at which they became effective. as a consequence, we only keep articles in the sample for the compliance analysis, if either we are able to determine the exact submission date of the article and know the journal’s data policy was effective at that point in time or if we know the journal’s data policy was implemented more than 24 months prior to the article’s publication. according to björk’s and ellison’s findings on the average https://doi.org/10.29173/iq6 https://www.aeaweb.org/journals/policies/data-availability-policy https://www.aeaweb.org/journals/policies/data-availability-policy http://www.edawax.de/wp-content/uploads/2012/07/data_policies_wp2.pdf 7/17 vlaeminck, sven and podkrajac, felix (2017) journals in economic sciences: paying lip service to reproducible research? iassist quarterly 41 (1-4), pp. 1-16. doi: https://doi.org/10.29173/iq6 publication times in economic journals, it is thus very likely that a journal’s data policy had already taken effect at the time the article was submitted. in a next step, we contrast the replication files we found in the journal’s data archive or in the supplemental information section of a paper to the specifications of the journal’s respective data policy. for instance, if the data policy requires the disclosure of data and program code but only the data has been provided, we consider such an article to be not compliant with journal’s data policy. special handling is needed for research based on restricted data. for this article type, we initially consult the data policy of the respective journal. if it mentions specific requirements (e.g. to post the program code and some contact information to access the dataset), we use this requirement to check whether the policy is fulfilled. in cases of data policies without such a procedure, we exempt the article from our analysis, because such an article does not fall under the journal’s data policy. by applying this approach, we are able to determine for any article and any data policy whether the policies’ requirements have been fulfilled – regardless of whether we consider the policy itself to be robust or weak, whether it is mandatory or voluntary. 3. findings of our study in the following subsections, we present the empirical findings of our study. first, we show the share of data-based papers in our full sample. subsequently, we depict the amount of databased articles for which we find replication data and present the respective outcomes for journals with mandatory and voluntary daps. second, we show the findings of the compliance analysis. we illustrate these results both on article-level and on journal-level. again, we pay attention to potential differences between mandatory and voluntary data policies. figure 1 graphically illustrates the different sample sizes of our analyses. 3.1 the share of data-based articles in total, we examine 599 articles published by 37 journals. because the sample of our study builds on journals with a dap, we expect the share of databased articles to be higher than normally anticipated for journals in economic sciences. based on our approach described in section 2.2, we classify 452 (75.5%) of all articles to be data-based (see figure 2). hence, these papers (at least those who fall under a mandatory dap) should have an entry in journal’s data archive or in the supplemental information section of the journal. when further examining these 452 data-based articles, we find that 284 of them (62.8%) employed non-restricted data for their analyses. after checking journals’ data archives and the supplemental figure 2: the share of articles categorised as ’data-based.’ figure 1: the research samples at a glance https://doi.org/10.29173/iq6 8/17 vlaeminck, sven and podkrajac, felix (2017) journals in economic sciences: paying lip service to reproducible research? iassist quarterly 41 (1-4), pp. 1-16. doi: https://doi.org/10.29173/iq6 information section of these papers, we find accompanying replication datasets for 104 (36.6%) papers and program code for 103 (36.3%). table 1. replication files and references to restricted data within our sample of data-based articles (n=452) type of article data available code available information on restricted data articles using non-restricted data (n=284 / 62.8%) 104 (36.6%) 103 (36.3%) n.a. articles using restricted data (n=168 / 37.2%) n.a. 51 (30.4%) 75 (44.6%) another 168 articles (37.2%) employ proprietary or confidential data to generate their findings. of these, 75 (44.6%) provide some references on the data utilised (see table 1). most often this information is available from the article only. we believe this information is frequently not sufficient to exactly determine which dataset was used for the analyses (although we did find some well documented examples12). in many cases, only the name of the dataset is mentioned, but no further references. to identify a dataset based on this information might work for some, but for the majority it will not (e.g. because there are several waves or corrections of a survey/dataset). for 51 (30.4%) of these papers, the program code is also available. the importance of mandatory data policies for data disclosure in our sample, 265 out of 452 (58.6%) articles examined fall under a mandatory dap. 155 of these (equates 34.3% of the sample) rely on non-restricted data. 98 of these (63.2%) possess accompanying datasets, and the program code is available for 96 (61.9%) articles (see figure 3). also for the 110 articles (equates 24.3% of the sample) based on restricted data, we ascertain that the program code exists for 46 papers (41.8%). 59 (53.6%) also give some information on restricted datasets employed in the research process. the remaining 187 articles have been published by journals with a voluntary dap. 129 of these (equates 28.5% of the sample) rely on non-restricted data. for these articles, replication datasets are available for six (4.7%) papers. program code is available for seven articles (5.4%). the shares for the 58 articles based on restricted data in journals with a voluntary data policy are only slightly higher: program code is available for five articles (8.6%) and 16 (27.6%) provide some information on the data employed (see table 2). figure 3: the sample of all data-based articles by type of data employed (restricted / non-restricted data) and type of data policy (mandatory /voluntary dap) https://doi.org/10.29173/iq6 9/17 vlaeminck, sven and podkrajac, felix (2017) journals in economic sciences: paying lip service to reproducible research? iassist quarterly 41 (1-4), pp. 1-16. doi: https://doi.org/10.29173/iq6 table 2. availability of replication files and information on restricted data by type of article and data policy (n=452) type of article data available code available information on restricted data articles using restricted data in journals with mandatory data policy (n=110 / 24.3%) n.a. 46 (41.8%) 59 (53.6%) articles using restricted data in journals with voluntary data policy (n=58 / 12.8%) n.a. 5 (8.6%) 16 (27.6%) articles using non-restricted data in journals with mandatory data policy (n=155 / 34.3%) 98 (63.2%) 96 (61.9%) n.a. articles using non-restricted data in journals with voluntary data policy (n=129 / 28.5%) 6 (4.7%) 7 (5.4%) n.a. these results indicate why it is crucial to obligate authors to comply with a dap. also our findings illustrate that voluntary data policies do not sufficiently work. 3.2 analysis of journals’ compliance rates of the 245 articles in the sample for the compliances analysis, 99 (40.4%) were published in 2013, and 146 (59.6%) in 2014. the biggest share of articles again comes from the journal of the american statistical association (53; 21.6%), followed by the review of economics and statistics (37; 15.1%), the american economic review (20; 8.2%) and the american economic journal: applied economics (19; 7.8%). therefore, these four high-ranked journals represent more than half of the articles in the sample.13 167 out of 245 articles (68.2%) employed non-restricted data. for 77 (46.1%) of these, datasets for replication purposes are available, and the program code is provided for 75 (44.9%). of the 78 articles (31.8%) that used restricted data, 50 (64.1%) give some information on the dataset(s) employed. the program code is available for 44 (56.4%) papers (see table 3). table 3. availability of replication data, information on restricted datasets employed and compliance rate on article-level (n=245)** type of article data available code available information on restricted data compliant with data policy articles using non-restricted data (n=167/68.2%) 77 (46.1%) 75 (44.9%) n.a. 73 (43.7%) articles using restricted data (n=78/31.8%) n.a. 44 (56.4%) 50 (64.1%) 41 (56.2%)** ** five cases have been removed from the compliance analysis, because these journals exempt papers based on restricted data from their data policy. these numbers are comparatively higher than for the sample of all 452 data-based articles – for both research based on restricted data (for comparison: program code 30.4%; information on datasets 44.6%) and on non-restricted data (datasets 36.6%; program code 36.3%). possibly, this higher share reflects that only articles remain in the sample for which the journal’s data policy was in effect at time of manuscript’s submission. https://doi.org/10.29173/iq6 10/17 vlaeminck, sven and podkrajac, felix (2017) journals in economic sciences: paying lip service to reproducible research? iassist quarterly 41 (1-4), pp. 1-16. doi: https://doi.org/10.29173/iq6 when examining the overall share of all articles that are compliant with journals’ dap, we find a compliance rate of 47.5% on article-level for both papers based on both restricted and non-restricted data (see figure 4).14 this percentage indicates that more than half of all articles did not honour the regulations of the respective data policy. interestingly, when subdividing the sample into papers utilising restricted data and those using non-restricted data, we find that articles employing restricted data are more often compliant to journal’s data policy than those using non-restricted data. we also find that journals can be divided into two groups (see figure 5): one group enforces its data policy (albeit to varying degrees) while another group apparently does not care much about its data policy. journals with the highest compliance rates are the american economic journal: applied economics (100% compliance rate), the american economic review (100% compliance rate), review of economics and statistics (91.9%) and the review of economic studies (90% compliance rate). on the opposite side 10 journals reach only low compliance rates up to 20 percent. in eight journals not a single data-based article of the investigated issues is compliant with the journal’s data policy. on journal-level, in total 10 out of 17 periodicals (58.8%) investigated fall into this group (see figure 5). this distribution is also apparent by the median and the mean of the journals’ compliance rates: whereas the mean compliance rate on journal-level is 36.9%, the median of the sample’s distribution reaches only 3.8%. however, the reservation must be made that for some journals only a few data-based papers have been analysed, so that these findings should be interpreted with caution. compliance rates and (non-) mandatory data policies 158 out of 245 articles (64.5%) were published in journals with a mandatory dap, while 87 (35.5%) were published in journals with a voluntary dap. articles published in periodicals with a mandatory dap reach a compliance rate of 70.9%. in contrast, the overall compliance rate of papers published by journals with a voluntary dap is only 2.4%.15 this discrepancy underlines the importance to obligate authors to honour the data policy of a journal. table 4 details these findings. figure 5: how much journals in economic sciences comply with their dap (means) figure 4: share of articles that comply with the dap of the respective journal. https://doi.org/10.29173/iq6 11/17 vlaeminck, sven and podkrajac, felix (2017) journals in economic sciences: paying lip service to reproducible research? iassist quarterly 41 (1-4), pp. 1-16. doi: https://doi.org/10.29173/iq6 table 4. availability of replication data, information on restricted datasets employed and compliance rate for voluntary / mandatory data policies on article-level (n=245) type of article data available code available information on restricted data compliant with data policy articles w/ mandatory dap using restricted data (n=58/24.2%) n.a. 42 (72.4%) 44 (75.9%) 41 (70.7%) articles w/ voluntary dap using restricted data (n=20/6.3%) n.a. 2 (10%) 6 (30%) 0 (0%)** articles w/ mandatory dap using nonrestricted data (n=100/41.7%) 73 (73%) 72 (72%) n.a. 71 (71%) articles w/ voluntary dap using nonrestricted data (n=67/27.9%) 4 (6.0%) 3 (4.5%) n.a. 2 (3.0%) ** five cases were removed from the compliance analysis because these journals exempt papers based on restricted data from their data policy. 4. summary and discussion when analysing the share of articles published in 37 journals with a dap, we notice that almost threefourths (452 out of 599) of all articles investigated are data-based and therefore fall under journals’ data policies. more than a third (37.2%) of these 452 articles employs restricted data. these numbers underline the importance to implement suitable data policies for this type of papers (e.g. by obligating authors to provide useful references of the dataset(s) employed and to post the program code of their statistical analyses or simulations). for articles utilising non-restricted data, we find replication datasets for 36.6% of the articles and program code for 36.3%. when more deeply examining 245 of the 452 data-based articles for compliance with the particular data policy of the journal, we find an overall compliance rate on article-level of 47.5%. thus, the share of articles that are compliant with journals’ respective data policy was less than half. on article-level, we also ascertain that the compliance rate for papers utilising restricted data is almost 9 percent higher than for articles employing non-restricted datasets (52.6% to 43.7%). at first glance, this seems to be surprising. we suggest two possible explanations for this finding: the data policies’ requirements for research based on restricted data are easier to fulfil than for articles using nonrestricted data. most often, only the program code and some information on the dataset(s) employed have to be submitted. also authors might feel more confident to provide these files and some references on employed datasets, because they know their published findings will be replicated in exceptional circumstances only due to the unavailability of the data. on journal-level, the average compliance rate is 38% and therefore significantly lower than the rate reported by mccullough, mcgeary and harrison for their study of four top journals in economics in 2008. our sample diverges into two groups: a minority of journals strictly enforces their data policies and achieves high compliance rates, while a larger group of journals does little in terms of ensuring data availability for published articles and fostering reproducible research. this majority of 10 out of 17 journals (58.8%) has a compliance rate of less than 20%. for eight journals, not a single article is compliant to the respective journal’s data policy. however, for the interpretation of these findings we have to bear in mind that for some journals only a few data-based articles have been investigated. https://doi.org/10.29173/iq6 12/17 vlaeminck, sven and podkrajac, felix (2017) journals in economic sciences: paying lip service to reproducible research? iassist quarterly 41 (1-4), pp. 1-16. doi: https://doi.org/10.29173/iq6 hence, when including more issues and articles of these journals in the analyses, we presume that the share of periodicals which do more or less pay lip-service to reproducible research diminishes. in sharp contrast, four periodicals in our sample reach compliance rates of 100% or slightly less. these journals publish many articles per issue16 and this is the only reason why our sample eventually reaches an average compliance rate of 47.5% on article-level. all of these journals have dap that are in effect since at least spring 2012. also, these journals are among the higher-ranked journals in economic research. our findings also provide evidence on the importance of mandatory data policies for journals when it comes to reproducible research. for 73 out of 100 articles (73%) that employed non-restricted data and have been published in journals with a mandatory dap, accompanying datasets are available. for 72 (72%) papers we find program code. the overall compliance rate for these papers reaches 71%. in contrast, for papers published in journals with a voluntary dap, we find only four out of 67 papers with accompanying datasets (6%). for three (4.5%) papers the programme code is available and the overall compliance rate reaches only 2%. these findings reconfirm the outcome of theoretical studies that emphasise the missing incentives for researchers to support the replication of their own work. also they illustrate why researchers like chang and li (2015) reported that articles published in journals with a mandatory dap can be much easier replicated. therefore, voluntary data policies do not seem to be a feasible instrument to foster replicable research. nevertheless, also for journals with a mandatory dap there is room for improvement: according to our findings, not all of these journals do continuously demand the replication files from their authors. for journals with a ’mandatory’ data policy we would assume a compliance rate close to 100% not 71%, as our findings suggest. considering that our investigation focusses only on the basic prerequisites for replications (because we did not try to reproduce the findings of the papers investigated), our predictions for the outcome of such replication attempts are not optimistic. if the fundamental prerequisites for replications are missing, how could a noteworthy share of these papers meet a basic criterion of the scientific method? furthermore, we only checked whether authors followed the policies’ requirements regardless of whether we consider the policy itself to be robust or weak. because not all of these policies can be regarded as ‘robust’ (cf. vlaeminck & herrmann, 2015a), the sheer compliance rate does not necessarily equate with the portion of articles that can be replicated. we also would like to discuss some limitations of our analyses. a first one covers the data base of our sample for the compliance analysis. to investigate one or two issues of 17 journals is not sufficient to determine robust results for journals’ general compliance with their data policies. nevertheless our analysis provides a useful snapshot of journals’ commitment towards reproducible research for the years 2013/2014. a second limitation refers to the fact that several journals in our sample publish much more articles per issue than others what results in disparate numbers of articles per journal in our sample. this also has consequences for the interpretation of journals’ compliance rates. the overall compliance rate of 47.5% on article-level is only reached, because some journals with a high compliance rate publish many articles per issue. future research on these topics should incorporate more issues per periodical to gain results which are based on a broader sample of articles. at the time of such a follow-up study, journals’ data policies will be much longer in effect and journals will have https://doi.org/10.29173/iq6 13/17 vlaeminck, sven and podkrajac, felix (2017) journals in economic sciences: paying lip service to reproducible research? iassist quarterly 41 (1-4), pp. 1-16. doi: https://doi.org/10.29173/iq6 more experiences with enforcing their data policies. whether journals’ commitment towards reproducible research also grows over time due to these experiences remains an open question. to conclude, we would like to give some policy recommendations for journals and suggest some measures for editorial offices to foster reproducible research. based on the outcome of our analyses, we recommend that editorial offices of economic journals tighten their data policies towards mandatory daps that requires both data and program code. also, journals should broaden their data policies on research which is based on restricted data. as our findings illustrate, more than a third of all papers investigated employ such data. to exempt these articles from journal’s data policy would result in excluding research based on proprietary or confidential data from basic scientific requirements. beyond, journals should be stricter in enforcing their data policies, because replicability of published research is a cornerstone of the scientific method. in the first place editors are accountable for enforcing journals’ data policies, but also the reviewers should feel a responsibility to take care of a periodical’s data policy. both, editors and reviewers, invest time in ensuring that authors comply with journal’s style sheet. to also invest efforts in ensuring that replication files are available according to journal’s data policy is a task that would strengthen the scientific reputation of the periodical. this is even more important when we take into account, that first and foremost scientific journals have a central place as a quality instance in science and research. thereby they are also crucial for promoting a culture of research integrity because published articles are the most visible output of research. https://doi.org/10.29173/iq6 14/17 vlaeminck, sven and podkrajac, felix (2017) journals in economic sciences: paying lip service to reproducible research? iassist quarterly 41 (1-4), pp. 1-16. doi: https://doi.org/10.29173/iq6 references anderson, r.g., greene, w.h., mccullough, b.d. & vinod, h.d. (2008) the role of data/code archives in the future of economic research. journal of economic methodology. 15. p. 99-119. björk, b. & solomon, d. (2013) the publishing delay in scholarly peer-reviewed journals. journal of informetrics. 7 (4). p. 914-923. bräuninger, m., haucap, j. & muck, j. (2011) was lesen und schätzen ökonomen im jahr 2011?, dice ordnungspolitische perspektiven [online] 18. available from http://hdl.handle.net/10419/49023 [accessed 2017-01-16] chang, a.c. & li, p. (2015) is economics research replicable? sixty published papers from thirteen journals say „usually not“’. finance and economics discussion series [online] 2015 (83). p. 1-26. available from http://doi.org/10.17016/feds.2015.083. [accessed 2017-01-16] clemens, m.a. (2015) the meaning of failed replications: a review and proposal. iza discussion paper [online]no. 9000. available from http://hdl.handle.net/10419/110735. [accessed 2017-01-16] dewald, w.g., thursby j.g. & anderson, r.g. (1986) replication in empirical economics: the journal of money, credit and banking project. american economic review. 76 (4). p. 587-603. duvendack, m., palmer-jones, r.w. & reed, w.r. (2015) replications in economics: a progress report. econ journal watch [online] 12 (2). p.164-191. available from https://econjwatch.org/file_download/866/duvendacketalmay2015.pdf. [accessed 2017-01-16] ellison, g. (2002) the slowdown of the economics publishing process. journal of political economy [online] 110 (5). p.947-993. feigenbaum, s. & levy, d.m. (1993) the market for (ir)reproducible econometrics. social epistemology. 7 (3). p.215-232. glandon, p.j. (2011) appendix to the report of the editor: report on the american economic review data availability compliance project. american economic review. 101(3). p.695-699. available from http://pubs.aeaweb.org/doi/pdfplus/10.1257/aer.101.3.684. [accessed 2017-01-16] hamermesh, d.s. (1997) some thoughts on replications and reviews. labour economics. 4(2). p. 107109. available from https://doi.org/10.1016/s0927-5371(97)00015-8. [accessed 2017-01-16] hamermesh, d.s. (2007) viewpoint: replication in economics. canadian journal of economics/revue canadienne d'économique [online] 40(3). p. 715-733. available from http://doi.org/10.1111/j.13652966.2007.00428.x. [accessed 2017-01-16] hamermesh, d.s. (2013) six decades of top economics publishing: who and how? journal of economic literature. 51 (1). p. 162-172. höffler, j.h. (2017) replicationwiki: improving transparency in social sciences research. d-lib magazine [online] 23 (3/4). available from https://doi.org/10.1045/march2017-hoeffler. [accessed 2022-09-09] king, g. (1995a) replication, replication. ps: political science and politics. 28 (3). p. 444-452. king, g. (1995b) a revised proposal, proposal. ps: political science and politics. 28 (3). p. 494-499. longo, d.l., drazen, j.m. (2016) data sharing. the new england journal of medicine [online] 374. p. 276-277. available from http://doi.org/10.1056/nejme1516564. [accessed 2017-01-16] mccullough, b.d. (2007) got replicability? the journal of money, credit and banking archive. econ journal watch. 4 (3). p. 326-337. mccullough, b.d. (2009) open access economics journals and the market for reproducible economic research. economic analysis and policy. 39 (1). p. 117-126. mccullough, b.d., mcgeary, k.a. & harrison, t.d. (2006) lessons from the jmcb archive. journal of money, credit, and banking. 38 (4), 1093-1107. mccullough, b.d., mcgeary, k.a. & harrison, t.d. (2008) do economics journal archives promote replicable research? canadian journal of economics/revue canadienne d'économique. 41(4). p. 14061420. mirowski, p. & sklivas, s. (1991) why econometricians don’t replicate (although they do reproduce). review of political economy [online] 3(2). p. 146–163. available from http://doi.org/10.1080/09538259100000040. [accessed 2017-01-16] https://doi.org/10.29173/iq6 http://hdl.handle.net/10419/49023 http://doi.org/10.17016/feds.2015.083 http://hdl.handle.net/10419/110735 https://econjwatch.org/file_download/866/duvendacketalmay2015.pdf http://pubs.aeaweb.org/doi/pdfplus/10.1257/aer.101.3.684 https://doi.org/10.1016/s0927-5371(97)00015-8 http://doi.org/10.1111/j.1365-2966.2007.00428.x http://doi.org/10.1111/j.1365-2966.2007.00428.x https://doi.org/10.1045/march2017-hoeffler http://doi.org/10.1056/nejme1516564 http://doi.org/10.1080/09538259100000040 15/17 vlaeminck, sven and podkrajac, felix (2017) journals in economic sciences: paying lip service to reproducible research? iassist quarterly 41 (1-4), pp. 1-16. doi: https://doi.org/10.29173/iq6 mueller-langer, f. & andreoli-versbach, p. (2014) open access to research data: strategic delay and the ambiguous welfare effects of mandatory data disclosure. max planck institute for innovation and competition research paper [online] no. 14-09. available from http://hdl.handle.net/10419/98845. [accessed 2017-01-16] savage, c.j. & vickers a.j. (2009) empirical study of data sharing by authors publishing in plos journals. plos one [online]. 4 (9). p. e7078. available from http://doi.org/10.1371/journal.pone.0007078. [accessed 2017-01-16] tenopir, c., allard, s. douglass, k., aydinoglu, a.u., wu, l., read, e., manoff, m. & frame, m. (2011) data sharing by scientists: practices and perceptions. plos one [online] 6 (6). p. e21101. available from http://doi.org/10.1371/journal.pone.0021101. [accessed 2017-01-16] vlaeminck, s. & herrmann, l.k. (2015a) data policies and data archives: a new paradigm for academic publishing in economic sciences? in: schmidt, b. & dobreva, m. (eds). new avenues for electronic publishing in the age of infinite collections and citizen science: scale, openness and trust. proceedings of the 19th international conference on electronic publishing. amsterdam: ios press. p. 145-155. available from http://hdl.handle.net/10419/121278. [accessed 2017-01-16] vlaeminck, s. & herrmann, l.k. (2015b) data policies and data archives: a new paradigm for academic publishing in economic sciences? (replication data). version: 1. dataset. available from http://journaldata.zbw.eu/dataset/data-policies-and-data-archives-a-new-paradigm-for-academicpublishing-in-economics. [accessed 2017-01-16] acknowledgements: the authors would like to thank the referee for valuable comments which helped to improve the manuscript. also the authors would like to thank dr. martina grunow and ralf toepfer for their helpful comments, derek kruse for his support in checking spelling and language of the article, and frank müller-langer for the kind provision of three implementation dates of journals’ data policies. these project results have been developed in the edawax project (european data watch extended, http://www.edawax.de). edawax was financed by the german research foundation (http://www.dfg.de). appendix: replication files and documentation the replication files for this paper are available here: http://journaldata.zbw.eu/dataset/journals-ineconomic-sciences-paying-lip-service-to-reproducible-research-replication-data. https://doi.org/10.29173/iq6 http://hdl.handle.net/10419/98845 http://doi.org/10.1371/journal.pone.0007078 http://doi.org/10.1371/journal.pone.0021101 http://hdl.handle.net/10419/121278 http://journaldata.zbw.eu/dataset/data-policies-and-data-archives-a-new-paradigm-for-academic-publishing-in-economics http://journaldata.zbw.eu/dataset/data-policies-and-data-archives-a-new-paradigm-for-academic-publishing-in-economics http://www.edawax.de/ http://www.dfg.de/ 16/17 vlaeminck, sven and podkrajac, felix (2017) journals in economic sciences: paying lip service to reproducible research? iassist quarterly 41 (1-4), pp. 1-16. doi: https://doi.org/10.29173/iq6 end-notes 1 sven vlaeminck was the project manager of the project ’european data watch extended’ (edawax) and works for zbw – german national library for economics / leibniz information centre for economics in hamburg/germany. since 2011, he is active in the field of research data management. he can be reached by email: s.vlaeminck@zbw.eu. 2 felix podkrajac is currently trainee in academic subject librarianship at the library and information system of the carl von ossietzky university of oldenburg. email: felix.podkrajac@uni-oldenburg.de. 3 we define an article to be ’data-based’ when the methods include creating, reusing, and analysing research data, other empirical methods (e.g. experiments) or self-compiled program code (e.g. simulations). 4 the growing importance of research data, its management and potentially also its integration in the academic publishing process is reflected in numerous statements and requirements of universities, learned societies, funding agencies and even political bodies. the nsf (us), the esrc (uk) and the rcuk (uk) all emphasise that data collected with public funds belongs in the public domain and highlight the advantages of data sharing: ’data sharing strengthens our collective capacity to meet scientific standards of openness by providing opportunities for further analysis, replication, verification and refinement of research findings’ (national science foundation, 2010). beside policies on data sharing for publicly funded research, others also explicitly mention publication-related research data. for instance, the european commission (ec) recommends that eu member states ought to implement policies to ensure that ’datasets are made easily identifiable and can be linked to other datasets and publications through appropriate mechanisms’ (european commission, 2012). 5 according to king (1995b) the primary purpose of a data archive is not necessarily to ensure replicability (this would be a high demand) but to enhance extensibility of published research (this presumes replicability). 6 the data availability policy of the american economic review (aer), for instance, lists varying requirements for econometric and simulation papers and for experimental papers. for econometric and simulation papers, the ‘minimum requirement should include the data set(s) and programs used to run the final models, plus a description of how previous intermediate data sets and programs were employed to create the final data set(s).’ for experimental research, ‘we normally expect authors of experimental articles to supply the following supplementary materials […] 1. the original instructions. […] 2. information about subject eligibility or selection. […] 3. any computer programs, configuration files, or scripts used to run the experiment and/or to analyze the data. […] 4. the raw data from the experiment.’ 7 according to the uc berkley, ‘restricted information describes any confidential or personal information that is protected by law or policy’. also data provided by third parties, that ’may not be redistributed or reused without the consent of the original provider’ can be characterised as restricted data, as the world bank points out. 8 the median of articles published by journals in our full sample is 13 (equates 2.2%) (means: 16.2/2.7%). 9 nevertheless, nine articles remain for which we cannot finally determine whether they are data-based. these articles are removed from the sample and are no longer regarded in our analysis. 10 for instance, this applies to the data availability policy of the journal of money, credit and banking (jmcb). 11 for an overview on how journals provide replication data to would-be replicators, please consult vlaeminck & herrmann (2015a). 12 exemplary, we provide two good examples for a useful documentation and well-structured replication files for research based on restricted data: (i) http://restud.oxfordjournals.org/content/suppl/2013/09/28/rdt030.dc1/supplementary.zip (supplement to doi: 10.1093/restud/rdt030) (ii) http://repec.wirtschaft.uni-giessen.de/~repec/repec/jns/datenarchiv/v234y2014i1/y234y2014i1p5_22/ (supplement to doi: 10.1515/jbnst-2014-0103) 13 the median of articles by journal in the ‘compliance sample’ was 9 or 3.7% (means: 14.4/5.9%). 14 five cases are removed from the compliance analysis’s sample because the journals exempt papers based on restricted data from their data policy. therefore, only 73 articles that rely on restricted data have been examined. 15 five cases are removed from the compliance analysis’s sample because the journals exempt papers based on restricted data from their data policy. 16 130 out of 245 (53.1%) articles in the compliance analysis have been published by just four journals: the journal of the american statistical association (53), review of economics and statistics (37), review of economic studies (20) and american economic review (20). https://doi.org/10.29173/iq6 mailto:s.vlaeminck@zbw.eu mailto:felix.podkrajac@uni-oldenburg.de http://www.nsf.gov/sbe/ses/common/archive.jsp http://www.esrc.ac.uk/funding/guidance-for-grant-holders/research-data-policy/ http://www.rcuk.ac.uk/research/datapolicy/ https://www.aeaweb.org/journals/policies/data-availability-policy https://security.berkeley.edu/what-restricted-data http://data.worldbank.org/restricted-data http://data.worldbank.org/restricted-data http://restud.oxfordjournals.org/content/suppl/2013/09/28/rdt030.dc1/supplementary.zip http://repec.wirtschaft.uni-giessen.de/~repec/repec/jns/datenarchiv/v234y2014i1/y234y2014i1p5_22/ lassist newsletter, vol. 3, no. 2 (spring 1979) user needs and confidentiality in sweden erika von brunken the issue of data confidentiality has become an important problem for discussion and legislation in many nations. legislation dealing with data confidentially can have an impact on social science research as well as on methods for archiving and retrieving data. erika von brunken, in a paper delivered originally in ottawa, discusses sweden's response to the issue . --ed itor . introduction since may 1973 sweden has a data protection act (datalagen, 1973) [1]. as questions related to registration of personal data have attracted considerable attention since the passing of this law, it became obvious that amendments would be needed in a near future. several parliamentary bills on this subject have been proposed during the past years. in may 1976 a swedish government commission on data legislation (dalk) was set up for a general review of the data protection law and of the activities and experience of the data inspection board. dalk examined the legal regulation of the protection of privacy in conjunction with registration of personal data, chiefly problems associated with the use of adp. as a result of this investigation the report "personregister datorer integritet" (person registers computers integrity) [2] was submitted (june 1978). emanating from this report, a government proposal[3] on certain amendments in the data act was submitted on march 22, 1979, which will come into force on july 1, 1979. dalk is now investigating the way in which computerization affects the principle of public access to official records, as well as the use of computers in public administration and by the business community as an international phenomenon. in the following i shall not dwell on the concepts of privacy, integrity, confidentiality, ranging from "the right to be left alone" to "being able to decide and act on one's own", but restrict myself to the legal aspects of privacy protection and their impact • on research in the social sciences. th e identity number and the principle of public access to official records . discussions on privacy and data protection in sweden revolve around two basic problems: the identity number, assigned to every person living in sweden, and the principle of public access to official records, confirmed by law in 1766. the identity number, in sweden called "person-nummer" , has existed since january 1, 1947, and comprises the birth date (year, month, day) and four digits, e.g. 650213-1193. the last four digits are coded information on country of birth, sex, and a control figure. as all information on an individual is stored by this identity number, it was easy to sort immigrants by their national origin. this discrimination has been cancelled 23 lassist newsletter, vol. 3, no. 2 (spring 1979) lately, so that now even immigrants get their identity number from the swedish series. the identity number, the name, address and family relations are entered in a personal file drawn up by the registration office of the parish in which the individual is registered. in the personal file instances of marriage, children, divorce, change of address and death are recorded. if the person moves to another parish, this file is transferred to its registration office. while in most countries the population statistics are still based on censuses, sweden now has a fully developed system for the continuous recording of population changes in local, regional and national registers. a vast amount of personal data has so been stored in sweden in machinereadable files, and the identity number makes it technically easy to merge information from different files, originally stored for other purposes. the principle of public access to official records was established by the press law ( tryckf r ihetsl agen) in 1766. according to this law any swedish citizen has the right to take part of, to read or to copy official records and to publish their content. [2] even the records of the nunicipal administration are official records due to this law. certain records, however, are not official, as documents concerning state security, central financial policies, interests to prevent crime or legal action against it, the economic interests of the society, etc. these documents are classified as secret material by the secrecy law and are not accessible. information stored on machine-readable media can be obtained in the form of printouts for a fee. personal data are not accessible. the press law has been amended several times. the latest amendment has been made in 1976 with special application to automatic data processing and other technical recording. the new rules came into force in january 1978. the data inspection board and its role in the protection ot pr i vacy . the data inspection board is a central administrative agency for examination of matters relating to licenses and supervision in accordance with the data act, the credit information act and the collection of debts act. [2] because of the large amount of personal data stored on machine-readable media since the beginning of the 1960s, the use of adp was considered to involve such risks of intrusion upon the privacy of a registered person, that special attention had to be paid. special legislation was therefore demanded for regulation of both public and private personal registers kept by adp. since july 1, 1973, every person, firm or authority, who wants to register personal data by adp, has to apply for a license at the data inspection board. now even for collection of personal data for automatic data processing at a later date a licence is needed. the swedish government commission on data legislation (dalk) recommends that the following considerations should be taken into account when licence applications are examined : it should still be permissible to start a personal register if it can be assumed that having regard to the various regulations which may be issued the register involves no risk of undue 24 lassist newsletter, vol. 3, no. 2 (spring 1979) encroachment upon the privacy of the registered persons the significance of the term undue encroachment upon the personal privacy of those registered cannot be decided in general, but must as now, continue to be judged from case to case in making this judgment special consideration should be given to whether the purpose of the register complies with the activity conducted or to be conducted by the responsible keeper of the register special atte paid also and quanti sonal dat wh i c h pe r included i and to whe and quanit the categ concerned purpose of ntion should be to the nature ty of the pera, as also to sons shall be n the register, ther the nature ty of data and ory of persons comply with the the register special attention, too, should be paid to whether the data to be included in the register were originally collected for another purpose than the register is to serve special attention, finally, should be paid to the attitude to the register held by or assumed to be held by those who may be included in it." dalk proposed additions to sector 3 of the data act on these lines. the proposed additions have been included in the amendment to the data act. dalk emphasizes the significance of public interest when a licence is examined, saying "that certain very delicate personal information may, under the data act, be registered if called for by a strong societal or other public interest". when a license has been given to set up and to handle a personal file by adp technique, the data inspection board gives instructions in accordance with section 6 of the data act on following points: 1. collection of information for the person register 2. how to perform the automatic data processing 3. the hardware 4. processing of personal data allowed by adp 5. notification of the persons concerned 6. the kind of personal data which may be made available 7. the handing out and other use of personal data 8. preservation and sorting out of personal data 9. control and security. concerning the registration of soft data section 6 was amended with the following: when considering if instructions are needed, special attention shall be paid in case the file contains personal data based on judgments or on appraised information on the registered person. 25 lassist newsletter, vol. 3, no. 2 (spring 1979) person registers o government or the par need a license, but by the data inspectio data inspection boar for the license proc ponding to the time n moment the fee is hour. research wor pay reduced fees. data inspection bo about 20,000 applicat died 18,000 of them, a simplified proc remaining 35 per cent lot of work. researc to this group. rdered by the 1 iament do not are supervised n board. the d takes a fee edures correseeded. at the skr . 315: per kers sometimes up to now the ard received ions and han65 per cent by edure. the often need a h files belong obi igations of the holder of a person register . every person, firm or authority who received a license for setting up or holding a person register on adp is must follow paragraphs 8 14 of the data act. i present these paragraphs in an abbreviated form: 8. if it can be suspected that some personal data in the file are wrong, the holder is responsible for immediate investigation and correction of the data in question. if a file c sonal data incomplete wi the aim of t and which by pleteness m undue intrus personal int individual legal implic holder has t the missing i ontai wh i th re he its ight ion eg r i t r mig ation o su nform ns perch are spect to eg ister , incomcause nto the y of an ht have s, the ppl emerit ation. requests it, the holder of the file has to inform the appl icant about the content of the personal data stored on him/her. even if no information has been stored, this has to be stated. once information has been given, no new information has to be confered to the same person before 12 months later. this kind of information is free of charge. certain information is yet excepted from this rule. 11. personal data may not be handed out if it can be suspected that the information will be handled by adp in conflict with the law. in case information shall be handled by adp in a foreign country, the consent of the data inspection board is needed. such a consent will only be given if no intrusion into personal integrity is involved. 12. the responsible holder of the file has the obligation to notify the data inspection board if the register shall be closed. the board will then give instructions what to do with the file. 13. the responsible holder or other persons working with a person register or collecting material for the file are not allowed to reveal information on an individual. the same is valid for persons who received information from a person register. 10. if a registered person 26 lassist newsletter, vol. 3, no. 2 (spring 1979) 14. if authorities ase adp would lead to infringement or prirecords for handling or vacy instead. it has been recomhearing a case, the matermended that the identity number ial shall be added to the should not be printed in data outrecords in readable form, puts if not necessary, if not special reasons give rise to another dalk says: "as regards other procedure. aspects of linking, it has above all been pointed out in the debate that public agencies, with the aim those who break the law can be admirable as such to fulfil their punished by fines or prison. an functions as justly and rationally individual has the right to claim as possible, have tended increasdamages if intrusion into personal ingly to make use of adp and the integrity has occurred. the data possibilities of linking registers inspection board may recall a in order to check the correctness license in case personal integrity of particulars submitted by the has been violated or cannot be individual in different contexts, secured. but it is not unusual that the private sector as well, e.g. insurance companies and credit information agencies, uses data in various merging information from official registers to check infordifferent files by use of mation submitted by the individual the identity number . relevant to a particular private sector. apart from these the linking of files or merging instances, the linking and other of information from di'ferent joint use of data would appear to files, originally set up fo. other be commonly desired in scientific purposes, has aroused public opinresearch, including community planion and drawn attention to the need ning and the production of statisfor protection of privacy. it is tics. these aspects of the linkage technically easy in sweden to merge problem involve, in dalk's opinion, selected information from different a broader political issue, namely file by using the identity number. which methods should be accepted all information on an individual is that public and private bodies use stored by this number. the tax for checking the correctness of office checks your declared income, particulars submitted by the indithe social welfare authorities make vidual, often on oath or in similar sure you have not received undue forms, in a specific administrative allowances, etc. the cancer-envimatter or a specific customer relaronment-reg ister , conducted by the tionship." national board of health and welfare, is the result of merging the question is to weigh the information from the cancer regisinfringement on privacy against the ter with census data on occupation, demands of the public interest, working place, living quarters, dalk proposed that governmental education, etc. a more restricted instructions should define mo-re use of the identity number, even clearly for which purposes data in its elimination, have been disofficial registers may be used, cussed, but the advantages are surthis might quiet the public uneasipassing the disadvantages. the ness concerning uncontrolled use of joint running of files would not be individual data. as easy as now, but name confusions 27 lassist newsletter, vol. 3, no. 2 (spring 1979) soft data and other sensitive be eliminated, research would be information . made impossible, which in turn would endanger society. the board the need for and the use of soft was afraid that dalk's way to data and other sensitive informabalance the demand for integrity tion and their handling has been against the demand for research discussed thoroughly by the might become fatal for future research community and the authoriresearch in the social sciences in ties. the point of view differs, sweden. janson stressed further depending on who is discussing it. that the data protection law largely the authorities agree that already has affected research in a soft data should be handled with negative way, and that the protecutmost care and should be judged tion of privacy has changed for the from case to case. the data worse. he foresees a strong inspection board emphasizes that it bureaucratic impact on research, cannot be said generally which kind and as a consequence reduced empirof information is sensitive, and ical research. janson proposed which is not. important is the that a distinction should be made feeling of the individual towards between administrative and pure it. it is that perception which research files, the latter ones should decide from case to case. should not need a licence. this the data inspection board attaches idea had already earlier been pregreat importance to the viewpoints sented by him during a symposium on of the ethical committees at the "forskning och integritet" respective faculties. these com(research and interity), arranged mittees investigate research proby the faculty of jurisprudence at jects of sensitive nature, and exastockholm university in march mine if they can be performed in 1978[5], where research workers accordance with ethical rules. thi from different fields in the social restrictive view of the data sciences had met and discussed inspection board is shared by dalk, their experiences. several partiwhich proposed the amendment to cipants of this symposium emphasection 6 of the data act already sized the need for soft data and mentioned earlier. the necessity to store them for later use in longitudinal and panel as expected, the strongest cristudies. the elimination of the ticism against the proposed amendidentity number would make such ments have been expressed by studies impossible. concerning research workers in the social sciarchiving and sorting out data ences. professor carl-gunnar janfiles it was pointed out that it son, sociologist and dean of the should be born in mind what kind of faculty for social sciences at data might be of value for research stockholm university, expressed the 20 years or more ahead, and that view of the board of the faculty in the demands of future research an answer to the department of jusshould be met. tice, which submitted dalk's report for consideration. [4 ] he said that the research council for the neither freedom nor right are absosocial sciences, well aware of the lute, neither the freedom to do need to preserve research files, research, nor the individual's set up a working group on data right for privacy. the interests archiving matters in march 1978. of the society had to be taken into the final report of this group has consideration. if all risks should just been presented, but no deci28 lassist newsletter, vol. 3, no. 2 (spring 1979) sion has yet been made. ulf advertisements filling their mailchristof fersson from the university boxes. of gothenburg, a member of lassist, belongs to this group. probably he still we are an open society, will report on this work at a later research workers, representatives date. the board of the research of the data inspection board, the councils, forskn ingsradsnamnden. research councils and the central set up a committee on longterm bureau of statistics have been most research for investigation of the helpful by providing me with inforneed for future access to data. mation. summarizing my impressions the work of this committee has been very crudely i could say: the presented in a report "forkningens farther away you are from research, framtida datatillgang" (future the less you are worried about the access to data for research) [6] by impact of data legislation on christer winberg and sune akerman. future research. references final remarks . 1. datalagen sfs 1973:389. sweden is known as an open 2. pe rsonreg i ster-da torer-in teg r i tet society where information is sou 1978:54, liber forlag allowed to flow freely. it might seem astonishing to people from 3. regeringens proposition other countries, that the swedes 1978/79:109 om andring i dataare willing to accept the accumulalagen (1973:289) riksdagen tion of a vast amount of personal 1978/79. 1 saml. nr 109. data on them in official files, which later on might be used for 4. stock lolms universitet, samhallsupervision by governmental or svetenskapl iga fakul tetsnamndlocal authorities. i think as the ens remi ssyttr ande over dataprinciple of public access to offilagsti fningskommi ttens cial records gives the citizen a betankande. 1978-11-01. possibility for insight into public administration, a feeling of reci5. forskning och integritet. symprocity is created. there have posium 9-10 mars 1978 anordnat been opinion polls after the census av juridiska fakulteten vid of 1970 and 1975, as some questions stockholms universitet asked in the census were seen as an 1878-1978. intrusion into privacy, but i think people are not yet aware what pos6. winberg c. and akerman s. forsibilities for control are given by skningens framtida datatillstorage of those files by adp techgang. samarbetskommitten for nique. it just begins to dawn on 1 angsi ktsmot iverad forskning. them. what is embarrassing people juni 1976. most, are the personally addressed 29 distribution of census data on cd-rom to depository libraries by juri stratford ' documents librarian shields library university of california, davis introduction the depository library program (dlp) was established by congress to inform the public on the policies and programs of the federal government. through the dlp, the government printing office (gpo) distributes publications to designated libraries. while government documents received through depository distribution remain the property of the federal government, depository institutions are responsible for the maintenance of, and providing public access to, the documents. this usually involves committing public service and technical service staff, and the purchase of bibliographic tools in addition to committing space. in the past, depository libraries have provided public access to the printed census publications while access to census machine-readable datafiles (i.e. tapes) has been available through state data centers and data archives. the census bureau is now beginning to distribute machine-readable datafiles on cd-rom to depository libraries. this affords depository libraries both new opportunities and new responsibilities. the depository library program there are almost 1,400 depository libraries. they are located in each state and congressional district in order to make government publications widely available. these government publications are available for the free use of the general public. for the purpose of depository distribution, a government publication is defined as "informational matter which is published as an individual document at government expense, or as required by law." [44 usc 1901] the origins of the dlp can be traced back to the 1790s when the state department distributed acts of congress to state governments and newspapers. funding was sought on an ad hoc basis until 1813 when congress passed a resolution authorizing "every future congress" to print additional copies of congressional publications for this purpose. in 1814, the american antiquarian society was designated the first depository library. responsibility for the program shifted between various agencies and departments, mainly the department of state and the department of interior, throughout 19th century. congressional resolutions in 1857 and 1858 affirmed the distribution of congressional materials to institutions such as libraries and colleges, and other organizations designated by members of congress. in 1895, a new printing act was passed, incorporating the old legislation and placing responsibility for the dlp in the office of the superintendent of documents at gpo; the act also specified that certain executive materials were to be included. 2 the present law, the federal depository act of 1962, increased the number of possible depository libraries; established a system of regional depository libraries which were to maintain a permanent collection, and provide interlibrary loan and reference service; expanded the variety of government documents available for distribution; and established a reporting mechanism, the biennial survey, to ascertain the libraries' condition. the 1962 act has been amended twice: in 1972 to exempt the highest appellate court of each state from the requirement of public access; and in 1978 to extend depository eligibility to law schools. 3 depository distribution of machine-readable datafiles gpo has reversed its position on depository distribution of machine-readable datafiles since the early 1980s. at their fall 1981 meeting, the depository library council (dlc), an advisory body to the public printer and the superintendent of documents, passed a resolution regarding the feasibility of the gpo providing free access for depository libraries to unclassified bibliographic data bases belonging to federal agencies. in response to the resolution, gpo general counsel garrett brown determined that "...the depository library act of 1962 does not direct the superintendent of documents to make published documents available in all possible formats to the libraries. it was the intent of congress that only printed publications be made available to depositories." 4 following the census bureau's plans to distribute data from the 1982 census of agriculture and the 1982 census of retail trade through the depository library program on cd-rom as test disc 2, the public printer requested approval through the joint committee on printing (jcp). in a march 25, 1988 letter to the public printer, congressman frank annunzio, chairman of the jcp, affirmed the committee's support of the census project and the gpo's authority to produce and distribute fall/winter 1990 government publications in electronic formats. 5 in 1989, gpo asked its general counsel grant g. moy, jr. to review the 1982 opinion. he concluded that "the earlier question presented to the general counsel concerned only the issue of access to unpublished information in a computer data base," and, this was still the case; but "the specific statement in the general counsel's 1982 opinion, limiting depository distribution to printed products, was disapproved."6 test disc 2 was distributed to 173 depository libraries as a pilot project in september 1988. the cd was available through regular depository distribution in 1989. at the march 1989 depository library council meeting, jan erickson of the government printing office reported on the initial distribution of test disc 2, and indicated that it was the census bureau's intent to distribute future cdrom products through the depository system.7 a second census cd, the city and county data book, was distributed to depository libraries in spring 1990. software accompanying early shipments of the city and county data book were "infected" with the jerusalem virus. test disc 2 the census bureau's cd-rom test disc 2 contains data from the 1982 census of retail trade and the 1982 census of agriculture on a single compact disk. the files are in dbase m format. files from the census of agriculture contain 1982 data by county, with comparable data for selected items from the 1978 census. the technical documentation for the census of agriculture describes the data as a single file with a logical record size of 40,320 characters containing 3,360 data fields. however, the dbase iii record structure only allows 128 fields. this large record structure is accommodated by storing the data in 28 separate dbase iii compatible files. the first file, ag82_geo.dbf, provides geographic information for each state and county. this information is a guide to the arrangement of data contained in the 27 numeric datafiles. for example, the geographic area indicated by record #170 in ag82_geo.dbf is "california." this means that state level data for california in the other 27 files is contained in record #170. each file is named ag82_nn.dbf where nn = 1 to 27. except for the last file, which contains the last 32 fields, each file contains 128 fields. data from the 1982 census of retail trade are available for each 5-digit zip code, including number of establishments, by kind of business, and basic data for retail trade total. the files for the census of retail trade have a much simpler record structure than the files for the census of agriculture. the 1982 census of retail trade data are stored in 51 separate files. each file is named rc82_xx.dbf where xx = the state postal abbreviation. data files for the retail trade data have a record length of 155 characters containing 19 fields. there are two files for each state, a dbf file and an ndx file. 8 county and city data book the 1988 county and city data book cd-rom contains the same data as published in the printed volume. these files contain data gathered from a variety of federal agencies and national associations. the disc includes data for states, counties, cities with a population of 25,000 or more, and places with a population of 2,500 or more. like the census of agriculture files on test disc 2, each record represents a geographic area, and subject fields are distributed across a number of datafiles. the datafiles range in size from about 200kb to lmb. the economic censuses and the census of population and housing almost all data from the economic censuses previously available on magnetic tape will be on cd. the economic census and census of agriculture will be available on 9 cds. the first disc, volume 1 , release 1a was released in early 1990, but has not yet been distributed to depositories. the disc contains data from the geographic area series for wholesale, retail, and service industries for selected states and includes the same statistics statistics as published in the corresponding report series: geographic area series for 1987 censuses of retail trade (all states), wholesale trade (selected states), and service industries (selected states); preliminary industry series for the 1987 census of manufactures (national with some state totals) and selected historical statistics. plans for the 1990 census of population and housing call for 20-30 cds to be released from mid1991-1993, including redistricting data and block statistics. 9 data extraction each of these cd-rom products is distributed with a program to display tables, but software is not provided to copy data subsets. the datafiles range in size from 30kb to 3mb; most of these files are too large to manipulate on a microcomputer without first creating a smaller data subset that can be copied to a floppy disk or hard disk. the texts documenting test disc 2 and the city and county data book discuss using dbase iu to work with the files. the documentation for the economic censuses describes the extract program. the extract program is a public domain program that was developed to create subsets from the large cd-rom databases and save them as files on a floppy or hard disk. extracted files can be created in dbase format, ascii fixed field format, or ascii comma-delimited format. version 1 of extract is slow and occasionally crashes. a new version of extract should be released shortly. extract is iassist quarterly available from the center for electronic analysis, university of tennessee. however, it has not been distributed with the cd-roms. conclusion in their paper, "government information in machinereadable data files: implications for libraries and librarians," ray jones and thomas kinney examine the requirements for the utilization of machine-readable datafiles in retrieval and reference services. two of their remarks can be paraphrased to apply to the utilization of the census cd-roms in depository libraries. first, when librarians have the responsibility of retrieving numeric information from cd-roms either they must know how to program or work with a colleague who programs; and second, librarians will require the critical judgment to determine when data retrieval from cdroms is needed to answer the patron's need most completely. 10 these skills are not widely held by depository librarians at the present time. this is evidenced by the fact that few depository libraries have successfully integrated these materials into their reference service. 11 the census bureau is currently reviewing the impact of the cd-rom distribution. the report, the role of intermediaries in the interpretation and dissemination of census data now and in the future, by census statistician sandra rowland examines these issues. the study credits the experience of librarians assisting in the understanding and use of census data; however, it concludes that "neither the gpo nor the libraries play a big part in the interpretation of data for users" and argues that role of depository libraries is "unlikely to change in the future unless librarians take a more aggressive role as information technicians." it also states that, while the regional depository libraries will acquire and hold all census products including data on high density optical storage devices, most depository libraries will acquire and hold fewer census products in the future than they do now. 12 at present, the census bureau appears to have a strong committment to the depository distribution of their cdrom products, and these materials are available to all depository libraries. the depository library community needs to work closely with data archivists to insure that effective use is made of the census cd-rom products. data archivists might try to meet informally with depository librarians within their own institutions to discuss how access to census data on tape, cd-rom, and paper copy could best be coordinated. data archivists might also consider coordinating presentations with state or national government document groups. finally, while the census bureau assures us that the cdrom production of the 1987 economic census and the 1990 census of population and housing will not be produced at the expense of the publication of the tape or paper products, sandra rowland's report suggests that we can expect to see fewer paper products and more electronic products for the 2000 census: "with respect to the year 2000 census, it is very likely that there will be a movement out of printed media and into electronic media for dissemination to the libraries." 13 while there are certainly instances where researchers' needs would best be served by census data on cd-rom, documents librarians and data archivists alike need to monitor the situation to insure that the cd-roms are not produced at the expense of other necessary census products. 1 presented at the iassist 90 conference held in poughkeepsie, n.y. may 30 june 2, 1990. 2 hemon, peter, charles r. mcclure and gary p. purcell, gps's depository library program: a descriptive analysis (norwood, new jersey: ablex publishing company, 1985), pp.3-7. 3 u.s. congress, joint committee on printing, a directory of u.s. government depository libraries (washington, d.c.: u.s. government printing office, 1988), pp.l3. 4 u.s. congress, joint committee on printing, provision of federal government publications in electronic format to depository libraries: report of the ad hoc committee on depository library access to federal automated data bases... (washington, d.c.: u.s. government printing office, 1984), pp.1 12-13. 5 u.s. congress, office of technology assessment, informing the nation: federal information dissemination in an electronic age (washington, d.c.: u.s. government printing office, 1988), p.143. 6 ibid. 7 "summary: spring meeting, depository library council, pittsburg, pennsylvania, march 8-10, 1989," administrative notes 10 (august 1989): 3. 8 stratford, juri, review of "cd-rom test disc 2 (machine-readable data file) and cd-rom test disc 2: technical documentation," government publications review, 16 (1989): 397-398; see also peter hemon and candy schwartz, "readers exchange: census product review," administrative notes 11 (march 1990): 13-21. 9 administrative notes 10 (august 1989): 3. 10 jones, ray and thomas kinney, "government information in machine-readable data files: implications for libraries and librarians," government publications review, 15 (1988): 25-32. fall/winter 1990 " diane smith examines the viability of federal depositories to deal with electronic information in her forthcoming paper, "depository libraries in the 1990s: whither or wither depositories?," government publications review, 17 (1990). 12 rowland, sandra, the role of intermediaries in the interpretation and dissemination of census data: now and in the future, (washington, d.c.: u.s. bureau of the census, 1989), photocopy, pp.32-33. 13 op cited, p. 14. data news here are two items that readers of the quaterly may find of interest. (submited by: jim jacobs jajcobs@ucsd .edu) 1 1) in the newest edition of the asis "annual review of information science and technology" (vol. 25, 1990, martha e. williams, ed., published for asis by elsevier science publishers pp. 3-54), karen j. sy of the university of washington and alice robbin of the university of wisconsin have written an article titled: "federal statistcal policies and programs: how good are the numbers?" in their concluding remarks they say, "...the 1980s saw a serious decline in the quality of our federal statistical system." and, "throughout the government, but especially in the key coordinating branch of the office of management and budget, we found policy makers who viewed statistics as a burden and not as indicators of who we are as a nation and where we should be going. in sum, the federal statistical system is in serious trouble." 2) the committee on national statistics and the social science research council with support from several agencies, have convened a panel on confidentiality and data access. the scope of the study includes publicly supported statistical data collection activities. the panel is in the middle of a two year study and is soliciting short statements from interested parties on the following topics: access problems, (examples of instances where confidentiality lawor policies have made it impossible to obtain data.) suggestions for improving access, (including suggestions tor improving access with appropriate safeguards to maintain confidentiality.) persons or businesses harmed by disclosure. you can submit statements to: george t. duncan, chair panel on confidentiality and data access committee on national statistics national academy of sciences 2101 constitution ave, nw washington dc 20418 if you have questions, or if you want a more detailed announcement ot the charge to the panel, you may call virginia de wolf, study director, (202) 3342550. 1 jacobs, jim. 1991 "quality of government statistics" [computer file]. edmontonl, alberta: or-l . electronic listserv. (listserv@ualtavm ). iassist quarterly sist newsletter vol.1, no. 4 new organizations/ reorganization unisist at the 19th session of the unesco general conference, a resolution was adopted concerning the future development of the unisist program. the unisist newsletter (volume 4, no. 4, 1976) notes that the "conference agreed to set up a general information programme concerning the activities of the organization in the field of scientific and technical information and of documentation, libraries and archives." part 5 of the resolution particularly concerns the data archive community: it "authorizes the director-general to facilitate the implementation of the general information programme, by seeing that activities are integrated with a view to: ... (c) contributing to the development of information infrastructures and to the application of modern techniques of data collection, processing, transfer and reproduction, [and] (d) promoting the training and education of information specialists and information users, with particular attention to the needs of the developing countries, especially the problems of transfer of information and data from the technologically advanced countries to the developing nations." the unisist newsletter issues which were sent to the lassist newsletter editor contain a great deal of information on upcoming meetings throughout the world which will deal with an assortment of information problems (such as transfer, documentation and systems development, bibliographic data bases, management, training of staff, resource sharing) and descriptions of information systems (such as popins [international population information system], isonet [international standards information network]). iassist members in the u.s. may request back copies of the unisist/general information program newsletter and ask to be placed on the mailing list by writing to: cistip secretariat, national academy of sciences, 2101 constitution avenue, n.w., washington, d.c. 20418. outside the u.s., lassist members should contact their respective national science academy or the unesco division of scientific and technological information and documentation, 7, place de fontenoy, 75700 paris, france. cistip stands for the u.s. committee for unisist. [committees with the same purpose must have been established in other countries. if lassist members have information about this, please pass it on through the newsletter .] unesco has also created an ad hoc committee on social science information, which will hold its first meeting in paris in november 1977. it is intended that this ad hoc committee will provide intellectual and policy guidance for the inclusion of social science information in unisist. [from a letter by fred w. riggs to judith s. rowe, 4 october 1977, pp. 1,2.] inssist, the international social science information services and technologies, is an action group which intends to serve as a catalyst for the tasks of liaison. inssist will meet to exchange information during the forthcoming isa [?] conference in washington (sheraton park hotel, february 22-26, 1978), 41 sist newsletter vol.1, no. 4 the director of the unesco division responsible for promoting the international development of social science, vladimir mshvenieradze, is expected to attend and report on the work of his division, including the new ad hoc committee. the time and place of the meeting will be announced in the final isa program. [announcement of inssist meeting from fred riggs to lassist editor, received 2 november 1977.] international association for statistical computing e. lunenberg, director of the permanent office of the international statistical institute, has sent wm. gammell (us coordinator of dom ag) the following announcement of the formation of an international association for statistical computing, a section of the international statistical institute. lasc will be open to membership among those interested in promoting the theory, methods, and practice of statistical computing. the objectives of the association will be to foster interest and knowledge in effective statistical computing through international contacts among statisticians, computing professionals, organizations, institutions, governments and the general public in different countries of the world. special attention will be given to the developing countries. it intends to promote collaborative efforts with international, national, regional and other organizations and institutions with similar aims; foster evaluations of statistical computing techniques and programs; and, facilitate the exchange of computer programs and their documentation. the association plans programs and meetings particularly in conjunction with sessions of the international statistical institute. the organization will formally come into existence and have its inaugural meeting in december 1977, during the 41st session of the international statistical institutes in new delhi, india. the program planned for the lasc meeting includes (1) statistical computing and future directions in the context of an international association for statistical computing; a constitutional/business meeting; and, interactions between statistics and computing. mervin e. muller will speak on the purposes and aspirations of lasc; ivor francis, on the portability, evaluation, and certification of programs by statistical associations and societies; tore dalenius, on international implications of computer privacy; graham wilkinson, on language requirements and designs to aid data analysis and computing; julio ortuzar, on editing census and survey data using small computers in developing countries, including a review of experience with concor; and, geoffrey thomas, on assisting data processing activities in developing countries. [all appear to be quite relevant topics for iassist members. let's hope that papers or proceedings will be available.] further information and details can be obtained by writing: international statistical institute, 428 prinses betrixlann, voorburg, netherlands. annual individual member's dues structure is (1) members from developed countries, $15 per year; and (2) members from developing countries who request a reduced rate, $7 per year. sist newsletter vol.1, no. 2 while 0yen's discussion focuses on norwegian developments, the reader will be able to draw parallels to his or her own national situation. 0yen's final caveat, "it would seem unfortunate if the increasinn need for social science research in the policy field, and the exoansion of public support and political control of social science developments, were to be linked to efforts to tie the hands of researchers through the introduction of rules which might give the social sciences a serious setback. in fact, such a development might be threatening another integrity requirement, namely, the right to understand." 0yen's review is an incisive look at the consequences of privacy protection legislation created without consideration of social science needs. it is a sophisticated analysis of issues which will continue to command the attention of researchers and data archivists in the next decade. recommended as essential reading--an insightful introduction to the problems of privacy protection for the social scientist and by extension, the data archivist. report on the committee of european social science data archives, january 22, 1977 meeting of the committee of european social science data archives, 22 january 1977 the committee of european social science archives was established at a i meeting in amsterdam in june 1976 and held its second meeting in paris on 22 january 1977. the meeting was attended by representatives of the seven member organizations: the adpss-milan, the bass-louvain-la-fleuve, the dda-cooenhagen, the nsd-bergen, the ssrc-sa-essex, the sa-amsterdam and the za-cologne. the meeting agreed to organize one conference a year and to encourage the establishment of a number of time-limited european working parties of 3-^ data organizations each. za-coloqne, offered to host the 1978 meeting: this will focus on privacy legislation . the committee also instructed bass to invite established, puhlic-service data organizations (archives and broader data services offering access for a range of universities, centres) all across the world to take part in the constituent meeting of an international federation of data organizations. this meeting is scheduled for 21-22 hay at louvain-la-fleuve. the new body would be based exclusively on organizational membership and would commit its members to support concrete projects of co-operation. the federation will, if established, serve to bring data archives, data services and data centres together as institutions and would offer a parallel structure to the individually-based lassist. there was full agreement that the two bodies should complement each other and help each other through joint activities. there was also agreement that the statutes of if-do should give full recognition to iassist and commit the federation to close co-oneration with the association. for further information please write to philippe laurent at bass. 25. iassist quarterly 99 cd-rom and the data archive: beyond retrieval optical storage offers data archives a new medium for storing data and a change in the information retrieval environmenl as with any new technological development, it is important to evaltiate the impact and tradeoffs of adapting to new eqtiipment and systems. by stepping back for an overview of information storage and access, better evaluations may be made within the context of the "big picttire" of data use. before embracing this new technology, it is important to understand where we are going and how optical storage can help us. optical media offer some solutions to storage problems, but have not yet shown themselves to be solutions to the overall needs of information providers and users. by aim gerken' universirv of california, berkeley 'presented to the association of public data users october 22, 1987, washington, dc first of all, thanks go to forrest williams, census bureau, and to george hall and courtney slater, of slater-hall information products, and to ed spar, of market statistics. these people have provided us with cdrom lest disks for our evaluation, and i sincerely appreciate their support some of the points i will discuss today have been included in the report of the state data center's subcommittee on cdrom. this repon is available from john kavalaimis of the data user services division. my ideas do not necessarily reflect all the opinions of that subcommittee. consider the advantages of optical media as storage devices. cdrom and worm disks provide a minimimi of 10 years' secure storage. the disks provide random access, and are portable and compact however, there are some problems. for most data archives who rely upon large centralized computing systems, the shift from storing data on tapes and disk packs to storing data on microcomputer peripherals means a dramatic shift in responsibilm'. how will data archives provide mulnple user access to data on microcomputers'^ can archives handle the equipment and network burden? another problem with cdrom is that archives cannot produce cdrom disk copies in house; there is no local recording capability. this fall/winter 1987 100 iassist quarterly means thai if data archives are to produce optical backups of their holdings, they must use worm disks, and will have to acquire and maintain two kinds of optical devices. other problems are access speed, and in some cases, limited storage capability (cdroms only hold 4 reels of tape, and stf4 for california takes up 14 reels of tape). another issue to consider is that the software to access the data on cdroms and worms has not kept pace with the technology. even though information on microcomputers appears to be closer to the end user, the interfaces and support for access are often limited. and finally, we must ask when will opdcal storage be cost effective, and when will the major data distributors begin to use this technology? information retrieval is the process of extracting specific data items from specific records in a data file. at this time, information retrieval software is produced for each cdrom product since standard retrieval software for statistical data files is not yet available. at its best, this retrieval process is informed, flexible, logical, and fast however, since each data file comes with its own custom software there is a wide range in the quality of these software interfaces. users will be required to learn multiple retrieval processes. also, stadsticaj cdrom products are more expensive to produce since producing custom software for each data file requires extensive programming. the premastering and production of the optical disks themselves are different from producing magnetic tape. the physical process itself is more expensive, and the nature of the files on the disk differ. since the mediimi is randomly accessed, "disk geography." or where on the disk various records will be placed, is more important than with tapes. disk geography can improve performance if done carefully. disk indexes should be produced and mastered with the data file if the disk is to prove efficient in sophisticated retrieval applications. error detection and correction codes must also be provided. in summary, "raw" data files on cdrom should be more carefully developed than those provided on magnetic tape. to exploit the technology and access software, optical products should be distributed with value added files. after information is retrieved, "beyond retrieval", lies the worid of post processing where the information is manipulated and packaged to user specifications. at its best, this processing should be supported by an integrated system providing clear steps for users to follow. information retrieved from cdrom products should be compatible with inhouse software, and transfer of the information should be standard and simple. in oiu' evaluations, we considered three cdrom products in relation to data product preparation, information retrieval, and post processing. the three cdrom products we evaluated were: census test disk #1, which includes data and software for zip codes, population estimates, and the 1982 census of agriculture; slater hall's 1982 census of agriculture disk; and "your marketing consultant" form market statistics. the census disk offers simple extract programs that demonstrate the usefulness of optical storage. the procedures for retrieval are straight forward and can be done by any novice computer user. a highlight of die slater hall software is its online help for each variable in the data base, as well as its numeric search capabilities. "your market consultant" offers the capability of sorting data by user defined criteria, and allows users to add their own data to the tables. our evaluation at the state data program emphasized the retrieval and post processing capabilities of these products, but did not take a close look at the premastering and production of the cdroms. the issues of disk geography and indexing are yet to be evaluated. however, in the area of retrieval, it is clear that a standard procedure has emerged among these fall/winter 1987 iassist quarterly 101 products. users progress through a series of simple procedures, first selecting geographic areas, then identifying a table or variables, then extracting data from the disk, and finally producing display, report or file output what more do archives and users need from this kind of product? given that a cdrom product provides the information that a user wants, a separate procedure is often required after retrieving information. a user must sort and create custom tables, graphics, or spread sheets in another system. obviously, there is a need to coordinate software compatibilit>and command language among in-house software, so that users are not faced with a confusing array of software and transfer methods. in addition, the lack of numeric manipulation and statistical analysis in the cdrom software requires an interface with statistical software ctirrently in use. data as well as textual information must be moved in efficient ways and in compatible formats. none of the current cdrom products contain microdata, i.e. data for individual households, persons or establishments. the data that are available are summaries for geographic units, useful for many applications, but are only part of the larger universe of data. extracts of microdata files are regularly produced in the archive environment, and products and procedures using cdrom technology should be considered. our suggestions are that 1) data distributors keep an open mind about which kind of optical media to adopt, so that archives are not forced to develop dual systems, one for data acquisition (cdrom) and one for data file backup (worm). we also are concerned about the technology becoming outdated. 2) optical products should be mastered to allow for optimal retrieval by a variety of software, including relational database management software. 3) retrieval software should provide output that is compatible with other software. and the transfer of data should be an integral part of the software product 4) data distributors should view their products not only as stand-alone retrieval systems, but as pieces of a larger information support system. some questions for the future: how can optical storage be integrated into the existing information fabric? who will build the interfaces for informed transfer of data from system to system? do we place responsibility upon the private sector to develop retrieval products, and how involved cjm we become? what information systems are in use that identify, retrieve, transfer, and analyze data? how may these systems influence the development of cdrom and worm products? data archives have a long history of responding to technological changes: from ibm cards to low density tapes, from fioppy diskettes and random access disk packs to optical media. eric tanenbaum wrote: "archives have a privileged role among information providers for they were among the first to computerize. thus they offer a rare perspective from which to view the changes." (iassist ouarteriv . spring 1986) i believe that archivists will be studying these changes and will offer a unique perspective in emphasizing the integration of new technology into existing services and procedures. archivists and other information professionals have an opporttmity to participate in the development of standard, reliable storage media, to assist in the design and evaluation of powerful fiexible retrieval software, and to help create dynamic post-retrieval environments.n fall/winter j 987 using the century of prose corpus by louis t. milic ' cleveland state university the century of prose corpus (copc) is one of a number of compilations of texts that have been developed during the last three decades to facilitate a certain kind of linguistic analysis with computers. unlike the corpus thomisticus (the whole of the works of thomas aquinas), for example, the kind of corpus i am talking about is a descendant of the brown corpus, devised in 1961 by henry kucera and nelson francis of brown university. the brown corpus consists of a million words of edited american prose, all published during that one year and taken from a great variety of kinds and genres of printed materials, from humorous fiction to articles in learned journals. the devisers assumed that their corpus was large enough to represent nearly every type of linguistic unit that might be of interest to scholars. although it was intended to be machine-readable, it has generated two large volumes of analysis and documentation in which alphabetic and rank-ordered word-lists provide a view of the american vocabulary at that period, among many other valuable pieces of information about the language, to say nothing of the other areas of knowledge thai arc served by this work. it would not be inaccurate to compare the brown corpus and other corpora that have sprung up since to anthologies, such as those that serve as textbooks in courses in literature, history and other fields. the anthology is more than anything else a sample of the writing of a field or period, representative and typical of the totality of the population. of course, it is not a statistical sample because of its preference for the best, the best-known, the most infiueniial..., but it is a sample nonetheless. someone who has read through an anthology has a grip on the writing and thinking, the preoccupations of a genre, time period, nation... similarly, anyone who had been living on another planet during 1961 and on his return read through the million words of the brown corpus would have a pretty complete idea of what went on during that year. but of course that was not the intention of francis and kucera: their compilation was primarily a tool for research in language. their corpus gave rise to similar ones of spoken english and of british english. but beyond that, it led others to create more specialized corpora. the copc is one such. the copc is intended to represent a norm for the study of the english of britain during the eighteenth century. its actual delimitations are the years 1680-1780 and its dimensions are approximately 500,000 words. it is composed of two parts: a) the major authors: addison, berkeley, bolingbroke, boswell, burke, chesterfield, defoe, dryden, fielding, gibbon, goldsmith, hume, johnson, locke, adam smith, smollett, steele, swift, temple, walpole; b) the 100 background writers. part a (the major authors) contains 15,000 words from each of the twenty most prominent authors in three selections of 5,000 words each drawn from various stages of each author's production. this part totals 300,0(x) words or 60% of the corpus. part b can best be visualized as a ten by ten matrix, in one dimension representing decades of years: 1=1680-1689 2=1690-1699 3=1700-1709 4=1710-1719 5=1720-1729 6=1730-1739 7=1740-1749 8=1750-1759 9=1760-1769 0=1770-1779. springysummer 1992 13 and in the other ten different genres: 1 biography (a) 6 history (g) 2 periodicals (b) 7 memoirsa^tters (h) 3 education (d) 8 polemics (k) 4 essays (e) 9 science (n) 5 fiction (f) 10 travel (q) it will be noticed that there is a blending here of genre and subject matter, which can be rationalized by the claim that subject matter dictates conventions that amount to genre. there is a text of 2000 words in each cell. consequently there are ten selections (20,000 words) for each decade (one from each genre) and ten for each genre (one from each decade), the whole consisting of one hundred selections of 2000 words each or 200,000 in all, 40% of copc. each sentence of each text is identified by means of a header block. an excerpt of one of the part b texts follows: 5n03(1728)0001/021-p1 language is a set of words which any people have agreed upon, in order to communicate their thoughts to each other. 5n03(1728)0(x)2/079-p0 the first principles of all languages, buffier observes, may be reduced to expressions signifying first the subject spoken of; secondly the thing affirmed of it; thirdly the circumstances of the one and the other: but as each language has its particular ways of expressing each of these; languages are only to be looked on as an assemblage of expressions, which chance or caprice has established among a certain people; just as we look on the mode of dressing, etc. the header block 5n03(1728)0001/021 -pi is analyzed thus: 5n03: identifier of text (decade 5, genre n, accession no. 03) / 728: date of publication 0001 : sentence number 1 of selection 027: number of words in sentence pi : sentence begins a paragraph. the entire corpus holds on three high-density 3 1/2" diskettes (or on tape) and may soon be available on cd. it can be used on mainframes or on 386-type personal computers. in its present form, it requires the user to have access to a program package (such as eyeball, arras, word cruncher...) or to be able to program in a string-manipulation language, such as snobol or one of its derivatives (e.g., spitbol). i have devised several programs with which i have analyzed the various texts for later statistical treaunent. i shall mention two of these. the letter program performs the following: 1. counts the length of each sentence 2. produces a sequential list of sentence lengths 3. calculates and prints a. mean sentence-length in words b. standard deviation of the sentence-length 4. displays a frequency distribution of the letters in the text, both raw scores and percentage 5. displays the rank-order of the letters according to frequency, compared to the brown corpus and other corpora 6. displays a frequency distribution of word sizes (assist quarterly 7. displays frequency distributions of word-initial and word-final letters for words greater than five letters in length 8. in a summary, provides the following: a. total words b. hyphenated words c. number of sentences d. number of interrogative sentences e. net number of letters in the text f. calculated vowel-consonant ratio g. mean sentence length 1. in letters 2. in words (by a method different from 1) h. mean word-length in letters. the indexer program does the following: 1 . alphabetical word-index with raw frequencies 2. rank-ordered word-index of the 100 most frequent lexemes, with raw frequencies and percentages 3. counts a. tokens b. types c. hapax legomena 4. calculates a. type-token ratio b. hapax-token ratio c. hapax-type ratio. and of course, the programs may be applied not only to individual texts, but to groups, to decades, genres. parts and to the whole corpus. as can easily be noticed, these two programs alone acting on each of the texts in copc generate a very substantial amount of data which can be analyzed or treated in a number of ways. to illustrate one possibility out of many, 1 shall follow newton's principle about the relation of data and hypotheses: for the best and safest method of philosophizing seems to be, first diligently to investigate the properties of things and establish them by experiment, and then to seek hypotheses to explain them. for hypotheses ought to be fitted merely to explain the properties of things and not attempt to predetermine them... inspection of the data that is, the texts themselves and the output of the programs had led me lo observe that writings in the same genre showed a consistent use of certain variables. in order to examine this possibility, i organized the data of part b into ten variables for each of the hundred texts in it and analyzed this by means of the spss statistical data analysis j)ackage. although the number of variables is of course unlimited, i chose ten more or less at random. these variables consist of two sets: "standard" and arbitrary. the five standard variables (often found in the literature): 1. mean sentence-length (msl) 2. mean word-length (mwl) 3. number of types (typ) 4. number of hapax (hap) 5. percentage sum of five most frequent function words (fw) the five arbitrary variables are: 6. frequency sum of the letters "t," "i," and "o" (let). spring/summer 1992 15 7. frequency of the leuer "s" in final position (spin). 8. frequency of the letter "d" in final position (dfin). 9. sum of the two most frequent function words (t0p2). 10. number of nouns in ranks 1-54 of each selection (nn). the correlation procedure in spss produces pearson correlation coefficients for each pair of variables, when these have been arrayed in an appropriate form, as follows: text msl mwl let srn dfin typ hap fw top2 nn 6q45 29.82 4.48 24.17 22.20 12.64 695 455 21.64 13.85 1 8q16 40.86 4.65 23.12 17.79 15.06 762 534 19.04 10.14 7 0b90 38.63 4.49 24.18 15.56 18.63 846 615 19.64 9.90 3 0n74 62.47 4.85 24.63 28.49 10.53 737 510 21.65 12.25 10 4d83 45.52 4.70 24.70 23.54 15.32 637 385 18.97 10.23 6 8k78 34.02 4.55 25.80 17.46 13.73 683 440 19.12 10.59 6 6k82 46.88 4.54 25.25 20.89 9.93 578 339 20.30 9.68 8 9b73 38.32 4.70 24.78 24.11 11.17 740 495 20.88 11.57 7 9k98 34.00 4.64 24.40 25.08 14.98 627 395 19.79 11.62 9 2n09 34.50 4.62 23.55 27.04 12.07 669 442 21.90 13.10 7 the pearson correlations are as follows: positive .01 .001 msl-fin .24 mwl-shn .46 msl-fw .29 mwl-typ .31 mwl-hap .28 mwl-fw .55 shn-typ .29 mwl-top2 .59 sfin-hap .28 mwl-nn .46 shn-fw .24 fw-nn .43 sfin-t0p2 .29 top2-nn .46 negative typ-nn -.28 let-typ -.31 hap-nn -.30 let-hap -.29 as can be easily seen, a good number of these are quite significant, some at the one percent, some at the next level, mostly positive, although a few are negative. of course, some of the conrelations are significant but meaningless, as they represent merely functional relationships, e.g., types and hapax, function words and "top two". but others suggest something factual and possibly important about the relationship of genre to the quantitative fabric of texts. to look into the possibilities of this relationship, we must go deeper and discover which genres select which variables. by subjecting the data to analysis of variance (anova), we find the pattern in the following matrix: {assist quarterly 1 2 3 4 5 6 7 8 9 10 variablea b d e f g h k n q msl + + + + mwl + + spin + + dfin + + + let + typ + + + + hap + + + + fw + + + top2 + + + nn + + + + total 7 6 7 3 9 7 4 5 8 6 it is plain that certain genres select more significant variables than do others. numbers 4 and 7 (essays and memoirs/ letters) seem less distinct than the others. numbers 5 and 9, on the other hand (fiction and science), are much more distinctive. following newton's recommendation, therefore, we are free next either to devise hypotheses about these relationships or try new experiments to deepen our understanding. a possible explanation might be that the conventions of fiction and science writing are much more strict than those of essays, memoirs or letters, and that this strictness manifests itself at the quantitative microlinguistic level. another might be that the term "genre" is not as rigorous or as easily defined as is generally believed. at any rate, to feel confident about such hypotheses would require further analysis of factors by means of regression or other advanced statistical techniques. this simple illustration is only intended to reveal a small fraction of the immense possibilities for study and research that are latent in a carefully constructed corpus of substantial size and extent. ' presented at the lassist 92 conference held in madison, wisconsin, u.s.a. may 26 29, 1992. spring/summer 1992 sist newsletter vol.1, no. 4 dpls requires that the user of the data agree to the following conditions for receipt of the file: (a) neither the transmitted file(s). listing (s), nor any copy thereof in whole or in part, shall be disseminated, or sold, to any further party; (b) the data provided shall be used solely for professional or other official purposes related to teaching, research, public service, public policy, and planning; cc) publlcauon or dissemination of data, or the results of analyses of che data provided, shall be based on a minimum of five individual subjects and/or .three organizations per distributional or tabulated cell and neither any individual subject or organization shall be explicitly identified. uhere exception to this requirement is requested, permission for access co the data shall be granted by the daj:a and computation center faculty policy committee upon review of the research needs of a particular project; (d) an abstract or copy of any published document upon which the analysis is based be provided to dpls; {e) dpls will not be responsible for losses or inconvenience resulting from delayed delivery or errors due to defects attributable to the requestor's own equipment. dpls will bear the cost of replacing the data in the event that the defect is attributable to dpls's work; (f) responsibility for the accuracy of the data and documentation rests with the donor unless dpls has itself been responsible for producing the file or documenflease sign below if you agr donor: na telephone dpls discussion paper/ donald f. harrison on work in progress: directory of directories donald f. harrison national archives washington, dc the us action group on process produced data files, in its mandate to prepare a "directory of directories", (lassist newsletter vol . i, no. 3, may 1977), has prepared the following initial entries and desires feedback from the membership. with two exceptions, entries describe only social science data files, printed and available in the united states. because of the nature of where files are created, plus the specialized knowledge of the list makers, this list is heavy on federal directories. the committee seeks information on: 1) additional directories not listed and available for description; and, 2) additional or different information that may be desireable in the entry format. please send any suggestions to: donald f. harrison chairman process produced data ag machine-readable archives division (nnr) national archives washington, dc 20408 28 sist newsletter vol.1, no. 4 association of public data users. (apdu). data file directory . august 1977. 644 pp. typescript. contains apdu individual and organizational membership directory as well as descriptions of contributing organizations. lists of selected data files available from each contributing organization with file name and keyword indices. available to membership only. for membership information, write: apdu, box 9287 rosslyn station, arlington, va 22209. inter-university consortium for political and social research. (icpsr. guide to resources and services , (annual publication). contains data files available for purchase, technical configurations, prices. write: icpsr, university of michigan, ann arbor, mi 48104. library of congress. machine-readable cataloging (marc); price announcements and subscription rates . washington, dc, undated, / pp. typescript. describes 1/ subscription service data bases in marc ii. available from l/c. prices, available technical configurations. write: library of congress, processing department, cataloging distribution service, washington, dc 20540. sessions, vivian s. (ed.) directory of data bases in the social and behavioral sciences . new york: social associates/international inc. 1974. hb 300 pp. badly out of date. describes files and other data configurations at 685 institutions/public and private, us and foreign. describes at the file level, availability, etc. texas christian university, drug abuse epidemiology data center (daedac) newsletter. describes files available from daedac, including the "original data file catalog". for further information, write: associate director, dr. laverne knezek, daedac, texas christian university, fort worth, tx 76129. u.s. bureau of the census. catalog , (issued quarterly). contains social and economic data compiled by the bureau on a quarterly basis. write: users' service staff, data user services office, bureau of the census, washington, dc 20233. u.s. department of commerce. national technical information service. directory of computerized data files and related software , (published periodically) . contains data files available from ntis or from federal agencies. technical configurations, prices. write: ntis, 5285 port royal road, springfield, va 22161. cy of federal agency education data tapes . by barbara a. feller. (pub. nces 76-206). washington: gpo 1976. 177 pp. apps. describes educational data bases available to the general public from federal agencies. prices, technical configurations. available gpo. u.s. department of health, education and welfare. public health service. directory of automatic data processing systems in the public health service . (published annually). 510 pp. typescript with indices. describes approx. 500 files in phs. directory not yet available to public. u.s. department of health, education and welfare. social security administration. some statistical research resources available at the social security administration . a description of the available lifetime overall earnings data files from the several -purpose research files making up the continuous work history sample (cwhs). 29 newsletter vol.1, no. 4 u.s. department of health, education and welfare. standardized micro-data tape transcripts. (dhew pub. no. 76-1213). washington: gpo 1976. 30 pp. describes data bases for health statistics available from the federal government. prices, technical configurations. available from gpo. u.s. department of labor. bureau of labor statistics. bls data bank files and statistical routines , (draft brochure). describes 30 data files available from the bureau of labor statistics. describes publications. write: commissioner, bls, dept. of labor, washington, dc 20212. u.s. general accounting office. 1976 congressional sourcebook series (opa 76-23) federal information sources and systems; a directory for the congress . washington, dc: gpo 1976. 456 pp. paperback. an indexed reference guide to over 1,000 federal sources and information systems in 63 federal agencies, which contain budgetary, fiscal and program-related data. available from gpo. u.s. national archives and records service. catalog of machine-readable records in the national archives of the united states . washington: gpo, 1977. 37 pp. describes 99 files created by federal agencies and retained in the national archives. prices, available technical configurations. write: national archives (nnr), washington, dc 20408. university of waterloo. leisure studies data bank. april, 1977. 23 pp. lists 24 files created by private and federal canadian agencies. in french and english. write: dept. of recreation, university of waterloo, waterloo, ontario, canada. n2l 3g1 . discussion paper/gary m. grandon using spss mult response to generate filtered marginals gary m. grandon social science data center the university of connecticut, storrs with the release of spss's version 7.0 this year, several additions have been added to its already extensive battery of analysis programs. mult response is one of these programs. it generates frequency counts and bivariate tables for "dummy" coded multiple response questions. these multiple response variables are a real nuisance to the analyst and the program at face value provides a means for their interpretation. considerations for the analysis of multiple response data are not the subject of this paper though, but rather the use of this program for quite another purpose; the generation of "filtered marginals." filtered marginals are typically basic frequency counts and percentages for specified variables for each of a number of subpopulations within a study. spss has in the past provided *select if statements to facilitate the processing of subpopulations. each such statement followed by frequencies procedure statements will generate filtered marginals. the inherent problem with this type of coding 30 ii^$sist newsletter vol.1, no. 4 first lassist west european meeting june 2729 , 1977 report of the west european lassist workshop held in copenhagen, june 27-29, 1977 submitted by per nielsen, danish data archives members present europe tomasz bankowski , zowar computing centre, warsaw, poland flemming bigom, danish data archives, copenhagen, denmark merete watt bool sen, danish data archives, copenhagen, denmark ulf christoffersson, university of gothenburg, sweden j.c. deheneffe, university of louvain-la-neuve, belgium bartlomiej gasiorowski, polish academy of sciences, warsaw, poland j^rgen grosb^l , danish data archives, copenhagen, denmark astrid bogh lauritzen, danish data archives, copenhagen, denmark cees middendorp, steinmetzarchief , amsterdam, the netherlands ekkehard mochmann, university of cologne, federal republic of germany per nielsen, danish data archives, copenhagen, denmark krzysztof ostrowski , polish academy of sciences, warsaw, poland karsten boye rasmussen, danish data archives, copenhagen, denmark north america e. m. avedon, university of waterloo, ontario, canada carolyn geda, university of michigan, ann arbor sharon chappie henry, data clearing house for the social sciences, ottawa, canada agenda june 27: june 28: june 29: what do we need to know about each european archive? how much detailed information is really required? can we design one questionnaire that will include all of the required elements and not be a burden to the respondent? should archive holdings be standardized with respect to documentation? if so, what should be included in these standards? how can we ensure that these standards are adhered to and used by primary research personnel prior to primary analysis? what are the sources of data? should archives classify and index these data in a standard manner to facilitate data retrieval as well as crossnational and international research? if so, how should this be accomplished? sist newsletter vol.1, no. 4 on behalf of the west european lassist secretariat and the dda, per nielsen welcomed the participants and expressed his gratitude that so many "guest invitees" from the east european and north american regions were present. to bring everybody up-to-date, information was provided on the latest "data-conferences": carolyn geda gave information about the lassist north american conference in florida (february 1977), where 35 data information professionals had attended. carolyn geda and sharon c. henry then discussed the second north american lassist meeting in toronto (may 1977), in which the number of attendees had doubled compared to the florida meeting. working papers and ideas from the above meetings were useful inspiration throughout the copenhagen meeting. ekkehard mochmann reported on the moscow conference on information and documentation in social sciences from which he had just returned; the meeting was an initiative of the vienna centre, hosted by inion and supported by unesco. the upcoming lassist meetings in chicago (february 1978) and uppsala (august 1978) were described. the issue of the first day in copenhagen was discussed at length. there was agreement that it would be possible to design a data organization registry form that included all the elements required for a full description of data service organizations. consequently, it was decided to construct such a data organization registry form and to recommend this form to ifdo for use; the intended audience for the form was to be the administrators of the various data organizations. as source material for the data organization registry form the following documents were used: 1. questionnaire on data acquisition policies and problems (marcia taylor, ssrc survey archive, essex) 2. questionnaire on data preparation procedures (eric tanenbaum, ssrc survey archive, essex) 3. preliminary list of data elements and subject terms for lassist data archive registry in canada (lisa lasko, institute for behavioral research, york university; and lana prokop, university of toronto) 4. preliminary outline for a "a guide to providing social science data services" (alice robbin, university of wisconsin; and laine ruus, university of british columbia) 5. request from alice robbin and laine ruus regarding service documents 6. input form (directory of data bases in the social and behavioral sciences) 7. data documentation form (data clearing house for the social sciences) the group constructed a preliminary questionnaire form consisting of material selected from the above documents; then modified, added, and deleted; and finally ended up with a handwritten version to be computerized before a final review during the wednesday session. the working group made an attempt to set up a list of elements describing a data file that would be useful from the user point of view: (i) library information (such as card catalogue) (ii) archive information/study abstract (iii) study description (in machine-readable form) (iv) list of variables (in machine-readable form) (v) codebook (in machine-readable form) (vi) classification/index (in machine-readable form) (vii) data (not part of documentation, but the basis for the file) (viii) special publications (either in print or machine-readable) sist newsletter vol.1, no. 4 each of the documentation items were subjected to detailed discussions, and some recommendations were agreed upon; a summary is listed below: (i): library information : the north american lassist classification action group {coordinated by sue a. dodd, university of north carolina) has set standards for cataloging machine-readable data files. data organizations and primary research personnel should adopt the library information recommendations of this action group. (ii): archive information/study abstract : each data organization has its own way of publishing its data holdings in an inventory; it seems difficult to make recommendations regarding production of study abstracts. however, it is recommended that data organizations as well as research institutions produce information at the abstract level. (iii): study description : it is recommended that the standard study description scheme (developed as a result of a meeting in copenhagen in june 1974) be tested by several data organizations. as a concrete step in that direction several data holders agreed to test this instrument during the next year: bass , louvain-la-neuve, belgium pda , copenhagen, denmark icpsr , ann arbor, michigan, usa lsdb , waterloo, ontario, canada steinmetzarchief , amsterdam, the netherlands zentralarchiv , cologne, federal republic of germany ulf christoffersson , university of gothenburg, sweden five study description questionnaire forms will be disseminated to each of the above holders along with an instruction in the use of the standard study description scheme. each data holder will send study descriptions (in english language) of five different files to the zentralarchiv. preferably, the study descriptions should be submitted on tape; however, the za would accept filled-in questionnaire forms from data holders unable to make the study descriptions machine-readable. attempts will be made to furnish the necessary software for printing the study description in the participating institutions. the outcome of this testing will form the basis for an extension and updating of the instruction manual which is presently being developed and distributed from the dda. at the end of this process, final recommendations can be made. (iv): list of variables : it is recommended that the list of variables be available in print as well as in machine-readable form. (v): codebook : the "ideal" codebook was outlined by the workshop as consisting of the following 15 items (the first three items referring to file level, the last twelve to be supplied variable by variable where applicable): (1) title of study (file and subfile names) (2) format of the data-file (3) comments at file level (e.g. concerning application of missing data codes; special weighting features; special precautions for use) (4) variable identification (number; label; short name; mnemonics) (5) variable source reference (6) variable location and length (7) variable type (alpha; alphameric; numeric; symbolic) (8) number of decimal places (scale of measurement) (9) source statements/texts/questions/scale descriptions/introductory statements related to the responses 9 sist newsletter vol.1, no. 4 10) code values (11) code descriptions (12) comments: coding and field work (coding and interviewer instructions) (13) variable contingencies (filter; skip; probe; control questions) (14) variable consistency (15) derived variables (vi): classification/index : it is recommended that future lassist meetings be concerned with this topic. a number of classification schemes (on file and variable level, respectively) are available; however, further testing and elaboration is required before final recommendations can be made. (vii): data : single punch data is preferable for archiving purposes. (viii): special publications : it was recognized that many data archives produce publications concerning specific files (examples were examined in hard copy). there was some discussion regarding machine-readable special publications and the future directions that this area of documentation may follow. the machine-readable preliminary data organization registry form was reviewed. questions were added, deleted, modified. this work was finished with agreement on a data organization registry form to be disseminated to the participants for completion; comments arising from the application of the form will be reported to the dda; such comments may make yet another editing process inevitable. the working group discussed sources of data. this topic had been subject to discussion during the data organization registry form and data documentation sessions of the preceding days; however, the classification/indexing problems regarding data from various sources could not at this meeting be operational ized to a level allowing for recommendations. on the contrary, it is evident that there is some confusion right down to the level of terminology. it would be very useful if the classification and process-produced data action groups of lassist could come up with a tentative taxonomy, taking into consideration the increasing number of different data sources relevant to the social science community. summary and prospect jt) the data organization reg istry form will be disseminated from the dda to the participants as soon as final editing is completed. (2) no later than october 1st participants will complete the form on behalf of their organizations and return the completed version to the dda with any comments and suggestions that may arise from the respondent role. (3) the data organization registry form will be recommended to the ifdo for use. (4) concerning the study description scheme participants will receive a manual (applications & instructions). (5) participants who have agreed to take part in the testing of the study description will receive five study description forms to be completed in english for five different files in their holdings. (6) no later than december 1st, the completed study descriptions will be transferred to the za (preferably in machine-readable form) along with a short report on the test experience. 10 machine readable archives: user survey by sue gavrel machine readable data archives public archives of canada this article is an abridged version of a stavko manoj'loviah of the social science western ontario, london, ontario, canada. background in november, 1981, the machine readable archives division undertook an analysis of the requests and inquiries it had received for its services over the last few years. it became evident during the analysis that although the division provided a great deal of information on the establishment of machine readable programs, procedures and practices used to archive machine readable records, the actual holdings of the division were not being used. there were several possible reasons for the lack of use of the files, but the main conclusion from the analysis was that the mra did not have sufficient information on who the user population was and therefore was not distributing information about its holdings to the appropriate areas. the users of the traditional archival records were not the same as those who would use machine readable data. it was decided that a survey of our existing users and potential users should be undertaken. the social science computing laboratory of the university of v/estem ontario was contracted to undertake the survey. several steps were involved: the drafting of the questionnaire, the collection of the data and the analysis. the project report written by s. paula mitchell and computing laboratory , the university of spanned the period february 1, 1982 to october 30, 1982 and the total cost of the project was $12,478.07, which included staff charges, computing services, printing, stationery and postage. the main objective of the survey was to obtain information on the needs of canadian social scientists who created or used machine readable data. the results of the survey would be used by the mra to help plan an expanded public service program and determine priorities for the allocation of resources. a secondary aim has been to identify sources of machine readable data in canada. specific objectives of the survey were as follows : • to provide a measure of the relative usefulness of a number of products and services to the user community; • to provide an indication of the preferred mode of acquisition for machine readable data files and the preferred type of codebook; • to identify the location of machine readable data files for the purpose of expanding a union list of machine readable data files in canada and identifying files of historical research value and national significance falling within the archival mandate of the mra; • to provide an indication of the public's awareness of the holdings and services of the mra; • to provide an indication of the need for a national organization to coordinate selected information services for users of machine readable data (including special workshops and/or conferences) ; • to provide a mailing list of individuals interested in receiving the i'tra's publications or announcements; and • to identify the need for products or services not covered by the first objective noted above. survey population and data collection the survey population was derived from two social science information systems, the canadian register of research and researchers in the social services (social science register) developed by the social science computing laboratory, and the canadian directory service of social scientists (socscan) , established by the social science federation of canada between 1976 and 1978. the register supercedes the socscan database. the target population was derived from these two data bases by selecting variables defining methodology and research orientation. although the target population does not include all social scientists who are current or potential users of machine readable data, the large number (4,065) does provide a good representation of users of social science data in canada. the questionnaire, designed to be as simple and short as possible, contained 15 questions. data were collected using a mail-out/mail-back process. a covering letter explaining the objectives of the survey was sent with the questionnaire. a follow-up reminder was sent to all non-respondents three weeks (15 working days) after the initial mailing. findings survey response as mentioned earlier, 4,065 questionnaires were sent out. of the 4,065 questionnaires, 82 were returned as undeliverable mail, thus making the actual survey population 3,983. response to the survey was very good with an overall response rate of 48.95 percent, or 1,950 completed questionnaires . survey results the following section provides a brief summary of results. findings are reported on a question-by-question basis. ql. prior to receiving this questionnaire, were you aware of the activities and services of the machine readable archives? the results of the survey indicate that awareness of the activities and services of the mra among social scientists is low. of the 1,930 respondents answering this question, 14.2 percent indicated that they were aware of the division's activities and services prior to receiving the questionnaire packet. there was no significant difference in the level of awareness between respondents in the academic and non-academic sectors (14.2 percent and 14.1 percent, respectively). 02. do you use machine readable (computer) data in your research, teaching, or other activities? of the 1,908 social scientists who responded to this question, 1,293 (67.8 percent) reported that they are users of machine readable data, and 615 (32.2 percent) reported that they do not use machine readable •~ ~~~,—^ continued data. given the quantitative orientation of the survey population, it is somewhat surprising that close to one-third do not use machine readable data. however, the very high interest shown by this "non-user" group in receiving information on the services and holdings of the mra (76.8 percent wished to be on the mra mailing list) suggests that the survey was on target in terms of reaching quantitative social scientists who are potential users of the mra. it is possible that some respondents in this "non-user" group found the question ambiguous: either limited the inclusive time frame of "use" to the present, or interpreted the question to mean use of secondary data rather than their own primary data. the distribution of users of machine readable data geographically and by sector of employment is proportional to the distribution of all respondents in these two respects. except for ontario and quebec, the percentage distribution by province of social scientists who use machine readable data does not vary by more than plus or minus 0.6 percent relative to the percentage distribution by province of all respondents. in ontario, there are 1.9 percent more users relative to the percentage of social scientists from ontario constituting the total respondent group, and 3 percent fewer in quebec. the distribution of users by employment sector (academic/nonacademic) varies by only 0.3 percent relative to the sectoral distribution of social scientists in the responding population. q3. is there a facility within your department or organization which provides information and/or other services for machine readable data? 1,269 users of machine readable data responded to this question. a very high percentage (93.2 percent) reported that there is a local facility which provides machine readable data services; 86 users (6.8 percent) reported that there is no local facility providing such services . 04. please identify the facility within your department or organization which provides information and/or other services for machine readable data. more than half (53.1 percent) of 1,172 users who are availed of a local facility are serviced by a computing center only. while the range of services provided by any one computing center may approach the full range of information, access, and utilization services provided by specialized local facilities, computer centers typically provide limited support services in these areas. it is a significant finding that such a high proportion of machine readable data users are availed of limited informational support services for finding and using machine readable data. 369 respondents (31.5 percent) indicated that more than one facility provides local machine readable data services. 25.1 percent of users have a specialized local data support unit (such as a data archives) available to them. for 23.3 percent of users, the local research library integrates some level of service for machine readable data with other services . q5. please indicate the potential usefulness of each of the following information products or services to your research, teaching, or other activities: a) catalogue of holdings of the machine readable archives; -^~continued on page 22 vol26no1 4 iassist quarterly spring 2002 iassist quarterly spring 2002 5 this reason, evaluation was an important component during the early stages of development. the criteria we used in these evaluations were robustness, performance, and institutional support. robustness means how much of the realm of all likely cases and circumstances the item could handle. for software, this concept includes its portability to a variety of platforms and programming languages. performance means: first, has the item been implemented at all? and secondly, how well, how simply, and how quickly does the item work? institutional support is a measure of how committed organizations are to the support and future development of the item in question. our second strategy for controlling duplication of effort was the dry principle, which is, as stated by hunt and thomas, “every piece of knowledge must have a single, unambiguous, authoritative representation within a system” (p. 27). that single representation is then used to generate data products and even software used in the processing and publishing of the data sets. the dry principle has helped reduce the amount of recoding and reformatting at cpanda, however imperfectly applied. evaluations i and other cpanda members spent a great deal of time evaluating standards, practices and software for possible adoption into our archive. they include: metadata (codebook) format standards, standards for controlled vocabularies to describe data sets, software for analysis and archival management, full text indexing software, and software used to create codebooks in the ddi format. case i: codebook formats one of the first decisions made by our group was what format to use to store the metadata for the data sets. the choice was not difficult, and we immediately selected the ddi format for the codebook presented to the users. its advantages are many and obvious. ddi uses xml, which has many positive features. as far as robustness, extensible markup language (xml) is well understood, uniform and extensible. regarding performance, there are many applications designed to parse and process xml. institutional support for xml is strong, with many groups working on using and improving the standard. the ddi abstract today, there are many customs, standards and standard applications which groups just beginning a data archive can choose to adopt or not adopt. this presentation discusses the strategies that the cultural policy and the arts national data archive (cpanda) team used in the development of the archive, including the criteria by which established standards and practices were evaluated. the choices made by the team as they worked to develop the archive will also be discussed. introduction this past year, the cpanda team undertook to design and implement a new data archive. in developing a new archive, one important consideration is to do so in a cost effective way. naturally, certain features are desired, but one does not want to waste effort through unnecessary duplication. like everyone else, we want to get where we want to go, without doing it the hard way. the primary guiding principle, then, in our data archive development was to design the archive with the features we desired while limiting the amount of duplication of effort required on our part. this is not a novel idea, but it is better to have a sound idea than a novel one. because we began our archive development long after many other individuals and groups, including many of you in this room, had already done so, our first strategy was to evaluate earlier efforts to see what we should adopt or build upon, and what we should create ourselves. our question was: which practices should be adopted, and which, like a mark twain classic, are better talked about than enjoyed? our second strategy was to follow, as much as possible, the dry principle as articulated by andrew hunt and david thomas (hunt and thomas, 2000, p. 26-28). dry stands for “donʼt repeat yourself.” this principle states that you should have only one canonical version of information about both the data and the software, and that all other forms of the information should be regenerated from the canonical form and should be disposable. these are the strategies we used for developing a data archive in a time when many other efforts had already come before us. the first strategy for controlling duplication of effort was, not surprisingly, to try to build upon earlier efforts. for by vernon leighton* developing a new data archive in a time of maturing standards 6 iassist quarterly spring 2002 iassist quarterly spring 2002 7 itself is robust in that it handles many different aspects of data sets. it has institutional support through many iassist members, including icpsr, nesstar, etc. these organizations are committed to supporting the ddi and extending it to cover even more types of data sets. performance is one area where the ddi has problems. because of the richness of its cross-references and structures, applications that use the ddi are easier to design if they treat it as a static object. this object orientation favors the use of tree representations of the document, such as the document object model (dom) over event driven representations, such as the simple api for xml (sax). because of the large size of these xml documents, the use of built-in dom methods for data retrieval can be quite slow and can use a great deal of memory. in our own operation, the naïve use of built-in dom data retrieval calls was unacceptably slow, and the use of superior data management resulted in a 300-fold increase in performance. for example, we have a program that pulls from each variable in a codebook the text of the question in the survey that relates to that variable. the question text is then loaded into a relational database for search and retrieval functions. if the function repeatedly obtained all variables through the dom method $codebook à findnode(“/codebook/ datadscr/var”), the algorithm runs very slowly. we solved the problem by creating a codebook object that loaded most variable-level information sequentially during initialization into a hash data structure. the repeated access to the hash was up to 300 times faster than the dom method as implemented in libxml2. other formats that we could have chosen for variable level metadata include the triple-s format in xml (hughes, jenkins and wright, 2001), and a variety of proprietary formats developed for individual software applications, such as spss and sda. all these formats suffer from a lack of elements for bibliographic information about the whole data set, lack a richness of structure for describing many possible types of data set layout, lack tools to parse them in many languages on many platforms and lack institutional support to address those shortcomings and extend them in the future outside of the groups that created them. on the positive side, these formats tend to be based on plain text, and can usually be parsed in a straightforward manner. because of their simplicity, some of them offer greater processing speed than some implementations of the ddi. skeleton in closet when i stated that we chose the ddi for our format to present to users, i confess that i was misleading you. when we began the project, we chose the survey documentation and analysis (sda) software to manage the online analysis of our data sets. at first, we had difficulty in getting xml parsing tools with acceptable features loaded onto our server. in order to get moving, we chose to represent the variable level data initially in sda̓ s ddl format. the data set processing tools were designed so that the codebook was an object accessed through a standard set of methods. the codebook object was originally implemented on top of sda̓ s ddl format, and the xml ddi format was generated by one of the programs as a transformation of the ddl. therefore, the authoritative xml version of the codebook was not the canonical version used by the archive software; instead, the xml version was a temporary product subject to being discarded and rebuilt. this situation had to change, because so long as the xml version was a disposable product, no changes could be made to the canonical version using its rich set of fields for data description. case ii: controlled vocabularies cpanda provides bibliographic descriptions of each of the data sets that we manage, which means that subjects and descriptive terms must be used to assist with the bibliographic control of the data sets. because we need a controlled vocabulary of descriptive terms, cpanda staff evaluated a variety of possible vocabularies. after discussions with icpsr, it was decided to work with them to add to their thesaurus new terms related to cultural policy and the arts. by working with them, we would be able to tap into a large vocabulary of terms related to surveys and social science research, and they would increase the base of organizations that use their terminology. it was also decided to use library of congress subject headings in addition to the icpsr terms, to allow bibliographic access by those familiar with that more commonly used vocabulary. the advantages of the collaboration with icpsr are 1.) we do not have to create our own thesaurus of terms, 2.) we can use a vocabulary already rich in terms related to social science data sets, and 3.) we can add to it, enabling us to customize it somewhat to our own needs. it is also robust and has institutional support. case iii: software for analysis and archival management several systems have been developed to perform analysis of data sets and manage data archives. because analysis is quite exacting and difficult to program well, it would require a great deal of effort to duplicate. there are also many logistical issues related to the management of a web site, and adopting software that would manage the site would again save effort. we examined the nesstar system (ryssevik and musgrave, 2001) for both its analysis and archive management functions. most important, it passes the first hurdle of performance: it has been implemented in a production mode, not just experimentally in beta mode. it is a 6 iassist quarterly spring 2002 iassist quarterly spring 2002 7 well built system which could readily accomplish our goals. it uses the xml ddi codebook. it allows full text searching for variable discovery. it deals with many archival management functions in a robust fashion and is institutionally supported by the ec. we chose not it adopt it primarily because it is on the microsoft nt platform, and we have committed ourselves to the unix platform. so, we did not select it, but not because it came up short in our evaluation. we examined the reports on harvardʼs virtual data center (altman, et. al, 2001), and were quite impressed with their plans. if they succeed in making the system robust and stable in the way they have envisioned it, it will be quite an attractive option on a unix-like platform. being open source, it will be robust in the sense that if it lacks a feature and the local archive has the ability to create that feature, the feature can be added. many useful archival management features and a flexible, distributed architecture are in development. the software will use the xml ddi. it has the institutional support of harvard, mit, and the nsf. in the performance realm, however, we feel that right now it is at too preliminary a stage of development to be adopted. it seems to only have been implemented in an experimental stage, without a track record of stability and portability. it does not seem to be robust yet. at last report, it could only run on some versions of linux. these details may have changed by now, but as of the time of the evaluation, we could not commit to its use. we remain interested and hopeful that we might adopt it in the future. however, it seems to be not quite there yet. we did not investigate data ferret from the u.s. federal government as closely as perhaps we could have. the interface presented to the users did not appeal to us, and we decided that we would not consider it further. the software that we in fact chose to license for the analysis of the data sets online is the sda software from the computer-assisted survey methods program at berkeley (http://sda.berkeley.edu:7502/). although it has limitations, such as not having source code available, it does have virtues, like the fact that it already works well in production mode. one can operate it within a data archive as a black box analytical engine and design oneʼs own interface to initiate online analysis. the results of the analysis can be captured and customized. it does not have archival management functions, but there is nothing to prevent those functions from being added to the interface that accesses sda. depending on how vdc is implemented, it is not inconceivable that sda will be complementary to it. case iv: full text indexing on our site, as was already explained, we have developed a feature that allows users to search the full text of questions, variable labels and other variable notes for terms of interest in the variable discovery process. in order to accomplish this, some software must index the appropriate text and provide operators for users to query the index. it is possible to create ones own full text indexing application, but doing so would require a fair amount of effort, while others have already created such tools. here again, we evaluated a variety of products before settling on one option. we may change our decision in the future. many database management systems have full text indexing capabilities. we decided against using a major commercial product like oracle because of price. the price of oracle is not just the licensing fees, but, because it is quite challenging to administer, one has to add the cost of hiring experienced personnel who are able to manage it. that personnel cost is perhaps higher than the licensing. otherwise, oracleʼs functionality, robustness and institutional support would make it acceptable. we settled upon mysql for our database management system. it is open source and free software. it is popular and has a large following in the open source community. it is fast and scales well, even if it does not implement all sql standard features. it has a built-in full text indexing function. unfortunately, the built-in indexing has very limited features and operators available. because the software that we chose has limitations, we evaluated products that might supplement the product used. several major internet search engines offer indexing software that can interface with database management systems. we discovered that these companies will license their products for modest sums if they index web pages, but they immediately begin charging much more if one uses the interface to index databases. perhaps they know that such interfaces can be quite commercially lucrative and charge accordingly. in any case, many of these search engines are not available on many platforms and with apis to many programming languages. as a result of these evaluations, we have settled for the moment on the built-in features on mysql and we hold out hope for improved functionality in the 4.0 release. note post-conference: in july of 2002, we found that the swish-e search engine was compatible with our web programming environment. we are now using the swish-e search engine for full text indexing of our mysql database, by extracting the records with a perl program and wrapping them in xml before handing them to swishe. the coordination of the swish-e index and the mysql database is handled using the swish-e php class developed by olivier meunier. again, we are able to leverage the sophisticated indexing technology of swish-e and the data retrieval power of mysql. 8 iassist quarterly spring 2002 iassist quarterly spring 2002 9 case v: codebook editors here i would just like to report that we evaluated and attempted to use the maddie software developed at the university of minnesota to create ddi conformant xml codebooks. the result of this experience is that we do not plan to use maddie in the near future for codebook editing. the primary problem as of six months ago with maddie is that it is not very robust. many features that one comes to expect from text editors in general are either not present or do not perform acceptably. i believe that the problem with the maddie project is that the project has tried to reinvent the wheel. they have built their own text editor to manage xml dtds. they have had to reimplement many of the features that come standard on a myriad of other text editors. because of the enormous amount of reimplementation involved, necessarily developed on a limited budget, the project has been overwhelmed and the product suffers from a lack of robust features. future efforts should perhaps try to build upon a text editor that already has a rich set of functions but which is open source and programmable. an editor like emacs could be given custom extensions to offer xml specific features based on a dtd or schema. one would then leverage the other editing features already developed by the open source community. that having been said, i would like to thank wendy thomas, bob wozniak, and hicham berrada for letting us use their code. they have put a great deal of work into its development, and they should be appreciated for their efforts, despite the results not being what one might hope for. evaluating evaluation in all of our evaluations, cpanda sought to find workable solutions to data archive needs by adopting or building on the work of others. those efforts need to be ongoing. we explain our criteria and findings in order to foster discussion and reevaluation and welcome criticisms. the dry principle the dry principle is quite common throughout computer science. programs have functions and included libraries to reuse code. object brokers such as corba, com and soap are designed to reuse executables. configuration files allow compiler options to be specified in one unambiguous location. the principle can be pushed to advanced levels in cases such as compiler compilers and automatic code generators, where the source code itself is a temporary duplication of the base information stored in one unambiguous location. the dry principle was successfully implemented at cpanda with regard to codebook metadata management and web page creation, but it was not so successfully implemented in terms of dynamic, automatic construction of the software itself. in the data realm, we have established a canonical version of the codebook and build all other codebook products dynamically from it. see chart 1 for a diagram of that process. on the web site, we use server-side database access to dynamically generate pages customized to the userʼs context. automated source code generation has not been systematically pursued, but should be placed on the schedule. chart 1 conclusion the concept of limiting duplication is one that is worth making a priority in the development of any software project. the two strategies of adopting previous efforts and minimizing duplication have allowed us to develop our data archive in an effective and efficient manner. 8 iassist quarterly spring 2002 iassist quarterly spring 2002 9 notes: altman, m., et. al. 2001. “a digital library for the dissemination and replication of quantitative social science research,” social science computer review, (v. 19, no. 4, winter, 2001), p. 458-470. hughes, k., jenkins, s., wright, g. 2000. “triple-s xml: a standard within a standard.” social science computer review, v. 18, n. 4, winter 2000, pp. 421-433. hunt, a., thomas, d. 2000. the pragmatic programmer: from journeyman to master (reading, ma: addisonwesley, 2000), p. 26-28. ryssevik , j., musgrave, s. 2001. “the social science dream machine,” social science computer review, (v. 19, no. 2, summer 2001), p. 163-174. * paper presented at the iassist conference, amsterdam, may 2001. h. vernon leighton, cultural policy and the arts national data archive, princeton university, vleighto@princeton.edu. mailto:vleighto@princeton.edu f1spe3ts ef dfltq [hfiilflseitieflt by ilona einowski, data archivist state department program university of california, berkeley like the traditional library, the data archive performs many different functions to meet the needs of users. these functions include data acquisition and cleaning, development of conventions and standards for description of the data, data processing and analysis, dessimination of information about the data, storage and maintainance of data tapes, development of an inventory system and inventory controls as well as a data retrieval system, diffusion of the data, training for archive users and program development. data acquisition obtaining new materials for archive holdings from some continuing sources of supply requires establishment of both formal and informal arrangements with institutions, departments or bureaus that produce data on a regular basis in order to obtain some or all of their productions. it is also necessary to establish priorities for the kind of data to be acquired. since the cost of processing and maintaining a data set is often greater than the cost of acquisition, selection must be made with great care. ephemeral or frequently replicated data sets should be acquired only when there is a concrete need for them since there is a high probability of being able to obtain popular data sets elsewhere if a local need develops. the cost of acquiring, cleaning, indexing, and maintaining a data set should be considered in relation to: (a) the likelihood of there being multiple users; (b) the possibility of acquiring at a later date if the need should arise; (c) its availability at a reasonable cost and with little delay from some other source; (d) the amount of overlap with the existing collection; (e) the intrinsic significance of the data. the form in which data arrives varies from supplier to supplier and from study to study. this can result in a great amount of time being spent figuring out just what it is you have received. the ultimate answer is to have funding sources or institutions conducting the survey require that arrangements for archiving be made prior to the actual funding or conducting of the research. this way, the archive can be involved from the beginning and provide guidelines and standards for researchers. a formal way to handle this would be the development of an institutional policy on minimal standards for data to be turned over to the archive. a policy statement of this type would insure that the datasets turned over to the archive meet the criterion of methodological adequacy. it has been the case that an archive decided to pass up a study of great substantive interest which appears to have been done in such a poor manner, utilizing such sloppy and shoddy techniques of data gathering or documentation that, despite the interest of the subject matter, the data set is not worth acquiring. general archive operating policy should include a list of criteria which studies should meet if the data set is to be considered acceptable. at very minimum the following documentation should be available for each study: -continued -continued from page 23 (a) complete and accurate codebook or description of the data structure; (b) description of the data format; (c) illustrations of structure and format; (d) total size of data set; (e) complete and accurate description of the organization of the files for the medium in which the data is stored; (f) precise definition for each data element; (g) complete explanation of all codes used; (h) sample of documents used in data gathering; (i) description of sampling procedures employed, with intended and resultant sample size; (j) summary of training provided fieldworkers and coders; (k) description of data collection procedures; (1) name and current address of study director. data cleaning in its most simplified form, data cleaning involves processes aimed at placing data into a format that is easily handled by computers. these processes include identifying and correcting possible discrepancies between the actual format of the data and the descriptions of that format. many archives employ specialized staff members who do this type of data cleaning. development of conventions & standards in order to facilitate the process of utilizing data initially prepared by others it is necessary to establish conventions for coding and standards for describing the data themselves in order to: (a) permit combining information from different collections in some reasonable way; (b) combine samples from different studies in order to increase the number of cases; (c) make comparisons among data sets; (d) facilitate later analysis. data processing and analysis the data processing and analysis function of the archive provides for the manipulation of data for the user's purpose. this may entail the reformatting of data for use at the user's local facility or providing specially prepared subset of cases or variables rather than a simple copy. some users may need a frequency distribution for the variables (if not provided in the codebook) or simple cross-tabulations. other users may need more detailed statistical analyses. dissemination of documentation the most important documentation produced by the archive is the codebook describing the dataset. archive staff also prepare abstracts of data sets for inclusion in a catalog and for advertising purposes. production of some type of archive catalog is almost mandatory since it provides not only an in-house listing of current holdings but is also the best way for a user to browse the contents of the archive. advertising the availability of the data can take many forms. some archives prepare and distribute their own newsletter announcing new acquisitions while others include a special data announcement section in an existing institution newsletter. archives should also strive to maintain a collection of published material related to the data sets in order to provide examples of how the data have been analyzed already and clarify ambiguities in the interpretation of the data. storeage and maintenance internal procedures must be established to identify the current storage location of all materials. magnetic tapes must be stored in a controlled temperature environment and protected from magnetic flux and iphysical shock. they must be recopied on a periodic basis in order to assure their continued utlity and, where usage is heavy, to protect against deterioration due to machine-induced wear. (see patricia reslcok's article for a detailed discussion of tapes.) inventory as with any collection, it is necessary that the archive maintain a catalog or index of holdings. the archivist might consider maintaining a "public" catalog and an annotated "private" catalog with additional information. the "private" catalog would include abstracts of studies added to the collection since the last published catalog update. it is also necessary to develop an internal inventory system to keep track of the current status of all studies in the archive including studies "on order" or being processed. other internal inventory materials would include a catalog of tapes by tape or storage number and a catalog of studies by study number. retrieval requests from users for access to data relating to their particular topic of interest often requires the archive to search not only its own holdings but those of other archives as well and where necessary, to obtain from other archives those materials required to serve the needs of the user. for this reason it is advisable for the archive to maintain a collection of catalogs from other archives and to become familiar.with the general class of holdings at other archives. most archivists find it helpful to maintain personal contact with other archivists through the network established by professional associations (like lassist) in order to facilitate the exchange of information about data holdings and to keep abreast of technological developments in this field. diffusion the data archive specializes in copying its own collection and making it available to the user at his convenience, in the form most suitable to his purposes. however, the archive must still maintain control over access to their materials in accordance with any wishes of the original donor. for this purpose the archivist usually develops a form letter which the user signs agreeing to archive terms. once the data have been copied, the archivist fills out a standard form to send along with the data tape which describes the files on the tape. ..number of files, logical record length, blocksize, number of records... as well as general tape characteristics. . .tracks, density, format, character set, and internal labeling. shipping data on magnetic tape requires that the tape be adequately packaged to prevent damage in transit and labeled on the outside of the package as being a magnetic tape with a warning to keep the package away from magnets and electric motors which could destroy the data set stored on the tape. tapes sent through the mails are often insured for the cost of recopying the data and have a return receipt included with mailing. training archives can perform a training function by teaching users how to make a query; where to make a query so that the appropriate data can be obtained; how to utilize the data once it is obtained; the devices available for processing; the strategies to be employed for analysis; and the kinds of interpretations that can be made from such analysis. the archivist may wish to prepare a user manual for distribution to potential users documenting how their archive is organized and including information on archive services and locally available -continued on page 29 iassist quarterly 17 historical research and cataloging using a combination of portable trs-80 model 100 and the pick database system on an ibm-xt by david l clark' 24851 piuma road malibu, ca 90265 (818) 888-9305 source mail bbj 949 the purpose of this presentation is to obtain your comments and assistance in devising a standard data entry form which coiild be used by both researchers and libraries for historical and other materials in the social sciences and humanities. the computer offers the potential of greatiy improving the efficiency and usefulness of the cataloging activities of libraries and the research activities of the individual scholar. yet the person or institution seeking to realize the promise and potential of this 'brave new world' of easy information interchange most often receives as a reply that terrible phrase which threatens to become the foremost cliche of our time: "but it's not compatible." the use of a standard form incorporated in a database system would allow both researchers and libraries to take much greater advantage of library resources and of cataloging efforts. for example, when a library sells copies of its historical photographs the institution could also sell copies of the database catalogue entries for those photos, including descriptions, identification of people and places, and library of congress subject terms for database searching. the library could sell the catalogue entries in electronic form, either on diskettes, over the telephone line or by transfer directiy to the memory of a portable computer such as tile radio shack model 100. 'paper presented at the international association for social science information service and technologn (iassist) conference held in marina dei rey, california. may 21-24. 1986 the library could also tap into the expertise of researchers much more easily and productively if the scholar's information and identification of dubious items could be transferred from the researcher's database into the librar\'s. most of a researcher's work never reaches the printed page. it is stored on 3-by-5 cards, handwritten in cryptic note-taking style and virtually beyond retrieval by the person who wrote them two years later. there has been no mechanism for treaung such work as an inventors upon which 10 draw after the current project is finished. computer database management offers such a mechanism. fall 1986 18 iassist quarterly securing the adoption of a common database entry form on a grand scale would probably prove as difficult as persuading all of the nato countries to adopt the same rifle. some will insist upon keeping incompatible weaponry, particularly if it is manufactured locally. however the concept of a common form can prove useful even on a very limited scale, among a group of scholars and/or libraries working in the same field. for example, there are perhaps a dozen libraries with significant holdings of historical photographs of southern california, and several hundred active professional and amateur historians. such a group could exchange the information of greatest interest to it in a common form. the idea would gain impetus when researchers and publishers who purchase photographs begin to prefer one library over another on the basis of its ability to provide electronic copies of database entries for its photographs. the program that i have written uses the model 100 for data entry in the field and the pick operating system on a larger computer for managing and utilizing the database. the radio shack model 100 is ideal for straight-forward data entry in libraries, archives and other field locations. this portable computer weighs only four pounds and can be powered by four penlight batteries. the screen is very readable, and may be used on a full time basis without eyestrain. with 32,000 characters of memory, the machine costs only $450. cost is a particularly important factor when considering the hazards to which a portable machine is subject data entered on the model 100 may be transferred to a small cassette tape recorder (cost: $35) or to a portable disk drive (cost: $200). the model 100 has standard serial and parallel printer outlets. the portable printer supplied by radio shack costs $200 and will print on either plain or thermal paper. the model 100 also includes a modem for transferring information over the telephone, and a bar-code reader outlet accompanying this paper, to illustrate the database entry process on the model 100, are examples from work on a history of ucla and the cataloging of historical photographs from the security pacific collection. i have included a copy of the procedure in the hope of obtaining your suggestions for improvement figures #11 and #12. contain a complete listing of the fields in the database in its present version. i would be grateful if you would, as you read the paper, circle those features, fields and cataloging terms that you would definitely want cross out those for which you would not have any use, and add any which you would wish to see included, and return the listing to me. my mailing address appears at the head of this paper. figure #1 presents the information and choices that first appear on the screen after entering the history database program on the model 100. screen #1 asks whether you wish to add new data, edit old data, print out data in the form of a report or transfer data to another computer or to a tape recorder, disk drive or telephone. the first screen is the "main menu" of the program. you will always return to this point after completing any task. note that the first item on a list is always the "default" choice. you may select that choice simply by pressing the enter key, without typing a number. screen #2 presents the "setup" which was in effect the last lime the program was used. the setup is a combination of the name of the project the name of the writer of the entries, and the type of data to be entered. (the term "data type" may be familiar to you as "relation" or "table".) a file must be created 10 hold your data until it transferred. the history database program creates a name for the file from the one-letter codes for project writer and data type. to those three letters, the program adds day of the month and an "a" fall 1986 iassist quarterly 19 for yor first file cteated on that day. your choices appear at the bottom of the saeen. if you choose "new setup" you will be shown the lists of choices available for projects, writers and data types. screen #3 (figure #2) contains an example of a list of projects. this list is in a text file that can easily be changed. this is also true of the lists of writers, data types, names and sizes of data entry fields, size of the largest data entry form, the requirements of any database program to which you are transferring, as well as many other parameters. all important parameters can be changed without changing the programs. screens #4 and 5 present lists of writers and data types. after choosing the new setup of project, writer and data type, the program will remm to screen #2, the "approve set" saeen, for your fimil approval. the following page (figure #3) presents a guide to the data entry and editing keys used for the history database program on the model 100. the purpose of this program is to support research which may involve long textual entries rather than short items such as would appear an address list therefore full editing capabilities are provided. screens #6, #7, and #8 (figures #4, #6, and #8) present data entry forms for data types photo caption, text and bibliography. each form encompasses one item. an item, sometimes called a record, is a collection of fields (or attributes). note thai the model 100 saeen displays 8 rows of 40 characters each; therefore the data entry screens as presented are divided into multiple screens on the model 100. as you enter each screen, the program presents the entries most recently made for thai setup of project, writer and data type. mosi data enir> work involves a great deal of repetition. the program allows you to repeal automatically any information which remains the same from one entry item (or record) to the nexl you can accept an entire field for automatic repetition, or edit the information in the field and subfields. the very long descriptive field which appears at the bottom of the photo caption and text entry forms is not repeated, it is left blank for new information. the field entries available for automatic repetition are taken from the most recent entries, even if the computer has been turned off or used for other tasks in the meantime. the program will always return to the point at which you left il thus, if you are resuming data entry after an interruption, you can return to your entr>' form with only two keystrokes: press enter on screen #1 for the first choice of add new data, then press enter on screen #2 to approve the same setup and data file name, and the program will present your previous field entries for approval or change. the first field, "temporary", (see figure #4) is a place marker. the program will automatically enter the word "new" in that field. when the file is transferred to the database program on an ibm-xt, the larger database program will automatically present for your approval the items which contain the word "new" in the temporary field. to approve the item for admission into the database, remove the word "new". when working with the database, the temporar>' field serves as a place marker, similar to a bookmark. for example, when selecting photo captions for printing, individual items might be coded "yes" or "maybe" in the temporar\field. compare the data entr>form to the following page (figure #5), which was printed with the model 100 and the portable printer. note thai the temporary field has not been printed, since it is only for internal use. the fields bibliography*. pari#, and subpari# correspond to entries in the bibliography and will be used to identify the sources of the photographs, the collections and libraries in which they are contained, and their fall 1986 20 iassist quarterly fonnat, without the need to type and store such repetitious information separately for each item. in the example (figure #4), the photo was part of the security pacific collection which is #1 on the bibliography list, the format of the photo is 5x7, and the subpart# field is not used, so a "1" has been entered as a dummy. the entry# designates the data item within a given collection. when comparing the entry form to the printed version, you will see that the program has combined bibliography*, part#, subpart# and entry* into an item id, separating the components with asterisks, and has made that item id the first field. the item id uniquely identifies an item within the database. unique identifiers are necessary for many database programs and desirable for others. when the file is transferred into a larger computer's database, item id can be handled with some fiexibility, depending on the requirements of the larger computer's program. the field ref# is intended to contain any identifying number that may already been assigned to the photo or other material thai you are describing. for example, when taking notes from a book, you might use the page number. the field "units" is intended for collections. a single photo description might apply to more than one photograph. in the example, rather than entering 17 separate items for 17 photos of the construction of the samson tire and rubber plant, all 17 have been included in one item. a library patron ordering a photograph would specify the item id and the individual number, for example "r57*l*1845 #13". after the dates on which the photos were taken, appear the fields "eratype" followed by "era". in the same manner, the field "location" is followed by "areatype", then by "area". in the printed report the history database program pairs the "type" field entries their counterparts in the descriptive fields. thus: "eratype: decade era: 1920; 1930" on the entry form has become "decade: 1920:1930" in the printed report, and "areatype: city; county area: city of commerce; los angeles" has become "city: city of commerce county: los angeles". the format of type-description pairs allows the accommodation of new and unexpected needs without requiring changes in the data structure. for example, "type" of geographic area might be a city, coimty, state, or other entity. in los angeles, most of us refer to ourselves as living in communities, such as hollywood or venice, which have no incorporated legal status. for some research purposes, the relevant area might be the congressional district or the census district for other purposes one might speak of an irrigation district, a federal coun district, or a parish. the variety of such designations defy all attempts to lock a particular term into a database within a reasonable total number of fields. therefore the type-description pair allows you to designate the type of unit that is relevant to your purpose while maintaining a structured format one could not simply have an area field and write in "los angeles," because no one would know whether you meant the city, the county or the state of mind. in addition to allowing you to designate the relevant "type" of imit the type-description pairs also allow multiple entries, such as "community: venice" followed by "city: los angeles". multiple entries in a single field are refened to as multi-values or sub-fields. in the example, the history database program has separated and paired off the sub-fields to produce a clear printed report a prime reason for choosing the pick system to manage the database after entry onto the ibm-xt is pick's handling of sub-fields. pick can be used to create "controlling dependent" relationships between fields and between their conesponding sub-fields. thus, pick can be used to select hems in the database from the city of los angeles but not including everything from the entire county of los angeles, even though there are no fixed city and county fields defined. you can designate: if city = "los angeles" fall 1986 iassist quarterly 21 when conducting a search and retrieve only those items in which you specified the correponding areatype as "city". pick will also treat the sub-fields as separate entities for global search and replace and for conepondence tables. when pick stores an item, the pick system keeps the fields separate and within each field also keeps the sub-fields separate. thus, if the library of congress changes a standard subject term from "agricultural utensils" to "farm tools" you can make that change throughout your database with one command, whether or not your subjects field contains multiple entries. other database programs, such as rbase, will treat the entire subjects field as one entr)-, even though you may have placed multiple subjects there. thus with rbase, the global search and replace command is only useful for single-entr>' fields. free text systems, which some researchers and hbraries have adopted in the mistaken belief that such programs are simpler to learn than database programs, will not perform global search and replace at all. to change "agricultural utensils" to "farm tools", you would have to make the change in each individual item. one alternative is to load the entire database into a word processor, and make the change with the word processor's replace command, but a word processor does not know one field from another, and would make other, unexpected, changes. the cataloging of materials, whether by a researcher or a library, is ver>' straight-forward database usage. to adopt a free text program for such a purpose would be inefficient the history database program which 1 have placed on the model 100 includes features to facilitate a changeover from free text to database, or from one database program such as rbase to pick. thus, after a data item has been created, during the transfer process, one can choose to send onh certain fields to one program and other fields to another. continuing with the entry form (figure #4), the field project indicates the particular project for which the research was conducted. in an institutional context, the entry might be the name of a department or funding source. the following two fields, filehead and topics, are dependent upon the field i*roject in the sort of controlling-dependent relationship mentioned earlier. in my example, the project is "security pacific" for the cataloging of the collection of security pacific photographs. the filehead "factories-rubber industry and trade" designates the label on the file folder in which the photographs have been placed within the security pacific collection and would not be confused with a similar file heading that might be used in the business department of the librar>'. a researcher might enter into the project field the title of a book on which he was working, and filehead could be used to indicate the chapter in which he expected to place the information. thus in another example, i named the project "ucla history", to indicate a book i am struggling to finish; the filehead is "graduate school founding-opposition by the regents" to indicate the relevant chapter and subchapter. a scholar might be working on several different books at the same time, and could indicate the relevance of the information for different chapters in different books. a library might also be combining funding from several sources or wish to designate where a photo might be used in an exhibit in addition to its placement within a collection. thus to the project field could be added the name of an upcomg exhibit such as "los angeles at work" and to the filehead field one could add, in a conesponding position, "factory architecture". the field topics contains subject headings which pertain only to a given project, and should not be confused with the standard library of congress subject headings in our subjects field. since standard subject terms are preferable, the use of the topics fields should be limited as much as possible. fall 1986 22 iassist quarterly because the pick system allows us to aeate conespondence tables, even when a field contains multiple entries, we can make qui work much faster and productive in several ways. for example, the researcher can indicate the order in which he expects the chapters and subchapters to appear. the table consists simply of his book outline, with rows of chapter names in one column and corresponding numbers in another. when the researcher is ready to write a chapter, he can bring the information from his database into the word processor sorted in the order of his outline. if he wishes to change the order, he need only change the outline. the program will also deliver the information with a footnote attached. to construct the foomote. pick will use information regarding sources from the bibliography, for which reason the text or photo caption item contains the bibliography# of the corresponding item in the bibliography lisl just as a researcher retrieves information for his book, in the same manner a library can easily give its staff a list of photos to be retrieved in the order in which they are filed, even when different ordering systems have been used for prints and negatives. to change the order on the computer, one need only change the table. a conespondence table can also be used to check the terms that are entered. thus, if you type a file heading or chapter name that does not appear on the approved list for the filehead field, the program will return it as an incorrect entry. this prevents the typing enors and other mistakes that plague cataloging. a single mis-typed letter will often place an item beyond retrieval. lists can be created with the name of each allowable project, filehead, topic, eratype, areatype, etc. pick also contains a feature called "soundex" which allows one to search for variations of names, as long as the sound of the name remains similar. the correspondence table used for error checking can also be used for abbreviations. enclosed are examples of each table. rather than writing "census district" as the areatype, one could type "cs". whether you typed "cs" or "census district" the program would lake up only the storage space required for "cs". when the data are printed or displayed on the screen, the program automatically converts "cs" to "census district". you could search for the areatype as either "cs" or "census district". the shorter format would be easier for someone entering tnany items, the longer format might be easier for someone who did only occasional searching. the pick system provides extensive facilities for such conversions between internal and external formats. pick can also create virtual fields or symbolic fields which contain no data of their own, but which display the results of manipulating other fields. thus although the photo caption entry contains no designations for source or collection, one could search for such fields, because pick would take the information from the bibliography, just as it did when constructing the footnote mentioned above. to finish the photo caption form, we enter under names the names of relevant individuals and companies, and under subjects, standard librarx' of congress headings. it should be stressed that the use of standard subject terms is absolutely essential. it is easy to believe, in the beginning, that you can simply create your own terms as you proceed, but you will soon find that you have used different terms at different times for the same purpose, making searching very difficult the final field. caption, contains a description of the photographs. you will notice that the description here of the building of the "assyrian rubber factory" is written in draft manuscript form, not in cryptic note-taking style. when working in a library or archives with the history database program on the fall 1986 iassist quarterly 23 model 100, it is actually easier and faster to type complete sentences than it was previously to scribble notes on 3-by-5 cards. when the information is retrieved a year later for use in a book or article, rather than hand-written notes which even the writer can barely decipher, one has instead a rough draft manuscript chapter, with foomotes. the next example (figure #6) is of the data type called "text" which is used for textual materials, such as books and collections of private papers. the entr>' form is identical to that used for photographs, except that the last field is labelled text rather than caption. the last field again contains the main descriptive entry, which in the example (figure #7) is a description of the struggle in the 1930s to found a graduate school at ucla. the description is wrinen in manuscript style, but took no more time to write than did scribbled notes previously. the final examples (figures #8, #9, and #10) of the bibliographic entries which support the photo caption and text entries shown earlier. the bibliographic material can be long and detailed, because it is only entered once, rather than repeated in each text or photo item. the top half of the bibliography form also includes location, project subjects, etc. please note that there is very litde penalty under pick for fields that often not used. pick uses variable-length fields: it will lake up room only for the data that you enter, plus a single character used to separate one field from another. a fixed-field system such as rbase requires that you specify the length of each field, and each field will take the predetermined amount of space whether or not there is anything entered. one rationale in the past for the adoption of free text programs has been that they often permit variable length fields. however, they do not allow you to do anything with the data once it is entered, other than to change it item by item. free text programs use the computer as a glorified typewriter and filing cabinet in addition, a free text program such as inmagic will require several times more space for field indexing than a fixed-field program such as rbase would have used. inmagic will not search fields that have not been indexed, so in effect indexing is required. pick does not require indexing, it does not require fixed-length fields, and it allows full manipulation of your data. the bottom half of the bibliography form (figure #8) contains the fields sourcetype, source and coltype, collection. these are type-description pairs similar to eratype and areatvpe shown eariier. thus you could indicate that the "tn^e" of source involved was an author, a photographer, a person who was interviewed, a correspondent, or an organization such as the census bureau or general motors. the "type" of collection could be archives or a collection of papers. the collection field indicates the relationship between an item of information and some larger body of information in which the item is contained. defining that relationship may call for some creative thinking, because the individual photo or piece of data will often be contained in other entities which are in turn contained in other collections and so on, several levels deep. the t>t3e-description pair will allow you to include all the levels, and indicate just what sort of object is involved in each case. the category field might indicate photographs, published materials, unpublished papers and ephemera, to provide separate listings in the bibliography of a book, or for a library organizing its materials. the format field will be most useful!) applied to photographs or other materials whose format determines physical storage location. if desired, the format from the bibliography can be used automatically as the pari# in the itemld. the note field contains extra remarks. publisher, place and fall 1986 24 — iassist quarterly volume. library and call# would help you find the photo, book or journal again. the last field contains any credit line that might be required for use of a photograph or other material. after the material is entered, the remaining saeens on the model 100 will guide you through the process of re-editing the data entries, printing the data, or transferring it to a larger computer. when the program transfers a file, it uses parameters previously indicated as required by the recipient database program. thus the same data can be transferred to a variety of different programs in a compatible matter. in the same manner, a library offering information in electronic form could use the pick system to export the information in the formal required by the researcher. thus it is not necessary to achieve absolute conformity in order to have compatibility, but it will be necessary to bring to researchers and librarians an understanding of the advantages of entering all information from the first step in electronic form, such as is made possible with the model 100. and then transferring that information into a powerful database management system such as pick.n fall 1986 iassist quarterly 25 fig. #1 1 f .........„„„.. ^ screen # 1 jj history database 1 data entry program for the model 100 1 copyright david l.clark 1986 i 2485 1 piuma rd maubu, calif 90265 | 1 (818) 888-9305 | l.add new data 3.print report 4.transfer file 5.end program pick a number [1]: ^ j r \ approve setup data type = p: photo caption data file = scp23a l.ok 2. change file 3. new setup 4. quit pick a number [1] screen # 2 fall 1986 26 iassist quarterly fig #2 screen # 3 screen # 4 screen # 5 r projects 1. a: avery international history 2. l: los angeles history 3. s: security pacihc photos 4. u: ucla history 4. x: not on list pick a number [1]: writers l c: qark, david l. 2. s: stone, brian 3. x: not on list pick a number [1]: data types l.t:text 2. p: photo caption 3. b: bibliography 4. r: reference 5. v: vocabulary pick a number [1]: y "^ fall 1986 iassist quarterly 27 fig. #3 history database data entry program for the model 100 copyright david l. clark 1986 24851 piuma rd malibu, calif 90265 (818) 888-9305 data entry news press key: to: enter send entry to the computer, move to the next field escape quit delete delete 1 character backspace backspace and erase 1 character tab move right 8 characters arrows: left left 1 character right right 1 character up up 1 line within field down down 1 line within field shlft-left left 8 characters shift-right right 8 characters shift-up previous field shift-down next field control-left start of field control-right end of field control-up previous page control-down next page fl not used on model 100 (on ibm-xt displays a help screen) f2 clear the field from the cursor f3 delete word f4 delete sentence f5 toggle between insert and over-type modes f6 delete sub-field f7 set new itemid f8 quit control-p power off, resume at that point when machine turned on fall 1986 28 iassist quarterly fig. #4 screen #6 data entry form: photo caption temporary:new bibliography#:1 part#:57 subpart#:1 entry#:1845 entrydate:05/22/86 researcher:clark, david l. ref#:lot 955 units:1 7 date:05/29/1 929; 1 2/26/1 929 eratype:decade era:1 920; 1930 location:5725 telegraph rd. area type:city; county area:city of commerce; los angeles project: security pacific filehead:factorles-rubber industry and trade topics: names:samson tire and rubber corp.; u.s. rubber corp. subjects:factories-design and construction; rubber industry and trade; tire industry; architecture, assyrian; wit and humor; constmction caption :constnjction of the samson tire and rubber factory. because the company name was "samson" a babylonian style was chosen for the building, giving it the popular name of the "assyrian tire factory." photos #1 -6 the start of constnjction, 05/29/1 929. photos #7-1 7 later construction on 12/26/1929 when the assyrian appearance had become evident. samson was a subsidiary of u.s. rubber. branch plants built by national mbber companies in la. in the 20s and 30s made the area 2nd to akron in rubber. fall j986 iassist quarterly 29 fig. #5 david l. clark 05/22/86 1 5:49 security pacific photos photo caption ltemld:1*57*1*1845 bibliography#;1 part#:57 subpart#:1 entry#:1845 entrydate:05/22/86 researcher:clark, david l. ref#:lot955 unlts:17 date :05/29/1 929; 12/26/1929 decade: 1920; 1930 location:5725 telegraph rd. city:city of commerce county: los angeles project:security pacific filehead:factories-rubber industry and trade topics: names:samson tire and rubber corp. ; u.s. rubber corp. subjects:factories-design and construction; rubber industry and trade; tire industry; architecture, assyrian; wit and humor; construction caption :constnjction of the samson tire and rubber factory. because the company name was "samson" a babylonian style was chosen for the building, giving it the popular name of the "assyrian tire factory." photos #1-6 the start of constmction, 05/29/1 929. photos #7-1 7 later construction on 12/26/1929 when the assyrian appearance had become evident. samson was a subsidiary of u.s. rubber. branch plants built by national rubber companies in l.a. in the 20s and 30s made the area 2nd to akron in njbber. fall 1986 30 iassist quarterly fig. #6 screen # 7 data entry form: text temporary: bibliography* part# subpart# entry# entrydate: ref# researcher: units: date: eratype: location: era areatype: area: filehead: topics: names: subjects: text: fa/1 1986 iassist quarterly 31 fig. #7 david l. clark 05/22/86 22:51 ucla history text ltemld:6*20*10*126 bibliography#:6 part#:20 subpart#:10 emry#:126 entrydate :05/22/86 researcher:clark, david l ref#:page1 units: date:06/12/1933 decade:1930 location: state:california project:ucla history filehead:graduate school founding topics: names:ucla graduate school; earl, guy c. ; regents of the university of california subjects:universities and colleges-graduate work; regionalism text:as the board of regents of the university of california continued to refuse to allow graduate work to begin at ucla, edward dickson, the regent most active in promoting ucla, expressed his reaction in a letter to his friend and fellow regent guy earl. dickson portrayed all of southern california as up in arms over the "outrage" perpetrated of the southland by the berkeley-dominated board, "the people here are indignant. there is growing resentment. they were led to believe that at last the stigma that attaches to the university here as not being able to undertake graduate work was to be removed... now we may look forward to years of warfare." fall 1986 32 iassist quarterly fig. #8 sreen #8 data entry form: bibliography temporary: bibliography* entrydate: ref# eratype: location: areatype: area: filehead: topics: names: subjects: sourcetype: source: coltype: collection: title: part# subpart# researcher: units: date: era: entry# category: format: note: publisher: place: volume library: call# credit fall 1986 iassist quarterly 33 fig. #9 david l. clark 05/22/86 1 5:53 security pacific photos bibliography !temld:r57*1*1 bibliography#:1 part#:57 subpart#:1 entry#:1 entrydate:05/21/86 researcher:clark, david l ref#: units:1 00000 date: decade: 1920; 1930 location: state :california region:southern california project:neh-security pacific filehead:caiifornia-history topics: names:los angeles chamber of commerce subjects :california-history; economic history; industry; factories organization :los angeles chamber of commerce donor: security pacific national bank collection :security pacific title: category :photos format:5x7 bw note: publisher: place: volume: library:l.a. public library call#:history dept. credit:l.a. public library/security pacific photo collection fall j 986 34 iassisl quarterly fig. #10 david l. clark 05/22/86 23:04 ucla history bibliography ltemld:6*20*10*1 bibliography#:6 part#:20 subpart:#10 entry#:1 entrydate :05/22/86 researcher:clark, david l ref#: units: date: decade:1910; 1920; 1930; 1940 location: state:california project:ucla history filehead: topics: names:dickson, edward a.; ucla subjects :universities and colleges source:dickson, edward a. collection :dickson, edward a., private paper title: category:unpublished papers format: note:the papers are stored in 27 boxes. the part# refers to the box, the subpart# indicates the folder. publisher: place: volume: library:ucla special collections call#:dickson papers credit: fall 1986 iassisl quarterly 35 fig. #11 history database codes david l. clark 05/08/86 page 1 circle the items that you would definitely want included. cross out those items that you would have no use for. add any new items that you would want. source: ac = architect ag k= agency at = artist au = author bd = builder bu = bureau cp = company cr = correspondent ct = collector de = department dv = division do = donor ed = editor gr = group in = interviewee ir = interviewer iv = inventor mf = manufacturer of = office or = organization ow = owner pa = painter ph = photographer rc= recipient sc = sculptor se = section so = source wr = writer collection; av = archives co = collection fi = file ne = newspaper jo = joumal si = series era: ce = century de = decade dm = dayofmonth dw = dayofweek er = era he = halfcentury mo = month sa = season sf = shift tm = time we = week ye = year area: ad = assemblydistrict ao = arrondisement ar = area as = assessmentdistrict bo = borough cc = councildistrict cd = congressionaldistrict ce = commune ci = city cm = community cs = censusdistrict cu = county de = department di = district fall 1986 36 iassist quarterly fig. #12 story database codes david l. clark 05/08/86 page 2 ds = diocese ed = enumerationdistrict fc = federalcourtdistrict id = lnigationdistrict id = legislativedlstrict ma = map nx = mapcoordinates md = marketdistrict mn = mapname mp = mappage ms = mapsection na = nation nd = nationalassemblydistrict ne = neighborhood pa = parish pc = precinct pk = pa(1< pr = province re = region sa = state/\ssembiydistrict sd = senatorialdistrict se = section sh = schooldistrict sp = supervisorialdistrict ss = statesenatorialdistrict st = state su = suburb te = territory to = town vi = village wa = ward wd = waterdistrict zc = zipcode fall 1986 iassisl quarterly 37 fig. #13 david l. clark 05/22/86 1 5:49 security pacific photos photo caption ltemld:r57*ri829 bibliography#:1 part#:57 subpart#:1 entry#:1829 entrydate :05/1 5/86 researcher:clark, david l. ref#:lot955 units:9 date:1 2/31/1 927 decade:1 920; 1930 location:1200 no. main st. city:los angeles county:los angeles project:neh-security pacific filehead:factories-steel industry and trade topics: names:llewellyn iron works; union iron works; columbia steel corp. subjects:factories; steel industry and trade; iron industry and trade; metal trade; boilers caption:the llewellyn iron works, which in 1 930 was absorbed into the union iron works. photo #1 administration building. #2 machine shop. #33-4 swinging fabricated steel trestles on an overhead track to the paint shop. #5 a 32 ton housing cast by the columbia steel corp. of torrance and machined by llewellyn. #6 an hydraulic riveter. #7 the forge shop making a huge drive shaft for steamships. #8 the vertical boring mill. #9 making boilers. fall 1986 iassist quarterly 2014 5 iassist quarterly editor’s notes data, the whole data, and nothing but the data … and the metadata, and the access to data welcome to the third issue of volume 38 of the iassist quarterly (iq 38:3, 2014). this issue is unquestionably about data. there are three papers on projects for improving delivery of data to users. the first paper is ‘distributing access to data, not data’ by david schiller from the institute for employment research (iab) at nuremberg (germany) and richard welpton at uk data archive, university of essex (uk). they focus on the problem that access to european microdata for researchers is restricted by national borders and the barriers for performing comparative analyses between the member states. the ‘data without boundaries’ project now has an initiative to build a ‘european remote access network’ (euran). the problem is that prevention of identifying respondents in the microdata conflicts with the importance for modern research methods of access to detailed data. some control is necessary and the paper describes remote access as the appropriate answer in the forms of job submission, remote execution, and remote desktop. as an example, one version of secure remote desktop access encrypts pictures of the desktop screens to make secure the transport over the internet. the authors reference a set of principles for access, e.g., that it is not desirable to physically move data and that access should come through a single point that can access multiple sources of data. the researchers’ need to analyse the data is supported by a ‘virtual research environment’ that includes software for generating and presenting results through the euran project. the next paper presents a two-year metadata project based upon two well-known series of studies: the american national election study (anes) and the us general social survey (gss). the goal is to improve their metadata and build demonstration tools to illustrate the value of structured, machine-actionable metadata as reported in ‘creating rich, structured metadata: lessons learned in the metadata portal project’. the authors are mary vardigan (inter-university consortium for political and social research (icpsr)), darrell donakowski (american national election studies (anes), university of michigan), pascal heus (metadata technology north america (mtna)), sanda ionescu (icpsr), and julia rotondo (norc at university of chicago). the article reports on their experiences, and also includes recommendations. the national science foundation funded the project under the ‘metadata for long-standing large-scale social science surveys’ (meta-sss) program. icpsr and anes are co-distributors of most of the anes studies while the gss is co-distributed by norc, the roper center, and icpsr. in the project metadata tools revealed small differences between supposed identical datasets, for instance in study titles, variable names, etc. the project also decided which types of content to include. both of the the series are huge collections as the 58 anes surveys contain 79,521 variables and the cumulative gss has 5,558 variables. marking up this legacy documentation is laborious and time-intensive and the future naturally lies in capturing the metadata at the source. in conclusion, the project learned a great deal about converting legacy documentation and identified several steps for documentation development, including the areas of paradata and versions of datasets. the concept of versions of datasets relates to the solution described in the first paper of not bringing data but access to data to the users. the third paper demonstrates further work in the project described above. in the paper ‘mapping the general social survey to the generic statistical business process model: norc’s experience’ the three authors scot ausborn, julia rotondo, and tim mulcahy – all from norc at the university of chicago present how they carried out the mapping of the gss workflow to the generic statistical business process model (gsbpm). an analysis of the business processes for the production of survey data was carried out with the intention of direct capture of survey cycle ddi-based metadata, thus avoiding the need to generate it retroactively. the work is based upon an internal survey of gss staff, asking them to explicate their respective roles on the survey in terms of the gsbpm. connecting aspects of the gss workflow to elements of the gsbpm produced a comprehensive and integrative view of the individual efforts that together produce the survey. of the lessons learned, i noticed that they later found that it may have been more fruitful to have held a workshop in which gss staff could discuss the workflow processes together, rather than having a survey with each person providing his or her input in isolation. they mention that they think an expert in gsbpm could have conducted the mapping of the workflow; however they did identify points for improvement in the workflow relating to both metadata and paradata. articles for the iassist quarterly are always very welcome. they can be papers from iassist conferences or other conferences and workshops, from local presentations or papers especially written for the iq. when you are preparing a presentation, give a thought to turning your one-time presentation into a lasting contribution to continuing development. as an author you are permitted ‘deep links’ where you link directly to your paper published in the iq. chairing a conference session with the purpose of aggregating and integrating papers for a special issue iq is also much appreciated as the information reaches many more people than the session participants, and will be readily available on the iassist website at http://www. iassistdata.org. authors are very welcome to take a look at the instructions and layout: http://iassistdata.org/iq/instructions-authors authors can also contact me via e-mail: kbr@sam.sdu.dk. should you be interested in compiling a special issue for the iq as guest editor(s) i will also be delighted to hear from you. karsten boye rasmussen march 2015 editor http://www.iassistdata.org http://www.iassistdata.org http://iassistdata.org/iq/instructions mailto:kbr@sam.sdu.dk vol21.3 32 iassist quarterly abstract increasingly information is available as images: colour, black and white from satellites, photography, scanning in the biomedical sciences or scanning of printed sources, like codebooks accompanying numeric data and historic records. these images have very specific computer formats but need to be easily embedded, identified, searched, transferred and browsed or reproduced to paper to make them useful as information carriers. principal media for distribution and access are cdrom and internet. briefly formats, standards and applications are discussed to indicate how much actual progress is being made in putting image information into the hands of researchers and interested users. for the last five years and not only in the social sciences but in many areas of research, in cd-rom and internet publishing and in library services, information is increasingly available as images. there are several reasons for this increase: 1. more applications that create images : data visualisation, analytic mapping, flow charting to represent questionnaire routing, gis systems, convenient document delivery 2. more applications that need images as content matter: computer assisted learning, distance learning, multimedia 3. increased transformation feasibility from other media (like printed or hand written sources) to images: including image enhancement and manipulation to correct imperfect originals 4. new types of hardware that can produce images: digitising video sources , digital photocameras, digitising to images in biomedical research and health care, remote sensing, forensic applications 5. new possibilities for distribution of series of images: improving internet transfer capacity, new multi-page image formats that fit better with internet client server approaches, easy creation of single or small series of cd-rom’s that can hold large numbers of images 6. new approaches that explore the next step from preservation to presentation by bringing images into structures like sgml and html designs or by betting on new formats like pdf 7. better availability of applications that also in the public domain view, browse series, and print images, including reproducing to colour on a low end desktop with colour printing ; new (pdf) reader software that can “search” images of textual information with help of shadow pages that hold text approximations of the original after “dirty ocr” 8. recognition of images as another electronic source of information that needs appropriate bibliographic description and identification in summary: more and different types of information are provided as images and at the same time there has been a shift in emphasis from long term storage with exact replication of the original and efficient compression towards giving access : finding ways to locate, search, browse and navigate the information content of images. in particular a format like tiff represents that longer established practice of preservation : it handles virtually any image original level of detail, colour schemes (including black and white) and the like has different compression choices (including a lossy one), is computer platform independent, is recognised as a standard and by consequence a format of choice for scanning stations and almost any application that one way or the other converts, transfers, shows, does ocr or reproduces images to print. in addition to long term storage it is efficient in simple distribution: like ftp, cd-rom’s with series of tiff images and document delivery of scanned journals. internet and tiff though proved difficult to match: the format has the multi-page feature but this allows by it self only a linear forward and backward browsing and tiff cannot be transferred to internet client software with features like an increasing level of detail or with “page on demand” . always a complete (multi-page) tiff with all original detail, has to arrive over the net before the image information can be viewed or printed. a possibly information that come as images: overview of issues by repke de vries * fall 1997 33 very time consuming exercise. 1 in terms of file formats and compression schemes did the internet and the relatively slow speed of parts of the electronic connection between user and provider, sparse improvements to bring down transfer time : jpeg for example reduces original image file size drastically but is a lossy technique: after decompressing there are differences with the original ; progressive jpeg and interleaved gif are web developments that help transfer by first presenting a rough impression to the user and gradually completing the image. 2 this development forces archival organisations that both have to preserve an image collection and are information provider with that very same collection, to keep their images in two formats and face a dilemma. a preservation choice should be tiff, probably off-line on cd-rom’s. but a different format would be needed to suit internet presentation and electronic publishing. though bringing down transfer time is needed where internet speed is not ideal , the real issue of access is a different one: 1. finding approaches to bring images into a structure together with other information that can have arbitrary other formats, like ascii text, raw data files or even sound samples when biological research has habitat photos of birds together with recordings of their calls and field notes. such a structure would also allow browsing and navigating with hyper-links and not just linear as in the single multi-page image file 2. finding ways to examine the information contained in images: scanned text can have a rough searchable real text duplicate after “dirty ocr” , other images (like photographs or historical sources) can have meta-data attached that allow for searching and locating. this meta-data would take the usual form of keywording, applying subject classifications or standard bibliographic reference. the two approaches go particularly well together when they complement each other : access to an image collection or single image after searching meta-data, with the images subsequently structured in a way that puts them into context with other available information and permits easy walk-abouts and retrieval. the main choices to realise this kind of access seem to be the following: for structure: 1. sgml and html as sgml applied to the internet pro’s: both meta-data and any further information type besides images, can be brought into this structure and at the same time keep their own specific (technical) format; it has hyper-linking; it is a general standard with good availability of software to code into this structure and to decode it; con’s: html web browsers which are themselves in the public domain may need additional commercial software to view and manipulate particular image-formats ; this forces the provider to choose a format that is either supported natively by the common web browsers or for which the additionally needed image-viewer comes free; the same software arguing holds for sgml • the data documentation initiative is an sgml structure initiative that explicitely defines the possibility to bring needed information into sgml in wahtever format it comes: also questionnaire pages scanned to images (url: http://www.icpsr.umich.edu/ ddi/ and the may 1996 paper by k rasmussen “convergence of meta data. the development of standards for social science data” (contact the author at boye@get2net.dk). 2. pdf with its additional hyper-linking and annotation features pro’s: public domain availability of pdf reader software; the reader software has all navigating, browsing, searching and image-viewing together in one application ; it has hyper-linking con’s: the structure is limited to images and text information (no raw data for example) that both first have to loose their original format and have to be imported into pdf with commercial software; while html and sgml know many applications that decode this structure and explore its information content, has complete pdf with mixed image/text content, searching and hyper-linking only the (free) adobe reader software for technical access • adobe, the initiator of pdf, can be found at url: http://www.adobe.com/ • “internet publishing with acrobat” by g kent , published by adobe press (1996) discusses “..creating and integrating pdf files with html on the internet ..” and marks adobe’s move to adapt pdf to internet presentation 3. a sgml and internet html approach supported by pdf, where pdf only holds the images but the (free) http://www.icpsr.umich.edu/ddi http://www.icpsr.umich.edu/ddi mailto:jtstratford@ucdavis.edu http://www.adobe.com/ 34 iassist quarterly adobe reader software brings sophisticated viewing and manipulation , certainly after the recent adaptation of pdf to internet requirements like “image on request” from a multi-page file. • the psid internet site can serve as an example: url: http://www.isr.umich.edu/src/psid/pdf.html it has to be noted that contrary to the above, the psid site and many other providers use pdf only as a distribution format or document delivery mechanism, which was made attractive by adobe putting pdf reader (and printing) software in the public domain. none of the expanded features of pdf are utilised. for searching information hidden in images: 1. database approaches several scenario’s exist: descriptive information is indexed and searched, with the images as search result; • wais makes this possible for the internet outcomes of “dirty ocr” on images from scanned documentation, are indexed and searched; a further linking scheme has to connect with the right image free-text indexing and searching, where the retrieved text has graphic-annotation links to images; • isys database software and the international social survey program cdrom are an example. url: http://www.za.uni-koeln.de/data/en/issp/index.htm pro’s : -the database or search engine approach quickly brings results, like any indexed free-text search con’s : for the same reason is precision in search outcome not always high ; additional effort is needed to create descriptive and image linking information, or with text from some source available at least the links to images ( like in codebook text and scanned questionnaire examples); 2. subject classifications and keywording • the cass question bank on the world wide web illustrates this: url: http://kennedy.soc.surrey.ac.uk/qb/ welcome.html it has in fact both different subject trees and a search engine. pro’s : a very precise search that takes time when browsing subject trees but can still be fast when assisted by a thesaurus approach con’s: even more than in rough free-text indexing, is human effort needed to apply keywords or classify each set of images (like a questionnaire in the cass example) 3. applying the commercial adobe acrobat software possibilities to expand the pdf format with searchable text that can be produced with internal “dirty ocr” . this search is linear and images that have a non-text content would obviously need other acrobat means to pinpoint particular images.: annotations and hyper-linking. the hyper-linking to relevant images within the single pdffile can interestingly enough be realised with the idea of a “clickable map” : particular clickable areas of an image, like a scanned table of contents or an overview picture of the human body, can be hot-linked to where further image information starts. this can create a very intuitive guidance in searching. • a very recent cdrom produced by the german zentralarchiv and the dutch steinmetz archive holding part of the eurobarometer questionnaires in the original languages, uses this type of clickable hyper-linking • icpsr produces cd-rom’s that have the search facility in pdf after internal “dirty ocr” : for example the cd0013 “health and well-being of older adults” prepared by nacda (url: http:// www.icpsr.umich.edu/nacda) this way both the original codebook page is available in pdf as image and for searching the same format holds a text equivalent where possible. pro’s : creating the descriptive information to search on is part of the adobe software the idea of “click and go” con’s: the search is slow ; non-text images cannot do with automated “dirty ocr” and need complicated, manually added hyperlinking and annotations ; clickable links need human effort to organise and implement ; the extended pdf format with the above mentioned features, probably has to be regarded proprietary and would always need the hitherto free adobe reader software conclusion. the newer, extended features of pdf are too recent to have had much evaluation . many services and products have been realised in pdf but predominantly as carrier for document delivery and distribution following the bringing in the public domain by adobe of the pdf reader software. sgml (html) seems to offer the better, more general http://www.icpsr.umich.edu/nacda http://www.icpsr.umich.edu/nacda fall 1997 35 structure for access, which structure can also hold the descriptive information , keywords or meta-data that a search engine could take for indexing to help locate particular images. html web applications have brought considerable illustration of giving access to information as images but not in any complicated way, other sgml examples like the ddi still have to make their point. this overview makes a distinction between preservation and presentation and has focused on the latter. it is an exciting new area of services, still open ended in direction but moving forward quickly. * paper presented at iassist/ifdo ‘97, odense, denmark, may 6-9,1997. r de vries, steinmetz archive, netherlands email: repke.de.vries@niwi.knaw.nl 1 the netscape web site “inline plug ins: image viewers” nevertheless feature several tiff applications 2 also the png image format is an interesting new development: url: http://www.w3.org/pub/graphics/ png/ mailto:repke.de.vries@niwi.knaw.nl http://www.w3.org/pub/graphics/png/ http://www.w3.org/pub/graphics/png/ iassist quarterly 2010 / 2011 77 iassist quarterly abstract at the time of the bremen workshop in 2009, there was no swiss institution responsible for archiving qualitative social science data collected in switzerland. since that time, the swiss institution fors (swiss centre of expertise in the social sciences) has assumed this role, with development of the infrastructure, policy, and know-how needed for implementation. the archiving of qualitative data at fors is now moving forward in close collaboration with swiss universities with established and strong qualitative research traditions. over time, the success of qualitative data archiving at fors will require the support of research funding institutions and policymakers, enhanced educational initiatives at universities, promotion of the value and potential of secondary data analysis, and a well-equipped and staffed archive that serves also as a network node and resource centre. keywords: secondary use; archives; data; qualitative research; research infrastructure; switzerland introduction during the last two decades, qualitative inquiries have gained increasing popularity among european researchers in the social sciences and related fields. however, despite the institutionalization and legitimization of qualitative inquiries, this form of research “has not yet reached the same significance and reputation in switzerland as it has in many other countries” (eberle & bergman 2005: 1). qualitative research in switzerland is “lagging behind with regard to networks and structures that could offer information, support, resources, quality control and advanced training” (eberle 2005: 4). whilst there exists federal, cantonal, and private archives in switzerland that provide access to a variety of historical data, at the time of the bremen workshop in 2009 there was still no swiss archive for data collected within the framework of qualitative research projects in the social and related sciences. neither was there a resource centre or formalized network for offering services, information, and advice for researchers working within the qualitative research tradition. since that time, fors (swiss centre of expertise in the social sciences) has assumed the role of central archive for qualitative data produced in switzerland, and in coordination with various institutions and researchers has begun to make available resources for qualitative work. in this article we describe some of the potential for and obstacles to qualitative data archiving in switzerland. since archiving and re-use of qualitative data has to be discussed within a wider framework of quality concerns of qualitative inquiries (bergman & coxon 2005; eberle 2005), we first describe recent steps in promoting qualitative research in switzerland. we then examine current challenges to maintaining a qualitative data archive, and close by discussing future prospects in switzerland. the current sitution in switzerland: steps toward promoting qualitative research and data archiving as a result of the contrast between the increasing numbers of qualitative studies on the one hand and the lack of institutionalization of qualitative research on the other, the swiss academy for humanities and social sciences (sagw/assh) has launched several initiatives to promote qualitative research in switzerland2 . the primary goals of these initiatives are to build and strengthen networks, work toward best practices in methods and teaching, and to assess the feasibility of an archive and resource centre for qualitative research. in cooperation with the former swiss information and data archive service for the social sciences (sidos, now part of fors, see below) and the social science policy council (a committee of the sagw), a workshop was conducted in 2002 to identify the experiences of active qualitative researchers in switzerland as well as key representatives of qualitative archives and similar institutions from other european countries3 .this event led to subsequent meetings and to three working groups, which were asked by the sagw to produce an informative and accessible document on (a) the possibilities and limits of qualitative research for the social and related sciences, (b) qualitative data archiving in switzerland by brian kleiner, claudia heinzmann, thomas s. eberle, manfred max bergman1 78 iassist quarterly 2010 / 2011 iassist quarterly quality criteria for assessing research results from qualitative research, and (c) recommendations on how to integrate qualitative research methods into the university curriculum. a document summarising the issues elaborated by the working groups was discussed at a final meeting with international experts, found strong support from the qualitative research community in switzerland, and a fully elaborated “statement” was published in 2009 (bergman et al). since the 2009 bremen workshop, fors has taken concrete steps to establish the capacity for archiving qualitative data, including a workshop of archiving and research experts in 2010, development of specific policies and procedures, as well as workflow and system adjustments to integrate qualitative data into its database. with the capacity now in place, fors is poised to introduce qualitative research data into its holdings. obstacles to qualitative data archiving in switzerland while fors has begun archiving data from qualitative research projects, it is not clear yet that the research community in switzerland will take notice, deposit their data, and make use of the qualitative data of others. specifically, there remain a variety of significant potential barriers: 1. the use of secondary data is considered to be more easily applicable for quantitative research than for qualitative studies. in general, qualitative researchers in switzerland are not familiar with the possibilities of secondary data analysis, and they do not know which research topics or potentially available data sets are suitable for secondary data analysis. more specifically, the idea of developing research questions based on the data of “someone else” seems challenging. this difficulty is related to the belief that it is necessary for qualitative researchers to go through the whole process of data collection in order to contextualize the material in an appropriate way (e.g., corti 2000: 26 for more details; for a critique, see moore 2007). 2. the various types of qualitative data make it difficult to decide which material should be archived and provided for re-use. among researchers, specific concerns arise in relation to supplementary materials, such as researchers’ field notes and personal notes made before and after interviews. these materials are quite important in providing an interpretive context, but they are often considered to be too private or sensitive to be shared with other researchers. 3. one of the key concerns is related to the ethical and legal implications of rendering qualitative data accessible. that is, there is the problem of ensuring anonymity, confidentiality, and data protection. can qualitative data be adequately and consistently anonymised without reducing the value of and interest in the data for re-use? if not, are there other ways to address adequately the need to protect confidentiality, such as informed consent or strict access conditions? beyond questions of anonymization and data integrity, it is not entirely clear how swiss law treats the subject of data protection. although there are federal and cantonal laws on archiving and data protection (e.g. confédération suisse 2008; grand conseil du canton du vaud 2007), there are no national policies specifically relating to qualitative data. these legal issues should be addressed by specialists knowledgeable about swiss law. 4. the various types of qualitative data (e.g. transcripts, field notes, audio and videotapes, pictures, etc.) present difficulties with respect to adequately preserving the data over a long period of time (see corti 2000: 20f for a detailed discussion). fors should have sufficient resources for ensuring the requisite infrastructure, staff, and knowhow for dealing with different and changing formats over time. 5. finally, there are financial challenges in developing and maintaining an archive for qualitative data indefinitely. in addition to these potential challenges, there is the problem of whether or not there will be significant interest in archived qualitative data. in switzerland, data re-use is not a deeply engrained part of the research culture. furthermore, secondary analysis is far more common with quantitative data. with respect to re-use and sharing of qualitative data, this happens only occasionally for individual projects. it is quite likely that the lack of work in this area is due to insufficient know-how and training on the part of most researchers, lack of available national data for re-use, a comparatively generous funding infrastructure for the collection of new data, and a general lack of awareness and appreciation about the value and potential of the re-use of high-quality data. it is clear that much work needs to be done to develop, advance, and promote re-use of qualitative data in switzerland, which should include training, networking, and support in research grant applications. development planning currently, fors is strongly interested in continuing in the direction of qualitative data archiving, dissemination, and the provision of additional resources to researchers. it has developed a set of institutional policies and procedures regarding qualitative data, as well as a guide to assist researchers in how to prepare their data for deposit. the social science research holdings within the archive at fors are currently mostly quantitative, but now include data from several qualitative research projects. in concert with researchers within switzerland (including several authors of this paper), future efforts will be devoted to promoting and encouraging secondary analyses of qualitative data and deposit of project data at fors. there are some gaps that should be noted. realising qualitative archiving at fors in the long-run will certainly require additional resources, including at least one new staff member and additional technical infrastructure and development. in any case, the continued development of qualitative data archiving in switzerland should be done in close collaboration with experienced researchers and with universities where qualitative research is firmly established. conclusion even in quantitative research, secondary data analysis is not well established among some research branches in the social and related sciences in switzerland. nevertheless, most stakeholders in the social science research domain in switzerland would agree on the various scientific and cost-benefit advantages of secondary data analysis of high-quality data from qualitative research projects. realising a sophisticated data archive for qualitative research in switzerland will have to include (a) the support of research funding institutions and research policy makers, (b) more systematic training in qualitative research at universities and the connected possibility of secondary data analysis, (c) a well-equipped and staffed archive that actively conducts outreach projects and serves as a network node and resource centre, and (d) the uptake of the use of an archive by the research community, encouraged and supported by funding bodies, an active research network, and their own methodological expertise. references bergman, m.m & coxon, a.p.m. (2005). ‘the quality in qualitative methods. forum qualitative sozialforschung / forum: qualitative social research. 6 (2). art. 34 bergman, m.m & eberle, t.s. (eds.). (2005). ‘qualitative inquiry: research, archiving, and reuse’ . forum qualitative sozialforschung / forum: qualitative social research. 6 (2) [online]. available at: (http://www. iassist quarterly 2010 / 2011 79 iassist quarterly qualitative-research.net/index.php/fqs/issue/view/12). [accessed 17th july 2009] bergman, m, eberle, t., flick, u, förster, t, horber e, maeder, c, mottier, v, nadai, e, rolshoven, j, seale, c, widmer, j. (2009). a statement on the meaning, quality assessment, and teaching of qualitative research methods.[online]available at. http://www.qualitativeresearch.ch/docs/manifest_qualitative_sozialforschung_online.pdf confédération suisse (2008). loi fédérale sur l’archivage . [online]. available at : http://www.admin.ch/ch/f/rs/152_1/index.html [accessed 17th july 2009] corti, l, witzel, a & bishop, l (eds.) (2005). secondary analysis of qualitative data, forum qualitative sozialforschung / forum: qualitative social research. 6 (1). [online]available at: http://www.qualitativeresearch.net/index.php/fqs/issue/view/13 [accessed 17 july 2009] corti, l, kluge, s, mruck, k & opitz, d (eds.) (2000). ‘text, archive, reanalysis’. forum qualitative sozialforschung / forum: qualitative social research,.[online] 1 (3) available at:.http://www.qualitativeresearch.net/index.php/fqs/issue/view/27>accessed 17 july 2009] corti, l. (2000). ‘progress and problems of preserving and providing access to qualitative data for social research. the international picture of an emerging culture’. forum qualitative sozialforschung / forum: qualitative social research. 1 (3). art. 2 eberle, t. (2005). ‘promoting qualitative research in switzerland’. forum qualitative sozialforschung / forum: qualitative social research. 6 (2). art. 31 eberle, t.s. & bergman, m.m. (2005).’ introduction’. forum qualitative sozialforschung / forum: qualitative social research. 6 (2) .art. 30 fielding, n.( 2005). ‘the resurgence, legitimation and institutionalization of qualitative methods’ forum qualitative sozialforschung / forum: qualitative social research. 6 (2). art. 32 le grand conseil du canton du vaud .(2007). loi sur la protection des données personnelles. [online]. available at : http://www.rsv.vd.ch [accessed 17th july 2009] moore, n.( 2007).’ (re)using qualitative data?’ sociological research online. 12 (3). notes 1. brian kleiner fors, swiss centre of expertise in the social sciences, lausanne, switzerland claudia heinzmann institute of sociology, university of basel, switzerland thomas s. eberle institute of sociology, university of st. gall, switzerland manfred max bergman institute of sociology, university of basel, switzerland contact: brian kleiner, brian.kleiner@fors.unil.ch 2. for detailed discussions and reasons on why to promote qualitative research, see eberle (2005). 3. updated versions of the conference presentations are published in bergman & eberle (2005). 14 iassist quarterly issues of privacy and access represent a threat to the survival of the data organizations and to a certain extent even of quantitative social research. consequently, in august of 1978, ifdo sponsored an international conference on emerging data protection and the social sciences' need for access to data . at this conference, which was hosted by the most experienced and biggest european data archive, the zentralarchiv in cologne, comparative national status reports were presented from 10 coimtries. the national reports were sent to the organizers who collected them in a volume of proceedings that was a tangible point of departure for the discussions at the 3-day conference. by per nielsen' early academic reactions to privacy and access regulations the ifdo resolutions, august 1978 for a few years in the mid-seventies, members of the international social science community could study the hessian and the swedish data legislation practices whilst preparing the viewpoints for which they foimd it necessary to fight on their national home groimd before the enactment of similar privacy legislation. to member institutions of the newly established international federation of data organizations (ifdo), the issues of privacy and access were of central importance, not just as an academic field of study, but as central issues that might the participants invited to the cologne conference unanimously adopted three ifdo resolutions which served the purpose of drawing attention to as many aspects and consequences of data legislation as possible. the ifdo resolutions are appended to this note as one of the first, outspoken, academic reactions to privacy and access regulations. they are included in the same form as that in which they were presented to the danish public in the danish data archives (dda) newsletter^ the bellagio principles, 1977 one year prior to the ifdo conference, a group of social scientists and senior administrators at national statistical biu'eaux had discussed access to statistical data in bellagio, italy. from this event, 18 so-called bellagio principles were circulated in the social science community. these principles were considered important because they represented a first compromise between social scientists on the one hand and senior statistics administrators on the other. 'presented at lassist/ifdo international conference may 1985. amsterdam. dda-nyt 9:12-14, december 1978. spring 1986 iassist quarterly 15 in an explicit statement, the ifdo conference endorsed the bellagio principles, which are reproduced below again in the same form as that in which they were presented to the danish pubhc in the dda newsletter^. the european science foundation statement, 1979-1980 in 1979, a working group of invited specialists in the social sciences, the medical sciences, and administrators from data inspection authorities, tried to reach agreement on a statement which was going to be subject to approval by the assembly of the esf. dtmng the working group meetings, i felt a peculiar distrust between each group and the other two. the medical experts asserted that their data-handling procedures were safe and felt that it was in the interests of patients (i.e. the public) to supply necessary information to their doctors without too much interference from the data inspection authority; on the other hand, the medical experts were sceptical about some of the data collection and handung procedures applied by social scientists! the experts with a social science background tended to hold that their own rationale for data collection, as well as their apphed data handling routines, were less dangerous to the public than most of the data collection ventures within medical science. and the data inspection authorities felt that both medical and social science projects involving confidential data should be rather rigidly controlled. in addition to these disciplinary variations in attitudes, the national differences were more outspoken in the esf working group than they had been in either bellagio (where canada, us, uk, west germany and sweden were represented by scientists and statisticians) or in cologne in which about a dozen countries were involved. furthermore, it took more than a year to reach agreement on the wording of the dda-nyt 9:14-16, december 1978. final texl after reworking the text as adopted at the conference, a slightly rephrased version was accepted by the esf assembly. it is this revised (official) version of the esf statement which is appended to this paper. reasons behind the diversification in attitudes i think it is fair to say that three or four major factors caused the change in attitudes to the issues of privacy and access during the last half decade of the seventies from consensus to a more diversified set of attitudes. first, the various groups of agents became more aware of their group interests in the course of the data legislation process as the latter proceeded in more and more countries. second, the discussions moved from a level of soft statements towards one of juridical phraseology in sections and subsections. third, the difterences in existing legal conditions between countries (e.g. in such areas as freedom of information) as well as practical set-up (e.g. a tradition for codes of ethics) implied associated difterences in the new legislation and in its actual implementation. this indicates that there is still a lot of research to be done in terms of comparing the conditions for quantitative research between countries as well as following the trend over time within a single country as practices are defined and acts are amended. as can be expected, substantial interest is devoted to this issue among social science data "pushers" and "addicts". since 1977, there has hardly been a conference of any size or generality which has not had issues of privacy and access on its agenda. concluding remarks and recommendations as a convenor of the ifdo/iassist 1985 conference session on issues of privacy and access, 1 thought that it might be useful to reprint some of these early deliberations, in spring 1986 16 iassist quarterly order to facilitate discussion along the following lines: what new issues (if any) have entered the debate in recent years, and what is the present-day situation, compared to expectations 5 or 10 years ago. finally, i should very much like to see a repetition of the 1978 ifdo conference. now that most coimtries have actually been living with enacted privacy bills, a new systematic comparison across countries would prove useful. ifdo international conference on emerging data protection and the social sciences' need for access to data resolutions in a plenary session the conference unanimously adopted the following three statements. social scientists' experiences with data protection. on the basis of evaluation of developments in data protection within eleven coimtries, and taking account of the general tendency for legislative measures to have unintended consequences, the conference expresses grave concern about some of the negative impact of data protection laws, regulations, and practices on the social sciences. while we recognize that it is essential to protect the privacy (integrity) of the individual, there is also a need to know and a need to secure the channels through which, under proper safeguards, a reliable and comprehensive understanding of the life situation of individuals and groups of individuals may be obtained. in the opinion of the conference the need to know and the need to secure a free flow of information constitute the other side of the issue of protecting the privacy of individuals. to a large extent this other side of the privacy issue has not been given due consideration in the process of enacting and implementing data legislation. the conference would like to draw attention to the fact that such legislation can and has become a vehicle for the protection of the vested interests of particulariy resourceful groups and organizations, thus contributing toward an infringement of the fundamental rights of other parts of society. it is recognized that the results of significant social research might jeopardize the interests of some of the groups or individuals about whom data are collected. however, it seems important to be sensitive to the possibility that because of this situation data protection measures can be utilized as a shield behind which socially significant issues are excluded from research. furthermore, developments in the field of information processing have resulted in very powerful instruments to control individuals and society. in most of the countries represented at the conference data protection laws are used by bureaucracies to monopolize the information necessary for the open discussion of public policies. the data fiow among government agencies has increased considerably during the last few years, although data protection has in some cases placed restrictions on this flow. however, researchers often find themselves excluded from the information necessary to enable them to contribute to public discussion by presenting independent opinions. this is especially dangerous in a situation where government policies are based increasingly on large data bases, including microdata. the conference is of the opinion that these issues have significant political implications and are associated with broad and general notions of the free and unrestricted flow of information in society. they should be given thorough political spring 1986 iassist quarterly 17 consideration in the future development of data legislation and practices. the conference has learned that with respect to data protection there are significant differences in the situations of the different countries. there are nations that have found an acceptable balance between data protection and access to data for research purposes. on the other hand, there are countries where data flow for research has come nearly to a standstill. in this situation it is necessary to develop guidelines for a general information policy. a fundamental aim of a modem information policy is to make information gathered by public (and private) institutions more transparent and visible in order to improve democratic control. within this broader framework, social research must be considered not only as a matter of interest to social scientists, but as part of that system of democratic control. a first important recognition of these problems at the international level came in 1977, when a group of social scientists and senior administrators of national statistical bureaus discussed the issue and drafted a set of recommendations, which are now known in the international social scientific community as the bellagio principles. we endorse these principles. we also hope that the pattern set by the bellagio conference of joint discussion of common problems between social scientists and governmental officers at all levels will be continued. in the perspective, the distinction between statistical and administrative data should not be used to make the latter less accessible to researchers. access to administrative data for scientific purposes should be regulated according to the principle or functional separation of research and administrative data incorporated also in the bellagio principles. the conference wishes to point to the high value placed on freedom of the press. the social science community might be in a better position to improve its services to society if its freedom and rights to do research were secured through similar principles, including the obligation to protect the source of informatioa preservation and accessibility in addition to these general principles the conference recognized other points of interest for the international development of social research. in particular it recommended: • that the data relevant to scientific investigations on human affairs should be preserved in readily usable forms; that with the sole limits of protection of privacy and confidentiality recognized in the first part of this statement, research data should be openly accessible to social researchers and the general public of all nations; that govenmients should work to eliminate barriers to general access to research data and should take appropriate action to facilitate their use under the principles established by the united nations charter and incorporated in unesco. codes of conduct finally the conference supported the following recommendations toward the adoption of codes of conduct by social researchers: social scientists collect information from and about individuals for research purposes. in doing so they have traditionally followed certain standards of behaviour: social research is conducted at all times so that no harm should come to spring 1986 18 iassisl quarterly individuals while being subjects of research. the current concern to beuer protect the privacy of individuals makes it necessary to increase awareness of difterences between administrative and research uses of information. to make this point better understood by the public, governments and researchers, it is recommended that in addition to the existing codes of ethics in various disciplines, codes of conduct should be developed for each research methodology. these can make explicit the rules that are already respected by the professional researcher. thus, by common practice in survey research, the anonymity of respondents, their right to be informed about the purpose of a study, their right to refuse cooperation at each stage of an investigation, and their right to know the identity of the researchers have been respected. the practical ground rules for the responsible research use of personal data will differ with the research method. each professional specialty should be asked to make its practitioners fully aware of the range of alternative techniques available to implement codes of conduct for survey research, as an example, such alternatives include randomized response methods, insulated data banks, and appropriate levels of aggregation. codes of conduct should have sanctions so that the public can be assured that such codes of conduct are more than mere declarations. the bellagio principles excerpted from david h. flaherty's report*. 1 national statistical offices should provide researchers both inside and outside government with the broadest practicable access to information within the bounds of accepted notions of privacy and legal requirements to preserve confidentiality. 2 legal and social constraints on the dissemination of microdata are appropriate when they reflect the interests of respondents and the general public in an equitable manner. these constraints should be re-examined when they result in the protection of vested interests, or the failure to disseminate information for statistical and research purposes (i.e., without direct consequences for a specific individual). 3 all copies of government data collected or used for statistical purposes should be rendered immune from compulsory legal process by statute. 4 in making data available to researchers national statistical offices should provide some means to ensure that decisions on selective access are subject to independent review and appeals. 5 the distinction between a research file, in the sense of a statistical record (as defined in the 1977 report of the u.s. privacy protection study commission), and other micro files is fundamental in duscussions of privacy and dissemination of microdata. all dissemination of government microdata " david h. flaherty. final repon of the bellaggio conference on privacy, confidentiality, and the use of government microdata for research and statistical purposes. statistical reporter 78-8:274-279, may 1978. spring 1986 iassist quarterly 19 discussed in connection with the bellagio principles is assumed to be a transfer of data to research files for use exclusively for research and statistical purposes. 6 there are valid and socially significant fields of research for which access to microdata is indispensible. statistical agencies are one of the prime sources of govenmient microdata. 7 public use samples of anonymized individual data are one of the most useful ways of disseminating microdata for research and statistical purposes. 8 techniques now exist that permit preparation of public use samples of value for research purposes within the constraints imposed by the need for confidentiality. countries with strict statutes on confidentiality have prepared public use samples. 9 there are legitimate research purposes requiring the use of individual data for which public use samples are inadequate. 10 there are legitimate research uses which require the utilization of identifiable data within the framework of concern for confidentiality. 11 other techniques of extending to approved research the same rights and obligations of access enjoyed by officers of the government agency need to be considered in terms of better access. 12 there is considerable potential for development of more economical and responsive customized-user services, such as: 1) record linkage under the protection of the statistical office, 2) special tabulations, 3) public use sample for special purposes. such services must often involve some form of cost recovery. 13 some research and statistical activities require the linking of individual data for research and statistical purposes. the methods that have been developed to permit record linkage without violating law or social custom regarding privacy should be used whenever possible. 14 professional or national organizations should have codes of ethics for their disciplines concerning the utilization of individual data for research and statistical purposes. such ethical codes should furnish mutually agreeable standards of behaviour governing relations between providers and users of governmental data. 15 users of microdata should be required to sign written undertakings for the protection of confidentiality. 16 considerable efforts should be made to explain to the general public the procedures in force for the protection of the confidentiality of microdata collected and disseminated for research and statistical purposes. 17 the right of privacy is evolving rather than static, and closely related to how statistics and research are perceived. therefore, statisticians and researchers have a responsibility to contribute to policy and legal definitions of privacy. 18 public concern about privacy and confidentiality in the collection and utilization of individual data can be addressed in part as follows: a. voluntary data collection, whenever practicable, b. advanced general notice to respondents and informed consent, whenever practicable, c. provisions for public knowledge of data spring 1986 20 iassist quarterly public education on the distinction between administrative and research uses of information. efs's statement on 'privacy' statement concerning the protection of privacy and the use of personal data for research (adopted by the assembly of the esf on 12 november 1980) ' preamble the necessity of safeguarding the individual against misuse of his personal data has been repeatedly emphasized, in the last few years, at both the national and the international level. this has been particularly the case in the countries with organizations which are affiliated to the esf. in austria. portugal and spain data protection is explicitly referred to in the constitution. specific legislation already exists in austria, denmark, france, the federal republic of germany. norway and sweden. draft laws are under consideration in belgium, the netherlands and switzerland, while an official report on the issue has been prepared in the united kingdom. there has also been considerable concern v/ith these matters at the international level. the council of europe has recently elaborated a convention for the protection of individuals with regards to automatic processing of personal data , while the oecd has prepared a series of guidelines concerning the protection of privacy and the movement of personal data across frontiers. mention should also be made of the discussions going on within the commission of the european communities about a possible directive and of the enquiry carried out by the european parliament which led to a resolution calling for immediate action. however, the implementation of data protection laws has led. in an increasing number of cases, to serious restrictions on access to personal data for research purposes. for example, problems connected with the collection and evaluation of information by means of questiormaires, access to information held by public authorities, particularly statistical offices, and the destruction of personal data by such authorities once the purposes for which they were collected have been fulfilled, have been creating considerable concern amongst the scientific community. this led to the drawing up of the bellagio principles in august 1977' and to an international conference on emerging data protection and the social sciences' need for access to data which was held in cologne in august 1978. sponsored by the international federation of data organizations (ifdo). these problems were also discussed at the 10th colloquy on european law organized by the council of europe at uege in september 1980. the esf fully endorses the necessity of protecting the privacy of the individual. it feels, however, that the attention of the legislators and international bodies conemed should drawn be to the researchers' case for special conditions for the use of personal data. these should ensure, under proper controls, dda-nyt 18:9-13, sommer 1981. 'contained in the final report of the bellagio conference on privacy, confidentiality, and the use of government microdata for research and statistical purposes, which was a meeting of representatives of the central statistical agencies of canada, the federal republic of germany, sweden, the united kingdom and the united states held at the rockfeller foimdation bellagio study and conference center in italy, 16-20 august 1977. spring 1986 iassist quarterly 21 access to such data when it is needed for specific research purposes. accordingly, a group of experts under the chairmanship of professor s. simitis, professor of civil and labor law at the university of frankfurt and data protection commissioner of the state of hesse in the federal republic of germany, was set up to draft such a statement after full discussion and revision within the esf the following principles and guidelines were adopted by the esf assembly at its meeting in november 1980. they are put forward to ensure both the protection of personaal data and the necessary access to such data for research purposes. j. goormaghtigh secretary general strasbourg 13 november 1980 basic principles 'personal data' are, in the context of this document and in accordance with the definition to be found in the coimcil of europe's convention for the protection of individuals with regard to automatic processing of personal data and also adopted by the oecd. any information relating to an identified or identifiable individual. data protection legislation must, in order to fulfill its task, which is to guarantee the respect of privacy, cover all uses of personal data and therefore include its use for research purposes. professional codes of ethics are a complement to legislative measures safeguarding the respect of privacy. the scientific communities concerned should encourage the development of such codes, within the fram.ework of the rules established by the legislator, in order to take into account the specific needs of the different disciplines. freedom of research presupposes the broadest possible access to information. legislation should, therefore, besides specifying the conditions tmder which personal data may be used for research, ensure access to the information needed. in order to ensure the respect of privacy, research should, wherever possible, be undertaken with anonymized data, following already accepted practices. scientific and professional organizations, together with pubuc authorities, should promote further development of techniques and procedures to secure anonymity. anonymity should be considered as given, whenever the individual can only be identified with an imreasonable amoimt of time, cost and manpower {de facto anonymity). guidelines any use of personal data for research purposes, irrespective of the aims for which they were or are to be collected, presupposes either the explicit permission of the legislator or informed consent unless the individuals concerned are not identifiable by the receivers. there is informed consent when the individuals concerned have been clearly informed: a. that the provision of data is volimtary and that a refusal to comply will have no adverse consequences on them; b. of the purposes and nature of the research project; c. by and for whom the data are being spring 1986 22 iassist quarterly collected; d. thai the data collected will not be used for any other purpose than research. with the approval of the data protection authority, or its equivalent, informed consent is not required in cases where the nature of the research project is such that: a. the informed consent of the individual would invalidate important objectives of research; b. informed consent could cause mental or physical distress to the individual concerned. for the sole purpose of selecting samples for research involving population-based surveys, legislation or other legally acknowledged procedures should permit the use of data concerning name, address, date of birth, sex and the occupation of individuals collected by state agencies for non-research purposes. personal data obtained for research should not be used for any other purpose but research. in particular, personal data obtained for research purposes should not be used to make any decision or take any action directly affecting the individual except within the context of research or with the specific authorization of the individual concerned. confirmation whether or not data pertaining to him are maintained, to challenge data relating to him and to have data erased, rectified, completed or amended should be limited to other research projects where it is intended that the data be used in an identifiable form. the leaders of research projects using personal data should be responsible for ensuring that the necessary technical and organizational measures are taken in order to guarantee the confidentiality and security of the data and for keeping these measures imder review in accordance with the latest scientific and technical developments. once the specific research purpose for which personal data have been collected has been achieved, these data should be depersonalized and the necessary measures (e.g. the deposit of identifying code numbers with a central research data archive) should be taken for their secure storage. the decision to destroy personal data held by public authorities should only be taken after consideration of their possible future use for research and after consultation with the central data archive or a similar organization. whenever personal data are used for research, they should not be published in identifiable form unless the individuals concerned have given their consent in the case of personal data used for research, the individual's right to obtain spring 1986 vol263 4 iassist quarterly fall 2002 iassist quarterly fall 2002 5 social science information service in poland: an attempt to present the state of art by teresa wildhardt & anna sokolowska-gogut* abstract the political, economic and educational transformation taking place in poland for several years has resulted in, among other things, ever-growing interest in social sciences and, consequently, in growing demand for specialized information in that field. this paper tries to give answers to the following questions: § can we observe any direct relation between transformation and growing demand for social science information service? if so, § what kind of information is searched for most frequently? § are social science information services ready to supply users with necessary information? a questionnaire has been sent to a number of representative scientific libraries. special regard has been given to the bureau of polish official statistics. the feedback of responding libraries has been thoroughly analysed and results presented in the paper. in the paper an attempt is made to describe a condition of contemporary social science information service in poland. a brief characterization of information sources, staff and categories of users is given. introduction great changes taking place in poland in the last decade due to political, economic and educational transformation has resulted in, among other things, ever-growing interest in social sciences and ever-growing demand for specialized information in that field. the number of students in social sciences doubled within the last six years, and new specialized studies have been launched. the process of transformation has drawn the attention of many research workers who examine and describe new phenomena. that process has been the object of interest for the authors of this paper as well, but only those aspects of it which are connected with information supply, take place in libraries and have influenced the social sciences information services activity. this paper tries to give answers to the following questions: § can we observe any direct relation between transformation and growing demand for social science information services? § what kind of information is searched for most frequently? § are social science information services ready to supply users with necessary information? survey process a questionnaire has been sent to 40 representative scientific libraries, including libraries of major polish universities, higher education institutions (of university level), governmental institutions, scientific institutes, and municipal libraries. the questionnaire consisted of four main sections: social science information services general information and introductory remarks; social science information services users; sources and forms of services; and information searched for most frequently. responses have been received from 32 libraries, however two questionnaires could not be taken into consideration because only one third of questions had been answered. five libraries have informed us that they do not run specialized social science information services, but inquiries relating to that topic may be addressed to their information desks. there has been no answer from three libraries. in all, feedback from 30 libraries has been analysed and an attempt has been made to present a condition of contemporary social sciences information service in poland. a brief characterization of information sources, staff and categories of users has been given. special regard has been given to polish central statistical office. in this context let us draw your attention to one fact. the ranking of the best state higher education institutions (of university level) reported recently in two outstanding polish papers (“rzeczpospolita”, “perspektywy”), included 75 schools, of which 45 run polytechnics, physical training 6 iassist quarterly fall 2002 iassist quarterly fall 2002 7 and medical studies. in view of the above, the sample (collected material) can be treated as representative. the analysis of feedback allows us to give a general description of contemporary social science information services in poland. the collected material is discussed briefly in the following sections: § social science information services launching § staff § categories of users § facilities and information sources § forms of services § information searched for most frequently § central statistical office § conclusions social sciences information services launching in poland, the first social science information services started their activities in the 1920s. the majority (50%) of them, however, began their work after the second world war and during the period 1960-1985. all respondents unanimously observed a radical increase in demand for social science information services after 1990, when big political, economic and educational changes began. this ongoing process of transformation, and preparation made by poland to integrate with the european community, allow us to expect that interest in social science information services will continue. staff according to received responses, all social science information services staff are well prepared to perform their duties. they have master degrees corresponding with their occupations, i.e., in the fields of economics, sociology, political science, psychology, pedagogic, library and information sciences. apart from that, 50% of examined staff attended special courses and training. they have knowledge of how to access and navigate databases (lan, man, www). they know how to download, print and operate with applications packages. they are always ready to supply users with necessary information, give advice and locate accessible data on a requested topic. users expect social science information services staff to be competent, and so they are. categories of users permanent visitors of social science information services are individuals, as well as institutions and organizations. individuals include students, research workers, persons preparing dissertations, teachers, journalists, managers, and entrepreneurs. the majority of individual users are students, teachers and research workers. institutions and organizations include governmental institutions, local government (self-government), educational organizations, and various foundations. moreover, any person who seeks information or advice related to social science information services is welcome. facilities and information resources all surveyed social science information services have at their disposal, reference collections consisting of printed materials such as encyclopaedias, dictionaries, bibliographies, and subject-oriented monographs, and materials issued by the central statistical office. since 1993/94, all social science information services have been able to provide access to information sources in electronic form and to the internet. they possess a wide range of polish and foreign databases on diskettes and cds, and online access to some foreign databases, mostly free of charge, issued by the european union (e.g., eur-lex, celex, europa). the most frequently quoted foreign databases subscribed to by libraries and the percentage of social science information services that have them are: § wilson social sciences abstracts (30%) § eric on silver platter (25%) § social sciences citation index with abstracts (15%) about 35% of social science information services elaborate their own databases. those databases refer to socioeconomic problems, european law, and education (e.g., “social sciences” in two languages (polish, english) and “economics on-line” issued quarterly by the main library of cracow university of economics). forms of services social science information services staffs offer a wide range of information dealing with socio-economic problems. they are capable to provide users with necessary information, give advice, and locate accessible data on the topic of interest. the forms of information delivery are different, but most commonly quoted are the following: § direct inquires on the spot at information desk § by telephone § via email § on-line services (topic-oriented searching) § provision of data to individuals using desktop pcs (exceptionally) § downloading at the userʼs request, social science information services officers elaborate subject-oriented bibliographies and statistical data, and prepare printed extracts from databases. access to databases and internet for parent institution users is free of charge but limited in time. 6 iassist quarterly fall 2002 iassist quarterly fall 2002 7 information searched for most frequently the very term “social science” designates the boundaries of information for which one might search. “social sciences” are meant to be a set of disciplines which deal with aspects of human society. it is commonly understood that “social sciences” include, first of all, economics, sociology and political sciences. the received responses allow us to identify precisely what kind of information is needed in poland now. we can identify the following topics (problems) which have been the object of special interest to social science information services users, and the information relating to these subjects that has been searched for most frequently: § quality of life § relation of economics to social values § economics of the elderly § child care § unemployment § demographic trends and forecasts § social reforms (of education, health, social insur ance and pensions) § state and local government § taxation § education § european union § integration with eu § european law § small business § marketing § analysis of social consequences of transformation § all kinds of statistical data in summary, information relating to socio-economic problems, reforms, integration with the european union and statistical data are the hot subjects of interest to polish social science information services users. central statistical office polish statistical system the statistical system in poland has been functioning since 1789, when the nationwide census and property state was carried out. the first in poland and one of the earliest statistical offices in europe was founded in 1810. the central statistical office (cso), which nowadays is the central statistical institution in the state, was created in 1918. it supervises the activities of the regional statistical offices. next year it will celebrate its 200th anniversary. the contemporary statistical system in poland consists of cso and 17 regional statistical offices collecting and processing various data, the statistical information system, classifications, registers, and the publication system. each calendar year, cso announces the statistical survey programme. the schedule for the year 2001, included 15 main categories, such as the population and demographic processes, labour market, wages and social insurance, education, science and the technical progress. cso information system the statistical information system is conducted by both the cso and the regional statistical offices, always using the nationwide data. individual and personal data collected during statistical surveys are confidential and subject to special protections. they are used exclusively for elaborations and analysis carried out by the official statistics. the legal basis of polish official statistics is the official statistics act of 1995. it guarantees indiscriminatory, equal, and simultaneous access to statistical data for all users. the cso system divides the information users into three groups, whose information inquiries are responded to by two units: data dissemination division (ddd) and information and publication bureau (ipb). ddd serves authorities and government administration, local government and international organizations. the services are free of charge. ipb serves mass media and other users. the detailed information and analyses are charged. it is also possible to obtain information by phone from the bulletin board system (bbs), automatical statistical information and the central statistical informatory. cso library the cso library the central statistical library was founded in 1918. the book collection is the oldest and the largest special collection in poland. it contains all the cso publications from the beginning of its origin, as well as more important publications of the local statistical offices in poland and a collection of current and retrospective worldwide statistical publications. the library collects the statistical yearbooks from about 150 countries and those of international organizations. the serials collection contains mainly the current journals dealing chiefly with statistics and demography. the cso publication system the data collected and elaborated by cso divisions are issued in printed and electronic form. we may distinguish different publication groups: statistical yearbooks, book series, bulletins and periodicals. some of these publications can be accessed directly from the cso web site at www.stat.gov.pl, as shown below. special regard should be given to book series published by cso divisions, such as: § studia i analizy statystyczne (statistical surveys and analysis) § studia i prace. z prac zakładu badań statystyczno-ekonomicznych gus i polskiej akademii nauk (studies and elaborations of research centre for http://www.stat.gov.pl 8 iassist quarterly fall 2002 iassist quarterly fall 2002 9 economic and statistical studies of the central statistical office and the polish academy of sciences) § informacje i opracowania statystyczne (statistical information and elaborations) § materiały źródłowe (statistical sources) § biblioteka ”wiadomości statystycznych” (”statistical news” library) cso information services cso offers processing and analysis of statistical data gathered within the confines of the annual statistical survey program of official statistics. results of the surveys are disseminated in the form of cso publications as well as surveys ordered by users. orders demanding distinct calculation and elaborations of statistical data are charged according to the price list. usage of the local data bank is free in the scope of topics specific to the cso regulations; access to the other data is charged. access to other on-line statistical sources (e.g., “information about social-economic condition” or “poland – main macroeconomic indicators”) is charged. conclusions 1. a direct correlation has been observed between transformation taking place in poland since 1990, and ever-growing demand for social sciences information services. 2. the interest in social science information services has intensified within last few years in view of the expected integration with the eu and reforms implementations. 3. information relating to socio-economic problems, reforms, integration with the eu and statistical data are searched most frequently. 4. social science information services staff have satisfactory qualifications to perform their duties and meet the demands made upon them. however, training and practice in well-organized specialized information centres could contribute to improvement of staff skills. 8 iassist quarterly fall 2002 iassist quarterly fall 2002 9 5. the main obstacle in further development of social science information services is scarcity of funding to support the high costs of necessary equipment and foreign database subscriptions. * paper presented at the iassist conference, may 2001, in amsterdam, the netherlands. teresa wildhardt, cracow pedagogical university, main library, poland (sbwildha@cyf-kr.edu.pl) and anna sokołowska-gogut, cracow university of economics, main library, poland (gogut@bibl.ae.krakow.pl). iassist quarterly 1 developing a social science microcomputer local area network: its uses and problems features.the microcomputer local area network is more than a technological fad, it is convenient and elegant way to add additional function to exisiting equipment for a moderate price. with proper planning and managment, local area networks' uses far exceed their problems. this paper will briefly discuss the basic definition of a local area network and review some of the points to investigate when choosing one. it will then cover the problems and solutions one might encountered in constructing and running a social science local area network. by fred nick center for social science computation and research university of washington rgure#1 abstract a local area network offers a social science microcomputer user many nice advantages but it can also create a host of problems for the individuals in charge of network administration. as social science microcomputer use continues to increase, local area networks will offer social scientists welcome communication, file storage and peripheral equipment sharing capabilities. for the network administrators problems arise in: initial funding for the additional equipment needed, the definiiion of which network to install, the added security problems in running a network, additional responsibilities for storage backup, the establishmeni of access policies, the day-to-day maintenance of the network and its software, and providing the additional help needed by the users of the system so that they can take advantage of the new network numb^of pcson multlus^ pc systems (thousands) 40000 30000 20000 10000 1985 year 1990 i tolal number of pcs in uje (thouwnds) i n urn b er of pc3 on multiuser pc sysletnd (thousands) fall 1986 8 iassist quarterly what is a microcomputer local area network? a brief definition of a miaocomputer local area network (lan) might be as follows: the equipment and software necessary to connect microcomputers to each other so that they can share files and peripheral equipment while establishing inter-microcomputer communication. some lans also enhance the operating characteristics of connected pcs (personal computers). a network usually includes a workstation which functions as a file or disk server, containing files stored by users and administrators of the network. this server is connected to the other network workstations by a cable. there may be more than one server on a lan. and often there are many printers and other peripheral equipment attached to the network. figure ;f 2 projected installations of shared processors & network file servers (thousands) 3000 2500 2000 1500 1000 500 i^m m -i i i lil^ j . 11 1985 y^^ 1990 industry analysts predict that the number of miaocomputer lans will be 4 to 6 times greater in 1990 thjm in 1985 (see figures 1-3)' in 1985, 26,500 networks were purchased; it is projected that there will be 121,270 networks sold in 1990' figure #3 anttcipated grovvth or pc network shipments hoood 120000 100000 8og0o 60000 40000 20000 1984 85 86 87 88 89 1990 year 'lilv, susan vendors ready network versions of sojware. pc week 2(25): 102, june 25. 1985. (graphics included as figures 1-2 above) lilv, susan "ibm token ring opens doors for independent net vendors" pc week 2(51):113, december 24-31, 1985. (graphics included as figure 3 above) fall 1986 iassist quarterly 9 initial enviroiunent the first stage in establishing a local area network is careful planning. networks can be ver>' expensive, and have a large number of differing capabilities; a bad choice could be costly. the first step in planning is to examine the environment in which the network will be placed. what equipment is presently being used? how much experience with microcomputers does the user community' have? what fimctions are needed, and what capabilities are patrons likely to anticipate having available. a primary' complaint about lans is that they have not performed as the users had anticipated. unfulfilled expectations can make a system appear to be a failure when it had merely not been designed to do what people requested of it as examples, i will draw on the installation of a lan in my organization, a computer support group which helps soda! scientists to use computers. it has existed for 13 years, and has traditionally been mainframe oriented. it is located in a state-funded university, and derives its entire budget from state monies. it has no mechanism for recapturing expenditures through charges. all services are free to users of the facility. all users of the system were familiar with mainframe-style interactions and capabilities. we had purchased an apple 11+ five years earlier, but ii had proven a failure. not enough software was available at the time; the operating system was too simple, and the functions demanded by the mainframe users who experimented with the microcomputer were not available. thai purchase was ahead of its time. three years ago we again became interested in introducing microcomputers to our patron community. a new generation of computers became available, major software packages were creating renewed interest, and microcomputer word processing was becoming popular. also, several major computer companies were starting to give equipment grants to universities. we therefore decided to try to develop a mictocomputer lan. the apple 11+ experience influenced the decision-making process when we began to plan for a miaocomputer lan. our user would demand certain 'mainframe' characteristics. there must be 'personal storage'. patrons were very imeasy about having files that any user could read, manipulate or destroy. they were accustomed to security precautions such as passwords and access modes. if our lan was to include shared storage, on the same disk, then these features would be requested. since the users had mainly mainframe experience, it was essential that we provide access to the mainframe so that files and programs could be moved freely to and from the lan. we decided that since our organization had many diverse groups of potential users, we should choose one group and develop the lan with those individuals in mind. because we had little chance of developing a large lan, we decided that a small lan would have the most far-reaching exposure if it were dedicated to faculty education and instructional development faculty and teaching assistants interested in gaining skills on miaocomputers or developing microcomputer courseware were allowed access to the lan. our staff developed free non-credit courses to educate novice users. several other campus microcomputer facilities had developed severe managmeni problems when word processing was allowed on them. the microcomputers had quickly become 24-hour-a-day word processing stations to the exclusion of all other uses. we therefore did not allow 'production' word processing. development of word processing skills was allowed if noi encouraged, however. having identified what you want, whom it will serve and a few of the operational details, how do you gel started? fall 1986 10 iassist quarterly the first two tasks are obtaining financing and finding an appropriate space for the lan. if you can charge for services, you should experiment with possible funding mechanisms to determine if the costs of developing and supporting the lan can be recovered. lans seldom have a convenient method of keeping track of useage. charging by the hour, or a monthly fee, are possible alternatives. it is difficult to differentiate between "heavy" users and "light" users other than by the clock time they spend using the equipment even consumables, such as the paper used in printing, is difficult to monitor on some systems; other systems record the number of lines printed. some granting agencies may be appropriate sources of funds to finance the lan. detailed advance planning will help organize and collect the information needed for gram applications. at many institutions of higher learning, physical space may be more difficult to obtain than money. each lan microcomputer requires approximately 16 to 20 square feet of space, when placed on a table with a chair in front of it. extra power outlets will be needed, and it is important to remember that microcomputers can generate lots of heat (as do their human users). proper electrical connections, ventilation and heating must be considered. also security must be considered when looking for space. microcomputers are very popular booty for thieves. consideration must be given to which fioor of the building the room is on, where the windows are, how the room is locked, who has access to the room now, and if it can be rekeyed. careful planning can prevent many managment problems later. in developing our lan both space and funding were critical points. our financial situation could not have been worse when we decided to pursue obtaining a lan. our budget came directly from the state legislature. the state was in a financial crisis. our operating budget had just been cut by 20%. there were no capitol funds available for equipment purchases. the dean considered the project an experiment, and there was a considerable amount of prejudice against microcomputers to be overcome. undaunted, we decided to proceed. we went directly to computer manufacturers, asking them for gifts or loans. to our delight and surprise, the ibm corporation agreed to loan us 5 ibm xts, 5 printers and some software to get things started. these were to be loaned to us for 6 months, but in the end we had kept them for one and one-half years. after 6 months of our having used the borrowed xts, the budget situation had improved greatly. use of the borrowed equipment had also quelled many criticisms voiced earlier. at that time we were given $40,000 by the university to purchase a lan. many projects had been started on the borrowed equipment, which had laid a solid foundation for the request for a lan. space was donated by another department, electrical outlets were nimierous, and the power was not polluted by other equipment on the circuit (having other electrical equipment on the same circuit could cause voltage fiuctuations which might affect the lan equipment). the room had no windows, and the door was rekeyed so that even the janitors did not have access. we did not plan for ventilation, however. the room would become terribly warm with just a few individuals in it since there was no venlilalion system, nor windows, users would open the door to allow the room to cool off, undermining our security and allowing access by unauthorized individuals. we eventually had to trade rooms to remedy the ventilation problem. we had no funds with which to have proper ventilation installed. fall 1986 iassist quarterly 11 which network to buy? choosing a networl: can be very confusing. there is a variety of technologies, hardware requirements and functional difterences. networking technology has its own set of "buzz words", or jargon, which is highly complex. there are many topologies or workstation layouts. most of the literature is difficult to read without first having learned the definitions of the most corrunonly used phrases. the best approach in choosing a network is to avoid becoming too involved in comparing all the performance figures quoted in the literature. first, define how the network will be used, and then determine which network characteristics are important if large files stored on a central disk are going to be used, transfer speed may be an important criterion. if users are to be able to converse from one station to another, check to see that the nodes on the network can address each other and send messages directly to the screen of another node. electronic mail is an attractive method of communication which is not reliant upon the recipient being at his desk when you want to communicate with him. electronic mail can store and organize conespondenses, forward messages to others, and send replies along with the original message. cost is an important criterion. networking hardware and software can be very expensive and is seldom inexpensive. if you are building a network of exisitng pcs, it may be possible to use some existing hardware to provide the comunicalion connections. low speed lans can use standard serial cards as communication hardware. networks can cost more than a $1000.00 per station to install. always try to balance cost and function. one network may cost a little more than another, but offer far superior performance or stability. lan technologies have length limitations between workstations, and usually have limitations on how far apart the first and last node can be. limitations on distance on a low speed lan will vary according to the speed at which the equipment is sel if the pcs are in different buildings, certain networks may not work at all. it is suprising how far apart two terminals can be if the wiring has to be run all over the building. distances must be measured accurately before the hardware is bought since different networks use different hardware, the life of a network can be increased by choosing one which allows one to change from one brand of software to another without changing the hardware. ethernet cards are used by several lans as well as serial cards. the reputation of the company is another good indicator in making the choice of network. if nothing else, having a popular network insures that one may be able to obtain help from another local user. what might you have to give up if you already have pcs and wish to install a network? some networks alter the operating system. when the next version of the operating system is released, it may not be usable. pcs may behave differently, thereby causing present users some problems. the machine which becomes the server may no longer be usable as a regular pc. what might you gain by installing a network? enhanced pc capabilities are one possible benefit some networks make it appear as if the pc has more hard disks than it realh has. you may gain electronic mail features. access to printers not directly connected to the pc is a common benefit of lans. files stored on the server may be protected by security access and logins with passwords. the abilit\ to isolate individuals' files and label printouts is another nice feature. fall 1986 12 iassist quarterly choosing a network: a case study my organization had 6 months experience using ibm xts when we made our decisions as to what equipment to buy and which network to install. the day the ibm xts had arrived no one but myself was suppose to know of their arrival. the equipment was delivered in large cases stamped "ibm" all over the sides. by the end of the first day, i had received 20 requests for keys to the room where the xts were placed. these 20 people were all interested in doing free word processing: a clear hint that word processing would a problem. other management problems quickly emerged. floppy disks were being stolen. faculty accidently reformatted the hard disks on a weekly basis. bootlegged software appeared on the hard disks even though we had instituted strict rules against its use. many policies had to be hurriedly created in order to solve these problems. we decided to buy ibm microcomputers and connect them with a 3com lan. the lan was viewed as a potential solution to many of the problems mentioned earlier. we had several users interested in using social survey data on the network, and therefore decided to try to obtain as high a transfer speed as possible. because of "ownership" problems experienced with more than one person using the hard disks, we needed a network which would have access passwords for each user, and provide exclusive control over certain administrative functions. we also wanted to enhance the functions of the network pcs, since we were buying expensive network hardware, rather than extra disk drives and printers. we purchased ibm single drive pcs with 512k of memory, color monitors, 8087 math coprocessors, and a multifunction card which supplied each machine with a serial porl, parallel port, and clock. the peripherals included a 6 pen table plotter, an ink jei plotter, a digitizer, and a dot matrix printer. we decided to buy a 3com network because it satified most of our criteria. it included a $1000 ethemet board and software for each pc (these now cost about $600). in addition, we had to buy thin ethemet cable and three pieces of software: file server software ($750), mail software ($750), and print server software ($500). 3com met our criteria in the following ways: a. it is very fasl it transfers information at 10 million bits per second. this was much faster than most other networks. b. it offers excellent security. each user has an account and a password. files belonging to one user can only be access or altered by another user if the owner gives permission. the adminisiative software and system library are protected from accidental erasure and the administrative menus are available only after a control password is repeated. c. the network dramatically enhances the performance of each pc. each individual machine has only one disk drive, but when logged into the network, performs as if it had 4 additional hard disks. each hard disk is emulated by the network and contains only information stored there by the owner. (each user creates his own "virtual" disks.) d. the network commands are separate from the operating system, not patches to it. they are easy to use and there is a good help feature. e. the network automatically captures all prim requests, including screen dumps, and directs them to the system primer where they are printed with a banner page on the front displaying the owners id. r the electronic mail facility is elegant and fall j 986 iassist quarterly 13 has many features. users cannot control or contact another user's tenninal except through electronic mail. we viewed this capability as a problem and were reheved that it was not available on the 3com network. at the time that we made our decision, 3com had the largest number of working networks in service, and after talking to several satisfied owners we felt that the network was reliable and that there were many possible sources of information and assistance. the installation of the network was straight forward and the installion instructions were ver\' clear. what were the negative aspects of 3com? the following: a. it is very expensive. we spent twice as much on 3com as we would have spent on other networks. b. 3com allows any user to generate another account for himself only two individuals have discovered this feature, which is apparently an oversight in the networking administrative software. c. 3com allows any user to delete another user's account but only if the account has no stored files. in practice, this is not as much of a problem as it may seem. every user on our system has a mailbox file, and therefore every user has at least one file. the uses of a social science microcomputer lan the social scientist can take advantage of a lan for both instructional and research applications. on the research side, a lan provides the user with a larger, embellished machine on which to perform analyses. access to 4 virtual hard disks allows versatility in the organization and storage of text and data. the speed of these virtual disks can increase performance when using software that accesses multiple files simultaneously. (there is no need for one of these files to be on a floppy disk.) in addition, software can be stored on the system disk and many users can use these libraries simultaneously. the lan workstation functions like a slow minicomputer rather than a single microcomputer. many of the frustrations involved in using a microcomputer are alleviated upon installation of the lan. the communication capability is a hidden blessing. leaving messages for other members of a research team becomes an addictive, useful tool. on the instructional side, the lan allows a single copy of assignments and data to be stored, instead of one on each machine. again, mail allows faculty and student interaction outside classroom or office hours. the software available for microcomputers is often easier and more fun to use than minior mainframe software. lan software is almost as elegant as mainframe software. it is also important to remember that microcomputers are now firmly entrenched in the business community. the microcomputer exposure that the students receive pays off not only in the classroom but also later when they are in the job market many interesting applications and programs were developed on our insiruclionai lan. the departments of politicai science, public affairs, geography, demography, computer cartography, economics, sociology, social work and psychology all took advantage of the lan. fall 1986 14 iassist quarterly in political science, a simulation was constructed to illustrate how community political pressure causes the development of a 911 response system. economics constructed telecommunication software to enable the lan mictocomputers to act as elegant graphics terminals interacting with large economic time series databases stored on the mainframe. in demography, population simulations were developed which graphically displayed population pyramids. geography/cartography developed a continous shading choroplethic mapping program. we trained 5 social work faculty members to leam basic so that they could program the ibms they had at home. there are many other experiments and many faculty have increased their miaocomputer skills. problems in using a social science microcomputer lan there are two major classes of problems with a lan: problems for the users and problems for the administrators. novice users are in need of extra support to help them begin. purchasing good computer tutorials for the software is a wise investment, as is creating locally written documentation which highlights what options to leam first and how lo use the help features of each program. the major problem we encountered was that the operating system for ibm pcs is terse and quite frightening to many novices. our solution was to obtain a gram to hire a programmer, for a summer, to write a menu driven control program for the lan. each user merely inserts a disk and turns on the microcomputer. from that poini on, a series of menus asks the user which programs he wishes to use. after he is finished using the selected program, the menu system again takes control, cleans everything up, and logs the user off the network. it is an elegant piece of software which works not only on a 3com lan but also on any single ibm xt computer. the software also allows instructors to create class menus which providing their student with access to only that software necessary to complete their class assignments. there were more administrative problems than user problems. the lan demands a great deal more adminstrative attention than i had envisioned. there are constantly updates to software to be installed. to install software requires taking the lime to become familiar with the program, its organization and its options. the installer is the first source for inquires about which version of the software is available, and is this or that option available. backups must be made of the system disks on a regular basis. if the hard disk needs to be reformatted, it may be necessary to invest several days of work to rebuild it, and confirm that it is now working properly. hard disk management is also time consuming. reorganizing storage on the disk will be mandatory several times during the first few months of operation. just to decide what software to buy can use days of time. there is no repair service that can be called in if the network stops working. if you are successful in developing a small lan, you will probably immediately be asked to build a deliver>' lan large enough to handle large classes. everyone will demand use of the equipment for word processing. if the resources are available, there is little reason not to provide this service, but if resources are limited, one must beware of setting an unfortunate precidenl everyone will want a key to the facility. but, one unlocked door can result in the loss of your entire lan to thieves. insurance on the equipment is a good investment, if the budget allows. we had noi anticipated the legal problems of software licenses and software theft it is the administrator of the lan who is primarily responsible for upholding software contracts, as well as making sure that state, local and fall 1986 iassist quarterly 15 university laws concering software theft are upheld. when reading software contracts, beware of clauses prohibiting use of the software on a lan. several major software producers have such stipulations. if the license states that the software can only be used on one machine, indicate to the users which" machine they must use to run that software. if possible, use of that software on more than one machine at a time shoiild be hindered. when buying software that will be used by several simultaneous users, one should buy the legal number of copies or inquire about lan versions of the software. the laws which concern software theft should be posted in the room containing the lan. it is important to ensure that siaft uphold these laws and that they project the right attitude about theft to the users of the system. do not turn your back on illegal copying. if someone is caught, it is the administrator who will be indicated as legally responsible for the protection of the software license. c. plan before you buy. carefully. identify needs plan to avoid obsolecence. try to purchase from companies which will help the network evolve and expand, especially as national standards for network structures appear. security should always be a primary consideration in developing budgets, buying equipment, and choosing space. beware of setting bad precedents. ask the network salesperson for a list of software and/or hardware that not work on their lan. having installed a successful lan, what next? problems versus uses trouble? is a lan worth the definitely! with proper plarming. the problems can be reduced lo a minimum and the value of the lan to the users increased. some steps which will streamline the installation of the lan: a. hire a lan manager. if the funding is available, this will have a dramatic positive effect on the installation of your lan. a bigger lan!! because of the ease of use and friendliness of microcomputers and the added utility of the lan, users will demand faster machines and more software. also, if you do not have a closed community of users, your lan converts will encourage others to use the system. use in an instructional setting can grow exponentially with just a few moderately sized classes moving their work onto the lan. my group is installing a 22 ibm at lan using the ibm ring network this summer. again a grant from ibm purchased the equipment. the experiment starts over again, but on a larger scale!! b. take advantage of student employees. especially in the beginning, the lan will involve a number of small tasks which are time consuming. student employees offer talented help at an affordable price. fall j 986 16 iassist quarterly references: bergin, thomas j. establishing a computer laboratory to support the teaching of the social sciences. social science microcomputer review 3(1): 14-27, spring 1985. chapman, dale. campus-wide networks: three state-of-the-art demonstration projects. t.h.e journal 13(9):66-70, may 1986. derfler. franl recent changes on the pc lanscape complicate buyers' choices. pc week sec. 1 l(43):69-82. october 29, 1985. jenkins, avery. with icons and menus, lan operating systems meet the demands of both file servers and network users. pc week 3{7):61 ff. february 18, 1986. lily. susan. vendors ready network versions of software. pc week 2(25): 102, june 25, 1985. lily, susan. ibm token ring opens doors for independent net vendors. pc week 2(51): 113, december 24-31 1985. rauch-hindin, wendy. universities are setting trends in data commimications nets, pp. 202-212 in: davis, george r. (ed.) the local network handbook. new york. n.y.:mcgraw-hill. 1982. thurber, kenneth j. and freeman, harvey a. the many faces of local networking, pp. 222-230 in: davis, george r. (ed.) the local network handbook. new york, n.y.:mcgraw-hill, 1982. way, dale. build a local network on proven software. davis, george pp. 230-233 in: davis, george r. (ed.)the local network handbook. new york, n.y.:mcgraw-hill. 1982. fall 1986 cd-rom publishing: review, developments and trendscdby paul t. nicholls & douglas g. link' social science computing laboratory and school oflibrary & information science the university of western ontario introduction: kuhn (1970) argues that major advances in science are not evolutionary, but revolutionary; they involve an unexpected change in perspective. the change does not abandon the previously valid model of research, but it establishes an alternative approach that often yields better results. academic based cd-rom publishing has the potential to influence dramatic changes in perspectives of education and research. for example, the university of california at irvine has produced a cdrom disc called thesaurus linguae graecae, a database, when complete, that will contain all the greek literature from homer in the eighth century b.c. to close to the sixth century a.d. this disc is an integral part of the ibycus scholarly computer (isc), a tool which is revolutionizing classical studies research. researchers, archivists and administrators have been assembling databases of encyclopedic proportions for decades, but the technology for cost-effective, do-it-yourself publishing of these enormous research databases have only just begun because of cd-rom technology. cd-rom information pubhshing has been the primary domain of commercial publishers. it is they who have successfully raised cd-rom to its present level of market appeal. products such as the educational resources information centre (eric) database. the new grolier electronic encyclopedia, oxford english dictionary and compton's multimedia encyclopedia stimulate the imagination of educators, researchers and students alike. business, government and libraries have been the first to embrace in-house cd-rom publishing as a cost-effective alternative for distributing and improving access to speciauzed textual, numeric and image information databases. the support market for in-house cd-rom publishing has given rise to companies such as innotech inc., meridian data inc., dataware technologies inc., onhne computer systems inc., oftim corporation, and knowledge access international. these companies specialize in providing professional cd-rom product development services and the sale of complete turn-key systems for supporting in-house cdrom publishing. cd-rom technology and publishing software is evolving to a point where an individual can actually design. build and produce a cd-rom disc with a home computer. at the sixth international conference & exposition on multimedia and cd-rom in san jose, california, sony and phiuips corp. announced their "orange book" standard which defines a way for worm drives to write to cd-rom format. users of drives based on this standard will be able to create a cd-rom with a worm drive and also play existing commercial cdrom discs. jvc information products will have available by the fourth quarter of this year, the first 5.25 inch, half-height, write-once cd-rom drive based on the new "orange book" standard. the new drive is expected to be oem priced at si (xx). cd-rom technology cd-rom technology emerged in 1983 as a joint effort of phillips and sony, who demonstrated their first cdrom drive in 1984. the first commercially available cd-rom database, bibliofile, appeared in 1985, and the number of available titles has been increasing exponentially ever since (optim 1990). carlos cuadra (1991) has documented the vigorous growth in portable, as opposed to online, databases of all types, including cdrom, magnetic tape, bernoulli cartridge and floppy disk: one year ago, online databases outnumbered portable databases by a 7-to-l ratio. the ratio is now about 3to1 , and closing fast these figures do not take into account the growing number of portable databases that are being produced or internal use, rather than for sale commercially. these numbers may be growing even faster than the commercial products... cd-rom in particular is responsible for much of this growth, and for several good reasons, not least of which is the medium's prodigious storage capacity in relation to other portable media: "while some high-density magnetic floppy disks hold an impressive one megabyte of data, a similarly sized optical disk usually holds 600 times this amount (lawrence 1990)" a 1990 survey conducted under the auspices of the canadian library association's cd-rom interest group (fox 1990) disclosed that about a third of canadian lassist quarterly libraries of all types had already implemented or ordered cd-rom systems. in the case of academic libraries, this proportion was already 44%. annual surveys by oclc in the united states have disclosed substantially higher rates of implementation in that country, approaching 100% in the case of academic libraries. cd-rom has found many types of applications in canadian libraries and other organizations (optim 1990): memorial university of newfoundland and the university of guelph have their entire library catalogue on cdrom. ford new holland uses cd-rom for its auto parts catalog. statistics canada offers bibliographies, directories and census data on cd-rom. the department of fisheries and oceans made its internal database of reports on fisheries and aquatic sciences available to the pubhc on cd-rom. industry growth the optical publishing association (columbus oh) estimates that cd-rom revenue from inhouse and commercial pubhshing and drive sales to be at least uss571 million in 1989, up 41% from uss406 million in 1988 (cd-rom 1990). "by 1993, according to market researchers frost & sullivan, the combined european market for optical disk drives and the optical media will reach us$900 million, up from uss37 million three years ago. (lawrence 1990)" according to information in tfpl publishing's annual cd-rom directory, the number of companies involved with cd-rom activities has risen from 48 in 1986 to 736 in 1989 and 1,840 m 1990. available cd-rom titles a survey of commercially available cd-rom titles was conducted in mid1990 based on the major printed directories to the medium and using comparative data from three previous annual studies (nicholls 1991). as of mid1990, 1,025 commercial cd-rom titles were identified. at the growth rate that has prevailed over the past few years, well over 2,(xx) titles are likely available at this lime. cd-rom. social sciences actually have a somewhat greater share, due to the business and legal databases that (along with medicine) account for 30% of all cd-rom titles. the majority of cd-roms (63%) are updated annually or even less frequently. the relative proportion of frequently updated titles, (quarterly, for example) has been declining steadily since 1987. this trend is related to the rise in the numbers of source databases, which require less frequent updates than index or abstract databases. of all cd-rom titles, 93% will run on an ibm pc/xt/ at/ps2 or compatible system, while only 13% will run on the macintosh. actually, 6% of the titles were designed to run on both platforms, and this proportion is now rising very rapidly. similarly, although only 15% of the titles incorporated multimedia content, this proportion is rising steadily. in the near future, multimedia titles in the new related optical formats cd-rom xa, cd-i and dvl will also begin to proliferate. standards standards were a serious problem until the recent adoption of the iso 9660 standard, as well as the appearance of the microsoft cd-rom extensions (mscdex) to dos. it is now possible for industry observers such as nancy herther (1991) to view standards in a different light: "standards continue to be the factor buttressing and fostering growth for cd-rom." nevertheless, other observers such as john dvorak (1991) points out that hardware compatibility problems definitely remain so far as 486 machines are concerned. software proliferation has raised another type of standards problem, the lack of consistent search interfaces. more than 1(x1 search software packages are now employed on cd-rom products, although a core of 1520 of these account for perhaps 60% of all available products. although it is not clear that a standard interface is actually desirable (at least not until we find a standard user and a standard database) this situation certainly has raised problems of several types: almost half of the titles identified were source databases (containing full text, numeric data, computer software, images or similar data) with indexes/abstracts and directory-type databases accounting almost equally for the other half. the overall proportion of indexes/ abstracts on cd-rom has been declining steadily since 1987, while the proportion of source databases has been rising steadily. the general/humanities, social science and natural science subject areas are represented almost equally on after the tenth disk 1 wondered if the vendors are missing the point. gads. each one has a different file format and a different search engine. each vendor wants you to dedicate a subdirectory and a bundle of megs to a given disk, and each disk requires a screwball path statement, a driver, or some memory-chewing crapola to operate properly. if you routinely used more than a few cdrom disks, your system would be chewed up by ancillary files and memory hogging tsrs. (dvorak 1991) summer 1991 these problems pose a particular challenge to multiple database workstations, public-access cdrom services and end-user instruction. the networking of cd-rom databases also carries its own problems; however, when networking is practical, it can suffice to greatly simphfy user access, addressing many of dvorak's expressed concerns. multimedia formats meridian data, california based developers of cd-rom publishing systems, with sales of us$7.2 million in 1989, expects that about a quarter of their clients will begin using their new cd-rom xa development system, introduced in 1990. (meridian 1990) the multimedia industry is expected to develop similarly to cd-rom: "two years following the sale of production units to developers, the end user market should emerge." (meridian 1990) multimedia applications require more powerful (and expensive) hardware than text-based cd-rom applications. the cost of a multimedia workstation with a 386 microprocessor, 40mb hard disk, optical drive, vga card and audio capability, estimated to be about us$3,000 in 1992. (cd-i 1990) cd-rom in libraries & information centres it is clear, particularly as earlier barriers to implementation have disappeared or decreased in importance, that libraries have implemented cd-roms extensively and will do so more extensively yet. libraries were one of the earliest markets for cd-rom products and continue to be one of the largest, as nelson (1991) observes: the face of academic computing is changing; it is now very much data driven. the most basic computeroriented tasks that teachers and researchers must now involve themselves with is the collection, analysis, annotation and organization of information. richard l. nolan suggests that we are in a economic transition, one which is taking us from being an industrial economy to an information/service economy. nolan (1990) argues that higher education must add information technology to its strategic equations for the 1990's. wether you agree with nolan that we are in the midst of a new industrial revolution, fuelled by information technology, there still remains an overwhelming awareness of the importance that must be given to information literacy for our future economic success. the american library association's 1990 g. k. hall award for library literature, was awarded to patricia senn breivik and e. gordon gee for their book, "information literacy: revolution in the library." the concept of information literacy involves examination of how information seeking an evaluation skills intersect with the need to understand and use information technology including software; systems design and hardware; access to information beyond that typically stored or accessed by libraries; understanding the complex interaction among stored or remotely accessed information; and finally one's own production of new information. libraries and information centres have the opportunity to be the facilitators of inhouse cd-rom publishing technology; however, useful and innovative applications must originate from all members of the academic community. for example, the following is a sample of cdrom applications that might originate from academic settings: it is no wonder that libraries lead the marketplace in cd-rom product acceptance. once they understood the potential of this precocious six-year-old technology, librarians resolutely tackled the several impediments especially its high price tag to full implementation of cd-rom into library processing and public services units. in doing so, they have rightly been accorded recognition as technological innovators. archaeological artifacts databases comprised of video images and full-text annotations. rare or significant local collections protected in backroom museums could become transportable and accessible for research and instructional purposes around the world. documentation, research and instruction of native heritage. databases comprised of scanned images, poems, songs, legends, symbohc writing, full-text analysis and audio tracks. it is also clear that these systems have received an enthusiastic acceptance by library patrons and end-users, and that the technology is not likely to be displaced in the near future by some alternative new medium. what remains unclear is whether the libraries and information centres should take the initiative to establish inhouse expertise and facilities to assist with "do-it-yourself cdrom publishing. the potential for these applications are universal and just beginning to emerge. geographic information databases which cater to municipal and rural communities where universities and college reside. virtually all information which has an underlying spatial aspect to it is a candidate. for example. master's level projects which create geographic image/text databases related to local urban planning issues, commerce, recreation, tourism, health, pollution, housing, population, etc. may find cd-rom an ideal medium. lassist quarteriy full-text databases comprised of large transcription projects involving international collaboration of scholars and editors. for example, the international scholars working on the transcription of the complete works of jeremy bentham (1748-1832). preservation and improved access to local history through the construction of specific photo image and full-text news databases. students in history and journalism may find their compilation and presentation of fulltext and image based research to be more effective, accessible and better suited for secondary analysis. as the cost of inhouse cd-rom publishing continues to fall and the tasks associated with inhouse publishing become even more pedestrian, we will see a proliferation of cd-rom information, research and instructional products from academia. 1990. nelson, n.m. "cd-rom growth: unleashing the potential," library journal 116(2): 51-53; 1991. nicholls, p.t. "a survey of commercially available cdrom database tides," cd-rom professional 4(2): 1991. seymour, j. "forecast '91: will the cd-rom market finally explode?" pc week 7(dec 24): 47; 1990. nolan, r. l. 'too many executives today just don't get it!" cause/effect, volume 13 number 4: 5-11; 1990. ' presented at the lassist 91 conference held in edmonton, alberta, canada. may 14 17, 1991. references campbell, b. 'the cd-rom industry in canada: trends, companies and products," optical infonnation systems 10(2): 99-103; 1990. "cd-i, dvi vie for spouight at microsoft cd-rom conference," idp report 1 l(march 9): 1-3; 1990. "cd-rom revenues reach s571 million," idp report ll(mar9):3-4; 1990. cuadra, c.a. "portable databases: bridging the gap between online and local databases," information today 8(1): 15-16; 1991. dvorak, j.c. "cd-rom: still a bust," pc magazine 10(1): 81; 1991. getz, m. "economics: storage technologies," the bottom line 4(1): 25-31; 1990. helgerson, l. "when it comes to cd-rom, it's still a dos world, but it's changing," cd-rom enduser 2(4): 70; 1990. herther, n.k. "cd-rom today: unanswered questions and unending possibilities," cd-rom professional 4(2): 10; 1991. herther, n.k. "cd-rom year six: change, growth and stabiuty," laserdisk professional 3(2): 5-7; 1990. lawrence, a. "optical discs finding market; still some hurdles to overcome," financial post (june 9): 38; 1990. "meridian data expects one-fourth of customers to upgrade to multimedia," idp report 1 l(may 5): 9-10; summer 1991 35 microsoft word 0_vol 41 1-4 lead.docx 1/2 rasmussen, karsten boye (2017) editor’s notes: the data is out there – just like the truth! iassist quarterly 41 (1-4), pp. 1-2. doi https://doi.org/10.29173/iq15 editor's notes the data is out there just like the truth! welcome to the first issue of volume 41 of the iassist quarterly. it has taken extra time for this issue to appear. the cause of this is not that we have been extra lazy. the paradoxical cause is that a great many people have been extra busy. thanks to the team of people in the editorial group of the iassist quarterly and not least the great help from sonya betz working as digital initiatives projects librarian at the university of alberta libraries in canada, the iassist quarterly has now moved to the open journal system (ojs) at the university of alberta. we believe this shift is going to benefit all stakeholders of the iq. it is mostly the inner workings of the production that has changed. as a potential author you are still very welcome to mail the editor. the first issue of volume 41 (2017) at the same time becomes the last issue. in order to get close to the real time we are catching up by jumping three issues. therefore, this issue is labelled as vol. 41:14 of 2017. next issue will be 42:1 of 2018. the first article concerns data for published articles in journals. the paper ‘journals in economic sciences: paying lip service to reproducible research?’ is by sven vlaeminck and felix podkrajac. sven vlaeminck works in research data management for zbw – german national library for economics / leibniz information centre for economics in hamburg, germany. felix podkrajac is an academic subject librarian at the library and information system of the carl von ossietzky university of oldenburg. some economic journals have a 'data availability policy', and vlaeminck and podkrajac are presenting a study of the compliance of actual research to such policies. half a century ago economic journals were mostly theoretical but now most articles are data-based. in their introduction, they present a good overview of literature and, among others, they cite mccullough (2009). i think it is appropriate to repeat part of that citation: 'replication ensures that the method used to produce the results is known. whether the results are correct or not is another matter, but unless everyone knows how the results were produced, their correctness cannot be assessed. replicable research is subject to the scientific principle of verification; non-replicable research cannot be verified.' therefore, not only do we need to know the methods, we also need availability and access to the data. their thorough investigation includes clear definitions and insights into methods and data. vlaeminck and podkrajac find that the overall compliance rate is a little below 50 percent and is very variable; some journals as we would expect when having a data availability policy reach compliance rates of 100 per cent, while 10 journals have compliance rates below 20 per cent because their data availability policy is voluntary. i don't need to add that the authors have made their replication files available. the second paper in this iq issue is titled 'designing the cyberinfrastructure for spatial data curation, visualization, and sharing' by the authors yue li, nicole kong, and stanislav pejša. all three authors are working at purdue university libraries as respectively gis analyst, assistant professor, and data 2/2 rasmussen, karsten boye (2017) editor’s notes: the data is out there – just like the truth! iassist quarterly 41 (1-4), pp. 1-2. doi https://doi.org/10.29173/iq15 curator. they argue that spatial data is an important component in many studies and has promoted interdisciplinary research development. in their development project at purdue they have streamlined spatial data curation, visualization and sharing by connecting the institutional research data repository with the library’s gis server set and spatial data portal. the purdue university research repository (purr) supports data deposits for all purdue faculty, staff, students and their collaborators. institutional repositories are becoming more numerous in universities. because the original dataset owner is often among the first users of the data, a local data repository makes local access an easy solution; however collaboration is important and standards must be in place to gain that benefit. the authors' focus is: 'to increase discoverability beyond the institutional data repository, we expect to ingest the spatial data into our geodata portal'. the last paper is also addressing data management. the paper 'research data management: a proposed framework to boost research in higher educational institutes' is a collaboration between bhojaraju gunjal at central library of the national institute of technology, rourkela, odisha, india, and panorea gaitanou of the department of archives, library science and museum studies, ionian university, corfu, greece. they begin with an abstract of research data management (rdm) issues where they promise 'a detailed literature review regarding the rdm aspects adopted in libraries globally'. their overview provides many links and resources for several aspects of rdm before they describe the implementation of rdm processes at the national institute of technology, rourkela. taking good care of data is worth writing articles about, and is also worth writing books about. in this issue we present two book reviews: 'databrarianship: the academic data librarian in theory and practice' by lynda kellam and kristi thompson is reviewed by chubing tripepi of columbia university, while 'the data librarian’s handbook' by robin rice and john southall is reviewed by ann glusker of the national network of libraries of medicine. submissions of papers for the iassist quarterly are always very welcome. we welcome input from iassist conferences or other conferences and workshops, and from local presentations or papers especially written for the iq. when you are preparing a presentation, give a thought to turning your one-time presentation into a lasting contribution. we permit authors 'deep links' into the iq as well as deposition of the paper in your local repository. chairing a conference session with the purpose of aggregating and integrating papers for a special issue iq is also much appreciated as the information reaches many more people than the session participants, and will be readily available on the iassist website at http://www.iassistdata.org. authors are very welcome to take a look at the instructions and layout: http://iassistdata.org/iq/instructions-authors authors can also contact me via e-mail: kbr@sam.sdu.dk. should you be interested in compiling a special issue for the iq as guest editor(s) i will also be delighted to hear from you. karsten boye rasmussen november, 2017 provider sophistication versus user simplicity: european servicing througli bridging the gap by per nielsen ' danish data archives odense, denmark introduction: working on the right problem when i first entered the "data archive movement" (february 1, 1974), everybody seemed to be very preoccupied with the creation of advanced software for the mainframe: report generators (even though there was littc to report on), search systems (even though there were few surveys to search among), and data base systems. later on, we realized that the defacto standards were developed at larger organizations either within (osiris from the icpsr, spss from norc) or outside (spss inc., sas in raleigh, and all the other business firms all over the marketplace) "our world" of social science research institutes and data archives. i am sure that we spent quite a lot of time during the early years working on the wrong problems given the needs of the time; however, the work initiated the habit of engaging in cooperative projects among the european archives. this good working habit has survived ever since; as institutions and as individuals alike, the european archives have a close and pleasant collaboration program, ranging from responding to incoming servicing requests over staff-relevant exjjert seminars^ to business meetings once or twice a year. one topic that was hardly ever discussed during these early years (even internally among data archivists) was the actual level of servicing provided by the data archives during a given year; it was the tacit understanding that the actual (quantitative) level of servicing was not a topic that would underline the raison d'etre of the new data organizations, data archives, and data libraries. now, almost twenty years later, we have vast amounts of data sets in custody, and the demand for data for secondary analysis has increased dramatically, not least with the advent of the pc. unfortunately, we are not so wellprepared to meet this demand as one would expect, given the advanced techniques developed and applied in the take-off phase. in babel-like europe, there are still huge obstacles to a free data how from data providers to data users. we are working on these, but there is a long way to go. at a recent seminar, two american scholars' told us (the data archive professionals) that we had been overtaken by the "ordinary" library people in terms of computer mediated communication. shame on usi the real problem during the sixties and seventies, looking in the rear-view mirror, was to localize and collect data, to develop standards for documentation, to teach data collectors "sound methodological/technical practices" during the data generating process, and to store the data safely with a long term archival perspective in mind. we did perform all of these tasks once we found out that we could import most of the software tools from the outside; however, then we worked little on the search systems and the other advanced tools that successively became relevant as we had thousands of data sets in our holdings; in many archives, the then advanced retrieval systems of the seventies were maintained and slightly developed; some of them are still in use in the early nineties. the "right problem" is simple: to find relevant data sets that meet the specifications of a user and transfer the data to that user with the shortest possible elapse time and with the least possible input of human and machine resources in both ends of the communication line. the topic: remote access and new user services my impression is that american scholars find it very difficult to get an overview of the european data marketplace; honestly, it is sometimes difficult even to europeans working in the market! in this paper. remote access is interpreted as "access from north america to european data" rather than looking at the technical notion of remote access (i.e. running jobs on a distant machine). the underlying philosophy is that the user is substantively oriented rather than technically fascinated; the user would rather have data available in a known environment than shuffle around in dozens of differently functioning systems to dig out what (s)he needs. given this interpretation. new user services will, to a certain extent, become equivalent with present user services. some 80% or more of the dda-servicing is domestic (i.e. national), so the new services will be developed for the national market first. consequently, i can claim to be knowledgable about the danish situation, only; and, being a native of a small country with a peculiar language, i realize that this situation may be of lassist quarterly minor interest in north america. finally, my personal feeling is that a paper on remote access and new user services in europe would be much better in 1995 (third year of the open internal market, cf. below); after a couple of decades with consolidation on the (national) archival side of the data organizations, we shall now move into an era where the (national and international) servicing aspects gain more weight. this gradual shift in emphasis is (among many other indicators) reflected by the fact that the ecpr council has accepted the theme of integrating the european data base for the ecpr joint sessions of workshops 1992 (limerick, ireland)*. the scene: the integrated europe (united states of europe?) being in north america, a notion about the eecgenerated phantom of european integration is perhaps necessary. with the introduction of the open internal market (end of 1992) and the cuaent plans regarding an economic and monetary union (three stages during the nineties), with or without a political (and maybe even foreign policy and defence) union, the brussels establishment (especially the commission) has succeeded in moving some frontiers of thinking, especially in business and maybe even more so in north america and japan than in europe! even though the united states of europe is being discussed, by supporters and adversaries alike, this vision will remain a phantom to most europeans for another couple of decades. in the perspective of integrating the european data base, the differences of language, ethnicity, culture, rehgion, wealth and political and social science tradition represent obstacles to the free data flow; the same is true of the differences in economic development and technological sophistication a gap which is more evident in the east-west dimension than on the north-south continuum. furthermore, it is important to notice that the cooperation between european research institutions (and hence social science data archives) has never been limited to european community members; it has been open to any institution with the interest and the capacity to participate. generally (and this is especially uue with respect to the data archives) the nation-state has been the represented unit. it has been difficult in some countries to find the relevant (i.e. nationally representative) institution; and this problem area will return to the scene as new states emerge (east) and as some countries develop specific instiuitions to deal with data archiving (e.g. specialized historical data archives in the larger countries). finally, with a landscape of "peer partners" among eurof)ean archives, it has not (yet?) been possible to establish the european data archive * that might be the gross dealing agent in europe, comparable with the icpsr in north america. needless to say, it would be easier from the outside to address one central agent that collected and disseminated all major european data sets of broader interest our response so far to this demand is: contact any one of the archives, and they will (ideally!) let the message pass to everybody else. as a matter of fact, this procedure has proven to be efficient on a number of occasions; but please be accurate and specific when elaborating the request! which are the heavy resource demanding servicing tasks? in the american context, the european archives should be understood as an amalgamation of the gross dealing agents (e.g. the icpsr) and the local retail servicing facility (university data libraries, state data archives, etc.) covering the whole set of archival as well as servicing procedures, we have a good overview of the costs involved in different parts of the whole process. at the archiving end, it is the cumbersome data processing to a standard archivalformat that digests heavy resources. even though standards may vary from place to place, most of us want to produce standard codebooks (in a format derived from the osiris dictionary-codebook format type 3 to have a clean ascii character set) most of us want to have a good study description (sub-structured in a more detailed way than the icpsr free text study description), and most of us want the data to be immediately accessible for the major analysis packages like sas and spss. at the servicing end, the gross-dealer functions take little time: users requesting specific data sets that have already been processed to the standard archival format can be serviced from one day to the other or even within hours; they will receive easily accessible data and can start off with their analyses immediately, cf. below. the resource-consuming customers are those that want to obtain data that meet certain search criteria (often too vaguely defined!) and this is especially time-consuming if the user wants to perform cross-national comparisons and/or if non-standard data sets are involved. obstacles facing the user of european data below is a list of existing obstacles that the data user may face when trying to get hold of data relevant for (cross-national) analysis. imagine that the user asks the local national archive in one country to faciutate access to data from several european countries, relevant to a certain topic; what is the process ahead? summer 1991 1 . the local archive sends out a "search warrant" to all other european archives. this can be done quite quickly via e-mail. however, most european countries still do not have a data archive; and some of the existing national archives are heavily under-staffed, so that the requesting archive will get a late reply or even no response at all. result with good luck, relevant data will be found in 3-5 countries, only. 2. some of the data sets localized may have been produced by central statistical offices (csos) or administrative agents; in these cases, data may not be available at all or the user will have to go to the country in question because data export is not allowed. another possible obstacle is an embargo period for secondary analysis, introduced by the primary investigator. 3. some of the relevant data sets may have an inappropriate format, i.e. they are not immediately available for analysis with sas or spss. it may take months or even years until the data archive has improved the technical availability of the data set. 4. the data user may have to sign undertakings with each of the primary investigators before the data can be delivered for secondary analyses. 5. the documentation of the relevant data sets may be available in the local language, only; and this, of course, is the rule rather than the exception. 6. remedying obstacles 1-5 may cost real money that the user may not have available. 7. the (few) data sets that actually pass the obstacles 1 6 will now be sent to the requesting archive; they in turn pass the data sets on to the user. 8. if, during the secondary analysis process, problems arise with the data or the documentation, the trouble-shooting will be quite difficult also. the above picture of the obstacles is quite pessimistic, some european data archive people might say; unfortunately, my feeling is that it is a realistic one. there is a long way to go for the archives united in cessda* (committee of european social science data archives) before we have an integrated european data base. with 15 years of fieldwork, we have not yet fully implemented the visions presented in a paper at the cessda founding meeting by one of the fathers of the data archive movement, the late professor stein rokkan: "our basic philosophy is very simple: we do not believe the archival movement in europe will get anywhere unless there is a real break with the tradition that archives are there simply to store, clean and reformat separate data sets. the future lies with active reorganization of data : linkage across files, build-up of time series sets, preparation of handy packages for use in the classroom, integration of packages with better computer routines for graphic display, cartography, visual modelto-data fitting." (rokkan's underlinings). remedying the obstacles: remedying the obstacles: state-of-the-art and planned activities we are doing our best to try to smoolhe the facilitation of european data to the user. let's reiterate on the obstacles mentioned above and see what is being done and what can be done within each of the "obstacle fields" identified in that section. doing so, we shall look at the european level first, and i shall add a few comments about the danish situation which 1 know best! localization of relevant materials most archives do have printed catalogues of holdings, with multiple indexes, from which you can figure out whether the other archives do have relevant data sets in a given field. however, the printing is expensive, and the paper-bound inventories tend to become outdated quite quickly both with the acquisition of new data and in terms of "processing classification," access restrictions, etc. it is possible, of course, to acquire the machine readable text from the catalogues of each archive and search these in your own retrieval environment; but this does not solve the updating problem. consequently, the most evident solution is either to integrate the primary cataloguing at one central location or to search in the catalogues of the other archives, via telecommunication, at their own computer installation. the first path, integrating the catalogues, has been worked on with some energy. the commission of the european communities (cec) had actually granted money that would allow catalogue integration in a project with dda, esrc-da, star and za as the major project partners. this project stranded because the cec demanded that the resulting data base had to be commercially viable after the 2-year project period. (it probably would not have been after 10 years; but the cec bureaucrats, preoccupied witii "commercialization," are not at all sensible to the special problems in the academic sector!) the second path, searching via telcommunications, is {assist ouarteriy probably a more realistic one. given that most archives now have tcp/ip and ftp facilities available at the installation where the catalogue information is stored, the searching as well as the actual exchange of data may take place using these communications and file transfer facilities. this procedure assures the user that the most updated version of the catalogues is searched and the most processed data set is transferred. even though many of the archives have retrieval facihties that are open to the user, most searching is still done by the staff of the archives on behalf of the user; this is probably going to be the case in the foreseeable future for all other than very heavy users of data. in the case of the dda, we have one central retrieval system, ddaguide, based on the study descriptions. it is available to most potential danish customers, located at uni*c (a national computing center for research and education). however, it is not yet open to users without an account number at uni*c. in addition to ddaguide, we have several in-house search systems (some mainframe-based, others pc-based) searching the contents of the machine readable codebooks. in order to integrate retrieval at the study description and codebook levels, we have designed an integrated system and applied for money from the ssrc to have the uni*cpeople implement the system; hopefully, there will be a remote user access to this integrated retrieval system''. (on the other hand, the language stored will be danish, cf. below, where the ddaguide stores english language texts!) handling access restrictions on the data in most european countries, the access to processproduced data (administrative and statistical data) is hampered mainly due to three factors: (1) the bureacratic traditions of government and a lack of freedom-ofinformation-tradition; (2) the privacy legislation; and (3) the wave of cutting in public spending indicating that the statistical bureaux (csos) and other data owners want to sell their data rather than offer the data for free (or at a quite low price) via the social science data archives. let us take a look at a few examples: in norway, where the relations between the central statistical bureau (cso) and the nsd has been better than in most countries, the nsd can disseminate a lot of statistical data for research and educational uses; this has mainly been done with regional data, where the nsd probably has the largest collection of commune-based data in europe. on the other hand, there are severe restrictions to the access to survey-data (data on individuals); for instance, the later election studies in norway have been collected by the cso; consequently, the data may not be taken out of the country, so that the user will have to go to norway to use such data sets. in sweden, also, the election study data sets and other cso data on individuals may not leave the country; the ssd has tried to apply the lis-model (luxembourg income studies) to gain indirect user access to such data*. in the united kingdom, the esrc-da is "re-selling" selected data series from statistical authorities to the academic community. also in hungary, tarki has quite close relations with the statistical office in budapest data from the academic community are usually more easily available than data from the csos. even so, it is in some countries (for some studies) necessary to ask the primary investigator's permission. it seems to be generally accepted that the primary investigator may impose up to a 2-year embargo on the data. in denmark, the cso (danmarks stalistik) tries to make a lot of time series data banks available on a commercial basis. four rather large "data banks" are available, covering national economic time series (dstb), commune statistics on the 275 local administration units (ksdb), labor market statistics (abba), and business related statistics (esdb). however, the danish cso is very reluctant to release survey data of any kind. the big public survey organisation. the danish national institute of social research, on the other hand, generously puts all its surveys (in a defacto anonymous form) at the disposal of the dda and her users free of cost. this fact demonstrates that it is the interpretation of the data legislation rather than the acts themselves that render access impossible. handling technical problems with the data there is little standardization across europe with respect to the processing classes (cf. icpsr's classes i-iv a data class structure which is presently being revised). some archives (e.g. dda, ssd, za) follow a strategy pretty much uke that of the icpsr, having machine readable codebooks for their top-class studies. other archives (e.g. nsd, star) try to make as much as possible available as spss export files not necessarily having all information from the questionnaire (or other instrument) in the machine readable codebook. the esrc-da has realized that they do not have the resources to process iheir several thousand studies to the "icpsr class r'-level; instead, they have developed a thesaurus and apply that to append relevant search entries to the study descriptions in order to have a search base without having fully-fledged codebooks for all studies. in general, most european archives aim at making the data sets available for analysis with spss and sas. summer 1991 23 however, most archives hold many data sets that have not (yet) reached this level of processing (icpsr classes ii-iv); for some of these data sets, the user will have to produce the setup on his own, based on a card-image data set and a paper documentation thereof. at the dda, we stick to the traditional "icpsr-type" of documentation with respect to the codebook, adding a (structured) standard study description to this codebook. streamlining the data documentation and processing work (mainly based on programs such as sas, kedit and rexx in an os/2 networking environment) we can more than keep pace with acquisitions, so that the number of "non-class i" data sets is diminishing. data ownership as a restriction to remote access whereas many archives have searching facilities available for remote access from the users, few archives have the data files as such available for immediate analysis. this is usually not due to technical restrictions; rather, it is based on proprietory considerations: the principal investigator is the official owner of the data, and sometimes her or his written consent is required in each individual case before the archive can offer access to the data set itself. this type of restriction should problably be removed in the future; given the fast technological development (with communications protocols like tcp/ip and file transfer protocols like ftp on the one hand, and distribution on cd-rom or other mass storage media on the other hand), the provision of free distribution should be granted to the data archives by primary investigators. vis-d-vis the researchers, the dda has not yet found the formula that will allow access to all stored materials without prior wriuen consent from the depositor. however, access is never denied, so we consider it feasible to get an agreement about unrestricted access with most donors, once we really need that either in order to allow remote access for analysis on our computing facilities or to distribute selected data sets on mass storage devices. the tremendous language problem in europe within each country, it is considered "normal" or even indispensable that the documentation be produced in the national language; the only general exception to this rule seems to be the netheriands, where sieinmetzarchief produces spss-setups for all files in english rather than in dutch. most archives do have catalogue information available in the english language, but the codebooks and/ or the questionnaires (or other instruments of data collection) are available in the national language, only. this is an obstacle to remote (in casu foreign) access to which there is no readily available solution. everybody who has been engaged in cross-national research projects will know that it is a tremendous problem to produce cross-culturally comparable data, in part due to the language problem. this is true even in the culturally relatively homogeneous european community (reflected in the euro-barometer surveys); but the difficulties are even greater if one goes to second or even third world nations (which, for instance, the isspprogram is doing). definitely, there is not enough resources within the european archives to produce all documentation in the national language and in a world language (e.g. english). all users of european data should be aware, consequently, that they will have to be able to read the language of the nation under investigation with the aforementioned exception of the netherlands. (some scholars would argue that you would have to know some language and culture prior to engaging in quantitative (or qualitative) investigations of a specific nation anyway, we shall not engage ourselves in that discussion here.) at the dda, we keep study descriptions in both danish and english; we can, therefore, inform about our data (in catalogues and ad hoc listings of selected topics) in either language. but with the codebooks it is different; even though we have produced some english language codebooks (in addition to the danish ones) for a few frequently exported data (e.g. the continuity guide to election studies, which is also disseminated through icpsr), the bulk of the codebooks are available in the danish language, only. it may be a future project among european archives to produce english language codebooks within areas where cross-nationally comparable data sets can be "constructed" by the archives. such projects have been successfully carried out in the past; for instance, at lot of national election studies (and continuity guides based on these) have been produced in cooperation with the icpsr and are now disffibuted from ann arbor to the membership. fee schedules among the european archives in some countries, the servicing of users is free of cost (apart from "media" such as paper-codebooks, diskettes, etc.); other archives have a fee schedule the size of the fee depending on factors such as size and complexity of the data sets delivered, staff time involved, etc. sometimes there is a discriminatory pay-schedule, where students are at the cheap end and business applications in the expensive end somtimes with researchers in a middle position. it is beyond the scope of this paper to try to spell out all the fee schedules; they change now and again, so the user has to ask in each specific case. 24 lassist quarterly at the dda, servicing of archived files is free (media cost recovery is demanded if the media are not returned). also, the staff and machine time consumed by performing searches for users is free. however, special services such as "super-quick processing" of dda-studies, processing of requestor's own data sets, or translation of codebooks will have to be paid for by the requestor. actual transmission of data to the user as a consequence of an agreement between all cessda archives, there are certain rules that apply to international data transfers; the major contents of this agreement can be summed up as follows: * each national archive is the primary repository regarding data from that nation. ("fishing-zone agreement"). * all requests from within one of the "cessda-countries" should be directed through the home archive. (this seemingly bureaucratic rule is administered liberally: if a foreigner requests danish data, for instance, we would always inform the relevant "home archive" and, if they so wished, send the data via that archive). the idea is, of course, that the users should not be able to circumvent pay schedules by going abroad to the free or low-cost archives. * cross-national data sets are processed and disseminated according to mutual agreements among the relevant archives. as mentioned before, the actual transfer of an akeady processed data set is not very time consuming. however, the actual procedures may differ from place to place; needless to say, in inter-archival transfers the technicalities would be agreed upon aforehand (if they are not already known from earlier transfers). at the dda, the d"ansfer is done according to the specifications wanted by the actual user. if the user works on a mainframe, we would normally send the data to that mainframe from our central archive (magnetic tapes at a central uni*c mainframe). if the user wants to work on a pc which some 90% of users do, we will send the data on a diskette, containing a dos bat-file that will do all the work necessary before the analysis: make the necessary directories on the user's hard disk, copy the files onto the hard disk, unzip the packed files, and, if necessary because of multi-volume delivery, put split files back together. let us assume that the study number dda-9999 was requested and sent on one or several diskettes; after the user has run a dda-copy.bat job, (s)he now has the following files in a directory with the name ocdda9999: dicb9999.0si (oseris-like dictionary-codebook, ascii) data9999.0si (osiris-like data fue, ascd) osi-spc.exe (dos-program for system nie, cf. below) list9999.prt (ascn-listing fue with sd and codebook) all the user has to do now is to run the osi-spc program (a dda-utility) which will ask (1) whether the user wants an spss or a sas file; (2) which variables the user wants to include in the systems file; and (3) the names of the dictionary-codebook input file and the setup output file. when osi-spc has finished (in seconds, even with very large files), a setup is ready to build the specified spss or sas systems file with all necessary variable and value labels in place. again, the user needs to specify only the data set names before running the systems file generating job. the list9999.prt file is just a stream of lines that can be printed on any type of printer; the idea is that the user can save the money for the printed documentation if (s)he prefers to run off a printed copy instead. the supply of data on diskettes takes place from a copy of the original archive which contains the zip-files; this "archive copy" is kept on a 1 gb traditional disk on a net server, so there is very quick access. needless to say, such files can be sent over the external networks to the user instead of using diskettes; however, the dda is waiting for the os/2 version of tcp/ip and ffp which is to be in the market very soon. our experience is that users (and we ourselves, when receiving foreign data like that) are very satisfied with the present procedure. problems with data or documentation during secondary analyses it is important to know which archive is responsible for the data sets that "drift around" in the international social science community; otherwise, when errors or omissions occur during the secondary analyses, it is close to impossible to remedy or clarify such problems. a couple of examples will demonstrate this problem: summer 1991 25 parties (venstre, i.e. the (conservative) liberals) are absent from data as well as from documentation. who "produced" the error: the principal investigators (jacques-rene rabier, helene riffault, ronald inglehart); the danish data collector (danish gallup); the international coordinator (fails et opinions); the zentralarchiv (where the file is first available); or the icpsr (where the final documentation is produced)? needless to say, the error is critical to a political scientist who will use this variable extensively during the analyses. a danish user participating in the repeated international value project finds out that there are problems with the "oversampling" of young people in the 1981 value project file for denmark. there were no danish researchers involved in the 1981 value project; all you can do is to ask the danish collector (observa, which has changed name and staff in the meantime); the coordinator (british gallup); or the involved archives (in casu the esrc-da, but we got the file from nsd). it is more than likely that you can never solve the problem! (which, in this case, derives from the fact that archives entered into the process of preserving the data long after the primary investigation.) in conclusion: what is achieved, and where are we heading? in the european countries with a national data archive, a huge data resource is immediately available for analysis to everybody who knows the human language of the country; the range of users has been augmented laterally (data in schools' and easily operated analysis packages'" ) and horizontally (the concept of "social sciences" is broadening with "new" disciplines all the time.) the users want the data "packaged" to their needs, on their own computer and for their own analysis package; it is my feehng that we have accomplished this task on most levels of user sophistication: the number of users is constantly rising, and the servicing is becoming very efficient in the archives. to move from diskeaes to ftp transfer of data is a technical detail of minor importance to the user; however, it may save resources among the data supphers and make the transfer across long distances quicker. the broadening of the topical area covered by the archives is a process that may take different paths in different counu-ies. take the historical data as an example: in germany, the zentrum fur historische sozialforschung started its carceer as an independent institute; later, the zhsf was moved to form a department at the za, in the netherlands, a historical archive is on the steps; right now, however, funding is lacking. in the uk, the situation is under review. in smaller countries, it seems likely that the existing archives will cover the "new" disciplines. in denmark, data from history as well data from social medicine are archived at the dda and have been for years. the dda hosts the up-coming ahc conference 1991 to demonstrate that fact to everybody in europe". obviously, the immediate future in europe will be devoted to the project of integrating the european data base. cessda is right now "incorporating" (with a formal constitution and membership fees) as one step in that direction. projects are underway that will facilitate searching across archives. data exchange will take even less time in the future using telephone lines for transmission rather than snail-mail with data media. let the remote users come; we shall give them access and demonstrate that our services are better and quicker than ever before! footnotes: • presented at the iassist 91 conference held in edmonton, alberta, canada. may 14-17, 1991. the danish data archives (dda) is a national social science data archive established in 1973. since 1978 the dda has been located at odense university; some time in 1991, the dda is likely to be relocated, most likely to be a department of the danish national archives. danish data archives, munkebjergv{nget 48, dk-5230 odense, denmark, e-mail: ddapn(a)neuvml.bitnetor ddapn@vm.uni-c.dk. telephone: (-1-45) 66 15 79 20 x 2810, fax: (-f45) 66 15 83 20 ^ the european archives arrange so-called expert seminars for staff-members, hosted by one of the archives, once or twice a year. it should be noticed that the data archive "milieu" is more institution-based than personbased compared to the situation in north america, cp. the large lassist constituency in north america compared to europe. ' professors harold clarke and mark franklin, texas, made these comments during the ecpr planning session integrating the european data base (essex, uk, march 22-26, 1991). ^this workshop, co-chaired by ekkehard mochmann (za) and eric tanenbaum (esrc-da) was being prepared in march of this year, cf. note 3 above. 'unfortunately, a local research institute in mannheim (frg) has adopted the name eda (european data archive); this may confuse some users inside and outside europe. the name is misleading also in the sense that the holdings of eda are mostly "second hand data" (i.e. data from other archives). lassist quarteriy * cessda was founded in 1976 in amsterdam (may 31 june 1) as an informal cooperation between existing social science data archives. cessda is the european branch of ifdo (international federation of data organizations). active european archives are with an asterisk (*) in front of the cessda founding members: adb (all-union data bank), moscow, ussr * adpss (archivio dati e programmi per la scienze sociali), milan, italy * bass (belgian archives for the social sciences), louvain-la-neuve, belgium bdsp (banque de donnees socio-politiques), grenoble, france * dda (dansk data arkiv), odense, denmark * esrc-da (economic and social research council data archive), essex, uk * nsd (norsk samfunnsvitenskapelig datatjeneste), bergen, norway ssd (svensk samhallsvetenskaplig datatjanst), gothenburg, sweden * star (steinmetzarchieo, amsterdam, the netherlands tarki (social science information center), budapest, hungary * za (zentralarchiv fuer empirische sozialforschung), cologne, frg wisdom (wiener institut fuer sozialwissenschaftliche dokumentation und meihodik, vienna, austria other countries, e.g. czekoslovakia, ireland, and switzerland, are expected to set up similar archives in the near future. undergraduate or even school students. in denmark, some 25% of all schools acquired a teaching package during the second half of the eighties. '" nsdstat from the nsd is presently being distributed in several of the other "cessda countries". nsdstat was presentedand demonstrated during one of the workshops prior to thelassist edmonton conference. " association for history and computing 6th international conference will take place in odense, denmark, august 28-30, 1991. usually, several hundred historians from all over europe (and some overseas guests) participate in the ahc annual conferences. '' this system represents the first stages of a larger project presented by karsten boye rasmussen at the lassist annual conference 1990. the remaining parts, having searched study descriptions and codebooks, include an automatic downloading of the relevant data sets for analyses. ' the lis-model aims at securing access to confidential data in the following manner instead of distributing the "real" data, a constructed "model data set" with the same distributional characteristics is disseminated. once the user has produced the setup that generates the right analyses, (s)he sends that setup to the data base administrator who then runs it against the real data. the administrator sends the output to the user after checking that no confidential information can be disclosed in the output. (setup and output can be sent via electronic mail.) ' most "cessda countries" have made teaching packages based on some of their more "popular" data sets for summer 1991 27 the depository distribution of cd-roms: a review of the first year by juri stratford' university of california, davis introduction a large part of the work of depository librarians is providing public access to the vast number of statistical publications produced by various agencies in the federal government including most notably the census bureau and the bureau of labor statistics. until 1989, the distribution of federal statistical data in machinereadable form was limited to programs administered by individual agencies, e.g. the census bureau's slate data center program and the bureau of economic analysis' local area data program, and institutions acquiring data directly from agencies for purchase or through consortiums and archives such as icpsr. the first cd-rom to be distributed to depository libraries was test disc 2. test disc 2 included state and county data from the 1982 census of apiculture and zip code data from the 1982 census of retail trade . this was distributed to a few libraries on an experimental basis in 1988 and was made available through regular depository distribution in 1989 shortly followed by distribution of the city and county data book cdrom in 1990. since then the census bureau has distributed a number of cd-rom products through the depository system including data from the the 1987 economic censuses , the 1987 census of agriculture , and the reapportionment data from the 1990 census of population and housing . other significant cd-rom products distributed to depositories include the national health interview survey produced by the national center on health statistics, the toxic release inventory produced by the environmental protection agency, and the national trade data bank produced by the commerce department a number of other agencies including the department of defense, noaa, and the geological survey are also beginning to distribute cd-roms through the depository system. cd-rom and papercopy distribution in some cases, the depository distribution of cd-roms complements the depository distribution of paper or microfiche products. for example, the epa's toxic release inventory was made available simultaneously to depositories on cd-rom and on microfiche. the cdrom distribution of the 1987 census of agriculture and the 1987 economic censuses followed the distribution of these pubucations to depositories in paper copy. this has also been the case, so far, with the cd-rom distribution of the city and county data book and county business patterns. in other instances, the cd-rom distribution provides depositories with materials that they might not have had otherwise. for example, the cd-rom distribution of the national health interview survey data, the pl94 data from the 1990 census of population and housing , and the zip code data from the 1982 and 1987 economic censuses represent data not available in paper copy. a final area, which should concern the depository community, is the replacement of the depository distribution of paper copy or microfiche with cd-rom. for example, the monthly import and export data from the census bureau is now only being distributed to depositories on cd-rom. also, the 1985 congressional record cd-rom was recently distributed to depositories on a trial basis as a replacement for the bound edition. advantages and disadvantages of data on cd-rom cd-roms offer some advantages both to end users and data producers. the electronic distribution of data provides the potential to enhance user access. libraries can provide access to a vast quantity of federal data in machine readable form allowing end users the ability to work with the data on their own microcomputers. where in the past researchers might have had to extract data subsets from tape, or to key in data by hand from printed sources, they can now copy the files directly from the cd-roms to diskette. data from these cd-roms are available free of charge and without copyright restrictions. cd-rom also offers some advantages to data producers. in many instances, it is less expensive to produce and distribute cd-roms than paper copy, and congress is anticipating a cost-saving through cd-rom. in a statement before the american library association legislation committee in january, robert w. houk, the public printer repwrted that gpo requested a fifteen jjercent increase for salaries and expenses primarily associated with the distribution of 1990 census publications. however, congress reduced the request from lassist quarterly s27.9 million to $26.5 million, projecting that the census bureau would distribute a greater proportion of documents in cd-rom formats, rather than in microfiche as originally anticipated^. when materials in paper copy or microfiche are replaced by cd-rom, the cd-rom distribution may be viewed as shifting the expense of production from the data producer to the data user. while many depository institutions are facing budget reductions, they are now also faced with the added expense of devoting cd-rom stations to provide public access to the federal data. other expenses will include acquiring the proper software to work with the data and providing paper for printouts. data format and software issues the census bureau data are being distributed in dbase format. for the demographic data, the census bureau is distributing software to display tables, usually for specific geographic areas, to the screen or to a printer. the software for the foreign trade data displays data for the current month and beginning of current year to date for the most specific commodity code. the census bureau is also distributing programs such as extract to create small data subsets for output as ascii files, dbase files, or lotus worksheet files. extract can also print tables from the data sets. the extract program requires a large number of dictionary files describing the data. the dictionary files for the economic census are about four megabytes and these files have been included on the cd-roms. however, in other instances these files have been distributed separately on floppy disks, and they must be installed on a hard disk for use with the cd-roms. the city and county data book files require 640kb of hard disk space. county business patterns files require 650kb, and the monthly import and export files require 3.2 megabytes. as the census cd-rom files are in dbase iii format, the data can also be accessed directly using third party database management and spreadsheet software. however, many of the census bureau files are large and are very difficult to work with on a microcomputer. the city and county data book files are small enough, and most of the economic census files are less than one megabyte, though a few are as large as two megabytes. however, some of the pl94 files are as large as 300 megabytes; and the foreign trade export files are currently about 258 megabytes while the import files are more than 550 megabytes. public access issues regardless of what software is used, whether a ubrary uses extract or dbase or some other software, data extraction from these files requires a lot of processing time. it appears unlikely that extract or dbase could be used to access these data files in a public reference area. in a recent article in government publications review, steven staninger documents the time required to conduct a few simple searches using dbase iii plus with the 1982 census of retail trade data on census bureau's test disc 2. in one example, it required twelve minutes to execute a search for a single type of business code in the california data file on an ibm pc. there appear to be two strategies to the problem of public access to numeric data files on cd-rom. a combination of both tactics will probably be necessary to provide adequate public access. the first strategy is to develop the personnel and physical resources necessary to effectively deal with the problem of data extraction. to adequately deal with electronic formats a depository will require at least one staff member with microcomputer expertise and a strong social sciences background. at a minimum, this person will have to be familiar with dos, database management software and spreadsheets; a familiarity with statistical programs such as sas and/or spss would also be helpful. this jjerson must also be familiar with the printed sources and have a close working relationship with the local data archives facility, if any, in order to know when cd-rom is more appropriate than paper copy or data on tape. there must be adequate physical resources as well. there appears to be a well established base of cd-rom stations in large depository ubraries. this is supported by the fact that at least half of all depository libraries have selected some census data on cd-rom'. many of these depositories are also using these cd-rom stations to provide end-user access to bibliographic files. however, as staninger's article suggests, the effective use of the census bureau cd-roms will not lend itself to this setting. libraries will need to provide cd-rom stations out of the reference area where either library staff or endusers can extract data from the cd-roms. these cdrom stations will also need to devote a large amount of hard disk space to work with the depository cd-roms. in addition to the requirements outlined above, many of the depository cd-rom products have their own front ends requiring a large amount of dedicated hard disk space, e.g. the national health interview survey requires about five megabytes and the toxic release inventory requires about seven megabytes of hard disk space. the second strategy is to look for solutions in the private sector. while there are a few products, such as ec stars , which are designed to work with the depository cd-roms, many commercial software producers are summer 1991 reselling the federal data. for example, space-time research's supermap, marketed in the u.s. by chadwyck-healey, contains data from the 1980 census of population and housing summary tape files . stfl-c and stf3-c . and additional county and land area data; and statmaster produced by cybersoft includes data from the county and city data book. while these commercial products might be expensive, each of these products provides enhanced access to the data. there appears to be a fear that this could make depository libraries more dependent upon the private sector for access to census materials; a census bureau report notes that librarians are "concerned by the need for userfriendly software to access census data and fear that it may be available from only the private sector at prohibitive cost."* but there is a long established, successful history of private publishers providing bibhographic access to depositories. as sir charles chadwyck-healey stated in a recent interview published in government publications review: "governments seem to be extraordinarily bad at distributing information efficiently. it is probably inevitable. i am not sure there is going to be an enormous advantage to having very cheap data available from government sources if those sources are not able to disseminate it in and efficient and effective way."* the depository cd-roms will have a great impact on library public services. in fact, if widely adopted, the cd-rom distribution has the potential to transform the typical depository as we know it perhaps the greatest impact of cd-rom in a hbrary is the increase in workload for the public services staff in whose area the cdrom is located. in a recent article, steven zink argues that the positive aspects of cd-rom have tended to overshadow the human resources required for its use. he notes that "a persistent administrative malady is the assumption that technology will decrease, or at least not require additional, demands on staff time." in fact, the opposite tends to be true. zink explains that "while technological advances may have resulted in personnel reductions in selected technical service areas, the use of automation where the public directly confronts technology has generally increased the need for user assistance."' depositories have yet to determine how best to deliver the additional assistance that users of electronic formats will require. generally, document librarians perceive that the data extraction from the depository cd-roms will follow the patterns established by mediated online searches. however, as elizabeth stephenson observes: while many librarians are experienced in handling bibliographic data on cd-rom, "few have any experience or training in the manipulation of numeric files." she believes that librarians will have to become familiar with the hierarchical structure of the files and statistical languages to effectively work with the depository cdrom products.* conclusion in 1988, diane smith of pennsylvania state university conducted a survey to determine the preparedness of depositories to provide access to electronic data products. she examined the extent to which depositories were experimenting with the provision of electronic services and looked for characteristics common to innovative libraries. her survey covered plans to include documents in local onune public access catalogs; the use of onhne databases, cd-rom, statistical software, and expert systems in depositories; and available hardware within reference areas. smith concluded that depositories were ill-prepared to deal with electronic formats and that there is a definite need for depository libraries to face the training and collection management challenges presented. she cautioned that "if this work is not done, it appears that there will be a major crisis in the ability of libraries to deal with electronic data; a crisis that questions the viability of the present situation."' as early as 1988, jones and kinney argued that when document librarians assume the responsibility of retrieving numeric, textual, or bibliographic information from computer tapes, they must either know how to program or work with a colleague who programs'". it is unlikely that a majority of depositories will be able to provide adequate programming assistance in-house. this means that many depositories may need to establish collaborative relationships with other units to adequately service the depository cd-roms. likely partners include computer centers, data archives and social science research units or teaching departments. at present, both product development and participation in the distribution of cd-roms through the depository program appears to be coordinated at the agency level. for example, within the commerce department, the census bureau is distributing cd-roms through gpo's depository program while the patent and trademark office is only offering their own cd-rom product, cassis, for sale or through their own patent depository program. each agency is approaching the data format and user-interface issues differently. while the census bureau has committed itself to distribute data to depositories in dbase iii format for use with privately-produced software, the office of the secretary is distributing the national trade data bank cd-rom with its own unique user-interface and data format. some cd-rom products will require that depository libraries acquire privatelyproduced software to work with the data; others will require that depository staff invest a substantial amount of lime mastering the userinterface provided by the agency. while the true magnitude of this trend has yet to {assist quarteriy be determined, the impact upon public service in depository libraries will be siginificant. ' presented at the iassist 91 conference held in edmonton, alberta, canada. may 14 17, 1991. ^ robert w. houk, "remarks before the american library association midwinter meeting: legislation committee, information update," administrative notes 12 (february 22, 1991): 1-6. ' steven w. staninger, "using the u.s. bureau of the census test disc 2: a note," government publications review 18 (march/april 1991): 172-3. * peter hemon and charles r. mcclure, "electronic census products and the depository library program: future issues and trends," government information quarterly 8 (1991): 61. * sandra rowland, 'the role of intermediaries in the interpretation and dissemination of census data now and in the future," reprinted in documents to the people 18 (june 1990): 81. * jean slemmons stratford, juri stratford and steven zink, "applying an "entrepreneurial attitude" to the dissemination of government information. an interview with sir charles chadwyck-healey, chairman of the chadwyck-healey publishing group," government publications review 18 (march/april 1991): 134. ' steven zink, "planning for the perils of cd-rom," library journal (february 1, 1990): 54. * elizabeth stephenson, "data archivists: the intermediaries the census bureau forgot. a review essay of 'the role of intermediaries in the interpretation and disemination of census data now and in the future," government publications review 17 (september/october 1990): 443. ' diane h. smith, "depository libraries in the 1990s: wither or wither depositories?," government publications review 17 (july/august 1990): 312. '° ray jones and thomas kinney, "government information in machine-readable data files: implications for libraries and librarians," government publications review 15 (january/february 1988): 30. summer 1991 31 vol183&4 4 iassist quarterly introduction as more organizations and institutions downsize computer facilities in order to make greater use of the ubiquitous and inexpensive desktop computer, the problem of how to get non-ascii data from there to here becomes increasingly common and pressing. archival requirements as well as data utilization are affected by the platform shift; additionally, users are expecting greater access to data than in past decades and devising access methods with and without the blessing of the mis staff. life expectancy of archival tape media from the 70’s and 80’s is diminishing. all of these issues draw us to ask the question: how do you get the data off the mainframe and onto the computer. let us break the big problem, the great need, into smaller and more manageable problems, in the spirit of the eating of the elephant. problem: determining whether the data is even suitable for conversion careful evaluation of the data will help you determine whether to shelve the project or to move onward. this information is critical in determining not only feasibility, but potential cost of the project. this evaluation will help you to discover those unpleasant exceptions to the rule that will require expensive programming and special processing that can drive costs for conversion out of the feasibility range. techniques: 1)check the physical condition of the media itself, especially if has been many years since cleaning and copying of the tape 2)try to evaluate the adequacy of documentation, so far as record format, field definitions and descriptions, and code tables. 3)do your best to determine that this data is really what you thought it would be, that it is suitable for your needs, or that isn’t already duplicated elsewhere in a more accessible format. 4)determine the tape density on older tapes, and for very old tapes, whether they are 7 track or 9 track2. 5)non-ebcdic data formats crop up on older tapes especially, and can greatly increase the effort and expense of conversion. look for packed decimal, zoned decimal, packed bit, or binary data, these data formats will need special conversion techniques3. 6) unusual file formats will also need special conversion techniques: for example, tapes from military sources may be in nips, those from medical facilities may be in mumps, and pick systems have been in wide use for many years4. solutions: if your facility has mainframe to pc connections, your tapes are in good condition, and your tapes are readable by your current mainframe or mini facility, you can run a sample of 3 to 5 megabytes from each file across the network. the pc interface cards necessary for the pc to mainframe connection will automatically convert mainframe ebcdic data to ascii data. this sample data will help you to determine the adequacy of the available documentation and the presence of unusual data and file formats, which we will discuss further in this section. data which does not convert directly from ebcdic to ascii can be readily identified. even in very large files, a sample of this size will almost always yield usable data in all fields. if, however, your data is truly historic, just accomplishing this task can be a problem in itself. unless your mis staff is familiar with older computers, tape, and data formats, this evaluation may be better left to professionals. the section on data conversion service bureaus addresses the issue of older tapes. data conversion service bureaus there are data conversion service bureaus in most cities that deal with old data on a regular basis. for very old and fragile tapes, consider contacting a disaster recovery service; many of these agencies have the techniques and equipment to do serious data recovery. be prepared to pay for this initial evaluation, and ask for a quote (based on the number of files you’ll want evaluated). you will utilizing mainframe data on pc platforms: problems, solutions, and techniques by carol wickenkamp1 wae 5fall/winter 1994 need an evaluation that will cover all the points discussed above, in 1 through 6. in addition to the initial evaluation, request a quote for providing a 3 to 5 megabytes sample ebcdic to ascii conversion from each file if the tapes are readable. unless the files are under 20 to 30 mb, ask if they can take two small samplings (500k), one from the middle of the file and one from the end of the file as part of the 3 to 5 mb sample, and find out how much extra it will cost you for these small samples. current price for ebcdic to ascii conversion is usually about $10 per megabyte. get cost quotes for your evaluations from more than one agency, and also ask if you can contact previous customers, as you would for any contract service. tape drive peripherals if you have neither mainframe to pc capabilities nor the funding for service bureau work, or for other reasons have decided to tackle the project in-house, consider rental or purchase of a 9 track tape drive that will interface with a pc. these tape drives will come with software that will perform simple ebcdic to ascii conversions, and some will have software will have software with even more capabilities. for example, qualstor’s drives come with software that will convert directly from ebcdic to dbase. drives are available that will handle varying tape densities; overland makes a tape drive that will handle even the very old 800 bpi density as well as the contemporary 6250 tapes. if you know that your tapes are not fragile and you can safely run them, you can use a tape drive peripheral to do your initial evaluation of your data, running the same 3 to 5 mb sample. data conversion service bureaus often rent drives, as do some of the larger computer equipment rental companies. the cost is usually about one teeth the purchase price; drives adequate for most conversion jobs will rent for around $600 per month. documentation and identifying unusual formats using the documentation you’ve gathered, and a print out of your sampleascii data (start with just a few records), you can begin the task of reading the raw data. this process will uncover gaps in your documentation as well as “funny” data. frequently unusual data and file formats will be easily discovered on initial examination, before you even begin to check your data against the documentation. see figure i for “funny data”; the fields that contain the curly brackets signal the presence of zoned decimal numeric fields, as do “/” characters and unexpected periods. zoned decimal will be converted incorrectly in a simple ebcdic to ascii conversion, as is obvious. other non-ebcdic numeric formats can also yield exotic results. if you find no indication of problem data, use the field descriptions in your documentation to mark off the fields in your data, as in figure 3. check your data fields one by one against both the field definitions and code table, if some of the data is coded. here in figure 3 we have clean data, with names, dates and julian dates, cities, etc. where they should be and in the proper format. make sure that the code values in coded fields are represented in the code tables. should you find codes that are not listed in the code table, but the rest of the data is clean and in agreement, you have probably encountered either an undocumented code (if there are many occurrences) or data entry errors. if you have undocumented codes, you can sometimes extrapolate the meaning from the data when the entire file is converted. often a further search for more documentation is necessary. (both the national archives and ntis retain copies of some federal computer documentation.) lack of sufficient documentation can doom your conversion project, unless you can be satisfied with either converting the portions of the data that you can identify, or just archiving the data in the hope that you can obtain the requisite documentation at a later date. take samples of 20 to 50 records from different places in your 3 to 5 mb sample and verify the data. if your are able to obtain records from the middle and end of your life, be sure to check them, as sometimes another file with a different format was appended to the first data file. should you find evidence of multiple files, you will want to make a note of it so that when you have the tape converted, the data can be run off in separate files during the conversion process. determining conversion costs using the evaluation information about your files, you can begin to calculate costs. for example, if you send 300 mb of clean ebcdic data to a data conversion service, and they charge $10 per mb, your charges will be $3000. to this you must add the cost of target media sufficient to store that volume of data. this figure will of course vary according to the media. should your facility plan to download the data from a mainframe to a pc, your in-house costs will, at a minimum, include target media costs and computer time, which may or may not include computer operator charges. coordination with your mis department will be essential in defining costs for in-house conversion. if you have data that requires special processing, costs may include data recovery fees for very old and fragile tapes, or programming costs to convert data that is in non-standard data or file formats. you will need to obtain a second round of quotes for this 6 iassist quarterly work, which will be more expensive than standard conversions, or negotiate with your mis department for programmers to do the work. doing the conversion yourself, for those without mainframe connections or funding for service bureau work, will be addressed in section converting data on a low budget. if your data will require the programming services, expect to pay a minimum of $50 per hour. programming costs in major metropolitan areas will be greater. as with other contract work, obtain more than one quote and ask to speak with previous customers. try to speak with customers whose programming and conversion needs were similar to yours, in order to ascertain that the programmers have actually dealt with this type of data or file format; you don’t want to pay for the programmer’s learning curve. problem: converting data on a low budget there are those facilities who will not have the resources of an eager to help mis department, or the budget to cover thousands of dollars for data conversion services. there are alternatives that can put the data conversion and migration process in the realm of the possible for even the most underfunded facility. techniques: hardware before we begin the “hands on” process of converting this data, we must have some repository for the finished product. depending on the volume of data, there are a number of target media that will be appropriate. high capacity hard drives are becoming very affordable, with prices dropping to around $1 per mb and even less for very high capacity drives of over 1 gigabyte. this drop in price put desktop mass storage within the reach of low budget facilities. the lowest cost storage media will be the inexpensive pc backup tape. qic 80 tape, which is becoming a standard for entry level backup, will store 250 mb of compressed data; this means that you will usually be able to store more than 250 mb of data on one tape. the drives are inexpensive, currently selling for under $200, and will function very well in older at class pcs. the media will cost about $20 per cartridge. the drives are adequate for short term archival storage (not recommended for a permanent solution), but are slow and inefficient if you plan to use the data frequently. removable media hard disk drives are available in either internal models or portable models that interface with the pc through the parallel (printer) port; these drives offer another attractive alternative. prices on these drives rapidly dropping; at the present, a drive in the 110-120 mb range can be purchased (with some judicious shopping) for about $400, including one cartridge; higher capacity drives are available. each cartridge contains a hard disk platter, and the user can easily switch cartridges. the media costs are about $65, and prices should fall rapidly. the advantages include very fast access to data for those who need frequent access and portability. these drives can be compressed with disk compression utilities, increasing the storage potential. they are an excellent choice if your data files are in the appropriate size range and you will require frequent access to the data. dat backup drives are more expensive starting at about $1000, but they are very fast, they store gigabytes of data, and the cartridges cost about half as much as the qic80 cartridges. solutions: data copy by data conversion service bureau service bureaus will make an exact copy of your data and write it to your media. the current cost for this service will be in the range of $1 to $1.50 per mb of data. for example, if you are using qic80 tape, request the bureau to make a copy of the data file(s) onto qic80 cartridge media, which you will then restore to a hard drive at your facility for do-it-yourself data conversion, or simply retain as archival storage. (a discussion of doit-yourself data conversion will follow in this section.) if you are using a tape backup medium, be sure to tell the service bureau the name brand of your tape drive, as cartridges written by one brand of tape backup equipment be readable by equipment manufactured by another company. it is wise to do a test run with a trial tape cartridge written by their equipment, to determine whether your equipment will read the tape. you will also want to request that the data file tape headers (preliminary system information written when the tape file was created) be stripped from the data, and that only data be copied onto your medium. if you have a large number of tapes, it will be wise to pre-determine a meaningful data file naming scheme, so that you will know which data file is which when you get them back. nine track tape drive rental your facility may decide that tape drive rental is the most feasible course. basics on pc peripheral 9 track drives were covered in an earlier topic. the company that rents you the tape drive may provide both installation and removal of the interface card if you have no one on site who can do it. as was earlier discussed, the software that comes with these drives will provide the option of converting the ebcdic data to ascii as it is copied off the tape and onto your storage medium. those who are 7fall/winter 1994 not familiar with tape conventions such as blocking, and fixed and variable length records, determine the degree of customer support available from the rental agency you may need some initial instruction. if you have no special conversion needs, this is a most cost effective solution to the data conversion. data conversion software service bureaus that do data conversion and rent 9 track tape drives often sell special data conversion software that has more features than the software that is bundled with their tape drives. typically, software of this type will handle the unusual data formats mentioned earlier, and can convert standard variable length records to fixed length records. expect to pay $200 and up for this software. do not count on conversion software to accomplish the task of converting the non-standard file formats discussed earlier; you probably will still require programming services. frequently the software interface is intimidating and may be hard to get used to, but the conversion process itself is not overwhelming. generally, you will be required to mark off the data fields (as you did with your sample, only on screen rather than on paper) and then define the conversion process that is to take place, i.e.. ebcdic to ascii, binary to ascii, or packed decimal to ascii. when you have defined your conversion instructions, your file is ready to be converted by the software. it is a good idea to run a partial conversion of 500 to 1,000 records to verify the accuracy of your field definitions. sometimes the process will require several tries before all the bugs are out of your conversion instructions, and it is far faster to convert 1,000 records for a sample than to convert 100,000 records. the speed of conversion will depend upon the processor speed of your computer, the complexity of your conversion instructions, and the length of your records. you can use the measure of 1 megabyte per minute as a rough rule of thumb. although most of these programs will operate on files residing either on the tape drive or a hard disk, it is much faster to copy your file onto a hard disk and do the conversion from disk. problem: the data is so heavily coded that it will be difficult to work with as a rule, database programming relies heavily on code table to hold frequently used values; old mainframe data can be coded in every field, thus yielding very compact files. the code values were replaced at processing time so that reports were understandable. this sort of data is very cumbersome to use, even with modern and easy to use database programs such as paradox, alpha four, access, etc. solution: given the low cost of hard disk storage, it is becoming more feasible to simply replace the coded fields in databases with their values, yielding a significantly larger, but easier to use flat file database. even with a two or three fold increase in file size, this solution can bring comprehensible, easy to manipulate data to the most unsophisticated user. it is far faster and more accurate to extract reports or meaningful data screens from a database that contains “lutheran” rather than “07”, “buick” rather than “15” or “ca” rather than “05”. expert programming skills are not necessary to accomplish these replacements, a moderately skilled inhouse programmer should be able to do the job. even if it is necessary to hire a programmer, it should not be a major expense, unless you have a large number of heavily coded files. conclusion although moving data from older mainframe generated tapes to a pc platform is a process that requires planning and attention to detail, the task is not insurmountable, nor is it always exceedingly expensive. with the exception of very old or non-standard tapes, much of the work can be done in-house and with a small budget, utilizing moderate computer skills. notes: 2. 7 track is an obsolete tape standard which used the 6 bit bcd (binary coded decimal) code together with a parity bit. the contemporary 9 track drives will not read 7 track tapes. 3. although data conversion software renders these numeric formats harmless to the non-technical user, a discussion of these formats is included for those who are interested. numeric data format which will not convert in a standard ebcdic to ascii conversion include: packed decimal with low order sign bit this is the normal ibm packed decimal field. zoned decimal with low order sign bit this format is generated by some cobol, pl/i and assembler systems; although not common, it is still in use in some contemporary installations. zoned decimal is a standard ebcdic numeric character field with the exception of a sign code in the high order nibble of the low order byte, with c hex and f hex being a positive sign code and d hex a negative sign code. this results in invalid ebcdic characters in the low order byte of some zoned decimal fields. binary with most significant byte first this is the format in which ibm mainframes normally 8 iassist quarterly process binary data; normal pc binary format is binary with least significant byte first packed with high order sign bit this is a binary format with the sign bit in the high order nibble of the high order byte. packed with no sign bit this is a normal packed field, except that all nibbles contain a significant digit (no sign field) and the field may begin and/or end on a nibble boundary. 4.these are all non-standard variable length file formats. mumps has been widely used in va hospitals and in medical clinics, and is still common. pick usage extends across the commercial spectrum. nips was designed specifically for use on ibm 360 computers, and is no longer in use. sources: * for further information on tape formats, labeling and file conventions, you can contact: american national standards institute, inc. 1430 broadway, new york, ny 10018. tel : (212)6424900 ask for publication x3.27, “magnetic tape labels and file structures” ibm tape labeling conventions are explained in the ibm publication “os/vs tape labels” (gc26-3795-3, file no. s370-30) and “dos/vse tape labels” (gc33-53741). dec information is described in “guide to vms files and devices” (aa-la06a-te), available from dec. * if your facility is not in a metropolitan area, you may find several reputable data conversion service bureaus advertised in pc magazine, which is available in most drug stores and supermarkets. * two companies that produce data conversion software, each with different capabilities, are: novastor 30961 aguora road, suite 109 westlake village, ca 91361 (818)707-9900 fax (818)707-9902 overland data 5600 kearny mesa road san diego, ca 92111 (619)571-5555 fax (619)571-0982 service bureaus may also have information on other data conversion software. 1. paper presented at iassist 1994 in san francisco.. reprints of this paper are available from: carol wickenkamp, wae, po box 349, clarkston, wa 99403 analyzing nursing home characteristics: issues in comparing state and federal data sources. by joel b. cowen' coordinator health services research college ofmedicine at rockford university ofillinois introduction aging of the population is increasing the importance of long-term care and nursing homes as a component of the health care system in the united states. policy planners need to fully understand the nature of the industry, as well as emerging trends. nursing homes, as we know them currently, are a relatively new entity on the health care landscape. the present longterm care system largely emerged in the early 1970s following the impact of medicare. earlier facilities hke board and care homes had sprouted up after the passage of social security, along with was activity by some churches and fraternal groups who had built retirement centers for their members. since nursing homes are only a couple of decades old and changes in regulations and financing impacting them are being seen with increasing frequency, nursing homes could well change in form or format into the next century, when the elderly portion of the population will move toward one-fifth as baby boomers edge into their senior years. planning for an essential service like long-term care requires accurate information. conjecture cannot take the place of data vital for shaping the nature of future nursing homes. however, despite the importance of nursing homes as part of health care, no comprehensive information system is in place nor are basic definitions of data elements accepted in a widespread manner. the quality of planning suffers when information is limited. for instance, the impact on nursing homes of recent changes in medicare reimbursement on hospitals needs to be known. in this paf)er, current data for nursing homes is analyzed. the two major federal data sets for nursing home facility and resident characteristics, plus similar informafion collected in one major state, illinois, are reviewed. data elements obtained and their definitions are compared and, finally, recommendations made as to how to improve the present situation. nursing home data sources national nursing home survey the national nursing home survey (nnhs) is a continuing series of national sample surveys of nursing homes, their residents and staffs by the national center for health statistics (nchs). three surveys have taken place, in 1973-74, 1977 and 1985. no plans have yet been announced for a fourth survey. the surveys employ a stratified two-stage probability design, first the selection of facilities, then the selection of residents and employees. data has been collected using both personal interviews and forms for self-completion. in 1985, added information was collected from relatives of residents, the "next of kin" questionnaire. in 1977 and 1985, a sample was also taken of persons discharged from the home in the past year, whether alive or dead. the nnhs is the only national survey which includes variables for individuals which can be analyzed against each other. other sources group data, which cannot be crosstabbed or correlated. all data is available on tape from the national center for health statistics. figure 1 shows the available variables. tapes are also available for the states of california, illinois, massachusetts, new york and texas. more cases were provided in these states so that reliable estimates could be made. results appear in written form in vital and health statistics . series 13, nos. 97, 98, 102, 103, and 1 15 plus advance dala, nos. 131, 135, 142, 147, and 152 from nchs. lassist quarterly figure 1 summary of 1985 nnhs data tapes by type of hle facility file resident file facility number ownership code number of beds (1985 and 1984) certification status per diem rates by certification status admissions (1984) residents days (1984) services offered to rsidents services offered to non-residents physician service arrangements fulland part-time staff part-time staff hours nursing staff hours volunteer staff (geographic region recode dhhs administrative regions msa recode fadbty weight facility number age sex race hispanic origin marital status at admission and currently living children date of last admission residence before admission hospital stays while a resident previous nursing home stays diagnoses at admission and currently mental disorders therapy services received vision and hearing status activities of daily living adapted instrumental activities daily living behavioral problems disorientabon or memory impairment deprssion, anxiety, fearfulness, or worry sources of payment at admission and last month total monthly charge for care last month amount paid by source last month resident weight record length block size 665 19,950 1,078 record length block size. 873 17,460 5;238 discharge file nursing staff file facility number age at discharge or date of birth sex race hispanic origin marital status at admission and at discharge date of admission and discharge discharge status (alive/dead) residence before admission residence after discharge for live discharges hospital stays while a resident nursing home stays before and after sample stay diagnoses at admission and at discharge mobility staff continence status sources of payment at admission and at discharge discharge weight facility number member of staff or other arrangement type of position length of work experience hours worked salary services performed employment conditions sex and age ethnicity marital status children living at home education staff weight record length block size number of records 544 21,760 6,017 record length block size number of records 307 21,490 2,760 exfjense file facility number expenses and revenues expense weight record length block size number of records 366 18300 731 fall 1992 17 inventory oflong-term care places the federal inventory of long-term care places (iltcp), formerly part of the national master facility inventory (nmfi), was conducted in 1986 by the census bureau for the national center for health statistics. this differs from the nnhs in that the entire sample of "nursing and related care homes" is covered. variables are limited to ownership, certification status, number of beds, residents, and race of residents. information on residents is tabulated in grouped data for reporting. one way that the iltcp differs from the nnhs is that special homes for the mentally ill, developmentally disabled, and other groups with special handicaps are included. also available on tape from nchs, the full set of variables is shown in figure 2. both the nnhs and iltcp are also available through the inter-university consortium for political and social research. figure 2 summary of national master facility inventory data tapes by type of facility hospitals nursing homes and other health facilities name name of administrator ownership tyf>e of facility number of beds days of care discharge admissions type of service outpatient visits employees facilities and service offered 1971 1972 1971 1974 1975 1976 name adress number of beds ownership type of facility ages served sexes served number of residents 1971 1973 1976 record length 840 block size 8,400 748 4,488 7,480 1 748 4,488 7,438 1 748 4,488 7,370 1 748 4,488 7,336 1 748 4,488 7,271 1 record length block size number of records number of reels 600 3,600 26,773 1 196 1,176 26,003 1 210 6,720 26,748 number of reels 1 1 illinois annual survey oflong-term care facilities most states collect data on their nursing homes and their residents. this is usually made necessary for the state's licensure certification or planning functions. illinois is typical. the illinois department of public health (idph) has conducted annual surveys of long-term care facilities since 1981. these surveys are carried out under the state's certificate of need act which authorizes data collection for planning purposes. in addition to using the data to create an inventory of long-term care services and bed need plan, idph makes survey data available to others who have an interest in long-term care, both public and private. other state agencies also constitute a major user. the luinois survey format is relatively standardized year to year with supplemental "special" studies each year. nearly a thousand licensed facilities (including mental heallh/dd) receive the questionnaires and participation is almost universal. uicom-r long-term care report the health services research office of the university of illinois college of medicine at rockford has created reports characterizing nursing homes in northwest illinois. these reports, completed in 1980, 1985 and 1990, utilize information from the idph long-term care facility survey. this effort is notable because few local area studies are conducted which track changes in the industry on a local area basis. state reports generally do not evaluate long-term changes or focus on regional differences. lassist quarterly data elements in this section, certain data elements have been selected to illustrate the availability of nursing home data, so as to reveal potential sources, their differing methods of gathering, commonalities and differences in definition. what is a nursing home? defining a nursing home is not a clear and simple task. the nnhs includes all types of "nursing homes" regardless of their level of care, participation in medicare or medicaid or licensure. no "board and care" homes were included or those providing residential care. they define nursing home as: facilities with three or more beds that provide to adults who require either nursing care or personal care (such as help with bathing, correspondence, walking, eating, using the toilet, or dressing) and/or supervision over such activities as money management, ambulation, and shopping. facihiies providing care solely to the mentally retarded and mentally ill are excluded. a nursing home may be either freestanding or a distinct unit of a larger facility. illinois relies on licensure for its survey definition. in illinois, a nursing home is defined as a: "private home, institution, building, residence, or any other place which provides personal care, sheltered care or nursing for three or more persons who are not related to the owner." long-term care institutions in illinois are further classified into two types of care, nursing and sheltered care. nursing care includes the provision of diagnostic, therapeutic and rehabilitative care under a patient's plan of care as prescribed by a physician. in addition to the medically oriented care given by nurses and the living assistance given by aides, other services commonly provided by a facility that provides nursing care includes physical therapy, speech therapy, occupational therapy and social activities. there are two levels of nursing care, skilled and intermediate , with the levels differing in the amount of available nursing expertise and supervision provided to residents. sheltered care includes the provision of personal care and support in daily activities with limited nursing consultation available, such as the taking of medications. the term "sheltered care" is somewhat unique to illinois. other states tend to utilize the terms "assisted living" or "personal care." the federal nnhs survey relies on medicare/medicaid definitions for designation of skilled or intermediate care levels. other beds are shown as "not certified." both units of government agree in that a facility must have three beds and provide nursing or personal care. illinois does not consider personal care to determine a nursing home unless there is nursing consultation, such as the taking of medication. terms common to definitions such as nursing care and personal care are not precisely defined. another area for possible confusion is whether facilities such as those for the mentally ill or developmentally disabled are counted. the nnhs excludes these. the iltcp includes them as does illinois. the iltcp, which begins with categories similar to the nnhs, goes beyond these categories to also include homes for unwed mothers, substance abuse, orphans and the terminally ill (hospice). in general, the definition of nursing home in this country is still imprecise, leaving board and care, congregate living and certain retirement centers and specialized facilities in a zone of uncertainty. beds beds are classified in various ways such as licensed or unlicensed, set-up and staffed, or occupied. the nnhs and iltcp both primarily use set-up and staffed, while illinois uses licensed beds. licensed beds usually reflect capacity whether actually used or not. the creation of swing beds and distinct unit skilled nursing beds at hospitals has increased in recent years. swing beds are those that can be used for either acute or extended care through reclassification of the patient who doesn't actually move from the bed. swing beds are certified by the medicare program. distinct units provide only extended care. for the most part they are utilized for hospital patients no longer needing acute hospital nursing care but who are not yet fall 1992 19 well enough to return home. under the "drg" reimbursement system, hospitals benefit financially from discharges to a lower level of service. the nnhs totally excludes hospital-based beds. on the other hand, the iltcp includes long-term care units in hospitals. presumably this includes both swing beds and distinct units. illinois counts distinct units as a "nursing home" category, but does not consider swing beds in this category. as is probably obvious, counts of beds can differ greatly according to the counting method. figure 3 shows some of these conflicting counts. figure 3 comparison of bed counts: nnhs, iltcp, and idph nursing homes in u.s. total nnhs iltcp nursing care hospitals not certified residential nursing 14,400 4,700 homes in illinois 16,388 • 734 9,258 total iltcp idph nursing care hospital based residenlia 744 25 48 898 734 admission, discharge and slay length of stay is important for policy decisions and planning of resources, not to mention actuarial calculations such as those for the long-term care insurance industry. average length of stay cannot be calculated with ease as is done in hospitals: patient days = average length discharges of stay nursing home stays more often cross multiple years and may involve intervening hospital stays. the iltcp collects annual admissions and number of "residents last night." average length can only be calculated if the assumption is made that "residents last night" times 365 = an estimate of patient days. idph collects patient days, admissions and discharges so that an estimate of average length of stay can be calculated. again, the year to year crossover can be a problem in such calculations. the nnhs is far more precise in its treatment of stay length. this is important because a good source is needed which differentiates the characteristics of stay types, especially the nature of post-hospital short stays from longer term chronic type stays. the differing nature of different "types" of nursing home residents is important for policy. the nnhs provides for a question on duration of stay at discharge. recently they have reconsidered this indicator of stay length in favor of the long-term care use history of individuals. an individual may be admitted and discharged several limes a year and, in fact, that pattern has been found to be common. additionally, nursing homes treat stay definitions differently, especially with regard lo temporary transfers to hospitals. some consider the movement to be a discharge with formal readmission on return, while others make no such change in assist quarterly status. some facilities include the hospital stay as part of the stay, while others exclude them or calculate "bed-hold" days. dlinois provides no instructions on how to deal with the reporting of "bed-hold" days. this is left to the institution. another element of admissions and discharges is the place of origin and discharge where are residents coming from? and where are they going? the iltcp does not collect this, but both the nnhs and illinois do. the nnhs collects "uving arrangement prior to admission and after discharge," while illinois asks for "admissions from" and "discharges to." the categories are similar except that the nnhs is far more precise in terms of home residence. a comparison is shown below in figure 4. figure 4 ltc admissions and discharges comparison of nnhs and idph nnhs idph own home or apartment private residences relative's home or apartment other private home or apartment retirement home boarding house, room other nursing home other nursing home general hospital general hospital mental hospital psychiatric hospital state mental health facility community mental health residences chronic disease hospice other other resident demographic characteristics recording the demographic characteristics of residents provides an essential component of any description or analysis of the nursing home industry. age, race, and sex constitute the core of these descriptors. birthdate of the resident is collected by the nnhs, allowing any age groupings. the iltcp utilizes three groups: 0-21, 22-64 and 65-t-. idph provides for seven groups: 0-17, 18-44, 45-59, 60-64, 65-74, 75-84 and 85-t-. sex (gender) is collected by the nnhs and idph, but not the iltcp. nursing homes tend to be dominantly female. like many of these data elements, race is classified three different ways. the nnhs applies the usual census format in which race (white, black, american indian, and asian) is a separate variable from ethnic origin (hispanic). illinois, however, uses combined racial/ethnic groupings (white, non-hispanic; black, non-hispanic; asian, non-hispanic; and hispanic). the iltcp collects only the number of black and hispanic residents, no other racial groups. two resident elements in the nnhs only are marital status (current and at admission) and number of living children. health and activity health of the individual is often expressed either through categorization of the major condition or disease resulting in the admission or indications of which activities of daily living (adls). the iltcp does not record diagnosis at all. illinois utilizes groupings of icd-9 codes such as "diseases of the circulatory system" or "musculoskeletal diseases." the nnhs usts certain common reasons for nursing home placement such as stroke, hip fracture and alzheimer's disease. the nnhs also collects indicators of health status prior to admission, including drg if hospitalized. fall 1992 figure 5 demographic characteristics of nursing home residents indicator age sex race/ethnic marital status living children? nnhs individual years collected yes white, black, amer. indian, asian hispanic yes, current and at admission yes iltcp 0-21 22-64 65-tgrouped only no black, hispanic only no no idph 0-17 18-44 45-59 60-64 75-84 85-h grouped only yes white, non-hisp. black, non-hisp. asian, non-hisp. hispanic no no figure 6 health and activity status variables for nursing home residents idphindicator nnhs iltcp activities bathing bathing bathing of daily dressing dressing dressing living eating eating eating (alds) walking inside walking mobility walking outside medication toileting toileting shopping or orientation transfer letter writing diagnosis drg if hospitalized no 15 groupings disease or condition of icd-9 resulting in admission 15 catergories such as hip fractures; prior health status 22 {assist quarterly payment source the ways that residents pay for care is an important issue in the delivery of care. the iltcp does not collect payment source. all sources of payment are indicated on the nnhs, while idph obtains only major payment source. the nnhs includes insurance within private pay, while idph breaks it out. nursing home trends despite their drawbacks and lack of standardization, the studies reviewed in this report can be used to form a picture of trends in nursing homes and their residents. the focus in this section is on recent changes in the industry. figure 7 payment source variables for nursing home residents idphnnhs medicare medicare medicaid medicaid private (own) pay private (own) pay va va state agency state agency other public pay other public pay church, foundation, agency insurance life care funds iltcp (1967-1986 unless otherwise noted) * the number of nursing homes nationally grew 18.2%. * the number of nursing home beds grew 105.0%. * the number of nursing home residents grew 106.5%. * persons 65+ in nursing homes grew 45.5% from 1967-1976, but then declined 6.2% from 1976 to 1986. * occupancy stayed around 92%. nnhs (1977-85 unless otherwise noted) * the number of elderly patients discharged from hospitals to nursing homes increased 36% from 1982 to 1985. * the proportion of nursing homes affiliated with a nursing home chain rose significantly 1977-85 from 28% to 41% of all facilities. * discharges rose 9.5%. * women in nursing homes aged 85-f rose from 34% to 4 1% of residents. * race and ethnic status was obtained for the first time in 1985. minorities were underrepresented and generally younger than the white population. * the proportion of individuals not dependent on help for mobility or continence dropped from 40.1% to 30.1%. * medicare-covered days in skilled nursing facilities per 1,0(x) beneficiaries dropped from 370 in 1977 to 320 in fall 1992 23 1984. * residency in a nursing home for persons 85+ dropped from 257 per 1 ,000 in 1974 lo 220 in 1985. * the average length of stay rose from 2.7 to 2.9 years. * the proportion of nursing home residents with mental disorders rose, while circulatory disorders fell. uicom-r gpph data for northwest illinois) * the ratio of beds per thousand population 65+ fell from 92.9 in 1980 to 85.9 in 1990. the elderly population is increasing more rapidly than the facilities for care. about 6.3% of persons aged 65 years and up reside in area longterm care facilities, down from 7.3% a decade ago. * the overall annual occupancy of general long-term care facilities in northwest illinois is 87.3% based on beds licensed. occupancy rates have increased since 1980 when the corresponding rate was 84.4%. * just under half (47.7%) of northwest illinois long-term care facilities are owned by for-profit concerns, down from 60.0% in 1981. the remainder are owned by churches (18.5%), government (13.9%) and not-for-profit corporations (20.0%). * hospitals are increasingly entering the long-term care "business." nine of the fifteen hospitals in northwest illinois have swing beds or distinct units. * most general long-term care residents are age 75 or older (81.0%), female (75.2%), and white (97.3%). residents aged 85 and over made up 47.7% of residents in 1990 but only 39.2% in 1980. the median age is now estimated to be 84.3 years, up from 82.1 in 1981. * leading admission diagnoses include the circulatory system (28.6%), nervous system (13.8%), and musculoskeletal (1 1.8%). one in ten residents is reported to have alzheimer's disease. mental illness, nervous system and musculoskeletal have risen, circulatory has declined. * nursing home residents are highly dependent on others for performing certain activities of daily living (adls) including bathing (62.8%) and toileting (56.6%). dependency rates have been increasing. * medicaid was the source of payment for 46.4% of the long-term care residents in 1989. private payers covered 48.9%. some of the remainder were covered by medicare (2.7%). medicaid coverage has been rising. * the average daily charge for a skilled care double bed in 1990 was s73, up from $48 in 1985. intermediate care averaged $56 for a double, up from $41 in 1985. many institutions make additional charges for supplementary services as rehabihtation therapies, bandages, or assistance in bathing. * 65.2% of area beds are certified for medicaid residents, but only 6.3% are certified for medicare. medicare certification has been stable or declining. improving nursing home data obra data requirements beginning this year, all nursing homes certified to provide care under medicare or medicaid must use assessment instfuments required by the state and approved by hcfa. the insffument must include a uniform minimum data set (mds) of care screening and assessment elements with common definitions. the mds contains a comprehensive set of data elements which describe most nursing home residents nationally. if modified to add certain elements, the mds could form the core of a standardized insu^ment which could be collected and analyzed on a national basis to analyze the characteristics of nursing home residents (see figure 8). other sources for homes certified by medicare or medicaid, financially-centered reports are filed with government or a fiscal intermediary. these "cost reports" generally include a great deal of information on the nature of the facility. (assist quarterly conclusion as has been shown in this paper, data describing nursing homes nationally is haphazard at best, and lacking in standardization of definitions. sources usually cannot be compared to each other, because of differing coverage, variables, and definitions. of the two federal sources, one is a f)eriodic sample which is relatively comprehensive, which employs several sub-component studies. the other is a periodic census with relatively sparse variables. neither follows a regular schedule. one was last completed in 1985 and the other in 1986. no immediate plans are currently in place for repeating these studies so as to yield more timely data. another data source is the surveys performed by individual states. most states have annua] surveys of the type exemplified by illinois. stales could form the framework for regular assessments of nursing homes as long as guidehnes for definition and collection are put forward by the federal government. much like the vital statistics system operates, states could report to the national center for health statistics, which could compile the information. long-term care is too important a component of health care not to have systematic data reporting. action is needed in this direction. 1 presented at the lassist 92 conference held in madison, wisconsin, u.s.a. may 26 29, 1992. fall 1992 sist newsletter vol.1, no. 4 discussion paper/ sue dodd titles: the emerging priority in bringing bibliographic control to social science machine-readable data files [mrdf] by sue dodd institute for research in social science university of north carolina chapel hill, nc two recent attempts to compile "catalogs" of social science data have encountered the lack of consistency among titles for the same data set. one attempt has been the recent cataloging efforts at the universities of north carolina, wisconsin, princeton, and yale, whereby traditional library cataloging records are created for social science data generated by academic research. the other has been the efforts of the association of public data users (apdu) to compile a directory of publicly available data files which represent primarily government produced data. both groups have experienced the same problem: variance of titles for the same data file. yet, without some control over titles and some mutually agreed upon primary source of title information, there can be no bibliographic control of social science data and none of the related products such as a union list of machine-readable data files. this paper will attempt to offer some suggestions for remedying the situation, including guidelines for transcribing titles; for creating a "title page"; for compiling a bibliographic reference; and for establishing an "authority list" for titles. origins of titles for social science data files unlike a book a social science data file may exist for a long time without a title. until it has been properly titled, it may be known only by a study number (£.£., study #5063), or by the name of the principal investigator (e.g^. , the stouffer study), or by the source of production (£.3.., the rand survey). if a data file survives the time period between data collection, data analysis, and data publication sans title, it is likely to assume the title of the primary publication (£.£., communism, conformity, and civil liberties ). the first appearance of a title usually occurs with the generation of early sources of documentation. documentation may include a questionnaire, coding instructions, codebook, manual, or project report. given the nature of mrdf, some type of accompanying documentation is required in order to "read" the data. titles recorded on documentation are also the most visible because "containers" (protective canisters) of mrdf have no identifying titles; or if they do, it is usually a shortened title given the space constraints of the container. however, an initially applied title of a mrdf is not necessarily the only or lasting one. during the life cycle of a social science data file, a title is frequently changed or modified as responsibility changes for file creation, processing, analysis and reporting. for example, one group of persons may be responsible for actual data collection plus the conversion to an "automated" format, while another group may be responsible for the data analysis and data reporting. such diversification of labor often leads to different titles for the same data. after the primary analysis, reporting, and possible publication by the principal 11 sist newsletter vol.1, no. 4 parties, a data file may be deposited with a data archive, center, or library for the purpose of secondary analysis. at this point, the data and documentation may go through further processing, including a new codebook and a new title. at about the same time, but not necessarily by the same person, a data abstract, study description or some type of informational notice is written to publicize its availability to the general public. if there are dual or multiple distributors of the same data file, titles could easily vary from distributor to distributor. for example, there are at least three known distributors for a particular harris survey with the following titles: violence in america the american public looks at violence harris 1968 violence survey, #1887 harris poll: "the american public looks at violence" in the case of the apdu directory, the overlap of mutually held and accessable public data files was impossible to determine, since members had listed the same data file under various titles. for example: city and county data book county and city data book 1972 county and city data book county and city data book tape the cataloging experience at the university of north carolina has revealed that out of approximately 500 separate social science data files from many different sources, close to 80 of these files had variant titles. in most cases, the title in the codebook varied from informational listings provided by the distributor of these data. as the life cycle of a data file continues, popularized titles begin to evolve and grow organically and are usually a modified version of the primary title. for example: modified title: french and german elite arms control data primary title: arms control in the european political environment: french and german elite responses, 1964 other titles are compressed into acronyms: modified title: the csep study primary title: the comparative state election project others take on the name of the principal investigator(s) : modified title: the matthews-prothro study primary title: the negro political participation study finally, variant titles may appear in "notes" or in bibliographic references in the various scholarly journals. without any guidelines on how to cite numerical data files and without any control over the proliferation of titles, title information will vary among scholars. often, the fault rests not so much with the person citing the data as it does with the distributing agency which has failed to provide proper bibliographic information on a particular data file. sist newsletter vol.1, no. 4 summarizing, the history of a data file usually reveals the various levels of title changes and modification. however, it is unlikely that a cataloger or a scholar will have access to this history. instead, he will be confronted with the problem of choosing or citing one title from among many for his respective uses. guidelines for transcribing titles in our attempt to offer suggestions on how this situation could be remedied, this section describes the basic components of a title and suggests guidelines on how to transcribe a title for social science data. components of a title for social science data files would include the following: 1) descriptive words indicating content; 2) geographic focus or unit; 3) chronological year(s) of data target or data collection; 4) source of data (e^.£. , court records); 5) producer, contributor or sponsor of data; and 6) study or series number (if important for ordering or for distinguishing one data file from another). guidelines include the following: i. make the title as descriptive and as complete as possible . a good title should be descriptive of the contents of the document or i data it is describing and should include as many of the components des' cribed above as are applicable. if there is one major theme or focus, ' then this should be mentioned in the title. if the data contain infori mation on many different topics, none of which appears to dominate, ' then a broader or more general subject approach may be taken (e.£. , harris 1972 public opinion survey; or the national opinion research 1 center 1974 general social survey; or the survey research center 1976 i social indicator survey). when transcribing a title, be aware that i the descriptive words contained in a title take on added significance i with the existing technology for keyword or full-text retrieval. for ' example, the only subject approach to social science citation index (whether it be by the printed reference work or by the on-line search capability) is via the descriptive words contained in the respective titles. ' ii. if at all possible, do not take a title from a publication based on the > data file , as this may cause copyright violations and problems with inf ternational coding schemes such as isbn (international standard book i number). if this cannot be avoided, then a qualifier should be attached at the end of the title . for example: i civic culture ( machine-readable data file ) communism, conformity and civil liberties (mrdf source documentation ) iii. for any data that are part of a predictable series (occurring at definite time intervals, such as the census or election surveys), titles should be consistent throughout the life of the series . iv. for data that are part of an on-going collection or series (collected at non-predictable intervals and with varying subject focus), one may consider the following sequential title arrangement : 1) organizational name of producer; 2) chronological date of data target or data collection; v 3) geographic focus (if unique); 4) descriptive content (including subtitles); and 5) study or series number. 13 sist newsletter vol.1, no. 4 an on-going series of data (such as public opinion polls) tends to be associated with the originating source or producer of these data. therefore, it is recommended that the organizational name of the producer come first. this arrangement also allows for a large collection of data from the same source to be grouped alphabetically for easy reference. for example: harris 1969 morals and values survey. no. 1933 harris 1969 science, sex and morality survey, no. 1927 in cases where there is more than one data collection per year on a given topic, the month or season could follow the year in parenthesis. for example: survey research center 1957 (fall) consumer attitudes and behavior survey, no. 3631 if the geographic focus is unique, it is recommended that it be included in the title. for example: american institute of public opinion 1975 japanese election survey, no. 7811 harris 1965 dallas sports survey, no. 1545 to indicate that data in a continuing series may have a varying subject focus, it is recommended that sub-titles be used. for example: survey research center 1963 detroit area study: a study of family-school relationships survey research center 1964 detroit area study: the measurement and validation of international attitudes study numbers should be included in the title if they are part of an on-going collection of data and are consequently helpful in distinguishing one data file from another. for example: harris 1967 public opinion survey, no. 1702 harris 1967 public opinion survey, no. 1718 national opinion research center 1963 (january) amalgam survey, srs-100 national opinion research center 1973 (december) amalgam survey. srs-4179 v. avoid beginning a title with articles (such as a, an. the. etc.). vi. avoid beginning a title with numerics (e.g., 1972 county and city data book) . with most computerized alphabetic listings, those titles beginning with numerics are placed either at the very beginning of a listing or at the very end. such placement may cause certain data files to be overlooked. vii. avoid using acronyms in titles . the full meanings of acronyms should be spelled out and if used at all. should follow full meanings enclosed in parenthesis. for example: world event/interaction survey (weis) viii. when applying titles to sub-sets of data files, indicate both the original data title and the fact that it is derived from a larger file. for example: comparative state election project: federal district sub-file 14 sist newsletter vol.1, no. 4 title control and an "authority list" of titles title control for social science data must be applied at one of two stages in the life of a data file: either at the production level or at the distribution level. ideally, the creator or producer of a mrdf should apply the "authoritative" title. however, in those cases where this responsibility has (for whatever reasons) defaulted to the distributor of the data, then he should provide the singular title. all other references to this mrdf should carry this title. one way to bring some order to the existing chaos among titles, is to establish an "authority list" of titles for social science data files. such an effort is being undertaken by members of apdu; and a similar "union list" of catalog records would have th'e same effect. again, the primary responsibility for establishing or determining the authoritative title rests first with the producer and then with the distributor. if the producer has abdicated that responsibility when depositing a data file with an archive or data center, then responsibility lies with the distributor. in those cases where there are multiple distributors of the same data, then a determination has to be made as to the one with the most "authority," or official status, or national prestige, etc. for data files that have been changed through major processing techniques or reformatted for a more efficient "reading," or have been changed in terms of content or observations, then the title remains the same but the data become a new edition. thus, an "authority list" of titles would include the various editions of mrdf, just as the national union catalog (nuc) carries the various editions of books. sub-files taken from larger data files should carry a distinctive title, and if not, the producer or distributor should modify the title with some type of qualifier (e^.c[. , sub-file a; selected sub-files, etc.). the major data producers and distributors of social science data would be responsible for publishing their respective lists of authoritative titles. these "authority lists" of titles could then be published in some appropriate newsletter or publication such as ssdata or the lassist newsletter . in establishing "authority lists" of titles there is also the need for establishing a concensus on the primary source of title information. if there is a title on the codebook and another title in a directory, which is the correct title? without having access to the history of a particular data file, how can the judgement be made as to the proper title? one answer would be to rate, in order of importance, the various sources of documentation. for example, the codebook or its equivalent would be the primary source of title information; the data directory or study description would be the secondary source-, the reporting source or publication would be the third, etc. some discipline has to be applied to social science documentation in general, but specifically to the bibliographic aspects of such documentation. such responsiblity should not end with titles, obviously, but should be extended to include all the components of a bibliographic citation. for example, the information required for compiling a bibliographic reference should come from the documentation accompanying a data file, and the most obvious place to derive this information would be from the "title page" of that documentation. 15 sist newsletter vol.1, no. 4 title page for social science mrdf in the past, very few data producers or distributors have taken the style or content of the title page of documentation seriously. however, if social science data files are to be readily accepted into the mainstream of bibliographic control, then more attention has to be given to these title pages. for example, the title page of a book becomes for the cataloger, the principal source of information. it is so respected by catalogers that the information contained on the title page becomes as "dogma" and cannot be deviated from in the transfer of information to the catalog record. however, the quality and amount of information provided on most title pages of social science documentation cannot be taken seriously by a cataloger. in truth, many sources of documentation for mrdf have no title page equivalent. information contained on a title page of mrdf documentation should consist of the basic bibliographic components including authorship; title; medium designator; edition; imprint; and series statement. the "medium designator" is a term used to denote the generic form or type of material listed or referenced. it is used to distinguish one type of medium from another and to provide clarity. the most universally accepted term for this medium is "machine-readable data file". the "imprint" includes the place of production; name of the producer; date of production; place of distribution; and name of distributor. the producer is defined as that party responsible for the collection, compilation, and physical production of the data (i.e., the mechanized process of transforming information into the format known as "machine-readable") and the distributor as that party responsible for disseminating the data to others upon request. the "series statement" would provide the reader with relevant information about an on-going collection of data (£.£., src/cps 1958 american national election study, no. 4). as mentioned earlier, an "editon" occurs when data files have been modified through major processing techniques or reformatted for a more efficient "reading" or when the data have been changed in terms of content or observations. edition statements appear in an abbreviated format on a title page (e.£. , dualabs ed., norc rev. ed., or 1st icpsr ed.). although, it is highly recommended that only the basic bibliographic information be placed on a title page, there may be situations where additional information would be helpful. examples would include: date of data focus, if not part of the title; source of funding, if appropriate; scope of documentation, if documentation consists of more than one volume; study number, if necessary for identification or ordering purposes; etc. information pertaining to unique classification schemes such as the international standard book number (isbn); the library of congress card number; and the catalog card facsimile should appear on the verso of the title page. for an example of a title page of a machine-readable data file codebook, see appendix. placement of the information is flexible. however, placement of a study number behind or immediately under a title will be construed as being part of that title. for example: 16 sist newsletter vol.1, no. 4 the src 1952 election study (s400) or the 1972 german election panel study (zentralarchiv nos. 635,636,637 -icpr no. 7102) bibliographic references for social science data files the title page, in addition to providing the basic information required for the catalog record, should also provide the information required to compile a proper bibliographic citation. the classification action group of lassist has been working to define the necessary components and structure for a proper bibliographic citation and has followed the guidelines provided in the forthcoming publication entitled: american national standard for bibliographic references . however, these guidelines have been modified slightly to represent the particular needs of social science numerical data and to conform with the forthcoming aacr ii cataloging rules. the basic components of the bibliographic reference would include authorship; title; medium designator; sub-title; edition; imprint (place of production, name of producer, date of production; place of distributor, and name of distributor); extent of file; notes; and series statement. the examples that follow were provided by the classification action group: title first: mexico's naturalized citizens, 1828-1931 [machine-readable data file]. principal investigators: harold sims, susan sanderson, and philip sidel. pittsburgh, pa : university of pittsburgh, 1975-76 [producer and distributor]. 1 data file (8,066 logical records). author first: shanas, ethel. the health of older people [machine-readable data file] : a social survey : public attitudes of older people . norc rev. ed. chicago : national opinion research center, 1957 [producer and distributor]. 1 data file (2567 logical records) and accompanying codebook (166p.). swidzinski, susan. syllabication [machine-readable data file] : a drill and practice lesson . bloomington, mn : control data corporation, 1976. on-line program lesson available only via the plato system. henry, neil. maxcls.bas [machine-readable data file] : a program for maxium likelihood estimation of parama ters of unrestricted latent class models . lafayette, in : gary income maintenance experiment, 1974 : pittsburgh, pa : social science computer research institute [distributor]. 1 program file (95 statements, basic) and accompanying manual (53p.). the information contained in the brackets, while highly recommended by the classification action group, are optional according to the ansi standards and the forthcoming aacr ii rules. the extent of file (logical records, program statements, etc.) and "notes" are also optional. many distributors of data already provide the user with some "data acknowledgement" information. guidelines or examples of how to cite these data in the literature should be part of this information. 17 sist newsletter vol.1, no. 4 conclusion in conclusion, we have discovered that it is not unusual for social science data files to receive many different titles in their lifetime. some titles are modified through data processing efforts; others grow or evolve from popular usage; and even others are erroneously recorded from one source to another. this lack of title control has proven to be detrimental to the efforts of those who are attempting to compile any authoritative listing or "catalog" of social science data. it is also apparent that there is an immediate need for an "authority list" of titles, and that some decision has to be made regarding the primary source of title information. it is hoped that this paper will bring these problems and needs to the attention of those parties who have both the responsibility for and the control over the situation. critical attention and immediate action, on the part of the major producers and distributors of data, is necessary to bring about true title control and better bibliographic documentation. without it, information on social science data files will remain in "elite obscurity". appendix machine-readable data file codebook berkeley radicals five years later: a follow-up survey of students who were arrested in the 1964 free speech movement at the university of california conducted by the detroit free press of kniyht-ridder newspapers, inc. under the direction of philip meyer and michael maidenberg social science data library university of north carolina chapel hill, north carolina 27514 18 "archival soundbites, footage, and photographs— past, present and future: the perspective of a documentary filmmaker and sociologist." by karl schonborn phd. ' professor ofsociology california state university hayward, california introduction documentary and educational film-makers trying to express sociological concepts have always had a great appetite for archival footage and photographs. in recent decades, this has come to include numbers and statistics and soundbites, too. this appetite will increase in a quantum fashion when hypermedia2 (also called multimedia) takes hold as an educational, instructional tool. i am a sociologist and the writer, producer and director of six major educational documentaries which have been rented and sold to universities and organizations across the united states. as a consequence, i have spent a good deal of time tracking down footage, photographs, and sounds to include in documentaries. therefore, i would like to share some thoughts about the storage and retrieval of such items in the past, currently, and in the future. my remarks will pertain primarily to documentaries and educational films, but they have relevance to many other instructional uses of such information. my favorite definition of the documentary is that it is a file or tape made with no love story, no plot, and no anticipation of profit. a more serious definition, befitting the archival focus of this paper, is presented by legendary documentarist/critic john grierson: documentaries entail "the creative treatment of actuality." relatedly, documentary pioneer dziga vertov claimed the documentary's task is to capture "fragments of actuality" and combine them meaningfully. radio and television news still use the term "actuality" to refer to the real world sights and sounds they record. documentaries and other educational films and tapes generally present "actualities" within a frame of reference or provide some interpretation. past regardless of the kind of documentaries they made, filmmakers have almost always spent time chasing down the images, words and sounds they needed to make their point(s). erik bamouw talks of the documentarian as biographer (e.g., p.m. adato and his "georgia o'keeffe"), historic chronicler (e.g., barbara kopple and her "harlan county, usa", explorer (e.g., robert flahrety and his "nanook of the north"), promoter (e.g., frederick leboyer and "birth without violence"), and guerrilla (e.g., peter davis and his "the selling of the pentagon"). "biographers" obviously utilized a good deal of archival data, whether more traditional ("paul robeson: tribute to an artist" 1980) or innovative ("wasn't that a time!" 1982 a celebration of the singing group, the weavers). these works are not to be confused with "docudramas" which likewise use archival data but which are really historical fiction rather than documentary (barnouw, 1983:309). and "historic -chroniclers" likewise were heavy users of archival material. since world war ii, film archives have proliferated because many new countries felt film archive collections would underscore their unique cultural traditions and beginnings. film-makers and collectors often gave their holdings to such collections. revisions of historic dogma often resulted as documentarists using these collections sometimes found "it just wasn't so." revisionist history got a boost from such debunking documentaries as "world at war" (1973), a re-examination of wwii by thames television; "men of bronze" (1977), a chronicle by william miles of a wwi unit of black americans who fought heroically for the french after being refused by general pershing who only wanted to command white soldiers; and connie field's "the life and times of rosie the riveter" (1980) which chronicled the important contribution american women made to winning wwii (barnouw, 1983:308). on a personal note, i remember in the 1970's going to the picture collection at the new york public library whenever i was "east"— being a californian— because i needed a picture of freud's mother or a shot of babe ruth's wife for my work. the remarkable thing about this collection besides its breadth was the easy access to it. an out-of-state person such as myself could get a lending card with no trouble and— even more amazing — could borrow numerous photos (usually mounted on pictureboard) for a long period of time, taking them 3000 miles away and mailing them back as i usually did. fall/winter 1990 in the past, before computers, keeping track of all the footage, still photo inserts, sound effects, etc., for a documentary was a monumental task. just coordinating the search for such materials was difficult since one often had to retrace ground all over if the image proved impossible to duplicate or the sound was filled with "noise" or other impedimenta. present the advent of the computer meant paper and pencil lists and scores of production assistants could be dispensed with. herewith is an accounting of how storage and retrieval works in many contemporary documentary projects. since the fdm making process is fairly well known (strips of film spliced together, editing tables, a & b rolling, etc.), the focus here will be on the videomaking process. computer list management and time-coding have made the handling of large numbers of actualities, words, numbers, images and sounds much more manageable for documentarians and makers of instructional videos. documentary-makers store and retrieve tape footage of actualities, etc. by means of "time-coding". 3 using a time-code generator, they put down an electronic signal along the length of the tape(s) on which they have recorded their needed information/data. this creates a unique "address" (in hours, minutes, seconds, and frames) for every bit of material on the tape(s). timecoding is often done when copies of the original tape are made. (generally, people work with copies while composing their work prints or "rough drafts" rather than risk erasing or damaging their original tapes. the originals are used again when the final print or "edit master" is created. when time-coded tape is played, the "address" is seen in a window— usually at the top of the tv screen. this 8digit address lets the documentarian manually or electronically find the exact place on the tape he or she is looking for. time-coding is a frame accurate version— with electronic markers— of the counter gadget seen on less sophisticated vcrs. lists of the "addresses" (locations) of the start and end of all tape segments can be put into a computer which then allows documentarians to pick and choose the footage, images, sounds, etc. he or she desires. a list of the desired segments can then be used to tell editors (or sometimes automated edit machines) how to assemble the documentary. the advantage of computers in present-day documentarymaking is that they greatly speed up the process of locating and retrieving moving images as well as still pictures, sounds, graphics and music. they also facilitate the ebb and flow of decision-making re.***sic*** what to include and exclude from the final cut of a documentary— heretofore, one of the most onerous tasks confronting the producer. some well-financed documentarians utilize videodiscs,4 especially for storage and retrieval of still images. they shoot pictures using a single frame 16mm film camera and then transfer the images to videotape with a telecine (essentially a film and video camera combination). most of them have then had to send the tape5 to a special lab to make the laser videodisc. this could take days or even weeks. mention should be made at this point of several sources of image data for contemporary documentarians. slide houses and stock footage libraries carry images of all sorts which can be bought unfortunately they are fairly expensive sources. a major newsmagazine also sells its hard-won still pictures on the open market and some news stations likewise sell footage. these latter sources — while also expensive— hark back to the "morgue" tradition started by newspapers. to be ready for any breaking story, and especially the deaths of famous people, newspapers created "morgues," i.e., files of photos and newsclippings of selected people and organizations. finally, sound effects are sold by the "needle drop" (a holdover from phono records); and, of course, canned music can be bought for a song and copyrighted music for an arm and a leg. future the future is already with us in some ways given that "digitization" of audio and video are current technologies. the main problem for documentarians is that digitized video is currently very expensive. in the high resolution, high fidelity age we live in, the shortcomings of analog 5 stored and retrieved video images makes digitization attractive as the technology of the future. (most analog video becomes fuzzy after several generations— copies of copies— but new "digitized" video continues to have sharp video resolution after scores of generations. digitization to digitize an image, it must be scanned using an electronic camera and a computer. each detail on the image— including light and dark variations— is assigned a number that is stored. if high resolution is desired, then digitization gets very expensive because millions of numbers must be stored. (a 5"" disc can only hold 2 or 3 images.) and color images require extra iassist quarterly memory, often a megabyte at the bare minimum. today, it is more cost-effective to use analog information. one can store 54,000 video pictures on a laser videodisc. but digital compression techniques and other technological breakthroughs should probably make digitization the technology of the future. digitized images can come from almost any source, but it helps to start with high resolution pictures. a character generator is used to put an identification "address" on each picture. this is entered on the original image as well as onto a computer using a list management program. documentarians are then able to enter, find, and extract data from the database. when recording still frames of video on a disc, care must be taken that there is no dropout, jitter, or chroma crawl. video encoders and filtering processes can sometimes deal with these problems. once an image has been digitized, one can play with it — erase parts of it, increase it, alter it any way ones wants. with high speed digitizers, it is possible to change things in a moving video scene. (high-speed digitizers— called frame-grabbers— can capture images in one-thirtieth of a second. conventional digitizers may take several minutes, especially if quality reproductions are desired.) this ability to partially erase or increase a digitized image will solve one problem documentarians face: encountering an image that is so cluttered with distracting elements that the key element is missed or one that is too small to be clearly shot even a micro lens cannot always capture the essential element for example, in my "stigma" documentary, being able to zoom in on a keloid scar (a thick fibrous scar) on a person would have eliminated my having to consult a medical journal for an extreme close-up picture of a keloid scar. organization the continuing explosion of audio and video information and "actualities" means that archivists and librarians will have to have highly organized systems for selection, storage, and retrieval of audio and visual information. like a family trying to find a shot of great-aunt betsy in a shoebox of color slides or in a cabinet of home movies, they might easily become overwhelmed. all this assumes that archivists will have the proper training or help to guarantee the "quality" of the images and sounds they store. (resolution, clarity, audio and video levels, lighting, composition, and security are part of the "quality" issue.) let us look closely at the matter of "selection," especially the challenge of selecting which "actualities" to store. while we might debate whether the number of truly "significant" events and verbal utterances has increased in recent times, there is no question that the record of events and words has. professional photographers, journalists, media departments, and— yes, amateurs— record countless scheduled and unscheduled events, speeches, and goings on. most amazingly, amateur recordings, such as the zapruder film of the kennedy assassination which was rare and freakish in 1963, are now common. tourists and everyday people with av gear record planes crashing, bridges collapsing and people being shot. there are important questions to answer: who will decide from the avalanche of possibilities which events, speeches, etc., are significant and therefore worth saving electronically? probably committees composed of historians, social scientists, journalists and the like. and what criteria will they use? will archivists automatically save the inaugural and farewell speeches of anyone who has attained a certain political level, say senator or governor? (this would probably be too status and stratification bound and would mean non-establishment speeches (like martin luther king, jr.'s "i have a dream" speech ) would not be stored.) will certain people and concepts be favored during certain periods because of their cultural centrality or importance? maybe early tv shows and personalities ("mickey mouse club," phil silvers,) would be emphasized for the 1950s, counterculture happenings for the 1960s, professional athletes for the 1970s, and entrepreneur and ceos for the 1980s. no matter what criteria are used, the selection procedure is bound to be difficult. deliberations over what is significant or not quickly get into issues of "reality." it seems as if more americans are able to identify certain tv commercials than identify the state of ohio on a map. (ditto for certain hollywood movies.) apropos of this, there are clio awards (clio was the muse of history) given each year for outstanding ads of various sorts. as a consequence, archivists will have no trouble documenting our "reel" history (according to madison avenue and hollywood). it's just our "real" history that will be difficult to document. some further thoughts and recommendations for future information storage and retrieval, though from the limited perspective of a documentarian (i base these recommendations partly on the assumption that everyone will someday be assembling words, pictures and sounds — for institutional use or for personal use (much as people assemble scrapbooks and photo albums today): fall/winter 1990 audio and video data should be stored digitally on videodisc. film and videotape are too fragile, necessitating recopying every so often because of oxidation and other chemical processes. audio and video information might best be cataloged and stored by social science discipline. these databases would resemble the cd-rom 6 databases which presently exist for some of the social sciences: e.g., psyclit (die apa's psychological abstracts), eric (educational resources information center), and infotrac. mechanisms similar to those used by the united nations to gather and report data might be used to insure that important audio and video data around the world are preserved and made accessible. the database might be organized along the lines of medline (published by the national library of medicine) which is international in scope. if possible, the databases should not overemphasize western and european audio and video. all databases might be updated quarterly and then reassessed at the end of each year and each decade— much the same way that newsmagazines summarize the year and decade in words, pictures, and the like. conclusion over the years documentarians and others making educational films have utilized increasingly sophisticated means to obtain the archival footage, photographs and sounds they needed for their work. the amount of this archival data has grown exponentially, and the need to access it will too as interactive learning and hypermedia find their own niche alongside documentaries. digital storage on videodisc probably represents the most realistic way for these immense amounts of data to be stored and accessed easily. much ongoing effort will be required to select, record, and organize the untold numbers of words, numbers, images and such. the cataloging, tracking, and safekeeping of audio and video data, though should be well worth it. nothing comes quite alive like the voices, gestures, and actions of the greats (and even the not-sogreats) of yesteryear. in a sense, archivists of the future will be in the business of freezing people and then bringing them back to life but electronically rather than cryonically. bibliography barnouw, erik, documentary history of the non-fiction film, oxford university press, oxford, 1983. hardy, f., grierson on documentary . praeger, praeger, 1971. st. lawrence, j., "still the image," videographv . february 1990. stevens, george, "applying hypermedia for performance improvement," performance and instruction . july 1989. 1 presented at the iassist 90 conference held in poughkeepsie, n.y. may 30 june 2, 1990.professor karl schonbom, california state university, hayward, california 94542,415 881-3173 2 hypermedia (or multimedia) refers to the coordinated use of more than one medium sound, video, animation, graphics and text for example. moreover, hypermedia can mean all these media plus "interactivity," where the medium responds to the user and vice versa. users, typically, explore hypermedia materials at their own pace, moving in different directions through the information, creating their own interpretations, involvements, experiences. 3 time-code refers to the 8-digit address code used to identify each video tape frame by hour, minute, second and frame number (allows frame-accurate precise editing). 4 videodiscs are like 33 1/3 phonograph records and they store video information. often used for instant "replays" in sports broadcasts, videodiscs allow for user friendly instant access to information as well as slow motion, freeze-frame, and interactive video effects. a major advantage of discs over tape, though, is there is no diminution of picture sharpness when dubs (copies) of the image are made. 5 in analog processing, an electrical signal varies over a continuous range to represent the original audio or video that is being reproduced. in digital processing, electrical impulses are either on or off just as they are in computers. one digit (0 or 1) is generally represented by the "on" state of an electronic device while the other is represented by the "off" state. one can store visual images on laser-based read/write optical discs as analog information or digital information, i.e., as video or as computer data. 6 a cd-rom (compact disc-read only memory) allows high density storage, having a capacity of 550 megabytes (roughly equivalent to 1,500 floppy discs or 100,00 pages with 5,500 characters per page). iassist quarterly vol26no1 12 iassist quarterly spring 2002 iassist quarterly spring 2002 13 by robert g. cromley & patrick mcglamery * abstract as more spatial databases are compiled and made available for dissemination via the internet, there is an increasing need for metadata descriptions of the downloaded data to be available on demand. this paper describes a system in development at the university of connecticut that not only allows users to select their data dynamically but also to prepare fgdc compliant metadata records for these data that can be downloaded in the same session. in this manner, users will have a more complete knowledge of how to use this information effectively. introduction thematic maps have been a part of research library collections for the past century. only since the later 1970s though has the choropleth map become a ubiquitous tool for visualizing demographic information in the united states and have these thematic maps found their way in growing numbers into the map library. the 1970 census and the dime (dual, independent, matching, encoding) files, dependent though they were on large mainframe computers, set the stage for our expectations of network delivered, on-demand demographic mapping and data dissemination. during the 1990s, radical innovations in technology increased the computing power of the personal computer (pc), developed the cd-rom as a distribution media, and created a burgeoning network infrastructure and browsing software, resulting in the worldwide web (www) as we know it today. while all of these technologies have resulted in the growth of data dissemination from map libraries as well as social science data libraries, perhaps the most significant technological change was the replacement of magnetic tapes by cds. this particular shift ʻdemocratized ̓social science data by moving it from the constraints of mainframe computing and putting the data in libraries and hence on scholarʼs workstations. in the meantime, geographic information systems (gis) have been changing the ways in which we look at geographies. technology has had a huge impact on gis, freeing it first from the mainframe, then from high-end unix workstations. powerful pcs and our networked integrating spatial metadata and data dissemination over the internet environment on the www have created a dynamic, online versions of gis. (green and bossomaier, 2002). now gis is expanding the user base of who is looking at geography… and maps. these innovations have put pressures on the map library. in the past two years over a dozen new “gis librarian” positions have been established and staffed in research libraries around the united states. in fact, the gis librarian is a geodata librarian focusing primarily on tiger (topologically, integrated, geographic, encoded, reference) line graph files, demographic data and mapping; this is a result of the association of research libraries ̓(arl) gis program that began in 1993 (arl, 1995). the goals of the arl gis literacy project were designed to meet the current needs of libraries and users while addressing the changes that libraries are facing during this time period of experimentation, transition, and transformation to networkedbased services. these goals include the following: • introduction of gis to a variety of libraries (e.g., public, state-based, academic, and university libraries in public and private institutions) to address diverse user information needs; • development of a team of gis professionals in the research library community who are willing to lend time and expertise to applications, user training, and education programs; • encouragement of connections among federal, state, and local gis users and information; • promotion of research, education, and the public right-to-know through improved access to government information; • initiation of library projects to explore new applications of spatially referenced data and to evaluate the introduction of these services in research libraries; and • implementation of programs to allow institutions that have invested in networking capabilities to leverage the sharing of resources via networks. working with esri, a gis software developer, the arl 14 iassist quarterly spring 2002 iassist quarterly spring 2002 15 program worked to reorient libraries toward considering the geoprocessing of social science data, especially census data. traditionally, map libraries have had the responsibility for the storage of and access to cartographic information. a library collects, catalogs, stores and provides access to information but does not produce the data itself. increasingly, the producers of spatial data are only distributing it as a digital database; not as paper or even as viewable information. in 1990, the u.s. bureau of the census stopped producing census tract and other printed maps (for the 2000 census some maps can now be downloaded as pdf files). this almost total conversion of information from a paper to an electronic format has forced map libraries to rethink how services are to be provided. metadata with the development of the internet in the 1990s, an opportunity was created for transmitting digital cartographic and social science information to a global user community. however, as more databases are compiled and made available for dissemination via the internet, there is also an increasing need for descriptions of the downloaded data to be available on demand. these ʻdata about data ̓or metadata traditionally play four roles in the archiving of data (fgdc, 1997): • availability – data to determine what data exists • suitability of use – data to determine proper data uses • access – data to acquire a set of data, and • transfer – data needed to process a set of data. metadata regarding the content, quality, and other characteristics of disseminated data are critical not only in transferring data from an external source but also in interpreting that data by the end user; that is, it acts as a codebook. codebooks describe the structure, contents and layout of a datafile. the purpose of this paper is to discuss the issues surrounding the construction of metadata records for customized databases, especially dynamic spatial databases with demographic attributes, disseminated over the internet. digital map formats in a recent book on web-based cartography, kraak (2000) defined two types of web maps: static maps and dynamic maps. static maps are images displayed on a browser whose elements cannot be changed. on the other hand, users can not only change elements of dynamic maps but also interact with them. however, although the visualization of some maps may be static, the process of their construction may be dynamic. a static map may exist as a predefined image or, using the common gateway interface (cgi), a server-side utility program can create a customized image (kobben, 2000). the same dichotomy can be applied to spatial databases being disseminated over the internet. some spatial databases may be static predefined files that are downloaded or, some may be dynamic customized files that are created by the user in real time. for static spatial databases, metadata descriptors are contained in separate files that can be viewed and/or downloaded with the database. dynamic spatial databases create new questions for metadata records that are not adequately covered by existing standards because the descriptors of the data set being transferred are in part created in real time. in constructing dynamic spatial databases, both the geography and/or the attribute information can be defined by the user, and the user can select one or both. one cannot anticipate what the user will select. therefore, descriptors need to be created a posteriori for dynamic spatial databases whereas descriptors for static spatial databases can be created a priori. the nature of metadata metadata is a dual-use concept. the metadata content in an html (hyper-text, markup language) or txt (plain text) formatted file is a codebook for the user. as an sml (standard markup language) or xml (extensible markup language) formatted file, it is tagged to create an index for search, query and discovery in a clearinghouse network. if the content elements are stored in a database, a print statement can generate metadata information in either an html, txt, sml, xml or any other format depending on the needs of the user. however, a complete storage of content elements is only possible for static spatial databases. for dynamic databases some content elements that are known a priori can be stored in a database whereas other elements that are defined by user choices and are created “on-the-fly”. since most users downloading data are mainly interested in the codebook aspects of the metadata, this paper focuses on the codebook concept for the a posteriori construction of metadata for dynamic databases. metadata standards this discussion is on the application of federal geographic data committee (fgdc) standards to dynamic spatial databases because these are the national standards used by most distributors of spatial metadata. there are seven major components of this metadata standard (fgdc, 1997): • identification information which contains basic characteristics of the data set including a description of its content, its spatial domain and its time period of content; • data quality information that provides a general assessment of a data setʼs quality and suitability of use, • spatial data organization information that describes the mechanism used to represent spatial information in the data set; 14 iassist quarterly spring 2002 iassist quarterly spring 2002 15 • spatial reference information that describes the reference frame used to encode spatial information; • entity and attribute information that outlines the characteristics of each attribute including its definition, domain, and unit of measure; • distribution information that identifies the data distributor and the options of obtaining the data, and; • metadata reference information that describes the currentness of the metadata and the party responsible for maintaining it. in addition to these main components, citation, time period, and contact information are important subcomponents that are repeated under different primary components. while the fgdc standards provide for extensive description of spatial characteristics, its abilities with entities is less robust. the social science data communityʼs data documentation initiative (ddi) standards are more definitive. while the developing ddi is actively adopting geodata descriptors, fgdc, bearing the burden of a seven-year legacy, has been less responsive. spatial data issues there are different issues associated with the spatial database if the underlying geography is area-class or choroplethic. for an area-class geography, the geo-units are defined by an attribute for which the associated attribute set is well defined and closed (probably better suited for a static database). for a choroplethic geography, the geounits are usually defined as political or administrative entities for which the associated attribute set is more openended (probably better suited for a dynamic database). these differences have implications for metadata records. for example, an area-class database is more likely to have the same originator for the geographic and the attribute information because these are intertwined. a choroplethic geography is more likely to have different originators one for the geography and perhaps multiple originators for the attributes. likewise, the geography and attribute lineage information for an area-class database should be the same, whereas these lineages should be different for a choroplethic database. there are numerous examples then where a clearer distinction should be made in the construction of the geography from the construction of the attributes. there are two additional considerations. given that the internet is a distributed network, the data behind a customized database can be distributed over many organizations and locations. the metadata for the customized database needs to capture the complexity of this system. some users will only want attribute data and not the geography for use in non-visual analyses. because the attributes are geo-referenced, though, some description of the geo-units is still necessary. constructing a dynamic system at the university of connecticut, we have been working on developing a system to generate customized spatial databases and their associated metadata records. the system is designed to create a full spatial database (in progress) or just a geo-referenced attribute database. in this system, metadata are generated both from existing metadatabases as well as the userʼs own query responses. the metadata in these databases are also used by the system to compile and retrieve the customized database. in addition the user defines: 1) a study area (the spatial domain), 2) a geographic unit of inquiry (the spatial resolution), and 3) the set of attributes. the tasks of meta-database organization involve: 1) determining the fgdc metadata elements relevant to a geography database and those relevant to a geo-referenced attribute database; 2) separating those elements with both geography and attribute descriptors into separate tags; and 3) deciding which elements cover a whole database and which cover elements of the database (figure 1). assigning the basic fgdc metadata elements is relatively easy. identification, data quality, spatial data organization, entity and attribute, distribution, and metadata reference information all belong to both, whereas spatial reference information is only relevant to the geography. separating elements is more difficult. for example, within identification information one can separate citation, description, time period of content, status, and native data set environment into spatial data and attribute data tags. on the other hand, spatial domain, keywords, access constraints and use constraints only have one tag. deciding which elements cover a whole database and which ones cover elements of the database is also more difficult. customizing the geography means that spatial reference information is not at the level of a whole geography database but at the level of the elements within the database whereas spatial reference is at a whole geography level. customizing attributes means that originator and lineage may belong at both the dataset and individual field level. while our current development has been for fgdc content standards, we are creating a meta-tag database for cross-walking between different metadata standards such as fgdc, ddi, the ferret projectʼs mif (metadata interface file), dublin core, and the library communities marc formats. for example, the name of an attribute has the following meta-tags: attribute_label in fgdc; attr name in ddi or m in mif. this will enable users from a variety of research communities to access metadata codebooks in familiar formats. it will also streamline 16 iassist quarterly spring 2002 iassist quarterly spring 2002 17 metadata creation for clearinghouse indexing. conclusions when designing a dynamic spatial database dissemination system, there are more considerations necessary in the construction of metadata records than for a static system. the nature of the dynamic system requires that a metadata codebook reflects the impromptu nature of the data query and extraction. some queries will be directed to obtain only attribute data that can be used in statistical analyses; other queries may want the full geography as well. the metadata needs to reflect these choices made by the user. in addition, demographic and other social science attribute data, may have different source organizations and locations. the metadata codebook needs to capture the complexity of the source information. by providing full narrative information for numeric datafiles, users have a value-added product. bibliographic references association of research libraries (1995), the arl gis literacy project, ftp://www.arl.org/info/gis/gis.descrip. federal geographic data committee (1997), fgdc standards reference model, http://www.fgdc.gov/ standards/refmod97.pdf. green, d. and t. bossomaier (2002), online gis and spatial metadata. new york: taylor and francis. kobben, b. (2001). publishing maps on the web. chapter 6 in web cartography, m-j. kraak and a. brown (eds.), new york: taylor and francis. kraak, m-j. (2001), settings and needs for web cartography. chapter 1 in web cartography, m-j. kraak and a. brown (eds.), new york: taylor and francis. * paper presented robert g. cromley+, department of geography u-4148, 215 glenbrook rd., university of connecticut, storrs, ct 06269 usa, phone (860)486-2059, cromley@uconnvm.uconn.edu. and patrick mcglamery, homer babbidge library, university of connecticut, storrs, ct 06269 usa +all correspondence should be directed to this author. data organization databases of databases databases of entities metadata pertaining to each geographic attribute database metadata pertaining to objects within a geographic attribute database figure 1. database organization ftp://www.arl.org/info/gis/gis.descrip fj^ssist newsletter vol. 1, no. 3 action group reports data archive registry canadalisa lasko, canadian consortium for social research, institute for behavioral research, york university, 4700 keele street, downsview, ontario m3j 1p3 europejoseph bonmariage, belgian archives for the social sciences, university of louvain, sh-2, 1348 louvain-la-neuve, belgium united statesjohn kolp, regional social science data archive, university of iowa, iowa city, iowa 52242 report of the joint canadian-united states action groups on data archive registry meetings in toronto, may 11-12, 1977 (par ag) submitted by lisa lasko institute for behavioral research york university members present canada lisa lasko, institute for behavioral research, york university sharon henry, data clearing house for the social sciences gerald prodrick, university of western ontario jana prokop, university of toronto the canadian data archive registry group made substantial progress when it met for a working session at the canadian lassist meetings. may 11 and 12, in toronto. the working sessions were attended by four canadian lassist members. no american members were present. the first issue that was addressed by the group was whether or not the data archive registry group still held a viable mandate. the presence of two new data archive directories— the upcoming 1977 september edition of vivian sessions' directory of data bases in the social and behavioral sciences and the unesco sponsored directory of data services now underway under the direction of jean meyriat, secretary general of the international committee for social science information and documentation (icssd) forced consideration of this issue. in addition, the canadian group had to consider that the data clearing house for the ibissist newsletter vol. 1, no. 3 social sciences intended to produce a hard copy directory of social science data centres for canada. the group agreed that the first two directories were less than satisfactory, but that the soon-to-be-produced directory of the data clearing house posed a real conflict. sharon henry, executive director of the data clearing house, reported that although she had to have the directory published by the end of the current year, in fact, no work had yet been started on the publication and no questionnaire had yet been designed. it was suggested that a formal liaison between the lassist data archive registry group and the data clearing house be established to produce the directory. more specifically, the questionnaire utilized for surveying canadian data centres would be jointly designed by both groups, and the data elements covered in the questionnaire would allow the results to be used by both groups for their various needs. one publication, the directory of social science data centres , would be published cooperatively by the data clearing house for the social sciences and the lassist data archive registry group. however, it was clearly understood that when and if an lassist international data archive registry did get underway, the relevant canadian entries would be pulled from the data base for inclusion in the lassist registry. the canadian directory, then, was to be viewed as a pilot project for the international lassist data archive registry. it was agreed that in order for this joint activity to work, a formal representative from the data clearing house for the social sciences should become a permanent member of the lassist data archive registry group. once this basic issue was resolved, the group immediately started to work on designing the questionnaire. both the session's questionnaire and the meyriat questionnaire were consulted. by the end of the two days, most of the necessary elements had been agreed upon and pierre lacasse, coordinator of the data acquisition group, made some valuable suggestions for elements concerning acquisition policies of archives. however, as work on the questionnaire did not get completed, it was decided that lisa lasko and jana prokop would meet in toronto to complete it, after which time, a meeting would be held with sharon henry to discuss it further. the group spent a substantial portion of its time compiling a list of broad subject headings to be used for describing categories of data in the questionnaire. during the conference, it was pointed out that the lack of a controlled vocabulary for descriptions of categories or holdings of data, was a major factor in the lack of good subject access to data archives. the group felt that its work in this area would be a significant and important contribution to useful descriptions of data archives. the group built upon the list of 105 subject headings compiled by david gerhan and loretta walker, and also consulted with vivian sessions' 26 broad subject categories. in addition, unique canadian subject terms were chosen. the final version differed substantially from the aforementioned lists, and, as it was still very rough, lisa lasko and jana prokop agreed to refine the list further at a later date in toronto. sue dodd, coordinator of the classification action group, generously offered to assist in the compilation of subject terms. it was agreed that the refined list would be sent to sue dodd for comments and suggestions as soon as possible. _ to summarize, then, the data archive registry group made significant headway in designing the questionnaire and developing a controlled vocabulary to be used for descriptions of data archive holdings. however, further work is' still required to complete both the questionnaire and subject heading list. for this reason, lisa lasko and jana prokop both agreed to meet in toronto in the next few weeks to take care of any work still outstanding. ^ssist newsletter vol. 1, no. 3 classification canadamohan sharma, humanities s social science library, university of alberta, rutherford north, edmonton, alberta europeekkehard mochmann, zentralarchiv fur empirische sozialforschung, bachemer strasse 40, 5 koln 41, federal republic of germany united statessue dodd, data library, insitute for research in social sciences, manning hall, university of north carolina, chapel hill , north carolina 27514 report of the joint canadian-united states action groups on classification (c ag) submitted by sue a. dodd, university of north carolina members present canada martha amschutz, canadian radio-television commission krystyna w. dynowski, university of western ontario sue gavrel , public archives of canada, ottawa katrin horowitz, library of consumer & corporate affairs hans g. schulte-albert, university of western ontario united states sue a. dodd, ch, university of north carolina gertrude lewis, rutgers university mimi schade, brookings institution philip sidel, university of pittsburgh agenda topics : current problems and associated tasks within the mandate of the classification action group include developing (1) examples and guidelines for bibliographic references for social science numerical data files, (2) a "cataloging-in-production" scheme for major producers of social science data files, (3) a more "universally based" classification scheme for social science data files; and, reviewing (4) the cataloging efforts to date and discussing any or all related problems and (5) existing printed thesauri in the social sciences in terms of their future applications to data files. the primary focus of the c ag at the canadian working conference centered around the first problem of how to cite properly a social science numerical data file in the published literature. ji^sist newsletter vol. 1, no. 3 problem : currently, there are no standards or guidelines for compiling such bibliographic references and the quality and amount of information provided when citing social science data files varies greatly among individual scholars. in many cases, the information is not sufficient to allow for direct access, and consequently, an interested party has to spend a considerable amount of time determining additional information that could easily be provided by a full bibliographic reference. complicating this problem is the american national standards institute's forthcoming work entitled "american national standards for bibliographic references." this ansi standard will provide detailed information on the theory, principles, and definitions underlying the technique of preparing bibliographic references for both print and nonprint materials, including machine-readable data files (mrdf). the problem with this forthcoming document is that it does not include examples of social science numerical data files which make up the vast majority of mrdf; nor does it attempt to be compatible with the forthcoming second edition of the anglo american cataloging rules (aacr ii), which also deals with bibliographic standards for mrdf. role of lassist c ag : the classification action group, in its continuing role to encourage uniform standards for mrdf and to see that social science data files are represented in the best possible light will provide a written critique on the forthcoming ansi standard, and at the same time provide more serviceable examples of how to cite social science numerical data files in the published literature. the c ag does not intend to create a new standard but rather to work within the framework of the ansi standard. the intent is to supplement it and recommend revisions where it is viewed as either necessary or helpful . the critique will focus on the omission of social science numerical data files and the lack of compatibility with other standards and related terminology dealing with the same medium. arguments to be presented in the paper are: that outside the field of quantitative research, social science numerical data files are probably the least understood category of mrdf and, consequently, are in tfie most need of attention and clarification; that the terminology and definitions used to describe mrdf are confusing and often misleading and, therefore, a uniform standard for this medium is necessary; that social science numerical data files make up the largest on-going collection of mrdf and, consequently, this body of information is in immediate need of bibliographic control; that social science numerical data files are cited more frequently in scholarly journals than the two mrdf examples listed in the ansi standard (i.e., computer programs and bibliographic data files), and therefore, the need for relevant social science examples seems justified; that this justification is magnified by the fact that the manner in which social science data files are cited in the literature plays an important role both in terms of access and in terms of secondary analysis; and, that the two examples given (i.e., computer programs, and bibliographic data bases) are significantly different from the characteristics of social science numerical data files as to be less than helpful. rai$sist newsletter vol. 1, no. 3 the first draft of this paper was reviewed by the participants of the canadian working conference and will likewise be reviewed by those c ag participants who were unable to attend the conference. the final work will be reviewed by the c ag's standards and quality-control review board. the resulting critique, along with an accompanying letter signed by the appropriate officers of lassist, will be sent to the american national standards institute, inc. in new york and beyond that, it is hoped that the recomnended revisions and resulting examples would be accepted by this body. it is also hoped that relevant examples could be sent to the editors and publishers of the various social science journals, who in turn can begin to implement conventions on the use, style and content of bibliographic references. tasks of c ag participants : (1) read and review the draft of the critique of the ansi standard, providing any comments, suggestions, or additions, etc.; (2) read an abbreviated version of the standard prepared by ellis mount and which appeared in the journal of the american society for information science ; (3) prepare a list of the necessary and descriptive bibliographic elements (e.g., title, author, edition, imprint, etc.) for social science data and at the same time indicate those that would be essential , recommended , or optional ; and, (4) utilizing the list of rated bibliographic elements, begin to compose bibliographic references of the respective data files. results of this task at the canadian conference : the first draft of the critique was reviewed and comments accepted. the beginnings of a list of bibliographic elements was compiled. each element was grouped and rated (i.e., essential, recommended or optional) according to its importance in identifying a data file. finally, some eleven examples were compiled by the participants, including the samples listed below: titlemexico's naturalized citizens, 1828-1931 [machine-readable data file], first: harold sims; susan sanderson; philip sidel , principal investigators. pittsburgh, pa: university of pittsburgh, 1975-76. 1 data file (8066 logical records). authorshanas, ethel. the health of older people [machine-readable data file]: first: a social survey: public attitudes on older people. norc rev. ed. chicago: national opinion research center, 1957 [producer and distributor]. 1 data file (2567 logical records) and accompanying codebook (166 p.). swidzinski, susan. syllabication [machine-readable data file]: a drill and practice lesson. bloomington, mn: control data corporation, 1976. on-line program lesson available only via the plato system. henry, neil. maxcls.bas [machine-readable data file]: a program for maximum likelihood estimation of parameters of unrestricted latent class models. lafayette, in: gary income maintenance experiment, 1974; pittsburgh, pa: social science computer research institute [distributor]. 1 program file (95 statements, basic) and accompanying manual (53 p.). the work at the canadian conference on this task was an important first attempt towards achieving the goal of providing relevant and acceptable examples of bibliographic references of social sciences data files. however, these examples should be viewed as only first attempts and not as the final recommended examples of bibliographic references; and, any and all comments are welcome. 10. i^aksist newsletter vol. 1, no. 3 data archive development canadalaine ruus, data library, computing centre, university of british columbia, 2075 wesbrook place, vancouver, british columbia v6t 1w5 europenot activated united statesalice robbin, data and program library service, 4452 social science building, university of wisconsin-madison wisconsin 53706 report of the joint canadian-united states action groups on data archive development meetings in toronto. may 11-12, 1977 (dad ag) submitted by laine ruus, university of british columbia members present canada laine ruus, ch, university of british columbia marilyn berry, university of victoria peter clinton, memorial university of newfoundland judy demaine, university of guelph ed hanis, university of western ontario alan kirby, queen's university elaine kozak, data clearing house for the social sciences united states john heddesheimer, national archives and records service, washington, d.c. alice robbin, university of wisconsin-madison judith rowe, princeton university ed vickery, research triangle institute, chapel hill, nc objectives of the meeting: (a) to revise and redefine the outline of the guide to providing social science data services ; (b) to elaborate and expand on the contents of the separate sub-sections of the guide ; (c) to finalize arrangements for contributions from the other action groups to relevant sub-sections of the guide , and set deadlines for abstracts and full-text contributions; (d) to allocate responsibility for contributing remaining sections and sub-sections of the guide to persons or groups with relevant expertise. the first priority of the action group was to write abstracts of those sections of the guide to be contributed by other action groups, so as to afford them a chance to consider more fully, during their meetings, their contributions and react to them. on the basis of these reactions, some sub-sections of the outline will be redefined, and some responsibilities reallocated. 11. ji^sist newsletter vol. 1, no. 3 specific responsibilities for remaining sub-sections of the guide were allocated to such persons of expertise in relevant areas as could be identified by the members of the action group, both within and outside the action group itself. those who could not be contacted during the course of the conference, or who had not previously been contacted, will be solicited for their collaboration immediately. a deadline for contribution of abstracts of all sections was set for june 30, 1977, and deadline for full -text for december 31, 1977. in addition, a system of readers, for the several sections of the guide , was established to provide a pre-editorial review of input, and to ensure that foci of the sections include minimal as well as maximal levels of service. it was further decided that preliminary copies of the glossary should be distributed to all contributors and readers for comment and ammendment, to ensure that all vocabulary necessary is included, and that definitions in the glossary agree with usage by the contributors. based on reaction from members of the data organization and management action group, it was found to be necessary to further refine and amplify the definition of the target audience, especially in terms of levels of expertise. a fuller redefinition will be distributed to all collaborators. the final order of business was a decision on a topic for the dad-ag symposium at the february 1978 meetings. various topics were considered, and the final choice was 'networking', as a topic of some immediacy for data archives/libraries. consideration will be given to a definition of 'networking', and its implications for all machine-readable data service units. process-produced data canadajohn devries, social science data archives, department of sociology, carleton university, ottawa, ontario kis 5b6 europepaul muller, institute for applied social research, university of cologne, greinstrasse 2, 5000-koln 41, federal republic of germany united statesdonald harrison, national archives (nnr), washington, d.c. 20408 12. sist newsletter vol. 1, no. 3 report of the joint canadian-united states action groups on process-produced data workshop in toronto, may 11-12, 1977 (ppd ag) submitted by john devries, carleton university members present canada john devries, ch, carleton university bill bradley, health and welfare canada, ottawa hy burshtyn, carleton university tony falsetto, public archives of canada, ottawa richard guttormson, ministry of state for science and technology, ottawa pierre lacasse, centre de recherche en amenagement regional, sherbrooke, que'bec united states harriet dhanak, michigan state university shirley gilbert, princeton university elizabeth powell, leaa/ncjiss, department of justice, washington, d.c. ed vickery, research triangle institute, chapel hill, nc we continued our discussions on the directory of catalogues which list machine readable data , which we had determined to have the highest priority for our ag. although we did not meet our self-imposed deadline (which expected us to produce a first draft of this proposed document at this meeting), some progress has been made. we have obtained copies of several existing catalogues and propose to continue this collection phase. we developed a listing of minimally required elements of information, which we hope each entry in a catalogue of data files would provide. the listing which follows was based on an examination of existing catalogues. the essential items, we feel are: a) file name, catalogue number or reference number (if any); b) a short description of the data. this description should especially indicate the time-span and the geographical area covered by the data; c) medium on which the data are held (e.g. magnetic tape, cards); d) size of file (number of records); e) references to available documentation and other publications pertaining to the file (e.g. a listing of reference numbers for working papers); f) the degree of accessibility (we propose: unconditional, conditional, never); g) cost of acquiring the file (if the file is accessible); h) department or agency where the raw data originates; i) the position of the person to contact for further information. i)3i$sist newsletter vol. 1, no. 3 on the last-mentioned item, we felt that the name of the person, and the telephone number, were additional items of importance. given the high degree of turnover of personnel, we felt that the position would generally be more enduring information than the name of the incumbent at a given point in time. we envision the eventual document — still planned to appear in february 1978— to be an annotated bibliography of catalogues, where the annotations are based on our proposed set of minimal information requirements. the bibliography would also state these requirements, and would, finally, contain a listing of governmental departments and agencies which, to our knowledge, have not (yet) produced a catalogue of their machine-readable data files. we are proposing to circulate our initial findings to the members of our ag before august 31, 1977. at this point each ag member will be asked to produce the annotations for a specified set of catalogues. we expect to review the eventual set of annotations at the projected lassist meeting for february 1978. with regards to our second priority: the inventory of existing guidelines in use by the originating institution or archival institution , we decided that little could be done at this stage, until we had an opportunity to examine the catalogues of data bases. in addition, tony falsetto will mail out (again) a package containing various documents related to the documentation and preservation of mrdf by the machine-readable archives division of the public archives of canada. we discussed the issue of communications between national governments and the community of researchers using process-produced data. we agreed that there is a great need for improvement and propose the following strategy: (1) comments, suggestions, queries and complaints about process-produced data would be sent to an lassist member who is at the same institution as the person requiring the information (or in a nearby institution). in short, we would propose that lassist endorse a network of "official representatives". (2) such questions, etc. would then be passed on to the appropriate ag coordinator (i.e., the u.s. coordinator if one were dealing with u.s. ppd, the canadian coordinator in the case of canadian ppd). (3) the ag coordinator would convey the question to the appropriate government agency and would be responsible for channeling the response back to the user (or, if the information were of wider relevance, to submit an item to the lassist news letter ). the ag also discussed ways in which other organizations working on related problems could be contacted to prevent any duplication of effort regarding process-produced data. we agreed that the two coordinators would get in touch with the chairman of such groups as the association of public data users, the committee on population statistics of the population association of america, and so on. 14. ^ssist newsletter vol. 1, no. 3 data organization and management canadagreg morrison, social science data archive, department of sociology, carleton university, ottawa, ontario kis 5e6 europeeric tannenbaum, social science research council survey archive, university of essex, wivenhoe park, p.s. box 23, colchester, essex, england c04 3s0 united stateswilliam ganrnell , social science data center, university of connecticut, storrs, connecticut 06268 report of the joint canadian-united states action groups on data organization and management meetings in toronto, may 11-12, 1977 (dom ag) submitted by greg morrison, carleton university members present canada greg morrison, ch, carleton university clement k. m. chan, mcmaster university janet chan, university of toronto rachel des rosiers, data clearing house for the social sciences robert logan, university of guelph paula mitchell, university of western ontario richard l. schnaar, public archives of canada, ottawa terry stewart, university of waterloo richard wolfe, ontario institute for studies in education united states bill gammell , ch, university of connecticut barbara aldrich, university of wisconsin gary m. grandon, university of connecticut pnina grindberg, columbia university sheldon laube, c. m. leinwond associates barbara noble, university of illinois sharon poss, duke university richard roistacher, university of illinois phil sidel, university of pittsburgh during the toronto sessions, members of the data organization and management group engaged in a reassessment of how various sorts of activities might best be carried out. it was felt that some kinds of projects will have to be done essentially by individuals or small sub-groups, in accordance with their particular areas of interest and expertise. draft documents, papers, proposals and the like will then be sent to the co-ordinators of the group, who will circulate them to group members for reactions and comments. after a review of possibilities, various activities were taken on by individual members. 15. sist newsletter vol. 1, no. 3 (a) the most important activity of the group during the next year will be the writing of sections of the guide to providing social science data services . in consultation with a member of the data archive development action group, eleven people agreed to draft material, and another two volunteered as sub-editors. (b) the development of a list of recommendations, addressed to potential researchers and data collectors, concerning study design as it relates to data management (the "do's and don'ts list"), first considered during the group's sessions in florida, has now been assigned to rachel des rosiers and terry stewart. (c) the proposal for an inventory of software relating to data management, also initiated in florida, has been rather drastically revised, in view of the existence of other groups outside lassist who are reviewing social science software, the need for clearer definition of what kinds of functions and programs ought to be included, and the formidable amounts of time, energy and other resources required for a comprehensive job. it seemed wiser to adopt an incremental approach and start with two clearly delineated projects: a survey of the programs used at one institution (columbia university), to be undertaken by pnina grinberg as a pilot study; and an investigation of methods and available programs for cleaning multi punched data, to be done by bill gammell, gary grandon and greg morrison. other such narrowly-focused efforts may be undertaken in the future. (d) the group felt it would be useful to gather information about the areas of technical expertise of the members of lassist, and to organize and publish the results. the aim would be to facilitate the use of the lassist membership as resource people or consultants for each other on technical problems. as a by-product, a profile of the members could be produced. the best procedure for obtaining the required information is under discussion. gary grandon and sheldon laube will be carrying out this activity. (e) another area of interest is the holding of workshops. sheldon laube will consider possibilities for organizing workshops at future lassist meetings, while bill gammell will investigate possibilities for regional workshops. (f) consideration of the inventory brought home to the group the importance of establishing communication with other organizations working on related problems. the co-ordinators. bill gammell and greg morrison, v/ill conduct these external relations. in addition to the above projects, for which individuals have taken on responsibility, two other issues were discussed. the first was the problem of information dissemination and exchange, crucial in the area of data organization and management, where relevant literature appears in a very wide range of publication from many different disciplines. we feel that this problem is particularly amenable to action by all lassist members, whether or not they are affiliated with the ag. we urge lassist members to send the co-ordinators of the data organization and management action group and/or the editor of the lassist newsletter relevant bibliographical references to any material which relates to data handling, broadly defined. these references could be to published articles or unpublished papers, to books, technical reports, computer programs, items in newsletters, and so on. the reference should include a standard bibliographical citation, information on where to obtain the item if not found in readily available sources, and preferably an annotation or abstract. this information will then appear in the book notices section of the newsletter . also welcome would be short technical notes, resumes of procedures, documents, etc.; these ought to be sent to the co-ordinators for scrutiny and possible publication in the newsletter . ji^sist newsletter vol. 1, no. 3 the second general issue discussed was that of lassist endorsements. while the group could see the value of lassist lending its support to projects in progress, it was not clear whether or not this should come as an official action and if so what process should be followed in determining what to endorse. a distinction was drawn between endorsements of good practices and endorsements of particular products; the latter possibility raises questions which might be considered by other action groups or by the steering committee. short, however, of formal endorsements, members can provide feedback for the originators of projects, and engage in lobbying activities on an individual basis on its behalf. data acquisition canadapierre lacasse, centre de recherches en amenagement regional, universite de sherbrooke, sherbrooke, quebec europemarcia taylor, social science research council survey archive, university of essex, wivenhoe park, p.o. box 23, colchester, essex, england i c04 3s0 united states[no report submitted.] data documentation canadadave l. salley, management and central services group, standards division, statistics canada, tunney's pasture, ottawa, ontario kia 0t6 europecees middendorp, steinmetzarchief , kleine-gartmanplantsoen 10, amsterdam-c. netherlands united statesjohn grasso, office of research and development, center for appalachian studies and development, west virginia university, morgantown, west virginia 26506 [no report submitted.] vol30-3.indd iassist quarterly fall 2006 by by carolyn l. geda* recollections of the formative years of iassist abstract before beginning, it is necessary to state that these are my recollections and in no way are meant to reflect recollections of other individuals involved with the formation of iassist. it is clear to me that we all have different interpretations of the past and come from different perspectives. i also apologize for the use of the first person at times, but that is necessary to assure that no one else receives blame for my interpretations and/or recollections. the flow of this presentation is not smooth, in part because so much was going on concurrently and the amount of effort that went into the establishment of iassist was too great to present coherently. it overwhelmed me then and it still overwhelms me. reliving the past is a very emotional experience for me, as is this 25th anniversary. it was certainly most unclear to me at the time that we would be able to establish iassist. as many of us know, toronto, canada is the origin of iassist. during august, 1974, the viiith world congress of the international sociological association met in toronto. at this congress, the international social science council (issc) standing committee on social science data (scssd) sponsored a conference on the problems of data archives and program library services. the conference was made possible through the very strong support of stein rokkan, president of the issc, and elina almasy, secretary of the scssd, and the endorsement of conveners erwin scheuch, chairman of scssd, michael aiken, university of wisconsin-madison, and hagen stegemann, the zentralarchiv in cologne, germany. participants of the conference reviewed common problems that confronted individuals affiliated with or utilizing data repositories. in general, these problems were related to the production, acquisition, preservation, processing, distribution and use of electronic data files of interest to the social science community. the problems were also relevant to the development and maintenance of data archives and data libraries. during this conference, many participants concluded that a communications network that provided a formal means for continued dissemination and sharing of information was needed. there were opportunities to meet each other through the interuniversity consortium for political and social research (icpsr), known at that time as the inter-university consortium for political research (icpr), annual meetings and later biennial meetings. but a broader format was needed that would permit individuals to discuss the problems they were experiencing, permit collective efforts toward resolving these problems and establish workshops that would provide basic training in data management as well as training with new technologies. toward this end, alice robbin, university of wisconsin-madison; judith rowe, princeton university; and i, icpr at university of michigan; proposed to icpr that a data management workshop be included in the summer program. when this workshop was offered, members of our peer group attended, thus demonstrating that there was indeed a need for an organization like iassist. obviously, there were additional needs for advanced workshops and seminars. a little later, the european archives addressed this need by offering expert workshops. in addition, icpr official representatives were predominately academics, while individuals staffing the archives and the user services were increasingly becoming data librarians. in order to facilitate dissemination and sharing of information, i suggested to icpr that a technical person as well as an academic be formally recognized for each icpr member institution. returning to the results of the toronto meeting, an ad hoc organizing committee (later to become the steering committee) for iassist was designated to draft a proposal for the establishment of an appropriate organization. since the organization was to be international, additional individuals were added to the committee to achieve regional representation from other parts of the world where interest in such an organization was felt to exist. ultimately, in order to satisfy the mandate of issc and assure proper representation for an international organization, the committee grew to 18 members. this was a large number of people with which to work, especially when all correspondence was to be sent to 16 iassist quarterly fall 2006 everyone on the committee via air mail. we did not have e-mail, certainly as we know it today, to assist us. many organizations were still using carbon copies as a means of distribution for documents and correspondence. although we made every effort to achieve an international organization, the expectations set forth by issc were completely unrealistic. organizations existed on paper for several of the regional areas, but we were not able to develop active regional secretariats in each of the areas designated. our attempts to comply even included reproducing additional copies of the newsletter to send to these secretariats for distribution to their interested individuals. the cost of these additional activities was borne by our home institutions as part of the commitment to iassist. eventually, we were forced to say that we could not expect an organization of individuals to continue to support such costs and the regional secretariats were scaled back to the areas that actually had members. as you might also expect, we were a very vocal group of people, each of whom obviously thought they were right about everything. nonetheless, i found myself chairing this organizing committee. it is unclear why i accepted this charge; all i recall is that i thought icpr had a mailing list that would form the base of individuals that might be interested in joining such an organization. in assuming this role, i tried very hard to prevent the feeling of american dominance, american imperialism and particularly american data imperialism. additionally, i continually struggled with a concern that icpr might be perceived as dictating to the committee or that iassist might become the puppet of icpr. the first task of the committee was to construct and agree upon an acronym. this acronym evolved at the infamous bar in toronto. once the acronym of iassist was agreed upon, we had the problem of finding appropriate words for it. although we were very pleased with the acronym, we frequently had and continue to have problems remembering the actual name of the organization! it should be noted that not all individuals felt iassist was a good idea. a well respected colleague and friend of mine thought it was a bad idea and vehemently spoke against it. several days later i received the following letter: carolyn, about the conversation in the bar—i’m really sorry that i spoke as i did. it was an unthinking reaction which did not become clear to me until later when i realized that, as opposed to most people i know these days, you really cared about the ideas and plans for iassist. real interest and concern is such a fugitive quality these days—one which i terribly admire—and thus am even more sorry to have listened and responded so negatively and so unfeelingly. in a way, though, i’m almost glad it happened since it made me realize that things still can get accomplished, if only because some people are ready to really work to achieve such accomplishments. in retrospect i should have considered this a serious warning. my friend, lorraine borman from northwestern university, recognized far better than i what was going to be involved. and, in fact, when i returned to icpr to discuss the creation of iassist, i discovered that some icpr staff thought we were forming a union! while in toronto in 1974, the committee met several times in an attempt to begin laying the foundation for iassist. some of the questions or problems raised during these discussions were: scope and type of membership objectives activities affiliation with other organizations newsletter governance participation of developing countries membership fees one of the most difficult issues was whether to establish an independent association with its own newsletter or to affiliate with an existing association whose objectives were consistent with those of iassist. included in this issue was the relationship iassist might have to the issc standing committee on social science data. the affiliation with another association was considered in anticipation of a relatively small membership. less time would have been required to activate iassist if the committee had decided to affiliate with an existing association. eligible associations were identified and considered, but the resultant lack of autonomy and independence for iassist was seen as too great a disadvantage. the committee felt the same way about a newsletter. as mentioned earlier, alice robbin, judith rowe and carolyn geda were among the united states prime movers in the founding of iassist. ultimately, we were branded as “american female functionaries” in part because we did not have phds. we were non-phd females breaking from the established issc composed of men who were professors at universities and directors of the archives. in april 1975, at the european consortium for political research meetings, the committee met again with a draft constitution and proposed a set of action groups. the constitution was not approved, but significant progress iassist quarterly fall 2006 17 was made on the mandates of the action groups. each action group was to address a problem area or set of related problems not being systematically addressed by another organization, and would look toward a final, hopefully publishable, product. any proposed action group that could not meet those criteria was deleted. two action groups were deleted: one dealing entirely with matters of archival policies and the other dealing with computer software. on the other hand, since data archive development was defined differently in europe and north america, a decision was made to create an additional action group to deal with data organization and management. emphasis was also placed on the desirability of developing action group products that were consistent with the ongoing professional activities of the membership. in this way the efforts devoted to iassist tasks by members would provide direct benefit to their home institutions. at that point in time each action group was perceived as having members from all regions in iassist. the action group structure was replaced by a regional structure because it soon became obvious that problems of international travel and communication precluded effective activity at this level. in addition, as the regional secretariats began to establish action groups, it was apparent that a regional or national focus and redefinition was required to meet the different needs, interests and levels of development. parallel action groups were recommended, in some instances with slightly different mandates. the chairpersons of the action groups were to be responsible for communicating with the chairpersons of the same or similar action group in other regions. action groups would be activated or dissolved according to the recommendations of the membership and/or action group chairpersons. action groups in existence were: data archive registry data acquisition data documentation classification data archive development process-produced data data organization and management regional secretariats were established to be responsible for the recruitment of members, receipt of membership fees, development of action groups, scheduling of regional meetings and handling of other associational activities. the regional focus of the association allowed members’ interests to be served more directly and helped to spread the burden of membership recruitment. membership meetings of regions could be more readily scheduled and better attended than could international meetings. regional secretariats were designated for the following areas: south and central america, canada, asia-west africa, west europe, east europe and the united states. australia was interested in joining pending sufficient membership. mass mailings were sent by the secretariats to approximately 1000 people– 123 in asia/africa, 71 in canada, 282 in europe, 96 in central and south america and 493 in the united states. at approximately the same time the european archives felt a need for greater recognition of their roles as national archives, particularly within icpsr. prior to this point, individual archives and universities within europe were members of icpsr. the formation of national membership in icpsr centrally organized around the national data archives permitted a stronger role by each of the archives. yet the europeans also felt a need for an organization of organizations within which they could continue to operate and develop policies. an organization of individuals was not the format in which they operated. it was imperative that iassist be an international organization, not just a north american organization, that would permit the europeans to affiliate with iassist. the issue of an organization of individuals versus an organization of organizations had to be resolved. the committee met again in edinburgh, scotland, in august 1976, where the international political science association (ipsa) was meeting. another constitution for iassist had been drafted and was presented. during this meeting, i endured a lengthy lecture on how to craft a constitution. we seemingly made no progress on reaching an agreement on anything. discouraged and drained, i left the meeting around midnight. (this had always been a very hard working group that frequently met into the night.) i arrived back at the dormitory where i was staying and called my director, jerome clubb. i said i had done my best to reach agreement but had failed. iassist was not going to be accepted. jerry gave me permission to return to the united states the next day in spite of the fact that he and i had additional meetings to attend after the ipsa meetings. he recognized that i had disintegrated and simply had no more stamina. his support was most encouraging. the next day the committee met again. somehow it had been concluded, in large part due to the extraordinary abilities of stein rokkan to work through very difficult international situations, that there could be an iassist and an organization of organizations, which became the international federation of data organizations (ifdo). the constitution was approved with the understanding that iassist was a fledgling organization and as the organization developed, the constitution would have to be amended. it was also acknowledged that the committee of 18 people might be too large to function effectively. we were all growing up professionally at this time and the growing pains placed strains in all directions. nevertheless, 18 iassist quarterly fall 2006 it was a golden era and a very special time to learn a great deal from each other and form lasting professional and personal relationships. throughout the formative years of iassist, which seemed to go on forever, many people worked untold hours. there are simply too many people to name individually, but i remain eternally grateful for their commitment and very hard efforts. one person in particular stands out, my good friend and colleague per neilsen who kept european iassist together and acted as their secretariat. others who acted as pillars of strength and are no longer with us are stein rokkan, warren miller, murray aborn, harold naugler and ed hanis, who was our first treasurer. in conclusion, i hope i have managed in some small way to convey the rich legacy of iassist. although it remains a small organization, it is one filled with deeply committed individuals who are always eager to assist others struggling in our field in any way possible. we continue to work towards the resolution of what sometimes seems to be the same set of problems we had 25 years ago, just packaged differently. iassist gives us an ongoing international format for listening to each other, of which we must continue to take advantage, and a format in which we can identify the similarities and differences of our approaches and work collectively on resolutions. * carolyn l. geda was the first president of iassist 19761979 and has served on many iassist committees since then. contact information: cg3@ix.netcom.com vol30-3.indd iassist quarterly fall 2006 by by karsten boye rasmussen repke de vries* abstract the purpose of this article is to investigate an interest in the concept of virtual communities and to consider whether iassist could be characterized as a virtual community. previously, the authors carried out an investigative description of the utilization of the iassist mailing list,, answering questions like: what are the characteristics of the top users of the mailing list? are there patterns of responses to initial submissions or to the subsequent requests mailed to the list? the work was carried out using the e-mails published on the iassist listserv over a period of time, and a descriptive report was most recently presented in iassist quarterly vol. 29-3 (rasmussen & de vries, 2006). the iassist web site is an equal part of the virtuality of iassist and was targeted for the next project. the foremost interest concerning the web site regarded how the web site could support the users’ needs in the future, with a focus on openness and the protection of privacy. this was carried out with a questionnaire from which the results are presented in this article. focus for the investigation of the web site a virtual community is a community that is to a large extent brought together and kept together by the use of electronic computerized media. quoting from an earlier article (rasmussen & de vries, 2006): “virtual organizations have been identified as real (davidow & malone, 1993) or real organizations sometimes viewed as imagined (hedberg et al., 1997). the concept of the virtual community has now existed for a good 10 years (rheingold, 1993). the virtuality emerges due to intense use of information technology corresponding to organizational arrangements that potentially and practically break the boundaries of time and space. in our time of virtuality, people no longer have to share the same space or be in the same time, as direct electronic communication can span the space, and relayed (asynchronous) electronic communication can span the time. a voluntary association – such as the iassist – is considered as both an organization and a community, and by applying electronic communications, such associations have the potential of growing into a virtual community.” the virtual community has an invested interest on the use of applications on the internet. the reason behind the investigation into the mailing list was to deliver an objective account of this aspect of the virtuality of the iassist organization. an equally objective attempt to describe the use of the iassist web site could have been accomplished by obtaining a web log for the site. these raw log data could then have been processed into more intelligible patterns of page views and their accumulation into user sessions with the intention of isolating characteristics of web behavior (kimball & merz, 2000). however, the objective analysis of the user’s actual behavior on the web site is postponed to a later time. while the actual web behavior would have sufficiently been explained by the log analysis, the investigation into the attitudes, expectations and wishes for the future implied that the request for information was to be directed towards users. the focus of the investigation was on what kind of openness and/or protection of privacy iassist should aim for. openness is a positive word and is considered to be a good characteristic, but even a good thing has its limits. the intention was to find where the audience placed these limits where the audience had gotten too much of a good thing. the questionnaire as explained, the soft issues of intention, values, attitudes, and future lead to the decision of using the questionnaire technique, and in order not to disturb the audience unnecessarily, a small, closed and heavily structured questionnaire – thus supposedly easy to fill in was posted on the web. the primary objective of the questionnaire was to obtain empirical information concerning the perception of iassist services, including the mailing list that is mentioned above. but several intentions lay behind the core of choice and design of the questionnaire. the attitudes of the membership towards existing as well as future services were to form a background for how the iassist web site should be developed. furthermore, the questionnaire explicitly addresses the ethical dimension by obtaining information concerning the disclosure of different types of open virtuality or virtually open? openness on the web as viewed by the iassist membership 20 iassist quarterly fall 2006 individual background information as well as the utilization of a questionnaire securing the “informed consent” of the participants, which often is disregarded in aggregate and more anonymous analysis; e.g. the e-mail analysis mentioned. the questionnaire was mainly addressed towards the members of the iassist organization but non-members were not excluded from answering the questionnaire. the rationale was that if the information about the iassist questionnaire actually did reach non-members and they reacted on that, then these people were probably in the periphery of iassist and therefore potential iassist members. both groups were invited to fill in the web questionnaire by an e-mail that included a link to the questionnaire. this invitation was sent to the iassist listserv as well as to some other relevant lists. methodology the degree of representativeness of the questionnaire can only partly be determined. the total membership of iassist (265 persons) was invited to fill in the questionnaire, but as the questionnaire should provide anonymity it is not possible to determine if a respondent that is stating to be a member of iassist is an actual member of iassist. one hundred and eight of the 138 respondents stated in the questionnaire to be members. acceptance of the validity of this statement resulted in an answer rate of 41 percent among the iassist membership. the answer rate is not considered exceptionally poor; however it is certainly not a random choice for the receiver of the invitation to fill out the questionnaire form. we will expect a very strong bias towards the group of iassist members that are most interested in iassist in general, and in the iassist web site in particular. the methodology presents a study with self-selection, which always is a peril for the validity of a questionnaire investigation (dillman, 2007). this construction can lead to overrepresentation as people outside the target group are answering the questionnaire and often these will fill in their e-mail address as “goofy” and “bart simpson”. many respondents (110) gave their actual e-mail address, which implies a high degree of seriousness. we have no indication that some were masquerading behind other real persons’ e-mail identities. luckily, the invitation went to communities that were not inclined to waste their time on spoiling questionnaires. however, this construction meant that we cannot calculate the overall significance of the shown figures; although the number of respondents is known, the magnitude of the population--including potential iassist members--is unknown. this methodological mishap was accepted in order to gain more insight into the preferences of the total audience. the data were collected from 13 march 2001 to 30 april 2001. announcements and reminders were sent to the iassist listserv, as well as to other lists for professionals where the iassist web questionnaire had been announced results from the questionnaire. the results presented here concentrate on background information in the form of country of origin; otherwise the tabulations relate to the experience of the web site, the need or ranking of the availability of further information, and the attitudes towards openness and privacy on the iassist web site country information in an anonymous global questionnaire, it would be unethical to reveal information about the country of a respondent, but this information can be obtained fairly accurately from logs, especially the ip address. however, as 110 respondents provided their e-mail addresses, the country information was available by a procedure similar to that carried out for the mailing list study (see the iassist 29-3 issue). it can be seen from table 1 that a large contingency of the respondents are from north america, as shown by the mapping of common and general url endings into this geographical area ��������� ������� ����� ��� ������ �� � � ����������� �� � � ������ �� � �� �� � �� � �� � �� � �� � �� �� ����� ������� �� �� �� ��� � ��� �� ��� � ��� � ��� � ���������������� �� �� ����� ��� ������������������������������������������������������ ������� iassist quarterly fall 2006 21 iassist conferences as we mentioned in the earlier article, which reported on the iassist mailing list, one of the questions we wanted to shed light on was, “is there iassist life between conferences?” it was confirmed that there is a virtual life. but the question was put in a negative manner, as if we expected there to be little life. we could also have stressed the high degree of iassist life found at the conferences and for the arrangers and presenters also in the time span up to the conference. the conference is also a recruiting platform, and one of the questions asked concerned when the respondent first participated in an iassist conference. the size of the bubbles indicates the number of first-time participants figure 1 shows that many members attended their first iassist conference several and even many years ago. it is stipulated that members, to a large degree, have continuous uninterrupted membership. however, there are also new members joining in the more recent years. a more accurate description can only be made with the use of the full member registry. in an investigation of virtuality, it seems to be strange to bother with the physical closeness that members experience at conferences, where attendees are in the same space at the same time. however, virtuality is not absolute. the interrelatedness of reality and virtuality seems to be a synergic fortification of both. thus, we postulate that it is the virtual activities between conferences that bring people to the conferences, and likewise the participation in conferences gives attendees the incentive to participate on the mailing list, to browse the web site, etc. openess to all iassist has members and non-members. if non-members receive the same information products as the members, what is the incentive to become a member? on one hand, this calls for reserving services to members. on the other hand; allowing non-members to utilize some services from iassist will act as an advertisement in addition to spreading the impact and supporting the mission of iassist. the problem for the organization is that it must decide on a reasonable differentiation of the products available for members and non-members or, phrased otherwise, to differentiate between the paying customers and the non-paying users. the product of the information age is information. in the digital age the information product is often reproduced, distributed and spread without any substantial extra cost to the organization. so one could argue for giving it away to potential customers (keen, 2001: 168). there are plenty of free offers: e-mail accounts, wikipedia encyclopedia articles, web pages, freeware, home videos, etc. however, since “there is no such thing as a free lunch,” the gratis offer is done for a reason. differentiating between the cost product and the free product often implies that there is an extra cost to the organization. the question is whether the extra cost is outweighed by the extra money raised through membership. there are examples where the premium product is the only product, and a new product with an inferior functionality is obtained only by some versioning effort from the organization (shapiro & varian, 1999: 62). what functionalities on the iassist web site should be available for members and non-members? these questions were asked of the respondents and the answers give an indication of the membership’s willingness to pay for features that can be used by all. in table 2, the different services that were investigated in the questionnaire have been ranked by the membership, as only the opinions from iassist members are included. the ranked order reflects whether the service should be available to members only or freely accessible for everyone. the members highly agree that non-members should be granted free access and searching capability of iassist figure 1. attending their first iassist conference, members only (n=108) 22 iassist quarterly fall 2006 quarterly (iq) as well as free access to search the pages of the web site (items 1 and 2). the interpretation of these findings is that the newsletter is considered an important vehicle for the spread of iassist influence and that the iassist pages should also be available as a promotional vehicle. the second item includes not only the search, and thus identification of relevant articles, but also the general availability of the iq (in full text pdf) that was readily available before this questionnaire was launched. there is a remarkable distinction between the different functionalities of free access. the evaluation is reversed from a liberal 20-80 distribution in favor of openness for items 1 and 2 to a restrictive 80-20 distribution for the items 3 to 5. one factor in explaining this dramatic shift may be the intention behind the content. while the web pages and the articles in the iq are written with the intention of being published for a wider audience, the other items – mails to the discussion list and the related archive are communication between members taking place in a more secluded fashion within virtual space and thus these contributions were not intended for non-members. openness about personal information this section concerns the disclosure of information about personal members. less than 25 percent are in favor of opening information on iassist members to non-members. the next question is: should the information about iassist members be available at all, i.e. available for the membership itself? a large majority of the membership found that this type of information should be made available (77 persons or 71.3 percent). openness is a broad concept, but how open? one way to look into this was to investigate the different attributes of the individual that were being disclosed in a membership database. the people who found that membership information should be available were then asked which attributes describing individual members should be made available: the percentage of respondents who agreed with the inclusion of specific personal information was expected to be very high. this is because only respondents who were positive towards a general description are included here. if a general description is acceptable at least some of the specific descriptive items themselves must be acceptable as well. a member’s name is the most obvious identification and the entry point for viewing other information concerning the person. in the list above (table 3), even name has not received 100 percent. this reveals some information about the precision of the measures obtained ������������������������������������������� ���������������������� ���������� ��� ���� �� �������������� �� ������������ �� �������������� �� ��������������� �� ������������������������� �� �������� �� ����������������������������������������������������������������� ����������������� table 2. opinions from the membership towards access for non-members, percentages (only members (n=108), missing data excluded, n between 101 and 106 iassist quarterly fall 2006 23 from quick answers. the person’s e-mail address being the prime address for contact is also high and above other attributes like work address and interest areas. lower down on the list is the description of the job and links to a personal web page. the reason behind this could be the instability of information due to changes in employment. lastly, the inclusion of a portrait ranges very low among these descriptive items. online communication has been analyzed for its effects on identity deception (donath, 1998) but in the mailing list, the identity of the person is well established by name and e-mail. even though identity is carried by physical appearance, and includes an online portrait, this portrait (according to members) is apparently crossing the line of privacy and adding too much “body” to the virtual community in addiction, adding an online portrait may counteracte with the traditional perception of a virtual community as being virtual only. and then came the future now six years later we can present what information is actual available on the iassist web pages that describes the membership. first of all, it should be stressed that the membership directory is only available to iassist members. secondly, the actual attributes comprise all of those mentioned above – with the exception of the picture. thirdly, the members provide their own attributes. a member not interested in revealing much information may be reluctant to fill in information. furthermore, some information can be hidden by a member because they can choose “no” in the “show contact information” field.. the feature of hiding information applies equally to the “profile” which is a text field for description of what could be both “job description” and “interest areas”. use and usefulness of the iassist mailing list the analysis of the iassist mailing list (rasmussen & de vries, 2006) is supplemented by the impression or opinion of the mailing list by its members. accordingly, the questionnaire included some questions as to the use of the mailing list. the perceived values of mailing lists are investigated in some earlier literature (hardie & neon, 1994;sproull & kiesler, 1992: 34). these values are often directly related to performance improvement via computermediated communication (rice, 1994). performance improvement was the focus for the question concerning the usefulness of the mailing list as related to the job of the respondent. among the 112 persons that responded as members of the iassist mailing list, 103 persons (92 percent) answered “agree” or “agree strongly” to the question concerning usefulness of the mailing list. even for the small group that did not find the mailing list as having any obvious, or recently experienced, direct job value (8 percent) there might be valuable aspects of the mailing list to the person, e.g. information concerning a future job area, affiliation with the area by pure interest, etc. the proof of the continued interest in the mailing list lies in the fact that these people have not unsubscribed to the mailing list even though these respondents did not identify any positive connection between the information and discussion on the mailing list and their job. the high account of usefulness of the mailing list among the membership can be taken as evidence for the mailing list being the “diamond” in the collection of functionalities available to the iassist members. to a high degree the mailing list can be considered to be the strongest incentive to becoming a member of iassist. a mailing list has also been treated as the vehicle for the analysis of virtual communities. a recent article (blanchard & markus, 2004: 67) cites and stresses the “sense of community” concept placed in a framework of four dimensions: “feelings of membership, feelings of influence, integration and fulfillment of needs, and shared emotional connection”. the figures above on the usefulness of the iassist list are an example of “fulfillment of needs”. the well distributed (democratic) activity on the mailing list shown in our earlier article can be interpreted as support for “feeling of influence”. to a certain extent the high percentage (64 percent) of respondents who answer positively to a question of “have you ever responded or reacted to information on the iassist mailing list by sending a reply directly to the sender or discussants so your comment did not show up on the iassist mailing list?” can also be taken as an emotional connection with the listserv uses – and even specific users. in the discussion of community, and especially of the virtual community, it should be noted that most membership applications take place in connection to the non-virtual conferences. the conferences are thus a strong aspect of non-virtuality supporting the sense of community in the membership. openness for others and towards self the title phrase of “open virtuality” is directed towards the aspect of opening the virtual community for people outside the membership. we were searching for the limits of openness and found them to be related to the mailing list and the membership directory. these areas were considered to be for members only. the concept of “virtuality” is also a synonym for “sort of” and including “but not quite”. “virtually open” means that the iassist offerings look as if they are open, but they are not quite all available. the “not quite” openness also extends to the opinion of the membership towards what attributes should not be included. it is “sort of” paradoxical that photos are not included, while on the other hand photos from the iassist conferences are considered to be popular and are often placed on a parallel site for amusement of the iassist conference attendees. conclusion 24 iassist quarterly fall 2006 the earlier investigation of the iassist e-mail list and this reporting of the questionnaire data have mostly been descriptive. the mailing list itself is considered a service that should be restricted to the membership. but other contributions that are directed towards a bigger audience (articles in the iassist quarterly or web pages) are considered suitable. the reason behind this consideration is the general attitude towards openness to information and research. but there might also be an aspect of promotion of the iassist organization. the bigger audience not belonging to the iassist community should be allowed to benefit from many of these services. the iassist organization and its members are evidence of a process where the boundary of the organization exemplified by the boundaries of the organizational services and of membership (and thus the boundary of the community of members) has become more blurred. the blurred boundaries concerning time and space transform the community into a virtual community. and the availability of at least parts of the offerings to non-members is also contributing to the blurredness. the joke of groucho marx that he would not join a club if it would have people like him as a member has become a fact. in the virtual community you don’t have to join the club in order to benefit from some of the services. * at the iassist conference in amsterdam in may 2001 “preliminary findings from the iassist web questionnaire 2001” were presented by karsten boye rasmussen. this was part of the collective presentation by karsten boye rasmussen and repke de vries: “professional associations in transition to virtual communities for collaboration: the case of iassist”. the finding have been discussed at some iassist administrative meetings, was latest used as powerpoint slides at the 2007 conference in montreal in the panel session “care and maintenance of a global knowledge community”, and are also to some degree reflected in the actual layout of the iassist website www.iassistdata.org that has expanded significantly in content and functionality. repke de vries, department of public services at the royal library of the netherlands. karsten boye rasmussen, department of marketing and management at university of southern denmark. contact about the articles should be directed to: kbr@sam.sdu.dk. references blanchard,a.l. & markus,l.m. (2004) the experienced “sense” of a virtual community: characteristics and processes. the data base for advances in information systems, vol. 35:1 65-79. davidow,w.h. & malone,m.s. (1993) the virtual corporation. harper business, new york, n.y. dillman,d.a. (2007) mail and internet surveys. the tailored design method (2. ed.), the 2000 version with 2007 update edn. john wiley & sons, new york. donath,j.s. (1998) identity and deception in the virtual community. in: communities in cyberspace, kollock,p. & smith,m. (eds.), routledge. hardie,e.t. & neon,v. (1994) internet: mailing lists. ptr prentice-hall, englewood cliffs, nj. hedberg,b., dahlgren,g., hansson,j. & olve,n.-g. (1997) virtual organizations and beyond: discover imaginary systems. wiley, chichester. keen,p.g.w. (2001) relationships. in: information technology and the future enterprise, dickson,g.w. & desanctis,g. (eds.), pp. 163-185. prentice hall. kimball,r. & merz,r. (2000) the data webhouse toolkit: building the web-enabled data warehouse. john wiley & sons, new york, ny. rasmussen,k.b. & de vries,r. (2006) self reflection of virtuality in a professional association: a compact description of mailing list data. iassist quarterly, 29 (fall 2005):3 14-18. rheingold,h. (1993) the virtual community. homesteading on the electronic frontier. addison-wesley, new york. rice,r. (1994) relating electronic mail use and network structure to r&d work networks and performance. journal of management information systems 11 (1) 9-29. shapiro,c. & varian,h.r. (1999) information rules : a strategic guide to the network economy. harvard business school press, boston, ma. sproull,l. & kiesler,s. (1992) connections: new ways of working in the networked organization. mit press. distributed database systems by raymondboard ' division ofresearch and statistics federal reserve board washington, dc 20551 abstract a distributed database system is a collection of logicallyrelated databases that are connected by a communications network, together with a software system for managing and accessing the data. a distributed database system is designed so that it appears to the user to be a single, unified database.this paper reviews the advantages and disadvantages of distributed database systems, and discusses issues relevant to their design and use. these issues include concurrency control, distributed query processing and transaction management, disaster recovery, reliability, and methods for distributing data. background much social science research involves the use of large datasets. these datasets are usually shared among many researchers, each of whom periodically wants to read, update, correct, and add to the data. managing such datasets presents significant problems, and maintaining them is not an easy task. frequently, specialized computer software systems, known as database management systems, are used to handle large datasets accessed by multiple users. a recent trend in computing has been to move away from the traditional mainframe environment toward a distributed computing environment. in a mainframe environment, all users share the same large, powerful computer. in a distributed environment, many computers— usually smaller machines such as personal computers (pcs) — are linked together by a communication network. such networks can connect users working on different types of machines and in different locations. this paper discusses an emerging computer technology that is intended to solve many of the problems associated with large datasets. this technology, called database management systems], takes advantage of a number of the benefits offered by distributed computing environments. as this is a relatively new technology, the reader should note that not all of the features of distributed database systems are currently available in commercial products. building such systems is a very active area of computer science research; although network database systems with many of the features herein described are already on the market, a complete disu-ibuted database management system is still a few years away. our discussion will be of a general nature; in particular, we will not discuss different data models, such as hierarchical, graph-based, and relational, currently used by database management systems. database management systems a database is an organized collection of data stored on a computer system. a database management system, abbreviated as dbms, is a software system for managing and accessing a database. these systems are typically used to manage data that is^ accessed by applications programs (e.g. packages for statistical analysis), or to answer direct user queries. it is certainly possible to store data in ordinary files in a computer's file system; in fact, for small datasets that are needed by only a single user, this is often the most practical solution. but for large datasets, and particularly those datasets that are accessed by multiple users, this practice has a number of drawbacks. dbms', on the other hand, are designed specifically to handle these sorts of situations, and to overcome the limitations presented by using ordinary files. in particular, dbms' are designed to handle large datasets that are accessed (often simultaneously) by many users, for both reading and writing. these datasets are usually of critical importance to the user, so protecting the accuracy and integrity of the data stored in them is of paramount importance. some of the crucial issues that must be dealt with by a dbms include the following: * what should be done if more than one person wants to access the data at the same time? * what happens if one person is changing data at the same time someone else is reading it? * is the data safe if the system crashes? what if someone was making changes to the data when it happened? * how can the data be accessed quickly, even when the dataset is very large? these are some of the problems that a database managelasslst quarterly ment system must address. network computing in recent years there has been a strong trend toward replacing mainframes with networks of smaller computers. the most common such arrangement is a localor wide-area network that connects a collection of personal computers or workstations'. this migration is largely the result of the more favorable price/performance ratios currently offered by pcs and workstations. the lower cost is partly due to the high servicing expenses (usually provided by the vendor) required by most mainframes; pc maintenance can frequently be handled by the customer herself. the main reason, however, is the low prices resulting from the intense competition among pc and workstation manufacturers. furthermore, the explosion in the number of these machines now in use has led to the availability of a wide range of application software to run on them. the proliferation of these smaller computers has put computing power and data closer to their end-users, often right on their desktops. another factor contributing to the spread of network computing environments is the growing popularity of the unix operating system. unix is particularly well-suited for networking, so its increased use has encouraged the implementation of network computing environments. unlike the case with mainframes, data storage and processing are distributed in a network environment. data can be stored at many different sites on the network, and computation can be performed on the machines bestsuited to the individual tasks. network file systems can make data location-transparent to users, so that they don't have to know where the data they're using actually resides. through the use of remote procedure calls, the same can be true of computation. software can be implemented so that users need not know which machine is running their applications. distributed database systems} the popularity of network computing environments has led to the development of distributed database systems. a distributed database is a collection of logically-related databases that are connected by a communications network. a distributed database management system (or distributed dbms) is a software system for managing and accessing a distributed database. a distributed database system consists ofa disuibuted database and its associated distributed dbms. the distributed database system is designed so that it appears to the user to be a single, unified database. typically, the data stored by a distributed database system is spread across a number of computers on the network. the data is disu-ibuted with the two following considerations m mind. first, data should be stored close to the machines that are most likely to run applications that use it, thus minimizing network traffic. second, the amount of data stored on the different computers should be well-balanced, to even out the data processing load and thus eliminate potential bottlenecks. a distributed dbms differs from a network file system in that its data is logically structured and is accessed by means of a high-level software interface. this interface is written so as to control concurrent access to the data and ensure data integrity. the data in a network file system is only organized into files, and can only be retrieved as files. note also the disunction between "distributed databases" and "distributed processing". the former refers to spreading databases across two or more computers, while the latter refers to spreading processing across multiple computers. distributed database systems typically perform distributed processing, but the two terms are not equivalent. the client-server model an increasingly popular network database configuration is the client-server architecture. this refers to a network computing environment in which one or more machines function as database servers. these computers store the data and manage all database operations. other machines, known as clients, run application programs. when a client application needs to read or write to the database, it sends an appropriate request to a server. the server processes the request and returns the result to the cuent, which then resumes running its application. it is important to point out that client-server database architectures are not necessarily distributed database systems. one reason is that there need not be more than one server on a network; thus the data is not necessarily distributed. another reason is that the client-server model does not require location transparency; in this model it is acceptable to require that the user know where her data resides on the network. in a distributed database system, the actual location of the data should be transparent to the user. client-server architectures are frequently used in network computing environments, and not just for managing databases. in addition to database servers, individual machines often function as file servers or mail servers for the network, handling client requests for those resources. issues the questions — described above in section database management systems— that "undistributed" dbms' must deal with are also important problems for distributed dbms'. in fact, these problems become more complex in a distributed environment. in addition, new issues arise because of the distnbuted setting. we discuss some of these issues in this section, and how they are handled by distributed dbms'. fall 1992 concurrency control recall that one of the problems that any dbms must deal with is what to do when more than user wants to access the same data at the same time. this can be especially troublesome when one (or both) of the simultaneous users wishes to change or update the data in the database. in the database world, these questions fall under the heading of (^em concurrency control) . as its name suggests, concurrency control means managing concurrent access to data by multiple users. two techniques commonly used by dbms's to provide concurrency control are data replication and locking. in one common data replication scheme*, each user wishing to access a particular set of data in the database is given her own copy of the data. (the copy is transparent to the user; to her it appears as though she's working with the database itself.) if she only wishes to read data, then she only has to refer to her own copy. but if she wishes to write to the database, all replicated copies must be updated; otherwise, users who are also working with that part of the database will no longer have an upto-date copy once she has made her changes. thus each time a user updates the database, all copies of that part of the data that are simultaneously in use must also be updated. this requires a lot of disk writes, which is a slow operation. thus system performance can suffer under this sort of scheme. in a distributed dbms, the updates to all replicated data will cause an increase in network traffic, in addition to disk operations. thus the performance degradation becomes even more of an issue in a distributed dbms. an alternative technique often used to enforce concurrency control is locking. when an application program (or a user making direct queries to the database) needs access to a particular piece of data, it requests a lock— a guarantee of temporary exclusive access— to that part of the database. if another process has already been granted a lock on that data, then the application program must wait until that lock has been released. thus at any time, at most one process has access to any piece of data in the database. locking has drawbacks also. the most obvious one is that when multiple processes need access to the same data, all but one of them must wait until they can obtain a lock. this problem can be alleviated somewhat by shrinking the granularity of the lock; that is, by making the lock apply to only a very specific piece of data. this makes conflicts less likely, but also increases system overhead, since processes will have to request locks more frequently. a similar tradeoff takes place when data replication is used. if the granularity of the replicated data is large, then the user needs to request additional data copies less frequently. however, this enhances the probability that she's working with data that is out-of-date. reducing the granularity, on the other hand, increases the number of data retrievals required, and thus the system overhead. another problem introduced by locking is the possibility oi deadlock.. suppose that two processes, p, and p^, are running simultaneously, and both need access to both of the data items d, and d^. suppose further that p, requests and is granted a lock on d,, while at about the same time pj requests and is granted a lock on d^. now neither process can proceed, since each is waiting for the other to release its lock. this situation is called deadlock, and is clearly something to be avoided. dbms' typically have subroutines that periodically check for this sort of condition; if it's found, one of the deadlocked processes is forced to relinquish its lock, so that the other process can proceed. deadlock detection is more difficult in a distributed dbms, since the locks that the processes are competing for may be at two different sites on the network. since each site typically handles the locks on its own data, it's possible that neither of the two sites realizes that it's waiting for a lock to be released on the other machine. this stalemate situation is known as global deadlock, and is much harder to detect, since there is no central program managing all of the locks. distributed dbms' often detect it by "timeout": once a certain period of time has elapsed without the locks being granted, a global deadlock condition is assumed to exist. concurtency control in distributed dbms' is further complicated by the possibility of communication or site failure. a message relinquishing a lock may be garbled or lost, or a computer may crash without releasing its locks. either of these situations results in dangling locks, which are no longer needed but have not been released. distributed dbms' must have contingency plans for detecting and dealing with these situations. as an example of concurrency conu"ol, consider an airline reservation system. this type of system is centered on a large database that can be accessed by thousands of ticket agents around the world. clearly some sort of concurrency control is needed to prevent two agents from simultaneously booking the same seat. when one agent tries to reserve a seat, she is granted a lock on that seat. the lock granularity should allow the agent to simultaneously lock two adjacent seals for a couple traveling together, but also allow other agents to book other seats at the same time. consistency maintaining database consistency is another important issue. consider a bank, and two customers (,4 and b) who have accounts there. suppose that a writes a check for si 00 to b, who deposits it in her account. recording this lassist quarterly transaction in the database thus requires two operations: the balance in a 's account must be decreased by si 00 and the balance in b"s account must be increased by the same amount. but what happens if an accounting program reads the balances after a 's account has been adjusted but not b's? the accounting program will be told that there is sloo less in the bank than is actually the case. the database will be in an inconsistent state. to prevent this sort of inconsistency, dbms' allow users to group database operations into transactions. these are sequences of database write operations that are treated as a single unit; either all of the operations in the transaction will be carried out (or committed), or else none of them will be. furthermore, no other changes will be made to the database between the time the first operation in the transaction is executed and the time that last operation in the transaction is executed. in the banking example, the dbms would guarantee that the accounting program could not gain access to the account information until after the adjustments to both of the accounts were made. the accounting program would thus see a consistent version of the database. in a distributed dbms it is important that, in transactions that affect data at multiple sites, either all of the sites are updated or else none of them are. this is usually ensured through the use of a two-phase commit protocol. in the first phase, the machine initiating the transaction sends a message to all of the sites that will be affected by the transaction, asking them to verify that they are prepared to commit the transaction. if each of these sites responds positively, then the initial machine sends a second message instructing the other sites to actually commit the transaction. if not all of the sites respond positively, perhaps because of a communication failure, then the initial machine sends out a second message cancelling the transaction at all of the sites. this ensures that each site maintains a consistent version of the database. this type of protocol requires a significant amount of system overhead. disaster recovery a very important issue is how to protect the integrity of the data during a system crash. normally, if the system goes down in between transactions, there is no major problem. this is because at the end of a transaction the database will be in a consistent state (assuming that transactions are managed as described above). of greater concern is the possibility that the system could crash in the middle of a transaction. in this case, some of the changes made inside the transaction may have been executed (i.e. written to disk), but not others. thus at the time of the crash the database could be in an inconsistent state. in the example above, if the system failure occurred after a 's account was debited but before b's account was credited, then the database would understate the bank's deposits by $100. dbms' frequently address this problem, known as disaster recovery, by keeping track of all transactions in a log file. then, in the event of a system crash, the log file is read automatically by the dbms (once the system is funcuonal again) to see if any transactions were in progress at the time of the failure. if there was such a transaction, all operations in that transacuon that were executed prior to the crash are "undone", so that the database is returned to the (consistent) state it was in at the time the transaction started. in a distributed database system each site typically maintains its own log file. another common technique for guarding against system failure is to make new copies of the disk pages where the data to be updated resides. the transaction of)erations are performed on the copy. when the transaction is completed, the updated disk page can be remapped to replace the old data in the database in a single, atomic operation. query optimization a primary goal of all dbms' is to provide fast access to information in the database, even when the database is very large. the speed with which a dbms can answer a query from a user or an application program relies to a considerable extent on the order in which it carries out the database operations necessary to extract the requested information. thus to ensure efficient performance, a dbms must be able to opumize the execuuon of queries. as a (rather simplistic) example of the importance of ordering the operations wisely, consider the "irish presidents" problem. suppose there is a database containing information on many thousands of current and past us citizens, and a request is made for a list of all people in the database who were both of irish ancestry and american presidents. one way to satisfy the request would be to first retrieve all people in the database of irish ancestry, and then select from this (very large) list those people who were also presidents. a much more efficient way to generate the list would be to initially retrieve all us presidents, and then select from this much shorter list those who are also of irish descent. many dbms' allow precompilation of queries. that is, if a certain type of query will be executed many times, the user can compile it into an optimized form that can then be used for all subsequent invocations; optimization is thus performed only once. for ad hoc queries that will only be executed once, the dbms must perform the optimization when the query is actually made. since optimization can be a time-consuming operation, a tradeoff can arise between the time needed to perform the optimization and the time saved by executing an optimized version of the query. in these situations it may be best to perform less extensive (and thus less timefall 1992 consuming) optimization. query optimization is especially important for distributed databases, particularly since individual queries may involve data stored at more than one site. communications overhead is a major concern in distributed environments, due to the relatively slow speed of network communication relative to most other operations. thus it is important to optimize queries so as to minimize the amount of network traffic, in terms of both the frequency of communication and the size of the messages passed. global optimization, which takes into account communication limes between sites, is needed for maximum efficiency, rather than just local optimization at each of the database sites. note that the more autonomous the individual sites are, the harder this will be to do, since effective global optimization requires that much information on the distributed data be available in a central location. data distribution a key aspect of the definition of a distributed database system is that the data is disffibuted among multiple sites on the network. the way in which the data is distributed can dramatically affect system performance. the question of how data is to be distributed can be divided into two pans,fragmentation and allocation. fragmentation refers to how the data is broken into pieces. once this has been determined, the data fragments must be allocated to various sites on the network. when large networks and databases are involved, finding an optimal (or near-optimal) fragmentation and allocation can be a very difficult problem. while it's important to put data close to users, it's also crucial to distribute the data so as to reduce netwoiic traffic, and to balance the distribution in order minimize the chances of bottlenecks appearing. thus the frequency of access to the data fragments must be considered. also, the structure of anticipated queries should be taken into account, with the goal of reducing the number of queries requiring data from multiple sites. a sound distribution design strategy should take into account all of these issues. in order to improve performance, the system designer may wish to store multiple copies of some data fragments, particularly the most heavily-used data. this can result in frequently-accessed data being stored near many or all of its users, as well as reducing contention for individual copies of this data. the cost of such a strategy is that this may incur substantial system overhead; updates must be made to all copies of the replicated data, resulting in more network traffic and additional concurrency concerns. the replicated data fragments will also require more disk space. another important feature of a distributed database system is that the data should be location-transparent. that is, no matter how the data is fragmented and allocated, the user shouldn't have to know how or where it is stored. the interface to the distributed dbms should be such that a user, or an application program that interacts with the database, views the distributed dbms as a single logical database. note that location-transparency implies that when the data in the database is moved around or restructured, existing application programs won't have to be altered to adjust for the changes. heterogeneous networks a practical consideration of some importance is that distributed dbms' should be able to work on heterogeneous networks. here "heterogeneous" refers not only to computer hardware, but also to the different types of network hardware, operating systems, communication protocols, and even database management systems that are commonly encountered. the last of these is critical, since much of the costeffectiveness and usefulness of a distributed dbms may result from its ability to link together a number of existing dbms'. a distributed dbms that is able to run on a wide variety of systems enables widespread sharing of data among databases in environments like universities, where different types of computer systems abound. another advantage to heterogeneity is that if a distributed dbms runs on many types of systems, then it's easier to add existing databases to it. this way system administrators can protect their investments in existing systems by being able to integrate them into a larger system, rather than having to replace them. hardware and protocol heterogeneity can be achieved through the use of low-level communication interfaces called gateways. once these interfaces have been established, there can still be communications problems if the distributed database system links together different types of dbms'. thus it may be necessary to translate between the two (or more) different dbms' query languages. the software programs that translate dbms requests into alternative query languages and send them to the appropriate sites are, confusingly, also known as gateways. note that in a distributed database system with many different types of machines, protocols, and dbms', the number of gateways (of both types) required can be very large. this problem could be ameliorated by the adoption of industrj'-wide standards for such things as data models, query languages, and concurrency protocols. such comprehensive standards, however, seem unlikely to be established in the near future. advantages and disadvantages as might be expected, there are both advantages and disadvantages to disunbuted database systems. perhaps the most obvious advantage is that such systems facilitate lassist quarterly the sharing of data among large communities of users— for example, among the faculties of different departments in a university— using existing, possibly heterogeneous, computer networks. thus more users can have access to more data, without having to know where or how the data is actually stored. user interfaces one advantage of a distributed dbms that can be very apparent to end users is the superior variety of user interfaces available on pcs and workstations as compared with most mainframes. graphical user interfaces (guis), allowing multiple windows and (often) bit-mapped displays, are commonly available on these smaller machines, and greatly enhance the enjoyment and productivity of the user. in a distributed database system, the user can work with a gui to access data stored elsewhere on the network without having to learn and use the less user-friendly style of command interface that still exists on most mainframes. performance a related advantage is that computationintensive applications can be moved off of the database server machine(s) and onto the users' individual pcs and workstations. this takes some of the processing load off of the servers, thus allowing all users faster access to data. in a mainframe environment, user applications compete with the dbms software for the computer's cpu. but by processing the data locally in a distributed environment, greater processing capacity is achieved by keeping many machines busy. another way that performance gains can be realized in a carefully designed distributed database system is by moving data closer to the people who use it by distributing data on the pcs or workstations of those most likely to use it, not only do those users benefit from faster data access, but other users benefit as well, due to the resultant lightening of the load on the other database servers on the network. note that, in addition to offering additional functionality such as concurrency control and data consistency, a distributed dbms may also achieve better performance than network file systems in certain applications. this is because distributed dbms' respond to queries, and thus need only send over the network the data that satisfies the specific query. in a network file system, however, only files can be transferred across the network. thus a much larger amount of data than is actually needed by the requester is likely to be sent, resulting in increased communication time. incremental growth distributed database systems also facilitate the incremental growth of databases. new machines, f)erhaps with new datasets mounted on their file systems, can be incorporated one by one into a distributed dbms. thus as the data to be stored outgrows the existing systems, new machines can be added to expand capacity. in a mainframe environment this type of incremental growth is generally not possible; the entire dbms would have to be replaced with a new system, a much more expensive solution. a related advantage of distributed dbms' is that they are well-suited for handling (nem legacy) systems. a legacy is a software program that, although today it might not be the best choice for its job, is so furoly entrenched in the user community (because of years of use, hundreds of applications that invoke it, etc.) that it would not be feasible to replace it. a legacy database system can be incorporated into a distributed dbms by making it one node in the system. applications requiring data from that dbms could still use it, while other programs and users could use data stored on other machines on the network. reliability robustness and reliability are other areas in which distributed dbms' display advantages. in most cases, a failure at one or more sites on the network will not crash the entire distributed dbms. in a well-designed disuibuted dbms, not only the data but also the control over query processing, concurrency, and disaster recovery is distributed. thus failure at one point will not render the entire database system useless. although some data may be temporarily unreachable in the event of such a failure, much of the data should still be accessible. in addition, if the system is designed with careful replication of data on multiple sites, all users may be able to continue working with no ill effects if part of the system goes down. disadvantages the most telling drawback of distributed dbms' is probably the increased complexity of administering and maintaining the database. instead of just managing a single database, the database manager must now also contend with the network, communications software, data that is replicated on multiple machines, and the backup and recovery of distributed data. there are more possible points of failure, including the machine requesting data, the network, and the machine hosting the data. testing new applications is harder, since there may be many different combinations of client and server machines that users will want to run the applications on. software updates are more of a chore; instead of installing an update on a single machine, database administrators must ensure that all machines in the distributed database system are running up-to-date software. finally, maintaining data security will be more difficult. there is an inherent conflict between granting wider data access to enable many machines on the network to use the database, and restricting access to sensitive data. in environments where sensitive data is stored, a careful balance must be struck between these competing interests. another potential disadvantage is that poody implefall 1992 menled distributed dbms' may exhibit worse performance than their centralized counterparts. a system with a very high rate of transactions, and data that is not distributed efficiently, could result in very heavy network traffic. this traffic, combined with the overhead of the software managing the distributed dbms' network communications, could cause communication delays that overcome the expected performance gains described above, and result in unsatisfactory performance. thus careful thought must be given as to how data is distributed and replicated among the machines on the network. concluding remarks although fully-functional distributed database systems, as described in this paper, are not yet a commercial reahty, they appear to be a promising means of handling large shared datasets in the near future. systems with many of the capabilities discussed are now available, and distributed client-server databases have been installed at a number of sites. distributed dbms' take advantage of many of the features that have made network computing environments increasingly popular. they distribute the processing load, move the data closer to the people who work with it, and allow cheap, incremental system evolution. one unknown aspect of distributed dbms' is how well the algorithms and protocols they use, such as two-phase commit, will scale up as networks grow larger and connect more and more computers. further research is also needed on data distribution strategies and on improving transaction management and query processing in a distributed environment. in spite of these obstacles, however, the future of distributed database systems seems bright, aided by the continuing growth in popularity of network computing environments. references r. dale. client-server database: architecture of the future. database programming and design, pages 28-37, august 1990. h. a. edelstein. lions, tigers, and downsizing. database programming and design, pages 39-45, march 1992. b. gold-bernstein. does client-server equal distributed database? database programming and design, pages 52-62, september 1990. m. krasowski. why choose a distributed database? database programming and design, pages 46-53, march 1991. d. mcgoveran and c. j. white. clarifying client-server. dbms, pages 78-90, november 1990. m. t. \"{0}zsu and p. valdurie. distributed database systems: where are we now? computer, 24(8): 68-78. august 1991. s. ram. heterogeneous disu-ibuted database systems. compiuer, 24(12): 7-10, decembet 1991. a. silberschatz and m. stonebraker and j. ullman. database systems: achievements and opportunities. communications of the acm . 34(10): 110-120, october 1991. 1. presented at the lassist 92 conference held in madison, wisconsin, u.s.a. may 26 29, 1992. 2. in the computer science literature, "data" is invariably treated as singular, rather than plural. this convention will be followed in this paper.) 3. workstation" here refers to a desktop computer, usually intended for single user operation, that features a faster processor, more memory, and a larger screen than a pc. most workstations are unix-based and offer elaborate graphical user interfaces 4. this is known as the "read one, write all", or "rowa", protocol. 10 lassist quarterly 40 iassist quarterly the new oed project at waterloo: old wine in new bottles by d. w. russell' waterloo centre for the new oed university of waterloo 'presented at the international association for social science information service and technology (iassist) conference held in vancouver, british columbia, canada on mav 19-22. 1987 in mid-1983, three years before the fourth and final volume of the supplement to the oxford english dictionary (oed) appeared in conventional print form, oxford university press (oup) had announced its intention of computerizing the oed. by the end of 1983 the decision had been made that oxford university press would manage this large imdertaking itself, and would enter into a series of contracts/agreements with other parties (such as the university of waterioo, ibm uk, international computaprint corporation (icc) of fort washington. pennsylvania, among others) to carr> out work on various aspects of the computerization. from a long-term perspective, the project's objective is to transform both the original 12 volumes of the oed and the 4-volume supplement into an electronic database; the work to be done in arriving at this objective has been broken down into several discrete phases: first, the initial sixteen volumes had to be entered into the computer, preserving the original text organization and presentation. this meant the tagging of structural and typographical elements in the dictionary as the data were keyed and transfered onto magnetic tape. one of the aims of this phase is to produce a printed version of the dictionary, integrating the supplement material with the body of the oed, and incorporating about 4000 new entries into this merged edition. this version will be printed in the spring of 1989. in a projected 22 volimies at a price of about £1,500. the next major phase involves the design of a database structure for the machine-readable data, so that alternative structures can be presented, and interactive querying will be possible. it is at this point that the new oed comes into being, leading to expanded, updated, and revised versions of the dictionary thai will allow several modes of data access, ranging from direct, online access, to the conventional printed version, and special, printed subsets of the dictionar,'. as can be seen, the possibilities are many, and the prospects for fall/winter j 987 [assist quarterly 41 lexicographical work seem to expand almost infinitely. the magnitude of the proposed task can be appreciated when one considers the size of the current source data, the sixteen volimies of oed and supplement: there are over 21,000 pages of three-coltimn print, with a total of about 500 million characters including pimcttiation and spacing. this breaks down to about 306,000 main entries, 163,000 subordinate entries, with over 2,350,000 illustrative quotations, and almost half a million cross references. in addition to the sheer size of the dictionar>', the material has, as anyone who has ever used the oed knows, an extremely complex structure, by which explicit and implicit information is conveyed to the reader through both the layout and the typography of the texl at the simplest level, each entry usually includes a headword lemma, a pronunciation key, a grammatical category label, a list of variant forms given by century, an etymological section, a series of sense definitions, each with a quotation bank; there is usually one quotation per century, with date, author, source and bibliographic reference. this structure is made more complex by words or forms that do not fit easily into the usual categories, and by the inconsistencies in method inevitable in a project that spanned many years and several generations of lexicographers. still, the capture of the dictionary' material, which had to be done by manual keying of the text, since it was too complex for optical scaimers, has been done, and done surprisingly successfully over a period of 18 months, ending in june 1986. the enoi rate, as noted by human proof readers hired by oxford, is seven or fewer errors per 10,000 keystrokes. working from enlarged copies of oed text, the key entry operators entered tags to identify all the typographical elements, as well as some of the structural elements of the source text one of the chief aims at this stage, was, after all, to be able to produce a new, typeset edition in 1989. but, since even this seemingly routine task becomes more complex when one is dealing with as much material as is foimd in the oed and supplement, a further refinement was introduced at this point in the process. computer science researchers at "waterloo developed a parser which automatically tagged structtiral elements not tagged dtiring keying, and which converted the icc tags into sgml codes.' this parsing permitted the automatic validation of the earlier tagging. the conversion to sgml codes was needed to permit a greater degree of automatic integration by computer of the oed and supplement, as well as to facilitate the lexicographical team's interactive integration of the two source texts. meanwhile, at the university of waterloo, research on database design for the new oed has begun, with the suppon of a 1.3 miluon dollar gran: from nserc. the project has the mandate to design a database for the new oed, and to develop software utilities for database access, maintenance and update. this work is to be carried out over three years by a team of seven full-time researchers, directed by a team of computer science faculty researchers, led by gaston gonnet and frank tompa. to date, two prototype software tools have been produced, to allow interactive access to, and sophisticated querying of, the database. these tools are named goedel and pat. although it is not within my competence to describe any of the technical specifics of these tools, i would like to offer some examples of the results currently possible, drawing mainly from my own research interests in the oed, namely the identification and smdy of the anglo-norman elements which were adopted by the enghsh language in the medieval period. 'see rick kazman, structuring the text of the oxford english dictionary through firute state transduction, (m. math, thesis) umversitv of waterloo, 1985. fall/winter 1987 42 [assist quarterly when the first tapes of the oed became available at waterloo in 1985, it was theoretically possible to begin searching the dictionary interactively. in order to find £dl words labelled as anglo-french or law french, i needed simply to ask the computer to list occurrences of the relevant strings, "af"', "onf", "law fr.", etc. at the time the data were moimted in a series of files, which meant the queries had to be repeated for each file. the results were not easy to read, and although one could scroll back or forward to identify the headword, in the end it proved easier to use the printed dictionary to locate the relevant entry. after the creation of goedel in 1986, a whole new range of possibilities made my querying of the data easier. i could now ask for the extraction of material according to structural categories in the dictionary entries, and according to specific strings within each category. 1 decided to limit my extraction to listing the headword lemma, the material within the etymological section, and the dates of the illustrative quotations, for all entries which could be considered to be derived from anglo-norman sources. within the structure of the oed, anglo-norman material is identified in various fashions in the etymology section: it may be labelled as anglo-french, it may be labelled as old french (or old norman french or old law french), or it may be found to be labelled as french but with quotation dales preceding 1500. the results of my queries using goedel could be printed in a formatted form which is eminently readable, and could be carried away for use elsewhere, freeing the researcher from being lied to a computer terminal. with the extraction power and fiexibility of goedel, the humanist researcher is faced with the new challenge of creating more sophisticated queries, based on previously unexamined possibilities, and building on the results of his or her ongoing interactive research. the second extraction tool, pat, allows a rapid and effective querying of a difterent sort, based in part on pattern matching; with pat, for example, 1 could extract all entries derived from af or of and which have supporting quotations from a particular author, such as chaucer or gower. or, to give another example, a researcher looking for infantine language was able to search the dictionary for all words whose sense definition included specific key words, such as "little" followed by "boy" or "giri" or "child" within a specified number of characters. increasing familiarity with the results of these searches led to more sophisticated querying of the database, and raised technical problems which were addressed by the data structuring group working on the project in a similar way, pat was used by a researcher interested in the source and frequency of quotations. it is possible to extract all quotations from a particular author, such as fanny bumey, and further, to extract only quotations from bumey from a specified work, such as cecilia. these short examples, from among the preliminary group of research projects underway at the waterioo centre, serve to emphasize the futuristic nature of the project; it is difficult to design a database for uses which may arise in the future, but which have not yet been imagined by humanist researchers. in an attempt to come to grips with this problem, oup and the university of waterloo conducted a user survey, to find out how individuals use the oed, to determine the principal facilities needed for the new oed, and to provoke considered responses about applications for the electronic version of the dictionary. over 1,100 individuals were contacted, of whom 60% were from the uk and europe, 40% from north america. the sample included both academic and non-academic users, with an emphasis on sophisticated users. the results have not yet been completely analyzed, but preliminary results give a broad outline of what users would like to see in the new oed. basically, most fall/winter 1987 iassist quarterly 43 users want everything currently in the oed, in a fashion that is simple to use, quick to give results, and cheap to access. ideally the data will be available publicly through an online data service, and privately via some disk formal the software must be user friendly, allowing quick access, both to a skeleton summary of each entry, and to a complete entry or selected details of an entry. the electronic version must not be prohibitively expensive, and yet be amenable to continuous updating and revision. and finally, the new oed should continue to be published in printed form. the survey results also suggest future applications, many of which will require revision of the oed: these include semantic field searches, frequency ratings by date or language of origin, the use of the oed as a thesaurus, and so on. one of the suggested applications, requiring the abilit>' to search phonetics elements, will be made much easier by the decision to replace murray's phonetic transcriptions with transcriptions based on the symbols used by the international phonetic association. this revision is to be incorporated in the 1989 print version of the new oed. searching on these phonetic elements will only be possible, of course, in the electronic version of the new oed. just what the electronic form of the new oed will be has not yet been decided. there are plans to market an exploratory electronic issue of the old oed. without the supplement, and without revisions, on cd-rom in late 1987, with the aim of validating some of the results of the user survey, and generating feed-back for the database design currendy in progress for the new oed. the database created to produce the print version of the new oed in 1989 is not now in a form which oup would offer to outside users, but the project does expect to market an electronic version of the 1989 and later editions of the new oed. the undertaking does present a number of technical and legal problems, the solutions to which will determine the final end product and the planned revision and enhancement of the new oed after 1989 will certainly carry the project well into the twenty-first century. judging from the unforeseen shifts and changes in plan which plagued james murray and the other editors involved in the creation of the oed from 1879 to 1933, it is perhaps imwise to promise completion of the new oed project by any given date. but there is room for cautious optimism. such is the position taken by tim benbow, oup's director of the new oed project, in his status report given at our second annual conference at waterloo in november 1986. he said:"at the moment the project is nmning on schedule and on budget we are not, however, complacent rectirrent nightmares — one, in which as a juggler one is constrained to keep an increasing number of dictionarx' voltmies of monstrous size in the air imder pain of lexicution, alternating with a sisyphean vision of ordering acres of dictionary slips only to have them taken by the wind as the last is about to be positioned — see to that!" it is, no doubt, significant that tim benbow's nightmares do not yet reflect the presence of any textual database monsters.n •tim benbow, "status report on the new oed project" paper given at 'advances in lexicology", second annual conference of the uw centre for the new oxford english dictionary, waterloo, nov. 9-11, 1986. fall/winter j 987 12 iassist quarterly 2015 iassist quarterly iassist quarterlyiassist quarterly abstract one of the roles the ddi standard can perform is to serve as a medium for the transfer of metadata and data across both space and time. a crucial component of this role is the ability to represent the data and metadata contained in common data analysis and management packages. this paper describes an experiment using the program stat/transfer to move datasets among five popular packages with ddi lifecycle as an intermediary format. we created a dataset in each of the five packages and then exported it to ddi lifecycle via stat/transfer. we also created a ddi lifecycle instance and an associated delimited dataset, containing as many of the metadata elements found in any of the five packages possible and then exported it to each of the packages. success or failure to transfer was recorded for a set of generic metadata elements identified in an earlier paper. using a commercial file transfer program helped identify which metadata elements were transferrable through a generally available machine actionable process. the experiment revealed some areas for potential improvements to ddi as well as suggestions for data analysis packages and research practices. keywords: ddi, data formats, metadata, statistical packages, stat/transfer, jmp, r, sas, spss, stata. introduction not uncommonly, datasets are initially produced in one software package specific format, whether proprietary, as in an spss dataset, or open source, as in an r workspace. as the data are reused, either at a later time or by other researchers in another place they will commonly need to be imported into a different software package. even within one organization a variety of software tools may be employed. researchers may have differing needs and personal preferences. different packages may have unique analysis tools. software and preferences for software also evolve over time. data retrieved from an archive after a period of dormancy may very well need to be represented in some new format.. an earlier paper (hoyle et al., 2010) enumerated a list of generic metadata elements that could be represented in at least one of a set of eight common data file formats. no one format was able to represent all of the metadata elements. ddi lifecycle came closest to being able to represent all of the metadata elements. in what follows, ”ddi lifecycle” refers to ddi version 3.1, the version used for the experiment described ddi as a common format for export and import for statistical packages by larry hoyle and joachim wackerow1 the role of ddi envisioned here is as an intermediate format in the process of moving data and metadata among software packages iassist quarterly 2015 13 iassist quarterly here. where appropriate, we will describe additional capabilities of ddi 3.2, published in 2014. some metadata elements were only representable in ddi 3.1 with a workaround (as r:note elements) making automated discovery challenging. a worthwhile goal for ddi is to be able to contain a machine actionable superset of the metadata elements across the most commonly used array of analysis software. for the “transport” role envisioned here, it is important for ddi to be optimized for clarity and completeness, not necessarily for speed or efficiency.. the packages this study used five of the formats which were able to contain the broadest array of metadata. each of the five packages is able to store metadata some elements not shared by all of the other packages. • jmp versions 8 and 10 (sas institute jmp) • r version 2.14 and 3.01 (r development core team, 2009) • sas version 9.2 and 9.4 (sas institute sas) • spss version 19 and 21 (ibm) • stata version 11 and 13 (statacorp) the role of ddi envisioned here is as an intermediate format in the process of moving data and metadata among software packages. figure 1 shows that role in moving data and metadata among the 5 formats investigated in this project. given its open nature and to the extent that it comes closest to handling a superset of the metadata managed by the analysis program formats, ddi has an advantage as the intermediate format. ddi in this role is not necessarily restricted to expression as an xml instance. while the formal specifications of ddi lifecycle 3.1 and 3.2 are xml schemas, ddi content can be stored physically as an xml file, in an xml database, or even in a relational database. for a discussion of the latter see (amin et al., 2011). the ddi alliance is also currently working on two rdf vocabularies (data documentation initiative. 2013a). future plans for ddi call for a model based specification which can be expressed both in xml and owl / rdf (data documentation initiative. 2013c). stat/transfer versions 11 and 12 were used for the experiment described below. transferring data there are several ways to transfer data and metadata to and from ddi and the five packages. while not practical for datasets of any size, ddi can be hand-edited with an xml editor. there are a number of tools listed on the ddi tools page (data documentation initiative, 2013b) which can convert metadata to and from ddi and at least one other format. as of version 11, stat/transfer can move data and metadata between ddi and 35 other formats. this breadth of coverage was a factor in choosing stat/transfer for this experiment. the intent here was neither to endorse nor critique stat/transfer, but rather to get a better sense of the current state of the ability to use ddi as a medium for automatic translation of data and metadata across software packages. since stat/transfer (circle systems) passes the information through its own internal model, in a sense it is also a sixth format as well as the transfer tool. the experiment our experiment was designed to address three questions. what metadata are currently possible to transfer with an automated procedure? what metadata elements does ddi support that stat/ transfer does not? what more could ddi support? stat/transfer relies exclusively on the g:resourcepackage element to contain the metadata in the ddi:ddiinstance it produces. no s:studyunit is produced. this usage is consistent with the approach taken by colectica when importing from spss and stata, the philosophy being that there are no specifically study-level metadata contained in the dataset. this does reveal a shortcoming in the native formats of all of the packages. these files all contain metadata about the structure of the file but almost no metadata figure 1 – ddi as an intermediate format figure 2 from packages to ddi validated with colectica figure 3 from ddi to packages 14 iassist quarterly 2015 iassist quarterly about the actual data or study. custom attributes on the dataset might offer a mechanism to remedy this shortcoming. a compatible ddi file used as a source for transfer into the other 5 packages was hand-entered into an xml editor. a separate comma separated variable file contained the associated data. the xml editor validated the metadata against the ddi xml schema. secondary level validation on the pair of files was performed using colectica reader version 3, and colectica express version 4. at this time neither stat/transfer nor colectica were capable of using embedded data in the ddi file. datasets like the one shown in figures 4 and 5 were created in r, spss, stata, sas and jmp. each of these datasets contained instances of all of the generic metadata elements they could represent. figure 6, for example, shows the addition of custom attributes named “concept”, “note”, and “universe” to the sample spss dataset, as well as metadata for level of measurement (nominal, ordinal, and scale), and role(input and target). the hand-edited ddi file was transformed using stat/transfer into each of the package’s native format. datasets created in each of the packages were also transformed into ddi. the resulting files were then reviewed for each of the metadata elements considered in the earlier study. grids like the one shown in figure 7 were filled in, with a “+” indicating successful metadata transfer, a “-” indicating unsuccessful transfer, and a “~” indicating partial success – such as metadata transferring to an unexpected ddi element. hatched shaded cells indicate metadata elements not supported by that software package. summary of results data, and basic metadata, transferred to and from ddi from all 5 packages. for the most part the elements representable in all of the packages transferred to and from ddi correctly. these elements include: • dataset name – for all of the packages this came from the name of the file in the host operating system. an r workspace file can contain multiple data frames and other objects. • variable names • variable labels – in jmp this becomes a note on the variable • variable order • important variable data type (such as date and datetime) – see below for issues related to time zone. • a missing indicator see below for issues related to multiple distinct missing values • labels for numeric values a few metadata elements common to most of the packages never transferred correctly: • display formats – this is really no surprise given that there are no standards for display formats across the packages. figure 4 the spss dataset with value labels hidden figure 5 the spss dataset showing value labels figure 6 variable properties from the spss dataset, including user defined properties concept, note, and universe iassist quarterly 2015 15 iassist quarterly • notes on datasets or variables– most of the packages have some facility for general notes on a dataset or a variable. these were not successfully transferred. • user defined attributes on datasets or variables – several of the packages allow for user defined attributes on the dataset or on individual variables. these did not transfer successfully. • measurement level – several packages allow for the specification of measurement level for variables. the vocabulary for measurement level varies across packages though. • weight – spss and jmp can store an attribute indicating that a variable functions as a weight. this never transferred. results in more detail dataset level successful • dataset name (except r) • date modified issues in most cases dataset labels, dataset date modified, and value labels for numeric variables transferred. value labels do not transfer to r, but this is not unexpected, since factors in r are somewhat conceptually different than a labeled variable in the other packages. for the 5 packages studied, dataset name is typically not included in the dataset file itself, but rather is contained in the name of the file (at the operating system level). if the dataset name is taken from the file name this is a potential problem for r where multiple data frames may be contained in the workspace file. a ddi file should contain that name in l:logicalproduct/ l:logicalproductname. additionally, the filename should be captured in pi:physicalinstance/pi:datafileidentification. metadata elements which are not common across the packages did not transfer well, even when possible. these include user defined characteristics of the dataset, notes (which are sometimes a user defined characteristic) and scripts embedded in the dataset (supported only by jmp in this collection). rule based integrity constraints also did not transfer. we did not include foreign key constraints or passwords on the sample datasets. variable level successful • variable name • basic data type • position • variable label • system missing values • value labels – numeric variables • value labels – text variables (where possible) issues at the variable level, display formats, such as a leading euro symbol, did not transfer at all. given that display formats are not standardized across packages, this is not surprising. other elements which did not transfer are: number of decimal positions, scale (not supported in most packages), measurement units, measurement level, variable as a weight, role, user-defined variable attributes, and notes on variables. some of these, such as weight, and measurement units are really essential for interpreting the data. others carry meaning beyond cosmetics. number of decimal places, for example, can connote the level of precision of measurement. missing values several packages have the ability to use multiple distinct values to indicate different categories of missing data. stata and sas both have a set of “out of band” values to represent distinct missing values. these are displayed as “.a” through “.z” and “._”. spss, instead, allows “in band” values to be chosen as missing values, a “9” or a “999”, for example. spss also allows one range of values to be declared missing, e.g. all values between 90 and 99 inclusive. r has only one missing value “na”, although figure 7 – example transfer results 16 iassist quarterly 2015 iassist quarterly it also has indicators for infinite numbers “inf” e.g. the result of 1/0 and “not a number” (“nan”) such as the result of 0/0 . transferring variables with multiple distinct missing values among packages is not straightforward. ddi3.2 added the managedmissingvaluesrepresentation element which allows adding a code scheme for missing to a numeric representation. this is consistent with the notion of “sentinel values” in iso 11404. one approach to avoid these issues in datasets to be transferred (or archived) is to just use the system missing value in the primary variable and then create a secondary dummy variable with labeled indicators to indicate categories of missing. these dummy variables could also be shared in a resource package. multiple value labels both sas and stata keep labels for values separately from variables. this allows multiple variables to share one set of labels. it also allows for alternative labels to be used in different contexts – e.g. longer labels for tabulation rows than tabulation columns, or labels in multiple languages. each of these packages allows a variable to have an association with one set of labels at a time. unaffiliated labels can exist in memory during a stata session but a stata “.dta” file only stores the sets of labels currently associated with variables. stata also has a dta xml export format that will export all sets of labels currently defined in a stata session. sas can export the sets of labels (called “formats”) into a separate dataset, a cntlin / cntlout dataset. with both packages the definition of multiple sets of labels can also exist in script files. an example of multiple formats in sas follows, with short labels for a variable “gender” in two languages, and a longer set of labels in english. the english value is attached to the variable. proc format cntlout=eddi.sas_fmts; value genderen 1=”male” 2=”female”; value genderde 1=”männlich” 2=”weiblich”; value genderl 1=”self identified male” 2=”self identified female”; … format gender genderen.; the stata example below show the same three sets as value labels. label define gendere 1 “male” 2 “female” label define genderg 1 “männlich” 2 “weiblich” label define genderl 1 “self identified male” 2 “self identified female” label values gender gendere in both of the preceding examples the language is represented by a user-defined convention. labels of a particular language cannot be selected in some general machine actionable way. ddi is capable of representing these multiple sets of labels in l:categoryschemes and l:codeschemes. ddi links to one l:codescheme from a variable, but each l:category in the l:categoryscheme linked from that l:codescheme may have multiple labels, distinguished by xml:lang and type attributes. there currently appears to be no way though, to indicate which label is the default, or currently associated label. in the ddi example below the labels for the male code are differentiated by both language, with the “xml:lang” attribute, and type, with the “type” attribute. ddi can associate the whole set with an l:code, but cannot indicate which label was the currently assigned label. male männlich self identified male none of the non-linked sets of labels exported from stata or sas to ddi in our experiment. role several of the packages have a defined variable attribute of “role”. this may be used to indicate which variables are eligible to be independent or dependent variables in an analysis, or to indicate more specific functions. given that the vocabulary for “role” is not standardized across packages, it is not surprising that this metadata element does not transfer. mapping against a commonly accepted controlled vocabulary would allow machine actionable decisions about comparable or incomparable proprietary terms. custom (user defined) variable attributes with the addition of extended attributes to sas version 9.4, all of the packages evaluated allowed the definition of custom attributes for variables. none of these were exported to ddi in our tests. ddi 3.2 has a facility for recording these name, value pairs in r:userattributepair elements. the adoption of a commonly accepted controlled vocabulary by data producers would allow mapping into defined ddi elements. labeled ranges both sas and jmp have the capability to label ranges of values for a variable. in this use sas formats act as an analog to a normalized structure in a relational database, allowing information to be recorded in only one place (variable). sas programs can use these formats dynamically to perform analyses on the categorized variables. since there is currently no good way to represent these in ddi, these did not transfer to ddi. perhaps it is not best practice to rely on them for archival datasets and instead create additional categorized variables. built-in display formats some display formats built in to the various packages serve mostly cosmetic functions – left or right alignment, the choice of decimal or thousands separator characters, and so on. others convey important meaning. currency symbols, for example, denote units of measurement. some sort of standardized representation for at least a subset of formats across packages would be very useful. date and time date and time conversion was not completely tested. date and time types are realized in the different packages in various ways. some (like sas) have the capability of exporting date and time data in iso formats. date and time formats according to iso 8601 should be used to assure interoperability of programs (wikipedia. iassist quarterly 2015 17 iassist quarterly iso 8601). a time value’s dependence on a time zone should be carefully checked and documented. offsets from coordinated universal time (utc) should be included where meaningful. stat/transfer provides a general means to specify the date and time format for export into a csv file where all values for a variable have the same offset. a format according to iso can be configured but is not provided by default. a fixed time zone offset can be added to all values for the variable. as an example, if all values for a variable were central european daylight saving time (cedt), a time value could be written as: 2009-06-30t18:30:00+02:00 i.e. 18:30:00 on 30. june 2009 (cedt). note that the time zone offset of +2.00 would be a constant across all values for the variable. if observations in the dataset might all have different offsets, the only option would be to add an additional variable containing the offset. r and datetime one issue we encountered with transferring datetime variables in and out of r, was that of non-explicit specification of whether datetime values represented utc or local values. when the different programs involved made different default assumptions datetime values got shifted by the local offset from utc. explicit inclusion of the time zone in datetime values would avoid this problem. here to there (and back again?) the grid like shown in figure 7 can be used to predict what metadata will survive a trip from one package to ddi and then to another package. we created new grids showing that evaluation for each of the packages. in the third to last row of figure 8, for example, labels for numeric values can be seen to transfer from spss through ddi (the yellow column, labeled “to ddi 3.1 from”) to spss, stata, sas, and jmp. they would not survive the journey from spss to r since r factors do not correspond exactly to labeled numeric variables. the complete set of these grids can be seen in appendix 3. suggestions for ddi this experiment generated several suggestions for ddi 3.1. these are listed below in rough order from the most specific to the most general. multiple labels for a category a mechanism to indicate whether one of a set of labels for an l:category was currently assigned to a variable (or not) would be useful for representing sas and stata data. it would also add clarity to the current representation of multiple labels for a category in ddi. this could perhaps rely on the “type” attribute of the label element. the following example shows a set of labels for a gender variable, both in multiple languages and with a “long form” in english. which “type” attribute should be selected by default when referencing the category? figure 8 – example here to there grid 18 iassist quarterly 2015 iassist quarterly kvinna weiblich female self identified female perhaps this could be specified within the l:coderepresentation element. one possibility would be for l:representation within l:variable to contain a “primarylabeltype” selecting a “type” attribute from the set of labels as in the example below. string gendershort 5c706c37-d19b-4b8e-ac6d40094024421f example.org 1 ranges a facility to assign categories to ranges of coded and numeric variables would allow representation of these features from sas and jmp datasets. it would also allow a range of values to be labeled as missing. ddi 3.2 allows this for numeric with a managednumericrepresentation. see the conclusions section for more details. precision vs. display information numeric output formats can convey an ambiguous amount of information about the precision of measurement of a variable. by convention, a value formatted in “scientific notation” as 1.2345 x106 would be considered to have been measured to five digits of precision. the level of precision for same value formatted as a decimal with zero digits to the right of the decimal point (1234500) is not clear. it would be useful for ddi to have an explicit representation of level of precision of measurement. integrity constraints several of the packages we evaluated allow for defining some sort of integrity constraint on a variable, either by a list of valid values or by logical expression. an expression like mod(age,1)=0, for example could be used to only accept integer values of age (even though stored as type float). expressions might also refer to multiple variables, as in an expression limiting years of education to age – 5. sas also allows foreign key integrity constraints, only allowing entry of values appearing in a column in another table. ddi 3.1 seems to lack an explicit representation of integrity constraints. ddi 3.2 has a workaround described in the conclusions section below. data types the data types defined in the statistical packages generally do not transfer perfectly. note that data type is distinct from display format. the most commonly used data types in the packages are integer, double precision, boolean, and character strings. the data types can be captured in ddi in p:physicallocation/ p:storageformat. the related documentation recommends the use of a controlled vocabulary. one way would be to develop it on the basis of xml schema data types (w3c 2004) and/or sql data types (wikipedia. sql -data types) which both seem to comprehend the most used forms of data types. without a standard for output formats against which to map the various proprietary formats, automated translation from one package to another via ddi is difficult. documentation for the ddi element r:genericoutputformat recommends the use of a controlled vocabulary. perhaps one could be developed based on java, c, or fortran. multiple distinct system missing values ddi could use a more explicit method of representing multiple “out of band” missing values such as used by sas and stata, or the inf, na or nan values in r. these values may sometimes also be labeled, necessitating links to l:categories. ddi 3.2 allows this with the managedmissingvaluesrepresentation element. text descriptions only several other generic metadata elements that can be represented in some of the packages can currently only be represented as r:description or r:note elements in ddi. more machine actionable representations of these elements would be useful. date created some packages (and operating systems) may track not only the date a dataset was last modified, but also its initial creation date. in ddi3.2 this can be recorded in pi:physicalinstance/pi:datafileversion/@ versiondate or pi:physicalinstance/r:citation/ dc:created. publication date can be recorded separately in pi:physicalinstance/r:citation/r:publicationdate scripts some packages can store scripts in the same structure as the dataset. currently these can only be represented in ddi as an l:logicalproduct/r:description element. a ddi element identifying the script as a script written in a particular language (e.g. r, jmp scripting language) would be useful. notes some packages allow a “note” attribute to be attached to a dataset or variable. while such a note can be preserved in an an l:logicalproduct/r:note, it would be more machine actionable to specifically identify the text as having come from a “note” in the original format. some packages treat a note as just another (name,value) pair. in ddi3.2 this could be recorded with a value iassist quarterly 2015 19 iassist quarterly of “note” in the r:attributekey element of a r:attributekey , r:attributevalue pair in either an r:proprietaryproperty or in an r:userattributepair. value colors some packages allow assignment of colors to particular sets of values or observations. colors may represent some manually assigned attribute of the data, such as suspected outliers, or, in the case of spss, imputed values. while in one sense this could be represented by a code and category scheme, it might be more machine actionable to flag the color values as representing colors in some way. filters (triple-s) triple-s allows a variable to point to another variable as it’s “filter”. when the filter variable has the value true, the variable pointing to it is available for that case. a workaround for numeric variables is described in the conclusions section below. suggestions for transport programs internal model for categories and codes some packages, like spss or jmp, store value labels as attributes of variables. others, like stata and sas, store sets of labels (formats in sas) separately and tie them to variables by reference. since the latter method can represent anything the former does, it should be the basis for the internal model of value labels (categories and codes in ddi). detect reuse of sets of value labels (categories and codes) once sets of value labels are represented independently from variables and used by reference, they do not need to be defined multiple times. when converting from a representation where a given set of labels might be defined multiple times (e.g. a likert scale in a survey) to one like ddi where a set can be reused, it is desirable to identify and eliminate the duplication. for one approach to this process see (wright, 2011). multiple sets of categories and codes applicable to a variable as discussed above, it is possible to have multiple sets of categories and codes applicable to a given variable, for example labels in different languages or of different lengths. with both sas and stata only one of the sets can be associated with a given variable. with ddi multiple languages and types can be associated with a variable. an internal model that allows multiple associations and specification of a default or primary set would allow representation of all of the possibilities. use the dataset name when distinct from the file name in cases like r where the name of the dataset is not necessarily the same as the name of the file containing it, the dataset name should be used. generic vs. proprietary information metadata harvested from a proprietary dataset can be classed into four categories: 1. generic information that corresponds to a ddi element 2. generic information for which there is no specific ddi element 3. proprietary information that corresponds to a ddi element 4. proprietary information for which there is no ddi element with ddi 3.1, information of type 2 could only be recorded in an unstructured r:note. information of type 4 could be recorded in a r:proprietaryproperty. with ddi 3.2, information of both types 2 and 4 can be recorded in r:attributekey , r:attributevalue pairs – in either an r:proprietaryproperty (category 4) or an r:userattributepair (category2). packaging structure ddi3.1 allows packaging many elements in either an s:studyunit or a g:resourcepackage. ddi 3.2 adds a third alternative with the ddi:fragmentinstance element. ideally a transport program should be able to handle any of the three structures. unfortunately this is not always the case. the ddi community should probably make a recommendation for a preferred structure for transport instances. suggestions for statistical and data management software packages more metadata all of the packages reviewed here are lacking in the ability to include enough structured metadata in a dataset to actually interpret the data. many packages allow the attachment of (name, value) pairs of attributes, but without those names and values coming from some sort of structure or controlled vocabulary the metadata have limited interpretability or searchability beyond the data’s creators. metadata elements such as concept and universe are relevant to a broad array of data. knowing that a variable was only measured on male children, for example, is important when drawing inferences based on analysis of that variable. for data captured by surveys, at a minimum it is important to know the exact text of the question asked of respondents. metadata about groups of concepts, questions, and variables are also important. geographers and others point out that all data are spatial. metadata about spatial coverage are proving increasingly useful. flexible structure instances of metadata from many metadata standards are expressible in xml. one possibility would be for statistical and data management packages to add the capability to attach an xml instance from a defined standard to their datasets. for well-known standards like ddi it would also be possible to integrate the use of that xml into their procedures. survey analysis procedures could incorporate metadata like question text and even question flow. with some software packages it might be possible to include ddi xml in a key value pair. suggestions for research practices and preparing archival datasets this experiment generated a few suggestions for the practice of preparing archival datasets. reason for missing – multiple missing types use an auxiliary variable to indicate reason for missing. these could be shared in a g:resourcepackage. figure 9 shows a variable “measuremissing” which differentiates type of missing for the variable “measure”. the pairing of these two variables could be documented with a l:variablegroup element. alternative formats create additional variables for data representable with alternative formats 20 iassist quarterly 2015 iassist quarterly • long labels • languages • coded ranges doing this for coded ranges requires some additional metadata though. the continuous variable body mass index (bmi), for example, has associated ranges indicating categories such as “underweight”, and “obese”. these ranges are further broken down in a hierarchy with multiple sub-categories. an additional variable could be recoded from a bmi measurement, but the rules for recoding would also need to be recorded (in a d:generationinstruction) in order to ensure that the exact ranges for each category were captured. user attributes use a controlled vocabulary for the names of user attributes (characteristics, properties) where available. if this practice were common, then a much wider range of metadata would be transferrable across packages, without the need for revisions to the programs. one possibility for such a controlled vocabulary might lie in a semantic data form of ddi elements. where available, also use a controlled vocabulary for the values of user defined attributes. an example would be to use an attribute named “analysisunit” with values taken from the ddi analysisunit controlled vocabulary. (ddi controlled vocabularies working group) time zones where possible, explicitly specify time zone information in datetime values. datetime values without explicit specification of time zone may be unpredictably interpreted as local times or universal times as they are converted from package to package, leading to changes in the data. conclusions a few general conclusions can be drawn from this experiment: adoption of ddi by tools like stat/transfer is encouraging. basic metadata is transferrable among all 5 packages via ddi. the current state still means that some important metadata that might be contained in proprietary format data files still must be either hand entered into ddi or harvested and entered by userwritten code. no one package has a superset of the other’s metadata. several desirable elements are not universally supported. some desirable elements like concept and question are not supported by any of the packages, except as user defined (name, value) pairs. the fact that all of the metadata typically recorded in a proprietary dataset file can be represented in a g:resourcepackage without an s:studyunit reveals a lack of a structured facility for recording information about the origin of that dataset. the development of best practice recommendations for using custom attributes of variables and datasets could be one approach toward remedying this situation. as mentioned earlier, ddi3.1 allows packaging many elements in either an s:studyunit or a g:resourcepackage and ddi 3.2 adds a third alternative with the ddi:fragmentinstance element. the ddi community should specify one preferred packaging structure to be used as a transport instance. ddi is almost a superset of the packages considered. being able to represent a superset of metadata elements across the most commonly used packages is a worthy goal for ddi. some missing elements like mentioned above should be added for this purpose. our suggested list of needs follows. • a facility to define an assigned set of value labels to a variable. this could be done by specifying a preferred type and language for subsetting a category scheme. • a facility to assign categories to ranges of coded and numeric variables. ddi3.2 now allows a managednumericrepresentation figure 9 an auxiliary variable for “measure” indicating type of missing iassist quarterly 2015 21 iassist quarterly to have a set of numberrange elements, each of which can assign a label to the range. manageddatetimerepresentation allows a label to be assigned to a duration. the managedtextrepresentation element does not have a corresponding “textrange” element. • a method of recording the level of measurement precision to variables, e.g. the number of significant digits. this is distinct from the number of digits to be displayed to the right of the decimal point. as an example the number written as 123000 has 0 digits to the right of the decimal point and carries no information about the precision to which it was measured. if expressed as 1.230e5, though, the convention is that there are four significant digits. • a facility for recording complex integrity constraints. in ddi3.2 a r:processinginstructionreference can be attached to a l:variablerepresentation which can then refer to a d:generalinstruction. this could be used to describe a constraint as a logical expression that must hold true for the value of the variable. this is a usable workaround, but the semantics are too narrow in that a constraint is a rule for the variable independent of how it is applied. a more general approach in future ddi version could be an instruction element having a type attribute. the type might take on values of ”derivation”, or ”constraint”, or ”processing”. note that foreign keys can be described with a recordrelationship. • integrity constraints can also take the form of ”unique” or ”not null” (the combination being required for a key). these specifications could be added to variablerepresentation or valuerepresentation to be inherited by substituted elements. • development of controlled vocabularies for data types and output formats • a more explicit method of representing multiple distinct system missing values. in ddi3.2 this is now possible using a managedmissingvaluesrepresentation. this is still a significant challenge when moving data from a package supporting multiple missing values to one (e.g. r) which does not. • a method for explicitly representing scripts stored in a dataset. currently scripts for specific data transformations can be stored in d:generalinstruction. scripts for other purposes or roles such as analyses or visualization could use some structure. in the future this may become part of a process model in ddi. • a facility for recording colors attached to either variables or sets of observations, including the color value and an associated concept or category. with ddi3.2 color assignments for a code could be recorded in a r:userattributepair attached to its associated l:category. for numeric variables a color gradient could be recorded by r:userattributepair elements attached to a r:managednumericrepresentation. a work-around for assigning colors to individual values could use a structured r:label of a r:numberrange. • facilities for a variable to point to a boolean variable as its filter (missing indicator) and to a categorical variable to indicate different types of missing. for a boolean filter of a numeric variable, a weight variable is equivalent to a filter indicator, where a weight of 0 indicates missing and a weight of 1 indicates valid. this is not technically correct for a string variable and does not work for the categorical case. what would be desirable is a reference to a variable for which a role could be specified. future work a longer version of this article is planned for publication in the ddi alliance working paper series. the paper will focus additionally on solutions which are developed in the ddi moving forward project on the next generation ddi. the appendix will have an extensive mapping table describing the metadata elements in each package and ddi. references algenta technologies 2012. colectica reader the free ddi 3 viewer ddi metadata and survey design software. amin, a., barkow, i., kramer, s., schiller, d. & williams, j. 2011. representing and utilizing ddi in relational databases. ddi working paper series: ddi alliance. (doi: http://dx.doi.org/10.3886/ ddiothertopics02). circle systems stat/transfer. (url: http://www.stattransfer.com/). data documentation initiative. 2013a. ddi rdf vocabularies | ddi data documentation initiative [online]. available: http://www.ddialliance.org/ specification/rdf. data documentation initiative. 2013b. ddi tools | ddi data documentation initiative [online]. available: http://www.ddialliance. org/resources/tools. data documentation initiative. 2013c. future plans for ddi development | ddi data documentation initiative [online]. available: http://www.ddialliance.net/ ddi-moving-forward-process-summary. ddi alliance expert committee. 2009. ddi 3.1 xml schema documentation (2009-10-18) [online]. available: http://www. ddialliance.org/specification/ddi-lifecycle/3.1/xmlschema/ fieldleveldocumentation/. ddi controlled vocabularies working group ddi controlled vocabularies overview table. ddi alliance. (url: http://www. ddialliance.org/specification/ddi-cv/). hoyle, l., wackerow, j. & hopt, o. 2010. ddi 3: extracting metadata from the data analysis workflow. in: vardigan, m., edwards, m. & hoyle, l. (eds.) ddi working paper series -use cases. ddi alliance, http://www.ddialliance.org/resources/publications/working/ usecases: data documentation initiative alliance. (doi: http://dx.doi. org/10.3886/ddiusecases04). ibm ibm spss statistics. (url: http://www.spss.com). r development core team 2009. r: a language and environment for statistical computing. in: computing, r. f. f. s. (ed.). vienna, austria. (url: http://www.r-project.org). sas institute jmp jmp statistical discovery software. sas institute. (url: http://www.jmp.com/). sas institute sas the sas system. (url: http://www.sas.com). statacorp stata. (url: http://www.stata.com/). w3c. xml schema part 2: datatypes second edition w3c recommendation 28 october 2004 (url: http://www.w3.org/tr/ xmlschema-2/#built-in-primitive-datatypes) wikipedia. sql -data types (url http://en.wikipedia.org/wiki/ sql#data_types) wikipedia. iso 8601 (url:http://en.wikipedia.org/wiki/iso_8601) wright, p. a. 2011. eliminating redundant custom formats. sas global forum 2011 sas institute. (url: http://support.sas.com/ resources/papers/proceedings11/217-2011.pdf ). appendices appendices and possible other related material may be found at http://hdl.handle.net/1808/19900. a pdf document containing the three appendices may also be accessed directly at https://kuscholarworks.ku.edu/bitstream/ handle/1808/19900/ddiasacommonformatiassistqappendices. pdf. 22 iassist quarterly 2015 iassist quarterly notes 1 larry hoyle is a senior scientist at the institute for policy & social research at the university of kansas and can be reached by email: larryhoyle@ku.edu. joachim wackerow is a metadata expert at gesis leibniz institute for the social sciences and can be reached at joachim.wackerow@ gesis.org. 19.3 31fall 1995 www: what do researchers want? summary of 44 responses to an e-mail survey february-march, 1995 by jim henderson1 maine state archivist introduction during february, 1995, i distributed a brief survey, “www: what do researchers want?”, to approximately thirty history and other listserves. others were approached, but not all allow non-subscribers to use their lists. a single “reminder” emailing was sent in march. each distribution generated just over 20 responses, for a total of 46. all responses were received electronically. in brief, researchers want clear guides to collections, supplementary information about the institution and its mission, and access information: rules for copying; mail, phone, e-mail information. they are far less interested in “cute” sample images (the olde map or photo) or sample text of selected collections. parochial items such as organizational structure or exhibits and upcoming events are clearly the lowest priorities among those listed. researchers most highly value “subject oriented keywords pointing to related collections,” and “detailed descriptions” and “finding aids” for major collections. next, they want to know the ways and means of access: rules about the cost and availability of copies, both traditional (mail, phone) and e-mail contact information. after the basics, and to get a view of the institution’s possibilities, researchers want 1) listings of collections by genre, 2) lists of guides, pamphlets, and other publications, supported by 3) reference room hours and procedures, and 4) a general description of the institution’s holdings and mission. following closely are interactive needs: the ability to leave messages for the staff and to find out “what’s new?” while given “some importance,” image databases of photos and maps were deemed slightly less useful than the proposed textual databases, which also were not highly sought after. selected sample items, by both typical content and format, were viewed unfavorably by one-third of the respondents. internal and local items characterized by “organizational structure” and “upcoming events” received rather negative reviews. current research listserve members find little interest in genealogical holdings, but this may say more about the respondents and the current availability of technology than about the potential broad interest in this information. the detailed responses: an analysis respondents were asked to rank the proposed features as very important (v), important (i), some value (s), not important (n), or forget it (f). the first three columns at the right below display the percentage of responses to the two lowest ratings (fn), the middle (s), and the highest ratings (iv). (rounded percents may not add to 100.) the next column reflects the average of all responses, with very important coded as 4; important, 3; some value, 2; not important, 1; and forget it as 0. the ranking of average scores sometimes differs from the order of the highest ratings because of 1) the varying portion of respondents choosing “very important, important, etc. as their selections, and 2) the fact that a few respondents did not respond to all items. the final column notes the standard deviation from the mean of all responses. the lower the number, the greater cohesion among respondents, with a tendency to cluster about one of the five choices. standard deviations ranged from .68 to 1.11. highest rated features the most desirable features center around three themes: collections level descriptions, availability of copies, and contact information. while “subject oriented keywords” were most valuable, “detailed descriptions” and “finding aids” of major collections where highly regarded. rules concerning the cost and availability of copies ranked second overall, while requirements to have both traditional (mail, phone) and e-mail contact information were equally valued as high priorities. all features in this group had average scores tightly clustered from 3.2 through 3.4. except for the “mailing, location,” item, they also had relatively low standard deviations. basically, these feature are highly rated, with over 80% endorsement as 32 iassist quarterly important or very important, and represent the relatively uniform opinion of researcher respondents. percent proposed feature fn s iv a v s d 9. subject oriented keywords pointing to related collections 2 9 89 3.2 .78 4. availability, cost, restrictions regarding copies of records 0 12 88 3.3 .68 8. detailed description of major collections 2 11 86 3.4 .69 19. finding aids for major collections 7 7 86 3.3 .79 2. e-mail addresses of site and key staff/departments 2 16 82 3.4 .75 1. mailing address, location, telephone number of the site 2 16 82 3.3 .95 table 1 majority supported features after a fairly clear break of 15 percent in the important/very important ratings, the following appear to be “helpful, supplemental” features. these features all rank from 3.0 (important) to 2.5 (mportant/some value). to get a view of the institution’s possibilities, researchers want 1) listings of collections by genre, 2) lists of guides, pamphlets, and other publications, supported by 3) reference room hours and procedures, and 4) a general description of the institution’s holdings and mission. at 3.0 and 2.9, these essentially rate important on average. following closely are interactive needs: the ability to leave messages for the staff and to find out “what’s new?” rather lower in this group’s ranking (and close in content and rank to the first two features in the next section) are requests for textual databases describing photographic and cartographic holdings publications articulate a cluster of desirable features. percent proposed feature fn s iv a v s d 21. list of collections by genre: text, map, photo, video, audio 9 23 67 3.0 .93 6. list of guides, pamphlets and other publications 0 34 66 3.0 .85 14. ability to leave message for archives staff 5 30 65 2.8 .81 20. what new? (acquisitions, finding aids) 9 29 64 2.8 .95 5. reference room hours, procedures, rules 16 18 66 2.9 1.02 7. general description of holdings and mission 1-2 pages 5 36 59 2.9 .99 15. database of textual description of photographs 20 25 55 2.5 .96 16. database of textual description of maps 20 25 55 2.5 .96 table 2 lowest rated features the final six features had no majority expression of combined important/very important responses. all rank below 2.5 on average and cluster around the some value rating of 2.0. image databases of photos and maps were deemed slightly less useful than the proposed textual databases in the previous section, though some comments indicated a “nice, but utopian” view of the proposition. selected sample items, by both typical content and format, were viewed unfavorably by a third of the respondents. 33fall 1995 internal and local items characterized by “organizational structure” and “upcoming events” received rather negative reviews. interestingly, while the overall ranking of these two items is similar, the standard deviation reveals virtual consensus (sd=.68) on the limited value of “exhibits, upcoming events,” but a wide disparity of views (sd=1.11) on the value of organizational structure information. current research listserve members find little interest in genealogical holdings, but this may say more about the respondents and the current availability of technology than about the potential broad interest in this information. percent proposed feature fn s iv a v s d 17. image database of photographs 20 34 45 2.4 .87 18. image database of maps 20 39 41 2.4 .87 3. description of organizational structure. 32 30 39 2.2 1.11 10. selected sample items, major collections, typical content 32 43 25 1.9 .91 12. genealogy holdings summary 32 43 25 2.0 .87 11. selected sample items, major collections, typical format 34 43 23 1.9 .91 13. exhibits, upcoming events 27 64 9 1.8 .73 table 3 ranking by average rating in yet an other arrangement of responses, this time by average rating, similar conclusions may be drawn. 3.5 important to very important 3.4 detailed descriptions / e-mail addresses 3.3 info about copies/major finding aids/postal mail, location, phone 3.2 subject oriented keyword searches 3.1 3.0 important list of guides, publications / list of collections by genre 2.9 reference room hours, rules / general holdings, mission 2.8 leave messages for staff / what’s new? 2.7 2.6 2.5 some value to important databases of textual description: photographs and maps 2.4 image databases of photographs and maps 2.3 2.2 description of organizational structure 2.1 2.0 some value genealogy holdings: summary 1.9 selected sample items indicating typical format and content 1.8 exhibits, upcoming events 1.7 1.6 1.5 not important to some value table 4 34 iassist quarterly the survey as sent sorry for duplications. this survey has been posted to over 30 history lists. i have not posted to lists focusing on non-north american history. feel free to post to other lists you think appropriate. please respond by february 24th. thanks. what do researchers want to know from archival sites especially regarding www design? the maine state archives is about to establish a www site and series of “pages.” the last few months have seen an explosion in this area. while we have reviewed many of the new sites, we wonder “what do researchers want to know?” here’s you chance! i will post this survey’s results on the archives listserve and anywhere else it may be helpful. your responses may apply to gopher design as well. keep in mind that not all wishes are granted if everything is “very important” then .... please reply to me hendersn@saturn.caps.maine.edu -and not to the list on which this is posted! —————— archives www design survey ———————— how important are the following in an archival electronic information site? (v)ery important (i)mportant (s)ome value (n)ot important (f)orget it! 1. mailing address, location, telephone number of the site 2. e-mail addresses of site and key staff/departments 3. description of organizational structure. 4. availability, cost, restrictions regarding copies of records 5. reference room hours, procedures, rules 6. list of guides, pamphlets and other publications by the site, including ordering info 7. general description of holdings and mission 1-2 pages 8. detailed description of major collections: title, inclusive dates, summary, scope 9. subject oriented keywords pointing to related collections 10. selected sample items from major collections indicating typical content 11. selected sample items from major collections showing typical format through displayed images 12. genealogy holdings summary 13. exhibits, upcoming events 14. ability to leave message for archives staff 15. database of textual description of photographs 16. database of textual description of maps 17. image database of photographs 18. image database of maps 20. finding aids for major collections 21. what new? (acquisitions, finding aids) 22. list of collections by genre: text, maps, photographs, video, audio places you have been (virtually), like, and why: ___________________ additional comments, suggestions: ___________________________________ thanks. please respond by february 24 (later march) to hendersn@saturn.caps.maine.edu jim henderson, state archivist cultural building, station # 84 augusta, maine 04333 (207) 287-5790 hendersn@saturn.caps.maine.edu maine archives bbs 207-287-5797 35fall 1995 comments from respondents: places you have been (virtually), like, and why: i use congress and related gophers to collect data and the text of documents and bills as well as information of votes. some servers provide campaign expenditure data which is useful to me. i sometimes connect with electronic collections (e.g. guttenburg project). so the availability of raw and secondary data is useful to scholars like myself. a second useful area involves access to catalogues and directories. often, it is sufficient to know that a document exists and what the source is without actually viewing the document. similarly, directories of various types are useful. ———— still exploring, but i liked the oregon state university site wpa exhibit, the cornell exhibit, and see incredible value to the johns hopkins gopher site. ———— u. of michigan mlink (gopher://mlink.hh.lib.umich.edu/) lots of information and links to other sources, sensibly arranged and easily accessed. ———— ukanshistory research ———— thomas, because you can go back and forth in your searches — nice menuing. the star trek site at paramount has some nice moving around tools, too: ———— places with a good, thought out design, graphics that transfer quickly or very few graphics at all. additional comments planning: i think it will be important to think through the various audiences you want to reach—not only now, but in the future. [perhaps a review of the site’s (organization’s) mission statement would be helpful.] the site should be designed with that (those) audience(s) in mind, and people need to recognize that one structure may not meet the needs of all audiences or users. having said that, let me add that collecting survey information from somewhat experienced web users is one good source of information. another source might be focus groups with intended users. also, if funds permit, you could draft a basic design, set up some pilot tests with various groups of intended users (librarians, teachers, students, others). then watch what they do. see what they like and don’t like. what’s confusing and what seems to flow more naturally, etc. then debrief through focus group interviews ————— the reason i think images have only limited importance is a) some people are still using text-based readers, b) images take a really long time to load, c) even at their best, you can’t always tell if an image is what you want. this is especially true for maps. also, if you are going to have messaging for staff, you need to have a really reliable system of responding. an explanation of the searching tools would be nice. ————— i would say that indexing and maybe full text of various state publications, periodicals, newsletters, and even local newspapers would be very helpful to many researchers. if full text is available i would say that indexing is unnecessary. including local newspapers would be extremely helpful as i am sure you know that national newspapers tend to not cover maine very much or very thoroughly. this would provide a gold mine for people who are doing research on maine. ————— search & preview capabilities should be maximized. ————— 15-18. i am a bit confused on this. if you mean all photographs and maps, then that would be wonderful. many researchers then would not even need to travel to a particular site. but that would also be a massive task for the people at a given site and could demand immense computer memory. if, on the other hand, you mean only certain photographs and maps, would that be any different from what you mean in #10 and 11? or, do you mean detailed descriptions of the “collections” of photographs and maps in a given archival site, and if so, would that be much different from what you mean in #8? ————— tourist information—hours, locations, costs, restrictions—are not necessary to me as a research scholar. i do need to know what you have. i may need to search your holdings by any word or combination of words in order to know whether i need to ask about hours and restrictions. even more wonderful, would be the ability to scan texts to determine the value of documents. 1. paper presented at iassist95 may 1995 quebec city, quebec, canada i i []r«sfi[iisi[is fi • [ifitfl liervflkli ann e. gerken, data archivist comell institute for social and economic research cornell university historical background and development the cornell institute for social and economic research (ciser) is an interdisciplinary organization of cornell faculty which seeks to support, strengthen and enrich the social and economic research community. in may 1981, ciser was founded to develop and support research programs and provide services and facilities required for research projects. the ciser data archive was established in february 1982 in cooperation with the cornell university libraries to provide central access and management for social science data to researchers on and off campus. data archive staff provide professional information services, technical consultation, and research support. ciser also sponsors workshops and seminar series, peer review of research proposals, grant management, computing facilities, newsletters, and a directory of research interests of faculty in the social sciences at cornell. the ciser data archive was established upon the recomendations of a committee made up of representatives from four colleges at cornell, the university libraries, cornell computer services, and the new york state cooperative extension service. the committee based its recommendations upon a survey of ten data archives located within research centers. information was gathered on staff, collections, funding, space, and computing consulting. the data archive's goals are to: 1. establish and maintain a centralized archive of machine-readable tapes and documentation; 2. acquire data and supporting documentation, coordinate buying consortiums, fill gaps in data file holdings, and assure the safekeeping of archival holdings; 3. provide an information center with professional reference and computer consulting in social science data, defining information needs and providing research services; and 4. support the research and service missions of the institute and the university. with the assistance of cornell's social science librarian, ciser opened the data archive in early february, 1982. a professional archivist joined the ciser staff and assumed administrative duties as well as the planning responsibilities for the development of the archive. data files were acquired, policies, mailing lists and ordering procedures were established, and a survey of faculty was taken to identify data files on campus and those needed. the survey was helpful in locating data files to incorporate into the archive, in establishing contacts with researchers, and in developing a collection policy. 4* since its opening in 1982, the data archive has grown significantly. staff, services, and the holdings of machine-readable data files have expanded. the archive is an increasingly important research support facility on campus, providing essential services to a wide range of users. in addition to walk-in information and consulting services, the archive offers workshops, seminars, and classroom lectures on data file contents and use. the integration of computer consulting with data file reference service makes the archive a unique resource at cornell. the staff is dedicated to the provision of continuous support to the educational and research activities in the social sciences, from data information to advanced analytical support. funding ciser and the data archive are supported by allocated funds from five colleges at cornell. the acquisitions portion of the budget is relatively small since most data sets are acquired through cornell's membership in the inter-university consortium for political and social research, through new york state data center affiliate status, and through other cooperative agreements with state and federal agencies. in addition, a number of files are donated to the archive by the faculty. a collection policy regarding acceptance of donations and purchase decisions is vital to the development of the archival holdings. computing costs and tape storage costs are separately allocated. staffing the staff of the archive consists of a professional data archivist, two fulltime computer consultants, and a half-time data manager. part-time student assistants are also available during the academic year. the data archivist's duties include administration of the archive, data file evaluation and acquisition, data file information services, policy making, and the development of new services for social science researchers. the archivist works closely with the cornell university libraries to develop integrated information services, and communicates with the computing services staff in regard to technical developments and services. the archivist is a professional information specialist with an academic research library background. the computer consultants provide statistical computing consultation and are responsible for the technical development of the data archive. they also work on contract for special data projects producing custom files, reports, analyses, and data management. the consultants have social science backgrounds with experience in statistical analysis and computerized research techniques. the data manager is a half-time employee with tape management responsibilities this person keeps inventories of tape contents and data users and oversees the addition and copying of tapes in the collection. fundamental knowledge of the computer and tape management systems is required. student assistants perform tape management tasks and edit inventory files. valuable management and secretarial assistance is also available. physical environment the archive is housed with the other offices of ciser. one large office houses the archivist, the data manager, and the library of technical documentation and reference materials. other offices house the consultants. each office has space for archive users to examine materials. print materials can be taken out overnight. the consultants' offices have enough space for small group instruction and storage space for printouts and other records. ^ equipment "^ as for local equipment, the archive has an ibm-pcl, will be acquiring additional microcomputers, and has two terminals which are used to communicate with the cornell mainframe computers. the microcomputers are also used as terminals for communication and data transfer, as text processors, and for social science workstation development, including database management, graphics, and custom programming. data on floppy disketts are being distributed. public computing facilities in the building offer state-of-the-art graphics equipment, high speed printers, consultants, and technical manuals. ciser also has a computing facility in the building, with terminals for use by ciser members and their research assistants. extensive microcomputer facilities are also located in the same building as the data archive including a software library and a demonstration facility. other hardware access includes an ibm 3081d, and a dec 1020. the holdings of the archive are stored at the mainframe facility. tapes are used on the ibm and also are exported for use on other systems accessible to cornell researchers. sources of data the ciser data archive holds machine-readable data in the areas of demography, vital statistics, health, social surveys, labor and employment, occupation international trade, business, service industries, education, agriculture, and life studies and aging. the archive comprehensively collects new york state data. ciser is an affiliate of the new york state data center, and acquires many of its files from that source. cornell is also a member of the inter-univeristy consortium for political and social research (icpsr) which provides the majority of non-census files to cornell. longitudinal data are acquired from the bureau of economic analysis. the archive receives data from numerous government agencies, both federal and state, and also acquires data files from other research institutions and survey centers. the contribution of research data files by cornell faculty members have been central to the development of the data archive. the archive seeks continuous data deposits to build longitudinal strengths. detailed collection development policies are developed in collaboration with faculty, librarians, and members of the institute. dissemination of information to make users of the archive aware of the holdings, a title list of data files in the archive is updated and distributed bimonthly. information about the archive, its holdings and services, is included in the bimonthly newsletter from ciser, called the syntheciser . the archivist meets with faculty, graduate students, and staff to advertise and promote use of the archive, teach methods of identifying data files, and establish a network of data users. future developments will include online directories of holdings and variable-level indexing of statistical data files. cooperative cataloging of machine-readable records is being investigated. in addition, a model relationship with one of the cornell university libraries has been established whereby professional staff development, acquisitions, and information dissemination is coordinated. data are disseminated through tape and disk access on the ibm mainframe, through file transfer to tapes and diskettes, and in special data files and printouts provided on custom bases. e users of the ciser data archive the users of the archive represent the many colleges and departments at cornell, and numerous off-campus organizations. services are available to faculty, graduate and undergraduate students, staff, off-campus service agencies, private corporations, government agencies, and the greater ithaca community. access to the holdings of the archive on the mainframe are limited to those with cornell computer accounts. information and data delivery services are available to others on a contractual basis, except when data are restricted to use by the cornell community. fee structures have been developed to recover costs of some tasks. the growth of the archive is evident in the increasing numbers of walk-in and returning users. future developments there have been a number of developments that present challenges to ciser and affect the long-term plans for the archive and the institute. among the goals is the expansion of the archive collection and services. the archive staff's expertise in census and other federal data products has brought an increasing number of requests for special workshops, tabulations, and data file extracts of federal data files. as an affiliate of the new york state data center, the archive frequently provides assistance to people throughout new york state. requests for assistance with federal data from the cornell community and off-campus institutions and organizations are expected to increase. researchers require subfiles and increased consultation for the larger longitudinal data sets and the microdata files in the data archive. an enlarged subsetting service and increased expertise in the management of heavily used data sets are being developed. in addition, technical support agreements, workshops, and classroom lectures on data file management and research computing techniques are being offered. the data archive must keep pace with the rapid changes in computer technology and the applications for social science research. especially important is the integration of microcomputer workstations into the research environment, with development toward multi-user and multi-level computing capabilities. the archive staff will be working with the cornell microcomputer evaluation and development facility in the application of microcomputers in social science research. the data archive also hopes to develop cooperative agreements with the cornell university libraries in efforts to integrate services and encourage extended participation in computerized statistical information services. the data archivist participates in the development of an integrated library system at cornell, and is a member of a cross-campus working group on statistical data resources. finally, the archive staff will be active in the development of a demographic, economic, and social computerized information system on new york state, to support information, research, and training activities. inc., hilary baker. each of these organizetions provides time-sharing network access to large-scale numerical, social, and economic data bases. topics covered were vendors' information services, content and scope of the data bases, data sources and update procedures, access and analysis capabilities, and costs. after the presentations, workshops and demanstrations were given on system access, retrieval, and analysis capabilities. social trend studies: a review essay richard sobel princeton university with the wealth of data now archived at centers like the inter-university consortium for political and social research (icpsr) and roper, it is not surprising that there have recently appeared a number of studies of historical trends in social attitudes. two sourcebooks from icpsr (via harvard university press) and a compendium from the national opinion research center (norc) are examples. american social attitudes data sourcebook presents trends in attitudes and behavior from 1947 to 1978. its companion, american national elec tion studies data sourcebook , includes election and demographic data from 1952 to 1978. a compendium of trends on general social survey questions examines the changes in issues, attitudes, and demographics from the late 1930s to 1970. the social attitudes sourcebook presents 84 repeated items drawn from a potential pool of 500. these variables are taken from 15 major studies archived at icpsr, including the panel study of income dynamics, the american national election studies, the surveys of consumer finances, the seasonal surveys of consumer attitudes and behavior, and the fall and spring omnibus studies. developed by philip converse, jean dotson, wendy hoag and william mcgee, the report runs from personal items to abstract, national issues. the presentation begins with attitudes toward self and others, toward racial issues, toward women and family living, and toward retirement. subsequently treated are government spending, war and peace, and outlooks on personal finances and on the national economy. tables are broken down by demographic variables including sex, race, age, and education. the election sourcebook has data from the 14 american national election studies conducted by the survey research center and the center for political studies at michigan during election years from 1952 to 1978. prepared under the direction of warren miller, arthur miller, and edward schneider, it is based on the expansion of a 1970 idea for studying "who votes and for whom." in attempting to locate electoral behavior and specific elections in historical and demographic context, the report is broader but less deep than the ameri can voter (1960), whose tradition it continues. chapters include the social characteristics of the electorate, partisanship, position on public policy issues, support of the political system, involvement and turnout, and, finally, the vote. the norc compendium is based on questions found in the general social survey, since 1972" a longitudinal study of social indicators and attitudes. while 44 concentrating on the gss years 1972-78, the compendium includes other studies which have asked comparable questions since the 1930s. trends in 238 attitudes, personal evaluations, and behaviors, and in 57 demographic items are shown in time series. topics ranee alphabetically from opinions on abortion to work attitudes. in addition, there are six clusters of issues: satisfaction with life, attitudes toward racial integration of the schools, job characteristics, qualities of children, opinions on the effects of pornography, and images of foreign countries. the study includes a methodology for analyzing the trends in the time series and indicates the type of trend for each variable. individually and in the aggregate, these studies provide an overview of stability and change in social attitudes and behavior over the last generation. each merits additional systematic study, but a few noteworthy points may be mentioned. for instance, the level of general happiness has declined slightly over time, so that today only one in three claims to be very happy. half of all employed people work in white collar jobs, and a similar proportion of all employees are very satisfied with their work. confidence in the president is low (13 percent) and falling. while only two percent have been robbed and seven percent burgled (unchanging figures since 1972), one in five has been threatened with a gun, and a similar number own hand guns. another 30 percent have rifles. newspaper readership on a daily basis has declined to about six people in 10 (57 percent in the compendium ; 73 percent in the attitudes sourcebook ). approval for women working is high (72 percent) and growing. there is a higher rate of approval for abortion than in the 1960s, but in the 1970s it leveled off. support for busing is limited (20 percent) but on the increase. in a paper on liberalism based on the gss, smith (1979) finds a general growth in liberal attitudes through the 1970s, with a slight leveling off later in the decade. (a similar rising and leveling trend appears in the research of beniger et al. on abortion.) in the aggregate the norc results indicate that about 56 percent of all items have shown change over time, and 44 percent have been approximately stable. the most change is seen in social attitudes. methods statistical methods for analyzing trends made of a limited number of points over time are not necessarily familiar to potential users of these studies--e.g. , to journalists, policy makers, and some social scientists. while apparently simple, the aggregation of survey data points into time series does not involve a simple analytic process. many groupings of points move up and down over time, and in graphic presentation, may appear to have a slope or shape. but it is not clear when these are indications of random oscillation, a linear trend, or a cyclical trend. when does a pattern actually represent change and when is it merely perturbations around a horizontal line? the rather large proportion of stability found in the fjorc analysis suggests that a null hypothesis of "no change" should be the initial assumption. a simple approach for determining trends from multiple points is to regress against time (weighting by the square root of the sample size to compensate for different numbers of cases). if the slope is significant, then the 45 sign indicates the direction of the trend. if not, no change is indicated. a nonlinear fit, such as a quadratic, may also be tested. the compendium includes another procedure for patterning points as trends and applies it to each of the series in the study. based on goodness of fit models, the procedure evaluates constancy, linear or nonlinear change, and indeterminate trends. the model determined for the series is stated at the bottom of each compendium table. the compendium and taylor (1975) present somewhat unclear explanations of the procedures. questions and concerns the three reports provide valuable information for the sophisticated user, but each has its own problems. first, no explanation is given of why the elec tion sourcebook does not begin with the 1948 election study. as variables in one sourcebook are sometimes excluded, sometimes included in the other, crossreferences between the two would have greatly enhanced their complementarity. the attitudes sourcebook includes in the same time series studies such as americans view their mental health (1957), based on a sample of the entire population, and the panel study of income dynamics (1968), based only on household heads. this shift lays interpretations open to question. in addition, it would have been helpful if actual election results were included in the election sourcebook for comparison. for instance, the number and proportions of registered voters (the eligible electorate) and voter turnout could easily have been incorporated. (see kelley and mirer, 1974:584-5, and tufte, 1977:310, on "overreporting" voting.) \/alidating census figures and current population survey data would have been welcome complements to the survey items. the sourcebooks fail to mention that the data may be obtained from icpsr for secondary analysis. in particular, no mention is made that each year's election data and the individual social attitudes studies listed in the icpsr guide to resources and se rvices are presently available for use. somewhat minor but irritating is the bulky form of the sourcebooks, whose pages will tear easily. it would have been better to put them into standard size books, like the norc study, perhaps at a lower cost. as the sourcebooks mention, there is significant cost in retrieving information, and this should be a strong incentive for further study of both the material in the sourcebooks and the series items retrieved but not reported in the published volumes. a system of identification and retrieval of items should be developed for icpsr, norc, and roper to assist in secondary research and to avoid the huge amount of time that went into producing the sourcebooks. it is also important to develop a series of articles which explain clearly the various methods of analyzing trends with limited numbers of points. the compendium series and analytic model should of course be made widely available. in sum, these books yield insights into social attitudes and behaviors based on a vast store of data. they begin to illuminate some significant social trends during the last 30 years. over the next few years there will be other presentations of trend data, as well as additional information on methods for analysis, and they can build upon the successes and failings of these timely reports. 46 references beniger, james, susan watkins and juan ruz. "trends in the abortion issue as measured by events, media coverage, and public opinion indicators." proceedings of the social statistics section, american statistical association 13, 1973, 118-123. campbell, angus, philip e. converse, warren miller, and donald stokes. the american voter . new york: wiley, 1960. converse, philip e., wendy j. hoag, and william h. mcgee. american social attitudes data sourcebook, 1947-1978 . cambridge: harvard university press, 1980. 441 pp. 520^ davis, james a., et al_. general social surveys, 1972-1978, cumulative code book . national opinion research center, 1978. davis, james a., ed. studies of social change since 1948 . 2 vols. norc report 127a. chicago, 1976. kelley, stanley, and than mirer. "the simple act of voting," american politi cal science review 68:2 (june 1974). miller, warren, arthur h. miller, and edward j. schneider. american national election studies data sourcebook, 1952-78 . cambridge: harvard university press, t980^ 388 pp. j20^ smith, tom w. , with guy j. rich. a compendium of trends on general social survey questions . chicago: norc, 1980^ norc report 129. 260 pp. $7.50. smith, tom w. "general liberalism and social change in post world war ii america a summary of trends." gss technical report 16, chicago, november 1979. "happiness: time trends, seasonal variations, inter-survey differences and other mysteries." social psychology quarterly 42 (march 1979). "sex and the gss." gss technical report 17, chicago, september 1979. taylor, d. garth. "procedures for evaluating trends in qualitative indicators.' in davis, 1976, v. 1. tufte, edward. "political statistics for the united states: observations on some major data sources." american political science review 71 (march 1977). lassist memberships lassist offers two types of memberships: regular, at $20 per calendar year, and student, at sio per calendar year. both include a subscription to the newsletter, a subscription to ss data, and special rates on other lassist publications. institutions such as libraries may subscribe to the newsletter alone for $35 per volume. further information on membership and on lassist' s current activities is available from the regional secretaries, listed opposite. members and subscribers are asked to remit payments directly to the treasurer, also listed there, by cheque, money order, or bank draft payable to lassist in u.s. or canadian funds. 47 alternative databases: the research resource division for refugees byles teichroew ' research associate research resource division for refugees carleton university alternative databases: the research resource divisionfor refugees part of the centre for immigration and ethnocultural studies, the research resource division for refugees (rrdr) is located at carleton university in ottawa, ontario, canada. it was established in 1985 to serve as an international archive and data collection agency for scholarly, governmental and field information on refugee resetdement and adaptation. in addition rrdr publishes a quarterly newsletter inscan (international settlement canada), each issue being devoted to a specific topic on refugee resettlement. rrdr also is actively involved in research on resettlement. to date three studies have been completed: the southeast asian refugee study a report on the three year study of the social and economic adaptation of southeast asian refugees to life in canada (1981-1983); the settlement of ethiopian refugees in toronto (1989); and, the settlement of salvadoran refugees in ottawa and toronto (1989). the resource division also publishes a series of working papers in immigration and ethnocultural studies addressing critical issues in the area. examples include: immigration and visible minorities in the year 2001 a projection by dr. john sammuel and the mosaic a generation later issues and trends by dr. frank vallee. the intent and form of rrdr's activities is largely a product of one fundamental characteristic of the field of refugee studies: namely, its state of considerable and continual development and flux. the composition and characteristics of groups of refugees and refugee claimants can change quickly, placing new and unforseen demands upon involved governmental, inter-governmental and non-governmental organizations, immigrant and refugee service centres, and sponsors. what these organizations and individuals require in order to provide services in an appropriate and culturally-sensitive manner is information; but this information needs to be produced quickly, and it should also be readily accessible. as it takes up to two years for information to begin circulating through traditional academic venues, so-called "gray zone" publications offer the quickest source of up-to-date information in the field. as a result, the holdings of rrdr are predominantly but not solely comprised of "gray zone" publications. our holdings also include a limited number of central, but specialized academic sources in the field. in addition, each of our documents has been entered into one of three on-line textual bibliographic databases in order to further enhance speed and ease of access to the required information. "gray zone" publications are produced by a wide range of organizations, including various government departments, inter-govemmental organizations, non-governmental organizations, ethnocultural associations, and other interested individuals, including some academics. their form is also diverse, including newsletters and periodicals, research monographs, field reports, occasional papers, conference proceedings, pamphlets, posters, and non-print forms, such as videos, films and photographic slides. there is a considerable breadth and diversity in the range of topics which one may find published in the field of refugee studies. however, rrdr has from the outset focused upon issues of resettlement as they relate to third-world refugees. areas of particular interest include the situation and experiences of refugee women, the health and mental health of resettled refugees, and the cultural background of refugee groups. however, numerous faculty members at carleton university can be consulted depending upon the area of research interest. we have made a conscious decision not to systematically collect documents relating to issues of human rights and refugee law, and we also limit the information obtained regarding the political conditions in refugee-producing countries which underlie refugee flight. the primary reason for not focusing upon these two areas is economic: collecting documents in these areas would require financial and personnel resources which we do not possess. fortunately, two years ago the federal government of canada established an organization whose mandate was to act as a documentation centre for information on issues of human rights and political conditions in countries of origin. the immigration and refugee board (irb), through its documentation centre (irbdc) in ottawa, canada and five regional offices, 2 has rapidly obtained a rather diverse and comprehensive collection of documents on these issues and others which are readily accessible to government officials, academics and other interested parties. to our knowledge they are the only government internationally to have invested the resources to establish such a collection of iassist quarterly documents. the resources in the irbdc are complementary to our own: indeed, we often exchange information. while rrdr purchases a portion of its holdings, many are obtained at no cost from governmental, intergovernmental and non-governmental organizations, ethnic community groups and researchers. rrdr also obtains a significant number of documents through exchange for our newsletter. we have explicitly sought to avoid obtaining documents which are already held at the carleton university library on-campus; fortunately, owing to our focus on gray-market documents there is little over-lap between the library holdings and our own. however, the library holdings provide easy and rapid online access to academic sources in the field as well as some government documents thereby extending the range of materials available to researchers in our office. in addition, a number of national libraries exist in the region, enabling rrdr researchers to search for (online) and obtain documents from a very wide range of sources. and, while we currently do not have any statistical data sets in our holdings, we would welcome the contribution of such data by any researchers, agencies or governments to our organization. although the holdings in rrdr are unique that is, textbased, predominantly "gray-market" publications on the resettlement of third-world refugees the format of our on-line records share features similar to those found in many other refugee documentation centres worldwide. this is because we exchange information with participating members of the international refugee documentation network (irdn) as well as other organizations and agencies. the format for our on-line data follows closely the convention established by huridocs. although initially intended for documents on the issue of human rights, huridocs has been adapted for use in the field of refugee studies, and also to meet our particular local needs. this format is somewhat different than that used in most libraries, owing partly to the different nature of our respective documents. a number of the participating members of the irdn have adopted the huridocs format, also with minor modifications. in numerical terms, our holdings currently comprise more than 6,000 items (i.e. publications) which have been entered into a number of on-line searchable (textbased) bibliographic databases. one of the databases focuses on the condition and experiences of women refugees in countries of origin, countries of first asylum and countries of resettlement. this bibliography is in the final stages of editing and will be published shortly. a second database contains items relating to the physical and mental health of refugees. a third database covers more general issues of refugee resettlement and integration, including economic, cultural, linguistic, and civic and social welfare facets, among others. ideally, each of the sources in the databases will be fully abstracted and keyworded using the international thesaurus of refugee terminology (developed by the international refugee documentation network, 1989). this should enhance the utility of the databases as the entry of standardized keywords will increase the speed and precision of literature searches. however, owing to limited financial and personnel networks only the bibliography on refugee women has been fully abstracted and keyworded. our holdings include items published by international and national non-govemmental organizations, governmental and inter-governmental bodies, local ethnocultural communities, research institutes, and interested individuals. examples of our international periodicals from nongovernmental organizations include icva news . published by the international council of voluntary agencies, the international catholic migration commission newsletter , by the international catholic migration commission, and refugee participation network , by the refugee studies programme, oxford university. refugees , from the united nations high commissioner for refugees, and monthly dispatch , published by the intergovernmental committee for migration are two instances of our inter-govemmental periodicals holdings. national periodicals include the canadian ethnocultural council , by the canadian ethnocultural council, update , by the u.s. catholic conference, migration and refugee services, and nv fremtid . by the norwegian refugee council. immigrant women of pel by the immigrant women's group in prince edward island, canada, is one example of more local periodicals from non-governmental organizations. our holdings also include a number of newsletters and periodicals from ethnic associations. our holdings of reports, field studies, occasional papers and policy analyses are as diverse. instances of documents from national governments include the hmong resettlement study , published by the u.s. department of health and human services, please listen to what i'm not saving: a report on the survey of settlement experiences of indochinese refugees. 1978-1980 . by the australian government, and a follow-up of the conditions of unaccompanied minors and handicapped persons among the refugees from vietnam resettled in sweden , by the swedish national board of health and welfare. our holdings also include provincial and state documents such as the training needs of settlement service workers , by the ministry of citizenship and culture, government of ontario, and documents from local governments, like the publication information for young refugee parents from the fresno county department of social services. our holdings also include a number of publications from intergovernmental organizafall/winter 1990 tions such as working with refugees in somalia towards a development perspective: a technical cooperation report, by the international labour office, and numerous reports from various united nations bodies, including violence against the vietnamese boat refugees: an assessment of needs and services from the united nations high commissioner for refugees. examples of documents from national non-governmental organizations include uprooted angolans , by the u.s. committee for refugees and voluntary repatriation programmes for african refugees: a critical examination , published by the british refugee council. making it on their own: from refugee sponsorship to selfsufficiencv . by the church world service, and helping refugee women help themselves: ywca's response are instances of publications from international nongovernmental organizations. zations, government departments, and assorted other nongovernmental organizations. 1 presented at the iassist 90 conference held in poughkeepsie, n.y. may 30 june 2, 1990. 2 these regional offices are in vancouver, calgary, winnipeg, toronto and montreal. 3 our fax number is (613)788-3676. inquiries can also be sent by electronic mail to "neuwirth@carleton.ca" and "rrdr@carleton.ca" . while one could provide many more instances of our publications holdings, the preceding recitation should have provided some indication of the range of documents inrrdr. one of our mandates is to make this information widely available to persons and organizations working in the area of refugee resettlement; we often receive requests from organizations world-wide for information in the field. the databases make it relatively easy to process these information requests and, following an on-line search, to forward our findings. in cases where the individual or organization would like to acquire specific documents in our holdings we may, with the permission of the authors, reproduce and forward copies of the publications. in events where this is not possible or feasible, we refer requests to the original publisher. unfortunately, owing to limited computer and software facilities in rrdr, the databases are not currendy accessible direcdy from terminals outside of the office and off-campus. on-line access to a read-only copy of the databases is planned for the near future, pending the availability of resources. requests for information are received and dealt with at rrdr through a number of routes. many information requests are forwarded to us by phone, fax and electronic-mail. 3 the simplest of these requests can often be answered over the phone. more complex information requests, or those requiring either a bibliographic listing of publications or the publications themselves, involve sending the information to the requesters by fax, electronic-mail, or through the postal service. a number of requesters also conduct research on rrdr premises. users ofrrdr facilities include students, academic and other researchers, refugee and immigrant service organiiassist quarterly lassist newslettert vol. 3« ^o. t (fall 1979) data archives retrospective amd perspectiven. k. nijhawan it is sometimes useful* in order to gain an understanding of the problems and opporunities for data archival efforts* to review the experiences of others. especially the problems encountered in establishing archival activities in "third world" nations may be illuminating for others* especially those used to the general level of technological deployment in the united states and europe. n . k. nijhawan has provided a description of current machine readable data archival efforts in inaia. editor. the out line counci l for the ence da attempt work do makes s the dev during years, deta -e 1 ssues f data ar factors establi institu during will be obj e the of dev ta s t ne s ome elop the b d th chiv th shtne tlon the bri ct pro sod elop arch exa fa tent ment cou ut disc e me es 1 at nt s a pas ef ly of t gram al ment ives mine r in ativ of r se bef uss i anin n ge ha v and ii t tw dea his pap me of t science of so in in criti this r e prop this of the re st on g and p neral a e led growth over t decad it with er he i res cial di a* call egar osa l prog nex arti f urpo s we to of he es is to ndlan earc h sc 1 it y the d and s for r amme t few ng a these se of u as the such world or so in the most general sense* a data archive is a library of data. like libraries of books* which concentrate on the acquisition and cataloguing of books in order to make them accessible to the academic community* these institutions* known variously as data archives* data banks* data resources' libraries* etc.* have been set up for the acquisition* or ga data need nity ever arch ma t 1 duri 1960 aval the n1ty gain deca usa r ap i comp and in s tate ot he data tist grow 1nst pres of t niza to s of . u * ives on ng s w labl sod ed n des and d de ut er use ocia d by r . res s ha th itut umab his tlon meet the nlike the ism of d the hen e to al sc thes oment or so ues velop tech of q i sci thi furt ource s al and 1 ons. progr and res soc usu ori g ore at a late the a s dis earc ial al i in rece arc 195 co izab 1 ence um d « p tern ment nolo uant ence s te her* s am so c deve t inf i amme inst urln arti eur s in 9/ o itat re chno the ong ont r lopm hese uenc in semi h an sci e ibra f t nt. hive os mput le esea 1tut g th cula ope* th n th 1 ve sear logy urg soc ibut ent dev ed indl nat 1 d tr nee r 1 es hese th s s and e r port rch 1 ons e pa rly d e fi e on tech ch f » e 1 lal ed of elop sett on of a i n 1 ng commu* howdata e f rtart ed early became ion of commuha ve st two in the ue to eld of e hand ni ques addon the share sdento the these ment s * ing up reprinted from icssr newsletter volume viii (3 & 4)* october, 1977 march, 1978* pp. 1-7* by permission of n. k. nijhawan. background the v.k.r.v. rao commltte which recommended the establishment of the council also made a recommendation that one of the functions of the indian council of social science research should be "to develop and supoort centres for maintenance ana supply of data." in pursuance lassist newsletter* vol. 3, no. 't (fau 1979) of th set u insti evolv progr group 1970 datio which data the i few count nated these by archl off 1c in 19 is r p a tuti 1 ng amme sub and ns « con cel cssr se le ry ne rec the ves e of 73. e c om wor ona l an e in mitt made th c e rn i" an ctea for twor omme cou was the menda t 1 on » king g roup a rrange ffectlve d the cou ed its rep a number e most 1 ed the set 1 n the se d the ass inst 1tut1 developing k of dat ndat 1 ons w nc 1 1 and estabus counc 1 l the to ment s at a a nt ry . ort 1 of re mport ting creta istin ons a a ar ere a the hed 1n ne counc 1 i suggest for r c h 1 va l the n march c ommenant of up of a riat of g of a in the coordi ch 1 ves • pp roved data in the w delhi and get better returns from this infrastructure. briefly stated* the council defined the functions of the data archives as follows: 1. to acquire* organize and maintain social science data sets 1n machinereadable form and make them available to interested social scientists for re-use* 2. to support suitable research institutions in different parts of the country to develop similar institution-based data archives* functions th sclen uncon t ions vent 1 so i. la posed semin for r tant be pe of th the must t ions of t insta train probl by ma equal of e adequ soc la of th that only of da tion f unct e in ce r vent 1 of 1 onal i i sci to a ate m e-use f unct r f orm e ics view assum 1 n v he i nce» ed m em of jor ly 1m asy ate i sci ese p the perf o ta ac but ions di an esea onal ts y« ence cqui achi ions ed sr. tha e s lew ndi a we anpo ace dat a port acce soft ence rob i data rm c qui s also to h co r ch vi data as m dat re* ne-r thes and by t th t th ome of t n hav wer . ess pro ant ss ware re ems* arc onve itio ta elp unc 1 1 has ew of arch i v ent lone a archi organi z eadab le e are v were p he dat e coun e data addit io he perc situati e a sh ther to dat due ing is the f comp facil search, it w hives nt ional n and ke up bulla a of s taken the es. d ab ve is e and dat a ery 1 ropos a arc cil w arc nal ul i ar on. ortag e is a pro agen que ut ing ities be as de shou i f unc di sse addit die oc 1 a i an f unc conove a supdisset s mpored to hives as of hives f unc ities for e of the duced c ies . st ion and for cause c i ded d not 1 1 ons m 1 na1 ona i nt e i e 3. to develop and maintain a computer programme library to handle the data management and retrieval problems of the data archives as well as data processing and analysis of problems of social scientists* 4. to organize training courses in survey research* and in data processing ana analysis* with emphasis on the use of computers and other mechanical devices and computer programme packages ; 5. to arrange for guidance and consultancy services for social scientists requiring assistance in data processing and analysis* 6. to act as liaison between official data producing agencies and social scientists in order to lassist mewsletter, vol. 3, no. *» (fall 1979) incorporate the research interests of social scientists in data gathering and tabulation plans of these agencies? and 7. to establish collaborative relations with social science data archives abroad. to this was adaed in 1976 the function of compiling a national register of social scientists in india to be revised per ioo i ca i ly . daia acquisition and oissemi nation the question of types of data to be acquired by the data archives and the sources of these data were first taken up. three main sources of data were identified: (a) data generated through the icssr-funded projects; (b) data generated by the official and semi-official agencies; end (c) data generated by scholars working in the research institutions and university departments. again* keeping in line with the general objectives of the councllt 1t was decided to acquire both survey as well as aggregate data relevant to the research needs of all the disciplines falling under the rubric of social sciences. however* in view of limited resources* both financial and otherwise* the council laid down certain priorities and gave the first priority to the acquisition of data sets generated by its own funded projects. the council being one of the largest single social science research funding bodies in the country haa at the time of the establishment of the data archives* funded trore than 300 projects of which nearly 5 per cent had been completec. since then* the number of the icssr funded projects have more than doubleo. it is* therefore* quite under s t anoab i e that the data archives established by the council would have a regular programme of acquiring and preserving data sets generated by its own funded or o ject s . during the past four to five years about 55 data sets have been acquired by the data archives. this also includes about ten data sets received from governmental agencies and other researchers who did not receive funds from the icssr. by any standard* this performance is not quite satisfactory though understandable. it should be appreciated that the pace of developments* which appear to have resulted in the growth of such an activity elsewhere* took a late start in india and has been rather slow. the tradition of quantitative research* based particularly on secondary data analysis* is relatively new. a very small number of researchers make use of the facilities of computer technology and other mechanical devices for recording* processing and analysis of data. the availability of the hardware and software facilities for social science data processing are quite inadequate. .hat is worse* even the meagre facilities available are not being fully utilized because the community of social scientists is not oriented to and trained in the use of these facilities. consequently* the majority of the scholars continue to analyse the data through hand tabulation and the raw data collected through the field sjrvey are rarely transferred on punched cares. therefore* only a limitec number of useful machine-readable data sets are available for acquisition. lassist newsletter, vol. 3» no. 't (fall 1979) general l tioned abo respons 1b i e acqulsl t ion case of da mental agen adequate re and otherwl buted to acquis it ion large numbe dest such survey org bank of ind tal agencle magnitude o data are social scl always aval ch they cou by soc i a i sc fore* extr impossible t t1on with meani ngf u i l and to ma interested y speak ve have for t to dat ta gene cles » sources set h the slo • for, r of g as t th anizat 1 1a« and s have f impor ext reme ence r lable 1 id be lentlst emel y d for an llmlte y orga ke th scholar i ng , i a been he slo a sets rated however f both ave als w pace over t overnme e natio on, th other genera tant d l y rel esea rch n a f easily s. it if f icul y s1ngl d res n1 ze t em ava s for r c t or pri w ra by g , i fin c of he y ntal nal e r gove ted ata< evan bu rm 1 ut is, t, e in ourc hese 1 lab e-us s menma r 11 y te of in the ve rnack of anc lal ont r1data ears a agensample e serve rnmena vast the t for t not n whl1 i ized t here1f not st itues to data le to c grad an c los data coun tack of lute ma jo cles nece ing f era simu sc le tlve shou the of t a p the thes inco of onsequ ua i ly all e coop gene try 1s le thi the op ly nee r offi the ssa ry their bly 1 itaneo nt 1 st s s of id wor nature hese d hased most e gro rporat social entl comi ut erat rati abs s ve inio essa clal urge ar ra own n ma us ly as va k t ata prog 1m ups e t he r scl ent 1 y« ng f ron ion ng olut x ing n t r y t dat nt n ngem dat chin t a wel r iou oget o lum and r amm port sh the t d th tal of t agenc e i y prob hat 1m a pro eed ent s a hoi e-rea grou 1 as s o her a e an help e f o ant ou id esear st s cou e v atta he 1 les esse lem . it i pres dud to m for ding dabl p o rep rgan nd d th in r a data als ch 1 in t ncil is lew that ck with mportant in the ntlal to it is s absos on all ng agenake the preservs , pree f o rm . f social resentai zat ions identify e format evolving cqu i r ing sets, o help nt erest s he data coll thes tion sigh ac qu usef of f the of t ever prog dire e c t i on e agenc s to t. t isition u i data actors, cont rol he acti , help ramme, ctly w1 and t abu l ie s . no these pr he entire and d 1 will dep not nee of the vi ties wh in st re direct i u be dis ation plans of immediate soluoblems are in programme of ssemlnatlon of end on a number essarlly within counc 11. some ich would, howngthening this y or not so cussed now. institution based £aia arghjves f coun izat in i f eas rese sity have host have the the fad in t in t info sy st i ar s tion tact appr less fund se le deve seal rom cil 1 on ndl a ible arch dep g of the pres abs litl he he m r mat emat out or s eci a r de s c ted lop the held of d is 1n artm ener the ne erva ence es a r oce a jor 1 on leal side bey f a ting cide ana ins dat very the v ata ar nelth for, stitut ents, ated se ins cessar 1 1 on of vast ss of ity of about ly av a pa ond t n ind thes d to ot he tituti a arch begl lew t chi va er n a lar ions ove r impor titut y f ac f the prop amoun deca the these ailab rt 1cu he pe 1 vi au e pro pr ov i r s ons t i ves nn 1 n hat i ac eces ge n and th tant ions hit se d er t of y an case dat le t lar rson al b lem de n erv i o h on g» cent 1 1v1 sary umbe uni e ye d do les ata. sto dat d de s, a is o s inst al s cho s , eces ces elp a s the ralties nor r of verar s, ata. not for in rage a is ath. even not choituconlar. the sary to them mall the inadequate financial resources of the icssr have been the main handicap for the development of this programme. these difficulties are likely to continue for quite some time to come. vast amount of financial resources would be needed to help these institutions purchase equipment needed for lassist newsletter, vol. 3, nio. t (fau 1979) the data organization and young teachers and ph.d. scholars dissemination. additional funds, in social science disciplines are on a regular basis would also be invited to participate in these necessa ry f or carrying out other courses of 3 t o a weeks duration, oata archival functions. it may these courses have proved to be not, therefore, oe possible for the quite useful and are expected to icssr to provide adequate funds for continue, setting up full-fledged institution based data archives. however, the council may consider providing at least some funds to selectea instigyj.danc: and consuklancy services tutions having a large volume of good data for upaating and organizquite closely associated with ing data and for preparing the the programme of training courses, necessary documentation. once this the data archives initiated in is done, one set of these materials 1974-75 a programme for providing may be made available to the icssr guidance and consultancy services data archives which will be made to the social scientists to tackle responsible for the data disseminatheir problems in data recoroing, tion task. gradually, these instiprocessing and analysis with the tutions can be developed to take help of mechanical devices. in over all the data archival funcorder to make these services ava1ltions. able to social scientists near their normal places of work, a number of institutions have been involved in this programme. curleiililng cc2ueses rently, these facilities are available, besides at the icssr data as statea earlier, the overall archives through (1) the indian success of the data arthival proinstitute of management, calcutta; gramme is closely llnkeo with the (2) the centre for the study of growth of quantitative techniques developing societies, delhi; (3) in social science research in genthe sardar patel institute of ecoeral and use of mechanical devices nomic and social research, ahmedain data processing in particular. daoi ('t) the centre for development these developments will generally studies, trivandrum; (5) the gokdepend upon the overall system of hale institute of politics and ecohigher education in the country. nomics, poonaj and (6) the tata the structure of university educainstitute of social sciences, bomt1on will take some time to change bay. the scheme is yet to gain and respond to these needs. in the momentum at all the centres, meantime, the icssr is trying to fill this void, to some extent, by organizing training courses in research -methodology and survey collibor atij^e ellailfihs kilt! fi*i* research techniques. keeping in archives abroad line with these objectives, the data archives decided to organize it would be readily admitted training courses in the application that the data archival programme of computer technology and other cannot be ceveloped in isolation of mechanical devices in social scithe developments in this field ence research with emphasis on use abroad. it has been, therefore, of computer programme packages in decided that the data archives oata processing and analysis. 10 lassist newsletter, vol. 3, mo. ^ (fall 1979) should maintain collaborative publish uddatea information periodrelations with similar data ically. archives abroad for the exchange of information* data» software and even data archival staff. for various reasons, not much progress has data inf^rmation services been made in this field. all these functions discussed here are quite important and are proposed to be continued and natio nal be'^^ster of sq£i^l further strengthened and expanded scien tists during the course of the next few years. however* these alone will in the beginning of 1976» the not be enough. the scope of the data archives took up the task of data archives needs to be further co«p1l1ng a national register of broadened for providing better sersoclal scientists in india. it was vices. systematic steps are necesalso decided that this register sary to build a data information should be revised periodically. system so that it can provide this attempt is aimed at filling referral services in the sources of the void in the area of basic and social science data in addition to comprehensive information about the the existing programme of physical background* research interests and acoulsition and dissemination of contributions made by the social important data sets. making relesclence community in india. inivant * clean and properly documented tlally* it was proposed to include data available to the interested in this register all the social scholar is an important task no scientists in university departdoubt. and* no less important is ments* colleges* research instituit to provide information on the tlons* governmental organization. availability of relevant data to a and private industry ana business scholar* even if the data archives and to cover anthropology* commight not be in a position to sermerce* demography* economics* eduvice that particular data to a cation* geography* history* interresearcher at a particular point of national relations* linguistics* time. management* political science* psychology* public administration* in order to systematically build sociology (including criminology)* a meaningful data base for these social work* communication (includservices two programmes are proing mass communication and journalposed to be initiated. first, we ism) and law. however* because of plan to prepare an inventory of practical difficulties the project current and recently completed had to be limited to the coverage social science researches in india of university departments* colleges and keep this information up to and research institutions. the date. specifically* we intend to first volume of this register comcollect information on the status prising about 7*000 social scienof current researches* type of data tists* information regarding whom utilized* method of data collecwas collected through a mailed tion* processing and analysis, questionnaire (covering the period efforts will be made to cover all up to december 1977) will be ready researches whether funded by the for publication soon. it is proicssr or not. second* we propose posed to keep this information up to initiate a series of projects' to date* extend the coverage and lassist newsletter, vol. 3» no. t (fall 1979) for preparing inventories on the initially, the data archives had type of data, periodicity of coldecided to develop this software on lection publication of data and its own. however, not much coulo unit of observation, etc., of be achieved in this regard for varsoclal science data generated by ious reasons. the most important official agencies. these two basic reasons being non-availability of sources will throw open such a vast adequate funds and skilled manreservoir of information that when power. even if the funds were to properly classified and organized, be made available, the problem of a would prove tremendously useful in non-ava1 i abl 1 1 1 y of properly discharging this function. moretrained programmers, who could overt this system would also help meaningfully interact with social the data archives in identifying scientists, understand their important data sets at the approrequirements and develop the softprlate time and facilitate their ware accordingly « will continue to acquisition. persist. therefore, this programme will have to be developed over a long time. in the meanwhile, a beginning in this direction could b^y^^opm^nt 2f sseiuare facitixx£§ be made by preparing an inventory of existing social science computer as has been emphasized earlier, programmes 1n the country and makall the data archives are supposed 1ng this information available to to acquire machine-readable data the interested scholars. such an and to organize them in such a form inventory would, decidedly, idenas would make the data retrieval tify gaps in the availability of and dissemination least cumbersone these facilities, both cross-secand time consuming. in other tlonally and in different problem words, data archives would generareas. once such information ally acquire raw data transcribed beccnes available, the gaps that either on punched cards, magnetic may exist, could be plugged in a tapes or any other mechanical devphased programme under given priorice. at the archives, the data ities. normally have to pass certain 'acid tests', such as checks for wild in sum, 1t may be recapitulated codes, inconsistencies in coding that the initial period has been a punching, standardization of data period of mixed experience. the formats and code categories, etc., achievements during this period, before they become ready for dissehowever, outweigh failures despite mlnation. for all these tasks, a no prior experience in this field data archive needs relevant compuin the country. this period has ter programmes (software). this been quite significant in many software is used for data cleaning, ways. during the period the necestransformation, organization and sary 1nf ras tuct ur e has been built, retrieval.' this type of software some of the programmes put on the is, therefore, necessary for the ground and sufficient experience icssr data archives. in addition, gained to help take a leap forward the data archives has to develop during the course of the next few and maintain another type of softyears, ware if it proposes to cater to the statistical data processing needs however, it may be reiterated of the social science community in that the success of the data archithe country, val programme in the country would depend on the proper development of 12 i assist news let t er , vol no. t (fall 1979 ) var i prop st re vent sage all the for gram t1ta in s deve come the in ment of t sure is 1 ous er s ng t he 1 onal d ea this sing the me is tiwe oc ial lopme abou uni ve the c s» s he da to c n ant functions out teps are n n these func or unconvent rlier or new is important le most impo development an extensive techniques a science rese nts are prima t through th rs i t y system ountry. th important f ta archives* ome about* icipation of i ined e c es s t i ons i ona i i y p rtant of th use nd au arch, rily e ch of e ese or t h are in ef thes above . a ry to : con« env i roposed . and y e t .t factor is proof quant omat ion these going to anges in ducat ion develope growth slow but feet* it e developments init iate gramme. quicken this pr initiate dance an training and res newly e the inve and so hopefull to this the leas the soci will det of devel india. that the c the dat however the pace ogrammet d its prog d consul t a courses i earch me nvisaged p nt ory of f t ware fa y » give a p rogr amme • t « it is t al science ermine t h opment of of d the r amm ncy n da thod rogr cur r cili ddit l he c cu e ra this i i dec r c h i va in or eve i op counc est i i servi ta pro ology . ammes « ent r ties i ona i ast « oopera mmuni t t e an progr ided to i proder to ment of i i has ke gu i ces and cess i ng the like esear ch would* support but not t ion of y which d level amme in 13 indexing machine -readable data files for a social science data archives by jacqueline mcgee rand corporation "it is still true that the best retrieval system is the expert human mind. " this paper was presented at the lassist annual conference, may 19-22, 1983, in philadelphia, pennsylvania. introduction in the recent past much has been written and discussed about the problems of cataloging and bibliographic control of social science data. many of these problems may have been resolved with the implementation of the angle-american cataloging rules ii, chapter 9 (aacrii) and the marc format for bibliographic control (2). however, there are a number of reasons these solutions may not yet be universally implemented. for instance, the aacrii and the marc format may be very familiar to library staff, but all archives are not staffed by librarians. many archives are suffering from a shortage of staff and financial resources. federal agencies produce a major portion of the data archived and used for secondary analysis and these agencies are also financially depressed. researchers and programmers who use these machine-readable data files (mrdf) are not as aware of the problems related to the acquisition or storage of data and their interests do not necessarily correspond to the interests of the data archivist. technological changes occur so frequently procedures may become obsolete by the time implementation occurs. and finally, so many new commercial firms are installing social science numeric data bases online and the interests of these firms do not lie in the same directions as those of the data archivist. to assist the novice who may be overwhelmed by some of these problems, it is the hope of the author this paper will provide some examples of simple record keeping. rowe and byrum previously described a user-oriented system for the documentation and control of mrdf (3). this system was comprised of four parts. first, a standard catalog entry, second, a data abstract or description form, third, documentation codebooks and lastly, the records of physical and logical characteristics of the data set. it is not the purpose of this paper to offer an alternative system for the documentation and control of mrdf but to provide a practical example of implementing such a system. this example will provide an illustration for the person who has just received responsibility for the safekeeping of a collection of mrdf or to establish an archive and isn't sure where to begin. the first item in the system by rowe and byrum, the standard catalog entry, was described before the anglo-american catalog rules ii, chapter 9 were implemented. data librarians located in a traditional library are already familiar with the rules for cataloging, but may not be familiar with the aacrii, chapter 9. the anglo-american cataloging rules ii, chapter 9 describes the standard rules for cataloging mrdf. it is not within the scope of this paper to argue the pros and cons of the acceptability of the aacrii. there can be no doubt that a uniform standard defining mrdf is necessary in order to alleviate present confusing practices and the proliferation of titles for one data file. certainly implementation of the aacrii and the agreement of the marc format were giant strides in the cataloging and bibliographic control of mrdf. standard catalog entry rowe and byrum state "standard catalog entries, constitute the primary records by which computerreadable data files should.be controlled and accessed." it is with this one area of their discussion that i disagree slightly. the standard catalog entry requires extensive staff time and financial resources and need only be considered as necessary under certain conditions; if the required resources are available and may be allocated to such an endeavor; if the data archive or data bank is situated in a library or a library is available and willing to participate; if the data holdings are original data from the institution responsible for the establishment of the archive. it is hoped non-originating archived data will be catalogued by the originating institution. however, the federal government is responsible for a major portion of the data files held in many archives and current fiscal restraints on most federal agencies probably will not permit such a project in the near future. there is, however, an ongoing cataloging project at michigan's inter-university consortium for political and social research (icpsb) which may resolve the problem of cataloging federal data (4). icpsr is certainly one of the largest, if not the largest, of the data archives in the united states. when this project to catalog their holdings is complete, it may be possible to consider a union catalog. data abstract or data description form it is the data abstract or data description form described by rowe and byrum which should be given priority in the development of an archive record system. the data abstract or data description form in a standard format is an absolute necessity and should be the core of the documentation for the archiving of mrdf. aldrich has proposed a similar abstract form for the documentation of federal mrdf (5). it was from a description by aldrich the following example was derived. changes made in the form were for the benefit of the user and do not reflect a disagreement with her proposed standards. since this form includes an abstract summarizing the data set or file being archived, this document shall be referred to here as a data base profile. with some slight variations, the items contained in the form generally should contain the items shown on the next page. if it is not possible or feasible for the archive or library to catalog the holdings of the archive according to aacrii at least the information supplied in the abstract or data base profile form will conform to the standards for describing mrdf. if at some future time cataloging is possible, the information for the catalog abstract form for the docwnentation of federal medf file #: an identifying number for the individual archive file name: a title file source: producer, distributor, processor principle investigator: primary researcher type of file: survey data, microdata, administrative records, process records, geographic records, software universe: total universe the records describe sample size: number of observations or records sample unit: household, person or unit being measured restrictions: none, or any restrictions placed on the distribution abstract: a summary description of the data set or file. each abstract held in the archive should contain the same information. care should be taken not to omit any portion of the required information and the information should appear as closely as possible in the same paragraph. this assures an easier search for the individual . looking for specific information as to size of the sample, purpose of the study, key variables, etc. references: a descriptive listing of the hard-copy documentation available for use with the data file, i.e., codebooks, survey instruments, dictionaries, etc. related printed reports: known reports where the data is described or where the data file has been used tape specifications: the physical characteristics of the data file, the tape numbers, the logical record length, the blocksize, data set names and density entry will be readily available. copies of the data base profile may be stored in computer format as well as in hard-copy. if the data is stored in computer format it would be possible to devise a simple online search capability. if the data librarian wishes to be bibliographically correct, study the aacrii and include in the profile the pertinent information from the aacrii as well as the information required or deemed necessary for the institution housing the data library (7). many of the elements of the data base profile may be utilized to produce a catalog. for instance, each abstract when extracted from the profile provides summaries of the archive holdings. the rand data facility catalog uses the abstracts in such a way. each abstract then written, therefore, should include the following (6): • data base identification number • source • name • date the information was collected • subject • geographic level of the data (lowest) • population or sampling unit • number of observations or number of logical records • key variables indicies may then be derived from the information given on a data base prof i te . source and name index the source and name index is derived from the file name and the file source as given in the data base profile. often these items are sorted as separate indices; an author index and a title index. an index should lead a user to the information he is seeking with as little effort as possible, and so we have combined these indices. on the data base profile and in data file records we use the most correct name for a data file. the correct name may be drived by using the aacrii rules. since we also wish our individual indices to assist the user in his search we also include in our local archive index those aliases or acronyms when they are commonly used, if the index is to prove useful. in order not to have a great many "see ." included in the index where aliases or acronyms or common usage names are listed, the correct identifying number of a particular data file is used as the pointer. keyword index this index is certainly one of the most difficult to construct. a thesaurus would be helpful; however, the keyword index discussed here was developed from individual data files. as mentioned earlier, one of the manditory sections of the data base profile is a list of key variables. at the time of archiving a new data file, the abstract is written and the key variables extracted. these variables or subject categories are added to the end of the keyword index. a copy of the keyword index is kept online. new keywords are added at the end of the old index and a short sas program sorts the keyword index by words or by identifying file numbers whenever necessary. geographic levels and major subject variables in the keyword index using census bureau designations, the lowest geographic level of the data is assigned as a second element of the keyword index. by using the lowest geographic level it is possible when searching for data to weed out those data files not useful to the researcher. a third element for the keyword index is a prescribed list of major subject variables. for each data file the keyword index will include at least one, but not more than three major subject categories. it is then possible to produce tables listing data files by geographic areas and major subject categories. librarians may want to use the library of congress subject headings (lcsh) list. the data base profiles may be produced as printed copies to be given to researchers interested in using a particular data file and can be used as documentation for bibliographic citations in research papers for the creating of a catalog of holdings. the data base profiles stored in a partitioned data set at rand are soon to be converted to a total online system using the ibm info/sys csd data retrieval system. the oz info/sys has the capability of handling multiple data bases and has been utilized at rand recently to create an information system for a library of software models as used in one department (8). with the data stored in oz it will be possible to search the system for a particular data base by title, by keword, or keyword combinations. requesting data bases by keywords produces a "hit list" of all data bases that contain the requested keywords. the hit list can then be accessed in either of two forms: one is an abreviated form that lists only a few specifics about the data bases on the list (including the data base number and title); the other is a full description of the data base. it is possible to get print copies of an oz screen, of the individual data base descriptions and of the hit list of keywords. references (1) michael phillipson, "the classification, storage and retrieval of survey data, reader in machine-readable social data,: h. white, ed. , information handling services, englewood, colorado. (2) see sue dodd, "cataloging machine-readable data files: an interpretive manual," american library association, december, 1982. (3) john d. byrum, jr. and judith rowe, "an integrated, user-oriented system for the documentation and control of machine-readable data files," library resources & technical services 16(3), summer, 1972. (4) carolyn geda, "marc formal applied to machine-readable data files: a pilot project of icpsr," paper prepared for delivery at the annual meeting of the society of american archivists, september 1-4, 1981, berkeley, california. (5) barbara aldrich, "proposed standards for bibliographic entry and abstracts for federal machine-readable data files," redraft 2, october, 1978. (6) don trees, "the rand computation center: rand's data facility: a guide to resources and services, " bcc-1555/18, santa monica, april, 1981. (7) anglo american cataloging rules ii, chapter 9, machine-readable data files, p201-216. (8) as described by william fowler, the rand corporation, santa monica, california, 1983. 10 sist newsletter vol.1, no. 2 training seminars summer 1977 workshop on management, library control and use of computer readable information jiilv 25 august 5. 1977 this two-week workshop is designed to meet the needs of individuals whose responsibilities may include providing data services or information about computer-readable data files to users. the objective of the workshop is to introduce individuals to data management, data library and data servicing procedures and techniques employed at established data service centers. specific attention will be given to the practical aspects of making data available to users. the workshop contains two entry points contingent upon the background, experience and interests of the participant. the first week of the workshop will consider the process of collecting data, documenting data collections, and processing (cleaning) data for primary analysis and use or storage centrally for public access. hands-on experience with data will be provided at each step of the data cleaning process. computer experience is not required. the second week will focus on data library procedures, user services, and the administration and organization of data service centers. data library procedures will include acquisition of data, transfer of data, accessioning data and bibliographic control. it should be noted that an intensive format for this workshoo will be used. scheduled sessions will be held both in the morning and afternoon. additional sessions may be scheduled for the evenings as needed. enrollment will be limited. for further information and application forms, contact: summer program, icpsr, p. 0. box 1248, ann arbor, michigan 48106 313/764-2570. summer training seminar in cross-national data analysis the international social science council (issc), with the institute fur soziologie at the university of vienna, in cooperation with the institute for advanced studies in vienna, is sponsoring a summer training seminar in crossnational data analysis, from 3 to 22 july 1977. for a number of years the issc has conducted a program of advanced training in methods and techniques of comparative cross-national analysis. under a grant from the stiftung volkswagenwerk, summer schools were organized from 1972 to 1974, for training in the use of data from several countries for comparative analysis. this g-ant has been renewed and the stiftung has agreed to finance the preparation of a series of workbooks for training in comparative analysis . two draft workbooks for data packages will be available at the training seminar. 29. sist newsletter vol.1, no. 2 one of the workbooks will cover the time-budget data for cities in four countries: canada, france, hungary, and the united states. (the cross-national survey research project directed by the hungarian sociologist, alexander szalai, has resulted in an enormous quantity of data, almost all of them unavailable to data archivists; thus, this workbook represents an important contribution to the research community.) a comprehensive report on the time-budget study is the a. szalai edited book. the uses of time (the hague: mouton, 1972, $48). [see also newsletter , volume 1, number 19, p. 6, 7, of the institute of social research, university of michigan, for a discussion on some of the sexual inequalities in the use of time.]) the data sets and workbook texts are being nrepared by andrew harvey of dalhousie university, halifax, n.s., in cooperation with philio stone and alexander szalai of the karl marx university, budapest. the second workbook covers data on intergenerational mobility for five countries: austria, the federal republic of germany, the netherlands, the united kingdom, and the united states. [ the flyer describes "interaenerational mobility" as a classic field of comparative research, an area which the international sociological association launched in the early 1950s. "cross-national data on mobility are of particular interest because they allow precise testing of several alternative models: blau-duncan path analysis, boudon transition matrices." ] thomas herz of the zentralarchi v fur empirische sozialforschung cologne, in cooperation with raymond boudon of the universite rene descartes, paris iv, is preparing the data sets and workbook texts. herz will test the data package and workbook at the training seminar in vienna. the training seminar will be open to graduate students and younger staff members with some training in statistics and some knowledge of computer processing. the seminar will take place at the institut fur hohere studien, stumpergasse 56, vienna. the institute's univac 1106 computer will be available for the seminar. all data sets will be available in spss formats and will have been checked for processing on the computer. the grant from the stiftung volkswagenwerk to the issc covers tuition costs. participants must apply to their own national research councils or other funding agencies to defray the cost of travel and per diem. participants will be housed in a nearby studentenheim . the cost of single rooms with breakfast will be about $10 per night, for double rooms, about $18. meals in nearby restaurants cost between $1.50 and $3.00. the total cost of the 24 days in vienna will probably be at least $400. graduate students and others interested in taking part in the seminar should write to professor rolf ziegler, institut fur soziologie, alserstrasse 33, a-1080 vienna, austria. application forms should be returned immediately because the deadline 1s 31 march 1977. those selected for participation in the seminar will be informed by 1 may at the latest. payment for the first 10 nights must be sent in by the end of may. departments and institutes in eastern europe should be aware that efforts are being made to obtain a grant for stipends in hard currency from unesco for students from their region. 30. sist newsletter vol.1, no. 2 tenth essex summer school in social science data analysis and collection session 1 session 2 session 3 15th july to 29th july 1977 30th july to 12th august 1977 13th august to 26th august 1977 the european consortium for political research will be sponsoring the tenth school, to be held at the university of essex in three continuous but independent sessions from 15th july to 26th august. special emphasis will be on introductory courses for participants who lack any training in statistics or computing. the introductory level courses include "absolute beginners course," "introduction to data analysis (spss-based)," and "mathematics for social scientists." intermediate level courses include regression theory and applications, basic scaling, computer-based simulation, cluster analysis, factor analysis, and content analysis. advanced level courses include multi-dimensional scaling, network analysis, analysis of nominal-level data, analysis of social processes, and analysis of systems and hierarchies. session : 15th july to 12th august 1977 two courses will be offered in survey design and analysis. this special module is organized in conjunction with the ssrc survey archive and runs parallel with the other courses. it is primarily designed for introductory and intermediate level students. instruction will particularly emphasize the application of these techniques to data collections held by the school. full supporting interactive computer facilities will be available. financial support may be available to participants from their own institutions or national research councils. the organizations wish to encourage attendance by graduate students, research assistants and junior staff. interested persons should write to: the organizing secretary, tenth essex summer school, department of government, university of essex, colchester c04 3sq, england. data needs center for the study of youth development at boys town the center for the study of youth development at boys town, omaha, nebraska, is a nonprofit organization conducting social science research on youth and youth development. ed meyers writes to ask lassist members' assistance in locating information about data collections (catalogued or uncatalogued) which focus on youth, adolescents, and youth development. center researchers have asked for longitudinal and cross-sectional data relating to the development of educational and occupational aspirations and attainments and to the antecedents and consequences of age at first marriage. in addition, there is interest in locating data on socialization and diffusion 31. 36 iassist quarterly 2010 / 2011 iassist quarterly qualitative research has been developing only for twenty years in the czech republic abstract the archiving of social sciences data is not a widely adopted practice in the czech republic, mainly due to the lack of a national policy on data archiving and sharing and the lack of feasibility studies for qualitative archiving. researchers provide their data on a voluntary basis. concerning quantitative research, however, hundreds of projects have been carried out that represent a large share of research projects in the czech republic. the quantitative section of our archive is well-developed and serves hundreds of users, mainly students. yet, the situation in the field of qualitative research is completely different. archiving is not a component part of the research culture and this will hardly change in the near future. the practice of the use of informed consent (particularly in a standardized form allowing archiving) is far from being adopted. therefore, our archive’s qualitative data library has only a limited number of data files and, as a consequence, of users. keywords: archive, czech, qualitative data, research infrastructure existing qualitative archiving infrastructure the infrastructure consists solely of our archive, ‘medard’, which, as an independent data library, is a part of the sociological data archive, institute of sociology at the academy of sciences of the czech republic. at present, the archive provides a small number of data files (seven), consisting of several dozens of completed interviews. several other data files are currently in the process of being archived and will be made available in the near future. the data archive is used about ten times a year mainly by the students of social sciences and, rarely, also by researchers. the archive functions as an electronic library, and this is why the data are archived in a digital form. presently, there are transcriptions of interviews and also audio and video recordings deposited there. as a part of the institute of sociology, the archive is financed through its research arm. however, there is only a small budget covering half of one employee’s salary. the archive also provides extended electronic counseling in archiving qualitative and qualitative longitudinal social sciences data in the czech republic by tomáš čížek1 iassist quarterly 2010 / 2011 37 iassist quarterly czech related to qualitative research, as well as basic information in english, on its web pages (see http://medard.soc.cas.cz). qualitative research has been developing only for twenty years in the czech republic. to the best of the author’s knowledge, longitudinal research projects are not widely performed. in the czech republic, the only instances of such research were longitudinal documentary films that are available in the form of compact works of art with video recordings (dozens and even hundreds of hours of recorded life histories) but these are not available for studying purposes. however there is a tradition of oral historical work, notably as a study of life histories somehow related to significant events in czech history. the oral history center has already made several dozens of interviews available as transcriptions in the form of books as well as in an authentic audio format obtainable from the center (see http://www.coh.usd. cas.cz/). development planning the archive has existed already for many years. so far, however, it has been a rarely used infrastructure. it is necessary to extend and revise the practice of informed consent that would allow archiving. the author has been giving lectures at universities and conferences aimed at promoting and supporting the practices, yet the result is only limited. in the near future, we are planning to establish a method of data handling and archiving in cooperation with charles university in prague that will be binding for the students of sociology conducting qualitative research. the archive is also involved in the project “czech sociology 1945-1968 oral history“. interviews performed within it are archived in the form of interview transcriptions, audio and video recordings. considering the necessity of significantly extending the contents of the data archive, it will be necessary to perform an extensive study of qualitative research in the czech republic and, based on its findings, to identify the files suitable for archiving. the cooperation with cessda or iassist should, first of all, provide us with the opportunity to become acquainted with “good practice” at other workplaces and, in this sense, inspire our further work. this would mean, for example, support for study visits to established workplaces, or perhaps collaborative involvement in international projects. notes 1. tomáš čížek tomas.cizek@soc.cas.cz institute of sociology, academy of sciences of the czech republic the sociological data archive (sda) jilská 1, 110 00 praha 1prague, czech republic http://medard.soc.cas.cz vol252 iassist quarterly fall 2001 5 the counting california project (http://countingcalifornia.cdlib.org)1 makes available statistical information about the state of california, using data from california state and various federal agencies. the initial project design called for a metadatadriven system, based on the data documentation initiative (ddi) document type definition (dtd). the goal is to provide consistency in data discovery and data display to the end user, regardless of the structure of the underlying data. proposed extensions to the ddi developed by wendy thomas of the university of minnesota (http://www.socsci.umn.edu/ ~wlt/ddi) allow for the definition and documentation of statistical summary data in the form of multi-dimensional tables. the provision for these matrix variables allow for the creation of metadata, which not only characterizes the fundamental content of the statistical data, but also allows for flexibility in the choice of data storage structures and information display of the stored data. this paper describes how the counting california project team utilized the proposed ddi extensions to implement a data system designed to be flexible in implementation and consistent in presentation. output tables multi-dimensional data tables produced by the bureau of the census present certain challenges. specifically, the relationship between the variables and the tables must be maintained. the proposed ddi extensions address this problem with a method for defining output tables. this definition is in xml and is referred to, in this paper, as the varmtx (variable matrix). while we found the varmtx to be extremely useful in working with the multi-dimensional tables, it could also be applied to other types of data files. in addition, we found that making the table our primary output object addressed multiple project goals. tables and project goals by selecting the pre-defined table as our initial dynamic data display, we were able to address the following project goals: • maintain consistent granularity of titles for data discovery • maintain consistent data display • control user interaction with the server for server-side data delivery • assign subjects and keywords at a logical and useful level in addition, the use of the varmtx to define dynamic output tables provided the ability to treat online data objects with different properties in a consistent fashion. we gained: • a structured metadata record for searching • harmonized representations of data using ddi extensions as an intermediary for data storage and data display by patricia cruse*, marsha fanshier*, fredric geyt and margaret low*1 6 iassist quarterly fall 2001 • powerful and flexible metadata available to the data delivery system granularity of titles a primary goal of counting california is to provide searching of metadata for the purpose of data discovery. variable-level searching of both standard rectangular files and multi-dimensional files presented a granularity problem. a standard rectangular file might have one variable (column) for race with five values: ex. 1: one variable race white african american native american asian hispanic the multi-dimensional data file would have five columns: ex. 2: five columns racerace-african race-native raceracewhite american american asian hispanic a search of variables on ‘race’ would present one hit in example 1 and five hits in example 2. a cross-tab of race by sex would have ten columns for race and present ten hits. utilizing the varmtx definition for both file types gave us the means to harmonize our definitions and present consistent discovery behavior to our users. by implementing discovery at the table level, a search on ‘race’ would yield both types of tables, but return one hit (table title) for each table. data discovery is done at the table level, not the variable level. xml and xml tools since the metadata is in xml format, it provided an opportunity for the team to look at a variety of methods and tools for data discovery and display. the xml-tagged metadata was created in two different manners. for datasets with many tables (stf3, usa counties), the xml code was generated using perl scripts and then manually edited using a commercial xml editor. for the smaller datasets, the xml code was manually created. the relationship between the tables and appropriate subjects and keywords was determined and then added into each of the varmtx xml tables. once the xml-tagged metadata was generated for each study, perl scripts and xsl transformations (xslt) with style sheets were developed to transform the structure of the xml into the data needed for the data discovery, search and dynamic data display. iassist quarterly fall 2001 7 the metadata for a selected table, contained in a single, structured record in the metadata database, can be fed into multiple software packages, including an rdbms database, text search engine and analysis software. only the analysis software cares about the type and structure of the statistical data; all others rely only on the metadata. harmonized representations of data while the varmtx provides a way of defining the multi-dimensional tables, it also can be used to define tables derived from standard rectangular files. one property of the multi-dimensional tables is that the aggregated data comes in short/fat files. traditional rectangular files tend to be long and skinny. detailed, in the next page, is how we used the varmtx as a middle-tier to work with both types of data. 8 iassist quarterly fall 2001 a data request is routed to either the vertical or the horizontal handler, depending on the type of data. each type requires the same parameters. minimal parameters are necessary since the analysis software runs its own lookup on table elements. outputs created include a table and cached data. further operations, such as graphing, subsetting and exporting, utilize the cached data for input. results from both horizontal and vertical data are discussed below: horizontal (short/fat) data example areaname h0460001h0460002 h0460003h0460004h0460005 h0460006h0460007h0460008 h0460009h0460010h0460011 h0460012h0460013h0460014 alameda 7413 9971 39964 76247 42186 17473 3454 567 956 6405 11121 5378 1831 378 alpine 25 18 59 51 5 0 23 0 2 0 0 0 0 1 the data for the table ‘hispanic origins by gross rent’ comes from the 1990 stf3a and is pre-summarized for display. to reproduce this table, the system needs to maintain the pre-defined relationships between variables as well as complex header information. hispanic origin by gross rent specified renter-occupied housing units this table has two dimensions (hispanic origin and gross rent). the varmtx specifies, in compact format, the dimensions of the table, labels as well as cell quantity and location information. the full record for this example, is available at http:// countingcalifornia.cdlib.org/moreinfo/ddi_ext_app.txt. hispanic origin by gross rent nigrocinapsihfoton nigrocinapsih tnerhsachtiw tnerhsachtiw emanaera ssel naht 002$ 002$ ot 992 003$ ot 004$ 005$ ot 947$ 057$ ot 999$ 000,1$ ro erom on hsac tner ssel naht 002$ 002$ ot 992$ 003$ ot 994$ 005$ ot 947$ 057$ ot 999$ 000,1$ ro erom on hsac tner ademala ytnuoc 314,7 179,9 469,93 742,67 681,24 374,71 -4,3 45 765 659 504,6 121,11 873,5 138,1 873 enipla ytnuoc 52 81 95 15 5 0 32 0 2 0 0 0 0 1 iassist quarterly fall 2001 9 2 14 1 hispanic origin 2 2 gross rent 7 the variable knows its location within the table. in this example, it would be the first cell in both dimensions (coordinate1=1, coordinate2=1). 1 1 less than $200 properties � data is pre-summarized � data is layed out in horizontal, linear order functional requirements � select variables for table � roll-up headers � cohorts repeat consistently (e.g., variable values) � sub-headers are attached to specific cohorts data statements the sql command for extracting this data would be: select h040001..h0460014 from stf3data ; counting california uses the sas procedure proc report to roll-up headers. vertical (long/skinny) data example name native sex year age alameda 208 male 1970 04 alameda 214 male 1970 5-9 alameda 208 male 1970 10-14 alameda 251 male 1970 15-19 10 iassist quarterly fall 2001 alameda 372 male 1970 20-24 alameda 209 male 1970 25-29 alameda 171 male 1970 30-34 alameda 155 male 1970 35-39 alameda 129 male 1970 40-44 alameda 114 male 1970 45-49 alameda 86 male 1970 50-54 alameda 82 male 1970 55-59 alameda 80 male 1970 60-64 alameda 61 male 1970 65-69 alameda 33 male 1970 70-74 alameda 19 male 1970 75-79 alameda 24 male 1970 80-84 alameda 31 male 1970 85 and up vertical data appears to be more complex given that values must be applied to give meaning to the variables as well as the tabulations performed. it is actually a simpler process, since analysis software is designed to accommodate this format. naciremaevitan elam sega 4-0 9-5 41-01 91-51 42-02 92-52 43-03 93-53 44-04 94-54 45-05 95-55 46-06 96-56 47-07 &08 pu ademala 802 412 802 152 273 902 171 551 921 411 68 28 08 16 33 42 the varmtx for this type of data uses sql commands to subset the data. in the example below the universe is native american males. limiting to native american is a matter of variable selection since that summarization has already been completed. the data must be subset to limit to males. that is accomplished with an sql where statement. the parameters for the where statement are coded into the derivation element. 2 sex 1 males only (sex=’male’) iassist quarterly fall 2001 11 properties • data may or may not be pre-summarized • data is not cross-tabulated functional requirements • select variables for table • subset data • run tabulations/cross-tabulations • apply values to variables • roll-up headers data statements the sql command for extracting this data would be: select name, native, sex, year, age from race1970 where sex=’male’ ; counting california uses the sas procedure proc tabulate to accomplish functional requirements 3-5, listed above. generating tables without modifying data or programs the new u.s. census public law data has four tables that contain more than 70 cells each. an example of a label for one cell is: white; black or african american; american indian and american native; native hawaiian and other pacific islander; some other race. the display of one of these tables is 71 columns wide, an unfortunate size for printing. in order to present a more useful and readable display to the end-user, a new representation of the data was created, selecting fewer data elements. the example below, race [14], shows an additional table that was created by modifying the metadata, but without modifying either the data or the program. this method provided the opportunity to display multiple tables, utilizing the basic 71 variables, via the metadata-driven program. conclusion the original ddi dtd, as developed by an international working group, was envisioned as an archival documentation standard for statistical microdata datasets. the idea was to encompass the bibliographic description level as well as details of file structure and layout. the ddi extensions developed at the university of minnesota extend the variable definition section to allow the definition of the logical structure of multi-dimensional summary data often encountered in census files. our contribution has been to recognize that these extensions and the ddi in general can also be utilized as an active agent in controlling both data storage layout for statistical summary data and in providing flexible data display of multi-dimensional noitalupoplatot]41[ecar ecarenofonoitalupop emanaera latot latot etihw enola rokcalb nacirfa nacirema enola nacirema naidni dna aksala evitan enola naisa enola evitan naiiawah dna rehto cificap rednalsi enola emos rehto ecar enola -pop -italu fono owt secar -pop -italu fono eerht secar -upop noital fo ruof secar -upop noital fo evif secar -upop noital xisfo secar ademala ytnuoc 147,344,1 715,263,1 433,407 895,512 641,9 812,592 241,9 -0,921 97 -6,47 40 -39,5 8 755 321 2 enipla ytnuoc 802,1 741,1 098 7 822 4 71 85 3 0 0 0 12 iassist quarterly fall 2001 tables. by separating the logical framework of the data from the layouts, it is possible to interface to multiple storage formats and to prepare a metatable display capability without having to hard-code display structure or order into the report preparation code. other suggested and possible uses of the output table definition include: • a means of communication between remote/disparate systems, i.e., web services • a way of transporting data between analysis packages • a dynamic method of creating time-series data, e.g., include variable/file definitions appendix please see http://countingcalifornia.cdlib.org/moreinfo/ddi_ext_app.txt for varmtx examples.* paper presented at the iassist/ifdo conference 2001 in amsterdam. footnotes: * california digital library, university of california. t university of california, berkeley 1 the counting california project is funded by the university of california’s california digital library and the state-based library of california. additional funding for continuing development comes through a federal grant from the library services and technology act (lsta), administered by the california state library. vol273.indd iassist quarterly fall 2003 17 1st international conference on e-social science manchester june 22nd 24th 2005 initial announcement and call for submissions < www.ncess.ac.uk > the vision of the ‘grid’ first emerged as a solution to the highly specialised computing infrastructure requirements of particle physics. the past five years, however, have seen the grid’s potential recognised by the wider scientific research community and the emergence of new forms of research practice now encapsulated in the notion of ‘e-science’. now, members of the social science research community in the uk and elsewhere are beginning to explore how they can use the grid and to explore the prospects for ‘e-social science’. this year, for example, has seen the creation in the uk of the national centre for e-social science (ncess). the opportunities presented by the grid for social science research are numerous and intriguing. the grid will make it possible for new computational tools to be brought to bear on a diverse range of social science research problems; it will make established social science datasets more readily accessible and easier to integrate; it will make feasible the collection and management new kinds of data on an unprecedented scale. beyond enhancing existing research methods, however, e-social science also brings with it the prospect of articulating a radically new research agenda and encouraging the formation of new forms of research community. realising the full potential for e-social science will be a major challenge and calls for a major collaborative effort from social scientists and grid developers. as a contribution to meeting this challenge, ncess is very pleased to announce the first international conference on e-social science. we invite contributions from members of the social science and grid research communities with experience of – or interests in – exploring, developing and applying e-social science research methods, practices, tools and technologies. vol21.3 fall 1997 25 rapid change in delivery method means that the information which is necessary to access data is also changing rapidly. the density at which a tape or cartridge was written is critical information for reading the data on it while if the data is online the location is critical. on way of handling such changes in required information is to store it in a relational database. what exactly is meant by a relational database? fully relational databases, satisfying all conditions set out in codd’s definition, may only exist in theory and would be more than needed for this discussion. here what is required is first that there be some degree of normalization of the data. as an example some studies have only one dataset while others have many, a record that tried to anticipate how many datasets is not likely to be very practical so put the study information in a study record and the dataset information in a dataset record with one record for each dataset and a key to the study record. then it is necessary that there be some way of accessing these records together so that it looks as if two (or more) separate records are really one. structured query language (sql) is the accepted (complete with an iso standard) way of doing this for a relational database. using the example of studies and datasets in a simplified version, let’s say we have two tables which is the accepted term in relation databases for the structure in which the records, called rows, are stored. the first table is called studies and has fields, columns in relational terms, study_num and title and the second table is called datasets and has the columns study_num, dataset_num, and name. now a very simple example to make sure we are all on the same page. you would like to have a list of all study titles and the names of the datasets for each study. you could issue the sql command select studies.title, datasets.name from studies, titles where studies.study_num = datasets.study_num; this will give you a list of study titles and dataset names which the title repeated for each dataset. since you did not request study_num or dataset_num you will not get them back. the experiences being talked about here used relational database management systems. these applications were started under ingres and migrated to oracle when the university wide choice of a database system made the switch. when dealing just with data some of our researchers have used sql under sas. the focus here is not on the specific relational database but rather on the concepts so specifics should be taken as concrete examples rather than the only way of doing it. although querying the database directly using sql is an option and it is possible to use sql scripts in place of some of the things used here what will be talked about are oracle forms applications (developed with oracle developer/ 2000), perl scripts, and c programs. the c could probably be replaced with perl but it was a very ambitious undertaking written by the system administrator. when a perl script has to interact with the database it is currently oraperl which is perl4. the perl5 scripts use dbi/ dbd::oracle and will go into production when the c application has been successful tested out against a more recent version of oracle. because they are probably only of interest to show the wide range of uses i will briefly touch on the unix administration applications. the c application (using pro*c precompiler to interface with oracle) is a print accounting program that keeps track of printing on our unix cluster and our nt network. users are given an allotment for printing each semester and must pay for additional printing. at the end of each semester a reconciliation is done from a cron job to zero out any allocated funds not used and put in the next semester’s allotment. the cron job is a perl script. the application for providing additional funds (user paying, refund because of bad printing, etc.) is a forms application. the other administrative application is for keeping track of our users. since not everyone on campus can have an account on our system this application must check with another database on campus to determine if a person is in the correct department and has the correct status (no undergraduates). this is accomplished by means of a tying everything together with a relational database by pat hildebrand * 26 iassist quarterly database link. i don’t know terminology for other databases but with oracle a database link is a means of accessing a table in another oracle database as if it were a table in the local database. the printing is related to system users by a column indicating what allocation of print funds they get. data access is a bigger issue as while we provide computing to a limited group we provide data to the entire university. since we require the use without an account on our system to show up in person and present then campus id we originally developed applications under ingres that were used on our unix system in character mode so that calling in from home did not present a problem. when the applications migrated to oracle and developed under developer/2000, we found that it didn’t make sense to develop in character mode any longer. even with ppp when people called in from home they were still using character mode as the campus software had vt220 emulation. the database had become even more a part of the application under unix as that is when we started putting data on line so that there was even more information that we were keeping track of although the user probably used less of it. when we designed the database we had been using tapes and cartridges for obtaining the data and at first the data was still used on the mainframe where the access was via a tape job. the tables included one for tape labels which also kept track of the tapes that were used for backing up accounts, temporary use, blank tapes entered into the system but not yet used, etc. another table contained tape information such as the density, the character set, and whether the tape was labeled. at first glance it might appear that these two tables are each one row per table but the labels table has text which could require more than one row. other tables are for the individual files on the tape. when we started putting data online we could redesign the tables dealing specifically with the tapes to include online information or use the fact that the database was relational and use an additional table for the online information. we choose the later since not all data was being placed online. data requests, information about what is online, and the library system for hard copy documentation all share tables about the studies as well as having their own tables. for data are current system is a mixture of perl scripts, forms applications, and even some cgi scripts for web access to the information. the web access takes care of the university wide access to the information. people on campus but outside those who have accounts on our system still have to present an id and request access in person but now they know if the data are on campus. if the data are on campus they are able to get some information such as the size of files that they are talking about before they make the request so that they can make sure they have the space. we still get people who only want a few numbers, often a statistic that they would have to calculate from a very large dataset, thinking that everything will fit on a floppy that already has a number of files on it, but the web seems to have found users who are better prepared to make use of the data when they come over to request access. the heart of our data system is a series of perl scripts. when new data are received information is entered into the database about the study and dataset numbers and the type of file (data, documentation, program, etc.). this is actually a forms application. as the information is entered a todo file is written with information from this table and a number for later identifying the file internally. please note that there is additional information entered form a master-detail relation in the form and programmatically. perl scripts to read tapes/cartridges or process files which have been ftped or are on other media such as cd-rom use the todo files to determine if a file should be read and do the processing. the todo files are also generated for existing tape/cartridge files to be put online and moving some files but not all from a cd-rom to disk so the issue of whether or not a file should be read is real. when processing is complete (some, unfortunately, is still manual checking) a file is written with information that should go into the database as well as information needed to move the file from the processing area to where it is accessible to the users. the names of these files are placed in another file that is read by a nightly cron job. the actual moving of the data and recording of the information in the database is done by this script. some of the other things done by the script are to check the type of file, find out from the database where that type of file should be moved, check that there is enough space for the file, if need be and there is one available set the database to use a new directory, check that there is not currently a file in the directory with that name, and send e-mail about problems. the library application is able to tell whether we have any hard copy documentation for a specific study and wether it is on hand or checked out, check thing out, and check them back in. if someone already has a copy of a specific piece of documentation checked out they are not permitted to check out a second copy of the same thing. also if someone has overdue documentation checked out the application will say what is overdue and no further check outs are permitted for that individual until the overdue documentation is returned and the situation cleared is some other way. all of the user information for requesting data and checking out documentation is the same table so that fall 1997 27 changes do not have to be entered in multiple locations. as much of the information as possible is look up information from other tables. this not only makes the entry simplified but it also makes for fewer errors’ in addition to avoiding typos this avoids ambiguous entries such as "student". there are a lot of relations that exist in the services that we provide. using a database allows us to restrict who can do what at what time. using a relational database for the necessary information has made for greater accuracy, a simplification of things since once we have entered an individual say for data we don’t have to do turn around and enter them again for checking out the documentation, and the ability to do automate some of the work. references: date, c.j. with hugh darwen. a guide to the sql standard, 3rd ed. reading, ma: addision-wesley publishing co. 1993. edelstein, stephen. learning oracle forms 4.5: a tutorial for forms designers. new york, ny: relational business systems. 1995. gundavaram, shishir. cgi programming on the world wide web. sebastopol, ca: o’reilly & associates, inc. 1996. musciano, chuck & bill kennedy. html the definitive guide. sebastopol, ca: o’reilly & associates, inc. 1996. wall, larry, tom christainsen & randal l. schwartz. programming perl, 2nd ed. sebastopol, ca: o’reilly & associates, inc. 1996. http://www.hermetica.com/technologia/dbi oracle developer/2000 documentation forms developer’s guide release 4.5 part no. a32505-1 forms reference manual, volume 1 release 4.5 part no. a32509-1 forms reference manual, volume 2 release 4.5 part no. a32510-1 forms advanced techniques release 4.5 part no. a32506-1 forms messages and codes release 4.5 part no. a32508-1 * paper presented at iassist/ifdo ‘97, odense, denmark, may 6-9,1997. pat hildebrand, social science computing, university of pennsylvania, pat@ssc.upenn.edu http://www.hermetica.com/technologia/dbi mailto:pat@ssc.upenn.edu vol263 16 iassist quarterly fall 2002 iassist quarterly fall 2002 17 by jackie carter, mark brown, cressida chappell * background through a number of strategic investments by the uk funding bodies for further and higher education.and the research councils, the uk academic community has access to an electronic collection of historical and contemporary census data and resources (chcc). individual datasets have been used extensively in research but they have been widely under-used in learning and teaching. to address this issue, the joint information systems committee (jisc), under its learning and teaching programme, has funded a project to develop learning and teaching materials based on the chcc. the jisc is a strategic advisory committee working on behalf of the funding bodies for uk tertiary education (fe and he). this role allows it to provide a centralised and co-ordinated direction for the development of the infrastructure and services needed to support this sector in exploiting ict for its learning, teaching and research activities. as part of its remit, jisc promotes the innovative application and use of information systems and information technology in tertiary education across the uk, by funding development projects and the creation of high quality materials for education. currently jisc, in line with other uk initiatives, is funding a series of development projects to create high quality resources, which can be directly integrated into learning and teaching. the outcomes from these jisc funded projects will feed into the information environment being developed by jisc, to deliver resources to students, teachers and researchers in meaningful ways (ingram, 2002). the chcc project is a multi-partner venture involving six partners from four uk universities; essex, glasgow, leeds and manchester. the project started in october 2000 and will run until september 2003. further details can be found at the website . for details about other projects that are funded under the jiscʼs learning and teaching programme see carter et al., 2002, or the website what is the chcc? the first modern census in the uk was taken in 1801, learning and teaching with the uk census with a census being conducted once every ten years since then, with the exception of 1941. the last census was taken on 29th april 2001, with plans to make the first statistics available in 2003. the national statistics web site is an excellent source of detailed information about the uk census . developments in data collection and provision in recent years have resulted in the uk academic community having free-at-the-point-of use access (subject to undertaking of agreements for terms and conditions of use) to a wide range of digital census datasets. the first paper in this session outlines issues concerned with access to these datasets in more detail. (bell, 2002). the collection of historical and contemporary census data and resources (chcc), as outlined below, is part of a much fuller collection. it is important to note that the datasets comprising the chcc are physically located in three different locations, and supported by three different support services. the chcc refers to the following datasets: • the census area statistics (cas) from the 1991, 1981 and 1971 census of population. these are aggregated counts for a variety of geographical areas, such as county, district, ward and smaller areas. the 2001 census area statistics will become available in 2003. • the samples of anonymised records (sars) released for the first time from the 1991 census. these are samples of individual census returns both for individuals and households. the sars for 2001 will also become available in 2003. for further information about the sars refer to the website at . • the historical censuses collection comprises individual-level data from the 1851 census (2% sample) and 1881 census (100% sample) and aggregate-level data from nineteenth and early twentieth century printed census reports. for further details see http://www.jisc.ac.uk/ 18 iassist quarterly fall 2002 iassist quarterly fall 2002 19 the chcc provides the most comprehensive and geographically detailed set of statistics relating to people and housing in the uk. the 1991 census covered a wide range of topics, which describe the characteristics of the uk population: demography, households, families, housing, ethnicity, birthplace, migration, illness, economic status, occupation, industry, workplace, transport mode to work, car ownership and language. the chcc thus provide a unique source of information for students wishing to explore important social, demographic and economic issues, in modern or historical contexts. the strategic importance of the census means that students from a wide range of academic disciplines will benefit greatly from a working knowledge of how to access and analyse these statistics. individual datasets from the chcc have been widely used in research (for example, see for list of sars related publications). however, the data have been less widely used in learning and teaching programmes. a number of factors have contributed to this: requirement for registration prior to data use; distributed nature of the census datasets; barriers to access due to hard-to-use interfaces; lack of knowledge about the existence of datasets and their potential for learning and teaching; lack of learning materials suitable for student use across disciplines. for an excellent summary of barriers facing teachers who use numeric data see the report by rice et. al, 2002. we also reported on barriers to the use of the census area statistics at iassist in 1999 (carter and bullen, 1999). the chcc project has developed clear aims, objectives and outputs to respond to these barriers, as detailed below. aims, objectives and outputs the primary aim of the project is to develop the chcc into a major learning and teaching resource. the key objectives to assist in achieving this aim are to: 1. promote increased and more effective use of network based data services for problem-based learning and student project work across a broad range of teaching programmes; 2. develop an integrated web-based learning and teaching system that links together data extraction and visualisation/exploration tools with comprehensive learning and teaching resources; 3. significantly increase the census user base by increasing use of the chcc in learning and teaching; 4. build new user communities by promoting increased awareness of the chcc and its learning and teaching potential; 5. improve the productivity of teachers by significantly reducing the overheads required to incorporate census data related resources into learning and teaching programmes; 6. improve access to key primary data related resources; 7. minimise delays in getting the key 2001 census outputs used in learning and teaching. four main outputs will be delivered from this project: 1. a range of learning and teaching resources for teachers and learners (tutorials, exercises, exemplarbased studies); 2. a census resource discovery system to facilitate discovery of census data and associated learning and teaching resources; 3. improved web-based interfaces for data extraction/ visualisation suitable for student use; 4. enhancement of the chcc through adding and linking other information. this paper reports on the development of these outputs, focusing on the learning and teaching materials and the census resource discovery system. learning and teaching resources a series of learning materials (ʻunitsʼ) are being developed both by the project team, and commissioned by distinguished authors who have used census data in their teaching. the units are designed to encourage a ʻpick and mix ̓approach; so that a teacher can choose from a portion of a single unit (perhaps a simple graphic illustrating a statistical distribution of, for example, unemployed males), through to several units built into a module. the units can be used for both classroom-based and online learning, and can thus be used both by teachers and directly by students for independent learning. such flexibility was a key requirement behind the development of the units and indicates our direct response to feedback from user consultation established at the start of the project. another key requirement elicited from user consultation was the need to allow for customisability of the units. authors provide the materials under a non-exclusive licensing agreement, which permits a teacher to alter the materials providing s/he credits the source and author of the original unit. due to conditions of funding the materials are available through restricted access to uk further and higher education sites only. http://www.ccsr.ac.uk/publications/sarpub.htm http://www.ccsr.ac.uk/publications/sarpub.htm 18 iassist quarterly fall 2002 iassist quarterly fall 2002 19 consultation the development of all learning and teaching materials is underpinned by a process of consultation and formative evaluation with potential end users (teachers and students). the first chcc user consultation workshop was held in january 2001, just after the project began. the feedback from this workshop was crucial in guiding our development of the learning and teaching units. several workshops have been held subsequently in which the first units have been piloted. the feedback received during these workshops is being used to improve upon existing units and assist in the development of future units. a period of piloting is currently taking place, in which the units are being tested in live teaching environments. evaluation from these pilot studies will be fed into a project report, and used to further enhance the materials. data specific units two types of materials are in development. the first are ʻdata specificʼ, that is, they are based on sets of data from the chcc (the cas, the sars or the hcc). for uk access these can be viewed at . it has been necessary to restrict access to users registered at uk he and fe institutions due to the terms and conditions of the projectʼs funding. cas units these resources provide a range of teaching and learning materials that cover many aspects of contemporary census area statistics (cas). each resource has been categorised into one of three types: census basics, census methods and census applications. a guide to the level of difficulty covering beginners, intermediate and advanced users has also been specified. this should help the teacher or learner choose an effective pathway through the resources. for further details about the units, see , sars units these units are being developed around two main themes 20 iassist quarterly fall 2002 iassist quarterly fall 2002 21 • introductory data analysis units covering a wide range of topics starting from basic exploratory techniques through to more formal methods of statistical analysis. • substantive issues units covering more substantive discipline-specific topics. all units have a common and flexible format, that can be readily adapted for use in classroom teaching and selfstudy. each unit comprises three types of materials: 1. overview designed for use by lecturers; suitable for a half hour lecture. 2. detailed notes more detailed notes with a greater degree of statistical content, designed to supplement teachers ̓notes and suitable for student self learning. 3. practical exercises generic and software specific exercises, suited to incorporation in lessons or for homework. figure 1 illustrates the structure of the sar methodological units, each of which incorporates a presentation (in this case, on the use of graphics), supported by detailed explanatory notes and exercises. in this example, a slide in the presentation introduces the use of bar charts. this provides the platform for a more detailed ʻtext-book ̓ type explanation of bar charts and some linked exercises that enable students to explore the construction and interpretation of bar charts with some real data. exercises vary in the level of difficulty and format, and include both computer and non-computer based exercises. for computerbased exercises, nesstar provides a user-friendly on-line interface by which to undertake exercises on subsets of sar data. for further information about these units, including details of which are ready for piloting, see hcc units units on the following topics are under development under the main theme of ʻbritish history and the censusʼ: understanding the nineteenth century census; household and family structure; migration in the nineteenth century; urbanisation; social status; work and employment; using the census for local history; skills. for further information about these units, see < http://chcc.gla.ac.uk/ie_index.php> combined units the second type of material that will be developed is based on common topics across the census datasets contained in the chcc. combined units include introductory resources that are developed jointly by the project partners, and include links to the datasets. these units are in the early stages of development. the topics fall into three main areas: 1. what is the census? these units include: overview; census topics; the census questionnaire; census geography; coverage; methods of data collection and coding; census outputs; the census elsewhere. 2. using the census these units include: confidentiality; data availability and dissemination; individual versus aggregate data; change over time. 3. uses of the census these units include: key users and uses of cas; key users and uses of sars; key users and uses of hcc; what future the census? web interfaces the efforts to develop teaching materials around the census data sets have been considerably enhanced by a dramatic improvement in the ease and flexibility by which users can access the underlying data. for example, when sars were first released, their use in teaching was seriously limited by the need for users to be familiar with a specialist statistical package, such as spss, and to access the data via mainframe computer. the situation got better as the improving specifications of pcs enabled users to put sar data files directly on to their desktops. however, the requirement to use a specialist statistical package still meant the data were inaccessible to many users. as part of the chcc project, it has been possible to incorporate new and user-friendly web-interfaces to census data, such as nesstar (see figure 1), opening up the teaching materials to a much wider user group. development work is also taking place to enable the cas datasets to be visualized online, using interactive exploratory data visualization software. this software, called descartes (see www.mimas.ac.uk/descartes/ for a sample dataset), has been developed at the fraunhofer institut autonome intelligente systeme (fhg ais) . it has been developed for the project, to improve access to the cas by providing online mapping of census datasets in a web browser. census resource discovery system the census resource discovery system will facilitate resource discovery of both census data and the associated learning and teaching materials. it will provide a flexible, consistent, integrated interface to the learning and teaching materials and a mechanism for effective resource discovery across the learning and teaching materials and the census data. the main goals of the system are to: allow students, teachers and researchers to discover and locate census data and associated learning and teaching resources; provide links to on-line census data and associated learning and http://www.mimas.ac.uk/descartes/ http://www.ais.fraunhofer.de/ 20 iassist quarterly fall 2002 iassist quarterly fall 2002 21 teaching resources; and provide a means for census data providers to publicise and propagate the use of their census data holdings and associated learning and teaching resources. the system will provide a web interface to ʻsearch ̓and ʻdiscover ̓census data and associated learning and teaching resources through the entry of any combination of a subject keyword, a geographic area (england, wales, scotland, or northern ireland), a time period (date or span of dates) or a resource type (aggregate-level data, individual-level data, or learning and teaching resource). users will also be able to browse hierarchical lists of census data and associated learning and teaching resources ordered by subject, geography, time-period or resource type. a list of items matching the browse or search criteria will be returned and users will be able to select and view the full record for one or more items simultaneously. tools for project partners to add, modify and delete metadata records are being developed and users will be able to maintain a user profile to store query definitions for future use. the ddi codebook standard is being used to describe census data sets and the ims learning resource metadata specification is being used to describe learning and teaching materials. the european language social science thesaurus is the source for subject keywords. an sqlserver database is being used to hold the metadata and java technologies are being used for the web interfaces. a cheshire z39.50 target will be used to ensure integration with the jisc information environment. the census registration service and athens single sign on will allow users to move seamlessly between the different census data sets regardless of census data provider. references bell, lucy “let us bring you to your census: recent developments in uk census data provision.” paper presentation in delivering the uk census; web-based access, iassist 2002, connecticut, us, 10-14 june, 2002. carter, jackie et al. “enhancing use of jisc data services.” vine 126, pp.40-47, 2002 carter, j and bullen, n. i “overcoming barriers to the use of the census through interactive visualization.” paper presentation in gis and social science data access over the net, iassist 1999, toronto, 17-21 may 1999. ingram, caroline “jisc development.” vine 126, pp.3-6, 2002 hayes, j. and carter, j. “web-mapping the census: developing interactive interfaces for learning and teaching.” society of cartographers bulletin, vol 35, no2, 2002 rice, r., burnhill, p., wright, m. and townsend, s. “an enquiry into the use of numeric data in learning and teaching” [computer file]. edinburgh: edinburgh university, september 2001 * dr jackie carter, chcc project manager, mimas, manchester computing, university of manchester, email: j.carter@man.ac.uk dr mark brown, deputy director ccsr (centre for census and survey research), university of manchester, uk cressida chappell, head of the history data service, uk data archive, university of essex, uk http://datalib.ed.ac.uk/projects/datateach/report/ http://datalib.ed.ac.uk/projects/datateach/report/ name: job title: organization: address: city: state/province: postal code: country: phone: fax: e-mail: url: iassist international association for social science information service and technology association internationale pour les services et techniques d'information en sciences sociales i would like to become a member of iassist. please see my choice below: options for payment in canadian dollars and by major credit card are available. see the following web site for details: http://datalib.library.ualberta.ca/membership/ membership.html $50 (us) regular member $25 student member $75 subscription (payment must be made in us$) list me in the membership directory add me to the iassist listserv membership form the international association for social science information services and technology (iassist) is an international association of individuals who are engaged in the acquistion, processing, maintenance, and distribution of machine readable text and/or numeric social science data. the membership includes information system specialists, data base librarians or administrators, archivists, researchers, programmers, and managers. their range of interests encompases hard copy as well as machine readable data paid-up members enjoy voting rights and receive the iassist quarterly. they also benefit from reduced fees for attendance at regional and international conferences sponsored by iassist. membership fees are: regular membership: $50.00 per calendar year. student membership: $25.00 per calendar year. institutional subcriptions to the quarterly are available, but do not confer voting rights or other membership benefits. institutional subcription: $75.00 per calendar year please make checks payable, in us funds, to iassist and mail to: iassist, assistant treasurer joann dionne 50360 warren road canton, mi 48187 usa r eturn u ndelivered m ail to: ia s s is t q u a r t e r ly c/o w endy treadw ell 1758 pascal st. n orth falcon h eights, m n 55113 u sa http://datalib.library.ualberta.ca/membership/membership.html http://datalib.library.ualberta.ca/membership/membership.html vol233 4 iassist quarterly supermarket:where do social scientists shop? by kirsti nilsen* this paper presents some findings of research which examined the statistics and data sources used by canadian social scientists, the formats in which they obtained the data, and their preferences with respect to data formats. five disciplines were the focus of the research: economics, education, geography, political science, and sociology, based on a literature review which is summarized below. the research was part of a larger study which examined the effects of government information policy on canadian social scientists. that research focused on policyinitiated price and format changes at statistics canada. (nilsen, 1996, 1997, 1998). in order to monitor the effects of the policy it was necessary to determine which statistics and data sources were used and any changes in that use over a period before and after policy implementation. using both bibliometric and survey methods to gather data, the study identified statistics sources used in published articles over the period 1982 to 1993, and supplemented those findings with a survey of authors in the fall of 1995. the terms “statistics” and “data” have unique definitions; however, for the purposes of this research, the terms tended to be used interchangeably as they are in everyday speech. in the survey, respondents were asked about their use of “statistical data (i.e. numeric information)”. literature review research on social scientists’ use of and demand for materials has confirmed that social scientists do use statistics and raw data. because governments collect, analyze and publish the largest amounts of data, social scientists will use government-produced statistics, along with other statistics sources. obviously not all social science disciplines use published statistics and data sets to the same extent. in order to determine which disciplines should be the focus of this study, published research on social scientists’ use and demand for materials was examined. it provided the data needed to identify those social science disciplines which use published statistics. where statistics were not specifically identified, use of government publications served as an indicator of use of statistics because, as hernon had shown, social scientists use government publications to obtain statistics more than for any other purpose (hernon, 1979, p. 10). no research was found which distinguished between use of statistical publications of governments versus those of other publishers. use of statistics by social scientists the first major study of users of social science materials was undertaken by the investigation into information requirements of the social sciences (infross, 1971) in the united kingdom. with 1,089 social science researchers responding to the infross survey, it has been described as the largest, most ambitious and influential study in the area (slater 1989, pp. 1). no research on a comparable scale has been done in north america. infross provided extensive data on the use of a variety of types and physical forms of information, along with data on information demand, by discipline, and with comparisons among disciplines. it specifically addressed the question of the use and perceived importance of statistics by researchers in each of the disciplines covered (anthropology, economics, education, geography, political science, psychology, and sociology). the infross study found that statistical, methodological and conceptual information was used by almost everyone, while historical and descriptive information was least used (line, 1971, p. 416). statistical material was used by 91% of respondents and over half used it frequently in their research. when asked to rate the importance of types of materials to themselves, infross found that 58% of respondents rated statistical material as very important, 20% rated it as moderately important, and 12% rated it as not very important (infross, 1971. vol.1, pp. 48, 50, 52). with respect to disciplinary differences in use of statistical materials, infross found that economists were the heaviest users of statistical data, followed closely by geographers. when asked to rate the importance of statistics, economists and geographers were much more likely than any other researchers to rate statistical material as “very important” and historians and anthropologists less likely to do so (infross, 1971, vol.1, pp. 43, 51). in its analyses of statistics use, infross did not discriminate between data which were self-collected and fall 1999 5 data gathered and published by someone other than the researchers themselves. however, the type of raw data used (e.g. interviews, experiments) was correlated with discipline of respondents (infross, vol. 2, table 20). the report noted that psychologists were more likely to use empirically derived data from experiments conducted by themselves than were other social scientists (infross, 1971, vol.1, p. 57). use of government information by social scientists in reviewing the literature on citation studies, hernon and shepherd determined that the percentage of citations to government publications ranged from 2% to 36% (hernon & shepherd 1983, p. 227). weech found that in various citation studies a median of 17.5% of total references were to government publications (weech, 1978, p. 179). the largest citation study of social scientists’ use of materials was design of information systems in the social sciences (disiss, 1979), a follow-up study to infross. disiss collected data from 140 social science serials, published mostly in 1970, for an examination of social science literature via citation analysis. out of 47,342 citations only 2.7% were to official (government) publications (disiss, 1979, p. 75). the variability in the findings on use of government publications among social scientists can be accounted for by disciplinary differences in the choice of disciplines included in citation studies. low percentages in general relate to the fact that statistical sources are often not cited in footnotes or reference lists (hernon & shepherd, 1983). the infross survey found that 34% of social science researchers used government publications “often”, while 23% never used this form of material (infross, 1971, vol.1, p. 53; line, 1971, p. 417). when use of government publications was examined by discipline, the investigation found that 53% of researchers in economics stated that they sometimes or often used them, followed by those in sociology (41%), education (29%), geography (22%), and political science (20%). fewer than 10% of researchers in anthropology, history and psychology used government publications (infross, 1971, vol.2, table 59). hernon (1979) investigated the use of government publications by faculty members from economics, history, political science, and sociology departments in american colleges and universities. he found a statistically significant difference among the four disciplines in frequency of document use, with economists and political scientists as the heaviest users of government publications, which was consistent with the infross findings (hernon, 1979, pp. 9,45). some research has shown which disciplines seek statistical information within government publications. hernon found that the “top priority of economists and sociologists [in using government publications] is to gather census and normative data,” and that historians used government publications for historical data more, while political sciences use them equally for statistics and current events information (hernon, 1979, p. 51). other studies by hernon and shepherd (1983) and hernon and purcell (1982) corroborated hernon’s earlier findings. determining the disciplines for this research on the basis of the infross and disiss research, which has been substantiated by other research, a typology of use of statistics and government publications was developed, as shown in table 1. based on this typology, and the research which supports it, five disciplines were identified which use primarily published statistics and sometimes or often use government publications. these five disciplines were economics, education, geography, political science, and sociology. thus, these five disciplines defined the domain of this research. methodology two methods were used to gather data on the use of statistics sources. bibliometric analysis provided objective evidence of use of statistics, while a survey supplemented the findings with more subjective data. a systematic, 1elbat tnemnrevogdnascitsitatsfoesufoygolopyt :snoitacilbup enilpicsidyb )1( :scitsitatsesurevenromodleshcihwsenilpicsid yrotsih,ygoloporhtna )2( tnemnrevogesurevenromodleshcihwsenilpicsid :snoitacilbup ygolohcysp,yrotsih,ygoloporhtna )3( :scitsitatsesunetforosemitemoshcihwsenilpicsid lacitilop,yhpargoeg,noitacude,scimonoce ygoloicos,ygolohcysp,ecneics 1 )4( :scitsitatsdetcelloc-flesesuhcihwsenilpicsid ygolohcysp 2 )5( :scitsitatsdehsilbupyliramirpesuhcihwsenilpicsid ,scimonoce lacitilop,yhpargoeg,noitacude ygoloicos,ecneics )6( esunetforosemitemoshcihwsenilpicsid :snoitacilbuptnemnrevog lacitilop,yhpargoeg,noitacude,scimonoce ygoloicos,ecneics 3 6 iassist quarterly stratified and proportionate sample of 360 articles was selected from a population of 5,414 articles in 21 canadian social science journals in the five disciplines noted above. the source journals were published in english or french in canada, covered primarily canadian topics, focussed widely in the discipline, were peer-reviewed, and published over the entire period 1982 to 1993. all journals which met these criteria were included. articles to be included in the population to be sampled were those listed in the tables of contents under “articles” or “research notes” or similar headings. in the final sample, the disciplines were represented in proportion to the amount of publishing in the 21 journals: economics 26.9% (97 articles), education 22.5% (81 articles), geography 7.2% (26 articles), political science 18.1% (65 articles), and sociology 25.3% (91 articles). the 360 articles in the sample were examined and data were collected from the text, tables, and citations. the bibliometric examination revealed the statistics sources used by the authors of the articles. all uses of statistics sources, whether documented or not, were recorded, whether governmental, nongovernmental, canadian or foreign. more detailed information about the use of statistics canada was gathered for the policy effects aspect of the larger study. data analysis dealt with the complete sample and, in more detail, with a subset of 207 articles which were identified as using published statistics and written with a canadian focus or setting. a survey questionnaire was sent in english or french to 163 authors (all who could be located) of these 207 articles. ninety-seven responded (59.5%). the questionnaire asked for background information, extent of use of statistics, statistics sources used, formats used and preferred, means of obtaining data, and opinions regarding prices and formats of data.. findings the 360 articles sampled for the bibliometric component of the research were categorized as to discipline, type, geographical focus or setting (if any), and language. the categorization was by discipline of the journal, (which was not necessarily the discipline of the author or of the subject covered). most articles (78.4%) could be categorized by type as either empirical (200, 55.6%) or descriptive (82, 22.8%), both of which were likely to use statistics. the remaining 78 articles (21.6%) were either historical, opinion, methodological, or theoretical, articles less likely to use statistics. the geographical focus or setting was canadian in 269 of the articles (74.7%), and the focus was not canadian in 34 articles (9.5%). the remaining 57 articles (15.8%) could not be categorized geographically, usually because of their methodological or theoretical focus. two-thirds of the articles (239, 66.4%) were in english, 121 articles (33.5%) were in french. as expected, not all of the articles used statistics. as figure 1 illustrates, 70 articles (19.4%) made no use of any statistics, most of these were categorized as theoretical or methodological. thirty-nine articles (10.8%) used only self-collected data derived by the author from experiments or other research methods. a few articles (13, 3.5% used only unidentifiable published statistics which could not be categorized as to source. the remaining 238 articles (66.1%) used identifiable published statistics. more than 70% of the articles in each discipline (excepting education) used identifiable published statistics. some of these also used self-collected data. a subset of 207 articles was identified which had a canadian focus or setting and used published statistics and this subset provided the data which follows. statistics sources used information was gathered on the use of the following broad categories of statistics sources: statistics canada, other canadian federal and provincial/municipal governments, foreign governments, intergovernmental, nongovernmental. more detailed information was gathered on the use of statistics canada in terms of formats used. it was found that social scientists used a wide variety of statistics sources and many used multiple sources. figure 2 illustrates the percentage of articles which used the various fall 1999 7 sources. as can be seen, statistics canada was used by 41.1% of the articles, and other canadian federal sources were used by an almost equal number (40.6%). american sources (governmental and nongovernmental combined) were used almost as much as was statistics canada, which was somewhat surprising in these 207 articles with a canadian focus or setting. nongovernmental sources were used by the highest percentage of articles (71%). nongovernmental sources include trade and scholarly books and journals, associations, universities, business, think tanks, polling organizations, etc. the survey respondents indicated higher use of all statistics sources (except nongovernmental) than was found in the bibliometric research. this probably results from the fact that the bibliometric analysis was looking at one-time use in a single article while the survey questioned life-time use. for example, 86.5% indicated that they had used statistics canada at some time, but only 41.1% used statistics canada in the articles. however, only 41.5% of survey respondents indicated that they used statistics canada often or almost always (i.e. more than 50% of the time in the years between 1985 and 1995), which is more consistent with the bibliometric finding. they also indicated less use of nongovernmental sources than was found in the bibliometric analysis. this difference might result because respondents might have been thinking of major nongovernmental suppliers such as polling organizations, rather than their use of sources from which they might obtain single facts such as a book or journal article. disciplinary differences there were statistically significant disciplinary differences in use of most statistics sources in the articles as is shown in table 2. note that variation by discipline was statistically significant for all sources but provincial government and foreign government sources. table 2 shows the importance of other canadian federal government sources in articles from economics and political science journals. those writing in education and political science journals were more likely to use other canadian federal government sources more than they used statistics canada. geography and sociology journal articles used statistics canada more often, and economics used the agency’s 2elbat nienilpicsidfotnecrepyb:desusecruosscitsitats )702=ngnittesrosucofnaidanacahtiwselcitra ecruosscitsitats oce ude oeg lop cos adanacscitsitats )100.=p( 3.85 3.32 2.27 2.91 1.24 laredef.ndcrehto )100.=p( 2.85 0.03 2.22 1.15 3.62 laicnivorp )821.=p( 7.21 0.03 2.22 7.72 3.33 lanoiger/lapicinum )520.=p( 5.5 3.31 8.72 3.4 0.5 tvogngierof )680.=p( 8.12 7.6 6.5 6.01 1.7 tvoglaredefsu )342.=p( 2.81 7.6 6.5 5.8 1.7 latnemnrevogretni )561.=p( 3.7 3.3 6.5 6.01 0.0 latnemnrevognon )810.=p( 3.76 3.37 6.65 3.98 2.36 tvognonnaidanac )300.=p( 9.05 3.35 6.55 9.08 9.34 tvognonnacirema )820.=p( 1.92 0.02 2.22 1.15 8.92 8 iassist quarterly statistics at approximately the same rate as they use the statistics of other federal sources. nongovernmental sources were important for all disciplines, particularly political science, which made heaviest use of both canadian and american nongovernmental sources. use of computer readable products in the survey (conducted in the fall of 1995), most respondents (81%) indicated that they had used computer readable formats at some time. there was statistically significant variation by discipline in these responses, with 100% of those who had published in economics and geography journals indicating prior use, while 77.8% of those in sociology, 73.7% of those in political science, and 56.3% of those in education indicating such use. however, when asked how they normally obtained statistics most still used paper (print) formats more than computer readable files. of the 97 respondents, 74 (76.3%) indicated that they obtained data in paper format, while 59 (60.8%) used computer readable formats, or both formats, as seen in table 3. questioned as to how they normally acquired the data they used, responses are shown in table 4: respondents were then asked to rank their first preference and their first three preferences of the various means of acquiring data, as shown in table 5. note that where 47% in table 4 used a library to acquire paper copies, for only 20% was that a top three preference. a larger percentage ranked purchasing computer readable files as a top three preference than had indicated normally acquiring data in this way. also, more preferred to use a data library. fewer preferred to collect their own data than actually did so. the bibliometric analysis provided objective data on the actual use of paper and computer readable formats. the determination of use of products by format focussed on the 85 articles with a canadian focus or setting which used statistics canada as a statistics source. using various statistics canada catalogues and other sources where necessary, the researcher determined the formats of the statistics canada issues used in the articles, if the author had not provided this information. seventy-one (84%) of these 85 articles used paper issues, while 29 (34%) used computer readable “issues”, with some articles using both formats. the format of some issues could not be determined in 11 articles. the ratio of number of articles using paper issues to the number using computer readable issues was 2.5:1. table 6 illustrates variation in the number of articles using issues by format. 3elbat yllamronstnednopseryevrushcihwnistamrof )79=n(scitsitatsdeniatbo stamrof on *tnecrep )repap,.e.i(tnirp 47 3.67 )mliforcim,ehciforcim(mroforcim 31 4.31 **selifelbadaerretupmoc 95 8.06 ***snoitalubatlaiceps 63 1.73 atadnwoymtcelloc 06 9.16 dluocstnednopseresuaceb%001deecxesrebmun* tamrofenonahteromesu yna"sadenifederewselifelbadaerretupmoc** cilbuprofdetaercselifretupmoc'flehs-eht-ffo` nihtiwnoitanimessiddetimilrofronoitanimessid .cte,ecremmocrossenisub otesnopsernidetaercerasnoitalubatlaiceps*** nirehtehw,esoprupcificepsaroftseuqercificeps mrofelbadaerretupmocrotnirp 4elbat deriuqcayllamronstnednopseryevruswoh )79=n(scitsitats scitsitatsgniriuqcafosnaem on *tnecrep atadnwoymtcelloc 65 7.75 selifelbadaerretupmocesahcrup 05 5.15 rorepaprofyrarbilaesu mroforcim 64 4.74 yrarbilatadaesu 73 1.83 seipocrepapesahcrup 53 1.63 snoitalubatlaicepsesahcrup 03 9.03 tenretniehtesu 91 6.91 rofnoitcelloclatnemtrapedaesu atadelbadaerretupmoc 51 5.51 rofnoitcelloclatnemtrapedaesu seipocrepap 01 3.01 seipocmroforcimesahcrup 9 3.9 rehto 2 1.2 dluocstnednopseresuaceb%001deecxeslatot* yehthcihwnoitisiuqcaatadfosnaemllaetacidni desu fall 1999 9 variation over the two time periods 1982-1987 and 19881993 in the number of articles using these formats was statistically significant for paper issues. it should be noted that while the year-to-year variations are not statistically significant, the percentage of articles in the sample, which used computer readable formats as a statistics (or raw data) source, increased in the last two years studied. these formats were used in 20% of the articles in 1992 and 36% of the articles in 1993. figure 3 illustrates year to year variation. this might be an indication of a trend which might have been evident in a larger sample and which could be examined in further research. there were statistically significant disciplinary differences in use of statistics canada’s paper ((r = .009) and computer readable formats (r =.014) among the 207 articles written with a canadian focus or setting. these are illustrated in figure 4. these data apply only to statistics canada formats, and as the discussion below indicates, other sources of computer readable information were used by these disciplines as well. political science for example, showed little use of statistics canada overall, but was a heavy user of nongovernmental materials, and made some use of computer readable sources, such as the national election studies. the decline in the number of articles which used paper formats might be attributed to the declining publication of 5elbat )97=n(scitsitatsgniriuqcarofsecnereferpdeknar scitsitatsgniriuqcafosnaem tnecrep gnitacidni tsrif eciohc tnecrep gnitacidni foeno eerhtpot seciohc elbadaerretupmocesahcrup selif 8.22 4.86 atadnwotcelloc 8.22 4.86 yrarbilatadaesu 3.02 9.15 seipocrepapesahcrup 1.01 6.62 repaprofyrarbilaesu seipoc 1.01 3.02 snoitalubatlaicepsesahcrup 1.5 0.91 tenretniehtesu 5.2 8.22 noitcelloclatnemtrapedaesu selifelbadaerretupmocrof 5.2 5.61 noitcelloclatnemtrapedaesu seipocrepaprof 3.1 3.6 6elbat adanacscitsitatsgnisuselcitraforebmun tamrofyb:stcudorp adanacscitsitatsdesuhcihwselcitrani )58=n( doirepemit repap )540.=p( elbadaerretupmoc )092.=p( 7891-2891 )93=n( 63 )%29( 11 )%82( 3991-8891 )64=n( 53 )%67( 11 )%82( figure 3 use of statistics canada paper and computer readable products in articles written on canadian topics (n=207) 10 iassist quarterly paper formats at statistics canada, rather than any absolute preference. however, the survey responses suggest that computer readable formats are preferred. it should also be noted that of the surveyed respondents who began to do research after 1980 (younger researchers?) 85% indicated that they had use statistics canada computer readable files at some time, while of those who began to do research before 1970, only 46.5% had used them. this suggests that in the future data users will rely on the computer readable files to an ever greater extent. machine readable data files used when computer readable files were used as major sources of data in articles, the titles of the mrdfs were recorded. because authors tended to cite these materials incompletely, if at all, the following discussion should be interpreted cautiously. statistics canada computer readable files were used by articles in education, economics, geography and sociology, with articles in economics journals using the greatest variety of files. special tabulations were used by economics and geography authors for census data, and by economics authors for family expenditures, manpower, manufacturing, and agriculture data. an education article used the labour market activity file; a justice database was used by one sociology article. public use sample tapes were used by one article from economics and two from sociology. cansim was mentioned by only one author. other canadian federal mrdfs were used in economics and political science articles. three economics articles used labour canada files, and the international trade data bank was used by a political science article. quebec provincial health databases were used by two sociology articles. one us government database was used by an economics article (dept. of agriculture cris), and two sociology articles used us government data obtained from icpsr. canadian universities were an important source for data for sociology articles, and to a lesser extent for political science and economics articles. here, the national election studies were used by an article in economics and one in political science. both york university’s quality of life survey and the university of western ontario’s canadian fertility study were each cited by one economics article and one sociology article. two francophone sociology journals cited sorep data on quebec population, while a third cited a database created at the école des hauts études commerciales. other databases used include one use of the fao trade tape, one use of the data from the correlates of war project (us university), and proprietary databases were cited by one economics article. the above information suggests a rather limited used of mrdfs by canadian social scientists. however, as noted above, authors do not cite these sources with any consistency. additionally unless an item could be clearly identified as an electronic file, it was assumed to be a paper product if such a product was available in print. thus, it is possible that some items that were recorded as paper products were in fact electronic files. bibliometric analysis of the use of electronic files suffers from inconsistencies in citation practices. this was noted as early as 1982 (white), but the situation has not improved. discussion the findings of this research are consistent with the findings of earlier studies cited in the literature review. social scientists do indeed use a wide variety of sources to obtain statistics and raw data. there are statistically significant variations in the sources used among disciplines. if any agency such as statistics canada wishes to expand its market, analyses by discipline can assist in identifying target consumers, or areas where its products are not meeting the needs of researchers. at the time period covered by the bibliometric research (1982-1993) and the survey (1995), social scientists still used paper products more than computer readable products to obtain statistics and data, but there was a statistically significant decline in the use of paper products. figure 4: use of paper & computer readable formats percent of articles in each discipline using statistics canada materials in each format fall 1999 11 additionally, respondents to the survey were enthusiastic about computer readable formats. these finding suggested that computer readable formats would be used more heavily in the future, indeed, in 1999, we see much more availability of information on computers. it is highly likely that future research will show a much stronger shift to electronic formats for data access. the findings of this research can provide baseline data for future comparisons. bibliography brittain, j.m. 1970. information and its users: a review with special reference to the social sciences. new york: wiley, 1970. [disiss]. 1979. design of information systems in the social sciences. the structure of social science literature as shown by citations. research report a no.3. bath uk: bath university library. hernon, peter. 1979. use of government publications by social scientists. norwood nj: ablex. hernon, peter, and clayton a. shepherd. 1983. “government publications represented in the social sciences citation index: an exploratory study.” government publications review 10: 227-244. hernon, peter, and gary r. purcell. 1982. “document use patterns of academic economists.” chapter 5 in their developing collections of us government publications. greenwich ct: jai press. [infross]. 1971. investigation into information requirements of the social sciences. information requirements of researchers in the social sciences. research report no.1. 2 vols. [bath uk]: bath university library. line, maurice b. 1971. “the information uses and needs of social scientists: an overview of infross.” aslib proceedings 23: 412-434. nilsen, kirsti. 1996. “the effects of electronic publication on social science researchers in canada.” in: electronic publishing: its impact on publishing, education and reading. pp. 1-26. proceedings of the 24th annual conference, canadian association for information science. toronto, 2-3 june, 1996. toronto: cais. nilsen, kirsti. 1997. social science research in canada and federal government information policy: the case of statistics canada. unpublished doctoral dissertation. university of toronto. umi danq28027 nilsen, kirsti. 1998. “social science research in canada and government information policy: the statistics canada example.” library and information science research 20 (3): 211-234. slater, margaret. 1989. information needs of social scientists. boston spa uk: british library research and development department. weech, terry l. 1978. “the use of government publications: a selected review of the literature.” government publications review 5: 177-184. white, howard. 1982. “citation analysis of data file use.” library trends 30: 467-477. footnotes: 1 based on infross results and including only those in which 60% or more of respondents indicated that they use statistics. 2 economists, social geographers, sociologists and some political scientists also use experimentally derived data, but to a lesser extent. 3 based on infross results and including those in which 20% or more of researchers claim they use government publications. * paper presented at the iassist conference, may 19, 1999, ryerson polytechnic university, toronto, ontario.. kirsti nilsen, faculty of information and media studies,university of western ontario, london, ontario n6a 5b7, canada using new technologies to provide easy access to research databases by andy covell ' manager, research data center syracuse university the research data center, a unit within syracuse university's central computer services organization, provides fee-based programming and data management services for researchers engaged in data-intensive research. research data center analysts routinely develop strategies for managing, processing, and analyzing research databases. mainframe based access to magnetic tape and disk is the prevailing strategy for providing access to large research databases. with recent technological developments a number of alternatives can now be realistically considered, and the research data center is investigating these alternatives. this paper summarizes the information obtained over the last 6 months of that investigation. introduction most research databases are created through the collection, entry and organization of data specifically for research (e.g. survey research) or are derived from data collected outside the research enterprise for reasons other than academic research (e.g. government databases). there are three basic types of research databases: raw files are electronically readable files which are not formatted for any particular software package so they cannot be analyzed directly. raw files are accessed very infrequently, and are usually stored offline once they are read into a master file. masterfiles are data files which have been formatted for some software package (e.g. sas) so they can be easily accessed for direct analysis or to obtain extracts. they are usually static, and they are typically accessed on a regular basis by some community of researchers over an extended period. storing master files off-line is common, although online access is clearly preferable. analysisfiles are data files created for a specific research analysis. they are formatted for some software package and contain the derived variables needed to answer a specific research question. analysis files, the working data sets of a research project, are created, modified, and analyzed with great frequency during analysis. they are rearely accessed once the research is completed. researchers always prefer on-line access to the analysis files which support current research, while the analysis files of previous research can be stored off-line. access to large research databases (raw files, master files, and even analysis files) is often limited commonly access is through mainframe attached magnetic tape drives. recent technological developments can significantly enhance access to large databases. high speed networks, an ever-increasing array of on-line and nearline storage alternatives, and user friendly, flexible database software are rapidly evolving technologies that can provide fast easy access to large research databases. networks high speed networks offer fast access to data stored on remote computers. with a few simple commands researchers can retrieve data from a remote computer, even if it is from a different model computer across the country. research databases no longer need to be written out to magnetic tape to be transferred from one computer to another. over the past five years high speed networks have become a fact of life at us universities. most have spent considerable sums on high speed campus networks, and millions have been spent on regional networks which link campus nets together and a national backbone which links the regionals. one result is the internet, a large conglomeration of interconnected campus, regional and national networks. basic, consistently implemented network services — remote login, file transfer, and electronic mail—are available on all internet computers (except some desktop computers which may lack electronic mail). researchers use the same procedure to access a file whether the file is on a computer located in another building on campus or on an internet computer located in another part of the country. file transfer is provided by a service called ftp (which stands for file transfer protocol). ftp, which is a basic service available on all internet computers, provides broad access to network data resources. unfortunately, ftp sometimes uses network capacity unnecessarily because ftp always transmits a complete file even if only a small extract is needed. 56 iassist quarterly another network service of interest to data-intensive researchers is network file system (nfs) which enables transparent remote file access. with nfs, researchers can access a remote file or directory as if it were on his or her local computer. thus, nfs overcomes the major shortcoming of ftp (nfs does not transmit an entire file). unfortunately, nfs is not yet as widely available as ftp. while ftp and nfs enable network access to data stored on almost all internet computers, they do not automatically convert binary coded numeric information between different computer models. thus, ftp does not support totally transparent data sharing. some database software vendors have solved this problem (see section iv). ftp is a nearly universal tool for providing network access to data resources. but do the internet and local campus networks really work fast enough to support network access to large research databases? or does network traffic and the existence of "weak links" (e.g. low-speed regional network connections that become a data transmission bottleneck) reduce performance to the point that transmitting large research databases is infeasible? to get a realistic assessment of network performance, an informal test was carried out with the help of jim jacobs of the university of california at san diego. a file containing roughly 10 megabytes (an extract of the general social survey) was repeatedly transferred, in roughly two hour intervals, through the national internet (from san diego to syracuse) and through the syracuse university internet (between workstations in separate buildings) on wednesday, may 9, 1990. the data transfer rate was recorded each time the file was transferred. the results of the test, summarized in table 1, indicate that campus networks which run at ethemet speeds, such as the syracus university internet, are indeed suitable for large data file transmission, while long distance transmission via internet is limited— for larger files magnetic tape is still an attractive method. of course within a couple of years the internet's national backbone network and many of the regional networks will be seeing a 24fold increase in performance. the day may soon come when magnetic tapes are used to transfer data only in rare instances. mass storage one of the most striking developments over the past few years in mass storage technology is the proliferation of high capacity storage products. not too many years ago, if one had a large research database there were only two realistic storage alternatives: it could go on mainframe mag tape or, with a little help from the mainframe systems folks, it might be put up on mainframe magnetic disk. today there are a slew of other alternatives suitable for a range of large research database applications— table 1 . results of informal internet data transfer test i£on transfer rate elapsedtime transfer rate elapsed time 6:33am 1 2 kbytes/sec. 14 min. 76 kbytes/sec. 2 min. 11 sec. 9:18am 11 kbytes/sec. 15 min. 74 kbytes/sec. 2 min.15 sec. 1 1 :27am 10 kbytes/sec. 17 min. 77 kbytes/sec. 2 min. 9 sec. 1:21pm 8 kbytes/sec. 21 min. 71 kbytes/sec. 2 min.20 sec. 3:09pm 8.6 kbytes/sec. 19 min. 76 kbytes/sec. 2 min. 11 sec. 5:03pm 8.8 kbytes/sec. 19 min. 76 kbytes/sec. 2 min. 11 sec. 7:47pm 8.6 kbytes/sec. 19 min. 71 kbytes/sec. 2 min.20 sec. 9:08pm 7.8 kbytes/sec. 21 min. 75 kbytes/sec. 2 min.13 sec. 11:24pm 11 kbytes/sec. 15 min. 71 kbytes/sec. 2 min.20 sec. fall/winter 1990 57 from 700 megabyte hard disks which cost a few thousand dollars and can attach direcuy to a pc or workstation to mainframe attached, terabyte capacity (i.e. a thousand gigabytes) optical jukeboxes. see the appendix for descriptions of a wide variety of storage products. with networks enabling high speed access to a heterogenous mix of computers, one is no longer constrained to a particular computer platform when evaluating mass storage alternatives. other characteristics of the application— e.g. required capacity, expected frequency of access, and life expectancy of the data— can be matched against the storage technologies available for a variety of platforms. while there are hundreds of options available for storing large amounts of research data, almost all storage products store data on one of three basic media: magnetic tape, magnetic disk, or optical disk. magnetic tape magnetic tape is the storage technology upon which many research data libraries were built, and it remains a dominant research data storage medium on many campuses. universities made the rather substantial investment to provide tape access some time ago, so researchers have ready access to central tape storage facilities and tape drives. magnetic tapes remain an attractive medium for storing large databases because the main cost is that of new tapes, which are relatively inexpensive. a standard reel of 9-track tape, which holds up to 180 megabytes of data, can be purchased for around $15. ibm mainframe 3480 cartridge tapes, which hold slightly more, cost under $10. tapes are often used to transport databases between institutions. tapes can be packed and shipped overnight, and standard tape formats exist which can be read and written on every major university campus in this country. one of the major problems with magnetic tape is slow data access— the operator intervention required to mount a tape and the sequential processing of magnetic tape results in access time measured in minutes. another problem with magnetic tape is limited archival life. tape is a relatively fragile storage media, with an average archival life of somewhere around five years. those charged with maintaining access to data on tape for extended periods must periodically "exercise" each tape to ensure readability and prevent print-through. compact, high capacity tape products commonly used to back up workstation and minicomputer hard disks have potential as storage media for research databases. 8mm tape store up to 2.3 gigabytes of data in compact cartridge, and an 8mm exabyte drive (exabyte is the only 8mm drive manufacturer although several companies sell exabyte drives) runs $4,000. 4mm tapes, also known as digital audio tapes (dat), are compact cartridges (smaller than a pack of cigarettes) which store 1.3 gigabytes of data. although 4mm tape drives do not have the installed base of the 8mm drives, buyers are attracted to dat because drives are made by more than one vendor. significant increases in the capacity of both 4mm and 8mm tapes are expected. magnetic disk the basics of magnetic disk technology have not changed significantly in twenty years, but continual improvements have resulted in steady increases in capacity and performance and a steady decrease in cost per megabyte. magnetic disks are the obvious choice if high performance on-line access to research databases is required. hard disk are now available which attach directly to pcs or workstations and hold around 700 megabytes of data. they can be purchased for as low as $2,500. for those with greater appetite for local storage, several drives can be daisy-chained together to provide access to several gigabytes of on-line storage. network servers are computers that are dedicated to the task of data access for network client computers. they typically have attached disks that are faster than those attached to individual desktop computers. network server performance is likely to be boosted in the near future with the introduction of raid (redundant array of inexpensive disks) technology. raid servers will achieve much faster transfer rates by transmitting data in parallel, using multiple disks and read/write heads. mainframe disk drives, also known as direct access storage devices (dasd), currently offer the best overall performance. mainframe dasd, which can provide online access to hundreds gigabytes to hundreds of mainframe users, are typically used for demanding timesharing and transaction processing applications. optical disk there are three basic optical technologies on the market today: cd-rom which is primarily a publishing media with data disks created and distributed by information providers, worm disks which enable a single write followed by unlimited read access, and erasable magneto-optical disks which allow unlimited read/write access. digital paper is a newer ultra-high density optical technology which is just becoming available in commercial products. 58 iassist quarterly cd-rom, worm, and magneto-optical are very dense optical disk storage technologies, with the disks typically available as removable cartridges or platters. optical disks are slower than magnetic disks, but the drives are not as susceptible to head crashes and other malfunctions. optical disks also have a lengthy archival life with some vendors claiming up to 30 years. cd-rom disks, developed originally for the audio industry, are 4.7" disks that hold roughly 600 megabytes of data. data is formatted and written on a cd-rom master disk (a process called mastering) which acts as a template for "stamping" copies. cd-rom disks can be mastered for one to two thousand dollars; disks stamped from the master run $2 to $3 per disk. cd-rom is popular medium for distributing textual and bibliographic databases and is beginning to catch on as a medium for distributing research data files, a development boosted by the census bureau's decision to distribute much of the 1990 data on cd-rom. its main attraction is on-line access to large databases from desktop workstations. in many instances, cd-rom is a good alternative to dial-up access to expensive information services such as dialog. however, the general suitability of cd-rom for research database access is questionable. cd-rom is relatively slow when compared to other on-line storage media (e.g. hard disks and worm disks) so it is far from ideal for supporting multi-user on-line access to research databases (though several cd-rom network server packages are on the market). furthermore, the computers which many cd-rom distributors target, often lack the computing resources to effectively handle the analysis files that are typically derived from the large master files distributed on cd-rom. worm (write once read many) is a high capacity, locally written storage media with a lengthy archival life. worm is faster than cd-rom, although worm drives are not nearly as fast as most magnetic drives. one of the striking things about worm technology is the wide range ofworm products. unlike cd-rom, worm is available in several sizes and configurations — from 5-1/4" disks which hold hundreds of megabytes to 14" platters with an 8.2 gigabyte capacity. most work drives act like magnetic disks, although mainframe attached drives typically emulate a tape drive. worm is available as a single drive removable cartridge drive in some products and in jukebox configurations in others. worm products are available for the complete range of computer platforms— from pcs to mainframes. magneto-optical disk, a recently introduced optical technology, is a high capacity, fully erasable, removable storage media. it shares many of the properties of worm drives (similar capacity to store data, comparable performance, long archival life), but it is not available in such a wide range of products— 5-1/4" disks which can store several hundred megabytes are the norm. digital paper is an ultra-high write-once optical media which can be produced in large sheets and reels. only one product using digital paper is on the market right now (the creo 1003 tape drive which stores a terabyte of data on a single reel of tape); and one product that was being planned has been dropped (a bernoulli drive based on optical paper). the future of digital paper is unclear, but if the technology takes off, it could become the storage media of the future. database sofware database software packages enhance database access by relieving the end-user from the burden of knowing the physical characteristics of each variable, for example where each variable is physically located, how long each variable is, and so forth. once a raw file has been read into a database package, end-users can simply access variables by name, usually with some a flexible, userfriendly query language. this is an excellent approach for master files which are used by groups of researchers. database access can also be enhanced by using the indexing capability built into most database software. a database index is essentially a computerized lookup table that speeds data access, similar to the way the index in the back of a book works. by indexing the variables which are frequently sorted on or used to select extracts, researchers can realize significant time savings. the ability of some database packages to provide transparent access to remote databases is another way database software can enhance access to research databases. several database packages (e.g. ingres and oracle) advertise remote access to data on different platforms through high speed networks. sas will soon offer an add-on product, called sas connect, which will provide the same capability for sas datasets. 1 paper presented at the 16th annual conference of the international association for social science information service and technology (iassist), poughkeepsie, new york, may 30-june 2, 1990. fall/winter 1990 59 19.3 12 iassist quarterly the incore server was set up by the university of ulster and the united nations university to act as a central resource for academics, policy-makers and others concerned with conflict resolution and ethnicity across the globe. the server utilises internet resources to make as widely available as possible information on ethnic conflict. however, like a lot of tools the internet can be used for more than one purpose. in one sense it is a very flexible and open resource that allows users to log into numerous systems around the world, transferring files, searching file systems and executing programs. however once you are connected to the internet it is also a way for other people to get into your system. if you restrict your workstation’s access only to parts of the internet , then you may find that there is a lot of useful resources and facilities that you can no longer access and use. in order to take advantage of the internet you must be a part of it. however, as this paper will try to demonstrate, by doing this you put your computer at risk, so you need to constantly keep abreast of the security holes in the network software installed in your system and protect it. background incore [2] was established by the university of ulster (uu) and the united nations university (unu) to provide a systematic approach to the study of ethnic conflict and to encourage links between research, training, policy, practice and theory. the incore metadatabase server [3] was set up to act as a central resource for academics, policy-makers and others concerned with conflict resolution and ethnicity across the globe, but particularly to those operating in areas of conflict. it is central to incore’s mandate that it serve those who have most difficulty in getting access to information. in this respect the internet can be both the most and least useful medium for connecting people to information. the advantage of placing information on the internet is that it is almost instantaneously available, uncensored, to anyone anywhere in the globe who is connected to the net. the disadvantage is that, unless one is connected to the internet, which involves expensive hardware, technical expertise and effective maintenance, this information is inaccessible. incore has to be concerned not merely to provide a state of the art server but also to ensure that as much of the information as possible on the server is accessible to those with the least connectivity to the internet, through e-mail and ftp as well as through gopher and world wide web. however in setting up this server we need to be aware of adopting a responsible attitude towards security in order to maintain a continued uninterrupted service. we need to consider where our system is at risk [4], where the possible threats and breaches of security are coming from and how to prevent these becoming problems. introduction as the volume of internet use increases dramatically (30,000 plus interconnected networks with 2.5 million or more connected computers daily swap gigabytes of information based on nothing more than a digital handshake with a stranger) [5, 6] and as we connect our organisation’s networks to thousands of other computer networks, serious security issues arise. there are a lot of considerations for the system administrator and the following list outlines some of the main issues : _ unauthorised users accessing data and information on our system _ authorised users causing inadvertent or malicious damage _ interception of data transmission _ virus attacks _ backup and storage policy it is difficult to fully understand and appreciate just how much the internet depends on collegial trust and a group effort by the whole community. one technique, for instance, that intruders use is to break into a chain of computers, e.g., break into a, use a to break into b, b to break into c, etc., thus covering their tracks. so, it would be very unwise and foolish to think that your little part of the network is safe because you believe that there is very little information there of a nature that would entice someone to break into it. even if there is nothing of use on your computer, it could prove a worthwhile intermediate staging post for someone who wants to break into other more useful systems. incore metadatabase server technical aspects and constraints by patrick curran1 incore, university of ulster, n. ireland. 13fall 1995 in the early days of the arpanet only researchers had access to the net, and they shared a common set of goals and ethics. data packets were forwarded along network links from computer to computer. a packet may have made a number of hops and every intermediate machine could read its contents. nowadays many internet packets start their journey on a local area network (lan), where privacy is even less protected. only a gentleman’s agreement assures the sender that the recipient and no-one else will read the message. the lack of security on the arpanet did not bother anyone, because that was part of the package. as the internet developed and expanded, the user population began to change, with a lot of the newcomers having little idea of the importance of the complex social contract guiding the use of this new and exciting tool. nowadays anyone with a computer, a modem and a small monthly connection fee can have a direct link to the internet and be subject to breakins or launch attacks on others. every day computer networks and hosts are being broken into with varying levels of sophistication. while it’s generally believed that most break-ins succeed due to weak passwords, there are advanced and sophisticated techniques that are more difficult to detect. the article looks firstly at the shortcomings of the password mechanism (concentrating on the unix system) and then discusses the more sophisticated techniques available to intruders, including some examples of the misuse of commonly used internet resources, e.g., ftp, sendmail, telnet, rlogin, rsh, etc. it also analyses the range of responses to network intrusion techniques, from software policing solutions like kerberos and cops to the hardware solution of firewall installation. security breaches system administrators can safely configure workstations on their network to allow connections to other workstations. they can also set up their network file system to export widely used file directories to “world”, allowing everyone to read them. it doesn’t take much imagination to see what can happen when such a trustworthy environment opens up its digital doors to the internet. suddenly, ‘world’ means the entire globe and “any computer on the network” means “every computer on any network”. there have been a lot of computer security problems and breaches of security in recent years, some more serious than others. some of the more widely known incidents include break-ins on nasa’s span network [7] and the ibm ”christmas virus”, but the most widespread breach of network security occurred in 1988 when the internet came under attack from within, later to be known as the internet worm incident. internet worm on november 2, 1988, a self-replicating program, called a worm* appeared on the internet [8] . this program copied itself from machine to machine, causing the machines it infected to labour under huge loads, and denying service to the users of those machines. the program spread quickly, and while many system administrators were aware that something like this could theoretically happen (the security holes exploited by the worm were well known) the scope of the worm’s break-ins came as a great surprise to many. the worm itself did not destroy any files, steal any information (other than account passwords), intercept private mail, or plant other destructive software [9]. however, it did manage to severely disrupt the operation of the network. several sites, including parts of mit, nasa’s ames research centre and goddard space flight centre, the jet propulsion laboratory, and the army ballistic research laboratory, disconnected themselves from the internet to avoid recontamination. in addition, the defense communications agency ordered the connections between the milnet and arpanet shut down, and kept them down for nearly 24 hours [10]. ironically, this was perhaps the worst thing to do, since the first fixes to combat the worm were distributed via the network. this incident was perhaps the most widely described computer security problem ever. the worm was covered in many newspapers and magazines around the country including the new york times and most computer oriented technical publications. security considerations the incidents above demonstrate quite clearly that computer security is an important topic. every day computer networks and hosts are being broken into with varying levels of sophistication. when you hear the term “security” the first thing that comes to mind is the password. however, while it’s generally believed that most break-ins succeed due to weak passwords, there are, moreover, a large number of unauthorised attacks that use more advanced and sophisticated techniques. less is known about these latter techniques because they are more difficult to detect. we will look firstly at the advantages and disadvantages of the password mechanism (concentrating on the unix system [11]) and then discuss the more sophisticated techniques available to intruders. 14 iassist quarterly passwords an underlying goal of the unix system has been to provide password security at a minimal inconvenience to the users of the system. for example, those who want to run a completely open system without passwords, or to have passwords only at the option of the users, can do so, whilst those who require all their users to have passwords gain a high degree of security against penetration of their system by unauthorised users. a good password system must be able not only to prevent access to the system by unauthorised users, but it must also prevent users already logged on from doing things that they are not authorised to do. the “super user” password, e.g., is especially critical because it allows all sorts of permissions and provides unlimited access to all system resources. passwords are important because they are generally the first line of defence against interactive attacks. in simple terms, if a cracker@ cannot interact with your system, and he has no access to read or write to the information contained in the password file, then he has almost no avenue of attack left open to break your system. unfortunately the unix passwd program [12] doesn’t place a lot of restrictions on what might be used as a password. it generally requires 5 lowercase letters ( or 4 characters) but if the user insists on using a shorter password (by entering it three times) the program allows it. the object when choosing a password is to make it difficult for a cracker to make educated guesses, thus leaving him with no alternative but to try every possible combination of letters, characters, special characters and numbers. a search of this sort, even on the most powerful computer would take at least 100 years to complete. robert morris and ken thomson carried out an interesting survey determining typical users’ habits in the choice of passwords [13]. out of 4,000 passwords 16% contained 3 characters or less, and 86% were what could generally be described as insecure. grammp & morris [14] in another experiment showed that by trying the 20 most common female names, followed by a single digit, at least 1 password was valid in each of the machines surveyed. they also found that by trying variations of the login name, user’s first and last name, and a list of nearly 1800 common first names, that up to 50% of the passwords on any given system can be cracked in a matter of 2 or 3 days. there are ways to improve the security of using passwords. according to schweitzer [15] a high quality password has the following characteristics : at least eight characters length the password is randomly generated the password has no personal relationship to the user, his family or job the password must be kept secret it must be changed at least every three months the internet worm, in trying to break into new systems attempted to crack user passwords [7]. first of all it tried simple choices (user names, names, etc) and then tried an internal dictionary of 432 words. if all failed it tried going through the system dictionary /usr/dict/words, trying each in turn. so, by adhering to the guidelines above, you will make it very difficult for an intruder to break into your system using the first line of defence. other problems associated with passwords are : expired accounts : accounts lying around due to people leaving the organisation. these cause problems because since nobody is using the account anymore it’s unlikely that a break-in will be noticed. a simple preventative measure is to place an expiration date on every account. if there is any doubt about an account, then replacing the encrypted password with an asterisk (*) will make it impossible for anyone to log into the account. guest accounts : these are usually made available for expediency. they should only be used for the period required and deleted immediately. it’s important not to give them simple passwords, e.g., “guest” or “temp”. network security besides cracking passwords on machines crackers have other means at their disposal, the success of which depends on your awareness and adoption of preventative measures. it’s important to be careful about techniques that bypass password requirements. there are two common ones, i.e., the .rhosts and the hosts.equiv files : hosts.equiv: the file /etc/hosts.equiv can be used to indicate trusted hosts. if a user remotely logs into your system using rlogin and does so from a host listed in this file, access is permitted without requiring a password. the default 15fall 1995 file has the entry '+’ in a single line indicating that every host should be considered a trusted host. this could prove a major security problem since hosts outside the local organisation should never be trusted. only specific host names should be included in the file, or the file deleted altogether. .rhosts: allows access to specific host-user combinations. each user may create a .rhosts file in his/her directory, allowing access to their account without supplying a password. for example, the entry host.com fred tells the computer on which the file resides to bypass password requirements when it sees someone logging in from the login name “fred” on the machine “host.com”. this means, obviously, that anyone who manages to break into fred’s account can also break into this machine. the internet worm made use of the trusted host concept to spread itself throughout the network [8]. secure terminals unix introduced the concept of a “secure” terminal, which prohibits ‘root’ from logging in from a non-secure terminal. the file /etc/ttyab controls which terminals are considered secure. the default is to consider all terminals secure allowing root to login from anywhere in the network. a more secure configuration would be to consider as secure only directly connected terminals, or only the console device. the most secure method is to remove the secure designation from all terminals including the console, so users with ‘root’ privilege must first login as themselves and then use the ‘su’ command. nfs the network file system (nfs) allows several hosts to share files over the network, generally used to provide file server access to diskless workstations in a small network. the file /etc/exports lists which file systems are exported. this file contains access specifications, e.g. : root=keyword specifies the list of hosts allowed "super-user" access to the files in the named file system access=keyword specifies the list of hosts that are allowed to mount the named file system. for example, the line /export/root/client -access=client,root=client allows the host “client” to access the named file system with root privileges. if the file isn’t properly configured then anyone on the internet may have access to your files via nfs, whether you trust them or not. e-mail whilst this is one of the net’s basic services, forging e-mail is a trivial exercise. an electronic letter consists simply of a text file with a header specifying the sender, receiver, subject, date and routing information, followed by a blank line and the message body. mail programs fill in the header lines with routing information, but there’s nothing to stop a malicious person inserting whatever he wants into the mail. protocols have been developed to verify the source of e-mail messages, but spoofers are also improving their techniques. some systems bar connection to their system from untrustworthy parts of the internet. the problem with this strategy is that the trusted “domain name servers” they rely on are just ordinary computers, and as such are also vulnerable to deception or intrusion. a cracker can modify the name server’s database so that it tells any computer querying it that the address belonging to e.g.,”cracker@breach.com” is instead that of”president@whitehouse.gov”. a computer allowing connection from whitehouse.gov will allow the cracker in as well. the “sendmail” bug has reappeared time and time again over the years, due to the fact that most mail programs make it possible to route messages not only to users but also directly to particular files or programs. people, e.g., forward mail to a program called “vacation” which sends a reply telling the correspondent that the recipient is out of town. other people route mail through filter programs that forward the message to various locations. this same mechanism can thus be subverted to send e-mail to programs that are designed to execute shell scripts. such a script could cause a copy of the receiving computer’s password file to be sent to an intruder for analysis, or simply wreak havoc on the recipient’s file system. intrusion techniques many system administrators are often unaware of the dangers presented by anything beyond the most trivial attacks. the purpose of this section is to present a few of the techniques available to the system cracker in his efforts to gain access to a shell# process on a unix host [16]. there are so many methods and techniques that it would be impossible to cover them all in this paper. however, i will try to outline some of the most common techniques. fig. 1 illustrates some of the more useful services and facilities that an intruder may avail of to ferret out information about a system and set up an interactive link with that system. 16 iassist quarterly fig. 1 : network intrusion techniques the finger services, provided by the finger program outputs information about the users logged on to the system, and the fingerd program extends this facility to remote hosts, e.g., host% finger@incore.ac.uk login name tty idle when where pat p.curran co 21 fri 10:38 i inf.ac.uk the most revealing information divulged are the account names, home directories and the host they last logged in from. this information can be supplemented by using the rusers command (with the -l flag), which produces, e.g. : login home-dir shell last login, from where ——— ————— ——— ———————————— root / /bin/sh fri apr 15 from inf.ac.uk pat /home/pat /bin/csh on since thur apr 14 from inf.ac.uk guest /export/guest /bin/sh never logged in ftp /home/ftp never logged in finger is one of the most dangerous services because it is very useful for investigating and receiving information about possible target machines, especially when used in conjunction with other data. a bug in the fingerd program was exploited by the internet worm to overrun the buffer that the daemon (background process that fingerd is intended to run as) used for input, thus altering the behaviour of the program. the showmount command can be used on an nfs fileserver to display the names of all the hosts that currently have something mounted from the server. running showmount on a target, e.g., reveals : host% showmount -e incore.ac.uk export list for incore.ac.uk /export (everyone) /var (everyone) since /export/guest is exported to the world, and this is the user guest’s home directory, we have a possible break-in scenario. the intruder could mount the home directory of ëguest’, and create a ëguest’ account in his local password file. by putting a 17fall 1995 .rhosts entry in the remote guest home directory he allows himself to login to the target machine without having to supply a password. if the target machine has a ‘+’ wildcard in its /etc/hosts.equiv file, then any non-root user with a login name on the target’s password file can rlogin to the target without a password. so the intruder’s next line of attack would be to login to the target host and modify the password file to allow root access : host% rsh incore.ac.uk csh -i host% ls -ldg /etc drwxr-xr-x 8bin staff 2048 jul 24 18:02 /etc host% cd /etc host% mv passwd pw.old host% (echo attack::0:1:instant root shell:/:/bin/sh;cat pw.old)> passwd host% ^d host% rlogin incore.ac.uk -l attack welcome to the incore server !! incore# the rsh -i gets one on to the system but doesn’t leave any traces in the wtmp or utmp system auditing files, making the remote shell invisible to the finger and who commands. going back to the finger output we can see that there is an ‘ftp’ account, which usually means that anonymous ftp is enabled. the file transfer protocol (ftp) allows users to connect to remote systems and transfer files back and forth. this can be an easy way to get access to a system as it is often misconfigured. in some cases the target machine has a complete copy of the / etc/passwd file in the anonymous ftp ~ftp/etc directory. if, however, this isn’t the case then there are other avenues for the would-be attacker. if, for instance, the ftp directory is writable, an intruder can remotely execute a command, e.g., mailing the password file back to himself simply by creating a .forward file that executes a command when mail is sent to the ftp account. if none of the methods above have worked then the attacker could use rpcinfo to see if the target is running nis or nfs. once you know the nis domainname of a server you can get any of its nis maps using a simple rpc query. also, just like easily guessed passwords many systems use easily guessed nis domainnames [17]. the showmount output usually divulges information on domainnames, which can then be checked with the ‘ypwhich’ command to see if the domainname exists, e.g. host% ypwhich -d incore incore.ac.uk domain incore not bound this proved unsuccessful, but if it was guessed properly it would have returned with the hostname of incore.ac.uk’s nis server. if an attacker has control of the nis master then he effectively has control of the client hosts, and can execute arbitrary commands, e.g., mailing password files to himself. security tools there are a lot of tools available that deal with the security management aspects of computer systems, more than can be adequately covered in this article [18]. because there are so many tools and techniques available to implement security controls you should, in the first place, identify the requirements of your network service and what you are willing to accept. the following sections provide only a flavour of the types of solutions available, from software based system auditing, authentication and checking techniques to the hardware solution of firewall installation. computer oracle and password system (cops) cops [19] is a unix security status checker, written as a suite of shell scripts which forms an extensive security testing system. basicially it checks various files and software configurations to see if they have been compromised (e.g., edited to plant a trojan horse& or back door$ ), and checks to see that files have the appropriate modes and permissions set to maintain the integrity of your security level (making sure that your file permissions don’t leave themselves wide open to attack). there’s a rudimentary password cracker, and routines to check the file store for suspicious changes in setuid programs, and to identify software behaving in ways which could cause problems. 18 iassist quarterly the current version of cops makes a limited attempt to detect bugs that are posted in cert advisories. also, it has an option to generate a limited script that can correct various security problems that are discovered. kerberos kerberos [20, 21] is a des-based encryption scheme that encrypts sensitive information, such as passwords, sent via the network from client software to the server daemon process. when a user logs in, kerberos authenticates that user (using a password), and provides the user with a way to prove her identity to other servers and hosts scattered around the network. this authentication is then used by programs such as rlogin to allow the user to log in to other hosts without a password (in place of the .rhosts file). the authentication is also used by the mail system in order to guarantee that mail is delivered to the correct person, as well as to guarantee that the sender is who he claims to be. the overall effect of installing kerberos and the numerous other programs that go with it is to virtually eliminate the ability of users to “spoof” the system into believing they are someone else. firewall a firewall is a machine which is usually attached between your site and a wide area network. it provides controllable filtering of network traffic, allowing restricted access to certain internet port numbers (i.e., services that your machine would otherwise provide to the network as a whole) and blocks access to pretty well everything else. they are an effective “all-ornothing” approach in dealing with external access security, and are fast are becoming very popular, particularly with the rise in internet connectivity [22]. the firewall doesn’t send out routing information about the internal network, making the internal network “invisible” from the outside. it doesn’t advertise routes which means that users on the internal network must log in to the firewall before accessing hosts on remote networks. also, in order to remotely log in to a host on the internal network from the outside, a user must first log in to the firewall machine, which may prove to be inconvenient, but, nevertheless, more secure. outgoing e-mail is forwarded to the firewall machine before being delivered outside the internal network and the firewall receives all incoming email, before redistributing it. it provides extra security by not mounting any file systems via nfs, or making any of its file systems available to be mounted. password security is rigidly enforced and the firewall does not trust any other hosts regardless of where they are. finally, anonymous ftp and other similar services is only provided by the firewall host, if at all. the purpose of the firewall is to prevent crackers from accessing other hosts on your network which means that security must be strictly and rigidly enforced. it is important to remember that a firewall can’t provide complete safeguards against intrusion if someone manages to subvert the firewall then he can subsequently break into any host behind it. many organisations, and more recently universities, have adopted firewalls to examine the packets entering and leaving a domain and to restrict certain internet connections. however, proposing a firewall and constructing it are two different things entirely. there are some things that you just can’t do securely. gopher and mosaic, for instance, are two programs of a trusting nature that defy the attempts of a firewall design to provide safety. additionally, a firewall must pass mail and fig. 2 : firewall set-up 19fall 1995 mailers can be very insecure. users also need to log into log into various public archive sites to receive files. one solution to provide this type of functionality is outlined in fig. 2, where two dedicated computers or gateways are used one connected to the local network and the other to the internet : the external computer examines all incoming traffic, forwarding only “safe” packets to the internal machine. an attacker , thus, could only break into machines on the local network by compromising the internal gateway. the internal gateway only accepts messages from the external computer, so that, if unauthorised packets do reach it they will not be able to pass. conventional password techniques, when used with firewalls, reduce the efficiency of the overall security provided. hackers gained access to panix, a public-access internet site in new york (oct. 1993), and installed “packet sniffers”. these programs watched data going by and recorded user names and passwords as people logged on to hundreds of other computer systems. connections, therefore, within a firewall require, for ultimate security, a different kind of authentication mechanism that cannot be recorded by sniffers, e.g., “one-time password” or “challenge-response password”. conclusions even if newcomers to the internet try to secure their systems it’s not always easy to find the information they need. hardware and software vendors are usually loathe to discuss their security problems. cert [23] generally issue advisories only after manufacturers have developed a definitive fix usually weeks or months later. spafford [10] points out that people don’t know the risks between half and three quarters of the security holes currently known to hackers have yet to be openly acknowledged. people only know the benefits. many of these benefits come from programs such as gopher, netscape, world wide web or mosaic, which allow simple menu-driven navigation of the internet. these tools are drawing thousands of people to the net. yet, the rapid evolution of these tools have bypassed steps that could lead to security breaches. the popular gopher problem, for instance, according to cert advisories, has security problems that make it possible to access not only public files but private ones as well. again, gopher is only insecure if it is misconfigured. whilst gopher servers can be confined so that they have access only to public information, by default they have a free rein. the internet attracts more recruits every day hoping to reap such rewards and benefits as connections with other people and organisations, file access and information interchange. yet, internet connection can also prove to be the source of very real risks and dangers to your workstation or network. while this article, hopefully, raises the reader’s awareness of the importance of security in a network situation, one shouldn’t get paranoid. it’s worth remembering that security, in most cases, is an elusive goal. don’t forget as well that the nature of unix and the internet helped to defeat the internet worm as well as spread it. the sensible approach is to secure your system according to its needs, keeping danger at a manageable level. in other words, don’t stop travelling but do wear a seat belt. references [1] paper presented at iassist95 may 1995 quebec city, quebec, canada. [2] “data on ethnic conflict and conflict resolution : the work of incore”, journal of ethnodevelopment, forthcoming publication [3] metadata on ethnic conflict and conflict resolution incore metadatabase project, iassist quarterly, forthcoming publication [4] kay, r. : “distributed and secure”, byte, pp.165-178, (june 1994) [5] mids, :”mids press release : new data on the size of the internet and the matrix”, matrix news, vol. 5, no. 1, 1.5. matrix information and directory services, inc., austin, (january 1995), available from mids@tic.com [6] lottor, dns zone survey (rfc 1296) through january 1995 [7] mclellan, v. : “nasa hackers : there’s more to the story”, digital review, p. 80. (nov. 1987) [8] spafford, e.h. : “the internet worm program : an analysis”, computer communications review, 19, no. 1, acm sigcom, (january 1989) 20 iassist quarterly [9] seeley, d. : “a tour of the worm”, proceedings of 1989 winter usenix conference, usenix association, san diego, california, (february 1989) [10] eichin, m.w., rochlis, j.a. : “with microscope and tweezers : an analysis of the internet virus of november 1988”, proceedings of the symposium on research in security and privacy, ieee-cs, oakland, california, (may 1989) [11] garfinkel, s., spafford, s. :”practical unix security”, o’reilly & associates (1992) [12] todino, g., strang, j., peek, j. : “learning the unix operating system”, o’reilly & associates (1991) [13] morris, r., thomson, k. : “password security : a case history”, communications of the acm, 22(11), pp. 594-597, (november 1979). reprinted in unix system manager’s manual, 4.3 berkeley software distribution, university of california, berkeley (1986) [14] grammp, f.t., morris, r. : “unix operating system security”, at&t bell laboratories technical journal, 63 (8), pp. 1649-1672, (october 1984) [15] schweitzer, j.a. : “managing information security. administrative, electronic and legal measures to protect business information”, 2nd. ed. massachussetts, butterworth publishers (1990) [16] farmer, d., venema, w :”improving the security of your site by breaking into it”, available by ftp from win.tue.nl as /pub/security/admin-guide-to-cracking.z [17] schuba, c. : “addressing weaknesses in the domain name system protocol”, purdue university (august 1993) [18] reinhard, r.b. :”an architectural overview of unix network security (specifically oriented towards internet connectivity)”, v.4 (feb. 1993), available from gopher.near.net in /security/papers [19] available via anonymous ftp from cert.org in ~/pub/tools/cops. [20] available via anonymous ftp from athena-dist.mit.edu in ~/pub/kerberos [21] bellovin, :”limitations of the kerberos authentication system”, available via anonymous ftp from research.att.com [22] ranum, m. :”thinking about firewalls”, proceedings of the 2nd. international conference on systems and network security and management, available via anonymous ftp from ftp.tis.com as /pub/firewalls/firewall .ps.z [23] the computer emergency response team (cert) advisories are available by ftp from cert.org footnotes * a worm is an independently operating program that actively propogates by spreading copies of itself throughout a network. @ in usenet parlance a ”hacker” is a person possessing a great deal of knowledge and expertise, and exercises this with great finesse, whereas a “cracker” is a person who persistently breaks into other people’s computer systems, for a variety of reasons. for further information refer to steele, g.l., woods, d.r., finkel, r.a., crispin, m.r., stallman, r.m., goodfellow, g.s. :”the hacker’s dictionary”, new york: harper and row, 1988. # a shell script is a unix program containing a series of commands that perform system functions. & a block of undesired code (intentionally hidden within a desireable block of code) which does things that the user does not intend, e.g., a program that simulates a computer’s logon procedure, but, rather than logging the user on, it simply records and steals the user’s id and password. $ a feature built into programs by the designers allowing them special privileges which are denied to the normal users of the program, e.g., a back door in a logon program would enable the designer to log onto a system without an authorised account. iassvol201 10 iassist quarterly bringing census data into the classroom: world wide web access and teacher networking by louis r. gaydosh1, the william paterson college of new jersey a social science framework for technology assessment. in order to carry out a complete assessment of the impact of computing technology in the modern world, social scientists should employ an organizing framework within which to place studies addressing the issue. such a schema should be comprehensive in order to allow for research across the entire range of affects which this technology can and does have on all societies in which it is found. at the same time, it should be simple enough to allow for a parsimonious explanation of the cumulative evidence regarding the consequences of computers on social structure and behavior. a paradigm satisfying both criteria would provide a common context for comparing different analyses and would allow for building a coherent body of knowledge on the subject. in this paper, i would like to suggest a frame of reference which provides for an inclusive yet uncomplicated organization of studies assessing the social impact(s) of computing technology and which can be used to generate hypotheses for further research in this area. we begin with the observation that technology can have two types of affects on human societies. one affect is quantitativethat is, it can influence the amount of activity(ies) in which people are involved, the time it takes to perform tasks, the economic costs and/or benefits of developing and/or adopting the technology, etc. the other type of affect is qualitative, which includes such factors as the types of social structures which encourage and eventually adopt technological innovation, the consequences of adopting the technology for the social environment, the moral, ethical, and legal implications of technology for the people and groups in society, etc. at the same time, the scale of the impact(s) of technology on society can be seen at two levels of analysis: a macrolevel and a micro-level. macro-level affects ramify through the entire society and/or its major institutions. micro-level affects are felt by individuals, either singly or in the interpersonal relationships in which they are involved. the intersection of these two independently variable dimensions produces a two-by-two table, in the cells of which any particular study assessing the impact of computer technology can be arranged. before applying this framework to the assessment of computing technology, we should point out a major pitfall of all such schemata. this formulation implies that there are categorical distinctions between both types and levels of affects which technology has on social structure. it is probably more valid to conceptualize these categories as endpoints of continua. kaplan (1964) comments on the differentiation between “quantitative” and “qualitative” phenomena that: in general, even if we are working with qualitative variables, the frequencies of their occurrence may be of importance to our inquiry, and these constitute a corresponding set of quantitative variables. similarly, the reliability of a classification into qualitative categories may itself be a quantitative matter. no problem is a purely qualitative one in its own nature; we may always approach it in quantitative terms. ( emphasis added) (176) for example, the term “information anxiety” has a primarily qualitative connotationreferring to a disorientation or malaise people feel when confronted with the overwhelming amount of massproduced information to which they are exposed in contemporary society. (wurman, 1989) however, as social scientists we may also determine the number of people who experience this conditionthis is its quantitative dimension. a similar criticism could be made of the distinction between macroand micro-level affects of technology on society.that type of affect scale of affect quantitative qualitavive macro-level micro-level 11spring 1996 is, phenomena which may be thought of as characteristic of entire societies penetrate the experiential world of individuals and small groups. people draw on such culturally “universal” concepts as gender, space, and time to structure their everyday roles, identities, and relationships with others. (see robertson [1987], or any other introductory sociology textbook for examples and summaries of supporting evidence.) having indicated the strengths (comprehensiveness and ease of understanding) and weaknesses (oversimplification of differences between types and scale of affects) of the proposed framework, we can conclude that it has value as a heuristic device. its principal virtue is that it presents a methodology for organizing research on the assessment of computing technology into a manageable number of categories which allow for the systematic development of a theory of the impact of technology on social structure. the framework applied to computing technology. when applied to the impact of computing technology on society, quantitative affects refer to the sheer amount of information generated and made available to people, the rate of growth of information and knowledge, the number of people who work in information industries, etc. as these are increased and/or accelerated by computers.qualitative affects include such phenomena as the structural characteristics of information societies, as well as the “quality of life” issues which have been raised in analyses of the impact of computers on employment, privacy, health, etc. as stated previously, macro-level impacts have societywide ramifications, including, in the case of computing technology, the transformation of contemporary social structure into an information society. individuals or small groups are the locus of micro-level affects of computers, as exemplified by the creation of “newsgroups”, “forums”, and other electronic channels of communication between people. in addition to providing a framework for organizing studies, this schema can also be used to generate hypotheses for research which can contribute to a theory of the effect(s) of the computer on the information society. the following figure is the previous table filled in with concepts summarizing potentially fruitful areas of study and/or generating hypotheses which can advance social scientific knowledge in this endeavor. the following sections offer commentaries on the types of studies which could be generated by this scheme. quantitative macro-level affects: gross information product (gip). information can be regarded as a commodity in the contemporary world. rosenberg (1992) notes that “information [is] a commodity ... a product in its own right”. (328) it is an object of commercial exchangethat is, it is produced, bought, sold, and used just as any other commodity. (an important distinction between information and other commodities is that information is not consumed [in the sense of being eliminated], destroyed, or reduced after it is used.) given the central importance of information and its (at least possible) commoditization in post-industrial society, it is appropriate that we develop some measure of how much of it we produce. wurman (1990) states that “[t]he amount of available information now doubles every five years...” (32) but, it is virtually impossible to find a precise quantitative measure (or estimate) of information available to people in the united states or any other contemporary society. in an attempt to provide such a gauge, i propose a concept termed “gross information product (gip)”, by which is meant the total amount of all final information output produced and disseminated in an information society during a given period of time, for example, each year. the concept is analogous to the “gross domestic product” (gdp) produced in an economy each year. gdp is defined as the “total market value of all final goods and services produced in an economy during a year”. (miller, 1994:170) however, it should be understood that gip is not intended as an assessment of the monetary value of information (for a methodology for developing such a measure, see porat [1977]), nor of its utility in facilitating decision-making, stimulating further research, etc. (these considerations are not unimportant; however, their use as components of a quantitative measure of information introduces complications which extend beyond the scope of this paper.) a society’s gross information product is nothing more than the total quantity of information which it generates in a year, irrespective of its economic worth or its practical consequences. nature of affect scale of affect quantitative qualitavive macro-level gross information structural properties product gip of information society micro-level personal productivity techno-social measures construction of reality 12 iassist quarterly i propose this measure in response to a “nagging” concern. in an industrial economy based on the manufacture of products, it is possible to ascertain the total number of units of output in any particular sector. for example, we can determine the number of automobiles, televisions, washing machines, etc. produced by companies in the united states, japan, united kingdom, etc. for a given year (or virtually any other time period). it seems fitting that we should have a comparable measure of output for a society in which the production, storage, and dissemination of information is a principal economic activity. however, like many indices, gip is more easily proposed than constructed. a major source of difficulty in constructing a measure is the multiplicity of forms in which information occurs in contemporary society. a starting point can be found in porat’s (1977) distinction between the primary and secondary sectors of “information activity”, defined as “the production, processing, and distribution of information goods and services” (24) in the economy. porat’s concept of primary information activity refers to any “good or service [which] intrinsically convey[s] information or [is] directly useful in producing, processing, or distributing information”. (porat, 1977:25) we might adapt this definition to gip to mean information goods or services made available to the general public, regardless of cost to either the producer or the user of the information. for example, most sites on the world wide web are freely available to anyone with a computer, modem, and web browser. research on the “primary sector” component of gip would include compilations of the number(s) of any or all of the following: 1) software programs, multimedia presentations, world wide web home pages, and other output intended for demonstration or dissemination via electronic media; 2) books, monographs, journal/magazine articles, reports, and other “hard copy” publications (including works of creative writing, such as novels, plays, etc.); 3) mass media broadcasts and productions. this enumeration includes references to information presented directly through the computer, in print, and over television and radio broadcast(s). the role of the computer in the production and/or distribution of information via these media is evidentit is the principal, if not the only instrument employed in this enterprise. the “secondary sector” of information activity is defined as “all the information services produced for internal consumption by government and noninformation firms”. (porat, 1977:4) the notion underlying this concept is that there are many individuals and organizations whose principal output is the production of goods and/or the provision of services which are not directly and immediately informational, but who still generate and rely on information to carry on their enterprise(s). the extent to which computing technology is utilized in this component of gip might be assessed through such indicators as: 1) the number of inter-office memoranda and other documents intended for circulationwithin an organization distributed via computing technology (e-mail, fax machines, “floppy” diskettes, etc); 2) the number of documents (letters, reports, etc.) exchanged between individuals and/or organizations via computing technology (e-mail, fax machines, “floppy” diskettes, etc.). this secondary component of gip points up the fact that computing technology is brought to bear on the production and dissemination of information within and between entities whose main objective is not to generate publicly available information, but which nevertheless create and exchange information in the course of their routine activities. together, these two components make up a society’s gross information product. what is called for is a composite measure of the total amount of primary and secondary information produced in a society in a specified time period. qualitative macro-level affects: structural properties of the information society. in the category of “structural properties of the information society”, i include analyses of the social structure of contemporary information societies, including examinations comparing this societal type with other social forms, such as agricultural societies, industrial societies, etc. any such discussion should begin with an acknowledgment that there is not a universal consensus among social scientists that what is commonly referred to as the “information society” is a distinct type. (kumar, 1995) bell (1981) estimated that, in 1970, only 28.6 percent of the civilian work force in the united states worked in the “industry sector” of the economy, while fully 46.4 percent worked in the “information sector” and another 21.9 percent were in the “service sector”. on the basis of these figures, bell concluded that we have become a “post-industrial society” and, given the ascendance of the information sector, it seemed appropriate to use the term “information society” to describe the new social structure. an alternative interpretation which has been proferred is that what is termed the “information society” is simply an adaptation of capitalism to a social context in which industrial production has been replaced by information generation and dissemination. (kumar, 1995) among other observations, proponents of this viewpoint have pointed out that the socalled information society is characterized by concentration of capital in a few corporate structures (microsoft has 13spring 1996 replaced general motors as the symbol of economic success), just as occurs in an industrial economy. while it may still be an open question as to whether the information society represents a new and different social structure, contemporary developments have brought about certain changes in our social behavior(s). these changes have the potential to react back on society and have an affect(s) on its structural characteristics. i have in mind one relatively indisputable fact: the dominant locus of social activity in the information society is the household. information technology, directed by a whole host of big business interests, has been increasingly put at the service of home-based consumption. entertainment is the most obvious example. ‘going out’ has been replaced by ‘staying in’. (kumar, 1995:155) television, vcrs, audio cassette players, and cd changers are obvious technological appliances which allow people to bring various forms of entertainment into their homes. computers can also contribute to this phenomenon by enabling people to play a wide variety of games, either alone or interactively with a small or large number of others, as well as allowing people to learn and/or play certain (simulated) musical instruments. but, it is not simply in providing entertainments that computing technology makes for a home-based society. other examples include ‘telebanking’, whereby a growing percentage of the population utilize electronic funds transfer to have their paychecks deposited directly into their accounts and pay their bills electronically either through a direct-debit payment arrangement or some variant thereof, or through software which allows them to write checks from their accounts. ‘tele-shopping’ makes it possible for people to purchase virtually the entire panoply of consumer goods available in the market without having to leave their homes. (see forester [1981] and rosenberg [1992], among others, for additional material on computing technology and “homecenteredness”. for illustrations of the variety of software applications, see almost any introductory textbook on “computer literacy, for example, capron [1992], laudon, traver, and laudon [1995], and shelley, cashman, and waggoner [1995].) in a different, but related, vein, ‘tele-education’ enables people to study a broad range of disciplines and subjects, either for institutional credit or not, through their televisions and/or computers. “distance learning” is becoming a popular mode of instruction in many institutional contexts. there are institutions at which students can earn baccalaureate, master’s, and doctoral degrees through various combinations of correspondence, videoconference, and other technologically communicated courses. the open university in england is an example; even the venerable london school of economics offers a limited number of degree programs in this mode. people interested in learning some foreign language(s) may do so through instructional software. those who want to prepare for certain college/ graduate school admission examinations can avail themselves of the relevant program(s). those who wish to improve their skills in selected fields of mathematics, science, or the humanities may use their computers to run software or communicate with instructional resources via communications programs. (see the “education” pages of any software distributor catalog for examples of the variety of programs available in this area.) however, it is not simply as consumers or recipients of externally generated information that we observe this tendency toward a home-based society. given the (growing) preponderance of the information sector of the contemporary economy, large numbers of employees are directly involved in information activity. it follows that many of these workers can, with a home computer, modem, and necessary software, do their jobs from their homes. such work arrangements are termed telecommuting. “telecommuters” may work entirely at home or they may simply take work home from their office(s). the category may include an extraordinarily wide range of employees, from computer programmers who must, of necessity, be in continuous contact with their offices, to part-time typists, data entry operators, and other clerical workers whose only contact with their employers may be limited to sending and receiving job assignments. in 1991, a national work-athome survey by link resources found that approximately 5.5 million part-time and full-time employees spend normal daily business hours working from their homes. the conference board has found that between 15 and 20 percent of the firms it studied offered formal telecommuting arrangements to at least some of their workers. moreover, almost 80 percent of the surveyed firms allow telecommuting on an informal basis. (filipczak, 1992) the growing numbers and prevalence of telecommuters in the work force has led some commentators to use the term “electronic cottage” to describe the trend toward home employment. (toffler, 1981) as the preceding paragraphs suggest, computing technology has the potential to allow for the creation of a social environment in which people can carry on most, if not all, vital socially relevant activities without having to leave their places of residence. this might lead us to conclude that we are witnessing a return to a society in which the family and kinship institutions are the dominant forms of social organization. but, what appears to be happening is that people are engaging in these various computing technological behaviors as individuals, rather than in terms of their group roles as family members. (kumar, 1995) one spouse may use his/her computer independently of or in isolation from the other; parents may not know about their children’s use of the “family” computer. in fact, many nonfamily households have and use computers. the implication of these comments is that, because it is conducive to 14 iassist quarterly individualized and private usage, computing technology in its present form may have the affect of weakening the bonds between people which are the foundation of any social order. if this proves to be the case, a distinguishing characteristic of the information society will be a relatively weakened solidarity in comparison with other types, such as agricultural societies, industrial societies, etc. on the other hand, computing technology may also contribute to the integration of the information society, at least in its political aspects. groper (1996) presents a rationale explaining how e-mail can bring about increased participation in the political process by facilitating communication between citizens and their elected officials. he cites two illustrative casesthe legislative information network (lin) in alaska and public electronic network (pen) in santa monica, californiaboth of which appear to have this affect, at least initially. (the qube project in columbus, ohio failed to have the desired affect on political participation (rosenberg, 1992) but it was not based on email which allows for interaction between people and their representatives. instead, it simply provided for a limited number of electronic responses to political speeches.) it is my opinion that probably the most plausible resolution of this matter is to recognize that, under some conditions, computing technology may weaken the solidarity of information society, while, under other conditions, it can contribute to strengthening social solidarity. what is needed is research documenting which affects of computing technology are associated with particular social structural variables. quantitative micro-level affects: personal productivity measures. studies of “personal productivity measures” would include research on the number of tasks for which individuals and/or small groups use a computer in the accomplishment of the task, as well as the number of times the computer is used for said tasks. in 1981, weizenbaum, citing another computer scientist, wrote: for home use, [computers] have potential for catalogue shopping, activity planning, home library and education, and family health... family recreation, including music selection and games; career guidance; tax records and returns ... and budgeting and banking. (weizenbaum, 1981:553) this quotation suggests several questions for research on the ways in which individuals and/or families use their home computers today. for example, how many people purchase consumer products through a computerized shopping service? what types of products do they buy most often? least often? how many household members use “personal information management” software to organize their own or their family’s daily (weekly, monthly, yearly, etc.) schedules? how many individuals use the computer to prepare and file their state and federal income tax returns? how many households keep track of their income and expenditures through a computer program? in a related vein, how many people do their banking and other financial transactions through a computer? these questions are a sampling of the possibilities by which individuals and families have come to replace earlier, non-electronic means of organizing their lives with computing technology. an entire set of questions for research is suggested by the growing prevalence of the internet. how many computer users are connected to the internet? for those who are connected, how much time do they spend on a daily, weekly, monthly, or yearly basis communicating with others via the internet? browsing the world wide web? using ftp, gopher, or telnet to glean information from the internet? how many users regularly participate in newsgroups? “chat” rooms? bulletin boards? how many individuals have more than one internet connection (for example, an on-line service, such as compuserve, and a local internet service provider) on their computers? the essential concept underlying this category of analyses is that individual people and small groups can and do employ the technology of the computer as a tool enabling them to carry out more tasks more frequently and in more areas of their lives than was possible before the mass production, distribution, and utilization of the technology. moreover, the tremendous interest in and growth of the internet demonstrates that they tend to use their computers to establish connections with others, albeit in different formats and for different purposes. it would be informative to devise a quantitative measure detailing exact patterns of computer utilization among individuals and households in the information society. this would be a measure of actual usage, not purchases of software, nor subscriptions to on-line services or other internet service providers. i have in mind a methodology by which a random sample of computer users would serve as a panel, analogous to the sample of television viewers whose program preferences are monitored by the a.c. neilson company. the procedure for such a study would follow these guidelines. 1) a device similar to the “people meter” through which the neilson ratings are compiled would be attached to the computer(s) in a random sample of households. 2) the device would be activated every time the user boots up his/her computer and would record the following: a) the date and time when the computer is operative; b) the operating system in use during the session; 15spring 1996 c) the application program(s) in use during the session, as well as the number of minutes each program is activated; d) the internet connection(s) in use during the session, as well as the number of minutes each connection is active; e) [for multitasking systems] the number and types of simultaneous applications and/or internet connections active during the session. 3) provision would be made for recording certain demographic data about the sample of computer usersage, gender, income, occupation, race, etc.which data could be correlated with the usage data. the technology needed to implement such a plan exists and could be adapted for the purpose of gathering the relevant information. what is needed is financing and an institutional structure to carry out the project. qualitative micro-level affects: “techno-social construction of reality. the “techno-social construction of reality” category is intended to include studies of the ways in which people come to rely on the computer to validate their “knowledge” of the world in which they live, as well as to justify their actions. from the perspective of the sociology of knowledge, we know that interaction with and feedback from others are integral components of the process by which people develop a body of knowledge or world-view. (see berger and luckmann [1966] and stark [1991] for detailed explanations of this process and its consequences for human society and behavior.) this is an ongoing process in which people create material and nonmaterial cultural products which become parts of “objective” reality and which are used to structure everyday interaction. through socialization, these cultural products are passed on to and accepted by new members of society. the reality we know is thus created, sustained, and transformed through the giveand-take of organized social activity. if we apply this line of reasoning to the information society, it raises potentially interesting questions. how much of an individual’s understanding of the world is, if not directly gleaned from the computer, at least mediated by his/her dealings with the technology? furthermore, if the individual receives contradictory information from the computer and other sources of behavioral cues (for example, other people or institutions), to which source does he/she accord greater weight in deciding on a response? allow me to sketch two possible (hypothetical) scenarios to illustrate the issues raised here. scenario 1 involves a bank officer who must decide whether or not to give a small business loan to a 35 year old africanamerican male who would like to open a music store on the border of a minority neighborhood in a large city. the manager presents the relevant information to a committee of bank personnel for review and evaluation. the committee recommends that the manager not offer the loan to the applicant. at the same time, the manager enters the same relevant information into an expert system computer program. its recommendation is that the manager should grant the loan to the applicant. the second scenario concerns an undergraduate student at an eastern university who must write a term paper for an english course demonstrating that a number of short stories written under several different pen names are in fact the work of the same author. (assume that the student has started this paper sufficiently early in the semester so that he/ she is under no time constraints; also assume that his/her primary motivation is not to get a good grade, but to write a correct analysis.) the student can take either of two approaches to this project. he/she can content analyze each short story him/herself, noting the occurrence of the same or similar phrases, terminology, or other evidence of the author’s writing style and conclude, on the basis of his/her own examination, whether all of the stories were written by the same author. or, he/she may enter the data into an artificial intelligence program designed to recognize word or grammatical patterns and base his/her decision as to the authorship on the results of the computer program. these two scenarios are, admittedly, fictitious, but they are not completely out of the realm of possibility. (indeed, readers of this paper may be familiar with comparable episodes.) in both cases, the individuals are in a position in which they face a choice as to whether to place greater confidence in their own or their colleagues’ judgment or in the output from computing technology. by virtue of placing the individuals in such a dilemma, these scenarios may be thought of as 1990s updates of asch’s (1952) famous studies of group influence on individual judgments. for the bank manager, the advantage of trusting the human actors in this situation is that he/she can be assured of the social support they provide. for the student, there is the satisfaction of knowing that he/she was able to complete the project on his/ her own. on the other hand, in both instances, the persons involved can point to computing technology as the basis for their behavior. such an “explanation” not only rationalizes the chosen alternative, it also can be viewed as absolving the individuals of any responsibility for their conduct. given the similarity between the nature of the subject matter here and in the asch studies, it might be possible to set up small group experiments in which subjects would be placed in a setting where they would have to choose between a computer-generated recommendation for action and a contradictory one emanating from other group members (confederates of the experimenter). we began this paper by presenting a heuristic two-by-two 16 iassist quarterly table which could be used to organize studies of the impact of computing technology on society and social behavior. we noted its value as a device for generating hypotheses for further research in these areas. i have tried to demonstrate the utility of the framework in the preceding four sections. what remains to be done is the empirical research which can clarify our understanding of this most complex topic. references asch, solomon (1952), social psychology, englewood cliffs, n.j:prentice-hall, inc. bell, daniel (1981), “the social framework of the information society” in forester, tom (ed.), the microelectronics revolution, cambridge, ma:the mit press, 500-549. berger, peter l. and luckmann, thomas (1966), the social construction of reality, new york:doubleday and co., inc. capron. h.l. (1992), essentials of computing, new york: the benjamin/cummings publishing co., inc. forester, tom (ed.) (1981), the microelectronic revolution, cambridge, ma:the mit press. forester, tom (1987), high-tech society, cambridge. ma: the mit press. groper, richard (1996), “electronic mail and the reinvigoration of american democracy”, social science computer review, 14, 157-168. kaplan, abraham (1964), the conduct of inquiry, san francisco, ca:chandler publishing co. kumar, krishan (1995), from post-industrial to postmodern society, cambridge, ma:blackwell publishers, inc. laudon, kenneth c., traver, carol g., and laudon, jane p. (1995), information technology:concepts and issues, new york:boyd and fraser publishing co. miller, roger l. (1994), economics today: the macro view (eighth edition), new york:harpercollins college publishers. porat, marc u. (1977), the information economy: definition and measurement, washington, d.c: u.s. department of commerce. robertson, ian (1987), sociology (third edition), new york:worth publishers, inc. rosenberg, richard s. (1992), the social impact of computers, new york:academic press. shelly, gary b., cashman, thomas j., and waggoner, william c. (1995), using computers: a gateway to information, new york:boyd and fraser publishing co. stark, werner (1991), the sociology of knowledge, new brunswick, nj:transaction publishers. toffler, alvin (1981), the third wave, new york:bantam books. weizenbaum, joe (1981), “once more, the computer revolution” in forester, tom (ed.), the microelectronic revolution, cambridge, ma:the mit press, 550-570. wurman, richard s. (1989), information anxiety, new york:bantam books. 1. paper presented at the iassist/computing in the social sciences conference, hotel radisson, minneapolis, mn, may 1995. louis r. gaydosh,18 leigh drive, florham park, nj 079322401. e-mail address: lefty-lou@worldnet.att.net telephone number: (201) 377 7454 vol223 12 iassist quarterly in the united states, privacy legislation generally has been limited to government records at the different levels of government, i.e., federal, state and local. these laws impose requirements and restrictions on how government agencies collect, maintain and use information about individual persons. the united states, with few exceptions, has not adopted legislation to regulate and restrict the collection, maintenance or use of personally identifiable information by non-governmental entities. in other countries, this type of legislation is frequently referred to as “data protection laws.” however, if a threat to the personal privacy exists today in the united states, it comes not from the regulated governmental data bases but from the non-regulated ones outside of government buildings. the information in these data bases has long be available in paper form without threatening individual privacy. but technological developments have altered the situation. first, with the spread of technology, data is increasingly available in electronic form. second, computer processing speeds have increased. with data mining tools, a data base query that took six minutes in 1994 now takes less than nineteen seconds. third, with the availability of increased computing speed, it is now easier to combine data from multiple sources and create comprehensive information products. fourth, the cost of storing electronic information — on-line, near-line and off-line storage — has dropped. fifth, personal computers are becoming more and more affordable and thus more wide spread. finally, with the spread of the personal computers, internet use is becoming commonplace from businesses and homes.1 technology, then, has permitted the growth of what vice president gore has called “profiling,” or the ability to build dossiers about individuals by aggregating information from a variety of database sources. these dossiers now have detailed information on the vast majority of the american population, including children. for example, acxiom corporation has information on 196 million americans on 700,000 data tapes with 350 trillion bytes. information america claims that it has employment and other demographic information on 160 million individuals, 92 million households, 71 million telephone numbers, and 40 million deceased persons. finally, medical marketing service, inc., offers its “patient direct” data base with information on more that 20 million households documenting an individual’s age, gender, income, educational status, and health condition, from allergies (4.3 million households) to yeast infections (1.4 million)2 . the sources of information are varied. first, governmental records provide a wealth of information on individuals and, under a variety of freedom of information laws, are readily available. while some invasive records are from the federal government, most are seemingly from state and local governments. these records concern real estate transactions and holdings, marriage and divorce, birth certificates, driving records, drivers’ licenses, vehicle registrations, civil and criminal court proceedings, paroles, postal service change-ofaddress forms, voter registrations, bankruptcies and liens, incorporation documents, workers’ compensation claims, political contributions, firearm permits, occupational and recreation licenses, and filings with the securities and exchange commission. for example, the federal aviation administration has a list of all individuals licensed to fly in the united states which includes certification class, medical certificate and date of last medical examination. real estate documents describe the property, dates of sales, selling prices, mortgage amounts, lender, and the names of the sellers and purchasers. social security numbers are readily available as well, most commonly from state departments of motor vehicles, which also provide individuals’ name, address, height, weight, gender, eye color, date of birth, and whether or not an organ donor. for example, new york’s publicly available drivers’ information includes vehicle ownership, accident reports, conviction certificates, police reports, complaints, hearing records, suspension and revocation orders. although government records are increasingly available in electronic form and electronic foia guarantees access to the records in that format, other records must first be digitized.3 non-governmental entities, but still publicly available information, offer additional sources of information. for example, newspaper and magazines identify and provide background information on individuals, and electronic editions are now frequently available from the internet. powerful search tools permit people to search these privacy and computerized data bases by thomas e. brown * fall 1998 13 newspapers and magazines and find all references to a given individual. many professional and alumni organizations make membership lists available on web sites. indeed, many web sites contain detailed personal information to anyone who clicks on the url, such as adoption pages where adopted children and birth parents post detailed personal information in hopes of connecting with their blood relatives.4 a third type is private proprietary information. maybe the most well known of this type is the information of the three major credit reporting agencies, trans union, equifax, and experian. but other even more invasive data bases exist with varying degrees of confidentiality associated with them. more than 3,000 corporations in the united states collect, maintain and sell information from their data bases for marketing purposes. most of the information is provided by the individuals themselves, often through warranty cards or product registration cards. a careful reading of these forms reveals the depth of personal information people supply. as an example, a recent card for a coffee bean grinder asked about sex and ages of all household members, marital status, occupation, income, educational level, credit cards, home ownership, anticipated changes during the next six or 12 months, and finally a list of sixty interests activities in which the respondent participates in regularly. another series of incredibly revealing data bases are those maintained by banks and credit card companies which list the individual transactions charged to each credit card. but far more sinister is the spread of supermarket customer cards. last year, a survey indicated that 60 per cent of the supermarkets in the country had started such a program or intended to do so soon. these cards are used to record in a data base each product which a consumer purchases. one’s shopping habits reveal a lot about that individual. the purchases of condoms or large quantities of alcoholic beverages connote a life style. patent medicines reveal aspects about one’s health, and purchases of certain brands and products betray one’s value systems and interests. every time somebody makes a telephone call, it creates a record in a database which offers much information to the phone company and other marketers. through billing records, local carriers and long-distance carriers learn whom people are calling and when and where calls are made. finally, the internet itself is a great collector of personally revealing information. commercial web sites collect personal information through a variety of means, including registration forms, user satisfaction surveys, contests, and order forms. for example, an online doctor-referral service asks users for their name, mailing address, e-mail address, insurance company, and whether they want information on a variety of health concerns, such as urinary incontinence, hypertension, cholesterol, prostate cancer, or depression. another web site is from a mortgage company for prequalification for home mortgages. potential borrowers provide their names, social security numbers, home and work telephone numbers, e-mail addresses, previous addresses, current and former employers, lengths of employment, income, savings, and credit histories including credit cards. in the spring of 1998, a quartermillion people completed a detailed survey at the espn web site to enter a lottery for tickets to the ncc final four tournament. even without telling the users, web vendors can collect detailed personal information as they can identify which pages the user visited, what the user bought, where the user linked from and where the user went on the “net” when he or she left the site, and in some cases, how long an individual was at the site and on each page of the site. recently reported, some of the largest commercial sites including lycos-tripod and geocities, had begun providing users’ reading, shopping, and entertainment habits to a centralized system which links that information the user’s age, income, zip code, number of children, and information obtained from on-line forms. to protect privacy, the system uses a unique identifying number associated with the hard drive of the computer accessing the sites. thus when the user connects to a participating site which recognizes the number on the drive, targeted sales messages will be sent. but already, entrepreneurs are trying to link the preferences in compiled computerized data bases with mailing lists of traditional direct marketers. the owners of the system have announced that it will not collect information about sexual interests or health related topics. but with money on the line, such a voluntary restriction is probably not universal or permanent. for example, some internet sites devoted to specific diseases have solicited data from site visitors and then sold that information, either directly or indirectly through data intermediaries, to companies marketing drugs or other therapies for the specific disease. as nat goldhaber, chair of a web vendor called cybergold, commented, “what the internet has done is make explicit what used to be implicit — namely that there dossiers on you than can be built up with great granularity.”5 as june 24, 1997, there were fifty-one database vendors and information bureaus which will sell detailed personal information on the internet about telephone use, assets, criminal histories, vehicle and driving records, aircraft, boat and gun ownership and usage, business materials, marriages and divorces, current and previous addresses, and information on neighbors and relatives. some have unexceptional names, such as discreet research whose slogan is “when you need to know.” others are more interesting, such as dig dirt, inc., whose saying is “because what you don’t know does hurt you.”6 in the collection, maintenance, and use of this information, the united states has adopted as essentially laissez faire approach. current american privacy law has been described as “sectoral,” that is “a handful of disparate statutes directed at specific industries that collect personal data.” in fact, the federal level has only six data protection 14 iassist quarterly laws in place. the first and clearest example is the fair credit reporting act which guarantees the individual the right to gain access to personal information which the credit reporting agency has and the right to dispute erroneous information and add corrective details. it also generally restricts access to those businesses to which the consumer has applied for credit, insurance, employment, or a lease agreement. the second is federal educational records privacy act (ferpa) or more commonly know as the buckley amendment. this limits the release of students’ educational records to educational personnel and educational institutions. two other acts are the cable communications policy act of 1984 which restricts cable television subscriber information and the telecommunications act of 1996 which governs customer proprietary network information. the last data protection law in the united states restricts the release of an individual’s video tape rentals. enacted as a reaction to publicizing judge bork’s video tape rental during his supreme court confirmation hearings, the law prompted secretary of health and human services donna shalala to comment, “our private health information is being shared, collected, analyzed and stored with fewer federal safeguards than our video store records.” in the absence of legislation, case law is working against data protection. in united states v. miller [425 u.s. 435 (1976)], the supreme court ruled that individuals have no constitutional protection of information which that they have voluntarily provided.7 but currently, some 80 bills to strengthen the data protection are pending in the congress. some would restrict the dissemination or use of the social security numbers; others would allow individuals to stop the post office from selling their change of address requests; and at least one would establish an independent regulatory commission to control the collection, maintenance, and use of individually identifiable information. according to one observer, the new york times’ nina bernstein, “but with few exceptions, the proposals seem to be going nowhere. beneath the surface of their popular appeal, most are mired in unresolved conflicts over contradictory goals: on the one hand, preserving personal privacy, and on the other, the advantages of quick, computerized access to personal information for fighting crime, fraud and waste, or promoting growth in the information economy.”8 the one exception will be data protection of health information. the health insurance and portability and accountability act of 1996, commonly known as kennedykassebaum act, required the clinton administration to propose legislation on the creation, maintenance and use of medical information on individuals by september 30, 1997. [ironically, this same legislation also required that the administration create a standard medical identifier so individuals’ medical records could be widely available, regardless of the health care program, from a database of health information.] if the congress does not enact legislation by august 1999, then the administration must issue regulations. so on september 11, 1997, secretary shalala proposed legislation. but the administration’s proposal is only one of several laying in the legislative hoppers. reaching a consensus will not be easy in the tugof-war between consumer groups, law enforcement agencies, and health care professionals. should law enforcement officials need a warrant or court order to get medical records? can records obtained investigating insurance fraud be used for other criminal prosecutions, such as information about illegal drug use lead to prosecution? if consumers have the right to change or delete medical information, then some argue that this could endanger a vital resource needed for medical research, for public health analyses, and for improved medical care.9 if a legislative consensus is difficult, self-regulation is an option which conforms to the laissez faire approach and has an honored tradition in the country’s history. last december, the federal trade commission announced an agreement amounting to self-policing by “individual reference services,” that are businesses which have been selling personal information to the general public. by the end of this year, fourteen of the largest of these organizations, including the three major crediting reporting agencies and the largest direct mail marketers, announced that they would no longer provide information to the general public. they would also limit the types of personal information they would provide to commercial users like marketers, banks, lawyers and journalists. yet they would still allow unrestricted access by law enforcement personnel, licensed private investigators, and corporate security staff. according to the new york times, this agreement embodied the administration’s strategy of selfregulation. “the agreement sets in motion the first meaningful trial of the clinton administration’s privacy policy, the stated goal of which is to protect individual privacy in the internet age without resorting to new laws and regulations.” privacy proponents have objected to the agreement as too little. while individuals can request that their records be erased from some data bases, they cannot access all of the information being maintained and disseminated about them, and to have themselves removed from selected data bases, individuals would have to contact each of the fourteen reference services separately. an obvious loophole will be for an individual to hire an attorney or private detective to serve as an intermediary.10 in this same vein of self-regulation, the 3000-member direct marketing association has issued “guidelines for personal information protection “ which stipulates that personal data collected for marketing purposes should be only for that purpose. it further maintains to its committee on ethical business practices to investigate the misuse of marketing information. the association has also announced that it will require, beginning next year, its fall 1998 15 members to disclose publicly how they gather and use data. the council of better business bureaus is working on a model for self-regulation that would impose sanctions on businesses that fail to follow a code of conduct to protect people’s privacy. a recent innovation to self-regulation is incorporating a seal onto a commercial web site to indicate that the site follows an established code of conduct regarding privacy. a nonprofit group called “truste” already has a system in place for about 120 companies. in july, another group, online privacy alliance, emerged with the same intent.11 these efforts at self-regulation have been haphazard at best. the federal trade commission tersely concluded earlier this summer, “to date, . . . the commission has not seen an effective self-regulatory system emerge.” the clinton administration began to hedge last may when vice president gore outlined a new administration initiative on privacy. it consisted of a renewed call for legislation regarding medical records, a federal web site to assist consumers in deleting their names from commercial data bases, creation of the privacy officer in every federal agency, and a conference to address the topic in june 1998. this speech, calling for an “electronic bill of rights,” was interpreted as an admission that self-regulation had not been as effective as the administration had hoped. according to janlori goldman, a noted privacy specialist at georgetown university, “this administration has been singing the praises of self-regulation for some time now, but this is an acknowledgment that there are significant limits to what the private sector will do on its own.” as far as the legislation concerning medical records is concerned, dr. goldman opined that the clinton administration is “using medical privacy as a signal to the public and a stick to industry to say that we have a history of abuse in this area and the administration wants to do something about it.” 12 but regardless of the mode of data protection — legislation, regulation or self-policing, discussions of data protection revolve around eight principles of fairness: openness, individual participation, collection limitation, data quality, use limitation, disclosure limitation, security, and accountability. while originally proposed by a united states federal advisory committee, the principles were most clearly codified in the council of europe’s convention for the protection of individuals with regard to automatic processing of personal data. for the united states to be full economic partners with europe, data protection efforts will have to conform to the council of europe’s convention. two statements in the convention should give pause to archivists, namely: “article 5 quality of data personal data undergoing automatic processing shall be: b. stored for specified and legitimate purposes and not used in a way incompatible with those purposes; . . . e. preserved in a form which permits identification of the data subjects for no longer than is required for the purpose for which those data are stored.”13 this concept that personal data can be used only for the reason for which it was collected is fairly common in the discussions of data protection. in the above discussion of the direct marketing association, its “guidelines for personal information protection” stated that personal data collected for marketing purposes should be only for that purpose. in discussing the clinton administration’s position on health care, secretary shalala stated that personal medical information should be “for health care and health care only” with very few exceptions. but from an archival perspective, this limitation on use has a potential problem. archival theory discusses the primary value of records as the purpose for which the records were created. this is in contrast to the secondary value of records which is the value of the records to someone other than the record’s creator for a reason other than that for which they were created, and it is the secondary value of records that justifies the archival retention of records after the record’s creator no longer needs them in the course of business, but if the use of personal data is limited to the reason for which it was collected and if the archives are not exempt from this limitation, then records cannot be retained in archives for their secondary values. indeed, under the council of europe’s convention, the records must be destroyed as soon as their primary value has ended. in terms of medical records, for example, if an exception is not made for archival retention of some personal medical information then the history of medicine, the history of technology and the history of science — all currently viable fields of historical study — will be severely curtailed. application of the convention’s beyond just medical records will threaten the continued existence of records important to future genealogists and family historians. obviously, a solution is to outline exceptions to the use of limitation provisions in any data protection effort — whether legislation, regulation, or self-policing. one of these exceptions must be for the archival retention of records with significant secondary values to be released only after the passage of time that would permit access without endangering the privacy of individuals. unfortunately, the possibility of records having secondary values seldom enters into the debates over data protection, but then it seems that the archival profession has seldom entered in the debates over data protection.14 16 iassist quarterly 1federal trade commission, individual reference services: a report to congress, december 1997, pp. 3-4. 2 “vice president gore announcee new steps toward an electronic bill of rights,” this week’s press briefings and releases, july 31, 1968, http://library.whitehouse.gov, august 3, 1998; robert o’harrow, jr., “data firms getting too personal?” the washington post, march 8, 1998, pp. a1, a18-a19; federal trade commission, individual reference services, pp. 36; sheryl gay stolberg, “the numbering of america: medical i.d.’s and privacy (or what’s left of it),” the new york times, july 26, 1998, p. wk3. 3federal trade commission, individual reference services, pp. 4-6, 37; nina bernstein, “high-tech sleuths find private facts online,” the new york times, september 15, 1997, electronic edition. 4federal trade commission, individual reference services, p. 5. 1federal trade commission, individual reference services, p. 3; o’harrow, march 8, 1998, p. a18; nina bernstein, “lives on file: privacy devalued in information economy,” the new york times, june 12, 1997, electronic edition; federal trade commission, privacy online: a report to congress, june 1998, p.3, 39; saul hansell, “big web sites to track steps of their users,” the new york times, august 16, 1998, pp. 1, 24; denise caruso, “who knows what about whom on the internet,” the new york times, april 13, 1998, electronic edition 5federal trade commission, individual reference services, p. 3; o’harrow, march 8, 1998, p. a18; nina bernstein, “lives on file: privacy devalued in information economy,” the new york times, june 12, 1997, electronic edition; federal trade commission, privacy online: a report to congress, june 1998, p.3, 39; saul hansell, “big web sites to track steps of their users,” the new york times, august 16, 1998, pp. 1, 24; denise caruso, “who knows what about whom on the internet,” the new york times, april 13, 1998, electronic edition. 6see and . 7federal trade commission, privacy online, p. 62; 15 usc 1681; 20 usc 1232g; 47 usc 551; 47 usc 222; 18 usc 2710; robert pear, “clinton to back a law on patient privacy,” the new york times, august 10, 1997, p. 22. 8peter maas, “how private is your life?” parade magazine, april 19, 1998, p. 5; nina bernstein, “goal clash in shielding privacy,” the new york times, october 20, 1997, p. a16. 9steven findlay, “prescription for patient privacy: administration today offers plan to ensure confidentiality,” usa today, september 11, 1997, pp. 1-2; editorial, “hhs identifier puts privacy at risk,” federal computer week, july 20, 1998, p. 24; arthur allen, “exposed,” the washington post magazine, february 8, 1998, pp. 11-15, 27-28. 10federal trade commission, individual reference services, passim; katherine q. seelye, “a plan for database privacy, but public has to ask for it,” the new york times, december 18, 1997, pp. a1, a24; john markoff, “guidelines don’t end debate on internet privacy,” the new york times, december 18, 1997, pp. a24. 11federal trade commission, individual reference services, p. 38; o’harrow, march 8, 1998, p. a18; robert o’harrow, jr., “white house effort addresses privacy,” the washington post, may 14, 1998, p. e1, e4; robert o’harrow, jr., “internet companies move to safeguard computer users’ privacy,” the washington post, july 22, 1998, p. a13. 12federal trade commission, privacy online, p. 41; john m. broder, “gore to announce‘electronic bill of rights’ aimed at privacy,” the new york times, may 14, 1998, p. a22; o’harrow, may 14, 1998, p. e4. 13robert gellam, “don’t fear privacy protection — arm yourself with fairness checks,” government computer news, may 4, 1998, p. 26; department of health, education and welfare, secretary’s advisory committee on automated personal data systems, records, computers, and the rights of citizens, (july 1973), council of europe, european treaties, ets no. 108, convention for the protection of individuals with regard to automatic processing of personal data, strasbourg, 28.i.1981, http://www.coe.fr/eng/legaltxt/108e.htm. 14federal trade commission, individual reference services, p. 38; pear, p. 22; t. r. schellenberg, modern archives: principles and techniques (chicago, university of chicago press, 1956), pp. 28-32; t. r. schellenberg, the appraisal of modern public records, bulletins of the national archives, number 8 (washington dc: u.s. government printing office, 1956), p. 6. * paper presented at the iassist conference, may 21, 1998, at yale university, new haven, connecticut. thomas e. brown, manager, archival services electronic and special media records services division national archives and records administration. http://www.dresearch.com/ http://www.pimall.com/digdirt/moore.htm http://www.pimall.com/digdirt/moore.htm http://www.coe.fr/eng/legaltxt/108e.htm vol 40 no 3 / iassist quarterly 2016 27 iassist quarterly image management as a data service by berenica vejvoda1 k. jane burpee2 paula lackie3 abstract across all disciplines, researchers are creating or gaining access to an ever-growing body of digitized images. since research data management includes the organization of ‘all materials’ intrinsic to a research project, a robust data management plan will include a path for images as well as data in the more traditional sense. while researchers across disciplines have a long history with the organization of numeric data, the inclusion of images as a resource set in research is only starting to take shape across the disciplines. this paper is intended for data librarians or academic support staff without expertise in image data management. the primary focus is to apply traditional data management practices to images and to discuss the challenges associated with managing image collections through the research data lifecycle. keywords data management, images, research data lifecycle, preservation, sharing, mixed-methods research introduction to move in small measure towards a greater understanding of image collections as a data management challenge, this article compares traditional numeric data management with organized image collections. by conceptualizing image collection management as a component of data service across the ‘research data lifecycle’ we hope to foster a better understanding of how data professionals can effectively transition their skills to include the management of images. this paper addresses images as data rather than as an object. the unique challenges associated with images (versus numeric data) will be highlighted through the points of the research data lifecycle which are most impactful for image management. although the principles outlined here may apply to any digital object identified as an “image”, this article will assume a format-based approach that includes any twodimensional digital image format. a research data lifecycle approach for researchrelated image collections the following stages images in the data lifecycle (see diagram at right) will be discussed in the comparison with numeric data and research-related image collections: creating or collecting images; processing images; analyzing images; preserving images; giving access to and re-using images (adapted to images based on dcc, 2012 and mantra, 2014). adapted from, create and manage data research data lifecycle. uk data archive, 2016. retrieved from http://www.data-archive.ac.uk/ create-manage/life-cycle. copyright 2002 -2016: university of essex. creating and collecting images as data managing data during the creation stage of the research data lifecycle can be challenging, most notably when the ‘data’ are in the form of images. images can be collected in a number of different ways, e.g.: in-house or outsourced scanning or photography; digital creation from the outset; or purchased from vendors (primary research group, 2013). just like any data gathering process, for collection methods to be successful researchers will benefit if they make decisions before they begin the process of collecting and capturing images. therefore, questions commonly asked when 28 iassist quarterly 2016 / vol 40 no 3 iassist quarterly collecting numeric data offer a useful framework for effectively collecting and organizing image as a data-style resource (dcc, 2013). • what type of data will be gathered? • in what ‘format’ will the data appear? (will it change through the life cycle of the project?) • what will be the expected ‘volume’ of data collected? types of data ~ types of images at the most fundamental level, digital images are numeric data. they are all ones and zeros stored on computer media. in practice however, images are often more like social science microdata, they can be both data as well as a datum at the same time. for example, a collection of images organized around a particular theme is comparable to a dataset and the individual image, a specific response source. when considering an image management strategy, simultaneously managing images with their metadata is like managing survey metadata, paradata, and data all at once. for example, both surveys and image collections may have metadata, paradata, and data important to the analysis or for reuse purposes. for the purposes of data management planning, it might be helpful to think of these example parallels: when working through choices in an image data management plan, comparing the process with data associated with a survey respondent can be a good place to start; both yield information elicited through data collection instruments, are inherently complex, and it is necessary to make choices regarding which aspects to focus on. of course, the analogy is limited since unlike survey respondents, an image database may also contain the complete image which may be flawlessly duplicated unlike humans. in addition to the need for understanding the complexity and structure of images as data, for images to be useful, to those who create them as well as for subsequent re-use, it is equally necessary to consider the format of image files as early as at the point of creation. attention to format will ensure long term access and functionality. data formats ~ image formats deciding on the best file format (i.e. the way information is encoded in a computer file) is understandably a question that applies to both numeric data and images. both rely on applications or programs that will recognize the file format in order to access information within the file. according to a report published by a digital preservation policy working group at cornell (2001) file format for images consists of the bits that comprise the image and the header information on how to read survey respondent digital image data survey answers o,en simply the viewable image but it may also be the direct analy8c content derived from that image metadata e.g. age, educa8on level, home address e.g. camera make & model, image 8mestamp as set by the camera, aperture, shu>er speed (exif data) paradata e.g. respondents clickrate through a survey, e.g. average image color, facial recogni8on material, image sequence in a set 1 and interpret the file. similar to numeric data files, images can be stored in a wide variety of formats, including: bitmap (bmp), joint photographic experts group (jpeg), jpeg 2000, and tagged image file format (tiff). these standard image formats vary in relative file size, image quality and flexibility, and compatibility with software programs. distinguishing between minimal requirements and recommended imaging requirements, the cornell report gives preference to tiff formats as a master image format since they do not compromise data, while jpegs, a “lossy ” format which compresses data, are included in the minimal criteria. one research domain in particular that has championed the adoption of open standardized image data formats is the imaging community in the biological sciences, namely the jcb dataviewer initiative. initiatives like the jcb dataviewer align with the conventions of numeric data that recommend the adoption of open source formats in order to retain the best chance for future readability. if open source isn’t an option then choosing formats that are in widespread use or agreed-on international standards will help achieve the objective of longevity and/or replicable research. volume: counts and file sizes related to formatting is the notion of volume. when data were first digitizable, disk memory was extremely limited. researchers were resourceful in how they encoded and managed their data. as expectations for robust data analysis were fed by moore’s law and a parallel rise in disk storage capacity enabled the rise in big data (e.g. moment-to-moment stock trade data or global social network data). likewise, expectations for big data in the form of images has also risen. like numeric data, a big concern for image data formats is related to a combination of storage space and computational power. just like with numeric data; the numbers of image collections, the numbers of images in collections, and the size of individual images are all growing. researchers should be aware of the trade-offs they are making when choosing either fewer images or images of lower quality than the original images as they were created. dealing with large numbers of lossless image files can quickly become unwieldy in terms of available storage space. compressing large image files to smaller files is most easily produced using lossy compression and while the images may still provide adequate information for the immediate intended purpose, their longevity may suffer. an advantage of the usual numeric data compression over image file compression is that they are lossless (e.g. .zip, .7z, and .gz.) on the other hand, choosing smaller image file formats (e.g. jpegs) that are lossy over lossless image files (e.g. raw, tiff) results in loss of image clarity; resolution, layers, and fidelity. as a result, the management of images becomes a more difficult decision when dealing with a large number of large files. in terms of image management, due to the usual lossy nature of compression, there is a clear preference for retaining uncompressed versions or for working to manage the balance of lossless compression against future format compatibility challenges. at the very least, it is recommended that researchers minimize the number of compression processes that need to be managed over the long term (cornell, 2001). vol 40 no 3 / iassist quarterly 2016 29 iassist quarterly the challenges associated with large image files are compounded by the sheer volume of images being produced across various disciplines and sectors. while the growth rate of images can be difficult to quantify and many claims appear as unsubstantiated hyperbole, the discourse surrounding the explosion of visual content agrees that it is undeniably large (kane and pear, 2016). kane and pear (2016) estimate that while 3.8 trillion photos were taken in all of human history until mid-2011, one trillion photos were taken in 2015 alone. in academia, specifically, the biological sciences, moore, allan, burel, loranger, macdonald, monk and swedow (2008) noted eight years ago that ‘most laboratories and imaging facilities do not have the means to store the volume of data generated by their microscopes in manageable and affordable way’ (p. 557). another testament to exponential growth in the biological sciences comes from the rapid progress in genome sequencing technology. in 2011 gross noted that second-generation machines like illumina’s genome analyzer ii create vast amounts of images and that the volume of these images was growing by five terabytes a day. the volume of images as data produced through medical imaging is also staggering. marketandmarkets (2016) estimate that medical image archives are increasing by 20-40% annually, and they predicted that by 2012, there will be 1 billion medical images stored. processing images as data the growth of digital images in size and number, the advent of powerful digital cameras and the willingness of libraries and archives to use them, has produced an overwhelming need for comprehensive image management software (roy rosenzweig center for history and new media, 2016). this need is documented across disciplines by researchers who struggle to manage collections of digital images. in response to this need, in 2015, the andrew w. mellon foundation announced funding for a new project to develop tropy an open-source software application that will help researchers collect and organize digital photographs, create metadata, and export photographs and associated metadata to other platforms (centre for history and new media, 2016). researchers in the sciences – particularly, in the medical and life sciences, are also expressing a need for image management systems. the open microscopy environment (ome) consortium has, for example, built a series of open source tools that assist researchers in managing large sets of complex images to support research in cell and developmental biology (moore et. al., 2008). researchers relying on medical imaging (e.g. ct, mri, x-ray, nm, mammography, ultrasound, radiology) are also in need of image management systems to keep up with unprecedented growth. a case in point comes from the critical role of medical imaging, specifically image biomarkers, in clinical trials for alzheimer’s disease. increased reliance on these medical images for study outcomes requires image management systems for effectively capturing, processing, analyzing, disseminating and archiving images (jimenez-maggiora, thomas, brewer, bruschi, hong, and aisen, 2012). jimenez-maggiora et. al. (2012) note that these specialized systems are complex, inflexible and resource intensive. recommendations for managing traditional numeric data files at this stage of the research data lifecycle can provide a useful framework whether the data are “big” or just awkward. key considerations for organizing numeric data include data carpentry functions, such as: versioning, naming, and renaming (mantra, 2014). as with a numeric data file, an image file name needs to be carefully considered for consistency, logic and predictability so that users can effectively browse and retrieve image data and also to avoid confusion when multiple researchers are working and naming shared files. in other words, ‘thisimage.tiff’ will be as problematic as ‘thissasfile.sav’. and similarly, images will present with multiple files in various formats, multiple versions, across differing methodologies, etc. a growing number of software tools exist to help organize images in a consistent and automated way through functions such as batch renaming. renaming files may be especially useful for image metadata in instances where digital cameras automatically assign base filenames of sequential numbers. in the realm of numeric data file management, tools include: using the grep command in unix or applications such as renameit . in addition, imagemagick can perform various batch processing functions on image metadata. in addition to batching renaming, image management software can support image workflows by assisting with recording location, generating thumbnails and storing basic associated metadata that are embedded in image files (jisc, 2016). analyzing images as data one of the defining features of data is that they are the raw material produced by primary research that is intended for analysis (geraci, humphrey, & jacobs, 2012). arguably, numeric research data are most often created for the intended purpose of analysis. in contrast, images may not always be intended for analysis at the point of creation. for example, many early digital library projects supported the creation of images to add to library collections, but in a large majority of cases such images were and continue to be used by researchers for making examples and illustrations versus serving as raw material for research. it is important to note that an image or collection of images may serve a variety of users and as such, they may also become data for analysis. recommended best practices for managing numeric data at the analysis stage of the research data lifecycle (mantra, 2014; dcc, 2012) include documenting analyses and file manipulations, managing versions of data files and deciding if analyzed data will be shared. with numeric datasets, it is fairly routine to document analysis and manipulation functions, usually in the form of programming code files. similarly, documenting the manipulation and analysis of images helps researchers with their own image processing and analysis workflows (e.g., logging the numerous steps taken to geo-reference an image) and will also produce greater transparency of techniques, critical to successful replicability, which have in recent years emerged as potential sources of controversy across many disciplines. according to mccook (2016), companies such as image data integrity (idi) exist to help journal editors, publishers, funding agencies and institutions screen and verify whether image manipulation (e.g., blots and micrographs in biomedical materials) compromises the interpretation of the images for scientific purposes. in addition to documenting the process of analysis and manipulation, 30 iassist quarterly 2016 / vol 40 no 3 iassist quarterly researchers working with numeric data also decide what form of data to ultimately share: i.e. raw, processed, analyzed, final. the analysis stage of the research data life cycle presents unique challenges for those managing image data when using image analysis software tools. by way of example, imagej , which has existed for over 30 years, is a general-purpose, extensible scientific image-analysis program that is used to capture, display and enhance images in the biological sciences (schneider, rasband & eliceiri, 2012). one key challenge noted by schneider et. al. (2012) occurs when one uses the software to open and parse the countless variety of image file formats. in the case of proprietary file formats tied to specific software (e.g. with some microscopes), using such software can be especially problematic for reproducibility or any further sharing. it is preferable to have images from research processes be available independently of specialty equipment. imagej, for example, is able to connect directly to matlab so that researchers are able to run statistical analyses as well as other tools such as imaris which supports 3d and 4d image analysis (schneider, et. al., 2012). in order to encapsulate diverse needs even within a particular domain w biological imaging, image analysis tools need to remain flexible and extensible. preserving images as data one of the most critical aspects of data management, regardless of data type is the pairing of metadata to accurately and sufficiently describe datasets. it is important to note that like most other data management functions, metadata hold a central role across the entire span of the research data lifecycle. so while it is discussed in reference to preservation for the sake of structuring this discussion, it applies to other stages as well. metadata for images is especially critical as a way to organize and search through growing libraries and repositories of images that are being produced by researchers and consumers alike. as is the case with numeric data files, decisions need to be made about how much detail to record in image metadata records. the internationally accepted metadata standard for describing social science numeric data, ddi (data documentation initiative) can be crudely parsed between study-level descriptions and variable-level metadata. this distinction can also apply to images. for example, low level descriptors such as ‘title’, ‘creator’, and ‘size’ are similar to ddi fields such as ‘title’, ‘abstract’, ‘producer’, ‘distributor’, and ‘time period’. for more detailed descriptions, ddi offers metadata fields at the variable-level – e.g. exact meaning of the datum (icpsr, 2016). this level of description is created directly from formatted datasets. while the need for fuller descriptions of image content are also equally necessary, a key difference is the degree of subjectivity involved in producing more abstract, higher level meaning, such as feelings portrayed by a particular image, ‘happy’, ‘sad’ (jisc, 2016). this challenge for describing images is discussed extensively by eadie (2008) in his description of the development of the jiscfunded dublin core images application profile. eadie (2008) aptly notes that unlike text-based materials that can be machine processed, images are not easily self-describing. this reasoning can also contrast with numeric data where machine-processing can quite easily produce meaningful variable and file-level metadata based on objective information embedded in formatted statistical data files. digital images, on the other hand, have a more complex relationship with machine-assisted metadata-extraction; while pixel data and bit depth are objective data points, they provide little by way of meaningful contextualization of images as data (eadie, 2008). so despite increased quality and quantity of camera sensor elements, it is neither practical nor meaningful to describe images to aid organization or querying based on millions of image pixels (metadata group, 2008). in addition to subjectivity, images can be complex to describe because they often have relationships with other objects which may even be embedded within them. so before you can even begin to say anything about an image you need to be very clear about which aspects of it, or its relationship with other objects, on which you are actually focusing (jisc, 2016). for example, images can be found in slides, photographs, books, manuscripts, lectures, and presentations. there may also be interdependencies between images. for example, in the area of geographic information systems (gis), different layers of gis data are superimposed to create a richer representation for spatial analysis. adding to the complexity of image description is that description consists of at least three different types (eadie, 2008). first, there is technical information relating to the image. this is usually pretty straightforward as capturing technical metadata is usually automated (e.g., captured by digital cameras) and resides in the image itself. the main metadata container formats for images are: exif (exchangeable image file format) (for device properties), iptc (international press telecommunication council), iim (information interchange model) (workflow properties), and adobe xmp (extensible metadata platform). each metadata container format has unique rules regarding how metadata properties are stored, ordered and encoded (metadata working group, 2010). while technical metadata are fairly easily captured, how they are structured is considerably more complex. even within the container format, metadata are stored, for example, according to various semantic groupings; within these groupings there can be numerous individual metadata properties. perhaps the biggest issue concerning the structural complexity of technical metadata is that different applications and devices handle these technical specifications in different ways; hence, creating challenges for interoperability. technical metadata also becomes more complicated for long term preservation as more metadata fields need to be added that are not normally captured by devices – for example, image format migration and versioning. the second and third types of description for images as data relate to the content in an image; the application of abstract principles to the description of the image. not only are content and abstraction difficult to describe in a standard way, but text-based descriptions will vary depending on the knowledge, culture, experience and point of view of the cataloguer (jisc, 2016). even more difficult is anticipating the needs of users in terms of what to describe in images. for example, in figure 1, one researcher may be drawn to the couple while another, looking for depictions of leisure activities would find the hoop-rolling relevant. one way to deal with this challenge is to balance the time and resources available for describing images with the anticipation of what level of detail users of the image will require for effective discovery (jisc, 2016). given the inherent subjectivity and richness of images, however, vol 40 no 3 / iassist quarterly 2016 31 iassist quarterly rarely in large image libraries and archives are there sufficient resources for in-depth description. in addition to the critical role of associated metadata, image preservation involves decisions about where and how to store them. as is the case for numeric data, depositing images in an archive or repository should facilitate discovery and preservation for the long-term. as mentioned earlier, some domains such as genomics produces such vast quantities of images (more than five terabytes a day; see gross, 2011) that the practicality and cost of archiving these images for preservation purposes is often not feasible. gross (2011) instead points to a motivation to alternatively invest in the development of real-time processing of the images ‘to output only the base calls and the quality values’ (p. r204). the distilled nature of the images for their originally intended purpose will limit future re-usability of these images but current technical and financial limitations mandate that some hard decisions are being made regarding precisely what to preserve. with the currently overwhelming volume of images as data we see additional kinds of re-use obstacles; image collections constitute “big data” in that they are larger than can currently be managed, adequate storage space is another issue, and even if storage were available, the transfer of such large files or sets of files is currently impractical. this should however be prefaced with a note that new infrastructure initiatives such as the pan-european project called elixir (european life science infrastructure for biological information) are looking for solutions that balance software compression with judicious data reduction. with elixir, the ultimate aim is to bring compression down to 0.1 bits (0.01 bytes) for every base stored which translates to a human genome taking up just 30mb of storage (gross, 2011) (as opposed to the 1.5gb without the compression). not all image collections are so unwieldy. for the more common, manageable collections of images, it is possible to decide where to best archive image sets and their metadata. because images as data are produced across a vast number of domains, repositories can range from individual solutions for photographers and other artists, to institutional photographic and slide collections, to archives and museum image repositories, and as institutional teaching and research archives. as previously mentioned in reference to open/standardized data formats, jcb dataviewer was the first open repository in the life sciences that allowed for archiving and sharing of original image datasets to support published scientific articles (linkert, rueden, allen, burel, moore, patterson, loranger, moore, neves, macdonald, tarkowska, sticco, hill, rossner, eliceiri and swedlow, 2010). in addition to archiving the original binary image and associated metadata, additional information captured by acquisition software includes: acquisition settings, image size, and resolution. while the imaging community in the life sciences already treats images as data and wherever possible has robust archiving solutions, this is not the case in all areas of research where digital images are produced. for these areas, the consideration for archiving numeric data can provide guidance for treating images as data in need of long-term preservation. arguably, university repositories that have traditionally focused on archiving text-based resources may benefit from examining numeric data archiving practices as they can assist in archiving images as data. as we know, preserving numeric data can be problematic due to the amount of data being generated (mantra, 2014). given that image files are substantially larger, on average, this becomes a key consideration. additionally, reliance on specific technologies for accessing anything digitized becomes problematic since the technologies change quickly. therefore, as with numeric data, it is essential to archive digital image data in a systematic way in order to minimize the chance of obsolescence or making images inaccessible over the long term. accessing and sharing images the sharing and accessing phase of the data lifecycle is perhaps where those working with images become easily overwhelmed and/or frustrated when a discrepancy emerges between needs and search results (chung & yoon, 2011). as noted in other phases of the lifecycle this challenge is exacerbated by the explosive growth and availability of images. in order to allow for effective access to images, the growing image collections must be organized in ways that allow for efficient discovery, browsing, searching and retrieval (rui, huang & chang, 1999). according to wang, mohamad & ismail (2010), an effective image retrieval system needs to be able to retrieve relevant images based on queries that conform as closely as possible to human perception. so unlike quantitative numeric data files, which are relatively straightforward to describe based on keywords that map to the represented measures, visual information is far more ambiguous and semantically rich (wang & ismail, 2010). relying on traditional keyword querying systems of access will not be sufficient. wang et. al. (2010) note that image retrieval based on keyword querying, popular in the 1970s, relies on keywords used as descriptors to index an image. while the jisc digital media guide (2016) discusses the need for advanced search features (e.g. boolean logic) to support relevant keyword image retrieval, wang et. al. (2010) would argue that assigning keywords manually to images is not only time consuming but that keywords alone are inadequate and grossly inefficient to describe the rich content of images. another trend in image retrieval that supersedes text-based image retrieval (popular in the 1980s), is content-based image retrieval (cbir). first used by ibm, this method of retrieval is based on extracting visual features from the image itself. while in theory cbir systems can include the extraction of low or high level features, even with sophisticated algorithms that can combine multiple visual features, elements such as colour, texture, shape and spatial relationships do not come close to mirroring the richness of image content that users have in mind when searching for relevant images. as a proposed solution, wang et. al. (2010) discuss at length the most recent trend of developing semantic-based image retrieval systems that allow users to query image data using high-level concepts. in short, this retrieval method maps the automated process of extracting low level visual features from images with semantic descriptions also stored in an image database. it is hoped 32 iassist quarterly 2016 / vol 40 no 3 iassist quarterly that this example of intelligent image retrieval, that can better represent the abstract concepts inherent to images, will help users discover and access relevant images. while semantic-based image retrieval, in theory, allows users to better discover relevant images based on higher-level meaning, creating semantic descriptions still requires human intelligence which is time consuming and expensive. one innovative way to tackle this dilemma is to consider the need for metadata creation by users during the access and reuse phases of the research data lifecycle. numerous web-based initiatives are testament to this type of metadata crowdsourcing (or ‘social tagging’). artuk, an online site for art from every public collection in the uk, recently launched, artuk tagger , which allows the public to add multiple tags to paintings. an algorithm then calculates which tags are likely most accurate and feeds these tags through to the art uk website. similarly, the philadelphia museum of art encourages online visitors to tag objects in the online collection in order to improve access to works of art. social tagging initiatives not only exist to provide better subject access to images but also to assist with quality control and processing functions. micropasts for example, encourages users to help with location accuracy of artifact findspots and photographed scenes as well as the masking of photos intended for 3d modelling. zooniverse is yet another platform that is designed to use volunteers to sort through and help classify excessive numbers of research images. “our goal is to enable research that would not be possible, or practical, otherwise.” while crowdsourcing initiatives and semantic-based search functionality can improve access to relevant images, building an image database based on shared standards remains a challenge given the diversity of image collections, widely varying budgets and differing individual requirements of user groups (bourne, 2005). regarding the varying requirements of users, chung and yoon (2011) found that users were more likely to search based on abstract meaning when images were intended as objects but not when used as data. another constraint that has similarly plagued the management of numeric data collections in university settings is that images (particularly slides) generate data that are very different from those handled by cataloguing systems created for books (bourne, 2005). re-using images in addition to browsing, search and retrieval challenges associated with managing access to images, a discussion of the closely associated re-using phase of the data lifecycle is not complete without attention to image copyright as well as issues regarding confidentiality or data sensitivity. generally speaking, factual numeric data in and of itself, represented in an obvious file structure, is not copyrightable in canada as a work needs to be original for copyright to exist. this holds no matter how much work goes into collecting the data (potvin, 2008). however, most numeric data need to be analyzed and processed and so the program code developed for these purposes is considered a ‘literary work’ (potvin, 2008). also, once data are formatted into, for example, a relational database, graph or dataset, it can be subject to copyright. as previously mentioned, images, unlike traditional numeric data files, are not always created for the purpose of data analysis. as such, most literature discussing images and copyright reference images as artistic works that are copyrightable. images considered to be artistic works in the uk, for example, include: blueprints, building plans, cartoons, charts, decorative graphics, diagrams, drawings, engravings, graphs, illustrations, logos, maps, moving images, paintings, photographs, sculptures and sketches. according to the government of canada images as artistic works include: patterns, art slides, maps, paintings, architectural drawings, plans, digital images, drawings, photographs, charts, and art prints. just as with copyrightable numeric datasets, it is necessary to clarify who has primary ownership of the datasets since copyright of a work comes into existence at the point of creation (i.e., authors/ creators). if images are treated as data then similar ownership and rights issues apply when figuring out how the images will be managed and disseminated. this will mean assessing whose rights need to be considered, for example: funders, institutions, research participants, collaborators, publishers and the public (mantra, 2014). copyright is undeniably complicated and there are obvious, notable exceptions to the principle of creator as copyright holder. for example, u.s. federal copyright law denies copyright protection for works produced by the us federal government so nasa’s images, for instance, are in the public domain, but individuals who create images based on data released by nasa can assert limited copyright because they have created derivative works or compilations. additionally, copyright surrounding images is context-specific. images may simply exist as facts (equivalent to a numeric data file), in which case, they are not subject to copyright. copyright may also not be an issue if copyright has expired or images are considered to reside in the public domain. some image owners may also allow reuse for non-commercial purposes (i.e., education) but require attribution (e.g., creative commons attribution). the educational sector may also find that images fall under fair use or fair dealing. in support of this, the visual resources association (vra) in the u.s. published a statement on the fair use of images for teaching, research and study (wagner & kohl, 2012). for teaching and study purposes, this statement covers preservation, use (both high-resolution and thumbnails), adaptations, sharing and reproduction. also, if images are photographs, they are likely to be treated as original works and subject to normal copyright restrictions. collections of images may be copyrightable if they exist as a database or as a result of researchers creating added value to images. ethical considerations, specifically privacy and confidentiality, also merit careful consideration when managing and sharing images. as with numeric datasets, researchers managing and disseminating images need to minimize the risk of disclosing confidential information and re-identifying study participants. one broad technique for safeguarding confidentiality of numeric data includes either collecting data without identifiable information or anonymizing data post-collection through de-identification processes. best practices for handling sensitive image data may include the anonymization of facial and location identifiers in digital photographs. the actual techniques for doing so, however, may pose unique challenges for images. for example, in a 2016 email thread on the jisc research data management listserv, the vol 40 no 3 / iassist quarterly 2016 33 iassist quarterly issue of anonymizing image data proved labour intensive and expensive for a use case where anonymization was required for a large collection of images. they couldn’t find a freely available tool that could effectively bulk-blur identifying characteristics in the images. in another comment from the same thread, a researcher noted that obscuring faces by pixelating sections of a video image could greatly compromise the usefulness of data. alternative strategies to anonymization noted by many researchers are to either gain consent to share, or to consider controlled access so that the usability of the images can remain unaltered. to maximize the effectiveness of these alternative strategies to anonymization, strong recommendations are made to consider and judge at an early stage the implications of depositing images with confidential information. figure 1 source: leblond & co. her majesty at osborne. regal series ca 1850. print collection, rare books and special collections, mcgill university. conclusion the unprecedented growth of research image collections across disciplines, coupled with increasingly powerful instruments and devices for image capture, have created challenges and new opportunities for managing images across the research data lifecycle. this paper offers some preliminary recommendations for managing images as data by looking to established research data management practices for traditional numeric datasets. analogously, images, like survey respondents will provide information in the form of data. but like people, images are inherently richer than the discrete slices of data that are extracted by research instruments. uniquely, images pose challenges in terms of size and volume, especially for storage and preservation. the creation of robust metadata is also complicated due to the subjectivity of image meaning and the difficulty in anticipating the search needs of users. also, automating the process of metadata creation is difficult and manual description remains necessary. some emerging solutions to these challenges are noted, including crowdsourcing initiatives for data processing and description as well as semantic-based retrieval systems. references avondo, j. (2010) bioformatsconverter. (available at cmpdartsvr1.cmp. uea.ac.uk/wiki/banghamlab/index.php/bioformatsconverter) bourne, m. image data. vra bulletin, 31(3), 26-29. (available at http://web.b.ebscohost.com/ehost/pdfviewer/ pdfviewer?vid=7&sid=106aa9eb-e435-4f6f-a71e-01e5c781e0f3%40s essionmgr106&hid=107) center for history and new media (2016). rrchnm to build software to help researchers organize digital photographs. (available at https://chnm.gmu.edu/news/rrchnm-to-build-software-to-helpresearchers-organize-digital-photographs) chung, e. and yoon, j. (2011). image needs in the context of image use: an exploratory study. journal of information science, 37(2), 163-177 (available at http://jis.sagepub.com/content/37/2/163.short) corti, l., van den eynden, v., bishop, l., and woollard, m. (2014). managing and sharing research data: a guide to good practice. los angeles: sage publishing, humphrey, c. (2006) e-science and the life cycle of research. (available at http://datalib.library.ualberta.ca/~humphrey/lifecyclescience060308.doc) ddi alliance (2013) data documentation initiative. (available at http:// www.ddialliance.org) data curation centre (dcc) (2012). (available at http://www.dcc.ac.uk/ resources/curation-lifecycle-model) edina. (2014). mantra (available at http://datalib.edina.ac.uk/mantra) eadie, m. (2008). towards an application profile for images. ariadne: web magazine for information professionals. http://www.ariadne. ac.uk/issue55/eadie geraci, humphrey, & jacobs (2012). data basics: an introductory text. unpublished. local pdf file. gross, m. (2011). riding the wave of biological data. current biology, 21(6), r204-r206. (available at http://ac.els-cdn.com/ s0960982211002818/1-s2.0-s0960982211002818-main.pdf?_ tid=e46374de-026b-11e6-b386-00000aab0f26&acdnat=1460657489 _27698d6d824b844bbc3b43f55c85558e) icpsr (2016) data management & curation: metadata (available at https://www.icpsr.umich.edu/icpsrweb/content/datamanagement/ lifecycle/metadata.html) jimenez-maggiora, g. a., thomas, r. g., brewer, j., bruschi, s., hong, p., and aisen, p. s. adcs electronic data capture (edc) integrated multi-modal image management for clinical trials in alzheimer’s disease. neurosciences department, university of california at san diego: la jolla, ca,(available at https://www.researchgate. net/profile/gustavo_jimenez-maggiora/publication/269038101_ adcs_electronic_data_capture_(edc)_-_integrated_multi-modal_ image_management_for_clinical_trials_in_alzheimer’s_disease/ links/548734e30cf268d28f071e7c.pdf ) jisc digital media (2016) (available at http://www.jiscdigitalmedia. ac.uk) kane and pear (2016) (available at http://sloanreview.mit.edu/article/ the-rise-of-visual-content-online) linkert, m., rueden, c.t., allan, c., burel, j.m., moore, w., patterson, a., loranger, b., moore, j., neves, c., macdonald, d. and tarkowska, a., (2010). metadata matters: access to image data in the real world. the journal of cell biology, 189(5), 777-782. 34 iassist quarterly 2016 / vol 40 no 3 iassist quarterly marketsandmarkets. (2016). rising volume of medical imaging data to increase the adoption of cloud computing in the healthcare sector. (available at http://www.marketsandmarkets.com/researchinsight/ north-america-healthcare-cloud-computing.asp) mccook (2016) retraction watch. don’t trust an image in a scientific paper? manipulation detective’s company wants to help. retraction watch. (available at http://retractionwatch.com/2016/02/24/ dont-trust-an-image-a-new-company-can-help) metadata working group. guidelines for handling image metadata. (2010). (available at http://metadataworkinggroup.com/pdf/mwg_ guidance.pdf ) moore, allan, burel, loranger, macdonald, monk and swedow (2008). open tools for storage and management of quantitative image data. chapter 24 (available at http://citeseerx.ist.psu.edu/viewdoc/ download?doi=10.1.1.461.9978&rep=rep1&type=pdf ) potvin, j. (2008). how is copyright relevant to source data and source code? technology innovation management review. (available at http://timreview.ca/article/121) primary research group. (2013). survey of best practices in digital image management. new york: primary research group. keeney, a. r. and rieger, o. y. (2001). report of the digital preservation policy working group on establishing a central depository for preserving digital image collections. (available at https://www. library.cornell.edu/preservation/imls/image_deposit_guidelines. pdf ) research data canada, (2014) (available at http://www.rdc-drc.ca) schneider, c. a., rasband, w. s., and eliceiri, k. w. (2012). nih image to imagej: 25 years of image analysis. nature methods, 9(7), 671-675. (available at https://www.researchgate.net/profile/kevin_eliceiri/ publication/228085958_nih_image_to_imagej_25_years_of_ image_analysis/links/0fcfd4fee589e852eb000000.pdf ) statistics canada. (2015). section 4: data. (available at http://www. statcan.gc.ca/eng/dli/guide/toc/3000276) ukda (2011) (available at http://www.data-archive.ac.uk/media/2894/ managingsharing.pdf ) wagner, g. and kohl, a. (2011). visual resources association: statement on the fair use of images for teaching, research and study. vra bulletin, 38(1), 1-18. wang, h. h., mohamad, d., and ismail, n. a. (2010). toward semantic based image retrieval: review. proc. spie 7546, second international conference on digital image processing, 754626 (available at doi:10.1117/12.853332). notes 1. berenica vejvoda (mist, university of toronto) is a data librarian at mcgill university. prior to mcgill, berenica worked as a data librarian at the university of california at san diego and at the university of toronto. berenica.vejvoda@mcgill.ca 2. k. jane burpee (mlis, mcgill university) is the coordinator, data curation and scholarly communications at mcgill university. she has been active in the area of scholarly communication since 2000 when she became a scholarly communication librarian at the university of guelph. she is a leading voice for open access and champions the transformation of scholarship. jane.burpee@mcgill.ca 3. paula lackie (ma, a.b.d., university of southern california) is the academic technologist for data at carleton college. she is longtime research data advocate and social science and humanities technologist. plackie@carleton.edu 4. the research data lifecycle (humphrey, 2006; dcc, 2012; ddi alliance, 2013, mantra, 2014). 5. lossy compressions transform and simplify the media information in a way that gives much larger reductions in file size than lossless compressions. while the file becomes significantly smaller, quality of the image is compromised during the compression process (e.g. jpeg). conversely, lossless compression results in no information loss, however, the image files are much larger (e.g. tiffs) jisc, 2016 (available at http://www.jiscdigitalmedia.ac.uk/infokit/file_formats/ lossless-and-lossy-compression) 6. jcb dataviewer (https://datahub.io/dataset/jcb-dataviewer), launched in 2008, for archiving and sharing original image data in the life sciences, allows users to download original image data in an open, standardized data format and preserves the original image metadata (ome tagged image file format [tiff]). similarly, the jiscfunded data management for bio-imaging project at the john innes centre developed bioformatsconverter software (avondo, 2010) to batch convert bio images from a variety of proprietary microscopy image formats to the open microscopy environment format, ometiff. ome-tiff, is an open file format that enables data sharing across platforms and maintains original image metadata in the file in xml format (ukda, 2011). 7. moore’s law refers to the long-standing pattern that computer processing power will double every two years. 8. big data is currently an ill-defined term that at its root simply refers to extremely large data files that require greater than average computational power to manipulate and/or analyze. 9. tropy: http://chnm.gmu.edu/news/rrchnm-to-build-software-tohelp-researchers-organize-digital-photographs 10. renameit https://github.com/wernight/renameit 11. imagemagick: http://www.imagemagick.org 12. a “thumbnail” is a very small version of the original image. thumbnail versions are useful as a kind of wordless summary of the image. 13. imagej: https://imagej.nih.gov/ij/ 14. artuk http://artuk.org/tagger 15. micropasts crowdsourcing: http://crowdsourced.micropasts.org 16. zooniverse is a “platform for people-powered research”: https:// www.zooniverse.org the development of which is funded by generous support, including a global impact award from google, and by a grant from the alfred p. sloan foundation. an analysis of cd-rom as a long term archiving solution by denis oudard' digipress before we go into this analysis, lei's take a look at the history of archiving. this table shows different systems of communication and the role of each element as well as their similarities. let's first tackle the 1st objective: access system longevity. simplicity is the best way to insure the longevity of an access system. such is the case for microforms. the magnifying glass is a simple system and we can rest assured that humanity will know how to use a magnifying data medium coding hieroglyph stone/papyrus rosetta stone text and bav image microforms language, alphabet, magnifying glass databases, raster 9 track tapes tape drive, computer images, ascii system sound, digital and compact disc cd-player all of this combined and/or fiuiher processed provides us with information. there is no doubt that at the eve of the information age, we will be called upon to archive, for the long term, huge amounts of information. in my first paper on this subject i stated that to provide a long term archiving solution, one needs to have two elements: 1. a retrieval or access system that will also endure the test of time. 2. a long lasting medium on which the data is stored. one without the other is as useless as a deck of cards without aces. glass for years to come. therefore, the survival of this access system is virtually guaranteed. unfortunately, when dealing with computer archives, simplicity is definitely not part of the equation, so we have to look elsewhere for longevity factors. the first one is momentum . in other words "how much acceptance or wide spread use does the technology have?" the law of large numbers is going to be a key ingredient consumer products have more momentum than professional products because they are sold in the 100s of millions rather than in the tens or hundreds of thousands. cdrom is the first computer media to be based on a widespread consumer item. therefore, we have 100s of millions of machines out there that contain 90% of the key components of a cd-rom drive. even if cd-rom technology, as we know it today, is abandoned in twenty 70 lassist quarteriy years or so, chances are we will find working cd-rom drives in a 100 years. another important way momentum can be measured is by the number of manufacturers which build the same product. today i can give you the names of at least 10 manufacturers of iso 9660 cd-rom drives. are you ready? here we go: chinon, digital equipment, hewleu packard, hitachi, nec, philips (lsmi), pioneer, sony, texel, toshiba; and i know i have left out some. compared to the momentum behind cd-rom, other media are very fragile. worm drives, for example, depend on the whims of a single manufacturer. each manufacturer makes a different type of worm, and the day the manufacturer stops making that drive, you have less than five years to transfer or lose your data. the second critical factor relating to long-term access is system independence. the reason this is important is because none of the computer systems we know today, none, will be available in 20 years, let alone 50, 100 or 200. in today's computer world we are dealing primarily with monolithic systems. that is, the hardware, operating system, appucation, and data are interdependent. ibm or vax or most any other computer hardware have their own operating systems, their own application, and their own data sets, which can work only within their own world. the perfect archive needs to be accessible not only to all the systems that exist today, but also lo all the systems which have yet to be created, all the super fast computers of tomorrow. thanks to the red book standard and iso 9660, cd-rom is a peripheral device that has already achieved hardware independence. diskettes, one of the most widely used media in the computer world, will never be able to make this claim. try pulling a ms/dos diskette into a macintosh; you can't even get a directory listing, let alone read the files. the 9 track tajje might be the only other medium which has achieved the same degree of hardware independence. hardware independence is key. an iso 9660 drive will read any iso 9660 disk, regardless of the drive manufacturer and machine environment. that is an amazing feat, and again only equalled by the 9 track taf)e. that is the good news. the bad news is that while hardware independence is a big hurdle, it is only the first hurdle. true system independence demands much more. system independence in today's world means flexible interaction between the following elements: 3. retrieval software. the infonnation itself. 1. presentation or process software. this independence can only be achieved by setting standards which define how one element works with another. in essence, this is what the red book standard does at the physical and opto-electronic levels. for logical software transactions, a similar set of standards is needed. some of these standards have been established, and still more are emerging at this time. i will hmit my discussion to these three standards: cd-rdx, sgml and tiff. more should most likely be included, but these will serve the purpose of explaining how such standards work. let's stan with cd-rdx. this standard was written for the primary purpose of enabling the user to utihze a single interface, his own. to achieve this, cd-rdx uses a client/ server approach. the client (the user interface software, or application software) is separated from the server (also known as the retrieval engine). the idea is beautifully simple. if you establish a set of rules by which the client can ask the server for specific "searches" (e.g. a boolean search), then any client using this protocol can interact with any server using the same protocol. moreover, cd-rdx is set up in such a manner that the client and the server do not need to be on the same computer, not even on two computers with the same operating system. so now we have system independence between the retrieval engine system and the presentation or application system. this represents important progress, but system independent data has not yet been achieved. the goal, remember, is to access the data on any disc via any retrieval engine. the solution is to agree on the structuring of information. tiff and sgml are such standards which could be used to structure information in a universal, non-proprietary way. then, a cd-rdx compliant search engine could be developed to access any set of information structured within the guidelines of sgml and tiff. there it is, not simple, but resilient for this to truly work, one also needs to take into account the problem of indexes, though the same reasoning applies, the question of standard index methodology is a truly thorny issue. what we would have is a structure where when one system changes, the information does not need to change and can remain on the same medium. if the user/client system changes, client software is rewritten to run on the new system, along the guidelines of cdrdx. if the server system changes, cd-rdx server summer 1991 71 software is rewritten to retrieve data structured along the sgml and tiff guidelines, keeping the data unchanged. this arrangement is being developed today for easy distribution of data. the multitude of clients is due to the multitude of users and potential users of the data distributed, multiplied by the number of titles they each use. in the case of long term archiving, the multitude of systems is even greater because time is a new multiplying factor. it must also be noted that this solution enables computer user interface and search engine to progress at their own pace, while "stable data sets" stay on one medium. of course, the characteristics of the medium remain intact. the time will come though, when today's advanced cdrom technology will be regarded as a bulky and slow medium. isn't it a lesser evil though, next to the alternative of losing access to the data content forever? while today, cd-rom is the perfect tool to have information on-line and on-site, to be used as an archive, the cdrom will be off-line. these archives will be fed into the super computers of tomorrow. after all, the cd-rom will not always be the medium where the information resides for processing. often the needed data will be downloaded to fast access, massive storage devices of future computers for processing as needed. this already happens when information is downloaded from a cd-rom onto a word processor or a spreadsheet to recap the longevity factors relating to access systems, we have established that hardware availability in the distant future is all the more likely if the technology is: 1 . widely acceptedthanks to the support of a product in the consumer market that uses the same technology. proliferation helping the survival of functioning hardware. 2. the technology must also be standardized which helps the consumer market become even bigger (see 1.) and documents the detailed working of hardware in a precise and widely available manner. logical access in the distant future is all the more likely if: 1. information is system independent. 2. information is formatted using a non-proprietary standard. indeed the wide acceptance of a few well chosen standards will foster the development of compatible software and ensure consistent data structuring. 3. the users can have access to this data and then manage it (for display, printing or processing) using the system of their choice through a system independent client/server type of protocol. more to the point, we have established that cd-rom is the medium that answers these requirements. the best 9 track tape, while very widely used, does not benefit from the momentum of a consumer market worm does not even come into the picture. this is not to say though that worm, tape and other media, which do not meet some or any of these criteria, will not continue to perform important tasks in our computer rooms. it is now time to move onto the longevity of the media itself. of course, given the above conclusion, if one could find a cd-rom that would last a hundred years or more, we would be home free. as some of you in the audience know, i am with a company that has dedicated large amounts of time and money over the last 5 years to develop such a cd-rom. conclusion as i see it, cd-rom, along with fantastic information distribution capabilities, has already more long term archiving features than any other mass computer storage available. only a few additional steps need to be taken to truly make it one of the top answers to data information archiving for the next 20 years. these steps are the focus of a committee being considered for creation by the commission on preservation and access. all this is very encouraging, and i hope that it can help solve some of your long term archiving needs. bibliography: 1 "taking a byte out of history: the archival preservation of federal computer records", house report 101-978, 101st congress, 2nd session "gao faults nasa for mismanaging storage of valuable u.s. space science data" james r. asker, aviation week & space technology, april 2, 1990 (enclosed) "lost in tapes" j. sniffen, associated press, january 2, 1991 (enclosed) "national archives needs better record-keeping technology, report says" ann m. mercier, federal computer weekly, november 1990 "'sgml like'and we could get 'sort of married' too.." william zoellick, disc magazine. premier issue, fall 1990. page 53-54. available from helgerson associates. tel: (703) 237-0682 72 lassist quarterly 6"cd-r0m read-only data exchange standard, version 3.0"; december 30, 1990; available from the opa. tel: (614) 793-9660 "an analysis of compact discs as a long term archiving solution" denis oudard, american library association midwinter meeting, alcts-plms' physical quality & treatment discussion group, january 12, 1991 "cd-rom as an archiving medium?" denis oudard, working paper, digipress, february 1991. available on demand. tel: (502) 895-0565presented at the lassist 91 conference held in edmonton, alberta, canada. may 14-17,1991. ' presented at the lassist 91 conference held in edmonton, alberta, canada. may 14 17, 1991. for further information on digipress's century disc contack the author at: 2016 bainbridge row drive, louisville, kentucky 40207 usa. tel (502) 895-0565. summer 1991 73 vol29-3.indd 4 iassist quarterly fall 2005 editor’s notes welcome to the third issue of the iassist quarterly vol. 29. iq volume 29 is still called 2005, while we are now enjoying being in 2006. from the iq-team we do not expect or even hope to be totally current, but we are happy to be gaining on some of our rather persistent backlog. in the past year our gaining has been explicitly due to the work of the co-editors louise corti and wendy watkins, but i also want to thank all the authors of articles for the iq as well as the reviewers / proof readers working behind the scenes plus all the people working on the technical side with layout, printing, distributing, etc. of the iassist quarterly. this work is all undertaken with good will and voluntarily! the first article in this issue is from the iassist conference in edinburgh in may 2005. bobray bordelon is the pliny fisk librarian of economics and finance/ data services librarian at princeton university and presented at the conference his paper: “cross-national & intergovernmental data: paying for one-stop shopping”. the point of departure for the article is that since most organizations lack the resources to build its own interface and provide the ongoing maintenance it is relevant to explore the choices for a commercial aggregator. bobray bordelon examines and characterizes four “aggregators”: datastream international, eiu world data, global financial database, and global insight. the second article is from louise corti associate director & head esds qualidata, outreach & training uk data archive – on “qualitative archiving and data sharing: extending the reach and impact of qualitative data”. louise corti argues that while there is a well-established tradition in social science of reanalyzing quantitative data, there is still not yet a well developed paradigm, nor a pervasive research culture of sharing or secondary analysis of qualitative data. so the article is a promotion effort in that respect and calls for further internationalization of the debate that has begun in some countries. to get started on this debate you can have head start by – among other things in the article – reading about the six main perceived barriers that have been identified through contact with researchers, together with some pragmatic solutions to confronting these barriers. lastly we have an article from our very own world. iassist maintains a very active list-server. the data – e.g. mails – from that list-server was collected some years ago and some categorization was performed on that collection. the article “self reflection of virtuality in a professional association: a compact description of mailing list data” is a description of the iassist organization moving in the direction of a virtual community. we try to offer an answer to the jesting question: “is there iassist life between iassist conferences?”. the ongoing iassist list-server, weblog and periodic iassist quarterly should be a positive answer to that matter. the investigation was carried out by repke de vries, department of public services at the royal library of the netherlands, and myself. a crew of people are taking care of the iassist website so it is constantly evolving. you can take a tour at http:// iassistdata.org and visit the iassist weblog (blog) iassist communiqué – at http://iassistblog.org. the iassist website contains information on previous and coming conferences, including past presentations, as well as easy access to the published articles of the iassist quarterly. articles for the iassist quarterly are most welcome. papers can be generated from iassist conferences, from other conferences, from local presentation, discussion input, etc. contact the editor via e-mail: kbr@sam.sdu.dk. karsten boye rasmussen, june 2006 lassist newsletter, vol. 3, no. 2 (spring 1979) a model for user's service: providing for information and data retrieval from an archival, user, and development library harriet a. dhanak christopher d. brown michigan state university although the newsletter has carred several articles dealing with the various operations of specific archives (the roper center, norc, etc.), the organizations reviewed were establishments devoted primarily to providing archival and library services. many of the members of lassist, in contrast, work alone or with limited assistance, within the context of an academic (or other) department. the article which follows was first presented at the 1979 lassist annual meeting in ottawa and is a description of the way in which a depar tmentally organized archive is confronting its work . --ed i tor . introduction characteristics of the user community and the michigan state univerthis paper will deal with the sity computing system plus our problems of information storage and tasks and needs and those of our retrieval associated with a machine user community, readable data archive located within a teaching/research department, namely the political science department of michigan state univlocal situation ersity. a model, in the process of development and implementation, is because the political data described that hopefully will archive is located within a teachalleviate the problems we have ing and research department, our encountered. the mode of delivery contact is with researchers, of copies of machine readable data faculty and students, during the files will be discussed as well as period that their work is in promodes of user information storage gress. through the pol itometr ics and retrieval. laboratory also located within the department, we will either conduct some subjects germane to the the computer runs or assist the topic, such as cataloging machine users in all aspects of the compureadable data files in a traditer applications. therefore, we tional library, have been dealt find it necessary to address their with extensively elsewhere and will needs as well as needs of the not be discussed here. while there archive, are generic problems facing the archives that deal with machine the present situation on our readable data files, this paper campus is: will focus on the ones that are central to our users and functions of the archive. points to be developed and discussed are the 30 lassist newsletter, vol. 3, no. 2 (spring 1979) 1. we are the only archive of machine readable data files of cross disciplinary material on the campus . 2. the main library is not in the immediate future planning to hold our types of files. at present the code books are not located in the main library, but are in a library located within the political science department. 3. the archive also serves the college of social science and the university community as the supplier of data sets from the inter-university consortium for political and social research. 4. the michigan state university computer laboratory maintains a control data corporation 6500, with an inter data that handles the terminal i/o. a cdc 6400 is available on a limited basis for batch work. a hewlett packard 2000 is maintained primarily for instructional purposes. 5. the computer laboratory offers consultation on system problems, computer language (fortran, cobol, etc.) execution problems and for analysis packages such as statistical package for social sciences and the msu stat package. 6. an interactive multipurpose computing package is not maintained by the computer laboratory. within the department, we need the use of special analytic programs that must be updated with computer system changes and modified for research needs. research problems one problem both for researchers and others using a computer system for information manipulation that differs from a traditional library use is that the functioning unit or system is under constant change, that is, the enviornment within which one works is constantly changing. the computer system itself may be changing with a new model or even worse a new vendor. the existing computer system is altered, hopefully upgraded, and is not necessarily upward compatible. new programs/methods are introduced on an existig system, and the changes may prevent previously executable program from running. this situation presents computer-oriented researchers with an important cost in personal and professional time just to keep up with the changes. this time is above the time consumed learning to use computers as a research tool. for those who constantly use a computer the information is easily recalled and current. however, most researchers will use computers episodically and the development of our model is partially dictated by the characteristics and needs of this type of user. those using a computer must familiarize themselves with some aspects of their use, at least analysis packages such as spss. we would hope to assist such users by enabling them to obtain a data set with a minimal knowledge of tape handling proce31 lassist newsletter, vol. 3, no. 2 (spring 1979) dures, assuming that the data set aspect that we will focus on is his is on tape. suggestion of a machine readable index of holdings. he details the ssdc method of keeping tape file information on a card image file. archive problems we have already used that method, and found that it was not adequate several authors have addressed for our needs. the card image file some of the issues of special could be augmented when new files interest to us, with most noting were created for a study, but basicommon problems. white (1974) discally it was laborious and could be cusses and cites literature that inaccurate, compares the varyng functions between libraries and archives. all archives, large or small. the discussions deal with acquiring with all our variations, are facing libraries and those concomitant a similar problem: the basic probproblems, not with data management lem is the storage and retrieval of problems associated with research various levels and types of inforor secondary analysis. he cites mation. we have probably all gone the difficulties facing researchers through the same process of first in gaining information on the producing a hard copy list of holdmachine readable data files availaings -typed at first, then later ble. placed in machine-readable form. many archival holdings and most the implementation of our system user data sets were on cards and will expand the information on were laboriously carried around available data-sets, and will also with notes and information written archive user library services to on the cards or across the top of the university community. our aim the deck. when data sets were/are is to break the pattern of informastored on tape or disk, a user can tion access noted by white that one no longer visually inspect the data needs personal contacts to gain and an uneasy feeling sets in. knowledge of data set availability. this is the period that technological change in data storage outferguson (1977) notes the types stripped the means used to document of questions data file users ask of files and studies, library and computing organizations. their interest is not in this detailing of developmental partitioned responsibility between problems in archiving does not even the library, computer center and begin to confront the enormous task archive, but in the availability, of indexing variables across stuaccess and documentation of data dies. since we do not have the files. resources to deal with these problems, we will move on to those that grandon (1978) discusses archive are presently manageable. we hold development stages in general and about two hundred and fifty stuthen details the system under dies, and have a library of about development at the social science three hundred tapes that probably data center/roper center. his contain about two thousand files, paper describes functions in which we also have the usual need for the all machine readable data file dissemination of information on the archives must participate to a studies available and have a progreater lesser extent. the one gram that lists them as card images 32 lassist newsletter, vol. 3, no. 2 (spring 1979) by subject and author within subject. archivists and researchers are faced with the problems of increased complexity of machine readable data files, and with multiple files associated with one study title. a survey might consist of one rectangular i zed file of one to two thousand card images, or may contain tens of thousands. the latter presents only problems of bulk. a survey may be updated (new editions) by cleaning or for other reasons. panel studies present additional difficulties by having the problems already noted as well as the addition of new waves. to this list of complexities and problems is difficulty of hierarchical files. at present our biggest problem is to store, retrieve, and have available information on tape files and a machine retrievable document for each study with a description of each file. turtle (1978) also notes that file documentation was deficient in the library at the tennessee valley authority and proposed procedures to improve the existing documentation system. we differ from her in that she states "a computer tape library ... contains only one physical format for information". a serious problem for our archive and users is that we may have multiple formats for files of one study. these may include forms usable by a batch-oriented analytic package while the file may also be stored diffecently for on-line terminal use. file tha dif f icul t by carte that is e ence quan universit levels of tation pi cate name as each e give our t may be updated is not and a method is suggested r and roistacker (1976) nforced by the social scititative laboratory at the y of illinois. however, information and documenus a prohibition of dupli(a search is conducted ntry point for dups) will users more flexibility. model the model that we are developing : 1. assumes minimal computerassociated knowledge by users. 2. will be a method for information storage and retrieval of documentation and/or numeric data. 3. will be a process that will allow modification of the stored indexes* and the programs executing number two above without having to redo the stored sets of archival information. this model can be viewed as a multipurpose system that will serve users who: 1. are browsing for information by author, title subject (general) and clerical procedures, for tracing files and problems with tape, (such as parity errors or the need for cleaning), require an inordinate amount of time and become insufficient and redundant. file documentation for studies with only one *this refers to archival documentation records as opposed to codebooks and other hard copy documentation describing a specific data set. 33 lassist newsletter, vol. 3, no. 2 (spring 1979) geography; 2. are searching for a specific author, subject, or geographical area; 3. may want to retrieve a file that they specify f r om t a pe or disc; 4. have the need to create their own library; 5. wish to hold information and files in the developmental section of the library for creating temporary files and the documentation associated with them. the system will enable archival maintenance and updating of all material, and logging and accounting records of all activities. after trying approaches previous became obvious that prehensive approach ment of information their associated fi enable us to actua procedures. one n ency of only doing multiple times, and dent help one needs tern that does no explanatory time time. some of the ly mentioned , i t we need a comto the manageon studies and les, that would 1 ly simpl ify our eeds the efficia job once, not when using stuan ongoing syst require more than executing frequently not enough time is spent in designing a system, which leads to implementation problems. we have spent much time, thought, and planning for a model that will serve well in the context within which it is anticipated it will operate. the development or use of a system such as riqs (1975) with tutorials would be very helpful to new users, but at the present time it is beyond the scope of the system we will be implementing. a cai is under development at our institution, and will be used if suitable to our application. we plan to develop a simple querying system, not a sophisticated one. our interest in a developmental library and in allowing users to create their own libraries comes from working with researchers and classes and attempting to keep track of all the tape and file activity. as with many aspects of research work, documentation of work in progress is one of our biggest problems. although the problem of classifying, cataloging and documenting studies has been dealt with in the appropriate literature, there is not, to my knowledge, any (or much) literature on systems designed to deal with the information problems associated with developmental file systems for users in an academic setting. those who have large grants may be able to hire their own staff to keep records, but most university researchers must keep their own records of their work as it progresses to completion through various files. at a later period the information on files could remain on a permanant record and accessible to them or it could be scattered . a user crea 1 ibrary wi 11 be g ity of entering a documentation tha the system and wi of limiting or access to any o information. fe noted the relucta to enter informat dies into libr records. our s them the privacy ting thei iven the fl 11 informat t is stand 11 have the pe rmi tting r all 1 ev rguson (197 nee of rese ion on the ary or ystem will they requir r own exibilion and ard for cho ice public els of 7) has archers ir stuarchive allow e and a 34 lassist newsletter, vol. 3, no. 2 (spring 1979) method of documentation they need and a history of their studies previously entered . the information to construct the indexes of information resides on several files. the files contain information in a hierarchical structure with imbedded fields to link them together. this allows us flexibility both with the files (used as data) and programming, that is, one or both can be modified independently. by using the hierarchical file structure of the holding files, we plan to include a list of studies availabe to us, but not located in the archive. this will enable the university community to know all the icpsr studies listed in the user's guide and could be easily updated when we receive periodic notice of new holdings. if we obtain a membership in the roper center, that information would also be entered . at the on-line querying stage any user will be given the information needed to submit a batch job to obtain a copy of the required data. tape numbers and their necessary information (tracks, density, character mode, etc) and file position will not be needed by the user to request a file. file location, whether disc or tape, will be retained in the index and the file retrieved with a simple command followed by pertinent information. for various policy reasons, the computer laboratory at msu will not allow a request for a tape mount to be executed when one is on-line at a terminal. although it is limiting when one needs a small file, it is not as restrictive as it appears because when one is working with a large data set, batch mode is generally preferred due to both time and cost factors. users superv i sio mitted for of a tape a po o 1 of ing their indexing i the job rel ieving entering m index at time (1979 informatio cal manner author izat file. runni n coul them file, tapes anal nfo rma is e them inimal a late ) we n 1 ink , plus ion f i ng und d have to ret we wi ava ilab ysis da tion e xecut io of t inform r date have f ed in a sub j le and er archive a job subrieve a copy 11 also have le for storta and the ntered while n, thereby he problem ation to the at this our files of a hierarchiect file, an a documents the holdings file contains information on the type of system (archive, user library, development) . a twenty character acronym will be the index reference for a study. authors, title, subjects, geography, and date are information that will be used for searching and printout. read and alter passwords can be set at this level. the mapping file contains information on the version (edition, set or subset) , and as many as necessary can be established for a study. this will cover multiple files, updates, etc. this level will also offer a restriction code if needed. date, document, and other user information will be held here . th on t binar etc) . numbe tion file) there versi users wi th archi forma other archi e access fil he form of y, card deck the file 1 r and file of reel in , and num will be as on as is ne will be r permi ssion) ve files tha t that mate form inform val informat e has a fi , book, ocat ion po s i t i o a mul ber of many cessary estr ict to ace t are h the ation ion onl inform le (c spss (tape n and t ipl e acce fo rms ge ed (e essi ng in sta codeb will b y ation oded , file, reel posireel sses . of a neral xcept only ndard ooks . e f o r 35 lassist newsletter, vol. 3, no. 2 (spring 1979) the tape file will contain the center, and main library choose to standard tape information, e.g. build an infrastructure similar to tracks, density, label, mode, those units found at stanford, wi sowner, plus a tape history of consin and northwestern among othaccesses, cleaning date, and number ers, we believe that the existing of mounts. files of information in combination with the executing program structhe authorization file will conture will serve as a base for tain access level, accounting expanded holdings and services, information, logging information and scheduling action if necessary. one is left with a frustrating this file will provide information feeling that there are developments for report generation. the scheoccuring elsewhere that would be of duling algorithms will be particuuse to us and coversely, others larly helpful for instructional might use our efforts to their files. benefit if we could surmount the major problem of program transportthe document files contain the ability to other vendor systems, standard information on each study: we seem to be "reinventing the number of cases and variables, samwheel" at many institutions, pling design, abstract, issuing archive, published material, etc. references we are still formulating a design for a version document section. we carter, m. c. and roistacher, r. c. plan to have documentation accessastatistical and data support to ble in several modes and are still a heterogenous user community, designing this phase. first international s. a. s. users meeting, 1976. those studies under archive entry and control will have all of ferguson, d. social science data the specified information entered. files, the research library and to store a file a user may enter the computing center. drexel all of the above information for a library quarterly 1 3 (1977) , permenant file, or enter only the 70-79. acronym, version and form for temporary files. in this case default grandon, g. m. an archive developinformation will be entered by the ment system: specification and archive system. progress. 1978 lassist annual conference. ittman, b. and borman, l. personconclusion alized data base systems. riqs—remote information query this paper has described an system (los angeles: melville information system still under publ ishing co., 1975), chapter development that will solve some of 2. our problems of maintaining a machine readable data file archive turtle, m. and smart, c. w. the in an academic setting. much planrole of the tva libraries in ning and programmming effort has digital data file documentabeen expended to date on this relation: a demonstration project, tively small problem. if the col1978 lassist annual conference, lege of social science, computer 36 lassist newsletter, vol. 3, no. 2 (spring 1979) white, h. sets : ph .d. dissertation d. social science data california: university of a study for librarians . california, 1974), chapters i, (berkeley, ii, and iii. 37 vol252 iassist quarterly fall 2001 25 introduction my purpose in this paper is to describe the progression of research material, mainly tape-recorded interviews, from the field to the sound archive of folklore and comparative religion i.e. the tku archive at the university of turku, finland. i will describe briefly the processes of cataloguing the interview material and the progression of the original analogue material to on-line use on the internet. also, issues in these key areas will be highlighted: the process of digitisation; the creation of databases; and associated with equipment, technology, and ethical questions. the tku archive the tku archive, founded in 1964, belongs to the department of cultural studies at the university of turku, finland. the department consists of four disciplines: archaeology, comparative religion, ethnology and folkloristics. they have vast collections of traditional material, mainly collected from different areas of finland, but the fieldwork projects in russia, estonia, peru, china, and india have also produced large collections of research material. (mahlamäki 2000.) in the mid-1960s new optimism concerning fieldwork became a significant part of folkloristic research in the nordic countries. it was also a time of new fieldwork ideology. in finland, the new anthropologically oriented folkloristics acquired a foothold especially in turku. the focus shifted from folklore texts to social and psychological aspects of folklore in living communities. many new questions arose. for instance, the problems of learning, performing, transmitting, and interpreting oral tradition became of interest to scholars, as well as discovering new genres of oral tradition. (tku/a/00/79: 1; see nyberg & al. 2000: 491-494.) during the 1960s, the section of folkloristics and comparative religion started several research projects that were directed towards developing research methodology, fieldwork and archiving techniques. the disciplines are very fieldwork-oriented and the main projects have been saami folklore; life-stories and oral history of ingrians; baltic-finnic laments; ethnomedicine in peruvian amazon; oral epics in india; and religious, ethnic and cultural groups in turku area. the tku archive is the third largest traditional archive in finland, and its material consists mainly of approximately 11,000 hours of interviews and numerous photographs, slides, manuscripts, and other original and copied material. (herranen & saressalo 1978, 65; tku/a/00/79:8; mahlamäki 2000.) from the field to the archive from the beginning, the cataloguing system of the tku archive differed from that in other traditional archives because the recording tape was defined as the basic archive unit with its own accession code. each archive unit was then divided, by the means of writing a recording protocol, into data units that were understood as analytical units separate from the entire tape. they were also considered to be the smallest coherent units to be referred to in research reports. usually the data unit was the unit of content which was most often determined by a single folklore motif. (tku/a/00/79: 3, 12; mahlamäki 2000.) for over two decades, each and every research project has, more or less, followed its own principles in archiving materials in the tku archive. by the end of the 1980s, the need arose to renew and standardise archiving practices. in 1988, researchers in the department edited a new kind of cataloguing system based on a filing card named collcard (collection card). the collcard filing system was first of all a medium for storing archive material on computer. its content and structure were based on researchers’ actual fieldwork experiences. collcard is an excellent tool for use in the field. it is a notebook in which one writes the basic data needed for archiving an item. with collcard it is possible to start the archiving process immediately during fieldwork. it also makes it possible to catalogue all kinds of material according to the same system. it contains the most important basic data that must be known about a folklore item of scientific value. collcard was first a paper version that was used in the field and later was converted to a wordperfect macro that was used when cataloguing the material into the tku database, created in 1989. now the collcard form can be found and filled in on the internet and can then be sent to the archive by e-mail. (huttunen 1992; huttunen & al from the field to the net: cataloguing and digitising cultural research material by tiina mahlamäki* 26 iassist quarterly fall 2001 1991; rajamäki 1989, 35; mahlamäki 2000.) the collcard filing system has proven its efficiency during the last decade. firstly, when using collcard during fieldwork, half the archiving process is done before returning home. secondly, the uniformity in the note-taking technique helps combine the observations and data of other researchers on the same project even if the language or culture is unfamiliar to them. thirdly, the collcard filing system can be used in training new and inexperienced researchers. by using collcard during fieldwork, the collector becomes practised in asking the informant the basic questions that might otherwise be forgotten or remain obscure. (rajamäki 1989, 38-39.) the saami folklore project i now describe in more detail the saami folklore research project launched in 1965 in different parts of the lapland of finland, sweden and norway. the ideological background was to prove the existence of living saami folklore (as opposed to the age-old literary sources), and save it in the archive. the main target of the in-depth research project that lasted for many years was the village of talvadas1 by the river of teno. three neighbouring villages were selected to serve as comparison villages. all the inhabitants over 16 years old in the village of talvadas were interviewed several times over several years. (the pilot project 2000; nyberg & al 2000.) during the 1960s and 1970s, the project collected approximately 1000 hours of recording tapes from the whole saami area. most of the sound tapes are in saami language, but some of them are also in finnish and in norwegian. all the material has been archived in the tku archive. additional fieldwork trips have also been made in the 1970s, 1990s, and at the beginning of the year 2001. thus, a remarkable amount of new material has been added to the old material corpus. (the pilot project 2000; mahlamäki & enges 2001.) in addition to the sound recordings, registers and transcriptions, the collection also includes photographs and slides along with a rich and varied collection of manuscripts and maps. during the fieldwork period all the material was catalogued manually. at the moment, all the photographs and slides and also over 200 sound recordings have been re-catalogued into the talvadas-database (see below). the process of cataloguing into the database is still in progress. (the pilot project 2000.) this collection is probably the most extensive sound collection of saami folklore in existence. it forms a valuable and irreplicable whole that would be impossible to recreate today because the oldest informants, born in the 1880s, are no longer living and many dialects of the saami language have now vanished. from the point of view of folkloristics, the material is a very rare ‘thick corpus’ that makes it possible to study organic variation and various meanings of folklore in a local community. (the pilot project 2000; mahlamäki & enges 2001.) from the archive to the internet the material of the saami folklore project is still actively used at the university of turku and there is a lot of interest in it and co-operation in connection with it from several universities in scandinavia. but there are two kinds of problems concerning the optimal use of the material: access and preservation. the material is situated at the tku archive and all researchers interested in it have had to come to turku. also, the original material is suffering from both the unsuitable archiving conditions and use and copying of the material. one solution for both the problems is digitisation of the research material. in a digitised form the material could be approached via an on-line connection all over the world and the original material would remain safe for use by future generations. (the pilot project 2000.) digitisation is, of course, not the only way to save the vanishing tapes. just recording them to new analogue tapes would do the same. but digitisation will provide many advantages in addition to saving the information. use of the material is not dependent on time and place, because the material is available on the internet. it will also be faster and easier to use large amounts of research material because it can be handled, searched and studied by using the archive database. several researchers can also use the research material at the same time. and last but not least, the original material will be safe from damage caused by copying and usage. (kurkela 1999; saarinen 2000, 119120.) talvadas-database the new talvadas-database, created in 2000 at the beginning of the digitisation project, is based on the trip highway program. it is now possible to search for and retrieve the collcards and the material of the saami folklore project via the internet. because of the legislation concerning personal data, and also for reasons of research ethics, access to the talvadas-database is not freely available to all users. the database is in a closed net environment and access requires a special password that is given to researchers by the permission of the tku archive. some indexes, registers and the documentation of the material in ddi model can be made available on the internet for everyone to use and search. the digitisation of the material of the saami folklore research project began autumn 2000 and is expected to take two years to complete. at the same time, the digitisation of all the material of the archives of ethnology, folkloristics and comparative religion has started. iassist quarterly fall 2001 27 in practice, the digitisation project started with the scanning of photographs, slides and transcriptions of the research material. the digitisation of the audio material started in the summer of 2001. the photographs, slides and transcriptions are scanned in two different versions. for the net the scanning resolution is 72 dots per inch (dpi) and the pictures are saved in joint photographic experts group (jpeg) format. the pictures are sufficiently clear to be studied easily in a www environment, but they cannot be printed or published. the second version of the pictures is scanned to the archive in tag image file format (tiff) and their resolution is 300 dpi. the archive version will also be copied on to two different cds — one for use and as a copy of the material and one as a security copy. the archive version of the picture, as well as the transcription, is sufficiently accurate for use in publishing, exhibitions and for other occasions2. problems to be solved. in addition to the technical solutions there is much food for thought in the area of research and archive ethics. one of the main problems is the protection of the confidentiality of informants. rules and solutions for this problem should be discussed together with other researchers and archivists who are concerned with these same problems. the archive databases belong to the area of legislation concerned with personal data (hetil 523/1999; julkl 621/99), because they include data about the informants, for instance, their worldview or religion that should remain confidential according to the law. so it is extremely important that the collected information not be used for purposes other than those for which it was collected: research and education. we also must remember that the archive material cannot be used for commercial purposes. the fast development of information technology has created new ways of using and disseminating the archived material; the ways are so new and so different from the old ones that it has been impossible to prepare for them during the collecting of the material. the informants may have been afraid that the interviews would be broadcast, for example, but global dissemination of the material via the internet is now a possibility for which it has been impossible to prepare. cooperation among traditional archives and researchers working with archived material is, and should be, extremely important. there is an urgent need for common rules and common procedures among those traditional archives that deal with digitised material. in principle, the names of the informants cannot be used for reasons concerning both legislation and research ethics. on the one hand, when we are dealing with older projects, such as saami folklore, the original names of the informants have been used in all the articles and research. so it may be frustrating for the researcher for the archive to start hiding facts that have previously been made available. on the other hand, the informant also holds ‘copyright’ for his or her narratives. it is the product of the informant and his or her community; they might want to be recognised and remembered by it. in such cases, permission for use of the informant’s name should be directly requested. when the digitised archive material is given to the researcher, in other words, when the researcher gets the username and keyword of the database, or the cd which contains the copied archive material, he or she must formally agree to use the material only for commonly agreed purposes and also not to make the material available to another person or institution. when new interviews are made, documented permission for archiving and use of the material should be made, but when we are handling old material, responsibility for the protection of the informant lies with the researcher and the archive. one of the main questions, we usually encounter is, of course, money. when starting the digitisation project, a primary issue was the provision of external funding for the project. but we have to keep the future in mind, too. when we use vast amounts of money for digitising our archive material, we may not want to surrender it and copy it for every possible researcher in the world at no cost. but is it ethically correct to collect money from colleagues? how much are they prepared to pay to consult the research material? how can you put a price on national cultural heritage? sources a) unprinted sources the project 2000. kansallisen kulttuuriaineiston digitointi ja on line -käyttöön saattaminen. hankesuunnitelma. [the digitising and on line -usage of the national cultural heritage. project plan.] the department of cultural studies at the university of turku. [http://www.utu.fi/hum/kultut/hanke/index.html] the pilot project 2000. saamelaisaineiston digitointi saamelaisen folkloren tutkimus. kulttuurien tutkimuksen laitoksen arkistojen digitointiprojektin pilottihanke. [digitising the saami folklore material researching the saami folklore. the pilot project of the digitising project of the department of cultural studies.] the department of cultural studies at the university of turku. [http://www.utu.fi/hum/uskontotiede/collcard/ pilotti.htm] tku/a/00/79. interview of professor lauri honko 16.6.2000. interviewers pasi enges and tiina mahlamäki. audio collections of the tku archive. 28 iassist quarterly fall 2001 b) bibliography arkistolaki 831/94. [the act on archives] henkilötietolaki 523/99. [the personal data act] herranen, gun & lassi saressalo (eds.) 1978. a guide to nordic tradition archives. turku: nordic institute of folklore. honko, lauri 1983: folkloristin keräilytaloudesta. [folklorist as a collector] kotiseutu 4/1983: 170-175. honko, lauri 2000 (ed.): thick corpus, organic variation and textuality in oral tradition. studia fennica folkloristica 7. helsinki: suomalaisen kirjallisuuden seura, 3-28. huttunen, hannu-pekka 1992: data-projektarkivet “abba” en trip databas. herranen, gun (red.): data och folklore. nif:s seminarium om databaserad registerhantering och kommunikation. nif rapporter nr. 7. åbo: nordiska instituitet för folkdiktning, 18-23. huttunen, hannu-pekka & jaakko isojunno & marianna kentala 1991: abba:n alkeet. tku-arkiston atkohjelma. turku. kurkela, vesa 1999: äänitteiden säilyminen on keräilijöiden varassa. [the preservation of the sound recordings depends on private collectors] helsingin sanomat 15.2.1999. laki viranomaisen toiminnan julkisuudesta 621/99. [the act on the openness of government activities] mahlamäki, tiina 2000. äänitearkisto-opas. tku-arkisto. [a guide to the sound archive] uskontotieteen toimitteita 3. turku. mahlamäki, tiina & pasi enges 2001 (in print) from the field to the net: analysing, cataloguing and digitising the material of the saami folklore project. ulrika wolf-knuts (ed.) input & output. the process of fieldwork, archiving and research in folklore. turku: norfa. nyberg, patricia & marjut huuskonen & pasi enges 2000: observations on interview in a depth study on saami folklore. honko, lauri (ed.): thick corpus, organic variation and textuality in oral tradition. studia fennica folkloristica 7. helsinki: sks, 489-536. rajamäki, maria 1989: introducing collcard. nif newsletter 4, 35-39. saarinen, jukka 2000: det digitala arkivet och internet. [digital archive and the internet] ulrika wolf-knuts (ed.) vägen till arkivet. åbo: åbo akademi, 119-132. saressalo, lassi & huttunen, hannu-pekka 1985: äänitearkisto-opas. [a guide to the sound archive] turku. suomen arktisen tutkimuksen nykytila ja strategian suuntaviivoja 1998. [the present state of the arctic research in finland] kauppaja teollisuusministeriön neuvottelukuntaraportteja 4/1998. 1 the village of talvadas is situated quite near of the village of outakoski. 2 you can read more about the technical details of digitisation from the archives website. * paper presented at the iassist/ifdo conference 2001 in amsterdam. tiina mahlamäki, department of cultural studies, university of turku, finland. lassist newsletter, vol. 2, no. 4 (fall 1978) toward creating the professional data librarian al ice robbin university of wisconsin madison abstract this article describes the need for training the professional data librarian, archivist, and information scientist in a framework of social science research and applications and library and information sciences. the university of wisconsin-madison intersession 1978 course was designed to meet this need. the course is described. evaluation involves the collection. since president lyndon johnson dissemination, and secondary analycalled for a "war on poverty" in sis of statistical machine readable his state of the union message of data files (mrdf) . some of these january 1964, we have seen an expofiles are produced in the course of nential growth in the production of individual research projects; othstatistical information in order to ers, by organizations in the course allocate resources at the national, of their operations; and still othstate, and local levels of governers, by ongoing data collection ment; to plan, audit and evaluate efforts funded by a consortium of the distribution of resources; and data users, to ensure adequate planning for human needs and an equitable delivwhile a substantial portion of ery of benefits. almost all fedthese data are probably not useful eral legislation has included for reanalysis, a great body of requirements to collect, analyze, data continues to be useful for and report the findings of data research related to public policy gathering. recent trends in fedanalysis, planning, and evaluation, eral reporting requirements suggest while some of the data have been an even greater increase in the transferred to national archives rate of production of statistical and data centers whose major funcdata during the next decade as the tions are to preserve, describe, need increases for more information and disseminate these data, many of by policy planners and analysts in these data files which are potenboth the public and private sectially rich sources of information tors. remain outside the public domain. most of the data have been colthe problems of access to inforlected as part of the administramation about mrdf, the quality of five record keeping process of data, and the need for good docugovernments, but a large portion of mentation describing mrdf have been the data gathering has been funded issues discussed by secondary anaas part of the research, policy, lysts and data archive staffs for a and evaluation activities of the number of years. one aspect of the federal government. in the social mrdf problem, however, has not been sciences, an increasing amount of sufficiently addressed, and that scientific research activity as has been the insufficient and well as policy planning and 95 lassist newsletter, vol. 2, no. 4 (fall 1978) inadequate concerted national efforts to facilitate access to mrdf through the development of training programs for professional librarians and information scient ists . when iassist was established in 1976, it was with the recognition that members of data archives and libraries needed a vehicle to communicate information about organizing, managing, and disseminating machine readable data files. in the iassist constitution, the 'objectives' include the establishement of training courses for data center personnel (newsletter, 1(1), 1976). this objective not only represents the recognition that data center personnel need assistance, but that established data services have the potential for providing training programs to assist others in understanding the nature of mrdf and the special problems of organization, management, and dissemination associated with this medium. the inter-university constortium for political and social research responded to this objective by holding two workshops on data library management in the summers of 1976 and 1977, as adjuncts to their regular summer program (rowe, 1977). the workshops were taught by carolyn geda. alice robbin, and judith rowe. participants included trained data center personnel and professional librarians who wanted to become more informed about mrdf and integrating them in a library collection. these workshops were a source of satisfaction to their instructors and a good deal of information was communicated and exchanged. but it became increasingly obvious that a workshop was not the best structure in which to communicate a conceptual framework for organizing and managing mrdf. there were few incentives to utilize the computing and data processing facilities, and within the time constraints there was little possiblity of dealing with major social research and applications concepts needed to understand statistical mrdf, data base management, and organizational behavior. in addition, if the ideas were to reach an audience who could most directly benefit from this learning exper ience--the professional librarian—the course had to be taught in a university library school environment, and integrated in the library school and social science departments' cur r icula. why two different departments and indeed ones which rarely communicate with one another? although social scientists continue to demonstrate negative attitudes toward the library profession, it is the library which has the expertise in our society to organize and disseminate information. libraries are a natural environment for mrdf, because it can be treated as an additional informational resource, albeit in a different storage medium. library schools train professionals to handle a variety of informational resources. courses in library automation, systems design, information storage and retrieval, and on-line bibliographic data bases are becoming integral offerings of library schools throughout the country. thus, a ' library school is the natural setting for introducing the concept of numeric or statistical mrdf. but while library schools routinely address problems of textual data in machine readable form, they have not addressed the problems of numeric mrdf and future professional 1 ibrar ians and information scientists are ill-equipped to serve the quantitative bent of today's social scientist. 96 lassist newsletter, vol. 2, no. 4 (fall 1978) on the other hand, most social data library specialists, and science disciplines offer at least computer specialists at the univerone course and at a growing number sity of wisconsin-madison. because of institutions, major course it would be offered during interofferings consist increasingly of a session 1978, it could be viewed as number of areas related to quantia potential course for in-service tative social research. social training for library and archive science departments train students professinals who wished to become in methodology, statistics, survey familiar with this informational research, modeling and simulation, resource, data handling, and the like, subjects which provide the basis for the course was cross-listed by understanding the construction and the school of library science and analysis of mrdf. in fact, persondepartment of economics, and for a nel of most of the data services in variety of bureaucratic reasons was north america and western europe entitled, "micro data collection have been (and continue to be) peomethods in economics." funding for pie who trained in one of the intersession 1978 was made possible social sciences. they have not with the generous support of the been professional librarians and uw-madison, which encourages its have been slow to recognize that faculty and staff to use intersesthe tasks they perform or the probsion as an opportunity to develop lems they encounter in organizing, new courses in response to permanaging, and retrieving mrdf (and ceived scientific and social information about mrdf) are tasks changes within and outside the which have been traditionally peruniversity. the course was formed by reference librarians who designed to meet the needs or work with other media (see, carmiinterests of social scientists, chael , 1978; and robbin, 1978). users, and generators of numeric machine readable data, and profesa course which integrates the sionals engaged in information sertheoretical foundations of library vices, whose present or future resand information science and social ponsibil i ties might include science research is therefore one managing large numeric data files which potentially speaks to the or providing data services or formation of a professional inforinformation about numeric mrdf to mation scientist and manager of users. the objective of the course numeric or statistical mrdf. with was to provide the student with the this in mind, in september 1977, i underlying principles of access to recommended that the data and comand management of mrdf in a library putation center, of which the data and archive setting. students were and program library service is a introduced to social science part, design a course which would research and applications, data respond to a perceived need to collection techniques, computing train professionals (both in the and data processing, statistical social sciences and library and analysis and file handling, policy information sciences) to deal with issues and problems regarding data the explosion of information in libraries, and bibliographic documachine readable form. it would be mentation and control of numeric an interdisciplinary, graduate mrdf. problem-solving was an level course, one semester in integral part of the course, and length, which would draw upon the included exercises in statistical expertise of social scientists, analysis and building a biblio97 lassist newsletter, vol. 2, no. 4 (fall 1978) graphic data base of numeric mrdf . reviews, and archive and library in addition, there was a heavy dose problems and their relationship to of daily readings on which the leemrdf). tures and discussions were based and a final paper (in lieu of an primary r esponsibl i ty for the examination). course was in the hands of martin david, department of economics and lectures included (1) introducdirector of the data and computation to social science research tion center; alice bobbin, head of methods and applications, introducthe data and program library sertion to statistical processing and vice; and al schubert, head of the file handling, orientation to data program consulting service. memlibrary and data processing and bers of the madison academic comcomputing facilities; (2) introducputing center (macc) contributed tion to systems analysis: data their expertise during the course, library as an informat ionmanagement as did experts in data base managesystem, complex data bases, data ment from the department of landsbase management systems (dbms) , cape architecture and the center networking; (3) selected policy for demography and ecology. le,cissues concerning mrdf and the data tures on social science research library and archive; (4) biblioand data collection were given by graphic documentation and control an economist, sociologists, and a of mrdf; (5) planning a data survey methodolog ist . library and information service for mrdf; (6) special issues of concern fourteen registered students and for mrdf and data librarians: auditors began and completed the copyright, strategies for file course. one half of the students preservation and handling. practiwere professional archivists (from cal exercises included (1) statisthe wisconsin state archives) and tical analysis of a mrdf specially library students and the other half prepared for the course, using the were graduate students (primarily spss package; (2) use of a data from developing countries) in ecobase management system; (3) use of nomics, sociology, political scinetworks; (4) data library proceence, business, and history. one dures; and (5) bibliographic docuof the students is a professional mentation and control (building a data librarian from the university bibliographic data base of informaof california at los angeles, tion for mrdf). the independent project (final paper) assignment most of the students had never was a choice of (1) designing a worked with a computer before, but research problem, carrying out 1 imwithin a few days were keypunching ited statistical analysis on the control cards and submitting stadata file prepared for the course, tistical runs. although the course and reviewing the findings in a was highly concentrated and very short paper; (2) writing a short demanding (students attended lecanalysis of the problems of creattures from 8-11 a.m., monday ing a bibliographic data base for through friday, and worked on mrdf; and (3) designing an indepenassignments with the help of a dent project with the approval of teaching assistant from noon, often the instructors (most chose to do until 10 p.m.), enthusiasm and comthis and selected a wide range of mitment never flagged. it was an topics on confidentiality and priexciting time for instructors and vacy, content analysis, book students. a detailed evaluation 98 lassist newsletter, vol. 2, no. 4 (fall 1978) instrument was completed by the students. the recommendation was that the course be integrated in the university's curriculum and extended into a full four week program during summer school or the regular academic year. the response was so positive that dacc submitted a request for refunding for the intersession 1979 program. (at uw , all intersession courses go into competition for funds.) we have recently learned that dacc has successfully competed for funds and the course will once again be offered and cross-listed by the department of economics and school of library science. several changes are anticipated on the basis of what the instructors learned last year. first, the title of the course has been changed to reflect the actual course contents. it will be officially titled, "management of machine readable numeric data for the social sciences." (people were more than a little mystified last year to learn that "economics 615, micro collection methods in economics," would deal with the subjects i just described.) the import of the title change and new course number(s) should not be underestimated: while we have no assurance that the divisional committee of the college of letters and science will approve a new course offering (these changes are very hard to come by), it does suggest the recognition of the need to respond to information and technological changes in our society; universities must continue to broaden their course offerings to meet professionals' needs and to respond to technological change. indeed, the concept of data services, libaries, and archives has too long remained the purview of special support facilities within the university and the notion that services to preserve and disseminate statistical mrdf are a function of the general or special library within the university or in both the government and private sectors has not been widely accepted. second, the course structure was too demanding, for both the students and instructors. the workload will be reduced. rather than building a bibliographic data base, time will be spent analyzing the structure and syntax of its contents. some of the time allocated to the data base will now be devoted to exercises in documentation and control and records management of mrdf (known earlier by its misnomer, "accessioning"). less time will be devoted to understanding the development of statistical software and more time to understanding problems (through statistical techniques) of statistical data (assaying the quality of data). the number of required readings will be reduced. however, the basic structure of the course remains the same. once again, a set of instructional materials will be created, but its size (242 p.) will be reduced . what is evident from the ethusiastic response of the students and instructors is that the course dacc offered last year is much needed. it responds to the recognition that the information explosion must be managed. access and retrieval of the vast quantities of information in statistical machine readable form are becoming important issues in the information and library sciences' professions and social science disciplines. the course also demonstrates that the complexity and dimension of information in numeric machine readable form are such that no one individual has the expertise to teach future professionals to organize and manage collections of statistical mrdf; rather, the approach to teaching 99 lassist newsletter, vol. 2, no. 4 (fall 1978) must be a pi inary o social re archiv ist communica t ise . i that scho mation sc numeric m the non-1 a place i n integ ne, whe searche , and c te thai n the ols of iences ach ine ibrar ia n their rated re th r and omput r par f utur libr will read n pro cur r , inte e quant analys ing spe ticular e i t i ary and recogni able da fession icula . rdisc ii tat ive t, data cial i st expers hoped inf orze that ta and al have note: for further information on the university of wisconsin-madison's course, "management of numeric machine readable numeric data for the social sciences," to be offered during intersession 1979 (may 29june 15, 1978), write the data and computation center, 4452 social science building, 1180 observatory drive, university of wisconsin-madison, 53706. references carmichael, nancy. "commentary on librarians and machine readable data files." special librar ies 69(8), 1978, 306-07. robbin, alice. "the impact of networking on the social science data library." lassist news letter 2(1), 1978, 3-13. rowe , judith. "training the professional data librarian." in h. d. white (ed.), drexel library quarterly 13 (1^) , 1977 , 100-108. 100 the regular labour force survey as a quality survey of the finnish 1985 census by aarno laihonen' central statistical office offinland the paper deals with experiences from using the samplebased regular monthly labour force survey as a quality survey of census data on the economic activity of the population. the method uses record linkage at the micro level between the data of persons in the labour force survey sample and the data of the same persons in the census file. an exact linkage is made possible by the uniform personal identifier used in the finnish population registration system. the use of the labour force survey as a reference quality survey of the census was possible because the survey week and the census week coincided, and because the survey and the census measured the same variables according to the same concepts and classifications. in this way, savings were achieved in the cost of the census quality survey. fot the purpose of the quality survey, after the regular survey interview a few additional questions were asked of a sub-sample consisting of onefifth of the labour fwce survey sample (about 2,300 persons). to ensure as errorless results as possible, the data of the sub-sample were reprocessed after being entered and coded as usual. the additional information on the subsample was utihzed when estimating the final errors on the basis of a micro comparison between the census data and the survey data of the whole labour force survey sample. introduction in finland, modem population and housing censuses have been carried out in 1950, 1960, 1970 and 1980, and so-called mid-decade censuses in 1975 and 1985. the latest census of 1985 was carried out within extremely tight budget constraints imposed on the central statistical office by the ministry of finance. the tight budget constraints also affected the production of census data. new cost-saving devices had to be used. a general outline of the 1985 census to give an idea of the cost frame, the direct costs of the 1980 census amounted to about 80 million marks (about 17 milhon us dollars) and those of the previous middecade census of 1975 to about 26 million marks, in 1986 prices. the total expenditure of the 1985 census was not to exceed 18 million marks. the central point of departure for the planning of the system solutions of the november 1985 census was to minimize the amount of manual work in census data collection and processing. first of all, data collection was minimized by an extensive use of registers and administrative records when gathering the basic census data. this was largely made possible by the comprehensive, high-quality population registration, taxation and social security systems characteristic of all the nordic countries. the use of registers and administrative records in population and housing censuses has increased steadily since the 1970 census. this development has been aided by the widespread use of the uniform personal identifier in different registers and administrative records. a significant improvement in the register situation, which helped a great deal in bringing down the cost of the latest census, was the establishment of a building and dwelling register for finland on the basis of the 1980 population and housing census. this register is operated in connection with the central population register, and it allows the unkage of persons and dwellings. therefore, no questions on housing were needed on the 1985 census form. the census of november 1985 used only one questionnaire, namely for gathering employment data. the central population register was used as the mailing list for the population of working age. the questions and instructions on the censusform are presented in annex 1. the questionnaires were sent out by mail from the central statistical office (cso) and were returned by mail to the cso. no local census organization was used. another special feature of this census was that the census form was preprinted, not only with the respondent's name and address but also, for about half the population, with the name of the respondent's workplace as it appeared in the 1980 census and with the respondent's occupational title obtained from the central population register. persons obligated to respond needed only to report any changes or errors that had occurred in this information. the census forms went directly to data entry, which was carried out as key entry. naturally, only changes and additions on the forms had to be keyed. in this way assist quarterly complete 'pictures' of the census forms were converted to machine-readable form before any other processing operations were performed. this enabled batch mode checking and correcting of the form data, leaving only about 10 per cent of the forms to be checked and corrected manually, on terminals. next, extensive automatic coding was applied to workplace and occupation data. thus, the number of forms requiring manual processing was drastically reduced in all phases of processing. this reduced the total cost of the census and allowed preliminary publication of the most essential census data as early as december 1986. in the final phase of data collection, register data were also used to obtain, by imputation, the census form data of non-respondents. in this way, a satisfactory 98.6 per cent total coverage was achieved for the central data of the census form. this also contributed to the relatively small regional variation in coverage, even though the census was carried out as a direct mail-out, mail-back system without any local census organization and with only one reminder sent to non-respondents. register imputation of questionnaire data was tried for about 139,000 persons (3.7 per cent of the population of working age), 84,000 of whom were non-respondents and the rest persons whose responses were incomplete. the regular labour force survey, the study week of which coincided with the census week, and the 1985 household survey were used as quauty surveys of the census. according to the quality surveys, the general quality of the 1985 census data is significantly better than that of the previous mid-decade census of 1975. the quality of employment data is in part slighdy inferior to the quality of employment data in the 1980 census. the setup of the quality surveys of the 1985 census because of the high-quality, up-to-date information obtainable from the central population register (cpr) in finland, the main purpose of the population and housing census is not to count the population, but to produce data on the economic activity and housing conditions of the whole population. the resident population of the country as registered in the cpr was the population of the census. thus, from the point of view of the quality of the census data, there was not, by definition, any undercount or overcount of the population. the aim of the census quality surveys was to analyze the quality of the data produced on different attributes of persons, dwellings and buildings. however the problem of underand overcount was still relevant for dwellings and buildings because of the shortcomings of these data in the cpr. the quality surveys of the 1985 census fall into two categories: those analyzing the quality of the cpr data (especially the data on household-dwelling units and dwellings) and those analyzing the census form data on the economic activity of the population. the data on household-dwellings units and dwelungs, for instance, were analyzed by comparing the census data with corresponding data from the 1985 household budget survey, an interview-based sample survey of 12,(xx) households. another source of dwelling data was a sample survey of dwellings registered as unoccupied in the cpr. this survey provided information on the overcount of dwellings in the cpr. the quality of the data on the economic activity of the population was analyzed with the help of processing error studies (data entry errors and errors in coding and editing) and a special quahty siua'ey in which the final census data on persons were compared with the checked and corrected data of the interview-based regular labour force survey. some experiences from this survey and the methodology of the survey will be discussed in this paper. a short general description of the regular labour force survey will be presented in the next chapter. the finnish monthly labour force survey the finnish labour fwce survey (lfs) is a sample survey based on a random sample of 12,000 persons selected from among the population aged 15-74 years. data collection takes place mainly in personal interviews carried out by the cso's interview organization. the person interviewed is asked questions about his labour force participation (current activity), employment, unemployment, workplace, occupation, industrial status, time use, days and hours actually worked, overtime and secondary jobs, and normal hours of work. about 94 per cent of the interviews are telephone interviews and five per cent personal interviews. about one per cent of the answers are obtained using a mail questionnaire. the average non-response rate of the survey is about 4.7 per cent structurally, the survey is a sor ma ;as 1 ex the .a 1 as app ipti icti boo y^9^ mg tor ce s ive /ma th les ar tri th r cti tri e a 1 al tio was ten in sc use rop r-g on, xea in in nera to o f use nter er vi ly a rc e in ot eas e val aad or on e val s an pro on nai aes si ve stit ienc s si nat and in n 10 lorm lorm 1 pur be uti toe dat by the •s tech ce l re cal proces ter nai the 3ps of inf and itiona 1 automat and int comp on-ii gram e n t e d a n a 1 y igned 1 iy revi ute f o e p r o g r m p 1 e en e comma online corpora gic pre ation d t ion c pose so lized i a are on car oil n nicai lu ibrary led: b sing f rccessi lie pri or ma tion report g program ed the er active lementin ne inter called bibiiog sis^ _ sy ccany a sea by m r eesear a j. n; 1 n g gi-ish la nds, pr tuton tes set cedures , or.line r f 1 1 n e cf the i for« a r c n; the taloc ing at has into a it its sbrod , it utes which it was connd is ternaraphic ential matted modate riable crdiug ft ware n tne es dea popformaand ibliosystem ng carnal iiy storeneras alsaur us subg this active tobias raphic stem) . r. d has embers en in staff, nguage ovides al mthcory disand which te as ce ow ca ch no wi is wh ut an a wh th ma bu at sd se th re ac me za ma co ve ine sent ers isti en , rpri a " ntra nea, roll uni apel eth thin ts ich iliz d in term ica 6 s inta t mo ions ecti rvic rod se adem nt a tion jor mput tuc cnal iang na arch >ag int tucc s a w witn cation resear ses. star n 1 comp opera na's t versit hill. carol this a time allows e the a co inal ( may be ystem ins a st of are ve com es an hout t ing 1 1 ic dep gencie s. u uni ve er ti import c are comp u le ins scienc cen te network commun ide and adverse varying degrees in the areas o eh, and commer tucc may be d etwork" built ute r facility ted and shared hree major univ y of north car duke universi ina state uni environment , t sharieij pt lo a number or omputer cone nversational ma with telephone remotely ioca installat ion. small service the information carried out by putation cente d by contact he network sys braries, data artaents, state s, and research sers outside t rsities must me from tucc. ant "commercial north carolin ting service: tit ute; and nor e and technol ity repset of of sof educac ia 1 e nescr ibed around a which is by north ersities olina at ty, and versit y. nere exn (tso) users to urrently nner via link-up) ted from tucc staff, al opertne rer's user persons tem repcenters, governorganihe three purchase three " users a educaresearch th caroogy f etne offi at 10 "dat tucc larl af te coop auto woui soft adai da ta vide he sour tucc da ciaiiy k n series d fixes commun i y known r some p erative mate tae d then b ware. a tional files for cat ce ta now do an ty, as rel eif e p iso inf wer alo of into holdings n as gen c u ji e n t n d data b "but w the tuc iminary ort was ucc dire r ocessed for the ormation e coilec ging inf rmat i is er al o. gi anks hat 1 c dir meet i laun ctor ym r irs on ted ormat on r or what is in for mr-045-3 in the s popuect or y. ngs , a ched to which the bps t time, these to proion and j5 lassist newsletter, vci. z, nc. 2 (spring 1978) for better access pcints tc the files including descriptors and inaex terirs. tne represe 1 n s t i t u duke un u c a 1 1 o n g 1 e u n 1 and the (unc) . tne unc joined pract ic the dir dii j.on. vided ana the gram av formati various on-line bv the created in the that t data ba bias as availab tucc ne in tdcc, 1 tralize tact p through ogerati or the sequent formati buffer the sys to mro the tu enhance staff c the aut fer morm the pertain need th same in cqope ntati tior.s ivers al co versi sen four sch this al ie €ct 10 al input loca ailab en wa cf fdisk bps the near he t se wi an o le to twork a net n whi d ref erson out t cna 1 ent ir on s fcetwe tern. rmati cc co the an p omate e fie ret ing t an ha forma rative ves frcm : unc ity. nor mputing ty comp o 1 of graduat 001 or effort arning e n of pr 1 re pre via re i qed {q le tnro s transf line dis and was "update" marc re future , ucc bib ii be c n-ime i all us communi e f f cr the at ch th ca cente utati libra e stu l iora as xperi otess senta mctfe uick ugh t erred ks t then pro cord it i liogr onver ntera ers ty. t in f oi apel roll r , on c ry s dent ry s part ence or ti ve t er edit so. ir o a pro gram data s ex aphi ted ctiv with work ch th erenc nel he st and i e net they pecia en t hav on on mmuni level rovid d sys xibil rieva o a rd co tion. env ere e ce at ate nf or work uoul list he " ing ava ty . or e to tem ity i. o part py v iron meat is no on nter, th key idc become c niation b system, d act a " or enduser on-line liable d would g service their would al and refi f infor icular ersions se data would ities tion graph inho ing, repor data ters data 1lecte cente perf ment they ic/ma use p ma int ting files and base d st rs orm ione wo rc d ur po aim info f libr for arr i throu the s d abo ud u ata b ses ng, c rma ti or ex ar les the f n 1 ghou ame ve , s6 ase such lass on ampl co olio ibrari t the respon but in the 3 lor c as a if ying on a va e, dat uld as wing ; ci jded lo wing hill, na ejinanenter , cience s from cience of a under martin s promi n a is ) p^oincm the master cessed which base. pected c/ilarc to toe file in the like e cene conations racial rokers cons "inas a " and access ata in reatly these users. so ofrement ma tion user • s cf the es and state eibiladdiiblioer tain cquir, and liable a cen— e the record management -including the inputting, editing, updating, and various listings and sort arrangements. cataloging cf records — including shared cataloging (reducing duplication 01 effort), verification of title, author, series, etc . classification of records -including the application of descriptors, construction of a thesaurus, and implementation of various subject classification schemes sucn as the library of congress subject headings. acquisition -including inrormation on contact person, location of data and f:ilt documentation, restrictions (if any) , cost, etc. report g eluding from the an inven local ho catalog ity list thors , shelf li mentatio with a formats display, microfic copy. ener atio the e master tory of lain gs ; records ; s of tit series, st of f n; etc. cnoice such as magnet ne, a n -inxtraction file of their own printed authorles, auetc. ; ile docu-all or print online ic tape j nd hara the rel technology including has cont changes i services p well defin miriam a. university out that technology communicat also had brary fun building, counting, ice" (drak at ion and how ribut n qua rovid ed i drak libr in a all e wit an "i ction orde catal e, 19 at in for data bl€ sy ste need that this bl iog er ate di f f e compu card the mati file thro m. for woul type r aph sue rent ter stoc prese on o s is n ugh a conse a bib d prov of in ic/mar h cata form gener ship libr on-1 ed t lity ed by n a r e (1 an es dditi owing a eac mpact s as r pro oqmq 77) . nt ti between onary functi ine techno o signifi and amount the iibrar ecent pape 977) or pu drake po on to onlibraries h other, it " on such "coilecti cessing and and user s ot cu ny 1 quent iiogr ide forma c dat leg ats , atea me, achin rrent ibr ar aphic libra tion. a bas recor mci per f o catalo eread j.y ava y net there data rians the e can ds in uding rated iiae ons, logy cant of y is r dy rdue ints line to has lions' acer vging able ilawork is a base with bigenmany t he 3 x5 36 newsletter , vcl. mc. 2 (iprin^g 1978) lct by hat by qce , a tai of ref inem sta jes the i"o cesses., stiii g si jned couj-d b lv otae tions c tional and te have mo which 1 avaiiab flies t ing a wna w:.a t th the ent s ct a jnda tn t ve i to fc e r jja cuia rtso chno ved s to ie c a 13 ii. t vou t vou e same iessoii made n y new t ion e woe k. opment bfe d xpanae rlies. be o arces io jica closer prov machi wider mu^^tiv .1 r 1 d 1 noi1 are ab 1 1 e c , s lea ii. eau ta r or d e s c l ai tu pr ct d and pre verccm -bo i, h too ide m ne-r ea audien p u r p g s tl .c le to t ne s r n e a a the 1 vour b future ibed n t it ctype 1 mpie sent 1 e with th oe qwevar ur od j format daoie ce. e b ae asur. nie v3 , pr odutn t5n d t a e nit la i eco.ties s ucere is is aewhich men tea iaitaaddirsonai , we €ctive ion on data euildibiioapni ik e liti auge t pr lenc € da tern hi ilt^ ude sour n v i r o n fcs ope s an a e vious e data ta bas at iona c repr auaien an y ce . data bast ment with o ns up ir. fo data acces ly existed f iles . e according 1 standards esentat icn, ce is exte user of ror a netn-line caparmation ex3 tnat have for social 3y building to existing for bibliothe potennded to ina library references draice, miriam a. impact on on-line systems on library functions. paper presented at the pittsburgh conference on the on-line revolution in libraries, pittsburgh, p a, november 14-16, 1977. nesvold, betty a. instructional applications of data archive resources. american behavioral scientist. it' (tio. t7 hafcn7ipfii 1976) , 445-67. weisbrod, david l. n uc reporting and marc redistribution: their functional confluence and it implication for d redefinition of tne marc format. journal of library automation. "tu 7no. 3, 37 vol21.1 8 iassist quarterly objective statistics norway (ssb) shall prepare and distribute statistical information on the norwegian society. the official statistics shall co-operate with research and analysis in order to monitor and analyse economic and social conditions as well as resource and environmental conditions. it shall thereby provide the public, the industry and trade sectors, and the authorities, with information on the society’s structure, developments and mode of operation. statistics and complementary analysis from statistics norway is a common property which should be used by as many people as possible. statistics and analysis shall therefore be easily available to all, that is, to individuals, authorities, political and other organisations, educational institutions, the media, establishments, etc. the institution statistics norway is a professionally autonomous institution. this means that statistics norway has full responsibility for the professional contents of the statistics and analysis. statistics norway is completely independent in deciding what official statistics shall be published and when and how this should be done. this independence from the authorities and interest groups is crucial as confidence and authority are essential prerequisites for official statistics and also for statistics norway to fulfil her role in the norwegian and international society. the situation during the last decade, statistics norway has had a real cut in her basic allocations from the national budget. the number of employment positions allowed by the national budget have gone down from over 700 to just under 600. in principle, the area of statistics covered by government funding should include a core of statistics and research subjects, defined from what is considered to be most relevant for the norwegian society. that goal is difficult to fulfil with decreasing funding. in certain cases, the establishment of new areas of statistics or the reinforcement and improvement of existing areas of statistics are not covered by current government funds. this will require binding long-term co-operation agreements with the ministries that have a special need for improvements in statistics and research for monitoring and for policy-making in their areas of responsibility. these ministries must cover costs incurred during the development and operating phase of this increased production of statistics. pricing policy the pricing policy of ssb as regards distribution and commissioned assignments shall be consistent with the ruling principle that statistics and analysis should be commodities that are easily available to various users. for commissioned assignments, which may consist of the development of new statistics, preparation of statistics or research and analysis, the guidelines require that the client covers ssb’s real costs, meaning full cost absorption. in certain cases, users may require more detailed or processed figures outside the scope of statistics produced by government funding. the extra work required to provide such information shall be priced at its marginal cost when the statistics is distributed. data collection at statistics norway is done by: n sending questionnaires to “all” persons and establishments n sending questionnaires to or interviewing a selection of persons or establishments n setting up at ssb, registers of persons and establishments n gathering information on persons and establishments collected by other public bodies. organisation of data the objective of the organisation of data is that all users shall have simple and good access to official statistics of norway. one way of answering the challenges from the major users would be to distribute data via databases where the user can make his own choice of pre-defined connections. the databases shall be user-friendly and shall contain metadata (descriptions and definitions of data etc.) while safeguarding the protection of privacy. in certain statistics norway and the social sciences by johan-kristian tønder* 9winter 1998 areas, such databases are under development. data protection in the production and distribution of statistics, confidentiality and the protection of privacy are strongly emphasised. safeguarding the anonymity of the client is a vital prerequisite for the activities of statistics norway. individual data on persons or establishments/enterprises should be treated confidentially and should only be used for statistical purposes and in accordance with the regulations laid down by the norwegian data inspectorate. research and statistics statistics norway carries out research on her own data. research activities shall cover research for social planning in the areas where statistics norway has a particularly central statistical responsibility. as such, we compete with others who use statistics in their analysis activities. however, the research is important and necessary for further development and improvement of the production of statistics in the various areas. distribution to researchers social researchers are major users of our publications and databases. in many cases, researchers provide feedback that contributes towards quality control and improvements of the statistics. however, research-projects often requires other statistics than what we present in standardised form. in such cases, researchers may order separate data selections to suit their purpose. sometimes, a researcher may like to combine his own data with our data. in such cases ssb combines the data and produces statistics or anonymised data files for the researcher. however, research often requires other types of data than normally produced in statistics. in particular, this concerns event history data where one follows the same individual over time and in different social situations such as education, employment, social security benefits, etc. as government funds allocated for the production of statistics are decreasing, it is difficult to get funding for event history databases within this form of finance. we therefore have to try to find other sources of finance. among others, the research sector itself. the nsd agreement in 1993 statistics norway and the norwegian social science data services (nsd) entered into an agreement, with the purpose of providing researchers and students with the easiest possible access to as much of ssb’s data as possible without violating the rules laid down by statistics norway and without reducing the degree of protection of sensitive data. the agreement covers two types of data. firstly, statistical data tables ready to be published from ssb’s statistical databases or specially prepared upon nsd’s request. secondly, anonymised individual data from ssb’s sample surveys or other data from persons or establishments. for these services the agreed price is lower than the marginal cost that is usual for ssb’s services. i must confess that due to inadequate capacity, ssb cannot always fulfil the agreement as quickly as desired. the agreement says that ssb is committed to keep nsd informed about the surveys and statistics that are produced and to inform students and researchers about nsd’s services. nsd shall in return inform them about the services that ssb can offer and ensure that ssb is cited as the source of data published in nsd’s own publications or presented in reports, periodicals etc, by researchers . the agreement has worked positively for both nsd and ssb. nsd has a reliable and reasonable access to important data for social research. ssb has been relieved of part of its services for students and researchers and ssb’s data and the possibilities lying in these data have become familiar for students and researchers. however, in ssb’s view, the arrangement has two disadvantages. firstly, when students are trained to use nsd, the result is that many of them do not approach ssb when they enter occupational life. they still want to use nsd. consequently, nsd’s subsidised service which is arranged for students and researchers, is demanded by groups for which it was not intended. this creates uncertainty about the division of roles between nsd and ssb. secondly, many of nsd’s users are not always giving reference to ssb as the source of their data. often, reference is only made to nsd. subsequently, there is less appreciation for how important a well-established statistical agency is for research and education as well as for the society. the official statistics are taken for granted. at a time when there is a battle for public resources, the lack of understanding may have the consequence that statistics looses priority. here, i refer to basic official funding of statistics norway as i mentioned before. however, altogether, the conclusion is that the agreement works positively for both partners. a non-authorised version of the agreement is enclosed this paper. data preparation at ssb for researchers as i mentioned many users need specially prepared tables for their purpose and these tables are not always available from nsd’s service. these users therefore have to contact ssb for special tables prepared. very often the researchers would like several data sources to be combined in order to produce tables. in such cases, the researcher has to deal 10 iassist quarterly with several divisions and in order to prevent this, ssb shall establish a fixed contact point for such services. this contact point shall be set up at our division for population and housing census 2000 as it has been decided that the norwegian population census shall be carried out by combining various registers. subsequently, this division has to build up the competency required for combining registers and this may also be of use to research. we suppose that in many cases nsd will be the institution that demand such data on behalf of many researchers and students. perspectives ssb has a three-pronged strategy for the provision of services for social research. the first is (of course) “more and better statistics” especially in fields of statistics that can fill out the uncovered areas we have today, both within national accounting and within statistics on demography and social welfare. secondly, to continue to direct our efforts towards cooperation with nsd in the distribution of statistics to students and researchers. we hope that this co-operation can be further developed in connection with the third prong of the strategy which is to increase the efficiency of our readiness to combine registers for research purposes. in this context, we also have to look into the development of databases for event history on persons. these databases should be built on many statistical sources and more annual volumes of these sources. here, however, ssb, social research and social planning probably have to come together to find an appropriate form of finance for what could become the «new gold mine» of norwegian social research. non-authorised version agreement between statistics norway (ssb) and the norwegian social science data services (nsd) on the distribution of data from statistics norway 1. purpose a) the purpose of this agreement is to give researchers and students the easiest possible access to the largest possible amount of ssb’s data without violating the current regulations and without relaxing the protection of individual data. the term researchers refers specifically to researchers affiliated with a university/ college/ regional college in norway or abroad and the institutional sector in norway. the social science institutional sector is to follow the current definitions established by the norwegian research council. enclosure 1 comprises all those on noras’ list at the turn of the year 1991. researchers from similar institutions in other sectors shall have the same rights. if there is doubt about whether an institution belongs to the institutional sector, the matter shall be resolved by the norwegian research council. all persons enrolled at a university/ college in norway or abroad are considered as students. b) the distribution of statistics to other public and private services is outside the scope of this agreement. however, nsd may distribute data on condition that nsd pays an additional fee to ssb (enclosure 3). individual data shall not be distributed to persons or institutions outside the primary target group without ssb’s special consent. c) in addition, ssb and nsd may prepare separate cooperation agreements for the delivery of integrated program and data packages intended for various user groups. 2. scope this agreement covers the delivery of the following types of data: 11winter 1998 group i. statistical tables ready to be published. group ii. individual data, rendered anonymous, from ssb’s sample surveys or other individual data (persons, establishments) also made anonymous. 3. delivery procedures group i data: if data can be delivered from ssb’s statistical databases, they are produced as standardised tables as soon as the underlying figures are registered in the database, and at an agreed price. enclosure 2 gives a summary of prices and the types of data to be regularly delivered in this way. nsd may choose not to have data delivered regularly, but may place an order in each instance. this is specified in enclosure 2. alternatively, nsd may be granted on-line access to retrieve data themselves from ssb’s databases. in the case where nsd orders a set of data that does not exist in the databases but requires special processing, an estimate of the additional work required to deliver the data shall be calculated. nsd shall confirm the order before the assignment is carried out. group ii data nsd shall apply to ssb for the transfer of data from individual surveys. ssb shall evaluate the application and give approval, subject to ssb’s concession from the data inspectorate. alternatively, ssb may propose changes in the selected data and prepare an estimate for the additional work required to deliver the data. researchers shall submit an application to nsd for access to the anonymous data. the application shall specify the project, etc. where the data shall be used. nsd shall process the application and approve or reject it. before access is given, the user must sign a special pledge of confidentiality. nsd shall keep a record of accepted applications. furthermore, this record shall contain project information and a list of any publications that may be based on the material. nsd shall send an update of these records to ssb every six months. for anonymous individual data that is not transferred to nsd, or tables with confidential information, nsd must apply for transfer in each instance and ssb shall determine the conditions that apply for the set of data in question. 4. ssb’s rights and obligations restrictions on distribution and use of any particular set of data and possible claims of return, shall be determined by ssb at any time. ssb shall keep nsd informed on all sets of data from sample surveys. as soon as the questions are finally agreed upon for each sample survey, ssb shall immediately dispatch a questionnaire and give a probable delivery date for the data. more detailed documentation on the data shall be delivered on nsd’s request, as far as ssb finds this possible in practice. ssb shall provide regular updates on contents and changes in ssb’s databases. ssb may demand that reports and documentation on ssb’s data, that are prepared and distributed by nsd, be submitted to ssb for approval. ssb may request copies of delivered data and complementary documentation, against reimbursement. ssb shall contribute to inform prospective user groups about nsd’s services. ssb is committed to deliver correct data that is released for delivery as well as documentation required to use and interpret the data. if any errors are detected during the transfer of data, ssb shall immediately arrange for correct data to be transferred. ssb may make data originally prepared for nsd, available also to other users without notice to nsd. ssb reserves the right to verify that the data is used in accordance with the conditions and requirements stipulated in this agreement. 5. nsd’s rights and obligations nsd has the right to distribute data to the user groups specified under point 1a. nsd may use data from ssb in their own analysis and assignments for users specified under point 1a. for distribution of tables to other users [administration, private users etc. (1b)] nsd shall pay a separate fee to ssb for the data used, in accordance with enclosure 3. when integrated as a part of teaching packages/ test data connected to nsd-stat, nsd may also distribute data from ssb to users other than those specified under point 1a. for distribution to non-primary users, reimbursement for the data shall be made in accordance with enclosure 3, or possibly in accordance with a separate agreement. for all use of ssb’s data, ssb shall be clearly cited as the source, both as a general reference and in association with individual tables where the data is used. this condition applies both for nsd’s own use and usage of data distributed through nsd. nsd shall establish protection for the data and prepare instructions for processing and data storage. data 12 iassist quarterly protection procedures and instructions shall be approved by ssb. nsd shall report to ssb any errors and deficiencies that are detected in the data or documentation. nsd shall contribute to distribute information to prospective user groups about ssb’s services. nsd shall give ssb free access to all teaching packages that are prepared with data from ssb. nsd may have on-line access to ssb-data by paying the same subscription fee paid as schools/ libraries. every six months, nsd shall send a report on usage in accordance with enclosure 4, which will serve as a basis for settling the fee for distribution to non-primary group users. 6. contact persons each of the partners shall designate a contact person. changes are to be continuously reported. 7. validity this agreement is valid from the date of signature and can be terminated with three months notice. enclosures 2 and 3 may be adjusted annually. * paper presented at the iassist/ifdo 1997 conference, may 6th may 9th odense, denmark. johan-kristian tønder, assistant director general, department of social statistics accessing city-county data book via dbase hi: census cd-roms from the ground up fred gey data archivist uc data archive & technical assistance (uc data) (formerly state data program) university of california berkeley, ca 94720 presented to the 1990 conference of the international association of social science information service and technology (iassist90), poughkeepsie, new york, may 30 to june 2, 1990. the bureau of the census has published the 1988 city-county data book data files on a cd-rom disk. physically the data are organized as dbase-hi database files and the bureau has supplied a computer program to display profiles of particular areas. however, accessing the data for analysis purposes (such as finding the ethnicity of american counties) can only be done by directly using dbase-ui, and doing so is a multi-step process. this paper describes how to use dbase-ui direcdy on ccdb88 to select subsets of data for statistical analysis, and compares and constrasts the ccdb88 structure with the 1982 census of agriculture files on test disk 2. some suggestions are made for how the bureau might organize cd-rom data and access software to better facilitate individual access and prepare for the 1990 census. fall/winter 1990 accessing city-county data book with dbase iii 1. introduction in the 1990s the bureau of the census will utilize the cd-rom as a major publishing medium for current and future census data. this process has been underway for several years and the bureau has been through several iterations of test disks before settling on a de-facto standard for distribution file formats of the dbase-hi database system by ashton-tate corporation. while some measure of standardization is imposed by this decision by the bureau, other standards and capabilities need to be added for the convenience of public data users of census data. it is the purpose of this paper to discuss the kinds of tasks which might be commonplace by data users and the effort required to accomplish them, as well as to draw some conclusions as to the structure the bureau might impose to ease the burden of accessing this data. 2. what is cd-rom? cd-rom is a new data storage medium based on audio compact disk technology. each cdrom disk holds about 650 megabytes of data (about the equivalent of 4 computer magnetic tapes as we know them on the ibm mainframe). since the production process is identical to that of compact disks, the cost of production is about the same (less than $5 per disk). cd-rom can be conceived of as an extremely large and slow disk that you can't write on. it operates like a cross between a hard disk (because it's so big) and a floppy (because the disks are removable). on the ibm-at machine in our archive, we have attached a cd-rom drive from denon corp, and designated it to be drive f: thus if we have the city-county data book disk inserted in the drive we can do a directory command to see what files are on the drive: d:\frfd> dir f: *.dbf/w directory listing vo 1me in d r i ve f ii ccdb 1988 of database files directory of f:\ on ccdb cdcif01 dbf cif01dct dbf cif02 dbf cif02dct dbf cif03 dbf rom cif03dct dbf cif04 dbf cif04dct dbf cif05 dbf cif05dct dbf cif06 dbf cif06dct dbf cif07 dbf cif07dct dbf cif08 dbf cif08dct dbf o0f0 1 dbf cof01dct dbf cof02 dbf cof02dct dbf cof03 dbf cdf03dct dbf cof04 dbf cof04dct dbf cof05 dbf cof05dct dbf cof06 dbf cof06dct dbf cof07 dbf cof07dct dbf cof08 dbf cof08dct dbf cof09 dbf cof09dct dbf cof1 dbf cof10dct dbf cof1 1 dbf cof1 1dct dbf cofi2 dbf cofi2dct dbf cof1 3 dbf cof13dct dbf cof1 4 dbf cof1 4dct dbf cof1 5 dbf cof15dct dbf cof16 dbf cdf16dct dbf cofi7 dbf cof17dct dbf cofi8 dbf cdf18dct dbf dct cou dbf dct cty dbf dct plc dbf dct sta dbf plf01 dbf plf01dct dbf stfo 1 dbf stfo 1dct dbf stf02 dbf stf02dct dbf stf03 dbf stf03dct dbf stf04 dbf stf04dct dbf stf05 dbf stf05dct dbf stf06 dbf stf06dct dbf stf07 dbf stf07dct dbf stf08 dbf stf08dct dbf stf09 dbf stf09dct dbf stf10 dbf stf10dct dbf stf1 1 dbf stflldct dbf stf12 dbf stf12dct dbf stf13 dbf stf13dct dbf 84 fil e(f) by ei free d:fred> 36 iassist quarterly 3. what data is available? we currently have data from the census bureau, and will soon obtain a great deal more. the following data files are on hand: disk file geography test1 ahs85 test2 ccdb 1980 census, stf3 1985 american housing survey, national core 1982 census of agriculture 1982 census of retail trade 1988 city-county data book zipcode nat/state state/county zipcode state/county/ city/place table 1 : cd-rom databases at uc data 4. accessing cd-rom data: city-county data book the census bureau has supplied some access software (in the form of canned profiles which can be applied to generate reports for particular areas), all of which is located in the cdrom subdirectory on drive c in the archive. one of these is the ccdb profile program, which can be accessed as follows: c:\ cd cdrom census bureau c:\cdrgm> dir cd-rom access software „ , . . ~ . ^rjr\ -130686 vo 1 ume in drive " direct •« c:\cdrcm ag82 7-12-88 9:34a retail 7-12-88 9:35a zips 8-02-89 8:56a readme 1 666 8-08-88 10:41a rhacme 2 747 8-08-88 10:46a readme bat 58 6-23-88 10:23a tab36 dbf 35815 6-29-88 9:54a agr exe 179692 7-12-88 l:20p ccdb exe 202788 8-24-89 11:26a cdreader exe 39067 8-08-88 10:30a retail exe 83026 7-09-88 l:20p data scr 4008 4-28-88 9:17a menu scr 4103 6-08-88 4:19p calif dbf 276 2-04-90 2:39p 16 file(s) 4771840 bytes free start ccdb profile c:\cdrcm> ccdb fall/winter 1990 37 cursor moves to choose california then bakersfield then land area ================ county and city data book 1988 ===== i states ^ i cities of 25,000 or more alakazarsicoctdedcfl la?u i ga hi id il in ia ks ky la me f bakersfielft .. mdmaminnmsmdmtnenvnh i baldwin park nj nm ny nc nd oh ok or pa ri i bell sc sd to tx ot vt va w w wi i be 1 1 flowe r w : california " ( i bell gardens subjects ww,^_^„ xand,ar«»,,«tid poputai ion vital statistics and health social welfare programs (* not covered in this geographic area) crime and education money income and poverty status personal income (* not covered in this geographic area) hous ing civilian labor force and employment agriculture (* not covered in this geographic area) manufactures construction (* not covered in this geographic area) wholesale and retail trade -cursor < enter-select pgup-page up pgdn-page down esc-reiet choosing the city of bakersfield and the land area and population subject (as shown by shading above) will give the following data about bakersfield: data display for city of bakersfield, ca county and city data book 1988 states __„ i cities of 25,000 or moi al ak az ar'ca*co ct de dc fl larusa ^_„^^^ ga hi id il in ia ks ky la me i> bakersfield m0mamimvmsm3mtnenvnh r "baldwin park ' nj nm ny nc nd oh ok or pa ri i bell sc sd tn tx ut vt va wa wv wi i be 1 1 flowe r •-• ^ufowiiai n-" ^-j i bell ga subjects 'land '"ktiy~azz"^"~~"'~~~~t~~ " land area. 1985 (square milr/ax land area. 1980 (square mi leo-' total persons, 1986 .'.'.'.'.'.'.'." rank of city population, 1986 persons per square mile, 1986 total persons, 1980 (corrected) net change, 1980-1986 percent change, 1980-1986 population characteristics, 1980: percent white percent black . . . . percent american indian. eskim3, and aleut. 78.3 73.6 150,400 109 1,921 105,611 4*.789 42.4 76.5 10.6 1.3 cursor p-print pgup-page up pgdn-page down esc-reset fl-fiag legendl city county data book has data for four levels of geography: states, counties within states, cities (of 25,000 population or greater), and census designated places (including unincorporated towns of 2,500 population or greater for which census data has been tabulated). the example above retrieved the first screen of items available for the city level of geography. 38 iassist quarterly 4.1. ccdb individual data as part of the uc data operations and services on ccdb, the following tasks might be desired: • obtain a data file of all cities and towns in california containing population and percapita income. • determine which counties in the united states have hispanic population greater than 15 percent of total population, and rank them by percent hispanic. in order to achieve these tasks we must use dbase-iti directly. assuming we have started dbase correctly, and then set the drive to f, and the following screen should appear: dbase-hi in assist mode set up create updite poiition retrieve organ i ze mod i fy tools 01 132 152 pm l==========_-,j i. database file . 1 1 ============= f------• i i cif01 .dbf i format for screen ii cif01dct.dbf i query ii cif02.dbf i ii cifoi.dbf i catalog ii cif03.dbf i view ii cif03dct.dbf i ii cifoi.dbf i quit dbase iii plus ii cif04dct.dbf i=============== || ci f05 . dbf i cif05dct.dbf i cif06.dbf i cif06dct.dbf i cif07.dbf i cif07dct.dbf i cif08.dbf i cif08dct.dbf i cdf01.dbf conmandl use fl assist l
l position loptl 1/84 ciion bar select dy. select a database file. 4.1. choosing california towns and cities in order to obtain data for towns and cities in california, we need to know (from the census bureau documentation [census 89]) that the data file we are looking for is the place file which contains incorporated and unincorporated (census designated) places with 2,500 or greater population in 1980. we must choose this file as our dbase data file. this is easiest done by pressing the esc key until we get the dbase command prompt (the dot in column 1) and typing the "use " command, and then using the dbase copy command to actually create the new file. fall/winter 1990 fvmftpmtmr^'g start place level display structured database and structure for database: f:plf01.dbf show structure number of data records: 9593 date of last update : 06/06/89 field field name type width dec 1 stoo character 5 2 mcdcode character 3 3 placecco character 4 4 level character 1 5 areaname character 36 6 \ny93080 numeric 9 7 nny93086 numeric 9 8 mny92079 numeric 7 9 mvy92085 numeric 7 10 mmy93186 numeric 6 1 11 nny92185 numeric 6 1 *• total •• 94 r ' ":-. " i v copy to d:plf01ca.dbf for stco='06' * cop; california 401 records copied subset corrmand line ::plf01 :rec: eof/9593 : enter a dbase iii plus corrmand. and then we can browse the new file to see its contents. to browse p use dtplfoica; -browse ,.,.. i california 1 1 ================ i ================ ========= i ===========| extract file ii cursor <•--> 1 up dow delete 1 insert model ins 1 1 1 char 1 recordi *x "y charl del 1 ex i t 1 end 1 ii field 1 home end 1 page 1 pgup pgdn fieldl "y 1 abortl esc 1 ii panl •<*•> 1 helpl fl recordi "u 1 set optionsl "home 1 1 1=========== 1 ================ 1 i istpl-level msapmsa central state plac1 • areaname state 1060000 1 06 0000 california ca 1060010 2 7362 5775 06 0010 al ameda ca 1,060025 2 4472 4480 06 0025 alhambr a ca 1060085 2 4 , 472 036° ' 06 007° anahe im ca 1060105 2 4472 itto °6 ° 085 1060175 2 4472 4480 ii ° 105 1060180 3 0680 1 06 l\l 5 ant ioch arcadi a azusa ca ca ca 1060185 2 4472 4480 06 0185 1060210 2 4472 4480 06 0210 1060215 2 4472 4480 06 0215 1 "-^rsfield ba ldwm . _ u bell bel 1 flower ca ca ca «4. 1 ibrcwse licif01 1 irecl 40/1008 1 1 ' view and edi t fields. 4.3. selecting hispanic counties mile obtaining california places only presented a mild challenge, the process of discovering us counties with substantial hispanic population concentration requires significantly more detective work. the census bureau does not make it easy because they don't include a comprehensive codebook to document the ccdb cd-rom file, and so we must search for the data element (field in dbase terminology) which has hispanic population. we can begin by examining what documentation the bureau does provide, as shown in the following tableiassist quarterly table b. counties population characteristics and households popuubon t^ace^jx3-co<\ hnmahqala ibm-con. 1980 1ss5 two '—percent— percent 198» l 585 now percent— county under 5 5 10 15 10 2« 25 10 34 yoara 35 10 44 45 54 55 to s4 s5 10 74 yaara 75 rearj and over can indian. esjuno. and alaui paaftc tatanoer hapane' female fanwy oneperson1 14 15 is 17 is is 20 21 22 23 24 25 28 27 2s 29 30 31 " 'hisptiruc paraora irwy b* of any i-icsl 7no spous* pr«»«nt. 'hotaonohmr flying tfasn*. table 2 : ccdb county data elements the numbers above the column definitions refer to a field called item in dbase, where is replaced by the actual column number. looking at the table, we find that percent hispanic population is item25. unfortunately, however, we don't know which of the 18 dbf files that 1tem25 is to be found. we can guess that it's probably in cof02.dbf, and display structure for this data subset, as shown below: p use f:cof82jibf -j 1 1 display structure ; structure of structure for database: f:cof02.dbf second county number of data records: 3191 dbf file dale of last update : 06/16/s9 field field name type width dec 1 stcd character 5 2 level character 1 3 nba character 4 4 p\ga character 4 5 areaname character 36 6 flag14 numeric 1 7 item14 numeric 6 8 flag15 numeric 1 9 item15 numeric 6 10 flag16 numeric 1 11 item16 numeric 6 12 flag17 numeric 1 13 item17 numeric 6 14 flag18 numeric 1 15 item18 numeric 6 16 flag19 numeric 1 press any key to continue... note number of comnand line ::cdf02 :8ec: t/3191 7 county records enter a drase iii plus conmand. but as we can see, this isn't the case. fall/winter 1990 fortunately, the census bureau has not left us completely in the lurch. several dictionary files have been constructed which describe the contents of the data items and where they are located in the many dbf files. ;r:dct oou enter a dbase iii plus cotrmand. i 1 cursor ii char: ii field: ii pan: up record : page: pgup help: fl pgdn i delete i insert mode: ins i char : del i ex i t : "endl i field: "y i abort: esc i i record: "u i set options: "home i i item desc file i item22 percent 75 years and over cof02 iflag22 gdf02 iitem23 iitem24 iitem25 i iitem25a total population used for computing percent's for race and age population characteristics. 1980: percent american indian. eskimo, and aleut percent asian and pacific islander percent hispanic total population (stf-1) used for computing percents for race and hispanic households : cof03 cof03 gof03 ::dct cou :rec: 36/391 view and edit fields. using this dictionary file tells us what data base file to use (cof03.dbf) and which item (item25) to use to search on and create our restricted file. iassist quarterly we can now begin the process of determining which counties had high concentrations of hispanics according the the 1980 census. open the dbf containing percent hlspanic use f:cofl)3.dbfy ' .list stco^reaname,itenj25 « copy to a new file if more than is percent hlspanic s t co 00000 01000 01001 01003 01005 01007 01009 01011 01013 01015 01017 01019 01021 01023 01025 cotttnand line a r c a name united states alabama. autauga, al baldwin, al barbour, al bibb, al blount, al bullock. al bu 1 1 e r , al calhoun, al chambers , al cherokee, al chilton, al choctaw, al clarke. al ::cof03 item25 6.45 0.86 1.13 1.01 1.08 1.12 0.56 1.51 1.33 1.11 0.91 0.53 0.55 0.97 rec: eof/3191 enter a dbase iii plus corrmand. copy to hispanic.dbf fields stco,level,msa,pmsa,areanarae,item23,item24, item25, item2sa for item25>15.0 197 records copied once we have obtained this data file we must sort in descending order using the dbase sort command, and then we can list the contents to find the highest hispanic concentrated counties in the united states. use dbase sort to tiispsrtdhf on ttem25/d '? sort command 100% sorted 197 records sorted to sort to a new * use hispart\dbf "."* " '* "^v.„ rj^"•" file in list stcojrean2me,item25aatem2s,heiii244teni23j descending order records' stco areaname item25a i tem25 item24 i tem23 1 48427 starr. tx 27266 96.93 0.05 0.10 2 48479 webb. tx 99258 91.52 0.09 0.11 3 48247 jim hogg. tx 5168 90.54 0.02 0.00 4 48323 maverick, tx 31398 90.34 0.11 2.35 5 48507 zavala. tx 11666 89.03 0.03 0.08 6 35033 mora. >m 4205 86.56 0.05 0.17 7 48047 brooks. tx 8428 85.99 0.04 0.06 8 48131 duval. tx 12517 85.76 0.0 corrmand line ::hispsrt :rec: 1/197 enter a dbase iii plus corrmand. 5. 1982 census of agriculture a task which one might wish to undertake is to draw information from the 1982 census of agriculture for counties with high hispanic concentrations and combine it with the county data book information just obtained. some of this information can be found in ccdb itself, but more detailed information would lead us to census bureau's cd-rom test disk 2 which contains the entire 1982 census of agriculture. if we pop this disk into our cd-rom fall/winter 1990 drive, we can see how the files are organized. directory of * display files like ag*.dbf ' ag82 data files ag82 0j.dbf ag82_02.dbf ag82_03.dbf ag82 04. dbf on test disk 2 ag82 05. dbf ag82_06.dbf ag82~07.dbf ag82 08. dbf ag82 09. dbf ag82~10.dbf ag82~11.dbf ag82~i2.dbf ag82~13.dbf ag82~14.dbf ag82 15. dbf ag82 16. dbf ag82 17. dbf ag82j8.dbf ag82 19. dbf ag82_20.dbf ag82 21. dbf ag82~22.dbf ag82 23. dbf ag82~24.dbf ag82 25. dbf ag82 26. dbf ag82~27.dbf ag82 doc. dbf ag82_geo.dbf 128747612 bytes in 29 files. bytes remaining on drive. comnand line :: : enter a dbase iii plus comnand. open the first % use ag82_0ldbf dbf file and % display structure 1 examine its structure for database: f: ag82_01 .dbf structure number of data records: 3177 date of last update : 02/03/88 field field name type width dec 1 tab_01_001 numeric 12 2 tab"01~002 numeric 12 3 tab~01 003 numeric 12 4 tab~01~004 numeric 12 5 tab~01~005 numeric 12 press any key to continue... 122 tab_02_010 numeric 12 123 tab_02~011 numeric 12 124 tab~02~012 numeric 12 125 tab~02~013 numeric 12 126 tab~02_014 numeric 12 127 tab_02~015 numeric 12 128 tab~02~016 numeric 12 •• total ** " 1537 comnand line : :ag82_01 :rec: 1/3177 enter a dbase iii plus comnand. what we find is that the agriculture data files, unlike the ccdb data files, don't have any geographic codes in their records! the sole place for the geographic codes is in a separate file ag82_geo.dbf. what this means to the unwary analyst is that the dbase command join can't be used to merge files from these two databases; a substantially more complex program will have to be constructed if we wish to put together data from these two sources, even though they are ostensibly collected for the same counties. we can look at this geographic file and find out what it contains. iassist quarterly browse the geographic reference dbf file for 1982 census of agriculture iise ag82_geadbf1 ..browse &ii*kis 11 cursor <---->! up doan 1 delete 1 insert mode: ins 1 || char: 1 record: 1 char: del 1 exit: "end 1 ii field: home end 1 page: pgup pgdn 1 field: "y 1 abort: esc 1 ii pan: "<"-> 1 help: fl 1 record: "u 1 set options: "home 1 | | ======== 1====== 1===== 1 ======= 1 01 000 alabavft 01 001 autauga county 01 003 baldwin county 01 005 barbour county 01 007 bibb county 01 009 blount county 01 011 bullock county 01 013 butler county 01 015 calhoun county 01 017 chambers county 01 019 cherokee county note number of broase ::ags2_geo :rec-. ' 1 /3177 : counties view and edi t fields . notice that the geographic codes don't have the same name for the two databases. state and county codes are separate on the 1982 agriculture geography file, while they are concatenated into a single stco code in the ccdb database riles. this further complicates the task of merging the two files. finally, the above screen shows 3177 counties in the data file, while looking at the ccdb screen at the bottom of page 6 shows 3191 counties in the ccdb file. while this difference can be explained by the independent cities in virginia and by the bureau's not releasing agricultural data for counties with fewer than ten farms, these facts further complicate the database compatibility problem. thus our preliminary investigation shows merging city-county data book with 1982 census of agriculture to be a complicated process beyond the scope of this paper. the key to making this process easier will be the development of standards for geographic coding. 6. conclusions and recommendations while the census bureau has gone a long way toward making census data available on inexpensive personal computers, certain additional features will make accessibility of cdrom census data to the average planner or statistician using these data. the census bureau, in constructing cd-rom products, should provide • a comprehensive codebook for data dictionaries which not only names and describes each data item, but also gives its universe, so items are not inadvertantly combined. a) an effort should be made to combine items having the same universe into the same dbase file. fall/winter 1990 • uniformity of file structure across databases: a) one record per geographic area b) the same geographic units over all databases (e.g. either one file for all counties in the u.s. or one file per state) c) the same geographic naming structure within each file and across databases (e.g. stco as in city-county data book or state, county as in 1982 census of agriculture). this is so that the dbase 'join' command can be used to connect data from different files or the 'locate for' commands will use the same selection sequence for files being connected. this purpose of this paper has been to give a glimpse of the effort necessary to do more than trivial tasks in retrieving data from the city-county data book on cd-rom. in doing so we have uncovered some of the issues in access to census data. the availability of 1990 census data on cd-rom will surely force social science information specialists to confront these issues in this new medium. the application of standards will resolve some problems, but others can be relieved with additional machine-readable documentation from the census bureau. iassist quarterly the merit networking seminars making your nsfnet connection count merit/nsfnet information services is committed to providing current information on national networking to all users of the nsfnet backbone. toward this end we will sponsor a two-day seminar in ann arbor, michigan, may 20 and 21. "making your nsfnet connection count" will be an informative seminar focusing on issues of interest to campus computing leaders, information systems and networking administrators, educational liaisons, librarians, and educators who want to learn more about national networking. day 1, "real people doing real things," will feature a numberof presentations concerning network applications in education from the elementary grades through the college level. the day 1 activities will begin with a keynote address by paul evans peters, director, coalition for networked information and will close with a tour of the merit network operations center. day 2, "how to get connected and stay connected" will provide local, state, midlevel, and national networking perspectives from the experts. day 2 will also be comprised of presentations on internetworking, information/user services, and network operations. the seminar will be held at the tenneco automotive training and development center in ann arbor. microcomputers connected to regional and national networks will be available on-site sothat attendees may access network resources discussed in the presentations. the registration fee is $395. an early-bird fee of $345 will be charged for registrations received before april 15, 1991. this fee includes the twoday seminar, a reception on sunday evening, lunch on monday and tuesday, all seminar material, and an optional tour of the network operations center. for further information send an electronic message to seminar@merit.edu or telephone 1-800-66-merit. 3480 cartridges. for those interested in 3480 cartridges: ntis (pb8 233135) is selling for $12.95 national archives technical information information paper no. 4 "3480 class tape certridge drives and archival tape storage: technology assessment report." call 703/487-4600 or write to document sales ntis, springfield, va 22161 . this paper covers the "mechanical & technological future of the systems" & "provides valuable information to data center managers data librarians, and archivists, in fact to all who are concerned about the long-term storage of machine-readable data." fall/winter 1990 the revised lassist constitutk harold naugler public archives canada prc^xjsed constitutional changes at the iassist annual meeting held in washington, d.c. , in may 1980, the general membership approved the establishment of a constitutional review committee to review the association's constitution of 1978 and to propose changes. as the chairman of the comnittee, i undertook a thorough review of the 1978 constitution, received recor.iraendations and changes proposed by association members, and incorporated changes approved by the membership at the 1979 and 1980 annual meetings concerning the establishment of standing committees as viell as action and interest groups. the proposed changes were sutsnitted to the ineinbers of the administrative committee for their detailed review in the spring of this year. the proposed constitution and by-laws vhich are described in the following pages incorporate mast of the changes suggested by adrninistrative committee members. for a caaparison of these proposals with the 1978 constitution, members are referred to the iassist newsletter, volume 2, number 1 (winter 1978). tlie suggestion has been made that the proposed changes do not ensure a sufficient degree of continuity in the governance of the association. it has been proposed that such continuity could better be attained as a result of the election of the administrative committee members on a rotational basis and through the vice-president succeeding to the presidency. in order to provide the members of the general assembly with a clear choice on this issue, i have prepared specific articles which are outlined after the main body of the constitution and eby-laws . as i shall be unable tx) attend the annual meeting of the association in grenoble, i would suggest the following course of action: that a motion be made, and seconded, that the proposed constitution and e3y-laws be approved in their entirety; that an article by article review of the constitution and by-lavs then follow, at whid-i time amendments can be made , seconded , and perhaps passed; and after the article by article review has been oonpleted, the membership approve the constitution and by-lavre as amended. under most rules of order, it is customary for the general assembly to move into comnittee of the whole in order to undertake a detailed review of an association's constitution and by-laws. i would like to express ny appreciatioi to all of those who have taken the tine and patience to review the proposals ard make suggestions. it is my hdpe that the result is a constitution witli related by-laws . vh ich reflect more accurately the views of the membership. harold naugler ottawa, canada july 1981 28 itjteraiatiohal associatiotl fdr social scieice iriformatiaj service aijd techiiokxy/l ' association iriteruatiotjale pour i£s services ejt techiiiques d'ltlforflatiai qj scihces sociales comstitutiai aio by-laws (proposed changes: septenber 1981) article i ijafle the narae of this organization shall be the iijterhatioual associatics^] fdr social scietrce nifortlatiaj service aijd techiiology/l'associatioij irjterilatiaiale pour les services et techinques d'ltlformatiom ni sciehces sociales, hereafter referred to as "lassist". article ii headquarters the official headquarters of lassist will be located with the treasurer. article iii objectives all activities of lassist will be based upon the following objectives: 3.1 to encourage and support the establishment of local and national information centers for machine readable data. 3.2 to foster international exchange and dissemination of information regarding substantive and technical developments related to nachine readable data, 3.3 tb coordinate international programs, pro3ects, and general efforts tliat provide a forum for discussion of issues relating to machine readable data. 3.4 tb protkdte the development of standards for machine readable data. 3.5 to encourage educational experiences for personnel engaged in work related to these objectives, article iv activities td accomplish the objectives of lassist, sane or all of the following activities may be conducted: 4.1 cotfllttees atx) groups committees rray be established and groups of members organized to undertake specific tasks, to find solutions to specific problems, to develop and ccmpile relevant materials for specific pro]ects, and to disseminate information en specific subjects. 4.2 ooiiferences, workshops, seminars, trainiic sessiqjs riemlaers may convene organized efforts on any subject consistent with lassist objectives. 4.3 publications a ifewsletter will be published and regularly circulated to all members, as well as to others wishing to subscribe. other kinds of publications may be produced on occasions. 29 4.4 cooperatiai with other orgaiiizatious efforts will be made to cooperate with other organizations in joint projects and activities v^en these are consistent with lassist objectives. 4.5 other other activities that advance the ob:ectives of lassist may be undertaken from time to time. article v meneership 5.1 ihe membership shall consist of regular and student members. 5.2 regular and student memberships shall be open to such persons as are interested in supporting the objectives of lassist. 5.3 membership m lassist shall include a subscription to the newsletter. 5.4 resignations of any members shall becane effective inined lately upon receipt by the treasurer of lassist. resignation shall imply forfeiture of the annual membership fee. article vi fltjallces 6.1 the fiscal year of lassist shall begin 1 january and end 31 december. 6.2 membership fees for regular and student inembers shall be paid annually to the treasurer by 1 march of each fiscal year. 6.3 ivie rate of membership fees rray he changed by a two-thirds vote of the members on a mail ballot, such ballots to be undertaken between october and december of any calendar year, the results of such ballots to go into effect on 1 march of the following year. article vii goveruaice 7.1 getjeral assembly lassist shall consist of a general assembly canposed of all regular and student members. the general assembly will be organized by geographic regions. 7.2 futlctiong asse^ely of the geijeral the general assembly will establish general policies for lassist and elect the members of the adninistrative canmittee, as veil as the officers of the association. each region will, in addition, elect its own administrative officer vho will be known as the regional secretary. 7.3 admihistrative oommittee the administrative conmittee will be the executive body of lassist, and shall be caiiposed of those members elected by the general assembly frcn its membership. the conposition of the adninistrative comnittee will reflect the geographic distribution of the members of 30 lassist and will be based on the number of members in each geographic region. the administrative committee will also include the regional secretaries. the members of the administrative canmittee, including the regional secretaries, will serve a three-year term and may be considered for re-election cor no more than three consecutive terms. 7.4 fui-lctiorb of the admitjistrative qommittee the administrative connittee v/ill iidplement policies, develop future directions, and coordinate activities for iassist. the administrative coimiittee will organize the general assembly into geographic regions, determine the number of mministrative committee members from each geographic region, and call i:>eetings of the general assembly at least once every year. the administrative committee will also establish coimittees and groups as required. 7.6 role of the officers the officers of lassist will be responsible for the conduct of business of the associatiaj between meetings of the administrative canmittee. article viii meetings 8.1 the annual meeting of the general assembly shall be held at a time and place chosen by the adninistrative cormiittee. 8.2 special meetings of the general assembly may be called by the adninistrative committee. 8.3 the secretary shall give notice to the members as to the time and place of the annual ireeting or special meeting not less than two months prior to the scheduled meeting . 8.4 a quorum shall consist of 20 members. article ix electigjs 9.1 a nominations and elections coimiittee will be appointed by the administrative canmittee. 7.5 officers of the associatiq] the general assembly will elect from arong its martoership the officers of lassist: presidein:, vice-presidetlt, secretary, and treasurer. the officers shall serve a three-year term and may be considered for re-election &dr no more than two consecutive terms. 9.2 the iternmations and elections canmittee shall conduct an election in each geographic region for officers of lassist, members of the administrative canmittee, and the regional secretaries. members withm each designated geographic region shall only be entitled to noninate and vote fidr the regional secretary in their home region. however, all members will be entitled to 31 nominate and vote for the officers of lassist and the other members of the administrative committee. social science council upon dissolution. article xii by-laws 9.3 a public call for naninations will be sent out by the nominations and elections i committee. voting will be conducted by mail ballot. elections will be held every three years. article x af-lhjdfletjts the constitution of lassist may be amended by a two-thirds vote of the members on a mail ballot, such ballots to be undertaken between october and december of any calendar year, the results of such ballots to go into effect at the following year's annual meeting of the general assembly, provided that: 10.1 notice of tlie proposed amendments shall have been given in writing to the standing conraittee on i constitutional review with the ' written support of at least ' five{5) members in good standing of the associatiq]; .' and 10.2 two month's notice of the proposed amendinents is given in writing to all members of the i', associatiom prior to the conduct of the mail ballot. article xi terhitjatigj lassist may be dissolved by a majority of t_he members. all property and funds of lassist will be transferred to the international sectig] 1 duties of the president 12.1 the president shall: (i) be the principal officer of lassist; (ii) provide leadership and guidance in the realization of lassist 's objectives; (iii) preside at all meetings of the general assembly and the administrative cormittee; (iv) be an ex-officio rtember of all standing contnittees and shall coordinate their activities; (v) represent lassist in its dealings with external bodies and agencies, particularly those at the international level; and (vi) report on the state of lassist at each annual meeting of the general assembly. sectiori 2 duties of the vice-presideijt 12.2 the vice-president shall: (i) perform the duties and exercise the powers of the president in the absence 32 or disability of the latter; (ii) assist the president in reccnutiending ineasures to further the ob3ectives of lassist ^hen and as often as requested; (iii) be an ex-officio member of all action and interest groups ard coordinate their activities, and be responsible for proposing the coordinators to t)ie administrative cormittee and maintaining regular contact v^ith such action and interest groups throughout the year; and (iv) in the event of the resignation, death, or incapacity of the president, succeed as acting president for tlie duration of the then pres ident ' s tena . sectioil 3 duties of the secretary 1/1.3 the secretary shall: (i) attend meetings of the adrainistrative ccsimittee and meetings of the general assembly and shall record all facts and minutes of all proceedings in the books kept for that purpose ; (li) be responsible for the maintenance of lassist 's records and for its general correspondence ; (iii) te an ex-officio meraber of the itaninations and elections caranittee to maintain lists of ncninees for office and to assist in the preparation and distribution of ballots; (iv) be an ex-officio member of the standing comnittee on constitutional review to maintain notices of proposed amendments to the association's constitution and to assist in the preparation and distribution of ballots; (v) give notice of all meetings of the general assembly and of the administrative committee; and (vi) perform such other duties as raay be prescribed by the administrative catimittee or president. sectiot] 4 duties of the tre.asijrer 12.4 ihe treasurer shall: (i) have the custody of the funds and securities of lassist and shall keep full and accurate accounts of receipts and disbursements in books belonging to lassist and shall deposit all monies and other valuable effects in the name and to the credit of lassist and in such depositories as may te designated by the administrative ccmmittee 33 from time to time; (ii) disburse the funds of lassist as may be ordered by the administrative canmittee ; (iii) render to the mtiinistrative cconittee at its various meetings, or whenever the members of the administrative committee may require it, an account of all his/her transactions as treasurer and of the financial position of iassist; (iv) prepare a written report for sulxiission to the general assembly at its annual meeting; (v) provide the standing gcmmittee on membership with up-to-date mailing lists of all members in good standing in each of the geographic regions; and (vi) perform such other duties as nay fran time to time be determined by the administrative committee. sectioti 5 duties of the rex^ioijal secretaries 12.5 the regional secretaries shall : (ii) provide leadership and guidance in the realization of lassist's objectives in their respective regions; (iii) represent lassist in its dealings with external bodies and agencies, particularly those at the national level; (iv) serve as manbers of the standing conr.iittee on membership; (v) attend all meetings of the general assembly and the administrative canmittee: and (vi) work closely with the program director of the annual meeting when the latter is scheduled in their particular region. sectiou 6 duties of appointive officials 12.6.1 the editor of the newsletter shall: (i) be appointed by the president of iassist, on the advice of the standing coiimittee on publications and with the consent of the administrative cortmittee, for a term of three calendar years which may be renewed ; (i) be the primary officers of lassist in their respective regions, working closely with the president of lassist; ( i i ) serve on the stand ing canmittee on publications; and 34 (iii) be responsible for t±ie regular preparation, publication, and distribution of lassist's official newsletter. 12.6.2 ihe program director of the annual meeting shall: (i) be appointed by the president of lassist with the consent of the adninistrative corrtmittee; (ii) set up and organize the next annual meeting following the appointment; (iii) be responsible for keeping the /^ministrative committee regularly informed of all preparat ions ; and (iv) vvork closely with the regional secretary in tlie region in v;hich the annual meeting is to be held. sectiai 7 cofuliitees 12.7.1 the adninistrative committee at the time of tlie annual meeting of the general assembly shall appoint and/ or confirm standing committees .and shall appoint and/or confirm chairpersons of t)ie said standing corrnittees. 12.7.2 standing corrnittees shall advise the administrative conmittee on matters of policy within their particular sphere, and shall have a chairperson appointed for a three-year term which may be renewed, two members drawn fran the regular membership of lassist appointed for a three-year term which may be renewed, one member of the administrative cormittee appointed for a three-year term which may be renewed unless representation from the mministrative conmittee is already included in the conposition of the standing conmittee in another capacity, and such officers as are designated ex-officio members. 12.7.3 the standing committees of lassist are the following: (i) cofistitutiaial review oomittee: responsible for receiving proposals for the enacting, amending, and repealing of the by-laws of lassist and for preparirg revised articles and by-laws for members' approval, as well as for undertaking an annual review of the constitution and by-laws and proposing amendments as it deems appropriate. (ii) educatiaj cdmiittee: resfxinsible for the development and advancement of professional prograr^is in education and training and for advisiny the administration conmittee on the criteria for the approval and certification of such programs. 35 (iii) metibership committee: responsible for recruiting membership in lassist, maintaining an up-to-date mailing list of all members in good standing in each of the geographic regions, and for recorrmending alterations in the classes of membership and dues. this committee's membership shall also include the regional secretaries. groups and for every action group so appointed a coordinator shall be named. 12.8.2 a minimum of three{3) nembers of iassist may make application to the administrative committee for the establishment of an action group at least one month prior to the annual meeting of the general assembly. (iv) tjomdjatiofb ai€) elections committee: responsible for receiving nan inations for the election of the administrative cormittee, the regional secretaries, and the officers of iassist, distributing ballots and electoral information according to regulation, and for recommending alterations in procedures . (v) publicatiotls committee: responsible for advising the administrative committee on general publications program policy and for reviewing manuscripts submitted for publication. this comnittee's manbership shall also include the editor of the ttewsletter. section actiaj groups 12.8.1 the administrative committee, at the time of the annual meeting of the general assembly, may appoint action 12.8.3 action groups shall be expected to undertake specific tasks, to find solutions to specific problems, or to develop and canpile relevant iraterials for specific projects. the mandate or terms of reference of action groups shall be clearly defined, including the resources and time required and the specific nature of the output or product . 12.8.4 action groups shall report to the adninistrative comnittee through the vice-president on matters relating to their particular sphere, and shall have a coondinator appointed for a one-year term which may be renewed, two or more members of lassist appointed for a one-year term which may be renewed , one irenber of the administrative comnittee appointed for a one-year term which may be renewed, and such officers as are designated ex-officio meml^ers . 36 gectiai 9 irjterest groups 12.9.1 ihe administrative ccnmittee, at the tine of the annual meeting of the general assembly, may appoint interest groups and for every interest group so appointed a coordinator shall be named. 12.9.2 a minimum of five(5) members of lassist may make application to the administrative canmittee for the establishment of an interest group at least one r.ionth prior to the annual meeting of the general assembly. 12.9.3 interest groups shall be expected to disseminate informaticn on specific subjects and to serve as a forum of discussion between as laell as during annual ineetmgs. 12.9.4 interest groups shall report to the administrative coiiaittee through the vice-president on itiatters relating to their particular sphere, and shall have a coordinator appointed for a one-year term which may be renewed, four or more members of iassist appointed for a one-year term which may be renewed, and such officers as are designated ex-officio inq±)ers . sectiotl 10 uomiiiatious alid electioiis procedures any regular member in good stand irq is eligible to hold office in iassist. 12.10.1 the administrative conn ittee (i) every three years, colttiencing in 1981, the administrative canmittee shall be elected fron a slate of candidates put forward by the standing canmittee on nominations and elections. (ii) during the first two weeks of october in any election year, any manber in good standing may submit in writing to the dominations and elections comnittee, tiie names of as many as five(5) persons for the adninistrative coirmittee regardless of the geographic region in which the ncninees reside. (iii) all naninations must be acconpanied by a written statement fran the ncminees declaring their willingness to stand for election; the signatures of two(2) additional members in good standing who have agreed to co-sponsor the nanination; and an outline of the qualifications of the ncrainees. (iv) the tlominations and elections conmittee will ccrnpile a list of ncminees and mail ballots to the membership during the first two weeks of november in 37 any election year. (v) all members in good standing, regardless of the geographic region in which they reside, shall be eligible to vote for a limited number of nominees froq each geographic region. ihe number of naninees fran each region will be specified on the ballot, based on each region's percentage of the total membership of lassist. voting will take place over a period of one month during any election year, but in no instance will it extend beyond mid-december. (vi) the results of the election shall be announced by the end of december in every election year. the results shall be published in the first issue of the newsletter following the election. (ii) during the first two weeks of october in any election year, any member in good standing in a particular geographic regicn may suhnit in writing to the naninations and elections committee, the name of a person for regional secretary vho must reside in the same geographic region as the nominator. (iii) a nonination must be accanpanied by a written statement from the noninee declaring his/her willingness to stand for election; a statement indicating that the noninee has institutional support to undertake the duties; the signatures of two(2) additional members in good standing froti the same geographic region uho have agreed to co-sponsor the nonination; and an outline of the qualifications of the noninee. (vii) itewly elected meriibers of the administrative committee shall take office after the annual meeting of the general assembly following the elections. 12.10.2 the regional secretaries ( 1 ) every three years , conmencing in 1981, the regional secretaries shall be elected fron a slate of candidates put forward by the standing conmittee on ifcminations and elections. (iv) the ifcninations and elections ccramittee will compile lists of nominees and mail appropriate ballots to the membership of each geographic region during the first two weeks of november in any election year. (v) all members in good standing in each geographic region shall be eligible to vote for the regional secretary for that particular geographic region. voting will take 38 place ever a period of one month during any election year, but in no instance will it extend beyond nid-december . (vi) the results of the election shall be announced by the end of december in every election year. the results shall be published in the first issue of the newsletter following the election. (vii) [fewly elected regional secretaries shall take office after the annual meeting of the general assembly following tl:e elections. 12.10.3 the officers of the association (iii) all nan inat ions must be accompanied by a written statement fron the ncninees declaring their willingness to stand for election; a statement indicating that the ncninees have institutional support to undertake the duties; the signatures of two(2) additional members in good standing who have agreed to co-sponsor the ncrainaticn; and an outline of the qualifications of the ncninees. (iv) the nominations and elections cormittee will conpile a list of nominees and mail ballots to the memldershio during tlie first twd weeks of ttovember in any election vear. (i) every three years, ccmencing in 1981, tiie officers of lassist the presidhjt, vice-presideiit, secretary, and treasupver shall be electee] from a slate of candidates put forv;ard by the standing comittee on tlominations aiid elections. (ii) during the first tv» weeks of octoter in any election year, any member in good standing may subnit in writing to the nominations and elections committee, the name of a person for the position of pricsidetit aixl/or vice-pl^sidettt and/or secretary and/or ti^asurer. (v) all members in good standing, regardless of the geographic region in which they reside, shall be eligible to vote for the presidetit, vice-pregidetrr, secretary, and itleasurnr. voting will take place ever a period of one month during any election year, but in no instance will it extend teyond mid-december. (vi) the results of tlie election shall be announced by the end of decei"aber in every election year. ttie results shall be published in the first issue of the itei-jsletter following the election. (vii) ilewly elected officers shall take office after the annual meeting of tlie general assembly follo'^ing the elections. 3q (other or alternative proposals ensuring greater continuity in the governance of the association) article vii govertlatjce 7.3 administrative committee the administrative committee will be the executive body of lassist, and shall be ocoposed of those members elected by the general assembly from its membership. the composition of the administrative connittee will reflect the geographic distribution of the members of iassist and will be based on the number of members in each geographic region. the administrative committee will also include the regional secretaries. the members of the administrative committee, including the regional secretaries, will serve a four-year term and may be considered for re-election for no more than two consecutive terms. the members of the administrative cormittee, excluding the regional secretaries, will be elected on a rotational basis, half of the members fran each geographic region being elected every four years and half being elected every tvro years between. 7.5 officers of the association succeeding to the presidency of the association after the expiration of his/her term. the secretary and treasurer shall serve a two-year term and may be considered for re-election for no more than two consecutive terras. article ix elections 9.3 a public call for naninations will be sent out by the nominations and elections coimittee. voting will be conducted by mail ballot. elections will be held every two years for the positions of vice-president, secretary, and treasurer, and for half of the membership of the administrative ccramittee. article xii by-lams section 2 12.2 the vice-president shall: (v) autoiiatically succeed to the presidency of the association after his/her two-year term of office. sectig] 6 duties of appointive officials 12.6.1 the erlitor of the ifewsletter shall: ihe general assembly will elect from among its mer.ibership, the officers of lassist: presidet^rr, vice-presideot, secretary, and treaslirer. the president and vice-president shall serve a two-year terra, tlie vice-president autonatically (i) be appointed by the president of lassiffr, on the advice of the starkling committee on publications and with tlie consent of the adninistrative conmittee, for a terra of two calendar 40 years which may be renewed; sectiai 7 cotmiltees 12.7.2 standing conmittees shall advise the administrative committee on matters of policy witliin their particular sphere, and shall have a chairperson at^pointed for a two-year terra which may be renewed, tvk> members drawn fron the reqular membership of lassist appointed for a two-year tenn which may te renev^ed, one raember of tlie administrative committee appointed for a two-year term v/hich may be renev^ed unless representation from the administrative corrrnittee is already included in the ccmposition of ttae standing ccnnittee in another capacity, and such officers as are designated ex-officio mer.itiers . scctiotj 10 rjomitlatiohs atro llectious pkxedores 12.10.1 the administrative committee (i) lvery two years half of the administrative canmittee shall be elected frcm a slate of candidates put forward by the standing connittee on nominations and elections. ixiring the 1981 election year, menbers will vote for a full slate of mministrative canmittee candidates, half of the elected candidates to serve for four years and half to serve for two years. the number of votes cast for each candidate will deterraine the length of office. thereafter, cx)nniencing in 1983, the procedure for a rotational adninistrative cormittee will be in effect. 12.10.2 the regional secretaries (ii) every four years, comnencing in 1981, tlie regional secretaries shall be elected fron a slate of candidates put forward by the standing committee on nominations and elections. 12.10.3 the officers of the association ( i ) every two years the vice-presideot, secretary, and treasurer shall be elected from a slate of candidates put forward by the standing cormittee on uaninations and elections. ixiring the 1981 election year, members will vote for a full slate of officers presidettt, vice-presideijt, secretary, and treasurer. all officers will serve for two-year terms, except the vice-presidein: who will automatically succeed to the presidency at that time. thereafter, corrmencing in 1983, the procedure for tlie vice-president succeeding to the presidency after his/her two-year term of office will be in effect. 41 iassist quarterly 2013 45 iassist quarterly the ddi matures: 1997 to the present by mary vardigan1 abstract the data documentation initiative (ddi) began in 1995 with a small international group coming together with a focus on social science metadata. the group moved quickly to develop a first specification, around which a community of practice emerged. that community, along with the ddi specification itself, has evolved over the last two decades to reflect developments in the social sciences, technological advances, and innovation in research practice. this paper chronicles the history of the ddi from its instantiation in xml in 1997 to its current status as the de facto standard for documenting data in the social and behavioral sciences. keywords documentation, metadata, standards note of acknowledgment the centrality of good documentation to effective social science has been a key tenet of the iassist vision throughout its history. since hosting the inaugural meeting of what was to become the data documentation initiative in quebec city in 1995, iassist has nurtured and promoted the ddi effort every step of the way. indeed, the ddi story cannot be told properly without acknowledging the support of the iassist community through the years. the pioneering efforts and leadership of sue dodd deserve special recognition. her work in data citation and structured metadata inspired a community and set the stage for developments like ddi. fittingly, in 1993 sue received the first icpsr-sponsored warren miller award for meritorious service to the social sciences, which recognizes contributions to essential infrastructure. we continue to benefit from her foundational work and wisdom. introduction when our story left off (see ann green and chuck humphrey’s article in this issue), the data documentation initiative (ddi) specification had just been translated from sgml to xml in 1997, and it was moving toward official publication. there was a strong sense in the community that ddi xml, given its rich and structured nature, could be used to drive process, and that its coverage could and should be extended to document more complex datasets. what happened next? how did the ddi committee go on to enhance and augment the specification to meet rising expectations? how did the community of practice grow to encompass users in over 70 countries around the world? and how did infrastructure surrounding the ddi, including sustainable support for the organization itself and tools to make use of ddi, come into being? this paper describes the ways in which the ddi initiative addressed these challenges from 1997 to the present, ending with a view into what the future holds for ddi as we move forward. publishing ddi version 1 as we have seen, ddi began as a volunteer effort, drawing on metadata expertise and interest from across the social science research community. in-kind contributions made it possible for ddi to be instantiated as a specification with a user community actively working around it. however, in-kind contributions cannot provide the type of sustainable structure needed to fund face-toface meetings and development work, and thus the ddi committee and its founders decided to pursue external funding streams. in 1997 the inter-university consortium for political and social research (icpsr) applied to the national science foundation (nsf) under 46 iassist quarterly 2013 iassist quarterly its “infrastructure in the social sciences” program and received an award that included funds for enhancing and testing the ddi specification. having convened the first ddi committee two years earlier, icpsr was the “home” for ddi, providing administrative and substantive support. the beta-test of the ddi dtd began in march 1999 and continued until august. betatesters2 received financial support to test the specification and to report on their findings. at the conclusion of the beta-test, a list of changes suggested by the testers was compiled and subsequently reviewed at a meeting of the ddi committee held in october 1999. version 1 of the dtd, incorporating these changes, was published march 24, 2000. extending version 1 with version 1 published and external support for ddi activities ending, the ddi committee needed to find new funding sources. while an independent review funded by nsf found that ddi was a “worthwhile scientific effort that filled an urgent need for standardization of social science technical documentation and interoperability” (indeed, one evaluator termed it “a strategic component of the infrastructure necessary to support the exchange of structured social research survey data” [icpsr 2001]), it was difficult to obtain funding for this type of endeavor. health canada came to the rescue, providing a substantial amount of financial support during 2001-2002 to enable the ddi committee to meet and to make improvements to the specification, most notably additions related to aggregate data and geography. aggregate data was the focus of a small working group meeting in april 2001 in voorburg, the netherlands. agreement was reached during that meeting on a draft aggregate data model (also known as “ncubes”), which was reviewed by the committee at a meeting in washington, dc, held in june 2001. committee members began testing the new version 1.02 of the dtd, with the extension describing aggregate/tabular data. several other changes were made to studyand variable-level elements. by march 2003, version 2 of the specification was published, with these enhancements: • aggregate extension elements (ncubes and location map) • internal formatting elements from the tei specification to permit formatting within elements • new geographic elements: geographic bounding polygon, polygon, point, g-ring latitude, g-ring longitude, and geographic map the ddi alliance emerges it was becoming clear that generating external funding to support continued development of the ddi specification would be challenging so another approach was considered: becoming a self-sustaining membership alliance, modeled along the lines of the successful world wide web consortium. in june 2002, the ddi committee met in storrs, ct, in conjunction with the iassist meeting. the main focus of this meeting was a draft charter, written by richard rockwell, to create a ddi alliance – a new membership structure and funding base that would provide support so that the initiative could continue. the charter document provided for an expert committee with representation from the ddi alliance membership, with each member of the committee having a vote and thus a say in the future of the ddi. a steering committee to provide oversight was also established via the charter. the final meeting of the original ddi committee3 was held in february 2003, in washington, where the group approved numerous changes to the dtd leading to the publication of ddi 2.0 (see above). an open meeting of the ddi alliance was held in conjunction with iassist in ottawa in may 2003. meeting participants discussed the new alliance structure and elements of a strategic plan for the next three years of the alliance. toward a lifecycle specification meanwhile, expectations around what the ddi could document were growing. a page on the ddi web site (“about the specification”) in 2003 provides this vision for the ddi: “the ddi aims to be the foundation for collection, distribution, use, and archiving of many future data collection projects in the social and behavioral sciences, across institutions, countries, and disciplines. it also aims to be the basis for retrofitting documentation of older studies for improved ease of use and stronger guarantee of archival preservation.” at its first meeting in 2003, the new ddi expert committee picked up this vision and set an agenda for the future that was ambitious and comprehensive. first, the committee discussed the need for a data model. it was generally agreed that the xml document type definition (dtd) for the ddi had limitations: it was not as modular and easily extensible as it should be and it had not been thoroughly reviewed for internal logic. having a model, most likely in universal markup language (uml), to reflect the underlying design and structure of the specification would represent a big step forward. with such a model, the ddi could be expressed as xml schema, rdf, a dtd, or possibly other formats. the health canada/nesstar partnership had already done some work on a data model for its dais/nesstar software that harmonized the ddi with iso 11179. employing xml schemas to express ddi was also discussed. moving the dtd to a schema to take advantage of the modularity in schemas, the capability for local extension, and the flexibility of namespaces was considered essential to the ddi’s continuing evolution. icpsr and harvard-mit data center had been working on a schema version of the ddi that incorporated all of the documentation found in the tag library as well as the dtd comments. it was pointed out that the alliance could not abandon the dtd since a lot of markup had been done that was compliant with versions 1 and 2. the alliance discussed the need to proceed on parallel tracks, moving the dtd along from version 2.0 to subsequent iterations in that development line at the same time that a modular version 3 was developed. the group also reviewed the statistical data and metadata exchange (sdmx), a project to develop an interchange format for time series data and metadata. the sdmx initiative was viewed as a natural partnership, and it would come to be seen as complementary to ddi (gregory and heus, 2007). aligning with the metadater and madiera projects in europe was another topic raised. iassist quarterly 2013 47 iassist quarterly to carry forward this ambitious program of work, new working groups were formed: • structural reform working group (later called the technical implementation committee [tic] and now the technical committee [tc]), which would take on the task of “schematizing” ddi • substantive content working group, broken out into: o group 1: aggregate data, geography & time o group 2: comparative data/families of datasets o group 3: complex files o group 4: instrument documentation • usability and outreach working group this ambitious agenda set the stage for ddi 3, which was five years in the making. at its meeting in 2004 in madison, wisconsin, the expert committee discussed the coverage and scope for ddi 3 and introduced the concept of a lifecycle model. this model (see figure 1) was innovative for its time; subsequently the notion of the data lifecycle became an integral part of the discourse around research data management. the ddi lifecycle approach would come to influence the generic statistical business process model (gsbpm) used by national statistical institutes as a framework for data production (vale, 2010). not everyone was on board with a move to this lifecycle model. a considerable investment had been made in ddi 2.x and people were understandably reluctant to support a new specification that would in effect shift attention and resources away from this version. however, in 2005, at the meeting in edinburgh, the alliance ratified the lifecycle model and ddi 3 began to take shape. a public review of ddi 3 took place in 2007 and the specification was published in 2008 as xml schemas. it was a radical departure from ddi 2.x in many ways. first, it was designed to be used by developers with machine-actionability in mind. while ddi 2.x could be understood and implemented by data librarians, ddi 3 was more complex, often requiring a higher level of technical expertise. the specification itself was designed to be modular and to document and manage different stages of the data lifecycle. it was predicated on the principle of reusing metadata to eliminate costly redundancies and support explicit comparison (vardigan, heus, and thomas, 2008). as green and humphrey note, “enter once and use many times” is a powerful paradigm, which ddi 3 exploited through referencing. as an example, response categories can be defined once and then used multiple times by both questions and variables. ddi 3 also aligned with several other metadata standards including iso 11179, sdmx, geographic and spatial standards, dublin core, and others the community expands as the ddi community of practice began to integrate ddi 3 into its work after 2008, interest from new audiences, including national statistical institutes (nsis), other data producers, and developers and implementers, became evident and new ddi projects sprang up, many of which were discussed in iassist presentations, workshops, and posters. uptake of ddi 2.x continued as well, resulting in ddi spreading across the globe. as a result of the world bank-supported international household survey network (ihsn) program and its incorporation of ddi into documentation tools, ddi came to be used in over 70 countries, many in the developing world (see figure 2). bringing users together ddi users were eager to meet in a forum where ideas, innovations, and knowledge of ddi could be shared. led by joachim wackerow of gesis-leibniz institute for the social sciences and nikos askitas of the institute for the study of labor (iza) in germany, the first european ddi user meeting (eddi), cosponsored by gesis and the iza, took place in bonn, germany, in december 2009. subsequent meetings took place in utrecht, netherlands; gothenburg, sweden; bergen, norway; and paris, france, with the 2014 eddi slated to take place in london. in 2013 the concept of a ddi user meeting spread to north america, with the first naddi conference taking place at the university of kansas, organized by larry hoyle with funding from the alfred p. sloan foundation. naddi 2014 will be held in vancouver. teaching about ddi ddi training, which had been taking place around iassist and in other venues since 2001, was expanded in response to ddi 3, with joachim wackerow of gesis organizing annual training and workshops at schloss dagstuhl, leibniz center for informatics, an it retreat center in wadern, germany, during the last quarter of each year. the first such training took place in 2007. typically, training in ddi 3 is held for a week, followed by a workshop on a dedicated topic. developing ddi tools tools are key to a metadata standard’s success: if markup cannot be produced efficiently, a standard may not find an audience. ddi owes much of its success to the parallel development of the ddi figure 1: ddi lifecycle model figure 2: organizations using ddi 48 iassist quarterly 2013 iassist quarterly markup and data analysis tool nesstar, created by the norwegian social science data service (nsd) for use with ddi codebook (nsd, 1999). nesstar publisher formed the basis for a toolkit designed to assist data producers in developing countries in documenting and disseminating data (see the ihsn discussion above). dataverse network, developed by harvard-mit, also adopted ddi, both as a standard for the study-level metadata entered at deposit and as a foundation for variable-level analysis. support for ddi 2 was also incorporated into the survey documentation and analysis (sda) online analysis system created at the university of california, berkeley. other innovative tools were developed with a focus on ddi lifecycle. for example, the michigan questionnaire documentation system (mqds) exports ddi lifecycle from blaise computer-assisted interviewing software. stattransfer, a commercial product that transfers data among software packages, now also exports ddi 3. the danish data archive developed a ddi editor and colectica developed a suite of tools that includes software to produce, view, and edit ddi lifecycle metadata with an interface to cai tools. colectica for excel was also added, permitting researchers working in excel to create study descriptions and document their data at the variable level. database tools for ddi like questasy, developed at centerdata in the netherlands, added new options for ddi users. the recently released sledgehammer tools suite developed by metadata technology facilitates the transformation of data across formats and enables the extraction and generation of ddi metadata. developers of many of these ddi-enabled tools have begun to meet periodically during the year at conferences to share ideas and to keep each other informed as new tools are created. stepping back: an evaluation takes place the flurry of activity after ddi lifecycle was released and the increasing diversity and expectations of new audiences led the ddi alliance to initiate an open and independent review of the ddi initiative to inform its evolution going forward. there was a sense that the alliance had matured since its inception in 2003 and that the organization needed to restructure to align with its accomplishments in order to be equipped to address new challenges. thus, in 2010, at the request of the ddi alliance members, the ddi steering committee initiated a thorough and independent review of ddi governance and ip issues. the steering committee contracted with breckenhill inc. to provide a review under the following terms of reference: • clarify the intellectual property rights to the ddi specification and how the alliance can best protect its ip • consider alternatives to the current alliance governance structure • review the structure of host institutions and associations described in the bylaws with a view toward opening up the alliance to others to participate in governance • review the bylaws and rewrite to be more specific on the above points • provide guidance on having a constitution that does not change and bylaws that are easier to change, separating the mechanism for revising the specification from the bylaws • review the membership agreement and suggest content • suggest content of a contributor agreement for those contributing products to the alliance • review the current conflict of interest form used by the alliance and provide guidance on how the alliance should approach this broad area after interviewing a large group of ddi stakeholders and consulting widely on legal issues, breckenhill provided a report to the steering committee in may 2011 detailing the findings related to the above questions (breckenhill inc., 2011). in response to the review, the alliance drafted a new charter and bylaws, which went into effect in july 2013. these new bylaws outlined an organization that is broadly representative of the membership and structured to support the effective development of the ddi specifications. there is an executive board elected by the member representatives, a scientific board that oversees the substantive development of the ddi specifications, and a technical committee that creates and stewards the specifications and ensures their usability. the revised bylaws also allow for the ddi alliance to be instantiated within the university of michigan as an organizational host. this arrangement permits the u-m to protect the intellectual property of the alliance and provides a home for the ddi alliance secretariat through the inter-university consortium for political and social research (icpsr). ddi also rebranded its specifications after the review to give an indication of their scope: ddi 1.x and 2.x became ddi codebook, while ddi 3.x became ddi lifecycle. ddi moving forward: what lies ahead attracting new audiences with the publication of ddi lifecycle, there was a surge of interest in ddi by many of the national statistical institutes and organizations around the world, and the ddi alliance now counts the u.s. bureau of labor statistics, statistics new zealand, the australian bureau of statistics, the french national institute of statistics and economic studies, the food and agriculture organization of the united nations, and eurostat as ddi members. the ddi alliance began working with nsis in some notable ways with important synergies emerging. the sdmx-ddi dialogue project helped to surface the similarities and differences between the two standards (many nsis mandate the use of sdmx), and the ddi alliance formally endorsed a collaboration with the sdmx community to enable the two standards to work together. the alliance has also supported development of the generic statistical information model (gsim), the first internationally endorsed reference framework for statistical information that nsis are using to inform the modernization of the production of official statistics. the alliance has made an offer of support to work together on an implementation model for gsim. integrating with the semantic web work is under way on two rdf (resource description framework) vocabularies, the ddi-rdf discovery vocabulary for publishing metadata about datasets into the web of linked data, and xkos, an rdf vocabulary for describing statistical classifications, which is an extension of the popular skos (simple knowledge organization iassist quarterly 2013 49 iassist quarterly system) vocabulary. the public review of both vocabularies is planned for 2014. developing the next-generation ddi while the existing ddi codebook and lifecycle specifications continue to be fine-tuned (ddi codebook version 2.5.1 [in schema form] and ddi lifecycle 3.2 were published in early 2014), the alliance has begun another ambitious project – to create a ddi specification based on an information model. the alliance supports this move to a model-based specification as it will provide greater flexibility: the model can be expressed in a variety of technical formats including xml schema, rdf/owl ontology, relational database schema, and other languages. also, having a model will make it easier to understand the specification, to interact with other disciplines and other standards, to develop and maintain it in a consistent and structured way, and to enable software development that is less dependent on specific ddi versions. interestingly, creating a data model was a component of the original agenda for development of ddi 3, so the initiative has come full circle. the alliance has other goals for this new model-based ddi: this is an opportunity to respond to community expectations by creating a new version of the specification that can transcend traditional disciplinary barriers to document data about humans and their impact more broadly. as an example, while data collection instruments in the social sciences have traditionally been surveys, we can also view blood pressure gauges and magnetic resonance imaging (mri) scans as new types of instruments that capture and export data. there is also a growing emphasis on documenting data from administrative registers and various internet sources. in addition, the next-generation ddi will ultimately add coverage in several new areas: • abstraction of data capture/collection/source with “plug-ins” to handle different types of data • new content on sampling, survey implementation, weighting, and paradata • new content pertaining to qualitative data • framework for data and metadata quality • framework for access to data and metadata • process (work flow) description across the data life cycle, including support for automation and replication • ntegration with existing standards like gsbpm/gsim, sdmx, cdisc, triple-s • disclosure review and remediation • data management planning work on the model began in october 2013 when a group convened at schloss dagstuhl to focus on gathering requirements for and modeling this next-generation ddi. as part of a paper summarizing the requirements (ddi working paper series no. 4), the group articulated a set of design principles for the information model that reflect what the alliance has learned over the years about effective standards and their development: 1. simplicity – the model is as simple as possible and easily understandable by different stakeholders. 2. user driven – user perspectives inform the model to ensure that it meets the needs of the international ddi user community. 3. terminology – the model uses clear terminology and when possible, uses existing terms and definitions. 4. iterative development – the model is developed iteratively, bringing in a range of views from the user community. 5. documentation – the model includes and is supplemented by robust and accessible documentation. 6. lifecycle orientation – the model supports the full research data lifecycle and the statistical production process, facilitating replication and the scientific method. 7. reuse and exchange – the model supports the reuse, exchange, and sharing of data and metadata within and among institutions. 8. modularity – the model is modular and these modules can be used independently. 9. stability – the model is stable and new versions are developed in a controlled manner. 10. extensibility – the model has a common core and is extensible. 11. tool independence – the model is not dependent on any specific it setting or tool. 12. innovation – the model supports both current and new ways of documenting, producing, and using data and leverages modern technologies. 13. actionable metadata – the model provides actionable metadata that can be used to drive production and data collection processes. conclusion from its modest start in quebec city in 1995 with 23 individuals around the table, the ddi initiative has accomplished some important objectives, producing two development lines to document social science research data. the work continues, with a new, more ambitious goal: to spread the next-generation ddi across the social and behavioral sciences and into new communities to ensure the effective documentation of research data and its future use. ddi will mark its 20-year anniversary in 2015. with almost two decades of experience, the ddi community has learned a lot about metadata standards development, and the lessons learned can inform what lies ahead. the journey is sure to be interesting and we welcome fellow metadata travelers, both within iassist and beyond. references “about the [ddi] specification.” ddi alliance web site (via the internet archive wayback machine): (accessed february 12, 2014) about the madiera project: (accessed february 12, 2014) about the metadater project: (accessed february 12, 2014) breckenhill, inc. (2011) ddi alliance external review: summary and recommendations july 8, 2011. ddi alliance original charter (via the internet archive wayback machine): (accessed february 12, 2014) 50 iassist quarterly 2013 iassist quarterly ddi tools catalog: (accessed february 12, 2014) green, ann, and humphrey, chuck (2013) “building the ddi.” iassist quarterly. vol 37. gregory, arofan and pascal heus.” (2007) ddi and sdmx: complementary, not competing, standards.” open data foundation paper, inter-university consortium for political and social research (icpsr). (2001) addendum to nsf final report “electronic preservation of data documentation: complementary sgml and image capture,” sbr-9617813: results of the evaluation of the data documentation initiative (ddi), april 24, 2001. minutes of first meeting of ddi alliance expert committee, october 12-13, 2003, ann arbor, michigan (accessed february 12, 2014) norwegian social science data services (nsd). (1999) “providing global access to distributed data through metadata standardisation – the parallel stories of nesstar and the ddi” (working paper #10). un/ece work session on statistical metadata (geneva, switzerland, 22-24 september 1999). participants in 2012 dagstuhl seminar on ddi moving forward. (2012) “developing a model-driven ddi specification.” ddi working paper series, paper no. 4. vale, steven, united nations economic commission for europe. (2010) “exploring the relationship between ddi, sdmx and the generic statistical business process model.” vardigan, mary, pascal heus, and wendy thomas. (2008) “data documentation initiative: toward a standard for the social sciences.” international journal of digital curation. vol. 3, no. 1, pp. 107-113 notes 1, mary vardigan is an assistant director at the inter-university consortium for political and social research (icpsr) and director of the ddi alliance. she can be reached at vardigan@umich.edu. 2. beta-esters of the first ddi specification included the following institutions: • centre for comparative european survey data (ccesd) – contact: richard topf • danish data archive – contact: nanna floor clausen • the (uk) data archive – contact: ken miller • harvard-mit data center – contacts: michael mcdonald, micah altman • niwi-steinmetz archive – contact: repke de vries • norwegian social science data services (nsd) – contact: jostein ryssevik • university of california, berkeley, survey research center – contact: juteh theresa cheng, jeff royal • university of giessen – contact: karsten d. wolf • university of ljubljana, social science data archive – contact: janez stebe • university of michigan, harlan hatcher library – contacts: bonnie dede, joann dionne, lynn marko, patricia dragon • university of minnesota, machine readable data center – contact: wendy treadwell • university of warsaw, institute for social studies – contacts: pawel morawski and jacek szamrej • university of wisconsin-madison, data and program library service – contact: cindy severt 3. members of the ddi committee when it met for the last time in february 2003 included: bjorn henrichsen, chair, norwegian social science data services; micah altman, harvard university; atle alvheim, norwegian social science data services; grant blank, american university; ernie boyko, statistics canada; bill bradley, health canada; cavan capps, bureau of the census; bill connett, university of michigan; cathryn dippo, bureau of labor statistics; pat doyle, bureau of the census; dan gillman, bureau of labor statistics; peter granda, icpsr; ann green, yale university; peter joftis, icpsr; ken miller, esrc data archive; tom piazza, university of california, berkeley; karsten boye rasmussen, university of southern denmark; richard rockwell, the roper center; jostein ryssevik, norwegian social science data services; merrill shanks, university of california, berkeley; peter solenberger, university of michigan; wendy thomas, university of minnesota; rolf uher, zentralarchiv für empirische sozialforschung; mary vardigan, icpsr appendix: past chairs, vice chairs, and dtd and schema authors ddi committee chairs • merrill shanks, university of california, berkeley: 1995-2002 • bjorn henrichsen, norwegian social science data service (nsd): 2002-2003 ddi alliance expert committee chairs and vice chairs • tom piazza, university of california, berkeley: 2003-2005 • hans jorgen marker, danish data archive (dda), chair, and ron nakao, stanford university, vice chair: 2005-2010 • chuck humphrey, university of alberta, chair, and mari kleemola, finnish social science data service (fsd), vice chair: 2010-2013 ddi executive board and membership chair and vice chair • gillian nicoll, australian bureau of statistics, chair, and ron nakao, stanford university, vice chair: 2013dtd and schema authors (in order of contributions) • david barber, university of michigan • john brandt, university of michigan • ann green, yale university • paul schaffner, university of michigan • nancy vlahakis, university of michigan • daniel pitti, university of california, berkeley • jan nielsen, danish data archive • jerome mcdonough, university of california, berkeley • perry roland, university of virginia • sanda ionescu, icpsr • mark diggory, harvard-mit data center • wendy thomas, university of minnesota • arofan gregory, metadata technology 6 iassist quarterly fall 2009 by kristi thompson1 torturing nurses with data: building a successful quantitative research module introduction/abstract this paper discusses two iterations of an effort to create a quantitative research module for a master's in nursing research methods course at the university of windsor. the first version involved a single three-hour class incorporating both lecture and hands-on practice, followed by an assignment to independently locate and analyze a dataset. the second version was both more extensive and more structured, with a three-hour lecture, an assigned reading, a three-hour practice session, and an analysis assignment using a preselected data set. this paper compares feedback from the two iterations of the module and explains what worked and what did not work. background in-depth support for data began at the university of windsor in 2006 with the opening of the academic data centre. the data centre was conceived of as a service operating within the library that would offer multifaceted support for quantitative research. it includes my position of data librarian as well as that of a data centre manager who runs a drop-in consulting service. as this was a new service without an established customer base we were given an open mandate to promote ourselves, find customers and serve them in any way that seemed appropriate. one of the main goals of this new service was to advance data and quantitative literacy on campus, but exactly how we were to accomplish this goal was unclear. getting a toehold in the classroom – any classroom – seemed like a good way to start, so we combed through the course catalog looking for classes that seemed like they might incorporate use of data, and then sent out custom emails to selected professors offering our assistance. somewhat to our surprise and relief, several of them took us up on our offer. and so we found ourselves doing various things – one professor needed help putting a teaching dataset together, others asked us to do presentations on available sources. one asked us to explain to her class the difference between qualitative and quantitative data, another wanted us to talk about survey construction. and one particularly adventurous professor of nursing invited us to design a unit to incorporate quantitative research into her research methods course. the challenge the data centre manager and i spent some time working with the professor to determine what our goals were to be for this unit. the students taking the masters of nursing program at windsor are mostly practicing nurses of varying ages who want to upgrade their credentials, and the majority of them do not intend to go on to further research. the program has a required statistics course which is separate from the methods course that is fairly math-oriented. many of the students had not yet taken the statistics course, and for these this unit was to be their first exposure to data analysis. there were about 15 students in the class. many students have difficulty grasping or retaining the details of statistical theory, but still need to grapple with the research results produced by the application of that theory. we determined that what we wanted to do for these students was to get them to think about quantitative research from a practical, real-world perspective. we thought their main need was to be able to understand and evaluate the quantitative research that they would come across in practice. "not only does practical knowledge about survey methods and secondary analysis teach students how research is actually conducted, it informs critical assessment of arguments based on the interpretation of survey data." 2 we also hoped to influence those who were going on to further research to consider more quantitative projects. first version we were given a single 3-hour block of class time to work with. this limitation forced us to carefully consider exactly what elements a quantitative unit needed to include. we ended up designing a program with three components; a lecture, a lab and an assignment. the focus of both the lecture and the lab was to prepare the students for the assignment. this assignment was in outline very straightforward: we told the nursing students to each come up with a research question that they could investigate by doing some sort of data analysis. they had to obtain the data, analyse it, and then write a research paper, complete with literature search and discussion. given the limited preparation the students would have, we set aside large blocks of time for walk-in help in the academic data iassist quarterly fall 2009 7 centre when both the centre director and i would be available to answer questions. the lecture our class period started with a lecture where i talked about some basic concepts of data and quantitative research – how data is collected, what types of data are collected, differences between population data, survey data and experimental data, and some of the types of research questions each type of data can be used to answer. one lesson i had already learned when given limited time to talk to a class of students with no real data background is to not spend the time talking about the interfaces to different archives. i’ve moved away from mentioning the details of specific sources as much as possible – i’ve found it usually isn’t retained, and whatever is retained, they could get as well from a class handout or web page. another of my main goals for a data information literacy session did not crystallize until after that first unit we did with the nurses; to cut down on the number of students who come to me hoping to find the impossible dataset. several of the students came in looking for datasets that i knew without needing to conduct a search were highly unlikely to exist: datasets that violated confidentiality rules, surveys that could be used to compare small geographic areas, surveys of very specific populations. for example, a number were hoping to find surveys of nurses in particular cities or working in particular specializations, while another was looking for information that could probably only be obtained by linking individuals and their hospitalization records. to cut down on these types of questions, and on the number of disappointed students who need to be told repeatedly to find another topic, i’ve found that the approach that works best is to spend some time discussing the realities of data collection, what sort of people and groups collect it, and why. the idea is to give researchers a conceptual framework for thinking about data sources – what is collected, what is released. researchers who have some understanding of the data collection and dissemination process have more realistic expectations. however, as i had not yet fully developed this line of thinking before that first unit, i instead found myself making these explanations to many of the students during the walk-in help sessions. the lab the lab component was handled by the data centre manager, dan edelstein, with me on hand to provide extra help to individuals who needed it. dan demonstrated how to do a basic set of analyses using spss, answering simple questions using a test dataset, and then walked them through doing the same analyses themselves on a different dataset. he stuck with showing them how to do descriptive statistics and frequencies, cross-tabulations, t-tests, anovas and linear regression. the students who’d already taken the stats course actually didn’t have much of an advantage here, as they seemed to have learned more about how to calculate statistics than how to read them. when i first presented this paper at iassist 2008, some of the session attendees questioned the idea that we could teach a group of students who lacked previous experience to “use spss” in the space of a single lab session. this is a valid question; spss is a package that can appear almost bewilderingly complex, especially to someone who is without much statistical background and is therefore unfamiliar with many of the terms used in analysis. the answer is quite simple: we did not teach them to use spss. we showed them how to do a very limited set of things using spss, and gave them a handout that laid out the steps for each procedure so they could be followed almost mechanically. our focus was on having them produce output which could then be interpreted. an equally cogent question would be how we expected to teach them enough statistics to interpret that output in the course of that same lab session. here the answer is a little more complex. what we did not do is teach them any statistical theory, nor did we really expect them to understand any. so given that we did not teach them either spss or statistics, what exactly were we teaching them, and what was its value? understanding how statistics are computed is not the same as being able to use statistical methods for research, and it was the application of statistics to research that we wanted to focus on. "the goal of quantitative research is to answer research questions or test hypotheses." 3 in this case that meant we told them to simply ignore most of the spss output and zero in on the numbers that would allow them to make a decision – to accept or reject a hypothesis. we wanted to abstract away from the details of the different statistical procedures and how they worked, and instead focus on what the procedures were doing, and what they all had in common: numbers that showed an effect and told them whether the effect was likely to be due to chance. this approach is discussed in an earlier article i co-wrote with dan edelstein, where we said "our aim is not to teach statistics in itself, but to provide users with the practical knowledge needed to carry out their research... teaching statistical theory is the professors’ role." 4 the assignment the assignment was the test of our approach. the professor had merely stipulated that the students had to write a research paper on a topic relevant to nursing. although elements of this program were based on previous work we had done with classes and individuals, dan and i had never attempted anything quite as ambitious as a threehour crash-course in quantitative analysis for students with little or no background. this left us uncertain as to exactly what level of work to expect on the assignment, so 8 iassist quarterly fall 2009 as a result we left it as general as possible. we hoped that telling the students to choose their own research question, leaving the choice of data and method of analysis open, and offering extensive one-on-one walk-in support would allow them to find and work at a level within their abilities. our expectations were not particularly high – we just wanted them to write papers using some form of data analysis to coherently and appropriately support an argument, demonstrating that they had acquired some understanding of the role of statistical methodology in the research process. the students took our injunction to find a research question that interested them and ran with it. many of them drew on their experience and training to come up with projects that were personally or professionally relevant to them one community practitioner looked at characteristics of groups that don’t access preventative care, while another investigated links between local service provision and hospital readmission rates. a military nurse who expected to work on medical programs in an international setting looked at international data on health interventions and mortality. many projects ended up being surprisingly sophisticated, and we were kept very busy during the month allotted to their assignment supporting our data neophytes as they used the simple analysis techniques they had been shown to conduct complex and interesting research. in short, our low expectations were not merely met or exceeded, but utterly blown out of the water. the professor had a similarly enthusiastic response, emailing us: "i just finished marking their papers for the quantitative assignment and they were superb. i have not had such an amazing outcome in papers such as these…" 5 the second version having judged the first iteration of our quantitative module a success, dr. snowdon invited us to implement an expanded version to her class the following year. instead of the single three-hour block of class time we’d been given before, she offered us up to three three-hour class sessions so that we would not be under such marked time constraints. this time, due to some quirk of scheduling, the class was more than twice as large as before, and even fewer of the students had taken statistics or used statistical software. we realized that the size of this class was going to be a real issue, as we’d had difficulty finding time to give the 15 students we’d had the previous time all the individual support they needed. we knew that our first version had worked well, but we were not exactly sure why it had been as successful as it had been. we had not thought to include a formal feedback mechanism, but we had spoken with many students, and the main request we received in this anecdotal feedback was for more class time on spss, as many of them found the analysis the most stressful part. as we had more time iassist quarterly fall 2009 9 to work with, we decided to cover much the same material we had previously, but more slowly and with more time for questions and practice. so we split the lecture and covered the introductory and data finding material in one class, and devoted the second to practice with statistics and spss. the third class we reserved for questions and one-on-one help. the largest change we made was to the assignment. we thought that the free-form assignment we had used the previous time, allowing the students to "find and work at a level within their abilities," had given us an idea of what level of work to expect. so given the larger class size, we decided to this time give the students a selection of prepared data sets and accompanying research questions to choose from. independently locating data was made an optional element of another assignment later in the term. we also decided that this time we would have the class fill out feedback forms so we could better judge what did and did not work. the numeric feedback we received on these forms was mixed. the students mostly reported that they found the content relevant and useful, but were divided on whether the unit would encourage or discourage them from conducting further quantitative research. many of them did feel that the unit had a positive effect on their understanding of quantitative research articles. however, dr snowdon this time did not rave about the quality of the assignments, and the students working on the assignments seemed less enthusiastic and engaged than the previous group had been. the majority rated the assignment "too difficult", even though we had designed it to be at a lower level of complexity than the work that many of the students had voluntarily taken on the previous term. the comments on the feedback forms helped clarify these results, and helped us to discern that the changes we had made to the program had had some unanticipated effects. particularly telling were comments like these: 6 • “(the quantitative unit) was statistics, not quantitative research” • “it was more like a stats class without having the theory” • “it seemed like a stats project instead of a research project” we also got a number of requests, both in the formal evaluation and informally, for even more time spent on the spss and statistics components. in other words, spending three hours instead of one on the spss and statistics material seemed to paradoxically increase the students' anxiety even more, causing them to ask us to spend still more time teaching it. discussion the two versions of this unit that we carried out formed an unintended, informal experiment on the use of a free-choice assignment in teaching students to conduct quantitative research. and the results of this experiment indicate that changing the assignment was a mistake. the comment “it seemed like a stats project instead of a research project” sums up what went wrong quite well. our original goal had been to get these students to think about quantitative research from a practical, real world perspective – to focus on how statistical findings are used to answer a question rather than on the details of how they are computed. but giving the students pre-selected research questions and prepared datasets left them with little to think about except the details of how the statistics were computed. not allowing them to choose their own research question squelched the interest and enthusiasm that had made the first iteration of the program exciting for both them and us; instead of being curious and excited about the results they were finding, they obsessed over doing the procedures correctly. in short, the unit turned into an exercise in torturing nurses with data, which is probably not the best way to encourage students to go on to further quantitative research! notes 1 contact: kristi thompson, data librarian, leddy library, university of windsor, windsor, ontario, n9b 3p4. phone:1 519 253 3000 x3858 email: kathomps@ uwindsor.ca this paper was presented at the iassist 2008 conference at stanford in the session "numeracy, quantitative reasoning and teaching about data". 2. corti, louise (2004). survey data in teaching project (sdit): enhancing critical thinking and data literacy. iassist quarterly vol. 28 no. 2-3 p 39-54. 3 vogelsang, joan (1999). quantitative research versus quality assurance, quality improvement, total quality management,, and continuous quality improvement. journal of perianesthesia nursing vol.14 iss.2 p 78-81. 4. edelstein, daniel m. and kristi thompson (2004). a reference model for providing statistical consulting services in an academic library setting. iassist quarterly vol. 28 no. 2-3 p 35-38. 5. anne snowdon, e-mail message to the author, december 13, 2006 6. comments taken from feedback forms filled out by students in nursing 63-583, research methods in nursing, december 2007. 19.3 36 iassist quarterly georeferenced population data by hendrik meij and robert chen1 consortium for earth science information network introduction. the primary mission of the socioeconomic data and applications center (sedac), at the consortium for international earth science information network (ciesin), is to develop new policy oriented applications and information products that synthesize earth science and socioeconomic data. the sedac policy applications development effort is the primary means by which the earth observing system data and information system (eosdis) program helps to ensure that the scientific investment embodied in nasa’s mission to planet earth (mtpe) program leads to tangible benefits. sedac’s activities closely link with ongoing activities related to land use and trace gas emissions at other eosdis distributed active archive centers (daacs) such as eros data center (edc) and oak ridge national laboratory (ornl). in addition, sedac’s efforts also play a critical role in the arena of integrated assessment of global climate change. its aim is making outputs of the u.s. global change research program (usgcrp) useful for policy and develop specific tools and mechanisms to enhance the use of integrated assessment models in the policy process. population dynamics and distribution have been consistently identified as key elements in understanding human interactions with the environment and in considering possible responses to environmental change. the national research council (nrc) has identified (1) population dynamics as one of five priority areas of research for the usgcrp. it also points out the key role of georeferenced social data in two other priority areas: 1) improving understanding of land use change, and 2) assessing impacts, vulnerability, and adaptation to global changes. in addition, the report highlights the need to address the full range of energy policy options in relationship to greenhouse gas reduction. the 1992 report of the human dimensions of global environmental change programme (hdp) on population data and global environmental change (2) emphasized the importance of georeferenced population data for global change research and applications and recommended development of data bases at several different levels of aggregation and resolution. georeferenced population data. georeferenced population data provide a critical link between data on the natural environment and data on human behavior and welfare. past and potential policy oriented uses of such data include natural resource management, famine early warning and vulnerability assessment, design and planning of sample surveys, damage assessment and associated disaster response, assessment of the impacts of environmental variation and change (e.g., coastal storms and sea level rise), public health and medical service applications, urban and regional planning, and estimation of pollution emissions and land use change. spatial and temporal data on population and environment can be useful in analyzing behavioral impacts. for example, energy analysts interested in emission reduction policies may want to consider population location and behavioral patterns in relationship to alternative energy sources, air and water resources, pollution control, work and recreational areas, and transportation infrastructure. urban planners may want to understand the effects of decentralized versus centralized population growth on energy use and emissions and assess the effects of alternative population and land use policies on population distribution. national policy makers are likely to be interested in the causes and impacts of internal and international migration, especially as they relate to environmental degradation and change in specific regions. educators may want to introduce such data sources early in the curriculum to provide a base for interaction with national and local processes. data products. sedac has been developing a number of different georeferenced population datasets: o a set of gridded population datasets for more than 120 countries originally developed by the center for international research (cir) of the u.s. bureau of the census for the department of defense (dod) and recently made available 37fall 1995 through sedac. the data files contain both urban and rural population density data at a resolution of either 20 by 30 minutes or 5 by 7.5 minutes. these files may be accessed via anonymous file transfer protocol (ftp). ftp ftp.ciesin.org user password cd /pub/data/global_population_db o a suite of high resolution, one tenth of one degree, population data sets for the globe. two gridded data sets are provided, one with smoothed population counts data, and one with smoothed population density data, both at the 1/10th of a degree resolution. these data are smoothed by applying a mathematical technique (pycnophylactic interpolation) for preserving areal data while redistributing such data on a sphere (3). these files may be accessed via anonymous ftp. ftp ftp.ciesin.org user password cd /pub/data/global_demog_project o a georeferenced population dataset for the conterminous u.s. at the square kilometer resolution level. the approach taken in the development of this product requires both input from the u.s. bureau of the census 1992 tiger/line and summary tape file 3a databases. the initial attempt will focus on the transformation of census blockgroup total persons counts and total housing units structures to the pixel based format of a square kilometer. several disaggregation methods are applied including majority rule, pycnophylactic interpolation, proportional allocation, and geostatistical estimations (kriging). prototype efforts may be found at: ftp ftp.ciesin.org user password cd /pub/census/usa/grid please email ciesin.info@ciesin.org for up to date information. o sedac has identified available sub national administrative boundary files for most countries of the world and is developing an integrated product using a public domain version of digital chart of the world (dcw) at a scale of 1:1,000,000. such boundary files are critical in making linkages between socioeconomic data (e.g., on population, land use, and energy production and consumption) and natural science data (e.g., on land cover, vegetation change, and pollutant levels). access products. a series of activities are underway to make data from the u.s. bureau of the census, and possibly other national census takers, more widely accessible and usable. this includes development of: o an interactive tabulation generator publicly accessible over the internet. this program enables rapid access to very large census datasets for user defined cross tabulations generated from microdata samples. currently, the u.s. 1% public use microdata samples (pums) files for 1980 and 1990 are available. results may be save and retrieved via kermit, ftp or email. please email ciesin.info@ciesin.org for up to date information. o an interactive extraction generator publicly accessible over the internet. this program enables the user to define, interactively, custom extractions to be performed on very large census datasets such as the u.s. 5% public use microdata samples (pums) files for 1980 and 1990. a custom extraction file is generated with a custom codebook and data dictionary. retrieval mechanism include only ftp. please email ciesin.info@ciesin.org for up to date information. o creation of an “ archive of census related products” involving the generation of usable products from pre tabulated datasets, such as the u.s. summary tape file 3a. boundary files (from tiger/line 1992). demographic data and boundary files are accessible via anonymous ftp. the demographic data records are uniquely linked to the appropriate area entities in the boundary files. the current archive contains 11,300 retrievable files spanning 4,000 mb. ftp ftp.ciesin.org user password cd /pub/census informational ‘readme’ files are echoed to the screen to guide the user in navigating this archive. www browsers, such as mosaic and netscape clients, can enter the archive via: http://www.ciesin.org:/datasets/us demog/us demog home.html look for the “/pub/census”’ hypertext link. o the ability for visual display, mapping, of census data by coverages, for browsing/ inspection purposes before retrievals. http://www.ciesin.org:2222/map.html o the generation of “dataset guides” for informational purposes. http://www.ciesin.org/datasets/us demog/ usdemoghome.html summary. the development of these products, providing rapid access to large demogrpahic databases, combined with visual presentation of the materials, might enable more focused response measures to catastrophic events and others. it is hoped that in the short term, users will express and identify databases of interest for merging into the prototype systems under developement. litertare cited. 1. paper presented at iassist95 may 1995 quebec city, quebec, canada. consortium for earth science information network 2250 pierce road university center, mi 48710 (hmeij@ciesin.org, bob.chen@ciesin.org) 2. national research council (nrc). 1994. sciences priorities for the human dimensions of global change. national academy of sciences reports. 3. human dimensions of global environmental climate change programme (hdp). 1992. population data and global environmental change. 4. tobler, w. 1979. smooth pycnophylactic interpolation for geographical regions. journal american statistical association, 74 (367): 519 536. iassist quarterly iassist quarterly 2013 6 abstract data about individuals and organizations are routinely collected across the member states of europe, through surveys, administrative and business transactions. yet access to these microdata for research purposes, particularly across national borders, is often restricted for confidentiality or legal reasons. despite the benefits that accrue to society from allowing comparative research to be undertaken using cross-national data sources, such as policy evaluation, researchers face significant barriers in making comparative analyses of data collected in more than one member state. legal restrictions on the dissemination and transfer of data, and the consequential cost of visiting the data within its country of origin, prohibit access. we present a new initiative to build a european remote access network, as part of the data without boundaries programme, which arguably represents one of the greatest efforts of recent years to allow researchers from across the european union to access data collected in more than one country. keywords european research infrastructure, comparative research, transnational research, data access, data security, remote desktop . introduction data about individuals and organizations are collected throughout the member states of europe, such as national statistics institutes and research institutes, for a variety of purposes. these include the production of aggregate population and economic statistics, such as unemployment measures, but also for a number of administrative purposes too. as an example, to calculate state-provided benefits such as retirement or unemployment payments. yet considerable demand also exists from the research community, to access such data for research purposes, for which microdata are required. some data sources are readily available to researchers, often via internet download. these data are heavily anonymised, by means of perturbation, e.g. removing variables, top-coding and aggregation, to protect the confidentiality of subjects within the data, for example, individuals and organizations. access to detailed, confidential non-perturbed data have previously been restricted, due to the potential risk of identifying the subjects. yet at the same time, accessing detailed data is becoming more and more important for modern research methods. in recent years, not only has access to confidential data improved within countries, but exciting developments are afoot in terms of international data access. the data without boundaries (dwb) project (dwbproject.org), involving data archives, national statistics institutes and universities from around europe, has published recommendations for cross-border data access (schiller 2013). the research data centre2 (fdz) of the german federal employment agency (ba) at the institute of employment research in germany (iab) has established access points in the usa, e.g. at the university of michigan, enabling researchers in the us to access detailed confidential german microdata (bender and heining 2011). without such initiatives, access to cross-country detailed and confidential data will remain constrained, forcing researchers to visit a research data centre (rdc) in person within each country, pertaining to the distributing access to data, not data providing remote access to european microdata by david schiller and richard welpton1 an rdc, often referred to as a ‘secure enclave’, is a centre where researchers can access detailed confidential microdata. iassist quarterly 2014 7 iassist quarterlyiassist quarterly data they require. an rdc, often referred to as a ‘secure enclave’, is a centre where researchers can access detailed confidential microdata, which are never removed from the rdc, but where researchers can undertake statistical analyses of the data. for example, a researcher wishing to compare the effects of trade unions on productivity in uk and germany would need to visit the uk to access the uk data, and then visit germany to access the german data. this paper explains how the proposed european remote access network (euran) will enable researchers to access data from different countries throughout europe, from a single location convenient to them. we begin this paper by explaining the concepts of remote desktop access and its alternatives. we then evaluate the demand by researchers for access similar sources collected in different countries, proceeding to explain the concept of the euran and how it might work. what is remote access? generally, remote access refers to controlling an application package remotely. it neglects aspects of security issues or what can be done from the remote locations. therefore a more detailed description is needed. in the context of this article we talk about the concept of a secure remote desktop to access confidential microdata, and we distinguish between remote desktop and job submission (or remote execution). however, the terms job submission and remote execution are often used synonymously. the table below provides a closer examination of the terms that illustrates their distinctions. remote execution would therefore describe the more automated solution, with reference to the table above, we take remote execution to include both job submission and remote execution. for the purpose of this text, the distinction between remote desktop and job submission (remote execution), both as sub-units of remote access, shall be enough. remote desktop: figure 1 presents a remote desktop solution. to enable work with a remote desktop solution the user needs an access device( this job submission remote execution remote desktop a user submits a program syntax (e.g. via email attachement) to an rdc and a staff member has to execute the program on the data server. in the us and australia, this may also be known as ’remote analysis server’. a user can start an application package (or at least an process) at the data owners' facilities by themselves, without seeing data. a user can log into a regular network computer account (e.g. windows), access familiar statistical software, and access data onscreen. could simply be a plug-in software for their internet browser on their own computer), or a ‘thin-client’, which is a small computing device configured to access a particular network, to send enquiries to a distant server via keyboard, mouse, or a touchscreen. a secure encrypted tunnel or virtual private network, vpn is established through the internet: enquiries (e.g. commands to analyse the data) travel in one direction; screen updates or results travel in the other. behind a firewall, secure servers for authentication and working with data are accessed. since data remain in a secure environment and data modification can only happen under controlled circumstances a secured and controlled environment to undertake research with confidential data is in place. only screen updates, such as pictures from a graphical user interface that is used to work with the data on the secured servers, e.g. the statistical package “r” are routed through the already mentioned vpn tunnel to the screen of the access device. this means that encrypted pictures are routed and scrambled across the internet. if a hacker were to crack the encrypted connection he or she will only see views of scrambled pictures of the graphical user interface that wouldn’t reveal anything intelligible, thus ensuring security. figure 1 secure remote desktop 8 iassist quarterly 2014 iassist quarterly the only potential security gap remaining is that the screen of the user can be ‘overlooked’. in general everybody within close proximity to the computer can look at the screen, no matter who has entered the accreditation credentials when opening the connection. it has to be kept in mind that, due to technical restrictions, no part of the screen can be extracted from the accessing device. the work undertaken on the server via the computer in use is completely sealed off. for example, it is not possible to copy and paste text or files from the server window on to the local computer. the ability to connect local computer devices such as hard drives to the server should be disabled (likewise, printers, usb sticks etc). intruders must take pictures of the screen or scribble down what is shown on the screen, a significant effort for little reward for researchers. this already implies an active action by the user that is, under normal circumstances, not to be expected, since proper training and management of researchers (see desai and ritchie 2009) will foster correct forms of behaviour by the user. additional security measures to control the remote desktop device can also be implemented. these may include technical controls such as incoming ip-address restrictions, validation of hardware certificates, gps information, as well as biometric authentication (e.g. finger-print recognition). safe rooms can also be set up in order to have a controlled environment around the screen (see brandt and schiller 2013). remote access using safe room: figure 2 presents a remote desktop solution using a safe room. if the data are deemed too confidential to allow remote desktop access from a researcher’s institution, a safe room is required to allow access to the data. a safe room is located in the facilities of a data owner, such as the ons virtual microdata laboratory or a trusted partner, who are enabled to run the service on behalf of the data owner, such as the uk data service secure lab. it is controlled and only devices to access the confidential data are located in the room. adding the physical secure room to the technical security measures creates a complete secured environment for the confidential data, since researchers are identified by staff before access to the data is permitted. in practice from a technical figure 2 secure remote desktop using safe room perspective, safe room access works identically to secure remote desktop (a computer securely connects to a server via encrypted vpn). what is different is that a physical wall is built around the equipment, and procedures for signing in researchers exist, among with other similar access protocols. in the diagram above, the only difference is highlighted by the square box around the terminal, which didn’t exist previously in our illustration of secure remote desktop access. job submission: job submission (or remote execution, see figure 3) differs from remote desktop access. the researcher cannot see the data. he or she can only submit program code (syntax) to a data owner, to be run on the data (see figure 3). the data remains in a secured environment, where data owner staff runs the syntax on the data. the user does not immediately see the results; they are sent to the user after confidentiality checks are undertaken by the data owner. depending on the specific realization of the job submission solution, the single steps can be automated and subsequently quickened. due to the setting of job submission solutions neither secured network connection, nor a secured access point is needed. both remote desktop and remote execution solutions have their assets and drawbacks. remote desktop access allows researchers to run calculations in “real time” and to browse the microdata; however, data owners often perceive risks about researchers being ‘overlooked’ by unauthorised individuals. job submission provides the highest level of data security because the microdata will never be seen by the user; at the same time analysing data is very inconvenient for the users, due to the fact that they can only “throw” their syntax into the black box and “guess” the next steps by looking into the provided output files. to make the job submission process efficient, investment in ‘fake’ or synthetic data must be made by the data owners, to allow researchers to write and practice their syntax. nevertheless, a large staff must be employed to receive and run the syntax, and process the requests for statistic outputs. according to the mentioned assets and drawbacks the best option to support the complete lifecycles of a research projects is a combination of both data access solutions; resulting in enabling researchers to choose their preferred access way depending on the current phase of their research project. iassist quarterly 2014 9 iassist quarterly in addition, data protection measures described within this article reflect only a part of the possible ways to protect confidential microdata. in order to find the best mixture between protecting microdata and supporting research a portfolio approach consisting of technical, organisational, legal measures and effective researcher management solutions are required. assessing the demand for data collected throughout europe in this section, we briefly summarise some of the reasons why access to data from more than one country is desirable, and consequently demanded from researchers. where data are collected by official agencies (e.g. government, national statistics institutes) the qualityassurance processes embedded in data collection methods satisfy the demand from the research community for robust data. in addition, data collected by these means usually provide large sample sizes, a particular feature of data collected through transaction and administrative processes such as taxation. but regardless of how the data are collected, whether by official agencies or research institutes/universities, we cannot understate the desire of researchers to compare the different institutions, markets and policy implementations across the different member states of the european union, and how the outcomes of individuals and organisations differ with respect to these different environments. we believe that more comparative research would be undertaken more frequently if detailed microdata from the different member states were more easily available. in the current economic climate for example, comparing the effect of unemployment insurance on the employment outcomes of the uk and german workforces requires comparative labour market data from both countries. in addition, where small-scale survey data are collected by each member state, researchers could capitalise on the larger sample sizes that may be obtained by pooling observations from different countries. data on business organisations that operate across the european union, currently fragmented only by access but not necessarily by sampling frame (eurostat regulations ensure that data about business organizations are consistently collected throughout the member states), could even be combined: a researcher could, for example, potentially compare the productivity of uk and german manufacturing plants that belong to the same company. finally, data collectors/producers can take advantage of a potentially free source of statistical validation and quality assurance: when researchers analyse data collected by the european member states, they will easily spot sources of varying quality (particularly in terms of documentation, metadata, and often the robustness of the data themselves). data collectors/producers can exploit this knowledge to improve collection methods, and enhance the quality of aggregate statistics that they are duty-bound to produce. for these reasons, access to sources of detailed confidential microdata collected throughout the european union is a prize highly sought after. the potential for pan-european research, and the implications to test an array of public policies, is enormous. distributing access to data, not data before the advent of secure network technology (providing secure access as we describe below) data could only be transferred to the researcher. yet one of the major barriers to accessing detailed microdata throughout the eu is the legal imposition that data cannot be transferred among member states. researchers are therefore forced to travel to the destination country where the data reside. however, researchers have many demands on their time. even without departmental and teaching duties, family commitments coupled with the expense of travelling, dissuade all but the most determined and well-funded researcher from accessing data from more than one member state. a similar situation confronts researchers even within their own country – access to the most detailed data may only be accessible by travelling to an onsite rdc; equipped with a lockable secure room where researchers can undertake their research. even within small countries such as the uk, travelling places a burden of constraint on the researcher. figure 3 job submission solution 10 iassist quarterly 2014 iassist quarterly however, the above mentioned recent advances in secure network technology have the potential to ease this problem considerably. the solution is to distribute access to data, instead of data. the main advantage can be seen in the fact that data owners responsible for confidential data are always in control of their data. the data remain physically in the place they want it to be and they can immediately cut the connection to the researcher, if deemed necessary. while this reassures data data owners that access is secure, researchers realise huge benefits from using remote desktop solutions. there is no need to have confidential data stored on private computers of researchers. they can do the same analysis remotely. sophisticated working environments enable user friendly solutions. in addition, data sources needed for modern research become available, where they were previously much harder to reach without remote desktop access and the approach of distributing access to data, not data. other examples of distributing access to data instead of data exist: for example in the big data world. the solutions are similar, even if the objectives are different. data remain in the same place during big data analysis because of network problems that would accrue when moving data from one place to another, primarily due to their size; confidential microdata for social science research are not transferred because of legal issues. this fact opens up the possibilities to exchange new developed solutions for data access, storage and analysis between the two worlds; and bring them closer together by doing so. for the time being developments from single remote desktop solutions to networks of remote desktops are needed. in the uk, a remote desktop access solution is provided to researchers by the uk data archive3 , and this is known as the secure lab. it currently provides access to detailed and confidential economic and social microdata to some six hundred researchers across the uk. and as previously mentioned, the iab in germany provides access to german microdata to american and german researchers via safe rooms located throughout the usa. remote access solutions such as these prove that it is possible, with collaborative efforts of the countries involved, to overcome the legal barriers of moving data, by distributing access to data, not data. principles for access to european research data since we cannot anticipate future changes to the technological and legal environments, we advocate that european data access is founded upon a set of principles (following ritchie 2005), which are designed to satisfy data collectors/producers and the researchers who wish to access the micro-data. providing these principles are met, future technological solutions can be implemented, in whatever form they take, and these may differ to the solutions we propose below. we set out the following principles: 1 access must be distributed. it may not be possible or desirable to physically move data between member states. but providing access is still possible, this should not matter. 2 access should come from a single point. the researchers should be able to access all available data from a single point, rather than accessing multiple sources of data from multiple points which is time consuming and costly. figure 4 euran iassist quarterly 2014 11 iassist quarterly 3 access must be secure. the connection between the researchers accessing the data, and the location of the data, must be secure. 4 access must be compatible. the access infrastructure that is provided to researchers should be compatible with technological systems used by researchers and data providers alike. 5 researchers must be able to work collaboratively we feel strongly that the last principle should be met. as we envision researchers from different member states will work together on the same research project. european remote access network (euran) in addition to the uk data service and iab in germany, there now exist rdcs that provide access to detailed confidential microdata in the netherlands, france, sweden and denmark, to name but a few member states. see also ‘data without boundaries deliverable 4.1: report on the state of the art of current sc in europe’ 4. the european data without boundaries (dwb) project announced work to create a european remote access network (euran). the aim of euran is to allow researchers based in one member state to access detailed confidential data from other member states, without the need to travel to those member states and to use services that support research on a european level. this section provides more information about how euran will work, and broadly describes its key features. many of the member states already have at least one rdc providing access to detailed confidential data. the euran can build on this existing infrastructures and experiences. the architecture of euran is illustrated in figure 4. suppose we have three member states (data providers a, b and c). the data would always remain within the member states, to comply with legal requirements. a researcher, in principle, located anywhere within the european union, would be able to access the data by a number of methods. in the figure above, these are specifically: from a safe room, from their institution, or even from anywhere (for example to look at data documentation). while the connection between the access points and the single point of access (spa) always uses a remote desktop approach, the connection between the spa and the data providers can consist of a remote desktop, remote execution or a job submission solution. how will the ‘user experience’ appear? imagine a researcher based at a uk institution (e.g. university). depending on the extent of confidentiality of the dataset, the researcher would log into their account either from a safe room, their institution computer, or from a laptop at any convenient location. by logging into their account, they will be provided with a simple account, for example a windows server account, and access data from one or more of the three data providers (which represent a different country). the data are stored in each of the three countries, but the account is set up such that the researcher, from the single account, can click onto one or more data storage devices that takes them to the server which resides in the respective country. for example, if the researcher is granted access to uk data (data provider a) and german data (data provide b), then they will be able to click and enter those respective storage devices from within their single account. we now describe the key components of euran, which are illustrated in the figure 4. access points for distribution from a technical perspective, working with data from many sources across the eu remotely is only limited by the possibility of using a device that provides access to a network, usually the internet. however, access nodes, the physical location where a researcher may access data, are often more narrowly defined by legal restrictions and the enforcement of data protection principles according to one or more member state. for example, in the uk, statistical legislation prohibits access to data only for government staff and ‘approved researchers’, a legal entity created which describes some trustworthy person who has been approved to access data collected under the legislation. therefore, one cannot access such detailed data using a laptop in a café where many non-approved individuals are located close by. the access nodes for the euran must be agreed by participating member states. as mentioned earlier using a remote desktop solution with access device located in a safe room offers a completely secured environment to access confidential microdata. euran supports this access solution. while this approach still forces the user to travel to a location where a safe room is available, a safe room network (brandt and schiller 2013) would reduce the need to travel. a more convenient way to work with data is the possibility to access from the home institution of the researcher or even from within their office. this type of access is now provided by a number of operational remote desktop services in europe, e.g. the secure lab at the uk data archive or the casd in paris5 . researchers can access the secure environment after they have gone through a two-factor authentication, e.g. password and finger print recognition. finally the euran also supports access from “anywhere”. this is possible and needed, when accessing nonconfidential services, like data documentation or a wiki shared with other project members. the principle here is that access is now distributed, the data only remain in the country from where they are collected, but the researcher can access data from all member states from their home country, rather than travelling to each country. this is essentially an exercise in minimising the movement of the data, while maximising access to the data, subject to legal, organisational and technical constraints. the euran solves this problem for the given set of constraints until further agreement among member states is such that access can be distributed further. we hope that the future access landscape will be more flexible, ‘disaggregating’ access to the point where researchers can access the data from a location convenient to them. single point of access (spa) and service hub a major design aspect of euran is that researchers should be able to access the detailed european data through a single point. this central access point enables the use of a number of functionalities, e.g. a centralized user authentication system, centralized information platform, workspace for cooperative works, such as projects or the scientific community, storage platform for multinational datasets, and a secure trusted third party environment. whether the principle of single access is achieved via a safe room or a more distributed access method (e.g. from the researcher’s own institution), is not important. however, the single point of access must first authenticate the researcher, and therefore needs a sophisticated rights management system in order to provide the 12 iassist quarterly 2014 iassist quarterly researcher with all functionalities needed and ensure security and compliance with the restrictions of the data owners. the user should experience a familiar and customizable working environment. ‘back office operations’ via a service ‘hub’ manages databases and applications that offer additional services demanded by researchers, such as documentation, metadata production etc. (see burghardt and schiller forthcoming for more details of such an operation). secure connection a third principle of establishing a euran is that the connection to the data must be secure. as noted above, it would be wise to take advantage of evolving technologies. implementing this in practice can be achieved by using encryption techniques and secure virtual private network (vpn) technology, which is currently used by e.g. the uk data archive secure lab and the iab. this provides a secure encrypted connection between the user at the access node and the computer server, which stores the data. such technology is widely used by the banking and military sectors which rely on confidential up-to-the-second data. in addition to meeting the secure connection principle, a compatibility principle must ensue, whereby the connection interface, the point at which a researcher logs into an account is compatible with all suppliers of the data and users of the data: otherwise it will be impossible to join the network with the various data suppliers, and researchers themselves will not be able to use the interface. the single point of access and its connection to the servers will always occur via a secure remote desktop solution. the euran will support a combination of remote desktop access, remote execution and job submission solutions as may be necessary depending on the researcher needs and data owner requirements. these solutions can work together via a single point of access. data storage the extent to which data can be stored outside of the member state in which they are collected depends upon national legislation and the disclosure risk of the data. typically, interpretations of national legislation prevent the distribution of confidential data outside of national boundaries. our final principle therefore is that data do not travel. but this is not necessary. modern technology, as shown above, allows access to data without the need for data to be physically moved. data can be accessed using existing storage facilities that are currently provided by the rdcs of the member states. using the model of a single point of access, the researcher simply authenticates themselves when logging into through their access point, and will securely access storage facilities at the various rdcs of the member states. if security restrictions ask for it, parallel data storage infrastructures can be established within the existing rdcs. this would result in two data storages: one isolated within the rdcs and the other one as part of the secured euran infrastructure. while data must remain in the member states where they were collected, it is up for discussion and agreement as to the circumstances that output files, containing statistical results from analyses are allowed to be moved to and stored in the single point of access. microdata computation centre (micoce) the euran can provide access to data storage systems of different rdcs, where researchers can work with confidential microdata and save their work. one of the services provided by the service hub in the single point of access could be a microdata computation centre (micoce). it would provide storage space, statistical software and above all computational power required to bring data together from different countries for comparative analyses in a secure it environment. as mentioned previously, confidential data cannot be transferred from one member state to another. when trying to promote european research instead of national research, a solution to send enquiries from one single point and run calculations on multiple data sources stored in different physical locations is needed. a workshop6 held by dwb (in late april 2014) demonstrated some of the potential approaches that could act as a solution. however, harmonization of data sources across europe is needed to push transnational research in europe, at least the storage of interim outputs from calculations with multiple data sources should be possible within the single point of access of euran. furthermore, the micoce could also function as a storage and computation centre for new rdcs that do not want to invest or do not have the resources to build the whole infrastructure by themselves. if legally possible these rdcs can use the capabilities of micoce. virtual research environment (vre) we understand that access to data by itself is not a means to an end, and that in order for researchers to undertake scientific enquiries of the data, they require tools in which to do so. these are encompassed within a ‘virtual research environment’, a workspace provided to each researcher who accesses the euran. this workspace can be protected by different security levels, depending on the disclosure risk of the data source involved. the basic requirement is a workspace which includes analytical software and applications to generate, prepare and present results. however, we believe that a euran must be built with collaboration in mind. we anticipate that researchers from different member states will work together on projects. a ‘collaboratory’ must be available, similar to that which will be available in the uk data archive secure lab, which allows researchers working on the same project to securely share and discuss results. iab is also involved in a project developing such an environment. this again is subject to evolving technology. at a basic level, a shared project area, accessible only by a group of researchers working on the same project, should be available. but more advanced solutions, including instant messaging to allow real-time communication between researchers, would aid research productivity and are therefore highly desirable. the availability of tools such as methodbox7 can provide a onestop data support solution for researchers. this tool can enable access to documentation and metadata relating to the data the researchers are using, can allow communication between researchers and data owners on queries, and provide support for one another. such tools foster understanding of the data, and the data producers can view community discussions about their data with a view to improving data collection and preparation methods for the future. in addition, the burden of research support for the data producers and rdcs will surely be minimised if researchers can support each other through online forums, as an example. iassist quarterly 2014 13 iassist quarterly this provides but a flavour of the virtual research environment. as technology develops, future devices for enhancing the working and support environment ought to be provided, indeed it is likely that researchers will drive the demand for collaborative tools.. information platform building a data infrastructure network such as euran must be complemented with dedicated support. by this, we mean user support functions that researchers can avail themselves of. these support functions provide help to users for accreditation, applying to access the data, including support for finding and selecting appropriate data sources and completing relevant application documentation; and support while analysing the data, which may include the production and promotion of documentation and metadata. these support functions constitute the ‘information platform’, and is likely to be the first point-of-contact a researcher will have with the network. this is also an opportunity for the member states to harmonize these support activities, which are currently provided by the individual member states. researchers should receive access to the same information and support, regardless of the member state where they are based. a platform such as this is described in ‘data without boundaries deliverable 5.1: report on the concept for and components of european service centre’ 8. developing euran to demonstrate how the euran could be established, three rdcs delivered by the casd (france), iab (germany) and uk data service, have begun a project to provide access to each others’ detailed, confidential data. this will begin with the installation of thin-client terminals in the safe rooms of each rdc. each thinclient will provide direct remote and secure access, using the vpn technology described previously, to each rdcs’ collection of detailed data. for example, a thin-client installed in the uk data archive safe room, can be used by researchers based in the uk, to access data available from the iab. this will be a pilot project, which can be developed into the full euran that we have described in this paper. the pilot project will provide pragmatic solutions to many of the technical, legal and organisational issues, which will surely need to be addressed as the euran begins to emerge. although the pilot will be established by these three particular member states, feedback on development and progress of this pilot will be shared and discussed with other data without boundaries participants and external experts. the network will therefore evolve with successive iterations, to help improve the concept. thus far, our overview of the development of the euran has focused on information technology. at this point, some digression is required because we recognise that an access network which relies on security, dependence on technology alone is not sufficient. as desai and ritchie (2009) point out, if researchers are treated as a risk by data producers and/or data providers, researchers have little incentive to consider themselves as responsible for the security of the data for which they are accessing. achieving ‘buy-in’ from the researchers accessing the network will therefore be crucial, not just in terms of establishing an effective easy-to-use network, but also for achieving data security. part of the role of the euran will be to foster such engagement by actively managing researchers to achieve the involvement of the community of researchers in protecting the confidentiality of the data. conclusion and outlook this paper has summarised the advantages of a more integrated network of access to detailed and confidential data can bring about. we have presented a technological solution, supported by principles of european microdata access to ensure that collectors of data and researchers who analyse the data are equally satisfied. technological solutions will evolve in the future: but the underlying principles required for secure and collaborative access can be met by an array of solutions. the data without boundaries project consists of many projects, of which the euran development is only one. other hard work, including an examination of legal issues, statistical disclosure control of results, and training of researchers, to name but a few, has also been undertaken by national institutes and data archives throughout europe, and will directly contribute to the future development of the euran. as a result of this project, we anticipate that the landscape for accessing detailed european data will soon be very different, to the advantage of the research community, and to society which benefits from the comparative research that can be undertaken using data collected throughout europe. references bender, s. and heining, j. (2011), „the research-data-centre in research-data-centre approach: a first step towards decentralised international data sharing“, iassist quarterly, fall 2011, p. 10. http://iassistdata.org/iq/research-data-centre-research-data-centreapproach-first-step-towards-decentralised-international brandt, m. and schiller, d. (2013), “safe centre network – need for a safe centre to enrich european research”, presented to the joint unece/eurostat work session on statistical data confidentiality, 2013, available at http://www.unece.org/fileadmin/dam/stats/ documents/ece/ces/ge.46/2013/topic_3_brandt_schiller.pdf burghardt, a. and schiller, d. (forthcoming),”introducing the service hub for remote data access”. ritchie, f. (2005), “access to business microdata in the uk: dealing with the irreducible risks”, presented to the joint unece/eurostat work session on statistical data confidentiality, 2005, available at http://www.unece.org/fileadmin/dam/stats/documents/ece/ces/ ge.46/2005/wp.29.e.pdf desai, t., and ritchie, f. (2009), “effective researcher management”, presented to the joint unece/eurostat work session on statistical data confidentiality, 2009, available at http://www.unece.org/ fileadmin/dam/stats/documents/ece/ces/ge.46/2009/wp.15.e.pdf schiller, d. (2013), “proposal for a european remote access network (euran) – main components”, presented to the joint unece/ eurostat work session on statistical data confidentiality, 2013, available at http://www.unece.org/fileadmin/dam/stats/ documents/ece/ces/ge.46/2013/topic_3_schiller.pdf acknowledgement we acknowledge the work and contributions of the data without boundaries work package 4 participants: atle alvheim (nsd), steve bond (ons), anja burghardt, iris dieterich (iab), leo engberts (cbs), kamel gadouche (casd), maurice brandt, christopher gürke (destatis), and roxane silberman (cnrs). the research leading 14 iassist quarterly 2014 iassist quarterly to these results has received funding from the european union’s seventh framework programme (fp7/2007-2013) under grant agreement n° 262608 (dwb data without boundaries). we also acknowledge previous unpublished work by felix ritchie and richard welpton presented at the 2011 iassist conference entitled “access without borders”. some of the ideas of this unpublished work are presented in this paper. notes 1. david schiller, institute for employment research (iab), nuremberg (germany); email: david.schiller@iab.de richard welpton, uk data archive, university of essex, colchester (uk); email: rwelpton@essex.ac.uk 2. http://fdz.iab.de/ 3. information about the uk data service secure lab is available at http://ukdataservice.ac.uk/use-data/secure-lab.aspx 4. visit http://www.dwbproject.org/about/public_deliverables/d4_1_ current_sc_in_europe_report_full.pdf 5. http://casd.eu/ 6. http://www.dwbproject.org/events/workshop-micoce.html/ 7. https://www.methodbox.org 8. visit http://www.dwbproject.org/export/sites/default/about/public_ deliveraples/d5_1_european_service_centre_report.pdf. iassist quarterly fall & winter 2007 5 i the 2008 iassist conference, “technology of data: collection, communication, access and preservation” included a session entitled “moving research data into and out of institutional repositories” from which several papers emerged. in “interoperability between institutional and data repositories: a pilot project at mit”, katherine mcneill describes a pilot project to enhance study discovery between two repository systems housed in the same institution, dspace and the institute for quantitative social science dataverse network, by enabling the harvesting and replication of metadata and content across the two systems. in a related project across the pond, libby bishop scales this discussion in her description of crossinstitutional collection sharing between the university of leeds and the uk data archive in the timescapes project. bishop asserts that coordination among multiple agents is likely to be challenging under any circumstances. challenges magnify when the trajectories of different life cycles, for research projects and for data sharing, are considered. robin rice echoes these sentiments in her article on the disc-uk datashare project, a collaboration between the universities of edinburgh, oxford and southampton and the london school of economics. rice provides visual evidence in a compelling diagram of the data sharing continuum based on storage, discovery, and preservation conditions of the digital research materials at each level along the scale -from the lowly thumb drive to the officious national archive. we see plainly that as one moves up the continuum, more and more human effort and intervention is required to craft the discovery, access, analytic and preservation environment. in other words, data curators matter. two other papers tackle these challenges by emphasizing the needs of data producers. luis martinez-uribe introduces the university of oxford’s scoping digital repository services for research data management project and the findings of a requirement gathering exercise. while the study results reveal researchers’ needs and workflows. martinez-uribe asserts that the study process itself made an impact on the participants. study participants reflected on and, as a result, fine-tuned how they work with data, why they create these materials in the first place and were able to articulate reasons for managing these resources the way they do. similarly, research data & environmental sciences librarian, gail steinhart, writes about the development of datastar, a data staging repository hosted by cornell university’s albert r. mann library. the project developed as a “managed workspace” where researchers contribute datasets they are still actively using in direct response to questions that have to do with sharing in the active research environment, rather than an archival one. while the authors in this issue describe projects going on in many different places and settings, taken together, these articles address common themes. all address the challenge of scaling data exchange between systems and then between institutions. this raises the perennial question of standards: by what mechanisms will we set them, and how well will we be able to follow them and still accommodate local needs? the importance of aligning repository services with researcher needs is another common thread. data managers must ask, “how will the active researcher benefit from curation efforts”? the answer may be that benefit is more than finding or accessing a particular resource (yep, i have downloaded the whole thing and all the bits are there), but instead being able to examine this resource in many ways (okay, lets run frequencies, now i want to see it on a map, and let’s include some other variables). this is a rich reuse experience, creating a real digital “laboratory.” finally, each contributor notes the expanding role of data manager. in its own way, each project described here moves data managers upstream, pre-publication, into the place where research is actively happening. though all of the articles focus on technological choices and architectures to support research data curation, it is striking to realize that each of these choices emerge from old-fashioned personal, social, and organizational relationships. what we can strive for as data and information managers is to work together as fellow researchers and to be ever curious about how these partnerships and the sharing of information back and forth can be enhanced by thoughtful information and technology design. some call this the digital plumbing, but i like to think of it as e-gilding. gretchen gano, new york university libraries guest editor’s notes ^sist newsletter vol.1, no. 2 planned several projects: (a) development of a list of recommendations addressed to potential researchers on study desiqn as it relates to data management (the "do's and don'ts" list). group members are to send in their suggestions to the dom ag coordinators. a consolidated list will be produced in toronto for publication in the newsletter and elsewhere, (b) collation of information on relevant monographs, technical reports, program writeups, etc., which the ag members would like to share. the intent of the collection is to publish this information in a "what's new in data flanagement" section of the newsletter . 5. prepared an agenda for the toronto meetings; items for that agenda have been referred to above. the dom ag coordinators would like to suggest the following revised mandate for future discussion at international lassist meetings: "this group addresses the problems of data organization and management confronting those archiving or usinq social science data. the ag will investigate and evaluate software and procedures for data and documentation preparation and management; recommend nuidelines for preparation procedures and software development; and, sponsor workshops and seminars for professional training and the exchange of information in these areas." [the original mandate includes "hardware" as an area to be addressed by this ag; see newsletter volume 1, number 1 for the full text of the mandate.] an overview of problems associated with proc ess prod uc ed data / paul miitler for the august 1976 lassist meetings, paul muller preoared a report which provided a broad overview of problems associated with process-produced data. what follows is an edited version of this report. [future issues of the news-letter will contain additional action group reports. the membership is encouraged to begin a dialogue on this and other issues of concern to the data archive community.] "administrative bookkeeping as a social science data base" by,, paul muller institute for applied social research university of cologne 1.0 [...] i will try to give a rather broad overview of problems associated with the "production, acguisition, preservation, processing, distribution, and utilization of machine-readable" process-produced data. this report is necessarily biased by my own viewpoints and experiences; other experiences may well be fundamentally different. but,' the function of this paper is to initiate discussion and later, joint actions. ^sist newsletter vol.1, no. 2 1 . 1 process-produced data 1.1.1 definitions within this action group we should be concerned with "process-produced data" as defined by rokkan, as well as with "official bookkeeping data" (e.g., administrative registers). a common problem with these data is that they are/were not originally collected for scientific purposes and/ or within explicit statistical routines, thus creating the research situation that the data are to a great extent already "given". the generic term, "process-produced data," would encompass all data that are/were not collected for statistical or scientific research (e.g., censuses or surveys), but are instead by-products or traces of the daily routines of private or public organizations or persons. 1.1.2 priorities as a first step, we should concentrate on already-created machine-readable data within public administration. this would imoly structured mass-data. 1 .2 problem areas 1.2.1 inventory of machine-readable administrative data bases the first task is to obtain an overview of the existing data bases within the restricted domain as defined in 1.1.2. in west germany, quantum plans a pilot study for morth-rhine-westfalia , "a continuous inventory of administrative machine-readable data bases." prior documentation projects have identified some 185 machine-readable data bases at the state level (northrhine-westfalia) , as well as some 655 files at the local government level (city of cologne) [...] as by-products of edp use in public administration. these inventories should be [...1 descriptions of the data bases: e.g., coverage, period of correction/update, variables, etc. the experiences with the project in germany show that updating would not occur without the active participation of the administration. the first step towards this end should be en overview of existing edp routines within public administration. this can be rather easily achieved by using the information of the existing coordinating committees/institutions within public administration and by jointly establishing a check list. 1.2.2 documentation existing documentation of machine-readable process-produced data e.g.. national archives and records service, catalog of machine-readable records in the national archives of the united states , (washington, d.c., 1975) or directory of computerized data files & related software , (national technical information service, 1974) are models far less ambitious than the envisaged codebook-1 ike documentation. good examples for this would be: department of health, education, and welfare, 1973 cur rent population survey summary earnings record exact match file code book, part i basic information studies from interagency data linkages , by f. scheuren, d. vaughan, and w. alvey. report no. 5 (1975) and dei^^sist newsletter vol.1, no. 2 partment of health, education, and welfare, 1973 current population survey summary earnings record exact match file codebook, part ii supplemental information, studies from interagency data linkages , by f. scheuren, b. kilss, and c. cobleigh. report no. 6 (1975). 1.2.3 preservation it is time to follow up the initiatives that were made by the ruggles-, kaysenand dunn report in the united states. specifically, we have to define research needs vis-a-vis national, state, and local archiving institutions. in so far as some 95% of data generated within public administrations are physically destroyed, joint actions must be launched to define worth-while material to be preserved. in germany, we have an ongoing discussion concerning an archive law which should take into consideration the specific interests of social science research (criteria for preservation, the archiving of temporal and/or cross section samples of linked files). the german data law will have an effect on the quality of the archived data, in so far as it allows for selective destroying of individual records within a register. it is yet very unclear whether this contamination effect can be avoided in some way. i propose beginning with an overview of existing archiving criteria (kassation, physically destroying of data) within governmental archives and investigating whether these criteria are compatible with the kinds of uses thai; are not case-studies or oriented towards identified persons (e.g., historical). 1.2.4 data laws the data laws will have an impact on accessibility to administrative data for research purposes, not only with reg.ird to access to identifiable individual data (very often necessary in the data collection and management phases). as these data laws do not provide for exceptions, serious research projects will be effectively hampered when these projects are not in the interests of a data providing agency. data laws (or their drafts) are increasingly used for the secretion of public administration data (even in the cases when anonymous information are required/souaht) . . . . 1.2.5 uses already made of process-produced data to ensure response to research needs for process-produced data, we should survey the uses made of administrative registers (i.e., those uses that were beyond the simple drawing of samples). in germany, we will take another look at around 6000 identified research projects within the informationszentrum social science research project. similar endeavors should be made in other countries as far as there exists information on research projects. 1-2.6 quality and characteristics of administrative bookkeeping there has been very little attention paid to problems with the quality of process-produced data (e.g., how are these data created, in what ways are the collection or validation processes biased) or to the characteris19. sist newsletter vol.1, no. 2 tics of official bookkeeping. there is a plethora of "validity studies," which compare sample survey data with official statistics, but as a specific kind of study are not very cumulative. the quality of process-produced data and characteristics of administrative registers are the objects of an ongoing research project at the institute for applied social research in cologne (wolfgang bick and paul j. muller) in which we analyse the laterality of representation of the individual's (client's) social context and the temporal changes in registration by public administration of everyday activities. 1.2.7 record linkage because these data are very often "meager" due to the limited administrative purposes for which they are collected, linked registers or data sets should be of the greatest interest. to the extent that administrative registers are organized within different life sectors, the potential of record linkage, (or family reconstitution as it is called within historical demography), for supplementing existing large scale, but limited registers is promising. a codebook-1 ike documentation of linked files should by envisaged (cf . 1 .2.2) research on record-linkage techniques increasingly concentrates on the problems associated with statistical versus exact links (cf. department of health, education, and welfare, some observations on linkage of survey and administrative record data, studies from interagency data linkages , by j. steinberg. (1973). the 60-odd exact-matching studies done in the last decades showed that without a unique standard indentifier (ssn, pk, person identification number) exact record-linkage would continue to be a very expensive task achieving only an average of 85-90% matches. we need a methodological breakthrough for synthesizing different data bases according to configurations of socio-economic characteristics. 1.2.8 instruments we will repeat the mistakes made in earlier phases of the archive "movement" if we neglect the need for specific computer programs to handle process-produced, often "ragged" data; therefore, we should consider the problems of analyzing strings of data (e.g. within crosstabs or troll, perhaps) hierarchies of relations (e.g. area-block-house-fanily-respondent), or "life histories". it is good to hear that the spss-survey has already brought these issues to the attention of the spss-people. [...] 1.2.9 interdisciplinary communication there is a real opportunity for intensified communication between those researchers working in quantitative history (the analysis of tax registers, birth certificates, marriage licences, etc.) and social scientists who have already worked with or are interested in using process-produced data for their research purposes. this communication, which is already institutionally organized within quantum, should enable us to document, 20 sist newsletter vol.1, no. 2 store, and distribute process-produced data in such a manner that the kinds of uses have not to be invented after completing all these tasks. in particular, the development of a "source criticism" for mass data, analogous to the development of the methodology of survey data, can only be achieved in an interdisciplinary way. such efforts are necessary for the envisaged descriptors for machine-readable process-produced data. similarly, cooperation with those people working on record-linkage problems (e.g., oxford record-linkage study on medical record linkage) or with large-scale process-produced data (e.g. criminal statistics utilizing court-records) should be initiated.... 1 .3 quantities it is not possible to deal with these problems solely within the existing social science archiving movement which has so heavily concentrated on survey data. other institutes must be brouaht into a network of archives and information centers in which coordination and a division of labor must be planned. (examples of these include the "sozialdatenbank" [social-information-system of the departiflent of labour and social affairs, west germany], the proposed "zentrum fur aggregatdaten" [centre for aggregate data of the german national science foundation], and the national archives.") in germany, the prospects for coordination are a little bit better because of the pioneering work being done within the informationand documentation program of the federal government. book n ot i c e s / kathleen m. heim introduction this column is a preliminary step in defining the literature of data archiving. those of us who have tried to assess the state of the art in order to formulate annual reports, write articles, or keep professionally informed have been frustrated by the lack of bibl iograohic control over our area of concern. indexing and abstracting services such as social science citation index , information science abstracts , library literature , social science index and resources in education are unsystematic in their assignation of subject headings to pieces of literature related to data archives. the problem is further confounded by the fact that seminal information concerning the establishment of data archiving has often been distributed informally at conferences or in unpublished papers. when our numbers were small we could depend upon an invisible college network to disseminate important information. however, as our numbers grow and as new archivists enter the field without access to the established network, it becomes mandatory that we define and organize the literature of our profession. 20 iassist quarterly building regional data files in the federal republic of germany by edwin ferger ' zeniralarchiv fur empirische sozialforschung universitial zu kbln the zeniralarchiv fur empirische sozialforschung, universitai zu kdln, has engaged itself in the creation of the german part of the 'joint european time-series data base'. the development of this data base for comparative social science research is a project by the ifdo (international federation of dau organizations for the social sciences). in the following paper, i will describe some of our experiences building regional data files, 1 will point out where we stand now and some problems we face. experiences in the past although the zeniialarchiv's primary tradition is 10 file surveys, there are several election studies in its holdings containing aggregate data; for the most part these are variables of basic demography and occupational suuclure, as well as political variables. many talks with social scientists have shed light on the fact that research interests are diverse. however, there is a strong desire for a common dataset on the lowest possible regional uniu some regional data files exist, but access is sometimes restricted; data in different files are seldom comparable because they are valid for different demarcations or lime points. data comparability also suffers from boundary or definition changes. some other problems apply to special circumstances in the federal republic of germany. during the past years many data files were created in germany. unfortunately the data are as heterogeneous as the projects and the underiying research interests from which they result although the data files contain precious data, they only have a few variables in common. knowledge of the existence of data files seems not to be widespread, as seen in the duplication of effort that is usual because of the many authorities and researchers that use large amounts of manpower to create machine readable data sets with more or less overiapping variables. usually, existing data files are not easily available. data sets under the jurisdiction of authorities may theoretically be accessible to the public, however, to get the data is another matter. the transmission of data to a third researcher or research institute almost always involves prolong negotiations or severe restrictions. 'presented at lasslst/ifdo international conference may 1985, amsterdam fall/winler 1985 iassist quarterly 21 one of out aims is to build a data pool that can be enlarged by the addition of data from different sources. there are some studies already in the zentralarchiv holdings which contain aggregate data. these sources should be usable for the pool, together with the data we transcribe from printed material to a machine readable format at the present time efforts are made to complete the ifdo list of variables as far as possible for the regierungsbezirk level and later on for the kreis level. we try to collect data from other researchers' projects and studies not only for these high levels of aggregation but also for smaller units. the intended combination of data for different regional levels makes data available for multilevel and contextual analyses. it seems to be reasonable to use data from these sources as the nucleus of the data pool and to complete the gaps that remain in the ifdo list of variables. most of the variables to be added concern settlement structure, education, income, production, health and household conditions. that existing files are faultless. considering the diversity of data files in the university sector, especially their regional and temporal coverage, efforts should be strengthened for the better use of the holdings of the statistical offices of the lander and the federal statistical office. the development of data banks or statistical information systems in most of these institutions during the last decade seems to open new possibilities for such improvement experience will show if these data banks can be used to construct regional data files covering the whole territory of the federal republic. there have been some disappointing experiences in the past the fact that the majority of files have been constructed by transcribing data from printed sources, demonstrates this. on the other hand, the willingness of the authorities to cooperate with the scientific community has increased recently. the data banks and information systems have been established with enormous financial and personal expenditures, and users are now made welcome; a large number of users will mean that their efforts have not been in vain. when the data are merged, they should be subjected to reciprocal control routines. regional data are hierarchical in structure across the different levels of aggregation and thus can be subjected to control procedures. these can use sum, line or row percentages. it is impossible, however, to execute checks "totally automatically". security restrictions and unavailability at lower levels also hinder one from coming in contact with the data. other data, for example relational data, cannot be aggregated. from our point of view, control routines are necessary in any case; whether data from existing files are merged or transcribed from printed sources. if machme readable data are merged, one should not start with the premise where we stand and the problems we face in our efforts to transcribe data from printed sources to a machine-readable format we have built four sas-files, followmg the 'list of ifdo variables' (loiv). we have created one file for each of the three census years: 1950, 1961. and 1970, and another file containing annual data for the variables, area, population, births and deaths. these files have been sent to the nsd. {editor's note: norwegian social science data service). accordingly, three questions arise: fall/winter 1985 22 iassisl quarterly 1. why didn'i we obtain a data file with the full set of ifdo variables for 1980? unfortunately the latest census in the federal republic of germany was conducted in 1970. about 1980 and for several years thereafter a broad public discussion on data protection took place, which prevented a new census. only recently the federal government decided to postpone the next census until 1990. the statistical offfices in our country have annual data, of course, but it has merely been updated by different authorities for various reasons. moreover, there is no single, definite time point for updating all variables. another point to be noted is that many data are registered in connection with special events only (with the exception of a census). let us take for example a student who has lived for the rest of his life in the same house in the city in which he attended university. according to the registration office, he will have died at the age of 75 stil! as a student, assuming no census was carried out in the intervening years and he does not move to another house. examples like this highlight the problems of reliability of data based on information updated by administrative authorities only on the occurrence of vital events. in the future, we will try to procure data of this type for 1980, as well as secure information about reliability limitations. 2. why has the data been stored in separate files? the answer is succinct and very important: great changes occurred in the numbers, boundaries and names of the regional units as a consequence of territorial reforms in germany. the main problem for the development of a regional data file consists in the instability of regional boundaries over time. from 1968 until the end of the seventies, a far reaching territorial reform took place, which affected not only most communes, but caused changes in regional units at higher aggregation levels as well. the territorial reforms were brought about to reorganize the communal self-administration at the local level in order to create an efficient system of common welfare and social security. walter christaller's central-place theory is basic to this concept. for ifdo purposes, the greatest difficulty is the comparison of variables over time. territorial reform did not only aftect communes as a whole, for example the incorporation of an entire commune into another, but many others were divided, and their parts merged with yet other communes. •' figures 1 and 2 give a sense of the amount of change that has taken place in the last decade. indeed, it's difficult to find any commune with stable boundaries over this period. the statistical offices of the lander are able to 'reconstruct' the old commune limits by a "backward recalculation" ("rlickwartsaktualisierung"). when the data on the newly delimited communes are converted into data of communes under the old demarcation, two different systems must be considered if the new commune was an amalgamation of old communes or parts thereof. under the old system, all data on the commune to be divided in the course of territorial reform, was transmitted to the commune that incorporated the largest portion of the population within its enlarged limits. the new system focuses on distributing area and fall/winter 1985 iassist quarterly 23 figure 1: enthicklunc des cemeindebestahdes i9s7 bis 1976 260o0 i 1 26<»1 19y 1953 1959 1960 1961 1962 1963 1944 1965 1966 1967 1968 1969 1970 1971 1972 1973 1974 1975 1976 jahresende figure 2 : entwicklunc der kreisfreien stadte und undkreise 1957 bis 1976 insgtsa til bndkre.\""»• v '**'"»»^ — k/ersfre t sladle brt swttoeise '***--.".. 1957 1958 19s9 i960 1961 1962 1943 1964 1965 1966 1967 1968 1969 i97d 1971 1972 19n 1974 1975 1976 jahf e jende source: staiistisches bundesamt wiesbaden (1977): devblkerungder gemeinden 1976. fachserie 1, bevblkerung und erwerbstatigkeil. reihe 1.2.2. stultgari: verlag kohlhammer. page 6. fall/winter 1985 24 iassisl quarterly population data exactly. other data are appropriated to the new enlarged communes according to the percentage of population they received from the old divided commune. a well-known conclusion is as valid here as elsewhere in regional research. the errors, resulting from boundary changes, that occur in comparisons over lime tend to vanish at higher aggregation levels. 3. do we need data which allow comparisons for regional units over time if the regional units have changed? changes in the boundaries of regional units are modincations of the unils of analysis. if one supports the idea that regions like gemeinde, kreis or regierungsbezirk influence people's lives, one must consider their geographic reality in present as well as earlier times. given the many far-reaching changes in regional units by the territorial reform, the aggregation or disaggregation of the old regions within the present-day boundaries means notjiing apart from getting a stable scanning pattern. it does not, however, have any influence on people's orientations in their geographic context from this point of view, we suggest that the several thousand boundary changes can be tolerated in germany during the territorial reforms. regional analysis at different time points are thus rendered possible. comparisons over time, however, are not only a problem of data availability in germany but also of adequate research design." fall/winter 1985 the changing nature of networking in the research library community by henriette d. avram ' associate librarianfor collections services library of congress networking among libraries is certainly nothing new; to the contrary, libraries have long been pioneers in networking activities. today, i want briefly to summarize that history, but in keeping with the theme of this session, my focus will be on future plans and prospects. we are in an exciting period now— one where the technology and our collective imagination are at such a confluence as to yield exhilarating results for libraries, especially the research library community. from the beginning, networking among libraries has been propelled by our strongly held tradition of resource sharing. the 1960s and 70s witnessed libraries turning increasingly to the computer and the possibilities it held for automating library operations. the development of the marc format at lc, the acceptance of the format by library practitioners, the adoption of the format structure as a national and international standard, and lc's distribution of its cataloging data in this format, marked the beginning of the true era of library networking as we define it today. there then appeared organizations that were new on the scene, namely, bibliographic utilities. these organizations, created to serve the needs of libraries desirous of a central source for cataloging records at a reasonable cost, built growing files of cataloging data to which access was limited to institutions who were members of the particular utility. by the mid-1970s several large databases— oclc, the research libraries group's rlin (research libraries information network), and wln (the western library network), along with the one at lc— coexisted in the united states, but could not be accessed and shared directly. to remedy this situation, efforts were initiated to enable libraries more easily to share data housed on dissimilar systems and thus was created the linked systems project or lsp. the international organization for standardization's open systems interconnection (osi) reference model was chosen by the lsp architects as the appropriate protocol package to run lsp because osi would substantially reduce future development necessary to accommodate new systems, new applications, and new standards being developed in accordance with the osi model. two lsp applications modules were developed— record transfer and information retrieval. record transfer enables records of any type and any number to be transported between systems. information retrieval permits users of one system to access a remote system and to view data found in the remote systems. information retrieval permits users of one system to access a remote system and to view data found in the remote system on their own system, invoking the familiar query command of their local system, effectively overcoming the problem of multiple syntaxes. the osi-based protocols for lsp were fashioned to be of general service and not application specific. thus lsp applications can be expanded to other purposes, e.g., the information retrieval protocol, which is now an american national standards institute standard, is the basis of several projects being planned in the u.s. for accessing remote databases of all kinds, e.g., full text, abstract and indexing, and others. currently, lsp is being used to support the exchange of authority records. lc has distributed via lsp connections more than 2.5 million authority records to rlin and over 1.5 to oclc. lc has received via lsp, over 88,000 authority records created by rlin and oclc libraries which have been added to the file at lc and distributed via lc's cataloging distribution service to libraries all over the world. the next step is to support the exchange of bibliographic records which will enable records to be searched between systems, retrieved and then added to a particular database for use by its patrons. by this augmentation of lsp, a cooperative operation located at lc, nccp (national coordinated cataloging program), will gain increased efficiency. nccp brings together eight research libraries that have agreed to contribute national-level cataloging records to the national database at lc for distribution to the nation's libraries. while the library community was availing itself of advances in technology to forge a national bibliographic network via lsp, using osi protocols, other networks were evolving in the u.s. using different standards. the academic and scientific community was busily laying the foundation for a supemetwork supported by tcp/ip (transmission control protocol/internet protocol) that fall/winter 1990 will support research and scientific investigations. this network, known as the internet, connects universities across the u.s. and links them to supercomputer centers. the internet is a long-haul network that provides national connectivity through the unking of regional networks which cover large geographic areas. nsfnet, the national science foundation network, acts as the backbone of the internet. because of the difference in standards used, the two networks being built were incompatible. the first step toward reconciling this incompatibility was to seek cooperation with educom (a consortium of u.s. institutions of higher learning). accordingly, in 1987, henriette avram invited ken king, president of educom, to come to lc for exploratory talks. these initial meetings opened our eyes as librarians to the enormity of what was happening, what was being planned on the academic side, and also what was missing, i.e., much of what libraries were already doing or had already accomplished. the aim was to connect scholars' workstations on the nation's campuses to each other as well as to supercomputer centers via the internet to support research needs. besides nsf, the players in this grand scheme were influential and represented big money interests— ibm, at&t, and new york telephone. as more was learned of what was envisioned, an additional strong concern emerged— the research libraries on university campuses being wired to each other and to supercomputer centers were also part of the lsp environment. these libraries are the keepers and organizers of much of the information and data that feed the research process. it is the ability of libraries to organize information for retrieval and end user access that makes sharing data over networks viable. we should not lose sight of the importance of the organization of information and the critical role standards play in this process of organization. indeed, it is this technical processing aspect that underpins and makes possible the research and reference functions that will become increasingly important as the supernetwork takes shape and the variety of data on it mushrooms. educom has been instrumental in pushing and tracking legislation currently before congress (s. 1067, sponsored by senator albert gore). if approved, the evolution of the internet will take on immense proportions. of keen interest to research libraries is title ii of the bill, which calls for the creation of a high capacity national research and education network (nren), which will interconnect over 1,000 colleges, universities, research organizations, and, we hope, their libraries. title ii further specifies that— working with other agencies— by 1996, nsf would establish a multi-gigabit nren capable of transmitting 100,000 typed pages or 1 ,000 satellite photographs in one second. the network is to be phased out when national, commercial high-speed networks can satisfy research needs. now that the networking infrastructure that will serve the nation in its various components has been described, it is well worth spending a few minutes discussing the "content" of the network, i.e., what kinds of information and data will be accessible over the network. as i've tried to make clear, in terms of describing and formatting bibliographic data for efficient searching and retrieval, we're pretty much there. the standards are well defined and broadly applied within the library networking environment. for non-bibliographic data, however, we have some way to go yet but strides are being made. having spent its first fourteen years focused primarily on networking of bibliographic data, the library of congress network advisory committee (nac, an umbrella group comprised of library and networking professionals) recently shifted its attention to non-bibliographic databases (which it defines to include full-text, numeric, and graphic data). nac devoted an enure program meeting to this topic last year. entitled "beyond bibliographic data," its goals were to gain a better understanding of the term "non-bibliographic" in the library network context and to begin to appreciate the range and potential of such electronic information. among other things, by the end of the meeting, it was agreed that: librarians must be able to cope with the multiplicity of forms of information; a user interface that will enable scholars and the public to access and display bibliographic and non-bibliographic data files must be devised; standardization and information selection issues remain outstanding; and libraries with local systems and how they affect the relationship of libraries to bibliographic utilities is emerging as a problem: at risk is resource sharing as librarians have known it as attempts are made for economic reasons to seek lower cost alternatives to cataloging on the utilities and thereby not adding expensive cataloging records to a large national database for sharing. the complete proceedings of the meeting are available as part of the library congress network planning papers series. educom has recently accepted a project proposal, "the library and the electronic document environment infrastructure," submitted by mrs. avram on behalf of the library of congress that calls for a full-scale effort coordinated by educom, in conjunction with the various stakeholder communities (libraries, researchers, information processors, publishers, professional organization, etc.). the proposal concentrates on three areas of activity: iassist quarterly 1) to carry out major studies in three core areas of the electronic document environment: a) technology and formats what information resources are needed and in what electronic format? how will the information be organized for retrieval, transmission, and exchange? will the technologies provide solutions to problems of storage, preservation, and presentation of information? b) economic issues and choices who pays for and owns information resources in the electronic document environment? c) roles and responsibilities for libraries how will the library exercise its role as organizer, classifier, and preserver of information and knowledge in die national network? 2) to devise several structural models conducive to the growth of electronic document environment; and 3) to offer programs, seminars, and publications which present findings about the library in the age of electronic research, production, and publishing. what has emerged from this discussion is the notion that through the application of technology to networking, libraries will become boundless and that users, by accessing networks, will become patrons of "libraries without walls." right here in this state, a plan has been issued by the new york state library which details how all libraries in the state— academic, school, public, or other— can become electronic doorways for citizens of new york. an electronic doorway library would make needed information available electronically to users from any part of the state via links to databases and resource sharing programs with computers. research libraries are moving ahead on several fronts through various organizations, both singly and collaboratively, in dealing with non-bibliographic data in a network environment. one of the most promising involves arl, cause (the association for the management of information technology in higher education), and educom forming the coalition for networked information. the coalition will consist of a large and influential group of institutions of higher education, not-forprofit organizations, corporate sponsors, and government agencies. it has set for itself an agenda that includes crafting a set of initiatives to deal with the provision of information resources on the national research and education network. the coalition will focus on issues related to intellectual property rights, standards, licensing, service arrangements, cost recovery fees, and economic models. so far, as of may 4, over sixty research libraries in the u.s. and canada, including the library of congress, have committed to joining the coalition. as we move into the 1990s, it is fair to say that the implications for research libraries of networking and the changing network infrastructure are immense. but, as can be seen, this final decade of the century holds great promise to be an exciting and innovative one as well. and while the task before research libraries in servicing non-bibliographic data in the network setting is staggering, some excellent first steps are being taken. 1 presented at the iassist 90 conference held in poughkeepsie, n.y. may 30 june 2, 1990. nara job announcement center for electronic records job announcement march 1, 1991 within the next two weeks or so, the national archives will announce one or more vacancies for senior archivists to deal with computer records. starting salary is $37,294 with all fringe benefits of u.s. government employment. u.s.citizenship required. if any one is interested, or knows anyone who may be interested, please contact the following: thomas e. brown chief, archival services branch center for electronic records national archives and records administration washington, d.c. 20408 (202)501-5565 tbrown@dcunsn.das.net fall/winter 1990 32 iassist quarterly promoting a computer conference, continued: the experience of the association of public data users by patricia c. becker .city of detroit planning department following publicaton of chuck humphrey's article in the summer 1985 issue of this journal', judith rowe suggested that readers might be interested in our experience with a computer conference for the association of public data users (apdu). ' c. humphrey, getting a turnout: the plight of the organizer. experiences in promoting a computer conference. iassist quarterly 9(2): 14-27, summer 1985. apdu is an organization of organizations, rather than of individuals, bringing together people with an interest in the development and use of public data. because these activities center around the federal government in washington, the membership is entirely american. almost everyone involved in apdu uses demographic data from the census, but many other kinds of data are of interest as well, such as economic data, health statistics, and data on specific populations such as the ageing. members are also interested in the software packages available for processing these data and, increasingly in recent years, in the potential for the use of microcomputers in their everyday work lives. cross-cutting all of this is a concern with federal statistical policy. the apdu electronic conference (or e-conf, as we refer to it) was the brainchild of ken riopelle, apdu board member. ken had previous experience as a conference organizer and promoter for the mott foundation, experience that included users signing on from around the country. i was a veteran conference participant, but had never been an organizer. the wayne state university computing services center in detroit is the "electronic home" for both ken and i, so it made sense to set up the e-conf there. the software of choice was confer, a sophisticated electronic conferencing package developed for use on the mts operating system. the original proposal called for a pilot project, for board members and a few others. this was a group of about 13 people. we began in the spring of 1984. a project account was established at the wayne state computing services center, the e-conf itself was created, and sign-on materials were sent to each member of the group. the initial group was registered into the userdirectory, and materials on the electronic messaging system were provided as well as on confer. since all were coming in on telenet spring j 986 iassist quarterly 33 or auionel, the appropriate phone numbers were provided to each prospective participant a wallet-sized "1-2-3" crib sheet with sign-on instructions was provided. in addition, a sample session of confer was created, printed and reproduced. how well did it work? of the initial group of 13 people, seven (including ken and myselo became active participants. active is defined, here, as signing on at least once a month. three board members signed on, joined, but rarely participated; the remaining three never actually joined the e-conf at all. two major factors seem to explain non-participation; to some degree they are interactive, one with the other. one was lack of equipment, and the other was a lack of of familiarity with using computer terminals. it appears to be necessary, to maintain active participation, to have a computer terminal available in the office, and to be in a position to use it frequently for other, routine work activities. most non-participants either had no access to equipment or had access only at home. however, the seven people who did participate had a great time. items were entered on internal apdu board issues, on federal information policy, on software and data access, and on use of the e-conf itself. enough was going on to keep people interested, so that activity levels did not drop among the seven active participants. after evaluation of the pilot project, the board decided to extend the e-conf to the membership. to promote interest, a $20 credit was offered to each organizational member. in the apdu membership structure, each member organization has a primary representative who is responsible for the dues; additional representatives can be added for a small fee. the $20 credit allowed members to sign up and tr>' out the conference at no cost expenditures in excess of $20, however, had to be paid for "up front", so that apdu would not get into financial difficulty. additional representatives were welcome to participate as well, but were required to arrange for funding from their primary organizational representatives. the computing center's accounting system allows us to control the amount of money available to each sign-on id, so there was no problem managing the accounts. in january 1985, the entire membership received a mailing which explained the opportunity to participate in the e-conf and included forms for signing up. unfortimately, the response can best be described as underwhelming. between the time of mailing and the aimual "people" conference in october, only ten sign-on ids were issued; of these, only three or four became active paricipants. we picked up a few more as a result of heavy promotion at the conference in october, as well as having added some new board members. at this point in january 1986, 37 sign-on ids have been issued but only 14 can be described as active participants. the number of items has grown to 121 and discussions continue to be lively. a great deal of information is being exchanged. several specific decisions have been made "electronically". draft resolutions, letters of comment and the like, representing proposed positions for apdu to take vis-a-vis the federal statistical establishment have been put into items to be reviewed by the board and other interested parties. and the cost? it appears that 1985 expenditures will be under the budgeted figure of $2500, primarily because there are fewer members than anticipated taking advantage of the $20 credit by budget category, costs can be broken down as follows: e-conf maintenance: disk space, on-line time for organizers, printing manuals, and project accoimting. these costs would spring 1986 34 iassisl quarterly have beeen higher had the organizers not had other wayne state accounts through which to participate and do some of the maintenance work. $425 use of system by executive secretary. $425 non-reimbursed use by board members. those who are in work situations in which it is difficult to obtain reimbursement are allowed unlimited access at present this item also includes use of the system by the organizers of the armual conference, for which two people in different cities accomplished most of their work via electronic messaging. $950 $20 credits (excluding accounts in the previous item). $250 overall, the consensus of the board is that the expenditure is worthwhile and justifiable within our total budget scenario. there has been some criticism of the project within the apdu membership, primarily of the cost to the members. some feel that they cannot afford participation (since $20 really doesn't go that far on telenet or autonet). as in any organization, members belong for different reasons and have different agendas; not all our members find interaction with other members to be a useful expenditure of their time and/or money. there also remains a significant problem with equipment access several members who wish to participate have been stymied by the lack of a terminal, or a good modem and communication software. this problem should decrease over time, as more and more organizations acquire microcomputers. we are plaiming to have an on-line demonstration of the system at our next annual conference (scheduled for october 1986), to promote interest in the system. we are also including a column entitled "from the e-conp in our monthly printed newsletter, both to provide information for those who are not conference participants and to encourage them to join. there is another participation problem of a different kind: some participants sign on regularly, but rarely contribute any responses. this is the reverse of the "habitual commentator" described in the humphrey articled two factors are at work here: one technical and one involving personalities. many participants sign on and simply "dump" the conference activity to disk, to be read later, or to a printer, without actually reading it while they are signed on. this has the negative effect of discouraging responses, since they must sign on again to enter them. the other factor is, as humphrey' described, the "implicit norm to say nothing." however, i think this is less a problem in our particular conference than it might be in others. the fact that most of the participants have met each other "in person" at the annual conference helps we generally know the people to whom we are talking electronically. all in all, apdu rates the e-conf a success and it has become an integral part of the organization's functional mechanism. we would like to see greater participation and will continue our efforts in that direction. what we have, though, is a system that works people are communicating, and that's what an organization is all about ibidem ibidem spring 1986 ilassist newsletter vol. 1, no. 3 book n ot i c e s / kathleen m. heim lucci, york; rokkan, stein; and meyerhoff, eric. a library center of survey research data: a report of an inquiry and a proposal . new york: columbia university school of library service. june 1957. [available in microform from the: columbia university, school of library service, new york, new york 10027] it has been twenty years since the lucci and rokkan report first appeared. in many ways the data archive movement may be said to have begun with its publication. many ideas presented therein have been built upon by later archivists and scholars. it is particularly interesting that while the columbia library school received the grant for the report, libraries have, until quite recently, been reticent to become involved in data archive activities. [see review of howard dalby white's ph.d. dissertation below.] background of the report rapid growth of opinion polling and survey research following world war ii made cross-national research a real possibility for scholars through the use of secondary analysis of comparable information collected in different countries. the difficulties of locating and obtaining even domestic researches, however, proved discouraging. the need for a library of survey research data which would assemble the more important survey data and make them readily accessible to scholars for secondary analysis prompted professor s. m. lipset then at columbia university to propose that an investigation be made of the need for and the problems of establishing an international library center of survey research materials. in response to this proposal the behavioral sciences division of the ford foundation awarded a grant in the spring of 1955 to the columbia university school of library service to undertake such an investigation. york lucci was given responsibility for the investigation in the u.s. and stein rokkan for europe. _ in addition, eric meyerhoff served as a consultant on the more technical archival considerations. the field investigators were charged with ascertaining and evaluating four areas: 1) potential utilization of a center for research materials; 2) availability of research material, its cost and difficulties in acquiring; 3) adequacy of available data in terms of their potential for secondary exploitation; 4) research and administrative requirements involved in such a library center. the results of the lucci u.s. investigation and rokkan european investigation appeared as separate reports since the situations in the u.s. and europe differed greatly with respect to the availability of survey materials and the extent of interest among scholars. 25. ji^sist newsletter vol. 1, no. 3 part i. the situation in the u.s. and recommended action background and purposes lucci discusses two chief reasons for a centralized archive: lack of funds for the collection of primary data and inadequate use of data already collected. he then outlines past efforts to collect and disseminate results of polls and surveys such as the 1938 public opinion quarterly 's aipo poll results, various periodical compilations published for short intervals, and the roper depository at williams college. he concludes that published releases do not permit systematic analyses and that any archives that exist have been too limited and have failed to develop archival practices permitting usage by a wide community of scholars. the "loss of data" receives much attention in the lucci report. a salient example is that of the canadian institute of public opinion where all but eight of 212 polls between 1941-1951 had been destroyed. with other examples of data loss, such as the experiences of the italian doxa organization, the swedish gallup institute, and the u.s. office of war information, lucci builds a strong case for the need to undertake a strong preservation effort. parallel developments in other disciplines of basic tools for research (cyclotrons for scientists, libraries for historians) are compared with the lack of basic preservation of the raw material for social science research. the human relations area files is discussed as the best known effort in the social sciences to assemble information from a variety of sources and make it available to a wide community of scholars. one hundred and thirty seven social science researchers in the u.s. were contacted for ideas and reactions concerning a central data library for the social sciences. their ideas are consolidated and set forth by lucci in the following sections. sources of data the massive accumulation of data gathered by universities, the government, independent agencies, and foundations cannot be readily documented, but lucci estimates that five to ten million interviews have been conducted annually in the post-wars years. one market research organization alone reported 25 million cards in storage. though much of the data may be insignificant, if even five percent offer the possibility of scientifically useful secondary analysis, a quarter to a half-million cards a year might be stored. urgently needed is some means of locating that fraction of theoretically significant material, assembling it, making its contents available to researchers, and encouraging its exploitation. queries to survey organizations and academic institutions produced an overwhelmingly positive reaction about sharing data. lucci details at length responses from prominent centers, polling agencies, governmental agencies, and commercial research organizations and includes a section on "conditions imposed by contributors." the main objections he found to a proposed center were feelings that one cannot work successfully with someone else's data unless there is 26. ^^sist newsletter vol.1, no. 3 close liaison with the person or agency initially responsible for collecting the data. lucci counters this by advocating complete documentation and noting that "if survey research is to make any claim to science it can hardly refuse, indeed it should welcome, having a set of facts subjected to close scrutiny and the possibility of alternate interpretation. utilization of an archive of survey data uses of survey data for secondary research are discussed and six major points elaborated upon: 1) as the only primary source of knowledge about certain facts or patterns of behavior; 2) as a means of testing hypotheses; 3) the cumulation of cases; 4) to prepare for new primary research; 5) to assist the work of the historian; and, 6) to train students. objections to the proposal for establishing a center of survey research materials related to the extent to which utilization could be anticipated. respondents to lucci 's questions felt that technical familiarity with methods of quantitative research was not widespread and that only large universities such as michigan, columbia, and chicago could make use of such facilities. lucci feels that availability will increase the skills of the entire social science community, and then reports comments of respondents, summarizing at length areas where respondents felt the center could be especially useful. further justification for the establishment of a center is given by a bibliography (pp. 134-38) appended to the report of published works based on secondary analysis, which represent different ways in which secondary analysis techniques can be used; expansion upon the use of secondary analysis by students and commercial organizations; and, the need by international organizations such as unesco for access to such a data center. conclusion and recommendations lucci summarizes by noting, "on the basis of the present inquiry, it can be stated flatly that the data are available. . .and it seems fairly clear that the possibility of their use is present and can be developed further." such a center would maximize the use of available social science talent (especially at universities with minimal facilities); maximize the use of research funds; promote the comparability of data collection; facilitate cross-national research; preserve data in a systematic manner; and, spearhead new studies where lacks are noted. a "specific proposal" is outlined with functions of the center noted at length: systematic collection and preservation of data; an index to data stored; information activities about the data to be disseminated to the research community; promotion of the center's materials; provision of duplicate sets of data for researchers with their own facilities; provision of tabulations where appropriate; maintenance of records of secondary research activities and publications; and, promotion of training and standardization. each of these functions is discussed in great detail along with possible locations for the center, staff considerations, equipment and storage needs, operating costs and sources of funding. 27. sist newsletter vol.1, no. 3 part ii: a review of the situation in western europe an d some recommendations for possible action rokkan sketches rough estimates of the data accumulation problem in western europe. he estimates that in the four major nations: france, italy, the united kingdom and germany, half a million interviews have been generated in each nation since 1951. among the smaller countries: austria, belgium, denmark, finland, the netherlands, and switzerland, the yearly production of interviews is on the average 75-100,000. thus, the total interviews for western europe are between 3 and 4 million annually. the problems of preserving, storing and choosing relevant data from this quantity is seen by rokkan as directly linked to problems of social science utilization. most polling agencies limit analysis to the presentation of overall response distribution, never analyzing the data as thoroughly as their methodological quality and their theoretical relevance seem to justify. surprisingly little has been done by western european social scientists to avail themselves of these analysis opportunities. the lag in utilization reflects the conservative reluctance of great numbers of social scientists to accept the new techniques as scholarly and scientific, the uneven distribution of statistical analysis skills among academic social scientists, and the difficulty of access to the primary materials in the polling agencies. during 1955-56 rokkan discussed problems of storage and data utilization with directors and officers of survey research organizations and with university teachers and research workers in twelve countries of western europe, as well as at regional meetings such as wapor, esomar, the congress of the international political association, and the international sociological association, in order to elicit reactions to the possibility of a solution through the development of an international archive of survey materials. the possibility of an international archive was presented to survey organizations by a questionnaire devised by rokkan. the response was unanimous in emphasizing the need for systematic encouragement of secondary analysis, although opinions varied about the steps to be taken to facilitate secondary analysis of data already assembled. the responses were capsulized by esomar president, leif holback-hanssen: western european materials should be stored and classified in a separate facility located within the western european region; development of the archive should be planned in close association with representatives of the survey profession; and, the archive should be built up gradually through a series of pilot analyses. rokkan's queries to academic social scientists are discussed in the context of the status of survey research in western europe: i.e., there was little being conducted. most social scientists tended to favor independent gathering of data in order to build gradually to the need for a regional or international repository. rokkan then discusses the format of pilot analyses and possible topics of these analyses. 28. insist newsletter vol. 1, no. 3 rokkan concludes: there is a need for fuller utilization of the bodies of information assembled by survey organizations through systematic analyses from the primary record; a need for the organization of a program of comparative secondary analyses of survey materials in selected substantive areas and, concurrently with them, the development of an archive of selected materials; and, a need for "groundwork" analyses and an archive within the western european region. cost estimates are given. appendices as noted above, a selected bibliography of works done using secondary analysis techniques is included. another appendix details methods for indexing and cataloguing archival materials. reviewer's comments the lucci-rokkan report is a fundamental document for the data archivist. the well formulated report gives the background of the data archive movement and predicts many of the developments we are now experiencing. this report is essential reading for all involved in data archives, because of the perspective it provides on the u.s. and western european research communities and because of its visions of the future. white, howard dalby. "social science data sets: a study for librarians." ph.d. dissertation. university of california, berkeley, 1974. as noted in the review of the lucci-rokkan report above, libraries have not generally accepted responsibility for acquiring data in machine-readable form although a number of scholars and librarians have argued that machine-readable data is a logical extension of a library's collection. white explores the relationship between libraries and archives, giving a historical overview of arguments that advocate the placement of data archives within libraries. he compares library and archival functions and although he explains that many tasks are similar, he concludes that archival tasks confront librarians with much that is unfamiliar, such as "opening tapes"; making back-up copies; updating originals; reproducing and distributing tapes; providing programs; and, training users in computer techniques; thus, making the merger of the two unworkable. in "data archives as publishers," (chapter iii) white discusses the most significant distinction between libraries and archives: major archives "publish" data sets. in his discussion. white gives a wealth of detail about the process of publication within data archives--material that is unavailable elsewhere. rather than merge libraries and archives white argues for an alternative: using libraries to make information about data more widely available by acquiring from publisher archives the human-readable materials by which the content of data files can be known. suggestions for better bibliographic control over machinereadable data through traditional library catalogs are given. because all social scientists are not employed at institutions with data archives. white argues that there is a need for an organization with cross-disciplinary responsibil ities--a library or social science information center--to function as a locale for codebook browsing or searching. ^ssist newsletter vol. 1, no. 3 part ii of white's thesis, "buying patterns at two publisher archives," is an analysis of the transaction records of the international data library and reference service of the university of california, berkeley (idl&rs) and the roper public opinion research center in will iamstown, massachusetts. the analysis is meant to serve as a basis on which libraries may predict the potential need for codebooks and data sets. white's intensive exploration of the type of materials bought (human-readable vs. machine-readable) also presents an interesting picture of the buyer of the data bases who he hypothesizes tend to be "high-status" persons (professors and researchers rather than students); from "high-status" departments (e.g., among the top ten in their fields); male; and geographically "near" to the archive from which data sets are purchased. after subjecting his data to statistical tests. white finds that only sex is a predictor and that although those close to idl&rs buy more, the reverse is true for buying patterns at roper. implications for libraries the major implication for librarians is that human-readable data sets are valued by social scientists in their own right (white demonstrates that information about data is as often purchased as the data themselves) and their purchase could be dele gated to libraries. the question of the prestige attached to codebook purchases, as indicated by white's exploration of who buys, is a fascinating insight into the sociology of the social sciences, and white has brought out a unique factor in acquisitions policy that libraries might well heed. reasons for putting human-readable tools, such as codebooks, in academic libraries include 1) introducing them into institutions that are more numerous and visible than existing data archives; 2) assuring that the tools are accessible to all disciplines and professional schools on campus; 3) deploying them with the other bibliographic tools and substantive literature available to social scientists; 4) improving bibliographical information on codebooks by bringing catalogue cards on them together with cards on associated monographs; and, 5) improving bibliographical information on data access tools generally by making them discoverable through the local catalogue, and possibly through union catalogues up to the national level. the final chapter of white's dissertation, "acquisition of data access tools," concentrates on guidelines for librarians interested in buying data access tools. appendices include correlation matrixes describing buyers at roper and idl&rs and a substantial bibliography on data archives and librarians. writing this review from the academic librarian's viewpoint, one is impressed with the wealth of information marshalled to provide a case for library purchase of human-readable access tools. if all areas of library buying had as much organized data as dr. white provides in this thesis, libraries would move from the realm of somewhat subjective buying to a more scientific policy. thus white's conclusions act at once as a model for acquisitions policy at a universal level, as well as in the specific case he addresses. 30. sist newsletter vol. 1, no. 3 data archivists, especially those working in isolation from traditional libraries, will want to consider the implications of white's work for their field. greater acceptance by libraries of access tools to data sets will generate a broader base of potential archive users and, since white argues cogently for the continuing housing of the actual data sets in archives, there is no reason to fear that the very specialized services of archives will be subsumed by libraries without staff expertise to promote their exploitation. we hope that dr. white will consolidate his groundbreaking findings into articles to be disseminated in the library and data archive press. the thesis itself is a mandatory purchase for library schools and a critical accumulation of data for archivists. we hope that dr. white's thesis is the first of many dissertations which will explore in greater depth the problems and characteristics of archives and their users. quantum completes survey quantum members in germany have completed a survey of completed, ongoing, and planned research projects in quantitative history. this survey has just been published as volume i in a new series of klett verlag, stuttgart, the hsf, which will be concerned with quantitative social scientific analysis of historical and process-produced data. the book entitled, the quantum documentation , is available from ernst klett verlag, rotebuhlstrasse 77, postfach 809, d-7000, stuttgart 1, at a cost of dm 39. the full bibliographic citation is: w. bick, p. j. miiller, h. reinke. quantum documentation: quantitative historische forschung 1977/quantitative history 1977. stuttgart: ernst klett verlag, 1977. new organizations/ reorganization bass hosts cessda meeting to create if-do belgian archives for the social sciences hosts committee of european social science data archives to create international federation of data organizations on the invitation of the committee of european social science data archives (cessda), representatives of data organizations from the united states, canada and europe gathered together at louvain-la-neuve, belgium on the 20th and 21st may, 1977 to discuss and formulate further means of mutual cooperation in the area of social sciences data archiving services. vol263 10 iassist quarterly fall 2002 iassist quarterly fall 2002 11 by james reid * introduction this paper will summarise work undertaken on behalf of the uk academic community to evaluate and develop a gazetteer server and service which will underpin geographic searching within the uk distributed academic information network. it will outline the context and problem domain, and report on issues investigated and the findings to date. lastly, it will pose some unresolved questions requiring further research and speculate on possible future directions. the context the joint information systems committee (jisc) is a strategic advisory committee working on behalf of the funding bodies for further and higher education (fe and he) in england, scotland, wales and northern ireland. geoxwalk – a gazetteer server and service for uk academia the jisc promotes the innovative application and use of information systems and information technology in fe and he across the uk by providing vision and leadership and funding the network infrastructure, information and communications technology (ict) and information services, development projects and high quality materials for education. fig. 1 provides an architectural summary of how jisc manages these areas within what is referred to as the jisc ʻinformation environment ̓ (jisc ie). the jisc ie provides access to heterogeneous resources for academia, including bibliographic, multimedia and geospatial data and associated materials. the geoxwalk project was conceived as a development project to build a shared service that would service the jisc fig. 1. the general architecture of the jisc information environment. source: information environment: development strategy 2001-2005 (draft). 10 iassist quarterly fall 2002 iassist quarterly fall 2002 11 ie by providing a mechanism for geographic searching of information resources. this would complement the traditional key term and author type searches that have been supported. ʻgeo-enabling ̓other jisc services would provide a powerful mechanism to support resource discovery and utilisation within the distributed ie. problem domain geographic searching is a powerful information retrieval tool. most information resources pertain to specific geographic areas and are either explicitly or implicitly georeferenced. the ukʼs national geospatial data framework (ngdf) estimates that as much as eighty per cent of the information collected in the uk today is geo-referenced in some form. geography is frequently used as a search parameter, and there is an increasing demand from users, data services, archives, libraries, and museums for more powerful geographic searching. however, there are serious obstacles to meeting this demand. geographic searching is often restricted because geographic metadata creation is excessively resourceintensive. accordingly, many information resources have no geographic metadata, and where it exists, it usually only extends to geographic names. search strategies based on geographic name alone are very limited, although they are a critical and often the only access point. an alternative is to geo-reference the resources using a spatial referencing system such as latitude and longitude or, in the uk, the ordnance survey national grids of great britain and northern ireland. the existence of a maze of current and historical geographies has created a situation where there is considerable variation in the spatial units and spatial coding schemes used in geographic metadata. many geographic names have a number of variant forms, the boundaries in different geographies do not align, and names, units and hierarchies have changed in the past, and will continue to change. in 1990, the uk data archive (university of essex) identified ninety different types of spatial units and spatial coding schemes in use in their collection and more have since been added. by way of illustration, the resource discovery network (rdn) resourcefinder (http://www.rdn.ac.uk/) does not currently have a specific geographic search function. instead it simply searches using a text based mechanism, so in order to find information referring to a particular place, the place must be referred to by name in the description field of the resourceʼs accompanying metadata. currently, searching by other geographical identifiers such as minimum bounding rectangle, postcode, or county is therefore not possible. clearly, no single system of spatial units and coding will suit all purposes, as people conceptualise geographic space in different ways and different servers deploy differing geographic naming schemes. ideally, users should not be forced to have to explicitly convert from one ʻworld ̓view to another. however, it is impractical for most service providers to support more than a few geo-referencing schemes, or develop the means to convert from one to another. a comprehensive gazetteer would be necessary to perform these translations (or ʻcross-walks ̓– hence the name geoxwalk). such a gazetteer could link a controlled vocabulary of current and historical geographic names to a standard spatial coding scheme, such as latitude and longitude and/or the ordnance survey national grid(s). the end result would provide the capacity for what might be referred to as ʻgeographical agnosticism ̓(see fig. 2.) this is what the geoxwalk project aims to deliver. background and rationale behind geoxwalk geoxwalk was funded by the jisc as a twophase development project, the principal aim of which was to assess the feasibility of developing and providing an online, fast, scaleable and extensible british and irish gazetteer service, which would play a crucial role in supporting geographic searching in the jisc ie. the project was a joint one between edina (data library, university of edinburgh), and the history data service (data archive, university of essex).fig 2. example case of geoxwalk to ʻcross-walk ̓different geographies 12 iassist quarterly fall 2002 iassist quarterly fall 2002 13 the general aim of the project is to investigate the practicability of a gazetteer service (a network-addressable middle-ware service), implementing open protocols, specifically the alexandria digital library gazetteer protocol (adl 1999, janée and hill 2001) and the open gis consortiumʼs (ogc) filter specification (ogc 2001), to support other information services within the jisc ie. firstly, the service would support geographic searching and secondly, it would assist in the geographic indexing of information resources. as the project has progressed, it has become clear that there is also a requirement for it to act as a general reference source about places and features in the uk and ireland. phase i was conducted as a scoping study to determine the feasibility and the requirements for such a service., phase ii, which commenced in june 2002, aims to develop an actual working demonstrator geo-spatial gazetteer service suitable for extension to full service within the jisc ie. technical implementation for conceptual purposes the geoxwalk project can be decomposed into a number of interdependent though to some degree independent features: i) a gazetteer database supporting spatial searches; ii) middleware components comprising apis supporting open protocols to issue spatial and/or aspatial search queries; iii) a semi-automatic document ʻscanner ̓that can parse non-geographically indexed documents for placenames, relate them to the gazetteer and return appropriate geo-references (coordinates) for confirmed matches – a ʻgeoparserʼ. i) the gazetteer database though the gazetteer database itself is of course crucial, geoxwalk extends the concept of a traditional gazetteer. gazetteers typically only hold details of a placename/ feature along with a single x/y coordinate to represent geographic location . in order to provide answers to the sorts of sample queries that geoxwalk will be expected to resolve (e.g., what parishes fall within the lake district national park? what is at grid ref. nt 258 728?, what postcodes cover fife?) each geographical feature held in the database needs to provide, as a minimum: a) a name for the feature, b) a feature type and c) a geometry. in the case of the latter, the simplest geometry would be a point location. significantly, however, geoxwalk accommodates recording of the actual spatial footprint of the feature, i.e., polygons as opposed to just simple points. by using a spatially enabled relational database management system (rdbms) , geoxwalk can then determine at runtime the implicit spatial relationships between features.this contrasts with the classic ʻthesaurus ̓ approach in which predetermined hierarchies of features are employed to capture the spatial relationships between features. holding such extended geometries in the database therefore permits more flexible and richer querying than could be achieved by only holding reduced geometries. furthermore, all polygonal features can be optionally queried as point features if desired. a feature type thesaurus has been implemented based on a national dialect of the adl feature type thesaurus (adl ftt) (http://www.alexandria.ucsb.edu/~lhill/featuretypes/ ) and provides a means of controlling the vocabulary used to describe features when undertaking searches. phase i of the project had identified the adl ftt as a suitable candidate given its relative parsimony balanced against richness and adaptability. the database schema used in geoxwalk is itself based upon the adl content standard (http: //www.alexandria.ucsb.edu/gazetteer/gaz_content_ standard.html). examination of the adl standard showed that it could embody the richness of an expanded gazetteer model. it further allows temporal references to be incorporated into entries and defines how complicated geo-footprints can be handled. explicit relationships can be defined, which is of particular use when a gazetteer holds significant amounts of historical data for which geometries do not exist, although implicit spatial relationships between features based on their footprints is arguably the more flexible approach to favour. the standard also recognises that data sources will not hold all the detail permitted by the standard. the minimal requirement for an entry and the holding of attribution information with each section,and in some cases, on fields within sections, means that features can be easily introduced into a gazetteer and be updated with additional information from other sources at a later date. potentially this could be save a significant amount of time when the gazetteer is constructed from a variety of data sources, as is in fact the case with geoxwalk. also, from a pragmatic point of view, the proposed standard has been implemented in a web based gazetteer service, containing over 10 million entries. the current implementation of the gazetteer resides in an ingres database with customised external support routines providing the spatial functionality. benchmarking of this solution against an implementation in oracle 9i is due to commence shortly. response times and scalability issues are important design goals and it is essential that the geoxwalk server does not become a bottleneck for other servers using it to answer their geographical queries. ii) the middleware components two protocols are currently supported by geoxwalk 12 iassist quarterly fall 2002 iassist quarterly fall 2002 13 – the adl gazetteer protocol (janée and hill 2001) and the ogc filter encoding implementation specification (ogc 2001). in order to process incoming queries a suite of servlets has been written that transforms the xml based queries generated by the client and maps these to database specific sql statements. additionally, in a sample demonstrator that has been developed, an ogc web map server (wms) is deployed to produce map based visualisations of the returned geographical features. the demonstrator is purely for illustrative purposes as the primary focus of the geoxwalk server is to service machine to machine (m2m) queries. a related project, go-geo! (http://www.gogeo.ac.uk), a geospatial metadata discovery service for uk academia, utilises the geoxwalk server to perform spatial searching as illustrated in fig. 2. simple object access protocol (soap) implementations have also been developed. iii) the geoparser much of the data and metadata that exists within the jisc ie, while having some sort of georeference (such as placename, address, postcode, county etc.) is not in a format that allows it to be easily spatially searched. one task of the project was to investigate how existing, non-spatially referenced documents could be spatially indexed. using the gazetteer as reference, a prototype rule-based geoparser has been implemented that can semiautomatically identify placenames within a document and extract a suitable spatial footprint (see fig. 3). the approach taken has not relied on a ʻbrute force ̓ approach as this is too resource intensive and in tests proved too unreliable. the rule-based approach takes account of the structure of the document and the context within which the word occurs. as a refinement to the basic geoparser, a web-based interface has also been developed to allow interactive editing/confirmation of the results of the geoparser (fig. 4). resultant output from this stage can fig. 3 a sample geoparsed document underlined words have been automatically identified by the geoparser as potential placenames and where possible a georeference (osgb national grid coordinates) derived from the geoxwalk gazetteer. 14 iassist quarterly fall 2002 iassist quarterly fall 2002 15 then be used to update the documentʼs metadata to include the georeferences, thus making the original document spatially searchable. outstanding issues the development of both the gazetteer and the geoparser have identified a series of questions that require further investigative research. amongst the more intractable of these are issues associated with map conflation and improvements to the geoparser. the first of these, referred to elsewhere as ʻmap conflation ̓(yuan 1999) essentially is concerned with the identification and resolution of duplicate entries in the gazetteer. in essence the problem is one of being able to discriminate between existing and new features and to be able to determine when two (or more) potentially ʻsimilar ̓geographic features are sufficiently ʻsimilar ̓to be considered as the same feature. as geoxwalk dimensions a geographic feature on the basis of: (i) its name; (ii) its feature type (adl ftt entry) and (iii) its geometry (spatial footprint). there are three distinct ʻaxes ̓along which the candidate features must be compared (see fig. 5). .furthermore, the question arises of how closely on all three axis must the features correspond and whether one axis should be weighted more highly than another. as geoxwalk relies on answering queries by using the implicit spatial relations between features using their geometries, one ruleof-thumb would be that afeatureʼs geometry is its principal fig. 4 an editing interface to the geoparser to assist in the semi-automatic georeferencing of documents. 14 iassist quarterly fall 2002 iassist quarterly fall 2002 15 axis and takes precedence over name and type. however, as more historical data is added to the gazetteer this problem intensifies. take for example the case of london (contemporary) and londineaum (roman); while both may be regarded as the same place, they have radically different spatial footprints and possibly feature types (town vs. city). the derivation of some metric that would allow confidence limits to be attached to features as they are added to the gazetteer is therefore a pressing research priority. a lot of refinement work could be conducted on the current basic implementation of the geoparser, specifically optimising it in terms of performance and accuracy in order to facilitate realtime geoparsing as well as minimising the number of false positives identified. additionally, the capacity to derive a feature type automatically as well as the footprint itself would be extremely useful. the future the project ends in june 2003 at which point evaluation by jisc may lead to an extension to a fully resourced shared service within the jisc ie. however, significant interest in the project exists outside the academic community that could forseeably lead to a more ʻpublic ̓service that is of relevance to a wider audience than just the uk academic sector. indeed, the model employed is sufficiently adaptable to scale to the development of a european geoxwalk comprising a series of regional geoxwalk servers. such a facility would form a critical component of any european spatial data infrastructure (esdi). references. (jisc 2001). joint information systems committee, information environment: development strategy 20012005. (adl, 1999). alexandria digital library project. [online] http://www.alexandria.ucsb.edu/ (01 october 1999) and adl. alexandria digital library gazetteer development information. [online] http://www.alexandria.ucsb.edu/ gazetteer/ (01 october 1999). (janée and hill 2001) janée g. and hill, l., (2001), the adl gazetteer protocol, version 1. available online: < http://alexandria.sdc.ucsb.edu/~gjanee/gazetteer/ specification.html> (ogc, 1999). open gis consortium inc.,filter fig. 5 the three ʻaxis ̓used to discriminate between new and existing features http://alexandria.sdc.ucsb.edu/~gjanee/gazetteer/index.html http://alexandria.sdc.ucsb.edu/~gjanee/gazetteer/index.html 16 iassist quarterly fall 2002 iassist quarterly fall 2002 17 encoding implementation specification version 1.0.0, (19 september 2001). available online at: (yuan, 1999). yuan s. (1999), development of conflation components, geoinformatics and socioinformatics ʼ99 conference ann arbor, 19-21 june 1999, pp. 1-13. geoxwalk phase i documentation available at: http: //www.geoxwalk.ac.uk. the project was presented at the 2002 joint digital libraries conference in oregeon portland * james reid, geoservices delivery team, edina, edinburgh university data library, george square, edinburgh, eh8 9lj, scotland. email:james.reid@ed.ac.uk. http://www.geoxwalk.ac.uk http://www.geoxwalk.ac.uk iassist quarterly 3 colonial control information available for inuit family research disko bay church and census records the following three papers were presented at the iassist '87 conference in a plenary session entitled research trends . the object of the session was to focus on the effect of new trends in research and data collection and the advance of technology on the use of data and data management techniques. by per nielsen^ danish data archives 'presented at the international association for social science information service and technology (iassist) conference held in vancouver, british columbia, canada, mav 19-22, 1987 ^i am indebted to kirsten elisabeth caning for the articles she sent me, and to jens ludvig wagner at the dda for assistance and support when i tried to get an overview of the project 1. historical sources from greenland indigenous people imder colonial rule were not left in peace! this was not because of any sincere interest in the ethnic peculiarities of the people, but rather because bureaucratic registration also served the purpose of ensuring that the path of development was in accordance with the stipiilations of the colonizing government in the case of the eskimo population in greenland, the government was in copenhagen, where difterent authorities requested very detailed information on personnel resources, production, buildings, etc. in most cases, the information was required in order to evaluate economic development; however, even health conditions, births and deaths, including reasons for death, were reported. and in realm of the church, baptisms, confirmations, weddings, and burials were reported to copenhagen. prior to 1774, the "colonial initiative" was mainly a private one, consisting of tradesmen on one hand, missionaries on the other. in 1774, the main administrative tasks were centralized within royal greenland trade, leaving the chiu-ch to take care of its own affairs. from 1782 to 1950, greenland was administratively divided into two inspectorate regions ,north and south, each with its own inspector; the division line was located between egedesminde and holsteinborg. needless to say, there were modifications introduced to this system during the last decades of the 19th century; and many administrative changes took place during the first half of the 20th century. however, a detailed description of these is beyond the scope of this short introduction.' 'the above outline is based on kirsten elisabeth caning, "personalhistoriske kilder i gr0nlandske arkiver." personalhistorisk tidsskrift 1979. fall/winter 1987 4 iassist quarterly during the period preceding 1951. all church matters as well as most issues concerning education were governed by the same administrative unit in copenhagen. it is from this clerical administration that most of the information dealt with in this article originates. 2. the church records chtitch records were maintained in the greenland parishes during the first half of the 18th century; most of them have disappeared, and a couple of preserved manuscripts are incomplete in that they include information on the top strata of the greenland poptilation only. from the second half of the 18th century, there are preserved church records from upemavik, godhavn and egedesminde. the lists of those baptized (christening lists) represent the majority of entries in the early years. the salvation of the souls of the heathens was one of the major preoccupations of missionaries; consequently, the number of baptisms performed was an indication of the efficiency of the missionaries. the church records from the beginning of the 19th centur>-, held in the danish archives, are more complete, especially those from many of the parishes in the north greenland inspectorate. the records from the south greenland inspectorate are not so complete, in part due to the loss of a lot of these documents during their transportation to denmark aboard the ship "hans hedtoft" (januar>1959)." the church records which have been preserved consist of four separate types of lists representing four different types of events: christening lists, confirmation lists, marriage lists, and burial lists. concommitandy, the data sets stored at the dda (the "raw data") are separated, in their initial form, into 4 different files: dda-0311: population history of greenland 1800-1930: christening lists dda-0312: population history of greenland 1800-1930: confirmation lists dda-0313: population history of greenland 1800-1930: marriage lists dda-0314: population history of greenland 1800-1930: burial usts "in the "danish titanic" case, "hans hedtoft" collided with an iceberg on her maiden vovage to greenland, january 30th, 1959. all 95 crew members and passengers died. 3. the census records the missionaries were also responsible for the registration of all people in their districts (including those who were not christians, and who, consequently, never appeared in the church records). the censuses in greenland contained the same information as those in denmark: name, age, occupation, and posiuon in household. in addition to these, the greenland censuses contained information on race (eskimo, mixed and european) as well as an indication of whether or not the person had been baptized. at the national archive in copenhagen, there are census records from: 1834, 1840, 1845, 1850, 1860, 1870. 1901, and 1911. those from the umanak and disko bay districts, for all censuses except the 1911 census (because of the 80-year access restriction to personal records at the national archives), have been made computer-readable and are available as: fall/winter 1987 iassist quarterly 5 dda-0645; population historyof greenland 1800-1930: census lists 4. the data collection process 3. geographic coverage and time period included in the datasets in the initial definitive stages of the "demographic histoi}of greenland, 1800-1930" research project, both geography and time periods had to be taken into account the above outlined existence of nearly complete archival series of church records and census records in archives situated in or near copenhagen, has some bearing on the geographic regions and time periods to be covered. geographically, the north greenland inspectorate had the most complete records. consequently, the rather isolated regions in northwest greenland, the disko bay district and the umanak district (isolated in terms of inand out-migration from/to the sunotmding regions, so that few individuals "disappear" via migration) were selected; from north to south (or rather, "around the bay") in the umanak and disko bay districts thus defined, umanak, godhavn, ritenbenk, jakobshavn, christianshab, and egedesminde are the larger township areas included in the project with respect to time period covered, the availability of records suggests that the registration should begin around the mm of the 18th century, i.e. aroimd the year 1800. for reasons of restricted access (and, of course, as a means of reducing the resources required for the data collection process), the registration was stopped at the beginning of the 20th century. coding was done from the original sources, or copies thereof by the principal investigator, kirsten caning, together with a smdent aide (see figure 1).^ the coding sheets were then sent out to have the information transformed to a computer-readable medium. the machine-readable data were deposited with the dda, where jens ludvig wagner (who was experienced in demographic data and family reconstitution methods from his work with several other historical projects) took over the transformation tasks as they were requested by ms. caning.' during this process, the following rectangular files were generated: dda-0311: 17,248 baptism entries, each with up to 53 variables dda-0312: 10,037 confirmation entries, each with up to 62 variables dda-0313: 4,720 wedding entries, each with up to 90 variables dda-0314: 13,217 burial entries, each with up to 84 variables dda-0645: 23,953 census entries, each with up to 37 variables these five files form the basic "raw data" of the project however, after this initial collection of raw data, a tremendous effort was devoted to data correction and family reconstitution. correction involved, for example, the deletion of double entries (e.g. the baptism of the same child in two parishes in the church records, or ^[figures and tables are collected together at the end of the article. ed. note] 'see jens ludvig warner. "datamaterialer med komplekse strukturer", in dd.^-nvt 25:64-70, 1983. fall/winter 1987 6 iassist quarterly enumeration of the same person in two households in the census records). those with experience in family reconstitution projects know the amount of preparatory work to be done before one can produce a centralized file. therefore, i shall go into some detail concerning the methodology applied in this particular family reconstitution project. how did we derive, from the above listed five raw data files within which all the events were sequentially numbered, the so-called "centralized file" containg the best possible reconstitution of famihes? 5. family reconstitution: manual and/or automatic procedures within the discipline of historical demography, there has been quite a long tradition of two "schools", one advocating automatic (i.e. computer based) family reconstitution procedures, the other maintaining that a lot of human thinking is necessary in order to get as close to the ideal of a complete reconstitution as possible. in the case of the data from greenland, a mixed automatic and manual process was used. with the raw data (or basic files) as the point of depanure, two tasks had to be completed: that of identifying the events (e.g. a baptism, confirmation, wedding, burial, or membership in a certain household in a specified census enumeration) as attributes, or descriptors, characterizing the "central persons" (or "actors") when the event took place, and that of establishing "pointers" among all the central persons based on their family relationships. who are the central persons (cps) in this project" in the event of a baptism or a confimiaiion, the baptized/confirmed person and his/he.'parents are the cps; in the event of a wedding, bride and groom as well as parents of both of the married persons are cps; when the event is a burial, the buried person and his/her parents as well as a spouse and children are cps. finally, in census enumerations, each person listed in each census is a cp. by machine generation, the basic files were thus expanded into a "theoretical centralized file" with more than 200,000 "central person in one event"-combinations. we shall call these central person-event records (cpe). a sequential nimiber was allocated to each cpe. the cpe-file was sorted on names; some auxilliary lists sorted on districts and other criteria were produced by jens wagner whenever kirsten caning needed and requested such lists during the work. the sorted cpe-file(s), written out as long paper listings, formed the basis for the next major state: the process of manually inserting personal identification nimibers (pin), (see fiqiue 2). even though the sorting based on names did help a lot, the task of numbering more than 20,000 individual persons (as it turned out later) on more than 200,000 cpe-records was a heavy consumer of both time and human memory! and now we may return to the question whether the machine or the manual reconstitution is "better", ll goes without saying that the choice of method depends heavily on the nature and quality of the raw data; so we shall talk about advantages and disadvantages, (see fiqure 3). .automatic family reconstitution has the advantage of being relatively quick and being well documented. however, in many cases a lot of persons are left "split" in two or more persons. for example, figures 1 and 2 (1901-census) compared to the extensive figure 3 (an excerpt from the cpe-file after all cps have been numbered) clearly demonstrates a number of problems with machine reconstitution: nikolaj jens andreas lange was found under many different names, partly due fall/winter 1987 iassist quarterly 7 to the fact that the first names may be written in any order (some of them may be missing), partly because of different spellings (some of which are not caught in the normalization process, e.g. nik(olaj) vs. nic(olaj)). figures 2 and 3 show that kirsten caning preferred to work with "normalized names" in order to minimize the spelling problem. but even with that precaution, the decision was made to perform the reconstitution manually. the reconstitution process was carried out by entering pin-numbers on all relevant locations in the cpe-file. during that process, a number of errors were detected and corrected this is probably one of the main advantages of manual reconstitution. after some iterations of pin-ntimbering, the theoretical cpe-file was reduced from more than 200,000 records to approximately 131,000 cpe-records. 6. the documentation problem in manual family reconstitution who "disappeared" from the theoretical cpe-file? approximately 70,000 entries vanished from the "theoretical" to the "reduced" version of the cpe-file. the records of type "own baptism" were reduced from 17248 to 16312; "own confirmations" were reduced from 10037 to 9169; "own marriage" from 9442 to 8388; and "own btmal" from 13217 to 11678. because some census registrations had been enoneously omitted from the first data entr)process, the number of census registrations increased from the original 23953 to 24076. (the first number is the size of file given in section 4 for raw data files). theoretically, one might expect that evenperson involved in a baptism, confirmation, marriage, or burial would have two parents. however, these cps have, in many instances. not been identified in the source records. therefore the number of cpe-records in which parents were registered in each of the four events was reduced from the "theoretical" numbers 34496, 20074, 18884 and 26434, to 31338, 13538, 4902 and 9181 respectively. predictably, it was in the events in which the "main actors" were adults (marriages and btirials) that the information on parents was most commonly missing. there is a major difference between the theoretical expectation that there was a spouse for each person buried (i.e. 13217 spouses) to the actual 2328 cases in which a spouse was actually identified. this is due, to a great extent, to the fact that many burials were of children (high infant mortality rate), and also in part to the death of immarried adults; also some of the missing spouses were due to lack of registration. finally, theoretically, it was expected that there would be information on at least one child for each buried person (i.e. an expectancy of 13217 persons); however, it turned out that this information was available in the source records only in 56 cases. does this mean, now, that we have lost the information in those of the cpe-records that were not merely theoretical, but in which the information was not given in the sotirce records? it does not! if a person was buried who was married at the time s/he died without an indication thereof in the burial record, we have data about the marriage from a difterent source. similarly, if one or both parents of a baptized person were not mentioned in the baptism file, we still may know of them from other sotirces, e.g. from a relevant census. this leads us on to the "record linkage", which in this project was done in an ad hoc data system. fall/winter 1987 iassist quarterly 7. finalization of the family reconstruction in a data base environment during 1984 and 1985, several mld/astrid' data bases were designed and tested on subsets of the data from the greenland projecl even though a number of these designs were acceptable from the substantive point of view, they had to be rejected because they would consume too much computer power when loaded with the full data base. ' with further computations and tests, a data base design was finalized by the end of 1984' and implemented with all the data from the project the basic design of the data base is that one part contains all persons and another pan contains all the events. each person is identified by the pin-number preceded by m (males) or k (females). there is no need for a lot of pointers because the searching is based on key fields in persons and events respectively. i shall not go into the technical details of the data base in this paper. however, 1 shall demonstrate the output from the data base using the two heads of household that we have followed in figures 1-3 above, (refer to fiqure 4). as can be seen from figure 4, not much has been done to present the output in an easily understandable way; thus far, only people 'mld/astrid is a data base language and system, defined by j0rgen grosb0l at dda, described in his manual: mld data bases and the astrid language . chensetddxt^sr" 'the basic design was described in icarsten boye rasmussen, "mld/astrjd database for gr0n]ands befolkningshistorie". working paper a460-kb. dda 1984-07-31. 'jens wagner, "mld/astrid database for grpnlands befolkningshihistorie design uden referencer", working paper a460-jw. dda, 10. december 1984. actively involved in the project have used the system, and they know what the encryptions mean. if the data base system were to be used by others outside the project, more text would have to clarify the output using this data base, the principal investigator is presently finalizing the family reconstitution task. also, special purpose functions have been designed for demonstration purposes; for example, it is quite easy to establish pointers to allows a user to move up and down along genealogical lines, drawing the complete pedigree of the person under analysis. it should be noted here that several features that are specific to the inuit culture are refiected in the data. people may be referred to their biological family, to an extended family or household, or to the particular house in which they lived. it was quite usual for the eskimo to live in so called longhouses.several households lived in the same longhouse during the winter (when censuses were taken); during the summer, they might move away from the house, living in tents or similar dwellings. the following year, the family might move to a different longhouse at the same or at a different location; the houses were not owned by anyone or, rather, they were owned by whoever happened to be occupying them. this tradition of living together in longhouses included a social security aspect; the data show that this tradition was slowly abandoned in favour of nuclear family houses diuing the 19th century.'" with the huge mmiber of cross-identifications in the data, it is possible to follow individuals or families (biological or extended) for generations. alternately, it is possible to examine a single location, and to describe how life changed in that particular location over the '"described in kirsten caning, "era bofasllesskaber til kasmefamihe", beretning fra carlsberfondet k0benhavn 1986. fall/winter j987 iassist quarterly — 9 years. the latter strategy was applied by the principal investigator in an article on sermermiui a small place with only two houses with 20 and 12 inhabitants respectively (3 households per house). ^' 8. access to the datasets for secondarj analysis despite the fact that the primar>investigator has not yet finalized the troublesome reconstituition process, it will be possible for secondar>analysts to have access to the data, with the consent of kirsten caning. all requests should be directed to the dda. needless to say, it is not quite as simple to address a user requesting this type of data as it is to send out a survey file. the user must define a priori the subsets or the formats s/he can handle. further, it is the nature of historical-demographic data of the txpe we have been discussing that their quality' improves over the years; during analysis, new findings concerning data relationships can be added. such changes are now reflected in the data base version of the data, but not in the raw data files mentioned in section 4. consequently, disseminable versions of the data should be extracted from the data base according to the specifications of the individual user. with the data stored in a data base management system, it is possible to generate rectangular files that can be analyzed with standard statistical software. however, to get the full personal histor. description capabilities, a data base management system environment would be needed by the user.n "tinna mobjerg and kirsten caning, "sermermiui in the middle of the nineteenth centurv", arctic anthropologv vol. 23{l-2):ntm, 1986. fall/winter 1987 10 iassist quarterly figure 1 .= ^ c>vl>oiwll< «« 53 ci suniidcii. ml. ill ju^uiy fin] jttxa^ fnyiaii^ jnii/if 72% figure 1: xerox copy of one of the source materials, viz. the census list from october 1901, with information from saraqaq in ritenbek. the upper household is that of nikolaj jens andreas lange, living with his wife and 4 unmarried children aged between 21 and 30. the lower household is an older son, lars jonas lange, with his wife and 4 children. hl!"^iu''^"^'^ u^<^p;"„du«.d from an illu'^tialjon in kirsicn bimhcth canine "om den eraniand^kebelolknmgs hisione". printed in l-orsknine i gipniand 1/82 . p. .s. ^' gmniandske fall/winter 1987 iassist quarterly u figure 2 134 71 9 118«223528 4 90337 11 11 99 3028nnrf1<;pfh .ihn « k'e 99'>1 7991 1 09139099 134925 5001423528 . 90338 11 11 •59 30" 'in pf« <;<;f'| rair < ar 90' 9991 4901 1 99909936104 35623528 '. 90339 11 11 99 301 »m pf/\>;<;f'.' pav : m r t f 9990"9911 9999o<" s89'. 9 jww' \ mi'i u u ?? jo] «nnrfa?;sfn ^o1 «n^r';^ssf•.' ja-: s-'r jw| ole :arl lr ')99735032 158725 4073223528 4110223 13 6 •)9 in?..a>ir,f -^irt per" 99964993? isbh'.is 6239423528 4 11022', 13 6 99 1c2-.af:rf gfrt flis aii; 99930993 62768 1292423528 4110225 1 5 6 99 lollai.'fe ihn 1 a c a l' t "992801 -^ 6302 3 129342352h '. 110226 1 3 h 79 im '.a'jgf ole j ms 9 7923013 158661 610212352m -.11022 7 13 6 37 10?'.atir,!; per', chrl a'je "9921993 62962 1294023528 4 110228 14 6 99 1 01 1 "'itf 1 ar.^ jotj ^993401 3, 1 5h62 4 4124823528 4110229 14 a -p? 1 o->i ^•^^,<^ a "i j fit: ci.ls ihm 0995599^; 630/"=! 1290323528 '. 1102 30 14 ^, 0> lolla-itf pet i-hr ha -is fari 0999^,99 51 62511 1288623528 4110231 14 ^ on 1 01 1 nf:rf fli p a 5 1 r p' 1 '. 039050931 1590!h 6238623528 4110232 4110253 14 /, 03 1 > l a n r r eli': r-rta 0990599^1 6117523528 14 'i -j 9 102'. a 10'^ larl sof oou'. 9 '90ir95i 1 6 70i'. 604742352" '.11023'. 15 3 t> 1 02:iiro'. ". i rr.n a-e let nort o99/,>,99i/. 52n'.4 1o8052352f 41102 3 s 1 5 ot 1 01 iinrh jac --av n.m" 0992 502 11 <: 16701 s 6368723578 411023'. 1 5 y 19 ^rp'^\^'^\^\<:^•\ «r,a :.'c i|l9091 ^9011 os 70130 140272352o '.110237 l^ 7 19 ni ••ar"ia<:';r'i mini «'i;l rl'c994 7':il2fo 1'.555 j 6239723528 4110238 16 7 99 1 ohmthia s-.rri h'^nu sof 'jorjt coo609j3-9.., 165473 6104523528 41 1 0250 1 6 7 09 1 o"-m 1:1 > "^^f;! ffc ai'lo <1r r9q1 7':9''1 coc. 7 014 7 140322352'^ '. 11024 l^ 7 9 9 1 "1 r-at" [ a<;<-fr' mie'. c s f i s '\ ' "991 s99''1 9^9 165 34 5 41 2292352t '.11024 1 16 7 09 lo^mat'li a«:';r'l aok, c.js '.ott^ c991 3"931 ^9"' 70137 1412223528 411024 2 1 f' 59 ni -1 rnii'-.s'--i ." i c 4 '. 1 r r a < ? '99109931 c09'; 70039 14 109 2 3 52 f^ 411 0245 1 1" / t9 101 >iathi 1 r'^fi 1 i\r;<; r h r car'. 09v9s "951 so"'r0463 2074523528 '.11024'. 17 1 '^ 10 30) nn<;9 \r\i pft is'.k too i7"s501 52'-9"9 1 7503n 6242^23528 4 110245 1 7 1 n 99 3n71.-0-.n„(;|. ^of 99os142329809 802^2 2"76623s2r 411 024 it 1 7 1 " 9 9 501 bft^nic ii .' ^^"; f :! 1 a • " 191 s9"31 9' 99 174951 4 2 3252 3s2'h '. 11024 7 1 7 1 n to 50"'rn':t 1:11 lum jul ri'.'.-^90 -90931 -^99' '. 186589 65284235'f^ '. 11024'' 1 7 1 09 ?o?zrf.o kali t ii-t irr, o-'927->9 :.10'-2r2 552?3 1143323528 411 02 4'^ 1 7 1 "^ o') 501 .if'.'pr-i .inf;f r:ici o'.r ' .<9 22" 1 1 1 ''99"9 17486^ 4221925528 '.11025'^ 1 7 i "^ 09 50:>ro<;-^ac " ri'r t a'lt f'. i:; o0935991 4-,'1 502 m035'. 2077123528 '.110251 17 1 c 09 501 ro';" ac ii .lom pi^t jaf 9 >91 390-1 99909 80171 3n7<. (.23528 -.110252 17 1 ri 99 301 k'nr.nac ma'i': pet "ic.nvi hi-i 31 .-909c 801 70 2074323528 11 3412352" 4110^55 1 7 1 ^ 9 9 501 posoat r a fj 1 m'^l fr u .0050051 9?0vs 5 5 1 n 4 1102 5'. 1" " '9 50] i c>..;e.| jhi;" 1a 7 l f 1 '/ 15 23t. 6106123528 '.110255 ri f •>9 202.1 e"i';e'' "".t i. c 1 s t l i • 9 '94799129"9'-/' 551 )4 11 '.3425528 4 1 1 0256 lb 99 201 .|r.|^.c-| ihf r )( jo oyv2 i'll 31 0997 r. 16'iy 6105<,?35pp 41102 5 7 18 9 09 20mf'1s'^m c aro ^r^t a"m •39. 1 091,^1 o',99 152256 ''117h2352" '.1102 5" 18 q 07 2c'jfn';?\' ^ft ••us t 1 ^ k ^:)'^ 59'j51 499' 55721 1150423528 4 1102 5 9 1m f09 201 i-^nsfs-p. ole ci;h r;ioi 1 o9.--'• 1 '59" ir ' 3597" 32"23528 41 1 n2'^5 1 9 ^ 9 9 201 allhr'^a';-.'-' .ifn: ,it: ott 030,. ilip1 34 761 4033223528 41102 64 i-? 3 99 2'^2a99 i'fa t'jf-i jup .ou o'o/, a991?r 114 50? ml 74 2 3 528 '. 1102',5 19 3 2r7to-i i a r.;!:" ;ara lj-:}. (iv;s'-991(. 54052 11 3572552! 4 1 1 02 ''.'•' 2" 5 nr 101 .ifnsf-i ha os 9;o4 501 315 20-^5 '.074923528 -. 1 1 02 4 7 ', 09 1 o-',|f>|,>.c., 1 a 11;: 'it ap ' 9 >; .'. s ' : " "> 55030 1150323';?t '. 1 1 02 f." ?r ; t> 1 -11 j f'l ''-. .1-7^; je'.'s "•••s jlr ..m '.co figure 2: the same persons as those listed in figure 1, now listed from the computer. to the left of the "normalized" names field, lots of numbers of identification appear, i.a. the pin-number appearing in columns 7-12. our head of household nikolaj jens andreas lange has pin-code no. 12914 in the ritenbenk district (code 4 in col. 19), the 1901-census (code 28 in cols. 16-17) from which the information was taken. he has that number in all other datasets as well, which we shall see later. fall/winter 1987 12 — iassist quarterly figure 3 ari 000»-000*-00 b t*.b u.b aro o wo "-tb »fco *3o xo >»^ <»o oco-o o < o — o o ooooo' ^oo >0 >c wo (.k3 _jc c >c cro> «» 0<> — o w w»0 vyfy oo" irto' 0>9z(> tt> t oowo *>> o* ->d -ic -lo ->o -»o ^c zo zo z-o ;co ux> uj^ u/o" luo luoluo' u/o o r^ qd tt x 03 o o o o -o c o o o 0do>o o c o »«^ o o c oooooo^o o ooooo*-oooooooooocoocrcco k» o o o • < ki •cd cc -o o oo o o o d ^ o o k •oo •o »n muodo^^o •^.»' oooo p^o ,-0 oo «0 *no — c»o-.-^ ofo-«o^ o^ o-^"'^ c;0| oon.o oo pwo oo r^o *>o ^'0 oo oo oo oo oo o_j oo ^o oo o^ -oo co -oo cc oo oo oo oo oc -co fall/winter 1987 lassist quarterly 13 figure 3 ...cont «o o ujo> ot> oo uooo-*> lo io zolo-«» «» «o ^o ^o _iojojo _jo jo5 s (jo wo (3 0> ujo. ii^o tr< « •-0 •~o rvio s s" s ceo s* s' (do*" o o o wffo rvoio 0.0 coo mo oo -oo so fall/winter 1987 14 iassist quarterly figure 3 ...cont su, fall/winler 1987 iassist quarterly figure 3 ...cont o^ __ _ _^ _^ _ _ _ _ _ vjo •-•<> «0» dooo< »-opo at *-»^ mo»^ ercft-w«*z«~ u ujoo xo-o «0-0 a ao'ki t.rf^*n a.o^^ t*,_ ^_ ^^ «._ o cc co> <» _)»m»ru e u. w o o ujoo. ^„ „ w — — — _ . . _ >. ujoo> vjo •-•<> «0» dooo< »-opo ac o-ki -"— zrtit-»<•z«»-» ujoo i*-0 qto o-o «o _j _( a ao'kl t-*?"*^ -^ c«vf4^ •»• < do o'o ^o <^o o wo «o z x t< z z»^»-»ni— »o too» •"•o < uj oc a. < xouo^o o o a.o> luo (aoco' -looo>^» ^ ooo « < < « < ujc u/o' ce o om ujouo' ujouo' u o uj ^ luzc zo' zozo^ ic z <0. «©. <0<0' < o_jo_jo_io_jojo oo' 00> oo' uo oouooo' uoz z z z z zozo' zoloic zr< o 8<^ 0> oo. ooi (d o(c cd 09 (t o£ ,t3 o 0> o o .k^o-o 0> o"^ ^f^ '+0' in-iof-^ory-»0c->0cdooo "4-^ t>u 'kio-^ tn t/\ o o o mo-o-o-t-o-^o^cocoz ^^ « cdz^>0 ••fvj l\l r•omr-inlm r\j ti"rm injunr.i l\n/^<\l oiu*/i 0_l 0>co'o o d o o o co o-o c—o o—lo o-^c c~i_i 00_j c_| o > d«o o d o o o 0> c jo0_j0 o _'0 d joofn*— oftjk o/vj -0 tnt-o i/%»-c w^^-b •nz'^j kioj *no(» k%o l^jo>' (m rvt rsj ini ^ ru«o r\»i«o i\««o "»• zo>o ko*^ «»^ ^(m^ uoo 0.0*0 uto'o so'o o*^ o'*^ of^ < xo'o ** o.x* (>>_l*^ 0_jo fv**" rviuio" (vjiu* "r; a jsi> ; o t! -a -o c " o s-oci 5'"af^ ^c.^= 5-i5p-ego. .^cu = „ 5"".. z £1 .i£ ,= "^ " "^s x^ 22 w 5 c « 9 ;£ 5 .2 e ,, c " ^ -d i5 "-*---! f3«° |s-='= 3£-o§« 8l"s ^ iz it i^^ e i iii |s:il" i§s.gal -: srs s-ij :2~; ? ; e-= -o ^ c .^^ « "s •_^ " £• = -j ?f^ < '^ s -^ -n -a e — =5 ~ " 2 t coc-^gs^ noo.sm ^5co -o o c fall/winter j 987 iassist quarterly 17 figure 4 •person: h01291x barn ic062355 «ni29 58 barn 'n0^29l7 m6128?1 1(062*29 1(062398 icq40998 hn128( mo12906 k06238* rg6?j97 xoblfo-* it061021 «0129; gift k0*n732 k040'12 it06?582 stah9ata 1825 jens nik andr lange 3 9 1905 816 transoata 8 01090190 5081601500 9 3467 19q90701 02 506 b063741f bq72u1871n722tq5q2 6063831 857n?1qln5q4 dogo**] 6(580701852091820352 t) 0999018 600 30720502 n10094l( ., , .^., d101801 8670*0210502 d 1 03361871 1 22 520002 d1 0* 931 85605u 1 072 t 1(009951894081520500 < 026921 889072 51 500 1(0310*1880072910502 k03187187a05302n502 1(0*5011842032801000 k051 391 866031 *20500 i^8im]§^i8^??i8i8s mmmimmm mmmimmm s70n*91 83*000003001 5 7007018 50000001722 t408001 901 000001 502 6106141858071120502 01062 718 61122920502 61 071 81 86909n51 q5q2 d16835187803o310o02 0658018 7*0104 10502 » 1 §' !« j g ^§043020502 d173721 86on317i0502 koos751 892031 31 0662 k009801 89203 1 320662 1(009951894081520500 1(031871874053020502 miiwmmmm s700491 83*000003001 1705081854000003001 tz05351650o000o1722 t708581 87000000150? vg21941 852012801000 •person: «01294n barn 1(040732 h01291* barn bh012903 m012886 h012916 k062386 k061175 gift k041248 k061998 stanoata 1867 3 1lars jon lange 3 11 . ,„.„„ ^„„ ,,.„„, , thansdata 01590189 5062 310012 01599189 7010710012 601 t-'o' 905091 71 001 2 6020551898111320012 0207119 00111120012 0101801867040201000 k019561911062510502 k02oo71 9200 5281 0502 '<021 151 91 5081 720502 k0212?1916o2252p5n2 1(055491909080110502 k09 559 1881 080601 000 tioanii 901000001012 t;'o858i87ooqogo5ggi vqq4n7i 894oeg50i on v007401916040910502 v 0076919 21 3271 0502 voofla 1 191602141 001 6 v008981 920042110502 v009091 9 2208221 0502 v009361 927022220502 v033851930101010502 figure 4: the information in the person archive on the same two heads of household as reflected in earlier figures. the person is identified by the pin-code preceded by m (males) or k (females); the second line gives the parents, third (and forth if necessary) the children. under the heading gift (= married) the spouses are identified. finally, under the heading transdata there are references to all events in which the individual was considered a cp. the initial letters identifying events are: b for burials, d for baptisms, k for confirmations, s for censuses with the extended family as the unit, t for censuses with a biological family as the unit, and v for the weddings. fail/winter 1987 28 iassist quarterly 2013 iassist quarterly iassist quarterlyiassist quarterly abstract the work and life of sue a. dodd had influence in its own right and her work was adapted and incorporated by others just as her work was influenced by others and part of a general evolution of social science metadata. from her focus on the catalogue description of machine-readable data that made users able to reference and identify data files as a research source, the description of social science data files gained further momentum. this paper centers on the fundamentals of social science data and their relation to metadata. there are levels of metadata in typical social science where the study, the variables of the study and the codes of the variables define a hierarchy. for each level there are many potential descriptive items that can be part of the full metadata. the work was initiated in the us but there was also work carried out in europe through the described period, mostly centered around 1975-1995. all of this can be considered the foundation of social science data metadata description that later evolved to become the work carried out within the data documentation initiative (ddi). keywords: ddi, data documentation initiative, metadata, social science, study description, codebook. introduction sue a. dodd identified the vacuum of library catalogue description for the new and growing area of machinereadable data (dodd, 1979), and provided guidance for using a standard bibliographic format to fill this vacuum (dodd, 1982). in following years she continued to communicate, discuss and elaborate upon the guidance. some of the other papers in this collection in the iassist quarterly (iq) will bring more focus to the writings of sue dodd and some papers will elaborate on the work carried out within the data documentation initiative (ddi). this paper sees the influence of sue dodd as her work was adapted and incorporated, highlighting some of the european work on the study description for social science data during the period 1975 to 1995. the term ‘machine-readable data’ for the materials being described was later to be renamed ‘computer files’ which comprises more materials than the data file that is the main object of this paper. this paper will bring descriptions and discussions of why and how the formats, processes, and technology developed in collaboration and hopefully demonstrate how this totality articulated the need for a documentation standard and formed the basis for a solution of elements for the data documentation initiative. for a tour of the developments in data archiving combined with the technical and political – as well as personal developments i’ll recommend ‘the decades of my life’ by judith rowe (1999). the digital age: useful data are useless without documentation the digital age signifies that everything is stored as numbers. obviously when only a number is communicated – ‘42’ is my favorite example it cannot alone carry any meaning for the recipient. a number has to be wrapped in explanation to convey any meaning. we know we have been fooled and we find the nonsense amusing when the computer deep thought after seventy-five-thousand human generations of calculating produces the answer ‘fortytwo’ to ‘the great question of life, the universe and everything’ (adams, 1986, p. 128). social science metadata and the foundations of the ddi by karsten boye rasmussen1 iassist quarterly 2013 29 iassist quarterly 29 iassist quarterly 2013 iassist quarterly data, information, knowledge the machine-readable data file consists of nearly endless series of numbers. let us visualize the data file as a database table (survey file of a questionnaire) consisting of many attributes (variables as columns) concerning entities (individuals as rows). many information scientists have attacked the problem of having more precise concepts for categories or levels of information. in everyday life we might use the term ‘‘‘data’ intermingled with the term ‘information’ without much bother. i must admit that i personally don’t have any problem with the less rigid use of the concepts. however, there are insights to be gained when entering a more meticulous definition of the terms. when working within systems analysis in britain the researcher peter checkland proposed some useful distinctions for his methodology of soft systems development. data are viewed as an unordered, formless, disorganized pool of facts. (checkland actually used the term ‘cloud’ that now carries the meaning of organized access to safe storage always available.) ‘information’ is facts situated in context resulting in ‘meaningful facts’. in relation to the data file the documentation is delivering the context. ‘knowledge’ is ‘larger, longer living structures of meaningful facts’. in our analogy this relates to the use of the data file mostly exemplified as the relationship between variables. in getting to knowledge the first issue is to select the individuals and the variables for our data collection. thus we select the data relevant to us. checkland proposed the term ‘capta’ for the selected data (checkland and holwell, 1998, p. 90) to be distinguished from the disorganized pool of data. the context in the form of the description of the selection process concerning both the selection of entities and the selection of attributes is a necessity for creation of ‘information’ from ‘capta’. these descriptive entries are the metadata or data description. i did warn you that i might not be rigid and consistent in my use of these concepts. the term ‘capta’ is in my view a way to understand what we normally call ‘data’. when we have data we require data documentation in order to produce information and knowledge. multiple benefits of data documentation during the 1960s and 1970s the research data file became a regular resource available to other users for secondary analysis and the benefits of data sharing and data documentation was addressed consistently in the 1980s (rasmussen and grant, 2007, p. 60). a conference in 1979 resulted later in a comprehensive report on ‘sharing research data’ supported by the national research council (usa) with extensive discussions and papers and a leading chapter of ‘issues and recommendations’. the number one recommendation is ‘sharing data should be a regular practice’ (fienberg et al., 1985, p. 25). in the journal american sociological review the benefits of data sharing published in the report was summarized as: ‘reinforcement of open scientific inquiry; the verification, refutation, or refinement of original results; the promotion of new research through existing data; encouragement of more appropriate use of empirical data in policy formulation and evaluation; improvement of measurement and data collection methods; development of theoretical knowledge and knowledge of analytical techniques; encouragement of multiple perspectives; provision of resources for training in research; and protection against faulty data’ (hauser, 1987). the benefits of sharing of machine-readable data are parallel to the benefit of having libraries sharing human-readable material. making the sharing possible implies the benefits of metadata describing the data file. the value of sharing data can be accredited to several dimensions of arguments: actors sharing data primarily implies the use of the data as secondary data when data are reused for other than the intended purposes by other researchers or by students for educational purposes. data archives have also often experienced the original investigator(s) among the requestors for their own dataset because the archives had not only preserved but also value-added to the dataset by elaborated and accessible metadata. resources collecting data is a money intensive action. research data collection is often financed by public money. sharing the data is the most cost-efficient way to carry out science. furthermore, some retrospect research initiatives are only possible to realize through secondary analysis often of an extensive kind where several earlier data sources are being used in combination to produce a more accurate account. naturally there are also costs involved in sharing data and producing useful metadata. it is possible to argue that not all collected data are worth the extra cost. however, when research projects are financed in competition (e.g. from types of science funds) it would be counter-intuitive if the board would not regard the future collected data as sufficiently valuable. the national science foundation (usa), the economic and social research council (uk) and the danish ssf (social science foundation) and probably many more like these funding agencies all had clauses in the contracts for archiving and documenting research data for reuse when i investigated this in the late 1990s (rasmussen, 2000, p. 169). controls a popular issue of research data being publicly available is an issue of being able to control that science is not infected with fraud. however, it is also mentioned in the citation above that the public through data access can control administration and evaluation of policies. furthermore the sharing of data also supports ‘multiple perspectives’. this is regarded as fundamentals of having a democratic and free science. the short hauser entry above reminds us that there is a distinction between survey data being ‘public’ and ‘usable’. this is a distinction depending upon metadata. hauser recommended that journals (e.g. the asa) would be keeping the tabulations related to the published papers in the journals (this was more than 25 years ago and the recommendation was practical and proposed the technical solution of storage on floppy disks). the recommendation also brought an attention for a recommended format. however, the discussion on what to archive and in which formats was in the 1980s already a continuing discussion and system decisions were implemented in data archives around the world and will be addressed later in this paper. the general improvement and development of research can also be placed under the dimension of control. metadata description of a research dataset will act as an evaluation of the study, e.g. a survey is described in terms of the population, the selection 30 iassist quarterly 2013 iassist quarterly process, the nonresponse, the response rate etc. these items will act as a checklist for new and not so experienced researchers as well as demand consistency and thus easier access to the data. archives and archivist collaboration some twenty years before the comprehensive reports on sharing of research data (fienberg et al., 1985) and of social science data (sieber, 1990) the sharing of data had already been implemented in institutionalized form by the establishment of the icpr (inter-university consortium for political research later icpsr, inter-university consortium for political and social research at the university of michigan). before that the roper center had been established as an archive for gallup and other commercial public opinion polls since shortly after the end of world war ii. with the eye on comparison of national statistics there were three (unesco-) conferences in the early 1960s also addressing data archives (rowe, 1999). the concept of data archiving was also institutionalized in europe by the creation of several national social science research data archives from the mid 1960s and onward. in germany the zentralarchiv für empirische sozialforschung (later included in gesis (leibniz institute for the social sciences)), in uk the uk data archive, in the netherlands the steinmetz archive (later included in dans (data archiving and networked services)), in norway the nsd, and in denmark the danish data archives. many data archives incorporating data files from research in many other european countries followed quickly after. when the iassist2 was founded there were a sufficient number of individuals within the profession of data librarians, data archivists and other advocates, researchers and technicians who were sharing machine-readable data to form an organization that has lived and has had influence for a span of time that was in older days a lifetime. the acronym iassist was created in advance of the name: international association for social science information services and technology (geda, 2006). both the acronym and the name continue to capture very well the focus of the association. the sharing of knowledge was institutionalized internationally through the iassist conferences and the included workshops as well as through workshops and the meetings of official representatives (ors) under the auspices of the icpsr. soon after the realization of iassist, the european cessda (consortium of european social science data archives) was formed as an umbrella organization for the national european data archives. issues were trans-border sharing of data and privacy in the different countries. cessda was also successful in knowledge transfer between individuals through themed seminars many of which addressed the issue of metadata and discussions on the use of different standards. metadata and levels of documentation the fundamental documentation of a data file as a whole is the identification of the data file as a research source. when we talk about the ‘american national election study, 2004: panel study’ others will be able to locate the study3 (and they will find a shorter form of identification as ‘icpsr 4293’). similar to citations from literary sources there should be a way to accurately identify the research data source. the documentation of the data source makes it possible for secondary users to give positive credit to the people behind the data source. among the most fundamental aspects of scientific research is the possibility of inter-subjectivity that is the closest thing to the unattainable objectivity. documentation has the potential to introduce discussions on the methods used and the research decisions. the researcher describes how the data were created and makes critical evaluation possible. provided the documentation reveals valid methods, the replication of the procedures described in the documentation should ideally lead another investigator to the same result. data equals documentation documentation equals data metadata for some of us the rectangular data file was and is the ‘normal form’ of social science data. the term ‘normal form’ brings us to relational databases and that is no coincidence. all kinds of complicated database structures are possible with the methodology of relational databases relating rectangular tables. others may be interested in more exotic examples of data for social scientific analysis, such as artifacts like images and sound bites. whatever type of data is to be analyzed using a systematic method, you have to define what your specific interest is and define and select your ‘capta’. in other words, the systematic and comprehensive documentation of a collection of machine-readable files can become data for a researcher investigating the collection. for instance a research objective could be to compare surveys carried out at different periods of time or at different geographical locations. the description of data is called metadata as it is ‘data about data’. we should note that metadata are not only ‘about’. metadata are also genuine data that can be analyzed. the levels of data documentation the primary user of a rectangular social science data file needs information on the variables (columns) as well as explanations of the codes as exemplified in the information: this codebook with information at the variable level and the code level is elementary yet helpful for the analyst who already is familiar with the overall background for the data file. however, for the secondary analyst ‘(i)nformation about variables is useless unless the population, sample, and sampling procedures are described’ (blank and rasmussen, 2004, p. 307). for this reason, many more precise items were introduced for study level description early in the history of data archiving and a standard for the study level description was agreed upon (nielsen, 1974 and 1975). it was discussed, refined and presented at several meetings, workshops, and conferences in the 1970s. however, although there was this continued presence around the ‘standard study description’ the standard was primarily the foundation for local implementations and never achieved the status of an international ‘de facto standard’. more important was the international agreement concerning the items described in the standard. several of these items were also implemented in other archives – in other local formats and systems – and the items were later included in the development of the ddi-standard. they are listed below: v6! sex  of  respondent’! 1.  ‘male’    !  2.  ‘female’ figure  1.  codebook  documentation  of  a  variable  with  codes   (a  minimum  example).   iassist quarterly 2013 31 iassist quarterly the items of the study description were included as material for the ddi development as explained further in the green and humphrey paper in this issue of the iq. machine-actionable metadata there was a great effect when documentation of machinereadable data had the positive experience caused by ‘taking its own medicine.’ naturally the decision on which items to include into the documentation was a very important and first decision. however, when that was settled and the human could read the documentation there was a revolution in having a well-formed documentation not as typed pages but in machine-readable form. the revolution happens when the machine-readable documentation becomes machine-actionable. when metadata collections were formed several data archives were looking into retrieval systems for supporting secondary researchers in their quest for finding appropriate data. in europe the german archive zentralarchiv in cologne (now gesis) was in the forefront of building retrieval systems. another earlier use of rigidly structured metadata was the translation of osiris codebooks into the more reduced control language used by other statistical packages like spss and sas. as system files were dependent upon the configuration platform of both machine and operating system the character-based solution of exchanging old-fashioned card images in the form of lines of text was the most general solution. furthermore, at the codebook level the machine-actionable documentation meant that a machine e.g. software on a computer would be able to read the documentation and be able to automatically interpret and perform calculations and controls on the data file. in addition, the machine-actionable documentation could, with great benefit, be created prior to the data file. the documentation could by a machine be the foundation for data collection, e.g. generating the screens and software for data-entry for telephone interviewing (cati, computer aided telephone interviewing). turning the complete documentation including the study level into machine-readable form meant that studies could be searched effectively and with the internet the accessibility to data files increased tremendously both regarding the number of studies as well as regarding the speed the users could access the selected information. having documentation in machine-readable form implies that there should be a defined format a standard. standards of data documentation the old joke about standards goes: ‘standards must be good since there are so many to choose from.’ in this section some of the actually used standards for documentation of social science research data before the ddi are briefly described. machine-readable cataloguing in short marc was developed at and institutionalized by the library of congress as a computerized ‘library card’ format to build a library catalogue that was machine-readable in contrast to the paper cards traditionally used in libraries. the latest marc format (marc 21) was finalized before the millennium. some can consider the format as old-fashioned and related to outdated technology. however, the format is still very much in existence in libraries all over the world. in regard to the legacy of sue dodd, the marc format for mrdf (machine-readable data files) stands central because her 1982 manual provided guidance for applying the marc/mrdf format to create standard catalog records for mrdf. her guidance based the catalog record upon file or study-level documentation. standards for documentation were discussed continuously throughout the 1975 to 1995 period. as mentioned earlier, there were during the 1970s international workshops in europe on the documentation at the study level. iassist accommodated for many years sessions on documentation with presentations, proposals, and discussions at the international yearly conferences. further information on the study level documentation standards standards including marc and ‘dublin core’ is found in the green and humphrey paper in this iq.. machine-readable documentation in use at the iassist conference in 1993 an action group for ‘codebook documentation of social science data’ was formed. in 1994 i carried out an investigation by mail questionnaire in order to obtain an overview of the amount and kind of documentation that existed at archives. the investigation also obtained preferences of documentation among the professionals. this was reported as a presentation at the iassist conference in 1995 as well as in the iq (rasmussen, 1995). the paper presented a snapshot of the situation now 20 years ago. the majority of the studies were then without any machine-readable documentation. the reported machine-readable datasets were from 19 answering archives. the distribution of datasets by type of machine-readable documentation is shown in table 1. the different machinereadable formats or level of standards are explained below demonstrating the development of metadata for social science. 001-­‐ general  informa/on 
 documenta/on  level,  subject  cluster,  keywords 101-­‐ iden/fica/on  and  acknowledgements   
 bibliographic  reference,  archive,  primary  inves/gator(s)   and  other  references 201-­‐ analysis  condi/ons 
 abstract,  kind  of  data,  data  sources,  type  of  unit,  number   of  units,  size  of  dataset,  /me  dimensions,  universe,   selec/on,  sampling,  data  collec/on  instruments,   weigh/ng 301-­‐ reuse  of  data 
 data  representa/on,  data  cleaning  and  controls,  access   condi/ons 401-­‐ references  to  relevant  publica/ons 
 primary  publica/ons,  secondary  publica/ons,  analysis   results,  references  to  other  studies  (data  files) 500-­‐ background  variables 
 personal,  (age,  gender,  ethnic  group  ...),  household   informa/on,  employment,  occupa/on,  income,   educa/on,  mass  media,  ... figure  2.  overview  of  a  selection  of  items  included  in  the  study   description  as  used  at  the  danish  data  archives  (summarized  from   rasmussen,  2000,  p.  287-­‐352,  and  rasmussen,  1981).   32 iassist quarterly 2013 iassist quarterly scanned images a doable and less time-consuming method of archiving and being able to distribute information was in the form of scanned images. the available documentation often existed in the form of a questionnaire and as sheets of paper that could be scanned and saved as images (de vries, 1992). as the processing often did not include ocr-processing this meant that the production did not deliver a searchable documentation nor was the information usable for input to the statistical packages so scanned images are considered the lowest form of machine-readable documentation. actually you can contest that they as images are readable by a machine. the information is not structured and an image is not data nor is it metadata. a considerable effort was demanded from the secondary user as information had to be keyed-in in order to analyze the data. sometimes the scanned images were additional to other formats and then the images were of convenience to the user and a safe storage solution to the archive. the investigation among individuals showed that most people would not be content with scanned images of a questionnaire for secondary use of a data file as they would likely prefer the structured information with the possibility of machine-action and the direct and accurate feed of the data file into a software package for analysis. text this documentation level implies that machine-readable documentation existed as unstructured (i.e. untagged) text. this could for instance be in the form of output from ocr or a questionnaire kept in wordperfect. the lack of structure implied that there was no easy solution for bringing the documentation into the analysis making it machine-actionable. however, the bulk of text could be searched just as we still often search within a text based file or in a collection of such files. dictionary the availability of a dictionary implied that the structure of the data file was reflected in the proofed dictionary and data could be loaded into standard packages like sas, spss or osiris. this brought down the time involved in setting-up a system for the secondary analysis as well as increasing the accuracy. the data quality could be greatly deteriorated by the hand-to-hand passing on of instructions. with the availability of a machine-actionable dictionary the secondary analyst would spend less time on controls and more on analysis. at this level see below the higher level ‘dictionary-plus’ only column information existed and often in a very restricted form. for instance variable labels would often only be able to carry 24 characters of documentation which people did not find sufficient. furthermore, this format did not include any information about the codes and categories. the user would for instance need scanned images in order to find explanation for the codes actually encountered in a variable. dictionary-plus for the category of ‘dict+’ the plus indicates that the documentation comprises category labels in osiris or in the form of ‘value labels’ in spss or ‘user formats’ in sas. often the documentation was stored in a system proprietary format and often the packages were forcing a ‘lock-in’ on the users so you could not for example directly analyze an spss system file with the sas package. the situation has improved but at that time even the change of version of system files within spss presented a problem with backward compatibility. having documentation embedded in system files also presented the archives with a load of migration tasks as information could be lost if files were left in old system formats just as they would be lost if they were left on old media. as rothenberg phrased it: ‘digital information lasts forever or five years, whichever comes first!’ (1995). dictionary plus codebook through the period in focus here (1975-1995) the sas and spss packages were the popular tools for analysis although they had no support for documentation at the study level apart from a filename and a short title for the study. the amount of study level documentation was very similar to the restricted documentation of variables and categories: a name or value and a label of limited length (grant, 1993; rasmussen, 1989). the osiris documentation format was early regarded outdated by being tied to a card-image format of 80-characters an inheritance from physical punched hollerith cards. when using physical cards placed in sequence it was very important to be able to reconstruct the sequence (with a counter-sorter) in case a stack of cards got dropped on the floor or was mixed into another stack. the osiris codebook layout was originally mostly a format for storing the electronic information for later printing a more nicely formed codebook. however, the osiris format was remarkable by being able to store unlimited amounts of text because a description of a variable could occupy multiple lines; this was also similar for comments, explanations, and code description. in osiris different types of documentation were identified through an alphabetical “tag” in the first field or column of each card. (see figure 3). remarkably, the osiris format included features for description at the study level (s-cards). osiris was developed over a period of time and i’m here referring to the osiris iii (icpsr, 1973). structure of the description at the study level was available as the format even included further distinctions on the study level for entries as title, original archive, etc. furthermore some ‘meta-metadata’ were possible as the general structural principles of the codebooklayout could be described within the standard itself using a meta explanation type card (e-cards). these features and other parts of the osiris format were extended at the danish data archives and the german zentralarchiv (rasmussen, 2000, p. 351). these were extensions for more elaborate description, data checking, and retrieval. the harsh backside was that the osiris system in itself in the analysis tasks – including tabulations totally ignored all available information apart from the limited dictionary information supplying variables with a number and a 24-character label plus some format information as well as information on missing data.  1.  scanning   505  2.  text   1,943  3.  dict   4,520  4.  dict+   3,656  5.  dict  +  codebook   2,003  total   12,627 table  1.  datasets  with  machine-­‐readable   documentation  (rasmussen,  1995).   iassist quarterly 2013 33 iassist quarterly although the osiris format was old-fashioned judged by the card-image layout, it was thus also very farsighted. the original osiris format and the extensions including some developed outside of icpsr had the capability to include a very high degree of the relevant documentation compared to the offerings by other statistical packages like spss and sas. the documentation types brought attention to the levels and structure of documentation and to the comprehensive items that were later developed into the ddi. the letter in the first column of the lines of osiris documentation was a very early markup of documentation through the ‘tagging’ by first-column letter. furthermore, the rigid card-image and fixed-columns osiris format was kept as the archival format, but it was generated from freely typed input with software generating the variable numbering and the required different indentions based on the tagging for the relevant cardtype (‘q’, ‘x’, ‘c’ etc. ) towards a standard some months after the formation of the iassist action group on codebooks a cessda seminar was held in gothenburg in august 1993 on ‘variable level documentation.’ the following year another cessda seminar on ‘networking and internet’ was held in grenoble. both of these workshops also had non-european participation, most significantly the participation by staff of the icpsr. the discussions not only collected the sum of identified elements that from the viewpoint of archives were considered important to include in a standard documentation package but also brought to attention the many functions that the documentation should support. they also introduced carefulness towards what could be termed independence. this independence was the guarantee that a standard could evolve and not be locked, as well as being available to all. functions of the documentation this paper has mentioned how documentation first of all delivered the printed study description and the codebook in a wellformatted human-readable form. this is believed to continue to be a relevant use though the human might read the information from another device than paper. another important function is that collections of documentation – especially in the form of well-structured computer files – are searchable. it has also been mentioned that documentation can deliver input for the validation of data previously collected and also control data being collected. lastly, the ultimate use of documentation is in the analysis and the presentations from statistical software of the documented data. but do users really need all the bells and whistles delivered by the ddi? when huge commercial companies like spss (now ibm spss) and sas deliver only a minimum of documentation facilities for a dataset should that not be taken as a sign of what the user community is interested in – and as a sign of what quality level of documentation the community is willing to pay for? osiris is still in existence and can be found as mircosiris for ms windows. however, this micro-version does not support the unlimited and elaborate osiris codebook format. mircosiris has only the minimum documentation facilities as described in figure 1 above. i believe that minimum documentation presents a one-eyed view with a narrow focus on machinery for analysis of your own data. the limited documentation will prove to be a problem for even the primary researcher if and when an older dataset is to be re-analyzed. naturally the problem will be even greater for the secondary analyst. in commercial settings they know within data management that elaborate and often very expensive documentation and management systems are necessary for gaining profit of the data warehouse. independent documentation when developing a new documentation format as set forth by the data documentation initiative it was considered very important that the standard be independent of commercial interests. public archives and university libraries would not be able to afford to tie themselves to a storage format that could imply a yearly user fee. independence was also the term used in connection to which systems should be allowed to analyze data described in the ddi-format so there should be no licensing. the ddi-format should also be independent of operating system platforms. it might be possible to obtain financial support for the development if a company could obtain rights for a proprietary format and system. however, data archives regularly service many users who have distinct preferences for this or that system. therefore the solution should be that the documentation format is open for use by all software developers or vendors. the archives already have great expertise in converting existing codebook/dictionary documentation to the reduced description used by sas and spss. when developers at archives would use the ddi-format the software for conversion to users’ preferred software would naturally follow. !! !! s study  level e meta  explana1on t dic1onary,  sub-­‐structured   including  missing  data  and  format q variable  descrip1on.  ques1on  in   ques1onnaires k con1nua1on   x explana1on c code  value  and  label,  sub-­‐ structured b for  grouping  of  more  cs  (higher   level  categories)   f frequencies,  agached  to  the  c   j temporary  comments g note  number m note  text figure  3.  the  original  card-­‐types  of  the   osiris-­‐codebook  (summarized  from   rasmussen,  2000,  p.  340-­‐352;  originally  in   icpsr,  1973).   34 iassist quarterly 2013 iassist quarterly along came the internet with a markup language the internet was early on seen as being of main interest to archives as a media for users searching for data and thus identifying relevant potential datasets. later the internet also became a media for direct deliverance of data. and later again the client could get thinner and direct analysis would be performed on the network servers. the internet was as such very promising for much faster and easier identification and access to data sources as well as much cheaper distribution and easier analysis. the use of the internet was accelerated with the popularity of the world wide web. for a coming standard of data documentation the display of documents through the use of html (hyper text markup language) was very stimulating. further stimulation came from the text encoding initiative (tei) that used sgml (standard general markup language) for marking up documents. consequently the proper move would be to make a document type definition (dtd) of data documentation using sgml. early on some were experimenting with html for their documentation but realized that if html was to become the standard they would tie the documentation to the presentation and thus commit an offence against the general principle of independence. quite similar to the reduction by software of documentation to sas or spss format or any other format it would be easy to reduce a complete standard data documentation like the ddi to an html page or to several selected forms of preferred html presentations. another technological development from the internet further paved the way for an effective solution for data documentation. the introduction of xml (extensive markup language) implied easier and more flexible work than sgml and with a strong connection to the internet. the slogan from jon bosak from sun microsystems one of the founders of xml – was that ‘xml gives java something to do’ (1997). internet, java, xml – all worked together. conclusion hopefully this paper has demonstrated that the origins of the ddi evolved from work related to social science data documentation issues by several institutions and people in the decades before the emergence of the data documentation initiative. during the period described in this paper several principles and levels of documentation were identified and further refined. re-using data is a community effort and the community encases the world. this was getting more and more attention with the fast expansion of the internet, with the spread of the world wide web. local inventions were put together and further developed in a distributed effort. that further story is another paper! actually the development of the ddi is discussed in several papers in this special issue of the iq. ann green and chuck humphrey describe the early years and mary vardigan offers a paper on the later years of development as well as a very useful ddi-timeline references adams, douglas (1986) the hitch hiker’s guide to the galaxy. london, heinemann. blank, grant (1993) codebooks in the 1990s; or, aren’t you embarrassed to be running a multimedia-capable, graphical environment like windows, and still be limited to 40-byte variable labels?. social science computer review 11(1): 63-83. blank, grant and rasmussen, karsten boye (2004) the data documentation initiative. the value and significance of a worldwide standard. social science computer review 22(3): 307-318. bosak, jon (1997) xml, java, and the future of the web. (). checkland, peter and holwell, sue (1998) information, systems and information systems: making sense of the field. john wiley & sons. de vries, repke and van der meer, cor (1992) exchange of scanned documentation between social scientists and data archives: establishing an image file format and method of transfer. iassist quarterly 16(1-2): 18-22. dodd, sue a. (1979) bibliographic references for numeric social science data files: suggested guidelines. journal of the american society for information science, march 1979. dodd, sue a. (1982) cataloging machine-readable data files. an interpretive manual. american library association, chicago. fienberg, stephen e.; martin, margaret e. and straf, miron l. (eds.) (1985) sharing research data. national academy press, washington, dc. geda, carolyn (2006) recollections of the formative years of iassist. iassist quarterly 30(3): 15-18. (). green, ann and humphrey, chuck (2013) building the ddi. iassist quarterly 37. icpsr isr (1973) osiris iii volume 1: system and program description. university of michigan. hauser, robert m. (1987) sharing data: it’s time for asa journals to follow the folkways of a scientific sociology. american sociological review 52(6): vi–viii. nielsen, per (1974) report on standardization of study description schemes and classification of indicators. copenhagen, danish data archives. nielsen, per (1975) study description guide and scheme. copenhagen, danish data archives. rasmussen, karsten boye (1981) proposed standard study description. the sd as a basis for on-line inventories of social science data. odense, danish data archives. rasmussen, karsten boye (1989) data on data. proceedings of the sas european users group international conference 1989, pp. 369-379. cary, nc: sas institute. rasmussen, karsten boye (1995) documentation what we have and what we want: report of an enquete of data archives and their staff. iassist quarterly 19(1): 22-35. (). rasmussen, karsten boye (2000) datadokumentation. metadata for samfundsvidenskabelige undersøgelser (data documentation: metadata for social science research). odense universitetsforlag, odense, denmark. rasmussen, karsten boye and blank, grant (2007) the data documentation initiative: a preservation standard for research. archival science 7(1): 55-71. rothenberg, jeff (1995) ensuring the longevity of digital documents. scientific american. january. rowe, judith (1999) the decades of my life, iassist quarterly 23(1). (). sieber, joan e. (1991) introduction: sharing social science data. in: sieber j.e. (ed) sharing social science data: advantages and challenges. sage publications, newberry park, ca, 1–18. vardigan, mary (2013) the ddi matures: 1997 to the present and ‘timeline’. iassist quarterly 37. vardigan, mary (2013) timeline. iassist quarterly 37. iassist quarterly 2013 35 iassist quarterly notes 1. karsten boye rasmussen is an associate professor of it and organization and a data scientist at the university of southern denmark. he worked in data archiving from 1974-1998. email: kbr@ sam.sdu.dk. 2. 3. keywords from vol 37 n0. 1-4. courtesy tagxedo.com vol233 summer 1999 19 this paper examines the role which historical gazetteers can play in webbased catalogues and data delivery systems. a gazetteer is a list of geographic names, which includes locational and other descriptive information. in this paper, the term ‘historical gazetteers’ is used specifically to describe gazetteers that incorporate both historical and modern geographical perspectives. in order to handle changed and changing geographical boundaries these gazetteers need to hold a wide range of information about geographic names, units, and hierarchies. this paper explains why gazetteers of this type are crucial for effective information retrieval and data browsing. in particular, it uses the history data service as a case study to describe how gazetteers of this type can be used to improve access to data via web-based catalogues and data delivery systems. this paper does not aim to describe the actual process of constructing and populating gazetteers (see harper 1997, hill et al. 1999, moss et al. 1998). the history data service (http://hds.essex.ac.uk) is funded by the joint information systems committee (http:// www.jisc.ac.uk) of the uk higher education funding councils to collect, preserve, and encourage the re-use of digital resources which result from or support historical research and teaching. the history data service is part of the uk data archive and is the arts and humanities data service (http://ahds.ac.uk/) service provider for the historical disciplines. the history data service collection covers a wide range of historical topics, and brings together over 450 separate data collections transcribed or compiled from original sources. the data collections cover a time period from the late tenth century to the mid-twentieth century, and the vast majority of data collections are either explicitly or implicitly geographically referenced. it is for this reason that the history data service is interested in developing and using gazetteers. explicitly and implicitly geographically referenced data correspond to a maze of complex geographies, which include administrative, electoral, census, and ecclesiastical geographies. these geographies are composed of a multiplicity of geographical unit types, which include amongst many others counties, wards, registration districts, and parishes. because of this complexity, gazetteers are crucial for effective information retrieval and data browsing. this holds true both in the context of an historical service provider like the history data service, and in the context of the wider social sciences and humanities community. gazetteers are needed to make sense of this maze of complex geographies for three main reasons. firstly many geographic names have a number of variant forms; secondly there are many incompatibilities between different geographies which means that boundaries do not align; and thirdly geographic names, units and hierarchies have changed in the past, and will continue to change. these problems are greatest with historical data, which are often associated with geographic names that have changed, or with geographical units that no longer exist, or with geographical units whose boundaries have changed significantly. it hardly needs saying that the disparity between modern and historical geographies increases with time. gazetteers improve information retrieval and data browsing by standardising geographic names and providing a controlled vocabulary of current and historical names within a system of preferred and non-preferred names. by linking disparate and changing geographies, gazetteers can help to integrate geographically referenced data collections, and deal with some of the incompatibilities when boundaries do not align. for example, gazetteers can make it easier to construct time series and other comparative data series by helping to identify those geographical units which, to a greater or lesser extent, correspond in different geographies. if gazetteers are to be used to improve information retrieval and data browsing, it is essential that we understand the needs and requirements of users. the history data service has an active and ongoing policy of consulting with actual and potential users, and we have established that many users from the historical community require web-based catalogues and data delivery systems which will allow them to perform sophisticated geographical searches in an fairly automated manner. users would like to be able to search changing boundaries: gazetteers, information retrieval and data browsing by cressida chappell* http://hds.essex.ac.uk http://www.jisc.ac.uk http://www.jisc.ac.uk http://ahds.ac.uk/ 20 iassist quarterly for data that cover a given place at a sufficient level of detail. for example, a user searching for the county of essex would like to recover not only data that are indexed by the geographic name essex, but also data collections that contain essex county-level data but which are indexed by a higher level geographical unit such as england. they might also wish to extend the search to include data that are indexed by geographical units within essex. users would also like to be able to search for any data that can be analysed at the level of a specified geographical unit. it is self-evident that a reasonably complex gazetteer, which holds information about geographical units and hierarchies, would be required if these types of geographical searches were to be supported. the history data service is working to improve and enhance access to its collection and a comprehensive uk historical gazetteer will be central to this work. historical gazetteers are attracting an increasing amount of interest from data providers, research projects, and traditional archives. in consequence, the history data service would like to develop a comprehensive uk historical gazetteer in collaboration with other services and projects. the history data service would use a comprehensive uk historical gazetteer both in web-based catalogues and data delivery services. it would use the gazetteer in web-based catalogues to support the types of geographical searches that users would like to be able to perform. information about the history data service collection is made available through three different catalogues, the uk data archive’s information retrieval system biron, the cessda integrated data catalogue and the arts and humanities data service gateway; however, of these only biron even adequately supports geographical searches. in biron geographical searching is facilitated by the geographical hierarchies in the humanities and social science electronic thesaurus, hasset (data archive, 1998). the geographical hierarchies in hasset have been built up over time by the uk data archive and the history data service, but they are not by any means comprehensive; the historical hierarchies in particular have been developed only as when they have been needed. the geographical hierarchies in hasset handle changing geographical boundaries by including geographical units in multiple hierarchies where necessary. the uk data archive and the history data service have increasingly recognised that the geographical hierarchies in hasset cannot fully support the types of geographical searches that users would like to be able to perform, and that in consequence a more complex and comprehensive uk historical gazetteer is needed. the history data service would also use a comprehensive uk historical gazetteer to help users to browse a web-based tree-structure, which will provide users with an alternative means of accessing information data. this will allow users to adopt a drill-down approach to locating data in addition to the more sophisticated geographical searching offered by web-based catalogues. in web-based data delivery services the history data service would use a comprehensive uk historical gazetteer to support geographical data subsetting. a geographical subsetting service has been developed for a large collection of nineteenth and twentieth century statistics assembled by humphrey southall as part of the great britain historical gis programme (southall and gregory, 1998). the great britain historical database online (history data service, 1998) allows users to search across 30 tables simultaneously to retrieve a geographical subset. users can select which variables are included in the subset, and they can also access online documentation. because the data collection included all the necessary gazetteers it was easier to develop a geographical subsetting service as part of the great britain historical database online. however, a comprehensive uk historical gazetteer is essential if the history data service is to extend this type of service to a wide range of other geographically referenced data. the history data service would also like to use a comprehensive uk historical gazetteer in web-based data delivery services to provide integrated access to historical data and appropriate digitised boundary data, which users could then utilise in a gis. the history data service and the ukborders service, located at the edinburgh data library, have been discussing the possibility of developing a joint interface which would provide integrated access to digitised boundary data held by ukborders and attribute data held by the history data service (such as the great britain historical database online). it hardly needs saying that it would not be possible to develop this type of service without a fairly comprehensive uk historical gazetteer. the history data service is confident that a comprehensive uk historical gazetteer can be developed in collaboration with other services and projects. we believe that it will enable us to respond to user needs and develop web-based catalogues and data delivery services which allow users to perform sophisticated geographical searches, and we believe that its use will result in improved information retrieval and data browsing. references data archive, 1998. humanities and social science electronic thesaurus: version 2. [online] http:// biron.essex.ac.uk/services/zhasset.html (14 may 1999). harping, p., 1997. the limits of the world: theoretical and practical issues in the contruction of the getty thesaurus of geographic names. ichim 97/eva 97, september 1997, paris, france. http://biron.essex.ac.uk/services/zhasset.html http://biron.essex.ac.uk/services/zhasset.html fall 1999 21 hill, l., frew, j. and zheng, q., 1999. geographic names: the implementation of a gazetteer in a georeferenced digital library, d-lib magazine, 5, 1. [online] http:// www.dlib.org/dlib/january99/hill/01hill.html (14 may 1999). history data service, 1998. great britain historical database online. [online] http://hds.essex.ac.uk/gbh.stm (14 may 1999). moss, a., jung, e. and petch, j.r., 1998. the construction of www-based gazetteers using thesaurus techniques. proceedings of international symposium on spatial data handling (sdh), july 1998, vancouver, canada. southall, h. and gregory, i., 1998. historical gis project home page. [online] http://www.geog.qmw.ac.uk/gbhgis/ (14 may 1999). * paper presented at the iassist conference, may 19, 1999, ryerson polytechnic university, toronto, ontario.. cressida chappell, history data service, uk data archive, university of essex http://www.dlib.org/dlib/january99/hill/01hill.html http://www.dlib.org/dlib/january99/hill/01hill.html http://hds.essex.ac.uk/gbh.stm http://www.geog.qmw.ac.uk/gbhgis/ vol21.1 4 iassist quarterly data archiving in africa: the south african experience by maseka a. lesaoana* introduction the south african data archive (sada) was established in 1993 by the centre for science development (csd) of the human sciences research council (hsrc) in pretoria. the first staff member of sada joined the hsrc in january 1994. as a data archive sada’s primary function is to locate, acquire, store and disseminate mainly quantitative machine-readable research data in the humanities and social sciences. south africa’s past has divided the research community in the country into two major groups: (i) the expert minority comprising mainly advantaged institutions (some universities, technikons and research councils; and the private sector); and (ii) the disadvantaged majority comprising largely universities, technikons and other institutions in the former homelands. in the short term, the expert community will tend to play the role of “depositor” of data at sada, while both the disadvantaged and expert communities will be users of the data. nineteen ninety-four (1994) was an epoch year in the history of south africa. apart from political renewal, it heralded a period in which researchers in south africa are faced with immense challenges in the development of a sound science and technology system. natural and social scientists are engaged in a number of large-scale local, national, regional and international research projects. however, the bulk of the government’s funding generally goes to natural sciences research and its support facilities. a unique challenge facing south africa in the transformation process following the first democratic election in 1994 is to implement the national system of innovation in science and technology, as defined in the recently published green paper on science and technology. the paper highlights national data gathering involving large sums of money, and suggests that the data be subjected to public scrutiny. south africa will soon establish a national research foundation (nrf) which shall promote and support research through funding, human resource development, infrastructure provision and capacity building in order to facilitate the creation of knowledge, innovation and development in all fields of science and technology. the nrf draft bill promulgated earlier this year indicates that the csd of the hsrc will form part of the social sciences and humanities division of the nrf. however, it is not clear whether sada and other research information facilities will also move to the nrf with the csd. the future of sada is shrouded in uncertainty. historical background in february 1993 per nielsen, the then director of the danish data archive (dda), was invited to the hsrc as a consultant to undertake a two-week feasibility study on the viability of establishing a data archive in south africa. he wrote a detailed report on his findings, including suggestions based on the experiences of data archives in other countries. nielsen found the hsrc a suitable location for the archive for the following two reasons: (1) the hsrc has the technical infrastructure needed to improve the quality of research and to provide training at all levels countrywide; and (2) the wealth of research data at the hsrc since its inception in 1969 can be placed at the disposal of sada for secondary analysis. some of the important issues raised in nielsen’s report included the following: n standards for proper documentation of data n rationale for establishing a data archive in south africa n implementation of the organizational structure n a variety of financial models including user payments n recovery of costs n staff composition n establishment of an interim advisory body n types of data at the hsrc that could constitute the first holdings of sada also discussed in nielsen’s report were the existing documentation standards, data-processing issues, international participation and co-operation, and implementation of bilateral co-operation. nielsen’s recommendations were used as a blue-print for the establishment of sada. challenges and developments location 5winter 1998 sada seems to be appropriately located at the hsrc, based on the experiences of other data archives, for carrying out its data-archiving activities. since its inception in 1969 the hsrc has accumulated a wealth of research data, mainly from numerous projects undertaken in collaboration with researchers at various universities countrywide. this provides an opportunity for sada to liaise with these researchers and their institutions. at the hsrc, sada is a division within the research information directorate (rid) of the csd. besides housing sada, rid also maintains a set of databases on human sciences issues through which it provides bibliographical information on current and completed research projects, research and professional organizations, researchers, and forthcoming conferences. the csd is committed to redressing the imbalances in research opportunities, and to empowering scholars from historically disadvantaged sectors by actively supporting their participation in the research structures, activities and decision-making processes of the broader research community. it also administers the allocation of various categories of research funding and scholarships to postgraduate students and researchers in the human sciences at south african tertiary institutions and ngos. the directorate: research capacity building of the csd supports the expansion of institutional research capacity by developing research skills among disadvantaged scholars. staff sada started its operations in 1994 with three staff members who joined the hsrc in january, september and november respectively. finding suitable staff in a country where data archiving and secondary analysis are new activities is not easy. furthermore, there are no data libraries or similar facilities in south africa, and methodology training courses are almost non-existent. the high staff turnover rate at the hsrc has also adversely affected sada’s progress. two of the three staff members left the organization (in february 1995 and december 1996). a fourth post of administrative assistant was filled in august 1996. sada staff were fortunate to be able to work with experienced data archivists. repke de vries of the steinmetz archive in the netherlands visited sada twice: (1) in october 1994 for a period of five months. the aim of this visit was to offer advice during the initial phase of the development of the archive; and (2) in may 1996 for a period of two weeks in order to study the developments implemented since his previous visit. shalane sheley of icpsr (inter-university consortium for political and social research) was contracted to sada for a period of one year ending in september 1996. in addition to the four posts originally approved, three new posts were added during 1996. two posts were filled in january 1997 and one in march 1997, and two are vacant. while many lessons can be learned through electronic discussions, appropriate technologies to meet the needs of the majority of researchers in a specific country can best be established through interacting directly and exchanging views with the researchers concerned. the research community a data archive needs to define clearly the research community for which the services are to be provided, both at supply and demand points. sada, as a national data archive, has stratified its research community into: (a) the academic community comprising universities, technikons, and training colleges (there are 21 universities and 15 technikons in south africa); (b) various government departments including the central statistical service, health, education, police, prisons and correctional services; (c) parastatal organizations such as the human sciences research council (within which sada is housed), the medical research council, the social sciences development forum, the development bank of southern africa, and several economic and social research institutes in the country; (d) research ngos such as the community agency for social enquiry (case) and the south african labour and development research unit (saldru); and (e) the private or commercial sector, including the mining industry and the association of market research organizations (amro). by defining the research community, sada can locate the type of studies undertaken and identify who should receive the archive’s services and what types of data are available or needed. knowing the research community also assists in the establishment of advisory committees and indicates the target groups to be consulted when assessing the delivery of services by the archive. meeting the needs of researchers data archives store data in such a way that they meet the needs of the majority of their researchers. often suppliers (depositors) of research data are also users at the demand point. however, at the demand point new researchers also become involved. an archive’s developmental stages should keep pace with researchers’ interests. computer technology plays a major role in data archiving. new developments in computer technology occur regularly, and data archives are faced with the challenge of keeping abreast of these technological advances. the recent switch from the mainframe to the unix environment is a typical example. because of the increasingly large amounts of data involved in research, new distribution and storage mediums and new statistical software packages are emerging. cdroms and cartridges have gained recognition, and due to their large storage capacity they are 6 iassist quarterly employed for storing large datasets in place of the still widely used diskettes. mainframes’ 9-track tapes are being phased out. data are stored and disseminated in a highly compressed form, and the internet has brought about new and user-friendly developments in computer technology. the data archivist’s job is to meet users’ needs effectively, while at the same time ensuring that the data in the archive are preserved for long-term usage. sada has devised a user survey form (also available on the internet) to determine what tools are most often employed by researchers in their analyses. this will assist sada in catering for its users’ needs. to date the response rate has been low, perhaps because in some cases researchers are not yet familiar with computer tools. since 1994 sada has organized workshops at major research centres in the country, mainly at historically disadvantaged universities (hdus) where the lack of infrastructure has proved to be a major problem. these workshops are aimed at creating awareness of the existence of sada and its activities. sada publishes a newsletter, sada news, and a draft sada guide (catalogue of holdings) to inform the research community of its activities and developments in data archiving. the study descriptions of sada’s collections are also accessible on the internet (world wide web). financial and time constraints financial constraints impede the establishment of any archive and the supply of data to the archive. the provision of cost-effective services by sada is not (yet) understood by funders and owners of data in south africa since secondary analysis itself is not well understood. even the data that are available are often not properly documented. data suppliers frequently request funds to enable them clean up their data. sada advisory committees after the visit by per nielsen in 1993, an interim advisory committee (iac) of sada was set up to offer advice on the establishment of a data archive. the iac was replaced by the sada board, also an advisory body, in february 1996. the board currently consists of 17 elected members representing the research community in south africa. the sada board reports to the hsrc council. an executive committee of the board was elected in november 1996. general progress sada has progressed well thanks largely to the help of the established data archives of europe and the usa. for instance, the above-mentioned visit by the then director of the dda, per nielsen, contributed greatly to sada’s understanding of data archiving. his investigations into research studies already undertaken in south africa, guided sada’s acquisition of potential holdings. he also assisted in registering sada as a member of the international federation of data organizations (ifdo), which led to a visit by the head of sada to a number of ifdo data archives: the icpsr in the us, the esrc data archive, the dda, the swedish data archive (ssd), and the steinmetz archive (star). the purpose of these visits was to learn how other data archives operate in order to design a suitable model for sada. it was during these visits that (i) sada became a member of the icpsr, and (ii) negotiations took place with repke de vries of star, resulting in his visiting sada for a few months to advise the staff at the take-off phase. he gave valuable tips on all aspects of data archiving and data management. during the planning stage, a cdrom drive, sas & spss, dbmscopy and folio views were acquired, and the internet facilities were set in place. conclusion the earliest data archives were established in europe and north america in the 1960s. through international associations such as ifdo, iassist and cessda these archives established guidelines for data archiving and shared their knowledge on the collection, storage, documentation and dissemination of data. newly established data archives are fortunate to be able to share this knowledge. however, for any data archive general considerations as well as unique considerations pertaining to the specific country have to be taken into account. general considerations include: location: in terms of wealth of and access to research data; institutional credibility; and technical and infrastructural support; staff: qualified staff to run the activities of the archive. initially fewer staff with broad knowledge base are needed; networking: networking with researchers and research institutions that produce data is vital. advisory committees formed with representatives from research institutions can help strengthen networking relationships; the science of data archiving: through international links such as ifdo, a new data archive can learn the science of data archiving, such as the procedures for acquisition, processing, storage and dissemination of data; and funding: this is a major problem. data archives are not profit-making facilities, and often depend on government funding. 7winter 1998 most specific lessons learned by sada stem from the economic and political changes currently taking place in south africa. data archiving activities are directly linked to the research capacity building activities at higher institutions of learning. due to apartheid, the education system in the country has been badly skewed. a few privileged institutions had the capacity to undertake research while the majority did not have the resources to do so. sharing of research between these two groups did not take place. sada is consequently faced with the challenge of rectifying this situation and also of promoting the new disciplines of data archiving and secondary analysis throughout south africa and africa in general. other issues to be considered by new data archives: research practices: these differ from country to country. some researchers do not want to give other researchers access to their data for reasons ranging from not being accustomed to the culture of sharing, to wanting to make money by selling their data; poor research training programmes: unlike in western countries, south africa has poor methodology programmes. summer schools in northern america and europe, for example, are held to strengthen the research capacity of data archives and data libraries; collaborative research in the region: joint activities by south african researchers are only now starting to take place. communication breakdown between african countries makes collaborative research difficult. in europe, for example, the eurobarometer studies and international social survey programmes are well-known research studies undertaken jointly by european countries. while some good lessons can be learned from visiting established data archives abroad, the best option is to invite experienced data archivists to the new archive at home. technology at developed data archives is already at an advanced level and it may be difficult, if not impossible, to implement new practices at the new archive where the infrastructure may be poor or not compatible. iassist and ifdo conferences have traditionally been hosted by canada, europe and usa in that order. the mission, objectives and activities of these bodies should be reviewed. one objective should be to draw in more participants, particularly from the third world countries and countries not previously included. efforts should be made to host such conferences on other continents and to provide funding for participation by researchers in third world countries, thereby increasing membership and building research capacity around the globe. * paper presented at the iassist/ifdo 1997 conference, may 6th may 9th odense, denmark. 14 iassist quarterly 2016 / vol 40 no 3 iassist quarterly abstract the effectiveness and impact of social science research is under constant review. from the sharing, reuse and archiving of social science research data to the outcomes, reach and impact of research, social science professionals are under increasing pressure to realise the maximum potential of their data collections and their research findings. the uk data service is playing a key role in supporting researchers in this process and is using detailed and well-received case studies to provide them with guidance on the best practice for sharing and reusing data and also identifying and capturing the impact of research. the impact of research is now routinely considered when examining the ‘success’ of funded projects, but the reality of identifying and capturing impact can be a challenge. publishing data in its own right is now recognised as being impactful to funders, yet exposing this narrative through ‘showcasing’ is under exploited. such narratives can incentivise others to share data and also improve the quality of the data and documentation. this paper seeks to explore the role that case studies of research can play in this regard. the paper also examines the role that depositor and user case studies can play in enhancing the reuse of a showcased data collection. to achieve this, a variety of illustrative depositor and impact case studies are discussed, highlighting the role that these can have on research projects. keywords case study, impact, data sharing, data management. introduction this paper explores the use of case studies in a variety of ways; first, it considers cases studies in regards to research and impact. second, the role they can play in assisting other depositors with depositing data at the uk data service. third, the role they can play in brand recognition and impact for the uk data service. finally, the paper concludes by identifying areas where we can develop the use of case studies further in the future. the uk data service has for a number of years been producing case studies on the data that we hold.3 these have traditionally been interviews or short vignettes of research papers, linked to the catalogue record of the data, and were mainly intended to show other researchers how particular data had been used. these short case studies described the research question posed, the data used, the methodology applied to the research and the findings and publications produced by the researcher. appendix 1 gives an example of this type of case study. case studies of this type are used primarily to give uk data service users accessing data information about research that had already used these data. they provide the basic ‘narrative’ for a piece of research, but do not look further at whether the research has been applied in practice, whether it has influenced policy or decision making, or whether any follow-up work has been completed. over the past few years’ we have been looking in more detail at how the uk data service is used and have been starting to track the impact of research using data held in the collection. a small team was created in the uk data service communications section, under the guidance of a director for communications and impact, which was tasked to identify and track research impact. identifying research impact is a time-consuming process and the team at the uk data service uses several methods to identify impactful research. the team follow the review documents produced for the national uk government and for local and regional government, policy reviews prepared by several leading uk think tanks and data use in the media. from these we can identify the research the role of case studies in effective data sharing, reuse and impact by rebecca parsons1 and scott summers2 vol 40 no 3 / iassist quarterly 2016 15 iassist quarterly publications that have been used for evidence and, if it has come from our collection, track the data use back to the uk data service. as well as this backwards identification of impact, the team are also in contact with think tanks and individual researchers about work that they are currently undertaking, which may be impactful in the future. this allows us to follow individual projects as they are happening. we also encourage researchers to submit details of publications and the team can then search to see if they have been used in policy. the focus on research impact is a relatively recent development in the higher education sector of the uk and has moved the definition of ‘successful’ research beyond simple metrics (for example a paper has been cited ‘x’ number of times, or a dataset has been downloaded ‘x’ times), to attempting to define how research has been used and what practical benefits it has brought. this change in measuring successful research has occurred at the same time as the expansion in providing open data, which anyone can download and use. for data which requires users to register with the uk data service, it is possible to track the metrics of data use, but for open data we do not record in detail how data are being used. there is now a greater emphasis on making data as open as practicable to allow greater use and reuse, but the focus of funders, research councils and the government is now not on the quantity of data being accessed, but on the quality of the research that comes from the data. case studies allow us to record the quality of research. we are still producing shorter case studies to document research, but we are supplementing these with case studies focussing on the impact that research has in the wider world, as well as the role that the uk data service plays in developing tools for data use and that share depositor stories for the benefit of other researchers. an example of an ‘impact’ case study can be found in appendix 2. the process of developing these case studies has been eye-opening, not just in seeing how widely the uk data service is used and the breadth of research that is coming from the data deposited with us and downloaded from us, but has also given us a greater understanding of the issues facing researchers in using and sharing data. case studies have enabled us to spread the word about the work that we do, but also analyse how the service is used and where we can improve our support for data users. research and impact a key feature of case studies is that they allow us to demonstrate the reach and impact of the research that is carried out using data downloaded from, or deposited with, the uk data service. within the uk there is a clear need for researchers to demonstrate the impact of their work, with ‘impact’ being a key indicator of success when universities are graded under the research excellence framework (ref) – a national system for assessing the quality of research produced by uk institutions.4 the ref uses case studies as a method of comparing the research output from universities and institutions who have to submit detailed case studies as part of their evaluation process. universities were last graded by the ref in 2014 and are now familiar with the developing case studies to show impact. case studies have become a familiar way of documenting the effect that research and projects have and most universities now employ staff specifically to promote research impact in preparation for the next ref in 2020. through talking with researchers that use the uk data service, we found that the expansion of these institutional impact case studies came with some issues for researchers. this is sometimes because the research has not yet produced any visible impact at the completion of the project, or because the researcher has moved on to new work and does not have the resources to identify the impact of their past work. this is particularly an issue for early career researchers, who may be building their portfolio of work and moving between academic institutions. uk universities are selective in which researchers are submitted to the ref for assessment, so researchers whose projects do not produce impact in a specific timescale can miss-out on institutional support for documenting the impact of their research. the uk data service identified this as one area where we could support researchers that use the service, providing useful evidence for researchers to use in their own career development and, at the same time, boost the profile and use of the uk data service. researchers can submit papers or projects that have used the service to us and the impact team will follow the outputs of the research to see what impact there has been. the case study is then prepared giving a summary of the research, information about the data that was used and the impact that has come from it. the case study is then used by the uk data service to promote the use of the data and is shared with the researcher for them to use too. the development of impact specific case studies has shown the breadth of research that has been undertaken using data from the uk data service. our case studies come from the fields of health,5 social policy,6 education,7 technology use,8 and business.9 examples from our most recent case studies include, dr ivy shiue who carried out a piece of research that compared health bio-markers in older adults with the temperature of their home and showed that living in a colder home has a detrimental effect on a person’s health, a finding which was used in the uk government’s winter planning policy.10 dr shiue has worked with the uk data service for several years and contacts the impact team when she publishes a piece of work that she thinks could be used by policy makers. another project entitled ‘onward migration,’ which investigated the experiences of refugees across the uk, was used as evidence by the scottish refugee council on the effect of forced relocation on refugee integration.11 this research was identified through media coverage of the project that prompted the impact team to study the data used by the researchers and contact them about their work. as a final example, we even have evidence of energy use for english wine production, a piece of research which led to advice for english wine producers on reducing their energy costs and carbon footprint,12 which was identified following the researcher submitting details of their research paper to the service. assisting other depositors through illustrative case studies another use that the uk data service has identified for depositor case studies is that of providing assistance to other depositors. put simply, depositor case studies allow us to point other depositors to illustrative examples of previous users that have deposited data with 16 iassist quarterly 2016 / vol 40 no 3 iassist quarterly us and how they managed – and dealt with – any challenges they may have had when depositing data. this has been highly informative not just for potential depositors, but also for the team who engage in research data management training. for example, one recent depositor case study helps to emphasise the message that we have been promoting for years at the uk data service, namely: the importance of planning for the sharing and managing of data early within the research project.13 dr karon gush’s study, which investigated how couples managed their households during recessions did just that, enabling the data to be effectively deposited and shared with the uk data service.14 dr gush and her team identified the potential challenges when it came to archiving and sharing the data early on in the project and sought to address them throughout the project so as to prevent future issues. for example, consent for sharing was gained from the participants when they were interviewed because the researchers had already identified that they wished to share the data and were aware that they would need ‘fully informed consent’ from the participants to do this. when it came to anonymising the data a careful balancing act had to be found so as to ensure that the data was as useful as it could be to future users, whilst still achieving the correct level of anonymity for participants. dr gush put it best in the case study interview when she said: “anonymisation should be thought about at the beginning and should be seen as ‘part and parcel’ of the whole project… and the anonymisation process should be completed as you go along and not left until the end.” 15 this is an important message that we have always promoted. now we can easily point depositors and users to the case study to highlight this. other depositor case studies have allowed us to explore and discuss a variety of other challenges; such as, the problems with archiving complex anthropological data ten years after the research is completed,16 and how one can gain consent for data sharing and archiving retrospectively.17 these case studies also help to illustrate the importance of adequate data management training, and planning being undertaken at the beginning of the depositor’s research project. by doing so, it ensures that most challenges/issues related to the research, the data, sharing and archiving are identified early on. thus, potential issues can be adequately addressed, which in turn leads to better quality sharable data. alongside this, there is a saving of resources for the researcher (both in terms of finance and time). utilising depositor case studies in this way has a variety of core benefits for us, our depositors and, our users; firstly, it allows us to actively encourage data deposits from users by showing how other data has been deposited and shared. secondly, it provides illustrative examples of how other depositors have addressed potential data sharing issues in practice. thirdly, it allows data depositors to build collaborations with users by providing further information on their research project and the steps that they went through to deposit their data. for example, this might be the anonymisation or consent procedure(s) they followed. fourthly (and finally), it allows us as a service to have impact in a variety of ways. specifically, it allows us to help other researchers to deposit data with us, share their research and experiences, have additional impact with their research, and provides examples of impact for our funders. raising awareness of the uk data service the final benefit to developing a range of case studies has been the collating of real-world evidence on the work of the uk data service. as with any publically funded organisation, the service has a need to prove to our funders and supporters the usefulness of the work that we do. in many ways, ‘data’ and ‘data archiving’ can be rather vague terms, particularly for those who do not work in the archiving or social science areas, and explaining the breadth of the work of the uk data service to people in the world outside our niche becomes much easier if we can show real examples of the data we hold actually in use. the uk data service is funded by the economic and social research council (esrc) and they have a remit to share the work that they fund with the general public.18 the case studies that we have developed for our funders focus both on individual pieces of research using data from our collection and also specific studies on how the service is being used. these more technical case studies look at the expertise and technology that the uk data service itself has and provide evidence for how the service is used. an example of this would be our case study on infuse – the interface that provides access to census data.19 we developed this case study to show how infuse was collaboratively developed, who was currently using it and some examples of research that had made use of it. this case study is now used by the esrc to illustrate the benefit of supporting data infrastructure like the uk data service.20 for use outside the data archiving community we worked with a data visualisation company to develop visual case studies, based on our impact and research case studies. these are an attempt to distil – on to one page – the research question, findings, methodology and policy implications of a piece of research, for example this case study looks at how long it takes elderly people to cross the road: vol 40 no 3 / iassist quarterly 2016 17 iassist quarterly this form of case study does have limitations: it is not suitable for very complex research questions and care needs to be taken that results or methodology are not over simplified by the visual representation. however, we have found these simple case studies useful when introducing the work of the service, particularly for school students aged 16-18 years old and undergraduate students, who are just starting to use social science data for themselves and need an entry level summary of how some of the data is being used. conclusions and the future in the future we would like to continue to engage with new uk data service users by developing case studies specifically for use in schools that can link to datasets suitable for teaching and are based on uk education curriculum. we also plan to develop case studies specifically focussing on early-career researchers, show-casing their work and helping them build a portfolio of impactful research. we will be continuing to add to our case study collection, with a particular focus on the innovative use of data. in addition, we plan to further develop the use of depositor case studies to assist our other depositors and complement our current research data management training and guidance. in particular, we will be identifying depositors’ collections which have ‘acutely sensitive data’ or are typically seen as ‘data that cannot be shared easily’, which we can point other users to. we will utilise these not only to illustrate how data can be effectively shared but also how these potential issues can be simply and effectively addressed in practice. appendix 1. example research case study health effects of industrial incinerators in england 21 about the research: research specific to waste incinerators has provided mixed evidence for the effects of proximity to incinerators on health. older incinerators have been associated with increased incidence and mortality from selected cancers, while more recent reports show little association. despite this, the effect of incinerator emissions remains a public health issue. this study assesses whether living close to industrial incinerators in england is associated with increased risk of cancer incidence and mortality. the results show no evidence of elevated risk for those living in areas containing an incinerator compared to those living in matched areas without an incinerator. moreover, within areas, there is little evidence of an increase in risk for those living in close proximity to an incinerator compared to those living further away. about the data: this research draws on aggregate data from the 2001 census for england. the data were key to the study for the identification of case circles around each incinerator, and for obtaining lower super output area level population counts, by five-year age groups and gender, for 2001. methodology: the researchers considered five regions with industrial incinerators in england, compared with five matched control regions, from 1998 to 2008. spatial and temporal trends in annual health outcome data within each circular region are analysed. initially, the researchers used a poisson log-linear model including age-standardised expected count as an offset and covariates for casecontrol status, matched pair and deprivation, fitted to circle level data to investigate temporal trends. they later modelled data at a finer resolution – at the lower super output area (lsoa) level. a poisson log-linear model was then used, as previously, with additional covariates for the effect of distance from the incinerator in case areas included as a multiplicative factor. population data was used in the calculation of the age-standardised expected count. 2. example impact case study do higher energy prices affect international trade?22 about the research: dr misato sato and dr antoine dechezleprêtre at the grantham research institute on climate change and the environment at the london school of economics and political science have been studying climate change policy and its effects on trade in a research project funded by the european union seventh framework programme under grant number no. 308481 (entracte research project), the global green growth institute, the grantham foundation and the economic and social research council (esrc) through the centre for climate change economics and policy. emissions trading policies, such as the eu emissions trading system (eu ets), are regulations implemented by many countries and cities to reduce industrial greenhouse gas emissions cost-effectively. these regulations are a key tool for achieving emissions reduction targets. however, according to the european commission on climate action they can result in carbon leakage and may affect the competitiveness of businesses. in fact, standard trade models suggest that policies that increase energy price put domestic firms at a disadvantage relative to foreign rivals facing lower energy prices. producers of energy intensive products respond to higher energy prices by producing fewer energy-intensive goods which may lead to a decline in net exports and the partial relocation of production to a region with low energy prices. 18 iassist quarterly 2016 / vol 40 no 3 iassist quarterly in this study dr sato and dr dechezleprêtre explored whether, and to what degree, changes in relative energy prices might influence trade and competitiveness. this is the first study to analyse the effect that energy costs have on global trade using historical data on trade and energy prices. it analysed 62 business and industry sectors in 42 countries over a 15-year period, from 1996 to 2011, using data that covers 60% of global merchandise trade. findings showed that changes in relative energy prices have a statistically significant but very small impact on imports. on average, a 10% increase in the energy price difference between two country-sectors increases imports by 0.2%. the impact is larger for energy-intensive sectors but even within these, the effect is minor changes in energy price differences across time explain less than 0.01% of the variation in trade flows. the authors calculated that a 30% increase in energy prices across europe would cause exports to fall by only 0.5% and would increase imports by 0.07%. they concluded that “contrary to some claims, rises in energy prices do not have much effect on the global competitiveness of businesses. even a sizeable difference in the price of energy relative to the rest of the world has only a very small impact on a country’s imports and exports.” methodology: the researchers estimated the short-term effects of energy price differences on bilateral trade at the sector level using a gravity model. trade between countries is not only influenced by energy costs but many other factors such as transport costs, labour cost, exchange rates, tariffs, trade agreements, common language and common currencies. the study therefore brings together a variety of datasets to analyse the impact of relative energy prices on trade. the coverage and detailed disaggregation of the data used goes well beyond previous work, allowing the first global ex-post analysis of the relationship between trade and energy prices. bilateral trade data were taken from the cepii’s baci database which contains detailed bilateral import and export statistics from the un commodity trade database. the researchers also used a unique and comprehensive dataset of industrial energy price indices at the country and sector levels covering 48 countries and 12 industry sectors for the period 1996 to 2011. the dataset was constructed in sato et al., 2015 and uses data from the international energy agency world energy balances, the international energy agency energy prices and taxes and the world bank, as well as other sources. the procedures used to construct the dataset including the methodology developed to reduce missing data-points, are documented in the working paper, as are the full set of original data sources. this dataset is made publicly available for download here. additionally, the researchers used data on gdp and population obtained from the international monetary fund’s world economic outlook (2012) and data on wages obtained from the united nations industrial development organization (2011). findings for policy: this study finds unique evidence suggesting that concerns about the risks of carbon leakage may have been overplayed. carbon pricing policies are intended to induce energy intensive industry to reduce carbon emissions by making it more expensive to pollute. however, if the carbon price is set too high, there is a risk that companies will respond by relocating production to regions with lax policies, rather than clean up their production. whether this ‘carbon leakage’ occurs is an important issue in the debates around how to design emissions trading schemes, and whether or not leakage occurs is a question in much need of robust empirical evidence. impact of the research: the study’s findings about the risks of carbon leakage have been included as research evidence in the ‘evaluation of the eu ets directive’ report carried out by the european commission in november 2015. the evaluation subsequently informed policy measures implemented by the commission regarding the revision of the eu ets directive, as set out in the framework of measures of the conclusions of the european council in october 2014. the study has also informed the following reports: • the oecd environment working paper no. 87, a review of literature on ex post empirical evaluations of the impacts of carbon prices on indicators of competitiveness. • the oecd economics department working papers no. 1282, “do environmental policies affect global value chains?” • the new climate economy working paper ‘implementing effective carbon pricing’, which was written as a supporting document for the 2015 report of the global commission on the economy and climate, seizing the global opportunity: partnerships for better growth and a better climate • methods for evaluating the performance of emissions trading schemes, a discussion paper by climate strategies, prepared as part of the project “evaluation of emissions trading scheme pilots in china” funded by children’s investment fund foundation (ciff) and executed by the tsinghua university. notes 1. rebecca parsons, uk data service, uk data archive, university of essex, colchester, essex, uk. rparsons@essex.ac.uk 2. dr scott summers, uk data service, uk data archive, university of essex, colchester, essex, uk. ssummers@essex.ac.uk 3. uk data service (2016), ‘case studies’ https://www.ukdataservice.ac.uk/use-data/data-in-use/case-studies 4. research excellence framework (2014), ‘ref 2014’ http://www.ref.ac.uk/ 5. uk data service (2016), ‘nearly a third of welsh adults struggling to cope with the pain of chronic health conditions’ https://www.ukdataservice. ac.uk/use-data/data-in-use/case-study/?id=192 6. uk data service (2015), ‘children with psychological distress are more likely to become unemployed’ https://www.ukdataservice.ac.uk/use-data/ data-in-use/case-study/?id=168 7. uk data service (2015), ‘exploring the ‘middle’ in gcse attainment’ https://www.ukdataservice.ac.uk/use-data/data-in-use/case-study/?id=179 vol 40 no 3 / iassist quarterly 2016 19 iassist quarterly 8. uk data service (2015), ‘screen-based media and well-being in adolescence’ https://www.ukdataservice.ac.uk/use-data/data-in-use/ case-study/?id=177 9. uk data service (2015), ‘investigating external and private benefits from investment in skills and training: uk innovators study’ https://www. ukdataservice.ac.uk/use-data/data-in-use/case-study/?id=194 10. uk data service (2016), ‘the impact of cold homes on health’ https://www.ukdataservice.ac.uk/use-data/data-in-use/case-study/?id=195 11. uk data service (2016), ‘moving on? dispersal policy, onward migration and integration of refugees in the uk’ https://www.ukdataservice. ac.uk/use-data/data-in-use/case-study/?id=191 12. uk data service (2015), ‘energy use within the english wine production industry’ https://www.ukdataservice.ac.uk/use-data/data-in-use/ case-study/?id=176 13. uk data service (2016), ‘karon gush – depositor stories’ https://www.ukdataservice.ac.uk/deposit-data/stories/gush 14 ibid 15. ibid 16. uk data service (2016), ‘pat caplan – depositor stories’ https://www.ukdataservice.ac.uk/deposit-data/stories/caplan 17. uk data service (2016), ‘maggie mort – depositor stories’ https://www.ukdataservice.ac.uk/deposit-data/stories/mort 18. economic and social research council (2016), ‘esrc shaping society’ http://www.esrc.ac.uk/ 19. uk data service (2011), ‘infuse’ http://infuse.ukdataservice.ac.uk/ 20. economic and social research council (2016), ‘opening up census data for research’ http://www.esrc.ac.uk/news-events-and-publications/ impact-case-studies/opening-up-census-data-for-research/ 21. uk data service (2014) ‘health effects of industrial incinerators in england’ https://www.ukdataservice.ac.uk/use-data/data-in-use/ case-study/?id=163 22. uk data service (2016), ‘do higher energy prices affect international trade?’ https://www.ukdataservice.ac.uk/use-data/data-in-use/ case-study/?id=206 sist newsletter vol.1, no. 4 (5) roll or scroll data is rolled up or down the screen, permits the user to scan a large volume of data (6) paging data is stored on pages (a full screen) user is able to review any selected page. editing functions editing features include: (1) character deletion (2) line insertion (3) line deletion (4) erase (5) character repeat. external i/o devices can add flexibility to the applications possibilities for display terminals. a cassette tape drive or diskette drive can be used to store display formats, data to be transmitted, or user programs. a printer can provide hard copy. selecting a terminal some questions you should ask yourself when selecting a display terminal are: (1) what are the essential parameters for a display terminal that will satisfy your needs? (2) who supplies the terminals with the features you desire? (3) maintenance provisions? (4) talk to users concerning problems encountered when installing it, failures that have occurred, and any incompatibilities. discussion paper/alice robbin the issue of confidential data: the need for formulation of policy by the data archive and library by alice robbin data and program library service university of wisconsin-madison the issue of confidential data i. during the last decade there has been increased concern about the problems of confidentiality involved in the collection and dissemination of individual microdata. concern has revolved around the government's perceived need to collect increasing amounts of information at a microdata level for social policy formulation and evaluation, types of information which potentially compromise sist newsletter vol.1, no. 4 personal privacy; lack of good security measures for protecting access to personal information; availability of this information to individuals, private firms and public agencies unrelated in any way to the primary producers of the statistical data; growing restrictions placed on the access to statistical microdata by the statistical agencies, restrictions which appear to be in response to public pressure; and, the impact of restrictions on access to these data which are needed for research and public policy development by the scholarly community and others engaged in statistical analysis. this concern has been manifested in a very large literature, formal conferences and commissions, and legislation.! recently, a meeting concerning the issues of privacy, confidentiality, and the use of government microdata for research and statistical purposes was held at the rockefeller foundation's bellagio study and conference center at lake como, italy, august 16-20, ig?/^.'^ it seems useful to summarize the draft report on this conference because of the importance of this issue for the data archive. ii. david flaherty, author of the draft report describes the origins of the bellagio meeting. during the last three years he and edward hanis of the university of western ontario have been involved in the privacy project studying the problems of privacy and confidentiality involved in the collection and dissemination of individual microdata by central statistical agencies in canada, the united states, the federal republic of germany, sweden, and the united kingdom. the privacy project identified key personnel in each central statistical agency and found that they face common problems in the formulation of policy on confidentiality and data dissemination. the privacy project found that although these individuals face common problems they have had surprisingly little contact with one another on these issues. the bellagio conference brought together leading individuals from each of the five statistical agencies to discuss common problems and solutions. the conference included sociologists, economists, and historians because it was felt necessary to demonstrate to the "custodians" of government data that a real need exists for access to individual microdata for research and statistical purposes. a series of general principles or propositions pertaining to research and statistical uses of individual information held by government agencies evolved during the meeting. i have reproduced the 18 principles in toto , since i feel it useful to add to the public debate on the privacy/confidentiality issue. many of the principles, while of potential interest to the data archive community do not directly affect data archivists at a local repository level; however, i have starred (*) those which i think may prove applicable to the archive's role. 1. national statistical offices should provide researchers both inside and outside government with the broadest possible access to information within the bounds of accepted notions of privacy and legal requirements to preserve confidentiality . 2. legal and social restraints on the dissemination of microdata are appropriate when they reflect the interests of the general public in an equitable manner. these constraints should be re-examined when they result in the protection of vested interests or the failure to disseminate anonymized information for statistical and research purposes without direct consequences for a specific individual . sist newsletter vol.1, no. 4 all copies of government statistical records should be rendered immune from compulsory legal process by statute. in making data available to researchers national statistical offices should provide some means to ensure that decisions on selective access are subject to independent review and appeals. the distinction between a research file in the sense of a statistical record (as defined in the 1977 report of the u.s. privacy protection study commission) , and other micro files is fundamental in discussions of privacy and dissemination of microdata. all dissemination of government microdata discussed in connection with the bellagio principles is assumed to be a transfer of data to research files. there are valid and socially-significant fields of research for which access to microdata is indispensible. one of the prime sources is government microdata, including statistical agencies . public use samples of anonymized individual data are one of the most useful ways of disseminating microdata for research and statistical purposes . techniques now exist that permit preparation of public use samples of value for research purposes within the constraints imposed by the need for confidentiality. countries with strict statutes on confidentiality have prepared public use samples. there are legitimate research purposes requiring the use of individual data for which public use samples are inadequate. there are legitimate research uses which require the utilization of identifiable data within the framework of concern for confidentiality . other techniques of extending to approved research the same rights and obligations of access enjoyed by officers of the agency need to be considered in terms of better access. there is considerable potential for development of more economical and responsive customized-user services, such as 1) record linkage under the protection of the statistical office, 2) special tabulations , 3) public use samples for special purposes. such services must often involve some form of cost recovery. some research and statistical activities require the linking of individual data for research and statistical purposes. professional organizations or national organizations should have codes of ethics for their disciplines with respect to the utilization of individual data for research and statistical purposes. these ethical codes should furnish mutually agreeable standards of behavior governing relations between providers and users of governmental data.^ users of microdata should be required to sign written undertakings for the protection of confidentiality . 22 sist newsletter vol.1, no. 4 16. considerable efforts should be made to explain to the general public the procedures in force for the protection of the confidentiality of microdata collected and disseminated for research and statistical purposes . 17. the right of privacy is evolving rather than static, and closely related to how statistics and research are perceived. therefore, statisticians and researchers have a responsibility to contribute to policy and legal definitions of privacy. 18. public concern with privacy in the collection and utilization of individual data can be addressed in part as follows: 1) voluntary data collection, whenever practicable 2) advanced general notice to respondents and informed consent, whenever practicable 3) provisions for public knowledge of data uses 4) public education on the distinction between administrative and research uses of information. flaherty summarizes the discussion on each of these principles in a concise form. appendix i includes the conference participants. appendix ii summarizes the themes which emerged from the sessions on the united kingdom (public relations, uses of information, and dissemination of data) and canada (dissemination and uses of data). appendix ii describes the agenda for the general sessions: goals, disseminating data, regulation of dissemination, forms of microdata dissemination, accountability of researchers, and privacy and data collection. appendix iv lists the 18 principles agreed upon by the participants. iii. data archives are concerned with the preservation of information. not only are there technical issues related to preservation and release of statistical information, but also ethical, moral and legal considerations about the release of confidential information on individuals. what should be the role of the data archive in accepting information of a confidential nature? should the archive agree to preserve this information? what sorts of restrictions on access should archives adhere to? if the archive is to distribute confidential information, what sorts of protection should be applied to this information? these and other questions concerning the maintenance and release of confidential information are questions which seem to have been raised by few data archives and organizations outside the federal statistical agencies such as the bureau of the census. most of the information which the data archive has dealt with has been of an anonymous nature and with few exceptions little confidential data have been deposited with the data archive. [this discussion does not pertain to the national archives, which has had a long history of concern with this issue and has developed various mechanisms for dealing with the problem.] nevertheless, confidential information is being collected and archives as such do play a role as preservors of this information. in addition, there are a number of data organizations, integral parts of (survey) research organizations, which are responsible for maintaining the data files created by their researchers; there are data libraries, which although do not have a mission to preserve original data, sometimes (because of their experience with data generation and processing) become involved in projects where confidential data are collected. 23 sist newsletter vol.1, no. 4 at the same time public concern about the release of confidential information by statistical agencies has led to restrictions on access to information needed for research and public policy development by the data archive's clientele. this suggests that the issue is of immediate and continuing attention by the data archiving community. what to do about confidential data, what is the role of the archive, and what are the archive's responsibilities in this area are difficult questions and have no easy solutions. the problem does suggest however that it would be wise for a data archive to formulate a coherent policy on the preservation and release of statistical information which potentially invades individual privacy. while i have not suggested any guidelines that a data archive can follow for formulating a policy on the archiving and dissemination of confidential data, i hope that this discussion will stimulate further thoughts by the archive community and a continuing dialogue in this newsletter . footnotes and references some introductory materials to the issue of privacy/confidentiality are: 1) the house and senate hearings held during the middle 1960's, which dealt with the creation of a federal (national) data center, contain a great deal of discussion on the confidentiality problems inherent in a government data center. u.s. congress. house. committee on government operations. special subcommittee on invasion of privacy. the computer and invasion of privacy . hearings, 89th cong., 2d sess. washington: g.p.o, 1966. 318 pp. (appendix 1: "the ruggles report," pp. 195-253; u.s. congress. senate. committee on the judiciary. subcommittee on administrative practice and procedure. computer privacy . hearings, 90th cong., 1st sess., on s. res. 25. washington: g.p.o. , 1967. 388 pp. ("the kaysen report," pp. 357-359.) u.s. congress. house. committee on government operations. privacy and the national data bank concept . 90th cong., 2d sess., h. rept. 1842. washington: g.p.o., 1968. 34 pp. 2) the alan westin book. privacy and freedom (new york: atheneum, 1967), provides a basic introduction to the problem. see also his book data banks in a free society (new york: quadrangle books, 1972). 3) it is useful to consult the various pieces of legislation which have been passed in the united states, such as the privacy act of 1974, the freedom of information act, federal reports act, fair credit reporting act, and the family educational rights and privacy act of 1974. summaries of the contents of these acts are contained in kent greenawalt's monograph. legal protections of privacy, final report to the office of telecommunications policy (executive office of the president). (washington, d.c.: superintendent of documents, u.s. government printing office). 24 l^^sist newsletter vol.1, no. 4 4) legislation has been considered and promulgated in other nations. orjar oyen, in his article, "social research and the protection of privacy: a review of the norwegian development" (acta sociologica : 19 [1976], pp. 249-262), discusses what is occurring in norway. the impact of the swedish data act is discussed in the proceedings of a symposium on personal integrity and the need for data in the social sciences , held at hssselby slott, stockholm, march 15-17, 1976, sponsored by the swedish council for social science research. other discussions have taken place and are reported in proceedings: stig stromholm, right of privacy and rights of the personality: a comparative survey (stockholm: p. a. norstedt and soners forlag, 1967), which is a working paper prepared for the nordic conference on privacy organized by the international commission of jurists, held in stockholm, may 1967. see also the first international oslo symposium on data banks and society , by universitetsforlaget in oslo in 1972. 5) for discussions about statistical techniques to protect confidential information, see the gwendolyn b. moore et al . report, accessing individual records from personal data files using non-unique identifiers , prepared for the institute for computer sciences and technology of the national bureau of standards in washington, d.c. (nbs special publication 500-2, u.s. department of commerce, national bureau of standards); appendix a ("confidentiality-preserving modes of access to files and to interfile exchange for useful statistical analysis" by donald t. campbell et al . ) in protecting individual privacy in evaluation research , published by the national academy of sciences (washington, d.c, 1975) (included in this monograph is an excellent bibliography); proceedings of the second midwest conference on confidentiality of health and social service records: where law, ethics, and clinical issues meet (chicago: university of illinois, 1976). robert boruch, director of the evaluation research program at northwestern university, has been very kind to supply me with a number of articles and references to techniques for ensuring confidentiality. some of the articles by him include "relations among statistical methods for assuring confidentiality of social research data ( social science research i, 403-414 [1972]); "strategies for eliciting and merging confidential social research data" (nejelski, p. [ed.], research in conflict with law and ethics (cambridge: ballinger, 1976); "statische unde methodische prozeduren zur sickerung der vertraluichkei t bei forschung" (eser, a. and k.f. schumann [eds.], forschung im konflikt mit recht un ethik [stuttgart: ferdinand enke verlag, 1976]); "educational research and the confidentiality of data: a case study", sociology of education 44, 1971, 59-85; "methodological techniques for assuring personal integrity in social research", september 1976 (prepared as background research for evaluation research program's project on secondary analysis. 6) most of the articles cited above also discuss the ethical issues inherent in the confidentiality problem. in addition to the boruch article published in the sociology of education , and the nejelski (editor) book, also useful is the boruch and joseph s. cecil article, "is a promise of confidentiality necessary? sufficient?", which appears as chapter 3 of methods for assuring confidentiality of social research data , prepared for the american psychological association task force on privacy and confidentiality and national academy of sciences committee on federal statistics, panel on privacy and confidentiality as factors in survey response (background research). 25 sist newsletter vol.1, no. 4 2 david h. flaherty. "report of the bellagio conference on privacy, confidentiality, and the use of government microdata for research and statistical purposes, lake como, italy, august 16-20, 1977." (draft.) university of western ontario, september 17, 1977. (mimeographed.) 3 for a brief description of the privacy project, see flaherty's paper. privacy and access to government microdata," (revised, november 8, 1976) prepared for the executive seminar on "expanding the right to privacy: research and legislative initiatives for the future," washington, d.c., october 14-15, 1976, sponsored by the washington public affairs center, university of southern california. the inter-university consortium for political and social research (icpsr) has developed a statement on the archiving, processing, and release of confidential data. i include it because it represents the fullest statement to date which describes what icpsr processing and dissemination responsiblities are in this area. preservation of confidentiality the issue of confidentiality arises prlradrily, but nol exclusively, in the case of data collections that include information that is or is seen as potentially damaging or threatening to respondents, oi when a promise of confidentiality has been given respondents in the process of data collection, and when information is included i hat allows or potentially allows identification of individual respondents. the ronto elite populations. such data are now increas ingly av^ijaolc, and elite are more easily identifiable by means of a few key variables than are respondents in a mass survey. also contributing to the gravity of the issue is the sensitivity of the public as a whole, and publii' figures more particularly, to the "potential" uses of data obtained tlirongh personal interviews. the increased research focus on elite populations and the resultant increased availability of data combined with the need to facilitate extended research use of data collections makes it necessary for the consortium to develop policies and procedures for ir-filinace data dissemination without jeopardizing the rights of respondtjnls to confidentiality. it should be noted, of course, that the issue of confidentiality is not confined to infornation collected through personal intnrviews nor ir. il limited to elite populations or even individuals. data collected through for example—also present the issue, and data collected from public record sources, such as court records, can be seen as unnecessarily and unjustifiably damaging to individuals unless anonymity is provided. data on organizations or particular groups can also present confidentiality issues as can data relevant to deviant populations and 1 he issue is also raised by mass surveys. while the following pollcie,-, are intended to apply specifically to individual level data for elite populations, they are also seen as applying in more general rermr to other categories of data such as those suggested above. to protect the identity of respondents, the followlnr guidelines will 1. the archive will not accept any documentation, list or da( a files that explicitly identify respondents by name except in the case of data that are in publicly available sources and which do not present problems of confidentiality. 2) in cooperation and consultation with principal invest i r.itor('l the consortium staff will assist in the idenr i t ic.it ion of "sensitive" variables but reserves the right to extend the list of variables to be deleted, aggregated or otherwise masked beyond those identified by the principal investigate tcs) 3) variables will be flagged when a) they allow identification of particular respondents, and b) when used in combination with ether variables, they also identify particular respondents. sist newsletter vol.1, no. 4 the list of sensitive variables will be presented to investlgator(s) with recommendations for uhich variab deleted, collapsed or grouped (such as age) or have t gory descriptions changed to generalize the category-. a version of the data which incorporates the above rccot 3r general distribut ion. sion of the data to be distributed6) the codebook for the include documentation for the masked variables, a descrip of the nature of the deletions and maskings and frequenci the masked variables. 7) requests for data reductions (for example, cross-tabulations) or aggregations involving the masked variables will be accepted, but the consortium staff will judge whether the requested data reductions or aggregations will preclude identification of individuals and will supply only those results that do so. decisions will be made in consultation with principal investigators in order to insure that protection of confidentiality does not unnecessarily limit researchers* access to analytic information. 8) the original version of the data will be maintained under security (see document entitled "physical file security"). 9) public record data will not be added or merged into a survey data file if that data allows identification of individuals. if public record data have been added to the survey data file and identification of individuals as a consequence is possible, these data will be retboved from the survey data file and retained separately without a linking variable that would allow the two files to be ed method for handling confidentiality atteirpts to protect the s without destroying the meaning of the data. if this process e data meaningless in any given case, the data could be mainta rity by the consortium and not be generally distributed. the ion could be distributed upon request and only data reduction ould be accepted. the da the data and program library service has attempted to identify donor , archive and user responsibilities in a simpler form, since it involves itself less frequently in processing a data file. appended is a first (working) version of the dpls "authorization form" which is agreed upon by the donor when depositing data at dpls, dpls when accepting responsibility for a file and the user when receiving a copy of the data. of particular relevance to the readers are lines #3-5, 10-18, 27-32. as the readers will note, dpls considers it essential that restrictions may be "relaxed" upon review and approval of data access by the dpls faculty advisory committee. data and program library service 4451 social science building university of wisconsin-madison madison, wisconsin 53706 608-262-7962 authorization form the data and progi related quantitative m£ when authorized as a pe sponsibility for the cc protection against the library service ser ne-readable data fo epository for >f the data, changes if information deemed the university of wi' e indicated by the donor of the data, distribu individuals or institutions which shall utiliz rch purposes. donor of a data file{s) ions placed upon dissemin dpls is responsible for n >e reproduced without the lit of not more than three onfidentiality constraint requested to supply dpls with information on a ion of the file. dpls adheres to these restric ifying the potential user of a file that the fi itten consent of the donor. dpls prefers that 3) years be applied to restrictions on the file limit in any way distribution of the data. in that case, unless the archive is assured by the donor of the da own evaluation of the file that confidentiality requirements ca cess to the file will remain restricted. be me publi sist newsletter vol.1, no. 4 dpls requires that the user of the data agree to the following conditions for receipt of the file: (a) neither the transmitted file(s). listing (s), nor any copy thereof in whole or in part, shall be disseminated, or sold, to any further party; (b) the data provided shall be used solely for professional or other official purposes related to teaching, research, public service, public policy, and planning; cc) publlcauon or dissemination of data, or the results of analyses of che data provided, shall be based on a minimum of five individual subjects and/or .three organizations per distributional or tabulated cell and neither any individual subject or organization shall be explicitly identified. uhere exception to this requirement is requested, permission for access co the data shall be granted by the daj:a and computation center faculty policy committee upon review of the research needs of a particular project; (d) an abstract or copy of any published document upon which the analysis is based be provided to dpls; {e) dpls will not be responsible for losses or inconvenience resulting from delayed delivery or errors due to defects attributable to the requestor's own equipment. dpls will bear the cost of replacing the data in the event that the defect is attributable to dpls's work; (f) responsibility for the accuracy of the data and documentation rests with the donor unless dpls has itself been responsible for producing the file or documenflease sign below if you agr donor: na telephone dpls discussion paper/ donald f. harrison on work in progress: directory of directories donald f. harrison national archives washington, dc the us action group on process produced data files, in its mandate to prepare a "directory of directories", (lassist newsletter vol . i, no. 3, may 1977), has prepared the following initial entries and desires feedback from the membership. with two exceptions, entries describe only social science data files, printed and available in the united states. because of the nature of where files are created, plus the specialized knowledge of the list makers, this list is heavy on federal directories. the committee seeks information on: 1) additional directories not listed and available for description; and, 2) additional or different information that may be desireable in the entry format. please send any suggestions to: donald f. harrison chairman process produced data ag machine-readable archives division (nnr) national archives washington, dc 20408 28 iassist quarterly 2013 57 iassist quarterly abstract the social science data community is fortunate to have a tremendous group of talented professionals. we work each day to build upon the ideals founded by those that came before us. this article recognizes the efforts of early iassist members whose pioneering efforts enabled our work today. in particular the work of sue dodd will be acknowledged. this article reflects on the many partnerships and data oriented projects the author has had the good fortune to be a part of over the last ten years in his work at the odum institute (odum, 2014). many of these projects are work performed under the aegis of the data preservation alliance for the social sciences “data-pass” (datapass, 2014). these projects are just a small subset of many achievements accomplished by the greater social science data community. keywords: collaborations, digital preservation, data management, metadata, catalog a personal prologue it is an honor to have the opportunity to reflect on the advancements our social science community has made in recent years toward metadata harmonization as well as the preservation of the materials we all hold so dear. i have had the pleasure of working for the odum institute, university of north carolina unc, for twenty-one years so it seems fitting for me to reflect on what has been accomplished as the institute celebrates its 90th anniversary this year. i was fortunate that my service here at the institute overlapped with the tenure of sue dodd, if only for a few years. the institute has always provided a home for researchers and staff who share a passion for service to the social science community. sue dodd exemplified this ideal, and today as we build on her work, the odum institute data archive is dedicated to serving the social science community and its customers around the world who are seeking critically important data and information to support their research and data management services. foundations libraries and archives have been organizing information long before the advent of digital records. in early 1970 the investigation into a set of rules to catalog machine-readable data files, or “mrdf” began. (dodd, 1982). recognizing the unique properties of digital materials and having a keen eye towards both the potential challenges and affordances of cataloging these materials, sue dodd was instrumental in the evolution of cataloging standards for mrdf that first made their appearance in the second edition of aacr2 published in 1978 (dodd, 1982). her work paved the way for the development of many tools that simplify the discoverability, accessibility, and usability of vast amounts of social science data. today’s advancements would have been tremendously more difficult without the development of standard cataloging requirements and descriptive methodologies used to define these mrdfs. my early work at odum was in the information technology arena. i knew nothing of these early foundations and the valuable work of my new colleague sue dodd. i did not know that one day i would be tasked with the migration of thousands of building on the work of colleagues: a moment of reflection by jonathan d. crabtree1 her work paved the way for the development of many tools... 58 iassist quarterly 2013 iassist quarterly catalog records from marc format (marc, 2014) to our current format the data documentation initiative “ddi” (ddi, 2014) (blank & rasmussen, 2004). it was the thoughtful design and planning during the early days of mrdf catalog records that made my job much easier. little did i know at that time, the standardization of mrdf catalog records and the early efforts of the social science community to adopt and embed these nascent standards into their workflows have provided the bedrock upon which we build today’s modern archive systems. dodd asked in her writings “where do we go from here” (dodd, 1982)? ever prescient, she speculated that researchers and scholars would need these records to enable shared cataloging, authority control, acquisition systems, private file creations, products and a union list. as we know, these services and products we now take for granted are offered around the world today for a vast amount of social science data. behind this mountain of data is a network of researchers, archivists, librarians, information scientists, and administrators like sue dodd who work tirelessly to safeguard and provide access to valuable social science data that has helped to guide everything from public policy to education. we owe credit not only to sue dodd, but also to the whole of our international social science community for building these remarkable tools and services that continually add to the legacy of pioneers in our field. i value this opportunity to reflect on the enriching collaborations i have been involved with over the past ten years working to fulfill these earlier visions, and i encourage readers to do the same. building a union catalog the odum institute was involved with one of the first projects following the founding of the national digital information infrastructure and preservation program “ndiipp” (library of congress, 2014). as part of the newly formed data preservation alliance for the social sciences “data-pass” (data-pass, 2014) led by the inter-university consortium for political and social research (icpsr), we became a member of a voluntary partnership to archive, catalog and preserve valuable social science data that were at risk, in support of the ndiipp agenda. the early data-pass partners – icpsr, the institute for quantitative social science at harvard (iqss), the odum institute, the roper center, the national archives and records administration (nara), and the murray center -were not strangers to one another. for many years, we had worked together on projects to provide access to quality social science data for our constituents. this familiarity, combined with a shared common goal, allowed the partnerships to grow and take root. once data-pass was established, the group immediately began to survey the landscape and take action. by building on existing relationships, the data-pass partners were able to expedite the process. (crabtree & donakowski, 2006). the datapass partners had four primary goals during the ndiipp project: (1) archive at-risk social science content, (2) build a shared union catalog, (3) provide replicated preservation, and (4) advocate for best practices in digital preservation. we began to identify at-risk content almost immediately and developed strategies and best practices to manage this task. jointly, we also began to develop a plan to take steps toward building the union catalog envisioned in the early days of mrdf catalog records. our strategy was to utilize standard harvesting methodologies like the open archives initiative protocol for metadata harvesting (oai-pmh, 2014) to collect metadata records from partner repositories. this would provide a standardized interface and allow the integration of existing and diverse technologies across the partnership into the new common catalog. we were very fortunate that our partners at harvard iqss (iqss, 2014) were already developing open source archive technology that used these standards and thus had experience in this area. the odum institute took advantage of harvard’s success in building and implementing their virtual archiving platform and became one of the first outside implementers of what is today the sophisticated dataverse network, dvn (crosas, 2011). because the data-pass common catalog design was platform agnostic, partners not having adopted the dvn could still contribute to the catalog via a simple oai-pmh interface server. this low barrier of technical entry was essential to the success of the partnership. of utmost importance to the success of the data-pass common catalog was the standardization of each partner’s metadata— made feasible by the groundbreaking work of sue dodd many years earlier. the odum institute during this period was also migrating marc records to the new ddi standard and looking for a replacement for the soon-to-be outdated version of the stanford public information retrieval system (spires) database (spires, 2014). this is not to say it was without challenges, but working with our partners at harvard we were able to complete the migration of the metadata and ingest the contents of the odum archive into our newly built dataverse network. aside from the odum institute’s collection, the data-pass partnership brought together a wide collection of social science catalog records from six major u.s.-based social science repositories and for the first time allowed discovery of these social science data from a single common catalog. the dataverse network has since expanded to include collections from all over the world and continues to grow daily. the power of a standardized metadata catalog record has been exploited to provide discoverability for a vast amount of social science data worldwide. replicating and preserving the catalog records of our joint institutions was an important first step, but our partnership also sought to create a distributed preservation system that our partners could leverage to provide geographically distributed preservation for the group. collaboration for preservation the formation of the data-pass partnership established the groundwork for a distributed preservation project. building on the success of the union catalog, the group identified the need to distribute our joint content in addition to our metadata as a means to better protect our data--despite disparities in repository size and resources among the partners. this is a challenge many organizations face. the expense of maintaining multiple machine rooms and backup systems in multiple geographic regions is prohibitive for small to mid-sized repositories. i would argue that it is equally a burden on larger repositories that would rather spend their ever shrinking resources in more fulfilling areas. both of these circumstances were present among data-pass partners, which created the need for our preservation system to deal with the asymmetrical size of the collections (altman et al. 2009). rather than reinvent the wheel, we decided to borrow from the work of other ndiipp partners working in this space. the metaarchive (metaarchive, 2014) project had been working on defining private lockss networks (pln) to adopt solutions already implemented at stanford university (lockss, 2014). we were able iassist quarterly 2013 59 iassist quarterly to build our preservation network using tested strategies. we had additional challenges along the way due to our content types and sizes but by leveraging the work of fellow ndiipp partners we were better prepared to tackle these challenges. the asymmetrical nature of our pln layered additional challenges on top of our more distributed administration approach. each partner had primary responsibility for running their independent lockss node, and because we had no one central administrator for the network, it was essential that we developed a reporting structure that would generate audit reports of the network. these tools did not exist, so we sought additional funding to build auditing tools for our pln. trust but verify data-pass members needed the ability to audit the new preservation network if it were to demonstrate compliance with standards for trustworthy repositories. the members all had diverse plans for preservation of content in place already, but the addition of a remote copy of each repository under the administration of other members is something that not only needed legal policies in place, but also the ability to audit the performance of the network. this prompted the design of the asymmetrical audit system prototype developed by data-pass during the ndiipp project extensions (altman et al., 2009). follow up funding from the institute for museums and library services (imls) had allowed the prototype to mature into the current open-source offering, the safearchive audit system (safearchive, 2012). utilizing the trac audit framework (crl, 2007) allows the safearchive to enable a pln to define preservation policies in both qualitative and quantitative means. these user-defined policies are stored in a schematized xml format and used to compare the actual performance of the lockss pln to their policies. the result is an audit report that can be provided for each of the members on the status of their content as it compares to the preservation policies they have specified. data management services it seems that today we are living in the “age of data management enlightenment.” everywhere you turn governments, funders, publishers and research institutions are seeking assistance for data sharing, data management, and data science (ostp, 2014). as i reflect on my time here at the odum institute, i want to scream “social science archives already do this!” when i calm down, i am thankful that the early work on mrdf has positioned the social science data community at the forefront of modern data management. our community is comprised of many individuals like sue dodd who have the insight, ingenuity, and enthusiasm to contribute to new initiatives. joint efforts to adopt a common metadata standard like the mrdf metadata grandchild, ddi, along with sophisticated approaches for handling confidential data and the experience of building partnership for preservation and access of complex data files, all provide a wealth of expertise as our society embraces open data policies (data transparency, 2013) and builds massive indexes of health-related data (nih, 2014). we should embrace new partnerships with libraries and library educators as they tackle the monumental task of managing a research output that is growing exponentially. the odum institute is currently working with the unc school of information and library science and the unc libraries in a joint effort to design data management curricula that are flexible enough to be delivered as online content via a massive open online course “mooc” yet grounded enough to allow students to develop local support networks within their own institutions. as a result of this new curating research assets and data using lifecycle education (cradle, 2014) grant, we hope students will share their new knowledge and experiences as they enhance their local data management networks. the international social science data community is graced with many great organizations that are working to educate researchers on proper data management practices. the current data-sharing climate has prompted the research community to seek these services around the world. journal publishers are encouraging and in some cases are requiring authors to submit data supporting their findings alongside their manuscripts. this push toward such a replication data requirement will provide a solid foundation for future scientific discovery as new research is designed around previous discoveries. the dataverse network is working with journal publishers to help satisfy this new requirement. efforts like the open source journal (osj) deposit api for the dataverse seek to streamline and simplify this process for the authors and publishers (ojs, 2014). where do we go from here: hello, big data in the spirit of the 1982 dodd manual (dodd, 1982), “which direction do we go from here?” i hereby declare that the social science data community has come a long way in standardizing data descriptions to make data accessible and understandable for secondary use. pausing to reflect on our past is indeed a worthy exercise, but we should not rest in our efforts to seek improvements for managing the growing collections of data under our stewardship. sue dodd would not be surprised that today the data we are entrusted to are increasingly larger and more diverse than those that came before them. the need for tools and services to visualize and analyze new data types has never been greater. social science is becoming more and more interdisciplinary, and the community will be facing more complex and larger data types like those used in social network analysis and mixed methods studies. relationships between social science datasets will become increasingly complex and require a complex object model to describe. this is not a revolutionary notion, and new standards like ddi version 3 are already designed to handle these relationships (ddi, 2014). the challenge will be to integrate these new models into large preexisting relationship among data sets within archives. as the sheer volume of research data becomes much more massive, we will be forced to seek the council of those in other disciplines that have become accustomed to handling dataset in the petabyte range. the odum institute has begun working with the data intensive cyber environment “dice” group to begin leveraging the irods (irods, 2014) rules-based grid system. tools like irods that have the ability to manage multi-petabyte collections and apply active policies will be needed as we begin embracing the new and larger data formats in the future. we have initiated the integration of the dataverse network and irods that seeks to provide data archiving at scale and allow the federated dataverse network access to discover the massive amounts of data existing in data grids around the world. through our work on the national science foundation datanet data federation consortium project we hope to link diverse communities of data users ranging from oceanographic and hydrologic disciplines to temporal dynamics and plant genomics communities (dfc, 2012). as social science researchers are encouraged or required to share their data, we must always remember our dedication to protecting human subjects. this will require archives to provide tools to 60 iassist quarterly 2013 iassist quarterly assist in this process. the odum institute is closely monitoring the progress of and learning from projects like the data privacy center at harvard’s data tags initiative, which will be critical in providing new tools to share these data while protecting our human subjects (privacy tools, 2014). we should also seek to partner with computer science and data science initiatives like the national consortium for data science (ncds, 2014) to better understand our security risks and provide input into the next generation of secure data transmission systems. data volumes are almost guaranteed to increase exponentially into the future. the social science data community will need to leverage as much as possible automated metadata generation technologies to help reduce the burden on depositors and archive staff. automated ingest tools that create variable level metadata are already being deployed in tools such as the dataverse network. projects such as the nsf-funded databridge project (rajasekar et al., 2013) seeks to use sociometric analysis techniques used in social networking to help determine relationship between users, data, and methods. these relationships could be used to produce multilevel object relationship models to aid in data discovery and population of ddi 3 object relationship models. tools like these, combined with advanced commercial indexing of datasets, will be important to the sustainability of data sharing. if the odum institute and other organizations dealing with data are to contribute to sue dodd’s legacy, we must recognize that the complex problems we contend with today often warrant complex solutions. these are solutions that likely cannot be generated by any one individual or organization alone. members of the social science data community must be willing to reach out beyond their own walls to forge partnerships that take full advantage of the vast amounts of talent that are dispersed throughout our community. to answer sue dodd’s question today, “where do we go from here?” i would suggest that “wherever we go, we go together.” conclusion the social science data community has made great advancements over the years since sue dodd and others began defining bibliographic control over computer information in the late 1970s. i have been fortunate during the past ten years to work with wonderful collaborators and colleagues to build on the work of early iassist members. our community has been fortunate to have strong foundations that have placed us ahead of the game. the social sciences are becoming increasingly interdisciplinary, and we should make every effort to help other disciplines that could learn from our experiences. sharing knowledge will enhance our ability to deliver quality data management and archiving to the diverse social science researchers we will encounter in the near future. building new relationships takes valuable time and effort, but the rewards are great. we have a wonderful data community and we should promote open exchange of knowledge to other disciplines. social science data specialists are seeing the demand for our assistance increase exponentially. our workflows are becoming increasingly complex with the introduction of innovative data formats and expanding data sizes. today’s modern services and tools for managing the outputs from social science research are grounded in the early works of iassist members like sue dodd. without these tools, we would not be equipped to handle our growing set of responsibilities as data stewards. new challenges for the social science data community evolve everyday. as we design services to address these needs, we should encourage new collaborations, encourage open exchange of knowledge, and build on past experiences. the social science data community has tremendous knowledge and experience in its ranks. we should share these with the world. acknowledgements i first would like to thank peggy adams and libbie stephenson for asking me to reflect on some of the current projects that are building on the early work of sue dodd. it would be overwhelming to mention all of the exceptional colleagues i have worked with on these projects but i feel compelled to mention a few. an overwhelming thank you goes to the odum institute that has encouraged me to build these relationships. working with the world class odum institute archives and information science staff has been a great pleasure. dr. kenneth bollen and dr. myron gutmann allowed me to work with the fabulous partners in the early days of data-pass. i owe a tremendous debt to these two great researchers and dear friends. i also would like to thank the current odum institute director dr. thomas carsey for his support and assistance as we continue to support archive development and data management services. a personal thank you also goes to our deputy director peter leousis and our current data archivist thu-mai christian for all the support. the projects i have described are the products of many great relationships and wonderful partner organizations. without the collaborations within the data-pass partners few of these projects would have been possible. many other great partners across the social science data community also deserve mention: educopia, california digital library, lockss, dice, renci, unc school of information and library science, unc libraries, and the australian data archive have been valuable partners in our efforts. none of this would have been possible without great funders. many thanks go to the library of congress, the institute for museums and library services, the national science foundation, the national institutes for health, and the sloan foundation. finally, i wish to thank the wonderful iassist community. we have a great history and community on which to build. references altman, et. al (2009). altman, m., beecher, b., crabtree, j., andreev, l., bachman, e., buchbinder, a., burling, s., king, p., & maynard, m. “a prototype platform for policy-based archival replication” against the grain 21(2). forthcoming. blank, grant, and karsten boye rasmussen (2004). “the data documentation initiative: the value and significance of a worldwide standard.” social science computer review (22): 307–318. doi:10.1177/0894439304263144. crabtree, jonathan and darrell donakowski (2006). “building relationships: ‘a foundation for digital archives.’” accessed on february 6, 2012, . cradle (2014). “dr. helen tibbo receives imls grant for cradle project | sils.unc.edu.” accessed february 2, 2014, . iassist quarterly 2013 61 iassist quarterly crl (2007). center for research libraries (crl), oclc. “trustworthy repositories audit & certification: criteria and checklist.” . crosas, mercè (2011). “the dataverse network®: an open-source application for sharing, discovering and preserving data.” d-lib magazine 17 (1/2) (february). doi:10.1045/january2011-crosas. . data-pass (2014). “about data-pass | data-pass.” accessed february 2, 2014, . data transparency (2013). “the future of open data policy.” accessed february 2, 2014, . ddi (2014). “welcome to the data documentation initiative | ddi data documentation initiative.” accessed february 2, 2014, . dfc (2012). “datanet federation consortium.” accessed june 12, 2012, . dodd, sue a. (1982). cataloging machine-readable files: an interpretative manual. chicago: american library association. iqss (2014). accessed february 2, 2014, . irods (2014). accessed february 2, 2014, library of congress (2014). “digital preservation (library of congress).” accessed february 2, 2014, lockss (2014). accessed february 2, 2014, marc (2014). “marc standards (network development and marc standards office, library of congress).” accessed february 2, 2014, metaarchive (2014). accessed february 2, 2014, . ncds (2014). “national consortium for data science.” accessed february 2, 2014, . nih (2014). “rfa-hl-14-031: development of an nih data discovery index coordination consortium (u24).” accessed february 2, 2014. . oai-pmh (2014). “open archives initiative protocol for metadata harvesting.” accessed february 2, 2014, . odum (2014). “the odum institute: advancing social science teaching and research.” accessed february 2, 2014, ojs (2014). “and so it begins: ojs dataverse plugin testing | pkp-dataverse integration project.” accessed february 2, 2014, . ostp (2014). “ostp announces plans to increase access to federally funded research.” accessed february 2, 2014, . pln (2014). “private lockss networks.” accessed february 2, 2014. . privacy tools (2014). “project description | privacy tools for sharing research data.” accessed february 2, 2014, . rajasekar, et.al. (2013), rajasekar, a, h. kum, m. crosas, j. crabtree, s. sankaran, h. lander, t. carsey, g. king, and j. zhan. . “the databridge,” science journal. ase (in press). safearchive (2012). accessed may 17, 2012, . spires (2014). “about spires.” accessed february 2, 2014, . notes 1. jonathan d. crabtree is the assistant director for information technology and archival research for the odum institute for research in social science at unc chapel hill and can be reached at jonathan_crabtree@unc.edu. the canadian experience witli post censal surveys by adele d. furrie , program manger post-censal surveys program and craig mckie ' editor canadian social trends, sttatistics canada abstract the canadian experience with post-censal survey is, at this point in time, limited to two surveys one which used the 1971 census returns as the sampling frame, and the 1986 survey, which used the 1986 returns as the sampling frame, utilized the census field organization to collect the data, and used the 1986 census data to supplement the data collected 'n the post-censal survey. post-censal surveys provide efficiencies in terms of overall costs because of the accessibility of the census data to identify relatively rare populations. the availability of a field organization to a) select the sample immediately following the collection of the census data and b) collect the data reduces the overall cost to the post-censal program since the hiring and some of the training costs are absorbed by the census. the availability of the census data, not only for the population of interest, but for the data base. respondent burden is reduced because many of the demographic and socioeconomic variables are included in the census. this is true also for family and household-related data. some consideration is being given to the conduct of a similar post-censal disabihty survey following the 1991 census. in addition, there is the possibility for a survey on the senior population and one on aboriginal persons. the canadian experience with post-censal surveys is limited to two surveys the highly qualified manpower survey (hqms) conducted following the 1986 census. while the methodology of the two surveys differed significantly, each provides an interesting application of post-censal methodology which could be considered by those who are planning such activities. the methodology for each survey will be described and an evaluation will abe provided of the methodology. the highly qualified manpower survey hqms was conducted in the fall of 1973, just over two years after the 1971 census was conducted. the objective of the survey was to assess past expenditures for education in relation to the utilized labour force status. the data were needed to assist in the formulation of policy relating to l(mig-range planning in the fields of education and manpower planning. the sample was selected from the census database based on the individual's age-sex-labour force status and level of post-secondary attainment. because name and address were not part of the census database, the original census questionnaires had to be accessed, and information such as telephone number and the name of the head of the household as well as the name and address of the selected person was transcribed. quality control procedures for the transcription and creation of the name and address file were employed to control for errors. the questionnaire covered a limited range of topics such as field of study, current labour force status and current earnings, and an employment profile over time. these data were supplemented by the information collected in the 1971 census, so that the combined database covered a wide range of topics, such as ethnicity, immigration status, marital status, etc. the survey questionnaire was mailed to approximately 138,000 selected persons (the number of persons meeting the selection criteria were approximately 720,000); about 72,000 or ahnost 70% returned the completed questionnaire. a tracing operation was conducted to establish the current address fw those selected persons who had moved. follow-up of non-respondents was conducted by mail, and in some instances, by personal visit, during the period september, 1973 to march, 1974. because the survey was conducted over two years after the census, there was some difficulty in locating some of the selected persons and some of the data retrieved from the census were out-of-date. for example, the individual may have married or additional children may have been bom, and these differences would not have been accounted for in the combined database. the 1986 post-censal survey, the health and activity limitation survey (referred to in this paper as hals) employed a different collection methodology, taking into account the experience of the hqms. hals was comprised of three distinct surveys two household surveys both of which used the disability question on the census as the screen to identify the sample and an institutions survey which used the census to identify the location, size and type of the institution. the need for a comprehensive database on disabled persons was articulated in 1981 in the report entitled obstacles, the report from the special parliamentary committee on the disabled and the handicapped. it spring 1990 noted that there were no national data available on disabled persons. not only were there no national estimates on the number of disabled persons in canada and the nature and severity of their disability, but, as important, there were no estimates on the barriers which these disabled canadians face in the conduct of their everyday activities. because many of the programs and services offered to disabled persons are the responsibility of provincial and local governments, it was important that these data be available for relatively small geographic areas. as programs and services differ for different age groups, d^e sample would also have to be large enough to be able to generate estimates within the geographic areas for each age group. the methodology of a post-censal survey was considered as the only viable alternative to obtain this level of detail. other options such as the use of existing survey vehicles the monthly labour force survey or the annual general social survey were considered but it was determined that neither would yield a sufficiently large sample of disabled persons. there were also limitations in the coverage for both of these survey vehicles. the labour force survey excludes the more remote areas of canada, residents of indian reserves, and residents of institutions. the general social survey generaoy utilizes a randomdigit dialing methodology to create a sample, therefore, households without telephones would be excluded from the survey. the general social survey is also household based so that residents of institutions would be excluded from the survey. the decision to use the census as the method to identify the samples for the household surveys necessitated the inclusion of a disability question on the census questionnaire. it was decided that this question would be added to the "long" questionnaire the one completed by one out of every five households. the question asked if the individual was umited in the kind or amount of activity he/she could undertake because of a health problem or condition. the second part of the question asked if the individual had any long-term disabilities or handicaps. households were advised through the guide that was included with their census questionnaire that the disability question was to be used to identify a population for a more-intensive survey on the issues facing disabled persons. it was determined through a pre-test that this question would identify most of lhe more-severely disabled population, but that additional questions would be required to identify all disabled persons. a copy of the census disability question is included in appendix a of this paper. the content of the hals questionnaire was determined through extensive consultation with representatives from, government departments that provided programs for and services lo disabled persons. consultation with organizations of and for disabled persons was also undertaken to ensure that their needs were reflected in the content. the questions used to identify the nature and severity of the individual's disability were, for the most part, developed by the o.e.c.d. these questions are known as the activities of daily living and were developed to identify physical and sensory disabilities. other questions were added to identify emotional, psychological and learning disabilities and persons who are developmental!y delayed. with the inclusion of the disability question on the census, the sampling frame was in place for the postcensal survey of disabled persons. to maximize the efficiency of this sampling frame, it was decided that an operation should be integrated with census processing to select a sample of individuals who had responded "yes" to the census disability question. this would allow for the conduct of the post-censal survey shortly after the census, thus minimizing the follow-up required because of inter residence moves taking place between the lime of the census and the conduct of the post-censal survey. it would also enable the utilization of the census field staff to conduct the face-to-face interviews. to accommodate the selection of the sample and to ensure that the field staff would be available for further work, geographic areas were identified prior to the conduct of the census. these areas were defined as the geographic area within which the workload for one census field staff (census representative) was located. census staff received additional training on the concepts and definitions used in hals and the face-to-face interviews were conducted immediately following the census collection. the reference day for the 1986 census was june 1. in most instances, the interviews were completed during august and september, 1986. there were an estimated 120,000 individuals selected for the follow-up interview; the overall response rate was in excess of 95%. the second household sample involved a sample of individuals who had responded "no" to the census disability question. this sample was necessary because the pre-test had indicated that some disabled persons may not respond affirmatively to the general disability question. a sample of approximately 80,000 individuals was selected during a later stage in the census processing, but before the census documents were returned to head office in ottawa. the same questionnaire was used and most respondents were contacted by telephone. those who did not provide telephone numbers on their census questionnaire were contacted in person. the survey was conducted from the regional offices of statistics canada during october and november, 1986 by interviewers who are part of the regular regional office staff. approximately 90% of the sample was contacted and agreed to participate in the survey. the data from, both household surveys was integrated with the census at the micro-record level, so that the linked database contains information from both the census and hals. the census data, for die most part, is for the selected individual, but included in the base are also some variables about the family and household within which the selected person resides. because the census and hals were conducted within six months of each other, the variables taken from the census such as marital status should not have changed significantly. 14 assist quarteriy another feature which adds to the richness of this database results from the sample being selected to represent both "yes" and "no" respondents to the census disability question. the hals sample can be divided into strata those who are disabled and those who are not for the non-disabled population, the data available on the linked database includes all of the census variables. for the disabled peculation, the data includes both census and hals variables. this affords the user the opportunity to make comparisons of the characteristics of the disabled and the non-disabled populations. the census methodology did not include the use of the long questionnaire in institutions; therefore, the disability question was not asked in institutions. to obtain information for residents of institutions, the census was used to identify the location, size and type of institution. penal institutions and correctional facilities were excluded because of operational difficulties. from the remaining list, a sample of institutions was selected, approximately 1,100 out of a total of approximately 5,300. each of the selected institutions provided a list of residents from which a sample was selected. of the 18,200 residents selected from the 1,100 institutions, less than 3% refused to participate in the survey. a personal interview was conducted with slightly over 50% of the sample. for the remaining sample of respondents whom the institution administrator deemed to be too ill or too disabled, the interview was conducted with an individual who provided the day-to-day care. data from hals was released in may, 1988 and has been used by both the public and private sectors. much has been learned concerning the conduct of post-censal surveys and the integration required with the census operations. planning is now underway for post-censal survey activity following the 1991 census. based on consultation with representatives involved in social programs, three potential topics have emerged and further consultation is now underway. the three topic areas are a survey of seniors with the focus on support networks, a repeat of the survey of disabled persons so that data are available over time, and a survey of aboriginal persons, both onand off-reserves. the possibility of one or more of these topics going forward is contingent on obtaining funding for them. the 1986 survey cost seven million dollars. that survey proved that a post-censal survey, closely linked to the census operation in terms of identifying the sample, utilizing the census field staff, and the census data is a viable option for surveys of relatively rare populations or which require significant geographic detail. q appendix a 20. (a) are you limited in the kind or amount of activity that you can do because of a long-term physical condition, mental condition or health problem: at home? no, i am not limited yes, i am limited at school or at work? no, i am not limited yes, i am limited not applicable in other activities, e.g. transportation to or from work, leisure time activities? no, i am not limited yes, i am limited (b) do you have any long-term disabilities or handicaps? no yes ' presented at the ifdo/iassist 89 conference held in jerusalem, israel, may 15-18, 1989. spring 1990 vol21.1 15spring/summer 1994 introduction the internet has awesome potentiality as a global network of information and the world wide web provides a potentially efficient and effective protocol for the delivery of information. the content of the resources made available are often of particular interest to those in university settings engaged in teaching social sciences. technologies will most probably continue to be innovative and technical advances will increase access to these resources. whilst this will be welcomed progress it also raises a series of issues and problems especially with regards to using internet based resources in social science university curricula. in this paper i will argue that whilst the use of internet resources is potentially good for university teaching and learning, the development of these resources has often been rather haphazard. the driving forces behind the development of these resources have been either from enthusiasts or technical specialists. my argument is that the internet is too precious and important a resource to be allowed to develop in such an unmanaged fashion. i will be advancing an argument that calls for the incorporation of non–technical knowledge in development of internet based resources that are appropriate for use in university social science teaching and learning environments. the ideas that i intend to express are to a large extent exploratory and intended to be evocative rather than the last word on the subject. my arguments are initially premised upon the conception that the internet can be considered simply as a piece of information technology. the internet is a global network of information which provides a potentially efficient and effective protocol for the delivery of information. in this regard the internet is special but if we consider the popular image of the internet as a ‘super highway’ then it is possible to think of it as a giant communications technology. whilst the internet has some specific features i wish to contend that if it is simply considered as a communications technology the issues that pervade its use in social science teaching also pervade the use of other information technologies in this area. given this, more general arguments about information technology and social science teaching also apply to internet based resources. the incorporation of information technology into teaching settings generally proceeds from a ‘technology is obviously a good thing’ approach. in some quarters this is considered as axiomatic. whilst i am sympathetic to the idea of incorporating technology into higher education teaching and learning it is the abandoning of this axiom that is essential if information technology is to be successful. the motivation behind information technology in the university setting has been from, enthusiasts on the one hand and technical specialists on the other. i will refer to this as the ‘technologist perspective’. my argument is that the ‘technology is obviously a good thing’ approach is a flawed departure point. due to the hegemonic domination of technical expertise held within the technologist perspective the design and implementation of new information technologies in teaching and learning environments will have very limited success. i will be advancing an argument against the ‘technologist’ approach that calls for the incorporation of non-technical knowledge in technological developments in university teaching and learning. this will draw upon some of the advances that have come out of computer supported cooperative work (cscw) approaches. i will also argue that a particular theoretical conception of ‘work’ and of the role of social science in technological design and implementation is appropriate. technology and teaching social science part of my disquiet with the technologist perspective’s approach to the design and implementation of new information technology in teaching and learning environments is the generality of the argument that technology is necessarily a good thing. the lack of specificity in this approach engenders a poor understanding of the particularities of university teaching and learning. in much the same way it would be easy to talk quite generally about technology and university teaching and learning, but to avoid this pitfall i will confine this discussion to the social sciences in particular. in the social sciences, students’ experiences of information technology will mostly be in the form of desk top computing of the pc or macintosh variety. their introduction to computing will often form part of research developing internet based teaching and learning resources for social sciences: a cscw approach. by vernon j. gayle* 16 iassist quarterly methods or study skills. the role of the computer in this instance is often not to deliver computer based or assisted learning but rather software and hardware are used as tools to undertake tasks rather than as learning technologies. in the british context, despite funding being directed towards computer based or assisted learning, in the social sciences there has been an absence of completed and useable bespoke software packages. the failure of these endeavours is evident insofar as there are few examples that are routinely incorporated into social science curricula. these bespoke software packages are not in widespread use in departments in british higher education! the alternative to bespoke software is the re-use of existing technology. in these endeavours software, and to a much lesser extent hardware, are directed toward a teaching and learning requirement in an attempt to add value to a specific learning experience. from my own experience and that of colleagues, these attempts to re-use existing technology are at best problematic. the re-used existing technology, is generally attempted on an ad hoc basis and its development is time consuming, and labour intensive. in the majority of cases the development has not proceeded from a clear pedagogical requirement and the end products lack the sophistication required to deliver the high quality educational experience that is the hall mark of universities. the ‘value added’ nature of these endeavours is not necessarily tractable, especially when cost is entered into the equation. despite the huge potentiality of technology for social science teaching and learning, attempts to introduce bespoke software and to re-use existing technology are ill conceived. this is due to the taken for granted assumptions about the actual nature of the teaching and learning environment and the lack of a comprehensive empirical understanding of the processes that are in motion. this in turn leads to a shallow understanding of the consequentiality of the introduction of new computer based technologies in social science teaching. it is highly likely that a similar unsatisfactory situation could arise with internet based resources. the internet offers social science teachers and students a global network of information resources. internet based resources are potentially good for university teaching and learning but the development of these resources has often been rather haphazard. once again the driving forces behind the development of these resources have been either from enthusiasts or technical specialists. the internet is too precious and important a resource to be allowed to develop in such an unmanaged fashion. in the world of commerce and industry many technical endeavours which are based around personal computing have not furnished adequate results. this has been due the technology failing to pay sensitive account to what i shall term as the ‘innate sociality’ of the environments into which they are being introduced. this parallels the situation in social science teaching in universities. one solution to the problem of new technology and the workplace has been the development of computer supported cooperative work (cscw) as a design paradigm. i do not wish to argue that the work environment in the world of commerce and industry is the same as it is in higher education, although there are some obvious similarities at a generic level. i maintain that a particular theoretical sociological conception of ‘work’ can inform cscw design and this is appropriate to the design and implementation of computer based learning technologies in general and this extends to the development of internet based resources for social science teaching. in the next part of the paper i will introduce the idea of computer supported cooperative work (cscw). the material relates to the development of technology more generally and is not restricted to either teaching and learning or the internet. this will provide a context within which my position can later be developed. the problem of human computer interaction fundamental to an understanding of the propriety of cscw is an appreciation of the problems of a human computer interaction (hci) approach. hci was a new and radical approach to systems design that achieved prominence in the 1980’s and it sought to provide a better cognitive coupling between human users and computers (bannon 1989). cscw can reasonably be considered as a response to the failings of human computer interaction (hci) approaches. hci is a general framework for innovation aimed at developing interaction techniques, analysis methods, software and computer systems within a controlled context in order to create enhanced products. hci endeavoured to go beyond simply providing improved surface characteristics, and hoped to address wider issues surrounding human interaction with computers. in this sense hci is a design and engineering science as it aims to produce artefacts of hardware and software within satisfactory frameworks of compromise that take functionality, performance and cost into account (brooks 1990). the hci perspective recognised that there were human consequences to the introduction of new technology, and that how technology was developed and applied could profoundly affect work. with regard to the introduction of new technology questions of usability, applicability and acceptability were being raised. from within the hci camp these issues were viewed as being of legitimate concern. what was considered necessary was an applied psychological dimension located within a problem centred approach which would enable hci practitioners to 17winter 1998 undertake research that would inform future designs1 (blacker & osborne 1987). despite hci being a radical new initiative the organisation of work is in fact endlessly richer and more complex than the majority of formal psychological models could have conveyed. due to the rigid frameworks that such systems imposed, human actors were not furnished with sufficient flexibility to make the system function (bannon 1989). another draw back which stems from the psychological foundation is that hci fashioned itself as a general paradigm for innovation and design in limited and controlled environments (brooks 1990). much of the early hci work was confined to rather small scale controlled experiments with the presumption that the findings could be generalised to other settings (barnard and grudin 1988). the hoped for contribution of hci to the design of computer systems and novel interfaces did not materialised in the 1980’s (carroll 1987). gray and atwood (1988) note the lack of examples of developed hci systems. this is largely due to the inherent deficiencies in hci approaches and what is required is an alternative theoretical and methodological orientation (luff & heath 1990). computer supported cooperative work the expression cscw is a comparatively new one in the information technology vocabulary and was first coined in the mid 1980’s by researchers in the usa. the term was most notably used by greif (1988) and has been applied as an umbrella term which takes in anything to do with computer support for activities involving more than one person. an alternative terminology to cscw includes the expressions ‘groupware’ and ‘workgroup computing’ (clark and o’donnell 1991). cscw is very much a generic term. bannon (1990) argues that despite disagreements about specific detail most cscw practitioners would agree with lyytinen (1989) who asserts that cscw is an attempt to place emphasis on both the distinctive qualities of cooperative work processes and on questions of systems design. cscw takes from its point of departure the visibly processual character of social activity (harper & randall 1992). the organisation of situated action is an emergent property of the interactions between actors and their environments (suchman 1987). settings engender a specificity unique to their social organisation. sociological inquiry within cscw must not be ad hoc or abstract and divorced from examination of the specificity of the setting. we must attempt to study settings and explore the coordinated tasks that computers might support in the context and settings which they occur (luff and heath 1990). it is fundamental to examine the natural settings where tasks and activities exhibit their sociality (bannon & schmidt 1991). the role of an effective cscw system as its name suggests is to support the cooperative nature of work. hirscheim and klien (1989) assert that the good system must not be designed in what they term as the ‘usual sense’, but has to be designed and developed within the framework of the social interactions that are embedded in the environment in which the technology is to be incorporated. the caveat that must be issued here is that in no sense is there an objective set of criteria that form a typology for an effective system. the system must be developed within what they term as the ‘user’s perspective’. a cscw system is not however simply an electronic cloning or duplication of a working environment. in contrast it is a pragmatic attempt to support the cooperative tasks of work in context within its natural social and physical environment. in terms of sociological inquiry, ethnography is the tool of sociological research most applicable in cscw endeavours. as with all research methods ethnography has advantages and disadvantages but the potency, in the cscw context, is that it depicts the activities of social actors from their own perspective. this challenges the preconceptions that alternative social science approaches often bring to phenomena (hammersley and atkinson 1983). ethnography is not simply description, rather, it is detailed explication. it is about capturing the real movement of experience in the concrete world. ethnography achieves something which theory and commentary in the majority of cases cannot namely it presents human experience without minimising it and without making it a passive reflex of structures, organisations and social conditions. ‘human productions are all of a piece, indivisible and always summed. the metal cannot be simply smelted out from the ore of experience in human affairs’ (willis 1978 p. 180). the use of ethnography is an attempt to ground the understanding of action in empirical evidence. an empirically based social interaction perspective is inductive from particular naturally occurring activities. ultimately, this will produce descriptions that are accountable to evidence. situation is crucial to the interpretation of actions but although this is fundamental and to some extent obvious, its importance could easily be overlooked. sociologists have long used ethnographic techniques to study work in general.2 what is required in cscw is an ethnographic analysis of settings that are due to be ‘technologized’, by which i mean where new forms of information technology are to be implemented. in the case under consideration the settings are where social science teaching and learning takes place. straightforward ethnographies, such as those developed in the sociology of work, are not sufficient however. what is required is what button and dourish (1996) term as a ‘technomethodologically’ informed approach. this i believe will lead to the successful design and 18 iassist quarterly implementation of new information technologies in social science teaching settings. the technomethodological approach departs from the desire to make conspicuous what actors are doing when they organise the activities that they do in particular settings. it draws heavily upon ethnomethodology which turns away from the structures and theorising of traditional sociology, and concentrates instead upon the details of the practices through which action and interaction are accomplished (button and dourish 1996). ethnomethodology poses the question ‘what the devil does the native think they are up to?’ (anderson, hughes and sharrock 1989). in the case of cscw the native is the individual in the setting about to be technologized. examining settings by technomethodologically informed ethnography is fundamental to the exercise of cscw if it is to be applied successfully to teaching and learning settings. the technomethodological approach is underwritten by the work of garfinkel3 who is concerned with that most pervasive sociological question; ‘how is it that actions recur and reproduce themselves?’ he insists that this orderliness be viewed as arising from within activities themselves and the work being done by the parties to that activity. garfinkel eschews the traditional sociological strategy of seeking to explain this orderliness and the organisation of social activities by attempting to identify causes and conditions out with the activities themselves (benson & hughes 1983). germane to this are garfinkel’s suggestions that it is evident from the availability of empirical specifics that there exists a locally produced order of work’s things that make up the enormous domain of organizational phenomena.4 he argues that the classical sociological studies of work, without remedy or alternative, depend upon these phenomena, make use of the domain and ignore it. that the domain is ignored is a systematic feature of the locally produced orderliness of work settings. therefore the reported phenomena are only inspectably the case, therefore they are unavailable to the art of designing and interpreting definitions, metaphors, models, constructions, types or ideals and cannot be recovered by attempts, no matter how thoughtful to specify an examinable practice by detailing a generality. the concept of the ‘egological organisation’ is advanced by anderson, hughes and sharrock (1989) and is an ethnomethodologically informed view of organisations that begins with a bottom-up understanding of them5. the conception of the egological organisation departs from an inquiry into the daily or routinized experiences of individuals. the value of this approach is that it places the actor’s point of view at the centre of the analysis. by employing the concept of the egological organisation, they develop the idea of the ‘working division of labour’. in working for profit they argue that it ought to come as no surprise that actors in work settings see themselves as part of an elaborate working division of labour. from the way that they talk about their work, both to each other and to outsiders it is clear that the notion of a working division of labour is one which they use to interrelate and explicate the things that they see going on about them, on a daily basis and on ordinary occasions. these accounts depict a body of activities marshalled by ‘a working division of labour’. technomethodology requires a sociological analysis of the organisation of social action and interaction and the organisation of work and work settings. the fullest possible description that captures the essence of the ‘working division of labour’ must be furnished. the thrust of technomethodology is the conception that sociological descriptions of the ways in which people routinely organize their actions and interactions can be furnished and compared to what is or is not possible using technology. in this sense the term ‘technomethodology’ is an identification of the need for the incorporation of ethnomethodologically informed accounts of the working division of labour that places the actor’s perspective at the centre. the object of the ethnographic exercise is therefore to provide what gertz (1975) terms as ‘thick descriptions’, which will inform the design and the implementation of the new information technologies.6 cscw can inform the development and implementation of new information technologies that are directed towards teaching and learning environments in the social sciences. i do not wish to argue that the these environments are the same as commercial and industrial settings which so far have been the foci of cscw. at a generic level university teaching and learning settings are similar insofar as they also require cooperation, coordination and collaboration to accomplish work tasks. and whilst the work carried out in university teaching and learning settings is arguably, often of a more individual nature, cscw has attended to the issue of individualistic work.7 if we treat the concept of ‘doing work’ as the active process of ‘sense-making’ that individuals undertake in settings, then the same kinds of issues that impinge upon actors in commercial work settings are also present in university teaching and learning settings. computer supported cooperative work and teaching in the social sciences it is certainly the case in britain that the desire to incorporate technology into higher education teaching and learning has been firmly placed on the higher education agenda. this situation is both desirable and essential to the future development of university education. if in the social sciences we wish to move from using computers as tools, to a scenario where we develop computer based learning 19winter 1998 technologies i believe that it is fundamental to incorporate non-technical knowledge. this will lead to a more circumspect and strategic development of information technology than could be achieved by enthusiasts or technical specialists. the development of internet based teaching and learning resources for the social sciences must proceed from clear sets of pedagogical requirements. this is not to argue that in any sense objective sets of criteria that form typologies for effective resources exist. the resource requirements in various settings will be context specific. the role of a cscw strategy, that is technomethodologically informed, is an attempt to uncover the pedagogical requirements of the internet computer based resource that is being developed. this is the level of sophistication that is required to develop high quality internet based teaching and learning resources that will be useful and used in social science departments. a cscw approach to the design and implementation of internet based resources for social science teaching and learning will be liberating. the need for a clear understanding of teaching and learning environments is critical. technomethodologically informed cscw, when brought to bear upon the design and implementation of internet based resources for social science teaching is an attempt to improve what gurdin (1988) and bannon & harper (1991) term as the ‘distinctly random success’ of new information technologies. bibliography anderson, r., hughes j. a. & sharock, w.w. 1989, working for profit: the social organisation of calculation in an entrepreneurial firm , gower aldershot bannon, l. 1989, ‘from human factors to human actores: the role of psychology and hci studies in systems design’, computing services university college dublin. bannon, l. & bodker, s. 1989, ‘beyond the interface: encountering artifacts in use’, department of computer science aarhus university. bannon, l. & schmidt, k. 1991, ‘cscw: four characters in search of a context’, studies in cscw: theory, practice and design, in j bowers and s benford, elsevier amsterdam. barnard, p. & gurdin, j. 1988, ‘command names’, handbook of human computer interaction, in helander, m. (ed), north holland press amsterdam. beynon, h. 1984, working for ford, penguin london. blackler, f. & osborne, d. 1987, ‘designs for the future: i.t. & people’, british psychology society conference leicester. brooks, r. 1990, ‘the contribution of practitioner case studies to human-computer interaction science’, interaction with computers, vol. ii no.1 april. button, g. & dourish, p. 1996 ‘technomethodology: paradoxes and possibilities’, technical report epc-1996101, rank xerox, cambridge. carroll, j. 1987, interfacing thought: cognitive aspects of human computer interaction, mit press cambridge. carroll, j. 1989, ‘evaluation, description and invention: paradigms for human computer interaction’, advances in computers, vol. 29. clark, b.m. & o’donnell, s. 1991, ‘computer supported co-operative work’, british telecom technical journal, vol 9 no.1 january pg. 47-56. daniel, j. 1993, ‘the challenge of mass higher education’, studies in higher education, vol 18(2) pp.197203. garfinkel, h. 1956, ‘some sociological concepts and methods for psychiatrists’, psychiatric papers, vol 6 pp.181-195. garfinkel, h. 1967, studies in ethnomethodology, prentice hall new jersey. garfinkel , h. 1986, ethnomethodological studies of work,routledge and kegan paul london. garfinkel, h. 1991, ‘respecification’ , in button, g. ethnomethodology and the human sciences, cambridge university press cambridge. gayle, v.j. 1991, ‘cscw and individual work: an investigation of work distribution in a london taxi firm’, unpublished masters thesis, department of sociology university of lancaster. geertz, c. 1975, ‘thick description’ in the interpretation of culture, hutchinson london. gray, w. & atwood, m. 1988, ‘review of interfacing thought’, interfacing thought: cognitive aspects of human computer interaction, in carroll, j. interfacing thought: cognitive aspects of human computer interaction, mit press cambridge. grief, i. 1988, “remarks in panel discussion on ‘cscw: what does it mean?’ cscw ‘88 proceedings of the conference on cscw, sept 26-29 portland oregon. 20 iassist quarterly hammersley, m. & atkinson, p. 1983, ethnography: principals in practice, tavistock london. harper, r. & randall, d. 1992, ‘rogues in the air: an ethnomethodology of ‘conflict’ in socially organised airspace’, technical report epc-1992-109, rank xerox, cambridge. harper, r., hughes, j., randall, d., shapiro, d., sharrock, w. (forthcoming), ordering the skies, routledge. hirscheim, r, & klein, k. 1989, ‘four paradigms of information systems development’, social aspects of computing, vol 32 no. 10 october pg. 1199-1216. hobbes, d. 1988, doing the business: entrepreneurship the working class and detectives in the east end of london, clarendon oxford. luff, p. & heath c. 1990, collaboration and control: the introduction of multimedia technology on london underground”, rank xerox. lyytinen, k. 1989, ‘information systems failures: a survey and classification of empirical litterature’, oxford surveys in information technology, vol.4 pg.257-309. pollert, a. 1981, girls, wives, factory lives, macmillan, london suchman, l. 1987, plans and situated actions: the problem of human machine communication, cambridge university press cambridge. thimbleby, h., anderson, s. & whitten, i. 1990, ‘reflexsive cscw: supporting long-term personal work’, interacting with computers, vol.2 no.3 pg. 330-336. willis, p. 1987, profane culture, routledge london. 1 it is important to note that there is no clear or coherent answer to the question ‘what is, or was, the goal of hci?’. one of the more traditional answers to this question is that hci intends to provide methods and matrices for evaluating the usability of computer systems. this stems from what can be loosely termed the ‘human factors’ approach. this is in contrast to cognitive scientists who argue that hci is a work bench for the application of cognitive psychology to a real problem domain. computer scientists assert that hci must help to guide the definition, invention and introduction of new computing tools and environments. this argument is to some extent a product of the exigencies of the computer industry. the point here is to illustrate that hci is a diverse discipline with fragmented foci and interests, a feature not often drawn out in hci literature. 2 for example beynon (1984) examined working at ford, hobbs (1987) studied detective work in east london and pollert (1981) treated the working lives of women in a factory setting. 3 garfinkel states that ‘the policy is recommended that any social setting be viewed as self-organizing with respect to the intelligible character of its own appearances as either representations of or as evidences-of-a-social-order. any setting organizes its activities to make its properties as an organized environment of practical activities detectable, countable, recordable, reportable, tell-a-story-aboutable, analyzable in short, accountable’ (garfinkel 1967 p.33). 4 see especially garfinkel (1986 & 1991). 5 the conception of the egological organisation is motivated by the desire to provide descriptions and analysis, but raises a deep methodological question. sociological descriptions like other theoretical accounts are thematically constructed. the methodological question at issue in this instance is that employing this approach is an attempt to provide a third person account of first person experience. this does not mean incorporating first person accounts into sociological depictions, rather a sociological re-constitution of that experience is required. the concern is not with particular people’s experience, but with the organisation of experience, as it is encountered in social life, as a readily accountable, known and shared schemes of interpretation (anderson, hughes and sharrock 1989). 6 an example of such an attempt is harper et al (forthcoming) which is an account of technology and air traffic control as part of an inter-disciplinary attempt to design and implement a technological system. 7 my own work on london taxi drivers (gayle 1991) and the work of thimbleby et al (1990) are two examples. * paper presented at the iassist/ifdo 1997 conference, may 6th may 9th odense, denmark. vernon j. gayle, department of applied social science, university of stirling, scotland. email: vg1@stirling.ac.uk mailto:vg1@stirling.ac.uk iassist quarterly vol 24 no. 3 4 iassist quarterly fall 2000 university information system russia: scientific and social challenge by tatyana yudina* an appropriate information base is the main challenge for research and education in social and human sciences in russia, especially in distant areas. due to an information shortage, the general level of teaching and applied investigations is decreasing university professors are unable to recommend as obligatory for study recently published books and periodicals: funding for book purchases is poor. as the ministry of science of rf (the russian federation) recently reported, only 67 scientific journals are available for 10,000 investigators in russia (408 in great britain, 186 in the usa). official government documents and reports, state statistics are also not available for educators and investigators. due to the lack of public domain state statistics, the new research methods based on processing of large sets of numerical data are not developed in russia. in the current situation, internet-based collective resource is not only the most rational but the only possible approach to arrange the information supply and build the research base for investigations and advanced education in russia. the moscow state university (msu) research computing center and non-commercial organization center for information research since 1994 have been working to meet the challenge and develop the university information system russia (uis russia). in january 2000 the uis russia (www.cir.ru) started operating on regular basis as a collective information base providing free access to all russian universities. during 2000-2002 the uis russia will compose an appropriate resource base for full-scale investigations in main human and social sciences. the internet access ensures equal opportunities to researchers and educators in all regions of russia. the universities of the former soviet union (fsu) countries are also granted free access. in 2001 the uis russia is planned to be available to local public libraries in russia and fsu countries. the current version of the uis russia includes the data and documents’ sources recommended as the first priority collections by the center for sociological research and economic faculty of moscow state university: • official data and documents (laws, presidential decrees and directives, governmental enactments, acts and regulations) since 1991; • stenogramms (daily records) of state duma of federal assembly of rf from 1996; • goscomstat of rf data (all available in electronic format); • election statistics of both federal and local levels since 1993, provided by central election commission of rf; • mass media sources (8 newspapers and 2 information agencies); • databases, publications and reports of leading analytical centers; • reference data on the russian political system (brief history, prerogatives, structure and personnel of federal institutions, political parties, churches, etc.); • extended reference information on the components of the russian federation. all data collections are obtained for free from official holders/producers under legal agreements with research computing center of msu. the provision to process the information, integrate the results into the uis russia and provide access to all universities of rf makes the uis russia a valuable resource for full-scale socially relevant investigations. information update full text documents official data and documents, stenogramms, mass media electronic versions are received electronically on daily basis, bulletins and analytical reports on weekly or monthly basis (upon publication). new full text collections will be added in 2000 : • international agreements, signed by rf since 1991, international agreements signed by ussr, • constitutional court of rf, supreme court of rf, arbitrary court of fr, decisions, • commonwealth of independent states countries multilateral and bilateral agreements. • local mass media sources. iassist quarterly fall 2000 5 all the documents are automatically processed – metadata created, classified, indexed, annotated and integrated into the uis russia. the nlp technology provides for 20 mb (equivalent of up to 10,000 pages) processed daily. appropriate retrospective coverage of each source will be realized during 2000. statistical data numeric data are the mostly used resource for social research. state statistics are the basis for socially-relevant investigations and sound recommendations for decision-makers. the uis russia legally obtains, stores and updates collections from state statistical committee of rf. under discussion are agreements with other main statistics-producing government institutions. the goscomstat of rf collections are received upon publication on monthly, quarterly or annually basis. statistical data are received in .doc format (digital versions of publication). to make the data available for internet search, the uis russia specialists convert the data into html format; and as a next step into excel spreadsheet format to make the data usable for secondary analysis. currently available are the following data collections : • russian annual statistical report, 1999, • industry of russia in 1999, • regions of russia in 1999, • national accounts of rf, 1991-1999, • environment protection in russia in 1999, • finance in russia in 1999, • prices in russia in 1999, • russia and commonwealth of independent states countries in 1999, • russian annual demographic report, 1999. the data collections are also indexed to make them searchable using the uis russia thesaurus. in 2000, all other goskomstat of rf collections will be added. for 2000-2001, there are also plans to obtain and integrate the data maintained by the centrobank of rf, the ministry of finance of rf, goscomimyushestvo of rf, the ministry of labor of rf, other ministries, committees and agencies of rf, regional statistics of components of rf, international organizations measurements, and the databases created under the foreign grants. the numeric data collections are complemented by methodological notes. part of the statistics of rf bloc are analytical reports prepared by leading “think tanks” in russia russianeuropean center for economic policy, bureau of economic analysis, fond for population sentiments index research, etc. reports of main russian and foreign foundations’ grantees are included. the documents are also indexed to make the analytical materials retrievable by thesaurus-based cross-search. electoral statistics are received shortly after elections. the current version stores all general elections results since 1993, and local election results. the data are converted into excel spreadsheet format and may be analyzed using standard software packages like statistika, spss, sas, etc. electoral statistics is region-tailored and displayed in a map format. nlp technology to process and integrate large scope of electronic documents the technology of automatic linguistic text processing (altp) is realized under the project. the altp performs: processing of electronic text corpora in main formats (ascii, html, ms word) in windows and operating as dll; morphological analysis of russian texts; terms’ recognition/disambiguation; thematic analysis event categorization, indexing, annotation/summarization; download of results on oracle database server. the main instrument of the technology is the thesaurus on contemporary russia (thesaurus), created under the uis russia project. in its current version it incorporates 18,500 concepts/descriptors, includes 6,500 geographic names, 39,000 synonyms, 70,000 relations between concepts, 200,000 inherited relations. the tool assists in detecting of main and subordinate topics in a document as a result of analysis of macroconcepts and relations between them. macroconcepts are modeled by groups of concepts semantically related in thesaurus. thematic representation provides for evaluation of weigh of each term in a text and performs event categorization and annotating/ summarization of a document. the thesaurus enables to determine up to 90 95 % of terms. the technology provides for up to 20 mb of electronic texts to be processed on each pentium200 pc and integrated into the university information system russia daily. the technology was evaluated by experts from nist and darpa in 1996 under the textretrievalconference-6 program and summarizationconference in 1997. the results are among the best in a group of 14 participants. the alpt ensures advanced search instruments. search engine being initially designed to serve scientific needs the uis russia provides for advanced search instruments: in 6 iassist quarterly fall 2000 addition to traditional tools it includes value-added elements the system of subject headings and thesaurus on contemporary life in russia. the system of subject headings consists of 200 topics (rubriks). all full text sources are filtered and event categorized according to system of subject headings. the congressional research service, lc legislative indexing vocabulary-based search is also available. the uis russia provides for research assistance, user services, metadata and annotation browsing, and thesaurus based query refinement. user-tailored automatic information update is realized. technical base: architecture the uis russia operates at the research computing center of msu server. mirror sites will be maintained in novosibirsk and st. petersburg to ensure more reliable access for universities in the northern part of rf, siberia and far east. analytical bloc educational activity in advanced research methods has been started. main research institutions were contacted and several software programs, workbench of sociologist, workbench of economist, etc., are presented by the authors. special training class is arranged by the research computing center of msu, where the analytical software is downloaded and made available for university faculty. special course is scheduled, it includes lectures analyzing main approaches to computer-based investigations and demonstration of working models, teaching and training in basic and advanced technique of social quantitative analysis. the main idea is to make available for investigators and educators the sound projects, to store the research results costly in both financial and human terms and to preserve them for future use. consultations of authors of the software will be available for the faculty ready to use the programs in educational courses and investigations. socially-relevant projects using the uis russia stuff will be initiated. not only msu faculty is informed and invited but other universities of moscow and regional universities. bilingual search instruments the uis russia has been initially designed as part of the international information structure to serve not only russian researchers but also foreign specialists on russia and general public. to meet the challenge, a special complex of the bilingual searching tools is being developed. the prototype of the bilingual complex is ready for testing and evaluation by a team of russian american specialists. funding to evaluate the bilingual search tools is currently being sought. the nlp technology and developed bilingual complex will produce an annotation in english on each russian document. this accomplishment widens the uis russia audience, helping foreign public to open russia and foreign specialists to investigate russia. work is underway on the global information service (gils)-profile to provide for the uis russia integration into the world information space assisting russian specialists in international cooperation in economic, social, political, human research. russian universities network the uis russia has been designed as a base for interuniversity cooperation in consorted and rational efforts to build a networked collective information infrastructure. the regional universities may actively participate. currently up to 50 local universities are technically and technologically equipped to take part in the cooperation. the open society institute (soros fund)-funded “russian universities’ internet centers” program provided the regional universities with the hardware-software platform compatible with that of the uis russia. the nlp technology developed under the uis russia project may be passed on for free to the regional universities to enable them to develop information systems of their own on local resources, integrated into the uis russia. the local stuff is important to make the social analysis relevant. in this respect the uis russia is close to the american universities’ internet2 initiative directed to rationally build internet-based far-reaching educational and research network. from the very beginning the uis russia project was developed in cooperation with the michigan interuniversity consortium for political and social research and european consortium for social research. specialists of both structures visited russia and have been of help to the russian researchers. program to become self-supportive maintenance of the system on self-supporting basis is a challenge of the project. the experiences of information structures in the usa, europe and other countries’ have been analyzed, and main elements of financial activity of those organizations will be realized institutional membership with annual dues for foreign universities and other organizations. preliminary discussions with american and european colleagues university professors, think tank specialists, government analysts, journalists prove that this way is the most acceptable for them as well. the dues will create a relatively stable and predictable financial base and allow the program to engage in longrange policy to develop the information resource and provide access for free to the high education institutions in russia. iassist quarterly fall 2000 7 the team the uis russia project began 1994. the key specialists have worked together since that time. the team includes 20 specialists from the research computing center, other faculties of moscow state university, academic institutions and other universities of moscow. several specialists are invited for half-time job and consultations. a group of american consultants provide their expertise of the project on the permanent basis. since 1994 the project has been supported by grants from russian fund for basic research, russian humanitarian scientific fund, ministry of science and technologies of rf “informatization of russia” program, macarthur foundation, usa, ford foundation, usa. * tatyana yudina, ph.d., leading researcher of moscow state university research computing center, director of university information system russia project. yudina@mail.cir.ru mailto:yudina@mail.cir.ru the contemporary jewry database: coordination and flexibility in bibliographical registration and retrieval by mira levine , mls ' director ofthe bibliographical center institute ofcontemporary jewry of the hebrew university jerusalem the bibliographical center in contemporary jewry was designed some five years ago to compile a computerized database of bibliographical information relating to various aspects of 20th century world jewry. a digital minicomputer (vax-750) was purchased and the aleph information retrieval program developed at the hebrew university was adapted for this purpose. in the fall of 1985 we began registering indexed and annotated descriptions of books, articles, films, taped interviews and ouier archival material at or available to the institute of contemporary jewry. the records are accessible both for publication as separate catalogues by the cooperating bodies and for on-line global searching via the aleph network of the hebrew university and other israeli academic institutions. to date, over 15,000 bibliographical items have been indexed, abstracted, and registered by seven cooperating but separate computer projects: bibliography in anti-semitism through the ages of the sasson international center for research in anti-semitism steven spielberg jewish film archive holdings and catalogues jewish filmography project listing of films on jewish subjects available throughout israel oral history division of the institute of contemporary jewry holdings and catalogues publications of the institute of contemporary jewry, and of its teachers and researchers bibliography of contemporary jewry 1984-1987, listing relevant books and articles published between 1984 and 1987 arriving at the jewish national and university library. (although discontinued for lack of funds, it will be resumed when funding is available.) studies in contemporary jewry in english and hebrew articles, book reviews and books reviewed in the two annual journals of the institute of contemporary jewry. slated for future inclusion are the library and bibliographical projects of jewish demography, syllabi of courses offered through the years at the institute of contemporary jewry, am^can and the holy land bibliographical and archival project, and other funded projects approved by the institute. once registered, all records are retrievable via a single master index by author, title, subject or words appearing in the title, subject headings of abstract. for example, by entering the search term "intermarriage" one summons a list of not only books and articles in which intermarriage is treated, but films and taped interviews as well. further sophistication in word searching is to be included this summer in the newest aleph version, thus facilitating access even more. having noted briefly the construction and composition of the contemporary jewry database, let us turn now to the various products and activities it generates at the institute of contemporary jewry bibliographical center. printed catalogues and bibliographies: several participating collections retrieve and publish their respective entries as catalogues, bibliographies or filmographies. the bibliogi^hical center staff designs the aleph application, advises, trains and supervises staff involved in all aspects of registration and retrieval right up to the preparation of photo-ready copy for the publisher. garland publishers in new york has undertaken the publication of sections of the database as part of its series garland reference library of social science. to date, one volume has appeared: anti-semitism: an annotated bibliography, volume 7, edited by susan sarah cohen for the vidal sasson international center for the study of anti-semitism, 1987. the second volume of the anti-semitism bibliography and three others are in the final stages of preparation films of the holocaust: an annotated filmography of collections in israel, edited by sheba skirball for the spielberg jewish film archive (in preparation) oral history ofcontemporary jewry: an annotated catalogue, compiled by institute of contemporary jewry. spring 1990 on-line database of course, the primary product of the bibliographical center is an efficient, easily used on-line database of bibliographical information on the various aspects of twentieth century jewry. subjects include the holocaust, zionism and the state of israel, jewish demography and other social science research, anti-semitism, israeldiaspora relations, jewish communities the world over and the arab-israeli conflict to name the most salient thesaurus generation a third product, an on-line publishable thesaurus for contemporary jewry, will be discussed in detail later on. network searching besides catalogue production, database maintenance and thesaurus development, the bibliographical center provides researchers at the institute with access to information relevant to their work via retrieval from our own database, libraries and databases on the network of israeli universities (aleph), and other israeli and foreign databases accessible via our facilities. individual search and subject updating requests are filled as much as time and budget allow. in coming months we hope to be extending these search services to include £k;cess to israeli databases of relevant materials outside the aleph system as well pertinent foreign databases and vendors such as dialog and brs. institute of contemporary jewry researchers are also assisted in computerizing their own research using the aleph system. experience in cooperative computerization one of the most interesting by-products of developing the database has been the experience and knowledge that has accumulated in the process of designing and implementing our cooperative computerization. we are frequently turned lo by libraries, archives and information centers interested in establishing similar or related projects, or in implementing the aleph system for nonlibrary and/or multi-media applications. we enjoy these opportunities to share both our knowledge and data. at the same time we benefit by learning about and often gaining access to related data registered at these institutions. thus, the spirit of cooperation has been extended beyond the parameters of our own institution's projects. aleph adaptation we have found the aleph information retrieval program particularly suitable for our independent/interdependent applications. first, its multi-lingual, multicharacter-type capacity is essential for registration of materials describing world jewry in hebrew and yiddish as well as latin character languages; arabic is also available, though not used by us; cyrillic and far eastern character use is still to be developed. aleph 's structure of several local libraries within a single global library allows, on the one hand, independence in design, cataloging and maintenance, searching and retrieval, and on the other hand, overall maintenance and control if desired. other invaluable aspects of the program for our purposes are its user friendly presentation (as we have over a dozen professionals and countless users at varying levels of computer proficiency), and its flexible thesaurus construction and maintenance capabilities. being part of the overall network of israeli academic libraries and databases is of great advantage, not only by providing access to our database from any of the 25 installations all over the country, but by allowing us to search their collections as well with but a simple 5-character command. finally, the excitement generated by participating in the development of such au courant, constantly growing and improving research project has stimulated creative applications, fruitful cooperation and productive commitment by our bibliographical center and participating project staffs. coordination and flexibility indeed, the execution of such a multi-faceted, multidisciplinary and multi-media database has been a fascinating, albeit challenging, exercise of coordination and flexibility. each stage of designing, implementing, evaluating, improving and expanding the database over the past five years has necessitated a careful balancing of desire for overall standardization while satisfying the requirements peculiar to each individual project cooperation began with sharing information on and the costs of hardware acquisition and maintenance. experimenting with aleph adaptations for bibliographical and archival applications also involved learning from previous and each other's insights and mistakes. common code assignments wherever possible facilitates maintenance as well as the sharing process. meanwhile standardizing divergence provides helpful searching and maintenance cues. for example, all codes peculiar to a given collection are preceded by the same character: p for the film archive, a for anti-semitism, while both ptl and atl are dumped into and accessible via a single title index. training for indexers, abstractors, catalogers, and editing is centralized as much as possible, and retrieval for publication, contacts with both our pubusher and computer facilities of the university are centrally coordinated. this avoids wasteful duplication of efforts while at the same time providing supportive, stimulating couegiality. peculiarities of different types of material, support organizations or lines of authority thus become fhiitful bases for comparison and adaptation. two activities at the bibliographical center illustrate particularly well the balance between inter-dependence and independence of the cooperating members of the contemporary jewry database: one is thesaurus control, the construction and maintenance of our controlled language for indexing; the other is catalogue preparation, the production of photo-ready copy of bibliographies/ catalogues/filmographies for publication. 50 lassist quarterly vol29-3.indd 14 iassist quarterly fall 2005 by karsten boye rasmussen & repke de vries 1 self reflection of virtuality in a professional association: a compact description of mailing list data abstract this article presents an investigative description of the utilization of a mailing list in our own professional membership organization: iassist. as an electronic form of communication, the mailing list has supported and supports the iassist in moving the organization in the direction of a virtual community. the mailing list offers an answer to the jesting question: “is there iassist life between iassist conferences?” this work contributes to methodology by offering a refined typology for the description and analysis of mailing lists in general, as well as a specific subject categorization for the iassist mailing list based on findings in the mailing list communication. the article gives a compact quantitative description of the key figures from the use of the iassist mailing list based upon the typology and categorizations. the analyzed data consist of emails from 32 months before the year 2000 millennium turn. a follow up on the analysis with more present email data is considered. introduction virtual organizations have been identified as real (davidow and malone, 1993) or real organizations sometimes viewed as imagined (hedberg et al., 1997). the concept of the virtual community has now existed for a good 10 years (rheingold, 1993). the virtuality emerges due to intense use of information technology corresponding to organizational arrangements that potentially and practically break the boundaries of time and space. in our time of virtuality, people no longer have to share the same space or be in the same time, as direct electronic communication can span the space, and relayed (asynchronous) electronic communication can span the time. a voluntary association – such as the iassist – is considered as both an organization and a community, and by applying electronic communications, such associations have the potential of growing into a virtual community. the aim: community and virtuality the object for investigation – iassist – is a small voluntary professional organization. the international association for social science information service and technology is an organization of professionals – typically from data archives and libraries – supporting research and education. the concept of “social science” is viewed in its most wide-ranging sense. iassist is a network and shows network externalities. “the more the merrier,” but naturally this is balanced with the group of members actually being a group with common issues. size has to be balanced with the necessary homogeneity within the group. the iassist organization was founded in 1974, and presently has about 300 members. the growth in members is mostly a result of the fact that “international” 25 years ago primarily meant “uswestern-european,” but now more fully encompasses the globe. however, the iassist organization has not had a strong intention of membership growth into areas outside its professional base of data archivists, data librarians, and some social scientists. all of these affiliated with mostly university and/or research institutions. the existing iassist network can be viewed as comprising most of the relevant potential membership, but there is small but constant growth in regarding data materials as a resource available through libraries and archives. a community is a togetherness that shares. historically, the sharing within communities can take different forms, from the communion (with focus on sharing idealistic/religious beliefs) to the commune (that also practices sharing of material effects). all communities share meaning through communication. the knowledge shared in the community can be exemplified or statistically described. in this context the starting point is a look at how the sharing takes place. historically, social groupings have shared the same geographical locality and togetherness in time. information technology breaks that barrier by making some sharing look and feel as reality, but it is “virtual reality.” virtuality is not reality, but it could be a dream or an imaginative creation. it can also be said that it surpasses reality – “it is surreal.” the surreal found in arts and literature surpasses the objective reality as it includes individual subconscious elements to provide a more full understanding of our reality – even an antagonistic image of reality. half a century ago the mass media turned the world into a “global village” (mcluhan, 1964). virtuality contains this implosion in a synchronous form; virtuality provides an even more elaborated ability to react and utilize the new media of communication in obtaining closeness across barriers of iassist quarterly fall 2005 15 time and space selecting media for observation the communication between members of iassist has, from the start of the organization in 1974, taken place through conferences and also a newsletter/periodical. the membership has been among the “first movers” in the use of information technology (it) for communication because their professional work included intensive use of it and new media have been added to the list. from early on iassist has adopted the use of a mailing list and, later on, also a supporting web site. these media are briefly examined here for their capability to support virtuality and thus for founding the basis of an investigation into the virtuality of iassist: conference: the conferences are where iassist members meet face-to-face. conferences are where people are situated in the same time and place and under a common heading and, furthermore, mostly detached from their obligations of everyday tasks. we may ask: “how and where do iassist members exchange views and share knowledge between the yearly conferences?” or this could be formulated: “is there life in iassist between conferences?” newsletter: some life is added to the organization four times a year. since 1977 there has existed a communicative channel for the membership associations in the form of a traditional printed and (surface) mail delivered membership newsletter (the iassist quarterly or iq). the content of the publication is primarily papers from presentations at the iassist conferences and the publication is not directly a medium for deliverance of data for research, although many articles have that subject as their focus. the articles can thus act as endorsements of data and directions for gaining access to data, as well as describe systems for data deliverance. secondly, the iq does not contain much communicative bi-directional interaction. some references between articles in the iq are found, but the periodical contains no actual debate. although the newsletter is now also available in electronic form at the iassist web site (http://www.iassistdata.org/), the newsletter is considered without substantial independent importance for support of the virtual community. the argument is that the communication is one-way and that it started – and still exists – in the low-tech form of printed paper. however, the ease of access facilitated by the availability of the iq on the web and the fact that iq documents the objectives that members of iassist are pursuing in their work life count as assisting factors for iassist being a community. web site: the iassist web site has been conventional by creating access to some formal documents and more permanent announcements, especially the conferences. a bigger move forward was made when the iq newsletter was published on the internet (issues are available from 1993 onward). furthermore, the web site features electronic copies of presentations made at the conferences, as the computer files (powerpoint) are collected and presented in the original conference structure of days and sessions. the authors of this article have carried out some empirical research considering preferences for services at the web site. (this work is intended to be published later.) after the collection of the empirical data utilized in this article and after the survey of preferences mentioned above, the iassist has also recently opened a web log as a publishing and discussion board. this is obviously relevant for the concept of virtual community, but is not considered further in this article that is concentrating on the facilities available at the turn of the century. mailing list: in the early 1980s, communication by email was already an established fact among the iassist membership. because of the international cooperation in the organization and the connection of data archives and libraries to mainframes for universities and research institutions, most members were connected to the forerunner of the internet (arpanet). the distribution of emails among members involved many copies to other members and was then consequently structured by establishing a mailing list for the membership. this mailing list has the central capability of permitting two-way communication. is the mailing list sufficient for constituting the infrastructure of iassist as a virtual community? the phrase “virtual community” was first used by howard rheingold (1993) in a book taking the point of departure in a mailing list (the well, “whole earth ‘lectronik link”), so there is precedent for virtuality obtained through a mailing list. consequently, we regard the iassist membership as a virtual community – in relation to the fact that the non-virtual communication is sparse – and we regard the mailing list as a valid medium for investigation. mailing lists a mailing list is basically a communication duplication facility for email. emails sent to the mailing list are being distributed or sent on to all members (“subscribers”) of the mailing list. hardie and neon (1994) distinguish mailing lists into three types based on the applied filtering of information. the first one is the unmoderated list where everything sent to the list immediately is replicated to the subscribers, with no waiting time, but some messages might be annoying. the second one is the moderated list, where a moderator has to accept the input to the list; this requires work and introduces a varying time buffering of the accepted messages. however, the editing is necessary for the list not to be overwhelmed by unwanted commercials (spam). the third type of mailing list is the 16 iassist quarterly fall 2005 digest list which sometimes resembles a newsletter by having several subjects included and commented on by an editor. there will be only a few emails and they will appear with some regularity (e.g., monthly). the digest list is a one-way distribution list because communication travels only from the editor to the members. the categorization above is based upon whether all subscribers can make postings to the list directly, indirectly, or not at all. this categorization is paralleled by database users having levels and combinations of “read” and “write” permissions. the next aspect to consider is whether subscription to the mailing list is open to everybody or to a defined group of people. the investigated mailing list of iassist is a moderated list, with only the membership as possible subscribers, and a subscriber can always both submit and receive emails from the list. literature mailing lists have been the subject of some earlier investigations. fox and roberts (1999) investigated the use of a mailing list amongst medical doctors in practice (gps). the article demonstrates some anecdotal content analysis of the emails as citations from emails are presented, however the article contains neither statistical analysis nor description. another approach was used in the study by hannah (1999), where questionnaires were emailed to the subscribers of a mailing list to investigate the observed benefits of the mailing list. this was carried out without investigating the activity on the mailing list itself. xu (1998) studied several mailing lists used by system librarians. emails from one of the lists were examined – for a two-month period of time – and questionnaires were sent to several categories of users (or user roles). an equal short period of mailing list traffic was examined by burton (1994). empirical later investigations of several mailing list and their members and in particular their non-participants (“lurkers”) are found in stegbauer & rausch (2002). available data the current investigation of a mailing list is also an endeavor into the investigation and demonstration of the obtainable level of information from the internet without actually asking for individual approval and consent from the subjects being investigated. often the data of the mailing list is placed on the internet and easily and directly available. in this case, clearance to access the iassist archives of the mailing list was given to the researchers by the organization. but many mailing lists are publicly open for retrieval and they often have searching facilities for the identification of threads of interest, and contributions to mailing lists are sometimes stored for easier retrieval and presentation through the use of a web-application. in the use of data in this article, precautions have been taken not to directly reveal the identity of people and their expressed opinions. however, the information for this kind of monitoring of activities on the internet and especially on mailing lists are available. king (1996) discusses the ethical aspects of the availability of communication data and burton (1994) also addresses the ethical aspects by sending out information to the mailing list under investigation. themes of analysis the analysis will present descriptive answers to the following dimensions and questions that became apparent when describing the mailing list from some obvious viewpoints of interest: active-passive: who is sending to the mailing list? the senders are all identifiable as email addresses, thus permitting determination and comparison of the active addresses. secondly, the members of the mailing list are also identified as email addresses (as the addresses being sent to) and these are the members of the iassist. the senders can be investigated and compared to the passive non-senders. (general communication and mailing list as media). nationality: with some accepted uncertainty, email addresses may indicate nationality. the main validity problem is that usa becomes the default used when a nationality is not directly given. but this is considered a minor problem in this context because most members outside the usa can be said to belong to educational institutions that in their email are directly attached to a nation. had it been commercial institutions, this solution might not have had sufficient validity. (internationality of an organization). officer-member: are certain membership groups more active in posting information to the mailing list? a list of persons performing official functions (officers) is available as another mailing list is used for administrative purposes. some differences in communication patterns between officers and regular members are expected. (hierarchy in the organization). modes: wwhat modes of communication take place on the mailing list? some emails stand isolated (“single”) while other emails are connected and can be combined into threads with regard to the same subject and within a defined period of time. within a thread, the single email is classified according to its role in the thread as “initiation” or “reply.” (communication initiator or follower). data and method the underlying unit of analysis is a single email and all emails are stored and retrievable from the list server from december 1991 and onward. a technical shift occurred in may 1997, so to secure the comparison issue the investigation period includes 32 months (from may 1997 until december 1999). because the same person could iassist quarterly fall 2005 17 have several email addresses and a person could have changed his or her email address during the 32-month period, a considerable process of match-merging by userwritten matching software took place for performing a valid aggregation and shift of analysis unit from emails to persons (participants on the iassist mailing list). findings and figures the total material consists of 691 emails that have been sent from 162 persons. the membership on the mailing list consists of 265 persons, i.e., 103 persons did not post email to the list during the investigated period of time. the 691 emails sent from 265 potential posters of mail results in an average of 2.6 emails per person (or approximately one email to the mailing list per year per person). when comparing the figures for within and outside north america, it appears that the ratio for mail per person is 3.1 versus 1.4. the higher figure among members from north america supports the findings of us-dominance in earlier mentioned studies (xu, 1998; burton, 1994). the membership of 265 persons can be subdivided into 31 persons belonging to the group of iassist officials and 234 regular members. seven of the officers have not participated in the mailing list. the ratio of emails from 24 active officers (280 mails) is 11.7, which is significantly higher than the 138 active nonofficers sending 411 mails (ratio 3.0). a regression model shows that the binary office variable is highly significant and explains 20.1 percent of the variation, and that the addition of the geographic variable and the interaction term only accounts for an extra 3.6 percent of the variation. in this investigation, the office variable is consequently considered a better explanatory variable for what otherwise appears as a regional dominance. the figures above account for the fact that 96 non-officer members are inactive in submitting to the list. inequality in electronic communication has in several contexts been the subject for studies (sproull and kiesler, 1991, p. 60), and terms like “quiet observers” (ha, 1997; xu, 1998) or “lurkers” (fox and roberts, 1999; stegbauer & rausch, 2002) have been introduced. however, it is only reliable to conclude that persons responding to the list are reading the emails – or more precisely just that email. however, the rational behaviour of a mailing list member who never reads the mails would be to unsubscribe to the list. inactive subscribers can be considered content and regard the list as providing useful information as in direct performance improvement via computer-mediated communication (rice, 1994). the emails were combined into threads or single emails, where a “single email” has no response from the list. the separation into threads was done by user written software that stripped the title field down to essential information, and when titles matched and emails were close in time they were considered to belong to the same thread. a thread with only one email is a “single.” so if an email is not a “reply” it is a starting email that can be categorized as either a “single” (without any later responding emails and not part of a thread) or an “initiation” (the start of a regular thread). three hundred of the emails were single and had no follow up; the remaining 391 emails were combined into 125 threads (as shown in table 1 below). the content of the emails are not analyzed in this context. however, some single emails never expected any response as they are often email announcements (e.g., emails announcing the availability of a new dataset). the officers were initiators of 213 (149+64) occurrences while the regular membership started 212 (151+61). the emails in table 1 are produced by 24 officers and 138 regular members. the officers are thus characterized as more frequent starters of emails (an average of 8.9 emails per officer) compared to the regular membership (with an average of 1.5 emails). with respect to posting replies to the list, the difference between the two groups is not that extreme (averages 2.8 and 1.4). furthermore, there is a clear relationship between being an active initiator and being an active replier. the number of replies correlated to the number of starts from the same person was much higher among the officers than among the regular membership (pearson 0.66 versus 0.20). this means that a small group of the officers are very active in their use of the mailing list. conclusion the article has demonstrated a utilization of data being available on the internet – the data driven approach – and we have not created or collected other data for this starting following all single thread initiation reply officer 149 64 67 280 regular member 151 61 199 411 all 300 125 266 691 table 1. distribution of emails to mode and membership type 18 iassist quarterly fall 2005 particular research task. we regard this as defendable ethical research as individuals are not being exposed, but we welcome a debate on surveillance of individuals without their given consent through the materials available or traces left on the internet. the findings of the descriptive analysis include explanation of email participation, and it was found that the crucial variable for high activity was whether a person had official duties within the organization. the mailing list forms a virtual community – and like for a virtual organization – a virtual community is also characterized by blurred boundaries. in further investigations we have used (but not yet published) questionnaire data to look for evidence of further virtuality in terms of boundary crossing where non-formal (nonpaying) members enjoy close to the same benefits as the regular membership. furthermore, a follow-up study on the mailing list (5 years after) is being considered * this article appeared in another form in the proceedings from the 11th international conference on humancomputer interaction (2005). some of the figures were earlier presented at the iassist 2001 conference in amsterdam as “professional associations in transition to virtual communities for collaboration: the case of iassist.” repke de vries, department of public services at the royal library of the netherlands. karsten boye rasmussen, department of marketing and management at university of southern denmark. contact e-mail: kbr@sam.sdu.dk. references batinic, b., reips, u.-d. & bosnjak, m. (2002). online social sciences. hogrefe & huber, göttingen. burton, p.f. (1994). electronic mail as an academic discussion forum; 1994; journal of documentation, 50(2), 99-110. davidow, w. h. & malone, m.s. (1993). the virtual corporation; harper business; new york, n.y. donath, j. s. (1996). identity and deception in the virtual community. in smith & kollock (ed.), communities in cyberspace; routledge fox, n.; and roberts, c. (1999). gps in cyberspace, the sociology of a “virtual community”; the sociological review (nov. 1999), 47 (4), 643-671. ha, l. (1997). active participation and quiet observation of adforum subscribers; journal of advertising education 1997, 2(1), 4-19. hannah, r. l. (1999). a case study of benefits-l subscribers; benefits quarterly; third quarter 1999, 24-29. hardie, e.t. & neon, v. (1994). internet: mailing lists; ptr prentice-hall; englewood cliffs, nj. hedberg, b.; dahlgren, g.; hansson, j. & olve, n.g. (1997). virtual organizations and beyond: discover imaginary systems; wiley; chichester. king, s.a. (1996). researching internet communities: proposed ethical guidelines for the reporting of results; 1996 information society, 12:119-27. mcluhan, m.. (1964). understanding media: the extensions of man, routledge classics. rasmussen, k.b.; and de vries, r. (2005). growing virtuality in a professional association – a data driven approach. 11th international conference on humancomputer interaction (pdf proceedings cd). rheingold, h. (1993). the virtual community; addisonwesley; new york. rice, r. (1994). relating electronic mail use and network structure to r&d work networks and performance; 1994; journal of management information systems 11 (1) 9-29. sproull, l.; and kiesler, s. (1991). connections: new ways of working in the networked organization; mit press. stegbauer,c. & rausch,a. (2002). lurkers in mailing lists. in: online social sciences, batinic,b. (ed.), 263-274. xu, h. (1998). global access and its implications: the use of mailing lists by system librarians; volume 35 1998 proceedings of the asis conference “information access in the global information economy”; 427-444 iassist special issue 6 iassist quarterly 2010 / 2011 iassist quarterlyiassist quarterly introduction in april 2009 the uk timescapes initiative, in collaboration with the university of bremen, organised a residential workshop to explore the nature of qualitative (q) and qualitative longitudinal (ql) research and resources across europe. the workshop was hosted by the archive for life course research (archiv für lebenslaufforschung, allf) at bremen and funded by timescapes with support from cessda (the council of european social science data archives, preparatory phase project). it was attended by archivists and researchers from 14 countries, including ‘transitional’ states such as belarus and lithuania. the broad aim of the workshop was to map existing infrastructures for qualitative and ql data archiving among the participating countries, including the extent of archiving and the ethos of data sharing and re-use in different national contexts. the group also explored strategies to develop infrastructure and to support qualitative and ql research and resources, including collaborative research across europe and beyond. background and context the bremen workshop can be seen as part of a much broader effort to co-ordinate research resources across europe. the impetus for the workshop was provided through cessda, a distributed research infrastructure that provides access to european research data and supports their use. cessda is currently a federation of national data dissemination and support organisations spread across europe, with a small, voluntary elected distributed executive. collectively they serve over 30,000 researchers, provide access to more than 50,000 data collections per year, and facilitate the exchange of data and technologies among data organisations through common authentication and access, cross-european resource discovery, secure data facilities, and the adoption of inter-operable metadata standards. a major upgrade is necessary, however, in order to strengthen and widen the existing research infrastructure and make it more comprehensive, efficient, effective and integrated. this was the key argument for placing cessda on the european strategy forum for research infrastructures’ (esfri) roadmap in 2006. work is now underway to establish and expand an upgraded cessda as a legal entity under the european council regulation 723/2009 as a european research infrastructure consortium (eric) (cessda 2011). to date, however, the data available through the cessda portal are predominantly quantitative (qn), including official government census data, social surveys, and quantitative longitudinal and cohort studies. while all the current infrastructure initiatives are vital, regardless of data format, there has been little development in building and harmonising infrastructures specifically for qualitative or ql data, and little account taken of the distinctive requirements for archiving and re-using these data. human data qualitative and qualitative longitudinal resources in europe: mapping the field and exploring strategies for development by bren neale and libby bishop1 2 editorial introduction iq iassist quarterly special issue 21 july 2011 iassist quarterly 2010 / 2011 7 iassist quarterly of the sort embodied in qualitative and ql research are challenging simply because they are endlessly varied, fragmented, complex, dynamic, multilingual, and historically, politically and geographically situated. preserving and disseminating the products of human culture and society is difficult and expensive, particularly for qualitative data. even so, new digital resources, including software and e-networks, are influencing the production of human records and how these are understood and communicated. it was in the context of this shifting european picture that the idea for the bremen workshop was first conceived. the workshop was framed in terms of identifying existing qualitative and ql resources and exploring ways of building a european network of qualitative and ql researchers and archivists committed to preserving and organising qualitative data resources for sharing and re-use. the endeavour was seen as complementary to the work being undertaken under the first phase of cessda. the esfri roadmap (2011) indicates the enormous potential of data—of all kinds—for understanding the profound social, cultural, political and economic life of europe, including social continuity and change. the roadmap also reminds us that the first step in developing such infrastructure is networking and co-operation, and it was in this spirit that the bremen workshop took shape. the bremen workshop the workshop participants were asked to produce a country report that would set out the nature of existing infrastructure for qualitative and ql archiving, policies and ethos for data sharing, an overview of key resources and collections of qualitative and ql datasets, and priorities for and barriers to future development. the reports were tabled at the workshop and, for the purposes of presentation, were grouped into three broad categories (from most to least developed in terms of infrastructure). one representative from each of the three groups presented a brief overview of developments within the group, pointing out areas of commonality across the countries, and important circumstances and features that distinguished them. the afternoon breakout sessions mixed members from all three groups. they were tightly focused on development planning and structured around these questions: • what enables and constrains data sharing? • how effective are existing models for sharing or archiving data? • what are the pros and cons of having a mixed infrastructure of data archives and collections, centralised and distributed, generic and specialised? • is there a case for developing separate infrastructure for qualitative and ql data resources or for merging these resources with existing quantitative and longitudinal resources? • what are the best ways of getting an archive started and what issues arise in developing and sustaining the resource? • would a european wide network for qualitative data archiving be beneficial and if so, how would archivists and researchers prefer to participate? we present here an overview of developments across the three groups of countries, the insights emerging from our workshop sessions, and some pointers for future developments group one – finland3, ireland and the uk this group has established national archives for social science data that include qualitative collections (esds qualidata in the uk funded from 1994, the irish qualitative data archive (iqda) from 2008, and the finnish social science data archive from 2003). in each case the archives include primarily interview data (with focus groups and other textual sources) and documentation. all three also have, or are planning to add multimedia formats (e.g., sound, images, and moving images) and analytical files. although funded as national resources, the three countries are characterised by patterns of decentralisation; in the uk, esds qualidata, for example, is a specialist service of the economic and social data service, led by the uk data archive, and qualitative data is fully integrated into its holdings. the qualitative collection is the most important but not the only hub in a vast network of independent and proliferating collections held by a wide range of organisations that are rarely co-ordinated. this ‘mixed’ infrastructure with specialist and generic resources existing alongside each other was seen as inevitable; though it may pose co-ordination challenges, there is also potential for innovative collaborations. ql research and resources are well represented across these three countries. in the uk a specialist timescapes archive for ql data, funded from 2007 by the economic and social research council, and developed at the university of leeds, has been established. it is based on a close integration of ql research, archiving and re-use and is useful as a platform for training in the secondary use of ql data. at the national level, the three countries in this group have policies promoting data sharing. there was growing awareness of qualitative datasets as important research outputs in their own right, and a growing appreciation, therefore, of the need to produce high quality data outputs for sharing and re-use. key national funding bodies in these countries all require data management planning and recommend archiving or data sharing as a condition of funding. despite these developments, however, support for data sharing in these countries remains uneven; complex issues surrounding data sharing have emerged that need to be taken into account. for example, in finland there is no established culture of promoting qualitative data re-use and an assumption remains that primary researchers are the only ones to understand and use the data correctly. in the uk such views are much less prevalent and researchers are beginning to explore the potential for combining primary and secondary data analysis in their work and, thereby, increasing the robustness of their evidence base. however, there are ongoing issues around balancing secondary access to data with the need to protect confidentiality and also to allow sufficient time for primary analysis to take place. in contrast to large scale survey and cohort data, qualitative and ql data are not generated solely for secondary use; they are generated, at the outset, by and for primary analysts to address particular research questions. the originating team therefore faces the challenge of balancing the potentially competing tasks of data gathering and analysis with that of preparing data for archiving. for ql research, where projects may run for many years with ongoing waves of data gathering and complex temporal analysis by the originating teams, this may prove a challenging task. the drive to archive in this context may be diminished unless sufficient incentives are provided by funders. whatever the ethos surrounding qualitative data re-use, these issues have important implications for the timing of archiving and the resources needed by originating teams for data preparation tasks. in the context of qualitative and ql data, then, it is clear that both primary and secondary use need to be accommodated and balanced in the strategic development of research practices and the provision of data infrastructures. priorities for development identified within this group of countries included technical development of the archives to include multimedia data, and the development of the specialist curation, data discovery and preservation procedures needed for ql data. despite the advances in these countries a need was identified in each case to build the culture of data sharing and re-use, and to strengthen policies and develop initiatives to support this aim, for example, through funding for secondary analysis of qualitative and ql data. in the uk, one encouraging move has been the economic and social research 8 iassist quarterly 2010 / 2011 iassist quarterly council’s announcement of a major strand of funding to support secondary analysis (2011). a need was identified for greater co-ordination of data resources across the mixed infrastructure, so that specialist and distributed collections could more easily be identified, searched and accessed. finally, funding was relatively fragile and there was a need to secure longer term funding to facilitate this work and make its outputs sustainable. group two – austria, czech republic, denmark, germany, norway and slovenia not surprisingly, this group was highly diverse with some members resembling group one in many dimensions, but others being more like group three. generally speaking, there is infrastructure in place for quantitative data archiving; all but the czech republic have existing national archives. in most cases, some fledgling effort is underway for these predominantly quantitative-orientated institutions to begin handling ql data. austria, for example, began archiving qualitative data at wisdom (wiener institut für sozialwissenschaftliche dokumentation und methodik) in 2007 and the danish data archives began handling qualitative data in 2009. the norwegian social science data services in bergen, norway is planning to incorporate qualitative data and the social science data archive in slovenia is in a similar situation. but these national infrastructures capture only a small amount of activity, as there are numerous qualitative and ql resources widely distributed in smaller institutions, departments, and held by individual projects. many of these are attempting to archive qualitative and ql collections, and some are seeking to form alliances or collaborations with quantitative institutions, where they exist. as with infrastructure, the situation regarding data sharing is also ambivalent. in terms of actual archive-mediated data sharing, levels of activity are rather low. but there is growing visibility of the issue and other indications of changing attitudes. formal feasibility studies (for archiving qualitative data) were done in austria, denmark and germany, revealing surprisingly positive attitudes toward both sharing data and using data collected by others. however, hurdles exist in translating these attitudes into more positive actions. where archives do exist – in denmark and austria for example – few datasets have been deposited and the rate of new deposits is low. major challenges remain in numerous areas: concerns about ethics and confidentiality; researchers’ continuing belief in exclusive ownership of data; technological and financial resources constraints; and complex infrastructure models. development priorities reflected the national situations, but all pointed to the need for networking with other institutions and countries. locating stable funding sources was also a high priority, as was engaging in activities to bring about cultural acceptance of data sharing—finding exemplar cases and teaching methods for re-using data, especially to post-graduate students. there are, perhaps, at least some reasons to be optimistic – in germany, the feasibility study, as well as publications and an annual workshop on secondary analysis, has encouraged more active debate about data archiving and sharing. and the commitment to developing appropriate infrastructure for qualitative and ql data and finding ways to harmonise datasets to facilitate wider re-use was evident across all the workshop participants in this group group three – belarus, hungary, lithuania, poland and switzerland members of group three reported only minimal infrastructure for curating qualitative or ql data, though there was obvious enthusiasm for developing such infrastructure among a subset of the academic community. of the five countries in this group, there are only two with national institutions for data archiving, the lithuanian humanities and social science data archive (lida) and the swiss foundation for research in social sciences (fors). where laws exist (e.g., in switzerland), these are general ones on archiving and data protection, with no specific provisions for qualitative or ql data. the culture of sharing is weak to non-existent, at least for qualitative and ql data. in poland, there is ‘no academic tradition’ of sharing qualitative data, perhaps partly because of a very strong prevailing positivist tradition in social research, although encouraging new initiatives began in december 2010. in hungary, there are some existing archives for particular surveys, but data sharing is not common, and the culture of re-using data is not widespread. in the case of belarus, there is no national infrastructure for archiving. data that are retained are held by individual organisations. secondary analysis is rare and occurs only after personal negotiations among primary and secondary researchers. in many cases, research data are not retained at all, even by primary researchers. the recent political climate has, in part, contributed to this situation. in contrast, lithuania does have some national policies promoting sharing, and in addition to lida, there is now access to online research data via electronic information for libraries (eifl.net), but this focuses more on research outputs and not raw data. as might be expected within this group, the list of development priorities is long and wide-ranging. basic work in establishing infrastructures is needed, with the concomitant requirements of appropriate technologies and financial resources. practical examples of archiving policies and procedures would be highly beneficial, and even with the adaptations required for specific national conditions, could avoid a great deal of work being reinvented. administrative advice is also needed, for example on the staffing of archives and what levels and specific skills of staff are needed. specifics include collections strategies (deciding what to archive), and rights management (consent, anonymisation, access controls, ipr, etc.). in one area, however, there was strong unanimity in group three, and across all the groups for that matter: the desire and need for stronger international knowledge exchange, joint projects, and resource sharing. workshop outcomes the bremen workshop produced an impressive collection of outcomes in three areas: short-term activities, agreed goals and objectives, and a strategic plan for future action. some aspects of the strategy outlined below have emerged in subsequent communications among the workshop participants. short-term activities the top priority arising from the workshop was to produce this publication, based on revised versions of all the country reports. additionally we have: • set up a network for qualitative and ql archivists across europe, known as equalan (european qualitative archiving network). • created forums for digital communication, including the bremen workshop webpage http://www.timescapes.leeds.ac.uk/ events-dissemination/past-events-presentations/bremen-workshop/ • produced a distribution list for the members of the network in methodspace, http://www.methodspace.com/group/timescapesqu alitativelongitudinalresearch?xg_source=activity. • revised a list of international data providers on the esds website—this is in progress here: http://www.esds.ac.uk/qualidata/ access/internationaldata.asp. • published a list of all ql collections and resources provided in the country reports—this has been developed through the timescapes website (www.timescapes.leeds.ac.uk) and will become available in the resources section of the site in the first half of 2012. iassist quarterly 2010 / 2011 9 iassist quarterly • investigated specific funding sources, including developing a proposal for infrastructure funding through eu framework programme 7, and a proposal for a panel at an international conference in 2012. • organised further meetings, including sub group meetings at iassist in june 2010 and a further workshop in brussels (october 2010), with funding from iqda in maynooth, ireland, and timescapes. • agreed to produce case studies from the most developed archives (iqda, finland, uk, and germany)—we are currently seeking funding to publish these reports. agreed goals and objectives there was broad agreement on the overarching goals and objectives of the network, as set out below. clearly, action in many of these areas is not specific to this network, and it was further recognised that many of these objectives need national or international co-ordinated action. nonetheless, the group felt it important to articulate explicitly how qualitative and ql archiving should become an integral part of these wider developments. strategies for pursuing this include the following: • active networking, in some cases with better-resourced quantitative partners and institutions. • promotion of metadata standards, including specific standards for qualitative and ql data, and encompassing new multi-media formats that characterise these data. • development of metrics for re-use and the technological systems to collect data for re-use. • lobbying funders for specific policy changes, including mandatory data deposit, funding for preparing datasets for archiving, and according equal merit to secondary analysis projects in funding decisions. • changing research output and reward systems to incorporate the production of qualitative and ql datasets. this requires reference and citation credits when using archived data; acknowledgements for data creators as joint authors; assigning digital object identifier (doi) numbers to archived datasets; and the inclusion of datasets as outputs within formal research review procedures (the research excellence framework in the uk and european equivalents). • promoting activities to accelerate a cultural shift toward data sharing. this may be achieved through work with professional associations; training and capacity building with postgraduates and early career researchers; and direct engagement with ethical debates over the re-use of data and the balancing of primary and secondary research. strategies for future development while all the above goals are vitally important, it was recognised that in most instances, these goals are not specific to qualitative or ql data. as noted above, cessda (both in the preparatory phase and in eric) is addressing areas of harmonised legal environments, a multiple language thesaurus, secure access to ethically sensitive microdata, and much more. what this makes clear is that equalan is well positioned to define and address issues that are particular to qualitative and ql data. when devising a strategic plan for archiving qualitative and ql data in europe, the central question is: in what ways are qualitative and ql data the same, or broadly similar, to quantitative data, and therefore able to be harmonised with existing data infrastructures to enhance comparability and enable different kinds of data to ‘speak’ to each other? conversely, in what ways are they distinctive, and thus potentially in need of customised treatment? answers are emerging from several directions. the timescapes initiative has built a specialist ql archive, and in doing so, is uncovering the special needs of ql data. in this instance, ql data archiving is being integrated within ql research practice and methodological developments through a stakeholder model of researcher and archivist collaboration. this is not simply a matter of bringing researchers to the archive but taking the archive into the world of research in a way that has had a significant impact on the impetus to archive and to cultures of data sharing and re-use (neale and bishop 2011 forthcoming). the experience of the uk data archive is also relevant because it was a well-established archive for quantitative data and incorporated esds qualidata into its existing infrastructure, proving that qualitative data can be processed in standardised ways. these experiences, along with related experiences in ireland, finland and germany, point to similar lessons learned. broadly speaking, qualitative and ql data are distinctive from quantitative data in three areas: metadata requirements, ethical considerations, and cultures of generation and re-use. in terms of the open archival information system (oais) model, the intermediate processes of data management, archival storage, preservation planning and administration are broadly similar regardless of data format. of course, provision needs to be made for different formats, large video files being one challenge. however, the processes for handling all data are broadly similar. it is in the early and later phases of the data life cycle where qualitative and ql differences matter most. two of these, metadata and ethics, lie in the ingest (or pre-ingest) phase while the culture of re-use falls within the access phase. by no means are these the only topics that could be chosen, and future strategic planning sessions may lead to a refinement in this list. however, the idea of defining distinctive aspects of ql and qualitative data, and the implications that follow for developing archiving infrastructures that support such data but also allow for harmonising with existing initiatives, seems like a sensible way forward. the first challenges posed by qualitative and ql data are for adequate metadata collection, in part because of the complex file formats involved. data need more extensive metadata and contextual material to render them “independently understandable” (a requirement of the oais standard) for those re-using the data. unlike much structured quantitative data with relatively standardised formats, qualitative research data and documentation are highly diverse. it is also generally accepted that qualitative data need extensive contextual information to enable effective resource discovery and re-use. much of this may fall into familiar metadata categories, but ideally context should also include information about the project background and the social and institutional conditions in the wider environment that might have shaped project design (bishop 2006; irwin and winterton 2011). ethics is the second area that distinguishes qualitative and ql data from quantitative data. on the one hand, ethical standards for the curation of much qualitative data appear relatively straightforward. consent for sharing is usually readily obtained and data can be protected through varied forms of anonymisation and controlled access. however, ethical concerns remain a major factor in debates among researchers about the re-use of qualitative data and every participant at bremen raised some topic related to ethical use of data. typical issues include: can consent be said to be informed when the topics of research for re-use cannot be known in advance? are there risks to participants if re-used data may be exploited or participants’ views misrepresented? are researchers exposed to unfair criticism when their work is made visible by archiving or where secondary interpretations contradict or challenge primary interpretations? these factors have the potential to limit the availability of data for archiving in the short term, even where consent has been obtained from research participants. in a ql context, this has implications for the way archivists work with researchers and suggests the need for involvement in the development of a research project from the outset to facilitate ethical archiving (bishop 2009) and the development of mechanisms to enable researchers to remain engaged in the re-use of data that they 10 iassist quarterly 2010 / 2011 iassist quarterly have generated (for a comprehensive review of debates on secondary analysis, see irwin and winterton (2011). despite rapid change in recent years, it is still the case that the culture of data re-use is weaker and less widely accepted for qualitative and ql data than it is for quantitative data. this is decidedly the case in the group two and three countries, as the country reports reveal. it also continues to be the case for finland, ireland and the uk, although as noted above, the focus of the debate in the uk seems to have shifted recently to the more practical issue of how best to balance the needs of primary and secondary research, particularly in the context of ql data. for data archives, the resource implications are that more effort and resources are needed to promote the re-use of qualitative and ql data. these range from preparation of focused outreach materials to the need for training and support that is customised to distinct audiences. nevertheless, successful qualitative and ql archiving is most important in this respect, because it plays a decisive ‘demonstrator’ role in alleviating researchers’ concerns and normalising the culture of archiving and re-use. future initiatives the bremen participants have stayed in regular communication since the workshop, primarily focused on revising articles for this special issue of iassist quarterly. informal meetings, usually conferences where a sub-group was attending, have taken place to exchange knowledge and explore future funding options. one such meeting was held at iassist in june 2010 at cornell university, where we mapped a strategy for a more formal meeting in brussels in october. the latter event was co-ordinated by the irish qualitative data archive and co-funded by the national institute for regional and spatial analysis (nirsa) at nui maynooth, ireland, and by timescapes. participants from nine countries were in attendance and efforts focused on developing an application for funding. the brussels meeting, and its aftermath, have provided significant progress toward our goals. at this meeting, we formally constituted equalan, our european qualitative archiving network. the remit of the network is to facilitate international data sharing and re-use by developing and implementing strategies for preserving, organizing and harmonizing qualitative and qualitative longitudinal data resources across europe (equalan 2011). more importantly, the formal constitution of equalan has given visibility to the network with the potential to bid for funding. to date the group has devised work packages for two fp7 funding initiatives for research infrastructures, working with dasish (data service infrastructure for the social sciences and humanities) and building on collaborations between archivists and social science researchers4. the work packages which cover areas such as metadata, ethics, and promoting a culture of re-use, can be tailored to specific funding calls. these are significant developments in a field where qualitative archiving has hitherto commanded little presence. equalan will use its considerable expertise to work across a range of local initiatives such as those below: • to develop standards for qualitative and ql metadata: several bremen participants are members of the ddi qualitative data working group that is developing a ddi compliant schema for qualitative data. • the timescapes initiative has produced a guide to the ethics of ql data archiving and re-use (bishop and neale, timescapes methods guides series www. timescapes.leeds.ac.uk). this needs further input from international sources, and the addition of international case study examples. additionally, the stakeholder model of archiving ql data, designed to build collaboration between researchers and archivists and encourage deposit of longitudinal data during the lifetime of a project, could be piloted and evaluated in a broader european context. • much technological development is still needed to create the complex access controls required for highly sensitive and confidential data. fedora software is under development and promises a more robust access system. there is a need to assess existing projects and work out strategies for further development. such work on access controls needs to remain aligned with ongoing work on similar services (such as the secure data service at the uk data archive) that are intended to enable sharing of potentially revealing microdata. • capacity building is needed for teaching the next generation of scholars about the benefits of data archiving and substantively grounded methodologies for conducting secondary analyses using qualitative and ql data. conclusion the development of a european wide network of qualitative and ql archives and resources that could fall under the cessda umbrella would be a step forward, with shared good practice for practical and technical development of resources (e.g., standards such as oais), common protocols for data sharing and kite-marking data, and portals that link qualitative and ql datasets internationally. it would be beneficial to investigate the large range of activities that are already underway in europe regarding digital repository infrastructure (driver 2010). strategies for advancing such a network could be developed, again with the support of organisations such as cessda and iassist. this could involve eu funding for shared activities or low cost alternatives such as web based networking through blogs or discussion lists. qualitative data is abundant across this mixed infrastructure, and has obvious value and potential as a knowledge base for addressing a range of social questions. realising this potential will depend on finding the means to more effectively manage and co-ordinate these rich resources of data. the bremen and brussels workshops have been highly fruitful, opening up a new and vital area for research archiving that is currently underdeveloped for the social sciences in europe. these efforts have highlighted the need to both recognise the unique situation of every archive, and also much shared intent over preservation, data management, and dissemination standards and practices. extending this to encompass the full range of data across the spectrum of the social sciences, with initiatives to create connections across diverse datasets, would be a significant step forward. the creation of the fledgling equalan, with a broad remit to put qualitative and ql archiving firmly on the map, is the first step towards this long-term goal. fp7 or european science foundation funding is a critical next step in securing resources and recognition for qualitative and ql data archiving. given the complexity and diversity of qualitative and ql data, the mixed and highly distributed infrastructure currently in existence, and the varied cultures of data sharing and re-use operating across the countries of europe, different models for the growth of qualitative and ql archiving and data sharing are undoubtedly needed. but notwithstanding these challenges, making such data ‘count’ in the spheres of archiving and secondary analysis will do much to enrich understandings of the social world references bishop, l. (2009). ‘ethical sharing and reuse of qualitative data’. australian journal of social issues. 44(3) spring. bishop, l. (2006). ‘a proposal for archiving context for secondary analysis’. methodological innovations online. 1(2). [online]. available at: http://erdt.plymouth.ac.uk/mionline/public_html/viewarticle. php?id=26&layout=html. iassist quarterly 2010 / 2011 11 iassist quarterly cessda. (2011). council of european social science data archives. [online]. available at: http://www.cessda.org / [accessed] 12th january 2011] economic and social research council. (2011). available at: http:// www.esrc.ac.uk/news-and-events/news/delivering-priorities-funding/secondary-data-analysis.aspx [accessed 20th july 2011]. esfri. (2011). esfri roadmap. european commission research and innovation – infrastructures. [online]. available at: http://ec.europa. eu/research/infrastructures/index_en.cfm?pg=esfri-roadmap. [accessed 12th january 2011] equalan. (2011). the european qualitative archiving network (equalan): an introduction. [online]. available at: http://www.timescapes.leeds.ac.uk/events-dissemination/past-events-presentations/ bremen-workshop/ [accessed 3 june 2011] irwin, s and winterton, m. (2011). debates in qualitative secondary analysis: critical reflections [online]. available at: http://www.timescapes.leeds.ac.uk/events-dissemination/publications.php [accessed 20th july 2011] neale, b. and bishop, l. (2011). the timescapes archive: a stakeholder approach to archiving qualitative longitudinal data. timescapes. (2011). timescapes: an esrc qualitative longitudinal study. [online]. available at: www.timescapes.leeds.ac.uk. [accessed 12th january 2011]. notes 1.bren neale: the timescapes initiative and archive, b.neale@leeds. ac.uk libby bishop: the timescapes initiative and archive, and uk data archive, ebishop@essex.ac.uk 2. acknowledgements: we would like to formally acknowledge all the contributors to this volume for their efforts. collaboration is always challenging, and in this case, it was made more so by the large number of participants, diversity of languages, and the wide disparity of resources available; limited indeed for some members of our group. their perseverance and patience has been deeply appreciated. several contributors also read each other’s papers—that editorial assistance was invaluable. some of our colleagues at the uk data archive read and commented on the introduction and uk report—any remaining errors are ours. jane gray deserves special mention for her continuing work to establish equalan and to obtain recognition and funding for our network. finally, we owe a great debt to esmee hanna who—in the final stretch—did the essential work of final reading and editing, even while preparing for her viva at the university of leeds. her professional efficiency deserves special recognition. 3. because the finnish social science data archive website provides much of the information addressed in the country reports, the article from finland takes a different format. it focuses on challenges to archiving and contributes significant new evidence in the form of research participants’ positive views of archiving data. 4. dasish aims to support social science and humanities infrastructures by providing solutions to common challenges relating to data quality, data archiving, data access and legal and ethical issues. dasish is funded under fp7 and brings together all five ssh research infrastructure initiatives on esfri’s roadmap (cessda, ess, share, clarin and dariah). 44 iassist quarterly the sampling bias in random digit dialing by a. dianne schmidley' bell atlantic corporation all research is based on information from data collection efforts, whether the data result from the use of qualitative approaches, such as, content analysis, focus group sessions, or in-depth interviewing by psychologists, social workers or ethnographers, or from methods which produce data more amenable to quantitative analyses, such as, sample surveys and censuses (technically, a census is a 100% sample stirvey). a study-, conducted by derek phillips in the early 1970s, concerning the preference of sociologists for either qtiantitative or qualitative data collection approaches, revealed that more than 90 percent of the research conducted by those social scientists resulted from the analysis of data collected through the administration of interview schedules and/or questionnaires developed for use in a survey setting. while sociologists are generally viewed as major collectors and users of the data resulting from surveys, they are not the only individuals interested in survey data. pollsters, advertisers, program administrators and others have important uses for survey data. in fact, there are several broad categories of surveys including: 1. attitude and public opinion polls, 2. marketing research, including advertising and public relations surveys, 3. government surveys, especially those conducted for the purposes of developing legislation and administering and evaluating the effectiveness of programs, 'presented at the international association for social science information service and technology (iassist) conference held in manna del rev. california on mav 21-24 1986 4. special surveys conducted by researchers in university settings, usually the basis of primary research, and often government funded, 5. other surveys, including internal ' knowledge from what? fall/winter 1987 tassist quarterly 45 organizational surveys conducted to monitor attitudes and opinions. individuals with training in survey research methods direct many of these collection activities; however, many surveys, particularly those in the areas of marketing, public relations and opinion research, are conducted by individuals who have no formal training and are not aware of the many pitfalls of survey data. once the ill-gotten information is translated into copy, it acquires a life of its own and the 'facts' become almost impossible to erase. the use of data from a survey, for some purpose other than that for which it was collected, is common. for example, data from govemmeni surveys such as the cunent population survey (cps) are often used as the basis of marketing decisions, while data from a public opinion survey may influence the decisions of government policy makers. the data from all five categories of surveys are utilized to support policy decisions affecting the expenditure of millions of dollars. given the penchant of policy makers for basing decisions on data collected through the use of surveys, the need for a critical appraisal of each of the various stages of the survey process has evolved. questionnaire/interviewer schedule design, sample selection, the adminisuation of the survey process, the collection, tabulation and interpretation of the data, and testing of the reliabilityand validityof the information collected from the survey, compared to some benchmark, have all become subareas of the survey research process and are carefully scrutinized by trained social scientists. the purpose of this paper is to shed additional light on the problem of selecting a representative sample of a population to be surveyed, using the procedure of simple random sampling. when a sample is selected for a survey, attention must be given to the parameters of the universe the researcher hopes to study, and the method of eliciting information from that universe. in the case where the universe to be studied is a human population, there are a limited number of ways of operationally defining the universe and collecting information from the individuals comprising that universe. each type of definition will affect the method of sample selection chosen. if the universe is operationally defined by street address records, the researcher can mail a questionnaire to each individual respondent if telephone numbers constitute the universe, the researcher can telephone respondents and administer a questionnaire/interview schedule. with street address records, the researcher can meet with each respondent and conduct interviews face to face. generally, evaluations of each of these three approaches suggest that the higher the quality of the data collected (given proper development of the survey instrument), the greater is the expense of the survey. mail-back questionnaires are the least expensive way in which to collect information and result in the poorest qualitydata with the lowest response rates, while in-person interviews are the most expensive way in which to collect the information, but produce better respondent cooperation. given this set of circumstances most researchers employ the telephone as a means of collecting survey information, since it produces medium quality data at medium cost i mentioned earlier that when drawing a sample from a universe, the researcher operationalizes the sampling process by giving the members of the populaton s/he hopes to study concreteness in the form of a street address or a telephone number. client/customer specific lists, collected by some agencies and firms, can be used when the researcher is interested in some subset of the population, and thereby has access to the specific names, telephone numbers and/or fall/winter 1987 46 lassist quarterly address records of that group. in most instances, however, the universe to be studied is only vaguely known. it may be comprised of all the individuals in some specific geographic location such as a trading area, county or school zone, or it may be a subset of the population, such as households with a certain income, or people of certain ages or ethnicity. at any rate, owing to a lack of specific information which could be used to contact and interview the individuals of interest in the population, the researcher more often uses a method of simple random sampling (srs) of available street records or telephone numbers in order to delineate those units which will be sampled for the purpose of collecting information which can then be ascribed to the larger universe from which the sample was drawn. the preference of survey researchers for telephone interviewing as a means of collecting information from respondents, coupled with the need to employ srs in order to to identify the individuals to be surveyed within the universe in which they are located, has led recently to the development of a technique called 'random digit dialing' (rdd). rdd is believed by the naive to be a cure for drawing biased or non-representative samples from a universe. misinformed advocates of rdd persist in believing that universal telephone service means that everyone has a telephone in their home. secure in this belief, rdd practicioners promote the idea that by simply dialing the telephone in a random manner one can draw a random sample of the population of any given geographic area. census data, which are used by experienced survey researchers, to design sampling frames and to calibrate survey results, indicate that universal telephone service does not mean that a telephone is found in every housing unit 'universal telephone service' is akin to the notion of 'full employment'. full employment is defined by many as the situation in which about 5% of the population in the labor force, ages 16 through 64, and desiring employment, is unemployed. there aie many persons between the ages of 16 and 64 who are not employed, and are not seeking employment; so full employment does not mean that 5% of this entire age group is unemployed. similarly, imiversal telephone service does not mean that every housing unit contains telephone network access. there are a variety of reasons why a imit may be without access to the telephone network. attachment a to this paper gives the reader the sources utilized to determine something called telephone penetration. as can be seen in this attachment, there are basically four ways in which to measure telephone penetration: (1) using decennial census data, (2) using information from the cps, (3) using gbf-dime records, and (4) comparing household counts with counts of telephone access lines in given geographic areas. examples of the decennial census and cps measures of telephone penetration of each of jurisdictions served by bell atlantic appear in table 1, entitled, "telephone availabilit}' for selected areas: united states and bell atlantic ser\'ed states, 1980-1985". figures 1 and 2 illustrate changes over time in telephone penetration in these geographic locations. table 2 provides and indication of the type of statistic that can be developed using the access lines/households ratio. several conclusions are possible from a careful examination of these various measures of telephone penetration: 1. data from different questions result in different statistics; i.e., the two censuses produced different results partly owing to question wording. these differences. fall/winter 1987 iassisl quarterly 47 although found in a comparison of data from the 1970 and 1980 censuses, are illustrated in table 1 in the 'unit' and 'availability' meastires from the cps, 2. in addition to these difterences, however, there are major difterences between geographic locations (figtires 1 and 2) and, 3. differences have occurred over time (table 2), 4. difterences result from the sample selection process. table la reveals the differences in sample drawn from 1980 census data (colunms a and b); table lb illustrates the range associated with each confidence interval computed for the various point estimates of penetration computed from cps data. thus fai, we have examined the penetration rate for the total nimiber of households. when differences in ty-pes of households is taken into account, the variation in telephone penetration becomes even greater. the data in in table 3 are from a 1980 census public use microdata sample (pums). table 3 contains a cross-tabulation of households by age of householder, household income, and telephone penetration for the state of virginia. these data indicate that the presence of a telephone in the housing unit is not a random event, but rather thai one is less likely to have a phone if one is poor and/or young. figure 3 illustrates the ratio of telephone access lines to households in the bell atlantic region from the early 1950s to the 1990s forecast period. a surprising event occurred in the late 1970s: the number of access lines increased to the point of outnimibering households. how can this be? the answer is quite simple, it's called multiple lines. in the state of new jersey alone, more than 10% of households have ai least two lines, and these households are not poor and they are not yoimg. can one draw a random sample, utilizing rdd? yes, but a random sample of what? certainly one will not draw a random sample of the households in a given geographic location. one will draw a random sample of telephone hnes and, as we have seen, these are not evenly distributed across the population. does this matter? it depends on what one is trying to accomplish. a pollster, trying to predict an election, will want to examine age specific voting patterns and calibrate these to the universe of households with telephones, allowing for multiple lines, of course. since the people who vote tend to be those with higher incomes, and voter participation increases with age, you will reach the group you wish to sample using telephone interviewing, although you raa\ overstate the case, especially if you are ttacking republican candidates. (table 4) a marketing researcher, must be careful to calibrate the resulting data with census data, because the upscale population will be over-represented and other groups imder-represented in the sample. an academic or government researcher, conducting a survey designed to produce information for the development of legislation or the operation of a social program, would need to calibrate the data with census or cps data in order to enstiie a representativeness. in conclusion, it seems highly unlikely that telephone interviewing is a reasonable replacement for the bureau of the census traditional decetmial census methods of data collection employing a housing address hst, mailout questionnaires and enumerators. given the problems associated with the universe of telephone access lines, if telephone interviewing becomes the mainstay of data collection for the u.s. government, we will have lost one of the most important means we have of analysing the composition of our population, n fall/winter 1987 iassist quarterly attachment a methods of measuring telephone avallabllltr/pesetration (state level) i. decennial census of population and housing questionnaire 1970 ihl is there a telephone on which people living in your quarters can be called? i i yes what is the number? phone number no 1980 fhzd do you have a telephone in your living quarters? yes n "° caveats: * timliness of the data • validity i reliability issues ii. current population survey interview supplementary questions regarding telephone availability asked in march, july & november of each year of a state based sample of households caveats : * sampling error at state level * coverage of non-bell atlantic served areas iii. geobased files/dual independent map encoding matching telephone customer address records with street address records in the gbf caveats : * non-urbanized areas not covered * new housing not covered, although it can be inferred from telephone company records iv. residence acess lines/household estimates; use telephone company access line information in the' numerator and household estimates prepared by company demographer in the denominator caveats : * second line development; failure of business office to identify second line • accuracy of household estimates; ok for internal purposes, may not stand up in legal/regulatory setting fall/winter 1987 iassist quarterly 49 table 1-a s. s s ^ 5 ^ oi oi /l o r£ « a f^ 1/1 o o f^ ri^ ct! « $ i » s 3: $ $ > oj cs >s— *3 e '£ s. i<. o >, e fall/winter 1987 50 iassisl quarterly table 1-b confidence intervals associated with point estimates of telephone penetration for bell atlantic served jurisdictions prepared demogr by d. schmidley aphic studies x8638 (1) state (2) unit measure nov. 1985 95.6 (3) cvx (4) sex 1.42 (5) range 68^ (6) range 95+ % dc 0.0149 94.2 97.0 92.8 98.4 del 93.4 0.0119 1.11 92.3 54.5 91.2 95.6 md 95.3 0.0120 1.14 94.2 95.4 93.0 97.6 nj 94.1 0.0093 0.88 93.2 95.0 92.3 95.9 pa 95.8 0.0059 0.57 95.2 95.4 94.7 96.9 va 92.0 0.0156 1.44 90.6 93.4 89.1 -94.9 wva 65.1 0.0183 i.58 84.5 87.7 82.9 -89.3 formula from u.s. bureau of census se ^ x • cv x x xi = 68% confidence interval x2 = 95% confidence interval se = standard error x = penetration rate cv ^ coefficient of variation fdl/winler 1987 iassist quarterly 51 table 1-c tins tdble is from the u.s. bureau of the census current population surve/ ferlentage ufhcjuslhqlds with a tellfhunl e-y hqubcholdek s ailtable i c all races white blac^ hisf -itj ur.i t ava 1 ur.lt ave 1 unit ava 1 ur. t. 8 mqnih average total h0usehuld5 91 .7 9r> b 9r. 7 95 ,;j b'j. 4 b4 b bl 1 16-lm yks old 77.4 b6 80.0 85 6 57. 9 69 7 62 2^-54 yrs old 91 .b 93 b 97.5 95 1 b.j . 2 84 5 62 7 s5-5'9 yrs old 94.9 96 1 96.0 96 9 87.2 89 6 87 4 60-64 yrs old 95. 1 96 96.0 96 b b7.9 b9 8 b8 a 65-69 vrs old 95.9 96 7 96. 9 97 5 87.9 90 2 90 4 70-99 yrs old 95.5 96. 6 96. 1 97 -) 89. 1 91 3 bl, 8 november 83 total households yrs old yrs old yrs old 60-64 yrs old 65-69 yrs old 70-99 yrs old 16-24 25-54 55-59 91 . 76. 91. 95. 95. 93 7 93 1 95 . 78. b 83. 9 bci 84 1 80 2 86. 2 49.9 68.2 64 93 7 93 4 95. 2 78.7 83.3 81 96 1 96 1 97.0 86. 3 88.5 89 96 4 96 4 97. 2 89.5 90.7 87 96 2 96 5 97.0 87. 2 89.0 90 96 5 96 97.0 90. 1 92.3 85 85. 89. 90. march b4 total households 91 .8 93 .6 93 . 3 94.9 80. 1 b4. 1 80 .7 83.6 16-24 yrs old 77 .6 84 . ':i 80 . 3 85.5 57.9 71.5 59 . 66. 2 25-54 yrs old 91 .9 93 . 7 93 . 5 95.0 bo. 4 84.0 83 -> 85. 6 55-59 yrs old 94 .9 95 .9 95 .7 96.6 87.6 b9.9 88 . 7 90.5 60-64 yrs old 94 2 95 . -^ 95 9 96.7 81.7 b5.0 67 4 89.6 65-69 yrs old 96 1 96 6 97 97.4 87.8 69. 3 85 8 87.8 70-99 yrs old 95 3 96 -^ 96 2 97. 1 87.2 88.8 82 = 85.5 july 84 total households 91 6 93 8 93 2 95.0 80.5 85.3 81 1 64. o 16-24 yrs old 77 83 3 79 4 85.3 60.4 70.0 62 9 7.j.a 25-54 yrs old 91 7 93 8 93 4 95. 1 79.8 84.9 83. 1 65.6 55-59 yrs old 95 1 96 3 96 1 97. i 87.5 90.2 87. 4 91 . 4 60-64 yrs old 95 96 2 95 8 96.9 87.7 69.5 bb. 1 9.;. . 5 65-69 yrs old 96 4 97 1 97 3 97.9 b9.3 91.3 66. 7 90.6 70-99 yrs old 95 96 ^ 95. 9 96.9 89.6 93. 1 64. 68.5 november 84 total households 91. 4 93. 6 93. 1 95.0 78.9 84.0 61. 1 64.5 16-24 yrs old 76. 1 83. 4 79. 85.4 56. 3 70.8 60. 8 70 a 25-54 yrs old 91. 4 93. 6 93. 3 95. 1 78.5 83.3 83. 1 85 b 55-59 yrs old 94. 9 96. 2 96. 3 97.5 84.7 87.4 85. 3 88 z 60-64 yrs old 95. 6 96. 5 96. 5 97.3 90. 3 92. 1 86. ij 67 2 65-69 yrs old 96. 96. 7 97. 1 97.6 86.7 69. 1 96. 2 96 2 70-99 yrs old 95. 3 96. 6 96. 1 97.2 88.0 90.7 87. 1 bb march 85 total households 16-24 yrs old 25-54 yrs old 55-59 yrs old 91 .6 77. 3 95 . 80. 1 64.4 81.2 ej4. 1 84.8 59.6 70.0 62.4 67. 1 95.2 79.5 83.9 83.0 b:,.z 96.7 67.3 89. 1 86.5 b9. 1 fall/winter 1987 52 — iassist quarterly table 2 •a of households with telephones in the c&p served area of virginia year (1) residence access lines-ln service (june 30) (2) bell served households (july 1) (3) telephone penetration (2)/(3)x100 (4) 1960 564.824 763.000 74.0 1961 588.783 782.000 75.3 1962 616.381 802,000 76.9 1953 640,932 823,000 77.9 1964 669.764 843,000 79.5 1965 702,113 853.000 81.4 1966 736.798 882,000 83.5 1957 776.020 904,000 85.8 195b 819.205 933,000 87.8 1969 860.518 975,000 88.3 1970 896.711 1,011,000 88.7 1971 933,814 1.027,000 90.9 1972, 972.258 1.058.000 91.0 ~ 1973 1.025.160 1.104.000 92.9 1974 1,069.475 1,134.000 94.3 1975 1.097.747 1,158.000 94.0 1976 1,132,060 1.202,000 94.2 1977 1,166,408 1.225,000 95.2 1978 1,212,158 1.263,000 96.0 1979 1,255,582 1,305,000 96.2 1980 1,288,870 1,355,000 • 95.1 1981 1,319.998 1.386,000 95,2 1982 1,343,961 1,416.000 94.9 source: (1) data for residence access lines-in-service are taken from the virginia monthly no. 7 report, years 1960-80, and taken from the virginia no. 2705 report for years 1981-82; (2) household data is based on household counts from the decennial census 1960, 1970 and 1980; and are based on the current population survey and p-25 population reports in other years. these reports originate at the u.s. bureau of the census. october 19£2 fall/winter 1987 lasit.u quarterly 53 table 3 t(le»io«t penetration dy hoe and incow, viroinia, l96t 16 to 24 25 to 29 3« to 34 35 to 39 48 to 44 45 to 49 58 to 54 55 to 59 be to ia £5 to 69 78 to 74 75 to 79 risrilkr hshlder hshlder hshlder hshlder hshlder hshlder hshlder hshlder hshlder mskjes hs-1j1es nousenoic ircoae less tnan ilb.eee t.ii 8.79 e.83 e.92 e.93 e.9i i.ee e.94 e.95 e.98 e.7a (.75 e.67 8.92 8.9t 8.9i 8.98 8.9i 8.97 8.99 t.si 1.69 8.77 8.68 8.94 8.97 8.98 8.96 8.96 1.08 8.99 1.88 8.92 8.79 8.68 8.94 8.97 8.99 8.99 i.ea 8.95 8.99 8.99 8.94 8.76 8.69 8.94 8.98 8.96 8.99 8.96 1.88 8.95 8.95 8.94 8.78 8.98 8.94 8.97 8.99 8.99 8.99 1.88 8.99 8.99 8,94 8.68 8.98 8.95 8.96 8.97 8.99 8.99 8.95 1.88 8.99 8.94 8.63 8.94. 8.97 8.96 8.99 8.99 8.99 8.96 1. 88 i.e« 8.95 8.66 8.95 8.96 8.96 8.99 8.99 l.»« 8.99 8.96 8.95 8.94 8.69 8.97 8.97 8.97 8.96 8.96 8.95 1.88 1.88 8.96 8.94 8.9e 8.97 8.96 8.99 8.99 8.99 1.88 1.88 8.97 8.99 8.93 8.91 ili.m h9s9 8.97 iimm 19999 e.95 »2«m £4999 8.99 $e5«ee 29999 8.98 am 34999 8.96 i35w!« 39999 8.97 »4»w 44999 1.88 »45eee 49999 1.88 »5«e8e or (tore 8.97 total 8.93 data s mc data bst generated from the 1980 public use microdata sample by the pennsylvania tate data c;nter, at the request of a.d. schmldley, bell atlantic fall/winter 1987 54 — lassist quarterly table 4 telephone penetration by age of householder (%) (1980 census) age rate age rate 18-2a 83.8 25-29 92.0 30-34 94.5 35-39 95.3 ao-44 95.5 45-49 95.8 50-54 96.0 55-59 96.2 60-64 96.2 65-69 96.0 70-74 96.2 75* 95.6 «»is data was generated from the public ut.e microdata samples fum3 for the bell atlantic served jurisdictions of washington, d.c., delaware, maryland, new jersy, pennsylvania, virginia and west virginia, combined. fall/winter 1987 lassisl quarterly 55 figure 1 o o q n o z hi o lij 'iw/vl (vol '1^^ data *». from the u. ?. decennial censuses of 1970 and 19h0 uj h< o i< hlu z bj d_ lli z o i d_ lij _j llj b'^^i q e-6b k\\\\\\\\\v\\\\^^\\\\^^m\^>^ ^ t56 £'56 cs k^^<\^^^-h^^^v^^vh\\; ^ fall/winter 1987 56 lassis! quarterly figure 2 co to (d (0 at cd r> fh > > o o z 2; 'tihl data ab-from the current populc survey cps for the dates shown m <: e:z; w 0. c-96^^mm^mm^mms l§ ^^^^^^^ zn 6'ie| tl6 ig fall/winter 1987 iassisl quarterly 57 figure 2 residence access lines vs. households bell atlantic region millions 1^11'^10-issi 9876households ^^^^ 5^^^ residence access lines ik 43'i i 1 1 1 1 1950 '55 '60 '65 '70 '75 '80 '84 '90 fall/winter 1987 \^. difltfl sekhiqe 11 f1 sditifltek seniter^ gertrude j. lewis project leader center for computer and information services rutgers university profile rutgers, the state university of new jersey, which was chartered as queen's college in 1766, and was designated a state university in 1945, is a large university offering a variety of learning environments. today the university has over 47,000 students enrolled in six separate colleges on the four campuses in new brunswick, and the campuses at neward and camden each of which is over 50 miles away from the new brunswick site. there are 24 instructional divisions and about 16 affiliated research units. facilities and support for academic computing are managed by the center for computer and information services (ccis) which provides services to students and faculty who use computing for instruction and research purposes. services include non-credit education courses, a reference center, a newsletter, maintenance of terminals and remote job entry facilities, equipment loaned to classrooms, program packages support, data archives and data base management, system programming, documentation, accounting and billing. in 1971, the princeton-rutgers census data project came into being through the combined efforts and finances of both universities. since there had already existed a tradition of cooperation between the two universities on special data collections, they decided to share the purchase of 1980 census data jointly. the project was organized with the support of the center for research libraries and financial contributions from the libraries of princeton and rutgers as well as interested departments on both campuses. through start-1 program, the data use and access laboratory (dualabs), a non-profit organization, purchased the 1970 census tapes as they became available from the census bureau, processed the tapes, condensed the data, and sold copies at a reduced cost to its members. along with data modifications, several computer programs, known as the mod series, were developed to access the data and were installed at rutgers and princeton universities. all the tapes were stored at princeton, and rutgers copied only those pertinent to its researchers. in order to administer the project, both universities were responsible for publicity, training, and physical tape maintenance. the project became self-sustaining by charging outside organizations for programming fees and computer cost which then covered the purchases of new tapes. the project has continued to promote collaborative efforts of cooperation and support between the two universities. icpsr and roper memberships in 1965, a member of the political science department requested membership in icpsr. a few years later, membership in roper was established by political science department and then transferred to sociology. as the membership in these organizations became known, an increasing number of researchers were discovering that the data produced by national archives have intrinsic research and academic value. as interest in these memberships increased, it became apparent that the individual departments could not handle the workload. since the princeton-rutgers census data project had been functioning successfully, the ccis decided to centralize other data bases in the same manner. the library assumed the operational control of transferring all the relevant information and materials from the individual departments and of developing adminstrative and ordering procedures to facilitate the acquisition of data. data base advisory committee (dbac) as to the administration of the roper and icpsr memberships, the ccis favored the creation of the data base advisory committee to insure adequate communication between the ccis, the libraries, and the departments, to determine university policy concerning future data acquisitions and to select official representation to the icpsr and roper organizations. the committee, established by the director of the ccis, consisted of representatives from the ccis, the library and the political and social science departments on the new brunswick campus. after the first meeting, it was expanded to include representatives from the camden and newark campuses, also. when a discussion of the budget for the roper membership led to a joint membership by the rutgers and princeton libraries, a representative from princeton joined the committee. although the committee is limited to a maximum of ten members, guests are invited and welcome. this committee, which convenes one or twice a year, discusses allocation of available resources in the departments, decides who shall represent rutgers at the icpsr conference, and awards any scholarship to icpsr science programs that become available. communications by mail and phone are conducted continually with committee members on relevant data matters as the need arises. the new jersey state data center in anticipation of the large amounts of data produced by the 1980 decennial census, the u.s. bureau of the census established a state data center program throughout the country to improve access to and use of census data products. rutgers university, as a primary participant of the new jersey state data center (njsdc), documents, distributes, and publicizes these materials. the ccis has also made available the census software package (censpac), which is an all purpose statistical and retrieval program created by u.s. bureau of census to be utilized with the census data. funding the data activities fall within the applications group of the center for computer and information services which provides the facilities and the support services for academic (instructional and research) computer users. no salary lines are designated specifically for the data archives. our programmers are responsible for computer expertise on our software and hardware for all our computer systems. travel requests are considered on an individual basis depending upon the overall requests for travel within the budget limitations. this fiscal year, i felt very fortunate to attend the state data center conference, the association of public data users, and this meeting of the international association for social science information service and technology. but our staff participation in such events varies from year to year. the rutgers membership in the interuniversity consortium for political and social research and the rutgers-princeton joint membership in the roper center are financed through the library budget. if some departments request the purchase of data outside of these memberships, the computer center coordinates the search for funding it. overhead expenses for office space, secretarial staff, mail, postage, telephone, etc., are not being considered here because these were already in existance when data archiving activity came into being. the cost and maintenance of the computer equipment and cost of data processing come out of our current operating budget. although most of the computing with the machine-readable data is used on our ibm mainframe, some is also utilized on the vax 11/780, which has spss and scss, and the dec 2060 which stores a citibase data file. the various departments are, allotted a specific dollar amount for computing time which is based on previous year's usage and future estimates of need. staffing at the time of the implementation of the rutgers-princeton census data project, a half-time programmer analyst line was created to carry it out. when all the data activities were centralized at the computer center, the responsibilities were expanded and half-time of another programmer position was included. unfortunately, this past year, because of many changes in personnel and the addition of a new computer, we lost ground in this area. now less than one full-time line, shared among three people, is devoted to machine-readable data file activities. and it is not enough. the first nine months of this academic year abour 250 consultations were recorded or two or three requests on an average daily basis. this figure does not include quick references in the libraries or computer-related problems which may go to the statistician or aid station. (there are aid stations on each of the campuses which are staffed by students and the computer center staff to aid in debugging all user problems. ) sources of data during this academic year we have serviced more than 23 different departments on campus. their data requests have referred to many different studies in many different fields. requests for our census service are just as likely to come from outside the university, particularly non-profit county and state agencies as from within the university. the level of sophistication in handling mrdf's ranges from zilch to familiarity with statistical packages on the computer. all manner of problems come to the ccis aid stations, the statisitical consultants, and our staff during any given day. in general, the procedure in handling inquiries is fairly routine. first, we check our rutgers university guide to machine-readable data files to see if the file requested is already on campus. if not, the catalogs of icpsr, the roper center, and the bureau of the census are searched for the particular file or subject requested. data from the first two are easily obtained because of the memberships we maintain with these groups. the census inquiries require a different approach. requestors are directed first to the printed reports. if the information is available only on tape, the researcher is assisted in ascertaining what tape contains the needed data, what census geographic area will most suit the needs of the project, and which program should be utilized. if the data needed is from a source which requires a cash outlay, the staff assists the researchers in finding funding, if at all possible. it is difficult to determine which files are heavily used. the number of tape mounts does not give an accurate picture of how frequently the data is accessed. most users, after accessing the tape once or twice, create a subfile on their own and continue their analytic studies on the smaller file. the sophisticated users know how to find out the tape information without checking with us. at the present time, the most frequently used studies appear to be the stf 3a tape from the 1980 census of population and housing, the norc general social surveys, the american national election studies from michigan, and the national longitudinal studies from ohio state. rh16 dissemination of information training/workshops/seminars. ccis is always looking for ways to reach more rutgers users. in the beginning of each semester, two session seminars are conducted on familiarizing the researchers with the content of 1980 census of population and housing and other census products. another class is held on the general data archives to describe the types of data available for research and study purposes. special seminars or workshops are conducted at the request of individual units within the university and are tailored to their particular interests and needs. the close association with the reference librarians is reflected in a special seminar for the reference special interest group on the resources at the ccis with special emphasis on the 1980 census. publications/articles/documents. articles on datarelated information appear regularly in the ccis bi-monthly newsletter, but ccis also publishes a number of technical documents pertaining to machine-readable data files and the computer programs available to access them. our publications are all geared towards making data use easier for the university community. as an example of this type of publication, information was extracted from the master area reference file for the 1980 census of population and housing data and sent to the reference librarians in all the libraries on all the rutgers campuses. this computer output included not only the census geographic codes and corresponding area names, but also included an index by countries, an explanation of the symbolic codes and total counts for population, housing, and families. complementing this will be our output, probably on microfiche, on county and mcd by zip code for distribution to the libraries. finally, our most important publication is the previously mentioned rutgers university guide to machine-readable data files, an index of all machine-readable data files on campus. consultation services. for guidance and assistance in using any of the machine-readable data, the ccis offers a consultaiton service, free-of-charge, to direct users of the data, to help users select the computer program best suited for the user's need, to provide necessary program and technical documentation, and to assist in analyzing computer error messages should they occur. computer reference center. under the auspices of the ccis is the computer reference center (crc), a library of computer-related materials. all the codebooks , manuals, and reference books are located in the data archive corner to facilitate accessibility to the widest possible range of users. these codebooks themselves can be useful tools in data analysis, sometimes eliminating the need to access the files by computer. in addition, catalogs of data holdings of several institutions which collect and disseminate data are available as well as computer-related periodicals. an information specialist who maintains and updates these materials assists users in finding the information they need. future developments out future plans center on professional development for the staff, improvement of our excellent census service, and expansion of the contract programming activity. in addition to in-house training workshops, staff are encouraged whenever possible to attend conferences and workshops which enhance their professional skills. it is hoped that more money can be made available for such attendance in the future. along with attendance at these functions, staff are also urged to participate in the related organizations which sponsor these meetings, such as apdu or lassist. these organizations do invaluable work in fostering increased awareness of mrdf's both on-campus and in the wider academic world. such participation continued on page 22 iassist quarterly iassist quarterly 2013 15 abstract this article uses the author’s recollections and some of sue a dodd’s own publications to provide a brief overview of various activities leading up to the publication of her cataloging machine-readable data files: an interpretive manual in 1982. of particular importance was the airlie house conference on cataloging and information services for machinereadable data files of 1978. the article also highlights events that accompanied the development of cataloging rules for data files in the 1970s. it concludes with the author’s memories of dodd’s commitment to professional involvement, the development of standards, and the role of libraries. keywords: libraries, data libraries, library cataloging standards, data classification. in 1977 sue a. dodd obtained a master of science degree in library science from the university of north carolina, chapel hill. it is unlikely that dodd, who had undergraduate and graduate degrees from the university of kentucky, needed the msls for her employment. she had been a data librarian at the institute for research in social science at the university of north carolina for ten years and was also the associate director of the louis harris data center, which was also housed in the institute. she was probably supported and encouraged to study library science by richard rockwell who was the director of the institute’s data library from 1969 until 1976. she had also been engaged in several efforts to bring bibliographic standards to social science data files. dodd had been a panelist in the session on “problems on inventorying data: classification schemes” at the 1974 toronto conference on data archives and program library services. she would become the us chair for the iassist action group on classification of data. in any case, her study of library science demonstrated a desire to obtain professional insight in the classification of data files. in march of 1978 dodd was the co-chairperson/ technical issues session leader for the conference on cataloging and information services for machinereadable data files held at airlie house in warrenton virginia, supported by a grant from the [u.s.] national science foundation. her contributions to the conference included a report of the working group on technical issues titled “characteristics of machinereadable data files” and a reprint of her drexel library quarterly (january 1977 no. 1) article “cataloging machine-readable data files – a first step.” that article was probably used by the technical issues working group as a starting point for their discussions. dodd sue a. dodd’s lasting influence: libraries, standards, and professional contributions by ann s. gray1 their final report would lay the foundation for a chapter on what were called machinereadable data files 16 iassist quarterly 2013 iassist quarterly often said that her award winning book cataloging machinereadable data files: an interpretive manual, published in 1982, was a result of the airlie house conference, but her recognition that library procedures for bibliographic control could be applied to data files preceded that event. . the drexel library quarterly (dlq) article mentioned above provides a fairly comprehensive history of libraries’ involvement in cataloging and data files, beginning with an american library association/resources & technical services division/cataloging and classification section (ala/rtsd/ccs) ad hoc subcommittee established in 1970. this subcommittee labored for six years and worked with various groups engaged in cataloging efforts for data files. their final report would lay the foundation for a chapter on what were called machine-readable data files or mrdf in the second edition of anglo-american cataloging rules (aacr ii), published in 1978. (although the rules were published in 1978, they would not be implemented until 1981 for reasons that had nothing to do with the inclusion of mrdf.) the same article also details the work of the iassist action group on classification and its goals. in 1976 dodd had prepared a “working manual for cataloging machine readable data files” based on her interpretation of the ala subcommittee’s recommendations. the iassist action group’s first project involved establishing the feasibility of cataloging data files by having several institutions actually catalog some data files, ideally ones unique to their collection. dodd gave each a copy of her working manual. this project was probably undertaken in 1976, as dodd cites an action group memorandum of that date in the dlq article. in the 1970s there were a number of important changes in standards for bibliographic description of print materials. in 1967 the anglo-american cataloging rules (aacr1) were implemented at the library of congress. aacr1 was not that different from the library of congress’s rules for descriptive cataloging, published in 1949, and allowed libraries to continue to use forms for older entries even though they conflicted with aacr1. this was called superimposition. only newly established entries would use the new forms set by aacr1. in 1974 a revised international standard bibliographic description (isbd) was published for monographs. standards and revisions for other types of materials would follow. the isbd provides standards for the form and content of bibliographic descriptions. in 1978 the second edition of the anglo-american cataloging rules (aacr 2) was published and it incorporated the revised isbd for both monographs and serials and included other types of materials, including machinereadable data files. but the change that caused many libraries to refuse to implement aacr 2 was that superimposition was no longer allowed. for example, “u.s.” became “united states.” “united states. department of commerce. bureau of the census” became “united states. bureau of the census.” no matter how many cards for older census bureau publications were in the catalog, all new publications would have to use the new format. furthermore, a serial—which might often change its title or publisher—would now be cataloged using the current title, omitting designations such as bulletin or magazine. in the age of card catalogs these changes would result in very different locations of the same materials published in different years or the library would have to re-catalog all of the older material to the new standard. librarians referred to this as “desuperimposition.” considering the problems this would cause for libraries, cataloging of mrdf was a minor concern and probably involved people who had not cataloged using the older forms, persons new to the field. of equal importance was the development of the machine readable cataloging (marc) programs at the library of congress. in 1966 the library of congress launched its pilot project for the marc system. using feedback from this test, the marc ii format: a communications format for bibliographic data was published in 1968. it would be adopted as a standard by several american library association divisions as well as other agencies. in 1971 it was given the american national standard designation ansi z39.2-1971. it is not necessary to go into the details of these standards, but it is important that they exist. standards are altered over time, but they provide a framework and a community that uses and applies them to make changes. ansi z39 and other communication standards provide a means to encode, organize, and retrieve information in a common or shared environment. in 1967 a group of libraries in ohio formed an organization to share cataloging and catalog information using a computerized network. by the early 1970’s this system, using the marc standards, was operational. it provided a shared database of catalog information and the production of catalog cards. the ohio college library center would expand its services to libraries outside of this network and change its name to oclc. having catalog records in a machine-readable format lead to the end of the card catalog. about the time the dodd manual was published, microcomputers, as they were called, were becoming common. the focus of the manual had always been social science data files. but school librarians were acquiring computer programs and other types of files for use on microcomputers; and thus there was a need to adapt aacr 2 chapter 9 for cataloging those types of materials. dodd worked with ann m. sandberg-fox on a follow-up manual for cataloging microcomputer files. in 1979 i began my own career as a student in the university of north carolina school of library and information technology. before that time i had worked for a number of years as a cataloger in a small academic library using the oclc system. there was a version of the marc system for use at the library school and i learned how to enter, change, and retrieve text as well as how to set up database structures and options, and all the system controls that went with a batch process using ibm 360 machines running mvs. the marc programs were very flexible; it was fairly easy to define fields, subfields, indicators, and output formats. but like many computer systems of the day it was merciless if there was an extra space or extra slash mark or missing comma. it was command driven and one had to know the commands. because i had somewhat mastered the system, i was offered a temporary position at the institute for research in the social sciences data library where the marc programs had been used for many projects. the job was to catalog the 1970 u.s. census files based on a machine-readable version of the census bureau’s directory of data files (abramowitz and aldrich, 1979). the machine-readable version of the directory contained print commands which could be used to identify separate sections as containing entry information, such as which part is the title, which part is the edition statement, the collation, notes, etc. i would use the marc system, design the fields, and add information if needed. there i met sue dodd who was finishing up her manual. iassist quarterly 2013 17 iassist quarterlyiassist quarterly sue and i, sometimes with others, would get together every day to review her interpretation of aacr2 chapter 9. she was very keen to know how a working cataloger would use the manual and what questions such a person might have. although her preliminary working manual had been used by the iassist action group project, she had not had the opportunity to actually talk to the catalogers since the project had been handled through the mail and only a few problems had been found. i had a lot of experience cataloging monographs and some training in cataloging serials. i don’t recall finding any problems with her interpretation and when she asked if something should be done differently, i always concurred with her decisions. often sue would explain to me why bibliographic control was needed, why libraries were important to social science data files, and who made what contributions to this effort. later i would have a staff position at the data library and work with sue for three years where my education would continue. the primary purpose of the iassist action group on classification had been to facilitate access and promote the use of social science data. bibliographic control would provide authentic identification of specific studies. within social science publications, data sources were not cited as published works, but were often mentioned in the text using general and ambiguous names, such as ‘the michigan study’ or ‘census data.’ writers could not be expected to cite their data if no publication information was available. cataloging data files was not the only tool sue promoted. she often spoke of cataloging in production efforts like those of patrick bova in the documentation for the general social survey. another project she was enthusiastic about was the generation of catalogs of data holdings. she had copies of all she could locate. sue was almost unique in her support of libraries as important resources for data. she believed that libraries wanted to be involved in data services and that they should be involved. in 1981 there was little evidence to support her opinion but she would be proved correct. sue also instructed me in two other principles she held dear – the importance of standards and the value of professional involvement. sue was a generous mentor and an important contributor to the professional development of data services within iassist and library organizations such as the american library association and the research library group. these professional communities allowed her to be part of the process by which standards were advanced. she understood that standards would only be applied if their complexity was justified by their utility. in library catalogs the term machine-readable data file for type of material was replaced by computer file which was replaced by electronic resource. but sue dodd’s legacy is not lesser because of changing terminology. besides, i don’t think the transition from mrdf to electronic resource is an improvement. i think sue would agree with me. references abramowitz, molly, and barbara aldrich. 1979. directory of data files. washington: u.s. dept. of commerce, bureau of the census. notes 1. ann s. gray worked in data services at the university of north carolina, chapel hill, institute for research in social science; cornell university, cornell institute for social and economic research; and princeton university, firestone library. she has been a member of iassist since 1984. she is retired. and the walls come tumbling down: the converging destinies of the rutgers university libraries and the center for computer and information services by linda langschied' & gertrude j. lewis on october 16, 1990, the following mandate was issued from the ofnce of the president of rutgers university... "complex interrelationships among print information, electronic data, resources, and telecommunications have opened enhanced possibilities for access to both information and communication for scholarly and management purposes. this sophisticated information and communications environment has, as a result, led to a convergence of many of the functions of libraries and computing services.... bringing computing facilities and libraries under the same management will enable us to build on existing strengths, avoid duplication, and coordinate planning, thereby enabling us to improve service to users." now the directors of the academic and administrative computing centers report to the associate vice president for information services, who is also the associate university librarian for technical and automated services. he reports to the vice president for information services and the university librarian. in the beginning.... two decades ago, representatives from rutgers and princeton universities met at firestone library, princeton, to discuss means for acquiring the 1970 census of population and housing data. the result of that meeting was the formation of a group calling itself the princetonrutgers census data project. a number of agreements were reached concerning costs, finances, billing for services, and procedures for acquiring data and software. three hundred reels of census data and a number of utilities to aid in accessing the data were purchased at that time. the census data was stored at the princeton university computer center and upon request copies were made and housed at rutgers university center for computer and information services (ccis). training seminars were made an integral part of the program. funding for the project came not only from the computer centers and libraries of the two universities, but also from some individual departments. although not all contingencies were considered in the original agreement, the princeton-rutgers census data project was founded in a spirit of cooperation and the belief that the primary objective was to make the 1970 census data available to members of the academic community as quickly and as economically as possible. as the census data project developed, a "census packet" was sent to rutgers hbraries and key academic departments. originally it was felt that interested people should go to the library first and not directly to ccis. library staff would help patrons to understand available census data in printed and magnetic tape form. at first, the census data tapes were stored at princeton, and the rutgers users paid for programming and computer time to princeton. thus began our road of cooperation between the ccis and the university libraries that continues today. the current status of macinne-readable data files... now the rutgers branch of the princeton-rutgers census data project houses its own data tapes and has been active in three major areas: education, consultation and data retrieval. to facilitate the use of census materials, the ccis has published many technical documents for faculty, staff and students that deal with locating, accessing, and analyzing census data. on the reference shelves of the libraries, are publications created by the ccis staff in which the researchers can locate the names and corresponding codes of the census geographic areas in new jersey or look up the index of machine-readable data files available to the rutgers community. information on the data acquisitions are publicized in the bimonthly ccis newsletter. however, census data is only a part of the extensive machine-readable data files (mrdf) collection services provided by the ccis in conjunction with the libraries. the roper center the libraries and computer centers of both rutgers and princeton university participate in a joint membership with the roper center's international survey library association (isla). the universities are entitled to the research services of searching the archives, producing tabulations of data analysis, and acquiring machinereadable data sets. as part of this venture, we also subscribe to the pubhc opinion location library (poll) database. 22 assist quarteriy the inter-university consortium ofpolitical and social research membership in the inter-university consortium of political and social research (icpsr) provides not only the acquisition of data and accompanying documentation, but also the opportunity to attend workshops in the icpsr summer session. over the years, the librarians, computer personnel, and academic researchers have attended these classes and brought information back to our researchers on icpsr's expanding resources. when the ccis orders a data set from icpsr, we receive a magnetic tape and an accompanying codebook in hardcopy. the tape is stored at the computer center and is available to anyone on or off campus who has a computer account the codebooks are shelved in the ccis computer reference center (crc), which is a small reference room, open about eighteen hours a week. data base advisory committee as an increasing number of researchers discover the enormous volume of data produced which have intrinsic research and academic value, they realize that these data are unmanageable without the use of a computer. in 1976 through the coordination of the ccis, the political science and sociology departments, and the library, the responsibility of the handling of icpsr and roper machine-readable data was transferred to the ccis. although the roper membership originated in the sociology department and the icpsr membership started with the political science department, the costs of these memberships come out of the library budget the ccis became responsible for the acquisition and maintenance of the various data sets. it was felt that a central clearing house for databases would eliminate duplicate purchases that had occurred. under the ccis, all communications from census, icpsr and roper that might be of academic interest would be forwarded to the librarians, the appropriate department chairperson or representative. a data base advisory committee was created to include participation with the various social science disciplines and with the libraries (again both rutgers and princeton universities) in order to determine general policies concerning acquisition and access. this committee meets periodically to keep up with the current activities in the field. as usage has expanded, representatives from other departments who wish to use the data can join the committee or participate as guests. the princetonrutgers census data project is no longer limited to its original mandate of providing the 1970 census information. it is now part of the overall data archive program which also incorporates the icpsr, the roper center and the new jersey state data center. the new jersey state data center in 1977, with the advent of computer sophistication and the use of the 1970 census data in machine-readable form, the anticipation of large amounts of 1980 census of population and housing data prompted the census bureau to develop plans for improved services to data users. these user services included access to census data in reports, computer tapes, microform, trainings, and consultation. the basic concept involved stale-related organizations operating data delivery and user service facilities with guidance and assistance from the census bureau. one of the more important resources and services to be made available was the federal depository library program, through which many libraries receive census bureau publications at no charge. during the time of the 1970 census data, these services had fallen short of users needs. not all processing centers offered training, not all states had census processing services, and not all locations offered consultation on technical matters. the state data center program was proposed to close the gap in these user services. rutgers university, as a primary participant of the new jersey state data center, has fulfilled its obligation to provide these services. continuing education... as part of the regular education series, conducted by the rutgers ccis, a general introductory seminar in data archives is given. in addition, special seminars are also available on an individual request basis. many of these seminars have been held with the librarians to help identify inquiries that we get in common, such as: • a class for the reference librarians of the university emphasizing how to find out what is available by making use of different reference materials; • a session on exposure to increasingly sophisticated techniques of research and manipulation of our machine-readable data files as part of a library instruction program to enhance the research capabihties of undergraduate honor students; • a session on data archives for the librarians and the researchers on how to bridge the gap between the traditional hbrary resource materials and the accompanying computer related material; • a seminar on how to direct prospective users through documentation, codebooks and accompanying printed material; • a class for the doctoral students of the school of communications, information and library studies on how mrdfs will help the researchers. in each case, the content of the lecture has been tailored to the interests and level of computer expertise of the group. these seminars enhance the interaction of the librarians and the computer center personnel because they often confer by phone when handling inquiries on the data archives. hundreds of researchers have been summer 1990 assisted in this way. we began to realize that as the university library utilizes the computer more and more for information retrieval, bibliographic searches, online cataloging and other functions, there will be an ever increasing working relationship between the library and the ccis. the current status of other projects... it is to rutgers good fortune that there were people within the two units who were keenly aware of the overlapping nature of our work, and willing to take on the extra work of communicating across departments. and so we went forward with a number of cooperative ventiuies in the early years. the census project got us going, and even before our merger, the ccis and library worked out several projects together: icpsr codebook one of the first joint reference-related projects we arranged was for the research libraries to acquire additional copies of icpsr codebooks provided at the ccis. any time a codebook is ordered by ccis, the icpsr automatically sends a second copy to the library; in the event that the codebook is only available on tape, the ccis runs off a paper copy for the library. thus, the ccis copies serve as a stable reference collection, and the library copies increase availability— through both our extended hours of operation and because we circulate the codebooks. the circulation policy enables our faculty and students in three distant campuses to have ready access to the codebooks through our document delivery systems. formerly, researchers from camden and newark had to travel to new brunswick just to see the codebooks. online end-user searching end-user online searching at rutgers has been addressed mutually, as well. programmers at ccis developed a special communications software for the libraries' online end-user search service, entitled "knightsearch." this software permits a masked password logon with an automatic self-destruct after one paid usage, and automatically terminates the session after the prescribed period of time. ccis contributed not only their programming expertise, but also their student microcomputer centers machinery as search terminals. the project was only partly successful: the microcomputer centers ultimately proved to be unsuitable environments for search; however, the software continues to be used in large science courses in the departments' own labs. local mounting ofdatabases as the library further sought to enable patrons to gain access to online information, we began to investigate the mounting of databases locally. as an initial test project, for which the bbraries invested seed money, we decided to mount isi's current contents database, through the brs onsite program. once again we turned to ccis for their assistance in determining the technical needs, and for use of their mainframe. having current contents available locally will allow researchers to search for titles directly instead of going through one of the commercial database vendors, like dialog. although both of ccis's mainframes, an ibm compatible and the vaxcluster were very heavily used, it was determined by the ccis systems staff that if we purchased an additional disk, we would be able to run current contents on the cms operating system of the ibm. however, when we tested the system, with two groups of about fifteen librarians each, using a bench mark program developed by the librarians, we brought the cms operating system almost to a halt by this time, we were already committed, by contract, to both brs and isi, and what we mainly had to show for all our efforts was a system with a response time that was too slow to be acceptable: to quote one member of the test team, "by the time you get a response, the contents aren't current anymore." most unfortunately, we had lost the chance for a trial period, where we could have detected the problem before committing ourselves in a contract, because of the amount of time that it took to communicate up two separate administrative ladders— the library's and the ccis's. this project serves as an example of how front-line efforts need coordinated support from administration in order to succeed. online interface to data the libraries have provided an online interface to ccisheld polling data from the roper center since the introduction a few years ago of poll, the public opinion location library. the library offers searches of the database to identify appropriate surveys, and the ccis provides data retrieval. similarly, the library can search the icpsr guide to resources and services for researchers online before they approach the ccis for data. furthermore, the library subscribes to and performs searches of some numeric databases produced by the u.s. government, and the state of new jersey. for example, the new jersey state data center/business & industry data center electronic bulletin board provides access to data prior to publication. data is downloaded in either ascii or lotus 1-2-3 format for post-search manipulation. again, we see a blurring of distinction between the kinds of information provided by the libraries and the ccis. one important addition will be the availability of the government information on cdrom. some of these data will be distributed to depository libraries by the government printing office. users can download the data and access with a commercial data base product. again lassist quarteriy librarians and computer programmers will investigate and support this media. student microcomputer project the creation of the student microcomputer project was another result of sharing the resources of the two units. about five or six years ago, due to a tuition supplemental, different university governing bodies, which included student representatives, voted on buying microcomputers for non-classroom use. one of the student stipulations was to have them placed in the libraries so access would be during normal library hours and library resources would be available to them. macintosh and at&t microcomputers were purchased and placed in four locations on the main campus. while the libraries agreed to provide precious physical space, the ccis agreed to maintain the computers and give general support as a result, this has proven to be one of the most successful projects that has greatly benefited students. the microcomputer areas are staffed by trained students; software and documentation are available on site. each semester, seminars are held on operation of equipment, word processing, spreadsheets, databases and graphics. after the initial expenditure, the university provided funds in its ongoing budget to support the program. software information center in almost any field, computers have become as essential as books, and in fact, in some instances, are even replacing books. to address this issue, the software information center (s.i.c.) was estabushed as a centralized forum for identification, evaluation, and sharing of software for the entire university. it also serves as a liaison to other consortia engaged in academic software development and exchange program. at the center, faculty and graduate students can preview various software packages; the range of courseware available is wide and holdings are constantly being expanded. the programs are available for both ibm and macintosh microcomputers. in addition, assistance in using authoring programs is provided to help faculty develop their own courseware. the software collection is cataloged in the integrated rutgers information system (iris) as a joint project with the university libraries. iris is a computer database that contains the records for books and other material cataloged for the rutgers libraries and networked by the ccis. in the furure ... networking networking, which is a scheme for connecting computers, is now on the horizon as the most important aspect of communicating between and among clients. we can reach most national and international networks but not all campus buildings. most academic buildings on the campus where the main computer center is located are linked by a campus-wide broadband system. in addition, there are networks at each of the four remote computer locations. now, with the merger of the libraries, academic computing and administrative computing, we are embarking on connecting all personnel electronically so that everyone will have access to an electronic mail box. this certainly will stimulate demand for computers. the logistics have to be worked out: significant upgrades will be required for existing systems; standards will have to be established to determine which system(s) will be used and how to handle capacity issues. we have a big job ahead of us. this can only be done through cooperation among those departments that provide information services to the university. national centerfor machine-readable texts in the humanities rutgers and princeton universities have received grants from the national endowment for the humanities, the andrew w. mellon foundation, and the new jersey committee for the humanities to joinuy undertake the planning for a national center for machine-readable texts in the humanities. project staff from rutgers university includes the associate university librarian, an associate director from ccis, and a member at large along witii similar personnel from princeton university and representatives from the research library group. during the course of the planning period the project staff will be investigating issues relating to the establishment of a cooperative center which will act as a central source of information on humanities data files and a selective source of data files themselves. the initial goals of the center are to continue the on-going inventory of machine-readable texts; to catalog and disseminate this information; to acquire, preserve and distribute the textual data files which otherwise become generally unavailable; to distribute such data files in an appropriate manner; and to establish a resource center/referral point for information concerning other textual data. other issues such as initial setup costs, administration of the project and the feasibility itself are also under investigation. the center plans to complement and enhance these collections by bringing bibhographic control to existing data files. to that end the project staff will be networking wiui other centers to establish appropriate means of collecting inventory data for the cataloging of archival holdings. library committee on cataloging machine-readable data files the ccis machine-readable data files committee was formed by the technical and automated library services to study how the cataloging of the data files housed at summer 1990 ccis should proceed. the purpose of this study is to make the university aware of this collection, to strengthen the existing liaison between the libraries and the computer center, and to make a contribution in the area of computer file cataloging. as yet, the mrdfs are not accessible to the rutgers community via the ubraries online catalog system neither are a large number of codebooks which accompany these files. computer personnel have and will continue to work with the librarians on creating these catalog entries. recommendations, as a result of this study, are that cataloging the data files and codes should be performed by librarians and administered by the special formats cataloging section. it is just a matter of time before this project will begin. problems and solutions researchers need information, and do not particularly care where the information resides within the university. as is implied throughout this paper, some of the distinctions that the library and ccis make between our services tend now to be rather artificial, maintained more out of habit than by design. we need to rethink our roles from the point of view of the patron in need of information, to break from tradition when appropriate, and create information systems that are responsive to our constituency. across the years, the libraries and ccis have both committed staff, time, machinery, and hard cash to common causes. we have done a great deal, voluntarily, and together. yet it must be said that in what we did, there were often problems; and moreover, there was so very much more left to be done. on thefront lines... yes we communicated. and yes, we did not the problem, i believe, on the library's part went to responsibility. while it is wonderful for individuals to voluntarily take on new and cooperative projects with another unit, the lack of formal responsibility led to things simply falling through the cracks. for example, when we began to collect a duplicate copy of the ccis copy of icpsr codebooks, back-ordering was assumed, and the availability was publicized enthusiastically in library and ccis newsletters. so, we were very red-faced when a faculty member from our newark campus, some thirty miles distant, responded to our much-publicized tout about availability of codebooks, and asked for the entire run of codebooks for the annual housing survey. we were able to provide only codebooks from the past few years, as our subscription turned out not to be retroactive. the oversight was not caught because there was no-one who's job it was to catch it! and this is just one example of a small detail which nevertheless hinders information access to the researcher, and erodes our own credibility, as well. one step that the alexander library, which is the research library fw social sciences and humanities research at rutgers, has taken to try to address this situation, is to designate a librarian to serve as coordinator for non-bibliographic database services. besides working with the above-mentioned numeric databases, that person's natural function will be to work in concert with appropriate ccis staff to see that researchers working along those "blurry" lines are guided along the most direct path to needed information. on the administrative end... rutgers university's former president, the late edward bloustein, was a primary mover in the merger of the university's computing activities, articulating the need for coordination of the complex interrelationships between print information, electronic data resources and telecommunications. at a time when resources are limited and budgets strained, this plan of pooling resources, while perhaps not showing actual dollar savings, is intended to produce "intangible" savings via a streamlined and more efficient organization. so, in order to coordinate the complicated, yet obviously converging, activities of the computing organizations on campus, a recent restructuring brings the libraries and the computer services under one organizational umbrella. all units now report to the vice president for information services at rutgers, and university librarian, joanne euster. president bloustein also appointed peter graham, associate university librarian for technical and automated services, to serve as associate vice president for information services, in addition to continuing with his current responsibilities. joanne euster explains the administrative rationale for restructuring in this way: "economies of scale, as well as an apparent fading of the distinction between administrative and academic computing suggest that those functions should at least share some of the same pool of exjjertise and infrastructure. the goal should include minimizing duplicate input of information, ensuring the integrity of shared databases, being cost efficient, and making possible optimum individual control of one's own data access and utilization." basically, this concept recognizes that the three affected units have tasks that are distinctive, and which are well-served within the unit, but that there are also certain aspects of the operations. for example, the academic and administrative sides of the computer share cpus; the library and academic side share data; and the administrative side provides its registrar's and personnel tapes to the hbrary for its online patron file. assist quarterly already, the administrative restructuring is having an effect on front-line services. our plans to mount the current contents database have been resurrected thanks to the associate vice president for information services' decision to purchase a new mainframe computer that will fit our needs. the organizational changes provide new opportunities to expand the cooperative tradition of the libraries and the ccis, building on existing strengths, and exploring new ways of improving our services to rutgers faculty and students. our aim is to provide a "seamless" system of information services to data users. in conclusion... computing services have changed. library services have changed. research methods have changed. whether it becomes the responsibiuty of the library or the computer center to respond to change and meet new challenges is a moot issue. the clients continue to require information and increasingly this information is available in machinereadable form. although the library and the computer center are independent of each other and deliver different forms and types of services, some of the information services overlap. the computing environment grows more complex each year and correspondingly, the responsibilities of our staff become more demanding. because new technologies require us to be technically proficient in these areas, the merger of the hbraries and the academic and administrative computing centers make sense. because of the current budget restraints, these tasks must be clearly defined. because of different perceptions of the type and level of service needed, there will be challenges in meshing these services. but with our long history of successful collaboration, the outcome of this new relationship is assured. and the faculty, students and staff will be the beneficiaries. ipresented joindy to the international association for social science information service and technology (lassist) conference on "numbers, pictures, words and sounds: priorities for the 1990's" poughkeepsie, new york on june 2, 1990. linda langschied, information services librarian, alexander library, rutgers university, new brunswick, new jersey, &gertrude j. lewis, deputy associate director, center for computer and information services, rutgers university, piscataway. new jersey 3480 cartridges. for those interested in 3480 cartridges: ntis (pbs 233135) is selling for $12.95 national archives technical information information paper no. 4 "3480 class tape certridge drives and archival tape storage: technology assessment report." call 703/487-4600 or write to document sales ntis, springfield, va 22161. this paper covers the "mechanical & technological future of the systems" & "provides valuable information to data center managers data librarians, and archivists, in fact to all who are concerned about the long-term storage of machine-readable data." summer 1990 message from the president the revised constitution as many members know, the review and revision of the constitution of lassist has taken several years to complete. the constitution was approved at the may 18, 1983 business meeting in philadelphia. several calls were made to the membership prior to the meeting for input, criticisms and changes, and many members took the time to provide written comments. some minor changes were suggested at the meeting last may, and these have been incorporated. as i have mentioned several times, work on the review of any constitution is difficult and time consuming. however, the importance of a good, well structured constitution cannot be stressed enough. the transfer of the treasurer this year from ed hanis at the university of western ontario to jackie mcgee at rand corporation in santa monica has underlined some of the difficulties which confront international organizations. some of these kinds of difficulties can be overcome if the constitution is clear and well structured. the revised constitution has defined the composition of the administrative committee, the officers of the association, and their functions. article xii provides a more detailed description of the duties of the president, vicepresident, regional secretariats and appointive officials. section 5 of article xii sets up five standing committees and outlines the composition of those committees. i would hope that at the next lassist business meeting members interested in serving on these committees could be identified and the committees could begin to take an active role in advising the administrative committee and the membership on matters within their scope of responsibility. this can provide an opportunity for more participation from the membership. each standing committee requires two members from the regular membership of lassist. any member interested in serving on one of the committees should contact either the secretariat in the region or a member of the administrative committee. the nominations and elections committee will be active in 1984 as the election of the administrative committee is slated for the fall of 1984. i hope that all members will take the time to read the constitution and will think seriously about taking an active role on the committees. many people were involved in the review and revision of the constitution, and i would like to thank all of them for their work. particular thanks go to harold naugler, who prepared the first draft, and to the constitutional review committee chaired by carolyn geda, who spent considerable time working through all the detailed changes. ms. sue gavrel president, lassist 17 article i name the name of this organization shall be the international association for social science information services and technology/association internationale pour les services et techniques d' information en sciences sociales, hereafter referred to as "lassist". article ii headquarters the official headquarters of lassist will be located with the treasurer. article iii objectives all activities of lassist will be based upon the following objectives: 3.1 to encourage and support the establishment of local and national information centers for social science machine-readable data. 3.2 to foster international exchange and dissemination of information regarding substantive and technical developments related to social science machine-readable data. 3.3 to coordinate international programs, projects, and general efforts that provide a forum for discussion of issues relating to social science machine-readable data. 3.4 to promote the development of standards for social science machine-readable data. 3.5 to encourage educational experiences for personnel engaged in work related to these objectives. article iv activities to accomplish the objectives of lassist, some or all of the following activities may be conducted with the approval of the administrative comailttee on a national or regional basis and the submission of an appropriate report: 4.1 committees and groups committees may be established and groups of members organized to undertake specific tasks, to find solutions to specific problems, to develop and compile relevant material for specific projects, and to disseminate information on specific subjects. 18 4.2 conferences, workshops, seminars, training sessions members may convene organized efforts on any subject consistent with lassist objectives. 4.3 publications a newsletter will be published and regularly circulated to all members, as well as to others wishing to subscribe. other kinds of publications may be produced on occasions. 4.4 cooperation with other organizations efforts will be made to cooperate with other organizations in joint projects and activities when these are consistent with iassist objectives. 4.5 other other activities that advance the objectives of iassist may be undertaken from time to time. article v membership 5.1 the membership shall consist of regular and student members, and shall be open to such persons as are interested in supporting the objectives of iassist. 5.2 membership in lassist shall include a subscription to the newsletter. 5.3 resignations of any members shall become effective immediately upon receipt by the treasurer of lassist. resignation shall imply forfeiture of the annual membership fee. article vi finances 6.1 the fiscal year of lassist shall begin 1 december. january and end 31 6.2 membership fees for regular and student members shall be annually to the treasurer by 1 march of each fiscal year. paid 3 the rate of membership fees may be changed by a two-thirds vote of the members on a mail ballot or during the business meeting of the general assembly. mail ballots will be undertaken between october and december of any calendar year. the results of such ballots or votes will go into effect on 1 march of the following year. in the event of a vote during the business meeting of the general assembly, the membership will be 19 informed prior to the business meeting and proxy ballots vn.ll be made available. article vii governance 7.1 general assembly lassist shall consist of a general assembly composed of all regular and student members. the general assembly will be organized by geographic regions. the establishment of a region must be approved by the administrative committee. 7.2 functions of the general assembly the general assembly will establish general policies for iassist and elect the members of the administrative committee, as well as the officers of the association. each region will, in addition, elect its own administrative officer who will be known as the regional secretary. 7.3 administrative committee the administrative committee will be the executive body of iassist, and shall be composed of at least 10 members elected by the general assembly from its membership. the composition of the administrative committee will reflect the geographic distribution of the members of iassist and will be based on the number of members in each geographic region; the regional secretaries; the immediate past-president of lassist; the president and vice-president; and the treasurer, the editor, and the secretary-archivist, the last three individuals having been appointed by the president with approval of the administrative committee. the elected members of the administrative committee, including the regional secretaries, will serve a three-year term and may serve no more than three consecutive terms. 7.4 functions of the administrative committee the administrative committee will implement policies, develop future directions, and coordinate activities for iassist. the administrative committee will organize the general assembly into geographic regions, determine the number of administrative committee members from each geographic region, and call meetings of the general assembly at least once every year. the administrative committee will also establish committees and groups as required. 7.5 officers of the association the nominations committee will propose candidates for the offices of president and vice-president, co be voted upon by 20 the general assembly. these officers shall serve a three-year term and may serve no more than three consecutive terms. 7.6 role of the officers the officers of lassist will be responsible for the conduct of business of the association between meetings of the administrative committee. 7.7 executive committee the executive committee will consist of the officers, plus other members of the administrative committee as required and designated by the officers. article viii meetings 8.1 the annual meeting of the general assembly shall be held at a time and place chosen by the administrative committee. 8.2 special meetings of the general assembly may be called by the administrative committee. 8.3 the secretary shall give notice to the members as to the time and place of the annual meeting or special meeting not less than two months prior to the scheduled meeting. 8.4 a quorum shall consist of 40 members. article ix elections 9.1 a nominations and elections committee will be appointed by the administrative committee. 9.2 the nominations and elections committee shall conduct an election in each geographic region for officers of lassist, members of the administrative committee, and the regional secretaries. members within each designated geographic region shall only be entitled to nominate and vote for the regional secretary in their home region. however, all members will be entitled to nominate and vote for the officers of lassist and the other members of the administrative committee. in the event that competitive circumstances do not exist for a regional secretary, the regional secretary may be appointed by the administrative committee. 9.3 a public call for nominations will be sent out by the nominations and elections committee. voting will be conducted by mail ballot. elections will be held every three years. 21 article x amendments the constitution of lassist may be amended by a two-thirds vote of the members on a mail ballot, such ballots to be undertaken between october and december of any calendar year, the results of such ballots to go into effect at the following year's annual meeting of the general assembly, provided that: 10.1 notice of the proposed amendments shall have been given in writing to the standing committee on constitutional review with the written support of at least five (5) members in good standing of the association; and 10.2 two month's notice of the proposed amendments is given in writing to all members of the association prior to the conduct of the mail ballot. article xi termination iassist may be dissolved by a majority of the members. all property and funds of lassist will be transferred to a branch of unesco to be determined by the administrative committee. article xii by-laws section 1 duties of the president 12.1 the president shall: (i) be the principal officer of lassist; (ii) provide leadership and guidance in the realization of iassist' s objectives; (iii) preside at all meetings of the general assembly and the administrative committee; (iv) be an ex-officio member of all standing committees and shall coordinate their activities; (v) represent lassist in its dealings with external bodies and agencies, particularly those at the international level; and (vi) report on the state of iassist at each annual meeting of the general assembly. 22 section 2 duties of the vice-president 12.2 the vice-president shall: (i) perform the duties and exercise the powers of the president in the absence or disability of the latter; (ii) assist the president in recommending measures to further the objectives of lassist when and as often as requested; (iii) be an ex-officio member of all action and interest groups and coordinate their activities, and be responsible for proposing the coordinators to the administrative committee and maintaining regular contact with such action and interest groups throughout the year; and (iv) in the event of the resignation, death, or incapacity of the president, succeed as acting president for the duration of the then president's term. section 3 duties of the regional secretaries 12.3 the regional secretaries shall: (i) be the primary officers of lassist in their respective regions, working closely with the president of lassist; (ii) provide leadership and guidance in the realization of lassist' s objectives in their respective regions; (iii) represent lassist in its dealings with external bodies and agencies, particularly those at the national level; (iv) serve as members of the standing committee on membership; (v) attend all meetings of the general assembly and the administrative committee; and (vi) work closely with the program director of the annual meeting when the latter is scheduled in their particular region. 23 section 4 duties of appointive officials 12.4.1 the secretary-archivist shall: (i) be appointed by the president of lassist with the approval of the administrative committee. (ii) attend meetings of the administrative committee and meetings of the general assembly and shall record all facts and minutes of all proceedings in the books kept for that purpose; (iii) be responsible for the maintenance of lassist's records and for its general correspondence; (iv) be an ex-officio member of the nominations and elections committee to maintain lists of nominees for office and to assist in the preparation and distribution of ballots; (v) be an ex-officio member of the standing committee on constitutional review to maintain notices of proposed amendments to the association's constitution and to assist in the preparation and distribution of ballots; (vi) give notice of all meetings of the general assembly and of the administrative committee or president. 12.4.2 the treasurer shall: (i) be appointed by the president of lassist with the approval of the administrative committee. (ii) have the custody of the funds and securities of lassist and shall keep full and accurate accounts of receipts and disbursements in books belonging to lassist and shall deposit all monies and other valuable effects in the name and to the credit of lassist and in such depositories as may be designated by the administrative committee from time to time; (iii) disburse the funds of lassist as may be ordered by the administrative committee; (iv) render to the administrative committee at its various meetings, or whenever the members of the administrative committee may require it, an account of all his/her transactions as treasurer and of the financial position of lassist; (v) prepare a written report for submission to the general assembly at its annual meeting; 24 (vi) provide the standing committee on membership with up-to-date mailing lists of all members in good standing in each of the geographic regions; (vii) maintain current membership lists which shall be published once per year and provided when needed for the official purposes of the association and (viii) perform such other duties as may from time to time be determined by the administrative committee. 12.4.3 the editor of the newsletter shall; (i) be appointed by the president of lassist, on the advice of the standing committee on publications and with the consent of the administrative committee, for a term of three calendar years which may be renewed; (ii) serve on the standing committee on publications; and (iii) be responsible for the regular preparation, publication, and distribution of lassist' s official newsletter. 12.4.4 the program director of the annual meeting shall: (i) be appointed by the president of lassist with the consent of the administrative committee; (ii) set up and organize the next annual meeting following the appointment; (iii) be responsible for keeping the administrative committee regularly informed of all preparations; and (iv) work closely with the regional secretary in the region in which the annual meeting is to be held. section 5 committees 12.5.1 the administrative committee at the time of the annual meeting of the general assembly shall appoint and/or confirm standing committees and shall appoint and/or confirm chairpersons of the said standing committees. 12.5.2 standing committees shall advise the administrative committee on matters of policy within their particular sphere, and shall have a chairperson appointed for a three-year term which may be 25 renewed , two members drawn from the regular membership of lassist appointed for a three-year term which may be renewed, one member of the administrative committee appointed for a three-year term which may be renewed unless representation from the administrative committee is already included in the composition of the standing committee in another capacity, and such officers as are designated ex-officio members. 12.5.3 the standing committees of lassist are the following: (i) constitutional review committee: responsible for receiving proposals for the enacting, amending, and repealing of the by-laws of lassist and for preparing revised articles and by-laws for members' approval, as well as for undertaking an annual review of the constitution and by-laws and proposing amendments as it deems appropriate. (ii) education committee: responsible for the development and advancement of professional programs in education and training and for advising the administration committee on the criteria for the approval and certification of such programs. (iii) membership committee: responsible for recruiting membership in lassist and for recommending alterations in the classes of membership and dues. this committee's membership shall include the regional secretaries. (iv) nominations and elections committee: responsible for receiving nominations for the election of the administrative committee, the regional secretaries, and the officers of lassist, distributing ballots and electoral information according to regulation, tallying the ballots, reporting on the results of the tally, and for recommending alterations ia procedures . (v) publications committee: responsible for advising the administrative committee on general publications program policy and for reviewing manuscripts submitted for publication. this committee's membership shall also include the editor of the newsletter. section 6 action groups 12.6.1 the administrative committee, at the time of the annual meeting of the general assembly, may appoint action groups and for every action group so appointed a coordinator shall be named. 12.6.2 a minimum of three (3) nembers of lassist 'nay nako application to the administrative committee for the establishment of an action group at least one month prior to the annual meeting of the general assembly. 26 12.6.3 action groups shall be expected to undertake specific tasks, to find solutions to specific problems, or to develop and compile relevant materials for specific projects. the mandate or terms of reference of action groups shall be clearly defined, including the resources and time required and the specific nature of the output or product. 12.6.4 action groups shall report to the administrative committee through the vice-president on matters relating to their particular sphere, and shall have a coordinator appointed for a one-year term which may be renewed, two or more members of lassist appointed for a one-year term which may be renewed, one member of the administrative committee appointed for a one-year term which may be renewed, and such officers as are designated ex-officio members. section 7 interest groups 12.7.1 the administrative committee, at the time of the annual meeting of the general assembly, may appoint interest groups and for every interest group so appointed a coordinator shall be named. 12.7.2 a minimum of five (5) members of lassist may make application to the administrative committee for the establishment of an interest group at least one month prior to the annual meeting of the general assembly. 12.7.3 interest groups shall be expected to disseminate information on specific subjects and to serve as a forum of discussion between as well as during annual meetings. 12.7.4 interest groups shall report to the administrative committee through the vice-president on matters relating to their particular sphere, and shall have a coordinator appointed for a one-year term which onay be renewed, four or more members of lassist appointed for a one-year term which may be renewed, and such officers as are designated ex-officio members. section 8 nominations and elections procedures any regular member in good standing is eligible to hold office in lassist. 12.8.1 the administrative committee and the officers (i) every three years, commencing la 1984, the administrative committee, president and vice-president shall be elected from 27 a slate of candidates put forward by the standing committee on nominations and elections. (ii) during the fall of any election year, any member in good standing may submit in writing to the nominations and elections committee, the names of as many as seven (7) persons for the slate of candidates regardless of the geographic region in which the nominees reside. (iii) all nominations must be accompanied by a written statement from the nominees declaring their willingness to stand for election and an outline of the qualifications of the nominees. (iv) the nominations and elections committee will compile a list of nominees which shall be reviewed by the administrative committee and will mail ballots to the membership during the fall/winter of any election year. (v) all members in good standing, regardless of the geographic region in which they reside, shall be eligible to vote for a limited number of nominees from each geographic region. the number of nominees from each region will be specified on the ballot, based on each region's percentage of the total membership of lassist. voting will take place over a period of one month during any election year, but in no instance will it extend beyond mid-december. (vi) the results of the election shall be announced by the end of december in every election year. the results shall be published in the first issue of the newsletter following the election. (vii) newly elected members of the administrative committee and the officers shall take office after the annual meeting of the general assembly following the elections. 12.8.2 the regional secretaries (i) every three years, commencing in 1984, the regional secretaries shall be elected from a slate of candidates put forward by the standing committee on nominations and elections . (ii) during the fall of any election year, any member la good standing in a particular geographic region nay submit in writing to the nominations and elections committee, the name of a person for regional secretary who must reside in tlie same geographic region as the nominator. (iii) a nomination must be accompanied by a written statement from the nominee declaring his/her willingness to stand for election; a statement indicating that the nominee has 28 institutional support to undertake the duties; and an outline of the qualifications of the nominee. (iv) the nominations and elections committee will compile lists of nominees and mail appropriate ballots to the membership of each geographic region during the fall/winter of any election year. (v) all members in good standing in each geographic region shall be eligible to vote for the regional secretary for that particular geographic region. voting will take place over a period of one month during any election year, but in no instance will it extend beyond mid-december. (vi) the results of the election shall be announced by the end of december in every election year. the results shall be published in the first issue of the newsletter following the election. (vii) newly elected regional secretaries shall take office after the annual meeting of the general assembly following the elections . 29 post censal surveys in great britain by robert barnes' office ofpopulation censuses and surveys london introduction censuses of population have been carried out in great britain every ten years since 1801 with the exception of 1941. in addition a mid-term census was conducted in 1966 on a sample of 10% of the population. the next census is being planned for april 1991. censuses in britain are still carried out in a conventional way in that data are collected through enumeration procedures desigjned for the purpose rather than from registers. specially recruited enumerators are used both for the delivery and the collection of forms. post censal surveys are of much more recent origin the first having followed the 1961 census. they have been of three kinds: post-enumeration surveys to evaluate the cover age and quality of the census; follow-up surveys which use the census as a frame from which to draw samples of groups of the population for more detailed enquiry; a longitudinal study which for a sample of about 1% of the population links data from successive censuses together with vital events registered during the inter-censal period. post enumeration surveys (pes) the first post enumeration survey was carried out in 1961. its aim was to assess the coverage and quahty of the 1961 census in england and wales. although separate from the main census, the checks were carried out by census personnel. the sample for the coverage check comprised a systematic random selection of 2,500 enumeration districts (eds) as first stage units. within each a 'plot' was identified containing roughly twenty addresses bounded by features that could l^ identified on the ground. a shortcoming of the system was that it did not provide adequate checks of coverage error arising at the boundaries of enumeration districts. the check produced an estimate of net under enumeration of 0.2% but the general report on the 1961 census^ said "the design of the enquiry was such that the quality of the result may be suspect but there is no information on this". the quality check sample was selected from the same 'plots' as the coverage check. in all about 17,500 addresses were revisited for voluntary interviews within about three weeks of the census date. census enumerators were used for this purpose but were not skilled interviewers. again the general report comments on the shortcomings of the study. 'this fact and the very limited instruction which it was practicable to give them were contributory factors in the failure of the postenumeration survey to give satisfactory answers to some of the questions that were included". in 1966 a quality check on the sample census was carried out by the social survey division of the central office of information. this was a government organisation separate from that responsible for the census itself. a sample of just over 5,200 households was selected from 300 eds in england and wales. there were several key features of this quality check survey which were to set the pattern for similar studies carried out in conjunction with future censuses. • whereas cooperation in the census was, and still is, compulsory, the post-enumeration survey was voluntary. all sample surveys of households and individuals in britain are voluntary. response to the 1966 quality check was 95%. • whereas the census employed a large force of temporarily recruited enumerators, the survey used highly trained and experienced interviewers. • the survey used detailed questionnaires to ascertain the 'true' answer to topics which in the census were covered by just one or two questions. • survey interviewers had copies of the informant's answers on the census form so that they could probe any discrepancies. this not only improves the quality of the check but also provided some reasons for the differences. • the survey sought interviews from each adult in a household and so relied far less on proxy responses than did the census. • in order to avoid discrepancies arising through genuine change between the time of the census and the time of the survey, it was important that the survey took place soon after the census. in fact the survey fieldwork was carried out between two and three months after the assist quarteriy date of the census. the results of the quality check were published' and indicated that answers to a number of questions on the census form were in error to a substantial degree. the worst item was the question on the number of rooms in a household this was misclassified in approximately one case in five. coverage and quality checks were again carried out on the 1971 census. as before the coverage check was conducted in england and wales by census officers. although the check showed an undercount of only 0.23%, the check was not considered entirely successful from a technical point of view. the general report on the 1971 census* said "it was carried out by census officers some of whom would still have been busy with other work on the census. some census officers might have viewed the check as a fault finding mission and in consequence would not have been fully motivated to ensure its success". in fact it was demonstrated that the accuracy of the coverage check was poorer than that of the census itself since there were more addresse: found in the census but missed by the coverage check than vice versa. neither did the coverage check provide a reliable measure of the extent of over counting that was provided by other means. taking account of the deficiencies of the check it was finally concluded that a more realistic figure for the net undercount was about twice that shown by the check. the 1971 quality check was conducted by means of a post enumeration survey, covering the whole of great britain, by interviews at just under 5,000 addresses. the response rate was 85%. the check was again carried out by the social survey which had in 1970 become a division of the same office which was responsible for the census the office of population census and surveys. as with the coverage check there were unsatisfactory aspects about the quality check. it was originally planned that fieldwoik should be completed within about two months of the census. in the event the work did not start till then and, because of difficulties in transferring census information from one division to the other, it proceeded slowly. eventually fieldwork was completed about five months after census night because of weaknesses in the design of the check and the length of time to complete fieldwork, a number of important aims of the check were not achieved. for example there was no assessment possible of the accuracy with which questions on household tenure and amenities were answered. in 1981 coverage and quality checks were carried out but this time both were undertaken by the social survey division of opcs. this provided a much more integrated approach to census evaluation. for this coverage check roughly 1,000 eds (about 1%) were selected in england and wales* and all addresses in them were thoroughly relisted by trained interviewers to see if any had been missed by the census enumerators. the eds were selected in blocks of four, adjacent to each other so that enumeration ofed boundaries could be checked. different samples of addresses were selected within the eds to see whether persons had been left off census forms for addresses otherwise correctly enumerated, to see whether anyone had been missed in addresses enumerated as vacant or non-residential and to check especially the enumeration of multi-household addresses. this was a much more rigorous approach than had been used in 1971. moreover in 1971 no attempt had been made in the field to reconcile discrepancies. in 1981 the interviewers were given copies of census listings and so could check by means of personal interview numbers of persons missed in non-enumerated addresses. also in 1981, unlike previous checks, discrepancies in the number of persons present in enumerated households were taken up with the informant, enabling the interviewer to form a more definite idea as to whether the census form or the post enumeration information was correct. there was one other way in which the 1981 check was superiw to that of its predecessors. all eds were graded on the basis of 1971 census data for expected difficulty of enumeration. the design used for the 1981 post enumeration survey over sampled, by a factor or two, these 'difficult' eds, since it was hypothesized that such eds would be likely to produce more errors. however in spite of the thoroughness with which the 1981 coverage check was carried out and the improvements made over previous such studies, there are inherent difficulties in checking the coverage of a census using re-enumeration methods. because of the cost of the approach, the samples have to be fairly small and the sampling errors are therefore relatively large especially at the sub-national level. also in spite of the thoroughness of the methods it is likely that some of the persons and addresses missed in the census will have been missed in the survey too for similar reasons. therefore in 1981, as in previous census evaluations, in addition to the post enumeration survey, checks were made also against independent administrative sources for particular groups of the population. for example checks were made for children aged 0-9 with data on registered births and deaths and making allowances for migration, for infants aged 0-1 with birth records, for children of school age with numbers on school rolls, and the census count of people of retirement age was compared with the records of the numbers receiving state retirement pensions. these checks confumed that the level of under enumeration in the 1981 census was small even though it may have been slightly higher than the half per cent found by the post enmeration overage check. in addition to the post enumeration check for coverage, a sample of about 5,000 households throughout britain within the eds selected for the coverage check was revisited for detailed interviews to check the accuracy of answers given on census forms. response to this post enumeration quality check survey was over 90%. all fieldwork was completed within three months of the date of the census. the 1981 census form was the shortest for fifty years. compared with many items which might have been included, and which are in other countries, the items in the british census might have been regarded as spring 1990 relatively straightforward and commonplace. nevertheless grass errw rate^ of 8% or more were found on five of the sixteen variables examined. for two of these the error was over 25% (and again the worst case was the question on number of rooms) and for some sub-groups the gross errors were even larger. the relative impwtance of these errors depends on the uses made of the information for example the errw rate reduced substantially when classifications were collapsed into fewer and simpler categories and, in any case, the net errors, in distributions, were smaller. improvements were made for the 1981 post enumeration survey in regard to publication of the results, compared with previous censuses. the full report* took a few years to produce but key results were published more quickly in the form of summary monitors both for the coverage check and the quahty check'. the 1981 census quality check also provided an opportunity to carry out a study to measure the coverage and quality of the electoral register. * this had first been done, although in a more limited way, in conjunction with the 1966 quality check. the 1981 study covered the whole of great britain and measured the extent of both persons who were eligible to be on the electoral register but who were not, and those who were on the register but for whom there was no census form or for whom census details suggested they were not eligible. the results showed that just under 7 per cent of those who should have been on the register were not and that the same proportion of names on the register should not have been. the under representation was especially high in inner london (14 per cent), and among those aged 17 (24 per cent), and those who had recently moved (27 per cent). these figures represented a deterioration since the register had previously been checked in 1966. it is planned to carry out a similar check in 1991. follow-up surveys the first survey to use the census as a sample frame for more detailed enquiry was a follow-up to the 1966 sample census. this was a study on diet and health concentrating especially on sugar intake because of the supposed relationship between that and myocardial infarction. this was a postal enquiry of some 20,(xx) persons conducted in 1967/68. 1966 sample census returns were used as a sample frame because the sample required was of men between ages 45 and 65 in the london area. moreover the subsequent fate of the sample members could be traced through death registers also maintained by the same office. considering the length of time between the census and the survey, the percentage of returns was high at 85%. of these 89% had completed the forms (75% response overall). of the remainder the majority were returned as gone away, deceased etc. important aspects of this survey were that data were collected by the same organisation that carried out the census and that procedures to ensure the anonymity of sample members were strictly adhered to before the data records were released to outside researchers for analysis. the tracing of sample members through death registers has continued ever since and, although no results have yet been pubushed on the relationship between sugar and heart disease the study has yielded other results relating diet and disease in particular between tea and coffee consumption and cancer.' following the 1971 census there were three such enquiries. the first was the income follow-up survey.'" there had been pressure to include an income question in the census itself but tests carried out in 1968 and 1969 indicated the severe problems of dealing with such a complex topic with simple questions suitable for a census form. it was also found that the inclusion of income as a possible topic in a compulsory census aroused hostility among some sections of the british public. therefore it was decided that the census division of opcs should carry out a voluntary survey, by post, on 1% of the population throughout great britain in such a way that the answers could be unked with the census forms of the sampled individuals. although it was intended to carry out the study as quickly as possible after the census, the adverse publicity which the 1971 census attracted caused the follow-up study to be postponed and it was eventually conducted over a year later. because of tliis and because of the subject of the enquiry, the response was only 40% and a substantial volume of imputation, using hot deck methods, was carried out the study was not repeated after the 1981 census. a second follow-up survey to the 1971 census was on the subject of qualified manpower. this also was a voluntary, postal enquiry carried out by census division. the object was to follow up a sample of those who on the census form had reported that they had academic, professional or vocational qualifications, to seek more information about the qualifications, their jobs and their employment income. however because of competing demands of other work, there were delays in completing this survey. in the end the need for the results was overtaken by other events and the report was not published. the third follow-up to the 1971 census was the so-called nursing survey." this was particularly important not only because of the topic of enquiry but because the criticism which the survey attracted had impwlant consequences for follow-up enquiries of future censuses. the purpose of the study was to obtain information from persons with nursing or midwifery quahfications, who were of working age but who were not practising nursing or midwifery. the census provided an ideal frame from which to select a sample of such persons who would otherwise would have been especially difficult to locate. a sample just over 7(x) throughout great britain was selected for voluntary face-to-face interview and of these 89% agreed to cooperate. the fieldwork was carried out within four to five months of the census date. the study was carried out by the social survey division of the then recently formed opcs. census division, of the same office, selected the sample and passed the names and addresses to the social survey. because both 10 lassist quarterly divisions were part of the same office there was no breach of the confidentiality undertakings given at the time of the census. but this fact was not always appreciated by the public, nor by some sections of the press or even the research communities and there was considerable criticism of the practice of using census forms for this kind of purpose. even an investigatory team from the royal statistical society commented in 1973 that "this use was, in our opinion, only doubtfully covered by the wording printed on the census form and by other public pronouncements on the confidentiality of information given in the census operation". nor did the reference to the nursing survey was still being made at the time of, and after, the 1981 census even though there were no follow-up enquiries of this kind attached to that census. as a result of the controversy, procedures changed after 1971. although there were no such studies following the 1981 census, had there been they would have had to have been announced to parliament before the census was carried out. similarly for 1991, the white paper published in july 1988'^ announcing the government's intentions to take a census 199^ , contained the following paragraph. "the census may be used as a source from which to select samples for further more detailed surveys, for example of people with particular educational qualifications. response to any such survey would be voluntary.... information would be treated in the same strict confidence as information given on the main census forms. it is too early to know whether there will be a need for any such census-linked surveys and the topics that might be covered, but parliament will be informed before the census is taken about the subject matter of any census-linked survey which it was proposed to conduct following the census and all those completing the census forms would also be informed of the possibility of being asked to participate in such a voluntary survey after the census." the white paper also emphasized that any such censuslinked surveys would have to be handled by opcs (and by the general register office in scotland). this means that the work cannot be contracted to commercial or academic research agencies. the longitudinal study " this is a rather special form of post-census survey. it is a data linkage exercise and involves no specific enquiry of the sample members. the study comprises a sample of persons in england and wales with birthdays falhng on each of four specific days in the year and therefore provides roughly a 1% sample of the population. the study started with the 1971 census records for the sample members. between 1971 and 1981 mortality and migration data were added to the records from registration data also held by opcs. also added to the records were live and still births to sample members, deaths of children under one year of age to sample members, death of a spouse and cancer registration. in addition children bom on any of the four sample days during any year between the censuses were added to the sample, thus maintaining it at approximately 1%. in 1981 the new intercensal data on migration and vital events has continued to be added since then. proposals are now being drawn up for further links to be made to the 1991 census. initially the longitudinal survey included about 530,000 people selected ftxjm the 1971 census. by the time of the 1981 census some 121,000 more had been added to the sample because of new births, immigration, ot because they were found for the first time in 1981. and 1 15,000 were removed from the sample because of death, emigration or because they could not be traced in the 1981 census. therefore the sample after the 1981 census records are available, numbers in particular subgroups are as follows: characteristics in 1971 number for whom 1981 details linked children under 16 teenagers 13-19 divorced 120,000 40.000 5.000 lone parents with dependent 5,000 dependent children men aged 50-64 40,000 persons aged 75 or over 10.000 unemployed 10.000 migrants 130,000 there are many imfwrtant uses for this longitudinal data set for such a large sample of the jjopulation. they are too numerous to list here but they include the relationship between socio-economic circumstances and mortality in relation to housing and household circumstances, sociodemographic differentials in cancer incidents and survival, the relationships between socio-demographic factors and migration (i.e. the propensity to change address about the time of events such as marriage, birth of children, birth or bereavement), fertility patterns of young women according to family background characteristics, and social and occupational mobility. apart from this kind of general analysis, the data set is also large enough to focus on particular sub-groups or geographical areas. conclusions post censal surveys have been carried out in great britain for nearly 30 years. they have been used to evaluate the census,' to take advantage of the census as a samphng frame and to provide longitudinal data for particular topics. they have not all been successful and problems both technical and ethical have arisen. however post censal surveys will almost certainly be a feature spnng 1990 of future censuses in britain and lessons learned from previous experiences should help to ensure that past difficulties are avoided. the main lessons are that field woric for follow-up studies should be completed as soon as possible after the census, that the field workers should be trained interviewers skilled in administering complex questionnaires, that the public and parliament should be informed in advance of the intention to use census data for this kind of purpose, and that census evaluation studies should be carried out by staff who were not directly involved with the census operation itself. 'presented at the ifdo/iassist 89 conference held in jerusalem, israel, may 15-18, 1989. general report on the 1961 census. kjray & gee. a quality check on the 1966 ten per cent sample census of england and wales. hmso, 1972. ktensus 1971. general report part 3, statistical assessment opcs, 1983. 'separate coverage checks were carried out in scotland. *britton & birch. 1981 census post enumeration survey. hmso, 1985. evaluation of the 1981 census. opcs monitor cen 82/3, 1982. evaluation of the 1981 census: post enumeration survey. opcs monitor cen 83/4, 1983. evaluation of the 1981 census: post enumeration survey (quality check). opcs monitor cen 84/3, 1984. *todd & butcher. electoral registration in 1981. opcs, 1982 1x0 kinlev et al. coffee and pancreas cancer; controversy is part explained. the lancet, feb 1984. l.j. kinlev et al. tea consumption and cancer. british journal of cancer, no. 58. macmillan, 1988. '"1971 census income follow-up survey. opcs studies on medical and population subjects, no. 38. hmso, 1978. "sadler & whitworth, reserves of nurses. hmso 1975. "1991 census of population. cm 430. hmso, 1988. '^e longitudinal study 1971-1981. cen81ls. hmso, 1988. brown & fox, opcs longitudinal study: ten years on. population trends, no. 37. hmso, autumn 1984. '^whitehead. the gro use of social surveys. population trends, no. 48. hmso, summer 1987. '2 lassist quarterly the promise of multimedia: data for every computer by janet vavra ' inter-university consortium for political and social research when the microcomputer arrived on many of our desks in the 1980's, few of us realized just how dramatically this relatively small (often mysterious) piece of equipment would change our lives. not only did the microcomputer alter the way we perform our daily tasks, it literally changed the way we view and interact with the world. sending messages to colleagues half a world away was something we did either through western union or the postal system; data were shipped on magnetic tape through the mail. one looked for a response to an overnight letter in several days, assuming the marine life did not nibble through the undersea cable in the meantime. today with the use of email facilities and public data networks, individuals send messages across the country and around the world with the same ease as they once picked up the phone and placed a call across town. today many users have desktop computers (workstations) that have more computing power than many mainframes had 10-15 years ago, and the machines do not require an entire floor to accommodate them and their peripherals. the user of a microcomputer, or workstation, can work with an enviable array of hardware and software performing data transfers, database searches, accounting tasks, data analyses, and receiving and sending messages without ever leaving the office. a researcher could conceivably conduct a project from beginning to end from the keyboard (or with a mouse) of a micro. by searching the library's on-line catalog, relevant publications can be identified thereby eliminating long hours spent in the library by the researcher or a graduate student. searches are not necessarily limited to local libraries, but rather identify publications available from a variety of sources. on-line databases can be used to locate machinereadable data containing the needed variables. if access to such databases is limited to selected users, a request asking for a given search could be sent via email to the individual authorized to search the databases. data could be ordered through electronic mail or electronic ordering systems. in some instances data could be downloaded through public data networks directly to the user's hard disk or to a local data library facility and then through a local area network to the user's machine. analysis can be performed with micro-based software and the results shared with project colleagues who may be at other locations, again using email or public data network capabilities. small files of information can be exchanged through bitnet, while larger files can be sent through networks such as internet. finally research reports can be prepared with wordprocessors, many of which have graphics capabilities. frequently the ease v/ith which changes can be made in any document can create a problem of another sort: one cannot resist making another "final" change or one more addition to the document one could go on citing similar scenarios in almost every field and activity. the point is that many of the capabilities that were only available to users at central computing facilities, or not at all, are now available on the powerful machines many people have in their offices. needless to say, this has impacted on organizations, such as the interuniversity consortium for political and social research (icpsr), that provide data and related services to many of these users. this paper will try to identify some of the changes brought about by the advent of the microcomputer. it will primarily look at those changes that impact both on researchers and on those organizations that provide machine-readable data to researchers and instructors. it will seek to identify some of the ways in which these organizations will and may provide their services in the future. the focus will be primarily on the challenges faced by the icpsr both in continuing to serve users with more traditional computing environments and those at institutions where the activities of the central computing facility have been largely replaced by personal computers. while the past two to four years have seen a growing interest in what can be termed "alternate media", data continues to be transmitted and exchanged among facilities on magnetic tape. however, the day is rapidly approaching where magnetic tape will not be the medium of choice but rather the medium of last resort magnetic tapes continue to be a medium that generally requires a mainframe. this is at a time when many central computing faciuties are starling to cut back on services they have traditionally 66 lassist quarterly provided, and many are also cutting back on staff that support these services. the reasoning frequently is that mainframes will serve as giant gatekeepers and servers while the microcomputers will take over the day-to-day computing needs of users. with the central computing facilities cutting back on individual user services and users finding themselves with impressive computing power on their desks, it is natural that demand will increase for data and other supportive services that are compatible with micros. however, given the variety of configurations one can have at the micro level and the differences in individual preferences, it is no easy task to come up with products that will meet everyone's needs. additionally archives face the very painful reality of the very high cost of converting all of their holdings from a mainframe-compatible format to a primarily microcomputer-compatible one. almost without exception, microcomputers have floppy diskette capabilities. however, not only do the diskettes generally come in two different sizes, they also can each be written in different densities. (sounds a bit like the old days of sevenand nine-track magnetic tapes.) one can try to identify the format used by the most users and write diskettes routinely with those specifications. while that may be the only thing that makes sense for an organization that supplies data to thousands of users each year, there will remain those users who absolutely cannot use the standard product and must have data at different specifications. it seems to make sense to have a standard product that most users can work with and to deal with those users that cannot handle the standard product on a case-by-case basis. unfortunately, while floppy diskettes are an excellent medium for transmitting small collections of data, problems begin to arise when large data collections that contain more than a couple of megabytes of data are involved. one solution is to supply the data in compressed format. most of the compression software around reduces the size of a file from 70%%-80%% of its original size. since the capacity of most diskettes ranges from an average of 350 kilobytes for a 5 1/4" low density diskette to 1.4 megabytes for a 3 1/2" high density diskette, it is easy to see that large files, even when compressed, very quickly become impractical for this medium. supplying one data file on numerous diskettes can create problems for software that may eventually have to manipulate the data. large compressed files have to be decompressed and space has to be available locally to accommodate the decompressed data. while the user has to be concerned with the amount of space available on their their microcomputer in order to be able to work with the data arriving on diskette, the data producer is further concerned with the effort that must go into preparing the data for diskette. it may be necessary to reformat the data to make it compatible with the micro environment. for example, pc-based software frequently cannot accommodate large record sizes with ease. additional problems may arise with large numbers of cases and/ or large numbers of variables. depending on the work that needs to be done, reformating could be an expensive profxjsition. accordingly, it may be necessary to identify only certain collections that can be provided on diskette and further to routinely provide these data in only a selected number of diskette formats. optical media go a long way toward solving storage problems when it comes to large collections of data. one of the more popular optical media is the cd-rom. on a disk no larger than a 5 1/4" floppy, a cd-rom can easily hold over 600 megabytes of raw, ascii data. while access on a cd-rom is slower than with a floppy, many producers bundle the data on their cd-rom with software which helps to reduce the retrieval time. in other instances, users are not concerned about the length of the retrieval time, since time spent on their microcomputer does not result in direct costs the way time spent on a mainframe does. instead they set up a batch job to run on their micro during the lunch hour or overnight. despite the high capacity of a cd-rom and some of its other attractive features, the cd-rom is basically not an inexpensive medium. users usually must purchase a cdrom drive. generally a cd-rom's performance depends both upon the drive and driver used and on the power of the machine on which the work is being performed. it may even mean purchasing a different micro, if the current micro is not suitable for cd-rom applications. from the producer's point of view, a cd-rom product can be a very exj)ensive undertaking. if the data are to be summer 1991 67 bundled with software, the producer must either identify existing software and then seek licensing agreements to use the software, or must write in-house software. licensing agreements can be costly; the preparation of in-house software may, however, involve an even greater financial commitment. data will additionally need to be prepared for input into the software or may need to be restructured for the microcomputing environment. another alternative is to not supply any software with the product and leave it up to the user to identify software to be used with the data. this lauer approach is more akin to using the cd-rom as a data transmittal and storage medium than as a complete data transmittal, storage, and retrieval system. while data can also be compressed on cd-rom, the large volume of information that can be stored on cd-rom usually necessitates special retrieval software for full or partial extracts. it is easy to visualize the problems that could arise if a user had to decompress a 550 megabyte file stored on cd-rom onto a hard disk, or other local storage media, before being able to manipulate the data. after the decision has been made regarding the nature of the cd-rom product, premastering and mastering must be done before copies can be made. normally premastering and mastering is done by service bureaus although producers can opt to purchase the necessary equipment and software for in-house capabiuues. the charges for such capabilities preclude most organizations from deciding to master their own cd-rom products. generally the costs for producing a cd-rom are such that only selected data collections can be considered for the medium. the transmittal of data over pubhc data networks has a great deal of appeal to both the users and the data producer. by simply identifying the data needed and giving the appropriate set of commands, the user can theoretically u^nsfer any data collection needed in a mauer of minutes. this can all be done without any direct intervention by the producer; the producer need only be notified in some electronic manner that the transaction occurred. as network speeds have been increasing from tl (maximum 1.544 megabits or 200 kilobytes per second) to t3 (maximum 45 megabits or 590 kilobytes per second), this mode of data exchange has created a great deal of interest. but as with all of the alternate media discussed so far, there is good and bad with this option. it is very attractive from a user standpoint to be able to simply give a few commands on your desktop and have megabytes of data arrive over the lines in a matter of minutes. there is no need to wait days for an order to be processed and then to always be concerned that it will not arrive in time either for a paper deadline or class assignment it would eliminate the waiting that lakes place when the user discovers that another data collection would have been a better choice than the one originally requested. it certainly is true that if pubhc data networks worked in practice as they sound in theory, our data exchange problems would be over. however, some of the same problems that impact other media are also at work here. the speed with which data arrive over the lines is the result of a number of factors, including the different routes they must pass through to get from the source to the user's machine. the speed with which the data make that journey will be only as fast as the slowest link along the network path. therefore, users never actually experience the maximum data ffansmitlal speeds quoted for any given network. while every effort possible is made by the public data networks to assure complete transfer of data sent, transmittal problems can arise, resulting in incomplete transfers. the machine receiving the data must have space to accommodate the information coming down the lines. finally, it is likely that not all data formats will readily lend themselves to public data networks. for example, since most of the data going over the lines are ascii, ebcdic binary data such as that found in osiris dictionary files and other similar formats will certainly not be usable on the receiving end. after looking at each of the several media available and the advantages and disadvantages of each, one may very well ask which is the best approach. we at the icpsr have been spending a fair amount of time exploring each of the different formats. unfortunately, we have not found any simple solution that will provide everything for everyone. instead, we have concluded that we will be fortunate if we can provide something for everyone. for the foreseeable future, icpsr expects that much of the data supplied to users will continue to be provided on magnetic tape. all surveys of our users indicate that magnetic tape remains the overwhelming preference as a transmittal medium by the majority of our users. (however, there is a great deal of effort going into determining the next generation of reliable storage and transmittal media.) additionally a significant number of our collections will not be suitable for transmittal by any other medium for the foreseeable future. this is largely due to the size of many of the collections which span several reels of tape. it is expected that the format of some collections will initially preclude their transmittal on other media. hierarchical data files will be best supplied on magnetic tape at least for the time being. nevertheless, the icpsr has been taking steps to move toward other alternate media. in february, 1991 nearly 100 copies of a two-volume cdrom containing the panel study of income dynamics data for waves 1-20 were distributed to official representatives at member institutions who had expressed a wish to participate in a field test of the product. the data were lassist quarterly supplied on cd-rom in raw, ascii format spss/pc and sas/pc statements were provided on each cd-rom. users could use the statements to prepare extracts from the main file or they were free to utilize their own software to perform the extracts. responses on the questionnaire that was provided with the field-test copies indicated that users overwhelming approved of this approach for the cd-rom. icpsr expects to produce a limited number of cd-roms in the future. it is expected that collections selected for this medium will be those that users have indicated they would like to see in this format, and those collections that have a high distribution volume. some of these additional cdrom products should be available within a year. the computer support group of the icpsr is in the process of conducting tests with a group of sites to evaluate the transmittal of icpsr data through internet. when these tests are concluded and relevant programming jiat supports this activity completed, we expect that access to icpsr data will be expanded to include public data networks. it is expected that while eventually all icpsr data will be accessible through the public data network, initially selected collections will be available in this manner. in order to make icpsr data available through internet, the thousands of files in our holdings will need to be moved from exclusively magnetic tape storage to optical disk storage. since this will be a relatively large task, not all data will be stored on optical media immediately. finally, icpsr is in the process of identifying collections that will additionally be available to users on floppy diskette. data collections selected for floppy diskette production will be those for which there is demand for the data to be provided on this medium. the data collections that will be provided on diskette will be those that do not require large numbers of diskettes to accommodate a given data file. additionally they will be collections that are available in raw, ascii format it is expected that the data will be supplied with self-extracting compression software with the appropriate readme files that provide users with relevant information about the contents of the diskette. as icpsr continues with the installation of both software and hardware that will allow us to move from magnetic to optical storage media, we expect to continue to explore the feasibility of adding new services and upgrading older ones. while the icpsr will monitor and respond to the technical changes many of our users are experiencing, we will also continue to remain responsive to those users who are not experiencing rapid technological changes. for the foreseeable future, icpsr will seek a balance that allows us to serve users spanning the full spectrum of technical capabilities. ' presented at the iassist 91 conference held in edmonton, alberta, canada. may 14 17, 1991. summer 1991 69 vol262 4 iassist quarterly summer 2002 iassist quarterly summer 2002 5 editor’s notes the iassist quarterly is not a newsletter. the frequency and production time means that the content cannot stay in the category of news. furthermore the content does not resemble a traditional newsletter. there are no letters to the editor nor replies or rejoinders from the membership. letters and replies are most welcome. but the “welcome” does not seem to have any effect, as the content of the quarterly is almost exclusively papers from presentations at the iassist conferences. therefore a name like “iassist quarterly studies” would better cover the function. is the name important? “whatʼs in a name”? “iq” is a beautiful name and most iassist members know the type of content to be found in the iassist quarterly. as you all probably are fully aware of the iq is also available in an electronic form on the internet (www.iassistdata.org). the iq is a media that is part of iassist as a virtual community. i find that the iq is important for iassist also as a traditional community. with the paper version of iq delivered to your door you receive the physical evidence of our professional affiliation. it reminds you that iassist does exist in reality, between conferences! this issue 26-2 of the iassist quarterly contains the following three articles: at the amsterdam conference in 2001 with the theme “collaborative working in the social science cyber space” julia paris from technikon witwatersrand in johannesburg (south africa) made a presentation in the session on iassist. to the conference julia paris brought original african artifacts (not software and data) and this added further color to the presentation on iassist and the technology collaboration in africa. (the iq can unfortunately only present the text). the article gives a more complex view of africa than the “poverty-stricken, politically unstable, war-torn continent” that is often presented in the news-media. there are technical initiatives and therefore involvement and further collaboration with iassist is recommended. at the session “tuning up your web site: increasing its usability” at the storrs 2002 conference stuart smith from the university of manchester gave at presentation on “creating accessible style and content for mimas social sciences web pages”. the presentation was based upon the paper that appears in the iq: “accessibility, social sciences and the development of content management systems”. the paper addresses how the use of the web has introduced a “digital divide” and how the data services will overcome these technical and cultural issues. the paper demonstrates how the presentation of web-pages on the internet has received a great deal of attention and shows how it is now possible for the content staff to concentrate on the content of the web-pages. at the same 2002 conference a session was dedicated to “delivering the u.k. census: web-based access”. in this session lucy bell from the uk data archive presented the paper “let us bring you to your census: recent developments in uk census data provision”. the paper describes the process of setting up a one-stop census registration service (crs). this is an online system for providing quick and simple user registration for access to the resources from the 1971, 1981, 1991 and 2001 uk decennial censuses. as the number of products associated with each census has steadily grown it has become important to make an easier access to the census materials. this became the project for “one-stop census registration service”. the article presents the questionnaire sent to users and summarizes the results revealing issues of dissatisfaction with the earlier process as well as preferences for improved functionality with the new system. be sure to enjoy the iassist web-site. you can locate many of the powerpoint presentations at www.iassistdata.org. (click “conferences” and then “multimedia presentations” under the “features” heading). if you have made a conference presentation that is missing on the web-site you should contact the collector lisa neidert (lisan@umich.edu). papers for the iassist quarterly are most welcome. please contact the editor (kbr@sam.sdu.dk) about submissions. karsten boye rasmussen, november 2002 mailto:lisan@umich.edu mailto:kbr@sam.sdu.dk iassist quarterly 3 archives law and machine-readable data files: a look at the united states files. how these corporate records — regardless of media — are created, maintained, preserved and accessed is specified in the organization's official policy statements. such policies will generally specify who in the organization has responsibility for each of these activities relating to the organization's official records. when the organization is a government entity, these policies are embodied in the laws or statutes of the govenmienl such laws are of obvious importance to government employees concerned with records since the statutes specify the basis for the activities relating to records by each agency and its personnel. individuals wanting information from a government agency should also be aware of these laws because they have direct impact on the accessibility of the informatioa this paper will review the provisions of the laws relating to archives in the united states, relate them to machine-readable data files in the federal government, and then will use the records of the bureau of the census to illustrate the legislatively mandated approaches. by thomas elton brown' national archives and records administration washington, d.c., united states of america within the united states, the federal government primarily controls the creation and disposition of record material through the federal records act of 1950 as amended. this statute defines records as: introduction in the strict sense of the word, archivists have responsibility for the official records of an organization, in contrast with manuscript curators who collect private documents accumulated by an individual person, or librarians who manage publications. the organizational records which the archivist is to manage may include a variety of materials, including machine-readable data 'presented at lassist/ifdo international conference may 1985, amsterdam. all books, papers, maps, photographs, machine readable materials, or other documentary materials, regardless of physical form or characteristics, made or received by an agency of the united states government under federal law or in connection with the transaction of public business and preserved or appropriate for preservation by that agency or its legitimate successor as evidence of the organization, functions, policies, decisions, procedures, operations, or other activities of the goverrmient or because of the informational value of data in them. [44 u.s.c. 3301] spring 1986 iassist quarterly one should note that the definition of the records in this statute specifically includes "machine-readable materials." the federal records act also includes a provision that states: the head of each federal agency shall make and preserve records containing adequate and proper documentation of the organization, functions, policies, decisions, procedures, and essential transactions of the agency and designed to ftunish information necessary to protect the legal and financial rights of the government and of persons directly affected by the agency's activities. [44 u.s.c. 3101] it is this provision that grants to the head of each agency the authority to determine what records the agency will create. thus it is the federal agency that determines what machine-readable information will be collected and processed. once an agency has created machine-readable records, they cannot be destroyed without the approval of the archivist of the united states. if the archivist determines that the machine-readable record has archival value and should not be destroyed, then disposition of the data involves their transfer to the national archives for continued preservation. when will the data be transfened? the timing of the transfer may best be described as a date negotiated between the agency and the archives. for all records regardless of media, the archivist: may direct and effect the transfer to the national archives of the united states of records of a federal agency that have been in existence for more than thirty years and determined by the archivist of the united states to have sufticient historical or other value to warrant their continued preservation by the united states government, imless the head of the agency which has custody of them certifies in writing to the archivist that they must be retained in his custody for use in the conduct of the regular current business of the agency. [44 u.s.c. 2107] because of the fragile nature of machine-readable records, a special provision for information on this medium has been added to the regulations which all federal agencies must follow: when the national archives and records service [administration] has determined that a file is worthy of preservation, the agency should transfer the file to the national archives as soon as it becomes inactive or whenever the agency can not provide proper care and handling of the tapes to guarantee the preservation of the information they contain. [41 c.f.r. 101-11.411-6] in addition, the national archives has the authority to establish the procedures which constitute proper care and handling. [41 c.f.r. 101-36.12] access by the public is governed by the freedom of information act, the federal records act, and individual statutes governing specific programs or data collection activities. the freedom of information act generally provides that a:ny person has the right of access, enforceable in court, to federal agency records except to the extent that such records (or parts of those records) are protected from disclosure by any one of nine exemptions. this statutory guarantee to access federal information applies equally to all record material — whether in the custody of the creating agency or in the national archives. thus the act of transferring the information to the national archives neither expands or limits the right to access the information. the limitations on access stem not spring 1986 iassist quarterly 5 from the physical location of the material but from the nine exemptions. one of these nine exemptions is "all matters specifically exempted from disclosure by statute." [5 u.s.c. 552] according to the federal records act, all statutory limitations and restrictions on the examination and use of the records while in agency custody are transferred with the records when they go to the national archives. again the physical custody of the records does not aftect any restricitons on access. these statutory restrictions: shall remain in force until the records have been in existence for thirty years unless the archivist by order, having consulted with the head of the transferring federal agency or his successor in function, determines, with respect to specific bodies of records, that for reasons consistent with standards established in relevant statutory law, such restrictions shall remain in force for a longer period.[44 u.s.c. 2108] thus the statutory restrictions acknowledged in the freedom of information act expire after thirty years unless extended by the archivist of the united states in consultation with the agency. responsibility to "provide guidance and assistance to federal agencies with respect to ensuring adequate and proper documentation of the policies and transactions of the federal government". [44 u.s.c. 2904] with regard to the census bureau, the national archives would intepret this provision as authorizing the national archives to provide advice on how to document how the census or siua-ey collected the information. it would not include advice on what information the census or survey should collecl however, if the census bureau determines that it will collect information which the national archives determines to have archival value, then the archives will advise the census bureau on how to process and maintain the information to ensure that the information is retained in a format that can be transferred to the national archives. tide 13 of the united states code is the legislation which authorizes the census bureau to collect and process its data and imposes three restrictions on the information gathered by the census bureau. the census bureau may not: use the information furnished under the provisions of this title for any purpose other than the statistical purposes for which it is supplied; or the machine-readable records of the bureau of the census can serve as an illustration of the management of federal records even though a specific provision of the federal records act governs access to some records in the census bureau. first the census bureau determines what material it will collect as part of its census and survey activities. in making its determination of what information to collect and how, the census bureau actively seeks advice from a variety of sources including other federal agencies. the national archives does not ofter advice to the census bureau on what questions should be asked or on how the censuses and sur\'eys should be conducted. the national archives does have the statutory make any publication whereby the data furnished by any particular establishment or individual -under this title can be identified; permit anyone other than sworn officers and employees of the department [of commerce] or bureau or agency thereof to examine the individual reports. [13 u.s.c. 9] to comply with these limitations and yet to provide users with needed data, the census bureau creates public use files, either extracts or microaggregations. in this way, the census bureau can release information which will not spring 1986 6 iassist quarterly identify a respondent — whether an individual person or economic establishment the national archives has the responsibility for determining which information has archival value and which information may be destroyed when no longer needed by the agency. this determination is what the archivist refers to as "appraisal." such appraisal of machine-readable information is done separately for microdata files with individually identifiable records, for public use extracts, and for microaggregations. the federal records act would normally limit title 13's restriction on the release of individual information to thirty years unless extended by the archivist however a provision in the federal records act stipulates that: [w]ith regard to the census and survey records of the bureau of the census containing data identifying individuals enumerated in population censuses, any release pursuant to this section of such identifying information contained in such records shall be made by the archivist pursuant to the specifications and agreements set forth in the exchange of correspondence on or about the date of october 10, 1952, between the director of the bureau of the census and the archivist of the united states... [44 u.s.c. 2108] the key to this agreement is that: [a] fter the lapse of seventy-two years from the enumeration date of a decennial census, the national archives and records service [administration] may disclose information contained in these records for use in legitimate historical, genealogical or other worthwhile research. [h.r. report 95-1522, august 21. 1978] the statute that makes reference to this exchange of correspondence also grants the two agencies the authority to amend the agreement, provided that they publicize the change in the federal register . the statutory clause which makes reference to the exchange of letters specifies "census and survey records of the bureau of the census containing data identifying individuals enumerated in population censuses". thus, this statute and the seventy-five year provision apply only to demographic information dealing with individual persons. they do not apply to the economic censuses and surveys which gather information from business establishments. what laws do apply to identifiable information on business establishments? under the authority of the federal records act, the archivist has appraised most of the microdata from the economic censuses and surveys as having sufficient value, primarily for economic time series studies, to wanant continued preservation in the national archives. however, title 13 restricts access to this information to census bureau employees only. as discussed earlier, the federal records act limits statutory restrictions to thirty years unless extended by the archivist of the united states. this statute also empowers the archivist of the united states to direct and effect the transfer of any of these records which not used in the regular, current business of the census bureau. since such old economic information is not needed in the regular current business of the bureau of the census, the agency has agreed to transfer the information to the national archives when the information is thirty years old. the mere transfer of material to the national archives for continued preservation does not necessarily mean that the information is available to the public; the national archives routinely accessions material to which access is denied for a period of time. of course, any such restriction on access must be sanctioned by one of the exemptions of the freedom of information act a statutory restriction can be extended beyond thirty years by the archivist in consultation with the agency "for reasons spring 1986 iassist quarterly 1 consistent with standards established in relevant statutory law." because of the permissiveness of this authority. census and archives personnel have from time to time discussed when the national archives would be able to release census-gathered machine-readable information concerning individual economic entities. to date, however, no agreement has been reached. this review can allow one to draw some conclusions about records administration within the united states government the legal provisions which relate to records and archives are "media non-specific" in that the statutes relate to all record material regardless of medium. however, as seen in the 1952 agreement regarding census material, these policies and responsibilities generally have been developed to deal with human-readable records and have later been applied to all record material. the statutes divided responsibility for the administration of the record material. but in this division of responsibility, the archivist has significant powers which have an impact on access to the information. the first of these powers is the exclusive authority to sanction the destruction of record material. obviously, the destruction of a document or a machine-readable data file effectively limits access to il the archivist has the authority to direct the transfer of non-current material to his custody after the records are thirty years old. finally, the archivist has primary responsibility to determine whether statutory restrictions will be extended past the thirty-year statutory limil while only having a minimal impact on current or contemporary records, these latter powers can be decisive in determining access to older information. yet, because of this division of authority, disagreements are possible among those sharing records management responsibilities. until these differences are resolved, open questions — such as the ones about access to microdata from the economic census — will remain." spring 1986 sist newsletter vol.1, no. 2 store, and distribute process-produced data in such a manner that the kinds of uses have not to be invented after completing all these tasks. in particular, the development of a "source criticism" for mass data, analogous to the development of the methodology of survey data, can only be achieved in an interdisciplinary way. such efforts are necessary for the envisaged descriptors for machine-readable process-produced data. similarly, cooperation with those people working on record-linkage problems (e.g., oxford record-linkage study on medical record linkage) or with large-scale process-produced data (e.g. criminal statistics utilizing court-records) should be initiated.... 1 .3 quantities it is not possible to deal with these problems solely within the existing social science archiving movement which has so heavily concentrated on survey data. other institutes must be brouaht into a network of archives and information centers in which coordination and a division of labor must be planned. (examples of these include the "sozialdatenbank" [social-information-system of the departiflent of labour and social affairs, west germany], the proposed "zentrum fur aggregatdaten" [centre for aggregate data of the german national science foundation], and the national archives.") in germany, the prospects for coordination are a little bit better because of the pioneering work being done within the informationand documentation program of the federal government. book n ot i c e s / kathleen m. heim introduction this column is a preliminary step in defining the literature of data archiving. those of us who have tried to assess the state of the art in order to formulate annual reports, write articles, or keep professionally informed have been frustrated by the lack of bibl iograohic control over our area of concern. indexing and abstracting services such as social science citation index , information science abstracts , library literature , social science index and resources in education are unsystematic in their assignation of subject headings to pieces of literature related to data archives. the problem is further confounded by the fact that seminal information concerning the establishment of data archiving has often been distributed informally at conferences or in unpublished papers. when our numbers were small we could depend upon an invisible college network to disseminate important information. however, as our numbers grow and as new archivists enter the field without access to the established network, it becomes mandatory that we define and organize the literature of our profession. sist newsletter vol.1, no. 2 currently, a comprehensive annotated bibliography on data archiving is being compiled by alice robbin (lassist newsletter editor) and kathleen m. heim ( newsletter , book notices). the methodology for identifying relevant literature has taken place in three modes: 1) conventional indexes and abstracts were searched both manually and by computerized bibliographic retrieval systems; 2) a call for relevant papers was made to lassist members; 3) all citations in papers identified from methods 1) and 2) were traced. we have at this time identified almost 200 papers and, judging from the completeness of our files in relation to citations, feel we are getting close to a total control over the articles relevant to the history and perspectives of social science data archiving. to insure comprehensiveness, however, we would like to request that you check your personal files and to send us any citations you have for published or unpublished pieces of literature concerning data archives--both retrospective and current. through a cooperative effort we hope to control this aspect of our literature. please send citations as soon as possible to the lassist newsletter editor. we are especially anxious for citations relating to data archiving outside the u.s. as a spin-off from compiling the bibliography we felt it would be stimulating to include reviews of data archiving literature as a regular feature in the lassist newsletter . we will include a series of "retrospective" reviews highlighting important early writing as well as current reviews reflecting recent publications. after the annotated bibliography is published, this column will function as an updating service to that volume. suggestions for reviews are solicited and any ideas concerning better bibliographic control of the data archiving literature would be appreciated; send them to lassist newsletter , book notices, c/o kathleen heim, data and program library service, 4452 social science building, university of wisconsin-madison, madison, wisconsin 53706. nasatir, david. data archives for the social sciences: purposes, operations and problems . unesco. reports and papers in the social sciences: ss/ch w. t973: isbn-92-3-101052-2 (english edition) isbn-92-3-201052-6 (french edition). price: us $3, t 1; 12 f. plus taxes, if applicable. u.s. distributor: unesco publications center, p.o. box 433, new york, n.y. 10016. u.k. distributor: h.m. stationery office, p.o. box 569, london sel 9nh. french distributor: librairie de i'unesco, place de fontenoy, 75 paris -7 .ccp 12508-48. german (fed. rep) distributor: verlag dokumentation, postfach 148, jaiserstrasse 13, 8023, munchen-pullach. (for addresses of other national distributors see back of unesco publications.) under unesco resolution 3.221 the director-general was mandated "to study the conditions required for the establishment within an international centre of card indexes of archives of investigations carried out in the domain of the social sciences." dr. david nasatir of the international data library and reference service, survey research center, university of california at berkeley was contracted to carry out the feasibility study. this unesco report is dr. nasatir's assessment of the purpose of data archives, operational considerations ^ and future recommendations for co-ordination of international data archive development. 22. sist newsletter vol.1, no. 2 the 126 page document provides the novice with a coherent introduction to the multi-dimensional problems of the field and the experienced data archivist with a thoughtful analysis of procedural considerations. the report is divided into three chapters and ten appendices. a brief recapitulation of these will give a clear idea of the scope of this important publication. an overview of the data archive movement the rationale for and a brief historical development of data archiving is presented in this section. nasatir treats the problems traditional librarians have had with understanding the archive movement and forsees that the u.s. bureau of the census interest in having libraries manage the summary tapes might well make librarians more cognizant of the use of archives. [nasatir's prediction has been proven correct; traditional librarians are beginning to manifest a greater understanding of archives in their professional literature.] funding strategies and location of archives also receive careful consideration by nasatir who provides a flow chart of data from "input through archive status for distribution to users" as well as a copy of the [now defunct] u.s. council of social science data archives "check-list for study processing." nasatir identifies the greatest problem of data archives as achieving inter-archival coordination. his treatment of the many sides of this problem bears special consideration by the lassist membership: notably, his account of data archive organizations throughout the world. operational consideration of archives in this section nasatir outlines the technical problems faced by data archivists, such as acquisition, cleaning, inventorying, retrieving, diffusing and training. he considers staff requirements, space allocation, archival security and administration. this is the only guide to the operations of archives that exists and though nasatir is brief in his treatment of each point, he nevertheless takes into consideration the major aspects of daily archival work. problems remaining in the creation of data library infrastructure nasatir identifies three major problems that must be overcome before it will be possible to create an infrastructure for the development and use of social science data archives: administration, technical and political. under administrative problems he discusses 1) creating new archives; 2) recruiting personnel; and, 3) allocating priorities. under technical problems: 1) data management; 2) data retrieval; 3) analysis; and, 4) inventories. under political problems: location of long-term stable funding and the need for an international organization. [ed. note: david nasatir is one of the founders of lassist. ] appendices the appendices of the report draw together important data: 1) sample operating budget; 2) list of archives; 3) codebook standards; 4) machine readable codebook; 5) cleaning notes; 6) key word listing of study titles; 7) analysis request forms; 8) timing of operations; 9) standard format; 10) set-up budget. 23. sist newsletter vol.1, no. 2 it is impossible within the scope of a brief review to do more than qive a cursory description of the wealth of material nasatir has compiled. the recommendations in this report, many of which are already in the process of implementation, will stand as a seminal document for the data archivist. in addition to the practical aspects of archiving its tone is thoughtful, provocative and will help archivists to reassess the role of their profession and its importance to scholarship. recommended. an absolute necessity for every data archivist. ijyen, 0rjar. "social research and the protection of privacy: a review of the norwegian development." acta sociologica : 19 (1976): 249-262. the concern about the individual's right to privacy has been the focus of legislation, conferences, and professional forums in nearly every western nation. in his review of developments in norway, 0r,iar 0yen comments on the universiality with which social science research policies seem to develop, "they appear in a number of different national settings almost simultaneously, in similar time sequences, and with the thrust of largely identical kinds of rationale." 0yen refers specifically to 1) an increase in the amount of government or political concern with and control of social science research and 2) a proliferation of efforts to regulate the social scientists' relations to issues of confidentiality, the protection of privacy, and the maintenance of the integrity of individuals. while 0yen's focus is on norwegian developments his perceptions of the ramifications of these issues are valid to social scientists and data archivists throughout the world. he notes, "the joint operation of the increased control of social science research and the restrictions placed upon researchers' access to and utilization of data may have far-reaching consequences." after raising questions about the norwegian data committee's proposal to institute a data inspection agency, 0yen comments upon the possible effects of such legislation on social science research. the effects on research include 1) access to personal data as a premise of social science; 2) failure of the data committee to recognize that the relationship between researchers and individuals furnishing data is different than that between individuals and public agencies; 3) social scientists' practice of not releasing data to public agencies; 4) social scientists' need for identifiable data for dynamic studies; 5) differential treatment of fund allocation; 6) possibility that data inspection will result in censorship of research; 7) possible shift of research from empirical to speculative; 8) augmentation of research proposals overemphasizing criteria of relevance and usefulness according to whatever goals the concession authority possesses; 9) rules governing privacy which might function as a mechanism for protection of the agency or bureaucracy claiming to be bound by such rules; 10). 1 imitations on the training of recruits and new researchers; and, 11) risk of an arbitrary definition of sensitivity. 24. sist newsletter vol.1, no. 2 while 0yen's discussion focuses on norwegian developments, the reader will be able to draw parallels to his or her own national situation. 0yen's final caveat, "it would seem unfortunate if the increasinn need for social science research in the policy field, and the exoansion of public support and political control of social science developments, were to be linked to efforts to tie the hands of researchers through the introduction of rules which might give the social sciences a serious setback. in fact, such a development might be threatening another integrity requirement, namely, the right to understand." 0yen's review is an incisive look at the consequences of privacy protection legislation created without consideration of social science needs. it is a sophisticated analysis of issues which will continue to command the attention of researchers and data archivists in the next decade. recommended as essential reading--an insightful introduction to the problems of privacy protection for the social scientist and by extension, the data archivist. report on the committee of european social science data archives, january 22, 1977 meeting of the committee of european social science data archives, 22 january 1977 the committee of european social science archives was established at a i meeting in amsterdam in june 1976 and held its second meeting in paris on 22 january 1977. the meeting was attended by representatives of the seven member organizations: the adpss-milan, the bass-louvain-la-fleuve, the dda-cooenhagen, the nsd-bergen, the ssrc-sa-essex, the sa-amsterdam and the za-cologne. the meeting agreed to organize one conference a year and to encourage the establishment of a number of time-limited european working parties of 3-^ data organizations each. za-coloqne, offered to host the 1978 meeting: this will focus on privacy legislation . the committee also instructed bass to invite established, puhlic-service data organizations (archives and broader data services offering access for a range of universities, centres) all across the world to take part in the constituent meeting of an international federation of data organizations. this meeting is scheduled for 21-22 hay at louvain-la-fleuve. the new body would be based exclusively on organizational membership and would commit its members to support concrete projects of co-operation. the federation will, if established, serve to bring data archives, data services and data centres together as institutions and would offer a parallel structure to the individually-based lassist. there was full agreement that the two bodies should complement each other and help each other through joint activities. there was also agreement that the statutes of if-do should give full recognition to iassist and commit the federation to close co-oneration with the association. for further information please write to philippe laurent at bass. 25. 30 iassist quarterly 2015 iassist quarterly abstract in october 2014 at the fifth ddi moving forward sprint a subgroup met2 to focus on adding structure to ddi4 to support enhanced citation of data. a principal question was how to record the role(s) and degree of contribution of those contributing to the creation and curation of data. we also considered the question of which information objects associated with data creation might need enhanced citation information. we chose to think broadly about this, moving beyond the notion of citing a dataset to explore other types of intellectual objects that might merit some form of citation or annotation and reuse – for example, a new data collection method or a constructed variable. in thinking about roles we reviewed the credit taxonomy (allen et al. 2014) and decided that it would serve as a good foundation in ddi4 for an extensible vocabulary for roles. further, we determined that all ddi4 versionable objects should allow for the attachment of an annotation supporting citation along with role and degree of contribution. as a result of the dagstuhl meeting the initial releases of ddi4 will have an annotation object allowing for the attribution of roles and associated degree of contribution for creators and contributors to the creation of versionable objects. attribution information has also been proposed as a cdisc odm-xml extension planned for development in 2015. keywords attribution, contributor role, data citation, ddi, dublin core, cdisc, enhanced citation, metadata introduction it is common to cite traditional scholarly literature such as conference papers and journal articles, and the mechanism for doing this is widely accepted and practiced. citing research data, which also represent significant intellectual effort, is becoming more common, but best practices and norms for citing data are not yet widely accepted (borgman 2012). a related about the data documentation initiative (ddi) the ddi initiative was established by the interuniversity consortium for political and social research (icpsr) in 1995 with support from nsf and in 2003 transitioned to become a self-supporting membership alliance with over 40 current members who contribute to shaping the standard. ddi has two major development lines: ddi codebook, intended to document simple quantitative survey data, and ddi lifecycle, which is broader in scope, covering the research data life cycle from conceptualization to collection and processing to data publication and beyond. ddi’s primary goal is to document research datasets and processes thoroughly so that data are independently understandable. advantages of the ddi approach are that metadata are machineactionable and reusable. with origins in the quantitative survey-based social sciences, ddi can be used by researchers in other disciplines and can describe other types of data, such as experimental, observational, biological, administrative, and transaction data. originally expressed in xml schemas, ddi is now evolving as a model-based specification that will enable a variety of renderings. ddi and enhanced data citation by larry hoyle, mary vardigan, jay greenfield, sam hume, sanda ionescu, jeremy iverson, john kunze, barry radler, wendy thomas, stuart weibel, michael witt1 iassist quarterly 2015 31 iassist quarterly issue is that to make data usable, there is a need to provide more than simple attribution and location information. the practice of citing other constructs related to data objects such as instruments, questions, and variables is rare and, in general, lacks the consensus needed for it to emerge as a practice. yet citing at this more granular level is increasingly viewed as important both in terms of provenance chains and awarding credit (cdc 2013). contributorship is another important part of the picture. processes and procedures for attribution around data are just getting established, and it is clear that the life cycle of research data presents new possibilities for how we view contributions to the creation of research data. the development of a dataset has many stages and can involve a multitude of actors who make significant contributions to the final product but have traditionally gone unacknowledged. the idea of extending credit beyond the principal investigators and authors to others who have played critical roles is also being explored in research done by the harvard/wellcome trust (allen et al. 2014). this synergy offers an opportunity to think about contributorship with respect to research data in new ways. the data documentation initiative (ddi) is an open structured metadata standard for documenting and managing research data, and as such, it needs to address these key issues of data citation and contributorship. the standard needs to include all of the metadata elements necessary to cite and describe a data object and to support that citation. ideally all of these citation-related metadata are machine-actionable and can facilitate additional data discovery. to explore these related issues, a group of experts on data citation funded by the national science foundation (#1448107) met in october 2014 at schloss dagstuhl and worked alongside the ddi4: moving forward sprint. representatives from the dublin core metadata initiative and cdisc, the clinical data interchange standards consortium, were also in attendance. the group sought to answer some key questions: • what objects documented by ddi should be citable? • what elements are needed in ddi and cdisc to cite data and describe data sources in a comprehensive way across the lifecycle? • given the lifecycle focus of ddi, how can we support broad attribution for contributions, and how can we describe the level of contribution? this paper reports on the group’s consensus around these questions and sets out a list of recommendations to improve citation coverage in ddi. to test our decisions and assumptions, we created a sample dataset and went through the exercise of citing the dataset and related information objects using enhanced citation information. the need to attach other kinds of annotations having a function similar to citation to objects was another focus of the meeting. this might include administrative information such as the ombrequired information about the provenance of survey questions or characterization information such as parameter settings for an instrument. the structure for such information is not generally well known enough to be formalized in ddi, as it is typically defined and revised by some community of interest. this highlighted the need for ddi4 to include a structured information object capable of being validated from some external vocabulary. plans are underway for the development of a ddi4 object to be structured by an external vocabulary. current status of data citation the changing nature of scholarly and research communication is not new; for example, a notable effort reporting on this subject met at dagstuhl in 2011 [bourne et al. 2012]. the rapidly evolving research landscape requires us to revisit some of the traditional paradigms that have characterized citation and attribution, particularly when it comes to publication and citation of nontraditional research products, such as research data. the overall purpose of data citation has been articulated in two similar sets of data citation principles [force 11, codata-icsti task group]. they both state that data should be considered legitimate, citable products of research and be accorded the same importance in the scholarly record as citations of research publications. noting that no single style or mechanism may apply to all data or all disciplines, the groups advocating for data citation also mention the following general requirements: • credit: data citations should facilitate giving scholarly credit and normative and legal attribution to all contributors to the data. • evidence: whenever and wherever a claim relies upon data, the corresponding data should be cited. • unique identification: a data citation should include a persistent method for identification that is machine-actionable and globally unique. • specificity, verifiability, and utility: a data citation should lead to the specific data subset used (timeslice, version, etc.), to sufficient context to verify that the data used were the same as the data retrieved (fixity, provenance, etc.), and to code, documentation, and methods adequate for making informed use of the data. • interoperability and flexibility: data citation methods should be sufficiently flexible to accommodate the variant practices about the clinical data interchange standards consortium (cdisc) cdisc is a global, open, multidisciplinary, non-profit organization that has established standards to support the acquisition, exchange, submission, and archiving of clinical research data and metadata. cdisc is member-supported by approximately 350 biopharma, academic, and service provider organizations. cdisc’s mission is to develop and support global, platformindependent data standards that enable information system interoperability to improve medical research and related areas of healthcare. cdisc standards are vendor-neutral, platform-independent and freely available via the cdisc website. cdisc standards cover the full clinical research lifecycle from protocol through analysis and reporting, including regulatory submissions. the cdisc data exchange standards provide support for data and metadata exchange and archiving of the foundational and therapeutic area content standards. 32 iassist quarterly 2015 iassist quarterly among communities, but should not differ so much that they compromise interoperability of data citation practices across communities. an international consortium called datacite3 was created in 2010 to establish citation of data and other non-traditional research products as normal, mainstream research activities and to provide an infrastructure for the registration of digital object identifiers (dois) for data. datacite has published a list of minimal elements that should be part of a data citation (while acknowledging that data publication practices can vary across disciplines). these include creator, publicationyear, title, publisher, and identifier with optional properties of version and resourcetype. while these citation elements may be formatted in different ways, datacite recommends the format: creator (publicationyear): title. version. publisher. resourcetype. identifier. (the major style guides each have specified formats for data citations.) orcid (open researcher and contributor id) is a complementary effort designed to disambiguate contributors’ names by assigning globally unique researcher identifiers to link researchers to records of their scholarly output that are accurate and complete. citable objects in ddi in our discussions of citations and annotations more broadly, the enhanced citation working group at the dagstuhl ddi sprint (october, 2014) identified several generalized use cases for annotations as well as the need to distinguish among varieties and purposes of such annotations. for example, we might think about attribution (manifested through citations), administration, and characterization as forms of annotation. we use the term description to signify structured annotations (as opposed to unstructured annotations such as notes). a description-type is a specific set of metadata elements intended to support an identified functional requirement. just as declared data-types rationalize the management of variables in a programming language, declared description-types will help rationalize the management of classes of metadata sets within a data management application. we suggest that this notion of description-types be explored as the ddi alliance models annotations. generalized use cases for description-types we identify four general use cases that we believe justify different description-types. it is expected that additional use cases will emerge. citation: a form of attribution, this description-type is the familiar bibliographic notion of establishing the relationship of one or more individuals with a manifestation of an intellectual product, disambiguating that intellectual product from others, and, where practicable, facilitating access. citation answers (at a minimum) the following questions: • who is credited with creating the product? • what is the product named? • when was the product created? • where can the product be accessed? • whether: are there constraints on access to the product? the intellectual product referenced by a citation can take many forms, including books, articles, and reports. the creation of other intellectual products may also be credited. examples could include data, an algorithm, an instrument, and more. a citation recognizes the creation of the intellectual product and may also serve to help locate information about the product. there are needs, though, for description types that go beyond the simple attribution and location use case (see below). sourcing: within the social science and biomedical research communities, other object types need to be referenced. for example, in the us, questionnaire questions administered to more than 10 persons by a federal agency must be assessed by the office of management budget (omb) for the degree of burden imposed on respondents. omb requires that each such question be ‘sourced’ to facilitate review. the structured data for such sourcing will look very much like a citation, but should be typed differently to facilitate discovery and administration. the functional requirement is not intellectual attribution, but rather administrative responsibility: • where does a question come from? • is it an accepted and tested component of an existing instrument? • does it require further analysis or vetting? instrument description: data collected across a large sample may rely on multiple physical instruments manufactured by different manufacturers. the set of all such instruments may be thought of as a conceptual instrument, but differences among instruments from different manufacturers may yield systematic differences in data that can be normalized after the fact based upon known operational differences among instruments. one can imagine a rich variety of such problems that require instrument-description metadata to identify the manufacturer, operational provenance, operational characteristics, and more. the characterization metadata in this case may also serve to document the use of the instrument to create data rather than just its creation. an infrared thermometer, for example might have a switch allowing for either fast response with less accuracy, or slow response with more accuracy. a second thermometer might also have a switch for fahrenheit vs. celsius. knowing the switch settings used to collect a set of data could be important. a questionnaire, another type of instrument, might be administered on paper, or on a computer. instruments can also be seen as intellectual products and thus can be cited to attribute them to specific creators. dataset description: the use case for a structured dataset description is broad-based and fundamental to the work of ddi and other data initiatives. promoting appropriate conventions for such descriptions is a core responsibility of ddi. we define a dataset as: a discrete collection of measurements collected via observation, experiments, or analysis, using specified methodologies and instruments, and structured in a manner documented by formal schemas. iassist quarterly 2015 33 iassist quarterly the unbounded diversity of datasets mitigates against any single means of characterizing or cataloging them. however, communities of practice exist and can be encouraged to coalesce around common conventions for structuring their data and dataset descriptions so as to promote discovery, reuse, and preservation. see, for example, the data discovery index that nih proposes as part of its larger bd2k initiative. such an index will provide pointers (actionable links) to dataset metadata that reside (and are managed) elsewhere. these pointers would form a ‘data catalog’. from the bd2k data discovery index workshop summary report (emphasis added): ‘a catalog could in some cases be a human-viewable database analogous to a traditional paper catalog and in other cases, could be a set of functions to serve both human users and, increasingly, machine interfaces (i.e., ‘computers talking to computers’) to support the needs of scientific data discovery, exchange, and analysis […] unlike the printed catalog of the past, it can be predicted that a new resource that enables locating, characterizing, and accessing nih-funded data will have to evolve in an agile way to serve both data producers and data consumers to keep pace with the ever-changing, networked world. technical approaches to describing, finding and providing access to the broad variety of ‘data objects’ that are the output of contemporary biomedical science will likely continue to evolve and improve.’ the abstract requirements of a dataset description might include the following: • who (institutional responsibility, authorship, funding sources) • what (title[s] and project description) • when (date of publication) • where (access points) • whether (management of access: who can use, edit, reference a dataset) • how (how was data collected, and what additional information may be necessary for its interpretation) • structure specifications (formal machine-interpretable schemas necessary for parsing the dataset) • provenance (reuse and modification history) the characterization information in this dataset description supports the traditional citation content and helps to facilitate data reuse. in cdisc, the define-xml standard (http://cdisc.org/definexml) provides the metadata to describe cdisc datasets such as sdtm or adam, and is required when these datasets are part of an fda submission. define-xml v2.0 meets most of the stated requirements with the exception of whether. we emphasize that these requirements are speculative and will evolve in the context of communities of practice. rationale for identifying and promoting typed references these generalized description-types (citation, sourcing, instrumentdescription, dataset-description) share a requirement for structured metadata, some elements of which will be common to several description-types. other description-types will require elements that may be specific to the particular description-type, or even to the specifics of a given instrument or experiment (e.g., sensor type or calibration history or a schema specifying the structure of a dataset). additional description-types will likely emerge as well. ddi should support the import of metadata elements from established metadata dictionaries such that description-types reflect the needs of existing projects and established communities. it should be noted that there is inherent tension between the objectives of (a) incorporating metadata practices established in existing communities and (b) encouraging the reuse of metadata elements by promoting standardization across communities. this is a socialization function that ddi can encourage, but cannot expect to achieve entirely. the metadata world is a messy place. as description-types evolve, the ddi community has a role in encouraging the reuse of elements from one description-type in another while avoiding the overloading of semantics that might create ambiguities. for example, a manufacturer can be mapped to a creator and date of manufacture can be mapped to publication date, but structured descriptions appropriate for an instrument-description will quickly diverge from the whowhat-when-where-whether of citation metadata. care must be taken to avoid conflation of semantics that will cause ambiguity or confusion. one of the most important benefits of description-typing will be to help create shared understanding among users. complicated systems such as ddi require shared understandings of functional requirements, component definitions, and relationships among the parts. designers, modelers, software developers, practitioners (system managers and creators of metadata), and end-users achieve shared understandings through natural language. to overload widely understood concepts (such as citation) with other description-types such as an instrument-description is to risk obfuscation of both description categories and violate shared user-models. finally, a search-view of ddi will benefit from distinguishing among the functional requirements implied in description-typing. the conventional notion of citation metadata leads users to expect to find such data in a coherent collection of item records of the who-what-when-where-whether form. those searching for instrument-description metadata or dataset descriptions will expect to find it in collections of records organized to reflect respective functional requirements. to summarize, we propose a description-typing approach that has the following characteristics: • metadata element sets exist in communities of practice, and should be welcomed into ddi, even when not formally sanctioned or managed by ddi. • ddi should promote, but cannot enforce, the standardization of metadata practices that support cross-community discovery. • description-types are discrete sets of structured annotations that serve specific functional requirements and are characterized by carefully selected metadata elements that meet these functional requirements and promote coherent discovery. we identify four such types here, but expect others to emerge. 34 iassist quarterly 2015 iassist quarterly w5 hsp proposed ddi property dublin core mapping ddi3.2 mapping what label (type = title) title label and/or citation/title   who   creator role degreeofcontribution  creator citation/creator when publicationdate date citation/publicationdate what userid identifier urn and userid and/or citation/ internationalidentifier and/or citation/ dc:identifier where publisher publisher citation/publisher and/or citation/ dc:publisher who contributor role degreeofcontribution contributor citation/contributor (with role) and/or citation/dc:contributor what language language citation/language and/or citation/ dc:language whether copyright rights citation/copyright and/or citation/dc:rights whether license rights? none? archive/item/access/ accessconditions, dc:accessrights? note: dc:license to be added in ddi3.3. what description description description and/or citation/dc:description, abstract? hsp userattribute dc:any, or none citation/dc:any or none, userattributepair what label (type = subtitle) subtitle citation/subtitle and/or citation/dc:subtitle what label (type = alternatetitle) alternatetitle citation/alternatetitle and/or citation/ dc:alternative when/p datecreated created citation/dc:created when/p datemodified modified citation/dc:modified what/p version isversionof ? physicalinstance/pi:datafileversion or r:version of the containing element what resource type   kindofdata w5 hsp pointer to metadata    physicalinstance/ r:datarelationshipreference and r:recordlayoutreference where actionable link to the dataset   datafileidentification/ datafileuri and/or r:location table 1. ddi proper*es suppor*ng data cita*on (w5 plus how, structure, and provenance) "1 iassist quarterly 2015 35 iassist quarterly we believe this approach can account for existing citation metadata approaches and provide a platform for developing emerging standards for referencing data objects. it also supports special purpose annotations (such as sourcing and instrumentdescriptions) and affords a flexible foundation for the evolution of description-types that are currently unforeseen. unresolved issues: • who has responsibility for introducing, defining, registering, and managing description-types? • how is a description-type declared in data instances? • do all description-types have a set of obligatory and optional metadata elements (as has been proposed for citations)? • are there constraints on the sources of metadata elements? • are there means by which ddi can help to ‘socialize’ metadata best practices within its domain? • how can description-types defined by a user community be validated? how can relationships among the metadata elements in the description-type be described? elements for citing and describing data ddi elements ddi lifecycle version 3.2 currently allows structured annotation metadata to be provided for several versionable object types, including studyunit and physicalinstance. however, we recommend that these fields be available on all versionable object types because, as noted previously, there is a growing need to recognize effort for non-traditional objects. moreover, having all versionable object types contain the same basic annotation properties increases consistency in the ddi standard and can also reduce redundancies in the current ddi lifecycle model, where, for example, the same identifying information can be recorded in multiple places. beyond the minimal set of metadata comprising a citation, there are other elements often used to administer, characterize, and validate data objects. for example, an author may want to provide copyright and license information for a question response scale. of course, the determination of when or what objects to cite is determined by researchers electing to re-use existing data, and is not a function of the standards. however, ddi needs to make it possible to cite and describe objects comprehensively. we propose the following properties be added to each versionable object. the properties are categorized as who, what, when, where, standard element odm-xml odm, study, protocol, studyeventdef, formdef, itemgroupdef, itemdef, codelist, codelistitem, enumerateditem, methoddef, condibondef, user, clinicaldata, subjectdata, studyeventdata, formdata, itemgroupdata, auditrecord define-xml valuelist, commentdef table 2. cdisc elements to be used to support data cita6on 1 whether, how, structure, and provenance (w5 hsp) as suggested by iso 19773. we also suggest that some additional properties and elements be further explored for possible inclusion. examples include a permanence or stability indicator (e.g., as in the national library of medicine vocabulary ) and data fingerprint. when adding structured annotation metadata, the attributes should be added at the highest applicable level in the hierarchy, and this information will apply to lower levels in the hierarchy that do not explicitly include such information. when an object does not have certain information, relationships can be followed to other objects in order to discover that information. for example, if a variable does not specify a creator, a user or system could look for creator information in the data file that contains the variable. if the data file metadata does not specify a creator, a user or system could look at the study description that contains the data file. this type of relationship traversal can be performed in a machineactionable manner. for example, a web application could show creator information on a variable’s page, even if the information was pulled from the study description. cdisc elements cdisc proposes to support data citations for key cdisc odm-xml and define-xml data and metadata elements (see table 2). the cdisc data exchange standards have not implemented the dublin core metadata element set, but do include several attributes that correspond to terms from dublin core including the operational data model (odm-xml) element attributes creationdatetime as date, originator as creator, description as title, and fileoid as identifier. the missing attributes will be added as an odmxml extension, based on the dublin core metadata terms where appropriate. to ensure that citation metadata does not demand the addition of redundant metadata that could cause integrity issues, the core attributes are populated using existing odm-xml attributes where available. in cases where odm-xml does not support the required attributes, an extension containing the new elements and attributes has been proposed. table 3 shows examples of odmxml elements and attributes that map to the proposed cdisc data citation properties. contributors and contribution the practice of attributing credit through the citation of datasets and other non-traditional scholarly objects is becoming more common but lacks the maturity of the traditional scholarly literature citation paradigm. one issue is that the number and 36 iassist quarterly 2015 iassist quarterly w5 hsp proposed cdisc properties element / attribute who creator /odm/@originator when date /odm/@creationdatetime what title /odm/@description what identifier /odm/@fileoid where publisher /odm/@dc:publisher who contributor /odm/study/metadataversion/dc:contributordef what language /odm/study/@xml:lang whether rights /odm/study/metadataversion/dc:rightsdef who creator /odm/study/metadataversion/itemgroupdef/@dc:originator what title /odm/study/metadataversion/itemgroupdef/@name what identifier /odm/study/metadataversion/itemgroupdef/@oid where publisher /odm/study/metadataversion/itemgroupdef/@dc:publisher whether rights /odm/study/metadataversion/itemgroupdef/@dc:rightsoid table 3. examples of odm-xml data cita6on informa6on 1 nuance of contributions to datasets are much greater than can be adequately expressed when reduced to an ordered list of authors, such as for a scholarly paper. pressure to give credit and attribution is expected to increase as federal research funders require data sharing and management plans with grant proposals and intend to track their outputs. the scholarship of data does not fit neatly within the model of traditional scholarly publication and requires recognition of new contributor roles and contributions. ddi lifecycle defines different stages of the research process from concept to data collection, processing, archiving, distribution, discovery, analysis, and repurposing of data9 the ddi controlled vocabulary group (ddi-cvg) has been developing a controlled vocabulary, the ddi controlled vocabulary for lifecycle events10, which names and defines a set of actions that constitute recognizable contributions to the entire data life cycle from project inception to data use and reuse (research process). the ddi lifecycle events controlled vocabulary primarily focuses on contributions related to research data in the context of the social and behavioral sciences. in addition, the cvg drafted the ddi controlled vocabulary for contributor roles, a controlled vocabulary of agents who perform specific actions that make up contributions11. to give an example using both vocabularies, the agent may be a data collector, whereas the ddi lifecycle event in which the action takes place may be data collection. similar activities have taken place outside of ddi. haeussler and sauermann (2014) analyzed patterns of contribution using the five-level categorization of contribution type requested by plos one (conceived, performed, analyzed, materials, wrote). allen et al. conducted a workshop and initiated a series of studies to begin to define and standardize a taxonomy (see appendix a) to help researchers identify their contributions to collaborative projects, primarily in the context of the preparation and publication of scholarly papers. categories of contribution were classified and defined by giving high-level examples of contribution actions. this taxonomy was evaluated by authors of scholarly papers and generally accepted with 85 percent of them saying that it was easy to use and covered all the roles of contributors to their papers. eighty-two percent of respondents reported that the taxonomy was at least the same or better in terms of accuracy than how contributions to their paper had actually been recorded. a follow-up study asked authors to reconstitute the submission of their original papers using these contributor roles, and this experiment was deemed successful. feedback indicated that a weight for contribution was missing, so a simple scale was created—lead, equal, and supporting—to augment the taxonomy, which was further revised and named the contributor roles taxonomy (credit). the project’s leaders are currently pursuing formal standardization through consortia advancing standards in research administration information (casrai) in conjunction with the national information standards organization (niso). if ddi were to adopt the credit taxonomy, it could increase and improve associations between objects in ddi and those outside of ddi that also share these terms. to investigate this possibility, similarities and differences between the ddi lifecycle events and credit were explored through a ‘stub’ mapping of the vocabularies iassist quarterly 2015 37 iassist quarterly to each other. early results suggested that ddi may benefit by adding some concepts to its vocabulary from credit such as software, formal analysis, resources, and funding acquisition. many if not most terms from ddi lifecycle events could be seen to fit into credit, but some gaps or mismatches were evident and warrant more thorough analysis. also the grounding of the credit taxonomy is a scholarly paper, and while it is not exclusive of data, there are potential gaps and alignment issues if the primary scholarly work is a dataset and is not assumed to be a paper. it is a good time for exploring these options because ddi is in the process of designing a new version (ddi4), and the ddi controlled vocabulary for contributor roles has been approved by the cvg but has not yet been published and could be extended. in terms of a mechanism to record role of contributor and weight of contribution in ddi, the following recommendations were made: 1. a citable object in ddi should provide sufficient information to build a citation that can include one or more contributors, e.g., the name of the contributor. 2. contributors can be classified by using a reference to the credit taxonomy, including weight of contribution (lead, equal, and figure 1. screenshot of the text miner process supporting) as properties of contributor. this is the recommended practice, although it should be possible for a different taxonomy that includes contributor roles to be referenced and used. 3. the same weight scale of contribution from credit can be added as a property of creator. creator may also have a property of role. 4. ddi should collaborate with the harvard-wellcome initiative and give input to close any gaps and align mismatches such that a shared taxonomy is applicable to the ddi lifecycle in particular and scholarship of data in general. 5. the ddi cvg should consider expanding its ddi controlled vocabulary for contributor roles and continue its examination of other similar lists of roles such as those created by datacite and onix. it should situate its controlled vocabulary for contributor roles into its lifecycle events for internally consistent mapping. a sas dataset use case to apply the principles and practices being discussed, the group created several use cases, including one in which a quantitative dataset was created from a qualitative dataset through text mining. for the latter we chose the raw minutes created as google docs during the first three days of the meeting. the derived dataset had figure 2. the topics table 38 iassist quarterly 2015 iassist quarterly as its unit of analysis the topics computed by the text mining software. we decided to also create an example variable to show how it could have structured annotation information useful for a citation. it became clear that the text mining procedure could itself serve as an example instrument, in that it is essentially a ‘black box’ with a set of inputs – data and parameters -and an output – a dataset. a structured annotation for this procedure would include documentation of all of these inputs. since the set of parameters for the text mining procedure is unique to this software (sas), this annotation would need its own instrument description-type. an external vocabulary could be developed and referenced allowing the recording of the parameters used to generate any other dataset with the same software. we exported the minutes of the first three days of the workshop from three separate google docs into microsoft word and concatenated them into a single text file in ultraedit. this process made each paragraph in the original documents into a single line in the text file. then we wrote a sas program to read the minutes into a sas dataset. a sas enterprise miner, text miner process produced a topics dataset and a clusters dataset from the minutes dataset using the default options (see figure 1). all of the parameter values for the default options were exported to an xml file to allow for future replication of the process. from the results window, we saved the topics results as a sas dataset (figure 2), and the resulting dataset was then modified in sas enterprise guide. a new variable was computed, combining the topic number, the number of documents using the topic, and the list of most highly weighted terms for the topic. since figure 3. enterprise guide process flow the variable names for the topic result table are standard, this is a reusable variable that can be recomputed. we used the topics2 dataset and the new variable (topicdescription) as objects to be cited. using an enterprise guide (eg) add-in (hoyle 2013), we added additional metadata to the sas dataset and generated a ddi3.2 instance and a codebook from that. this additional metadata included attribution information (creator, contributors, funding information, etc.) and other descriptive information (coverage, methodology, etc.). see appendix 1 and appendix 4 of the full use case and documentation, available in ku scholarworks. the complete set of extended attributes for the dataset is listed in appendix 5. extended attributes for the variable topicdescription are listed in appendix 6. figure 3 shows the enterprise guide process flow diagram. text miner is a separate application so that is represented in the flow by a note. representing structured annotation information in ddi3.2 the ddi3.2 instance for our example dataset (topics2) includes the following elements, reflecting the need for information on attribution (who, what, when, where, whether), and characterization (how, structure, and provenance). who – creator, contributor, fundinginformation we did not include institutional responsibility or a reference to a persistent researcher identifier. ddi3.2 allows both creator and contributor to reference a structure which iassist quarterly 2015 39 iassist quarterly can contain references to external persistent identifiers (like orcid). it is not clear, though, how to document institutional responsibility for different phases of the data lifecycle in ddi3.2. what title, description, abstract, version when creation date, modification date, and publication date in some datasets there may be other date/ time references required. retrospective studies may ask respondents about some time period in the past. these can be documented in r:temporalcoverage. where publisher, pointer to metadata, actionable link to dataset we provided a handle (hoyle et al. 2014) pointing to a landing page having links to the data files and associated metadata. this (scholarworks) landing page is not really structured to provide a persistent actionable link to each dataset, although it does provide a separate url to each. whether access rights, copyright, license, permanence it is not clear whether accessrights and license are both required. how processingdescription, generationinstruction, language, collectionmethodology, relatedresource the metadata includes descriptive text about the method used to generate the dataset, its source, and collection methodology. structure logicalproduct, physicalinstance a ddi3.2 description of the data was harvested from the sas dataset. this should allow machine-actionable interpretation of the structure of the dataset. provenance – provenance is present only as unstructured descriptions in collectionmethodology, processingdescription, and generationinstruction. citation-related information for the dataset in ddi3.2 all of the elements we propose for versionable objects (see table 1) could be documented in ddi3.2 with the following caveats. • contributor was entered as a structured string including both role and degree of contribution. ddi3.2 would allow multiple elements, each of which could include a r:, but not a degree of contribution. • copyright information was not provided but could be structured in a element. • license appears as . in ddi 3.3 will be available. • in many small research projects userid may not be clear in the context of a dataset which is being developed. a dataset has a name, unique within a file system folder, but not necessarily unique outside of that context. a ddi identifier is not certain to remain the same during development of the file. once archived, a dataset will probably have a unique identifier. • the description also included coverage information – spatial, temporal, and topical (subject). note also that creation date, modification date, and publication date were all included. citation-related information for a variable in ddi3.2 we created a new variable (topicdescription) that could be reusable with the topic result table from any sas text miner text topic node instance. the citation information for that variable is different from the dataset. ddi3.2 doesn’t allow an r:citation element to be attached to a variable, so the citation information was structured in a set of r:userattributepair elements. these are listed in appendix 6. note that the creator and contributor information for the variable is different than for the dataset as a whole. allowing structured annotation information to be attached to any versionable object in ddi4 will make the structure of this information more consistent. instrument description the text mining procedure used to produce the topics dataset can be considered as a ‘black box’ instrument that takes a text dataset and produces a quantitative dataset. documenting the use of this instrument to allow someone to reproduce the results requires recording all of the parameter choices made in using this instrument. this instrument description description-type can require a large number of information objects unique to the particular instrument. appendix 2 shows the values of the 96 properties set for the run of text miner used to generate the topics dataset. at the time of this analysis these were the ‘default’ choices, but there is no guarantee that the default values will remain the same for future versions of the software so listing them is important for replication. sas enterprise miner allows the export of diagram properties as an xml file. the tables shown in the appendix were processed from that file. in the case of text miner many properties are relevant to only one node in the process. these properties, for example, are only relevant for the textparsing node: • delimit = std, • bcapitalize=y, • stoplist=sashelp.engstop, each of the nodes in this process has its own set of inputs and outputs and might be considered sub-instruments linked by their inputs and outputs. some of these properties, like stoplist, point to a data file (sashelp. engstop). this is a list of terms that will be ignored in the computations within and following the text parsing node. in this case, then, an input parameter can be complex – e.g., the contents of another dataset. sample citations16 here we show how citations for the dataset and the variable might be listed in three common styles. 40 iassist quarterly 2015 iassist quarterly use case dataset apa – hoyle, larry (2014). topics generated from minutes from nsf1448107 group at dagstuhl event 14432 [data file, codebook, ddi metadata] http://hdl.handle.net/1808/15746. mla hoyle, larry. topics generated from minutes from nsf1448107 group at dagstuhl event 14432. university of kansas, 2014. web. 17 nov 2014. chicago hoyle, larry. topics generated from minutes from nsf1448107 group at dagstuhl event 14432. lawrence kansas: university of kansas. 2014. http://hdl.handle.net/1808/15746. all three styles for citing a dataset leave out contributors and cited author roles: contributors: larry hoyle (conceptualization, lead; methodology, lead; software, lead; formal analysis, lead; data curation, lead), mary vardigan (conceptualization, equal), sam hume (conceptualization, equal), sanda ionescu (conceptualization, equal), jay greenfield (conceptualization, equal), jeremy iverson (conceptualization, equal), john kunze (conceptualization, equal), barry radler (conceptualization, equal), wendy thomas (conceptualization, equal), stuart weibel (conceptualization, equal), michael c. witt (conceptualization, equal) variable – topicdescription apa – hoyle, larry (2014). topic descriptor combining sequence number, number of related documents, and terms list from a sas text miner text topics node result table [variable]. http://hdl. handle.net/1808/15746. mla hoyle, larry. topic descriptor combining sequence number, number of related documents, and terms list from a sas text miner text topics node result table. university of kansas, 2014. web. 17 nov 2014. chicago topic descriptor combining sequence number, number of related documents, and terms list from a sas text miner text topics node result table. lawrence kansas: university of kansas. 2014. http://hdl.handle.net/1808/15746. in the three examples above only the apa style indicates that the object being cited is a variable. the mla style doesn’t yield a persistent identifier. the handle shown above points to a landing page (hoyle et al. 2014) with a description of the collection and more than a dozen urls to objects within the collection (original raw data, software code, a codebook, a ddi instance). an explicit link to the data file and another explicit link to the structured metadata for the data would be much more machine-actionable. none of these styles allow for designation of a role or degree of contribution for the creator or a listing of contributors and their roles. if the standard citation styles included an explicit reference to structured metadata, including some mechanism for identifying the structure style, both of these problems could be handled by machine-actionable searching of the metadata. implications for ddi4 following the dagstuhl meeting the ddi4 model will have the following features supporting enhanced citation: • all objects in ddi4 are identifiable (except for a few primitives) • all versionable objects will have an annotation • an annotation can have attribution information including creator and contributor • creator and contributor can have a list of role, degree of contribution pairs • the annotation should include the possibility of an additional set of information structured from an external vocabulary (scheduled for release 2 of ddi4) this last feature addresses the need for ddi4 to have a mechanism to allow the incorporation of a set of information objects with a vocabulary drawn from an external controlled vocabulary. ideally this mechanism will include the capability of validating those objects and also indicating relationships among the objects. designing this mechanism will be a task for the ddi4 modeling group. this need comes up both for instrument parameters and for the vocabulary for creator and contributor roles. the latter might have a hierarchical structure. at the top level of this hierarchy we propose using the credit taxonomy (appendix a). in a hierarchy a ‘software’ role, for example, might be more specifically described as ‘algorithm development.’ ddi3.2 allows for attaching role to contributor but not creator. in a large study co-principal investigators may have specialized roles that should be documented. in each case role should also be paired with a ‘degree of contribution’ measure. we propose using the credit taxonomy as the top level of a taxonomy for describing role and a three-level category (e.g., ‘lead’, ‘equal’, and ‘supporting’) for degree as in the credit proposed standard. input parameters may also be complex objects, including datasets, as noted for the stoplist dataset. parameters might also come from stream sources at specific times. the group recommendation to allow structured annotation information to be attached to any versionable object will yield a more consistent structure for this information. for citation type information the addition of role and degreeofcontribution to creator and contributor, along with the elements already present in ddi3.2, should allow for a usable set of information. for other description types, though, ddi4 will need to support external controlled vocabularies for attribute names and complex data types (including datasets) for attribute properties. conclusion we began the meeting at dagstuhl with a set of questions related to enhancing the citation information available in ddi, with the goal of informing the initial releases of ddi4. our initial thoughts were focused mainly on attribution of credit for intellectual creation (a traditional citation). along the way, however, we realized that there are other classes of information that can be referenced like a citation. some of these are fairly well understood, like the information characterizing a dataset. this is the information that the ddi initiative has been developing over its 20-year history – what the data represent, when they were collected, how they were collected, why they were collected, and whether they can be used. structure for other information may not be so well developed, or may be known only to a special community. we realized the need for ddi to be able to point to external structures and to incorporate that information for specific cases as needed. there may be many cases in which attribution information will need iassist quarterly 2015 41 iassist quarterly to be recorded along with this additional, externally structured, information. other questions cannot be addressed by the structure of ddi. common adoption of conventions for contributor role and degree of contribution will result in a significant expansion of the information expected to be provided in a citation. imagine reference sections and curricula vitae in fields where large scale multi-authorship is common if all the roles and degrees of contribution were to be listed. clearly some sort of common infrastructure for locating structured annotation information is needed. there are such efforts under way (e.g., casrai17, datacite). requesting or requiring additional attribution information will increase the demand on researchers to provide metadata, already often seen as a burden. future work might address the kinds of tools that could lessen this burden. incorporating the collection of metadata, including attribution-related information, into the research workflow rather than considering it an additional task to be undertaken at the conclusion of a project might help ease the friction and improve the quality of the metadata. properly designed tools might help. training in research practices might also encourage better practices (e.g., long 2009). finally, we leave it to others to develop algorithms and software to generate metrics for enhanced citation and to determine the best way to store, harvest, and display this information in a machine-actionable way. multi-dimensional information could be collected – roles by degrees by numbers of collaborators, as well as the traditional creator vs. contributor distinction. will a univariate metric be adequate, or should this complex of information be represented in a more nuanced simplification? contributorship is clearly a fruitful topic for further research. contributors contributors to the project leading to this paper are listed below along with their roles. attribution of roles was self-assigned using a web based form. roles and degree of contribution are structured using the credit taxonomy. micah altman – conceptualization, supporting; methodology, supporting; software, none; validation, none; formal analysis, none; investigation, none; resources, none; data curation, none; writing – original draft, none; writing – review & editing, none; visualization, none; supervision, none; project administration, none; funding acquisition, none jay greenfield conceptualization, equal; methodology, equal; software, none; validation, none; formal analysis, none; investigation, equal; resources, none; data curation, none; writing – original draft, equal; writing – review & editing, equal; visualization, none; supervision, none; project administration, none; funding acquisition, none larry hoyle (orcid http://orcid.org/0000-0002-8262-2393) conceptualization, lead; methodology, lead; software, lead; validation, none; formal analysis, lead; investigation, equal; resources, none; data curation, lead; writing – original draft, equal; writing – review & editing, lead; visualization, lead; supervision, lead; project administration, lead; funding acquisition, lead sam hume conceptualization, equal; methodology, equal; software, none; validation, none; formal analysis, none; investigation, equal; resources, none; data curation, none; writing – original draft, equal; writing – review & editing, equal; visualization, none; supervision, none; project administration, none; funding acquisition, none sanda ionescu conceptualization, equal; methodology, equal; software, none; validation, none; formal analysis, none; investigation, equal; resources, equal; data curation, none; writing – original draft, supporting; writing – review & editing, none; visualization, none; supervision, none; project administration, none; funding acquisition, none jeremy iverson (orcid http://orcid.org/0000-0003-30029245) conceptualization, equal; methodology, supporting; software, none; validation, supporting; formal analysis, none; investigation, none; resources, none; data curation, none; writing – original draft, supporting; writing – review & editing, supporting; visualization, none; supervision, none; project administration, none; funding acquisition, none john kunze (orcid http://orcid.org/0000-0001-7604-8041) conceptualization, equal; methodology, equal; software, none; validation, none; formal analysis, none; investigation, equal; resources, none; data curation, none; writing – original draft, equal; writing – review & editing, supporting; visualization, none; supervision, none; project administration, none; funding acquisition, none nancy cayton myers conceptualization, none; methodology, none; software, none; validation, none; formal analysis, none; investigation, none; resources, none; data curation, none; writing – original draft, none; writing – review & editing, none; visualization, none; supervision, none; project administration, none; funding acquisition, supporting barry radler conceptualization, supporting; methodology, supporting; software, none; validation, none; formal analysis, none; investigation, none; resources, none; data curation, supporting; writing – original draft, equal; writing – review & editing, supporting; visualization, none; supervision, none; project administration, none; funding acquisition, none wendy thomas conceptualization, supporting; methodology, supporting; software, none; validation, none; formal analysis, none; investigation, supporting; resources, none; data curation, none; writing – original draft, none; writing – review & editing, equal; visualization, none; supervision, supporting; project administration, none; funding acquisition, none mary vardigan (orcid http://orcid.org/0000-0002-6168-6531) conceptualization, lead; methodology, lead; software, none; validation, none; formal analysis, none; investigation, equal; resources, none; data curation, none; writing – original draft, lead; writing – review & editing, lead; visualization, none; supervision, equal; project administration, lead; funding acquisition, lead joachim wackerow conceptualization, equal; methodology, none; software, none; validation, none; formal analysis, none; investigation, none; resources, supporting; data curation, none; writing – original draft, none; writing – review & editing, 42 iassist quarterly 2015 iassist quarterly supporting; visualization, supporting; supervision, supporting; project administration, none; funding acquisition, supporting stuart weibel conceptualization, equal; methodology, equal; software, none; validation, none; formal analysis, none; investigation, equal; resources, none; data curation, none; writing – original draft, equal; writing – review & editing, equal; visualization, none; supervision, none; project administration, none; funding acquisition, none travis weller conceptualization, none; methodology, none; software, none; validation, none; formal analysis, none; investigation, none; resources, none; data curation, none; writing – original draft, none; writing – review & editing, none; visualization, none; supervision, none; project administration, none; funding acquisition, supporting michael witt (orcid http://orcid.org/0000-0003-4221-7956) conceptualization, equal; methodology, equal; software, none; validation, none; formal analysis, supporting; investigation, equal; resources, supporting; data curation, supporting; writing – original draft, equal; writing – review & editing, supporting; visualization, supporting; supervision, suppporting; project administration, supporting; funding acquisition, none references allen, l., scott, j., brand, a., hlava, m. & micah altman (2014) publishing: credit where credit is due? nature. 508 (april). p. 312–313. doi:10.1038/508312a. available from: http://www.nature.com/ news/publishing-credit-where-credit-is-due-1.15033. [accessed: 28/01/2015] borgman, c. (2012) why are the attribution and citation of scientific data important? in ulhlir, p. (rapporteur). for attribution - developing data attribution and citation practices and standards: summary of an international workshop. isbn: 978-0-309-267281. available at: http://nap.edu/catalog.php?record_id=13564. washington, dc: the national academies press. [accessed: 28/01/2015] bourne, p.e., clark, t.w., dale, r., de waard, a., herman, i., hovy, e.h. & david shotton (2012) improving the future of research communications and e-scholarship (dagstuhl perspectives workshop 1133). in dagstuhl manifestos. isbn 2193-2433. 1 (1). doi: 10.4230/dagman.1.1.41. casrai (2015). credit -an open standard for expressing roles intrinsic to research. available from: http://credit.casrai.org/. [accessed: 28/01/2015] center for disease control (cdc)--behavioral risk factor surveillance system (brfss) (2015) suggested citation styles. http://www.cdc. gov/brfss/questionnaires.htm#citation. [accessed: 29/10/2014] codata-icsti task group on data citation standards and practices (2013) out of cite, out of mind: the current state of practice, policy, and technology for the citation of data. data science journal 12 (2013). p. cidcr1-cidcr75. available from: https://www. jstage.jst.go.jp/article/dsj/12/0/12_osom13-043/_pdf. [accessed: 29/10/2014] datacite. datacite metadata kernel (2014) available from: http:// schema.datacite.org/meta/kernel-3.1/doc/datacite-metadatakernel_ v3.1.pdf [accessed: 29/10/2014] force11-the future of research and communications e-scholarship (2014) data citation principles. available from: https://www.force11. org/datacitation [accessed: 29/10/2014] häussler, c., & sauermann, h. (2014) the anatomy of teams: division of labor and allocation of credit in collaborative knowledge production (may 7, 2014). available from ssrn: http://ssrn.com/ abstract=2434327 or doi10.2139/ssrn.2434327. hoyle, l. (2013) using extended attributes in data analysis software controlled vocabularies, tools and ddi. in: proceedings of the 5th annual ddi users conference december 2013, paris, france. available from: http://www.eddi-conferences.eu/ocs/index.php/ eddi/eddi13/paper/viewfile/86/86. [accessed: 28/01/2015] hoyle, l. (2014) a sas dataset use case for enhanced data citation. available from: https://kuscholarworks.ku.edu/bitstream/ handle/1808/15746/asasdatasetusecaseforenhancedcitation. pdf?sequence=19&isallowed=y. [accessed: 28/01/2015] hoyle, l., vardigan, m., hume, s., ionescu, s., greenfield, j., iverson, j., kunze, j., radler, b., thomas, w., weibel, s. & witt, m.c. (2014) comprehensive citation across the data life cycle using ddi work products from the nsf1448107 sponsored group attending dagstuhl event 14432 in october 2014. available from: http://hdl. handle.net/1808/15746 long, j. scott (2009) the workflow of data analysis using stata. college station, tx: stata press. mooney, hailey how to cite data: general info. available from: http:// libguides.lib.msu.edu/citedata. [accessed: 28/01/2015] notes 1. jay greenfield began his professional career developing artificial intelligence applications in a university environment. these apps could both listen and speak. they were expert systems that assisted in medical diagnosis and circuit board trouble shooting. in 1992 jay joined westat where eventually he became the leading technologist. at westat jay used his knowledge of artificial intelligence to create actionable metadata. he used actionable metadata to spawn survey research questionnaires, conduct data management and create data dictionaries. while at westat jay led the development of data collection and data analysis systems for many major health and nutrition studies including the medicare current beneficiary survey (mcbs), the medical expenditure panel survey (meps), the integration of the continuing survey of food intake by individuals (csfii) with the national health and nutrition examination survey (nhanes) and the development of the original legislation for the children’s health insurance program (chip). in 2008 jay joined booz allen as a study and data management expert on the national children’s study. currently, jay is a health informatics architect working with data standards, data standard groups and terminologies to annotate medical data with metadata in order to facilitate search, research and discovery. larry hoyle (orcid http://orcid.org/0000-0002-8262-2393 ) is a senior scientist at the institute for policy & social research at the university of kansas and was principal investigator for the nsf grant (1448107) which funded the enhanced citation group at dagstuhl event 14432. he is a member of the ddi moving forward advisory group. he was also the first chair of the north american data documentation initiative conference (naddi). correspondence may be addressed to larryhoyle@ku.edu or 1541 lilac ln. suite 607 blake, lawrence ks 66045-3129. sam hume is vice president of share technology and services at cdisc. at cdisc he leads the share project and co-leads the xml technologies team. sam has over 20 years of work experience in clinical research informatics. previously, he worked as director of is architecture at astrazeneca, vp of technical operations at phoenix data systems and chief technology officer at cb technologies. sam iassist quarterly 2015 43 iassist quarterly has an ms in information science, ms in telecommunications, and is completing his doctorate in healthcare informatics. sanda ionescu has been with icpsr since 1999, working to implement and support the data documentation initiative (ddi) — an xml-based specification for social science data documentation. she manages ddi-related projects and serves as the icpsr representative in the expert committee of the ddi alliance. at icpsr she is also part of a team that provides user support for data and documentation issues. within the ddi alliance she participates in the efforts to develop and promote the ddi standard, and leads the controlled vocabularies working group that produces classifications for studyand variable-level metadata. she holds an ma in communication from the university of massachusetts-amherst, and a ba in english and french from the university of bucharest, romania. jeremy iverson (orcid http://orcid.org/0000-0003-3002-9245 ) is a co-founder and partner at colectica where he helps build software to document statistical data using open standards. previously, he was a programmer at the wisconsin longitudinal study, working to process, document, and disseminate data for the longrunning study. jeremy is currently an invited expert on the data documentation initiative’s technical committee. john kunze (orcid http://orcid.org/0000-0001-7604-8041 ) is an identifier systems architect at the california digital library. a former bsd unix hacker who helped standardize urls and dublin core metadata, his current work focuses on the ezid service, the n2t resolver, ark identifiers, dataset citation, and lightweight metadata dictionaries. barry radler is a researcher at the university of wisconsin institute on aging. his research interests focus on understanding how human beings process information, make decisions, and behave in social, political, and marketing contexts. for the last 20 years he has explored, advocated, and implemented the use of information technologies to improve research processes and data. past and ongoing projects include an optical character recognition system for survey data entry, investigating mode effects between mail and online surveys, using custom computer applications to identify the processes and output of different cognitive systems, and, most recently, the application of web-based metadata standards to document complex longitudinal datasets. wendy thomas is the director of the data access core in the minnesota population center (mpc) at the university of minnesota and has been active in the data and information technology community for over 25 years providing data research and support services to academic, governmental, non-profit, and for-profit researchers. she has been a coordinating member of the state data center program since 1990 and is a former president of the u.s. association of public data users. she has been active in the work of the data documentation initiative (ddi) since 1997, chairs the ddi technical committee and is a member of the ddi moving forward advisory group. her work in the mpc covers the preservation and documentation of historical census data and supporting materials for the ipums international projects. her major publications focus on data documentation and the impact of standards on data dissemination and preservation. for more information, http://users. pop.umn.edu/~wlt/ mary vardigan (orcid http://orcid.org/0000-0002-6168-6531 ) holds the position of archivist at the inter-university consortium for political and social research (icpsr) where she directs the icpsr collection delivery unit, which involves oversight of activities in the areas of metadata, publications, web site development, user support, and membership development. she also serves as director of the data documentation initiative (ddi), an international collaboration to establish a metadata standard for the social and behavioral sciences. she is involved in other projects related to data stewardship, including the data seal of approval, the research data alliance, and various efforts to promote data citation. she was co-pi for the nsf grant (1448107) that funded the enhanced citation group at dagstuhl event 14432. stuart weibel worked in oclc research for 25 years, where he contributed to research in web standards for libraries, digital libraries, and convened workshops that led to the formation of the dublin core metadata initiative. michael witt (orcid http://orcid.org/0000-0003-4221-7956 ) is the head of the distributed data curation center (d2c2) and an associate professor of library science at purdue university. he is also the project director for the purdue university research repository and editor-in-chief of databib. witt serves on the organizational advisory board of the research data alliance, the editorial board of information technology and libraries, and the dmptool steering committee. sponsors for his research include the institute for museum and library services, microsoft research, and the alfred p. sloan foundation. for more information, http://www.lib.purdue.edu/ research/witt. 2. this meeting was funded in part by nsf grant 1448107. facilities support was provided by schloss dagstuhl – leibniz center for informatics, the site of the meeting (http://www.dagstuhl.de/14432). 3. datacite https://www.datacite.org/ 4. orcid. http://orcid.org/ 5. http://www.whitehouse.gov/sites/default/files/omb/inforeg/ statpolicy/standards_stat_surveys.pdf 6. iso 19773:2011, http://www.iso.org/iso/catalogue_detail. htm?csnumber=41769 7. national library of medicine permanence levels, http://www.nlm.nih. gov/psd/pcm/devpermanence.html 8. universal numeric fingerprint, http://thedata.org/book/ universal-numerical-fingerprint 9. ddi lifecycle, http://www.ddialliance.org/what 10. ddi lifecycle events cv, http://www.ddialliance.org/specification/ ddi-cv/lifecycleeventtype_1.0.html 11. draft ddi roles cv, [unpublished] https://docs.google.com/ file/d/0b5as3-mimlfkq3dptnziq2hxtnm/edit 12. report on the international workshop on contributorship and scholarly attribution (16 may 2012) http://projects.iq.harvard.edu/ attribution_workshop/files/iwcsa_report_final_18sept12.pdf 13. nature 508, 312–313 (17 april 2014) http://dx.doi. org/10.1038/508312a 14. appendices numbered appendices are available in hoyle 2014. appendices a and b follow in this document. 15. url for data http://kuscholarworks.ku.edu/bitstream/ handle/1808/15746/topics2.sas7bdat?sequence=11&isallowed=y url for extended attributes companion file http://kuscholarworks. ku.edu/bitstream/handle/1808/15746/topics2.sas7bxat?sequence= 12&isallowed=y 44 iassist quarterly 2015 iassist quarterly url for ddi3.2 metadata instance http:// kuscholarworks.ku.edu/bitstream/handle/1808/15746/ nsf1448107topicsusecase2014_11_09. xml?sequence=10&isallowed=y 16. example styles taken from from how to cite data: general info, http://libguides.lib.msu.edu/citedata 17. casrai http://casrai.org/ iassist quarterly 2015 45 iassist quarterly a note about appendices: all numbered appendices are available online in the document: hoyle 2014. project data are archived at: https://kuscholarworks.ku.edu/handle/1808/15746. appendix a harvard/wellcome trust taxonomy (credit taxonomy) a classification of the diverse roles played in the work leading to a research output. the classification includes, but is not limited to, traditional authorship roles. when there are multiple people serving in the same role a ‘degree of contribution’ should be further specified as either ‘lead’, ‘equal’, or ‘supporting’. roles are intended to apply to all those who contribute to a project — and it is recommended that, if possible, all contributors be listed, whether or not they are formally listed as authors. it is also intended that multiple roles be assigned to a single person where appropriate. roles and their descriptions are listed below from http://credit.casrai.org/proposed-taxonomy/. #1 conceptualization ideas; formulation or evolution of overarching research goals and aims. #2 methodology development or design of methodology; creation of models. #3 software programming, software development; designing computer programs; implementation of the computer code and supporting algorithms; testing of existing code components. #4 validation verification, whether as a part of the activity or separate, of the overall replication/reproducibility of results/experiments and other research outputs. #5 formal analysis application of statistical, mathematical, computational, or other formal techniques to analyse or synthesize study data. #6 investigation conducting a research and investigation process, specifically performing the experiments, or data/evidence collection. #7 resources provision of study materials, reagents, materials, patients, laboratory samples, animals, instrumentation, computing resources, or other analysis tools. #8 data curation management activities to annotate (produce metadata), scrub data and maintain research data (including software code, where it is necessary for interpreting the data itself ) for initial use and later re-use. #9 writing – original draft preparation, creation and/or presentation of the published work, specifically writing the initial draft (including substantive translation). #10 writing – review & editing preparation, creation and/or presentation of the published work by those from the original research group, specifically critical review, commentary or revision – including preor post-publication stages. #11 visualization preparation, creation and/or presentation of the published work, specifically visualization/data presentation. #12 supervision oversight and leadership responsibility for the research activity planning and execution, including mentorship external to the core team. #13 project administration management and coordination responsibility for the research activity planning and execution. #14 funding acquisition acquisition of the financial support for the project leading to this publication. 46 iassist quarterly 2015 iassist quarterly appendix b citation related objects in the ddi4 model the diagram below shows the elements added to the ddi4 model during the dagstuhl sprint and its immediate follow-up. all objects except for primitives and complex data types inherit from annotatedidentifiable which, in turn, has an annotation. an annotation contains attributes of creator, contributor, and publisher of type agentassociation. an agentassociation has a role attribute of type pairedcodevaluetype that allows a codevalue (a role from a specified set of roles) to be paired with an extent (a degree of contribution) also drawn from a specified vocabulary. class annotationview annotation title :internationalstring [0..1] subtitle :internationalstring [0..n] alternatetitle :internationalstring [0..n] creator :agentassociation [0..n] publisher :agentassociation [0..n] contributor :agentassociation [0..n] date :annotationdate [0..n] identifier :internationalidentifier [0..n] copyright :internationalstring [0..n] language :codevaluetype [0..n] typeofresource :codevaluetype [0..n] informationsource :internationalstring [0..n] versionidentification :xs:string [0..1] versionresponsibility :agentassociation [0..n] abstract :internationalstring [0..1] relatedresource :resourceidentifier [0..n] provenance :internationalstring [0..n] rights :internationalstring [0..n] annotatedidentifiable concept question conceptualvariable datastore most objects inherit from annotatedidentifiable. examples: agentassociation agent :bibliographicname [0..1] role :pairedcodevaluetype [0..n] agent pairedcodevaluetype extent :codevaluetype [0..1] codevaluetype codevalue :xs:string [0..1] codelistid :xs:string [0..1] codelistname :xs:string [0..1] codelistagencyname :xs:string [0..1] codelistversionid :xs:string [0..1] othervalue :xs:string [0..1] codelisturn :xs:string [0..1] codelistschemeurn :xs:string [0..1] e.g. extent codevalue = conceptualization e.g. role = lead individual organization machine the annotation object will also have an additional property capable of containing administrative, characterizing, and other information structured by an external vocabulary 0..1 agentassociation 0..* hasannotation sist newsletter vol. 1, no. 3 papers presented at the second lassist north american meeting, may 11 12,1977 [editor's note: the following are abstracts of papers presented at the second iassist north american meeting. the full proceedings of the conference will be available in the near future. all participants of the conference will receive a full copy of the proceedings ; lassist members will receive a copy at a reduced rate; all other individuals should contact sharon henry of the data clearing house for the cost to non-iassist members.] lassist progress towards solving problems in data archiving carolyn geda inter-university consortium for political and social research the paper reviews the establishment and development of lassist including the rationale for such an association. problems encountered during the formation of the international organization are presented with interim proposals for dealing with them. the governing body, regional structure, and functions of the action groups are discussed. a brief review of regional activities is given and potential directions for the association are projected. canadian secretariat report rachel des rosiers data clearing house for the social sciences the report summarizes the membership campaign for lassist in canada. it outlines the results as well as expected plans to increase the membership in canada, e.g., by contacting related associations in the field and publicising lassist information in professional journals and at conferences. it also discusses its relationship with the american secretariat, especially in regard to action groups' projects and plans. finally, it invites comments and suggestions from the canadians present concerning its intended role and functions. ibi^sist newsletter vol.1, no. 3 lassist secretariat report: united states judith rowe princeton university the report of the united states secretariat covers the current status of membership enrollment as well as proposed membership recruitment activity. specifically, it addresses the goal of recruiting for action group activity all of the professional staff of each data archive and data library. it includes a general summary of both secretariat and action group activity, as well as a report on the first north american action group conference in cocoa beach. other topics include an initial report on the plans for next year's conference and some activities and programs planned by other organizations which would be of interest to lassist members. among the latter are the annual conferences of special libraries association (sla), association of public data users (apdu) and the american association for information science (asis), as well as the inter-university consortium for political and social research (icpsr) workshop for data librarians. data archive registry-survey of past effort, suggestions for the future lisa lasko canadian consortium for social research the mandate of the lassist action group for data archive registry is described. the major portion of the paper is devoted to a brief review, description, and evaluation of some of the more important existing directories. these are: social science data archives in the united states , published by the council of social science data archives in 1967; a directory of information resources in the united states: social sciences , 2nd edition, published by the library of congress in 1973; the second edition of the encyclopedia of information systems and services , edited by anthony kruzas and published in 1974; and the directory of data bases in the social and behavioral sciences , edited by vivian sessions and published in 1974. recent developments in data archive registries, such as the unesco sponsored directory of data services and the directory of data centres , to be published by the data clearing house for the social science in canada are also discussed. essential elements and user requirements of social science data archive directories are dealt with in reference to these directories. finally, a summary of the problems facing the data archive registry group, both abroad and in canada, is given, along with several proposed options include a discussion of the viability of a data archive registry group in canada; unique contributions to be made in this area, one of which might be a compilation of a list of subject headings for social science data archives; and, the production of a directory of data archive personnel. 20. laissist newsletter vol. 1, no. 3 problems in handling process-produced data john devries social science data archive carleton university the paper is an attempt to outline the problems one encounters in the handling of process-produced data, starting with their definition through the stages of obtaining information about them and acquiring them, to the phases of transforming and analysing them. in addition to enumerating and discussing the often unique problems with which the user of process-produced data is faced, i will try to mention various projects which have already begun, and indicate future plans by the lassist action groups on process-produced data. definitions . in the last few years, various definitions have been suggested. although there is considerable overlap between them, the remaining differences make it important that we come to some agreement about the proper "domain" for an action group on process-produced data, and thus to a clear definition which unambiguously specifies this domain. i will discuss the definitions of which i am aware, and present some arguments in support of the definition which was developed at the cocoa beach conference. information and acquisition . generally, we will find that process-produced data have been generated by governments and other public organizations. problems in this phase are thus to a large degree subsumed under the general heading: "the relationship between the research community and the government-" specific issues i will discuss include: privacy versus the right to information, i.e., the government as a research resource; and, obtaining information from governmental agencies. documentation . in many non-trivial ways, process-produced data differ from surveys. some of the unique aspects regarding documentation are: definitions of the universe to which the data pertain, description and estimates of the errors of coverage and content, and requirements for a "codebook". data transformation . once a data set has been acquired by a researcher or an archive, there are various difficulties one faces in transforming and analyzing the data. many data sets are "ragged" rather than rectangular, thus requiring either intricate transformations or specialized statistical packages. aggregated data require a high level of statistical sophistication on the part of the analyst. finally, one will frequently wish to link data from various data sets. in doing so, one often runs into problems of comparability: stimuli may not be standardized across data sets; for identical or comparable stimuli, response categories are frequently not identical across data sets; different data sets may not relate to exactly identical universes; finally, where spatial and/or temporal delimiters are involved, they are often not identical across data sets. )i^slst newsletter vol. 1, no. 3 data acquisition for archives pierre lacasse university of sherbrooke the purpose of this paper is to raise a number of issues about data acquisition in order to narrow down the subject to the most important avenues for further discussion. with so diversified stock of social data being created, we need to ask "who will keep what, and on what basis"? to clarify this question we will particularly discuss three major inputs to the subject: (1) the collectors of data: who are they? what types of data are they collecting? what data are we interested in? (2) the archives: what are they? what are their purposes in holding data files; and for what types of secondary users? what mandate do they have and what is their scope of interest? (3) the users for whom data is held: who are they? where do they stand? on what issues? what are their needs? how will they use the data held by the archives? after discussing these questions, we will take a look at what has been done in the data acquisition action group. in particular we will discuss the questionnaire prepared by the european section of lassist on "archives data acquisition policies and problems." consider the group's mandate, which requires proposing ways of linking collectors and archives, with the stress on defining the group objectives and conditions which must be met in order to close the gap between collectors and archives. finally, we will try to suggest some possible courses of action in the short and long term. organization of data archives laine ruus university of british columbia the mandate of the lassist data archive development action group is two-fold: the creation of a "procedures manual consolidating current archival organizational, administrative, and personnel structures, procedures, and pol icies. . .to aid developing archives", and the organization of training workshops and seminars to aid in personnel training and professional development. the efforts of the data archive development action group have been entirely concentrated on the development of a guide to providing social science data services . while response to past workshops has indicated the need for this guide, plans are for the moment in abeyance. problems in data archive development occur at three levels. intra-archival problems are primarily administrative: planning and policy making, staff, acquisition, data management, technical and user services. intra-institutional problems, impinging on the former, occur in the areas of planning, inter-departmental cooperation, support services, and funding. problems at the inter-archival level result primarily from lack of adequate standards and conventions and from underdeveloped formal channels of information dissemination. l^$sist newsletter vol. 1, no. 3 standards for data documentation dave salley statistics canada in recent years, a number of initiatives related to the concept of data documentation have been undertaken within the canadian federal government. the motives for these initiatives have centered around the need to inform users on data availability, the need to monitor duplication of data collection, the need to control questionnaire content and the need to assess the costs incurred in the collection, compilation and dissemination of federal information. while these motives are not exhaustive, they do represent a significant class of problems for consideration by the lassist action group. the shortcomings of most efforts to date have centered on the particular or special purpose nature of the systems. in many cases, a particular subject area or class of users restricted the approach severely. in short, little consideration has been given to the development of standard components of a data documentation system with a view to meeting general requirements. the reason for this problem may well be that information managers in the federal government have not yet developed a real appreciation of the need for standard documentation procedures, both as a tool for managers and planners of information activities and to assist "end data users". thus, most initiatives have been quite ad hoc and narrow in scope. however, one should not underestimate the useful technical approaches and methodologies developed to date. as a departure point for the action group, it is suggested that terms of reference be developed which contain a strong "standards" ingredient. furthermore, an intensive needs and benefits study should be undertaken with respect to the whole data documentation issue. certainly, the group should bring together current work in a systematic fashion, but the development of standard approaches in the scope, technique, uses and purposes of data documentation systems will have far more useful results in the long run. cataloguing and classification of machine-readable data files; a preliminary report on the us lassist classification cataloguing project sue dodd university of north carolina the primary emphasis of the us classification action group of lassist has been on establishing standards and on the study of library information systems as they may apply to social science data files. some of the recent developments within the library system which the classification group is examining include: 1) the development of rules and guidelines for cataloguing machine-readable data files (mrdf); 2) the development toward the acceptance of the marc (machine-readable catalog) record format as a universal standard for the automated bibliographic record; 3) the development of networks and on-line information systems which allow for multiple input and immediate retrieval of information; 4) the development of thesauri and controlled vocabularies for social science terms; and, 5) the development towards future considerations of a national union list of available mrdf and their location. i»i$sist newsletter vol. 1, no. 3 given the current work and interest in cataloguing mrdf, the first task of the lassist classification group v/as to participate in an organized project designed to test the feasibility of cataloguing social science data files according to the american library association (ala)'s designated subcommittee to recommend rules for cataloguing mrdf. to facilitate the test, which was conducted by mail, a "working manual for cataloguing machine-readable data files" was compiled based on an interpretation of the ala subcommittee's recommendations. the task required that the participants select six data files, either numerical, text or program files, and proceed to catalogue these data using the information and guidance in the manual; apply subject descriptors on the content of the files; and complete an evaluation form. the outline of the project, the actual test, and the preliminary results will be discussed in this paper. data organization and management applications in data archiving greg morrison social science data archive carleton university the formal mandate of the data organization and management action group of iassist is used as the starting point for consideration of various kinds of problems with which the group might concern itself -problems which most data archivists will confront sooner or later. these include: the transfer of datasets; the cleaning, editing, transforming, merging and sub-setting of data; the organization of data files for the above activities and for statistical analysis; the technical aspects of the documentation of individual datasets by codebooks; and, the technical aspects of the documentation of collections of datasets by catalogues or inventories. while an obvious organizing focus of the group is computer software, it will try to define appropriate procedures in these areas, and to encourage the exchange of information about programs and procedures. the activities of the group to date are reviewed, and some suggestions are made for the future. the impact of microcomputer technologies on dissemination of integrated social, demographic and related statistics by robert johnston' and graham templeman introduction the relation between integration and dissemination the united nations has been concerned with general issues in measuring development and levels of hving and related social, economic and environment conditions since the beginning of the organization, pursuant to the promotion of "higher standards of living, full employment, and conditions of economic and social progress and development" as set forth in the charter of the united nations (anicle 55). over the past five years this work has received new stimulus in the united nations statistical office fi-om the great interest at national and international levels in statistics and indicators on women and special population groups such as youth, elderly and disabled persons, and from new interest at national and international levels in the compilation and use of social statistics and indicators to design and implement social policies, to monitot the achievement of social objectives and to monitor the social impact of economic adjustment policies. in response to these emerging interests and priorities, the statistical office has undertaken a substantial reorganization of its methods of compilation, presentation and organization of social and related demographic and economic statistics and indicators. this work follows up the development in the 1970s of the united nations framework for integration of social, demographic and related statistics (fsds), and the preliminary guidelines on social indicators published by the united nations in 1978.^ it has been particularly oriented to compilation and dissemination of integrated indicators drawing on a wide range of sources and aimed at non-specialists, rather than compilation of primary data, which are detailed, technical, and speciauzed by field. this work has been greatly facilitated by the rapid development and now nearly universal availability of highly standardized microcomputer hardware and software technologies for statistical work. as described in the united nations handbook on social indicators^ , the development of integrated social statistics and indicators is a wide-ranging and multifaceted process which aims to bring together basic statistics from many different fields and data collection programmes and recompile them for many different purposes. the handbook provides a basic core of structure, concepts and methods for use in this process and thereby promote the compilation and dissemination of social statistics and indicators to better meet a wide variety of user needs through more effective integration and use of the basic data. unfortunately, the cost and methodological difficulties of bringing together social statistics and indicators from many disparate and often intractable primary and secondary sources have held back work on social indicators in many countries and internationally. even where statistical services have built up a considerable volume of basic data, this has by no means ensured the ready availability to users of indicators relevant to their specific purposes and concerns, including policy issues. for such purposes, close collaboration between users and producers of indicators, and detailed attention to data requirements for indicators at the stage of designing data collection programmes, and co-ordination within a framework such as fsds are needed. thus much of the work on social indicators in the 1970s and early 1980s was concerned with precise identification of user interests and requirements and their translation into well-structured statistical methods and concepts. co-ordination and integration of social and related classincations for dissemination oflntegrated statistics and indicators one of the technical features which works on fsds and has been stressed from its inception has been the development and harmonization of social and related classifications for integration and for indicators. however, the development and harmonization of social and related classifications is a complex process for many of the same reasons that social statistics and indicators present statistical offices with so many difficult problems of organization and methodology. in general, the subject matter is extremely heterogeneous and the relevant statistics come from a very wide range of sources, each with established traditions, procedures and objectives and often administered more or less independently of the central statistical service. these circumstances are similar at national and international levels. the initial development of fsds in towards a svstem of demographic and social statistics and in the preliminary guidelines on social indicators established a basic spring 1990 37 subject-matter and classifications framework for the further development of classifications and indicators. these early reports served to clarify what classifications were relevant to integrated social statistics and in what ways. with the current interest in improved multidisciplinary compilation and dissemination of social statistics and indicators and given the new technical possibilities of microcomputers for bringing together data in microcomputer daui bases, the importance of harmonized classifications emerges all the more clearly. a basic jminciple of work in the statistical office in this area continues to be the importance of the close linkage between so-called basic statistics and statistics for integration, and for indicators that are, basically for general disseminations. thus it has never been suggested that new classifications should be developed for integration or for indicators, which would in any case be a technically and organi2ationally impossible task, given the degree or decentralization of responsibility for statistical classifications at national and international levels and the large number of competing interests and technical problems that must always be delicately balanced in preparing any kind of recommendation on classifications. what the statistical office undertook in the preliminary guidelines on social indicators and has now been made much more explicit in the handbook of social indicators , is to recommend abstracting from existing classifications shorter forms which are needed for integration and for indicators. as the draft handbook states, once the fields and topics for indicators have been outlined in an indicators programme at the national or international level, basic statistical classifications for use in indicators should be developed. these must, of necessity, be based on the classifications used in the basic data but for purposes of indicator compilation these source classifications often require careful adaptation. the process of adaptation should be undertaken with three objectives in mind: (a) meeting specific indicator requirements; (b) abbreviating classifications as much as possible to simplify compilation and presentation of indicators; (c) devising classifications into which data from a variety of sources often using differing classifications or variants of classifications, can be fitted as consistently as possible; (d) identifying population groups of special policy concerns. all of the classifications referred to in the illusu-ative series and basic data tables for indicators in the handbook are usted in the table below which also shows the fields in which they are used. sixteen of these are considered basic classifications in the handbook . five of these concern demographic and social characteristics (sex and age group, national or ethnic group, household size and composition, household headship and level of education); three are geographical (urban and rural areas, cities and urban agglomerations, and geographical regions); four concern activity characteristics (occupation, status in employment, socioeconomic group and time-use); and four are classifications from economic statistics (percentage distributions of household income and consumption, kind of economic ^tivity (industry), functions of government and institutional sector). these basic classifications can be used to provide a firm foundation for the development of indicators in all of the fields covered by the handbook . they were selected for discussion as basic classifications on the basis of (a) their substantive importance fw indicators, usually in more than one field and drawing on multiple data sources, (b) the extent of their importance and use for indicators in national and international experience, and (c) the relative detail and complexity required in their use for compiling statistics for indicators. all but one of the basic classifications are shown and discussed in the illustrative formats for basic data tables of the handbook, drawing on the relevant international recommendations. the exception is classification by national or ethnic groups. in this case, national and experience and circumstances are so diverse that no international recommendations are feasible and even an illustrative classification could not serve any useful purpose. principles of integ:ration applied to dissemination of statistics and indicators on women and special population groups interest in the development of statistics and indicators on women and other population groups that are considered to be of special relevance for policy planning has given considerable impetus to a range of activities concerned with statistics and indicators on these groups. the principal groups on which work has bosn concentrated in the united nations statistical office are women (beginning with the world conference of the international women's year in 1975), disabled persons (beginning with the international year of disabled persons in 1982), youth (in connection with international youth year in 1985) and children. there has been interest in the development of statistics and indicators on the elderly (in connection with the world assembly on aging in 1982 and the international plan of action). in international compilation and dissemination of indicators on women, for example, a substantial quantity of data is being routinely collected in international statistical services and supplemented, in many cases, with standardized international estimates and projections. the rapid spread of microcomputers and the ease of use of spreadsheet techniques have now made it feasible to compile iliese data in one source, using the fsds framework, disseminate them to users cheaply and quickly on diskettes, and prepare user-oriented software and documentation for reference, analysis, table-generation and similar uses. a special project with these objectives was established in the statistical office in 1984, and this work was basically completed in 1987. the united nations women's indicators and statistics data base (wistat) consists of 72 microcomputer {assist quarterly spreadsheet files (currently using lotus 1 -2-3) ranging in size from approximately 2wcb to 150kb and totalling about 12mb. wistat is available from the statistical office on 22 microcomputer diskettes complete for 178 countries and areas or for specific regions, using the forms p-ovided in the printed user's guide (currently available, in part, as a statistical office working paper). wistat will be fully documented in the user's guide, to be issued in final form as a sales publication of the united nations. a listing of statistical series and topics in this data base is given in the annex below. using quite different underlying technical methods of organization and compilation but identical microcomputer hardware and software an international disability statistics data base was also completed by the statistical office in 1987, comprising detailed statistics on disabled persons from censuses and surveys in 55 countries and areas between 1975 and 1985. like the women's data base, the disability statistics data base (distat) is disseminated on diskettes. it consists of 34 microcomputer spreadsheet files ranging in size from about 7kb to 314kb and totalling about 3.3 mb. the files are described in detail in united nations disability statistics data base. 1975-1986: technical manual ." the complete data base is available from the statistical office on 12 microcomputer diskettes using the forms provided with the technical manual . version 1 of the data base (as of 3 1 december 1987) contains (a) information on sources and availability of statistics on disability for 95 countries or areas for various years between 1960 and 1986, and (b) detailed statistics on disabled persons from national censuses, surveys and other data sources from 55 of those countries or areas for the period 1975-1986. finally, the basic strategy and framework for organizing social statistics for social indicators, as set out in the handbook on social indicators , were adopted by the statistical office for preparation of the compendium of statistics and indicators on the situation of women 1986 and the compendium of social statistics and indicators 1986 .' that is, highly simplified basic data were compiled from primary international sources into microcomputer spreadsheets. once in the spreadsheets, new series and indicators could be calculated and data transferred within and among spreadsheets with great flexibility and minimal time and effort and, once final table formats were agreed, they could be tested and then generated in final form very quickly. on this basis, series and classifications such as those given in the handbook have been prepared for these two compendiums, with the possibility of recalculating percentages, rates, ratios, distributions, reaggregations and the like and of juxtaposing series from different sources that may be of interest almost at will. a short, preliminary version of the social compendium was prepared using tliese techniques for the united nations interregional consultation on social welfare policies and programmes held in september 1987* and generated considerable interest among delegates with no special statistical background. overall, it appears that microcomputer hardware and software for spreadsheets, data bases and analysis are at the leading edge of basic changes in the development of social statistics and indicators at national and international levels. the effects are now beginning to be seen on a wide scale and at the same time the technologies are advancing and spreading so rapidly throughout the world that the (section and full implications of these changes are still not completely understood or fully appreciated the potential role in dissemination of spreadsheets and data bases on microcomputers. spreadsheets and their users spreadsheets are used extensively by pet^le interested in statistics, and are a useful tool in most cases. spreadsheets are characterized by a row/column cellular ak)ro^h to data in which each cell may be (typically) a number, a character string, or a numeric or logical function of the values in other cells. cells are named by a column/row coordinate system. for example: spreadsheets are capable of displaying or printing the cells in row/column format, and manipulating the cells individually or by rows, columns, or groups of these. they are, effectively, the "cell processing" equivalent of a word processor. a word processor imposes minimal structure on its atomic units (words) except to organize them within given boundaries such as paragraphs and margins. a spreadsheet, on the other hand, maintains a positional relationship between its atomic units (cells) which is capable of being displayed as a two-dimensional row/column table. the facilities for manipulation offered by spreadsheets and word processors have many parallels, such as "cut and paste", insertion and deletion, and search and replace. word processors have some facilities peculiar to themselves, such as reformatting between redefined spring 1990 margins, text flow from line to line, and so on. spreadsheets also have some peculiar features such as row and column transposition, default cursor movement by row or column first, and so on. it is not surprising, therefore, that people dealing with statistical tables are drawn to spreadsheets in much the same way that typists and writers are drawn to word processing programs, by the facility far direct interactive visually-based control. beyond these "cell-{kxxessing" capabilities, some spreadsheets offer what they call "database capability". (see, for example, lotus 1-2-3 tutorial manual , v.2.01, p. 5.1) spreadsheets "database capability" comes from an analogy with some aspects of relational database theory. this type of data wganization can be simulated using a spreadsheet, by using the relation name as the spreadsheet name, and by treating rows as records or instances of the relation and columns as fields or attributes, as long as the user sets up the data in the format of a single relation. processing. a second consequence of the affinity of spreadsheets for human readability is that spreadsheet designers tend to follow the make use of the horizontal left to right direction of reading. this results in the expansion of more variables hwizontally than vertically, creating hierarchies of column headings. in the wistat data collection, for instance, almost every spreadsheet uses one row per country or area, with column hierarchies up to four levels deep. for example see table below: some simple arithmetic shows that this hierarchy generates 24 columns and that the "adult" heading will appear 8 times. an interrogation asking for listing of those countries for which, in 1975, the number of adults 1975/1980 persons prosecuted/convicted male/female total/adult/juvenile 1 1975/1980 prosec/conv maie/female tot/adult/juv 1 1 1 when the user sets up the data in "flat file" format, i.e. with one column representing the key field, all columns with unique names, and so on, the row and column manipulations available to the spreadsheet user parallel some of the simpler facilities available to a relational database user. for example, sorting rows by the key field, searching a column for a particular value or for values falling within a range. in some spreadsheets such as lotus1-2-3 it is possible to use foreign key values located by a search such as a range search to extract rows from another spreadsheet. in this way the "cut and paste" spreadsheet operations simulate the linking of different relations. when such operations are exp-essed as a macro, i.e. a named sequence of operations which can be invoked by its name, the result can be quite efficient for some purposes. it is important to realize, however, that with spreadsheets there is only a limited connection between the column name and the data. the true column name is its coordinate. if a column is removed or added to a spreadsheet, thus pushing other columns to new coordinates, any operation which refers to coordinates will have to be redefined. this is not a problem for perfectly stable data. spreadsheet statistical tables in presentation format tables with hierarchies of field names simply do not lend themselves to database-style manipulations other than those which can be simulated by "cut and paste" cell 1 1 1 1 1 1 1 11111 1 1 1 1 111 11111111111111 prosecuted exceeded a certain number, would require two of the twenty four columns to be identified, then summed, then compared to the given reference number. the first step, identifying the columns, can only be done by a person or by software capable of dealing with hierarchical field structures. for similar reasons, the task of sorting this particular data by number of persons prosecuted within age-group or sex would h^ daunting. in the wistat data collection the tendency to horizontal expansion of variables has an even more direct problem. many indirect users want to extract data for some or all subject areas for one country or region. country or region is typically the only vertically expanded variable, i.e. it identifies the rows. when the relevant row from each spreadsheet is identified and extracted it is not possible to combine these rows into a new spreadsheet or an integrated listing because each row has a different column structure. the result of extracting all data for one country is approximately 70 one-line spreadsheets. if titles and column headings are extracted as well, then we have approximately 70 spreadsheets each with a dozen or so rows. thus the most appropriate application for sprcadlassist quarterly sheets is the manipulation of cells and blocks of cells after the table has been formed in the desired way by tabulation or data management software. database theory and statistical data when people talk about a "database" they are usually referring to the type of organization of data which allows the retrieval and display of items of subgroups of the data by name or description rather than by reference to where or how they are stored. strictly speaking the software which carries out access and retrieval is an essential component of the database. a typical database query would be "list the names and addresses of all respondents who are hostile to interviewers". or fw aggregated data: "list the coimtries for which female constitute more than 50 per cent of the population". queries for an aggregate statistical database would also be expected to extract and format tables, either for printing or for expwt to spreadsheet files. the particular way of expressing the question depends on the query language defined by the database software in use. one type of database organization is known as "relational". this approach is very popular, having received excellent coverage in computer magazines, and many data management packages claim to be relational. the most wellknown aspect of relational data management is that its fundamental data structure is a "table". this makes it attractive to people who like to think of their data in tabular format the relational "table" is, however, very different from a statistical table, and in many ways the relational approach is not suited to statistical data. in fact the fundamental data structure is a relation. a relation is like a pattern. any type of entity for which data is to be held is given a relation ot pattern, which is a list of named place-marker. for example, if we are to hold data about staff we would define a staff relation specifying the items of data to be held: staff (id-number , name, department, position, date-hired, type of contract, etc.) one of these items, or a group of them, must function as a unique identifier, or "key". it is shown underlined here. the actual data consists of "instances" of the relation, e.g.: (65sss2,templeman,diesa„6 .tun 88,consultant,etc.) when a set of instances is listed it looks like a table. such a set may be stored in traditional computer terms as a file with a key field and a simple set of fields, sometimes known as a flat file. when data is to be stored about entities which have a specific relationship to each other, relational theory specifies the mechanism for relating them. for example. there may also be a department relation, e.g.: dept (dept-id.dept-name.locationjiame-ofhead,phone-of-head) there is a specific relationship between staff and departments. this is expressed by including the key of the department relation as an ordinary field of the staff relation. it is known as a "foreign key": staff (id-number,name,dept-idjocation,...) for data to function relationally it has to be set up specifically to do so. data about one entity should not be embedded within the relation for another entity, multiple field values are not permitted, hierarchical field structures are not permitted, and so on. when data comes naturally with such impurities of structure it must first be converted into a logic^y equivalent set of relations of the acceptable type. this conversion process is known as "nonnalization".q ^presented at the ifdo/iassist 89 conference held in jerusalem, israel, may 15-18, 1989. "the authors are, respectively. chief of the social and housing statistics section of the united nations statistical office and consultant to the statistical office on the united nations women's indicators and statistics data base (wistat). the views expressed are those of the authors' and not necessarily of the united nations. ^see towards a system of social and demograr)hic statistics. studies in the integration of social statistics: a technical reixjrt. improving social statistics in developing countries: technical report . studies in methods, series f, nos. 18, 24 and 25 (united nations pubucations, sales nos. e.74.xvii.8, e.79.xvii.4, e.79.xvii.12), and social indicators: preliminary guidelines and illustrative series . statistical papers, series m, no. 63 (united nations publications. sales no. e.78.xvii.8). 'series f, no. 49 (united nations publications, in press). ^series y, no. 3 (united nations publications. sales no. e.88.xvii.12). 'series k, no. 5 (united nations publications, in press), and series k, no. 6 (united nations publications, in press). ""compilation of selected statistics and indicators on social policy and development issues" (e/conf.80/ crp.l). spring 1990 iassvol201 19fall 1996 the archive the lijphart elections archive, housed at the university of california, san diego campus, is a research collection of districtlevel election results for twenty-seven countries. until 1994, the collection focused on post-world war ii democracies in western europe, but included the united states, canada, india, israel, japan, australia and new zealand. recently, costa rica and the european union were added. future plans call for expansion of the archive to more than 70 countries— including many new democracies from central and eastern europe, latin america, and africa. the collection includes all national elections for the lower house ( in some cases the only house) of the legislature. where the legislature is bicameral, the upper house is included only if it is directly elected by the voters. the volumes in the collection are the detailed, district-level election results that are usually published by government statistical offices in one or more volumes for each election. in some cases, non-governmental publications are included if they are the more complete source. the data includes the number of votes by party for each election district, including detailed lists for minor and major parties, and the number of seats or “elected candidates” by party for each election district. complete election data was sought for those countries where there is preferential voting (e.g., australia and ireland), where there is a first or second ballot (e.g., france), and where there is proportional representation (e.g., where one candidate may be elected from several districts). until 1994, the preferred format for building this library collection had been the original hard copy. during the last year, prompted by requests from graduate students, efforts have been made to acquire data in machine-readable form. historic origins the lijphart elections archive is named for arend lijphart, research professor of political science at the university of california, san diego, and the man responsible for its establishment. 2 prof. lijphart, is a world-renowned expert on elections. when he began researching comparative electoral systems back in 1984, no library anywhere collected all the detailed statistical data that he needed. after consulting with international colleagues as to the need for such an archive, professor lijphart and the university library agreed to create such a collection at ucsd. the task of finding materials that in many cases was already out of print, proved to be difficult at times. much of the success of the collection must be attributed to professor lijphart himself. many of the volumes in the collection were acquired and donated by him. it should also be noted here, that the collection was named in his honor by the university’s librarians. present and future the establishment of the internet and the emergence of many more democracies has infused a new energy into the lijphart archive. although adding machine-readable data to the collection was always a consideration, it wasn’t until a critical mass of technology, people and funding formed, that it became a reality. new people have now joined the team. jim jacobs as the university library’s data librarian and i, as the political sciences bibliographer along with professors gary cox and matthew shugart, of the political science department, hope to enlarge both the content and the access of what we see to be a worldwide resource for scholars. others apparently share our vision, since establishing the lijphart elections archive as a permanent internet web site has recently received funding from the national science foundation. one result of this initial funding has been the creation of a home page in march of 1995. 3 additional funding will be sought from nsf and other sources. enlarging the archive to include pre-1994 data has received some discussion. the possibility of scanning existing paper holdings is being considered. ocr reliability, the condition of the original materials, the different languages and type fonts and lastly, cost, are some of the concerns that will need to be dealt with. democratic elections on the internet: the lijphart elections archive by renata g. coates1 , university of california, san diego 20 iassist quarterly web site the current home page for the lijphart elections archive is still under construction and is part of ucsd’s social sciences data collection. presently, the archive’s holdings can be searched in a number of different ways: by material type (eg. paper holdings, vs. electronic holdings), by country, by electoral system and by keywords. the program can also show the library catalog entry for the item— a feature that provides information suitable for generating an interlibrary loan transaction. in addition to providing the actual data, the lijphart home page will link users to other election information on the internet. a current example, is slovakia. the lijphart archive never collected election data on slovakia, but since the information has been made available through eunet, we have added a linkage for the convenience of our researchers. conclusion the ucsd social sciences and humanities library is still actively acquiring election returns in hard copy. although our plans are to expand the archive’s holdings to all the worlds’ democracies, much of our success will depend on funding from outside the university. until we are assured of the stability of the internet site, we will probably be duplicating the data in hard copy—to maintain the integrity of the lijphart archive for future researchers. plans are underway to submit another proposal to the national science foundation. simultaneously, efforts are being pursued for international agreements with other agencies for sharing both information and more important, staff and data. creators of the home page envision it as a major resource for locating and reflecting other internet election resources. the on line version of the lijphart election archive however, will continue to have the same goal as the paper collection —to provide researchers with the actual data of elections down to the district level. we know this is an ambitious project, but we also have come to realize that promoting the results of democratic elections worldwide is a worthwhile effort for all of us. 1 submitted for the iassist conference held in quebec, canada. may 1995 2 prof. lijphart served as president of the american political science association during 1994-95 and is the author of electoral systems and party systems: a comparative study of twenty-seven democracies, 1945-1990 oxford: oxford university press, 1994. 3 lijphart elections archive url: vol281.indd 22 iassist quarterly spring 2004 iassist / ifdo 2005 edinburgh iassist/ifdo 2005 evidence and enlightenment 24-27 may 2005, edinburgh, scotland http://datalib.ed.ac.uk/iassist/ the 2005 iassist conference (international association for social science information service and technology) is being held in association with ifdo (international federation of data organisations) at the holyrood hotel in edinburgh, scotland from wednesday 25th to friday 27th may, hosted by edinburgh university data library & edina. the conference, which takes place in europe on a rotating basis, is relevant to data managers, creators, and users, as well as librarians and other information professionals who now have responsibility for providing support for the use of research data, especially in the social sciences and other observational disciplines. the conference will include plenary sessions, parallel sessions, a poster session and workshops. see http://datalib.ed.ac.uk/iassist/programme.shtml for the full programme and social events. the conference will be preceded by optional skills-building workshops on tuesday 24th may. (conference attendance not required see http://datalib.ed.ac.uk/iassist/workshops.shtml for descriptions.) a weekend in the scottish highlands will follow the conference for those wishing to add to their experience of scotland and network further with delegates. please register for the conference in advance, at http://datalib.ed.ac.uk/iassist/registration.shtml or email iassist05@ed.ac.uk with any questions. we look forward to welcoming delegates to edinburgh. iassist quarterly 35 topically-focused data archives: a new paradigm for the codification of social science research by josefina j. card president, sociometrics corporation 3191 cowper street palo alto, california 94306 tel. (415) 321-7846 the "information explosion" has become a distinguishing feature of modem science. both the published scientific hterature and its supporting data files continue to grow at unprecedented rates. more than ever, it has become important that efficient ways be found to store available information on a given topic and then retrieve relevant portions of that information as they are required. in the 1970's enormous progress was made in the development of procedures to store and retrieve bibliographic information. the dialog, eric, and medlars databases are but a small sample of the growing number of computerized bibliographic databases available to social scientists. less significant progress has been made in the development of analogous procedures to store, catalog, and retrieve elements common to the numeric (or raw data) information underlying the published studies. enormous productivity and cost savings could result from such development for little additional cost relative to the data collection costs already incurred, a substantively-focused data archive with indexing capabilities could: accelerate the growth and dissemination of scientific knowledge about a topic of contemporary interest; encourage corroboration and replication of newly reported findings; provide policymakers and practitioners with a larger scientific base on which to build their work; and stimulate investigations by new investigators without access to the substantial funds required for new data collection. this paper introduces the data archive on adolescent pregnancy and pregnancy prevention (daappp), to illustrate features of an emerging information resource: the special-purpose social science data archive. the accumulation of knowledge about himian reproduction, coupled with the development of relatively safe, eftective, and inexpensive contraceptive methods, has made it possible for human beings to seize control of their biological destinies, and to plan the size and spacing of their families. difterences continue to persist. spring 1986 36 iassist quarterly however, in the degree to which various groups of people have been able to avoid unplanned and unwanted pregnancies. rates of unplanned and unwanted pregnancies are higher in the developing than in the developed world. in a given country, young unmarrieds and the socially and economically disadvantaged have generajly been found to be more vulnerable. the rate of out-of-wedlock pregnancy and childbearing among u.s. teenagers is among the highest in the world. daappp was established by the u.s. office of population affairs of the office of the assistant secretary for health to encourage the conduct and dissemination of research on these important social issues. daappp identifies, selects, acquires, and archives the most valuable databases dealing with u.s. adolescent fertility and u.s. family planning. database identification refers to the systematic identification of all machine-readable data sets capable of addressing these topics. database selection refers to the selection, from the identified universe, of the most outstanding data sets to include in the collection. technical quality, substantive scope, and policy relevance are considered simultaneously by a national advisory panel of scientists in making selection decisions. database acquisition refers to obtaining selected data sets from their holders. the raw data, the documentation, and completed reports and publications are all acquired. database archiving refers to the processing and documentation of acquired data sets by archive staff, so that standardized, easy-to-use products are produced and disseminated. daappp then makes its data and documentation publicly available, for the cost of reproduction. the following products are now publicly available for each of the 45 data sets currently in daappp (see table 1): a computer tape for use with mainframe computers with two machine-readable files: •the raw data; •spss-x program statements to convert the raw data to an spss-x system file (spss is an acronym for the widely-used statistical package for the social sciences). floppy diskette(s) for use with microcomputers (in either 360-kilobyte or 1.2-megabyte format), with two machine-readable files: •the raw data; •spss/pc program statements to convert the raw data to an spss/pc microcomputer system file. a printed and boimd user's guide, with five standard sections: •an overview of the original purpose for which the data were collected, and a description of the file's processing history; •a description of the machine-readable files available for the data set; •a categorization of the variables included in the data set by their topic and type, followed by a listing of all variables, sorted by topic and type; •a report on the completeness and quality of the data; •a bibliography of representative publications based on the data sel the codebook and instrument from the original investigation, where available. users of statistical package programs know that part of the routine procedure in the development of system files for analysis is the assignment of names and labels to variables in the file. we use the output of this routine procedure—lists of variable names and values—in an innovative way to give the archive indexing capabilities. each variable in daappp is given an eight-character name for use with spss-x or spss/pc, to standardize variable names across all files, and to provide the user with quick reference to certain useful information about the spring 1986 iassist quarterly 37 variable. characters 1-2 encode the variable's topic, the main subject matter of the variable. character 3 encodes the variable's type, further classifying the subject matter into one of the many variable types commonly used by social scientists. characters 4-5 are a reference to the data set id, indicating the original source of the data. characters 6-8 contain the variable sequence number, indicating the sequential position of the variable within the source data set table 2 contains the list of topics, table 3 the list of types. definitions of each topic and type have been developed that provide an inter-rater categorization reliability of over 90%. the list of data set ids is shown in table 1. each list can be altered easily to suit archives focused on other substantive topics. the variable naming scheme encodes information both on what each variable has in common with other variables in the archive (its topic and type), and what is unique to the variable (its sequence number writh a given data set id). daappp staff members have written a simple computer program that uses the topic and type characters of the variable names as input to produce a matrix that depicts, at a glance, the topical emphasis of each data set for example, daappp data set no. 2 is the 1976 u. s. national survey of young women (john kantner and melvin zelnik, john hopkins university, principal investigators). table 4 contains the topic^by-type matrix for this data set the matrix allows the user to see at a glance where the "areas of richness" of the data set lie. for example, we can see in the last column of table 4 that this particular data set has a total of 386 variables; the data set is rich in information on family characteristics (142 variables), contraceptive information (56 variables), and child-bearing related information (42 variables). there are seven items relating to abortion, the first topic in the alphabetically-ordered topic lisl the first row of table 4 shows that all seven of these variables are attitudinal. information of the type contained in tables 4 and 5 can be extremely helpful in ascertaining whether a given data set can be used to answer a particular research question. it is important to note that virtually no extra processing time, beyond the routine procedures used by any social scientist to create an spss-like set-up for his data set. is required in order to produce and display the information. when variable names and labels from all the data sets in daappp are used as input, the same program provides a matrix and variable listing that depicts the state-of-the-archive. the 45 data sets currently in the daappp collection contain 14,216 variables, characterized as shown in table 6. while a listing of these 14,216 variables is too long to print here, such a list exists, is publicly available (for the cost of reproduction), and is updated quarterly. although the daappp project is by no means over (the collection is currently growing at the rate of about five data sets per quarter), we see from table 6 that there appears to be relatively little empirical data on important topics such as sexually transmitted disease and substance abuse (in the context of adolescent pregnancy studies). at the end of the daappp contractual period in september 1987, we will be in a position to ' evaluate the amoimt and types of information available on adolescent pregnancy, pregnancy prevention, and family planning, and to identify significant gaps. social science archival data can be used in many different ways: (1) for secondary analysis (the analysis of data for purposes other than those for which the information was originally collected); (2) for meta analysis (the analysis of data common to a number of data sets to investigate similarities and differences in the patterning of relationships); (3) for longitudinal analysis of panel data (such as that found in daappp data sets 20-24, the national spring 1986 38 iassist quarterly longitudinal survey of youth); (4) for cross-sectional trend analysis of related surveys (such as analysis of trends in information found in daappp data sets 11-18. the 1977, 1980, and 1982 current population surveys); (5) for provision of contextual variables to add to an individual-level data file (for example, one could add all or part of the information contained in daappp data set 8 on state policy determinants of teenage childbearing to one's individual-level data file to study the additive and interactive effects of individual versus environmental factors in producing fertility-related behavior); (6) for derivation of comparison group data against which to compare data from clinic patients or service program participants; and (7) for instructional purposes, as an exciting aid in the teaching of statistics and research design. acknowledgements the data archive on adolescent pregnancy and pregnancy prevention is funded by contract 282-84-0083 between the office of population affairs, office of the assistant seaetary for health, and sociometrics corporation (j. j. card). it is our hope that those interested in studying problems of adolescent pregnancy and family plaiming will use daappp, and that the daappp experience will be helpful in stimulating and facilitating the formation of other, special-purpose data archives containing the best scientific data on important issues facing us all. spring 1986 iassisl quarterly — 39 table 1 table 1 list of data sets currently in daappp data set id data set name (investigators) 01 1971 u.s. national survey of young women: selected variables (m. zelnik & j.f. kantner) 02 1976 u.s. national survey of young women (j.f. kantner &. m. zelnik) 03 project talent: consequences of adolescent childbearing for the young parents' future life, 1960-1974 (j.j. card) 04 detroit mother-daughter communication patterns: mother file, 1978 (g.l. fox) 05 detroit mother-daughter communication patterns: daughter file, 1978 (g.l. fox) 06 philadelphia collaborative perinatal project: economic, social, and psychological consequences of adolescent childbearing. 1959-1965 (j. marecek) 07 nashville general hospital comprehensive child care project, 1974-1976: selected variables {h>m> sandler) 08 state policy determinants of teenage childbearing. 1979 (k.a. moore) 09 1980 u.s. survey of services provided by adolescent pregnancy programs (jrb associates) 10 1982 evaluation of dapp adolescent pregnancy programs (m. burt) 11 1980 u.s. current population survey: selected variables — women (bureau of the census) 12 1980 u.s. current population survey: selected variables — men (bureau of the census) 13 1980 u.s. current population survey: selected variables — children (bureau of the census) 14 1982 u.s. current population survey: selected variables — women (bureau of the census) 15 1982 u.s. current population survey: selected variables — men(bureau of the census) 16 1982 u.s. current population survey: selected variables — children bureau of the census) 17 1977 u.s. current population survey: selected variables — women (bureau of the census) 18 1977 u.s. current population survey: selected variables — men (bureau of the census) 19 first u.s. health and nutrition examination sruvey (hanes). 1971-1975 (national center for health statistics) 20-24 national longitudinal study of youth (nlsy). 1979-1982: selected variables (waves 1-4). and supplementary variables (ohio state university) 24 1981 u.s. survey of title x funded family planning clinics (r. herceg-baron) 26 1982 national survey of family grovi/th (nsfg). cycle iii — women aged 15-44 (national center for health statistics) 27 1982 national survey of family growth (nsfg). cycle iii — women aged 15-44 (national center for health statistics) 28 1979-1980 u.s. survey of unmarried women under 18 in family planning clinics (a. torres) 29 effects of organized family planning programs on u.s. adolescent fertility (j.d. forrest) 30 johns hopkins study of repeat adorescent pregnancy. 1976-1982 (j.b. hardy) 31 1972-74 ventura county of unmarried pregnant women aged 13-20 (m. eisen) 32 1982 san jose, california study of adolescent perinatal risk behavior (p. a. hensleigh &. n. moss) 33 1981-1982 evaluation of dapp adolescent pregnancy programs: individual level data 1 (m.r. burt) 34 1981-1982 evaluation of dapp adolescent pregnancy programs: individual level data 1 (mr. burt) 35 1979-1981 philadelphia study of psychological factors associated with adolescent fertility regulation — females (e.w. flaherty & j. marecek) 36 1979-1981 philadelphia study of psychological factors associated with adolescent fertility regulation — males (_e.w. flaherty & j. marecek) the national survey of children, 1976 (child trends. inc.) 39 florida-puerto rico study of adolescent pregnancy and neonatal behavior, 1978 (b.m. lester) 40 maricopa county, arizona study of child maltreatment risk among adolescent mothers, 1976-1978 (f.g. bolton, jrj 41 1955 growth of american families: married women (a. campbell, p. k. whelpton, & j.e. patterson) 42 1955 growth of american families: single women (a. campbell. p. k. whelpton, & j.e. patterson) 43 1960 growth of american families (a. campbell, p. k. whelpton, i j.e. patterson) 44 1979 u.s. national survey of young women (m. zelnik &. j.f. kantner) 45 1979 u.s. national survey of young men (m. zelnik &. j.f. kantner) 37-38 spring 1986 ^ iassist quarterly table 2 jgjjjg ' list of topics and their two-letter codes ab abortion mp ac agency charactermh ad adoption me ag age nu bf biological function oc cb childbearing ot cr childrearing ow cl clinical activities pe cm communication ra cn contraception rc ci crime rl ed education rs fh family and household se fs friends and social sx gr gender and gender sd go guidance and sa hl health un in intellectual function wf iv interview table 3 marriage patterns mental health istics meta level nutrition occupation and development other out-of-wedlock parenthood personality race/ethnicity recreation religion residence /location sex education characteristics sexuality activities sexually transmitted role disease substance abuse counseling undocumented wealth, finances, and material things list of types and their one-letter codes a attitudes b behavior c cognitions e emotions h history i intentions m motivations other p program/policy r reasons s status t traits u undocumented x meta y aggregate z household spring 1986 iassisl quarterly 41 table 4 overview of contents. the 1976 national survey of young women /icsicnchit olhietcst i ihiemtiqihortvi i ins iciis t^ut^jiyt 1 71 ai 01 01 01 01 ol ol 1 ii 1 2 1 lo 1 1 ii 1 t 1 o i « 1 c0rjil.7ac 1 31 01 ol ol 01 01 01 tl ol c:i.rs>:£?r 1 o 1 o i o 1 <.; 1 o 1 o 1 o 1 i2 1 t 1 el chl.d-f^ 1 ll ol ol 01 01 ol ol 11 ol ol £gucii:cll 1 11 1 1 1 3 1 1 1 '• 1 11 1 finllt c^a9 1 2 1 ii ii 25 1 2 1 1^ 1 ' 1 111 i 1 f^if.r.5 1 2 1 2 1 ii 1 1 1 1 1 1 a 1 ' 1 r^ta .1 1 1 1 1 9 1 1 1 1 iii r.ji-iliie 1 ii ol 01 i*! 21 01 ol si tl ol occjpkricl. 1 21 ii 01 01 01 21 01 ll ]| 01 onie^ 1 1 1 1 1 ii 1 ii 4 1 1 1 ojl ksljlcck 1 21 0! 51 01 ol ol 01 ol 01 01 pire e'hi 01 01 01 01 ol ol oi ol il ol e£llc:c>r 1 ii ll 01 ol 01 01 0! 01 21 01 i!-:::e ice 1 t 1 o i o i o 1 o i o i o i o i s 1 o 1 i;_.u:'r l 31 21 oi i".! ol ol ol oi 01 oi i is |£utu3 iheta i i spring 1986 42 (assist quarterly table 5 list of variables. by topic. national survey of young women page 1 of 10 -topic;abortion newid abad2110 abad2110 abad2112 abad2113 abad2n4 abad2115 abad2116 label abort ok if the woman had been raped abort ok for very young person abort ok if pc endang womans health abort ok if child born defrmd or mently defec abort ok if the woman couldnt afford it abort ok any reason important to her views about having an abortion -topic:abortion adad2099 adad2100 attitudes attitudes if unable have wanted chidrn. wld adopt'' would r adopt chld instead of having own topic:agenewid agh02021 agh02022 ags02003 ags02061 label yr of birth month of birth age screen—age topicibiol functnewid bfh02132 age list period topic :childrenbearnewid cbad2101 cbad2133 cbad2135 cbho2006 cbho2230 cbh02232 cbh02234 cbh02237 cbh02238 cbh02239 cbho2240 cbh02242 cbh02243 type label attitutes ideal age for a girl to have 1st baby cognitions know when preg is most likely to occur cognitions when preg is most likely to occur history pregnancy status at marriage ind history ever been pregnant history number of previous pregnancies history outcome of 1st pgat marriage ind history what first pregnancy think good chance history yr of outcome 1st pg history month of outcome 1st pg history age at outcome 1st pg history outcome 2nd pg history what 2nd pregnancy think good change spring 1986 iassist quarterly 43 table 6 the state-of-the-archive (9/85) topic type |limtvdtlteh»»io«|ax»itio|diort:ohs|histopi |iir:tl(tio|hfftiv«il|l |5 |s ins i i \»i |(1« i iim iwaions |5t«tus \m:n |ui.-axrjkt|krt» |a3cpti»t|z i i 1 |bt£3 -j-h spring 1986 iassist quarterly 2013 51 iassist quarterly ddi timeline by mary vardigan1 abstract this timeline lists key developments and events in the history of the ddi, beginning with selected foundational developments that set the stage for ddi and facilitated its creation. note that while many archives have played a significant role in ddi history, only the ones established early on are listed here. 1946 roper center established. 1960 zentralarchiv established. 1962 inter-university consortium for political [and social] research (icpsr) established. 1964 steinmetz archive established. 1965-1968 marc (machine readable cataloging) standards developed at [u.s.] library of congress. 1967 united kingdom data archive (ukda) established. 1967 osiris (organized set of integrated routines for investigations with statistics) statistical software package developed at university of michigan, institute for social research. 1968 first version of spss (statistical package for the social sciences) released. 1970 american library association appoints subcommittee to recommend rules for cataloging machine-readable data files. 1971 norwegian social science data service (nsd) established. 1974 international association for social science information services and technology (iassist) formed following the conference on data archives and program library services, held in conjunction with the 8th world conference of sociology in toronto, canada. 1974 report issued on “standardization of study description schemes and classification of indicators,” a product of a meeting held in copenhagen, denmark, at the danish data archives with nine archives from six countries in attendance. 1975 iassist establishes classification action group. 1976 council of european social science data archives (cessda) established. 1978 anglo-american cataloguing rules, second edition (aacr2) published by the american library association (includes a new chapter 9, for machine-readable data files). 1979 sue dodd publishes “bibliographic references for numeric social science data files: suggested guidelines.” journal of the american society for information science, volume 30, issue 2, pp. 77–82, march 1979. 1980 a style manual for machine-readable data files and their documentation published by r.c. roistacher with contributions from sue dodd, b.b. noble, and alice robbin. 1982 sue dodd publishes cataloging machine-readable data files. an interpretive manual. american library association, chicago. selected foundational developments that set the stage for ddi 52 iassist quarterly 2013 iassist quarterly 1985 sue dodd and ann m. sandberg-fox publish cataloging microcomputer files: a manual of interpretation for aacr 2. 1985 fienberg et al. publish sharing research data. committee on national statistics. commission on behavioral and social sciences and education. national research council. washington, dc: national academies press, 1985. 1985 nsfnet, the precursor to the internet, created. 1986 sgml (standard generalized markup language), descended from ibm’s generalized markup language (gml) developed in the 1960s, becomes an iso standard. 1989 world wide web invented. 1989 international federation of library associations (ifla) recommends an international standard bibliographic description for computer files (isbd/cf). 1993 iassist forms action group for “codebook documentation of social science data.” 1993 cessda holds seminar on “variable level documentation” in gothenburg, sweden. 1994 dublin core metadata initiative established. 1995 first sgml codebook committee, constituted by icpsr director richard rockwell, meets in quebec city, canada. members develop a draft list of codebook elements. 1996 first ddi specification prepared at university of michigan library. an sgml document type definition (dtd) was produced by david barber and john brandt (university of michigan), ann green (yale university), and the ddi committee. 1996 first working draft of xml (extensible markup language) published. 1997 u.s. national science foundation (nsf) funding received to enhance ddi and betatest it. this award (sbr-9617813) funded ddi development and icpsr codebook digitization. final report (pdf ) 1997 sgml ddi specification translated to xml. this work was done by jan nielsen, danish data archive. 1998 committee meets in new haven, ct. prepares for betatesting. 1999 betatest of ddi dtd takes place (see related article for testers’ names). 2000 ddi version 1 (dtd-based) published. view version 1 and successive iterations 2001 formal ddi evaluation, funded by nsf, takes place. evaluators praise the effort evaluation report 2001 funding for ddi development received from health canada during 2001-2002. covers costs of meetings until alliance established 2001 working group on aggregate data meets in voorburg, netherlands. group develops a proposal for ddi coverage of aggregate/tabular data 2001 first ddi training held at iassist in amsterdam, netherlands. bill block and wendy thomas lead the workshop on “creating ddi compliant codebooks” 2002 bjorn henrichsen, nsd, becomes chair of the ddi committee. 2002 committee meets in storrs, ct, to draft ddi alliance charter. view the charter 2003 final meeting of original ddi committee held in washington, dc (see related article for members’ names). minutes 2003 ddi alliance established with tom piazza, uc-berkeley, chair. view “about the specification” from original alliance web site, march 2007 2003 ddi alliance steering committee meets for the first time. minutes 2003 ddi 2 published. ddi now covers aggregate data and geography view version 2 and successive iterations dtd version history http://www.ddialliance.org/sites/default/files/final.pdf http://ddi-alliance.cvs.sourceforge.net/viewvc/ddi-alliance/ddi/dtd/version1.dtd?view=log http://www.ddialliance.org/sites/default/files/evalsummary.pdf http://www.iassistdata.org/downloads/ddi_workshop.pdf http://www.iassistdata.org/downloads/ddi_workshop.pdf http://web.archive.org/web/20120405100126/http:/www.ddialliance.org/alliance/charter http://www.ddialliance.org/ddi/committee-info/minutes/2003-02-07.html http://web.archive.org/web/20070307225957/www.ddialliance.org/codebook/index.html http://www.ddialliance.org/ddi/committee-info/minutes/2003-02-08.html http://ddi-alliance.cvs.sourceforge.net/ddi-alliance/ddi/dtd/ http://web.archive.org/web/20060711204401/http:/www.icpsr.umich.edu/ddi/dtd/version-history.html iassist quarterly 2013 53 iassist quarterly 2003 ddi expert committee meets for first time in ann arbor, mi. committee discusses transition from dtd to schemas; new working groups on structural reform and substantive issues formed minutes 2004 ddi expert committee meets in madison, wi. committee discusses requirements for version 3 minutes lifecycle model 2005 hans jorgen marker, dda, becomes chair of the ddi expert committee ; ron nakao, stanford, vice chair. 2005 ddi expert committee meets in edinburgh, scotland. committee ratifies life cycle model and ddi 3 begins to take shape minutes 2006 ddi expert committee meets in ann arbor, mi. committee approves the scope and timeline for version 3 minutes 2007 public review of ddi 3 takes place. 2007 ddi expert committee meets in montreal, canada. committee approves candidate draft of ddi 3 minutes 2007 first ddi training workshop takes place at schloss dagstuhl in wadern, germany. course description 2008 ddi 3 published as xml schemas. version 3.0 2008 expert committee meets at iassist in palo alto, ca. plans for tools development discussed minutes 2008 ddi develops best practices for ddi 3. best practices 2009 ddi lifecycle 3.1 published. link to specification 2009 first european ddi users conference (eddi) held in bonn, germany. program (pdf) 2009 ddi expert committee meets in tampere, finland. committee discusses tools and outreach to nsis minutes (pdf ) 2010 chuck humphrey, university of alberta, becomes chair of the ddi expert committee; mari kleemola, fsd, vice chair. 2010 second eddi conference held in utrecht, netherlands. program (pdf ) 2010 ddi developers group meets for first time in utrecht, netherlands. view minutes and reports 2010 ddi expert committee meets in ithaca, ny. rebranding ddi 2 and 3 as ddi codebook and lifecycle approved minutes (pdf ) 2011 external review of ddi takes place. review report 2011 ddi alliance publishes first set of controlled vocabularies. view vocabularies 2011 expert committee meets in vancouver, canada. results of external review discussed minutes (pdf ) 2011 third eddi conference held in gothenburg, sweden. program (pdf ) 2011 ddi alliance releases tools catalog. view catalog 2011 ddi alliance establishes agency registry. view registry 2012 ddi expert committee meets in washington, dc. plans for a model-based specification discussed minutes (pdf ) 2012 ddi codebook 2.5 published as xml schemas. specification 2012 first dagstuhl workshop on model-based ddi held. description paper http://www.ddialliance.org/ddi/committee-info/minutes/2003-10-12.html http://www.ddialliance.org/ddi/committee-info/minutes/2004-05-29.html http://www.ddialliance.org/system/files/concept-model-wd.pdf http://www.ddialliance.org/ddi/committee-info/minutes/2005-05-22.html http://www.ddialliance.org/ddi/committee-info/minutes/2006-05-27.html http://www.ddialliance.org/ddi/committee-info/minutes/2007-05-19.html http://www.dagstuhl.de/en/program/calendar/evhp/?semnr=07432 http://downloads.sourceforge.net/ddi-alliance/ddi_3_0_2008-04-28_documentation_xmlschema.zip?modtime=1209367396 http://www.ddialliance.org/ddi/committee-info/minutes/2008-05-31.html http://www.ddialliance.org/resources/publications/working/bestpractices http://www.ddialliance.org/specification/ddi-lifecycle/3.1/ http://www.iza.org/conference_files/eddi09/2011_01_17_online_program.pdf http://www.ddialliance.org/sites/default/files/minutes/2009-05-25.pdf http://www.iza.org/conference_files/eddi10/program_2010-12-04_jw.pdf http://www.ddialliance.org/alliance/working-groups#ddc http://www.ddialliance.org/sites/default/files/ddi alliance expert committee meeting minutes(2) 5-31-2010.pdf http://www.ddialliance.org/sites/default/files/ddi alliance review summary 2011-07-08.pdf http://www.ddialliance.org/controlled-vocabularies http://www.ddialliance.org/sites/default/files/ddiallianceexpertcommitteemeetingminutes2011-05-30_0.pdf http://www.iza.org/conference_files/eddi2011/call_for_papers/eddi11_program_2011-11-26.pdf http://www.ddialliance.org/resources/tools http://registry.ddialliance.org/ http://www.ddialliance.org/sites/default/files/ddiallianceexpertcommitteemeetingminutes2012-06-04.pdf http://www.ddialliance.org/specification/ddi-codebook/2.5/ https://www.dagstuhl.de/en/program/calendar/evhp/?semnr=12432 http://dx.doi.org/10.3886/ddiworkingpaper04 54 iassist quarterly 2013 iassist quarterly 2012 fourth eddi conference held in bergen, norway. program 2013 ddi rdf discovery and xkos vocabularies published. description 2013 first north american ddi users conference (naddi) held in lawrence, ks. program 2013 ddi membership/scientific board meets in cologne, germany. transition to new governance structure discussed minutes 2013 first ddi executive board (successor to steering committee) meets with gillian nicoll, australian bureau of statistics, as chair and ron nakao, stanford university, as vice chair. view members and minutes 2013 ddi “sprints” launched to work on model-based ddi. more information 2013 fifth eddi conference held in paris, france. program 2014 ddi lifecycle version 3.2 published. references “about the [ddi] specification.” ddi alliance web site (via the internet archive wayback machine): (accessed february 12, 2014) block, bill, wendy thomas, robert wozniak, and joshua buysse. (2001) creating ddi compliant codebooks: breckenhill, inc. (2011) ddi alliance external review: summary and recommendations july 8, 2011. ddi alliance agency registry. (accessed february 12, 2014) ddi alliance original charter (via the internet archive wayback machine): (accessed february 12, 2014) ddi best practices for ddi 3. (accessed february 12, 2014) ddi [codebook] dtd version 1 and successive iterations. (accessed february 12, 2014) ddi [codebook] dtd version 2 and successive iterations: (accessed february 12, 2014) ddi [codebook] dtd version history (via the internet archive wayback machine): (accessed february 12, 2014) ddi codebook version 2.5 (xml schemas). ddi controlled vocabularies. (accessed february 12, 2014) ddi lifecyle model. (2004) (accessed february 12, 2014) ddi lifecycle version 3. (accessed february 12, 2014) ddi lifecycle version 3.1. ddi tools catalog: (accessed february 12, 2014) ddi rdf discovery and xkos vocabularies. (accessed february 12, 2014) european ddi users conference (eddi) (bonn, germany), 2009. (accessed february 12, 2014) european ddi users conference (eddi) (utrecht, netherlands), 2010. (accessed february 12, 2014) european ddi users conference (eddi) (gothenburg, sweden), 2011. (accessed february 12, 2014) european ddi users conference (eddi) (bergen, norway), 2012. (accessed february 12, 2014) european ddi users conference (eddi) (paris, france), 2013. (accessed february 12, 2014) inter-university consortium for political and social research (icpsr). (2000) final report to national science foundation “electronic preservation of data documentation: complementary sgml and image capture” sbr-9617813 [august 1 1997-july 31, 2000] inter-university consortium for political and social research (icpsr). (2001) addendum to nsf final report “electronic preservation of data documentation: complementary sgml and image capture,” sbr-9617813: results of the evaluation of the data documentation initiative (ddi), april 24, 2001. minutes of final meeting of original ddi committee (washington, dc), 2003: (accessed february 12, 2014) (accessed february 12, 2014) minutes of first meeting of ddi alliance expert committee (ann arbor, mi), 2003. (accessed february 12, 2014) minutes of ddi expert committee meeting (madison, wi), 2004. (accessed february 12, 2014) http://www.eddi-conferences.eu/ocs/public/conferences/1/schedconfs/1/program-en_us.pdf http://www.ddialliance.org/specification/rdf http://www.ipsr.ku.edu/naddi/program.shtml http://www.ddialliance.org/system/files/ddiallianceannual meetingofmembersandscientificboardminutes2013-05-27.pdf http://www.ddialliance.org/alliance/structure#steering http://www.ddialliance.org/alliance/minutes http://www.ddialliance.org/ddi-moving-forward-process http://www.eddi-conferences.eu/ocs/index.php/eddi/eddi13/schedconf/program iassist quarterly 2013 55 iassist quarterly minutes of ddi expert committee meeting (edinburgh, scotland), 2005. (accessed february 12, 2014) minutes of ddi expert committee meeting (ann arbor, mi), 2006. (accessed february 12, 2014) minutes of ddi expert committee meeting (montreal, canada), 2007. (accessed february 12, 2014) minutes of ddi expert committee meeting (palo alto, ca), 2008. (accessed february 12, 2014) minutes of ddi expert committee meeting (tampere, finland), 2009. (accessed february 12, 2014) minutes of ddi expert committee meeting (ithaca, ny), 2010. (accessed february 12, 2014) minutes of ddi expert committee meeting (vancouver, canada), 2011. (accessed february 12, 2014) minutes of ddi expert committee meeting (washington, dc), 2012. (accessed february 12, 2014) minutes of ddi membership/scientific board (cologne, germany), 2013. (accessed february 12, 2014) minutes of ddi executive board. (accessed february 12, 2014) minutes of ddi steering committee. (accessed february 12, 2014) minutes and reports of ddi developers group. (accessed february 12, 2014) nielsen, per (1974). “report on standardization of study description schemes and classification of indicators.” association for computing machinery (acm) sigsoc bulletin, volume 6, issue 2-3, fall-winter 1974-1975, pages 39-46. north american ddi users conference (naddi) (lawrence, ks), 2013. (accessed february 12, 2014) participants in 2012 dagstuhl seminar on ddi moving forward. (2012) “developing a model-driven ddi specification.” ddi working paper series, paper no. 4. schloss dagstuhl workshop (first dagstuhl training) (wadern, germany), 2007. http://www.dagstuhl.de/en/program/calendar/ evhp/?semnr=07432 (accessed february 12, 2014) schloss dagstuhl workshop on ddi model (wadern, germany), 2013. (accessed february 12, 2014) “sprints” to develop model-based ddi. (accessed february 12, 2014) notes 1. mary vardigan is an assistant director at the inter-university consortium for political and social research (icpsr) and director of the ddi alliance. she can be reached at vardigan@umich.edu. 2. members of original sgml codebook committee: • merrill shanks, uc-berkeley, chair • richard rockwell, icpsr • atle alvheim, nsd • martin appel, census • david barber, michigan • grant blank, university of chicago • bill bradley, health canada • pat doyle, ahcpr • terry finnegan, national center for supercomputing applications, university of illinois • peter granda, icpsr • ann green, yale • stephan greene, maryland • lynn jacobsen, columbia • john price-wilkin, michigan • karsten rasmussen, dda • rolf uher, zentralarchiv, koeln • mary vardigan, icpsr 3. full urls for all material linked in the time line can be found in references. 56 iassist quarterly 2013 iassist quarterly keywords from vol 37 n0. 1-4. courtesy tagxedo.com vol21.2 28 iassist quarterly abstract roads (resource organisation and discovery in subject-based services) is a uk higher education funded project to design and implement a user oriented resource discovery system. the project is investigating the creation, collection and distribution of resource descriptions to provide a transparent means of searching for, and using resources on the internet. the system is being piloted on a number of internet subject gateways, namely adam (art, design, architecture and media), biz/ed (business education on the internet), ihr-info (institute of historical research), omni (organising medical networked information) and sosig (social science information gateway). the paper will discuss the background to the project, the type of metadata being collected by the subjectbased gateways and the possibilities of cross searching distributed databases, with specific references to sosig. the project uses a standard template for recording information about resources (this was originally based on the internet anonymous ftp archive (iafa) template) which is a simple text based record using attribute-value pairs. the simplicity of the roads format also provides possibilities of mapping to and exchanging data with other metadata formats, for example the dublin core or other standards. one of the aims of the subject based gateways is to encourage information providers to become involved in the creation of records about their own data in order to make their information as useful and accessible as possible; an approach to this will be discussed. background roads1 (resource organisation and discovery in subject-based services) is a collaborative project funded by the electronic libraries programme2 (elib) in the uk, to design and implement a user oriented resource discovery system. the roads partners are: n ilrt (institute for learning and research technology) at the university of bristol responsible for user liaison and project management n loughborough university (department of computer science) responsible for the software development n ukoln (office of library and information networking) at the university of bath responsible for co-ordinating metadata requirements and issues roads has been created to provide a set of software tools and standards for building and maintaining catalogues of internet resources. the system allows resources to be catalogued and indexed and provides a searchable and browsable interface to the resource descriptions. the roads system is primarily being piloted on a number of elib funded subject information gateways (under the access to network resources (anr) programme) who feed into the development of the software. the gateways using roads are: n adam (art, design, architecture and media)3 n biz/ed (business education on the internet)4 n ihr-info (institute of historical research)5 n omni (organising medical networked information)6 n sosig (social science information gateway)7 the system is being used or evaluated by a number of other projects in the uk and interest in its use has also been expressed outside the uk. in addition, roads (along with sosig) is involved in the eu funded desire8 project (under the fourth framework telematics programme). each of the elib anr gateways is building a subject specific catalogue of internet resource descriptions. the roads software attempts to be as modular as possible to allow the gateways to configure the ‘look and feel’ of the services; for example, in the way they present browsable listings of resources and search results. it also allows the gateways to pick and choose parts of the system appropriate to their service and ‘plug in’ their own applications; sosig plans to use this capability to add a thesaurus tool to the standard roads search facility. each service may also differ on issues of selection policy, classification, etc. however, the gateways share a common metadata format for collecting information and end-users roads to metadata by debra hiom* summer 1997 29 will ultimately be able to cross-search the gateways (using the whois++ directory technology). roads templates in addition to software development, the roads project is concerned with issues of metadata. the elib anr gateways are creating records for selected quality resources on the internet using a standard template. the metadata format used by roads is based on the internet anonymous ftp archive (iafa) template definitions9; which as the name suggests were originally designed to describe resources available through ftp sites. with the growth of the web, these templates were extended to cover internet resources in general. the templates have been extended further by roads based on the implementation experiences of the subject gateways (therefore the templates will be referred to throughout the paper as roads templates rather than iafa). the template format was originally designed to be created by site administrators and therefore the emphasis is on simplicity and ease of creation. this simplicity also means that information skills are not essential and subject gateways can use the expertise of subject specialists as well library and information professionals to build their catalogues. the template is a text-based record composed of a series of attribute-value pairs that describe a resource content, format and location. a number of different resource description template types exist: n document n image n mailarchive n organization n project n service n software n sound n usenet n user example of search results in sosig 30 iassist quarterly n video within a roads template, there are three kinds of attribute; plain, variant and cluster. plain attributes describe the basic characteristics of a resource, such as title, description and keywords. they contain information about a resource that is only required once. variant attributes are repeated for multiple versions of a resource. examples of variant attributes are language and uri; if a document is available in english and french two sets of variant attributes would be used to record the language and url of each version. other examples of variant attributes include format and size of the resource. cluster attributes record information that may be common to a number of resources, for example name, address and email details of individuals or organisations. the original iafa templates have been extended slightly based on requirements from the subject gateways. for example, the attributes subject-descriptor and subjectdescriptor-scheme have been added. these attributes allow the resources to be classified using an appropriate classification scheme and this information is used to form the basis of browsable listings on the gateways. a range of administrative attributes was also added. the subject gateways can choose which template types and attributes to use according to the requirements of their endusers. however, a minimum set of attributes exists to ensure a level of interoperability between the gateways. the core attributes are: title, description, keywords, uri, subject-descriptor and subject-descriptor-scheme. this does not include any of the administrative attributes such as the record creation date as these are generated automatically. individual gateways may choose to make other attributes mandatory according to the requirements of their user community. for example, sosig is currently extending its coverage of european resources and has made language and country mandatory attributes in order to support this. a registry of the templates10 is maintained at ukoln. example of browsable listing in sosig summer 1997 31 this allows roads gateways to register the need for new template types or new attributes within existing templates if or when required. the registry will also contain some basic cataloguing rules to assist with interoperability. example of a roads template template-type: service handle: sosig472 category: database title: ibss online alternative-title: international bibliography of the social sciences uri-v1: telnet://bids.ac.uk uri-v2: http://www.bids.ac.uk/ibss admin-handle-v1: admin-name-v1: admin-work-postal-v1: admin-country-v1: uk admin-work-phone-v1: +44 (0)1225-826074 admin-work-fax-v1: admin-job-title-v1: admin-department-v1: admin-email-v1: bidshelp@bids.ac.uk publisher-handle-v1: publisher-name-v1: bath information & data services publisher-type-v1: publisher-work-postal-v1: university of bath, bath, ba2 7ay publisher-country-v1: uk publisher-work-phone-v1: publisher-work-fax-v1: publisher-email-v1: description: ibss online provides electronic access to the database of the international bibliography of the social sciences. it contains the bibliographic details of journal articles, book reviews books, and the chapters from selected multi-authored monographs. the database contains over 680,000 records covering publications appearing between 1981 and the present day, and is growing at the rate of approximately 100,000 items per annum. subject coverage is based on the four principal disciplines of anthropology, economics, political science and sociology, but it also reflects the interdisciplinary nature of the social sciences. material can be found which covers, for example, agriculture, archaeology, business studies, criminology, education, environmental issues, history, law, social policy, social work, and statistical methods. there is extensive coverage of international material. records come from over 100 countries, and 95 different languages are represented in the database. the database is mounted at bath information & data services (bids). all members of uk higher education institutions (heis) funded by the hefcs are eligible to use ibss online free at the point of use. users must register with their own institution’s library. keywords: social science, sociology, politics, economics, anthropology authentication: access to most databases is by username and password. registration: users are required to register with a representative at their own he institution. access-policy: the ibss online data may only be used by an employee, student or other person authorised by the institution or organisation which has taken out a licence to use the service. access-times: all bids services are normally available 24 hours a day, 7 days a week. copyright: subject-descriptor-v1: 3,301,32,33,572 subject-descriptor-scheme-v1: udc language-v1: en language-v2: en issn: source: to-be-reviewed-date: record-last-verified-email: record-last-verified-date: destination: uk,world record-last-modified-date: mon, 28 apr 1997 17:21:46 +0000 record-last-modified-email: ecdh@aubergine.ilrt.bris.ac.uk record-created-date: wed, 15 jun 1995 13:22:00 +0000 record-created-email: ecdh@ssa.bris.ac.uk mapping roads templates one of the main project objectives of roads is to ‘implement and test emerging standards and to improve uk participation in international standards making activity’11. to this end the project closely monitors metadata developments and is very active in metadata standards initiatives, in particular the dublin core and warwick framework. mapping roads templates to whois++ templates has been done as part of the planned developments for distributed searching of roads databases (these are actually very similar in format to the roads templates). in addition ukoln has produced several textual mappings from the roads templates to other metadata formats such as usmarc, dublin core, soif and the z39.50 bib-1 attribute set12. the templates map reasonably well on to these other formats although there may be some difficulties with syntax. it would seem fair to assume that there will continue to be several metadata formats in use to describe internet resources. however, roads has committed to providing telnet://bids.ac.uk http://www.bids.ac.uk/ibss mailto:bidshelp@bids.ac.uk 32 iassist quarterly conversion tools if they are required by the elib gateways and some proof of concept work has already been carried out converting the roads templates into usmarc and other formats. as part of another project, an experimental z39.50/whois++ gateway has been built which allows users to search a roads database in parallel with a range of z39.50 databases13. roads developments roads is presently in version 1 of its development cycle; this provides all the tools and software to build and maintain a gateway of internet resources. roads version 2 (already in alpha development) will continue to enhance these tools in response to requirements from the subject gateways. in addition to this, the next version will include: cross searching distributed databases currently each roads subject gateway is searchable in a standalone format although some experimental work on cross searching has already been carried out between some of the gateways. the project is using the whois++ search and retrieval protocol developed by bunyip information systems (who provided some industrial consultancy on the project). this allows distributed databases to be queried over the network and the next version of the software will fully support this cross searching mechanism. in addition to linking the databases together, the project will be investigating the use of a related technology the common indexing protocol (cip) to provide a method of routing search queries to appropriate databases14. cross searching will be particularly useful for the sosig and biz/ed gateways whose subject areas (social sciences and business and economics) overlap; potentially causing confusion for end users trying to identify which gateway they should use. once cross searching is implemented, sosig will no longer continue to catalogue business or economics resources but users of the sosig gateway will still be able to search for and find economics related resources. because of the overlap, the two projects are also looking at ways of presenting browsable lists across the two projects. as part of the desire project sosig is also hoping to collaborate with social science institutions or libraries in europe who want to set up national databases of networked resources. european institutions would be able to make use of the tools and documentation developed by roads and desire to create national gateways. this model is currently being piloted by the koninklijke bibliotheek (national library of the netherlands) who are building a roads database of dutch social science resources. harvesting resources the elib anr gateways concentrate on cataloguing high quality internet resources and it is this human input that distinguishes them from other web search tools such as altavista. users of the gateways are not overwhelmed by thousands of matches to their queries but are presented with a small number of resources which have been through a careful process of selection and description. however, this process means that there is a high cost associated with the creation of the catalogue records. consequently, the gateways tend to catalogue at a server level rather than at the level of individual documents or pages. roads is looking at incorporating a harvest-type technology in order to try to bridge the gap between the ‘hand picked’ approach of the gateways and the so called ‘vacuum cleaner’ approach of the web search engines. one approach is to use a harvested database to supplement the qualitycatalogued records and roads is investigating ways to integrate and present the two. one of the aims of roads and of the subject based gateways is to ‘encourage information providers to become involved in the creation of records about their own data in order to make their information as useful and accessible as possible’15. typically, information providers supply little or no metadata with their resources, due in part to a lack of standards or direction in this area. roads is promoting the idea of ‘trusted information providers’ (tips) who would be identified by the individual subject gateways. the tips may be services or institutions whose information had been previously validated by the gateways that would provide metadata with their resources to be collected automatically. this second level of approach to support the tips idea is to develop a tool that can be used to pre-populate roads databases by harvesting metadata from resources and inserting them into templates. the cataloguers can then ‘add value’ to the automatically generated template before finally submitting it to the database. the roads harvester can be used to generate a single template based on one url or it can be run recursively across a range of urls as a ‘bulk harvest’. the harvester is still under development but in the longer term should help the gateways to redress the imbalance between quantity and quality. contact details debra hiom is a research officer on the sosig and desire projects at the institute for learning and research technology, university of bristol in the uk. she can be contacted at the following address: institute for learning and research technology, university of bristol, 8 woodland road, bristol bs8 1tn, uk. tel: +44 (0)117 928 8443 fax: +44 (0)117 928 8478 email: d.hiom@bristol.ac.uk acknowledgements mailto:d.hiom@bristol.ac.uk summer 1997 33 this paper references the work of the roads project team; in particular jon knight and martin hamilton at loughborough university, rachel heery, michael day and andy powell at ukoln and chris osborne and paul hofman at the university of bristol. any inaccuracies are the author’s own. for more information about roads if you would like more information about the roads project and availability of the software, contact paul hofman at: references 1. resource organisation and discovery in subject-based services 2. electronic libraries programme 3. adam (art, design, architecture and media information gateway) 4. biz/ed (business education on the internet) 5. ihr-info (institute of historical research) 6. omni (organising medical networked information) 7. sosig (social science information gateway) 8. desire 9. publishing information on the internet with anonymous ftp 10. roads template registry 11. roads objectives 13. zexi: z39.50 experimental implementation 14. the common indexing protocol 15. roads objectives * paper presented at iassist/ifdo ‘97, odense, denmark, may 6-9,1997. mailto:roads-liaison@bris.ac.uk http://www.ukoln.ac.uk/roads/> http://www.ukoln.ac.uk/elib/> http://www.adam.ac.uk/> http://www.bized.ac.uk/> http://ihr.sas.ac.uk/> http://www.omni.ac.uk/> http://www.sosig.ac.uk/> http://www.nic.surfnet.nl/surfnet/projects/desire/> http://www.roads.lut.ac.uk/system-docs/internet-drafts/draft-ietf-iiir-publishing-03.txt> http://www.roads.lut.ac.uk/system-docs/internet-drafts/draft-ietf-iiir-publishing-03.txt> http://www.ukoln.ac.uk/roads/templates/> http://www.ukoln.ac.uk/roads/ http://www.ukoln.ac.uk/metadata/interoperability/> http://www.roads.lut.ac.uk/zexi/> http://www.roads.lut.ac.uk/system-docs/internet-drafts/draft-ietf-find-cip-new-00.txt> http://www.roads.lut.ac.uk/system-docs/internet-drafts/draft-ietf-find-cip-new-00.txt> http://www.ukoln.ac.uk/roads/ iassist quarterly spring 2012 5 iassist quarterly editor’s notes data bring maps, archive brings data, and accreditation brings research this issue (volume 36-1, 2012) of the iassist quarterly (iq) is the first of the 2012 issues. this editorial is written in march 2013 when many iassist people have received acceptance for their papers at the upcoming conference iassist 2013 in cologne. i am certain there will be many interesting presentations at the conference. however, presenters can reach a greater audience by having their paper published in forthcoming issues of the iq. the three papers in this issue bring reports on the presentation and availability of data in a geographical portal for geospatial data, the collection and dissemination of holdings in a data archive, and on access to trans-national data and the accreditation involved.. the first paper is scholars geoportal: a new platform for geospatial data discovery, exploration and access in ontario universities. the authors are elizabeth hill and leanne trimble (formerly hindmarch) of university of western ontario and scholars portal, toronto. the paper was presented at the iassist 2011 conference data science professionals: a global community of sharing at simon fraser university and the university of british columbia in vancouver, in the session power of partnerships in data creation and sharing. sharing was the theme for the conference, the session, and certainly also the paper on the scholars geoportal. data collections are no longer only numeric data collections. this paper focuses on the use of geospatial data for learning, reporting a project carried out for universities in the province of ontario on the initiative of the ontario council of university libraries (ocul). the paper describes the background the need and the vision and the components of the geospatial portal project. the project needs to have very good metadata handling in order to provide valid results and the paper demonstrates how the geoportal presents the various different types of data. the portal is also used for research and has further involved policy makers and legal experts. the second paper was presented at the iassist 2012 conference data science for a connected world: unlocking and harnessing the power of information in washington, dc hosted by norc. in the session national data landscapes: policies, strategies, and contrasts, the paper strategies of promoting the use of survey research data archive was presented by the meng-li yang from the center for survey research, rchss, academia sinica, in taiwan. the paper is a report from the largest data archive in asia: the survey research data archive, taiwan, whose collection includes government statistics raw data. the data archive was established in 1994 and now has 1400 members who can draw on the facilities of the archive including direct downloading of datasets. in 2011 a survey showed that about 20% of a relevant group of researchers were members of the data archive. this result and other findings led to strategies on improving the search facility and active promotion of the service. the paper goes into details on the tasks that were necessary to improve the search facility. these details and other observations and experiences are provided for others in the data archive arena paola tubaro, university of greenwich and cnrs has, with marie cros, université de lille and roxane silberman, cnrs réseau quetelet, written the paper access to official data and researcher accreditation in europe: existing barriers and a way forward. the authors perceive that the barriers against trans-national access to data in europe are based upon accreditation and that there are major inconsistencies across the countries. one obvious barrier is that some descriptions are available only in the national language, other barriers are at the policy-level and will require negotiation and coordination. the paper presents the information collected on european accreditation procedures based on the trio: eligibility, application, and service. accreditation is found to involve the criteria of eligibility (who is a researcher etc.), the procedure of application (how to request access etc.), and the level of service (who approves applications etc.). this work is part of the data without boundaries project in the ec 7th framework. the positive conclusion is that almost all european countries provide research access to micro-data, enabled by the open data movement and other factors. but there still remain areas where improvement is needed for better trans-national data access.. articles for the iassist quarterly are always very welcome. they can be papers from iassist conferences or other conferences and workshops, from local presentations or papers especially written for the iq. authors are permitted “deep links” where you link directly to your paper published in the iq. chairing a conference session with the purpose of aggregating and integrating papers for a special issue iq is also much appreciated as the information reaches many more people than the session participants, and will be readily available on the iassist website at http://www.iassistdata.org. authors are very welcome to take a look at the instructions and layout: http://iassistdata.org/iq/instructions-authors. authors can also contact me via e-mail: kbr@sam.sdu.dk. should you be interested in compiling a special issue for the iq as guest editor(s) i will also be delighted to hear from you. karsten boye rasmussen march 2013 editor vol21.2 58 iassist quarterly the electronic freedom of information amendments of 1996 [1] in september, 1996, the u.s. house of representatives and the u.s. senate passed the “electronic freedom of information act amendments of 1996” on a voice vote, with no debate, and with support from both republicans and democrats. president clinton signed the e-foia bill into law in early october. the e-foia amendments revise the text of the freedom of information act (5 u.s.c., sec. 552) or, as it is commonly known, “the foia,” by addressing the subject of electronic records for the first time. many of the amendments took effect after a 180-day period, on march 31, 1997. others do not take effect until november 1, 1997, and still others at a later date. while the foia and its e-foia amendments pertain solely to u.s. federal government records, the influence within the u.s. of this law is such that the ramifications of these new amendments can be expected to be closely watched among electronic records creators, providers, and users in both the public and private sectors of u.s. society, and potentially elsewhere. the purpose of the freedom of information act, as enacted in 1966 and amended subsequently in 1974 and 1986, is to “require agencies of the [u.s.] federal government to make certain agency information available for public inspection and copying and to establish and enable enforcement of the right of any person to obtain access to the records of such agencies, subject to statutory exemptions, for any public or private purpose.” when he signed the e-foia amendments into law, president clinton noted the important role foia had played in the previous 30 years in establishing an effective legal right of access to government information. he underscored the crucial need in a democracy for open access to government information by citizens. he offered his hope that as the government uses electronic technology to disseminate more information, there will be less need for citizens to use foia to obtain government information. the legislative history prepared by the [u.s.] house committee on government reform and oversight as background for the e-foia amendments identifies the purpose of the amendments as providing for “public access to information in an electronic format, and for other purposes...” their history highlights several key aspects of the efoia amendments, from the legislators’ perspective: -electronic records: the amendments make explicit that electronic records are subject to the foia. furthermore they acknowledge the increase in the government’s use of computers and encourage federal agencies to use computer technology to enhance public access to government information. — format request: with implementation of the e-foia amendments, requestors may request records in any form or format in which an agency maintains the records. also, agencies must make a “reasonable effort” to comply with requests to furnish records in [any] formats specified by the requestor, “if the record is readily reproducible by the agency in that form or format.” this change reverses a legal opinion that dates from 1984 and which formed the basis for federal agency practices since that time. this change will be discussed further, below. -redaction: agencies redacting electronic records (deleting part or parts of an electronic record [or an electronic records file] to prevent disclosure of material that is exempted from release), shall indicate the amount of information deleted on the released portion of the record, unless doing this would harm an interest protected by the exemption. further, “if technically feasible, the amount of the information deleted shall be indicated at the place in the record where such deletion is made.” — expedited processing: the amendments establish that certain categories of requestors would receive priority treatment of their requests if failure to obtain information in a timely manner would pose a significant harm. the first such category are those who might reasonably expect that delay in obtaining the information could pose an imminent threat to life or physical safety of an individual. the second category includes requests made by a person(s) primarily engaged in the dissemination of information to the public, e.g., the media, and involving a compelling urgency to inform the public. -multitrack processing: the writers of the amendments an early perspective on the “electronic freedom of information act amendments of 1996” by margaret o. adams* summer 1997 59 created an incentive for requestors to submit narrowly specified information requests to federal agencies by allowing agencies to establish procedures to process foia requests of various sizes on different tracks rather than on a first-received, first-responded-to order. the assumption here is that requests for specific or smaller amounts of information can be completed quickly, so responses to such requests no longer need to be in a queue with more general or larger-volume requests. -agency backlogs: in an effort to ameliorate the phenomenon of significant backlogs in responding to foia requests in many federal agencies, the amendments stipulate that agencies can no longer delay responding to foia requests because of “exceptional circumstances” if such circumstances simply result from a predictable agency request workload. — deadlines: the amendments extend the deadline for agencies to respond with an initial determination to a foia request to 20 working days, from the previous deadline of 10 working days. from the perspective of the executive branch of the government, the part of the federal government that has to implement the e-foia amendments, the effects of this bill are highlighted somewhat differently. according to the justice department’s newsletter, foia update, a major change of the amendments concerns the maintenance of electronic access in agency reading rooms. prior to the efoia amendments, agencies were required to make three categories of records routinely available for public inspection and copying: final opinions rendered in the adjudication of administrative cases, specific agency policy statements, and administrative staff manuals that affect the public. the amendments add to the categories of reading room records and also establish a requirement for electronic availability of reading room records. the new category of records that agencies have to make available in their reading rooms as of march 31, 1997, includes any records processed and disclosed in response to a foia request that “the agency determines have become or are likely to become the subject of subsequent requests for substantially the same records.” by the eve of the new century, december 31, 1999, agencies are to have a general index to foia-released records and are to make this index available by computer telecommunications. theoretically the idea is that making records in greatest demand accessible in an agency reading room should satisfy most future demand for those records. but, even for federal agencies that maintain public reading rooms in their regional offices throughout the country, this expectation may not be met. it suggests that the bill drafters assume that most of the public’s demand for records under foia can be satisfied by having an agency reading room where researchers can come to “read” records. yet the amendments stipulate that anytime an agency receives a foia request for records, the agency must treat the request in the standard foia fashion, regardless of whether it also makes these records available in its reading room or online. in other words, even though it already makes such records available in an agency reading room or online, it must respond formally to the requestor within a 20-day time period, and provide the records under whatever guidelines and fees it has established for processing requests under foia. note here: a foia request is any request for records that invokes or mentions the foia, or freedom of information act. the e-foia amendments, as suggested above, also expand the concept of an agency reading room to what some are referring to as “electronic” or virtual reading rooms. the amendments require that agencies use electronic information technology to enhance the availability of their reading room records. and, for any “newly created reading room records [i.e., records created on or after november 1, 1996 that are in the category of “reading room records”], agencies must, by november 1, 1997, make these records available to the public by electronic means. preferably this new electronic availability should be via computer telecommunications, i.e., in the form of on-line access, such as from a world wide web site(s) established to serve as “electronic reading room(s).” while the amendments do not explicitly state that agencies are to continue to maintain their conventional reading rooms, the advice in the justice department’s foia update, is that they are to do this. in other words, the three categories of administrative and policy records traditionally maintained by agencies in their public reading rooms, plus any records released under foia that the agency determines are likely to become the subject of subsequent requests, must be made available for public inspection and copying in the agency public reading room. further, any of these reading room materials created after november 1, 1996 must also be made available to the public by computer telecommunications by november 1, 1997. finally, there are two additional new requirements that may have significant implications for agencies seeking to comply with both the spirit and the letter of the e-foia amendments. the first has already been mentioned earlier: agencies are to make records available in any form or format requested by the person if the record(s) is(are) readily reproducible by the agency in that form or format. the second new requirement was somehow not highlighted in the legislative history. yet compliance with it may require substantial reorientation in the way federal agencies treat foia requests for information in federal records, when those records are in an electronic format. this is the requirement that states: “an agency shall make reasonable efforts to search for the records [responsive to a request] 60 iassist quarterly except when such efforts would significantly interfere with the operation of the agency’s automated information system.” “search” is defined as “to review, manually or by automated means, agency records for the purpose of locating those records which are responsive to a request.” clearly, the impetus for the e-foia amendments just described comes from evolution in the use of electronic computer technology by both government agencies and the federal records-seeking public. it also reflects growing expectations for public access to more and more government information that has accompanied the proliferation of computer technology, especially personal computers. some aspects of the amendments, however, reflect more than natural evolutionary change. the new right for requestors to choose the format in which they expect to receive federal records reverses long-standing legal opinion. the requirement that agencies use computer technology to search for records in electronic form reverses widespread federal practice rooted in a series of court rulings. to understand how or why these changes came to be law, it may be helpful to consider the historical context from which they emerged. “the freedom of information act in the information age...” our colleague tom brown recently published an article that examines the historical context of the foia. he discusses the case law on how foia related to computerized records prior to enactment of the e-foia amendments.[2] he notes that foia guarantees any person the right to gain access to records unless the records contain information on matters specifically excluded under one or more of the nine exemptions identified in the foia. further if a portion of a record is exempt from disclosure then a reasonably segregable portion of the record is to be provided after deletion of the portions which are exempt. brown notes that in interpreting the foia statute, the courts have consistently ruled that agencies are not required “to create records in order to respond positively to a foia request,” i.e., to provide records in response to a foia or to segregate exempted portions of records in order to release them in response to a request. even prior to the e-foia amendments, the courts seemed to have resolved the question of whether electronic materials were subject to the foia. in a 1982 case cited by brown, a federal appeals court ruled that “[c]omputerstored records, whether...in the central processing unit, on magnetic tape or in some other form, are still records for the purposes of the foia.” further, a 1989 department of justice survey of federal agency practices found that government-wide practice also affirmed that electronic records may be records under the foia. the question of whether the foia required federal agencies to provide requestors with records in the formats of the requestors choosing had generally been decided in ways that allowed the federal agencies to make that choice. for example, brown discusses a 1984 case, dismukes v department of the interior. the case centered on a foia request to interior’s bureau of land management (blm) for a list of participants in oil and gas leasing lotteries in california. dismukes, on behalf of the national wildlife federation — a private organization, had filed a foia for these records and requested that they be provided on 9track, 1600 bpi magnetic tape, in an ibm-compatible format, and with file dumps and file layouts. blm’s office of surface mining provided computer printouts of the requested records. the national wildlife federation appealed this response, arguing that “it is impossible to work with such volumes of data without having it in computer form.” the court of appeals ruled that “the computer printout was fully responsive to the...request.” as a result of this ruling, the department of justice advised federal agencies that they, not requestors, had the right “to choose the format of disclosure, so long as the agency chooses reasonably.” this same principle was upheld a few years later when a requestor appealed the response of the central intelligence agency (cia) to a foia request for an index of documents that the cia had previously released. in response, the cia provided 5000 pages of printouts, arranged by date of release of the item in the index. in its appeal, the requestor asked that this information be made available on tape or disk. the court ruled however, that the “information [the cia had provided] was in a “reasonably accessible form” and furthermore, that the agency was not obligated to provide in electronic format, records it had already provided in paper copy.” these rulings, brown suggests, provide evidence that the courts were in denial of the computer age. while the courts may have been in denial, or in ignorance of the computer age, brown also makes clear that their rulings were predicated on the basic premise that agencies are not “expected to be private research firms, ...subject to every beck and call of a requester.” to bolster this assessment, he cites a number of rulings. one, by a federal district court in pennsylvania in 1988, clarified that foia does not require agencies to write new computer programs to search for data not already compiled for agency purposes. another, a federal appeals court, determined that “the foia dictum to release reasonably segregable portions [of records] does not require creation of a ...summary file because it is not functionally analogous to manual searches” for records that contain information responsive to a request. another concluded that the foia “in no way contemplates that agencies, in providing information to the public, should invest in the most sophisticated and expensive form of technology.” on the issue of format, brown also shows that as early as 1988, the administrative conference of the united states (until october 1995, the federal government agency summer 1997 61 responsible for studying and recommending improvements in administrative procedures to executive branch agencies) had recommended that in responding to foia requests, “agencies should provide electronic information in the form in which it is maintained or, if so requested, in such other form as can be generated directly and with reasonable effort.” taken together, these rulings of the courts and the practices of some agencies responding to foia requests for information in electronic formats point to the rationale for the e-foia amendments. so, now, what can be their anticipated impact? application of the e-foia amendments in general, it is far too early to know how federal agencies will adapt their practices to be in compliance with both the spirit and the letter of the e-foia amendments. similarly, it is much too soon to have any idea whether the cumulative effect of compliance with the e-foia amendments will result, as president clinton said he hoped, in less need by citizens to use foia to obtain government information. there are a few things that we can suggest at this time however. j. timothy sprehe, writing in the federal computer week (january 6, 1997) suggests that the principal benefits of the law “lie in the fact that efoia overturns two bad court decisions.” one of these was a ruling that the national library of medicine’s on-line information systems were not agency records for the purposes of foia. the other was the dismukes case discussed above, where the court had ruled that an agency had the prerogative to decide the format in which it fulfills a foia request. but beyond this, sprehe does not consider the e-foia to be “a great leap forward,” except in the sense that the “fact that it was passed at all may be an important reminder that the public has rights of access to electronic as well as paper-based information resources.” another article, this one in a trade newspaper, washington technology (april 24, 1997), quotes an analyst for the federation of american scientists as saying “agencies must undergo a cultural transformation to accommodate [to] the requirements of [e-foia]...the law does not change reality...but it provides an incentive to modernize.” what he is referring to was further enunciated in this same article by a washington lawyer who raises the question of whether federal agencies have the hardware, software, and personnel that will enable them to be in compliance with the e-foia amendments. as he states, the amendments “raise the issue of the availability of suitable software and the equipment to handle it, such as a client-server system with sufficient storage capacity and database software with advanced features. it also raises the issue of being able to hire and find personnel who are sufficiently trained to use what is sophisticated search and retrieval software to comply with requests.” the principles that underlie the foia and now the e-foia amendments are firmly based on the principles of the u.s. constitution and its bill of rights. yet, the iassist community, sophisticated and knowledgeable as it is in matters regarding maintenance and access to electronic records and information, might well ponder the implications for federal government agencies seeking to be in compliance with the foia as amended. keep in mind that in general, the data community that iassist represents, offers or supports services for researchers knowledgeable in the structure and use of electronic data. social science researchers and those who offer support services to social scientists, such as data archivists and librarians, are among those most likely to expect increasing expansion in access to electronic government information. yet the need by social scientists, generally, for access to administrative and programmatic databases to use as sources for rigorous research and analysis purposes, are quite different than the needs reflected in a significant portion of the requests that, for example, the center for electronic records at the u.s. national archives and records administration, receives. since our holdings reflect the records of the entire federal government, we can assume that at least in some senses, the requests we receive mirror those received by all federal agencies. in our case, for the first six months of the current fiscal year, approximately one-third of the inquiries (over 1100 — and they generated over 1800 separate responses) requested specific items of information from records in our holdings. very few of these invoked the foia, an indication perhaps, of the well-known practice of the [u.s.] national archives and records administration (nara) to treat all requests as if the requestor had invoked the foia, thus negating the need for requestors to use it in order to receive the records or the information in records that they seek. that is, it is the policy of nara to respond to all inquiries in a timely manner, and as responsively as possible. the extent to which it can successfully do this is in large measure a reflection of how informed the request is — how specifically it identifies the information sought, and whether the manner in which the specificity of the request reflects the description that nara has for the relevant records. in our particular case, most requests for specific records or for information from specific records, pertained to the casualty records from the korean and vietnam conflicts, two of our best-known electronic records files. we have long experience in responding to such requests, and the efoia amendments will have virtually no impact on the manner in which we handle responses to these inquiries — from printouts of the full files, or by using some pc-based versions of the files, with retrieval software, that a small business vendor has created and provided to us in beta-test 62 iassist quarterly format. two individuals have developed web sites with these records where anyone can access them. interestingly, such widespread availability of these records seems to have had no impact on the continuing demand that we receive for specific information from them. in microcosm at least, this experience suggests that the ready availability of records does not stem the direct demand for them. so, our real challenge in complying with the e-foia requirement to search electronic records for information responsive to a request will not be in regard to the casualty records. rather it will be in responding to requests for information that may reside in any of the other 30,000 (and growing) electronic records files in our holdings. and this is but a reflection of the challenges facing federal agencies as a whole. while each agency only receives requests for information from its records, whereas nara receives requests for information from archival records of the entire federal government, the records in agencies are more current and thus they are in more urgent demand than archival records usually are. there is no arguing amongst us that technological innovation makes access to public information more efficient and varied in ways that none of us could have imagined even just a few years ago. but, we also know too that with each innovation, new complexities and possibilities have also presented us with the reality of a seemingly infinite rise in the level of expectations for what we can do or offer. keeping pace with such expectations, or rather, determining the nature of the “reasonable” effort by which we measure the limits for responding to those expectations, is perhaps the greatest challenge for those seeking to provide service to citizens that is responsive both to the spirit and the law embodied in the foia, as amended. the experiences of the data community can assist and influence these determinations. outreach by the data community to the larger universe can perhaps also help to inform their expectations. notes 1. the first portion of this paper is based primarily upon the committee on government reform and oversight, u.s. house of representatives report 104-795, 104th congress, 2d session, electronic freedom of information amendments of 1996: report [to accompany h.r. 3802], september 17, 1996. this paper also draws upon u.s. department of justice, office of information and privacy, foia update, vol. xvii, no. 4, fall 1996. 2. thomas elton brown, “the freedom of information act in the information age: the electronic challenge to the people’s right to know,” american archivist, vol. 58, spring 1995, pp. 202-211. much of the analysis in this section of this paper draws upon brown’s article. * presented at iassist-97 (may 9, 1997), odense denmark name: job title: organization: address: city: state/province: postal code: country: phone: fax: e-mail: url: iassist international association for social science information service and technology association internationale pour les services et techniques d'information en sciences sociales i would like to become a member of iassist. please see my choice below: options for payment in canadian dollars and by major credit card are available. see the following web site for details: http://datalib.library.ualberta.ca/iassist/ mbrship2.html $40 (us) regular member $20 student member $70 subscription (payment must be made in us$) list me in the membership directory add me to the iassist listserv membership form the international association for social science information services and technology (iassist) is an international association of individuals who are engaged in the acquistion, processing, maintenance, and distribution of machine readable text and/or numeric social science data. the membership includes information system specialists, data base librarians or administrators, archivists, researchers, programmers, and managers. their range of interests encompases hard copy as well as machine readable data paid-up members enjoy voting rights and receive the iassist quarterly . they also benefit from reduced fees for attendance at regional and international conferences sponsored by iassist. membership fees are: regular membership. $40.00 per calendar year. student membership: $20.00 per calendar year. institutional subcriptions to the quarterly are available, but do not confer voting rights or other membership benefits. institutional subcription: $70.00 per calendar year (includes one volume of the quarterly) ○ ○ ○ ○ please make checks payable, in us funds , to iassist and mail to: iassist, assistant treasurer joann dionne 50360 warren road canton, mi 48187 usa http://datalib.library.ualberta.ca/iassist/mbrship2.html http://datalib.library.ualberta.ca/iassist/mbrship2.html r eturn u ndelivered m ail t o: ia s s is t q u a r t e r ly c/o w endy t readw ell 1758 p ascal s t. n orth f alcon h eights, m n 55113 u s a newsletter vol. 4 no.2 user services in a data library laine g.m. ruus data library, university of british columbia there seems to be some confusion as to just what is meant by user services. although the term is relatively new to the library literature, the concept dates back to at least 1876 when, amongst others, samuel swett green was beginning to argue "the disireableness of ..personal intercourse between librarians and readers " dictionaries of library science or librarianship do not yet define the term, nor do the library administration texts that 1 have consulted david nasatir (1973) has outlined the components of user services in a data library as consisting of dissemination of data files; analyses on demand for users; training of users; consultation on such subjects as mathematics, statistics, methodology, data analysis, etc; the conducting of training programs; a "current awareness" function acquainting local users with parallel research being done elsewhere as reflected in the catalogues of holdings of other archives; and the creation of machine-readable codebooks. alice robbin (1977) has defined user services as consisting of reference (or finding the right data for the right user) training; data and documentation reproduction and dissemination; data preparation, processing, and analysis; project planning; and instruction and orientation. these then are my terms of reference when i speak of user services in the data library/archive context. samuel rothstein (1%1) defines reference service as "the personal assistance given by the librarian to individual readers in pursuit of information..." given this definition, his "reference" = our "user services." his three levels of reference service are minimum, middling, and maximum. his "minimum" or conservative level of service is that in which the librarian is merely a guide to the use of the collection, and the user is encouraged to be as selfsufficient as possible through user instruction and orientation techniques—a policy of "laissez-faire." "middling" service offers a great deal more service to the individual user (especially, in an academic library, to faculty and graduate students), including in-depth searching of the literature, compiling bibliographies, and generally doing a fair amount of the user's work for him. on the other hand, "conservative" service is given to the undergraduate student, on the assumption that learning to use the library and its resources is part of the educational process, and the student should therefore not be spoon-fed "maximum" or liberal service relates primarily to the situation of a special or corporate library, where the librarian's raison d'etre is to do that part of the research that involves the literature, leaving the part that mvolves the laboratory (or whatever) to others qualified in other fields. that is, the emphasis is on the delivery of information, rather than on delivery merely of books, journals, etc., which might contain the information. the information delivered is authentic, relevant, and founded "on the impeccable scholarship of the librarian" (rothstein, 1%1) ti.e application of this scheme to the data library/data archive spectrum is not difficult. on the conservative side there is, for example, the new, local-service data library, quietly growing in the bowels of some university structure--in an academic department or computing centre. because at first it must concentrate on developing internally (you must have a collection before you can provide services based on it), this data library offers only basic services to users: acquisition of machine-readable data files (mrdf) on request, epedally if the researcher knows already where a file is to be had; maintenance of data files as they are received from outside sources without much, if any, cleaning or checking for coding, wild punches, inconsistencies, etc.; access to codebooks as they are received from the supplier; and some basic consultation on problems involving data processing and statistical procedures. difficult problems are referred to experts: computing problems to the programmers in the computing centre, statistical problems to the statisticians in the statistical centre, data identification problems perhaps to the librarians in the library. the middling level of service is exemplified by the same local service data library several years later, when it has grown both internally and in its services. because the collection is now large, with many massive and newsletter vol. t, no.2 intricate mrdf, the data library has developed an online inventory describing in great detail the collection, to assist not only users but the data library staff as well in identifying and locating specific files. data files acquired in a "dirty" state are cleaned, because it is easier to clean them at once, while one still has contact with the principal investigator, than several years later when a researcher needs the data and it is found to be unuseable. codebooks are routinely converted to machine-readable form, because this is the most satisfactory way of ensuring that any user can get access to a copy at any time. (and how else does one economically supply a copy of a gallup survey codebook to a political science class of 60 on one day's notice?) special programs are developed, and specialpurpose subfiles of data files are prepared, to make things easier (and cut down on hand-holding) for novice researchers, and more espeaally, as the only way to provide adequate service for 200 freshman commerce students who every year descend on the data library for their annual excerrise with crsp stock price data. (although the staff has grown, it still cannot provide consulhng to 200 students, and the only way to give them sahsfactory service is to make the procedure as "idiot-proof" as possible.) regular orientations, and some impromptu ones tailored to special courses, are given to introduce novices to the facility and its services, and there may even be a manual describing how to search the inventory, how to mount data files, special programs developed for often-used files, etc. there is now a variety of staff experhse that can be called on for consultahon but the "toughies" are still referred to the experts in other departments, and there is shll stnct adherence to the basic principle that the user should do the actual work himself, because this is after all part of the educational process, and the "true" researcher learns to be selfreliant. and at the 'liberal" far end of the spectrum, i envisage a special purpose data library embedded in a research institution, or a corporation, or a government department. much of the work of this data library is involved in the actual creation of new data files, the maintenance of on-going data bases, and the secondary analysis of existing files. the data library saff here are experts in computer programming, statistics, sampling, etc.; these experts are part of the research team of any project, handling such details as the technical aspects of research design, the research instrument, actual data gathering, and, later, data analysis and interpretation. here there is no question of the user doing his own work. the level of service given depends on the expertise of the staff; there is no need for such ancillary services as orientation tours and courses, of "idiot-proof" data files and documentation for novices, for there are no novices. the data base management system describing the collection is designed for maximum efficiency, and data files and documentation are machine-readable and very clean, because this is the most efficient means of maintaining and updating them. each of these levels of service has its own immediacy of purpose, which level of service a data library or archive approaches is dependent on its user community and the constraints placed on it by available funding and staff certainly the level of sophistication possible in the maximum service archive is far beyond the developmental capabilities of the small, local-service data library operating on two and a half people but the techniques can be transported, and techmques which result in greater efficiency for the end user will also generally result in greater efficiency for the data archive staff as well. my dream, from the point of view of our small, local-service data library serving an academic institution, is eventually to make user services so efficient that no user need darken our door again, except to persuade us to acquire a file that we do not already have. there are, 1 think, several components to this. the first is to make information retrieval as efficient as possible, then to make documentatior access as efficient as possible, and finally, to maxe data access and software access as efficient as possible these efficiencies need not, of course, be implemented in quite this order; one normally does the easiest things first. but 1 am going to treat them in the order in which most users approach them efficiency of information retrieval t)ecomes cntical when the data library's collection grows beyond the point where any staff member can remember what every file in the collection contains. it then becomes necessary, not only for the sake of users, but also for the sake of staff, to be able to quickly find all particulars about any given data file—not just the pnndpal investigator, title, or date of collection, but the individual variables contained in it. the consensus of practice seems to point to some form of data base management system (dbms). it is a little difficult to ascertain who has such systems operating. we know that roper center has had an in-house information retrieval system in operation for some years. we also know that icpsr has offered (in october of last year) on-line access to its inventory through telenet. spires seems to be favourite among mts installations, being in use at the university of alberta computing centre, stanford university libraries, and u.b.c. library. other systems are being used by the university of wisconsin-madison and the university newsletter vol. k no.2 of washington the data clearinghouse for the social sciences in canada also had developed a data base management system, although its present fate is unclear. among these systems, the amount of information offered on any given file varies immensely. none that i have seen, however, includes quite the detail envisioned for the data documentahon system that ifdo is supporting, which contains 16 pages of (mostly optional) fields per record. the characteristics most desirable in a dbms for this purpose would seem to be that it; • be flexible enough to handle a variety of file formats and interrelationships. • be easy and cheap to maintain. • have a powerful and flexible searching capability • be easy to teach users to search. • be available at all times that the computer is up. in other words, as a user service, it should be capable of informing a user at a remote location of what files a . collection contains, and should provide sufficient information to access the data without his visiting the data library there is another aspect to the question of efficient information retrieval, the retrieval of information vis-avis data files not held in the local collection--that is, the identification and location of data files held elscwhere-for eventual acquisition this is not a problem that can be solved at the level of the local service unit. present services are dependent on local collections of the inventories of other data libraries or archives, government departments, etc; literature searching; and the intuition of the local data library staff. what is needed is a union catalogue of the holdings of all known dissenvinators of mrdf, and some efficient means of access to information on what new data files are being created. the movement by icpsr and the roper center towards on-line remote access to their inventories is a mojor step towards information retrieval. the data base developed by the data clearing house for the social sciences in canada, which included the holdings not only of canadian data libraries/archives, but also the holdings of government departments, the commercial sector, and private individuals, did for a time provide a union catalogue of canadian holdings of canadian data files, albeit in a restricted subject area. the data base was unfortunately not made remotely available, and the dch is now defunct (no cause and effect relationship is implied). what is needed is a similar product at the international level. efficiency of documentation retrieval is more involved. documentation comes in a variety of forms and formats some data files have no documentation at all, others have as their documentation only the second paragraph of a personal letter addressed to person x, and yet others have exceedingly complex documentation that is much longer than the data file that it describes. documentation can be in hard-copy or machine-readable. the utility of machine-readable documentation for user services is obvious. it is the only way that documentation can be made maximally accessible, since in the optimum system, all machinereadable documentation is accessible at all times that the computer is up. placing extra copies of documentation in a local library does not serve quite the same purpose, since i know of no library that is open as long hours as the computing centre. and a local computer system with a network of remote terminals scattered throughout an institution gives more immediate access than the library, which is invariably at least a twenty minute walk from where the user is (a very important consideration on the west coast where it rains a great deal). however, to make the documentation machine-readable is not quite good enough. no one wants to read through the whole of the six nation project codebook (inkeles, n.d), which occupies 9 linear inches of sh ,'lfspace, to find the twenty variables wanted in order to maximize the efficiency of documentation retrieval, what is needed is some computer-assisted means of searching the codebook, so that one can retrieve just those portions describing variables containing specified key words. the roper center's proposed on-line system is designed to retrieve at this level of information. but there are many codebooks which are not amenable to this type of treatment, such as that describing the cansim system, which contains over 300,000 time series and as many variable labels. maintaining a data inventory is one thing, but maintaining this level of information in an on-line data base is, for a small localservice operation, at present quite unfeasable. maximizing data access efficiency is dependent on the efficiency of tape mounting procedures in the local operating environment. however, it should be possible to develop a system such that, based on the information given in the information retrieval system, it is possible to gain access to all data files at all times that the host computer is functioning, not just at those times that the data library is open. desirable features of data access would include; a minimum number of commands to access a file (whether on disc or tape); and maximum simplicity of commands--i.e., mechanical details such as mode, blocking factor, label, reel number, and position should be transparent to the i$t newsletter vol. 4 no.2 user. in addition, a data access system can be designed to maintain data security (e.g., check the permit status ot computer ids), collect statistics on tape mounts, perform an sdi funcbon (referral to updated editions of files), and reinforce the moral conditions of data access (such as by cautioning against unauthorized copying for dissemination of files by users). of course, efficient access to data and documentation is of little value if one does not also have access to appropriate software. software acquisition, creation, and support is very much dependent on the institutional structure within which the data archive/library exists. the smaller a local-service data archive/library, the more likely it is to be heavily dependant on the software resources of the host computing centre, with little possibility of influenang the computing centre's decisions as to hardware and software procurements, software to be supported, and future software (or hardware) developments. if the computing centre in question is user-oriented (which is the exception rather than the rule), if it provides comprehensive documentation, user orientation, introductory courses on commonly used statistical packages and programming languages, and extensive programming consultation, then the data archive/library need not provide these services. it, on the other hand, the computing centre is not useroriented, the problem of providing effiaent access to appropriate software becomes serious, and the solutions are neither easy nor cheap as the amount of computing expertise dedicated to the data library/archive (or employed in it) is increased, however, an individualized level of service can be provided. this may extend to special-purpose housekeeping programs, to improve the efficiency of such invisible services as data file cleaning, rationalizing inventory, tape mounting and security. it may also extend to visible services such as writing a formula for the retrieval of data from a complex database, or creating a program to provide uniform access to and manipulation of aggregate data across several censuses. as the amount of computing expertise dedicated to user services is increased yet further, and basic software for common problems has been provided, yet more elaborate services become possible, such as the development of special purpose software once one has all these systems operating, of course, one must teach one's users to use them. seminars, special class presentations, and short courses seem to be the normal approach. a good addition is a user's manual, wntten not for the computer programmer (as most computing documentation is), but for the novice user. a history professor, with almost no computer experience, should be able, with the help of your manual, to sit down at a terminal, search your inventory for any files containing, eg , illegitimacy rates in ontario in the 19th century, have printed a copy of the (short) codebook on the local high-speed printer, and mount your data file, all on a rainy sunday afternoon. and then, of course, there arc the special purpose subfiles, teaching packages, and special purpose documentation for the 200 commerce students. designing this kind of service requires more effort to be put into instructor-education than into studenteducation it is essential that the instructor have designed the assignments so that they are appropriate to the data files he wants to use the data library staff should be given ample warning of the assignment it must know the expectations of the instructor, and the level of consulting he or she will give the students as part of the course with adequate preparation by the data library staff, the onslaught can be relatively painless, and will have a marvelous effect on tape mount statistics throughout, i have been considering only "minimum" and "middling" service levels of the data library/archive. at both levels, the user's loing the job is considered part of the educational experience, and self-reliance is something to be fostered 1 shall not consider here the extent and nature of special services offered in the maximum level data archive each of the above components is composed of transparent services and apparent services. apparent services i define as those which the user can see and appreciate. in this category i include the basic provision of information services; acquisition on demand of mrdf; consultation services of all kinds; provision of some kinds of dissemination services (e.g., relieving a principal investigator of the responsibility of maintaining and disseminating his data file himself); insuring efficient access to documentation, data, and software; orientation activities; provision of a user's manual, and so on. transparent services are those which the user seldom notices or consciously appreciates, including the extensive level of reference services involved in the identification and location of fugitive mrdf, "on spec" acquisition of mrdf, rationalization of tape library and inventory procedures, routine data file cleaning and machine-readable codebook creation, and creation of special-purpose files to support special-purpose software, such as a map file to support graphic retneval of census data. newsletter vol. 4 no.z why this distinction? user services are usually the raisoii d'etre of the facility, and the reason for its conhnued funding. apparent services are directly user-oriented transparent services, although often refinements of apparent services, are more often stafforiented, and only indirectly contribute to that desirable phenomenon--the satisfied user my conclusion is that, at the outset, a data archive/library should concentrate on apparent user services, in order to culhvate an active and supportive user community, and only later in its development, when this has been accomplished, develop transparent user services. reference tools for machine-readable data files john g. kelp laboratory for political research university of iowa it would appear to be a rather awesome task to try to summarize the current state-of-the-art in data reference; for surely after 16 years of steady growth in archives, reference matenals giving access to the data in such archives should be exhaustive yet this is clearly not the case, although some attempts to provide useful reference tools in this area have appeared over the past decade. in the brief report which follows, i have singled out for discussion three categories of reference tools for machine-readable socal science data which seem well-established those reference tools which appear to play the most prominent role in directing users to appropriate machine-readable data are: (1) data catalogues describing the contents of individual social science data archives and data libraries; (2) directories describing the contents of more than one archive, or directories within special topical areas; and (3) penodicals, like .; s data, which have attempted to report at regular intervals informahon on the holdings of social science data archives an attempt will be made to examine each category of reference tool from two very personal perspectives: (1) the user consultant who is continually asked to locate data files which must meet a number of very special condihons; and (2) the editor and compiler who must try to ickate all known data files relating to certain topical areas and acquire informabon on the most recent acquisitions of data archives in the former role, one is frustrated by the inadequacies in reference tools; in the latter role, one is amazed that we have come as far as we have. data catalogues lists of holdings, guides to resources, archive directories, or data catalogues are available from most individual data repositories probably the most wellknown of those documents which describe individual archives would be icpsr's guide ic resources and scn'ices, issued on a nearly-annual basis since sometime in the i960's. others of this j(enre would include the recently-issued ssrc survey archive data catalogue, the sleiiimctz archwes: catalogue and guide ((1978), the b-a. s.s. invenlairc des archives disponsibies (1975), catalcf^ue of machine-readable records in the national archives of the united states (1975), and the lokahseringsoversi^l of the danish data archives (1978), to mention just a few. physically, documents of this type are soft-cover, book-length descriphons of the holdings of a major archive. these are real publications, meant for broad distribuhon to a national or international clientele. entries are often arranged according to some broad subject classification scheme which might include such major headings as community and urban studies, elites and leadership, mass political behavior, social welfare, religion, the international system, legislative and deliberative bodies, etc. the individual entries usually include title, author or data collection agency, population covered and/or sampling scheme, time period of study, number of cases and variables, distribution restrictions, and a brief abstract summarizing the purpose of the study and the focus of major categories of variables. iassvol201 18 iassist quarterly data and social science rhetoric: policy and instuction by kenneth p. jameson1, university of utah the 1996 american economics association presidential address was given by health economist victor fuchs(1996) and frames the issue of this paper quite well. after quoting a 1965 article by george stigler that "the age of quantification is now full upon us" and that it will be "a scientific revolution of the very first magnitude"(p.6), fuchs goes on to note that the revolution is still in the future: "(b)ut the shallow and inconclusive debate over health policy in 1993-94 contradicts (stigler's) expectation that this research would narrow the range of partisan disputes and make a significant contribution to the reconciliation of policy differences."(p. 6) there are many other examples of contemporary social issues, and related policies, whose resolution seems immune to the insights claimed by social sciences: environmental disputes ranging from the northwest salmon to the utah wilderness, welfare reform, in particular the treatment of teenage welfare mothers, or even the inheritability of intelligence with all of its racial and social darwinist implications2. does the continued intransigence of social issues and the stubborn intractability of social policy imply that all of the developments in data collection, data storage, and data analysis have come to naught, that the age of quantification is a bust and that the millennium should see social science move in a different direction? i don't draw that conclusion, though i accept the problem as quite real. i believe that social science and empirical investigation can make important contributions to our understanding and to resolution of policy issues, but only if we are clear on the nature of social science and the role of quantification. in particular we must admit the limits of our truth claims, their communal nature, and the possibility of their being utilized to serve vested interests. we must then be very clear about our potential contribution and must educate our students and the public about what we can offer. finally, i think that we must find ways of broadening access to our basic data and analyses in order to include a wider array of interests in the dialogue. visualization techniques graphics-based policy simulations may be fruitful in this regard. if so, we social scientists and social science data managers will still be active participants in the issue/policy discussions and our data and analysis could have an even greater effect on those debates. this view implies a very different set of challenges for social science data librarians and for social science data users, one which will be interesting and potentially quite helpful for social scientists as well as for social policy. the nature of social science: rhetoric fuchs offers three possible explanations for the unsatisfactory state of affairs he describes: that health economists cannot agree among themselves, that the results were not disseminated adequately to influence policy, and that differences in values could not be bridged by empirical research. he developed a questionnaire to examine the issue and concluded that on "positive(i.e. logical positivist) issues," there is substantial agreement among health economists(seventy-two percent gave the same answer to seven of his questions). there is less(thirty-four percent) agreement on "policy-value" questions. so he concludes that value differences account for the irrelevance of economic analysis to the health care reform debate, though the inability to convince policy-makers about the "positive" results also contributed.(p. 15)3 fuchs settles comfortably into the mainstream understanding of what economists do and even quotes its central document, milton friedman's essays in positive economics(1953). it is based on a distinction between positive or scientific statements by economists and normative or value-laden statements. that seventy-two percent of health economists agree with fuchs on seven propositions is taken as evidence of the possibility of positive economics; fuchs claims that this realm should be expanded to find those positive components of policy-value issues, i.e. that economists can solve social and policy issues in the degree that they can succeed in attaining positive scientific results. there is a different and more satisfactory way of understanding the very real problem that fuchs highlights. it starts again from a particular understanding of the nature of economic science, in this case that economics uses "rhetoric" to arrive at conclusions whose status and limitations can best be understood within that context. let us examine this approach. in an important article in one of the central journals of the economics profession, d. mccloskey(1983) argued that economics is best understood as a form of argumentation or persuasion rather than the value-free scientific endeavor that logical positivists would have us believe. this effort does allow economists and other social scientists to "make knowledge," but it is a contingent knowledge which depends greatly on the operation of the scientific community of economists, the times, the biases or 19spring 1996 ideologies of researchers, the historical development and context of the issues, and the technical capacities of the scientists. rorty(1987) describes this as "pragmatism" and suggests that the aspiration of scientists should be to find mechanisms to bring about "unforced agreement" among themselves, rather than to reach some objective truth. examination of the main economics journals indicates that mccloskey's perspective is gaining a very slow and gradual acceptance(sims, 1996). however, this methodological perspective remains controversial (maki,1995), and, for the most part, economists remain minimally introspective about their methodology. when economics is viewed from the perspective of rhetoric, the shortcomings of the age of quantification are not surprising. databased empirical analysis is only one among many approaches to persuasion/knowledge. indeed for aristotle, data are "extrinsic" to making an argument, i.e. not an inherent part of the process. as crowley(1994) notes: ancient philosophers seem to have had a clearer understanding of the limited usefulness of empirical facts than moderns do. perhaps because of their skepticism about the nature of facts, ancient rhetoricians were equally skeptical about the persuasive potential of facts. aristotle wrote that facts and testimony were not truly within the art of rhetoric...he considered extrinsic proofs to be outside of the art ofrhetoric because a rhetor only had to pick them up and display them to an audience.(p. 6) for the greeks the intrinsic components of arguments were the proofs and the canons or principles. proofs can be logical or "logos," pathetic (emotional) or "pathos," and ethical or "ethos," all of which can and do play a role in persuasion, even in economics. the canons prescribe how a persuasive argument is structured, its arrangement, style, delivery, and memory(or links with other pieces of shared knowledge)(covino and jolliffe,1995). making a convincing argument about a social issue or policy is very complex from this perspective. data and quantitative analysis play a role, one which has certainly increased in importance since the nineteenth century. but many other elements enter into any research, influence what is accepted as true, and determine what is persuasive in policy. indeed, much of mccloskey's original article is concerned with illustrating how metaphors, appeals to authority, analogies, etc. are immanent in good economic argument. to understand how we might approach data and empirical analysis differently and enhance its role in arguments over social issues and policy, let me illustrate the claims that i am making with several specific cases. social science rhetoric observed i have chosen three case studies to illustrate different aspects of the claim that rhetoric best describes what we do in social science. one comes from "psychology" considered very broadly, one reports on a recent treatment of advances in macroeconomics and economic policy, and the final example is an overview of the debate on nafta (the north american free trade agreement) and the role of economic analysis. a. social darwinism in modern clothes darwin's evolutionary theory with its mechanism of natural selection provided a powerful metaphor for viewing society.4 it easily lent itself to categorizations of societies and of societal groups as superior and inferior. herbert spencer's "survival of the fittest" aphorism provided a handy shorthand for this supposed process and was the basis for "social darwinism," the belief that the elite of society had attained that status because of their evolutionary superiority.5 the development and application of the iq test around the turn of the century provided a new quantitative measure which could be used to document the differences between the superior and the inferior, be it in terms of economics or race or gender. after this analysis was carried to its logical extreme, in the eugenics movement, all became aware that there were many confounding factors in the environment which could account for most of the variation across groups. advances in social scientific knowledge combined with a realization of the potentially terrible implications of eugenics to discredit such simple stratifications. this seems an excellent case in which quantitative analysis resulted in reaching a truth, i.e. that individual variation has a strong biological basis, though most differences across broad groups are better explained by cultural and environmental factors. however, the debate is newly joined, history is in the process of being repeated. biology and evolution have once again become quite dynamic, and genetic determinism is being explored in every area of the individual, from the prevalence of cancer to aggression and criminal behavior to homosexuality. evolutionary ecologists use the biological framework to examine the whole gamut of societal differences from children's food preferences to marriage behavior. and though biologists and evolutionary ecologists are careful to delimit their claims, it was inevitable that the new biology would spawn a return of social darwinism and of scientific racism. the latter had continued to exist in the backwaters of social science and in non-standard journals, and it received much more widespread consideration through its identification with arthur jensen of university of california and richard herrnstein of harvard. the more recent and more interesting case, from a social science perspective, is the social darwinist manifesto the bell curve. (murray and herrnstein,1994) the argument is simple: that the demands of the modern economy and society require higher levels of cognition and, as a result, a cognitive elite has emerged. 20 iassist quarterly entry into the elite is largely determined by genetic inheritance of iq(sixty percent roughly). while "social murrayism," i.e. "rule by the fittest," is not new, it is presented in a thoroughly modernist manner, i.e. with reams of data and statistics. indeed the 552 pages of text are accompanied by over 100 pages of statistical appendices and tables of regressions and other tests, symbolic of "the quantitative age." at the same time the book illustrates the failure of that age. it was a best seller and reached a wide audience. it was widely reviewed, and, as noted by gould(1994), most of the reviewers immediately disqualified themselves from assessing the quantitative claims of the book. it gained a great deal of credibility simply for its many pages of tables, for its quantitative argument. the very presence of tables became an important part of the book's rhetoric on the issue. and they disqualified many participants from the discussion because of their selfadmitted inability to assess the quantitative basis for the argument. goldberger and manski did examine the quantitative analysis, and they found that the authors' claims were not supported by the data and statistical analyses. they wrote: "(w)e conclude the bell curve. is driven by advocacy for hm's vision, not by serious empirical analysis. america may or may not be on the way towards a custodial state. policy interventions may or may not be effective. we know no more after studying the bell curve. than we did before."(p. 775) although herrnstein was a psychologist, the book can be seen as a political tract that has consciously and extensively adopted the quantitative rhetoric of social science, with notable success. in any case, the issues which the book focused on forced the american psychological association to issue a report of a task force on "intelligence: knowns and unknowns"(neisser,1996) which followed an earlier statement of the american association of physical anthropologists. these are reminiscent of anti-social darwinism/racism statements issued by the american anthropological association in 1938 and unesco in the 1950s(degler, 1991, pp. 203204). the apa report supports the work that has been done in measuring intelligence while at the same time noting its limitations. they conclude: "what is responsible for(group differences in test performance)? the fact is that we do not know. various explanations have been proposed, but none is generally accepted."(neisser, p. 94) in any case this debate illustrates that one of the central issues that was posed in the 1880s--and seemingly solved through quantitative analysis--remains open to dispute. this is despite the incredible increase in availability of data and in technical sophistication of quantitative analysis. the advances of the quantitative age have not been successful in resolving the century old issue. if we are to reach any conclusions, we must bring to bear a much wider range of mechanisms for discourse, including ethos and pathos as well as the empirically based logos which is accepted as the substance of contemporary social science. b. modern macroeconomics sims(1996) provides another excellent example, in this case from contemporary macroeconomics, which again illustrates the nature of economic discourse and the inability to reach agreement based solely upon logical positivist canons. his conclusions are quite different from my own, however. macroeconomics, the study of economic processes at the national level, is dominated today by a theoretical approach termed "new classical" economics. one of its founders, robert lucas, was honored with the nobel prize for economics in 1995. there are two outstanding elements of new classical economics: it is consistently based on the reigning deductive economic theory of markets and maximization, and it allows little role for government policy in affecting the macro economy. the newest work in this genre deals with one of the remaining puzzles, the existence and explanation of business cycles. its approach is to use "computational experiments," computations which are not based directly on empirical or econometric work. this is a departure from traditional quantitative approaches in economics, though it remains fundamentally quantitative. the sims article is a critique of this approach and makes the case that normal empirical investigation of economic phenomena can lead to scientific progress. his concerns have many parallels with fuchs's. the underlying question is why social science disciplines seem to be turning away from empirical investigation, moving toward anthropological ethnography on the one hand or toward non-empirical quantitative model solving and calibration on the other. he concludes that "the popularity of the critiques(of traditional empirical work) probably arises from the excesses of enthusiasts of statistical methods."(p. 109) the promises of empirical investigators have gone beyond what can be delivered, which contributes to the isolation of economists from policy debates. my conclusion is that we should be more careful of our claims, and we should realize that they are simply one input into the argument, the rhetoric, about significant economic and policy issues. sims's suggestion is quite different and diametrically opposed to fuchs's. he uses the metaphor of economic researchers as a priesthood or guild whose purpose is to perpetuate a given body of knowledge.(p. 107) that knowledge he terms "data reduction," his term for advances in natural science and for potential advances in social science. traditional data managers and users will find support from sims. he catalogues many advances in empirical macroeconomics gained by applying probabilitybased inference to the new class of dynamic, stochastic, 21spring 1996 general equilibrium models which are at his frontier. this may be tempered by his suggestion that those who persist in "technically demanding forms of theorizing and data analysis" should spend less time criticizing and more time reading each other, i.e. should join their priesthoods and guilds together. from my perspective, such a step would only reinforce the separation of economic researchers from economic policy discussions, the problem highlighted by fuchs. and given the theme of this paper, that we need to find mechanisms to open up our discourse to a wider community, the direction that sims suggests is inconsistent. from the standpoint of data managers and social science computing specialists, creating greater solidarity among guilds would simply extend and expand the isolation that troubled fuchs. it would cause data to be further removed into the realm of a very narrow "discourse community" insulated from broader discussion and from participation in public policy debates. c. the nafta debate and economic analysis the debate over the north american free trade agreement was characterized by incongruous alliances, shifting alignments and deep divisions over the merits of the pact. fast track authority was approved in response to fears of an unmanageable contest among special interests. nafta was approved by a narrow margin after supplemental agreements over labor and environmental issues were added and after last minute bargaining by the clinton administration. the debate was acrimonious and often gave way to polemics. orme (1993:2) has argued that the debate was not about the agreement itself but was instead about "competing domestic political agendas and irreconcilable world views." however, most of the debate was not conducted in these terms. combatants presented their arguments as scientific facts, based upon sophisticated empirical analysis, above reproach and self-explanatory to all who would honestly examine them. a series of articles appeared from both sides whose objective was to dispell the myths and fallacies of opponents.(orme 1993) opposition was equated with faulty thinking, incomplete reasoning or plain stubbornness. particularly divisive was the debate over the employment and wage effects of the nafta. economists entered this debate in an unprecedented manner and seemed to be integral to the debate in contrast with the health care debate. seemingly, every position required an economic model churning out specific supporting numbers. the model of choice during the debate became the computable general equilibrium (cge) model. its highly mathematical nature tended to recast the debate in terms of who had the best numbers rather than addressing the multifaceted divide separating opposing viewpoints. in congressional hearings, little mention was made of the various structural considerations within the models nor was attention given to the implications of various assumptions. instead, numbers of jobs to be lost or gained were quoted back and forth. because of the lack of transparency of cge modeling, the policy discourse tended to focus on the sheer volume of studies supporting a particular position, or on the source of the studies. as the debate reached its finale, ideas and observations seemed to subside in favor of an endless numbers game. various models generated the number of jobs which would be lost and the number of jobs which would be created under the proposed agreement. wild variations existed between the most optimistic and the most pessimistic projections. the clinton administration eventually settled on a figure of 200,000 job gains. almost all of the modelers advocated or opposed nafta. there was little discussion of the possibility that an agreement could have positive impacts under certain conditions, with negative effects in other circumstances; the collapse of the peso showed this to be a major failing. the debate, then, was over whose numbers were better and which study was more scientific and impartial. indeed, discussions of the jobs issue often incorporated phrases such as "every reputable study" and "a distinguished economist." a joint economic committee report recently concluded, "the predictions of the studies [of the effect on jobs of nafta] are widely contradictory and the utility of the studies in reaching policy conclusions on nafta is extremely limited."(glenn, 1993, p. 1) the arguments based on cge models tended to obscure rather than illuminate the policy debate. their complexity and sensitivity to specification tended to focus arguments on the quality of the model, rather than on its policy significance. the most advanced cge and econometric models represent the state of the art in terms of internal consistency and mathematical elegance. however, these models contain a huge number of equations and entail many hidden assumptions about unknown parameters: elasticities of supply and demand, crosselasticities of demand, substitution rates between capital and labor, expenditure functions, and so forth. the solutions require high-powered mathematical algorithms. often the results look as if they came from a classic black box: only the authors of the models, and perhaps a few other scholars, understand all the ingredients (hufbauer and schott 1992, p.51). economists were central to debate over nafta, however the debate was cast in terms of these highly mathematical models. this effectively limited discussion within policy circles to the results of these models, with legislators quoting numbers back and forth amongst themselves. the result was a debate filled with studies and statistics, all of which seemed to add little to effective communication. further evidence of the limits on the role of economic analysis was the 51-49 vote for nafta despite the heavy weight of 22 iassist quarterly economists and their studies on the pro-nafta side of the debate. so nafta provides a third example of the inadequacy of empirical analysis for resolving public policy debates. in this case the most advanced and sophisticated approaches to economic analysis were marshaled for the debate. the end result of the effort was a complete analytical stalemate, with the resolution of the issue depending upon politicians' commitment rather than the persuasiveness of economic analysis. indeed the widely varying projections and the clearly interested participation of economists was probably counterproductive. what conclusions can be drawn from these three experiences? and what direction might we go as economists, as social scientists, as data managers and data users, to change the manner in which we enter the public policy debates and the manner in which we teach social science? what lessons for instruction? mccloskey's original article advocated the rhetorical stance because economists would write better, teach better, have better relations with other disciplines, make better arguments, and have better dispositions--quite the promise!(1983, pp.512-515) it is not clear that increasing use of rhetoric has notably changed the discipline, i.e. there is no evidence that economists have become more even tempered in the last thirteen years. nonetheless, mccloskey's claim about teaching should be taken seriously. he argues that: economics is badly taught, not because its teachers are stupid, but because they often do not recognize the tacitness of economic knowledge, and therefore teach by axiom and proof instead of by problem-solving and practice...it is frustrating for students to be told that economics is not primarily a matter of memorizing formulas, but a matter of feeling the applicability of arguments, of seeing analogies between one application and a superficially different one, of knowing when to reason verbally and when mathematically, of what implicit characterization of the world is most useful for correct economics.(p. 507) this perspective has very important implications for instruction and, implicitly, for democracy as well6. the key to defining the difference is that rhetoric as an approach to knowledge is based upon persuasion, is based upon discourse, and strives to reach "unforced agreement." thus to teach a discipline requires more than simply amassing a set of axioms, proofs and facts that are then transmitted and embodied in explicit knowledge of the students. it requires engagement and active knowledge-making on the part of students. the difference from traditional teaching is seen quite clearly in instruction in social science. if there are few truths to be transmitted, students must be empowered to become actively involved in investigating issues and in reaching the level of agreement that they can. of course the best results and the best techniques of social science should be used, and the best and most extensive sources of data should be brought to bear. but all of the most sophisticated approaches available must be combined with the broader issues of persuasion, with the compelling metaphor, with the ethical stance of the argument and those making the argument. this makes a very different challenge out of teaching and forces a reworking of the goal of the educational process. the process is likely to involve much more collaboration among students and many more efforts to transmit tacit knowledge of a discipline. students must first be convinced to become involved in the discourse, in the effort to "make knowledge." while they will not reach an irrefutable truth, they can gain greater knowledge about issues, using data and other extrinsic proofs, and they can then defend a position and contribute to reaching some better resolution of the issues involved. when they have a stake in the outcome, their attention to the data and the techniques increases. teaching based upon rhetoric also may alter the definition of the task of the data manager, the data librarian. the challenge of finding information, finding data, and knowing the methods with which to peruse the data remain. but the challenge now it to enable student interaction with the data, student research or investigation or knowledge-making. and the data can only be part of the argument. so active access to and interaction with the quantitative element of economic discourse becomes the goal, and learning is facilitated to the extent that is achieved. the new technologies are challenging the very meaning of data, for the usual organizational categories agreed upon through the library of congress are becoming virtually irrelevant. the new organization is keywords and thesaurus based, implying that students can create their own organization of data and redefine it in that fashion. of course the meaning of thesaurus comes into play here, i.e. storehouse or treasury. one of our social science librarians had her students do a search using a new web based search engine, and each student returned over 500,000 references for the term selected. how can they organize that data? what lessons for public policy? the policy problem is more complex. very rarely can a policy be advocated without support of a social science analysis. even the utah state legislature attempts to base its parochial attempts to return to an earlier age upon social science analysis, albeit research which is used in a way opposite its author's intent. (wilson, 1996) and there are now mechanisms in place which allow each side to have its own analysts, e.g. the whole industry of think tanks in washington and in state capitals who collaborate with and often serve legislators and even lobbyists. in this regard policy analysis has come to resemble forensic testimony 23spring 1996 more than science. from a logical positivist standpoint, this would be a very unsatisfactory state of affairs; truth should not depend on the discourse framework. from a rhetorical perspective, however, this is understandable and speaks to the importance of social science. at least each position does have to have a justification that can be understood in terms of social science. this is to the good and a testimony to the importance of social science analysis. that opposing viewpoints can have their defenders often simply reflects the partial and contingent nature of social science knowledge, though at times disputes may represent an abuse of knowledge and research7. however, the inability of social science to give definitive conclusions may contribute to cynicism about the process and to dismissal of social science as a contributor to the debate. so in the long run, this situation may turn unfavorable to the social sciences. in summary, the use of data and social science analysis in instruction and in policy illustrates both the positive and the negative of current approaches to social science. in the case of instruction, adoption of a rhetorical understanding is indeed likely to improve instruction and to give to students, particularly undergraduate students, a better sense for how social science relates to their lives. in the case of policy, the opposite trajectory seems to be underway. while we can understand why all sides have their experts, in the long run social science may be sullied and will have much less importance in policy debates. how might social science and data(quantitative analysis) be used more effectively in policy debates? what avenues are open? i believe that the combination of "rhetoric" and the new technologies opens up new avenues of linking social science approaches with policy issues. the task is to make our approach and analysis more accessible to wider groups of persons. how can this be done? we are beginning the attempt to use "visual rhetoric," i.e. to present the analyses in visual terms rather than in our more common statistical terms8. this approach of visualization is becoming more accepted and used in science. there are major projects at ncsa and argonne in illinois and at cornell. the former are termed "caves" and combine three dimension projection of audio and video with high performance computing power. they have produced a number of simulations of complex processes in visual form, e.g. the cooling of molten metal running down an inclined plane. to my knowledge they have not simulated economic or social science processes, though cornell does talk of sociological problems. we are experimenting with different display devices for exhibiting data and for allowing the viewer to maneuver through the data to investigate relations that may appear. this is a very inductive approach which we hope may broaden access to data and data interpretation. we will see if the conclusions drawn differ from those the statistical procedures had indicated. as you can see there is no need for knowledge of sophisticated statistical techniques, that anyone with visual acumen could examine the information and search for patterns. the second approach that we will be using is development of simulations of social science phenomena, e.g. the role of education in expected incomes and health outcomes of children in utah. here we hope to put in sets of transitional probabilities and allow the observers to change the amounts of education and see the differing outcomes for categories of children, based upon previously calculated statistical relations. this draws much more upon the deductive framework used in economics. it should allow a much broader range of participants into the discourse and open up active interaction with the issues and the underlying analysis. we hope that this could be a useful input into public policy discussions and could even guide some decisions. while this work is only in its beginning phases, there is evidence that the impact of "visual rhetoric" can be substantial. what is needed now is its application to closing the breach between social science and social policy by broadening access of the public to the results of social science analysis. finally, this effort may provide a new role for the data librarian and the data analyst, one which places emphasis on the interaction with data as much as with its location and access. for to the extent that the information can be presented visually it should allow much wider access to the data and to the possibility of its interpretation, and therefore it should broaden the range of persons who can be included in the discourse on a particular issue. whether this will result in a partial response to the query of victor fuchs and will allow better use and more influence of economic analysis remains to be seen. and whether policy will be better made, is yet another question that is far from being answered. references covino, william and d. jolliffe. 1995. rhetoric: concepts, definitions, boundaries. boston: allyn and bacon. crowley, sharon. 1994. ancient rhetorics for contemporary students. new york: macmillan college publishing company. degler, karl. 1991. in search of human nature: the decline and revival of darwinism in american social thought. new york: pantheon books. friedman, milton. essays in positive economics. chicago: university of chicago press, 1953. fuchs, victor. 1996. economics, values, and health care reform. the american economic review 86#1(march 1996): 1-24. glenn, john. 1993. opening statement of senator glenn. in nafta job claims: truth in statistics , s.hrg. 103-386. washington, dc: united states government printing office. goldberger, arthur and c. manski. 1995. review article: the bell curve by herrnstein and murray. the journal of economic literature 33(june 1995):762-776. gould, stephen jay. 1994. "curveball." the new yorker (november 28, 1994):139-149. herrnstein, richard and c. murray. 1994. the bell curve: intelligence and class structure in american life. new york: free press, 1994. hufbauer, gary clyde, and jeffrey j. schott. 1992. north american free trade: issues and recommendations. washington, dc: institute for international economics. maki, uskali. 1995. diagnosing mccloskey. the journal of economic literature 33(september 1995): 1300-1318. mccloskey, d. 1983. the rhetoric of economics. the journal of economic literature. 21(june, 1983): 481-517. neisser, ulric et.al. 1996. intelligence: knowns and unknowns. american psychologist 51#2(february 1996): 77-101. orme, william a. jr. 1993. continental shift. free trade & the new north america., briefing book. washington, dc: the washington post company. rorty, richard. 1987. science as solidarity. in john nelson, et. al., eds. the rhetoric of the human sciences: language and argument in scholarship in public affairs. madison:the university of wisconsin press. sayre, kenneth, ed. values in the electrical power industry. notre dame, ind.: university of notre dame press, 1977. sims, christopher. macroeconomics and methodology. the journal of economic perspectives 10#1(winter 1996): 105120. wilson, anne. 1996. "legislature fudges facts in justifying club ban. the salt lake tribune (may 12, 1996):a1,a5. endnotes 1. paper presented at the iassist/computing in the social sciences conference, hotel radisson, minneapolis, mn, may 1995. k. p. jameson, department of economics, buc 308, university of utah, salt lake city, utah 84112. email: jameson@econ.sbs.utah.edu tel: 801-581-4578 fax: 801585-5649. my thanks to david plante for his assistance, and to the members of the "rhetoric in the disciplines study group" at the university of utah for continuing stimulation and support, especially to chris oravec, susan miller and mary reddick. 2. the earliiest controversy i worked on was the safety and desirability of nuclear power. we organized a multidisciplinary team to examine its various dimensions, work which resulted in a book.(sayre, 1978) the subsequent decimation of that industry indicates that we had a better sense of the complexity of the issue than the firms involved in the industry at that time. 3.there are no criteria for differentiating "positive" statements from "policy-value" statements. indeed, there often seems little distinction, aside from the level of agreement among the economists. 4. purcell(1973) convincingly traces the late nineteenth creation of modern social science and our social science disciplines to the intellectual ferment created by darwinian's evolutionary theories. 5. karl degler(1993) provides an excellent history of social darwinism and its demise, and then the recent resurgence of biology which again opened the door to social darwinism. 6. purcell's(1973) treatment of the relation of democracy and social science in the early twentieth century is an excellent point of reference on this important issue. 7. one current case in point comes out of "medical science," which used the statistical experimental techniques also used in social science. a drug test of "bio-equivalence" of four thyroid drugs was suppressed by the contracting company which apparently did not like the results and therefore raised a series of objections.(wall street journal, april 25, 1996) 8. there are a number of internet web sites related to this issue. much of the external information for this section was taken from them. iassist quarterly 2010 / 2011 23 iassist quarterly sharing qualitative and qualitative longitudinal data in the uk: archiving strategies and development by libby bishop and bren neale1 these trends in qualitative research are part of a wider movement to enhance the potential of research data abstract over the past two decades significant developments have occurred in the archiving of qualitative data in the uk. the first national archive for qualitative resources, qualidata, was established in 1994. since that time further scientific reviews have supported the expansion of data resources for qualitative and qualitative longitudinal (ql) research in the uk and fuelled the development of a new ethos of data sharing and re-use among qualitative researchers. these have included the timescapes study and archive, an initiative funded from 2007 to scale up ql research and create a specialist resource of ql data for sharing and re-use. these trends are part of a wider movement to enhance the status of research data in all their diverse forms, inculcate an ethos of data sharing, and develop infrastructure to facilitate data discovery and re-use. in this paper we trace the history of these developments and provide an overview of data policy initiatives that have set out to advance data sharing in the uk. the paper reveals a mixed infrastructure for qualitative and ql data resources in the uk, and explores the value of this, along with the implications for managing and co-ordinating resources across a complex network. the paper concludes with some suggestions for developing this mixed infrastructure to further support data sharing and re-use in the uk and beyond keywords: qualitative data, qualitative longitudinal data, data archive, data policy, secondary analysis, re-use, uk 1 introduction over the past two decades significant developments have occurred in the archiving of qualitative and qualitative longitudinal (ql) data in the uk, supported by two major funders of social science research and of data archiving: the economic and social research council (esrc) and the joint information systems committee (jisc). the first national archive for qualitative resources, qualidata, was established in 1994. this followed feasibility studies that set out the case for gathering such resources together for preservation and encouraging data sharing and secondary use by enabling access to publically funded research data. since that time further scientific reviews have supported the expansion of data resources for qualitative research in the uk and fuelled the development of a new ethos of data sharing and re-use among qualitative researchers. as part of these developments, in 2007 a qualitative longitudinal research initiative, the timescapes study and archive (timescapes 2010b), was funded by the esrc to scale up ql enquiry and create a specialist resource of ql data for secondary use. these trends in qualitative research are part of a wider movement to enhance the potential of research data of all kinds to support robust research, to inculcate an ethos of data sharing, and to provide both generic and specialist infrastructure to facilitate their use. in this paper we trace the history of these developments and provide 24 iassist quarterly 2010 / 2011 iassist quarterly an overview of the main data policy initiatives and recommendations that have set out to advance qualitative data sharing across the social sciences. we provide an overview of the mixed infrastructure in place for qualitative and ql resources, focusing in particular on qualidata as a key generic resource and timescapes as a specialist distributed resource. the paper concludes with some suggestions for developing this mixed infrastructure to further support data sharing and re-use in the uk and beyond. 2 the development of qualitative archiving in the uk qualitative datasets are rich and varied in nature, based on in-depth interviews and a range of ethnographic methods (including participant observation and the generation of fieldnotes, case studies and aural and visual data) to capture the contexts and complexities of real life experiences (hammersley and atkinson 1995; mason 2002). the establishment of qualidata in 1994 as the national archive for the curation of such data was a major landmark in the uk. much of the recent history of developments in qualitative archiving in the uk equate with the history of this initiative. at its inception, qualidata was not a place of deposit itself, but acted as a clearing house by locating existing data collections and arranging for their deposit in suitable institutions such as archives, libraries, museums and other repositories. in 1997, the major funder of qualitative social science research in the uk, the esrc, made it a condition of funding that researchers should deposit their datasets with qualidata. this policy change was a critical factor in enabling qualidata to accelerate the acquisition of its own holdings of qualitative data. in 2000, qualidata was incorporated into the uk data archive, itself established in 1967 and curator of the largest collection of digital data in the social sciences and humanities in the uk. in 2003, the national infrastructure was further bolstered through the establishment of the economic and social data service (esds), a national data service of the uk data archive, which provides access to and support for an extensive range of key economic and social data, both quantitative and qualitative. esds qualidata became one of the core components of the new service. the qualidata quinquennial report 1994-1999 (corti and thompson 1999) provides a comprehensive summary of the early years of qualidata and addresses existing archives, cataloguing procedures, dissemination, re-use, management and funding. more recent developments are documented in subsequent reports for the uk data archive and esds (esds 2009)2. 3 the development of qualitative longitudinal (ql) research and archiving in the uk and elsewhere ql research is well established. qualitative researchers have a long history of engaging with time, through a wide range of methods and from different disciplinary perspectives, most notably, anthropology and oral history (elder 1981). time is built into these studies in a complex variety of ways. retrospective studies capture change through biographical, historical or inter-generational accounts. recently there has been growth in the number of projects that re-visit classic studies of communities, institutions or groups to understand changes and continuities and to re-interpret past findings. prospective studies, on the other hand, track individuals or groups over time in order to capture changes in the making and to revisit changing perceptions of the past and future. individuals may be tracked intensively through particular transitions, or extensively across different periods of their lives. prospective tracking is valued because it captures the immediacy and complexity of real lives as they unfold (neale and flowerdew 2003; saldana 2002; thomson and holland 2003). over the past decade ql methods have begun to gain legitimacy as an integral part of the methodological canon. researchers are increasingly seeking to incorporate ql methods into their research design and a range of studies are now being funded by government, the esrc and the main uk charities (the nuffield and joseph rowntree foundations); e.g., on lone parenthood, families after divorce, the life trajectories of offenders and probationers, passages through primary school or the benefits system, and the life histories of migrants or people living in poverty). until recently ql datasets tended to remain the preserve of the originating researchers. archival collections that bring such data sets together to facilitate re-use remain scarce. notable exceptions are the oral history archives at the british library and the mass observation archive at the university of sussex. mass observation is a key historical resource, a paper archive accessible in person or through an online catalogue, of popular accounts of every day life in the uk, produced by a panel of recorders who respond to thematic directives (e.g., rationing, family food, life during the war, birthdays). the archive is seeking funding to digitise parts of its considerable collection that date back to the 1930s, and a range of secondary analysis projects have been funded to use materials from the resource. recent developments have placed ql archiving more centrally on the map. this began in 2006 when esrc funded the archiving of case studies from the inventing adulthoods study, a nine year study tracking a sample of young people from different regions of the uk (inventing adulthoods 2010). the archiving of the case studies has continued with further funding under the timecapes initative (described below) and the data are held at both esds qualidata and the timescapes archive. a feasibility study into the development and scaling up of ql research and resources (holland, et al. 2005) led to funding for timescapes under the esrc qualitative longitudinal initiative 200712. the timescapes study is resource-led as well as having a strong substantive and conceptual focus. the archive, which was launched in october 2009, is being developed as a resource of ql data, with the current collection focused primarily—although not exclusively—on studies of personal lives and relationships across the life course. as well as data from the inventing adulthoods study, the archive is collating data from seven core projects that span the life course and from a growing number of separately funded ql projects that are affiliated to the initiative. in this way, timescapes aims to build up a range of ql data collections on life course themes across diverse substantive fields in the social sciences, as well as encouraging re-use through secondary analysis initiatives. its holdings are primarily digital and are multi-media, including audio and written data, as well as still and moving images. currently the archive is supported by an institutional repository (ludos: leeds university data objects store) which uses digitool proprietary software. documentation about the technical and procedural development of the timescapes resource is available on the website (timescapes 2010a). these include consent forms, guidelines for interview transcription and anonymisation, user registration documents, a depositor licence and the multi-media metadata schema. further documentation will be added as it becomes available. priorities for the near future are to further develop the timescapes resource as a working archive across a broader range of projects, and to encourage and assist secondary use of the data. timescapes is innovative in encouraging archiving as an integral part of the research process rather than an administrative task relegated to the closing phase of a project. this feature is important because of several characteristics of ql research. firstly, since qualitative researchers usually generate their own data and, in the process, build up relationships with their participants, they have a uniquely personal affinity with and ‘feel’ for the data and the context within which it was generated. secondly, ql research often involves the generation of highly sensitive data that is contextually rich, difficult to anonymise, and therefore runs higher risk of disclosing identities. particular care is therefore needed to preserve confidentiality. thirdly, ql projects are often the product iassist quarterly 2010 / 2011 25 iassist quarterly of individual or small team scholarship that can last over many decades. the originating teams control and have exclusive access to the population samples which make up a study, and determine how and when they are followed up over time. such projects may have a continuously provisional feel; they are never quite finished, either in terms of the potential for further data generation or the endless possibilities for complex analysis and ‘reworking’ data to produce new insights and interpretations. finally, unlike quantitative longitudinal data, which is gathered solely for secondary use, ql data is generated, at least initially, by and for primary researchers to enable them to address particular research questions. archiving for secondary use in this context must therefore run alongside the tasks of ongoing data gathering and analysis by the originating team. this has implications both for the resources needed to attend to these tasks and for the timing of archiving within the project life cycle. given these characteristics, ql data needs specialist curation to encourage deposit and sharing while a prospective study is ongoing. timescapes has developed an innovative stakeholder model of data sharing that enables archiving to be seen as an integral part of the research process. researchers who deposit their data with timescapes are stakeholders in the resource and are encouraged to re-use data as well as depositing, thereby combining primary and secondary analysis to raise new questions and produce new insights. researchers can thereby continue to use their data and link it to other related data as their research progresses. the commitment to good data management planning at the research design phase enables data to be generated and organised for archival as well as primary use and prepared according to international archiving standards along with appropriate documentation or metadata (i.e., data about data that provides important context for the resource). depositors are in the best position to provide rich and descriptive metadata that is aligned with the requirements of temporal analysis. as users of the dataset, depositors have a vested interest in ensuring that the resource is fit for purpose, with accurate metadata and refined thematic search and retrieval functions (e.g., through assigning key words to interview transcripts). crucially, the timescapes archive enables finely granulated controls on the re-use of data by building in different levels of access: public, registered, case-by-case approval, and embargo. for example, permission to use highly sensitive or un-anonymised data can only be given by the originating researchers, and the proposed use needs to be specified in consultation with the originating team. in this way, the archive opens up the potential for the sharing of data that might otherwise remain unarchived (bishop 2009a). in essence the archive builds the necessary infrastructure to enable a more personalised mode of data sharing, which (as will be shown below) has been the primary way in which researchers have chosen to share their data in practice. timescapes is not a stand-alone archive; it is a distributed satellite of the uk data archive which also holds the data for long-term preservation purposes. it is simultaneously a part of the canon of longitudinal resources in the uk, and of qualitative resources, and needs to develop in both directions. part of the remit is to encourage the linking of timescapes data with data from other longitudinal resources, both qualitative and quantitative. these have included, for example, understanding society (the uk household longitudinal study), the national child development study, mass observation and the oral history collections at the british library. there is also evident scope for comparative research and secondary use projects with international collaborators. these are beginning to emerge through equalan (see introduction to this issue). distributed archives such as timescapes can play a crucial role as brokers between specialist research communities (whether defined in terms of data genre, methodology or thematic content), and generic data centres with broader remits, such at the uk data archive (bishop 2009b). the specialist infrastructure being developed in timescapes has the potential to form a valuable bridge between the research and archiving fields, and between primary and secondary research, that would enable these to be seen as iterative and reciprocal processes. 4 data sharing: uk policies, practice and ethos the development of infrastructure to support the re-use of qualitative data goes hand in hand with the development of an ethos of data sharing; both are necessary if data are to be made available for sharing and valued as a resource for re-use. the process of enabling data sharing is developing in a wide variety of ways when viewed comparatively across europe. this is shown clearly by ruusalepp (2008) who comprehensively reviews developments across the 30 countries of the oecd. he shows that organisations such as the oecd, unesco, esfri (european strategy forum on research infrastructures) and codata (the committee on data for science and technology) have policies that promote or recommend data sharing, and that these policies have influenced the policies of numerous uk organisations (e.g. office of science and innovation (e-infrastructure), jisc strategy 2007-2009, and the research information network’s strategic plan). these policies stop short of recommending mandatory data sharing. to date there are no national policies across the countries of the oecd that mandate data sharing in this way, although there is an increase in recommendations for “data management plans” which ask researchers to take into account data sharing and curation, most notably in 2011 by the research councils uk. even so, the ethos of data sharing is strongly endorsed within these policies and is beginning to have a discernable impact at the organisational level. since 2000, the esrc data policy, for example, has required all award holders to offer for archiving and sharing copies of both digital and non-digital data, and similar obligations continue in the 2010 policy (esrc 2000; esrc 2010b). the recently revised esrc framework for research ethics notes that data should be collected with the expectation that others will re-use it (esrc 2010a). (see research information network (2011) for a detailed review of funders’ policies and the funded sherpa juliet project (jisc 2009b) for an international inventory of such policies). while at the level of uk policy there is a clear and growing commitment to data sharing, the extent to which this translates into practice among the wider community of researchers is less clear cut. one way of gauging the ethos of re-use in the uk is through academic debate on this issue which has taken place primarily among a small community of sociologists and archivists. publications that promoted qualitative data sharing began to appear shortly after 1994 (corti 1995; corti and thompson 1998). these inspired rejoinders about the value and ethics of re-use as a research strategy (mauthner et al. 1998; perry and mauthner 2004) that were then taken up in special issues of a number of journals. these debates have done much to open up the issue of qualitative data re-use to the research community. in the introduction to a special issue in sociological research online, a leading qualitative researcher in the uk reflects this changing ethos: what is particularly refreshing and useful about the articles contained in this special issue is the way that they push past the more moralistic overtones of the ‘re-use’ debate to focus instead on what happens, what is involved, what can and cannot be achieved, when sociologists get on and do it. in the process the articles give grounded and finely grained insights into the challenges but also the potential for qualitative ‘secondary’ analysis. in their different ways, the articles are qualitatively analytical about ‘re-use’ and they are engagingly reflexive in their arguments. they make the case for 26 iassist quarterly 2010 / 2011 iassist quarterly using any qualitative data carefully, revealingly, and reflexively, rather than arguing that a specific set of rules applies to so-called data re-use (mason 2007: 1.3). a further important metric has been the willingness of funders to underwrite projects that are fully or partially engaged in data sharing. these include continuing support for the uk data archive and esds, specific initiatives designed to encourage secondary analysis of micro data (e.g., understanding populations trends and processes and the collaborative analysis of micro data resources), and funding for a number of qualitative data sharing projects (see the data exchange tools and conversion utilities (dext) 2006-8; and quads-qualitative archiving and data sharing scheme 2005-06 (uk data archive 2010). these initiatives have taken place alongside the advent of core funding for research methods and infrastructure initiatives, in particular the national centre for research methods. new depositors, especially major research centres, are also an important signal about attitudes toward data sharing. the national centre for social research is the largest independent social research institute in britain, with major holdings of public policy data. in 2009, it began discussions with the uk data archive to plan for depositing its qualitative data. as a further example, over the past three years a steady stream of ql researchers have sought to affiliate their research with timescapes and pursue secondary analysis of the archival resources, or to combine their primary research with secondary analysis as a way of broadening the scope of their data and providing a more robust evidence base. however, there remain many qualitative resources that are not archived centrally but held as independent datasets by the originating teams or institutions, and, it is also the case that much archived data remains under-utilised. this is evident from a number of surveys that have been conducted to attempt to gauge the level of support for data sharing, both in principle and practice. while these often have low response rates and attitudes towards data sharing are in any case difficult to discern from these sources, they do give some indication of prevailing trends. a recent feasibility study into the co-ordination of uk data resources (uk research data service, 2008) found that: • although only a minority of researchers share data via a data centre, almost half need to access others’ data and most share data by informal means, usually peer networks.”. the current picture is clearly mixed. evidence of sharing includes the fact that approximately 1000 data sets are downloaded each year from esds qualidata and this represents only a fraction of the re-use of qualitative data, much of which still takes place informally. also, panels on re-use are becoming more frequent in mainstream academic events such as the biennial research methods festival. however, criticisms continue to be voiced. for example, a recent special issue of the australian journal of social issues (2009) devoted to data archiving reports the views of australian researchers, some of whom oppose data sharing, and such views continue to hold sway among some uk researchers as well. overall, the current picture reflects an uneasy tension between pressures to share data, for the benefit of the wider research community and public good, countered by requirements to protect data and confidentiality, for the benefit of the subjects of research and also for the qualitative researchers who both generated and analysed the data. while it is no longer seen as legitimate to protect data because the originating researchers wish to have exclusive use of it, valid concerns remain about how to protect sensitive and confidential data and how to accommodate the sometimes conflicting demands of conducting primary research with the production of archive ready datasets for secondary use.. the close nature of the relationship between researchers and the qualitative or ql data they generate outlined above needs to be taken into account in the way such data is curated and its re-use facilitated. key factors in the strong move towards data sharing include: making “unmined” data available, avoiding duplication, reduced burden on research participants, greater transparency of research procedures, alignment with open access principles, and recognising that outputs of publicly funded research are public assets (fry et al. 2008). equally important are the concerns for protection, codified in the uk data protection act 1998 (and international laws) which intend, rightly, to assure that all data sharing is done ethically. overall, then, the current environment is challenging and complex, with many general laws, little applied case law, and researchers often subject to contradictory advice (e.g., archives demanding data sharing and research ethics committees calling for data destruction). (for key reports in this debate, see thomas and walport 2008; swan and brown 2008). 5 complex infrastructures for qualitative and ql resources mapping the field of qualitative and ql data resources in the uk is a complex and seemingly never-ending task. given the wealth of resources and their scattered nature, it would take a dedicated project to provide a truly exhaustive inventory. our mapping exercise is therefore highly selective . it is probably safe to say that esds qualidata, as a national resource, is a central hub in this network, especially since its incorporation into the uk data archive, but it is by no means the only hub, and the network is vast. in part, this is because what might count as qualitative data is so diverse – ranging from open ended responses on otherwise quantitative surveys to large holdings of historical materials, to newly emerging blogs, twitter and other “born digital” resources. the forms of these data are also highly diverse, ranging from written and other paper resources, visual and audio materials, film and photography, through to web-based and other digital materials. furthermore, qualitative data for social research is available in a growing number of organisations in the uk. these span libraries, museums, funders’ archives, universities, government departments, broadcasting and media archives, independent institutions and organisations, and localised collections held by community or special interest groups (for a comprehensive review see foster 2004). one reason for the complexity of the network is its interdisciplinary nature and the obvious attraction of bringing thematically or methodologically linked data together in special collections to increase their visibility and enable specialist curation and ease of re-use. for example, oral historians have produced extensive resources of qualitative and ql data. the oral history society website provides information about archives in the uk, including regional collections. a related discipline, discourse analysis, produces its own collections e.g., childes – child language data exchange and talkbank. two major resources with esrc funding are regionally based, with a remit to develop holdings of locality data for re-use: the wales institute for social and economic research and data and ark: access research knowledge on northern ireland. as a further example, the british film institute holds an extensive collection of social historical documentary films that is international in scope5 . a dramatic change since 2000 is the proliferation of digital content held in institutional repositories, most often affiliated with universities (jisc 2009a). timescapes, as a specialist resource of life course data, is one such example. it is currently part of the ludos repository at the university of leeds, which also holds the extensive disability archive. clearing houses for open source repository data are also emerging (e.g., open doar). while such repositories offer immense potential for the curation of specialist or locally generated data, our iassist quarterly 2010 / 2011 27 iassist quarterly initial investigations indicate that, as yet, holdings of qualitative social scientific data in such repositories are limited, and that, where they are held, limited metadata and searching options make these data difficult to locate. this brief overview reveals the great complexity and diversity of qualitative data in the uk and of the infrastructures in place to manage and facilitate access to these data. the picture is one of a ‘mixed economy’ with centralised, generic, national level resources existing alongside distributed specialist resources. the latter are valuable in enabling the specialist curation, discovery and re-use of data that take particular forms, are generated through distinctive research methodologies, or that have a particular thematic or substantive focus. the uk research data service recently carried out a feasibility study (ukrds 2009) for a national shared data service, as part of which they considered a number of options for the future management of uk’s research data outputs. these included, firstly a continuation of the current proliferation of data services and resource with little change in management or co-ordination; secondly, the creation of a highly centralised agency to provide for and manage all new capacity; and thirdly a co-operative service by which the ukrds would be an enabling framework, working across a range of uk stakeholders and acting as a catalyst for new services and partnerships. the report found substantial research infrastructures existing in ‘islands’, with limited coherence and communication among them. the authors recommended a cooperative model for future development that would enable good co-ordination of existing data resources and maximum value from infrastructures and services already in place. this recommendation recognises that data is held, and will continue to be held, at a variety of levels, including project level data sets held by the originating researchers, institutional repositories, specialist archives that focus on particular kinds of data (defined by format, methodology, or thematic content) and in national level data centres. alongside this, however, the need for some rationalisation of existing data services has been recognised, and the development of a more integrated national data service, that could encompass and oversee both generic and distributed resources, is a likely next step in the uk (esrc 2011). 6 future developments the overview presented here suggests a number of priorities for the future development of qualitative and ql archiving and data sharing. firstly, there are practical considerations in building capacity in this area in the uk. there is a skills shortage in data curation and management, which is particularly evident given the scale of the uk network. standards for the management of data and the production of metadata across the network are currently lacking, with metadata remaining inadequate for easy resource discovery. however, a new working group for qualitative data exchange within the data documentation initiative is a promising development. technical challenges also arise, for example, in curating and organising complex forms of qualitative data, such as audio and video formats, and in developing adequate protection of confidentiality as part of the broader ethical challenges of data re-use. more broadly, our review of the developing ethos of data sharing and the mixed infrastructure of qualitative and ql resources suggests a number of challenges. the inherent tensions between protecting and sharing data are evident with qualitative data, and become even more acute with ql data. various strategies may help to further the ethos of data sharing. further mandating by funders will help, but just as important may be the rise of a new status for data as bona fide citeable research outputs, even enabling researchers to receive recognition in the research excellence framework for the production of data sets for sharing (such moves are being explored in the jisc managing research data programme (2010) and in datacite (2010), among other places). perhaps just as critical in encouraging sharing and re-use is the further development of new approaches to archiving that are more closely integrated with research processes and that build on dialogue and collaborative models of sharing. this will depend on the specialist archiving of data with distinctive formats, content or modes of generation, to run alongside and complement generic archiving and to act as brokers between the research community and the national level facilities. such a model has been developed in timescapes but will require follow on funding to be properly realised and tested. there are challenges, too, in working across the mixed infrastructure identified above, to ensure that distributed resources develop in consultation with centralised resources such as esds qualidata, enabling special requirements to be met but without re-inventing the wheel. the development of effective co-ordination between generic and specialist resources and across the network of resources is important, and this needs to include the development of key portals so that data resources can be easily identified, described and located. nonetheless, there are many reasons to be optimistic, even in the face of complex challenges. common principles for managing and sharing data across all uk research councils (and other funders) signal a demonstrable shift toward an ethos of sharing. nor are such signs only in high places. recent workshops on managing and re-using data offered by both timescapes and esds qualidata attracted hundreds of participants. the potential for combining primary and secondary analysis to broaden the scale and historical reach of qualitative and ql research and produce robust evidence for policy and practices is an exciting development that is likely to flourish over the next decade. the provision of well co-ordinated generic and specialist infrastructure to support this development is a vital next step references australian journal of social issues. (2009). spring special issue. 44(3) bishop, l. (2009a). ‘ethical sharing and reuse of qualitative data’. australian journal of social issues, 44(3) pp. 255-272, spring. bishop, l. (2009b). ‘moving data into and out of an institutional repository: off the map and into the territory,’ iassist quarterly iq, 31(3&4). [online]. available at: http://iassistdata.org/publications/iq/iqvol31. html. [accessed 4th march 2010] corti, l. et al. (1995). ‘archiving qualitative research data’. social research update, 10. corti, l. and thompson, p. (1998) ‘are you sitting on your qualitative data? qualidata’s mission’. international journal of social research methodology, 1(1), pp. 85-89. corti, l. & thompson, p. (1999). quinquennial report to the economic and social research council, october 1994 september 1999. university of essex: qualidata, datacite. (2010). datacite international initiative to facilitate access to research data. [online]. available at: http://www.datacite.org/. [accessed 4th march 2010] driver. (2010). digital repository infrastructure vision for european research. [online]. available at: http://www.driver-support.eu/index. html. [accessed 4th march 2010] 28 iassist quarterly 2010 / 2011 iassist quarterly economic and social data service. (2003). accessing and analysing social and economic data: users needs uncovered. a web based survey. [online]. available at: http://www.esds.ac.uk/news/publications/ usersurveyrep.pdf esds survey 2003. economic and social data service. (2009). publications. [online]. available at: http://www.esds.ac.uk/news/publications.asp. [accessed 4th march 2010] economic and social research council. (2000). ‘data policy’, economic and social research council. [online]. available at: http://www.esrcsocietytoday.ac.uk/esrcinfocentre/images/datapolicy2000_tcm6-12051. pdf . [accessed 4th march 2010] economic and social research council. (2010a). framework for research ethics (fre). [online]. available at: http://www.esrcsocietytoday.ac.uk/esrcinfocentre/images/framework%20for%20research%20 ethics%202010_tcm6-35811.pdf. [accessed 4th march 2010]. economic and social research council. (2010b). ‘research data policy’, economic and social research council. [online]. available at: http://www.esrc.ac.uk/about-esrc/information/data-policy.aspx . [accessed 3rd june 2011] economic and social research council. (2011). ‘the uk data service’, economic and social research council. [online]. available at: http://esrc.ac.uk/funding-and-guidance/funding-opportunities/16116/ uk-data-servicecore.aspx [accessed 21st july 2011] elder, g. (1981). social history and life experience in d.h. eichorn. j.a clausen j. haan, honzik, m, and mussen, r. (eds). present and past in middle life. new york: academic press pp 3-31. foster, j. and sheppard, j. (2004). british archives 4th edn. basingstoke: palgrave macmillan. fry, j., lockyer, s., oppenheim, c., houghton, j. and rasmussen, b. (2009). ‘identifying benefits arising from the curation and open sharing of research data’, uk higher education and research institutes’. [online]. available at: http://ie-repository.jisc.ac.uk/279/. [accessed 4th march 2010] hammersley, m. and atkinson, p. (1995). ethnography: principles in practice. 2nd ed. london: routledge holland, j. thomson, r. and henderson, s. (2004). feasibility study for a possible longitudinal study: discussion paper. [online]. available at: (www.lsbu.ac.uk/inventingadulthoods/feasibility_study.pdf ) inventing adulthoods. (nd). [online]. available at: www.inventingadulthoods.lsbu.ac.uk. [accessed 4th march 2010] joint information systems committee. (2009a). digital repositories. [online]. available at: http://www.jisc.ac.uk/whatwedo/topics/digitalrepositories.aspx. [accessed 4th march 2010] joint information systems committee (2009b) sherpa juliet: research funders’ open access policies. [online]. available at: http://www.sherpa. ac.uk/juliet/. [accessed 4th march 2010] joint information systems committee. (2010). jisc grant funding 14/09: managing research data programme. [online]. available at: http://www.jisc.ac.uk/fundingopportunities/ funding_calls/2009/12/1409researchdata.aspx.[accessed 4th march 2010] mason, j. (2002). qualitative researching. 2nd ed. london: sage. mason, j. (2007). ‘re-using’ qualitative data: on the merits of an investigative epistemology, sociological research online 12(3). [online]. available at: http://www.socresonline.org.uk/12/3/3.html. mauthner, n., parry, o. and backett-milburn, k. (1998). ‘the data are out there, or are they? implications for archiving and revisiting qualitative data’. sociology, 32(4), pp. 733-45. neale, b. and flowerdew. j. (2003). ‘time texture and childhoods: the contours of longitudinal qualitative research’ in international journal of social research methodology: theory and practice.6 (3): 189-199 parry, o. and mauthner, n. (2004). ‘whose data are they anyway? practical, legal and ethical issues in archiving qualitative research data’. sociology, 38(1), pp. 139-152. ruusalepp, r. (2008). ‘a comparative study of international approaches to enabling the sharing of research data’ v1.6. [online]. available at: http://www.jisc.ac.uk/media/documents/programmes/preservation/ data_sharing_report_main_findings_final.pdf. swan, a. and brown, s. (2008). ‘to share or not to share: publication and quality assurance of research data outputs’, research information network. [online]. available at: http://www.rin.ac.uk/our-work/datamanagement-and-curation/share-or-not-share-research-data-outputs. [accessed 4th march 2010] timescapes. (2010a). the data archive. [online]. available at: http:// www.timescapes.leeds.ac.uk/the-archive/. [accessed 4th march 2010] timescapes. (2010b). timescapes: an esrc qualitative longitudinal study. [online]. available at: www.timescapes.leeds.ac.uk. [accessed 4th march 2010] thomson, r. and holland, j. (2003). introduction to longitudinal qualitative research in international journal of social research methodology: theory and practic.e 6 (3). thomas, r. and walport, m. (2008). ‘data sharing review’, ministry of justice. [online]. available at: : www.justice.gov.uk/reviews/datasharing-intro.htm. [accessed 4th march 2010] uk data archive. (2008). manage and share data.[online]. available at: : http://www.data-archive.ac.uk/sharing/.[accessed 4th march 2010] uk data archive. (2010). completed research and development. [online]. available at: http://www.data-archive.ac.uk/randd/completedrd.asp.[accessed 4th march 2010] uk research data service. (2008). the data imperative: managing the uk’s research data for future use: a summary of the uk research data service feasibility study. [online]. available at: http://www.ukrds.ac.uk/. [accessed 4th march 2010] research information network (2011). ‘data centres: their use, value and impact’. jisc (september). available at www.rin.ac.uk/data-centres [accessed 13 november 2011]. iassist quarterly 2010 / 2011 29 iassist quarterly notes 1. dr. libby bishop e.l.bishop@leeds.ac.uk; senior research archivist, timescapes study and archive, university of leeds; http://www.timescapes.leeds.ac.uk/ esds qualidata, uk data archive, u. of essex; http://www.esds.ac.uk/ qualidata/ professor bren neale; b.neale@leeds.ac.uk director timescapes initiative and archive, university of leeds uk http://www.timescapes.leeds.ac.uk/ 2. documents are available on the esds and uk data archive websites that give details of archiving policies and procedures for these major services (e.g., preservation, back-up and storage information) and extensive information on managing and sharing data (confidentiality, ethics, consent, documentation, etc.) (esds 2009; uk data archive 2008). 3. additional articles can be found at http://www.esds.ac.uk/qualidata/ support/reusearticles.asp and for further publications see http:// www.disc-uk.org/publications.html#data_sharing). 4. uk and international qualitative data providers are listed here: http:// www.esds.ac.uk/qualidata/access/otherdata.asp and uk and international ql resources are here: http://www.timescapes.leeds.ac.uk/ methods-ethics/international-qualitative-resources/. links to these and other data providers can be found here: http:// www.esds.ac.uk/qualidata/access/otherdata.asp. by 30 iassist quarterly 2008 carina carlhed and iris alfredsson1 swedish national data service’s strategy for sharing and mediating data practices of open access to and reuse of research data resources for researchers to document and make their data accessible for others as the most important obstacle. concerning interventions to enhancing reuse of digital data, the majority of the doctoral students and the professors thought it should be effective to get more information about accessible research data in data archives or databases. nearly 100% in both groups reported that more training in research methods, digital research databases, and information about accessible e-tools would be effective interventions. the most effective interventions for enhancing accessibility to digital data were that research grants should include funds for preparing the data for sharing and archiving and that archiving data for use by the scientific community is acknowledged to be of scientific merit. surprisingly, when it comes to the degree of urgency in sharing their own data, the professors seem to be a bit more eager to share data than the doctoral students. the results are compared with the results from the parallel study of the professors and from a recent survey targeted at professors in various social sciences and humanities disciplines at finnish universities (kuula and borg, 2008). 1. introduction 1.1 building a swedish research infrastructure the swedish research council (vr) has, since its start in 2001, been focusing on the need to build a research infrastructure. as a part of this work the committee for research infrastructure (kfi) was set up in 2004. the main purpose of kfi is to formulate long-term strategies and handle resource allocation for expensive scientific equipment, large research facilities, and extensive databases. the committee also deals with swedish interests in, and funding of, various national and international research infrastructures. the overall aim is to provide better conditions for swedish researchers by ensuring access to high-quality infrastructures. the committee is also the producer of the swedish roadmaps for research facilities to meet future scientific demands. the first swedish research council’s guide to infrastructure was published in 2006 and the second by the end of 2007 (the swedish research council, 2007). as part of the swedish research council’s major abstract this paper begins with a description of the current key actors in sweden, which are promoting research infrastructure and accessibility to research data, put into context. the swedish national data service’s (snd) organization, mission, and strategy to promote data sharing is also described. snd’s strategy is a combination of top-down and bottomup activities. an example of a top-down activity is to influence research funders to put higher demands on future open access data when studies are completed or to support researchers through the whole research process by providing guidelines on ethical and legal issues. examples of bottom-up activities are to be present in different research contexts and to inform about the benefits of sharing data. one example of this is a joint project with snd and four university libraries. snd has conducted a national inventory survey, initiated in the fall of 2008, of existing databases and database research, as well as attitudes towards data sharing among researchers and university managements within social sciences and humanities departments at swedish universities and university colleges. in addition to the inventory process, two survey studies have been carried out in spring 2009, one targeting professors and the other doctoral students in the same domains of disciplines at swedish universities and university colleges. the questionnaire contained 80 items covering the researchers’ affiliations; domain of discipline; gender; age; familiarities with research policies and ventures; and use, reuse, and archiving practices of digital research data. furthermore, there were questions about possible reasons for not using digital data, interventions and barriers to enhanced reuse and accessibility to data, possible agents in overcoming barriers, and willingness to share their digital research data. the surveys were carried out through email questionnaires sent to professors (n=549) and doctoral students (n=1147). the results from the surveys show that doctoral students in general expressed great uncertainty about questions of amounts of reusable digital data and effective interventions to enhance accessibility to digital research data. they identified research ethical aspects as important barriers to sharing digital research data, while professors emphasize lack of iassist quarterly winter summer 2008 31 infrastructure initiative, the database infrastructure committee (disc) was founded in 2006 (www.disc.vr.se). disc’s mission is to promote the development of an effective infrastructure for sharing research data resources in sweden and it aims to ensure that researchers have rapid, easy, and free-of-charge access to research databases of high quality. high quality here means up-to-date, relevant, quality-assured, well-documented, and standardized databases meeting high international standards of quality, comparability, and security. the mission also includes creating new joint research data and facilitating access to existing data. one of disc’s first key issues concerned transforming the swedish social science data service (ssd) into the swedish national data service (snd, www.snd.gu.se). the matter was studied during 2006, and in autumn 2006 there was a call for applications to host the new data service. the university of gothenburg was, in competition with four other universities, appointed to host the snd. a five-year agreement to support the organization was signed by the research council and the university in november 2007. the new organization should not only replace ssd, but also take responsibility for a broader area. according to section 3 of the agreement, snd “shall meet the needs of the research community for data on empirical research in the areas of social science, humanities, and medicine. actions include providing technical, legal, educational, and other administrative resources for collecting, storing, and distributing data for research.” snd is governed by a steering board and by a national reference group. the board of snd consists of five national representatives for the above sciences, appointed by the swedish research council, the national reference group, and the university of gothenburg. formally the new organization started 1 january 2008. however, ssd performed the operational tasks during the first six months of 2008. on 1 july 2008 most of the ssd staff was transferred to the new organization. at the same time ssd was closed down, and the ssd data collection and equipment were taken over by snd. 1.2 the swedish national data service (snd) according to the guiding principles of snd, the main purposes of the data service are to mediate information on databases and other digital material collections for research, to facilitate access to research databases, and to serve as a knowledge node for documenting and managing research data in several knowledge fields. thus, a very important task for snd is to strengthen the altruistic approach of the importance of data sharing and open access among researchers. experiences from snd’s predecessor ssd, show that this is not an easy task. only a small proportion of data produced by swedish social science researchers were deposited at ssd. snd’s conditions are, however, better than ssd’s: increased economic resources and an organization placed within a bigger network of infrastructural resources. however, an important factor for the result is the general attitudes towards data sharing among researchers. is there simply no culture of data sharing and reuse of data among researchers in sweden? or does sharing and reuse exist, but not via a data service? we have identified two major barriers for reaching our goals: legal barriers and possessive barriers. the legal barriers are obstacles in swedish current laws and regulations. the possessive barriers are thresholds connected to unconscious attitudes of researchers. 1.3 legal issues issues surrounding shared data infrastructures have important legal implications. for this reason, disc has surveyed the legal regulations that apply. the survey includes an inventory of relevant regulations in the areas of integrity protection, copyright, and archiving (disc, forthcoming). the report will provide a basis for determining the actions needed to facilitate the creation of a common data infrastructure. a working hypothesis at disc is that the issues involved are so fundamental that they require a public investigation. an example of legal obstacles pointed out by disc is the regulation concerning the use of the swedish personidentified population registers on health and social conditions. this very important source for swedish research gives unique opportunities to create research databases for longitudinal research in medicine and social sciences. however, the current ban on creating a common research infrastructure with personalized data limits the use of these resources. the personal data act, the secrecy act, and other regulations allow the use of research material only for specific projects. universities may not collect and store data intended to serve a wide number of researchers in the same scientific area. disc also calls attention to the fact that the central swedish administrative agencies, mainly statistics sweden, are not given the basic instruction to provide the research community with data from registers. instead they sell research data to individual research projects as the need arises. this results in the fact that research funders over and over allocate funds to purchase the same research data. while disc is looking into the need for new regulations, snd will work on the task of informing researchers on legal issues. the impression is that there is a great deal of uncertainty among researchers when it comes to the legal aspects of data sharing. starting in june 2009, one of the university jurists will support snd with legal advice. this project will include training the snd staff in legal matters 32 iassist quarterly 2008 concerning the collection of research data, as well as the use and reuse of data. the jurist will also represent snd in the cooperation between disc and snd on legal matters. 1.4. possessive barriers data collected by a researcher or a research team are often considered to belong to the original investigator(s). this is not the case, as the ownership of the data most often is connected to the university where the researcher is employed. nevertheless swedish researchers often bring the data with them when changing workplaces. experience shows that data not used and taken care of rapidly get obsolete. documentation gets lost and old data formats become unreadable. when asked if they would consider depositing their data at snd, researchers often doubt that their data are of interest for other researchers. other reasons for not sharing data with others are that data are not properly documented and organized or that reuse of data needs a lot of information from the principal investigator. 1.5. activities to promote data sharing the snd strategy to promote data sharing is a combination of top-down and bottom-up activities. an example of a topdown activity is to influence research funders to put higher demands on future open access data when studies are completed. to make the researcher aware of the complete life-cycle of data, the research plan always should include a plan for how to preserve and share the data. another way of encouraging data sharing is to regard it as a merit to make your research data available for other researchers. another activity is to support researchers through the whole research process by providing guidelines on ethical and legal issues, on how to store and document data, etc. the snd web site will be the central place for this information, but it will also be published in different printed versions. examples of bottom-up activities are to be present in different research contexts and to inform about the benefits of sharing data. one example of this is a joint project with snd and the university libraries in gothenburg, lund, linköping, and malmö. financed by the royal library’s open access program, the aim of the project is to look into open access within the humanities and arts. the one-year project, starting in september 2009, will try to answer the following questions: where and how to store research data? which parts can be published as open access? how to connect the open archives and the swedish national data service? how to connect research data and publications? 1.6. feedback from the research community when working on a strategy to promote data sharing, you need to know the opinion of the target group. inspired by our colleagues at the finnish social science data archive (fsd), we decided to ask the professors within the humanities and social sciences about their opinion on open access and data sharing. to compare with another target group within the research community, we also asked the same questions of the ph.d. students within the humanities and social sciences. 1.7. outline the swedish national data archive (snd) is currently an operative key actor in conducting a national inventory survey, initiated in the fall of 2008, of existing databases and database research as well as attitudes towards data sharing among researchers and university managements within social sciences and humanities departments at swedish universities and university colleges. the aim of this ongoing inventory survey is twofold: first, to identify and coordinate existing data resources; and second, to identify barriers and enablers to using and depositing data to open repositories. some preliminary findings from this inventory survey and follow-up interviews with researchers based at the participating departments have revealed a number of important issues that need further investigation: the general unwillingness towards sharing information about research data with coordinating institutions (such as snd); the reported time scarcity preventing researchers from collecting, coordinating and delivering information about research data; and the ethical concerns of how to handle commitments to research subjects and how to protect sensitive information. a number of issues also clearly related to the fact that there was a wide distribution among social sciences and humanities disciplines represented in the inventory survey. there were, for example, quite different views among the respondents on fundamental issues such as the nature of and purpose of research, research ethics, ownership of research data and research results, and how to best enhance research infrastructures. in addition to the inventory process, two survey studies have recently been carried out – one targeting professors and the other doctoral students at swedish universities and university colleges with input from the above mentioned national inventory study by snd, and from a finnish survey, which was carried out 2006 by the finnish social science data archive (fsd) targeting professors in various social sciences and humanities disciplines at finnish universities and the practices related to open access to research data (kuula and borg, 2008). the empirical part of this paper is based on two recently conducted survey studies, targeted at professors and doctoral students within humanities and social sciences at swedish universities and university colleges, with the broader aim of investigating existing practices and attitudes when it comes to availability and reuse of research data. the results are tentative and descriptive and are discussed in a theoretical context in another conference paper (axelsson & carlhed, forthcoming). iassist quarterly winter summer 2008 33 2. procedures of the surveys the two surveys, one directed to swedish professors in humanities and social sciences and the other directed to swedish doctoral students in the same domains of disciplines, contained 80 items covering the researchers’ affiliations; domain of discipline; gender; age; and familiarities with research policies and ventures, and use, reuse, archiving practices of digital research data. furthermore, there were questions about possible reasons for not using digital data, interventions and barriers to enhanced reuse and accessibility of data, possible agents in overcoming barriers, and willingness to engage in promoteing changes in this area and to share their digital research data. the surveys were carried out through e-mail questionnaires and with lists of respondents based on retrievals from databases at the universities’ offices for it or personnel administration. in some of these lists it was easy to recognize respondents’ disciplines; others were sorted by thematic or interdisciplinary departments and no information about discipline was accessible. therefore, even departments that were within science and technology, educational sciences, and social medicine were included, but only departments that described themselves as interdisciplinary on their websites. nevertheless, most of the departments were within humanities and social sciences. because the population was broad and had somewhat non-distinct boundaries, we asked respondents to reply to us if they did not use perspectives of social science or humanities in their research. in those cases they were cancelled from the survey. initially, the survey was sent to 1589 professors from 35 universities/university colleges, and after the cancelling procedure of non-social science or non-humanities researchers (by the definition above) there were 1436 professors. the response rate was 38%, with 549 responses. the same procedure was carried out with the population of doctoral students. however, the lists from the universities that formed the respondent list had minor inaccuracy problems, due to some “natural” conditions, namely doctoral students becoming doctors. this affected the update status on information in the university personnel information systems, which had in some cases inaccurate information about the doctoral students. in addition, doctoral students at university colleges could also appear at a list from another university, hence with double mail addresses. a check up was made before the distribution of the e-mail questionnaire in order to avoid obvious doubles; however in some cases the e-mail addresses were abbreviated and impossible to relate to the names of the doctoral students. similar to the professor survey, the population was broadly defined, which called for a similar procedure for cancelling, by respondents’ reply stating their non-social sciences or non-humanities affiliation. initially, the doctoral student population included 4697 potential subjects and after the cancelling procedure (mentioned above), 4065 remained. the response rate was 28%, with 1147 responses. when comparing how the professors’ response rate patterns related to the distribution among a selection of the universities that received the largest proportion of questionnaires, we can conclude that the response rate from the larger respondent groups’ universities alternated between 22 to 44 %. see table 1 below. table 1 shows response rates based on the initial number of questionnaires sent before the cancelling procedure of non-social sciences or non-humanities affiliation. because our method of selection was somewhat unstable, we found it necessary to investigate our precision further. the swedish national agency for higher education produces statistics about the universities and university colleges2. xxcomparing statistics of professors and doctoral students and their affiliation to university and disciplinary domain from 2008 and our response patterns gives a view of how our survey succeeded in targeting the population. it seems that the population of professors (constructed from statistics, i.e., number of professors in different domains of disciplines and university), is well-covered by our group of professors who have participated in the survey. in concordance with this one can argue that our response frequencies are higher in reality (see table 1). table 2 shows the doctoral students’ response rates 34 iassist quarterly 2008 related to the distribution among a selection of the universities that received the largest proportion of questionnaires. like the professors’ response rate patterns, the table above shows response rates based on the initial number of questionnaires sent before the cancelling procedure of non-social sciences or non-humanities affiliation. as argued above, the actual response rate is higher when comparing it with the statistical population, which we constructed for comparison reasons. for some universities, however, the response rate was lower in this comparison. it signals distortion in our precision about the doctoral students. in conclusion, our generalization opportunities are limited due to these aspects that have been discussed above. it seems that the ground for conclusion about the group of professors is more stable than the group of doctoral students. nevertheless, a large number of professors (n=549) and doctoral students (n=1147) have participated in our studies, which implies considerable opportunities to assume valid conclusions. 2.1 generalizability in the professors’ group, a majority of men answered the questionnaire, 73% compared to 27% women. this reflects the demographics of the larger population, whereas 23% of the professors in social sciences, humanities (and law) are women. in the group of doctoral students the conditions were opposite; 61% of the doctoral students in our survey were women. in comparing with the statistics from the swedish national agency for higher education, the larger population consisted of 56% women. in both cases we can conclude that women were slightly a bit more inclined to participate in our surveys than men. considering age, with our survey we seem to engage a larger part (25%) of the younger doctoral students (younger than 29 years old), than expected (16%). the same counts for the group of professors, but there were only a minor difference. two percent more of professors participated that were younger than 50 years old (18%), compared to statistics from the swedish national agency for higher education (16%). according to domains of disciplines, it seems that our groups of professors and doctoral students reflect the structure of the larger population (table 3). based on the discussion above, our conclusion is that the results from our surveys could be treated as fairly valid, in spite of the relatively low response rate. the amount of responses from professors and doctoral students in different domains of disciplines, age, and gender corresponds to the official statistics that have been described and discussed. 3. results the swedish research council has in a current venture made a long-term strategic plan a roadmap the swedish research council´s guide to infrastructure (the swedish research council, 2007). in the questionnaire we asked the researchers about their knowledge about this venture and their opinions about it. eleven percent of the professors were familiar with the venture and the guide and only 1 % of the doctoral students. half of the professors’ group did know about the venture but not its details and 40% did not have knowledge about it at all. this was also true for iassist quarterly winter summer 2008 35 the majority of the doctoral students (82%). professors were more inclined to express positive opinions about the venture and the doctoral students followed the same pattern (figure 1). the knowledge about the oecd guidelines on open access to research data from public funding (2007) was generally low; 75% of the researchers (both groups) did not know about it at all. surprisingly, 61% of the professors were not aware of its existence (figure 2). breaking down results by domains of disciplines, it seems that professors within social sciences are the most informed about the oecd guidelines, and the group which was least informed was the doctoral students within law. considering the situation of being a doctoral student, we are not surprised at the large amount of them not having knowledge about the guidelines and/or the research venture mentioned above. 3.1. archiving practices and reuse of digital research data the primary condition of archiving and reusing digital research data is that data are collected and compiled in some way. 73% of the professors stated that digital empirical data are used in research and 16% stated that the use of digital empirical data is unusual or are never used. the major part that did not use empirical digital data was the professors within humanities (56%). among the doctoral students, they declared that digital empirical data was used (42%), but they expressed a great uncertainty about these questions generally. according to the professors, the digital data are often kept by the researcher after analysis and reporting, without any actions to documentation (46%) but according to 15% of the researchers, they occasionally keep the digital data without any further documentation. archiving practices where digital data always are kept and documented in a catalogue/database at the university is quite rare (11%). almost half of the professors’ group stated that these practices were unusual or never occurred. the same tendency showed concerning facilitating availability of digital research data at a data archive, that is, unusual practices. however, it seems that the research data are not regularly destroyed after analysis and reporting; at least 49% stated that destruction is rare 36 iassist quarterly 2008 and only 3% reported that it was common. the reuse of digital data are relatively common; 59% of the professors stated that data are reused in ph.d. works or other research projects and only 3% reported that it never happened. the use of reused digital data in teaching is also quite frequent according to 59% of the professors. reusing all kinds of empirical data is most common in situations when researchers use the data themselves; approximately one-third stated that they pass data on to other researchers who are studying similar kinds of areas. five percent of the professors reported that this never occurred. regarding the amount of the digital data that are reusable, professors are more optimistic in general than the doctoral students, who seemed very uncertain and had difficulties expressing opinions of estimates (figure 3). there were small differences between both groups and the domains of disciplines in these issues. the wide range of empirical research data within social sciences and humanities is shown in figures 4 and 5. in both domains of disciplines researchers use a broad empirical base, especially when it comes to use of nondigital empirical materials, where 74% of social sciences researchers and 86% of humanities always use several empirical sources. when considering the use of digital data it seems that the empirical data are less varied, according to 66% of the researchers in social science and 68% of the researchers within humanities. important reasons for not reusing digital data are mentioned by the professors as uncertainty about the quality of data (62%), ethical aspects (57%), technical issues (53%,) and juridical issues (49%). however, the professors’ group is divided in opinions and the other part does not think that these factors are crucial (38%-50%). according to those who think ethical aspects are important, we found that these professors were mainly from social sciences. that is true iassist quarterly winter summer 2008 37 also for doctoral students in the same domain of discipline. the importance of juridical aspects is represented by the doctoral students in law, but not the law professors. both professors and doctoral students in humanities deviated in general from the others in these issues, i.e., the technical issues were considered important. they also report other reasons for not reusing digital data, such as not using empirical or/and digital data at all, lack of knowledge and routines, decontextualized data having weak relevance for others, etc. concerning interventions to increase reuse of digital data, 95% of the doctoral students and 93% of the professors thought it should be effective to get more information about accessible research data in data archives or databases. nearly 100% in both groups reported that also more of training in research methods, digital research databases, and information about accessible e-tools would be effective interventions (89%-95%). it seems that professors and doctoral students in humanities are most positive towards more education interventions and researchers in social sciences are the least positive, but all groups are generally positive to the interventions proposed. 3.2 obstacles to sharing digital data our seven suggested obstacles to sharing digital data have been ordered in degree of perceived difficulty by the respondents. the professors regard deficiency of resources for researchers to document and arrange their data for reuse as the most difficult obstacle to sharing digital data. they also ranked lack of other resources like guidelines and directions for documentation as an important issue. another obstacle highly ranked by the professors was doubt about the correct use of their data, i.e., risks of mistakes and misuses. an additional impediment was the fact that their respondents were not informed that their contributions should be used in the research society in general, only for a particular study. juridical obstacles and loss of one’s own advantage of competition in keeping data to oneself were not considered as crucial. the least important obstacle, according to the professors, was ethical aspects such as threats to confidentiality and delicate information. the doctoral students, however, thought that the ethical aspects mentioned above were the most difficult obstacles of all. after that, they considered the information given to the respondents and the use of their contributions to the research society in general was an important issue. deficiency of resources for researchers to document and arrange their data to for reuse were also ranked as important, followed by juridical aspects. the least important obstacles according to the doctoral students were lack of other resources like guidelines and directions for documentation, loss of one’s own advantage of competition in keeping data to oneself and doubt about the correct use of their data, i.e., risks of mistakes and misuses of data. the response pattern did not change depending on the researchers’ use of digital data or not. on the other hand, researchers in social sciences and women were more concerned with research ethical aspects and threats to confidentiality, etc., while researchers in humanities and men tend to stress lack of resources to document and arrange their data for reuse. according to age, older researchers tend to emphasize lack of resources and juridical issues. younger researchers pointed out ethical aspects, threats to confidentiality, and doubts about incorrect use of their data. thirty-eight percent of the professors and 34 % of the doctoral students stated that these obstacles did not prevent them from sharing data with the swedish national data service (snd). however, 65% of the researchers indicated that these obstacles did prevent them from sharing data to snd. when we asked if, for example, snd could help them to overcome these obstacles, 43% (of the 65%) responded positively. the research funders were also regarded as important agents in overcoming the obstacles by 43% of those who expressed that the obstacles prevented them to share their data. researchers who did not feel prevented to share data believed to a greater extent that snd could be of help (55 %) and research funders as well (59 %). there were however, minor differences, where the older researchers and the researchers in humanities were more optimistic about overcoming the suggested obstacles. there were very small differences between professors and doctoral students. on the other hand, if the researchers would consider engaging in promoting alterations in these areas, the doctoral students tend to embrace issues of research ethics and changing values in accessibility and practices in sharing data, while professors were inclined to issues of jurisdiction. surprisingly, when it comes to the degree of urgency in sharing their own data, the professors are a bit more eager to share data (30%) than the doctoral students (24%). in total, there were 53% of the researchers who thought it was urgent to share data (55% of the professors and 52% of the doctoral students), but only 26% in the total group reported that they intended to share their data. a large proportion of the total group expressed doubts about sharing data (40%). researchers in law were the least keen on doing it and thought it was not so urgent. researchers in humanities however, were those who distinguished themselves as potential “sharers.” according to gender and age, men and older researchers expressed more willingness to share than others. 3.3 promoting accessibility to digital research data the results show which authorities the researchers point out as important key actors in promoting accessibility to publicly-funded digital research data and also to actively participate in shaping guidelines. the universities and university colleges were considered as the most important key actors according to 82% of the researchers and the second was the two largest research funders for social sciences and humanities, the swedish research council 38 iassist quarterly 2008 and the swedish council for working life and social research, with 80% of the researchers’ responses. fifty percent stated that statistics sweden would also be participating in shaping such guidelines. no significant differences between doctoral students and professors, or age, in this matter were observed. according to domains of disciplines there were very small differences, except for the opinion about participation of statistics sweden, where researchers within social sciences emphasized this as a key actor to a larger extent than the others. it seems also that male researchers are a bit more pessimistic about the importance and role the proposed key actors could play. furthermore, the most effective interventions for enhancing accessibility to digital data were that research grants should include funds for preparing the data for sharing and archiving (88% of the doctoral students and 83% of the professors) and that making data accessible for the use by the scientific community is acknowledged to be scientific merit (87% of the doctoral students and 83% of the professors). generally, the doctoral students were more optimistic about the efficiency of interventions proposed, especially the issue of acknowledgement of promoting accessible data to be of scientific merit and except for those mentioned as top-ranked above, where the professors expressed beliefs in their efficiency in a higher degree. there were no or very small differences in response pattern among the domains of disciplines in these issues. according to gender it seems that women researchers believe in more education about life cycles of digital data (74%) and research ethics (77%) to a higher degree, than men. 59 % of the male group thought that more education about life cycles of digital data is needed and 63% of the male researchers believed that more education about research ethics is necessary. 4. discussion our results are descriptive and have been presented tentatively in this article. further statistical analyses are needed concerning impact of differences and in addition, scrutinized examination of all openended questions, where a lot of interesting comments are made by the researchers. overall we interpret the researchers’ attitudes towards current ventures and strivings in research infrastructures as predominantly positive. a key actor is the swedish research council that has, in a current venture made a long-term strategic plan a roadmap the swedish research council´s guide to infrastructure (2007). the researchers’ knowledge about this venture was minor. professors were more inclined to express positive opinions about the venture and the doctoral students followed the same pattern. the knowledge about oecd guidelines on open access to research data from public funding was generally low and somewhat discouraging; professors within social sciences were the most informed, however. the least informed were the doctoral students within law. in comparison with the finnish survey (kuula & borg 2008), where 81% of the professors did not know about the oecd recommendations, compared to 61% of the swedish professors, it could be encouraging from the swedish perspective. nevertheless, time has passed between thses surveys and perhaps have the finnish professors got more informed than they were in 2006. a conclusion based on our results is that it seems important to raise issues of guidelines in social sciences and humanities concerning accessibility to digital research data and to engage researchers and relevant authorities in creating arenas for discussing and shaping research infrastructure for the future. according to the researchers, the universities, university colleges, the swedish research council, and the swedish council for working life and social research are the most important key actors in promoting accessibility to digital research data from public funding through participation in shaping guiding principles. the most effective interventions for enhancing accessibility to digital data that were identified were that research grants should include funds for preparing the data for sharing and archiving and that making data accessible for the use by the scientific community is acknowledged to be of scientific merit. more education about life cycles of digital data and research ethics were expressed as needs. considering archiving practices, use and reuse of digital research data, 16% of the swedish professors stated that the use of digital empirical data is unusual or are never used. eighteen percent of the finnish professors reported a similar amount of digital data non-use. when comparing between the countries what happens to digital data after analysis and reporting, it seems that it was more common for finnish professors to keep digital data without any further actions to documentation (56%) compared to swedish professors (46%). data are destroyed to a larger extent in finland (20%) than in sweden (3%). however, the saved data is reused by the researchers themselves to a greater extent in finland (94%) than in sweden (54 %). the opinions about amounts of reusable digital data differ also; 50% of the swedish professors stated that more than half the amount of produced digital data is reusable, compared to 21% of the finnish professors. in analyzing responses to important reasons for not reusing digital data, it appears that the swedish researchers emphasize ethical, juridical, technical aspects, and quality of data as more problematic than the finnish researchers. again, it seems that there are a considerable amount of issues that need to be clarified and solved in order to develop a wellfunctioning research infrastructure with a high degree of re-using practices within social sciences and humanities. as information of importance, we think that the researchers’ beliefs that promoting accessibility of their own data to iassist quarterly winter summer 2008 39 be acknowledged as a scientific merit and that research grants should include funds for preparing the data for sharing and archiving, points out legitimate measures with both force and enticement like the stick and the carrot. the last mentioned intervention was also one ranked high by the finnish professors (80%), but their top priority of effective interventions was establishment of guidelines and principles by the finnish universities together (84 %). the swedish professors also point out other obstacles to sharing digital data, and regard deficiency of resources for researchers to document and arrange their data for reuse as the most difficult obstacle to sharing digital data together with lacking guidelines to documentation, while the finnish professors reported that it was the situation when the respondents were not informed that their contributions should be used in the research society generally. they share this concern with the swedish doctoral students. these results relate to the results mentioned in the former paragraphs, namely the need of different types of guidelines (ethical, technical, and juridical), earmarked resources to documentation and education in this area. as we mentioned in the results section we found that the professors seems to be more eager to share data than the doctoral students. a large proportion of the total group also expressed doubts in sharing data, probably because of uncertainty and lack of sufficient guidelines. the finnish questionnaire did not have a pushing question like we had, but the finnish professors were asked their attitude to open access to digital research data collected in their own research and 76% of them expressed positive attitudes. one might conclude that professors in social sciences and humanities in sweden and the finnish professors differ a lot in opinions about digital research data. however, two years have passed with increasing focus on open access issues in research policies in these countries. it would be interesting to see if the finnish professors have changed their minds since 2006. about our own results, it is always interesting when the research is surprising. we were surprised that the professors were more positive and humble towards sharing and promoting accessibility to digital research data than were the doctoral students. but on the other hand, being a doctoral student means a lot in terms of keeping on one own’s track, concentrating on the ph.d. work, and having little time to orient oneself to ventures, research policies, and university practices. in conclusion and in spite of many prejudices about “conservative” professors, it seems that one has to acknowledge their positive orientation about e-science and put forward these survey results of barriers and opinions to be able to support and realize sharing of digital research data in the future. at last, the results of these surveys have to be acknowledged and seriously taken care of in understanding the obstacles and challenges we face, in order to achieve a sufficient and approved research infrastructure, adapted for the distinctive features of social sciences and humanities with their wide range of empirical materials. 5. acknowledgments the authors wish to thank the doctoral students and professors who participated in the surveys. references axelsson, as. & carlhed, c. (forthcoming). next generation e-researchers: doctoral students in social sciences and humanities in sweden and their attitudes towards open access and open repositories. paper presented at ncess national centre for e-social science, the 5th international conference on e-social science. cologne, germany, 24th 26th june 2009. database infrastructure committee (disc). (forthcoming). juridical conditions for a research infrastructure. högskoleverket. (swedish national agency for higher education). http://www.hsv.se/statistik/statistikomhogsklan/ personal.4.6df71dcd1157e43051580001770.html kuula, a. and borg, s. (2008). open access to and reuse of research data – the state of the art in finland. finnish social science data archive 7, 2008. http://www.fsd.uta.fi/ julkaisut/julkaisusarja/fsdjs07_oecd_en.pdf oecd. (2007). principles and guidelines for access to research data from public funding http://www.oecd.org/ dataoecd/9/61/38500813.pdf the swedish research council. (2007). the swedish research council’s guide to infrastructure. http://www. vr.se/download/18.76ac7139118ccc2078b800011940/ rapport+5.2008.pdf notes 1 carina carlhed, mälardalen university, sweden, email: carina.carlhed@mdh.se, and iris alfredsson, swedish national data service, sweden, email: iris.alfredsson@ snd.gu.se. 2 http://www.hsv.se/statistik/statistikomhogskolan/per sonal.4.6df71dcd1157e43051580001770.html. in this paper, comparisons have consistently been made between statistics from the swedish national agency for higher education for year 2008 and the background information about the participants in our surveys. abstract this paper is a summary of a panel session which consisted of four presentations by individuals from or affiliated with the uc curation center (uc3) at the california digital library keywords: uc3, data lifecycle, dmptool, dataup, data one.. uc3: developing tools and services for the the first presentation, given by carly strasser was originally scheduled to be given by patricia cruse. it covered the basics of the current data management landscape, and how the uc3 group at the cdl is working to meet the data management needs of uc libraries and researchers. the presentation also described the suite of uc3 tools and how each fits into the research data life cycle. this presentation laid the groundwork for all following presentations since the three tools described (dmptool, dash, and dataup) are all part of the uc3 suite of offerings. dmptool the second presentation was given by marisa strong who provided an overview of the dmptool, which was recently updated thanks to funds from the alfred p. sloan foundation and the institute of museum and library services. strong described the new features of the tool, including an administrator interface, and walked attendees through how to create custom templates for researchers at their institution. she also showcased new offerings on the dmptool website, including a library of public plans, general guidance for data management, and a suite of resources for promoting the dmptool. an audience member asked if the dmptool connected to or suggested repositories to those submitting plans, and strong and strasser answered that although this would be useful, the difficulty of enabling this functionality has not been explored. the dmptool does, however, provide references to re3data.org and databib.org, both searchable registries of data repositories. dash the third presentation was also given by strasser, and described the uc-wide dash project. dash is an application that provides an easy, self-service way for researchers to publicly share their datasets. the major functions that researchers can perform using dash include uploading datasets, providing datacite metadata for those datasets, obtaining an identifier, and publishing the data so that it is accessible to the public. dash began as datashare, which was a collaborative project with uc san francisco. the uc3 group has since explored the expansion of dash so that each individual campus in the uc system may have their own locally branded version of dash. dataup finally, susan borda of uc merced presented the dataup tool. dataup is a complete system for researchers to upload, describe, and share tabular iassist session 5s summary: tools and services for supporting research data management, june 5, toronto, ca by carly strasser1 re3data.org databib.org 8 iassist quarterly 2014 iassist quarterly datasets via a web-based application. it is openly available for anyone to use, and is affiliated with the dataone repository, oneshare. borda described the first version of dataup, and then provided an overview of the new version, including highlighting improvements such as expanded sign-in options. discussion several audience members asked questions about connecting dash and dataup to local repositories at their institutions. borda mentioned that a developer working on dataup had begun to explore connecting dataup to dataverse, but had not completed the project before their term ended. notes 1. .carly strasser, california digital library, carly.strasser@ucop.edu mailto:carly.strasser@ucop.edu iassist quarterly introduction an archivist's challenges: adapting to changing technology and management techniques over twenty years ago, the national archives of the united states embraced the concept that automated records were actually records which could be considered permanent within the meaning of the federal records act and set about collecting them. since then it has confronted problems incident to finding these automated records, acquiring them, preserving them and making them available to the public. previous papers have discussed access to public automated records in the normal sense; that is, the ability of the researcher to get at them. in this paper i wish, however, to discuss the national archives' acquisition process as a form of access. by donald fisher harrison' national archives and records administration washington, d.c., united states of america this paper addresses three threats to the acquisition of machine-readable records: the threat of an onslaught of hardware and software incompatibility, the threat of discontinuity within textual records series brought about by end-users with microcomputers and the threat brought about by new management techniques from the paperwork reduction act of 1980. archivists ought to view these threats as challenges. when overcome, the challenges will have presented the archives with the opportunit>' to create a better collection of automated records. 'presented at lassist/ifdo internationa! conference may 1985, amsterdam. ** the views expressed in this paper do not necessarily correspond with those of the national archives and records administration. i wish to acknowledge that some of the material has been developed out of long standing collaboration with fellow archivist dr. w. jon heddesheimer, but any mistakes in concept or fact are mine. software and hardware dependency the first challenge to the national archives is well publicized and needs no significant introduction in this treatise. the archivist of the united states, confronted with the research community's complaint that valuable data were being created by federal agencies without any consideration for their preservation or dissemination to the public, established in the spring 1986 (assist quarterly 9 1960's the forerunner of today's machine-readable branch. this branch was given the task of inventorying federal data bases and deciding how best they should be preserved for posterity. we accessioned a number of machine-readable data files created in the 1960's. some of these files were dependent on other outside factors and could not be read on their own. three examples of software and hardware dependency illustrate our initial problems. the first example came early in our organizational being. we received over thirty-five machine-readable data systems from the office of the secretary of defense and the office of the joint chiefs of staff. these systems were encoded in a data base management system called the national information processing system (nips). they caused serious problems in access and handling and a considerable backlog in the accessioning workload. nips was devised for generalized file handling using languages designed to support user requirements in six components. it afforded any data center the capability of reporting long and involved statistical manipulations on extremely short notice to a variety of users. however, the software was compatible only with ibm computers. the presence of nips files suggested serious difficulties in providing a uniform reference service to researchers and brought up the whole question of software dependent files. to retain the files in nips would constitute a precedent since researchers by and large preferred to use their own utility software, transportable files would afford a range of options that encoded files would not last but not least maintaining large inventories of software would add to the preservation costs and require more shelf space. for all these reasons we decided to decode the files. it appears now, with hindsight that despite the fact that these files were unique and very valuable, we should have insisted that the material be transportable before being accepted by the national archives. the second example was the national archives' accessioning of a microfilm series of records containing pictures of captured north vietnamese documents. these were filmed in saigon during the war by the combined document exploitation center on 94 oversize (13-inch) rolls of 35mm microfilm, each roll 1(xx) feet long. the documents were on one side of each frame, with digital bar codes on the other side to provide indexing and control information. soon after we received the microfilm we discovered to our chagrin that the material was hardware dependent in a system known as "file search." four configurations of this machine had been manufactured and sold to federal agencies in the 1960's. the last model (generation four) had a small computer in it it could therefore provide a printout by reading the bar code on the film strip, transferring it to magnetic tape, which in turn could be manipulated and dimiped on to paper. the machines cost $250,000 new and were used only by military and intelligence agencies, as far as is known. it was only after this information was made available to us that we discovered that other file systems were known to exist in this environment and were equally unreadable without any machines in existence to retrieve data. these included some important files in the navy sea systems command (in arlington, virginia) and the navy oceanographic command (in bay st louis, mississippi), including the defense intelligence agency. recently we have discovered the existence of an intact file search model in salvage channels. we have requested that it be turned over to the national archives, and we think we have the technical expertise to restore the model to operating condition. spring 1986 10 iassist quarterly the third example entailed the 1960 decennial census, offered to the national archives by the census bureau in the mid 1970's. these records were created by a univac ii-a computer, of a generation that had been effectively phased out of use in the federal government after the tapes had been created. it has been reported that once the tapes became available for transfer to the national archives, only two such machines existed to read them, one in japan and one in the smithsonian institution.' eventually a reasonable approach was agreed to by the national archives and the census bureau, to convert the data into a compatible format, making them available for preservation in our vaults. these three examples are illustrative of the long term problems aeated by hardware and software dependency of records created in the 1960's, when computers were maintained in relative isolation from each other. it was a period in which data managers were concerned with the ctcation and the use of computer products and were by and large ignorant of the long term value of these products as federal records. it can be said that the letter of the law — the fact that the tapes were handed over to the national archives — was carried out the fact that the tapes were unreadable because of software and hardware dependency was a new problem that had never been faced with paper acquisitions. for their part, agencies were understandably reluctant to dispense funds solely for the benefit of depositing these records in the national archives. thus reason has had to prevail in our dealings on transfer of the tapes, and no one solution can be applied in all cases. small computers and office automation the second challenge to the smooth flow of records into the archives stems ironically from the very machine meant to facilitate administrative operations in the modem office. for several years, most federal agencies have been extending the advantages of their word processing pools by placing terminals at the hands of management officials, giving fingertip control to their own records creation. office automation (ao), more aptly termed "the paperless office", is based on a series of compatible, menu-driven programs utilizing a centralized data base for common shared-use data and unique smaller databases for individual users. these systems have the ability to transfer data and information between data bases through a network or a distributed system. the advantages of such a system are obvious. federal managers frequently need information suddenly and immediately, and often the demands for this information come after the staff has left for the day or the weekend. managers would like the ability to search for the data or reports they need through an indexed automated bibliographic/numeric data base, access and use the appropriate software to perform simple to moderately complex analyses of this data (e.g. forecasts, conelations, etc.), use graphics to illustrate their results, access word processing/office automation tools to produce a memo in the appropriate format, and finally send this report/memo electronically to the recipient's office, all without the necessity of using the phone, typewriter, or staff that are not avjulable. 'commiaee on the records of government, report washington, dc, march 1985, pp. 86-87. keeping all this in the system can cause an archival "log jam". the designers and the users of paperless office systems are frequently ignorant of the paper systems they are replacing and the archival need for intellectual continuit>'. outside contractors compound the problem. in spring 1986 iassist quarterly 11 the absence of any other information the hardware and software dependency problem has reemerged in the small computer world, and agencies are finding that transportability cannot cross the boundaries between offices. software now provides end users with ultimate fingertip access. this allows handcrafted programming and instant manipulative gratification. the same person who creates data on the system can now dispose of it with equal ease. by closing the gap between the user and the machine, the system eliminates the apparent need for the data middleman, lo say nothing of the records manager who, under other circumstances, looked after standardized formats, ensured traditional records disposition practices and provided for a continuity of records series in the agency. thus the danger inherent in oa is that the practice concentrates on the information as it is used immediately after creation without making a record of actions taken. it is said to parallel the dangers of telephone use when first introduced. with that instrument, mjinagers needed go through no intermediate device for communication. telephones assured privacy of communication and were sheltered from the public record. the comparison with oa is evident just as managers could converse at the push of a telephone button, so they do now with electronic mail. further, if one of the parties is absent, there need be no callback, because the mail has already been delivered electronically. like the telephone, the oa challenge is to find a way to record the communication. with the small computer, software must be devised to ask the user for a determination of the ultimate value of the information before it is ever keyed into tlie system. this software has been integrated into the planning for oa systems in most federal agencies. whether or not it solves the problem in practice remains to be seen. information resources management the third challenge to a smooth transfer of records to the archives now comes in the form of an application of new management techniques to the cteation and the use of information within the federal establishment. this new methodology typically accommodates the reality that government must function with less personnel and with individuals of lesser skill and training by altering the way agency missions are carried out the paperwork reduction act of 1980 was rightly concerned with a problem that had existed for some time in that the federal government was preoccupied with the physical problems associated with the large volume of paper records created. the authors of the bill reasoned that managers should have been concerned with how the information was being controlled and how it could be shared with the maximum number of sources. thus the new law espoused intellectual control vis-a-vis physical control, regardless of the medium on which the information had been stored. in order to do this a number of organizational changes have taken place in federal agencies, each a bit different from the next, in which an "information czar" has been placed at the highest levels to control access and dissemination of all information, regardless of the medium. this new arrangement has now been entrenched for four years. a typical arrangement has been established to combine the former functions of "automation, communications, office automation, records management, publications, audiovisual activities and other information activities, services and facilities." an information management plan is usually mandated beginning with a problem analysis, designing a model information system, constructing the "architecture" which produces a program and provides guidance for a budget request under this concept, every information system will have centralized management the spring 1986 12 iassisl quarterly "single manager" concept has been extended to encompass all information, defined as "... all processes by which the user may receive, display or project desired information... (including) voice, text, graphics, audiovisual, video teleconferences, micrographics, files, records management, optical discs and other forms of published information."' in many ways, the single manager system makes a lot of sense. the information manager is in a unique position to disseminate information within an agency to avoid duplication of eftort — or better, to avoid disparate and conflicting data creation. by being organizationally placed at the highest level, the irm provides information for important decisions and controls a sizeable portion of the agency's budget furthermore, the concept will ease the path of liaison between the agencies and the national archives. as we began to accession records in machine-readable form in the 1960's, we became increasingly aware of the presence of the data manager as a viable records aeator and manager. between 1961 and 1980, the machine-readable branch frequently communicated with the data manager directly when it was not able to get required information any other way. furthermore, in the first half of this decade, we became more and more . concerned with deahng directly with government managers since they were creating (and destroying)' information without acknowledging either the federal records act or their agency records administrator. with the advent of the irm principle, however, the archives need only deal again with one official, who, if properly briefed on the urgency of the problem, will coordinate the actions of the records manager, the data manager and the end user. 'draft army regulations 25-1 and 25-5. 1984. conclusions technology has created new solutions to old problems, but in the process has itself aeated new problems. the archival community is thus confronted with unique challenges to its traditional role as keeper of the records, which requires our attention. some measures come to mind as actions to stem the tide. first, the archivist must keep professional pace with the proliferation of computing technology, not only as it is practiced in federal agencies in this decade, but also as many writers envision that it might be practiced 25 years from now. reading the literature is not enough. it requires a shrewd selection of educational services and an on-going dialogue with other archivists. this must include at a minimum the study of software, hardware and storage media as trends develop. an archives must be capable not only of receiving machine-readable records in various modes and written on various media, but also of serving its users with a multiplicity of anangements. this leads to the second measure. the archivist must determine far enough ahead in time in what mode and on which medium these new records will appear as candidates for acquisition. to do this, archivists must assert their professional needs to the creators of records throughout the life-cycle of the records. furthermore, the requirement to deposit tapes and other media in the national archives should be anticipated and budgeted by federal agencies. third, the archivist must reach end users by some means, to ensiu-e standardization of practices and procedures. it is vitally important to overlay records management practices on the uses and outputs of small computers and of office automation systems. this might include commimicating with procurement ofticers and spring 1986 iassist quarterly 13 irm officials to standardize hardware and software packages which would be interchangeable within and between federal agencies. the fourth, and by no means the least important, point is that the irm developments in federal agencies, formed as a result of the paperwork reduction act of 1980, must be influenced by direct communication with archival officials. irm managers have been imbued with the immediate needs of the agency information program in mind — the here and now concept there is always the danger that not enough planning will be conducted for the ultimate fate of records. by the way they maintain certain modes of information, irm officials can influence the disposition, and in turn, the configuration of future holdings of the national archives." records acquistion liason for flow of records between the national archives and federal agencies /' adminlstftaton before i960 x records \ '' wichivibt ' ftominietratofl^ mdp 1 i ctio / \_ maiiacer user / \^ y —— _ 1 -"''^ 's ihfortution j \ re'sdurces / \ rw-i^ger / 1981 1985 spring 1986 iassist quarterly iassist quarterly 2013 9 abstract this paper is a summary of three presentations given on metadata and related topics at iassist 40 in toronto. some of the text is adapted from the presenter’s published abstracts; other parts were contributed by the session chair. keywords: metadata, identifiers, orcid, doi, datacite, blaise, mqds, xml, ddi, best practices, open journal.. using identifiers to connect researchers, authors and contributors with their research data ithe first speaker in the session was elizabeth newbold from the british library who was filling in for a colleague who had proposed the paper but then took another position elsewhere. she gave an update on the orcid and datacite interoperability network (odin) project which was the subject of a session at iassist 2013. the project is collaboration between the british library, cern, orcid, datacite, dryad, arxiv and the australian national data service with the aim of using persistent, open and interoperable identifiers for people and for datasets to connect researchers, authors and contributors with their research data. the presentation outlined some key results from the first year of the project including the proofs of concept (pocs) in the humanities and social sciences (hss) and high energy physics (hep). they faced challenges in a few areas: access, discoverability, interoperability and sustainability. the project allows for the identification of contributors as well as authors. for the pocs, there was an extreme dichotomy between the hss and hep communities: there can be as many as 100 names associated with a paper in hep and in some cases an entity and not an individual is associated with a paper. these practices are quite foreign to the hss world. the project is, however, looking to outline commonalities between the extremely different disciplines.. the first year of the project centered on building the conceptual model for connection creators, curators, contributors, and data sets. this approach was more straightforward for new data being cataloged but much harder for data already being held. people are even harder to retrofit: assigning identifiers to inactive researchers may be problematic, especially if they are dead. and when researchers change institutions, assigning a new id to a researcher that affiliates them with their current institutions causes problems if the institution they were at when the research was done still wants credit. the second year focus was on identifying generic workflow and how to use it as a framework for implementing workflows for assigning dois and orcids. when assigning metadata, there is no iassist session 5p summary: big picture metadata, june 5, toronto, ca by san cannon 1 these best practice descriptions will be modular with a homogeneous format, allowing reorganization in multiple ways 10 iassist quarterly 2014 iassist quarterly equivalent to the data documentation initiative (ddi) in hep. in order for the assignment of orcids and datacite to work well, there needs to be interoperability. the project includes a tool for claiming datasets within your orcid file. work is also being done to link international standard name identifiers (isni) to orcids. other work includes interacting with many stakeholders, including funders and policymakers. another update is planned for after the fourth plenary of the research data alliance (rda) in september 2014. the humanities and social science proof of concept report (http://dx.doi.org/10.6084/m9.figshare.824317) by john kaye (bl), tom demeranville (bl), steven mceachern (ada) was published in july 2013. rich metadata from blaise the second presentation was by beth-ellen pennell of the institute for social research and university of michigan, to which gina cheung also contributed. this presentation focused on the process and challenges faced during the harmonization and preparation of the metadata and data files of the collaborative psychiatric epidemiology surveys (cpes) http://www.icpsr.umich. edu/cpes/index.html. the cpes joins together three nationally representative surveys of adults living in the united states: the national comorbidity survey replication, the national survey of american life, and the national latino and asian american study. these data were collected face-to-face using the blaise software, a product from statistics netherlands often used by statistical agencies and others. the michigan questionnaire documentation system (mqds) was used to extract metadata from blaise using ddi standards. this system is free to blaise users and allows for data transformation from the blaise format to other such as sas, spss, and sql. in this case, the blaise data were transformed into xml capturing the rich metadata available in blaise data models. the combined cpes dataset contains approximately 20,000 interviews. the initial combined dataset had 9.400 raw variables distributed over 92 sections of the three surveys which needed to be harmonized across the datasets. a survey instrument crosswalk was needed because the same instrument wasn’t used in each survey. the sections appeared in different orders and even within a section, question order may vary. initial cleanup of the data was required before documentation could be created. the final dataset contains approximately 5,600 harmonized variables, 400 constructed variables and 14 separate weights. the website contains rich metadata including an interactive cross-walk of all harmonized variables with question text in 5 languages, response options, missing data codes, descriptive statistics (frequencies, etc.), universes, detailed documentation of all constructed variables, and descriptive statistics of all variables, among a wide variety of other products. future work includes a move to ddi3.2 to cover the full survey lifecycle and dealing with the release of blaise 5. ddi handbook – overview and examples of recommended best practices the final speaker in the session was joachim wackerow from gesis – leibniz institute for the social sciences. this discussion introduced the ddi handbook project. the use of ddi is increasing both the number of users and producers of ddi materials as well as the number of new projects at gesis that use ddi. the use of ddi is heterogeneous and many users find ddi to be complex because the subjects are complex. there is some documentation already available but coverage is incomplete, some documents are outdated, and for others the understanding has changed. building upon previous efforts at various ddi workshops, the project plans to produce a collection of best practices on using ddi using a community approach. the goal is compile a set of best practices aimed at a broad audience so that the output is useful and accessible information providing a balance between a book and a list of frequently asked questions. these best practice descriptions will be modular with a homogeneous format, allowing reorganization in multiple ways. the primary structure for the collection will be organized in alignment with the ddi lifecycle. a goal will be to involve the ddi community in producing a shared body of resources for all organizations and individuals using the ddi specification. the format is to create an independent open access journal that is published twice a year in coordination with, but not published by, the ddi alliance. the platform will use the open journal system which provides online presentation and editorial management workflow. the system will provide a structured paper template and publish under the creative commons sharealike license. multiple formats will be available for reuse (docbook or dita) and additional materials can be provided as xml files. the submitted best practice documents will be reviewed by a team of editors and reviewers and published on a dedicated website. the editorial board is now being formed with 4-5 members who will set policies and make final decisions on papers. there will be an open peer review process, partly because the ddi community is small and partly because open peer commentary could result in a new article or new version of the original article. in addition, the content could be used as the basis for tutorials or other teaching materials. some initial topics may include guidelines for archives introducing ddi into their workflow and other institutions already using ddi codebook and shifting some of their workflow to ddi lifecycle. another area of interest will be utilizing ddi for data discovery. the project is looking for outside involvement from the community in the form of comments, papers, and reviewers. one audience member asked why the user community had to write their own documentation instead of having experts write it for them. the discussion centered on how the broader community can provide a larger selection of use cases from which others can learn. a suggestion to include “what not to do” or “lessons learned” articles as well as best practices was well received. notes 1. .san cannon, deputy chief data officer, federal reserve board, washington, d.c. 20551 sandra.a.cannon@frb.gov http://dx.doi.org/10.6084/m9.figshare.824317 http://www.icpsr.umich.edu/cpes/index.html http://www.icpsr.umich.edu/cpes/index.html mailto:sandra.a.cannon@frb.gov vol31_1.indd iassist quarterly spring 2007 by by taina jääskeläinen and tuomas j. alaterä* multilingual web services of data archives: possibilities and pitfalls our article discusses the kinds of issues data archives and other data providers need to consider when setting up multilingual websites. the article, based on a presentation at the iassist 2007 conference (jääskeläinen and alaterä 2007), draws from our experience at the finnish social science data archive (fsd), which provides web services in three different languages, and from the examination of other european data archive websites. first things first when planning a multilingual website, we recommend that the purpose and scope of the site be discussed and decided first. website developers need to ask themselves the following questions: why are we making the website multilingual, what kind of functions or information do we wish to offer, who will use the service, and what are the users’ needs. having web services in multiple languages means more work, so they must also consider how much money and staff resources can be allocated. organisations often have to balance between maintaining an adequate level of services and having too much information to update. clear decisions regarding goals and content should be the first step in the planning process. for example, the fsd provides web services in finnish, english and swedish, but has different goals for each. the language of the main site is, of course, finnish. since five percent of finns count swedish as their mother tongue, we also offer rudimentary web services (such as information on how to order and deposit data) in swedish. our english website is fairly comprehensive in order to provide services for all nonfinnish speakers. one of our main goals is to provide information on data, therefore we have study descriptions in english for all archived quantitative datasets. roughly 80% of the web pages in finnish have a corresponding page in english, though the content is not identical. priorities for data archives when we look at data-rich, multilingual websites from the user’s point of view, we must ask ourselves what are the main reasons for users to visit such sites? it is a fairly safe bet to assume that most users will be looking for data and information on data (hansen and richardson 2006). however, in our experience browsing the multilingual websites of european data archives (mainly the english versions), it sometimes seemed easier to find information on the archive than on data! we found, quite often, that links from an english page led, without any prior warning, to a web page in another language, often to the data search or data catalogue page in the dominant language. this can be considered a major flaw in web design as well as in usability. an arrangement like this can work if many users have at least a passing knowledge of the dominant language; however, when it is reasonable to assume that foreign users do not understand the dominant language at all (as is the case with finnish), some additional information is needed. we recommend having a declaration of content that provides details on the web services that are available in each language. this declaration will prevent users from having to browse through the whole site in order to find out what services and information are available in a particular language. the declaration need not be long but it must be to the point. a couple of well-placed sentences will most likely be sufficient. since many users of multilingual websites are looking for data, the declaration of content for the non-dominant language pages should answer the following questions: 1. is it possible to search data in that language? 2. are there study descriptions in that language? 3. if there are no study descriptions, who is the contact person for information on the data? even without the declaration, users will appreciate finding answers to these questions quickly and easily. if registration is required at any point, providing information on who can register is a good idea. if it is available for all, say so. what may be self-evident to those who are setting up a website is often much less so to users. let us assume that the user has found data that are relevant to him/her. now he/she needs to know how to proceed. a 6 iassist quarterly spring 2007 functional website provides adequate and easily-accessible information on how to order data, for example, by providing a link from the study description page to ordering information and order forms. users need to know whether (and how much) it may cost to order data, whether the archive can provide translation if requested, and how much the translation will cost. if the archive does not provide data translation, stating this will make the situation clear to the user without further enquiries. if the data archive does not have the resources to provide study descriptions in the non-dominant language(s), one alternative is to translate the data search interface, including ddi field names. this may be useful for users who have a passing knowledge of the dominant language. if they have access to the search interface in their stronger language, they may be able to use dictionaries or multilingual thesauri to find search terms for study descriptions in the dominant language. we also recommend providing detailed contact information. that is, more than just one email address and telephone number for the whole archive. users appreciate knowing who does what and who they can contact regarding a particular issue. it is worthwhile to carry out a user test before launching a multilingual website. a simple and quick one may be sufficient. recruit a couple of people not familiar with the creation process of the site and ask them to spend up to an hour trying to find, for example, information on data on a particular theme. have a staff member sit beside the tester and take note of his/her comments. it is crucial that testers try to use the services offered on the website and not simply review how the site looks. that being said, a critical look at colour etc. may also be useful. in all probability, the hours spent on this type of simple testing will turn out to be very useful for the eventual functionality of the site (alaterä 2000). about translation when considering translation of a website into another language, it is important to remember the differences that may exist in a particular language that is spoken in more than one country. for example, at the fsd we had to consider whether we were designing our web pages in swedish for users living in finland who count swedish as their mother tongue, or people living in sweden, or both. people living in another country may not be familiar with your society and systems. this fact is particularly relevant when dealing with abstracts in study descriptions, as some terms related to health care, taxation, educational systems, etc. may not be understandable to users from other countries, even if they speak the language. it may be better to globalise some terms instead of merely translating the abstract as it is. the procedures for ordering data may also be different for researchers working in your country and for researchers working in another country. ordering information should clearly reflect this. the people who are paying most detailed attention to the content of a multilingual website are the translators. therefore, if translation is contracted out, the translators should be well-informed as to the goals and target audience for the site. this will make it easier for them to decide when the content needs to be adapted to fulfil the goals (nielsen 1993, 242–245). even if the person doing the translation is a staff member, process writing and teamwork is always useful. we suggest pre-writing, consultations with other staff members to discuss the result followed by revision, user tests, and further revision. many heads are better than one. easy navigation clarity of navigation is a must, whatever the language. it is a well-known fact that internet users are impatient. it has been said that “if a man from mars doesn’t figure out your navigation in four seconds, your web page sucks” (flanders). even though somewhat provocatively expressed, these are wise words and well worth keeping in mind! (if you would like to see other web design tips like this, go to http://www.webpagesthatsuck.com/.) two main pitfalls in navigation are making the navigation meet the needs of the organisation rather than those of the user and assuming that users know more of the organisation and its services than they in fact do. if the website does not provide one specific link to the data catalogue, but has several links to different types of surveys (for example, different survey series), it may confuse users. they will have to go to the page of each different type of survey and try to figure out how they are categorized and what the content is. navigation problems like these are easily revealed by a simple user test. again, some kind of declaration of content will be beneficial (hansen and richardson 2006). regarding links between language versions, frequently used practices are the best (w3c working group 2007). we recommend putting language links where users expect to find them and always in the same place on every page. if the language links are flag-based, adding or replacing these with text (english, español, deutsch) is strongly recommended. people do not remember flags, and, in the case where one language is spoken in several countries – which flag would you choose? when providing links from one language version to another, the optimal solution is to provide a direct link to the corresponding page in another language. however, this is often difficult or impossible since the content of different language versions differ. it also means more work. linking to the home page in the other language is the most common iassist quarterly spring 2007 7 solution, which again stresses the importance of a clear and understandable navigation design. users will be keen to get back to “that one important” page instead having to test their search skills trying to figure out where they might find it. this task is made easier if bread crumbs are used on each page and, if possible, each language version should be given similar structure (nielsen 2000, 325–331). web design we recommend that plenty of time be dedicated to discussing web design. concentrate on achieving your desired goals. web design should be based on the goals and target audience chosen for each language version. what functionalities might target users expect (felke-morris 2006)? in all web design it is important to remember that internet users are likely to be impatient. avoid splashes and unnecessary animation, especially on the index page of your website. has anyone ever seen an animation that they would like to see again, again and again? it is advisable to think twice before choosing to rely solely on for example flash techniques or additional plug-ins. difficult or nonconventional navigation (which means that users get lost and frequently have to return to the main page) combined with time-consuming animation on the main page usually results in very frustrated users. this type of design does not work for information-rich websites where it is crucial to be able to find a particular piece of information without too much effort. this is why popular web services like google rely on a simple and fast-loading front page. applying some kind of a template system reduces the burden of running websites in multiple languages. when changes are needed in navigation or in other fairly constant elements of the site, templates make life much easier. fonts, colours, titles, navigation links and so forth are controlled by one style sheet and one design template, meaning that only one alteration per language is sufficient. using templates also tends to have the effect of pushing the web design toward an easily-maintained form, which is good for everyone. it is a good practise to put navigational elements (search interface, database field names, etc.) into language packs, and to keep them as separate as possible from the text content of the site. this is preferable to coding these elements into the source code directly. with language packs, it is easier to translate functionalities like navigation, search interface etc., and to make changes in all language versions simultaneously or even to add completely new language versions. when a website is up and running it should not be left on its own. checking the log files periodically gives valuable information on where users are coming from, how long they spend on the site, if there are pages that are more popular than others, and which pages (if any) tend to drive people away to information sources outside your site (rasmussen 2007). analysing search logs helps to find out how and what people are searching for, and whether they are looking for information that does not exist on your site, or whether they are looking for existing information but in a wrong place. log file information can be used to improve your web service. conclusions the main point of our discussion is that multilingual websites should be designed as if they were separate websites developed in separate languages, rather than multiple translations of a single language site. what we are talking about here is creating new websites in other languages. what the content and actual technical solutions are depend very much on the goals set for each site. decide on your goals, be clear about what services you are providing, and never lose your users to difficult navigation. you might ask do we at the fsd have optimal multilingual web services? the answer is probably not, but we did recently make changes to our multilingual websites in order to follow our own recommendations. there is always room for improvement! * this article was originally presented at the iassist 2007 conference. the authors taina jääskeläinen and tuomas j. alaterä work at the finnish social science data archive, 33014 university of tampere, finland. contact: taina. jaaskelainen@uta.fi references alaterä, t. j. 2000. ”www-suunnittelun periaatteita” (lecture presented at the institute for extension studies at the university of tampere). felke-morris, t. 2006. web development & design foundations with xhtml. 3rd ed. boston: addison wesley. best practices checklist also available online at http:// terrymorris.net/bestpractices/index.htm. hansen, s. e. and m. richardson. 2006. evaluation of web sites: what works and what doesn’t. paper presented at the iassist 2006 conference, ann arbor. also available online at http://www.iassistdata.org/conferences/2006/ presentations/c2_hansenrichardson.ppt. jääskeläinen, t. and t. j. alaterä. 2007. multilingual web services: possibilities and pitfalls. paper presented at the iassist 2007 conference, montreal. also available online at http://www.edrs.mcgill.ca/iassist2007/c3_fsd.ppt. nielsen, j. 1993. usability engineering. san diego: morgan kaufmann. nielsen, j. 2000. www-suunnittelu. jyväskylä: it press. 8 iassist quarterly spring 2007 ———. designing web usability. indianapolis: new riders publishing. rasmussen, k. b. 2007. f2: data beyond numbers: using data creatively for research. paper presented at the iassist 2007 conference, montreal. also available online at http://www.edrs.mcgill.ca/iassist2007/presentations/ f2(1).ppt. w3c working group. 2007. internationalization best practices: specifying language in xhtml & html content. world wide web consortium. http://www.w3.org/ tr/i18n-html-tech-lang/. flanders, v. does your web site suck? flanders enterprises. http://www.webpagesthatsuck.com/does-my-web-site-suck/ does-my-web-site-suck-intro.html. vol281.indd iassist quarterly spring 2004 15 by by zoltán lux * this internet-accessible multimedia oral history database (http://server2001.rev. hu/oha/index.html) forms a slice, or cross-section, of the database of contemporary history held by the 1956 institute in budapest. its basis is the oral history archive (oha) established by some of the founders of the institute decades ago. there are about a thousand life interviews, divisible into three main groups. about 500 interviews were made with participants in the 1956 hungarian revolution. many of the rest were done with their children. others were life interviews made under a leadership research programme in 1981–5, with those thought likely to have a career ahead of them in the communist party or state hierarchy. the interviews were tape-recorded. transcripts were made from these recordings, which up to the end of the 1980s meant that they were typed. in the 1990s, the transcription was made using computer word processing, so that the texts have survived in digital form as well. in 1991, to assist orientation among the interviews, we developed a database using 3–5-page abstracts whose information could be searched and recovered in detail. the purpose of this database was to record the interview content as accurately as possible. it was not possible to archive the full texts or the sound materials in digital form because of the memory constraints at that time. in 2003, the institute’s successful application for competitive funding under the evilág project of the ministry of informatics and telecommunications allowed a sizeable proportion of the interview materials to be digitized. this funding will allow the institute to digitize the texts of interviews available only on paper as well as old sound materials. the project involved more than archiving, as the public part of the database is being placed in the public domain via the internet. before the project began, we were developing a database handler into which we could transfer the old oral history database. this oracle-based database can now receive the full interview texts and the sound documents. this database not only assists with making the content and technical data of the interviews searchable, it is becoming increasingly suited to fulfilling the complete archiving function and digital storage of all the related text and audio-visual documents. the oral histories comprise only part of our database of multimedia oral history database 1 http://server2001.rev.hu/oha/index.html http://server2001.rev.hu/oha/index.html 16 iassist quarterly spring 2004 contemporary history. the other elements of the database include the photo archive (some 15,000 digitalized photographs with technical and content descriptions); the historical chronology; the biography archive; the trial documents archive; the video archive; and the bibliography. each of these database elements is closely associated with the interview subjects. there may be a photograph of the person interviewed, a biography of the subject, photographs of events, or a chronological description. ‘private history’ content: ’56 and the kádár period based on these interconnecting database ‘elements’, we produced a content service entitled ‘private history. 1956 and the kádár period’. the content basis for this is a partial set of the interviews in the oha, the basic data in these interviews, and details and sound documents edited from the full interviews. in these interviews, the subjects speak from the viewpoint of people once condemned for their actions during the kádár period. the internet content service also demonstrates the connections between the interviews and persons, perhaps linking an edited biography from the biography archive with photographs from the photo archive. based on the database, we are able to produce a book-like block of information via the internet, whose ‘technical organizer’ is a subset selected for a specific attribute of the elements of the interview set. within this system and based on such attributes, successive new ‘books’ will be possible. main features of the framework system: 1. the system allows ‘almost complete’ digital archiving of the documents, based on oracle database handling software. the internal database maintenance program allows for the loading of the full texts of documents (which will also be archived), digitized files of related audio-visual documents (sound or video recordings), details of the full text edited according to various criteria, and descriptive data on the documents’ content and technical characteristics. 2. it provides an immediate framework system for internet publishing, fine tuning of a grading system for user accessibility (fully public, or full access only to researchers or staff of the data owners) and secure handling of ‘sensitive’ non-public data. 3. in developing the data structure of the documents, attention was paid to offers to do with qualitative data archiving from ddi, dublin core and tei, and it is suitable for preparing xml outputs compatible with further development of the system. problems and solutions during development: 1. the life interviews were conducted in hungarian, which confines them to a relatively small audience. however, the basic data in the interviews and the biographies of the subjects have been translated into english.. we have also translated some of the interviews themselves. of course it would be desirable in the long term for such tasks to be performed by an automatic translating program, but such devices are not yet sufficiently developed for long texts in the hungarian language. native english-speaking translators are being used for the time being and the translations are loaded into the appropriate english fields. (developed translating programmes would confine the native-speaker translators to checking and editing the texts linguistically.) 2. the life interviews contain references to many events. these can sometimes be linked with actual chronological events, but in each case, what appears is a subjective interpretation. the connection with the actual event has to be signalled, but the data discrepancies also need to be recorded. 3. the many connections between the edited interview texts and other elements of the database (photos, biographies, events) should ideally be found automatically by the program among the requisite objects in the databases, allowing for further exploration through these references. some of the photographs have been linked singly to the appropriate references, but these references are also embedded in the database. there is a complex problem with footnotes to interpret and explain the text. in some cases, a section of the text has to be explained, elucidated or possibly interpreted. such footnotes are individual and of no use elsewhere in the text. other sections of the text call for footnotes that expand on a concept, institution, event or person. these can be applied elsewhere in the text as well and often exist already as data records elsewhere in the database (as entries in lists of concepts or biographies or in a chronology etc.) in the latter case, a cross-reference is needed to the appropriate record. this facility has not been created under this project, partly for technical reasons, and partly due to incompleteness in the content. however, the problem has been addressed with data storage between meta-references in the individual footnote texts. it may even be possible to pick these passages out automatically from their texts, during later development (for instance, with a separate data table). 4. some parts of the database (oha interviews, trial documents) also contain confidential information. it was very important for the internal data maintenance module to operate according to a well-defined, graded userentitlement system, so that such information cannot be accessed from the internet in any way. 5. the interview texts and interview subjects provide much data of a statistical nature for qualitative sociologi iassist quarterly spring 2004 17 cal researches. we would like to ensure the possibility of making the data available to various analytical programs partly through requisite exchange formats (with coding if need be) and partly directly. for the time being, direct access is ensured for programs capable of linking directly with the database. (we are experimenting with mineset.) 6. the problem of data archiving. the really desirable solution would be for all objects for archiving to be loaded into the database. however, the databases still contain only ‘simplified’ versions of the digitized photos and sound documents. full-sized digitized photo files with authentic colour can approach 100 mb in size, not to mention the size of long sound documents. these original files are kept at present in external data stores (cd–rom, dvd–rom, dlt). photographs are available within the database as 300 x 300 thumbnails and 600 x 600 viewing-size pictures. the sound documents are stored in highly condensed wma format versions. 7. we intend to install an xml format output for connecting with other databases, for data conversion and for data exchanges. this is not yet ready. we are seeking other institutions to provide similar historical or social scientific content, so that we can develop a modern-history portal service that is as broad as possible or join a similar service ourselves. acknowledgements the inspiration for the online multimedia database of the 1956 institute’s oral history archive (http://server2001. rev.hu/oha/index.html) was the edwards online project (http://www.qualidata.essex.ac.uk/edwardians). special thanks are due to marcia freed taylor, director of ecass, for providing a scholarship in 2003, and to louise corti, associate director of uk data archive, for arranging, during the scholarship period, for me to examine digital archiving and content service practice of the uk data archive and qualidata. no less important to implementing the project were the contributions of the oral history archive staff—the framework system for the digital content service crystallized in its final form after long, intensive and fruitful discussions with them. of the institute staff, thanks are due in particular to adrienne molnár, head of the oral history archive, and to judit m. topits, who acted as the daily operative coordinator and organizer for the whole project. references 1. edwardians online: http://www.qualidata.essex.ac.uk/ edwardians. 2. baker, emma j., and louise corti, ‘edwardians online’. in: iassist quarterly, 26:4 (winter 2002), 5–7. 3. ddi. http://www.icpsr.umich.edu/ddi/org/index. html. 4. kõrösi, zsuzsanna, and adrienne molnár, ‘carrying a secret in my heart... 5. children of the victims of the reprisals after the hungarian revolution’. in: 1956. budapest/new york: central european university press, 2002, 200 pp. 1 this work was done with financial support from the improving human potential (ihp) and knowledge baseenhancing access to research infrastructures programmes. a research scholarship was provided in 2001–3 by the european centre for analysis in the social sciences (ecass). * paper presented at the iassist conference, may 2004, in madison, wi, usa. zoltán lux, institute for history of the 1956 hungarian revolution, h 1074 budapest, dohány u. 74., email: luxz@helka.iif.hu, url: www.rev.hu. http://server2001.rev.hu/oha/index.html http://server2001.rev.hu/oha/index.html mailto:luxz@helka.iif.hu http://www.rev.hu a survey of icpsr member institutions ann janda vogelback computing center northwestern university introduction at the meeting of the icpsr official representatives (ors) held at ann arbor on november 11-14, 1983, a brief survey of member institutions was handed out to attending members. the survey was drawn up at northwestern university by lorraine borman and ann janda (vogelback computing center) in an effort to find out what other universities were doing in planning for a distributed computing environment. the questionnaire asked for information about what services are provided to access icpsr data, mainframe and software usage, current extent of micro usage, and a final open-ended "future plans" question. out of approximately 150 meeting participants, 38 responded to the questionnaire. this report discusses the results of the survey and provides two appendices: (a) an index to universities and their mainframes, and (b) a listing by university describing hardware, software, services, and future plans. (appendix b available from ann janda, vogelback computing center, northwestern university, evanston, illinois 60201.) commentary on survey results by question 1. availability of icpsr hardcopy codebooks: 18 at departmental site 15 at university library 9 at central computer site 6 at some other location the above results show that most institutions housed their codebooks at the department with the university library running a close second. the actual totals to this question ran higher because from the 38 responding institutions, 7 located their codebooks at more than one site, generally with a department/library combination. research institutes largely fell into the "other" category, accounting for 4 of these 6 responses; one research institute identified with a department. 30 computer programs such as database management systems to identify and access datasets? if yes, which? 13 yes 25 no this was probably an ambiguous question. for those who answered "yes", the responses fell into two categories: (1) software that provided tape access to the data or (2) software that identified the studies by title or abstract. no two programs mentioned were alike. the software programs are listed in appendix b. personal assistance in identifying and accessing datasets? 23 at departmental site 12 at central computer site 8 at some other location 7 at university library according to the above results, by far the greatest assistance in identifying and accessing datasets is given in the department. as in question 1, more than one answer was checked, and indeed, 7 departments shared this function with the library and computing center. most of the responses in the "other" category pertained to social science research laboratories or institutes. 4. computer programs for graphic analysis of data? if yes, which? 24 yes 11 no of those who responded 'yes', most used sas/graph. spss plot was another popular package as was tel-a-graf display (spelled 3 different ways on the questionnaires). the graphics programs are also enumerated in appendix b. 5. consulting on statistical analysis of icpsr data: 25 at central computer site 24 at departmental site 9 at some other location 1 at university library the above figures splitting the statistical consulting load between the department and the computing center accurately reflect the number of multiple responses falling into the combined department/computing center pattern. slightly more than 31 half of the institutions provided this combination of statistical consulting. most of the research labs ("other" location) shared this function with both the computing center and the department. how are your icpsr data stored? 37 magnetic tape 20 system files, e.g., spss or sas files \l raw data files disk packs almost all of the member institutions store their data on magnetic tape—probably as they are originally received from the icpsr. in addition to storage on magnetic tapes, more than half convert at least some of the data to spss or sas system files. disk packs also seem to enjoy considerable use. do you offer subsetting services for large data files? 20 yes 17 no the responses indicate that a little more than half of the institutions provide some form of subsetting for their users. although the responses were straightforward, it is not clear whether the services provided only consulting and programs enabling the user to subset the large dataset—or actually "doing the subsetting" for . the user and delivering a smaller dataset based on user specifications. the latter option was intended. is documentation available online? 17 yes 19 no although the responses to this question showed a fairly even split among icpsr members in providing online documentation, some of the additional comments to this question indicated that this may have also been ambiguously phrased. while a few comments referred to availability of "abstracts", a few others referred to the use of machine-readable codebooks and dictionaries on tape. obviously, online documentation could be construed either as study descriptions/abstracts—or as actual data documentation, e.g., codebooks and data dictionaries. while the online study descriptions were the target of this probe, the online codebooks have more interesting and powerful applications. what mainframe icpsr data? 18 ibm 13 dec 8 cdc computer(s) are used for storing and processing 3 prime 3 others 2 univac the above figures indicate that leaders in hardware are ibm, dec, and cdc— in that order. within the "other" category are two amdahls and a hewlett-packard 3000. a more detailed enumeration of specific models by university is indexed in appendix a. 10. if more than one mainframe is used, how are data transferred? in all, there were 12 responses to this question, ranging from the use of decnet, ethernet, tapes, to locally developed software and utilities. specifics are found in appendix b. 11. what mainframe analytical packages are used with icpsr data? mos tly sometimes seldom never spss: 35 2 1 sas: 10 11 1 14 bmdp: 1 13 15 7 osiris: 8 10 18 sir: 4 5 27 other: 1 6 2 1 it is no surprise that spss and sas are the most commonly used statistical packages for the mainframe. bmdp comes in as the next contender with high usage in the used "sometimes" and "seldom" categories. the comparatively low frequency of usage for sir could most likely be explained by its relative newness in the market. among packages named by respondents in the "other" category were minitab, abc, tsp, scss, midas, datatext, and troll. from this group scss was noted as most commonly used by one institution; the remaining packages fell largely into the "sometimes/seldom" category with minitab usage noted at 5 institutions. 12. what use is made of microcomputers in analyzing icpsr data? a great deal 8 some 30 none or virtually none micros have not made much headway as consistently used tools for statistical analysis—yet. although the number of responses in the "none or virtually none" category seems high, a few were 33 qualified by remarks referring to the early arrival of additional micros, pending changes to increase the use of micros, and current use by individual faculty and user groups. also as more statistical packages become available, more use is anticipated. 13. if microcomputers are used for analyzing icpsr data on your campus, please describe how the datasets are downloaded for use: there were a total of 9 responses to this question. two institutions used kermit; the remainder variously responded with comtty, smartterm pc, locally developed software, public domain software, an interactive mainframe program via superwylbur—on through "the problem is currently being addressed." 14. if microcomputers are used, what are the most common machines and analysis programs? from 11 responses, the most frequently mentioned micro was the ibm pc followed by the apple he. ibm-xts appeared among the sprinkling of victors, zeniths, teraks, radio shacks (probably trs-80s), and tl professionals. no clear pattern of software use emerged. comments referred to uses other than statistical, e.g., as spreadsheets and as terminals hooked to mainframe statistical programs. it appears that presently micro stat packages are largely under consideration for use/purchase, e.g., one of the institutions is reviewing statgraphics , statpak, and other programs. 15. are you planning to change your existing method of access and distribution in the next two-three years? if so, how? twenty three respondents answered "yes" to this question. their elaborations mainly pointed to the increased use of micros. for many this increase is linked with networking, workstations, shared storage, and subsetting datasets for micros. a few mentioned a move toward creating a social science laboratory. the complete references to future plans are printed in appendix b. summary: what have we learned? we have reviewed information on hardware, software, and services provided to users of icpsr data at 38 member institutions. for most institutions, the department is the most important site providing services to data users, e.g., codebooks and personal assistance in identifying datasets. the central computing site is next most important, and the library's role 34 is limited largely to storing codebooks . spss is by far the most common program for analyzing icpsr data, and about half the institutions store data prepared for analysis in spss or sas system files. icpsr data are still analyzed primarily through computers at the central site, despite the popularity of microcomputers. indeed, no respondent reported extensive use of microcomputers, and 30 said that they made little or no use of microcomputers at present. although most respondents reported future plans for analyzing icpsr data with microcomputers, the revolution has not yet occurred . appendix a: universities and mainframes index arizona state university auburn university australian national university baruch college, cuny bowdoin college bryn mawr college california state university, fresno carnegie-mellon university central michigan university colby college cornell university florida state university hunter college, cuny indiana university loyola university, chicago north texas state university northwestern university ohio wesleyan princeton university rutgers university ssrc data archive swarthmore college temple university texas tech university trinity college, yale federation university of alberta university of california, los angeles university of florida university of iowa university of oregon university of pittsburgh university of southern california university of utah ibm, model: cdc, model: ibm, model: dec, model: 1020 cdc, model: 170 ibm, model: 3081 ibm, model: 3033 dec * univac ibm dec, model: 10 other: hp 3000 cdc, model: 170 series, 720 locally— 730 and 760 centrally in los angeles dec, model: 20 's and vaxs cdc * we hold tapes and use as needed dec, model: vax 3081 * dec, model: : 760, 730 370 system, 3081 ; vax & pdp 11/44 /855 (most use) * prime ibm, model: 30335 ibm compatible, model: nas 8040 dec, model: vax 11/780 * cdc, model: cyber 170/730 dec, model: vax 750 and two 730' s ibm, model: 3081 ibm compatible, model: nas/ 9000-2 * dec, model: vax 730, vax 780 dec, model: 10 prime, model: 750 cdc, model: 172, 174 (soon, 750) ibm, model: 3033 ibm (at yale) other: amdahl 5860 ibm, model: 3033, and 4341 ibm ibm, model: 3033 * prime, model: 750 ibm, model: 4341 * dec, model: 1051 dec, model: 1099 ibm, model: 360 univac 35 university of washington university of windsor university of wyoming utah state university west virginia university cdc, model: 750-150 ibm, model: 3031 cdc ibm, model: 4341 * dec, model: vax 11/780 (5) other: amdahl 420 books' political terrorism a research guide to concepts^ theories^ data bases and literature by alex p. schmid, centre for the study of social conflicts, state university of leiden, leiden, the netherlands 19 84 xiv +5 86 pages pricei us $40.50 ( in usa & canada) dfl. 95.00 (rest of world) isbn 0-444-85602-1 papeisack approx. publication date 12/84 this extensive handbook surveys contemporary social science thinking on political terrorism. as a reference work it provides the reader with the largest bibliography on the subject--a computer-based, partly annotated, 4000+ item, multi-disciplinary, multi-lingual literature survey covering aspects of theory and practice. apart from regional and country entries, it carries subdivisions on such varied aspects as nuclear terrorism, hostage saving measures, state terrorism, etc. divided into 21 major and 46 minor categories, this author-index bibliography covers legal, psychological, sociological, military and ideological aspects of political terrorism. the handbook also includes a 130-page "world directory of 'terrorist' organizations and other groups, movements, and parties involved in political violence as initiators or targets of armed violence," which has been compiled by a.j. jongman. an 80-page survey of current thinking on the origins of terrorist violence in various contexts provides insight into the sociological, psychological, conspiratorial theories on the subject, and attention is given also to the theories of regime terrorism and those of the terrorists themselves. a further section in the volume discusses and evaluates available data bases for the study of terrorism such as those of the cia, rand, etc. the accessability and the reliability of data on terrorism are discussed and data requirements for social science research are indicated. finally, the present state of the literature on terrorism is discussed, and research desiderata and strategies pointed out. contents: forward by i.l. horowitz, introduction. parts: i. concepts. ii. theories. iii. data bases. iv. literature. for information or to order, write: faxon europe p.o. box 197 1000 ad amsterdam the netherlands 36 20 iassist quarterly summer 2007 san cannon* snippets of data at a glance: using rss to deliver statistics introduction for most researchers, more data are always better. the federal reserve board, like other statistical institutions, has a longstanding tradition of publishing tables of data in statistical releases and has developed applications to aid users in downloading large quantities of data. for some number watchers, especially in the economic and financial realm, an observation or two is all that is needed—but it is needed the moment it is available. how can data providers serve these clients as well as those who want every observation? our answer: provide statistics as rss feeds for simple, immediate viewing of individual observations, publish tables and serve data via applications. in developing mechanisms for delivering statistical content to the public from our website, we have noticed a recurring theme: the electronic information environment is shifting so that the presentation of our information is far less important than it once was. new methods of disseminating and repurposing information mean that we have less control over the appearance of our data and news than we have had in the past. emphasis on instant access to information on a variety of devices means that fewer of our customers are content to wait until we have information posted on a website where it will look good in a traditional browser. in response to these changes, the federal reserve board, in concert with the federal reserve bank of new york and several other central banks, have created rss-cb, a technical specification for representing common central bank data. “problems” to be solved the proliferation of gadgets that deliver information means typical delivery methods are no longer sufficient. cell phones, blackberrys, and pdas present new challenges for delivering information; users want their content delivered to them rather than having to go get it. and the new hardware on which they want the content has new software with which to access the information. in addition, it is becoming more and more common for our content to be “harvested” or accessed by automated processes. the federal reserve board’s data download program (ddp) was designed specifically to allow this type of access: interested users can write programs to download our data without ever opening a browser. not all users are interested in downloading a spreadsheet with just one or two numbers. these users are likely to have “screen scraping” programs to pull the latest exchange rate or commercial paper rate from an html table. this an inefficient method of aggregating such information and we have little control over what information is chosen or how it is used and attributed once it is extracted from our website. the fact that our electronic content is being processed by machines and not just read by humans highlights another challenge: how to address the needs of two different audiences. data management staff at the board too often have to extract individual data observations from pdf files and therefore are acutely aware of this dichotomy. traditional methods of providing data meet the needs of one group or the other. formatted html tables that humans readily understand are not the best format for machine processing of data. the board uses sdmx as a download format for data from the public website which is easily managed by automated processes but difficult for humans to consume. first step: an alternative format for the human readers while it is likely that modifying the solution for one audience would have better served the other, we decided to take a different approach and look at rss as a format that could more easily be adapted to both retain some control and fit two different sets of consumers. the board joined with the federal reserve bank of new york and other central banks and central banking organizations to investigate the plausibility of using rss to represent data. initially, these institutions were the bank of canada, the banco de méxico, and the bank for international settlements. the european central bank (ecb) and swiss national bank then joined the project. this group is now organized as the central bank online communication group (cboc). iassist quarterly summer 2007 21 as the most common use of rss is to deliver news stories, rss might not seem to be an appropriate tool to deliver statistics, as it is not immediately clear how the elements that so readily accommodate news dissemination could represent data. it is relatively straightforward to determine what the title of a news story should be, but it is not at all straightforward to discern what title a data observation should have, or whether it even makes sense to title such an item. for example, an exchange rate has many components, including base currency, target currency, time of the observation, and the value itself. what might the title be? the group created recommended formats for rss representations of common central bank data. to continue with exchange rates as an example, the group recommended that the title comprise, in order, a code for the country producing the rate, units of a target currency, a target currency code, units of a base currency, a base currency code, a date, an institutional identifier, and a rate name — us: 10.6925 mxn = 1 usd 2008-03-28 nyfed noon buying. this title allows a human reader to take in an exchange rate at a glance, no matter what central bank was the source. a banco de méxico publication of their rate appeared as mx: 10.6957 mxn = 1 usd 2008-03-28 bm for payment the group compiled an application guide for data feeds that outlines the details for order of the elements, specification details, and issues to consider. for example, implementers of rss-cb feeds should be aware that many rss readers, including live bookmarks, have a 50 character limit so publishing exchange rates with lengthy rate names or many decimal places may not provide the best information for that observation. second step: a standard for the machines while specifying the details for the title allows human users to be able to consume a single observation in a live bookmark or rss reader page, it doesn’t make automated processing much easier if machines still have to parse text strings to obtain the pertinent information. the group also created atomic metadata elements for the components of data that central banks report, to be used in common by the group’s different members. to that end, the group decided to use rss 1.0 as the “flavour” of rss on which to base the standard. rss 2.0 is more widely used owing to the simplicity of its implementation but the complexity of extending the model made it unsuitable for our purposes. each piece of information in the title has a corresponding atomic element, drawn from an existing metadata standard where possible and created as an extension where the existing standards were not appropriate. the clarity of the tags for the atomic elements helps to increase the likelihood that the information will not be misunderstood or misrepresented when repurposed by aggregators or other users. for example, the intended use of the value in an element called should be clear to users programming the automated processes. elements created as particular to central bank information are in the rss-cb namespace and are spelled out in the standard. these elements, like the target currency example above, are identified by cb indicating that they are in the cb, or central bank, namespace. elements from the dublin core namespace (dc) or the dublin core terms namespace (dcterms) are also allowed and are in some cases required by the specification. use of dublin core elements without specific reference in the specification is also allowed. the specification also provides for institution-specific extensions but to date no institution using rss-cb has developed any such extensions. the application guides and user guidelines written for the standard also give examples and requirements for the inclusion and specification of atomic elements and, where applicable, their relationship to the information displayed in the title. for example, the application guide for statistics spells out how to specify the institutional abbreviation in an interest rate feed: required. this element contains the abbreviation that signifies the identity of the institution. it contains no blanks. it is identical to the fifth field of the title. and the feed for the overnight nonfinancial commercial paper rate from the board follows that specification where frb is the institutional identifier for the federal reserve board. us: cp 2.24 2008-04-07 frb overnight aa nonfinancial commercial paper rate for exchange rates the recommendations in the application guide are slightly different: required this element contains the abbreviation that signifies the identity of the institution. it contains no blanks. it is identical to the seventh field of the title. the previously mentioned exchange rate examples show that the new york fed and the banco de méxico have the institutional abbreviations nyfed and bm, respectively. 22 iassist quarterly summer 2007 the result: rss-cb to continue with the exchange rate example, here are the rss-cb elements for the ny fed exchange rate example: <![cdata[us: 10.6925 mxn = 1 usd 2008-03-28 nyfed noon buying]]> http://www.newyorkfed.org/markets/ fxrates/noon.cfm/mxn 2008-03-28t12:00:00-04:00 en us nyfed 10.6925 usd mxn noon foreign exchange rates 200803-28t12:00:00-04:00 for the automated process, the information to be collected is clearly identified and easily repurposed. such transparency helps central banks to retain some control over their information by reducing the probability of error. for the human consumer, the title appears in their rss reader of choice and they can discern the value of the exchange rate between the u.s. dollar and the mexican peso at a quick glance without needing to download any files. in a google reader window, a list of such exchange rates from the new york fed appears as such: it is straightforward to see the information you want quickly. this particular feed is for all the exchange rates published by the new york fed but single rate feeds have also been created. users can choose to subscribe to the same rate but published by different institutions and differences will be easy to discern in such a display. subscribing to special reader services is not necessary to consume this information, although they are very popular. rss-cb data feeds can be displayed in a live bookmark in firefox or ie 7 such as these commercial paper rates from the board’s feed. notice the effect of the character limit on the rate name. our users can add a feed to their my yahoo! page which has no character limit so the rate names are displayed in full and the information appears with other information in which they are interested. iassist quarterly summer 2007 23 status and future plans the rss-cb standard is now on version 1.1 and is fairly stable. there are a few recent changes and developments that are being incorporated into all aspects of the specification and supporting schemas and some existing rss-cb feeds will need to be updated to validate with the current specification. there are several central banks besides the original developers of rss-cb who currently publish data feeds using the specification. some institutions adhere more closely to the standard than others. this is an unordered list of institutions that produce rss-cb feeds, with urls of the locations of those feeds. banco de méxico (http://www.banxico.org. mx/sitioingles/rss/suscripcionrss.html) bank of canada (http://www. bankofcanada.ca/en/rates/rss_fx.html) bank of finland (http://www.bof.fi/ en/suomen_pankki/ajankohtaista/muut_ uutiset/2007/uutinen_27092007.htm) bank for international settlements (http://www.bis.org/rss/index.htm) bank negara malaysia (http://www. bnm.gov.my/index.php?ch=107) european central bank (http://www.ecb.int/ stats/exchange/eurofxref/html/index.en.html) federal reserve bank of new york (http://www.newyorkfed.org/rss/) federal reserve board (http:// www.federalreserve.gov/feeds/) swiss national bank (http://www. snb.ch/en/ifor/media/id/media_rss) reserve bank of australia (http://www.rba. gov.au/rss/rss_cb_exchange_rates.xml) the central bank online communications (cboc) group meets annually as well as collaborating online at www. cbwiki.net; the details of the specification can also be found on the website. the first meeting was in october 2007 at the new york fed and included many of the institutions listed above as well as others who participated in our discussions on how best to deliver rss and to whom we are grateful. these additional institutions include the centro de estudios monetarios latinoamericanos (cemla), the reserve bank of australia, the deutsche bundesbank, the reserve bank of india, the reserve bank of new zealand, the central bank of nigeria, the bangko sentral ng pilipinas, and the monetary authority of singapore. the next meeting will be at the banco de méxico in october 2008 where we will continue the newly begun discussions for work on version 1.2. references rss 1.0. (2001). rdf site summary (rss) 1.0. retrieved january 23, 2008, from http://web.resource.org/rss/1.0/spec. rss-cb application guides. (2008). application guides. located at http://www.cbwiki.net/wiki/index.php/ application_guides. rss-cb specification. (2008). specification 1.1. located at http://www.cbwiki.net/wiki/index.php/specification_1.1. rss-cb user guide. (2008). user guide 1.1. located at http://www.cbwiki.net/wiki/index.php/user_guide_1.1. * prepared for the united nations economic commission for europe’s dissemination and communication work session, geneva, switzerland, on may 12, 2008. an early version of this paper was presented at iassist in 2007. san cannon, chief, economic information management, federal reserve board, washington dc, usa. contact: scannon@frb.gov. the opinions are of the author and not the federal reserve board. 4 iassist quarterly spring summer 2009 editor’s notes transformations slowly into the future welcome to this double issue volume 33 issues 1 and 2 (2009) of the iassist quarterly. this is a special issue centered on the developments of the data documentation initiative or its now familiar acronym: the ddi. we have slowly made some changes to the iassist quarterly (iq). our distribution has changed to being ‘web only.’ we have stopped printing but continued to make pdf versions available on the iassist website and you can access from the website http://iassistdata.org/iq the generated pdf versions of the journal. at the same time as producing new issues of the iq we have been helped by other people scanning old issues of the iq. right now michele hayslett at chapel hill is producing scans of the old issues. if you take a look at the iassist website you will notice that the look and feel of the website have also changed; it looks bright and sharp and you should hopefully find the interface to be more intuitive. another recent change is that the iq is now a fully reviewed journal and the iq also in the last years started to run special theme issues like this ddi double issue. these changes have demanded more work from more people and have only been possible because many people are doing iassist work in addition to their day jobs. finally access is in the process of being broadened as we have just signed an agreement with ebsco to make the iq accessible through their platform. in this issue the special editors mary vardigan from icpsr and joachim wackerow from gesis have collected from recent workshops and iassist conferences important papers on the development of the data documentation initiative as it reaches ‘ddi 3.’ in the following pages they present the articles: metadata-driven survey design by jeremy iverson, questasy: online survey data dissemination using ddi 3 by marika de bruijne and alerk amin, implementing ddi 3: the german microcensus case study by andias wira-alam and oliver hopt, metadata creation, transformation and discovery for social science data management: the dames project infrastructure by jesse m. blum, guy c. warner, simon b. jones, paul s. lambert, alison s. f. dawson, koon leai larry tan, kenneth j. turner, ddi 3 development at dda by jannik jensen and dan kristiansen, and controlled vocabularies for ddi 3: enhancing machine-actionability by taina jääskeläinen, meinhard moschner, and joachim wackerow. as iq editor i would like to thank all the authors and the special editors for their work. instead of introducing the papers this editorial will look backwards at some of the history of the ddi. a perspective could be riding on how bruno latour explains actornetwork-theory (1987, 2005). the ant-method is exemplified in latour's description of the development of the diesel engine. when mr. diesel presented the design of the engine he did not present a prototype. that took years to develop and relied heavily on many other technical skills and routines. early on, diesel took out a patent but nearly 10 years later the idea was close to collapse as the machines needed much attention and were continuously being modified. it took what latour terms as ‘translations’ and a complicated mixture of connections that translated the problems and solutions before the diesel engine was truly functional. before entering the description of these ‘translations’ latour instructs us: ‘in this technoscience game we are watching, the object is modified as it goes along from hand to hand. it is not only collectively transmitted from one actor to the next, it is collectively composed by actors’ [1987, p. 104]. similarly the ddi has slowly developed towards maturity. we find the roots of the development of the ddi in several data archives and individuals, but we also find them in collaborative organizations – amongst which iassist was one of the most important. before the foundation of the ddi, iassist established the ‘codebook action group’ in 1993. before that an even earlier development of a description standard at the study level had taken place. a report of the documentation activities at a selection of data archives was presented in the iassist quarterly (rasmussen, 1995). funding for further developments on the issues of standardization of social science metadata was established by icpsr in ann arbor and the ddi committee was formed and had its first meeting in may 1995. in 2000 i described and reported the progress of the ddi in a danish book; however, danish is not the most widespread language. it might be fruitful to remember that the ddi was focused on ‘independence.’ as a standard the ddi should be independent of platform, media, presentation, applications, and independent of commercial interests by being a standard without royalties. in latour terms we can say that the then current dependencies were transformed into independencies. much work and experiments were carried out with the ddi, but nearly a decade went by before articles directly addressing the ddi appeared in reviewed journals like historical methods (block & thomas, 2003), social science computer review (blank & rasmussen, 2004), and archival science (rasmussen & blank, 2007). when jacobs and humphrey in 2004 in communications of the acm presented their viewpoint on ‘preserving research data’ they also referenced the ddi. many people and organizations have contributed to the development or ‘transformation’ of the ddi, and many people have reported in journals, at workshops and at conferences notably the iassist conferences! one of the members of iassist quarterly spring summer 2009 5 the first ddi committee, mary vardigan, has also done a good job in making this special issue of the iq come into being. so a very special ‘thank you’ to mary and to her coeditor joachim wackerow. this month (july 2010) i met joachim (achim) in germany where gesis was having celebrations, one of these being the 50th anniversary of the zentralarchiv in cologne. there have been tremendous developments during its lifespan. technical developments are changing the possibilities in archiving and also in the relationships between people and roles e.g., in the collaboration between depositors, archive staff, and researchers. new people are also entering the scene and they will be the next ‘transformers.’ things take time. we can at times become impatient and this impatience can act as an extra driving force behind the slow transformations towards more mature solutions. with this latest report on the ddi development we can hope that the message will reach even more people. whether you are somewhat familiar with the ddi or a newcomer in the field of data documentation i hope you will find articles of interest in this issue. and you might be among the next people developing and disseminating the standards further. references blank, grant & rasmussen, karsten boye (2004): the data documentation initiative: the value and significance of a worldwide standard. social science computer review vol. 22-3, p. 307-318. block, william & thomas, wendy (2003): implementing the data documentation initiative at the minnesota population center. historical methods, vol. 36-2, p. 97 101. jacobs, james a. & humphrey, charles (2004): preserving research data. communications of the acm, vol. 47-9, p. 27-29. latour, bruno (1987): science in action. harvard university press. latour, bruno (2005): reassembling the social: an introduction to actor-network-theory. oxford university press. rasmussen, karsten boye & blank, grant (2007): the data documentation initiative: a preservation standard for research. archival science, vol. 7-1, p. 55-71. rasmussen, karsten boye (1995): documentation what we have and what we want: report of an enquete of data archives and their staff. iassist quarterly, vol. 19-1, p. 22-35. rasmussen, karsten boye (2000): datadokumentation. metadata for social videnskabelige undersøgelser. odense university press. articles for the iassist quarterly are very welcome. articles can be papers from iassist conferences, from other conferences, from local presentations, discussion input, etc. contact the editor via e-mail: kbr@sam.sdu.dk. karsten boye rasmussen july 2010 6 iassist quarterly spring summer 2009 guest editors’ notes welcome to a special double issue of the iassist quarterly featuring articles focused on the data documentation initiative (ddi), a metadata standard for the social sciences. we are proud to present these six articles, which explore various projects related to ddi 3 and its enhanced features. the articles draw on previous presentations and papers created in connection with the 2009 “expert workshop on implementation of ddi3 -advanced topics” held in wadern, germany; the 2009 european ddi users group (eddi) meeting held in bonn, germany; and the iassist conferences held in tampere, finland (2009) and ithaca, new york, usa (2010). jeremy iverson’s article on metadata-driven survey design highlights the reuse of metadata starting at the very beginning of the research data life cycle and also discusses the benefits of using metadata to drive the process of collecting, visualizing, and analyzing survey data. this is a powerful and efficient approach that should be taught in survey methods courses in order to save costs and to enable data producers to leverage the metadata they create across the life course of research data. also related to data collection is the article on the questasy online survey documentation tool by marika de bruijne and alerk amin. questasy permits internal users to document longitudinal data and to make this documentation available to external users on the web. a benefit of ddi 3 for this system is that it facilitates tracking of question items across waves in the study, where each wave can have question constructs and variables that refer to the same question item. this system was developed for the liss panel online survey at the university of tilburg in the netherlands. “implementing ddi 3: the german microcensus case study” by andias wira-alam and oliver hopt looks at using ddi 3 to document the german microcensus through a customized ddi 3 editor and a web view providing different perspectives for the end users based on the same ddi 3 items. interestingly, andias and oliver discuss basing some of their decisions about software design on jannik and dan’s use case describing the development of the ddi 3 metadata authoring tool – see building a modular ddi 3 editor. “metadata creation, transformation and discovery for social science data management: the dames project infrastructure” by jesse m. blum, guy c. warner, simon b. jones, paul s. lambert, alison s. f. dawson, koon leai larry tan, and kenneth j. turner shows the wide variety of data management tasks that ddi 3 can support and document, including recodes, merging, and data cleaning. using ddi 3 to document these phases of the data life cycle is an exciting development. “ddi 3 development at dda” by jannik jensen and dan kristiansen of the danish data archive provides a fascinating look into the development of an authoring tool for ddi metadata, a tool that is being designed to play a central role in the work flow at the dda archive. it focuses as well on the underlying reusable middleware and general considerations on open source software development for ddi. this article provides the reader with an up-close view of strategic decisions made at dda to incorporate the functionality of ddi 3 into the architecture of the dda archive. with its focus on machine-actionability and data typing, ddi 3 needs a strong system of controlled vocabularies to supplement the creation of metadata. the article on controlled vocabularies by taina jääskeläinen, meinhard moschner, and joachim wackerow presents the case for using controlled vocabularies and the ways in which they benefit the user. the article also showcases the work of the ddi controlled vocabularies group and its efforts to create vocabularies for ddi 3, which will be made available as separate products using a format called genericode. we hope you enjoy reading these articles, and we offer our thanks to all of the authors. we also want to express our appreciation to iassist for the opportunity to publish this work in the iq. we are grateful for the ongoing support of the iassist community and its nurturance of ddi from the very beginning. sincerely, mary vardigan and joachim wackerow iassist quarterly 25 a network system for the flow of african administrative information materials by edward seth asiedu, ' chief. division of documentation, caprad, tangier, morocco. introduction nearly three decades ago, the "wind of change" which blew over the african continent left in its wake spates of political agitation, and social movements and pressure groups seeking autonomy from their, then, metropolitan and colonial masters namely, the british, the french, the belgians, the portuguese, and the spanish. the granting of independence to these countries brought with it great challenges to their civil services. a consequent change in the variet)'. number and complexity of government functions has emerged, necessitating a thorough examination not only of the administrative machinery to facilitate the required expansion and modernization for accelerated economic and social development, but also, of training and orientation policies so that civil servants would be belter able to meet the challenges posed by the accelerated momentum of change. durable administrative reforms and innovations, however, often depend upon research in theories and practices of government through administrative surveys, observations, analysis of research questionnaires, interviews, experiments, etc., data concerning inefficiency and corruption can be obtained. data-collection or fact-finding is the first step towards bringing about administrative reform and creating innovations. if modem government in a developing setting is to accomplish its task and mission effectively and successfully, the administrators who will be required to assimie these new governmental responsibilities should have adequate access to organized information to broaden their outlook and training in order to equip them with the capabilities they need. studies have shown that administrative sciences are basic to national development. the development of adminisuative and managerial skills and the application of modem management techniques in an african society faces a lack of adequate and reliable administrative information for making policy and administrative decisions and for analysis of problems. in most cases, the proper administrative information and documentation channels are lacking, thus hindering the uansfer of the right amount and kind of information at the right place, to the interested user. 'presented at lassist/ifdo international conference may 1985, amsterdam fall/winter 1985 26 iassist quarterly sources of administrative information what is administrative information? briefly, let us defme the parameters of administrative information. "administration" may be defined as the achievement of an organizational objective through the use of men, money, materials and machines. "information" is "a statement that describes an event (or an object or a concept) that helps us distinguish it from others".' by extension, therefore, "administrative information" is a meaningful statement about how men, money, materials and machines are managed to achieve the objectives of an organization. in other words, administrative information could be described as information related to administration, designed in general terms, as a set of data necessary to understand institutions, make decisions and to participate in them, in order to produce the goods and services required by the community. in the african context, however, information on public administration is in reality attached to information on "development administration". despite the disagreements on an exact detmition of "development administration", most students, scholars, and practitioners involved in the development process in africa agree that information on the following subjects fall within the mandate of "development administration: namely. development planning and plans; administrative reform; public enterprise management; training of administrators; research in adminisuative systems; and socio-economic data. it would also include materials dealing with the administrative and management aspects of the functional areas of the productive sectors e.g. agriculture, industry, eonomic development, etc... administrative information sources are usually in the form of pamphlets and report literature, surveys, statistical data, case studies, issuances and other such like documents. a large number of these forms are government publications. government documents or publications are the living record of the efforts of people to govern themselves. the responsibilities of government extend over many aspects of a citizen's life, work, and leisure and, in the course of administrative control, the agencies of the government collect a vast amount of information. much of this collected data is of potential value to research workers in economics, social, scientific and technological subjects. government publications range, in size, from pamphlets to ponderous volumes; and in content, they vary from articles with popular appeal to technical treatises of value. taken as a whole, they constitute a great library covering almost every held of human knowledge and endeavor. some of these publications may be transcripts of original records and therefore constitute primary source material in the history of government administration and activities. others contain accounts by administrative or executive officers on the work and activities under their direction, e.g. annual reports of departments, bureaux, etc... voluminous series published by other agencies also present statistical pictures of socio-economic conditions and afford bases for measuring social and economic change. ' eilon. s. management control. new york, macmillan, 1971, pp. 98. fall/ ivinter j985 iassist quarterly 27 the flow of administrative materials data cannoi exist if not generated. there is evidence tjiat african governments are becoming more and more prolific in their output of official documents, published and unpublished. the remarkable transformation of postwar africa, accelerated in the past two decades or so, has created many problems in the field of government publications. there have been extensive constitutional, governmental and administrative changes, ministerial and departmental "re-re-organizations", and alterations in policies governing publishing of official documents. it is becoming increasingly difficult to identify what is available. whilst in the colonial days, it was easy to acquire the bulk of the official publications from african countries, now it is quite a problem to trace, not to mention acquire, a government document from each other. the reason is obvious the absence of a plausible, systematic bibliographic control. in-depth research^ carried out by the speaker in 1977 aimed at studying the generation, production, control, availability and utilization of african government documents provides ample information on the situation referred to above. another research survey*, carried out on the provision of library facilities in government departments in a typical african country identified one major draw-back to the improvement of the situation namely, the lack of financial support and/or importance accorded by african government's to this area. an ' asiedu, e.s. african governments: the case of west african english speaking countries: gamiba, ghana, liberia, nigeria & sierra leon. tangier, morocco, cafrad, 1978. * asiedu, e.s. "organizing the flow of administrative information in a developing setting: a proposal for the ghana public services". in greenhill journal of administration, vol.1, #3, ocl-dec.74, pp.39-56. examination of budget estimates^ for a four-year period (1971-72/1974-75) revealed some interesting facts. whilst a few of the ministries, departments etc... had huge amounts respectable enough for the establishment, maintenance and running of a library, many others run on votes which look very miserable indeed. in most cases the library votes are hidden under such headings or codes as: "other oftice expenditures"; "office expenses"; "newspapers";... worse still, others are found under "toilet papers, stationary and library"; "newspapers and rent of office soap" a clear indication that libraries, or for that matter the organization of administrative information and literature, are very much frowned upon by the civil service the very people who need data in their day-to-day occupations. albeit, it is encouraging to observe that in the last five or so years some improvements have been made in this direction. the financial situation alluded to above explains the paltry and disorganized condition of the libraries of african government departments. with such an appalling state of affairs, how can proper information be obtained, which information may be vital for the formulation of a major governmental policy? it is pathetic that often most ministries and departments lack complete collections of their own publications! very often, a belter organized library outside the governmental system is called upon to satisfy inquiries from ministries about their own publications! it is amazing to find that a foreign library has everything that african governments have issued published or unpublished from the earliest times to the present african scholars and research fellows have had to travel outside the continent to consult such materials, at great expense to the taxpayer. this situation cannot continue forever; the time is overdue for something more concrete to be initiated. ^ see annexure fall/winter 1985 28 iassist quarterly the "anai" concept the last twenty five years have witnessed a continued growth of educational and training institutions of public administration and management. also civil service commission and agencies have been inaugurated to carry out different phases of administrative reform in their respective countries. these are manned by capable african experts who are mainly involved in applying and disseminating their know-how, experience, and knowledge to the african administrative environment over the past few years their activities have resulted in the production of vast quantities of documents often mimeographed reports, government documents and other sorts of publications. these documents describe existing conditions, define goals, describe programmes and projects, and evaluate the consequences of any adminisiiative action and decision. in addition, african administrative science researchers and trainers study administrative processes, generate training materials and publish their findings and conclusions in journals, research reports and other media issued in africa and elsewhere. it is quite fair to say that this growing literature of adminisuative sciences is not yet under adequate bibliographic control and it is exceedingly difficult to know how to obtain access to the information and experiences which have been investigated in prinl with this growth of african administrative information comes an increasing demand to know about african experiences in the various aspects of administrative and managerial development it is an awareness of this situation, as well as the realization that no one african country can ever hope to become self-sufficient in the face of an expanding universe of administrative information, that the african training and research centre in administration for development (cafrad) has, in keeping with one of its constitutional mandates, and in concert with african governments, initiated the african network of adminstrative information (anai) a description of which follows: objectives of anai anai is aimed at bringing together materials from different african perspectives on administrative sciences as they relate to the african continent and bringing together literature which may have been gathered by different african libraries and/or documentation centres, but which provide in their combination a new and much greater resource, and synthesizing and disseminating the collected data at regular intervals and upon request thus, the objectives of anai are: to improve co-ordination between the existing african administrative information services i.e. libraries/ documentation centres of institutions and agencies of public administration and management; to foster the improvement and building of national administrative information services or systems; to supplement local collections by drawing more effectively on african external resources; to avoid unnecessary duplication and waste of resources; to provide a referral service for enquiries from organizations both inside and outside the african continent; to serve more users and improve national. fall/winter 1985 iassisi quarterly 29 regional and iniemalional access to african adminisualive and managerial experience and knowledge; 10 make administrative information uniformly available and to change the image of the african library from that of a store of books to that of an active and dynamic information centre. anal, as proposed, would imply a degree of diffusion and popularization of administrative information services, a steady increase in the ability to serve at all points of service, and co-operative sharing without constraint of time, distance or form of data. also, it is intended to minimize the financial pressures which are facing african administrative information activities, and to consider ways of sharing rather than duplicating materials and other resources. it encourages building appropriate local administrative collections to meet immediate needs. it is also devised in such a way as to make readily available distant collections in other african countries. anai is aimed also at motivating the utilization of modem miniaturization and computerization techniques in the storage and retrieval of administrative information. research workers in the domain of the administrative sciences. also, it is expected that the system will extend its services to any user interested in african administrative and managerial development anywhere in the world, without any restriction on access. with regard to the subject coverage, anai is designed to deal with specialized materials on public administration and management with respect to the african continent it will encompass data pertaining to topics such as: local government, rural and urban development, administrauvc reform, project management, financial management, human resources development, etc... as a mission-oriented system, anai would have to include relevant multiand inter-disciplinary information sources in as much as it has a role in the amelioration and improvement of african administrative and managerial development. the administrative information materials to befed into anai may be recorded in a wide range of media formats: printed forms; audio-visual forms; machine-readable forms (punched cards, magnetic tapes); microphotographic forms (fiches/films). these materials should be mainly generated in africa, or elsewhere if they are confined to the subject coverage in the african context scope and coverage anai, as seen from the preceeding paragraphs, is a specialized system limited to special categories of users and confined to specific locations. it is anticipated that anai will serve a wide variety of users who fall into many categories such as: governmental and other administrative managers; policy makers; educators and trainers; tramees and students; organizational structure anai is envisioned to function as an african region-wide decentralized information system with a network of (for the time being) three interrelated lines of operation, namely: fall/winter 1985 30 iassist quarterly 1. regional central clearing-house (cafrad) 2. national administrative information focal points 3. local libraries and documenialon centres of ministries, departments, bureaux etc... it is envisioned that a fourth operational line or organizational level could be inserted between regional and national levels to represent a "sub-regional" level, but this development will depend on the magnitude of growth of anai in the future. brief descriptions of the functions of these levels follow; regional central clearing-house: will function as a coordinating agent of the whole network. it will develop and set out technical methods, procedures and standards for the whole network. it will also develop, revise, maintain and unify its indexing language in thesaurus of administrative informatin descriptors; forms for input data, and dissemination channels for output products. it will maintain a data base of information needed, perform miniaturization, microfilming and microfiching of textual input data; and issue guides, directories, abstracts and indexes. it will develop training programmes for the information and documentation personnel of member states and also provide specialist consultancy services to governments as and when required. national administrative information focal points: these national focal points represent the national participants of anai. they would be libraries and/or documentation centres designated by their governments. the role of these national focal points will be to furnish the regional clearing-house with local data on administrative and managerial development these will include input data on institutions and agencies, training programmes, consultants and trainers, training materials, and current research activities which will constitute the data base of the regional central clearing-house. each national focal point would assist in mobilizing and co-ordinating the national resources and services of administrative information systems. this policy is in conformity with the recent recommendation promoting the creation of national information systems (natis) made by the "inter-governmental conference on national libraries, documentation and archives infrastructure" which was organized by unesco, ifla and fid and held in paris in september 1974, and also with the uneca conference of ministers' resolution n" 359 (xiv) of the 14th session held in rabat, morocco. national focal points will be the main suppliers to the regional clearing-house of the relevant textual documents, acquisition lists and abstracts. in return, the ncps will get all anai publications, directories, guides, bibliographies, indexes, microfilming services and other dissemination tools pertinent to administrative information. local libraries & documentation centres: these will include libraries and other such information units within the civil service organizatons like ministries, bureaux, departments and other stale agencies and institutions of public administration and management these local centres will feed the national focal points with their materials and also collabourate with them to fall/ winter j 985 iassist quarterly 31 idcnlify and supply ihcir inpul coniribulions to the regional central clearing-house. diagramalically. the organizational structure of anai looks like this: regional level regional central clearing-house (cafrad) international information system e.g. devisis.padis, etc. national level national administrative information focal point ( country " a" ) local level \ s > v \ \ local library local library local library national administrative information focal point (country "b" ) national administrative information focal point (country "c" ) local library local library local library this kind of organizational relationship could assist in bringing together the african administrative literature under national and regional bibliographic control and this in turn will surely assist and facilitate the identification of resources and access to them. it will also ensure the success of co-ordinating, managing and operating anai. fall/winter j 98 5 32 iassist quarterly anai output anai would ensure output of required data in the required form at the required time to a specific user. this would include the production of specialized guides and directories represented in the different data files stored in tlie data base, as well as periodic regular abstracts, indexes and specialized bibliographies which represent the current awareness services to users. it will also furnish answers to multiple queries. there will be back-up service of microfiche copies of the full texts of documents identified in the indexes, abstracts and other bibliographies disseminated through the output operation. diagramatically this output will look like this: direciones. & files insuiuijons/agencies training programmes training materials consullanis & tramers current research projects ouipui producis current awareness monthl> indexes service quarterly absiracls specialized bibliographies reprography services photostats/xcros microriching pnnung conclusion an improved network of african administrative and managerial development needs a high-speed, reliable information system. the conventional means of professional communication, i.e. libraries, provide coverage that is neither complete nor timely, especially as regards research proposed or on-going, as well as decision-making process. the evolution of a voluntary association of existing libraries and documentation centres of administrative information services in africa, each preserving its own management autonomy, offers a realistic basis for creating anai. therefore, the anai project can be viewed as a regional movement towards increased voluntary cooperation among national partner focal points with varying degrees and modalities of interconnection. this regional network could be further developed into a world administrative information system that provides cunent reports on all significant findings, activities and planning affecting any part of the network. for that ultimate goal, the development of anai could emerge in conformity and in line with experience gained from the existing international information systems and within the conceptual framework of unesco's worid science information system (unisist) programme. early this year (january 1985) i.d.r.c. (editor's note: international development research centre) of canada responded to c.afrad's call for financial support to enable it to implement a pilot project code-named "anai experimental" for a three-year period to test the operational system of anai. we are very grateful to the centre for this demonstration of confidence in us. however, there is room for further assistance. may i, therefore, crave your indulgence to make an appeal, on behalf of cafrad: to this august body; to development institutions here present; to other donor agencies world-wide, to collabourate with cafrad, to assist cafrad, to share with cafrad their expertise and financial resources to ensure the germination, growth and fruition of the anai seed. fall/ ivinter j985 [assist quarterly 33 library note statistics depanmem 1971/72 1972/73 1973/74 1974/75 remarks transport & communication auditor-general trade & tourism agriculture town & country planning national fire service social welfare & com. dev. police service internal affairs (central admin.) prison service foreign aftairs posts & tele-communications local government (cent. admin.) nrc (general administration) nrc (press secretary) nrc (chieftaincy) national council for higher edu. nrc (annex) nrc (msd) nrc (pib) regional organizations parliament house establishment secretariat (training) public service commission information (central admin.) information services department informaion services (regions) attorney-general's department judicial services registrar-general's department central revenue finance and economic planning 1000.00 100.00 500.00 600.00 300.00 130.00 2000.00 800.00 100.00 1250.00 3500.00 55.900.00 100.00 70.000.00 100.00 6000.00 6000.00 including regions 146.00 3400.00 4400.00 1700.00 5000.00 3000.00 800.00 700.00 760.00 2000.00 500.00 500.00 900.00 1500.00 2200.00 including education instruction materials 400.00 400.00 400.00 subscription to ghana library board & prison library service 12000.00 8800.00 10000.00 660.00 4000.00 3000.00 3000.00 10000.00 includes capital vote of 2730.00 5000.00 5000.00 1000.00 3000.00 4500.00 5000.00 1300.00 . 4000.00 6000.00 6000.00 4000.00 2000.00 380.00 1000.00 1400.00 1920.00 5220.00 6600.00 no vote for ashanti & upper regions foi 1974/75 500.00 1000.00 3200.00 1000.00 200.00 200.00 for newspapers onl>^ 1900.00 1900.00 2000.00 2800.00 7640.00 6000.00 1400.00 1000.00 3000.00 32,000.00 8000.00 8000.00 including regions 53.000.00 22,140.00 25,820.00 400.00 1.000.00 1,000.00 1.410.00 10.000.00 60.000.00 9 9 56.500.00 including 3 other divisions fall/ winter 1985 vol271 10 iassist quarterly spring 2003 iassist quarterly spring 2003 11 by hannele keckman-koivuniemi & mari kleemola* data processing in fsd: challenges in a new archive the finnish social science data archive fsd began operation in 1999 as a separate unit of the university of tampere. it is funded by the ministry of education. at present the archive has thirteen full-time employees. fsd has archived about four hundred studies, of which one third are international in scope. fsd obtains about one hundred new studies annually from a variety of sources: icpsr, public and private research organisations and individual researchers. studies from other archives (for example icpsr) are documented in finnish but the data are not processed. the data that fsd archives must be from an area or discipline in the field of social sciences. it should also fulfil certain technical and legal requirements: aspects of copyright and ownership should be clear, there should be no legislative impediments to archiving (e.g. data protection, privacy protection), the original purpose of data collection should not prevent archiving and, last but not least, the information content and technical properties should make the data suitable for archiving. data are processed intensively the archiving process begins when the depositor delivers machine-readable data to fsd. the data are most often received in spss or excel format and occasionally in sas or ascii format. depositors are asked to give detailed information about the collection procedure, resulting publications and the research project in general. fsd prefers to obtain supplementary documentation both in electronic and paper formats. the completed and signed material description and material deposit agreement forms are necessary as well. all national studies acquired by fsd are processed intensively – a very time-consuming undertaking. there are two reasons for doing this. first, the number of archived studies so far is manageable. second, researchers do not usually offer their data. rather, fsd staff identify studies from, for example, scientific journals and then request the data from the researchers. this way all archived studies are by presumption “important” and deserve intensive processing. a goal of intensive processing is to make the content of the archived data correspond as closely as possible to that of the original questionnaire. documents received from the depositor, such as the questionnaire, original observation matrix, variable lists, printed books and articles play a key role in this work. all alterations are carefully documented in the syntax. fsd uses spss in data processing and preserves data in portable format. we have chosen spss because it provides at least at present the best option for preserving data and identification information (that is labels) in the same file. since spss is widely used, converting material to future formats should be manageable. most of our customers are familiar with spss and about 90 % of researchers wish to have data in this format. checking the data the original data file is preserved without modification as long as necessary for the archiving process. so far we have not deleted any original datasets. a copy of the original file is used to produce a version suitable for secondary use. we focus on the content, not the “look and feel” of the studies. during the checking process mistakes are corrected and verifications, amendments and additions are made. we confirm that the number of variables and cases match the documentation supplied. we check frequencies to verify variables, valid and invalid values (e.g. missing data, not applicable responses etc.). variables are renamed corresponding to question numbers. background variables without their own question numbers are renamed bv1, bv2 etc. fsd has standardised labels of some background variables (e.g. original question: how old are you? --> variable label: respondentʼs age). variable and value labels are constructed based on the questionnaire. variable labels often become quite long as they may contain the whole question text. we try to keep variable names and labels consistent within studies of the same series. questionnaires often include questions that are directed only at respondents who meet certain requirements. the archive checks these filter conditions. if the data include answers from people who do not belong to the specified 12 iassist quarterly spring 2003 iassist quarterly spring 2003 13 target group, the responses are classified as missing data. in order to make sure that the content of the archived data corresponds as closely as possible to that of the questionnaire some variables may have to be dropped or added. a variable is dropped if it is undefined or data security aspects so require. constructed variables, such as combined variables and sum variables, are usually dropped. however, those constructed variables that are integral to the usability of the data, especially weight variables, are kept providing that the documentation provided is explicit enough. new variables are added only if usability so requires. data protection confidentiality aspects require that personal data are deleted. it is recommended that depositors remove these types of data (names, addresses, birthdays etc.) before delivering the material to the archive. under certain circumstances, fsd stores materials which contain personal data. in these cases, the archive anonymizes the data according to its own guidelines and depositors ̓instructions. variables indicating place of residence and business are also problematic. on one hand there is always a risk that a single respondent might be identified, on the other hand the deletion of these types of variables prevents secondary users from conducting regional comparisons, especially if no other regional variables are used. usually these types of variables are dropped. if necessary, they can be restored. variables of larger regional units (provinces, districts) are kept. version control we aim to process the datasets only once. still, some of them have to be reprocessed. sometimes additional information about variables is given by depositors, errors are detected or processing procedures updated. for example, datasets processed in the early days of fsd already need repairing. the first final version of a dataset is called version 1.0. if subsequent changes are more or less cosmetic (typing errors etc.) the new version will be called 1.1. in the case of significant changes (e.g. a variable added) the new version will be named 2.0. we add these alterations to the end of the same spss syntax file used in processing the dataset in the first place. this is not a long-term solution but has worked so far. we track changes and versions as well as all files in our operational database. documentation fsd uses ddi standard for creating data documentation. at present fsd produces study descriptions in finnish and english and pdf codebooks in finnish for all our national studies. all documentation is available on the internet. datasets are translated into english on request. the data and documentation are fully compliant with the search engine nesstar. fsd also takes part in the madiera project (multilingual access to data infrastructures of the european research area) which started in december 2002. challenges for the future we need to and will review and update our data processing instructions and procedures in the near future. it is obvious that fsd will not have the resources to continue processing all datasets intensively, but will have to introduce several different levels for processing and documenting datasets. first we need to define the minimum level of processing required to preserve dataset quality for the long-term. one problem common to all data archives is how to get principal investigators to provide enough details about the data and documentation. the intensity of our processing requires a large amount of information about each study. researchers are often unwilling to take the time necessary to dig up and assemble the details and the lack of information slows the archiving process significantly. using different processing levels would mean a faster archiving process and quicker publication of data. another challenge is version control. we cannot continue controlling versions by amending the syntax file originally used to moderate the dataset because the original data might and probably will be unreadable in the future. the spss syntax should merely be a tool, not documentation needing to be preserved. also, as our purpose is to preserve the content, not the “look and feel” of the data, we are not planning to preserve the original datasets forever. the data for a growing number of studies are collected by computer-aided interview systems like cati and capi. at the moment we process these datasets manually. this needs to be automated as the number of archived cati and capi studies grows. computer aided interviews also mean that there are no traditional-style survey questionnaires but our current data processing procedures assume that a printed questionnaire exists. nesstar and madiera will enable online data downloading in the future. this may create an entirely new set of problems. another challenge will be data protection. the measures we take today may not be sufficient in the future. the scientific community is increasingly aware of these challenges. last but not least is the question of longterm preservation of data and other electronic documents. in this article we have only touched the surface, the list of challenges surely does not end here. fsd has received a lot 12 iassist quarterly spring 2003 iassist quarterly spring 2003 13 of information of good or best practices in data processing from other data archives in europe and north-america. co-operation on the international level has been, and will continue to be, crucial for all data archives, especially new archives like fsd. * paper presented at the iassist conference, may 2003, in ottawa, canada. address: fsd, university of tampere, 33014 tampere, finland, fax: +358 3 215 8520, internet: www.fsd.uta.fi. email: hannele.keckman-koivuniemi@uta.fi tel: +358 3 215 8530. email: mari.kleemola@uta.fi tel:. +358 3 215 8528. http://www.fsd.uta.fi mailto:hannele.keckman-koivuniemi@uta.fi mailto:mari.kleemola@uta.fi evaluation and appraisal of university administrative computing datasets by mark conrad ' data archivist pennsylvania stale university introduction in september of 1990 the pennsylvania state university archives in cooperation with management services, a division of the university responsible for administrative computing services, began a two year grant project funded by the national historical publications and records commission (nhprc grant #90-095). the objectives of this project are to appraise, preserve, and make available electronic records created, stored and used in the management services division of penn state university; to develop ongoing procedures for the appraisal of administrative computing data in the future; to develop protocols for the use of the data by institutional and outside researchers while enforcing restrictions on access for privacy and confidentiality purposes; and to provide recommendations based on the project for the preservation of archival data from on-line administtative database systems. the project was divided into four phases. the first phase lasting two months was used for the orientation of the data archivist to the operations of the university archives, the university records management program, and management services. phase two, the phase we are currently operating in, is used for the appraisal of datasets. this phase is scheduled to last 18 months. phase three and four are scheduled to last the final four months of the two years. in phase three recommendations will be made for the identification and preservation of future archival datasets and protocols will be developed for the research use of the datasets. in phase four reports on the project will be prepared and circulated. in actuality, much of the work scheduled for phases three and four has already begun and is being carried out concurrently with the appraisal process. this paper will focus on the second phase of the project i will discuss the appraisal process as it is currently being carried out, difficulties encountered and lessons learned to date. the appraisal process for the purposes of this project the appraisal process will be confined to a finite number of datasets from the university's administrative computing mainframe. we will not be examining electronic records from other mainframe, mini or microcomputer systems. rather, the appraisal process will be limited to some 3,000 datasets recorded on "history" tapes. history files, in management services' parlance, are usually copies of master files at a particular point in time (often the end of a semester or academic year). the datasets would be very difficult to recreate if they were destroyed because they are copied from files that are constantly updated. these files are kept for possible reuse ^. the datasets date from the late 1960's to the present. these datasets contain information for fourteen areas of administrative responsibility at the university. the areas are: accounting, payroll, bursar, student aid, agriculture, planning and analysis, budget and resource analysis, management and systems engineering, admissions, regisd-ar, testing services, development, graduate school, and physical education. each area has a data steward who "develops the coding structure of the data, insures the data's accuracy, determines the frequency of updating, and establishes data use and protection requirements."' the data steward is usually a senior administrator, such as the registrar, or that persons' designate. the appraisal process begins by selecting a data steward's area of responsibility. the data archivist must have the written permission of the data steward to examine the datasets under his/her control. in some cases datasets are jointly "owned" by more than one data steward. in such instances the disposition of the datasets must be discussed with all interested parties. in starting the appraisal process i chose data stewards that had only a few history tapes to test the procedures developed for the appraisal process and the database system to track appraisal information. having chosen an area of responsibility, the next step is to identify those datasets that belong to the data steward from the (management services) tape library listing of the history files. once identified the datasets are grouped by common dataset name. this is done because datasets with the same name usually contain the same types of data and can be initially evaluated as a group. the next step is to locate as much information about the datasets as possible. i am to find the procedure and summer 1991 program that created or used the information, any documentation, a file description, a record description, record counts, or any samples of input or output. records management program retention schedules are checked to see if similar records in another format have already been scheduled. if similar records have been scheduled for destruction and the datasets do not have some additional value by virtue of their being in electronic format and thus more manipulable, the datasets may be recommended for destruction. (junk is junk no matter what the format.) while the search is on for information about the datasets, some of the tapes are read and printouts are made of a sample of records from each file. this serves several purposes. firstly, it verifies whether or not the tapes are still readable. some of these tapes have been in storage for a very long time under less than ideal environmental conditions. secondly, the dump can be used for comparing what's actually on the tape with what the documentation says should be on the tape. if the two don't match, other documentation must be located or the file may be recommended for disposal. a file is of no value if a determination cannot be made as to where one field ends and the next begins or as to what a particular value in a field indicates. having gathered as much information about the datasct as possible, the next step is to interview the data steward and/or a contact designated by the steward about the datasets under his/her con&ol. a standard list of questions has been developed to help the data archivist gather all the information necessary to make an informed appraisal decision. those questions are: do you have documentation for these files? do you have samples of input and output for these files? where did the data come from? what was it used for? is it still being used? is it updated? how often? are the records maintained in another format? has the other format been scheduled for retention or disposal? are there any requirements for the retention of this data that you are aware of? are there any restrictions on the use of this data that you are aware of? at this point the data archivist should have enough information to begin making the appraisal decision. the decision-making process is not that different from the process for more traditional records. it is certainly not very different from the process used by archivists working in the electronic records programs in government archives. does the dataset have legal, evidential, or informational value? is this dataset unique? is it the most desirable format for keeping the information? is the data hardware/software independent? if the answer to all of these questions is yes, all datasets that share the common dataset name and structure will be recommended for accessioning by the archives. retention schedules arc developed for all datasets, regardless of their status, in concert with the data steward. management services, the records management program staff and the university archives/records management advisory committee. once the decision has been made that a group of datasets are archival, each dataset is read and a data dump is obtained. as with the sample of datasets read previously this is to verify that each dataset is readable and the data is valid. file structures and record descriptions change over time so the data archivist must insure that data from each dataset is adequately documented so that a researcher or the archives staff can use it. assuming the datasets are readable and understandable, copies are made of each dataset and the relevant documentation. the datasets are accessioned by the university archives and the data archivist turns his attention to the next set of files. problems encountered you may have already gathered that one of the biggest problems has been locating adequate documentation for many of the datasets. the record description for many files seems to change on a regular basis. often the documentation is not updated to reflect these changes or conversely when the documentation is updated previous versions are discarded despite the fact that files still exist that were created using the previous documentation. it is not unusual to find a number of files with different file (assist quarterly structures, the same name, and one set of documentation that may not match any of the files. in talking with other archivists working with electronic records, i have been assured that penn state is not alone in this predicament. management services is currently exploring the possibility of recording documentation for a dataset directly onto the first label of the tape a dataset is recorded on. as long as the documentation is copied along with the dataset whenever the dataset is transferred to new media, the proper documentation should be available for the life of the dataset another related problem occurs when trying to appraise an older dataset often there are no employees still working in the office that used the file who remember what the file was used for. sometimes the office itself no longer exists! the turnover on a university campus and the restructuring of administrative units can make it difficult to find someone who can tell you how a file was originally used or what dataset replaced the one you are evaluating. the only solution to this problem is to carry out the appraisal early in the life cycle of a dataset. lessons learned one of the most important lessons we have learned is that the shorter the time lapse is between the creation of a dataset and its appraisal the easier it is to identify archival datasets and insure their preservation. the archivist can interview all the players involved in the creation of the records to better understand why the records were created and under what circumstances. documentation can be evaluated to insure it adequately explains the data so that it will be useful to future researchers. datasets that are identified as archival can be marked for special handling to insure the data will still be readable 10, 25, or 100 years from now. another lesson we have learned is that the cooperation of the administrative computing center is essential to the success of an electronic records program. the archivist needs to understand how data is manipulated at the center to meet the informational needs of the institution. the administrative computing center personnel must have an appreciation of the potential value of the data beyond the purposes for which it was originally created. any archives considering implementation of an electronic records program would be well advised to begin building relationships with their institution's administrative computing center(s) now. the most important lesson we have learned is that more records are being stored in electronic format all the time. if we do not identify and preserve the archival datasets a large portion of our institutional memory will be lost. at penn state we have begun the process of insuring these valuable records will be preserved, we encourage other institutions to join us, and we are happy to share information about our project. footnotes ^pennsylvania state university. management services division. standards and procedures manuals (on-line manual). ' pennsylvania state university. administrative policy, ad-23. references pennsylvania state university, management services standards and procedures manual #3, chapter 5, section 1. (on-line manual) pennsylvania state university administrative policy, ad-23, p. 1. 'paper presented at lassist 91, 15 may 1991, edmonton, alberta, canada. summer 1991 essda 6 iassist quarterly summer 2010 iassist quarterlyiassist quarterly abstract estonia was one of the pioneering empirical data researchcentres in the soviet union since the 1960s, with computing centres located at the largest research institutions. after the restoration of estonian independence, a team of social scientists developed an initiative to save unique data produced in the preceding decades. higher education support program from open society foundations supported early data conversion projects. establishment of the estonian social science data archive in 1996 at the university of tartu provided a focal point for data acquisition and support efforts, but its activities have been constrained by lack of funding and the size of the estonian social science research community. essda has broadened its user base by offering services to employees in the public sector and teaching general principles of data use and archiving to social science students development of social science and data archiving in estonia before founding of essda estonia was incorporated into the soviet union only in 1940 after a period of independence. for that reason, it had a relatively open ideological atmosphere and many of its citizens spoke foreign languages. for these reasons, new ideas were accepted more easily than in the rest of the soviet union. thus it was not surprising that estonia became (along with the russian cities of leningrad and moscow) one of the pioneering empirical data research centres from the 1960s (titma 2002). two of the largest mainframe computing centres were housed by estonian radio and the university of tartu. the research environment in soviet period produced relatively good quality data products in robust computing environments. however, shortcomings in university education in sociological disciplines caused data analysis methodologies to be weak. improvements in sociology education in universities started only in 1989 when the soviet union was very near to its collapse. that milestone sharply accelerated the development of quantitative methodologies and expanded the number of potential secondary data users in estonia. after the restoration of estonian independence, a team of social scientists from the university of tartu made up an initiative group in 1993 that had two primary goals: to create a social sciences data bank and to develop a strategy for saving and encouraging use of the research materials collected by the estonian social scientists during the previous decades. their application for support to the higher education support program (hesp) was successful; a grant gave a chance to “save” unique data from previous decades. in 1996, the estonia social science data archive was officially established as an interdisciplinary centre of the faculty of social sciences of tartu university and as a national social science data bank. at the end of that year, essda hosted an international conference on data archives and their functions in social research in eastern europe (murakas and rämmer 2001). essda became a full member of the european council of social science data archives (cessda) in 1997. developments of data archiving in estonia after the establishment of essda lack of the funding seriously restricted development of the estonian social science data archive in the late 1990s (murakas and rämmer 2002). the situation was quite critical and the main goal for essda in that period was survival. several applications for funding were rejected, so the faculty of social sciences of the university of tartu has been our only regular funding source. the archive was also supported by the national programme collections for the social science data archiving and needs of the public sector: the case of estonia by rein murakas and andu rämmer1 a grant gave a chance to “save” unique data from previous decades. iassist quarterly summer 2010 7 iassist quarterly humanities and natural sciences” in 2005–2008, especially in activities connected with soviet-time data. despite very limited resources, essda has continued its activities compiling new data and organizing existing data. our electronic data collections were supplemented with complementary non-electronic archives of soviet period studies (study reports, unpublished analyses) from estonian radio and the former laboratory of educational studies (that continued its activities as department of sociology) at university of tartu. the number of users of the essda has increased. in addition to estonian users, some are foreign social scientists and journalists. a growing number of international publications (kaplan 2006, kaplan and brady 2009, brady and kaplan 2011) are based on the archived estonian data; essda itself is noted as an important institution in various overviews about the state of estonian social science (kroos, murakas and veski 2009, murakas and rämmer 2000, rosimannus and titma 2004, titma 2002). at the same time interest among estonian researchers in studies housed by foreign archives is growing. most are interested in eurobarometer surveys. martinotti and stefanizzi (2000) stressed the importance of eurobarometer as one of the most comprehensive and continuous academic survey programs made available to the academic community thanks to the data archiving. essda has developed some important initiatives. accompanying printed questionnaires of soviet-era datasets were scanned and now make up an essential part of our online database. we tried to launch a regular international electronic journal “estonian social science online”2. due to lack of finances, only two first issues in english were launched. however, two subsequent numbers were issued with articles in estonian. these featured the presentations of the estonian annual conferences of social sciences. essda also faces many challenges. the biggest limitation that emerged in acquiring new data is the lack of trust. like most spheres of life in post-communist societies (howard 2003) the lack of trust also characterizes potential providers of social data. for example, they can be worried that local customers at tartu university may enjoy advantages of access to the data not available to others. it appears that the tradition of sharing data among social scientists is remarkably weak and depends on personal relationships. secondly, support from the public sector is often cooperative but comes with no financial assistance. thirdly, the permanent lack of funding generally reduces chances of finding additional means of financing. often the work of the data archive was regarded as some small local initiative of enthusiasts and it was not financed from state level structures. in the framework of structural reform at tartu university essda has been re-organized to a consortium in 2008. essda’s statute was also renewed. this organizational reform would probably give us better possibilities for direct cooperation with different institutions involved with using and producing of social science data. current data sources and data usage in estonia estonia is one of the smallest eu countries with a total population of 1.34 million at january 1st, 20103 . although membership in the eu boosted international cooperation among social scientists (murakas et al 2007), the total number of social scientists in estonia is one of smallest among member countries. on the basis of very discussable statistical data, in estonia in 2009 947 people are working as social science researchers in non-profit institutions. in full-time equivalents the same number is 4744 . those data are based on broad definition of social science. if we exclude law and economy professionals, not directly connected with social science survey data stored in essda, we can say, that only a few hundred social scientists are potential users of our archive data. they are organised into 30-40 small research groups that use and need data in different ways. if essda data collections consist of only local data, they are not a good resource base for publications in international peer-reviewed journals, the main criterion for evaluating academic contributions among the social sciences in estonia. the ministry of education and research together with the estonian academy of sciences launched a process of compiling the estonian research infrastructures roadmap. to develop possibilities of international data usage in social sciences, essda made a proposal to include the council of european social science data archives (cessda) as an important international infrastructure in the roadmap. essda’s proposal was not accepted with suggestion to apply again in the next round. but it is important that from social sciences, the proposal “estonia in european social survey” was accepted for the roadmap. archived estonian data can be essential for the narrow branch of historical analysis of the former soviet union. for example we can mention estonian science foundation grant “changing cultural dispositions of estonians through the four decades: from the 1970s to the present time” headed by professor marju lauristin, started in 2009. the historical part of the project is based on the re-analysis of the soviet and transition-time estonian data from essda. however, research groups use mostly public data from comparative surveys (like european social survey) that are available online, but also official statistics collected by estonian statistical office that is stored in their online database and data collected by joint projects with international partners. however, we can conclude that there cannot be very high demand for data from essda’s traditional clients for the foreseeable future. to enlarge the circle of potential customers, reorganization of the activities and change of paradigm were needed. raising the importance of alternative customers from the public sector in 2009, the estonian public sector employed about 159, 000 people, about 47% of them with higher education, master’s or phd degree5. about 37, 000 people are working in the state-level institutions (branch public administration and defence, compulsory social security in estonian classification of economic activities6 ). it is difficult to say how many of public sector employees can be considered analysts who need social science data to support their work but in any case their number is much higher than that of academic social scientists. fast economic development triggered new divisions in society, and the need to encourage use of social science research information by public sector staff is urgent. this information is of growing significance to estonian public policy and its development programs in recent years. based on surveys (kasemets 2009), the real situation suggests that the use of social science data is rather modest in the public sphere. possible rise in interest toward data resources should be connected with the need to raise the effectiveness of the public sector (especially due to budget cuts) and also with the realisation of different projects that are directed to the development of public sector. the number of projects financed from european structural funds that require increasing of quantitative methodologies based on social analyses has increased sharply after joining the european union. at present, the focus of sharing the fruits of applied research is production and distribution of research. these fulfil the legislative mandate that the results of applied research projects that are publicly funded by the estonian government be widely distributed by government agencies. however, such reports are generally published only on web sites of these agencies. there is no coordination among agencies to provide a comprehensive view of public research investments, for example to establish centralized database of research reports. 8 iassist quarterly summer 2010 iassist quarterly to help address this, essda submitted a proposal to establish an inter-agency metainformation database with information on these dispersed research projects. however, the proposal was not funded because of emerging institutional barriers. every institution needs for its purposes specific information, for example, an overview about research projects connected with higher education or multicultural issues. as a result, the current solution is that essda started establishing metainformation database by topics depending on the available funding sources. this means that single metainformation database is accessible via different user interfaces via different user interfaces like separate interface for educational or for cultural surveys. in 2009, essda already developed all-estonian survey information database about higher education and also small database about multicultural problems like for educational surveys or for cultural surveys. plans for the future are to build a large nesstar-based meta-information database that will cover all research areas. our nesstar server is launched and the number of resources in our virtual library is increasing. early data descriptions are in estonian as we try to extend the nesstar community in the estonian public sector. however, even if research reports are sometimes insufficient, there is currently little further interest from the public sector representatives to analyse data themselves. as a general rule, they do not need complex secondary analysis but quick answers to very concrete questions. in addition, they sometimes lack relevant skills and appropriate software expertise. these are difficult to address when public sector budgets are being cut. essda may have a future role in providing a service to analyse and interpret survey results for many different research datasets to support public policy work. as part of the faculty of social sciences, essda could use graduate students as part-time analysts. in the short term, there is no funding for this. but in the longer term, the data service tradition can be institutionalised if it can become self-financed. raising the profile with teaching and training a large percentage of social science students start their careers in the public sector after graduating from the university. as a result, working with students can supply them with the data use and archiving skills they may need when they enter the workforce. at the university of tartu, almost all social science students (excluding law students) get at least basic information about social science data archives and learn fundamentals of secondary analysis in the introductory social science courses (“introduction to social sciences”, basics of social analysis”). we also teach advanced data archiving study courses (“social science data archives”, databases in social sciences”) or sociology students on a regular basis and offer a special course “re-contextualization of sociological inquiries from the soviet era” about data resources and interpretation difficulties of soviet-period data to phd students. when our alumni attain key positions in the public sector, they are familiar with the potential of secondary analysis based on data archiving. they will also be more competent in using resources offered by the data archive. dangers associated with misuse of data and disregard of rules for data protection can be serious obstacles to accepting the products of data analysis (jagodzinski and moschner 2007). we initiated activities directed toward raising the knowledge and competency of the public sector in the field of social data analysis. to that end we developed specific components in special continuing vocational training (cvt) courses about the methodology and research design. we already have positive experiences about providing such training with education and criminal justice analysts, secondary school leaders etc. conveying knowledge of the advantages and strengths of data libraries is not an easy job, but such problems are no unusual and characterize the current state of social sciences in general. references brady, h. e., and c. kaplan. 2011, forthcoming. gathering voices: political mobilization and the collapse of the soviet union. new york: cambridge university press howard, m. m. 2003. the weakness of civil society in post-communist europe. cambridge: cambridge university press. jagodzinski, w., and m. moschner. 2007. archiving poll data. in: w. donsbach, m., w. traugott (eds) sage handbook of public opinion research. london: sage, 468-476. kaplan, c., s. 2006. setting the political agenda: cultural discourse in the estonian transition. in esherick, j. w., h. kayali and e. van young. (eds) empire to nation: historical perspectives on the making of the modern world. boulder colorado: rowman & littlefield publishers, inc.: 340-372. kaplan, c., s. and h. e. brady, 2009. conceptualizing and measuring ethnic identity. in abdelal, r., y. herrera, a. i. johnson and r. mcdermott. (eds) measuring identity: a guide for social science research. cambridge: cambridge university press, 33-71. kasemets, a. 2009. the rift between legislative drafting standards and the facts in presenting information on evaluating impacts and involving interest groups; in estonian. riigikogu toimetised, 19, 104 115. kroos, k., r. murakas, and l.veski. 2009. the need for integrating higher educational policy and research studies; in estonian. riigikogu toimetised, 20, 159 167. martinotti, g. and s. stefanizzi, 2000. data banks and depositories. in borgatta, e. f., and r. j. v. montgomery (eds) encyclopedia of sociology. ny: macmillan reference, 573-581. murakas, r. and a. rämmer. 2000. essda spreads information about changing society; in estonian. riigikogu toimetised, 1, 282-283. murakas, r. and a. rammer. 2001. estonian social science data archives: past and future perspectives. iassist quarterly (iq), 25 (2), 9 10. murakas, r. and a. rämmer. 2002. empirical research and situation of data archiving in estonia. in hausstein, b., p. de guchteneire. (eds). social science data archives in eastern europe: results, potentials, and prospects of archival development. berlin: ferger verlag, 177 – 180. murakas, r., i. soidla, k. kasearu, i. toots, a. rämmer, a. lepik, s. reinumägi, e. telpt, and h. suvi. 2007. researcher mobility in estonia and factors that influence mobility. tartu: archimedes foundation. rosimannus, r.; and m. titma. 2004. estonia. in geer, j. g. (ed) public opinion and polling around the world, vol. 2. santa barbara: abcclio, 570 – 576. titma, m. 2002. sociology estonia. in: kaase, m,; v., sparschuh, and a. wenniger. (eds). three social science disciplines in central and eastern europe. handbook on economics, political science and sociology (1989-2001). budapest: social science information centre, 425 – 436. notes 1. rein murakas and andu rämmer, university of tartu, estonia and estonian social science data archive (essda). contact: andu rämmer, andu@ut.ee. 2. http://www.psych.ut.ee/esta/online 3. database of estonian statistical office (http://www.stat.ee) 4. database of estonian statistical office (http://www.stat.ee) 5. estonian labour force survey 2009 data, estonian statistical office (http://www.stat.ee) 6. database of estonian statistical office (http://www.stat.ee) vol271 4 iassist quarterly spring 2003 iassist quarterly spring 2003 5 by editorʼs notes welcome to the iassist quarterly iq vol. 27 issue 1. three articles are presented in this issue. the article from ken reed, betsy blunsdon, nicola mcneil (deakin university, victoria, australia) and steven mceachern (university of ballarat, victoria, australia) was presented at the iassist conference (may 2003, ottawa) in the session with the title “strengthening numbers: measurement, aggregation and policy”. the reed, blunsdon, mcneil & mceachern article has the title “integrating public domain data to construct community profiles”. the paper describes the construction of a database of community characteristics at the postcode level. a community is described by variables in the areas of economic, social, political and cultural characteristics. these publicly available data from a wide variety of sources are in the project collated and integrated. the data are supporting multi-level research enabling the contextualization of individual behavior by data on the context in which people live their daily lives, such as the neighborhood, school, community or region. there is a great deal of data available in the public domain, but these are collected for a wide range of purposes, and with great variation in the units of analysis. the project on local areas in victoria, australia relates to other neighborhood projects in the usa (los angeles and chicago). the database has information on the years 1996 and 2001, which are years in which there was both a national census and a national election. from a variety of sources data is included on businesses, crime and suicide, licensed premises, information about religious institutions, schools, recreational facilities, cultural organizations and services. the second paper is also from the may 2003 iassist conference in ottawa. the paper gives an insight into a new archive by the authors hannele keckman-koivuniemi and mari kleemola, both research officers at finnish social science data archive. the title of the paper is “data processing in fsd: challenges in a new archive”. the finnish data archive started in 1999 situated at the university of tampere and funded by the ministry of education. the archive has about thirteen employees and has archived about 400 studies. the archive is a social science data archive and receives data files in spss, excel, sas or ascii formats. intensive data processing is carried out at the fsd in order to make the data corresponding the original questionnaire. the data is documented in spss files, the integrated documentation is obviously considered important, but a more extensive documentation is also made through the use of the ddi. in the last article from one of the iassist key figures chuck humphrey past president of iassist and an active person at almost every iassist conference – you are invited to take a look into the major accomplishments of iassist in the past decade. this article is for both newcomers and old members of iassist. while looking back chuck humphrey is at the same time into the process of setting the goals for the future of iassist. in the past iassist has been active in helping to set up data archives and in obtaining members outside the traditional geographical areas of iassist (north america, europe (western) and australia). the outreach program that started in 1996 has brought in many new members to iassist conferences and spread the information about iassist. in the communication room the website on the internet was the biggest move in the past. apart from delivering services directly to the members, iassist has also been active as participants and initiators of metadata standards like the data documentation initiative (ddi). many more points are made in the article. read the article and contribute by articles to the iq or by stating your opinion on the iassist mailing list. remember to visit the iassist website on www.iassistdata.org. papers for the iassist quarterly are most welcome. please contact the editor (kbr@sam.sdu.dk). karsten boye rasmussen, november 2003. http://www.iassistdata.org/ mailto:kbr@sam.sdu.dk newsletter vol. 4 no.z why this distinction? user services are usually the raisoii d'etre of the facility, and the reason for its conhnued funding. apparent services are directly user-oriented transparent services, although often refinements of apparent services, are more often stafforiented, and only indirectly contribute to that desirable phenomenon--the satisfied user my conclusion is that, at the outset, a data archive/library should concentrate on apparent user services, in order to culhvate an active and supportive user community, and only later in its development, when this has been accomplished, develop transparent user services. reference tools for machine-readable data files john g. kelp laboratory for political research university of iowa it would appear to be a rather awesome task to try to summarize the current state-of-the-art in data reference; for surely after 16 years of steady growth in archives, reference matenals giving access to the data in such archives should be exhaustive yet this is clearly not the case, although some attempts to provide useful reference tools in this area have appeared over the past decade. in the brief report which follows, i have singled out for discussion three categories of reference tools for machine-readable socal science data which seem well-established those reference tools which appear to play the most prominent role in directing users to appropriate machine-readable data are: (1) data catalogues describing the contents of individual social science data archives and data libraries; (2) directories describing the contents of more than one archive, or directories within special topical areas; and (3) penodicals, like .; s data, which have attempted to report at regular intervals informahon on the holdings of social science data archives an attempt will be made to examine each category of reference tool from two very personal perspectives: (1) the user consultant who is continually asked to locate data files which must meet a number of very special condihons; and (2) the editor and compiler who must try to ickate all known data files relating to certain topical areas and acquire informabon on the most recent acquisitions of data archives in the former role, one is frustrated by the inadequacies in reference tools; in the latter role, one is amazed that we have come as far as we have. data catalogues lists of holdings, guides to resources, archive directories, or data catalogues are available from most individual data repositories probably the most wellknown of those documents which describe individual archives would be icpsr's guide ic resources and scn'ices, issued on a nearly-annual basis since sometime in the i960's. others of this j(enre would include the recently-issued ssrc survey archive data catalogue, the sleiiimctz archwes: catalogue and guide ((1978), the b-a. s.s. invenlairc des archives disponsibies (1975), catalcf^ue of machine-readable records in the national archives of the united states (1975), and the lokahseringsoversi^l of the danish data archives (1978), to mention just a few. physically, documents of this type are soft-cover, book-length descriphons of the holdings of a major archive. these are real publications, meant for broad distribuhon to a national or international clientele. entries are often arranged according to some broad subject classification scheme which might include such major headings as community and urban studies, elites and leadership, mass political behavior, social welfare, religion, the international system, legislative and deliberative bodies, etc. the individual entries usually include title, author or data collection agency, population covered and/or sampling scheme, time period of study, number of cases and variables, distribution restrictions, and a brief abstract summarizing the purpose of the study and the focus of major categories of variables. newsletter vol. 4 no.2 another type of data archive catalogue is common to a group of archives which serve primarily as regional, provincial, or state-viide resource facilities these are designed for consumption by individuals beyond the local computing environment as well as scholars from many departments on the local campus. such catalogues may group entries according to the subject classifications mentioned above, although length and detail of information contained in the entries may vary considerably. for example, the directory issued in 1978 by the data and program library service at wisconsin lists studies according to a reasonably detailed classification scheme, although each individual entry has very limited information. the university of british columbia data library catalogue (1974) has data files arranged alphabetically by title and individual entries are described wnth considerable detail the annotated lislin;^ of data holdinf^s of the soaal science data library at north carolina organized entries according to a fairly detailed subject classification and at the same time included lengthy descriptions of each data file the final type of data archive catalogue is the one usually generated by the data library primarily to serve users of the local computing environment. like those mentioned above, method of organization and detail in terms of individual data file descriptions vary considerably. most are mimeographed, multilithed, or reproduced by some inexpensive method; many in fact are produced by line printers from machine-readable bibliographic files of various types. in my filing cabinet are such documents as the indiana university political science data archwe holdings as of may i, 1971, annotated listing of data holdings, polimetrics laboratory. october 1975. file inventory (latin american data bank, february, 1973), compendium of data holdings (center for quantitative studies in sodal science, univerity of washington, 9/76), data file descriptions (public affairs information service, university of missouri, 5/2/74), "untitled printout" (project impress, dartmouth college, november 1978), and data holdings (1976) with updates (university of iowa) that is a brief review of the types of documents describing individual data archives. now let us try to determine how useful each of them would be if we had to use them to search for a data file with previouslydefined characteristics in my view, the usefulness of these catalogues is solely dependent upon (1) the scheme used to arrange the order of the entries in the catalogue; and/or (2) the types of indices that are appended to the list of data files. data file descriptions or indices can be arranged by different schemes and below are listed twelve such schemes that have been used in various data catalogues. the schemes are listed according to my own judgement as to their usefulness as a reference tool; "1" is most useful; "12" is least useful. 1 subject or substantive categories (and subcategories) produced from keyword descriptors assigned to the data file independent of words in the title 2 unit of analysis or universe sampled 3 kwic or kwoc of words appearing in the title 4 geographical area or location of sampled units 5 year or time penod to which the data refer 6 author or pnndpal investigator 7 data collection agency 8 depositor 9 first word of title 10 year data collected or archived 11 sample size or number of units 12 study number or archive accession number no doubt, one may wish to disagree about the priorities suggested in this list, although 1 do hope that most would agree that "subject or substantive categories produced from keyword descriptors assigned to the data file independent of words in the title," is one of the most important reference schemes. of course, if we would all heed sue dodd's advice on the construction of titles. no. 3 could be equally as important. it should be noted that i have placed author, principal investigator, data collection agency, and depositor indices some distance down the list. this is for a very good reason. in the old days (the late 60' s) it was frequently the case that data files were very closely identified with a particular individual or research team, and thus references so arranged made a good deal of sense today, however, the number of data sets residing in archives is so large that personal references are starting to fade and the need for references which key on subject classifications, unit of i $t newsletter vol. k no.2 analysis, and geographical area seem clearly of more importance. a second area of note is "year or time period" which i have ranked no 5 my academic training as an historian is, no doubt, part of the reason for the place this item holds on the list, although time has become a more important factor in sodal science research in the past few years time is important because of advances in the way we ask survey questions and in the way samples are drawn from populations; older surveys may be less likely to contain the kinds of questions a researcher wants on a certain kind of population. also the hme periods over which certain kinds of surveys have been taken, or certain kinds of questions asked, continue to grow longer, and we are now starting to build up impressive sets of files for longitudinal analyses. in addibon, the analyses undertaken on these sets of longitudmal files can themselves be used to build up chronologically-ordered, aggregate files for sophisticated time-series analyses at the bottom of my list are things like "the year the data was collected" and "study number" and "archive accession number." these may be important things to know about a machine-readable data file, but they are not, in my view, of any help as a reference item in locating such files. these references are extremely useful for internal archive purposes, but do little to help those outside the archive find appropriate data they are tools related to the acquisition process, not the data reference process. the aforementioned priorities in data reference, however, would seem to be somewhat contrary to the actual practice of individual data archives. upon examination of the data catalogues from thirteen different archives prepared over the past six years, one finds that six of the archives have used the internal study number as either the method of ordering the entries in the catalogue or as the distinguishing feature of a separate index. another six have used the first word of the title, nine have used author or principal investigator indices, five the geographical area, and only two have used date of the study or time penod on the other hand, five have used a good subject classification as the basis for ordering entries in their catalogues, while only another three have included useful subject or substantive indices. only three provided indices to the unit of analysis—a reference which i think is extremely useful. one does not wish to name names here, but it seems appropriate to award a first pnze to the steinmetz archive for the most indices--nine to be exact five of the archives provided no indices al all with their data catalogues, although three of these ordered their data entries in such a way as to make the publication somewhat more useful than they might otherwise appear. this discussion of the usefulness of various indexing schemes should not be taken as specific condemnation of the data catalogue of a particular archive, but has been developed to demonstrate that individual archives may not have always given adequate thought to the way in which descriptions of their holdings will be used by those outside their local computing environment in addition, some of this talk may prompt further discussion along these lines and perhaps eventually some recommendations directories a data directory is distinguished from the data catalogues or the lists of holdings just described by one important feature -it attempts to provide a useful reference tool to the machine-readable holdings of more than one social science data archive or data library. by this definition, few so-called directories would remain in this category one would be selections from the "directory of directones," which appeared in the lasslst newsletter in 1977; others are the national technical information service's directory of compulerued data files and related software, the directory of federal agency education data tapes, and vivian sessions' directory of data bases m the soaal and behai'ioral sciences. all of these directories apparently contain references to data held by more than one repository, although in some cases these separate repositories may all be in the united states federal government. the sessions volume is probably more widely known among data librarians, so it may be appropriate to discuss its utibty as a data reference tool. published in 1974, and clearly out-of-date now, it is nonetheless of considerable interest because of what it attempted to do. the "major thrust of this directory," according to the introduction, was "the identification of the nonhibliographic data bases." the title of the volume would suggest a concern with the sodal and behavioral sdences, but a number of factors undermine the promise of the title. first, data bases were apparently defined as any systematic collection of data, primarily but not exclusively in machine-readable form second, newsletter vol. 4 no.2 whether these data bases were in the social and behavioral sriences was left to the reporbng agency third, those data centers that were included in the volume (and there were over 600 of them) were also self-ascribed fourth, major subject classifications in the index represented a strange mixture of academic disciplines and sub-categories with the practice of pubbc administration. and fifth, "it was inevitable that the primary organization of this directory is by data center." when these five factors are evaluated alongside indices concerning institutional names, data center personnel, and geographical location, one comes up with a rather limited reference tool the volume has little to do wnth what would normally be thought of as the "social and behavioral sciences" and in fact concentrates on municipal, county, regional, and other governmental agencies directly involved in the planning or administration of public programs. it is also not really a directory of the data bases, but rather a directory of places that collect and/or store data for purposes other than those normally ascribed to the physical sciences. to my knowledge, no directory of this type has been attempted since publication of the sessions volume had one been published, we can imagine that we would have wanted it to be organized and indexed in much the same manner as a catalogue describing the holdings of a single data library. the entnes might be organized according to some general subject classification scheme, with four or five indices focusing on: (1) a more detailed subject classification scheme; (2) unit of analysis; (3) words appearing in title; (4) geographical area; (5) time period of study; etc anyone considering a project to provide a general guide to machine-readable data in the sodal sciences or a directory on some special subject area would be welladvised to study the sessions volume as well as the methods of organization and indexing found in data catalogues of the major data archives and libraries. data periodicals periodicals issued on a regular basis provide one mechanism whereby information on the recent acquisitions of social science data archives can be transmitted to users of data reference matenals s s data is the prime example of such a periodical and the one that will be discussed in the remainder of this paper for those not aware of its history and purpose, let us begin here. s s data began publication in september, 1971, under support from a two-year grant from the national science foundation its purpose, as stated by g.r. boynton in the initial grant proposal, was to fill a much-needed gap in reference information on the holdings of sodal sdence data archives m the united states and abroad. the grant proposal also envisioned that s s data was only a stop-gap measure—something to fill the information void for two or three years until a better system was developed by leading professionals in the data library field. the plan suggested that s s data would reference on a quarterly basis the new acquisitions of all academic socal sdence data archives in the united stales and as many canadian and european archives as possible. it was estimated that this would be 40 or 50 new data sets each quarter or about 180 per year. while this figure probably over-estimated the acquisitions of academic archives in the united states (exclusive of the roper center), it was clearly an under-estimation of potential world-wide acquisition activity. the audience for this newsletter was thought to be first and foremost, sodal saentists, and second, data librarians and individuals involved in reference activities that is what s s data was supposed to be; what was it initially and what has it become over the past 8 1/2 years? first, the periodical could not cover comprehensively the acquisition activities of all acaderruc data archives m the united states for two basic reasons: (1) it was never possible at any one point in time to know which data archives were in existence and which ones were not; and (2) among those identified at any one time, the degree of cooperativeness in providing appropriate reference materials for publication varied a great deal of the approximately 25 archives who were initially contacted about their partidpation in the newsletter, only 18 provided a positive response to that request. some never were heard from and apparently had gone out of business; others simply refused to answer their mail, although it was clear that the archive was still in operation. those archives that did agree to partidpate provided information on their holdmgs at irregular intervals or in spurts. during the first few years, for example, archives at wisconsin, iowa, illinois, york, and icpsr and the roper center provided information on a fairly regular basis as other archives joined the list of partidpants it became more common for an archive to simply send a copy of its latest data catalogue and entries would be selected from these for each issue of the newsletter a ist newsletter vol. 4 no. 2 ffw. like till' st.itf d.it.i progr.im .it berkeley, the drug abuse epidemiology d.itn center, and icpsr established a regular procedure for transmitting information on new aequisitions, althougfi all except the drug abuse archive have stopped doing so. new acquisihons at the consortium, for example, are now only identified through new catalogues, news releases, announcements, etc., which are received second-hand from other faculty and staff at iowa most archives do send the most recent issues of their data catalogues, and it is through these that most of the entries in s s diilit are obtained occasionally, however, batches of new acqusihons and announcements come dnftmg in from archives which are nol otherwise heard frtim on a regular basis. because of the various problems invoked in obtaining up-to-date informanon on the recent acquisitions of data archives, the newsletter has evolved into something a bit different than was onginally envisioned. while s s diilii still tries to identify and report new acquisitions, increasingu' the focus has shifted to the reporhng on data files related to common areas of investigahon two years ago, this shift in emphasis was offically noted in the newslelter--previouslv data sets had been reported according to the discipline of the pnnapal inveshgator now they are reported according to sub|ect areas within disciplines, or by areas of concern in public policy formulahon, or bv new fields of study that represent interdisaplinarv approaches for example, rather than simply grouping a set of files under ;h'/i/ira/ sciciicv, they are now listed under such headings as elections, political attitudes, or lei^nlalivc cltlei--the latter being an interdisciplinary grouping of studies conducted by both histonans and political saentists. some recent public policv listings have included the elderly, /lousniv;. ccnscmatioii. and iransportatioii and travel; while some addihonal interdisciplinary listings have included mobility studies, legal studies, and demography in addition to listing data files under topical headings, an attempt has been made to reference (when possible) related data files which have appeared in previous issues a recent issue, for example, included a section on national legislative elites which not onlv included descriptions of nine data files but also referenced sixteen other files appeanng in past issues. a few readers have commented that they find this new method of listing and referencing past entries to be very helpful, but the majority of readers have remained silent on the issue. readership is another factor which has changed dramatically over the past eight years during the twoyear grant period (1971-1973), the number of subscnbers surpassed ickx) the newsletter was distributed free to subscnbers in the united states during this period, and most readers were individual social saentists attached to colleges and universities. less than a quarter were probably inshtutional subscriphons when nsf support ended in 1973, subscnption fees of $2.00 for individuals and $4.00 for instituhons were initiated and these have remained the same since then. the number of subscnbers dropped rather markedly as soon as an actual fee was charged and in recent years has stabilized at about 400. the character of subscribers, however, has changed considerably. about one-third are lassist members who receive subscriptions as part of their annual membership in that organizahon another 50% arc institutional subscribers, including primarily university and college libraries, public and private research centers, data archives and libranes, and a few metropolitan public libranes. the remaining 20% are individual subscnbers, such as social saenhsts, information specialists, and planners and researchers in the private sector. today, s s data serves the data reference community and not pnmarily the individual researcher, social scientist, or community planner. in concluding this paper, what remains to be done is to reflect on the usefulness of s s data and the potential role of such penodicals in the future, s s data has not been and never wnll be able to cover comprehensively all the acquisihons of data centers in the united states, let alone those in canada or europe or the third world second, even when references appear in s s data they rarely contain enough detail to immediately initiate the data acquisition process, if desired; numerous technical details are missing as well as specific information on all variables. third, reference centers with a fairly complete set of data catalogues might find little need for periodicals like s s data fourth, although s s data has published descriptions of nearly 1000 data files over the past eight years, one would need to have a complete set of back issues and a cumulative index to easily locate references to particular kinds of files. given these criticisms (and all of them and more have been leveled ats s data since it began publication), what is the present and future role of such a penodical? the only real contributions that s s data can make in its present formal are: (1) it can highlight the recent acquisitions of archives who are willing to provide such information; (2) it can gather together data files relating to specific disciplinarv, interdisciplinarv, or 1st newsletter vol. a no.2 public policy areas; and (3) it can relate sets of data files to those of similar focus that have been described previously. if those functions are enough to justify its conhnued publication, then the newsletter can go on for some time in the future however, it is important to raise the possibility that the print medium is now an out-dated mode of communication in the field of data reference. on-line systems at individual archives and projects whereby archives exchange data descriptor tapes certainly hold the promise of a much improved data reference system. networking and other developments, such as improved cataloguing of machine-readable data files, also offer interesting possibilities. it is hard to imagine that one system will dominate the field of data reference in the future, and therefore one important function that organizations like lassist can play is to recommend ways in which new and old reference systems can be effectively integrated in the future. on-line reference tools for the hard sciences gordon h. wood canada institute for scientific and technical information national research council of canada ottawa, ontario i ir toduction it is one thing to know or suspect that collections of machine-readable numerical data pertinent to one's discipline or problem may exist, it is something else to find that data, gain access, and use it profitably in ones research. the purpose of this paper is to review briefly the methods presently used to generate and access numeric data bases relevant to the so-called "hard saences " speaal emphasis will be given to the areas of on-line data retirieval and manipulation--areas where the state of the art in the "hard sciences" is generally conceded to be ahead of that in the "soft sciences." the reader wishing an inventory and description of the many scientific/technical data bases that are available worldwide is referred to the references at the end of the paper. for the sake of clarity, it is useful to define a few terms as they will be used in this paper. a. (scientific/technical) numeric data base an ordered collection of numbers whose values: 1. correspond to various properties, parameters or attributes of elements, substances or systems. 2. are critically evaluated by experts prior to their being included in the data base. b. numeric data base system a numeric data base system consists of one or more machine-readable scientific/technical numeric data bases as defined above plus: 1. programs for searching, retrieving and organizing the data according to user selected criteria and, usually, 2 programs to manipulate the data. (in general the latter property is what sets a numeric data system apart from a simple handbook or compendium. for example, a search routine may retneve data giving the co-ordinate positions of the atoms in a given crystal a simple command permits the user to calculate the various interatomic distances and the angles between the bonds joining the atoms. another command generates a two dimensional dravking of the crystal projected along any desired axis or plane.) 4 iassist quarterly 2016 iassist quarterly editor’s notes our world and all the local worlds welcome to the first issue of volume 40 of the iassist quarterly (iq 40:1, 2016). we present four papers in this issue. the first paper presents data from our very own world, extracted from papers published in the iq through four decades. what is published in the iq is often limited in geographical scope and in this issue the other three papers present investigations and project research carried out at new york university, purdue university, and the federal reserve system. however, the subject scope of the papers and the methods employed bring great diversity. and although the papers are local in origin they all have a strong focus for generalization in order to spread the information and experience. we proudly present the paper that received the 'best paper award' at the iassist conference 2015. great thanks are expressed to all the reviewers who took part in the evaluation! in the paper 'social science data archives: a historical social network analysis' the authors kristin r. eschenfelder (university of wisconsin-madison), morgaine gilchrist scott, kalpana shankar, and greg downey are reporting on inter-organizational influence and collaboration among social science data archives through data of articles published in iassist quarterly in 1976 to 2014. the paper demonstrates social network analysis (sna) using a web of 'nodes' (people/authors/institutions) and 'links' (relationships between nodes). several types of relationships are identified: influencing, collaborating, funding, and international. the dynamics are shown in detail by employing five year sections. i noticed that from a reluctant start the amount of relationships has grown significantly and archives have continuously grown better at bringing in 'influence' from other 'nodes'. the paper contributes to the history of social science data archives and the shaping of a research discipline. the paper 'understanding academic patrons’ data needs through virtual reference transcripts: preliminary findings from new york university libraries' is authored by margaret smith and jill conte who are both librarians at new york university, and samantha guss, a librarian at university of richmond who worked at new york university from 2009-14. the goal of their paper is 'to contribute to the growing body of knowledge about how information needs are conceptualized and articulated, and how this knowledge can be used to improve data reference in an academic library setting'. this is carried out by analysis of chat transcripts of requests for census data at nyu. there is a high demand for the virtual services of the nyu libraries and there are as many as 15,000 annual chat transactions. there has not been much qualitative research of users' data needs, but here the authors exemplify the iterative nature of grounded theory with data collection and analysis processes inextricably entwined and also using a range of software tools like filelocator pro, textcrawler, and dedoose. three years of chat reference transcripts were filtered down to 147 transcripts related to united states and international census data. the unique data provides several insights, shown in the paper. however, the authors are also aware of the limitations in the method as it did not include whether the patron or librarian considered the interaction successful. the conclusion is that there is a need for additional librarian training and improved research guides.. the third paper is also from a university. amy barton, paul j. bracke, ann marie clark, all from purdue university, collaborated on the paper 'digitization, data curation, and human rights documents: case study of a libraryresearcher-practitioner collaboration'. the project concerns the digitization of urgent action bulletins of amnesty international from 1974 to 2007. the political science research centered on changes of transnational human rights advocacy and legal instrumentation, while the libraries’ research related to data management, metadata, data lifecycle, etcetera. the specific research collaboration model developed was also generalized for future practitioner-librarian collaboration projects. the project is part of a recent tendency where academic libraries will improve engagement and combine activities between libraries and users and institutions. the project attempts to integrate two different lifecycle models thus serving both research and curatorial goals where the central question is: 'can digitization processes be designed in a manner that feeds directly into analytical workflows of social science researchers, while still meeting the needs of the archive or library concerned with long-term stewardship of the digitized content?'. the project builds on data of urgent action bulletins produced by amnesty international for indication of how human rights concerns changed over time, and the threats in different countries at different periods, as well as combining library standards for digitization and digital collections with researcherdriven metadata and coding strategies. the data creation started with the scanning and creation of the optical character recognized (ocr) version of full text pdfs for text recognition and modeling in nvivo software. the project did succeed in developing shared standards. however, a fundamental challenge was experienced in the grant-driven timelines for both library and researcher. it seems to me that the expectation of parallel work was the challenge to the project. things take time. in the fourth paper we enter the case of the federal reserve system. san cannon and deng pan, working at the federal reserve bank in kansas city and chicago, created a pilot for an infrastructure and workflow support for making the publication of research data a regular part of the research lifecycle. this is reported in the paper 'first forays into research data dissemination: a tale from the kansas city fed'. more than 750 researchers across the system produce yearly about 1,000 journal articles, working papers, etcetera. the need for data to support the research has been recognized, and the institution is setting up a repository and defining a workflow to support data preservation and future dissemination. in early 2015 the internal center for the advancement of research and data in economics (cadre) was established with a mission to support, enhance, and advance data or computationally intensive research, iassist quarterly 2016 5 iassist quarterlyiassist quarterly and preservation and dissemination were identified as important support functions for cadre. the paper presents details and questions in the design such as types of collections, kind and size of data files, and demonstrates influence of testers and curators. the pilot also had to decide on the metadata fields to be used when data is submitted to the system. the complete setup including incorporated fields was enhanced through pilot testing and user feedback. the pilot is now being expanded to other federal reserve banks. papers for the iassist quarterly are always very welcome. we welcome input from iassist conferences or other conferences and workshops, from local presentations or papers especially written for the iq. when you are preparing a presentation, give a thought to turning your one-time presentation into a lasting contribution. we permit authors 'deep links' into the iq as well as deposition of the paper in your local repository. chairing a conference session with the purpose of aggregating and integrating papers for a special issue iq is also much appreciated as the information reaches many more people than the session participants, and will be readily available on the iassist website at http:// www.iassistdata.org. authors are very welcome to take a look at the instructions and layout: http://iassistdata.org/iq/instructions-authors authors can also contact me via e-mail: kbr@sam.sdu.dk. should you be interested in compiling a special issue for the iq as guest editor(s) i will also be delighted to hear from you. karsten boye rasmussen june 2016 editor http://iassistdata.org/iq/instructions mailto:kbr@sam.sdu.dk 48 iassist quarterly the development of a canadian union list of machine readable data files (culdat) this article is an abridged version of the final report, "pilot project for the development of a canadian union list of machine readable data files (culdat)," prepared by edward h. hanis, social science computing laboratory, university of western ontario for the machine readable archives, public archives of canada. a survey of canadian social scientists, undertaken in 1982, indicated that a need existed for an inventory or union list of data files available for secondary analysis. in the mid-seventies, the data clearing house for the social sciences (dchss) had spent considerable time and effort in the development of an automated inventory. the loss of dchss, due to lack of funding, unfortunately also involved the physical loss of the magnetic tape which held the descriptions of these files. the results of the 1982 survey indicated strongly that the research community continued to feel thai a union list of data files would be a valuable resource. in response to this need, the machine readable archives division (mra) [of public archives canada] established a contract with the social science computing laborator>' of the university of western ontario to develop an online inventory describing computer files held by canadian data archives and libraries. the overall purpose was to develop organizational, technical cind informational foundations for maintaining and disseminating a computerized inventory. specific objectives involved: the establishment of a standard for describing mrdf for entry into the data base; the design and implementation of the pilot data base containing a partial inventory; and the definition of the organizational roles and mechanism to effect routine and cost-effective flow of descriptive information from data archives and other organizations to the imion list beyond the conclusion of the pilol the pilot project was carried out over a fourteen-month period. in january of 1985, a committee of data archivists and data librarians established a list of elements which were to be used to describe the holdings of the institutions. these elements were taken from those defined in the marc format for data files. a data dictionary was developed to aid participants in the entry of descriptive information. the social science computing laboratory was involved in six major activities: the aeation of the pilot data base; the set-up of online access with basis on the lab's vaxll/785; the set-up of datapac and standard dial-up communications; conducting an evaluation of the online system; a survey of potential contributors; and production of a hard copy reference doctimenl contributors to the data base were from the university-based archives and libraries and included: data library, university of british columbia; the institute for social research, york university; data resources library, university of western ontario; institute for social and economic research, university of manitoba. the mra also contributed descriptive entries. in all, 753 records were entered into culdat. evaluation of the data base was extended to more participants than those listed above and included both frequent users of online systems as well as infrequent users. although a number of suggestions have summer 1986 iassist quarterly 49 been made as tohow to improve the online inventory, the general consensus was that the data base was very useful and shotild be continued. it is not siuprising that the most crucial component of the data base was the description of the data file. a number of difficulties were experienced with the lack of consistent terminology used and the detail of the description itself. the problems encountered are simunarized in the following paragraphs. the resolution of these difficulties have formed the basis of the culdat work plan for 1986-87. the choice of data elements to be included in culdat was based on the fields in the marc format for data files. a limited number of elements was chosen as it was felt by the conomittee that the intention of the data base was to include only sufficient information to identify a imique data file, to aid researchers in selecting files of interest, and to locate archived copies of the file. the resulting culdat data element dictionary contained the field names and a brief description. during the pilot project, it was noted that in some cases the data dictionary did not provide sufficient guidance to the archivist or librarian to allow him to adequately describe data files, and presumed a knowledge of the marc format and anglo-american cataloguing rules ii. this cteated some difficulty in mapping out the information received for input into culdat. the consequences of a weak data element dictionary are inconsistent presentation of the information which can make the descriptions difficult for the end user to interpret weak data descriptions yield inefficient indexes, which, in turn, require that the user anticipate all possible variations of a term in order to find all relevant records in the data base. specific problems were fotmd in the following data elements. 1. investigators : the differentiation between principal investigator and other investigators caused some difficulties for both cataloguers and the users. the determination of principal investigator for a data file is difficult, if not impossible, at times. the separation of these fields requires searching two fields rather than one for the user wishing to browse the index. the distinction between investigator (personal) and investigator (corporate) was considered essential. the lack of authority control in the corporate investigator field was a problem which could be overcome through the use of canadiana to control the use and spelling of names. producer: generator. distributor : a tendency to repeat the same data in these fields was found. this may have been due to the inadequacy of the data dictionary. abbreviations and acronyms were used. the adoption of an authority file for corporate names would apply to these fields as well. file size: number of cases : some difficulty was experienced in the data provided in this field. again, this was due to lack of guidance in the data dictionary. access restrictions : as all institutions have their own access regulations, it was felt that this field shotild only be completed when the distributing organization has contributed the record. abstracts : information contained in this field was found at times to repeat information found in other fields. the vocabulary used varied widely which made control of the field extremely difficult the types of variables used in a data file is vital information for the prospective user. in order to provide improved access to this field, it would be summer 1986 50 iassist quarterly preferable to separate the abstract from the variable list variables could then be left unindexed. such a change would significantly reduce the indexing overhead and improve the quality of the printed keyword index by using variable names instead of individual words. the online system could continue to index variables as individual words as well as expressions. geographic coverage : the pattern adopted by the pilot was as follows: site, city, region, territory, province, state, country (quahfier) continent the pattern worked well in most cases and ensured that the user interested in data about a particular province could retrieve information on a file which covered only a city in that province. the only records which do not conform to this pattern are physical data where orbital coordinates are submitted. chronological coverage : the format of the dates recorded in this field was inconsistent, rendering the retrieval of data ineffective. the data dictionary should prescribe one acceptable format to which all dates would be converted. a standard format will provide the possibility of performing systematic retrieval on time periods by scanning the text, even though every unit of time within a range is not actually recorded in the field. the difficulties which have been encountered will provide valuable information to allow us to improve the quality and guidance required for the data dictionary. the second version should improve the consistency of the descriptive entries. the conuibutions made by the data archives and libraries were extremely useful in building the pilot data base and allowing us to identify specific areas for improvement in the data dictionary. user evaluation and potential contributors the original project design called for online testing and evaluation of the pilot culdat data base by project participants and constructing a list and contacting potential contributing organizations in order to learn about their holdings and interest in submitting entries into culdat in the future. three important additions were made to enhance the project the establishment of a datapac service reduced usage costs and significantly improved convenience to remote users. in addition, the survey of contributors was expanded to include questions on evaluation as they were potential users as well. the third activity was to include three local university of western ontario groups (students in the school of library and information science, the university's reference librarians, and social science researchers who use the lab's support services). these additions increased the use of culdat during the pilot phase. the evaluation of the data base was very favourable and many respondents expected to benefit from the availability of culdat in the future. considerable information from prospective contributors and users was acquired. this information and experience provide a sound foundation for the design and planning of the next stages in the development of culdat. the major activities planned for 1986/87 will include: 1) the revision and expansion of the culdat data dictionary in order to provide more guidance on the description of holdings for entry into culdat; 2) continued support to university based data archives and libraries to ensure their holdings are included in the inventory; and 3) the redesign of the formatted hardcopy version to make it available as a reference document at less costn summer 1986 24 iassist quarterly 2014 iassist quarterlyiassist quarterly abstract data play a critical role in fulfilling the federal reserve board’s mission across a broad range of functions, including monetary policy, financial stability, supervision, consumer protection, and economic research. the current data environment was designed to allow the various business functions to manage relatively small and predictable data sets that required limited sharing across departments or functions, effectively creating business-based silos. the office of the chief data officer was created in may 2013 to address the data needs of the board, postfinancial crisis, with an enterprise focus and a clear set of mandates to enhance data governance, data management, and data integration. the ocdo started operations with a small staff of data management professionals who traditionally supported the research function and now must shift gears to provide data services to a broader base of users and a wider range of analytical work. new infrastructures, programs, processes and staffing are being developed and deployed to ensure that data needs across the lifecycle are met and that a variety of analytical approaches can be supported. keywords: research, data management, governance, dissemination introduction data are the lifeblood of the board of governors, and the federal reserve system has always been a data-driven organization. data play a critical role in fulfilling the board’s mission across a broad range of functions, including financial stability, monetary policy, supervision, consumer protection, and economic research. as the board’s mandate has expanded in the wake of the financial crisis and the passage of the dodd-frank act, so has the need for an optimized data environment to meet the breadth and depth of analytical and supervisory challenges the board is addressing. the office of the chief data officer (ocdo) was created in may 2013 through the 2012-2015 strategic framework to address major data challenges at the board, to implement a data governance framework, to improve data management operations cross all divisions and board-delegated functions, and to strengthen the board’s information sharing environment to optimize the investment in data assets. there are a variety of projects in three strategic areas that must be undertaken to shore up the foundation of our data management operations and position the ocdo to provide the strategic thought leadership around data management and governance for which it was created. inventory and metadata critical to the board’s vision for improved data optimization is knowledge of the body of data and content within the board and across the system. more importantly, we need knowledge not just that data exist, but we need an understanding of their relationship to the organizational mission as well as the internal functions, services, and processes. this includes more complete and better-documented knowledge of information flows, system interconnections, and data security classifications. the overarching vision of the ocdo enterprise data inventory program is to provide the visibility and knowledge of data and their relationship to the people, processes and technology environment in order to support critical data needs and improve data management in a dynamic and building an enterprise data service at the federal reserve board by san cannon1 iassist quarterly 2014 25 iassist quarterly complex ecosystem. existing system data—collected, contracted, and created—have not been systemically or strategically cataloged. moreover, data that have been cataloged have not been cataloged consistently as there are no system-wide standards for doing so. the first stage of the enterprise data inventory (edi) project was recently completed, identifying necessary metadata elements for capturing relevant information about our data assets. these elements describe physical and logical datasets, roles and responsibilities of the individuals that interact with the data, the business processes and functions that include those data and the applications and technologies where the data reside. acquisition of a repository to hold these elements and a user interface to make things easy to find are part of the next project steps. data management data management is the business function that develops and executes plans, policies, practices, and projects that acquire, control, protect, deliver, and enhance the value of data. the board’s data reside in a number of operational and analytical systems distributed across the board and system, in varying states of modernization. there are industrial-strength systems that collect and manage data for some departments using time-proven methods and relational database technologies. in other areas, there are quick queries to collect data from financial institutions or other government agencies via spreadsheets that are stored on file shares. other departments are developing data warehouses and using new “big data” storage technologies. all these repositories meet local and departmental needs but none was developed with the goal of broader access as a key requirement. there has been a pressing increase in the post-crisis needs of the board to manage a much wider variety of data types across a multitude of data management platforms and to do a better job of sharing those data regardless of their storage location or technology. there have been suggestions in the past that a single enterprise data warehouse would solve all our storage and sharing problems. however, we have determined that there is no one-size-fits-all technology solution to the wide variety and volumes of data that the board is now responsible for managing. we are engaging in fundamental work will that look at appropriately matching data management technologies with data assets. the ocdo thinks strategically how to best manage the challenges facing data operations across the board. as an enterprise service provider, it also needs to build a strong foundation for its own data management operations. the goal is to provide scalable data management services to the board, and the system where appropriate, that can meet the increasing data capacity demands of the organization. the current staff have historically provided service primarily to the research community at the board only, but now need to manage data for and disseminate data to other departments with very different business models. the legacy systems for doing the basic intake and output will not scale to support the entire enterprise and a the ocdo has engaged consultants to best determine the technologies and work flow needed for this new level of service is well underway. data governance data governance is a core part of the ocdo’s mission. data governance is the exercise of authority, control, and shared decision-making (planning, monitoring, and enforcement) over the management of data assets. data governance supplies the discipline to deal with both the predictable and unpredictable nature of new data acquisition and data distribution across the organization. today, data governance activities occur in business lines with little or no formal coordination across the board and across board-delegated functions. little of this work is automated, and even then generally not beyond the use of sharepoint sites as document repositories. the ocdo will facilitate a variety of data governance activities through its data governance program (dgp). the dgp serves as a conduit for coordination, communication, and decision-making across and within business lines and departments within the board. the dgp has responsibility for developing the body of policies, processes, standards, and metrics; communicating progress and informing the organization about ongoing and completed activities; and, completing specific dgp actions, deliverables and activities. once foundational projects in these functional areas are underway, the board will be better positioned to undertake more strategic work to provide discovery services. to improve the board’s ability to find, layer, and explore a wide variety of data in unique ways is an important goal to better serve the board’s mission critical work which is more multi-disciplinary and cross-functional than it has ever been. we are planning work that supports the notion of discovery through the development of an enterprise-wide information architecture that shows the conceptual and logical relationships between different types of data regardless of the technology or application that supports the data. this includes the development of subject area models and taxonomies as well as the development of analytics, visualizations and knowledge management services. some of this has been done at the departmental level in the past but the ocdo focus is to provide such services for data users across the board. additional services other services will be needed to support the growing scope of work for an enterprise data service. the creation of a formal information architecture practice will lay the foundation for more complete data modeling and documentation of the logical and conceptual relationships between data assets. business analysts will help map business processes and evaluate workflow efficiencies. formal change management and communication will help staff to better understand what the ocdo can do for them, rather than worrying about what the office is trying to do to them. project management and training complete the service portfolio and will help keep work moving forward and users better informed about what data we have and how it can or should be used. conclusion while the federal reserve board is midway through the strategic plan that brought centralized data management and governance into being, the work will extend far beyond this strategic plan and will likely be an important theme in the next plan. when fully staffed at just under 50 staff members, the office of the chief data officer will have to be efficient, effective, collaborative and communicative to be able to serve as an enterprise service and work effectively in an institution with long standing business line traditions and work flows. notes 1.san cannon, deputy chief data officer, federal reserve board, washington, d.c. 20551 sandra.a.cannon@frb.gov 1.san mailto:sandra.a.cannon@frb.gov 4 iassist quarterly spring summer 2009 editor’s notes transformations slowly into the future welcome to this double issue volume 33 issues 1 and 2 (2009) of the iassist quarterly. this is a special issue centered on the developments of the data documentation initiative or its now familiar acronym: the ddi. we have slowly made some changes to the iassist quarterly (iq). our distribution has changed to being ‘web only.’ we have stopped printing but continued to make pdf versions available on the iassist website and you can access from the website http://iassistdata.org/iq the generated pdf versions of the journal. at the same time as producing new issues of the iq we have been helped by other people scanning old issues of the iq. right now michele hayslett at chapel hill is producing scans of the old issues. if you take a look at the iassist website you will notice that the look and feel of the website have also changed; it looks bright and sharp and you should hopefully find the interface to be more intuitive. another recent change is that the iq is now a fully reviewed journal and the iq also in the last years started to run special theme issues like this ddi double issue. these changes have demanded more work from more people and have only been possible because many people are doing iassist work in addition to their day jobs. finally access is in the process of being broadened as we have just signed an agreement with ebsco to make the iq accessible through their platform. in this issue the special editors mary vardigan from icpsr and joachim wackerow from gesis have collected from recent workshops and iassist conferences important papers on the development of the data documentation initiative as it reaches ‘ddi 3.’ in the following pages they present the articles: metadata-driven survey design by jeremy iverson, questasy: online survey data dissemination using ddi 3 by marika de bruijne and alerk amin, implementing ddi 3: the german microcensus case study by andias wira-alam and oliver hopt, metadata creation, transformation and discovery for social science data management: the dames project infrastructure by jesse m. blum, guy c. warner, simon b. jones, paul s. lambert, alison s. f. dawson, koon leai larry tan, kenneth j. turner, ddi 3 development at dda by jannik jensen and dan kristiansen, and controlled vocabularies for ddi 3: enhancing machine-actionability by taina jääskeläinen, meinhard moschner, and joachim wackerow. as iq editor i would like to thank all the authors and the special editors for their work. instead of introducing the papers this editorial will look backwards at some of the history of the ddi. a perspective could be riding on how bruno latour explains actornetwork-theory (1987, 2005). the ant-method is exemplified in latour's description of the development of the diesel engine. when mr. diesel presented the design of the engine he did not present a prototype. that took years to develop and relied heavily on many other technical skills and routines. early on, diesel took out a patent but nearly 10 years later the idea was close to collapse as the machines needed much attention and were continuously being modified. it took what latour terms as ‘translations’ and a complicated mixture of connections that translated the problems and solutions before the diesel engine was truly functional. before entering the description of these ‘translations’ latour instructs us: ‘in this technoscience game we are watching, the object is modified as it goes along from hand to hand. it is not only collectively transmitted from one actor to the next, it is collectively composed by actors’ [1987, p. 104]. similarly the ddi has slowly developed towards maturity. we find the roots of the development of the ddi in several data archives and individuals, but we also find them in collaborative organizations – amongst which iassist was one of the most important. before the foundation of the ddi, iassist established the ‘codebook action group’ in 1993. before that an even earlier development of a description standard at the study level had taken place. a report of the documentation activities at a selection of data archives was presented in the iassist quarterly (rasmussen, 1995). funding for further developments on the issues of standardization of social science metadata was established by icpsr in ann arbor and the ddi committee was formed and had its first meeting in may 1995. in 2000 i described and reported the progress of the ddi in a danish book; however, danish is not the most widespread language. it might be fruitful to remember that the ddi was focused on ‘independence.’ as a standard the ddi should be independent of platform, media, presentation, applications, and independent of commercial interests by being a standard without royalties. in latour terms we can say that the then current dependencies were transformed into independencies. much work and experiments were carried out with the ddi, but nearly a decade went by before articles directly addressing the ddi appeared in reviewed journals like historical methods (block & thomas, 2003), social science computer review (blank & rasmussen, 2004), and archival science (rasmussen & blank, 2007). when jacobs and humphrey in 2004 in communications of the acm presented their viewpoint on ‘preserving research data’ they also referenced the ddi. many people and organizations have contributed to the development or ‘transformation’ of the ddi, and many people have reported in journals, at workshops and at conferences notably the iassist conferences! one of the members of iassist quarterly spring summer 2009 5 the first ddi committee, mary vardigan, has also done a good job in making this special issue of the iq come into being. so a very special ‘thank you’ to mary and to her coeditor joachim wackerow. this month (july 2010) i met joachim (achim) in germany where gesis was having celebrations, one of these being the 50th anniversary of the zentralarchiv in cologne. there have been tremendous developments during its lifespan. technical developments are changing the possibilities in archiving and also in the relationships between people and roles e.g., in the collaboration between depositors, archive staff, and researchers. new people are also entering the scene and they will be the next ‘transformers.’ things take time. we can at times become impatient and this impatience can act as an extra driving force behind the slow transformations towards more mature solutions. with this latest report on the ddi development we can hope that the message will reach even more people. whether you are somewhat familiar with the ddi or a newcomer in the field of data documentation i hope you will find articles of interest in this issue. and you might be among the next people developing and disseminating the standards further. references blank, grant & rasmussen, karsten boye (2004): the data documentation initiative: the value and significance of a worldwide standard. social science computer review vol. 22-3, p. 307-318. block, william & thomas, wendy (2003): implementing the data documentation initiative at the minnesota population center. historical methods, vol. 36-2, p. 97 101. jacobs, james a. & humphrey, charles (2004): preserving research data. communications of the acm, vol. 47-9, p. 27-29. latour, bruno (1987): science in action. harvard university press. latour, bruno (2005): reassembling the social: an introduction to actor-network-theory. oxford university press. rasmussen, karsten boye & blank, grant (2007): the data documentation initiative: a preservation standard for research. archival science, vol. 7-1, p. 55-71. rasmussen, karsten boye (1995): documentation what we have and what we want: report of an enquete of data archives and their staff. iassist quarterly, vol. 19-1, p. 22-35. rasmussen, karsten boye (2000): datadokumentation. metadata for social videnskabelige undersøgelser. odense university press. articles for the iassist quarterly are very welcome. articles can be papers from iassist conferences, from other conferences, from local presentations, discussion input, etc. contact the editor via e-mail: kbr@sam.sdu.dk. karsten boye rasmussen july 2010 6 iassist quarterly spring summer 2009 guest editors’ notes welcome to a special double issue of the iassist quarterly featuring articles focused on the data documentation initiative (ddi), a metadata standard for the social sciences. we are proud to present these six articles, which explore various projects related to ddi 3 and its enhanced features. the articles draw on previous presentations and papers created in connection with the 2009 “expert workshop on implementation of ddi3 -advanced topics” held in wadern, germany; the 2009 european ddi users group (eddi) meeting held in bonn, germany; and the iassist conferences held in tampere, finland (2009) and ithaca, new york, usa (2010). jeremy iverson’s article on metadata-driven survey design highlights the reuse of metadata starting at the very beginning of the research data life cycle and also discusses the benefits of using metadata to drive the process of collecting, visualizing, and analyzing survey data. this is a powerful and efficient approach that should be taught in survey methods courses in order to save costs and to enable data producers to leverage the metadata they create across the life course of research data. also related to data collection is the article on the questasy online survey documentation tool by marika de bruijne and alerk amin. questasy permits internal users to document longitudinal data and to make this documentation available to external users on the web. a benefit of ddi 3 for this system is that it facilitates tracking of question items across waves in the study, where each wave can have question constructs and variables that refer to the same question item. this system was developed for the liss panel online survey at the university of tilburg in the netherlands. “implementing ddi 3: the german microcensus case study” by andias wira-alam and oliver hopt looks at using ddi 3 to document the german microcensus through a customized ddi 3 editor and a web view providing different perspectives for the end users based on the same ddi 3 items. interestingly, andias and oliver discuss basing some of their decisions about software design on jannik and dan’s use case describing the development of the ddi 3 metadata authoring tool – see building a modular ddi 3 editor. “metadata creation, transformation and discovery for social science data management: the dames project infrastructure” by jesse m. blum, guy c. warner, simon b. jones, paul s. lambert, alison s. f. dawson, koon leai larry tan, and kenneth j. turner shows the wide variety of data management tasks that ddi 3 can support and document, including recodes, merging, and data cleaning. using ddi 3 to document these phases of the data life cycle is an exciting development. “ddi 3 development at dda” by jannik jensen and dan kristiansen of the danish data archive provides a fascinating look into the development of an authoring tool for ddi metadata, a tool that is being designed to play a central role in the work flow at the dda archive. it focuses as well on the underlying reusable middleware and general considerations on open source software development for ddi. this article provides the reader with an up-close view of strategic decisions made at dda to incorporate the functionality of ddi 3 into the architecture of the dda archive. with its focus on machine-actionability and data typing, ddi 3 needs a strong system of controlled vocabularies to supplement the creation of metadata. the article on controlled vocabularies by taina jääskeläinen, meinhard moschner, and joachim wackerow presents the case for using controlled vocabularies and the ways in which they benefit the user. the article also showcases the work of the ddi controlled vocabularies group and its efforts to create vocabularies for ddi 3, which will be made available as separate products using a format called genericode. we hope you enjoy reading these articles, and we offer our thanks to all of the authors. we also want to express our appreciation to iassist for the opportunity to publish this work in the iq. we are grateful for the ongoing support of the iassist community and its nurturance of ddi from the very beginning. sincerely, mary vardigan and joachim wackerow the european voters study 1989 by cees van der eijk ' manfred kuechler herman schmitl theoretical perspectives the european voters study (evs) 1989 is a study of behavior, motivations, attitudes and perceptions of the electorates of the member states of the european community in the european parliament election of 1989— the third of its kind after 1979 and 1984. the objectives for designing and conducting a european voters study are twofold. first, a european voters study can be mainly looked at from the perspective of studying european elections and their place in the process of european integration. second, more generally, it can be viewed from the perspective of comparative electoral research. the perspective of european integration . protagonists of european integration have always showed great interest in the direct elections of the european parliament which took place for the first time in 1979. those who had lamented the slow pace of development of the european community, hoped that a directly elected parliament would provide a powerful stimulus to further integration. unlike the other institutions of the community, the parliament would have its own popular mandate and would exemplify by its very existence the desire of the citizens of \he member states to live in a unified europe. some of these expectations reflected a certain degree of naivety with respect to the immediate political effects of these elections. yet, the actual turnout disappointed not only the protagonists, but also startled more neutral observers. it was widely assumed that abstentions reflected a considerable degree of indifference or even opposition to the idea of european integration. no 'popular mandate' for further european integration could be inferred. in most countries the campaign was dominated by other, mosdy national political issues and concerns. the few exceptions to this general rule offered litde comfort from a pro-integration perspective: predominandy in denmark and to a lesser degree in great britain, party choice appeared to refiect a sizeable amount of anti-ec sentiment. the experience of 1979, reinforced in 1984, raised a number of questions concerning both turnout and party choice of european voters. reliable answers were needed in order to properly evaluate the implications for the future course of european integration. does low turnout reflect just a widespread lack of familiarity with the european parliament, is it just a visibility problem? or does it reflect a more fundamental feeling that the european parliament, and possibly the european community at large, is irrelevant or detrimental to the individual citizen's interests and concerns? are those abstaining from the european elections decidedly critical about, ot even downright hostile towards european integration in general and towards the european parliament in particular? what part do the political parties play? are they unable or unwilling to put europe on the national agenda, to channel and represent the ec related interests of their clienteles? to which extent, then, is the voters' choice between the parties an acknowledgment of specific party goals with respect to european integration? does party choice reflect different ec policy preferences or is it predominanuy determined by domestic considerations? obviously, contingent upon the answers to these questions, very different conclusions concerning the future course of european integration can be drawn. for most, if not all of these questions, survey data representative of the electorate at large are necessary to obtain answers solidly grounded in empirical evidence. the perspective of comparative electoral research . the study of elections and individual voting behavior is a very well developed area of empirical political science. in virtually all western democracies large scale surveys are conducted during election times to uncover the forces which shape voting behavior and thereby election results. however, there is considerable national variation in die depth (over time), quality, and accessibility of these data. the united states, great britain, and west germany have long standing traditions of scholarly election surveys which are generally available for secondary analysis. the situation in a number of otiier countries is less fortunate. still, a number of valuable attempts have been made to utilize national election studies from various countries for cross-national comparisons (see e.g. budge, crewe, and fairlie 1976; crewe and denver 1985; dalton, flanagan, and beck 1984). on the one hand, the volumes which document these efforts exhibit the strong common strands in the design and conceptualization of the various election studies. yet, on die other hand, they clearly reveal die discrepancies between tiiem. summer 1990 national election studies are indeed strongly national in character. to some degree this is unavoidable. the diversity reflects real differences with respect to systemic arrangements (e.g. electoral rules) and political culture. but this diversity is also due to (false) economy: questions which have litde explanatory value in a strictly national study, but which are essential to establish comparability with other countries are the first ones to be cut if such questions are considered at all. incompatibilities in overall research design, in choice of concepts, in manner of operationalization, in question wording and format, and — last but not least— in the demographics section are likely to continue for the noble cause of preserving national comparability over time. the situation, then, is somewhat paradoxical: while the field of electoral research is among the oldest, and certainly most developed areas of empirical social research, it has not generated the kind of large scale cross-national survey projects which have been so pivotal in the development of other areas of comparative mass political behavior (see e.g. almond and verba 1963; barnes and kaase 1979). background with the purpose of designing and organizing a truly comparative eiu^opean voters study to be conducted in 1989. subsequent meetings were held in mannheim in may and october 1987, which resulted in the formation of a group of six scholars serving as coprincipal investigators: roland cayrol, cees van der eijk, mark franklin, manfred kuechler, renato mannheimer, and hermann schmitl though not a formal member of the group, karlheinz reif was essential to the success of the project in providing good scholarly as well as very practical advice from the very beginning. most members of the core group were intimately involved in earlier studies of the european election. following precedence, cooperation was (re)established with other research teams focusing on the campaign (coordinated by oskar niedermayer, at the university of mannheim, west germany) and on the communication process (see e.g. blumler 1983). during the two intensive meetings in mannheim, the group hammered out a design of the european voters study to be, drew up a strategy for securing funding, and decided on some division of labor. previous work in the past, the eurobarometer surveys have been utilized in various ways to generate data related to the process of european integration. questions concerning electoral participation have been included in the surveys prior to and following the european elections of 1979 and 1984. questions relating to affective and evaluative orientations towards european integration, the european community, and its various institutions and policies have been included frequently in eurobarometer surveys and constitute an important pan of the 'trend' questions which are included in each wave. still, in spite of the wealth of material which has been collected, a number of important lacunae remain. these originate partly from the fact that certain questions were never included (e.g. questions assessing factual knowledge), and partly from the fact that the regular eurobarometer surveys take place too far before (march), and too late after (november) the point in time at which the european elections actually take place (june). likewise, various slu^'eys conducted at the occasion of previous european elections do not fill this void. they have focused on media effects and on various kinds of elites including party candidates running for seats in the european parliament (see e.g. blumler 1983; reif 1984, 1985; reif and schmitt 1980), but they did not center on the voting behavior of the electorate at large. internal organization and cooperation during the ecpr joint workshops of april 1987, first contacts were established between scholars of various in terms of internal organization two factors were essential. first, most valuable support was provided by the university of mannheim which made it possible for hermann schmitt to serve as the coordinator for the group. second, ample use of electronic communication via earn/bitnet compensated for the very limited opportunities for personal meetings of the entire group. geographical dispersion of its member and the lack of sufficient travel funds could not have overcome otherwise. not just with respect to travel, funding was a major problem continuously haunting the group. funds were secured from various sources, in various amounts, at different points in time. a major portion, covering the costs of the field work for the post-election wave, was supplied by the british economic and social research council (esrq. other funding sources include several national newspapers which were given priority publication rights of elementary, but timely analyses of part of the data. unfortunately, we could quite meet our funding objectives. this required several cuts and modifications in our original question program. in particular, some questions could not be replicated in all three waves as planned. still, the core of the original plan could be carried out. a series of questions were added to the core questionnaires of the eurobarometer surveys #30 (november 1988), #3 1 (april 1989) and #31 a (june 1989). matter of fact, the close cooperation with the eurobarometer proved to be an indispensable asset. without it the study could not have been completed. it gave us — and the hopefully 10 lassist quarteriy many more researchers to come— access to the standard eurobarometer questions and with the special edition of july 1989 (#31a) it provided a base for the post-election wave. design and contents with our theoretical focus on mass behavior, there was no alternative to a cross-national survey design. in addition, we felt that a purely crosssectional design would be inadequate (though much more feasible) in order to study the process of cognitive, attitudinal and behavioral mobilization. the choice, then, was between a genuine panel design and a series of repeated crosssections. without entering the sometimes vivid debate on the advantages and disadvantages of panels in contrast to repeated cross-sections, we quickly determined that a panel design was not fundable; that only buying into an established european siu^fey like the eurobarometer would bring cost for data collection within a feasible range. while not denying these very practical concerns, the repeated crosssections design does match our theoreticaj and conceptual interests. our emphasis was not on the dynamics of individual vote choice but on the preferences of groups and segments of voters, on the change of these group preferences over time, and on patterns of association. below, we will briefly outline the sets of variables included in the study. in terms of our prime target, turnout and party choice in the european elections 1989, a broader set of questions needed to be included. previous research had convincingly suggested that electoral behavior in european elections is to a large extent determined by national factors. consequently, intended national electoral behavior was to be probed as well. furthermore, drawing on theories on voter behavior and party competition developed in the context of the dutch national election studies (see e.g. van der eijk and niemoeller 1983), a more comprehensive assessment of the electoral attractiveness of all major parties was called for— with respect to both european and national elections. explanatory or independent variables fall into five categories. the ftfst category consists of variables which describe the voters' social situation; in particular, their location within the cleavage structure of each country. these are necessary for explaining behavior in terms of the traditional cleavage theories. these theories have come under attack in recent years, but the scholarly debate over the persistence of established cleavages is not over yet. also, these variables are needed as controls in assessing the effects of attitudes, perceptions, experiences, and general political behavior on turnout and vote choice. these variables do not attract much attention in national studies, they are mostly part of an estabushed demographic section. however, for a cross-national study they constitute a major problem. to deal with the pervasive problems of (in)comparability which traditionally plague researchers working with these characteristics, we drew on the ongoing work of another group (franklin, mackie, and valen, 1990). with a few additions, the set of demographic variables used in the eurobarometer met our needs. a second block of independent variables deals with substantive issue concerns. obviously, to the extent that issues play a role in voters' decision-making, they may arise from different contexts. at the least, the following kinds of issues have to be distinguished: a. community issues (extending ec membership, common agricultural policy, payments to and subsidies received from ec, etc.), b. supra-national issues (issues pertaining to all member states but not, or only partly related to the ec like defense, unemployment, etc.), and c. country-specific issues (the most salient of these were determined in close cooperation with additional country specialists). it is desirable to tap absolute and relative saliency as well as perceptions of party competence for each one of such issues, but this would require an inordinate amount of question time. as a compromise, we constructed a list of 12 issues (4 each of the three types mentioned above). each item was individually rated as 'very important' or 'not very important', then the respondent was asked to name the three most important ones. for these (up to) three issues, we further established which party was seen as best able to handle this problem. funding problems restricted the full approach to the second wave, while individual salience ratings were obtained in all three waves. the third block of variables comprises european orientations, which deal with the european community, its institutions, the idea of european integration, etc. many of the indicators which are regularly included in the eurobarometer questionnaires capture the affective components of such attitudes. in addition, we also tapped the cognitive and evaluative aspects of european orientations. a fourth block of explanatory variables deals with specific perceptions of the political parties contesting european and national elections. one set of such perceptions deals with the parties' position on europe, summer 1990 another with perceptions of the parties' location on a left-right scale. additional questions establish the respondents' own location or preference. a fifth and final block of questions deals with media exposure and information. here, we closely cooperated with another project focused on the communication process in the electoral campaign (see above) and followed their lead. most of these questions were replicated from the 1979 communications study (see blumler 1983). apart from some cuts within these five sets of questions due to funding problems, other aspects originally discussed had to be shelved altogether. these include questions dealing with possible candidate effects on party choice. no attempt was made to measure the elusive concept of party identification beyond the standard item in the eurobarometer questionnaire. however, the battery of questions in which the electoral attractiveness of all pardes is to be rated (see above) offers new options to construct possibly more valid operationalizations of this concept strategies for analysis and publications a number of initial analyses on the data from the first wave (eb30) have been presented and discussed in an ecpr workshop during the joint sessions in paris in april 1989. special panels at the annual meetings of the midwest political science association (chicago, april 1990) and the american political science association (san francisco, august 1990) have and will provide other opporiunities to present and discuss findings from this study. a special issue of the european journal of political research (planned for the second half of 1990) will contain a first set of cross-national comparative analyses by members of the core group. this will be followed by an edited volume with chapters on each of the ec member countries to which additional country specialists will contribute. it will also contain a second round of comparative analyses. to conclude this presentation, we will briefly discuss the general analytic strategy behind these pubhcation plans. at the same time, this discussion may also further productive use of this data base by other researchers in the future. as argued in more detail elsewhere (kuechler 1987), mass survey data provide an invaluable, but also inherently limited base for the study of mass (political) attitudes and behavior. in general, survey data do not just speak for themselves, they require a careful interpretation within the context in which they are generated. this holds for any (national) survey, but it becomes even more apparent in a cross-national setting. a question may have a different meaning in a different political and cultural system, even when great care is exercised in aiming at 'functional equivalence'. a comparison of marginal distributions across nations has some heuristic value, but it does not lead to meaningful theory construction. it is more useful to look for patterns of associations, e.g. the impact of degree of jrolitical interest on issue evaluations, and to compare on the level of these relationship patterns. in a way, we can look at such an analysis as an instantaneous eleven-fold replication of a relational hypothesis. our first round of analyses has produced few, if any hypotheses which can be successfully replicated this way. matter of fact, particular in the area of issue voting, we have come across a surprising number of sign reversals, i.e. the same two variables show a positive relationship in some countries and a negative one in others. this strongly points to the need to assess the survey data in the light of other countryspecific sources of information. detailed countryspecific analyses (the second stage in our analytic strategy) then go way beyond mere idiosyncratic description. their prime objective is a "cross-nationally informed country-specific" analyses which will focus on singular and deviating patterns. in turn, these will provide the base for a second, higher level of comparison. at this point it is premature to predict the possible returns from this three stage comparative strategy. we may find a considerable amount of higher level communality, or we may conclude that idiosyncratic systemic factors tend to dominate, severely curtailing efforts of location-independent theory building. but even if our group fails, a valuable host of data will be available to the other researchers with all sorts of brilliant ideas in the very near future. the social science community is fortunate to have the services of many fine data archives available. their supporting role is vital for the further growth of the social sciences. references almond, g. and s. verba (1963). the civic culture. political attitudes and democracy in five nations. princeton, nj: princeton university press. barnes, s., m. kaase et.al. (1979). political action: mass participation in five western democracies. beverly hills, ca: sage. blumler, d. (ed.) (1983). communicating to voters. television in the first european parliamentary elections. london: sage. budge, i., i. crewe and d. farlie (eds.) (1976) 12 lassist quarteriy party identification and beyond: representation of voting and party competition. new york: wiley. crewe, i. and d. denver (eds.) (1985). electoral change in western democracies: patterns and sources of electoral volatility. london: croom helm. dalton, r. j., s. c. flanagan and p.a. beck (eds.) (1984) electoral change in advanced industrial democracies: realignment or dealignment? princeton, nj: princeton university press. eijk, c. van der and b. niemoeller (1983). electoral change in the netherlands. empirical results and methods of measurement. amsterdam: ct-press. franklin, m., t. mackie and h. valen (eds.) (1990, forthcoming) electoral change: responses to evolving social and attitudinal structures in seventeen democracies. cambridge: cambridge university press. kuechler, m. (1987). the utiuty of surveys for cross-national research. social science research 16: 229-244. reif, k. (ed.) (1984) european elections 1979/81 and 1984. berlin: quorum. reif, k. (ed.) (1985) ten european elections: campaigns and results of the 1979/1981 first direct election to the european parliament. aldershot, uk: gower. reif, k. and h. schmitt (1980) nine second order elections, european journal of political research 8: 3-44 'paper prepared for lassist 90, annual conference and workshops, may 30 june 2, 1990, poughkeepsie, new york. cees van der eijk department of political science, university of amsterdam, netherlands manfred kuechler hunter college and graduate center, department of sociology (cuny) herman schmitt zeus, university of mannheim, west germanyyork. principal co-investigators (in alphabetical order): roland cayrol (france), cees van der euk (netherlands), mark franklin (united kingdom/usa), manfred kuechler (west germany/usa), renato mannheimer (italy), hermann schmitt (west germany). coordinator: hermann schmitt bitnet addresses : van der eijk (a7150204hasarall), franklin (polsljl@uhupvml) kuechler (makhc@cunyvm), schmitt (fs91@dmarum8) summer 1990 ^ sist newsletter vol.1, no. 1 i united states secretariat report judith rowe computing center i, princeton university the united states secretariat has received positive letters of intent from 170 people, more than half of whom are already involved in action group activity. by popular demand, we are organizing a north american working conference in florida this february. the accomodations will be pleasant and inexpensive and the meeting will provide an opportunity for the ag's to work intensively on their projects. full details will be sent out shortly. i would like to take this opportunity to thank each of our ag coordinators for the enthusiastic efforts they have already expended and to encourage all of you not only to join lassist, but to participate in its activitites. action group reports data archive registry canadalisa lasko, canadian (.onsortium for social research, institute for behavioral research, york university, 4700 keele street, downsview, ontario m3j 1p3 europejoseph bonmariage, belgian archives for the social sciences, university of louvain, sh-2, 1348 louvain-la-neuve, belgium united statesdavid nasatir, behavioral sciences graduate program, california state college, domingus hills, california 90747 mandate a directory containing names, addresses, types of holdings, and dissemination policies of existing data archives and libraries throughout the world will be compiled. a supplementary directory listing archival personnel and other individuals with relevant expertise would be developed. in addition, a central location would be designated to maintain all published lists of archival holdings and study descriptions. these documents would be produced as reference documents for users of archives or institutions interested in consulting with individuals in the field of archiving. activities and plans with respect to the first part of the mandate, the action group will focus initially on developing the directory of data archives and data libraries. iai)c3i5t newoietter, vol. d. no. i (summer lyvb) tkchiucal stanuahus fok magnetic tape exchange between data organizations karsten boye hasmussen danish data archives intkoduction tape labelling (nl, no label ) at the lassist-sessions in uppsin principle, aii products t rom aia august ly 10, lyvo, there was tne major computer lirms should be some discussion of the technical capable of reading ansi labels aspects of data exchange. in (american national standard recent years, the danish data labels). however, dilferent sysarchives (uua) has had considerable terns may produce slightly dillerent experience concerning exchange of ansi-labels (due to frequent change data files on magnetic tape -ol operating systems) , and some internationally between data organsystems may not process ansi labels izations as well as inside denmark correctly. bor these reasons, we between computer installations would propose that all exchange (notably ibm, cdc and univac). the tapes be written without labels da therefore accepted the invita(nl, no label), as nl-tapes can tion to write a note on these probdelinitely be processed by any comlems proposing some usable stanputer center. leading tapemarks dards. although the dda as a do should be avoided. (data orgnization) emphasizes the importance ol documentation, this note deals with the technical aspects of data transfer on magthacks ( 9-track , ibou bpi , pe ) netic tape in general . at most installations, 9-track the correct procedure for tape drives are availaole; normally exchange of magnetic tapes naturthis is true even tor computing ally includes preparation ol a comcenters where the system uses plete technical description of the /-track tape drives. at present, tapes. however, the standards prothe most commonly used density is posed in this paper, while they may ibuu bpi (bits per inch) with phase not always be the most effective encoded (pe) magnetization, ones or the easiest ones to use, should assure the possibility ot reading the tape at other locations, even in cases where tne tape data kohmat ( blocked , fixed record description for some reason is length , tfu char . , ascii ) lacking. the philosophy behind the standards is that of "simplicamost system and/or machine tion". you should never let a dependent files should be aban"data maniac" convince you to ship doned , as even a very good a multi-reel, machine and operating system dependent spss-file on seven track tape; that is, unless you are on really bad terms with the receiver ol the file. yb ia^sist newsletter, vol. ^, no. i (summer ly/b) programmer would have to spend months converting them to the local tile tormat. some installations support conversion programs, but these programs may not exist at the receiving computer center, and even if they do, the program 'level' may de different. it is advisable to snip data files in the most conservative and simple format possible, i.e. card image character format . given the tact that many computers will have trouble handling large blocks, the blocksize should not exceed 204a bytes. finally, we would recommend using the ascii character set, which can normally be converted by standard software (or the operating system). lunllusiun for file 1 the fo istics labels the f record (ascii exceed lation plied ported readin (in is urgani pleted nation zation case , sent capabl exchange n magnet llowing : the ta (nl), yile sho s , be 1 ; and h ing 204a s will h by the m by the c g and con fact is r zation h by the m al federa s) . sho tiles ot only a m e program purposes ic tape s technical pe should track, 160 uld have n charact ave a bio bytes. mo ave softw anufac tur e omputing c verting s ellected i egistry fo embers of tion of da uld this this kind inor prob mer . , a data hould have characterbe without bpi, pe. au byte er tormat cksize not st instalare (supr or supenter) for uch tiles, n the data rm , comthe interta urganinot be the will prelem to a the file produced by wkite ultint-uj already tollow the standards outlined above as tar as data format is concerned; and if, in the future, a standard for documentation files is defined -e.g. the "data interchange tile" as proposed by hichard koistacher -the standard file will almost certainly be in character card image format too. hhb fktncts cue. ums-iyu vehsion 1 no. bu^yt) icfsk. usih arbor : michigan , ibm os utilit tu. (gc nie, norman, fuitiun. hill, lyy hoistacher , intfhchan 2b>y. un febrauary , cl^bfr kecurd manager reference manual , yoo, i^yb. is 111, vul. 1. ann the university of 19yb. its (itbgkner), ibth 2b-bbbb-lt5; . april, hi al. spss stcunu new )(ork: mcgrawrichard c. the data gt h ilt. cac doc. no . iversity of illinois, , lyya. system supported software for gen erating and reading " standard " files : ibm: us utilities (lebgener) order no. gc^'a-bbaoi 5 , ibth ed . april lyyj cdc: dms:1y0, cyber record manager version!, reference manual, no. bu4yt)y0u, lyyb as mentioned above, this note does not distinguish between data and documentation files, as this distinction is irrelevant for purposes of generating and reading the tape. indeed, documentation files trom the most commonly used packages (osiris: the codebook; spss: yy iassist quarterly 3 educating the data user introduction finding information today need not mean going to a librar,-. one can put a data diskette into a microcomputer, dial a remote mainframe through a modem, or turn on a cd-rom drive from a desk at home, from the office, or from a college lab. and the information sources accessed can be as varied and as complex as the collection of a college library-. but has the ease of access to information sources increased "information literacy", the ability to define a search strategy, identify good access sources, retrieve appropriate materials, and evaluate sources located. probably not information skills in use today are too often identical to those used when the gateway to information was the librar)-'s card catalog and the reader's guide! new information access skills must be developed. retrieval skills which can fully exploit the potential in contemporary information access by combining traditional print sources with electronically generated media in an innovative synthesis. to meet this challenge, a growing number of american colleges and universities have added modijes on electronic information access to their bibliographic instruction programs. many articles in the literature describe their successes and failures. but an examination of this literature indicates that no specific attention is given to pubhc data in machine-readable format despite the growing importance of datafiles in business, public administration and academia research, public data is still discussed primarily in its print form in bibliographic instruction. obviously this doesn't mean college students don't learn about data. quite the contrary. college courses in business and the social sciences are ver>' quantitative in keeping with the reality of the way business and research is conducted today. but the objective of most quantitative courses in these disciplines is the development of technical and analytic skills, not the acquisition of information proficiency. often instructor-created, the datasets used in cours^work offer analytic problems but bear no relation to actual public data sources. data sources which these same students will surely use on a regular basis in their careers in business, in public agencies or in academia. the failure to teach them to design search strategies which include numeric datafiles. to evaluate the usefulness of these files alongside the same information in other formats, and to acquire these files in the most efficient manner possible, leaves a large gap in their training. identifying and locating good numeric data sources and choosing among the storage formats available are important information skills. at baruch college, city university of new york, the objective of the library instruction division is to educate faculty and students, graduate and imdergraduate. to the vast possibilities of the contemporary information summer 1988 4 iassist quarterly environment we have included all the varied sources the library presently collects, including machine-readable datafiles. we treat datafiles as an as an information resource, leaving to others the analytic training of data users. the ptirpose of this panel is describe the efforts that have been made at the college to incorporate numeric datafiles into a very varied bibliographic instruction program and to underline the importance of bibliographic instruction in the training of data users. the panel consists of members of the library instruction division each of whom, besides their instructional responsibilities, has responsibility for another data-related library function. each of the panelists will discuss their individual roles and the work they have done in developing the methods being used in this program. but first some information about baruch college. baruch college is a four-year college, predominantly underpaduate. it is part of one of the largest public university systems in the united states: the city university of new york. il consists of three schools. liberal arts, education and educational services and business although its primary strength, and the majority of the student body, is in the business fields. most of the graduate programs, which include a ph.d. are in business fields. to meet the needs of its students and facult>-, the library began an instructional program more than 15 years ago. its goal was and is the improvement cf the level of "information literacy" among baruch graduates and faculty. the baruch programs emphasize the responsibility of the researcher or information consumer to develop appropriate strategies for finding information, to become knowledgeable about major sources in a field, and to choose access methods most appropriate to the problem at hand. extremely sophisticated and complex, the course offerings have responded to changes in the information environment students are expected to work in. the program offers levels of education and training suitable for students and faculty with differing needs and differing backgrounds. the ofterings include library orientation exercises library research workshops — one or two lecttire modules given within another course to meet specific information objectives within that course bibliographic instruction courses — 3 credit courses designed to teach the conceptual aspects of information access as well as the specific skills of information retrieval computerized information services training seminars — programs designed to teach online information access; data resources seminars — seminars devoted to machine-readable numeric datafiles specialized study center to provide resources and training for graduate business students. the most recent addition to this multi-faceted program is a curriculum, currently under development, for an information studies major and minor intended for those students in liberal arts, business or education who are interested in pursuing information, not necessarily library, careers. the papers presented as part of this panel will detail how numeric datafiles, public data sources, have been incorporated into the varying parts of this program.n summer 1988 lassist newsletter vol. 4 nos.3&4 3) at present the canadian historical association and the canadian pohtical science association represent the research community on the advisory council on public records. the public archives is giving serious consideration to widening the diversity of representation of the research community 4) the full text of part i v of the canadian human rights act can be found as an appendix to the 1980 index of federal information banks which is available for reference in every post office and canada manpower office across canada. 5) kissinger vs reporters committee for freedom of the press et al.. supreme court of the united states. march 3. 1980. no 78-1088 and 78-1217. the case involved summary or verbatim transcripts of kissinger's conversation notes and telephone notes which he "unlawfully removed" from the state department and al a later date deposited al the library of congress with severe access restriction. in short, the supreme court ruled that the plaintiffs had "no standing" and that only the national archives and records service and/or the state department could pursue kissinger for return of the notes. 6) senate report to the federal records act of 1950. s rep no. 2140. 81st congress. 2nd session, at 4jou£nal. oi and perceptions of the st. social. history (winter, 1968), louis police, 1899-1970," in 156-163; "crime and the indusjohn conley, ed.. current trial revolution: british and l££ilds id criminai. ^lyiiice: american views," j ourn al of itie2£x and riliilik* (ctncinsgclai. hi£torx (spring, 1974), natlt anderson, ~1979) . 2b7-303. 16. eric h. monkkonen, "toward a 22. colin loftin and robert h. dynamic theory of crime and hill, "regional subculture and the police: a criminal justice homicide: an examination of systems perspective,' his.torithe gas t i i -hac kney thesis," £.al met hods neji^letter (fall, ijieri£.an so£i2i2alcal review 1977)7 157-165; "systematic (oct., 1974), 714-724. criminal justice history: some suggestions," journal, of int^r23. for instance, douglas greendiiciaiinarii his,to£2 "^winteti burg, c£im£ and law £nf2ci£19 797, ~4 51-4 64 ." ~ hent iq the col.oni of new york, 1^51-1775 (ithaca, ny: cornel michael d. maltz, "crime 21 lassist newsletter, vol. 3» no. t (fall 1979) univ., 1976), and harvey j. criminal prosecutions of slaves graff, "crime and punishment 1n in antebellum south caroline," the nineteenth century: a new iiournal of. *21£.i.il2.q !lislo£x look at the criminal," journal toec, 1976), 575-599. see of interdisciplinary tliilor^ also, hindus, "the contours of (winter, 1977), 477-491, both crine and justice in massachuexempllfy the dangers of ad hoc setts and south carolina, theorizing. 1767-1878," the al!l££i£.3n jourdsk £l k£.a£i. tlis.t£lz (july, 24. harvey j. graff, "pauperism, 1977 ) 7~212-237. misery, and vice: illiteracy and criminality in the nine25. monkkonen, the d^naerous class: teenth century," journal of £.li!l£ §.u.i e.2i£i.li iq co i umbus , soclai. h1.ii££x (winter, 197777 ohtot 1860-188^ tcambridge: 245-268. michael s. hindus, harvard univ., 1975). "black justice under white law: iassist quarterly 2010 / 2011 71 iassist quarterly some social scientists are skeptical about sharing their own data abstract this paper gives a brief overview on practices of archiving and re-using social scientific data in poland. we start from drafting our diagnosis of the field, then move to some concrete examples of relevant initiatives. most of them are of non academic character and are focused on oral history documentation and research project (which we count among qualitative research); some examples of longitudinal studies are also mentioned. finally we are trying to formulate key problems impeding development of the field – without loosing conviction, they will be overcome. keywords: data sharing, archiving, qualitative data, qualitative longitudinal research, oral history, poland. to what extent is data sharing or archiving for re-use part of the research culture in poland? based on the example of probably all social sciences it might be said that data sharing is not a part of research culture at all. there has been no academic tradition of such kind in social sciences in poland. although the problem is being brought up at least occasionally no significant change has been noticed within the rather conservative sociological main stream. furthermore, whereas some social scientists are skeptical about sharing their own data, others, e.g. anthropologist, are reluctant even to use qualitative data not gathered by themselves (their motto: the researcher has to soak in the atmosphere of the field”, “only he/she knows what is the real meaning of the collected materials, because he/she was there, usually for a long period, he/she experienced the situation and as a consequence only he/she can properly interpret the data”). therefore, it is more likely that using someone else’s qualitative data in a project (obviously with a permission of the author) might cause deep methodological concerns of more experienced colleagues rather than arouse academic interest or approval. as a result, on the level of everyday practice most of the data collected in the field has a “disposable” character. once the publication is ready they find their place in the researcher’s desk drawer. exceptions are being made for close co-workers, assistants, selected phd and ma students, mostly in the situations when they join particular research teams and/or continue the work of their predecessors (very often their supervisors). apart from what paul thompson called “sitting on data” there is also an important question of reliability of sources. for instance most of the contemporary polish historians treat sources like recorded interviews as untrustworthy. the dominant approach in historiography has been deeply positivistic and, as such, too strict to develop research based on oral history interviewing. this group of researchers had to deal also with a problem of political character. oral history has been often understood as some sort of data sharing and archiving qualitative and ql data in poland by piotr binder and piotr filipkowski1 72 iassist quarterly 2010 / 2011 iassist quarterly “history from below” or as “giving voice to the voiceless”, which before 1989 was ideologically uncomfortable and treated as suspicious in the official historiography in poland. what is interesting is that institutions and projects that are oriented towards archiving qualitative data do exist. this report provides many examples of such initiatives. although some of them are relatively advanced, they do not compose a network, but constitute rather a group of dispersed and diversified projects facing similar problems and very often competing with each other especially in the area of funding. however what must be stressed – and this is crucial for this document – is that archiving institutions function practically outside academia. contacts between both sides are loose and based on individual rather than inter-institutional cooperation. the existence of one only – and until now not a very successful one – official archive of qualitative data affiliated at the academic institution in poland proves that changing the existing patterns and introducing new ones is not an easy task to do. nevertheless, the authors would like to believe that with efficient cooperation and exchange of information as well as with effective promotion of good practices some “qualitative change” will be possible even within very conservative polish academic circles. existing qualitative archiving infrastructure: after what has been stated in the introduction one can expect not to read much about existing infrastructure for archiving qualitative data of an academic character in poland. to a vast extent those intuitions are true, because at the present time there is only one enterprise of such kind. qualitative data archive [archiwum danych jakościowych adj] 2 adj is an initiative of a group of researchers from the institute of philosophy and sociology of the polish academy of sciences (ifis pan) to establish an archive of qualitative data. unfortunately, after almost five years of experiences it is still at its very initial stage. problems with financing as well as with the common practice of not sharing gathered data have limited development and expansion. the milieu of researchers involved in this initiative grows slowly, however, working pro-bono they keep collecting qualitative data. apart from those coming from their personal research project (e.g. over 200 interviews with young people gathered within a project generation 1989) they managed to save the entire collection of fieldwork materials from 1970s and 1980s (from tape-recorded interviews and fieldwork notes to photographs and family budgets) of the team of professor andrzej siciński (which is discussed below). in the last months there has been more intensive discussion at the polish academy of science on the need to develop and expand qualitative archiving. we would like to start with data collections produced in different project realised within the academy, but also start collecting ‘external’ data. the key obstacle is lack of additional funds for this initiative. existing qualitative archiving infrastructure of nonacademic character oral history archive of the karta centre and history meeting house3 in warsaw this was the first polish initiative that started to record, collect and archive qualitative data in a systematic way on a wider scale. until now it is also the biggest and the most diversified collection of qualitative data in poland. karta started in the early 1980’s as a milieu of young dissidents who decided to oppose the official system through gathering, preserving and publicizing documentation (including: memoires, diaries, interviews, pictures, documents etc.) regarding individuals and groups “forgotten” or rather neglected in the political mainstream. in the 1980’s karta managed to establish the so called eastern archive, where approximately 1200 interviews with poles repressed by the soviet state were stored. additionally, karta’s activists were conducting interviews with polish dissidents, people who, in various ways, actively opposed the communist system before 1989. these two experiences were of a decisive character. in 2002 a decision was made to invite karta to participate in the mauthausen survivors documentation project. it was the biggest european oral history project devoted to a single nazi concentration camp. over 860 interviews were conducted all over the europe in the united states and israel. karta recorded and elaborated over 160 biographical narrative interviews with polish mauthausen survivors. the strong position of karta was highlighted by the fact that, unlike in other participating countries, the polish part of the project was coordinated by a nonacademic institution. the experience of the mauthausen project forced karta to establish a permanent oral history programme and create a modern oral history archive. the process that started at that time has had a long and difficult history. however thanks to generous support of institutions like the european commission, the polish senate (the upper chamber of the parliament), german foundation “remembrance, responsibility and future” and the polish committee of scientific research, within the past few years karta has completed several significant oral history projects. nevertheless it needs to be stressed that the financial side of creating an entirely new archiving infrastructure without stable institutional or state support is particularly challenging. the archive could not function and develop so intensively without the support of the history meeting house (hmh), a municipal institution of culture established in 2006 on karta’s initiative. the archive itself belongs to both institutions (physically it is located at the history meeting house). thanks to this co-operation the oral history collection could be digitalised (including the interviews collected in 1980’s), catalogued and partly published on the website: www.audiohistoria.pl (full access available on the premises of hmh). each interview is accompanied with a questionnaire, biographical data of the interviewee and short description of the narrative together with interviewer’s remarks regarding the interview situation. unfortunately, due to the lack of sufficient sources only selected stories could be transcribed. museum of warsaw uprising4 shortly after its opening in 2004 a separate oral history unit was established to conduct, collect and make accessible interviews with polish soldiers of the warsaw uprising. up to now approximately 2000 video interviews have been recorded. edited transcripts are available online, with full recordings at the premises of the museum. in 2008 the museum published a catalogue of its oral history collection. the museum of warsaw uprising is a state funded institution with a stable budget. its oral history archive is additionally supported on a regular basis by one of the biggest polish banks. brama grodzka – teatr nn in lublin5 iassist quarterly 2010 / 2011 73 iassist quarterly the institution itself has existed since 1990, and its oral history programme since 1998. brama grodzka is active in the lublin area (in the south-east of poland) the region that until wwii was particularly multi-ethnic and multicultural. although its pre-war past constitutes the main focus of interest, the institution is also running other related projects e.g. to wwii or democratic opposition in poland. brama grodzka has collected and archived approximately 800 audio (and some video) interviews. partial access to this collection is possible over the internet, full at the premises of the institution. brama grodzka cooperates with public radio lublin, which created the oral history studio. as a municipal cultural institution it is maintained by the city of lublin. these are the biggest and most advanced (also in archiving and reusing) oral history initiatives in poland. since several years, there are many smaller projects operating at the moment. some of them are listed at the end (see appendix ii). these various activities in the field of oral history are not only dispersed but there is also very little discussion and cooperation between them. however the need to integrate the milieu has already been perceived. in autumn 2007 the first international conference on oral history in poland was organized in kraków (oral history: the art of dialogue). this event was followed by the setting up of the polish oral history association in january 2009. the platform for discussions, exchange has appeared. one of the goals is to create a database covering all polish oral history initiatives and introduce common standards of description and archiving of collected material. qualitative longitudinal (qll) data (i.e., data of diverse formats (interviews, videos etc.) addressing time and temporality gathered in follow-up and repeat cross sectional studies, and retrospective studies such as life and oral histories) due to the fact that there is no access to data collections of most of academic research projects we have decided to provide some examples – although not many –of: longitudinal studies or projects involving follow-up studies, repeated studies or retrospective studies. additionally, where possible, contact details of the heads of the projects or people who possess collected data sets are provided. 1) repeated monographs of local communities: monographs of local communities along with the analyses of personal documents and competitions of memoirs once gained the title of being the “calling card” of polish empirical sociology (to mention only “the polish peasant in europe and america” edited by f. znaniecki and w.i. thomas, 1918-1920). chronologically the first monographs in poland were written by ethnographers. the greatest achievement of the 19th century was the people, an 86-volume work by oskar kolberg. the beginning of the 20th century was a time of ‘economic’ monographs e.g. by franciszek bujak (1901, 1903). those monographs that are regarded in polish sociology as “sociological” appeared only after that and were usually focused on a particular problem – so called problem monographs e.g. the polish-german antagonism in the factory settlement “kopalnia” in upper silesia by józef chałasiński (1935). although neglected in the 1970s, they have been recently gaining interest, together with the growing popularity of the ideas of localism and local problems in social sciences in poland after 1989. examples: bujak, franciszek. 1901. maszkienice village of brzeg county. economic and social relations. bujak, franciszek. 1911. maszkienice village of brzeg county. from 1901 to 1911. zawistowicz-adamska, kazimiera. 1948. rural community: experiences and delibeations. (based on fieldwork conducted in zaborów village in 1938) wieruszewska, maria. 1978. transformation of local community. zaborów after 35 years. bujak, zbigniew. 1903. żmiąca -village of the limanowa county. economic and social relations. wierzbicki, zbigniew. 1963. żmiąca half a century later. łuczewski, michał. 2007. national experience in everyday life. problem monograph of żmiąca village: 1370-20076. 2) research on life styles7 in the 1970’s and 1980’s the team of professor andrzej siciński from the institute of philosophy and sociology of the polish academy of science was running a project on life styles of individuals and families in polish cities. throughout this time each team member was cooperating with several families. a wide variety of qualitative methods were applied: from participant observation, in-depth interviews and documents to photographs and home budgets. the main outcome of the project was six volumes edited by siciński between 1976 and 1988. 3) research on poverty8 since the mid-1990’s three sociologists, namely elżbieta tarkowska (polish academy of sciences), kazimiera wódz (university of silesia) and wielisława warzywoda-kruszyńska (university of lódź) have jointly conducted research on various aspects of poverty in poland. part of the interest is devoted to the inhabitants of the former collective state farms. after conducting a wave of interviews with several members of each selected family in 2003, elżbieta tarkowska came back to her respondents in 2007. in effect two edited volumes were published in 2004 and 2008. additionally, one of the books of the team was composed mostly of the selected interviews preceded with a short commentary (tarkowska, warzywoda-kruszyńska, wódz, poor people about themselves and their lives, katowice-warszawa 2003). examples of quantitative longitudinal research conducted in poland 1) research on primary and secondary school pupils9 longitudinal research was conducted between mid-1970’s and mid1990’s on whole cohorts of school children in the toruń and włocławek regions by the team of sociologists from university of nicolaus copernicus in torun headed by professors zbigniew kwieciński and ryszard borowicz. until the mid-1990’s the dominant technique was audience questionnaire. after the introduction in poland of the personal data protection act in 1997, the research team decided to change the character of the project and base it on representative samples. 2) social structure in poland. dynamic analysis in the international context10 a research team of sociologists from the institute of philosophy and sociology of the polish academy of sciences (kazimierz słomczyński 74 iassist quarterly 2010 / 2011 iassist quarterly head, henryk domański, bogdan mach, krystyna janicka et al.) between 1988 and 2008, conducted 5 waves of their panel research project. this group of scientists has managed to continue their work on the project in spite of difficulties related to the new law regulations (personal data protection act from 1997). they publish regularly the outcomes of their research in the form of books and articles (both in polish and english). development planning: although thank to various initiatives a significant amount of qualitative data have been collected over recent years, there are serious gaps within the existing framework: a) data sharing or archiving for re-use is not a part of the research culture within the field of social sciences in poland; b) social scientists are not only reluctant to share the data collected by themselves, but also to use the available data collected by other researchers. as an effect only a very little involvement of the academic researchers in the initiatives devoted to collecting qualitative data can be observed; c) there is no balance between collecting data and analysing it. what is more, some of the institutions mentioned in this report focus entirely on collecting data. the task they fulfil is very important; however it would be very good if the qualitative material gathered with so much investment of time, effort and financial sources aroused interest of the analysts as well. d) there is still very little communication between institutions collecting qualitative data on various levels. this lack of exchange of sometimes basic information causes situations when several projects overlap with each other or even repeat the same type of interviews with the same people over a short period of time. e) institutions collecting qualitative data constantly suffer shortages of financial sources to the extent that very often they do not have enough money for the creation of professional catalogues and computer databases, and furthermore, not enough for good quality recordings that could be used later as audio/video materials (and not only the basis for transcriptions). f ) the personal data protection act from 1997 constitutes a separate question. despite the fears of many researchers, its introduction was not the end of social research in poland. institutions and research teams have to follow certain regulations, however conducting research (including research of a longitudinal character) and collecting data is still very much possible. therefore it is not entirely clear why the group of sociologists from torun claims that they were “forced” to reshape their longitudinal project whereas the team of professor słomczyński from the polish academy of sciences continues its work (discussed above). in terms of priorities for development over the next three years, the main priority is probably a gradual change in the dominant research patterns and practices within the area of social sciences. this naturally will be a long-term process, nevertheless certain actions could be taken. a strategy of building closer relations with academia based on already existing infrastructure (although its condition is not satisfactory) seems to be a natural direction. financial support is always very much appreciated, however existing organizations (iassist, cessda) could provide something equally important, namely intellectual support. polish institutions dealing with archiving of the qualitative data are still very little advanced, therefore assistance and advice of more experienced partners might help avoid many mistakes and thereby have a significant influence on future development. references bujak, f. (1903). zmiaca wies powiatu limanowskeigo. montanta: kessinger chałasiński, j. (1935). the polish-german antagonism in the factory settlement “kopalnia” in upper silesia. publisher unknown. kolberg, o. (n.d.). the people. (86 volumes). publisher unknown. tarkowska, e, warzywoda-kruszyńska, w and wódz, k. (2003). biedni o sobie i swoim życiu. katowice: śląsk. znaniecki, f and thomas, w. (ed). zarestsky, e. (1996). the polish peasant in europe and america. illinois: university of illinois press appendix i national policies on data archiving and sharing 1) personal data protection act from 1997 [ustawa z dnia 29 sierpnia o ochronie danych osobowych ]. full text in polish available on the website of the polish parliament: http://isip.sejm.gov.pl/servlet/search?todo=open&id= wdu19971330883 the most important legal act concerning protection of personal data in poland is the act of 29 august 1997 on the protection of personal data. the act establishes the inspector general and determines framework of the personal data processing. the inspector general is competent in the issues concerning personal data protection and may inspect any subject who processes personal data. entities which belong to the eea (european economic area) are obliged to comply with the act on the protection of personal data only if they operate in the territory of poland. this means that almost every company registered in polish national court register have to comply with the act. the main requirements for the data controller are: to process personal data on the basis of legal prerequisites to inform data subjects about their rights and the data controller status to register the data files in the inspector general office to secure the personal data from uncontrolled access to remove the personal data in case of a request from the data subject personal data usually cannot be transferred outside of eea (european economic area) without prior approval of the data subject. however, it can be transferred without restrictions in the european union and few other countries (e.g. norway, island, lichtenstein). not applying the provisions of the polish act on personal data protection can lead to criminal responsibility, in some cases even to three years of imprisonment. more often it leads to administrative proceeding. source: law firm czuchaj and partners www.czuchaj.pl iassist quarterly 2010 / 2011 75 iassist quarterly 2) the civil code (within the scope of research among children). full and unified text in polish available in on the website of the polish parliament: http://isip.sejm.gov.pl/servlet/search?todo=open&id= wdu19640160093 3) icc/esomar international code on market and social research http://www.esomar.org/uploads/pdf/professional-standards/ iccesomar_code_english_.pdf 4) qualitative research consultants association (code of ethics and guide to professional qualitative research practices) http://www.qrca.org/displaycommon.cfm?an=1&subarticlenbr=21 5) control program of the quality of work of pollsters (polish association of public opinion and marketing research firms) http://www.ofbor.pl/index.php?i=pkjpa&id=1&pg=1 appendix ii other selected oral history initiatives that protect collected sources/ data museum for the history of polish jews (dozens of interviews with poles who rescued jews during wwii – project “righteous among the nations”) contact details: ul. warecka 4/6, 00-040 warszawa, web: www.jewishmuseum.org.pl tel: +48 22 833 00 21, email: łucja koch lkoch@jewishmuseum.org.pl center for citizenship education (45 interviews with poles who rescued jews during wwii) contact details: ul. noakowskiego 10, 00-666 warszawa, web: www.ceo. org.pl tel: +48 22 6220089, marianna hajdukiewicz (coordinator) marianna@ ceo.org.pl christian association of auschwitz families – project memento (45 interviews with auschwitz survivors, 200 hours, 1000 pages of transcriptions). contact details: ul. partyzantów 1, 32-600 oświęcim, web: www.auschwitzmemento.pl tel: +48 508-099-030, +48 503 078 357, biuro@auschwitzmemento.pl lower silesian forum of cultural background „milenium” (130 interviews with people relocated to lower silesia after 1945) contact details: ul. kowalska 58/28, 51-424 wrocław, tel: +48 888 315 334, juliusz woźny juliuszw@wp.pl museum of warsaw praga district (several interviews with the oldest inhabitants of the district) contact details: ul. targowa 45, 03-728 warszawa, tel:+48 818 10 77, +48 695 645 501, muzeum.pragi@mhw.pl european solidarity center („solidarity movement in my memory” 20 interviews) contact details: wały jagiellońskie 1, 80-853 gdańsk, tel: +48 58 3237056, monika bogdanowicz m.bogdanowicz@gdansk.gda.pl pedagogical academy in cracow, institute of history (150 interviews with people expelled from polish eastern borders (kresy) after 1945) contact details: ul. podchorążych 2, 30-084 kraków, tel: +48 604 135 930, hubert chudzio, phd hubert@ap.krakow.pl center „remembrance and future” in wrocław (interviews with poles relocated to polish western borderlands after 1945) contact details: al. gen. j. hallera 8, 53 – 318 wrocław; tel: +48 71 33490 44, +48 663 901 767, biuro@pamieciprzyszlosc.pl fundacja kobieca efka (feminist organization conducting interviews with women in different gender-focused projects) contact details: ul. krakowska 19, 31-062 kraków, tel:+ 48 12 430 19 70, efka@efka.org.pl and – last but not least – there are memorial sites, especially of former concentration camps, which collect, archive and make accessible oral history interviews with survivors. most of these interviews are thematic ones – they focus on the camp experience. most of these, however, treat texts of transcripts as the primary source and neglect the recordings. auschwitz memorial (3000 written accounts of former inmates, plus over a 100 audio and video interviews) contact details: ul. więźniów oświęcimia 20, 32 603 oświęcim, tel: +48 33 8431934 majdanek memorial (interviews with former prisoners) contact details: droga męczenników majdanka 67, 20-325 lublin, tel:+48 81 744 26 47, sekretariat@majdanek.pl stutthof (550 accounts of former prisoners – text, audio and video) contact details: ul. muzealna 6, 82-110 sztutowo, tel: +48552478353, stutthof@stutthof.pl gross-rosen museum (206 video and 97 audio interviews with former prisoners) contact details: ul. szarych szeregów 9, 58-304 wałbrzych, muzeum@ gross-rosen.pl treblinka museum (47 interviews) contact details: 08-330 kosów lacki, tel: +48-25-781-16-58, biuro@ muzeum-treblinka.pl bełżec museum (80 interviews – all with transcriptions) contact details: ofiar obozu 4, 22-670 bełżec, tel:+48 846652510, muzeum@belzec.org.pl notes 1. contact details of country reporters: • piotr binder pbinder@ifispan.waw.pl , tel: +48 505289998 institute of philosophy and sociology, polish academy of sciences nowy swiat 72, 00-330 warsaw, poland 76 iassist quarterly 2010 / 2011 iassist quarterly • piotr filipkowski p.filipkowski@karta.org.pl, tel: +48 694699470 institute of philosophy and sociology, polish academy of sciences nowy swiat 72, 00-330 warsaw, poland and karta center, narbutta str. 29, 02-536 warsaw, poland 2. contact details: qualitative data archive [archiwum danych jakościowych] http://www.ifispan.waw.pl/ archiwum_danych_jakosciowych/o_archiwum// nowy świat 72, 00-330 warsaw, tel: +48 22 657 2852 professor hanna palska hpalska@poczta.onet.pl artur kościański, phd akoscian@ifispan.waw.pl 3. contact details: karta, ul. narbutta 29; 02-536 warszawa, tel. +48 22 848 07 12, web: www.karta.org.pl; email: p.filipkowski@karta.org.pl history meeting house, ul. karowa 20, 00-324 warszawa, tel: +48 22 826 25 78, www.dsh.waw.pl ; www.audiohistoria.pl ; email: ahm@dsh. waw.pl 4. contact details: ul. grzybowska 79, 00-844 warszawa; tel: +48 22 539 79 38; www.1944.pl; email: kontakt@1944.pl 5. contact details: ul. grodzka 21; 20-112 lublin; tel. +48 81 532 58 67; www.tnn.pl; email: teatrnn@tnn.lublin.pl 6. contact details: michał łuczewski phd, warsaw university, institute of sociology, 18 karowa st., 00-927 warsaw, luczewskim@is.uw.edu.pl 7. contact details: qualitative data archive [archiwum danych jakościowych] http://www.ifispan.waw.pl/ archiwum_danych_jakosciowych/o_archiwum// nowy świat 72, 00-330 warsaw, tel: +48 657 2852 professor hanna palska hpalska@poczta.onet.pl artur kościański, phd akoscian@ifispan.waw.pl 8. contact details: professor elżbieta tarkowska, institute of philosophy and sociology, polish academy of sciences, department of theory of culture, research group of poverty studies nowy świat 72, 00-330 warsaw, etarkows@ifispan.waw.pl 9. contact details: monika kwiecińska-zdrenka, phd, university of nikolas copernicus, institute of sociology, fosa staromiejska 1a, 87-100 toruń, monika.kwiecinska@umk.pl 10. contact details: professor kazimierz m. słomczyński, institute of philosophy and sociology, polish academy of sciences, research group of comparative analysis of social inequalities, nowy świat 72, 00-330 warsaw, kms0543@aol.com by 16 iassist quarterly spring summer 2009 implementing ddi 3: the german microcensus case study abstract this paper shares experiences in developing an application for documenting the german microcensus at the variable level. first, we developed an editor in compliance with the ddi 3 standard to improve and simplify the process of documentation. second, we developed a web information system in order to provide the end user with various views on the metadata. the scope of the work depicts the development cycle of applications based on ddi 3. introduction the german microcensus is a representative annual population sample containing structural population data for 1 percent of all households in germany (bohr et al., 2006). the german microdata lab (gml), the service centre for microdata of gesis leibniz institute for the social sciences, created a project to build an information system to provide information on the german microcensus2 for public needs. this project, called missy or “mikrodateninformationssystem”, was begun in july 2003 (bohr, 2007). as a pilot project, missy succeeded in documenting the german microcensus for the years 1995 and 1997 based on ddi 2.1 and also in presenting it on the web (janssen et al., 2006). following those accomplishments, the project expanded in 2008 to become a collaboration between gml and ips3 this ongoing project has the long-term vision of documenting the data life cycle4, specifically for the german microcensus at the variable level. further, the project focuses on accessibility and reuse of the metadata for other purposes in the future. since ddi 2.1 is targeted at documenting unrepeated surveys, we decided for the expanded project, called missy ii, to use ddi 3 and to draw on its support for maintaining historical versions of surveys. our current task deals with the metadata of the census years between 1973 and 20075 . because ddi is an xml standard, we considered creating the documentation (a) with a common xml editor and (b) with an editor customized for ddi. the first option had several disadvantages. first, it demands skill and practice in xml, which are unlikely for common users. second, it takes too much time as each documentation file contains thousands of lines. in contrast, the main advantage of using a customized editor is that it simplifies and accelerates the process of documentation so neither skill in xml nor knowledge of the ddi standard is required. in addition, we have to ensure that the metadata are being well rendered and displayed for public access. therefore, we provide not only the metadata in ddi format, but also an easy way to view the metadata on the web (as a continuation of the previous project). ddi 3 editor qdds foundation since there are only a few tools for ddi, we decided to develop a ddi 3 editor for our own needs. the missy editor is based on the same architecture as the questionnaire editor software called qdds, which uses ddi as the main storage format (hopt, stempfhuber, et al., 2009). qdds is a collaborative project run by the university of duisburgessen and gesis6. qdds itself is a proven editor based on the ddi standard, has already been used to document many surveys, and is still being enhanced continuously (hopt, amin, et al., 2010). the qdds questionnaire editor was built using the java™ programming language and classes for the document object model (dom). the architecture is a user interface implemented in java swing which is connected to a questionnaire manager. this class allows access to questionnaires loaded by providing manipulators. the manipulator classes all implement a defined interface for loading ddi nodes and for reading and setting named fields. they work directly on the xml structure of ddi and are instantiated by name. the user interface just has to know which sort of manipulator it needs for a special task and then ask the manager for it (e.g., “question”). the manager then knows about the metadata format, ddi 2.1, and creates the requested manipulator. as a result of this architecture, new versions of ddi or even new metadata formats can be supported by implementing a new set of manipulators and changing the format information in the manager class. this also includes validating the xml against, for example, ddi 2.1 or ddi 3. figure 1 shows the by andias wira-alam and oliver hopt 1 iassist quarterly spring summer 2009 17 figure 2. ddi 3 editor at the variable level figure 1. data manipulation within the editor data manipulation mechanism within the editor. missy ii system as an overview, missy uses the following main elements from ddi 3: datacollection questionscheme questionitem logicalproduct categoryscheme variablescheme variable physicalinstance statistics variablestatistics 18 iassist quarterly spring summer 2009 the editor uses the possibilities of webdav7 for authentication purposes. additionally, webdav allows editing files on a remote server. one census year is stored as a single ddi 3 file which is located in a shared folder. when a particular census year is edited by a user, it can be used by others in read-only mode. however, other users are able to edit the other census years that are idle. another advantage of using webdav is that the editing process of the metadata becomes location-independent. as shown in figure 2, the editor displays a list of all variables from the selected survey and period. the variable selected from this list is then displayed in an edit form, in which the user can easily enter and edit the metadata. this form again is arranged with two tabs. the first tab contains all content describing the variable in general. the second tab contains a single table to display all answer values, the corresponding labels, and their frequencies in absolute numbers and percent (overall and valid). the only column that is editable in this table is the value labels. all other metadata (like variable statistics and variable names) are imported from files generated from the raw data. metadatabase the ddi 3 editor produces plain text files (xml files), each of which represents a census year and contains thousands of lines. in a plain text file, the metadata of a census year is considered a single record, although it is well structured hierarchically using ddi 3 xml. suppose we want to determine whether a certain variable from a given census year also appears in other census years. this can be handled by searching through all census years (except one to be precise) to match the equivalent variable. of course, this leads to a long computation and causes a performance problem. the problem, however, can be reduced by using an indexing technique. suppose that all variables are being indexed and each variable is mapped to its corresponding census year. through the index, the precise location of each variable is clearly described, e.g., by line number or xml node. using this indexing strategy, we transformed our ddi files into a metadatabase, related to the solution described in jensen et al., 2010. we use dbclear, which is a generic, platform-independent clearinghouse system, whose metadata schema can be adapted to different standards (hellweg et al., 2002). dbclear has an open architecture and reusable components that make it easy to customize and to enhance depending upon the requirements in compliance with the mvc (model-view-controller) design pattern. in general, this design pattern gives us a quick and effective solution to the frequently occurring problems in the software development process (gamma et al., 1994). the mvc design pattern is typically used for developing web-based software applications. to be more specific, figure 3 gives an overview of how mvc works in a simple manner. the model represents our metadata stored in the metadatabase, whereas the views are equivalent to html pages, and dbclear acts as the controller. this design pattern is therefore a strong choice for our web information system. drawing on the mvc design pattern, we can build the architecture of our application software as shown in figure 4. we can now see “what does what,” and clear distinctions between parts of the software. transforming ddi 3 into such a metadatabase format (in this case, dbclear format) is actually a process of flattening the hierarchical structure of ddi 3 into a tabular structure. a detailed explanation of the hierarchical structure of ddi 3 can be found in ionescu 2007. however, we focus here on our particular method of transformation. figure 3. mvc design pattern iassist quarterly spring summer 2009 19 figure 5 depicts an overview of the ddi 3 structure. we transform that structure into a tabular one as shown in table 1 where each record is considered a resource. indeed, this format is quite similar to the two-dimensional data model, hence its representation is easy to understand. at the lower level, this tabular structure is written in dbclear’s xml metadata schema. we transform the ddi 3 into dbclear’s metadata schema using an xsl transformation, which is then translated by dbclear into its own rdbms (relational database management system) schema (a more detailed explanation can be read in hellweg et al., 2002). the schema is compatible with several types of rdbms, but the current system we use is postgresql. web information system the second main part of this project is to build a web information system to present the end user with various views and perspectives on the metadata in a simple but effective way. one useful feature for resource discovery on the web is “faceted browsing.” faceted browsing allows the user to explore the metadata and its additional information by filtering unnecessary parts. during the browsing, users have a short overview of the available information and they can delve into more detail if required. this kind of figure 4. software architecture figure 5. short overview of the ddi 3 structure 20 iassist quarterly spring summer 2009 technique is a well-known and preferred method for most users as studied in yee et al., 2003. for our application, it is more or less like a guided search with predefined categories (controlled vocabularies), where its precision and recall are 100 percent. dbclear uses apache lucene as a search engine library to index and retrieve the data. based on the previous project, we applied faceted browsing to show a list of variables and their details ordered by (a) census years (”variablenliste”), (b) hierarchical subjects8 (”thematische gliederung”), and (c) time line matrix of variables (”variablen-zeitpunkte-matrix”). we currently are staying with these three main aspects, despite the fact that other orders can also be applied. at the low level, the dbclear application produces an xml document for each request by default. this xml document, which we call a “raw page,” is not easy to read by common users and therefore we transform it into a valid html document and integrate it into typo3 as a basic content management system for our web application. we use xsl transformations to transform xml into html documents, which allows us to customize the html documents according to the requirements without changing the source. we also enhance the transformation using java and groovy embedded in the xsl. one of the important results of what we built is shown in the screenshot in figure 6. it shows the time line matrix of variables for the current available census years, with an enhancement that users can select the subject(s) as well as census year(s) they are interested in. this matrix represents the core part of the information system, where users can immediately see the comparable and potentially comparable variables over time. as shown in figure 6, users currently select the main subject “bildung und qualifikation” (education and qualification) and see the comparable variables of all variable census year question number question text ef21 ef22 ----1980 f19 what is your name? 1980 f20 how old are you? ---------------table 1: tabular structure as basic schema of dbclear formattable 1: tabular structure as basic schema of dbclear formattable 1: tabular structure as basic schema of dbclear formattable 1: tabular structure as basic schema of dbclear format figure 6. time line matrix of variables iassist quarterly spring summer 2009 21 included census years. the variables in the blue cells indicate possible comparability of the particular subjects. as mentioned, since the subjects are hierarchically outlined and grouped, all sub-subjects belonging to the selected main subject are shown as well as the corresponding variables in the matrix. the next screenshot, as seen in figure 7, shows a detailed view of a variable9. we selected variable ef50 for the year 2007 as an example. a short overview (a combination of variable name, question number, and label) is given in the first line. fields describing the subject hierarchy and related variables are also provided at the beginning to help users in the search. in addition, other fields are also important as users can also see the question text, filter assignment, or frequency count and can jump to the pdf files related to the variable: key directory, questionnaire, and interviewer’s manual (if available). the missy web site is available at http://www.gesis.org/ missy. conclusions, discussion, and future work we have successfully demonstrated the implementation of ddi 3 for documenting the german microcensus at the variable level. the ddi 3 editor allows us to manipulate ddi 3 “on the fly,” which brings advantages in improving and simplifying the process of data documentation. we emphasize the importance of using ddi as a documentation standard for managing the data life cycle. moreover, we also strongly recommend the use of a metadata repository to address performance issues, or in other words to increase the speed in the searching. as a further achievement, our web information system permits end users to access and browse the metadata in simple ways. since we use the flexible mvc design pattern, it is easy to add features according to the requirements and without changing the metadata. currently, we are making plans to develop our application to cover other ddi 3 features, such as comparison/grouping and filters. to address these issues, more efforts are needed to research the complete data life cycle. an appropriate data model is also needed to handle, for example, the implementation of reusable schemes. our current approach in the development process is xml-centric, which is more straightforward and adequate for a short time project. figure 7. detailed view of a variable 22 iassist quarterly spring summer 2009 we also have not yet integrated the regional microcensus into the current application because the required data are not yet available. we are currently in the process of enhancing the performance of our application to generate a list with a large number of elements. further, we are exploring the abilities of the editor to be used not only for the microcensus but also for other studies. acknowledgments we thank jeanette bohr, andrea lengerer, and julia schroedter for their intensive and cooperative work on this project. we also thank dr. maximilian stempfhuber, prof. dr. christof wolf, and prof. dr. york sure for their kind supervision. finally, we especially thank joachim wackerow for his great work on ddi and for his extensive help with the implementation of ddi. this project is funded by the bmbf germany (01uw0707 “servicezentrum für mikrodaten der gesis / missy ii”). references bohr, jeanette. abschlussbericht missy-nutzerstudie. zuma-methodenbericht, 2007. bohr, jeanette, andrea janssen, and joachim wackerow. “problems of comparability in the german microcensus over time and the new ddi version 3.0.” iassist quarterly, 2006. gamma, e., r. helm, r. johnson, and j. vlissides. design patterns: elements of reusable object-oriented software. addison-wesley professional, 1994. hellweg, h., b. hermes, m. stempfhuber, w. enderle, and t. fischer. “dbclear: a generic system for clearinghouses.” 6th international conference on current research information systems, 2002. hopt, oliver, alerk amin, arofan gregory, jannik jensen, dan kristiansen, and mary vardigan. “questionnaire management and ddi: the qdds case.” ddi working paper series, 2010. doi: http://dx.doi.org/10.3886/ ddiusecases05 hopt, oliver, max stempfhuber, rainer schnell, and anja zwingenberger. “qdds documenting survey questionnaires throughout their lifecycle.” fifth international conference on e-social science, 2009. ionescu, s. “introduction to ddi 3.” presentation slides at cessda expert seminar, 2007. janssen, andrea, and jeanette bohr. “microdata information system missy.” iassist quarterly, 2006. jensen, jannik, et al. “building a modular ddi 3 editor.” ddi working paper series, 2010. doi: http://dx.doi. org/10.3886/ddiusecases02 yee, k.-p., k. swearingen, k. li, and m. hearst. “faceted metadata for image search and browsing.” acm sigchi: human factors in computing systems, 2003. notes 1. andias wira-alam and oliver hopt. contact: andias. wiraalam@gesis.org gesis leibniz institute for the social sciences. 2. raw materials kindly provided by the german federal statistical office (statistisches bundesamt deutschland). 3. information processes in the social sciences which is also a scientific section of gesis – leibniz institute for the social sciences. 4. documenting the complete data life cycle cannot currently be completed since gesis does not have access to materials of the earlier stages. 5. but note that several survey years are missing. 6. for further information, see http://www.qdds.org/ 7. web-based distributed authoring and versioning 8. we have to take into account that each variable has a particular subject; the subjects are hierarchically outlined and grouped into 11 main subjects. 9. note that the detail information is currently only available in german; for this screenshot we translated it into english. by 12 iassist quarterly 2008 john kallas & apostolos linardis1 a documentation model for comparative research based on harmonization strategies abstract this paper deals with studies of fixed design, where comparability is feasible either for culture or for time and culture. for the time dimension, studies are divided into cross-sectional and longitudinal, and for the cultural dimension they are divided into monocultural and cross-cultural. since modern societies are mainly organised into nation-states, cross-cultural studies are carried out more often on country level. however this does not necessarily mean that the organisation of cross-cultural studies in the same country is not possible. comparative cross-cultural studies follow strategies of data harmonization such as the «ex ante input harmonization», the «ex ante output harmonization», and the «ex post harmonization», as well as mixed strategies. these strategies of data harmonization are complex procedures for which success or failure is reflected in the final study product. in this paper, a documentation model is proposed for both longitudinal and cross-cultural studies, for which documentation is particularly difficult. all the remaining study types can be considered sub-cases of the longitudinal and cross-cultural study and consequently are covered by the proposed model. three documentation models are proposed according to the different harmonization strategies examined. finally, the different documentation models are integrated into one. all the models examined are drawn as entity relationship diagrams based on the harmonization methods that govern comparative research. keywords: research documentation, data modelling, comparative research, harmonization strategies. introduction comparative cross-cultural studies follow different strategies of data harmonization such as the «ex ante input harmonization», the «ex ante output harmonization», and the «ex post harmonization», as well as mixed strategies. these strategies are complex procedures for which success or failure is reflected in the final study product. in this paper, a documentation model is proposed for both longitudinal and cross-cultural studies. three documentation models are proposed according to the different harmonization strategies examined. finally, the different documentation models are integrated into one. 1. cross–national studies as a sub-case of cross-cultural studies the original pattern of comparative research brings together at least two social formations. some researchers identify cross-national with cross-cultural studies. thus, for hantrais (1995), comparative research is a research pattern of the social sciences that aims to conduct comparisons between representations that result from two or more social formations. globalisation and the revolution in communication technology on one side and the developments in europe and the course of its unification on the other side, bring in question our current understanding of a «nation–state» as well as the convention to consider comparative research simply as cross-national research. globalisation began in the economic sector and next swept across sectors of policy, culture, and knowledge production. the consequence of this evolution was the gradual delimitation of relations, and of the role of the nation-state (albrow 1998). this change became evident in social science literature as the loss of territoriality followed by the reduction in the sovereignty of nation-states and denationalization (zurn 1998). despite the controversial nature of the «nation-state», most cross-cultural studies are still organised as cross-national ones. however, in order to also cover the case of crosscultural studies that are not cross-national, we will use the term «cultures» instead of the term «countries» which is commonly used in cross-national studies. moreover, as previously stated, cultural discrepancies may exist in the same «nation–state» either on a regional or local level. a basic cultural difference is language. countries such as belgium, finland, and luxembourg are obligated to carry out any national study in multiple languages. the action plan of the european science foundation reports explicitly: «translations should be made into any language iassist quarterly winter summer 2008 13 which is used as a first language by five percent or more of a country’s population» (1999, 10). this leads us to consider cross-cultural research to have a wider nature than cross-national research and cross-national research as a sub-case of cross-cultural research with a determined culture, namely the nation–state. 2. formal description of the data element in crosscultural research in empirical fixed-design studies (robson 2007), data production is organised based on a data schema. this data schema is constructed on the basis of statistical ontology, which predicts that the examined population is constructed of similar units of observation, each of which is described by concrete, distinguishable attributes, each of which is represented by a variable. the adoption of statistical ontology as an organisational model of the data schema of fixed-design studies has two basic advantages. the first advantage is that it allows the application of statistics as a method of data analysis. the second advantage is that it allows for the development of a documentation initiative for studies of fixed design. each population attribute is formally described by a data element that is defined by one concept and by one pattern of value determination (kallas and linardis 2009). each concept is defined by one term and one definition. each value determination pattern is defined as a) a mono-dimensional classification, b) a number that results from the direct measurement of a concept, or c) text. the unit of observation, as it is introduced by statistical ontology, is a mathematical schema that does not always correspond to real objects of observation. that is, to social objects that are presented in social practices independently from the observer. the relation of a unit of observation with a real object of observation occurs in the context of each concrete study. in certain cases where the examined social phenomenon consists of the relationships between more than one object of observation (for example in the case of a household in which we usually have at least two objects: the household and the members of the household), the unit of observation does not correspond to real objects of observation but simply represents the total set of all attributes of individual objects (kallas 2005). in cross-cultural studies where partial recordings describe different societies (the objects of observation based on which the formal descriptions of social phenomena are constructed) it is possible to differentiate from society to society and consequently from recording to recording. this difference means that either the corresponding objects between two recordings cannot be described by the same data elements, or that certain data elements are not precisely the same. consequently, in cross-cultural research, a data element should be described in reference to the object of observation that it describes. the object of observation in each study is related to one unit of observation, which should be documented based on the following: a) the definition of objects of observation that make up the unit of observation; b) the social system in the context of which the objects of observation are constructed. each object of observation is notionally defined in the context of a concrete social system. when it is used in the context of another social system then it is differentiated notionally, and consequently it is described via a different pattern. these differences also concern the data elements whereby the object is defined; c) the territory where the recording is organized. the universe is defined by a) a territory, b) a unit of observation and, c) other partial determinations (such as age, sex, marital status, income etc.). the universe refers to both the data elements as well as to the objects of observation. vice versa, each data element or object of observation refers to a universe. a documentation schema of a data element suitable for the cross-cultural research is shown in figure 1 (see next page). it is possible to define a territory by a multilingual, controlled vocabulary that includes country and region names (for example the nomenclature of territorial units for statistics or nuts classification) as well as other predefined values such as international, european, etc. it is also possible to define the object of observation and social system by a controlled vocabulary. this standardisation further helps in comparative research since the objects of observation are often used by researchers as basic search and comparability criteria. the documentation schema of a data element includes three basic structural components: the concept, the universe, and the classification. it is possible to reuse the same data element for the documentation of one or more studies. it is also probable for some other data elements to be identified partly by reusing a subset of the structural components that compose a data element. in addition, components such as the universe or category schema can also be reused for the determination of other entities. for example, the category schema can be reused either for the determination of a classification, for the determination of a question, or for the determination of a variable. additionally, the universe can also be reused for the determination of study «wave instance». however, the universe of each data element usually constitutes a subset of the general universe of a wave instance. in most cases, both universes are similarly identified at the territory and unit of observation level but differ in their other determiners. 14 iassist quarterly 2008 the data element, the concept, the universe, and the classification are study components for which even small changes in the content should be documented since they may affect the overall comparability. this is the reason why they are defined as versionable objects. these objects are identified by one code but also by a version number (complex key). both major and minor changes in the lower level objects simultaneously alter the version of parent objects. the version change documentation should include the following: a) the new version date, b)who made the change, c) why there was a change, so that the users can comprehend if the version change influences the analysis of data (ddi alliance). 3. approaches in data harmonization harmonized data can be achieved either by using strategies for collecting harmonized data from the beginning (in the study design) or by using harmonization strategies for existing data (granda, hadorn and wolf 2008). 3.1 strategies for the production of harmonized data in cross-cultural studies the following is a short description of strategies followed for the production of harmonized data in cross-cultural studies. • ex ante input harmonization ex ante input harmonization means that the institutions that participate in the study have agreed on common concepts, common measurement patterns of the concepts and also on common questions based on a common source questionnaire. the ex ante input harmonization is used mostly in cross-national studies but is also applied in studies that are conducted in the same country. for example, a question relative to the underground in greece, such as «how many times do you use the underground per week? » is asked only to athenians but not to all greeks. consequently, an agreement is required for the concepts, the measurement patterns, and the questions even in the same country in order for the questions to have meaning for all participants. in crosscultural studies where ex ante input is applied, no country-specific variations are allowed except the ones that are absolutely essential such as the language used in the questionnaires (ehling 2003). • ex ante output harmonization in ex ante output harmonization the institutions that participate in the study have agreed on common concepts and common measurement patterns. the objective is fixed and the choice of suitable questions is left to participating research groups who adapt the questions to the cultural particularities of the universe that they study. each research group determines its own concepts and measurement patterns; however, they must correspond to the common concept via transformation routines. for example, let us assume that in a cross-national study the participants are asked to indicate their highest level of education. a figure 1: documentation schema of a data element iassist quarterly winter summer 2008 15 common measurement pattern for education level is the use of an international classification such as the international standard classification of education (isced). the measurement of education level via isced may serve international needs but not national ones, since a country would serve itself more if its national data were detailed enough. another reason to follow this strategy is that the same data can be collected simultaneously for different studies. more concretely lene mejer reports: «output harmonization means to give a common internationally agreed definition for a variable and then leave to each single member state to decide on its implementation. each member state decides what is the best national source for the variable (for example from already-existing surveys and / or registers) » (2003, 69). this strategy is used mostly for cross-national studies; however, it can be used as methodology in cross-cultural studies. • mixed strategy some studies, while they follow the strategy of ex ante input harmonization (such as the european social survey2 or ess and the international social survey programme3 or issp), in certain selected data elements they apply the ex ante output harmonization strategy. for example the highest education level in ess is ex ante output harmonized while the study is ex ante input harmonized. in this case, the harmonization strategy should be placed at data element level per wave and not at study level. there is another case of «mixed strategies» where the literal question is agreed upon by the harmonization committee but the category schema is fixed by each country separately. 3.2 harmonization strategy for existing data: ex post harmonization ex post harmonization is a harmonization strategy where the total study results from already-existing studies. in ex post harmonization, the institutions that participate in the study agree on common concepts, on common measurement patterns, and on common universes (common data elements). they also agree on alreadyexisting studies that have to be ex post harmonized using transformation routines. the achievement of harmonized data via this process is not guaranteed, even if it has been optimally designed, because of the diversity of concepts and measurement patterns in existing studies. since no new questions are created, the basic structural elements of these studies are the data elements. new data elements are created that reference already-existing ones. for the implementation of such studies (that resemble research programs more than studies) transformation routines are required, which are written with statistical software. the difficulty of implementing studies following the strategy of ex post harmonization lies in the localisation of common concepts and measurement patterns between the universes. a very useful tool for such studies would be a bank of concepts, classifications, and universes for the localisation of similar data elements. 4. data archives and the documentation process: a comparative perspective in recent years, new organizations have been created in europe called data archives (da). da deal with the accumulation, documentation, and dissemination of data. these organizations support secondary analysis and comparative research, and act as mediators between the producers and analysts. the european council is called the council of european social science data archives (cessda). each da must document its own studies based on a common strategy that ensures the following: • reuse of common structural study components of a simple study in the same da each study component can be constructed from other structural components. for example, study components such as classification, question, and variable use some common structural components such as codes and categories. often the category schema of classification, question, and variable coincide. for example, the codes and categories that are used for isced classification (for the corresponding question but also for the corresponding variable in a statistical data file) may coincide. in this case the study components’ common structural components should be imported just once and then reused (ddi alliance) even in the case of a simple study that is neither longitudinal, nor crosscultural. each person who documents a study should follow these rules so that double entries are avoided. to aid in this laborious documentation work, certain processes for localisation of common structural components can be automated. • comparability of a longitudinal study in the same da. • the comparability of a longitudinal study in the same da lies in the reuse of study components between waves. components such as concepts, classifications, universes, questions, and variables are principal components for comparability between waves of longitudinal studies and they should be reused in the various waves. consequently, each da should maintain local banks of all these study components. • comparability of different studies in the same da comparability of different studies is also based on the reusability of the same principal study components 16 iassist quarterly 2008 between the various studies, as in a longitudinal study. • comparability of a cross-cultural study in the same da the proper documentation of a cross-cultural study involves all the participating organizations and it differs depending on the harmonization strategy that has been followed. the documentation of a crosscultural study based on the harmonization strategy followed is developed analytically in section five. while the documentation of a cross-cultural study often occurs in different da’s, in this work we will deal with the documentation of a cross-cultural study in the same da (or data-metadata repository). • study comparability in different da’s study comparability in different da’s is a very critical process for wider comparative research but this will be analyzed in a later work. summarizing the above, it is immediately evident that the documentation procedure is a difficult and laborious process. on the other hand, the result of this procedure will be useful for the wider research community, particularly for those researchers who want to carry out comparative research and secondary analysis. the comparative documentation further strengthens the role of da’s. the documentation process is best carried out in collaboration with the primary data producers as well as with the statistical institutes. 5. the documentation of a cross-cultural, longitudinal study the documentation process is rendered particularly difficult and laborious in the case of cross-cultural, longitudinal research. the collaborating institutions should follow the documentation of the coordinating institution, since the resulting documentation will be based on common agreed concepts, measurement patterns, questions, and universes. the documentation completed by the coordinating institution should not be changed by the participating institutions. the documentation language of the coordinator is the common agreed language (usually english). in the cases that follow, the model is presented first and then a description of how the model should be used by participants based on the agreed-upon harmonization strategy. it takes into consideration the most complex study type, the cross-cultural, longitudinal study, since all other studies can be documented based on this. the diachronism relies on the reuse of study components for each wave. the multiculturalism lies in the creation of references between the source and universe study components. it should be noted that the models that follow concern the most complex study type – the cross-cultural, longitudinal study – but they document just one study. another limitation is that the documentation takes place in the same da and not in distributed documentation systems. below is a short description of the main entities used in the models: • study: the entity that is used to store the general study information such as title, more general objectives, summary etc. the main purpose of adopting such an entity, beyond the storage of general information, is that it aims to unify all study waves. the study level documentation is completed by the coordinating institution in the common agreed language. translation into other languages occurs only for dissemination reasons. this entity is not reusable but can be referenced by other studies or by other study components. • wave: a longitudinal, cross-cultural study takes place in many time and universe instances. the time instance of a study is called a wave while the universe instance of a study wave is called a wave instance. while documenting, but also while a study is conducted, it is common practice to establish the time and, for each time period, to receive snapshots for the various universes that participate in the study. in addition, at this level, the general wave title, any special objectives perwave, and the total duration of the study wave are all recorded. information such as the universes or institutions that participate in the study may be recovered automatically from the wave instance documentation level so that no differences in the aggregated fields of wave level exist. the most crucial documentation at wave level has to do with the determination of the study harmonization strategy. it is also crucial this be selected from a controlled vocabulary where the user chooses between the following options: a) ex ante input harmonization, b) ex ante output harmonization, c) ex post harmonization, or d) mixed strategies. the harmonization strategy is determined at wave level, not at study level, because it is possible (although rare) that the harmonization strategy may change from wave to wave. the wave entity is also used for the grouping of source data elements, of common source questionnaires (when they exist), and of the harmonized statistical data files. the documentation at wave level should be completed by the coordinating institution in the common agreed language. translation into other languages is done only for dissemination reasons. the wave entity cannot be reused but can be referenced by other study components.. iassist quarterly winter summer 2008 17 • wave instance: the entity that is used for storing information concerning wave snapshots per universe. the documentation of wave instance level is completed by all the research groups that participate in the study, in the languages that have been decided per group but also in the common agreed language for dissemination purposes. the wave instance includes extensive information such as universe, sampling methods, participating institutions by role (local coordinator, financiers, data producers, organizations responsible for data dissemination), researchers, sampling frame, data collection method, time of data collection, and description of weights (accompanied by weighting methodology). the wave instance is not a reusable object but can be referenced by other study components. • source data element and universe-specific4 data element: entities used to store information concerning the data elements that were introduced in section two. the implementation of these two entities in a database does not necessarily require the creation of two tables for the two types of data elements; however, both data elements are presented as separate entities in the entity relationship diagrams in order that the required relationships are evident. the same holds for the questionnaires, the questions, the data files, and the variables. the data element and its structural components are reusable entities for different studies or study waves. • source questionnaire and universe-specific questionnaire: entities used to store general information concerning the questionnaire such as the number of questions, type of questionnaire (standardized versus non-standardized), abstract, and link to the questionnaire file. this is also a grouping entity for questions. questionnaires are not reusable entities but have to be defined again in each wave or wave instance. • source question and universespecific question: according to kallas and linardis (2009), the questions are composed of some or all structural elements presented in figure two. the question is also a reusable object. • harmonized data file and universe-specific data file: entities used to store general information about the statistical data files such as the number of variables, the number of cases, and likely a link to the statistical data file. data files are also used as grouping entities for the variables. the data files are not reusable entities and have to be defined again for each wave or wave instance. • harmonized variable and universe-specific variable: the variables consist of structural elements such as name, description, type, measurement level, and category schema. the variable entity, as it is described here, consists only of metadata and not of data. the same variable may have a number of data depictions but in different statistical data files. the variable is a reusable entity. • finally, the transformation routine is the process that describes the necessary transformations of the universe-specific data element to source data element. 5.1. case 1: study documentation following ex ante input harmonization strategy the documentation process for ex ante input harmonization is portrayed in figure 3 ( on next page). the left parallelogram portrays the documentation that should be completed by the coordinating institution while the right one portrays the documentation that should be completed by the participating institutions. figure 2; : documentation schema of a question 18 iassist quarterly 2008 the documentation may be published in intermediary stages (indicated below). the basic principle of ex ante input harmonization is the standardization of data elements and questions between the research groups. the stages of documentation are as follows: stage 1: documentation of the general context of the study and wave, and documentation of data elements (documentation provided by coordinating institution). 1. documentation of the general context of the study: documentation is completed by the coordinator at the beginning of the study. 2. documentation of the general context of the study wave: study waves have to reference the corresponding study. 3. documentation of common data elements (common concepts, measurement patterns, and universes): data elements should reference the study wave and not the study because the data elements may differ from wave to wave. for example in the ess, there are some data elements that are used in all study waves while others are added or removed periodically. fundamental practices to ensure comparability: dissemination of stage one documentation to the participating institutions. stage 2: translation of the context of study and study wave; translation of the data elements; determination of each wave instance (documentation provided by participating institutions). 1. after the dissemination of stage one documentation to the participating institutions, each participating institution should translate the general context of the study and of the study wave, and the data elements, for dissemination reasons. 2 each participating institution should then document the wave instance based on the universe it represents. each wave instance has to reference the corresponding wave. (first intermediate phase of study publication) stage 3: documentation of source questionnaire/s (documentation provided by coordinating institution). 1. documentation of source questionnaire/s: each questionnaire should reference the corresponding wave. 2. documentation of source questions: the reference between source questions and the corresponding data elements as well as between source questions and the source questionnaire is required. fundamental practices to ensure comparability: a) dissemination of stage three documentation to the participating institutions, b) creation of a statistical file template with common variable names (based on the data elements) for all participating institutions, c) dissemination of the template to the participating institutions. stage 4: translation of source questionnaire/s, leading to universe-specific questionnaire/s (documentation provided by participating institutions). 1. documentation of universe-specific questionnaire/s: the reference between the universe-specific questionnaire and the corresponding wave instance is required. 2. documentation of universe-specific questions: the questions specific to each universe are created via translation of the source questions in language or languages decided by each research group. the reference between universe-specific questions and source questions as well as with the corresponding universe-specific questionnaire is required. fundamental practices to ensure comparability: a) each institution conducts the research, collects the data, and submits the statistical data files to the coordinator, according to the template already sent by the coordinator, b) at the same time, each institution preserves the data files for the stage six documentation procedure. (second intermediate phase of study publication) stage 5: documentation of harmonized statistical data file/s (documentation provided by coordinating institution). 1. documentation of harmonized statistical data files that have come from the merging of universe-specific data files: the reference between harmonized statistical data file/s and wave is required. 2. documentation of harmonized variables: the reference between harmonized variables, the harmonized statistical file, and the corresponding source questions is required. fundamental practices to ensure comparability: dissemination of stage five documentation to the participating institutions. iassist quarterly winter summer 2008 19 stage 6: documentation of universe-specific data files (documentation provided by participating institutions). 1. documentation of universe-specific statistical data files: the reference between universe specific statistical data files and the corresponding wave instances is required. 2. documentation of universe-specific variables: the reference between universe-specific variables and the statistical data file they belong to, universe-specific variables and the corresponding universe-specific questions, as well as reference between universespecific and harmonized variables is required. (final phase of study publication) 5.2. case 2: study documentation following ex ante output harmonization strategy. figure 3: study documentation model following ex ante input harmonization strategy 20 iassist quarterly 2008 a basic difference between ex ante input harmonization strategy and ex ante output harmonization strategy is that the second presupposes the determination of universe-specific data elements by the participating institutions. consequently, the participating institutions should document the universe-specific data elements and reference them to the common agreed data elements via transformation routines. also, there is no source question, just a source data element. on the other hand, there are universe-specific questions but these are considered mostly as additional documentation of universe-specific data elements not as fundamental structural study components. the documentation process for ex ante output harmonization strategy is portrayed in figure 4. for simplistic reasons we have not drawn the structural elements of the data element again (concept, measurement pattern, and universe). the left parallelogram portrays the documentation that should be completed by the coordinating institution while the right one portrays the documentation that should be completed by the participating institutions. following this strategy, there are five study documentation stages instead of six. this occurs because stage three of ex ante input harmonization does not make sense here since there are no source questionnaires or source questions. the figure 4: study documentation model following ex ante output harmonization strategy iassist quarterly winter summer 2008 21 differences between the documentation procedures of the two strategies are summarised as follows: • stage 2. this stage includes two additional documentation actions to be performed by the participating institutions: 2.3) documentation of universe-specific data elements; and 2.4) documentation of transformation routines of source to universe-specific data elements. • stage 3. as was mentioned before, stage three does not exist. nevertheless, the coordinating institution can establish some fundamental practices to ensure comparability such as: a) the creation of a statistical file template with common variable names for all participating institutions, based on the data elements; and b) the dissemination of the template to the participating institutions. • stage 4. phase 4.2 is different because there are no source questions. consequently, the universespecific questions are developed from scratch instead of being produced as translations of the source questions. reference between universespecific questions and questionnaire as well as between universe-specific questions and data elements is required. • stage 5. phase 5.2 is different because there is no reference between harmonized variables and source questions. reference between harmonized variables and harmonized statistical file is required. 5.3. case 3: study documentation following ex post harmonization strategy the basic difference between this strategy and ex ante output harmonization lies in its relationship to the common concept. in ex ante output harmonization, its relationship to the common concept is guaranteed because the researchers design the wave instances from scratch, keeping in mind the common data elements. in ex post harmonization, the individual studies have been designed autonomously by the researchers without adhering to a common concept. thus, the relationship to the common concept is not guaranteed. on the other hand, the two harmonization strategies have a lot of similarities related to methodological issues. the documentation process is similar to the one that was described based on figure 4. the basic difference is that in ex post harmonization, the new study derives from already-existing studies. consequently, it should initially be documented using the already-existing study waves or wave instances from which the new study derives. it would provide great relief from the excessive documenting load for researchers working on an ex post harmonized study if documentation of existing study waves or wave instances was already available. let us assume that a new study is designed following the ex post harmonization strategy. the new study concerns attitudes for a set of countries (two of which are: cyprus and russia), for the time period 2004-2005. the coordinating committee decides to harmonize ex post the second wave of the ess. according to jowell et al. (2007), cyprus and russia did not participate in the second wave of the ess. nevertheless, the coordinating committee is aware of the existence of other national attitude studies for cyprus and russia during 2004-2005 and decides to harmonize them ex post. at the same time, the coordinating committee has to decide on the common data elements of the new study. after the common data elements have been defined, the data elements of the existing surveys have to be transformed via routines to the common ones. it is common for different groups to undertake the transformations of different studies. in our example, three groups will undertake the burden of transformations of source data elements to the common agreed data elements: one group for the ess, one for the cyprian study, and a group for the russian one. the documentation process of the three groups includes: a) the documentation of each new wave instance; b) the reference of already-documented data elements (ideally) to the wave instance; or c) the documentation from scratch of all previously conducted studies in the research program, if their documentation does not exist; and d) the application of transformation routines to universe-specific data elements and source data elements. the documentation process in ex post harmonization is portrayed in figure 5. the left parallelogram portrays the documentation that should be completed by the coordinating institution of the research project, while the right one portrays the documentation that should be completed by other institutions that participate in the research project. these institutions will have undertaken the documentation of concrete wave instances from alreadyexisting studies. figure 5 differs from figure 4 in the following ways: • each new study wave can be designed based on existing study waves and/or existing wave instances. consequently, suitable documentation at wave level is required. • the relationship between source data element and universe-specific data element is a “many to many” relationship, since the same universespecific data element may correspond to more than one source data element in the same system. for example, a universe-specific data element may correspond both to the source data element of the 22 iassist quarterly 2008 initial study and to a new source data element that was created during the design of a study following ex post harmonization. 5.4. common documentation model for all harmonization strategies as already mentioned, the harmonization strategy followed should be determined at wave level. in the case where the strategy followed is ex ante input harmonization, the documentation model is the one in figure 3. the documentation models of ex ante output harmonization and ex post harmonization are depicted in figure 4 and figure 5, respectively. in the case where the harmonization method is determined as «mixed strategies», the harmonization strategy has to be determined for each data element per-wave. the entity relationship diagrams 3, 4 and 5 are unified in figure 6. in using this model, the documentation process will differ depending on the choice of harmonization strategy. the documentation model in figure 6 also serves the needs for the documentation of different studies based on a comparative perspective. this is feasible because entities such as concept, classification, universe, data element, question, and variable, that constitute the basic study components for comparative research, are reusable entities for all studies. the reusability of these study components, in a documentation system of a specialized architecture, aims at comparative documentation between different studies. another useful outcome of such a documentation figure 5: study documentation model following ex post harmonization strategy iassist quarterly winter summer 2008 23 model is that it is feasible for a researcher to locate universe-specific study components derived from source study components. when they are referenced, the study components referred to above can never be deleted or changed. these components are identified by persistent identifiers (pids). if one of these components has to change then a new version of this component has to be created. multilingualism of study components is applied in two cases: a) when a component is translated by an institution in order for its translation to be an “active component” of the study (for example translation of the source questions to universe-specific questions); and b) just for dissemination reasons (for example translation of a study’s abstract). the reasons for translation of a study component should be declared in the documentation. in the first case, both major and minor changes may lead to version change of the study component, not so in the second one. 6. conclusions the general documentation process in a da or in a metadata and data repository is based mostly on ex post harmonization procedures. the institutions that document a longitudinal, cross-cultural study should do so based on already-existing documentation. in the case of longitudinal studies, this should be done by repeating study components from other waves, and in the case of cross-cultural studies, by referencing source and universe-specific objects. figure 6: : study documentation model used for every harmonization strategy 24 iassist quarterly 2008 consequently, the documentation procedure not only involves the typical description of the components derived in the context of a single research project, but also all of the a posteriori references (additional documentation) between independently-designed studies. the proposed study documentation procedure is laborious for the researchers that are making the documentation. additionally, s/he should know the specifics of each study, which presupposes a close collaboration with the primary investigators. moreover, the “golden super rule” of metanet research project states: «metadata are as important as data, and metadata need as much work as data» (sundgren, 2003, 129). the study documentation with suitable metadata can provide particular advantages in the identification of «equivalent» or «equal» study components either for longitudinal or cross-cultural studies, or even studies with similar subjects. documenting studies based on the proposed documentation model, allows the researcher to: • search and locate questions that use either common concepts, common classifications, questions that are addressed in common universes, or even questions that use common data elements; • search all the variables that are derived from the same questions; • locate data element transformation routines, so that the sequence from one data element to the other is clear; • locate all the translations of a source question, likely accompanied by qualitative criteria such as validity and reliability but also non-response rates (sarris et al. 2007); • locate data according to the following criteria: a) the concepts that the data imply, b) the universes the data refer to, c) the time period the data refer to or the time the survey was conducted, d) concrete classifications based on which the data have been produced. this model can also be extended and modified so that it may cover the documentation of comparative research in distributed environments. in this case, not only the model would be of particular interest, but also the flexible and functional architecture of the overall documentation system. the extension of the model in distributed environments as well as the architecture of such a system will be analyzed in a future paper. references albrow, martin, and raimund fellinger. 1998. abschied vom nationalstaat: staat und gesellschaft im globalen zeitalter. frankfurt am main: suhrkamp. ddi alliance. ddi 3.0. http://www.ddialliance.org/ddi3/ index.html ehling, manfred. 2003. harmonizing data in official statistics: development, procedures, and data quality. in advances in cross-national comparison: a european working group for demographic and socioeconomic variables, ed. jurgen h.p. hoffmeyer–zlotnik and christof wolf, 17-31. new york: kluwer academic / plenum publishers. european science foundation (esf). 1999. blueprint for a european social survey. strasbourg: author. granda, peter, rheto hadorn, and christof wolf. harmonizing survey data. paper presented at the international conference on survey methods in multinational, multiregional, and multicultural contexts (3mc), june 25-28 in berlin germany. hantrais, linda. 1995. comparative research methods. social research update 13, (summer). http://sru.soc.surrey. ac.uk/sru13.html jowell, roger, max kaase, rory fitzgerald, and gillian eva. 2007. the european social survey as a measurement model. in measuring attitudes cross–nationally: lessons from the european social survey, ed. roger jowell, caroline roberts, rory fitzgerald, and gillian eva, 1-31. los angeles: sage. kallas, john and apostolos linardis. 2009. questionnaire documentation model on the needs of comparative research. paper under review for publication. kallas, john. 2005. data modelling and the formation of a grid. in the node for secondary processing, ed. john kallas, 32-45. athens: national centre for social research. mejer, lene. 2003. harmonization of socio-economic variables in eu statistics. in advances in cross-national comparison: a european working group for demographic and socio–economic variables, ed. jurgen h.p. hoffmeyer– zlotnik and christof wolf, 67-85. new york: kluwer academic / plenum publishers. robson, colin. 2007. real world research: a resource for social scientists and practitioner-researchers. oxford: blackwell. sarris, willem e. and irmtraud n. gallhofer. 2007. can questions travel successfully? in measuring attitudes cross–nationally: lessons from the european social iassist quarterly winter summer 2008 25 survey, ed. roger jowell, caroline roberts, rory fitzgerald, and gillian eva, 53-74. los angeles: sage. sundgren, bo. 2003. developing and implementing statistical metadata systems. http://www.epros.ed.ac.uk/ metanet/deliverables/d6/ist-1999-29093-d6.doc zürn, michael. 1998. politik jenseits des nationakstaats. frankfurt am main: suhrkamp. notes 1 kallas, ioannis ,university of the aegean, department of sociology ,academic field: methods & information techniques of the social sciences.address: tertseti & mikras asias str., 81 100 mytilene, lesvos, greece./ tel. (+30) 22510 36559. email: i.kallas@soc.aegean. linardis, apostolos . national centre for social research . address: 14-18 messoghion av., gr-115 27, p.o.b 142 32, athens,greece.tel. (+30) 210 7491656.email: alinardis@ ekke.gr 2 http://www.europeansocialsurvey.org/ 3 http://www.issp.org/ 4 the documentation of all universe-specific entities is completed by all participating research groups, in the languages that have been decided by each group but also in the common agreed language. the u.s. public use census microdata files as a source for the study of long-term social change by steven ruggles ' department of history university of minnesota the united states public use microdata samples are machine-readable hierarchical files consisting of individual-level and household-level records drawn from the federal decennial censuses. samples covering nine census years between 1880 and 1990 are currently available or in preparation. taken together, these microdata comprise the richest source of quantitative information on long-term changes in the american population. because these samples were created at different times by different investigators, however, they have incompatible documentation and a wide variety of record layouts and coding schemes. these differences among the samples inhibit their use as a time-series. at the social history research laboratory of the university of minnesota, we are planning to convert the series of public use samples into a single coherent form. the success of this project will depend on the usefulness of the data scries to a broad range of social scientists. this essay describes the history of the public use samples and some of their potential applications for time-scries analysis, in the hope of stimulating interest and suggestions at an early stage of our work. background social scientists have increasingly recognized the need to study society as a process. if we confine our analyses to the stale of society at a single moment, we cannot hope to understand the sources of social change. sociologists, economists and demographers have developed a variety of quantitative data sources to study social change, including retrospective surveys, repetitions of early social surveys, and longitudinal surveys. although such data sources are essential, they are usually limited to the analysis of changes during the past thirty years. the study of longer term change— over the past 100 or 150 years— has been sharply constrained by the limited availability of consistent data series. analysts of nineteenth-century society have often turned to institutional and bureaucratic records, such as those generated by churches and the military, but these sources are typically available only for the distant past and they are limited to the study of specific population subgroups. the decennial census is the most consistent general source of information about the american population over the past two centuries. quantitative studies of longterm social change have always relied on the published tabulations of the census, but these data have substantial limitations. in each period, the topics addressed by census publications have focussed on contemporary concerns, and these concerns have shifted dramatically over the past century. for example, the early twentieth century census volumes include a wealth of data on immigrants, but virtually nothing on family composition. moreover, the high costs of tabulation before the introduction of modem data processing equipment meant that few cross-classifications of census data were possible, and much of the information collected by the census was never tabulated at all. even for recent census years, the published census volumes have significant limitations for the study of social change. despite the dramatic increase in the quantity of published census data in recent years, the census bureau cannot anticipate all the questions social scientists want to ask. the census bureau has addressed these problems by producing individual-level public use samples of the census (u.s. bureau of the census 1972, 1973, 1982a, 1989). the first public use sample was created as a byproduct of the 1960 census (u.s. bureau of the census 1954). in an effort to meet the needs of scholars who needed specialized tabulations, the census bureau created a 1 in ickx) extract of the basic data tapes they had used to create tabulations for the published census volumes. to preserve confidentiality, the census bureau removed names, addresses, and other potentially identifying information. the 1960 public use sample was an immediate success. not only did it allow researchers to make tabulations tailored to their specific research questions, but it also allowed them to apply new methods to the analysis of census data, especially multivariate techniques. but the sample did have two significant limitations. first, the sample size was relatively small. the 1 in 1000 sample density yielded about 180,(xx) person records. given the modest capacity of computers in 1964, this was a lot of cases, but as researchers began to use the sample for detailed analysis of small population subgroups, its limitations became apparent. second, the 1960 public use sample provided highly limited geographic information. in its zeal to preserve confidentiality, the census bureau stripped off all information on places below the lassist quarterty state level. this meant, for example, that it was impossible to extract a subsample of the new york city population. both of these problems were addressed by the 1970 public use samples. the 1 in 1000 density of the 1950 sample was increased dramatically; the census bureau provided six independent public use samples for 1970, each of which had a 1 in 100 density. users who required an exceptionally large number of caaes could combine the samples to obtain a six percent density, or about 12 miuion person records. in addition, the 1970 samples provided a variety of alternate geographic codes, although the census bureau still did not identify any places of less than 250,000 population. in conjunction with the 1970 public use samples, the census bureau released a new version of the 1960 public use sample. they enlarged the sample density from 1 in 1000 to 1 in 100, and at the same time reorganized the coding schemes and record layouts to be compatible with the samples from 1970. this compatibility made it relatively easy for investigators to pool data from 1950 and 1970, and thus incorporate change into their analyses. by the late 1970s, the public use samples had become one of the essential tools of american social scientists. it was in this climate that two separate teams of researchers independently came up with the idea of creating historical public use samples for earlier census years. samuel preston directed projects at the university of washington and the university of pennsylvania to produce a l-in-750 sample of the 1900 census and a l-in-250 sample of the 1910 census (graham, 1980; strong et al., 1989). meanwhile, halliman winsborough and a group of others at the university of wisconsin and the census bureau created 1 in 100 samples for the censuses of 1940 and 1950 (u.s. bureau of the census 1984a, 1984b). a fifth historical public use sample is now underway. at the university of minnesota, we are creating a 1 in 100 sample of the 1880 census. that project is about half done, and a preliminary 1 in 1000 subsample is already available (ruggles and menard, 1990; social history research laboratory, 1990). in addition, we have applied for funds to create a public use sample of the 1920 census; if that project is funded, the 1920 sample will be complete by 1997. in the meantime, the census bureau has released public use samples for the 1980 census, and has scheduled a 1993 release date for samples of the 1990 census (u.s. bureau of the census, 1982a, 1989). these samples include greater geographic and subject content detail than either the 1960 or 1970 public use samples. what all this means is that we can anticipate a series of public use microdata samples of the u.s. census covering the years 1880, 1900, 1910, 1920, 1940, 1950, 1950, 1970, 1980 and 1990. this data series will constitute a resource of unprecedented power for the study of longterm social change. the availability of the historical census files is especially important, because few national microdata files of any sort exist for the period before 1960. furthermore, as one goes farther back in time the published tabulations of the census become increasingly sketchy and the problems of comparability increase. table 1 summarizes the availability of variables for each of the census years ciurently available or in preparation. eleven basic questions were asked in all census years, and twenty-two inquiries are available for at least seven of the nine census years. there are a significant number of variables omitted from table 1 that are available in only one or two census years. note that in addition to the differences in available variables across census years, there are also multiple versions of the samples for recent years that incorporate slightly differing variables. a detailed discussion of comparability problems can be found in ruggles (1991). applications of the public use microdata series the range of potential topics that can be addressed with these data is far too great to describe within the page limitations of this paper. the following paragraphs are intended only to suggest some of the most obvious topics of investigation. 1) household composition. american living arrangements have been radically transformed since the late nineteenth century. in 1880, for example, 77 percent of the elderly lived with their children or with extended kin, compared with 24 percent in 1980. the frequency of primary individuals has increased about eight-fold, and residence as secondary individuals or extended kin has dropped almost as dramatically. these changes began shortly after the turn of the century, and accelerated after 1940 (ruggles 198b; ruggles and king, forthcoming). we are only beginning to understand the dimensions of change in family structure over the past century, and the analysis of the determinants of that transformation has yet to be seriously undertaken. the public use samples are the only detailed national source of information about changing living arrangements in the nineteenth century and first half of the twentieth century. all the public use samples provide sufficient information to construct fully compatible and highly detailed measures of household composition and family interrelationships. 2) fertility. between 1850 and 1940 the total fertility rate for while americans declined from about 5.4 to 2.2 (coale and zelnik 1963:36). research on early fertility summer 1991 trends in america has relied for the most part on childwoman ratios (forsterand tucker 1972; yasuba 1953) and backward projections of age distributions in the published census volumes (coale and zelnick 1953; mcclellan and zeckhauser 1982). neither of these techniques allows close analysis of marital fertility or fertility differentials. analyses of fertility using ownchild techniques were among the earliest and most fruitful multi-sample studies carried out with the two original pubuc use samples produced by the census bureau (e.g. rindfuss and sweet 1977). the public use microdata series will permit study of differential marital fertility patterns over the period of greatest fertility decline, comparing characteristics such as race, occupational class, region, literacy, size of locality, family structure, and a wide variety of other variables. the richness of these data will greatly enhance our ability to analyze the determinants of early fertility decline in a developed country, and this may in turn lend insight into the onset of fertility control in developing countries. 3) life course analysis. long-term changes in the timing of major life-course transitions— such as leaving school, leaving home, starting work, marrying, and establishing a separate household— have been studied using both cross-sectional data (modcu, furstenberg, and hershberg 1975) and retrospective survey data (hogan 1981). both approaches reveal that american society has become more age-graded during the twentieth century: people tend to pass through the major transitions to adulthood at increasingly prescribed ages and in an increasingly prescribed sequence. recently, stevens (1991) suggested that the heterogeneity of the early decades of the twentieth century was a short-term phenomenon brought about by rapid urbanization and immigration from southern and eastern europe. the public use microdata series will provide the opportunity to test this hypothesis through cohort analysis of both the timing of change and of differences among subpopulations. 4) household economy and female labor force participation. much of the research on late nineteenth and early twentieth century social structure has focussed on patterns of employment within the household. some investigators see a fundamental transformation of the household economy with the rise of wage labor; others fxiint to the continued strength of preindustrial modes of informal family labor (katz et. al. 1982; anderson 1971; barron 1984). since the existing studies are based on small local samples of census data, regional variation may explain much of difference in interpretation. the hierarchical organization of the proposed census scries is well suited to study of the household economy. female labor force participation is a closely related and equally controversial issue (bose 1987; conk 1981; folbre and abel 1989: goldin 1980, 1983; openheimer 1970; jaffe 1955). changes in census definitions of employment and labor force participation have complicated such analysis. the public use microdata series will allow researchers to minimize the effects of such changes, since labor force participation can be allocated according to the procedures proposed by abel and folbre (1990); such adjustments are impossible with aggregate data. analysis of the determinants of female labor force participation and child labor during the late nineteenth and twentieth centuries should prove especially revealing. 5) ethnicity and immigration. the questions on nativity in the public use samples makes them a rich lode of information for immigration historians. throughout the period 1880-1970 the census asked about parental birthplaces as well as the respondent's birthplace. most of the census years also provide information on mother tongue and year of immigration. this makes it possible to analyze patterns of acculturation for a wide variety of cultural groups. understanding the varied experience of immigrants in the late nineteenth and early twentieth centuries has taken on a special relevance in light of the recent resurgence of immigration. these topics are intended only as representative examples of the sort of research that can be carried out with the public use microdata series. other key areas of investigation include the transformation of industrial and occupational structure, urbanization, internal migration, nuptiality, and education. the large size of the public use samples increases their versatility by permitting analysis of small population subgroups. consider, for example, some of the topics addressed by minnesota graduate students using the historical public use .samples: -the professional ization of nursing american indian fertility patterns -race differentials in the living arrangements of the elderly labor force composition in minneapolis and st paul the adaptation of scandinavian immigrants -changes in the gender composition of clerical workers -the household structure of early black migrants to lassist quarterly northern cities -italian immigration to the southern u.s. -living arrangements of parentless children these research topics could not be pursued using a general social survey of the scale ordinarily undertaken by academic social scientists. indeed, even the largest social survey carried out by the government— the current population survey — is too small for the detailed analysis of topics like american indian fertility or the professionalization of nursing. the pubhc use samples are the only general source af microdata with sufficient cases to study such small population subgroups. the large scale of the public use samples also makes them the most suitable source of microdata for policy analysis at the state and local levels. policy analysts have traditionally focussed on short-run change, but there is increasing recognition of the need to distinguish longterm secular trends from temporary fluctuations. the public use samples also allow policy analysts to set their investigations of state and local conditions in a comparative national context in summary, the decennial enumerations of the population include a great deal of information on demography and socioeconomic structure that can only be taken advantage of through the public use samples. we presently understand just the broad outlines of the social transformation that has taken place since the late nineteenth century; pubhshed sources provide only limited information on topics such as fertility behavior, urbanization, immigration, household composition, and occupational structure. the public use microdata series allows the construction of comparable cross-tabulations on a wide range of topics that were not covered by census publications or were incompletely tabulated. perhaps even more important is the potential for pooled multivariate analyses opened up by the availability of microdata, used in combination, the nine data sets spanning a century of cataclysmic social and economic change will comprise our most important resource for the study of changing social structure. integration of the public use microdata series despite the enormous potential for time-series analysis of the public use samples, to date only a small proportion of the research based on these data has fully exploited the potential for the study of change over time. many investigators are using the samples as isolated crosssections. a preliminary bibliography of recent research using the public use samples compiled by the social history data archives at the university of minnesota reveals that 178 of 220 studies use only one of the eight public use samples currendy available. it is difficult to use more than one of the public use samples at a time because each sample has a different format, different coding schemes, and different documentation. six separate research teams have been involved in the creation of the samples, and each of them has had their own ideas on how to organize the data. we are faced with eight different occupational classifications with a total of 3200 different categories, and seven incompatible classifications for variables such as birthplace, household relationship, and institution type. in fact, the only variable that is readily comparable across census years is age, and even there the samples differ widely in treatment of missing, illegible, and inconsistent data and in the coding strategy for the very old. documentation for the eight existing samples is contained in eight separate volumes totaling about 3000 pages. these volumes are for the most part organized differently from one another, and their treatment of comparability issues is often cursory. only for the 1950 and 1970 public use samples— where the record layout and coding schemes were made to be reasonably compatible— has there been substantial multi-sample research. indeed, most of the research using more than one public use sample has focussed on these two census years. this suggests that the incompatibilities of the other samples have been a significant barrier to research on long term social change. the incompatibility of the public use samples in their present form means that multi-sample studies require a large initial investment to prepare the data for use. the number of investigators using multiple public use samples is growing rapidly. most have proceeded by creating a set of special-purpose semi-compatible extracts containing a limited number of variables and minimal documentation. this ad hoc approach has already led to increasing duplication of effort moreover, given the complexity of the files and the often subtle differences among them, the potential for eiror is large. the social history research laboratory plans to convert the public use samples for 1880, 1900, 1910, 1940, 1950, 1960, 1970, 1980, and 1990 into a single consistent format and to prepare an integrated set of documentation oriented to the use of the samples as a series. in the long run, we anticipate adding data for all the remaining census years for which individual-level census enumerations survive; these years are 1850, 1860, 1870, 1920 and 1930. we are cunently applying for funding to create a sample for 1920, and plan future applications for the 1850, 1860, 1870 and 1930 census years. we already have had extensive experience with the entire series of public use samples. indeed, the creation of common-format extracts of the samples has been a major preoccupation of the social history research laboratory summer 1991 table 1 summary of availability of selected variables: public use samples, 1880-1990 blank = variable not available n = neighborhood samples, 1970 pus y = variable available st == state samples . 1970 pus c = can be constructed sm = smsa samp es, 1970 pus s = sample-line individuals, 1940 and 1950 a = a" sample, 1980 pums 5 = five-percent sample onl) , 1970 pus b = 'b" sample, 1980 pums 15 = fifteen-percent sample only, 1970 pus c = c" sample, 1980 pums 1880 1900 1910 194c 1950 1960 1970 1980 1990 geographic information state y y y y y. y n,st y y(l) urban/rural residence y y y y n,st c y farm identifier(2) c y y y y y y y y large cities y y y y y sm a,b y modified sma c c c sma y y smsa sm a,b y county or county group y y y y y sm a,b y personal characteristics age y y y y y y y y y sex y y y y y y y y y race y y y y y y y y y marital status(3) y y y y y y y y y household relationahip y y y y y y y y y duration of current marriage y y s(4) age at first marriage s y 5 y number of marrbages y s(5) s(5) y 5 y married in past year? y y y c,s c,s c c c children ever bom y y s s y y y y children surviving y y surname code y y y y subfamily relationships y c c y y y y y y secondary fam. relationships y c c y y y ethnicity and migration birthplace (country, stale) y y y y y y y y y citizenship/naturalization y y y y 5 y y parental birthplace (country) y y y s s y 15 parental birthplace (stale) y y y s s residence five years ago y y 15 y y year of immigration y y 5 y y mother tongue y s s y 15 y y speaks engush? y y y y spanish surname y y y y y y y lassist quarterly table 1 (oontinued) 1880 1900 1910 1940 1950 1960 1970 1980 1990 economic status and employment wage and aalary income y y y y y y total income y y y y y occupation y y y y y y y y y industry c c y y y y y y y home ownership y y y y y y y mortgaged? y y y y rent/home value y y y y y class of worker y y y y y y y period worked in census year y s y y y y hours worked last week y y y y y y period unemployed y y y y s y year last worked y y y y currently unemployed y y y y y y y education and veteran status school enrollment y y y y s y y y y can read y y y can write y y y years of schooling y s y y y y veteran stauis y(5) s s y 15 y y 1. not all geographic information indicated will be avaj lablc for all versions of the 1990 sample. 2. definition of farm varies. 3. the "separated" category of marital atatus is not available before 1950; however, the similar category of married, spouse absent can be constructed for all census years. 4. duration of current marital status. 5. the 1940 and 1950 censuses indicated whether married more than once. 6. civil war veterans only. summer 1991 over the past five years. these files are custom designed to meet the research and teaching needs of minnesota faculty and graduate students. increasingly, we have been receiving requests for common-format extracts from investigators at other institutions. we currently prepare about 25 common-format extracts a month for a broad range of users. in the course of our work, we have become intimately familiar with the intricacies of the public use samples. our staff has invested hundreds of hours in the reconciliation of variables such as occupation and birthplace. it has become obvious, however, that our current procedures— which are duplicated at various institutions across the country — are highly inefficient. what is needed is a complete reworking of all the existing public use samples into an integrated format with complete documentation. this would allow most users to construct their own specialized extracts, and thus dramatically reduce the costs of research. we arc presently in the process of developing a detailed prospectus for the design of such an integrated public use microdata series. it is our hope that prospective users of the data series will provide us with as much feed-back as possible before the design is cast in stone. copies of the prospectus are available upon request. references abel, marjorie and nancy folbre (1990). "a methodology for revising estimates: female market participation in the u.s. before 1940." historical methods 23: 167-178. american economic association (1899). the federal census. new york: macmillan. estimates offertilbty and population in the united states. princeton, nj: princeton university press. conk, margo a. (1980). the united states census and labor force change. ann arbor: uml research press. (1981). "accuracy, efficiency, and bias: the interpretation of women's work in the u.s. census of occupation, 1890-1940." historical methods 14:55-72. edwards, a.m. (1943). "comparative occupation statistics for the united states, 1870 to 1940." sixteenth census of the united states: 1940, population. washington, d.c.: u.s. government printing office. folbre, nancy and marjorie abel (1989). "women's work and women's households: gender bias in the u.s. census." social research 56: 545-569. forstcr, colin and g.s.l tucker (1972). economic opportunity and while american fertility ratios. 18001860. new haven: yale university press. goldin, claudia (1980). "the work and wages of single women, 1870-1920." journal ofeconomic //wtory 40: 81-88. (1983). "the changing economic role of women: a quantitative approach." journal of interdisciplinary history 12: 707-733. x graham, stephen n. (1980). 1900 public use sample: user's handbook. seattle: center for demography and ecology, university of washington. hogan, dennis p., 1981. transitions and social change: the early lives of american men. new york: academic press. anderson, margo j. (1988). the american census: a social history. new haven: yale university press. anderson, michael (1971). family structure in nineteenth century lancashire. cambridge, england: cambridge university press. barron, hal (1984). those who stayed behind: rural society in nineteenth-century new england. new york: cambridge university press. bose, c. (1987). "devaluing women's work: the undercount of women's work in 1900." in c. bose, r. feldberg, and n. sokoloff (eds). hidden aspects of women's work. new york: praeger, 95-1 15. coale, ansley j. and melvin zclnik (1963). new jaffe, aj. (1956). "trends in the participation of women in the working force. " monthly labor review 79(5): 559-565. katz, michael b., michael j. doucet, and mark stem (1982). the social organ ization of early industrial capitalism. cambridge, mass.: harvard university press. mcclelland, peter d. and richard j. zeckhauser (1982). demographic dimensions of the new republic: american interregional migration, vital statistics, and manumissions, 18(x)-1860. new york: cambridge university press. modell, john, frank furstenberg and theodore hershbcrg (1975). "social change and the transition to lassist quarterly adulthood in historical perspective," journal of family history 1:7-32. openheimer, valerie k. (1970). the female labor force in the united states: demographic and economic factors governing its growth and changing composition. population monograph 5, university of california, berkley. rindfuss, ronald r. and james a. sweet (1977). postwar fertility trends and differentials in the united states. new york: academic press. ruggles, steven (1988). "the demography of the unrelated individual, 1900-1950." demography. 25:4. (1991). "comparability of the public use files of the u.s. census of population. social science history. ruggles, steven and mbriam king (forthcoming) fragmentation of the family: living arrangements in america, 1580-1980. ruggles, steven and russell r. menard (1990). "a public use sample of the 1880 census of population." historical methods 23: 104-1 15. (1984a). census of population, 1940: public use sample technical documentation. washington, d.c.: u.s. government printing office. .(1984b). census of population, 1950: public use sample technical documentation. washington, d.c.: u.s. government printing office. _(1989). 1990 census of population and housing. tabulation and publication program. washington, d.c.: u.s. government printing office. yasuba, yasukichi (1963). birth rates for the white population of the united states. baltimore: johns hopkins university ' presented at the lassist 91 conference held in edmonton, alberta, canada. may 14 17, 1991. s. ruggles, department of history, university of minnesota, minneapolis, mn 55455 ruggles@legohead.histumn.edu (512)624-4081 stevens, david (1991). "life-course transitions to adulthood, 1900-1970." journal of family history. strong, michael a., samuel h. preston, ann r. miller, mark hereward, harold r. lentzner, jeffrey r. seaman, henry c. williams (1989). user's guide: public use sample, 1910 census of population. philadelphia: population studies center, university of pennsylvania. u.s. bureau of the census (1954). the 1960 census of population and housing: 1960. two national samples of the population of the united states. washington, d.c.: u.s. government printing office. (1972). public use samples of basic records from the 1970 census description and technical documentation. washington d.c.: u.s. government printing office. (1973). technical documentalbon for the 1960 public use sample. washington, d.c.: u.s. government printing office. (1982). public use samples of basic records from the 19b0 census: description and technical documentation. washington, d.c.: u.s. government printing office. summer 1991 50 iassist quarterly 2010 / 2011 iassist quarterly the oecd (2007) declaration on open access is on the national priority list abstract this report starts from recognition that the archiving of qualitative raw materials that achieved a status of national cultural heritage, has been part of long established and well elaborated slovenian national policy on preserving historical materials of national importance. a well established network of regional and national museums and archives operates a service of preservation, and of access for scientific purposes. on the other hand, despite a rich and flourishing tradition of academic (and even private) qualitative research in slovenia, in most cases all that remains after a project is finished is a research report. still, besides official traditional archives and museums, there is a range of topic-specific public qualitative data resources, e.g. archives of slovenian radio and television, and some emerging infrastructure for preserving the qualitative data originating from social science research projects (etnoinfolab [http://www. etnoinfolab.org/]). qualitative longitudinal resources for re-use were identified with similar problems as are found elsewhere. in conclusion we observed that the problems of qualitative data archiving are not insurmountable, as generally there is willingness, and much valuable material identified. international and national collaboration could help in optimising resources allocation, and an explicit national policy on data archiving would be a prerequisite for future success. keywords: qualitative data, qualitative longitudinal research, oral history, archiving, data sharing, slovenia introduction to report for slovenia the state of qualitative data archiving and sharing in slovenia2 is found to be mixed: on one hand, there is an organised network of well operated national archives and museums under state patronage, systematically preserving what remains of slovenian historical national heritage; on the other hand, there are numerous independent, unrelated, and dispersed research groups and traditions outside of direct state regulation (e.g. university research centres, private research institutes, marketing research agencies) producing a variety of original qualitative data collected for different purposes that lack any form of systematic preservation. after consulting officials and public documents at the slovenian research agency (arrs)3 and ministry of higher education, science and technology (mvzt)4 we have not identified any national policies governing the archiving and sharing of qualitative research data produced by public or private organisations. also, to date, we are not aware of any feasibility studies for qualitative data archiving, outside of the state owned and managed national archives and museums. nevertheless, there is a sign of willingness from the mvzt to regulate the area of public domain scientific data sharing and reuse in general, following the principles of oecd declaration on open access, of which slovenia just recently become a full member. archiving and re-using qualitative and qualitative longitudinal data in slovenia by janez stebe, jože hudales, boris kragelj1 iassist quarterly 2010 / 2011 51 iassist quarterly the archiving of the materials directly important to slovenian national cultural heritage follow a well elaborated national policy on preserving historical materials of national importance. these qualitative data materials are stored by following strictly defined archiving procedures, well established international standards, and clear policy directions for preserving, archiving and sharing. these materials are of potential interest for various aims of qualitative research. they are available to the public under certain legal rules for their access and re-use5. despite the fact that a centrally coordinated and standardised professional documentation (information) system is missing, the network of slovenian museums is currently establishing a national project for the registering of movable cultural heritage, which will finally bring all their collections systematically together, in one place, with a single information service point. slovenian museums are also following the convention for the safeguarding of the intangible cultural heritage (ich 2003): they are systematically documenting oral tradition, traditional crafts and skills, knowledge and praxis connected with natural environment, living heritage, life stories in the form of oral histories, etc. as the mainspring of slovenian cultural diversity. on the other hand, similarly organized initiatives of archiving are widely missing among state-independent producers of qualitative data: e.g. private, non-government, or scientific research organisations that continuously generate original qualitative data as part of their research projects. in this particular research arena any form of regulation of qualitative data preservation, archiving and sharing for the purpose of its reuse is almost non-existent. the culture of qualitative data archiving and sharing, as part of the research culture, is estimated to be rather low, with little awareness of the potential of data sharing and archiving for its reuse. at best, it is in its early stages, with growing awareness of the richness and value of qualitative data but missing action for its systematic management. when one of researchers (fsd) contacted for the purpose of this report was asked about concrete experiences of re-use of original material from the qualitative studies, only few could be imagined. these few cases are based on the involvement of a researcher who is willing to make materials accessible from his own past research projects, and make the data available to a new research team. another example mentioned (memo) was about partner research agencies sharing their already available qualitative resources for their common projects. more typical were negative experiences, such as when a researcher would be willing to share the original material from past research projects with colleagues, but it proved to be impossible simply because the data were not properly preserved. as a general practice, the extent of qualitative data preservation here depends on researcher’s own standards rather than any established organisational or institutional guidelines. the most common way of preserving qualitative data, if done at all, is by keeping the raw materials in the form in which they were originally collected, or in the form in which the analysis was concluded (fsd, fdv (faculty of social sciences), memo, aragon). even in the highest profile research studies, usually there is no systematic or institutionally regulated data preservation policy: raw materials are dispersed between offices and homes without specially designated spaces for their storage. such non-systematic preservation means that qualitative data collections often lack methodological and contextual details (study and data documentation, metadata) of the process of data collection that are much needed for judging the reliability and validity of the original data, and the possibility of its reconstruction for the purpose of re-use6. accounting also for the variety of mixed data formats as one of the most commonly mentioned problems of qualitative data, ranging from handwritten notes, audio and video files, to word processor transcripts, which are not consistent either between or within a single qualitative study, it is often felt that reconstruction of the original research framework and its organisation for the purpose of re-analysis is too challenging a task (fsd). in sum, an overriding estimate across the research community indicates that original qualitative research materials are very hard to reach in a form suitable for reuse, either because they were not properly preserved, or because relevant metadata information is missing. agreement from funder agencies as research commissioners, and ethical concerns regarding the privacy of participants, are often mentioned issues that additionally prevent further data exploitation (fsd, fdv, memo, aragon). considering the fact that secondary data are rarely completely appropriate to address a new research problem, starting a new process of qualitative data collection often still seems to be the option preferred over considering the possibilities for reuse of existing qualitative data resources. however, the prospects are becoming more optimistic: the oecd (2007) declaration on open access is on the national priority list. the aim is to build relevant national policies for the long-term securing of data resources from publically funded research, and to glean examples of good practice from other countries as the basis for organising qualitative data resources in slovenia in the future. existing qualitative archiving infrastructure in slovenia the existing qualitative archiving infrastructure in slovenia can be divided into three distinct segments with respect to the content of its holdings (subjects, methods and data formats covered). these segments range from (a) several centralised, highly regulated and state managed national archives, (b) more or less well organised but devolved topic-specific qualitative data collections to be found in particular public archives and museums, and (c) loosely managed archiving facilities preserving original qualitative data production, collected for the purpose of genuine social science research projects. a. state (national) archives a range of archives run under state control with management and funding provided directly by the government of the republic of slovenia7. they all share the same competences and fall under the same jurisdiction. they are managed through one central national archive of the republic of slovenia and several additional regional archives that cover their particular geographical areas. they hold various original resources about the life of institutions, state, society and individuals in slovenia preserved in the form of written, drawn, printed, photographed or otherwise documented archive materials that are estimated to be of permanent importance for slovenian national cultural and historical heritage: historical documents on parchment, paper, film, and magnetic or optical data holdings, witnessing important information about the nature, objects, places, phenomena and people relevant to slovenian national history and contemporary life (e.g., certificates, birth and death records, census data, cadastral registers and maps, and other personal data collections). these qualitative data materials are preserved and used for national research purposes and historical evidence, as well as for administrative and jurisdiction purposes. recently there has been a strong initiative to convert them into electronic format and provide them online8. there is also a strong emerging initiative to make these materials available to 52 iassist quarterly 2010 / 2011 iassist quarterly the public in digital form through an online service: some parts of the archive are already served electronically9. all materials of these archives are carefully selected and professionally preserved following well-established (international) rules and procedures that are shared by all state archives10. some of these rules and procedures are (in a limited way) also relevant to other (non-state, specialist-focused) archives and can serve them by providing a possible model of good practice (see appendix 2, table 1). b. other topic-specific public qualitative data archive resources among other topic-specific public qualitative data archive resources we list particular data collections (see appendix 2; table 2) which do not fall directly under the state archiving policy regulations, although the government of the republic of slovenia still provides the main source of their funding11. they are more or less autonomous and are independently managed by scientific, education or information institutes such as institutes or museums that run under special government licence. these independent public qualitative data resources share distinct competencies and are preserving material in their specific area of activities, covering substantive areas of interest. they follow their own idiosyncratic12 archiving rules based on their specialised subjects. each of them holds their own original subject-specific evidence, in subject-specific formats, for subject-specific use (primarily social and biographical historians, or the general public at occasional exhibitions). for example, the archives of slovenian radio and television (rtvslo) can be seen as a very rich and relevant source of qualitative data. this archive does not fall under the same policy rules as the national archives, but is otherwise systematically well-organised, preserving video and audio materials of all the tv production broadcast on slovenian national radio and television. the accessibility of these multimedia materials is not under any strict rules and much of the national radio and television broadcast is directly accessible through their internet multimedia archive centre, while an advanced search engine within the video content is under development but can be tested online13. among relevant organized initiatives of archiving of national heritage on the state side three other projects of digitization of national cultural heritage supported by the ministry of culture must be mentioned14: dlib – digital library of slovenia is part of the national and university library (nuk), kamra – digitalised cultural heritage of slovenian regions and sistory – digitization of slovenian historical literature and historical sources as part of slovenian cultural heritage. sistory is managed by the institute for contemporary history in ljubljana and takes the form of a database available for the whole research community, with historians as the primary users. it also takes part in dariah – digital research infrastructure for the arts and humanities which tries to provide a coordinated infrastructure for supporting preservation of cultural heritage in europe (lazarević and vodopivec, 2007). dlib is probably also the place for a future internet archive. c. archiving infrastructure for preserving the production of qualitative data originating from social science research projects this section explores qualitative data archiving facilities that are especially devoted to preservation of qualitative data from social science research projects. these do not cover the artefacts relevant for slovenian national heritage, but rather all independently produced qualitative data sources emerging from public, academic or scientific research initiatives. first attempts of practice in archiving special, research-oriented qualitative data collections can be noted at the department for ethnology and cultural anthropology at the university of ljubljana by a project, etnoinfolab, which strives to organise and centralise the department’s documentation service system (including all qualitative data collected by the students during their study or by researchers working at the institute) and provide the available qualitative data sources for its staff and interested researchers as well as the wider public, over a single computer system together with applications for data storage and re-use. these are small in scale and are managed and financed autonomously by the research institutes alone, following their own particular and idiosyncratic qualitative data archiving rules. they hold rather limited numbers of qualitative data collections as they are still in the process of developing the proper qualitative data archiving and sharing infrastructure: mainly they contain ethnographic research materials (transcripts of interviews, copies of photographs) and limited amounts of other diverse data formats resulting from qualitative social science research. re-use of these data is in general limited to a closed circle of academic researchers (most commonly used by students and professors for teaching). social science data archives (adp arhiv družboslovnih podatkov)15, as the main quantitative data repository in slovenia, endeavours to extend its services more intensively into the qualitative domain: it is already holding data from a few qualitative studies and plans to broaden its activities in this direction. the analogous processes for quantitative data are used for qualitative data. digitised primary research material of a qualitative study, at this time limited to textual data, is preserved as is raw data in quantitative studies. studies are catalogued using the data documentation initiative metadata standard (ddi)16 for study description. terminology used is method specific, e.g. when describing kind of data and research instrument used. adp is subsidized by arrs long-term research infrastructure grant. its mission is to be responsive to the demands of the general social science research community both nationally and internationally. it has been under development for more than ten years, and has thus acquired extensive professional knowledge on data archiving rules upon which to build its services. it offers data free for scientific and educational purposes and uses internationally recognised standards and procedures in acquiring, preserving and disseminating the data. it is also a member of international organisations and partners in their projects (council of european social science data archives (cessda), international federation of data organizations (ifdo)). these attempts have the potential for establishing the foundation for a genuine qualitative data infrastructure upon which to build in the future. while only these two facilities are already active in providing qualitative data sources for reuse (either in electronic or hard copy), there are also several other existing public (qualitative research oriented) organisations generating an important amount of qualitative data without any form of systematic preservation. these are listed in appendix 2 as a relevant source of qualitative data to be potentially integrated into a proper qualitative data archiving infrastructure in the future. d. private archives or facilities for preservation of qualitative data of private research agencies the remaining category in our classification is another important source of qualitative data production: private research organisations iassist quarterly 2010 / 2011 53 iassist quarterly such as marketing research agencies, advertising companies, or public relations offices. they produce a mass of ad-hoc qualitative studies generating a large quantity of data without any real consideration for its integral, systematic and organised preservation which would enable its later re-use. despite the existing legal framework for its establishment, we have not found any evidence of private venture activities for qualitative data archiving17 that would enable data to be shared more widely. the main reasons for this are always given as personal privacy protection laws and the problem of property rights of the data, which belong to private companies that are not willing to share what is often perceived as confidential business data (argon, memo), even if examples exist of quantitative data from private sector, stored in adp. some private research agencies (aragon, for example, from 2008 onwards) have started to build their own private data archives, systematically preserving all the data originating from quantitative as well as qualitative studies conducted by them. however, these data sources in any case remain unavailable for further exploitation by external users. the majority of qualitative data sources originating from the private sector, which are rich in content, diverse in formats and large in number (but of variable quality) now remain dispersed without any central node for their coordination, preservation and potential reuse. qualitative longitudinal data in slovenia by their nature, the majority of state managed and specialist archives of qualitative data that are concerned with the preservation of slovenian culture are by definition longitudinal. this type of qualitative data indirectly addresses social temporality in terms of bearing witness to its time. for example, since 2003, a museum of contemporary history (kokalj kočevar, 2005) is systematically collecting oral life stories and memories (in video and audio formats) about 20th century slovenian history (the project is mainly aimed at capturing the period of the first and second world wars). some of these materials are already accessible to the public, in digital format and directly over the internet, for example mojazgodba, www.radio.ognjisce.si and www.ushmm.org/research/ collections/oralhistory/search. as another example, the sistory project18 was launched at the institute for contemporary history to provide a coordinated infrastructure for supporting preservation and sharing of national cultural heritage (lazarević and vodopivec, 2007). relevant national historical documents are becoming widely available from a single common access point for the research community at large. the information system allows for browsing or searching for scientific papers, reports and discussions published in slovenian historical publications (printed and electronic) on various periods of slovenian history. there is also open access to some of the historical databases containing visual materials and maps. a similar recent project that issued a public call for lay memories of tito’s funeral promises to archive the material collected19. on the other hand, these are not conceptually designed, longitudinal data that would have resulted from problem-oriented, qualitative research studies directly addressing time, temporality, or prospectively tracking changes over time. in the area of originally produced data by social science research from the autonomous social science or private research organisations described above, we are aware of two longitudinal infrastructures for the management and re-use of longitudinal data. these are the already mentioned social science data archives and etnoinfolab, the first holds mainly quantitative data sources, some of which are longitudinal. adp holds numerous repeated cross-section studies (time use studies, media use studies, social values studies), some of which are in the form of internationally harmonised, continuous longitudinal datasets20. occasionally there are minor follow-ups on the same sample21. finally there are also limited panel or continuous cross-section data available22. however, adp does not hold data from any kind of qualitative longitudinal study. something close to qualitative longitudinal data can be found in the archiving infrastructure of the etnoinfolab documentation service, provided by the department of ethnology and social anthropology at the faculty of arts – ul. etnoinfolab (ethnographic information laboratory) started in 2006 as the project for systematically regulated digitization of extant ethnographic data collections and their integration into a single common database, mainly comprising interviews (voice transcriptions in text format) with accompanying photos. the collections cover topics that range from material, social and spiritual culture in slovenia as reflected in the everyday life of individuals, and encompassing the time period from the 19th century until the present day. the most important collections in etnoinfolab were created by: • ivan benigar, a south american anthropologist working among the mapuce indians in patagonia, argentina; • joel m. halpern, a prominent american cultural anthropologist who has conducted community studies in two slovenian villages; • vekoslav kremenšek (1930) and vilko novak (1910-2002), who have conducted locality and economic migration studies in slovenia. further details of these important qualitative longitudinal resources are listed in the appendix and on the web23. all of these qualitative longitudinal data collections are already available via etnoinfolab documentation and information service. since august 2008 all the collections are listed on the internet in detail, and many digitised images of original documents, field notes and photographic material are freely available for public use24 . in addition, etnoinfolab and its electronic information service provide students with practical knowledge and skills to use the necessary computer applications, as well as processing techniques for data, sound, and film. outside of the etnoinfolab only occasional individual research projects involving qualitative longitudinal data were identified such as a small scale student research project collecting family genealogy histories is currently ongoing at the faculty of social sciences as a part of the study process, but its results are still uncertain25 . such dispersed individual research projects based on the collection of longitudinal qualitative data are obviously not yet a part of any of the archiving infrastructures mentioned above, but should in any case be kept in mind for future development of this particular area. development and planning two major problems in qualitative data resources in slovenia have been identified: a general lack of qualitative longitudinal resources and missing infrastructure for archiving the existing qualitative data resources originating from independent research projects conducted by public and private organisations. the main reasons for such a situation can be found in the rather low value attached to qualitative research data in slovenia in general, and a low level of awareness of the potential of data archiving and sharing that prevails within the research community. also, in general there is a very low level of 54 iassist quarterly 2010 / 2011 iassist quarterly knowledge about what qualitative data is available and what can be counted as data archiving infrastructure facilities. identified problems and gaps: an important reason for the limited systematic organisation of qualitative data resources are, above all, ethical concerns, as is often emphasised by data producers (fsd, aragon, memo, and fdv). the consent forms used for participation in qualitative research normally allow use of primary research materials for the sole purpose of generating reports for the particular research case, and prohibiting exposure of personal information that may identify individual participants. it is believed (wrongly) that these conditions automatically preclude making the original research materials available to others: strong doubts persist that even with very sophisticated qualitative data manipulation, anonymity could not be achieved for the purpose of privacy protection for the original research material that might be shared. this is one area where established institutions could provide best practice guides and training for researchers to start thinking about data archiving early in project conceptualisation in order to remove any ethical or legal obstacles for future re-use26 . another reason for resistance to sharing may be found in a tendency toward monopolisation of original research material for the advantage of the primary researcher or agency, which is again in tension with a data sharing. in the case of private (marketing) research companies, which are likely to represent one of the largest resources of qualitative data, the research material is legally owned by the company paying for the research, which makes data sharing even more complicated. agencies conducting this research have little motivation to resolve such complications to enable sharing. an additional problem facing a stronger qualitative data sharing initiative is anchored in the lack of knowledge and professional sound practice among active researchers of how to deal with the task of organising, managing, archiving and sharing data. despite the availability of data, procedures for archiving are not wellspecified and therefore challenging to fulfil. a main obstacle for qualitative data archiving is often simple lack of human resources in the research organisations where there is often very little administrative support for research activities, such as, for example, data preservation after the completion of a project. researchers working alone or in small teams do not tend to engage in data preservation activities as they lack the motivation and often the time for such work. archiving is unfortunately not perceived as an activity that has the same level of academic prestige and relevance to professional identity as conducting primary research. there is currently no clear professional track of rewarding archiving activities, and people with the capacity to pursue academic careers don’t take enough interest in it. there are also institutional barriers to the development of more dedicated, specialised social science qualitative data archiving infrastructures. for example, existing infrastructures and networks around qualitative research that are centralised at the university of ljubljana (by far the largest academic institution in the country) were not perceived to be effective enough for taking over the task of archiving. in discussions, it was mentioned that there is a need for stronger coordination and networking of data infrastructures in an interdisciplinary fashion, but currently there is no real interest in such an initiative. adp was often considered as an institution that could be asked for support by data producers in the past (fsd), but no further steps were taken in the direction of establishing any formal contract. considering the state of archiving and sharing of qualitative data resources in the future there is a need to work on the following development priorities: promotion of the idea of qualitative longitudinal data collection and raising the awareness of the value and benefits of qualitative data archiving and sharing. one possible step in this direction is adp’s organisation of seminars setting out the important aspects of data archiving as part of the research process, and providing an overview of guidelines for efficient data archiving. for this to take place the collaboration of research organisations is required! establishment of a formal appraisal and selection policy that would reflect the future re-use potential of qualitative research material compared to the cost of production, preservation and access. a detailed methodological description of the research process itself should be always undertaken to allow for judgement of data quality for this purpose. it can be estimated that only a limited number of the most valuable qualitative studies (up to 5 per year) would qualify for archiving in adp. those remaining could be deposited in a general public research digital repository such as dlib. preparing a package of procedures for consistent qualitative data archiving, specially dedicated to preserving original materials from autonomous qualitative research projects, taking into account ethical considerations and international standards. this will be available on demand for research organisations to help them organise and preserve their future qualitative research projects throughout the data life-cycle. iassist (international association for social science information services & technology) and cessda assistance is required for advising on qualitative data archiving regulations, guidelines and standards27. trying to develop a motivation scheme as well as formal requirements to be set out for research organisations and individual researchers to systematically report and organise their qualitative data materials for the possibility of being included in the larger infrastructure for qualitative data archiving. for example, this would include introducing additional scientific criteria that would properly value and reward archiving activities in terms of academic careers. this would require further changes in established institutional rules. acquiring specialisation of human resources in qualitative data archiving within existing adp. this would require devoting a post for qualitative data resource management, as a point of reference and advice, as well as practical support and management of existing qualitative data from various sources. additional funding would be required. taking a first step towards coordination of dispersed producers and sources of qualitative data by moving towards establishing a common catalogue of all the existing qualitative data materials publicly available. this should slowly lead to general rationalisation of national qualitative archiving activities and step-by-step establishment of a general single point of access to all qualitative data resources. a set of common standards and rules should be agreed upon to achieve inter-operability and professional soundness of the whole endeavour. iassist quarterly 2010 / 2011 55 iassist quarterly conclusion: how might existing organisations be of assistance? support for professionalization and training in the field of qualitative data archiving is needed. the usefulness of esds qualidata training and support materials is already recognised as a baseline for international integration of activities in the realm of qualitative data archiving and sharing. establishing a stronger organisational foundation within cessda could strengthen support for members and spread good practice, e.g. in how to fill in gaps in collections in particular countries, by enabling visits and mentoring facilities in development of new services, and advice and training for researchers and archivists. in particular, the support from the international organisations would be important in the following areas: internationally coordinated fund raising and country research policy formulation activities, e.g. inclusion of cessda and dariah projects in the future esfri (the european strategy forum on research infrastructures ) funding scheme. development of professional profiles and establishment of training schemes to build expertise for qualitative archiving, data sharing and its re-use in terms of secondary analysis. synchronisation of common activities for development of needed tools and processes that would support qualitative archiving and sharing. development of common standards and training in their use (e.g. a ddi “lite” template suited for qualitative data). establishment of harmonised rules based on established protocols and preservation policies that effectively guarantee the long-term access to qualitative data resources. extension of the role of adp is often mentioned as one of the best solutions for preservation and sharing of the valuable qualitative data produced by these research centres (fdv, fsd). adp is the only science data archiving infrastructure in the country, with a well-established repository of data mainly from large-scale social surveys. as such it has a solid foundation from which to broaden its service to qualitative data archiving in the future. however, its existing data deposit system and data sharing regulations should be properly extended for qualitative data archiving28. in this way adp could serve as a possible node, with experiences and existing resources in the area of quantitative data, which could be extended or adapted for the special needs of qualitative data, and for building a centrally integrated mixed-methods social science data archive. an alternative for preservation of these dispersed qualitative data sources is their integration into existing the etnoinfolab service of ethnographic documentation system. each solution has its benefits. in the first case, a single access point to mixed types of data in the form of a common and centrally coordinated management system for archiving and sharing qualitative and quantitative data that can be reused for any purpose can be seen as an advantage. in the second case the existing qualitative data might be simpler to integrate into already established information systems, thus avoiding the problem of bridging quantitative and qualitative data storage standards. this may result in easier implementation as well as faster service provision to the end users. closer familiarity with the substance and format of data could be an advantage for this latter option. in any case these do not need to be mutually exclusive alternatives. cross-overs between the systems, harmonisation and co-ordination of activities would serve best the common interests. as it is not expected to anticipate an abundance of human and financial resources, all those willing to spend their energy in establishing an expanded slovene qualidata service would be welcomed. references kokalj kočevar, m. (2005). ‘time in memory: lived experience; what is remembered and what is forgotten’ in: adding of memories of ww2, using the war: changing memories of ww2. annual conference of the ohs: london lazarević, ž. and vodopivec, n. (2007). (inštitut za novejšo zgodovino/ institute for recent history, ljubljana), 2007, beograd, sistory: digitalizacija istorijskih sadržaja (sistory: digitization of historical sources), преглед нцд 11, 31–36. oecd. (2007)). ‘oecd principles and guidelines for access to research data from public funding’ [online]. available at: http://www.oecd.org/dataoecd/9/61/38500813.pdf [accessed 12th march 2010] žumer, v. (2001). arhivi v rs – zakladnica virov za rodoslovna raziskovanja. drevesa : bilten slovenskih rodoslovcev issn: 1318-6221.vol. 8, no. 3 (2001), str. 29-35. [online]available at: http://www2.arnes. si/~krsrd1/conference/speeches/zumer_ars_rodoslovje.htm appendix 1: list of other individuals, organisations and their abbreviations for the purpose of the report • etnoinfolab hudelja, mihaela, documentalist etnoinfolab oddelek za etnologijo in kulturno antropologijo ff (department for ethnology and cultural anthropology, faculty of arts, university of ljubljana) e-mail: mihaela.hudelja@ff.uni-lj.si internet: http://etnologija.etnoinfolab.org/en/default.asp • mnz-si kokalj kočevar monika, ma, museum adviser muzej novejše zgodovine slovenije / national museum of contemporary history (mnz-si) e-mail: monika@muzej-nz.si interent: http://www.muzej-nz.si/ • fsd rihter liljana, phd, senior lecturer fakulteta za socalno delo / faculty of social work, university of ljubljana (fsd) e-mail: liljana.rihter@fsd.uni-lj.si internet: http://www.fsd.si/faculty_and_staff/2008050813072178/ • memo perčič eva, ma, research director, memo memo institute creative research d.o.o. e-mail: eva.percic@memo.si internet: http://www.memo.si/index.php • aragon prešeren jana aragon, d.o.o., research and planning e-mail: jana.preseren@aragon.si internet: http://www.aragon.si/eng/ 56 iassist quarterly 2010 / 2011 iassist quarterly • fdv tivadar blanka, phd, and kamin tanja, phd center for research on social psychology, institute of social sciences, faculty of social sciences, university of ljubljana (fdv) e-mail: tanja.kamin@fdv.uni-lj.si; blanka.tivadar@guest.arnes.si internet: http://www.fdv.uni-lj.si/english/research/research_c. asp?id=11 http://csp.fdv.si/ adam frane, phd faculty of social sciences, university of ljubljana e-mail: frane.adam@fdv.uni-lj.si internet: http://www.fdv.uni-lj.si/kontakti/osebne.asp?id=1 appendix 2 (žumer, 2001) table 1: list of slovenian state archives holding remains of national heritage arhiv republike slovenije (archives of the republic of slovenia) address: zvezdarska 1, 1127 ljubljana, p.p. 21 phone: (01) 24 14 200, (01) 24 14 250 fax: (01) 24 14 269, e-mail: ars@gov.si internet: http://arhiv.gov.si/ note: slovene film archives keeps slovene documentary films, cartoons and feature films from 1905 (when the oldest slovene film was made) onwards. more than 90 % of all slovene films or 5,100 titles are preserved. zgodovinski arhiv ljubljana (ljubljana history archive) address: mestni trg 27, 1000 ljubljana phone: (01) 30 61 306 fax: (01)42 64 303 e-mail: zal@zal-lj.si internet: http://www.zal-lj.si pokrajinski arhiv koper (koper province archive) address: goriška 6, 6000 koper phone: (05) 62 71 824 fax: (05) 62 72 441 e-mail: arhiv.koper@guest.arnes.si internet: http://www.arhiv-koper.si/ pokrajinski arhiv maribor (maribor province archive) address: goriška 6, 6000 koper phone: (05) 62 71 824 fax: (05) 62 72 441 e-mail: slavica.tovsak@pokarh-mb.si internet: http://www.pokarh-mb.si/ pokrajinski arhiv nova gorica (nova gorica province archive), address: trg e. kardelja 3, 5000 nova gorica phone: (05) 30 27 737 fax: (05) 30 27 73 e-mail: pang@guest.neticom.si internet: http://www.pa-ng.si zgodovinski arhiv ptuj (ptuj history archive) address: muzejski trg 1, 2250 ptuj phone: (02) 78 79 730 fax: (02) 78 79 740 e-mail: zgod.arhiv-ptuj@guest.arnes.si internet: http://www.arhiv-ptuj.si/ table 2: topic-specific public infrastructure of qualitative data resources zgodovinski arhiv in muzej univerze v ljubljani (history archive and museum of university of ljubljana) address: kongresni trg 12, 1000 ljubljana phone: (01) 42 54 055 fax: (01) 42 54 053 e-mail: info@uni-lj.si internet: http://www.uni-lj.si/o_univerzi_v_ljubljani/univerzitetni_arhiv/ amsu.aspx note: documentation about the work, life and function of university of ljubljana dokumentacija rtv slovenija (documentations of slovenian national radio and television) address: kolodvorska 2, 1000 ljubljana phone: (01) 475 36 16 fax: e-mail: tvdokumentacija@rtvslo.si internet: http://www.rtvslo.si/modload. php?&c_mod=static&c_menu=1053436918 note: audio and video material of slovenian national television production in electronic form narodna in univerzitetna knjižnica rokopisni oddelek (national and university library – manuscript division) address: turjaška 1, 1000 ljubljana, phone: (01) 200 11 10 fax: (01) 4257 293 e-mail: uprava@nuk.uni-lj.si internet: http://www.kud-logos.si/rokopisi/rokopisni-anglesko.htm note: central national and state-owned collection of original manuscript material from the fields of literature, linguistics and broader humanities dating back to 1774 (available on microfilm, paper, photographic and digital copies); narodna in univerzitetna knjižnica digitalna knjižnica slovenije (national and university library – the digital library of slovenia) address: turjaška 1, 1000 ljubljana, phone: (01) 200 11 67 fax: (01) 4257 293 e-mail: dlib.si@nuk.uni-lj.si internet: http://www.dlib.si/ note: a web portal providing ready access to slovenian knowledge and cultural treasures with offering free searching of text (books, periodicals), visual (photos, maps) and sound resources with respect to slovenian cultural heritage; new project under development: building a complete archives of slovenian (electronic) documents on the world wide web. iassist quarterly 2010 / 2011 57 iassist quarterly arhivi katoliške cerkve (archives of catholic church) address: krekov trg 1, 1000 ljubljana phone: (01) 43 37 044 fax: (01) 43 96 435 e-mail: arhiv.lj@rkc.si internet: http://lj.rkc.si/?id=11&fmod=2 note: valuable historical collection of books with birth recordings; important for genealogy research muzej novejše zgodovine slovenije (national museum of contemporary history) celovška 23, 1000 ljubljana tel. 00386 1 3009637 fax. 00386 1 4338244 e-mail: info@muzej-nz.si internet: www.ushmm.org/research/collections/oralhistory/search note: huge collections of life stories and memories (in video and audio format) on 20th century history (focus on first and second world war). muzej novejše zgodovine celje (museum of recent history in celje) prešernova ulica 17, 3000 celje tel.: +386 3 428 64 10 fax: + 386 3 428 64 11 e-mail: mnzc(at)guest.arnes.si internet: www.muzej-nz-ce.si note: huge collection of audio and video documentation about craft and craftsmanship in celje, their life stories, etc. inštitut za novejšo zgodovino (institute for recent history) kongresni trg 1,1000 ljubljana phone: +386 1 200 31 20 fax: +386 1 200 31 60 e-mail: info@inz.si internet: www.sistory.si/ note: large and rich internet based electronic service point of historical sources in slovenia; total access to print and archive sources, databases, visual material and maps. table 3: archives of originally produced qualitative data originating from autonomous social science research projects etnoinfolab oddelek za etnologijo in kulturno antropologijo ff (department for ethnology and cultural anthropology, faculty of art) address: zavetiška 5, 1000 ljubljana phone: 00386 1 2411 520 internet: http://www.etnoinfolab.org/ note: valuable and rich ethnographic data collections (interviews, photos) collected by the members and students of the department during their research and pedagogic activity dating back to 1956. national territory of slovenia, austria, italy, and hungary, as well as the republics of the former yugoslavia and some other parts of balkan are covered. adp arhiv družboslovnih podatkov (social science data archives, faculty of social sciences, university of ljubljana) address: kardeljeva ploščad 5, 1000 ljubljana phone: (01) 5805 292 fax: (01) 5805 294 e-mail: arhiv.podatkov@fdv.uni-lj.si internet: http://www.adp.fdv.uni-lj.si/ note: preserving a few (around 10) subject specific qualitative studies that are not survey or statistical data originating from autonomous research projects table 4: important research centres as producers of qualitative data without the infrastructure for its preservation institututum studiorum humanitatis – faculty for postgraduate studies in humanities address: slovenska cesta 30a, 1000 ljubljana phone: + 386 1 425 18 45 fax: + 386 1 425 18 46 e-mail: ish@ish.si internet: http://www.ish.si/en/ scientific research centre of the slovenian academy of sciences and arts address: slovenska cesta 30a, 1000 ljubljana phone: 386 1 470-6-100 fax: 386 1 425-52-53 e-mail: zrc@zrc-sazu.si internet: http://odmev.zrc-sazu.si/zrc/ peace institute – institute for contemporary social and political studies address: metelkova 6, 1000 ljubljana phone: + 386 1 234 77 20 fax: + 386 1 234 77 21 e-mail: info@mirovni-institut.si internet: http://www.mirovni-institut.si/main/index/en/ university of ljubljana – faculty of social work (delinquency, youth, and mental health) address: topniška 31, 1000 ljubljana phone: +386 1 280 9240 fax: +386 1 2809 270 e-mail: info@fsd.uni-lj.si internet: http://www.fsd.si/eng/ university of ljubljana – institute of social sciences address: kardeljeva ploščad 5, 1000 ljubljana phone: + 386 01 5805 200 fax: + 386 01 5805 213 e-mail: fdv.idv(at)fdv.uni-lj.si see in particular (internet): center for research on social psychology (contemporary ethnology of everyday life) [http://csp.fdv.si/; http://www.fdv.uni-lj.si/english/ research/research_c.asp?id=11] social communication research centre (media content studies) [http://www.fdv.uni-lj.si/english/research/research_c.asp?id=7] centre for spatial sociology (remains of large scale qualitative studies conduced by prof. zdravko mlinar, now retired) [http://www.fdv.unilj.si/english/research/research_c.asp?id=14] 58 iassist quarterly 2010 / 2011 iassist quarterly centre for methodology and informatics (meta analysis of qualitative studies) [http://english.fdvinfo.net/index.php?fl=0&p1=300&p2=301 &p3=&id=103] appendix 3: qualitative longitudinal resources in slovenia important cultural historical studies held in etnoinfolab • ivan (juan) benigar (1883-1950): a famous slovene anthropologist who lived in south america among indians mapuče in patagony (argentina) for more than 40 years and became one of the best known anthropologists in south america. the collections of his work consist of digital copies of hundreds of pages of all of his field notes, excerpts and different kinds of manuscripts that relate to his research work. • joel m. halperen (1929): halpern, a prominent american cultural anthropologist who has worked as a field researcher in slovenia in the period of 1962/63, has gathered field data covering two slovenian villages (šenčur and gradenc) that consists of field notes (more than 2000 pages of typescript) and more than 1000 photos and slides of everyday social life. • vekoslav kremenšek (1930) and vilko novak (1910-2002), in doing extensive research projects with their students, have gathered a huge collections of data covering the following area and time period: • collection etseo / ethnological topography of slovene ethnic territory. 1975-90 • collection galjevica (urban suburb of ljubljana). 1970-75 • collection vitanje. 1975-80 • collection izseljenstvo (economic emigration from slovenia) notes 1.contact details of country reporters: the country report on qualitative data resources for slovenia was written by the following team of experts, each of them covering a particular area of the field under investigation. janez štebe area: national policies, data archiving and sharing procedures, quantitative and qualitative research resources, development policies, iassist and cessda assistance university of ljubljana social science data archives [http://www.adp. fdv.uni-lj.si] kardeljeva ploscad 5, 1000 ljubljana [janez.stebe@fdv.uni-lj.si] jože hudales area: ethnographic qualitative data resources, qualitative data in museums, qualitative longitudinal data university of ljubljana – faculty of arts, department for ethnology and social anthropology [http://etnologija.etnoinfolab.org/en/default.asp] askerceva 2, 1000 ljubljana [joze.hudales@ff.uni-lj.si] boris kragelj area: state archives with national cultural heritage, private and public qualitative resources university of ljubljana – faculty of social sciences [http://www.fdv.uni-lj.si/english/office_ic/] kardeljeva ploscad 5, 1000 ljubljana [boris.kragelj@fdv.uni-lj.si] 2. the report reflects the particular institutional background of the reporters: it mainly builds upon the circumstances of (predominantly quantitative) data archiving at social science data archives of slovenia (adp), that covers broad domain of social sciences in a country, supplemented with the perspective from a more humanities oriented domain, covered by department of ethnology and social anthropology of the faculty of arts at university of ljubljana (ff). in addition, to survey overall existing qualitative data resources and collect the relevant information about the general situation concerning the data sharing and data archiving culture, a number of research groups and organisations with qualitative research profiles were briefed, either orally or in writing (a complete list of external collaborators involved in the preparation of the report is presented in the appendix 1). among all surveyed qualitative research centres, the following in particular proved to be valuable for the information presented in the final report: faculty of social work (fsd) with its long and well established tradition of qualitative research on delinquency, youth, and mental health; the national archives of the republic of slovenia (ars), a preeminent institution for systematic preservation of historical national heritage concerning the slovenian state, and the national museum of contemporary history (mnz), a state museum, which provides a central public service in the area of the movable heritage of contemporary history, preserving studying and communicating the material and non-material heritage in the sphere of the history of the slovene ethnic space from the beginning of the 20th century. 3. for further information see, http://www.arrs.gov.si/en/ 4. see, http://www.mvzt.gov.si/en/, 5. for more information on the collections, policies, procedures and laws guiding and regulating the archiving and sharing of these materials see national archives of the republic of slovenia – ministry for culture [http://www.arhiv.gov.si/en/]. 6. some descriptive metadata is usually presented in the final research report alone and not prepared as an independent report accompanying raw data (fsd, fdv). 7. see list in appendix 2. 8. for further information see, http://www.arhiv.gov.si/ si/e_hramba_dokumentarnega_gradiva/. 9. see http://arsq.gov.si/query/suchinfo.aspx 10. for further information see http://www.arhiv.gov.si/si/ zakonodaja_standardi_in_dokumenti/. 11. the only obvious exception is archive of the catholic church. 12. see, www.rtvslo.si 13. see http://www.rtvslo.si/odprtikop/ 14. weblinks for other projects of digitization of national cultural heritage supported by the ministry of culture , see http:/www.dlib.si/; http:/www.kamra.si/; http:/www.sistory.si/; http://www.dariah.eu/ 15. see http://www.adp.fdv.uni-lj.si/ 16. see http://www.ddialliance.org/ 17. archives of the catholic church can be seen as the exception 18. see www.sistory.si 19. see http://www.tovaris-tito.si/, accessed on 1.3.2010. 20. for an example of an event history data study, see lol94 http:// www.adp.fdv.uni-lj.si/opisi/lol94/ 21. for example, preand post-election re-interviews: sjm90 http:// www.adp.fdv.uni-lj.si/opisi/sjm90/ 22. for example, slovene labour force survey – http://www.adp.fdv. uni-lj.si/opisi/serija/ads/. 23. for further details see under the heading “etnološke raziskave slovenskih in drugih kultur” at http://www.etnoinfolab.org/ 24. for details see the following internet addresses: [http://www. etnoinfolab.org/; http://etnologija.etnoinfolab.org/sl/informacija. iassist quarterly 2010 / 2011 59 iassist quarterly asp?id_meta_type=75&id_informacija=199]. the parts of data collections which are still undergoing research or publication are only available through a special information system (http://193.2.104.52/ studsistem/) for students, members of department and researchers who must get special permission to log in to the system. 25. for further details contact: prof. anton kramberger; fdv: anton. kramberger@fdv.uni-lj.si 26. see e.g. manage and share data section of esds – economic & social research council. http://www.data-archive.ac.uk/ 27. overcoming a problem of diversity of qualitative research methods resulting in variety of different data forms and their particular formats shows itself as one of the biggest problems for realisation of more systematic preservation of qualitative data. 28. see, http://www.adp.fdv.uni-lj.si/za_uporabnike/ izrocanje_podatkov/ major post censal redesign of household sample surveys in the united states by preston jay waite' u.s. bureau of the census acknowledgements i wish to acknowledge the work of many staff members in the statistical methods division of the u.s. bureau of the census in the preparation of materials for this paper. special thanks to dr. charles alexander and gary shapiro for their contributions to the editing and review and to pat curran for her work in preparing the manuscript. post census surveys (pcs) are utilized in a variety of ways in the united states. a large sample (one household in six) is imbedded in the census collection itself. this sample allows for the collection of detailed housing and persons information not covered by a complete count. we also use sampling for the measurement of undercoverage by selecting a large post-enumeration survey (pes) of approximately 1 50,000 households. the results of this survey are then matched to the census enumeration to measure the extent and characteristics of census undercoverage. major census follow on surveys of residential finance and of scientists and engineers are conducted immediately following the census. frames for these surveys are constructed by screening units with particular characteristics from census questionnaires. all of these survey collections are major operations and complete papers suitable for this conference could have been produced for each of them. i would like to focus my remarks today, however, on an additional use of the census that being for a frame for selection of the major household surveys conducted by the united states government this paper will discuss our ongoing plans to redesign our current household surveys based on the 1990 census. i will discuss our general methodology and the challenges we face by trying to simultaneously select several surveys simultaneously. 1 will also mention briefly some of the planned uses of new technologies in the reselection of our survey samples. using the census as a frame for continuing household survey the united states decennial census address lists are used as a sampling frame for many of the government's major continuing household surveys. the principal household surveys using the census as a frame are: 1. the current population survey, (sponsored jointly by the labor and commerce departments; the basic labor force survey). 2. the consumer expenditure surveys, (sponsored by the labor department; used as input to the consumer price index). 3. the current point of purchase survey, (sponsored by the labor department; consumers are interviewed to generate a frame of retail outlets for measuring prices for the consumer price index). 4. the survey of income and program participation, (sponsored by the commerce department; a longitudinal survey which follows persons every four months for twoand-one-half years to collect information on income dynamics and use of government transfer payments programs). 5. the national crime survey (sponsored by the justice department; collects information from victims of crime). 6. the american housing survey (sponsored by the department of housing and urban envelopment; a biennial longitudinal survey of housing which updates the census data for sample units, while adding in new construction). the census address lists are the primary source of the sample for all of these siu^eys. these lists give us a frame for the united states as of census time 19sk). since these are continuing surveys, the sampling frame derived from the census must be kept up to date between censuses. this is done mainly by sampling building permits, which are required for new construction in most parts of the country. permits are listed and sampled every month from selected building permit offices. where such permits are not required by local governments, new construction is represented through area sampling. in the area sample, a list of all the addresses for selected areas on the map is made by the field representative. area sampling is also used for both old (existing in 1990) and new construction in some mostly rural areas where permits are not required for new construction and for areas where the census addresses are hard to locate. the census address list is an inexpensive source of sample, compared to an area sampling approach, and it 16 lassist quarterly gives more complete coverage of the peculation than any other available list of addresses. but even with the census list we find coverage to be a problem. potential coverage problems can be of two types; coverage of households and coverage of persons within households. evaluation of the coverage of households in the census shows that at most a one to two percent overall undercoverage at the time of the census, although undercoverage for households of minority races is known to be substantially worse than for the population as a whole. coverage of households by the continually updated census frame is more difficult to measure, but estimates of missed households range from about one to five percent. estimates of missed persons are more reliable, since the survey estimates of the number of persons may be compared to updated census estimates. this comparison shows on the order of a 10 percent undercovwage of persons, with the worst coverage for young males. there is evidence that for young males most of this is due to failure to obtain complete lists of household members, rather than to missing households. the updated census estimates ot persons by age, race, and sex are produced by inflating or deflating the census counts for births, deaths, immigration and emigration. most of the survey estimates are calculated using poststratification to bring the final survey estimates of persons into agreement with the updated census estimates. the surveys using the census as a frame are all conducted by the bureau of the census, although the data may be analyzed and published by other government agencies or research organizations. by law, no one outside the census bureau may have access to the actual census addresses, nor to information which would permit the identification of any sample household. this places limits on the amount of detail which can be included on data files intended for public use. it also means that only sworn census bureau agents may contact the sample households for any survey which uses the census as a frame. to avoid these restrictions, another major household survey conducted by the census bureau, the national health interview survey, uses only an area sample. the operations for this survey are coordinated to some extent with the other household surveys, and it is redesigned in conjunction with them, but the frame for this survey is created independent of the census. having a single field staff conduct the interviews for all survey is extremely cost-effective. the administrative and office costs can be shared. also, in many cases, the same field representative can conduct interviews for several surveys, since the surveys take place at different times of the month. this sharing of field representatives reduces the number of field representatives who have to be recruited and trained. detailed coordination of the sampling operations also saves effort and money. the lists of building permits used to keep the frame up to date can be shared among the surveys. in the area frame, the later surveys in a particular map area can make use of the lists made for the earlier surveys. the cost savings from sharing listings are substantial. since the sample addresses for the different surveys are kept close together whenever possible. redesigning the sample after each census the census address list was first used as a sampling frame for the ongoing household surveys following the 1960 census. in theory, this frame could have been kept up to date perpetually by adding new construction from building permits and area listings. in reality, the sample was reselected after the 1970 census and again after the 1980 census. the sample will be reselected after the 1990 census. one reason for reselection is the likelihood that in spite of our best efforts, continued updating of the old frame inevitably will lead to a gradual deterioration of coverage. additionally, as time goes on, a greater proportion of the sample would come from the mwe expensive permit frame. a basic reason foe redesigning the sample after each census is to use information from the new census in improving the design. the census information collected for each household includes household size, race of the occupants, whether the unit is a farm, whether the dwelling is rented or owned, and the rent or value of the dwelling. a sample of about one-sixth of the households receive a "long form" during the census, which asks for many additional details about the dwelung and its occupants, including income and labor force status. all this information is used in the redesign to restratify the primary sampling units, so as to reflect changes which occur between the censuses. the economic characteristics of many metropohtan areas have changed in recent decades, and there has been a shift of population to some formerly less developed parts of the "sunbelt" in the southern portion of the united states. areas which were similar 10 to 20 years ago may be very different today. taking these changes into account results in more efficient and reliable samples. prior to the 1980 redesign, all the surveys used the same general purpose sample design. this general purpose design was modified somewhat for the cuirent population survey (cps) to improve the measurement of labor force data. the other surveys had smaller sample sizes than the cps, but used the same psus, the same cluster size, and the same within-psu stratification. indeed, the other surveys were merely allocated a portion of the extra "reserve" sample which was selected for cps along with its regular sample. as the importance of the other surveys grew, greater attention was paid to their sample designs. following the 1980 census, the surveys were redesigned individually, in an attempt to optimize them based on their separate objectives, rather than using a common design modified only for measuring labor force characteristics. the expenditure and housing surveys sort units based on census characteristics which are highly correlated with the variables measured by the survey. the expenditure surveys concentrate on rent or value of housing, which is asked of all units in the census. the housing survey takes a subsample of census "long form" households, for spring 1990 which detailed housing characteristics are available. the other five surveys do not sort individual households using census characteristics, either because the relevant questions are not asked in the census, or because the relevant variables are not stable over time and the benefits of sorting would quickly dissipate. another reason for not sorting separately for all surveys is that some surveys interview clusters of adjacent households to reduce travel costs. some sorting, such as separating urban and rural areas within each county, is used for all the surveys. the cps is still our largest survey and as such still has some effect on the others. the cps is the only survey that attempts to measure data for states as well as for the united states as a whole. in 1980, the cps sample was selected as 51 independent state samples, one from each of the 50 states and the district of columbia. this was necessary because there was a reliability requirement fw the unemployment estimates for each state. the need for reliable state data was in response to the allocation of federal funds determined in part by the estimated state unemployment rates. the other surveys use sample designs aimed at making national estimates, and therefore their primary sampling units may cross state lines. although each survey now has its own stratification of primary sampling units, steps were taken in the 1980 redesign to maximize the selection of common sample areas across surveys. this will be done again in the 1990 redesign. this allows field representatives to be shared, and allows the permit and area samples to be better coordinated. the largest metropolitan areas are automatically in sample fw all the surveys. several of the surveys select their sample psus as a subset of the cps psus. the crimes survey used a "maximum overlap" method, in which each psu was given the appropriate unconditional probability of selection, while the expected amount of overlap with the cps was maximized. (this maximum overlap sample selection is still commonly called "keyfitzing" after nathan keyfitz, who developed one of the original methods). a different mathematical method is now used, but the objectives are the same as keyfitz addressed. the same maximizing technique will be used to select the new 1990 cps psus from their new strata while maximizing the expected overlap with the old cps sample psus. this was done to reduce the need to tfain new interviewers in some areas while having to lay off interviewers in old areas. effect on new technologies the last 10 years have seen increasing automation of census data products and'survey interviewing techniques. this affects the sample redesign directly because automation offers potential new efficiencies in the sampling operations, and indirecdy because changes in design may be needed to make the most efficient use of the new interviewing technologies. the most dramatic change in interviewing techniques has been the inu-oduction of computer-assist^ interviewing, combined with increased use of telephones by the interviewers in the field. the census bureau has recently opened a centralized computer-assisted telephone interviewing (cati) faciuty, located in hagerstown, maryland. interviewers at the facility work from a computer terminal which automatically selects the next case for interview, schedules callbacks, and displays the questionnaire on the terminal's screen. the computer program determines the path through the questionnaire based on the answers which are entered, checking the responses for consistency as it goes. we expect this system to improve the control of data quality, both because of the control provided by the computer and because interviewers can be closely monitored by the supervisors at the centrahzed faciuty. the system ehminates labor-intensive data entry and some steps in processing, which are needed for paper questionnaires. cati has been tested successfully for the national crime survey and is being tested for the cps. we expect the cati methodology to be in full-scale use in the early 1990's. its use for household surveys has so far been restricted mainly to follow up interviews for panel surveys. an address sample is still used and the first visit to a household is made in person. even when full cati is being used, there will still be a need for field interviewers. not every housing unit has a telephone and some that do request a face-to-face interview. testing the use of computer-assisted personal interviewing to complement cati is now underway and is also well underway. some testing using samples based on randomly selecting telephone numbers has been done by the census bureau. however, the response rates in these tests were much lower than when the first visit was made in person. because of this, along with the need to represent households without telephones, a purely telephone sampling approach is no longer considered for most of the household surveys. an exception is the current point of purchase survey (cpp), which is just completing a test of a combined telephone-list/random-digit-dialing approach, with promising results. this could eventually remove cpp from the ust of surveys using the census as a frame. we plan to select some cpp sample in the 1990 redesign as a backup strategy in the event that the random-digit dialing approach proves unsuccessful. one aspect of the 1990 sample redesign will be to modify the sample designs to make more efficient use of cenffalized cati. cati removes some of the follow up interviewing from the dispersed field representatives. this means that the field interviewers will be underutilized unless they are given greater workloads initially. thus, for a constant total budget, the optimal design using centralized cati will have fewer psus, wi3i a larger initial workload in each psu. the increased use of telephoning also reduces the relative importance of travel costs, which may reduce the optimal cluster size for those surveys which select clusters of households. the increased use of computers in the 1990 decennial census will make it easier than ever before to use census data in constructing a sample frame. the most important innovation is a geographic database known as the topologically integrated geographic encoding and 18 assist quarterly referencing (tiger) system. this computerized system (developed jointly by the census bureau and the u.s. geological service) will produce a map of any city block or comparable rural "block," with roads and natural features correctly represented. addresses from the census will be linked to the correct block on the map and in most instances the location within the block will be indicated. the tiger maps have the potential to revolutionize the area sampling q)erations. in the past the devel(^ment of maps has been a particular problem and staff members have struggled with maps and information of inconsistent quality from a variety of sources. the computer generated maps will also simplify locating those new units whose building permits have been selected. the tiger maps and data (without detailed census address information) will be available to the general public and will be useful for area sampling by survey organizations outside the census bureau. another 1990 census product which will facilitate using the census as a sampling frame is the automated address control file. this contains a record for each census address, with basic information about the housing unit at that address. for about 95 percent of the records, the actual address will be included as text on the file. in previous censuses, the addresses could only be obtained by going to the handwritten register completed by the census enumerator, which necessitated an expensive address keying operation before the surveys could use these addresses. computer technology was also used to advantage in implementing the mathematical methods for stratifying and selecting psus in the 1980 redesign. similar methods will be used in 1990. the cps strata before the 1980 redesign were formed by writing key psu characteristics on 3x5 index cards and grouping the cards manually to form intuitively homogeneous strata of roughly equal stratum population. for the 1980 redesign, a multivariable clustering algorithm was modified to form strata so as to minimize a measure of total variance for a set of specified variables, subject to constraints on the stratum peculation. also, for the 1980 redesign, an improved "maximum overlap" method was developed, which selected a probability sample of psus while maximizing the overlap with some other survey's selected areas. this method used a linear programming algorithm to maximize the expected overlap, subject to constraints on the probabilities. this gave a greater percentage of common psus than methods used previously. coordinating sampling with different sample designs a central theme of the 1990 redesign research is to better coordinate the sample selection operatiions for the different surveys. as i have described already, in the 1980 redesign we allowed different surveys to use different psus and different ways of sorting, stratifying, and selecting households within the psus. at the same time, every effort was made to keep the surveys' sample units close together to save on the cost of keying addresses, listing for the area sample, and sampling building permits. this task of linking different sample designs turned out to be quite compucated. an example is the coordination of listings for the area sample. if surveys are to share each othct's lists of households in sample blocks, then it is necessary to keep a cross-referenced index so that each survey can find out who else has previously made a list of the block and where that list is being kept this sort of thing was much easier in the 1970 redesign where there was only the one cps design, so that it was only necessary to find out whether the block had been listed for the previous cps sample. extensive recwd-keeping is also needed to avoid duplicate selection of the same address by different surveys. united states government statistical policy, as set down by the office of management and budget, is that a single address should not be included in more than one census bureau survey. with different surveys selecting sample from the same universe, using different method, it was not easy to avoid such duplication. particular problems were in the health interview survey, which used an area sample where the other surveys were using list sampling, and the housing survey, which selected specific longform units where the other surveys were using area sampling. all this complex record-keeping and cross-referencing is amply justified by the large savings from coordinating the survey operations. however, the complexity becomes a liability if at any time between redesigns, one of the surveys has its sample reduced, expanded, or has a change in the scheduled interview dates. when such changes are made, all the references to the changed survey anywhere in the reference system must be checked and updated, to avoid duplication and other operational problems for the other surveys. because the system was not designed with updating in mind, this causes even small changes in a survey's sample to be time-consuming and expensive, even when the changes can be made by computer. the sampling of building permits is especially inflexible. permits are sampled every month as they are issued by the permit offices. (there are over 10,000 of these offices throughout the country.) many of the offices destroy their old records after a few years, so it is impossible to go back and select more permits. to try to simplify the record-keeping in the 1990 redesign, we will closely examine the details of all the clerical and computer procedures, and standardize these procedures whenever possible. some of the major research issues ccmcem whether specific siuveys would incur significantly higher field costs or higher variances by simplifying on certain (^rational details in order to standardize their procedures with the other surveys. in the 1980 redesign, four separate sets of computer programs were us^ to select the sample for the seven surveys. (some surveys were able to share programs). these programs were developed separately and there were minor differences in defmitions and data formats spring 1990 which tumed out to be major barriers to coordination. in the 1990 redesign, the plan is to use one integrated set of computer programs for the entire sampling operation. in developing these programs, our programmers propose to use a computer-assisted software engineering approach. besides providing a common logical framework for the computer algorithms, the computer-assisted planning produces a common data dictionary for all the programs. this will enfwce standardization of definitions across the surveys. selecting several surveys from the same frame can add complications to the mathematics of sample selection. a basic example is that if a survey removes units from the universe with probability proportional to size, then the remaining universe tends to under-represent large units. such concerns need to be kept in mind as the sample selection methods are designed. a goal of the 1990 redesign is to leave a "clean" universe, so that future surveys can be selected from the census address lists, building permit offices, and the area frame without having to make special adjustments because of the sample of units which has been "removed" by the redesigned household surveys. one final challenge in coordinating sample selection for multiple surveys is getting agreement on a time for selecting the sample. it is most economical to do the bulk of the sample selection work at the same time for all surveys. however, this requires all the different sponsoring (xganizations to complete their research on the new sample design in time. some agencies prefer to have their redesigned sample introduced as soon as possible after the 19w census, to take advantage of the new design. others would benefit by having more time to use the 1990 census data in research on special topics affecting their survey, before deciding how to design their survey. "as soon as possible" after the 1990 census tiutis out to be nearly four years later, the 1990 redesign sample will start being introduced in april 1994. part of this lag is due to the census processing; the last 1990 census data file used in sample selection becomes available about 18 months after the official april 1, 1990 census day. once the design has been specified, computer processing to prepare materials for the clerical work takes about 12 months, and the various clerical activities and related processing take about 9 months. this leaves about 9 months to use thel990 census data to specify the design, including selecting psus, deciding how to stratify units within psus, and deciding on the sample size at each stage of selection. obviously most of the basic research, planning, and software design has to take place prior to the availability of the 1990 census data. 'presented at the ifdo/iassist 89 conference held in jerusalem, israel, may 15-18, 1989. notice to lassist readers to all iassist members we are trying to collect as many photos as possible taken at lassist conferences or other official functions. if you have pictures please send a copy to: sue gavrel 129 blackburn ave ottawa, canada kin 8a6 we hope to have an album or two for the edmonton conference 20 lassist quarteriy vol30-1.indd 4 iassist quarterly spring 2006 editor’s notes welcome to the first issue of the iassist quarterly, vol. 30, the first 2006 issue. from my seat, it is that time of year when days are gray. danish poet henrik nordbrandt has said: “the year has 16 months: november, december, january, february, march, april, may, june, july, august, september, october, november, november, november, november.” but november is a good time for planning, for writing abstracts, and for preparing to go to montreal in the sun of may 2007 for the next iassist conference. read more about the conference in this issue. iassist conferences are international and the first article, “setting up acquisition policies for a new data archive,” is international and based on two presentations at the iassist conference in edinburgh, may 2005. in the session, “enlightened policies: improving collections and acquisitions,” finnish and slovene data archivists found that their two presentations could be combined. consequently, they compiled a joint article from the presentations. the authors are helena laaksonen and sami borg from the finnish social science data archive, and janez stebe from the social science data archives at the university of ljubljana in slovenia. when discussing data quality, central concepts are: accuracy, timeliness, accessibility, comparability, and coherence (eurostat). furthermore, the archive must include the cost and burden of acquisition, processing and distribution by the archives. after the initial phase of archiving mostly quantitative social science data, the archives have moved to include qualitative data and also expand the coverage area into educational and health sciences. it is also found that some archives do not reject datasets, but there might be very little processing done if the data are considered to be of low quality. the article not only concerns the two archives, there is also a short report on “what kind of support social science funding organisations give to archiving and data sharing.” a solid kind of support from data producers is when data are “understandable,” have “clarity,” or to put in archivist-lingo, are “well-documented.” that leads us to the second article in this issue, “problems of comparability in the german microcensus over time and the new ddi version 3.0,” by jeanette bohr, andrea janssen and joachim wackerow from the centre for survey, research and methodology (zuma) in mannheim, germany. the article starts with a short introduction to the data documentation initiative (ddi). the problem with earlier versions of the ddi metadata standard was that it did not have an option for recording changes over time in repeated surveys. the german microcensus is repeated annually and is undergoing changes from year to year, and these changes must be documented. ddi 3.0 addresses problems of describing the data life cycle. this includes describing groups of studies as well as the relationships within collections of comparable studies. the ddi version 3.0 is not yet a frozen standard and this article is considered a contribution to the exploration of the coming standard. the article demonstrates the hierarchical grouping model with examples of the structure of the german microcensus. the conclusion consists of some pros and cons concerning the 3.0 version of the ddi. ultimately the new version is regarded as a better instrument for managing and processing of metadata. the third article is also about metadata: “everything but the kitchen sink: building a metadata repository for time series data at the federal reserve board.” this was presented in the session, “the essential role of metadata in resource discovery,” at the iassist 2006 in ann arbor. the authors are san cannon and meredith krug from the federal reserve board. the federal reserve board uses a variety of time series data for both research and forecasting. as stated in the preface to the article: the tasks consist of collection, maintenance, and upkeep of more than 50,000 time series from many sources, plus documenting the metadata for the compilation and use of the data. the metadata repository will link three kinds of metadata about the time series: structural metadata, reference metadata, and operational metadata. structural metadata is the short text that brings meaning to what would otherwise just be a number without value. the reference metadata contain more detail about the calculations behind the number, such as data collection, sampling, etc. operational metadata are the exact processing instructions. while in the past the information was in different formats and on different platforms, this project is building a repository, which includes, among other things, a hierarchical nomenclature system in the structural metadata that is exemplified in the article. also addressed is the “challenge of time” an obviously important area for these data. “everything but the kitchen sink” typically means that the kitchen sink would be too much, but i think it also means that a few other things incorporated probably already are too much. however, in my evaluation, the title is supposed to be a catcher of attention not a statement about having too much metadata. the iassist website is always open http://iassistdata.org and you can look at conference information and visit the iassist blog the iassist communiqué – at http:// iassistblog.org. articles for the iassist quarterly are most welcome. articles can be papers from iassist conferences, from other conferences, from local presentations, discussion input, etc. contact the editor via e-mail: kbr@sam.sdu.dk. karsten boye rasmussen, november 2006 vol19.1 9fall/winter 1994 abstract a valuable meteorological data archive collected by the alberta research council over the course of the hail studies project in central alberta is in jeopardy of becoming unusable as the digital data stored on magnetic tape degrade over time, and expertise in the data collection, calibration, and interpretation becomes scarce. the overall goal of this project was to preserve the digital radar, aircraft, upper air and surface precipitation data along with supporting calibrations and documentation; to transfer this archive to the university of alberta; and to make the archive available to the scientific community. there were three distinct operations carried out to ensure the long-term preservation of the archive; retreival of the digital data and all supporting (secondary) datasources; transfer of digital data from magnetic tape to compact disk; and the collection and preparation of relevant documentation describing the data. the archive will provide researchers with a documented dataset to support further research in radar meteorology, climate change, hydrology, cloud physics, mesoscale meteorology and severe weather phenomena. 1. introduction the acquisition of atmospheric data is an expensive endeavour, and the data are usually irreplaceable. the subsequent research uses of good data are often not contemplated by the original data collectors. for example, data can be re analyzed to test new hypotheses, or can be used for comparative analyses with other geographic areas. data may become unusable when supporting documentation is lost or destroyed, or when the physical media on which these data are stored become no longer readable through degradation over time, or through the lack of equipment capable of reading the physical medium due to its obsolete format. the alberta hail studies project (1956-1985) was established to study hailstorm physics and dynamics and to design and test means for suppressing hail. central to these activities was the alberta research council’s (arc) radar facility located at the red deer industrial airport in central alberta (figure 1). a vast amount of data was collected from several platforms to conduct research into precipitation mechanisms, severe storm development, hail suppression, hydrology and microwave propagation. since the termination of the alberta hail project in 1986, numerous research projects have demonstrated the value of using the alberta data archive. during the period 1990-1994, 23 archive-based publications have appeared in refereed journals and conference proceedings and 4 scientific reports have been prepared. there have also been nine graduate theses (2 ph.d., 7 msc) awarded at 3 universities during this period. the areas of study have included radar meteorology, cloud physics, hydrology/hydrometeorology, computer science, instrumentation, and synoptic, dynamic and mesoscale meteorology. scientific research and collaborations continue to this day. recognition that this valuable meteorological data archive was in jeopardy of becoming unusable as the digital data stored on magnetic tape was degrading over time, and expertise familiar with the data collection and calibration procedures, and their interpretation became scarce, prompted an effort to save this unique dataset. 2. objectives the specific objectives of this project were: 1. to transfer the computer-readable radar, aircraft, upper air and surface data from the existing shortterm storage medium (magnetic tape) to an archival medium (cd rom). 2. to collect all available supporting (secondary) data sources and develop the necessary documentation to describe the computer-readable data files. 3. to coordinate these efforts with the university data library and develop appropriate mechanisms to make the archive available to the scientific community. 3. approach and work plan there were three distinct operations carried out to ensure the long-term preservation of the archive; retrieval of the digital data and all supporting (secondary) data sources; transfer of digital data from magnetic tape to compact disk; and the collection and preparation of relevant documentation describing the data. 3.1 the data archive the radar facility near red deer consists of a unique polarization-diversity s-band (10 cm) weather radar, a standard c-band (5 cm) weather radar, and an x-band (3 cm) radar used to track aircraft through a transponder data rescue: experiences from the alberta hail project by b. kochtubajda 1, 2 c. humphrey 2 m. johnson 3 10 iassist quarterly system. with the addition of computer interfaces in 1974, a systematic archive of radar data was initiated. this archive now includes close to 200 gb of data, representing approximately 12, 000 hours of multi-parameter radar data. in addition to the radar archive, an extensive archive of aircraft, surface precipitation, and upper air data has also been collected. approximately 18 gb of data were recorded between 1983 and 1985 aboard an instrumented research aircraft flying through convective storms and cumulus clouds. also, quantitative precipitation reports (hail and rain) were obtained from approximately 500 ground stations within the radar coverage, between 1974 and 1985 and in 1989. these data exist on several media, including 800, 1600, 6250 bpi magnetic tape and 8 mm cassettes. comprehensive radar, aircraft and upper-air software packages are also available for data analysis and display. 3.2 retrieval of the digital data and all supporting data sources a number of assumptions were made before the digital data were recovered. to provide the broadest range of research potential of the data, unprocessed aircraft and radar data would be provided. this would yield a quicker retrieval of the data and (given the time and budgetary constraints of the project) would result in more of the data being recovered. to maximize the research potential, we would archive all data types and work backward by year, from 1985 towards 1974, thus ensuring a complete multiplatform data set. 3.2.1 aircraft data physical experiments designed to explore the potential of hailfall suppression, and rain augmentation through airborne glaciogenic seeding on convective cells were conducted in central alberta between 1983-1985, as described in humphries et al (1986). these studies emphasized in-situ aircraft measurements to investigate natural and artificially modified precipitation processes. the primary observational platform used in these studies was the intera/ alberta research council cloud physics instrumented research aircraft, a cessna 441 conquest, pressurized twin-engine turboprop aircraft. data from the instruments were managed by a computer based data system which provided data acquisition, recording, and real-time calculations and display (johnson et al, 1987). approximately 300 magnetic tapes were processed for all the research flights conducted from 1983 to 1985. data were segmented into unique files for each hour of the day. the file names follow the iso 9660 level 1 standard, and are of the form yymmddhh.adb where yy year; mm month (numeric to help in sorting); dd day; hh hour of the day; and adb represents the file extension identifier for aircraft data block. aircraft data were recorded with coordinated universal time. yearly index files summarizing the amount of information collected for each hour of the research flight were prepared. the first line provides a brief description of the purpose of each research flight, including the date, start and end time of data collection (hh:mm utc) and the type of study carried out. the subsequent lines describe the filename, file size (in kbytes and mbytes) as well as the number of records collected (including the 2-d imagery). 3.2.2 polarization radar data the s-band polarizationdiversity radar, installed in 1967, operates at 2.88 ghz and has a parabolic reflector antenna with a 6.67 m diameter dish that produces a 1.15_ beamwidth in both azimuth and elevation. the radar sweeps out a helical volume scan rotating at 48_s1 (7.5 s per revolution) and rising 1_ in elevation for every 360_ in azimuth, up to a maximum elevation of 8 or 20_ selected depending on the proximity of the storms to the radar. the radar records data with an approximate azimuthal resolution of 1_, for 147 range gates, from 3 km to a distance of 157 km from the radar, with a range resolution of 1.05 km per range gate. the radar can transmit any polarization, but has usually transmitted left-hand circular (lhc) polarization with 450 kw peak power. the receiver circuitry digitally records four measurements from the figure 1 11fall/winter 1994 12 iassist quarterly lhc and rhc components from each range bin. these are the rhc co-polar signal power, the lhc cross-polar signal power, and the correlation and phase between the lhc and rhc signals. approximately 240 s-band radar tapes (previously copied to 6250bpi) have been examined for the period 1980 to 1985. radar data were extracted into files containing one complete 3d volume scan, representing either 1.5 minutes for the 8_ scan, or 3 minutes for the 20_ scan. a quality control report was produced with each data file containing information about the azimuth and elevation time histories of the volume scan. the file naming convention adopted for the radar data uses all 11 characters (yymmddhh.mmr) where yy year ; mm month (numeric to help in sorting); dd day; hh hour of day (24 hour clock); mm minute of first data in file; and r represents the data type (s for s-band, c for cband, q for quality control reports). radar data were recorded with local (mountain daylight) time. 3.2.3 surface hailfall and rainfall data volunteer observations of precipitation events were used to supplement quantitative surface measurements obtained by specially equipped vehicles, which were directed beneath thunderstorms. after each storm, telephone surveys were conducted to collect hail and rain reports. report cards were also received by mail from volunteer farmers. this information was useful to develop hail climatologies and correlate hailstone characteristics with crop damage. the surface data collection includes the digital hail and rain report files (yyhail.dat, yyrain.dat), from the telephone surveys for the period 1957 to 1985 (except those files missing from 70-73); selected time-resolved hail and rain truck observations (yymobile.dat); and daily precipitation measurements collected by approximately 500 volunteer farmers during the months of june, july and august for the period 1975 to 1983 figure 2a 2b table 2: summary of data compact disks data type year no. files no. cds total mb aircraft 1983 349 5 2947.5 1984 226 3 1706.8 1985 166 2 1284.8 radar 1980 13 3 1433.1 1981 9 1 467.9 1982 20 3 1657.8 1983 48 8 3681.5 1984 56 9 4775.2 1985 51 11 5093.3 13fall/winter 1994 (yyprecip.dat). 3.2.4 upper air data (limex) a mesoscale upper-air study, limestone mountain experiment (limex-85) was carried out over the foothills and mountains of southwestern alberta during july, 1985 (strong, 1989). the objectives of the field experiment were focused on mesoscale convective processes, orographic effects, and interactions with synoptic processes, with particular emphasis on severe storm forecasting applications. the archive data includes two-hour soundings from nine upper-air sites with an average spacing of 50 km, continuous sodar profiles, research aircraft soundings at 20-km intervals, surface data from eight automated systems, and an extensive cloud photo set. the compressed upper air data and accompanying analysis software are currently archived on 2 high density diskettes. 3.2.5 supporting data sources an equally important component of the retrieval process was the collection of the secondary (supporting) data sources including aircraft mission scientist notes, radar plot summaries, various operational log books, checklists, calibration notes, cloud photographs and video tapes. the data were boxed and transferred from the alberta research council to the meteorology division at the university of alberta. subsequently, the boxes were itemized and given a location identifier. the contents of each box were stratified into one of five categories (photos and maps; data summaries and timelines; operational logs, checklists; and calibrations and data. a listing has been prepared which stratifies the data according to the type of information and the year of collection. a subset of the listing (for 1983-1984) reproduced in table 1, illustrates the variety and richness of the materials available. 3.3 transfer of digital data from magnetic tape to compact disk the procedures used to produce the aircraft and radar data compact disks are depicted in figures 2a and 2b. a series of programs were used to copy unprocessed aircraft data from tape to file (copyadb), and to generate hourly flight files (timesel). these hourly files were backed up on a series of 8mm exabyte data tapes and high capacity sony compact tapes. complete 3d volume scan data files and associated quality control reports were generated from the radar tapes (tapedd) and also backed up on exabyte and sony tapes. compact disks were produced using a pinnacle micro rcd1000 recordable compact disk recorder with macintosh authoring software and a quantum 1 gb fast scuzzi 3 hard drive for preparing a cd image. the requirements on the hard drive included an average seek time of 12 milliseconds or faster, a transfer rate of 1.2 mb per second or better, and an intelligent calibration feature (ie: thermal recalibration not performed during a continuous read) to avoid a write interruption which would render the cd invalid. the iso 9660 level 1 standard was selected as the format for the cd-roms. this standard allows the same cd to be read and interpreted on mac, ms-dos, unix, vax/vms, and other computer platforms. this includes the restriction of file names to 8 characters, with a 3 character “extension”. an inventory of the compact disks produced is summarized in table 2. there are 46 cds in total, including 10 research aircraft data cds ; 35 cds containing the s-band polarization data from 197 days between 1980 and 1985; and 1 cd containing the surface precipitation files, dataset documentations, and the radar and aircraft software source codes. each aircraft cd contains a series of sequential hourly data files and 3 text folders (mac, dos, and unix) containing the file summaries and a disclaimer. the directory structure of a radar cd is described in figure 3. a radar cd contains a series of daily files (yymmdd). each file contains 4 sub-directories (data, quality, calib, logs). the data sub-directory contains the sequence of radar volume scans. the quality control reports for each radar file are located in the quality sub-directory. the calibration text files and the daily radar and transmitter log files are found in the calib and logs sub-directories, respectively. 3.4 documentation describing the data a series of documents have been gathered and/or prepared to describe the various datasets. aircraft experiment descriptions including study objectives; flight procedures; aircraft description and instrumentation list; 2d image processing; aircraft tag descriptions; daily flight assessments (instrument evaluations); and sensor calibration files accompany the digital aircraft files. digital radar and transmitter logs for the period 1977-1985; as well as descriptions of the radar characteristics and scan protocols; data structures; and calibration files accompany the digital radar files. the primary documentation for the surface hail and rain dataset is a coding sheet describing the file format. the daily farmer precipitation files from 1975-1983 has an accompanying text file. an extensive bibliography of hail project related papers has been compiled. the original hard copies are currently stored in the meteorology division at the university of alberta. software packages to analyze and display radar, and upper air data, developed at university of essex and at aes in saskatoon, have been obtained and can be shared by users. a summary of the archive as it is currently configured is presented in appendix 1. 4. conclusions potential users of the archive have indicated that the radar and ground measurements dataset would be used to continue severe hailstorm and rainstorm studies; to provide input to 14 iassist quarterly distributed hydrologic models; to carry out radar-based precipitation climatology studies; and to validate numerical models being developed during mags, boreas, gcip, or base. the aircraft archive would be used to improve our understanding of the chemical composition of cloud water and the processes which affect it, and for icing research. the limex upper air dataset would be used for the atmospheric correction of noaa-avhrr data in estimating regional evaporation, as well as in moisture budget estimates and evaluation of evapotranspiration studies. a documented archive of radar, aircraft, surface and upper air data has been provided from which further research in these areas can be carried out. the retention and preservation of the archives through the university of alberta will ensure the continued accessibility and long-term survivability of these datasets. acknowledgements this project was financially sponsored by the university of alberta, alberta research council, and the atmospheric environment service. mr. s. kozak and mr. f. bergwall assisted in the data retrieval. the authors also acknowledge the contributions and support of drs. ep lozowski, gw reuter, and t gan (uoa), dr. a. holt (uoessex), dr. dr rogers (colorado state univ), drs. ga isaac, p. joe, gs strong, (aes), and dr. bl barge, and mr. cf richmond (arc). references humphries, r.g., m. english, and j.h. renick, 1986: weather modification research in alberta canada. 10th conf. planned and inadvertent weather modification, arlington, ams, 357-361. johnson, m.r., l.e. lilie, and b. kochtubajda, 1987: a data structure for acquisition, analysis, and display of meteorological data. 6th ams symp. on met. obs. and instr., new orleans, ams, 397-400. strong, g.s., 1989: limex-85: 1. processing of data sets from an alberta mesoscale upper-air experiment. climatological bulletin, 23, 98-118. 1 paper presented at the iassist confernce 2 university of alberta edmonton, alberta 3 alberta research council edmonton, alberta appendix 1: archived database summary data type period archive filename calibration documentation digital s-band radar 1980-1985 yymmddhh.mmr (n = 1-4) radar data structure yymmddna.txt radar characteristics rlyymmdd.txt calibration procedure tlyymmdd.txt (n = 1-2) yymmddnx.txt digital aircraft data 1983-1985 yymmddhh.adb expt desc data index aircraft + instruments daily scores video logs surface reports: 1957-1985 yyhail.dat hailcard coding form hail and rain yyrain.dat daily rainfall reports: 1975-1983 yyprecip.dat file description 500 stations (june 1 sept 1) mobile reports yymobile.dat file description hail and rain upper air data (july 4 23, 1985) limex-85 9 upper-air stns. 8 auto surface stations analysis software: radar / aircraft / upper air 15fall/winter 1994 data rescue: experiences from the alberta hail project addendum the earth and atmospheric sciences department in collaboration with the university of alberta data library, and the information systems department of the alberta research council have just completed a 15 month effort to rescue the alberta hail studies project dataset. the project included the organization, retrieval, and formatting of the digital data and all supporting (secondary) data sources; the transfer of digital data from magnetic tape to compact disk; and the collection and preparation of relevant documentation describing the data. there are 62 cds in total, including 10 cds of research aircraft data collected between 1983-1985; 47 cds containing the s-band polarization data from 287 days from 1979 to 1985, and in 1989 and 1991; 4 cds of coincident c-band radar data collected on those days when both radars were operating simultaneously (44 days) between 1979 and 1991; and 1 cd containing the surface precipitation files, aircraft transponder files, dataset documentation, and the radar and aircraft software source code. the archive also includes the collection of supporting data (such as; operational log books, manuals, photographs, slides and videos) and 2 diskettes of upper air data and accompanying analysis software. example access software (in the c programming language) and documentation has been developed for quick inspection of the original unprocessed aircraft and radar files from the cds, and as a demonstration of data access. a set of world wide web pages has been developed and is now available on theinternet via browsers such as netscape and mosaic. the “alberta hail project meteorological and barge-humphries radar archive” can be accessed through network services provided by the data library at the university of alberta, by opening the url: http:// datalib.library.ualberta.ca/ahparchive/ data can be accessed in one of three ways. researchers can obtain hail and rain data files directly from an anonymous ftp site: (ftp://datalib.library.ualberta.ca/ahparchive). to obtain aircraft and/or radar data from the compact disks, click on the order form and submit a specific request. for small amounts of aircraft and/or radar data (e.g.: a single case study), the set of hourly aircraft files or the daily radar directory will be transferred from the cd library and placed in the anonymous ftp site for subsequent retrieval. requests for larger amounts of data will result in the production of customized cds and shipment to the researcher for minimal cost. use of the archive, is subject to the following conditions: 1. these data are to be made freely available only to the scientific research community, whether national or international. 2. these data are provided for the exclusive purposes of teaching, academic research and publishing, and/or planning of educational services and may not be used for any other purposes without the explicit written approval, in advance, of the data library at the university of alberta. 3. the alberta research council, the atmospheric environment service and the university of alberta will be acknowledged in any anticipated presentations and papers associated with the archive. 4. the citation to be used for the archive is: alberta research council. the alberta hail project meteorological and barge-humphries radar archive: [computer files], edmonton, alberta, canada. alberta research council [producer], university of alberta data library [distributor]. august 1995. . 13 may 2003. 22. statistics canada. “history of the census of canada”. . 13 may 2003. 23. statistics canada. “history of the census of canada”. . 13 may 2003. 24. national archives file 9580-rg31, volume 1, “letter to mr l fry, acting chief statistician, from wi smith, dominion archivist”, 25 april 1980. 25. national archives file 9580-rg31, volume 1, “letter mr m swift, director general from da worton, assistant chief statistican”, 10 may 1982. 26. national archives file 9430-50/s5, volume 3, “fax to claude beaulé from jerry oʼbrien”, 12 december 1991. 27. terry cook, “an appraisal methodology: guidelines for performing an archival appraisal” (national archives of canada: 31 december 1991). page 4. * paper presented at the iassist 2003 conference in ottawa, canada. cara downey, library and archives of canada, cara.downey@archives.ca data development for international research (ddir) ddir ii: event data research by rchardl. merrit ' and dina a. zinnes university ofillinois at urbana-champaign numerous scholars of international relations have recently sought to improve the empirical quality of their research. they feel that quantitative approaches, properly designed and applied, can significantly enhance our ability to understand international events and interactions among nation-states. one result has been a plethora of analytic techniques that rely on mathematical bases. global modeling is an example of this direction. another result is the generation of new data sources. this article focuses on the latter tack: data development for international research. growing emphasis on quantitative data has not been without problems. for one thing, some researchers flat out reject their usefuhiess or validity. such intransigence obfuscates a central fact: our growing comprehension of social scientific knowledge is linked inexuicably to the computer-based information revolution. whether we like it or not, whether we comprehend it or not, we cannot avoid their implications for political analysis. both developments—new analytic techniques and data sources—demand greater sensitivity. another problem is weaknesses in early data collections. we cannot deny the fact that some important datasets were flawed, just as we cannot ignore criticisms about some analytic methods researchers have used. such weaknesses have contributed to misunderstandings, skepticism, and even occasional hostility. this article describes a particular research project undertaken in the field of international and cross-national relations by a community of u.s.-based social scientists. the data development for international research (ddir) project seeks to maintain, extend, and develop new data banks for the study and analysis of crossnational and international political phenomena. it was the outgrowth of three years of discussion, correspondence, and seminars involving both data collectors and data users. funding for 1986-89 by the national science foundation enabled the project's first phase (ddir i) to focus on four tasks: datasets in the areas of national attributes and interstate disputes, data planning, research organization, and international broadening. new nsf funding for 1991-93 permits a second phase (ddir ii) to concentrate on the area of event data. the article describes how ddir began, what it has done, and where it is heading. it seeks neither to assay the often sterile debate on the usefulness of quantitative approaches, nor to offer a definitive answer to the question of what analytic techniques and data sources are appropriate for what purposes. its concern is rather how the ddir community envisages the stauis of quantitative research in international and cross-national relations. it summarizes ddir's organizational background, philosophic orientation, and goals. origins: need for quantitative data four trends in the social sciences are particularly relevant for understanding the need to develop data for international research: • an explosion in the scientific study of national development and processes of interstate interaction has characterized the last five decades. questions concerning the relationship between national attributes and the domestic and foreign policy behavior of nations, the evolving structure of the international system, causes and consequences of international crises and war, and the dynamics of interstate interaction both confiictual and cooperative have come under careful and systematic scrutiny. many of the cherished maxims of international behavior have been shown to be false; and new insights into causes and consequences of national and international processes have been observed. • the awareness has grown that datasets are crucial within the context of the entire research process, and integral in the continuing feedback relationship between theory and research. contradicting the often trite argument that we allow our data to shape our questions, having large datasets that researchers know exist—and which continue to be maintained—opens up the range of research questions and continues development of theory in international relations and comparative poutics. • funding for data development has been at best sporadic. this has meant an inability to mount a concentrated and coordinated attack on fundamental problems facing the field. assist ouarteriy cuirently existing datasets are largely the work of a few dedicated researchers scattered throughout the country, who have been dependent on the vicissitudes of changing national funding strategies. there is no guarantee that these data collections will be continued and certainly no clear opportunity for extending and further developing them in response to the evolving needs of the research community. furthermore, while data collectors are generally aware of one another, there is no overarching mechanism to integrate and compare their results. this has led to unfcxtunate duplications of effort, differences in definitions, and differences in usage of sources. • the data movement of ihe past several decades has enhanced the methodological expertise for the extraction of data from public sources, development of indicators for basic concepts, and quality control through reliability checks. this, together with the extensive technological advances of recent years in computer technology, makes feasible the future development of considerably more valid and reliable datasets. these facts—the research record; recognition of the need for systematic datasets; the currently scattered, ad hoc nature of data collection activities; and the available methodological/technological expertise—point to the desirabihty of a large-scale, integrated effwt that can contain, extend, and further develop the data resources available to the research community of international relations scholars. such perspectives on the state of the art in international and cross-national relations generated an interest in taking action to improve the field's quality. a series of informal meetings, piggybacked on to professional conferences, and workshops at the university of dlinois at urbana-champaign and elsewhere led to a remarkable degree of consensus among several dozen researchers. these meetings and workshops eventually honed in on a basic decision. if those interested in using quantitative data did not take action, the participants argued, then opportunities to have such data would atrophy. accordingly, an effort to organize the relevant community of scholars was warranted and, indeed, long overdue. the researchers then focused on the overall strategy that such an organization. data development for international research, should pursue: should ddir serve solely as an interest group, or should it encourage and seek funding for research activities? and, if the latter, which kinds of relevant research should have ddir's initial attention? the organizational task was easily resolved provided that some colleagues were willing to devote some of their time and energy. the point of departure for ddir was in a sense the national election study project as a largescale, long-term data collection project for the enhancement of social science research, the nes clearly stands as a model. in another sense, however, important differences distinguish, on the one hand, the theoretical framework, goals, and structure of the nes and, on the other, the needs of the research community studying international and cross-national phenomena. the two research communities are diverse in the questions they ask, degree of consensus regarding fundamental methodological issues, and sheer number of researchers currently relying on the data collections. this diversity suggested the need for more decentralization in the data-collection efforts and communications framework than has been needed in the nes. ddir thus supports not a single, massive project, but rather individual researchers at different universities carrying out separate—though clearly related— projects. the diversity should be seen as a major strength of ddir i; and it is this orientation that guides ddir ii since it also points to a multiplicity of research agendas. while the pitfalls of decentralization are apparent, these dangers do not obviate possibihties for successful coordination and integration. the task of choosing areas for research focus proved to be more difficult simply because the potential areas are many and the competition for needed funding and other scarce resources is even greater. some of the principle supporters, none of them with any immediate claim for ddir-relaled resources, distributed questionnaires and carried out other research to ascertain how members of the potential community evaluated data priorities (mcgowan etal., 1988). a solicitation of research ideas, further consultation in meetings and workshops, and much telephoning eventually produced substantial if not complete agreement on a particular strategy. the informal consensus saw three activities: first of all, ddir would seek funding to carry out a discrete number of projects in two research areas, national attributes and interstate war, that its growing number of members considered most relevant and likely to be carried out. second, those administering ddir, in conjunction with an advisory committee, would assess the pwospects for similar research projects in two other areas, event data and international political economy (ipe) data, which ddir might wish to initiate later. third, ddir would also try to improve communications among scientists interested in quantitative research in international and cross-national relations. this meant on the one hand setting up a regular newsletter, ddir-update, and, on the other, scheduling at professional conferences both research sessions and organizational meetings ddir i: national attributes and interstate war ddir's first task, aimed at improving the quality of data spring 1991 on national attributes and interstate war, proceeded from a rich background. significant international and crossnational data collections were developed well before worid war ii (merritt, 1990). not until the late 1950s and early 1960s, however, did the large-scale data movement begin. as part of the general behavioral movement in political science away from assessments based on intuition or folk wisdom and toward more rigorous, systematic analyses, scholars became sensitive to their need fw adequate data bases to study key questions. is inequality in the world at large becoming more or less intense? do alliance configurations and power distributions enhance or decrease the probability of war? do internal domestic problems have specific effects on foreign policy behavior? to move beyond simple, impressionistic answers to such questions, it was necessary to begin collecting data on the attributes of nations and events characterizing their interactive behavior. it is possible, in retrospect, to trace three broad data collection efforts that sought to provide the evidence necessary to facilitate the scientific study of international processes: data collections that focused on the quantitative and qualitative characteristics of (1) nationaj attributes, (2) major conflicts and wars, and (3) interactive events within and between nations. intriguingly, these projects all began within a year or two of one another and spread rapidly across the scientific geography of the united states and even abroad. ddir i: national attributes dimension historical background. in the late 1950s and early 1960s karl w. deutsch at yale university was arguing for the use of data to confirm or disconfirm hypotheses about international and cross-national politics. he had demonstrated the practicability of the search for such data and their analytic value in his studies of nationalism and social communication (deutsch, 1953) and political community and the north atlantic area (deutsch et ai, 1957), and in a series of articles (most notably deutsch, 1960, 1961) that showed how important questions were not being addressed because of the absence of valid, comparable indicators based on reliable (that is, replicable), impersonal, and quantitative data. the year 1963 was a watershed for these innovative ideas. at the yale data conference (merritt and rokkan, 1966) held in september international scientific researchers gathered to discuss systematic means to compare nation-states, outline organizational efforts to further such research, and learn at first hand of three major data-collection activities reaching fruition in the united states. • russettetal.(1964): yale political data program. with financial support from the national science foundation, deutsch and harold lasswell created the yale political data program, which in 1962, under the direction of bruce m. russett, had begun the crossnational collection of political, social, and economic data (see deutsch et al., 1966). its immediate result was the world handbook of political and social indicators, known as "world handbook i"; and in later years world handbooks ii and m appeared (taylor and hudson, 1972; taylor and jodice, 1983). • banks and textor (1963): cross-polity survey. arthur s. banks, a political scientist, and robert b. textor, an anthropologist, had combined forces to classify 115 polities according to 57 sets of carefully operationalized criteria. • rummel (1964): dimensionality of nations. at northwestern university, initially as a component of harold guetzkow's inter-nation simulation (ins), rudolph j. rummel had compiled data characterizing nation-states (see rummel, 1979). (he also—and we shall return to this later—systematically searched the new york times index and other sources to record domestic-political and foreign-conflict events.) years subsequent to this burst of creativity saw three important developments. the first was the growing use of data already collected to examine theoretically interesting propositions. second, the efforts of the 1960s were continued and expanded during the 1970s. particularly important here were (1) taylor et al.'s world handbooks ii and iii, (2) gurr's (1974, 1978) research on polities and segmental groups, and (3) the national characteristics assembled by the correlates of war project directed by j. david singer (see singer and small, 1972; small and singer, 1982). third, researchers became more sophisticated in both measurement and data-collection techniques. they also met increasingly frequently to discuss their work, as well as to exchange preprints and sometimes tables of data. a "community" of quantitative international relations (qip) scientists was emerging. these years of an emerging qip community were heady ones for scientific advances. data-generators and users did not simply rest on the intellectual platforms given them in the 1950s by karl deutsch, harold guetzkow, harold lasswell, and others, but used them to push understanding forward. the analyses themselves were not always elegant, but nevertheless established clearly at least two points. first, "the activities of nations," in rummel's (1966:205) words, "are highly patterned behavior," and, second, it is both possible and intellectually profitable to establish data banks to help ascertain what these behaviors are, how they are structured, and what impact this structured behavior has on such vital lassist quarterly issues as war and peace. demographic, and related variables. by the mid-1980s the need for aggregate data was growing but so was the reality that past data sources were aging and not being kept alive. it is thus not surprising that research scientists saw in ddir the opportunity to resuscitate and significantly improve this field. a coordinated effort by both data generators and data users assembled a research design on national attribute data and, through ddir, submitted it to the national science foundation. nsf support enabled ddir i to carry out the following projects: ddir i-l. correlates of war civil war datasets. principal investigator: j. david singer, university of michigan. updating and revalidating the correlates of war (cow) dataset on civil wars, 1816-1988. ddir 1-2. correlates of war national capabilities dataset. principal investigator ted robert gurr, then university of colorado, boulder, and now university of maryland at college park. cooperation with j. david singer, university of michigan, to produce an integrated dataset, 1816-1988, on the cow project's national capability variables (population, urban population, iron/steel production, energy consumption, military expenditures, and military personnel) plus government revenues and expenditures for all states at one time members of the central-state system and, insofar as possible, peripheral systems. ddir 1-3. correlates of war dyadic relationships dataset. principal investigator: j. david singer, university of michigan. collaboration with michael wallace, then university of british columbia and now simon eraser university, to update and revalidate the cow dyadic relationship dataset, 1816-1988, on shared membership in foreign alliances, diplomatic representation, and shared membership in international bodies. these projects are now for the most part complete, reports included in ddir's newsletter, ddir-update, and the datasets sent to the inter-university consortium for political and social research (icpsr) in ann arbor, michigan, for access to the scientific community. ddir i: international conflict dimension historical background. just as the national-development datasets had their predecessors in the many efforts of individual scholars, government agencies, and international bodies, so, too, current efforts to generate data on international confiicts can look back on a tradition of earlier projects. though flawed, these early studies pioneered the path that modem researchers, with their richer sources and computer-based operations, continue to treat. among these are four deserving particular attention: • woods and baltzley (1915): is war diminishing? frederick adams woods, an eminent biologist, and a young political scientist, alexander baltzley, provided a list of wars and their participants for most of the major european states for 145019(x) (and back to 1 1(x) in the cases of england and france). chapters on each of eleven countries indicated the years of initiation and teimination of national war, and hence the duration of war, and statistical graphs showed national percentages of interstate, imperial, and civil wars. • sorokin (1937): social and cultural dynamics. pitirim a. sorokin identified "almost all the known wars" for the major european states from antiquity to 1925—including internal disturbances as well as interstate, civil, and imperial wars. he gathered the dates of initiation and termination, the war's duration for each major state, and estimates of the average army size, percentage of casualties, and total number of casualties for each state. ddir 1-4. political structures dataset principal investigator: ted robert gurr, then university of colorado, boulder, and now university of maryland at college park. for all international-system members, 1816-1985, development of a complete and updated dataset on rdgime characteristics, collaboration with mark lichbach, university of illinois at chicago, to transform into time-series data coding on each authority dimension. ddir 1-5. world handbook national attributes dataset. principal investigator charles lewis taylor, virginia polytechnic institute and state university. expanding and filling in yearly data, 1950-1985, on readily accessible economic. • wright (1942): a study of war. for 1480-1936 quincy wright hsted balance of power, civil, imperial, and defensive wars involving each major or minor state. the dataset includes for each war its initiation and termination dates, identity of participants, their individual day of entry, and number of important battles. wright also assembled data on the frequency and types of battles, casualties, and internal systemic disturbances. • richardson (1960): statistics of deadly quarrels. lewis fry richardson's compilation of conflict data includes all deadly quarrels—imperial wars, civil wars, and other forms of domestic confiict^between 1820 and 1949 that caused death of spring 1991 humans. the war is the unit of analysis, and wars are organized by magnitude and then chronologically within magnitudes. the data on each war include the magnitude, dates of initiation and termination, participants, identity of the initiator, and ostensible cause or issue at stake. each of these data collections, of course, had problems. that by woods and baltzly was remarkably incomplete, especially by modem standards. sorokin gave no explicit operational criteria for interstate, civil, and imperial wars. wright's use of legalistic criteria for including and excluding wars was questionable. richardson's research raises serious questions about the reliability of the data and validity of the categorical codings. but each was a significant beginning. each, in a sense, set the groundwork for the subsequent ones (although the lateness of richardson's publication makes it easy to ignore the fact that his research efforts were contemporaneous with those of wright). more to the point, however, each has been superseded by modem datasets that have not only initiated major data collections but also spawned the use of those datasets to conduct significant, empirical research on questions of war and peace. some recent intemational conflict datasets are: singer: correlates of war (cow). j.david singer's intellectual focus was on the correlates of war, those factors that seemed to covary and thus be associated with the occurrence, duration, and magnitude of wars. his cow project, which began formally in 1963 under the auspices of the national science foundation, initially concenu^ated on obtaining information on the attributes of the intemational system that theorists had argued were the "cause" of war. cow researchers systematically culled historical texts to obtain as complete a listing as possible of all wars since 1815, together with major identifying characteristics, such as the number of participants, battle deaths, and durations (singer and small, 1972; small and singer, 1983). subsequent data collections expanded cow's horizons—to such independent variables as population, iron and steel production, energy consumption, military expenditures, and military personnel, and, with several colleagues, to forms of intemational behavior, such as crises. siverson and tennefoss (1982, 1984): interstate conflict unaware of cow's shift in focus and data gathering, randolph siverson and michael tennefoss independently developed a dataset on major intemational crises since 1815. they classified into three types the implicit/explicit level of war: threats to use force, unilateral uses of force, and reciprocated military interactions. levy (1983): great power war. jack levy's dataset overlaps that of cow but extends the latter back to 1495 for the great powers. concentrating on interstate, great-power wars (excluding civil and imperial wars) which had more than 1,000 battle deaths, it provides data on their magnitude, severity, and intensity. overlapping but not identical with these efforts were several other projects. robert butterworth (1976, 1980), michael brecher and jonathan wilkenfeld (1982), and hayward r. alker, jr. and frank l. sherman (1982; cf. sherman, in progress) collected data on significant attributes of intemational crises since world war 11. along a somewhat different but related dimension, frederic s. pearson (e.g., 1974) collected data on international interventions in the post-world war ii period. as was the case along the national development dimension, participants in the ddir enterprise saw a clear need to update and expand data on the intemational conflict dimension. major insights had come from the analyses based on the earlier datasets. but, as was true with respect to the national development dimension, inadequate coordination had led to duplication and incomparability. members of the emergent ddir community responded to the need by preparing research proposals that eventually formed a component part of ddlr's main task: research to develop adequate measures of intemational conflict and systematically to collect relevant data. nsf funding supported the following projects: ddir 1-6. great-power war dataset. principal investigator: jack s. levy, then university of texas and now rutgers university. collaboration with t. clifton morgan to revalidate and fill in missing data in levy's dataset on participation, casualties, and initiation/termination dates for all wars among great powers, 1495-1815. ddir 1-7. international crisis behavior dataset. principal investigator: jon wilkenfeld, university of maryland at college park. revalidating the intemational crisis behavior (icb) dataset, 1929-79, and updating it through 1987. ddir 1-8. interstate war catalog. principal investigator: claudio cioffi-revilla, then university of illinois at urbana-champaign and now university of colorado, boulder. completion of a master catalog comparing (with reliability indicators) existing datasets on interstate wars (see cioffi-revilla, 1990). ddir 1-9. interstate war dataset. principal lassist quarterly investigator: j. david singer, university of michigan. updating for 1980-88 the cow dataset on the initiation of interstate wars, participation (in nationmonths), and casualties; defining and coding additional variables for 1816-1988 on interventions by third parties, war phases, monthly casualty rates, and characteristics of war terminations. ddir i-io. interventions dataset. principal investigator: frederic s. pearson, university of missouri-st. louis. filling in the dataset on unilateral, multilateral, and international organization interventions for 1816-1988 on interventions by third parties, war phases, monthly casualty rates, and characteristics of war terminations. these projects are now for the most part complete, reports on most included in the ddlr's newsletter, ddir-updale (and one, cioffi-revilla's [1990] interstate war catalog, published), and datasets sent to the icpsr for access to the scientific community. ddirh: event data a second, and equally important, ddir 1 activity was planning future data-gathering activities on two dimensions: interstate events and international political economy (ipe). for the field of international relations to keep up with and anticipate data needs deriving from new theoretic growth requires imaginative and sustained attention to such matters as conceptualization, indicator validity, and collection procedures. ddir's organizational goal was to hold separate sets of conferences on the two dimensions, at which active scholars would discuss needs, priorities, and procedures. the long-term hope was that conferences would produce specific research programs which could be developed for future funding. accordingly, with respect to the event-data dimension, planning conferences took place in may 1987 in columbus, ohio (hermann, 1987), november 1987 in cambridge, massachusetts (alker, 1988), and march 1990 in chicago, illinois. what emerged was a two-year proposal to the national science foundation that included researchers at seven different academic institutions who will carry out distinct but generally integrated research projects. nsf funding, awarded in january 1991, permits the realization of ddir ii. and, as in the past, the merriam laboratory for analytic political research, located at the university of illinois at urbana-champaign, serves as ddir ii's administrative umbrella. from political arithmetic to event-data research narratively oriented diplomatic historians generally view the course of international relations as a series of events—d6marches, protests, treaties, crises, wars, conferences, and the uke. an event in this sense is an occurrence that stands out against the gray background of everyday living. in principle an event is a discrete unit of action, with its own beginning and ending points. in practice we often view events as nested sequences of yet smaller events. thus an historian may view the francoprussian war of 1870-71 in the light of inter alia bismarck's wars against denmark and austria, the ems dispatch, declaration of war, miutary hostilities, siege of paris, conclusion of a peace treaty, and such consequences as indemnification, territorial transfer, and formation of the german empire; and each of these in turn comprises a congeries of lesser events. is there another, more systematic, way to look at international events? analysis have devised various ways to study the events they define as important in our individual and social lives. indeed, modem statistics fmds one of its main roots in the "political arithmetic" used in the 17th century by john graunt and william petty to examine mortality tables. sickness and death are individual events. and yet knowledge of how many of a society's members suffer from particular illnesses and die of particular causes tells us something about the society itself, and enables us to predict the need for medical services and the proper price for insurance. similar considerations led petty and other social philosophers to argue for the collection of criminal statistics (walker, 1971; coumann, 1973), and the occasional monarch or cabinet minister undertook a survey from time to time. such studies had individuals as their unit of analysis. not until the late 19th century, with the flowering of labor unions throughout the industrialized west, did government agencies begin to gather data on social events. the target was the strike or lock-out, industrial disputes leading to stoppage of work in some fum or branch of industry. nor is it surprising, given the general attitude then prevailing toward labor unions as a whole, that data on strikes took on the character of criminal statistics (international labour office, 1926). in the united states, the department of labw's bureau of labor statistics combed newspapers and other sources to identify work stoppages, sent questionnaires to key participants to ascertain the dimensions of these events, and reported on the number of strikes, workers involved, duration, days idle, and so forth (see u.s. department of labor, 1976: 195-202). the 1930s saw three major social scientific efforts to collect data on social events. the fu-st, described earlier, focused on aspects of wars. a second was harold d. lasswell's (1936) intentionalist/instrumental view of pohtics in terms of "who gets what, when, and how." the third was lasswell and blumenstock's (1939) study of social unrest and world revolutionary propaganda in chicago from 1919 to 1934. they recorded the number spring 1991 and characteristics of communist-sponsored meetings, demonstrations, parades, and other social gatherings; strikes; group and individual complaints about violations of civil rights; and evictions, foreclosures, and arrests of "radicals." lassweu and blumenstock concluded among other things that communist propaganda was most successful during tough economic times and when it incorporated american symbolism instead of harping on soviet accomplishments. but at the same time, by giving hardstrapped citizens an outlet to vent their frustrations and business a scapegoat to blame for the country's economic woes, communist agitation worked ultimately to deflect any truly revolutionary spirit and hence to strengthen the c^italist system. two decades later scholars of international relations renewed their interest in systematically studying events. one starting point was growing concern with processes of political development and the place of violence in them. cross-national studies using data for a single year (that is, synchronic) aimed at discovering the correlates of unrest and violence; longitudinal (diachronic) studies traced patterns over time among some more limited set of countries. the nation-state was the unit of analysis. researchers tabulated such events occurring within a state's boundaries as demonstrations, coups d'etat, and revolutions. another starting point for event analysis centered on foreign-policy decision-making. scientists conducting simulations of international processes—whether using people only, computers only, or some combination of the two—discovered they needed hard data both to feed into the simulation itself and/or to check the realism of their findings. eventually the focus shifted from the nationstate as the unit of analysis to interactions between pairs of nation-states: ongoing processes such as trade and diplomatic exchanges as well as more or less distinct occurrences such as a threat or militarized intervention. from there it was a short step to taking seriously the new emphasis on the international system qua system (kaplan, 1957) and tabulating the attributes of that system as a whole and the events taking place within it. still a third and doubtless the most important starting point was a growing concern with international crises and war. in the nuclear age, the possibility of war cannot be taken lightly. if analysts had had the correct tools, scientists asked, could they have recognized the pnabable outcome of the sequence of events in m id1 9 1 4 or in the 1930s early enough to have prevented the outbreak of war? is there some means to ascertain when international crises are reaching the boiling point? what steps can governments lake to de-escalate crises? answers to such questions seemed to require detailed information on the course of events occurring in the global arena. progress in developing event datasets initial efforts to assemble data about the events of nationstates electrified the discipline of international politics. they were, broadly speaking, of two types. first, global studies defined events of interest, specified coding rules, and, in such universal sources as the new york times or facts on file, coded every single occurrence of such events. (regional studies pursued the same procedures but focused jmimarily on regional issues and sources.) second, event-specific studies proceeded from the opposite direction. that is, they identified critical events of interest, such as the suez crisis of 1956, and searched a wide variety of newspapers and historical treatises to describe, in detail, their characteristics and the chronology that preceded the key event. not only did these event studies set the standards that subsequent researchers would use and contend with, but they resulted in empirical studies that opened scientists' minds to new modes of research. as the data movement captured the field of international politics a series of datasets were compiled by different researchers. limitations of one dataset for a new research question being posed led to the development of new datasets. a glance at the history of this evolution suggests at least seven major compilations. • dimensionality of nations (don). rummel, as we saw earlier, generated one of the original collections of national-attribute data. he also focused his research on interactions within and among states. he used five sources to assemble data for 1955-57 on the domestic-politics and foreign-confiict behavior of 77 nation-states (rummel, 1964, 1967, 1972). among other things exdn tabulated the presence or absence of guerrilla warfare, number of assassinations, and seven other domestic conflict events. rummel's thirteen foreign-confiict variables were, besides the presence or absence of military action, the number of anti-foreign demonstrations, negative sanctions, protests, countries with which diplomatic relations were severed, ambassadors expelled or recalled, diplomatic officials of less than ambassador's rank expelled or recalled, threats, wars, troop movements, mobilizations, accusations, and people killed in all forms of foreign conflict behavior. • world event interaction survey (weis). at roughly the same time charles a. mcclelland initiated at the university of southern california an unrelated data enterprise. this collection focused on the events, or interactions, that took place over time between pairs of countries (and in this sense was not dissimilar to the foreign conflict events coded by rummel). weis consisted of a very detailed set of coding categories (63 mutually exclusive and exhaustive categories) designed to capture the type of 10 lassist quarteriy hostile or cooperative action that one country directed toward another, but not the intensity of hostile or cooperative behavior. relying on reports published in the new york times, mcclelland and his colleagues (1971) recorded such acts in terms of initiator, target, type of act, and date of occurrence, covering the period after 1946. the extensive historical chronicle of interstate interactions that resulted made it possible to observe patterns in the activities of states and to determine uie degree to which special patterns preceded major crises or wars. that such a monitoring system might facilitate forecasting of the onset of future crises was an integral part of mcoelland's overall research design. • conflict and peace databank (copdab). edward e. azar's particular interest in recurring middle eastern conflicts led him to develop a new and somewhat differently focused event dataset (see azar, 1970, 1980a, 1980b; azar and sloan, 1975; azar and havener, 1976; azar and lemer, 1981). building on earlier work by robert c. north, lincob e. moses, and their collaborators (moses el al., 1967; choucri and north, 1975), azar defined events as occurrences between or within nation-stales that were sufficiently distinct from the constant flow of "transactions" (such as trade or mail flow) to stand out as reportable or newsworthy against this background. the coding categories were very similar to those of mcclelland (see howell, 1983; mcclelland, 1983; vincent, 1983), but the sources azar used for coding the events went far beyond the new york times to include a variety of international as well as local reporting sources. • comparative research on the events of nations (creon). yet another important event dataset, developed by charles f. hermann and his colleagues (1973) primarily at the ohio slate university, sought to examine the correlates of foreign policy behavior. it focused on events that characterized different foreign policy positions of slates. the coding categories were therefore somewhat different from those developed for weis or copdab. further, since the central question concerned the relationship between certain attributes of states and types of foreign policies, extensive and costly time-series were not necessary. creon rather provided snapshots at various points in lime of the foreign poucy behaviors of states. • world handbook of political and social indicators (1983). in the late 1970s, charles lewis taylor and david a. jodice (1983) significantly expanded the data-gathering approaches originally developed, as noted above, at the yale political data program by russett et al. (1964) and taylor and hudson (1972). world handbook iii provided daily event data for domestic political events only, for 136 nation-states for 1948-77. the event categories include political unrest (e.g., protests, riots), state coercive behavior (e.g., government sanctions, political executions), and governmental change (e.g., elections, executive transfers). the number of deaths from events involving domestic violence is also recorded, and additional codings for event duration, intensity, scale, and impact are included for events from 1968. world handbook iii also separately compiles for each state statistical indicators of poutical, economic, and social change, thus helping to define the broader context within which coded events these five event datasets, despite their apparent differences, share two important similarities. first, the definition and coding of an event are in terms of actors (national or subnational) and actions; and events are classified into a set of predetermined categories which provide descriptors of the event. second, they pursue global coverage, that is, they are concerned with the entire international system. these were not, of course, the only event datasets to emerge after the 1950s. for the political instability data bank, ivo k. and rosalind l. feierabend (1966a, 1966b) codified 28 types of events occurring for 1955-61 in 84 countries. in his comparative study of civil strtfe, ted robert gurr searched standard sources for the occurrence in l%l-68 of civil violence in 1 19 polities; this data collection, which he analyzed in various forms and made available to the scholarly community, provided the empirical basis for gurr's impwtant, prize-winning theoretic work. why men rebel (gurr, 1970). two other datasets are event specific and thus differ from the others in significant ways. in effect, two levels of "events" characterize these datasets. one is the identification of a key event, for example, an international crisis. the other is a minute examination in considerable detail of all preceding events, where "event" in this second instance is considerably more fine grained. • behavioral correlates of war (bcow). the bcow dataset, developed by russell j. leng as an offshoot of the correlates of war {h^oject, starts with leng's earlier data on miutarized interstate disputes (mid)—defined in terms of disputes in which parties on both sides threaten, display, or use miutary force — but focuses only on a subset of mwe intense disputes, called militarized crises (leng and singer, 1988). it then provides for the time period prior to each militarized crisis a fine-screened description of all events. unique features of the bcow coding scheme (beyond the core coding of who does or says what to whom and when) include: location of each event; spring 1991 duration and variations in intensity of multi-day events; assignment of physical events to one of 103 categories of military, diplomatic, economic, or unofficial behaviors; and detailed analysis of sequential verbal interactions (allowing identification of bargaining strategies). this fme-grained coding of verbal actions allows for a detailed analysis of interstate bargaining and the development of an "hierarchical choice tree." • sherfacs. using criteria to select and merge conflict cases from the facs dataset (farris, alker, carley, and sherman, 1980) and nearly 40 other studies, frank l. sherman's sherfacs produced a combined file of 730 international disputes and 980 domestic quarrels that provide data on, among other things, the identification of conflict phases, means of referrals to management agents, and nature of actions taken by all parties (see alker and sherman, 1982; sherman, 1987a, 1987b, in progress). sherman then developed a phase structure for domestic quarrels similar to the gascon structure for international conflicts. some related datasets were mentioned earlier: butterworth (1976), brecher and wilkenfeld (1982), and pearson (1974). then, loo, empirical studies of conflict management, such as sherfacs, have a rich tradition: ernst b. haas's (1968) disputes referred to the united nations for management, joseph s. nye's (1971) added conflicts referring similarly to regionaj international organizations, the joint effort by haas, robert butterworth, and nye (1972) added to the existing set three new types of conflicts—interstate disputes in which some kind of international wganization, e.g., the united nations security council, sought involvement; civil strife in which one side of the dispute enjoyed the support of another government; and "non-managed" interstate conflicts in which fatalities occurred—and the cascon phase structtire developed by bloomfield and leiss(1969). these event-data projects saw enormous use by scholars. this was particularly the case with azar's copdab, which continued until 1979 to collect data, and, like other event-data collections, was made generally available to users. but these fffojects—and hence the fundamental idea underlying event datasets—also came under fu-e by critics, both friendly and hostile. complaints ranged from the usefulness of particular sources, such as the new york times, to the modes of categorizing the data. the level of hostihty had multiple effects. it diminished funding and shifted intellectual concerns. it discouraged previous and emerging event-data researchers from either undertaking new collections or updating the old ones. the scientific progress of the 1960s soon began to languish. but, at the same time, challenging the past value and uses of event data encouraged researchers to spend time thinking through various dimensions of previous projects, exploring new ideas, and, particularly, adapting their research plans to take advantage of modem computational capabiuties. ddir 11: developing new event-data research ddir's three event-data conferences sought first of all to assess the state of the art, then to review new data priorities, and fmally to develop an effective research strategy. several considerations shaped a decision to pursue a mixed strategy: the need to (1) generate a rich and general, core dataset; (2) improve the capabilities of key specialized event datasets that akeady exist; (3) enhance software so as to minimize the time and cost of expanding datasets in the future; and (4) explore the possibilities for new styles of event-data research. enhancing existing and generating new event datasets. if we are to enhance the quality and quantity of some existing datasets, which ones should they be? our survey of the literature (mcgowan et ai, 1988) together with a study of each event dataset's time-span and comprehensibility across a wide range of theoretically interesting issues strongly suggested a central focus on the copdab file. not the least reason for this is the fact that, of the five global event datasets—ix)n, weis, copdab, creon, and worid handbook iii— copdab best met the combined criteria of past scientific usage, availability over a long time series, and attention to a broad range of new styles of computeraided, event-data research (starr, 1987). other factors included copdab's compatibility with case-oriented datasets (most notably bcow and sherfacs), the needs of those initiating regional event datasets, copdab's apprqjriateness for testing new software, and, by no means least significantly, the fact that the center for international development and conflict management (cidcm) at the university of maryland at college park was planning to update and expand the copdab dataset. thus the global event-data system (geds) project at maryland became the natural focal point for organizing ddir ii's core data-generation part. the cidcm's research team will establish geds for computer-assisted identification, abstracting, and coding of daily international and domestic events, as reported primarily in comprehensive, on-line news sources such as the reuters news service. geds thus aims at developing a core event-data stream from 1979 forward it will include: • the actions vis-d-vis each other of (1) nation-states, (2) major nonstate communities, and (3) international organizations, • detailed event summaries and coding, including 12 assist quarterly direct quotations and cross-referencing, and • information allowing users to access those full-text source articles which are available on-line. geds software will permit partially automated, continuous updating after 1990 of the core event-data stream. in the discussion that follows, the term geds refers to the event-data stream generated by using maryland's computer-assisted coding procedures on on-line news sources. each of the projects described below produces a specialized dataset based on geds. ddirii-l. university of maryland: updating and extending existing datasets. as part of its larger geds effort, the maryland team—john l. davies, ted robert gurr, and chad k. mcdaniel—will update to 1990+ the existing copdab dataset, and incorporate updated weis and, as they become available. world handbook iii (and bcow and sherfacs) event data. the updated dataset will be compatible with each of these previously-coded datasets, but expanded to include new foci (e.g., inclusion of nonstate actors) and new sources made available through computer-assisted coding. ddir n-2. american university: foreign policy behaviors of southeast asian states (sas). llewellyn d. howell will use the geds computerassisted procedures on regional sources to produce a data bank on 10 southeast asia states. the sas event-data stream, lo be added to the maryland core event-data stream, will thus enrich the latter and provide a check on the comparability of global sources vs. regional sources. ddir n-3. university of kansas: kansas eventdata sources (keds) for central europe and the middle east philip a. schrodt, ronald a. francisco, and deborah j. gemer have two tasks. first, they will extend their existing software for automated coding. using the geds files as inputs, the current software automatically generates weiscoded data. resources permitting, the software can be expanded to produce copdab-coded data. second, the kansas team will assemble a high-density, international, event dataset for central europe and the middle east. it uses specialized journals and government publications around the world to increase regional coverage without the time and expense involved in working with regional journalistic sources such as newspapers. like howell's sas project, the use of regional sources will provide the basis for comparing alternative, global vs. regional sources of events. ddir n^. middlebury college: behavioral correlates of war (bcow). for 40-55 militarized crises occurring in 1979-90, and starting with the core data provided by geds, russell j. leng will apply bcow data-collection procedures to produce a finescreened dataset. the bcow coding manual specifies as many as 103 descriptors of each action (such as alert, mobilization, or evacuation) that could take place during a militarized crisis. each such event action is categorized according to the date of occurrence, actor, target, location, whether the actor was acting unilaterally or with another state, and "tempo" of the action. ddir n-5. miami university: nonstate actors in interstate connicts (sherfacs). frank l. sherman at miami university of ohio will enhance and bring up to date the sherfacs dataset, which comprises fine-screened accounts of several kinds of episodic conflict situations. inclusion is global, but limited to international conflicts and domestic quarrels, especially those involving collective management (e.g., un mediation) and nonstate actors. the expanded event summaries generated by geds will increase the number of international conflicts and domestic quarrels that will be coded using the sherfacs template. and, like the bcow dataset, the sherfacs dataset will augment the analytic capabilities inherent in the expanded copdab dataset to be developed by cidcm at the university of maryland. ddir n-6. massachusetts institute of technology: data development for interpretive analysis. hayward r. alker, jr., at mit, will develop methods for the interpretive analysis of detailed event summaries by adding narrative depth and varieties of interpretive perspectives for specific conflict episodes in the geds dataset. the three data components to be studied are (1) explicidy coded weis/copdab/bcow/ sherfacs event data, (2) humanly constructed narrative summaries of each event, and (3) quotations attributed to principal actors/ interactors of the event being described. in addition, original and secondary source stories will be made conveniently accessible, possibly as part of each record, for the purposes of detailed textual and interpretive analysis of both quantitative and qualitative, political data. these various data-collecting activities can significantly improve the quality of research in the field of quantitative and textual international politics. first, they will bring up to date and expand the more important event datasets identified by publications and by quantitative and textual scientists. second, they will provide procedures fw routinizing future such event-data collections. this will sharply reduce the need to turn to funding spring 1991 13 agencies every five years or so in the search for new support to update the datasets. third, they aim at achieving an integrated event dataset. interaction among the principal investigators through ddir's aegis can ensure that interchangeable datasets are in the public domain. fourth, the coordinative thrust nevertheless permits maximum flexibihty among these principal investigators to carry out their individual research strategies. software developments to aid data collection and analysis. recognizing the need for a core data-collection effort such as geds was only one step. researchers in recent years also began to appreciate the important role that computerized methods could play. with major international news sources, such as reuters, associated press, and united press international, as well as local news reports (as translated, for instance, by the fbis reports) either now or soon to be accessible on-line, the retrieval of source stories begs for automation. moreover, the enormously expanded storage capacity, processing speed, and programming flexibility at the microcomputer level now makes it possible to develop an event-coding system which sacrifices neither the comprehensiveness of global coding efforts nor the depth and diversity of coverage of the episodic coding projects. dder ii proceeds from the conviction that the development of computerized methods for the collection of data is not only a desirable but a necessary innovation. it includes several projects in this area: ddir ii-7. university of maryland: computerassisted and partially-automated coding in geds. with a grant from ddir and backing from their institution, the maryland team has developed and tested a preliminary version of software for computerassisted entry, coding, and editing of reuters on-line source stories to produce geds event records. as a significant product of its software development, the team will set in place at cidcm a process for continuously coding geds records. ddirii-8. university of kansas: partiallyautomated procedures for the keds machine coding systems. the keds machine-coding systems will be enhanced to permit continued development of event-data generating software, which will use inexpensive, machine-readable data sources and personal computers. the keds-x rule-based coding system will (1) add a practical english parser to handle grammatical tasks associated with text analysis, (2) experiment with non-english source text, and (3) implement a parallel processing network for increased coding speed. schemes for coding timedependent datasets, such as bcow, will also be explored. the software developed at kansas will provide inexpensive, up-to-date, and easily customized datasets on international and domestic confiict and cooperation, and will also aid in developing the partially automated coding software being written by the maryland team. in addition, machine-assisted coding procedures will be implemented by two other projects. howell's sas project will make extensive use of the computer-assisted (and ultimately partially automated) methods that the maryland team will develop. some of these methods are even now in use in the sas project. in addition, leng's bcow project will use machine-assisted coding software recently developed as a part of that project. this software is specifically designed to use as input for the detailed data records produced by the geds project. the software component of ddir ii also focuses on softwarefor data analysis. included are four projects at the participating institutions as well as an evaluation to be carried out in illinois: ddir n-9. massachusetts institute of technology: computerized textual and interpretive analysis of conflict episodes. alkcr is exploring software development for the interpretive analysis of event histories. this will allow subsequent validityand reliability-oriented comparisons of original sources, geds codings, human narrative summaries, speech fragments, and such computational interpretations as would be produced. central to redefining available software routines for computational text analysis in the schank-abelson tradition are developing and implementing an "event description framework" motivated by lasswell's work on interactions, and a translation scheme for "filling in" this framework using, in particular, sherfacs data. the interpretive routines would then operate on this framework to produce event interpretations computationally. ddirn-10. middlebury college: extension of computerized procedures for the analysis of bcow data. leng is modifying and enhancing two currently existing software packages developed for analyzing bcow data. because of the richness of bcow coding categories, software is the only efficient way for aggregating the data for subsequent analyses. one program, crisis, permits users to select, count, and scale events along various dimensions. another, influence, is designed specifically for analyzing crisis bargaining. both programs currently exist only in the environment of a (vax) mini-computer, and the goal is to increase their functionality and availability by converting them to microcomputer environments. assist ouarterty ddirii-11. miami university: computerized preparation of sherfacs data for interpretive analysis. sherman will also explore means to fit the sherfacs coding schema into lasswellian frames, which alker proposes to use for interpretively describing conflict episodes. computer-assisted or partially automated coding sequences are needed to transform into lasswellian categories the sherfacs information (and, by extension, the associated event summaries and event categories of geds). the software will be compatible with the geds datacollection system. ddirn-12. university of maryland: geds user software. the maryland team will develop geds end-user software for browsing, data selection, temporal and spatial aggregation, graphic display, and to interface with related databases with full-text sources as well as statistical and interpretive software packages. the merriam lab is considering the possibility of enhancing the utility of software developed by the various projects. for example, its numerous computer language compilers (e.g., c, pascal, lisp) for several different operating system environments (e.g., ibm, macintosh, umx) are available for coordination tasks; and it can develop simple macros designed to link processing across the different executables so as to reduce the amount of time needed by users to perform multiple research tasks. though focusing primarily on data collection, ddir ii can creatively enhance software facilities that expand the usage of such data. to some measure it banks on enhanced hardware and software technologies. an ideal and very "futuristic" automated system for handling unstructured data would provide multiple interpretations of one unstructured data stream—just as ordinary citizens, political activists, and scientists working within different research traditions while looking at the same ordinary language texts might draw different interpretations. while several experimental parsers already exist, more basic research is needed before they can become reliable components of a data development infrastructure. although a multiple-interpretive parser for ordinary language text will probably not be available for some time, we recognize the need to anticipate future technological advances in the more modest coordination outlined here. future technological developments undertaken by other researchers will eventually permit some further extensions such as semi-automated technologies for processing unstructured, that is, ordinarylanguage, text. ddir ii itself can also contribute to enhancing the hardware and software technologies that are needed. it is also essential, however, to look more closely at the degree to which coding judgments stray from case-study level understandings. the merriam lab will thus include some general comparisons across the basic event datasets (copdab, weis, bcow. and sherfacs) to assess their relative validity against original source texts regarding, say, the crisis leading up to the persian gulf war. the point is not that these datasets are invalid, but rather that their quality will reflect coders' perceptions, and that, therefore, independent analysts would have to take this fact into account in using the data for their own research. toward the future: ddir iii on international political economy data about a dozen years ago, international relations scholars rediscovered the importance of international political economics (ipe). it had of course remained alive and well in some quarters, particularly in great britain where the field of political economy was nurtured some two centuries ago. but it tended to interest economists, not political scientists, just as such issues as social change in developing countries tend to interest sociologists. political scientists, even those concentrating their studies on international relations, by and large treated economic considerations as peripheral to the main struggle for national power and global order. especially in recent decades the main thrust of their scholarship and instruction had been power politics, with its emphasis on military security. east-west confrontations, and guiding the political development of new nation-states. the long-standing tradition of political economy paled in the perspectives of all but a few of those who were shaping the post1945 directions of international political research. the renewed interest in ipe caught empirical researchers in a state of acute embarrassment as we have seen, qip scientists had focused on national characteristics, conflict, events, and a wide variety of other topics. by the end of the 1970s, when they looked in the larder of systematically evaluated ipe data, they found the cupboard bare. a curious sequence of events then took place. the availability of ipe-relaied data from the united nations and other agencies posed a delicious dilemma. on the one hand, a wide variety of such data sources existed but, on the other, they were of mixed quality for the type of analysis condikted by qip scientists. the data were not always compatible, nor did they address some of the key questions relating to the broad domain of ipe research. this led to dismay in some circles. perhaps scientists had become too accustomed to readily available, reliable, and paradigmatically similar data from such agencies as the icpsr, the european consortium for political research (ecpr), and the zentralarchiv fur empirische spring 1991 sozialforschung at the university of cologne. the absence of any comparable storehouse of ipe data may have led these scientists to ignore the fact that such rich data sources were not the product of a single day's labor. the virtual lack of appropriate data had different consequences in other circles. some researchers—possibly following admiral david farragut's injunction during the american civil war to "damn the torpedoes: full speed ahead!"—wrote treatises based on existing data sources, however disparate they may have been. the predictable result was sharp criticism from their colleagues, and especially from those who were fundamentally disposed to favor data-based research. (those opposed to the basic idea of such research merely found their predilections confirmed!) viewed from a more distant perspective, such studies could be described as courageous but flawed efforts to make sense out of a complicated field. still other researchers explored the means to generate new, more sophisticated ipe data (bomschier and heintz, 1979; groenick, 1988; miiller, 1988). what they quickly discovered is that such reliable data, especially those encompassing long time series, are as scarce as the proverbial hen's teeth. this discouraged the faint of heart. the result was that, though many researchers called for better ipe data, few proved willing, in the words of the famous american challenge, to put their money were their mouths were. the ddir community held two workshops—in october 1987 in new haven, connecticut (russett, 1988), and april 1988 in tempe, arizona (mcgowan, 1988; pollins, 1988)—to address three questions about data important for studying pohtical dimensions of international economic transactions. • what is the current status of ipe data? of particular importance are their availability and quality, and differences among datasets generated by national and the international institutions, commercial firms, and university research institutes. the concern is a very pragmatic one: to what extent can qip researchers interested in a broad range of ipe issues actually use existing datasets? • what ipe data do active researchers need? two issues are problematic here. first, given unhmited resources, including funding, computational facilities, and qualified research assistance, which datasets are the most significant in terms of probable intellectual or scientific payoff? second, given the fact that such resources are not unlimited, how can we prioritize among competitive claims of significance? • how can we enhance ipe data development? the assumption sometimes seems to be that desired datasets will drop from the clear blue sky. to the cchitrary, they must be developed. the question thus focuses on two issues—especially in an international framework. one is. how can we enhance institutional arrangements to facilitate data development? the other is. how can we support or persuade leading ipe researchers to take on leadership roles in these endeavors? ddir's plan to initiate a third research phase on ipe data remains in its pre-planning stage. an earlier effort to organize a team of researchers interested in generating data programs proved to be premature. the reason for this may have been simply that scientists invited to participate were too involved in other projects to undertake new, time-consuming ones. it may also be that the most active ipe researchers view their own roles as chiefs rather than braves, as theoreticians willing to recommend and eventually to use improved data files rather than practitioners willing to dig out the data. but, whatever the cause, the result is that any ddir effort to encourage ipe data collections will require renewed vigor. in the meantime, word of mouth and conversations at professional meetings have revealed a number of younger and perhaps less well-known scientists with a keen interest in assembling new data collections so that they can use them for their own research. this suggests a revised ddir strategy. it should doubtless solicit requests for proposals for ipe data programs, fu-st to ascertain the extent to which the community of ipe scientists is interested in undertaking data-gathering activities and, second, if this proves to be the case, to work out joint procedures to coordinate these activities and seek funding. a key element of a projected ddir iii will be the internationalization of any joint data-gathering activities. ddir i and ii have been directly related to datasets generated and carried out predominandy in the united states. it thus made sense to seek initial funding from the u.s. national science foundation. in the future, of course, given the international response to data on national capabilities, interstate conflict, and international events, we may expect more data-gathering activities to emerge in other countries. accordingly, it will make sense to enhance international collaboration and seek international funding. these conditions already exist in the field of ipe, for both data-producers and data-users; and, indeed, the most significant ipe datasets to be created in recent years came from west europe (groenink, 1988; muller, 1988). going it alone, either for individual researchers or those at a single country's academic institutions, may continue to be feasible but is not the best research strategy. 16 lassist quarterly clearly, international collaboration is needed. in april 1989 a study group on qip data was established within the framework of the international political science association (ipsa). ipsa's 15th world congress, to be held in buenos aires, argentina, in july 1991, provided an opportunity for the study group to hold sessions on ipe data development and data uses. letters sent to several dozen u.s. and foreign scientists, however, found virtually no response—and only one expression of interest in participating in such a session (and this a u.s. scientist). establishing the basis for better international cooperation appears to be something yet in the future. the scientific field of international political economy is clearly in an exciting state of fiux. while it is burgeoning in an intellectual sense, its data needs continue to be substantial. governmental and nongovernmental agencies create many datasets, of course, but, for theoretic research carried out at academic institutions, these clearly need assessment to ascertain their value and sometimes much reworking to ensure consistency across time and space. an increasing number of scientists working in the field has recognized the need for ipe data to carry out their research activities. also important is the fact that some of these scientists express interest in improving existing datasets and/or generating new ones. multiinstitutional and multinational organizations can facilitate such research activities. if ddir's current organizational efforts can be carried out—or modified so that they function more effectively—the prospect is for a new era of data-based research on ipe that can significantly address important human issues. references alker, hayward r., jr. (1988) "second (mit-cis/) ddir conference on event data." ddir-update 2:2 (january), 2-5. and frank l. sherman (1982) "collective security-seeking practices since 1945," pp. 1 13-145 in managing international crises, ed. daniel frei. beverly hills, calif., london, and new delhi: sage pubucations, inc. azar, edward e. (1970) "analysis of international events." peace research reviews 4,1 (november), 1-113. (1980a) codebook of the conflict and peace data bank (copdab). ann arbor, mich.: interuniversity consortium for political and social research. (1980b) 'the conflict and peace data bank (copdab) project." the journal of conflict resolution 24:1 (march), 143-152. and thomas n. havener (1976) "discontinuities in the symbolic environment: a problem in scaling." international interactions 2:4,231 -246. and steve lemer (1981) "the use of semantic dimensions in the scaling of international events." international interactions 7:4, 361-378. andt. j. sloan (1975) "dimensions of interaction: a sourcebook for the study of the behavior of 31 nations from 1948 through 1973." international studies association, occasional paper #8. banks, arthurs. (1971) cross-polity time-series data. cambridge, mass., and london: them.i.t. press. and robert b. textor (1963) a cross-polity survey. cambridge, mass.: the m.i.t. press. bloomfield, lincoln p. and amelia c. leiss (1969) controlling small wars: a strategyfor the 1970' s. new york: alfred a. knopf. bomschier, volker and peter heintz (reworked and enlarged by thanh-huyen ballmer-cao and jiirg scheidegger) (1979) compendium ofdatafor world-system analysis: a sourcebook ofdata based on the study ofmncs, economic policy and national development. zurich: university of zurich, sociological institute. brecher, michael and jonathan wilkenfeld (1982) "crises in worid politics." world politics 34:3 (april), 380-417. butterworth, robert lyle (1980) "managing interstate conflict, 1945-79: data with synopses." final report, february. with margaret e. scranton (1976) managing interstate conflict, 1945-74: data with synopses. pittsburgh, pa.: university of pittsburgh, university center for international studies. choucri, nazli and robert c. north (1975) nations in corrflict: national growth and international violence. san francisco, calif.: w. h. freeman. cioffi-revilla, claudio (1990) the scientific measurement ofinternational conflict: handbook of datasets on crises and wars, 1495-1988 a.d. boulder, colo., and london: lynne rienner publishers. coumann, hans-jiirgen (1973) internationale kriminalstatistik: geschichtliche entwicklung und spring 1991 17 gegenwartiger stand. stuttgart: ferdinand enke verlag. deutsch, karl w. (1953) nationalism and social communication: an inquiry into the foundations of nationality. cambridge, mass: the technology press of the massachusetts institute of technology; and new york: john wiley & son, inc. (1960) 'toward an inventory of basic trends and patterns in comparative and international politics." the american political science review 54: 1 (march), 34-57. (1961) "social mobilization and political development" the american political science review 55:3 (september), 493-514. , sidney a. burrell, robert a. kann, maurice lee, jr., martin lichterman, raymond e. lindgren, francis l. lowenheim, and richard w. van wagenen (1957) political community and the north atlantic area: international organization in the light ofhistorical experience. princeton, nj.: princeton university press. , harold d. lasswell, richard l. merritt, and bruce m. russett (1966) "the yale political data program," pp. 81-94 in comparing nations: the use of quantitative data in cross-national research, ed. richard l. merritt and stein rokkan. new haven, conn., and london: yale university press. farris, lee, hayward r. alker, jr., kathleen carley and frank l. sherman (1980) "phase/actor disaggregated butterworth-scranton codebook." cambridge, mass.: the massachusetts institute of technology, center for international studies, working paper. feierabend, ivo k. and rosalind l. fcicrabend (1966a) "aggressive behaviors within polities, 1948-1962." the journal of conflict resolution 10:3 (september), 249-271. and (1966b) "the relationship of systemic frustration, political coercion, international tension and pohtical instability: a cross-national study." paper prepared for delivery at the annual meeting of the american psychological association, new york city, 2-6 september. groenick, ronald j (ed.) (1988) data on europe, 1945-1980. bilthoven: prime press. gurr. ted robert (1970) why men rebel. princeton, nj.: princeton university press. (1974) "persistence and change in political systems, 1800-1971." the american political science review 68:4 (december), 1482-1504. and associates (1978) comparative studies of political conflict and change: cross-national dataset. ann arbor, mich.: inter-university consortium for political and social research. haas, ernst b. (1968) collective security and the future international system. denver, colo.: university of denver, monograph series in world affairs, vol. 5, no. 1. , robert l. butterworth and joseph s. nye (1972) conflict management by international organizations. morristown, n.j.: general learning press. hermann, charles f. (1987) "first ddir conference on event data." ddir-update 1:5 (july), 4-5. , maurice a. east, margaret g. hermann, barbara g. salmore and stephen a. salmore (1973) creon: a foreign events data set. beverly hills, calif., and london: sage publications, inc., sage professional papers in international studies series, no. 02-024. howell, llewellyn d. (1983) "a comparative study of the weis and copdab data sets." international studies quarterly 27:2 (june): 149-159. international labour office (1926) methods of compiling statistics ofindustrial disputes. geneva: international labour office, studies and reports, scries n (statistics), no. 10. kaplan, morton a. (1957) system and process in " international politics. new york: john wiley & sons. usswell, harold d. (1936) politics: who gets what. when, how. new york and london: mcgraw-hill book company, inc., and whittlesey house. and dorothy blumenstock (1939) world revolutionary propaganda: a chicago study. new york and london: alfred a. knopf, inc. leng, russell j. and j. david singer (1988) "militarized interactive crises: the bcow typology and its applications." international studies quarterly 32:2 (june), 155-173. 18 lassist quarterly levy, jack s. (1983) v/ar in the modern great power system, 1945-1975. lexington: university of kentucky press. mcclelland, charles a. (1961) "the acute international crisis." world politics 14:1 (october), 182-2(m. (1983) "let the user beware." international studies quarterly 27:2 (june), 169-177. , rodney g. tomlinson, ronald g. sherwin, gary a. hill, herbert l. calhous, peter h. fenn and j. david martin ( 1 97 1 ) the management and analysis ofinternational event data: a computerized system for monitoring and projecting event flows. los angeles, calif.: university of southern california, school of international relations (september), mimeographed. mcgowan, patrick j. (1988) "second ddir international political economy data conference." ddir-update 2:3 (april), 4-8 , harvey starr, gretchen hower, richard l. merritt and dina a zinnes (1988) "international data as a national resource." international interactions 14:2, 101-113. merriu, richard l. (1990) "data in international research: confluence of interest and possibility." ddir-update 4:3 (april), 1-11. and stein rokkan (1966) comparing nations: the use of quantitative data in cross-national research. new haven, conn., and london: yale university press. moses, lincoln e., richard a. brody, ole r. holsti, joseph b. kadane and jeffrey s. milstein (1%7) "scaung data on inter-nation action: a standard scale is developed for comparing international conflict in a variety of situations." science 156:3778 (26 may), 1054-1059. miiller, georg p., with the collaboration of volker bomschier (1988) comparative world data: a statistical handbookfor social science. baltimore and london: the johns hopkins university press. nye, josephs. (1971) peace in parts: integration and conflict in regional organization. boston, mass.: litde, brown and company. pearson, frederic s. (1974) "geographic proximity and foreign military intervention." the journal of conflict resolution 18:3 (september), 432-460. pollins, brian m. (1988) "ipedata: a preliminary survey." ddir-update 2:3 (april), 8-12. richardson, lewis fry (1960) statistics ofdeadly quarrels, ed. quincy wright and c. c. lineau. pittsburgh, pa.: boxwood press; and chicago, 111.: quadrangle books. rummel, rudolph j. (1964) "dimensions of conflict behavior within and between nations," pp. 1-50 in general systems: yearbook of the societyfor general systems research, vol. 8, ed. ludwig von bertalanffy and analol rapoporl ann arbor, mich.: society for general systems research. (1966) "a foreign conflict behavior code sheet" world politics 18:2 (january), 283-296. (1967) "dimensions of dyadic war, 1820-1952." the journal of conflict resolution 1 1 :2 (june), 1 76183. (1972) the dimensions ofnations. beveriy hills., calif., and london: sage publications, inc. (1976) 'the roots of faith," pp. 10-30 in /n search of global patterns, ed. james n. rosenau. new york: the free press. (1979) national attributes and behavior: data. dimensions. linkages and groups, 1950-1965. beverly hills, calif., and london: sage publications, inc. and hayward r. alker, jr., karl w. deutsch, and harold d. lasswell ef a/. (1964) world handbook of political and social indicators. new haven, conn., and london: yale university press. russett, bruce m. er a/. (1988) "first ddir international political economy data conference." ddir-update 2:3 (april), 1-4. sherman, frank l. (1987a) "four major traditions of historical events research: a brief comparison." paper presented at the second ddir event data conference, boston, mass., the massachusetts institute of technology, 13-15 november. (1987b) partway to peace: the united nations and the road to nowhere. state college, pa.: ph.d. dissertation, the pennsylvania state university. (in progress) recognizing and responding to spring 1991 19 international disputes. singer, j. david and melvin small (1972) the wages of war, 1816-1965: a statistical handbook. new york: john wiley & sons. siverson, randolph m. and michael r. tennefoss (1982) "interstate conflicts, 1815-1%5." international interactions 9:2, 147-178. and (1984) "power, alliance, and the escalation of international conflict, 1815-1965." the american political science review 78:4 (december), 1057-1069. small, melvin and j. david singer (1982) resort to arms: international and civil war, 1816-1980. beverly hills, calif., london, and new delhi: sage publications, inc. sorokin, pitirim a. (1937) social and cultural dynamics, vol. 3: fluctuation of social relationships, war, and revolution. new york: american book company. starr, harvey (1987) "what is a national resource in international data?" ddir-update 1:3, app. (february), 5-6. taylor, charles lewis and michael c. hudson (1972) world handbook of political and social indicators (2d ed. ) new haven, conn., and london: yale university press. taylor, charles lewis and david a. jodice (1983) world handbook of political and social indicators (3d ed.). new haven, conn., and london: yale university press. united states, department of labor (1976) bls handbook of methods. washington, d.c.: u.s. department of labor, bureau of labor statistics, bulletin no. 1910. vincent, jack e. (1983) "weis vs. copdab: cotrespondence problems." international studies quarterly 27:2 (june). 160-168. walker, nigel (1971) crimes, courts and figures: an introduction to criminal statistics. harmondsworth, middlesex: penguin books ltd. woods, frederick adams and alexander baltzly (1915) is war diminishing? study of the prevalence of war in europefrom 1450 to the present day. boston, mass., and new york: houghton mifflin company. wright, quincy (1942) a study of war. university of chicago press. chicago, 111.: ' richard l. merritt is professor of political science and research professor in communications, and dina a. zinnes is the charles and ethel merriam professor of political science, both at the university of illinois at urbana-champaign. ddir is housed at: merriam laboratory for analytic political research, university of illinois at urbana-champaign, 512 east chalmers street, champaign, illinois 61820-3696, u.s.a. (tel: 217-2440739; fax: 217-333^369). its newsletter, ddirupdate, is currently distributed without cost; as its subscribers expand in number, however, we can expect a nominal fee to be charged. subscriptions are available from the merriam lab. lassist quarterly iassvol201 4 iassist quarterly bringing census data into the classroom: world wide web access and teacher networking by william h. frey and cheryl l. first1, population studies center the university of michigan once upon a time as jackie, a college sophomore, puts together her fall schedule, her interest is peaked by courses which address current social issues, such as sociology 202, which focuses on marriage and childbearing trends. her enthusiasm for soc 202 wanes slightly when she sees that it meets at 8:30 am. karen, jackie’s roommate, reminds her that she needs to fulfill the statistics requirement and that stat 402 meets at 11am, a definite plus. jackie groans, “that class will put me back to sleep anyway.” like many students, jackie dreads the statistics course because her math skills are not strong. furthermore, word on campus is that it is a dry course that has nothing to do with real life. jackie opts for soc 202; she will put off stat 402 for as long as possible. from the same old story to a new perspective any resemblance to the living or dead in the above story is not accidental, it is inevitable. this dilemma is played out on most campuses every semester. to many, the dilemma may not seem especially problematic because “jackie” must eventually complete the statistics requirement. however, two major problems are created by a curriculum which presents theory and data analysis as separate entities. first, this curriculum structure leads the student to believe that the compelling questions and possible solutions to today’s pressing social issues are somehow distinct from quantitative reasoning. in other words, why and how are presented, and consequently understood, as two completely different questions. the second problem is of a more practical nature, but nonetheless troublesome. when students put off the quantitative element of their degrees, they are postponing their opportunity to take the more substantive upper level courses which require data analysis skills. consequently, the quality of their undergraduate training, and its relevance to their eventual careers, decreases. in light of these problems, the social science data analysis network (ssdan) project seeks to make empirical data analysis explorations an accessible, available, and desirable component in introductory social science courses. it combines engaging course material on american society with basic data analysis exercises which utilize data from the u.s. census and other sources. issues that can be addressed with u.s. census data include: immigration and the increasing diversity of the american population changes in the roles of women and the structure of the family industrial restructuring and the shrinking middle class the civil rights movement’s impact of black-white inequality by “marrying” theory and data analysis in an active learning setting, these courses illustrate that quantitative reasoning skills are relevant to social issues. furthermore, the courses prepare students for applied upper level courses and careers that utilize data analysis skills. in addition to designing course material and preparing datasets, the ssdan project provides support for instructors who are committed to, but not necessarily experienced in, incorporating “hands on” data analysis in their courses. this support includes in-person and “virtual” internet accessible workshops, a world wide web homepage (see exhibit a for ssdan homepage) and electronic e-mail groups. through these mechanisms, instructors not only receive support from ssdan staff, but also have the opportunity to network with each other. a more comprehensive explanation of these mechanisms, and other ssdan materials, will be provided later in this paper. the most innovative objective of this project is to introduce networking capabilities that link data and research expertise at the university of michigan population studies center with faculty at two and four year colleges. this link will enable interactive feedback on development of curricular materials over the internet. the computer network will be used to: (1) aid in the creation of datasets and curricular (i.e. faculty will suggest exercises and relevant datasets to be produced at michigan); and (2) to provide continuous sharing of these materials and feedback among faculty via conferencing. the nuts and bolts of building a new approach in order to introduce faculty to this new approach and help them create data exercises appropriate for their courses, the ssdan project has implemented both traditional in-person workshops and a “virtual workshop” — via the internet — that enables social science faculty at two and four year colleges to exchange data and ideas through the project’s world wide homepage and electronic e-mail groups. through these channels, we have already developed an extensive network of over four hundred interested faculty 5spring 1996 social science data analysis network census data and exercises for college classes located at the population studies center, university of michigan in collaboration with the great lakes colleges association. contact:william frey, project director population studies center, university of michigan william.frey@umich.edu what is ssdan? ¥answers to questions about ssdan ¥ send me a "start-up package"/put me on mailing list summer workshop in ann arbor ¥ workshop for faculty of glca colleges ¥ workshop for all others ssdan census data sets ¥about ssdan data sets ¥ data sets for downloading classroom resources ¥course exercises ¥ other course aids email ssdan-staff@umich.edu for technical assistance. check back soon for new features on our web site for the 1996 fall semester! this project is supported by the national science foundation and the department of education fipse with additional funding provided by the alfred p. sloan foundation. ssdan world wide web homepage address: http://www.psc.lsa.umich.edu/ssdan/ exhibit a 6 iassist quarterly around the country. we have also published a workbook which includes over two hundred student exercises and covers ten american society topics that can be incorporated into many social science courses. the workbook is bundled with a diskette which contains specially tailored u.s. census data for 1950-1990 and the chipendale program, an extremely user-friendly data analysis software for novices. in-person workshops the primary goals of the in-person workshops are: (1) to expose the participants to curricular materials we have developed through lectures, discussions, and extensive “hands-on” use; (2) to make them familiar with computer conferencing and data access features of our computer network; (3) to work with them, individually or in small teams, toward creating data analysis exercises that they will use during the next academic year; and (4) to expose participants to other resources that complement their use of the data analysis exercises they will produce in our workshop. our annual in-person summer workshops in ann arbor are open to a national audience of instructors. the two workshops, each a week long, that were held during the summer of 1996 introduced 28 college teachers to the ssdan materials. the participants were selected from some 80 applicants, and represented a variety of colleges and social science disciplines. primary consideration was given to highly motivated faculty interested in adding a data analysis component into a lower level substantive course they already teach. they can adapt any of the project’s current census data analysis exercises to their classes or, with our assistance, develop new exercises. participants were introduced to the resources of ssdan in “hands on” training sessions, had seminar discussions, worked with ssdan staff to begin developing classroom exercises specific to their own courses, and practiced exploring the ssdan materials. participants were also exposed to pdq-explore, a “cutting edge” instructional computer tool which allows users, to conduct u.s. census data analysis directly over the internet. the pdq-explore program is currently being developed in cooperation with the university of michigan and will be discussed later in this paper. overall, the workshop got high marks from the 1996 participants due to the information presented and the human networking possibilities that were set in place. we are now working with these faculty to design classroom exercise/ dataset modules that they will use. these modules and datasets will be posted on our homepage in order to facilitate sharing among any interested instructors. ssdan world wide web homepage : http://www.psc.lsa.umich.edu/ssdan/ the project homepage describes the project (see exhibit a), makes exercises available, and facilitates downloading of census datasets that can be accessed with chipendale software in both ibm and mac format. the datasets are indexed according to the variables and the kinds of courses for which the datasets are most suited. the web page also serves to update instructors on developments related to the project, and provides links to other teaching resources. the homepage has also been a useful way to advertise the project, and as a result, we have received requests for additional information and datasets among social science faculty in a wide variety of institutional contexts around the world. although the homepage is free to browsers, we do ask that they “register” with us first. the homepage enables project participants and all others who wish to access it, with the ability to send requests and communicate with the project staff. the project staff is currently developing more features for the web page that will assist both professors and students. our “new and improved” homepage will include: new datasets and exercises for downloading, references to current popular and scholarly articles relevant to ssdan; a form for faculty to submit classroom exercises and corresponding dataset ideas to ssdan staff; and a quicksurvey for feedback on ssdan. in future months, we will create a section which targets students. this section, in addition to being a bit more “hip”, will include links to articles which can help them with their coursework, data they may find interesting, and surveys. the “virtual” internet accessible workshop the most innovative aspect of this project is the implementation of a “virtual” workshop. this means that faculty from any undergraduate program can participate in the conferencing, data access, and exercise creation activities of this project by communicating with our core faculty and staff over the internet, or using e-mail. our “virtual” workshop revolves around two components: (1) an e-mail discussion group for active participants; and (2) a data analysis exercise “bank” that contains text and datasets for individual class exercises which can be retrieved by the instructor over the internet or via e-mail. to use our materials, it is not necessary that entire universities, classrooms, or student audiences have internet access, but rather that the instructor has access, somewhere on campus, to an individual internet connection, or e-mail account. hence, a large body of faculty participants in this network can use this system to retrieve data exercises for use in their classes (from the “bank”). all instructors who wish the michigan staff to work with them in producing exercises must agree to “try them out” in their own classes and provide feedback to us regarding their effectiveness. this feedback will be recorded into the notes that are attached to the data exercises, deposited in the general bank. these exercises, in addition to those created by 7spring 1996 example 1 look at the percentage of blacks and nonblacks who were never married, from 1950 to 1990. how has the percentage of each group who have never been married changed over time? what might account for these changes? (marr5090) n create a line graph with separate lines for blacks and nonblacks indicating the percentage of each group who have never been married for each year. cross tab= race / marital (control for year); percent across year = 1950 curmr widow divor sepra nevmr total black 57.6 10.5 2.3 7.8 21.7 100% nonbl 67.5 8.0 2.2 1.2 21.0 100% all 66.6 8.2 2.2 1.9 21.1 100% year = 1960 curmr widow divor sepra nevmr total black 56.0 10.0 3.2 7.7 23.2 100% nonbl 68.7 7.8 2.5 1.2 19.8 100% all 67.5 8.0 2.6 1.8 20.1 100% year = 1970 curmr widow divor sepra nevmr total black 49.1 9.8 4.3 7.6 29.1 100% nonbl 64.8 8.0 3.3 1.3 22.6 100% all 63.2 8.2 3.4 1.9 23.3 100% year = 1980 curmr widow divor sepra nevmr total black 39.4 8.6 7.7 7.3 37.0 100% nonbl 60.2 7.5 6.0 1.6 24.6 100% all 58.0 7.6 6.2 2.1 26.0 100% year = 1990 curmr widow divor sepra nevmr total black 35.2 8.0 10.1 6.6 40.1 100% nonbl 58.1 7.3 8.1 1.7 24.7 100% all 55.6 7.4 8.3 2.3 26.4 100% the answer can be plotted as follows: 50 40 30 20 10 0 key black nonblack 1950 1960 1970 1980 1990 ▲ ❇ ▲ ▲ ▲ ▲ ❇ ❇ ❇ ❇ ▲ ❇ blacks and nonblacks never married 1950 to 1990 exhibit b 8 iassist quarterly our in-person workshop members, should result in over 300 data exercises at the conclusion of the project. “investigating change in american society” workbook our workbook, investigating change in american society: exploring social trends with us census data and studentchip has several topics (chapters) that are appropriate for almost any social science course. the workbook is flexible enough to be used with a variety of texts, or additional readings, and can easily be integrated into existing courses. the single most important feature of our workbook is its adaptability to a wide range of social science courses in which the instructor wants to introduce one or more “hands on” data analysis modules. see exhibit b for a sample exercise. in order to make data analysis interesting and engaging to students, the workbook includes a wide variety of interesting issue-oriented topics. see exhibit c below for a listing of all ten investigation topics. these investigation topics (chapters) are self-contained. in other words, they can be “mixed and matched” according to the instructor’s course sequence. within each topic (chapter), the difficulty level of the exercises increases as the students read through the topic. at the end of each topic there are “think tanks”, broader questions which aim to generate class discussions. these questions are ideal for team exercises. the students complete the data analysis modules with userfriendly, engaging chipendale software which comes with the workbook. a full tutorial at the beginning of the book can get teachers and students “up to speed” in one classroom period. the “look and feel” of our book is not of a statistics or methods book and our exercises are centered on having the students explore data to examine interesting and engaging social issues — rather than to learn more advanced statistical methods or jargon. this is reflected in the heavy use of graphics, and attention to understanding basic sociological concepts and measurements rather than focusing on statistical or specialized methodological concepts. the world on a diskette the workbook includes a diskette (students can choose between an ibm or mac compatible edition) containing chipendale contingency table software. this software was selected because of the ease of use, low expense for students, and appropriateness for straightforward contingency table analyses of census data. the program is menu driven and extremely user friendly, but also sufficiently unfriendly to allow students to recognize when they have made poor decisions. the program makes it easy for students to recode variables, select controls, and make graphs that highlight their comparisons. any standalone pc can use the software in virtually any configuration. the chipendale program is efficient to use on any pc because it takes “input” data as tables or matrices, rather than individual cases. this allows students to manipulate large aggregate-type datasets, such as those from the u.s. census, in small compact files. a whole range of topics related to american social, economic and geographic issues can be explored with investigation topics •population structure: cohorts, ages, and change •gender inequality •race and ethnic inequality •households and families •immigrant assimilation •poverty •labor force •children •marriage, divorce, cohabitation, and childbearing •the older population exhibit b 9spring 1996 ssdan datasets. the datasets are drawn from the u.s. census and include variables such as: race-ethnicity, gender, immigration status, earnings, education, occupation, cohabitation, and workhours. some allow analyses of trends over the census years 1950, 1960, 1970, 1980, and 1990. others permit more in-depth comparisons across social and demographic groups and geographic areas from the 1990 u.s. census. u.s. census data the most comprehensive dataset needed to assess the kinds of over-time changes that we have been discussing are available from the u.s. census. the wide range of statistics collected by the decennial census is especially useful in social science research. this is because this information is collected for a large number of people and detailed social and economic information can be gathered for tiny population subgroups and small geographic areas. unlike many small surveys, the census information is rarely limited by having “too few observations” to be statistically representative. ssdan participants this project especially targets groups and institutions that have often been neglected in advances of quantitative work in the social sciences. specifically, this means concerted efforts to include women, minority and disabled faculty, as well as two-year undergraduate institutions and historically black colleges and universities. this special targeting is made explicit in the recruiting for our in-person workshops in ann arbor, and in our efforts to engage off-site instructors, interactively, in creating exercises for their classes. funding ssdan is currently funded by the national science foundation undergraduate faculty enhancement grant and a us department of education fipse project, building upon earlier funding from the alfred p. sloan foundation and an undergraduate initiatives award granted to dr. frey who first developed this approach in his university of michigan course. the fipse project demonstrated the feasibility of incorporating interactive u.s. census data analysis via the internet into existing undergraduate curricula in colleges of the great lakes colleges association. the current project extends this approach to a national community of social science instructors, that has come to be called ssdan, the social science data analysis network. the future in the next year, the teaching approach of this project will be enhanced by utilizing a student version of the pdq-explore program and class-room exercises will be made available. as mentioned previously, the pdq-explore, a “cutting edge” instructional computer tool where students can conduct direct u.s. census data analysis interactively over the internet, is an avenue which the ssdan will be taking. pdq-explore allows users to request us census data tables that are immediately delivered to the users computer screen over the internet. eventually, this program, along with the chipendale software, will be modified so that our approach can be available to high school students, as well as undergraduate college students. the social science data analysis network looks forward to continuing to grow, as our network of interested faculty and our “bank” of class-room exercise modules and datasets expands. please contact us via e-mail or through our world wide homepage at: http://www.psc.lsa.umich.edu/ssdan/ if you or one of your colleagues would like more information about the project, or would like to participate in developing exercise modules and us census datasets for your own classroom. 1. paper presented at the iassist/computing in the social sciences conference, hotel radisson, minneapolis, mn, may 1995. dr. william h. frey, director of ssdan (social science data analysis network) is on the faculty of the population studies center, and the department of sociology at the university of michigan in ann arbor. ssdan draws from an approach he developed with his michigan course, and later disseminated to faculty of the great lakes colleges association. cheryl l. first msw is the project manager of the social science data analysis network. she coordinates the inperson and virtual workshops for the project, and also maintains the ssdan world wide web homepage. acknowledgments: the authors are grateful to bridget fahrland for her editorial contributions. oecd iassist quarterly summer 2010 9 iassist quarterly abstract in recent years there has been an explosion of web 2.0 technologies such as facebook, twitter, blogs and wikis. the world is becoming ever more connected with great technological strides which are increasingly allowing for more access to those who previously were disconnected. web technologies are fostering mass collaboration on a scale that has never been seen before, allowing for a collective intelligence to develop on specific topics and themes. the crucial element to this participation is the openness of new web platforms, engaging everyone from policy makers to local ngos to interested individuals. recent technological advances are very timely as the world is also becoming more complex, and society has to face a number of global challenges that it can no longer address without connecting to others: from climate change to the financial crisis to sustainable livelihoods, these problems are global and therefore require global thinking. an important part of the problem-solving process involves the collection of information and data from diverse sources. this paper is concerned with the idea that collaborative platforms such as wikis along with advances in data visualisation are a way forward for the collection, analysis and dissemination of data across countries and societies, and presents two wiki platforms that can support such a vision: wikiprogress and wikigender . the paper first explains the wider movement behind the two wikis, namely the oecd-hosted global project on measuring the progress of societies, then explores the importance of fostering collaboration and data dissemination through wikiprogress and wikigender. finally, this paper concludes that collaboration is required for the shift from measuring economic production to measuring well-being to occur. such collaboration is facilitated by web 2.0 technology. introduction societies are becoming increasingly connected thanks to advances in new technologies. global problem-solving has been made easier through the new forms of collaboration this connectivity allows. a plethora of web 2.0 initiatives have flourished around the world, including issues like tackling climate change to solutions for sustainable development, the reduction of poverty and hunger, and decreasing the wage gap between women and men. the range of topics tackled on the internet reflects societies’ increasing concern with their quality of life along with their material well-being. leading economists studying the measurement of those factors which contribute to a person’s welfare conclude that “gdp is an inadequate metric to gauge well-being over time particularly in its economic, environmental, and social dimensions, some aspects of which are often referred to as sustainability.” numerous initiatives worldwide, such as the oecd-hosted global project on measuring the progress of societies, amongst others, have shown an interest in going beyond the mere economic prosperity of a nation – claiming that we should look beyond traditional measures such as gross domestic product (gdp) in determining the well-being of societies. wikiprogress, the official platform of the global project on measuring the progress of societies is a collaborative tool for measuring progress beyond gdp that exemplifies this new form of collaboration and the potential web 2.0 technologies offer in solving global problems. it is important initially to understand the underlying movement and what “measuring progress” actually means. examining the background of ‘measuring progress’ and identifying the key players will clarify the oecd’s choice of a ‘wiki’ platform and what impact this has on the measuring progress movement. wikiprogress and wikigender a way forward for online collaboration by angela hariche, estelle loiseau and philippa lysaght box 1: wikiprogress: goal and purpose goal: to create a web community around the vision of measuring the progress of societies by creating a place where progress data and research articles can be loaded, visualised and analysed so good decisions about societies can be made at the local, national and international levels. purpose: to invite and inform all parts of the progress community, citizens and policy makers alike to the debate on progress-related initiatives by creating a robust wiki of related research and statistics. 10 iassist quarterly summer 2010 iassist quarterly 1. “measuring progress”: an overview of the movement for the last century the most commonly used indicator of economic performance has been gdp, also used as an indicator of national performance. it is commonly assumed that gdp growth is synonymous with progress, that a rise in gdp mirrors an increase in the quality of life and well-being of a particular society. this, however, is not the case: gdp only measures what is ‘produced’ – it is a very narrow economic measure that fails to take into account other elements which are important in the daily lives of citizens such as security, pollution and health care. gdp calculations include economic activities that reduce well-being such as traffic congestion, which increases fuel consumption, or those which remedy the costs of economic growth like pollution abatement. currently there are many organisations who are looking at new measure of progress and calling for indicators going beyond gdp to measure the well-being of individuals, of our societies and our environment. a key message from the stiglitz-sen-fitoussi commission on economic performance and social progress in 2009 shows an urgency for developing such measures, declaring that now is the time for our measurement systems to shift the focus from measuring economic production to measuring people’s well-being. furthermore, such measures of well-being should be considered and analysed within the context of sustainability. the movement is not concerned solely with developing progress indicators but also the development of a collaborative community, working to determine how we measure progress. this kind of knowledge sharing and creation can aid statistical organisations or government by a nurturing community that encompasses diverse actors, including individuals worldwide. for this reason, the global project on measuring the progress of societies chose to use a wiki platform, called wikiprogress. the box below outlines the goal and figure 1. visualisation of the www.my.genderindex.org tool iassist quarterly summer 2010 11 iassist quarterly purpose of wikiprogress, reflecting the importance of fostering a collaborative online community. 2. why a wiki? according to the world’s largest online encyclopaedia, wikipedia, a wiki is “a website that allows the easy creation and editing of any number of interlinked web pages via a web browser using a simplified mark-up language or a text editor.” in other words, a wiki is a type of web platform that allows user-generated content. the content in a wiki is created by a community of users working together to develop information on a particular subject or within a particular field. it is essentially a database of information that can be browsed, searched, created and edited. figure 2. visualisation of the gross enrolment ratio (ger) and gender parity index (gpi) datasets, global education digest 2010, unesco institute for statistics. global education digest 2010 12 iassist quarterly summer 2010 iassist quarterly figure 3 preview of a country note on benin. but wikis, blogs and other web 2.0 technologies are much more than simply tools to facilitate collective contribution. they are the force behind a new era of mass collaboration and participation that is changing the way we deal with global issues. information is now being created, updated, discussed and made public by an engaged community from all over the world. rather than an expert or team of experts assembling information on a particular topic, access is now open to worldwide resources in developing knowledge around a subject. a recently coined term, “wikinomics”, defines this movement as “the new force that is bringing people together on the net to create a giant brain” . openness and collaboration using web 2.0 functionality provides for a democratisation of innovation. the speed, scope and manner in which innovation occurs has radically and rapidly changed since web 2.0 tools, including wikis, can provide “low-cost collaborative infrastructures… in ways only large corporations could in the past” . hence, the collective knowledge and ingenuity of individuals and businesses is fully exploitable. 3. what are wikiprogress and wikigender? wikiprogress (www.wikiprogress. org) and wikigender (www. wikigender.org) are interactive communication tools that centralise information, data, initiatives, publications, events, media coverage and research networks that are all part of the international movement to look beyond gdp in measuring the progress of societies. wikiprogress brings together many dimensions of progress (ecosystems condition, human well-being, economy, governance and culture). wikigender specifically looks at gender equality and women’s empowerment, with a particular focus on developing countries. 3.1 wikiprogress wikiprogress is the official online platform for the oecd-hosted global project on measuring the progress of societies. it was announced and presented in beta version at the 2009 oecd world forum in busan, korea. since then, wikiprogress has grown significantly: it now has over 450 active members, over 840 articles and over 8,400 unique visitors a month. the global project is a partnership project that aims to develop a large “network of networks” of experts in the field of research into progress, including academics, civil society and ngos, to work together on gathering information and data related to measuring progress. the wiki platform facilitates this collaboration and participation, ensuring the movement is inclusive and ever-growing. it is not only the community that makes wikiprogress effective, but also the content diversity. so far, the term “information” has been used to cover all of the content on wikiprogress, but to understand the extent of the information available, it can be divided into key sections. these divisions will change and develop over time as the wikiprogress community expands and enriches the content. • information by topic – information on a particular dimension of progress (e.g. peace, trust, pollution) is gathered in one article with internal links to related pages, external links to relevant organisations and a list of reports and publications for further reading. • information by country – each country has its own article listing the main national and regional initiatives. the country articles also link to relevant data and list country-specific publications and reports. • progress initiatives – initiatives from around the world that gather information and/or data on measuring progress have their own article including the aim, function and background of the iassist quarterly summer 2010 13 iassist quarterly particular group. these articles also contain links to relevant data, videos of presentations and speeches, publications and reports. there are many different initiatives, some working within specific dimensions of progress and others working across all dimensions. • progress publications – wikiprogress highlights reports and papers as they are published, on the condition that they are not for sale. the reports can be searched by topic or by country. • progress-related events – events are held worldwide focusing on different dimensions of progress, the measurement of progress and the development of better indicators. • media coverage – a good measure of a movement’s success, and certainly a good way to monitor the development of the movement, is by tracking what the media is covering and examining media treatment of key players and events. the wikiprogress community portal gathers news items and blog posts on the movement as they are released. • finally, one of the most significant aspects of wikiprogress: progress-related data. this component of wikiprogress is detailed in the last part of this paper which focuses on data. 3.2 wikigender wikigender is an online platform that was developed by the oecd development centre and launched on 8 march 2008, on the occasion of international women’s day. it is the first wiki platform hosted by the oecd, and acted as a pilot project before the launch of wikiprogress. wikigender is therefore naturally linked to wikiprogress. furthermore, measuring the progress of societies without paying particular attention to women, who represent more than half of the population, would not be offering a comprehensive picture. since its launch, wikigender has seen its community grow exponentially: it now consists of over 960 registered users coming from around 170 countries and comprises more than 1,120 articles. added to these statistics, the site recently attracted over 15,000 visits during the course of a single month, a number that increases monthly. wikigender, is an interactive online tool that focuses on gender equality issues, as well as on data and measurements in the area of gender equality. like wikiprogress, the site allows its registered users to add, edit or discuss information and/or data provided on the site. the platform therefore easily allows for up-to-date and accurate information. the information available on the site is also organised by topic/ category, country, organisations/initiatives, publications, events and statistics. it also has a community portal that lists media coverage on timely gender equality issues. in terms of specific projects, wikigender has started to work with paris-based universities, engaging students to add content to the site and benefitting from the networking opportunities of wikigender. this programme, called “wikigender university”, will be extended to selected universities worldwide, further enriching wikigender with content related to the measurement of progress made in gender equality, and at the same time empowering students through information-sharing and capacity-building in the area of information technology. similarly, “wikigender impact” is project under development examining the establishment of a bridge between the community of donors and policymakers and the community of practitioners on the ground, by developing a database of case studies and becoming a clearinghouse for funders and those seeking financial support for projects. 3.3 the importance of understanding data developing better indicators and awareness is not enough. we also need to ensure that the measures become widely used, and widely understood – not only by statisticians, but by all those interested in societal progress. visualising data in a time series allows us to see the impact of the progress – or regress – according to the particular data set. it also makes the data easier to understand, so that a wider audience has access to understanding its significance. there follows an introduction to the data presentation and visualisation tools available on the wikis. a) wikiprogress.stat it is one thing to have good statistics. it is another to make sure these statistics are used. new information communication technology (ict) tools have the ability to not only make data easily available for anyone with an internet connection, but also allow for the data to be presented in a way that can be understood. both wikiprogress and wikigender have developed ways in which data could be better understood. wikiprogress has developed a statistical tool that facilitates the task of data collection, analysis and dissemination: wikiprogress.stat. data from a variety of sources is progressively populating this database of progress indicators and is being used to support articles in both wikiprogress and wikigender. the database also complies with the wiki model; users can upload their own data and metadata, therefore democratising the process of data collection. through wikiprogress. stat, anyone can upload data and access the information. after a quality control procedure, the data will be uploaded into wikiprogress. the overall aim of wikiprogress.stat is to create a robust database of progress indicators for a wide range of regions, over the longest possible time-scale. there are currently over 100 data sets in the wikiprogress.stat database, including the global peace index, world bank economic indicators, and the world database of happiness, amongst others. tables, graphs and charts created on wikiprogress.stat will then be used in articles on wikiprogress to illustrate how different progress dimensions have impacted the well-being of particular societies. additionally, the oecd is working with norrköping communicative visual analytics (ncomva) on the implementation of a new data visualisation tool, explorer, to allow users to create stories using data sets in wikiprogress.stat and visualise the material in an animated time series. each visualised story, a vislet, is embeddable in wikiprogress articles and on external websites, blogs, etc. furthermore, the explorer will be an embeddable tool not only in wiki articles, but also in external blogs and online journals. this will broaden the scope of discussions and reports that can use the visualisation to assist in explaining and/or supporting ideas and arguments on social progress. b) wikigender gender equality datasets are also uploaded through wikiprogress.stat. apart from presenting oecd data such as the gender, institutions and development database (gid-db) or the social institutions and gender index (sigi), wikigender links to www.my.genderindex.org, a tool allowing users to manipulate sigi data for their own analyses. furthermore, any new gender-related data included in wikiprogress. stat automatically appears on wikigender. this is the case, for example, of unesco data accompanying the 2010 global education digest. finally, wikigender also presents information on the gender equality situation by country. each country page includes quantitative and qualitative data from the sigi as well as data from additional sources to offer a more comprehensive view of gender equality in a given country. for example, sigi data examines discrimination against women from the point of view of social norms, cultural values, and traditions. but some country notes also cover data from wikigender partners the world bank and the international finance corporation (ifc), looking 14 iassist quarterly summer 2010 iassist quarterly at laws and regulations affecting women’s prospects as entrepreneurs and employees. more data will also soon be added, for example, examining women’s and men’s access to land. wikigender will also continue to collaborate with other organisations in order to give the most accurate and comprehensive picture of gender equality across countries. such an approach gives the advantage to the user to find information on one topic, in one country, looked at from various angles and perspectives. . the above tools offer various entry points to the data, allowing any site users to access, upload, manipulate, or analyse data and draw analytical conclusions. both wikiprogress and wikigender have large and diverse audiences which include policymakers, non-governmental organisations, think tanks, international organisations, statistical offices, students, academics, practitioners, and donors. conclusion the movement to look beyond economic indicators in measuring well-being can be successful by working society-wide. in order to measure what is important to societies, a collaborative effort is required. currently, wikis are the correct web platform for collective brainstorming, collaborative problem-solving and for centralised global efforts and initiatives working towards a common goal. in this regard, wikiprogress is the appropriate tool for the successful collection, analysis and dissemination of information across countries and societies, as wikigender is similarly for collaboration and knowledge sharing in the area of gender equality and gender statistics. we invite members of the iassist community to join the community and participate in the ongoing debates on www.wikiprogress.org and www.wikigender.org. we are always looking for new ways to collaborate with different organisations. if you are interested in either of these initiatives, we invite you to contact us at info@wikiprogress.org. notes 1.hariche, angela, estelle loiseau, et philippa lysaght. «wikiprogress and wikigender: a way forward for online collaboration.» 2010 iassist conference. ithaca, 2010. 2. organisation for economic co-operation and development (oecd). wikiprogress. 2010. http://wikiprogress.org/. 3.organisation for economic co-operation and development (oecd). wikigender. 2010. http://wikigender.org/. 4. stiglitz, joseph e., amartya sen, et jean-paul fitoussi. report by the commission on the measurement of economic performance and social progress. pdf, paris: commission on the measurement of economic performance and social progress, 2009. pg. 8. 5. ibid. .wikipedia contributors, “wiki,” wikipedia, the free encyclopedia, http://en.wikipedia.org/w/index.php?title=wiki&oldid=411214821 (accessed february 10, 2011). 7.tapscott, don, et anthony d. williams. wikinomics: how mass collaboration changes everything. new york: portfolio, 2006. 7. williams, anthony d. «wikinomics and the era of openness: european innovation at a crossroads.» the lisbon council e-brief issue 05/2010. 2010. http://www.lisboncouncil.net//index. php?option=com_downloads&id=316. 8. organisation for economic co-operation and development (oecd). wikiprogress.stat. 2010. http://wikiprogress.stat/. 9. wikigender.org contributors, “the 2010 global education digest: datasets,” wikigender.org, http://www.wikigender. org/index.php?title=the_2010_global_education_digest:_ datasets&oldid=15654 (accessed february 10, 2011). 10. other wikigender partners; américa latina genera, destatis, fao, fidh and africa for women’s rights campaign, ips genderwire, la halde, mujeres del muldo + 2nd international documentary festival on gender, united nations womenwatch, un-instraw vol25s.1 iassist features quarterly volume 25 number 1 spring 2001 the iassist quarterly represents an international cooperative effort on the part of individuals managing, operating, or using machine-readable data archives, data libraries, and data services. the quarterly reports on activities related to the production, acquisition, preservation, processing, distribution, and use of machine-readable data carried out by its members and others in the international social science community. your contributions and suggestions for topics of interest are welcomed. the views set forth by authors of articles contained in this publication are not necessarily those of iassist. information for authors: the quarterly is published four times per year. authors are encouraged to submit papers as word processing files. hard copy submissions may be required in some instances. word processing files may be sent via email to: kbr@sam.sdu.dk. manuscripts should be sent to editor: karsten boye rasmussen, department of organization and management, university of southern denmark, sdu-ou, campusvej 55, dk-5230odense m, denmark the first page should contain the article title, author's name, affiliation, address to which correspondence may be sent, and telephone number. footnotes and bibliographic citations should be consistent in style, preferably following a standard authority such as the university of chicago press manual of style or kate l. turabian's manual for writers. where appropriate, machinereadable data files should be cited with bibliographic citations consistent in style with dodd, sue a. "bibliographic references for numeric social science data files: suggested guidelines". journal of the american society for information science 30(2):77-82, march 1979. announcements of conferences, training sessions, or the like, are welcomed and should include a mailing address and a telephone number for the director of the event or for the organization sponsoring the event. editor karsten boye rasmussen, department of organization and management, university of southern denmark, sdu-ou, campusvej 55, dk-5230 odense m, denmark phone: +45 6550 2115 email:kbr@sam.sdu.dk production william block, walter piovesan minnesota population center, maps/data/gis library, university of minnesota, simon fraser university, 537 heller hall burnaby, b.c 271 19th avenue south. canada v5a 1s6. minneapolis, mn 55455. phone: (604) 291-5869. phone: 612-624-7091 email:walter@sfu.ca email:block@socsci.umn.edu title: newsletter international association for social science information service and technology issn united states: 0739-1137 © 2002 by iassist. all rights reserved. c o n t e n t s 5 understanding barriers to the use of numeric data in learning and teaching robin rice 10 the virtual training suite: internet skills for teaching and learning heather dawson 15 mission (multi-agent integration of shared statistical information over the [inter]net) the data archive perspective joanne lamb 21 research for building a better data community charles k. humphrey the articles in this issue are based on presentations from the iassist conference in amsterdam in 2001. the articles show great involvement and enthusiasm in data archiving and dissemination. we begin with a paper from the session on "learning and teaching" by robin rice from the data library at edinburgh. the title is "understanding barriers to the use of numeric data in learning and teaching" and starts by stating as a fact that "data resources are under-used in the learning and teaching environment". robin rice quotes that there is a lack of "statistical literacy". the joint information systems committee (jisc) funded this project on – among other issues – "the extent of use and the practicalities of using editor's notes continued on page 4 understanding barriers to the use of numeric data in learning and teaching the virtual training suite: internet skills for teaching and learning mission (multi-agent integration of shared statistical information over the [inter]net) -the data archive perspective research for building a better data community mailto:kbr@sam.sdu.dk mailto:walter@sfu.ca mailto:block@socsci.umn.edu 4 iassist quarterly spring 2001 data in teaching". a survey was carried out and i have observed that one of the results was that only "one-quarter of those who teach with data" were familiar with the national data services. the late per nielsen from denmark often mentioned that data archivists also should be "data pushers" – a marketing effort is needed! the article concludes by wondering how information technology will change into learning beyond the traditional classroom, the "virtual classroom" concept comes to mind. this leads to the next article – from the same session – where heather dawson with assistance explains about "the virtual training suite: internet skills for teaching and learning" aiming to "support lecturers, students, and researchers in finding and using resources on the internet". again the jisc has funded the project, which has developed 40 free web based internet tutorials that cover a wide range of the academic subjects taught in uk universities and colleges aiming to provide a structured learning environment. in the session "tools for data services" joanne lamb from edinburgh talked about mission (multi-agent integration of shared statistical information over the [inter]net", in this article " – the data archive perspective" has been added. this time the project is funded by the european commission to develop software "to allow consumers of statistics to access these data in an informed manner with minimum effort". the article also includes some technicalities about the architecture of the system as well as showing the use of metadata and connects to iassist well-known acronyms as xml, ddi and faster. you could call it "navel-gazing" or "praxis related research" when we had a session on "iassist" at the iassist conference. from canada charles k. humphrey writes on "research for building a better data community". and i simply have to quote him on: "i was exited by the research carried out by karsten boye rasmussen and repke de vries about iassist as a virtual community". in this community we are (still) facing the problem of "how to get researchers to conduct their projects so that their data products meet archival standards". in a study on practitioners use of medical research findings one objective was to identify data products from the researchers. a survey with attitudinal items shows that (only) around 80 percent regard "data a valued by-product" and "secondary analysis as a valid research method". do we have a 20-80 problem here? furthermore charles humphrey shows that about half (!) of the researchers regard it as a "waste of funds to save" data and do not agree in "archiving is integral". indeed, a marketing effort with argumentation is required towards both producers and consumers. presentations from the conference are available for view at the iassist web-site http://www.iassistdata.org. from this toppage you can then click "multimedia presentations at the iassist 2001 conference". enjoy! karsten boye rasmussen http://www.iassistdata.org gesis 6 iassist quarterly summer 2012 iassist quarterlyiassist quarterly abstract since its initial publication over a decade ago, the oais reference model, its concepts and terminology, have become essential to the digital preservation discourse. in this discourse, the topos – or myth – of “oais compliance” continues to play a central role as archives and repositories seek to demonstrate their fitness for the challenge of digital preservation. after briefly considering what oais is (and can be used for) and what it is not – namely, an abstract reference model, but not an architecture that can be implemented directly –, we will use the gesis data archive for the social sciences as an example of mapping oais to an existing archive. we will then explore positive effects and benefits, as well as difficulties of completing this process. thus, such a mapping can be taxing for an established archive: as most of the workflows have grown and proven their adequacy over a considerable period of time, taking a step back and viewing these processes from a new perspective is a challenge in itself3. keywords oais, standards, trusted digital repositories introduction the importance of the reference model for an open archival information system (oais), which since the releases of its first draft versions in 1997 and 1999 has shaped and influenced digital preservation discourse like hardly any other model, is undisputed (see, for example, lee, 2012; allinson, 2006; oßwald, 2010). the oais standard has not only provided us with a common language – and thereby a common understanding of what it is that archives do when they preserve digital information objects; is has also given important impulses to move towards greater standardization in the field of digital preservation, including the development of criteria and procedures to analyze and assess archival preservation and dissemination practice (e.g. iso 16363:2012 “audit and certification of trustworthy digital repositories”). despite – or possibly because – of the model’s influence, the ubiquity of its terminology and concepts, one frequently encounters misconceptions as to what oais is and what it is for. often, these seem to be linked to a misunderstanding of what a reference model is. on a more concrete level, it is the notion of oais compliance and its – sometimes seemingly unreflected – use in archive self-portrayals or in the description of software packages which appears problematic. oais is a reference model often, one will hear or read about “oais being implemented” in some organization or another. what is usually meant by this is that a system is being built or adapted which conforms to the oais model in some way. such statements are misleading, however, because as a reference model, oais can de-mystifying oais compliance: benefits and challenges of mapping the oais reference model to the gesis data archive by natascha schumann1 and astrid recker2 reference model for an open archival information system (oais) ... iassist quarterly summer 2012 7 iassist quarterly by definition not be directly implemented: it is an abstract and highly generic conceptualization of a preservation and dissemination environment. thus, as defined by the organization for the advancement of structured information standards, “[a] reference model is an abstract framework for understanding significant relationships among the entities of some environment...” (n.d.). as such, it ”is not directly tied to any standards, technologies or other concrete implementation details, but it does seek to provide a common semantics that can be used unambiguously across and between different implementations” (ibid.). this is in accordance with the purpose of the oais model as given in the standard itself, which among other things states that the model: • ”provides a framework, including terminology and concepts, for describing and comparing architectures and operations of existing and future archives” • ”provides a framework for describing and comparing different long term preservation strategies and techniques” (ccsds 2012, p. 1-1) the oais reference model can be compared to a language operating on a meta-level, allowing us to speak about archives, their architectures and processes. therefore, oais will not make a certain preservation strategy or technique a requirement. it will define the characteristics of a strategy it deems successful, but it will not prescribe a concrete, monolithic solution. this means that oais cannot be used as a check list which can be ticked off as one builds an archival information system. instead, to make this meta-language useful in building such a system, a translation process is required to create an architecture from the reference model which can then in turn be implemented. in this process, abstract oais concepts have to be translated into concrete system elements and processes tailored to work in a specific environment. it is for this reason that to speak of an oais implementation is misleading. while this may seem quibbling over details, it is important to understand that the oais reference model will not translate into a real-world system seamlessly, and that this has an impact on the notion of oais compliance as put forward in the model, and as interpreted or translated by a given archive or preservation service provider. a mythical creature: oais compliance because of the oais model’s abstract character, the notion of oais compliance is, as has been pointed out repeatedly, “necessarily vague” (lavoie, 2004, p. 17). to comply with the oais model means complying with a set of very abstract requirements which themselves need interpretation, translation, and concretization if they are to be useful. thus, the standard itself makes only two requirements for compliance: 1 “support” (itself a rather vague notion) of the oais information model described in chapter 2.2, including among other things the concept of information packages composed of content information and accompanying metadata. 2 fulfill the set of mandatory responsibilities described in chapter 3.1 of the standard (see ccsds, 2012, p. 1-3). the latter (see box 1) are high-level requirements that, as beedham et al. note, “it would be difficult for any functioning archive not to comply with” (2005, p. 10). box 1: oais mandatory responsibilities the oais shall: • negotiate for and accept appropriate information from information producers. • obtain sufficient control of the information provided to the level needed to ensure long term preservation. • determine, either by itself or in conjunction with other parties, which communities should become the designated community and, therefore, should be able to understand the information provided, thereby defining its knowledge base. • ensure that the information to be preserved is independently understandable to the designated community. • follow documented policies and procedures which ensure that the information is preserved against all reasonable contingencies, including the demise of the archive, ensuring that it is never deleted unless allowed as part of an approved strategy. there should be no ad-hoc deletions. • make the preserved information available to the designated community and enable the information to be disseminated as copies of, or as traceable to, the original submitted data objects with evidence supporting its authenticity. (ccsds, 2012, p. 3-1) regardless of this vagueness, “oais compliance” has almost become a topos in digital preservation discourse, a label that is applied to repositories and their hosting institutions “to underscore [their] trustworthiness” (ccsds, 2011, pp. 1-1). yet this label remains largely meaningless without context and specification. thus, for any organization or repository labeling itself as oais-compliant, it has to be clear what this is taken to mean – that is, how the vagueness of oais compliance has been translated into a concrete set of criteria in a given case. these criteria might be something so complex as those laid down in the above-mentioned iso standard. but, as lavoie explains, to be oais-compliant could also quite simply involve using “oais concepts, terminology, and the functional and information models” when building a digital archive or preservation system; or oais-compliance can be the result of a mapping process in which “the various components in the archival system [are matched with] the corresponding features of the reference model” (2004, p. 17). but oais compliance could also mean “explicit application of oais concepts, terminology, and the functional and information models” or ”that the oais concepts and models are ‘recoverable’ from the implementation – in other words, it is possible to map, at least from a high-level perspective, the various components in the archival system to the corresponding features of the reference model” (lavoie, 2004, p. 17). we would argue that there are good reasons to include the oais functional model in compliance testing as suggested by lavoie. thus, in particular, it can be assumed that in order to fulfill the oais mandatory responsibilities, an archival information system also has to perform the functions described in the standard. accordingly, it is the second approach described by lavoie that the gesis data archive adopted in testing oais compliance; it thus followed in 8 iassist quarterly summer 2012 iassist quarterly the steps of the uk data archive and the icpsr (see beedham et al., 2005; vardigan & whiteman, 2007). mapping the gesis data archive to oais the gesis data archive was originally founded in 1960 at the university of cologne as the central archive for empirical social research (zentralarchiv für empirische sozialforschung), europe’s first data archive in the social sciences. in 1986, it became a member of the newly founded gesellschaft sozialwissenschaftlicher infrastruktureinrichtungen (gesis). since 2007, the data archive is one of five scientific departments of gesis – leibniz-institute for the social sciences, germany’s biggest research-based social sciences infrastructure institution. it is also a member of cessda, the council of european social science data archives, dedicated to improving standardized access to social science research data in europe (see http://www.cessda.org). since its foundation, the gesis data archive has undertaken continual responsibility for preserving social science research. collecting primarily digital data from empirical social research, the data archive currently holds more than 5,100 studies equaling over 600,000 files. to consolidate and demonstrate its status as a trustworthy provider of preservation services, the data archive has embarked on a series of self-audit and certification activities. the first step, now almost completed, is the application for the data seal of approval (http://datasealofapproval.org/). from these activities resulted a decision to test oais compliance by carrying out a mapping of the gesis data archive to the oais reference model. the objectives of this mapping are the following: • gain a more structured overview of workflows and preservation/ dissemination processes; • identify and close possible gaps in these workflows and processes; • introduce oais terminology and concepts to support communication within the archive and with other organizations. to achieve these goals, a mapping between the archive and the oais functional model, as well as an application of the concepts from the oais information model, are currently being carried out. in the following, we report briefly on the procedure and first results of our functional model mapping. functional model mapping the main tool to carry out mapping was a simple spreadsheet listing oais functions and the different processes/responsibilities that these comprise. for each of these sub-processes we then determined the following: • who is responsible within gesis (organizational unit down to team level)? • is the process carried out by a human staff member and/or is it supported by a technical system? • how is the process incorporated into archive workflows? (e.g. is it a routine activity carried out on a regular basis?) • are our activities sufficient? • any open questions or comments at the same time we created a simplified diagram of the current archive workflow containing the main top-level functions performed as data are acquired, deposited, archived, and disseminated (see figure 1). this helped in creating a general overview of where and when processes were taking place, and to match these with the functional entities of the oais model. we then started increasing the granularity of the different sections of the overview diagram by spelling out the steps carried out in a given phase and by matching them to oais functions. for the pre-ingest and ingest phase this resulted in the realization that the ingest process at the gesis data archive (which as a social sciences archive puts a strong emphasis on extensive quality control, data processing and enhancement) cannot be adequately captured by oais in this form and detail (see also vardigan & whiteman, 2007). thus, quality controls carried out during ingest, include: disclosure control; technical control of the files (format, readability, presence of malware, etc.); control of completeness; plausibility, consistency and weightings; as well fiqure 1 gesis data archive digital preservation workflow iassist quarterly summer 2012 9 iassist quarterly as format conversions. if any problems are discovered, further communication with the data depositor may be necessary in order to clarify the discovered issues and to correct mistakes. although this does not pose a problem for oais compliance, as the standard itself acknowledges that “[t]he complexity of this ingest process can vary greatly from oais to oais, or from producer to producer within an oais” (ccsds, 2012, p. 4-52), it does complicate the mapping process. it further became clear that some of the activities performed during ingest at the data archive are placed in different functional entities in the oais model. this caused us to “re-allocate” some of the functions to accommodate the actual data archive workflow4. as a consequence, our ingest comprises of functions from the oais functional entities ingest and administration among others, which are performed by several members of archive staff. it should be noted that this, too, is accounted for by the standard, which clearly states with regard to the functional model: “however, this is not to be taken as a recommended design or implementation, and actual implementations are not expected to have a one-to-one mapping to the functions shown, and may for example choose to combine functions or break out functionality differently” (ccsds, 2012, p. 4-3). yet, this makes the mapping less straightforward and hence more time-consuming. benefits and challenges we primarily benefited from the mapping in three areas: communication, self-reflection, and process evaluation. as noted in beedham et al. (2005, p. 82), oais – with its clearly defined vocabulary – can support communication within, or between organizations, by offering a common language. thus, one stated purpose of the oais standard is to provide the digital preservation community with a vocabulary composed of terms “that are not already overloaded with meaning so as to reduce conveying unintended meanings. therefore it is expected that all disciplines and organizations will find that they need to map some of their more familiar terms to those of the oais reference model” (ccsds, 2012, p. 1-5). this mapping process has begun at the gesis data archive, and while the introduction of this new vocabulary and its establishment in everyday communication is a gradual process taking its time, oais terminology’s potential to help ensure that staff are speaking about the same things is already apparent5. at the same time, mapping the elements and processes of the gesis data archive to the oais functional model and vocabulary has fostered self-reflection. as mentioned above, the data archive has grown over decades, and while we are certain that our digital collections are in expert hands at the archive, mapping to oais gives us the opportunity to – figuratively speaking – take a step back to analyze our daily routines and procedures. looking at these routines through “oais glasses” and applying oais terminology to them, has helped us gain a more systematic understanding of the workflows that take place as we go about our daily work, the communication processes supporting them, and the roles and responsibilities of the actors involved. finally, undertaking the mapping helped us not only to identify and name processes, it also allowed us to evaluate them. thus, we were able to spot gaps in our routines and to plan and implement strategies to close them. as was to be expected, the mapping process was not without challenges. leaving aside the difficulty that the scrutinizing and questioning of accustomed daily routines can pose for any established organization, some features of oais functional model itself contribute to making the mapping more difficult. while the following account is certainly not exhaustive (it does, for example, leave aside the problem of vagueness mentioned earlier), some of the problems we identified can be summarized under the headings ‘simplicity vs. complexity’ and ‘formalism vs. pragmatism’. simplicity vs. complexity as already discussed for the ingest functional entity, mapping the data archive to the oais functional model was complicated by the fact that often oais did not seem complex enough to model functions and processes performed by the data archive and its staff. on the other end of that scale, we find oais functional entities and functions which seem relatively simple and straightforward, but turn out to be highly complex in mapping. such unexpected complexity can occur particularly when single functions are performed jointly by several teams and/or departments. in this case, mapping the function entails documenting all the communication processes taking place between the actors and the systems involved in this process. as beedham et al. observe for the data management function, this leads to “an ‘explosion’ of mappings to all the different systems and processes that an archive performs” (2005, p. 47). another example of this is “establish standards and policies,” which is part of the administration function, and which at the gesis data archive cuts across different teams and departments (see table 1) formalism vs. pragmatism as beedham et al. observe with regard to the uk data archive and the national archive, “the oais standard can sometimes be overly bureaucratic and overconcerned with processes. realistically organisations like ukda have to be more pragmatic in their approach to decision making. . . . the oais reference model only provides a formalised view of the functions of digital archiving; it does not prescribe implementation strategy or management style. nevertheless, a real archival organisation never operates quite as ‘cleanly’ as the oais model envisages” (2005, p. 53). 10 iassist quarterly summer 2012 iassist quarterly the same is certainly true for the gesis data archive, as problemsolving, planning, and decision-making processes can be less formalized than those described in the oais functional model. thus, we will often find that the data archive performs all the processes of which an oais function is composed. however, many of these processes take place as part of routine, team (or department) internal communication, which may take many different forms (as well as degrees of formality) and which will not always be explicitly labeled as pertaining to a given oais function. thus, the fulfillment of a certain oais function becomes a sideeffect of certain communication processes. without question, the mapping process helped us identify possibly critical processes where we need to introduce more formality; for example, in the case of the monitor technology function. greater formalization may also be needed in cases where functions, which according to oais communicate with each other, are fulfilled by one and same person. as in this case no real communication (e.g. in form of requests and responses to these requests) takes place, the use of additional documentation tools (e.g. check lists) will have to be considered to create more transparency and to ensure that all necessary steps are taken. however, as beedham et al. point out, the oais reference model only lists and describes functions without specifying how exactly they should be implemented (see 2005, p. 53). this means that it is really up to an archive to decide how frequently, and in which form, a process should take place. does fulfilling the “establish standards and policies” function require regular meetings (on team, department or institutional level)?; or, can certain processes be adequately addressed in ad-hoc communication between the staff members involved? as already mentioned, what helped in assessing our current practice in the course of the mapping process, was to identify the level of “formality” with which a certain oais function is fulfilled by the data archive. this was achieved, for example, by indicating • whether a function is performed routinely (that is on a regular basis, or for every dataset submitted to the archive), or only on request/as required (and hence reactively rather than proactively; on this aspect see also beedham et al., 2005, p. 53); • whether checklists, minutes, or other records exist to document the process; • and whether in our opinion the level of formality was sufficient or not. however, in some instances the mapping may also result in a conscious decision to deviate from oais entirely for reasons dictated by the specific environment in which the archive operates. degrees of compliance as the gesis data archive’s experience with the ongoing mapping process illustrate, oais compliance – in contrast to compliance with a certification standard or criteria catalog such as iso 16363:2012 – is not 1 or 0, yes or no. rather, we would argue, compliance comes in degrees. thus, what oais compliance means is really a matter of interpretation, an act of filling in the gaps that the reference model necessarily leaves with context information from of our own organizations. sometimes, this context may even lead to the decision to not comply with a certain aspect of oais. we would argue that such decisions, as long as they are well-founded, do not necessarily compromise oais compliance. however, this illustrates once more that without providing enough of the context information specific to the archival system (which as a reference model oais must necessarily ignore), and without spelling out and making transparent the acts of interpretation performed in translating the reference model into something like a checklist, the statement “we are oais compliant” remains utterly meaningless. references allinson, j., 2006. oais as a reference model for repositories: an evaluation. available at: [accessed 20 august 2013]. beedham, h., missen, j., palmer, m. and ruusalepp, r., 2005. assessment of ukda and tna compliance with oais and mets standards. colchester, essex: uk data archive, university of essex. available at: [accessed 28 june 2013]. ccsds, 2011. audit and certfication of trustworthy digital repositories. recommended practice. available at: [accessed 28 june 2013]. ccsds, 2012. reference model for an open archival information system (oais). recommended practice. available at: [accessed 28 june 2013]. lavoie, b.f., 2004. the open archival information system reference model: introductory guide. dpc technology watch series report 04-01. available at: [accessed 28 june 2013]. lee, c.a., 2010. open archival information system (oais) reference model. in m. j. bates & m. n. maack, eds. encyclopedia of library and information sciences. taylor & francis, pp. 4020–4030. oßwald, a., 2010. kapitel 4: das referenzmodell oais open archival information system. in h. neuroth et al., eds. nestor handbuch. eine kleine enzyklopädie der digitalen langzeitarchivierung. version 2.3. available at: . [accessed 20 august 2013] vardigan, m. and whiteman, c., 2007. icpsr meets oais: applying the oais reference model to the social science archive context. archival science, 7(1), pp.73–87. notes 1. natascha schumann is affiliated at the data archive for the social sciences at the gesis leibniz institute for the social sciences in cologne. the main focus of her work is on digital curation of social science research data and audit and certification in this area. her contact email is natascha.schumann@gesis.org. 2. dr. astrid recker works at the gesis data archive where she is responsible for the design and delivery of digital preservation workshops for the “archive and data management training center.” she is also involved in the eu-funded project data service infrastructure for the social sciences and humanities (dasish). her contact email is astrid.recker@gesis.org. 3. this paper is an updated version of a presentation given at the iassist 2013 conference in the session ”beyond bits and bytes: the organizational dimension of digital preservation.” 4. a similar observation is made in beedham et al. for the administration function (2005, p. 48). 5. it is similarly clear, however, that oais terminology will not become the only vocabulary with which the data archive operates. for example, the communication with stakeholders – data producers iassist quarterly summer 2012 11 iassist quarterly and users in particular – requires the use of a different, less ”technical” vocabulary. similarly, internal communication, too, will continue to use non-oais terms and concepts where we regard these as (more) appropriate. yet, it is important to be able to use oais terminology as a point of reference in cases of ambiguity or unclarity. moving to distributed computing: experiences from the minicomputer transition by halliman h. winsborough ' social sciences computing cooperative university of wisconsin, madison where we are going in these remarks, i take the phrase "distributed computing" to indicate the expected computing environment of the next several years rather than its more technical and narrow meaning. most of us will want to move in the direction of this expected environment in order to do our work with competitive efficiency. i think this environment has four elements: 1 . powerful processing is accessible from all users'desks. that power is likely to be many times greater than that available in the past. many computers may be involved in making that happen. they are accessible from every desk that needs access. they are also accessible from home, hotel room, laptop, and god help us from the car. 2. powerful connections are available from the user's desk. the desk top, lap top, car top, machine, whatever it may be, is connected through the electronic networic to other machines locally and to the national and international elecffonic networks. electronic mail is the "normal" mode of communication locally, nationally, and internationally. in principle, data at remote locations can be accessed easily. software that is legally available lo the user can be accessed remotely. the user may run programs on her own machine or on the remote machine. in the best of these ideal worlds, running locally doesn't requires recompiling. 3. maintenance of all these wonders is invisible to the user. machines are connected, repaired, and replaced. files are backed up. important new files are added and potential users informed of their availability. documentation is maintained and improved. programs are checked for accuracy. network addresses are updated. network protocols and even physical connections are changed. all this behind the scenes, as it were. 4. openness prevails. there is standardization of operating systems, editors, and programs. as aresult, a user can work on a new machine or on a remote machine with only modest additional training. in the best of these worlds, standardization pertains to data as well as systems. in this world, one would retrieve data from, for example. dialog, cendata, and icpsr using the same "language." no doubt some of this description seems hopelessly utopian, even to the most enthusiastic among us. but a good deal of it is currently in place. powerful machines are here. we just proposed a sparcstation 10 for a faculty member. at about $10,000, it will compute at 85 or so mips. that is mainframe speed. several competitors do as well. but even that kind of power is not sufficient for one of the faculty members i serve. he routinely ships jobs from his desk to a supercomputer in san diego. communications improvements abound. electronic mail is a commonplace. i suspect the organizers of this conference wonder how they could have done their job without it. i also suspect that remote access to data is an ongoing theme in this association. i will have more to say subsequently about what we must do in order to make data access fit the new computing environment. when it comes to maintenance and support, things get more speculative. a lot goes on without the user knowing about it but sometimes the behind-the-scenes machinery creaks pretty loudly and an occasional flyer falls on the cast. that is because distributed processing can get pretty complicated. the tangle of things can get so dense that it is hard to see the bug before he bites. openness and standardization is in process but not very far along. as of today, trends are mixed about how well this user-demanded principle will stand up to corporate proprietary urges. a year ago all the big players were marketing openness. but recent events suggest a reircnchmenl the ace consortium looks moribund. sun doesn't even make a c compiler for its new machines, so it is a bit harder to eschew solaris for bsd than it was. and so on. over all, then, there is a lot of progress toward the ideal disunbuted computing environment but a lot of room for uncertainty as well. as i visit my colleagues at other universities, i sense quite a lot of uneasiness about the transition from whatever kind of computing they currently have to the new environment. my informal survey lassist quarteriy suggests that how awesome, impractical and distant the norm of distributed processing seems depends a good deal on where you start from. where we are coining from people are facing the transition to distributed computing from a number of different current environments. all of those environments have elements of the future in them some more than others. in the following i will distinguish three types of startpoint environments mainframe shops, personal computer shops, and minicomputer shops and discuss how the transition to distributed computing looks from each vantage point. the mainframe shop by a mainframe shop, i mean a group that depends on a large, centrahzed, computing "utility." people from this environment are used to quite powerful machines and find nothing very exciting about a computer that turns out 85 mips. it is what they expect. they are also used to a pretty high level of invisible maintenance and technical support; so good, in fact, that it leads to change-resistant users, as we will see. the organization of access to data in such a shop can be superb. but often it is not. connectivity is less familiar to people from the mainframe environment it is unusual for everyone in a mainframe shop to have a terminal on their desk. batch processing remains a main mode of work. although interactive computing is available from mainframes, it is pretty pallid stuff. you go to a terminal to create and submit a batch job. ibm has introduced profs recently to f)ermit local communication, but it doesn't have the same presence as e-mail does when everyone has a connected machine on their desk. openness doesn't exist in mainframe shops. enough of them use the same vender's equipment, though, that movement from one mainframe shop to another is fairly easy. thus monopoly substitutes for openness and the only victim is price. i think it is people from mainframe shops who react most violently to the prospects of a transition to distributed computing. those who haven't begun the transition are most resistant those who have made serious strides toward distributed computing are the most ecumenical. partly, i think, it is because the mainframe mavens have done such a good job of making things transparent in so doing, the mainframe priesthood has shielded social science users from the grubbier aspects of computing by making them appear an esoteric mystery so much so as to produce a kind of learned helplessness in the users. an important part of making the transition to distributed computing is to take some things into your own hands. that prospect can look remarkably dangerous, even sacrilegious, to oldline mainframe users. once convened. well, it is like the old saw. besides, the new environment is worlds better. the personal cwputer shop pc shops, until recently at least aren't really shops. the big thing about a personal computer is that it is personal. it's yours. it's on your desk. you take care of it buy software for it, install the stuff, decide when to upgrade the operating system and do it yourself, back it up, defragment its little disk, change the battery for its clock and install new boards, interfaces and disks. the idea of doing it yourself isn't daunting to people from the pcworld. it is just a bore. invisible maintenance can seem like a dream, especially when your disk crashes and you realize you forgot to back up last night people from the pc world are also pretty comfortable with interactive computing. they expect "standards." they also are often quite interested in more computing power, sometimes to a level of fixation that raises my freudian eyebrows. i think it is the connectivity of distributed computing that gives pc {people the most trouble. it is all so un-personal. connectivity and the consequent standards reduce the user's freedom to do anything on "their" machine that they wish. but connectivity is beginning to catch on even here. witness the success of compuserve. the result of all this is that pc users are a lot more eager than mainframe people to make the transition to distributed computing. most pc people are eager to have a powerful, networked unix box on their desk. they just insist that the desk be big enough to hold their pc, too. the minicomputer shop in a classic minicomputer shop, users have terminals on their desk that are connected to a rather modest computer. such shops start off closest to distributed computing. one accesses computing cycles from the desk. computing is interactive. communications with one's own work group are quite facile. wider area access to cycles, data, and software has been in place for some years. maintenance is pretty invisible. if your mini run unix, many of the things listed under my openness rubric were there, too. if you run one of the proprietary operating systems, such as vms, openness has been a lot slower in coming. minicomputer types generally feel that the transition to distributed computing is just a bit more of what they have been used to for a long time. the big attraction is the increased power and, for those stuck in proprietary operating systems, increased openness. one of the reasons that people from minicomputer shops face the transition to distributed computing with a bit spring/summer 1992 more equanimity than people from mainframe shops or pc shops is that they have already made important parts of the transition. because of this history of change this slower transition the experiences of one minicomputer shop may be of some use in thinking about making the transition lo distributed computing in other places. the experience of one minicomputer shop. the social sciences computing cooperative at the university of wisconsin, madison, where i work, has been operating a computing facility for social science research since 1972. for the first eight years, we operated in the "mainframe" model; complete with glass enclosed shrine, an ibm iron god, and batch processing. in 1980, we made the minicomputer transition when we gota vax 1 1/780. before long, nearly every faculty office and most of the research rooms had terminals connected to the vax. we didn't (and still don't) charge for resources used. the mail system was pretty good. suddenly our clients had copious interactive computing and were connected in an instant communications network. that was the most dramatic subjective transition that we have made. we started technical distributed processing in about 1988 when we began to distribute tasks among several vaxes that were previously independent network partners. now we operate two client-server unix systems as well as a local area vaxcluster the direct progeny of the vax 1 1/780 and a growing pathworks network of pc's and mac's. from the latter, it is easy to connect to any of the former networks. there are four aspects of these transitions that were somewhat unexpected for us. 1 pass them along in the hope that they will be of some help to those of you are just beginning the transition. 1 . it costs a lot to service fancy equipment in people'soffices. 2. teaching becomes an increasingly important activity. 3. rapidly dropping costs means that plans and policiesmust stay flexible and be reviewed regularly. 4. the social organization of computing becomes asimportant as its technical aspects. in the remaining pages i will discuss each of these findings as we experienced then. then 1 will discuss problems associated with the transition we haven't made, the transition to distributed, on-line data. equipment in offices. when we moved from the mainframe to the minicomputer, our operations people proposed the policy that our responsibility for equipment should go from the machine room to the wall plug and no further. the terminal on the user's desk was the user's problem. our organization has always been a consumer's co-op, so that policy lasted about a week. diagnosing, repairing, and replacing faulty terminals became a standard task for us. initially, the co-op provided fairly simple terminals. as time went on, people wanted fancier machines and bought them from grant funds. we took care of those, too. as pc's became more popular, many users bought one for home. before long they wanted to use terminal emulation software and call in from their home pc. so we got in the modem business and even took over some maintenance of home pc's. the emulation software worked well, and some users decided they wanted pc's in their offices rather than terminals. some place in there we should have reared back and passed a policy about what kind of equipment we would service and what we wouldn't but we didn't. so we got into the business of repairing nearly any kind of pc computer, printer, or storage device and ensuring that it worked in a civilized way with the other computers in the system. it was foolish of us. a faculty member saved s75 by buying an unfamiliar laser printer and we spent s750 in time making the thing work properly on our networks. the advent of workstations brought some order to our policies. we decided that the co-op had to apree to service a non-standard workstation before its purchase or the user was on his own. we have extended that policy to other equipment as well. of course, that meant we had to decide on what was "standard" in the pc equipment business. that is taking some time, but we expect it will have good results for both users and co-op staff. teaching teaching rather sneaked up on us, too. initially, we gave occasional lectures as introductions to our systems and to provide some training on software we had wntten. of course we have always provided fairly extensive consulting. since we are in a university, we get a fairly large batch of new users every year. before long we were doing more extensive training of new users training designed to reduce the burden of answering the same question over and over again in consulting. then the people who teach statistics decided they wanted us to take over more of the training in how to use the statistical software. so that got added to our teaching portfolio. with the addition of unix to our operating system mix, we are doing more short courses in the operating system and its editors. lassist quarterly a new addition to the list next school year will be instruction in sql. we have taught about relational ideas and data normalization for several years but instruction in sql will be a new addition. one result is that over the years we have added personnel in the consulting and teaching part of the staff. user services, as we call these functions, are about 1/3 of our staff activities. we did not expect it to grow to such a large fraction. of course it would have been possible for our organization to have avoided doing many of these things. but they represent rsal user needs. if we didn't satisfy them, they wouldn't just go away. things are cheap it is wonderful that the price of computing equipment has fallen so dramatically in the past decade. keeping up with the changes can be a problem for a computing organization, however. not only do the people in charge of buying things have to keep their information refreshed but also one must re-think policies on a regular basis to see if they were made contingent on a particular price environment. take disk space, for example. we initially allocated new users 2000 blocks of disk space on the vax. that was when a 75 meg disk for the vax costs $20,000. it became a kind of rule of thumb that lasted much too long into the dramatic decline in disk prices. we now try to regularly review policies to see if they are outmoded. new employees can be especiallyhelpful in detecting these residues of previous price regimes. the social organization of computing the flexibility of technical computing arrangements has grown so dramatically in the past several years and the price of computing has gone down so dramatically that we cunently believe that the greatest leverage in computing efficiency can be achieved by using the new flexibility to modify the social organization of computing. three organizational modifications have been particularly useful to us. first, we have become a consumer's co-op user owned as it were. second, we deal with money in a special way. we don't charge for computing. co-op members agree to contribute to co-op costs from their budgets. third, we use the flexibility of modem computing to "fit" the unique work-group style of users. we don't have much pride of invention about these arrangements. like many opportunities for organizational change, they rather happened lo us and we tried to keep the ones that looked promising. initially we were the computing arm of the center for demography and ecology. in the mid-1980's, several other organizations on campus came into some computing money and decided they wanted to join with cde in providing services to their members. since there was a very considerable membership overlap between cde and these organizations, it made a lot of social and political as well as economic sense to try to achieve the expected aggregation economies. the growing flexibihty of computing made this organizational arrangement possible. that's when we formed the co-op. in this new organization, each of the "sustaining" organizations has a more or less equal say in what goes on. policy decisions and oversight are performed by a "steering committee" made up of representatives from each agency. the budget is decided annually by the chairs and directors. it has worked pretty well so far. the non-faculty computing director has the committee as boss. when agencies' needs conflict, he can ask the committee to decide how to play fair rather than making it up himself. as you can see from the foregoing, we deal with money and accountability in a special way. agencies decide each year how much they should conoibute to the expenses of the co-op. agencies own some of the machines that we run and pay the attendant software, maintenance, and supply costs for those machines. other machines are held in common. each agency pays a share of the cost of those machines. we have used the flexibility of the various operating systems to keep the permissions straight in this arrangement. users are authorized on machines belongingt o agencies they are members of and on common machines. the accounting system keeps pretty good track of who's doing what on all the machines and what agency is responsible for the time. the notion of common machines is more flexible than one might initially suspect. certainly servers are common machines. but we also retain some older and smaller vaxes as common machines because software is cheap on them. we have them loaded up with software that is used only occasionally by any one group but is cost-effective to license on a small machine for the whole co-op's use. finally, we use the flexibility made available to us by distributed processing to fit a work group's computing as closely as possible to its special needs and style. for example, most co-op members have been fairly happy with our system for using tapes. operators are on duty about 18 hours a day and do the tapemounts. the institute for research on poverty, however, has a group of programmers that do quite a lot of work with large filescps and the like. they very much like to mount their own tapes. so irp has a tape drive on one of its machines in a room accessible to its programmers and they do their own mounting. data access in a distributed environment the last issue i want to address is the one of data access spring/summer 1992 in the distributed environment. i think this is an issue we all face. a crude way of putting it is, "what will we ever do without round tapes?" some people seem quite far along. jim jacobs with the social science group at san diego has a wonderful jukebox/menu-interface. it is the neatest thing i have seen. al anderson in the demography group at michigan has a plan for data to be delivered from a campus data utiuty over local fddi to a risc machine with an enormous main memory space for buffer. it is the most ambitious thing i have seen. in the co-op we are moving fairly slowly to rid ourselves of round tapes. at the same time, we aren't buying replacements for the nearly worn out ones we have. after several years of thinking, visiting other installations, and trying things out, we have come to an important conclusion for our shop. it was really tom flory's insight it looks like the big issue is the media you will use next; whether to go to worms, mo's, dat's, or 3480's. but that's probably unanswerable without knowing how you are going to use the equipment. we think that the place to start is with the interface. what should the user's access look like? what kind of tools for extracting data should be available? do you need to do complex joins as well as restriction and projection? how frequently? how is information about the data to be coordinated with the access process? how is one to implement solutions to these problems in a way that is reasonably open and standard? these questions and the others that arise in answering them are bedeviling us currendy. when the only media was round tape, the answers to these questions were fairly constrained because serial access is fairly constraining. we can now debate about the most amazing things: is it more "standard" to preserve archival provenance and keep the data in the form we get it from the distributor or is it more "standard" to rearrange and decompose files to satisfy, say, third normal form? should we use a commercial data base, say ingres, to organize the data and make relational joins possible? or can we get along with what you can do in sas and spss? we haven't come to any grand solutions to these problems. we lean toward normalizing the files and keeping them as ascii files. for the moment, our solution to the media problem is to buy quite a number of scsi drives. we will keep the most used data online, probably in compressed form, on these devices. our interface decisions will be made assuming that whatever media eventually is favored, it will be possible to make the machine think it is just anodier directory. it is an exciting lime for all of us in the computing business right now. it is probably most exciting for those of us who deal in data. for the first lime, there are the facilities out there at a reasonable price for us to serve our users much more effectively. if we can now just manage to do it in an open and standard way, all will be well. 1 presented at die lassist 92 conference held in madison, wisconsin, u.s.a. may 26 29, 1992. the center for demography and ecology receives core support for population research from the national institute for child health and human development (p30 hd050876). lassist quarterly vol252 iassist quarterly summer 2001 29 the first internet connection in latvia was created in 1991 during the so-called “putsch” when the connection via satellite from the university of latvia to sweden was the only mean of communication with the outside world. since then the internet has been developing in latvia. still there is a shortage of public financing and it specialists that hinders further progress in this area. in 1999 “the national program of information” and in 2000 a social and economic program “e-latvia” were approved. the aim of those and several other specific programs is to promote the establishment of the national information structure. the number of internet users in latvia at the end of 2000 was approximately 150 thousand, i.e. 10% of working age population. internet access is available in 14% of schools and only 5% of public and research libraries. there is also a great lack of computers and still more money is spent on buying hardware than on purchasing software. at the same time the number of internet users was 50% more that at the beginning of 1999 and the number is increasing rapidly (economic development of latvia. report. ministry of economy, republic of latvia, riga, december, 2000, p.92 93). still we are far from reaching the standards of the leading countries where 60% of population are it users. all the above mentioned problems affect the development of the electronic archiving and exchanging of data in latvia. the establishment of an electronic on-line archive reflects a general trend to move away from paper-based information exchange among the researchers who specialize in social sciences and humanities. compared to traditional archives, the development of electronic data bases and other electronic services do not develop dynamically in latvia. there are several reasons for it: 1. the results of many social studies in latvia are preserved only partly, but due to the insufficient documentation the use of them is very limited. 2. even sufficiently documented data are accessible only to specialists, as the use of them requires expensive software. 3. the opportunities of scientists to inform the society on the results of their research and the opportunities of the society to obtain the necessary social information are limited. the latvian social science data archive (lszda) is currently being developed with the financial support of the latvian science council (the program ‘social development and social security’, grant no. 93.330) and the soros foundation latvia (transformation of education program). compared to other data archives, the technical aspects of the lszda are, for financial reasons, rather limited. the lszda is supposed to serve the following purposes: 1. to gather and maintain social science data that characterize latvia and the latvian social scientists; 2. to facilitate the development of the latvian social sciences and to open access to research results, facilitating the dissemination of data and documentation; 3. to make data available for secondary analysis, to improve the quality of research and the credibility of results; 4. in the area of educational aspects, to serve as an information base for training in the social sciences as well as to provide the analysis opportunities of the sociological data to a broader range of users. these activities will facilitate the more efficient disposition of resources that are granted for research purposes. the archive also tries to provide latvian scientists and other interested parties with access to data archives in other countries, as well as provide foreign researchers with opportunities to learn more about latvia and its social sciences. at the present moment a number of articles (mostly in english) is available. the publications of the latvian social scientists who worked outside latvia during the soviet occupation are represented very widely in the bibliographic file which has more than 1,000 entries covering the time period from the beginning of the 20th century to the 1980ies. there are no limitations for access and all information is available free of charge. the latvian social science data archive by ausma tabuna 1 30 iassist quarterly summer 2001 until 1999 the data collections were accessible only to authorised people who used spss software. one of the urgent tasks of the archive was to create a data analyses possibility for a broader range of the people concerned and especially for the students of sociology and other social sciences. in 1999 the data documentation system of the swedish social science data service (gothenburg) was mastered, which allows the analysis of the data collections to those users who have no access to the spss software. these data collections can be processed on-line or by maintaining them on a cd-rom for the users without any background knowledge. in 1999 the data from three modules of the international social survey programme (issp) were documented (national identity, role of government, religion), and an agreement has been achieved on the possibility to use this system for the documentation of further studies. contact: ausma tabuna, lszda, akademijas laukums 1, riga, lv-1940, latvia. phone: +371 7227110, fax: +371 7210806. internet: http://www.lszda.lv. e-mail: ausmat@lza.lv (an outline of) the impact of future social and technological trends on the dissemination of census bureau information by donald l. day ' abstract this study examines social and technological trends that may impact the dissemination of u.s. census information via the depository library program in the year 2000 and beyond. the study looks beyond currently emerging systems to examine a limited list of future issues in technology, regulation, funding, access, and user demand. it examines information dissemination in the broad, societal context, rather than concentrating narrowly upon the means of delivery. its main objectives are to pinpoint key issues, to stimulate an appreciation of the inextricable nature of information in postindustrial society, and to recommend policies and directions for further research. contact: donald l. day, 4-284 center for science and technology, syracuse university, syracuse, ny 132444100 usa. bitnet d01dayxx@suvm. study method elite interviewing was interspersed with a review of literature in the future studies field. key research questions the key research questions drafted from an analysis of the literature and during the interviewing process were as follows. technology 1 . what will be the leading edge information technologies in the first decade of the next century? 2. when and to what degree will depository library materials (especially census data) be distributed via cdrom or other machine-readable media? 3. what technological developments will affect patrons' remote electronic access to depository libraries? 4. what software and data structures will be required for electronically disseminated census data, to facilitate rapid and effective searches and retrieval? regulation 5. what will be the sponsorship and impact of standardization efforts to facilitate network access to federal government information? 6. to what extent will anti-trust concerns inhibit development of data integration protocols and telecommunications software necessary fw widespread network access to federal government information? 7. how will the distribution of government information be controlled, under whose auspices and with what objectives? 8. how will data integrity be maintained without impeding widespread electronic dissemination of information? funding 9. which sponsors of information production, dissemination and use will support high technology access, under what conditions and with what goals? 10. what are the prospects that congress will choose to privatize depository library distribution? what impact would that have upon the quality, quantity, availability and cost of census bureau information? access 11. what will be the minimum skill levels required of users and depository librarians in accessing electronically disseminated information? 12. to what degree might user fees and other costs of accessing electronically disseminated information disenfranchise individuals? user demand 13. what impact will changes in work force composition and employment arrangements have upon the types of census information sought by users? overview of future trends • strong rise in "knowledge industries" • increase in non-english speakers • swelling ranks of citizens over 65 18 assist quarteriy • spending will continue to shift toward service industries (88% of work force in 2000) • job retraining programs for 4% of work force • seventy percent of u.s. homes may have computers in 2000 • people changing careers an average of every 10 years • ranks of the self-employed will grow at a faster rate than salaried workers • more mid-career professionals will become entrepreneurs • do-it-yourself activities will be popular, because a 32-hour work week will create more leisure time and due to the high cost of services • massive increases in storage technology • inferred major policy issues > privacy > the part government plays in information dissemination > intellectual property rights > functional literacy knowledge in postindustrial society social organization will be shaped by intellectual technology in postindustrial america in accordance with what is known as the "knowledge theory of value". knowledge, even when it is sold, remains with the producer. it is a "collective good"— once it has been created, it is available to all. there is little incentive for any single person or enterprise to pay for the production of knowledge unless a proprietary advantage (such as a patent or copyright registration) can be obtained. thus, government policy in regard to intellectual property and contractor marketing of publicly funded products will be key in the management of future information dissemination technology. a reduction in incentives for individuals or companies to produce knowledge will cause the responsibility for and costs of satisfying information needs to fall to government whether information dissemination is "privatized" and in what manner may affect the availability of that information significantly. the u.s. as postindustiral state the optimistic view 1. centrality of theoretical knowledge as the basis of innovation. 2. creation of new intellectual techniques to engineer solutions to economic (and even social) problems. 3. the spread of a (technical and professional) knowledge class. 4. the change from goods to human services. 5. a change in the character of work (people must learn to live with one another, since interaction among groups will be key). 6. the employment of women in expanded human services. 7. science as the societal standard bearer. 8. political units comprised of either vertical organizations of individuals into scientific, technological, administrative, and cultural centers, or of institutions arrayed as economic, government, university, or social complexes. 9. meritocracy (an emphasis on education and skill). 10. scarcities of information and of time. 1 1 . the economics of information. the pessimistic view new technology ... 1 . will be highly beneficial to some segments of society, but detrimental to others. 2. will have a positive impact primarily in the middleclass suburbs, with a negative impact in central cities. 3. will not be properly understood and regulated until considerable damage has been done in major urban development 4. will reduce the economic viabihty of the central city by accelerating delocalization of business and commerce. 5. will affect the service sector most, because its processes involve paper transactions that are particularly sensitive to technological substitution. five major areas that may affect future dissemination 1. technology 2. regulation 3. funding 4. access 5. user demand • technology > leading edge technologies > machine-readable media > remote electronic access > software and data structures ' regulation > standardization > anti-trust concerns > control of distribution > data integrity summer 1990 • funding > sponsors > privatization • access > skill requirements > disenfranchisement • user demand > the aging population > changes in the work force > multilingual services key issues should future information dissemination be oriented toward individual users or toward businesses and institutions? should joint ventures with private industry be pursued as a means of funding future dissemination in the face of a shrinking federal budget? what policies should be adopted regarding intellectual property rights in data analysis, access software development and copyright protection? how and where should advanced indexing and retrieval software be procured for access to machine-readable data? what role should be played in coordination of federal information dissemination policy to eliminate fragmentation of jurisdiction over media, content and formats? how should demands for multilingual presentation be addressed? to what extent is the census bureau liable for ensuring the integrity of data disseminated in machine-readable formats? would the census bureau be accountable for invasion of privacy or threats to defense or industry confidentiality that might result from the ability to manipulate data in machine-readable format (the "mosaic" issue)? to what extent should the census bureau be involved in establishment of network protocol and human interface standards both within government and within industry? how should responsibility and costs be divided for creation and maintenance of on-line access networks? should the census bureau abandon the depository library program in favor of alternative means of data dissemination, or be a driving force in effecting a restructuring of the program in keeping with new information needs and dissemination technology? what types of training should be provided for depository library staff to better enable them to deal with the challenges of new technology? to what extent should collection acquisition, operating and other funds be diverted to the purchase of hardware and software to support the use of electronically disseminated information? should the census bureau decide upon the medium and content of information disseminated based upon extent and type of use research? conclusion this study was fielded under the presumption that government will be required to continue providing public access to federal information as part of its commitment to maintaining the informed citizenry that is central to participatory democracy. the nature of that access, however, is entwined in a host of social, economic and technology issues that must be addressed promptly if the pace of change is not to overwhelm policymakers as well as information intermediaries and users. references bell, d. (1978). the postindustrial economy. in j. fowles, "handbook of futures research". westport, conn.: greenwood press, 507-514. bezold, c. & olson, r. (1986). 'the information millenium: alternative futures". washington: information industry assn. bortnick, j. & relyea, h. (1990, march 30). interview at the madison building, library of congress, washington, d.c. caddy, d. (1987). "exploring america's future". college station, tex.: texas a&m univ. press. cetron, m. (1988). into the 21st century. futurist", 22:4, 29^0. 'the clarke, a. (1978). communications in the future. in j. fowles, "handbook of futures research". westport, conn.: greenwood press, 637-652. cornish, e. (1985). the library of the future. "the futurist", 19:6, 2, 39. 20 lassist quarterly diebold, j. (1985). new challenges for the information age. 'the futurist", 19:3, 68. eldredge, h. (1978). urban futures. inj. fowles, "handbook of futures research". westport, conn.: greenwood press, 617-636. federal information center program (1990). [a background paper.] (available from the gsa information resources management service, washington, dc 20405). freese, r. (1988). optical disks become erasable. "beee spectrum", 26:2, 41-45. government printing office improvement act of 1990. [hr3849]. (available from the superintendent of documents, washington, dc). harm is averted by quick response to computer virus (1990). "administrative notes", 11 (april 13), 1 . newsletter of the federal depository library program. haub, c. & von cube, a. (1987). the united states population data sheet (6th ed.). washington, d.c.: population reference bureau. in marien, m. (ed.), "future survey annual" (item nr. 8369). washington, d.c.: world future society. 23:6, 53-60. powell, elizabeth (1990, february 16). interview at the hart senate office building, washington, d.c. tenner, e. the revenge of paper. 'the new york times", march 5, 1988, 27. weiner, e. & brown, a. (1989). human factws: the gap between humans and machines. "the futurist", 23:3,9-11. weinstein, s. & shumate, p. (1989). beyond the telephone: new ways to communicate. "the futurist", 23:6,8-12. addmonal sources the following sources were identified in the literature search for this study, but are not referenced in the final report cornish, e. (ed.) (1982). "communications tomorrow". bethesda, md.: world future society. didsbury, h. (ed.) (1982). "communications and the future". bethesda, md.: world future society. dowlin, k. (1984). "the electronic library". new york: neal-schuman. hemon, p. & mcclure, c. (1987). "federal information policies in the 1980's: conflicts and issues". norwood, n.j.: ablex. lamm, r. (1985). "megatraumas: america at the year 2000". boston: houghton mifflin. longman, p. (1988). the challenge of an aging society. "the futurist", 23:5, 33-37. marshall, c. & rossman, g. (1989). "designing qualitative research". newbury park, calif.: sage. mcgee, milton (1990, march 30). interview at the adams building, library of congress, washington, d.c. ferrarotli, f. (1986). "five scenarios for the year 20(x)". new york: greenwood press. gorman, m. (ed.) (1984). "crossroads". (proceedings of the first national conference of the library and information technology assn., sept. 1721, 1983, baltimore, md.). chicago: american library assn. naisbitt, j. (1982). "megatrends". new york: warner books. pasqualini, b. (ed.) (1987). "dollars and sense: implications of the new online technology for managing the library". chicago: american library assn. "the new american boom". (1986). kiplinger washington letter. washington, d.c: the kiplinger washington editors, inc. "1988 ten-year forecast". (1988). menlo park, calif.: institute for the future. outlook '90 and beyond. (1989). "the futurist", 'presented at the lassist 90 conference held in poughkeepsie, n.y. may 30 june 2, 1990.donald l. day, school of information studies, syracuse university summer 1990 iassist quarterly 2010 / 2011 47 iassist quarterly abstract in norway, history is the research field that has the most experience with digitizing qualitative data, and several norwegian historical milieus have established infrastructure for qualitative data collections. in addition, data on the political system at norwegian social science data services (nsd) is an example of a large collection of qualitative resources, however thematically limited. a more general central archive for qualitative social science research does not currently exist in norway. there are many examples of scattered qualitative longitudinal projects, but as far as we have learned no large archives/ collections are made available. nsd is about to increase focus on data acquisition and storage of qualitative data by expanding its archiving routine to all data that are digitally storable, starting by concentrating on the projects financed by the research council of norway. a majority of these projects have an obligation to deposit research data at nsd. one possible direction for nsd in the future is to establish general archiving agreements with major research institutes. whether project research data can be stored and/or reused is to a large degree determined by whatthe respondents have consented to prior to the data collection. therefore it is important that researchers are aware of the possibilities for archiving and reuse early in the research process. the privacy ombudsman for research is in direct contact with a large share of research projects in their initial phases to give advice both with regard to the development of questionnaires and interview guides and formulation of information letters to respondents, and can play a role in this respect. privacy protection is a challenge, especially for qualitative data archiving. nsd envisages a common solution for archiving quantitative and qualitative data. existing routines for collection, documentation and presentation of data will be expanded to include qualitative data. a challenging, but rich new area for nsd, as we see it, is first and foremost, to offer qualitative data prepared for analysis. keywords: archiving; qualitative data; qualitative longitudinal data; data sharing; secondary use; norway introduction the content of this article was presented as one of the country reports at the bremen workshop held at the university of bremen, germany, on 24th of april 20092. the main aims of the workshop were to map out qualitative datasets and resources across europe and to develop plans for a european network of qualitative data collections, researchers and projects, with particular focus on qualitative longitudinal resources. delegates from data archives, universities and institutes across europe attended the workshop representing 14 countries. a few countries, i.e. finland, ireland and the uk, had quite extensive experience in central archiving of qualitative data, while the rest divided themselves equally between the groups of “developing resources” and “potential resources”. norway placed itself in the “developing” group. as far as we know, in norway no feasibility study regarding archiving of qualitative data has been conducted. therefore comprehensive knowledge within this area is limited. however, norwegian social science data services (nsd) (n.d) is about to launch an initiative concerning this issue the research council of norway (n.d) is promoting a national data archiving policy based on the oecd guidelines (oecd, 2007) that state that all publicly funded research data that can be digitally stored should be archived for future dissemination. the research council supports these guidelines by requesting that a majority of the projects they finance should be contractually obliged mapping out qualitative data resources in norway by gry-hege henriksen, maria bakke orvik, trond pedersen1 48 iassist quarterly 2010 / 2011 iassist quarterly to archive data. data from the research projects shall, as a default, be archived at nsd within two years of the projects’ termination. these contract terms give nsd the legitimacy to actively collect data from already finished research projects. central in the collecting process are emails to project leaders reminding them of their obligations, including uploading of metadata via an online archiving form. this is a routine that covers both qualitative and quantitative research data. sharing research data is relatively well developed in norway for projects financed by the research council of norway. until the project leader or data owner has published his or her findings, the data are usually under embargo. existing qualitative archiving infrastructure history is the research field that has the most experience with digitizing qualitative data, and norwegian historical milieus have established infrastructure for qualitative data collections. examples of these are the digital archives, museum of cultural history, the norwegian historical data centre and the documentation project3. data on the political system4 at nsd is an example of a large collection of qualitative resources, which focuses thematically on various aspects of the norwegian parliament, government, political parties, civil service, etc. to our knowledge, a more general central archive for qualitative social science research does not exist in norway as of today. nsd is now about to increase its focus on data acquisition and storage of qualitative data by expanding its archiving routine to all data that are digitally storable, starting by concentrating on the projects financed by the research council of norway. we will also update the information about archiving and the archiving form at nsd’s homepage. archiving of qualitative data can, technically, use the same routines that today are used for quantitative data. but adapted metadata, standardising actions and new tools for analysis and processing are needed. routines for updates and maintenance will probably stay the same. preliminary intentions are to use ddi or a modification also for this kind of data (data documentation initiative, 2009) qualitative longitudinal data there are many examples of qualitative longitudinal projects but they are decentralized. several of these have youth as a central group for study: e.g. their life stories, participation among youth with disabilities, youth and violence/abuse. the life cycle is another central theme. however, as far as we have learned there are no large archives or collections of data made available – nor is there national cooperation at this point5. development and planning within social sciences and humanities, the share of qualitative research is considerable. for instance, 35% of the 216 projects in the norwegian programme on welfare research (1999-2008) used qualitative data, either solely or in addition to quantitative data. of these, the in-depth (and semi-structured) interview was the most frequent method, used in almost 90% of the qualitatively oriented projects. one quarter of the programme’s qualitative projects were anthropological studies. qualitative research takes place in different milieus in norway – one example of a fairly large research institute is nova norwegian social research6. one possible direction for nsd in the future is to establish general archiving agreements with major research institutes like nova. such a general agreement will include both quantitative and qualitative data – and the “data catch” will be expanded accordingly. in a longer perspective it is important for the data archive to encourage more research institutes to archive data at nsd. archiving research data is a topic that is discussed at national level with the research council of norway taking the lead (see their work to implement the oecd guidelines). whether a project’s research data can be stored and/or reused is to a high degree determined by what the respondents have consented to prior to the data collection. therefore it is important that researchers early on in the research process are aware of the possibilities for archiving and reuse. in norway, the privacy ombudsman for research is in direct contact with a large share of research projects in their initial phases to give advice both with regard to development of questionnaires and interview guides and formulation of information letters to respondents. the main task of the privacy ombudsman for research (2011) is to disseminate knowledge of the legal and ethical guidelines regulating research privacy protection of the respondents is a challenge, especially for qualitative data archiving. projects that have not obtained consent from the respondents for long term archiving with identification have to anonymise data before archiving. anonymisation will, in many cases, reduce the utility of data when it comes to reuse. this is especially an issue for longitudinal data. up until now nsd has mainly focused on archiving quantitative data. one reason for this is technology – quantitative data have to a higher degree than qualitative ones been suited for digital archiving through well-developed technology and metadata standards. as the technology has evolved, it is now possible to digitally store large quantities of text, photos and video. there is also a lot of development when it comes to tools and software to analyse this kind of data. nsd envisages a common solution for archiving quantitative and qualitative data. existing routines for collection, documentation and presentation of data will be expanded to include qualitative data. nsd uses nesstar7 as a presentation utility; this gives nsd a wide range of options regarding how to present data. data can be published as part of a portal solution for all available data at the archive, as a research programme specific server, or as a separate qualitative collection. the welfare web portal (2011), which is under construction, will be an example of a research specific programme server, with many different types of data, texts, qualitative interviews and quantitative data. a public server will mainly be restricted to metadata, as this kind of individual micro-level data requires an elaborate security system. the challenge for nsd, as we see it, is first and foremost to offer qualitative data prepared for analysis, as we do not yet have much experience in this field. what is meant by prepared qualitative data? what is the routine of other archives within this field? the specialised competence that is required has to be defined and developed. nsd is a member of, among others, cessda (council of european social science data archives) and some members of staff have joined the iassist (international association for social science information service & technology)8. through these memberships nsd seeks to stay informed about developments pertaining to technology as well as maintenance of qualitative data (software, organization and routines). these organizations have also secured agreements among the archives regarding access to and exchange of data that benefits the users. in iassist quarterly 2010 / 2011 49 iassist quarterly addition, more specialised initiatives focusing on qualitative data, both acquisition and curation and dissemination are of interest and are welcome. references data documentation initiative. (2009). [online]. available at: http:// www.ddialliance.org/ norwegian programme on welfare research. (1999-2008). [online]. available at: http://www.forskningsradet.no/servlet/satellite?c=pag e&cid=1222932203400&p=1222932203400&pagename=vfo%2fho vedsidemal norwegian social science data service. (2011). [online] available at http://www.nsd.uib.no/nsd/english/index.html privacy ombudsman for research. (2011). [online]. available at: http:// www.nsd.uib.no/personvern/om/english.html oecd (2007) ‘oecd principles and guidelines for access to research data from public fund-ing’. http://www.oecd.org/dataoecd/9/61/38500813.pdf the research council of norway. (n.d). [online]. available at: http:// www.forskningsradet.no/en/home+page/1177315753906 the welfare web portal. (2011). [online]. available at: http://www.nsd. uib.no/velferd/. (note: only in norwegian) notes 1. gry-hege henriksen*, norwegian social science data services (gry. henriksen@nsd.uib.no) maria bakke orvik*, norwegian social science data services (formerly at nsd) trond pedersen*, norwegian social science data services (trond.pedersen@nsd.uib.no) * the content of this article was originally presented at the bremen workshop held at the university of bremen, germany, on 24th of april 2009. the authors gry-hege henriksen (adviser) and trond pedersen (adviser) work at norwegian social science data services (nsd), harald hårfagres gate 29, n-5007 bergen, norway. e-mail addresses: gry.henriksen@nsd.uib.no, trond.pedersen@nsd.uib.no. maria bakke orvik is no longer employed at nsd. 2. for further information see, http://www.timescapes.leeds.ac.uk/ about/bremen-workshop/ 3. for further information see, http://www.digitalarkivet.no/cgi-win/ webfront.exe?slag=vis&tekst=meldingar&spraak=e; http://www. khm.uio.no/ ; http://www.rhd.uit.no/indexeng.html ; http://www. dokpro.uio.no/engelsk 4. polsys, a collection with search possibilities for norwegian political party documents includes election programmes, press releases, statutes and the political parties’ platforms. note: only in norwegian. http://www.nsd.uib.no/polsys/index.cfm?urlname=parti&lan=&inst itusjonsnr=3&arkivnr=9&menuitem=n1_3&childitem=&state=coll apse 5. bibsys (http://www.bibsys.no/norsk/english.php) contains information on projects and pub-lications, but not data. contact information to researchers/research milieus (nova or others) is easy to find. 6. for further information see, http://www.nova.no/?language=1 7. the nesstar server is built as an extension to a normal web server. as well as providing all the usual facilities for publishing web content, this server provides the ability to publish statistical information that can be searched, browsed, analysed and downloaded by users. this is done either by using a standard web browser or using nesstar webview. see more at http://www.nesstar.com/ 8. for further information see,http://www.cessda.org/; http://www. iassistdata.org/ vol21.3 4 iassist quarterly changing the way the united states measures income and poverty: a progress report1 by daniel h. weinberg & charles t. nelson * this paper reports the general results of research undertaken by census bureau staff. the views expressed are attributable to the authors and do not necessarily reflect the views of the census bureau or the u.s. government. the authors would like to acknowledge and thank the following people for their comments and suggestions—nancy gordon, edward welniak and the coauthors of the other papers cited in this report; they bear no responsibility for any errors that remain. i. background — the official definition the united states census bureau has been compiling income estimates annually since 1947. these estimates are from the current population survey (cps), a nationwide random sample of households, whose primary purpose is to collect labor force information monthly. in march of each year (april prior to 1956), data are collected on the household’s income for the previous calendar year. the official definition of income is not specified in law or regulation. in effect, what is included in income depends on the questions asked. as survey researchers know, the more questions one asks about income by source, the better able respondents are to identify all income. initially, there were only two questions asked of each adult:2 (1) “how much did ... earn in wages and salaries in 1947?” and (2) “how much income from all sources did ... receive in 1947?”. in 1949, self-employment income was asked separately and in 1950 farm and nonfarm self-employment income was asked separately. in 1962, the census bureau began systematically assigning values to missing income items (based on reported characteristics using the “hot deck” method). in march 1967, the number of income questions was again expanded, from four to eight categories. these additional items dealt with social security, interest, dividends, and rent. in 1968, interest, dividends, rents, and royalties were combined into one question and separate questions were added on public assistance and on unemployment and workers’ compensation. in 1975, the number of income questions increased from eight to eleven through addition of a question on the supplemental security income program, a question on aid to families with dependent children and general assistance, and private and government pension income. a major change took place in 1980 — the questionnaire was expanded to identify over 50 sources of income and recording of up to 27 different income amounts, including receipt of numerous noncash benefits, such as food stamps (coupons used as cash for qualified food purchases), and housing assistance. except for minor wording changes, those questions are still in use today. the survey was converted to a computerassisted interviewing mode in 1994. the data on income thus cover money income received (exclusive of certain money receipts such as capital gains) before payments for items such as personal income taxes, social security payroll taxes, and union dues. money income does not reflect the fact that some families receive part of their income in the form of noncash benefits, such as food stamps, health benefits, rent-free or subsidized housing, and goods produced and consumed on the farm. in addition, money income does not reflect the fact that noncash benefits are also received by some as fringe benefits, e.g. the use of company cars, and full or partial payments by business for retirement programs, medical insurance, and educational expenses. moreover, for many different reasons, there is a tendency in household surveys for respondents to underreport their income. from an analysis of independently derived income estimates, it has been determined that income earned from wages or salaries is much better reported than other sources of income and is nearly equal to independent estimates of aggregate earnings (coder and scoon-rogers, 1996). among the least well-reported sources are interest and dividends. the detailed components of money income are presented in the appendix. ii. alternative measures of income because money income is but one measure of economic well-being, the census bureau also reports on 14 other definitions of income (the series begins in 1979). while not exhaustive, they do illustrate different perspectives on what could be included. definition 1. money income excluding capital gains before taxes. this is the official definition described above. fall 1997 5 definition 2. definition 1 less government cash transfers. government cash transfers include nonmeans-tested transfers such as social security payments, unemployment compensation, and government educational assistance (e.g., pell grants), as well as means-tested transfers such as aid to families with dependent children (afdc), temporary assistance to needy families, and supplemental security income (ssi). definition 3. definition 2 plus capital gains. realized capital gains and losses are simulated as part of the census bureau’s federal individual income tax estimation procedure. while the census bureau has access to some income information on individual tax returns that can be matched (with substantial time lag) to survey data, actual capital gains or losses or tax liability are not known. definition 4. definition 3 plus imputed health insurance supplements to wage or salary income. employer-paid health insurance coverage is treated as part of total worker compensation; no other benefits paid for or provided by employers are estimated. definition 5. definition 4 less payroll taxes. payroll taxes are payments for social security old age, survivors, and disability insurance, and for hospital insurance (medicare). definition 6. definition 5 less federal income taxes. the effect of the earned income tax credit, targeted to low-income workers, is shown separately in definition 7. definition 7. definition 6 plus the earned income tax credit. definition 8. definition 7 less state income taxes. definition 9. definition 8 plus nonmeans-tested government cash transfers. nonmeans-tested government cash transfers include social security payments, unemployment compensation, workers’ compensation, nonmeans-tested veterans’ payments, u.s. railroad retirement, black lung payments, and pell grants and other government educational assistance. (pell grants are income-tested but are included here because they are very different from the assistance programs included in the means-tested category.) definition 10. definition 9 plus the value of medicare. medicare is counted at its fungible value.3 definition 11. definition 10 plus the value of regular-price school lunches. definition 12. definition 11 plus means-tested government cash transfers. means-tested government cash transfers include afdc, ssi, other public assistance programs, and means-tested veterans’ payments. definition 13. definition 12 plus the value of medicaid. medicaid is counted at its fungible value. definition 14. definition 13 plus the value of other means-tested government noncash transfers. including food stamps, rent subsidies, and free and reduced-price school lunches. definition 15. definition 14 plus net imputed return on equity in one’s own home. this definition includes the estimated annual benefit of converting one’s home equity into an annuity, net of property taxes. table 12 is a reproduction of a table from u.s. bureau of the census (1996a) illustrating the different distributions of income that these definitions imply.4 table 5 (u.s. bureau of the census, 1996b) illustrates this effect on poverty estimates. these alternative definitions illustrate the dilemma faced by official statisticians when presenting income statistics. different definitions serve different purposes. money income has its uses — it represents command over the resources available to purchase the necessities of life in the open market, including meeting the obligations of citizenship (taxes). definition 4 probably comes closest to measuring what resources would be available in the absence of government, except that some benefits paid for or provided by employers are not included and others are mandated by the government, some benefits are not provided by employers because they are provided by the government, and work effort is presumably reduced by the existence of a tax on earnings. definition 8 is closest to after-tax income. disposable income tries to take account of the effect of taxes and transfers on the household’s command of resources — definition 14 probably comes closest to that approach. finally, in definition 15 there is an attempt to include the income equivalent value of owning one’s own home in that such an asset reduces the need for additional expenditures on shelter. iii. considerations in measuring poverty formal measurement of poverty in the united states is less than three decades old. not since the adoption of official poverty thresholds by the federal government in the late 1960’s has there been such a great interest as now in examining and possibly respecifying the thresholds and the income compared with them. the official poverty thresholds in use today by the u.s. bureau of the census to measure poverty have their basis in work by orshansky 6 iassist quarterly (1963, 1965). orshansky started with a set of minimally adequate food budgets calculated for families of various sizes and composition by the u.s. department of agriculture for 1961. based on evidence from the 1955 household food consumption survey, she determined that expenditures on food represented about one-third of aftertax income for the typical family. this relationship yielded a “multiplier” of three, that is, the minimally adequate food budgets were multiplied by a factor of three to obtain 124 poverty thresholds that differed by family size, number of children, age and sex of head, and farm or nonfarm residence (ad hoc adjustments were made for families of size one and two). in 1969, the u.s. bureau of the budget (now the u.s. office of management and budget — omb) adopted the orshansky measure using pre-tax income as the standard government poverty measure, mandating that thresholds be adjusted for inflation using the consumer price index (cpi) published by the u.s. bureau of labor statistics. with only minor modifications since then (mostly reducing the number of categories, now 48), the orshansky thresholds still form the basis for the official poverty statistics.5 when considering the adequacy of the official poverty thresholds, it is critical to realize that one cannot separate the issue of income measurement from poverty definition. when one defines the level of resources needed to be nonpoor, one must also determine which resources are to be counted. therefore, the discussion below covers both income measurement and poverty definition issues; income measurement is discussed first.6 whatever poverty thresholds are chosen should be the result of a carefully specified process that cannot be changed arbitrarily from year-to-year, and should be capable of being updated at reasonable intervals as the economic circumstances of the society and the behavior of its demographic and economic components change. a. defining income for measuring poverty the key measurement issues are three — valuing and counting noncash income, subtracting taxes, and reducing survey underreporting and nonsampling errors. also of interest is whether to continue to publish official estimates based on the cps or switch to a newer survey designed to collect better income information, the survey of income and program participation (sipp). a.1. noncash income the issue of valuing noncash income spans the income distribution. a more comprehensive income measure, such as definition 14 above, would place a value not only on noncash government transfers, such as food stamps, which typically go to low-income families, but also on elements of nonwage compensation (from employer-paid health insurance to company cars) that typically go to earners at all income levels or only at high levels. the noncash income of u.s. families has grown substantially in the past 25 years. in the 1990’s, over half of government transfer spending for the poor is in the form of noncash benefits (u.s. bureau of the census, 1996a), whereas the only noncash benefit program that predated the 1960’s “war on poverty” was subsidized (public) housing. this growth of benefits to the poor has been paralleled by a growth of nonwage compensation to wage earners, induced in part by tax laws exempting such compensation from income and payroll taxes, and by growth in health benefits for the elderly. by 1996, employer costs for nonwage compensation had grown to over one-quarter (28.4 percent) of total compensation costs, up from 19.4 percent in 1966.7 further, nearly two-thirds of households own homes, which provide them with additional noncash income in the form of housing services. of key concern to understanding the well-being of u.s. households is the valuation of medical benefits, both the government health programs—medicare (medical aid to the elderly and severely disabled) and medicaid (medical aid to a portion of the poor)—and employer-paid health insurance. the valuation of medical benefits is particularly difficult since coverage of high medical expenses for people who are sick does nothing to improve their poverty status (although the benefits clearly make them better off). even if one imputes the value of an equivalent insurance policy to program participants, these benefits (high in market value due to large medical costs for the fraction who do get sick), and cannot be used by the recipients to meet other needs of daily living. accordingly, the census bureau developed a not-altogether-satisfactory method, termed fungible value (described in footnote 2), to avoid giving too high a value of these benefits to those toward the low end of the income scale. note that this is not a problem for countries with universal health care systems. a.2. disposable income even though orshansky’s original calculations were based on post-tax income, poverty has always been calculated for the official statistics using pre-tax income because of the limited information collected on the cps. after-tax income is a better measure of the ability to meet the daily necessities of life than is money income. also important, in calculating disposable income though, is to address the advisability of deducting work expenses for wage earners such as child care, uniforms, and transportation costs. a.3. other issues as noted earlier, research matching household survey responses to federal income tax returns and comparing them with national income accounts has revealed substantial areas where the level and receipt of certain fall 1997 7 income sources is underreported. attempts to reduce underreporting were made by revising the language used in the cps questionnaire (and using a shorter reference period) when the sipp was launched. this was only partially successful, and response errors remain. while current procedures of the census bureau reweight the data for full interview nonresponse and impute appropriate income responses for individual unanswered questions (item nonresponse), these corrections are insufficient to fully resolve the problem. procedures to enhance the data through microsimulation or other means are being investigated, along with continued improvement in imputation for nonresponse. in most societies, “underground,” “nonmarket,” or “black market” income from legal or illegal activities is typically poorly reported by household respondents to government surveys (or not even collected) and consequently is substantially omitted from official income statistics. this income ranges from barter transactions to home production (e.g., home gardens) to illegal income. researchers are a long way from measuring this activity accurately, however, so including this income in official statistics would be quite difficult. it has been suggested that consumption is a better measure of well-being than income (see cutler and katz, 1991, and slesnick, 1993). if a family can maintain its consumption through judicious use of assets when income falls, is it truly poor? unfortunately, it is difficult to collect accurate annual data on consumption or even expenditures. further, consumption reflects choices on how to allocate resources, rather than need. nevertheless, fuller investigation of a consumption-based measure would be useful. a final issue of income measurement is the choice of surveys to use. as mentioned briefly above, the sipp questionnaire design, as crafted to reduce income underreporting, does succeed for almost all income sources.8 yet, when compared with the cps, it has historically had several drawbacks—a smaller sample size (one-third as large) and necessarily slower data release because of its much greater complexity. these defects are compensated for by the sipp having greater income detail, both in number of sources and in time segments (by having monthly as opposed to the cps’s annual statistics,) and lower underreporting. the new version of the sipp, as implemented in 1996, increased the sample size substantially (to 36,700 households) and oversampled lowincome households. national estimates from the sipp will then be comparable to or better than (in terms of sampling error) those from the cps (reduced to 48,000 households but inefficient for national estimates because it uses a statebased design). one drawback for obtaining a consistent time series of annual national income or poverty estimates from the sipp, though, will be sample attrition and time-insample bias as current plans call for only one sipp panel to be in the field during any one four-year period. the cps sample is constantly refreshed by new households. while the timeliness issue may never be resolved fully in sipp’s favor, the sipp can provide a preliminary estimate on much the same schedule as the cps. still, it is desirable to view the surveys complementarily. if modeling using administrative records can correct underreporting errors in both surveys, they would then give the same aggregate statistics. the cps could be used for a quick snapshot, consistent with data collected since 1947 (the sipp began in 1983), while the sipp would be used for more detailed estimates, for subannual and multiyear estimates, and for understanding other dimensions of poverty (assets, disability, gross flows, and other dynamic aspects).9 b. setting thresholds to define poverty with an absolute measure of poverty, there are key decisions to be made about determining the appropriate level for poverty thresholds. the key research issues addressed here are minimal consumption levels for specific commodities, ways of correcting for differences in family size and composition, and ways of correcting for cost-ofliving differences across time and among areas. b.1. minimal consumption standards minimal consumption standards for all necessary commodities could in theory be established, perhaps by an expert panel, but doing so would raise difficult ethical issues about which commodities to include (e.g., is a telephone a necessity?). one alternative is to define minimal consumption standards for a limited number of necessities (e.g. food, clothing, shelter) and obtain a poverty threshold by using a multiplier to account for necessities not measured.10 b.2. equivalence scales the relationship embodied in the current u.s. poverty thresholds among families of different sizes (termed the equivalence scale) is supposed to represent the different relative costs of supporting those families at a minimally adequate levels. in fact, the relationship is based solely on the relative food costs as they existed in 1961 and include some unfortunate anomalies (see ruggles, 1990, pp. 6468). while it is possible to develop minimal budgets for every type and size of family separately and thus eliminate the need for equivalence scales entirely, in practice it is difficult to do so. no one scale now exists that is generally accepted. issues in developing equivalence scales include which distinctions in family circumstances (e.g. owner/ renter) should lead to different thresholds, how resources are shared within the family or household, and whether a more useful basis for determining poverty is the household (those living in one housing unit) rather than the family 8 iassist quarterly (those in one household related by blood or marriage). see betson (1996) for a further discussion of these issues. b.3. cost-of-living differences in as large and diverse a country as the u.s., there are significant differences in the cost of living among localities. unfortunately, there are no currently available data upon which to estimate interarea price differences reliably. (see kokoski et al., 1992, and moulton, 1992, for some work in this area.) a related price issue is how to adjust for inflation. the u.s. poverty thresholds now use the cpi to adjust thresholds over time. if the measurement of minimal consumption is used as the basis for new thresholds, presumably this should be the basis every year, with components, prices, and multipliers reestimated as often. clearly this is not practical. a reasonable compromise might be to respecify and reestimate the minimal consumption bundle at prespecified intervals as market baskets become outdated, say every ten years, and use the cpi for interim adjustments. the market basket used for the cpi itself is typically reviewed and respecified once every ten years or so.11 c. the committee on national statistics report the national academy of sciences’ committee on national statistics (cnstat) released a report in may 1995 entitled measuring poverty: a new approach (citro and michael, 1995). in that report, the committee recommended that the federal government redefine the way it measures poverty. omb has requested that experts from the census bureau and other agencies examine technical methods for doing so. the key changes they recommend are threefold: change the income measure, change the poverty thresholds, and change the survey used. to change the income measure from the current money income definition, they propose to add noncash benefits, subtract taxes, subtract work expenses, subtract child care expenses, subtract child support paid, and subtract medical out-of-pocket expenses (moop). the poverty thresholds are to be based on food, clothing, shelter, and “a little bit more” (75-83% of median expenditures on these items multiplied by 1.15-1.25), a new equivalence scale, an allowance for geographic variation, and are to be updated annually based on growth in median expenditures. finally, the panel recommended that the government use the sipp instead of the march cps to collect the basic income and poverty-related data. among the technical issues to be resolved before implementing such a new measure are the following: 1. reestimating the valuation methodologies for government noncash transfer programs including school lunches, food stamps, and housing benefits; developing new estimation methodologies for additional programs and possibly developing a new methodology for valuing medicare and medicaid (depending on whether the subtraction of moop is adopted or not); 2. completing development of a tax simulation model for sipp; 3. developing a methodology for estimating moop (e.g. a statistical match of the national medical expenditures survey to sipp) or reestimation of employer contributions to health insurance using more recent data; 4. estimating and imputing work and child care expenses; 5. redesigning the sipp sampling scheme to maximize reliability of a time series of cross-section estimates while maintaining some longitudinal estimation capabilities, taking account of the need for state-level estimates, and minimizing the attrition bias; 6. reviewing the consumer expenditure survey to improve its effectiveness for its new dual role (defining the market basket for the consumer price index and the poverty thresholds) and possibly preparing for consumption-based rather than income-based poverty estimates in the future; 7. creating a time series of poverty estimates from the sipp and developing methods to impute additional variables to the cps to develop comparable time-series data for that survey; 8. doing substantial further work on income underreporting and imputation models; 9. adding child support and alimony paid questions to cps; and 10. developing and adding “medical care risk” and possibly medical expenditures questions to sipp to supplement the poverty measure if medical care costs and benefits are excluded from the measure. even if these technical issues can be resolved expeditiously, there are still policy issues that must be debated and resolved before a new measure is adopted. these include: 1. including or excluding medical costs and benefits. on the one hand, the cnstat recommended excluding moop, employer contributions to health insurance, and benefits from medical transfer programs from income. on the other hand, adopting as official the current (experimental) practice of including them would require fall 1997 9 improving the current method for valuing medical transfer program benefits, measuring medical needs more accurately, and updating the methodology for imputing employer contributions to health insurance. 2. basing thresholds on a pre-specified fraction of median expenditures. how might the public and congress react to a new poverty threshold that showed millions more poor persons than the current measure? are we confident enough about the quality of (i.e. lack of biases in) the consumer expenditure survey data to use it as the arbiter of the poverty level? it may be that the likely acceptance of any new definition would be enhanced if the new index were “chained” to the old by matching the overall poverty rate obtained (but allowing the distribution to vary). 3. developing geographical cost-of-living variations. it is clear that the cost of living differs substantially from place to place, and different choices of methodology to reflect this fact would have different implications. if geographic variation is to be incorporated, some method for periodically updating the thresholds for relative price changes among areas would also need to be established. 4. annual inflation updating. the panel proposed using the rate of growth in expenditures to index the thresholds. this is an attempt to introduce some deliberate “relativity” into the measure and would have quite different ramifications from using the consumer price index. 5. choosing the equivalence scale. choice of the scale would inevitably alter the distribution of the poor. 6. underreporting. if the technical issues about how to do so are resolved, should the income statistics from the survey be adjusted for underreporting based on administrative data and modeling? 7. review and revision. should any new definition include a regular cycle of review and revision based on pre-specified criteria (cnstat recommended once a decade)? open debate of these issues seems the most likely way to resolve them, potentially leading to a new way of measuring poverty that omb would approve and that other policy makers would accept as an improved methodology for measuring poverty in the united states. d. census bureau poverty redefinition research in order to provide a basis on which some of these issues can be resolved, the census bureau and other u.s. government agencies have begun research studies. d.1. census bureau-bureau of labor statistics study the cnstat report on redefining poverty contained sweeping recommendations for changing the way poverty is defined in the u.s. recent joint research by the bureau of the census and the bureau of labor statistics (bls) (garner et al., 1997) examined two of these issues — changing the income definition and modifying the poverty thresholds. in formulating poverty thresholds, bls researchers started by implementing the basic recommendations from the cnstat report. some of the cnstat panel recommendations regarding thresholds were given as ranges. thus, some simplifying assumptions were made. first, the panel recommended a range of thresholds, with a lower bound based on 78 percent of median expenditures for food, clothing, and shelter and a multiplier of 1.15 to account for other needs. the upper bound was based on 83 percent of the median and a multiplier of 1.25. in the garner et al. paper the midpoint of this range was used. the other simplifying assumption was for the equivalence scale (the relationship between thresholds for different family sizes). the panel recommended a range of economy scale factors of 0.65 to 0.75 and again they choose the midpoint — 0.70. thresholds were computed for the years 1990 through 1995. on the resource side, the panel’s recommendations were followed to the extent possible. the only recommendation not followed (because of a lack of data) was their recommendation to subtract child support paid from income when computing a poverty resource measure. though the panel recommended changing the official source of poverty statistics in the u.s. from the cps to the sipp, the initial work was based on the cps. at this time, the cps is the only survey with a working tax simulation model and in-kind benefit valuation procedures, both necessary ingredients for producing a resource measure based on the panel’s recommendations. the report found that the threshold computation methods as recommended by the panel result in relatively stable thresholds over time (at least over the 1990-1995 period measured in this study), and the resulting poverty rates based on applying the panel’s basic resource definition to these thresholds also showed relatively stable results. in fact, though the panel’s recommendations result in significantly higher poverty rates than the u.s. official estimates, the trends based on the official estimates and the panel’s recommended method show very similar trends over the 1990-1995 period (see figure 1). differences across subgroups were also found to be stable over time. however, the key change under the proposed definition of poverty is in the composition of the poverty population. consistent with the panel’s findings, poverty rates under the recommended poverty measure are significantly higher among groups with relatively low official poverty rates (for 10 iassist quarterly example, whites or those living in married-couple families). groups with relatively high poverty rates, on the other hand, did not tend to have very different poverty rates under the revised measure. thus, an effect of moving to the recommended poverty measure would be to narrow the gaps that now exist in the u. s. between highand lowpoverty groups (married-couple and single-parent families, whites and blacks, etc.). put another way, under the revised measure, the poverty population looks more like the total population in terms of demographic and socioeconomic characteristics. (see table 1 and figures 24.) other, slightly different poverty thresholds were also examined in the census-bls study. one modification, which was suggested by the panel, was to define shelter costs by their rental equivalent value. this technique resulted in higher poverty thresholds (and higher poverty rates), and appeared to have some effect on the composition of the poverty population (further narrowing the gaps, for example, between high-and low-poverty groups). another set of thresholds was based on alternative multipliers that were computed more precisely than those used in the panel’s report. this modification resulted in little change in the composition of the poverty population. d.2. other census bureau poverty research the panel recommended changing the source of official u.s. poverty estimates from the march cps to the sipp. as noted earlier, the sipp is a longitudinal survey with: 1) a more detailed set of questions than the cps, 2) a shorter reference period (4 months versus 12 months for the cps), and 3) increased flexibility sufficient to add the questions required to measure poverty based on the broadened resource definition recommended by the panel. questions have already been added to sipp to collect some of this additional information, and a sample design change, in order to make sipp a better cross-sectional survey (a requirement for measuring annual poverty changes) has been proposed, though not yet adopted. the census bureau has also examined the panel’s recommendations on work-related and childcare expenses (the panel recommended subtracting these costs from income when computing the poverty resource measure and has suggested alternative methods for imputing such costs). this research showed that using a definition of resources that excludes child care and other work-related expenses has a significant effect on poverty rates. in both cps and sipp-based analyses, the effect of using a resource definition that excluded these expenses was to raise children’s poverty rates by about 3 percentage points. (see short et al., 1996.) another area of research at the census bureau is on the housing subsidy valuation method. the value of public or subsidized housing is included in the recommended poverty measure, and the current census bureau method for imputing such subsidies (on the cps) is badly outdated. current methods are being reviewed, and ways to implement this imputation on sipp are being explored. a paper is planned for presentation in august (eller and naifeh, forthcoming). the one major element of the panel’s recommended resource measure not included in the census bureau-bls study was the subtraction of child support paid, since this information was not available in the cps. data from sipp indicate that the inclusion of such payments would increase the poverty rate by 0.3 to 0.5. questions were added to the april 1996 cps supplement on child support to examine the feasibility of capturing this information on a regular basis on the march cps. data on child support paid are regularly collected on sipp. as already noted, the treatment of medical benefits and expenditures in defining poverty is a difficult one. staff are currently examining the treatment of medical out-of-pocket expenditures in the definition of poverty (see doyle, forthcoming(a)). to come up with a definition of income that excludes these expenditures, our current thinking is that statistically matching sipp to another federal government survey that includes detailed information about these expenditures (the medical expenditure panel survey) holds the most promise. in addition, staff are working on a proposed medical care risk index to complement the new poverty measure (to address another recommendation of the panel). (see doyle, forthcoming(b).) since the panel recommended an after-tax income definition for its poverty measure, one problem with transferring the official poverty measure from the cps to sipp is the lack of a working tax simulation model based on the sipp (since the early 1980’s, the cps has employed a model to estimates taxes). the census bureau, along with several other federal agencies, supported the development of a sipp-based tax model, and we are now in the process of exploring how to best incorporate this model into the census bureau’s processing system. equivalence scales are an important issue in the formulation of poverty thresholds. betson (1996) provides compelling evidence that the choice of equivalence scales has a significant effect on the composition of the poverty population. he also pointed to the need for continued research in this area. in another paper, betson (1995) examined the issue of home ownership and whether the flow of housing services from owner-occupied homes should be taken into account when defining poverty status. he found that counting the value of housing services would change the distribution of the poor, primarily by counting fewer of the elderly as poor. fall 1997 11 e. concluding remarks we believe that prospects for developing a consensus around a new measure of poverty in the united states are the highest since the current measure was adopted in the l960’s. converting the measure to the sipp is not costless, though, and budgetary pressures may cause a delay even if a broad methodological consensus is reached. furthermore, delicate negotiations over broad policy issues must ensue before any change is made. readers are welcome to follow further developments as they happen. visit the special poverty measurement web site at http://www.census.gov/hhes/www/povmeas.html. appendix: definition of money income the current official u.s. definition of income is based on questions which are asked of each person in the cps sample household 15 years old and over.12 these questions cover the amount of money income received in the preceding calendar year from each of the following sources. table 5. poverty rates: official and experimental by race, hispanic origin, family type and age: 1992 official experimental percent difference all persons 14.8 19.9 34.5 white 11.9 17.1 43.7 black 33.4 37.1 11.1 hispanic origin (of any race) 29.6 41.5 40.2 married couple 7.7 13.7 77.9 female household 39.0 42.8 9.7 under 18 years old 22.4 27.1 21.0 18 64 years old 11.9 16.3 37.0 65 years old and over 12.9 22.5 74.4 earnings from longest job (or self-employment) and other employment earnings can be classified into three types: (1) money wage or salary income is the total received for work performed as an employee during the income year. this category includes wages, salary, armed forces pay, commissions, tips, piece-rate payments, and cash bonuses earned, before deductions were made for items such as taxes, bonds, pensions, and union dues; (2) net income from nonfarm self-employment is the net money income (gross receipts minus expenses) from one’s own business, professional enterprise, or partnership. gross receipts include the value of all goods sold and services rendered. expenses include items such as costs of goods purchased, rent, heat, light, power, depreciation charges, wages and salaries paid, business taxes (not personal income taxes);13 and (3) net income from farm self-employment is the net money income (gross receipts minus operating expenses) from the operation of a farm by a person on their own account, as an owner, renter, or sharecropper. gross receipts include the value of all products sold, payments from government farm programs, money received from the rental of farm equipment to others, rent received from farm property if payment is made based on a percent of crops produced and incidental receipts from the sale of items 12 iassist quarterly such as wood, sand, and gravel. operating expenses include items such as the cost of feed, fertilizer, seed, and other farming supplies; cash wages paid to farmhands; depreciation charges; cash rent; interest on farm mortgages; farm building repairs; and farm taxes (not state and federal personal income taxes). the value of fuel, food, or other farm products used for family living is not included as part of net income.14 unemployment compensation includes payments received from government unemployment agencies or private companies during periods of unemployment and any strike benefits received from union funds. workers’ compensation includes payments received periodically from public or private insurance companies for injuries received at work. social security includes social security (old age) pensions and survivors’ benefits and permanent disability insurance payments made by the social security administration prior to deductions for medical insurance. medicare reimbursements for health services are not included. supplemental security income includes payments made by fall 1997 13 federal, state, and local welfare agencies to low income persons who are 65 years old or over, blind, or disabled. public assistance or welfare payments include public assistance payments made to low-income persons, such as aid to families with dependent children, temporary assistance for needy families, and general assistance. veterans’ payments include payments made periodically by the department of veterans affairs to disabled members of the armed forces or to survivors of deceased veterans for education and on-the-job training, and means-tested assistance to veterans. survivor benefits include payments from survivors’ or widows’ pensions, estates, trusts, annuities, or any other types of survivor benefits. payments can be reported from ten different sources: private companies or unions; federal government (civil service); military; state or local governments; railroad retirement; workers’ compensation; “black lung” (miners’) payments; estates and trusts; annuities or paid-up insurance policies; and other survivor payments. disability benefits include payments received as a result of a health problem or disability other than those from social security. payments can be reported from ten sources: workers’ compensation; companies or unions; federal government (civil service); military; state or local governments; railroad retirement; accident or disability insurance; black lung payments; state temporary sickness; or other disability payments. pension or retirement income includes payments reported from eight sources: companies or unions; federal government (civil service); military; state or local governments; railroad retirement; annuities or paid-up insurance policies; withdrawals from special (tax-favored) retirement accounts such as individual retirement account (ira’s); or other retirement income. interest income includes payments received (or credited to bank accounts), from bonds, treasury notes, ira’s, certificates of deposit, interest-bearing savings and checking accounts, and all other investments that pay interest. dividends include income received from stock holdings and mutual fund shares. capital gains from the sale of stock holdings are not included as income. rents, royalties, and estates and trusts include the net income from the rental of a house, store, or other property, receipts from boarders or lodgers, net royalty income, and periodic payments from estate or trust funds. educational assistance includes pell grants; other government educational assistance; any scholarships or grants; or financial assistance from employers, friends, or relatives not residing in the student’s household. child support includes all periodic payments made by parents for the support of children, even if these payments are made through a state or local government office.15 alimony includes all periodic payments to ex-spouses. one-time property settlements are not included. 14 iassist quarterly financial assistance from outside of the household includes periodic payments from nonhousehold members. gifts or sporadic assistance is not included. other income includes all other regularly received payments that are not included elsewhere on the questionnaire. some examples are state programs such as foster child payments, military family allotments, and income received from foreign government pensions. receipts not counted as income include capital gains received (or losses incurred) from the sale of property, including stocks, bonds, a house, or a car (unless the person was engaged in the business of selling such property, in which case the net proceeds would be counted as income from self-employment); withdrawals of bank deposits; money borrowed; tax refunds; gifts; and lump-sum inheritances or insurance payments. references betson, david m. 1996. “‘is everything relative?’ the role of equivalence scales in poverty measurement.” university of notre dame. march. betson, david. 1995. “effect of home ownership on poverty measurement.” unpublished paper. november. citro, constance f. and graham kalton (eds.). 1993. the future of the survey of income and program participation. washington, dc: national academy press. citro, constance f. and robert t. michael (eds.). 1995. measuring poverty: a new approach. washington, dc: national academy press. coder, john and lydia scoon-rogers. 1996. ”evaluating the quality of income data collected in the annual supplement to the march current population survey and the survey of income and program participation.” working paper. u.s. bureau of the census. july. cutler, david m. and lawrence f. katz. 1991. “macroeconomic performance and the disadvantaged.” brookings papers on economic activity no. 2, pp. 1-74. doyle, pat. forthcoming (a). “how can we deduct something we do not collect? the case of out-ofpocket medical expenditures.” u.s. bureau of the census. doyle, pat. forthcoming (b). “who’s at risk? designing a medical care risk index.” u.s. bureau of the census. eller, t.j. and mary naifeh. forthcoming. “housing subsidies: effect of estimates on poverty.” u.s. bureau of the census. fisher, gordon m. 1992. “the development and history of the poverty thresholds.” social security bulletin vol. 55 no. 4 (winter), pp. 3-14. garner, thesia i., geoffrey paulin, stephanie shipp, kathleen short, charles nelson. 1997. “experimental poverty measurement for the 1990’s.” prepared for the allied social science meetings, session sponsored by the society of government economists, january 1997; and the sge session: measures of well-being from the consumer expenditures survey, january 1997. kokoski, mary, patrick cardiff, and brent moulton. 1992. “interarea price indices for consumer goods and services: an hedonic approach using cpi data.” u.s. bureau of labor statistics, january. moulton, brent r. 1992. “interarea indexes of the cost of shelter using hedonic quality adjustment techniques.” u.s. bureau of labor statistics, october. orshansky, mollie. 1963. “children of the poor.” social security bulletin v. 26 (july), pp. 3-13. orshansky, mollie. 1965. “counting the poor.” social security bulletin v. 28 (january), pp. 3-29. ruggles, patricia. 1990. drawing the line. washington, d.c.: urban institute press. short, kathleen, martina shea, and t.j. eller. 1996. “work-related expenditures in a new measure of poverty.” prepared for the 1996 meetings of the american statistical association. slesnick, daniel t. 1992. “gaining ground: poverty in the postwar united states.” journal of political economy vol. 101 no. 1 (february), pp. 1-38. u.s. bureau of the census, money income in the united states: 1995, current population reports p60-193, washington, dc: u.s. government printing office, september 1996[a]. u.s. bureau of the census, poverty in the united states: 1995, current population reports p60-194, washington, dc: us government printing office, september 1996[b]. watts, harold w. 1993. “a review of alternative budget-based expenditure norms.” prepared for the panel on poverty measurement and family assistance of the committee on national statistics, revised (may). welniak, edward j., jr. 1990. “effects of the march current population survey’s new processing system on estimates of income and poverty.” prepared for american statistical association annual meeting, august. fall 1997 15 table 12. income distribution measures by definition of income: 1995 (numbers in thousands. households as of march of the following year. for meaning of symbols, see text) characteristic money incomem before taxes after taxes money incomemdefinition 1 less taxes plus capital gains (losses) excluding capital gains (current official measure) without eitc with eitc definition 1 less government transfers definition 2 plus capital gains (losses) definition 3 plus health insurance supplements to wage or salary income definition 4 less social security payroll taxes definition 5 less federal income taxes definition 6 plus earned income tax credit 1 1a 1b 2 3 4 5 6 7 all households total ................................ 99 627.......... 99 627 99 627 99 627 99 627 99 627 99 627 99 627 99 627 recipiency status with income as defined .................. 99 032.......... 99 032 99 032 93 004 93 009 93 009 93 009 93 014 93 014 with addition or deduction................ (x).......... (x) (x) 42 392 15 918 54 312 75 096 73 158 14 860 mean addition or deduction dollars................. (x) (x) (x) 8 879 8 512 3 897 3 193 7 719 1 250 standard error dollars.......................... (x) (x) (x) 51 308 14 13 99 12 mean total income dollars........................ (x) (x) (x) 23 715 85 353 64 598 51 682 47 964 20 696 standard error dollars.......................... (x) (x) (x) 269 1 309 419 343 252 232 income levels percent ............................ 100.0.......... 100.0 100.0 100.0 100.0 100.0 100.0 100.0 100.0 under $5,000........................... 3.7.......... 3.9 3.7 16.5 16.5 16.4 16.7 16.8 16.4 $5,000 to $9,999 ........................ 8.6.......... 9.3 8.8 6.4 6.3 6.1 6.5 6.9 6.6 $10,000 to $14,999...................... 8.7.......... 10.1 9.6 6.5 6.4 6.0 6.5 7.0 6.7 $15,000 to $19,999...................... 8.3.......... 9.9 10.4 6.5 6.5 6.1 6.5 7.2 7.6 $20,000 to $24,999...................... 7.6.......... 9.4 9.7 6.5 6.6 6.1 6.4 7.3 7.5 $25,000 to $29,999...................... 7.4.......... 9.0 9.1 6.2 6.2 5.9 6.3 7.1 7.3 $30,000 to $34,999...................... 6.8.......... 7.9 8.0 6.0 6.1 5.8 5.9 6.5 6.5 $35,000 to $39,999...................... 6.3.......... 7.2 7.3 5.6 5.5 5.3 5.5 5.9 5.9 $40,000 to $44,999...................... 5.6.......... 6.2 6.2 5.1 5.1 4.9 5.0 5.3 5.3 $45,000 to $49,999...................... 5.0.......... 4.8 4.8 4.5 4.4 4.4 4.5 4.8 4.9 $50,000 to $59,999...................... 8.3.......... 8.0 8.1 7.8 7.7 8.0 7.7 7.8 7.8 $60,000 to $74,999...................... 8.8.......... 6.7 6.7 8.2 8.2 8.6 8.1 7.8 7.8 $75,000 to $99,999...................... 7.7.......... 4.3 4.3 7.3 7.4 8.3 7.4 5.3 5.3 $100,000 and over ...................... 7.1.......... 3.4 3.4 6.8 7.1 8.1 6.9 4.4 4.4 summary measures median dollars.................................... 34 076 29 093 29 219 30 931 31 082 32 819 30 793 28 393 28 535 standard error dollars............................ 197 135 134 166 171 215 193 173 170 mean dollars...................................... 44 938 36 729 36 915 41 160 42 520 44 644 42 238 36 569 36 756 standard error dollars............................ 246 181 181 251 279 286 277 207 206 gini ratio ............................... .444.......... .418 .414 .503 .511 .509 .514 .490 .486 standard error ........................ .0039.......... .0039 .0039 .0038 .0040 .0039 .0040 .0039 .0039 quintile measures lowest quintile: upper limit dollars............................... 14 420 13 408 13 921 7 654 7 679 7 851 7 410 7 351 7 756 percent of households ................. 20.0.......... 20.0 20.0 20.0 20.0 20.0 20.0 20.0 20.0 with type of addition or deduction ..... (x).......... (x) (x) 17 144 697 412 4 814 423 2 794 mean amount dollars......................... (x) (x) (x) 9 666 –110 1 386 314 443 546 standard error dollars...................... (x) (x) (x) 73 112 70 5 142 16 second quintile: upper limit dollars............................... 26 966 23 610 23 831 22 950 23 086 24 400 22 891 21 450 21 834 percent of households ................. 20.0.......... 20.0 20.0 20.0 20.0 20.0 20.0 20.0 20.0 with type of addition or deduction ..... (x).......... (x) (x) 10 031 1 653 6 299 15 137 13 583 6 683 mean amount dollars......................... (x) (x) (x) 9 354 795 2 054 1 197 1 017 1 650 standard error dollars...................... (x) (x) (x) 102 89 22 8 9 18 third quintile: upper limit dollars............................... 42 012 35 288 35 397 39 659 39 940 42 235 39 619 36 021 36 127 percent of households ................. 20.0.......... 20.0 20.0 20.0 20.0 20.0 20.0 20.0 20.0 with type of addition or deduction ..... (x).......... (x) (x) 6 685 2 489 13 412 17 691 19 353 3 799 mean amount dollars......................... (x) (x) (x) 7 806 1 258 2 807 2 292 2 528 1 057 standard error dollars...................... (x) (x) (x) 130 89 17 10 14 23 fourth quintile: upper limit dollars............................... 65 258 52 481 52 520 63 123 63 970 67 767 63 639 56 502 56 551 percent of households ................. 20.0.......... 20.0 20.0 20.0 20.0 20.0 20.0 20.0 20.0 with type of addition or deduction ..... (x).......... (x) (x) 4 774 3 575 16 710 18 532 19 909 1 095 mean amount dollars......................... (x) (x) (x) 7 189 2 310 3 877 3 566 5 187 1 302 standard error dollars...................... (x) (x) (x) 165 98 20 14 23 46 fifth quintile: percent of households ................. 20.0.......... 20.0 20.0 20.0 20.0 20.0 20.0 20.0 20.0 with type of deduction ............... (x).......... (x) (x) 3 758 7 505 17 479 18 923 19 890 488 mean amount dollars......................... (x) (x) (x) 8 073 16 371 5 477 5 997 20 037 1 201 standard error dollars...................... (x) (x) (x) 196 623 28 30 327 66 48 valuation of noncash benefits 16 iassist quarterly table 12. income distribution measures by definition of income: 1995 mcon. (numbers in thousands. households as of march of the following year. for meaning of symbols, see text) characteristic after taxesmcon. definition 13 plus other means~tested governmentm definition 7 less state income taxes definition 8 plus nonmeans~ tested government cash transfers definition 9 plus medicare definition 10 plus regular~price school lunches definition 11 plus means~tested government cash transfers definition 12 plus medicaid noncash transfers noncash transfers less medical programs definition 14 plus net imputed return on equity in own home 8 9 10 11 12 13 14 14a 15 all households total ................................ 99 627.......... 99 627 99 627 99 627 99 627 99 627 99 627 99 627 99 627 recipiency status with income as defined .................. 93 022.......... 97 510 97 629 97 646 99 041 99 041 99 224 99 224 99 419 with addition or deduction................ 64 827.......... 37 786 23 259 12 663 8 306 10 207 15 750 30 101 65 139 mean addition or deduction dollars................. 2 296 8 930 5 004 88 4 690 2 796 1 876 4 815 3 370 standard error dollars.......................... 26 54 26 1 68 38 22 26 30 mean total income dollars........................ 44 052 31 024 34 655 57 171 19 596 31 942 21 925 15 056 50 829 standard error dollars.......................... 245 232 298 655 403 454 206 342 256 income levels percent ............................ 100.0.......... 100.0 100.0 100.0 100.0 100.0 100.0 100.0 100.0 under $5,000........................... 16.4.......... 6.0 5.8 5.8 3.6 3.6 2.7 2.7 2.2 $5,000 to $9,999 ........................ 6.7.......... 7.6 6.4 6.4 7.4 7.1 6.4 7.8 5.6 $10,000 to $14,999...................... 6.9.......... 8.6 7.1 7.1 7.4 7.2 7.6 9.9 7.3 $15,000 to $19,999...................... 7.8.......... 9.1 8.9 8.9 9.1 8.9 9.3 9.6 8.6 $20,000 to $24,999...................... 7.9.......... 8.8 9.0 9.0 9.1 9.2 9.4 9.2 9.1 $25,000 to $29,999...................... 7.5.......... 8.6 8.7 8.7 8.7 8.9 9.1 8.8 8.7 $30,000 to $34,999...................... 6.7.......... 7.6 8.1 8.1 8.2 8.2 8.3 7.7 8.1 $35,000 to $39,999...................... 6.1.......... 6.9 7.4 7.4 7.6 7.7 7.8 7.1 7.8 $40,000 to $44,999...................... 5.4.......... 6.1 6.5 6.5 6.5 6.6 6.6 6.2 6.9 $45,000 to $49,999...................... 4.8.......... 5.3 5.7 5.7 5.7 5.8 5.8 5.3 6.1 $50,000 to $59,999...................... 7.9.......... 8.5 8.8 8.8 8.9 9.0 9.0 8.6 9.4 $60,000 to $74,999...................... 7.3.......... 7.9 8.2 8.2 8.2 8.3 8.3 7.9 9.1 $75,000 to $99,999...................... 4.7.......... 5.1 5.3 5.3 5.4 5.4 5.4 5.2 6.4 $100,000 and over ...................... 3.8.......... 4.0 4.1 4.1 4.1 4.1 4.1 4.0 4.8 summary measures median dollars.................................... 27 772 30 892 32 549 32 563 32 761 33 149 33 306 31 280 35 259 standard error dollars............................ 163 156 146 146 144 142 143 153 154 mean dollars...................................... 35 262 38 649 39 817 39 828 40 219 40 506 40 802 39 347 43 006 standard error dollars............................ 192 188 188 188 187 187 186 187 190 gini ratio ............................... .481.......... .424 .412 .412 .404 .400 .394 .409 .388 standard error ........................ .0038.......... .0039 .0038 .0038 .0038 .0038 .0038 .0039 .0038 quintile measures lowest quintile: upper limit dollars............................... 7 700 13 785 15 382 15 384 15 855 16 219 16 758 14 816 17 933 percent of households ................. 20.0.......... 20.0 20.0 20.0 20.0 20.0 20.0 20.0 20.0 with type of addition or deduction ..... 2 323.......... 10 441 4 785 354 4 823 2 776 7 014 8 110 7 327 mean amount dollars......................... 98 6 802 2 015 81 4 161 1 222 2 244 2 705 1 885 standard error dollars...................... 8 47 28 3 62 28 34 26 66 second quintile: upper limit dollars............................... 21 354 24 957 26 564 26 570 26 837 27 195 27 429 25 434 29 127 percent of households ................. 20.0.......... 20.0 20.0 20.0 20.0 20.0 20.0 20.0 20.0 with type of addition or deduction ..... 13 247.......... 9 233 6 282 1 215 1 536 2 875 4 439 8 870 10 540 mean amount dollars......................... 389 9 588 4 630 81 4 994 2 703 1 721 5 099 2 448 standard error dollars...................... 5 86 26 2 169 46 41 41 48 third quintile: upper limit dollars............................... 35 008 37 682 38 937 38 950 39 096 39 410 39 537 37 948 41 760 percent of households ................. 20.0.......... 20.0 20.0 20.0 20.0 20.0 20.0 20.0 20.0 with type of addition or deduction ..... 15 857.......... 7 446 5 238 2 528 1 000 2 009 2 624 6 081 13 685 mean amount dollars......................... 1 014 9 584 6 259 85 5 268 3 620 1 447 5 965 2 725 standard error dollars...................... 8 116 49 1 259 85 49 59 45 fourth quintile: upper limit dollars............................... 54 274 56 093 56 986 57 002 57 110 57 330 57 363 56 239 60 300 percent of households ................. 20.0.......... 20.0 20.0 20.0 20.0 20.0 20.0 20.0 20.0 with type of addition or deduction ..... 16 649.......... 5 828 3 845 3 953 548 1 397 1 301 3 900 15 898 mean amount dollars......................... 1 960 9 517 6 486 90 6 382 4 015 1 421 5 912 3 150 standard error dollars...................... 12 160 61 1 382 131 75 80 50 fifth quintile: percent of households ................. 20.0.......... 20.0 20.0 20.0 20.0 20.0 20.0 20.0 20.0 with type of deduction ............... 16 751.......... 4 839 3 108 4 612 399 1 151 371 3 141 17 689 mean amount dollars......................... 5 657 10 553 6 412 89 6 142 3 908 1 393 5 872 5 231 standard error dollars...................... 87 255 72 1 484 172 133 87 82 valuation of noncash benefits 49 fall 1997 17 table 12. income distribution measures by definition of income: 1995 mcon. (numbers in thousands. households as of march of the following year. for meaning of symbols, see text) characteristic money incomem before taxes after taxes money incomemdefinition 1 less taxes plus capital gains (losses) excluding capital gains (current official measure) without eitc with eitc definition 1 less government transfers definition 2 plus capital gains (losses) definition 3 plus health insurance supplements to wage or salary income definition 4 less social security payroll taxes definition 5 less federal income taxes definition 6 plus earned income tax credit 1 1a 1b 2 3 4 5 6 7 households with female householder, no husband present, with related children under 18 total ................................ 8 751.......... 8 751 8 751 8 751 8 751 8 751 8 751 8 751 8 751 recipiency status with income as defined .................. 8 670.......... 8 670 8 670 7 653 7 653 7 653 7 653 7 659 7 659 with addition or deduction................ (x).......... (x) (x) 4 467 666 3 630 6 728 4 455 4 648 mean addition or deduction dollars................. (x) (x) (x) 6 188 5 590 3 327 1 683 3 241 1 622 standard error dollars.......................... (x) (x) (x) 117 1 254 42 27 258 20 mean total income dollars........................ (x) (x) (x) 13 513 59 529 39 149 26 601 34 049 20 493 standard error dollars.......................... (x) (x) (x) 442 5 353 840 646 671 483 income levels percent ............................ 100.0.......... 100.0 100.0 100.0 100.0 100.0 100.0 100.0 100.0 under $5,000........................... 10.8.......... 11.5 10.3 27.2 27.2 27.0 27.9 27.9 25.9 $5,000 to $9,999 ........................ 17.0.......... 17.8 15.5 10.4 10.3 9.7 10.2 10.3 8.8 $10,000 to $14,999...................... 14.8.......... 16.5 14.7 11.1 11.1 10.2 10.7 11.0 10.5 $15,000 to $19,999...................... 11.5.......... 12.8 15.0 9.2 9.3 8.8 9.0 9.7 10.9 $20,000 to $24,999...................... 9.0.......... 9.8 11.3 8.6 8.3 7.9 8.1 9.1 9.8 $25,000 to $29,999...................... 7.6.......... 8.3 8.8 7.1 7.0 7.1 7.4 7.7 8.9 $30,000 to $34,999...................... 6.6.......... 7.0 7.2 6.2 6.4 6.5 6.4 6.7 6.7 $35,000 to $39,999...................... 6.1.......... 4.2 4.7 5.3 5.0 5.1 4.7 4.6 4.9 $40,000 to $44,999...................... 3.7.......... 3.7 3.9 3.5 3.7 3.8 4.0 3.4 3.6 $45,000 to $49,999...................... 3.0.......... 2.3 2.5 2.5 2.5 3.4 2.5 2.2 2.4 $50,000 to $59,999...................... 4.2.......... 2.9 3.0 3.7 3.5 3.9 3.8 3.6 3.7 $60,000 to $74,999...................... 2.9.......... 1.5 1.5 2.7 2.7 3.3 2.5 1.9 1.9 $75,000 to $99,999...................... 1.5.......... 1.0 1.0 1.3 1.6 1.9 1.7 1.1 1.1 $100,000 and over ...................... 1.3.......... .7 .7 1.2 1.2 1.3 1.1 .8 .8 summary measures median dollars.................................... 17 936 16 600 18 039 15 584 15 651 16 783 15 693 15 400 17 191 standard error dollars............................ 409 303 287 395 393 456 431 395 367 mean dollars...................................... 24 508 21 504 22 365 21 349 21 774 23 154 21 860 20 210 21 072 standard error dollars............................ 466 363 362 473 534 549 534 417 416 gini ratio ............................... .454.......... .433 .415 .525 .532 .532 .534 .516 .496 standard error ........................ .0134.......... .0134 .0132 .0129 .0135 .0133 .0135 .0127 .0127 quintile measures lowest quintile: upper limit dollars............................... 14 420 13 408 13 921 7 654 7 679 7 851 7 410 7 351 7 756 percent of households ................. 41.3.......... 40.6 37.3 33.0 33.1 32.7 33.1 32.7 30.7 with type of addition or deduction ..... (x).......... (x) (x) 2 413 21 38 1 271 25 744 mean amount dollars......................... (x) (x) (x) 6 513 (b) (b) 272 (b) 972 standard error dollars...................... (x) (x) (x) 141 (b) (b) 9 (b) 31 second quintile: upper limit dollars............................... 26 966 23 610 23 831 22 950 23 086 24 400 22 891 21 450 21 834 percent of households ................. 24.8.......... 24.9 27.1 30.1 30.5 30.0 29.5 29.0 29.4 with type of addition or deduction ..... (x).......... (x) (x) 1 080 86 1 045 2 389 1 202 2 224 mean amount dollars......................... (x) (x) (x) 5 406 1 069 2 487 1 035 682 1 999 standard error dollars...................... (x) (x) (x) 269 386 54 13 24 26 third quintile: upper limit dollars............................... 42 012 35 288 35 397 39 659 39 940 42 235 39 619 36 021 36 127 percent of households ................. 18.9.......... 18.5 18.9 21.7 21.2 21.7 21.5 21.5 22.7 with type of addition or deduction ..... (x).......... (x) (x) 607 209 1 374 1 757 1 785 1 209 mean amount dollars......................... (x) (x) (x) 5 789 1 667 3 080 2 041 1 706 1 381 standard error dollars...................... (x) (x) (x) 297 303 47 24 33 43 fourth quintile: upper limit dollars............................... 65 258 52 481 52 520 63 123 63 970 67 767 63 639 56 502 56 551 percent of households ................. 10.6.......... 10.9 11.5 10.6 10.8 11.2 11.5 12.0 12.3 with type of addition or deduction ..... (x).......... (x) (x) 246 185 844 936 1 034 378 mean amount dollars......................... (x) (x) (x) 6 778 2 699 4 015 3 247 3 812 1 501 standard error dollars...................... (x) (x) (x) 577 342 84 55 81 75 fifth quintile: percent of households ................. 4.5.......... 5.1 5.2 4.5 4.4 4.4 4.4 4.7 4.9 with type of deduction ............... (x).......... (x) (x) 121 165 328 376 408 93 mean amount dollars......................... (x) (x) (x) 7 501 16 944 5 390 5 005 16 231 1 453 standard error dollars...................... (x) (x) (x) 1 006 4 794 213 172 2 600 165 50 valuation of noncash benefits 18 iassist quarterly table 12. income distribution measures by definition of income: 1995 mcon. (numbers in thousands. households as of march of the following year. for meaning of symbols, see text) characteristic after taxesmcon. definition 13 plus other means~tested governmentm definition 7 less state income taxes definition 8 plus nonmeans~ tested government cash transfers definition 9 plus medicare definition 10 plus regular~price school lunches definition 11 plus means~tested government cash transfers definition 12 plus medicaid noncash transfers noncash transfers less medical programs definition 14 plus net imputed return on equity in own home 8 9 10 11 12 13 14 14a 15 households with female householder, no husband present, with related children under 18 total ................................ 8 751.......... 8 751 8 751 8 751 8 751 8 751 8 751 8 751 8 751 recipiency status with income as defined .................. 7 660.......... 7 944 7 954 7 966 8 675 8 675 8 739 8 739 8 742 with addition or deduction................ 4 188.......... 2 361 531 1 828 2 964 2 476 5 294 2 709 3 139 mean addition or deduction dollars................. 1 015 5 536 3 933 80 4 915 2 797 2 659 3 327 2 317 standard error dollars.......................... 81 171 160 1 102 81 48 86 125 mean total income dollars........................ 30 900 24 111 34 364 37 830 14 355 25 399 19 893 12 885 36 398 standard error dollars.......................... 625 639 1 724 1 037 398 799 386 980 683 income levels percent ............................ 100.0.......... 100.0 100.0 100.0 100.0 100.0 100.0 100.0 100.0 under $5,000........................... 25.9.......... 21.4 21.3 21.3 10.2 10.0 3.5 3.5 3.2 $5,000 to $9,999 ........................ 9.0.......... 10.0 9.8 9.9 15.0 13.6 10.9 11.2 10.5 $10,000 to $14,999...................... 10.5.......... 10.7 10.6 10.6 12.7 12.0 14.8 18.0 14.6 $15,000 to $19,999...................... 11.3.......... 11.8 11.8 11.8 13.6 13.0 15.8 16.0 15.4 $20,000 to $24,999...................... 10.5.......... 10.4 10.6 10.5 11.1 12.2 13.4 12.7 13.2 $25,000 to $29,999...................... 8.6.......... 8.7 8.4 8.4 8.8 9.1 10.6 10.0 10.5 $30,000 to $34,999...................... 6.8.......... 7.6 7.7 7.7 7.8 8.1 7.8 7.9 7.3 $35,000 to $39,999...................... 4.8.......... 4.9 5.0 5.1 5.7 5.9 6.5 5.6 7.3 $40,000 to $44,999...................... 3.4.......... 3.8 3.6 3.6 3.7 4.1 4.4 4.0 4.4 $45,000 to $49,999...................... 2.6.......... 3.0 3.3 3.3 3.4 3.4 3.4 3.0 3.5 $50,000 to $59,999...................... 3.4.......... 3.7 3.8 3.8 3.9 4.1 4.3 4.1 4.8 $60,000 to $74,999...................... 1.7.......... 2.0 2.1 2.1 2.1 2.3 2.3 2.1 2.6 $75,000 to $99,999...................... 1.0.......... 1.1 1.2 1.2 1.3 1.4 1.5 1.2 1.8 $100,000 and over ...................... .7.......... .7 .7 .7 .7 .8 .8 .7 .8 summary measures median dollars.................................... 17 086 18 306 18 527 18 539 19 400 20 569 21 786 20 529 22 360 standard error dollars............................ 357 342 337 336 312 329 285 299 300 mean dollars...................................... 20 587 22 081 22 319 22 336 24 000 24 792 26 400 25 370 27 231 standard error dollars............................ 386 390 392 392 381 383 372 366 382 gini ratio ............................... .491.......... .470 .470 .470 .421 .413 .367 .370 .368 standard error ........................ .0125.......... .0125 .0125 .0125 .0130 .0128 .0129 .0131 .0129 quintile measures lowest quintile: upper limit dollars............................... 7 700 13 785 15 382 15 384 15 855 16 219 16 758 14 816 17 933 percent of households ................. 30.6.......... 39.3 42.7 42.7 40.5 38.8 35.6 31.8 37.4 with type of addition or deduction ..... 154.......... 874 143 177 2 104 876 2 686 648 559 mean amount dollars......................... 68 3 889 1 467 81 4 570 1 433 3 154 1 614 1 029 standard error dollars...................... 8 158 152 4 95 49 67 63 172 second quintile: upper limit dollars............................... 21 354 24 957 26 564 26 570 26 837 27 195 27 429 25 434 29 127 percent of households ................. 29.1.......... 25.0 24.2 24.1 25.8 26.5 28.4 30.4 28.1 with type of addition or deduction ..... 1 292.......... 557 99 433 485 830 1 572 1 135 764 mean amount dollars......................... 249 5 117 3 932 76 5 629 2 917 2 281 2 962 1 518 standard error dollars...................... 9 293 256 3 322 92 83 90 156 third quintile: upper limit dollars............................... 35 008 37 682 38 937 38 950 39 096 39 410 39 537 37 948 41 760 percent of households ................. 22.9.......... 19.2 17.6 17.6 17.8 18.2 18.8 20.8 18.1 with type of addition or deduction ..... 1 502.......... 483 113 580 229 428 659 529 843 mean amount dollars......................... 660 5 964 4 624 77 4 814 3 847 1 951 4 965 2 185 standard error dollars...................... 17 387 190 2 439 199 132 236 189 fourth quintile: upper limit dollars............................... 54 274 56 093 56 986 57 002 57 110 57 330 57 363 56 239 60 300 percent of households ................. 12.5.......... 11.4 10.6 10.6 10.8 11.4 11.9 11.8 11.3 with type of addition or deduction ..... 894.......... 291 94 401 89 220 308 245 650 mean amount dollars......................... 1 366 7 988 5 398 85 8 353 4 615 1 866 4 843 2 733 standard error dollars...................... 44 575 376 3 1 102 388 216 415 255 fifth quintile: percent of households ................. 4.9.......... 5.0 4.9 4.9 5.1 5.2 5.3 5.2 5.1 with type of deduction ............... 346.......... 157 82 236 57 123 69 152 323 mean amount dollars......................... 4 925 10 337 5 587 86 (b) 4 800 (b) 5 224 5 939 standard error dollars...................... 900 1 207 444 4 (b) 832 (b) 521 787 valuation of noncash benefits 51 fall 1997 19 table 12. income distribution measures by definition of income: 1995 mcon. (numbers in thousands. households as of march of the following year. for meaning of symbols, see text) characteristic money incomem before taxes after taxes money incomemdefinition 1 less taxes plus capital gains (losses) excluding capital gains (current official measure) without eitc with eitc definition 1 less government transfers definition 2 plus capital gains (losses) definition 3 plus health insurance supplements to wage or salary income definition 4 less social security payroll taxes definition 5 less federal income taxes definition 6 plus earned income tax credit 1 1a 1b 2 3 4 5 6 7 households with members 65 years old and over total ................................ 23 732.......... 23 732 23 732 23 732 23 732 23 732 23 732 23 732 23 732 recipiency status with income as defined .................. 23 592.......... 23 592 23 592 20 124 20 124 20 124 20 124 20 124 20 124 with addition or deduction................ (x).......... (x) (x) 22 374 3 572 4 251 7 673 10 500 1 013 mean addition or deduction dollars................. (x) (x) (x) 11 414 6 168 3 100 2 225 6 116 762 standard error dollars.......................... (x) (x) (x) 64 483 50 41 224 41 mean total income dollars........................ (x) (x) (x) 18 631 54 360 57 737 42 439 36 990 20 747 standard error dollars.......................... (x) (x) (x) 347 2 051 1 446 1 036 584 884 income levels percent ............................ 100.0.......... 100.0 100.0 100.0 100.0 100.0 100.0 100.0 100.0 under $5,000........................... 3.3.......... 3.3 3.3 40.7 40.7 40.5 40.9 40.9 40.8 $5,000 to $9,999 ........................ 16.7.......... 16.7 16.7 12.6 12.6 12.5 12.6 13.2 13.3 $10,000 to $14,999...................... 16.0.......... 16.6 16.6 8.9 8.7 8.6 8.7 9.1 9.1 $15,000 to $19,999...................... 13.0.......... 13.4 13.4 7.0 6.9 6.7 6.9 7.5 7.5 $20,000 to $24,999...................... 9.5.......... 10.2 10.2 5.6 5.6 5.6 5.6 6.1 6.2 $25,000 to $29,999...................... 8.1.......... 8.8 8.8 4.1 4.0 4.0 4.1 4.0 4.1 $30,000 to $34,999...................... 6.1.......... 6.7 6.7 3.3 3.3 3.3 3.4 3.8 3.8 $35,000 to $39,999...................... 4.9.......... 5.4 5.4 3.0 3.0 3.0 2.9 2.7 2.7 $40,000 to $44,999...................... 3.9.......... 3.8 3.8 2.2 2.1 2.2 2.1 2.0 2.0 $45,000 to $49,999...................... 2.8.......... 2.6 2.7 1.8 1.8 1.9 1.9 1.7 1.7 $50,000 to $59,999...................... 4.4.......... 4.3 4.3 2.8 2.7 2.8 2.6 2.6 2.6 $60,000 to $74,999...................... 4.1.......... 3.3 3.3 2.6 2.8 2.8 2.7 2.4 2.4 $75,000 to $99,999...................... 3.3.......... 2.7 2.7 2.5 2.6 2.8 2.6 2.0 2.0 $100,000 and over ...................... 3.9.......... 2.1 2.1 2.9 3.1 3.3 3.0 2.0 2.0 summary measures median dollars.................................... 20 503 19 959 19 994 8 427 8 447 8 552 8 348 8 231 8 277 standard error dollars............................ 236 204 206 226 231 231 226 207 207 mean dollars...................................... 30 934 27 745 27 777 20 173 21 101 21 656 20 937 18 231 18 264 standard error dollars............................ 369 287 287 365 408 416 404 307 307 gini ratio ............................... .470.......... .436 .435 .655 .664 .665 .664 .639 .639 standard error ........................ .0087.......... .0084 .0084 .0088 .0091 .0090 .0091 .0088 .0087 quintile measures lowest quintile: upper limit dollars............................... 14 420 13 408 13 921 7 654 7 679 7 851 7 410 7 351 7 756 percent of households ................. 34.0.......... 31.5 32.9 48.5 48.5 48.7 47.9 47.8 48.8 with type of addition or deduction ..... (x).......... (x) (x) 11 213 514 100 1 011 59 293 mean amount dollars......................... (x) (x) (x) 10 747 90 1 355 290 (b) 316 standard error dollars...................... (x) (x) (x) 84 128 152 11 (b) 36 second quintile: upper limit dollars............................... 26 966 23 610 23 831 22 950 23 086 24 400 22 891 21 450 21 834 percent of households ................. 28.0.......... 26.4 25.3 24.3 24.2 24.8 24.5 24.8 24.2 with type of addition or deduction ..... (x).......... (x) (x) 5 478 838 842 2 255 4 032 365 mean amount dollars......................... (x) (x) (x) 12 286 1 196 1 861 971 854 870 standard error dollars...................... (x) (x) (x) 126 104 59 21 16 73 third quintile: upper limit dollars............................... 42 012 35 288 35 397 39 659 39 940 42 235 39 619 36 021 36 127 percent of households ................. 17.4.......... 18.2 18.0 12.2 12.2 11.8 12.5 12.7 12.2 with type of addition or deduction ..... (x).......... (x) (x) 2 649 757 1 186 1 757 2 942 191 mean amount dollars......................... (x) (x) (x) 11 657 2 032 2 411 1 916 2 893 880 standard error dollars...................... (x) (x) (x) 206 173 55 39 40 98 fourth quintile: upper limit dollars............................... 65 258 52 481 52 520 63 123 63 970 67 767 63 639 56 502 56 551 percent of households ................. 11.0.......... 12.7 12.6 7.7 7.5 7.4 7.5 7.5 7.5 with type of addition or deduction ..... (x).......... (x) (x) 1 622 549 1 010 1 270 1 757 120 mean amount dollars......................... (x) (x) (x) 11 547 3 524 3 171 2 965 6 226 1 190 standard error dollars...................... (x) (x) (x) 258 266 80 64 101 140 fifth quintile: percent of households ................. 9.6.......... 11.2 11.2 7.2 7.6 7.3 7.6 7.2 7.2 with type of deduction ............... (x).......... (x) (x) 1 412 913 1 113 1 381 1 710 45 mean amount dollars......................... (x) (x) (x) 12 728 19 171 4 863 5 403 24 155 (b) standard error dollars...................... (x) (x) (x) 319 1 713 125 134 1 153 (b) 52 valuation of noncash benefits 20 iassist quarterly table 12. income distribution measures by definition of income: 1995 mcon. (numbers in thousands. households as of march of the following year. for meaning of symbols, see text) characteristic after taxesmcon. definition 13 plus other means~tested governmentm definition 7 less state income taxes definition 8 plus nonmeans~ tested government cash transfers definition 9 plus medicare definition 10 plus regular~price school lunches definition 11 plus means~tested government cash transfers definition 12 plus medicaid noncash transfers noncash transfers less medical programs definition 14 plus net imputed return on equity in own home 8 9 10 11 12 13 14 14a 15 households with members 65 years old and over total ................................ 23 732.......... 23 732 23 732 23 732 23 732 23 732 23 732 23 732 23 732 recipiency status with income as defined .................. 20 126.......... 23 426 23 501 23 504 23 596 23 596 23 626 23 626 23 701 with addition or deduction................ 10 540.......... 22 030 20 707 459 1 838 2 354 2 617 20 748 18 737 mean addition or deduction dollars................. 1 559 11 268 5 063 79 3 894 2 140 1 491 5 296 4 636 standard error dollars.......................... 50 64 27 2 133 62 34 29 58 mean total income dollars........................ 30 749 27 582 34 904 67 312 22 190 32 571 18 819 14 762 40 487 standard error dollars.......................... 507 286 313 3 685 824 913 553 409 376 income levels percent ............................ 100.0.......... 100.0 100.0 100.0 100.0 100.0 100.0 100.0 100.0 under $5,000........................... 40.9.......... 5.3 5.0 5.0 3.3 3.3 2.9 2.9 1.8 $5,000 to $9,999 ........................ 13.4.......... 15.5 11.1 11.1 11.9 11.7 10.9 15.4 8.1 $10,000 to $14,999...................... 9.3.......... 16.3 10.0 10.0 10.2 10.2 10.7 17.9 10.1 $15,000 to $19,999...................... 7.6.......... 13.2 12.2 12.2 12.4 12.3 12.9 13.5 11.1 $20,000 to $24,999...................... 6.2.......... 9.7 10.6 10.6 10.6 10.5 10.5 9.9 10.5 $25,000 to $29,999...................... 4.4.......... 8.6 9.0 9.0 9.1 9.2 9.2 8.7 9.0 $30,000 to $34,999...................... 3.6.......... 6.5 8.6 8.6 8.8 8.6 8.6 6.6 8.2 $35,000 to $39,999...................... 2.7.......... 5.2 7.2 7.2 7.4 7.5 7.5 5.3 7.9 $40,000 to $44,999...................... 1.8.......... 4.1 5.3 5.3 5.3 5.5 5.5 4.1 6.8 $45,000 to $49,999...................... 1.6.......... 2.6 4.2 4.2 4.2 4.2 4.3 2.7 5.6 $50,000 to $59,999...................... 2.7.......... 4.3 5.5 5.5 5.6 5.6 5.6 4.4 6.7 $60,000 to $74,999...................... 2.2.......... 3.5 4.8 4.8 4.8 4.9 4.9 3.5 6.0 $75,000 to $99,999...................... 1.8.......... 3.0 3.7 3.7 3.7 3.7 3.7 3.0 4.6 $100,000 and over ...................... 1.8.......... 2.3 2.7 2.7 2.7 2.8 2.8 2.3 3.5 summary measures median dollars.................................... 8 214 19 897 25 556 25 556 25 828 26 035 26 106 20 205 29 611 standard error dollars............................ 203 205 262 262 258 251 251 232 276 mean dollars...................................... 17 571 28 031 32 448 32 450 32 752 32 964 33 128 28 498 36 789 standard error dollars............................ 288 295 305 305 304 305 304 294 319 gini ratio ............................... .633.......... .448 .420 .420 .414 .413 .409 .435 .393 standard error ........................ .0086.......... .0084 .0080 .0080 .0080 .0080 .0080 .0084 .0078 quintile measures lowest quintile: upper limit dollars............................... 7 700 13 785 15 382 15 384 15 855 16 219 16 758 14 816 17 933 percent of households ................. 48.8.......... 33.2 27.0 27.0 27.5 27.9 28.8 35.4 26.4 with type of addition or deduction ..... 1 324.......... 7 197 4 012 34 1 005 788 1 645 6 019 3 633 mean amount dollars......................... 80 7 656 2 005 (b) 3 012 689 1 622 2 927 2 392 standard error dollars...................... 3 49 28 (b) 115 29 40 30 86 second quintile: upper limit dollars............................... 21 354 24 957 26 564 26 570 26 837 27 195 27 429 25 434 29 127 percent of households ................. 24.2.......... 26.6 24.7 24.7 24.3 24.2 23.8 24.8 23.0 with type of addition or deduction ..... 3 964.......... 6 051 5 718 28 314 516 496 5 736 4 252 mean amount dollars......................... 336 11 896 4 637 (b) 4 254 1 955 1 223 5 927 3 483 standard error dollars...................... 7 84 27 (b) 311 57 71 45 68 third quintile: upper limit dollars............................... 35 008 37 682 38 937 38 950 39 096 39 410 39 537 37 948 41 760 percent of households ................. 12.4.......... 18.1 20.8 20.8 20.8 20.5 20.1 17.8 20.1 with type of addition or deduction ..... 2 351.......... 4 058 4 772 53 245 400 236 4 087 4 206 mean amount dollars......................... 1 105 12 964 6 325 (b) 5 794 2 844 1 375 6 517 4 389 standard error dollars...................... 23 131 51 (b) 590 114 145 62 80 fourth quintile: upper limit dollars............................... 54 274 56 093 56 986 57 002 57 110 57 330 57 363 56 239 60 300 percent of households ................. 7.4.......... 12.0 15.0 15.0 14.9 14.9 14.8 11.8 16.6 with type of addition or deduction ..... 1 446.......... 2 618 3 432 118 134 325 157 2 662 3 561 mean amount dollars......................... 2 068 14 116 6 517 76 5 011 3 660 1 261 6 464 5 447 standard error dollars...................... 46 230 66 4 498 212 151 83 120 fifth quintile: percent of households ................. 7.2.......... 10.2 12.5 12.5 12.5 12.6 12.6 10.2 13.9 with type of deduction ............... 1 456.......... 2 106 2 772 226 139 325 82 2 243 3 085 mean amount dollars......................... 6 459 14 993 6 397 83 5 021 3 565 1 276 6 426 8 267 standard error dollars...................... 287 379 75 4 603 238 242 92 239 valuation of noncash benefits 53 fall 1997 21 weinberg, daniel h. 1996. “changing the way the u.s. measures income and poverty.” prepared for the canberra group on income statistics, december. * paper presented at iassist/ifdo ‘97, odense, denmark, may 6-9,1997. daniel h. weinberg chief, housing and household economic statistics division and charles t. nelson assistant division chief for economic characteristics housing and household economic statistics division u.s. bureau of the census washington, dc 20233-8500 usa may 1997 phone: (301) 763-8550 facsimile: (301) 763-8412, e-mail: daniel.h.weinberg@ccmail.census.gov, charles.t.nelson@ccmail.census.gov 1 this paper is largely based on weinberg (1996) and garner et al. (1997). 2 the history of income questions asked on the current population survey is from welniak (1990). 3 the fungible approach for valuing medical coverage assigns income to the extent that having the insurance would free up resources that would have been spent on medical care. the estimated fungible value depends on family income, the cost of food and housing needs, and the market value of the medical benefits. if family income is not sufficient to cover the family’s basic food and housing requirements, the fungible value methodology treats medicare and medicaid as having no income value. if family income exceeds the cost of food and housing requirements, the fungible value of medicare and medicaid is equal to the amount which exceeds the value assigned for food and housing requirements (up to the amount of the market value of an equivalent insurance policy — the total cost divided by the number of participants in each risk class). 4 these tables also include three additional variants (denoted 1a, 1b, and 14a). 5 see fisher (1992) for more historical detail on the development of the poverty thresholds. 6 also critical to the definition of poverty is whether to use an absolute or relative measure. a relative measure sets the poverty standard at a fixed fraction, say 50 percent, of some measure of the population’s well-being such as median family income. thus, under a relative poverty measure, only if the incomes for the families at the bottom of the income distribution improve relative to the rest of the distribution would poverty decline. the alternate method of measuring poverty and the one currently in use in the u.s., at least in theory, is more or less an absolute measure. when constructing an absolute measure, one attempts to measure the minimal consumption levels of as many goods as possible. the cost of that consumption bundle is then increased to account for necessary goods not included by use of a “multiplier.” orshansky measured only the cost of a minimally adequate diet. other proposals have suggested adding shelter, clothing, and medical care to the list. we restrict the discussion here to absolute measures; most observers expect the u.s. poverty concept to retain this feature. 7 data are from the compensation and working conditions branch, u.s. bureau of labor statistics. the 1966 percentage is not strictly comparable to the 1996 figure. 8 exceptions are wages and salaries (we suspect that respondents sometimes report net instead of gross earnings) and workers’ compensation (payments for injuries on the job.) there are early indications that changes to the sipp questionnaire in 1996 have ameliorated these problems. 9 a national academy of sciences panel on the future of the sipp recommended moving toward the use of the sipp for official income and poverty measurement (citro and kalton, 1993). 10 a full review of budget-based approaches is in watts (1993). 11 there is also an issue about whether to use the official cpi or an experimental cpi created to correct for inaccurate measurement of housing costs in the official cpi prior to 1983. the next cpi market basket revision is scheduled for 1998. 12 this section drawn from appendix a of u.s. bureau of the census, 1996a. 13 in general, inventory changes are considered in determining net income from nonfarm self-employment; replies based on income tax returns or other official records do reflect inventory changes. however, when values of inventory changes are not reported, net income figures exclusive of inventory changes are accepted. the value of saleable merchandise consumed by the proprietors of retail stores is not included as part of net income. 14 in determining farm self-employment incomes, inventory changes are usually considered in determining net income only when they were accounted for in replies based on income tax returns or other official records which reflect inventory changes; otherwise, inventory changes are not taken into account. 15 child support paid and other inter-household transfers should theoretically be subtracted from income to avoid double counting, but the data necessary to do so are not collected. mailto:daniel.h.weinberg@ccmail.census.gov mailto:charles.t.nelson@ccmail.census.gov vol23/1 14 iassist quarterly overview while climate influences many social and behavioral phenomena, it is often poorly or incompletely represented in social science research. studies of elderly migration, for example, often rely on a single variable to represent the full set of climatic conditions found across the united states (walters 1994b). moreover, there is no reliable guide to the selection of the most appropriate climate variables. any single construct such as winter temperature can be represented by a variety of indicators — minimum daily temperature, average daily temperature, number of freezing days, number of below-zero days, number of heating degree-days, etc. although observed variables are essential in climatological research, statistically constructed indices may be more useful for many social and behavioral applications. this report describes the use of factor analysis to create five climate indices from a set of 37 original (observed) variables. these indices represent all the major components of near-surface climate variation within the united states. in addition, they offer at least three advantages over the original variables: 1) while any individual observed variable may be affected by measurement error, each index incorporates the variance common to more than one of the original variables. for instance, the difficulty of obtaining accurate snowfall measurements will produce more error in the observed variable (snowfall depth) than in an index that incorporates both snowfall and a number of related measures. 2) the five indices are uncorrelated and represent nearly 90 percent of the variance within the original set of 37 variables. there is no need to select a subset of the variables for use in multivariate studies since all five can be used together without danger of multicollinearity. 3) the data set is readily accessible to scholars whose primary interests lie outside climatology. (appendix a presents the complete set of indices for almost every first-order weather station within the coterminous united states.) in contrast, many of the data files distributed by noaa require expertise in the use of complex and sometimes discipline-specific data formats.1 along with the climate indices (factor scores), factor analysis produces a set of factor loadings that reveal the relationships among the original variables. the results of this analysis confirm that american climates are dominated by strong seasonal influences. in particular, summer air moisture and temperature are not closely linked to the corresponding winter conditions. previous research factor analysis, developed for use in psychometric research, has since achieved widespread application in the field of climatology — most often in the construction of climate classification schemes. r-mode factor analysis, a variant of the usual technique, can be used to reveal the relationships among a set of observed climate variables and to represent those variables through a smaller number of factors.2 the resulting indices (factor scores) are useful whenever it is necessary to represent the full range of climate variation through a limited number of variables, or whenever the underlying components of climate are more important than the observed values themselves. as a predictor of retirement migration, for example, an index of winter climate severity is probably more meaningful than the number of snow days or the average january temperature (walters 1994a). richman (1986) reviews the use of factor analysis in climate research. he describes six modes of analysis, which can be used to (1) classify geographic locations according to climate, (2) identify time periods in which climatic conditions remained stable, and (3) represent a large number of climate variables through a smaller number of factors. while many authors have focused on the first two goals, only a few have conducted the r-mode analyses that meet the third objective. micklin and dickason (1981), for example, found that 16 climate indicators for the soviet union could be adequately represented by just four factors. these factors — aridity, continentality, atmospheric turbidity, and thermality — captured 85% of the variance within the original set of variables. similar analyses have been undertaken for australia (puvaneswaran 1990), canada (powell 1977), greece climate indices for use in social and behavioral research by william h. walters spring 1999 15 (bartzokas and metaxas 1995), nigeria (olaniran 1986), and pakistan (oliver et al. 1978). using data for the state of maine, briggs and lemin (1992) found that 37 climate indicators could be represented by just three constructed indices. the climate of midland, texas, is apparently more complex, involving up to ten distinct factors (ladd and driscoll 1980). only two studies have presented r-mode factor analysis results for the entire united states. davis and kalkstein (1990) focus on weather rather than climate, however, while walters (1994a) uses pre-1970 data and evaluates only those sites near metropolitan areas. the r-mode analysis presented here is based upon more recent data and represents the full range of climate variation within the coterminous united states. data and methods data for 216 first-order weather stations were taken from the local climatological data series of the national oceanic and atmospheric administration (wood 1996). eighteen stations were excluded due to insufficient data. the temperature and precipitation data are site-adjusted averages, 1961 to 1990. all other variables are based on measurements made prior to 1994. the length of record varies by site and phenomenon but is typically 30 to 50 years. principal components analysis (pca) with varimax rotation3 was applied to the 37 variables shown in table 1. these variables include all the meaningful components of climate: annual, summer, and winter values of temperature, precipitation, humidity, cloud cover, wind speed, storm days, fog days and precipitation days; as well as related indicators such as snowfall, wind chill, and heat stress. pca, like other types of factor analysis, is an objective, empirical procedure that reapportions the variance within the original set of variables. the results reflect the pattern of correlations among these variables so that each factor usually represents a cluster of related measures. in this instance, 87.8% of the total variance can be represented by just five factors (five indices). these factors were rotated and interpreted according to the criteria suggested by cattell (1958), rummel (1970) and thurstone (1947). results varimax rotation always produces independent (uncorrelated) factors. in this case, each factor is conceptually distinct as well. that is, each has a unique and readily identifiable meaning. (see table 1.) the first factor, f1, represents winter temperature and snowfall. locations with high values of f1 tend to have mild winters, relatively few freezing days, little snowfall, and only modest seasonal temperature variation. in contrast, sites with low values of f1 can expect severe winter temperatures and heavy snowfall. to a lesser extent, f1 represents annual and summer temperatures. (high values of f1 correspond to high temperatures throughout the year.) factor 1 is not a straightforward indicator of summer temperature, however, since (1) another factor, f4, represents maximum daily temperature throughout the summer months and (2) the summer temperature variables most closely associated with f1 are strongly related to f4 as well. while winter temperature is fully represented by f1, summer temperature fails to emerge as a single, independent component of the climate system. the second factor, f2, is a summer air-moisture indicator representing summer precipitation, cloud cover, humidity, and storms. while summer temperature and humidity are often thought to occur in tandem, these results show that the two phenomena are not necessarily related. in particular, only one of the variables most closely associated with f2 (heat stress — humiture) is strongly related to both f1 and f2. the third factor, f3, is much like f2 but represents winter rather than summer conditions. locations with high values of f3 tend to have many rainy days, heavy cloud cover and high humidity throughout the cooler months. in contrast, places with low values of f3 are distinguished by relatively clear, dry winters. while the annual air moisture variables have high loadings on both f2 and f3, factor 3 is the best single indicator of year-round precipitation, cloud cover, and humidity. the fourth factor, f4, represents those aspects of summer temperature not included in factor 1. in particular, summer maximum daily temperature is most closely related to f4. (high values of f4 correspond to cool summers.) table 1 shows that the other summer temperature variables are also closely linked to f4 even though their primary association is with f1. the fifth factor, f5, is primarily a wind-speed indicator. it incorporates all three wind-speed variables (annual, summer, and winter) as well as the number of days with dense fog. taken together, the factor loadings confirm that american climates are dominated by strong seasonal influences. rather than forming a single precipitation factor, for instance, the various precipitation variables combine with other air-moisture indicators (cloud cover and humidity) to create two distinct seasonal factors, f2 and f3. likewise, summer temperature is at least partly independent of winter temperature. of the several components of climate, only wind speed and fog (factor 5) fail to display strong seasonal independence. 16 iassist quarterly spring 1999 17 the climate indices (factor scores) for each weather station are presented in appendix a.4 by mapping the highest and lowest scores, we can identify the spatial pattern associated with each factor. figure 1 reveals that each factor is spatially coherent — nearby locations have similar values — and that each has a distinctive geographical pattern. winter/annual temperature and snowfall (f1) vary with latitude, for instance, while summer air moisture (f2) is highest in the southeast and lowest in the west. figure 1 also helps illustrate why the summer temperature variables are associated with both f1 and f4. factor 1 shows the influence of latitude, primarily, while f4 best represents the distinction between continental and marine climates. summer temperature is therefore a function of both latitude and continentality. in contrast, winter temperature and snowfall can be adequately represented by a single factor (f1) that varies chiefly by latitude. conclusions the american climate system can be represented by just five indices — five sets of factor scores. because these scores are uncorrelated, all five can be used together — as explanatory variables, for instance — without danger of multicollinearity. the results of this analysis are consistent with previous research on the factor structure of american climates. in particular, five of the six factors identified in an earlier study (walters 1994a) can be seen here as well. this suggests that the factor structure has not changed over time and that it does not vary when new locations are added to the analysis. the relationships observed here are not necessarily valid for other countries or for particular regions of the u.s., however. the climate of queensland, australia, for example, does not display strong seasonality (puvaneswaran 1990). likewise, the climates of nigeria (olaniran 1986), pakistan (oliver et al 1978) and maine (briggs and lemin 1992) are dominated by regional and local factors not present in the united states at the national level. notes 1. see, for example, the first order summary of the day (http://www.ncdc.noaa.gov/onlineprod/tfsod/climvis/ ftppage.html). 2. richman (1986) provides a good overview of this technique. 3. several oblique and orthogonal rotation methods were evaluated empirically. while each method generated a similar set of factors, varimax gave the most robust results — the results that changed the least when random variation (representing error) was added to the original climate variables. 4. a machine-readable version of appendix a is available from the author. references bartzokas, a., and metaxas, d.a. 1995. “factor analysis of some climatological elements in athens, 1931-1992: covariability and climatic change.” theoretical and applied climatology 52: 195-205. briggs, r.d., and lemin, r.c., jr. 1992. “delineation of climatic regions in maine.” canadian journal of forest research 22: 801-811. cattell, r.b. 1958. “extracting the correct number of factors in factor analysis.” educational and psychological measurement 18: 791-838. davis, r.e., and kalkstein, l.s. 1990. “development of an automated spatial synoptic climatological classification.” international journal of climatology 10: 769-794. ladd, j.w., and driscoll, d.m. 1980. “a comparison of objective and subjective means of weather typing: an example from west texas.” journal of applied meteorology 19: 691-704. micklin, p.p., and dickason, d.g. 1981. “the climatic structure of the soviet union: a factor analysis approach.” soviet geography 22: 226-239. olaniran, o.j. 1986. “on the classification of tropical climates for the study of regional climatology: nigeria as a case study.” geografiska annaler 68a: 233-244. oliver, j.e., siddiqi, a.h., and goward, s.n. 1978. “spatial patterns of climate and irrigation in pakistan: a multivariate statistical approach.” archives for meteorology, geophysics, and bioclimatology 25b: 345-357. powell, j.m. 1977. “climatic classifications of the prairie provinces.” atmosphere 15: 27. puvaneswaran, m. 1990. “climatic classification for queensland using multivariate statistical techniques.” international journal of climatology 10: 591-608. richman, m.b. 1986. “rotation of principal components.” journal of climatology 6: 293-335. rummel, r.j. 1970. applied factor analysis. evanston: northwestern university press. thurstone, l.l. 1947. multiple-factor analysis. chicago: university of chicago press. walters, w.h. 1994a. “climate and u.s. elderly migration rates.” papers in regional science 73: 309-329. walters, w.h. 1994b. “place characteristics in elderly migration research.” bulletin of bibliography 51: 341-354. wood, r.a. 1996. weather of u.s. cities. 5th edition. new york: gale research. * william h. walters, albert r. mann library, cornell university, ithaca, ny 14853, usa. whw2@cornell.edu. (607) 255-7192 http://www.ncdc.noaa.gov/onlineprod/tfsod/climvis/ftppage.html http://www.ncdc.noaa.gov/onlineprod/tfsod/climvis/ftppage.html mailto:whw2@cornell.edu 18 iassist quarterly table 1 . rotated factor loadings a variable f1 f2 f3 f4 f5 h2 freezing days (annual) -0.97 — — — — 0.95 min daily temp (winter) 0.97 — — — — 0.95 avg daily temp (winter) 0.97 — — — — 0.96 heating degree days (annual) -0.96 — — — — 0.98 zero-degree days (annual) * -0.95 — — — — 0.92 avg daily temp (annual) 0.94 — — — — 0.99 snow days (annual) * -0.94 — — — — 0.92 snowfall (annual) * -0.94 — — — — 0.92 wind chill (winter) 0.93 — — — — 0.96 seasonal temp variation -0.80 — — -0.41 — 0.85 cooling degree days (annual) 0.79 — — -0.40 — 0.92 storm days (winter) * 0.72 0.44 — — — 0.75 avg daily temp (summer) 0.70 — -0.30 -0.51 — 0.92 heat stress — thi (summer) 0.68 0.58 — -0.33 — 0.92 ninety-degree days (annual) 0.66 — -0.33 -0.53 — 0.84 precipitation (summer) — 0.91 — — — 0.91 precipitation days (summer) — 0.88 — — — 0.88 storm days (annual) — 0.84 — -0.30 — 0.87 storm days (summer) — 0.81 — — — 0.78 cloud cover (summer) — 0.75 0.37 0.38 — 0.87 heat stress — humiture (summer) 0.48 0.74 — — — 0.86 humidity (summer) — 0.67 0.45 0.41 — 0.86 precipitation (annual) 0.32 0.63 0.47 0.34 — 0.86 humidity (winter) — — 0.89 — — 0.80 cloud cover (winter) -0.37 — 0.88 — — 0.92 precipitation days (winter) — — 0.81 0.41 — 0.86 cloud cover (annual) -0.44 0.35 0.74 — — 0.91 humidity (annual) — 0.49 0.71 0.30 — 0.85 precipitation days (annual) -0.31 0.44 0.69 0.36 — 0.89 precipitation (winter) 0.42 — 0.53 0.51 — 0.74 fog days (summer) * — 0.42 — 0.70 — 0.79 max daily temp (summer) 0.55 — -0.39 -0.65 — 0.91 wind speed (annual) — — — — 0.92 0.94 wind speed (summer) — — — — 0.88 0.89 wind speed (winter) — — — — 0.87 0.92 fog days (winter) — — 0.37 — 0.61 0.68 fog days (annual) — — — 0.58 0.58 0.75 % variance explained 39.3 25.5 11.4 8.2 3.5 cumulative % 39.3 64.8 76.2 84.3 87.8 a principal components analysis with varimax rotation. annual = average for all months. summer = average for june, july, and august. winter = average for december, january, and february. values in bold type are the highest loadings for each variable. loadings between -0.30 and 0.30 are not shown. variables marked with an asterisk (*) were entered in cube root form to maintain linearity. communality (h2) indicates the proportion of the variance within each variable that is shared with the other variables in the set. spring 1999 19 appendix a climate indices (factor scores — regression method) for 216 first-order weather stations in the coterminous united states. sixteen stations were excluded due to insufficient data. each factor has a mean of 0.00 and a standard deviation of 1.00. weather station state f1 f2 f3 f4 f5 birmingham al 0.74 0.88 0.28 -0.07 -0.95 huntsville al 0.66 0.78 0.38 0.06 -0.50 mobile al 1.36 1.67 0.22 -0.14 0.30 montgomery al 1.18 0.81 0.22 0.02 -0.74 fort smith ar 0.51 0.43 -0.02 -0.66 -0.52 little rock ar 0.79 1.02 0.23 -0.49 -0.28 flagstaff az -1.10 -0.32 -1.93 1.17 -1.32 phoenix az 1.49 -1.43 -2.25 -1.17 -0.56 tucson az 0.95 -0.56 -2.57 -0.60 -0.20 winslow az -0.29 -0.94 -1.85 -0.67 -0.20 yuma az 1.69 -1.83 -2.78 -0.90 0.18 bakersfield ca 1.49 -2.86 0.01 -1.36 0.03 fresno ca 1.58 -2.92 0.87 -1.58 0.46 long beach ca 1.65 -1.97 -1.25 2.39 -0.53 los angeles (airport) ca 1.63 -1.90 -1.36 3.03 -0.31 los angeles (civic center) ca 1.70 -1.79 -1.76 2.41 -0.99 redding ca 1.16 -2.17 0.23 -0.82 -0.18 sacramento ca 1.55 -2.77 0.95 -0.85 0.69 san diego ca 1.59 -1.85 -1.23 2.66 -0.67 san francisco (airport) ca 1.33 -2.46 0.07 1.68 0.72 santa maria ca 1.33 -2.14 -1.46 3.96 -0.22 stockton ca 1.55 -2.98 1.03 -1.14 0.75 alamosa co -1.59 -0.43 -1.54 0.37 -0.56 colorado springs co -1.07 0.40 -2.86 1.49 -0.19 denver co -0.92 -0.14 -1.87 0.46 -0.55 grand junction co -0.48 -1.22 -0.72 -1.34 -0.30 pueblo co -0.74 -0.23 -2.32 -0.03 -0.04 bridgeport ct -0.13 0.06 -0.38 1.19 0.74 hartford ct -0.53 0.14 -0.15 0.97 -0.46 washington (dulles) dc -0.11 0.29 -0.02 0.70 -0.72 washington (national) dc 0.20 0.19 -0.18 -0.02 -0.15 wilmington de 0.02 0.17 -0.06 0.70 -0.06 daytona beach fl 1.56 1.58 -0.09 0.01 0.06 fort myers fl 1.71 2.44 -0.60 -0.57 0.01 jacksonville fl 1.40 1.43 0.06 -0.03 0.01 key west fl 2.01 1.18 -0.32 -0.55 0.61 miami fl 1.77 1.91 -0.45 -0.12 -0.02 orlando fl 1.61 1.98 -0.30 -0.40 0.18 pensacola fl 1.46 1.38 0.26 -0.15 0.24 tallahassee fl 1.44 1.91 0.06 0.18 -0.39 tampa fl 1.62 1.85 -0.32 -0.46 0.03 west palm beach fl 1.72 1.87 -0.17 -0.29 0.06 athens ga 0.86 0.67 -0.18 0.71 -0.37 atlanta ga 0.80 0.65 -0.13 0.59 0.06 augusta ga 0.91 0.85 -0.11 0.24 -0.79 columbus ga 1.11 0.86 0.26 -0.19 -0.76 macon ga 1.05 0.80 0.04 0.00 -0.41 savannah ga 1.16 1.35 -0.28 0.30 -0.12 des moines ia -0.75 0.59 -0.09 -0.70 0.38 sioux city ia -0.89 0.41 -0.15 -0.75 0.46 waterloo ia -1.02 0.51 0.09 -0.53 0.33 20 iassist quarterly appendix a cont... boise id -0.21 -2.00 0.71 -1.04 0.07 pocatello id -0.84 -1.56 0.74 -1.39 0.35 chicago il -0.72 0.29 0.44 -0.41 0.08 moline il -0.70 0.64 0.00 -0.39 -0.02 peoria il -0.52 0.48 0.41 -0.56 0.14 rockford il -0.83 0.47 0.27 -0.31 0.06 springfield il -0.37 0.45 0.37 -0.75 0.45 evansville in 0.05 0.37 0.52 -0.54 -0.56 fort wayne in -0.58 0.25 0.86 -0.48 0.01 indianapolis in -0.33 0.42 0.75 -0.39 -0.06 south bend in -0.67 0.31 1.16 -0.42 0.06 concordia ks -0.49 0.63 -0.33 -1.07 0.42 dodge city ks -0.21 0.26 -1.67 -0.29 0.74 topeka ks -0.42 0.86 0.07 -0.87 -0.06 wichita ks -0.07 0.13 -0.19 -1.33 0.23 jackson ky 0.10 0.80 0.32 1.31 -0.66 lexington ky -0.05 0.54 0.54 -0.03 -0.24 louisville ky 0.02 0.47 0.42 -0.25 -0.65 paducah ky 0.28 0.69 0.42 -0.33 -0.44 baton rouge la 1.42 1.36 0.36 -0.18 -0.22 lake charles la 1.59 1.09 0.83 -0.58 0.47 new orleans la 1.56 1.27 0.71 -0.45 0.01 shreveport la 1.17 0.41 0.46 -0.78 -0.09 boston ma -0.26 0.02 -0.50 1.21 0.80 worcester ma -0.56 0.04 -0.48 2.21 0.61 baltimore md 0.07 0.11 -0.27 0.55 -0.06 caribou me -1.79 0.39 0.35 0.69 0.23 portland me -0.83 0.12 -0.39 1.94 -0.37 alpena mi -1.30 0.06 0.83 0.26 -0.81 detroit mi -0.66 0.04 0.85 -0.32 0.14 flint mi -0.85 0.05 0.83 -0.16 -0.08 grand rapids mi -0.83 0.09 1.32 -0.40 -0.07 houghton lake mi -1.24 -0.05 1.07 0.03 -0.46 lansing mi -0.88 0.10 1.15 -0.41 -0.10 muskegon mi -0.82 -0.11 1.34 -0.25 0.07 sault ste. marie mi -1.49 0.09 1.17 0.68 -0.38 duluth mn -1.69 0.48 -0.21 0.90 0.42 international falls mn -2.07 0.47 -0.04 0.09 -0.64 minneapolis-st. paul mn -1.28 0.38 -0.13 -0.58 0.12 rochester mn -1.26 0.47 0.23 -0.47 1.27 st. cloud mn -1.53 0.28 -0.21 0.03 -0.68 columbia mo -0.21 0.52 0.13 -0.43 0.19 kansas city mo -0.29 0.61 -0.35 -0.51 0.54 springfield mo 0.01 0.63 -0.08 -0.44 0.44 st. louis mo -0.05 0.40 0.38 -0.90 0.02 jackson ms 1.11 0.92 0.61 -0.52 -0.48 meridian ms 1.12 0.78 0.39 0.03 -0.93 tupelo ms 0.83 0.58 0.37 -0.16 -0.74 billings mt -1.11 -0.64 -0.89 -0.11 0.39 glasgow mt -1.45 -0.61 0.01 -1.12 0.41 great falls mt -1.29 -0.54 -0.64 -0.13 0.76 helena mt -1.28 -0.66 -0.23 -0.47 -0.85 kalispell mt -1.14 -1.12 1.31 0.02 -0.95 missoula mt -0.95 -1.22 1.29 -0.49 -0.96 asheville nc 0.20 0.86 -0.52 2.18 -0.42 cape hatteras nc 0.98 0.69 0.40 0.37 0.57 spring 1999 21 appendix a. cont... charlotte nc 0.58 0.46 -0.31 0.63 -0.59 greensboro nc 0.37 0.62 -0.36 0.89 -0.55 raleigh nc 0.49 0.64 -0.41 0.91 -0.51 wilmington nc 0.90 1.09 -0.13 0.49 -0.17 bismarck nd -1.62 -0.03 -0.27 -0.71 -0.01 fargo nd -1.66 0.14 -0.09 -0.83 0.70 williston nd -1.61 -0.31 -0.08 -1.00 0.02 grand island ne -0.86 0.45 -0.03 -1.06 0.43 lincoln ne -0.77 0.46 0.19 -1.35 0.17 norfolk ne -0.95 0.45 -0.63 -0.69 0.58 north platte ne -1.02 0.24 -0.92 -0.36 0.13 omaha (eppley) ne -0.71 0.55 -0.35 -0.69 0.29 omaha (north) ne -0.73 0.55 -0.45 -0.41 -0.17 scottsbluff ne -1.07 -0.01 -1.22 -0.36 0.11 valentine ne -1.22 0.19 -1.06 -0.72 -0.18 concord nh -1.02 0.13 -0.40 1.55 -1.01 mount washington nh -1.10 0.34 1.00 4.13 11.90 atlantic city (nafec) nj -0.01 0.21 -0.18 1.10 0.20 newark nj -0.05 0.17 -0.16 0.34 0.13 albuquerque nm -0.18 -0.63 -2.34 -0.04 -0.22 roswell nm 0.22 -0.46 -2.18 -0.05 -0.05 elko nv -0.87 -1.61 -0.24 -0.64 -1.28 ely nv -1.26 -1.20 -1.11 -0.59 0.06 las vegas nv 0.86 -1.90 -2.77 -0.95 0.31 reno nv -0.44 -2.10 -0.97 -0.20 -0.93 winnemucca nv -0.67 -1.92 -0.46 -0.89 -0.51 albany ny -0.92 0.27 0.34 0.34 -0.52 binghamton ny -0.89 0.15 1.01 0.77 0.17 buffalo ny -0.81 0.05 1.56 -0.34 0.47 new york (central park) ny -0.02 -0.02 -0.30 0.35 -0.38 new york (jfk) ny 0.04 0.10 -0.46 1.17 0.75 new york (la guardia) ny -0.02 0.07 -0.51 0.70 0.69 rochester ny -0.88 -0.06 1.29 -0.17 -0.38 syracuse ny -0.99 0.15 1.41 -0.28 -0.52 akron-canton oh -0.61 0.25 1.00 0.03 -0.14 cincinnati oh -0.26 0.46 0.60 -0.06 -0.29 cleveland oh -0.67 0.14 1.23 -0.42 -0.03 columbus oh -0.46 0.39 0.72 -0.07 -0.67 dayton oh -0.40 0.26 0.74 -0.27 0.03 mansfield oh -0.58 0.23 0.95 -0.08 0.42 toledo oh -0.70 0.20 0.85 -0.27 -0.29 youngstown oh -0.72 0.22 1.25 0.09 -0.17 oklahoma city ok 0.42 -0.01 -0.37 -1.11 0.94 tulsa ok 0.37 0.30 0.00 -1.25 0.30 astoria or 0.77 -1.09 2.33 2.61 -0.35 eugene or 0.87 -2.00 2.58 0.79 0.11 medford or 0.65 -2.52 1.85 -0.60 -0.49 pendleton or 0.00 -2.29 1.18 -0.82 0.26 portland or 0.57 -1.65 2.16 0.52 -0.34 salem or 0.56 -1.91 2.32 0.51 -0.49 allentown pa -0.34 0.28 0.02 0.59 -0.21 erie pa -0.72 0.18 1.50 -0.42 0.16 middletown/harrisburg pa -0.22 0.18 -0.02 0.49 -0.79 philadelphia pa -0.02 0.17 -0.13 0.49 -0.04 pittsburgh pa -0.61 0.23 0.87 0.15 -0.54 wilkes-barre/scranton pa -0.64 0.19 0.42 0.52 -0.64 22 iassist quarterly appendix a. cont... williamsport pa -0.55 0.49 0.26 0.86 -0.79 providence ri -0.33 0.08 -0.33 1.26 0.16 charleston sc 1.08 1.35 -0.09 0.18 -0.03 columbia sc 0.87 0.96 -0.20 0.29 -0.68 greenville-spartanburg sc 0.66 0.58 -0.38 1.07 -0.62 aberdeen sd -1.43 0.10 -0.27 -0.82 0.46 huron sd -1.30 0.21 -0.31 -0.92 0.57 rapid city sd -1.16 -0.01 -1.14 -0.13 0.37 sioux falls sd -1.17 0.33 -0.20 -0.82 0.55 bristol tn 0.08 0.55 0.14 1.25 -1.42 chattanooga tn 0.57 0.79 0.25 0.43 -1.09 knoxville tn 0.34 0.59 0.30 0.58 -0.88 memphis tn 0.79 0.48 0.38 -0.73 -0.14 nashville tn 0.38 0.63 0.33 -0.19 -0.55 abilene tx 0.73 -0.23 -1.01 -1.09 1.03 amarillo tx 0.00 0.00 -2.08 -0.18 1.68 austin tx 1.40 -0.21 0.06 -0.93 0.45 brownsville tx 2.03 -0.52 0.93 -1.51 1.55 corpus christi tx 1.82 -0.36 0.85 -1.38 1.69 dallas-forth worth tx 1.02 -0.15 -0.13 -1.41 0.74 del rio tx 1.29 -0.60 -0.63 -1.28 0.75 el paso tx 0.44 -0.66 -2.64 -0.31 -0.19 houston tx 1.41 0.67 0.76 -0.91 -0.03 lubbock tx 0.25 -0.08 -1.77 -0.38 1.21 midland-odessa tx 0.64 -0.54 -1.68 -0.61 0.94 port arthur tx 1.59 1.05 0.96 -0.96 0.75 san angelo tx 0.80 -0.47 -1.02 -1.05 0.58 san antonio tx 1.41 -0.30 0.20 -1.19 0.53 victoria tx 1.63 0.29 0.81 -1.21 0.98 waco tx 1.18 -0.21 0.20 -1.73 1.08 wichita falls tx 0.67 -0.06 -0.62 -1.43 1.07 salt lake city ut -0.36 -1.30 0.42 -1.68 0.05 norfolk va 0.56 0.47 -0.18 0.50 0.34 richmond va 0.26 0.56 -0.10 0.51 -0.57 roanoke va 0.01 0.42 -0.58 0.95 -0.63 burlington vt -1.34 0.25 0.51 0.23 -0.60 olympia wa 0.54 -1.67 2.56 1.89 -0.19 quillayute wa 0.76 -0.73 2.68 3.43 -1.17 seattle-tacoma wa 0.55 -1.53 1.78 1.51 -0.07 spokane wa -0.42 -1.99 1.53 -0.35 0.51 yakima wa -0.34 -2.24 0.82 -0.84 -0.52 green bay wi -1.18 0.23 0.16 0.14 -0.05 la crosse wi -1.10 0.56 -0.07 -0.15 -0.48 madison wi -1.06 0.37 0.25 -0.14 -0.04 milwaukee wi -0.85 0.21 0.20 0.22 0.51 beckley wv -0.47 0.68 0.68 1.15 -0.38 charleston wv -0.01 0.73 0.21 1.71 -0.86 elkins wv -0.70 0.93 0.70 1.79 -1.32 huntington wv -0.07 0.62 0.36 1.24 -1.00 casper wy -1.38 -0.51 -1.04 -0.42 0.85 cheyenne wy -1.23 0.19 -2.21 1.07 0.77 lander wy -1.40 -0.89 -1.30 -0.11 -1.12 sheridan wy -1.40 -0.51 -0.51 -0.33 -0.90 38 iassist quarterly 2010 / 2011 iassist quarterly dda hopes to take on the obligation as the national repository for data archiving of qualitative data archiving and disseminating qualitative data in denmark kick off for a new archival service by anne sofie fink kjeldgaard1 abstract the article describes the situation for archiving qualitative data in denmark. presently denmark does not have infrastructure for preserving or sharing qualitative data. a pilot project aimed at developing documentation standards is being carried out at the moment. based on this project and recommendations from fellow data archives with established services for qualitative data, a launch of a service for qualitative data will be made in the near future. keywords: data archiving, qualitative vs. quantitative data, secondary use, barriers for re-use of qualitative data, research culture, data services, denmark introduction the danish data archive (dda) was established in 1973 as a national data service for human centred quantitative research carried out primarily in the social sciences but also in medical science and in history. in 1993 the dda became part of the danish state archives. the dda acquires, preserves, and disseminates machine-readable research data. the issue of archiving qualitative data has been discussed in the dda since 2000 (fink, 2000a; fink. 2000b). the reason for the long deliberation is to a large extent the distinctive nature of qualitative data as well as the researcher’s relationship towards the data. setting the scene – qualitative data archiving qualitative data is unstructured, without common format, personally sensitive and so on. “qualitative data are normally relatively messy, unorganized data.” (mccracken, 1988: 19) these messy, unorganized data are the product of a personal encounter between research object/ respondent and researcher. due to the personal involvement in data production the researcher feels responsible towards the data with significant consequences for data archiving and data dissemination (mauthner at al, 1998; gillies and edwards, 2005). this is in contrast to the well-defined structure of quantitative survey-based, de-personalised data material (kuula, 2000; rasmussen, 2000). at the moment dda holds only a few qualitative studies archived in original format. there is no existing infrastructure for qualitative data archiving in denmark. however the dda has for years taken part in cross national discussions, meetings and workshops concerning archiving of qualitative data often led by esds (economic and social data service) qualidata, at the university of essex (corti, 2000). iassist quarterly 2010 / 2011 39 iassist quarterly the dda hopes to take on the obligation as the national repository for data archiving of qualitative data by setting up a unit for qualitative data alongside the unit for quantitative data. at the moment activities concerning qualitative data are funded by the dda. it will have to be considered as the dda is moving in the direction of becoming a data archive for both quantitative and qualitative data if sufficient resources for this development are available. in denmark the danish council for independent research for the social and medical sciences requires data archiving in dda as a prerequisite for funding research activities incorporating collection of survey data. these requirements on behalf of the independent research councils are core to dda’s activity. at the moment qualitative data is not mentioned in the policies of the research councils. it is our hope and expectation that an initiative concerning a service for qualitative data from the dda will motivate the research council to expand their requirement to embrace qualitative data as well. data re-use in academic literature data sharing of major quantitative data materials such as the election studies or the world value studies is carried out informally among researchers as well as formally through the dda. more and more often data can be retrieved freely available as web resource. however, data re-use is a research practise suffering from a complete lack of literature describing and discussing data re-use as part of the researcher’s methodological palette in a danish research context. in the danish research context literature promoting and describing data sharing and re-use as a scientific approach is lacking (kjeldgaard et al, 2008). this alone is not an obstacle to sharing or re-using data. however, as long as data sharing and re-use remains an informal practice data re-use/secondary analysis will not gain the academic legitimacy that the approach deserves. obviously the dda will have a pivotal role in the promotion of data reuse by adding supporting infrastructure, standards, tools, etc. hopefully this effort will stimulate articles and textbook chapters presenting reuse analysis as a methodological approach to be positioned as a viable alternative to traditional approaches for empirically based research activities. preparatory steps a preparatory step in the direction of archiving qualitative data was made in an article by fink (2000a) in which she points out how researchers’ involvement in the data production process (as interviewer, as observer, etc.) is creating an obstacle to data archiving. as a consequence qualitative data is of a nature that does not comply with data archiving and re-use in the way structured quantitative data do. the issues raised in the article were further developed by six in-depth interviews with researchers collecting qualitative data. the interview material was reported by fink (2000b)2. findings in the interview data were in fact strikingly similar to findings reported by broom et al (2009). the article concludes with a list of suggestions about how to handle the relationship between researcher and data archive that complies with the special nature of qualitative data. the following suggestions for a service for qualitative data archiving were made: handover of data just after data collection. in this way the risk of losing data or mixing up different versions of the data set is minimised. additionally, data is handed over at the stage in the research cycle where the researcher is exclusively focused on data and data quality rather than on analysis or publication as will be the case when he has moved on in the cycle. publications integrated in the archival unit of the qualitative data set. enlarging the archival unit for qualitative data sets encourages secondary users of the data to become informed about the interpretations of the primary researcher. in this way the secondary users are guided towards potential interpretation ‘span’, that is, taking into account the context of the original research as well as documenting the context of the re-use as well. distinction between access and re-use. due to potential concerns about misuse of data among primary researchers a possibility of limiting access for secondary users can be offered. access restrictions only allow users to view/browse data and restrict them from performing actual re-analysis of data. privileged access. to mirror informal data sharing among colleagues and research partners, primary researchers could be allowed to name persons who could be allowed free access to data. personal acquaintance. part of the resistance towards qualitative data archiving is due to the feeling of personal insecurity of moving into unfamiliar territory. being personally acquainted with the data archivist will to some extend compensate for this. dialogue. the personal contact between the researcher and the data archivist should be founded on an on-going dialogue between relevant research environments and the data archive. supervision. both paper-based and web-based resources should be at hand for informing researchers about data archiving of qualitative data at dda. in particular, it should be explained that data documentation be made an integrated part of the research process. these suggestions will guide and inspire the service the dda will set up for qualitative data. future development at the moment the dda has taken on a pilot project together with the danish national centre for social research (sfi). the project objective is to develop a documentation standard for qualitative data sets that is in line with the needs of depositors, re-users and the data archive. obviously the dda will seek inspiration, recommendations and best practices from fellow data archives experienced in the field of qualitative data archiving, especially uk data archive and finnish social science data archive (fsd). as mentioned above archiving of qualitative data faces different kind of challenges. the most striking of these are described below. research culture for sharing or re-using qualitative data neither qualitative data sharing nor re-use is practiced formally in danish research environments. as mentioned, data sharing in general seems to be something that is carried out informally. therefore the adoption of data sharing and re-use through a data archive can be expected to be slow. promotion of data sharing and re-use of qualitative data is a responsibility the dda has to take on. 40 iassist quarterly 2010 / 2011 iassist quarterly resistance from researchers resistance from researchers is often due to ethical and methodological concerns, concern for respondents, handling of sensitive information, disclosure issues, the researcher’s active role, etc. (e.g. fink, 2000b). furthermore, the data is viewed by some as the personal property of the researcher (van den berg, 2005). obviously consent for data archiving and data sharing by respondents is also an issue related to resistance. often it seems that researchers take on a role as the respondents’ protector against potential complementary research interest from other researchers. sometimes it seems that researchers are more protective than the respondents actually expect them to be. it should be noted that danish legislation does not prevent re-use of data originally collected for research purposes (daasnes, 2000). an argument that is missing in this debate is the point that it may be a way of respecting to respondents’ efforts to make sure that their data can be re-used for further research purposes (hakim, 1982). economic resources at the moment initiatives concerning services for qualitative data archiving will be sponsored by the dda. but additional funding is an important issue to be addressed as it is necessary to sustain the initiative with sufficient resources in the future. formal organisations like iassist and cessda as well as informal groups like the bremen group might be of assistance by providing information, materials and web resources, e.g. presentations of success stories for qualitative data archiving. additional examples could be from the established world of data archives e.g. from the uk data archive or finland (fsd) or it could from theme centred organisations such timescapes (see www.timescapes.leeds.ac.uk/) or the archive for life course research in bremen (see www.lebenslaufarchiv.unibremen.de/). additionally journal articles making discussing data archiving and re-use of qualitative data sets as well as presentations of actual research projects and theoretical discussions of the issue will be helpful. to conclude, data archiving of qualitative data is still in an early phase in denmark. these challenges mentioned above – research culture for sharing or re-using qualitative data, resistance from researchers and economic resources – will be taken up by the dda. but it is critical to remember that dda is part of important organisational network that will be able to support its efforts. thanks to international, europeanbased cooperation and the preparatory steps the dda has taken already we feel well prepared to take up the challenge of archiving and disseminating qualitative data alongside with quantitative data in the near future. references broom, a, lynda c and michael, e. (2009). ‘qualitative researchers’’ understandings of their practice and the implication for data archiving and sharing’. sociology. 43. corti, l. (2000). ‘progress and problems of preserving and providing access to qualitative data for social research—the international picture of an emerging culture’. fqs forum qualitative social research. 1 (3) daasnes, c. (2008). ‘persondataloven – regler og praksis for god databehandlingsskik’. metode & data. (94). fink, a. (2000a). ‘the role of the researcher in the qualitative research process a potential barrier to archiving qualitative data’. fqs forum qualitative social research. (1) 3. fink, a. (2000b). ‘kvalitative data i dansk data arkiv’. metode & data. 83 (2). gillies, v and edwards, r. (2005). ‘secondary analysis in exploring family and social change: addressing the issue of context’. forum qualitative social research. 6 (1) hakim, c. (1982). secondary analysis in social research. london: george allen & unwin kjeldgaard fink, a, bredahl, l and stenvig, b. (2008). ‘genbrug af forskningsdata overset tilbud eller anvendt mulighed’. konferenceudgivelse: symposium i anvendt statistik (2008) p. 291-298. kuula, a. (2000, december). ‘making qualitative data fit the ‘data documentation initiative’ or vice versa?’ forum qualitative sozialforschung / forum: qualitative social research .1(3). [online] available at: http://www.qualitative-research.net/fqstexte/300/3-00kuula-e.htm. mauthner, n, parry, o and backett-milburn, k. (1998). ‘the data are out there, or are they? implications for archiving and revisiting qualitative data’. sociology. (32) mccracken, d. g. (1988). the long interview: qualitative research methods vol. 13. newbury park, ca: sage rasmussen, k., b. (2000). datadokumentation – metadata for samfundsvidenskabelige undersøgelser. gylling, dk: odense universitetsforlag van den berg, h. (2005). ‘reanalyzing qualitative interviews from different angles: the risk of decontextualization and other problems of sharing qualitative data’. forum: qualitative social research. 6 (1) notes 1. contributor details: anne sofie fink kjeldgaard asf@dda.dk, data archivist and senior researcher, ph.d., danish data archive 2. a short summary of the article in english is available by request to the author. iassist quarterly 2010 / 2011 41 iassist quarterly bremen workshop: qualitative longitudinal research and qualitative resources in europe: mapping the field and exploring strategies for development april 2009 from census to integrated population data or sociodemographic accounts by mathieu vliegenl netherlands central bureau of statistics introduction the census of peculation taken in the netherlands in 1971 seems to be the last one in a series which started in 1830. in 1981 the government postponed the population census which according to the 1970 census law had to be taken in that year. meanwhile, the government has announced to parliament a proposal to revoke the 1970 census law. at the same time the government submitted an alternative statistical programme consisting of a set of register-based enumerations in combination with survey research during a period of circa ten years. first reactions indicate that parliament will be in favour of such revocation. in several publications attention has been paid to some underlying factors with regard to this development (redfem 1986, choldin 1987). relevant for the topic under study here post-censal surveys are the public and parliamentary discussions on the privacy issue in relation to a census, in which pleas in favour of an absolute anonymity at taking a census are a central topic: data collection should take place without names and addresses. in this way it was the argument the privacy of the individual citizen would be guaranteed. furthermore, it is worthwhile reminding the demands in these discussions for a voluntary participation in the census by the individual citizen. every legal obligation in respect to this participation should be rejected, in particular a legally imposed penalty for not cooperating. such obligations tlie argument was would infringe the fundamental rights of the individual citizen with respect to his willingness to give information about himself to others. ultimately those pleas and demands were not yet effective on the 1971 census as such. however, they affected tlie possibility laid down in the executive regulations of the 1971 census with regard to the keeping of a 10% sample from that census. by linking this sample to the 1981 census it was intended to obtain longitudinal data on among others occupational and educational mobility as well as on changes in household status. to carry out this longitudinal study it should be necessary to keep the names and addresses of the sampled persons during a period of more than ten years. more and more, however, this procedure was considered in public opinion and in parliament as an encroachment on the personal privacy of the citizen. although the importance of such longitudinal research for statistical purposes was recognized by parliament, utlimately it was staled that this kind of research connected with a census had to be subordinated to the interests of the individual citizen. consequently, a few weeks before census date, the minister politically in charge of the 1971 census was obliged to cancel a possible keeping a 10% sample from the 1971 executive regulations. since the 1971 census various sources and methods have been come into use for the construction of a system of population statistics as comprehensive as possible from a demographic as well as from a social and economic point of view cvliegen and van de stadt, 1988). the following developments should be mentioned specifically: • starting enumerations from the municipal population registers and their respective enlargements in content; • more extensive exploitation of the municipal reporting with regard to vital events and changes of residence; • conducting regularly large-scale sample surveys on the labour and housing market; • ^plication of methods to obtain estimates at the level of the total population; • the accomplishment of an automated (and yearly updated) register of all addresses in the netherlands with geocodes, a so-called geographic base file (gbf), which jointly with the municipal population registers also is in use as sampling frame. these developments have brought a statistical programme into practice which in view of the results produced is highly comparable to that of a conventional census in combination with post-censal surveys. some parts of this programme show characteristics analogous to a conventional census. in particular this regards the system of demographic statistics which is based on periodical enumerations from the population registers in combination with the processing of monthly municipal reports on vital events and migrations. to a certain extent this applies also to the large-scale sample surveys which generate benchmark data in the spring 1990 social and economic field. these surveys have to fulfill this function, since the content of the municipal population registers is restricted to purely demographic data. of course, in delivering the social and economic benchmarks they cannot completely compete with a census. for example, they cannot provide these data with the same regional detail as a census does. however, the extent of regional detail is sufficient for the level of which national policies regarding the labour and housing market are made. at the same time, these large-scale sample surveys show characteristics inherent to so-called post-censai surveys. first of all, they provide in-depth information on some specific groups of the population. secondly, the results of the system of demographic statistics (comparable to the results of a census in uie classical sense) are used for making the relevant estimations on the level of the total population from the survey results. finally, the sources for compiling the demographic statistics (the municipal population registers) themselves are sometimes used as frame for the selection of the sampling units. these points will be discussed more deeply in the next sections. it should be pointed out, however, that in the system of population statistics built up uniii now, no use has been made of the technique of record linkage. in the near future the application of this technique will presumably not be used either. reasons of public and political nature rather than technical impossibilities prevent such af^lications at present. therefore, the above sketched statistical programme cannot be qualified as a register-based census (redfem, 1986). rather it is an approach in which both kind of statistical instruments are jointly applied in such a way that (a) societal changes relative to the field of inquiry in question can be discerned almost continuously by the statistical information provided, and (b) actual developments can be taken as a subject of inquiry in the research programme almost instantly. the system of demographic statistics 2.1. the continuous population accounting: the basis of the system the main soiu^ce for compiling demographic statistics are the municipal population registers. these registers have been introduced in 1 850 and set up using the data collected at the population censusof 1849. since then these registers have been continuously updated according to the regulations laid down in the system of population accounting (van den brekel, 1977). until world war ii the registers consisted of fam.ily documents in which all members of the family were listed. since then the personal card has been introduced. this personal card is made out at birth and follows the individual person during his whole life time. an essential feature of the population registration system is its decentralization. this implies that each municipality keeps its own population register. persons are registered in the population register of the municipality in which they normally reside. the regulations for the systematical updating of the municipal population registers refer among others to the registration of all changes by birth, death, marriage and dissolution of marriage by the local registrar of the civil registration to whom they have to be reported. furthermore they consist of detailed rules with respect to taking permanent residence in and removals from a municipality as well as all changes of residence within a municipality (see figure 1). all these regulations are intended to guarantee the completeness and accuracy of the municipal population registers. 2.2. system ofdemographic statistics and register-based enumerations. the system of population accounting also includes regulations concerning the municipal repxarting of the various vital events and changes of address to the netherlands central bureau of statistics (see also figure 1). this reporting enables the cgs to compile continuously statistics on natality, mortality and nupiialiiiy as well as statistics on internal and external migration. the municipal information on vital events and migrations is also used at the cbs for updating a statistical file with aggregated data on the demographic composition of the population. this file set up for the first time using the 1947 census data enables ijie cbs to compile annually statistics on the size and composition of the population for every municipality. the file with data on the demographic composition of the population has to be revised regularly. in the course of tim.e deviations from the real situation are inevitably introduced due to the above mentioned method of updating this file. consequently, the relevant statistics are becoming less reliable over time. the revision of the 1947 file took place in 1960, still based on the results of the census taken in that year. since 1971, however, complete enumerations from the municipal population registers are used for revising purposes. the underlying factors for this switch were among others satisfactory results from checks on the quality of the data in the municipal population registers obtained at the 1971 census as well as technological developments in data processing. at every enumeration the amount of characteristics in the file has been augmented (see figure 2). at present by means of this file statistical infomiation is supplied annually for each municipality on e.g. the total population by age, sex and marital status, and the alien population by age, sex and marital status. 2.3 the system ofpopulation statistics: its reliability the basic demographic data in the municipal population registers have a high degree of accuracy. it is in the interest of the citizen that his data have been accurately recorded in the municipal population register in view of, among others, a request regarding a resident permit, a 22 assist quarterly driving licence, a passport, as well as several benefits or grants. moreover, it is in the interest of the municipalities that the data of its citizen are accurately registered. the financial contribution of the central government to the municipalities depends on, among others, the number of its inhabitants. by law, these figures have to be determined annually by the cbs after mutual control with the municipalities. the procedure has been laid down in the above mentioned system of population accounting. therefore, the reliability of the demographic statistics is in general high. this can also be concluded from the results of the register-based enumerations carried out for revising purposes (for figures see table 1). the deviations found from the confrontation of the results from these enumerations with those from the yearly updated statistical file are usually very small for various age groups. they are higher for categories of marital status and for the alien population, preponderantly due to imcompletenesses in the reporting of the relevant changes to the cbs. 2.4. register-based demographic statistics: some prospects in addition to the data used in register-based enumerations for revising purposes, the municipal population registers contain still other data which from a statistical point of view are of great value. up till now these data were not a topic for regular register-based enumerations due to the lack of proper automated processing systems with regard to these registers. until recently, some municipalities had set up duplicates of their registers in an automated fwm, but the systems developed were for a great part different from each other. other municipalities had duplicates of their population register in a mechanised form: either punch-cards or address-plates. still other municipalities (mostly the smaller ones) had no duplicates at all. at present an automated system of municipal population accounting called the municipal administration of the population (map) is developed under the supervision of the ministry of home affairs. not only the municipal population registers itself is taken into regard, but also the reporting of the relevant changes between the municipalities mutually and between a municipality and its clients (including the cbs). this reporting will take place by means of an electronic network. the implementation of the whole system has been planned to take place within a few years from now. it is obvious that the automation of the municipal population register in a similar way as described by the map will give rise to further developments on the statistical field (verhoef and van de kaa, 1987). for example, by means of a register-based enumeration on the population by status in the family (spouse/lone parent, child, not living in a nuclear family) it will also be possible to compile regularly benchmark statistics on families and (groups oq persons not living in a family nucleus. already in 1987 such statistics have been compiled. from an organizational and financial point of view this enumeration had to be restricted to municipalities with an automated register^ moreover, analyses from the results with respect to the number of 'family units' (i.e. nuclear families and persons not living in a nuclear family) living at one addi^s have shown that by means of such a register-based enumeration it is also possible to compile statistics on households, provided that complementary statistical data bom other courses (e.g. survey research) are available. decisions on the organization of such additional registerbased enumerations have to be taken yet. these are dependent on the definite fwm the municipal reporting to the cbs on vital events and migration will take in the new system of population accounting. in consultation with the minisffy of home affairs several alternatives are discussed at the moment. finally, the implementation of the map will also offer possibilities for improving and enlarging the current demographic statistics. first, individual demographic events could be linked, so that statistics on life-cycles could be compiled. secondly, demographic statistics could be presented on the territorial subdivision of municipalities formerly used in censuses. large-scale surveys 3.1 introduction since the seventies two large-scale sample surveys have been conducted periodically by the cbs: fi'om 1975 the labour force survey (lfs) and from 1977 the housing demand survey (hds). both surveys aim at describing regularly the situation at a specific field of inquiry (the labour market, the households and the housing market respectively) as completely as possible, as well as monitoring the developments which are taking place at those fields over time. consequently, the statistical information provided by these surveys includes both data which can be used to provide some benchmarks on the relative field in question and so-called in-depth data on those fields. these data are obtained by grossing up the survey results to the level of the total population, the annually compiled demographic statistics being the basis. the principal characteristics of both surveys are described below. 3.2. the housing demand survey (hds) 3.2.1. topics:benchmark data and in-depth data since 1977 the housing demand survey (hds) is conducted every four years, partly at the request of the ministry of housing, planning and the environment the principal aim of the hds is to provide statistical information on the present housing situation of the population, the expenditures of the population for housing, realised residential moves in the two years following the survey date (everaers, 1987a). by means of this survey benchmark-data can be provided on (a) households, and (b) the dwelling stock and other housing units. benchmark-data on households concern characteristics such as their size, type and demographic spring 1990 composition. benchmark-data supplied with regard to the dwelling stock are type of dwelling, period of construction, number of rooms and type of ownership amongst others. in-depth data relate mainly to characteristics of households which arc of importance for the statistical description of the housing situation. in this respect special attention is paid to ownership or tenancy and the expenditures of the households on housing. other topics on which in-depth data are provided, are residential moves and potential households, i.e., persons in private households of 18 years and over and personnel living in institutional households who want to move to another dwelling or another housing unit. the benchmaik data are only available for regions with circa 100 000 inhabitants or more. fot the housing policies of the government this regional detail is sufficient, since mainly big cities and so-called housing market areas are the target areas in these policies. as already has been pointed out, in the near future more regional detail in the benchmark data on families (and perhaps on households) will be obtained by enumerations from the municipal population registers. furthermore, annual statistics are compiled on the size of the stock of dwellings at the level of the municipality. these statistics are based on a yearly updated statistical file set up at the 1971 census. in the coming years this file will be regauged by building up an automated register of dwellings by address. 3.2.2. sampling procedures at present two sampling frames are available: the decentralized municipal population registers and the geographic base file (gbf). the gbf a joint project of the postal service, the central bureau of statistics and the government physical planning service is an automated (and yearly updated) register containing all addresses in the netherlands including codes for the postal district, the grid square (500 by 500 meters) and the territorial subdivision of municipalities formerly used in censuses. theoretically the address together with its known occupants as sampling unit would be the best representation of the target populations of the hds (living quarters, private households and potential households). in this respect, however, both sampling frames have disadvantages. at present it is not possible to draw such a sample from the decentralized municipal population registers, due to organizational and budgetary problems. moreover, addresses with vacant dwellings are not included in the population registers. on the other hand, the gbf contains only an indication of the number of postal deliveries and type of building for each address. weighting the disadvantages of both sampling frames in connection with the target populations of the hds, the municipal population registers have been chosen as sampling frame and the person as sampling unit the disadvantage of having no information on vacant dwellings counterbalances strongly the disadvantages of having a very high underrepresentation of sub-tenant households (particularly one-person households). such an underrepresentation was discovered from analyses of results of the 1971 census with those of surveys carried out around 1971 with the address as sampung unit. however, using this sampling procedure the probability of being selected into the sample is not necessarily the same for all households and dwellings. this probability is twice as high for a dwelling occupied by a household with a spouse as the corresponding probability for a dwelling occupied by a household without a spouse. this is a consequence of the research-design: questions regarding occupied dwellings are only to be answered by the main occupant or his (married or unmarried) spouse; questions regaurding the composition of households only by the reference person or his spouse. therefore, corrections are made afterwards for a great deal of the survey-data on more-person households and dwellings (see the next section). under the given budget, the sample fraction (1:150) is chosen in such a way that statistics can be compiled with sufficient precision for the big cities and the housing market areas. the sample is drawn using a two-step method. in the first step a selection of municipalities is made by the cbs. the criteria used are the size of the intended sample fraction and requirements of the fieldwork. in the second step the selected municipalities draw a sample of persons aged 18 years or older from their population register according to written instructions by the cbs. until now these instructions have to be made separately for municipalities with automated and mechanized data processing systems as well as for those municipalities which still administer their register by hand. names and addresses of the selected persons (and some registered characteristics) are sent to gbs. in spite of these instructions there are always some problems in getting samples of sufficient quality from the different municipalities. some municipalities which have no automated system are unable to draw the sample. sometimes preselection of persons has been taken place in drawing the sample. many of these problems can only be solved adequately, when all municipal population registers will be automated. 3.2 j. method ofestimation the grossed up hds estimates are based on weighted observations. the respective weights are determined by using the method of post-stratification (i.e. stratification after selection of the sample). in this method the survey population is partitioned into a number of subpopulations, called strata, and all selected persons within a stratum are given the same weight per stratum the weight is calculated as the ratio of the size of the total population which is partitioned in the same way and the number of selected persons in the survey (bethlehem, 1987). by applying this method both the non-response bias can be reduced and the precision of the estimates at the level 24 lassist quarterly of the total pofxilation can be improved. it is known that these effects are only obtained if a relationship exists between the target variables of the survey and the variables used to construct the strata. the data available for constructing the strata satisfy this condition to a great extent. the respective weights are determined according to the following procedure (everaers, 1978b). first, weights are calculated for correcting the overand underrepresentation of population categories in various areas due to selectivity in non-response. the relevant strata for these areas are obtained using the information on a number of registered characteristics of all selected persons received in the sampling stage from the municipalities, such as sex, year of birth, marital status and family status. the selection of the areas is primarily based on the urban/ rural distinction. secondly, weights are calculated indicating for the various areas the number of persons the selected person represents. the partitioning in strata is based on the municipal information on all selected persons and the results of the demographic statistics on age and marital status. the areas used in this reweighting are the areas for which data of the hds are published. finally, the definitive weights are determined, first, by multiplying the above two weights and, then, by dividing the obtained results by two in those cases where the probability of being selected was twice as high (see section 3.2.2). 3.2.4. main results and their reliability the main results of the last hds are presented in table 2. their reliability can be checked by comparing these estimates with results from other statistics. such comparisons can be made with respect to the estimated figures of occupied dwellings and households. the estimate of occupied dwellings can be compared with the corresponding figure to be derived from two sources, namely: the already mentioned updated 1971 file on the stock of dwellings and the regularly published figures on vacant dwellings. it is found that the hds estimate significantly deviates from die last figure. however, it has been already noted that the updated 1971 file will be regauged. there are indications that the information on some of the changes in this stock sent monthly by municipalities to the cbs, is unreliable. hds estimates on hosueholds can be compared with similar information derived from the labour force surveys which up to 1985 have been held every second year. such a comparison shows a significant difference in the estimated number of one-person households between the two surveys. it is not clear yet which figure could be considered as more reliable. an indirect check on the reliability of the household estimates can be performed by comparing the hds population estimates calculated by means of the frequency distribution of the household-size with the corresponding figures from the population statistics. in table 3 the several figures are given for the total population and for age groups. looking at this figures one may conclude that the hds estimates on households on this point seem to be reliable. 3.3. the labour force survey (lfs) 33.1. topics: benchmark data and in-depth data from 1975 until 1985 the labour force survey (lfs) has been regularly conducted every two years; since january 1987 continuously. data collection and data processing in the continuous labour force survey (clfs) are completely automated (van bastelaer, 1987). from the beginning these surveys have been designed to provide statistical information on the labour force, educational attainment and qualifications of the population and commuting of the currently active population. the benchmaric data which can be supplied by the lfs refer to among others the main categories of engagement (such as economically active, educational training, engagement in household duties) and not-engagement (such as retirement or disablement); the size and sociodemographic composition of the labour force (including educational attainment) as well as some economic characteristics of the employed persons such as occupation, branch or economic activity, status in employment and place of work. due to the sample fraction (circa 2.5%) these data can only be published for the administrative areas of the regional labour exchange (64 regions) as the lowest level of regional detail. however, this regional level suffices for national policy purposes widi regard to the labour market. in-depth data are compiled for the employed labour force (for example on several aspects of the time worked as well as of retirement of working; secondary occupation), and for the unemployed labour force (for example on job seeking, registration at a regional employment exchange and social security benefits received). moreover, in due time fiow data can be provided, since the continuous labour force sample collects data on the labour history in the preceding year for the population of 15 years or older. this regards among others the dates employment started or ended in the previous twelve months; the main characteristics of the jobs performed during this period such as occupation and branch of economic activity; the reason of terminating a job as well as job seeking activities for every period of unemployment in the previous twelve months. as yet, this kind of data are collected by means of retrospective questions. at present plans are worked out to use a panel for it. finally, in the near future further in-depth data will be collected on various additional topics according to a rotating system. every year a specific tq)ic will be chosen on which monthly information will be collected at the moment plans are worked out for collecting data on not regular education and training next year. in the long run the new design of the labour force survey offers the possibility to provide a complementary spring 1990 set of stock and flow data on behalf of which better insights in the dynamics of the labour market can be obtained. for the present annual figures are published, whilst the compilation of three month moving averages is worked on. 3.32. sample procedures the geographic base file is used as sampling frame and, therefore, the address as sampling unit. organizational factors and budgetary reasons prevent the use of the municipal population registers as sampling frame, although the person as sampling unit fits the target population the best. at present households living at ca. 12 000 addresses are visited monthly. this number is halved in the holiday season. in view of requirements of the fieldwork a stratified multistage sample is used. the first stage consists of a monthly revolving sample of municipalities stratified in ca. 80 geographical areas. this stratification is applied in order to obtain reliable annual figures at the levels of the relevant territorial sub-divisions (i.e. areas of the regional labour exchange and areas covered by the regional subdivisions used by the european communities). the revolving system is applied to municipalities with less than ca. 20 000 inhabitants only and is chosen in such a way that the territorial distribution of the sample over the whole year is as adequate as possible. the municipalities with more than this number of inhabitants are drawn every month. in the second stage addresses in the selected municipalities are systematically selected. in drawing the sample a double selection probabiuty is given to addresses with more than one "jxjstal delivery". this procedure is applied in order to reduce evental cluster-effects, since households living at the same address are expected to resemble each other. therefore, at addresses with a single delivery all households are interviewed; at addresses with more than one postal delivery only half of the households are interviewed. addresses of institutional households are excluded. it should be noted that only 4% of the addresses are addresses with more than one postal delivery. the greater part of these addresses regards addresses at which more than one household is hving in a dwelling or another housing unit. for the lesser part it concerns addresses with two or more dwellings: not surprising, after all, since the municipalities are recommended to address each dwelling separately. 3.3j. estimation method the grossed up lfs estimates are likewise obtained by assigning weights to the observations using the method of post-stratification. in the biennial surveys roughly the same procedure has been applied as the one mentioned in the section on the hds. the calculation of weights for correcting non-response effects was based on information from the respondents and for the non-response on information from the municipalities sent to the cbs in connection with the fieldwwk which was carried out by municipal civil servants. for estimating annual figures from the clfs the weighting procedure used in the biennial survey had to be adjusted. in the adjusted procedure the continuous character of the survey had to be taken into account, in particular the halving of the number of observations in the holiday season. furthermore, the cbs does not have the relevant municipal information for correcting nonresponse effects at its disposal any more. therefore, the number of steps in the revised weighting procedure has been extended. the definitive weights used for grossing up the sample results are calculated as the product of five intermediary weights. first, a weight dependent on the monthly protebility of being included in the sample is given. the second and third intermediary weights are calculated for correcting non-response effects. the second for correcting seasonal differences in the non-response; the third for differences in the non-response by various population categories. in calculating the correction weights for nonresponse in step two and three the same territorial subdivision is used; the population categories are determined by a combination of the characteristics sex, age and nationality. the partitioning in strata for these areas is based both on the characteristics of respondents and the corresponding demographic statistics. the calculation of the weights in the fourth and fifth step is intended to get estimates which are representative for detailed populations categories (forth step) as well as for geographical areas on a detailed level (fifth step). in both calculations the same characteristics (sex, age and marital status) in determining the population categories are used, whilst one combination of those characteristics is reducible to the other. this principle of reducibility also apphes to the geographical subdivisions used in both calculations. the calculation of both weights takes place simultaneously by iteratively proportional fitting. tlie strata for the various areas are obtained by using information both from the respondents and the system of demographic statistics. the estimates are calculated as averages for the whole year. the averages on the level of the total population for the various areas necessary to perform the calculations are obtained by linear extrapolation of demographic figures on the first of january of the relevant year. the extrapolation is based on the demographic developments during the preceding year. the method of extrapolation is applied, since the annual results of the clfs ought to be published only a few weeks after the fieldwork in december has been finished. 3.3.4. reliability of results the reliability of the estimates from the labour force survey as far as they relate to persons in employment can be checked by comparing these estimates with data on employed persons which are regularly obtained from (partly integral) surveys among private enterprises and public services. it should be noted that the last-men26 lassist quarterly tioned data refer to jobs; the labour force survey, however, to persons having a job. moreover, in the lfs estimates data on the arm^ forces and persons employed in households are included; in the results of the establishment-based surveys they are not. taking these differences into account the main results of both kind of statistics did not significantly deviate from each other during the period 1975 to 1985. this situation changed at the introduction of the coninuous labour force survey. in comparison with the lfs 1985 the results of the clfs 1987 show a higher increase in the number of persons employed than could be expected from the increase over this jjeriod derived from the establishment-based statistics. this extraordinary increase is but exclusively concentrated under part-time workers with less than 20 hours worked a week. probably changes in the wording of the questions on employment and a better probing of the cbs-interviewers have led to these results. toward integrated population data or sociodemographic accounts 4.1. separate collection of various benchmark data and coherency in statistical iriformaiion on the population. the preceding sections have shown that demographic, social and socio-economic characteristics of the population are collected in connection with the statistical description of a specific field of research and policy. this proceeding has the advantage that coherent statistics can be provided on certain benchmark data and in-depth data on a distinct field simultaneously. in applying this procedure it turns out that for the greater part the data are not tuned to each other. when data obtained in one field are also collected (usually as background information) in another field, very often the relevant figures differ from each other. incoherencies in statistical information on subpopulations also exist between results from the above-mentioned large-scale sample surveys and data regularly collected from surveys among private enterprises or institutions of public services. some examples of the last kind of data are: the data on employed persons already mentioned in the last section, enrolment data obtained from educational establishments as well as data on persons in institutional households based on various surveys among e.g. health care institutions, homes for the aged and other social welfare institutions. some examples of such incoherencies are given in tables 4 and 5. the incoherencies are considered unsatisfactory by users of statistical information on the socio-demographic situation of the population. in order to meet the demand for more coherent information on this field the cbs recently started the compilation of socio-demographic accounts (koesoebjono, 1987). the underlying aim in compiling these accounts is to provide a coherent statistical description of the sociodemographic composition of the total population in a twofold way. firstly, on the level of stock data, reflecting size and structure of the population at a certain moment in time; secondly on the level of flow data, expressing changes in the size and structure of the population between two moments in time. achieving data coherency in these accounts necessarily imphes a process of adjustments in existing data and of additional estimates for lacking data. a coherent system of stock and flow data requires one and the same reference period, uniformity in concepts and operationalizauons as well as an identical target peculation, i.e. the total population of the counuy. in this respect it should be mentioned that the existing data (a) relate for the most part to different observations periods, (b) are often based on different operationalizations of concepts and sometimes even on conceptual differences, (c) show differences due to the appucation of sampling jh-ocedures (precision of sampling results, possible sampling errors) and of different estimation methods, and (d) refer to different population categories. 4.2. methodology of the integration: meanfeatures as yet the stock data in the accounts refer to the situation at the fu-st of january of each year; the flow data to the period between the first of january of two successive years. the accounts are presented in a matrix form: the stock data in the distributions of the marginal distributions relate, therefore, to the beginining, respectively the end of the period under review; the flow data to the transitions between the categories in these distributions. as a consequence, intermediate transitions are not taken into account. in compiling the accounts the basic principle in the population accounting is followed: that is, the size of the population at the beginning of a period plus the number of persons entering the populations in the course of the period equals the size of the population at the end of the period plus the number of persons who left the population in the course of the period. this rule is consequentiy applied for each category which has been distinguished in the matrix. the process of data integration occurs in various steps. first, the stock data (the marginal totals in the matiix) are compiled. for this purpose quantitative analyses are carried out with respect to the differences mentioned earlier in available data, and adjustments in data as well as minor additional estimates are made. in compiling the stock data the various figures are arranged in order of reliability. in all matrices the population figures are treated as the most reliable ones. second, the flow data (the cells in the maoix) are established analogously to the compilation of the stock data. however, in this step more estimates have to be made. not all data are available, and if they are, they are not directly related to the categories used in the stock data. in the next step stock and flow data are confronted with each other in the matrix. explanations for differences and contradictions between stock and flow data are sought for. thereafter, the relevant figures (on flow and even on stock data) are revised. spring 1990 27 finally, a procedure of iterative prqwrtional fitting is applied in order to obtain a matrix which is internally consistent (that is: the basic accounting principle is valid for each category of the matrix) and which deviates as less as possible from the original matrix (established after the third step). during this process, the stock data are assumed to be fixed, only the flow data change. 43. matrix construction and main results at present two kinds of socio-demographic accounts have been compiled. one consists of coherent statistical data on the population with reference to type of engagement or non-engagement, the other one with reference to its status in household. the basic matrices for men and women separately are very detailed since they contain the data by engagement (status in the household respectively) and age group. the data on sex and age composition of the total population is the framework whereupon the data on type of engagement or not-engagement, and the data on household status are gauged. therefore, the first step in the matrix construction consists of the compilation of the demographic data matrix by sex and age. table 6 shows an aggregation of the demographic matrix for the year 1984. following the compilation of that matrix, the definitive matrix can be compiled step by step fw each demographic population category (by sex and age group). the final residt is a matrix by age and type of engagement, respectively household status for men and women separately. an aggregation of the first mentioned data matrix is given in table 7. both matrices reflects the composition of the population at two successive moments in time, as well as changes which take place between these two moments. consequently, the destination of persons belonging to a certain category can be traced at the end of the period. furthermore, the origin of persons belonging to a certain category at the end of a period can be derived. in this way the respective flows the outflow and inflow of each category can easily be calculated. this also applies to a calculation of the turnover flow for each category, that is the numbers flowing into and out of a certain category. from this point of view the data in the respective matrices can serve as a basis for projections, as among others the destination percentages can, with due reserve, be considered as probabilities of transitions. the availability of a series of such figures over time allows to formulate hypotheses about the future developments with respect to processes of change, and in connexion with this, a projection of the various categories. 4.4. some prospects at present integrated stock data with regard to type of engagement or non-engagement are abeady available for five successive years (1980-1985); integrated flow data for three years. further developments are directed towards (a) the extension with other categories of type of engagement (e.g. engagement in household duties or in voluntary work) and relevant categories of non-engagement (e.g. retirement, disablement); (b) the construction of such matrices for specific population categories, e.g. the alien peculation; and (c) the compilation of quarterly accounts in connexion with the development in compiling stock and flow statistics on employment and nonemployment based on the clfs. the matrices with regard to household status will soon be available provisionally for only one year (1985), due to the lack of relevant data at present. in particular annually compiled statistics on households analogue to the population statistics are missing. therefore, work is underway to compile such statistics for the short term using different sources, especially the demographic statistics and the clfs. on the long-term it is expected that household statistics can be compiled regularly by register-based enumerations on family status together with survey data. finally, studies are progressing with regard to the presentation of transition within the population in order to have a better insight on its mobility. concluding remarks in the preceding sections a broad outline of the statistical system in the netherlands has been given as far as this system contains elements in reference to the general topic of post-censal surveys. it has been pointed out that various research and statistical techniques are jointly applied to obtain the relevant demographic, social and economic data on the total population. the instruments for compiling the basic demographic statistics, i.e. the system of population accounting (including the municipal rejxmting to the cbs) and enumerations from the municipal population registers have been described. next to this special attention has been given to some methodological aspects regarding the large-scale surveys as being the relevant sources for providing data on the social and economic situation of the population. finally, a description has been presented on the first efforts to generate coherent statistical information on the level of the total population by means of the development of socio-demographic acounts. this broad outline is summarized in figure 3. the system of population statistics, the mean features of which have been presented above, may be considered as an alternative statistical programme to a population census and post-censal surveys. however, it should be emphasized that this system is not a substitution thereof in the sense that it aims at obtaining exactly the same statistical information. on the contrary, it is to be considered as a procedure of bringing up-to-date the formerly used instruments within existing possibilities and hmits posed by (the dutch) society. within these possibilities and limits, the systems aims at producing the statistical information users in general are looking for. references bastelaer van, a.m.l., 1987, the continuous labour force survey. netherlands official statistics, vol. 2, no. 4, pp. 30-32. 28 assist quarterly bethlehem, j.g., 1987, weighting sample survey data. netherlands official statistics, vol. 2, no. 1, pp. 17-18. brekel van den, j.c, 1977, the use of the netherlands system of continuous population accounting for the population statistics, netherlands central bureau of statistics, voorburg/heerlen, the netherlands. choldin, h.m., 1987, statisticians' responses to the privacy issue. paper presented at the meeting of the american statistical association, chicago. everarers, p., 1987a, the housing demand survey 19851 1986. netherlands official statistics, vol. 2, no. 4, pp. 42-46. everarers, p., 1987b, the pitfalls of secondary data analysis with special reference to the dutch housing demand survey. paper presented at the fifth european colloquium of theoretical and quantitative geography. bardonnechia (italy). kosoebjono, s., 1987, socio-demographic accounts: framework to measure the dynamics ofpopulation. netherlands central bureau of statistics, voorburg/ heerlen, the netherlands. redfem, p., 1986, which countries willfollow the scandinavian lead in taking a register-based census of population? journal of official statistics, vol. 2, no. 4, pp. 415-424. redfem, p., 1987, a study of thefuture of the census of population: alternative approaches. statistical office of the european communities, luxembourg. verhoef, r., and van de kaa, 1987, population registers and population statistics. population index, 53 (4), pp. 633-642. vliegen, m., and h. van de stadt, 1988. is a census still necessary? experiences and alternatives. netherlands official statistics, vol. 3, no. 4, pp. 27-34. ' presented at the ifdo/iassist 89 conference held in jerusalem, israel, may 15-18, 1989. the author expresses his acknowledgements to santo koesoebjono for his valuable comments on an earlier draft of this paper. ^ notwithstanding this restriction (nearly 75% of the total population has been enumerated), results have been presented for every municipality by generalizing the results obtained from the automated municipalities to the non-automated ones using their respective composition of the population by age, sex and marital status as a base. given their geographical position and degree of urbanization, the family composition within the automated municipality was supposed to be equal to the family composition within the similar non-automaled municipality. spnng 1990 29 figure 1 : from reporting of the population to statistics on the population population birth i death i marriage i dissolution of marriage municipal registrar civil registration -request forcltlzenshlp ministry of justice vital events change of residence change of naclonallcy municipal population register migration cbs population statistics assist quarterly figure 2 synopsis of (planned) entiaeretlons from the ntunlclpal registers for revision purposes, 1971-1990 1971 total population by sex, year of birth and marital status (yearly updated) 1976 : alien population by sex, year of birth, marital status, and country of nationality (yearly updated) 1983 : total population by sex, year of birth, marital status, and country of nationality (yearly updated) 1990 : total population by sex, year of birth, marital status, country of nationality and country of birth spnng 1990 figure 3. from rescrlcted to in-depch information on th« population municipal population registers (incl. population accounting) population statistics (poststratification) cbs geographic base file (gbf) (sampling frames) survey i in-depth information on e.g. labour (clfs»') survey in-depht information on e.g. education survey in-depth information on e.g. household (clfsi>) (hds2>/clfsi') survey in-depth information on e.g. housing (hds^)) (integration) little in-depth information on restricted number of variables stock data only (other sources) coherent in-depth information on various number of variables stock and flow data *' continuous labour force survey ^' housing demand survey lassist quarterly t*bl« 1. dlffai*nc*a bctmcan th* r**ult« of th« r*tiit*r»d-bu*d (numratlon 1ss3 wid th* upd4t«d 1b71 and ia7t tlla aa a p«rcadta«a of the rslavant catafoiiaa froa tha updated filaa ac* toib 20 48 so s4 ts yaara tal yaara yaara yaara or oldar i mala -0.0 0,0 0,1 -0.2 faoala 0.0 0.0 0,1 -0.2 total 0.0 0.0 0,1 -0,2 marital itatua navar barrlad ifldowad dlvorcad barrlad z mala -0.8 1,0 -1.2 -2.8 faaala -0.3 0,3 -0,2 -o.i total -0,6 o,e -0.4 -1,* hatlona llty dutch alian z mala -0.1 l.s famala -0.0 0,0 total -0,1 1,3 tabla 2. population. houaaholds and houalnf aituatlon. bos 1s8s/1888 to~ occuplad othar inllvlnt atltutal doalllnta quartara ttona x 1000 uouaaholda ona-paraon houaaholda 1 s30.7 1 280.8 167,4 nultl-paraon houaaholda 4 034. s 3 003.6 28.2 total 3 383.2 i 283.4 103.8 numbar of paraona 14 402,1 13 086,6 234,2 231.3 1) 1) paraona of 18 yaara and oldar spring 1990 tabl* 3. populttlon •itlbsttt by m*, bds itts/lsst uid itmagctfkilc •tatlitio by •(•, 1866 toa«* < ij 13-28 30-48 so-64 6} yaari t*l years y*»xi y**ra yaars or oldar dacdographlc alatlatlcs 14328,4 2786,2 372s,0 4107,1 2140,0 1768,2 bos: population in -prlvata houaaholda 14240,6 2621,1 3637.0 4032,4 2127,9 1362,3 -initltut.houaaholda 1) 231,3 13,0 20,1 13,7 204,4 total 14482,1 2621,1 3670,0 4072,3 2141,6 1786,8 dlffaranca with ragard i to tha daoiosr. atatlat. 0,3 + 1.2 1.3 0,8 + 0.1 + 1,0 1) population 16 yaara and oldar tabla 4. population in prlvata houaaholda by atatua in houaahold, lfs 1863 and hos 1863/1866 onahultl-paraon houaahold paraon houaahold rafaranca apouaa child othar paraon 2) paraon z 1 000 1 434 4 008 3 368 2 016 202 1 331 4 033 3 622 4 633 188 1) population of 13 yaara and oldar 2) inel, living in conaanaual union lassist quarterly tabi* s. populttion of is ytari u>d oldtr in full-tin* adueatlon by »»x wid m*. es 1864 (••pt«ad>*i) 1) and lfs istj (aprll) 2) to*«• tal 15 24 25 yaara yaars or oldar s 1 000 es isa* (••ptaobsr) hals e72 621 51 fasala i«3 513 26 total 1 21j 1 136 70 lfs 1985 (aprll) mala 632 306 34 famala 371 32t 47 total 1 203 1 122 61 1) educational stattatlca 2) labour forca survay tabla 6. total population by a«a, data natrlz 10e4/'63 of which on 1-1-1985 in population not in population 1-11984 total 0-14 yaara 13-64 yaara 63 yaara aolor oldar daath ^ration stock on 1-1-1b83 14 434 2 850 9 673 1 730 of which on 1-1 -1984 in population total 14 395 14 222 2 661 9 832 1 726 0-14 yoart 2 930 2 914 2 661 252 15-64 yaara 9 756 9 691 9 380 112 65 yaars or oldar 1 708 1 617 1 617 not in population birth 173 173 iamlgration 60 16 42 1 spnng 1990 table 7. total population by type of engasement/non-engataoiant, data matrix 198'i/'e5 (provisional figures) stock of which on 1-1-1985 on in population not in population stock on 1-1-1965 of which on l-l-igs* in population 1-1toprefull time full-time othereml1961) tal school education eotployment wise death grstion 14 395 u 222 pre-school ft. education f.t. employment otherwise 709 70a 3 421 3 405 4 582 4 547 5 582 5 565 180 2 1 4 3 162 138 104 1 15 5 4 273 268 15 21 11 215 5 339 102 14 not in population birth lomlgration 173 60 173 5 lassist quarterly vol29-2.indd iassist quarterly summer 2005 by by katrina stierholz * economic data as snapshots in time the federal reserve bank of st. louis has initiated two new projects, fraser and alfred. these two projects share one goal: to provide better access to historical economic data. but the projects differ in their audiences and uses. fraser is an image archive of economic statistical publications, from government or near-government sources (aka the fed). alfred is a machinereadable archive that allows researchers to pull real-time data in handy formats such as excel. these two projects build on fred. fred, which stands for federal reserve economic data, is our most heavily-used data product. fred offers over 3,000 economic and financial time series drawn from government and commercial sources. reason for project several years ago, bob rasche, the research director at the st. louis fed, began searching for economic data that would tell him exactly what economists knew at a particular point in time. he wanted to be able to create a database that would give him answers to historical realtime data questions. researchers at the philadelphia fed had also been looking at this issue, and they had created the real-time data set for macroeconomists, the first realtime datasets for researchers to use. this dataset is useful, but the number of variables is relatively small. the st. louis fed seeks to build on that work. real-time data represents a point in time, either the moment we are in or sometime in the past. the st. louis fed’s fred database is real-time data, but only for the current moment. and while the data in fred is historical, it does not contain the actual numbers that economists saw in the past—it contains the data that has been revised for the historical time period. that’s because many economic time-series are revised (and revised again and again). in order to see the data at a particular point in time, as it was seen by contemporaries of that time, the researcher needs some way of accessing that original data. unfortunately, those data are much more difficult to obtain. economists have seen the need for this data, to answer questions about economic research and economic policy. 1. replicating other economists’ work fraser and alfred will offer the ability to reproduce other economists’ work. while researchers in other fields have been documenting their studies and offering other researchers the chance to replicate their work, economists have been less likely to replicate other economists’ work. a significant reason is the difficulty in replicating both the program and the dataset. when economists publish papers, they often do not publish the datasets and the program code used to generate their results. without this information, it is virtually impossible for other economists to replicate their work and check for errors. [see dewald 1986.] the st. louis fed became a leader in replication when dewald become research director in 1992. since that time, for our publication, the review, we require that the data that supports a paper be published (on the web) alongside the article. 2. sensitivity of results to vintages of data another possible use for this historical real-time data is checking the results of studies using a variety of vintages of the same series, all of which contain the preliminary data. an economist could check a model’s usefulness by using the real-time data and have the final data to use as a check. it is also useful to examine what economists knew at the time, to see if their policy decisions made sense. that is, based on what we knew then, would this model predict the way things turned out or does this policy decision make sense? [croushore’s august 2004 manuscript details this nicely.] only real-time data can answer that question. 3. expectations and policy making with initial data release yet another reason for using real-time data involves issues that surround expectations. the information that economists and other policy makers have at the time that decisions are made may change. sometimes once, sometimes many times. policy decisions may be made on these “not great, but all we have” numbers – and so it is important to preserve the historical information used for those decisions. looking back at decisions made, it is important to be aware of what the data at the time said, not what they say now. evaluating the changes in these revisions will also help economists evaluate/properly weight the reliability of the first run numbers. 6 iassist quarterly summer 2005 as an example, consider the numbers published for gdp. gdp is measured quarterly. new measurements, for a variety of quarters, are published each month. most hotly anticipated, of course, is the data for the most recent quarter. the first release, or measurement, for that quarter is the “advance” release. for the 1st quarter of 2004, this number was released on april 29, 2004. the second published number is known as the “preliminary” estimate — and it was released on may 27, 2004. the third published number is known as “final”, and was released on june 25, 2004. unfortunately, the “final” value isn’t final! a revised number often is published the following month (although it lacks a name) and thereafter revisions are published annually. eventually, the numbers are revised further roughly every five years. [see croushore 2004.] in the process of gathering release dates for his own research, bob rasche saw the value of this information for other economists. he wanted to build a database that would include several data series, along with the information from each revision, and that would have the dates that the information was good for: think of it as an expiration date for the data. he was very interested in knowing both the data and the date those data were released. rasche decided to make it available to the rest of the world; first as the publication (a scanned document via fraser), and then as a database with data points and release dates (alfred). so, for these reasons, rasche decided to go with a twopronged approach to preserving this information and making it widely available. alfred and fraser have been conceived as a pair, but they are very different data products, and will, in the end, serve different audiences. both build on the original fred database concept. fred the federal reserve bank of st. louis has fred (federal reserve economic data), which offers over 3,000 economic time series, including banking, financial, employment, monetary, interest rate data, and more. the data is collected from a variety of sources — the board of governors and the us government, as well as some commercial sources. it is presented in useful ascii or excel formats, and the data can be downloaded in large data sets or as a single item. fraser fraser — our internet image library — is a natural outgrowth of gathering all of that lost information from the press releases. as we discovered how difficult it was to find press releases, it became clear that information was being lost. and, if it hadn’t disappeared yet, it would soon, as libraries hurry to clean off their shelves. the easiest method for providing this information is to scan serial economic publications that have a variety of economic data. so we began scanning all those press releases for various economic time series. we then spent weeks and months hunting down all the gaps in our information, taking documents from every librarian that would let us beg and borrow their documents. fraser, as an image archive, includes tools that allow for sophisticated retrieval of statistical tables linked over many years of publication. we started with the monthly economic indicators, a publication of the joint economic committee, and we added some federal reserve and federal government titles to the mix—things like banking and monetary statistics, all bank statistics, and business statistics. these are basic titles, with lots of economic data. much of this data will not be available on alfred anytime soon; there’s just too much to enter, and it would require so much work. but the data are available for anyone who wants it. fraser as part of the gpo digitization project fraser will give historians access to what policy-makers and economists (and the public, for that matter) knew at the time. it will also be a part of the national bibliography of government publications that the gpo is coordinating in our natural niche of economic statistical publications. we have scanned the material in the manner requested by gpo. fraser fulfills the need for wide access to historical economic data, at a relatively low cost. we have added more publications to our list of scanned periodicals: the business conditions digest, survey of current business, the economic report of the president, and the h.6 money stock measures (a release from the board of governors); these will be posted soon. to make the data in these publications more accessible, we’ve done a couple of things. one, we’ve added keywords to the metadata so that searching is useful. another is that we’ve linked all of the tables within a publication, so that if you find a table that answers your research need, you can download the same table over time. we’ve ocr’d the documents, so the text is searchable for terms ( for example, “gold”). we have added title continuation information and sudoc numbers to help users. and, finally, we’ve made it all searchable via the autonomy search engine. this combination gives users all the access that i can imagine they need, except for one thing. this combination gives users all the access that i can imagine they need, except for one thing. while we have ocr’d everything, including the data, we do not allow users to extract or copy the text or numbers, because the ocr has not been verified. users can save the pdf, or print it out, but the ocr has been locked so that it isn’t accessible. not that we want to deny users this but, because we haven’t corrected the ocr, we’re concerned about naïve users who might copy the information into a spreadsheet and not realize that some numbers are incorrect. it is particularly difficult to spot errors in uncorrected ocr tables. correcting the ocr is time iassist quarterly summer 2005 7 consuming and tedious (which translates to expensive). if the use of the material is high enough to warrant ocr correction, we’ll put it into alfred, because alfred is our database of real-time information that has been verified. the user will have more access. the user will have more access, in a better format, than if we leave it in fraser. any really intense user can easily ocr the pdf files themselves, using a variety of readily available commercial packages — in which case the user is responsible for all errors, not us. fraser has been built on the jstor model in many ways. we have scanned each publication at 600dpi in tiff format, and then digitally “cleaned” it of non-printed marks (such as handwritten notations, creases, or stapler marks) or imperfections created during the scanning process. the clean file is then run through an optical character recognition (ocr) software application that images the file, introduces metadata, and converts the file to adobe portable document format (pdf), as well as compressing the scanned image from 600 dpi to 300 dpi. we have complied with the united states government printing office requirements for scanning, because if and when the day comes that we need to migrate this material (should pdf become an obsolete format), we want to be in a very common format, so that there is a common solution. if we did something fabulous but unique, we might not have good options for migrating the information. fraser is a low-cost way of providing a large amount of statistical data, and it allows for uses of the data in ways that we have not imagined. fraser is useful to a wide variety of audiences—not just economists but historians who want to see a contemporaneous look at economic data. fraser recognizes that the largest cost in many research projects is locating the published data; entering the data into the computer is the easier part. the second project we are undertaking is called alfred. it builds an archival feature onto our foundational fred database, but with less data than is available on fraser. alfred (archival fred) data alfred will provide sophisticated data for economic researchers in a much more useable format than fraser. the data will all be available in text and excel spreadsheets, and will allow the user to pull multiple versions of the same data set, using different points in time as a reference. alfred will be populated initially with the archived fred data. at first, the data will go back to december of 1996, as we have taken a snapshot of our data every friday of each month since 1996. it is this data, along with the information compiled about the release dates, that will populate alfred. the release dates for economic data have become standardized over the years, and holiday information is easier to get. so verifying release dates for the fairly recent data has been relatively straightforward. the first release of alfred will contain the data from these snapshots and the release dates. because the data are being verified, the first release will contain the information for employment, cpi, and ppi; not for every single series in fred. a user will be able to download the data for any one of these series for many different dates. for example, she could download the employment numbers available at several different points in time. so, if she wanted to know what employment numbers were available before every fomc (federal open market committee) meeting, where they target the federal funds rate, she could download the employment numbers that were available each time they met... and see exactly what those policymakers saw when they were making policy. not the revised numbers that came out later, but the information they actually had. bob rasche has also collected the release dates and data for 25 data series going back before december 1996. these have been entered by hand (by an intern, who probably regrets ever taking that position). these numbers were then verified by a research analyst at the bank. these will also be available in the next release of alfred. later, more historical data will be included, and we plan to provide links between fraser and alfred, so that if alfred doesn’t have the information in its database, the information can be retrieved from a document stored on fraser. our initial release of alfred to customers on the internet will occur in july 2005 as a feature within fred. alfred will work in two steps. first, a user locates a needed data series in fred; then, second, s/he has the option of getting the most recent data and/or additional historical data. the user can select the range of desired historical data. to our customers, alfred will look very much like fred, but with additional available data. internally, however, the fred database has been reconstructed in an entirely new way, in order to make this project scalable. the database was created by george essig, a senior web developer at the st. louis fed, and he devised an ingenious way to store the data so it wouldn’t take up a ton of server space, and so that it would contain all the information necessary for researchers. george has used a concept found in richard snodgrass’ publication developing time-oriented database applications in sql (2000) that suggests a sophisticated way to store the data. in addition to the number that represents the observation for that moment, each data point also has two separate, additional pieces of information about it. the first is a measurement of the time interval that the data covers (for instance, if the data covers march of 2005, it would have 8 iassist quarterly summer 2005 a time interval measurement of 3/1/2005 to 3/31/2005). then, it has a second bit of information: the dates for which this information was valid. so for the first release of the march 2005 employment number, the period of validity would be from april 1, 2005 (the date of the first release) until may 6, 2005 (the date of the second release). it would have an open-ended date until the data was updated, so if it is never updated, the number stays valid forever. however, if the data changes, the program caps off the end date of the expired number and adds the new number. this feature keeps the back-end of the program relatively simple, and the total file size relatively small. for more details on what george has done to create the database, and the time interval measurement, validity intervals, and transaction intervals, please read a paper written by one of our economists, richard anderson: “replicability, real-time data, and the science of economic research”, march 2005. dick anderson has several papers on this topic. fred and alfred are not two separate databases, but one database. absent historical data, fred would have 1.1 m data points; in the new combined fred+alfred, the database has 2.2 m data points. alfred will start small, and live within fred for a while. the first version will allow users to locate a dataset, and then choose the date they want. we anticipate that alfred will serve economic researchers who are looking at modeling their work using real-time data from a variety of vintages, and also economists who are interested in replicating other economists’ work. fred is used by a large number of people with a wide variety of purposes. by comparison, alfred will reach a small audience with a narrow focus. also, while alfred will have excellent retrieval methods, useful for gathering multiple vintages of a series in a single shot, creating the data sets in alfred is a labor-intensive and time-intensive project. this will limit both the number of data sets that will go back before december 1996, and the time period that they will go back to. there may also be some series for which we are unable to gather the information before a certain date. library/librarian issues there are also some important library issues that arose during the work on this project. it was sometimes very difficult to find the release date for data. virtually every library threw out the oldest press releases. in many cases, we found only a few libraries that owned the oldest press releases. even the issuing agency and library of congress did not have these documents. as well as losing the print press releases, or other “current news” that was subsequently revised, the electronic versions were also often lost. because disk space was scarce, files were replaced rather than versions added, a loss that is permanent. while we would sometimes find the paper press releases in libraries as a result of benign neglect, in the case of the electronic files there might only be the most recent copy (if that), as all the others had been overwritten. as data librarians, consider the potential use of snapshots of your data, particularly if your data are subject to revision. i am sure that there are other fields where these issues might apply, and even if you aren’t sure if it will apply, consider taking snapshots if your data changes over time. for many researchers, it important to analyze data as it was available at the time, rather than the perfected data released much later. important policy decisions are made with this imperfect data. it is important to see the data using the same imperfect lens as economists had at the time it was released; and determining the problems of that data as well as its usefulness is key to developing good models. this data may also help economists create better models, or provide information on the usefulness of other economists’ models. we hope that by providing fraser and alfred, the st. louis fed will be able to contribute to that work by economists. references anderson, richard g. “replicability, real-time data, and the science of economic research” manuscript, federal reserve bank of st. louis, march 2005. (forthcoming, federal reserve bank of st. louis review.) anderson, richard g., william h. greene, bruce d. mccullough, and h. d. vinod, “the role of data & program code archives in the future of economic research” federal reserve bank of st. louis, working paper 2005-14. croushore, dean. “forecasting with real-time macroeconomic data” manuscript, university of richmond, august 2004. croushore, dean, and tom stark. “a real-time data set for macroeconomists” journal of econometrics 105 (november 2001), pp.111-130. croushore, dean, and tom stark. “a funny thing happened on the way to the data bank: a real-time data set for macroeconomists” federal reserve bank of philadelphia, business review (sept./oct. 2000), pp. 15-27. dewald, william g., jerry thursby, and richard anderson. “replication and scientific standards in empirical economics: evidence from the jmcb project” american economic review, september 1986, 76:4, pp. 1255-57. iassist quarterly summer 2005 9 rasche, robert, katrina stierholz, robert suriano, and julie knoll. “as it happened: economic data and publications as snapshots in time” presentation at the federal depository library conference, october 19, 2004, washington, d.c. snodgrass, richard t. developing time-oriented database applications in sql. morgan kaufman, san francisco. 2000. endnotes 1 contact: katrina stierholz, federal reserve bank of st. louis, po box 442, st louis mo 63166 314-444-8552, katrina.l.stierholz@stls.frb.org the author is grateful for the comments of robert rasche, richard anderson, and george essig. the views expressed are those of the author and do not necessarily reflect the official positions of the federal reserve bank of st. louis, the federal reserve system, or the board of governors. * the paper was presented at the iassist 2005 conference in edinburgh. iassist quarterly 39 tabulations on the dda study description by k.arsten boye rasmussen 1 danish data archives odense university the object of this paper is to give a brief introduction to the standard study description and to add a few remarks on its recent history and development. for those already familiar with it, the first section will also answer the question: 'whatever became of the access project?' the main part of the paper concentrates on a presentation of the holdings at the danish data archives (dda) in the form of crosstabulations based on a data file compiled from the contents of the study descriptions at the dda. the study description for those not familiar with the term 'standard study description', allow me to clarify: the phrase is meant to be read backwards. first of all, the standard study description is a 'description'. it is a machine-readable document written in a specific format it is like a library catalogue card, except that the object is not a book but a machine-readable data file, or, to take another step backwards, a study. the data file or study is, for the purposes of this paper, within the broad area of the social sciences. the format of the standard study description (ssd) is somewhat complex. the ssd contains a large number of items, which can be compared to variables. each item consists of a numeric identification code and an entrycontaining specific information. what makes the format complex is that the different entries may contain different types of information. at the ifdo/iassist conference in grenoble in 1981, i presented a working paper which extensively expounded the format of the ssd. in this paper, a few examples should suffice to illustrate the complexity of the ssd. for example: 1. • item 101 contains the title of the study. it is an unstructured text item. 2. • item 212:04 contains the unweighted number of cases in the data file. a numeric subitem (:04) inside a structured item. 3. • item 222 describes the target population using predefined codes. a precoded item. 'paper presented at the 1986 iassist conference, santa monica, calif.. mav 22-25, 1986. : karsten boye rasmussen: "proposed standard studv description". working paper presented at the ifdo/iassist conference, grenoble. 1981. summer 1987 40 lassist quarterly thus the study description contains text items, numeric items and precoded items; of these, the precoded item is the most complex. in precoded item 222, an entry "01" will signify that the target population is restricted by "age limits". there may be further restrictions specified by further codes. when a precoded item is viewed as a variable, what we have is indeed complicated (like a multipunch column on an old punch card). to complicate things further, the ssd format allows a text explanation of unlimited size to follow any type of item, whether it be a text, numeric, or precoded item. the ssd is like a data file: it is incomprehensible without the proper documentation describing the significance of the item numbers as well as a 'codebook' describing the codes used in the ssd. this documentation is, of course, machine-readable as well. with a computer program, one can merge the ssd data file and the codebook information to produce a human-readable printout describing the study. working backwards we have now reached the last word in the term 'standard study description'. the ssd has not yet been accepted as an international standard. given that this year is the 12th anniversary of the ssd, i shall not insist on the word "standard". rather, in the remainder of this paper. i shall refer to the 'study description' or the sd. on the other hand, it is a standard of sorts, in so far that some of the european social science data archives use the sd extensively, with only minor differences. those interested in the true story of the standard description are advised to read the paper "standard study description as a meta research data base" 3 given by per nielsen at the 1983 iassist conference. that paper outlines the development of the ssd and includes references to historical papers on the subject. this paper is an updated version of per nielsen's paper. the function of the sd from the beginning, the object of the sd has been to fulfil several functions as a tool for the data archives and the social science community. some of the functions have changed, primarily because new technology has facilitated the achievement of new goals. 1. data abstracting and catalogue production the main function of the sd is (as described above) to produce human-readable printouts of the study descriptions. this function is what was originally" termed "data abstracting and catalogue production". since 1978, the dda catalogues have been produced using computer programs to generate phototype setting instrucuons from the sds and a 'skeleton file'. at the same time, a considerable amount of indexing of the descriptions is done automatically. the same procedure has been used at the zentralarchiv in cologne, at the steinmetzarchief in amsterdam, and is presently being used at the esrc data archive in essex. 5 (it should be noted that other data organizations use similar techniques, they are 'per nielsen: "standard study description as a meta research data base". paper presented at -'(cont'd) the iassist conference, philadephia, 1983. (reprinted in dda-nvt ni. 26, 1983, pp. 5-23) "per nielsen: "study description guide and scheme". copenhagen, dda, 1975. !esrc data archive bulletin, january 1987, no. 33, p.l. summer 1987 lassist quarterly 41 not mentioned here because they do not use the sd as the basis for their catalogue production.) at the dda we are now completing the 1985 catalogue of holdings. only a subset of items from the study descriptions is being printed, but a number of comfiche containing the complete descriptions will be supplied with each copy of the catalogue. the production of catalogues has many drawbacks. first of all, it is very expensive. of course all the sds are available, as they are produced as part of the documentation process when data are deposited in the archive. nonetheless, a lot of proof reading is necessary before the catalogue is ready for printing. another drawback is the problem that the catalogue is outdated before it even reaches the market thus, online access to a computerized catalogue is necessary for up-to-date information. 2. mapping and methodological research base the collection of sds has long been regarded as the perfect object for methodological research or presentation of holdings. but the perfection resides in the very detailed information they contain, and not the process of compiling the information from the sd format to a rectangular data matrix ready for analysis. because of the complex format of the sds, such compilation requires some computing, to produce a rectangular numeric file without text information. this paper includes a chapter presenting the holdings al dda with tables based on such a rectangular file. all the programming, including the extraction of information, was done with the software package sas. 3. data (re)analysis prerequisite at the dda, the documentation of studies deposited in the archive includes a machine-readable codebook with complete questionnaire text as well as coding instructions plus one-way distributions of each coded variable. the codebook describes the data at the variable level. it also includes the total study description describing the background, objectives and outcome (publications) of the data collection. thus the sd is a prerequisite for the process of secondary analysis. 4. intraarchival loggin this function, which was to supply the archives with a tool for keeping track of the processing of their data, has completely lost significance in the technological race of the last decade. the development of interactive data bases with immediate updating facilities, as opposed to a sequential and both timeand costconsuming method, has. at least at the dda, led the archive to implement data base applications which have the power to keep the most important information ready at hand ("just a pf -key away"). 5. interarchival exchange however the sd is still a standard for inter-archival exchange. it is an exchange format which is simple enough to be read into any machine (including microcomputers). special software is needed, however, to process the sds. summer 1987 42 tassist quarterly it is therefore my impression that the sd will remain a standard format for exchange purposes, and that its list of items will be a check-list of the kinds of information which should be provided for each study. the actual storage mode of the sds at individual archives may differ, depending on what data base facilities are available, but the data base application should be able to both 'export' and 'import' sds in the standard format information retrieval data base the archives in cologne, amsterdam and odense have developed retrieval systems for searching their own sds. the most comprehensive system is the zar system at za in cologne, which has been described earlier in iassist surroundings. the data archive at essex has recently announced 6 that they are setting up an online information retrieval system. other archives have information retrieval systems as well, but the archives mentioned above are all using the sds. the access project the idea of using the sd-format as an exchange format as well as the possibility of setting up retrieval systems on the basis of the sds, led to a project amongst the european archives within cessda (committee of european social science data archives). under the projeci heading "access: integrated european archive inventory" a catalogue was to be published, sds to be exchanged, and a data base retrieval system to be set up with common 6esrc data archive bulletin, january 1986, no. 33, p. 1. access (available through euronet). the eec was to finance the project, but the demands of the bureaucrats in brussels made the project much less attractive, and it was finally abandoned. the project has instead developed into an ongoing effort at the four archives (mainly the dda). but due to a lack of funding, this project competes with regular activities of the archives and has therefore often been postponed. the tables in this paper are based on the collections of a single archive, the dda. it is my hope that within a reasonable time period i shall be able to present similar tables comparing the holdings of the four european archives. at present the dda has received sds from the steinmetzarchief; with the completion of the esrc data archive catalogue, the dda will receive a new batch of sds, and finally the sds from za. setting up a retrieval data base as described in the access project will demand a considerable amount of work. at present, network facilities are still not sufficiently effective to supply online access to other computers. when these techniques have been improved, the online integrated european archive inventory will become a reality. as mentioned above, the four archives use slightly different formats for the sd. as long as the differences are fully documented, this presents only minor problems. the za and dda formats are very similar, although the za does not use as many items as the dda. at the steinmetzarchief. the numbering of the items is different, but the mapping of the formats is the same. at the esrc-da, depositors, etc. are identified by a number from a special file. at the dda. the latest change in the sd format has introduced an item (220) pertaining to historical data materials, which shows the time period covered by the data, and an item (225) to show to what regional area or countrv the data describe. summer 1987 iassist quarterly 43 tabulations on the dda sds the tables in the remainder of this paper describe 962 studies. these tables will not be extensively commented upon nor compared with the findings of per nielsen in 1983' as the purpose of this section is not primarily to comment on the development of the data holdings of the dda since the 645 datasets were investigated in 1983. the tables refer to the datasets described in the forthcoming dda catalogue and give an overview of the contents of that catalogue. there has been no exclusion of datasets with missing data. for some items, missing data should not exist, but for others the existence of missing data is perfectly alright. some of the variables are coded with multiple codes. to illustrate this, some of the tables have information on the number of codes. if, in a precoded item, a study has more than one code, the entries are weighted accordingly (e.g. 2 codes, weight=.50, 3 codes, weight=.33 etc.). therefore even in these multiple response items, all the tables still total 962 studies (apart from rounding error). contents of the sd data bank the sds at the dda are very extensive in their description of the background of the study. table 1 s contains the distribution of sds by the number of lines they contain. it shows that a typical study has between 51 and 200 lines of information in its study description. the mean of the 962 studies is approximately 160 lines of information. the comparison between an sd and a library catalogue card is therefore not a very good one. the sds at the dda are more like a library catalogue card plus a very elaborate abstract of the study. the magnitude of information shows the potential for making a retrieval data base with the sds as input data. in an ongoing archival process, not all studies have optimal documentation; many studies are not yet fully processed. a few remarks about the dda level of documentauon may be useful here. at the dda, two categories are of especial interest to users. studies in class "d" contain a complete machine-readable codebook. from this documentation, we generate setups in spssor sas-format for the user as well as deliver published documentation on the study. studies placed in class "c" do have some machine-readable documentauon, but are not as "polished" as studies in class "d", and do not contain a machine-readable codebook. the study description item '001' contains the status of the study. however, at the dda, this item is updated by extracting information from our data base which keeps track of the processing of studies. as this process had not taken place at the time that i computed these staustics, i have computed table 2 directly from the processing data base for the purpose of showing the status of the studies. this total differs from the number of studies drawn from the sds. this difference is due to the fact that all the other staustics are made on the basis of the sds to be published in the catalogue of holdings. since the cut-off of new additions to the catalogue, 141 studies have come to our attention: these studies are typically placed in class l. op.cil 5 (tables have been collected together at the end of the article. ed's note) summer 1987 44 iassist quarterly description of the studies this section will illustrate the kinds of studies available from the dda. table 3 shows the subject headings or areas covered by the dda holdings. the table shows a heavy bias towards the traditional areas of the social sciences. election studies, general sociology and political science together constitute more than 60 percent of the studies. a similar picture is displayed when looking at the kind of data on which studies are based. as shown in table 4, 83 percent of the studies consist of survey data. earlier, approximately 76 percent of all studies had been generated by the old-fasioned oral interview. since then, there have been an increasing number of mail surveys. it is possible that the rising cost of conducing the traditional personal interview has led investigators to use mail surveys. furthermore, it is my impression that many surveys are now carried out by telephone, but this is not yet reflected in the distribution of these data. typically, the time span between data collection and the deposition of data in the archive is between 3 and 4 years. (table 5). the studies stored in the dda also appear to be conservative with respect to the cases on which the data are based (the "units of observation"). close to 90 percent of all studies are based on individuals. (table 6). in table 7 the definition of the universe is shown. a central variable is age. most election surveys concentrate on those over 18 (which is the age at which one obtains the right to vote in denmark). please note, that this table indicated a frequency total of 1303. of the 831 studies described in this item, many have more than one code specified. of interest are also those categories which are lacking in the table below. in the sd 'skeleton' file, provision is made for definition of the universe in terms of 'race' or 'religion'. these do not occur in any of the dda studies. when we consider the applied sampling procedures, we find that most surveys fall into one of two categories. those coded 'no sampling' are typically studies carried out at a distinct location (e.g. a working environment). the other major category is the many studies (25 percent) which are based on some kind of multi-stage sampling. (table 8). without performing detailed cross-tabulations, it is nonetheless easy to outline the characteristics of a typical study in the dda collection: it would seem to be an election study, carried out as a multi-stage sample survey, with individuals as participants and as units of analysis. the dda data bank as potential for analysis it is interesting to note that, according to the frequencies on item 211, approximately 23 percent of the studies are panel studies. this high percentage is due to an agreement between dda and the public opinion and marketing bureau observa a/s to the effect that all their political panels from 1967 to the present are being stored in the dda". apart from the observa studies, the remainder of the panel studies are typically election studies also. it is also worth noting, in table 9, that the the largest number of studies are in the category of cross-sectional sectional studies with replication the observa project was described by karsten bove rasmussen and lone borgersen in dda-nyt rir. 34, 1985 summer 1987 iassist quarterly 45 of one form or another. these, together with the large number of panel studies, comprise a majority of the studies at the dda which are connected as part of a panel study or having variables which are very similar, and which therefore present great potential for secondary analysis. the european archives have often tried to identify studies carried out within the same period of time in a number of the european countries. at the dda we have introduced a new item (225) to indicate area of coverage. table 10 shows that the vast majority of the studies at the dda are national and therefore cover all denmark. about 40 of the studies, however, are cross-national. these studies are typically the euro-barometers conducted for the european economic commission, but within the last few years the dda has identified some other cross-national studies. it should be noted here that the dda does not publish information about studies also stored in other data archives, unless they contain information on danish matters. to obtain access to a dataset deposited in the dda, the user is asked to fill out a requisition form and send a one-page description of how the data are to be analyzed. this information is then sent to the depositor or other person authorized to permit access. most of the studies (65 percent) are without any access restrictions for the typical user working in the social sciences. surprisingly, a large number of studies are categorized as being available only by special arrangements with the access-granting authority. ai the dda. we prefer not to store studies thai are not available to users. the studies in this category may possibly be those of which the principal investigators have not yet finished analysis, but it might be worth rechecking the access conditions of studies in this category. (table 11). time of study and the data matrix in this secdon we investigate the hard facts of some numeric items from the studies placed at the dda. a discussion of whether this material may in any way be regarded as being representative of social science data in denmark can be found at the end of this paper. tables 12-14 are believed to show nothing but the distribution of the studies placed at the dda. one of the key variables is the starting year of the time period covered by each study (item 220:01). this new variable is not be confused with item 231:01, the data collection date. for surveys, these two dates will of course be identical, but for historical studies based on old documents (e.g. parish registers or census lists) item 220 is indispensible. subtracting item 220:01 from item 220:02 (end year) shows that 75 percent of the studies start and end in the same year. on the other hand, 15 studies cover a time period of more than 100 years. in table 12 the start year is cross-tabulated with a grouped variable containing the number of cases in each data file. although the dda was founded in 1973, more than 20 percent of the studies in the archive deal with a time period previous to that year. of the 62 studies covering the period before 1950, 42 studies are concerned with a time period before 1900. these historical studies have typically a large number of cases, but are problematical in that they are not easily sampled. the cases are inter-related (i.e. family reconstitution data) and therefore all cases must appear in the data material so that the relationships can be determined by computer. item 212:01, the number of cases, is missing in approximately 28 percent of the studies. most of these studies are new. the number of cases is missing because for many of these studies summer 1987 46 iassist quarterly only a limited description is as yet available; dda has as yet received neither the data file nor precise information about the file dimensions yet. the typical data file has between 800 and 2500 cases. these newer, partially documeted studies are also a major portion of the great number of missing cases in table 13 below showing the number of variables in the datasets. the studies dealing with the time period before 1950 (the 'historical' studies) contain, as expected, a low number of variables. the information concerning the historical cases is not very full, but the number of cases is as shown in the previous table often huge. in the course of the last 25 years there seems to have been a rise in the number of variables in a single study. the average number of variables is approximately 160 variables per dataseu table 13 shows the cross-tabulation of number of cases by number of variables. as has already been mentioned, there seems to be a weak reverse relationship. on the one hand, the datasets with few variables have a large number of cases. on the other hand, the datasets with more than 40 variables are concentrated around the 801 to 2500 case size. studying data archives or social science the tables in this paper have shown the contents of the data bank at the dda. the interrelationships between the studies, and their accessibility . at the same time, the tables have served as a test of how the sds at the dda are being completed by presenting an overview which is verv difficult to obtain when the sds are viewed as isolated entries. the collection of descriptions of danish social science studies provides an opportunity to examine this collection as a sample of danish social science empirical research. but is it possible to draw valid conclusions from this sample? i do not intend in this paper to present a solution to this problem. but i should like to discuss some of the problems of bias that must be discussed before any conclusions can be drawn from the collection of sds. it is my intention, in raising these problems, to stimulate a discussion which will be of benefit to the future comparative analysis of the characteristics of the data holdings of the other european archives. it is indeed questionable if the collection of studies in the dda is representative of danish social science research. first of all, the studies represented consist of data collections or empirical studies. secondly the studies must be machine-readable, which will normally mean that computer analysis has been performed on the data. the target of the analysis will therefore at least be iimited to danish machine-readable empirical social science studies. the most serious threat to the validity of the analysis is whether or not there has been a change in the dda's criteria for incorporating a study into the data archive. given the limited resources available at the data archive, it is to be expected that over the years some changes in the basic criteria may have taken place. furthermore, it is to be expected that such a 'drift' in criteria may have happened unobserved and without being part of explicit archival policy. one way to prove or disprove this hypothesis would be to compare the holdings of the dda with a complete inventory of danish social science research. but the dda catalogue of summer j 987 iassist quarterly 47 holdings is the only available source which is close to being a complete inventory. such a comparison would result in tautological nonsense. instead, a few limited areas could be compared. the dda has deposition agreements with some research institutions as well as with the danish social science research council. thus the dda could check whether these agreements are being fulfilled, or if some studies are for any unknown reason not brought to its attention. until there has been a further, more thorough investigation into the representativeness of the studies placed in the data archive, the tables above cannot be construed to be representative of social science research. on the other hand, even if the representativeness of the sample is not tested, we can argue as follows: bias is inherent in the selection of social science data files for archival storage. one major source of bias is technical. a simple 'rectangular' survey file is the archetypic data file in data archives. these files are practically ready for storage on receipt by the archive, while a hierarchical study demands more data processing by the archive. both types of studies, of course, need the production of the proper machine-readable documentation. the other major source of bias in the selection of social science data for archiving lies in the nationalistic characteristics of the social sciences. because of the similarity in standards and technical capabilities of the four european archives (in germany, great britain, the netherlands and denmark), the technical bias can be isolated. based on these assumptions, a table showing the differences in the holdings at the european archives may serve as a guideline for the actual differences in social science research being carried out in the four respective countries.n tables 'nu nber of lines in sds" 100000 frequency percent 1-50 56 5.8 51-100 395 41.1 101-200 458 47.6 201-300 44 4.6 301 + 9 09 total 962 table 1 class/status freq uency d: fully machine readable documentation 285 c: no codebook 56 b: available from primary investigator 72 o: being processesd/ongoing acquistion 456 l: only preliminary donor contracts 234 total 1103 table 2 "area" item 002 frequency percent missing 20 0.2 organizational 61 5 6.4 general sociology 134.8 14.0 history, demography 53.1 5.5 law & criminology 250 26 political science 91.1 95 social physics 33 7 3.5 social medicine 758 7.9 welfare & leisure 48.7 5.1 socialization 41.7 4 3 election studies 3632 37.8 macroeconomics 17.0 18 microeconomics 138 13 (total # of codes = 1 358) table 3 summer 1987 iassist quarterly "kind of data" item 202 frequency percent missing 90 0.9 survey 797.5 829 census data 9.1 1.0 statistics 52 5.4 legislative roll 4.0 0.4 clinical data 87 09 textual data 32 0.3 coded textual data 23 2 2.4 coded documents 55 3 58 (total # of codes = 1042) table 4 method of data collection" item 232 frequency percent missing oral intreview 100.0 452.7 104 47.1 telephone survey mail survey pencil & paper psychological test other 13.8 317.5 73.5 05 39 1.4 33.0 7.6 0.1 0.4 (total # of codes = 1019) table 5 "definition of target population" item 222 frequency percent missing 131.0 13.6 age limits 535.5 55.7 sex 5.7 0.6 marital status 1 1 ethnic group, nationality 1.8 0.2 language characteristics 2 00 location of unit 69 7 7.2 housing conditions 57 0.6 postion in family 05 1 occupation 47 1 4.9 education 170 1.8 physical conditions 8 2 0.8 mental conditions 3.2 0.3 time limits 116.4 12.1 other 19.0 2.0 (total # of codes = 1303) table 7 "units of observation' item 211 frequency percent missing individuals 80 8580 8.3 892 families/household 160 1.7 groups other 5 5 2 5 0.6 03 (total # of codes = 972) table 6 "sampling procedures" item 223 frequency percent missing 329.0 34.2 no sampling 178.4 183 quota fample 24 0.2 simple random number 91.4 9.5 stratified random sample 36 2 3.8 area-cluster sample 13.0 1.4 multi-stage sample 245.7 255 other 66.2 69 (total # of codes = 1019) table 8 i "time dimensions" item 221 frequency percent missing 59.0 6.1 cross-sectional 322.6 33 5 as above with partial replication 3500 36.5 panel study 218.6 22.7 trend study 100 1.0 other 1 1 (total # of codes = 974) table 9 summer j 987 iassist quarterly 49 "geographical code" item 225 frequency percent missing local regional natioanl cross-national (total # of codes = 985) table 10 2 133.5 34.5 750.5 41.5 02 13.9 36 78.0 43 "accessiblity" item 331 frequency missing no access restrictions 169 no restrictions to scientific usuage no publication without permission no use of data without permission available after special arrangement other access conditions 3 (total # of codes = 965) table 11 40 47.2 1 4 154 14 1 04 d213:02 (dimensions of data number of variables) d220:01 (time period begin year) frqncy missing before 1950 19501959 1960 1969 19701979 1980total missing 4 5 1 43 171 120 344 1-40 34 25 12 12 45 6 134 41-100 11 17 16 33 56 26 159 101-250 1 i 8 2 4 154 46 215 250+ | 1 j 7 1 7 69 25 110 99 495 223 962 d212:04 (dimensions of data number of cases) d213:02 (dimensions of data number of variables) frqncy missing 1-40 41-100 101-250 250+ missing 244 8 6 6 5 269 1-800 80130 24 28 44 31 157 2500 250137 39 86 128 56 346 20000 23 35 31 28 15 132 20001 + 10 28 8 9 3 58 d212:04 "dimensions of data number of cases" d220:01 "time period begin year" frqncy misbefore 1950 1960 1970 1980total sing 1950 1959 1969 1979 missing 5 3 2 20 151 88 | 269 1-800 9 6 11 14 70 47 157 8012500 6 14 17 47 203 59 346 250120000 3 30 2 12 57 28 132 20001+ 28 9 6 14 1 58 total 51 62 32 table 12 99 495 223 962 total 344 134 159 215 table 14 summer 1987 iassist quarterly fall winter 2012 23 iassist quarterly abstract the need to prove and improve trustworthiness is an issue not only for new archives but also for those who have been doing this job successfully for a considerable time. but what happens when an established data archive faces the challenges of certification and audit processes? founded in 1960, the gesis data archive for the socials sciences has been “in the business” for more than 50 years and is one of the oldest archives in germany to preserve electronic resources for the long term. driven by a growing awareness of the needs of its stakeholders, who have to be sure that the data they produce, use, or fund is treated according to common standards, the data archive started a process of audit and certification within the european framework for audit and certification. after giving an overview of the gesis data archive and the european framework for audit and certification, this article describes how existing workflows were evaluated with regard to the requirements of the chosen level of certification. while the workflows themselves are already in place, in some cases the evaluation process showed a lack of appropriate documentation. suitable documents have to be created and made available to the public. in some cases, this process has to be accompanied by discussions within the institution about the mission and goals of the archive. in our experience, these can be very productive and lead to a common understanding and an improvement of services2. keywords: digital preservation, trusted digital repositories, certification. tried and trusted experiences with certification processes at the gesis data archive by natascha schumann1 gesis figure 1: research data life cycle 24 iassist quarterly fall winter 2012 iassist quarterly the gesis data archive and its organizational context as an infrastructure institution for the social sciences, gesis not only carries out research but provides services in all phases of the research data life cycle (see figure 1), from the conception to the archiving and re-use of social science research. among others, it provides consultation in methodology, develops software tools for research, and offers information services. the gesis data archive for the social sciences is one of five departments of gesis and has been providing comprehensive data services for national and international comparative surveys for several decades. one of its main tasks is to make research data available for re-use. to support this goal, the data archive has an explicit mission for long-term preservation, which is also laid down in gesis’s by-laws. accordingly, among gesis’s primary objectives is the “archiving, documentation, and long-term preservation of social sciences data, including the indexing of data as well as the high-quality enhancement of particularly relevant data to prepare them for re-use” (gesis constitution § 2).3 workflows within the archive are organized according to an archival life cycle, ranging from pre-ingest (incl. acquisition) to ingest and processing to archival storage up to the dissemination of data. the central functions of the oais reference model (ccsds, 2012) can be mapped to the existing structure of the archive (see schumann and recker, 2013). however, although the gesis data archive already has working procedures and processes in place to ensure the preservation of its data, there is still a need for further activities. for example, in some cases documentation of workflows and defined interfaces between different steps of the preservation process are lacking, and some definitions of information packages are not up to date. by addressing these issues in a systematic fashion, the archive aims to further increase its trustworthiness. a need for trust a definition of a trusted digital repository is given by the rlg/oclc working group: “a trusted digital repository is one whose mission is to provide reliable, long-term access to managed digital resources to its designated community, now and in the future” (research libraries group, 2002, p. i). but what does that mean in detail? audit and certifications standards such as the nestor catalogue of criteria for trusted digital repositories on the one hand add aspects from an it security perspective – for example, “authenticity, integrity, confidentiality and availability” (nestor, 2009, p. 1). these technical aspects are relevant issues for trust, but beyond that organizational aspects are just as important. as pointed out in audit and certification of trustworthy repositories (ccsds, 2011), “[c]onstant monitoring, planning, and maintenance, as well as conscious actions and strategy implementation will be required of repositories to carry out their mission of digital preservation” (p. 2-1). these quotations show that building trust depends on more than one factor. a trusted digital repository has to ensure that the digital objects it preserves are not corrupted by accident or intentionally, and that access is given – not only physically, but also in appropriate digital formats. another criterion of trust is if and how the organization demonstrates its know-how in digital preservation and, for example, if succession plans exist for the case that the institution ceases to exist. thus, transparency is very important in the context of trust. all stakeholders should have the opportunity to ascertain the statements made by the institution. accordingly, to appear trustworthy, the gesis data archive has to provide stakeholders – data depositors, data users and funders – with sufficient information to demonstrate that their data is treated according to the agreed standards of the social sciences and digital preservation communities. because existing certification standards and audit tools support archives in the building of trust, we decided to start a process of audit and certification within the european framework for audit and certification (see below). our decision to do so coincided with similar efforts initiated by the council of european social science data archives (cessda), of which gesis is a member. as cessda has been transformed into a new organization and legal form, cessda as, and is on its way to becoming a european research infrastructure consortium (eric), it is necessary that all member institutions agree on the same standards regarding trustworthiness. to start off this process, during 2013 all archives carried out a self-assessment based on the guidelines of the data seal of approval (dsa; see below). all of this illustrates that the need to prove trustworthiness is not only an issue for new players, but also for established ones like the gesis data archive. however, the challenges such “established players” face are somewhat different from those that new archives have to deal with: it is a different kind of procedure to set up a completely new service or to conduct a certification process in an existing system. thus, when setting up a new archive it is possible to take into account the requirements for trusted digital repositories from the outset. what is more, new archives can benefit from other institutions and their experiences and avoid mistakes. in contrast, an existing archive may have gained a lot of expertise and know-how over time, but it can be very complex and challenging to adapt established workflows to new requirements. european framework for audit and certification of trusted repositories over the years, many different approaches and standards have been developed in the field of audit certification for trusted digital repositories. the most established among them are • the data seal of approval (dsa), originally initiated by dans, • the nestor catalogue of criteria for trusted digital repositories, which became a german din standard (din 31644) in 2013 and will also be available in english, and • the trusted repository audit checklist (trac), which is also an iso standard (iso 16363). to achieve greater harmonization between these different initiatives and criteria catalogues, a memorandum of understanding (mou) was signed in 2010 for a european framework for audit and certification4. this process was accompanied by the european commission and the alliance for permanent access to the records of science (aparsen) 5. the mou defines three levels of certification (see figure 2): 1. basic certification is granted by obtaining the dsa. 2. extended certification requires completing the dsa and an externally reviewed self-audit based either on iso 16363 or din 31644. 3. formal certification requires completing the dsa and a full external certification based either on iso 16363 or din 31644. data seal of approval the target audience of the dsa are repositories committed to longterm preservation. working from the assumption that data quality is dependent on “aspects related to the creation, storage and (re-) iassist quarterly fall winter 2012 25 iassist quarterly use of digital data” (dsa, 2013, p. 5), the dsa contains 16 guidelines reflecting different roles: data producers, data repository and data users. although the main focus of the dsa is on data repositories, it is open to other digital archives as well. there are different levels of compliance for each guideline: to be awarded the dsa, the minimum compliance level as stated in the dsa guidelines has to be reached by the applicant (see dsa, 2013, p. 6). the archaeology data service has published a best practice report to support other institutions in obtaining the dsa (mitcham and hardman, 2011). nestor seal/ din 31644 the din 31644/nestor seal is based on the nestor catalogue of criteria for trusted digital repositories (2009). in 2012 it was accepted as the german standard din 31644. it contains 34 criteria covering the following thematic areas: organizational framework, handling of information objects and their representations, infrastructure and security. the level of compliance is measured on the following scale: rac/iso 16363 the repositories audit checklist (rac) was developed from the audit checklist for the certification of trusted digital repositories (2005) and trustworthy repositories audit and certification: criteria and checklist (2007). it became an iso standard in 2011 and information about the standard can be found at the primary trustworthy digital repository authorisation body (iso-ptab)6. rac consists of 50 main criteria and has 109 criteria in total. their structure is orientated towards the nestor catalogue and accordingly rac criteria cover the following areas: organizational infrastructure, digital object management, infrastructure and security risk management. our approach as stated above, the gesis data archive decided to commence its certification activities with the dsa. several reasons contributed to this decision. first of all, the dsa is a good starting point for certification activities because it addresses all basic aspects of trust but is not as detailed as either din 31644 or iso 16363. thus the dsa is an immensely helpful tool to gain an overview of the processes within our archive. it therefore helps us lay the groundwork for follow-up activities in audit and certification within the context of the european framework. getting started as the dsa is a self-assessment, the applicant completes a selfassessment statement for each of the guidelines including links to the relevant documentation or evidence (see dsa, 2013, p. 6). this will then be reviewed by a peer reviewer appointed by the dsa-board. as a first step we evaluated the existing workflows with regard to the requirements stated in the guidelines. in this manner we obtained an overview of those workflows already in place and supported by sufficient documentation. this initial evaluation showed that the majority of our workflows comply with the dsa guidelines. however, a lack of appropriate documentation, especially on our website, became apparent. in consequence, the main tasks at this stage were: − the detection and subsequent creation of missing documentation and documents, e.g. policies or recommendations. − the revision of existing documentation and documents if those were not up to date or not yet ready for publication. e.g. a versioning policy had to be updated. an example for a newly created document is our preservation policy. it contains information on the organizational context of the gesis data archive, states our mission, and describes the main figure 2: three levels of certification within the european framework for audit and certification 0 not applicable 1 no. we have not considered it yet 2 theoretical. we have a theoretical concept 3 in progress. we are in the implementation phase 4 implemented. this guideline has been fully implemented for the needs of our repository 0 no  concepts  are  in  place 3 a  concept  exists 6 well-­‐elaborated  concept 10 implemented 26 iassist quarterly fall winter 2012 iassist quarterly principles of our approach to the digital preservation of data (see also friese in this issue). the process of developing the policy not only meant creating the content, it was also a process requiring a good deal of coordination: staff members from different teams had to be involved as well as the head of the department. as it is insufficient for the certification process to publish the policy only in german, an english translation was also created. as explained above, an important requirement for trustworthy digital repositories is transparency: all relevant information should be available to the public. in our case this meant redesigning the archive’s website and deciding which information it should include. in this process, we reconsidered the whole structure of the website. it is now organized along the steps of the data lifecycle and refers to the functional entities in the oais reference model. creating the website content was another challenge and involved more than simply adding some documents. we had to answer the question of how to make the website helpful for the different stakeholders and user groups: data producers, data users, funders, and other interested persons – both in terms of content and the language used. experiences and benefits one of the (first) benefits of the preparatory work on the dsa was that we gained an overview of what we already have in place and what will have to be amended or improved. it became clear that we would have to reconsider some of our workflows. traditionally, the focus of the gesis data archive has been the indexing and processing of empirical social research data. over the past years, more and more digital preservation issues became relevant – for example, questions of how to implement procedures to ensure authenticity and integrity, or the need to define different archival packages referring to the oais model etc.. but not only did the archive have to update workflows; an important part of this process (which is not completed yet) was the development of a common understanding of digital preservation: is it an “an added value” that we somehow create on top of everything else that the archive does? or is it not rather the sum of everything we do in the archive: the bundle of measures, that is, that we employ to ensure access and use of the data in the long term, including, for example, the creation of metadata and documentation, doi registration, etc.? in addition, the strengthened focus on digital preservation has consequences not only with regard to processes and procedures but also for the level of transparency. stakeholders have an increasing interest in learning how their data is curated and preserved, and the archive has to provide the respective evidence in order to maintain its stakeholders’ trust. however, the dsa application not only required us to address external stakeholder communication; it also prompted a review and evaluation of our internal communication and documentation procedures. for example, the archive staff uses a wiki for internal documentation purposes. it was built and filled with content over the last years, but our review in relation to the dsa application showed that it was not up to date – neither with regard to the contents nor to the structure of our workflows. we are now in the process of restructuring and updating it, which also includes agreeing on the current procedure of maintaining the wiki and keeping it up to date. the processes set in motion by our decision to apply for the dsa have helped to create awareness of the capabilities and strengths as well as of weaknesses or gaps within our archive. the systematical compilation of existing and relevant information and documentation required by the dsa is a helpful step in itself, and this gap analysis has helped us gain a more concrete idea of our workflows and their documentation instead of the vague feeling that “surely, everything works as it is supposed to.” preparing the dsa application served as an incentive to continue some projects – e.g. for new services or the adoption of standards – that had been planned for some time but had been neglected in the face of “more pressing” problems arising in our day to day work. some of the required measures are easily created and implemented, but others need more time and discussion to be realized. the fact that so many activities are linked to each other entails that it may take some time to implement new processes: the necessary changes concern different applications or workflows cutting across different teams and cannot be made without implications for other parts of the system. the process of (preparing) an audit or certification is very time consuming. but in our experience, the process of preparing the dsa self-assessment statement produced many valuable effects in that it helped us establish a common understanding for the mission and goals of our archive. the first steps of complying with the guidelines of the dsa have been made and we have gained an overview of our capabilities as well as existing gaps. accordingly, we now know where we stand and what our tasks for the near future are. our next step will be to hand in the dsa application. after this is completed and the dsa will have been granted, we are already planning to take the next step in the european framework for audit and certification: the extended certification, which will take the form of conducting a self-audit for the nestor seal, based on din 31644. references ccsds, 2011. audit and certification of trustworthy digital repositories. [pdf ] available at: [accessed 03 december 2013] consortium of european social science data archives (cessda-as). homepage. [online] available at: [accessed 03 december 2013] data seal of approval: homepage. [online] available at: [accessed 03 december 2013] din, 2012. din 31644: kriterien für vertrauenswürdige digitale langzeitarchive. available at: [accessed 03 december 2013] friese, y., 2014. how to develop a preservation policy? guidelines from the nestor working group. [pdf ] in: iassist quarterly [include information for this issue]. mitcham, j. and hardman, c., 2011. ads and the data seal of approval – case study for the dcc. [online] available at: [accessed 15 november 2013] nestor, 2009. nestor criteria: catalogue of criteria for trusted digital repositories, version 2. nestor materials 8. [pdf ] available at: [accessed 03 december 2013] iassist quarterly fall winter 2012 27 iassist quarterly schumann, n. and recker, a., 2013. de-mystifying oais compliance: benefits and challenges of mapping the oais reference model to the gesis data archive. [pdf ] iassist quarterly, 36 (2), pp. 6-11. available at: [accessed 03 december 2013] research libraries group, 2002. trusted digital repositories: attributes and responsibilities. an rlg-oclc report. [pdf ] available at [accessed 03 december 2013] crl and oclc, 2007. trustworthy repositories. audit & certification: criteria and checklist. [pdf ] available at: [accessed 03 december 2013] yakel, e., faniel, i., kriesberg, a., and yoon, a., 2013. trust in digital repositories. [pdf ] the international journal of digital curation, 8.1 (2013), pp. 143-156. available at: [accessed 03 december 2013] notes 1. natascha schumann is affiliated at the data archive for the social sciences at the gesis leibniz institute for the social sciences in cologne. the main focus of her work is on digital curation of social science research data and audit and certification in this area. her contact email is natascha.schumann@gesis.org. 2. this paper is an updated version of a presentation given at the iassist 2013 conference in cologne. 3. gesis constitution (in german): http://www.gesis.org/das-institut/ der-verein/satzung/ 4. http://www.trusteddigitalrepository.eu/site/welcome.html 5. http://www.alliancepermanentaccess.org/index.php/aparsen/ 6. primary trustworthy digital repository authorisation body (isoptab): http://www.iso16363.org/ sist newsletter vol.1, no. 1 activities and plans a comprehensive questionnaire on archive acquisition policies has been drafted by marcia taylor and is now being reviewed by action group members in north america and europe. the questionnaire addresses the criteria for selection including contractual obligations, medium of storage, geographic scope, field of study, quality, selection procedures, and incentives and sanctions for encouraging/discouraging deposits. whereas other action groups are concerned with various means of making data accessible to the user, this action group will focus on the problems relevant to the initial archival acquisition. related to this issue is the possibility of joining with other organizations in providing incentives for both academic and governmental data producers to make their data files more generally usable and more widely available. the action group intends to address questions of deacquisition as well as acquisition policies to the needs of a local service data library. data documentation canadadave l. salley, management and central services group, standards division, statistics canada, tunney's pasture, ottawa, ontario kia 0t6 europecees middendorp, steinmetzarchief , kleine-gartmanplantsoen 10, amsterdam-c. , netherlands united statesjohn grasso, office of research and development, center for appalachian studies and development, west virginia university, morgantown, west virginia 26506 mandate this group will develop standards for the measurement of variables, i.e., the definition of constructs and their operational isation and measurement. in principle, these standards should be applicable cross-nationally, although this may not always prove to be possible. standards will be developed for "simple background variables" used in surveys, i.e., educational level, age, head of household, as well as constructs such as job satisfaction, anomia, political interest (i.e., to be measured by a scale or index). thus, the work of this group will be closely linked to that which is going on regarding the development of social indicators. the codes will be incorporated into source books to provide researchers with a resource tool for coding and organizing their data consistently. i^ssist newsletter vol.1, no. 1 ^ activities and plans cees middendorp suggests that the work of this action group could perhaps be linked to the older inventories by robinson, et al . of the university of michigan survey research center. he notes the difficulties in solving the problems and implementing the recommendations, that standards can only be developed on the basis of research-evidence, and suggests that revisions will be necessary on the basis of either new theoretical insights or new empirical evidence. at the steering committee meeting in edinburgh, erwin k. scheuch of the institute for applied social research and the zentralarchiv presented a description of face sheet problems. he described his analysis of 19 studies in which he looked for comparability of demographic characteristics across studies. his findings indicated that the broad collapsing of categories masked various life cycle and income changes which have occurred over time. on the basis of his preliminary results, he suggested that lassist focus on the need for standardizing demographic variables. he also suggested that further study could be carried out by lassist members and the results be communicated to survey agencies. lassist recommendations could be used by these agencies to perfect their survey techniques and would provide lassist with an excellent opportunity for impact in the area of standardization of demographic variables. this action group has begun reviewing already available material in the field, which includes the work carried out by scheuch and van dusen and zill at the center for coordination of research on social indicators in washington, d.c. this work will form the basis for a project which will test the value of these independent variables in actual research and provide a basic list of such items with recommended coding schemes. in time, this should lead to the developmert of general recommendations for the design of codes. classification canadanot yet activated europeekkehard mochmann, zentralarchiv ftir empirische sozialforschung, bachemer strasse 40, 5 ktlln 41, federal republic of germany united statessue dodd, data library, institute for research in social sciences, manning hall, university of north carolina, chapel hill, north carolina 27514 mandate existing classification schemes for both studies as a whole and variables within studies would be examined. the results of such a review and recormiendations would be published. additionally, information required for the cataloguing of machine-readable data in libraries would be researched and concrete guidelines for such catalogues would be formulated. required classification schemes for the production of indicator source books and variable level retrieval systems will be investigated. the necessary components of an adequate study description will be identified. iassist quarterly 2013 5 iassist quarterly editor’s notes special issue: a pioneer data librarian welcome to the special volume of the iassist quarterly (iq (37):1-4, 2013). this special issue started as exchange of ideas between libbie stephenson and margaret adams to collect papers relating to the work of sue a. dodd. margaret adams (peggy) acted as the guest editor and the background and content of this volume is described in her preface to this volume on the following page. as editor i want to especially thank peggy and libbie for pursuing and finalizing their excellent idea. i also want to thank all the authors that contributed to produce this volume. as one of the authors i can witness that peggy did a great job. articles for the iassist quarterly are always very welcome. they can be papers from iassist conferences or other conferences and workshops, from local presentations or papers especially written for the iq. when you are preparing a presentation, give a thought to turning your one-time presentation into a lasting contribution to continuing development. as an author you are permitted “deep links” where you link directly to your paper published in the iq. chairing a conference session with the purpose of aggregating and integrating papers for a special issue iq is also much appreciated as the information reaches many more people than the session participants, and will be readily available on the iassist website at http://www.iassistdata.org. authors are very welcome to take a look at the instructions and layout: http://iassistdata.org/iq/instructions-authors authors can also contact me via e-mail: kbr@sam.sdu.dk. should you be interested in compiling a special issue for the iq as guest editor(s) i will also be delighted to hear from you. karsten boye rasmussen april 2014 editor 16 iassist quarterly spring 2012 iassist quarterly abstract the survey research data archive (srda) is the largest data archive in taiwan and in asia. it collects not only survey data in social sciences but also raw data of major government statistics. these archived data have made significant contributions to research. data and remote access service are provided without charge. in addition, an english website along with the english version of the data and their metadata will be available by mid-2013. to improve the search efficiency and promote itself among domestic researchers, srda began to launch a series of projects around june of 2011. these include the revision of abstracts, the construction of new search functions, and the compiling and circulating of a power point concerning the use of srda. this paper documents the endeavors, reports the current progress, and reflects on the experiences learned from the developments. keywords: asian survey data archive, archive management, archive development introduction data sharing is an important trend internationally. researchers deposit their data to archives after the research project is completed, while others go to the archive to identify and use these secondary data to support different research interests. however, optimal data sharing requires continuous hard work and innovation. data archives not only have to actively persuade data owners to deposit their used data, but also have to promote their services to potential users. most importantly, while making efforts to secure data confidentiality, archives must make data as accessible and “transparent” as possible, so that researchers can find out if any data suit their needs as easily as possible. such efforts increase the user-friendliness of the data and thus increase the possibility of data sharing. in 2011, survey research data archive (srda, https://srda. sinica.edu.tw/) of the center for survey research (http:// survey.sinica.edu.tw/) of academia sinica (http://www. sinica.edu.tw/index.shtml) in taiwan began a series of projects aimed at achieving such goals. in this paper i share our experience in conducting these projects. after an introduction to srda and its mission, i describe the challenges we faced and how those challenges were met. the paper concludes with a discussion of the experiences. the survey research data archive srda was established in 1994 by the center for survey research (csr) of academia sinica, and managed by the data division of csr. it is the oldest and the largest survey data archive in taiwan and asia. it currently has almost 1400 members, who can be faculty, research staff and graduate students at universities, or researchers in government agencies and research institutes. while any visitor can review documentation and summary statistics online (available in 2012), members can download datasets directly from the website, request assistance from srda staff, and use remote access for secure data. these services are all provided without charge. srda archives both the raw data collected by major government agencies for the production of important government statistics and data collected by academics. each dataset released by srda is carefully cleaned and documented, and is released with detailed metadata. by the end of 2012, srda has released 423 datasets collected by the government agencies and 1,132 datasets collected by the academics, totaling 12 gb. these data are a very important resource for research and teaching. in 2012, members initiated a total of about 13800 downloads. access to these data is responsible for a large number of publications in major national and international journals. although we are still trying to build up the database of publications based on data archived in srda, according to records available now, data from only the six major survey projects in taiwan archived in srda are the basis of 384 journal articles up to mid-2012. strategies of promoting the use of survey research data archive by meng-li yang1 srda iassist quarterly spring 2012 17 iassist quarterly srda continues to improve its services. in 2009, srda inaugurated on-line analysis service using the networked social science tools and resources (nesstar) software developed by the norwegian social science data services (the service was seriously underutilized up to the end of 2011, though, because only several datasets were uploaded to nesstar due to limitations to be specified later). to insure appropriate data security, the information security management system (isms) protocols (based on iso27001) were introduced in 2010. in april of 2011, srda, along with the data division of csr that manages it, was certified by the british standards institution (bsi) (iso 27001:2005). to enlarge the audience for its holdings, an english-language version of the srda website along with the datasets and their metadata will be online around mid-2013. most major datasets should have english versions available by then. the data division organizes activities and produces communications to promote srda holdings and services to scholars in taiwan. the division holds at least one workshop each year on important themes. these themes include skills for collecting and cleaning survey data, using important longitudinal survey data series, sampling methods, and advanced statistical analysis techniques. the division also issues a bi-weekly newsletter and a monthly e-digest to announce srda’s newly released data and activities. despite these successes, however, there was a sense that the data services needed to be more user-friendly and that information about srda should be disseminated more widely. a motivation for need of improvement some promotion strategies were already forming in june of 2011, but results from a survey to some extent confirmed the need to promote srda and its service. in october 2011, the national science council (nsc) conducted a web survey2 to solicit the opinions of scholars in selected fields. the survey target was scholars and researchers in humanities and social sciences who had submitted a grant proposal to the nsc in the previous five years (but who were not necessarily srda members).3 to gauge researchers’ use of srda, i took advantage of the opportunity to add several items4 to the survey. the results of the srda-related items confirmed our original impressions. of the 3019 respondents, 52.7% (1590 persons) had not heard of srda, and only 18.4% (556 persons) are or had been srda members. among the 28.9% (873 persons) that had heard of srda but had never been a member, 616 had never even visited srda web site, and 257 did visit but did not apply for a membership. among these 257 people , 17% said that information about datasets was insufficient for an effective search, and 10% similarly said that it was difficult to find needed datasets, although 64% had no data need. among those who are or once were srda members but never used srda datasets for research (n=295), 17% said they did not know how to find what they needed and 16% said they could not find what they needed, whereas 52% did not had data need. the two most important messages from the survey were that 1) more than half of the researchers who might find srda valuable were not aware of its existence; and that 2) among those who tried to obtain data from srda, about 30% were frustrated with the process. the messages reflected the difficult position srda was in. although srda strives to promote its collection and services to researchers, it seemed that only groups that are already familiar with survey data or, more specifically, with srda, can benefit from the activities. for example, attendance at srda’s survey data workshops was limited to those who could participate in person. from my own contacts with colleagues in the social sciences, many colleges in non-northern parts of taiwan were not aware of srda. although others wanted to receive additional training, csr’s limited resources forced us to refuse requests to hold on-site workshops at universities in other regions of taiwan. clearly, we needed to promote srda more actively. second, the survey results demonstrated that the site’s search efficiency was poor. for many reasons, the search function within the original srda archive was rather old and inefficient. for example, data produced by government agencies require a separate user application process, whereas datasets produced by academics do not. the result is that each type of data requires a separate search. in addition, whereas the variable-level search of nesstar is not functioning, the most powerful search in the original archive function searches only the abstract; the other search areas being the project name, the name of the pi, the serial number of the datasets, keywords, and the subject domains of the project. however, srda relied on data depositors to provide abstracts and keywords. unfortunately, depositors are not always aware of how important the abstract is in archive search and retrieval functions. poorly written or cursory abstracts minimized the effectiveness of srda’s original search functionality. therefore, except for searching within abstracts, the other search functions require the users to already know specifically what datasets they are looking for. even the on-line analysis function of nesstar was seriously underutilized. by the end of 2011, only several academic datasets were uploaded to nesstar because nesstar does not keep records of people who make downloads but csr needs such records. government data were not considered for nesstar at all, for the same reason that the use of government data requires further application. strategies the survey results provided evidence that srda should improve search and discovery efficiency and also actively promote its services to researchers across the country, building on promotion efforts that began in june 2011. in sum, three strategies were aimed at improving search efficiency, and one was to promote srda among all potential users. 1 improving the search efficiency 1.1 revising abstracts the first project launched was revising abstracts so that they included more information from questionnaires and accurately reflected the content of their datasets. the overall goal was to improve the effectiveness of searches. each revised abstract should contain the purposes (and history if applicable) of the survey project, contents of the questionnaire, the survey mode, the survey period, the target population, the sampling frame, the sampling method, and the sample size. i asked all the data division members to review the project proposal/report and the questionnaire for such information, and to revise the abstracts from the view point of the dataset. this was done in early 2012 for 22 waves of a longitudinal survey projects, totaling 44 datasets. checking and editing the revisions proved to be more daunting than anticipated. in the beginning, i doubted the value of including only the title of the questionnaire sections. ideally, concepts would be the most helpful for searching. however, the questionnaire of a social survey contains measures of all kinds of concepts. it is impossible to include them all in the abstracts. in addition, assigning concepts requires expertise in fields relevant to the goals of the survey, although staff members assisting with this project specialized in statistics. the result was to compromise and use only titles of the questionnaire sections 18 iassist quarterly spring 2012 iassist quarterly to describe the contents. although this compromise may decrease the potential use of abstracts in improving search efficiency, highly detailed abstracts are just not feasible. however, it is still important that abstracts contain all the other pieces of information, so that users quickly have a concise idea about a dataset by reading the abstract. therefore, later in the middle of 2012, i recruited a doctoral student good at writing. i worked with him on revising abstracts for several longitudinal survey projects, after which he began to work independently. 1.2 constructing a new search function members of the division offered much better ideas. around september of 2011, they suggested that we model our search function after the survey question bank maintained by the uk data archive at the university of essex.5 the question bank has a variablelevel search function. by using the ddi (data document initiative) format, with which nesstar is compliant, we can create a search function that allows users to find out the items of interest along with the data file by specifying key words in the items. this way, users can quickly find the exact data by specifying words/phrases of items that they need. as long as researchers know what items to find, they do not have to go through every possible dataset. a search function like question bank makes up for the deficiency of the original srda search features and accomplishes what searching in abstracts cannot achieve. so we decided to construct a question bank for srda. for this new search function, i recommended that the government data should be also made within the search area. so they have to be put in the nesstar. however, to decrease the risk of exposing any level of confidential information in on-line analysis, we allowed nesstar to perform only univariate analysis for the government data. the programming work began around the end of 2011. a division member undertook the system analysis (sa), and the programmer of the division did all the programming tasks. during this time we also uploaded all datasets, government as well as academic, to nesstar. in september of 2012, we put the question bank on line for service under the function name “search by item contents” (http://140.109.171.171/ bank/). this new search function eliminates all the hassles of searching in the original srda and takes full advantage of nesstar software. that is, by linking the original srda database with nesstar, the search function searches in nesstar the contents (variable labels and variable concepts) of government datasets and academic datasets at the same time and presents the results separately. users are linked back to the original srda for downloading or requesting datasets. for nesstar’s on-line analysis functions, users can use all the functions for academic datasets, but only the univariate analysis for government datasets. more importantly, the new search function offers two types of search. the first type is called “search for datasets.” any dataset is listed that contains all the texts/concepts of variables entered by the user, whether or not these appear in the same variable. users can also limit the search by specifying the range of years when the data were collected, the range of the sample size, keywords, words in abstracts, the project name and the name of the pi. the second type is “search for a specific variable,” which is actually a method transplanted directly from the computer program used by the data division to construct the concept bank (explained later). using this option, one can enter up to five phrases to identify a variable in mind. any datasets that contain a variable which includes all the texts entered is listed. one powerful feature of the new search function is that results of each search method can be modified by either of the two methods. finally, as we are also constructing an english version of the archived data, both english and chinese versions of an identified dataset are always linked together. this way, researchers will be able to use the english translation of the data directly if they wish to submit the analysis to an international journal. the english version of the srda website will also have this search function with all the features available, where english texts and english concepts are used for searching. the features of the new search function are summarized in table 1. search  for  a  specific  item search  for  a  dataset type  of  search  term   • variable  text • variable  text • variable  concept maximum  number  of   terms 5 3 applicable  search   restric:ons   none • collec3on  year  range • lower  limit  of  sample  size • name  of  pi • name  of  project • words  in  abstract • project  keywords logic  of  search intersec3on:only   variables  containing  all   search  terms  are   displayed   union:datasets  mee3ng  the   search  criteria  and  containing  all   the  search  terms  are  displayed;   search  terms  do  not  necessarily   appear  in  an  item. search  further?   yes,  and  can  use  the   “search  for  a  dataset”  to   do  further  search.   yes,  and  can  use  the  “search  for  a   specific  item”  to  do  further   search. table  1.  the  two  types  of  search  in  the  “search  by  item  contents”  func:ontable  1.  the  two  types  of  search  in  the  “search  by  item  contents”  func:ontable  1.  the  two  types  of  search  in  the  “search  by  item  contents”  func:on iassist quarterly spring 2012 19 iassist quarterly 1.3 constructing a new item-level search option—the concept bank the division started to develop a “concept bank” in 2010 but abandoned the project in early 2011, before i became the advising researcher of the data division. the idea of a “concept bank” is to assign concept(s) to every variable for the archived data, so that users can also use concepts to search for variables. however, in 2010 the division did this by translating an english thesaurus for the social sciences to chinese, a “top-down strategy.” when this was almost done, they met with three obstacles. first, they found concepts that are not applicable to taiwan’s situation and vice versa. second, there are concepts that seem to have more than one translation. third, they could not find resources (expertise) to assign these concepts to items. while i believed in the value of building the concept bank, i thought the top-down strategy was not efficient. instead, i proposed a bottom-up strategy by asking experts to assign concepts for variables, in both english and chinese. my idea was that the structure of concepts that a thesaurus offers may not be essential when used in variable-level searching. especially, compared to the bottom-up strategy, the top-down strategy may require a great deal more manual labor—to link the concept for each variable back to the thesaurus— in addition to the expertise required for the assignment. even asking experts to assign concepts from the thesaurus had limitations, and, after all, the thesaurus did not always incorporate concepts unique to taiwan’s situation. also, concepts that do not apply to taiwan’s situation should be no concern at all because the archived data would not contain such concepts. the problem of one single concept corresponding to more than one translation does not need to be a concern either, since experts may know most of the translations, and future users should also know the several chinese translations of a certain english concept and vice versa. the division’s greatest concern for the bottom-up strategy, and also the reason why it adopted the top-down strategy before, was that different scholars are likely to define different concepts for identical variables, which would diminish the bank’s value for searching. anticipating this problem, i proposed designing a computer program to check for the inconsistency. people working on the project can use the program to check if variables that are supposed to have identical meanings do have identical concepts, and if not, they can output the concepts along with the variables and have the concepts revised by some other experts. furthermore, i suggested that once we have concepts for a variable, we can save time and resources by assigning the same concepts to variables almost identical except for small differences in non-substantive words. in short, we needed a computer program that allowed us to input and output concepts to nesstar, and also to identify variables with identical substantive meanings except for some non-substantive words. such a function was quickly designed and programmed by division members. this program allows one to specify up to five phrases to locate a variable. the staff member uses it to identify variables that contain these phrases, check the concepts for consistency, and, if necessary, output them for revision, and, afterwards, input the revised concepts. when consistency is assured, the staff member uses the program to assign the processed concepts to all the other variables that are identical in meaning. i am responsible for revising concepts for items with inconsistent concepts. my principle of doing the revision is to include all the concepts unless they are obviously wrong. after all, concepts can be very specific or very general; including them all may serve researchers’ different needs. thanks to resources from nsc, we invited 45 scholars to assign concepts for 45 surveys in 2012. in the first round of the invitation, we selected studies of a wide range of topics from several longitudinal projects. each topic was represented only by one study. for example, only the most current one, rather than all, of the surveys on political science was sent for concept assignment. the same was done on topics of family studies, citizenship, secondary school students, teachers, etc. nevertheless, even surveys supposedly on different topics have many overlapping variables, and there is rather low consistency among the concepts assigned to identical variables, as the division has expected. further, the timeliness with which scholars returned the completed material varied widely. consequently, even though we completed consistency check for variables within several studies, those received later introduced more inconsistencies. to accommodate this problem, we decided not to check until all the studies of the same project are returned. and then, after we have checked the consistency for variables in all the studies that were sent out, we can start to assign concepts to identical items in all the other datasets. by the end of 2012, we had completed the assignment for the 45 datasets of one longitudinal project (the goal in the grant proposal), and several other surveys of different series on a variety of themes. the concepts are put in nesstar for use in the new search function. we are still waiting for more scholars to send back their work so that we can resume the consistency check. in 2013, we obtain additional funding from the nsc to invite scholars to work on additional surveys with different themes. learning from the experience, we will be working with a smaller number of studies as we expect the load of consistency checking will increase as more studies are assigned concepts. 2actively promoting srda across the country as mentioned earlier, the 2011 survey results indicated that many potential users still did not know about srda and efforts to promote the archive to researchers were needed. early in 2012, i decided to put together a power point program that introduces srda and also contain some examples of how archived survey data could be applied to research projects. the introduction of srda itself was easily completed, but the demonstrations had to be designed from scratch. i chose to demonstrate the usefulness of survey data in two ways. the most obvious one is to point out research articles that analyze survey data. i wanted to focus on articles that are easily found in the internet and would encourage use of srda datasets. since the data division did not have a spare hand for such a job, i asked my own part-time assistant to find such articles in tssci (taiwan social science citation index) journals, and to compile an abstract for each. i spent a significant amount of time revising the abstracts. however, it became clear when the abstracts were inserted in the power point file that much of this effort was unnecessary. because the power point file is for self-viewing, an article is easier to understand if it is discussed on one slide. however, a long abstract is too long to fit into one slide. the result was that i shortened the abstracts into research questions for each of the studies (22 studies) included in the power point file. another way of highlighting the potential value of less well-known datasets was to write up some analysis using the data to answer simple research questions. this strategy is adapted from the icpsr on line learning center (http://www.icpsr.umich.edu/icpsrweb/olc/). another part-time assistant of mine (a doctoral student) wrote the analyses. from those, i selected seven works for the power point file. i ran into the same problem as in the abstract case, and had to shorten the 20 iassist quarterly spring 2012 iassist quarterly skill development and innovation. after all, working for data is a very special application of their non-statistical skills. without a nurturing environment, such people may soon feel frustrated and leave the organization. i myself actually had spent some time and the other two members also spent time listening and talking with such people, so that they felt they had someone (if not all) in the division to rely on when frustrated. from the experience, i found that designating a mentor for them not only helped keep them in the organization but also enhanced their performance. the mentor does not have to be as skillful as the new colleague in the specialized area. the mentor just needs to be kind enough to be willing to help a completely new learner and to provide information on matters that are related to where the skill is to be applied. notes 1. center for survey research, research center of humanities and social sciences, academia sinica. address: 128 sec.2 academia road, nankang taipei taiwan 11529. contact via email: mengliya@gate. sinica.edu.tw. this is an expanded version of a paper presented on june 7, 2012 at the iassist 38th annual conference in washington, dc., 2. the web survey was implemented by csr by sending an invitation via email to the researchers, email addresses being provided by the nsc. the email gave url links and asking the receiver to answer survey questions on the web. there were three follow-up emails for people who did not respond to the survey. . 3. since the nsc is the most important, if not the only, agency that supports academic research, these researchers may be considered constituting almost all of the scholars in taiwan that do research. 4. the items used to gauge about use of srda are as follows. the first number in the parentheses following each response option is the frequency that chose the option. the percentage is the percentage that these people account for of the total number of respondents to the question 1. are you currently an srda member? (n=3019) (1) yes, i am. (go to q2) (365, 12.1%) (2) no. i was before, but the membership is not valid now. (go to q2) (191, 6.3%) (3) no, but i heard of srda and that it provides free access to data. (go to q3) (873, 28.9%) (4) no, i have never heard of srda. (1590, 52.7%) (for those who are or were a member) (n= 556 =365+191) 2. have you ever used data archived in srda for research? (1) yes. (261, 46.9%) (2) no. (go to q2-1) (295, 53.1%) 2-1. what is the reason that you did not use data from srda for research? (n=295, those who answered “no” to q2) (1) i do not have data need. (153, 51.9%) (2) i don’t know how to find the data i need. (50, 16.9%) (3) i cannot find the data for my research. (46, 15.6%) (4) i downloaded some data before but then i found that they did not fit my research purpose. (38, 12.9%) (5) others. (8, 2.7%) (for those who heard of srda before, n=873) 3. have you ever visited srda website? (1) yes, i did. (go to q3-1) (257, 29.4%) (2) no, i did not. (stop) (616, 70.6%) complete analysis reports to include only the research questions, the title of the data, and a short description of the results. the power point file was completed in january of 2013. although we would have preferred to complete it earlier, the delay allowed us to include a brief introduction of the new search function (question bank and concept bank). the file was sent via email to all college professors in humanities and social sciences across the country in march 2013. in addition to informing the professors of srda, we suggested in the cover letter that they show the file to students in class. we hope that this will encourage more students and researchers to use the archived data conclusion since june of 2011, srda has been engaged in such developments to promote the likelihood of data sharing. it has completed a variablelevel search function with variable concepts as a new search option, and compiled a power point file to promote its use. the other two projects, those of abstract revision and concept assigning, are still going on. to write a good abstract turned out to be more difficult than originally thought. as we can give only section titles, rather than major concepts, as the contents of a study, the potential value of abstracts for search may be reduced, unless there is a clear description of the theories or purposes to be tested by the study. nonetheless, as we have a variable-level search function, users probably do not need to rely much on abstracts to find data. the assignment of concepts is the most resource consuming because it requires scholars’ contributions as well as staff members’ continuous checking for consistency. however, the construction of concepts has to continue if it is to contribute to the search efficiency. concepts are valuable not only in searching for data but also in designing questionnaires when it is necessary to include a measurable theoretical concept. during all this time, as an advising researcher to the data division, i have voluntarily involved myself in the developments, sometimes even using my own resource. whereas the division focuses on their routine tasks of data cleaning most of the time, i work closely with two or three of the members, who are more skillful in designing and programming. i regularly enquire about the details and progress of the projects and hold discussions, to make sure the projects are in the right track or to seek solutions to problems. when projects are in a good preliminary shape, ideas are also solicited from the division or the csr, which results in more improvements. such close supervision had helped with the construction of question bank in two different stages. from the experience, i realize the importance of organization and of the leader’s active involvement when the business is just developing. to ask a researcher to oversee an archive will probably lead the archive nowhere because the researcher cannot pay too much attention. therefore, for an archive to develop, it is important to have someone with research experience as its own director. a full-time director will be able to devote all the attention to the archive. the director will be able to not only learn about researchers’ needs, learn about development policies and strategies from other archives, but also carefully plan for projects, and supervise closely the progress of the projects. it is also important to equip the director with a team with various skills. for example, skills such as project designing, computer programming, formal document writing, in addition to data processing skills and data preservation knowledge and techniques, are necessary in the above projects. without these skills, it is very difficult, if not impossible, for an archive to implement improvement projects. however, such people need substantial orientation and an environment that encourages iassist quarterly spring 2012 21 iassist quarterly 3-1. why didn’t you apply for an srda membership? (multiple choice) (1) i do not have data need. (164, 63.8%) (2) the application procedures for a membership require too much. (38, 14.8%) (3) the amount of data archived is not large enough. (23, 8.9%) (4) the information provided on line is not sufficient enough to find data easily. (43, 16.7%) (5) the interface on line is not easy to use to find data. (26, 10.1%) (6) the application procedures for using the data i need (government data or secure data) require too much. (58, 22.6%) (7) i cannot find data that meet my needs. (45, 17.5%) (8) others. (8, 3.1%) 5.http://surveynet.ac.uk/sqb/ rdm@ emory 16 iassist quarterly summer 2012 iassist quarterly abstract academic libraries are increasingly engaging in data curation by providing infrastructure and services to support the management of research data on their campus. efforts to develop these resources can benefit from a greater understanding of the social factors that affect how researchers manage their data during and after their research projects. in particular, the age or amount of experience of researchers is often thought to be an important factor influencing their viewpoints on research data sharing and preservation. in this study, we categorized faculty members who responded to our campus-wide survey on research data management into four ranks—professor, associate professor, assistant professor, and non-tenure track— and analyzed differences in their patterns of survey responses. we found statistically significant differences among faculty ranks in familiarity with funding agency requirements for data management plans, reasons that might prevent data sharing, and interest in potential research data services. these findings reveal key distinctions among different ranks of faculty members in their outlook toward research data management, which can help guide academic librarians and data curation professionals to develop research data services that are tailored to the unique needs of specific populations of researchers. keywords:data curation, research data management, data sharing, researchers, faculty rank, library introduction academic libraries are increasingly providing support for the management and dissemination of research data by offering infrastructure (e.g., insititutional respositories), services (e.g., consultation on data management plans), and education (e.g., best practices in data management) to campus researchers (acrl research planning and review committee, 2012; fearon, et al., 2013; heidorn, 2011; monastersky, 2013). to build research data management support systems that are both effective and desireable by researchers, academic librarians and other information professionals have conducted surveys and interviews with researchers to further understand how they manage data throughout the research lifecycle and their opinions on issues such as data sharing (e.g., bardyn, 2012; jahnke & asher, 2012; scaramozzino et al., 2012; wells parham et al. 2010; westra, 2010; witt et al., 2009). these investigations indicate that researchers exhibit a myriad of approaches to managing their data depending on their discipline, research topic and methodology, source of funding, data privacy concerns, and collaborative networks. the age or amount of experience of researchers is another factor that may influence data management actions and attitudes. within conversations among information professionals, two assumptions are often made: (1) younger researchers have been raised in a culture of greater openness of information and therefore are more willing to share their data, and (2) researchers nearing retirement are concerned about their research legacy and therefore are more eager to preserve their data. these assumptions, however, are primarily based on anecdotal evidence or small numbers of researcher interviews (e.g., office of policy and analysis, 2011). moreover, the few formal investigations on this topic have yielded contradictory results. for instance, kuipers & van der hoeven (2009) differences among faculty ranks in views on research data management by katherine g. akers1 and jennifer doty2 iassist quarterly summer 2012 17 iassist quarterly found that less experienced researchers (< 10 years of experience) were more willing to deposit their data into a disciplinary respository than more experienced researchers (> 20 years of experience). likewise, piwowar (2011) found evidence suggesting that younger researchers are more likely to share their data than older researchers. by contrast, tenopir et al. (2011) found that younger scientists (< 50 years of age) were less likely to make their data available to others without restrictions than older scientists (> 50 years of age), and andreoli-versbach & mueller-langer (2013) found that junior-ranking economy professors were less likely to share their data than full economy professors. therefore, the nature of the relationship between researcher age/experience and tendency to preserve or share their research data remains in question. to futher explore how the age or amount of experience of researchers is related to their views on managing data throughout the research lifecycle, we took faculty rank into account when analyzing the results of our campus-wide survey of researchers’ practices and perspectives on research data management. specifically, we categorized faculty member respondants into four different ranks—professor, associate professor, assistant professor, and non-tenure track—and searched for differences in their patterns of survey responses. methods in the fall of 2012, emory university libraries, in cooperation with the office of institutional research, planning, and effectiveness, administered an online, 13-question survey on research data management practices and perspectives using qualtrics software. a link to the survey was sent via email to all emory university employees with faculty status according to human resource records (n = 5,590). the survey was open for 4 weeks, and three email reminders were sent at 1-week intervals. the survey was initiated by 456 faculty members (~8% response rate). our analysis focused on respondents who answered ‘yes’ to an initial question of whether they conducted research that generated some type of data (e.g., spreadsheets, text, images, videos, audio files, instrument files, photographs, physical samples/ specimens, etc.; n = 330). due to difficulties in equating rank among tenure, clinical, and research tracks, faculty members with clinical or research track designations were not included in the analysis. the remaining faculty members (n = 210) were divided into four groups based on human resource records: professor (professor, professor emeritus, or dean), associate professor, assistant professor, or non-tenure track (instructor, lecturer, visiting scholar, or adjunct professor). these faculty members were predominately based in the emory college of arts and sciences (50%), with others from the school of medicine (25%), rollins school of public health (10%), goizueta business school (4%), oxford college (4%), candler school of theology (4%), nell hodgson woodruff school of nursing (3%), and school of law (1%). differences in survey responses among the four ranks of faculty members were evaluated using chi-square (χ2) tests. statistical significance was set at p < 0.05. data are shown only for survey responses for which there were statistically significant differences among ranks. complete survey results, including differences among arts & humanities, social science, basic science, and medical science domains, were previously reported (akers & doty, in press). results data management planning we found no significant variations among different ranks of faculty members in the amount of research data they were storing or their methods for data storage and back-up (e.g., computer hard drive, external hard drive, instrument hard drive, university server, internet-based storage, lab notebooks, discs/tapes). however, we did find variations among faculty ranks in their familiarity with federal funding agency requirements (e.g., national science foundation (nsf), national institutes of health (nih), national endowment for the humanities (neh)) for data management or data sharing plans as components of grant applications (χ2 (3, n = 210) = 13.5, p = 0.004; figure 1). the majority of full and assistant professors stated that they were either somewhat or very familiar with data management plans, and over half of associate professors also expressed familiarity with these requirements. by contrast, most non-tenure track faculty members were not familiar with data management plan requirements, which may reflect a greater focus on teaching and less reliance on research grants. data sharing faculty rank did not predict faculty members’ willingness to share their research data with other people (e.g., researchers working on project, researchers outside of project, funders, instructors, general public) or their method of sharing research data (e.g., email upon request, supplementary material linked to journal article, data repository, university or personal website). figure 1 18 iassist quarterly summer 2012 iassist quarterly however, different ranks of faculty members expressed different opinions on why they might not share their data. specifically, full and associate professors were more likely than assistant professors and non-tenure track faculty members to state that it takes too much time or effort to share their research data (χ2 (3, n = 199) = 10.1, p = 0.018; figure 2). this finding may reflect that seniorranking faculty members may simply feel they are too busy to organize, document, and compile their data into shareable data packages that can be understood and used by others. alternately, junior-ranking faculty members may feel that sharing their data with others is an expected part of the research process and thus may not perceive preparation for data sharing as an imposition on their time. there were no differences among faculty ranks in other reasons that might prevent data sharing, including having data that contain private or patentable information, having data that require restricted access, fear of not getting credit for their data, fear of possible misinterpretation or misuse of their data, or belief that their data are of little use to others. different ranks of faculty were also equally likely to deposit their data in data repositories or express familiarity with data documentation and metadata. interest in data services in the final survey question, we offered a list of ten potential research data services and asked faculty to select which services they would use if available. the service garnering the most interest was faculty workshops on general data management. this service was desired by non-tenure track faculty members more than by assistant, associate, or full professors (χ2 (3, n = 191) = 11.6, p = 0.009; figure 3). there were no rank-related differences in interest for the other potential services, including assistance preparing data management plans, consultation on data confidentiality and/or legal issues, personalized consultation on research data management for specific researchers or research groups, an institutional repository for research data, assistance with data documentation or metadata creation, research data management workshops for trainees (i.e., graduate students or postdocs), digitization of physical research materials, assistance identifying appropriate disciplinary data repositories, or methods for data citation. discussion it is often assumed that younger researchers are more supportive of open data and therefore more likely to share their research data with others via websites or data repositories/ archives (e.g., johnson, 2008; lin, 2013; boulton, 2013). however, empirical studies have not consistently provided support for this assumption. although evidence from kuipers & van der hoeven (2009) and piwowar (2011) suggests that younger researchers are indeed more willing to share or archive their data than older researchers, our survey failed to find differences among ranks of faculty members in their willingness to share research data or their preferred method of data sharing. moreover, tenopir et al. (2011) and andreoli-versbach & mueller-langer (2013) found that younger researchers were less likely to share their data than older researchers. therefore, younger age or less research experience may not always be a predictor of data sharing. rather than a simple correlation between researcher age and willingness to share data, findings by tenopir et al. (2011) suggest that the situation is more complex. their survey revealed that younger scientists were less likely than older scientists to place their data in a central repository without restrictions. however, younger scientists were slightly more likely than older scientists to make their data available if they could place conditions on data figure 2 figure 3 iassist quarterly summer 2012 19 iassist quarterly re-use, such as requiring legal permission for re-use of their data or the receipt of a complete list of products that make use of their data. andreoli-versbach & mueller-langer (2013) speculate that young researchers might be hesitant to openly share their data because this could enable other researchers to use their data before they can be fully exploited for additional publications. in other words, young faculty members who have not yet secured tenure may act more competively than their tenured counterparts, choosing to withhold their research data or impose restrictions on data re-use. as we found no rank-related differences in faculty members’ tendency to state that they might not share their data due to fear of not getting credit, our results do not directly support this possibility. nevertheless, junior-ranking faculty members may be the ideal target population for outreach on ways of turning datasets into citeable outputs of scholarly research to increase personal research impact, including assigning digital object identifiers (dois) to datasets, depositing data into disciplinary or institutional repositories, or publishing data papers. younger researchers might also be particularly receptive to evidence indicating that openly sharing research data increases the citation rate of associated journal articles (bueno de mesquita et al., 2003; dorch, 2012; henneken & accomazzi, 2011; ioannidis et al., 2009; piwowar et al., 2007; piwowar & vision, 2013; sears, 2011). our survey did not contain questions about data re-use, but previous studies indicate that the age of researchers may be an important indicator of their likelihood to re-use or re-purpose other people’s data. kuipers & van der hoeven (2009) found that less experienced researchers are more eager to re-use data from other disciplines than more experienced researchers. similarly, tenopir et al., (2011) found that younger scientists are more likely to consider lack of access to data as a barrier to scientific progress that has restricted their ability to answer research questions. these findings underscore our suggestion that junior-ranking faculty members, in particular, could benefit from learning about ways to publically disseminate and thereby open their research data for re-use. also, younger researchers may be more interested in receiving assistance with discovering and accessing pre-existing datasets. although we observed no differences among faculty ranks in willingness to share data, we found that senior-ranking faculty members were more likely to state that they might not share their research data due to the amount of time and effort involved. indeed, tenopir et al. (2011) found that insufficient time was the top reason that scientists did not make their data available to others, and others have also recognized this potential barrier to data sharing (cragin et al., 2010; peters & riley dryden, 2011; williams, 2013). however, vickers (2006) makes the case that a fundamental responsibility of data producers is to create clean, accurate, and well-annotated datasets, after which it should take little time to delete extraneous variables, remove personal identifiers, convert files into accessible formats, and briefly describe the dataset contents. therefore, objections to data sharing on the premise that it takes too much time suggests that researchers do not always develop clean and well-annotated datasets (savage & vickers, 2009). we speculate that this is because researchers often have no clear incentive to do so; instead, they may be motivated only to develop datasets that are sufficient to support their own analyses for a particular project and not to invest time to make their datasets understandable to other researchers or to themselves at future dates. therefore, our results suggest that senior-ranking faculty members, in particular, may benefit from increased university investment in the management and dissemination of research data, including the creation of electronic research data management systems that could automatically organize data, generate metadata and documentation files, and push data packages into open repositories after project completion. finally, we found that non-tenure track faculty members were least familiar with funding agency requirements for data management plans but most interested in taking advantage of faculty workshops on general data management practices. at our university, faculty members are hired onto one of three different tracks: tenure track, research track, or clinical track. due to difficulties in equating rank status across tracks, however, we removed research and clinical track faculty respondents from our analysis, meaning that the professional focus of the remaining non-tenure track faculty members was more likely to be teaching than other types of scholarly activities such as research. therefore, their unfamiliarity with data management plans may be a result of their lack of dependence on external funding for research projects. nevertheless, these faculty members also answered ‘yes’ to a question of whether they performed research that generated some type of data, indicating that they were indeed engaged in research to at least some extent. as such, our finding that nontenure track faculty members expressed the highest desire for learning about “best practices” in research data management is very interesting and suggests that this faculty contingent, which is continuing to increase in size across academia (curtis & thornton, 2013), may be an overlooked population of researchers that might welcome greater outreach from academic librarians and other data curation professionals. acknowledgments we thank vincent carter for his critical role in creating and administering the survey. we also thank members of the emory university libraries research data management team for their feedback on survey design and analysis references acrl research planning and review committee. (2012). 2012 top ten trends in academic libraries: a review of the trends and issues affecting academic libraries in higher education. college & research libraries news, 73, 311-320. retrieved from http://crln.acrl.org/content/73/6/311.full akers, k.g. & doty, j. (in press). research data management practices and perspectives: differences among the arts and humanities, social sciences, medical sciences, and basic sciences. international journal of digital curation. andreoli versbach, p. & mueller-langer, f. (2013). open access to data: an ideal professed but not practiced. ratswd working paper series no. 215; max planck institute for intellectual property & competition law research paper no. 13-07. retrieved from the social science research network: http://papers.ssrn.com/sol3/papers. cfm?abstract_id=2224146 bardyn, t.p., resnick, t., & camina, s.k. (2012). translational researchers’ perceptions of data management practices and data curation needs: findings from a focus group in an academic health sciences library. journal of web librarianship, 6, 274-287. retrieved from http://www. tandfonline.com/doi/abs/10.1080/19322909.2012.730375?journalco de=wjwl20#.uypahlxcaso boulton, g. (2013). open data: why and how it matters to the future. knowledge exchange berlin, april 2013. retrieved from http://webcache.googleusercontent.com/ search?q=cache:vhc50qd1ubcj:www.knowledge-exchange.info/ 20 iassist quarterly summer 2012 iassist quarterly admin/public/dwsdownload.aspx%3ffile%3d%252ffiles%252ffile r%252fdownloads%252fprimary%2bresearch%2bdata%252fmakin g%2bdata%2bcount%2bworkshop%252fboulton_presentation%2b%2bberlin.pdf+&cd=2&hl=en&ct=clnk&gl=us bueno de mesquita, b., gleditsch, n.p., james, p., king, g., metelits, c., ray, j.l., russett, b., strand, h., & valeriano, b. (2003). symposium on replication in international studies research. international studies perspective, 4, 72-107. retrieved from http://onlinelibrary.wiley.com/ doi/10.1111/1528-3577.04105/full# curtis, j.w. & thornton, s. (2013). here’s the news: the annual report on the economic status of the profession, 2013-2013. american association of university professors. retrieved from http://www.aaup.org/report/ heres-news-annual-report-economic-status-profession-2012-13 cragin, m.h., palmer, c.l., carlson, j.r., & witt, m. (2010). data sharing, small science and institutional repositories. philosophical transactions of the royal society a, 368, 4023-4038. retrieved from http://rsta.royalsocietypublishing.org/content/368/1926/4023 dorch, b. (2012). on the citation advantage of linking to data. hprints & humanities, hprints-00714715, version 2. retrieved from http:// hprints.org/hprints-00714715/ fearon, d., gunia, b., sherry, l., pralle, b.e., & sallans, a.l. (2013). spec kit 334: research data management services. association of research libraries. retrieved from http://publications.arl.org/ research-data-management-services-spec-kit-334/ heidorn, p.b. (2011). the emerging role of libraries in data curation and e-science. journal of library administration, 51, 662-672; retrieved from http://www.tandfonline.com/doi/abs/10.1080/01930826.2011. 601269#.uyvnl7wg2hs henneken, e.a. & accomazzi, a. (2011). linking to data – effect on citation rates in astronomy. arxiv:1111.3618. retrieved from http://arxiv. org/abs/1111.3618 ioannidis, j.p., allison, d.b., ball, c.a., coulibaly, i., cui, x., cuthane, a.c., falchi, m., furlanello, c., game, l., jurman, g., mangion, j., mehta, t., nitzberg, m., page, g.p., petretto, e., & van noort, v. (2009). repeatability of published microarray gene expression analyses. nature genetics, 41, 149-155. retrieved from http://www.nature. com/ng/journal/v41/n2/full/ng.295.html jahnke, l.m. & asher, a. (2012). the problem of data: data management and curation practices among university researchers. council on library and information resources. retrieved from: http://www.clir. org/pubs/reports/pub154/problem-of-data johnson, c.y. (2008) out in the open: some scientists sharing results. boston globe, august 21, 2008. retrieved from http://www.boston.com/news/local/articles/2008/08/21/ out_in_the_open_some_scientists_sharing_results/ kuipers, t. & van der hoeven, j. (2009). insight into issues of permanent access to the records of science in europe. parse. insight. retrieved from http://www.parse-insight.eu/downloads/ parse-insight_d3-4_surveyreport_final_hq.pdf lin, t. (2013) imagining data without division. quanta magazine, september 30, 2013. retrieved from: https://www.simonsfoundation. org/quanta/20130930-imagining-data-without-division/ monastersky, r. (2013). publishing frontiers: the library reboot. nature, 495, 430-432. retrieved from http://www.nature.com/news/ publishing-frontiers-the-library-reboot-1.12664 office of policy and analysis. (2011). sharing smithsonian digital scientific research data from biology. smithsonian institution. retrieved from http://www.si.edu/content/opanda/docs/rpts2011/11.03. datasharing.final.pdf o’reilly, k., johnson, j., sanborn, g. (2012). improving university research value: a case study. sage open, 2, 1-13. retrieved from http://sgo.sagepub.com/content/2/3/2158244012452576.full. pdf+html peters, c. & riley dryden, a. (2011). assessing the academic library’s role in campus-wide research data management: a first step at the university of houston. science & technology libraries, 30, 387-403. retrieved from http://www.tandfonline.com/doi/pdf/10.1080/01942 62x.2011.626340. piwowar, h.a., day, r.s., & fridsma, d.b. (2007). sharing detailed research data is associated with increased citation rate. plos one, 2, e308. retrieved from http://www.plosone.org/article/ info:doi%2f10.1371%2fjournal.pone.0000308 piwowar, h. & vision, t.j. (2013). data reuse and the open data citation advantage. peerj preprints, 1, e1v1. retrieved from http://dx.doi. org/10.7287/peerj.preprints.1v1 piwowar, h.a. (2011). who shares? who doesn’t? factors associated with openly archiving raw research data. plos one, 6, e18657. retrieved from http://www.plosone.org/article/info:doi/10.1371/ journal.pone.0018657 scaramozzino, j.m., ramirez, m.l., & mcgaughy, k.j. (2012). a study of faculty data curation behaviors and attitudes at a teaching-centered university. college & research libraries, 73, 349-365. retrieved from http://crl.acrl.org/content/73/4/349.full.pdf+html savage, c.j. & vickers, a.j. (2009). empirical study of data sharing by authors publishing in plos journals. plos one, 4, e7078. retrieved from http://www.plosone.org/article/info:doi/10.1371/journal. pone.0007078. sears, j.r. (2011). data sharing effect on article citation rate in paleoceanography. american geophysical union, abstract# in53b-1628. retrieved from http://adsabs.harvard.edu/ abs/2011agufmin53b1628s. tenopir, c., allard, s., douglass, k., aydinoglu, a.u., wu, l., read, e., manoff, m., & frame, m. (2011). data sharing by scientists: practices and perceptions. plos one, 6. e21101; retrieved from http:// www.plosone.org/article/info%3adoi%2f10.1371%2fjournal. pone.0021101 vickers, a.j. (2006). whose data set is it anyway? sharing raw data from randomized trials. trials, 7, 15. retrieved from http://www.trialsjournal.com/content/7/1/15. wells parham, s., bodnar, j., & fuchs, s. (2010). supporting tomorrow’s research: assessing faculty data curation needs at georgia tech. college & research libraries news, 73, 10-13. retrieved from http:// crln.acrl.org/content/73/1/10.full westra, b. (2010). data services for the sciences: a needs assessment. ariadne, 64. retrieved from http://www.ariadne.ac.uk/issue64/westra williams, s. (2013). data sharing interviews with crop sciences faculty: why they share data and how the library can help. issues in science and technology librarianship, spring 2013. retrieved from http:// www.istl.org/13-spring/refereed2.html#8 witt, m., carlson, j., brandt, d.s., & cragin, m.h. (2009). constructing data curation profiles. international journal of digital curation, 4, 93-101. retrieved from http://www.ijdc.net/index.php/ijdc/article/view/137 notes 1. katherine g. akers was the escience librarian and a council on library and information resources (clir) postdoctoral fellow at emory university libraries. she is now the escience librarian and a clir postdoctoral fellow at the university of michigan libraries. she can be reached by email: kgakers@umich.edu. 2. jennifer doty is the data management specialist at emory university libraries and can be reached by email: jennifer.doty@emory.edu. 60 iassist quarterly 2010 / 2011 iassist quarterly currently there is no common practice of archiving data abstract the authors briefly describe the current situation in the field of quantitative research in belarus and analyse the existing qualitative resources in this country. the major finding of the paper is that sociologists in belarus have not yet adopted a professional culture of sharing research information. therefore, the resources of most surveys are not available for observation and use for anyone else except for the “owner” of this information. the authors conclude that a specially elaborated strategy for involving belarus in european-wide research activities and the inclusion of belarusian scholars in all-european professional associations is needed: it may help to bring the professional culture of sharing data into professional practice in belarus, and make qualitative research more transparent to all scholars and the public in belarus. keywords: data resources, data archiving, qualitative data, longitudinal and repeat research, belarus introduction the republic of belarus is a typical post-soviet state that inherited a tendency for soviet sociology to be non-transparent and with poor public funding for quantitative research. being an independent state for almost 20 years (since 1991) belarus has still not joined any professional sociological structures in europe in order to help its own social scholars to adopt and follow european research standards. in the case of quantitative research it means that belarusian sociologists are not members of any groups or organisations that can help to inculcate such norms of professional work. in belarus there are no social research data archives available for external users. each research organisation archives its own data and materials for internal re-use only. usually the head of this organisation has the right to decide who can use the data, and for which purposes. research culture is therefore rather closed. data sharing is not prevalent as there are no traditions or official (or professional) policies to propagate research information – whether this relates to the research topics, data qualitative and qualitative longitudinal research and resources in belarus by prof. larissa titarenko and asst. prof. olga tereschenko1 iassist quarterly 2010 / 2011 61 iassist quarterly received or potential users of the data. regardless of the fact that the generation of such data is publically funded (i.e., should serve the needs of a whole society) such data are viewed as the intellectual property of particular organisations or even individuals. therefore, as a rule, nobody, except the putative owner of the data can check this information and use it for any purposes. secondary analysis of the data collected by someone else is very rare: it can take place only in the case of personal negotiations between the owner of received data (research organisation or a researcher) and the potential user (also research organisation, researcher or other person). however, such negotiations take place only in relation to quantitative data. as for qualitative data, they are considered exclusively “personal” and not available for any “external” re-use. additionally, qualitative data are often kept in a format that makes their re-use impossible by anybody else except for the person who collected them. for this reason, university staff do not even ask each other to share data within their own department in order to use such data for teaching sociology. therefore, students often know qualitative methods more “in theory” than in practice. currently, there is no common practice of archiving data in belarus. the raw data sets were sometimes destroyed after the completion of the research report because the research organisation simply did not have space to keep them. this reason dwindled during last years. another reason for such action can be “professional security”. some heads of sociological firms or other organisations are afraid of any “loss of data” and their external use without his or her permission, especially data related to the political sphere or research that was “ordered” by the state. under the current conditions, the loss of such data or their re-use by someone else can be punished. last but not least, a lot of quantitative data were lost in the early 1990s, when belarusian sociologists started to use personal computers and didn’t know how to shift the “old” data to the new electronic platforms or din’t have resources for this because of the deepest economic crisis. national policies on research data archiving and sharing the national policies on research data archiving and sharing apply mainly to the national statistical committee of the republic of belarus (2011). the national statistical committee collects the current statistics (according to the approved list of the indicators and spheres to research), conducts some special demographic, social and economic surveys, and publishes the main results in statistical books. interested organisations or private persons may buy some unpublished results, but they never have access to the raw data and full statistics. the national statistical committee has its own archive of population and sample surveys. however, it does not conduct qualitative research. the belarusian government provides some financial support to a number of research organisations such as: • the science and research economic institute that belongs to the ministry of economics2 • the science and research institute for labour and social defense that belongs to the ministry of labour and social defense3 • minsk science and research institute of social-economic problems of the minsk-city government 4 • the institute of sociology of the national academy of sciences.5 all these institutions conduct research for the state (a particular ministry or department or the academy of sciences) and/or their private customers (in the case of commercial research). they archive the collected data and raw research materials, and – again – do not share their data with any researchers from outside (regardless of the affiliation with the state institution or private firm). similar practices are followed by all other research organisations, regardless of their ownership (i.e. state-owned or private-owned). currently, there is no rule or habit to keep, archive and share the data that is gathered with state funding. therefore, no organisation does it under its own initiative. qualitative data sources there is no infrastructure for archiving and sharing qualitative data in belarus on either the national or regional levels. therefore, in this report we would like to review the main qualitative sources that may be potentially available. • the majority of qualitative research projects are carried out by the marketing agencies, whose customers are usually private companies (quite often international and transnational trade companies). • the main fields of research are the following: marketing, consumer behavior, and media functioning/ effectiveness. • the main methods of research are focus groups and in-depth interviews. • the methods used in organisational and corporate research are non-structured and/or participant observation. • in the field of political research the prevailing method is in-depth (deep) interview. all the data gathered for this marketing and political research is considered to be the property of the customers. besides, most of the marketing agencies are included in the international marketing networks and can’t make their own decisions concerning data sharing. they follow some corporate ethics within their network. however, in some cases the state institutions, international organisations/ foundations (unicef, undp, unesco, etc.), and non-government organisations (especially those with international links) can commission qualitative research and pay for the information. we consider this kind of research as the most appropriate for data sharing and archiving, because these international organisations are familiar with the legal rules on archiving and sharing in europe and world-wide. therefore, they can provide the data—probably—after some negotiations or after the official agreement with some belarusian organisations or the government. the main potentially available qualitative sources are: institute of sociology, the national academy of sciences (n.d). the state-funded institute of sociology conducts quarterly the so-called quantitative repeat cross-sectional study. usually they use a mixed methodology: population survey and in-depth interview. their research includes the following topics (among others): (1) “social and political situation in the republic of belarus” (annually). 62 iassist quarterly 2010 / 2011 iassist quarterly the commissioner is the national government. methodology: crosssectional sample survey. however, some additional information for this study is collected with qualitative methods. (2) “means of private ownership, wealth, and prosperity” (2007) method used: case study in a rural settlement. sociological and political research center at the belarusian state university (n.d). according to their own estimations, qualitative research comprises up to 20% of all research done by this centre. marketing research prevails. however, when the national elections take place, the centre is always commissioned by the political parties to organise qualitative research. they have their archive of qualitative data (papers, typed records, and films), although currently they are not ready to share their data. laboratory for axiological research novak (n.d). the laboratory regularly conducts qualitative studies for commercial organisations, state bodies and international organisations and foundations. the most used methods are focus groups and in-depth interviews. the main fields are marketing, media, and social studies. the topics of their social research in 2008–2009 were the following: • “social contract” (method: focus groups, customer the institute of belarusian studies, vilnius) • “aids” (method: in-depth interviews; customer: the ministry of health of belarus and the world health organisation) • “integration processes between russia and belarus” (method: focus groups, customer eurasia foundation) • “children with special needs” (methods: focus groups and in-depth interviews, customer unesco faculty of philosophy and social sciences, belarusian state university (n.d) the departments of sociology, social psychology, and information and communication affiliated with this faculty, regularly conduct some small-scale qualitative research for their own needs – primarily for teaching. methods used: in-depth interviews, observations, case study, and focus groups. however, their staff is often involved in some out-ofuniversity consultancy. the research information is usually available for the department’s staff (only those involved in the collection of data) and for students for writing research papers, diplomas, dissertations, etc. additionally, there are some phd studies conducted in belarus on the basis of qualitative research and data collected by the authors: • oksana shelest (institute of sociology, national academy of sciences) “arising of new kinds of religiosity” (2005–2006), method in-depth interviews. • ludmila yakusheva (bsu, department of sociology) “women trafficking in belarus” (2005-2006), methods: life histories and in-depth interviews concerning trafficking problems. • julia lahvich (bsu, department of psychology) “social and psychological factors of successful adoption” (2008-2009), mixed methodology: sample survey, sentences completion method, narrative interviews, in-depth interviews. longitudinal data sources there is no infrastructure for the management and re-use of longitudinal data in belarus. nevertheless quantitative panel research studies are rather widespread, especially in economics and marketing. for instance: • the households panel of the national statistical committee of the republic of belarus. • the media audience panel and the customers panel of the laboratory for axiological research novak. • the stores panel of the a.c. nielsen belarusian branch, etc. in 1983–1998 the four-waves sociological longitudinal project “paths of generation” related to the high school graduates of 1983 has been conducted by the institute of sociology (1st and 2nd waves) and the faculty of philosophy and social sciences, bsu (3rd and 4th waves), under supervision of elena borkowskaya. overall, it was a part of a bigger soviet (after 1991 – a cross-cultural, international) research study in which three baltic republics were involved, together with belarus. the panel belarusian-american thyroid cancer project has been conducted regularly since 1999. it was devoted to the population of belarus exposed to ionizing radiation due to the chernobyl accident (1986). repeated qualitative and mixed methods research the only repeated qualitative study we know is “gender aspects of violence in refugees’ midst” conducted by the belarusian young christian women association in 2004 and 2009. the method of this research was focus groups. the commissioner: the office of the un high commissioner for refugees (unhcr). the same organisation ran a research project on “home violence against women” in 2010 using qualitative methods (50 in-depth interviews). this study was funded by international sponsors with the intent to repeat it in a few years time. as for mixed methods research, they are also not common and depend on the public needs and financial opportunities. thus, in 2001, the first wave of the international project (12 countries involved) organized by the minsk city narcological dispensary and funded by the world health organisation, was arranged in minsk. the topic was called “expressestimation intravenous drugs consumption in minsk”: it included of 2 surveys of drug users, 7 focus groups with drug users, 3 structured observations, and a series of in-depth interviews with experts in this sphere (medical doctors, policemen, ngo members and other specialists). the second wave (in 2003) included only two qualitative surveys (400 drug-users). the third wave of research is expected in the coming years, also with mixed methodology. additionally, there is publicly available information about some repeat research studies run by the institute of sociology: • “civil society in perception of belarusians”, repeat study of the same object and design: the first wave was in 2005; the second wave was carried out in 2009. mixed methodology: survey and in-depth interviews with selected respondents. • “belarusian national identity”, repeat study of the same object: the first wave was in 2004; the second wave was in 2008, sentences completion method. development planning priorities for the next three years are: • creation of the first open data archive at the belarusian state university (faculty of philosophy and social sciences). • drawing up some rules and/or procedures to enable archived data and materials to be shared. iassist quarterly 2010 / 2011 63 iassist quarterly • promotion of the idea of data sharing in the professional communities of researchers and data owners. the main barriers seem to be the following: lack of experience, keeping the old (soviet) tradition, conservative culture that forbids data sharing, archiving and re-using; and deficiencies in information about data sources and owners. the existing foreign archiving organisations could be of some assistance in sharing experience in data archiving, procedures of data sharing, and ways of persuading research organisations and owners to give data for archiving and sharing. they could also help in establishing contacts with such international organisations as unicef, undp, unesco, which would enable work in belarus to develop in tandem with best practice elsewhere and so that archivists can learn from the experience of others. conclusion in common with all other european countries, belarus badly needs research information of high quality about the social, political, and cultural processes that take place in the country. for these purposes the scientific community has to provide the state branches and the research organisations with objective social information. however, the republic of belarus is not rich enough to provide its researchers with the modern instruments that are necessary for collecting, archiving and sharing qualitative information in the social sciences. additionally, since the country was politically isolated for almost fifteen years, the community of sociologists did not have a real possibility to join the international professional organisations or to develop their work with the support of such organisations. therefore, there was no professional demand to adopt the existing professional rules and professional culture of sharing research information. the existing situation in the social sciences is characterized by the lack of national public and/or scientific data archives (both quantitative and qualitative), the low level of awareness of the need for sharing the research data, their professional control (checking their quality), and the potential for data re-use. there are some researchers and even centers that try to follow the international standards in data collecting; however, they do not share their data. overall, belarus badly needs to start the creation of data archives and join the european community of qualitative archives of social data. however, currently, the country does not have enough experience, resources and professional information to start the process. the help of the international community of social scholars and international organisations is badly needed. in march 2009 dr. kevin schürer, director of the uk data archive, paid a short visit to minsk as part of cessda project activities. he met the leading persons at the state committee on science and technology – the so called “ministry of science” and had a talk with the staff. then he visited the faculty of philosophy and social sciences, bsu where he discussed the possibility of starting the archive (both qualitative and quantitative) at this university. dr. schürer promised to provide some necessary information and organisation support to bsu in organizing the data archive. however, there was no more information from his side as it was necessary to await funding decisions for cessda. currently the faculty expects that some information on how to start the data archive will be provided from gesis and cessda. references belarusian-american thyroid cancer project. (1999). [online]. available at: http://www.minzdrav.by/en/tools/adm/trap.php?id=174. belarusian young christian women association (2004, 2009). ‘gender aspects of violence in refugees’ midst’. office of the un high commissioner for refugees (unhcr). [online]. available at: http:// bywca.iatp.by/ belarusian young christian women association. (2010). ‘home violence against women’. [online]. available at: http://bywca.iatp.by/ borkowskaya, e. (1983-1999). ‘paths of generation’. longitudinal research project carried out by the institute of sociology (1st and 2nd waves) and the faculty of philosophy and social sciences, bsu (3rd and 4th waves). faculty of philosophy and social sciences: belarusian state university. (n.d.). [online]. available at: http://www.ffsn.bsu.by institute of sociology. (2005-2009). ‘civil society in perception of belarusians’. repeat research study: institute of sociology, belarus. institute of sociology. (2004-2008). ‘belarusian national identity’. repeat research study: institute of sociology, belarus. institute of sociology, belarus. (n.d). [online]. available at: http://www. sociologyby.org laboratory for axiological research. (n.d). [online]. available at: http:// www.novak.by/. lahvich, j. (2008-2009). ‘social and psychological factors of successful adoption’. unpublished phd research: bsu, department of psychology minsk city narcological dispensary. (2001). express-estimation intravenous drugs consumption in minsk. who. national statistical committee of the republic of belarus. (2011). [online]. available at: http://belstat.gov.by/homep/en/main.html shelest, o (2005-2006). ‘arising of new kinds of religiosity’. unpublished phd research: institute of sociology, national academy of sciences. sociological and policy research center at the belarusian state university. (n.d). [online]. available at: http://www.bsu.by/main. asp?id1=20&id2=820 yakusheva, l. (2005-2006). ‘women trafficking in belarus’. unpublished phd research: bsu, department of sociology notes 1. prof. larissa titarenko, larisa166@mail.ru faculty of philosophy and social sciences, belarusian state university http://www.ffsn.bsu.by ass. prof. olga tereschenko, oteresch@tut.by faculty of philosophy and social sciences, belarusian state university http://www.ffsn.bsu.by 2 for further information see: http://w3.economy.gov.by/work_web/ niei.nsf/ 3. see: http://www.instlab.org/ 4. see: http://www.minsk.gov.by/cgi-bin/ org_ps.pl?mode=ind&k_org=3636 5. see: http://sociologyby.org/structure.html iassist quarterly 3 techniques for secondary analysis: unfolding analysis of "pick k/n" and "pick any/n" data by wijbrandt h. van schuur' faculty of social sciences university of groningen oude boteringestraat 23 9712 gc groningen the netherlands introduction in survey research we regularly encounter the following type of question: "which of these stimuli do you prefer most? which of the remaining ones do you now prefer most?" etcetera. sometimes a full rank order of preferences is asked in this way, more often only a partial rank order is obtained. 'paper prepared for presentation at the ifdo/iassist conference, workshop on techniques for secondary analysis, amsterdam, may 20 23, 1985. comments are welcomed by the author. sometimes the question asked is only: "which k of these n stimuli do you prefer most?" or, even more generally, "which of these n stimuli do you prefer?" such questions can be referred to as 'ramk n/n', 'rank k/n', 'pick k/n' and 'pick any/n' data, respectively. stimuli may be political parties, candidates, career possibilities, or brand names of some consumer good. rather than asking about 'preference', the questions may also be phrased in terms of other evaluative concepts, such as 'sympathy' or 'importance'. in this paper i will be concerned with analyzing data of the form 'pick k/n' or 'pick any/n'. generally these data types are difficult to analyze. often responses to such data are only reported in the form of frequency distributions of the number of times a stimulus is mentioned as most preferred, second most preferred, etcetera. trying to find structure in these responses with the help of standard techniques, such as factor analysis or cumulative scaling, is not possible either because no full set of responses to all stimuli is available, or because the responses given are not independent it is then difficult to determine whether or not all responses given were based on the same underlying criterion. in this paper an analysis technique is presented that allows one to look for structure in the responses to 'pick k/n' or 'pick any/n' questions. since complete or partial rank orders can always be recoded to the 'pick k/n' form, and since survey questions with independent responses, such as five-point likert items, can be recoded to the 'pick any/n' form, the type of data analysis presented here can have very general applicatioa the data analysis technique presented here is a dichotomous version of the unfolding model, proposed by coombs (1950, 1964), as 'parallelogram analysis'. it differs from coombs' original proposal in the following ways: the technique proposed here allows for some error (i.e., it conforms to a stochastic model), and it is an exploratory technique to search for summer 1986 4 iassist quarterly maximal subsets of stimuli that can be represented in a unidimensional unfolding scale. in both these aspects the 'parallelogram analysis' model proposed here resembles the stochastic unidimensional aunulative scaling technique developed by mc^en (1971). the reader should be warned that the technique presented here is not an all-purpose technique for analyzing 'pick k/n' or 'pick any/n' data, but only for those types of data which can be expected to conform to the unfolding model! the perfect unidimensional unfolding model for complete rank orders of preference in this section i wiu first simimarize some basic ideas behind unfolding analysis, by using an example from meerling (1981). in an investigation by ritzema and van de kloot, preference rank orders were collected for the following statements: : people can be changed in any conceivable direction, provided that the environment is manipulated in the proper way (o = omgeving, envirormient); 1 : the major condition for people to change is for them to have a clear understanding of their situation (i = inzicht, understanding); e : behaviour is determined much more strongly by emotions than by rational considerations (e = emoties, emotions); a : inborn characteristics determine to a large extent what kind of person someone becomes (a = aangeboren, inborn). these four statements were shown to psychologist colleagues, and the following six types of preference rank orders were foimd: oiea. loea. leoa, hoa. eaio, and aho. in applying the unfolding model we assume that there is a latent dimension on which each of these statements can be represented. meerling suggests for these statements and these preference rank orders that a 'nurture-nature' dimension may be appropriate, in which the statements are arranged in the order oiea. when the location of each of the statements on this dimension is established, the dimension can be divided into two areas for each pair of statements (i,j): the first area, in which the first statement is preferred over the second, and the second area, in which the second statement is preferred over the first the boundary between these two areas lies in the middle between these two stimuli, and is called the 'midpoint of the pair of stimuli', m(ij). this midpoint allows us to locate individuals who give their preference rank order along this dimension. an individual, pi, who prefers statement o to statement i will be located to the left of midpoint m(oi), whereas another individual, p2, who prefers statement i to statement o, will be located to the right of that midpoint (see figure 1) figure 1 midpoint m(oi) divides the dimension areas into two pi ; p2 ! 1 m(oi) i the four statements, together, have six midpoints. these divide the dimension into seven areas, the areas that are separated by the midpoints. each of these areas is characterized by a special preference rank order, and is called an 'isotonic region', (see figure 2) summer 1966 iassist quarterly 5 figure 2 4 stimuli, 6 midpoints, and 7 isotonic areas or subject types a subject is usually represented on the scale by a single point, called his 'ideal point'. the preference order of the subject is called his 'individual scale', or 'i-scale' for shorl the representation of all subjects and all stimuli jointly on the same dimension is called the 'joint-scale', or 'j-scale' for short the i-scale gives the order of the stimuh in terms of their distance from the ideal point of the individual. in other words: the i-scale has to be 'unfolded' at the ideal points to produce the j-scale. unfolding analysis is designed to find a joint representation of stimuli and subjects in one dimension, that is, to find a unidimensional j-scale on the basis of the preference rank orders of the individual 1-scales. finding a j-scale brings us two things. the first is an unfoldable order of the stimuli which can generally be used to infer the criterion used by the subjects in determining their preference order (e.g., the nurture-nature criterion). secondly, having a j-scale allows us to combine a subject's answers to the n survey questions in a single rank order which can be used to measure the preference of the subject in terms of his ideal point on the criterion dimensioil by measuring a subject's preference in this way we can create a new variable which can be related to other characteristics of the subject the purpose of creating such a new variable is to try to explain why people differ in their preferences, or to explain other attitudes or behaviours on the basis of scale values on the j-scale. if we have perfect data, such as we usually find in textbooks on scaling (and by perfect data i mean i-scales that can be perfectly represented in a unidimensional unfolding scale) it is no problem to find the j-scale that represents the i-scales. problems only arise when the data are not perfect, which is in most cases. the major reason why the unfolding technique has so far been relatively unpopular and why it has as yet not been incorporated into most standard statistical packages, is that up to now we have not been able to imfold imperfea data in a satisfactory way. if we could find a usable unfolding technique, interest in it should be great, since the model is plausible, and there is a great deal of interest in measuring the preferences of subjects. discussion of some alternative proposals for unfolding models before introducing my own model, i will first consider five strategies that have been developed in the literature and which attempt to find useful and interpretable unfolding results. these strategies are all derived from a description of the ideal type of unfolding analysis, namely the perfect representation of a complete rank order of preferences in a unidimensional space, in which all stimuli and all individuals can be represented. these five strategies are: 1. analyze the i-scales after they have been dichotomized into the k most preferred and n-k least preferred stimuli; 2. relax the criterion of perfect representation to allow stochastic representation; 3. find a representation in more than one dimension; summer 1986 6 iassist quarterly 4. find a representation for a maximal subset of the stimuli; 5. find a representation for a maximal subset of the subjects. the first strategy is to dichotomize full or partial rank orders of stimuli into the k most prefened and the n-k least preferred stimuli. the unfolding analysis of such data, parallelogram analysis, can be defended with the argument that the stimuli a subject prefers most will be the most salient ones for him, and a subject will therefore be able to single them out more rehably than the remaining ones. moreover, although the imfolding model assumes that successively chosen stimuli are in a sense substitutes for the subjects' most preferred stiumlus to a deaeasing degree, graudally, in the course of giving a full rank order of preference, a subject may begin to use other criteria. coombs (1964) talked about the 'portfouo model* in this respect, and tversky (1972, 1979) suggested an 'elimination by aspects' model, in which different criteria for preference are hierarchically ordered. if we are interested in finding the dominant criterion that is used first by all respondents, then we should restrict ourselves to analyzing only the first few most preferred stimuli, lest we run the risk of introducing idiosynaatic noise. two more practical advantages of this strategy can be mentioned. first, if applying an unfolding model in which the distinction between the k most preferred and n-k least prefered stimuli does not lead to a good-fitting representation, it is no use trying more sophisticated models that require the full rank order, or that may even require metric preference information. second, the unfolding of dichotomous data implies that essentially all types of data can be used in a preference analysis, as long as the most preferred responses can be distinguished from the others. the second strategy is to relax the criterion of perfect representation to allow stochastic representatioa i regard it as obvious that preference judgments reflect so many idiosyncratic influences, that we should be happy to fmd that a rather heterogeneous group of subjects agrees on at least a dominant criterion. stochastic models have been proposed before (sixu, 1973; zinnes and griggs, 1974; bechtel, 1976; jansen, 1981). i regard tiiese proposals as inferior to the model 1 propose for at least two reasons. firstiy, many of the probabilistic unfolding models assume that the order of stimuh along the j-scale is already known, and only parameter estimation of subjects and stimiili on the basis of the known order is needed. in many cases such an approach is begging the question, as often the order of the stimuli is not known in advance. secondly, other stochastic unfolding models require that for each subject, we need the probability of his preferring one stimulus to another. in many practical applications this information is impossible to obtain: it is expensive and time consuming enough to ask respondents one single time to compare all pairs of stimuli with respect to preference. a third strategy to analyze data that are not unfoldable in one dimension is to try to represent them in more than one dimension. it is possible that subjects did not use a single criterion in making their preference judgments, they may instead have used two or three criteria simultaneously. multidimensional models have been proposed by bennett and hays (1960), roskam (1968), schonemann (1970), carroll (1972), young (1972). gold (1973), kruskal et al (1973), and reiser (1981), among others. they are appealing, because the use of more than one dimension implies the possibility of using a number of additional models that differ in the way in which the various dimensions are combined: the vector model, the weighted distance model, or the compensatory distance model, to mention only a few. summer 1986 iassist quarterly 7 there are at least four possible problems with the multidimensional unfolding model. first, in applying a nonmetric multidimensional unfolding model, we may fmd an almost degenerate solution, in which most subjects are close together in the centroid of the space, and most stimuli lie in a circle around il secondly, also with respect to nonmetric multidimensional unfolding, there is a fundamental difference between the nonmetric analysis of similarities data and the nonmetric analysis of preference data, even though both models are based on the same principle. in the multidimensional analysis of siniilaiities, the isotonic region in which a stimulus falls becomes so small that for a sufficient number of stimuh each stimulus can only be represented by a point in the space, rather than by a region. but in multidimensional unfolding, the representation of some respondents in the form of such isotonic regions is different; some isotonic regions do not shrink to points, but remain open. such respondents cannot be uniquely represented by one point in the space. thirdly, multidimensional unfolding assumes that all dimensions are used simultaneously, rather than in a hierarchical order. this is an empirical question, rather than an untestable assumptioa fourth, the assumption that all dimensions are appropriate for all stimuli is equally an empirical question, rather than an untestable assimiption. we are told that reality is not unidiraensional. indeed, a chair has a colour, a weight, and a nimiber of sizes. a person has an age, a sex, and a preference for certain drinks. and a pohtical party may be large, religious and right wing. still, we never analyze reaht>'. we analyse aspects of reality! we do not compare chairs, subjects or political parties, but sizes of chairs, ages of subjects and ideological positions of political parties. that objects or subjects have more aspects than the ones in which we are interested, does not at all imply that our analyses need to be multidimensional. they may be, but that is an empirical question, and not an untestable assimiption from the outset i do not fundamentally object to a multidimensional representation of the preferences of a group of subjects. there may be instances in which this is indeed the best model. but the utility of different models will have to be shown in their practical applicability. with respect to the last two strategies for salvaging the imfolding model, selecting a maximal subset of stimuh and selecting a maximal subset of subjects, it is estabhshed practice in multidimensional unfolding analysis to assign stress values to subjects. this implies that any difficulties in fmding a representation can be explained by pointing at suljjects who used different criteria, or who perhaps even behaved completely at random. a possible procedure, given this assumption, is to delete respondents whose stress values are too high. however, it may be the case that large stress values occur because one or more stimuli caimot be represented since they do not belong in the same imiverse of content as the other stimuli. subjects are allowed to differ in their evaluation of the stimuh, but for unfolding to be applicable, they must agree on the cognitive aspects of the stimuh; whether gentlemen prefer blondes or brunettes is a different matter from estabushing whether marilyn is blonde or brunette. if there is no agreement among the subjects on the characteristics of a stimulus, differences in preference will be difficult to represent often, subjects are selected as representatives of a larger population. deleting subjects lowers the possibility of generalizing from a sample to a population. stimuli, on the other hand, are often not so much a random sample of a population of stimuli, but are more often intended to serve as the best and most prototypical indicators of a latent trait; we are often not so much interested in the actual stimuli, but rather in their implications for measuring subjects along this latent trait this summer 1986 iasast quarterly means that we generally can delete stimuli with less harm that when we delete subjects. the discussion of these strategies is intended to justify the strategy adopted in the technique to be described below, of finding a stochastic representation of a maximal subset of stimuli and all subjects in one dimension, using the first few preferences of each subject unfolding dichotomous data: the concept of 'error' we generally do not know in advance which stimuli can be represented in an imfolding scale, nor in which order they can be represented. the approach used here is a form of hierarchical cluster analysis, in which first the best, smallest unfolding scale is found, and then is extended by more stimuli, as long as they continue to satisfy the criteria of an unfolding scale. the smallest unfolding scale consists of three stimuli, since it takes at least three stimuli to falsify the unfolding model. if stimuli a, b, and c form a perfect unfolding scale in this order, then subjects who prefer a and c but not b, do not exist for the unfolding scale abc the response patterns in which a and c are prefened but b is not, is defined as the 'error pattern' of that triple of stimuli. but since we do not know in advance in what order the stimuli form an unfolding scale, we must take into account the three permutations in which each of the three stimuli is the middle one: bag, abc, and acb. if a subject prefers a and b, but not c, for example, he makes an enor according to the unfolding scale acb. for each triple of stimuli, given a dichotomous response to each stimulus, eight response patterns are possible: 111, 110, 101, oil, 100, 010, 001, and 000. if these stimuli form pan of an unfolding scale, then one of these eight patterns cannot occur: the pattern '101' (see table l)^ this pattern is called the 'error response'. for each triple of stimuli, in each of its three possible permutations, the frequency of occurrence of the error pattern can be counted. counting frequencies of enor response patterns can be extended to larger response patterns, in which each subject evaluates more stimuli. table 2 gives five response patterns in which two or three stimuli are preferred from a set of four. in the first two response patterns only one triple is in error. in the last three response patterns two triples are in error. the amount of error in a response pattern is defined as the number of triples in that response pattern that are in error. the last three response patterns therefore contain twice as much error as the first two. in the second example, four subjects prefer six out of seven stimuli. it makes an enormous difference to the amount of error in their response patterns whether the stimulus not preferred is d, c, b, or a. in the case of d, the amount of error is maximal, whereas in the case of a there are no errors at all. stochastic unfolding the stochastic aspect of the unfolding strategy proposed here lies in comparing the amount of error observed with the amount of error expected under statistical independence. in the deterministic unfolding model, the k stimuli that are preferred by a subject are found within the symmetric closed interval around the subject's ideal point the probability of preferring a set of stimuli (e.g., two, three, or more) will be '1' if all stimuli fall within the subject's preference ' editor's note: tables are gathered together at end of article summer 1986 iassist quarterly 9 interval, and '0' if at least one stimulus falls outside this interval. the null model differs from the deterministic model in two ways. first, local independence is assumed among preference responses for different stimuli. this means that for each subject the probability of a preference response pattern to a set of stimuli is the product of the positive (preferential) response to each of the stimuli. second, the null model assumes that there are no individual differences in the probabilities of giving a positive preference response to the stimuh. the expected frequency with which a set of stimuli is preferred will therefore be the product of the relative frequencies with which each stimulus is prefened times the number of cases, if subjects are free to selea as many 'most prefened' stimuli as they wish: exp.freq(ijk,101) = p(i).(l po)).p(k).n where p{i) is the relative frequency with which stimulus i is preferred and n is the number of cases. the expected number of errors under the null model for 'pick k/n' data is first explained for 'pick 3/n' data. it consists of two steps: 1. determine the expected frequency of the '111' response pattern by applying the n-way simple quasi-independence model (e.g., bishop et al, 1975); 2. from the '111' responses to each triple, other response patterns like 110, 101, or oil cjm be deduced. in a data matrix, in which each of the n subjects picks exactly 3 of n stimuli as most prefened, the relative frequency p(i) with which each stimulus is picked can be found. in the null model, these p(i)'s are derived from the addition of the expected frequency of triples (ijjc) for all combinations of j and k with a fixed i. this expected frequency of triples ijjc, a(ijk), is the product of the item parameters f(i), f(j), f(k) times a general scaling factor f. without interaction effects: a(ijk) = f fi;i).flj).f(k). the values of f, and each f(i) are found iteratively. (see table 3) the details of this procedure are given in van schuur (1984). once the expected frequency of the '111' pattern of all triples is known, the expeaed frequency of the other response patterns can be foimd, given that each subject picked exactly 3 stimuli as most prefened. for example: consider the situation in which there are five stimuli. a, b, c, d. and e, and each subject chooses three stimuh as most preferred. for the unfolding scale abc the error response pattern is the pattern 101, in which stimuh a and c are picked, but stimulus b is not if b was not one of the subject's choices, then d or e must have been. we can therefore calculate the expected frequency across all respondents of the response pattern 101 for the triple abc by summing the expected 'hi' responses of the triples acd and ace in general: exp.freq.(ijk,101) = ff[i).n:k).i fi[s) this procedure can easily be generalized to the 'pick k/n' case, where k = 2, or where k > 3. first, the expeaed frequency of each k-tuple, ranging between 1 and (") is found. second, the expected frequency ot the enor response pattern of an unfolding scale of three stimuli is foimd by calculating: exp.freq.(ijk,101) = ff^i).fik).q where q is the sum over all (?_-) k-2 tuples of the product of their f(s)'s, where s is not equal to i, j, or k. once we know the frequency of the error response observed, obs.freq.(ijk,101), as well as the frequency expected imder the null model. summer 1986 10 iassist quarterly exp.freq.(ijk,101), for each triple of stimuli in each of its three essentially different permutations, we can compare the two using a scalability coefticient analogous to loevinger's h (loevinger. 1948; mokken. 1971): h... = 1 obs.freq.{ijk.l01) exp.freq.(ijk,101) for each triple of stimuli (ij, and k), three coefficients of scalability can be found: h(ijk), h(ikj). and h(jik). perfect scalabihty is defined as h = 1. this means that no error is observed. when h = the amount of error observed is equal to the amount of error expeaed under statistical independence. the scalability of an unfolding scale of more than three stimuli can also be evaluated. in this case we can simply calculate the sum of the error responses to all relevant triples of the scale, for both the observed and expected enor frequency, and then compare them, using the coefficient of scalability h: 3 j3 obs. freq. (i jk , 101 ) h = 1 _„iii>l=li 3 j; e;;p. freq. (i jk , 101 ) < v jk= 1 > mudfold: multiple unidimensional unfolding, the search procedure after having obtained all relevant information about each triple of stimuli in each of its three different permutations (e.g., obs.freq.(ijk,101), exp.freq.(ijk,101), and h(ijk), we can begin to construct an unfolding scale. this is a two-step procedure. first, the best elementary scale is found, and second, new stimuli are added, one by one. to the existing scale. the best triple of stimuu that conforms to the following criteria is the best elementary scale: 1. its scalability value should be positive in only one of its three permutations, and negative in the other two. this guarantees that the best triple has a imique order of representation; 2. its scalability value must be higher than some user specified lower boundary. this guarantees that if the scalability value is positive, it can be given a substantively relevant interpretation. 3. the absolute frequencies of the perfect patterns with at least two of the three stimuli (i.e.. ill, 110, and oil) is highest among all triples fulfilling the first two criteria. this guarantees the representativeness of the largest group of respondents. the scalability of single stimuli in the scale can equally be evaluated, by adding up the frequencies of the enor patterns observed and expected, respectively, in only those triples that contain the stimulus under consideration, and then comparing these frequencies using the scalability coefficient for each stimulus separately. once the best elementary scale is found, each of the remaining n-3 stimuli is investigated to determine whether or not it might make the best fourth stimulus. the fourth stimulus (e.g., d) may be added to the three stimuli of the best triple (e.g., abc) in any one of four places: dabc, adbc. abdc, or abcd, denoted as place 1 through place 4, respectively. the best fourth or, more generally. summer 1986 iassist quarterly 11 the criteria: p+l-st _ stimulus must fulfill the follomdng all new (p) triples, including the p+l-st stimulus and two stimuli from the existing p-stimulus scale, must have a positive h(ijk)-value. this guarantees that all stimuli are homogeneous with respect to the latent dimensioa the p+l-st stimulus should be uniquely representable, in only one of the p possible places in the p-stimulus scale. this guarantees the later usefulness and interpretability of the order of the stimuli in the scale. the h(i)-value of the new stimulus, as well as the h-value of the scale as a whole, must be higher than some user-specified lower boimdary (see second criterion for the best elementary scale). if more than one stimulus conforms to the criteria mentioned above, that stimulus will be selected which leads to the highest overall scalability value for the scale as a whole. the dominance and adjacency matrices: visual inspection of model conformity once a maximal subset of unfoldable stimuli is found, a final visual check of model conformity can be performed by inspecting the dominance and adjacency matrices. the dominance matrix is a square, asymmetric matrix which contains in its cells (ij) the proportion of respondents who preferred stimulus i but not stimijus j. if the stimuh are in their order along the j-scale, then for each stimulus i the proportions p(ij) should decrease from the first column toward the diagonal and increase from the diagonal to the last column. the adjacency matrix is a lower triangle that contains in its ceus (ij) the proportion of respondents who preferred both i and j. if the stimuli are in their order along the j-scale, then for each stimulus i the proportions p(ij) should inaease from the first column to the diagonal and deaease from the diagonal to the last row. this pattern is called a 'simplex pattern*. stimuli that disturb these expected characteristic monotonidty patterns should be considered for deletion from the scale.(see table 4). this procedure, of extending a scale with additional stimuli, can continue as long as the criteria mentioned above are met if, however, no stimulus conforms to these criteria, the p-stimulus scale is a maximal subset of unfoldable stimuli. a new procedure then starts which begins by selecting the best triple among the remaining n-p stimuli. this procedure, in which, for a given pool of stimuli, more than one maxima] subset of unidimensionally unfoldable stimuli can be found, is called 'multiple scaling'. scale values once an unfolding scale of a maximal subset of stimuli has been found, scale values for stimuli and subjects must be foimd. the scale value of a stimulus is defined as its rank number in the unfolding scale. the scale value of a subject is defined as the mean of the scale values of the stimuli that the subject chose as most preferred. subjects who did not pick any stimulus from the scale cannot be given a scale value, and must be treated as missing data. an example of the assignment of scale values is shown in table 5. summer 1986 12 iassist quarterly respondents may have different response patterns, but be assigned the same scale value. this can be seen by comparing subjects 1, 2. and 3. subjects 4 and 5 show that a scale value for a subject does not need to be an integer value. respondent 6 shows that a scale value is assigned to a subject regardless of the amount of error in his response pattern, which in his case is maximal. subject 7 does not pick any of the 7 stimuli and therefore cannot be represented on this scale. an example: pick the 2 most sympathetic of 6 european party groups as part of the middle level elite project (e.g.. van schuur, 1984), sympathy scores for six european party groups in the european parliament of 1979 were elicited from party activists from 50 political parties in the european community. the responses of 1786 subjects about their two most sympathetic party groups were analyzed. the six party groups are, with the letter by which they will be denoted, and with the frequency with which they were mentioned as sympathetic in brackets: a: communists (359); b: social demoaats (747); c: european democrats for progress (366); d: european liberals and demoaats (662); e: christian democrats (792); and f: conservatives (646). the frequency with which each pair of parties was mentioned as most sympathetic is: ab(341) ac(9) ad(3) ae(2) af(4) bc(106) bd(202) be(86) bf(12) cd(124) ce(50) cf(77) de(217) df(116) ef(437). on the basis of this information, a labeled matrix can be constructed that contains, for each triple of stimuli in each of its three essentially different permutations, the values obs.freq.(ijk,101), or e(o), exp.freq.(ijk,101), or e(e), and h(ijk). this information is given in table 6. table 6 provides all the necessary information for constructing an unfolding scale. first, the best elementary imfolding scale is found among those triples that have a positive scalabihty value in only one of its three permutations. this leaves the ordered triples abc, abd, bcf, bde, cde, dcf, cfe, and def. triple abd is the best triple, since the sums of the pairs (a3) and (bj)) is highest. the h-value of triple abd is 0.96, which is well above the recommended default user specification of 0.30. on the basis of scale abd, stimulus c cannot be represented in this scale in any position, since the triple b,cj) has negative h-values in all three permutations. stimulus e is uniquely representable in place 4, forming scale abde, whereas stimulus f is representable in either place 1 (scale fabd) or place 4 (scale abdf). stimulus e is selected because it is the only one uniquely representable. the four-stimulus scale is abde, its h-value is 1 93/485 = 0.81. which is acceptably high. for the best fifth stimulus, we need only consider stimulus f. this is now only representable in place 5, which gives the final scale abdef. its h-value is 1 245/1185 = 0.79. in the process of scale construction, the h-values of individual stimuli are also calculated. for the triple abd these values are the same: h(a) = h(b) = h(d) = h(abd) = 0.96. for the fourand five-stimulus scales these values must be computed separately. the resulting h-values for the final scale are shown in table 7, along with the dominance matrix and the adjacency matrix for the stimuli in the order of the final scale. neither matrix shows any violation of the expeaed characteristic monotonicity pattern. five of the six european party groups can be included in an unfolding scale based on party summer 1986 iassist quarterly 13 activists' sympathy scores for these party groups. the scale can be interpreted as a left-right dimension, with the communists represented in the left-most place and the conservatives in the right-most place. to corroborate this interpretation, i have correlated subjects' scale scores for this imfolding scale with their scores on a left-right self-placement scale. this correlation was 0.66. the european demoaats for progress (edp, stimulus c) was not incorporated in the scale. this party group consists of the french gaullists (rpr), the largest irish party fiaima fail (ff). and the danish progress party (frp). this party group is not represented in many ec countries, so it is probably less well know than other party groups, and did not, therefore, receive high sympathy scores from respondents who might have been expected to be sympathetic, based on their positions on the scale. concluding remarks the procedure described above for the analysis of 'pick k/n' data can be extended to apply to 'rank k/n' data. such procedures have been independenuy proposed by davison (1978) and by van schuur and molenaar (1982). using partial rank order information might provide more precise measurements for both the stimuli and the subjects. however, since in this procedure all six permutations of a triple of stimuli have their own observed and expected error patterns, the accuracy of estimation with the same data set decreases sixfold. as table 7 already shows, the h(ijk)-value of some triples is based on a comparison of rather small numbers, and such comparison will therefore be even more difficult in the 'rank'-case. moreover, for small k the increase in measurement precision is minimal, and 1 have already expressed some doubts about the rehability of the k-th preference judgment, when k gets large. a computer program (mudfold) has been devised to perform a multiple unidimensional unfolding aiialysis on complete or partial rank order data, 'pick k/n' or 'pick any/n' data, or on the usual attitudinal data, such as likert items or thermometer scores. the program is interactive, self-explanatory, and very user-friendly. the user may define a startset rather than use the best elementary scale to find a larger unfolding scale, or test the unfoldability of a given set of stimuli in a given order. in either case, if a triple of stimuli in the user defined order has a negative h(ijk)-value, this triple will be flagged, along witii its e(o>-, e(e)-, and h(ijk)-values. the output not only consists of the hand h(i)-values of the final scale, but also gives an overview of which stimuh at which places were candidates for selection at what step of enlargement, and the hand h(i>-values of the stimuh in the scale at which step of enlargement moreover, the output contains a variety of additional information which may help the researcher either find a better scale, or explain why certain stimuh did not fit in the unfolding scale. the computer program is available from the university of groningen. the development of the unfolding model presented above, together with more than twenty applications, is described in more detail in my dissertation (van schuur. 1984).q references: bechtel, g.g. (1976). multidimensional preference scaling . the hague: moutoa bennett, j.f. and w.l hays (1960). multidimensional unfolding: determining the dimensionality of ranked preference summer 1986 14 iassist quarterly data. psvchometrika. 25. 27-43. bishop, y.m.m.. s.e fienberg. and p.w. holland (1975). discrete multivariate analysis: theory and practice . cambridge, mass.: mit press. carroll, j.d. (1972). individual differences and multidimensional scaling. in: r.n. shepard et al (eds.). multidimensional scaling. vol. i: theory. new york: seminar press. coombs. c.h. (1950). psychological scaling without a unit of measurement psychological review. 57, 148-158. coombs, c.h. (1964). a theory of data. new york: wuey. davison, m.l (1979). testing a unidimensional, qualitative imfolding model for attitudinal or developmental data. psvchometrika. 44, 179-194. gold. em. (1973). metric unfolding: data requirements for unique solutions and clarifications of schbnemann's algorithm. psvchometrika. 38, 555-569. heiser, w.j. (1981) unfolding analysis of proximity data. university of leyden: impublished dissertation. jansen, p.g.w. (1983). rasch analysis of attitudinal data . catholic university of nijmegen rijks psychologische ekenst, den haag: unpublished dissertation. kruskal, j.b., f.w. young, and j.b. seery (1973). how to use kyst. a very flexible program to do multidimensional scaling and unfolding. murray hill: bell labs., mimeo. meerling (1981). methoden en technieken van psvchologisch onderzoek. deel 2: data-analyse en psychometrie . meppel: boom. mokken, r.j. (1971). a theory and procedure of scale analysis, with application in political research . the hague: mouton. roskam, ee. (1968). metric analysis of ordinal data. voorschoten: vam. schonemann, p.h. (1970). on metric multidimensional unfolding. psvchometrika. 35, 349-366. sixtl. f. (1973). probabiustic unfolding. psvchometrika . 38. 235-248. tversky, a. (1972). elimination by aspects: a theory of choice. psychological review. 79. 281-299. van schuur. w.h. and i.w. molenaar (1982). mudfold. multiple stochastic unidimensional unfolding. in: h. caussinus. p. ettinger. and r. thomassone (eds.) compstat 1982. part i: proceedings in computational statistics . vienna: physica-verlag. pp. 419-424. van schuur. h. (1984). structure in poutical beliefs, a new model for stochastic unfolding with application to european party activists. amsterdam: ct press. young. f.w. (1972). a model for polynomial conjoint analysis algorithms. in: r.n. shepard et al (eds.). multidimensional scaling: theory and applications in the behavioral sciences, vol. i: theory . new york: seminar press. zinnes. j.l and r.a. griggs (1974). probabihstic multidimensional unfolding analysis. psyrhnmemka 39, 327-350. summer 1986 iassist quarterly 15 table 1 table 1: parallelogram analysis of perfect 'pick 3/11' data 1 : subject prefers stimulus o: subject does not prefer stimulus subjects : stimuli : 123456789 -i i i 1abcdefghijk subject nr. response pattern: 1 11100000000 2 01110000000 3 111 4 111 5 ' 111 6 111 7 111 8 00000001110 9 00000000111 table 2 table 2: two examples of response patterns that contain error example 1: example 2: a b c d error in triples a b c d e f g error in triples 1110 111 ade adf adc bde bdf bdg cde cdf cd 110 1111 acd ace acf acg bcd bce bcf bcg 10 11111 abc abd abe abf abg 111111 none 10 10 abc 10 1 bcd 10 1 abd acd 110 1 acd bcd 10 11 abc abd summer 1986 16 — iassist quarterly table 3 table 3: observed data matrix and matrix with expected frequenc observed data matrix: stimuli: sum subjects: a b e ...i j k ...n 1 1110 3 2 110 1 3 n 10 10 1 p(i\) p(b) p(£:) p(d) p(i) p(j) p{k) p(n) matrix with expected frequencies : stimuli: triples: a b c d ...i j k ...n (^«^) ^abc ^abg ^abc ° a^^ (^^^ ^abd ^abd ° ^abd a^^ (ijk) a. j^ a..^ a..^ (3) p(a) p(b)-'ptc) p(d) p(i) p(j) p(k) p(n) ijk : expected frequency of triple (ijk) = f .f (i) .f ( j) . f (k) (i.e., no interaction) "ijk the values for f and f(i) are found iteratively summer 1986 iassist quarterly 17 table 4 table 4: dominance and adjacency matrix for a perfect 4-stimulus unfolding scale data matrix a b c d frequency 1 p 1 q 1 r 1 s 1 1 t 1 1 u 1 1 v 1 1 1 w 1 1 1 x dominance matrix: a b c d a p p+t p+t+w b q+u+x q+t q+t+u+w c r+u+v+x r+v r+u+w d s+v+x s+v s adjacency matrix a b c d a b t+w c w u+w+x d x v+x table 5 table 5: assignment of scale values to stimuli and subjects stimuli rank number subject nr . 1 2 3 4 5 6 7 a b c d e f g 12 3 4 5 6 7 1 1 10 10 10 10 10 1 10 10 1 1 1110 111 scale value of subject 3 3 3 2.5 5.67 4 (missing datum) summer 1986 18 — iassist quarterly table 6 table 6: labeled h-matrix for 'pick 2/6' european party groups, scale jik scale i]k scale1 ik3 e(o) e(e) h(ijk) e(o) e(e) h(ijk) e(o) e(e) h(ijk) abc 106 88 -0.21 9 36 0.75 341 86 -2.97 abd 262 177 -0.14 3 73 0.96 341 86 -2.97 abe 86 226 0.62 2 93 0.98 341 86 -2.97 abf 12 171 0.93 4 71 0.94 341 86 -2.97 acd 124 75 -0.66 3 73 0.96 9 36 0.75 ace 50 95 0.4 8 2 93 0.96 9 36 0.75 acf 77 72 -0.07 4 71 0.94 9 36 0.75 ade 217 192 -0.13 2 93 0.98 3 73 0.96 adf 116 146 0.20 4 71 0.94 3 73 0.96 aef 437 186 -1.35 4 71 0.94 2 93 0.98 bcd 124 75 -0.66 202 177 -0.14 106 88 -0.21 bce 50 95 0.48 86 226 0.62 106 88 -0.21 bcf 77 72 -0.07 12 171 0.93 106 88 -0.21 bde 217 192 -0.13 86 226 0.62 202 177 -0.14 bdf 116 146 0.20 12 171 0.93 202 177 -0.14 bef 437 186 -1 .35 12 171 0.93 86 226 0.62 cde 217 192 -0.13 50 95 0.48 124 75 -0.66 cdf 116 146 0.20 77 72 -0.07 124 75 -0.66 cef 437 186 -1.35 77 72 -0.07 50 95 0.48 def 437 186 -1.35 116 146 0.20 217 192 -0.13 table 7 table 7: final unfolding scale for 'pick 2/6' european party groups a communists b social democrats d european liberals and democrats e christian democrats f conservatives dominance matrix adjacency matrix p(i) h (i) 0.20 0..96 0.42 0,.85 0.37 0,.71 0.44 0..72 0.36 0,.79 n=1786 h = .79 a b d e f a 1 19 19 19 b 17 25 31 35 d 30 19 18 24 e 41 37 29 17 f 32 31 25 7 17 11 5 12 1 6 summer 1986 vol21.1 21spring/summer 1994 changing technology has created many challenges for today’s data suppliers and this paper will begin with a brief introduction regarding recent technological changes. it will then go on to look at the role of the ‘user services’ section in the uk data archive and how this has now been split to create a new user support role before finally raising some questions as to the direction user services will be taking in the future. introduction we are all aware that technology has advanced rapidly over the last few years with a huge growth in dispersed computing systems, such as personal computers, and the use of the associated networking facilities especially the world wide web. along with this has come a much wider range of software for use by pcs and changes in media for data delivery with cd-roms in particular becoming popular. the following graph shows the media used for the delivery of data from the uk data archive between 1993 and 1997 and how it has changed over this period of time. the uk data archive has seen the number of orders supplied increase significantly over the years. this has brought with it an increase in the number of enquiries dealt with by the user services section. there has also been a disproportionately higher increase in queries from less experienced users, who require much more help and guidance than other types of users. in order to provide an efficient service to users within the archive’s constrained resources it was decided to divide user services into ‘pre’ and ‘post’ data delivery. thus user services staff now specialise in assisting users before they order data and offering support to users who have queries after their data and documentation have been delivered. pre-order queries in the uk data archive two user services staff focus upon enquiries before any data or documentation have been delivered. these tend to fall into the following categories of ‘pre-order’ enquiries: n general information n ordering data n the forms required and how to complete them n the datasets held by the archive n the costs which may be involved n the formats available n the media available these sort of enquiries have been received by the archive regardless of the technological changes that have taken place, but the detailed nature of these enquiries has changed as technology has changed. although there has always been a choice of media on which to receive data this choice has now expanded and users need advice on which would be most suitable for the dataset they are thinking of ordering. they may also ask about the formats in which a dataset can be delivered. can it be converted to a format the user is familiar supporting data users in a world of changing technology by margaret ward* media used for data delivery 1993 to 0 5 10 15 20 25 30 35 40 45 50 cd-rom diskette digital tape mag. tape ftp media % 1993/94 1994/95 1995/96 1996/97 22 iassist quarterly with? will it be easily read by their favourite spreadsheet package? these types of questions mean that user services staff need to be knowledgeable about the various media that are available and the limitations of each. they also need to know more about the various formats appropriate for different data and whether the dataset can be converted to this format for the user. this means that it is important for user services staff to keep up to date with changes in technology and that they should have an understanding of the various formats and media which are currently available. post-order’ queries in the uk data archive. ‘post-order’ queries tend to be more broader than other types of queries and often need some investigation to resolve. therefore in order to provide a better ‘post-order’ service to the users of the data archive it was decided that one person should have responsibility for handling this type of query and take on a more supportive role. one of the advantages to users of this change is that they would have one key person to contact. this person has responsibility for ensuring that all queries received are dealt with directly or by redirecting them to the most appropriate person within the archive. if a query can not be answered by archive staff, and the data depositor needs to be contacted, then the post-order support person is also responsible for contacting the relevant depositor and also for keeping track of the progress of the query with them. the user is kept informed at all times of the progress of their particular query. examples of ‘post-order’ queries? n “i’m having problems importing some export files from the cd you recently sent me. what should i do?” n “i seem to have more categories for one of the variables in my file than there are labels. could you tell me what this one means?” n “could you give me some more information about this variable, i’m not sure exactly what is included in it?” the queries database to enable queries to be tracked through the archive and in order to ensure that we improve our service in future by learning from the queries we receive, it was decided that all of them should be logged into a database. as all the queries are channelled through one person, this person has responsibility for ensuring that all the necessary details are entered before being assigned a query number. this person is also responsible for initially examining every query and dealing with it when possible or deciding who is the most appropriate person to pass it to. the database is accessible by all members of staff. this is to provide everyone with the ability to check the current status of any query and to enable them to add any relevant information they may have as to the current status of a particular query. this is particularly important if a query has been re-assigned as any relevant information must be added to the query so that anyone can find out the current status of it. it was decided that a microsoft ‘access’ database would be used to record the information for each query. this was decided in part because other databases within the archive were also to be written in access, including the new order tracking system, and it would therefore be easier to link in any common information such as names and addresses. but what information should be recorded? discussions took place with different members of archive staff as to the use which would be made of the information recorded in order to determine the choice of fields. the information recorded for each query is as follows: n date query is logged onto the database n name of person logging the query n name and address of person reporting problem n priority of query (assigned by the archive on a high, medium or low basis) n study number n order number n brief details of the problem n name of person to whom the query has been assigned as more information is gathered regarding each query further comment fields can be added, which are also dated, thus allowing the progress of a query to be seen at any time. when a query has been resolved details of the resolution are entered. a brief note of any action that has been taken is also recorded, examples of which are: n referred to depositor n advice given n re-order entered n data re-acquired a field containing a category relating to the type of query will be added shortly. this was not included when the database was designed as we wished first to monitor the types of queries we received. we have used this information in order to decide on a standard list of query types. a standard list will be used for analysis purposes in order to identify the types of problems we receive. the queries database has been operating for the past seven months and the following graph shows the number of queries logged per month. 23winter 1998 advantages of, and information provided by, the queries database logging ‘post-order’ queries enables problems to be tracked, provides the archive with information which can be used to enhance and improve the support given to its users, and also results in statistics on the performance of the archive as detailed below: n no query is lost. as all queries are logged in one database none can be lost within the archive as could happen if the query details are not recorded and passed on orally, or emailed, from person to person. the information is also being maintained in a database accessible to all staff. n depending on the resolution to a query additional information may be needed to enhance the documentation for future users of that particular dataset. an example of this would be where a user has identified a variable with no variable or value labels. once information has been received from the depositor this is added to the documentation supplied with that particular dataset so that other users do not experience the same problem. n the database also provides information on frequently encountered problems (feps) as opposed to frequently asked questions (faqs)! by logging all queries it is possible to identify the types of problem users are having and enables us to make appropriate improvements to the archive’s services in order to reduce future problems. for example if a number of users are having difficulties reading files from cd-roms then perhaps there is a problem with the way they are being written, or the way a users’ hardware or software handles cds. problems with extracting compressed files from floppy disks may be due to users not understanding the instructions they have been given, or possibly to inadequate instructions provided by the archive. these types of problems are investigated and action taken if it is thought necessary. the database allows a range of problems to be identified, monitored and resolved. n particular datasets which result in a number of queries can also be identified. these could be examined further to see if there is anything ‘peculiar’ to these datasets. perhaps they are only provided in a certain format and this is proving to be problematic for users. alternatively these could be known to be difficult to use so perhaps additional documentation is needed to make them more ‘user-friendly’. n queries which have had to be referred to depositors can also be identified. this allows the archive to identify which datasets have needed extra information from the depositors before they can be used more easily by the archive’s users. hopefully this can assist us to improve the acquisition process for data in the future. n the archive can also identify which of its users repeatedly report queries! perhaps some users need more support than we can reasonably provide and we might involve local organisational representatives to assist them. using the information provided by the database as mentioned above, a vast amount of information can be gleaned from the queries database. but how can the archive use the information obtained effectively? creating additional notes to be added to existing documentation has already been mentioned, however perhaps this information should be made more widely available. perhaps we could utilise the archive’s web pages more effectively to distribute information. these would be accessible by everyone and could be promoted as a place to look to first before contacting the archive. this may also be particularly useful for datasets used by a large number of users in several countries; the imf databank, supplied by icpsr for example. these pages would have to be organised in such a way as to enable users to find the information which was relevant to them quickly and easily, but would be effective in providing information to a large number of queries logged in database 0 5 10 15 20 25 oct-96 nov-96 dec-96 jan-97 feb-97 mar-97 apr-97 month no 24 iassist quarterly number of people. the data archive could link additional useful information to the biron (our on-line catalogue) entry for a dataset. users who are using a particular dataset could then look at the relevant part of biron and see if there is any new information relating to that study. however it is important that any information relating to a dataset is also included with the main documentation supplied with the data. users should not have to have to check different web pages for vital information necessary for analysing the data which have been supplied to them. it should also not be forgotten that some users do not have access to the world wide web, but are entitled to the same support as those that do! the future? what type of service will data suppliers be offering in the future? will all data formats and media be available for users or will availability be limited? just how much support should be given once the data has been delivered? what is a reasonable amount of time to spend on one query? should this time be limited and should a charge be made for the help given? these are questions which i think will become important in the future. should this support be monitored? it has been argued that it takes longer to log a query than to answer it! in some cases this is true but i believe the information which can be gained by monitoring queries received far outweighs the time taken to record the details. so much can be discovered about users and the problems they have which can be used to make sure that we, the data supplier, provide them with a good support service and so ensure that as a data supplier we do indeed have a future! * paper presented at the iassist/ifdo 1997 conference, may 6th may 9th odense, denmark. iassist quarterly 3 downloading problems? in discussing some of the problems which might occur in these four areas, i will use two specific downloading applications attempted to solve conference board problems. one of these involved the downloading of numeric data from our own mainframe computer, a burroughs; the other was a test downloading of mixed numeric and text data from an external on-line database. in both cases we were downloading into an ibm pc-xt. by kay worrell' director, survey research center the conference board, new york in the past two years, with the trend toward decentralization of many computer services and with the spread of microcomputers, new applications and new problems have arisen in research organizations. "downloading," or moving files from a mainframe computer to microcomputer, olters some new opportunities and some solutions to old problems. but as with other solutions, new problems are also presented. there are four areas in which problems may occur: as background, i should say something about the conference board and our research. the conference board is a not-for-profit business and economics research group, based in new york and with offices in brussels. the conference board of canada is headquartered in ottawa. we are supported primarily by subscription income from associate members and by conference fees. my department, the survey research center, processes questionnaires for researchers surveying business practices and economic trends, and assists research staft with computer systems and software. it is important to note that for most of our survey data, the "case" is the corporation or company, and the "background variables," or demographics, are company traits such as sales, assets, number of employees and type of industry. the mainframe "source" of the data the data itself — its form the communications package and modem — the vehicle(s) with which the data are to be moved the receiver — the software (or system) into which the data are to be stored. 'presented at iassist meeting 1985, amsterdam our specific applications the first and simplest application was downloading coded numeric data from a dataset we had developed in our mainframe computer. the data was set up in 80-column records to be used with spss, as is most of the data we process from our questionnaires. the purpose was to provide research staft an opportunity to work on the data interactively in a pc statistical package, and to experiment with pc spreadsheet summer 1985 4 iassist quarterly and database management packages. the latter would offer more flexibility in using nonnumeric data such as names of companies, titles of indivuduals, and responses to open-ended questions. after downloading, we would add this information to the coded numeric files in the micro. the second application was to try to download company information from an external online database as a possible source of background information for our questionnaire data. in the past this information was requested on the questionnaire, and checked upon receipt, or was added to the questionnaire after it was returned from printed sources such as the standard & poor's directory of corporations or fortune magazine's annual list of 500 largest u.s. corporations. the information was then key-entered with the rest of the questionnaire data. sometimes the information was encoded on the questionnaire label, prior to mailing, from our own mainframe computer list, where much of this information is also kept and updated regulariy. again, it would be re-keyentered into the specific dataset when the questionnaires were returned. we decided to try to download company information from the disclosure ii database, available online through dialog. we wanted to download company names, sales, number of employees, and primary sic (standard industrial classification) number, used to indicate industry group. we were then going to explore ways to link this information with that in our data files, or organize it in such a way that it could be easily referenced by clerks in an off-line mode. the information in the disclosure ii database is derived from the forms that publicly-held companies in the u.s. are required to file with the securities and exchange commission. disclosure has an exclusive contract with the sec to computerize this data, so that it is the most complete, authoritative, and up-to-date source available. i should note that this downloading was purely exploratory, and that permission would be arranged before the conference board would implement use of such an external source. thus our second application might offer solutions to several problems, by allowing us update our own mainframe list with the most current and reliable information, match this information with questionnaire data, and avoid considerable clerical work and redundant keyentry. interface considerations in downloading beginning in the mjiinframe dataset, the size of the file should be one of your first considerations. this will be influenced by the number of records as well as by the volume or size of each record. the critical question at this point is: will the file fit on a fioppy disk? format of the dataset is another important concern. is it fixed field? sdf (standard delimited format) convertible? is the data in column format? if so, is it sdf or dif convertible? is it numeric, alpha, or mixed? are there multiple "lines" or "records (cards)" per case? are there variable length fields or variable length records (a different number of fields possible on each)? what are the host system characteristics? are there line numbers? is numbering an option? how long is each line, or record? what type, of end-of-line or end-of-record character is used? can the host system transmit anything besides ascii files? what is the configuration of the communications hardward? full or half duplex, synchronous or asynchronous line; speed of communication — baud rate? what communications package will be used? what will it do for you? can you summer 1985 iassist quarterly move system files or jusl asqi files? will the receiving unit be a hard disk or diskette? how much space is available? what is the method of transfer, or "protocol?" finally, after the data is downloaded to the pc or microcomputer, there are the following questions. can you load the information directly into the package you wish to use it with? will it be desirable to load it into a word processing package? as an intermediary, for reformatting? for word processing uses? if it is to be used in a database management package or a spreadsheet package, will you wish to add additional data? to merge with other files? the downloading we used linkit to communicate between the two systems, and most often used keepit, a database management package, to receive the files. both are produced by itsoftware of princeton, nj and distributed by martin marietta data systems. the main advantage of keepit is its interfacing capabilities, allowing it to receive and reformat data and generate output directly in sdf, dif or other ascii format the steps to download are the same for all applications,and are: sign into the communication package. call up the host computer in which the "source" daiaset is stored. (set or reset communications parameters, if necessary. most often these can be saved in the communications package used.) sign into host computer and call up dataseu do not issue a "ust" or "display" command yet. you may want to check on the size of your dataset before downloading. indicate to communications package that you wish to "receive" a file, this will probably be done with a function key. (linkit leads the user to such an option with a menu.) the package will ask you to name the file. this may be done using normal naming conventions, including designating the disk drive address. be sure there is sufficient space available for the file. return to host system and issue a "list" or "display" command. the "listing" should then be "received" by the communications package. note: if the host system allows an option of listing "unnumbered," that is, without line numbers, use this option. otherwise you will want to remove the line nimibers after the file is downloaded. when the listing is completed, the downloading will be completed. most often you will see this indicated by the end of a count of characters received appearing on the screen. sign off the host system. then exit the communications package. the downloaded file will be stored on your hard disk or diskette, with the name you supplied in step 5. you may then load the file into the package of your choice. it may be called directly into most word procesing packages. if you use a database management package such as keepit or dbase iii, you must first define the file parameters, including all fields and the length and type of each. summer 1985 iassist quarterly the first problem we encountered downloading data files from our mainframe was that of the line numbers, mentioned above. this can be avoided by using an "unnumbered" option. another, related problem was some unnecessary information repeated on each "card" of a case or observation — questionnaire identification number and "card" number, this information took up the first nine columns on each "card" or line, and was easily removed using a word processing package. or, loading the whole file into a database management package, that information could be defined as "dummy" variables or fields, to be omitted later. we have tried to keep the size of data files downloaded rather small — less than three "cards" per case and only a few hundred cases. larger files are not very efficiently processed in a pc. even so, downloading can be time-consuming at normal baud rates of 300 or 1200. to speed up the communicating of larger files, our edp department provided us with a special line to transmit data to the pc at 9600 baud. this required changing the modem from hayes to "direct" and some changing of the plugs, besides changing the baud rate setting. also, as noted above, it is important to estimate the size of the file to be received and to be sure that the disk on which it will be received has sufticient space. if this is not the case, you may lose much of the data you tried to download, and waste considerable time. as our data was in 80-column records and was already in fixed-field format, we expected no problems defining the fields to the database management packages. using keepits menu, we just "read data." we discovered, however, that keepit requires each record to have 80 columns, and no less. some of our records were shorter than 80 columns, and cande, the burroughs operating system command-and-edit language did not fill those records in with anything recognizable by keepit. so our staff had to place a character in the 80th column of each record before downloading into keepit. this was done with a "replace" command in cande, but it meant we could only use 79 columns for data. we have used keepit as an intermediary, to format files for dbase, lotus 1-2-3 and wordstar's mailmerge. in dbase, after defining the file parameters, we can use "append". if the file contains addresses, we can use either keepit or dbase to aeate a mailmerge input file. we can produce a die output file with keepit, using spreadsheet intereace, and then "import" this file into lotus 1-2-3. or we can use the dbase-lotus interface for this purpose. to add data to any of thse files, we need only create the additional fields, copy existing data in, then edit the file to add the new data. or we might create parallel files with a linking id number for the new data, then merge the files. on moving data into statistical software packages, we found some common good points and bad points. in their favor is the fact that virtually all pc statistical packages seem capable of accepting ascii data files. we have experimented with spss-pc, statpac and statlt. some time can be saved if spss is to be used on the pc, as much of the labelling used for spss in the mainframe can be downloaded and adapted for use in spss-pc. of course if you are using an ibm mainframe with kermit, downloading will be even more convenienl a disadvantage is that pc statistical packages have a smaller capacity, as would be expected, than mainframe packages. thus only partial files could be used. another problem encountered was that statpac will not read multiple-line or multiple-card files. the data must be in a continuous stream ending with a carriage return after each record. a record may not exceed 255 characters. a utility file, which comes with statpac, must be used to concatenate 80-coumn records to form a statpac-readable summer 1985 iassist quarterly 7 record. trying to download data from an outside system created considerably more problems. as was noted at the beginning, this was an experimental application. again using linkit, we signed into dialog, and then into disclosure ii. we discovered there were three basic forms in which data could be presented from disclosure. the first did not include the variables we wished to see — sales, assets, number of employees and sic number. the second, the corporate resume, contained considerably more information than we wished — about two full screens of data for each company. the third contained much, much more — whole company records, including text from annual reports. we decided to try to download the corporate resimie for a small number of companies, just to see how long it would take. selecting only companies with upwards of $40 billion in net sales, we narrowed the search to 17 companies. we then downloaded these companies. even at 1200 baud, it took nearly 10 minutes to complete this download. based on the time it took, and especially considering the line charges for dialog and disclosure, we determined that this would not be an efficient updating mechanism. further inquiries revealed that dialog will, under certain circumstances, download large files for the user onto a 9-track tape, to user specifications. we also found that we could purchase the entire disclosure database on 9-track tape from disclosure, including updates. (and since this paper was presented. 1 have received promotional information on microscan, a software product from disclosure to assist the pc user in searching and downloading from that database.) other problems that became evident, but which we did not seek to solve after the downloading, were the very large size and variable length of each record. a positive factor we discovered about disclosure was that, in addition to virtually all the company data that is publicly available for corporations, each record contained the d & b identification number, as well as that used by standard & poors' and others. ticker symbols used on the stock exchanges were also included. thus data from this database could easily be merged with data from any other, where one of these common numbers were in use. on the basis of this experience, we have decided to consider purchase of disclosure on tape to update our mainframe files. we have also decided to use one or more of the common identification numbers listed in disclosure as an identifier on all futtire datasels we plan. i should add that, although i will not discuss it at length here, we have also foimd it convenient to enter data into the pc on occasion, and upload it to the mainframe. we have done this using lotus, to allow us to compute new variables during the data entry step. it should be noted that for large datasets the reaction time for lotus becomes quite slow. we then uploaded one "card" of data and merged it with nine other "cards" that had been keypunched in the traditional way. we have also entered some questionnaire data direcdy into a "screen" set up in dbase iii, and uploaded some of this information for use with spss. the problem we had to deal with in this type of transfer was transmitting a pause after each record, or "line," to allow the system time to assign a line number for each. it is expected, of course, that future developments in specific software packages will include interface enhancements. these, along with improvements in communications hardware and software, will greatly facilitate our ability to move data among packages and systems." summer j985 iassisl quarterly linkit — screen one linkit 1.2 (c) copyright 1983 vm personal computing offline (c) copyright 1983 it software your pc id is: the conference board inc, f1 = call a computer named confbd f2 = answer a call from a pc f3 = review the directory of computers f4 = set personal computer options f6 = edit a file f7 = edit a file f8 = run a program f9 = stop printing esc = exit f10 = help summer 1985 iassisl quarterly 9 linkit — screen two directory of computers name telephone number speed type notes and comments a pc compserv conf bd 83,456 dowjones source tso tymshare unattend vm 300 pc ibm pc using linkit 300 host compuserv service 1200 host the in-house burroughs sys 300 host source timesharing service direct call to a tso system tymshare or equivalent leaves linkit unattended direct call to a vm system 300 host 300 host 300 host 300 pc 300 host use pgdn and pgup to scroll the directory f1 = call name at cursor f2 = answer name at cursor f3 = add a new name in directory by copying entry at cursor esc = quit f4 = review connect options for name at cursor f10 = help summer 1985 10 — lassist quarlerly linkit — screen three linkit 1.2 your pc id is: the conference board online f1 = return to terminal screen alt fl = redial or reanswer the telephone alt f2 = hang up and return to main offline menu f3 = send files to another computer f4 = receive files to your pc f5 = set current connect options f6 = edit a file f7 = print files f8 = run a program f9 = stop printer or file transfer fio = help summer 1985 iassisl quarterly 11 keepit — screen one records: keepit main menu file: c:day data entry file definition ed enter data df define the file pd post data dc define constraints rd read data data maintenance di define indexes interfaces cd change data ci calcit interface vd view data gi graphics interface cf compute/fill data si spread sheet interface ml mail list interface reports fi forms interface st statistics interface pr print a report wo write output file tr tabulation report sh showit interface sr summary report wi move reports to wrilt quit housekeeping qf quit file do direct printer output qd quit to dos fm file management qa quit to askit mc maintain catalogues please select an option summer 1985 12 iassisl quarterly keepit — screen two fields: define the file file: c:day a create fields for a new file b insert a field c design/edit fields on screen d delete a field e display a field f scan the fields and make changes g display the data entry screen h print the field specifications i print the data entry screen j set file parameters m return to main menu please select an option: keepit — screen three fields: 1 (1 ) prompt 1 (2) name (3) page ,pageend (4) row (5) column (12) define the file specifications for field: (6) (7) (9) (10) (11) file: c:day type of field max lengtht (8) lower limit upper limit input spec format default/formula summer 1985 iassist quarterly 13 keepit — screen four records : mail list interface file: c:day a writ it/multimate b wordstaf/mailmerge c easywriter/easyfiler d peachtext e wordplus-pc f wordperfect g edis+wordix h spellbinder/eaglewriter i quote marks/comma delimited j fixed length please select an option: summer 1985 vol221 4 iassist quarterly econdata by albert bots * what is econdata? in july 1996 niwi’s steinmetz archive started the project econdata to establish a dutch data service for economic data. this service will be integrated with the current activities of the archive. econdata builds on previous feasibility studies conducted by the economic and social institute (esi) in amsterdam and the economic institute tilburg (eit). both of these studies have been funded by the netherlands organization for scientific research (nwo). for econdata the steimetz archive receives additional funding from nwo. this grant follows a reccomendation by the social science counsil (swr) of the royal netherlands academy of arts and sciences (knaw). aim econdata aims at broadening the scope of the steinmetz archive. new services will be established to support economic research, including macro-economics, business economics and economic modeling. in addition the more traditional functions of a data archive, econdata puts strong emphasis on data brokerage. the data service will act as an intermediary between suppliers of economic data and data users. this will include suppliers of international data and users of dutch data abroad. for this purpose the project plan includes the establishment of an online register of available data sets, irrespective of whether theses data sets are available from the steinmetz archive of from other sources. econdata will be evaluated in september 1998. registration econdata will establish a public register of economic data files. owners of data are invited to register data files which are suitable for secondary analysis. in addition to actual information on the data files, including information on the data owners and the conditions for use, special attention will be given to key information which will allow users to quickly and efficiently locate relevant data files. furthermore, summaries on the contents of data files will be given to provide users with quick impressions of the possible uses of the data files. the register will be offered to users in two forms: a publication (from the steinmetz archive) and an online database (accessible through the internet) which will be updated and expanded on a regular basis. mediation service the public register of economic data files will serve as a starting point for econdata’s mediation service which will facilitate negotiations between data owners and users. econdata will not only promote contact between these two groups but also represent users in negotiating data access with data owners. furthermore, through mediating access to dutch economic data files, econdata also hopes to expand its archive. in addition to mediating the acquisition of economic data files in the netherlands, econdata aims to play a role in negotiating access to international economic data files. for this, econdata-data utilizes the extensive experience of the steinmetz archive in the field of international data archive cooperation. the steinmetz archive represents the netherlands as coordinator within the inter-university consortium for political and social research (icpsr), the managing body for the world’s largest collection of scientific data files which include many economic data files. the steinmetz archive is also an active member of the council of european social science data archives (cessda) which oversees a cooperative network of european data archives on the internet. data guide the data guide is the first inventory of economic data files from a number of important archives. this inventory also includes approximately 100 relevant files which are available from the steinmetz archive. other archives covered by the data guide are those in the united states (including icpsr), australia, the united kingdom, germany, and sweden. from these archives, a selection was made of data sets which would be of interest to dutch economists. the data guide provides key information on each file as well as references to information on economic data files available on the internet. furthermore, the data guide contains information on micro economic data files of statistics netherlands which are, in part, available through the scientific statistics agency (wsa). finally, the data guide provides information on belgian and german panel data files. the data guide can be ordered from the steinmetz archive (price: 25 dutch guilders plus shipping). please post or fax orders to the address listed at the end of this paper. spring 1998 5 further possibilities econdata aims to stimulate data owners of registered data files to place their data in the care of econdata’s archive. such an arrangement offers the following advantages for data owners: secure storage of data and corresponding documentation description and documentation of data according to international standards systematic collection and documentation of publications based on secondary data analysis exposure of the research to the national and international scientific community increase of the yield of data collection through secondary analysis handling of inquiries from users and administrative management results so far up until april 1998 more than 200 economic data sets have been registered. descriptions of most of these data sets are available in the on-line database. since the beginning of the project more than 40 economic data sets have been deposited at the steinmetz archive. these data sets are available under standard conditions to users of the steinmetz archive. niwi/steinmetz archive the econdata project is carried out by niwi’s steinmetz archive. the steinmetz archive currently manages approximately 2700 data files in the field of social science research and has 30 years of experience in collecting and providing access to social science research data files. standard procedures have been developed by the steinmetz archive for the acquisition, protection of privacy, storage, provision of access and the documentation of data files. the steinmetz archive has an extensive network of contacts with data suppliers who can also supply economic data files. the archive has exchange contracts with data archives outside of the netherlands. on the web site of niwi one can find links to diverse archives for extended search and access possibilities. furthermore, search facilities are available for data file users at the archive. each quarter, the steinmetz archive issues data news, a newsletter which among other subjects, announces new additions to the archive. three databases are available free of cost, of which the most important database is star. this database contains descriptions of the studies from the main collection of the steinmetz archive. star is accessible both at the steinmetz archive and on the internet. there is also an order form for data files available at the web site. contact person since july 1996, albert bots has been appointed by the steinmetz archive as project manager for econdata. in addition to supervising the activities of econdata, albert bots is also active as a lecturer at the faculty of economics of the free university in amsterdam. albert bots can be reached at the steinmetz archive for further inquiries. address niwi/steinmetz archive po box 95110 1090 hc amsterdam tel: 020-4628625 fax: 020-6685079 e-mail: albert.bots@niwi.knaw.nl web site: http://www.niwi.knaw.nl * paper presented at the annual iassist conference, new haven, connecticut, may 19-22,1998. mailto:albert.bots@niwi.knaw.nl http://www.niwi.knaw.nl http://www.niwi.knaw.nl 6 iassist quarterly - -- -- --- ----- - --- --- ---- ----- ----- - ------------ ---- --- ---------- ---the international association for social science information service and technology (iassist) and the canadian association of public data users (capdu) announce their joint 1999 conference, "building bridges, breaking barriers: the future of data in the global network". the conference will be held may 16-21, 1999 on the university of toronto campus in toronto, ontario and will address issues of computing and information services in social science research, teaching, and data management. this is iassist's 25th annual conference, and the ninth capdu conference. http://datalib.library.ualberta.ca/iassist/ http://nexus.sscl.uwo.ca/assoc/capdu/index.html http://www.yorku.ca/org/iassist/ vol192 5summer 1995 introduction for the first time in a british census, the 1991 statistical output included samples of anonymised records (sars). known as census microdata or public use sample tapes in other countries, sars differ from traditional census output of tables of aggregated information in that abstracts of individual records are released. the released records do not conflict with the confidentiality assurances given when collecting census information since they contain neither names or addresses nor any other direct information which would lead to the identification of an individual or household. essentially three per cent of records have been released in two samples. the sars offer users the freedom to import individual-level census records into their own computing environment and the ability to produce their own tables or run analyses which are not possible using aggregated statistics. background to the release of the sars requests had been made for sars to be released from previous censuses in great britain. the principal stumbling block in the past had been an argument as to whether sars could be considered a statistical abstract for release under section 4.2 of the census act 1920 at the request and expense of user(s). furthermore, in the past, requests for sars had failed to reach a compromise between those (often geographers) wanting fine grain areal detail and those (often sociologists and demographers) wanting fine grain detail on other variables such as occupation. the 1991 census white paper (her majesty’s government 1988), however, announced: “the government intends that results from the 1991 census should wherever practicable be made available in a convenient form to meet users’ needs” legal advice having been received that sars could be deemed statistical abstracts, the white paper went on to say: “requests for abstracts in the form of samples of anonymised records for individual people and households ... would also be considered, subject to the overriding need to ensure the confidentiality of individual data”. the economic and social research council (esrc) set up a working party to negotiate with the census offices and present a formal request. their report, presented to the census offices in 1989 (subsequently published as marsh, skinner et al. 1991) concentrated on the benefits of releasing sars, the uses to which they would be put, and also an assessment of the confidentiality risks involved in releasing sars. the request was mentioned by ministers during the debate on the census order in parliament at the end of 1989. having considered the request, the registrars general for england and wales and for scotland announced in july 1990 that they had agreed in principle to the release of sars from the 1991 census. there then followed detailed work by the census offices and esrc in developing the statistical specification. an independent technical assessor, professor holt (university of southampton), was appointed to advise the registrars general on the confidentiality aspects and to write a report to ministers. following receipt of the report it was announced in march 1992 that two sars from the censuses in england and wales and in scotland would be produced and released to esrc. similar sars for northern ireland have also been made through an esrc purchase. these allow the production of harmonised sars for the whole of the united kingdom. details of the sars two sars have been extracted from the gb censuses: 1 a two per cent sample of individuals in households and communal establishments; and 2 a one per cent hierarchical sample of households and individuals in those households. samples of anonymised records from the 1991 census for great britain by angela dale1 census microdata unit, university of manchester 6 iassist quarterly the two per cent sar has finer geographical detail and the one per cent sar has finer detail on other variables, thus providing a solution to the conflict between users’ demands discussed above. the two per cent individual sar contains some 1.12 million individual records (1 in 50 sample of the whole population enumerated in the census). it was selected from the base which lists persons at their place of enumeration. details are given as to whether or not the person was a usual resident of that household, and if so (and enumerated in a household) whether they were present or absent on census night. the following other information is given for each sampled individual: details about the individual ranging from their age and sex to their employment status, occupation and social class; details about the accommodation in which the person is enumerated (such as the availability of a bath/shower and the tenure of the accommodation) or, if they were in a communal establishment, the establishment type (hotel, hospital, etc.); -information about the sex, economic position (in employment, unemployed, etc.), and social class of the individual’s family head; and -limited information about other members of the individual’s household (such as the number of persons with long-term illness and numbers of pensioners). in effect, all the census topic variables listed are on the file; the only exceptions are variables either suppressed or grouped to maintain the confidentiality of the data. in all, there are about forty pieces of information about each individual, and the size of the raw data file, before any new variables have been derived and before any data compression techniques have been applied, is around 80 megabytes. the one per cent household sar contains some 240,000 household records together with sub-records, one for each person in the selected household. information is available about the household’s accommodation together with information (similar to the two per cent sample) about each individual in the household and how they are related to the head of the household. the raw data is supplied as a hierarchical file in non-software specific character format (one line of information about housing and household, followed by one line of information about each individual in the household). the full details of the information provided in both sars are given in the codebook and glossary files produced by the census microdata unit. table 1, however, provides summaries by describing the information collected on the census form, the detail of coding of that information on the census database, and in how much detail that information is being released in the sars. the sampling procedure used census data goes through two separate coding processes. the easy to code information such as housing details, sex, date of birth, and country of birth is processed for all forms (100 per cent). the harder to code information such as occupation and industry is only processed for 10 per cent of forms. both sars were drawn from the 10 per cent sample so that they contain information from the whole of the census form. a detailed description of the sampling scheme for the sars is given in dale and marsh (1993, chapter 11). confidentiality protection in the sars the census offices in some european countries have refused to release microdata because they believe, on the basis of research such as that conducted by paass (1988) and bethlehem et al. 1990), that the risks of disclosing information about respondents’ identities are too high. much of this work is concerned with how many people have unique combinations of census characteristics which would make them open to identification. the economic and social research council working party which negotiated the release of the sars took the view that uniqueness was only one part of a four-stage process of disclosure: data in the microdata file would have to be recorded in a compatible way to that in an outside file, the individual in an outside file would have to turn up in a sar, the individual would have to have unique values of a set of key census variables and the matcher would need to be able to verify this uniqueness. rough estimates of the size of risk at each stage were made; when cumulated, the risks of disclosure appeared very low; multiplying the various probabilities together, the working party concluded that the risk of anyone in the population being identifiable from their sar record were extremely remote; their best estimate was something of the order of 1 in 4 million. (for more details of such calculations, consult marsh, skinner et al. 1991, marsh, dale and skinner (1994) and skinner, marsh et al. 1992.) the arguments put forward were important in persuading the census offices to release the sars suitably modified to protect anonymity where this was 7summer 1995 felt at risk. in this section the various disclosure protection measures taken are described. sampling as protection the low sampling fractions of the sars offer a strong source of disclosure protection for sensitive data. it not only reduces the actual risk that a particular individual can be found in the census output, but it probably has its greatest effect by reducing the chances that anyone would make the attempt at identification by this means. the two sars (a one per cent sample of households and a two per cent sample of individuals) are sufficiently small to offer a great deal of protection; the samples do not overlap so that the detailed household or occupational information available on the household file cannot be matched with the detailed geographical information available on the individual file. restricting geographical information one of the key considerations which may affect the possibility of disclosure of information about an identifiable individual or household is the geographical level to be released (i.e how much detail is given about where the person was enumerated). the full census database holds information at enumeration district level (about 200 households or 500 persons in each ed) and even at unit postcode level (about 15 households). if released, such detailed geography would obviously pose a confidentiality risk. empirical work and comparisons with sars released in other countries showed that a sensible level for release would be areas equivalent to large local authority districts for the individual (2%) sar. to be separately identifiable, the decision was taken that an area had to have a population size of at least 120,000 in the mid1989 estimates. the primary units used were local districts; only one geographical scheme was permitted, or smaller areas could be identified in the overlap, say between a local district and a health district. a population size of 120,000 is slightly higher than the lowest level of geography permitted in the us sars (100,000), but it still has the advantage of allowing all non-metropolitan counties in england and wales, most scottish regions, all london boroughs (except the city of london), and all metropolitan districts to be separately identified. smaller local authority districts (under 120,000 population) were grouped to form areas over 120,000. several rules were used to decide how districts should be amalgamated where this was necessary. first, the integrity of county/scottish region geography was always maintained, where possible. secondly, districts which achieved the minimum population threshold on their own were left intact, where possible; and smaller areas were grouped with each other. thirdly, grouping was done on the basis of contiguity. and finally, if there was a choice left once the above criteria had been met, areas were grouped on the basis of their apparent social and historical similarity. the one per cent household sar, because of its hierarchical nature (i.e. statistics about the household and all its members), is more of a disclosure risk. for this reason it was decided that, for this sar, the lowest geographical detail revealed would be the registrar general’s standard regions, plus wales and scotland. the only exception is that the south east is split into inner london, outer london, and the rest of the south east region. it should be noted that the order of records in both sars has been re-arranged before the census offices release them. this is to prevent any possible tracing of individuals or households back through a region or district. suppression of data and grouping of categories some alterations have been made to the data to reduce the number of rare and possibly unique cases. the extent to which the variables on the local base have been either suppressed entirely or modified by grouping small categories before release in sars is shown in table 1. information which is unique in itself, such as names and addresses, has been omitted altogether; (technically these variables have not been suppressed since they are never put on the computer). precise day and month of birth have been suppressed. the thresholding rule the degree of detail permitted on other variables was the subject of a thresholding rule which ensured that the expected value of any category at the lowest level of geography on any file was at least 1. the threshold, when operationalised, dictated that a category must have 25,000 cases in it in the gb file before it could be released on the individual sar, or 2,700 cases before it could be released on the household sar. with some other variables, the smaller categories have been grouped, either across the entire range of the variable or only at the extremes (a process know as “top coding”). the rule used to decide the level of detail to be released was to group information categories to a sufficient detail so that, on average, the expected sample count would be at least one for each 8 iassist quarterly category of each piece of information for the lowest geographical area permitted on each sar. some justification for restricting attention to the distribution of the univariate categories of each variable in turn was given by marsh et al (1994). they demonstrated that the risk of an individual having a unique combination of values of a set of variables could be predicted with a high degree of certainty simply from knowledge of their membership of rare categories of each variable taken singly. the precise cut-off at an expected value of 1 was set at a value sufficiently high to give reasonable protection of anonymity. the rule was applied to each census variable. expected counts were obtained by using 1981 census frequency counts (supplemented by more recent surveys, for example the labour force survey) at the national level for the whole population. to obtain expected counts, the count of 1 per category per sar area was grossed up to the national level: c = 1/x * (y/z) where c = expected count at the national level x = sampling fraction (1/50 for individual sar and 1/100 for household sar) y = national population (56 million) z = smallest geographical area population (120,000 for individual sar and 2.1 million (east anglia) for household sar thus 25,000 and 2,700 were the two thresholds used for the individual and household sars respectively. in theory, a small amount of random noise could have been added to certain variables in a manner analogous to the procedure adopted for the small area statistics. a technique similar to this has been used in the 1990 us census for example: geography has been subject to a degree of perturbation by switching a small number of similar households between nearby areas (navarro et al. 1990). however, the natural levels of noise in the data, combined with the analytical difficulties of minimising bias to both measures of location and spread by such techniques in a multipurpose file led to perturbation not being implemented in any form for the sars. grouping of variables when expected frequency counts fell below the threshold, categories were grouped. with some variables, grouping was only required at one end of the distribution: thus rooms were top-coded above 14 and the number of persons in the household was top-coded above 12. two variables were both grouped and top coded; with age, 91 and 92 were grouped, 93 and 94 were grouped and 95 and over was top-coded; with hours of work, 71-80 hours per week has been grouped and the rest top-coded above 81. when variables were not measured on a numeric scale, judgments had to be made about which categories to put together. classifications for census data are often hierarchical. for example, for the standard occupational classification there are 371 unit groups, 77 minor groups, 22 sub-major groups, and 9 major groups. in cases such as these, small categories could be amalgamated to the next level in the hierarchy. in other cases, detailed advice was sought from subject experts about how the groups should be formed. in the case of three variables in the two per cent individual sar, it was deemed necessary to further group categories, even though they contained numbers which fell above the threshold: occupation, industry, and subject of qualification. as a result of advice received from the technical assessor, occupation was reduced from the 220 categories proposed (out of a possible 371) to 73; similarly industry was cut from a possible 334 to 60 and subject of educational qualification from a possible 108 to 35. (almost full occupational detail remains on the one per cent household sar, however.) there were other factors which determined the detail to be released: categories of occupations and industries in the public eye were grouped further than mathematically necessary to guard against disclosure; for example, actors/actresses and professional sportsmen/women; large households were seen as a disclosure risk in the household sample. applying the frequency rule to size of household, a large household in the 1981 census was estimated to be one of 12 persons or more. consequently, only housing information is given for households containing 12 or more persons. no information about the individuals in the household is given. 9summer 1995 table 1 details of the information in the two samples of anonymised records from the 1991 census of great britain item household (1%) sample individual (2%) sample no. of other details no. of other details categories categories (maximum*) (maximum*) geographical area of 12 standard regions of england 278 local authority districts over renumeration (with split of south east into 120,000 population. others inner london, outer london amalgamated to form areas over and rest), wales and scotland 120,000 housing/household information accommodation type 14 (14) detached, semi-detached or as household sample terraced house; purpose built flat in a commercial or residential building; converted or not self-contained accommodation in a shared house or flat availability of amenities ù bath/shower 3 (3) exclusive, shared or no use as household sample ù inside wc 3 (3) exclusive, shared or no use as household sample ù central heating 3 (3) full, part or none as household sample cars (number of) 4 (4) 0, 1, 2, 3 or more as household sample floor level (lowest), of 7 (101) basement, ground, 1st/2nd, as household sample accommodation (scotland only) 3rd/4th, 5th/6th, 7th to 9th 10th or higher number of household 4 (35) top coded: 4 or more not included (accommodation) spaces in dwelling number of persons 12 (99) top coded: 12 or more not included (enumerated) in household number of residents in derivable 4 (99) 0, 1, 2 to 5, 6 or more household number of dependent children derivable 2 (99) 0, 1 or more in household number of pensioners in derivable 2 (99) 0, 1 or more household number of persons with derivable 2 (99) 0, 1 or more long-term illness in household number of persons in derivable 3 (99) top coded: 2 or more employment in household number of rooms 15 (19) top coded: 15 or more not included 10 iassist quarterly number of persons per room derivable 5 ranging from less than 0.5 to more than 1.5 tenure 10 (10) owner occupier or rented as household sample (public sector or private) wholly moving household 2 (2) yes (all resident household not included indicator members are migrants from the same address) or no individual information age 94 (111) single years 0 to 90, 91/92, as household sample 93/94, 95 and over status in communal not applicable 3 (4) visitor, resident staff or establishment resident non-staff type of communal not applicable 15 (35) hotal, hospital, nursing establishment home etc. country of birth 42 (102) as household sample migrants _ distance of 13 5, 10, 20 and 50 km bands; as household sample move (km) top coded above 200 km distance to work (km) 8 10 km bands; top coded as household sample above 40 km; 0_9 km band split 0-2, 3-4 and 5-9 economic position primary 10 (12) employee, self-employed, as household sample unemployed, student, retired etc. secondary 7 (10) as household sample economic position of family derivable 3 (12) employed, unemployed or head inactive ethnic group 10 (10) as household sample family head indicator 2 (2) yes or no not included family number 5 (5) used to identify individual’s not included family family type 8 (8) married or cohabiting couple as household sample family with or without children or lone-parent family gaelic language 5 (8) ability to speak, read or as household sample (scotland only) write gaelic hours worked weekly 72 (99) single hours 0_70, 71 as household sample to 80, 81 or more industry of employees and 185 (334) mainly third digit (groups) 60 (334) mainly second digit (classes) self-employed of 1980 sic of 1980 sic 11summer 1995 limiting long-term illness 2 (2) yes (individual has illness) as household sample or no marital status 5 (5) as household sample migrant geographical area 13 standard regions of england as household sample of former residence (with split of south east), wales, scotland, outside gb occupation 358 (371) mainly unit groups of 73 (371) mainly minor groups 1990 soc of 1990 soc number of higher 3 (7) 0, 1, 2 or more as household sample educational qualifications level of highest qualification 3 (3) higher degree, first degree, as household sample above gce a-level subject of highest 88 (108) mainly third digit of 35 (108) mainly second digit of qualification standard subject classification standard subject classification relationship to household 17 (17) 8 (17) head resident status 3 (3) present resident, absent, as household sample resident, visitor sex 2 (2) as household sample sex of family head derivable 2 (2) social class 8 (8) as household sample social class of family head derivable 8 (8) socioeconomic group 19 (20) as household sample term-time address of 4 inside or outside region of as household sample students and school children usual residence transport to work (mode) 10 (10) as household sample visitor _ geographical area 13 standard regions of england as household sample of residence (with split of south east), wales, scotland, outside gb welsh language (wales only) 5 (8) active use of (speak, read as household sample or write) workplace 5 inside or outside region of 5 inside or outside sar area of usual residence usual residence * the maximum number of categories as available on the full census database. 12 iassist quarterly given. geographical information for such items as workplace and migration (address one year before census) has been heavily grouped. this is because of the high likelihood of uniqueness of such information when used in conjunction with area of residence. dissemination the licensing and distribution of the sars is the responsibility of manchester university who have a contract with the esrc. the sars may be used for both academic and non-academic purposes. all higher education institutions (hei) are required to sign an end user licence agreement which makes the hei responsible for those members of their institution who are using the data. users within each institution must be either members of staff or students and must sign a further individual registration form which contains a binding undertaking to respect the confidentiality o the data. specifically, users have to guarantee not to use the sars to attempt to obtain or derive information about an identified individual or household, nor to claim to have obtained such information. furthermore, they have to undertake not to pass on copies of the raw data to unregistered users, and the census microdata unit has the responsibility of auditing their use of the data. they must sign a statement that they understand that the consequences of any breach of the regulations on the part of any user in a specific institution can lead to the withdrawal of all copies of the data from that institution. non-academic organisations sign a similar end user licence agreement and undertake not to allow the data to be user other than by their employees. the data is free for the purposes of academic research; to get the data free the researcher must be doing the research in an institution qualified to receive an esrc award, and the research must be funded either by the universities funding council or one of the research councils. when the data is used either by those outside the academic sector or by researchers in universities for sponsored research, a charge is made for the data. in order to encourage a high volume of usage of a product whose advantages may not yet be well appreciated in britain, these charges are being kept extremely low; an entire national sar can be bought for £1,000 + vat, and subsets of a county or local district for £500. 1 paper presented at iassist 21st annual conference may 9-12, 1995, quebec city, canada. references barnett, v. (1991) sample survey principles and methods, london: edward arnold bethlehem j g, keller w g and pannekoek j. (1990) disclosure control of microdata, journal of the american statistical association, 85: 38-45 breton, r, isajiw, w, kalback, w and reitz, j (1990) ethnic identity and equality, toronto: university of toronto press goldstein, h (1987) multilevel models in educational and social research, london: charles griffin and company her majesty’s government (1988) white paper (cm 430), 1991 census of population, hmso li, p (1999) ethnic inequality in a class society>, toronto: wall and thompson marsh c, skinner c, arber s, penhale b, openshaw s, hobcraft j, lievesley d and walford n. (1991) the case for samples of anonymised records from the 1991 census, journal of the royal statistical society (a), vol 154 (2): pp 305-340 marsh, c, dale, a and skinner, c (1994) safe data versus safe settings: access to microdata from the british census, international statistical review, 62,1, 35-53 paass g. (1988) disclosure risk and disclosure avoidance for microdata, journal of business and economic statistics, 6(4): 487-500 skinner,c.j,holt,d. and smith t.m.f. (eds)(1989) analysis of complex surveys, new york: wiley skinner, c, marsh, c, openshaw, s and wymer, c (1992) disclosure control for census microdata, university of southampton, mimeo. wolter,k.m.(1985) introduction to variance estimation, new york: springer verlag 18 iassist quarterly 2014 iassist quarterly discovering and accessing subnational statistics and geospatial data of east asian countries: trends and obstacles by jungwon yang1 abstract the increasing use of geographic information systems (gis), combined with the wider availability of sub-national statistics, has recently opened up new possibilities for more interdisciplinary academic research in social sciences. social science researchers have become more and more interested in combining geographic analysis, traditional quantitative and statistical methods in order to test hypotheses and present arguments in more effective ways. however, discovering, accessing, and using the international geospatial data and statistics is still a challenge for the researchers. as the organisation for economic co-operation and development (oecd) global science forum report (2013) noted, information about the existence of micro-data and the availability for the re-use is often difficult to find. the language barriers, as well as, legal, cultural, and technological obstacles often exacerbate the difficulties of re-using the discovered data. in this paper, i will review types of geospatial data and sub-national statistics of east asian countries have recently developed via central and local governments, and academic institutions. also, i will discuss obstacles researchers encountered while using such data in their research. keywords: : gis, sub-national statistics, east asia, geospatial data, china, korea, japan introduction the increasing use of geographic information systems (gis), combined with the wider availability of sub-national statistics, has recently opened up new possibilities for more interdisciplinary academic research in the social sciences. social science researchers have become increasingly interested in combining geographic analysis and traditional quantitative and statistical methods to test hypotheses and present arguments in more effective ways. researchers are more likely to be interested adoption of data citation and in the promotion of data sharing and its benefits. iassist quarterly 2014 19 iassist quarterly in interdisciplinary studies as visualizing data can facilitate interpretation and understanding of social scientists’ results. people often have some difficulty to interpret the results of quantitative analyses if they do not have a solid grasp of statistical analysis, such as the p value, r square, or confidential levels. using gis technology and geospatial data in interdisciplinary research helps people to understand their research even if they do not have comprehensive knowledge of the research methods. moreover, the share of the united states’ local governments adopting gis rose steadily from 20 percent in 1990 to nearly 88 percent in 1997. as of 2005, more the 60 percent of municipal web sites had begun to provide interactive gis features (ganapati, 2011). a wide variety of population and housing census data is freely available in various data formats and for a range of geographies and time periods from the united states census bureau website . the university of minnesota’s national historical gis project is processing and making freely available u.s. census boundaries in gis format as well as aggregate census data from 1790 to 2012 (https://www.nhgis.org). given the increased availability of geospatial and sub-national statistics, social science researchers can test their hypotheses and explain them in more effective ways. for example, sinclair et al (2011) uses gis-coded flood-depth and census data to examine the voting behavior of registered voters in new orleans before and after hurricane katrina. to measure the quality of the urban environment around public housing buildings in montreal, apparcicio et al (2008) uses multiple years of census data and geospatial data, such as the 2001 montreal urban community land use map, a landstat tm 7 image, the geobase, and the quebec topographical database. researchers can also create their own geospatial data to use in their research, potentially combining it with existing geospatial or other data. neckerman (2009) uses data both from gis measures and field observation in new york city to identify disparities in neighborhood conditions, by aggregating geospatial data that represents low-income neighborhoods. even though interdisciplinary studies using this gis technology and geospatial data are prevalent in the western countries, such as the united states, canada, and some european countries, the discovery, accession, and use of international geospatial data and statistics is still a challenge for researchers. part of the reason is noted in the oecd global science forum report (2013): “information about the existence of micro-data and availability for re-use is often difficult to find”. moreover, the language barrier, legal, cultural, and technological obstacles often exacerbate reusing the data. in this paper, i will review the kinds of geospatial and sub-national statistics on china, korea, and japan that have been recently developed by the central and local governments, and academic institutions. i will also address what obstacles researchers have encountered in using the data in their research. geospatial data and sub-national statistics in east asian countries macro-level economic and population data of east asian countries have been collected by the intergovernmental organizations, such as the united nations statistics division, the international monetary fund (imf), and the organization for economic co-operation and development (oecd). the korean and japanese central and local governments tend to collect and distribute sub-national statistics, including census, housing and economic data, and industrial enterprise data, as well as geospatial data, such as hydrology, elevation and land cover data. most statistics collected by the chinese government, however, are not available from the government website. china the primary department of the chinese government that collects sub-national statistics is the national bureau of statistics of china (nbs) . the nbs organizes collection of census data, which is held every ten years. it also collects, processes, and tabulates basic economic, social, and industrial enterprise statistics. most of the data collected by the nbs, however, is not open to public. on the chinese version of the nbs website , only brief summaries of the statistical yearbook for the prior three years are available. the chinese government does not collect geospatial data. subnational statistics and geospatial data for china is available from some fee based databases, such as the china knowledge resource integrated (cnki) , and the china data center . the cnki database provides china’s statistical yearbook (19812013), the chinese health statistics yearbook (2003-2012), the china financial yearbook (1949-2013), as well as provincial statistical yearbooks. yet, the datasets are written in chinese only. the cnki database does not have geospatial data for the region of china. english versions of statistical and geographic data for china have been collected by the china data center at the university of michigan. the china data center’s databases consist of two parts: china data online and china geo-explorer. china data online provides the census data (2000 and 2005) on provincial, county, and township levels, provincial and city economic statistics, and yearly and monthly industrial data. china geo-explorer , which is the web-based spatial data service of the china data center, aggregates government statistics, such as census, economic, and industrial data in a spatially integrated system. this spatial data service supports the creation and export of thematic maps based on built-in data. unlike arcgis software, however, users cannot combine the data which they create or acquire from other resources with the data of china geo-explorer. from geo-explorer ii, users can export geospatial data and statistics in the database as a shapefile for use in gis software. several projects from the u.s. academic sphere aim to build historical geospatial data. for example, the university of washington’s china in time and space (citas) data sets provides vectorized county level base maps of china and georeferenced socioeconomic data for the period 1982-1993. the citas data is currently archived in the socioeconomic data and applications center (sedac) ’s china dimensions data collection . the sedac provides citations, copyright and privacy policy information to users. even though the original china geo-explorer database is a fee-based database, the sidney gamble photo digital collection , created by duke university and applied to basic china geoexplorer software, is open to the public. the database currently features photographs dating between 1917 and 1932. the china historical gis website of the harvard yenching institute offers datasets such as time series geospatial data called chgis13 (221bc -1911ce), the 1820 data which includes spatial data of the qing dynasty territory for the year 1820, the chinaw data (1820-1893) created by the uc davis regional systems analysis project , and the 1911 data which contains spatial data for the provinces of anhui, fujian, gansu, guangdong, hebei, henan, hubei, hunan, jiangsu, shaanxi, shandong, zhejiang, and zhili for the year 1911. the chgis website also contains the 1990 citas data, the 1990 https://www.nhgis.org 20 iassist quarterly 2014 iassist quarterly gns place names, the 1997 citas provinces data, topographic images and raster data derived from gtopo-30 digital elevation model data, and other supplemental datasets. as thompson (2010) notes, non-governmental organizations tend to collect sub-national data targeted toward a particular context or issue. for example, the harvard yenching institute focuses on collecting geospatial data for the pre-modern period of china. the citas is more likely to focus on collecting and providing relatively current geospatial data. another interesting finding is that academic institutes which provide open geospatial data are more likely to archive their data into other bigger and stabilized academic institutions. for example, the citas data, aggregated by the university of washington, is currently deposited in the sedac and the chgis datasets. as a result, even if an academic organization is no longer able to provide their data to users, the original data, metadata, and citation information are available from other sources. korea korean central and local governments produce many sub-national statistics, including population, household, employment, prices, health, environment, agriculture, mining, energy, transportation, business, and education data, which are freely accessible from the korean statistical information service (kosis) web site . in addition to domestic statistics, the kosis website also provides international statistics compiled by the imf, oecd and un. the english version of the kosis website offers statistics at the provincial level, including special self-governing provinces (teukbyeoljachi-do), special cities (teukbyeol-si), and metropolitan cities (gwangyeok-si). however, municipal level data , such as cities (si) , counties (gun), districts (gu), towns (eup), townships (myeon), neighborhoods (dong), and villages (ri), are only available from the korean version kosis website. in the case of geospatial data of korea, the national geographic information institute (ngii) of korea provides aerial photos (19662012), satellite images (1973 -2004), orthophotos (2005-2011), and dem files (2005-2009) to users for a small fee . aerial photos from the 1940s and 1950s are freely available to the public. the standard image format is national image exchange (nix), which contains the image files compressed by jpeg 2000 as well as associated metadata. based on the information disclosure act of korea , all korean citizens, state agencies, local governments, governmentfunded institutions, and public authorities, as prescribed by presidential decree (for example schools, construction companies, non-profit corporations related to social welfare), can request to access official government documents (including electronic documents), drawings, photographs, films, tapes, slides and other similar recorded mediums. access to ngii’s data from a foreign ip address is strictly prohibited , so researchers cannot acquire the ngii data from overseas. the majority of provincial governments (gun) and metropolitan cities in korea provide gis information for their regions. for example, the seoul metropolitan government provides a city information map service in korean and english languages. the english version of the map service is an interactive online map which contains basic information for tourists, such as government office building locations, hospitals, schools, transportation information, and historic sites. the korean version of the interactive mapping service for the city of seoul additionally provides administrative boundaries, buildings, transportation, facilities related to welfare services, environment, land use and road information. users can overlay multiple kinds of geospatial data onto a map and freely download the customized map to as an image file. the city of seoul also offers open api service to korean citizens but based on article 21 of the public survey and overseas export ban act, international map mashup services are strictly prohibited . the map service web sites of incheon city , gyeonggi , south gyeongsang , and gangwon provinces also offer customized map services for overlay of aerial photos, a base map, transportation and other types of statistical information. the customized map can be downloaded to as an image file without fee. other provincial governments, such as jeju and south chungcheong provinces and metropolitan cities, such as busan, daegu, gwangju, and daejeon cities, also provide a gis map service, but it requires the up-todated internet explorer browser to work. it is quite cumbersome for other internet browser users, such as chrome and firefox to access their data. moreover, even if users use the internet explorer browser, they cannot access the data if they use out-of-dated internet explorer program. in sum, the korean central and local governments have a strong interest in developing and distributing geospatial data as well as sub-national statistics to the public. however, the gis services are mainly provided to serve korean citizens. overseas researchers often encounter difficulty to acquire and use the geospatial data. japan the statistics bureau and ministry of internal affairs and communication of japan collect both sub-national statistics and geocoded census data, and distribute via the statistics bureau website . current national level statistics can be found in the japan statistical yearbook series section of the website, in both pdf and excel formats. historical statistics from 1868 to 2011 are available from the historical statistics of japan section of the website. all the data from the website is available in excel format. the current sub-national statistics report called social indicators by prefecture 2014 can be accessed via the social indicators by prefecture section of the web site. it contains 608 social indicators and 571 items of basic data for the sub-national areas . time series data of sub-national statistics is available from e-stat , the official statistics site of japan. the time series data for prefectures, however, are not accessible on the english version of the site. geocoded census and socio-economic data also can be downloaded from the e-stat web site . currently, geocoded census data (2000, 2005, and 2010), the establishment and enterprise census (2001 and 2006), the economic census data (2009), and the census of agriculture and forestry data (2005 and 2010) can be downloaded in shapefile format. the japanese version the e-stat website also offers an interactive map service, where a user can download customized images. the statistics bureau of japan states that their data can be used for “compiling social indicators, by research institutes, universities and colleges for regional characteristic analysis, modeling to analyze regional development plans, and modeling to measure administrative performance, as base data for compiling welfare indicators by region, and for investigating and improving social statistics” . in other words, there is no restriction for accessing the data from overseas. another resource of geospatial data is the global map japan database of the geospatial information iassist quarterly 2014 21 iassist quarterly authority of japan (gsi) which is part of the ministry of land, infrastructure, and transport of tourism of japan. this database provides the land cover, vegetation, administrative boundary, population, and transportation data for the year 2000, 2006 and 2011 without fee. they also do not prohibit the access from overseas. based on the copyright laws of japan, as well as an international treaty, people can download data without consent from the gsi if they use the data in small quantities for noncommercial purposes. historic geospatial data for japan is available from harvard university’s japan data archive . currently, administrative boundaries for the tokugawa and meiji periods are available in shapefile format. these data are also available from the geodata@tufts database . in addition, elevation data for 1996 and administrative boundaries for the 1990s can be found on the website. all of japan data archive’s data is open to public; however, citation and copyright information is not available on the website. in sum, the collection and distribution procedures of government data in japan is highly centralized. they are more willing to share their geospatial data with overseas users, as compared to the korean and chinese governments. trends and obstacles data collection and distribution the central governments in all three east asian countries tend to collect sub-national statistics. most of the chinese sub-national statistics are not available from the government website, but all sub-national statistics for korea and japan are freely available from their statistics bureau websites. the chinese government does not collect geospatial data. most of the available geospatial data for china is therefore produced by overseas academics or institutions. korea’s geospatial data is collected and distributed by the ngii. the local governments of korea are eager to provide the geospatial data of their reasons as well. based on the law, the korean citizens and academic institutions located in korea can access this data without difficulty, but the data from ngii, the main provider of geospatial data, cannot be accessed from overseas. in the case of japan, all the sub-national and geospatial data is collected and distributed by the statistics bureau of japan. in addition, the geospatial information authority of japan (gsi) provides the geographic data. restriction of access to japanese geospatial data is quite minimal. copyright and open access willingness to support open data access differs across these countries. the japanese government puts little restriction on either domestic or overseas use of their geospatial data. most of the korean statistical and geospatial data created by the central and local governments are open to domestic users. but, the acquisition of geospatial data from overseas is restricted. both korean and japanese governments provide citation information, the data collection method, and copyright information from their websites. chinese statistics and geospatial data, however, are not open to public. rather, researchers must subscribe to a fee-based database to access the data. the historic geospatial data, created by academic institutions in the united states, is free to use. most of the academic institutions provide citation, copyright and disclaimer information for the data. reliability of data all three countries’ governments provide sufficient enough information about data collection procedures. the korean and japanese governments also state that they follow the imf’s data classification standard in the process of data collection. therefore, we assume that the reliability of data created by these governments is quite high. most of the academic institutions in the united states also provide information about their data collection procedures. yet, the china data center does not provide the sources of data or information regarding data collection procedures. given these conditions, the question of the reliability and the accuracy of data (economic data in particular) of the china data online have long been raised (chua, 2012) . language barrier as currently available sub-statistics and geospatial data of china are developed in academic institutions in the united states, the accompanying information, including metadata, is usually written in english. thus the language barrier to use the chinese data is relatively low for users, compared to korean and japanese data. the majority of sub-national statistics and geospatial data for korea and japan are only available in their respective languages. specifically, menu and the mapping options for the interactive map service website, where users extract customized maps, often cannot be translated into english (see figure 1 and 2). people who cannot understand korean and japanese languages will obviously be unable or have difficulties using these mapping services. figure 1 seoul metropolitan city’s gis portal (korean version). 22 iassist quarterly 2014 iassist quarterly technical concern some korean central and local government websites only can be accessed via the internet explorer browser. moreover, even if a user uses the internet explorer browser to access the website, the user cannot access the data without installing a current version of internet explorer. this technical requirement may add a small burden to overseas researchers’ attempts to acquire the data, especially as many united states libraries and computer labs at academic institutions do not allow to users to install new software without administrative access due to security issues. this may be a minor obstacle as many computers come with internet explorer installed, or researchers could most likely locate a computer with it installed simply to download the data. in the case of the china geo-explorer database, its thematic map service does not have an overlay option, so a user can only create a thematic map with a single set of data. also, census data from 2010 is now available in the china geo-explorer database, but the statistical data is not available from the china data online database. as a result, researchers have to find an additional database to get the numeric data. conclusion after investigating these data resources, i conclude that we can find quite a large amount of sub-national statistics and geospatial data from government, as well as academic institutions in the united states. however, the unreliability of data (china), the language barrier (korea and japan), copyright and legal restrictions (china and korea), and technical issue (china and korea), make it difficult for researchers to find and use data for east asian countries. in the academic library setting, collaboration among specialists and librarians in area studies, copyright, data, gis, government documents, and maps, will be critical to overcome these legal, cultural, and technical issues and to support social scientists’ interdisciplinary studies. reference apparicio, p., seguin, a. & naud, d. (2008) the quality of the urban environment around public housing building in montreal: an objective approach based on gis and multivariate statistical analysis. soc indic res. 86. p. 355-380. doi 10.1107/ s11205-007-9185-4 bosak, k. & schroeder, k. (2005) using geographic information systems (gis) for gender and development. development in practice. 15(2). p. 231-237 chua, h. (2012) indiastat and china data center online: an evaluation and comparison. reference reviews. 26(2). ganapati, s. (2011) uses of public participation geographic information systems applications in e-government. public administration review, may/june 2011. p. 425-434. ghiradeli, a. quinn, v. & foerster, s. (2010). using geographic information systems and local food store data in california’s low-income neighborhoods to inform community initiative and resources. american journal of public health. 100(11). p. 2156-2162. imai, h., keiko i., kazuo i., & koichi k. (2003). gis infrastructure in japan— developments and algorithmic researches. nontraditional database systems. 5. p. 130-145. li, y.(2012). the spatial variation of china’s regional inequality in human development. regional science policy & principle. 4(3). p.263-278 neckerman, k. et al.(2009). disparities in urban neighborhood conditions: evidence from gis measures and field observation in new york city. journal of public health policy. 30. p.264-285. sinclair, b., hall ,t., & alvarez, m. (2011). flooding the vote: hurricane katrina and voter participation in new orleans, american politics research. 39(5). p.921-957. doi:10.1177/1532673x10386709 thompson, k. (2010). data in development: an overview of microdata on developing countries. iassist quarterly. winter/spring 2010. ubaldi, b. (2013), open government data : towards empirical analysis of open government data initiatives. oecd working papers on public governance. no.22. oecd publishing. http://dx.doi. org/10.1787/5k46bj4f03s7-en west, amy. (2010) sources for international trade, prices, production, and consumption. iassist quarterly.winter/spring 2010. available online: http://www.iassistdata.org/downloads/iqvol334_341west. pdf notes 1. jungwon yang, international government information and public policy librarian. 240c clark library, hatcher south, university of michigan, ann arbor, mi 48109-1190. yangjw@umich.edu figure 2 e-state gis portal (japanese version). http://dx.doi.org/10.1787/5k46bj4f03s7 http://dx.doi.org/10.1787/5k46bj4f03s7 quarterly.winter/spring http://www.iassistdata.org/downloads/iqvol334_341west.pdf http://www.iassistdata.org/downloads/iqvol334_341west.pdf mailto:yangjw@umich.edu iassist quarterly 2014 23 iassist quarterly 2. http://www.census.gov/prod/www/decennial.html 3. http://www.stats.gov.cn/english/ 4. http://www.stats.gov.cn/ 5. www.cnki.nci 6. http://chinadataonline.org/ 7. http://www.chinadatacenter.org/ 8. http://chinadataonline.org/cge 9. http://citas.csde.washington.edu/data/data.html 10. sedac is one of the distributed active archive centers (daacs) in the earth observing system data and information system (eosdis) of the u.s. national aeronautics and space administration (nasa). sedac focuses on human interactions in the environment. its mission is to develop and operate applications that support the integration of socioeconomic and earth science data (http://sedac. ciesin.columbia.edu). 11. http://sedac.ciesin.columbia.edu/data/collection/cddc/sets/ browse 12. http://chinadataonline.org/gambleapp/cityclient35/ 13. since 2006 the statistics korea has been integrating the kosis that are produced by statistical organizations according to the project “integrated national statistics db”. initially 462 kinds of national statistics from 113 organizations were being provided through kosis. in 2009 an additional 165 kinds of statistics from 44 organizations were integrated into the national statistics databases. the project “integrated national statistics db” was completed after integrating an additional 50 kinds of statistics produced by 6 organizations, including the ministry of public administration and security into the integrated database in 2010. 14. http://kosis.kr/eng/ 15. for more detailed information about administrative division of south korea, please use the following link of the wikipedia (http:// en.wikipedia.org/wiki/administrative_divisions_of_south_korea). 16. the list of available geosptial data files can be found in the land area image information service system website of ngii (http://air. ngii.go.kr/info/sub01.do). 17. https://www.open.go.kr/pa/paretrieveinfodisclosureguide. laf?menuflag=11 18. https://www.nsic.go.kr/ndsi/ 19. http://gis.seoul.go.kr/seoulgis/englishmap.html 20. http://gis.seoul.go.kr/seoulgis/metroinfo.jsp 21. http://gis.seoul.go.kr/guide/flex_map.jsp 22. http://imap.incheon.go.kr/icmap/map jsp?viewtheme=basemap_airex 23. http://gris.gg.go.kr/ 24. http://gis.gndo.kr/ 25. http://map.gwd.go.kr/ 26. http://www.stat.go.jp/english/data/nenkan/index.htm 27. http://www.stat.go.jp/english/data/chouki/index.htm 28. http://www.stat.go.jp/english/data/shihyou/index.htm 29. http://www.e-stat.go.jp/sg1/estat/listedo?bid=000001052147&c ycode=0 30. http://www.e-stat.go.jp/sg1/chiiki/selectmapdispatchaction.do# 31. http://e-stat.go.jp/sg2/estatgis/page/download.html# 32. http://e-stat.go.jp/sg2/estatflex/ 33. http://www.stat.go.jp/english/info/guide/2011ver/05.htm 34. http://www.gsi.go.jp/kankyochiri/gm_japan_e.html 35. http://www.fas.harvard.edu/~chgis/japan/archive/ 36. http://geodata.tufts.edu/opengeoportalhome.jsp 37. the data sources information of china geo-explorer is available when a user downloads a shapefile from mapexport section. the shapefile package contains a codebook which explains the data source. http://www.census.gov/prod/www/decennial.html http://www.stats.gov.cn/english http://www.stats.gov.cn www.cnki.nci http://chinadataonline.org http://www.chinadatacenter.org http://chinadataonline.org/cge http://citas.csde.washington.edu/data/data.html http://sedac.ciesin.columbia.edu http://sedac.ciesin.columbia.edu http://sedac.ciesin.columbia.edu/data/collection/cddc/sets/browse http://sedac.ciesin.columbia.edu/data/collection/cddc/sets/browse http://chinadataonline.org/gambleapp/cityclient35 http://kosis.kr/eng http://en.wikipedia.org/wiki/administrative_divisions_of_south_korea http://en.wikipedia.org/wiki/administrative_divisions_of_south_korea http://air.ngii.go.kr/info/sub01.do http://air.ngii.go.kr/info/sub01.do https://www.open.go.kr/pa/paretrieveinfodisclosureguide.laf?menuflag=11 https://www.open.go.kr/pa/paretrieveinfodisclosureguide.laf?menuflag=11 https://www.nsic.go.kr/ndsi http://gis.seoul.go.kr/seoulgis/englishmap.html http://gis.seoul.go.kr/seoulgis/metroinfo.jsp http://gis.seoul.go.kr/guide/flex_map.jsp http://imap.incheon.go.kr/icmap/map http://gris.gg.go.kr http://gis.gndo.kr http://map.gwd.go.kr http://www.stat.go.jp/english/data/nenkan/index.htm http://www.stat.go.jp/english/data/chouki/index.htm http://www.stat.go.jp/english/data/shihyou/index.htm http://www.e-stat.go.jp/sg1/estat/listedo?bid=000001052147&cycode=0 http://www.e-stat.go.jp/sg1/estat/listedo?bid=000001052147&cycode=0 http://www.e-stat.go.jp/sg1/chiiki/selectmapdispatchaction.do http://e-stat.go.jp/sg2/estatgis/page/download.html http://e-stat.go.jp/sg2/estatflex http://www.stat.go.jp/english/info/guide/2011ver/05.htm http://www.gsi.go.jp/kankyochiri/gm_japan_e.html http://www.fas.harvard.edu/~chgis/japan/archive http://geodata.tufts.edu/opengeoportalhome.jsp vol271 by 14 iassist quarterly spring 2003 iassist quarterly spring 2003 15 a reflection on the past decade by chuck humphrey* the following discussion offers a perspective about the major accomplishments of iassist over the past decade. particular attention is paid to areas that are recognized as strengths of the organization and in which there have been notable successes. many of the actions and outcomes in these areas flow from the goals embraced by iassist at the beginning of the 1990ʼs. it is hoped that this discussion will assist in reviewing these goals and in setting directions for the future. most of the evidence in this summary is found in the business and program of the annual conferences of iassist. for many of our members, the iassist conference is a time to learn about new products, standards, services, and technology. these meetings keep iassist members at the forefront of our field. concomitantly, the direction of the profession both is reflected in the issues of these conferences and is shaped by the discourse and outcomes of these meetings. therefore, iassist conferences are an important source of evidence when considering the progress of this organization. six topics are discussed below that touch on the values expressed in the goals of iassist. while these topics are not strictly a re-expression of these goals, one can nevertheless find elements of the original goals in each of these six areas. the pulse of the membership the size of the membership has remained fairly stable over the past decade. while some have allowed their membership to lapse, overall there have been slightly more gains than losses. new members came from two primary sources. first, there has been an increase in the number of europeans joining iassist, many of whom work in national data archives. secondly, professionals new to data services, many working in academic libraries, account for the other membership gains. a consorted effort was made in the early 1990ʼs to expand membership among the staff of data archives. in 1993 under the direction of the uk data archive director, many archive staff attended iassist in edinburgh. this experience demonstrated the value of iassist to data archive staff and equally important, these new members brought a fresh enthusiasm to the organization. in subsequent years, several national data archives have supported their staff to attend the iassist conference. one encouraging indicator of this new commitment is that european attendance at iassist conferences held in north america has increased in recent years. in conjunction with this growth of european members, many participants joined iassist after attending the icpsr summer program workshop on social science data services. for this group, iassist serves as an organization where they can continue their professional development and can network with other data professionals. many of these new members are the only individuals on their campuses providing data services. iassist is the one organization to which they belong where they find colleagues doing the same kind of work that they do. during the past decade, structural changes occurred in several universities that altered the institutional location of data services. computing centres that had been actively supportive of social science data services, for example, princeton university, northwestern university, and the university of alberta, began divesting themselves of client services. the paradigm of central computing centres shifted from a services-based to a utility-driven organization. under this model, computing became another outlet in the offices and labs on campus. along with the light switch and the power and phone outlets, institutions added a computing network outlet. while computing centres became uninhabitable for data services, academic libraries welcomed the it expertise of data services staff. in many instances, libraries absorbed data services that had become orphaned by the movement toward utility-driven computing centres. iassist members played a supportive role in the relocation of data services in several universities during these times of structural change. our members work in a rich variety of organizational settings and consequently, there is a wide range of models for providing data services. one model that was imitated by a number of institutions was an amalgamation of government information, maps, by 14 iassist quarterly spring 2003 iassist quarterly spring 2003 15 and data into a single administrative unit. this saw the convergence of support for statistical and spatial data. another significant environmental change having an impact on the institutional location of data services was the digital library initiative. this movement started to take root in academic libraries at the same time that several social science data services were moved into libraries. the digital library seemed to legitimize the incorporation of data services in the eyes of some academic library directors. other directors recognized data as a valuable research resource that belongs in the library. the confusion between data services and the digital library unfortunately has never been resolved. for one thing, the digital library has been largely dominated by digitization projects and has failed to establish a strong connection with data services. consequently, attention to the creation of digital collections has often overlooked data services as part of this movement. as interest in social science data increased in areas of the globe outside of australia, europe and north america, iassist introduced two new initiatives to support these activities. a spin-off from these programs has been the addition of new members to iassist. first, an outreach program, which is discussed in more detail below, was formally established during the 1996 conference in minneapolis. this initiative seeks to support staff in countries where social science data services are just taking root. the second initiative was to create a new region for africa within iassist. with a leadership base in the sada and a growth in social science data activities on this continent, the prospects for new members from this region are encouraging during this period, some turnover occurred in the membership due to career changes for some people and retirements for others. a testimony to the quality of the people in iassist is that contacts and friendships have been maintained with many of those who have changed careers. we also have the exemplar role model of a few retired members who have remained active in iassist! as an organization, iassist needs to support professional data staff through training and upgrading, promoting social networking, contributing to standards development, disseminating information and knowledge, generating collaborative work, initiating and coordinating research, and offering a forum for professional issues. many of these activities are conducted on a peer-to-peer basis, which makes iassist membership all the more important to professionals in data services. communications communication occurs on many fronts. there is memberto-member contact, official communiqués from the leadership of iassist, knowledge dissemination through the iassist quarterly, and the public face of iassist. these communications require tools and organizational structure to occur, and the work and leadership of the publications committee has been essential is this regard. iassist has provided an electronic means for members to communicate with one another through an email discussion list, which started in 1991 initially hosted by princeton university before moving to yale and then to columbia, where it currently resides. this is a closed list to members and is used primarily by most members as an extended reference service where advice or information is sought from the wide realm of expertise among the membership. in recent years, a membership directory has been produced annually from the records of the treasurer and distributed to members in good standing. the administrative committee of iassist has decided not to give or sell mailing labels from this directory to advertisers, protecting the membership from unwanted junk mail. the primary purpose of the directory is for member-to-member contact. in 1995, the iassist website was introduced to provide the organization with a public presence on the internet. iassist members are good citizens of the internet and the website was intended to reflect this spirit through the organizationʼs support of open access to information. the website has undergone two phases of development. the initial phase focused on providing a description of the organization to the public. it tried to communicate what iassist is and why it is important. the second phase remodeled the site and incorporated material that members will also find useful. this design tried to balance the promotion of the organization with tools created by members that are important in the work that we do. as development continues with this phase, creative ways of using this medium are being sought to strengthen the organizationʼs outreach to the wider data community and to deliver training and educational programs. the iassist quarterly (iq) continues to be a print publication that is mailed to members in good standing. the production of the iq, however, has long involved electronic stages in its preparation. in recent years, this electronic copy has been converted to pdf format and made available on the iassist website. the digital version of the iq on the website is made available at approximately the same time as the print edition to the membership. open access to the iq contributes to the wider value of the internet (as a public service of iassist), provides the organization with greater visibility, and shares the expertise of the membership more globally. other structural changes have happened in the organization that also improved communications within iassist. in 1998, a treasury group was organized creating three assistant treasurers located in three regions: the u.s., 16 iassist quarterly spring 2003 iassist quarterly spring 2003 17 europe and canada. this new structure within iassist allowed for a flexible way of dealing with the bridge financing of conferences in these regions and for collecting membership fees. beginning in 2000, the membership committee was restructured around the regional secretaries. this change was introduced to provide a more direct contact with members in a region. both of these initiatives were begun to establish a closer relationship between activities within iassist and its members. the communication initiatives discussed above have been important for the life of this organization. another assessment is found in the research conducted by karsten boye rasmussen and repke de vries where they investigated iassist as a virtual community. their findings suggest that iassist can make better use of network communication technology. physical distances on this planet are an obstacle to our members meeting together, which is further discussed below under the topic of regionalism. discovering a balanced use of network communications with in-person conferences is a priority of the organization. metadata, standards, and technology technology is a means to an end for those of us working in data services. there is no doubt, however, that technology at times seems to be in the driverʼs seat. this was particularly true in the early 1990ʼs when many data services struggled with the migration of their data and documentation from mainframe environments to network technology. iassist conferences throughout this past decade have held sessions dedicated to the issues of technological change in data services. in 1994 at the san francisco iassist conference, prototypes of data extractors using a web interface were demonstrated. these early examples confirmed the need for data documentation standards to facilitate web extraction services. discussions around codebook standards evolved from roundtable discussions into interest groups in 1993 when a working group on codebook documentation of social science data and a group to create documentation guidelines for data producers were established. in 1995, the data documentation initiative (ddi) merged earlier working group interests and became the focal point for discussions about documentation standards within iassist. many iassist members have contributed to the ddi standard and continue to lead in its development. the maturation of ddi is one of the major success stories of this organization. the work of our european members through cessda and national data archive initiatives made further outstanding contributions in metadata and extraction tools. among these projects was the integration of a major social science thesaurus with data catalogues, multi-catalogue searching over the web using multiple languages, and nesstar. these major contributions have been the focus of discussion at many iassist conferences over the past decade. like ddi, they are successes to be celebrated. the ottawa iassist in 2003 held a session in which major data archives described how recent technology changes have been integrated into their data processing operations. this is another example of the convergence of changes in metadata, documentation, and technology. these new tools are changing the daily business of the major data archives. the intellectual contributions of iassist members to these very significant developments over the past decade are achievements in which the organization can take great pride. confidentiality, intellectual property and privacy the issue of confidentiality and access is a long-standing topic in iassist. a delicate balance exists between the protection of the privacy of the individuals from whom data have been collected and claims for access to such data for legitimate research purposes. the millennium bug seemed to draw new public attention to the vulnerability of the massive amounts of data existing on individuals. one consequence has been a negative public reaction to the potential uses of digital information on individuals. protective legislation has appeared recently in many jurisdictions that pose serious threats to legitimate research use of such data. these concerns have been the topic of many plenary and concurrent sessions at iassist conferences. institutional practices and procedures have been presented, such as the norwegian model of a research data ombudsman. technological fixes have been debated, including synthetic files and research data centres. at the ottawa conference the issue of privacy was expanded to incorporate spatial as well as statistical data. these topics will not go away any time soon and iassist needs to continue addressing them to better inform the public, policy-makers, researchers, and the rest of the data community. data commodification surfaced over the past decade with the commercial success of the internet. while the recent collapse of the dot-com industry has temporarily tempered interests in commercializing data, the issue itself will not go away. the next resurgence of e-business will drive a new wave to commodify data. open scientific research is threatened by ownership claims to data. consequently, this organization will need to steer the discussion about data ownership to one of data stewardship, and to be prepared to address ownership issues in terms of the barriers that it presents both to data access and to preserving data. 16 iassist quarterly spring 2003 iassist quarterly spring 2003 17 new frontiers in research data the frontiers of social science research data continue to expand. one important new area has emerged from advances in computational methods in qualitative research. in fact, developments in this area have led to the formation of data services dedicated to qualitative data. of particular note is qualidata in the uk data archive, which is responsible for the acquisition, preservation, and dissemination of qualitative research data. recent iassist conferences have held workshops and sessions on both the archiving of and providing data services for qualitative data. for example, a workshop was presented at the ottawa conference in 2003 entitled, “everything you ever wanted to know about preparing qualitative data, but were afraid to ask.” in a sense, we have witnessed qualitative data come of age over this past decade with an expectation that secondary uses of this type of data will increase the overall value of qualitative data. with the popularization of pc-based geographic information systems in the 1990ʼs, spatial data have become another rapid growth area in social science data. the affiliated use of geo-referenced statistical data with corresponding spatial data has added new demands for aggregate statistics from data services. starting slowly in the early 1990ʼs, gis is approaching tidal wave strength as applications for spatial analysis sweep across social, health, economic, business, and educational research. iassist conferences have incorporated spatial data as a focal topic in recent years. the theme of the 2000 conference in evanston was “data in the digital library: charting the future for social, spatial and government data.” in 2003, a plenary was dedicated to spatial data in additional to concurrent sessions that dealt with gis data issues. historical data have always been part of social science data interests in iassist. more recently, the development of public use microdata files from historical censuses has taken on new prominence. following the success of the ipums project at the university of minnesota, a new international historical census microdata program has emerged, ipums-international. these projects are contributing a wealth of new data for researchers. an additional spin-off from these projects is a network of collaboration between researchers and professionals in data services. all three of these new data frontiers share metadata, preservation and access issues with existing social science data. the developments in ddi have been applied in the ipums project. furthermore, enhancements to the ddi standard have resulted from the challenges of documenting these historical data sources. similarly, the lessons, methods, and tools used in documenting quantitative data are being adopted with qualitative data. again, unique aspects of qualitative data are contributing to metadata practices established for quantitative data. the archiving of spatial data lags behind other social science data types. however, many gis researchers are now looking upon the preservation of spatial data as an extension of other social science data. iassist can play an important role in this development. these three data areas warrant special mention because of their growth and contribution to new research in the social sciences. this does not, however, detract from the continued developments in quantitative data in the social sciences. projects in these areas continue to see major advancements and include issues such as synthetic data, probabilistic linkage of large administrative databases, and the management of large consumer expenditure databases. activities in these areas will continue to challenge iassist members to find appropriates ways of preserving and providing access to these data. outreach and regionalization beginning in 1996, the iassist administrative committee initiated intentional outreach to support participants financially from countries just developing data archives and services. post cold war changes in europe have resulted in the emergence of a number of new national data archives. similarly, the post apartheid period in south africa saw the development of the sada. iassist has been supportive of the staff from these new archives. some financial assistance has been made available to encourage their participation in our conferences. in addition, conference programs have included opportunities for speakers from new data archives. this has helped build contacts for those new to data services and archiving as they work to establish a network of professional connections. the work of the international outreach committee, in collaboration with ifdo, unesco, and the conference host-institutions, has been another iassist success story. with social science data services and archives expanding around the globe, the reliance on an annual conference to keep the community connected is not as effective as it once was. while the size of the community remains relatively small, the distances of worldwide participation create serious difficulties. one approach worth examining is the formation of strong regional iassist communities structured around the membership committee and regional secretaries. in a sense, europe has been functioning like this since the formation of cessda and its introduction of expert seminars. similarly, the biennial meeting of the icpsr official representatives has served a similar function in convening mostly north americans outside of iassist conferences. ways of engaging members to meet within regions without eroding the attendance at annual iassist conferences need to be explored. conclusion iassist remains a membership-based organization committed to institutional solutions to preserving and 18 iassist quarterly spring 2003 name: job title: organization: address: city: state/province: postal code: country: phone: fax: e-mail: url: iassist international association for social science information service and technology association internationale pour les services et techniques d'information en sciences sociales i would like to become a member of iassist. please see my choice below: options for payment in canadian dollars and by major credit card are available. see the following web site for details: http://datalib.library.ualberta.ca/membership/ membership.html $50 (us) regular member $25 student member $75 subscription (payment must be made in us$) list me in the membership directory add me to the iassist listserv membership form the international association for social science information services and technology (iassist) is an international association of individuals who are engaged in the acquistion, processing, maintenance, and distribution of machine readable text and/or numeric social science data. the membership includes information system specialists, data base librarians or administrators, archivists, researchers, programmers, and managers. their range of interests encompases hard copy as well as machine readable data paid-up members enjoy voting rights and receive the iassist quarterly. they also benefit from reduced fees for attendance at regional and international conferences sponsored by iassist. membership fees are: regular membership: $50.00 per calendar year. student membership: $25.00 per calendar year. institutional subcriptions to the quarterly are available, but do not confer voting rights or other membership benefits. institutional subcription: $75.00 per calendar year please make checks payable, in us funds, to iassist and mail to: iassist, assistant treasurer joann dionne 50360 warren road canton, mi 48187 usa providing access to data. the mission of iassist has not diminished over the past decade. if anything, its mandate has expanded as a result of the changes discussed above. professional development will remain as important tomorrow as it has been over the past ten years. iassist clearly has a significant role to play in this area. furthermore, a “voice for data” will be needed on the international scene tomorrow as much as it is today. the representation by ifdo and iassist this past year to get research data incorporated within the unesco charter on the preservation of digital heritage is an example of the type of leadership required at the international level. currently, advocacy for research data at the international level has very few voices, and the voices that do exist are without coordination. the committee on data for science and technology (codata) and iassist hold many of the same interests but rarely if ever share the same platform. iassist needs to look for strategic alliances with other organizations concerned about the preservation and access to research data. the talents and capacity of the membership of iassist have made this a successful organization. we may be small in numbers; but everyone in our number, counts. the recruitment of gifted new members, whose vitality will carry organizationʼs mission forward, is crucial to the future of iassist. http://datalib.library.ualberta.ca/membership/membership.html http://datalib.library.ualberta.ca/membership/membership.html iassist quarterly spring summer 2009 31 ddi 3 development at dda abstract the danish data archive (dda), a national data bank for researchers and students in denmark and abroad, is dedicated to the acquisition, preservation and dissemination of machine-readable data created by researchers from the social science, health science and history communities. the dda has a need to convert existing osiris3 and data documentation initiative (ddi) version 2 documentation to ddi 3 and to integrate information from various metadata sources across the data life cycle. health sciences data in particular require an enhanced metadata structure. our efforts to make these transitions and improvements and to build useful tools have required a thorough grounding in ddi and we provide our perspectives in this paper. open metadata structure approach since there are only a few tools for ddi, we decided to dda is a member of the data documentation initiative alliance (ddi alliance) and has been involved in the development of the ddi 3 standard from the beginning. in general the dda finds the ddi 3 standard a very flexible mechanism to capture the metadata structures defined by the social science and health science communities. as a conceptual model ddi 3 can provide inspiration to research communities on how to structure metadata most efficiently for reuse across the data life cycle. as ddi 3 is an open, widely accepted standard across data archives, there is an opportunity to influence its development and to draw upon the experience of the community around it. and the community around the ddi 3 standard is reaching out to other communities. for example, ddi 3 was developed in line with other widely accepted standards, including the statistical data and metadata exchange (sdmx) standard and the metadata registry standard (iso/iec 11179), thus facilitating metadata interoperability4. a proposal to define controlled vocabularies using genericode has also drawn upon the expertise of the wider ddi community through a crossarchive working group. and work on using the resource description framework (rdf) together with ddi 3 is also being undertaken.5 all of this is leading towards more tools in the ddi tool box to apply to metadata creation and use. advantages for archives the primary business model of the dda is cleaning and enhancing the quality of data and metadata without changing the layout of the metadata and data deposited by researchers. the ddi 3 standard offers some key benefits in this process and possibilities that were not available with previous metadata models. ddi 3 brings with it strong referential and versioning features for fine-grained metadata elements. the dda will implement these options by reusing metadata structures across studies with the intent to focus more attention on the content of the surveys while also marking them up to fulfill long-term preservation goals. in time reuse can extend to data mining across the metadata collection. with the osiris standard not being able to document hierarchical, panel, cross-sectional, or follow-up studies, ddi 3 provides a useful alternative, especially given its focus on extended reuse and inheritance. the health sciences unit within the dda is particularly interested in these features of ddi 3. regarding long-term preservation planning and storage, the ddi 3 and its community are joining forces both with technologies like fedoracommons6 and archiving standards such as the metadata encoding and transmission standard (mets)7 and preservation metadata implementation strategies standard (premis)8 to incorporate additional metadata about archiving. all components come together when defining an oais9 implementation of a data archive for the social science, health and history domains. open development with the many standards and stakeholders in play across the current metadata landscape, no one organization can be expert in all subjects, and thus the straightforward solution is collaboration. of course, collaboration itself is often not straightforward but more like a wave or a moving target. the same can be said for the standards themselves as they develop, mature and are constantly being enhanced with new features and as additional standards emerge; this is the joy of information technology. over time we have seen a greater focus on information exchange and tons of exabytes of it. the strategy at dda is to upgrade our data for this by jannik jensen1 and dan kristiansen2 32 iassist quarterly spring summer 2009 exchange, make it available within various networks and capture its reuse and relationships. developing metadata upgrade and data ingest systems built upon ddi 3 technology is key in this process.10 to execute a software project in the current environment characterized by high complexity and uncertainty, the dda has decided upon an open source approach aligning with the community and the standard itself. this approach brings potential collaboration, knowledge exchange and product hardening informed by the feedback of others into the end product and offers wider integration possibilities with other open products in the domain, for example, the questasy project at centerdata in the netherlands11. even though the dda is creating open source software, we acknowledge the closed source initiatives such as colectica by algenta technologies12 as clear assets for the ddi community. the more vendors that produce tools and applications for ddi 3 the greater the adoption is likely to be. the synergy effect of both the closed and open camps developing ddi 3 software is a real benefit for both end products and the evolution of ddi. open and configurable reusable solution in 2007 the dda took an active role in the ddi foundation tools program13, a collaborative endeavor with stakeholders from several institutions contributing time and money to developing ddi 3-based tools using an outline of common it development tools14. specifying an operatingsystem independent approach based on java15 with incorporation of various open source projects, the project resulted in ddi 3 tools licensed under an open source license16. the past and current it developments in the ddi 3 field are following an updated version of these recommendations. the first step for the dda was to develop a common object model in java using the apache xmlbeans technology17. around this object model the dda has built tools for secondary validation, urn element generation, and a mechanism to extract studies contained within a grouped structure18. the first ddi 3 developments led to the design and developments of a centralized suite approach instead of relying on pooling separate tools into a system. the aim of the suite approach is to minimize the footprint when dealing with large amounts of xml. work on an editing tool for ddi 3, which is being built using the eclipse rich client platform19 as the front end, began in the fall of 2008 with architecture design reviews provided by the open data foundation. the editing suite is primarily designed for configuration and reuse/extension. the rationale for these design decisions arose out of the need for flexibility in the implementation of metadata in ddi 3 and the need to tweak system components for customized needs 20. this flexibility makes it possible to, for example, change the xml persistence layer or how ids are generated for identifiable ddi 3 elements. conclusion so far the dda has built a sound basis for its ddi 3 strategy, and the major components have been identified and constructed. the work ahead is to harden and extend these components. to lead this process, the dda has compiled a functional requirements document following the ieee std 830-1998 21 . on the agenda are several topics of interest, including functionality to help create and update longitudinal metadata and functionality to facilitate the review process of a study. to enhance the reuse functionality, an internal lookup and resolution service for ddi 3 urns is scheduled. internally the dda has allocated two additional user representatives within the dda for the project to ensure end user collaboration, interaction and acceptance of features and the user interface layout, leading developments in a more agile direction. as ddi 3 is growing in use and becoming part of production systems, more services will evolve around it, including services not currently identified. the software approach the dda is taking is to deliver components that are very close to the standard but not tied to a particular vendor, thus ensuring agility in implementation and reuse in planned and future systems. notes 1 jannik jensen is a software developer at the danish data archive in odense, denmark. 2 dan kristiansen is a software developer at the danish data archive in odense, denmark. 3 anne sofie fink et al. (2003). “preservation of knowledgedata processing in the danish data archives.” iassist quarterly, 2003. 4 arofan gregory et al. (2009). “metadata.” ratswd working paper 57, 2009. 5 patrick carmichael and agostina martinez garcia (2009). semantic technologies to support teaching and learning with cases: challenges and opportunities. university of cambridge, 2009. 6 http://www.fedora-commons.org/ 7 http://www.loc.gov/standards/mets/ iassist quarterly spring summer 2009 33 8 http://www.loc.gov/standards/premis/ 9 http://www.iso.org/iso/iso_catalogue/catalogue_tc/ catalogue_detail.htm?csnumber=24683 10 jannik jensen and dan kristiansen, (2010). “building a modular ddi 3 editor.” ddi working paper series, 2010, doi:10.3886/ddiusecases02. 11 http://centerdata.nl 12 http://www.colectica.com/ 13 http://tools.ddialliance.org/ 14 open data foundation (2008). guidelines for tools development and recommendations for operating environment, 2008. 15 http://www.java.com/ 16 http://www.opensource.org/ 17 http://xmlbeans.apache.org/ 18 jensen, jannik; pascal heus; joachim wackerow; jeremy iverson; dirk roorda; rene van horik. “ddi and related tools: next generation tools for converting, displaying and visualising data.” chair wendy thomas. presented at the annual meeting of the international association of social science information service and technology (iassist), palo alto, ca, may 2008. 19 http://wiki.eclipse.org/index.php/rich_client_platform 20 jannik jensen and dan kristiansen, (2010). “building a modular ddi 3 editor.” ddi working paper series, 2010, doi:10.3886/ddiusecases02 21 ieee std 830-1998 ieee recommended practice for software requirements specifications – description http:// standards.ieee.org/reading/ieee/std_public/description/ se/830-1998_desc.html vol192 20 iassist quarterly the term “national archives” usually conveys an image of a large organisation with a staff numbering several thousands, as in the national archives of canada or the united states, or several hundreds, as in most of the national archives in europe. it should be made clear from the outset, however, that the national archives of ireland must be considered on a much smaller scale. ireland is a small country on the periphery of europe with a small population (just over 3.5 million in the republic of ireland) and the national archives of ireland in dublin can be seen to reflect the size of this population base. not only are we smaller than most national archives, we are smaller even than the specialist divisions of many national archives. we are smaller, for instance, than the center for electronic records in the us national archives. our total staff numbers 35; our total professional staff numbers 13. we are, therefore, comparable in many ways to some of the state archives in the united states. in fact on the evidence of richard cox’s recent study, the first generation of electronic records archivists in the united states, there are many points of similarity between the situation obtaining in state archives in the united states, and the situation obtaining both in the national archives of ireland and among the archival profession generally in ireland.2 the national archives of ireland has existed under this name only since 1988 when the national archives act (1986) came into effect, though the constituent parts of our organisation, the state paper office and the public record office of ireland, have existed separately since 1702 and 1867 respectively and have been part of a de facto amalgamation since the late nineteenth century. the national archives act has radically transformed the role of our organisation, however, and has given us responsibilities similar to those of the national archives of canada and australia. we now have a thirty year rule of access for government records, and no such documents may be disposed of without the written consent of the director of the national archives. our act placed an enormous burden on us, with accumulations of documents dating from the beginning of the state and before, and formerly not covered by legislation, having to be processed. at the same time as our responsibilities have expanded so dramatically, our traditional business has also been increasing significantly. we have an annual readership of 17,000. this may not be huge by the standards of most national archives (according to a recent notice posted on the “archives” listserv, the number of people accessing the new york state archives gopher in january 1995 was 17,000, the same as our readership for the whole of last year) but we have experienced a huge increase in public access in a generation amounting to a tenfold increase in the last twenty-three years. our user profile is very different to that of a data archives or library. over 50% of the readers’ tickets which we issued in the first three months of this year were issued to people undertaking genealogical research on their own families. this statistic has a bearing on the sort of service we must provide and how priorities are addressed. tourism is ireland’s second largest industry (after agriculture). there is a huge irish diaspora in north america, australia and the uk, and it is from this that most of the tourist traffic comes. the roots factor is an important element in all of this and we are, whether or not we would wish to be so, part of the roots industry. some 37% of our readers come from abroad, most of them tracing their roots, and they form a constituency which we must be careful to service. apart from genealogists, amateur and professional, the remainder of our readers are divided between academic researchers, local historians, teachers and trainee teachers, and a considerable body of legal searchers. as to the documents being produced, the emphasis here is also heavily on genealogy. the household returns of the 1901 and 1911 censuses (which are, respectively, the earliest irish census for which full household returns are extant, and the latest census for which the household returns are open to inspection) accounted for 42% of all documents produced to readers in the first quarter of this year. the 1901 census alone accounted for 26% of all documents produced in this period. far behind the census, the next largest categories were: modern departmental records (22%) eighteenth and nineteenth century state papers (11%) and testamentary records (7%) like most national or state archives, we must face two ways at once. we are expected to provide a service to a research sober ways, politic drifts and amiable persuasions; approaching the information highway from the dusty trail. by ken hannigan1 national archives of ireland 21summer 1995 public which is largely composed of genealogists, and we must provide a service to government, to appraise its records which must be authorised for disposal or accepted for transfer. we must balance our obligations to the research public and to government with our obligations to a third constituencyposterity. we must preserve an adequate record of our own time and continue to preserve the records of previous ages which have been entrusted to our care. we have 13 professional archivists on our staff. this is a small enough number, but in relative terms these 13 constitute a sizeable proportion of the professional body of archivists in ireland. total membership of that body at present numbers 67. increasingly, candidates for jobs in archives are required to have a post-graduate qualification in archival studies . in the past twenty-five years archivists have professionalised, indeed it could be said that it is only in the last twenty five years that the profession has been defined in ireland. most current holders of the diploma in archival studies are graduates of the only archives school in ireland, that in university college dublin, and so that school, to a large extent, controls entry to the profession. however, many of us in mid-career, particularly in the state sector, have no specialist archival qualification. we are all arts graduates, however, most with history degrees. because of our low numbers, there is no separate irish professional organisation for archivists; we form an irish region within the society of archivists, the bulk of whose members are in the united kingdom. our professional focus and contacts, therefore, have tended to be with our colleagues in the united kingdom with whom we have much in common. and so to the dusty trail. dust is certainly a metaphor with which traditional archivists in ireland and the uk are familiar, though not, perhaps, entirely comfortable. dust and decay are an essential part of our popular image and this image is one of the problems which we face in approaching the superhighway. it is likely that a word association test administered to the average person in the street in ireland would result in a string such as “archives, dust, decay, dead, buried”. “buried in the archives” is a phrase we frequently hear used in relation to documents, or even in relation to ourselves as archivists! thus the following statement which a national daily newspaper in ireland recently published as part of an interview with one of the country’s leading popular composers is probably fairly representative of popular attitudes: “i honestly do believe that merely sticking with the past is for archivists. forging new forms for the future, on the other hand, is for the living”.3 well, certainly there is a sense in which archivists are seen to be, if not actually dead, then as having escaped from life. we hear frequent reports of people being told by career guidance counsellors or teachers that a career in archives is an option for those of a shy retiring nature, or timid disposition, who might find an alternative, teaching, for instance, or career guidance counselling, perhaps, too hard on the nerves. we tend to have a cobweb-enshrouded image largely based (as richard kesner has identified it) on the popular notion that archivists are antiquarians, that we are a little removed from everyday life.4 we are not entirely blameless in this regard. some of us have cultivated the image of the antiquarian, perhaps many of us are attracted by this selfimage and have even been attracted to the profession by it. so there may be something of a self-fulfilling prophesy at work here, as the world of traditional archives has attracted those who have consciously not wanted to be part of a thrusting, aggressive, brash, profiteering, macho world. we are mostly history graduates; we are people who put posterity above profit and power. the world of archives is also a very stable one. within the archival profession in ireland today most us who have been there for ten years or more are doing the same jobs which we were doing ten years ago -and in the same organisations. few of us have experienced anything else in our professional lives. it is not typical of the organisations with which we do business, the organisations for whose records we are responsible. it is certainly not typical of the it people with whom we come in contact but who disappear out of our orbit again with bewildering speed. this stability has left many of us locked into practices and perspectives which are antidynamic. and as most of us are burdened by the daily demands of keeping a public service going and overwhelmed by backlogs of unlisted and unappraised records, it is frequently not until systems break down that we consider change. there is a large element of this present in our response to computers. there was a time, not so long ago, when archivists could get away with a statement like “i know nothing about computers” and even make this sound like a virtue. we were helped in this by the fact that our favourite constituency of readers historians -by and large also tended to spurn computers. it is of course no longer fashionable for archivists to admit that they know nothing about computers. even the most obdurately antiquarian of us have by now realised that computers are, or should be, essential tools of the trade. but we are not yet really at home with them. we have not as a profession come fully to terms with the impact of automation. it is a fact that the largest special interest group within the society of archivists is the it group, but within that group to date we have tended to concentrate very narrowly on a single aspect of computerisation, and the most popular events organised by that group are software demonstrations. we are terribly interested in learning how computers can help us to continue doing the things we have always done in the ways we have always known and loved. we have come to the conclusion that computers are probably a good thing, we certainly want to know a little more about them, but really, we are not technical people and we still tend 22 iassist quarterly to revel a little in the fact. these attitudes put us at a considerable disadvantage in coming to terms with the wider aspects of computerisation. automation has implications for the specialised functions of “traditional” archives in three main areas. firstly there is the question of automating the archival tasks, accessioning, repository management, and so on. this should not pose any difficulties for traditional archives. we are basically talking about stock control here, something which is eminently suited to automation. secondly there is the obligation to provide an efficient and reliable service to readers and potential researchers, including the obligation to provide and disseminate information about our holdings. we are in the information business, though we do not all see it this way, and computers are tools for information management. thirdly there is the increasingly worrying question of what to do about the records generated by computers. these three aspects cannot be divorced; our failure to come to terms with the first two leaves us ill-equipped to deal with the third. many people from outside the world of archives, and even some archivists, are surprised at the failure of archives in europe to automate more rapidly. in a recent issue of the american archivist, ronald weissman expressed astonishment at finding a newly-created series of handwritten finding aids at the new state archives in florence5 . there would be no difficulty in finding similar instances of archives all over ireland and the uk tenaciously holding on to the old methods. there are two main problems which we face in automating and which partly, though not totally, explain our slow progress. one is quite simply the question of resources. it seems that archives everywhere are low on the priorities of governments and funding agencies. the country will not grind to a halt if the archives fail to function efficiently. the business of archival management does not generally attract largescale commitment of resources. our very modest degree of computerisation in the national archives of ireland has been achieved in a piecemeal manner and without the benefit of a specialist it unit. the second problem attaching to automation is potentially more difficult to resolve. effective automation of archives demands consistent descriptive standards, ideally ones which are universally accepted. unfavourable comparisons are frequently made between the extent of our computerisation and that obtaining in even fairly modest county libraries where users see the benefits of online catalogues and barcoding systems. there are of course some fundamental differences between archives and libraries, though these are not perceived by an impatient public, despite the efforts of both the professional librarians and professional archivists to delineate the two professions. the fact that our collections come readymade, that rather than being a continuous series of single-level items our collections sometimes involve complex arrangements, and that retention or recreation of the original order is a cardinal rule of archival description, these have all posed problems for traditional archives the world over in their attempts to computerise their services and exchange information on their holdings. despite some heroic, some would say quixotic, efforts, there is no universally accepted standard of archival description in europe or even within the society of archivists in the uk and ireland, nor is there any widely-used or agreed software for archives such as the dynix system for libraries. to computerise the archives is to plough a lonelier furrow. there have undoubtedly been some fairly sophisticated archival automation systems in europe. the public record office in london has since the nineteen seventies operated a computerised ordering system which is still far ahead of what is available in most other archives in ireland or the united kingdom. the current updating and extension of that system will put the pro very much ahead of the field again. in france, computerisation allows not only for online searching of finding aids but also for remote access and advance ordering, something which is made possible by widespread use of the minitel videotext system in that country, a degree of use unparalleled in any other european country (france’s minitel system accounted for 87.41% of all european videotext terminals in 1993)6. the historical archives of the european union in florence has, since 1993, provided online access to its database finding aids on the european commission’s echo co-host. spain is also well advanced towards linking its various state archives in one network which will allow remote access to all of them7 . online access to finding aids is still very much the exception rather than the rule for european archives, however, and most computerbased projects have tended to be exclusive to each institution. there has been little or no co-ordination among or between archives, no sharing of information other than what is already available over publicly accessible channels, no cross-fertilisation. the systems are mostly not compatible with each other and do not lend themselves to the sort of inter-institutional exchange of information that is now the norm for libraries8. there is a commitment at high level to do something europe-wide about automation and there is in existence a group of experts, comprising the heads of all national archives in the european union, charged with coordinating archival policy and practice including archival automation, but a large part of the problem is that the senior managers, the heads of archives, who are attempting to formulate common policies in this area, are in general themselves not terribly comfortable with technology and, therefore, not sure what it is they wish to do. despite a commitment to harmonisation and co-operation at the top, there has been little contact or co-operation among archives 23summer 1995 and archivists further down the hierarchy across national and linguistic boundaries. a major part of the problem in europe is also of course the difficulty of language. one indication of this is evident on the internet. the archives listservs in north america are not parallelled in europe (though a small “archives and the internet” discussion group has just this year been established within the it group of the society of archivists in the uk and ireland and may develop into a listserv. [author’s note:since this paper was presented, the “archives and the internet” discussion group has become a very vigorous forum for exchange of information among archivists in the uk and ireland.]). while there has been criticism of the american “archives” listserv from within the profession in the united states, it represents a very useful forum of over 2000 archivists exchanging information on matters of common concern. the us and canadian listservs are certainly of considerable benefit to those of us who access them from outside north america. it is significant, though perfectly understandable, that those subscribers to the listservs who are outside the united states and canada are mainly in the english speaking world, and predominantly in australia and new zealand. on the “archives” listserv there are, for instance, only two subscribers from germany (the country which accounts for 28% of the it market in europe) and none from france which, in terms of archival automation, is arguably the most advanced of the larger countries in europe9. it is obvious that the internet as a whole is still overwhelmingly a north american phenomenon. but this area is developing rapidly in the uk and ireland. in europe the number of computers directly accessible on the internet has doubled every year for the last three years, but in ireland within the last year, the number has tripled, and all the signs are that this is continuing to mushroom10. the tendency until now within ireland and the uk has been for internet access to come mainly from the academic community. it is not common for government employees to have access to the internet as part of their work, so there is no “.gov” element in our addresses. high telephone charges in europe compared to those in the united states and the disparate nature of the telephone systems, which have coincided fairly rigidly with national boundaries, have inhibited access to the internet by private individuals. also household computer ownership in europe is only about a third of that obtaining in the united states11. nevertheless, just by looking around one can see that things are changing. the fact that the next version of microsoft windows will come bundled with an internet access program (microsoft itself functioning as an internet access provider) will almost certainly result in a huge new wave of irish and uk connections from outside academia. for those archivists who connect, there will probably be a gravitational pull, at least initially, towards north america rather than into europe. despite commitments to further cooperation and harmonisation in europe, it is likely that the real dynamic will exist, for the moment, on the internet. given that there has been no listserv for archivists in the uk and ireland, presence on the american “archives” listserv is probably a reasonable guide to the number of archivists who are using the internet in these countries, and the number of archivists who are on the internet in these countries is probably in turn something of an indicator of the extent to which archivists have themselves embraced the new technologies [author’s note:since this paper was presented, the “archives and the internet” discussion group has become a very vigorous forum for exchange of information among archivists in the uk and ireland.]. relative to the size of their populations, australia and new zealand are leagues ahead of the uk, and ireland hardly figures. in this context it may also be significant that more than half of those appearing on the “archives” listserv with uk addresses are in university archives rather than state or official archives. the internet has huge potential for satisfying one of our primary needs, the need to disseminate information on our services and holdings to potential readers, and particularly to that diaspora of roots enthusiasts which we must cultivate. some traditional archives have already started to run gophers or to put up web pages. although our own computerisation is not very far advanced, we have considered it important to establish a presence on the internet and now have some pages on the world wide web by courtesy of a neighbouring third level college which has kindly afforded us space on their server12. there is clearly going to be growing demand for us to provide more and more information online. there will be growing pressure from europe to service a free information market to match that being developed in the united states. academics will surely soon start demanding that we use the available resources to improve access for them. it is rather surprising that they have been so reticent to date. in the light of the statistic of 17,000 people accessing sara’s gopher in january this year, we await with some trepidation the consequences of our own heads appearing above the parapet of the superhighway. as “traditional” archivists, we have much to learn from the pool of available knowledge on the internet in many areas, but particularly in relation to the problem of electronic records, which represents one of our greatest challenges, if not our greatest challenge, but has as yet has caused very few ripples to appear on the surface of the archival waters in europe. we in ireland have a national archives act as strong as most comparable archives acts and one which gives us statutory powers in respect of digital data. our act specifically defines “departmental records” to include magnetic tapes and discs, optical or video disks, and other machine-readable records. in fact there has been some debate over whether our act, in specifying types of media, such as tapes and disks, has rather missed the point and concentrated on the medium rather than the message (a major part of the problem being of course that you can happily preserve mountains of disks and tapes but this will 24 iassist quarterly not guarantee that the data remain accessible). however, we are confident that such definition is not exclusive and we regard the terms “files” and “other documentary or processed material” mentioned in our act to be mediatransparent. it is the message that we are charged with preserving. the main problem however, is not one of definition, it is the problem of what we do to give effect to our act. we have not yet managed to seriously address the challenge posed by electronic records, but we are not alone in this regard. although traditional archives in europe are aware of the challenge posed by digital data, progress to date in addressing it has been very slow. according to a recent study presented to the canberra conference on electronic records last november, no national archives in europe has yet got beyond the stage of holding the output of anything other than database systems, and many of us have not even got that far13. in the united kingdom and ireland the strongest player on the archival field and therefore the one that leads the way in many respects, the public record office in london, despite a number of high level studies of the issue going back over twenty five years, has yet to decide a policy on electronic records14. things now seem to be moving in britain, however, with the appointment in late april 1995 of an information manager in the public record office specifically charged with the task of developing a strategy for handling electronic records, and the appointment of a powerful committee of senior officials to ensure that he functions with the necessary support. it also seems certain that a formal decision will be made that archival electronic records in the form of structured datasets will be lodged with an existing agency rather than in the public record office itself and that preservation of digital data will continue to be outsourced15. in fact the existence of the esrc data archive in essex as the de facto place of deposit for official electronic archives in the united kingdom has probably allowed the public record office the luxury of time on this issue. most of the large datasets which might have been identified for preservation by the public record office have probably been preserved in essex. elsewhere in europe surprisingly little has yet been achieved. per nielsen has outlined exciting developments in denmark which may offer a blueprint for some other countries16. of the other national archives in the european union, it seems that only those in finland, france, germany and sweden have themselves accessioned electronic records and these mostly consist of datasets17. the national archives of the netherlands, however, has taken the initiative in attempting to bring the question of electronic records onto the archival agenda in europe18. traditional archives seem to have suffered a paralysis in confronting this issue which has presented them with problems of two types. firstly there are obvious problems associated with the preservation and future accessibility of such records instability of storage media necessitating regular migration of data, rapid hardware and software obsolescence. there is no need to recite these to an audience of data archivists. it is possible that we in ireland have already lost some of the large datasets created in our large information-gathering departments. we do not know, and our very preliminary efforts to find out, based as they are on our own ignorance of systems, have been inconclusive to say the least. the responses we have received have tended to be blandly reassuring, disturbingly so in the context of what we know to be the practice of some of these agencies in relation to their paper records. given that we do not yet know how we are going to address this problem, we have not yet probed too deeply. that said, we have found the level of response to our preliminary questionnaires to be disappointingly low, the lack of response indicating, perhaps, a belief among it managers that we are not there to help them. the second area of concern for traditional archives relates more to what has been termed the second generation of electronic records, the records of the electronic office, and to what has been called the distributed environment in which electronic records are being created. alongside the spread of computers has gone the breakdown of central file registries and filing systems. everyone creates their own documents and files them on the hard drives of their pcs or on personal directories or even on floppies. we find a multiplicity of systems, a multiplicity of software packages being used on them, a multiplicity of drafts and duplicates being stored in them. finding our way through this maze will be a colossal task. the traditional practice of traditional archivists, appraising records when the records have reached the end of their lifecycle is clearly not appropriate in the case of electronic records. if we wait until the records cease to be current or until they are released into the public domain in thirty years, or even twenty years, time there may be nothing left to appraise. there is a coincidence of developments here which is alarming. the last twenty five years or so, a period which has seen and is continuing to see the transition from paper to digital records, is also the period which has seen a generation of archivists professionalise. we are in the process of climbing into our professional fortresses and pulling up the drawbridges behind us, making it more difficult for those from other than a very narrow spectrum of training to enter the profession. but it is ironic that this generation of archivists, which has been so careful to professionalise, to define standards, may be the generation which will fail most spectacularly to leave behind a record of its own time. the options for traditional archives faced by the problem of what to do about electronic records are threefold. we can decide to use existing data archives and libraries as places of deposit and even perhaps develop an organisational link with these archives along danish lines; we can try to establish our own data archives as an integral part of the existing archives; 25summer 1995 or we can insist that archival electronic records be maintained by the creating agencies, with our organisations providing an inspectorate to ensure that such records are adequately catered for by the creating agency. it is unlikely that the deposit of official digital data with an existing data archives will be the strategy followed in ireland, despite the fact that this seems to be about to happen in the united kingdom. there are various reasons why this is unlikely to be our route but the strongest one is that there is no such entity as a data archives currently existing in ireland. as to our becoming a data archives, it has to be asked, and it has been asked, if it is at all appropriate for “traditional” archives to accession electronic records other than as a last resort? would we be placing ourselves on a treadwheel to maintain access to these records, something which may be done only by relegating other aspects of our responsibilities? would the archives be able to administer whatever privacy laws may regulate access to such data in the future? with so many systems current throughout the organisations for whose records we are ultimately responsible, would we have to become museums of software and hardware systems? the last question scarcely bears thinking about. as it is, we “traditional” archivists can barely master our own software and hardware. there is a compelling logic to the arguments advanced by david bearman and margaret hedstrom in favour of a noncustodial approach by traditional archives to such records19. the fact that the australian archives will now opt for this kind of approach, as set out in recently published guidelines, will weigh heavily in its favour with those of us who have yet to make a decision in this area20. this question will be addressed at european level in the spring of 1996 when a major multidisciplinary forum will be called in brussels to be attended by representatives of archives as well as it specialists from throughout the european union. this meeting will be held under the auspices of the european commission and will attempt to co-ordinate policy on machine readable records. it seems very likely that this forum will be influenced by decisions taken by the australians and by the very forceful arguments emanating from pittsburgh. yet there is a huge caveat which must be entered here, as edward higgs has recently warned elsewhere21; our previous experiences with some of the agencies which would have to become custodians of archival data do not inspire total confidence. yet it seems at the moment that, even with this caveat, local retention is the only practical option open to us in ireland though this of course may change. as mentioned above, traditional archivists in ireland are not computer people. we in the national archives do not at present have the resources to manage these records. it is unlikely that we will be given them in the short term, not on the sort of scale that would make the job feasible, and there is little merit in embarking an a project with a better than even chance of failure. it is simply very difficult to force the issue of electronic records onto the archival agenda or indeed onto any agenda. few people are interested. there is no pressure group or no constituency outside the archives demanding that something to be done about electronic records. historians in ireland have not seriously begun to use such records (some of them are now engaged in setting up databases of economic statistics or even online textual databases, but they have not yet begun to lobby on behalf of existing machine readable records). the late john blackwell who addressed the amsterdam conference of iassist in 1985 made some attempts to raise the issue in ireland, but seems to have met with little support22. if we were to close our reading room in order to stocktake, were we to withdraw a heavily used series of records and substitute microfilms, we could be fairly sure of a loud and unfavourable reaction from our research public. but if we choose to do something which will actually result in catastrophic consequences, if we ignore electronic records, no one will notice for a long time. no one outside the world of archives is currently lobbying about electronic records. this is something that we in the archives have to worry about for the moment on our own, sure in the knowledge that if we continue doing nothing will have left a shameful legacy. we must seek allies in attempting to give electronic record keeping a higher priority. there are some developments which indicate where we might find these allies. freedom of information legislation is imminent in ireland. there is a strong political commitment to this at present and the legislation currently promised looks set to be a far-reaching measure with radical effect. there will be major consequences both for the archives and for the holders of official information. for the archives, freedom of information, together with data protection, may eventually supplant the national archives act and the 30 year rule as the regulator of access. there are, anyway, moves in europe to have the norm for access reduced to twenty five or twenty years23. the gap, therefore, between current records and noncurrent records is likely to diminish. as for the informationcreating agencies, they will have to be more accountable for the information they create and hold, in whatever form it is held. something like the traditional registry system will have to be reinstated, but perhaps with routes of access from the outside world. and this system will of course have to encompass electronic records. perhaps a government information locator system may be used in the future as a route into unpublished official information or archival information, or at least into the finding aids for such information, and may ultimately support a gateway for online access to archival electronic records. whether traditional archives become non-custodial regulators of electronic records or custodians of such records, or, more likely, become a combination of both, we will clearly have to acquire the knowledge and skills which will allow us to make intelligent and correct decisions on the 26 iassist quarterly scheduling of such records. given our background and training and what has been to date an unimpressive track record with computers, it is unlikely that we traditional archivists will easily turn ourselves into electronic archivists. no-where within the profession in ireland at the moment are there the skills required to tackle this job. we are, however, greatly heartened by the news that one of the staff of the center for electronic records at nara, mark conrad, has been selected under the fulbright scheme to spend the next academic year teaching in the archives department of university college dublin. this is a hugely significant development in terms of archival formation in ireland and we may soon see the emergence of a generation of irish archivists with some skills in the management of electronic records. perhaps we in the traditional archives also need to make more radical plans now for a period of transition, and look outside our traditional recruiting pool to train new archivists for a new age. we should, to the extent that we can, encourage into the profession some from a technical rather than an arts background. and certainly “traditional” archivists must seek to forge stronger links with the data archivists and librarians, for it seems that we are now on the same road, having travelled to it from very different starting points. 1 paper presented at iassist 21st annual conference may 9-12, 1995, quebec city, canada. 2 richard cox, the first generation of electronic records archivists in the united states: a study in professionalization (primary sources and original works, volume 3, numbers 3/4), (new york, 1994). 3 the irish times, 10 february 1995, p12. 4 quoted in cox, op. cit., p.40 5 ronald f.e. weissman, “archives and the new information architecture of the late 1990s” in the american archivist, vol 57, winter 1994, pp 20 -34. 6 emerging technologies: information networks and the european union (european parliament, directorate general for research, working papers, economic series wll), (luxembourg, 1993) 7 archives in the european union (report of the group of experts on the coordination of archives), (luxembourg, 1994). 8 ibid. 9 list of subscribers to “archives” and “arcan-l” listservs supplied on 26 and 28 april 1995. 10 communications today, vol. 2, no. 2, (dublin, march 1995), pp 22-23. 11 emerging technologies, p.5 12 the address for the national archives of ireland home page at the dublin institute of technology is . 13 edward higgs, “information highways or quiet country lanes? accessing electronic archives in the united kingdom”, paper read to the “playing for keeps” conference on electronic records held in canberra, australia, 8 -10 november 1994. 14 ibid. 15 unpublished lecture by alexandra nicol delivered to the it group of the society of archivists, at the public record office, kew, london, 30 march 1995. 16 see per nielsen, “merging cultures: danish integration of academic data services into a traditional archival system” elsewhere in this volume. 17 higgs, op cit. 18 t. k. bikson and e. j. frinking, preserving the present: towards viable electronic records, (the hague, 1993). 19 for a discussion of these arguments see especially david bearman and margaret hedstrom, “reinventing archives for electronic records: alternative service delivery options” in electronic records management program strategies (archives and museum informatics technical report no. 18), (pittsburgh, 1993). 20 greg o’shea, managing electronic records: a shared responsibility (canberra, 1995). 21 higgs, “information highways” 22 john blackwell, information for policy (national economic and social council report, no. 78), (dublin, 1985) pp. 105-107. john blackwell, “public data in use: a case study of ireland”, iassist quarterly, fall/winter 1985, pp. 3-13. 23 archives in the european union (report of the group of experts on the coordination of archives), (luxembourg, 1994). 10 iassist quarterly undergraduate education and data: the entry level information specialist experience by eleanor langstafp baruch college, city university of new york "the medieval historian m. t. clanchy has illustrated the reluctant acceptance of written documentation in place of first-person witness over more than three centuries of early english history. 'documents,' he tells us, 'did not immediately inspire trust' people had to be persuaded that written documentation was a reliable reflection of concrete, observable events" (zuboff 77). although it is true today that at one level people will accept anything in writing as probably authoritative, and all statistics as incontrovertible, in most of our working environments a higher level of sophistication obtains. from time to time there may be an urgent desire to replicate surveys available in archival form, a survival from the 12th century where the spoken word, direct experience, had a 'presented at the international association for social science information service and technology (iassist) conference held in washington, d.c., may 26-29, 1988 validity not available in documentary evidence. the real problem, however, is more substantial, and deals with the nature of information, with those basic characteristics of information that make it what it is: transferability without diminution, growth each time it is used, its accumulation at meteoric rates, the increasing dependence of each new management decision on prior cases. in the federal government, for instance, according to wilson dizard, as early as 1979 there were 600+ database management systems installed (dizard 81); by 1990, it will cost nearly $10 billion to maintain the computers and $2 billion for modernization. of course, these levels of growth are reflected in the private sector. the impact of information technology on society has been variously described in terms of lessening labor or improving final quality. shoshana zuboft, in her in the age of the smart machine characterizes it as resulting in a comprehensive "textualization" of work, the creation of a new symbolic medium, an "electronic text" that increasingly mediates between workers and their work, between the body and the task. we are well aware of this phenomenon: the surrogate level, in the index, abstract and full data levels. in the information process, work becomes abstract, the manipulation of intangible symbols rather than concrete objects. zuboffs concern is with process, not with the accumulated data it engenders, but i suggest that where work is increasingly done in a technological mode, no matter what attitudes people may have about information, it continues to increase and require management a kind of management which cuts across present functional lines and which most futurists, such as shoshana zuboff and harian cleveland in his the knowledge executive , assert to be destroying the hierarchical structure of organizations. although i have not observed this phenomenon, except as it apparently engenders anxiety in certain groups when discussed, it is one of the justifications for the academic preparation of summer 1988 iassist quarterly 11 end-users of infonnation (atkins, passim; debons, passim; porat, passim). so loo the permanent growth of data is justification for providing new kinds of information intermediaries (duffy, passim; schmidt, passim; spivack, passim; topics, 77). what i would like to do this morning is to describe to you, for the purpose of discussion, the infonnation studies program now in the design stage at baruch college, in the hope that a such discussion will provide material with which we can fine-time our work, to make it more responsive to the workplace. the infonnation studies program is based on three premises based on perceived need: 1. that infonnation has value to society, to organizations and to individual professionals. 2. that personnel are needed who imderstand and are able to organize and utilize information effectively. 3. that to reach enough information users, programs in information studies are needed at the undergraduate level. what is the information studies program? this inter-disciplinary program focuses on the use and users of infonnation as well as the technologies involved. it provides students with the conceptual bases they need to work in an information environment and some professional training. when taken as a minor concentration it complements a major—business, one of the social sciences, or one of the natural sciences. baruch college, offering as it does fully accredited business, health care administration, education and liberal arts programs, is in a unique position to offer an information studies program which is responsive to the needs of a changing economy. one of the senior colleges of the city university of new york, it has an enrollment of about 14,000 and houses the business programs of the university. there is a growing need for information specialists in business, science and education; the marketing of information products is one of a growing number of examples that dramatically illustrate the viability of the baruch meld of information studies and business. in 1981 baruch college began to develop and ofter courses in the discipline of information studies and at present has five advanced courses approved by the board of trustees, and four courses approved by the school. the design of these courses is based on our view of the nature of information studies. in spite of the growth of programs for training information professionals, and for some, the concentration on that part of the spectrum of information science which can be styled information studies, the discipline is still in what can be described as a pre-paradigmatic stage in which the producers are characterized by training in other, usually related, disciplines. in common with the outer reaches of all disciplines, confiicting views are held and discussed among peers and with a general but educated, thinking audience (kusack, passim). the product, the texts produced, are treatises characterized by full discussion of the whole subject in so far as it is known (for instance summary results...). the assumptions of the discipline are the assertions of the treatises. information studies also has characteristics common to the next phase of developmenl the paradigmatic phase partakes of patterns common in academic departments in which research is carried out in response to specific questions posed by the discipline. thomas kuhn's paradigms—recognized scientific achievements that define acceptable problems and methods—provide a generally accepted conceptual context for further investigation. the audience contracts to one of peers only. communication moves from the treatise, a summer 1988 12 iassist quarterly literary fonn, to brief articles, written in highly technical tenninology for that specialized audience (paulson, chapter 2). information studies in the united states exhibits many of the characteristics of the pre-paiadigmatic stage, especially in the kind of publication it engenders—think pieces, forecasts, theoretical approaches, essentially all assertions. have you counted the number of assertions i've made so far? let's move to the more practical, more descriptive kind of approach to the subject: general objectives of the program are: 1. to contribute to the preparation of undergraduate students for critical and eftective participation in the complex structure of today's information society; 2. to prepare undergraduate information studies students for employment as information specialists in a variety of agencies and institutions; 3. to orient undergraduate information studies students towards graduate education in information studies or related fields; 4. 10 contribute to the academic preparation of undergraduate students who will enter careers or graduate education in related fields; 5. to provide course offerings that might be utilized by other segments of the university to augment their curricula. specific objectives of the program are to prepare students who will, within the limits of a minor subject specialization: 1. understand the nature and role of information in society in general, and in a variety of organizational settings; 2. understand how to find, evaltiate, and use recorded information; 3. tmderstand how to organize and retrieve recorded information; 4. acquire applied skills in information processing technologies; 5. be able to analyze the information needs of individuals, organizations and other social entities; 6. be able to design and manage information systems which meet specific information needs; 7. be able to instruct others in the use of retrieval systems; 8. be able to evaluate the eftecdveness of information systems; 9. understand behavioral aspects of information transfer including communications theory and communication skills. what valorizes information studies'" the valorization process, still in the earliest stages, stems from politics and from economics. information poses problems both political, in the sense of well-ordered organizations becoming chaotic in the face of too much information, and economic in developing feasible ways of dealing with information. thus, we have talked about the specific and practical ends of such training, with the assumption that such training will be part of the solution, not pari of the problem. the following are some of the courses we are teaching, or plan to teach in conjunction with such subject majors as international business, managment, biology, education or business communication. the latter concentrations we might call the content courses; what follows are summer 1988 iassist quarterly 13 the information studies courses: online information retrieval . juniors and seniors learn database searching using several databanks, and employ advanced strategies. they download, edit and format bibliographies and abstracts. a variety of software—gateways and frontends—is examined and used. advanced information retrieval . students learn to prepare material for input into databases. an indexing component presents automated indexing using standard software packages for file, periodical, and back-of-the-book indexing. an abstracting component explores the writing of indicative and informative abstracts, as well as other forms of terse writing. students prepare an index for the alumni magazine, a permanent responsibility of the class. information and society . a discussion and reading course covering policy matters and a general range of information technology and effects on society. the impact of telecommunications, electronic media, transborder data flows, etc., are studied. students also visit state-of-th^art workplaces. informational writing and editing in computer environments . efficient use of computer-generated information depends on its presentation. students learn to reprocess bibliographic, text, and numeric data in prim and graphic forms usable in business, government and non-profit organizations. science information retrieval . this course teaches basic principles of information retrieval in science and technology to students in pre-medical, health sciences and natural science programs using various interactive systems. in addition to bibliographic databases, students gain experience using databases to search for patents, to track technological developments, and to identify chemical substances. the purpose is to develop, in a scientific environment, those diagnostic, prescriptive and evaluative skills needed for information-searching at entry level in the modem science environmenl information technologies . an overview course designed to acquaint the student with several categories of information technology: computers, telecommunications and satellites, and video/print reproduction/graphics. representative technologies are examined in terms of fimctions, roles and design, and are related to the management of information resources. term projects explore in depth examples from each category. field trips to technologically advanced worksites acquaint the student with the latest applications. management of information resources and records . general principles of information organization: classification and filing, coding and indexing, routing and copying are examined, then applied to specific formats in print and computer environments. working with a computer model of an information center for both internal and external data, smdents make information management decisions and test them for relevance to organizational needs. management of external numeric data bases . managerial principles and practice by which external numeric data bases such as economic time-series, surveys and polls are eftectively handled in business, academic, and public organizations provide the substance of this course. acquisition, organization, service, and dissemination are considered. students gain familiarity with a variety of data sources available from government agencies, data archives, research institutions, private vendors and scholars. they use mainframe computers and microcomputers to work with actual data in a laboratory setting to gain expertise with secondary data files and the technical aspects of data storage and retrieval. this then is the program: what have we omitted that should be added? what have we summer 1988 14 lasstst quarterly included that could be deleted? it is an ambitious program. does it respond to the needs of an information societytn sources undergraduate preparation." journal of education for library and information science 26 (summer 1985): 50-51. topics in health care financing 14 (winter 1987):77-88. zuboft, shoshana. in the ase of the smart machine: the future of work and power . new york: basic books, 1988. atkins, t. v. survey of information industry needs 1987. new york: baruch college. part 1-2. part 2 incomplete to date. clanchy, m. t. from memory to written record: england 1066-1307 cambridge: harvard up, 1979. quoted in zuboff cleveland, harlan. the knowledge executive: leadership in an information society . new york: talley books, 1985. debons, anthony, et al. the information professional: survey of an emerging field. new york: marcel dekker, 1981. dizard, wilson p. the coming information age: an overview of technology, economics, and politics . new york: longman, 1985. duffy, james and william j. jeffrey. "is it time for the chief information officer?" management review 6 (november 1987): 59. kusack, james m. "librarians and the information age: an affair on the rocks?" (part 1). asis bulletin 14 (dec. 1987/jan. 1988): 26-27. paulson, william. the noise of culture: the literary text in a world of information . ithaca, ny: cornell up. 1988. porat, marc uri. the information economy: definition and measurement washington, d.c.: u.s. department of commerce/office of telecommunications, 1977, ch. 7. schmidt karen a. "the other librarians: undergraduate library science programs and their graduates." journal of education for librarianship 24 (spring 1984): 223-32. spivack, jane f., ed. careers in information . white plains: knowledge industrypublications, 1985. "summary results of propositions related to summer 1988 vol23/2 26 iassist quarterly introduction this paper firstly will review the literature with a view to describing the current state of the art of cai, secondly it will describe a scottish housing project which used capi and consider the quality of data output, and thirdly, it will draw conclusions about the use of ca(p)i on the basis of the authors’ own findings but also placing the discussion in its wider research methodology and research environment context. the questions which this paper seeks to explore are broadly: * is capi suitable only for large projects ? * is capi good with sensitive data? * does use of capi improve data quality ? * can capi substitute for qualitative research in a contract research environment ? these types of questions have not been asked before as, arguably, studies on data quality have been too restricted. furthermore, it is suggested that the potential of computer-assisted data collection methods has not been fully utilised. this paper explores whether quantitative research could be regarded as a universal solution. part 1: literature: previous users of capi compared to even a few years ago, computer aided interviewing (cai) is now relatively widespread and mature. a move to cai can, for example, lead to improvements in data quality and turnaround times; it can even make possible surveys that would not otherwise be contemplated. for these and other reasons, many survey organisations and clients have been persuaded that cai is where the future of survey research lies. indeed for almost every traditional approach to survey data collection there is now a computer assisted alternative. the two most widespread are computer assisted telephone interviewing (cati) and computer assisted personal interviewing (capi) (collins & sykes, 1998). just because capi/cai is so widespread these days, there is no longer much research computer assisted personal interviewing: a method of capturing sensitive information by emma forster and alison mccleery* abstract this paper will discuss how computer assisted personal interviewing (capi) coped with collecting sensitive data in a difficult interview situation. a recent research project, funded by scottish homes (a uk government housing agency for scotland) used the capi technique to collect information on home ownership at the margins of affordability. the project used an innovative joint approach between the academic sector and a leading uk survey consultancy. it could be argued that a more sensitive method of collecting this sort of information would have been indepth interviews, which could then have been analysed using qualitative research methods. the paper will discuss the outcomes of using capi and quantitative research methods in such a sensitive project. it is suggested that the use of capi has achieved a better response rate on sensitive questions than other techniques would have. the use of capi has a number of well-known advantages, such as improvements in data quality and turnaround times. this paper will assess whether capi can deliver in a number of interview conditions or if its potential benefits will be realised only under certain conditions. it will critically review how the quantitative method worked in this specific situation before placing the discussion in its wider research methodology and research environment context. summer 1999 27 which compares them as modes to the more traditional modes. instead the focus of the literature has shifted to improving cai, as for instance bulmer et al. (1998), away from questioning the intrinsic comparative value of the mode itself. however, due to the shift which has occurred in the way survey data are collected with telephone surveys and, to a lesser degree, mail surveys now being more extensively used, this has stimulated a limited amount of empirical research on the influence of the data collection method on data quality. de leeuw et al. (1996) compare a mail, a telephone, and a face-toface survey and found that the different data collection methods did have an effect. there is less research however on the question of whether to adopt cai but more on what effect it has and how this might be mediated. disadvantages of capi and cai in general there is conflicting evidence on the impact of capi on data quality. collins & sykes (1998) find little evidence of clear improvement as yet. furthermore, despite general acceptance of cai in the literature, it is recognised that paper and pen interviewing (papi) still has its role. the promise of cai will be realised only under certain conditions and capi is not a panacea. face to face interviews and capi as a method of data collection also have weaknesses associated with their usage. for instance blyth (1998) lists the hours needed to be worked by the interviewers; falling response rates obtained by this method; and personal safety considerations amongst others, all factors that would militate against the increased use of this method. to this list blyth (1998) adds equipment cost and software/hardware/data interaction as well as the need for batteries, as major issues to be considered in capi. although, as will be seen later, there are ways round the capital investment and overheads barriers to the use of cai, nevertheless these types of hardware and software issues associated with capi must be fully taken on board before the decision is taken to use capi as a data collection method. finally, the quality and accessibility of the output data from cai has been questioned. this is because cai does not eliminate human error, which where capi is concerned, simply manifests itself in a different way as compared with papi. routing mistakes may mean that whole questions are missed in every interview, whereas in papi the interviewer may inadvertently turn over two pages together and may miss a whole page but not in every case or at least not the same error in every case i.e. papi is likely to be associated with a series of random inconsistent errors, while capi is likely to be associated with a single consistent serious error affecting every interview schedule. however, it is suggested that the latter is highly responsive to improvement as a result of piloting the questionnaire and training of interviewers. currently one of the disadvantages of capi is that the non-technical reader finds the content, structure and workings of a casic questionnaire much more difficult to understand (bulmer et al., 1998). the growing possibilities of computer hardware and software associated with technological advance have made it possible to develop very large, and complex electronic questionnaires. as a consequence, it has become more and more difficult for developers, interviewers, supervisors, and managers to keep control of the content and structure of cai instruments. various attempts have been made to render these cai questionnaires intelligible to non-specialists. recently a more comprehensive attack on the documentation problem was launched. this project has been named tadeq (a tool for analysing and documenting electronic questionnaires), manners & bethlehem, 19991. the tadeq project proposes to develop a flexible tool for documenting and analysing electronic questionnaires. this tool aims to be neutral so it will be possible to use it in combination with existing computer assisted interviewing (cai) systems. as a documentation tool, it must be able to produce a human-readable presentation of the electronic questionnaire. tadeq will produce two types of output: a paper version of the documentation, which can be used e.g. by either interviewers or managers; and an electronic version of the documentation, in some kind of hypertext format, allowing developers or researchers to scrutinise the contents and structure of the questionnaire. this development poses a conundrum for data archivists: should they adjust their systems to include the tool to convert cai questionnaires into this standard format once they have received them, or should they get data depositors to do this before they submit their questionnaire? the result of such a tool is that it produces standardisation of survey documentation and so is advantageous for data archivists in the long run. advantages of capi dent (1999) does recognise that the well-known advantages of technology are that they provide an improved research quality, a quicker turnaround and easier integration with other activities, for instance management reporting and marketing approaches. he further states that the main reasons for the use of cai include managing complexity of data, of samples and of reporting and maintaining records and files. these have been comprehensively illustrated elsewhere. bulmer et al. (1998) in their finding that the introduction of computer assisted survey information collection (casic) has made technically 28 iassist quarterly feasible a much higher level of questionnaire complexity than was possible in paper and pencil days. this confirms it is literally the case that much of today’s more ‘serious’ work could not be otherwise achieved without cai. de leeuw et al. (1995) reviews the evidence of the effect of computer assisted interviewing on data quality and find that there are clear advantages of cai in two main areas: survey data quality; and acceptance of the computer by respondents and interviewers. their main conclusions are that computer-assisted data collection methods are accepted by both respondents and interviewers, and that survey data quality improves, especially when complex questionnaires are used. by 1999, in a discussion of the widespread market penetration of cati and to a lesser extent capi, de leeuw has added lower costs to the previously identified advantage of improved data quality. although a study by hox & de leeuw (1994) shows that response to mail surveys has been improving recently, nevertheless face-to-face surveys continue to achieve the highest response rates. given dillman et al.’s (1993) finding that asking potentially difficult and/or objectionable questions lowers the response rate, a capi approach which combines good results on sensitive questions with the traditional high response rate of face-to-face interviews should now be the method of choice. furthermore, the use of technology gives a professional image, and capi is generally liked by respondents, while it intrigues many elderly respondents. research should therefore concentrate on further reducing human error associated with capi. this present paper offers a step along the way to producing more research on face-to-face methods, producing additional research into non-response in face-to-face surveys. already a major advantage of capi is the way in which it is able to reduce potential interviewer and respondent error. routing errors are eliminated because the script automatically routes to the correct questions. this ensures that data are generally more complete, can considerably reduce the number of ‘non-responses’ and, correspondingly, the need for corrective editing (or even re-contacting respondents) later on. capi also has the ability to ‘range-check’ data and carry out logic and consistency checks during the interview. either ‘hard’ or ‘soft’ checks are set, to query or confirm key pieces of data with the respondent. in this project for scottish homes capi was used to check financial data during the interview and clarify whether figures relate to pounds or pence. ranges were set for key pieces of financial data and discrepancies queried with the respondent if they fell outside certain ranges. although we did not use this facility in this study, capi can also be used to calculate and provide derived variables during the course of the interview (which can then be fed back to the respondent or queried with them, as appropriate). returning to the matter of sensitive questions, there is limited evidence that capi is better than papi for sensitive questions. few studies have looked at this. those that have compared the same questions and different collection modes have tended to compare between cai methods and rather than between papi and capi. when looking at various modes of survey data collection which had to ask difficult or sensitive questions, it was found by tourangeau & smith (1996) that the mode of data collection did indeed affect the level of reporting of sensitive behaviours. this study compared three methods of collecting survey data about sexual behaviours and other sensitive topics: computerassisted personal interviewing (capi), computer-assisted self-administered interviewing (casi), and audio computerassisted self-administered interviewing (acasi). it was found that the three mode groups did not differ in response rates, but both forms of self-administration tended to reduce the disparity between men and women in the number of sex partners reported. self-administration, especially via acasi, also increased the proportion of respondents admitting to the use of illicit drugs. this study also highlighted the importance of the closed answer options in determining the response. thus it is suggested that open answers, although time-consuming to code, may produce more unbiased answers in answering difficult or sensitive questions. other evidence (de leeuw, 1999) points to that perceived confidentially playing a role in obtaining higher response rates. earlier research by de leeuw et al. (1995) concluded that this is an under-researched area. the authors’ own survey, while it does not give definitive evidence as it does not use a comparison of methods, nevertheless does lend weight to the argument that use of capi gets a good response rate on sensitive questions. de leeuw (1999) established from her review that there is a greater willingness to report extreme views using capi. mori scotland have found, when comparing data from paper based and capi surveys, that the problem of non-response to specific questions is significantly reduced, and that some sensitive information is collected more fully (such as household income) with capi. for example, in transferring the national mori omnibus survey from paper to capi, it was found that the proportion agreeing in principle to being re-contacted rose from 77.4% to 80.8%, and the proportion refusing to declare a household income declined from 16.8% to 14.1%. however, there is no evidence that, overall, using capi has a significant summer 1999 29 effect one way or another on respondent willingness to participate in surveys. what is, however, certain is that keying errors associated with data entry are very much reduced since data entry is done once, during the interview, rather than by coding onto paper and subsequently transferring responses to computer. even with 100% verification, errors of 0.05% on some variables may be expected where one is interpreting hand-written numbers; experience to date suggests that fewer errors are associated with capi. summary of advantages and disadvantages of capi thus it appears from the literature that capi is good for capturing sensitive data, has a fast turnaround time (good for the short deadlines that are associated with contract research) and generally improves data quality (although there is conflicting evidence on this). one of the main disadvantages is cost: the hardware and software is expensive and then there is the interviewer training over and above. this is likely to limit the use of capi to commercial survey companies with a large throughput which can afford overheads by spreading the capital investment. furthermore, it is worth stressing at this point that capi and more generally cai do not in themselves address many of the common problems of collecting survey data for the quantitative studies and furthermore cannot pretend to emulate the painstaking detail of qualitative interviewing. nevertheless, it is also, paradoxically, time to say that the conclusion reached from the available literature is that many projects/surveys could not be done with such accuracy, and some not done at all, without cai. in particular those surveys which require complex routing but need to be carried out in a fairly short interview time. part 2: project description and results brief description of project the purpose of this research was to: 1. develop an understanding of the issues concerning poorer owner occupiers, and identify what greater role information could play in attaining more successful housing outcomes; 2. undertake a detailed case study of owners in south rogerfield on glasgow’s eastern periphery. specifically the aims of the case study were to: * analyse the factors, both immediate and contextual, which influence the decision to buy; * produce an evaluation of the retrospective understanding of the owner’s responsibility in respect of common repairs; * determine what information had been used in the process of buying, and of that identify information deemed to have been useful and information considered misleading or unhelpful at each stage in the decision; * assess the role that information or the lack of it played during any difficulties regarding repair bills, financial difficulties, redundancy; and * identify examples of best practice where provision of information helped avoid common pitfalls and where lack of information obscured common pitfalls which happen to those buying on limited budgets. 3. catalogue the available information for homebuyers before and after they buy their home. the actual results of the project are reported elsewhere in forster and mccleery (1999a, 1999b). south rogerfield was identified as the survey area by scottish homes on account of the high number of repossessions: between 1987 and 1993 49 houses in the estate were repossessed by lenders: this is equivalent to 20% of the total number of properties. together south and north rogerfield form one of the fourteen neighbourhoods which make up the sprawling peripheral housing estate of easterhouse on glasgow’s north-eastern perimeter, as shown on location maps 1 and 2. comprising circa 300 housing units, south rogerfield consists of 3-storey tenement blocks built in the late 1950s and arranged in either rows or quadrants, predominantly with continuous frontages. on the instructions of glasgow district council, the properties were sold off and improved in about 1985 by two firms of property developers, crudens and barratt, although twenty or so properties pepper-potted throughout are still council-owned, while six are in housing association shared ownership. the modernisation carried out by the developers included storey height reduction, new roofs, double glazing, new bathrooms and central heating. 30 iassist quarterly map 1: location of survey area within scotland map 2: location of easterhouse estate, glasgow, scotland source: after http://www.scotland.gov.uk/library/documents3/fs12-34.htm local government in scotland, fact sheet 12, scottish office http://www.scotland.gov.uk/library/documents3/fs12-34.htm summer 1999 31 establishment of the sample was not straightforward, but the eventual response rate was very satisfactory for an area of this type. after initial confusion as to how many flats there were in south rogerfield, 138 full interviews with owner-occupiers were obtained. however, 48 more households, all occupying rented housing, were asked a much more limited list of questions. if the 41 sheltered and vacant properties are subtracted from the total count, the number of ‘valid and occupied’ flats comes to 308. from these 186 interviews were obtained, 138 with owners and 48 with renters, giving a total response rate of 60%. a fuller breakdown of response types is provided in table 1. table 1: response rate total addresses identified 349 sheltered or vacant 41 ‘valid and occupied’ (used as base) 308 interview with owner 138 45% interview with tenant 48 16% overall response rate 186 60% refusal (including not interested 45 15% and too busy) too ill to take part 2 1% insufficient english 1 1% no contact after 4 calls 74 24% the data from the owner-occupiers was collected by mori scotland using computer aided personal interviewing (capi). choice of method this section begins by defending the choice of capi as the survey method. thereafter the experience of using capi is described, prior to a discussion of the extent to which the reality of using the method matched expectations. first the choice was made between qualitative and quantitative. sensitive questions are often dealt with in in-depth interviews qualitative methodologies. the following section explains why qualitative was considered unsuitable for us in this project. a method was needed that produced a quick turnaround. the whole project ran only for 3 months and the data collection phase was allocated only 3.5 weeks, with a draft of the survey results needed shortly after the data collection was finished. qualitative interviews would have surely discovered more in-depth information on the topic in question but could not have been collected nor analysed in the time allowed. quantitative method was chosen primarily because of the time-scale. due to the constraints of research work in a research contract environment, the future of qualitative research in that setting is questionable. increasingly today, the nature of contract research demands quantitative data which can be analysed statistically and held in data archives for future comparability. moving on to discuss the specific choice of capi. capi has a good response rate on sensitive questions (tourangeau & smith, 1996; de leeuw, 1999). yet sensitive data is in the past thought to be better collected in a self-completion method. yet, casi has a lower response rate than face-to-face, and has a slower turnaround time. as is seen above in the literature review, the fast turnaround time is one of capi’s main advantages. our expectations in choosing capi for this survey were that: * it would catch the client’s imagination to win the contract in the first place * capi would help achieve first time round a 60-70% response rate needed in connection with the blanket coverage in a small area, (it would not be possible to re-sample to improve the response rate therefore it needed to be possible to get a fairly high response rate straight off.) 32 iassist quarterly * fast turnaround time, in particular it would help speed up the stage between the end of the survey and obtaining the data * would cut down on data input errors * allow complex routing through debt, income, mortgage/endowment payments and repair sections yet it was envisaged that each interview should not last more than 30 minutes. in summarising the aims of chosen survey method, the nature of the survey demanded a high response rate. the area under study and therefore size were very small and 100% blanket coverage sought. furthermore, an important objective of the study was to collect as much data as possible on sensitive questions relating to income, mortgage and other house-related payments, debt etc. choice of joint approach a collaboration was chosen between a survey company and academic institution, primarily because it was felt that this difficult interview situation asking sensitive questions would need very highly trained and very experienced interviewers. the tight deadlines necessitated this approach, as there was no time for training of interviewers. not only professionalism and timesaving but cost was an important factor too. this strategic approach avoided the initial expensive investment required in hardware and software i.e. the up-front costs mentioned earlier which are problematical for resource-lean uk academic institutions. outcomes of choice of capi in fact, our expectations of use of capi as a data collection method were exceeded as we got better response rates, and lower refusals rates than we expected2. capi proved successful in terms of data quality and a high response rate on sensitive question. our survey differs from the common experience that asking potentially difficult and/or objectionable questions lowers the response rate as found by dillman et al. (1993). the refusals to answer sensitive questions did not rise above 6.5%. the question that had the biggest refusal rate was the income question with 9 households refusing to give this information. income was asked only in bands so as to improve the response rate and this seems to have worked as there was a high level of response to this question. an examination of the pattern and number of refusals reveals that it was the same few households that account for most of the refusals. a full list of the questions where refusals occurred is in appendix 1 and the breakdown of the number of refusals is found in appendix 3. a further examination was made of the characteristics of these people and it was found that they did not fall into one demographic or economic group but were scattered between these. full results of the testing of characteristics of the people who refused to answer some of the questions are to be found in appendix 2. overall, it was shown that there were very limited refusals in the income and debt questions. so, in summary, the total refusal rate remained low throughout the survey. possible reasons for this include professionalism of the highly trained interviewers, the perceived higher confidentiality that capi gives the respondents. however, there was not a question on the interviewees attitude to capi in the questionnaire and so it is only possible to speculate on this. while the fairly high response rate overall (table 1) was pleasing for the researchers, it was not on the whole unexpected and was in conjunction with the fast turnaround time the reason that this method was chosen. however, what was thought to be unusual and not mentioned elsewhere in the literature was that the survey was a fairly small-scale one. normally capi is presented only as cost-effective in large-scale data surveys. also unexpected was the high response rate to the sensitive questions. lessons to learn/unavoidable problems as the literature made clear, capi is not foolproof and is only as good as the interviewers who administer it and the programmers who set up the routing. in this particular survey two questions were completely missed from every interview due to a fault in the routing. because of the nature of contract research, with limited budgets and tight time-scales there was no pilot of this survey and only limited testing of the routing before the interview schedule went into the field. capi can arguably improve on papi by improving turnaround without any loss of data quality or even an improvement in data quality, in particular with sensitive questions. but for reasons stated earlier use of capi has to date been favoured for large studies only. blyth predicted in 1998 that in the future capi would be most relevant for big international players and for very large surveys. he felt that non-capi will gravitate more quickly to telephone and that this will result in only a small number of field-only capi companies. however, our survey offers a less limited future for capi in that it proved successful for a smallish study in which an academic institution sub-contracted the data collection to a commercial concern. summer 1999 33 conclusions in summary: * capi not panacea (although better than papi) * not a replacement for qualitative methods * appropriate for certain data in certain contexts e.g. can achieve good results with sensitive data * further research into how even better advancements in capi is now appropriate * involves high-level of investment up-front costsand so is not for everyone although strategic studies to this problem are possible (as we have shown). * previously better for large-scale studies, but especially when handling sensitive data, possible for small-scale studies too (as we have shown). capi was the most suitable method for use in this contract research project primarily because it gets data quickly. however, additional advantages also emerged. the surprising finding was capi was useful and cost-effective in small-scale data collection situations too. the combination between academics and market research survey company worked well. the resultant high response rate and low refusal rates commends this method for use in other similar interview situations. but of course, without conducting the same study again using papi, it is impossible to conclusively say whether it was capi alone that led to high response rate and low refusal rate. blyth (1998) points out that with increasingly fragmented populations and busier lifestyles we need to use a mix of technologies to obtain maximum response from our survey population. as blyth (1998) rightly points out the focus should be on the answers and not the media, however, it is important to consider media in the context of which media gives the most (highest quality) answers. it is also necessary to adjust the media according to the research environment context, although in doing so we must not lose sight of issues of research philosophy and quality. acknowledgements the authors would like to thank scottish homes, the funders of this research work ‘housing information and advice and home ownership at the margins’, and also mori, scotland who carried out the survey work and to simon braunholtz in particular. glossary computer assisted personal interviewing capi hand-held assisted personal interviewing hapi computer assisted telephone interviewing cati computer assisted self interviewing casi computer assisted interviewing cai computer assisted data input cadi paper and pen interviewing papi computer-assisted data collection methods cadac audio computer-assisted self-administered interviewing acasi computer assisted survey information collection casic appendix 1: the questions where refusals occurred: qe1. what was the purchase price of this property? qe3. and roughly how much do you think you could sell this property for now, if you put it on the market? qe4. from which of the sources on this card did your household get the money to buy this property? qe8. apart from money to move in, how much have you borrowed since you moved into the flat for costs to do with the house for instance repairs, improvements, cookers/fridges etc.? qe13. at the moment, how much does your household pay each month in mortgage or loan payments? qe14. how much, if anything, does your household pay in the additional separate endowment part of the mortgage each month? qf3. how easy or difficult is it for your household to pay the mortgage payments? qf4. have you been more than two months behind with your mortgage payments at any time in the past two years ? please look at this card and tell me the letter next to the band in which you would place your total household income for the year from employment or benefits. qsv. at the moment do you (or your partner) have any money saved or invested? qsv2. showcards: how much do you (and your partner) have saved together? please tell me the letter on this card for the group in which you would place your total savings? 34 iassist quarterly appendix 2: who refused to answer ? in total 33 refusals to 11 questions by 14 households. case number 113 refused 7 questions case number 41 refused 5 questions case number 72 refused 4 questions case number 94 refused 3 questions case number 26 refused 2 questions case number 55 refused 2 questions case number 99 refused 2 questions case number 130 refused 2 questions case number 68 refused 1 question case number 69 refused 1 question case number 70 refused 1 question case number 100 refused 1 question case number 109 refused 1 question case number 123 refused 1 question .snoitseuqynarewsnaotdesuferohwsdlohesuohtnednopserfosnoitapuccodlohesuoh trap emit dna -lluf emit rekrow -esuoh efiw dna -lluf emit rekrow owt deriter elpoep ruof lluf emit srekrow eerht lluf emit srekrow eno lluf emit rekrow eno -lluf emit rekrow enodna kcistl eno -trap emit rekrow ylno -lluf2 1,emit -trap emit 1dna tneduts latot latot 3 3 1 1 1 2 1 1 1 41 ohwsdlohesuohtnednopserfosezisdlohesuoh noitseuqynarewsnaotdesufer ezishh 1 2 3 4 5 latot latot 2 3 4 3 2 41 sdlohesuohtnednopserninerdlihcforebmun noitseuqynarewsnaotdesuferohw forebmun nerdlihc 0 1 2 3 latot latot 8 3 2 1 41 noitseuqynarewsnaotdesuferohwsdlohesuohtnednopserfosepytdlohesuoh dlohesuoh epyt nosrepelgnis dlohesuoh owt-ylimaf dlihcdnastluda owt-ylimaf eromdnastluda dlihcenonaht -elpuocredlo tnedneped-non nerdlihconro latot latot 2 4 3 5 41 summer 1999 35 appendix 3: the number of refusals ?ytreporpsihtfoecirpesahcrupehtsawtahw1eq ycneuqerf tnecrep dilav tnecrep evitalumuc tnecrep desufer 2 4.1 7.66 7.66 wonkt’nod 1 7. 3.33 0.001 latot 3 2.2 0.001 gnissim 531 8.79 latot 831 0.001 ?wonrofytreporpsihtllesdluocuoyknihtuoyodhcumwoh3eq ycneuqerf tnecrep dilav tnecrep evitalumuc tnecrep desufer 1 7. 1.9 1.9 wonkt’nod 01 2.7 9.09 0.001 latot 11 0.8 0.001 gnissim 721 0.29 latot 831 0.001 ?ytreporpsihtyubotyenomehttegdlohesuohruoydiddracsihtnosecruosehtfohcihwmorf4eq ycneuqerf tnecrep dilav tnecrep tnecrepevitalumuc ton 731 3.99 3.99 3.99 desufer 1 7. 7. 0.001 latot 831 0.001 0.001 otstsocroftalfehtotnidevomuoyecnisdeworrobuoyevahhcumwoh,nievomotyenommorftrapa8eq esuohehthtiwod ycneuqerf tnecrep tnecrepdilav tnecrepevitalumuc desufer 2 4.1 7.1 7.1 wonkt’nod 4 9.2 4.3 0.5 enon 311 9.18 0.59 0.001 latot 911 2.68 0.001 gnissim 91 8.31 latot 831 0.001 stnemyapnaolroegagtromnihtnomhcaeyapdlohesuohruoyseodhcumwoh31eq ycneuqerf tnecrep tnecrepdilav tnecrepevitalumuc desufer 2 4.1 3.33 3.33 wonkt’nod 4 9.2 7.66 0.001 latot 6 3.4 0.001 gnissim 231 7.59 latot 831 0.001 36 iassist quarterly tnemwodnelanoitiddaniyapdlohesuohruoyseod,gnihtynafi,hcumwoh41eq ycneuqerf tnecrep tnecrepdilav tnecrepevitalumuc desufer 2 4.1 4.5 4.5 wonkt’nod 4 9.2 8.01 2.61 nidedulcnisiti,gnihton detouqerugif 13 5.22 8.38 0.001 latot 73 8.62 0.001 gnissim 101 2.37 latot 831 0.001 stnemyapegagtromyapotdlohesuohruoyroftisitluciffidroysaewoh3fq ycneuqerf tnecrep tnecrepdilav tnecrepevitalumuc ysae 901 0.97 0.97 0.97 seitluciffidevahsemitemos 32 7.61 7.61 7.59 seitluciffidevahnetfo 3 2.2 2.2 8.79 seitluciffidevahsyawla 1 7. 7. 6.89 wonkt’nod 1 7. 7. 3.99 desufer 1 7. 7. 0.001 latot 831 0.001 0.001 sraey2tsapehtniemitynatastnemyapegagtromruoyhtiwdnihebshtnomowtnahteromneebuoyevah4fq ycneuqerf tnecrep tnecrepdilav tnecrepevitalumuc sey 01 2.7 2.7 2.7 on 621 3.19 3.19 6.89 desufer 1 7. 7. 3.99 wonkt’nod 1 7. 7. 0.001 latot 831 0.001 0.001 puorgemocni ycneuqerf tnecrep tnecrepdilav tnecrepevitalumuc 000,5£rednua 3 2.2 2.2 2.2 000,01£-000,5£b 11 0.8 0.8 1.01 000,51£-000,01£revoc 63 1.62 1.62 2.63 000,02£-000,51£revod 53 4.52 4.52 6.16 000,52£-000,02£revoe 91 8.31 8.31 4.57 000,03£-000,52£revof 9 5.6 5.6 9.18 000,03£revog 9 5.6 5.6 4.88 wonkt’nod 7 1.5 1.5 5.39 desufer 9 5.6 5.6 0.001 latot 831 0.001 0.001 summer 1999 37 ?sgnivas)rentrapruoyro(uoyevahvsq ycneuqerf tnecrep tnecrepdilav tnecrepevitalumuc sey 24 4.03 4.03 4.03 on 88 8.36 8.36 2.49 erusnu/desufer 8 8.5 8.5 0.001 latot 831 0.001 0.001 ?devasevah)rentrapruoydna(uoyodhcumwoh2vsq ycneuqerf tnecrep tnecrepdilav tnecrepevitalumuc 000,1£rednua 8 8.5 0.91 0.91 999,2£-000,1£b 01 2.7 8.32 9.24 999,4£-000,3£c 8 8.5 0.91 9.16 999,9£-000,5£d 1 7. 4.2 3.46 000,61£-000,01£e 3 2.2 1.7 4.17 000,61£revof 6 3.4 3.41 7.58 wonkt’nod 2 4.1 8.4 5.09 desufer 4 9.2 5.9 0.001 latot 24 4.03 0.001 gnissim 69 6.96 latot 831 0.001 references blyth, bill (1998) the current and future use of technology in european survey research, presentation at association for survey computing (asc) compstat satellite meeting, new methods for survey research, 21 22 august 1998, chilworth manor, southampton. bulmer, m. thomas, r. & donagher, p. (1998) survey documentation: representation of casic questionnaires, in new methods for survey research, edited by a. westlake et al., association for survey computing, pp 37-48. (http:// www.assurcom.demon.co.uk/events/c98/) collins, m. & sykes, w. (1998) the impact of computer assisted interviewing on uk survey research, in new methods for survey research, edited by a. westlake et al., association for survey computing, pp 3-12. (http:// www.assurcom.demon.co.uk/events/c98/) de leeuw, e. (1999) the effect of computer assisted interviewing on data quality: a review of the evidence, social statistics meeting, royal statistical society, march 16th, 1999. de leeuw, e.d.; hox, j.j.; snijkers, g. (1995) the effect of computer-assisted interviewing on data quality a review, journal of the market research society, 37(4): 325-344. deleeuw, e.d., mellenbergh, g.j. & hox, j.j. (1996) the influence of data collection method on structural models a comparison of a mail, a telephone, and a face-to-face survey, sociological methods & research 24(4): 443-472. dent, t. (1999) the impact of cai on large and complex surveys: taking advantage of the new technology, presentation at the impact of cai on large and complex surveys, 8 january 1999, at imperial college, london. http:// www.assurcom.demon.co.uk dillman, d.a., sinclair, m.d., clark, j.r. (1993) effects of questionnaire length, respondent-friendly design, and a difficult question on response rates for occupant-addressed census mail surveys, public opinion quarterly, 57 (3): 289304. http://www.assurcom.demon.co.uk/events/c98/ http://www.assurcom.demon.co.uk/events/c98/ http://www.assurcom.demon.co.uk/events/c98/ http://www.assurcom.demon.co.uk/events/c98/ http://www.assurcom.demon.co.uk http://www.assurcom.demon.co.uk 38 iassist quarterly forster, e. & mccleery, a. (1999) housing information and advice and home ownership at the margins, report to scottish homes. http://www.scot-homes.gov.uk/ hox, j.j.; & de leeuw, e.d. (1994) a comparison of nonresponse in mail, telephone and face-to-face surveys applying multilevel modeling to metaanalysis, quality & quantity, 28(4): 329-344. manners, t. & bethlehem, j. (1999) tadeq: a tool for analysing and documenting electronic questionnaires. presentation at the impact of cai on large and complex surveys, 8 january 1999, at imperial college, london. http:// www.assurcom.demon.co.uk scottish homes (1999) housing information and advice and home ownership at the margins, scottish homes: precis no. 85. http://www.scot-homes.gov.uk/indexnext5.html tourangeau, r. & smith, t.w. (1996) asking sensitive questions the impact of data collection mode, question format, and question context, public opinion quarterly, 60(2): 275-304. 1 the tadeq project is funded under the european commission’s esprit programme. it is led by statistics netherlands and the other partners are the office for national statistics, uk; statistics finland; instituto nacional de estatística, portugal; max planck institute, saarbrucken, germany. more information on the tadeq project can be found at:http:// www.blaiseusers.org/tadeq/abouttdq.htm 2 the survey was both of renters and owner-occupiers but only the owner-occupiers were asked a full interview using capi. the responses of the owner-occupiers only will be dealt with in this paper. emma forster and alison mccleery, dept of psychology and sociology, napier university, redwood house, 66 spylaw road, edinburgh, eh10 5br,scotland, u.k. telephone: +44 131 455 5139 fax: +44 131 455 5141 e-mail: e.forster@napier.ac.uk * paper presented at:international association for social science information service & technology, building bridges, breaking barriers: the future of data in the global network, toronto, may, 1999. http://www.scot-homes.gov.uk/ http://www.assurcom.demon.co.uk http://www.assurcom.demon.co.uk http://www.scot-homes.gov.uk/indexnext5.html http://www.blaiseusers.org/tadeq/abouttdq.htm http://www.blaiseusers.org/tadeq/abouttdq.htm mailto:e.forster@napier.ac.uk iaiisist newsletter, vol. 2, no. 3 (summer iy7tt) nutf.s on the uistribution uf labuk in a social sciencks data infohmation netwokk. haul tvan peters social sciences intormatin utilization laboratory university center lor international studies university of pittsburgh fhamewukk tne point-oi-view wnicn i apply imlhuuucllon to mkufs and networks can be expressed fairly simply. within ine concern of these brief the last five to ten years we have remarks is to contribute to the witnessed a very important adjustaeliberations having to do with ment in the aspirations of those networks and machine-readable data involved in the "data library movefiles (mkufs; by calling attention raent," to use a single descriptor to the fact that one way to proceed to refer to those activities navng is to aistribute labor judiciously to do with secondary uses of social among network participants. to science data. we used to hear distribute labor is to distribute a rather frequent mention of national very large cost, and, therefore, and even international data librarinvestment, associated with network les. indeed, with varying degrees operation. distribution of investof success, some nations have ment often shows an added benefit embarked on such a course. more ot distributing commitment with as and more, though, we are focusing many people as possible having a our attention on perceived stake in network success national/ international efforts or failure. the distribution of directed at organizing mtorraation labor also invites specialization about data rather than data thereby network participants so that selves. 1 agree with this change those in the best position to of "target" and am optimistic that accomplish certain tasks are this dream can be realized, encouraged to do so without distraction irom other matters. this the sort of thing 1 have in mind modularity fosters incremental when 1 think of a "data information growth and is thus well attuned to handling system" (ulhs; can be smiting tinanciai arrangements and sketched as tollows: a changing population of participants. most of all, though, care1. the information flow ful distribution of labor is an begins with the primary organizing style uniquely appropriresearcherc s) . through ate to grass-roots movements strugthe efforts of this agent gling to legitimize their intera set of social science ests. it is style that can be data is created. if we found in the history and operation are lucky, a codebook is of lasslst itself. for all of also prepared. the flow these reasons 1 would like to take ends with a secondary a moment to pursue this thought to researcher, who saves see what it has to offer. time and money by learning of another by laiisiot newsletter, vol. i", no. i lijuinmer lyyts) researcher's data and using them. within "organizational reach" ol these researchers things are consideradiy simpiitied il some data handling agent (uha) is available. 'ihis agent could be a data archive or library, a statistical consulting service, a computer center, an academic/research library, etc. assuming the availabilty ol such an agent, either knowledge ol the researcher's data or actual responsibility lor them can lollow. 11' such an agent does not exist then knowledge ol' the existence ol research data is more problemmatic; both the dissemenation and acquisition ol data and inlormation thereabout is more haphazard , resting upon an uncertain set ol prolessional incentives. ol system design, operation, and user groups . j. uhas ol' mation economy scope kather all pos as need ually s tern ne with a tlhis simple simpler tage 1 justi f i as both ducer s wnateve unit d this e and qui ler a d handling in the ol than con sible r ing to b urveyed , ed only 11 possi is by problem .; anot s that ably be the pr and con r bib rives th stablish te poten ata sys ar cov ceiv esea e co th ke ble o me uha re imar sume iiog e s es a t id inlortem an ea of erage . ing 01 rchers ntione sysep up uhas. ans a merei y ad vancan garded y prors ol' raphic ystem ; rare entity a tio est "sh pre by led kee 1m the sys kno my was tha ent cth b) res all to que tag ice unn lar "ro tho to ing to minimal da n system ablished ared" inlo sent state word-ol-mo ic catalog ps abrea dings and problems tern are a wn to thi favorites te of uha t all hav ries in m e tip of t a waste earcher el have to uha askin stion ain , the berg; ; a ecessar y ger , bust" uh se with e under wri /promotion operations ta inlormawould be if the uhas rmation, the of affairs, uth and perupdates one st of new acqusitions. with such a 11 too well s audience, are : a ) a effort in e redundant ailing lists he iceberg) ; of secondary fort in that go from uha g the same repeatedly tip of the nd , c j an bias toward financially as , i.e., nough money te marketin addition we have beg to think 1 national or data inform agent cdiha tive record by some mea to the uiha would be e central int (. computer-b course ; . 1 bank would issue prod provide se major produ publ ication un , therefore, n terms of a international ation handling ; . uescrips of some type ns would flow ihere they ntered into a ormation bank ased , of he inlormation be used to ucts and to r vices. the ct would be a containing yu iaii^j.i>t newsletter, vol. ^, no. j lijumraer lyytt) tne descriptive records searcn service vendor, and a set ol' suitable such as lockheed, iiuu , indexes. ijucti a publicabh3, etc., or both. tion could take the lorm of a directory, whereby each edition replaced the it may seem that the problems faced previous one, or a perby llhijs are simpler than those lodical, whereby each faced by dihss. 1 submit that this issue reported new impression can be traced to the records and cumulative lact that llhbs have become a famindexes are printed once iliar teature ol academic lite and per year. two services that this fact can cloud our recolwould be provided, both lection of how things used to be in of them involving customthe 1s)5us and even, for many, the designed searches of the 19faos through the early 1970s. 1 information bank: a; believe that the similarities are desired data could be strong enough to invite coraparision described and the entire so that the "light" uk lihs experibank could be searched ence can be shed on uihs developithe "retrospective" raent. perhaps even a co-developsearch mode); b) desired ment of these two types of systems data could be described is in the otling; this is certainly and recent additions my hope and 1 don't believe 1 am could be searched on a unique in this respect, periodic basis tthe "current awareness" search mode) . analysis the dihs schematic just outlined resembles the structure of literamy analysis of the distribution ture iniormation handling systems of labor in dlhss begins with the ilihbsj as tne combined talents of observation that ihss always have library, computer, and mtormation one or more "bibliographic units" scientists have brought them into cbus) which are the life-blood of being: their operation. 1 believe that three types of bus can be found in 1. a primary researcher can a dihs: be viewed as an author while a secondary 1. the most complete type is researcher can be viewed that ot the actual data as a reader. and codebook themselves. these can be collectively '^. a dria can be viewed as an referred to as a "study academic journal or a documentation." it may book publisher. both be that these have been collect materials from transferred by a primary authors and make them researcher to a dha. the availabe lor use by readdiha snould not assume ers. that this is the case and should provide for partii. tne dlha can be viewed as cipation by individual a secondary publisher, researchers who desire to such as sociological maintain discretionary adstracts, or a computer control over their data yl lao^iu'i' newsletter, voi.. c", no. i liiummer lyva; i. even tnough they are willing to release information about them. i'he second type can be called a "study description" cijuk by this 15 meant a somewtiat exhaustive char acter izt ion ol the data: the research problem lor which they were produced, notable aspects ol' the methodology, time-1'rame, population, etc., and comments about condition, accessibility and so i'orth. 1 propose that the preparation 01 this be regarded as the responsiblity ol the individual primary researcher, uha , or both, i wish that it was reasonable to think in terms of a "standard bu scheme" but i think that the most we can hope for in the near future is the promulgation of a set of guidelines, much as academic journals offer to their contributors. 1 regard the 5u an dihs muc that of journal l1h5. t the bibl sent to t uiha the lyzed and tion" (iic pared . constitut entry t informati would be the ent lihss , vi article citation , keywords . this d its h as the artic he 5u lograp he uih i)u wou a " s ) wou these e th o the on ban very r les z , au title , abst ihe concep role 1 re indiv i le in would hic a. at id be tud y c id be 3cs e ac gro k. s i rn i 1 a bound thor n jou r act , uiria w t of in a gard dual a be unit the anaitapr ewold tuai wing they r to in ame , r nal and ould have its liha jour var i cont expe for . ble crea file had 5u r this oneindi coll libr othe thes to be a receipt is in 1 nal art ance in ent cted an it wo for t tea mi of all received eproduct could at-a-tim viduals , ection ar les , r basis e two al s fl of ts r icie fo shou d uld he crol th and ion be d e ba a basi or in tern ex ibl sus eceip s ; rmat id accou be po uiha iche/ e iius to o servi one sis , n en s , on bet ative e in as a t of some and be nted ssito film it ffer ces . on a for tire for some ween s . 1 have refe without nav in specific for the pur concern is believe an are forthc efforts in one, has areas. my ing what 1 with these red to ing sa they poses not a swers oming proces actio inter thin bus. the id m wou of t pre to from s an n gr est k se iius and uch about id look 1 his paper ssing one. this ques a number d laiiiilst, oups in is in desc hould be ijcs what ike . this 1 tion of for both ribdone this discus their processi to mind anothe sion of labor have presented arrive at a ul filed alter a and entered i bank. the r that all aspec the direct c ratner than their own scs that a valuab uniformity wo this means, the case insot tne :5cs consis sion ng b r que impl thin hi) w 5c h nto eason ts of ontro havin , is le i uld ithis ar as tent of y a stio icat gs, here ad b the wh the 1 g tha f no be is the wit the bus and ulhb brings n with diviions. as 1 a i)u would it would be een composed information y 1 propose sc be under of d uiha, uhas submit t 1 believe t necessary achieved by particularly "coding" of h the proviy^ iaiiiiist newsletter, vol. t! , no. i tiiuraraer lyvbj sions ol a vocabulary control simplil ication ol bu preparation protocol are concerned. iiucn proand transmission would hasten tocols are now regarded to de a development ol' sd standards. the necessary requisitie of reliable most advanced possibility along ihs operation.; 1 also believe, on this line would entail the sc the other hand, that the majority information bank being available in ot the inl'ormation provided by the much the same way that the uhio bl should be imbedded in the bu. uollege library tenter cuclcj has this would ease the i>c preparation made available library ot congress labor requirement at a uiha. tven clc) cataloging data. these ideas though this is the case it is are exciting and technologically important that a uiha be responsifeasible. nevertheless, without a ble tor the accuracy, completeness, firm base of support 1 fear that and general quality of the scs. setting our sights on these equipspecific, identified accountability ment cont igurations as a tirst step in this area is extremely useful. of development is ill-advised. l believe, to say it once again, that 1 have two additional observawe are involved in grass-roots tions on this subject. part of my organizing and should use distribuposition vis-a-vis the division ot tion ot labor concepts appropriate labor attendent with su and iic preto this, paration is based on my experience with journal article abstracts in there are also division of labor the l1h3 represented by united concerns to consider at the output states political political science end ot the system. in my scheme a documents lui>p:sl)}. this experience uiha would broadcast the entries in shows that, in general, there is its information bank by means of quite a difference between publication, directory or periodiabstracts written by authors and cal style, and search services, those written by document analysts retrospective and current awarein a l1h3. my belief is that the ness. to every extent possible, abstracts ot analysts are better the preparation ot the publication than those ot authors in the senses should be done at a uiha using that they are more accurately desadvanced information-base managecriptive ot an article's actual ment concepts such as those we use contents and that they are more for uspsu. a sophisticated proadequate predictors of source docugramming system, for which many ment relevance. my second observamodels are available, can produce a tion is that even though l have sc listing and all desired indexes described a situation in which sus from the same "linear" file as is arrive at a uiha by mail, l believe used tor searching purposes. this that the uihs should encourage submeans that searching and publishing mission of sus in machine-readable can conveniently interact so that torm. this may also have a bearing the intormation bank can be packon the standardization ot i>us aged and repackaged in a very issue. it is possible to conceive dynamic fashion with little or no of uhas presenting iiu data to a additionak keyooarding labor. ihe uiha through use ot an interactive most desirable output from such a program housed in the uiha compusystem would oe "camera-ready copy" ter. it is also possible to confor the publication m question; ceive ot a uiha making a program given this, the only "external" available to accomplish the same cost to be incurred by the publicathing using the computers of the tion would be that of printing and uhas. whichever, the 1i irtssist newsletter, vol. d. no. i uummer lyyttj dindi inter compu raent , would encod who a nave cond i an ar the such costl logic ng. nal tert be ed in oes ; pri tion rang pudl a th y an ally it the control based t hen the to prod agnetic most nters wti s should ement wh ication ing wou d it is necessa uiha or ypese best uce tapes metro do the ereby is id be no 1 ry. did not have some sort ol' tting equipal ternati ve appropr lately for someone politan areas under no uiha accept the text of rekeyboarded ; inord inately onger technoihe reasoning applying to provision of search services differs from that of publication. here the co-operation ol a uiha with other agencies is not only possible, it is desireable. the minimal method ot service delivery would entail the contacting ol the uiha by someone interested in having a search performed. a discussion ot the characteristics of desired data would be held with a "retrieval specialist" who would compose "search strategies," review results, and lorward what seemed to be the most relevant entries in the information bank to the inquirer. the experience of llhijs leads me to believe that, depending upon volume, tne "turn-around" time for such a system would run irom lu to ly days and would occassional ly take even longer than this. evidently, a more time-responsive approach would be preferrable. the manner in which this could be accomplished would entail the provision of access to the bu file by way of an on-line, interactive search and retrieval systeiri. iwq uiha could engineer such a system itsell or it could transmit its information bank to a vendor of searcn and retrieval services, someone like lockheed, iiuu, or bhi> to name the current three major commercial actors. 1 favor the latter approach primarily because it woud reenlorce what 1 consider to be opmen secto it w which ulhb this opera tant ind iv their resul form and r answe an ex tint r of t ould would in it sort tion . thing idual own ts as of onetr lev r . trem he he i also hav e ot s he wou or sear soon line al ely promi 1 iteratur n formatio reduce e to be b tlorts to earch an gardless , id be to a uha ches and as possi , interac system wo sing devele searching n industry. the costs orne by the estadlish d retrieval the imporenable an to pertorm to receive ble. iiome tive search uld be the two footnotes can be added to this discussion. birst, 1 believe, consistent with my reported oeliefs throughout this paper, that the place to start is with the provision 01 search services through a retrieval specialist attached to a uiha. it is a tried and true technique. it concentrates investment in labor rather than equipment, it can be initiated quickly, and it provides a powerful evaluative framework. nevertheless, the eye of the ulhb would always be on the on-line, interactive approach and as soon as adequate financing decame available and oroad popular interest and support became evident the ulhi> would move in that direction. it is also important to note that such a movement would not render obsolete the retrieval specialist role at the uiha. my experience with llhss, again, provokes this observation. some individuals and uhas , tor various reasons, will never want or be able to assume responsibility for their own searching and retrieving activities. others will perform searches so infrequently that they will desire to consult with someone with more experience. in short, retrieval specialists attached to any ihs always add value to computer-based search and retrieval procedures in a way analagous to y4 iaiisist newsletter, vol. ^, no. i liiummer i^ybj tne value added by a carpenter to a 1 am lelt with tne impression tnat nammer. mucn more needs to be said. my intention has been to direct attenmy second footnote is tnat the tion to the question ot who will do uiha should nave the capaoiiity to what in the ulhbs we are all envitransmit the iic intormation bank sioning. this seemed to be worth into any requesting agent. this doing because much of what i have means that suitably equipped and been reading and hearing recently motivated individuals and uhas has been distionctly oriented could acquire the entire informstoward equipment concerns. it tion bank and updates thereto so occured to me that it was once that tney could process their own again time for someone to raise the searches locally rather than issue of the people dimension ol remotely. for agents with a high the data information production, volume ol' searching activity this acquisition, organization, and dispossibility will allow greater cost semination process under discuscontrol than the other arrangements sion. this is what 1 have discussed so far. it will also, attempted to do. l believe that need l say it again, distribute this concept of "distnoulabor in a useful way. tion/div ision of labor" which has occupied my attention in this paper is extremely important and can provide a very concrete principle cunululalun guiding how we choose among competing courses of action. i hope that these "brief" remarks have wound its articulation and examination up as rather "extended" remarks but will lead to further discussion. ytj vol18172 14 iassist quarterly the advent of the microcomputer has introduced great uncertainty to those persons responsible for ensuring the viability of information created on a computer. yes, magnetic tape, round tape, has been recognized for years as a fragile medium. but at least there were recognized standards for this media. if someone orders a file from icpsr and they sent a tape to the requestor’s university, the information would be accessed easily by that university’s mainframe. but oh, what choices have appeared in the last five to ten years. i am willing to wager that each of you have come into contact with people who are using cd-rom, worm, exabyte cd-recordable, diskette—i could go on. since for most of us as archivists or librarians, the purchase or acceptance of files on these various media imply a commitment to ensure their accessibility to users for a certain period of time, the decision as to what media to use for storing information has a great impact on the computer operations of our institution. deciding what types of storage media to use involves determining the length of time your institution expects to have this information available, the computer resources currently available at your institution and your comfort level. what i would like to discuss this morning are the general steps that should be considered by all institutions in seeking to preserve the information stored on electronic or optical media and then to briefly summarize the preservation requirements for the media currently in wide use. what do we mean when we say we are preserving information? how long will a product, a document or the information be preserved? i believe in many cases, there are underlying assumptions by the various groups that employ terms such as “archiving data” that are not necessarily shared by the larger community. permit me to use a personal story. when i began to work as an archivist with the office of presidential libraries, i was responsible for preliminary preservation work for the color negatives from the carter administration. since i had no previous experience in working with still photographs, i attended a workshop on preservation of photographic materials at the eastman kodak institute. the information provided was timely and instructive. my most vivid memory, however, was the demonstration of what happened to color film produced in the 1960’s. we were viewing slides from the movie “west side story.” as the lecturer explained the dye process used in manufacturing the film and showed a slide taken from one of the stills, the audience audibly gasped at what they saw: the projected image from the slide projector had only one color-magenta. one could almost feel the dismay as the audience realized that this famous cultural icon had deteriorated to the point that the only color on the film was red. now, i will grant you that the basic information was still recorded, but for the millions of consumers who had been urged to buy this film to record all those wonderful family memories, would a film of children on vacation at the beach mean much to them if the entire photograph was red? in many ways this is an unfair example, but it points to the fact that for many the concept of preservation is abstract and the question of longevity, unacknowledged. it is the purpose of this paper to discuss how general preservation practices can and should be applied to electronic media and to suggest that for many institutions or organizations there is a need to carefully consider how long the information currently stored in an electronic format will be retained by that institution. that decision will influence the media used to store the information electronically. first, let us discuss general preservation activities, as they relate to electronic records. when i began to research the basic preservation requirements for electronic records i was struck by how the requirements for textual materials seemed to mirror those for electronic records. i will be the first to admit that electronic records pose their own unusual problems, but , and it is a large but, the general maintenance and environmental requirements are very similar for textual and non-textual materials. many of the preservation policies that are constructed around attempts to prevent deterioration are just as relevant to electronic records as they are to paper based records. in their discussion of implementing an archival preservation program, norvell jones and mary lynn by fynnette eaton1\ chief, technical services center for electronic records u.s. national archives and records administration electronic media and preservation 15spring/summer 1994 ritzenthaler detail the interrelated factors that cause archival records to deteriorate: the chemical and physical stability of specific materials, storage under adverse environmental conditions, and external causes such as excessive or careless handling, and loss or destruction brought about by human-induced or natural disasters. in every case the factors enunciated on this list are factors that must be considered in the preservation of electronic records as well. an understanding of the physical properties of electronic records and the environmental conditions that they should be stored under are essential for ensuring that the information stored on these records are preserved. the seven elements of a preservation program: environment, storage, handling and use; microreproduction and reformatting, exhibition, disaster planning and treatment must be considered by an institution charged with preserving electronic records. the only element that does not have real importance in electronic records is exhibition. as with other media, perhaps the single most important factor in the preservation of electronic media is the environment. electronic records, like other audiovisual records require temperatures between 62-68 degrees fahrenheit, with an optimum of 65 degrees, which is probably within the range required for textual records. the humidity requirements, however, are different for magnetic tape than for paper. lower humidity between 35 and 45 percent, with an optimum of 40 percent is the recommended level according to the national institute of standards and technology (formerly the national bureau of standards), but this is less than the 50 percent recommended for paper records. according to george cunha, the commonly accepted view currently held is if audiovisual materials (including magnetic tape) cannot be isolated in a mini-environment, then the overall humidity in the building should be kept between 40 percent and 50 percent.” successfully attaining the optimum environment recommended can be difficult. most institutions have conflicting requirements for staff and various media. one must recognize the difficulty of creating the perfect environment with competing interests, and take to heart what one conservation authority has learned: “it is far more important to stabilize both temperature and humidity at points as near as possible to the optimum conditions than to strive for optimum conditions with heating and cooling machinery that is unequal to the task and likely to produce constantly fluctuating temperature and humidity levels.” i would like to emphasize this point as well. studies indicate that one of the major contributions reducing the life expectancy of magnetic tapes is fluctuating temperature and humidity. strive for the best conditions possible, but emphasize stability rather than occasional optimum conditions. the proper storage and handling of archival materials is an essential element in a preservation program, particularly for paper records; but, again, this is also applicable to electronic records. proper storage includes placing open reel tapes in plastic canisters and storing these tapes or cartridges vertically in shelving constructed specifically for open tape reels or tape cartridges. unlike paper, which can be stored indefinitely if placed the proper containers, reels should be exercised periodically (there is discussion as to how often this should be done) and there should be a periodic inspection of a random sample of files, to test the readability of the media. an interesting theory proposed by margaret adams, who oversees the reference activities at the center for electronic records is that, unlike paper records, reference activity actively promotes preservation in electronic records, because the staff uses the files, thereby determining the readability of that specific file and the media is cleaned and rewound after use, thus ensuring proper tensioning of the media. improper handling can have disastrous effects on magnetic tape. dirt can create read errors. if the tapes are not tensioned properly, stretching can occur, which would create misalignment, leading to the inability of the computer to process the tape. any distortion of the data due to improper tension or shrinking or expansion of the tape, or erasure of the tape can lead to the loss of the information stored on the tape. improper handling of magnetic tapes or tape cartridges can cause edge damage as well. thus procedures for ensuring the proper handling of electronic media must be an integral part of a preservation program for electronic records. reformatting, the next element in a preservation program is absolutely essential with electronic records. the requirement of moving electronic records to new formats is to keep up with the ever-changing technology. as the national research council pointed out in their study “preservation of historical records” and the national institute of standards and technology (nist) has confirmed, the recording media in use may well outlast the hardware, thus making it necessary to recopy the electronic file every 10 to 20 years to ensure access to the information. this recopying process simply reformats the information to avoid obsolescence. the information is not changed in any way. disaster planning must be a part of any preservation program. electronic records are susceptible to water and 16 iassist quarterly fire damage. the best way to protect the information in electronic format is by making a second copy of any file and storing it offsite. the costs of a second copy are minimal compared to the expenses that would be incurred in trying to recreate the data. it is highly recommended that there be a second copy of any file stored in any electronic format, even ( and i would say particularly) diskettes. treatment, the last element discussed by jones and ritzenthaler, does not figure as prominently with electronic records, although the national archives recently encountered problems with some of its older tapes and is working with the national media lab, in minneapolis, minnesota to find a way to salvage as much of the information from these tapes as possible. generally, the best method of treatment is prevention, recopying electronic files before serious problems develop. these then are the basic elements of a preservation program for archival materials. i have focused on their relevance to magnetic media. but what about the other media available on the market? do cd-roms, optical disk systems and diskettes require the same type of program? generally, i would say the answer is yes. optical disks and cd-roms have been touted as being extremely durable. perhaps yes, perhaps, no. there has been little empirical testing performed on these media. what you have heard are vendor claims and some horror stories. in certain cases, the seal on the cd-rom was not perfect, so oxidation occurred and information was lost. there are indications that information stored on the outer layers of optical disks tend to have greater proportions of errors. nothing is failsafe. clean environments should be required for any media. temperature and humidity should be controlled for best possible results. disasters must always be planned for, so there should be a backup copy of any file that you are required to preserve. i must admit however, that there is no requirement for cleaning and rewinding of optical media. is the data permanent on these media? no. but the reason is not necessarily the medium. some of these disks could well last 100 years. it is the technology that will fail. as i was preparing this paper i received a publication entitled government imaging, which claims the title of “the national newspaper for government imaging technology.” in an article about standards for optical disk storage systems there is the clear acknowledgement that optical disks are not necessarily the best media for archival storage. in discussing the various standards used within document imaging systems, the author (harvey spencer) states the problem being “ . . . that we are relying on these disks being available to us in twenty or maybe more years time and it is highly likely that the drives, and formats, that we are writing in will no longer be supported. . .” he goes on to explain that the only standard that has survived from the 1960s is the 1/2" magnetic tape. the reason for its durability was the domination of the computer industry by a handful of suppliers that everyone used; the amount of information stored on these tapes is so great that manufacturers can not abandon this format. for optical disk systems, this situation does not exist. cd-roms look more promising, because of the number of files being published on this media, and the acceptance by the library community as a means of information distribution. there have been questions for a number of years about the longevity of the polycarbonate cd media. an organization which is interested in promoting the use of cd-roms by government agencies, sigcat or special interest group on cdrom applications and technology, is trying to collect information on this issue. one member, ron kushnier, a storage specialist with the naval air warfare center in warminster, pa reported to sigcat members last spring about his extensive environmental tests of cd media from about 100 manufacturers. what he found was that “all cd-roms are not created equal.” some disks came out of high-humidity, high temperature chambers in as good shape as they went in; others failed miserably. sigcat is continuing its efforts to determine longevity for this medium. yet there is again the issue of standard, or i should say, the lack thereof. charles dollar, a member of the archival research and evaluation staff at the national archives has argued that disk longevity takes second place to “a much more important and pervasive issue—how to deal with technology-dependent records,” dollar used as an example relevant for cd’s—data compression. although there is an international standard for data compression, many vendors use proprietary compression techniques “that in essence become an encryption tool that only one vendor’s software can open or close.” the point that i am trying to make with this discussion of the limitations of various media, is that you and your institution should consciously decide how long you intend to preserve information in an electronic format and base the decision of the which format on the length of time you will need access to the media. if it falls within ten to twenty years, then optical disk or cd-rom is a valid choice, although you must monitor changes in technology in the marketplace and the condition of the equipment you use to access this information. if the requirement is for longer-term preservation, you can still use cd-roms and optical disks, but you must plan to reformat the information onto a technology that can be accessed in the future. there is no panacea for electronic media. it is a very small cost of migrating files to newer 17spring/summer 1994 format, to preserve previously unimaginable amounts of information and making this available to a world community. 1. paper presented at iassist 93 in edinburgh. sources norvell m.m. jones and mary lynn ritzenthaler, “implementing an archival preservation program,” in managing archives and archival institutions, ed. james gregory bradsher (chicago: the university of chicago press, 1988), 188. george martin cunha, “current trends in preservation research and development,” the american archivist, vol. 53, no.2, spring 1990, 195. sidney b. geller, care and handling of computer magnetic storage media, national bureau of standards special publication 200101, (washington, d.c.: national bureau of standards, 1983), 86. cunha, p. 195. jones and ritzenthaler, 191-2. bruce i. ambacher, “managing machine-readable archives,” managing archives and archival institutions, ed. james gregory bradsher (chicago: university of chicago press, 1988), 124-5. national research council, preservation of historical records, (washington, d.c.: national academy press, 1986), 61-2. harvey spencer, “standards for optical disk storage systems,” government imaging, vol. 2 no. 3, may june 1993, 15. florence olsen, “experts debate longevity of polycarbonate cd media,” government computer news, june 8, 1992. ibid. 10 iassist quarterly spring summer 2009 marika de bruijne and alerk amin1 questasy: online survey data dissemination using ddi 3 introduction based on the ddi working paper ‘questasy: online data information and dissemination using ddi 3’. special thanks to the co-authors of this paper: michelle edwards, oliver hopt, jannik jensen, dan kristiansen, olof olsson and joachim wackerow abstract: questasy is a web application developed to manage the dissemination of data and metadata for survey projects. it was primarily developed for the liss data archive, but was designed to be repurposed for other archives as well. the application has been operational since 2009 and its external web interface can be viewed at: www.lissdata.nl. questasy manages both metadata and data and provides an easy-to-use data entry module for administrators to create metadata. the external web interface allows researchers to browse and search the metadata and download datasets. the questasy system also manages files, tracks downloads, and creates web pages for viewing documentation including studies, concepts, questions and variables. due to the longitudinal nature of many of the liss panel studies, the ability to track questions and variables throughout a study was a key requirement of the system. to support this, ddi 3 was chosen as the basis for the structure of the application, from the underlying database to the generated web pages. this paper describes the main functionality of questasy and how we designed and implemented the system. 1. background: liss panel the liss (longitudinal internet studies for the social sciences) panel is the principal component of the mess (measurement and experimentation in the social sciences) project. it consists of 5000 households in the netherlands, comprising 8000 individuals. panel members complete online questionnaires every month, totaling approximately 30 minutes per month. half of the interview time available in the panel is reserved for the liss core study. this longitudinal study is repeated yearly and is designed to follow changes in the life course and living conditions of the panel members. the other half of available interview time per year is used to collect data for external research projects in different disciplines within social sciences. this is cost-free for purely scientific research. researchers from both the netherlands and abroad can submit survey proposals. the panel has been in full operation since the end of 2007. next to the liss core study, many of the other liss studies are longitudinal. in these studies, questions are repeated in new measures (waves) to the same respondents in order to measure changes over time. one of the goals of the liss panel is to make the collected data available to the international scientific community. due to the longitudinal character of the data, we needed to pay special attention to the way the data and metadata of all measures of a study would be presented. 2. the project questasy was developed with the researcher in mind, but in two capacities: one as the internal employee (referred to as “administrator”) and second, as an external user of the metadata and data (referred to as “researcher”). administrators clean the collected data and prepare data files in spss and stata formats, then load them onto questasy for immediate access by researchers. once the data are available in questasy, researchers from around the world can search and browse the metadata for each of the surveys available. 2.1 start in order to disseminate the data and metadata for the liss panel surveys, we started with relatively simple application requirements. soon we realized that if we wanted to provide researchers with a better tool for browsing and searching the metadata of our largely longitudinal studies, we would need a more advanced solution than, for instance, simply publishing pdf or word codebooks. at the beginning of the project, we evaluated several existing applications. we investigated software packages, such as nesstar, but also looked at other custom-built websites. our requirement to support longitudinal studies eliminated most options. the options which did support iassist quarterly spring summer 2009 11 longitudinal studies were not flexible or comprehensive enough for our needs, so the decision was made to build our own system: questasy. 2.2 choice for ddi 3 the choice for ddi 3 was initially not an obvious one. the main reason that ddi 3 piqued our interest was its support for longitudinal studies. when we started studying the ddi 3 way of managing each detailed entity of a survey project as its own controllable element, this first seemed to cause a large amount of extra work in terms of data entry, which we hadn’t expected in advance. in particular, the separation of question and variable metadata, and full documentation of both, required ‘out-ofthe-box’ thinking from our administrators who were used to think about data dissemination in terms of codebooks, containing mainly variable metadata. although we believed the structure of ddi 3 would support our needs, it was in fact only later through practical examples of more complex survey data, that the full advantage of this structure was truly experienced. finally, we had found a system that enabled documenting the questions the way they had been presented to the respondents and the data variables the way they were contained in the final dataset, without losing information from either side. 2.3 ddi elements in questasy after having made the decision to use ddi 3, questasy developers created a proposal outlining a system that met the internal user requirements and used ddi. administrator involvement at the beginning was crucial since implementation of the project would affect their workflow. with no ddi experience, questasy developers approached the administrators with diagrams and “english” vs. ddi translations. talking the same language is a great benefit. during the initial planning phase, we found fields in ddi for the metadata that were already being documented and included in codebooks. we also looked through ddi for ideas about new metadata that could be delivered, including metadata that was already available internally, but not being distributed to data users. word documents with tables and diagrams travelled back and forth between developers and administrators to result in a comprehensive list of items the administrators felt met their current and immediate future needs. technical fields that were important for the developers to maintain were also included in the project implementation. the ddi elements that were chosen to be used by questasy are: • citation • code / code scheme • coding • collection event • concept / concept scheme • conceptual component • data collection • funding information • group • organization / organization scheme • other material • physical data product • physical instance • question construct / control construct scheme • question item / multiple question item / question scheme • representation • response domain • study unit • variable / variable scheme 2.4 ddi 3 as architectural basis ddi 3 was chosen as the architectural basis of questasy. after having analyzed the ddi hierarchy to determine which elements were required to document the liss metadata, the chosen elements were then converted to a relational database schema. figure 1. figuring out relations between questions and variables 12 iassist quarterly spring summer 2009 the major ddi elements, such as question items and variables, were mapped to tables in a relational database. the fields in the tables correspond to the fields in ddi. the relationships between ddi elements were extremely important to the system. the biggest benefit is the tracking of question items across waves in the study, where each wave can have question constructs and variables that refer to the same question item. relations between ddi elements can occur in 2 ways: through the normal hierarchy, or through references. both of these types of references were mapped to one-to-one, one-to-many, and many-to-many relationships between the tables. for manyto-many relationships, join tables were used. a couple of fields were added to the database for internal use, and do not map to the ddi hierarchy. an example of this is in the control construct schemes table, to which we added the name of the original source file for the questionnaire. this working name is used only for internal purposes and therefore not shown in the researcher interface, nor is it part of the ddi standard. as these extra fields will not be part of a ddi export, it was important to find a ddi equivalent for all the fields we wanted to publish. substitution groups in ddi presented some problems, specifically for response domains. in ddi, a response domain is placeholder for a text domain, numeric domain, code scheme, or other domain. this type of inheritance can be resolved in several ways in the database. we chose single-table inheritance. the response domains table contains all of the fields required for all of the various domains, and a flag to signify which type of domain it is. the database design does not support versions of elements, with one exception. sometimes, errors are found in datasets after they are released. these errors are fixed, and new datasets are released. to support this activity, variable schemes can be versioned, to keep track of the various releases of a dataset. however, only the latest version of a dataset is downloadable by researchers, and only the metadata for the latest version is displayed on the website. some ddi elements require a tree structure, such as groups/study units, concept hierarchies, and control constructs. these were implemented in the database schema using left-right trees. these are supported by the application framework and have very good performance for query operations. 2.5 system design once the database schema was determined, the web application was built using a php framework. the application queries the database to save/retrieve data (such as questions or variables). it then formats the data as html, which are then delivered to the end user’s web browser. both the web forms for data entry as well as the views for researchers are created this way. we decided to use a php framework on top of a relational database due to several factors. the most important was previous experience within centerdata. xml databases were considered, but were rejected based on performance considerations. we determined that a relational database could easily scale to the usage we required. as ddi is a file-format standard, we decided that we could easily interoperate with other systems via a ddi import/export implementation. thus, our choice for the internal storage would not affect other systems. database transactions are an important part of the system, to maintain the integrity of the database. many data entry screens can affect multiple tables in the database. for example, entering a new question can create new question items, response domain, code schemes, and codes. all of these data are collected via a single web form, and then processed at once in a transaction. if the inserts are successful, the transaction is committed. if there is an error in processing the data, the transaction is rolled back, and the errors are shown to the administrator, who can then fix the errors and resubmit. this ensures the integrity of the database, especially that all of references between elements remain valid. the search is a very important feature of questasy. we wanted to make it easy for researchers to search for text that might appear in different tables, such as question and answer text. for this reason, we decided to use a search engine that would index across tables. the sphinx search engine was chosen because of php support, performance, and flexibility. 2.6 overall timeline developing the framework for the project with the researchers took approximately two months, followed by another twelve months of development and refinement. a large portion of the design time was spent on determining how to use ddi as a basis for the project. it took approximately 8 months to develop a working system and another 4 to tweak and massage the system to match everyone’s requirements. during the development phase of the project there were 1.5 – 2 fte on the project, with 1 fte currently available for both development and support of the production system. 3. website usage the liss data archive website went live in early 2009, for internal administrators to begin entering data and metadata. the site went live for external researchers in march 2009, and has been well received. traffic to the website has been better than expected, and web statistics show that researchers are making use of all of the functionality available. iassist quarterly spring summer 2009 13 3.1 data entry in questasy, one can capture metadata from many aspects of the lifecycle of liss surveys. from the data production process, metadata about the concepts, collection, and processing steps are collected. and of course, the resulting data files can be stored in questasy. for the administrator interface, we have implemented a full system of web forms to enter and manage the data and metadata. the web forms handle the ddi relationships automatically, making it easy for administrators to enter the metadata without prior knowledge of ddi. metadata is currently entered into questasy by a student working approximately one day per week. to speed data entry, we have implemented an automatic import for variable metadata captured in spss. in spss, the administrator first creates an xml file in an oms session (output management system) for the spss dictionary information. this can be imported into questasy to import the variable metadata, including representations. the administrators then enter the questions and additional metadata associated with the variables via web forms. 3.2 external website on the public website, researchers can browse studies that have been conducted in the panel. in the study unit view, they can see the metadata for the studies, including information about the abstracts and data collection. they can also download the data files and other materials associated with each study. to download datasets it is first required to fill out an agreement about the use of data. researchers can also browse the concepts. the concepts are organized into a tree. from the concepts, researchers can view the associated variables and continue to navigate through the metadata. when viewing question items, the variables that are associated with the item are listed. for longitudinal studies, this includes the variables across all the waves of the study. while the questasy database contains questions items, control constructs and question constructs, we don’t present all these relations on the external website. here, we have simplified the information about questions showing an integrated view for question item and question construct information. in fact, question constructs, and related question item information, are shown as a list, in the order that they are asked in the questionnaire. we have also integrated a search engine with full text indexing. the advantage of the search engine is that it can index complex items, including fields combined from multiple tables. the search results are ranked according to relevance, providing the researchers with the best results possible. the searches also execute significantly faster than against the mysql database. all surveys for the liss panel are conducted in dutch. in the liss data website, question texts are distributed with the original dutch text, as well as the english translation. all other metadata are distributed only in english. the questasy application is capable of supporting multiple languages throughout its interface, although this is not used in the liss data website. 4. issues and restrictions looking at ways in which our use of ddi is somewhat restricted, the use of grouping could be mentioned. grouping currently occurs at the top level for studies and is mainly limited to concept and organization schemes. grouping also occurs within longitudinal studies, to allow individual waves to share a common question scheme. questasy takes advantage of ddi inheritance to reuse items within individual studies. figure 2. questasy supports data entry via webs forms and automatic spss metadata import 14 iassist quarterly spring summer 2009 currently, only variable-level metadata can be automatically imported into the system, via spss oms files. in the future, we would like to automate more of this process. especially importing of question and response domains from the questionnaire engine would make a big difference in required data entry effort. full question flow from the blaise system, which we use in data collection, is also not captured at the moment, but is of great interest to developers, internal administrators and external researchers. the current solution uses the universe element to give some routing information, but investigation is underway in how to implement the full questionnaire flow. finally, while questasy was developed primarily for one application, the liss panel, it was designed to be customizable and extendable, so that it could support other studies in the future. 5. outlook and future developments the questasy system has received very positive feedback from both internal and external researchers. however, developers are looking into the future for further improvements and developments. questasy can deliver the information in several ways. currently, the web interface for researchers is operational. once the metadata and data are in questasy, the system functions as a local archive, but the data could be exported to a permanent archive. the development of an export function to generate ddi xml is currently underway, to deliver the structured metadata to other applications. there have also been some experiments on delivering content to portable devices such as ipods. in the future, paper codebooks in pdf/word format might be generated directly from questasy. the great benefit of a relational database in which the ddi structure is embedded is that this makes it relatively easy to create new forms of ddi compatible exports. a customized basket for picking out variables from waves across the years is a feature that would enhance the researchers’ ability to download data. while current downloads are restricted to the dataset level, the researchers’ questasy experience would be improved by adding variable level sub-setting. current projects underway are looking at the integration of enhanced publications, with up to variable-level metadata about the publications. this involves giving researchers the ability to list the publications they have written based on questasy data, and link the publication to the studies and variables they used in their research. finally, as the second wave of the liss core study is being released through questasy, including changes to some questions, harmonization and the relationship between and among similar variables have become of great interest to us. additional views may be implemented to provide researchers with a clear overview of how the questions and variables are comparable across waves. the comparison module will be investigated for this. 6. summary: main benefits to summarize, the main benefits we have experienced figure 3. questasy provides possibilities for delivering metadata in various ways iassist quarterly spring summer 2009 15 while using questasy and ddi 3 for disseminating our data are: 1. separating the documentation of questions and variables and enabling many-to-many relations between these fully documented variables only point to the survey questions they are based on. this enables full documentation of both entities within a survey project, without losing any information such as original question formulation. for example: the data based on a question where multiple answers are possible are often processed into several ‘dummy’ (0,1) variables, one variable containing the answers to one answer option. in questasy this can be presented by creating a single question and several variables which each point to the same question. 2. longitudinal data comparison via questions within longitudinal studies questions can be reused within a longitudinal study. this enables creating a connection between variables in different waves that are based on the same question. 3. several options for searching data on detailed level not only information on study level, but also metadata of both questions and variables can be searched by entering keywords. the website visitor can choose in which fields to search: study units, concepts, dutch or english question text or variable metadata, or all of these. it is also possible to browse the contents of the database in different ways: by studies or concepts, for example, and in the future by topics and year. 4. relational database enables presentation of metadata in many ways thanks to separating pieces of metadata on a detailed level into their own manageable fields in a relational database, questasy provides a flexible basis for creating various different views on the website to data users. further, since the metadata of questasy is structured according to ddi 3, it is relatively easy to create new forms of output that are ddi compatible. notes 1. marika de bruijne (m.debruijne@uvt.nl), alerk amin (a.amin@uvt.nl) . centerdata, university of tilburg .po box 90153, 5000 le tilburg, the netherlands.phone: 013 466 8325 / 8326. fax: 013 466 2764 iassist quarterly fall & winter 2007 by robin rice* introduction disc-uk (data information specialists committee united kingdom) is a forum for data professionals working in uk higher education who specialise in supporting their institution's staff and students in the use of data for analysis (primarily statistical and geo-spatial). this partnership, led by edina, is carrying out the disc-uk datashare project (march 2007 march 2009)1 that aims to explore new pathways to assist academics wishing to share their data over the internet. with three institutions taking part – the universities of edinburgh, oxford and southampton – plus the london school of economics as an associate partner, a range of exemplars will emerge from the establishment of institutional data repositories and related services. it is part of a wider programme to develop institutional repositories funded by the uk’s joint information systems committee (jisc). this project brings together the distinct communities of data support staff in universities and institutional repository managers in order to bridge gaps and exploit the expertise of both to advance the current provision of repository services for accommodating datasets. the project's overall aim is to contribute to new models, workflows and tools for academic data sharing within a complex and dynamic information environment which includes increased emphasis on stewardship of institutional knowledge assets of all types; new technologies to enhance e-research; new research council policies and mandates; and the growth of the open access / open data movement. this article will summarise the work of datashare in the following areas: defining the institutional data repository in the broader landscape and within the ‘data sharing continuum’; investigating deposit of research data in institutional repositories including metadata and policy development; understanding and improving data management practice through partnering with academic departments in the use of the data audit framework; and licensing issues including ‘open data’. data and institutional repositories according to the jisc-commissioned digital repositories review (heery and anderson, 2005), a repository is differentiated from other digital collections by the following characteristics: • content is deposited in a repository, whether by the content creator, owner or third party • the repository architecture manages content as well as metadata • the repository offers a minimum set of basic services e.g. put, get, search, access control • the repository must be sustainable and trusted, well-supported and well-managed. institutional repositories are those that are run by institutions, such as universities, for various purposes including showcasing their intellectual assets, widening access to their published outputs, and managing their information assets over time. these differ from subjectspecific repositories, such as arxiv (for physics papers) or repec (research papers in economics). the project – along with others funded simultaneously – will help to realise the vision of the digital repositories review of a “coherent aggregation of content from a network of institutional repositories”, and more particularly of the digital repositories roadmap, e.g. the milestone under data: “institutions need to invest in research data repositories” (heery and powell, 2006). there are of course some notable centralised data archives and centres serving particular disciplines in the uk, such as the uk data archive/economic and social data service (ukda/esds) for the social sciences and the natural environment research council (nerc) data centres for natural and environmental sciences. other disciplines have created vast online databases on the internet or over e-research grid networks, which is the logical place for ‘publishing’ data outputs in those domains. (digital archiving consultancy et al, 2005; swan and brown, 2008a.) this project does not aim to challenge these nationally funded organisations that have set internationally recognised high standards in data archiving, management and curation, nor the model of domain-specific data archives/centres. it does, however, aim to explore the role of filling in the gaps left open by the paucity of coverage disc-uk datashare project: building exemplars for institutional data repositories in the uk 22 iassist quarterly fall & winter 2007 of dedicated data archives, and in doing so, gain leverage from being able to work closely and directly with potential depositors at one’s own institution. indeed, the lifecycle approach to data sharing encourages intervention at the earliest stages of a research project to ensure adequate consent, documentation etc., are achieved for the data to be usable by others (humphrey, 2000). one of the first tasks of the datashare project was to learn what repository managers had done in earlier projects and to what extent institutional repositories (irs) in the uk were already dealing with data. a disc-uk member therefore conducted a thorough state of the art review, which included depositors’ motivation and barriers for depositing data in an ir (gibbs, 2007). the data sharing continuum institutional data repositories are only one possible response to data publishing requirements of creators and funders, and they have limitations, such as lack of ability to visualise or manipulate the datasets online (using dspace, fedora, or eprints software as it exists at present). the dataset, e.g. data files plus documentation, must be downloaded from the repository and analysed on a desktop computer using requisite software. this marks the “zip and ship” level of data sharing that irs are well-suited to host. by adding value in terms of interacting with users to enhance metadata, documentation, and to reformat data into suitable sharing and preservation formats, we see our institutional data repository services as sitting comfortably in the middle of the data sharing continuum shown below. the increased level of human effort required to curate data at the highest levels may be reserved for the most important, special, or highly-used datasets, as is done at national data archives. on the other hand, by raising awareness, providing local services, and offering a repository with a simple deposit interface, there is scope for the numerous datasets languishing on portable drives or with minimum bitstream backup only, to be moved up a notch or two on the scale, and therefore not lost to potential new uses and to the scholarly record (see figure 1 page 24). in some cases it is the data creators themselves who wish to re-use the data later on, after they have moved onto other research projects. if they have documented and deposited their data for sharing purposes, they will not have the experience of many researchers of not being able to find, read, or interpret their data at a later time. the project is looking at assisting researchers not only with depositing their data in an institutional repository or data archive, but also with using web 2.0 tools to “mashup” their data for online visualisation through numeric-based applications such as swivel and geo-spatial tools such as openstreetmap. there are of course advantages and disadvantages of using these type of “cloud” computing applications to publish academic data, which is covered in two briefing papers produced by the project (macdonald, 2008a & 2008b). metadata and policy development application of appropriate metadata is an important area of development for the project. datasets are not different from other digital materials in that they need to be described for discovery and also for preservation and re-use. the grade project found that for geospatial datasets, dublin core metadata (with enhancements such as drawing a bounding box to enter geospatial coverage) give sufficient context for discovery within a dspace repository, though more in-depth metadata or documentation is required for re-use after downloading (seymour, 2007). the project partners are examining other metadata schemas such as the data documentation initiative (ddi) versions 2 and 3, used primarily by social science data archives (martinez, 2008). crosswalks from the ddi to qualified dublin core are important for describing research datasets at the study level (as opposed to the variable level which is largely out of scope for this project). datashare is benefiting from work of the dryad project -a repository for evolutionary biology2 (carrier, et al, 2007) and gap3 (geospatial application profile) in defining interoperable dublin core qualified metadata elements and their application to datasets for each partner repository. the solution devised at edinburgh for dspace makes use of just twenty fields from dublin core (simple) and dcterms (qualified) and attempts to follow dublin core metadata initiative (dcmi) recommendations in applying them to datasets (rice, 2008). one innovation introduced is to use a look-up to the open utility geonames for ensuring consistency in entry of placenames for geographic coverage. datashare also is developing a briefing paper4 to provide a range of requirements that repositories can consider as they plan to add research datasets to their digital collections. the briefing paper discusses the scope of a data repository, its content policies for types of files and data sets held, its metadata policy for descriptive information about items in the repository: its submission policy, and its quality, copyright, and preservation policies. background material was gathered from the online opendoar policies tool5 maintained by sherpa at the university of nottingham, the oais information model and the trac checklist, and other sources. data management partnerships the stewardship of digital research data report, (research information network, 2008) examined the responsibilities of research institutions, funders, data managers, learned societies and publishers in turn. for example, research councils may choose to fund a domain data archive, as the economic and social research council iassist quarterly fall & winter 2007 23 figure 1: data sharing continuum 24 iassist quarterly fall & winter 2007 does for the uk data archive, or they may require grant applicants to include a data sharing plan, as the medical research council has been doing since 2007. similarly, the data quality seal of approval, developed by dans--data archiving and networked services—in the netherlands, stakes out roles and responsibilities of different players for assuring data quality for data in repositories (sesink, et al 2008). intriguingly, they include users as part of the equation. our experience so far shows that even where academics are not interested in sharing their data publicly, they do recognise the importance of data management and are interested in the possibility of getting institutional support for improving current practice. the project has benefited from the input of its consultant, digital life cycle research & consulting, with regard to the need for a life-cycle approach to data curation and the partnerships that could be forged throughout that life cycle (green and gutmann, 2007). one of the findings from such a perspective is that librarians – in their roles as either data librarians or repository managers – could find ways to move “upstream” in the research process, i.e. get involved in the pre-publishing stages where data is created and processed, rather than the usual librarian’s comfort zone of dealing with post-published materials “downstream” (gold, 2007). (see figure 2 below) identifying and describing the data management requirements of digital collections is a central part of understanding what roles and services are required for research data management. one of the additional deliverables taken on by the three datashare partners for the second part of the project is to conduct data audits in partnership with academic departments using the tools and methodology developed by the digital curation centre as part of the data audit framework6 development project. jisc funded the project in response to one of the many figure 2: partnerships in the data & research lifecycle (courtesy digital lifecycle research & consulting) iassist quarterly fall & winter 2007 25 recommendations in the dealing with data report: a framework must be conceived to enable all universities and colleges to carry out an audit of departmental data collections, awareness, policies and practice for data curation and preservation. (lyon, 2007). open data and open licenses open access repositories allow any user to access their content via the www as well as allowing other servers to access and harvest their metadata, e.g. google or scholarly search engines. the following definition is from the budapest open access initiative (2002): by 'open access' to this [scientific and scholarly journal] literature, we mean its free availability on the public internet, permitting any users to read, download, copy, distribute, print, search, or link to the full texts of these articles, crawl them for indexing, pass them as data to software, or use them for any other lawful purpose, without financial, legal, or technical barriers other than those inseparable from gaining access to the internet itself. the more recent open knowledge foundation definition equates open access with open content, rather than just research literature: “a piece of knowledge is open if you are free to use, reuse, and redistribute it”7 (similar to the open source movement for software). this is not to be confused with the ‘open data definition’ used for copying personal data from one social networking site to another.8 as peter murray-rust has stated, “where the open access movement is concerned only with ensuring that scholarly papers are human readable, the open data movement requires that they are also machine readable.” (poynder, 2008) in other words, for data to be useful, they need not only be read, but manipulated, re-used, re-coded, mined, and merged (or mashed, if you will) with other data. for a chemist like murray-rust who analyses chemical structures by mining chemical literature for tables, charts, images containing data about molecular structures, not only is a pdf document not sufficient to re-use the data embedded within, but the restrictions on the re-use of the content by publishers and even supposedly open access repositories are too stringent. for this reason, and to generally encourage the proliferations of mashups on the web, the open data community developed an “open data license” for data publishers (i.e. those responsible for making their data available) to set their data free. first, science commons, a project under the banner of creative commons that deals with licensing copyrighted materials, undertook the development of a protocol upon which any open data license would be based: science commons’ protocol for implementing open data9 1. the protocol must promote legal predictability and certainty. 2. the protocol must be easy to use and understand. 3. the protocol must impose the lowest possible transaction costs on users. following this, a number of players including law scholars at the university of edinburgh were responsible for bringing about the public domain dedication and license, which attempts to include wording that either waives ipr altogether or in cases where that is not legally possible, dedicates it to the public domain. key concepts covered by the pddl are: • ‘converge on the public domain’ by waiving all rights based on intellectual property • take into account “sui generis” database right (in european jurisdictions, e.g. database directive rights) • avoid attribution stacking (as the “attribution” norm is a burden when merging from highly numerous sources of data) the edinburgh datashare repository offers an option to attach a pddl to deposited datasets. where depositors do not wish to freely give away their data, or are prevented from doing so based on agreements with subjects, funders, or research ethics boards, they may fill out a rights statement field spelling out the terms of use, or fill out a metadata-only record where potential users are free to contact the depositor to request access. conclusion data management, curation and publishing are getting much attention globally and in the uk at present. there are a number of nagging problems that have not been solved over the years, such as the lack of career reward for publishing data as opposed to papers and the lack of career paths for ‘data scientists’ (national science board, 2005; swan and sheridan, 2008b). related to this is the poor practice of data citation, especially where data are shared only informally between peers or downloaded from a website rather than obtained from an archive or repository with a complete metadata record. the scattered infrastructure for data curation across disciplines and institutions is being addressed in australia through the ands national data service10 and a feasibility study for a more modest shared research data service in the uk11 . canada has stepped forward with the research data strategy working group to address the challenges surrounding the access and preservation of research data12 . two american universities, cornell and mit, are pursuing ground-breaking, library-led data curation services for their users. these are described in sister articles in this issue of iq. perhaps the current critical mass of attention and 26 iassist quarterly fall & winter 2007 effort will help break through the remaining barriers that prevent data from being cared for, shared, used to develop new knowledge, and preserved as an essential part of the scholarly record. it has been an exciting time to be involved in a project such as datashare. we are tracking as many of these developments as we can keep up with on our website. we welcome any and all feedback. * contact: robin rice, edina and data library, university of edinburgh, uk. e-mail: r.rice@ed.ac.uk references (2001). budapest open access initiative, february 14, 2002, budapest: open society institute. http://www.soros. org/openaccess/read.shtml carrier, s., j. dube and j. greenberg (2007). the driade project: phased application profile development in support of open science. in proc. int’l conf. on dublin core and metadata applications 2007. dcmi. http://www.dcmipubs. org/ojs/index.php/pubs/article/view/39/19 the center for research libraries and oclc (2007). trustworthy repositories audit & certification: criteria and checklist. version 1.0. february 2007. oclc.http:// www.crl.edu/pdf/trac.pdf the digital archiving consultancy, the bioinformatics research centre (university of glasgow) and the national e-science centre (nesc) (2005). large-scale data sharing in the life sciences: data standards, incentives, barriers and funding models (the joint data standards study). http://www.mrc.ac.uk/utilities/documentrecord/index. htm?d=mrc002552 gibbs, h. (2007) disc-uk datashare: state-of-the-art review. disc-uk, august 2007. http://www.disc-uk.org/ docs/state-of-the-art-review.pdf gold, a. (2007). cyberinfrastructure, data, and libraries, part 2. libraries and the data challenge: roles and actions for libraries, d-lib magazine 13(9/10). http://www.dlib. org/dlib/september07/gold/09gold-pt2.html green a. and gutmann, m. p. (2007) building partnerships among social science researchers, institution-based repositories and domain specific data archives. oclc systems and services, 23 (1), 35-53. http://deepblue.lib. umich.edu/handle/2027.42/41214 [open access version] heery, r. and anderson, s. (2005). digital repositories review. http://www.jisc.ac.uk/uploaded_documents/digitalrepositories-review-2005.pdf heery, r. and powell, a. (2006). digital repositories roadmap: looking forward. bath: ukoln/eduserv. http://www.ukoln.ac.uk/repositories/publications/ roadmap-200604/ humphrey, c.k., estabrooks, c.a., norris, j.r., smith, j.e. and k.l. hesketh (2000). archivist on board: contributions to the research team. forum qualitative sozialforschung / forum: qualitative social research 1(3). http://qualitativeresearch.net/fqs/fqs-eng.htm lyon l. (2007) dealing with data: roles, responsibilities and relationships, consultancy report. june, 2007, bath: ukoln. http://www.jisc.ac.uk/media/documents/ programmes/digitalrepositories/dealing_with_data_reportfinal.pdf macdonald, s. (2008a). data visualisation tools: part 1 numeric data in a web 2.0 environment. disc-uk, january 2008. http://www.disc-uk.org/docs/numeric_data_ mashup.pdf macdonald, s. (2008b). data visualisation tools: part 2 spatial data in a web 2.0 environment and beyond. disc-uk, september 2008. http://www.disc-uk.org/docs/ spatial_data_mashup_v2.pdf martinez, l. (2008). the data documentation initiative (ddi) and institutional repositories. disc-uk, february 2008. http://www.disc-uk.org/docs/ddi_and_irs.pdf national science board (2005) long-lived digital data collections: enabling research and education in the 21st century. washington, dc: national science foundation. http://www.nsf.gov/pubs/2005/nsb0540/ poynder, r (2008). the open access interviews: peter murray-rust. open and shut [weblog]. january 21, 2008, http://poynder.blogspot.com/2008/01/open-accessinterviews-peter-murray.html rice, r., macdonald, s. and g. hamilton (2008). applying dc to institutional data repositories. international conference on dublin core and metadata applications (dc-2008,) 23-25 september, 2008, berlin. http://dc2008. de/wp-content/uploads/2008/10/12_rice_poster.pdf (2002). reference model for an open archival information system (oais). consultative committee for space data systems. january 2002. research information network. (2008) stewardship of digital research data: a framework of principles and guidelines. january 2008, london: rin. sesink, l., van horik, r. and h. harmsen (2008) data iassist quarterly fall & winter 2007 27 seal of approval: quality guidelines for digital research data in the netherlands. may, 2008, the hague: data archiving and networked services (dans). http://www. datasealofapproval.org seymour, r (2007). user based evidence for the requirements and functionality of a repository capable of managing licensed geospatial assets (public version), april, 2007, edinburgh: edina. http://edina.ac.uk/projects/ grade/formalgisrepositoryfeedbackfinal.pdf swan, a. and s. brown (2008a). to share or not to share: publication and quality assurance of research data outputs. june, 2008, london: research information network. http://www.rin.ac.uk/data-publication swan, a. and s. brown (2008b). the skills, role and career structure of data scientists: an assessment of current practice and future needs. july, 2008, london: joint information systems committee. http://www.jisc. ac.uk/publications/publications/dataskillscareersfinalreport. aspx footnotes 1.http://www.disc-uk.org/datashare.html 2.http://ils.unc.edu/mrc/dryad 3. http://edina.ac.uk/projects/gap_summary.html 4. this will be available on the project deliverables page http://www.disc-uk.org/deliverables.html before march, 2009. 5. http://www.opendoar.org/tools/en/policies.php. 6. http://www.data-audit.eu/ 7. http://www.opendefinition.org/ 8. http://www.opendd.net/about.php 9. http://sciencecommons.org/projects/publishing/ open-access-data-protocol/ 10 http://ands.org.au/ 11. http://www.ukrds.ac.uk/ 12. http://cisti-icist.nrc-cnrc.gc.ca/media/press/rds_ group_e.html vol23/2 4 iassist quarterly abstract gaining access to information on the health sector in bangladesh, and in many other developing countries, can sometimes be very hard. although a considerable amount of data is collected by government departments, non-governmental organisations (ngos) and other agencies, it is not always easy to find out what information has been collected or to gain access to this information. these difficulties can reduce the potential value of the information, slow the decision-making and planning process or cause it to be based on less reliable information. with the current trend towards involving all stakeholders, in developing countries, in a health sector wide approach to policy-making, planning and programme implementation, the need for coordination in information gathering and access is greater than ever. the health economics unit, of the ministry of health and family welfare, has initiated the development of a health economics data archive (heda) for bangladesh, which aims to address the problems of access to information for policy-makers, planners, researchers and others involved in the health sector. amongst the aims of the project are: providing a tool for dissemination of research results; a standardised approach from which to improve methods of data collection; the development of a health data dictionary for bangladesh; encouraging data security; and fostering a culture of information sharing. use of the archive can also prevent duplication of research activities and encourage improved or standardised methodologies the needs and suggestions of the potential users and holders of an archive were obtained through a process of workshops, seminars and consultation. the archive itself was then started as a small entity holding the databases, and supporting documentation, for health economics unit studies. a user-friendly front-end screen was designed in access 97 software, enabling searches by subject area, key word, geographical area and free text to identify databases held on the archive. at present, it is possible to hold and use the archive on a standard pc computer using microsoft office 97 software, thus requiring no extra capital investment in the initial development period. the creation of an operational archive in a short space of time and at minimal cost has allowed potential users to see the immense benefits of such a tool. the flexibility of the archive design will allow it to expand to meet the demands of more databases and users with few technical problems. the next steps will see wider dissemination so that more databases related to the health sector will be entered on the archive, and users will expand to a wider audience in the gob, donors, ngo’s and research institutions. the process of institutionalisation and mechanisms for cost-recovery are now being addressed, to ensure the maintenance and sustainability of the archive. background there are a large number of organisations working in the health sector within bangladesh, including aid organisations, government of bangladesh (gob) and a variety of non-governmental organisations (ngos). many of these organisations are involved in data collection and all of them need relevant and up to date information to carry out their work. however, although a considerable amount of data has been, and continues to be, collected, it is not always easy to find out what information has been collected or to gain access to this information. information can be over-protected, located in numerous sites and difficult to track down. these problems are exacerbated by limited computing facilities. problems in data access, and lack of information exchange and co-ordination between organisations carrying out research, often lead to a duplication of data collection efforts and can limit the opportunities for improving the process of information gathering through collaboration and dialogue. these difficulties can slow the decision-making and planning process or cause it to be based on less reliable information. however in bangladesh, as in some other developing countries, there is a current trend to move towards involving all stakeholders in a health sector wide approach to policy-making, planning and programme implementation. this means that the need for coordination in information gathering and access is therefore development of a health data archive for bangladesh an example of a cost effective and sustainable approach to information sharing in a developing country. by deana leadbeter & lorna guinness* summer 1999 5 greater than ever. in the united states and europe overcoming these information problems is assisted with the use of a depository of information, called a data archive. this is a database that holds metadata i.e. information about the data that is held on other databases, as well as holding the actual data for some of these databases. the wealth of data concerning the health sector in bangladesh continues to grow. although these data are of potentially no less value than that in western archives, central archiving has not been a common practice and so the location of these data remains dispersed and difficult to access. in a series of workshops, held by the health economics unit of the ministry of health and family welfare, bangladesh, in 1996, a serious problem in awareness of past and present research activities and also of the location of key health sector databases was identified. this situation is of concern due to increased costs of accessing information, duplication of research efforts and limited opportunities for improving information collection process. in addition, if there is a lack of knowledge concerning data or difficulties in accessing data, the potential benefits are not fully realised. the data are used only for their primary purpose and then either discarded, or stored but not re-used. since data collection is usually very costly, if the data can be used for other analytical work (secondary data analysis), or if they can be used to inform the design of future studies or routine data collection exercises, then considerable savings and additional benefits could occur. in bangladesh, this potential is currently not being fully realised which results in a resource waste, that resourcepoor bangladesh can ill-afford. to address this problem and facilitate collaboration and information sharing amongst researchers and stakeholders, it was suggested, at the heu training workshops, that bangladesh should begin to develop a central depository for health economics relevant data, in the form of a database archive. the health economics data archive (heda) was therefore proposed to bring these data together to a central location, providing the similar functions to database archives operating in the u.s. and europe, thus allowing for data to be re-used and wider dissemination of their key findings. in other words, adding value to the primary research carried out. this paper describes the process of the development of heda in bangladesh and the particular method used, which was low cost and user-friendly. thus, it suggests a possible model for other resource-poor nations, where the full value of the wealth of primary data and generated research information may not be fully realised at the moment. aims of the health economics data archive (heda) in the initial phases of development, the establishment of heda for bangladesh had two primary aims: 1. improving accessibility to data by documenting health sector databases and using a standardised approach to documentation, as well as providing electronic searching facilities, heda provides easy access to data. 2. dissemination of health economics research findings an archive makes the work of any research activity or organisation available to a wide audience in more detail than is possible through the publication of research papers or other reports. in the case of the health economics unit, it was felt that heda would widen access to the primary and secondary research findings, within and outside the ministry of health and family welfare. in addition to the primary aims of heda, there are several important additional benefits that arise from the process of developing heda that were considered as critical outputs of the project: 3. improving study designs and methods of data collection before a study can be entered on an archive, a study description form has to be completed. the completion of this form requires the lead investigator to describe the study design in a clear and consistent manner. experience has shown that this not only provides valuable information for anyone wishing to carry out secondary analyses on the study data, but it is also a useful checklist of the issues that need to be considered when designing a study. using the study description form can therefore serve as an ongoing training exercise in study design for all staff involved in the process. completion of the form at the start of each study, rather than when the database is complete and ready to be entered on to an archive, is helpful in ensuring high quality study designs. having a clear and well-documented design is also likely to be useful when collaborators in several different organisations are involved in a study. 4. development of a data dictionary the study description also requires the studies to have clear definitions for all data items for the study, which are easily available to anyone wishing to access the data. a data dictionary should therefore be set up, within or in parallel to the archive, which includes 6 iassist quarterly data definitions across all studies in the archive. this helps to identify where data from different studies can be combined or compared because the same definitions have been used and also where different definitions have been used for similar data items, and hence direct comparisons are not valid. as with the study descriptions, if this process of documenting data definitions in full is carried out at the start of the study it will lead to improved study procedures, particularly in the area of data collection and analysis. setting up this data dictionary helps users, who wish to carry out secondary analyses on the data or to combine data from different studies. it is also a valuable resource when designing future studies, particularly where these follow on from, or need to be compared with, the results of other studies. further, it can enable improved data quality by allowing cross validation between databases. 5. fostering data security another challenge, when storing data, is that of data security both in terms of not allowing unauthorised access or inappropriate use and in terms of ensuring that the data are maintained in good condition. data can be lost at the flick of a switch, or may get corrupted because of problems with the power supply or physical environment, and databases can be manipulated without permission. the updating, maintenance, back-up and security problems usually faced with storing data can be placed under the responsibility of those responsible for the archive, thus saving time and money for the original data producers. 6. enabling bibliographic citations of databases by establishing databases as bibliographic entities and “publishing” them as such, as well as offering advice on citation, archives play a major role in extending research and scholarship, giving recognition and acknowledgement in the same way as printed piece of research work. 7. fostering information sharing experience in other countries has shown that initiatives for information sharing can lead to greater summer 1999 7 understanding and collaboration between organisations across all their activities not just those related to data collection and analysis. it also facilitates more comprehensive data analyses by linking data collected by different departments or agencies. this is of value at any stage of health sector development but is particularly relevant in bangladesh or those countries introducing a sector wide approach, which requires a greater co-ordination within the ministry of health and family welfare and between all those involved in the programme, including the donors. how heda can add value the value of a resource such as heda can be demonstrated by two examples. in the first example, the secretary of the ministry of health and family welfare may make an ad hoc request to his assistant for information on the current level of household expenditures on health as opposed to government expenditures. what does the assistant do? the required processes of data collection and analysis are shown in figure 1, both with a data archive available and without a data archive. a second example of the value of the archive can be shown by the steps taken by the ministry of health and family welfare, bangladesh, to develop a new approach to revenue generation in the health services. the government officer designated to assist in the gathering of background information will require data on the income levels, health expenditures and health seeking behaviour of the population, other health services provided by ngos and the private sector and methods and current levels of revenue generation. how does the government officer do this? the officer may have to locate and approach a number of different sources: • the bangladesh bureau of statistics (bbs) for household incomes and levels of health expenditures; • find and consult or carry out a surveys on healthseeking behaviour, a survey of ngos active in the health field, a survey of private clinical services; • locate and consult mohfw financial information on current levels of revenue generation. • health care facilities and mohfw for unit costs of health services a phone call or visit to a central archive would establish whether this information was available and, if so, could provide the officer with the necessary information, saving both time and money for the officer. an archive also prevents duplication of research activities and encourages improved or standardised methodologies. under the second scenario, the government officer could have consulted the archive to discover that a survey of private clinics had been completed and therefore the planned private clinic survey was not necessary. alternatively, it could be that a survey of private clinics had been completed but was out of date. using the instruments and results of the old survey available on the archive, the officer could make improvements based on the problems encountered in the first survey and collect information in a standardised way to create a time series. development of the health economics data archive identification of need and appropriateness in june 1996 a series of workshops, meetings and training sessions regarding health databases were organised by the health economics unit (heu). the purpose of these sessions was: • to promote greater knowledge of existing databases in bangladesh, • to help create a programme of co-operation among different data providers and data users and • to agree a way forward for developing a metadatabase or data archive for health data in bangladesh. the participants in these workshops were asked the following specific questions: 1. what and where are the existing databases in bangladesh? 2. what are the means of access to these databases? 3. how, and how often, are the data collected/ updated for each of the databases? 4. how can the needs of consumers be fed into future database design and management? 5. what areas of mutual collaboration could be pursued, and how can this collaboration best be carried out? the training sessions that were linked to these workshops were designed to orient non-specialists in the utility of databases and to explain best practice, associated with their design, handling and use. the training was aimed at mid-level managers in government and those with a specific interest in databases and the use of information it was proposed at the workshops that a database could be set up to hold information about data already collected which could potentially be of use in health economics and other health related studies. this would function as a data archive holding data as well as ‘metadata’ concerning a particular study. holding ‘metadata’, rather than the data itself, allows the data holder the option of retaining control over the specific purposes for which the data can be released, where the data held are confidential or sensitive. 8 iassist quarterly it was expected that both the workshop and the training sessions would help to assess the feasibility of setting up an archive, and to identify the data items that could be included in such an archive. it was planned that participants of the workshops and training sessions could pilot the collection of these data items. a programme for setting up a health economics data archive (heda) for bangladesh could then be drawn up, taking into account the information provided and views expressed at the workshop and training sessions and also the experience with the proposed pilot. since the participants at the workshops and the training sessions were mainly from government departments, individual meetings were also held with other organisations to discuss the heda project and obtain their views. once the pilots were completed and development was underway, potential users and heu staff was consulted about the design of heda and the methods of access. based on this a specification was drawn up for the technical support required and the need for both a computer programmer and a database manager with extensive experience in the health sector to join the heda was identified. in consultation with the heda technical team and potential users, the structure of the front-end screens and the search criteria were agreed. the design of heda was tested with examples of user queries, given by potential users of heda identified by the heu, then considering how the use of the heda might assist in answering these queries. this ensured that the development of the heda was following a model appropriate both to the future users and holders of the data archive. issues that were considered in the design and development of heda since the concept of a data archive was new to many people within the health sector in bangladesh, some basic principles were agreed at the outset. these principles are outlined below. the term ‘archive’ can be used for a repository or store of any material, although it is most commonly used for a store of information. the purpose of keeping anything in store is so that it is available for use when required. there is no point in keeping anything in any type of store unless: • you know it is there • you can get access to it when you need to use it • it is kept in a good state an important feature of a data archive is therefore to facilitate use of the information as well as holding the data. this means that a data archive should provide: • a secure place to hold data, and also to hold information about the data or information derived from the data • information on what is held on an archive • mechanisms to find the information or data when needed • mechanisms to access the information or data when needed this requires, when archiving data, that sufficient and accurate data documentation is provided both on the data background, including sampling methods, sources of data, investigators etc, and the data characteristics, including data definitions, data relationships and coding systems. this is discussed further in the section on database documentation below. in addition to these guiding principles, there were some key criteria, which were agreed on for heda in order to address the two common causes of developmental failure: • when projects are over-ambitious so that they often fail to deliver within agreed deadlines • when users, whose future participation is essential, do not see any benefits for themselves after the initial promises, and so they lose confidence and interest in the project the key criteria were as follows: 1. timeframe the initial stage of the project should be designed so that it is manageable within a short timeframe and produces a product that can be demonstrated to users within that timeframe. this means that the first phase should cover only a limited set of databases. priority for inclusion in this initial set should be given to those databases that are already well structured and documented so they can be brought in to the archive with a minimum of effort in order to demonstrate benefits within the short timeframe. 2. maintenance of user interest to maintain user interest the databases selected for inclusion at the initial phase should be those that are considered to be of most interest to potential users 3. limited technical resource requirements although specialist technical skills are needed for the initial development, the data archive should be designed so that it can be easily updated and developed by an in house team after the initial phase has been completed. 4. ease of access the archive should be accessible on a user-friendly front-end screen that requires minimal training and should be located in a central location with easy summer 1999 9 physical access. documentation of databases the concept of two types of data documentation were agreed – macro or data background and micro or data characteristics. the data background covers the context within which the data were collected and issues relating to how the information can be used, including: • supplier and user documentation • original forms and instructions • reports on data collection and usage • original output • minutes of meetings or policy documents • levant to the collection and use of the data • information on data quality and usefulness the data characteristics cover details about the data items and how they are held, including: • data types, field descriptions, data ranges etc. • data relationships • coding schemes • missing values • system information without this level of documentation it would be impossible to achieve many of the aims of an archive, including the basic premise of using the databases. a standardised tool for the documentation of each database to be held on any archive is required to facilitate this process. technical requirements in the initial stages of development, support from a computer programmer is essential. however, the archive should be designed such that once the software has been developed, it can be easily maintained and updated and archive queries easily answered by in-house staff with a minimum of computer skills. in the long run, as the archive expands it would be expected that a health information specialist will be required to maintain and update the database as a permanent member of staff. this is discussed further in the section on future directions. clearly staged development process experience with projects involving the development of software has shown that a critical factor in the long-term success of such projects is ensuring that the software is available at the same time as users are made aware of its potential uses. raising expectations before the product is available can be counter-productive. in addition, starting small but with flexibility can encourage a stronger foundation and demand for the project and allows the project to grow with the demands placed upon it. for these reasons, a clearly staged development process is necessary. design of heda the heda was designed to provide users with access to: • data documentation for establishing the history of data collection and analysis process for any data held on heda • raw data • a data dictionary, and • some key results from analyses of the studies, where available. it is also planned to distribute lists of the databases, key findings of newly acquired databases and developmental news of heda, both on a regular basis to regular users and on request. this will be provided on either floppy disks (probably containing excel tables), or hard copy. heu will also provide additional results tables in response to ad hoc requests from users. following discussions, it was decided that the first set of databases to be included should be drawn from studies carried out by heu. the rationale for this was that these databases were likely to be more easily and quickly accessible to the heu staff involved in the development, and heu staff would be more familiar with the data structures and the results. this approach would allow heu staff to test out the procedures for documenting the databases, using their own data. any problems could then be identified and rectified before other organisations were asked to complete the documentation. another advantage of this approach was that potential contributors would be able to see heda in operation before being asked to complete the documentation for their studies. this should help them to understand more clearly some of the requirements of the documentation, and also to see the value of contributing information to heda. at present heda itself contains the databases, tables of key results from those heu databases and tabulations created in analysis of secondary sources of data. it is planned that, at later stages, contents will include databases containing health financing and expenditure data (e.g. national health accounts) and socio-economic information (e.g. from the bangladesh bureau of statistics), as well as information from research and ngo projects. heda also contains documentary information about each database, and this is described in more detail below. it is expected that some of this information will be of interest in itself e.g. study design, as well as providing background information to aid the selection and use of data within heda. the design of heda includes the provision of facilities to search not only using pre-specified lists, but also using the keywords in the study design and, if 10 iassist quarterly necessary, in the text within the documentary information, including the data dictionary. documentary information for each database as stated above heda contains documentary information for each database held on heda. in order to record this documentary information and create the searching facilities within heda, it was necessary to use a standardised questionnaire. at the training sessions, in 1996, a questionnaire was presented, covering the metadata collected by the uk’s national economic and social research council (esrc) database. the data items on this questionnaire were discussed and participants agreed that, with only a couple of exceptions, which could easily be amended, all the questions were suitable for use in bangladesh. all the participants agreed that this information could be collected about their databases for inclusion in heda. the esrc data archive was approached to check that they had no objections to the use of their form, and to obtain advice on the use of the data documentation. the response from the esrc was very positive. they were happy for their form to be used and made some useful comments in relation to the proposed development, and thanks are due to them for their help and encouragement throughout this project. thus, for each database the following information should be included in heda: a) study description this is based on the study description questionnaire used by the esrc data archive in the uk, amended to suit local circumstances. this includes summary information such as study name and topic areas, as well as more details on the study design. these details will need to be known by any future user of the study data, as well as being used for searching for studies within heda which cover the user’s particular area of interest. b) data dictionary for each data item (or each variable, in statistical terminology) a set of information is required. this is specified at the end of the study description questionnaire. for each data item included in a database the following are needed: i) data item identifier (probably a summary name) ii) full name of data item iii) description of data item iv) data type v) field size vi) coding system (if used) vii) any particular comment about the data item e.g. parts of the list of responses and codes may only be relevant in specific organisations. viii) whether this is being used as a proxy for another data item other information that is expected to be included, either within the data description or within the comments, relate to whether it is a computed data item and, if it is, how it was calculated, whether it is raw (primary) data or derived data, and any relationships between data items. c) survey form where data were collected by survey, a copy of the questionnaire form used on the survey will be available. it is possible that, for some databases, some of the information required for completion of the questionnaire may not be immediately available. therefore, to collect this information, it was suggested that the questionnaire forms should first be sent out to the participating organisations, and then a visit should be arranged to review the forms completed and clarify any points of confusion. a specific appointment would be made for this visit to ensure that the person with the knowledge of the database is available to answer any queries. entry of data documentation user-friendly electronic data entry forms were created within the heda software, for entry of data background and data characteristic information on the archive. accessibility it was agreed that heda should be easily accessible and comprehensible to a range of different users. this requires the use of software that is readily available and userfriendly. after discussions with it experts familiar with database programmes, it was decided that heda should be developed using access 97 which comprises both a programming language for development purposes and user friendly facilities for use by non specialists. the design of the front-end screens and the search facilities has made full use of the facilities already available within the access software. this has allowed for rapid, flexible and cost effective development of the system. the approach has been to provide a mixture of using userfriendly menus and ‘buttons’ for selecting the chosen options for searching and/or viewing the contents of heda, as well as using standard access or windows features. all the features used will already be familiar to archive users who use other windows based software, and an instruction sheet will be provided for those unfamiliar summer 1999 11 with windows. in these initial stages, heda is stored on one central computer where it can be accessed both by heu staff and by other users. two options were considered for the final location: remaining within the offices of the heu and moving to the national resource centre for health economics. this is discussed further later in this paper. once a run time version of heda is available it should be possible to hold a version at both sites, providing easy access for both gob officials and external researchers. search facilities heda design includes the provision of a user-friendly interface. this ‘front–end’ should help direct the user to the most useful databases according to his or her work. this is through clear subject categorisation, plus an index and explanation of the databases contained within. in developing search procedures a compromise had be reached between the speed and efficiency of the search and the amount of freedom the searcher is allowed in specifying the search criteria. if the search is restricted to using preselected terms then the database can be set up to allow this to be carried out quickly and easily. the drawback of this approach is that the terms selected by those setting up the database may not cover the specific interests of the full range of potential users, or be suitable for categorising additional study databases when these are added to heda. an alternative approach is a ‘free text’ search, which allows the user to enter any word, or combinations of words, and search for a mention of these in the study documentation. this gives the user freedom to pick terms that reflect their area of interest. however searching through the full documentation for all studies can be slow, particularly as more studies are included in heda. also there may be different ways of describing a particular topic. if the terms chosen to describe the topic of interest, for the purposes of the free text search, are not those used within the study description then the search will not pick up this study. the approach taken to deal with these issues was to provide a mixture of search options. the front-end system to the data archive provides for several different searches these use: (i) study name where this is already known. (ii) pre-selected topic areas (iii) geographical areas (iv) keywords or key topics within the study description (v) free text within supporting documentation e.g. any mention of immunisation for example, using the pre-selected topic areas (option ii above) provides a quick route to finding studies covering general areas. within the study description there is also the opportunity for the investigator to specify key topics and keywords that describe the areas covered by the study. these are then available for the user of heda to use in their searches (option iv above). the user can either type in a topic or keyword describing an area in which they are interested, or can select from a ‘pick list’ which gives all the topics or keywords recorded in the study descriptions held on heda. this is a slightly slower search method than using the pre-selected topic areas, but is much quicker than free text searching, and still allows considerable flexibility and specificity in the search criteria. also the ‘pick lists’ of topics and keywords can be automatically updated every time a new study is entered on to heda. for very specific queries these search methods may not be sufficient and so the user would then need to use the free text search facilities (option v above). this searches for the occurrence of a specified word, or combination of words, within the study description and also within the data definitions. the study description questionnaire asks whether the study is national or district specific, and asks for the districts covered by the study. this allows the user of heda to search for studies relating to a specific district within bangladesh. in this case, instead of selecting from a pick list of the districts, the selection is made using an annotated map of bangladesh it is expected that, as well as using the search facilities to find studies satisfying a specifically defined search criteria, users will find it helpful to use the front-end system for browsing the information on heda. at any stage the user can browse through the information relating to all studies, or to a selected set of studies, to learn more about their design, and to view some of the results of analyses on the study database. browsing through some of the lists created from the study descriptions, such as the lists of key topics and key words that have been recorded, or browsing through the data definitions will also be helpful in identifying studies of interest. queries because of the different needs and technical capabilities of users, a flexible approach is needed in terms of methods of access. this includes the need for a user friendly front-end to give direct access to heda, and the issuing of both short bulletins in hard copy and copies of key results on floppy disk. a query service will also be offered, where heu staff will access heda on behalf of users. it is expected that this query service will be required where key results tables available through the front-end system do not provide sufficient information to answer the queries therefore requiring additional analysis. these key results tables will initially be tables giving the results of the 12 iassist quarterly analyses that were carried out when the study was first analysed. as heda develops, the range of these tables will be increased to cover the more common queries. to run these queries the user will use the searching facilities to identify the appropriate database and extract the data for analysis, to create tables and print or save to floppy disk. it will also be possible to cross-reference the databases and, where compatibility allows, create links. technical support required for the development of heda in phase 1 technical support was required in two main areas: (i) in access97 programming (ii) in system design and standards for data definitions and coding the staff contracted to provide technical support was given the opportunity to discuss, and comment on, the draft specification and programme before this was finalised. this was important since it gave them an understanding of the overall aims of the project, and they were therefore able to participate in ensuring that the work they were carrying out provided the best technical solution to the requirements of the project. the data archive was designed so that it could be easily updated and developed by an in house team after phase 1 had been completed. the technical specialists provided an element of training to the in house staff during phase 1 (mainly through advice and support on tasks that in house staff will be carrying out). this should ensure that routine updating can be carried out by in house staff with additional support being required for ad hoc technical inputs. however, once the scale of the project requires it, a database manager will be required as a permanent member of staff. information requests in the form of queries from outside heu will also require heu staff input, particularly where synthesis or evaluation of the results is needed. security heda has been developed to enable all those who are familiar with windows environments to gain access to databases for downloading and to examine the techniques used in data collection. in order to prevent misuse on heda, security passwords have been built in for different levels of users, and access to original data files for general users will be ‘read only’ which will allow copying but not amendments. to prevent data loss, there is a cd-rom back up system, and back ups will be taken on a regular basis updating and development the effect of taking the in-house, staged development approach was that the initial phase of the project focused mainly on the inclusion, within heda, of databases available within the heu itself. the further development of heda therefore includes adding additional databases as they become available as well as providing enhancements to the front-end screen and the pre-prepared analyses to meet user needs. in the early stages heu staff will be responsible for updating and developing heda to meet these and other user requirements as they arise. these developments will include modifications to the frontend screens and to the pre-prepared analyses available, as well as additions to the databases themselves. it is expected that some assistance will be needed, even in these early stages, from a health information specialist on a part time basis, for example on the development of the data dictionary. it is recognised that, as heda grows and develops, some more dedicated support will be required for updating, maintenance and continued development. staged developmental process it was important to ensure that the timetable for the heda activities were agreed upon before any further approaches were made to the potential contributors and users of the proposed heda. the developmental process was planned in a series of stages, to allow review and dissemination and key points. this was planned to create demand for heda and obtain maximum input from potential users. the expected outputs of the three planned development stages are as follows: phase 1: design and programming of heda • heda, held on a single desktop computer, including database background information for all heu research activities as well as the data characteristics and the database itself for at least three heu studies. • key findings of heu research documented and held in electronic distribution form • heu personnel trained in using heda • full documentation of heda development • draft gob-approved protocol covering rights of access to heu databases • launch seminar for heda for potential users and contributors phase 2: development of heda contents and long term plan • heu personnel trained in updating and developing heda • instruction manual • full set of heu databases on heda • final gob-approved protocol covering access to heu and other gob databases held on heda, or for which study descriptions are held on heda. • agreed programme for inclusion of selected nonheu databases • agreed plan for institutionalisation, cost recovery, summer 1999 13 ongoing maintenance and support, and future developments, of heda • further dissemination seminars phase 3: institutionalisation and implementation • implementation of institutionalisation activities planned in phase 2 • staff recruited to provide ongoing maintenance and support to heda • promotion of heda use for mohfw, ngo, university and other research personnel through training, a newsletter and briefing seminars. analytical tools in addition to the databases available, heda also provides access to a suite of statistical and modelling packages. this allows the user to carry out analyses or modelling using the data they have selected and copied from archive databases. it is expected that this will be of use where the packages provided are not available on the user’s own systems. also, even if a user is intending to take a copy of the data for analysis on their own systems, they may wish to carry out some preliminary investigations using the software available on heda. this may be helpful in case the results of these preliminary analyses suggest that some amendments may be required to the data selected, for examples including some additional data items. hardware and software requirements during the process of specifying the database structure and the costing of the development, it was suggested that the esrc data archive should be approached to see if the software used in running their metadatabase could possibly be transferred for use in bangladesh. this was discussed with the esrc archive and it was found that their software was not suitable for transferring, although it appeared that the structure proposed could be set up fairly quickly in microsoft access 97. the advantages of access are that it is a user friendly package and it is easy to transfer extracts from an archive written in access into other windows based packages e.g. for inclusion in reports written in word. it is also easy to view tables previously prepared in word or excel from such an archive. in addition, ms access is available as part of ms office97 packages and already available at the heu. as a result, in the first phase, no software or hardware upgrading was necessary, thus keeping initial development costs to a minimum. progress to date the programming of the front-end screens and the search facilities were completed in 1998, within three months of the start date for development. initial testing was carried out and any necessary amendments made to the software. testing of all the data and facilities continues and will be an ongoing process. documentation of the software is available within the programme but an information sheet for users has yet to be developed. the pilot study, subsequent to the 1996 heu workshops, found that the esrc documentation form was, subject to minor amendments, suitable for use in bangladesh. data international, a consultancy group working closely with the heu, was then asked to complete the study description questionnaires for the heu studies, in liaison with the heu staff responsible for the individual studies (usually the principal investigator for the study). this information was entered on to heda using the electronic data entry forms. experience with completing these forms led to some minor amendments being made to the questionnaire, but no major problems were found with the questionnaire content or design. a list of potential keywords was prepared to assist the principal investigators in identifying keywords relevant for their studies, although the principal investigators were free to specify whatever keywords or topics they felt best described their study. the study descriptions were entered on to heda for all completed heu studies, even if the databases and other supporting information were not yet available for the study. in order to prepare the heu databases for inclusion in heda a list of all heu studies was prepared, together with their current location and state of documentation. a list of outputs (tables of results) available for each of these databases was also prepared. these outputs will be made available to users via heda, and will also be available on request on either floppy disk or hard copy the demonstration of heda at a formal launch seminar at the end of the three month development period in 1998 (phase 1), and in a series of individual demonstrations following the launch, showed that heda was already a product that is both easily accessed and useful. subsequent activities have focused on adding more databases to heda, and on plans for institutionalisation and appointment of staff which are discussed later (phases 2 and 3). . lessons learnt data documentation apart from the difficulties in actual physical access to information, the major problem in sharing data is that of complete and comprehensible data documentation. this is usually one of the greatest challenges in developing any archive and the development of the heu data archive has been no exception. as well as documentation for users of the output tables, documentation is also needed for those wishing to use the databases held on heda for secondary data analysis. in order to carry out an analysis on any database it must be clear exactly what is the meaning of the terms that are used in the study. often there is a lack of common 14 iassist quarterly understanding on data items. for example, if one wants to talk about bed capacity within a hospital, what is a bed? alternative views of the definition of a bed could be: • a fully functional bed in a hospital • a space in a hospital that is available for a bed or mattress in a hospital • a broken bed lying in the hospital storeroom or, to what does revenue refer? • are we talking about the revenue allocations of the gob? • are we talking about the revenue budget of the ministry of health? • are we talking about revenue collected from user fees? in addition, data items are often used as a proxy for another data item, which can cause confusion if this is not clearly specified. for example, allocations are sometimes used as a proxy for expenditure, or utilisation as a proxy for demand. ideally, these data definitions should be decided and documented before the study can be carried out and possible proxy data items identified, but as with the other supporting documentation that has been discussed, this is not always the case. retrospectively completing documentation involves interviews with principal investigators and examination of survey questionnaires and codebooks. this is a time consuming task but the benefits are multiple, leading to the ability to re-use data and process or learning of the investigators contacted, therefore resulting in value added on research and improvements in methods of work. the work so far on data definitions has focused mainly on clarifying and documenting data definitions within individual studies. however, if the data are to be linked or compared across studies, a common data dictionary is needed which covers the data across all the studies and which uses a common name for data items that are used in more than one study. at the moment the data definitions are held in one data dictionary, but no work has been done to identify common or proxy data items. this requires additional work to review those data items that appear to be similar, and to identify whether the same data definition has in fact been used and whether the data items can therefore be linked. once this review has been carried out then a linking table can be set up which includes the study data item name and the data item name from the common data dictionary. this can be used to list all data items within a study, or to look at all studies that include a particular data item in the data dictionary. having a common data dictionary is an essential part of heda, so the steps that need to be taken to achieve this will need to be agreed. as with other parts of heda development, the technical task of merging the individual dictionaries into one is likely to be more straightforward than the non technical issues i.e. the task of checking across studies for consistency of definitions and clarifying the situation where different definitions have been used. search facilities during phase 1 a question was raised as to whether the search criteria should link only to studies, or whether it should be possible to identify individual output tables within a study. it was decided that, as far as the documentation and search procedures are concerned, the study description (in particular the keywords and topics), should give sufficient indication of the areas addressed by a particular study and its related tables. there should therefore be no need to have additional search criteria linked to individual tables within the file of output tables. also any derived data items in the output tables should be in the data dictionary, and so searching on a particular data item will identify the studies (although not the individual output tables) in which it has been used. once a study has been selected using the various search criteria, then the user can scan, ‘by eye’, the supporting documentation, including the list of descriptions of the output tables. it is expected that this will, in most cases, be sufficient for the user to select the tables of interest. if heda grows and develops then it may be possible to consider introducing some more sophisticated search procedures although, as has already been mentioned, the introduction of more complex and flexible search procedures can lead to slower, more cumbersome searching. it was therefore felt that, at least in the short to medium term, the best approach would be to keep the search procedures relatively simple by linking the search criteria to studies, and not to individual tables. it is possible that, on reviewing the output tables, it may be felt that a user would not be able to find the table they want. in this case, the table title in the ‘pick list’ could be made a little clearer, and the topics covered by the table could be included in the study description. presentation of key findings and preset queries one issue that needed to be addressed during phase 1 was the format in which the output tables should be held. the outputs held on heda include tables giving the results of analyses already carried out using the study data. there are two methods of holding these outputs. the first option is to hold a specification of the calculations that were carried out, and use this to recalculate the results from the data whenever the tables are required. the other option is to hold the results of these calculations i.e. the actual outputs. this second option saves having to spend time recalculating the results each time they are required, but may need more space to hold the tables. summer 1999 15 for the initial databases entered on to heda the decision was taken to hold the actual outputs, usually in word or excel tables, rather then re-calculate the results each time. in many cases what can be viewed (and copied for further manipulation if required) is just an electronic copy of the tables of results from the published reports. this means that entering the tables of results does not involve any new data entry, just taking a copy of the electronic version of the existing documents and then setting up the necessary references and linkages within heda. one of the strengths of the design of heda is that different approaches can be taken for different studies or for different sets of tables within a study, and so a different approach can be taken for future studies if this is preferred. this includes the option of having hard copies available if the tables of results are not available electronically. in this case asking to view these tables within heda would simply lead to a message indicating where and how the hard copies can be viewed. an important issue also raised was the supporting documentation required for the output tables, including definitions of the derived items. this documentation should be available to anyone wishing to use the output tables. ideally this text should have been prepared at the time the tables were produced. however, if the principal investigator for the study had not prepared their report with a more general readership in mind, then the existing explanatory notes may not be sufficient for the purposes of heda. it was therefore found that some additional work was needed to enhance the existing documentation. the output tables for each study can be held in one file or held in a series of files – one for each table or subject related group of tables. each file containing a set of tables is listed separately in the ‘pick list’ that is viewed when a particular study has been selected. thus, if each table is set in a separate file their identification is more immediate than if all tables for one study are held in a single ms word or excel file. if it was felt necessary for a particular study, every table could be held on a separate file, but this could involve a considerable amount of additional work in setting up these individual files. for each study entered on to heda, the benefits of having more detailed listings of individual output tables will therefore need to be weighed against the extra work involved in setting this up before a decision is made on how the tables should be held. training training for heda users will be carried out through a process of in-house workshops and learning by doing. however, heda can also act as a training tool itself. by providing a series of databases on various health economics related issues, along with the full and detailed documentation of the data and the data collection processes and analytical software packages, it provides a facility for training in: • questionnaire design • sampling methods • statistical analysis • health economics analysis success of following the basic principles the basic criteria followed for the development of heda were to start small, limit the timescale and ensure users were involved and could recognise the need and relevance of the project at all stages. following these criteria meant that in a short space of time and with very limited resources heda was able to stand alone as a useful and technically easily accessible package. the potential users and contributors have shown interest in its further development as they can see results at this early stage and visualise the benefits in the future. continued involvement with these users is essential towards maintaining the momentum already created. future directions expanding the heda user population the driving forces behind the development of heda have been the issues of co-ordination and information sharing. it has started small, but this is not for want of ambition. smallness has given greater flexibility to adapt and make amendments during the development phase. it is hoped that, unlike many initiatives that have started big, the enthusiasm and interest will not wane after the first phase since the project can already demonstrate benefits. there have been many incidental benefits during the development phase, as has been discussed above, but the main benefit is that the heda database is useful as it stands. it already provides access to both modelling and analysis tools and a series of comprehensive databases on health economics in bangladesh, with the room and flexibility for growth and expansion at low cost. in the future, it is expected that there will be further development of the model and a steady increase in the number of databases held on it. the developments will be based on feedback from all potential users, in particular those who attended the launch and other demonstrations of heda. as well as developing the model, the aim is to ensure that heda is used by all those with an interest in information about health and health services whether government, donors, and researchers, and to see the numbers of users growing over the years. further workshops, to demonstrate heda, are planned. these will focus on the donor community, who are expected to be both contributors to, and users of, the information held on heda. in addition to formal group sessions such as this, and informal individual demonstrations during the early stages of the project, it is important that users are kept up to date on progress with 16 iassist quarterly heda. it is therefore planned to issue a newsletter to let interested individuals or organisations know what new features or new data are available on heda. the frequency of this publication will depend on the speed of development of heda but it is expected that it will be issued quarterly. addition of further databases the initial phases of the project within bangladesh involved including only heu databases in heda. however, the way heda has been designed means that it is relatively easy to bring in data from other organisations, and the inclusion of databases from other organisations is now underway. the results of workshops and discussions with those involved in collecting or using health data, have indicated that several organisations would be interested in having their data included in heda. this will improve dissemination of results from many different sources. also the wider the range of data included in heda the more useful it will be in providing answers to users’ ad hoc queries. towards this end, study description questionnaires were made available to all those attending the launch, and are also being sent to other organisations who hold health sector relevant databases and may be interested in contributing data to heda. it is planned that the completed questionnaires will be entered on to heda even if the database itself is not to be held on heda. heda users will then have a reference to sources of information in addition to those held on heda. hardware and software requirements the technical requirements of heda will need to be reviewed in the light of the proposed developments. it is expected that the computer currently being used for heda will need to be upgraded in terms of memory, speed and disk space available as more databases are added to heda. also increased security facilities would be available if the operating system was changed from windows 95 to windows nt. the front end systems which have been written in access7 would not need to be changed since they will run under windows nt and so this change would only involve minor programming amendments. staffing as has already been mentioned, existing heu staff have been responsible for updating and developing heda in the early stages with some clerical and data input support from staff at data international. if heda is to build on its successful initial phase and develop as planned, then some more dedicated support will be required for updating and maintenance, and for liaison and support for user and data producers wishing to enter databases on to heda. it is suggested that this person should be a health information specialist who can carry out analyses to support user queries as well as being responsible for maintaining and developing the database. the recruitment process is currently underway. multi-user networks and dial-up access if the windows nt operating system were used, this would also allow dial up access if it were wished to include this in later developments. alternatively a new microsoft product has just been launched for multiuser access. this may be more appropriate for use with heda since it is designed to need less powerful facilities at the remote sites. it should be noted however that both these options do require reliable and high quality telephone connections. it is also not clear how well remote access, without an heu staff member available to answer queries, will work in practice. it is therefore suggested that dial up access is not considered until heda has been in use for some time and there has been an opportunity to assess the level of support required by users. one way of making information available to users on what is held on heda, and also possibly giving access to some of the data or results tables, is via a web site. one of the benefits of developing an open web site is that the information is then easily available to anyone with access to internet, whether in bangladesh or elsewhere. this could be particularly useful in developing collaborative links between health sector researchers and analysts in bangladesh and those working in other countries, particularly in the asian region. however there are several difficulties with this sort of development. there is no control over who can access the information, and whether they are then using the information appropriately. also specific technical skills are needed both to set up and to maintain the site. it is therefore suggested that, if a web site is to be developed, a staged approach should be taken to this development. the first stage should focus on providing textual information summarising the activities of the heu, and the databases that it has available, plus e-mail contact details. summary tables of results or other relevant study information could then be e-mailed to interested enquirers. e-mailing information on request would be much simpler to do than setting up a front-end system to access information via a web site. it would also allow records to be kept of all those who have received specific study information and provide some control over access. institutionalisation consideration needs to be given to the most suitable place within the organisation for housing and maintaining heda. although the heu has been responsible for setting up heda, and will be supporting it in the short term, this may not be the most suitable option in the longer term. the statutes for an institute of health economics at dhaka university have recently being drawn up, and it is proposed that this institute should house a national resource centre summer 1999 17 for health economics. it is expected that heda will play a central role in the development of any health economics resource centre. it has therefore been agreed that the resource centre that is being set up should house and take on the responsibility of the running of heda in the longer term. plans for setting this up are currently underway. the issue of accessibility by the main users of heda needs to be taken into account when making any recommendations regarding the future siting of heda. since it is not expected that dial up access will be available for some time then ease of physical access will be an important factor. housing heda at the university will make it more accessible to academics, researchers and other interested organisations outside the gob, such as donors. however this option would create barriers to access by gob staff and so may have the effect of reducing the use of heda, and the valuable information held in it, by policy makers on the gob for their decision making. it should be noted however that heda has been designed to run on any reasonably powerful pc, and both the software and the data will be copied onto cd-rom on a regular basis for back up purposes. it is therefore be relatively straightforward to house heda at the university and carry out any updating there, but to have a copy of heda also running at a site within the mohfw. this could be regularly updated via cd-rom. if this option was followed then an information analyst, based at the mohfw, could be responsible for supporting gob users of heda and liaising with gob data producers, as well as carrying out analyses using heda to answer ad hoc queries from policy makers sustainability if heda is to be properly maintained, suitable funding arrangements need be agreed both in the short term and in the longer term. one of the main aims of heda is improved dissemination of information, and this will be achieved by encouraging data collectors to deposit information about their databases in heda and by encouraging use of the information and databases held on heda. any fees introduced for either depositing information, or for using heda, will therefore need to be carefully considered to ensure that they are not creating barriers to effective development and use of heda. funding mechanisms used elsewhere usually include some ‘block’ funding to cover the basic cost of maintaining and developing the archive, with only a proportion of the overall cost being recovered through user fees. these user fees are usually linked to a registration fee for an organisation or individual wishing to sign up as an ‘archive user’, rather than being linked to amount of use which can be difficult to monitor. an additional charge is usually made for use of archive staff time to access and analyse information on behalf of a user, unless it is a routine query, which can be dealt with quickly. opportunities for obtaining ‘block’ funding from different sources are currently being considered and will be explored further once the institutionalisation process has been completed. acknowledgements thanks are due to all the participants at the workshops, training sessions, development meetings and launch for heda for their time and helpful comments on the design and development. in particular thanks are due to the following for their commitment and assistance to the project: mr shafiq hussein, computer programmer, hb consultants, dr indrani haque, independent university of bangladesh, farrah hannan, azizur rahman and tahmina begum of data international. we would also like to express our appreciation to professor sushil howlader, head of the institute of health economics for his vision and commitment in the setting up of this institute, and in taking the data archive forward as a key part of the institute’s planned resource centre. * paper presented at:international association for social science information service & technology, building bridges, breaking barriers: the future of data in the global network, toronto, may, 1999. deana leadbeter, international health information specialist, south east institute of public health, tunbridge wells, uk. lorna guinness, maxwell stamp plc, london, uk, and associate economist, health economics unit, ministry of health and family welfare, government of the people’s republic of bangladesh. iassist quarterly 2016 27 iassist quarterly digitization, data curation, and human rights documents: case study of a library-researcher-practitioner collaboration by amy barton, paul j. bracke, ann marie clark1 abstract at purdue university libraries, a project involving the digitization of amnesty international urgent action bulletins from 1974-2007 combined the strengths of political science and library science researchers. the political science research was centered on transnational human rights advocacy and legal instrumentation changes over time, while the libraries’ research related to data management, data lifecycle and curation, metadata, and collaborative research modeling. the conceptual framework of this case study is rooted in the literature on embedded librarianship and lifecycle models of data curation. we investigate the intersections and alignments between scholarly workflow and curatorial workflow, and the implications of these intersections and alignments in collaborative research and curatorial lifecycles. the case study also examines how library resources supported research, and how library science and political science experts collaborated in research through the development of a conceptual model. a research collaboration model was developed specifically for the human rights texts project, but was then generalized to be applicable for a variety of practitioner-librarian collaboration projects. the research resulted in data production, data curation, data management, data publication, and scholarly communication and dissemination. keywords: data curation lifecycle, metadata, human rights, amnesty international, research collaboration, digitization, archives introduction partnerships among librarians and faculty members that develop ways to preserve and create digital access to research information have the potential to open up new avenues of teaching and inquiry for both faculty and librarians. a project at purdue university to create a digital research collection of human rights documents has piloted this sort of innovation and collaboration. by its nature, the project required close communication at key points to make the most of faculty and library expertise. this paper explores the process of creating a digital collection of international human rights documents in a way that is integrated into the research workflows of faculty. academic libraries have been exploring new roles in recent years that improve engagement and more tightly couple the activities of the libraries with the activities of their users and institutions. this has been manifest in several ways. significant attention has been paid to better integrating digital collections with discovery and linking technologies, and ‘getting into the flow’ of faculty and students (dempsey 2012). there has also been considerable focus on revitalizing liaison librarian roles as a way of developing stronger relationships and partnerships with faculty, particularly in support of information literacy, scholarly communication, and support for digital scholarship (auckland 2012; jaguszewski & williams 2013; kenney 2014). information literacy and informed learning, for example, imply a deeper level of integration into the curriculum than bibliographic instruction efforts of the past (jaguszewski & williams 2013). libraries have also been active in recent years in exploring opportunities for supporting scholars in their research processes. there has been considerable interest among research libraries, for example, in supporting digital humanities and e-science services. these often involve partnerships between libraries and researchers to find ways to better integrate library collections, library science expertise, and library services into research processes, and to better integrate primary research outputs, such as datasets, into curatorial processes (brandt 2007). the digitization project for the amnesty international urgent action 28 iassist quarterly 2016 iassist quarterly bulletins collection at purdue university libraries, referred to below as the human rights texts project, developed out of such a partnership between faculty and the library. the goal was not merely to develop information products that are of general use to scholars, but also to align the development process with lifecycle models of data use and scholarship support. conceptual framework and research questions the conceptual framework of this case study is rooted in the literature on embedded librarianship and lifecycle models of data curation. embedded librarianship is an emerging model of librarianship in which librarians apply expertise in information organization, management, and use in the context of library users – embedded within teaching, research, or clinical environments (schumaker & talley 2009, p. 8). rather than a passive model in which librarians react to the needs of their constituents, embedded librarianship presents an active model in which librarians are partners in addressing instructional or research needs (kesselman & watstein 2009; rankin et al. 2008; carlson & kneale 2011; clyde & lee 2011; rudasill 2010). while many librarians, especially those at purdue university libraries, have focused their efforts in support of information literacy and instruction, there is increasing interest in developing partnership models for research collaborations (brandt 2007). numerous discussions in the literature engage in how librarians can develop relationships with scientists and social scientists related to data management, and describe pilot programs to do so (brandt 2007; garritano & carlson 2009; walters 2009). there is also a history of librarian partnerships in the digital humanities, often through the creation of digital humanities centers, in which scholars may receive support for research projects. for example, assistance with digitization or text analysis (vandegrift & varner 2013; posner 2013; gold 2012; svensson 2010; zorich 2009; zorich 2008). this study also is framed by lifecycle models of data management and research collaboration. there has been significant interest among librarians and others involved in data management in recent years about data curation lifecycles, a phrase that refers to how data are acquired and cared for throughout their production, acquisition, use, and preservation (carlson 2014). one of the most cited of these models is the dcc curation lifecycle, developed at the university of edinburgh’s digital curation center (dcc). this model is designed to aid the planning of curation and preservation activities, and the long-term continuity of access to digital materials (higgins 2008). accordingly, it is largely focused on activities directly related to curation and preservation, and presents a model to which more granularly-defined local practices could be mapped. the model, seen in figure 1, places the digital object at the center, with central rings representing curation actions applicable across the entire lifecycle of the object and outer rings representing sequential or occasional actions that are performed upon the data. while this is a robust, generalized framework for planning, it is also a framework that was created with a specific purpose – digital curation. while curatorial activities and responsibilities were well represented in the model, the scholarly processes that drive the creation and use of digital objects are under-represented. other models (ddialliance.org 2013; green & humphrey 2013; vardigan et al. 2008) have inverted this focus on developing a figure 1: dcc curation lifecycle model figure 2: ddi combined lifecycle model iassist quarterly 2016 29 iassist quarterly model of data-based research that feeds curatorial processes. taking the data documentation initiative (ddi) combined lifecycle model (figure 2) as an example, the lifecycle begins with the genesis of a research study and progresses through a number of stages of data preparation and use, with a step for data archiving and distribution. this view is researcher-centric and underrepresents curatorial processes. designing a practice to integrate two different lifecycle models presents a challenge. for the project presented in this case study, the practice must serve both research and curatorial goals: first, by providing ongoing research access to primary source documents, facilitating qualitative coding of the documents and ultimately the creation of qualitative and quantitative data sets; and second, by simultaneously incorporating the curatorial steps required to create the lasting collection of primary source documents, a research database, and data that are appropriate for long-term curation and preservation. the researcher and the library units involved in this project envisioned mutual benefit in building a process that could use research steps to facilitate curation and, in addition, employ curatorial techniques to facilitate research. accordingly, this case study sheds light on a question central to research on integrated models of the data lifecycle: can digitization processes be designed in a manner that feeds directly into analytical workflows of social science researchers, while still meeting the needs of the archive or library concerned with longterm stewardship of the digitized content? answering this question, from the standpoint of an academic library, leads to two subquestions. first, what are the intersections and alignments between scholarly workflow and curatorial workflow? second, what are the implications of these intersections and alignments on research and curatorial lifecycles? the project since the mid-1970s, as an advocacy technique in support of human rights, amnesty international (ai) has produced periodic, almost daily urgent action (ua) bulletins (rydkvist 2013). these bulletins are shared with members of its ua network, who are asked to write direct appeals on behalf of individuals whom ai believes to be at risk of human rights violations. examples have included action on behalf of people who may be tortured, disappeared, or detained illegally, or severely mistreated while in detention. the bulletins themselves are oneto three-page informational alerts advising members to write quick messages directly to officials in violating governments. in addition, the bulletins advise members on how to raise issues of concern in each specific case. concerns may range from protecting the affected person’s physical well-being to other forms of compliance with relevant principles of human rights protection. the ua bulletins collection provides a detailed record of a major human rights organization’s transnational advocacy on behalf of individuals over more than four decades. data can be drawn from the documents to serve as indicators of how human rights concerns changed over time, as well as the nature of human rights threats in different countries and different periods. the documents also provide a unique window into the mobilization of the human rights movement during a crucial time period in the development of international human rights law (clark 2001). the digitized collection will be of use to researchers, as well as journalists and attorneys who can refer to the documents as primary evidence from the historical record. the project to digitize the ai ua bulletins was initiated for several reasons. the researcher, a faculty member in purdue’s department of political science (principal investigator [pi]), had served for a number of years as a volunteer on amnesty international-usa’s (ai-usa) archives advisory committee. in this capacity, she learned of the need to preserve the early records of ai-usa, the united states branch of the global human rights organization. the ai-usa offices, however, housed additional documents with significant research value that were publications of ai’s international secretariat (ai-is), the organization’s international headquarters in london. these documents included the ua bulletins and associated documents as they had been processed for distribution to members of the network in the united states. while the preservation needs for documents created by ai-usa were being addressed with the establishment of its organizational archives at columbia university in 2007, digital preservation of the ua documents originally created by ai-is were ultimately deemed out of scope for columbia’s collecting practices. between 2007 and 2011, the researcher engaged in extended conversations with staff from ai-is, ai-usa, and the amnesty international usa archives (columbia university, n.d.) about developing joint digitization and research as a stand-alone project related to the ua bulletins processed and issued by ai-usa. after permission to pursue funding for digitization and research with the documents was approved by ai-usa, the researcher contacted the purdue libraries. with the libraries, discussion centered around how to prepare the ua bulletins as a publicly available collection in conjunction with the pi’s research process, as well as how best to incorporate them into the libraries’ digital collections. a funding proposal was developed that sought to combine the strengths of the librarians’ archiving, technical, metadata and research data expertise with the researcher’s scholarly human rights expertise, and contacts at ai-usa and ai-is, who saw the benefit of a digitized collection of these particular documents. ai-usa agreed on the general outlines of the project as envisioned and agreed to share the documents. a competitive grant from the purdue university office of the vice president for research funded the project (clark & bracke 2012). the proposal was to combine library standards for digitization and digital collections, as well as additional researcher and practitioner-driven metadata and coding strategies. the collaborative research would result in a searchable, e-archive primary source collection that would also benefit the human rights organization, as well as a human rights dataset for researcher use, with the potential for expansion into a numeric dataset compatible with other international data sources. working with ai-usa required extensive consultation, both before and after funding was secured, to ensure that each of the parties’ needs and interests, and sensitive information concerns, would be taken into account. variation in the timelines and organizational approaches of each of the three parties (library, researcher, and human rights organization) posed further challenges. while these logistical issues slowed progress as the project got off the ground, in retrospect it was a necessary feature of the collaborative process. research process processing workflow archival best practices were established for the handling of the documents during digitization and post-processing reassemble. this ensured the documents could be sent along to their 30 iassist quarterly 2016 iassist quarterly permanent home at the amnesty international usa archives at columbia university in a good archival state. the digitization process (figure 3: 1.a digitization) was standard, with the exception of the creation of the optical character recognized (ocr), full text pdf derivative. high resolution and high quality ocr were necessary for accurate topic modeling by the pi and her research team. nvivo software (qsr international 2014) was used by the pi to code the digitized documents. the quality of the processed pdf documents was crucial for text recognition and topic modeling in the nvivo software for accurate and comprehensive coding. digitization was driven by the need for processed documents to feed into the coding progression. in this scenario, it became a challenge to keep up with the need for high quality derivatives, which were time and labor intensive to produce, and the need for documents to code. ultimately, these activities aligned such that the digitization and coding co-occurred in real time (figure 3: 1.b. coding). data and metadata ai-is and ai-usa each provided a data source documenting records of ua bulletins and associated documents from 1974-2013. these sources had been independently maintained, resulting in differences in the underlying data. each contained some unique information, and some records were incomplete in one source or the other. combined into a single metadata master file along with technical metadata captured during digitization (figure 3: 1.c data figure 3: processing workflow merge), the two sources provided comprehensive information such as the organization’s internal document numbers; dates and countries addressed; document types; and some details of the case. the combined file is being used as master metadata file for the public-facing primary source digital collection and as a basis for a research database for the pi (figure 3: 2. metadata & 3. collection). in adherence to library science best practices, we identified three human rights authoritative vocabularies and used them, in combination, to create a controlled vocabulary for the project that would supplement ad-hoc keywords assigned by ai-is and ai-usa in each data source. the combined authoritative vocabularies include huridocs’ ‘micro-thesauri: a tool for documenting human rights violations,’ ‘witness media archive topic terms’ published by witness, a human rights ngo, and the ‘universal human rights index research guide’ produced by the united nations office of the high commissioner for human rights (huridocs 2010; united nations office of the high commissioner for human rights n.d.; witness n.d.). the pi and her research team combed the vocabularies to identify semantics that closely align with her research interests as well as the descriptive terms used in the two amnesty international data sources. the development of the controlled vocabulary was used to develop coding structures for content analysis within nvivo. this allowed the research team to create an export file within nvivo (‘node extracts’) that listed each code applied to a document. these node extracts were, iassist quarterly 2016 31 iassist quarterly in turn, added to the metadata master file by the libraries. this aligned indexing terms between the researcher’s nvivo research files, the master metadata file, the digital collection, and the research database. the controlled vocabulary we developed will be shared with ai for possible applications in their own data sources. interoperability and metadata exchange opportunities in any future projects with ai are more feasible if the concepts and terms are standardized in all parties’ data sources. findings and discussion this research project was intended to explore whether digitization processes can be designed in a manner that feeds directly into content analysis workflows of social science researchers, while still meeting the needs of the archive or library concerned with long-term stewardship of the digitized content and data. the authors sought to understand the intersections between scholarly workflow and curatorial workflow and their implications for research and curatorial lifecycle. the project illuminated both opportunities and challenges for advancing such collaborations. intersection of workflows one of the key findings of the project was that, while some of our ideas about integration were correct, there were significant difficulties in aligning workflows between library and researcher from a scheduling point of view. areas of success in aligning workflows included developing shared standards for file naming, application of the controlled vocabulary, and extraction of code structures from nvivo to provide subject enrichment of descriptive metadata being processed from ai-is and ai-usa data sources. the successful extraction of code structures for descriptive purposes demonstrated that qualitative coding processes can be leveraged in metadata creation. despite these points of overlap between the two, a fundamental challenge was achieving alignment in the context of a grant-funded project with time constraints. the timelines associated with a grant-funded project required both parties, library and researcher, to begin their work as soon as possible. the first issue that arose was the need for a formal agreement between purdue and amnesty international prior to shipping documents for scanning. while this process went relatively smoothly, it did result in a delay of several months in having the ability to begin scanning of documents. while graduate assistants working for the researcher were able to begin their coding process with newer documents, available in borndigital formats from ai-is (figure 3: 1.b coding) that could later be matched with scans of the same documents from ai-usa, the delay in beginning the scanning process would prove problematic later in the process as it was difficult to maintain a steady supply of scanned documents for coding. the second issue that arose, from a project planning point of view, was that initial planning was done based on photographs of sample documents from the collection taken by the pi during a visit to the ai-usa archive. once documents were received, it became apparent that ocr processes would be slower than anticipated due to document condition. a third challenge resulting from the grant-driven timeline was with the use of student labor in scanning. when the documents were figure 4: conceptual partnership model based on the human rights texts project 32 iassist quarterly 2016 iassist quarterly received for scanning it was mid-semester, making the hiring of new student workers more challenging than it would have been at the beginning of the semester. additionally, there are inconsistencies in productivity inherent in the use of student labor. students are often less available for work, at short notice, during exam periods and during semester breaks. intersection of research and curatorial lifecycle working through the phases of the research project, we discovered that the intersections and alignments of research and curatorial lifecycles include collaborative research, the alignment of research and curatorial processes, and the development of a new research partnership model. due to the way the project unfolded, with delays, digitization challenges, and labor and time constraints, much of the research had to be done in real-time. the research model worked well for the project. the pi and her research team, and the faculty from purdue libraries and their digitization team, were able to identify research problems and processing issues that impacted research, then work together to produce the intended research outcomes. the libraries supported the project with services such as digitization, digital content management and digital collection development. however, our experience in this project went beyond typical services. two libraries faculty collaborated with the political scientist in research that resulted in the development of a dataset (populated master metadata file) that further facilitated the development of a research database and a public-facing, primary source digital collection (in development). we felt this was a fairly unique situation in that the libraries faculty was engaged in research alongside the political scientist in the context of research data development, metadata application, data management and curation. below, we present a model that explores the process and components involved in a successful research collaboration based on the human rights texts project. the conceptual model (see figure 4, page 31) represents scholarly collaboration between political science and library science experts, along with libraries services and resources, to conduct research that resulted in metadata and data development, data curation throughout the data lifecycle, digital collection development, and scholarly outputs and dissemination. in the model, the dark blue represents the research continuum from start to completion. the light blue circles on the left of the model represent libraries services and resources that supported the research process. the circle in the middle represents political science and library science expertise converging in real-time research. finally, the white circles on the right of the model represent outcomes. the seeds of the project were sown by the pi as early as 2007, but ‘project development,’ in the first dark blue circle, began after the libraries and the pi had discussions and committed to the project. funding was sought and awarded. the next point on the continuum, ‘domain expertise,’ the second dark blue circle on the continuum, overlaps the first set of light blue circles representing figure 5: generalized conceptual partnership model iassist quarterly 2016 33 iassist quarterly libraries services and resources. at this point, archives and digital programs established standards and workflows for digitization, while the libraries established data management and curation standards and workflows in consultation with political science and library science experts. this is also the place where the pi and metadata specialist developed the controlled vocabulary to be used throughout the project. in the model, here the continuum undulates to represent the pi working specifically with her research team to develop processes for nvivo coding, while the metadata specialist researched the development of the master metadata file and the data sources merge workflow. the libraries faculty and the pi come back together in the middle circle to conduct research interacting as researchers, addressing real-time problems, and producing research data. here again, the continuum undulates to represent the pi evaluating her coding results on a subset of data, and the libraries faculty working with the libraries it to develop the research database based on the data-populated master metadata file. finally, the pi and the libraries researchers converge once again to evaluate the outcomes of the project: the refinement of the ––research database; the development of the digital collection; long term data preservation and curation; scholarly communication via dissemination of research results; and the publication of the raw dataset through the purdue university research repository (purr). having described the development of a research collaboration and partnership model for a very specific research project, we now explore how the model could be generalized and applied to other academic libraries and research collaborations, inclusive of libraries services and expertise. the generalized model (see figure 5, page 32) represents on a research continuum from start to completion the major components of the process that can provide guidance in establishing a libraries-domain expertise research collaboration. we previously discussed in detail how this conceptual model applied to a specific research project. taking a step back, one can see the three major components of this model: • research project development, • applied research collaboration, • tangible and intangible research results, which includes scholarly impact and dissemination. the research project development component includes relationship building; project development and planning; identifying domain expertise combined with library expertise to address the research problem; and the library services and resources that will support the research continuum from start to completion. the applied research collaboration component occurs after all the project planning, workflows, and processes have been established. in this component, combined domains apply research methods to address the research inquiry and to produce data. this component leads into the last component, research results. this includes the intangible results solidifying relationships and continued collaborations and the tangibles data production, data curation, data management, data publication, and scholarly communication and dissemination. based on the successful outcomes of the human rights texts research project, further exploration of this model and the application of this model perhaps even a component to operationalize the model for project management will be pursued. references auckland, m., 2012. re-skilling for research. london, uk: research libraries uk, (january 2012). available at: http://www.rluk.ac.uk/wp-content/uploads/2014/02/ rluk-re-skilling.pdf. [accessed 25/01/2016]. brandt, d.s., 2007. librarians as partners in e-research. c&rl news. 68 (6). p. 365–368. carlson, j., 2014. the use of lifecycle models in developing and supporting data services. in: ray, j.m. (ed). research data management: practical strategies for information professionals. west lafayette, in: purdue university press. carlson, j. & kneale, r., 2011. embedded librarianship in the research context. college research libraries news. 72 (3). p. 167–170. clark, a.m. 2001. diplomacy of conscience: amnesty international and changing human rights norms. princeton, nj: princeton university press. clark, a.m., & bracke, p., 2012. human rights texts for digital research: archiving and analyzing amnesty international’s historic ‘urgent action’ bulletins at purdue university. office of the vice president for research, purdue university, award no. 206400; granting period january 1, 2012 december 31, 2014. clyde, j. & lee, j., 2011. embedded reference to embedded librarianship: 6 years at the university of calgary. journal of library administration. 51 (4). p. 389–402. columbia university. center for human rights documentation and research. n.d. amnesty international usa archives. new york, ny. available from: http://library.columbia. edu/locations/chrdr/archive_collections/aiusa.html. [accessed 25 january 2016]. ddialliance.org, 2013. ddi data documentation initiative. welcome to the data documentation initiative. available at: http://www. ddialliance.org/. dempsey, l., 2012. thirteen ways of looking at libraries, discovery, and the catalog: scale, workflow, attention. educause review online. [online] december, 2012. p. 1–14. available from: http://er.educause. edu/articles/2012/12/thirteen-ways-of-looking-at-librariesdiscovery-and-the-catalog-scale-workflow-attention. [accessed 25/01/2016]. garritano, j.r. & carlson, j.r., 2009. a subject librarian’s guide to collaborating on e-science projects. issues in science and technology librarianship. (57), p. 5. gold, m.k., 2012. debates in the digital humanities. digital scholarship in the humanities. [online] june 2014. p. 504. available from: http:// dhdebates.gc.cuny.edu/debates. [accessed 25/01/2016]. green, a. & humphrey, c., 2013. building the ddi. iassist quarterly. 37 (1-4). p. 36–44. higgins, s., 2008. the dcc curation lifecycle model. international journal of digital curation. 3 (1). p. 134–140. huridocs, 2010. micro-thesauri: a tool for documenting human rights violations | huridocs, geneva: human rights information and documentation systems, international. available at: https://www. huridocs.org/resource/micro-thesauri/. [accessed 11/10/2015]. jaguszewski, j.m. & williams, k., 2013. new roles for new times: transforming liaison roles in research libraries, washington, d.c. available at: http://www.arl.org/storage/documents/publications/ nrnt-liaison-roles-revised.pdf. [accessed 25/01/2016]. kenney, a.r., 2014. leveraging the liaison model: from defining 21st century research libraries to implementing 21st century research universities. ithaka s+r, p.11. available at: http://www.sr.ithaka.org/ blog-individual/leveraging-liaison-model. [accessed 25/01/2016]. 34 iassist quarterly 2016 iassist quarterly kessleman, m.a. & watstein, s.b., 2009. creating opportunities: embedded librarians. journal of library administration. 49 (4) p. 383–400. posner, m., 2013. no half measures: overcoming common challenges to doing digital humanities in the library. journal of library administration. 53 (1). p. 43–52. qsr international. 2014. nvivo 10 for windows. published by qsr international pty ltd. available at qsrinternational.com. rankin, j.a., grefsheim, s.f. & canto, c.c., 2008. the emerging informationist specialty: a systematic review of the literature. journal of the medical library association : jmla. 96 (3). p.194–206. rudasill, l.m., 2010. beyond subject specialization: the creation of embedded librarians. public services quarterly. 6 (2-3). p.83–91. rydkvist, s., 2013. still relevant after all these years. [online] march 19, 2013. available from: http://www.amnesty.org.uk/blogs/ urgent-action-network-blog/still-urgent-after-40-years. [accessed 25/01/2015]. schumaker, d. & talley. m., 2009. models of embedded librarianship: final report. alexandria, va: special libraries association. available from: http://hq.sla.org/pdfs/embeddedlibrarianshipfinalrptrev.pdf. [accessed 25/01/2016]. svensson, p., 2010. the landscape of digital humanities. digital humanities quarterly [online] 4 (1). p. 1–31. available from: http:// www.digitalhumanities.org/dhq/vol/4/1/000080/000080.html#. [accessed 25/01/2016]. united nations office of the high commissioner for human rights, universal human rights index research guide. available from: http://uhri.ohchr.org/search/guide. [accessed 11/10/2015]. vandegrift, m. & varner, s., 2013. evolving in common: creating mutually supportive relationships between libraries and the digital humanities. journal of library administration. 53 (1). p. 67–78. vardigan, m., heus, p. & thomas, w., 2008. data documentation initiative: toward a standard for the social sciences. international journal of digital curation. 3 (1) p. 107–113. walters, t.o., 2009. data curation program development in u.s. universities: the georgia institute of technology example. international journal of digital curation. 4 (3). p.83–92. witness, witness media archive topic terms. available at: http://www2. witness.org/vocab/topics/. [accessed october 11, 2015]. zorich, d.m., 2008. a survey of digital humanities centers in the united states. available from: http://www.clir.org/pubs/reports/pub143/ contents.html. [accessed 25/01/2016]. zorich, d.m., 2009. digital humanities centers: loci for digital scholarship. working together or apart: promoting the next generation of digital scholarship. p.70–78. available from: http://pdf. aminer.org/000/246/607/the_next_generation_distributed_object_ web_and_its_application_to.pdf#page=78. [accessed 25/01/2016]. 12 iassist quarterly summer 2012 iassist quarterly abstract efforts towards internationalization have become increasingly important in scientific environments. as for content-based indexing of scientific research data, however, standards leading to internationally coherent indexing which is vital for retrieval purposes are not yet sufficiently developed. even concerning the concrete use of indexing instruments, launched by initiatives on an international scale, there are still no binding policies and guidelines. against this backdrop, essential criteria which internationally applicable indexing systems should meet will be outlined. these will be illustrated through the multilingual european language social science thesaurus (elsst), originally based on the uk data archive’s (ukda) humanities and social science electronic thesaurus (hasset) and ultimately developed by the council of european social science data archives (cessda). additionally, the general pros and cons of using international versus national indexing languages will be weighed using the elsst and the thesaurus for the social sciences (tss) developed by gesis – leibniz-institute for the social sciences. in this light, the benefit of vocabulary crosswalks for supporting a combined use of international and national indexing systems will be discussed. keywords: research data, cataloguing, thesaurus, internationalization, social sciences. introduction over the past several years, multiple efforts pertaining to the standardization of workflows, working instruments and working methods have been undertaken in various scientific domains. these efforts have been at national levels and, increasingly, on an international scale because standardization both supports and facilitates interoperability between cooperating institutions. in the field of the social sciences the data documentation initiative (ddi) metadata specification has been developed to serve as an international standard for describing data from the social sciences and related disciplines. by using this standard, coherent documentation across institutions in different countries is ensured and data exchange facilitated. internationalization and standardization efforts can also be observed in the context of subject indexing. the use of commonly applied indexing systems or the mapping of dispersed terminological resources is an attempt to support subject retrieval across distributed collections. subject indexing of research data in the social sciences in europe the council of european social science data archives (cessda), founded in the 1970s, is an umbrella organization for european social science data archives. membership is comprised of data archives and other organizations which archive and provide social science data for secondary use. cessda currently provides access to 25,000 datasets with the collection growing thesaurus-based indexing of research data in the social sciences: opportunities and difficulties of internationalization efforts by katrin baum1 and andreas oskar kempf2 elsst iassist quarterly summer 2012 13 iassist quarterly by approximately 1,000 datasets annually. among other functions cessda is responsible for the development and maintenance of the european language social science thesaurus (elsst) and the topic classification, both of which are used for subject indexing of research data by the member organizations. the cessda catalogue enables retrieval of data stored at cessda archives throughout europe and provides besides free-text search options for searches by topic of studies indexed with the topic classification or searches by keyword for studies indexed with elsst. the european language social science thesaurus (elsst) is used for subject indexing of research data by the cessda member organizations. it is based on the subject thesaurus hasset which is hosted by the uk data archive and is being further developed by the cessda thesaurus management team. it is a multilingual thesaurus for the social sciences and has been translated from english into danish, finnish, french, german, greek, norwegian, spanish and swedish. it consists of approximately 3,300 concepts extracted from hasset. these concepts aim to be culturally neutral thereby reflecting a european perspective instead of one that is country-specific, thus allowing international applicability of terms. furthermore, elsst allows for the addition of local extensions, which means that concepts of local importance can be added to meet institutional needs. currently, indexing practices using elsst vary widely across the participating institutions due to the lack of binding indexing guidelines. for example, indexing specificity ranges from the description on a very general level with only some descriptors for one study to a very precise and deep indexing with more than a hundred descriptors for one study. this can lead to the result that the same issue is very differently described. this again has implications on retrieval as the same issue can only be retrieved by using different search terms – a fact that users will not be aware of. additionally, due to the requirement that concepts be internationally applicable, fine-grained local issues as well as historical, juridical, religious, political and other country-specific aspects cannot be displayed if using solely elsst. consequently, retrieval is limited to internationally valid concepts. thesauri in subject indexing thesauri are being used for verbal subject indexing in documentation. consistently applying the same, controlled descriptor for specific issues results in consistent documentation and facilitates retrieval. indexing systems are usually based on specific collections, meaning that content and structure of systems even in the same domain can differ considerably. as well, levels of abstraction and hence specificity can vary among different thesauri depending on local conditions and needs. moreover, different classification aspects following from a variety of perspectives on a topic can lead to different semantic relations between concepts. requirements of an internationally applicable thesaurus one of the most important requirements for an internationally applicable thesaurus is that it be free of bias. the concepts it contains need to exist in every participating culture and have to be displayed in a hierarchical and semantic structure that fits all cultures and languages. terms for concepts have to be multilingual to allow access in all of the languages in use. however, the characteristics of an internationally usable system such as this include numerous limitations and constraints. finegrained issues at both the institutional and country-, respectively language-specific level cannot be displayed; thus retrieval is limited to internationally applicable concepts. local indexing systems local indexing systems are able to reflect the scope of the local collection very accurately and with respect to cultural characteristics. this allows for more precise indexing. beyond that, they are easier to maintain as there are no cross-institutional agreements to follow. on the other hand, the exclusive use of a local indexing system has its own deficiencies. since it remains a solely locally applied system, without further measures there can be no access points for unified subject retrieval across dispersed collections that have been indexed using different terminological resources. recommended indexing model one possible way to offer uniform subject access to heterogeneously indexed collections in a dispersed environment is the mapping of institutionally used indexing systems (doerr 2001). applied on retrieval, mappings aid the user in finding documents indexed no matter which indexing system was used by being able to only employ search terms in the system he or she is familiar with. though, mappings often carry a certain amount of intrinsic vagueness due to incomplete congruity between concepts of different indexing systems. for this reason we propose an aggregate of local thesauri with common, internationally applicable core concepts (boteram & hubrich 2008; gödert 2008). it should contain concepts that exist in any language and its hierarchical structure should fit all languages as close as possible ensuring that it is free of bias. concepts already existing in local systems should be mapped to concepts of the core system. as well, concepts missing in the local systems should be added. figure 1: interplay between local and international indexing system. 14 iassist quarterly summer 2012 iassist quarterly to reiterate, we would argue for a direct interplay between international and local indexing system in a way that will permit internationally applicable concepts standing for the core vocabulary to be fully integrated, i.e., represented, in all of the different local indexing systems attached to it. consequently, the whole content of the core vocabulary is simultaneously part of every single local indexing system. this is the reason that the universal core system, i.e., the elsst, must contain all central concepts which exist in all of the included languages. and, vice versa, these central concepts must be integrated into local indexing systems. for instance, the key indexing tool for german-language social sciences, the thesaurus for the social sciences (tss), translated into english and french, contains more than 8,000 concepts and, like the elsst, includes a wide range of subdisciplines of the social sciences. looking at this direct interplay between local and international indexing system in practice, we explicitly advocate for further use of the local system as key indexing system. taking “secondary school” as an example from the education sector, this concept, referring to the local indexing system introduced so far, needs to be, if not already the case, incorporated into the tss. to summarize, missing concepts which are part of the internationally used core vocabulary must be created in the local indexing systems. in addition to these internationally applicable concepts, the local indexing system must also contain any locally distinctive specificities. for example, the german concept “gymnasium”, stands for a certain type of secondary school, one with a strong emphasis on academic learning. it is comparable to the british grammar school system or preparatory schools in the united states. moreover, the local indexing system contains collection-specific concepts, e.g., the geographic subject heading “nordrhein-westfalen”, a federal state in germany, indicating, for instance, the provenance of a dataset. it becomes clear that a connection between internationally used core concepts and local indexing systems is necessary. in our judgment, this linkage could be best achieved by terminology mapping between international and local indexing system. referring again to the two thesauri mentioned above, we like to hint at the major terminology mapping initiative conducted by gesis leibniz-institute for the social sciences as part of the project competence center modeling and treatment of semantic heterogeneity (komohe) (mayr & petras 2008). carried out shortly after the elsst extension in the framework of the european commission project multilingual access to data infrastrctures of the european research area (madiera) its main objective was to create crosswalks between various controlled vocabularies and also between the elsst and the tss. approximately 2,300 equivalent relations were built up in each direction which could be reused when translating elsst vocabulary for inclusion into the tss. a search using the elsst-concept “secondary schools”, would directly create a link to datasets indexed with the german term “weiterführende schule”, and respectively into their english and french translations. additionally with the help of these crosswalks, the extent to which the international core system is already part of local indexing systems becomes apparent. moreover, looking at these crosswalks from the opposite direction, in our case from the tss to the elsst, and looking at non-equivalent relations of this mapping, which had also been built up in the past, gives a hint at local specificities being part of the local indexing system. for example, the german concepts “gymnasium”, “realschule” and “hauptschule” are narrower terms of the above mentioned concept “weiterführende schule”. retrieval aspects we suggest there are significant information retrieval benefits to be obtained as a result of the direct interplay that occurs between internationally applicable concepts and local indexing systems. first of all, using an integrated retrieval system, e.g., the cessda catalogue, the researcher is able to use proper terminology for core concepts included in the multilingual elsst. due to the hierarchical semantic structure of the thesaurus, the researcher is aided in the search for narrower subject terms. at this point the researcher has access to all the data collections of the cessda member organizations indexed with those commonly shared international core concepts. for fine-grained regional datasets with existing vocabulary mappings between international and local indexing system, a linkage between both vocabularies could be established. similar to the cessda catalogue’s tree-like hierarchical search structure, additional local specific terms could be connected to the broader terms in the international indexing system. thus, datasets indexed with the specific subject headings become searchable and accessible. it will prevent locally embedded information from being buried under broad general indexing terms. figure 2: conclusion in a period when efforts towards standardization and internationalization of subject indexing of research data have become increasingly important, there is a pressing need to determine ways to integrate local indexing systems into widely launched internationally applicable vocabularies. even though work on this began in the 1970s, there are still no binding and coherent indexing guidelines. vast differences in cross-country indexing of research data remain. with this as our backdrop, our aim was to present a concrete proposal on how to create an interconnection between local and international indexing system. hence, we argued for an aggregate of local thesauri with a common core vocabulary of internationally applicable figure 2: application of the indexing model for retrieval. iassist quarterly summer 2012 15 iassist quarterly key concepts. these concepts would exist in all of the participating languages represented in cessda. the vocabulary would be kept free of bias as much as possible. concepts already integrated in local systems would be mapped to the core system and any missing concepts in the local systems would be added. doing this would achieve a coherent and unified subject indexing of dispersed collections of research data. information retrieval would be significantly improved. fine-grained institutional and locally distinctive datasets that have been indexed with collection-specific subject headings will become searchable via crosswalks built up to the international indexing system. the local indexing system will remain the key tool as it is much easier to maintain. concrete efforts to move forward will include as a first step the review of elsst to ensure no bias in the concepts and the interconceptrelations. following this, the second step is to adapt the local systems. subsequently, universally accepted indexing guidelines for the core system need to be developed. references boteram, f, hubrich, j 2008, towards a comprehensive international knowledge organization system, available from: < http://linux2.fbi. fh-koeln.de/crisscross/vortraege.html >. [19 july 2013]. council of european social science data archives 2013, available from: . [19 july 2013]. cessda catalogue 2013, available from: . [19 july 2013]. doerr, m 2001, ‘semantic problems of thesaurus mapping’, journal for digital information, vol. 1, no. 8. available from: . [19 july 2013]. european language social science thesaurus 2013, available from: . [19 july 2013]. gödert, w 2008, ontological spine, localization and multilingual access: some reflections and a proposal, available from: . [19 july 2013]. mayr, p, petras, v 2008, cross-concordances: terminology mapping and its effectiveness for information retrieval, available from: . [19 july 2013]. thesaurus for the social sciences 2013, available from: . [19 july 2013]. notes 1. katrin baum is a librarian at gesis – leibniz-institute for the social sciences in cologne, germany, where she works in the data archive department in the area of study descriptions. her professional focus is on subject indexing and information retrieval. katrin can be reached by email: mailto:katrin.baum@gesis.org. 2. andreas oskar kempf is a research associate at gesis – leibnizinstitute for the social sciences, cologne, germany. he holds a phd in sociology from goethe university, frankfurt am main, and received a master´s degree in library and information science from humboldt-university, berlin, germany. andreas conducts applied research on gesis authority data for content cataloguing (i.e. thesaurus for the social sciences) of social science research literature, projects and data. he can be reached by email: mailto:andreas. kempf@gesis.org. lassist newsletter* vol. 3t no. 3 (summer 1979) a library-based reference service for machine-readable census data; the canadian experience slavko manojlovlch university of windsor windsort ontario the cens source for econom i c an social scie small propo census data copy form, growing use analysis w machi ne-rea the census more import library is information pus with t h avai lable i patron's i the i ibrar paper will place of t h academ i c i followed b potent i al which i i b.r alike may of the m-r us of c the dem d hous nee re rt ion is pu if we of co e can dable ( is beco ant to the ma on the e libra nte rmed nf ormat y • s r at temp e m-r i br ary » y a di problem ar i ans encount census . anad ogra 1 ng sear of t blis ad mput se m-r) mi ng rese jor un r i an i a ry i ona esou t to cens t scus s and er i a i phi c data c h. he hed d to ers inc arch depo i ver act be i n r ces ex us w his s 1 on and re s a prime « social* used in only a avai lable in hardthis the for data that the er s i on of reasingly er s . the sitory of s i t y c aming as an tween the eeds and this amine the ithin the will be of the pitfalls searchers aking use ih£ mze census and jhe academic limaei the census of canada first became available in m-r form in 1961 with the production of the user summary tapes (u.s.t.). in 1971 the public use sample tapes (p.u.s.t.) were produced in response to the researcher's need for micro-data. for those libraries which had access to the m-r census tapes the referencing of the patron's request for census information required a consideration* on the part of the librarian* of both the hard-copy and the m-r census resources. the importance of the latter is evident in the following situations: 1. the requested data is not abailable in a printed publication but may be retrieved from a census tape. since the census bulletins provide summary data for a limited number of standard geographic regions* they are unable to satisfy data requestes for small municipalities* subdivisions of these municipalities* or usercreated geographic areas. city planning districts are an example of the latter. the major feature of the u.s.t. (enumeration area series) is that the record is based upon the enumeration area (e.a.) . this is the smallest geographic unit of analysis allowed by the principle of confidentiality which governs the dissemination of census data. furthermore* since e. a. boundaries never cross those of larger recognized statistilassist newsletter, vol. 3, no. 3 (summer 1979) complete a description as inaicated by the statistics canada informat ion. c) furthermore* these respondents were asked* "who processes the user's request for information from tape?". the majority response was the computer centre (33%), followed by the data arch the ef f o and tre ing a v user with of f cert (i.e soc i ive (2 rt of the c (13). 33% re ar i a t i h i mse out th acuity ai n ge i ogy ) 0%) co-o the ompu the spon on if, e as me depa ogra t hrough perat 1 ve library t er cenrema i nded with of the with or s i stance tubers in r t ment s . phy or table 1 do you have any mac h 1 nere adab i e census tapes on your campus? size of holdings major other total response yes no 8 7 56% (15) 7 5 44% (12) 100% (27) t cate of act1 acce whi c f aci the exi s or, data camp libr cana tape libr eith i nd i hese o that ac adem ve ly i ss to h are uty's libra tence if it from us fa ary . da inf s wer ary . er t h v i dua l bserva the ov ic i nv i ve the m locat tape ry is of th is, th tape c i l i t y a c c or o rma t i e rare r equ e c omp f a c u l 1 1 on er wh i b ra d in -r c ed i lib not e me pr i s ot ding on , ly es t s u t er ty s c le e i m i n r i es p rov ensus n th ra r y . a wa r ce ocess per f o he r to the order were c en membe a rl y g ma are idin res e co re nsus i ng rmed tha stat m-r ed b m a tre r s . i nd i j r i t y not g user our ces mput er either f the tapes of the by a n the i s t i c s census y the de by or by thi s cou i libr the it mere libr t ow a data the data the rule will libr m-r of t d p ary' m-r is ly ary' r ds f il ado i f il an s ( i ngn ary r eso he i ossibly s posit census, more i i a re s tradi m-r in es and t i on o e s in t g lo-ame aacrii ) ess on to acce ur ces a i b rary ' ac ion on kely flee t i on form c omp f a he i r i c a s th pt s a s ho coun with the th t i on al at i uter cha a t es n hou i e p the leg idin t re oth at indi n i pro p t e r t ed cat d f art un 1 v i t i m gs. for the spec t to er hand, this is of the f f erence nc ludi ng grammes. on m-r i t ion of a logui ng oster a of the ersity's ate part 57 lassist newsletter, vol. 3, no. 3 (summer 1979) expand their services primary data reference w reasonable in light of to acquire adequate supp ment at i on. file the of in vide he d the 1 1 em i uld thei ort s tea data seconeither d with e s i red i nf orpt to nc lude not be r need docuthe library has the personnel and the resources to provide a primary data reference service for the m-r census. the success of this venture depends on the following condi t i ons : 1. the university community must recognize thelibrary as the resource centre for all census information* m-r and hard-copy. thust the library should be prepared to promote its expanded service by conducting several information and training seminars for faculty* staff and students on the availability and use of the m-r census . 2. the li take a m-r c rent ly pus a na t e acquis stat is ma t i on of the major acqui r of the sar y a cons i d are s t t e r l i b r a r b ra r n in ensu ava nd s i t i o tics ind i n h ed m tap dded er or ed c e y . y will v en t or s t a i lable hou id the n of canad i c ated st itut i di ng ultipl es * an e xpen that t in t n t r e s at nee y of pes on co-o fu ta a in that ions s e co unne se if he t he co d to the cu rc a mr d i t u re pes . f r2 5% with had pies c esy ou apes mp ut ape h e se institutions tape requests were made by as many as 8 departments. the libr copy wh ic ac ce c omp p r oc r equ some e lem an i a p such t han expe mass a mi of a such help ma j ar ies com h th s th uter . essi n ests st a enta r nt rod rogra as sat i rt i se aging ni -se st a as s ful. or i t y have a put er ey can e un i v howev g of will f f trai y prog u c t or y c mm i ng pl/1 wi sf y the necess data f r ss i on on tistical as would of the hardterminal use to ersity's er , the user require ni ng in ramm i ng . ourse in language ii more computer ary for om t ape . the use package also be a vice tine user is d copy sec a re tied t i on wi i i may use i i b rar for th t advan . firs i rec t ed and m nd, the f erence in the the focus u arise i of the m y-ba e mt age t, t to a -r requ libr art rema pon n t -r c sad r c s f he p n i c ens est aria of inde the he r ensu refer ensus or t at ron nt egr us c is pr n who quer r of prob e f ere enc e has he • s r ated olle oces is y ne this letns nc i n serdiscensus equest har dct i on. sed by qua i igo t i apaper which g and "solutions problems and thelogical analysis of a typical census query into its component parts can provide us with a framework for comprehending the underlying problems associated with referencing census summary data. statistics canada has developed the pqrst classification as the basis of an indexing and retrieval system lassist newsletter* vol. 3» no. 3 (summer 1979) —' — • o — — o co co co co ol ^j en en o en co co en *. ud en en o o en cj coo o co *njo en o o ro ro -c» co co en o o en to n> ol -b -^ u3^ 4en o en o cu a—t r+ ,-i ^ o en eno o o o d. n ro a> c»j ro (t> o en o o (o o en o o en en en o o o o o 61 lassist newsletteri vol. 3, no. (summer 1979) major realignment o't boundaries. witness the recent emergence of the "regional municipality" which replaced certain counties in the province of ontario. code a geographic area is identified in the m-r census via an assigned code. the introduction of the "standard geographical classification" in 1975 produced a complete revision of these codes in the 1976 census. l 16 ] t of chan the grap i ys i the cont numb bod i e.a. boun mine c ana c i pa with trol no w h i c geog the det a to c pr i a the m en t a gen he the ge s fact hie s m ro i er es . an da r i a b da » i b in p sin h de r aph requ i i onsu t e p re sp ce c i es i den se i s h th i eve ay i j ur i of of fo c e e s a y whe ound ro v i sine g i e s c r i i c c i red one it ubli ec t i part t if i pot ampe at a i o i e sd i c one gove r ex n su s r e stat rea s a r i e n c i a e th do bes hang le i s the c at i ve g men t cation en t i a i red by geof anawit h i n t i on a l of a rnm ent amp i e t tract deteri s t i c s m u n i s are i c oner e is cum ent these e s in ve i of forced app r oons of ove r ns and d) a.'s.c theor e wou id create area years . corres for en reveal of the f i cat i mit th of equ all c t i s t i c p ro v i a i at i on c i f i ed census fi le become 17] t ical enab i ove unf ponde umera that bou ons e i de i va le a s es . s ca e spe s for area ma s ho ne c e ly» e a com r ort u nee t ion the ndar o n nt if n t u lie: nada e i a i us s f s t er uld ssa r this user to parable census nat e ly » lists areas nature y modi ot peri cat i on nits in stadoe s t abue rsper om the data this y« time: a is conduct years with sus occur 5-year int next major be in 196 sufficient that the contains o demographi and some housing o questions. major c ea eve a mini ring at e r va i . census 1. i to m i n i c n ly the c oues addit r eco ensus ry 10 -centhe the will t is not e ensus core t i ons i ona l nomi c conclusion the -"official lists" of each census identify geographic areas in terms of e. the first sec tion of this paper described the current si tua t i on sur roundi ng the accessibi i ity of the m-r census at canadian un i ve rsities and made a case f or the creation of a centralized census cat a reference service with in the library. the discussion of the problems related to census data use unoer scored the need for a library or information professional to assist the researcher in the ac qu i lassist newsletter, vol. 3. no. 3 (summer 19 79) round ii delphi questionnaire instructions please read carefully in the last issue of the ne w s,j^et.t.ejr you were asked to respond to round i of a delphi study designed to aia our thinking about future developments relating to the distribution and archiving of machine readable data. that questionnaire was completely open-ended and was structured to elicit the ideas and concerns (in a number of categories) of the members of lassist. the responses from round i were usee in the development of the questionnaire presented in this issue -a closed questionnaire consisting of forty-five "event" statements. for this round of the delphi we are asking that you respond to each event statement with three different ratings or "questions" (designated question 1, question 2t and question 3). questions 1 and 3 are seven point rating scales extending from "low" to "high". question 2 requests that you make an estimate of the year in which the event will take place or begin. with these three questions we are attempting to do the followina: 1. assess the importance of each event (question 1). to the members of lassist 2. forecast the approximate date by which the event will occur or begin to change (question 2). 3. estimate the likelihood that the event will actually occur (questions). when the data are evaluated, median ratings on each of the three questions will be used as an estimate of the collective viewpoint of the lassist membership. when using the seven point ratings scales scales should be interpreted as follows: for questions 1 and 3 the 1. very low importance or likelihood. 2. moderately low importance or likelihood. 3. low importance or likelihood. 4. neutral importance or likelihood (a 50/50 chance). 5. high importance or likelihood. 6. mooerately high importance or likelihood. 7. very high importance or likelihood. to use the scales, simply circle the number which best approximates your viewpoint concerning the event statement. our primary concern is in forecasting directions in the nineteen eighties so your estimates of aates shoulo generally be in the time period 1980 1990. it may be, however, that for a particular event you believe that the oate will be later than 1990 -if so, indicate the date. if you think the event will never happen, enter "5999" as the page t question 1 question 2 how important is about what year will this event to the this event take delivery of service? place or begin? question 3 what is the likelihood this event will take place? key key 1 loi=low importance 3 lol=low likelihood hii=high importance hil=high likelihood 1. a significant increase in the use of oroprietary restrictions on the dissemination of data. loi hii 12 3-4567 your estimate ( ) lol hil 12 3 4 5 6 7 2. expanded use of restrictions (copyrights* restrictive contracts) on the use of software usea for archival and analytical purposes. loi hii lol hil your estimate1234567 ( ) 1234567 3. an expansion of the use of "non-standard" formats for disseminating data. loi hii 12 3 4 5 6 7 lol your estimate ( ) hil 12 3 4 5 6 7 4. a decline in the ability of archivists to promote the use of standardized systems for defining data elements (variables). loi hii 12 3 4 5 6 7 your estimate ( ) lol hil 12 3 4 5 6 7 5. a decline in the ability of archivists to promote the use of standardized systems for classifying and retrieving machine readab le data. loi hii lol hil your estimate1234567 ( ) 1234567 6. a precipitousincrease in the amount of machine readable data. loi hii 12 3 4 5 6 7 your estimate ( ) lol hil 12 3 4 5 6 7 page 6 question 1 question 2 question 3 how important is about what year will what is the likelithis event to the this event take hood this event will delivery of service? place or begin? take place? key 1 loi^low importance key 3 lol=low likelihood hii=high importance hil=high likelihood 13. expansion of storage space requirements for data files. loi hi! 1 2 3.'t 5 6 7 lol hil your estimate ( ) 12 3 4 5 6 7 14. standardization of criteria for evaluating the importance of data files for inclusion in an archive. loi hii 12 3 4 5 6 7 lol hil your estimate ( ) 12 3 4 5 6 7 15. the rising cost of traditional methods for disseminating information (paper» printing) alter methods and techniques of dissemination. loi hii lol hil your estimate1234567 < ) 1234567 16. expansion of the variety of users requiring machine readable data files. loi hii 12 3 4 5 6 7 your estimate ( ) lol hil 12 3 4 5 6 7 17. expansion of commercial archival services for machine readable data. loi hii 12 3 4 5 6 7 your estimate ( ) lol hil 12 3 4 5 6 7 18. development of more varied funding techniques for the development of machine readable data archives. loi hii 12 3 4 5 6 7 your estimate ( ) lol hil 12 3 4 5 6 7 page 8 question 1 how important is this event to the delivery of service? question 2 about whatyear will this event take place or begin? question 3 what is the likelihood this event will take place? key key 1 loi=low importance 3 lol^low likelihood hiuhigh importance hil=high likelihood 25. increased need for access to machine reaaable aata files through technical systems which are transparent to the end user. loi hii lol hil your estimate12i't567 ( ) 1234567 26. expanded training of potential end users concerning the availability and use of machine readaole data files. loi hii 12 3 4 5 6 7 your estimate ( ) hil 12 3 4 5 6 7 27. dissemination of data files through networks rather than by tapes or other magnetic media. loi hii 12 3 4 5 6 7 lol your estimate ( ) hil 12 3 4 5 6 7 28. establishment of centralized and hierarchical data networks to improve access and to reduce cost of access to machine readaole data. loi hii lol hil your estimate1234567 ( ) 1234567 29. expanded use of microfilm or microfiche techniques in the dissemination of documentation and other relevant informat ion. loi hii lol hil your estimate1234567 ( ) 1234567 30. use of cassette tapes or diskettes used in conjunction with microcomputers (or with larger mainframes) for the a i s sem i na t i on of information. loi hii lol hil ycur estimate1234567 ( ) 1234567 page 10 question 1 question 2 question 3 hou important is about what year will what is the likelithis event to the this event take hood this event will delivery of service? place or begin? take place? key key 1 loi=lou importance 3 lol=low likelihood hii=high importance hil=high likelihood 37. greater concern for problems oi human engineering in the production of equipment designee! for data retrieval. loi hii 1 2 3 't 5 6 7 your estimate < ) lol hil 12 3 4 5 6 7 38. development of software and hardware for dealing more automatically with natural language materials. loi hii 12 3 4 5 6 7 lol hil your estimate ( ) 12 3 4 5 6 7 39. expanded deployment of f ont -i ndependent readers for converting printed materials into machine readable form. loi hii 12 3 4 5 6 7 lol your estimate ( ) hil 12 3 4 5 6 7 40. expanded concern on the part of machine readable data archivists with commercial and governmental data needs. loi hii 12 3 4 5 6 7 lol your estimate ( ) hil 12 3 4 5 6 7 41. expansion of the scope of lassist to meet the demands of users other than researchers. loi hii 12 3 4 5 6 7 lol your estimate ( ) hil 12 3 4 5 6 7 42. substantial resources devoted to the development of less expensive methods of information delivery. loi hii 12 3 4 5 6 7 lol your estimate ( ) hil 12 3 4 5 6 7 page 12 51« in what region do you reside' c d 1» western europe c d 2. eastern europe c ] 3. canada c ] 4. the united states [ ] 5. other (where) 52. institutional affiliation; c 3 1. academic institution c 3 2. commercial organization c 3 j. governmental agency c ] 4 . independent data archive l 1 5. other (what) 53. regardless of the type of agency for which you work» are you employed within a department which has: c 1 1. primary objectives other than data archiving? c ] 2. does primarily data archiving and servicing? 54*55. how many employees devote substantial amounts of time to data archiving? c j comment s lassist newsletter, vol. 3, no. 3 (summer 1979) official 117. statistics canaoa list, 1971. statistics canada. correseoqaina kllkzllkk equmsnaiifin area numbers, 65 iassist quarterly vol 22 no2 4 iassist quarterly establishing a data resource centre experiences at the university of guelph by bo wandschneider & doug horne * introduction: the following paper outlines the process of establishing a data resource centre (drc)1 . the paper documents the experiences at the university of guelph, where such a service was established from scratch, and where gains have been made relatively quickly. prior to the fall of 1996 guelph was in a situation similar to many other research/teaching institutions. there were no formal procedures in place for acquiring, distributing and analyzing data in an electronic format. it was the responsibility of individual faculty, researchers and students to develop the necessary skills to make use of data, there was limited statistical support, and overlap existed in acquiring data resources. all of this resulted in duplication of effort with respect to the use of electronic information on campus. it is hoped that an account of these experiences can be of use to others currently in the process of establishing a drc, as well as those considering undertaking such an endeavour. to that end, this paper is written in an easy to follow manner, with limited technical details. certain goals and objectives were set, and this paper looks at how these goals are being achieved and some of the obstacles encountered. issues such as motivation, targeted audience, teaching needs, research needs, levels of service, staffing, hardware, software, security, and delivery tools will be discussed. establishing a data resource centre (drc) can be a very complicated process. for the people involved in the front line delivery and use of electronic information, the needs and benefits have previously been laid out and shown to be substantial2 . the challenge to the manager of a data centre is to express these benefits in such a way that the administrators who control funds see the need to commit scarce resources to this type of service. background: the university of guelph is a major research institution in canada, with approximately $81 million dollars in research grants per year. the undergraduate enrollment is approximately 10,500 students with another 2,000+ graduate students. the university is broken up into six separate colleges including the college of applied and human sciences, ontario agricultural college, ontario veterinary college, college of biological sciences, college of physical and engineering sciences and the college of arts. recently several new remote campuses were added that deal specifically with agriculture. there are also significant ties with the ontario ministry of agriculture and rural affairs. the omafra head office was recently relocated to guelph. at the moment, the biggest users of drc services seem to come from the first two colleges, although clients are spread amongst all groups. the nature of drc holdings and background of staff members has lead to heavier use from departments such as economics, geography, sociology, rural planning and development, agricultural economics and consumer studies. as the service grows and new contacts are made, and a new population of users is developed, it is expected that this will change. it has been found that one of the most important tasks, and a possible problem, is informing the user community about what is being done. experience to this point has been that once contact is made with people, and there is an actual demonstration of the capabilities of the drc, response is extremely positive. the process of establishing a drc at the university of guelph actually began in the fall of 1993. the pilot project did not begin until december 1996. in order to get support for the project a detailed proposal was written, outlining all of the possible options for running a drc. a great deal of this information was developed from a workshop offered during the summer program at icpsr3 . this was augmented with tours of the university of toronto data library and chass facilities, university of western ontario’s social science data centre, and input from the canadian association of public data users (capdu). the proposal was very detailed, defining the users, the benefits, the possible levels of service, considerations for a suitable computing environment, the departments capable of managing the service, the costs, staffing, and what other institutions in canada were doing. this document was used as background information to justify and explain options of how a drc could be run. summer 1998 5 until this point in time, individual faculty members had been responsible for managing their own data needs. this usually entailed hiring a research assistant and spending time and resources getting these individuals ‘tooled-up’ to using whatever data they needed4 . the net result was that there was frequent duplication of efforts, especially for major data sets, and there was a great deal of frustration for many researchers. the ideas presented in the initial proposal were well received by various groups and individuals on campus. generally the response was that such a service had been needed for a long time and would be of use and welcome on campus. the problem was to not only find the resources to get this project to run, but to find sufficient resources to make it function properly. after a long period of preparation and waiting for the right circumstances, it was clear to those involved in the project that initial impressions of the functioning service must be positive. if the service lacked support from the beginning, and did not manage to impress the various stakeholders, it would be very difficult to attract and maintain the client-base needed to develop a commitment in the long term. the pilot project, it was clear, was going to be an all or nothing situation. it took approximately four years to move from the initial ideas to the start of the actual pilot, during which there was a great deal of lobbying done by the interested parties. in the mid 1990’s the university was developing a detailed strategic plan and the need to address electronic information was included as a small paragraph in the plan. this was an important step, as it became clear that the idea of a data centre was becoming central enough to the university’s plans that it was being discussed in various high-level committees. partly in response to this, a beta version of the web retrieval system was developed over a few days5 . this was extremely useful for presentations, and faculty were able to clearly see the potential and the possible applications of this system. with a working prototype in place, it was much easier to generate enthusiasm about what was being proposed, and a number of presentations to potential users where very successful during this period. this was also about the time that the data liberation initiative was being established by statistics canada, with the basic idea of making the data more available and universally accessible. it was very clear that the university needed something like a drc to take advantage of the opportunities that were being presented. at the same time there was a movement at other institutions in canada to establish data centres. all of these factors helped to make the data centre seem like a feasible, and particularly timely, project. the initial proposal was a for a collaborative effort between the library, computing and communications services and the college of social science. the basic resources being committed included the secondment of staff, infrastructure money (which was very limited), and physical space to house the centre. central computing facilities such as a unix system and software, already in place, were also used. the collaboration between the library and computing services (the college of social sciences dropped out early on) was very important in that it brought together a diversity of skills that is still reflected in our current staff. our planning had always taken into account that useful skills could be drawn from the computing and library fields, and that input from the user community is vital. the latter includes the group that uses the information, be it researchers or teachers. as pointed out by kroeker (1997) you need to know your patrons and what their needs are. figure 1 gives an example of the environment that existed prior to the establishment of the drc. the university can be simplified by dividing everyone into 4 groups. assume students gain access through any of the four groups. these 4 groups communicate in an ad-hoc fashion, as depicted by the dashed lines. note that there is no direct communication between type a and type b researchersteachers. the distinction between the two groups is that type b have direct access to incoming data. this may be due to certain skills they posses, or their access to resources needed to acquire this data. the problem lies in the fact that data flows into either the library, or type b researchers/ teachers. it is not clear where data will finally reside and the information flow related to data is poor. this is especially true between group a and b. data management university of guelph prior to drc teachers/ researchers a incoming data figure 1 computing services library teachers/ researchers b 6 iassist quarterly services a paper by jacobs (1991) gives a very good outline of the needs and ways to deliver data. jacobs lists users expectations and the services associated with these expectations, breaking it down into general library services, references services and computing services early on in the proposal stage there was a need to clearly define the services that were going to be offered at guelph. figure 2 outlines different levels of service associated with a data library (see ruus (1990)). basically, the drc was prepared to undertake most aspects associated with collection care and user services. however, there was no commitment to archiving services. over the first year this decision was reconsidered. it has been found that there is a demand for these services and in the summer of 1997 there was a grant to hire a graduate student to begin archiving historical census records from 19th century canada. it is believed that the drc will continue to evolve in this direction. another interesting development that occurred was related to the user community. initially it was believed that the heaviest users would be faculty, researchers and graduate students. applications related to teaching, particularly in applied undergraduate courses such as statistics and upper year research courses, have made use of the drc. this tied in nicely with the services being provided in the government documents section of the library, where the traditional paper-based statistical sources are housed, and where many of the supporting documents that drc users would require could be found. it was suspected that there would still be a large number of ‘one-off’ type questions that would be more efficiently answered using the traditional hard-copy sources. for example, questions like the population for a given cd, csd, or ea might best be addressed using traditional sources. in these cases the user was often looking for one number, or even just a few numbers. however, the ease and flexibility of the web retrieval system has opened up the drc for these types of questions. one of the challenges has been deciding when users should refer to the drc in order to get the quickest and most efficient response and when they should refer to traditional sources. the drc is centered around the www6. expectations are that users will become self-sufficient in finding and extracting information. this is similar to the objectives of other data centres (see kroeker (1997)). the feeling was that if we established a drc, demand was so high that we could easily spend all of our human resources answering requests without expanding the information available. initially the drc did not have any public hours for walk-in consultation, and once it did, these hours were limited. this allowed staff to get a jump on providing selfhelp information for users, get comfortable with the drc’s services themselves, and have some time for some early fine-tuning of the service before declaring the service fully major functions of a data library identification acquistion storage verfication documentation collection care location indexation cataloging system file generation subsetting special purpose software general user documenation consultation orientation training (user and staff) user services cleaning inventorying archiveing promoting national standards inter institutional cooperation archives figure 2 summer 1998 7 functional. the drc web site has gone through several iterations since it appeared in the first week of operations having, at that point, been created with minimal content behind it. it has developed “on-the-fly” in response to demand, feedback, and experiences in a live web environment, with the guiding principle being that web sites should be dynamic, and regularly edited to fit the changing demands or the latest ideas of users or staff. recently there was a major overhaul to simplify the site. most of the changes were a direct result of user input, and studying other sites on the net. the main page links the user to 5 major areas: the first area takes users to the on-line data holdings, which includes access to the web retrieval system (discussed later), cd-rom products available over the net, access to gis data from the census and any on-line services that are subscribed to, such as cansim. the second area takes the user to information and links to all the cd-rom holdings. if they have not been made available over the net then there are instructions on how to access the information. the third area links the user to information on data from consortia agreements which include icpsr and dli. there are direct links to sites where the user can search holdings. only a small fraction of this data is stored locally and is usually obtained on a request basis. the fourth area deals with external data sources. as time permits staff gather links to sites that provide free access to data (some commercial sites are also linked) and logically order these links to help users find what they need. this is also a valuable resource for reference staff within the drc. more often than not this list is expanded as sites are discovered on routine reference questions. the page is divided into categories such as economics, agriculture, science and others. the final area links the user to other data centres and data related sites that may be useful. the office of the drc is located adjacent to the government documents section of the library. staff in the centre have workstations and desk space within this office, aside from the librarian who has his own office in this area. there is also one additional workstation that users have access to, if needed. the idea is that this is a point of reference within the library where users can contact staff in person or by phone. there is also a work area to hold meetings and discuss data related problems. there are also computer pools within the library to which users can be directed to work on their problems, and recently three more workstations were added just outside the drc that can be used by clients (they double as library catalogue machines). the web server for the drc is split into two components. the main portion runs on an nt server located in the drc office. this server stores data that come with a prewritten pc/windows interface7 . this server also stores cd-rom’s that can’t be loaded onto the www. there is a cd-rom tower attached to the server. the bulk of the data holdings, in terms of both size and number of data sets resides on a central unix server that is used for running statistical applications. this server was already in existence,figure 3 library teachers/ researchers a teachers/ researchers b computing services drc ashton lab incoming data data management univiversity of guleph after drc 8 iassist quarterly and the drc web retrieval system runs along with other services. this makes it very convenient for more experienced users to by-pass the web retrieval system and work directly on the data with centrally maintained software. this will be discussed more later. pilot project under the pilot project, a format similar to the one in figure 3 was undertaken. essentially, the drc would be a service offered through the library, and as such, was physically located in the main library. there was direct communication with all researchers, ccs and the library. all data would be channeled through the drc, so that everyone was aware of what was available and where to find it.. the ashton lab already existed, providing advanced consultation with researchers dealing with data collected in the field. clients also have access to sas/ spps help through central computing services. the drc does not provide these services. staff for the drc were assigned from both the library and ccs. for a more detailed break-down of tasks, refer to appendix a. currently, the drc has a full-time systems analyst assigned by ccs as project leader. this person is essentially responsible for coordinating the project and participates in all aspects of the drc, including interaction with researchers using larger data sets available through the drc. ccs has also supplied a 0.5 fte systems analyst, whose main task is managing and writing the web retrieval system. the library has assigned a 0.5 fte librarian who has experience in government documents and working with cd-roms. this person coordinates the drc within the library and handles all cd-rom issues. the library has also assigned a 0.5 fte library associate with experience in the government documents section. this person functions as a reference person for clients coming into the drc, and helps to bridge the gap between the traditional collection and the drc collection. these assignments are best case situations. there is rarely 2.5 fte’s available in any given week. on occasion funds are secured for students to work on specific applications such as 1991 census gis files, historical census, and hife files . web retrieval system the system a large portion of the efforts in the drc are centered around the development of a web retrieval system . a perl script has been developed to provide a web-based interface with sas. this allows an enormous variety of data to be easily mounted, distributed and analyzed on-line. in the 14 months since the first iteration of the script, over 200 surveys have been mounted and made available. there are many objectives in running a www interface to the data. the interface allows simple point and click access to a variety of data sets. users are able to select a subset of variables, draw a sample based on conditions of certain predefined variables8 , output to approximately 30 different formats9 and perform simple statistics on their subsets. at this point this includes frequencies, crosstabs, means and simple regressions. however, most importantly the interface is consistent across all the data sets. several data suppliers are developing very good interfaces to their own data. one of the problems with this arrangement is that there is always a cost, no matter how intuitive, to learning a new interface. users seem to appreciate the consistency and the ability to easily move from one data set to another that is provided with our single web interface. in addition to the above features the retrieval system also gives users access to electronic codebooks, record layouts, users guides, sas contents files10, and sample sas and spss programs. these programs can be transferred to the user’s own pc or central unix account. these programs can be used as a template on the user’s own unix account to run and read the data as defined by them. the system also points the user to the raw data files, the sas datasets, the associated format and index files. all of the data is stored in compressed sas datasets that are fully readable by anyone with a central unix account. variables in these data sets normally have labels and many have associated value statements. the degree of completeness depends on what is available from the supplier of the data. there is a large variety in dli and icpsr data. staff are currently in the process of indexing the larger data files to significantly improve retrieval times. as with the subset variables it is not efficient to index by every possible term , so an attempt is made to try and choose the 2 to 6 variables that satisfy 90% of the requests. currently the raw data file is also stored with the sas data set for users who prefer to write their own programs. disk space is relatively cheap, but there may be constraints in the future. the users essentially the script serves two types of clients. the first are those individuals who simply want a table or even a single number and are not willing to wait or perform a very complex procedure to get output. in many instances it may be faster to look up the table or number in a hard copy source. it must be kept in mind, however, that the webretrieval system has a seemingly infinite number of custom tables available, whereas in hard copy the number, and nature of tables made available is decided upon by the publisher. the second type of user is the researcher, who is looking for data on which to perform some analysis11. experiences at guelph suggest that many researchers have a problem when dealing with empirical work. there is an initial cost to getting a feel for the data, figuring out the record lay-out, writing a program to read in relevant data, and then performing some simple summary statistics on the data. summer 1998 9 once this is done, and the data is sufficiently massaged into a format they can use, the more detailed analysis begins. the drc concentrates on helping with the first part of this process. the web interface makes it extremely easy for anyone to go in and ‘play around’ with the data before they decide what they want. essentially there are now decreased costs, which allows more time for the more complicated analysis, as well as expanding what the user may attempt. once the detailed analysis begins, staff forward problems to the data analysis support group or the ashton statistical laboratory. the administrator from an administrative point of view the system is extremely simple and flexible. one script handles all the different formats of data. when a data set is added to the system a series of form files are created containing information on available variables, labels, variables to subset by, and their possible values. these files can be easily generated from the sas program and contents file using a variety of simple editor commands. once created, information such as directory location, data name, weights, whether to use scroll, input, or pull downs, and how to place boxes on the screen is quickly added. there is no modification of the perl script necessary to deal with data sets. data such as the 1992 famex survey have been added in as little as 1.5 hours. this includes downloading the data, creating the sas dataset, and mounting it on the web. in terms of security, access to the system is protected by ip address at the directory level. the use of apache software to drive the web server allows us to place .htaccess files in any directory to limit or give access to users. the perl script is also capable of checking ip addresses based on the data being downloaded. these options allow us to open and close access to different data sets as the need arises. this will be particularly useful when we move to a shared system among universities with differing licencing arrangements. an area that is still being developed is the search capabilities. this was intentionally avoided during the first year of development as the priority was to give access to as much data as possible. recently, the ability to do searches by keyword and strings was enabled on the contents files. the results of the search are presented in a tabular format where the user is linked to the contents files and the associated ‘readme’ file for that data set. shortly there will be a link directly to the web retrieval forms for this data set. the files that are searched will also be expanded. integration with library one of the major objectives of the drc was to integrate data identification and retrieval services with other services already in the library. the web-based nature of this service means that every workstation in the library (and on campus) is a potential contact point with the drc, and library staff will be faced with data that is integrated with the other, traditional, library services. the fact that many reference staff in the library have had little or no experience with data means that training will be a time-consuming process, and developing a reasonable level of comfort with this type of information will have to occur gradually. this is an entirely new resource for the reference staff member to consider, and education will involve not only instruction in statistics, but more basically the understanding of the appropriate uses of data. the obvious place to start training was with staff in the government documents section, where staff had experience with the paper-based version of the data. there have been a few general sessions for library staff on what happens in the drc and what data is available, with more detailed sessions planned for the future. a great deal of emphasis has been placed on the notion that we are trying to make users self-sufficient, and to minimize the need for consultation for relatively straightforward queries. training sessions tend to consist of giving general outlines, asking staff to try to use the drc web pages, and then to come back to us with questions. with lots of hand-holding this seems to be working, although, as is often the case, the technology itself rather than the nature of the data causes problems. the feed-back from these staff members is extremely useful in helping set a direction. in other words, we get the users to tell us what works and what doesn’t work. training will slowly move into more detailed and specific sessions as services are better defined. up until this point the drc has been evolving rapidly and changes are frequent. the general view is that there will always be a need for specialized assistance that probably will not reside in the general reference staff of a library. the first year success and areas for improvement overall the drc has been extremely successful and in most areas we have progressed well beyond expectations. there was a lot of uncertainty with respect to how we were going to do things, but everything seemed to fall into place very well. this can be attributed to a few things. the first was the level of technology in terms of software. perl, sas, dbms copy and apache delivered what was needed. the second related to the centralized computing environment that existed at guelph. although there were some rough spots associated with system security and access to configuration information by drc staff, the resources were again well suited for what was needed. delivering the data from a centralized unix system, accessible by everyone on campus, is extremely efficient. the final point was related to the staff involved. as mentioned earlier, it is important to have individuals from the computing fields, the library, and the user community, working together. the drc was lucky to get a group of 10 iassist quarterly people who not only technically complimented each other, but also worked very well together. one of the biggest surprises has been how easy it is for administrators to add data to the retrieval system. the way the perl script has been written allows for data in a wide variety of formats to be easily mounted for retrieval. at best, expectations were that a few dozen data sets could be mounted during the first year and it would be difficult to determine which ones. currently there are well over 200 data sets on the system and the effort necessary to mount these diminishes as staff comfort levels increase . it takes between 1 hour and 2 days (large multiple-file data sets) to prepare data. as we begin moving back in time and mounting ‘old’ data, that lacks sas/spss code and doesn’t have electronic codebooks, the process gets more difficult. it is so easy, we are starting to train library staff with limited sas and unix skills to mount data. another surprise has been the speed of extraction. the current system runs on a fairly slow unix server. however, by indexing the data the response time can be improved significantly. the functionality of the retrieval system, in terms of output formats, and summary statistics, is also beyond expectations. some of the areas we still need to improve are related to publicity. it is still difficult to get people to understand what is being done in the drc. as soon as we get 1 on 1 contact, or demonstrate the system to a class, it becomes easy, and the users become largely self-sufficient. a goal is to get out to classes more often, and to make the drc a standard tool for the completion of assignments. we are also continuing to publish a newsletter each semester outlining developments and what people are doing. the service is also not without it’s detractors. some researchers who are experienced with working on large data sets, and using sas, spss do not always see the benefits. it is very difficult getting these users to understand that the structure of the whole system can help even those who want to do their own extractions12. in some cases we spend more time with these experienced users than with ‘new’ users who readily accept the system. what is next? as mentioned above, we are progressing much faster than we initially expected, but there are several things that still need to be worked on. some of these were mentioned in section 7, related to publicity and staff training. other areas include increasing the functionality of the web retrieval system, expanding on-line statistical options, particularly related to graphing, and possibly interfacing better with gis systems. work also needs to be done with on-line keyword searching for information. the current system works well if you know what data you are interested in. one area in which much consultation is still needed, however, is in the identification of data to suit a query. it is hoped that an efficient search capability for the system may also make the selection of data possible for the inexperienced user. we have started in this direction and hope to make progress over the summer. possibly the biggest project to date will be the sharing of this resource with other institutions. the web-based system is ideal for taking advantage of the potential efficiencies of a joint service. discussions are well under way to develop this service into a seamless, shared resource between the university of guelph, university of waterloo and wilfrid laurier university. the end result will be a better service for all parties involved. it his hoped that eventually other such centres will appear in canada and the workload and overhead can be shared between even more institutions. appendix a brief job descriptions project leader 1 fte systems analyst, ccs tasks: coordination of project; liaison between ccs, library and user community; participate in management group; report to sac; periodically report to college it committees; liaison with outside groups such as capdu, dli, icpsr, statistics canada; coordinate joint ventures with wlu and waterloo; control inflow of data from outside sources ie download from dli and icpsr; participate in data purchase agreements and consortiums (work with other data centers on national issues); provide user support and consulting for staff, researchers and students; participate in development of front end applications; participate in production of newsletter and annual report; assist other drc staff as needed. www resource person .5 fte systems analyst, ccs tasks: incorporate, develop and maintain web retrieval interfaces for electronic data resources; implement a process for controlling access to data; implement a process for measuring usage; limited user support and consulting services for end users; manage nt server and various data products summer 1998 11 librarian .5 fte -library tasks: overall coordination of “library specific” side of project; user support and consultation; addition of data to web site, coordinate cd-rom products and acquisition; establish and develop communication between drc and the rest of the library; participate in planning of layout and design of service point; develop and participate in training of library staff (classes and production materials); work with library staff on publicity and information; work with acquisitions staff to bring data acquisition in line with acquisition of other library materials; work with cataloguing staff to develop workable method of cataloguing electronic data sets; keep up to date with development of data resources available on the internet. library associate .5 fte carol perry library tasks: link between ‘hard-copy’ reference and drc collection; user support and consultation; adding data to web site; publicity (newsletter); backup strategies; web design and graphics; www searching and inventory (collect resources, data sites, and useful sources of related information); participate in training of library staff; maintain hard copy collection of codebooks; keep statistics of patron traffic and data use. references horne, d., and mccaskell, p., and wandschneider, b. (1996), university of guelph data library/centre proposal, mimeograph, university of guelph, september 1996. (http:drc.uoguelph.ca/ jacobs, j. (1991), “providing data services for machinereadible information in an academic library: some levels of service, public-access computer systems review, vol. 1, issue 2, 144-160. kroeker, b. (1997), data services assists teaching and research: delivery of data services via the world wide web, iassist quarterly, vol. 21, number 1, spring 1997. lubanski, a. (1996), “ social science data services during the last five years of the millenium”, iassist quarterly, vol. 20, number 4, november 1996. ruus, l.,g.,m. (1990), “planning a data service facilty”, mimeograph,university of toronto. 1 in this paper we define a drc as a point of service where users are able to get access to, and assistance with the use of, electronic information such as the census, general social survey, survey of consumer finances and so on. we do not include resources such as electronic journals and books. experience suggests that users are frequently confused about this distinction. 2 it is felt that the benefits are well defined and understood, and as such will not be discussed in detail. the emphasis will be on the process rather than the justification. however, it is essential to have a thorough understanding of them and for more detail see horne et al. (1996). lubanski (1996) gives some justification for increased funding, highlighting the increases in empirical research. 3 the workshop was put on by diane geraci (suny binghampton), chuck humphrey (university of alberta) and jim jacobs (uc san diego). 4 during a presentation of the web retrieval system at another university there was a very positive response in this direction. the faculty member was elated because they could now sit down at their workstation and retrieve subsets of the data in a matter of minutes. up to this point they would spend valuable resources on graduate students and many times they ended up with nothing at the end of the project. 5 the basic ideas for this were obtained from sample perl scripts and sas programs written at kansas. they were available over the www. this was a very crude implementation of a interface between sas and the www. 6 see http://drc.uoguelph.ca 7 many of statistic canada’s data products now come in a format that uses ivision’s beyond 20/20 browser. we find this a very useful interface for many queries, but stress it is only a compliment to our web retrieval system. it is not an efficient format to deal with may research type questions that are posed in the academic environment. 8 staff decided early in the process to keep the retrieval forms as simple as possible. as such, each data set has a small pre-determined set of variables to subset by. in the case of time series, this would include months and years, whereas in the case of most cross-sections something like age, region, education and gender are chosen. the objective is to try and satisfy the most possible requests with the smallest, most commonly used subset variables. if a user determines they would like to subset by some other variable, such as income, this can easily be added. the way the script is written it takes about 2 minutes to add another subset variable. the necessary ‘form’ file is created from the ‘values’ section of the sas program. 9 examples include: sas, spss, spps portable, gauss, http://drc.uoguelph.ca/ http://drc.uoguelph.ca/ 12 iassist quarterly stata, lotus, excel, quattro, dbase, ascii and many others. 10 these files are used to easily access information on variable names, labels and formats, as well as sample size. 11 this researcher could be very experienced and not need much assistance or they could be a student in an applied course who has never worked with this type of data before. in the later case the web retrieval system allows the user to easily (and painlessly) obtain the data and move on to more important tasks, central to the course. 12 for example; codebooks, user guides, record layouts, sas and spss code ready to read the data and even the data already in sas data sets with formats and labels are all available through the web site. * paper presented at iassist 1998 yale university may 1998. bo wandschneider systems analyst data resource centre university of guelph, doug horne librarian data resource centre university of guelph. http://www.yorku.ca/org/iassist/ a user's perspective on electronic data archival: tlie importance of standards by annette jones watters ' and carl e. ferguson, jr. the university ofalabama revolution probably wins the prize for the most overused characterization of rapidly changing non-violent events. however, few words better characterize the rapid rise of the microcomputer as the dominant technology of the 1980s. the device has revolutionized the workplace, bringing the power of electronic digital computing lo the desktop. and, with speed and storage capacity increasing extremely rapidly, the microcomputer continues to transform every task associated with the acquisition, maintenance, and use of information. this paper offers a brief look at the impact of the microcomputer revolution on data distribution and archiving standards, then attempts to chart current trends and conditions in the rapidly changing technological landscape. historical perspective digital document archival has traditionally been directed by considerations of space and convenience.^ space was a consideration because traditional library or reference facilities simply could not accommodate copies of all the historical information. although magnetic tape offered relative high storage densities, even tape storage quickly became problematic. the space savings achieved by going from tape canisters stored in wire racks to hanging tape seals was quite significant however, every unit eventually ran out of room—no one could ever buy enough tape cabinets. convenience. seldom used historical data could not compete successfully with current information for shelfspace. as a result, data progenitors and librarians soon developed usage rules to help establish shelf-life and retention standards for data sets. the limits ofspace and accessibility limited space dictated that out-of-date items be compressed and/or relegated to less expensive (albeit less accessible) mediums. numeric data was frequently transcribed from a character format (typically ascii or ebcdic) to a much more dense binary format. the resulting files were then written to the highest density magnetic tapes available. standard tabulations based on these data were candidates for microfiche, 35mm microfilm, or paper microform products. binary tapes and microform products do offer significant storage densities. however, retrieval has always been a tiresome process. in all cases, accessibility and space savings were the primary considerations and the end-user was frequently the loser. the end-user usually played little or no direct role in the determining the method of compression or archival.^ indeed, the end-user generally worked through an intermediary who selected the archival strategy. that strategy frequendy was not based on the needs or retrieval skills of the end-user. a critical aspect of these archival strategies was the skills and tools available to the person archiving the data, not the skills and tools of the researcher. archival responsibilities frequently fell to computer programmers who had no sense of the practical value of the data involved. machine time prior to the advent of the microcomputer, most data analysts worked with paper products developed and maintained by a group of modem-day alchemists called the programmers. working patiently, the analyst communicated the nature of the application to the wizard, who with cards in hand communicated with the machine. this was a most serious relationship, for usually there was only one machine in the organization — one computer to be used by all. competition for its time and attention could be intense. trial and error was expensive. research strategies requiring alternative methods of analysis were expensive. researchers were allocated a limited amount of machine time and they learned patience. they conceptualized the table or statistical procedure to be run, gave it to the programmer, and waited. two or three turnarounds in the morning— maybe the same in the afternoon — meant that they did not spend too much time trying alternative methods or procedures. with the development of the statistical packages, spss, sas and others, the role of programmer as the analyst's interpreter began to fade— though not yet disappear. the canned packages greatly facilitated the analysis of these data and they confronted the analyst with a new challenge. the intellectual cost to the analyst in time and commitment to develop programming skills in fortran or some other higher level language was almost always viewed as excessive. however, the statistical packages were different. using surprisingly few procedural commands the analyst could now read the data and spring/summer 1992 actually do the analysis. most researchers immediately realized how much time could be saved by skipping the intermediate step of using a computer programmer. that time might now be used to explore alternative methods or forms of analysis. consequence it was the limits imposed by space, time, and accessibility, that profoundly directed the data distribution and archival standards of the 60s and 70s. programmers working for data progenitors prepared distribution tapes for programmers working for data analysts. and, programmers chose the archival standards and formats for the day that someone would want to look at old data sets. end-users rarely read tapes. end-users worked through programmers and it was the programmers who decided how they would communicate with one another. the evolution of desktop computing while apple and others were offering microcomputers in the late 1970s, the introduction of the ibm personal computer (pc) must be regarded as the beginning of the workplace revolution. ibm's entry into the market gave the microcomputer credibility. it was no longer a toy or experimental device for hobbyists— it was made for work and from the first day it began to recreate the office. one could say, "and the rest is history!", but there is too much to be learned from this transformation of the workplace to move on too quickly. at first office workers were given a machine, and little else. many quickly learned two new words — hardware and software. they learned that without software the hardware did not do very much! software these microcomputers were fast and could remember things! and they did like numbers. however, they were business machines— they liked documents and numbers. the numbers they liked best were of the financial variety (spreadsheets) and the documents were correspondence. lotus, microsoft word, and others quickly found their way into the market— and the world would never be the same. hardware as more software and data applications became available, the 10mb hard disk quickly filled up." although the earliest microcomputer chip— the 8086— was fast, more complex applications quickly called for more speed. today, the fastest machine uses an 80486 d^itel chip running at 50mhz, 8 mb of ram, and is typically packaged with a 350mb hard disk.' such a machine will operate hundreds of times faster than the original 8086 and offers more total computing power than large mainframe computers of less than a decade ago. socialization by the end of the decade the vlsi (vary large scale integration) sihcon chip, that thumb nail sized computer, could be found in every office and on almost every desk. it was no longer a curio down the hall but rather an extension of the worker. in the decade of the 1980s it was ok for men to type and for senior executives to get their hands dirty with data. scientists captured data via analog ports while specialized software, running in the background, conducted the analysis in real-time.'' survey research introduced cai. spss, sas, and bmd for the pc were not far behind.' never before could the analyst get so close to so much data— manipulate it, manage it, analyze, and interpret it. microcomputer based analysis and text processing (eventually to be called desktop publishing) skills were fast becoming an integral part of every data user's personal skill set. whether the analyst was a social scientist, music historian, paleontologist, or greek mythologist, the power of the micro was sweeter than the songs of the sirens. user groups, first formed to provide aid and comfort to practitioners of the infant technology, disappeared as help became available from the officemate next door. power users began talking to software and hardware developers, offering (frequently demanding) new features, more power, more speed. and, as the size of the market continued to grow, the software developers listened; their craft was now a multi-billion dollar business. and so, what has become of the programmer? who now sets the distribution standards? what has become of the limits of space and time? microcomputer hardware, application software, and enhanced user skills have dramatically altered the traditional role of the programmer data analyst. the installed base of ms dos microcomputers is now measured in the hundreds of millions and the market potential for a good applications software package can quickly exceed a million dollars.' microcomputer applications software developers have atu^acted exceptionally bright and creative systems designers and programmers with training and interests in many functional fields. as a result, researchers now have computer based tools unimaginable less than a decade ago. standards the uses of standards electronic data distribution and archival standards serve the user community in several way. standards promote ease of communication among users and between users and data providers; {assist quarleriy equitable global access to data opportunities through improved documentation and communication environments; and the convergence of distribution and archival media. ease of communications between data users and data provides is critical to both analyst and provider. frequent providers include governments (national, state and local), universities and other research organizations, and businesses. while each has a unique mission in our society, as data providers, they and their user community can benefit from improved communications— improvements through mutually agreed upon standards. the user community, public and private social and physical science researchers and analysts, share in this responsibility. all too frequently, a me versus them mentality sets in. if there were a common understanding of the technical standard for providing data and common understanding of what is reasonable for the end-user lo bring to the table, the level of antagonism would be reduced. these standards of expectation do not now exist. improved, jointly developed standards, are a major step toward equitable global access. global communications today is no more exotic then a hard-wire link to your local mainframe in an adjacent building. however, to be most useful, data providers and user worldwide must work closely together lo insure not just interagency or national standards but rather international (universal) agreements on media and form. such standards must transcend multiple platforms and operating systems. microsoft dos machines must be able to easily communicate with unix (xenix), macintosh, and others. communication standards are desperately needed to allow word processing (desktop publishing) software to easily share a document and its complete formatting. microsoft, with its rich text file (rtf) concept is offering the market one such standard for consideration. and, of course, the need for standards can be found for spreadsheet, database, cad, and other systems. what has changed is the role of the user. the size and sophistication of the user community both commands the attention of the developers and shares with them the responsibility to develop and adopt global standards in each of these functional areas. poorly written documentation and a lack of common understanding on what documentation is supposed to cover disrupts international data distribution, even without intervening language barriers. the future ofstandards while it may seem contradictory, standards are dynamic. distribution and archival standards will continue to be affected by technological change. for all the progress to date, we have yet lo achieve fully error-free exchange of information, easy retrieval of archived data sets, or interoperability of hardware and software. more technological changes are inevitable. the size and sophistication of the user community, high-density storage media, and an unparalleled apphcalions software development effort have rewritten rules for data standards. and, in the judgment of these authors, it is the combination of the three that has had the greatest impact. the globalization of data uses necessitates communication on the issue of standards. international business, international academic research, and united nations programs are examples of sophisticated uses of data sets requiring new disdibution and archival standards. the integrated european market; the political changes in germany, eastern europe, and the former u.s.s.r.; potential tariff agreement in the western hemisphere; and strengthened copyright laws in the pacific rim countries will accelerate the push for easy communication among data users and promote the development of international standards. hands-on users without a common understanding of distribution and archival standards and principles, forthcoming changes may not all be improvements. as the largest producers of research data, government agencies worldwide face the unprecedented challenge of serving a rapidly growing end-user market now numbering in the millions. today, data end-users number in the tens of millions and all of them expect lo interact directly with the data product. during this period of flux, the development of distribution and archival standards can benefit from input by a knowledgeable user community. and, lassist members are on the leading edge of understanding the need for governments and analysts alike to practice "safe data." as the traditional technical role of the programmer has faded, end-users have acquired a new responsibility for safe-guarding the welfare of their own data resources and distribution channels. yet, it is unclear how much of this responsibility users will co-opt for themselves and how much they will hand back to the professional data processing community. what will be role of the data archivist as we move into the 1990s? how many of the distribution, retrieval, and archiving functions will be done by business professionals, research analysts, and computer professionals? iassist members will be facing these questions headon in the coming years. spnng/summer 1992 selected bibliography chartrand, robert lee, ed. critical issues in the information age. metuchen, nj.: the scarecrow press, inc., 1991. snowhill, lucia, and meszaros, rosemary. "new directions in federal information policy and dissemination" microform review 19:4 (fall 1990): 181-185. u.s. congress. office of technology assessment. critical connections: communication for the future. ota-cit-407, 1990. u.s. congress. office of technology assessment. informing the nation: federal information dissemination in an electronic age. ota-crt-396, 1988. wells, norman e.; chow, ivan, and johnson, linn d. document formal considerationsfor a document tracking and storage system. ntis, 1991. 1 presented at the lassist 92 conference held in madison, wisconsin, u.s.a. may 26 29, 1992. 2 digital documents are defined to be those that reside, are distributed, or primarily archived in digital form — traditionally on magnetic tape. 3 the term archival is used throughout this paper to mean the transition or transformation of digital-data from that form normally associated with daily or regular use to an alternate compressed form intended for less frequent use and/or long-term historical storage. 4 mb is an abbreviation for mega-byte or million bytes (characters) of storage. to put this in context, an average single-spaced page of text contains approximately 3,200 characters. thus, a 10mb hard disk could store the equivalent of approximately 3,100 pages of text. 5 the original 8086 operated at an internal clock speed of approximately 4.7 mhz (4.7 million cycles/second). a 350mb hard disk can hold the data traditionally stored on approximately 10 magnetic tapes recorded at 6250 bpi, the current recording density. 6 scientific experiments now frequently incorporate instrumentation that measure such information as temperature which is automatically captured by pcs as the experiment occurs. other programs running on pc analyze these data continuously as the experiment occurs. 7 cai is an acronym for computer assisted interviewing. spss, sas, bmd are all statistical packages that began on the mainframe and which are now available for use on the pc. 8 ms dos is an acronym for microsoft disk operating system. it is the dominant operating system (control program) used by microcomputers. the chief rival to ms dos is the apple macintosh. lassist quarterly iassist '84 conference report by sue gavrel, president over 90 data archivists, librarians, researchers and other users and creators of machine-readable data attended the iassist '84 conference in ottawa, may 15-18. the week began with a day of pre-conference workshops designed to provide practical information and experience on specific aspects of machine-readable data. workshop topics included: complex data files; microcomputers; data library management; and cataloguing microcomputer data files. the conference opened with an international panel of representatives of statistical agencies from canada, sweden, england and the united states. the topic of the panel was "issues confronting statistical agencies: an international perspective." members of the panel included christopher denham, office of population census and surveys; edmund rapaport, statistics sweden; lome rowebottom, statistics canada; and paul zeisset, bureau of the census. each representative outlined the activities of his respective agency addressing specific problems which are and will be confronted in the next few years. similar issues were addressed and the audience was able to see how different countries were confronting similar issues. the session was very informative and stimulated many questions from both the panel and the audience. the sub-theme of privacy and confidentiality began in the afternoon with a plenary session focusing on the researcher's view of these issues. professor john bossons of the institute for political analysis, university of toronto, addressed the problems experienced in obtaining and using microdata. two discussants, thomas brown, vice-president of the association of public data users and david norton of statistics canada, responded to the issues raised by professor bossons. the plenary session lead into three consecutive sessions, each dealing with a specific aspect of privacy and confidentiality: anonymization techniques, which addressed techniques used by statistical agencies and private companies in their collection and use of data; legislative aspects associated with machine-readable data in an international perspective; and the implications on archives of privacy and confidentiality issues, in particular the effect on distribution and standardization. the second sub-theme, the advance of technology, addressed the advances in statisical software packages and their adaptation to microcomputers. nancy morisson, spss, inc. and gary anderson, sir/dbms, inc. were the major speakers and provided insights into the changes being made in these packages. it was suggested that a useful addition to this type of session would be a user of these packages who could discuss problems encountered in using the software. three concurrent sessions followed, each one dealing with a specific aspect of the topic. "to minis and micros" dealt with the adaptations or approaches required by archivists and researchers in analysing data in hierarchical and relational databases. the "on-line bibliographic systems" session discussed the different approaches that can be taken in develping on-line bibliographic systems. the third sub-theme. changing roles and responsibilities, began with a session on "the growth of the information elite." nancy brodie, national library of canada; joseph paradi , dataline systems limited; and erika -continued conference report continued von brunken, karolinska institute library and information centre, sweden, debated the question of whether a new elite was being created due to the cost of equipment and knowledge in the use of computers. three concurrent sessions followed: "data archives and libraries: new challenges," which assessed the impact of government policies on the establishment and maintenance of data archives; "data collection and use: new directions," which addressed new trends in data processing and their impact on governments, universities and private companies; and finally, the "collection and use of international data," which outlined the problems which the collection of data from cross-national sources poses to both archivists and researchers. lunch hour meetings were held for two groups: data link and the lassist working group on the preparation of an international standard bibliographic description for mrdf. the latter met to discuss a draft working paper and approve reconmendations on changes to specific cataloguing rules. this report will be submitted to the isbd (nbm) review group. discussions carried over into the social activities of the conference. both statistics canada and the public archives sponsored receptions for participants. a 10th anniversary buffet/reception was held on the last night followed by a digital disco. a tour of the city in a double decker bus with a stop at statistics canada to see a cansim demonstration was organized. the week provided an excellent opportunity for participants to meet with colleagues from north america and europe and discuss issues and problems which are confronting the profession. the program and local arrangements committees would like to thank all the participants for their enthusiastic ~^pcoming elections elections will be held this fall, and all paid lassist members will receive information shortly. new officers will assume their positions at next year's conference in amsterdam. the nominations and elections committee is composed of three people: sue gavrel machine readable archives public archives of canada 395 wellington street ottawa, ontario canada kia 0n3 jackie mcgee the rand corporation 1700 main street santa monica, california usa 90406 ekkehard mochmann zentralarchive fur empirische sozialforschung university of cologne bachemer strasse 40 d-5100 cologne 41 federal republic of germany please address questions or comments to this group. participation in the conference. next year's conference will be in amsterdam, may 20-24. the theme of the conference is "public access to public data." vol262 4 iassist quarterly summer 2002 iassist quarterly summer 2002 5 abstract from a research-collaboration perspective, african researchers have many opportunities to learn from the international community. the international association for social science information services and technology (iassist) possesses the resources and intellectual capacity to play a more meaningful role in information sharing and exchange of best practices in quantitative research service delivery. the aim of this paper is to make recommendations with regard to the way in which iassist can bring about closer collaboration between its international members and the african region. closer collaboration could boost the interest of researchers in quantitative research and secondary data analysis, using the new technologies and the services of the existing african national data archives. introduction this paper aims to address closer collaboration between the international association for social science and information service technology (iassist) and the african region. firstly, it gives an overview of the political and socio-economic conditions in africa and how these conditions hamper the development in the region. secondly, it deals with the attempts made at regional stability, and finally the role iassist can play assisting social science data utilisation in africa. political and socio-economic problems hampering connectivity in africa the most familiar landscape painted of africa is one of a poverty-stricken, politically unstable, war-torn continent. this picture is not entirely exaggerated, but it is beginning to show signs of positive change. some african scholars believe that there is a “dawning realisation that the impetus for long-range social and economic development is informed by scientific knowledge, technological innovation, entrepreneurial management, and good leadership to transform the harsh conditions of the continent and how it is perceived by itself and others”. (odhiambo, t.p.2). the african continent has much to learn from the international community with regard to innovation in iassist collaboration in africa by julia paris* the application of scientific expertise and in service delivery. iassist has the intellectual capacity, as well as the global network to bring about closer collaboration in the area of quantitative machine-readable social science data. closer collaboration between iassist and africa implies many challenges. as the second largest continent with an estimated population of 700 million people, africa is marked with many contradictions. although richly endowed with natural resources, it is plagued with abject poverty, economic underdevelopment, and shifting political agendas, those problems are compounded by high rates of illiteracy. articles in the news media constantly highlight the pessimism with which local and foreign journalists paint the african political and economic landscape. africaʼs current situation and prognostications for the future are reported in terms of general desolation, ruthless state brutality, unbridled corruption and social dysfunction. african political leaders are held responsible for selfinflicted failure and poor political leadership prevalent in the region. the political leaders are accused of selfaggrandisement, procrastination, pilfering of state funds, and indulging in empty rhetoric. this allegation seems to find substance in ambitious projects with no economic utility that is to be found in some of the most severely poverty-stricken african countries. a well-known example is the notre dame de la paix basilica in the ivory coast, which was built to rival the grandeur of the st. peterʼs basilica in rome. (africa news, 2001) studies on africa reflect contrasting viewpoints of the current african situation. they sketch a picture of gloom, with occasional rays of hope breaking through as depicted in the following example. studies show that many universities in africa have been in a state of crisis for a long time. this is exacerbated by limited budgets, a shortcoming which leads to low morale among faculty and students. an increased sense of isolation and lagging behind in education makes it difficult for universities to attract and retain top quality professors. too few graduates are produced from programs that are relevant to development, and therefore, too little knowledge related 6 iassist quarterly summer 2002 iassist quarterly summer 2002 7 directly to urgent african needs is generated. on the positive side, in angola, the roman catholic church established a catholic university as a substitute for the public university that had been destroyed in civil war. this enterprise was an attempt to rebuild research infrastructure aimed at restoring access to higher learning. there also seems to be a flicker of hope in the case of the university of sierra leone. this university decided to follow an orderly set of procedures to put them on their way to reform. these procedures include a clear definition of the objectives for reform, a detailed examination of all system elements, and a selection of key elements that need to be influenced. african theorists are of the opinion, however, that only managerial commitment and a sound injection of financial resources will ensure the success and sustainability of this initiative. even in difficult circumstances innovations are emerging in africa to improve the enabling environment for the higher education system. connectivity in africa solving the problems experienced in the african region requires a great deal of scientific investigation. scientific problem solving, however, requires valid and reliable information. africa lacks the research capacity to generate data, which is comparable, consistent, reliable and appropriate in this regard. james dean, director of programmes at the information institute in panos, is of the opinion that africa will not be a fully ʻwired ̓society for a long time. however, there are a few countries on the continent that have done well in promoting technological networking, namely, senegal, zambia, zimbabwe, mali, uganda, mozambique, egypt, kenya and tunisia. the ivory coast used to have connectivity through bitnet, but those links were discontinued due to lack of sustainability planning. noteworthy initiatives are the examples of the kenya computer institute (kci) and the african regional centre for technology information system (arctis) serving 11 african countries. the kci uses electronic mail as a tool for information exchange in the developing world. at present, it is the main means of facilitating collaboration between kenya, some other african states, and the international community. the kci provides communication infrastructure and computing facilities, management support, human resources, sustainability, and motivation. international memberships include the united states, canada, the united kingdom, italy, sweden, turkey, australia, new zealand, and singapore. from the african region, only kenya and uganda are active members. the kci seems to be the largest and fastest growing electronic network of african scientists. this constitutes a unique pool of kenyan expertise that could become the means to national development of information technology application in the region. the institute has experience with cuttingedge and mainstream technology, including tcp/ip (protocols) which provide internet connectivity. the other example is that of arctis. it is an african regional network supported by the united nations development programme (undp) and the international development and research centre (idrc) in canada. this network handles a wide range of numerical and non-numerical, graphical, and full-text technological information relevant to socio-economic development of africa. it is linked with several national and international facilities and information sources, and serves as a network using hardware based on local area network and wide area network configurations. another noteworthy factor is the existence of other small communities of dedicated electronic networks in africa, namely, the non-governmental network (ngonet) and the health network (healthnet) to name a few. the initiatives cited serve to confirm that africa holds great possibilities despite the many challenges facing it. it is with a sense of encouragement that one perceives the perseverance of some african scholars and the donor community support to sustain these initiatives. the documentation of these success stories is important for future benchmarking and follow-up. few as the existing pockets of research networks in the african region may be, there are clear indications that the region is moving towards workable solutions to its vast problems. attempts at regional stability and integration the southern african development community (sadc) several collective attempts at regional stability and integration are underway in the african region. this is evident from initiatives such as the southern african development community, treaty for enhanced east african co-operation, and the economic community of west africa. the main aim of these initiatives is to pursue integration and greater co-operation between economies in the african region. (brümmerhoff, 2001). another objective is to establish a culture of self-sufficiency that could serve as an impetus for social renewal on the continent. the role of south africa in the renewal of the african region as a member of sadc, south africa is also viewed as the economic ʻpowerhouse ̓of the african continent. consequently, it is required to make substantial contributions locally and internationally within the area 6 iassist quarterly summer 2002 iassist quarterly summer 2002 7 of scientific investigation, to assist in the renewal of the continent. it should be borne in mind, however, that south africa is also a society in transition, one that is grappling with serious socio-economic issues like the hiv/aids pandemic, malnutrition among certain population groups, unresolved conflicts with their roots in the apartheid era, an urgent need for economic transformation and a diverse cultural heritage. as a result of these societal demands there is a growing need for social science research data in south africa. this is verifiable through reports generated by the national research foundation in pretoria. the application of social science data in africa. south africa is deemed to be the only african country with a well-established information infrastructure. without going into the dubiousness of its motivation, the previous nationalist party government ensured that south africa had the necessary structures and research institutions to execute research in social development issues. here, the human sciences research council, the medical research council and the water research commission are cited as examples. despite the strong infrastructure, however, there exists a dire need for awareness and demystification of quantitative analysis. a workshop held by the south african data archive in 1997 reflected practical problems experienced by researchers in the analysis and interpretation of data, as well as the manipulation of the data using computerised statistical packages. it has been noted that the rapid advances in computer technology exacerbate these problems. addressing these problems would be beneficial for any future research efforts not only in south africa, but also in the broader region. the national research foundation has identified research focus areas to accommodate the investigation of issues such as hiv/aids, malnutrition, compact resources, etc. as mentioned above to make use of the unique opportunity available to south african social scientists. researching these issues presents an opportunity for south africa to obtain a better understanding of the african situation and the role south africa could play in social and economic development in the region. however, it is not known how many current attempts at capacity building in social science research are addressing these problems. impact of collaboration between iassist and africa in keeping with social, economic and technological development, collaboration and global exchange of information are becoming more important for the development of underdeveloped countries. iassist plays a pivotal role in bringing together social science users and producers of data. the linkage already existing between south africa and iassist could only serve to benefit the whole of the region. south africa, viewed to have the most developed information infrastructure, seems to be in the best position to be the first port of call for strengthening of collaboration between iassist and africa. this could become the benchmark for operation outreach initiatives in the region. the iassist vision and mission already include a collaborative focus. this is found in the outreach objectives and programmes of the association. the objectives are to: • encourage and support the establishment of local and national information centres for social science machine-readable data; • foster international exchange and dissemination of information regarding substantive and technical development related to social science machine-readable data; • co-ordinate international programs and projects and general efforts that provide a forum for discussion of issues relating to social science machine-readable data; promote the development of standards for social science machine-readable data; • encourage educational experiences for personnel engaged in work related to these objectives.” (iassist constitution) these objectives tie in very well with the social science research needs of the african region. recommendations the following list of recommendations serves to provide options for closer working ties between iassist and the african region. the list is neither comprehensive, nor prescriptive. its purpose is to provide iassist and the african regional representation with a starting point for further discussions. the recommendations are: • expand the regional secretariat into an action committee to manage the iassist/ africa affairs in the african region. infrastructural assistance could be sought through strategic discussions with the south african data archive and the national research foundation, as well as institutions of higher learning. • organise an iassist/africa awareness campaign, with an initial focus on south africa, to create awareness of the global network among south african social science individuals. the idea is to target social science graduate students, social scientists and information professionals. this could be done through an iassist/africa marketing pamphlet and a questionnaire to stimulate their awareness of quantitative 8 iassist quarterly summer 2002 iassist quarterly summer 2002 9 research resources and facilities, and of the existence of iassist and possible membership. • create an electronic mailing list to disseminate iassist information, news and upcoming events, with links to the iassist homepage. • market iassist and the data centres at the annual conferences of the south african online users group (saoug), the library association of south africa (liasa), and the standing committee of african libraries (scescal) in an attempt to reach information professionals that could be instrumental to establishing of data libraries at their residential academic or research institutions. • build capacity through research methodology and data analysis training. arrange regular symposia and workshops on data archiving and secondary analysis, addressing the problems as identified earlier in the paper, targeting universities and technikons. sponsorships could be sought from the national research foundation, medical research council, and statistics south africa. • circulate the available training material on the iassist web site to members with the permission of the authors and the iassist administrative committee, to assist in the capacity building exercise. • convene symposia, workshops, seminars and training sessions on any subject consistent with the iassist objectives for networking purposes. • publish an iassist/africa one-page newsletter customised for the african region, and reflecting iassist news and events, e.g. regional news and important announcements from the administrative committee, funding opportunities, and upcoming conferences on machine-readable data analysis to be held in the region. • tap into the network of the various consortiums to market iassist, the national data archive and other data centres • form a data library interest group under the guidance of the iassist objectives and joint co-operation of the liasa leadership. • encourage national archives similar to that of the south african data archive in the rest of africa, starting with the southern african development community and sub-saharan countries. • strengthen linkages between african universities and iassist-linked institutions abroad. • create opportunities for individuals to attend iassist/ifdo conferences to interact with colleagues both in africa and abroad. some of the recommendations require iassist intervention, whereas the rest depend on local awareness, support and regional commitment. conclusion despite the serious obstacles in africa, there is cause for optimism. the seriousness of the african plight is not irreversible. what is needed, though, are more resources, structural analysis, evaluation and scientific investigation by african scholars. affiliations between iassist and the south african data archive have made huge contributions towards machine-readable social science data archiving in south africa under the leadership of dr. maseka lesaoana. closer collaboration could only assist in enhancing the existing professional working relationship. references africa overview http://mbendi.co.za/land/af/p0005.htm brümmerhoff, w: south africa in the finance and investment sector of the southern african development community (sadc). http:www.resbank.co.za/economics/articles/art1298/ articles.html dike, alaezi, d.: scarcity of textbooks in nigeria: a threat to academic excellence and suggestions for action. journal of librarianship and information science, 24 (2) june 1992, pp. 79-85 distinct south african research opportunities. http://www.nrf.ac.za/focusareas/distinct/ ekong, donald: regional co-operation in graduate education and research as a key element in revitalizing higher education in africa, science in africa: innovations in higher education, aaas sub-saharan africa program, washington, d.c. 1992 nageri, michael w.: the african regional centre for technology information system. electronic networking in africa: advancing science and technology for development. workshop on science and technology communication networks in africa, nairobi, kenya, august 27-29, 1992 odhiambo, t.: the design and launching of afrand, the african foundation for research and development (keynote address), science in africa: the challenges of capacity building, aaas sub saharan africa program washington, d.c. 1992 http://mbendi.co.za/land/af/p0005.htm http://www.nrf.ac.za/focusareas/distinct/ 8 iassist quarterly summer 2002 iassist quarterly summer 2002 9 *paper presented at the iassist conference, may 2001, in amsterdam, the netherlands. julia paris, technikon witwatersrand, johannesburg, south africa, juliap@twrine t.twr.ac.za. mailto:juliap@twrinet.twr.ac.za mailto:juliap@twrinet.twr.ac.za vol281.indd iassist quarterly spring 2004 5 by by stuart macdonald * about edina 1 edina, based at edinburgh university data library, is a jisc-funded national data centre. it offers the uk tertiary education and research community networked access to a library of data, information and research resources. all edina services are available free of charge to members of uk tertiary education institutions for academic use, although institutional subscription and end-user registration are required for most services. services include spatial data services; abstract and indexing bibliographic databases; multimedia and images databases; in addition to a number of geo-related development projects such as geoxwalk, e-mapscholar and go-geo! the spatial data services offered include ukborders (boundary datasets of the united kingdom), digimap (ordnance survey maps and mapping data) and the edina agcensus which allows downloading and the visualisation of grid square agricultural census data from as far back as 1969. history of the agricultural census collecting livestock and crop information goes back as far as the domesday survey commissioned in december 1085 by william the conqueror who invaded england in 1066. in the middle ages governments did not collect statistical information for the benefit of the population, or even as a guide to policy but simply for tax and administrative purposes. in1801 an agricultural inquiry coincided with the first british population census; the reporters were local clergy and a standard form was used to record areas of main crops. this continued into the first half of the nineteenth century where various attempts were made to conduct agricultural inquiries. in general these did not produce a good response. about the same time, the first statistical account of scotland (which was published in twenty-one volumes between 1791 and 1799) was undertaken under the direction of sir john sinclair of ulbster. based on detailed parish reports, the statistical accounts enumerate and describe such topics as agricultural and industrial counting cows and cabbages – web-based extraction, delivery and discovery of georeferenced data production. the second new statistical account was published between 1834 and 1845. in 1864 parliament agreed to the collection and publication of agricultural statistics in great britain and in 1865 allocated £10,000 to cover the cost. thus the census began in its modern form in 1866. the agricultural census the agricultural census is conducted annually in june by each of the united kingdom agriculture departments to help form, monitor and evaluate policy by providing information on the distribution and extent of crop and horticultural production and rearing of livestock. each farmer is obliged to declare the agricultural activity on the land via a postal questionnaire. the respective government departments collect the 150 items of data and publish information relating to farm holdings for recognised geographies. farm holdings above a certain economic or physical threshold are regarded as major holdings with the rest being regarded as minor. the data provided from the respective government departments and converted by edinburgh university data library correspond to major holdings only. prior to 1998 data for england and wales was provided by defra. farmers from both territories completed the same questionnaire. since 1998 welsh farmers complete a questionnaire supplied by the welsh assembly depc. farmers in scotland complete a questionnaire as supplied by seerad. thus, due to each government department having responsibility for their respective questionnaire there is not complete comparability over the three territories with regard to census questionnaires. similarly, due to changes in agricultural policy the content of the questionnaires within the three territories has changed over time. in addition certain census items have been aggregated in recent years to address issues concerning disclosure. conversion algorithms areal research in relation to the distribution of agricultural census data was carried out at the university of edinburgh in the 1970’s and 1980’s by professor terry coppock and http://datalib.ed.ac.uk/" \t "_blank http://datalib.ed.ac.uk/" \t "_blank 6 iassist quarterly spring 2004 jack hotson, with the co-operation with maff and adas. the level of publication of the agricultural census data was the parish summary. coppock and hotson developed algorithms to redistribute the parish summary data into 5km and 10km grid square estimates, taking into account potential land uses. this was done for the following reasons • the geographies (e.g. scottish parishes, welsh communities) vary in size, shape, the land use capability • parish summaries may underor over-report agriculture activity e.g. a farmer need make only one return, even if some land or livestock are remote from the main holding • the census returns are a 'snapshot' of activity on 1 june • grid square data aligned to the national grid facilitate analysis of data over a number of years and with other data sets. • the format of data obtained from the government departments is potentially disclosive the key to transforming the agricultural census data into grid square data was the definition of each geography (parish) as 1km squares. this framework was used in conjunction with a 7-fold classification of the land-use of the same 1km grid squares called the land-use framework. the resulting distribution of the data gave a good estimate at 5km level of "what was likely to be where", as well as protecting farmers' confidentiality. migration to a web service up until recently data had to be extracted from a mainframe or local server, using a set of command driven extraction programs. the data could then be mapped using the interactive gridmap utility, written in-house by alison bayley, building on the raster-type camap software written by jack hotson for data retrieval and grid mapping. however the emergence of desktop gis and web technologies has enabled the service to develop further, making it potentially available to a newer and wider audience. new edina agcensus in spring 2004 preliminary work began on processing and reformatting existing data into a grid square format suitable for importation into and delivery from a mysql database. data from census years 1969, 1976, 1981, 1988 and 1994 was reworked from existing grid formatted data, while year 2000 data was converted directly into a suitable format from the areaspecific data provided by the respective government departments. this allows analysis of change over time with intervening and more recent census data to be processed in due course. existing data processing algorithms were re-written in java to allow for easy maintenance and migration between machines. the interface uses java servlets and jsp technology to enable clear presentation and access to the data. the options presented in the interface reflect the variation in censuses from country to country and from year to year. the post-processing database holds a table for each country, year and grid resolution combination, for which metadata is held in two lookup tables. the lookup tables hold information about what tables are available, which census items occur in each country/year combination and text descriptions, groupings and units of analysis. the data visualisation component of the new service uses image generation servlets which present the data in map form for the area, census item, year, grid size and bandings specified. decisions were made (and problems resolved) regarding the provision of an alternative visualisation for context mapping, and the use and development of the area select tool; allowing for both javascript and non-javascript functionality in addition to accessibility. it was also decided that the service would be more responsive if data were held in aggregate form rather than aggregated for each data request, after weighing up the issues of storage versus ‘on the fly’ delivery. two bartholomew’s raster datasets (1:200,000 and 1:800,000) were used as the context mapping for great britain. land use data from the 1980’s (for scotland, england and wales) were used to provide an alternative context. the new edina agcensus service was launched on october 1st 2004 offering grid square agricultural census data to 3 client communities: the academic community (via an annual athens authenticated subscription); commercial organisations, and research/policy makers (both via an edina controlled authorisation and authentication on a ‘per project’ basis) (see fig 1). iassist quarterly spring 2004 7 fig 1: the edina agcensus interface offering access to academic and non-academic customers in addition to a free demonstration version of the service a free visualisation demo service containing all census items at 10km resolution for the most recent census year held enables potential users to preview distribution maps of chosen census items although no data download is offered. authorised users are presented with a simple interface offering two routes into the data, namely data download (ascii delimited comma separated values) and data visualisation (distribution maps of census items) via a 7 step access procedure: • select country from scotland, england, wales (also gb for small subset of items) • select year from1969, 1976, 1981, 1988, 1994, 2000 • select census item(s) – at present only one census item can be visualised at a time. all (or a subset there of) of the census items for a chosen year at a chosen resolution can be downloaded • select grid size (level of aggregation) from 2km, 5km and 10km • select extent of data coverage (see fig. 2) http://agcensus.edina.ac.uk/demo/index.html 8 iassist quarterly spring 2004 fig 2: for small area analysis use the extent tool or enter british national grid co-ordinates. toggle between context map and land use map. • data selection summary (allows user to change chosen parameters) • download or visualise selected census data. as an example, cattle distribution for the south west area of scotland has been visualised (see fig. 3). the image itself can be downloaded as a gif file by right clicking on the image or printed off as a hard copy. in addition the unit bandings can be customised and the land use data for the specified area can be viewed. after visualising the chosen item for the selected area the corresponding data can be downloaded by clicking on the appropriate icon. iassist quarterly spring 2004 9 fig 3: distribution map of dairy cows and heifers in scotland, 1988 at 2km resolution. service issues prior to service launch the following key points were addressed: authentication. this was achieved via the athens access management system which provides users from the uk tertiary education community with single sign-on to numerous web-based services. non-academic authorisation and authentication to the service is managed by the edina helpdesk. documentation. availability of online user guides, questionnaires and publicity material (and hardcopy on request). field trialing. field trials of the service were conducted by the scottish agricultural college using nielsen-norman usability tests to precipitate feedback. feedback from in-house testing was also implemented into the interface and functionality of the service. subscription model. unlike other edina services agcensus is an exception with regard to subscription in that commercial/ policy/research organisations can gain access to the data. the jisc banding structure was employed for the academic audience. however those from the aforementioned non-academic institutions subscribe on a per project basis (allowing institutional access). individuals from both academic and non-academic organisations unable to raise relevant subscription costs also have the option to pay on a per project basis. training/outreach. structured training events and workshops will be organised to both publicise and demonstrate the service. this will be done in conjunction with the edina training officer through liaison with external institutions. online training materials will be made available via the edina agcensus website. user support. the edina agcensus service is supported by the edina helpdesk which adheres to a set of service level definitions. enquiries are dealt with via telephone and email with service downtime, alerts, upgrades etc being posted on the edina agcensus website. technical and in-depth service support are also available. accessibility. an accessibility statement explains edina’s policy of working towards maximum accessibility for all users to all services. 10 iassist quarterly spring 2004 further developments there are a number of developments being investigated for version 2 of the service. these include: • introducing data that complement the grid square agricultural census datasets such as gridded meteoro logical data, species data, historic census data (at present data for the 1871 scottish agricultural parishes are being digitised for processing) • online visualisation of change over time for a chosen census item • statistical reporting to provide summary statistics and enable rudimentary numerical analysis of the census data and predictive modelling • combining more than one census item together for visualisation or download • a teaching dataset for use in learning and teaching in the classroom, laboratory etc • other visualisations such as histograms, pie charts, bar charts the bigger picture as a uk national data centre, edina engages in both projects and services, the former being geared to development activities which inform and develop the operation of edina national services, either producing new services or improvement in existing services. projects are generally externally-funded and often in partnership with other institutions. three projects funded under the jisc 5/99 programme relate to web delivery of spatial data to the uk he/fe sector, and build on the national significance of digimap (which delivers access to ordnance survey mapping) and ukborders (digitised boundaries). these are go-geo!, geoxwalk and e-mapscholar. go-geo! with increasing amounts of spatial data being created within higher education, demand for managed access to this data is growing with gis tools becoming more commonly available. however, two barriers confront the potential user of spatial data: • how to find out what datasets exist • how to ascertain their quality and suitability for use. these barriers can be overcome by comprehensive, standardised metadata, available through the web-searchable portal. go-geo! is a jisc-funded project run jointly by the edina and the uk data archive. as a z39.50 compliant resource discovery tool it allows identification and retrieval of metadata describing the content, quality, condition and other characteristics of spatial data within and beyond the uk he community. the metadata profile originally employed by the go-geo! project was based on uk ngdf guidelines and the iso 19115 geographic information metadata standard which was adopted in march 2003 and mapped recently to dublin core. go-geo! also acts as the academic node of the uk gigateway service (hosted at edina) and can go beyond discovery to provide direct access to data in some cases (see fig. 4). iassist quarterly spring 2004 11 fig 4: go-geo! portal offers both simple and advanced search options in addition to a ‘library’ of geo-resources including news items, case studies, learning materials, data providers, training courses and discussion groups. a simple keyword search for e.g. agricultural census retrieves a number of results allowing the metadata for chosen records to be viewed (see fig. 5). with the advanced search facility searches can be restricted by data type (maps, images, datasets, reference material, projects etc), location, text, data range. in addition to cross-searching spatial databases from major data producers go-geo! also extends the data discovery function by providing access to a ‘library’ of other related resources of use to the user. these resources can be either local to the portal or found by searching the jisc information environment and other online information services. 12 iassist quarterly spring 2004 fig 5: retrieved results and individual record complete with ‘what, when, where’ metadata tabs geoxwalk the main purpose of this project was to provide a shared service within the jisc information environment (ie) that can support geographic searching. at present each information provider or service adopts different geographic coding principles (such as postcode, place name, grid reference). thus the creation of an online z39.50 compliant british and irish gazetteer would facilitate a unified entry point into geographical searching in addition to providing researchers and teachers with an online reference tool. the geoxwalk gazetteer itself contains a list of place names with their associated spatial location expressed in several ways (e.g. latitude and longitude co-ordinates etc). it also classifies features into types such as cities as areas, rivers as lines etc and stores an appropriate spatial ‘footprint’ against each feature. this introduces transparency to a geographic search by allowing the ‘cross-walking’ of these different geographies in addition to allowing searches to be conducted on a proximity/distance basis. integral to such a project was the need for a geoparser, software than can ‘read’ and automatically identify place names in an electronic document (e.g. a resource description or digitised historical document). such identified place names can then be compared against the geoxwalk gazetteer entries thus providing access to ‘alternate’ geographies by the assignment of geotags (e.g. a grid reference) to implicit geo-referenced material (e.g. a place name). thus the combination of the parser and the digital gazetteer has potential for powerful geographic based searching across a range of otherwise disparate resources such as those contained within the jisc ie. iassist quarterly spring 2004 13 e-mapscholar the third of edina’s jisc-funded spatial projects, e-mapscholar, is a learning and teaching project. from an earlier project (the jisc electronic libraries (elib) digimap project) it was identified that there was a skills/ concepts gap between creating a map and downloading and using digital map data in a gis. thus the aim of e-mapscholar was to fill this gap by developing tools to promote the use of spatial data, including os digital map data available from the edina digimap service, initially within tertiary education. however the model could be applied to other levels of education. it would support both those learners who need to progress to using a gis in addition to those whose needs are more straightforward. this spatial data literacy project has four components: teaching case studies which consist of the data and materials used by the learners, along with descriptions of how the data and learning materials have been integrated into a variety of disciplines, and evaluations by staff and students (see fig. 6). fig. 6: an example of a teaching case study “an introduction to digital mapping in archaeology” online learning and teaching materials such as tutorials have been developed with interactive tools which enable users to develop skills in the use of digital map data and knowledge of spatial data concepts such as integration and visualisation. a teaching content management system has been developed which allows teaching staff to customise and re-purpose the online learning materials to suit their curriculum. 14 iassist quarterly spring 2004 a virtual work placement has been designed in which students can carry out an assessment of the visual impact of wind turbines at the nant carfan development in wales. this provides the opportunity for learners to develop workplace-related skills in the use of spatial data, using problem-based learning techniques. at present the jisc are funding a follow-on project to look at ways in which they might be able to make the products from the e-mapscholar available in a service environment. summary the edina agcensus service forms part of a growing number of geo-data resources utilised within uk academia and has evolved using web technologies to enable data access to an expansive audience. projects related to spatial data services such as go-geo!, geoxwalk and emapscholar aim to raise awareness of geo-data and associated resources. additionally they enhance access to, and use of geo-resources to those both within and beyond the academic sector. such resources highlight the role of data service providers, such as edina, in offering and strengthening networked access to a collection of data, information and digital materials to the uk tertiary education and research community. comments on the edina agcensus service are welcome. email the author at: stuart.macdonald@ed.ac.uk. footnotes 1 all urls and acronyms are listed here in the appendix. appendix: urls edina national data centre: http://edina.ac.uk edina digimap: http://edina.ed.ac.uk/digimap edina ukborders: http://edina.ed.ac.uk/ukborders edina agcensus: http://edina.ac.uk/agcensus the domesday book online: http://www.domesdaybook.co.uk statistical accounts of scotland: http://edina.ac.uk/stat-acc-scot go-geo! http://www.gogeo.ac.uk geoxwalk http://www.geoxwalk.ac.uk e-mapscholar http://edina.ac.uk/projects/mapscholar/index.html gigateway http://www.gigateway.org.uk acronyms jisc joint information systems committee defra department of the environment, forestry and rural affairs (previously maff) seerad scottish executive environment and rural affairs department depc department for environment, planning and countryside maff ministry for agriculture, fisheries and food adas agricultural development and advisory service gis geographic information systems ngdf national geospatial data framework os ordnance survey * paper presented at the iassist conference, madison, may 2004, by stuart macdonald. contact: stuart macdonald, edina national data centre & edinburgh university data library, tel. 0131-650-3304 | e-mail: stuart.macdonald@ed. ac.uk http://edina.ac.uk http://edina.ac.uk/agcensus http://www.gogeo.ac.uk/ http://www.geoxwalk.ac.uk 26 iassist quarterly 2008 sharon bolton & matthew woollard1 strengthening data security: an holistic approach abstract in the light of heightened concern around data security, this paper highlights some of the measures that can be used to develop and strengthen security in data archiving. the paper includes discussion of the different approaches that can be taken towards the construction of firm and resilient data and information security policies within the social science data archiving communities. while international standards can provide theoretical guidelines for the construction of such a policy, procedures need to be informed by more practical considerations. attention is drawn to the necessity of following a holistic approach to data security, which includes the education of data creators in the reduction of disclosure risk, the integration of robust and appropriate data processing, handling and management procedures, the value of emerging technological solutions, the training of data users in data security, and the importance of management control, as well as the need to be informed by emerging government security and digital preservation standards. keywords data security; data archiving; information security; data handling; user training. new legislation and the data security ‘climate’ during 2007, a series of high-profile data losses by uk government and associated organisations took place, involving reputedly 37 million items of personal data2 in a climate of increasing concern over data security and identity theft, legislation in the form of the statistics and registration services act 2007 (srsa) was at that time already making its way through the uk parliament, its provisions intended to become effective from 1 april 2008. as a result of media furore over the data losses, in november 2007, the cabinet office was instructed to review data handling within government departments and make recommendations for their improvement where needed. this review was to be conducted under the direction of robert hannigan, head of security, intelligence and resilience. an interim report3 was produced in december 2007, followed by the final report4 in november 2008. together, the srsa and cabinet reports had a marked effect on how data are now handled across uk government departments and culminated in the ‘mandatory minimum measures’ for data handling5 the srsa also established the uk statistics authority, and for the first time introduced criminal penalties for the unlawful disclosure of confidential information (i.e. data relating to, or identifying, a particular person or business held or disclosed by the statistics authority). the uk data archive (ukda) is the curator of a large collection of uk government social science research data, held for use by the academic community. it is also licensed as a legal place of deposit by the national archives, allowing the ukda to ingest and preserve public records. therefore, while not strictly covered by the statistics and registration services act, by the nature of its business the ukda is intimately concerned with data integrity and security. while the ukda has developed and maintained robust security practices over the years since its inception, it was felt that the time was right to review practices in response to the challenge of new legislation. during the course of its work, the ukda acquires data and associated materials from data creators, conducts ingest processing to prepare those data for secondary use, and supports users once they have received the data. therefore, a similar holistic approach to the audit and refinement of data security was taken, to ensure that all the ukda’s data acquisition, ingest, access and support activities are supported by coherent and consistent procedures that work at all stages of the data archiving life-cycle. data creators and security enhancement firstly, the initial part of the process was reviewed – the work that the ukda undertakes with data creators – and an assessment was made of how the new uk legislation on data security may affect practice. the ukda will advise data creators at all stages of their project: from the planning stage, throughout the data gathering process and after completion and deposit of the data. this helps to ensure respondent confidentiality while maintaining sufficient detail within the data to enable effective research. the work is wide in scope, as data are acquired from a range of sources including large, well-funded government organisations, established research centres, and independent iassist quarterly winter summer 2008 27 small-scale academic research projects. the first tangible effect of the new uk legislation that the ukda experienced was a marked tightening of physical data transfer process from government, including the increased use of file encryption, more secure methods of data delivery, use of courier services, etc. this was unsurprising given the recommendations of the cabinet office report. the ukda accordingly streamlined and secured acquisition and internal data transfer procedures to accommodate these developments. beyond the security of data transfer, new developments in the nature of data released from the office for national statistics (ons) also quickly became apparent. since before 2006, the ons microdata release panel (mrp) have overseen the release of ons datasets by providing advice and testing for statistical disclosure, and as a result of the srsa, the panel updated their policy for the release of data at certain access-controlled levels. the ukda were primarily concerned with two ‘levels’ covered by this data access control strategy, which have been developed by negotiation between the ons and the ukda: the end user licence and the special licence. datasets are compiled by ons to these respective levels of detail prior to deposit with the ukda. to be able to download and use data from the ukda, each user must register an account with the economic and social data service (esds for which the ukda is a service provider), via the ukda website, and agree to the terms of the end user licence. this includes an undertaking that the user must preserve at all times the confidentiality of information pertaining to, and not to attempt to identify, individuals and/or households in the data collections, nor must they share data with others who are not registered users. in accordance with the terms of how the srsa defines levels of information, ons data made available for research access at end user licence level must not reveal or have the potential to reveal the identity of an individual6 however, it is also recognised that researchers may sometimes need access to more finelydetailed data to conduct their research, and for this reason the special licence was set up with ons. to gain access to data held under a special licence (which may include data designated as ‘personal information’ under the srsa, but must have had some degree of protection applied), the prospective user must apply via the ukda to ons for approved researcher status. the candidate must provide extensive details of their academic background, status, and prospective research, and prove that they are a ‘fit and proper’ person to receive special licence data. there are significant differences between the end user licence and special licence versions of a dataset. as a practical example, the lowest geographic level permitted for end user licence data is generally government office region (gor, or nomenclature of territorial units for statistics (nuts) level 1) and education or employment and other demographic data may be banded or aggregated. special licence data, by contrast, may include geographic data at a finer resolution, such as unitary authority and nuts2 or nuts3 geographies, and more detailed education, employment and demographic data. of course, the advent of the special licence does bring an added administrative burden; separate holdings of special licence and end user licence versions of a dataset require doubled ingest processing work to be undertaken at the ukda, and the administration of approved researcher applications is resource-intensive for both the ukda and ons. however, these steps must be taken to ensure that in a climate of heightened concern around security, data of a sufficient level of detail remain potentially available to the research community while respondent identity is still protected. further to this (whilst it lies outside the remit of this paper) the ukda will launch the secure data service (sds) in late 2009, where more potentially disclosive data will be available to selected approved researchers, who will then undergo further levels of security and system training. however, not all government data acquired by the ukda originate from ons. while government organisations primarily concerned with data have robust procedures in place, other departments, especially those who have only recently begun to release data to researchers, are still coming to grips with the provisions of the srsa, and do sometimes need guidance. as a result of consultations with ons over end user licence and special licence data, the ukda are uniquely placed to offer such guidance should potential confidentiality issues arise, and often do so. as well as offering practical advice on how to balance disclosure risk while maintaining useful detail within data, by means of data edits, access control or a mixture of both, the ukda (alongside esds government) are currently involved in the process of facilitating discussions between the ons mrp and other government departments, so that all may benefit from the mrp’s expertise. additional negotiation and guidance is likely to occur as a result of the secure data service with new licence agreements between the ukda and the ons and between the ukda and researchers. in the ukda’s experience, academic researchers also vary widely in data security expertise. some are attached to large research centres with established data managers and procedures, and others may undertake small-scale team or solo projects. while the data ukda receive from government are largely quantitative, the data generated by academic projects may be quantitative, qualitative or a mixture of both. a considerable amount of work is undertaken by the ukda to educate and inform the academic community on all aspects of data management, including obtaining consent; maintaining confidentiality and security; holding well-publicised regular workshops 28 iassist quarterly 2008 covering both quantitative and qualitative data are held around the uk; and providing advice to individuals and organisations at all stages of the research process. the data deposited at the ukda as a result of unique academic projects may present unusual challenges, and any potential edits or levels of access control that may be needed to reduce the risk of disclosure are discussed with researchers prior to the commencement of full ingest processing. the recent managing and sharing data 7 guide gives straightforward and plain english advice on security issues that affect researchers. internal data handling the second part of the holistic approach the ukda took to strengthen data security was to audit all aspects of ‘inhouse’ data handling. members of staff at the ukda have a wide range of expertise and experience, and work on diverse data tasks. therefore, it was essential as part of the audit process to examine and scrutinise internal procedures, both human and technological, for the handling and storage of dataset files and associated administrative materials. as a result of this, existing good practice was identified and additional methods developed. these were collated into a comprehensive set of data security procedures, which were then distributed to all ukda staff. in addition to the requirement that all staff register as ukda users and are thus bound by the same conditions of the end user licence, a confidentiality agreement has been also introduced. this details the responsibilities of staff with regard to data confidentiality, and requires the signature of the individual staff member. the smooth introduction of such stringent security measures needs to be carefully handled, and staff professionalism, awareness and expertise must be respected and acknowledged. explanations as to why security measures were to be strengthened were given to staff, and all background information on the srsa and the cabinet reports have been made available, including legal aspects regarding differential treatment of data at the end user licence and special licence levels. staff members were encouraged to ask questions and actively feed into discussions throughout the process, and where needed, training was provided. feedback from staff at all levels has resulted in some excellent suggestions that have since been incorporated into the security procedures. at the same time, a ukda security plan8 has also been developed in conjunction with the planning of the secure data service. this plan brings together all aspects of information security within a single document and is based on the iso 27001 standard. while the plan has been developed explicitly for the secure data service, it has ramifications which extend across the whole of the ukda’s procedures. it is within physical and environmental security and operations management that the key changes are likely to impact, but there are also technical changes that will be necessary and will influence the way in which staff carry out their daily work, even if their work is entirely unrelated to the secure data service. the plan will also be revised to take account of the more recent her majesty’s government (hmg) security policy framework9 and the provisions made in that framework will have to be applied to the ukda’s interactions with government departments as well as researchers. while this paper does not cover technological solutions in detail, it must be emphasised that the ukda servers and system architecture and associated infrastructure are maintained to broadly conform to iso 27001/2, and are updated periodically in step with technological advances. regular systems testing is undertaken to identify any vulnerabilities, and other practical measures include the inclusion of integrated checksums in the names of downloaded files to prevent sql injection, and processes put in place to prevent cross-site scripting and os command injections. these ‘development’ and technical solutions are are under continuous development, and there will also be additional changes before the commencement of the secure data service. the security plan also dovetails with other ukda policies and procedures including the established ukda preservation policy,10 which covers data authenticity/integrity arrangements. all plans, policies and procedures need regular review and updating. all the external factors that influence internal ukda procedures and policies will continue to develop, and the ukda must keep pace with these developments to both maintain internal standards and promote external confidence in the work of the ukda. this ongoing process of review and revision will remain collaborative; all staff will continue to be encouraged to reflect on issues relating to data security that affect their roles and to use that knowledge to improve working practice. the successful introduction and implementation of new and sometimes tiresome procedures related to security needs should have considerable staff support, and it is only by allowing staff to contribute to the process that this support can be effectively harnessed. data security: the user’s perspective the ukda’s work with data users comprises the third major area for data security review and development. as there are currently over 50,000 registered ukda/ esds users, providing data security advice and support to users is a considerable task. the ukda alongside its esds partners, acts to represent the interests of data users and accordingly facilitates dialogue with data creators to ensure users are given access to sufficiently-detailed data to permit useful research to take place. however, this brings an associated responsibility to promote safe data practices and train users in effective data security. the ukda have developed measures to facilitate this, including regular workshops and training, and a web-based guide to good practice on micro data handling and security 11, which has been available for some years and is regularly updated to take account of new developments. these measures will be iassist quarterly winter summer 2008 29 extended considerably for the secure data service. it is the ukda’s opinion that mandatory training courses for users of this service will provide one of the main planks in the strategy of allowing researchers desk-top access to secure data. the combination of detailed training, implementation of secure technologies, strong penalties and conformance with relevant standards will provide the necessary checks and balances to provide the most viable user experience. training in best practices in data management leads to training for researchers about the risks of disclosure. for the secure data service it will be imperative that researchers are aware of what constitutes disclosure, how to they can identify disclosure risks in their own analyses, and what the legal and practical consequences of disclosure are. training users will reduce the risk of security breaches. as part of the audit of data security, further measures are currently under consideration to improve the range of tools to educate data users on robust data security practice. however, an effective policy has to be in place to deal with any sanctions and breaches of data confidentiality by users, who are made aware of potential measures that can be taken against them or their host organisation in event of a breach. again, the planning for the secure data service has identified new risks which will be covered by a new multi-tier procedure. conclusion to summarise, it must again be emphasized that the work undertaken by the ukda to strengthen data security has of necessity taken a holistic approach. the three fronts on which data security work is of prime concern (data creators, internal practices, and data users) are all interdependent. work with data creators, including government departments at the beginning of the process, inculcates strong internal practices and data security awareness at the ukda, which in turn leads to the better safeguarding of data and excellent educational work with data users. this paper has shown how data security at the ukda has been strengthened in response to a changing external environment, and how the work will continue to develop as that landscape changes further. notes 1. dr. sharon bolton, data services manager, uk data archive: email sharonb@essex.ac.uk.dr. matthew woollard, associate director, head of digital preservation and systems, uk data archive: email matthew@essex. ac.uk 2 harrison, d. (2008) ‘government’s record year of data loss’, daily telegraph, 7 january. retrieved 15 may 2009, from http://www.telegraph.co.uk/news/newstopics/ politics/1574687/governments-record-year-of-data-loss. html 3 cabinet office (2007) data handling procedures in government: interim progress report, cabinet office, december. retrieved 15 may 2009, from http://www. cabinetoffice.gov.uk/media/65934/data_handling.pdf 4 cabinet office (2008) data handling procedures in government: final report, june (published november). retrieved 15 may 2009, from http://www.cabinetoffice.gov. uk/media/65948/dhr080625.pdf 5 cabinet office (2008) cross government actions: mandatory minimum measures, retrieved 20 may 2009, from http://www.cabinetoffice.gov.uk/media/cabinetoffice/ csia/assets/dhr/cross_gov080625.pdf 6 abrahams, c. and mahony, k. (2008) new policy and procedures governing the release of microdata derived from ons social surveys, paper presented at the 13th gss methodology conference, london, june 23rd. 7 uk data archive (2009) managing and sharing data. 8 the ukda security plan is currently an internal document only. 9 uk cabinet office (may 2009) hmg security policy framework, v.2.0 , retrieved 20 may 2009, from http:// www.cabinetoffice.gov.uk/media/207318/hmg_security_ policy.pdf 10 uk data archive (2008) uk data archive preservation policy. retrieved 15 may 2009, from http://www.data-archive.ac.uk/news/publications/ ukdapreservationpolicy0308.pdf 11 economic and social data service (2008) guide to good practice: micro data handling and security . retrieved 15 may 2009, from http://www.esds.ac.uk/news/ publications/microdatahandlingandsecurity.pdf the wisconsin longitudinal study: adults as parents and children at age 50 * by robert m. hauser ^' william h. sewell, john a. logan, taissa s. hauser, carol ryff, avshalom caspi, and maurice m. macdonald, institute on aging and adult life and centerfor demography and ecology the university of wisconsin-madison summary we are can7ing out a survey of more than 9000 american men and women who were first interviewed as seniors in high school in 1957 and have subsequently been followed up in 1957, 1964, and 1975; they will be about 53 years old when they are interviewed in late 1992 or early 1993. each interview, about one hour in length, will be followed by a shorter mail questionnaire. we shall also interview a randomly selected sibling of each respondent, using a slightly shorter version of the telephone interview. we also hope to obtain a waiver that will permit us to link our survey records to information from the social security system, but this part of the design is ciurendy under negotiation with the social security administration. finally, we expect to obtain enough information to link our records to the national death index. data from the wisconsin longitudinal study (wls) will be a valuable pubhc resource for studies of aging and the life course, inter-generational transfers and relationships, family functioning, social stratification, physical and mental well-being, and mortality. the study has 5 specific goals: (1) to extend models of occupation and earnings and to elaborate the roles of aspirations in adolescence and at mid-life, of previous achievements, and of familial responsibilities in current economic and social standing, subjective well-being, mental and physical health, disability, and wealth; (2) to identify and measure local effects on opportunity, that is, specific characteristics of a person, firm, or economic sector that directly influence the chances of obtaining a job or a limited range of jobs; (3) to extend and elaborate models of sibhng resemblance that will elucidate influences of the family of origin on the life course; (4) to investigate self-assessments of well-being in the context of aspirations, accomplishments, and social relationships with significant others; (5) to measure social and economic exchange relationships with parents, children, and siblings and assess the consequences of those relationships for well-being. we are planning a follow-up survey of more than 9000 american men and women who were first interviewed as seniors in wisconsin high schools in 1957 and have subsequently been followed up in 1958, 1964, and 1975; these individuals will be approximately 53 years old when they are interviewed in 1992. ' at the same time, we will interview a randomly selected sibling of most respondents. approximately 2000 of these siblings were previously interviewed in 1977, and we have sufficient resources to interview approximately 4000 more siblings during this round of the study. the data collection process will include a 1-hour telephone interview, followed by a self-administered mail questionnaire, and a waiver that will permit us to link our survey records to information from the social security system. also, we are collecting sufficient data to link our records with the national death index. these new follow-up data, combined with our existing files, will become a valuable public resource for studies of aging and the life course, inter-generational transfers and relationships, family functioning, social stratification, physical and mental well-being, and mortality. we expect that it will be possible to enhance the value of the sample and data with additional data collection and data linkages. we believe that the cost and effort of this project are fully justified by five specific goals outlined herein: (1) to extend the series of measurements and models of occupational achievement and earnings of the members of this cohort that have been obtained in their younger years and, in particular, to elaborate the roles of aspirations in adolescence and at mid-life, of previous achievements, and of familial resfxjnsibilities in current economic and social standing, subjective well-being, mental and physical health, disability, and wealth; (2) to identify and measure local effects on opportunity, that is, specific characteristics of a person, firm, or economic sector that directly infiuence the chances of obtaining a job or a limited range of jobs;* (3) to develop models of sibling resemblance that will elucidate infiuences of the family of origin on the life course, including social and economic achievements, social participation, subjective well-being, menial and physical health, success in childrearing, provision for retirement and old age, and patterns of morbidity and mortality; (4) to investigate selfassessments of well-being in the context of comparisons with previous aspirations and accomplishments, social statuses of parents, childhood friends, siblings, spouses. spnng/summer 1992 23 and children, and in the context of past and current social relationships with those significant others; (5) to measure social and economic exchange relationships with parents, children, and siblings and assess the consequences of those relationships for well-being. background and significance the wisconsin longitudinal study (wls) is a long-term study of a random sample of 10,317 men and women who graduated from wisconsin high schools in 1957. survey data were collected from the original respondents or their parents in 1957, 1964, and 1975. these data provide a full record of social background, youthful aspirations, schooling, military service, family formation, labor market experiences, and social participation of the original respondents. in 1977 the study design was expanded with the collection of parallel interview data for a highly stratified subsample of 2000 siblings of the primary respondents. the wls data the wls is a rich source of data on life-cycle processes that is of continuing interest to scholars in sociology, education, psychology, and economics.' the interview data have been supplemented by mental ability tests (of primary respondents and siblings), measures of school performance, and characteristics of communities of residence, schools and colleges, employers, and industries. the wls records for primary respondents are also linked to those of three, same-sex high school friends within the study population. the measurement of social background includes earnings histories of parents obtained from wisconsin state tax records, and the data on the socioeconomic careers of men in the main sample are supplemented by social security earnings histories from 1957 through 1971. the wls is widely recognized as one of the most useful bodies of longitudinal data on the lives of americans because of the quality of the survey measurements (and our efforts to measure that quality), extremely high retention of panel members, complete, multi-layered documentation of the data, and multiple linkages to personal and institutional records. research based on the wls the wls panel has been used to develop the well-known "wisconsin model" of social and psychological factors in socioeconomic achievement. we have located more than 800 ssci citations to 7 core wls publications since 1972. in addition, or in extensions of this central line of research, the wls data have been used in studies of geographic constraints on college access; recruiunent into teaching, nursing, and other occupations; choice of marital partner; differential family formation and fertility; gender differences in market participation and success; religious and ethnic differences in achievement processes; birth order effects on ability and achievement; effects of high schools anu colleges on aspirations and achievements; and inter-firm and inter-industry differences in compensation. also, the project has been the locus of many useful methodological developments built around the design, collection, or analysis of data from the wls. these include successful methods for tracing respondents over long intervals; the analysis of unit record data from the social security administration without compromising confidentiality; structural equation models of achievement processes; methods for comparative analysis of social mobility; models with errors in the reporting of social and economic variables; and models of common family factors in the achievements of siblings * our last direct contact with the primary wls respondents took place in 1975, when they were about 36 years old. at that time, most of the women were completing childbearing and were participating in the labor market or planning a return to it; men were well established in their occupational careers, but because they married younger women were not as far along in family formation. using these data, we have analyzed the process of socioeconomic achievement from adolescence to midlife and compared the socioeconomic achievement processes of men and women. wls siblings varied widely in age, but 80 percent were between 27 and 45 years old in 1975, and for adult sibling pairs, we were able to conduct studies of family resemblance and intrafamily differences in education, occupation, earnings, and fertility. among our main research goals are to extend our models of social and economic achievement and participation of primary respondents and their siblings. planned follow-up surveys in summer 1992 we began to interview the 9000 primary respondents and, whenever possible, a randomly selected sibling of each. the primary respondents will be 53 years old, and four fifths of their siblings will be 44 to 62 years old. at those ages, the wls respondents and their siblings will be anticipating their own retirement and aging as well as managing relationships with one another, their adult children and their elderly parents: (1) in 1975, 92 percent of respondents had at least one living sibling, and 71 percent had at least two. moreover, because of their position at the leading edge of the baby boom, siblings tended to be younger than primary respondents. thus, we expect that an overwhelming majority of wls respondents will still have at least one living sibling. (2) in 1975, 93 percent of the respondents had at least one living child; since child-beanng began around age 18 (for women) and was not yet complete in the cohort, we expect that almost all respondents will have at least one adult child. (3) survivorship is much less among the respondents* parents. we ascertained the 24 lassist quarleriy father's year of birth in 1975, and we estimated that 18 percent of respondents will have a living father in 1992, while 42 f)ercent of respondents will have a living mother. we estimate that about half the respondents will have at least one living parent, and the age of these parents will be around 80 years.' thus, we believe that our respondents are ideally suited for a study of aging and of intergenerational relations among adults. in our 1992 interviews, we are updating our measurements of marriage and divorce, child-rearing, education, labor force participation, jobs and occupations, social participation, and future aspirations and plans among primary respondents and siblings. in addition, we are expanding the content of the study by obtaining data about psychological well-being, mental and physical health, wealth, and social and exchange relationships with parents, siblings, and children. in designing the new measurements, we have attempted to maintain an appropriate balance between comparability with our own previous concepts and methods (which are similar to those used in the ciurent population survey and the 1973 occupational changes in a generation survey) and comparability with other significant research efforts, e.g., the new survey of health and retirement, the national survey of families and households, nih surveys of work and psychological functioning, and the norc general social survey; in addition, we have coordinated our design efforts with those of members of the macarthur foundation research network on successful midlife development finally, we plan to obtain information and waivers that will eventually link our survey data to social security records and the national death index. we have considered whether the collection of new data for the wls is warranted, given the existence of other longitudinal studies and the possibility of collecting similar data for a new national sample. the latter alternative may be desirable for some purposes, but it would be most difficult, and probably impossible, to provide the wealth of background and life history data that are available from the wls or other longitudinal studies. the more serious question is whether the wls is worth further investment, relative to other longitudinal studies of similar vintage. we think it is, for several reasons: (1) the wls data on the hfe course are unique in richness and quality. (2) major national longitudinal studies that began in youth cover more recent cohorts. these cohorts are of interest in their own right, but none has reached the pre-retirement years. for example, members of the national longitudinal study of 1972 will be around 38 years old in 1992, and there are currently no resources for further follow-up activities. those in the hsb samples of 1980 and 1982 will be 26 to 28 years old in 1992; members of the two younger panels in the 1967-68 national longitudinal studies of labor market experience will be 40 to 50 years old in 1992; the oldest cohorts covered in the monitoring the future surveys will be about 35 years old in 1992. (3) other longitudinal studies are restricted in similar ways to the wls, which covers high school graduates from wisconsin, almost all of whom are white. for example, members of the career development study were juniors or seniors in the state of washington in l%5-66, and they will be about 43 years old in 1992. the members of the norc survey of 1961 college graduates are essentially the same in age as those in the wls, but the sample is substantially more restricted with respect to educational attainment. (4) inject talent may provide a national sample that is just 3 years younger than the wls. however, it lacks the linkages of the wls to socioeconomic data, and there have been serious problems of sample coverage and data access throughout the history of project talent.' new directions for research in the wls we have considered several ways in which the research agenda of the wls could be extended. we have decided to focus on three of these opportunities in our initial work, without foreclosing the development of others at a future date. we believe that each of these is scientifically important and that they are complementary to the design and content of the wls: (1) effects of special preferences, skills, and attachments; (2) mental and physical health at midlife; and (3) social and economic exchanges and well-being. local effects most of the previous analyses of the wls data have used continuous measures of outcomes — particularly, years of education, occupational status, and earnings— as dependent variables in structural equation models. this has improved our understanding of the relationships of a number of background and social psychological variables to education, occupation and earnings. however, the amount of variance explained by these linear models has always been relatively modest this has been attributed to the operation of "luck" in individual outcomes (jencks et al. 1972), but it may also arise from a systematic neglect of factors that are not easily captured by linear models. the planned new wave of data collection will attempt to measure persistent effects of some of these factors, called "local" effects. a local effect on occupational opportunity is any characteristic of a person or of a firm or economic sector which directly influences the chances of obtaining only a limited range of jobs. it is contrasted with a "general" effect, such as the effect of general education, which influences chances of employment in a wide range of jobs. an example of a local effect would be a particular skill or aptitude, such as mechanical aptitude. there are some jobs, mostly in the middle range of prestige and income, which demand high mechanical aptitude. possession of this aptitude should have a local effect on spring/summer 1992 25 an individual's occupational chances, raising the probability of landing the jobs requiring it, but not raising the probability of landing good jobs in general. aside from specific skills or aptitudes, three other main types of local effects can be distinguished, namely, preferences, contacts and structural shifts: (1) individuals may, for reasons subject to empirical study, have preferences for certain kinds of work, such as outdoors jobs, jobs with less than usual amounts of direct supervision, or jobs with high creative or artistic potential. jencks, perman and rainwater (1988) have examined nonmonetary, non-prestige attributes of jobs and found them highly predictive of individuals' reports of job satisfaction, yet only poorly related to demographic measiu'es such as age, sex, and education. we want to examine the relationship of such non-monetary, non-prestige preferences to particular hfe course developments, rather than to the measures just named. for example, preferences for certain job characteristics may vary with interand intra-generational family responsibilities, and stage of the life course. (2) individuals may have direct or indirect personal contacts among those making hiring decisions in certain jobs. the desire of a parent to pass along a business to a child, the preference of a union for enrolling the children of its members, and the general social contacts of a parent or child which may be useful job leads for the child are all examples. appropriate methods, which we expect to apply and refine, will make it possible to estimate the magnitude of these social network effects in a well-defined, general population. (3) finally, the economy as a whole may experience contractions or expansions of opportunity in certain types of work. such structural shifts cause transitory increases or decreases of opportunity in limited ranges of jobs. we are asking for retrospective descriptions of the first and last occupations held by respondents in their first two and last two businesses or organizations where each respondent has worked since 1975; in most cases, this will give us a complete employment history. thus our occupational data will cover years witli widely different levels of overall economic activity. it is important that models of local effects in occupational outcomes are not confounded with local structural effects; the broad temporal scope is intended to aid in distinguishing the two. to put the overall point most simply, measuring and modehng "local effects" may explain more of the variation in occupational (and related) outcomes than can be done with regression, and may increase the qualitative detail of the explanations associated with multivariate studies of life course achievement in general populations. as individuals age and experience changes in their priorities and responsibilities in the posl-childrearing, pre-retirement years of their fifties, qualitative aspects of the choices they make— as reflected in local effects — may produce more concrete explanations of behavior. mental and physical health we plan to examine the influence of educational and occupational pursuits on mental and physical health. the inclusion of detailed measures of psychological and physical functioning to the telephone and mail instruments will strengthen the multidisciplinary significance of the wls by linking the attainment process, typically the domain of sociology, to mental and physical health, typically the domain of psychology. the proposed linkage affords significant strides in several research domains. first, prior studies of well-being have documented connections with education and income for men and women in american society (diener 1984; veroff, douvan, and kulka 1981), but the effects have been small. however, previous studies have used single-item measures of well-being that are of questionable reliability and validity (larsen, diener, and emmons 1985). these measures have shown no connection to theories of psychological health (coan 1977; jahoda 1958; lawton 1984; ryff 1989a) nor to related empirical measures (ryff 1989b). the wls employs a differentiated, multifactorial concept of positive functioning that incorporates not only global happiness and satisfaction, but also the respondents' assessments of their effectiveness in dealing with the external world (autonomy, environmental mastery), and their sense of direction and progress in life (purpose in life, personal growth). with additional measures of physical health status, it will thus be possible to map the effects of educational and occupational attainment on an array of mental and physical outcomes. previous research on the relation of social structural factors (e.g., education, income) to psychological functioning has also been largely descriptive. previous studies chart the magnitude and direction of linkages between demographic characteristics and subjective well-being, but do not specify the mechanisms through which educational and occupational achievements affect self-evaluations. two central social-psychological mechanisms will be explored in the research: (1) we will examine how social comparisons with significant others influence subjective well-being in midlife. the parallel sibling sample in the wls provides vital comparative data about the respondents' attainments relative to a key group of significant others. this question constitutes a significant departure from prior psychological research on siblings, which has focused on effects of sibship variables (e.g., number of siblings, birth order) on achievement, intelligence, and personality (zajonc 1976), as well as on disaggregating the comparative effects of genetic and environmental factors on behavior (plomin and daniels 1987). few studies have examined the nature of sibling relationships in adulthood and later life (cicirelli 1989) or the consequences of these relationships for psycholassist quarteriy lopical well-being. it is likely that adults use their siblings as "measuring sticks" to evaluate their lot in hfe (troll 1975). the wls thus provides a compelling data set with which to study the influence of sibung relationships— and their inherent social-comparative features— on self-evaluations, subjective well-being, and physical health in midlife. the specific cognitive mechanisms through which such comparisons influence well-being are derived from a synthesis of relative deprivation theory (suls 1986), tesser's (1988) self-evaluation maintenance model, and various strands of attribution theory (mirowsky and ross 1990). additional comparative data will be obtained on the attainments of the respondents' parents and children. adults who have accompushed less than their parents may be at greater risk for psychological distress. alternatively, the "american dream" suggests that parents hope to have children who do at least as well, if not better, than themselves, so negative discrepancies with children (i.e., when children have accomplished more) may be conducive to positive self-evaluations. these expanded self-other comparisons offer important new directions to research on intergenerational relations, which has neglected the midlife era when one's children are becoming young adults and one's parents are growing old (hagestad 1987). (2) the second proposed social-psychological mechanism through which educational and occupational attainments influence self-evaluations is temporal comparisons. those who have advanced considerably beyond their starting resources are expected to show more positive self-evaluations than individuals who have made little gain or have lost ground. it may not be absolute levels of education, income, or status that predict psychological well-being, but the magnitude of those attainments relative to the resources with which one began. the wls provides significant advances over prior studies because we can operational ize temporal comparisons in a behavioral, performance-oriented domain (educational and occupational achievement). previous research has examined only subjective perceptions of personality change (markus and nurius 1986). in sum, the planned study combines a theory-guided view of psychological well-being with fresh ideas about the relevant social-psychological processes by which people evaluate their accomphshments within the context of enduring family bonds across the life course. the design weaves data on three generations and multiple siblings, enabhng us to explore the dynamics of individual development in the context of family histories and generational succession. social and economic exchanges eggebeen and hogan (1990, p. 4) have nicely slated the case for improved measurements of social and economic exchanges among parents and children: "in small-family societies, ... [tjheory thus suggests that parental investment will be diluted when a large number of children compete for resources, and that it will be more heavily concentrated on children who bear them grandchildren. ... these hypotheses have been difficult to evaluate for modem societies because of the paucity of data documenting patterns of exchange." our new data will include measures of exchanges including patterns of kin contact, financial assistance, and the provision of services and care-giving. in the context of extensive wls information on family origins, marital and fertility histories, earnings records, and status attainments, measuring these exchange variables should permit major advances in our abiuty to test a variety of hypotheses about inter-generational u-ansfers. following the family sociology tradition of adams (1968), recent findings from the national survey of famihes and households (sweet, bumpass, and call 1988) have again demonstrated that patterns of kin contact, care-giving, and financial support tend to be intertwined. relatives who help each other with one type of assistance tend to provide the others as well, and to communicate more frequently in person and via mail and telephone contacts. although the most intensive caregiving assistance is provided for severely ill or disabled persons, help in the form of child care remains very important despite the trend toward purchasing that care on the market. as opposed to receiving aid, giving follows a u-shaped pattern by age, and in a manner that is particularly salient for persons of the wls sample's ages: young and elderly adults get more aid, and middle-aged adults are much more likely to provide aid (lee 1979; morgan 1982). additionally the female members of the sample are more likely to exchange services with their kin, with males involved more in financial exchanges (eggebeen and hogan 1990). along with the likelihood that wls members are heavily involved in family transfers and exchanges because of their stage in the life-cycle and the interaction effects of gender and age, their social exchanges should also tell us more about the impact of several recent trends. these include increased female labor force participation, rising divorce rates, increased demands for government spending on programs for children, and the slow growth in wages since the early 1970s. furthermore, the extent to which prime-age parents may have been able to offset the effects of these influences on their children may have been restricted by their own economic and personal difficulties, as well as by the need to plan for new obligations that arise from increased life expectancy for themselves and their parents. the current sources of family financial support for wls sample members will include gifts and loans from older spring/summer 1992 27 parents and other relatives, as well as actual and expected bequests. however, many wls members are likely to donate substantial time and financial support to their parents, as well as to their own offspring. also, donalions by grandparents to the children of wls respxjndents may alleviate financial pressures for some. economists who have studied these inter-household transfers (its) emphasize them as potentially important determinants of economic status that may have substantial redistributive effects (cox and raines 1985; kurz 1984). furthermore, the literature about the effect of social security on savings and retirement behavior has long recognized the potential of its to either complement (cox 1987) or offset income opportunities from public ffansfer programs (barro 1974, lampman and smeeding 1983). whatever their effect on the mix of support from family and public sources, it is clear that if its substantially augment the resources available to pre-reiirees, they are likely to affect work behavior via wealth effects on labor supply and savings decisions (kodikoff 1987). consequently, there is a need to identify which wls respondents receive substantial its, to improve understanding about their role as a potentially important reason for heterogeneity in work behavior, social exchanges with kin, and psychological well-being. additionally, analyses of the circumstances that motivate wls respondents to donate substantial its to their parents, children and siblings can help to elucidate how the financial pressures of those responsibilities influence earnings and other economic status variables. previous research ' previous research with the wls developed comprehensive social psychological models of socioeconomic achievement from adolescence through age 36. in recent work, our aim has been to incorporate estimates of response errors in variables entering into the models and to account analytically for similarities and differences between siblings. we studied the extent to which the parameters of our stratification models were distorted by random and correlated errors in reports of parental status, social influences, and educational and occupational aspirations and attainments. we incorporated our estimates of errors in these variables into attainment models both for men and for women. we have made considerable progress in our research on sibling similarities and differences in socioeconomic careers and in family formation and fertility behavior.'" socioeconomic achievements ofmen and women throughout the project one of our principal efforts has been to develop models of social and psychological influences on educational and occupational attainments of men and women that incorporate our best estimates of enrors in parental status variables, social influences, educational and occupational aspirations and attainments. we began our efforts by developing a model for men that incorporates some 26 measured variables into a recursive system of 14 unobservable (latent) constructs; the functioning of 9 of the latter variables is further simplified by postulating 3 other unobservable variables (hauser, tsai and sewell 1983). briefly, the model specifies that social origins and ability affect postsecondary schooling and occupational careers by way of aspirations and social influences in late adolescence. the analysis asks whether the wisconsin data are consistent with the modified causal chain hypothesis proposed in the original formulation of the model (sewell, haller, and portes 1969), rather than with models that incorporate many more lagged effects. the causal chain hypothesis receives far greater support than in previous analyses of the data that have not taken account of survey response error (and other stochastic components of latent variables in the model); that is, the lag-1 effects postulated in the model are far stronger than has been found in the past, and few delayed effects are present. for example, the model accounts for 69 percent of the variance in post-secondary schooling, for 73 percent of the variance in the status of first jobs, and for 69 percent of the variance in occupational status at age 36. the model identifies random response errors and correlations among responses obtained on the same occasion, from the same person, or using the same method. the model also allows analysis of the contamination of retrospective reports of social influences and aspirations by intervening events. thus, the analysis provides new evidence about the stratification process, about the validity of retrospective and contemporaneous reports of status variables, and about the social psychology of retrosf)ection. this model has also been estimated for women in our sample in order to compare the educational, occupational and economic achievements of men and women (tsai 1983; also, see sewell, hauser, and wolf 1980). we find that, although women have gained parity in educational attainment, their labor force activities and outcomes are still restricted. whatever occupational equality may exist at any one stage of the life cycle, women have fewer opfxjrtunities for gains in occupational status over the life course. women obtain smaller returns on their earlier occupational achievement than men do. whereas women are forced to rely more on academic performance and formal education for occupational placement, men increase their occupational status over the life cycle mainly as a result of their earlier occupational experiences. moreover, parents transmit direct occupational and economic advantages across generations to their sons, but not to their daughters. on the other hand, women receive larger earnings returns to educational attainment and occupational status than do their male counterparts. the comparisons also indicate that men's earnings are primarily determined by their occupational status, whereas women's earnings are primarily dcterlassist quarterly mined by the amount of labor supplied to the market. finally, marriage and childbearing have positive effects on men's earnings, but negative effects on women's earnings. however, the negative effects of marriage and childbearing for women disappear when labor force participation is controlled. effects offamily structure and sibling resemblance we have examined the effects of birth order and size of sibship on educational attainment for the full sibships of our primary respondents (hauser and sewell 1985). we have undertaken this analysis because of the recent revival of interest in birth order effects resulting from theories proposed by zajonc and markus (1975) and by lindert (1977). in our suidy of the 30,000 men and women in the full sibships of our 9,000 primary respondents we find no effects of birth order on educational attainment when size of sibship and other relevant variables are controlled, whether we look at selection into the sample of high school graduates, post-secondary educational attainments of those graduates, or educational attainments within full sibships. educational attainment appears to increase with birth order when family size is controlled but this happens when secular increases in schooling have occurred within as well as across families. thus, when we control birth year and parental education, there is no significant association between birth order and educational attainment there are no linear or non-linear effects, there are no effects of being first or last bom, and there are no statistically significant or patterned differences among ordinal positions. thus, there is no need to invoke any of the more complex theories of child development or intra-familial resource allocation to explain the effects of birth order on educational attainment because there is nothing to explain. retherford and sewell (1991) have carried out a comprehensive test of the confiuence model using data on the mental ability of wls primary respondents and siblings, and there, too, the findings have been clear and negative. we have studied sibling resemblance in education, occupational status, and earnings and in age at marriage and fertility (clarridge 1983). we find little resemblance between in fertility between sisters, but there is a great deal of family resemblance in socioeconomic achievement and its antecedents. for example, we estimate that family origins are associated with 49 percent of the variance in measured ability, 46 percent of the variance in educational attainment, 4 1 percent of the variance in the status of first jobs, 38 percent of the variance in status of current jobs (in 1975), and 27 percent of the variance in earnings. much of this research has involved the development of su^uctural equation models of sibling resemblance in educational and occupational status (hauser 1984; hauser and mossel 1985; hauser and mossel 1987; hauser 1988). in this work multiple measurements of educational attainment and occupational status for male high school students and their brothers are used to develop and interpret skeletal models of the regression of occupational status on schooling that correct for response variability and incorporate a family variance component structure. these analyses have provided a methodological template for the specification of more complete models of stratification (hauser and sewell 1986), and we are very excited about the prospect of extending them to cover the later achievements of wls respondents and siblings. we have not yet exhausted the possibilities for analyses of sibling resemblance in the existing wls data, and we are continuing to work on several topics: inter-sibling influence on educational attainment (lee 1989); the factorial complexity of schooling; sibling resemblance in social participation; and the social psychology of adolescent status attainment much of our previous analytic effort has been spent in developing models and methods for these analyses. with the combination of the sibling pair design and multiple measurements obtained from selfand proxy reports by sibhngs, we believe that it will be possible to make dramatic progress in modeling effects of family background, of individual differences in achievement, and of cross-sibling effects on achievement (hauser and wong 1989). we believe that similar models and methods will also help to elucidate social influences on the broader array of outcomes that will be measured in the 1992 survey, especially those pertaining to physical and mental health. mental and physical health psychological well-being will be assessed with a multidimensional formulation of positive functioning based on the integration of clinical, mental health, and life-span developmental theories (ryff 1989a). the points of convergence in these theories constitute six key dimensions of well-being (autonomy, environmental mastery, personal growth, positive relations with others, purpose in life, self-acceptance), which have been operationalized with sdiictured self-report scales (ryff 1989b). preliminary research indicates that the scales have acceptable psychometric properties, and that certain of them, particularly positive relations with others, personal growth, autonomy, and purpose in life, account for additional and independent variance beyond that covered by earlier measures of well-being (e.g., life satisfaction, happiness, self-esteem). in addition to these instruments, the multifactorial assessment of well-being will include global, single-item indicators as employed in prior survey research (veroff, douvan, and kulka 1981), measures of psychological distress (i.e., depression), and physical health status. these instruments will enable comparisons with other data sets (e.g., isr surveys) as well as afford more precise evaluation of the impact of the attainment process on multiple aspects of mental and physical health. spnng/summer 1992 29 depression will be measured by the center for epidemiological studies' depression scale (ces-d) (radloff 1977). physical health will be assessed with the oars (duke university 1978) checklist of illness, measures of height and weight, and items regarding subjective health evaluations and perceived changes in health since age 40. additional items, as developed by the macarthur midlife research network, have been included to assess activities and time devoted to health maintenance. the self-evaluation maintenance (sem) model (tesser 1980) predicts the conditions under which people will react with either jealousy or pride to the success of comparison others. specifically, closeness/likeness (e.g., in age, sex) is hypothesized to moderate the effect of relative performance. thus, the interaction effect predicts that if the sib performs better than the self and the sib is more like the self, there will be greater friction. using this general model, we can examine the implications of social comparisons in the family for subjective well-being as well as for patterns of support and assistance between family members. we will also examine how sibs cope with discrepancies in their individual achievements. within-family achievement differentials are hypothesized to be a source of psychological discomfort. to alleviate distress individuals may reconcile their achievements relative to their sibs by discounting the success of comparison of others and relinquishing responsibility for their own shortcomings, e.g., denying responsibility for failure: "i've had little control over the things that happen to me." just as people compare their attainments to those of others, they also compare their present selves with their past attainments. our objective here is to first predict people's present achievements (occupational status, earnings) using the wisconsin model. our interest lies in the effects of achievements, net of individual endowments, education, and achievement in the early career. first, we arc concerned with the implications of these differing relative locations for mental and physical wellbeing in midlife. for example, individuals who have gone beyond their original resources are predicted to show positive self-evaluations (self-acceptance), a sense of effectiveness (autonomy, environmental mastery), and personal progress (purpose in life, personal growth). second, we are concerned with respondents' aspirations for the future and for their offspring as a function of where they are relative to where they began. of special interest is whether the success of children helps to mitigate the adverse psychological effects of underachievement among midlife adults. finally, we will examine how people revise their past as a function of what they have or have not attained. among the hypotheses derived from control theory is that people rewrite the past. for example, they may look back to an earlier period and recall having lower aspirations than they actually reported at the time. they can thus exaggerate change when in fact little change has occurred (ross 1989). social and economic exchanges in work on household economics with the national survey of families and households, we have been studying the determinants of inter-household transfers (its) by focusing on the respective roles of family background, hfe-course events, and government transfer income opportunities. our analysis has established that its may be particularly important for certain types of households, and especially after age 45. beyond that age, most nsfh respondents give more in gifts and loans than they receive on average. however a substantial minority reported receiving much more than they gave during the nsfh's 1982-86 recall period. for the 45-59 year subgroup, 12 percent received gifts averaging nearly s9,(xx), with 7 percent reporting loans that averaged about $10,000. after age 60, the average of gifts and loans received are substantially less—roughly half of their pre-retirement levels. although the percentage receiving bequests is about 2 percent for all persons over 45 the average amounts of these its are large—at about $20,000 for the 45-59 group, and $50,000 for older respondents. hence although mature adults continue to help their own children via gifts and loans, those who receive help from their relatives (primarily parents) get substantial support from them, and a few can expect to receive very large bequests. as in our work with the nsfh, we plan to use wls histories on demographics, earnings, and other experiences to construct variables that describe whether and how recently the respondents had experienced events that would increase (or decrease) their needs for gifts and loans. events such as the onset of a severe illness influence the timing of these transfers. given that a transfer occurs, family background, respondent's earnings, and government income opportunities determine how much help gets provided. for younger persons its seem to be associated with major life-course events such as births and marriages that create need for help with basic living expenses. however after age 45 nsfh respondents tend to report that gifts and loans they received were more often for homebuying and other invesunents, i.e., intended to help them accumulate wealth. accounting for wealth transfers and their potential influence on well-being and the rigidity of the class structure requires a more comprehensive model that links prior transfers, such as those to fund educational achievement, to current transfer behavior. to study that process we intend to adapt the wisconsin status attainment model, by elaborating it to include its as an influence on status achievements. an nsfh result that motivates our interest in its as a potentially imporlassist quarteriy tant intervening variable is that the net effect of respondent's education on gifts and loans received is highly positive in models that control for father's education and current earnings. education may be tapping otherwise unmeasured influences of family wealth operating through educational achievement however families that provide help to educate their children may continue to support them throughout the life-cycle, for which they presumably get better support in their old age—in which case educational differences tap differences in preferences, not wealth. the family background measures in the wls are more complete for the purpose of analyzing it effects than in any other data set. specifically the parent's income data from wisconsin tax records will control much better for initial differences in ability to provide transfers. finally, we plan to analyze whether and to what extent both receiving and giving its and care-giving assistance influence wls respondent's psychological well-being. douthitt and macdonald (1990) have been using the wisconsin basic needs survey to study the relative contribution of life-cycle variables and alternative measures of financial status to variation in the andrewswithey deughted-terrible scale on subjective well-being. as part of that work for nimh, they have been able to separate the effects of wealth and measures of net worth from current earnings. a wls follow-up that included a match with social security earnings would permit better analyses of the financial sources of variation in global satisfaction measures. additionally placing perceived well-being as the ultimate dependent variable in a model that includes family background, current economic statuses, and measures of inter-family transfers would yield information about the relative importance of these transfers, as gauged by measures of satisfaction and not merely in dollar terms. in particular, we note that although economists have been very active in modeling the determinants of family assistance, they have not been very explicit about the importance of that assistance—either in economic terms, or as otherwise evaluated by the recipients themselves. design and methods the study is based on a telephone interview and selfadministered mail-out, mail-back questionnaire of wls primary respondents and their siblings. it will build on information about the life course previously obtained in surveys in 1957, 1964, 1975, and 1977 and from various public records. this section describes the means by which new survey information is being collected, integrated with the existing data (excepting confidential social security records), subjected to preliminary analyses, and made available to other cooperating researchers as core information on which additional data collection efforts and analyses may be based. the wls sample is large and heterogeneous, and it is broadly representative of white american men and women who have completed at least a high school education. the sample is mainly of german, enghsh, irish, scandinavian, polish, or czech ancestry. some strata of american society are not well represented in the wls. everyone in the primary sample graduated from high school; about 7 fjercent of their siblings did not graduate from high school. we have estimated that about 75 percent of wisconsin youth graduated from high schools in the late 1950s. minorities are not well represented; there are only a handful of african american, hispanic, or asian persons in the sample. the wls sample is sometimes criticized for over-representing persons of farm origins. that is not correct. about 19 percent of the wls sample is of farm origin, and that is cmisistent with national estimates of persons of farm origin in cohorts bom in the late 1930s. in 1964 and in 1975, about two thirds of the sample lived in wisconsin, and about one third lived elsewhere in the u.s. or abroad." there has been very little attrition in the course of the wls. response rales, relative to the full, initial cohort sample of 10,317, were 86.5 percent in 1964 and 88.6 percent in 1975. (that is, in the 1975 follow-up we did not drop individuals for whom no response had been obtained in 1964.) in the current round of the study, we originally planned to include only the 9138 individuals who participated in the 1975 survey and a surviving sibung (if any) of those individuals. in addition to individuals who died by 1975, this excluded about 3 percent of the original sample who were dropped from the 1975 survey because they could not be found, about 6 percent who were dropped from the sample because they could not be interviewed by telephone (because of illness, institutionalization, or residence outside the u.s., or because they could not be reached by telephone), and another 4 percent of the original sample who refused to participate in the 1975 study. before the tracing began (in july 1991), we knew that we would achieve substantial success in tracing the 1975 respondents, for we had found 92 percent of a pilot sample of 184 respondents in the 1975 survey. for the production tracing operation, we divided the sample into three broad strata: 1975 respondents for whom no brother or sister had been drawn into the 1977 sibling survey (about 65(x) persons);'^ 1975 respondents for whom a brother or sister had been drawn into the 1977 sibling survey (about 25(x) persons); and 1975 nonrespondents who were not known to have died (about 1000 persons). each of these groups was divided into 10 stratified random replicates. the main lines of su-atification are the sex of the respondent and his or her selected sibling, and the educational attainments of the respondent and sibling. we carried out production tracing one spnng/summer 1992 subsample at a time, beginning with the non-sibling subsamples, followed by the sibling subsamples. the subsample design gave us rapid and reliable feedback about our overall success rate, and it also smoothed the flow of easyand hard-to-find cases. as of september 1992, we have successfully located between 96 and 98 percent of both members of each potential respondent sibling pair in each of the first five replicates of both the non-sibling and sibling-pair samples. we are continuing the tracing operation to complete the remaining subsamples and to relocate respondents who move between the initial trace and the attempted telephone interview. our tracing efforts are carried out almost entirely by telephone. we begin with a direct call to the primary respondent or selected sibling at the last known telephone number, and we continue with a call to the parents' last known number. those methods yield sufficient information for about half the cases. we find the remaining cases using a variety of methods, based on previously known addresses, siblings' and childrens' names, high schools or colleges attended, and places of employment. two key tools have been a commercial credit union database (in which we have no access to financial information) and a national database of names, addresses, and telephone numbers on cd rom. we count a case as completed only when we have confirmed names, addresses, and telephone numbers (or lack thereoo for both members of a sibling pair with a responsible adult in their family. given the success of the main tracing effort, we decided to carry out a pilot effort to find persons who did not respond in 1975 and were not known to be dead. using our standard methods we were able to locate 86 percent of a random pilot sample of 99 persons, and— after considering the need to collect additional background material — we decided to include 1975 non-respondents in the new follow-up. study design the wls cohort of men and women, bom about 1939, precedes by about a decade the bulk of the baby boom generation that continues to tax social institutions and resources at each stage of life. for this reason, the study can provide early indications of trends and problems that will become important as the larger group passes through its fifties. this adds to the value of the study in obtaining basic information about the life course as such, independent of the cohort's vanguard position with respect to the baby boom. in addition, the wls is also the first of the large, longitudinal studies of american adolescents, and it thus provides our first large-scale opportunity to study the life course from late adolescence through the mid-50s in the context of a complete record of ability, aspiration, and achievement. past waves of the study have provided multiple, often overlapping measures of factors affecting life-course aspirations and outcomes. in addition to the fundamental advantage of obtaining true longitudinal measures for causal modeling, multiple measures have been valuable in estimating the effects of measurement error on the parameters of causal models of aspiration and attainment. of all the multiple measurements, however, the most fruitful have perhaps been the parallel questions asked of core respondents and their siblings. a recent series of papers, described above, has shown the power of this design for discovering the effects of unmeasured factors which operate within famihes to influence a variety of outcomes in later life (hauser and mossel 1985, hauser and sewell 1986). unfortunately, as the possibilities of this feature of the study design have come to seem ever more promising, the smaller size of the sibling sample (about 2000) compared with the core sample (9,138), has become a limiting factor. some analyses cannot be done with the low statistical power available at this sample size, for example, when we work with subsamples of sibling pairs defined by the sex of the primary respondent and his/her brother/sister (lee 1989). for this reason, and because of the substantive importance of investigating family effects, we proposed that a randomly designated sibling of every primary respondent, an additional 55(x) persons, be interviewed in this round of the study; at this writing, we expect to have enough support to interview about 4000 of these brothers or sisters, so we will exclude some of the replicate samples from this part of the study." it is important to note that the existing sample of 2000 siblings was augmented to include all twins of the core sample members, whether or not they had been drawn in the 10,(x)0 original cases of the high school sample. there are 1 16 distinct pairs of twins, a sizeable number for a sample from a general population, followed for a long period of time. timing ofactivities previous experience with the wls provided a sound basis for planning the sequence of our activities. the past year was spent primarily on sample tracing, instrument development and pretesting, and the creation of selected abstracts of data from the project files or from the 1975/1977 questionnaires that are being used directly in the 1992 interviews. for example, aside from identifying information, the telephone interviews use prior data on marital status, job in 1975, children, and siblings; we ask the respondent about his relationship with a best high school friend only in the 20 percent of cases where each member of a dyad in the sample named the other as among his or her best high school friends. there are three instruments: the core respondents' interview schedule; the siblings' interview schedule (possibly with some modification for siblings who were not previously 32 lassist quarterly interviewed); and the mail-back questionnaire to be sent both to core respondents and siblings. the first two questionnaires will be very siniilar.%% the instruments have been developed and pretested thoroughly with the help of persons in the class of 1957 who are not in the wls sample. this year will see the collection of the data, by both telephone and mail, together with additional tracing activities for previously-located respondents who cannot be relocated at the time of the survey. data will be merged and loaded into a preliminary file, and cleaning operations will begin. as soon as the data are clean, the preliminary files will be made available to interested researchers outside the group. we expect to prepare these files for replicate subsamples on a flow basis, so some data will become available before the fieldwork is complete. in the third year extensive data merging and variable construction will lead to preliminary data analyses and publications. these early efforts will probably be straightforward exploitations of the new data, and will extend the time horizons in standard sociological and social psychological models of the life course. also during this year, plans and proposals will be formulated for additional analyses of the data and for the record linkages that will be possible with them. the wls data the planned research will make use of detailed information already obtained for earlier periods of the life course. previously collected data span more than 3600 columns of coded items per case, and they cannot be described in detail here. an overview of the existing data may help indicate the potential usefulness of the planned new survey data and linkages.'* in 1962, william h. sewell obtained data from a 1957 survey of wisconsin high school seniors in public, private, and parochial schools. a random sample of 10,317 cases (approximately one-third of the seniors), was selected for further study. information on the measured mental ability of each student was added to the cards from the files of the wisconsin state testing service, which at that time conducted a testing program covering all eleventh graders in the state. a number of indexes based on information from the survey were developed and added to the cards for each student, including the socioeconomic status of the student's family, the student's attitudes toward higher education, educational and occupational plans, and perceived influence of significant others on educational plans. relevant measures of school, neighborhood, and community contexts for example, the socioeconomic composition of each senior class, the percentage of its members who planned on going to college, the size of the school. the size and degree of urbanization of the community of residence, and the distance of the student's place of residence from the nearest public or private college or university were constructed from secondary sources. in the spring and summer of 1964, seven years after the students had graduated from high school, we undertook a follow-up study of the original sample. using a questionnaire on a double postal card, information was obtained from parents on the post-high school education, current occupation, military service, marital status, and present residence of over 87 percent of the sample (sewell and shah 1967). with the cooperation of the wisconsin department of revenue (and following their strict arrangements to guarantee the privacy of individual records), information on tlie parents' occupations and income was obtained from their 1957 to 1960 state income tax returns. still later, we obtained information on earnings for the males in our sample from the social security administration for each year of covered employment from 1957 to 1967. this phase of the project required an elaborate linkage procedure to protect individual identity. the earnings record was later extended to cover the period from 1957 to 1971. our data were further enriched by addition from several published sources of information on the characteristics of secondary and post-secondary schools, colleges, and universities. during 1975, we carried out 1 hour telephone interviews with the sample and obtained the following information from our sample: (1) composition of family of origin: age, sex, and education of each sibling, the occupation and address of a randomly selected sibling, and the parents' ethnic and religious background; (2) the education of the respondent: content, timing, and location of all post-secondary schooung, including vocational, collegiate, and military schooling; (3) labor force experience: dates and types of military service, first civilian job, occupation in 1970, current (1975) job, longest job in 1974, earnings in 1974, weeks and hours worked, location, size, and type of work organization, work satisfaction, work authority, occupational aspirations, labor force participation and jobs held before marriage and in each birth interval (women only); (4) characteristics of family of procreation: marital status, marital history, a roster of children by age and sex, and educational and occupational aspirations for a randomly selected child; spouse's work status, education, occupation, and 1974 earnings; (5) selected retrospective information: aspirations while in high school and names of best high school friends; (6) social participation: membership in organizations, church attendance, visiting behavior, voting. we obtained similar information from interviews with 2000 randomly selected siblings (including all twins) during 1977, and at that time we also searched the records of the state testing service for spring/summer 1992 33 mental test score )r the siblings. data collection although no attempt was made to obtain an agreement to be interviewed as part of the 1989 trial study, the 1975 survey obtained responses from 88.6 percent of the primary respondents, and the 1977 survey obtained responses from 87.4 of the randomly selected siblings. the project has attempted to cultivate the good will of the sample, through reports made to the respondents and by other means, and we expect that excellent response rates will again be obtained. at this writing, about 700 interviews have been completed in the first two random replicates of the main sample, and these reflect about a 95 percent completion rate among all direct telephone contacts with respondents. within the first random replicate, the overall response rate is akeady more than 80 percent. survey operations the questionnaire will be administered in two parts. core items dealing with social and demographic characteristics and changes in them, self-assessments of health and well-being, social participation, and relationships with parents, siblings, and children, along with future aspirations and plans, are obtained in the telephone interview, which is being conducted by the university of wisconsin letters and sciences survey center (lssc). items were selected for the telephone interview if there administration required many logical branches or if the items were not grouped with a long list of similar questions. the interview may be somewhat shorter, perhaps 45 minutes, for siblings who participated in the 1977 survey. interviewers are using computer-assisted (cati) techniques, with which the lab has long experience. the project staff provides initial telephone numbers to lssc from its separate tracing activity, and is standing by to do additional tracing when numbers prove to be out of date. other information essential to the telephone interview, such as rosters of children's and sibhngs' names needed for information updates, have been transcribed from the original 1975 questionnaires and entered in the computerized interview schedule database. responses to occupation and industry questions, which are especially difficult to code, are routed from the field to our occupation coding section, and cases with incomplete responses are returned to the field within a day or two for a callback. one useful feature of the cati interview is the ability to introduce alternate forms or to sample selected questions at different rates in different internal replicates. for example, we are using two different series of questions about job authority, each administered to half the sample; we are asking a lengthy set of questions about depression and alcohol use of 80 percent of the sample; and we are asking about the current income of surviving parents in half the sample. we may alter sampling rates of these and other questions as the fieldwork proceeds. because the telephone interview should not be too long, some of the social psychological, health, occupational, and social exchange data are being obtained with a mailed, self-administered questionnaire. mail items tend to be groups of closely related questions with few logical contingencies and similar closed-ended response alternatives. the mailed questionnaire requires about 30 to 45 minutes to complete. lssc is providing two remaihngs to encourage respondents to mail back their questionnaires. however, a subset of the items in the psychological scales is administered in the initial telephone interview, to avoid a total loss of information from those not returning the mail questionnaire. appropriate statistical technjques will allow the resulting partial information to be included in structural models with measurement error, correcting for biases that would otherwise be intractable (allison 1987, allison and hauser 1991). after two pretests of preliminary mail questionnaires, we carried out a final pilot test of the mail instrument with three waves of mailing, and we obtained an 80 percent response rate. as explained above, we hope that respondents will grant us a limited, written waiver for the use of their social seciuity numbers (ssn's) to obtain social security earnings data. this will permit us to use our existing files of social security data directly, and, more important, it will permit us to build earnings histories of women, to complete the earnings histories of men in the wls, and to obtain additional data from social security records on disability, dependency, and death. aside from the administrative requirement to have written waivers for access to these data, the ssn will also be important in linking the wls to the national death index in future studies of differential mortality. we have ssn's for almost all of the males but for none of the females in the wls; we have no waivers at all. we want to obtain waivers and additional ssn's without losing the good will of the sample. our tentative plan for obtaining waivers is as follows: during the telephone interview, we ask the respondent to give us his or her social security number (ssn). if the response to this request is negative, the matter will be dropped; if it is positive, as it is in some 92 percent of the interviews completed thus far, we will mail a waiver form after completion of the mail interview. the mailed waiver itself will be accompanied by a note inviting the respondent to call the principal investigator directly with any questions. we had originally planned to obtain waivers before completing the mail interviews, but delays in reaching an agreement with the social security administration have precluded this design. assist quarterly the replicate samples will be introduced sequentially into the interviewing, mail survey, and waiver processes, just as in the tracing operation. aside from the advantages already mentioned, a smooth work flow and feedback on response rates, this design makes it possible to terminate data collection prematurely if costs run above budget; that is, it will be possible to reduce costs by lowering the size of the final sample without jeopardizing the quahty of the data or permitting nonresponse rates to rise to an unacceptable level. the two thousand matched sibling pairs for whom we already have sibling interviews (from 1977) are in pre-existing subsets of the wls; they will be introduced into the field operations near the beginning of the process, but not at its very beginning. that is, we want to be sure that everything is working smoothly before we begin to work on these key segments of the sample, but we do not want to wait so long that there is any chance of our terminating the field operations before their data have been collected. a similar internal sampling procedure was used successfully in the 1975 followup survey. 1. the research described herein is supported by grants from the national institute on aging and the national science foundation and by the graduate school of the university of wisconsin-madison. preparation of this paper was supported in part by the spencer foundation, the william f. vilas trust, and the kenneth and carolyn brody foundation and by a training grant from the national institute on aging to the center for demography and ecology at the university of wisconsinmadison. the opinions expressed herein are those of the authors. please address all correspondence to robert m. hauser, department of sociology, the university of wisconsin-madison, 1 180 observatory drive, madison, wisconsin 53706. 2. presented at the lassist 92 conference held in madison, wisconsin, u.s.a. may 26 29, 1992. 3. we began this round of study with the intention of following only those individuals who had been interviewd in 1975. however, we found it possible to locate previous non-respondents, as well, and we now plan to follow and interview all surving members of the original sample. 4. one might think of mental ability, educational attainment, or occupational prestige as general effects, whereas aptitude, personal contact with an entrepreneur, or training in cosmetology are local effects. 5. these data (with identifiers removed) have been placed in the public domain through the data and program library service of the university of wisconsin-madison. one exception is social security earnings histories of men in the sample from 1957 through 1971, which were obtained under conditions which preclude their distribution (or the direct access of the investigators to the data files in which they are contained). two other exceptions are files of detailed characteristics of colleges attended and of the employers of the primary respondents in 1975; these are not confidential, but we maintain them separately from the master file. 6. the wls data have been used in 4 research monographs, 23 doctoral theses, 1 1 masters theses, and more than 1(x) research articles or chapters in books. sewell and hauser (1992) review the study from the early 1960's to the present 7. in the first 400 completed interviews, the rartes of parental surviorship far exceeded our estimates: 55 percent of respondents had a living mother, and 26 percent had a living father. these cases represent the first 62 percent of persons to respond within a stratified random subsample of 650 primary respondents. 8. a pilot effort to relocate members of the project talent sample, carried out in parallel with our initial feasibility tests, provided ubsatisfactorily low coverage. 9. the wisconsin longitudinal study was supported continuosly by the national institute of mental health (mh-6275) from 1962 through 1982. during that period, the wls also obtained support from the social security administration (social and rehabilitation service grant no. 314) for linking and analyzing earnings histories and from the spencer foundation for the 1977 survey of siblings. from 1980 to 1986 the project was supported by nsf grants for studies of sibling resemblance (ses 8010640) and for the documentation of machine-readable data (ses 83-19879). the wls had no federal support from 1986 to 1991, and we have continued to work on analyses of family effects on achievement with support from the guggenheim foundation, the volkswagen foundation, the graduate school of the uw-madison, the brody foundation, the spencer foundation, and the use of core facilities of the center for demography and ecology at the uw-madison, which are supported by grants from the national institute of child health and human development (hd-5876) and the william and flora hewlett foundation. 10. this text covers only a few of the issues in recent wls research. for a full review see sewell and hauser (1992). 11. the 1991-92 respondent tracing activity show a similiar distribution of respondents between wisconsin and other locations. 12. we had selected a brother or sister of these persons during the 1975 survey, but we could not afford to spnng/summer 1992 35 interview them at that lime. 13. these individuals were designated in the course of the 1975 interview with the primary respondents, and at that time their full name, address, age, sex, educational attainmen, occupation, and industry were ascertained, along with the name of the last wisconsin high school they were known to have attended. the last piece of information is helpful in finding mental test scores. thus, while the records for these individuals will lack the selfreported information obtained in the 1977 sibling interviews, there is already some baseline information about them in the wls files. funding for this phase of the study has not been obtained. 14. copies of the mail questionnaire and a list of questions in the telephone interview are currendy available from the authors. at this time, there is no complete written text for the telephone interview, other than the script for the catl program used in the survey (cases). 15. with the exception of identifiable or confidential material, these data are now available from the data and program library service of the university of wisconsinmadison, 1180 observatory drive, madison, wisconsin 53706. we expect to release the new edition of the data through the inter-university consortium for political and social research. references adams, b.n. 1968. kinship in an urban setting. chicago: markham. allison, paul d. 1987. "estimation of linear models with incomplete data." pp 71-103 in sociological methodology 1987, ed. clifford c. clogg. washington, d.c.: american sociological association. allison, paul d., and robert m. hauser. 1991. "reducing bias in estimates of linear models by remeasurement of a random subsample." sociological methods and research 19 (4)(may);466-92. barro, robert 1974. "are government bonds net wealth?" journal of political economy 82 (november/december): 1095-117. cicirelli, v. g. 1989. "feelings of attachment to siblings and well-being in later life." psychology and aging 4:211-16. clarridge, brian r. 1983. "from family of orientation to family of procreation." diss. university of wisconsin—madison. coan, r.w. 1977. hero. artist. sage, or saint? a survey of views on what is variously called mental health. normality. maturity. self-actualization, and human fulfillment. new york: columbia university press. cox, donald. 1987. "motives for private transfers." journal of political economy 95 (june): 508-46. cox, donald, and frederick raines. 1985. "interfamily transfers and income redistribution." in horizontal equity. uncertainty, and economic well-being, ed. m. david. nber studies in income and wealth, vol. 50. chicago: university of chicago press. diener, e. 1984. "subjective well-being." psychological bulletin 95: 542-75. douthiu, robin, and maurice macdonald. 1990. 'the relationship between subjective well-being and the relative income hypothesis." in quality oflife studies—proceedings of the 3rd quality-of-life conference, ed. h. lee meadows. blacksburg, va: omni press. duke university center for the study of aging and human development. 1978. multidimensional functional assessment: the oars methodology. durham, nc: duke university. eggebeen, david j., and dennis p. hogan. 1990. "giving between the generations in american families." paper presented at the 1990 meetings of the population association of america, toronto. hagestad, g.o. 1987. "parent-child relations in later life: trends and gaps in past research." pp 405-34 in parenting across the life span: biosocial dimensions, ed. j.b. lancaster, j. altmann, a. rossi, and l.r. sherrod. new york: aldine. hauser, robert m. 1984. "a computing environment for the social sciences." berichl der edv kommission. geisteswissenshaflliche sektion, max planck gessellschaft: 44-59. . 1988. "a note on two models of sibling resemblance." american journal ofsociology 93 (may): 1401^23. hauser, robert m., and peter m. mossel. 1985. "fraternal resemblance in educational attainment and occupational status." american journal ofsociology 9\ (november): 650-73. hauser, robert m., and peter a. mossel. 1987. "some structural equation models of sibling resemblance in educational attainment and occupational status." pp 108-37 in structural modeling by example: applications in 36 (assist quarterly education and the social and behavioral sciences. cambridge: cambridge university press. hauser, robert m., and william h. sewell. 1985. "birth order and educational attainment in full sibships." american educational research journal 32 (spring): 1-23. . 1986. "family effects in simple models of education, occupational status, and earnings: findings from the wisconsin and kalamazoo studies." journal oflabor economics 4 (spring):s83-s115. hauser, robert m., and raymond sin-kwok wong. 1989. "sibling resemblance and inter-sibling effects in educational attainment" sociology of education 62 (my): 149-71. hauser, robert m., shu-ling tsai, and william h. sewell. 1983. "a model of stratification with response error in social and psychological variables." sociology ofeducation 56 (january): 20-46. jahoda, m. 1958. current concepts of positive mental health. new york: basic books. jencks, christopher, marshall smith, henry acland, mary jo bane, david cohen, herbert gintis, barbara heyns, and stephan michelson. 1972. inequality: a reassessment of the effect of family and schooling in america. new york: basic books. jencks, christopher s., lauri perman, and lee rainwater. 1988. "what is a good job? a new measure of labor market success." american journal of sociology 93{6){may): 1322-357. kotlikoff, laurence j. 1987. "intergenerational transfers and savings." nber working paper no. 2237. cambridge, ma: national bureau of economic research, may. kurz, mordecai. 1984. "capital accumulation and the characteristics of private inter-generational transfers." economica (february). lampman, robert j., and timothy m. smeeding. 1983. "interfamily transfers as alternatives to government transfers to persons." review of income and wealth (march). larsen, rj., e. diener, and r.a. emmons. 1985. "an evaluation of subjective well-being measures." social indicators research 17: 1-17. lawton, m.p. 1984. 'the vaneties of wellbeing." pp 67-84 in emotion in adult development, ed. cz. malatesta, and c.e. izard. beverly hills, ca: sage. lee, g.r. 1979. 'the effects of social networks in the family." in contemporary theories about the family, ed. w.r. burr, and et al. new york: the free press. lee, mehng-lum. 1989. "some structural models of family and inter-sibling effects on educational attainment: findings from wisconsin data." master's thesis. university of wisconsin—madison. lindert, peter. 1977. "sibling position and achievement" journal ofhuman resources 12: 198-219. maricus, h., and p. nuruis. 1986. "possible selves." american pyyc/io/ogj5r 41(9)(september): 954-^9. mirowsky, j., and c.e. ross. 1990. "control or defense? depression and the sense of control over good and bad outcomes." journal of health and social behavior 31 (11-^6). morgan, james n. 1982. "the redistribution of income by families and institutions and emergency help patterns." in five thousand american families, vol. 10, ed. martha hill. ann arbor. institute of social research. plomin, robert, and denise daniels. 1987. "why are children in the same family so different from one another?" behavioral and brain sciences 10(1): 1-59. radloff, l. 1977. 'the ces-d scale: a self-report depression scale for research in the general population." applied psychological measurement 1: 385^01. retherford, robert d., and william h. sewell. 1991. "further tests of the confluence model." american sociological review 56 (april). ross, m. 1989. "relation of implicit theories to the construction of personal histories." psychological review 96: 341-57. ryff, cd. 1989a. "beyond ponce de leon and life satisfaction: new directions in the quest of successful aging." international journal of development 12: 35-55. . 1989b. "happiness is everything, or is it? explorations on the meaning of psychological well-being." journal of personality and social psychology 57: 1069-081. sewell, william h., and vimal p. shah. 1967. "socioeconomic status, intelligence, and the attainment of higher education." sociology of education 40 (y^inlet): 1-23. sewell, william h., archibald o. haller, and alejandro portes. 1969. "the educational and early occupational attainment process." american spring/summer 1992 37 sociological review 34 (february): 82-92. seweu, william h., and robert m. hauser. 1992. "the wisconsin longitudinal study." working paper 92-01, center for demography and ecology, the university of wisconsin-madison. sewell, wilham h., robert m. hauser, and wendy c. wolf. 1980. "sex, schooling and occupational status." american journal of sociology 86 (november): 551-83. suls, j. 1986. "comparison processes in relative deprivation: a life-span analysis." pp 95-1 16 in relative deprivation and social comparison: the ontario symposium, vol. 4, ed. j.m. olson, c.p. herman, and m.p. zanna. hillsdale, nj: erlbaum. sweet, james, larry bumpass, and vaughan call. 1988. "the design and content of the national survey of families and households." nsfh working paper no. 1. madison, wi: center for demography and ecology. tesser, a. 1980. "self-esteem maintenance in family dynamics." journal ofpersonality and social psychology 39:77-91. . 1988. "a model of self-esteem maintenance." in advances in experimental social psychology, ed. l. berkowitz. new york: academic press. r. troll. 1975. early and middle adulthood. monterey, ca: brooks/cole. tsai, shu-ling. 1983. "sex differences in the process of stratification." diss. university of wisconsin—madison. veroff, j., e. douvan, and r.a. kulka. 1981. the inner american: a self-portraiifrom 1956 to 1976. new york: basic books. zajonc, robert b. 1976. "family configuration and intelligence." 5ae/ice 192: 227-36. zajonc, robert b., and gregory b. markus. 1975. "birth order and intellectual development." psychological revie\i' s2: 74-88. lassist quarterly iassist quarterly 2016 35 iassist quarterly abstract the federal reserve system3 has a longer tradition of doing economic research than disseminating data from economic research. each of the 12 reserve banks and the board of governors have research departments that together publish nearly 1,000 working papers and journal articles annually. unfortunately, researchers have not often made the data from their papers publicly available until recently. a new program at the federal reserve bank of kansas city aims to correct this imbalance and make such data available for reuse in other research. as a pilot participant in a new dissemination platform, we have educated economists, built metadata specifications, recruited contributors, collaborated with technology and legal staff, and coordinated and built coalitions across multiple functions at our institution and others. this paper outlines the challenges faced and obstacles overcome as we created the infrastructure and workflow and took steps toward making the publication of research data a regular part of the research life cycle. keywords data dissemination, metadata, usability testing, research data management introduction and background as researchers are deluged with data, the need for them to share or disseminate their research data widely may not seem pressing—there’s plenty to go around. (the data deluge, 2010). researchers now have more choices when compiling data to support their research, and reusing other researchers’ data is an important option. a wide range of available data—and advances in analytical capabilities and processing power—means researchers can replicate or build upon a broader variety of research than was possible in the past. recent mistakes in (and fabrication of ) research data have only highlighted the importance of research replicability and the sharing of research data. the federal reserve bank of kansas city, like other reserve banks and the board of governors in washington d.c., conducts research to support its monetary policy mission, contribute to the safety and soundness of banks, and promote financial stability. the bank shares its research products with policymakers, other researchers across the system and in academia, and the public. increasingly, this research requires analyzing vast quantities of data and employing substantial computational resources to address increasingly complex questions. in kansas city, more than two dozen researchers and research associates produce about 50 research products (journal articles, working papers, etc.) each year. they are part of a larger community of more than 750 researchers across the system who produce nearly 1,000 such works each year. as the need to acquire data inputs to this research has increased, so has the pressure on support staff and the budget. to help alleviate some of the pressure, the 12 reserve banks and the board have been formally collaborating to acquire source data as research inputs. first forays into research data dissemination: a tale from the kansas city fed by san cannon1and deng pan2 36 iassist quarterly 2016 iassist quarterly while the system has established a set of services to bring in data, there are no such services for pushing out data once the research is complete. the federal reserve currently has no coordinated approach to preserving and disseminating research data across the system, and differences in strategies and resources have precluded a consolidated approach. other academic domains face similar challenges (borgman, 2012). challenges and opportunities setting up a repository or an archive, or defining a workflow to support data preservation or future dissemination, are not just technology decisions. for many disciplines, these activities require a fundamental change in researchers’ perceptions of the research process. for most researchers, the ultimate goal is publication; fed economists are no exception. all data-related work is simply in support of that goal: ‘time and money spent on documenting data for use by others are resources not spent in data collection, analysis, equipment, publication fees, conference travel, writing papers and proposals, or other research necessities’ (borgman, 2012). effecting change in such circumstances is difficult but not impossible, and research funders may lead the charge. for example, some funding agencies now require publicly funded researchers to make their underlying data available to the public. although these requirements do not affect federal reserve researchers, they will undoubtedly help change the culture of empirical research as a whole (arzberger et al., 2004). even without funding requirements, a few banks in the federal reserve system have begun considering how to disseminate their research data sets. the federal reserve board, for example, publishes data for select working papers along with the papers on their website4 . the data are being disseminated, but researchers must know with which paper they are associated and must then go to that page. the federal reserve bank of new york takes a similar approach but is also compiling a separate page for such data5 . the federal reserve bank of kansas city, however, does not currently disseminate research data sets on its public website. for many reserve banks, the major means by which research data is disseminated is individual requests to the author. this process, though widespread even in the academic community, is taxing on the author and does not encourage broader reuse. a coordinated approach to data acquisition in the federal reserve system began in late 2011 and was aided by the creation of a data librarian role in each of the reserve banks. the federal reserve bank of kansas city took research support one step further by creating the center for the advancement of research and data in economics (cadre)6 in early 2015. cadre’s mission is to support, enhance, and advance data or computationally intensive research in economics. bank leadership identified research data preservation and dissemination as important support functions for cadre. as these functions were being designed, the kansas city fed was offered, and accepted, an opportunity to join a pilot program to help provide use cases and product suggestions for a new research data dissemination platform being developed by a nonprofit academic partnership outside the federal reserve system. becoming a pilot participant the new publication platform is designed for researchers to submit data directly for dissemination. the workflow in the platform allows for two stages: an initial submission by the researcher and a curation step to verify, edit, update, or clarify the contents of the initial submission before publication. to participate, cadre needed to specify the details of each stage in the workflow and build use cases for our research community. to do so we considered the following questions: • what kind of collections? in the dissemination platform, a collection is a group of data sets for which similar policies and access controls can be set. because metadata and access rights are controlled at this level, we created a public data collection and a restricted access data collection based on our evaluation of the data sets that had been used by researchers in the federal reserve system. we needed to be able to store data to which access was limited and test if the access controls worked sufficiently to meet our information security requirements. we also wanted to test the submission workflow and applicability of common metadata across the collections. • what kind and size of data files? although the platform was built to accommodate very large files stored in a central location, the typical data files in the pilot range from 1mb to 5gb, and are stored on the researcher’s desktop or on the kansas city fed’s high-performance computing cluster. • what kind of workflows? in the initial phase of the submission workflow, submitters fill in required metadata fields to describe the data sets and then assemble data files. in the subsequent curation phase, curators review and possibly modify the metadata or files before approving or rejecting the submission. in addition, we had some practical questions about how the pilot itself should proceed. • who would be the testers? we expected four to six test users, mostly economists and their research associates at the kansas city fed, to start this pilot and provide initial feedback. we also hoped to expand the test-user base to include users in a few other reserve banks as well as their co-authors at academic institutions. • who would be the curator? while cadre had plans to hire a data curator, the position had not yet been filled. the curator role was thus temporarily filled by two staff (specifically, the authors). metadata creation the most critical decisions for this pilot involved choosing which metadata fields the data set submissions should capture. opinions vary on how much information is sufficient to adequately describe any item. we investigated several existing specifications to evaluate what others view as necessary information for finding or discovering data. first, we looked at the requirements for depositing a data set at the university of michigan’s inter-university consortium of political and social research (icpsr), which maintains a data archive for social science research data7. although our data-dissemination pilot is not iassist quarterly 2016 37 iassist quarterly meant as an archive, many of the issues for data discovery are the same. the submission process for icpsr doesn’t explicitly require any fields, but we believe title, principle investigator, and description are the minimum metadata fields that are practicable. in addition, icpsr requests another dozen or so categories of metadata ranging from mode of collection to geographic coverage. because the icpsr archive handles a large number of data sets that are primary data collections, many fields, such as response rate and weights, rarely apply to data used in federal reserve research projects. next, we reviewed one metadata schema specifically designed to help users discover data. for federal agencies that provide data, the executive order omb 13-13 specifies a particular metadata schema to catalog data assets8. the metadata elements defined for that order are published as part of project open data9 comprise 12 required fields, some common to other specifications (for example, title, description, keyword) and others specific to government agencies (bureau code, program code). in addition, six fields are required by the schema if applicable, including license, rights, and spatial and temporal metadata. finally, we examined the metadata requirements that an internal workgroup had developed to catalog data assets across the federal reserve system. this unpublished specification lists more than 30 metadata elements including 14 mandatory elements. some items are common to other specifications (such as name or description), whereas others describe access restrictions (for example, security classifications). after carefully considering the user burden for metadata entry, we decided on the following metadata elements. required optional data set name additional information contact author name key words contact author information journal of economic literature classification description geography update information unit of observation category date(s) frequency documentation access restriction security classification we also added three metadata fields to the curation workflow: file type(s), file size(s), and article information. the file specification information is important both for technical staff managing the file space and for users who initiate a download. incorporating these into the curation workflow rather than the submission process reduces burden on the depositor and allows for possible changes to file types and sizes (through compression, for example) before the data are published. infrastructure choices and challenges once the interface and workflow were developed to our specifications, we performed initial testing before involving our users. we faced early infrastructure challenges; specifically, getting the dissemination platform to work well with our environment. the dissemination platform is meant to interact with our existing file system and storage infrastructure. we had some difficulty getting the file system and publication platform to interact smoothly. because we need to allow access by outside users, the file storage and platform live in an external zone of the kansas city fed’s intranet. this positioning made it easier for external users to get to the files, but made it more challenging for internal users to load the files from their desktop to the endpoint. the challenge was not insurmountable, but it did frustrate us as we worked through a typical user experience. another major infrastructure difficulty involved authentication and identity management. as part of a pilot project, we needed to have accounts on the platform infrastructure. these identities were then used to manage access to the collections through group definitions. for our initial work, adding individual users to the appropriate roles and access groups was fairly straightforward. however, in planning for a more robust long-term implementation, we do not want to maintain identities and security groups separate from existing information security infrastructure. usability testing the goal of this pilot is to evaluate whether this data dissemination platform meets our expectations from both the technical and user perspectives. how useable bank and system staff will find any tool is one of our primary concerns. while we were not able to engage in 38 iassist quarterly 2016 iassist quarterly any formal usability testing, we did want to get informal feedback from potential users. we worked with three volunteers to address the following questions. • how long do users need to complete the entire submission workflow and do they consider the process burdensome? • do the metadata fields defined in the pilot describe the data sufficiently? • are users willing to provide supporting documents and further descriptions, such as a data dictionary, in the submission process? • how effective is the user interface? do users suggest any improvements? • what is the role of the data curator and how could this individual help improve the efficiency of the workflow? each user completed the test independently to diminish peer influence. during the test, we instructed each participant to walk through the submission workflow: logging in, entering the necessary metadata, and then assembling and submitting the data file. once the test was completed, the participant was asked to provide feedback on his or her overall impression of the tool and the effectiveness of the workflow, as well as suggestions on how to improve the user interface. users found some steps in the workflow challenging at first. they received registration emails for access to the platform that they weren’t sure how to handle and for which we had failed to prepare them. even once they understood the registration process, they were somewhat stymied by a technical difficulty peculiar to this pilot: the submission process relied on two separate websites for different parts of the workflow. during the pilot, the two sites were not well integrated, resulting in some confusion as users navigated between two similar looking, but disconnected, web pages. we do not expect this problem when the platform is commercially available. once the users understood where to start, we observed them while they completed the two parts of the submission workflow: metadata entry and data file assembly. metadata entry required the users to fill out all of the mandatory fields and presented optional fields as well. by limiting the number of required entries, we hoped to improve efficiency and reduce the cost to researchers for publishing their data. none of the users we observed seemed to notice the distinction between required and optional, so they simply filled out all of the blank fields. the next step, data file assembly, required users to upload data from their computers to the kansas city fed staging area for the platform. as previously mentioned, users encountered some technical difficulties with the connection to the storage system, and only two of the three testers were able to successfully construct and upload a data file. user feedback overall, the users reported that the amount of time spent on submission was not burdensome. they also reported that a majority of the metadata fields were easy to fill out. however, the users offered a few suggestions regarding metadata and user interface at the end of the test: • enhance the metadata selection interface users were asked to select terms from a controlled vocabulary specific to economics and finance. to encourage consistency across the federal reserve system, we used a list maintained by colleagues at the federal reserve board of governors to describe their research publications. the list was a flattened hierarchy of more than 70 lines and was difficult to navigate in a drop-down menu. users suggested retaining the hierarchical structure but splitting the long list into multiple drop-downs to improve navigation and readability. • add detail to some fields the journal of economic literature (jel) maintains an alphanumeric, hierarchical classification system that is the standard for classifying scholarly literature in the field of economics10. the interface for submission used the text description for the economic field without including the 2-3 digit identifier. because economists are so familiar with the jel scheme, users suggested adding the code designation next to the description. • clarify metadata fields the users found a few metadata fields ambiguous and suggested clarifications and examples to improve the workflow. for instance, we named one metadata field ‘data restriction audience’ to contain information on time restrictions (when can the data be shared) and access restrictions (with whom can the data be shared). users weren’t sure what the name implied, and they certainly did not understand our intended usage. the field labeled ‘documentation’ also confused users. we expected users to provide supporting documentation such as a data dictionary. one economist expressed willingness to do so but was unsure how much detail was required. he also pointed out that as the definition of certain data variables has changed over the years, providing information on how the data was constructed and elaborating on the difference would add high value to data dissemination. • allow options for restricted data many of the researchers in the federal reserve system are assigned to work with restricted data that can only be shared with certain audiences. the user who volunteered to test the platform using restricted data noted that some of the data could still be made available after sensitive information was extracted. creating versions of the data that can be shared more broadly will definitely iassist quarterly 2016 39 iassist quarterly require more work on the researchers’ side, but will eventually be beneficial to the research community within the federal reserve system as well as to the public. • ensure persistent identification all three test users expressed concerns with the permanence of the url for their research output. as the purpose of data dissemination is to make data available, the economists wanted to know how we would ensure consistent access to the data once published. cadre staff were already working on a program to assign digital object identifiers (dois) to research output, which would ensure current location information for files and provide identifiers to published datasets. curation testing in the curation workflow, the curator reviews the submitted metadata and the data file, adds additional information as needed, and approves or rejects the submission. though we had early success in curating entries in the pilot, we were not able to curate the data the three users submitted during this phase of testing. there appeared to be some technical issue that neither the platform developers nor the kansas city staff could identify or resolve. after a seemingly unrelated patch application, the problem disappeared. all curation attempts thereafter were successful. overall, the curation workflow functioned as anticipated and worked well. we are considering changing the workflow so that the curator assembles the data file instead of the data submitter but have not yet tested this possibility. next steps to accomplish the objectives set at the beginning of the pilot, we need to take the following steps: • repeat the usability testing with the three economists to ensure all of them are able to submit metadata and upload data files from their computers. we will modify the curation testing to determine whether the curator is able to assemble data files for the data submitter and to ensure the submitted metadata and data are successfully curated. • finalize the dissemination workflow to include doi assignments which were created locally and temporarily during the pilot. we will investigate how the publication platform integrates with our doi registration service. • work with technical staff to verify the security settings of the platform. ideally, we would implement a data dissemination platform for both public data and restricted access data. one important aspect of this pilot is evaluating the access settings of the platform to ensure they meet our security requirements. • expand this pilot to other interested participants such as researchers in other federal reserve banks and the federal reserve board of governors. we anticipate involving more potential users and collecting feedback from them, particularly on their experiences with metadata entry, data submission, and identity authentication. references arzberger, p. et al. 2004, ‘promoting access to public research data for scientific, economic, and social development’, data science journal, vol. 3, no. 29, pp. 135–152. available from: . [24 june 2015]. borgman, c. 2012, ‘the conundrum of sharing research data’, journal of the association for information science and technology, vol. 63, no. 6, pp. 1059–1078. available from: . [24 june 2015]. the data deluge. 2010. economist. available from: . [24 june 2015]. notes 1, san cannon, federal reserve bank of kansas city. email: sandra.cannon@kc.frb.org. 2, deng pan, federal reserve bank of chicago. email: deng.pan@chi.frb.org. 3. http://federalreserveonline.org/ 4. http://www.federalreserve.gov/pubs/feds/2005/200533/200533abs.html 5. http://www.newyorkfed.org/research/staff_reports/sr493.html 6. https://www.kansascityfed.org/research/cadre/ 7. http://www.icpsr.umich.edu/icpsrweb/deposit/ 8. https://www.whitehouse.gov/sites/default/files/omb/memoranda/2013/m-13-13.pdf 9. https://project-open-data.cio.gov/v1.1/schema/ 10. https://www.aeaweb.org/econlit/jelcodes.php sist newsletter vol.1, no. 1 certain definitional problems are now being addressed by members of the us action group: what do we mean by a social science data archive? how does it differ from a social science data library? joseph bonmariage has suggested that there are different levels of production and dissemination. the issue of what constitutes "social science" has been raised. david nasitir suggested that "everything is {or can be) grist for the social scientific mill." he has proposed that by social science we mean "used by social scientists" rather than "produced by social scientists." this broadens the definition to include, for example, original or secondary distributors of process-produced data, a major social science resource. the directory would be developed in machine-readable form in order to facilitate updating and would either be published annually or available on a subscription basis. it will include entries for both archives and libraries and would include an appendix listing known organizations for which no further information is available. a questionnaire is being prepared for distribution by the first of the year and a preliminary version of the directory should be ready in april, 1977. in the meantime, a "primitive biliography" of existing data archive "registries" is being compiled which might appear in the directory in the form of an annotated bibliography. data acquisition canadapierre lacasse, centre d-^ recherches en am^nagement regional, universite de sherbrooke, sherbrojke, quebec europemarcia taylor, social science research council survey archive, university of essex, wivenhoe park, p.o. box 23, colchester, essex, england c04 3s0 united statesdonald harrison, national archives (nnr), washington, d.c. 20408 mandate this action group addresses the problems of data acquisition for archives with particular emphasis on the necessary relationship between the collectors and creators of data and the archives. recommended procedures for the acquisition of data would be developed with the intent of assisting researchers at critical points during the data collection process to ensure and promote the transfer of high quality data to the public domain for further academic investigation. the group will also survey different acquisition policies already in use, with attention to both legal and procedural problems connected with the acquisition and de-acquisition of data . [editor's note: the original mandate included the statement, "a survey of confidentiality laws, implications, and problems existing in machine-readable data would be conducted. a report summarizing the survey and resolving the problems when possible would be produced." this activity has been transferred to the action group on process-produced data. the under-lined section of the mandate of the data acquisition action group reflects an enlarging of the scope of activities to be undertaken.] 1 0. sist newsletter vol.1, no. 1 activities and plans a comprehensive questionnaire on archive acquisition policies has been drafted by marcia taylor and is now being reviewed by action group members in north america and europe. the questionnaire addresses the criteria for selection including contractual obligations, medium of storage, geographic scope, field of study, quality, selection procedures, and incentives and sanctions for encouraging/discouraging deposits. whereas other action groups are concerned with various means of making data accessible to the user, this action group will focus on the problems relevant to the initial archival acquisition. related to this issue is the possibility of joining with other organizations in providing incentives for both academic and governmental data producers to make their data files more generally usable and more widely available. the action group intends to address questions of deacquisition as well as acquisition policies to the needs of a local service data library. data documentation canadadave l. salley, management and central services group, standards division, statistics canada, tunney's pasture, ottawa, ontario kia 0t6 europecees middendorp, steinmetzarchief , kleine-gartmanplantsoen 10, amsterdam-c. , netherlands united statesjohn grasso, office of research and development, center for appalachian studies and development, west virginia university, morgantown, west virginia 26506 mandate this group will develop standards for the measurement of variables, i.e., the definition of constructs and their operational isation and measurement. in principle, these standards should be applicable cross-nationally, although this may not always prove to be possible. standards will be developed for "simple background variables" used in surveys, i.e., educational level, age, head of household, as well as constructs such as job satisfaction, anomia, political interest (i.e., to be measured by a scale or index). thus, the work of this group will be closely linked to that which is going on regarding the development of social indicators. the codes will be incorporated into source books to provide researchers with a resource tool for coding and organizing their data consistently. vol252 iassist quarterly summer 2001 9 a peculiarity of the development of social sciences in the former soviet union, and especially in estonia was that conducting of empirical social studies was possible from the 1960s on, while teaching of social sciences and participating in the activities of the world sociological community was hardly possible. teaching of sociological disciplines as well as engaging of students in the research was extremely limited until the end of the 1980s. in 1993 a team of sociologists, psychologists, political scientists and human geographers from the university of tartu made up an initiative group for creating a data bank on social sciences and began to work out the strategy of saving and usage of the research material collected by the estonian social scientists during the previous decades. in summer 1994 a project for creating a data bank was presented to the open estonia foundation, or more concretely to the higher education support project (hesp). the application for support received a positive response; a grant aimed at “creating social science data bank for modernization of education in social sciences and training in investigative journalism” was awarded for years 1994 1996. in summer 1995 the office of the data bank was opened in the faculty of social sciences of the university of tartu. the data bank was officially formed as an interdisciplinary centre of the faculty of social sciences in early 1996, and it began to function as a national social science data bank the estonian social science data archives (essda). the nine-member essda council includes three persons from tartu university, three representatives from academic centers outside tartu university and the same number from non-academic institutions. essda co-operates closely with the academic union of estonian sociologists as well as non-academic institutions that are conducting social research. the work done in the initial period of founding essda was discussed at an international conference “ data archives and their functions in social research in eastern europe” held at tartu university in december 1996. representatives from 7 countries attended this conference. in 1997 essda became a full member of the european council of social science data archives cessda. the data collection of essda consists of more than 200 empirical social studies covering such research directions like opinion polls, youth and media studies, entrepreneurship, rural sociology and many more, carried out between 1971 and 2000. essda holdings are increasingly being used for academic purposes by undergraduate and graduate students as a basis of secondary analysis. it should be mentioned that up to now sociology undergraduate and graduate students have dominated among the archives customers, but there are also graduate students from the departments of journalism and political science, the faculties of philosophy and medicine, and the estonian agricultural university situated in tartu. training courses based on the essda’s holdings have been included in the curricula of various disciplines taught at the faculty of social sciences. in 1998 essda was awarded a hesp grant to introduce the estonian social science data archives systematically in the teaching of the social sciences at the bachelor, master and doctoral levels, and in the open university by using the facilities of the internet and creating an electronic journal. the use of the holdings of essda for practical purposes has begun to grow considerably. here the most important direction is to meet the needs of the institutions of public policy. permanent contacts have been established between essda and the economic and social information department (esi) of the chancellery of riigikogu /parliament of estonia/. first steps have been made in meeting information needs of local authorities. essda will be acting as a mediator between the institutions of local policymaking and the social scientists’ community. in 2001 the preliminary achievement was reached to deposit state-financed surveys at essda. essda’s international contacts of different scope have been established with the data archives in different counby rein murakas & andu rämmer1 estonian social science data archives: past and future perspectives 10 iassist quarterly summer 2001 tries. representatives of essda have visited data archives in the united states, sweden, germany, hungary, denmark, spain, norway, and finland. in april 1997 a cooperation agreement was signed with the world’s largest research center of public opinion the roper center at the university of connecticut. the agreement foresees the exchange of data and training possibilities for estonian colleagues at the roper center. cooperation with the finnish social science data archive (fsd) is speeding up. in connection with the preparations for the creating a common database about estonian-finnish joint social science research projects fsd equipped essda with the new www server to promote essda’s transition to the nesstar solutions in 2001. essda’s home page in the internet helps to spread information about estonian social research to the international community. in 1998 essda moved into new rooms provided by tartu university. essda’s activities have been supported by the chancellery of riigikogu /parliament of estonia/. by the help of the open estonia foundation essda began to compile the first estonian social science electronic journal “estonian social science online”. the assistance from the german data organization gesis enabled the translation of a number of study descriptions of significant studies into english. in december 1998, modernized essda’s internet home page was positively mentioned at the competition of civil aid and info pages organized by the oef. due to the chronic lack of finances in 1999 and 2000 essda could only continue the description, systemization and cataloguing of studies. in addition the second issue of the electronic journal “ estonian social science online” was published. essda’s internet page has steadily been improved and complemented. mutual meetings and development of co-operation projects with colleagues from the finnish social science data archive (fsd) should be mentioned. 1. essda, room 217/218, 78 tiigi street, 50410 tartu, estonia. phone: +372-7-375 931. fax: +372-7-375 900. e-mail: socarch@psych.ut.ee. internet: http://psych.ut.ee/ esta/ vol252 4 iassist quarterly summer 2001 at the iassist/ifdo conference in amsterdam in may of 2001 a session on “new archives (forum)” was chaired by paul de guchteneire (unesco) and brigitte hausstein (gesis) with participants from new and emerging archives as well as some from the “old world” of existing data archives. the chairs of the session have succeeded in publishing the collection of papers. the reason for re-publishing the session in the iassist quarterly is to spread the word about the archive movement in eastern europe to a broader audience of the full iassist membership and iq readers. in this brief introduction i will take the opportunity to thank the chairpersons and editors brigitte hausstein and paul de guchteneire for their effort and also to thank the authors from several countries in the eastern europe. for a more detailed introduction to the papers you may refer to the introduction by the chairpersons. karsten boye rasmussen march 2002 editor's notes introduction editors: brigitte hausstein, gesis branch office berlin/ central archive cologne and paul de guchteneire, unesco/most, paris. the “new archives forum” at the 2001 iassist/ifdo conference “a data odyssey – collaborative working in the social science cyberspace” the international association for social science information services and technology (iassist) held its 27th annual conference with the international federation of data organizations (ifdo) from may 14 19, 2001. the conference was convened in amsterdam and hosted by the niwi (nederlands instituut voor wetenschappelijke informatiediensten netherlands institute for scientific information services) and the scientific statistical agency of the netherlands. during the last quarter of the 20th century, iassist has held its annual international conference in europe about every four years. iassist did so again, this year in amsterdam in collaboration with ifdo. about 200 data and archive specialists from the usa, canada, eastern and western europe as well as from africa and asia attended the conference. the conference was structured in three different domains of potential sharing of information, experience and expertise: first were organizational matters, like acquisition policy, access to data, and setting up new national or topical archives. second was meta data: standards, tools, and new developments. and third was content: various (new) data types like qualitative, aggregate data, multi-media or geo-data, or combining data from registries. the pre-conference workshops which are a tradition at iassist/ifdo conferences, introduced the participants to the a new standard for meta data documentation – the data documentation initiative (ddi). the most often discussed issue at the conference centered also on the ddi and special attention has been devoted to the applications of the ddi and to the experiences in the data archives. iassist conferences also bring together colleagues from new data archives around the world. in amsterdam a special session for these new archives was organized for the first time. the main focus of the session was on data archives in eastern europe but also representatives of new archives from japan, finland, ireland and greece joined the meeting. the session chaired by paul de guchteneire unesco/most (management of social transformations program), paris and brigitte hausstein gesis (german social science infrastructure services) branch office berlin/central archive cologne comprised two case studies from slovenia and south africa and a forum of representatives of data archives from hungary, estonia, latvia, czech republic, slovakia, slovenia, russia and romania. i^ssist newsletter vol.1, no. 1 data archive development canadalaine ruus, data library, computing centre, university of british columbia, 2075 wesbrook place, vancouver, british columbia v6t 1w5 europenot activated united statesalice robbin, data and program library service, 4452 social science building, university of wisconsin-madison wisconsin 53706 mandate a procedures manual consolidating current archival organizational, administrative, and personnel structures, procedures, and policies as well as recommended guidlines will be compiled to aid developing archives and to act as a resource guide for existing archives. workshops and seminars at various levels will be sponsored to provide professional training in the skills necessary for effective operation of a data library, data archive or social science information center . [editor's note: a new action group on data organization and management has been constituted to focus on the substantive aspects of data base organization and management. this action group is described more fully below. the focus of the data archive development action group will be on the more general functional aspects of establishing, managing and operating a data library, archive, or information center. the underlined section reflects this change.] activities and plans this past summer, a very successful workshop on data library organization, management and user services was held under the auspices of the inter-university consortium for political research as a pilot module of the icpsr's summer program. it was organized and jointly sponsored by lassist. it is hoped that a similar effort will be sponsored by icpsr and lassist next summer. [editor note: see p. for full description of the workshop.] members of this action group are developing a detailed outline for the proposed procedures manual. the outline will be circulated to action group and steering committee members for connent and revisions by february, 1977. the action group is also collecting administrative and record-keeping forms from existing archives. the collection of forms will be annotated for inclusion in the manual as well as available separately. 1 4. i^ssist newsletter vol.1, no. 1 at the lassist-ipsa panel, martinotti presented an analytic framework for the historical and present uses of process-produced information by governmental, administrative and research institutions and the implications of these data for the european data archives and of the growing demand for policy-oriented data. some of the implications he identified included the linking of different data bases, quality of the data, and the political and institutional nature of the consequences of publicly supported social science research. at the lassist-ipsa panel, erwin k. scheuch provided the session's attendees with an historical overview of the development of information systems for storage and management of large data bases. he described the problems encountered with these information systems, some of the large data bases which have been organized and the problems utilizing them, the relationship between the data collectors who have been primarily the public agencies and other communities such as researchers, and the relationship between the data archive and public agency. [editor's note: his paper will be forthcoming.] further input from action group members will help the group coordinators decide which projects are of most immediate interest, but tentative plans call for a continuation of the study of confidentiality laws; a preliminary pilot survey of available files in order to define the magnitude of this potential resource; a documentation and coding manual specifically for aggregate data; and, a survey of completed resenrch using process-produced data. paul mdller has suggested the following priorities: (1) an overview (survey) of process-produced data for each country of interest; (2) decisions and guidelines on which information should be put into machine-readable form; and, (3) recommendations to public administrators for preservation of materials for social scientists. it is likely that given the enormity of the tasks outlined by mdller and leavitt and other issues identified in the moller memo, these issues might well be handled by other action groups, but these matters are still fluid. this action group will be interacting with quantum, the us association for public data use, and other organizations to address subjects of mutual interest. data organization and management canadagreg morrison, social science data archive, department of sociology, carleton university, ottawa, ontario kis 5e6 europeeric tannenbaum, social science research council survey archive, university of essex, wivenhoe park, p.o. box 23, colchester, essex, england c04 3s0 united stateswilliam gammell, social science data center, university of connecticut, storrs, connecticut 06268 vol21.2 52 iassist quarterly introduction researchers who work with large sequential datasets are often limited in the kinds of analytic strategies they can use because of the sheer size of the data. automated techniques for analyzing sequences were developed in the 1960s by scientists studying dna, rna, and proteins. in a classic volume on sequence analysis, sankoff and kruskal (1983) demonstrated its potential application for subjects as diverse as bird songs and macromolecules. in other work, andrew abbott developed “optimal matching” for sequence analysis in the field of sociology. in this paper, we describe a technique for analyzing sequences using regular expression matching (rem). this technique allows researchers to examine patterns in longitudinal data by condensing sequences of events into smaller, more tractable units. we also briefly discuss the development of a database structure that facilitates this kind of analysis. although all sequence analyses compare linear arrangements of symbols, whether in human behavior or dna, they differ in their assumptions about what makes two sequences similar or different. sequences in their original form often contain too much detail for useful comparison, since the possible permutations of occurrences can be limitless. therefore, in all cases researchers must create the rules that define sequence similarity for their analyses. methods for determining sequence similarity are often referred to as sequence-matching algorithms. these algorithms are mathematical, and compare sequences without reference to the semantic or theoretical structures that created them. when using such methods, researchers who wish to place their analyses in an appropriate context must carefully define what events represent the phenomena of interest. the technique described in this paper was developed as an alternative to existing algorithms and allows researchers to identify sub-patterns of events within sequences at the start of their analysis, based on theoretical or practical considerations. because this technique operates on a single sequence at a time, it is faster than processes that require comparing many sequences to one another. the project to illustrate rem, we will describe how we used it to analyze the sequence of events that led to a child’s placement into foster care in three states: illinois, michigan, and missouri. we were looking for systematic demographic and geographic differences among children that correlated with the events they experienced in the child welfare system. rem was developed to describe and compare the pathways the children took through this system. the data were derived from the administrative data systems of the illinois department of children and family services, the michigan family independence agency, and the missouri department of social services. preliminary data processing we received two data extracts from each state: one covering investigations of child abuse and neglect in the child protection system (cps), and the other, services such as foster care to children in the child welfare system (cws).1 we began by creating a project database for each state with the same essential structure. each state’s database contained tables for cps and cws data and one table for demographic information on the children. next, we created an event table that contained all of the administrative events for all of the children in each system. we then transformed each child’s events into a sequence variable or “history.” finally, we used regular expression matching to formulate “careers” by reducing the history sequences. at each step in the process, we preserved enough information from the previous step to retain flexibility in the subsequent steps. as the categories became broader at each step, the comparability of the data across states increased. categorizing event sequences using regular expressions by lisa sanfilippo & john van voorhis* summer 1997 53 creating the event table in this analysis we focused on four key administrative events: (1) indicated investigation, an investigation in which credible evidence of abuse/neglect was found, (2) unfounded investigation, an investigation in which no credible evidence of abuse/neglect was found, (3) case opening, when a case was opened for child welfare services, and (4) placement, when a child was placed in a foster home or institution. we created one record for every event a child experienced in either the cps or the cws. we then coded every record with a number denoting a particular event type (see “event codes” in table 1). these records contained the child’s id, an event date, and an event type code (see table 2). creating the history sequences we transformed each child’s event records into a single sequence of codes, since as separate records the table structure was not appropriate for sequence analysis. to make the programming and its interpretation easier, we used only single-character codes in the history sequence. although each history code represented a single event, a given code value could represent more than one type of event (see “history codes” in table 1). 54 iassist quarterly we first reviewed a frequency distribution of the history sequences to identify the most common sequences and to see the repetition of patterns within and among sequences. this review also revealed data entry errors that we could correct or eliminate, such as children receiving services before their birth or children being born multiple times. although we had anticipated that the variation in the patterns between sequences would make them unsuitable for analyses in their present form, we had not foreseen the amount of variation in the length of the sequences. for example, examining the distribution of event sequences revealed that many children experienced only one event, while others experienced up to fifty. this wide variation in length made it difficult to make meaningful comparisons among cases and suggested that we needed a method that would not rely solely on whole-sequence comparison. therefore, we focused our attention on identifying the sub-patterns which we had observed in the sequences. creating the career sequences one goal of our research was to elucidate the connection between cps investigations and a child’s subsequent placement in foster care. we had three initial questions: (1) what sequences of investigations never resulted in a child welfare case opening and placement? (2) what sequences of investigations resulted in the child’s first placement? and (3) what sequences of events resulted in the child entering the system without an investigation? because of our extensive work with the illinois data and our contact with all three states regarding current and past practices and policies, we had some knowledge of what the most common patterns of events might be. the following examples illustrate how this prior knowledge provided us with clues about what patterns to focus our attention on: • we understood that the number of investigations a child experienced was not a critical factor in the caseworker’s decision to place the child in foster care. we knew that children with histories composed solely of unfounded investigations were almost never provided with services, despite repeated contact with the department. therefore, we believed that the number of indicated investigations would predict placement better than the raw number of investigations. • we knew that, in one state, caseworkers were reluctant to remove children from their homes after only one indicated investigation unless they were in imminent danger. thus, we expected that a child with one indicated investigation would be less likely to be placed into foster care than a child who had two or more indicated investigations. • in all three states, we knew it was possible for children to experience a case opening and placement without an investigation of abuse or neglect, but we had no information on the frequency of such occurrences. • our prior analyses of the foster care data indicated that once in foster care, a child could move between placements numerous times before being returned home. although the placements could be of different types, the child was still living away from his or her parents. as a result, we chose to treat a series of placements without a return home as one career event. regular expression matching it became apparent in looking at the sub-patterns that they could be represented by regular expressions, a notation used widely in the computer science field for specifying and matching sequences.2 (see appendix.) we created a file listing the regular expression patterns we had decided to analyze along with a “career” code for each pattern which is shown in table 4. we grouped the patterns in passes because we knew that certain patterns occurred only at the very beginning of the history and we needed to control the generation of the matching program. the first pass was used to summer 1997 55 remove any events that occurred before a child was born. since we were especially interested in the first series of investigations, we created a pass that only matched to initial investigation sub-sequences. the last pass, which was applied repeatedly until the history was exhausted, contained all of the sub-patterns we were investigating. from this pattern file we generated a series of programs to transform the data. we used the awk programming language for both our program generator and the matching programs themselves. an awk program is composed of a series of pattern and action pairs. it automatically reads through data files one line at a time, and each line is matched against the patterns in the order they are listed in the program. when a line contains data that matches one of the patterns, the action associated with that pattern is executed. the patterns may contain regular expressions, while the actions are written in a language similar to the c programming language. in our project, the program generator read the pattern file containing the sub-patterns of interest to us and generated a series of programs that used those regular expression patterns to process the history data. each program in the series corresponded to a particular pass in the pattern file. if a pattern matched to the beginning of a history sequence, the matching characters were removed and the career code for that pattern was appended to the career sequence. the child’s id, history, and career were then passed to the next program for the next pass. the final program passed the data back to itself until the history sequence was empty or until a fixed number of passes had been run. if the history sequence was completely matched, a lower case ‘x’ was appended to the career to indicate completion. an upper case ‘x’ was appended if more history remained after the maximum pass limit had been reached. 56 iassist quarterly analyzing the career sequences since our analysis was limited to examining the sub-patterns that led to a child’s first placement, we did not analyze children’s entire careers. instead, we only analyzed the first four career events after a child’s birth. because the rem approach simply recoded the original history sequences, it preserved the unit of analysis, thus allowing us to attach explanatory variables such as year of first entry into the system, sex, race, and region3. once this information was stored in one file, we aggregated the data by creating a crosstabulation which contained frequencies for every combination of the career sequences and the explanatory variables. these files were relatively small (fewer than 1,000 records) allowing us to import them into a spreadsheet program for final analysis and presentation. conclusion the rem technique described in this paper departs from more common pattern matching methods in that it incorporates theory and practice into the actual matching process. using this technique, researchers can test their assumptions about the structure of a sequence. it is an iterative technique that allows the analyst to explore patterns in the data and to compare them across populations simply and quickly. because the process of developing the career file is split into several steps (i.e., creating the event table, creating the history sequences, and pattern matching), it provides many opportunities to check the data and to ensure that the processes are transforming the data correctly. rem allows the researcher to take a very large dataset and to represent it in a much smaller form, while maintaining the critical details of event order and sequence. for example, in our illinois database we began with an event file of over 5 million records. transforming this file into history sequences, career sequences, and finally into a crosstabulation, decreased the size of the file by a factor of 5,000, making it significantly easier to work with. the rem technique, as written in awk, can save the researcher hours of processing time, in large part due to: 1) the way awk reads data files (i.e., it automatically reads a file one record at a time) and 2) the minimal programming it requires. performing the same analyses using a statistical software package would have required much more extensive programming and perhaps more important, would have restricted the kinds of questions we could have asked in exploring the original data. future directions clearly, rem has a much wider application than what we have illustrated with our project. our analysis did not utilize rem to its fullest potential. for example, instead of analyzing just the initial sequence of sub-patterns, rem could be used to analyze full careers. we could run a similar process against the career sequences to further shrink the number of categories. finally, we did not explore the sub-patterns in as much detail as we could have. for example, we included specific placement event types in our event table and history sequences but did not treat them as separate types. in the future, we can easily compare differences in children’s histories following specific types of substitute care placements (e.g., home of a relative, private foster home, group home, etc.) based on this project’s current database. appendix regular expressions in general, a character in an awk regular expression matches itself. some characters with special meanings in our pattern file are listed below along with some examples of their use. see the references for more details. special characters used in regular expressions: summer 1997 57 regular expression examples: references abbott, andrew. 1995. sequence analysis. annual review of sociology, 21:93-113. abbott, andrew & alexandra hrycak. 1990. “measuring resemblance in sequence data: an optimal matching analysis of musicians’ careers.” american journal of sociology, 96(1): 144-185. abbott, andrew & john forrest. 1986. “optimal matching methods for historical sequences.” journal of interdisciplinary history, 16(3): 471-494. aho, alfred v., brian w. kernighan, & peter j. weinberger. 1988. the awk programming language. reading: addisonwesley. aho, alfred v., jeffrey d. ullman. 1979. principles of compiler design. reading: addison-wesley. forrest, john & andrew abbott. 1990. “the optimal matching method for anthropological data: an introduction and reliability analysis.” journal of quantitative anthropology 2:151-170. friedl, jeffrey e. f. 1997. mastering regular expressions. sebastopol: o‘reilly & associates, inc. sankoff, david & joseph b. kruskal eds. 1983. time warps, string edits, and macromolecules: the theory and practice of sequence comparison. reading, ma: addison-wesley. notes 1 . child protection systems: in illinois, the child abuse and neglect tracking system; in michigan, the protective services management information system; and in missouri, the child abuse and neglect data system. child welfare services systems: in illinois, the child and youth centered information system; in michigan, the children’s services management information system; and in missouri, the alternative care tracking system. 2 . regular expressions (res) can recognize patterns which are left linear. patterns, such as a balanced sequence of parentheses, cannot be recognized by res because such patterns require “going-backwards” or maintaining information outside of the re. for further information see aho, kernighan, and weinberger 1988 in the references. 3 . in all three states we differentiated the major urban area from the balance of the state. paper preseneted at the iassist/ifdo 1997 annual conference odense, denmark may 7, 1997 by 24 iassist quarterly summer 2007 lynn woolfrey* the establishment of the african association of statistical data archivists (aasda) aasda represents practitioners in survey data curation in africa and was established to facilitate co-operation among them with regard to the development and use of best practices in the preservation and sharing of survey microdata in the region1. this association was established with the assistance of international organisations promoting optimal management of survey data. these included the international household survey network (ihsn) and the international association for social science information service and technology (iassist). background and role of the ihsn/adp the ihsn was established in 2004 as a recommendation of the marrakech action plan for statistics, an international initiative to build statistical capacities in developing countries2. the network aims to improve the quality of survey data in developing countries, and to promote data usage for research and policymaking in these countries. its membership is comprised of organisations that provide funding and technical support for survey programmes in these countries, and includes among others representatives from the uk’s department for international development (dfid), paris21, the un statistics division and the world bank. coordination of the network is undertaken by the world bank’s development data group (decdg). the ihsn has developed a programme of action to counter obstacles to survey data production and utilisation in developing countries. problems with data curation in these countries include limited funding, and a lack of technical resources and expertise to collect and manage survey microdata. lack of co-ordination of international programmes designed to support statistical development in these countries has hampered their effectiveness. the ihsn’s plan of action includes the co-ordination of donor programmes supporting data curation, and the provision of guidelines and technical tools for optimal survey data management to promote the production of quality data in developing countries. further plans to assist data management include conducting and maintaining data audits for relevant countries, and establishing regional data archiving organisations for networking to support data curation and data sharing among these countries3. guidelines and tools designed by the ihsn and other agencies are available via the ihsn website. they deal with the full survey life cycle, from sampling and questionnaire design to data archiving and dissemination. included is a question bank (the ihsn q-bank), a central, xml-based repository for international classifications, and links to international surveys. the website supplies information on data anonymisation as well as open source software for this purpose, and publications on metadata standards for data documentation. software designed by the ihsn includes the microdata management toolkit for documenting and disseminating data according to international standards. the network has also designed the national data archive (nada) toolkit, which is an open-source web application that allows the creation of searchable online data catalogues and data downloading via the web4. the ihsn does not provide technical or financial support to countries. the accelerated data programme (adp) was thus initiated in 2006 to expedite the goals of the ihsn. this project, financed by the world bank and implemented as a satellite program of the paris21 secretariat at the oecd, provides training and assistance in data curation to national statistics offices (nsos) in developing countries, utilising the tools designed by the ihsn.5 the ihsn, iassist and data management in africa in africa the work of the ihsn/adp partnership builds on the incremental advances in survey data management fostered by previous statistical capacity building projects of regional and international development agencies. however, the work of the partnership has led to major advances in the curation of national survey microdata in several african countries. the project has the potential to realise the goal of data sharing in the region. the project team have achieved this partly through identifying and co-opting key organisations to assist them in this endeavour. important links were made when the ihsn co-ordinator attended a conference of the international association for social science information service and technology (iassist) in montreal, canada, in 2007. iassist represents data professionals working with information iassist quarterly summer 2007 25 technology and data services for social science research. the ihsn funded the participation of delegates from several african nsos at this conference. there the ihsn co-ordinator made contact with staff from datafirst, a survey data archive and training facility based at the university of cape town in south africa. this meeting led to the ihsn’s contracting datafirst to undertake nada installations and training in african nsos from 2008.6 the ihsn/adp/datafirst team are installing the nada data cataloguing and dissemination software in african nsos and providing training on the system. the nsos in gambia, ethiopia, liberia, nigeria and uganda have nada installations. the national institute of statistics in cameroon is at the testing stage of implementing the system. the mozambican nso is also being supported by this project, and aims to have a nada set up by the end of 20087. the software will soon be installed at nsos in ghana, niger, senegal and mali (this includes an installation at the afristat regional offices)8. the nigerian national bureau of statistics is the second african nso (after statistics south africa) to provide public use datasets and this has been accomplished by the implementation of the nada software. requests for these datasets have been received from researchers abroad and in africa, providing evidence of data demand within the region.9 a comparison of web searches conducted in 2007 with a recent (august 2008) search reveals that the technology and training provided by the ihsn/adp/datafirst team have served to revitalise the websites of african nsos involved with the project, which has further promoted data discovery via these sites (nso websites, 2008). the establishment of aasda the ihsn/adp’s links with iassist helped the ihsn accomplish their goal of creating a regional organisation to facilitate networking among data curation organisations in africa. this was initiated when african delegates at iassist’s 2007 conference formed an interest group concerned with data management in africa. founder members of this group included representatives from nsos in cameroon, ethiopia, gambia, mozambique, niger and uganda, and datafirst and the human sciences research council (south africa). forms of collaboration among group members initially involved exchanging ideas on data management in africa, and discussions regarding ways to collaborate on staff training. the idea for an association for african data managers was initiated by this group.10 the ihsn had identified a growing need among nsos and other data management organisations in africa for the establishment of a community of practice to work towards regional co-operation in the preservation and sharing of african survey microdata. thus, they were amenable to providing financial and logistical support for the establishment of an association of survey data managers in africa. the ihsn/adp met with the iassist africa group in kampala to initiate a meeting of representatives from african nsos and the two african survey data archives11 to establish an association of african microdata managers. iassist members and the ihsn/adp team were willing to provide support for this initiative in the form of technical and professional skills transfer. the inaugural meeting of aasda was hosted by the datafirst survey data archive at the university of cape town, south africa, on 10-12 april, 2008. it was attended by delegates from several african nsos12, the adp/ ihsn13, iassist, uneca’s african centre for statistics (acs), the uganda statistical society (uss), the african development bank (afdb), and afristat14. the meeting was opened by the south african statistician-general, pali lehohla, who discussed the problems involved in archiving survey data in africa. he expressed concern over the lack of indigenous data preservation organisation in africa, and bemoaned the fact that often the only extant datasets from some african surveys are currently archived in countries outside of africa. he felt that aasda could play a key role in supporting the preservation and reuse of high quality african survey microdata. the meeting represented the first occasion where data managers in nsos from francophone, lusophone and anglophone african countries had met to form a collaborative grouping. the meeting revised and adopted a constitution for the association, and elected office bearers. data managers from the nsos exchanged experiences on the use of the ihsn microdata management and dissemination tools, and discussed their future data curation plans. the ihsn/adp team gave an overview of work done with african nsos, and reiterated their support for data curation organisations in the region. delegates emphasised the need to obtain support for aasda from the senior management of nsos to encourage them to become aware of the advantages of sound data curation and to promote data sharing in the region. members of the association also stressed the need to cater to all language groups represented, and it was agreed the association should make all important documentation available in both french and english. delegates concurred with the ihsn/adp representative when he spoke of training as the key task of the association. among other decisions, it was agreed that aasda should undertake an audit of the training requirements of data managers and all regional data curation training programmes available in africa15. the aasda website has been implemented, with the technical assistance of the ihsn/adp team.16 with technical and financial support from international agencies such as the ihsn/adp, the association is mandated to work towards overcoming organisational obstacles to data production and data sharing in the region. they aim to 26 iassist quarterly summer 2007 advocate for the creation of good metadata for african datasets, based on international standards, and for the incorporation of best practices for the preservation and dissemination of data into their organisational agendas. members receiving training in the use of data curation tools can play a mentoring role on the continent, providing further training, nationally and regionally. the creation of an association of survey data archivists in africa can be seen as a step towards formalizing data sharing on the continent through fostering linkages among survey data managers in the region. the association will also be in a position to support the establishment of data management facilities on the continent. at this stage it appears that these will be housed in the nsos in african countries, as it is to these organisations that the financial and technical support is being directed. the formation of a group of dedicated practitioners in the field can serve to highlight the need for co-ordinated microdata curation in africa. the initiation of a network of survey data managers in africa represents the first step towards such collaboration. *contact: lynn woolfrey, data manager, datafirst resource centre, university of cape town, south africa. email: lynn.woolfrey@uct.ac.za. footnotes 1 aasda constitution, 2008. 2 http://www.mfdr.org/documents/marrakechactionplanfor statistics.pdf 3 ihsn website: http://www.surveynetwork.org 4 ihsn website: http://www.surveynetwork.org. see also the adp website at http://www.surveynetwork.org/adp. 5 adp website: http://www.surveynetwork.org/adp. 6 datafirst oecd contract, 2008. 7 tomas bernardo, mozambique instituto nacional de estadistica, at the aasda meeting, 2008. 8 olivier dupriez, ihsn, 2008. 9 information from olivier dupriez, ihsn, 2008. 10 iassist notes and discussions with yakob mudesir, ethiopian central statistical agency. 11 the south african data archive (sada) and datafirst, both based in south africa. 12 these included the nsos of cameroon, ethiopia, gambia, ghana, liberia, mali, mozambique, niger, nigeria, south africa and uganda. 13 representatives from the ihsn present at the meeting included staff from the uk department for international development (dfid), the paris21 secretariat at oecd, and the world bank. 14 afristat is an umbrella body supporting statistical development in francophone africa. 15 aasda minutes, 2008. 16 aasda website: http://www.aasda.net. vol30-2neu.indd 4 iassist quarterly summer 2006 editor’s notes welcome to the second issue of the iassist quarterly, vol. 30. winter came late in denmark, but suddenly the situation was normal: snow came and traffic stopped. and now 14 days later spring is here. i sat outside in the sun reading articles. still it’s more difficult to write articles outside; we are waiting for improvements to the computer screens. however, some people have stayed inside to write articles for the iassist quarterly. they are presented below. the first article, “microdata information system missy,” is written by andrea janssen and jeanette bohr from centre for survey research and methodology, zuma, at mannheim in germany. the missy was presented by andrea janssen, jeanette bohr, and joachim wackerow at the session on “effective strategies for metadata management” at the iassist conference in may 2006 in ann arbor. the data in the system are from the german microcensuses for 1995 and 1997, which contain a sample of one percent of all german households. the microcensus has been carried out since 1957, and parts of the microcensus are available for research. the researchers need extensive metadata on both the study and variable level, e.g., the microcensus uses complicated classifications of professions, sectors, and household arrangements. the system is based on the standard from the data documentation initiative and documentation includes general information such as questionnaire, codebooks, interviewer guides, frequencies and also some tips and recommendations on use of the data. missy is a german system, in german language, and for researchers in germany. however, all are free to gain from the experiences presented in the article. the second article was also presented at the iassist 2006 conference, at the session “innovations in data dissemination.” the title of the article is “user-centered design and innovation in the sociometrics social science electronic data library (ssedl).” the authors, josefina j. card, tamara kuhn, and thomas wells, are all at sociometrics corporation. the article describes the sociometrics data archives ssedl as being a rich source of data for those in the public health, medical, nursing, social work, and social science professions. in the “product package,” datasets come with several data input files for sas and spss. purchasers can acquire data on cd-rom or download from the web site, and downloading has had a significant rise in usage. in order to help users identify relevant datasets, sociometrics has launched a topic-based drill down system showing “areas of richness,” which helps users identify and reach datasets and variables and the appropriate documentation. the last article is “overview of a proposed standard for the scholarly citation of quantitative data,” by micah altman and gary king from harvard university. this is an extended abstract summarizing a proposed standard for citation. this was presented at the iassist 2006 conference at the session “new standards in statistics and data citations.” the authors mention that, at a minimum, citations should include author(s), date of publication of the data set, and the data set title. these fields have been discussed before during the close to 50 years of documentation of data sets, and the fields are not so unambiguous now not to be discussed further. for instance, a great many people can be called “authors” in the production of a data set, and the same data set can have several other relevant dates attached besides “publishing date.” the authors recommend using the ddi elements, but their main purpose is to propose some novel fields that are directly linked to the use of modern technology. first of all, a unique global identifier. the authors mention a naming resolution service and that brings to mind the technology of the internet with name servers for looking up the correct ip address; however in this context, more than one copy of the dataset can exist at different physical locations. secondly, technology is applied by adding a universal numeric fingerprint. this should guarantee that the data set has not been changed even though the data set might exist in different software. this should probably apply to the documentation as well. the iassist is always open at its website, http:// iassistdata.org, where you can look at conference information and visit the iassist blog (iassist communiqué, http://iassistblog.org). articles for the iassist quarterly are most welcome. articles can be papers from iassist conferences, from other conferences, from local presentations, discussion input, etc. contact the editor via e-mail: kbr@sam.sdu.dk. karsten boye rasmussen, march 2007 vol29-2.indd 4 iassist quarterly summer 2005 editor’s notes welcome to the second issue of the iassist quarterly vol. 29. this issue contains three articles from the iassist conference in edinburgh in may 2005. at the session called "using national data" the paper on "economic data as snapshots in time" was presented by katrina stierholz from the federal reserve bank of st. louis. they are into names at this bank, their data is called fred (federal reserve economic data), and they have a fraser, that is an image archive of economic statistical publications, and an alfred that is a machine-readable archive with access for researchers to pull real-time data. katrina stierholz explains in the paper that while the fred data is a revised time-series, alfred and fraser contain the historical real-time data. with fraser and alfred the ability to reproduce other economists' work becomes possible. "looking back at decisions made, it is important to be aware of what the data at the time said, not what they say now". the article describes the systems fred, alfred, and fraser and what they contain and what functionality they offer. at the session "discovering a profession: the accidental data librarian" a presentation "looking for data directions? – ask a data librarian" was given by stuart macdonald (edinburgh university data library) and luis martinez (london school of economics data library). this presentation has been turned into the article presented here: "the local data support landscape in the uk". the article focuses on specialized national data centres. they start with a tax assessment from 7th century, mentions the domesday book from 1086, but quickly moves on to the establishment of the start of uk data archive in 1967, and others to follow. the office for national statistics produce many key statistics used in policies as well as in research and so does "the cousin" the general register office for scotland. the article describes the centers ukda, esds, edina, mimas, ahds and several more. connections between these centers exist such as the "data information specialist committee-uk" (disc-uk) that is a sort of national iassist for data librarians and data managers. in the session on "enriching metadata: the lifecycle perspective" the presentation "providing context for understanding: the data life cycle" was given by elizabeth hamilton from the university of new brunswick, canada. elizabeth hamilton has turned the presentation into the paper "providing context for understanding: insight from research on two canadian health surveys". she uses the national population health survey (nphs) and the canadian community health survey (cchs) as cases for evaluation of documentation using the ddi-format (data documentation initiative). the conclusion is that there is a need for placing the survey data in context, and that some information (metadata) from the earliest part of the data life-cycle is "integral contextual information and as such should be identified, described and preserved, in addition to the formal data collection itself". can we say we can trust the data? elizabeth hamilton want to be able to give a clear positive answer with the addition: "we have a complete record of evolution of that question, from concept to analysis." the iassist website is constantly evolving so remember to pay a virtual visit to http://iassistdata.org and to the iassist weblog (blog) iassist communiqué – at http://iassistblog.org. at the iassist website you can find information on previous and coming conferences as well as easy access to the articles of the iassist quarterly in the form of pdf-files. papers for the iassist quarterly are most welcome. papers can be from iassist conferences, from other conferences, from local presentation, discussion input, etc. contact the editor via e-mail: kbr@sam.sdu.dk. karsten boye rasmussen, february 2006 vol252 6 iassist quarterly summer 2001 the czech sociological data archive the sociological data archive (sda) of the institute of sociology in prague collects computerized data files from quantitative sociological surveys. its main objective is to make czech sociological data publicly available for academic, educational and other non-commercial purposes. other activities of the sda include the promotion of data dissemination and secondary data analysis, and support for special research projects. the sda is the only institution of its kind in the czech republic. institutional settings the sda has been open to the general public since september 1998. it was established within the project “social trends” (research archives publication graduate training) of the institute of sociology. the project was headed by dr. petr mateju and sponsored by the grant agency of the czech republic. since 1999 the sda has been an independent department of the institute of sociology, a non-profit social research organization operating within the academy of sciences of the czech republic. between 1999 and 2000 the grant agency of the czech republic financed a new sda project. thanks to it, the sda has developed into an infrastructure capable of providing data and other services on the customary level. in spring 2001 the sda became a member of the cessda (council of european social science data archives). at present the sda team includes three younger researchers; the archiving work is also supported by co-operation with other research teams of the institute of sociology. current institutional settings of the sda are defined in the statute of the institute and operational costs are covered from its budget. www links: sda: http://archiv.soc.cas.cz/ institute of sociology: http://www.soc.cas.cz/ social trends: http://www.soc.cas.cz/trends/ history of the czech social survey research and data archiving the czech social research has a long tradition, but the continuity of its development was deeply affected by the communist regime (1948 1989). the history of empirical surveys in former czechoslovakia dates well before the 2nd world war (the 1930s). in 1946 the american gallup institute inspired the creation of the institute of public opinion research which started systematic opinion polling. in the 1950s, with the exception of official statistics all activities in the field of social research were stopped on account of ideological reasons. sociology itself was denounced as a “bourgeois pseudoscience”, and all sociological institutes and faculties were closed down. in the 1960s research activities had been reestablished, but after 1968 were restricted again during “normalization” and put under the control of the communist party. the transformation of czech academic institutions of social sciences began soon after the “velvet revolution” in 1989, but the results in some areas have not yet reached satisfactory levels. furthermore, commercial survey research and public opinion polling on social topics has developed after 1989. archiving data from survey research in the czech republic the sda is the only institution which systematically provides access to data files from quantitative sociological surveys. before its formation, data files from sociological research projects were usually under the control of individual research teams. there was no systematic index of existing files, and poorly protected data were at risk of being lost or damaged. people interested in a specific data file had to negotiate with its owners. the lack of documentation and unsuitable formats often made it difficult to gain immediate access to data. a debate concerning the formation of a social data archive in former czechoslovakia was opened in the late 60s. at that time, the first efforts to establish an archive were related to the re-birth of czech sociology and the growing popularity of survey research. unfortunately, after the russian invasion in 1968 the communist control over social sciences was gradually reinforced. in the period of 1970 by jindrich krejci 1 iassist quarterly summer 2001 7 1989 the idea of systematic social data archiving reemerged, but it did not find a wide support among social scientists in view of the danger of abusing data for the purposes of the communist regime. a serious discussion on the establishment of a publicly accessible data institution continued after 1989. there were several attempts to establish the archive, and finally, the project “social trends” was successful in concentrating the necessary financial, personnel and institutional means, brought former ideas to fruition and founded the archive. official statistics in the field of official statistics data services are provided by the czech statistical office (csu). the csu publishes statistical yearbooks, regular reports entitled the indicators of economic and social developments in the czech republic, booklets “the czech republic in figures”, and a series of publications devoted to the labor market, family budgets, demography, etc. it is possible to access csu’s micro databases but it must be negotiated on an individual basis after contacting the office. several data modules from the surveys organized by the csu are part of international programs of eurostat, ceps/instead, cestat, etc. qualitative data in recent years two qualitative data archives have also been established. the czech archive of qualitative data and documents at the school of social sciences of masaryk university in brno was opened in 1999. its objective is to establish a publicly accessible information database of qualitative data collected in the czech republic. in 2000 a qualitative research project on “alternative culture” resulted in the establishment of the digital archive of soft data medard at the virtual institute in prague. information on both archives is available online. www links: czech statistical office (csu): http://www.czso.cz czech archive of qualitative data and documents: http://www.fss.muni.cz/qarchiv/ medard: http://medard.institut.cz/ sda: archived data the sda brings together computerized data files from quantitative sociological surveys. at present the data catalogue includes approximately 160 titles, some of which are, however, english versions of original czech data sets. data holdings include data files collected by the institute of sociology and other czech organizations conducting statefinanced sociological research, data from czech publicly available opinion polls and from international surveys with czech participation. the main access to the data library and other sda’s services is provided on the internet. the data processing launched by the “social trends” project focused on comparative projects in which the czech republic has been participating since 1990 and on research monitoring the main trends of the development of social structures in the czech republic. the further collection of data files has been aimed at earlier research projects of the institute of sociology, and the sda has also developed a co-operation with several other research institutes. the current largest project concerns the archiving of data from the public opinion surveys of the former institute for public opinion research (ivvm) from the period 1990 2000. since 1990 the ivvm has been organizing monthly public opinion surveys on attitudes to political, economical and social issues. 120 data files from regular surveys and approximately 50 data files from other ivvm research projects have been transferred to the sda’s library, the data have been transformed into the spss format, checked, cleaned and documented. a number of questionnaires from older surveys have to be scanned. the data files from surveys conducted in 2000 and 1999 have already been made available to the general public. other files are under preparation. www links: sda’s data holdings: http://archiv.soc.cas.cz ...than continue into “data archive” cvvm centre for public opinion research (former ivvm): http://www.soc.cas.cz/cvvm/ promotion of data dissemination and secondary data analyses sda was founded relatively recently and the tradition of using data services has not yet fully developed in the czech republic. international networks of data services are also little known. therefore, the archive has to pay great attention to promoting secondary analysis and employment of the existing data sources. sda publishes information on available data services in scientific and other periodicals, provides information to universities and public sector institutions, and organizes public presentations of the archive and data service networks. the team of the archive also participates in the educational programs of charles university in prague. in 1999 a course “social data archives” was lectured at the faculty of social sciences of charles university. the course will be taught again in the fall 2001. sda info archive’s information bulletin is issued four times a year in the czech language and is distributed free of charge. sda info provides an outline of available services, gives a more detailed overview of the stored data and research projects, provides references to other social data sources and is dedicated to promoting secondary data analysis. 8 iassist quarterly summer 2001 sda’s internet services provide an online access to a number of analytical publications and offer the option of ordering publications from the institute of sociology. the directory of references located on the sda’s server includes useful internet links to other social data sources, czech social science information, international survey researches, and general information on the czech republic. in addition to these data services, the archive has also become a source of more general information on czech society and social science research. in some areas the czech social science infrastructure has not yet been developed, and until recently the english language sources of general information on the czech republic were also limited. especially for foreigners, it is sometimes hard to orient themselves in the available sources. as a result, the archive often answers questions from completely different fields of interest than sociology. support of special research projects sda has co-operated in organizing research projects prepared within the institute of sociology especially the czech portion of international projects such as the international social survey programme (issp), the second international adult literacy survey (sials), and the european value study (evs). international collaboration the sda has co-operated mainly with the zentralarchiv für empirische sozialforschung (za) in cologne and with the gesis branch office in berlin. contacts have also been developed with the centre for the study of public policy at the university of strathclyde in scotland, tarki budapest and with the department of sociology at ucla. additionally, members of the archive’s team participate in several international research projects. at present the sda intends to join international networks of data services. in april 2001 the sda became a member of cessda (council of european social science data archives) and the application for membership in the ifdo (international federation for data organisations) is planned for the near future. references: krejci, jindrich. 1999. “sociological data archive in prague”. za information 45: 142-152. issn 0723-5607. krejci, jindrich. 1999. sda – sociological data archive. a library of computerized data files from sociological surveys. institute of sociology, prague. krejci, jindrich. 1999. ñsociological data archive“. in: ten years of rebuilding capitalism: czech society after 1989. eds: jiri vecernìk, petr mateju. academia. prague. krejcì, jindrich. 1998. “novy zdroj sociologickych dat” [a new source of sociological data]. sociologicky casopis 1998/3. sda info 2000: no. 1 4. issn 1212-995x. sda info 1999: no. 1/2 4. issn 1212-995x. vecernìk, jiri. 2000. “social reporting in the czech republic since 1989: the present state of art”. eureporting working paper no. 11. 1. jindrich krejci, sociological data archive, institute of sociology of the czech academy of sciences, jilska 1, 110 00 praha 1. phone: +420 2 22220098-0100 ext. 231. fax: +420 2 22221658. e-mail: krejci@soc.cas.cz vol19.1 36 iassist quarterly introduction the focus of cataloging the inter-university consortium for political and social research (icpsr) online codebooks is to provide users in a timely fashion adequate bibliographic information on virgo, the university of virginia library’s computerized library system. many catalogers today are cataloging materials that cannot be held in hand. gathering bibliographic information for electronic formats can be a bewildering and monstrous experience. the author shares her experience on how the fear of working with computer files was reduced to a minimum with the help of the computer support department, and the sense of triumph and accomplishment she felt when patrons successfully retrieved what they needed through the online catalog! background the icpsr is one of the world’s leading repositories and data dissemination organizations for machine-readable social sciences data. icpsr receives, processes, and distributes machine-readable data on subject matters covering over 130 countries. the content of the icpsr archive extends across the spectrum of economic, sociological, historical, organizational, social, psychological, and political concerns. icpsr was founded in 1962 as a partnership between the survey research center at the university of michigan and 21 universities in the united states. currently, membership extends to over 370 colleges and universities world-wide. the university of virginia (uva) faculty and graduate students (undergraduates with the permission of an instructor) may order icpsr data free of charge by contacting the official representative in the social science data center. there are currently over 350 studies available at uva. most of the studies consist of more than one dataset, and often include a study description and a codebook. the codebooks which are in print or machine-readable format describe the location of the variables in the data record. virgo, a notis-based online catalog provides online access to the library’s holdings through keyword, author, title, subject, and call number searches. it also offers online access to: nine periodical indexes published by the h. w. wilson company; current contents (an index to recent issues of over 6500 scholarly journals); abi/inform (provides citations and abstracts from business and management journals); and newspaper abstracts (indexes and abstracts articles in 28 major newspapers). cataloging project for online codebooks although, alderman library began a cataloging project for paper codebooks in the summer of 1994, the project of cataloging the machine-readable codebooks did not take place until january 1995 with the arrival of the original cataloger for electronic resources. before the project began for the machine-readable format, there were several issues that needed to be considered — 1) the number of titles incoming and in backlog, 2) the status of acquisition, 3) the procedure of retrieving bibliographic information, 4) the cataloging procedure, namely the procedure from getting actual text file to oclc marc format, and 5) the limitation for access of the materials. currently, the online codebooks reside in the machine named maggie under the directory, /archive/public/icpsr at the university’s information technology and communication (itc)2. they are accessible via all networks at the university. the files are sub-organized by the icpsr series number from 0001-9999. each individual series contains at least two basic files —the codebook and statistical data files. to catalog a codebook, the following information needs to be obtained —the actual title, statement of responsibility, edition, file characteristics, physical description, series information, publication or distributor, special note and terms of availability, etc. the chief sources for the bibliographic description are taken from the title screen and table of contents. take for example, census of population, 1910. united states: public use sample (icpsr ; no. 9166). to obtain the title proper, one can either pull up the title screen of codebook 9166 online from the sub-subdirectory of series number 9166 (see figures 1 and 2), or consult the paper format reference book, icpsr’s guide to resources and services, 1994-1995 . some of the series only have online codebooks, while others have both online and print versions. the social science data center at the university does not have the complete collection of neither online nor print version. series titles are added as the faculty place subscription orders which result in the increase of the database, thus the icpsr is a growing collection. once the bibliographic information is recorded, the title is searched against virgo and the national bibliographic utility, oclc. if a record is found in virgo for paper tackling icpsr online codebooks with success by jackie shieh 1 esrc data archive university of essex, 37fall/winter 1994 format (even if it is a different edition or release), the original cataloging for the computer file is created using the der (derived) feature from the existing bibliographic record. if no virgo record is available and an oclc record is found for a different format or edition, the oclc printout is used to create an original record on virgo. all original cataloging for machine readable datafile are created locally on virgo. currently, tape loading of machinereadable marc records without field 856 from virgo to oclc is not available. filter program using perl the initial stage of the icpsr project required the gathering of bibliographic descriptions on individual series numbers from the icpsr directory. that task seemed far more cumbersome than the creation of the marc record. it became even more overwhelming when it was discovered that new titles are checked-in and constantly added to the database. therefore, in order to manage the existing and incoming codebook files tracking uncataloged materials became the priority. jeff herrin, of the library system office wrote and explored a filter program using perl (see figure 3) to read all current files in the directory of /archive/ public/icpsr, from file 0001 to 9999. perl copied the first 70 lines of text files, which corresponded to ‘*.codebk*’, then copied the readings into an output file. the program would also generated an e-mail message notifying me even if there was no successful hit after the run. the first time perl was run we had little difficulty. a few minor adjustments were made, especially when two codebooks for the same series number were present. for instance, series figure 1 arch ive/pub l ic/ icpsr% d i ra rch ive/pub l ic/ icpsr% d i ra rch ive/pub l ic/ icpsr% d i ra rch ive/pub l ic/ icpsr% d i ra rch ive/pub l ic/ icpsr% d i r drwxr-xr-x 4 3180 public 8192 nov 23 09:07 0001-5999/ drwxr-xr-x 14 3180 public 8192 nov 23 09:07 6000-6999/ drwxr-xr-x 9 3180 public 8192 nov 23 09:07 7000-7250/ drwxr-xr-x 3180 public 8192 nov 23 09:08 7251-7500/ drwxr-xr-x 7 3180 public 8192 nov 23 09:08 7501-7750/ drwxr-xr-x 10 3180 public 8192 nov 23 09:09 7751-7999/ drwxr-xr-x 11 3180 public 8192 nov 23 09:09 8000-8250/ drwxr-xr-x 15 3180 public 8192 nov 23 09:09 8251-8500/ drwxr-xr-x 5 3180 public 8192 nov 23 09:09 8501-8750/ drwxr-xr-x 7 3180 public 8192 nov 23 09:10 8751-8999/ drwxr-xr-x 9 3180 public 8192 nov 23 09:10 9000-9250/ drwxr-xr-x 12 3180 public 8192 nov 23 09:10 9251-9500/ drwxr-xr-x 12 3180 public 8192 nov 23 09:10 9501-9750/ drwxr-xr-x 18 3180 public 8192 nov 23 09:10 9751-9999/ figure 2 archive/public/icpsr/9000-9250/9166% dir -rw-r--r-1 3180 public 408240 nov 28 11:22 i9166.codebk -rw-r--r-1 3180 public 50965936 nov 21 11:36 i9166.odata 1 8166 had two online codebooks, i8166.codebk and i8166.codebk01. when examining the output of each file, i found that it did contain sufficient information for creating a cataloging record— the size of the file, title proper, author, publisher and publication date information (see figure 4). although, the filter program was a success, the social science data center cautioned us that incoming materials are being deposited to their appropriate subdirectories regularly. in response to this, a cron daemon which runs the filter program automatically is scheduled every other month3. the virgo template (see figure 5) for icpsr online materials was created to facilitate the cataloging process. faculty, students and staff from the university can retrieve the information as soon as the bibliographic record is created. conclusion after implementing the perl filter program and creating the virgo template (similar to constant data on oclc), cataloging online icpsr series became more manageable than previously perceived. yet, tapeloading from virgo to oclc still presents a challenge for the library. oclc currently does not allow tapeloading on machine-readable original cataloging. since the library is committed to information and resource sharing, it means that contributing these bibliographic records requires additional manual work. each record must be re-created for oclc holdings. the library is looking forward to the day when cataloging records for machine-readable format can be transferred 38 iassist quarterly figure 3 #!/usr/bin/perl $user=”userid@virginia.edu”; $base=”/lv2/users/userid/icpsr”; $tmpfile=”$base/invent”; $pages=”$base/pages”; $lastscan=”$base/lastscan”; open(new, “find /archive/public/icpsr/*/* -name *codebk* -newer $lastscan -type f -print|”); open(tmp,”> $tmpfile”) || die “can’t open tmp file!\n”; print tmp “here’s the new icpsr codebooks:\n\n”; while($file=) { print tmp $file; chop $file; $number = $file; $number =~ s/.*i([0-9]*)\.*cod.*/$1/; $edition = $file; $edition =~ s/.*\.codebk(.*)/$1; $size = -s $file; open(codepg,”> $pages/$number.$edition”) || die “fail to open!\n”; printf codepg “size: %s bytes\n”, $size; chop $file; $number = $file; $number =~ s/.*i([0-9]*)\.*cod.*/$1/; $edition = $file; $edition =~ s/.*\.codebk(.*)/$1; $size = -s $file; open(codepg,”> $pages/$number.$edition”) || die “fail to open!\n”; printf codepg “size: %s bytes\n”, $size; open(codebk,$file); while() { print codepg; last if $. > 80; } close(codebk); close(codepg); } `touch $lastscan`; printf tmp “\nthe first pages are in $pages\n”; close(tmp); close(new); system(“mailx -s \”new icpsr items\” $user < $tmpfile”); system(“rm -f $tmpfile”); 39fall/winter 1994 figure 4 size: 408240 bytes 1 census of population, 1910 ^munited statesy: public use sample (icpsr 9166) principal investigator samuel h. preston university of pennsylvania first icpsr edition spring, 1989 inter-university consortium for political and social research p.o. box 1248 ann arbor, michigan 48106 1 bibliographic citation, acknowledgment of assistance and data disclaimer all manuscripts utilizing data made available through the consortium should acknowledge that fact as well as identify the original collector of the data. in order to get such source acknowledgment listed in social science bibliographic utilities, it is necessary to present them in the form of a footnote or a reference. the bibliographic citation for this data collection is: preston, samuel h. census of population, 1910 ^munited statesy: public use sample ^mcomputer filey. philadelphia, pa.: university of pennsylvania. population studies center, 1989 ^mproducery. ann arbor, mi.: inter-university (end) figure 5 ulalk8984 fmt d rt m bl m dt 01/04/95 r/dt 01/25/95 stat nn e/l dcf a d/s d src d place miu lang eng mod t/aud d/code ? s/stat ? dt/1 ???? dt/2 df/typ d mach freq reg govt 040: : a va@ c va@ 049: : a va@@ 090/1: : a h62 b .i25 no. 100:1 : a 245:10: a 256: : a computer data (1 file : ca. kilobytes). 260: : a ann arbor, mich. : b inter-university consortium for political and social research, c <year>. 490/1:1 : a icpsr ; 500/1: : a codebook to accompany related data tape. 516/2: : a text. 516/3: : a <numeric (summary statistics).> 520/4: : a <optional.> 500/5: : a <also available in paper format.> 580/6: : a issued also in paper format, titled: 537: : a hard copy documentation (year) transformed into machine-readable text utilizing optical character recognition (ocr) scanning, date. 650/1: 0: a <subject> 700/1:10: a <personal author> 710/2:21: a inter-university consortium for political and social research. 710/3:21: a <corporate author> 830/1: 0: a icpsr (series) ; v 856/2:7 : m social science data cender and icpsr services, (804) 982-2630 u gopher:// gopher.lib.virginia.edu:70/11/socsci/icpsr 2 gopher. 4. tapeloading is available for the internet cataloging project participating libraries under certain guidelines, erik jul’s building a catalog of internet-accessible materials: project overview. url:http://www.oclc.org/oclc/man/ catproj/overview.htm. successfully via either internet ftp or tapeloading. the less editing required on one record, the more reliable the information remains. 1. paper presented at iassist 95 quebec city, quebec. jackie shieh is original cataloger for electronic resources at alderman library, university of virginia library, e-mail ejs7y@virginia.edu 2 the itc is the equivalent of computer center in other institutions. 3. in unix, the cron daemon runs shell commands at specified dates and times. regularly scheduled commands can be specified according to instructions contained in the crontab files. the cron daemon examines crontab files and at command files only when the cron daemon is initialized. iassist quarterly 5 educating the data user: the data archivist and bibliographic instruction by bliss b. siman' associate professor, data archivist baruch college, city university of new york information specialists have long been aware that graduates usually leave academia with onh a rudimentary knowledge of information strategies and resources. the expansion of information access opportunities has not changed this situation radically. in the area of information about public data sources, the gap between what students know and what they could know, is tremendous. at baruch, this problem is particularly important because there are few areas more dependent on information access and utilization, particularly in machine-readable formats, than business, and few academic disciplines, therefore, in which this "information gap" has greater significance. since the librar\' at baruch college provides 'presented at the international association for social science information service and technology (iassist) conference held in washington. d.c., may 26-29, 1988 access to information in almost all cunently available formats: print, microform, audiodisc, microcomputer diskette, cd-rom, and machine-readable data files, its instructional program has tried to include all of these information technologies in its workshops and courses. sometimes this has come about quite by accident when the library began to collect machine-readable data files (mrdf), the activit)was assigned to me based on my expressed interest since it was not a full-time assignment, i continued teaching in the bibliographic instruction program. in retrospect it was a fortuitous combination of responsibilities. working with data users, i became aware of the gaps in their information skills, the ver>skills i was teaching in my sections of the "information research in business" cotirse. most data users were capable of using sas or spss.x to analyze their data, and many could write elegant cobol programs, or download data into lotus, but almost none had training in how to search for and identify qualitydata. few graduate students or faculty members were sufficiently aware of the vast potential of public data for their research or teaching. surprisingly, many sophisticated faculty researchers continued to use datasets first introduced to them by their ph.d. mentors simply because they were unaware of alternatives. few were familiar with the varied storage media for data and how these cotild be eflectively combined. for example, a facultv' member working on a project using the census of population and housing on magnetic tape might be totally unaware that portions of the work cotild be better accomplished using the same data in print clearly there were many exceptions to this bleak picttire, but data users were not able to use the wide variety of sources available due to lack of information skills. finding the means to overcome this deficiency became an important objective of a combined summer 1988 6 iassist quarterly data archives/bibliographic instruction program. at the same time that the data library was being developed, other sections of the instructional program were beginning to include numeric data as part of the information process being taughl online computerized retrieval services began to provide access to numeric data, and instruction in print sources of data was increasing in sophistication. naturally, with all these programs being developed in the same division, there was a good deal of beneficial "cross-fertilization" in the planning and implementation of the instructional programs as well as the data service. recognition of the competitive importance of numeric, quantitative information in business reinforced our desire to equip baruch graduates, many of whom are the first generation of their families to go to college, with the data access skills they lacked. with respect to information literacy, business requirements were growing and we wanted to be sure that our students were prepared. our method of attacking this problem also reflects a basic philosophy that information resources in general are underutilized due to the public's lack of training and education in research resources and research skills. consistent with this philosophy, the data resources service immediately organized a seminar series to introduce users to numeric data sources. these seminars, which are still oftered on a regular basis, were organized by subject field and publicized to faculty and graduate students. unlike many similar seminars given in computer centers and data archives, these presentations included not only discussions of important datafiles, but also information on the same data in print, or the major print reference works in the same field. often the data being discussed was also available through online vendors or on microcomputer diskette, and the criteria for using the various media were presented with practice problems to illustrate important points. the objective of the seminars was not to provide a "shopping list" of datafiles, but rather to equip the audience with the skills to find and evaluate the data needed for a particular project most importantly, no seminar failed to include information on important sources, whether data archives, government agencies or private vendors, of new data in the field, both general and specialized resources. models for locating data in a new field were discussed. these strategies paralleled those taught in the "information research in business" course, in which students are equipped with the ability to identify sources in new fields as a basic part of modem information retrieval skills. the reception this information received underscored the need the modem data user has to identify quality data when beginning research in an unfamiliar area. for new graduates in entr>'-level positions, such skills can be invaluable. many of these seminars included an online demonstration of some dataset using, where appropriate, scss (the interactive version of spss). the object of these seminars was to present mrdf in the context of other information sources, as well as to educate users in the available resources. initially, the target audiences were those who were already data users, individuals who had used data and were probably aware of the limitations of their knowledge of public data file availability. because they were knowledgeable about data, they particularly appreciated training in strategies for finding data sources when the usual avenues were unproductive. these users were also receptive to information on selection of format of data, because they were also not aware of the many choices that could be made. widening the scope of data knowledge and information retrieval skills among attendees was the primary objective of the first seminars. this objective remains cenual, although the seminars today are often presented to those who don't know very much about data. of necessity, lists of data files are distributed when the seminar is attended by new data users. although analytic summer 1988 lassist quarterly 7 issues are sometimes discussed, research methodology is never the focus. when baruch became the coordinator of the icpsr membership for the cit>' university*, these seminars were extended to the entire university with equal success. building on this basic format, we have experimented with workshops which include specific demonstrations of ciata sources that appear in online format, perhaps in print, and as machine-readable data files. for example, the trinet database, which contains market share data for companies, is issued online and on magnetic tape. the computer search services librarian and i conducted this seminar to demonstrate the pros and cons of using nimieric data online versus on tape. baruch faculty and graduate students have become informed users of online information services, partly because these services have been free, and partly due to the excellent assistance that has been available to online users. but the consequence has been an expectation among the clients of the online services that all information, bibliographic and nimieric. will be foimd neatly set up in database format, easily retrieved by a packaged quer\language. semmars in which the nuances of using different formats has been presented, have increased the sophistication and information proficiency of faculty and students, especially the graduate students. on the other hand, one can be too successful in encouraging data users to become aware of the wide variety of data sources available. the demand for data increases and the multiplicity of data requested creates budgetary problems. since our graduate students are primarily in business fields, their need for expensive financial and economic data caimoi always be met within the data resources service budget finding ways in which to meet their needs is a constant challenge. recently the number of very specific, very limited financial daiaseis being requested (example: 10 years of currency prices for five specific countries) went far beyond the capacity of the baruch computerized information services. methods of creating datasets, rather than purchasing them from private vendors were explored. using an online vendor, i. p. sharp, we downloaded small amounts of data to meet specific needs. in order to make the data more widely available, they were also uploaded to magnetic tape on the mainframe and listed in the data archive holdings. as another alternative, microcomputer diskette data sets in lotus 1-2-3 format were also created. we are hoping that knowledge of these data will encourage users to plan their graduate theses, etc., around data that are available rather than devising projects dependant on costly new data sets. however, philosophically, it was important not just to create these data sets ourselves, but to educate users to the benefits of this technique, a technique that could be very useful to faculty in their current work and to students in their future business roles. consequently, the entire process was demonstrated at a data seminar. the demonstration included how the data were identified, accessed online, downloaded to diskette, uploaded to magnetic tape and then accessed using sas. at the same time, the use of the data at each intermediary step was discussed, providing greater depth to the users' understanding of the pros and cons of each technique. as the variety of formats increases, we see increased possibilities for this type of seminar. for example, seminars on the census, could present alternative methods of searching for data files, beginning with bibliographic searches online, as well as the choices to be made among formats: print, microfiche, cendata, machine-readable data files, diskettes, cd-rom, etc. census data are very much underutilized at many colleges of the city university, and we believe such instruction will assist users in identifying valuable data not previously considered. summer j988 iassist quarterly most of this discussion has focused on the kind of training given sophisticated data users to improve their information skills. but the library instruction division and the data resources service have been equally engaged in a dialogue concerning the competence that undergraduate students should master in order to attain "information literacy" with respect to public data sources. actually, this dialogue is part of a continuing discussion and re-evaluation of the entire bibliographic instruction program, necessary in a rapidly changing information environmenl there is not yet consensus on what represents an adequate set of information skills. although the library instruction division (lid) offers a basic course in information research in business (library 1016, information sources in business) the core of materials covered has changed over time. in fact, although all instructors teaching the course use the same text and give uniform exams, there is considerable flexibility in planning the curriculum. instructors are encouraged to experiment with materials and share their successes or failures with colleagues. as the importance of public data sources became recognized within the lid, the department began discussing methods of including this information resource in the curriculum. frankly, a complete answer has not been arrived at, although several configurations have been used. cleariy, undergraduates do not need sophisticated knowledge of public data, but they do need some awareness of the role of raw data in the information process, of the availability of data for secondary analysis, and where data can be obtained. the inclusion of data sources in the curriculimi reinforced some of the conceptual goals of the course as well. master}' of the development of search strategies is an important objective of the course, and data files, due to the lack of bibliographic control, provide an excellent example of non-standard information search strategies. the process of searching for data is, of necessity, very different from the index searching model which undergraduates come to assimie is relevant in all situations. awareness of raw data files and of what constitutes an authoritative soiu^ce in this field reinforces another objective of the course, that of enhancing the students' ability to evaluate the information they consume on a daily basis. polls and government reports based on data are constantly presented in newspapers, magazines, etc. questioning the data on which these reports are based is an important attribute of the educated citizen in private or professional life. by including data as an information resource in the research course, we are equipping students the better to handle this important task. working with my colleagues, i developed several different instructional modules which presented numeric data and the different formats in which they come. the lectures contained explanations of machine-readable data files and the nature of secondary analysis. basically, 1 wanted students to understand how data are used in business and academic research and the retrieval methods used to find data files. these lectures were tied to different units of the course, depending on the instructor: marketing in one case, social science research in another. in each case, the lectures built on what the students had already learned about information access and retrieval as well as emphasizing the evaluation of sources. the lectures always included a demonstration of data use in an area that would pique students' interest during these demonstrations, students were encouraged to participate in the development of a hypothesis and the testing of it with data at hand. although this technique was borrowed from modules developed for sociology courses, the emphasis was placed on information dissemination and evaluation issues rather than summer 1988 iassist quarterly — 9 on research methodologies. these demonstrations were very effective in making students aware of the whole process of research transmission, but it is not clear that the best combination of subject year and lecture format has yet been found. despite their success, these lectures have not yet become a standard part of the course for several reasons. as with every academic institution, we are understaffed. it is not always possible for me to do these lectures at the appropriate time in each instructor's syllabus. nor do the instructors necessarily have the time in a single semester course to devote a full lecture to data files. the amount of material we would like to cover in a course is far greater than can be managed in a semester. topics are constantly being juggled as we search for the optimum mix. however, the department has a commitment to including the identification and retrieval of data files in the basic information course because we feel that use of these sources is an important information skill. data users must first be able to find data, and the bibliographic instruction course can equip them with the skills to do the job efficiently and cost-effectively.n summer j988 vol18172 18 iassist quarterly you should “be able to see the library of congress at a glance” declared ben shneiderman, head of the human-computer interaction laboratory at the university of maryland. he was one of many presenters at the advanced information processing & analysis symposium which was held at tyson’s corner, virginia between march 22nd and march 24th, 1994. the symposium was organized by the advanced information processing & analysis steering group which represents the intelligence community. a primary goal of this organization is to provide liaison between intelligence analysts and potential contractors. analysts are responsible for assimilating and synthesizing information from a variety of sources and producing digests of the request in timely fashion. in effect, they face the same overwhelming flood of information that all researchers do. they need some technology which allows them to visualize the search environment at a glance, filter out irrelevant information and focus in on what is critical to their tasks. text visualization is one of the technologies being explored. shneiderman provided an example of how the university of maryland library holdings could be represented by their macintosh-based treeviz program as a series of 100 boxes on a computer screen. each box represented a dewey decimal-based classification, such as history, science or philosophy. the size of the box was proportional to the size of the particular holdings. boxes were color coded to indicate rate of use. boxes were arranged in hierarchies, so that one could descend from a general box, such as history, to a more detailed box-representation of books in various fields of history. not all intuitive user interface designers agree that more is better. the screen design suggested by shneiderman could be overwhelming for some users. other designs and retrieval technologies are examined below. text retrieval and visualization systems since the majority of information requested by users is textual in nature, a great deal of research is being done to find improved ways of clarifying the researcher’s information needs (queries) and comparing them with text content in the database. the effectiveness of any text retrieval system is the degree to which it satisfies “an information need” [croft, 9]. does it retrieve most of the relevant documents from the prospective document pool? to satisfy these criteria, any system must be capable of “representing a user’s information problem or need, representing the content of text documents, and comparing these representations to decide which documents should be retrieved” [croft, 9]. there are two approaches: statistical and knowledge-based. the former is based upon the notion that words occurring less frequently in documents are more important than those frequently found. knowledge-based approaches are more concerned with the human aspects of information retrieval [croft, 10]. these text retrieval approaches are fundamental to text visualization. before a document collection can be visualized, each document must be retrievable. to be retrievable, each must be represented within the retrieval system. conventionally, documents are represented in an inverted index file. each keyword is individually indexed and cross-referenced with the documents containing it. an inverted file is the “set of indexes for all allowable terms and attribute values” [salton, 232]. this file might be imagined as a spreadsheet with column headers consisting of document identifiers and intersecting rows of keywords. at each column-row intersect would be a “o” if the terms does not appear in the document or a “1” if it does. a further refinement is the addition of term locations to the index giving the frequency of keyword occurrence in a document, and the precise location of the term; e.g., document 35o, paragraph 11, sentence 3, word 7. since not all terms are created equal, keywords are weighted by their importance within the document. frequently occurring words such as “the” or “and” receive lesser weights than infrequently occurring keywords such as nouns. keyword weights, therefore, are also added to the index [salton, 231 239]. indexing may be done manually by human experts or automatically by a program such as those described below. when indexing is done automatically, keywords are identified, their weights computed and combined to produce a document vector that may be compared with other document vectors to compute the degree of similarity and provide a “the library of congress at a glance”: text visualization and reference rooms without walls by lee a. gladwin1 center for electronic records, national archives and records administration 19spring/summer 1994 basis for graphical mapping of document relationships [salton, 275 290, 304 308]. following the description of both documents and queries in terms of their vectors, retrieval is executed on the basis of a “computation of query-record similarities” [salton, 275]. tasc’s textviz program follows the knowledge-based approach. multiple documents are first scanned via an optical character recognition (ocr) device into the database from which “features” (symbols, keywords, phrases such as “leaders, “house”, “bogota”) are then extracted using a natural language processor. features are then converted to “feature vectors” or unique identification codes which allow the system to compute “the degree to which a specific concept is correlated with a document” [textviz, 2]. text features or concepts may then be “mapped to points in a graphical space (text map)” where documents dealing with similar concepts are clustered together. dissimilar documents are spaced farther apart on the screen. in tasc’s textviz program, key words appear on the map. in other systems, documents or concepts may be represented as alpha-glyphs, tadpole-like icons with attached lines or “tails” pointing in the direction of similarity. hnc, inc took a statistical approach. their matchplus program employs neural network technology to learn database vocabulary by first discovering “similarity of usage at a word level, in a language-independent manner, without the need for external dictionaries, thesauri or semantic networks” [caid & carleton, 2]. the neural network generates word vectors which represent document content. words learned in a given context point to or cluster with other related terms used in such contexts as weather, finance or government [caid & carleton, 3]. the next step is to represent documents in terms of “the weighted sum of the context vectors associated with words in the document” [caid & carleton, 5]. documents dealing with similar topics are clustered or indexed for easier scanning. document retrieval is accomplished by converting the user’s natural language query into a context vector and searching for the nearest matching document vectors [caid & carleton, 7]. a statistical approach was also adopted by the national security agency’s acquaintance program which employs a “language-independent n-gram method of sorting and retrieving documents by language and topics. n-grams refers to “sequences of n consecutive characters” [damashek, 39]. parentage, a visualization program, is used to explore the retrieved documents (see below). although not exhibited at the conference, it should be noted that ge research and development center’s nldb textbased information system is perhaps the first to employ a hybrid approach to text retrieval. nldb imitates human indexers and “automatically assigns categories to news stories for dissemination, retrieval, and browsing” [jacobs]. based on recall and precision criteria (retrieval of a high proportion of all relevant documents), their tests show that a combined knowledge-based and statistical approach to term categorization is superior to using either method alone. intuitive user interfaces (iui) regardless of the technology used to process, index and extract information from textual sources, it is the interface with which the user must interact. james a. wise, battelle pacific northwest laboratory, stated that an iui is only “‘intuitive’ to the degree that it exceeds the bounds of syntax”. he emphasized that it must capture and communicate “the essence of meanings in messages through both the verbal and nonverbal domains”, utilizing “analogs of the kinds of things that inform intuitions in everyday life”. this interface must allow the user to instantly grasp the breadth of the information environment, easily filter out irrelevant material and focus search upon the most promising areas. several approaches to displaying the information domain were described above: proportional-colored boxes, concept maps and alpha-glyphs. starfields, appearing like explosions of multicolored confetti, may be used in conjunction with a 20 iassist quarterly legend at the bottom of the screen to indicate what the colors represent. in a library setting, for example, “blue” could indicate a mystery novel or film. viewers of the national security agency’s parentage program first see a screen covered with various concentrations of dots. this is the overview. to filter out some of these points, the user may click on a label in a menu beside a given dot cluster. this results in a cross-sectional view resembling a concept map in which related document nodes are linked by lines of varied thickness depending upon the strength of the relation. document clusters may be searched by using a query template containing one or more labels and setting that label equal to a specific value, such as profession s engineer (see figure 1). this results in a list of documents or a display of objects from which to make final selections [cohen, 115]. logicon’s browser is designed for the user who says, “i don’t know what the evidence is likely to be. i’ll know it when i see it.” a document folder metaphor is used to aid researchers in grasping quickly how the system works. the user begins by looking at a list of folder names. opening a folder will display a list of subject lines. clicking on a subject line provides information about a specific document. the user may select a document and then click on a word highlighted in a text in order to see “a list of folders that include documents containing the term, bring up a list of documents in any folder on that list, and open any document”. all of these approaches sounded terribly futuristic to attendees until they stepped into the exhibit hall and saw actual demonstrations of the systems discussed in the presentations. a paradigmatic shift in how we view information for those using these new interface designs, a major shift in how we think about information is required. in the gutenburg galaxy, information was organized linearly. text and pictures followed in an orderly succession and was searched from beginning to end. in the electronic age, paper, pictures, movies, and sound recordings are collections of objects adrift in cyberspace which can be manipulated by individuals. information is reduced to related data chunks which may be viewed contextually through text visualization [hilbing, 54]. order is initially imposed by various algorithms, saving the user great time and effort. in a way, the concept of information as inter-related chunks of data is analogous to the organization and storage of information in the human brain’s neural network hierarchies. learning is the formation and modification of neural synapses in the process of forming more complex structures. hnc, inc’s matchplus forms text associations in this manner. visualization techniques transform associated data chunks into a cohesive visual representation that can be understood by the user. a paradigmatic shift in reference support services text visualization and associated intelligent systems will transform not only how the researcher locates information, but traditional reference support services as well. since users will be able to search library holdings and documents on their own, only the most difficult reference work will need to be performed by staff [hayes, 4]. through text visualization and related technologies, researchers will be able to view vast amounts of data at a glance, focus their search and explore areas relevant to their interests. this will allow them greater time to interpret, analyze and report information than they have ever had. reference staff will be free to assist researchers with more challenging reference problems. the intelligent technologies under development for the intelligence community today will be in our reference rooms tomorrow. while there are still problems to be solved before these systems become generally available, the day of the electronic reference room without walls is closer than we may wish to think. 21spring/summer 1994 22 iassist quarterly references. caid, william r. and joel l. carleton. “context vector-based text retrieval” (hnc, inc. nd). working paper supplement to paper presented at symposium on advanced information processing & analysis, march 24 26, 1992. carlotto, mark j. “text visualization” (tasc, 1992). paper presented at symposium on advanced information processing & analysis, march 24 26, 1992. cohen, jonathan, “parentage”. paper presented at symposium on advanced information processing and analysis, 22-24 march, 1994. combs, nathan h. “large text database visualization” (tasc, 1992). paper presented at symposium on advanced information processing & analysis, march 24 26, 1992. croft, w. bruce. “knowledge-based and statistical approaches to text retrieval”, ieee expert (april, 1993). damashek, marc. “acquaintance”. paper presented at symposium on advanced information processing and analysis, 22-24 march, 1994~hayes, phil. “knowledge-based systems for the information industry”, ieee expert (april, 1993). hilbing, capt. john f. “electronic production in the post paradigm shift world”. paper presented at symposium on advanced information processing and analysis, 22-24 march, 1994. jacobs, paul s. “using statistical methods to improve knowledge-based news categorization”, ieee expert (april, 1993). meadow, charles t. text information retrieval systems (ny: academic press, 1992). proceedings. symposium on advanced information processing and analysis, 22-24 march, 1994. sponsored by advanced information processing and analysis steering group, intelligence community. salton, gerard. automatic text processing: the transformation analysis, and retrieval of information by computer (reading, ma: addison-wesley, 1989). sasseen, robert v. and william r. caid. “docuverse: a context vector-based approach to graphical representation of information content” (hnc, inc. nd) textviz, “visualization of large text document databases: suggested statement of work for text visualization (textviz) development and demonstration (27 february 1992). paper accompanying presentation at aipasg symposium, 1994. note: descriptions of exemplars of these systems should not be construed as endorsement of any particular system or retrieval-visualization methods by either nara or the author. all opinions expressed are those of the author and do not reflect nara policy or programs. 23spring/summer 1994 the isbd(cf) review group meet at the library of congress april 24-26, 1995 to consider a revised version of the text of the _international standard bibliographic description for computer files_ (1990) prepared at the chairman’s request by ann sandberg-fox who is serving as principal editor of the second edition. in attendance at this meeting were group members sten hedberg (uppsala universitetsbibliotek); catherine marandas (bibliotheque nationale de france); ms. sandberg-fox (colchester, vermont); chairman john byrum (library of congress) as well as corresponding members laurel jizba (michigan state university libraries) and lucy evans (british library) as well as observer claire vayssade (bibliotheque nationale de france). the meeting was made possible by a subsidy from ifla and a grant from the research libraries group (rlg). the first day was devoted to discussion of several issues-papers which ms. sandberg-fox had prepared. these covered the topics most in need of reconsideration in the light of the rapidly developing technology which has influenced the creation and dissemination of computer files: interactive multimedia; the general material designation (gmd); sources of information; reproduction and multiple versions; designation of file; and, published versus unpublished remote texts. in addition other aspects, such as preliminaries, type and extent of file, physical description and notes were thoroughly discussed, as were a number of proposals received by the chair prior and subsequent to the formation of the review group. on the second and third days, the members focused on a close reading of the revision prepared by ms. sandberg-fox, with the result that an agreed upon text emerged from the meeting. the draft will now be updated to incorporate decisions taken at this gathering and, with permission of the sections on cataloguing and on information technology, presented for world-wide review on or about september 1, 1995. following a six-month comment period, a final version of isbd(cf) second edition will be readied for ifla approval and publication; in addition, the text will be shared with the authors of national and international cataloguing codes, such as the joint steering committee for aacr. following is a brief summary of the most important outcomes of the april 2426 meeting and which will be reflected in the revised isbd(cf), presented in terms of the objectives that were set out to guide this project: (1) to take into account the emergence of interactive multimedia, a new and still developing technology that combines and stores products of audio and video technologies, together with text and graphics, on optical discs. regarding interactive multimedia, the review group concluded that all such resources be incorporated into the new version of cf. this conclusion was reached because no existing isbd covers these materials (which entered the mass market beginning in the mid-1980’s), and because user-manipulated, non-linear navigation using computer-controlled technology are hallmarks which characterize interactive multimedia. (these materials are distinct from multimedia/kits that are covered by the stipulations of isbd(nbm).) as a result, the new version of cf will add or amend provisions regarding sources of information (0.5), edition (area 2), type and extent of file (area 3), dates (area 4), physical description (area 5) and the notes (area 7) to show treatment of interactive multimedia as a subset of computer files. examples will be added to illustrate such files. (2) to consider the impact of developments in optical technology, as new and improved optical discs are replacing magnetic disks as primary storage devices. the review group decided to improve cf to cover not only cd-roms (compact disc read-only memory) but also cd-i’s (compact disc interactive), and other emergent forms such as photo-optical compact disc. as a result, the new version of cf will add or amend provision regarding sources of information (0.5), edition (area 2), physical description isbd(cf) review group meeting of april 24-26, 1995 ~ summary report ~ john d. byrum 24 iassist quarterly (area 5), and notes (area 7). the term “disk” (spelled with “k”), currently used throughout area 5 to describe both optical and magnetic devices, will now apply only to magnetic devices, while “disc” (spelled with “c”) will be used in relation to optical manifestations. (3) to provide for the availability of remote electronic files on the internet, a global network of networks that allows users access to a vast wealth of remote electronic files, including books, journals, articles, reference sources, and even library catalogs. since, at the time cf was first formulated, this was a new area especially designed to treat these files, caution was exercised as to the kind and amount of detail to be given. designations of the type of file are limited to general terms only—”data” and “program” and their combination “data and program.” the review group decided that these terms are not adequate for the purposes of identifying the many different types of data files and software on the internet. indeed, the whole treatment of the designation of file was thoroughly reworked and developed, with area 3 emerging as the one most thoroughly changed in revised cf. consequently, the second edition of cf will propose several levels of specificity as appropriate. the current terms “data” and “program” will continue to be authorized, but data files can alternatively be indicated as “numeric”, “text”, pictorial”, “representational” or “sound”, while programs can be identified as “utility”, “application” or “system”. most of these categories are further delineated for more specific designation when appropriate; for example, a bibliographic database may be so identified, as may be a game. as before, the combination “data and program(s)”will continue to be used when applicable. however, alternative identification as to particular types of data and program(s) may be taken from the authorized listing and be used in conjunction with the following terms: “interactive multimedia” or “online service.” these latter terms also function as designations when terms from the authorized listing are not appropriate. where, in the case of combinations, the program or the data may be incidental to the whole, the primary term only is to be given. as for the general material designation (gmd), the group decided to retain “computer file” in the absence of a better alternative. further addressing internet resources, the revised cf will provide better treatment of the networking environment where an electronic file may be accessed by several methods, reside in many directories, and require more detailed information, enabling users to locate and retrieve these files. specifically, cf will be updated to include provision for url’s, gopher and ftp sites. c (4) to deal with bibliographic problems arising from reproductions of computer files such that many cf titles are now available in a variety of physical formats. although such problems are not easily resolved, the cf review group did authorize changes to areas 2 and 5 to better distinguish between an “original” and other versions thereof. reformatting changes were moved from inclusion in the definition of edition to inclusion, instead, into the definition of what would not constitute a new edition. also, output medium and display format are newly reworked phrases to better reflect cf technology. in addition, the review group agreed to significant modifications of the provisions concerning sources of information (0.5). area 4 (“publication”) will be amended to require treatment of all remote cf as published materials. in addition, the glossary and examples will be updated and increased. in the course of its meeting, as requested, the group considered the official draft proposal of the ifla division of bibliographic control study group on the functional requirements for bibliographic records, with barbara tillett, one of the three consultants to that project, present for part of the discussion. it was decided that as a medium, computer files would provide a good test of the draft, and the group agreed to undertake an in-depth study. specifically, 1) the use of the words “item” and “work” in the _functional requirements_ document will be examined in relationship to related terminology in the isbd(cf); 2) an experiment will be conducted to apply the suggested model using several types of computer files in several library environments; 3) the results of the experiment will be analyzed; and 4) a summary document, including any potential recommendations for the isbd(cf) will be written. laurel jizba will coordinate this study for presentation by november 1, 1995. national archives and electronic records: where are we going? by sue gavrel^ electronic records: the challenge to archives introduction the information society has had a major impact on the activities of traditional archives. the prediction of the "paperless office" in the eighties has not materialized as yet and in fact the amount of paper has increased substantially over the past decade. rather than reduce paper, the introduction of computer technology has increased the number of products and copies of those products. the work of the archivist in the identification of the archivally valuable records has increased due to the paper burden. one must sift through far more records to identify those of historical value. added to the increase in the number of records created is the pressure of the research community to retain more rather than fewer records. prior to the mid-seventies, the major factor in the appraisal of records was that of evidential value the evidence the records contain of the organization and functions of agencies. archives (and i restrict my comments mostly to north american and particularly canadian archives) have acquired many more records based on their informational and research value over the past fifteen years. in canada, this coincides with the growth of social and economic programs of the federal government. the computerization of many government programs began in the early sixties and has increased ever since. the centralization of edp expertise was very evident during this time and continued until the arrival of the micro computer. the large database systems were, in most cases, built by the edp experts and used to service the program managers' needs. archival programsfor machine readable records the use and importance of computers was recognized by the large national archival repositories in many countries in the establishment of machine readable record programs. over the years standards were developed for the appraisal, acquisition, processing, conservation and servicing of machine readable records. due to the small number of archivists involved in these programs, a great deal of co-operation and sharing of information lead to the develqjment of procedures to handle these new records. only a minimum amount of success was achieved in the identification and subsequent transfer of computer records of archival value from government agencies. the control of these records was in the hands of the edp area and outside the normal channels of the control of paper records (records managers). efforts to mimic the systems in place for the control of paper records met with limited success, mostly due to the lack of familiarity with the archives by those in charge of the development of the systems. machine readable programs, in traditional archival settings, although recognized as important, lacked the focus and strength required to affect the overall organization of records. trends in recent years, this trend is changing due to a variety of reasons. in the discussion which follows, i will outline some of the changes and trends which will have a profound impact on archival repositories. these changes range from new understandings of the importance of information; technological change; and major changes in the ways data are created, stored and used. the next decade will require archives to focus on electronic records or risk losing the electronic cultural heritage. technological change it is not my intention to provide an overview or history of the changes we have experienced in technology in the past decades. it is, however, important to review some of these changes in light of how records are created, why, and how archives must adapt to these changes. the major trend which has affected the way records are created results from the rapid penetration of microcomputers into the market in the last five years and, in particular, into government departments and agencies. the centralization of edp services is disappearing with the use of microcomputers in the office environment. managers, officers and clerks have now as much computing power on their desks as the mainframes of the seventies provided. the ability to create, manipulate, access and disseminate data has been decentralized. the trend to purchase "off the shelf' software has taken away from the centralized edp shops in the creation of in-house software and database systems. linked to the penetration of microcomputers is the development of local area networks. the linking of staff provides for the creation 32 lassist quarleriy and revision of documents which can be done on-line with only the final version being available in either hard copy or electronic format the ability to provide for the development of documents relating the evolution of the policy, to changes in administration, or to data is now in the hands of the creators of those documents. in most cases, lan's do not provide for the traditional records management approach to control of records. individual workers make decisions regarding the disposition of the recotds. such a system existed for the development of large database systems under the control of edp professionals. the difficulties experienced in gaining control over what information is being created and destroyed can be magnified as all employees become responsible for their records. the program of economic restraint experienced in most western countries is also leading to the use of microcomputers, as it is seen as a way of increasing productivity and decreasing personnel costs. a major effort can be seen in the development and interest in communication standards. the lack of compatibility between hardware has been, and continues to be, a major problem to the increased usage of data. the efforts now seen in the development of international standards is encouraging. the trend to the open systems interconnection standard protocols provide for the possibility of connecting systems with different hardware. other standards will have an impact of the increased sharing of data. map and chart data interchange format, or macdif, data is an attempt to provide a standard format for the transfer of chart and graph data from and to a variety of systems. office document architecture/office document interchange format (oda/odif) provides for similar transferability of text. the work towards developing and implementing such standards must be followed closely by archivists, as it is through such efforts that some of the technical issues such as making valuable data accessible in the future may be resolved. new techniques for software development such as fourth generation languages and expert systems techniques are becoming important tools providing faster and more flexible software development and more user interfaces. no longer must the design and development of databases be the sole responsibility of the edp professional. databases can be created and used by those who have access to d-base or other such software. expert systems potentially pose major problems for the archivist in the past the acquisition of data tried to steer away from software dependent systems. expert systems which can be defined as " an intelligent computer program that uses knowledge and inference procedures to solve problems that are difficult enough to require significant human expertise for their solution" are only in the early stages of practical application. their potential to assist managers with complex planning and scheduling tasks, diagnose diseases, etc. is great the impact of such systems on the documentation of the decision making process is evident how archivists will respond to such systems is a major challenge. types ofdata part of the technological changes, but one which should be highlighted, is the trend towards integrated systems and appucations. two specific types will be discussed: the geographic information system (gis) geograjaic information systems are beginning to play a major role in the information society. today's systems only superficially resemble the automated mapping systems of the sixties. gis are increasingly being used to conserve and manage a wide variety of data from natural resources to environmental pollution as well as in the planning and management of cities such land information systems cross organizational and sectorial boundaries and represent an opportunity to develop new information based products and services. compound documents the move to more integrated systems is seen in the "compound document". the integration of voice, data documents and graphics oversteps the traditional media boundaries. all are reduced to the common language of binary code. not only is it feasible to create the compound document, but it may also have been created from information which was only accessible on the screen for a brief period. the source of that information cannot be traced. the accessibility of data from other systems through local area networks and the merging of data from a variety of data bases will create documentation problems for the archivist. the ease with which such information becomes available and usable will be reflected in the move to adopt communication standards, more user oriented software, and more computing power. ir^ormation as a resource the information society has created a new awareness of information as resource. the major expenditures on hardware and software development of the seventies has created a new awareness of the value of the information which these systems manipulate and store. in canada, access to information and privacy legislation led to the acceptance of a computer based record as a record. in the definition of a record for the purpose of the legislation, machine readable is included. the requirement to account for information regardless of the machine on which it was stored was an important step in the recognition of computer records. the new national archives act passed in 1987 also uses the same definition of record. the act stipulates that no records of spring 1991 the government of canada can be destroyed without the consent of the national archivist it further stipulates that those records deemed to have archival value must be transferred to the national archives. the definition of record to include machine readable records ensures that electronic records are part of the national archives' responsibility. more recently two new policies have added to the importance given to information: the information holdups policy (just recenuy approved) which will direct government dejjartments to manage their information in a holistic manner and account for that information through the development of directories to it; and the information management plan which ouuines to departments the importance of the planning of information management rather than the justification of new equipment purchases. all of these policies, acts, and planning strategies are the result of an awakening to the importance of information as a resource. another interesting trend in the field of information is the move by the public and private sectors to develop cooperative databases. the geographic information systems are an example of this type of co-operative effort. similar to the compound document such cooperative databases will have an impact on the archival organization of records. much more could be said on the trends and changes which technology will initiate. it is important to look briefiy at the impact such changes will have on the traditional functions of archives. the challenge to archives "the digitization of information through the common language of the binary code is bringing about the convergence of voice, usage and data and of the telecommunications, electronics and computing industries based upon them". over the last decade, archives have tended to expand accch-ding to media-based responsibilities; textual records, cartographic, film and television, photographic and machine readable. practices and procedures were developed to acquire, process, store and service the different forms of information as each had its special requirements. the fundamentals of archival theory were common to all media. appraisal criteria, were based on the principles of evidential, infcxmational and legal value. it was in practices for arrangement and description where the differences became more evident. procedures and practices for the long term preservation of machine readable records were developed in the seventies. the procedures were based on large mainframe systems and proved successful ioc the conservation of data in systems. technology is now the driving force behind the integration of the different types of records. just as media divided archives into specific units, it now will play a large role in integrating these units. with the use of electronic technology in the creation of all types of records, the media on which the information resides becomes the common element. information is created and transmitted in so many forms that archivists must now look at program activities as a whole and identify those records which have archival value as well as the most appropriate form in which they should be stored. electronic records provide many research possibilities. as more types of records are created in electronic form more such records are likely to be of archival value. the major obstacle is, of course, the long term accessibility of the information in electronic form. for other records-paper, photos, maps estabhshed techniques have been developed which preserve the records for future use. technology did not have a major effect on long term preservation except for improving the techniques. with electronic records, the media, the software and hardware are constantly changing evidence of this can be seen in the experience gained to date. electronic records created in the seventies are different from what is now being created. procedures, valid for data in systems, must be modified and reevaluated to cope with compound documents and gis. the technological requirements put pressure on established procedures. the focus of any archival program for electronic records must focus its resources on resolving the technical issues of how best to transfer records to the archives, how to process these records to ensure their accessibility. archivists will be required to support efforts to ensure standards; to become involved with systems as they are being created; and to keep abreast and knowledgeable about the changing technology and how it affects the creation and use of records. these are major changes for institutions which have traditionally dealt with the past. finally, new methods in records creation may have fundamental effects on traditional archival theory and principles. archivists must become active participants in the creation of information, in many instances identifying elements of archival value before they are created, in order to ensure the preservation of the historical record. archives have, to date, been concerned with documenting the activities of an organization or business. how does this new role affect the documentation of activities when the archivist has participated in the creation stage? as more organizations undertake co-operative efforts in information creation, how do we determine which records originate with which organization? who has the ultimate responsibility or control of the records? such lassist quarterly systems break down the barriers between public and private sectors; federal, provincial and municipal governments. the clear lines of origin become blurred. conclusion today's presentation can only briefly mention the issues and resulting challenges to traditional archives. issues such as these are of utmost importance to the archival community. the "paperless office" is not yet a reality but signs of its existence are much more evident today than they were two to five years ago. effcmts to resolve the problems are imperative if records documenting the nineties are to be available for future generations. presented at the ifdo/iassist 89 conference held in jerusalem, israel, may 15-18, 1989 forester, tom. high-tech society. mit press, 1988. p.l spring 1991 64 iassist quarterly 2010 / 2011 iassist quarterly abstract in hungary, like in many of the former socialist countries, regardless of ongoing research projects based on qualitative data, archives of such research data are sporadic and fragmented. the strong tradition of sociological research produced a vast amount of qualitative social data but no infrastructure was developed to capture, organize and make available these data sets for further analysis. in most cases the secondary analysis or the longterm preservation of these valuable collections is hindered by institutional and disciplinary boundaries as well as a culture of unwillingness to share data in the humanities and social sciences. in hungary future progress can only be expected if legal barriers are removed or addressed at the policy level; if the underlying infrastructure offers an attractive set of tools for future data collectors; if potential researcher’s become aware of the improved access to data; and finally, if preservation best practices, standards, and benchmarking are spread among custodians of the content. it remains an open question which approach is more appropriate: creating a central data repository or setting up a decentralized network of institutions and individuals which can lead to inter-operable platforms to share content and know-how. keywords: fragmented archives, legal barriers, data preservation, lack of benchmarking, hungary background in the past 50 years in hungary, tens of thousands of interviews have been conducted across widely divergent topics of sociology. this giant scientific source is scattered, idle and is gradually perishing. even though publications and articles referring to the original data sets do not cease to come out, the raw data is unlikely to be found with ease. to our knowledge, surveys, interviews, and transcriptions collected by researchers were often merged into personal collections of known scholars. sometimes these files were simply discarded due to lack of space, preservation problems, or, in more fortunate situations, they were donated to libraries, archives or museums but without being described and made available for research. inadequacy of cataloguing and preservation by all three types of institutions is due to their lack of expertise to describe these materials in depth, e.g., archival descriptions do not go beyond the functional grouping of files. for example, before the transition period the research projects related to roma minority issues in hungary were conducted under various umbrella projects for fear of exposing serious social issues. in the communist era terms such as roma problems, unemployment, and social integration were not part of the official research discourse, which was controlled and censored by the party. moreover, legitimate research results had to be hidden in the drawers of dusty filing cabinets, and published reports had to avoid sensitive issues as well. in the past two years, the pilot project called “voices of the 20th century archive and research center” (for further information on voices see, www.voicesofthe20century. hu), financed by the hungarian state research fund, has been creating in an inventory of existing resources of qualitative social scientific data, especially interview material gathered in communist and post-communist periods. this comprehensive inventory (now freshly available on the website) is the first step towards setting up an open and public online archive. our aim is to make the textual and audio-visual heritage of the history of the 20th century accessible to broader audience, both scholars and the qualitative longitudinal research and qualitative resources the hungarian case by judit gárdos and gabriella ivacs1 iassist quarterly 2010 / 2011 65 iassist quarterly public. broad public access would enable informed citizens to learn about the 20th century and experience it in a sympathetic way. on the other hand, advanced users of the material – researchers, teachers, media workers and students – could analyse the context in which data was collected, how the results were built into scientific knowledge, and how they could be re-assessed to generate new scientific knowledge. initially, qualitative audio and audiovisual social scientific collections are being uploaded to the webpage. it is also to be noted that segregated resources based on disciplinary silos can offer less value to researchers who are interested in broader subjects, periods, and phenomena, rather than a particular social data set from a given period. it is highly recommended to evaluate the possibility of integrating qualitative data archives in hungary into the existing archival infrastructure, if an adequate one already exists. archival management, preservation methods, and technology issues can be easily adapted to the special content; in exchange, integrated resources are more attractive to researchers. for archives and information professionals it is always a challenge to work with legacy data; contextual information attached to collections gathered for various purposes in various ways is as valuable as the records themselves. preserving provenance information will remain a real problem in the case of scattered qualitative data sets in hungary as well. the chain of custody in many situations is not clear: nobody knows how certain sets of files migrate to particular members of research teams, who holds the most comprehensive collection, how many copies exist with whom, and so on. at the same time ownership and intellectual property rights issues also ought to be addressed: research projects financed by state funds are often considered private initiatives, and regulations imposed by the funding organs to share collected data with ‘secondary’ users are rarely enforced. secrecy, restrictiveness, fragmentation, isolation, and lack of awareness of cooperation, sharing and preservation continues to characterise the hungarian research culture, and we can state that in the last 50 years no real progress has been made in this respect. meanwhile quantitative social sciences and quantitative data archiving are in a somewhat more favourable situation. there is no the doubt that the internationally known institution, tárki social research institute (the only hungarian member of cessda) has had a crucial role in introducing the practice of re-using quantitative data from empirical social research, but the use of their archive is not widespread either. the hungarian data protection law seems to be another obstacle to interdisciplinary data sharing: the current version of the legislation does not provide straightforward guidelines on how to make digitized resources containing private data available online, how to define private data in the technology enhanced environment, or how to anonymize digital content in a less difficult way. the law is deeply rooted in the outdated structure of traditional paper archives. a model of distributed archiving, sharing a common infrastructure and with the participation of several players from organisations to individuals is not approved by the legislation. rum harupic imusdae perchil iquaeped et landaepro delest molorat aquiatiam remquas ventiberiti vollupt atiassimpere voloruntus peroratem invelecta cor samendi a que quatibe rsperrovit the voices project the project is overseen by the institute of sociology of the hungarian academy of science, in cooperation with the open society archives at central european university and the hungarian national audiovisual archive. the voices team has developed a two year project plan ending in spring 2011 to examine through a pilot how different types of interview material can be re-used for further research, with their various formats enhanced by technology solutions such as digitization, ocr, voice recognition techniques and so on. having completed the pilot, a sample collection containing mixed media are being published online along with the contextual information about the methodology employed during the data collection process. prospective students, researchers of science studies, cultural studies or social studies will be involved in the early project stages to explore the content and incorporate the data into their early works: theses, publications, and other projects. other expected outcomes of the pilot project are a mechanism to regulate online access according to privacy levels, added features like tagging, annotating, searching and browsing on the web site, and published resources on the methodology of research. the future archive has three pillars: preservation, research and access. as a result, the voices team has direct and continuous contact with new qualitative research projects and collections as well as undergraduate, graduate and post-graduate university programs. the planned website gives an opportunity to co-operate with similar domestic and international initiatives, e.g., other sub-archives, user networks, links, and partner institutions (voices is now part of the equalan network of european qualitative archives). our research team, being well-integrated in international research and archival networks, will be able to inform the wider international scientific public about the work of the voices of the 20th century archive. a plan about data protection and confidentiality for our new archive has been completed (in hungarian), in accordance with the hungarian ombudsman for the protection of personal data. such a plan proved to be crucial in gaining state funds, since there has been considerable concern about data protection in the reviewers’ reports. data protection is a sensitive issue in many post-socialist countries, hungary among them. in particular, the role of key personalities in the former socialist secret service has been a huge topic in the hungarian public sphere. some historians–using state archive material−reveal from time to time the real names of secret service agents, some of them prominent individuals in contemporary politics or arts. this is made possible by a law text (§32 of the data protection law) stating that if it is necessary for the description of the scientific results concerning events of a historical period, personal data can be published without the consent of the person. in some cases, there has been considerable public discussion whether the identification of a specific person was really needed, and there have been lawsuits against historians for publishing personal data. lately, there have been lively discussions among quantitative social scientists about to re-use and merge statistical databases generated in the administrative branch (government organizations, municipalities, censuses, etc.). the question of anonymity is crucial here, since two anonymous databases, if combined, can possibly lead to a new one in which identities can be traced. there is the desire for countrywide guidelines concerning the merging of mainly state-owned databases, and short proposals for a new law have been issued in may 2010 by the fényes elek research center for statistics and econometrics of the hungarian sociological association. voices in the context of archiving in hungary this is the context in which the voices archive has to be established. we have been able to investigate the attitudes of some other archiving 66 iassist quarterly 2010 / 2011 iassist quarterly institutions and we are becoming aware of the fears of individual researchers in hungary, which are yet to be fully understood; we have just begun our work with them. the institutions are very protective about their material. there seem to be several reasons for this attitude. for the big state archives such as the state radio, easy (online) access even to the catalogue is not an important principle. they seem to be content with their huge holdings and are not very eager to cooperate substantially either with other institutions or with individual users. the audiovisual archives of the national széchényi library cite reasons to explain their data protection policies: there is no existing public catalogue of the interviewees−although such a restriction is not required by the current laws. the researcher can find out whether a certain person has been interviewed, but he cannot ask for a complete list of all interviewed persons. we have only some numbers available: the collection of historical interviews in this archive has collected and produced approximately1500 interviews with the principal figures of hungarian history, public life and culture since 1985. in the case of research institutions, we encountered a somewhat similar attitude, maybe typical of countries with weak qualitative archiving and sharing traditions. research institutions seem to guard their findings and are not particularly supportive of “alien” researchers. there are some peculiar means used to restrict information about their material: in the case of a research institute with extensive oral history interview materials, the catalogue describing which interviews are also available in audio format is not made public. but if a researcher asks for a particular interview, he or she will be able to get this information. there seems no rational reason for such conduct. these different attitudes of researchers and institutions would make an interesting research topic and would reveal much about the post-socialist research culture. archival holdings in hungary most archives collecting scientific interview material focus on topics related to the humanities, especially world war ii and the socialist era. the first topic is represented in hungary by centropa, an interactive database of jewish memory. the project combines old family pictures (a database of 25,000 digitized images) with their accompanying stories (abridged and summarized interviews presented as texts on the homepage and parts read by an actor). centropa has interviewed more than 1,350 elderly jews living in central and eastern europe, the former soviet union, and the sephardic communities of greece, turkey and the balkans. it has a separate hungarian sample and institution. the second topic is covered by the oral history archive (oha) of the institute for the history of the 1956 hungarian revolution, a collection of oral history interviews focusing on the revolution and the era of jános kádár. the open society archive (osa) also holds data on the socialist era and archives research resources for the study of communism and the cold war, particularly in central and eastern europe, as well as issues of human rights. the osa has launched a digital archive, with the aim of broadening access to primary sources by overcoming technical, legal, geographic, and socio-cultural barriers. they developed a strategy which includes large-scale digitization, multilingual description, and the implementation of open-source solutions and open standards. they also seek to meet current international benchmarks by becoming a trusted digital repository. to our knowledge, there are no qualitative longitudinal data collections in hungary. all the above mentioned collections contain material addressing life courses and oral histories. there are some follow-up studies, but only on the level of the individual researchers. one of the aims of our new archive voices will be to investigate who has done qualitative longitudinal projects, as they are not sufficiently documented even at the level of scientific publications. the only hungarian institution that is a member of cessda is tárki. the tárki data archive is the national data repository for empirical research data in hungary. it collects’, stores and publishes a great number of surveys for the social science and business communities nationally and internationally, focusing on quantitative data. the aim of voices is to create an internationally acclaimed research database for qualitative social scientific resources in hungary as well, since –unfortunately− there are currently many gaps in qualitative archiving in hungary. the main priorities of voices for the next three years are to map existing qualitative resources and to develop a data workflow of creation, capture, deposit, archiving and long-term curation based on international standards. we are engaged in exploring cost-effective technological solutions for managing digital data, in building a network of partners (such as potential data donors, experts and users) and in shaping the hungarian research culture of not sharing data. there are some major problems to be solved during our work, such as gaps in hungarian legislation and outdated legislation, ipr issues, the handling of personal data, the lack of a community approach in the sharing of scientific data, and of course the shortage of long-term funding and the need for sustainability. existing organisations (iassist, cessda) might be of assistance in many ways, among them with their know-how on inventorying, collecting data, and depositing content, their expertise in data archiving systems, technology solutions and in the methodology for developing the policy framework for managing and redistributing data sets. to secure their future, qualitative archives all over europe depend on constant and reliable financial help. being embedded in a community of archives brings new opportunities for mutual assistance in issues of both science and funding. further reading for a bibliography about oral history literature in hungarian language see: www.replika.hu/system/files/archivum/58-07.pdf one of the rare analyses of hungarian oral history: lénárt, a. (2007). történetgyűjtés. az oral history tudományos műhelyei magyarországon 1945 után. [collection of history. oral history labs in hungary after 1945] aetas 2. pp5–30. székely, i. (2009.) ‘positive discrimination and data protection: a typology of solutions and the use of modern information technologies’. in: szabó, m. (ed). privacy protection and minority rights. eötvös károly public institute, budapest. 27-62. note 1. judit gárdos hungarian academy of sciences, institute of sociology, “voices of the 20th century archive and research center”, the work of judit gardos has been partly enabled by fund no. 77566 of the hungarian scientific research fund (otka). gardos.judit@socio.mta.hu gabriella ivacs open society archives, central european university, ivacsg@ceu.hu ocul 6 iassist quarterly spring 2012 iassist quarterlyiassist quarterly abstract the need to support and promote the use of geospatial data and available collections has grown at ontario universities in recent years. these data, along with numeric data collections, are used to enhance research and expand the skill set of students graduating from a broad array of disciplines. providing access to these types of data collections has proven challenging, and access points available to large groups of people inside academia, in government and public domains have been made possible through online portal implementations. the defined need for a geospatial portal at ontario universities is outlined in the first part of this paper, followed by components of the geospatial portal project vision and specific requirements and technical aspects of the project. how metadata is handled is also discussed, as is the implementation of the portal’s web application. additional project components include health data collections that are available to researchers and which could be used in conjunction with the portal. finally, project governance and future growth beyond the ontario user base are also discussed.2 keywords: geospatial data, data access platforms, geospatial metadata, software development, collaborative projects introduction this paper describes a presentation given by elizabeth hill, leanne hindmarch (now trimble), and jennifer marvin at the june 2011 iassist conference (vancouver), entitled ocul’s geospatial portal project: from vision to reality. this paper will elaborate on the progress of the project, which is an initiative of the ontario council of university libraries (ocul). ocul is a consortium of all twentyone university libraries in the province of ontario, and is involved in collective purchasing, storage, and delivery of library resources and services. the technical infrastructure to accomplish this is provided by ocul’s scholars portal program. since 2002, scholars portal has hosted shared collections and has been involved in building, maintaining, and/or supporting a range of platforms for collection delivery to end users.3 ocul member institutions all contribute funds toward the operation of scholars portal, including its staff of librarians, developers, and systems specialists who develop and support its services. established scholars portal platforms include electronic journals, electronic books, and a data delivery system named odesi (ontario data documentation extraction service and infrastructure; http://odesi.ca). the international data which feeds into this type of social science research is produced by intergovernmental organisations (igos) such as the international monetary fund, the international energy agency, oecd, the united nations and the world bank. these organizations have a presence in every country in the world, the authority to create international standards and the technical and financial capacity to support the development of national statistical infrastructures. they have long produced high quality, regularly updated time series databanks for their own internal use which typically contain a huge range of macro-economic and social indicators aggregated to national or regional level and collectively cover virtually every country in the world. the academic research community needs access to these unique datasets in order to contribute to and comment on policy responses to global issues. furthermore, access to these data resources give students the opportunity to work with real world data. the need for a geospatial portal ocul member libraries support the use of maps and geospatial data in many ways, including the licensing and management of geospatial data collections and software packages; distributing these geospatial data collections to users; and providing instruction on geospatial literacy and data use. ocul map and gis librarians recognized that there was a significant duplication of effort across institutions providing the first two of these services to their individual constituencies. geospatial data collections were stored locally at each institution, resulting in challenges obtaining sufficient storage space given the vast size of many geospatial datasets. in addition, each institution had its own scholars geoportal: a new platform for geospatial data discovery, exploration and access in ontario universities by elizabeth hill and leanne trimble1 since 2002, scholars portal has hosted shared collections ... iassist quarterly spring 2012 7 iassist quarterly practices in place for distributing data to users, ranging from home grown online data delivery systems, to burning data onto dvd for each individual request. the resources available to provide geospatial services varied widely from institution to institution, with many of the smaller ocul libraries finding themselves challenged to support geospatial data given fewer staff resources and smaller budgets. few institutions, large or small, had the resources or expertise to create and maintain a sophisticated, automated data delivery system. as a result, a significant amount of time was being spent on repetitive data management and distribution tasks, leaving staff fewer opportunities to devote time to work with faculty on the development of strategies aimed at improving students’ spatial literacy and information skills it was because of these challenges that ocul embarked on a new project in 2008, to create a platform for the online delivery of geospatial data resources to the ocul community. by leveraging the shared infrastructure and expertise already available at scholars portal, the geospatial portal project could reduce duplication of effort and improve access to geospatial data by offering a suite of tools that individual institutions would not have the capacity to support alone. together, odesi and the geospatial portal would facilitate the growth of ontario’s research capacity by making statistical and geospatial data more consistently available. developing the vision in 2008, ocul formed a working group to assess the state of geospatial library services across ontario universities, specifically with respect to data collection management, including an audit of how geospatial data files were being used and delivered. the working group also familiarized themselves with geospatial data management in current, international context. for example, the literature on spatial data infrastructures (sdi) provided excellent information on best practices for geospatial data sharing and access. the concept of spatial data infrastructure (sdi) is one which has been around for nearly two decades, and implies “a reliable, supporting environment, analogous to a road or telecommunications network, that, in this case, facilitates the access to geographically-related information using a minimum set of standard practices, protocols, and specifications” (nebert, 2004, p.8). the idea of using established standards for the sharing of data and metadata proved to be important to this project. the working group also investigated the role of geospatial portal applications, and examined a number of existing geospatial portals. two important examples that existed when this research was being conducted included the us government’s geodata one-stop (now known as geo.data.gov), and the canadian geoconnections discovery portal (http://geodiscover.cgdi.ca/). in canada there are also various regional portals that were examined. in 2004, the open geospatial consortium developed a guide to implementing geospatial portals as part of standards-based spatial data infrastructures (rose, 2004). the guiding principle used in this document is the concept of serviceoriented architecture, in which software can communicate with content from disparate sources through a web services model. it was also important for the working group to examine the topic from an educational perspective. at that time, most geospatial portal initiatives were affiliated with government, however, some international organizations had examined the impact on the educational community. the findings of an inventory of geospatial repositories conducted by edina at the university of edinburgh by stuart macdonald, titled data visualizaton tools: part 2 – spatial data in a web 2.0 environment and beyond, served as a guiding document for the working group. this report, published in september 2008, highlights a series of “web 2.0” initiatives that were developed with the primary purpose of permitting users to visualize and manipulate data geospatially, in some cases in an open source environment. macdonald points out that utilities such as these “empower the novice user by enabling the creation of spatial representations and visualisation with a minimal knowledge of the underlying technology” (p.6). as geospatial web computing becomes more advanced, more and more applications focus on the user experience; allowing users to visualize data, perform queries and analysis, create, annotate and share maps, and build a social experience around the use of geospatial data. this is sometimes referred to as an extension of the “web 2.0” trend, or “geo 2.0”. macdonald’s paper validated the urgent need for ontario academic institutions to develop something of this nature. the proposed geospatial portal, a partner to the existing odesi platform for statistical data delivery, would increase use of numeric and geospatial files by students at all levels, thus improving statistical and geospatial skills of students in ontario. since, as macdonald states, “spatial data lends itself to visualisation”, the tool would encourage novice users to increase their comfort level working with data by doing so in a visual way (p.17). in addition, the portal would provide new opportunities for libraries to engage with both students and faculty about the teaching & learning of geospatial concepts. finally, as web technologies continue to improve, it would be possible to offer sophisticated data manipulation and extraction tools for advanced users, meeting the needs of those constituents as well. in 2008, ocul held a “geo-visioning day,” which brought together researchers, students, and librarians to synthesize the research about the current landscape for web gis in teaching and research, and to assess interest, need, and support for a centralized geospatial file delivery service at ontario universities. the geo-visioning day led to the fleshing out of the project’s goals, which reflected the dual need of addressing unequal levels of service at ontario university libraries and improving the use of limited library staff resources, as well as capitalizing on the opportunities of geo 2.0 to develop a suite of tools accessible to both novice and advanced users. in particular, support for novice users was to be integrated throughout the final product. finally, existing data collections were assessed and a need was identified for improved access to health-related data; this became another important aspect of the project. based on these decisions, a project scope was developed, and a successful application was made for support from the government of ontario through its ontariobuys initiative. the project’s official name is the geospatial and health informatics cyberinfrastructure portal project, however the portal itself has since been branded “scholars geoportal”. project governance scholars geoportal has a project governance structure, which oversees the project planning and implementation. the participating committees include an external advisory committee, a project management group, and three working groups:: 8 iassist quarterly spring 2012 iassist quarterly • technical, standards, and collections. this group’s primary task has been to develop the functional requirements and provide ongoing feedback on development progress. in addition, a subcommittee developed a metadata best practices document based on the north american profile of the iso19115 metadata standard. the team also developed a model license for geospatial data, and established a priority list for the loading of geospatial data into scholars geoportal. • health data collections. responsible for investigating health data collections for inclusion in scholars geoportal or in odesi (this is a joint committee between the data odesi and map scholars geoportal groups) • teaching & learning. responsible for developing training and instructional tools of various kinds, and supporting scholars portal in the development of relevant scholars geoportal features. functional requirements for scholars geoportal based on the project goals as developed from the geo-visioning day exercise, and with additional feedback from the project working groups, a comprehensive list of needed functionality was drafted, which was grouped into the following major categories: standards-compliant metadata like in other realms of the library world, standardized metadata is vital to facilitate search and discovery. there are several metadata standards that are commonly used for describing geospatial data. the most established standard is the content standard for digital geospatial metadata (csdgm), created and maintained by the federal geographic data committee (fgdc) and commonly referred to as simply the fgdc standard (1998). however, this metadata standard is gradually being superseded by a newer iso standard known as iso19115: 2003 – geospatial information – metadata (international organization for standardization, 2003). this standard offers a detailed and comprehensive set of metadata elements for describing geospatial datasets, and offers the added benefit of allowing for the creation of “profiles” for specialized user groups or jurisdictions. the north american profile (nap) was created by the american incits technical committee l1, geographic information systems and the canadian general standards board committee on geomatics (cgsb-cog), with feedback from the user communities in both canada and the united states, and was released in 2009. one of the project working groups, the metadata standards working group (which later merged into the technical, standards, and collections working group described above), conducted a thorough review of the relevant metadata standards, and recommended that the project adopt the (at the time) brandnew north american profile of iso 19115, which is also gradually being adopted by provincial and federal government entities in canada. therefore, a metadata editor was needed which met the following criteria: • compliant with the nap standard and allow for the export of nap-compliant xml. • supports the use of multiple controlled vocabularies (selected by the metadata standards working group) for describing topics and places. • accessible to multiple metadata creators/editors who might be located at different ocul libraries. • able to interface seamlessly with the scholars geoportal web application. the project’s metadata standards working group developed a comprehensive best practices guide, which describes how the nap metadata standard is being used for the scholars geoportal.4 online mapping tool the online mapping tool is the main scholars geoportal component that users interact with (available at http://geo.scholarsportal.info). this tool provides ocul students, staff, and faculty with access to the range of consortially-licensed data collections that are available to them through their libraries. for many ocul universities, this is the first time that these data collections are available for online download, both on and off-campus. in planning for the technical implementation, it was decided that the online mapping tool needed to incorporate the following fundamental features: • offer a range of options for searching the metadata collection and displaying search results. • allow the user to display and/or download a complete nap metadata record if desired. • allow the user to preview each dataset in a dynamic fashion within the mapping tool. • allow for the validation of users and the presentation of end-user terms of use agreements. • allow the user to query features and/or view the attribute tables associated with each dataset, to assist in assessing whether the dataset meets their research needs. • allow the user to select an area of interest and download their datasets, which would be clipped to the area of interest (“clip & ship”). downloads needed to be available in a range of file formats and projections. additionally, allow the user to download full datasets when they do not need to clip an area of interest. teaching & learning support in addition to the core data access functionality described above, it was important that additional features be incorporated which would support the teaching of spatial literacy and geospatial data use skills. these features would provide added value to faculty who wish to incorporate scholars geoportal into their lectures, assignments, or other course work, and to students who wish to create simple maps as part of their academic work, without using complex desktop gis software. • integrated help tools. scholars geoportal provides contextsensitive tips and information to assist users in becoming familiar with the terminology used and tools available (for example, what is a projection and how does one decide which projection they wish to use?). this is important for users who may not have gis expertise but still wish to make good decisions when creating a simple online map. • printing and exporting. this feature allows users to generate maps with title, legend, and data source list, and output the map to a pdf or image file (e.g. for inclusion in an assignment). • user accounts. this feature allows users to save maps, searches, and user-defined area-of-interest polygons into a personal user account in order to access them again at a later date. • permalinking. this feature generates a permalink that replicates the user’s map state, opening scholars geoportal and loading their selected layers, zoom and extent information. for example, a professor could use this to share a map or dataset list with their students. • map annotation. this feature allows users to enhance their online map by adding their own points, shapes, and lines onto the map and associate notes or images with these markers. scholars geoportal components a request for proposals (rfp) was issued by ocul in 2009 that resulted in the selection of the arcgis suite of software from esri inc. to serve iassist quarterly spring 2012 9 iassist quarterly as the platform for serving geospatial data and creating a web portal application. esri technology is used in conjunction with other software already available at scholars portal (such as marklogic xml database), and is supported by a server and storage infrastructure as described in figure 1. figure 1. scholars geoportal system architecture diagram in order to accomplish the functional requirement described above, scholars geoportal consists of the following main components. metadata database and editor the metadata collection is housed in a marklogic database. marklogic is a powerful database management system which is excellent for storing xml data. scholars portal uses marklogic to store xml metadata for many of its other services and its staff has considerable expertise with the software. for scholars geoportal, metadata records are loaded into the database and then edited using a custom-built metadata editor, which communicates with the database using the xquery query language. the scholars geoportal web application itself also queries the database directly using an api. the metadata editor allows ocul librarians to load and edit napcompliant metadata records in an easy-to-use interface, as well as export the completed metadata record in xml format. wherever possible, metadata records are obtained from data providers, converted to the nap/iso 19115 standard, and loaded into the database, where additional information is then added (such as controlled theme and place keywords).5 currently, metadata is being edited by scholars portal staff, but plans are in place to build an entitlements management component to the editor, that will allow metadata editors from ocul libraries to login in order to edit and manage metadata records, permitting distributed metadata management. (see figure 2. ) spatial database the vector data included in scholars geoportal are housed in an esri sde geodatabase. raster data are stored on a separate server and are referenced from esri mosaic datasets, which are stored within a file geodatabase. web map services arcgis server allows scholars portal to publish web mapping services which permit users to interact with ocul dataset collections in an online environment. the scholars geoportal web application includes esri map services and image services; arcgis server also supports the publishing of ogc services including wms, wfs, and wcs. scholars portal serves the ocul licensed data collections through map and image services, which are designed and published in-house. these are then invoked by scholars geoportal when a user clicks an “add” button in the search results display, and are added to the map viewer over a base map. (see figure 3: ) web application when making plans to develop an online mapping tool that would communicate with scholars portal’s arcgis server data and services, there were a number of esri application developer frameworks (adfs) and apis to select from. around this time, esri announced that the web adfs for .net and java were to be deprecated in the near future, and that support was moving entirely to the three rest apis, for javascript, flex, and silverlight. since applications built using either flex or silverlight require the end user to download and install a browser plug-in (which is not possible on many locked-down campus computers), it was decided to use the javascript api. the arcgis javascript api allowed scholars portal to build a web application with an embedded map view, which can display a range of base maps hosted by esri, as well as those created by our own web mapping services. users can search the metadata repository and view detailed information about the datasets contained in scholars geoportal (see figure 4). they can then “add” the data to the map by clicking a button that invokes the scholars portal-hosted web mapping service containing that dataset. once data has been displayed on the map, there are a number of ways the user can work with the data. from the “map” tab, they have the option to toggle layers on and off, re-order layers, and change their transparency. in addition, clicking on the map provides attribute information about the features at the location clicked on (see figure 5). in addition, the application offers “clip & ship” functionality, whereby the user can select layers representative of an area of interest, and an output format/projection and the system will generate a zipped file for download (see figure 6). for users without desktop gis software expertise, it is also possible to create, export, save and share an online map from scholars geoportal. the “share” feature generates a permalink to the map created by the user, while the “export” feature provides the option to print or save the map in a range of formats. the exported map includes a user-defined title, important map elements like a legend and scale bar, and a data credits (see figure 7). a future release will include full citations to each dataset on the map, encouraging students to include data citations in their assignments. scholars geoportal also offers map annotation 10 iassist quarterly spring 2012 iassist quarterly features, allowing users to draw their own markers on the map and label them. these can be saved, along with the map, to an individual user account for future access (see figure 8). help tips are sprinkled throughout the portal to answer questions that users may have about the options available and terminology used (see figure 9). in addition, the teaching & learning working group have used springshare’s libguides to create a help guide which offers stepby-step instructions for the different tasks which can be performed in scholars geoportal (see figure 10). other project initiatives additional teaching & learning support while teaching and learning support is a key implementation consideration in the design of scholars geoportal, there is quite a bit more that can be done to encourage spatial literacy and use of gis in a wide range of academic disciplines. to that end, the project’s teaching and learning working group is involved in creating a variety of online resources to build upon the context-sensitive help provided within scholars geoportal. the first of these initiatives is the creation of an online user guide for use with scholars geoportal (http://guides. scholarsportal.info/geoportal). this will be expanded to include screencast tutorials, videos, and other learning resources. scholars geoportal is intended to serve the entire range of users, from novice to expert. the integrated teaching and learning modules will enable novice users, particularly those from non-traditional gis disciplines, to accomplish tasks with ease, developing confidence and expertise with gis tools. these resources will also be important teaching tools, available for use within course offerings and by librarians providing reference support. health data collections the core collections available to ocul students, staff and faculty include licensed collections from both government and commercial data producers. however, as mentioned above, one of the objectives of the scholars geoportal project was to explore ways to increase access to health data collections. this has been an area of weakness within ocul’s licensed collections because many health data stewards are not currently mandated to provide access to anonymized or public use files, and there are many roadblocks in place for researchers, particularly students, wanting to obtain access to health data. figure 2. scholars geoportal metadata editor figure 3: map documents are created in arcmap and then published to arcgis server using arccatalog. the web application makes requests to arcgis server in order to display the service on the map. iassist quarterly spring 2012 11 iassist quarterly in parallel to the technical implementation part of the project, a health data collections working group was formed. this group has undertaken a range of initiatives towards the goal of improved health data access, including organizing a health data summit in march 2011. this day-long event brought together researchers, librarians, data providers, policy makers, and legal experts to contribute to the discussion on access to and use of health data. the summit was a learning experience for ocul, helping the group become more knowledgeable about the range of issues affecting the availability of health information, including privacy and confidentiality, and the challenges associated with the resource-intensive process of anonymizing data so that individuals cannot be re-identified. health data collections will be a long-term initiative for ocul, and we will continue to communicate with data producing organizations to identify ways to overcome these barriers and increase our health data holdings for academic use.. model licensing members of the project’s collections working group (which later merged into the technical, standards, and collections working group), developed a model data license agreement, which is intended to serve as a model for how agreements for licensing geospatial data collections should be structured (the model license is available on the ocul website at http://www.ocul.on.ca/node/114). while it is recognized that data providers/licensees develop their own license templates and requirements, it is hoped that the ocul model license can help inform discussions and negotiations. assessment a number of assessment initiatives have been undertaken as part of the scholars geoportal project. one of these asks each ocul gis/ map library to track statistics about the usage of their services and the time spent on various types of activities (, systems support, data and metadata management, data reference, instruction, etc.). as scholars geoportal usage grows, these data will be useful in assessing what impact scholars geoportal has on the nature of geospatial data library services in ontario university libraries. in addition, ocul has conducted a survey of researchers, which asked them about their use of geospatial data for research and teaching, their sources for acquiring data, their knowledge of geospatial library resources, and what features they would most like scholars geoportal to offer. the results of this survey are currently being analyzed. ocul’s goal is to conduct a second survey once scholars geoportal has been in use for some time, to begin to assess figure 4: scholars geoportal metadata detail view. 12 iassist quarterly spring 2012 iassist quarterly figure 5: scholars geoportal – map with several layers added. clicking on the map provides information about the feature selected (from the attribute table). figure 6: scholars geoportal – download options and list of files ready to be downloaded. iassist quarterly spring 2012 13 iassist quarterly figure 8: scholars geoportal – map annotations and map saving options figure 7: scholars geoportal – exported map. 14 iassist quarterly spring 2012 iassist quarterly whether access to the tool has had an impact on research and teaching using geospatial data. finally, ocul conducted a usability study with student, faculty, and librarian participants, to identify issues with specific aspects of the scholars geoportal web application. the results of this study will inform future development of the user interface. benefits & challenges of consortial projects this project has been collaborative from start to finish. inevitably, there are challenges to such an approach. as rewarding as the work is, all professionals who have served on committees know that tasks can be more difficult to coordinate in a timely fashion when many individuals are involved, particularly when, as in the case of ocul, they are distributed over a very large geographic area. while there were occasional opportunities to meet as a group, most of the decisionmaking for this project was done by conference call. because of the number of parties involved (project administrators, technical staff, librarians, funding agencies, software vendors, etc.), decision-making was at times slower than anticipated, and staying on our timeline was a continual challenge, but deadlines were indeed still met. another challenge of the project was how new this area was to all participants involved. while there was considerable programming and systems expertise at scholars portal, and significant gis expertise within the ocul libraries, no one involved had participated in a largescale web gis development project before. there was a learning curve in a number of areas, including discovering how to communicate effectively between stakeholders with differing areas of expertise. despite these challenges, it is truly the collaborative nature of the scholars geoportal project that enabled its success. without working together within the ocul framework, the consortial data licenses that enabled this project to happen would never have existed in the first place. in addition, the centralized expertise and infrastructure at scholars portal could only be supported by working collaboratively. moving forward on launch (march 1, 2012), scholars geoportal was a robust data access tool that fills a distinct gap in ontario library services. there’s still much more to be done, however. in addition to the ongoing initiatives described above, ocul has an interest in exploring the following areas: expanding the tools the scholars portal team will continue to work on developing all of the features described in the functional requirements list, as well as embarking on new projects that further enhance the tools offered by the geoportal. one area of interest is to find better ways to link the geospatial and statistical data available through both odesi and scholars geoportal – for example, developing data visualization tools that allow users to select census or other data variables, in order to map them to a chosen administrative or statistical boundary level. expanding the collections in addition to supporting the core collections described above, ocul is interested in exploring which other data collections should also be included within scholars geoportal. the project’s technical, standards, and collections working group is embarking on developing a collections policy that will help prioritize the possible initiatives, and this will be a topic for discussion as the initial development project comes to a close. one initiative under consideration is to develop a model for the loading of local data collections (which are licensed by only one or a few ocul libraries). scholars geoportal is able to manage entitlements so that only the appropriate users can gain access to the data, and license agreement terms are respected. for some ocul institutions this would mean that scholars geoportal can become a “one-stop-shop” whereby they do not need to maintain a local system. another area to explore is how we might provide discovery tools for the range of open geospatial data collections now available. conclusions ocul’s scholars geoportal is poised to become an important tool supporting the use of geospatial data within ocul universities. it meets a range of needs, both supporting ocul data, map and gis libraries by centralizing the management and distribution of consortially licensed data collections, as well as offering online search, preview, and download tools for students, staff, and faculty across the province. the project would not have been possible without many hours of hard work contributed by ocul librarians, the scholars portal development team, and the community members who participated on the external advisory committee. while scholars geoportal remains a work in progress, it has immense potential to support future development of new tools, and expansion of data collections. it will be exciting to see how scholars geoportal grows. references canadian general standards board. (2009). north american profile of iso 19115:2003 – geographic information – metadata (nap – metadata), can/cgsb-171.100-2009. gatineau, quebec: author. figure 9: scholars geoportal – example “help tip” – clicking on the question mark causes a small definition to appear. iassist quarterly spring 2012 15 iassist quarterly federal geographic data committee. (1998). content standard for digital geospatial metadata, fgdc-std-001-1998. washington, d.c.: author. international organization for standardization. (2003). geographic information – metadata, iso 19115:2003. geneva, switzerland: author. macdonald, s. (2008). data visualisation tools: part 2 spatial data in a web 2.0 environment and beyond. jisc. retrieved from http://www. disc-uk.org/docs/spatial_data_mashup_v2.pdf nebert, d.d. (ed.). (2004). developing spatial data infrastructures: the sdi cookbook. global spatial data infrastructure. retrieved from http:// www.gsdi.org/docs2004/cookbook/cookbookv2.0.pdf rose, l.c. (ed.). (2004). geospatial portal reference architecture: a community guide to implementing standards-based geospatial portals. opengis discussion paper ogc 04-039 (draft). retrieved from http:// portal.opengeospatial.org/files/?artifact_id=6669 notes 1. contact: elizabeth hill, data librarian, map and data centre, university of western ontario, london, ontario, n6a 5c2. email: ethill@uwo.ca; leanne trimble (formerly hindmarch), map and data librarian, scholars portal, toronto, ontario, m5s 1a5. email: leanne. trimble@utoronto.ca. 2. this paper is an updated version of one presented at the iassist 2011 conference at university of british columbia in the session “power of partnerships in data creation and sharing “. 3. for more information on scholars portal, visit http://scholarsportal. info 4. the scholars geoportal metadata best practices guide is available on the scholars geoportal wiki at http://spotdocs.scholarsportal.info/ display/geospatial/ocul+geospatial+portal (under “documents”) 5. the metadata standards working group recommended two thematic keyword thesauri: the government of canada core subject thesaurus (from library and archives canada) and the lio-mnr thesaurus (from the ontario ministry of natural resources). in addition, they recommended two place name vocabularies: ceonet (used in the geoconnections discovery portal) and the global change master directory’s location keywords. to date, all but the ceonet thesaurus have been implemented within the metadata editor. in addition, french-language versions will be implemented where available. vol252 4 iassist quarterly summer 2001 at the iassist/ifdo conference in amsterdam in may of 2001 a session on “new archives (forum)” was chaired by paul de guchteneire (unesco) and brigitte hausstein (gesis) with participants from new and emerging archives as well as some from the “old world” of existing data archives. the chairs of the session have succeeded in publishing the collection of papers. the reason for re-publishing the session in the iassist quarterly is to spread the word about the archive movement in eastern europe to a broader audience of the full iassist membership and iq readers. in this brief introduction i will take the opportunity to thank the chairpersons and editors brigitte hausstein and paul de guchteneire for their effort and also to thank the authors from several countries in the eastern europe. for a more detailed introduction to the papers you may refer to the introduction by the chairpersons. karsten boye rasmussen march 2002 editor's notes introduction editors: brigitte hausstein, gesis branch office berlin/ central archive cologne and paul de guchteneire, unesco/most, paris. the “new archives forum” at the 2001 iassist/ifdo conference “a data odyssey – collaborative working in the social science cyberspace” the international association for social science information services and technology (iassist) held its 27th annual conference with the international federation of data organizations (ifdo) from may 14 19, 2001. the conference was convened in amsterdam and hosted by the niwi (nederlands instituut voor wetenschappelijke informatiediensten netherlands institute for scientific information services) and the scientific statistical agency of the netherlands. during the last quarter of the 20th century, iassist has held its annual international conference in europe about every four years. iassist did so again, this year in amsterdam in collaboration with ifdo. about 200 data and archive specialists from the usa, canada, eastern and western europe as well as from africa and asia attended the conference. the conference was structured in three different domains of potential sharing of information, experience and expertise: first were organizational matters, like acquisition policy, access to data, and setting up new national or topical archives. second was meta data: standards, tools, and new developments. and third was content: various (new) data types like qualitative, aggregate data, multi-media or geo-data, or combining data from registries. the pre-conference workshops which are a tradition at iassist/ifdo conferences, introduced the participants to the a new standard for meta data documentation – the data documentation initiative (ddi). the most often discussed issue at the conference centered also on the ddi and special attention has been devoted to the applications of the ddi and to the experiences in the data archives. iassist conferences also bring together colleagues from new data archives around the world. in amsterdam a special session for these new archives was organized for the first time. the main focus of the session was on data archives in eastern europe but also representatives of new archives from japan, finland, ireland and greece joined the meeting. the session chaired by paul de guchteneire unesco/most (management of social transformations program), paris and brigitte hausstein gesis (german social science infrastructure services) branch office berlin/central archive cologne comprised two case studies from slovenia and south africa and a forum of representatives of data archives from hungary, estonia, latvia, czech republic, slovakia, slovenia, russia and romania. iassist quarterly summer 2001 5 janez stebe & irena vipavc (social science data archive at the faculty of social sciences university of ljubliana adp) discussed how the new archives could take advantage of experiences of the already existing data archives in the world and heston phillips & patience tshose spoke about how the data service was managed and organized in the south african data archive (sada) which has been existing since 1994. the forum discussion introduced to a variety of topics in the design and implementation of new data archives in eastern europe. the forum member reported about their experiences in, results and the difficulties of creating a data infrastructure in their countries. in short reports they gave an overview of the current situation. jindrich krejci (head of the sociological data archive in prague, czech republic) mentioned that the czech archive had been established in 1998 within the framework of the project “social trends” funded by the grant agency of the czech republic. now it is part of the institute of sociology of the academy of sciences and since 2001 a member of cessda (council of european social science data archives). ildiko nagy (data archive department of tarki, budapest) introduced the hungarian social research informatics center, the first data archive in eastern europe. ludmila khakhulina (deputy director of the research center vciom, moscow) and larisa kosova (head of the information department of the vciom) presented their project “creating a public data archive in russia”. nina rostegaeva (head of the data bank of sociological researches dbsr) reviewed the historical background of the data archive at the institute of sociology of the russian academy of sciences in moscow, that had been founded in 1986 and describes the future plans. the estonian social science data archive in tartu (essda) was introduced by andu rämmer. the archive was set up in 1994 and became a member of cessda in 1997. at the moment it is struggling very hard to keep alive because of the lack of permanent funding. ausma tabuna (head of the latvian social science data archive, riga) described the same situation in latvia. the idea of establishing a data archive in slovakia and romania is relatively new, therefore katarina strapcova, (institute of sociology of the academy of sciences, bratislava, slovakia) and adrian dusa (institute for quality of life research, bucharest) presented their first views on this issue and informed about their plans for the future. ekkehard mochmann (ifdo president, za cologne) opened the discussion by noting that obviously all data archives old and new ones are facing the same problems. they have to cope with financial restrictions, insufficient technical equipment as well as with the lack of well-trained archive specialists. he emphasized that beside financial support these new archives needed also special training in data processing and documentation techniques. in this respect it would be necessary to provide a platform for new archives in eastern europe to exchange experiences in producing and storing metadata. he offered to set up a discussion group on new data archives on the ifdo homepage. paul de guchteneire appreciated the support for new archives in eastern europe provided by gesis/central archive cologne, the finish and swedish data archive. he expressed his hope that this kind of support (he called it “twinning”) would be growing and he pointed out that the unesco/most would also promote both the establishment of further data archives in those countries where such facilities are weak and the establishment of a network of data archives from eastern and western europe. in her closing speech brigitte hausstein underlined: ”the forum has been the first opportunity in the last years for data archive specialists from eastern europe to meet and share experiences. on the one hand the forum provided a comprehensive view of the progress achieved in the field of establishing data archives in eastern europe. on the other hand there is a common understanding of the fact that the new archives still need support provided by the international data and network organizations and experienced data archives. the gesis branch office will continue to foster the cooperation between data archives from eastern and western europe by offering workshops and training facilities”. published at: http://www.gesis.org/en/data_service/ eastern_europe/news/naf2001.pdf i^ssist newsletter vol.1, no. 1 ^ activities and plans cees middendorp suggests that the work of this action group could perhaps be linked to the older inventories by robinson, et al . of the university of michigan survey research center. he notes the difficulties in solving the problems and implementing the recommendations, that standards can only be developed on the basis of research-evidence, and suggests that revisions will be necessary on the basis of either new theoretical insights or new empirical evidence. at the steering committee meeting in edinburgh, erwin k. scheuch of the institute for applied social research and the zentralarchiv presented a description of face sheet problems. he described his analysis of 19 studies in which he looked for comparability of demographic characteristics across studies. his findings indicated that the broad collapsing of categories masked various life cycle and income changes which have occurred over time. on the basis of his preliminary results, he suggested that lassist focus on the need for standardizing demographic variables. he also suggested that further study could be carried out by lassist members and the results be communicated to survey agencies. lassist recommendations could be used by these agencies to perfect their survey techniques and would provide lassist with an excellent opportunity for impact in the area of standardization of demographic variables. this action group has begun reviewing already available material in the field, which includes the work carried out by scheuch and van dusen and zill at the center for coordination of research on social indicators in washington, d.c. this work will form the basis for a project which will test the value of these independent variables in actual research and provide a basic list of such items with recommended coding schemes. in time, this should lead to the developmert of general recommendations for the design of codes. classification canadanot yet activated europeekkehard mochmann, zentralarchiv ftir empirische sozialforschung, bachemer strasse 40, 5 ktlln 41, federal republic of germany united statessue dodd, data library, institute for research in social sciences, manning hall, university of north carolina, chapel hill, north carolina 27514 mandate existing classification schemes for both studies as a whole and variables within studies would be examined. the results of such a review and recormiendations would be published. additionally, information required for the cataloguing of machine-readable data in libraries would be researched and concrete guidelines for such catalogues would be formulated. required classification schemes for the production of indicator source books and variable level retrieval systems will be investigated. the necessary components of an adequate study description will be identified. ^sist newsletter vol.1, no. 1 [editor's note: underlined phrase indicates addition to mandate based on steering committee discussion. another section dealing with teaching materials and modules... for data comparable across nations or within nations has been deleted.] activities and plans ekkehard mochmann distributed a status report describing the present situation regarding study description and classification schemes for social science data as a first step toward problems to be addressed and priorities to be established for various projects. he briefly described the development of standards for study description schemes based on efforts of the danish data archives, steinmetzarchief and zentralarchiv, whereby the study description would serve as a data abstract base, mapping and methodological research base, information retrieval base, data (rejanalysis prerequisite, intra-archival log and inter-institutional exchange instrument. routines are underway at steinmetz and zentralarchiv for linking study description with the individual file retrieval. he suggested that priority should be assigned to implementing a structure which supports efficient information exchange about data between archives and then between archives and potential users. he noted that although classification schemes had been developed, joint efforts in the application and further development of these schemes are still lacking. he recommended that a basis for implementing a scheme could be the "alphabetic subject classification of data holdings of the data library at north carolina, the "item index file" of the roper center, and the zentralarchiv "classification scheme for contents, form and function of survey questions," and his approaches to indicator retrieval from survey archive data bases. in the united states, this action group has focussed on several related issues, but the emphasis has been primarily on the library cataloguing of machine-readable data files in public multi-media catalogues"! sue dodd has used the rules recommended by the american library association's subcommittee on the cataloguing of machinereadable data files to prepare a draft version of a working manual for cataloguing machine-readable data files which will be tested by members of the us action group. she will chair this cormnittee. howard d. white will chair a committee to address the possible structure of and necessary steps involved in the establishment of a national union listing of catalogued mrdf. the other two committees being organized will investigate the marc ii record format for the purpose of storing and retrieving study level information and will prepare a bibliography and critical review of existing thesauri and controlled vocabularies for use in information retrieval systems. it is this latter committee which will initially interact most closely with the european classification action group. 1 3. vol192 27summer 1995 abstract: the historical outline of the danish data archives as an academic service facility is outlined. the reasoning underlying a recent and globally unique organizational affiliation of the dda, viz. to the group of traditional archives, is presented; and the stages of the merging cultures process are outlined. the appropriateness of the archives integration is demonstrated in a presentation of projects that were not feasible in the old university affiliation of the dda. an outlook towards future projects is also given background and history of the dda the danish data archives (dda) is slightly older than iassist. founded in 1973, this author joined the dda in 1974 early enough to be there (in toronto) when iassist was established as a “grass-roots organization” of individuals working in or using data archives, data centers, data libraries, or what these academic service facilities were named in each country or state. in the following, we shall refer to such installations as data organizations (dos). to some extent, iassist was set up as the response from the old boys’ network to the claim of the 1968-generation of more influence or power; to some extent iassist was set up to bridge the gap between the (predominantly male staffed and dominated) european academic data archives and the (astonishingly female influenced) north american data libraries. iassist represented a merging of cultures according to generation, gender, and geography. be it as it may: iassist has survived with astonishingly small adjustments in a changing environment, borne by the enthusiasm and energetic work of (especially north american) individuals. during the same period, many dos have undergone substantial changes. this report provides an overview documentation of some of these changes in a small country (denmark, 5.2 million inhabitants) and refers in the form of parallels to changes in a number of other countries, predominantly in europe. 1.1. feasibility project of the danish ssrc 19731976(1978) the dda was established on april 1st, 1973, after several years of preparation within the danish social science research council (ssrc), as a feasibility project dealing with archiving and servicing problems related to three major data types: 1.1.1. political and social survey data, i.e. questionnaire-collected research data resources in a de facto anonymous form. this was the “typical do activity”, known from e.g. the icpsr, the zentralarchiv, and the esrc data archive. 1.1.2. economic time series and to some extent regional data, both in terms of contents and methods (harmonization, adjustments to regional changes, etc.) and computer handling systems. this area, especially the regional data aspect, was known from norway, where the nsd was started in 1971. 1.1.3. population register data, i.e. identifiable data on individuals. there were no known dos active in this area, but it was expected to be central in the future. needless to say, it was the advent of the computer and the challenges inherent with its use that was the rationale behind the project. during the feasibility project period (19731976), the staff (predominantly engineers!) were occupied with all the technicalities of the computer age; there was, unfortunately, less knowledge (or even ignorance) vis-a-vis the substantive issues within the research disciplines potentially contributing data to and using data from the project. it is symptomatic for the situation that a steering committee consisting of former researchers, now research administrators (viz. the ssrc chairman, an organization professor from a business school; the director of the danish national institute of social research (isr), a government research facility of considerable magnitude and influence; the director of danmarks statistik (the danish central statistical office, cso); and the director of the national archives (also heading the provincial archives) would establish the project with almost exclusively engineers as staff members. it illustrates the attitude that the computer age was still so young that only technical specialists were able to deal with the matter. technicians were the priesthood of the time. when the author of this article (an economist by training, but rather a sociologist by practice) was accepted as a staff member (february 1st, 1974), he was the first non-engineer in full-time employment as an academic staff member within areas 1.1.1 and 1.1.3 above (there were economists in 1.1.2, merging cultures: danish integration of academic data service into traditional archive system by per nielsen1 danish data archives 28 iassist quarterly which lived its own life); only one half-time student had a social science training. many years were to pass until the technical education and skills were considered the “side product” and the social science background was the focus of the staff qualifications. by the end of the three-year feasibility project, the first “culture clash” emerged within the staff: the (few) social scientists felt that the (many) engineers were not appropriately contributing to the development of the organization and, especially, to its integration in the research milieus of the universities and other schools of higher education within the social sciences (broadly conceived). already at this early stage did we (the economists) demand that all staff members reported their time spent on different (detailed) subprojects; of course, the time-use statistics calculated showed that there were too many engineers and a lack of social scientists if we were to fulfill the plans defined by the steering committee. 1.2. interim ssrc period of transition 1976-1978 given that the danish ssrc (contrary to the situation in e.g. norway and the uk, where the ssrcs finance the dos to a great extent even to-day) had a formal limitation on the period of time in which the council was allowed to run projects (three years), negotiations were carried out by the mid-seventies to find the lasting host of the dda. little by little, it was realized that a strong base in a social science research environment was more important than the technicalities; therefore, negotiations were carried out with the relatively large, public “midwife-institutions” (the isr, the cso, and the national archives) to urge one of these established organizations to adopt the techno-baby. given that not enough breeding monies were offered to keep the baby alive at its present size, the negotiations failed. internally, partly because the ssrc gradually shrinked the money sack and partly because the technicians ran projects according to their own interest rather than to the benefit of the baby (shown by the subproject time registration referred to above), a change in staff policy had to take place. more social science trained staff were employed when vacancies appeared (which they did frequently, because highly qualified computer people were in high demand everywhere); and, more importantly, the scope of the dda was narrowed considerably: both the time series subproject and the population register subproject (1.1.2 and 1.1.3 above) were abolished, and only the survey subproject (1.1.1) was kept in the final model. furthermore, the second “culture clash” emerged, this time involving external agents: the understanding and confidence between the dda director (an engineer) and the ssrc members (social scientists) deteriorated; and, in 1977, the directorship moved to the social science side when the former director returned to concentrate on his own private consultancy firm. the final organizational belonging of the dda ended up being decided by opportunistic political/bureaucratic considerations rather than substantive research concerns: the ministry of research and education (whose minister happened to come from and be elected mp in odense!) found it relevant to support the smaller university centers rather than the big universities; consequently, odense university was urged (it cost them money!) to take the baby into custody. 1.3. independent national institute of odense university 1978-1988(1992) formally by april 1st, 1978 (five years after establishment), the dda was moved to odense; the physical move took place at the turn of the year 1978/79. looking in the rearview mirror, this turned out to be the beginning of the consolidation decade, the happy childhood of the baby: after initial fightings over relative budget sizes, the dda ended up in a stable and acceptable economic situation. organizationally, the dda was set up with a double reference structure: on one side, as employees of the university, the dda had to follow the rules of the rector and board of the university. on the other side, the dda had an external board of overseers (five persons) who took care of the more narrow inspection of the activities and the development of the organization. in practice, to be honest, the dda director and the staff took most of the strategic decisions during this decade of consolidation; the baby was free to mature according to its own qualifications and cumulation of experience. more and more, the dda staff identified with the “do culture” (acquired and supported from international cooperation on many different levels and in many different projects) rather than anything else. (this is the kind of “data archive movement” culture that has kept iassist going strong for so many years.) all was well; nobody questioned the relevance of the do culture or the utility of the dda activities, and most staff members considered the odense university affiliation a permanent one. but alas! the centre-right governments of the mid-eighties saw it as their major task to shrink the public sector, and the universities had reductions in their budgets at the same time as there was in increase in student enrolments. universities had to critically inspect their resource allocation; and, needless to say, the eyes of the odense university administration fell on the dda during that process: the university demanded that the 3 academic staff members of the dda (all with titles of associate professors) should participate in the normal social science curriculum of the university, teaching in the same amount of time as all other professors at the university. the dda staff argued that (1) formally, the dda was an institute with national coverage, not an institute with special contribution to odense university; (2) the teaching obligation of the academic staff was fulfilled in national 29summer 1995 training programs rather than in the odense university curriculum; (3) the ministry of research and education gave the budget of the dda directly, exactly in order to make sure that the archive could fulfill its national obligations. 1.4. dispute period and review and negotiation process 1988-1992 in fact, this was the fourth “culture clash”, viz. between more and more strangled university administrators and dda’s relatively “anarchistic do culture” (in the best sense of the term). when it turned out that the “stubborn rector and top administration” of the university were not willing to listen to the arguments of the dda director and staff, we told them that we had to discuss the situation and our future with the dda board of overseers. of course it was annoying and frustrating for the university top management to see that their subordinates did not just obey orders (which they were supposed to do under the “visible management model” which was in fashion). the dda director and academic staff told the university administration that we would opt for a review process if the dda board was in agreement. fortunately, the board members were in agreement; they were even enthusiastic about such a step, because the evaluation or review mania had floated over the country as a politically correct measure in the years of budget cutting a review committee of six established researchers was set up (nominated by three research councils social science, humanities, and medicine and three important research institutions). the review committee report was generally favourable seen from the viewpoint of the dda board and staff; they presented a number of recommendations among which the organizational ones are of interest in this context: the uncertain leadership structure should be abolished; it had been inadequate right from the outset and was critical in times of crisis. the dda should be relocated institutionally, and six possible solutions to the organizational setting were proposed for the dda board to further negotiate. (odense university was not among the institutions recommended; they had been so negative in the review process that they disqualified themselves in the eyes of the review committee.) after discussions with the involved research councils (for social science, humanities, and medicine) and the major research milieus within the same disciplines the board could start negotiating a final placement for the dda, now an adolescent. there were three organizational belongings that were considered interesting, viz.: 1.4.1. the dda as a unit within the danish cso (danmarks statistik). it took only one meeting to be turned down: the director of the cso held that the two cultures could not be merged, especially due to two incompatible phenomena: (1) where the dda had always tried to push their (anonymized) data on as many users as possible, the cso had the principle of keeping their (identifiable) data strictly within the organization itself. (2) where the dda had always succeeded in keeping their services free of charge, the cso tried to earn a big fraction of their total budget by user payments. [needless to say, the dda interest in a cso placement was exactly to change that big organization in a more service-oriented direction.] 1.4.2. the dda as a unit within the danish computer center for research and higher education (uni*c). the uni*c director was interested; she felt that the center should add substance to its predominantly technical services, and they were under transformation so that the integration would be feasible at short notice. the dda could choose between copenhagen and aarhus if they were to go for that model. the transaction was bureaucratically simple, because the dda would stay within the realm of the same government department, viz. the ministry of research and education. 1.4.3. the dda as a unit within the national archives. here again, the dda board was met with relatively open arms (i.a. because the outgoing director of the danish national archives had been functioning two periods (6 years) on the dda board during the mid-eighties). the dda could stay in odense, because the archives were spread over the country anyway. the transaction was bureaucratically more complicated, because the dda would have to change government department, moving from the ministry of research and education to the department of culture. there was an incalculable risk of losing money during such a transfer. all in all, we were quite satisfied with the negotiations. getting a “yes” in two out of three proposals is not all that bad! several rounds of negotiations were carried out with the management of the two possible hosts; i think it is fair to simplify the matters to the following decisive elements that destinguished the two: continuity: because the dda could continue its activity in odense in the national archive-model, there was no risk of loss of professional capacity in that model. there was a risk of losing substantial parts of the “do culture” in a geographical move and thus a risk of assimilation with the new culture 30 iassist quarterly (maybe even annihilation of the “do culture”) rather than integration into the new culture with dda’s own cultural identity relatively intact. permanence: the national archives, being several hundred years old already, and being one of the only institutions mentioned in the constitution, will survive new centuries. uni*c, on the other hand, was already undergoing severe changes in business plans in transition from being predominantly a mainframe host to having a wider agenda: mainframe host (parallel processors and other very expensive equipment), facility management host, network administrator, and value added services agent. substance: the major argument, however, was that the substance dealt with in the traditional archives and in the dda was the same: both are information agents, the major difference being the data-carrier which will change in the traditional archives anyhow. many avenues of dda development were more easily passable in the national archives model than in the uni*c model. the choice having been made, only the bureaucratic work remained; and even though this process took considerably longer time than expected it ended succesfully: as of january 1st, 1993, the dda was a unit in what had, in the newly enacted archives act, been named the danish state archives (sa). we could thank our board members (whose assignment period had twice been prolonged with one year because the transition took so long to carry through), and we were cast in the arms of a new host 1.5. independent unit in the danish state archives group from 1993 the “anarchistic do culture” had to be integrated into the “bureacratic civil servant culture” according to the decisions taken. as always when you move in with new people, there was some reluctance and cautiousness from both sides: from the dda point of view, we insisted on staying separate for some time to secure (reassure) the independence; we were not going to be “swallowed” by this, as we considered, somewhat “dusty” system ten times larger than we were. the entry avenue was paved with a number of lucky circumstances: (1) a new director of the national archives entered the arena a couple of years before us, and he came from the university and research circles, too; (2) yet another unit had been adopted in the state archives only three months before us; (3) a modernization process had been started within the archives themselves. partly due to these circumstances, the entry into the new world (which is a very old world!) was succesful and seems to develop to the benefit of both sides. before looking into that, however, we shall make a short digression to a description of our new “family” and then return to a specification of the potentials of the new affiliation of the dda. short description of the danish state archives before the advent of the archives act of 1992, the state archives were referred to as “the national archives and the provincial archives” most of which were century-old. in the archives act of 1992, the state archives (sa) was defined as a group; we shall very briefly introduce these institutions and the rest of the archives complex in the country. the danish state archives have less than 200 man-years at their disposal; quite a considerable number of the employees, furthermore, are not regular employees; rather, they are unemployed or disabled persons undergoing training or rehabilitation programs on behalf of social authorities. the staff-size og the danish state archives in comparison with the size of the state administration that they serve is considerably lower than in the other nordic countries, a fact which has been demonstrated to the politicians again and again. 2.1.the national archives the national archives (rigsarkivet) and its predecessors (i.a. geheimearkivet, the secret archives) date back some 400 years. the institution is located face-to-face with the danish parliament (folketinget). with approx. 80 man-years available, the national archives is obliged to make an appraisal of all documentary material in central government and archive what is deemed necessary from legality considerations and to document the present for future researchers. the national archives is divided into an appraisal branch (incl. a private archives unit, a military archives unit, and an mrda unit) and a servicing branch; also, the institution hosts the secretariat of the whole group of the danish state archives. 2.2. the provincial archives there are four provincial archives. their purpose is to provide archival facilities for government agencies spread over the country. also, voluntarily, the county and municipality administrations may deposit their archives with the provincial archives; however, they have to pay. even so, they have to abide by the principles for appraisal defined by the state archives (formally: the director of the national archives). three of the provincial archives (for zealand and the other islands east of the great belt, in copenhagen; for the island of funen in odense; and for northern jutland in viborg) are exactly 100 years old here in the mid-nineties. the fourth provincial archives, that of southern jutland in aabenraa, is only about 60 years old. it was established some years after 31summer 1995 the referendum in 1920 which brought southern jutland back under the danish crown; it cooperates closely with archives in schleswig which remained german as an outcome of the referendum. two provincial archives (in copenhagen and viborg) are “big” (approx. 35 man-years), two others (in aabenraa and odense) are small (approx. 10 man-years). 2.3. the danish national business history archives founded as an independent state-financed institution in the fifties, the danish national business history archives tries to reflect all aspects of business life: it holds archives from firms and business units as well as from organizations (employers’ organizations, employee’s organizations, private organizations and associations) as well as from individuals with a certain standing. needless to say, before as well as after the entry of the danish national business history archives into the state archives group (entry per october 1st, 1992), there has been a need to define the functional dividing lines between that institution and the private unit within the national archives. opposite the major volume within the national archives and the provincial archives, the danish national business history archives has to rely exclusively on voluntary depositing of material (much like the dda); they have no legal claim that donors shall archive their administrative remains. the danish national business history archives has less than 15 man-years at its disposal; within that frame, it also serves as a municipality archive for the city of aarhus where it is situated. 2.4. the danish data archives the dda entered the “family” on january 1st of 1993; it had about 10 man-years of staff-time at its disposal in the operating budget when entering. due to the uncertainties regarding affiliation in the late eighties and early nineties, it had become extremely difficult to attract research grants to augment the total level of activity. needless to say, there are donors of computer archives that may either deposit at the national archives (mrdf unit) or at the dda; we shall refer to the “functional integration” in some detail below. 2.5. other archive groups (not in the danish state archives) outside the “family”, a number of archive institutions are of interest in terms of collaborative projects (private archival material) as well as because they rely on the definitions of the sa in terms of appraisal (city and local archives). the major groups are: 2.5.1. the national library: as per tradition, many private papers (especially from writers, artists and other actors in the cultural realm) end up in the national library (next nabour to the national archives in copenhagen). 2.5.2. the labour movement’s library and archives: financed by the labour unions, this library and archive documents the labour movement in denmark and is thus also predominantly in the private archives sector. 2.5.3. the city and local archives: according to the archives act of 1992, counties and municipalities have an obligation to keep their records according to the decisions taken by the director of the danish national archives; however, they do not have to deposit the records with the provincial archives. more than a dozen of big city municipalities have established city archives with a professionally trained archivist (usually a historian) as the head. in many minor municipalities, the local archives have been staffed only with amateurs in the past. from the archives act of 1992, however, local archives have to be part of the municipality administration and professionally managed; otherwise, the records shall be deposited with the provincial archives of the relevant region (paid for by the municipality). 3. advantages and disadvantages of archives integration from the national viewpoint, the archives act of 1992 explicitly regulated that all public authorities shall deposit their archives in a “professional” archive institution. this is, of course, an important step in the direction of securing future historical research at all levels, the national, regional, and local. in this section, however, we shall return from the digressional “family description” and look at the advantages and disadvantages of the integration of the academic service facility (the dda) into the traditional archive system (the sa). without doubt, the viewing angle is that of the dda due to the fact that the author is placed there, and because that is the “natural” iassist platform for evaluation. 3.1. the “laissez-faire period” as already touched upon above, the first year or so in the new family was characterized by a “laissez-faire” state of affairs in the sense that all parts did what they used to do without much interference. it was a period of gradual confidence-building. however, the period was also one where the activities of all units in the sa were thoroughfully documented in a 3-volume action plan. when the dda entered the sa, they were in the middle of this documentation process; so they could immediately add the dda resources, products, and services to those of the 32 iassist quarterly other sa-units so that the final report presented to the ministry of culture provided an overview of the whole new group of the danish state archives. based on the sa action plan 1994-1998 that was published in three volumes by the end of 1993, a so-called performance contract was undersigned between the ministry of culture and the sa in 1994. the idea is that the archives get more resources (approx. 10 man-years) in return for specified improvements in performance (efficiency, servicemindedness, productivity). the first performance contract is running in the period 1995-1996, only; however, it is anticipated that a new contract be designed for the period 1997 through 2000 by the end of 1996. during the “laissez-faire period” there were not many advantages or disadvantages of the new host situation. life went on pretty much as in the past; the dda was left with the same resources and the same tasks as under odense university. however, on the positive side, this generated confidence that the sa system was not going to “swallow” the dda; on the negative side, some resources had to be spent on statistical reporting and planning activities that were not immediately to the benefit of the dda and our “traditional” user clientele. 3.2. the integrationist period gradually, as the work with the action plan 1994-1998 proceeded, it became necessary to define what was labelled “functional integration” (in fact meaning specialization) within the sa group. in short, this means that, opposite to the century-old tradition, not all units can upkeep all the specialties of the archival business. for instance, all the production and distribution of microfilm and micro-fiche will take place at one “virtual unit” (which happens to be located within one physical unit, viz. the provincial archives in viborg). similarly, the conservation activities are being collected in another “virtual unit”, in this case spread over 2-3 physical units. furthermore, we work with the notion of “specialist archives/archivists”, meaning that one unit (and one archivist within that unit) is the sa specialist vis-a-vis one type of authorities (e.g. police authorities, county archives, hospital patients’ files). turning to the mrdf material, there are two centers in the sa system: the dda takes care of everything from the “private sector” (incl. research). also, the dda is responsible for research remains from many public authorities (e.g. the isr and an institute for clinical epidemiology) and for a number of semi-public institutions (e.g. the cancer register, which is now being moved from the de jure private danish cancer society to the public realm). it took tough negotiations to define these functional division lines between the dda and the mrdf unit of the national archives. a fifth “culture clash” appeared between dda’s service-oriented activity, international orientation, and informal contact methods on one side and the mrdf-unit’s acquisition-oriented activity, relative isolation, and formal contact methods. furthermore, the mrdf unit of the national archives was stuck with very old equipment whereas the dda has been trying to be at the technical frontline. so what’s the difference, the sceptic might ask; hasn’t the dda held the danish omnibus surveys, the danish welfare studies, the danish time budget data, and other material from the danish isr all the time? yes! but there is a difference, and the difference is twosided: firstly, the dda now holds not only survey materials that are de facto anonymous as before; the dda can now hold materials that are registers according to the danish acts on public and private registers. secondly, with respect to public authorities, the dda is not dependent on the willingness of the agent to understand the importance of archiving; if the director of the national archives and the director of the data surveillance authority agree that a register shall be archived, the dda staff can collect that register from the data owner in a capacity as an archive authority. even in terms of research registers (especially medical registers) there was some reluctance to give very sensitive patient information to an archive that was a university institute. being part of the “official archive system” improves the chances that single researchers and research groups are willing to deposit their materials. as a consequence, more data materials will be available from the dda for future research under the new model. the advantage for the dda (or rather for our traditional user clientele) is that more research relevant information will be available for secondary analysis. the disadvantage, seen through the glasses of the dda staff, is that more tasks are placed on our shoulders without a corresponding inflow of personnel resources. furthermore, the dda senior staff is heavily involved in tasks (appraisal of computerized stuff from public authorities that are not immediately of interest for our research users; modernization of the other units in the technical sense, incl. establishment of a new version of “their” archives data base on a new platform) that make life busier without augmenting the service level towards our primary users. 3.3. the immediate future like in many other countries, the politicians and the broader public are very interested in the so-called “information society”. in denmark, a government committee report (“the information society in the year 2000” was published in the autumn of 1994. it was immediately followed by the 33summer 1995 establishment of a separate ministry of research under which the national it-strategy was located (in accordance with the recommendations in the bangemann report from the european commission which appeared a few months earlier than the danish info-2000 report). so, in march of 1995, the government produced its annual it-plan “from vision towards action: the information society in the year 2000”, which in some respects looks like the clinton/gore initiative in the direction of information superhighways, in other respects is encompassing a lot more due to the special character of the danish society. the government it-plan for 1995 and the sa performance contract with the ministry of culture require a lot of decisions from the sa. just to mention a single challenge with a long-range perspective: before mid-1995, the sa is going to define the rules and procedures that we deem necessary in order to allow the authorities to adopt the practice of “the paper-less office” from the beginning of 1996 (paper-less, because incoming paper-mail is scanned and saved (e.g. in a tiff-format or equivalent), and where in-coming e-mail as well as outgoing mail of all types are saved in searchable format, e.g. in the sgml-format with a well-defined dtd or in other expectedly long-term viable formats). 3.4. the longterm perspective the advantage for the dda of the placing within the state archives system is, of course, that we “archived the institution” within a long-term viable institutional structure, forming part of the national information strategy. as many iassisters will realize (more or less horrified!), the whole raison d’‚itre of many data libraries may vanish within a very foreseeable future due to the fact that end-users can download their research resources directly from the producers or other facilitators on a global scale. in the near future, academic data service organizations will face a strong competition from private and quasi-public vendors trying to monopolize their services not unlike the way that many (european) csos have done in the past. the information society involves rapid institutional changes even to the information specialists. to establish a condition with freedom of information (and equal access) is no longer a question of some academic institution-building, only. national, and in turn international (in europe e.g. within the european union) information strategies will be developed from the political level, and they will severely influence the survival conditions for most of our academic service institutions. 4. projects facilitated by archives integration the functional integration of the academic data service and the traditional archive system has already had an impact on the “palette” of activities of the dda. below, we shall touch upon a few projects that are facilitated by this integration. 4.1. the source entry project like the other scandinavian countries, denmark has excellent demographic sources. in order to ease the access to those sources that may account for so much as 80% of the use of traditional archival material, the traditional archives have had large projects (in part jointly with the mormon church) producing films and fiches with these sources. the film/fiche versions of the sources have in turn been distributed to city and local archives, thus releasing the increasing pressure on the sa reading rooms. it goes without saying that such demographic sources invite computerized treatment. and, indeed, many amateur historians and genealogists (organizationally cooperating within the association dis-danmark) have been entering a lot of these sources into computer programs. these source entry initiatives, however, were scattered in coverage, differing in quality, and more often than not non-transferable because of technical limitations. in 1992, dis-danmark formed a cooperation committee for source entries (danish acronym: saki), and several staff members from the danish state archives (incl. hans hans j¯rgen marker from the dda) were invited to serve on that committee. during less than one year’s work (1992-1993), this committee completed a set of recommendations (called the saki model) for the creation of machine readable source editions of structured sources (published in a special issue of the dda quarterly newsletter dda-nyt). the recommendations should secure higher quality of the products from this huge amateur project. in order to improve the transferability of data, a special source entry program (kip) is offered to people who want to serve as source entry personnel. furthermore, the dda serves as the central archiving facility and distributing service for all these computerized sources. finally, a coordination committee (koki) keeps track on who is doing what to avoid duplication of effort. at the dda, the computerized sources are standardized and documented. and the dda can supply copies of sources (usually in paper-form, but if needed also on film or microfiche) free of charge to people who are willing to and capable of making contributions to the program. by the end of april, 1995, more than 5.1% of the 1845 census is available, with the 1787 census in second place (3.4%). in a not too distant future, such frequently used demographic source material as censuses and church registers may be available in data form as well as in the form of scanned images (based on the film/fiche versions). this will revolutionize the nature of use of such sources and open a lot of new projects: person recognition and family reconstitution based on neural networks; automatic movements up and down family trees in a graphically based environment; etc. 34 iassist quarterly 4.2. the computerization “rightsizing” as mentioned above, the “old sa-family” used technical equipment from the mid-eighties (mini-computer technology from norsk data); the system is completely closed from the outside world, because this was considered necessary to secure confidentiality at the time of installation. in august-september 1995, all units within the sa (except the dda which will be on that platform already) get new client-server equipment after specifications laid down in a group where the dda has held the chairmanship. this means that the sa-units will be able to benefit from the resources on internet and other communication networks, and it implies that a strategy can be adopted where the descriptions of the materials in the archives can be brought to the users electronically. dda is heavily engaged in a rescue operation where an existing (hierarchically organized) archival data base is going to be transferred to the client-server environment and entered into a relational data base system (viz. ms nt sql server). although these technical cooperation projects have drained resources from the dda, they do hold a perspective for the future: because the dda has a longstanding experience with user contacts (the mrdf unit in the national archives has only served about a dozen users since its inception in the early seventies; the dda has several hundred user requests a year), we may well be disigning the user interfaces for the whole sa “family” in the future. 4.3. the register research facilitation as an initiative of the danish research foundation, a working committee (where the author of this article was a member) has been defining a model that might facilitate the use of personal registers (incl. registers in the cso) for research purposes. the recommendations of the committee, to establish a register research center, adjacent to but independent of the cso, and to establish a register archival facility at the dda) are being implemented right now. the dda could not have played an active rùle in this project without having an authorization to hold identifiable personal records. the idea is, furthermore, to take on a medically trained staff person to make sure that the many registers in hospitals and medical departments be rescued, to the benefit of contemporary and future research. to people outside scandinavia a comment may be relevant: the danish society is administered almost completely via computer registers; the citizens are registered with the cprnumber (central personal number) as the unique identification code. this implies that all types of personal information may be merged in research projects, and this is of great importance so far especially within medical research. this register-based research potential is considered to be unique for the scandinavian countries. 4.4. the government system contacts the placing of the dda within the danish state archives seems to have brought us closer to the government system than we were under odense university. this implies that the dda has been represented on numerous committees and working groups where “the future is designed.” this being said, we still try to keep the “anarchistic do culture” as our life-style and the equal access to information as our distribution principle. to do so is facilitated by a comment from the political system in the report leading to the archives act of 1992: the leading principle in the administration of the archives act is going to be to secure “the greatest possible openness.” 5. closing notes: merging cultures in less than a quarter of a century, a lot of “culture clashes” have been experienced by the dda internally and in the contacts with the outside world. my guestimate goes that we are now going to see a “reverse process” a merging of cultures where there are not so many “computer-nicks” or research-discipline monopolists who claim their superiority. so much information will be readily available that technical and human network-building as well as inter-disciplinary sensibility, cooperation and understanding will be much more important than media-oriented or discipline-based exclusiveness. in the danish case, the merging cultures are visible in two respects already demonstrated in the project descriptions above: firstly, from being a “traditional” social science data archive holding survey data relevant for the political and social sciences, the dda is rapidly moving into a position where historians and medical researchers are added as new user groups. secondly, since we had to abandon the population register data subproject (cf. 1.1.3 on p. 3) in the mid-seventies, the activities have not included register data. there is no crucial difference in method analysing survey or register data; the two should complement each other rather than being seen as two different approaches. more often than not, register research projects will contain a process where subpopulation data held by the researcher have to be merged with register data held by some public authority; therefore, it seems logical to have the services and the data resources collected in one place in a small country which cannot afford to have several, discipline-specific data service organizations. on the danish data arena, only the economic time series (incl. the regional data, cf. subproject 1.1.2 above) are not yet incorporated in the service “palette” of the data service unit; and, to be honest, i think that they should not be! time series data should be available from the main producers, viz. the csos. needless to say, they will be entered into the archives for historical research in due time; 35summer 1995 but as far as contemporary research is concerned, the time series data should be distributed by the producers and if they introduce obstacles, we should concentrate our energy on removing these. the major reason why i find that contemporary (economic) time series and regional data are unappropriate in academic dos is that they are constantly changing in the course of time (new weekly/monthly/quarterly/annual figures should be added) and because of changes in administrative regions (which necessitates a backward harmonization). in conclusion: the technological development will have a crucial effect also on the institutional landscape a decade from now. we already face the rapidly changing conditions of our activity brought about by the internet and www services; so far, we (the do personnel) can feel easy at the frontier because we know more about these advanced technical information interchange facilities than most of our users. but take care: new generations of users are entering the professional scene; they know “the computer age” because they already grew up in it, and they will ask for services in terms of selective information facilitation that we are not yet able to produce. there are plenty of challenges for iassisters for the next couple of decades. after that, many of the iassist pioneers can sit back in their homes, living on their pension schemes, communicating with each other about the rapid-changing world and the oddities of the younger generations. the topics old people always communicated about ... but we shall be in the favourable position to communicate electronically and globally! 1 paper presented at iassist 21st annual conference may 9-12, 1995, quebec city, canada. vol252 16 iassist quarterly summer 2001 description of the data bank of sociological researches the data bank was established in the institute of sociology of the academy of sciences of the ussr in 1985. in the first place, the bank’s principal purpose was the storage of empirical data from sociological studies in a way suitable for repeated use. in 1987 the data bank was given the all-union status under the sponsorship of the soviet sociological association. the bank was set up by a number of organizations interested in the joint use of accumulated empirical data. these organizations represented practically all regions of the former soviet union. after the collapse of the soviet union, the territorial scope of the bank was reduced, and new organizational and financial problems appeared. however, the material accumulated by dbsr reflect more than 30 years of the society’s life and therefore it is of tremendous scientific and practical value. since 1992, because of the political developments and the resulting breakdown of communication links between the former soviet republics, the data bank had to confine itself to russia only. it still contains essential information on the russian society as well as on the other parts of the soviet union, from central asia to the caucasus as well as on the baltic states, moldova, ukraine and belarus. today taking into account the fundamental changes which have swept our society since 1991, we consider that these data are unique and have historical importance. the secondary analysis of them will allow researchers to trace and locate the origins of many social and economic processes in modern russia. in contrast to tendencies of isolation shown by some sociological research centers, our goal is to preserve and improve old links and establish new ones with former member organizations of the data bank in other countries. our position is that scientific and informational space that was formed in the course of many years must not be torn apart by the requests of individual politicians. the creation of dbsr was preceded by years of methodological studies and applied scientific research into questions of the accumulation, storage and use of sociological information. these problems are reflected in the works of major russian scientists such as v.g. andreyenkov, v.i. molchanov, and others. these studies were conducted in co-operation with leading foreign experts in the field such as r. bisko, g. heiman, d. nasatir, g. marx. e. mochmann, g. kjabb etc. therefore a concept has been developed that has the idea of consolidation and integration of sociological data at its core. this is the principle on which the bank is based. consolidation refers to the accumulation of a large quantity of empirical data, which provides the basis for conducting secondary and comparative analyses of data sets produced in various sociological research centers. it opens wide opportunities for social modeling, numerical experiments, verification and evaluation of methodology, formulation and solution of methodological and methodical problems. data integration requires the integration of separate data into a system of interdependent indicators, describing the society as a whole and in its parts. in spite of long drawnout discussions on the structure of the indicator system for separate sociological phenomena, many problems of methods and methodology remain. the bank possesses expertise in establishing the systems of indicators. the problem of storing empirical sociological data has arisen in our country in the late ’60s at a time when the first large-scale sociological studies were carried out. only in the early ’80s however, with the wide spreading of computers, the question of automated information system to serve the needs of sociologists moved into the practical realm. the problem was solved when the fundamental principles of data banks as integrated information systems had been worked out. “the data bank of sociological researches” refers to multi-functional information and analytic systems aiming at one goal: the accumulation of various sociological information in a systematic form, including empirical data, for increased efficiency in its use. among the bank’s many functions we emphasize the following ones: • improvement of methods, means of accumulation, and analysis of sociological information; • methodological research with the aim of standardizing the methods and separate blocks of development and prospects of the data bank of sociological research by nina rostegaeva 1 iassist quarterly summer 2001 17 indicators for subsequent comparative analysis of the data obtained at different times, from different territories and in different social and cultural environments; • information and reference resource for sociologists; • co-ordination of sociological research by informing users of recent research and new empirical data; • exchange of primary empirical data access; • providing conditions for secondary and comparative data analyses; • calculations, based on commercial agreements, including sociological modeling, experimenting etc. during the development of the data bank various factors such as technological progress, continuous increase of the information volume to be stored and growing information needs of sociologists were taken into account. the bank’s goals and functions determine how information is stored. it consists of the following three databases: • summaries of studies (topic, time, methods of gathering data, sampling procedures etc.); • summaries of documents (forms, questionnaires) used in gathering data • data from empirical studies, stored in computerized form (at present these data do not form one database, but the system in use still allows an unlimited access to this information). dbsr’s users are research groups conducting theoretical or practical work in sociology. they provide the bank with materials from their own research and obtain in reverse information they require. in addition, the bank informs its users about new data received by the bank and publishes a frequently updated reference: “bank of sociological data” (bank sotsiologicheskikh dannyikh). the files of empirical data are divided into three classes based on their accessibility for users. authors wish to reserve their rights to the results of their studies for a period of time, after which their data are placed into a more accessible class. more than 650 studies conducted by the institute of sociology and other research centers from 1966 until 2000 are stored in the bank. these include over 20 union-wide studies. this research reflects all aspects of society’s life, including the political and structural changes that have taken place over the last few years. the dbsr encompasses investigations of social tension, national conflicts, stratification of society, emergence of new classes, strata and groups, problems of the transition to a new political and economic structure, problems of family and children, the status of women, problems of youth and education, structural changes in society, ecology, and demography etc. most of these data sets are available to those investigating soviet and post-soviet periods of the society. prospects with advances in the information science and computers as well as the communication technology getting ground, geographical distances are no longer barriers to the communication between researchers from different countries and continents. modern technology has given a great impetus to the modernization of the data bank. the improvements which are being planned are based on conceptions of universal services with interactive access. these improvements will stimulate the researchers’ interest in developing a national network of sociological information and attract investments. modern communication technology between various research centers will reduce the costs of creating and maintaining archives. thanks to it researchers will have access to the bank’s computing and information resources. the main efforts in the course of the practical implementation of the above conception will be directed towards the following goals: • universal and accessible services for users: the creation of an information search system, the establishment of a special database for registered studies, the implementation of interactive information access facilities; • mobility of data, which presupposes the observation of the storage standards, permitting input, retrieval and information exchange as well as the protection of intellectual property; • unified telecommunications network connecting various centers where sociological data are stored and analyzed and giving access to the powerful technical and informational resources of dbsr; • international co-operation and co-ordination of the effort aiming at including the bank into the world-wide network of archives and banks of sociological data. achieving these goals will allow the modernization of the existing system which serves sociologists’ needs, eliminate incompatible database formats, create a network infrastructure for the collection and distribution of knowledge on the society. let us discuss some of these goals in greater detail. accessible service reaching this goal will ensure a convenient and efficient interactive information access and exchange. it will 18 iassist quarterly summer 2001 revolutionize the information search and permit easy access to information resources. it is intended to achieve these goals by creating an interactive information retrieval system. yet, well-organized resources and effective means of access can yield the expected benefit only when every researcher is sure that he will find the requested information. for this reason the existing three-level access system of dbsr will be reviewed and changed in favor of greater accessibility (while observing the authors’ rights). business-like openness of dbsr is the principal idea of universal services. the information environment is quite variable. data from empirical studies come in different amounts and in a variety of types. code-books with distribution tables are also stored there. the machine-readable catalogue connects all types of information contained in the bank. the consolidation of stored information and its integration into a database gives full and coherent information on any of the registered projects. such an environment supports many types of information objects and establishes connections and relations between objects easily discernible. this approach is based on the strategy towards an open architecture of the dbsr, which refers to a collection of a variety of independent information sets united in integrated databases and to the ready access to this information and automated search. an information retrieval system is intended for serving the integrated information environment. data mobility creating the integrated system requires the unification of the information stored. this can be achieved by developing standards for new data and adjusting the existing data to make them conform. the specifications for the standards for new data is based on established world standards, and this standardization is an impetus for an extensive information exchange. the mobility of data will allow the distribution of initial and secondary analysis in the telecommunications network. users are thus guaranteed a conflict-free interface. the new standards will cover the elements of the catalogue and the elements of databases making up the integrated medium. first of all, the standards are applied to research data stored in the dbsr as spss portable files exported from the mainframe to a pc in a wrapped-up format. the new standard must take into consideration various software options, allow reliable storage and presentation of information, and assure compatibility with the world standards for free export and import of information. in this paper we do not try to describe the particulars of standard levels, but only remark on the range of problems in standardizing data and on the general direction of realizing such a project. when it comes to ensuring the mobility of data, the question of authors’ rights and the related issue of regulating the access to data is another important consideration. registration and identification of users, as opposed to the three-level access system, seems to be the most constructive approach. the mobility of data will promote international cooperation which will in turn, significantly improve the dbsr’s position in russia. international co-operation at the core of all foreign and russian information centers is the idea of international cooperation. it is especially relevant for russia, where the support for dbsr’s efforts in transforming data to conform to international standards and following the strategy of the bank’s development will allow the preservation of unique data on the society in the socialist period and about its subsequent transformation. international cooperation will be of great benefit for the process of collecting new empirical data on the formation process of the society and will play a decisive role in the integration of local national centers. references andreyenkov. v.g. & cherednichenko, v.a.: k voprosu o sozdanii banka sotsiologicheskoi informatsii. sosiologichesklye issledovaniya, 1993, no. i molchanov, v.i.: sistemnyi analiz sotsiologicheskoi informatsii. moscow, nauka. 1981. garskova i.m.bases and data banks in historial researches. moscow, 1994. rostegaeva n.i. the data bank of sociological researches: the invitation for co-operetion. in: “sociology: methodology, methods, mathematical models”, 2000, 12. 1. contact: russian academy of sciences, institute of sociology, krzhyzhanovskogo 24/35, building 5, moscow 117 259, russia. phone: +7 095 7190940. fax: +7 095 7190740. email: bank@isras.rssi.ru vol221 spring 1998 19 the world’s communications and exchange have changed dramatically, bringing us closer to the notion of global village. globalization has become the byword of our era. libraries seem to have lost their clarity of definition. where a library exists is no longer important, but what a librarian performs counts. librarians have long been experienced in organizing knowledge and serving the user. their roles have been changing with social advances. from the bookkeeper and custodian in ancient times to the reference librarian and the chief information officer in late 20th century, the scope and meaning of the term librarian is expanded. the primary driving force is the information and communications technologies. as a result, many new titles are facing librarians, such as information navigator, information broker, information engineer, etc. to sum up, three major roles are waiting for librarians to assume with the coming of the new millennium: global information provider, educator and trainer, knowledge manager. global information provider the changing characteristics of global information environment can be summarized as: automation of the information infrastructure of the whole society; the increasing popularization and deepening of computer and communications networks; and the multimedia dissemination of information. library managers in global information environment as we reach global information environment, library managers should have cross-cultural management competencies [nicholson and rochester, 1996]: • transformational management skills—shifting from attitudes and behaviors that are ethnocentric to ones that are cross-cultural and mastering new drivers of competitive success • interactional management skills—understanding how leadership, motivation and staffing practices are addressed in differing locations • transactional management communications skills— understanding how to market operations successfully and work well with colleagues and mastering a complex, fast changing and possibly unfamiliar competitive environment information provision in global information environment. the traditional acquisition, organization and distribution of information is no longer enough for both information users and librarians given the exponential growth of information technology and ever-increasing demand for information service. the future is one of users accessing electronic data and catalogs of electronic and printed collections anywhere in the world from workstation unfettered by local, institutiona, national or geographical consideration. so, information service meets challenge and threat at the same time. librarians should learn to be proactive rather than respond to changes. librarians are disconcerted by the marginalising of much of their roles as guardians of intellectual heritage and share a common concern with all others in information provision. as librarians, we are accustomed to viewing ourselves as primary information providers. libraries house a wealth of information, collected with substantial knowledge of our clients’ evolving needs. but, a new information environment is facing us. online catalogs offer much greater access to library collections; computer-based indexes and databases are more comprehensive than printed ones; and interactive multimedia can provide wellstructured independent instructions in information retrieval and other skills [mclean, 1996]. the use of internet has pervaded every domain of library work. making collections available to people regardless of locations—“the library beyond the walls” will increasingly become a focus of activity. librarians should not only bring collections to the user, but service as well. for example, email reference service is possible now [hardy, 1996]. in the field of global information provision, the internet is a constantly evolving global network of networks, which is transforming the way we communicate, research and live. users can take on many kinds of roles: world traveler, foreign correspondent, explorer, publisher and so on. the internet is comprised of many different information resources, and numerous applications are available for the the expanding roles of librarians for the new millennium by jinhong tang * 20 iassist quarterly purpose of internet searching. because of massive volumes of heterogeneous information sources on the internet and browsing as a major access paradigm, the process of matching may not be effective or efficient. the consequences of such an environment may be: unused information; tiring retrieval process; low recall and low precision. information provision in an electronic environment is not an easy process. the challenge to librarians is to create an agreeable environment for electronic information retrieval. librarians can fulfill this task by facilitating in electronic information retrieval and consummating indexing. • facilitating electronic information retrieval to come to grips with the internet resources is a promising mission for library world nowadays. many libraries are trying their best to act as pathfinders. let us take hong kong university of science and technology library(hkust) as an example [yip, 1997]. early in december 1994, a pilot group was set up to work on the library world wide web project at hkust. its objectives are: (1) to develop selection guidelines for the internet resources; (2) to build the internet navigation skills among selectors; and (3) to select free internet resources and make them accessible to the hkust community through the library catalog or hyperlinks on the library web server. a conclusion can be drawn from the experience of hkust library: far from sounding the death knell of librarianship, the internet is the best thing the library world has ever had. the internet needs the library and the library needs the internet. • consummating indexing the traditional way to represent information documents in large collections is by indexing. each document is assigned one or more index terms selected to represent the best meaning of the document. these index terms are then searched to locate documents related to queries expressed in words taken from the index language. indexing is the oldest technique for identifying the contents of documents to assist in their retrieval. the indexing process is typically performed by professional indexers associated with library organizations. the objective of indexing has changed with the evolution of information retrieval [kowalski,1997]. now let us look at the environment for indexing which exists today. the explosion in the availability of information through computers, which allows people to explore databases in other places by systems such as the internet and world wide web has created new opportunities in information organization and retrieval. indexing is the ancestor of such activities and professional indexers or librarians must be flexible enough to adapt their skills in indexing to information management and retrieval. indexes need to convey more intelligence than the contents of text, in the sense that it should be able to structure or indicate every possible route a user might take into the text. discussions and writings about searching the internet frequently mention the difficulty of finding what is wanted and the need for good indexing. a prepared index to a site with a specific depth, should offer a thesaurus of controlled entry terms closely related to the outline and bring together material from disparate sources. indexing the internet is an exciting and challenging project. it is exciting to develop approaches to search the growing amount of online materials. there remain key issues on indexing for librarians today: how will indexing techniques have to change to stay relevant? librarians should be finding out how they can participate in and contribute to information access and the internet. people demand better result, which is partly a feature of the search facility and partly of indexing. librarians should grasp the new opportunities coming with the internet. in short, librarians have the expertise and skills in content analysis to contribute to the control of information on the internet [macdougall, 1996] educator and trainer there are many forms of user education from library tour to bibliographic instruction. many of those programs are now working with or evolving into information literacy programs with emphasis on the internet instruction [martin, 1997]. information literacy education literacy, beyond embracing the basic abilities of reading and writing, now embodies the general ability to understand and perform functions successfully. the term is often paired with areas such as media, computers, culture and information. the goal of information literacy is to ensure that people understand how to, and why they need to, learn about sources in the information society. some of these sources will be in the library, others will be in the world at large. the definitions of information literacy varies slightly from source to source, though the focus is helping users gain a broad understanding of information sources and enhancing their ability to deal with that information. the american library association gives this definition: to be information literate, an individual must recognizes when information is needed and have the ability to locate, spring 1998 21 evaluate and use effectively the information needed [american library association, 1989]. there is a growing recognition of the need to train users in information literacy skills. libraries should take on the role of imparting information literacy skills. for example, the digital information literacy program at the university of texas at austin is dedicated to promoting electronic resources to the library users and encourages users to examine the internet information [martin, 1997]. three steps in user education end user training is an evolving area of research. learning to use a library was once a fairly simple activity for the user. library searching was relatively straight forward, and the end-product easily retrievable. users are becoming more independent and effective, search engines in both print and electronic formats are coming into being. access to information is not enough and librarians will be encouraged to become editors and create filters to help users select what they want. • identify what librarians need to impart to users the mission of user education aims to help users at all levels learn how to identify the information they need from the morass of information in electronic and other forms. • recognize that learning to search is a progressive process users will have rudimentary skills in their first weeks of training. what should be kept in mind is that they will need help every later time they use a library, search the internet or databases. • have an important part in playing extension support to users people go in for do-it-yourself activities. that is true with information seeking. librarians may prepare users to deal with the complexity of information environment and encourage users to search what they need on their own given the availability of technological support. knowledge manager we are embarrassed with the problem of information overload. longing for information gives way to knowledge seeking. knowledge related to specific problem solving is crucial. knowledge, not merely information, is the major competitive factor in life. knowledge assets become the key assets (with its emphasis on concepts such as intellectual capital, intangible assets, intellectual property, etc.). the shift from distributing information to managing knowledge is becoming an independent production factor next to labor, capital and natural resources. knowledge is evolving into intellectual assets on which business organizations around the world are dependent for their survival [bonaventura, 1997]. tacit and explicit knowledge are the two basic forms in which knowledge can be operative in an organization. tacit knowledge resides in people’s heads. explicit knowledge stores in books, journals, cd-roms, the internet, etc. explicit knowledge is formal knowledge that can be packaged as information and can be found in the documents of an organization: reports, articles, manuals, patents, pictures, images, video, sound, software, etc. tacit knowledge is personal knowledge embedded in individual experience and is shared and exchanged through direct, eyeto-eye contact. tacit knowledge is practical knowledge that is key to getting things done, but has been sadly neglected in the past [borghoff and pareschi, 1997]. many factors contribute to the rise of knowledge management. they can be grouped into the following areas: increasing popularity of learning organization; growing importance of knowledge; technological availability; the transition of economy; and growing interests in knowledge management research. library practice has been updating with the societal advances. from book warehouse to information center, the principle that library is an ever-growing organism is embodied. what can librarians do to realign their focus from the old world of “information management” to the new paradigm of “knowledge management”? librarians have excellent skills in organizing and codifying information sources and making these accessible to others. this represents the top layer of the knowledge map— information—rather than tacit and explicit knowledge. librarians are involved in a continuing search for excellence in organizing and codifying information sources, networking, etc. all these activities are important for knowledge management, but not sufficient [broadbent, 1997]. new knowledge often begins with the personal. the fact that a reference librarian knows something of why services aren’t utilized the way the organization desires isn’t of itself organizational knowledge. it becomes organizational knowledge when there are management processes in place which capture that often personal, tacit, front-line information from which others in the organization learn and make decisions. that represents a quantum shift for most organizations to a focus on using human expertise for business advantage. so, knowledge management is a form of expertise-centred management. it is characterized by variety and exception rather than routine and is performed by professionals or technicians with a high level of skill and expertise. knowledge management is about the 22 iassist quarterly acquisition, creation, packaging and application or reuse of knowledge [broadbent, 1997]. we will see in 1998 how much further management experts can go along the road that leads towards knowledge management, that is direction that librarians should be looking if we want to be abreast of the next century. how to make the transition from “information management” to “knowledge management”? as an information manager or a librarian, we are best placed to be the driver of knowledge management within our organisations [bonaventura, 1997]. there is no common-held model for knowledge management. knowledge practitioners are responsible for accumulating and generating both tacit and explicit knowledge. library information service processes can be viewed from the perspective of knowledge management. if you are an information and technology professional, you may look at it from information technology perspective; if you are a business manager, you may look at it from business perspective. knowledge management means different things to different people. knowledge management isn’t only for senior management, librarians at junior or middle management levels are having to deal with the complex task of knowledge management [kinnell, 1996]. library managers: knowledge coordinator the role of the chief librarian in a library is that of a designer, teacher, and steward who can build a shared vision and challenge prevailing mental models. the chief librarian needs to be responsible for knowledge coordination. he or she needs to have an understanding of developing the human and cultural infrastructure within a library which will facilitate information sharing, specifically the conversion of tacit knowledge of librarians into explicit knowledge that may be shared across the library. the person should have the combined capabilities of a business strategist, technology analyst, and a human resource professional. a consulting background, particularly a background involving similar roles of liaison and consultation, would be desirable. communication skills are also crucial. if a library is managed in this way, the tacit knowledge of librarians can be made best use of , the library service strategy can be most appropriate, the benefits of a library can be most accomplished. the major tasks of a knowledge coordinator fall into several areas [cronin and davenport, 1988]: • identify users’ needs faced with an exponential increase in the amount of published information in a variety of forms, the knowledge coordinator will have to identify and respond to the particular needs of local clientele and to structure , market and deliver services and knowledge to these users. • emphasize access to knowledge academic and theoretic knowledge is fundamental to teaching and researching at a university, and priorities should be given to such areas. • exploit new technology and local networks the availability of personal computers and campus networks presents both an opportunity and a challenge as faculty, students and research staff become more information literate and their needs more sophisticated. • link library program to academic program with limited resources and ever growing universe of information, there must be a close mapping between library collections and services and the educational and research priorities of the university. • market library service a demanding area for academic libraries is the need to acquire marketing and public relations skills. it is no use developing library service without being publicized and marketed. professional skills coupled with business acumen are important part of academic library operation. librarians: knowledge creator from the library’s perspective, knowledge creation implies participating more in users’ reading and studying by identifying users’ information needs and developing appropriate service and products to meet the needs. while a user is seeking knowledge, explicit knowledge stored in documents either in print or electronic format is needed to be implanted in the user’s head and to be analysized and synthesized with the user’s tacit knowledge and/or his previous experience, etc. that is the process of internalization of knowledge. in the meantime, new ideas, concepts and methods can be generated by the user. these new ideas, concepts and methods are the new contribution to human knowledge and need to be formalized and organized for transmission and storage. that is the externalization of knowledge. during this knowledge transfer process, a librarian acts as a node in internalization and organization of knowledge. only by participating in the teaching and research activities, can academic librarians become part of knowledge creation. (see figure 1) it can be inferred that knowledge management call for usercentred library service. briefly stated, a librarian , analogous to the medical professional model, must be able to diagnose needs, prescribe a service or remedy, treat or design an appropriate remedy, and evaluate the treatment or remedy [glazier and powell, 1992]. librarians need to lay the foundation to serve the users of tomorrow. librarians need to plug into the information super highway and spring 1998 23 exploit the relevant technology to make knowledge available to users in a fast and reliable manner. librarians need to move towards a borderless information environment and remove all boundaries and barriers to knowledge transfer and share. knowledge management will be playing a key factor in the future of libraries. references 1. american library association presidential committee on information literacy. final report. chicago: american library association, 1989. 2. bonaventura, m. “the benefits of a knowledge culture”. aslib proceedings 49, 4(1997), pp 82-89. 3. borghoff, u.m. pareschi, r. “technology for knowledge management” journal of universal computer science information 3, 8(1997), pp 835-1021. 4. broadbent, m. “the emerging phenomenon of knowledge management” the australian library journal 46, 1(february 1997), pp 7-23. 5. cronin, b. davenport, e. post-professionalism: transforming the information heartland. london: taylor graham, 1988. 6. glazier, j. d. powell, r. r. qualitative research in information management. englewood: libraries unlimited, inc. 1992. 7. hardy, g. “the web and libraries: the victorian experience” , pp 92-96 in reading the future: proceedings of the biennial conference of the australian library and information association 1996, ed. by alia. canberra, act: alia, 1996. 8. kinnell, m. “management development for information professionals” aslib proceedings 48, 9(1996), pp 209-214. 9. kowalski, g. information retrieval systems: theory and implementation. norwell: kluwer academic publishers, 1997. 10. macdougall, s. “ rethinking indexing: the impact of the internet” australian library journal 45, 4(1996), pp 281-285. 11. martin, l. the challenge of internet literacy: the instruction-web convergence. new york: the haworth press, inc., 1997. 12. mclean, n. “reading the future: access and technology” , pp 13-20 in reading the future: proceedings of the biennial conference of the australian library and information association 1996, ed. by alia. canberra, act: alia, 1996. 13. nicholson, f. rochester, m. “reading the management future for libraries: implications for library management education”, pp 75-84 in reading the future: proceedings of the biennial conference of the australian library and information association 1996, ed. by alia. canberra, act: alia, 1996. 14. yip, k. f. “selecting internet resources: experience at hong kong university of science and technology library” electronic library 15, 2(1997), pp 91. lecturer, department of information resources management, nankai university. visiting scholar, department of information studies, university of technology, sydney. vol29-1.indd 26 iassist quarterly spring 2005 the international association for social science information service and technology (iassist) invites your participation in its 32nd annual conference entitled data in a world of networked knowledge on may 22-26, 2006 in ann arbor, michigan. the conference will be preceded by workshops and followed by optional weekend activities in the ann arbor area. details about the conference and the association may be found on the iassist website at: <http://www.iassistdata.org> proposals for papers, sessions and poster/demonstrations should be submitted by 16th january 2006. the 2006 conference theme, data in a world of networked knowledge, highlights the role of empirical data in a society that wishes not only to know itself, but also to build an enduring, interconnected storehouse of knowledge for learning and research. once again iassist offers a time and place to explore, enlighten, and energize the participation of data professionals in the networked information world. we seek submissions of papers, poster/ demonstration sessions, and panel sessions on topics that address the full range of digital data life cycle issues, including those that focus on access, documentation, dissemination, preservation, data use and current empirical research activity. additional topics might also include information and statistical literacy, data confi dentiality and statistical disclosure, geographic information systems (gis) and spatial data, as well as publication, annotation, curation and authentication of networked knowledge assets. for other key topics see previous iassist conferences at <http://www.iassistdata.org/conferences/index.html>. about iassist iassist is an international organization of professionals working in and with information technology and data services to support research and teaching in the social sciences. the organization also explores issues of access, stewardship and the interconnections among social science, behavioral, biological, and health data. typical workplaces include quantitative and qualitative data archives/libraries, statistical agencies, research centers, libraries, academic departments, government departments, and non-profi t organizations. see the iassist website at <http://www.iassistdata.org> for further information. iassist conferences bring together data professionals, data producers, and data analysts from around the world for presentations and workshops covering new and persistent issues relating to access to data, its documentation, and digital preservation, with special emphasis on the social sciences. the social sciences have a long history of data sharing activity which will make the conference of interest to colleagues in disciplines where improving data access practices is on the policy agenda, and where there are clear overlaps with digital curation, data publishing, e-science/ cyberinfrastructure initiatives, and new interdisciplinary collaborations. the iassist quarterly (iq), available online from the iassist website and in print, is another important means of communication for the data community. each year, iq features the papers associated with conference presentations. of special note is the iassist publication award, involving a cash prize for the winning paper. for further details see: <http://www.iassistdata.org/publications/pubaward.html >. the iassist outreach committee accepts applications from data professionals in countries with emerging economies for funding to attend iassist 2006. more information about the outreach committee’s work, including funding criteria and online application form, can be found at <http://www.iassist.ucdavis.edu/> procedure iassist 2006 call for papers iassist quarterly spring 2005 27 iassist 2006 the deadline for paper, session, and poster/demonstration proposals is 16th january 2006. the conference program committee will send notifi cation of the acceptance of proposals on or before 10th february 2006. individual presentation proposals and session proposals are welcome. proposals for complete sessions, typically a panel of three to four presentations within a 90-minute session, should provide information on the focus of the session, the organizer or moderator, and possible participants. the session organizer or moderator will be responsible for securing session participants, some of whom may submit paper proposals independently. all proposals, including proposed title and an abstract (recommended length 150 words), should be submitted using the link on the following website: <http://www.icpsr.umich.edu/iassist/call.html> alternatively, proposals may be sent via email to <iassist06@gmail.com>. in this case, please use a subject heading of “paper proposal your name” or “session proposal your name” replace “your name” with the name of the session organizer. further information on travel and accommodations will be available at links from the iassist ‘06 conference website: <http://www.icpsr.umich.edu/iassist/>. online registration is scheduled to open on 1st february 2006. make plans to come to ann arbor for the iassist 2006 conference may 22-26, 2006 28 iassist quarterly spring 2005 submissions due january 10, 2006 announcement of winner march 1, 2006 award: $250 us and one year membership in iassist http://www.iassistdata.org content focus: education. as an organization, iassist has a history of working to educate its members about matters of common professional interest to the social science data community. traditionally, this education has taken the form of professional development opportunities available in member-initiated and member-taught workshops at the annual iassist conference. yet as the world of social science data grows increasingly complex, staying abreast of new developments in the profession is likely to present an ever increasing challenge for iassist members. as a result, the education committee of iassist is receptive to recommendations for employing new instructional methods and technologies as we strive to meet iassist’s educational mission. at its annual conference in may, 2004, the iassist membership approved a 5-year strategic plan that focuses upon three strategic directions: education, outreach, and advocacy. this paper competition has been established as a means of exploring, articulating, and documenting topical issues related to education. for further information about the strategic plan, see: http://www.iassistdata.org/membership/plan_june2004.pdf iassist seeks papers that address one or more of the issues, principles, and strategies for engagement in the following subject areas: 1. iassist-related educational initiatives 2. professional development and educational opportunities of interest to iassist members 3. educational outreach to research communities, including and beyond the social sciences, to promote data preservation and access. papers should include recommendations for action or suggestions of specifi c projects that can be undertaken (or are underway) to further the education goals of iassist. prospective submitters may wish to review the discussion on the iassist blog about the educational issues arising under the topic of the accidental data librarian (available at http://iassistblog.org/?cat=5). criteria for evaluation: call for papers: 2006 iassist strategic plan publication award (competition is not limited to current iassist members) iassist quarterly spring 2005 29 2006 iassist strategic plan publication award --relevance of the paper to one or all of the themes in the iassist strategic plan (specifi cally strategic direction i: improve and expand the educational component of iassist both internally and externally. --inclusion of specifi c suggestions for action or specifi c projects that will further the education of iassist members or otherwise encourage progress in support of the iassist strategic plan --potential in building a base for future iassist activity --quality of writing -bibliographic content including references to related materials --clarity in presenting issues and viewpoints as outlined above competition details: all papers are to be submitted in english. the winning paper will be announced on the iassist list-serve on or around march 1, 2006. in addition to being designated as the winning paper in the iassist quarterly, the author of the winning paper will receive the monetary award and a one-year membership in iassist and be recognized at the iassist conference in ann arbor. all other submissions meeting the criteria for evaluation will be published in the iassist quarterly (iq) (online and print). papers that are submitted to this strategic plan publication award competition may also be submitted for inclusion in the iassist conference in ann arbor in may 2006. papers must be a minimum of 5 pages in length, including bibliography and graphics as appropriate. we strongly prefer that for publication purposes all documents be submitted in word format and each graphic be submitted as a separate fi le in one of the following formats: .gif .jpg .tif .bmp .png papers must not have copyright limitations; iassist quarterly (iq) rights will apply upon publication. all papers must be submitted by january 10, 2006 to the competition web host: david sheaves <sheaves@vance.irss.unc.edu> questions (not papers, please) may be sent to: <iassist-reviews@mailman.srv.ualberta.ca> competition is not limited to current iassist members. the review committee for the iassist strategic plan publication award will be announced on the iassist website www.iassistdata.org and on the iassist list serve. vol25s.1 iassist quarterly spring 2001 21 research for building a better data community i have come to believe that iassist members must seriously consider the value of conducting research about our own profession and field. if we do not initiate and value such research, the expectation that anyone else will undertake this task for us is unrealistic. i am not necessarily referring to “theorydriven” research, although this would be welcomed. rather, i have in mind “issue-driven” research, that is, the kind of research that helps us understand the relationships, norms, and behaviours within our information and science cultures, including our own data subculture. for example, we need to conduct research on the issues behind data preservation and access. i am not talking about “how” to preserve data or provide access but instead investigating the norms and behaviours underlying the activities of data preservation and access. a few recent events have led me to this new conviction. first, i was asked a couple of years ago by a major research council to review a grant application in which the principal investigators were proposing to study the economics of archiving data. in my enthusiastic endorsement of this application, i wrote that the principal investigators should expand their scope to explore the “data economy”, which i characterized as who gets access to which data, when and how. i thought that this would be a wonderful project for our profession. we would have research conducted about data archives and their role in the data economy. here data archives would be the object of research rather than the sources of data for research. unfortunately, the project was not funded. nevertheless, it did stimulate my thinking about the possibilities of such research. the next recent event took place at the 2000 iassist conference. i was excited by the research carried out by karsten boye rasmussen and repke de vries about iassist as a virtual community. their use of the iassist e-mail discussion list and the log files for the web site and on-line issues of the iassist quarterly clearly demonstrated ways of doing research about our organization and profession. a third and even more recent experience has been the research that i have been conducting in conjunction with the consultation underway in canada about creating a national data archive. before talking about specific research findings, let me briefly describe through an example one way in which “issuedriven” research might be performed in our field. one challenge we face, which has not changed over the last forty years, is how to get researchers to think about archiving their data at the beginning of a project rather than at the conclusion of their research. and more than just thinking about archiving data earlier, how do we get researchers to conduct their projects so that their data products meet archival standards – rather than having to build an archival data product after the research has concluded. in other words, how do we mainstream data archiving in the research process? this is neither a new nor novel idea. however, is it an idea whose time has arrived? if our profession better understands the dynamics of current research practices that inhibit data archiving, can we be instrumental in bringing about the necessary changes to mainstream data archiving? i am currently a co-investigator on a nationally funded project studying research utilization in the field of nursing. specifically, we are studying how practitioners eventually use medical research findings about pain and pain control. how do research outcomes end up in practice? i introduced the principal investigator of this project to the value of data sharing and stewardship when she was a graduate student several years ago.1 i also helped her preserve the research data from her dissertation. subsequently, she has become a successful researcher who is a key proponent of data preservation and data sharing in nursing research. this is how i came to be invited as a member of her team. out of this project, i was the lead author of a paper presented at the 2000 conference on social science methodology in cologne, germany entitled, “archivist on board: contributions to the research team”.2 this paper presented the role and value of an archivist on the research team. while i was not at the conference to present the paper, a colleague and co-author read it on my behalf. ekkehard mochmann was present at this session and came to the rescue of my colleague when one researcher, after hearing the presentation, protested that we were trying to turn researchers into archivists. my colleague reported that ekkehard explained to the researcher that what we were proposing was to initiate partnerships between researchers by charles k. humphrey* 22 iassist quarterly spring 2001 and archivists. when first hearing of this exchange, i was struck by the immediate reaction of the researcher. what were the attitudes and values behind such a response? why were the attitudes of this researcher so seemingly different from those of the principal investigator with whom i work? turning to some findings from the research we are conducting in conjunction with the national data archive consultation, some insights into attitudinal differences underlying the principles of archiving data can be found. one of the four surveys that were administered was a sample of researchers who received a grant from the canadian social sciences and humanities research council between 1998 and 2000. one objective of this survey was to identify the number of researchers who produce data products as part of their research and to determine how many researchers have ever archived or intend to archive their data. a second objective was to investigate researchers’ attitudes about data sharing and archiving. eleven items were used to gauge these attitudes. these questions touch upon the legitimacy of secondary analysis as a research method, on the value of data as a by-product of research, on the issues of data ownership and data sharing, on research council funding to prepare data for sharing, and on the impact that ethics review boards have on data sharing (see list 1). because the wording of five items (5a, 5d, 5e, 5h, 5j) does not support the principles of data sharing or archiving, the response categories for these items were recoded to correspond with the direction that supports the archiving principle.3 figure 1 shows the combined percentage of respondents ‘agreeing’ or ‘strongly agreeing’ on each item. the items in this figure have been arranged in decreasing order of support. eighty-one percent agree that data should be a valued byproduct of research, while only 21 percent agree that data do not belong to the researcher as her or his intellectual property. this decreasing order of agreement represents an increasing difficulty in support of sharing and preserving research data. six steps of itemdifficulty can be seen in this figure. eighty-one and 78 percent of the respondents accept the first two items, data as a valued by-product and secondary analysis as a valid research method, respectively. the second step consists of the items about research councils covering the costs to prepare data for sharing and figure 1 attitudes underlying support for data archiving data valued by-product (n=114) secondary analysis (n=115) research councils fund (n=115) researchers as trustees (n=113) data should be shared (n=108) educate ethics boards (n=110) waste of funds to save (n=113) archiving is integral (n=112) ethics make a barrier (n=109) share data regardless (n=115) data don’t belong to pi (n=112) percent agreeing or strongly agreeing 100806040200 21 27 28 48 50 62 64 68 71 78 81 list 1 attitudinal items underlying support for data archiving 5a. secondary data analysis is not a valid research method. 5b. data should be considered a valued by-product of .. research. 5c. data should be shared with other researchers, assuming it has been appropriately anonymized. 5d. data belong to the principal investigator as her or his intellectual property. 5e. data should only be shared if the principal investigator decides to share it. 5f. archiving data should be an integral part of conducting research. 5g. researchers who obtain information that cannot be easily reproduced from respondents are, to a degree, trustees of the data. 5h.spending resources to prepare the data from my research so that other researchers can use it would be a waste. 5i. research councils should include funds to cover the costs of preparing data for sharing. 5j. ethics review boards make it impossible to share confidential data on human subjects. 5k. ethics review boards need to be educated about the need to preserve data. iassist quarterly spring 2001 23 about researchers serving as trustees of data that cannot be easily reproduced from respondents. seventy-one and 68 percent endorsed these items, respectively. the third step is made up of the item stating that data should be shared if it has been properly anonymized (64 percent) and the item that ethics review boards need to be educated about the need to preserve data (62 percent). a slightly larger step occurs with the next two items. fifty percent agree that spending resources to prepare the data from their research would not be a waste and 48 percent agree with the statement that archiving data should be an integral part of conducting research. an even larger drop occurs with the fifth step. twenty-eight percent disagree that ethics review boards make it impossible to share confidential data on human subjects, while 27 percent disagree that data should only be shared if the principal investigator decides to share it. as mentioned above, the smallest percentage of agreement (21 percent) is that data do not belong to the principal investigator as her or his intellectual property. a scale was constructed based on the total number of items on which each respondent agreed with the principles of data archiving (see figure 2). a score of zero indicates that the respondent did not support any of the items endorsing data archiving, whereas a score of 11 represents someone who supported all of the items. seventeen percent were low supporters of data archiving (those with scores from zero to three), while 24 percent were high supporters (those with scores from eight to 11). fifty-nine percent are in the middle. the correlation between this scale and a question asking how important it is for canada to establish national services for the preservation of research data is 0.50, which is corroborative evidence that this scale measures some aspect of support for data archiving. what do we make of these findings in light of the question asked earlier about how to mainstream data archiving in the research process? first, only around a quarter of canadian researchers in this study appear to be strong advocates of archiving data. while only 17 percent seem to be protagonists, close to 60 percent are in the ambivalent middle. apparently, this rather substantial group requires further education on the principles of archiving. secondly, one of the eleven items rather succinctly summarizes the notion of mainstreaming data archiving. this is the item that states, “archiving data should be an integral part of conducting research.” looking at the results of this item, 12 percent agreed strongly, 36 percent agreed, 30 percent were unsure, 18 percent disagreed, and 4 percent disagreed strongly. the 30 percent who are unsure is as alarming in this distribution as the 22 percent who disagree. one concern raised by this finding is that at least half of the researchers do not view data archiving as part of the normal practices of conducting research. the concept of archiving may be generally understood, but archiving as part of the research process has not become routine. an explanation for these results may be directed at incomplete training of researchers in their graduate school years or at senior researchers who are not mentoring junior researchers about data archiving. another possible explanation is the failure of our profession to promote archiving as part of the research process. if the importance of the practice is not taught as part of the research method, data archiving will not be discussed or perceived as a generally important activity. turning away from this specific example, i would like to conclude with a couple of observations about changing our thoughts in iassist toward research. first, our organizafigure 2 distribution of respondents on the scale constructed from the 11 items 11 10 9 8 7 6 5 4 3 2 1 0 percent 2520151050 3 2 8 4 9 20 13 19 13 6 4 24 iassist quarterly spring 2001 tion is well positioned to conduct comparative, crossnational research. we are an international organization with oportunities to investigate the generalizability of national findings. are the attitudes described above held only by canadian researchers, or are these attitudes commonly found among researchers across societies? will we discover underlying attitudes about archiving data that are held by researchers around the globe? these are challenges that we can undertake together. secondly, we live in a world increasingly calling for evidence-based decision-making. we should approach the issues we face in our profession by building evidence through research. not only will we be stewards of data, but we will be contributing to the knowledge about the research process. footnotes 1 while completing her doctorate, she and a fellow graduate student published the following article about data sharing: estabrooks, c.a. and romyn, d.m. (1995). data sharing in nursing research: advantages and challenges. in canadian journal of nursing research, 27(1):77-88. 2 humphrey, c.k., estabrooks, c.a., norris, j.r., smith, j.e., hesketh, k.l. (2000). archivist on board: contributions to the research team. in forum qualitative sozialforschung / forum: qualitative social research [online journal], 1(3). 3 “strongly agree” and “agree” are the responses supportive of these principles. * paper presented at the iassist/ifdo conference 2001, amsterdam charles k. humphrey may 2001. iassist quarterly 15 educating the data user: classroom support services by bobbie pollard' baruch college, cik universiu' of new york "you people are really in the business of selling information" was a recent comment from a faculty member at baruch college after a presentation to his class of 98 students. we certainly are in the business of marketing information. indeed, our classroom support services program is an active and proactive one. during the past academic year we conducted nearly 200 library research workshops. over 6000 students were spoken to directly and many more were reached through printed literature such as resources for research and access guides. currently, twenty-two departments in the college use our services, including: marketing, management, business communications, marketing, education, speech, and english, to name only a few. the goals of the classroom support services are to inform assist students in their search for information for assignments in the various classes, and, equally important, to provide them with an imderstanding of how information is organized in the various disciplines. these information seeking skills help students develop confidence in the use of bibliographic and quantitative sources, not only for their college assignments but throughout their their later careers. 'presented at the international association for social science information service and technology (iassist) conference held in washington, d.c., may 26-29, 1988 the library research workshops are assignment driven and given at the request of faculty. therefore the success of the program depends on how well faculty are informed about the importance of students learning how to do research. we advertise our program to faculty through a flyer and by attending departmental meetings. a lot of our publicity is done by word of mouth; a faculty member who is pleased with the program tells his colleague. most of the requests come from otir undergraduate faculty. we only do library research workshops for classes with research assignments (sometimes these assigimients are made up by the librarians and faculty member together), and we require faculty members to be present at the workshops. summer 1988 16 iassist quarterly the content of a typical library research workshop includes the following: (1) the importance of information and an outline of a research strategy with examples of tools appropriate to the subject, discipline, or topics that the students are researching; and (2) practice in the library or online classroom using the materials and techniques discussed in the workshop. practice is essential because it is effective in clearing confusion and misunderstandings about the information presented in the workshop. most workshops are seventy-five minutes in length. the majority of the subjects taught at baruch require students to use some type of public data. the basic bibliographies and keys to finding information which are tailored to the content of the assignments and/or the course are called resources for research. most of the resources for research include government docments or an access tool that refers to government documents as sources of information. for example, the most heavily used resources for research . "the basic research strategy", lists the public affairs information service bulletin (pais) which indexes several government documents. the one entitled "company and industry" includes many other public data sources as well. others that include listings of public data sources are: "international marketing," "statistics", "marketing", "education", "international business", and "business journalism". public data are presented to students as being plentiful, reasonable, and generally easily accessible. because baruch college is one of the largest business schools, it is not surprising that our students need information on companies, industries, marketing, and statistics. even the english classes tend to do research on social science issues which required access to government data. in the majority of workshops some information on government data is presented, but it is a crucial source of information in the following classes; business communications, marketing, international business, and business journalism. below is a sampling of the kinds of information sought most by students in these classes, with some representative public data sources about which students are informed. the usual assignment in the business communications and marketing classes requires students to research companies, industries, products, and demographic data. they need information on sales, maketing, a financial profile, market share information, etc. in these classes, students are introduced to a variety of public data sources such as annual reports of publicly held companies, and u. s. department of commerce publications such as the u. s. industrial outlook . for statistical information on both products and industries students are advised to begin with the statistical abstract which serves as a summary and guide to most federal statistics including census data. predicasts f & s index of corporations and industries . pais, and other indices such the business periodicals index identify governmental data in periodicals. for information in books, students are shown how to look up information by subject, e.g. u. s. industries, or by governmental agency, e.g. u. s. department of commerce. in some market research classes, students are required to develop questionnaires. the instructors suggest that they look at many different types of questionnaires. the inter-university consortium for political and soda! research (icpsr) codebooks such as the qualitv of american life are very helpful to students in completing these assignments. in international business and international marketing classes, students study conditions for business in foreign countries and the import and export trade. for assignments in these classes, students are introduced to the numerous publications produced by the u. s. department of state and u. s. deparmenl of summer 1988 iassist quarterly — 17 commerce. background notes . overseas business reports. marketing in .. . and foreign economic trends and their implications for the united states are but a few of the public data sources taught in these classes. the business journalism class produces a national periodical called dollars and sense . the students research and write articles using mostly primarysources. therefore, the>need general backgroimd information, names of experts to interview, and a great deal of statistical data to back up their theses. the topics in the latest issue of this magazine, aids, women entrepreneurs, illiteracy in the workplace. west indian businesses, are representative of the articles that students write for this publication. public data sources on national and local levels are used heavily. students are introduced to the following indices which are excellent for identifying statistical information to support a point of view or to document a trend. the indices are the american statistical index which identifies statistical information in over 400 federal governmental agencies, and statistical reference index , an excellent source of statistical data published by local governments. for example. business statistics by the new york state department of commerc conttained important statistical data on west indian businesses in new york. students are also made familiar with many other government publications through the use of the monthly catalog . cis index , and a multitude of other government directories.n summer 1988 6 iassist quarterly spring summer 2009 guest editors’ notes welcome to a special double issue of the iassist quarterly featuring articles focused on the data documentation initiative (ddi), a metadata standard for the social sciences. we are proud to present these six articles, which explore various projects related to ddi 3 and its enhanced features. the articles draw on previous presentations and papers created in connection with the 2009 “expert workshop on implementation of ddi3 -advanced topics” held in wadern, germany; the 2009 european ddi users group (eddi) meeting held in bonn, germany; and the iassist conferences held in tampere, finland (2009) and ithaca, new york, usa (2010). jeremy iverson’s article on metadata-driven survey design highlights the reuse of metadata starting at the very beginning of the research data life cycle and also discusses the benefits of using metadata to drive the process of collecting, visualizing, and analyzing survey data. this is a powerful and efficient approach that should be taught in survey methods courses in order to save costs and to enable data producers to leverage the metadata they create across the life course of research data. also related to data collection is the article on the questasy online survey documentation tool by marika de bruijne and alerk amin. questasy permits internal users to document longitudinal data and to make this documentation available to external users on the web. a benefit of ddi 3 for this system is that it facilitates tracking of question items across waves in the study, where each wave can have question constructs and variables that refer to the same question item. this system was developed for the liss panel online survey at the university of tilburg in the netherlands. “implementing ddi 3: the german microcensus case study” by andias wira-alam and oliver hopt looks at using ddi 3 to document the german microcensus through a customized ddi 3 editor and a web view providing different perspectives for the end users based on the same ddi 3 items. interestingly, andias and oliver discuss basing some of their decisions about software design on jannik and dan’s use case describing the development of the ddi 3 metadata authoring tool – see building a modular ddi 3 editor. “metadata creation, transformation and discovery for social science data management: the dames project infrastructure” by jesse m. blum, guy c. warner, simon b. jones, paul s. lambert, alison s. f. dawson, koon leai larry tan, and kenneth j. turner shows the wide variety of data management tasks that ddi 3 can support and document, including recodes, merging, and data cleaning. using ddi 3 to document these phases of the data life cycle is an exciting development. “ddi 3 development at dda” by jannik jensen and dan kristiansen of the danish data archive provides a fascinating look into the development of an authoring tool for ddi metadata, a tool that is being designed to play a central role in the work flow at the dda archive. it focuses as well on the underlying reusable middleware and general considerations on open source software development for ddi. this article provides the reader with an up-close view of strategic decisions made at dda to incorporate the functionality of ddi 3 into the architecture of the dda archive. with its focus on machine-actionability and data typing, ddi 3 needs a strong system of controlled vocabularies to supplement the creation of metadata. the article on controlled vocabularies by taina jääskeläinen, meinhard moschner, and joachim wackerow presents the case for using controlled vocabularies and the ways in which they benefit the user. the article also showcases the work of the ddi controlled vocabularies group and its efforts to create vocabularies for ddi 3, which will be made available as separate products using a format called genericode. we hope you enjoy reading these articles, and we offer our thanks to all of the authors. we also want to express our appreciation to iassist for the opportunity to publish this work in the iq. we are grateful for the ongoing support of the iassist community and its nurturance of ddi from the very beginning. sincerely, mary vardigan and joachim wackerow vol29-1.indd iassist quarterly spring 2005 by john adams, hasan almadfai, and ray thomas 1 the production and presentation of statistics of unemployment. comparability issues. abstract the united nations publishes unemployment statistics for 123 countries. most of these statistics are based on international labour office (ilo) criteria for the definition of unemployment. many countries also produce unemployment statistics based on insurance records and on the basis of registered unemployment. this paper aims to compare the main features of the different methods. the dimensions compared include the conceptual basis for the definition of unemployment, boundaries of employment and inactivity, entry statistics and duration of unemployment, use of denominators for production of unemployment rates, and the cultural influence of the statistics. the paper identifies conflicts between achieving international comparability and national needs. survey statistics that underpin international comparisons do not support geographically detailed analysis within countries. the value of unemployment statistics based on ilo criteria is limited by a failure to recognise the concept of entry to unemployment and the difficulties of integration with administrative unemployment statistics. the standard labour force survey (lfs) questionnaire should be modified to support the production of statistics for entrants to unemployment. the sampling frame should be modified to ensure consistency with nationally produced unemployment statistics derived from administrative records. introduction – three conceptual bases of unemployment statistics three types of systems – insured unemployment, registered unemployment, and unemployment measured by sample social surveys – provide the basis for most unemployment statistics. the first statistical series started in 1886 when the board of trade asked the trade unions to provide monthly statistics of the number of their members who were not in employment (garside, 1980). this series led to the idea of insured unemployment. the uk and many other countries have offices that help people find work or pay benefits to those without work. such systems provided the basis for statistics of registered unemployment. in the united states, concern about mass unemployment in the 1930s led the government to develop household surveys in order to measure the extent of unemployment (see anderson, 1988). the international labour office (ilo) gives details of registered unemployment systems for 75 countries (see http:// laborsta.ilo.org/). nearly all the countries of eastern and western europe have insurance and/or registered unemployment systems. the us uses insurance based statistics to help make unemployment estimates at sub-national levels (see section 5 below). but the system that has increasingly dominated in recent decades is the sample survey. since 1948 the monthly current population survey (cps) has been the dominant method of measuring unemployment in the us the main focus of the cps is employment and unemployment, and nowadays the cps would be described as a labour force survey. the cps defines unemployment in terms of the numbers seeking work. in the 1980s, when the time came for an international standard, the cps provided a model. the 13th international conference of labour statisticians in 1982 adopted the seeking-work criterion of the cps presumably because it could be applied in any country independently of any existing national systems for dealing with unemployment. ilo criteria for conduct of labour force surveys and the definition of unemployment (hussmans et al.,1990) count the numbers seeking work in almost exactly the same way as the cps. the ilo provides details of labour force surveys conducted in 109 countries. the standard labour force survey (lfs) does not use the word unemployment. the crucial question in the uk lfs, for example, is “thinking of the 4 weeks ending on sunday. were you looking for any kind of paid work at any time in those four weeks?” by avoiding the term unemployment, the ilo criteria aim to produce statistics that are independent of national insurance and other systems that give benefits to the unemployed and therefore use the term ‘unemployment’ in a variety of different contexts. but this pursuit of independence makes comparison with other datasets difficult or impossible. ilo unemployment statistics for the uk, for example, are not comparable to uk claimant unemployment statistics. 12 iassist quarterly spring 2005 the next section of the paper discusses the categorisation of unemployment – the boundaries between employment and unemployment, and between unemployment and inactivity. section 3 focuses on entry to unemployment – important because entry to unemployment is not recognised by labour force measures of unemployment, but is demonstrably important in developing policies that go beyond seeing exits from unemployment as the exclusive solution to unemployment problems. the choice of denominators is a key theme of sections 4, 5, and 6. sections 4 and 6 refer to problems with the conventional economically-active-population denominators for measurement of unemployment rates at the national and local levels, and the main population alternatives. section 5 identifies a flaw in the use of current level of unemployment as a denominator in the usual measure of long-term unemployment and introduces the idea of population at risk denominators as a superior alternative. section 6 discusses problems associated with measuring unemployment on a local scale and section 7 points to the increasing need for such measures. section 8 points out that the seeking-work criterion of the ilo definition of unemployment conditions users of unemployment statistics to view unemployment as a matter that belongs exclusively to the unemployed, although this runs against many cultural traditions. the section suggests modification of the ilo criteria for the definition of unemployment and modification of labour force survey sampling methods in ways that would add value to unemployment statistics at local and national levels without reducing international comparability. fixing the boundaries the definition of what is considered as employment is generous both in the cps and the standard lfs. the crucial question in the cps questionnaire is ‘last week, did you do any work for pay or profit?’ (http://www.bls. gov/cps/cps_htgm.htm). the lfs questionnaire in the uk asks first ‘did you do any paid work in the seven days ending sunday as an employee or as self-employed?’, and later ‘(in) the seven days ending sunday, how many hours did you actually work ..?’ these questions support the production of statistics of employment as defined by the ilo as paid work of one hour or more per week. the motivation for the generosity of this definition is the wish to link production to total labour input (hussmanns et al., 1990, p 71), or in other words, to produce statistics of labour productivity in terms of output per person-hour rather than just output per person. it is unlikely that this differentiation is understood or accepted by most survey respondents. the lfs in the uk, like the cps, establishes employment with questions that elicit the amount of paid work of more than one hour. but in the 2001 uk census of population respondents were asked if they were in employment. the resulting census statistics gave an employment level of 640 thousand, or 2.5%, below that of the corresponding lfs estimate, and an unemployment level of 204 thousand, or 14%, above the lfs figure (heap, 2005). it seems unlikely that the line between employment and unemployment is drawn at this one hour boundary in most systems of insured or registered unemployment statistics. in the uk the rules specify that claimants for unemployment benefits cannot work for more than 16 hours per week. it can be expected that each insurance or registered unemployment system will have individual regulations on the matter. a similar variety of regulations can be expected to apply at the other boundary. ilo criteria specify that respondents must be seeking work within the reference period. the search period supported by the organisation for economic co-operation and development (oecd) is four weeks. in the uk both lfs and the census ask about two weeks. surveys in japan and taiwan ask only about one week. those expecting to take up a specific job can also be classified as unemployed. the cps in the us includes as unemployed those laid off from employment who are expecting to resume their former work. ilo criteria specify that respondents must be available to take up work within two weeks in order to be classified as unemployed. but the conditions for eligibility for job seekers allowance (jsa), the name given to claimant unemployment in the uk, is tougher. a number of welfare groups describe jsa as being “about hassling people off the dole into low paid work by making it tougher to sign on” (for example, urban 75, undated). being available for work to qualify for jsa means being ready to start permanent or temporary work immediately. the standard lfs questionnaire supports a major conceptual extension of unemployment by including a question ‘... even though you were not looking for work, ... would you like to have a regular paid job at the moment ...?’. this question in the uk lfs is addressed to respondents after they have been identified as economically inactive by questions that have established that they are not classified as in employment nor as unemployed. the question is also part of the standard lfs in europe. statistics for the number wanting work but not classified as unemployed are published by eurostat. there is a follow-up question on the main reason for not seeking work. in the uk lfs the pre-coded answers are given in this order: student; waiting to take up a job; looking after family; sick; believes no jobs available. the interpretation of statistics resulting from this question may be influenced by the ordering of these pre-codes and the precise instructions given to interviewers. iassist quarterly spring 2005 13 that last pre-code reminds us that the ilo definition is tight in that it excludes respondents who are not looking for work because they believe that no suitable jobs are available. this clash between the logic of the individual and that of the ilo criteria is honoured by describing such respondents as discouraged workers. (hussmans et al., 1990, p 107-8). statistics for discouraged workers are published by the oecd (see http://www1.oecd.org/scripts/ cde/-queryscreen.asp?), but are rarely subjected to detailed appraisal. according to the oecd there were more than two million discouraged workers in japan in 2000. entry and short-duration unemployment statistics relating to the duration of unemployment commonly include figures for unemployment of up to four weeks. but such figures usually relate to uncompleted spells of unemployment. it does not include completed spells of unemployment of less than four weeks. the distribution by duration is subject to left-hand censoring or truncation. this limitation is avoidable. it is not, however, avoided by the ilo criteria for labour force surveys, nor by the cps questionnaire. questions on unemployment are addressed only to those who are unemployed at the time the survey is conducted. cps and labour force survey questionnaires establish the number of those who became unemployed during the previous four weeks, but do not include those who became unemployed during the previous four weeks but had exited from unemployment by the date of the survey. the sample is biased against being representative of the whole population of working age. newly unemployed who have re-entered employment or become economically inactive within four weeks are not included. kiefer (1998) described the resultant statistics for duration of unemployment as subject to ‘length-biased sampling’. systems that record entry to unemployment, such as the uk system of claimant unemployment, can be used to produce the numbers exiting before four weeks. the office for national statistics (ons) makes available such statistics extending back to 1983 through the nomis database. chart 1 (see pg 14) indicates that monthly exits before four weeks are small relative to the total stock of unemployment but represent a substantial proportion of monthly entrants. monthly exits before four weeks show strong seasonal variation, but 12-month moving averages range between 20-30% of entrants. that range is small relative to the variation in levels of unemployment in this period. crosssection analysis also shows a relatively small variation in exits before four weeks. the coefficient of variation (cov) among 659 parliamentary constituency areas (pcas) in 2004 was only 15% compared with the cov for unemployment rates of 54% (calculations by the authors). the omission of unemployment of less than four weeks could be considered a matter of poor survey design. we could make an analogy with a hypothetical survey of incidence of the common cold. it is to be expected that a survey of the common cold would ask people when they last suffered from a cold in order to get information relevant to catching a cold. it is unlikely that the survey would be limited to those who had colds on the day the survey was conducted. it is unlikely that respondents would first be asked ‘do you feel healthy?’, and if they answered ‘yes’, discarded from the sample! but the standard labour force survey creates an analogous situation by addressing questions on unemployment only to those unemployed at the time of the survey. the omission of unemployment of less than four weeks does not affect statistics for unemployment of more than a month’s duration. but the oecd regularly publishes statistics for member countries for the percentage of unemployment of less than a month. the statistics are footnoted with the misleading comment ‘these percentages only take into account those persons for whom the duration of unemployment is known’. in fact, lfs duration statistics are based on uncompleted spells of unemployment – the difference between the date previously worked or the date started seeking and the date the survey was conducted. the only statistics of duration collected by the lfs relate to periods longer than the specified period. exit statistics are necessary to measure known duration and exit statistics by duration are not obtainable from lfs surveys. countries with systems of insured and registered unemployment can be expected to have records that support the production of statistics of entry to, and exit from, unemployment. the bureau of labor statistics (bls) in the us publishes weekly statistics for initial claims and continuing claims for insured unemployment. but the bls web site does not include any breakdown of continuing claims by duration. the bls also produces monthly statistics for mass layoffs. the mass layoff numbers come from establishments which have at least 50 initial claims during a 5-week period. extended mass layoff statistics, issued quarterly, relate to a subset of such establishments where employers indicate that 50 or more workers were separated from their jobs for at least 31 days. european countries seem to make little use of administrative data on entry to unemployment. some implications of the failure to recognise the concept of entry to unemployment can be illustrated with statistics for claimant unemployment in the uk. chart 2 (see pg 15) illustrates that statistics for the 659 uk parliamentary constituency areas (pcas) in 2004 show a 90% correlation between entry to unemployment and the unemployment rate. it could be said that the chart only demonstrates 14 iassist quarterly spring 2005 iassist quarterly spring 2005 15 16 iassist quarterly spring 2005 iassist quarterly spring 2005 17 the obvious – that the main cause of unemployment is becoming unemployed – but the scale of geographical variation is remarkable and notable. statistics of entry to unemployment have not been widely used in the uk. over the past decade government labour market policy in the uk has been focused almost exclusively on exits from unemployment. labour market policy in the uk has been dominated by programmes and slogans such as ‘new deal’ and ‘welfare to work’ as if unemployment were solely a matter of the unemployed making themselves employable. as a result, authorities know little about the causes of unemployment or the factors that are leading to growing inequality in the geographical distribution of unemployment. in light of the large variation between areas, it is not surprising that the emphasis on exits has been associated with an increase in inequality in the geographical distribution of unemployment (adams and thomas, 2005). characteristics of the unemployed and long-term unemployment labour force surveys can be expected to provide profile information on the unemployed on the same basis as that for the employed so that it is possible to make comparisons between the unemployed and employed population. the cps questionnaire also includes questions on the previous job, including: ‘did you lose or quit that job, or was it a temporary job that ended?’ these questions support the production of statistics on six alternative reasons for unemployment: temporary layoff; permanent job losers; completed temporary jobs; job leavers; re-entrants; and new entrants. but guidelines for the standard lfs do not include questions that would elicit this information. the information available from administrative systems can be expected to vary according to the nature of the system. the bls does not publish details on insured unemployment except for those on federal programmes – presumably because of variations between different state schemes. claimant unemployment in the uk included occupation until 2000. although the jsa system appears to require a lot more information from claimants, little gets through to the domain of published statistics. labour force surveys and administrative systems produce statistics for duration of unemployment. but there is a flaw in the way those statistics are usually presented – both by national and international organizations. the usual form of publication has been to express the numbers in duration groups as a percentage of total unemployment. the right hand cell in such tables is typically the numbers employed for a year or more as a percentage of all unemployment. that figure has become a standard measure of long-term unemployment. webster (1996 and 1997) called this measure long-term unemployed as a percentage of unemployment (lapu). lapu is a misleading measure – especially in time series analysis. numerator and denominator are incommensurate. the size of the denominator is determined by the number who became unemployed in the previous twelve monthswhich is not directly related to long-term employment. this problem is well recognised by the ilo (see http://www.ilo.org/public/english/employment/strat/ kilm/kilm10.htm) and has been succinctly described in a report of the royal statistical society: at a time of rising unemployment the number of shortterm unemployed will be increasing, and consequently, the percentage of long-term will be decreasing. this might mistakenly be read as an improving situation. conversely when unemployment is falling the percentage of long-term unemployment will increase if most of the slack is taken by those recently out of work. (working party, 1995, p 387379) use of a population at risk denominator provides a straightforward solution to this problem. in the case of year-or-more unemployment the population at risk (par) is the number unemployed a year earlier. the number unemployed a year earlier are all at risk of being unemployed a year later, and no-one not unemployed a year earlier is at risk of being unemployed a year later. but recognition of the problems with lapu has not prevented a generation of economists from seizing on lapu statistics to assert that the long-term unemployed have become insulated from the labour market. stephen nickell, a member of the bank of england monetary policy committee, writes: “long-term unemployed still form a substantial and important group … this has a significant macroeconomic impact because the long-term unemployed tend to lose skills and motivation as well as being discriminated against by employers. this weakens their attachment to the labour market... they become ineffective in holding down wage inflation and this leads to the impact of adverse shocks to the economy … (nickel, 1999, page 23). use of a population at risk (par) denominator reveals that year-or-more unemployment is actually more sensitive to changes in the state of the labour market than par rates for less than year unemployment groups (adams and thomas, 2004 and 2005). but stephen nickell was misled by lapu statistics. webster (forthcoming) found that lapu lags total unemployment by six quarters. chart 3 shows that lapu statistics lag the par rate by up to two years. chart 3 (see pg 16) demonstrates that the par rate for year-or-more unemployment moves parallel to the trend in unemployment for less than a year. the parallelism suggests that levels of less-than-year and year-or-more 18 iassist quarterly spring 2005 unemployment are influenced by the same set of factors. denominators for unemployment rates in 1886, when the board of trade in the uk asked trade unions for the number of their members who were unemployed they also asked for the total number of members. the number of members provided an obvious denominator to support unemployment rates that could be used to make comparisons over time and between industries. when the uk introduced compulsory unemployment insurance in 1911, the insured population provided an obvious denominator. nowadays the standard denominator for unemployment rates is the number in employment plus the number unemployed – usually described as the economically active population. employment, unemployment, and inactivity are usually thought of as three alternative labour market states. but the use of the economically active population as a denominator for unemployment is not consistent with the way employment, activity, and inactivity rates are usually measured. employment rates and activity rates are usually expressed as a percentage of the working age population or, of the population in the specific age group under consideration. use of a common population denominator would support direct comparison of unemployment rates with employment, activity, and inactivity rates. use of the number of trade union members as a denominator in the 19th century could be justified on the grounds that it could be assumed that the major flows over time between employment and unemployment were accounted for by trade union members. a recession could be expected to reduce employment and increase unemployment. the ratio of unemployment to employment could be expected to be an appropriately sensitive monitor of changes in labour market conditions over time. but it cannot be so easily assumed in the 21st century that the dominant flows between employment and unemployment are limited to the economically active population. flows between employment and inactivity, and between unemployment and inactivity, detract from the value of the unemployment rate as an economic indicator. not taking economic inactivity into account also limits the value of the conventional unemployment rate for comparisons between different areas, age groups, or social groups. it can be expected, for example, that the scale of unemployment is often correlated with the scale of economic inactivity (for example, beatty et al., 1997 and 2000). where this happens comparisons based on the conventional rate could be expected to systematically understate the differences in economic or social conditions in different regions. a systematic relationship does not preclude a lot of individual variation, and in making comparisons between two particular areas or two particular groups, the conventional unemployment rate can be misleading. there is, for example, wide variation in activity rates for women between different countries, especially in the older age groups. in many cases it would give a false picture to make comparisons of conventional unemployment rates without taking into account the differences in activity rates. the ilo acknowledges this problem in regard to youth unemployment where there is great variation in the scale of economic inactivity due to variation in the proportion classified as inactive because of training or full-time education and to cultural matters such as social expectations about women working outside the home. the ilo response has been to produce statistics entitled “youth unemployment, share of youth unemployed to youth population”, or in other words, unemployment rates with population denominators – in this case the population aged 15–24 years. unemployment rates with population denominators reveal substantial differences between countries. youth unemployment rates using the conventional economically active denominator are particularly high for a number of east european countries – bulgaria, poland, and slovakia. but the impression given by these statistics is mitigated by expressing unemployment using a population denominator. it might be assumed that a substantial proportion of the youth in these countries are economically inactive because they are investing in human capital by remaining in the educational system. one of the problems with measuring unemployment rates at a local scale is that statistics for the economically active population are not available; statistics for employment are usually produced only by place of employment. statistics for residents in employment in local areas are available only from censuses or surveys. the following section discusses the solution adopted for claimant unemployment in the uk since 2003 – to use population of working age (pwa) denominators. there do not appear to have been any disadvantages with unemployment rates measured in this way except for lack of comparability with the conventional rates used at national and regional levels. local unemployment statistics there is no contest at the local level between the quality of survey statistics on one side and insurance or registered unemployment statistics on the other. sample size limits the accuracy of survey statistics for local areas. but administrative statistics from insurance and registered unemployment systems can be produced on a 100% basis. the united states combines cps statistics with unemployment insurance statistics to produce unemployment rates for more local areas. the cps sample is 60,000. according to local area unemployment statistics (laus) as displayed on the bls website there are 31,792 series to query for. these include series iassist quarterly spring 2005 19 relating to states, counties, parts of cities divided by county boundaries, and minor civil divisions. the coverage and rules of insurance schemes vary between states. not all of those who become unemployed are eligible for insurance benefits. benefits do not usually extend beyond 26 weeks and the average duration is about 16 weeks. the insured unemployment rate is typically about one third of the cps rate. the bureau of labor statistics, with state authorities, makes estimates of labour force, employment, and unemployment on the basis of the cps, the current employment survey (ces) and the unemployed insurance statistics. the scale and sophistication of the estimation processes needed to produce estimates for local areas is formidable and impressive. the ces ‘place of work’ estimates are adjusted on the basis of commuting data to ‘place of residence’. separate estimates are implied for those who have come to the end of their period of insured unemployment, and for entrants and re-entrants to unemployment who are not covered by the insurance system. there are integral seasonal adjustment programs. the statistics are controlled to state totals. in the uk, statistics for claimant unemployment for local areas are publicly available in considerable detail through the nomis database at the university of durham (http://www.nomisweb.co.uk/). but these statistics are not reconcilable with those for ilo unemployment from the labour force survey (lfs). the lfs includes questions on claimant unemployment. but the grossed up statistics from the lfs are typically about 20% below the level of the administrative count of claimants (jenkins and laux, 1999). the local labour force survey in the uk has an enhanced sample in low population areas to increase geographical coverage. this supports the production of unemployment statistics based, for example, on parliamentary constituency areas (pcas) – that have on average a population of working age of 43,000 within a fairly narrow range. but little reliability can be given to most of the unemployment statistics. in 2003 the level of unemployment in 40 pcas was too low to support any estimate of the annual average, and it was not possible to give any confidence level to the estimate of the unemployment rate for 2003 for more than half of the remaining 600 pcas. statistics for claimant unemployment are produced on a 100% basis from administrative statistics, are available monthly, and are more up-to-date. the lfs, at the time of writing, can only give patchy annual unemployment statistics for pcas for 2003. the claimant system provides monthly statistics that, at the time of writing, support analysis of the pattern of seasonal variation for individual pcas in 2004. the pca showing the greatest seasonal variation in 2004 was dorset south, on the south coast, with a coefficient of variation (cov) of 22%. unsurprisingly seaside areas show the greatest seasonal variation. seven pcas have covs of more than 20%. but at the other extreme seven pcas, all in major cities, have covs of less than 2% (calculations by the authors). statistics of unemployment in the uk have been available on a place of residence basis since 1983. their development depended upon the system of postcoding that was completed in 1974 and upon computerisation in the early 1980s of the unemployment statistics based on employment office areas. an account of this development, including explanation of the abandonment of statistics of registered unemployment, is given in brimmer (1981). the incompatibility noted between claimant statistics and the lfs does not provide a sound basis for the production of unemployment rates with the conventional economically active population denominator. the statistics of unemployment rates for local areas first published in 2003 and available back to 1996 have, as noted in section 5, used population of working age (pwa) denominators. geography and full employment the concept of full employment as well as the concept of unemployment was more of less invented in britain. william beveridge’s full employment in a free society published in 1944 remains the most comprehensive single study of unemployment problems in industrial societies. beveridge distinguished frictional, structural, and demand deficiency unemployment. frictional unemployment was conceived as unavoidable unemployment between ending one job and starting another. frictional unemployment can be assumed to be mostly short-term. beveridge would not have been surprised at the small variation in the proportion of exits before four weeks that is revealed by statistics of exits from claimant unemployment. such short-term unemployment would have been classifiable as frictional unemployment which can be expected to exist independently of the state of the labour market. frictional unemployment for beveridge would constitute the minimum level of unemployment achievable. it would set the level of unemployment compatible with full employment. beveridge identified demand deficiency and structural unemployment, and noted the difficulty of making the distinction between them. structural unemployment could be regarded as a form of demand deficiency unemployment. making the distinction and identifying appropriate remedies depends upon the availability of regional and local statistics. several generations of economists have elaborated on the idea of full employment in a theoretical way with concepts such as non-accelerating inflation rate of unemployment (nairu). the central point is that an optimum level of 20 iassist quarterly spring 2005 full employment is achieved when labour market pressure for higher wages and salaries is not sufficient to lead to runaway inflation. the concept of nairu demonstrates that full employment is inseparable from the geographical distribution of unemployment. it cannot be assumed that labour market pressures that lead to wage inflation, or labour market vacuums that lead to unemployment, are likely to occur equally in all labour markets in all parts of a country. if unemployment is unequally distributed geographically, inflationary labour market pressures will be reached first in areas of low unemployment. areas of low unemployment will have achieved full employment or over-full employment while other areas continue to suffer from high unemployment. the proper functioning of the labour market as well as the management of the labour market by the government requires statistical information on areas within a country. ilo/lfs statistics could be said to be adequate at a regional level, but not at local level. over the past few decades in the uk, for example, there has been persistent growth of inner city unemployment. the main unemployment problem has become intra-regional rather than inter-regional. survey based statistics, such as those of the lfs, are inadequate for measurement and investigation of the relatively finely-grained variation in unemployment levels now evident in every sizable urban area. chart 4 shows the distribution of unemployment in england among pca areas. the map divides pcas into quartiles according to the claimant unemployment rate. the lightly dotted pcas are in the lowest quartile with the lowest unemployment rates. the pcas coloured black are those in the top quartile with the highest unemployment rates. in between light grey shading denotes pcas in the second quartile with below average levels of unemployment, and the the dark grey denoted the third quartile with above average levels of unemployment. the map shows that high levels of unemployment are concentrated in urban areas. every city and major town contains major concentrations of unemployment. with a small number of exceptions there are no major concentrations of unemployment that are not urban areas. ilo measures of unemployment are inadequate for investigation of such a fine grained geographical distribution. chart 4 demonstrates the need to combine ‘whole population’ information from the lfs with administrative data on unemployment, as was used to construct this chart. the cultural influence international labour office criteria define unemployment in terms of seeking employment. in other words unemployment is a condition found among the population. at first sight that seems unobjectionable. how can anybody be unemployed if they are not looking for employment? one feature of this definition is that it puts the onus of being unemployed upon the individual. if individuals are unemployed, it is implied, it is their own fault. but the idea that individuals should have the right to work is a component of a number of belief systems. islam, for example, recognises a right to work. the cairo declaration on human rights in islam (1990) states that “work is a right guaranteed by the state and the society for each person with capability to work (http://www.humanrights. harvard.edu/documents/regionaldocs/cairo_dec.htm). the catholic church teaches that “the obligation to earn one’s bread presumes the right to do so. a society that denies this right cannot be justified, nor can it attain social peace.” (centesimus annus, 1991, para 43). the former soviet union managed to achieve full employment by insisting that everyone should work. the un-habitat human settlements programme has a charter of human rights that specifies that male and female citizens have the right to work through worthy employment with sufficient resources to guarantee the quality of their lives. the practical consequence of the ilo exclusive emphasis on seeking work is a lack of acknowledgement of factors that contribute to unemployment. defining unemployment as a condition does not require investigation of cause. the lfs does not, like the cps, include questions on reasons for unemployment, and does not allow for the production of statistics for entry to unemployment that give indications of cause (see thomas, 2005, for elaboration). the easiest reform would be to modify ilo guidelines for the conduct of labour force surveys. modification would require the inclusion of a question on unemployment addressed to all respondents – not just to those unemployed on the date of the survey. for example, ‘have you been unemployed at any time during the past three months?’ or, to more fully comply with other ilo criteria, ‘have you been without paid employment and seeking work at any time during the past three months?’. such questions would recognise the concept of entry to unemployment and would provide statistics on the number of entrants. a follow-up question on dates of unemployment would support the production of statistics for unemployment in the four weeks prior to the survey date. statistics for the number of entrants in the previous four weeks would provide support for the production of accurate statistics on duration of unemployment, and so deal with ‘length-biased sampling’. questions identifying entry could well elicit reasons for unemployment along the lines of the cps. data on reasons would allow for better international comparisons and would iassist quarterly spring 2005 21 chart 4 unemployment rates for pcas in england in 2004 22 iassist quarterly spring 2005 support more comprehensive analysis of time trends in unemployment than is possible with the statistics produced in accordance with current ilo criteria. ilo/lfs statistics are also of limited value in investigating the geographical distribution of unemployment. the inescapable problem with ilo criteria is that they are based on survey statistics that cannot be expected to provide adequate information on local unemployment. they are difficult to integrate with national statistical systems that have the geographical detail and data on entry. ilo guidelines for the conduct of labour force surveys followed the pattern set by the us current population survey more than thirty years earlier. it is ironic that nowadays the cps provides less information on unemployment in the us labour market than do insured unemployed statistics. weekly statistics on entry to insured unemployment provide information on current trends. monthly mass layoff statistics provide information on an important cause of unemployment. the local area unemployment statistics system demonstrates the value of combining administrative systems with survey statistics with the production of statistics that combine information of a few thousand cps respondents identified as unemployed with statistics for around three million continuing claims for unemployment insurance. but the value of such combining is not expressed in the ilo criteria for the conduct of labour force surveys. in the case of the uk the nearest equivalent insurance statistics – the claimant statistics – account for a much larger proportion of unemployment (as defined by ilo criteria) than us insured unemployment statistics. but it is known that the uk lfs data does not provide accurate information on claimant unemployment. the production of estimates of ilo unemployment by means of statistical estimates on the lines of the bls would not be the best solution. the general solution would be, not for labour force surveys to ignore administrative unemployment systems, but for the standard lfs to embrace administrative unemployment systems. the administrative records of insurance or registrant based unemployment systems could be used as part of the sampling frame for labour force surveys. weighted sample figures could be grossed up to national totals in accordance with standard statistical practice and there would be no loss of representativeness. the focus on unemployment could be achieved without reducing comparability between lfs statistics for the employed, unemployed, and inactive populations at the national level, and without jeopardising international comparability of statistics relating to any of these categories. the ilo guidelines for the conduct of the standard lfs could be extended to give detailed guidance on the methods that might be followed. such a development could be expected to contribute to the quality of unemployment statistics as defined both by ilo criteria and by national administrative systems. comparison of the survey results for the administrative sample with that of the general population could be expected to yield information that would support the production of estimates of ilo unemployment for local areas that would be of more ascertainable quality than the laus estimates in the us. the addition of a range of ilo personal profile variables to administratively defined unemployment statistics could be expected to add significant value to these statistics for national policy and decision making. addendum many of the points made in this article are supported by statistical evidence that is included here only in highly summarised form. for a more detailed report see john adams and ray thomas ‘patterns and trends in unemployment in scotland 1985 to 2004’ to be published by scotecon at the universities of stirling and strathclyde. acknowledgement is made to the royal statistical society for the award of a campion fellowship to ray thomas that has supported the research for this article. acknowledgment is made to scotecon for a grant to john adams for the ‘patterns and trends in unemployment in scotland 1985 to 2004’ study that has also supported the research underlying this article. references adams, john and ray thomas (2005) patterns and trends in unemployment in scotland 1985 to 2004 a report to scotecon at university of stirling, march 2005. adams, john and ray thomas (2004) ‘dynamic measures of unemployment’ paper given at the ‘statistics: investment in the future’ conference, prague, september 2004. http:// www.czso.cz/sif/conference2004.nsf/i/dynamic_measures_ of_unemployment anderson, margo (1988) the american census – a social history, yale university press. beatty, christina, stephen fothergill, tony gore, and alison hetherington (1997) the real level of unemployment, centre for regional economic and social research, sheffield hallam university, march. beatty, christina, s fothergill and r macmillan (2000) ‘a theory of employment, unemployment and sickness’ regional studies, 34, pp 617-630. beveridge, william h (1944) full employment in a free society, george allen and unwin. brimmer, m.j. (1981) review of the statistical services of the department of employment and manpower services commission, department of employment, april (part of iassist quarterly spring 2005 23 rayner review). garratty, john (1978) unemployment in history – economic thought and public policy, harper & row. garside (1980) the measurement of unemployment – methods and sources in great britain 1850–1979, blackwell. heap, daniel (2005) ‘comparison of 2001 census and labour force survey labour market indicators’ labour market trends, january 2005, pp 33–48. hussmanns, ralf, farhad mehran and vijay verma (1990) surveys of the economically active population, employment, unemployment and underemployment: an ilo manual on concepts and methods, ilo, geneva. jenkins, james and richard laux (1999) ‘evaluation of new benefits data from the labour force survey’, labour market trends, september, pp 505–515. kiefer, nicholas (1988), ’economic duration data and hazard functions,’ journal of economic literature, 26, 646–679. nickell, stephen (1999) ‘unemployment in britain’ in gregg, paul and jonathan wadsworth (eds) (1999) the state of working britain, pp 7–28. urban 75 (undated) job seekers allowance a survival guide, http://www.urban75.com/action/jsa/jsa1.html thomas, ray (2005) ‘is the ilo definition of unemployment a capitalist conspiracy?’ radical statistics, vol 88 summer (forthcoming). webster, david (1996) the simple relationship between long-term and total unemployment and its implications for policies on employment and area regeneration, glasgow city housing working paper. march. webster, david (1997) ‘the l-u curve’, university of glasgow, centre for housing research, 36. webster, david (forthcoming) ‘long-term unemployment, the invention of hysteresis and the misdiagnosis of structural unemployment in the uk’, cambridge journal of economics. working party on the measurement of unemployment in the uk (1995) ‘the measurement of unemployment in the uk (with discussion)’, j. r. statist. soc. series a, 158. part 3: 363–418. endnotes 1 this paper was presented at the iassist 2005 conference in edinburgh by john adams (j.adams@napier. ac.uk), hasan al-madfai (hmadfai@glam.ac.uk), and ray thomas (r.thomas@open.ac.uk), faculty of social sciences, open university, 35 passmore, tinkers bridge, milton keynes mk6 3dy, england. email: r.thomas@open.ac.uk vol281.indd 4 iassist quarterly spring 2004 editor’s notes welcome to the first issue of the iassist quarterly vol. 28. when talking about publication of the iq we are still in last year as vol. 28 is 2004, and the articles presented here are also from 2004, mainly from the iassist conference. three articles are presented in this issue. from the iassist conference in madison, may 2004, is presented a paper from the session on “mapping the past with gis”. after the w3 a conference presenter has here creatively come up with c3. the paper from stuart macdonald at the edinburgh university data library has the title of “counting cows and cabbages” with the subtitle of “web-based extraction and delivery of geo-referenced data”. the geographic information systems are a growing area for the iassist community of data professionals. the paper demonstrates the scope of the scottish edina’s data services, and readers can here find information and expertise in setting up the gis services. several areas that are added to the normal data library services were presented in the conference session on “the diverse world of digital libraries”. one of the papers addresses “oral history”. so we have now moved from the clean-cut quantitative numeric databases of the old days. the article is “multimedia oral history database” by zoltán lux from the institute for the history of the 1956 hungarian revolution in budapest. the oral archive has tape-recordings and transcripts of about a thousand interviews with participant in the 1956 hungarian revolution or their children. the archive has since 2003 been carrying out some digitization of the materials (incl. photos). judged by the publications presented on the website www.rev.hu there are many projects founded in these materials. the last article is also from the iassist 2004 conference. a session presented data from asia and contained presentations on vietnam, china, and korea. daniel c. tsang is the author of “reflections on a quest for social science data in vietnam”. dan tsang is social science data librarian and bibliographer at the university of california, irvine. the article is from his work during 2004 as a fulbright research scholar in hanoi, socialist republic of vietnam. dan tsang was based at the institute of sociology, vietnamese academy of social sciences. the article mentions the many social science data sources that are available in vietnam, most importantly a wealth of statistical data available in yearbooks (cd-roms) as well as surveys. being a country in fast development not all features are as we expect them to be from a myopic western viewpoint. dan tsang mentions that from researchers he heard that “data is power” – but actually as it turns out “data is money” is more suitable, as many datasets are being marketed for a highest bidder. pay a visit at the iassist website on www.iassistdata. org. you can find information on previous and coming conferences, and the upcoming conference is in edinburgh in 2005 (24-27 may). among other features of the website is the possibility to access the articles of the iassist quarterly as pdf-files. papers for the iassist quarterly are most welcome. papers can be from iassist conferences, from other conferences, from local presentation, etc. contact the editor via e-mail: kbr@sam.sdu.dk. karsten boye rasmussen, march 2005 [assist quarterly 23 attrition and the national longitudinal surveys of labor market experience: avoidance, control and correction by dr. patricia rhoton' data archivist center for human resource research the ohio state university since 1966 the center for human resource research has been analysing the longitudinal surveys conducted by the census bureau for the department of labor. the main purpose of 'this report was prepared under a contract with the employment and training administration, u.s. department of labor, under the authority of the comprehensive employment and training acl researchers undertaking such projects under government sponsorship are encouraged to express their own judgments. interpretadons or viewpoints stated in this report do not necessarily represent the ofiicial position or policy of the u.s. department of labor. these surveys is to study the labor force activity of different population groups. the original groups included men who were 45-59 years old in 1966, women who were 30-44 years old in 1967, men who were 14-24 years old in 1966, and women who were 14-24 years old in 1968. in 1979, a new survey, conducted by the national opinion research center {editor's note:) (norc) in chicago, was added for young men and women who were 14-21 in that year. each of the five surveys is designed to collect information on all phases of the respondent's labor force activity and on other characteristics such as educational attaiiunent, health, family composition, and financial status that are known to be related to such activity. the original plan in 1965 was to interview the same respondents each year for a period of five years. because of the usefulness of the data and the relatively small sample attrition, a decision was made at the end of the first five-year period to continue for another five years. the interview pattern was changed at that time from a face to-face yearly interview to a 2-2-1 pattern. each respondent was contacted by phone every two years, then again in person one year after the second phone interview. this pattern was used again both during the third five-year extension obtained in 1976 and during the fourth five-year extension, obtained in december 1982. at the time of the most recent extension, a study looking specifically at attrition within the different cohorts was carried oul longitudinal studies in general have several advantages over the more frequent cross-sectional studies. while longitudinal studies are very expensive, the data are collected in great detail over time, with respondents reporting events and attitudes as they occur rather than retrospectively. collecting the data in this way also enables the researcher to go beyond issues of correlations to address the more urgent issues of causalit>'. the main advantage of a longitudinal survey, following the summer 1986 24 iassist quarterly same set of respondents year after year, aeates two major problems, however. the first is the difrciilty of relocating respondents for subsequent interviews, and the second is maintaining respondent cooperation over repeated interviews. attrition in the nls table p shows the numbers and percentages of respondents for all interviews up to and including the 1983 questionnaire. the base year row shows only those respondents who were interviewed that first year. between the original scteening and the first interview, some of the eligible respondents were lost: 9.0 percent of the older men, 5.5 percent of the older women, 8.3 percent of the young men, 5.8 percent of the young women, and 11.5 percent of the new youth. while table 1 shows the distribution of interviews between and among the five cohorts. tables 2-5 show interview/noninterview status of the four older cohorts by reason for noninterview. while there are difterences between the cohorts in the distribution of reason for noninterview, within each cohort the distribution of reason remains consistent across the years. the method of interview, whether face-to-face or by telephone, does not seem to affect the attrition rate. some of the losses in the sample are unavoidable. in the survey of mature men (table 2), for example, an increasing percentage of sample losses are due to respondent deaths. the mature women's survey (table 3) has the second highest retention rate among the four older cohorts. this high rate is probably due to the fact that this group is very stable and has low geographic ^editor's note: tables are gathered together at end of article mobility. the young men's cohort has the lowest rate of retention and has been the test case for new attempts to stop the gradual decline in sample size. a variety of factors account for the difficulty in locating these respondents: completion of school, acquisition of new jobs, formation of families, and movement in and out of the military services. the higher rates of attrition in the earuer years were attributed to influx into the mihtary since the sample was drawn, and initial interviewing done during the vietnam war. however, rates remained high even as these respondents returned from the military. the yoimg women's cohort, which is similar to the young men's with respect to completion of school, acquisition of new jobs, and formation of families, posed the added challenge of name changes accompanying changes in marital status, yet the overall response rate has remained high. the new youth cohort has benefited greatly from the lessons taught by experience with the four older cohorts. in 1983, the response rate for this group was 96.3 percent a comparison between this cohort and the first five years of the young women's cohort, which had the best retention rate of the older cohorts, shows that different procedures and techniques can substantially decrease attrition. not only does norc have a higher overall interview rate, but also the organization seems to be better at retrieving respondents. in 1982, 96.0 percent of the original 1979 sample were interviewed. some of these had not been interviewed in previous years: 2.2 percent in 1980, 1.1 percent in 1981, and 0.5 percent in 1980 or 1981. only 165 respondents (one percent) of the original sample had had only one interview after four roimds of the survey. in 1983, the number of respondents who had had only one interview dropped to 115. over eleven thousand (90.7) respondents had been summer j 986 iassist quarterly 25 interviewed every year, and 5.5 percent had completed four out of the five interviews. the impact of attrition od representativeness the gradual decline in sample size over time becomes very important if it results in a biased sample. while each cohort was checked at the end of the first five-year series of interviews, and smaller checks were made in the context of reports on occupational distribution, educational attainment, age distributions, and marital status with published national data, no one looked at all the cohorts systematically until 1982. at this point the issue of representativeness had to be addressed as part of the proposal to extend the cohorts for another five years. such a study could essentially be carried out in either of two ways. first, the remaining sample could be compared with some outside group, such as the decennial census or the current population survey. comparison with an outside sample was difficult given time constraints and the fact that census data were not yet ready for release. while the cps data were available, differences between the cps and each of the four older cohorts had already been documented in the first year. the second alternative was to compare the characteristics of the respondents who were left after ten years with the characteristics of all respondents interviewed in the initial year to see how much difference, if any, there actually was. each cohort was checked for differences in the age distributions, educational attainment levels, employment status, industry and occupation distributions, marital status, smsa status, annual income distribution, and wage and salary' distribution. the yoimg men and young women were also checked for differences in enrollment status. a separate evaluation was done by race for each of the four cohorts. table 6 is an example of the type of table constructed for each group. the ten-year sample was weighted using two methods: the entry level weight and a ten-year weight, which includes successive adjustments for each year's noninterviews. for au cohorts except the young men, the relevant comparison was between the entry year weighted figures and the ten-year sample using the ten-year weight in the young men's cohort, the 1966 sample using the 1966 weight was compared to the 1976 sample using the 1966 weight because the 1976 weight had been adjusted to include individuals formerly in the military. since young men already in the mihtary had been deliberately excluded from the young men's sample, using the 1976 weight could have created apparent differences where none existed. for this group alone, it was more appropriate to use the 1966 weight table 7 summarizes the distribution of differences by cohort and shows that for most characteristics the difference between the two samples was less than two percentage points. after the differences were identified, statistical tests of significance were computed for each of the comparisons. table 8 shows the number of statistically significant differences at various levels for each cohort by race. while the number of differences was higher than would be expected by chance, several were based upon small sample cases in the initial year and characteristics with only two values. in the latter cases, a statistically significant result in one category means the other category will also be statistically different [sic]. after reviewing the entire set of tables, it was clear that noninterviews had not seriously distorted the representativeness of the sample. given this finding, and the ability to apply weights to eliminate any potential bias, the decision was made to continue all four surveys for another five years. summer 1986 26 iassist quarterly it is unclear, however, how further erosion of the samples will affect representativeness. concern with this issue, together with the higher noninterview rates that norc was having with the new youth sample, led to an evaluation of the rules that had been established in the original five-year period and an attempt to see if it was possible to retrieve some of the noninterview cases. retrieving former noninterview cases since the yoimg men's panel had lost the most respondents, it became the target for the first attempt at retrieval. respondents from the 1975, 1976, 1978 and 1980 survey years, who normally would not have been included in the workload (i.e., attempted contacts) because of noninterview status in those years (refused, unable to contact, institutionalized, moved outside the u.s.) were sorted, and a sample of 279 selected. several changes were made in the procediu^es for contacting these special respondents. no restrictions were placed on the number of telephone calls, mileage, or time spent locating and retrieving these respondents. each interviewing packet included the respondent's most recently completed interview and household record card, as well as the most recent questionnaire, and all record cards for any other household members participating in any of the other cohorts. in addition, an expanded list of methods of locating respondents was included. as a result of these additional steps, 104 (37.3 percent) of these respondents were interviewed. these interviews have been flagged and will be checked as soon as the data tapes are available from the census bureau to determine if they differ in any way from the rest of the respondents. if these respondents remain in the sample for the next round of interviews in the latter part of 1983. a concerted effort may be made to use these procedures during the regular interviews and in similar attempts to retrieve noninterviews in the other three cohorts. differences between census and norc one of the major differences between census and norc is the amount of location information obtained from the respondent norc obtains more information, and request information on other individuals with specific relationships to the respondent, depending upon the respondent's circumstances. the interviewer begins by asking the name, relationship, address, and phone number of the person most likely to know where the respondent is. if the respondent is uving in a dormitory, fraternity, sorority, hospital or other temporary situation, the interviewer is instructed to obtain the name and relationship of a householder at a permanent home address. if the respondent is married and living apcirl from a spouse, the spouse's address and telephone nimiber are requested. if the respondent is not living with a parent and has not provided a parent's name, this information is obtained, including whether or not the parents live together. the name of another relative with whom the respondent is in contact, and the names of friends and places to which the respondent goes when not spending spare time at home, are also obtained. respondents are also asked for nicknames, maiden names if they are married women, and whether or not they expect to move in the next 12 months. this extensive list gives the norc interviewer a real advantage when contacting someone on the list, since the ability to mention the respondent's parents, relatives, friends, hangouts or nicknames demonstrates that the interviewer summer 1986 iassist quarterly 27 knows the respondent to some degree and may make the reference more willing to give out information about the respondenl another major advantage that the norc interviewer has over the census interviewer is the existence of a centralized locating shop in chicago. the person working at the locating shop has access to all previous questionnaires, original copies of locator documents and information about the respondent's brothers and sisters. working with this additional data, the respondent can usually be located by phone and reassigned to the same or another interviewer. the census interviewer starts out with less information with which to locate the respondent. s/he has a questionnaire with a label indicating the respondent's name and most recent home address. in addition, there is a household record card for each respondent which contains the telephone numbers, all addresses at which the respondent has lived since the survey began, the names of all persons who have lived with the respondent, and the names, addresses and telephone numbers of only two persons who will always know where s/he can be reached. besides the more extensive locating supplement that norc builds in the interview, several other differences appear. each respondent in the new youth cohort is paid $10.00 for a completed interview, since many researchers believe that even a small amount of money helps in obtaining cooperation, especially among younger respondents. the new youth respondents also had an opportunity' to take a series of tests for the department of defence, which needed to evaluate tests given to individuals in the military. for these tests, which take several hours, the respondents were paid $50.00. when the four older cohorts were first interviewed, paying respondents was not as well accepted. now there are fears that starting this procedure with the older cohorts would cause concern on the part of respondents. they will be interviewed each year for the next several years and are therefore aware that they will be contacted about the same time each year. the census interviewers are told only that they may be conducting additional surveys, and should not tell the respondents that this is the last time s/he will be interviewed. the lack of an answer to give the respondent, in addition to the 2-2-1 pattern, probably leaves the respondent without a sense of when or if s/he will be contacted again. while this ambiguity may not have an impact on their cooperation in the survey, the norc approach leaves the respondent with a greater feeling of certainty about the interviewing schedule. revising the rule for dropping respondents after the first year, respondents in the four older cohorts who refused to participate or had died, were dropped from the census sample. those who were not reinterviewed for any reason for two consecutive years were also dropped. the only exception was made in the young men's sample with those respondents who were in the armed forces. since the sample was to represent the national civilian, non-institutionalized population, young men were not interviewed while they were in the armed forces but were retained in the sample and reinterviewed in the first interview after they had left the services. however, norc's success in retrieving respondents even after they had refused and the success of the young men retrieval effort resulted in a change in these rules. cunently, no respondent is dropped except those who have died. norc goes back each year and attempts to interview all living respondents. another procedural difference is that new youth cohort respondents are told up front that summer 1986 28 [assist quarterly maintaining respondent cooperation conclusions while both census and norc send out advance letters about the entire survey, stressing the importance of the respondent's cooperation, norc also sends out a newsletter that tells respondents in a very "chatty" format about some general results of the previous survey. the census bureau had a short, formal fact sheet that went out with the cover letter, but interviewers reported that respondents did not feel it was very useful. in the 1982 young women's survey, a more extensive description of the surveys and a list of the research results from the survey were sent to any respondent who filled in and returned a postcard requesting additional informatioa over one-third of the respondents interviewed in that wave mailed in the postcard. a variable will be created identifying these respondents and if distribution of the handbook inaeases the response rate in the next round, the handbook will be offered to the respondents in all three cohorts. the new youth survey has, at this time, a considerably better response rate than any of the four older cohorts. much of its success can be traced to the solution of problems that developed over time in the fouj older cohorts. while the necessity of maintaining the same measures over time prevented change in the handhng of the four older cohorts, these problems were conected in the first wave of the new youth cohort questions that the respondents or the interviewer had difficulty with in the four older cohorts were altered so that there was no confusion from the very beginning. perhaps most importantly, given the highly mobile nature of the younger age group, much more detail was obtained on individuals who would always know where the respondent was. in addition, more information about the survey was given to the respondent before, during and after each interview. all these factors combined have resulted in a response rate that is very good for any survey and exceptional for a longitudinal survey in its fifth year, n summer j986 iassist quarterly 29 table 1 m summer 1986 30 iassist quarterly table 2 c — 32£summer 1986 iassist quarterly 31 table 3 £.: 1. ni? &, »± i ii. £ a— £ o tl summer 1986 32 iassist quarterly table 4 • ^1 e-s §o ii summer 1986 iassist quarterly 33 table 5 s § summer 1986 34 iassist quarterly table 6 table 6 selected dibracterisli in 19r.6 of original srnple and simple inlerviewd in 1976 milurc w-n mii tcs only qinracleristics nirrt>or of in 1966 respondcnij in 1566 « potcntiiil ly el ipible for lacfi sn. |ile l\c i killed m 1 ng unweirhled 19611 weirlit' « % (i % (nop) unweighted 19g6 wc « % i (oop) are 45-49 1329 1202 39.3 50-54 1230 1043 34.1 55-59 1041 811 26.5 educational btlairment less then u yrs. 2038 1679 55.3 12 years 885 778 25.fi nbre than ij yrs. 655 583 19.2 ehplo^tnent stat fhployed 3348 2897 94.9 unsiployed 46 38 1 .2 out of labor force 206 121 4.0 industry' a^r iculturc 335 293 10.1 mining 30 25 .9 construction 351 292 ip.l ^vlnu^6ctu^lnk 1000 866 29.9 transporlatlon 315 2f.3 9.1 public adnin. hvirltal status never morrled occupation^ professional 359 nv\nager lal 582 clerical 173 sales 176 crafts 828 operatives 572 household services 180 fanwrs 255 farm laborers 60 laborers 154 s^ca status in smsa 2487 out of snea 1112 annual incone fqual to lero 5 1-2,199 249 3,000-9,999 1418 10,000-14,999 712 15,000-19,999 211 * 20,000 156 wages and salary e<5ubi to zero 777 1-2,999 243 3,000-9,999 1690 10,000-14,999 447 15,000-19,999 81 • 20,000 60 'excludes death, nnilitarv end ^hose orployed sijrvey week. 80.3 81.2 82.6 82.6 90.4 92.0 82.4 71 .1 91 .0 flfi.o 1329 1230 1041 36.9 4996 36.5 57.0 7645 24.7 3385 18.3 2561 39.0 3691 34.4 3239 26.6 2610 652 25.9 2486 26.1 335 10.0 1171 9.2 30 0.9 lie 0.9 351 10.5 1335 10.5 loop 29.9 3805 30.0 315 9.4 1213 9.6 12.9 1644 90.1 5.3 4.6 12278 732 624 90.0 5.4 4.6 10.8 1406 11.1 17.4 5.2 5.3 2245 667 688 17.7 5.3 5.4 24.8 3153 24.9 17.1 2198 17.3 25.9 2765 777 23.5 2876 23.0 1834 19.3 4482 38.8 3917 33.9 3152 27.3 6294 54.6 3002 26.1 2222 19.3 3348 93.0 12709 93.0 2395 95.0 10.2 ilpo 10. p 702 29.3 2660 29.4 375 15.7 1441 15.9 1094 10.0 3220 29.4 1036 9.5 1746 15.9 12.6 1142 12.6 1388 12.7 2298 91.3 8695 91 .3 10526 91.3 10.6 999 11.1 1218 17 7 1623 18.0 1969 4.9 446 5.0 542 5.2 491 5.4 597 24.9 2249 24.9 2722 16.7 1544 17.1 1867 22.3 20p9 21.9 2420 21. b summer 1986 iassist quarterly — 35 table 7 table 8 table 7 nintoer and percentage of differences by panel ~~ ~ absolute differences c^) panel 0^ 2^3 3+ total mature men black 34 (73.9 8 (17.4) 4 (8.7) 46 (100.0) white 43 (95.6) 2 (4.4) 45 (100.0) mature wcmen black mii tp young men black 30 (73.2) 5 (12.2) 6 (14.6) 41 (100.0) white 43 (97.7) 1 (2.3) 44 (100.0) young wonen black 33 (82.5) 6 (15.0) 1 (2.5) 40 (100.0) white 40 (95.2) 2. (4.8) 42 (100.0) 42 (93.3) 3 (6.7) 45 (100.0) 45 (100.0) 45 (100.0) table 8 nurber and percentage of statistically significant differences by panel level of significance panel 190 2% 3% mature men black 4 (9.1) 7 (15.9) 12 (27.3) ^>iite 4 (9.1) 7 (15.9) 14 (31.8) mature women black 2 (4.5) 2 (4.5) 3 (6.8) white 1 (2.3) 4 (9.3) 5 (11.6) young men black 1 (2.6) 4 (10.3) 6 (15.4) vn'hitc 2 (4.7) 4 (9.3) 6 (14.0) young wonen black 1 (2.6) 3 (7.9) 4 (10.5) v-liite 1 (2.6) 2 (5.1) 2 (5.1) summer 1986 vol18172 11spring/summer 1994 the center for electronic records is the unit of the national archives and records administration (nara) that is responsible for, among other things, appraising the electronic records of federal agencies; accessioning those electronic records identified as permanently valuable; preserving the records once they have been acquired; and providing reference services for the records. this paper will offer an overview of the procedures used in carrying out these responsibilities. particular emphasis will be placed on the steps involved in accessioning electronic records. first, we need a definition of records. records, as defined in 44 u.s.c. 3301, ..include all books, papers, maps, photographs, machine readable materials, or other documentary materials, regardless of physical form or characteristics, made or received by an agency of the united states government under federal law or in connection with the transaction of public business and preserved or appropriate for preservation by that agency or its legitimate successor as evidence of the organization, functions, policies, decisions, procedures, operations, or other activities of the government or because of the informational value of data in them. library and museum material made or acquired and preserved solely for reference or exhibition purposes, extra copies of documents preserved only for the convenience of reference, and stocks of publications and of processed documents are not included. there are several ways in which the center for electronic records may have its first contact with an agency concerning a dataset. three of the most common ways are: first, an agency may contact nara to schedule the final disposition of records. under 44 u.s.c. 33, no federal agency may dispose of federal records without the approval of the archivist of the united states. agencies usually obtain approval for the final disposition of their records through the use of a standard form 115, request for records disposition authority. this form is often referred to as a “schedule”, “records schedule”, or “sf 115”. a records schedule contains a description of the records to be disposed of and the proposed disposition. if a records schedule contains electronic records, the center is contacted. second, the agency may contact the center with a direct offer of unscheduled records. this sometimes happens if the agency has produced a dataset that they consider important to documenting the mission of the agency. third, the center for electronic records may initiate contact with the agency in an effort to acquire particular datasets with lasting value. if the center staff learns that an agency is producing important datasets, we contact that agency in an effort to get them to schedule those records. unlike textual records which can sit for thirty to fifty years or more without significant deterioration, it is important to acquire electronic records soon after their creation to ensure their preservation. while there are a number of different ways that the center for electronic records can become involved with the acquisition of electronic records, once the process begins, the same steps are usually followed. the records are first appraised to determine if they have sufficient value to warrant their acquisition and preservation. it would not be practical, possible, or desirable for the center for electronic records to acquire a copy of every electronic record produced by every agency of the united states government. appraisal is the process of determining which records will be retained and preserved and which records will be discarded. records are appraised as permanent if they have high legal, evidential or informational value. records have high legal value if they are required to preserve legal rights of the government or individuals affected by the government. records have high evidential value if they document how a government agency carries out its mission, develops policy, or makes key decisions. records have high informational value if they contain data that might be valuable to a researcher for reasons by mark conrad 1 archivist national archives and records administration reeling them in:accessioning the electronic records of the united states government 12 iassist quarterly other than the reason the federal agency gathered the information in the first place. i do not intend to offer a complete explanation of the appraisal process in this paper. the appraisal of electronic records could be the subject of a paper by itself. the appraisal archivist must weigh a number of factors related to the electronic records at hand in order to determine whether they should be preserved. some of the questions to be answered are: • where did the data come from? • how was it used? • did this dataset have a major impact on federal policy? • is the dataset unique? • is the electronic format the most desirable format for keeping the information? • is the data available in a software/hardware independent format? • is there adequate documentation of the dataset to make the data accessible to a researcher? • is this dataset likely to be used by researchers? once the appraisal archivist has reached a decision and that decision has gone through an extensive review process, the center for electronic records will begin the negotiations with the agency to accession those datasets that have lasting value. accessioning is the process of transferring legal and physical custody of the records from the agency to the national archives. standard form 258, request to transfer, approval, and receipt of records to the national archives of the united states (sf 258), is the document used for transferring title of the records to the national archives. the agency sends the datasets, documentation for the datasets, and the sf 258 to the center for electronic records to begin the accessioning process. the center for electronic records requires the agency to transfer the electronic records in a prescribed form. all records must be transferred in a software/hardware independent format. records must be transferred on 7 or 9 track, 1/2 inch open-reel tape at 1600 or 6250 bpi, or on 3480 class cartridges. records must be in ascii or ebcdic with no internal control characters and blocked no higher than 30,000 bytes. these specifications are used to ensure that the center will be able to transfer the electronic records to new media as current media become obsolete. when the tapes arrive, they are checked for readability. if they contain data errors that cannot be corrected by cleaning the tape, the agency is contacted for replacement tapes. if the tapes can be read with no problems, tape maps and dumps of a limited number of records from each dataset are printed. the accessioning archivist compares the tape map with the documentation to verify that the tape described in the documentation is, in fact, the tape she/he is looking at. the next step is for the archivist to manually compare the printout of several records with the documentation to ensure the record layout and codebook match the actual records. if, for example, a record layout indicates that columns sixty through sixty-five should contain the “respondent’s date of birth” and the archivist finds “peoria” in those columns, clearly there is a problem. the archivist consults the agency for a new record layout or new records. this validation can be very time consuming. the archivist usually validates less than ten records per dataset, but an accession may contain several hundred datasets and individual records may be several thousand bytes long. the center for electronic records is currently developing a system to automate the validation process. the application will allow a staff member to enter information about the record layout and codes for a particular dataset. the dataset can then be checked by the computer, byte for byte, against the entered description. the dataset descriptions can be saved and used for analysis of other datasets with the same record layout and codes. the application should allow the center to preserve the links between individual files from a relational database system. researchers may eventually be able to formulate and execute queries against the data. this will enable them to identify diverse files sharing common attributes that could be used to link the files. this capability would allow researchers to perform analyses that were not contemplated by the agency or agencies that produced the files. the application should be able to produce public use files for those files that contain restricted data. in addition to validating the datasets, the accessioning archivist must process the documentation that accompanies the data. the archivist must first verify that the documentation is adequate. while a record layout, code book, and technical data about the tapes used for transferring the records may be all that is necessary to validate 13spring/summer 1994 the datasets, these files are electronic records of the agency that produced them. the archivist must make sure there is enough information to document who used the data, how they used the data, and what impact the use of these records had on the agency. sometimes this information has not been recorded in any form. in this case the archivist may have to interview agency personnel to gather this information. the archivist must take steps to ensure the documentation will be preserved for the long term. documentation in paper form is placed in acid-free folders and boxes. preservation photocopies are made of unstable documents such as newspaper clippings and fax pages. electronic documentation is subject to the same preservation regimen as the electronic records themselves. once the accessioning archivist is satisfied that the data matches the documentation and the documentation is adequate, two copies are made of each dataset. these copies are compared with the original records to verify that they are true copies of the files. the original tapes are then returned to the agency. at this point the accessioning archivist prepares a final accessioning report. this report lists all the datasets received and processed, identifies any problems that remain to be resolved, suggests how the records should be described, identifies any restrictions on the use of the data, and recommends that the sf 258 should be signed by the appropriate parties to complete the transfer. once the sf 258 has been signed and copies distributed to the appropriate parties, the accession is completed. agencies are required to take steps necessary to ensure the preservation of all unscheduled or permanent electronic records in their custody. when an agency transfers their records to the center for electronic records, the center assumes responsibility for their preservation. under the provisions of 36 cfr 1234.28, unscheduled or permanent electronic records must, among other things, be stored under tight temperature and humidity controls. tapes have to be rewound under controlled tension every 3 1/2 years. the agency must test a sample of their tapes for data loss annually. all records must be copied onto new media at least once every ten years. these requirements serve as an incentive for agencies to schedule all their electronic records and transfer their permanent records to the center for electronic records on a timely basis as well as ensuring against loss of data in the agencies. at the same time they provide for the preservation of the records until the agency does schedule or transfer their records. once electronic records have been accessioned, the center staff takes steps to provide access to the records. the records are described. the center forwards descriptions of new accessions to archival publications and accessions control. this unit maintains a centralized database of descriptions of records accessioned by units of the office of the national archives. the office of the national archives is currently developing a new centralized computing system. the archival information system (ais), as presently planned, will include a module for detailed description of electronic records. when ais is fully operational researchers will be able to query the database interactively to find records related to their research by using one or more sophisticated search paths. the reference staff maintains the “partial and preliminary title list of holdings”. this list contains the titles and some information about the availability of some of the datasets that the center has accessioned. this list is available upon request from the reference staff. in addition to the “title list”, the reference staff produces finding aids to particular collections of data that may be of interest to a wide audience. the reference staff has in depth knowledge of the center’s holdings and offers personalized service to the records. the center for electronic records does not presently provide on-line access to its holdings. at this time a researcher may purchase a copy of a dataset of interest on 1/2 inch open reel tape. the center plans to make datasets available on 3480 cartridges this summer. the center is also investigating the possibilities for making datasets available on media other than 1/2 inch open-reel tape and 3480 cartridges. the center does plan to make some of its datasets available for on-line access within the next three years. this has been a brief overview of the procedures used by the center for electronic records in appraising electronic records of the united states government, accessioning those records determined to have lasting value, preserving those records that have been acquired, and providing reference services for the records. these procedures are constantly being re-examined and revised in an effort to better meet our responsibilities. 1. paper presented at iassist 92, madison, wisconsin. integrating data resources and library resources: the spires experience by slavico manojlovich memorial university ofnewfoundland st. john's, newfoundland introduction the cansim database, maintained by statistics canada, contains over 400,000 socio-economic time series. the cansim university base, a subset of the main cansim base, contains approximately 35,000 numeric time series and is made available to academic institutions on magnetic tape for instructional and research purposes. prior to 1991, access to the cansim university base at memorial was provided by the economics department via a locally developed fortran program which retrieved time series from tape for specified databank numbers. in 1991 the library, for budgetary and other reasons, took over the cansim subscription. cansim has since been loaded into spires, the library's database management system, which is also used to maintain the library catalogue and other library and departmental databases. cansim is now accessible to the university community using the same interface as other library databases. it has also been linked to the library catalogue for the display of holdings information associated with related print publications. the remainder of the article will describe various issues regarding loading a numeric database in a library environment. selecting a database management system the spires database management system, developed by stanford university, has been used for over 15 years to manage a variety of information resources including bibliographic, numeric, full-text and image databases. memorial has been using spires for several years to maintain its library catalogue. various components of an integrated library system including acquisitions and cataloguing were developed locally, whereas, circulation was obtained from rensselaer polytechnic institute, a member of the spires consortium. in addition to the library catalogue memorial has mounted a number of locally developed and commercial databases using the folio interface available in spires (screen no. 1). spires has a powerful set of development tools which enable you to tailor the system to accommodate a variety of data types. since cansim was the first numeric database which the library was loading into spires an important first step in the loading process was lo identify the dbms functionality which was required in order to support adequate access to a numeric database. the following functions, although not critical for bibliographic database support, were important for accessing cansim: 1) since end-users typically retrieve time series data for specified time periods spires must prompt for the time series start and end dates. 2) spires must format the data for tabular display or for input to statistical analysis software. 3) spires must output data to a file on the mainframe or on the end-user's microcomputer. the cansim database in spires the cansim file supplied to memorial contains 1.5 million card images. the logical record describing one time series is comprised of various fixed fields (codes and text strings) and a variable number of data values (figure 1). spires provides a utility for loading data in its original form thereby alleviating the need to write loader programs. no problems were encountered in loading the cansim file. spires has a library of processing functions which allows you to read in data in any form, store it as you like and output it in any form. this inherent fiexibility of spires made it easy to implement prompting for start and end dates and the various output formats. spires was also able to accommodate the requirement to output a variable number of data values for specific years when retrieving weekly time series. the various features of cansim in spires are illustrated in the following sample search session (please refer to screen displays at the end of the article): 1) a variety of help screens navigate users through a search. (screen no. 2). 2) users can search cansim directly using the find command on various indexes (screen no. 3). terms can be combined from various indexes using boolean operators. the brief record display output from a keyword search on "newfoundland and women and unemploy#" is illustrated in screen no. 4. 3) the full record display (screen no. 5) includes in lassist quarterly the source field not only source publication information as supplied by cansim but also information regarding memorial's holdings for that title. for each displayed record spfres looks up the holdings information in the library catalogue and includes it in the source field. the addition of holdings information enables the user to easily consult the corresponding print publication for additional information describing the time series or to obtain older data not included in cansim. spires performs a corresponding look-up for users searching the library catalogue and notifies them of related information in the cansim file (see the notes field in the record display from the library catalogue— screen no. 6). 4) users have the option of displaying data in tabular form (screen no. 7) or as raw data formatted for input to a variety of time series analysis softwarepackages (screen no. 8). the user selects the desired format by issuing a dis table or dis data command. spires will then prompt the user for the time series start and end dates. output in either tabular or raw data format may be directed to a file on the mainframe or to the user's microcomputer by issuing either the save table or save data commands. in save mode spires prompts the user for the output file name. 5) users who are uncertain of the appropriate search terms may use the spires browse command to scan entries in various indexes (screen no. 9 and no. 10). users can selectively display time series from the hit list of displayed terms. all text indexes are may he browsed in the cansim application. 6) the above session used spires menu-driven folio guided mode. a command mode option is also available for more experienced users. the integrating role of spires the concept of integration in library automation literature typically refers to linkages between various modules in an integrated library system. the cansim application described above illustrates how spires expands upon the traditional meaning of integration. integration in a spires environment also includes: 1) interface integration: the use of a common user interface for accessing avariety of databases. formats for direct input to other systems (eg. sas). 4) workstation integration: support for saving data to a file on the mainframe or on the user's workstation. all of the above further integrate the user into his/her information environment. cd-rom version of cansim statistics canada is also distributing a larger portion of the cansim database in cd-rom format. although cd-rom is an excellent cost-effective medium for the distribution of large quantities of data it falls short in terms of "integration"as described above, esf)ecially in a university environment. the cd-rom medium forces the user to leam a new user interface. it is not directly linkable to the library catalogue and other campus databases. access to the database may be restricted to a single workstation. if the cd-rom is mounted on a campus network, access may be restricted to users with particular hardware. remote access to a cd-rom by an end-user working on an old vt-loo terminal may not be possible at all. migrating data from the cd-rom to the end-user's statistical analysis software (which is probably on a mini or mainframe) may also be quite cumbersome. conclusion thanks to the integrating pxjwer of spires, access to cansim at memorial university is the same as accessing the library catalogue or any other bibliographic database. based on the success of mounting cansim, memorial is planning on loading census and other fact databases in spires thereby further expanding access to the world of numeric data resources for the library user. notes 1. additional information on spires is available from: spires consortium office, stanford university, stanford, california. ' presented at the iassist 91 conference held in edmonton, alberta, canada. may 14 17, 1991. slavko manojlovich, assistant to the university librarian for systems and planning, memorial university of newfoundland, st. john's, newfoundland. 2) database integration: the provision of automatic linkages between databases thereby expanding the user's knowledge base (eg. source publication link between cansim and the library catalogue). 3) system integration: the output of data in a variety of summer 1991 figure 1 sample cansim rkcord as supplied by statistics canada add b 1 90-1102 195319901210 5 2 310 1 199999999* (12x, 4f17.0) bank of canada assets and liabilities , weekly series (101 -103) and monthly series (1-3), wednesdays and average of wednesdays, unadjusted , millions of dollars . | b.of c-statement/ave total assets dollars scalar factor 6 source bank of canada review average of wednesdays cansim series identifier 000911. 1 note data published in the bank of canada review approximately 30 calendar days after end of reference month . b 1 1 2348. 2318. 2332. 2352. b 1 2 2354. 2352. 2410. 2408. b 1 3 2371. 2364. 2429. 2444. b 1 4 2390. 2404. 2355. 2357. b 1 5 2427. 2431. 2309. 2284. b 1 6 2326. 2304. 2382. 2420. b 1 7 2369. 2226. 2278. 2310. b 1 8 2316. 2357. 2433. 2468. b 1 9 2464. 2476 2532. 2547 b 1 10 2509. 2368. 2421. 2472. b 1 11 2467. 2511. 2528. 2531. b 1 12 2519. 2543. 2550. 2571. b 1 13 2514. 2406. 2429. 2492. b 1 14 2519. 2580. 2604. 2629. b 1 15 2632. 2645. 2696. 2670. b 1 16 2606. 2540. 2574. 2646. b 1 17 2652. 2719. 2800. 2855. b 1 18 2885. 2997. 2956. 2951. b 1 19 2800. 2753. 2768. 2809. b 1 20 2838. 2857. 2857. 2928. b 1 21 2880. 2848. 2943. 2869. b 1 22 2822. 2728. 2736. 2816. b 1 23 2830. 2842. 2902. 2905. b 1 24 2860. 2895. 2950. 2927. b 1 25 2906. 2824. 2876. 2896. b 1 26 2920. 2909. 2981. 2998. b 1 27 3030. 3066. 3064. 3066. b 1 28 3062. 2940. 2990. 3075. b 1 29 3105. 3227. 3242. 3309. b 1 30 3178. 3205. 3215. 3221. b 1 31 3136. 3012 3072. 3167 lassist quarterly screen no . 1 menu of public access databases at memorial folio contains 17 files. press the return key to see the rest of the list. public information files mun library online catalogue canadian socio-economic time series database archival records of the centre for nfld. studies division of extension resource library canadian labour bibliography labrador institute of northern studies info cen . canadian research and report literature mun folklore and language archive ocean engineering information centre newfoundland periodical article bibliography grenfell college fine arts slides database ===> press the return key to see the rest of the list, or select a file by typing its name or number which file? cansim 1. biblio 2. cansim 3. cns archives 4. extension 5. labbib 6. linsic 7. microlog 8. munfla 9. oeic 10. pab 11. slides screen no . 2 initial cansim help screen -cansim selected cansim is statistics canada's computerized data bank and information and retrieval service. the cansim university base on folio contains 32,425 of the most popular time series in the main base. subject areas covered by cansim include system of national accounts, population, labour, prices and international trade. the database is currently updated once a year (the last update was december, 1990) . time series may be output as either tabular displays or in raw data format suitable for input to a variety of time series analysis software packages. raw data or tables may be output to a file at nlcs or downloaded to the user's microcomputer using kermit communications software. the cansim full display includes the mun library holdings of the associated print publications. enter help cansim output for more information on cansim display formats. for additional information contact joy tillotson, information services, qe ii library, ext . 7427. for more information on this file: type help. to search this file: type find. to search headings in sequence: type browse. to select a different file: type select. to see all your options: type options. your response: find summer 1991 31 screen no . 3 menu of cansim indexes you can search the cansim file for any of the following information: type of search matrix matrix numbe source keyword index to source publications maword keyword index to matrix titles seword keyword index to series titles keyword keyword index to titles and notes freq frequency db databank numbers example 911 or 000911 bank of canada review bank of canada assets mortgage* consumer price monthly bio indicate below the type of search you want by typing the name or names of the type of information you have, e.g. matrix. use the browse command to examine entries in various indexes. type of search: find word newfoundland women unemploy* screen no . 4 cansim brief record display cansim / search: find keyword newfoundland women unemploy# result: 6 series 1) series: unemployment rate women 20-24 yrs . (monthly, 1975-1990) (databank no: d774089) matrix: newfoundland, basic labour force characteristics, monthly from jan 75, unadjusted (flaw) in thousands . selected series are linked to previous surveys of jan 66 or jan 70. (no: 002078) 2) series: unemployment rate, 25 yrs and over, women (monthly, 1975-1990) (databank no: d772686) matrix: newfoundland, basic labour force characteristics, monthly from jan 75, unadjusted (raw) in thousands. selected series are linked to previous surveys of jan 66 or jan 70. (no: 002078) _series continue; press return to see next page_ to see a full series: to begin a new search: to select a different file: to get more information: your response: df 1 type display full followed by a number. type find or browse. type select. type help or options. 32 lassist quarterly screen no . 5 cansim fitll record display (databank no : cansim / search: find keyword newfoundland women unemploy* result : 6 series 1 unemployment rate women 20-24 yrs . {monthly, 1975-1990) d774089) newfoundland, basic labour force characteristics, monthly from jan 75, unadjusted (raw) in thousands. selected series are linked to previous surveys of jan 66 or jan 70. (matrix no: 002078) ; monthly labour force data (71-001), stc mun holdings: the labour force. la population active 71-001, location: mugd, holdings: v. [ 10-11] [33] 1954source materials covering backgrounds of previous lfs revisions and modifications of definitions, of concepts and of lfs design, as they introduced in the current revision may be obtained from the division. requests should refer to 1) 'labour force information', cat. no. 71-oolp, feb76 2) 'research paper' #2 and #3 3) 'methodology of the canadian labour force', statistics canada, cat. no. 71-526, ottawa, 197 6 scalar factor: 00 data output format: (lox, 4f17.1) missing values = 9999999999. or equiv. secure data = all asteris)cs. series series : matrix : source : notes : are screen no . 6 library catalogue database record display biblio / search: find title labour force result: 16 titles title 16 the labour force. la population active labour force bulletin. main d' oeuvre bulletin 1945-19 ottawa, statistics canada, labour force survey division title former title published description dates issn notes notes : subject (s) : other entries: other entries call number rsn ==> notes v. [1]194503806804 v.32, no. 1-2 not published vol. numbering begins with v. 6, no.l. mar. 1960 issues for 1945-49 called no. 1-13 vols, for 1945-1971 issued by the dominion bureau of statistics. ; 197 -19 by statistics canada, labour force surveys section labor supply—canada—statistics—periodicals . statistics canada. labour force surveys section canada. dominion bureau of statistics labour force bulletin 71-001, location: mugd, holdings: v. [ 10-11 ]-[ 33] 195475348973 *** note: numeric data associated with this record are summer 1991 33 screen no . 7 cansim tabular record display 1) series: unemployment rate women 20-24 yrs . (monthly, 1975-1990) (databank no: d774089) matrix: newfoundland, basic labour force characteristics, monthly from jan 75, unadjusted (raw) in thousands. selected series are linked to previous surveys of jan 66 or jan 70. (no: 002078) year jan feb mar apr 1985 32.7 30.0 29.2 ' 31.0 year may jun jul aug 1985 32.7 33.8 25.8 27.0 year sep oct nov dec 1985 25.4 26.0 25.9 23.3 screen no . 8 cansim raw data output format series: unemployment rate women 20-24 yrs. matrix: newfoundland, basic labour force characteristics, monthly from jan 75, unadjusted (raw) in thousands. selected series are linked to previous surveys of jan 66 or jan 70. (matrix no: 002078) (lox, 4f17.1) (monthly, 1985-1990, scalar factor: 00) 32.7 30.0 29.2 31.0 32.7 33.8 25.8 27.0 25.4 26.0 25.9 23.3 26.9 25.1 28.0 26.1 25.9 25.9 30.1 23.8 30.0 29.2 33.8 25.8 26.0 25.9 25.1 28.0 25.9 30.1 lassist quarterly screen no . 9 menu of cansim browse indexes you can browse the following types of information in the cansim file: type of browse example mat i matrix titles banlc of canada seti series titles short term loans soph source publication bank of canada review word list of all terms in text fields consumer indicate below what you wish to browse by typing the name of the type of information you have, e.g. soph. use the find command to perform a direct search. type of browse: browse word fish screen no . 10 cansim browse keyword index cansim / search: browse word fish result filed under the following headings: -3) word firms (690 series) -2) word first (4328 series) -1) word fiscal (129 series) 0) word: fish (332 series) 1) word: fishery (1 series) 2) word: fishing (3688 series) 3) word: fitting (2 series) 4) word: fittings (7 series) 5) word: five (926 series) 6) word: fixed (1053 series) 7) word: fixture (63 series) 8) word: fixtures (103 series) headings continue; press retufuj to see next display followed by a page heading number.to see a heading's series: type to see a full series: type display full followed by a number. to begin a new search: type find or browse. to get more information: type help or options. your response: dis 1 summer 1991 iassisl quarterly 3 debates and directions in the future of opinion polling data the following two papers were presented at the 1assist '87 conference in a session entitled: the uses of sociopolitical data. the session focused on the comparability of electoral data, public opinion data and other comparative data projects, including technical and political factors affecting secondary analysis. (ed. note). by neil guppy 1 department of anthropology and sociology university of british columbia topical issues, and the mass media broadcast polling results in an incessant stream. daily we see or hear new polling results which reveal how our contemporaries rate the politicians, or the postal service, or the latest soft drink. recently, pollsters have taken to asking the public how they feel about polls. since pollsters are concerned with assessing the images and opinions of the population, it is hardly surprising that polls on polling, or surveys on surveys, have increasingly found their way into tne polling literature (see e.g., roper, 1986; goyder, 1986). the irony of using polls to evaluate polls is not lost on the pollsters, and a good deal is revealed by these self-assessments. this paper reviews this recent literature in an attempt to gain some leverage on the potential directions of public opinion research in the next decade or so. 1 begin with estimates of the sheer volume of polling data now being collected in different countries. this pervasiveness of polling, though, has generated substantial controversy and conflict in the practice of polling. in an attempt to understand the possible ramifications of current practices and techniques for the future of opinion polling data. i review general criticisms levelled against opinion polling. these criticisms are used to organize a discussion of future directions for both the industry, and by implication, for those who rely on poll-generated data. a social invention of this century, opinion surveys arc now commonplace in liberal democratic societies. major political parties cannot afford to ignore polling data, market researchers tap public opinion on a plethora of 'this is a revised version of a paper originally presented at the annual international association for social science information service and technology (iassist) conference held in vancouver, british columbia, canada on may 19-22, 1987 the prevalence of polling it is difficult, especially on an international scale, to ascertain exactly how much polling data is currently being collected. no central registry of polling data is available, so a variety of proxy estimates must be employed. i rely on summer 1987 4 iassist quarterly three — the percentage of the population reporting participation in opinion research studies, the amount of money spent on polling activities, and the frequency of publication of polling results. one method of assessing the prevalence of polling is to determine how many members of the general public have been involved as respondents in polling or survey research. table 1: trends in respondent involvement, 1978-1984 participation 1978 1980 %ever 47 59 %lastyear 19 25 1982 59 23 1984 54 23 source: schleifer, 1986 table 1 shows trend results from a series of u.s. polls where respondents were asked, first, to report whether they had "ever participated in a survey before", and second, whether they had previously been "interviewed for a survey in the past year" (schleifer, 1986). since 1980 the majority of americans have been contacted at least once, and almost one-quarter have panicipated in at least one poll or survey in the past year. table 2 presents similar findings reported by roper (1986). his findings parallel the results reported in table 1, demonstrating that most americans now have first-hand experience with polling and survey research. furthermore, many americans have multiple experiences as participants. results from other countries are less systematic. goyder (1986) reports that canadians in a mid-sized city have experienced levels of participation roughly equal to the u.s. findings. table 2: respondent involvement in public surveys, 1985 percent never 41 once 17 twice 16 3-5 16 6+ 9 d.k. 1 source roper, 1986 estimates of the amount of money spent on polling are available only for the u.s., where advertising age annually reports financial data for polling firms. for the fiscal year ending in 1985, u.s. research firms in the marketing, advertising, and polling sector billed for $1,785.3 million, up 11.5% from the previous year, and more than three umes the annual rate of inflation. of this, approximately 78% was from u.s. based work (see honomichl, 1986). in the mid-1970s, paleu et al. (1980) reported that the new york times ran news stories containing polling data on an average of one in even three days. a rough count of news, editorial, and feature stories in the 1985 new york times index reveals 278 items under the heading "public opinion polls" (a non-election year in the u.s.). worcester (1980) reports that in the united kingdom (as elsewhere), opinion polls dominate the front-page headlines during the build-up to national elections. polling is pervasive, and every indication suggests that growth has continued to this day. recent advances in random digit dialing and computer assisted telephone interviewing have served to extend the pollsters' reach even summer 1987 iassist quarterly 5 farther. rather than dulling criticism, this growth has occurred in the face of skeptical commentary. criticisms of polling several very general criticisms of polling are heard frequently. often the charge is made that polls are, at best, only a superficial barometer of public beliefs. in its strongest guise, this argument suggests polling results are frequently wrong. a weaker version claims opinion polling gives only a perfunctory account of facile opinions. others argue that polls are invasive, trampling public privacy by asking for personal information (e.g., political preferences). yet others complain that polls have fundamentally altered the political process such that substantial policy matters are unduly influenced by popular and often ill-informed opinion rather than by thoughtful deliberation. these are very basic arguments on which substantial ink has already been spilled. rather than add to this area of the debate, i will attempt to look behind some of these general objections to more specific issues in the practice of polling. 1. distortion — one concern is that people don't give true responses when answering the pollsters' questions (lewis and schneider, 1982). individuals lie, or as the pollsters say, respondents "misreport". sometimes this appears to be deliberate, as when people are asked whether they voted in the last election (evidence shows that more people claim to vote than actually do vote). on other occasions, it is less certain whether people actually lie or whether they are generally confused; a classic example here is a poll conducted by the german magazine der spiegel in which a fictitious cabinet member came sixth in popular rankings, ahead of ten real-life ministers of the crown. 2. non-attitudes — if pollsters ask people questions about which they have no opinion, some people feel pressure to respond and instant opinions may be invented. evidence suggests that the more remote an issue is from a respondent, the more random is the response. it is especially on this basis that critics claim polls are superficial. 3. opinion change — individual attitudes are often not stable or deep-seated. snap-shots from opinion polls may be as interesting as yesterdays news, but they are known to be poor launching pads for general social forecasts. 4. issue complexity — few issues are so clear-cut that single attitude questions can capture the essence of the matter. the black and white image of the world that one may acquire from reading opinion poll results does not do justice to the full array of public sentiment. 5. words and deeds — the ease with which people may express an opinion on a topic is no guarantee of the direction their actions may take. the link between attitude and behaviour has been probed repeatedly, and we still have less than perfect knowledge of when any congruency between the two will hold. 6. question wording — social scientists have known for some time that subtle changes in question wording can influence response patterns. asking respondents whether they would "forbid" or "not allow" something leads to very different results, with a swing of some 20% in response frequencies. so too, the sequence of questions in an interview can influence responses. 7. impersonality — just as students complain that multiple choice exams do not adequately assess the depth of their knowledge, so some summer 1987 6 lassist quarterly argue that polls similarly distort reality because people are forced to respond in fixed categories which rarely allow them any self-expression. the frame of reference for the entire polling exercise is determined a priori and this can easily disqualify certain questions and certain responses. 8. sampling — the ability of samples of several hundred people (up to about 2,500) to accurately reflect the diversity of opinion in an entire nation has often been doubted. polls seldom tap the rich or the poor in any society, thereby predominantly reflecting the views of the middle class. the more general criticisms, and the eight more specific objections to polling listed immediately above, are likely to continue to surface in debates over polling in the forseeable future. the veracity of these claims is often less compelling than the volume of their elucidation would suggest. both polling experts and secondary users of polling data are conversant with the limitations involved. it does not follow from this, however, that these criticisms will gradually dissipate, or even more importantly, that they will have no consequences for the future of opinion polling. public attitudes about opinion polling data are as likely to influence decisions about the future of polling as they are to influence the future of political parties. what i turn to now is evidence pertaining to criticisms of opinion research in an attempt to develop some perspective on future directions in the polling marketplace. directions and tendencies i) assessments of accuracy one recurrent question concerning polls is the frequency with which the pollsters accurately reflect public sentiment william buchanan (1986) has examined the results of election polling in several western democracies in an attempt to examine the precision of election forecasts based on opinion surveys of voter intentions. in analyzing 155 polls from 68 national elections, he found that on 22 occasions the wrong party was predicted as being victorious. beyond forecasting the wrong victor in 1 out of 7 attempts, there appears to be no trend of improvement since 1949. similar findings are reported by worcester (1980) for the u.k. in general, erroneous forecasts are made by a group of pollsters for particularly close elections. it is not true that the polls are always wrong, although it is the case that when the polls are wrong, they all tend to be wrong. a second issue, linked to accuracy, is the actual reporting of polling results. here the quesuon is not so much whether the polls are correct or incorrect, but whether they are properly reported by the press. reporting is crucial, because it is via press reports that most people form their perceptions of the practices of pollsters. smith and verrall (1985) undertook a critical evaluation of australian television coverage of election opinion polls and they claim "poll coverage is extensive, superficial, and inaccurate". typical errors included "temporal transposition" (incorrectly attributing past or present results to some future point), overgeneralization (extending claims to beyond the sample universe), overstatement (exaggerating the strength of findings), and making ambiguous contrasts (comparisons between poorly conceived groups or time periods). summer 1987 iassist quarterly 1 related to this is yet a third aspect, the completeness of press reports on opinion polls. here the concern is with whether or not the press meets basic reporting standards so that consumers can make informed judgments regarding polling results. table 3 contains a listing of the basic standards, showing the percentage of poll reports (either election or non-election polls) which comply with these basic levels of reporting adequacy. as the percentages reveal, the three papers under examination (l.a. times, chicago tribune, atlantic constitution) do not do a particularly good job of providing basic information for informed judgments about polling results. these figures are from the 1970s, and there may have been improvements since this time, especially as the media assign specific people to do all their polling reports. no systematic evidence is currently available to assess possible improvements. table 3: polling standards versus polling practice 1972-79 % reported in electionstandards non-election sample size 89 81 sponsor 80 87 wording 71 34 sampling error 31 2 population 91 66 method 62 38 timing 76 48 n 61 55 source: miller and hurd, 1982. [in norwegian] finally, pollsters have asked the public about their perceptions of polling accuracy. andrew kohut (president of gallup) reports that in 1985, 68% of respondents thought the pollsters got election forecasts correct most of the time (kohut, 1986). this was an increase in public confidence from 57% in 1944. when asked about non-election polls, however, only a slim majority of people felt the polls were generally right in tapping the public mood (52% in 1944; 55% in 1985). conversely, negative sentiments about non-election polling accuracy seem to have increased, with 21% saying they felt the polls were "not right at all" (up from 12% in 1944). the british appear to be a little more sceptical about polling results, with only 32% believing that "the opinion polls are normally right" (46% thought they were normally wrong, with 22% giving other responses — see worcester, 1980: 561). if public opinion is as influential as the practice of polling implies, then pollsters need to be conversant with the public images of polling. when gauged by specific measures of accuracy, polling does not have a massive majority of support ii) bogus polls polling got a bad name in the 1936 u.s. presidential election when the literary digest, a magazine for affluent americans, asked readers to write in with their choice for president. on the basis of these responses, the digest predicted that alf landon would win the election, opening itself and the prestige of polls to ridicule when franklin roosevelt won by a landslide. pollsters have insisted that only surveys with "scientifically" selected random samples should be called polls, and for some lime this seemed to be accepted practice. recently, however, phoney or bogus polls have become more prevalent for example, after the 1980 carter-reagan debate, abc asked viewers to phone and report who they felt won the debate. on the basis of some 727,000 calls. abc reported that their "poll" showed people felt reagan had won 2-1. this phenomenon of summer 1987 [assist quarterly self-selected "samples" appears to be growing in the polling marketplace. qube is a more recent invention, allowing cable subscribers to send digital signals back through the video cable to record their "vote" on various issues. political parties in several countries have taken to doing "surveys" of the public, asking people first to rate the current government, and then to donate money to help the party doing the survey. the prevalence of these bogus surveys is hard to detect, but concern is mounting that sales people are using this technique to identify potential customers. schleifer (1986) reports that in 1980 some 13% of u.s. respondents reported having been exposed to false surveys, a percentage that increased by 4 points to 17% in 1984. with almost 1 person in 5 being confronted with bogus surveys, the reputation of the industry could be tarnished quickly and decisively. in) respondent burden knowing that 1 in 5 people are approached by phoney surveys, one wonders about the frequency with which people have been approached by legitimate pollsters or survey researchers. estimates vary, as shown in tables 1 and 2, but over one-half the population in the u.s. appears to have been involved in a survey or poll at some time. schleifer (1986) reports that of the 23% of respondents who had been involved in a survey in the past year, almost 1 in 5 had participated in four or more polls. these latter individuals are known to survey researchers as "professional respondents" and the inclusion of the same people in multiple surveys has pollsters worried about the "freshness" of their samples. the growth of polling raises the possibility of 'over-kill' — people will be 'turned-oft polls by too many requests for their help. one way of assessing respondent burden is to examine empirical evidence of possible overexposure. steeh's (1981) results, shown in figure 1, chart refusal rates between 1952 and 1980 for two national u.s. samples, both conducted by the university of michigan survey research center. as the graph shows, refusal rates in both the election and consumer attitude series are rising, although whether or not this is due to respondent burden per se is difficult to determine. as steeh notes it could be caused by one factor or a combination of factors, including overexposure, disillusionment with the use or accuracy of survey results (see above), or heightened concern about privacy and confidentiality. goyder and leiper (1986) report similar trends based on an analysis of polling and survey response rates in the u.k., the u.s., and canada, and they point to rising criticism of census practices, especially in canada and the u.k. iv) exit polls in 1980, jimmy carter conceded defeat before the polls had closed in the american west one reason for this was that the television networks were using exit polls to predict the winner before everyone had had an opportunity to vote. exit polls (or 'same-day polls' as they are called in britain) are conducted by standing outside selected polling places and asking those leaving for whom they had voted. based on these reports, the networks have been able to forecast with accuracy the eventual winner. the state of washington was so upset with this practice that they banned exit polls, making it illegal for people to conduct surveys within 300 feet of a polling station. the media challenged the law in court, losing an initial verdict and then winning on appeal — the state is currently appealing the appeal. whatever the eventual outcome, the concern remains that the techniques and the process of polling have fundamentally altered the practice of politics. whether exit polls actually alter the outcome of elections is debatable (see sudman, 1986), although they do appear to have a summer 1987 iassisl quarterly 9 marginal impact on voter turnout when a landslide has occurred. the key point here, however, is not whether exit polls actually have any effect, but that people believe they have an effect as pollsters themselves have shown, it is the image that is important v) polling initiatives pollsters have recently been expanding their craft at a rapid rate. the use of polls for marketing is an old and established pastime (labatts brewery in canada has opinion data dating back to 1910 in canada). more recently, polling has had an influence in the courtroom where survey results have been used in judgments over trademark protection (the nfl, coming glass works), advertising claims (pepsi vs coke), and jury selection (ford, ibm, mci communications). furthermore, various departments of government charged with regulatory functions have begun to use polling data to assess the impact of certain initiatives. listerine was required to engage in corrective advertising to dispel the myth they had created that the mouthwash would prevent people acquiring colds and sore throats. to assess the effectiveness of the correction, the u.s. regulatory agency that was responsible for enforcement used a poll to examine changes in opinion (see crespi. 1987; dutka, 1982). as image and knowledge grow in importance, the pollster's craft is more in demand. but as the demand rises, the value of information escalates, and polling agencies in the private marketplace are less willing to freely relinquish their data. beyond cost, the sheer volume of information frequently makes archiving data a burden to avoid — profit lies with the next project conclusions u.s. respondents continue to report that they feel polls are "a good thing" (73% in 1944 and 76% in 1985 — see kohut 1985), although the british are less sanguine, with a majority feeling that polls were "pointless" or "not very accurate" (worcester, 1980: 560). potentially, fatal dangers for the polling industry would appear to lurk in the areas of bogus polls and respondent burden. the very prevalence of polling may undermine the craft as individuals feel inundated with strangers asking dubious questions about issues which people increasingly define as nobody else's business (see goyder and leiper, 1986 for an analysis of increasing objections to the census). serious, but probably not fatal, dangers would appear to lie in the possibility of disastrous election predictions in several countries simultaneously, or the use of polls in a way so as to make people feel their personal freedoms or rights are subverted (as seems to be the case with exit polls). finally, several signals suggest that more and more public attitude data from polling firms will become off-limits. currently the vast majority of opinion polling data is not publicly available. increasingly, polling data will be kept secret as polling agencies realize the economic value of trend projections. as the value of information grows, pollsters will protect their investments and profitability. several companies in north america now conduct omnibus surveys which they keep confidential. in addition, several countries now have governments collecting general social survey data. thus there will be less pressure on the pollsters to serve the academic interest by releasing the poll data. summer 1987 10 — (assist quarterly references buchanan, william. election predictions: an empirical assessment. public opinion quarterly 50(2):222-227, 1986. dutka, s. bringing polls to justice. public opinion october/november, 47-49, 1982. goyder, john. surveys on surveys: limitations and potentialities. public opinion quarterly 50(1):27-41, 1986. goyder, john and jean leiper. the decline in survey response: a social values interpretation, unpublished paper. waterloo, onl: university of waterloo, 1986. honomichl, jack. the nation's top marketing/advertising research companies. advertising age may 19, 1986. kohut, andrew. rating the polls: the views of the media elites and the genera! public. public opinion quarterly 50(l):l-9, 1986. lewis, i. a. and w. schneider. is the public lying to the pollsters. public opinion april/may, 1982. miller. m. and r. hurd. conformity to aapor standards in newspaper reporting of public opinion polls. public opinion quarterly 46(2):243-249, 1982. roper. burns. evaluating the polls with poll data. public opinion quarterly 50(1): 10-16, 1986. schleifer, stephen. trends in attitudes toward and participation in survey research. public opinion quarterly 50(l):17-26, 1986. smith. t. and d. verrall. a critical analysis of australian television coverage of election opinion polls. public opinion quarterly *9 2j:58-79, 1985. '..:::. c. trends in non-response rates. public opinion quarterly 45:40-57, 1981. aorcester, robert. pollsters, the press, and poiiucal polling in britain. public opinion quarterly 44(4): 548-566, 1980. summer 1987 vol23/1 spring 1999 11 democratizing access to data: the american religion data archive by roger finke, jennifer mckinney and matt bahr * democratizing access to data: the american religion data archive (www.thearda.com) from its beginning, the american religion data archive (arda) was developed to provide immediate access to the best data on american religion at no charge. starting in 1997, the arda was created as an internet-based archive and was designed to serve a highly diverse audience. but serving a diverse audience, including many with little or no background in the social sciences, required arda to meet the rigorous methodological standards of the social science community and still be easily used by those without a knowledge of statistics, research design or data management. since its inception, the arda has attempted to democratize access to data, without compromising the integrity of the data being archived. this essay will review our efforts. we begin by giving a brief overview of the data we archive and the audience we serve. next, we will review the goals of the arda and how we attempt to achieve each goal. although the goals are similar to many other archives, we will highlight how we have developed features that allow us to achieve these goals in new and creative ways. religion data sources and users when arda was initially conceived, the 1995-96 icpsr guide to resources and services reported on more than 40,000 data files from over 3,000 social research studies. even a topic such as education, which had comparatively few entries, reported 119 data files from 65 studies, with 34 of these studies being conducted since 1980. by comparison, the subheading of religion reported only 9 data files from 9 studies, with only two of the studies being conducted after 1980. yet, this paucity of archived data on religion does not mean that data are not being collected. over the last 10 years alone, lilly endowment has funded over 150 grants with a data collection component, the pew charitable trusts has funded several major national and international surveys, and many denominations support research divisions that collect large amounts of data each year. unlike, education, health care and other substantive areas, where most studies are funded by government sources, nearly all of the data collections on religion are funded by private endowments or religious organizations. most funding sources have either wanted the data to remain “in house” or have not required principal investigators to place the data files in a public archive. in the mid-1990s, however, the lilly endowment began a major initiative for improving dissemination. one component of this initiative was the american religion data archive. after awarding a planning grant to roger finke in 1996 to study the feasibility of starting a religion archive, lilly endowment funded the start up and operation of the arda from 19972000. recently, they extended the support until 2003. thus, the funding sources for the collection and archiving of data on american religion remain private sources. the arda currently holds 120 data files and the number should approach 150 by the close of 1999. these studies include national samples of the united states and canada, regional samples of selected communities or areas, and samples of selected religious groups or professionals. although all surveys include the topic of religion, the survey items span a wide range of other topics (e.g., from involvement in small groups and politics to attitudes on race relations and professional development). in addition to the surveys, the arda also distributes data on american religion by ecological units, such as the association of statisticians of american religious bodies’ data on churches and church membership by counties and states for 1980 and 1990. both the size and the diversity of the collection will continue to grow. once established, the greatest challenge for the arda was appealing to the diverse audience interested in american religion. initially, we were most aware of the social scientists from research universities who frequently conduct and report on the major data collections. for this group, the arda was a valued repository of past data collections and a source of new data for future research studies. but this audience, often sophisticated in research methods and statistics, represents only a small portion of the total audience. many, and probably most, of our users have little background in the social sciences and are not located at research universities. instead, many are faculty members and students located at small universities, colleges, and seminaries that previously had little access to http://www.thearda.com 12 iassist quarterly data on american religion. based on our web site reports, seminaries have made more contacts and referrals to the arda site than any other type of educational institution. and, though we have no record of individual users, our most frequent e-mail and telephone inquiries come from journalists and students. several instructors have informed us that they have incorporated arda data files and software into class assignments. rather than limiting access to a small group of researchers, arda has democratized access to the data, and a very disparate audience is taking advantage of this access. below we review how we appeal to this disparate audience as we strive to achieve standard archiving goals. goals of the american religion data archive the goals of the arda are similar to those of many other archives. arda was established to: 1. preserve data 2. improve access to data 3. increase the use of data 4. allow comparison across data files to achieve these goals we combine proven archiving practices with new attempts to serve a diverse audience. the first goal, preserving data, is the foundation of virtually all archives, and in the case of data on american religion, it was the most essential. of the first 150 data files we received for the arda only three were previously held in a public archive. preparing the data for the archive follows many of the same procedures developed by other scholarly archives. after we receive the data files, we verify the accuracy of the data by comparing our variable frequencies with those of the principal investigator and we begin collecting summary information, or metadata. for each of the files we offer a brief abstract of the study and we provide information on the number of cases, number of variables, the year it was conducted, sampling techniques, sources of funding, principal investigators, collection procedures, any related publications and additional information on the construction of indices or the use of weight variables when appropriate. in our effort to “democratize” access to the data, however, we have gone beyond the standard procedures used to prepare data files for scholarly research. we have added a couple features that make the data files more accessible and easier to use. first, we recreate the original survey instrument within the data set. using the original questionnaire, we record the complete variable description and all response categories. users are not forced to keep a codebook by their side to decipher variable names or truncated descriptions. moreover, when the files are downloaded as microcase files the entire survey wording remains.1 second, we have designed the web site so users are forced to review the metadata before they download files, and they can easily link to the metadata whenever they are reviewing questions or data from the file. this is handy for experienced researchers and essential for those with less experience. improving access to data, the second arda goal, was primarily achieved by adding an easy download feature to the site. thanks to the support of the lilly endowment, anyone with access to the internet can download the data free of charge. once users find a data file they want to use, they can easily download it to their own pcs as an spss, ascii or microcase file. they also have the option of downloading a codebook without the data. once again we have added a feature that allows the data to be used by non-specialists. microcase corporation’s statistical software, explorit, can be downloaded free of charge and is fully compatible with the microcase data files available from our site. the explorit software is used by thousands of social science students each year throughout the united states and canada and is remarkably easy to use. the explorit version offered from the arda site holds fewer statistical options than the version typically distributed for classroom use, but it offers an important option for non-specialists who do not have a statistical package readily available.2 many professors have found this to be an especially attractive option for their students. for the third goal, increasing the use of the data, we wanted to allow users to conduct basic analyses of the data files on-line. yet, from our own classroom experiences with undergraduates, we knew how confusing bivariate cross-tabular analysis can be for those not familiar with statistics. first, constructing the table requires students (or any user) to fill in boxes that ask for an independent and dependent variable — unfamiliar and unfriendly words for most. second, they must select variables with an appropriate number of categories. for example, when a student tries to set up a table with age by income, the resulting table might offer an incomprehensible 70 columns and 20 rows. and, even if they are successful in constructing an appropriate table, they need to know which way to percentage the table. choosing to percentage in the wrong direction leads to meaningless or often misleading results. we have avoided this quagmire by working with microcase corporation to develop a simplified version of their auto-analyzer for our web site. when users find a question of interest, they can click on a button called “analyze” and tables are constructed using preset variables. the tables are percentaged correctly and typically include standard demographic variables such as age, gender, income, marital status and education. this avoids the potential problems of choosing an independent and dependent variable or deciding which way to percentage.3 for example, if a question is selected that asks “which spring 1999 13 party’s candidate would you be most likely to support if a federal election were held tomorrow?,” the user would first see a table summarizing the number and the percentage of respondents who would vote for each candidate. then a series of tables would follow, showing how these percentages and numbers vary by age, gender, income and so forth. the user has received a series of meaningful tables on the question of interest, without struggling through a series of commands. the fourth goal, allowing comparisons across data files and over time, is achieved through standard searches. the user can search for a topic of interest within a single data file, a selected group of data files, or all arda data files. after locating questions of interest, the user can quickly compare the results for each question by conducting on-line analysis or they can compare the data files from which the questions were selected. thus, users can quickly compare similar questions to see if they offer equivalent results, and they can review information about the data files to better understand why the results might differ (e.g., the samples might vary by location, time or religion). once users receive questions from their searches, they can also place the questions in their own question bank. in other words, they can start saving questions for their own survey. during the planning phase of the arda, we were encouraged by prospective users to establish an archive of questions as well as an archive of data. because the complete survey questions are entered and stored in the data file, however, the data file represents a complete record of the survey instrument. hence, when data collections are submitted, arda serves as an archive for the questions used and the data received. by combining the question bank feature with the search feature, the arda becomes a rich resource for constructing a new survey as well as using a previous one. summary we recognize, of course, that arda’s initial efforts to democratize access to data are simply that: initial efforts. still, we are encouraged. the support of the lilly endowment has made the archive possible and has eliminated the barrier of financial cost for using the data. the availability of downloading microcase’s explorit software and using their on-line analysis tool has greatly reduced the barrier of data analysis for a larger audience. and, providing data files that offer complete question wording, detailed metadata, verified data, and muliple download formats, renders a rich resource to the experienced and inexperienced user alike. reducing each of these barriers, and extending the services offered, has helped to increase the use of the data and expand the diversity of the audience using the arda. finally, we want to close with a gentle reminder to ourselves and others. democratizing access to data and metadata are noble goals made possible by recent advances in technology. yet, we should remember that metadata are often an empty promise unless the data are available; and, easily accessible data can still be useless (and misleading) unless they are carefully conducted and prepared data collections. a data archive will still be judged by the quality of the data it provides. hence, just as evangelists close each revival with an invitation to submit to the message just heard, we end each essay and presentation with an invitation for submitting data. if you have data on american religion to submit, or you know of data that should be submitted, contact the arda (archive@sri.soc.purdue.edu) or download a submission form from our web site (www.thearda.com). 1 due to the character limitations of spss for variable descriptions, some of the questions will be truncated when spss portable files are downloaded. 2 the simplified version of the explorit software, downloaded from the arda web site, provides univariate statistics with the appropriate bar graphs and pie charts, crosstabs with the appropriate statistics, and a complete list of survey questions that can searched for a topic of interest. 3 if the variable has too many categories for constructing a table, the user receives a message with this information. * paper presented at the iassist conference, may 17, 1999, ryerson polytechnic university, toronto, ontario. roger finke, jennifer mckinney and matt bahr, the american religion data archive, department of sociology, 1365 stone hall, purdue university, west lafayette, in 47907-1365, 765-494-0081, archive@sri.soc.purdue.edu, www.thearda.com mailto:archive@sri.soc.purdue.edu http://www.thearda.com mailto:archive@sri.soc.purdue.edu http://www.thearda.com vol28-4.indd 12 iassist quarterly winter 2004 by linda f. powell1 data archiving at the u.s. central bank abstract as the central bank of the united states of america, the federal reserve system consumes vast quantities of economic, financial, and organization structural data. these data are used for making monetary policy, conducting banking supervision, performing economic research, and implementing consumer protection policies. the focus of this paper is on micro data archived at the board of governors of the federal reserve system. the paper discusses the types of data used by the central bank, how data are collected and edited, data documentation and meta data, and data purchased from commercial vendors. the paper discusses the challenges faced by archiving a diverse pool of data including communication and coordination, user access across various computer platforms, and meeting the diverse needs of a variety of end users. finally, the paper discusses some of the solutions to the challenges faced and how technology is facilitating the growth of data archiving. introduction as the central bank of the united states of america, the federal reserve system (frs) consumes vast quantities of economic, financial, and organization structural data. these data are used for making monetary policy, conducting banking supervision, performing economic research, and implementing consumer protection policies. the federal reserve system is comprised of the federal reserve board of governors (the board) and twelve regional federal reserve banks (reserve banks). the focus of this paper is on micro data archived at the board. however, much of the data archived at the board is supplied by and used by the reserve banks as well as board staff. both micro and macro data are housed at the board. micro data refers to institution level data whereas macro data refers to sector, industry, or economy-wide aggregated data. industrial production, which is a principal indicator of economic activity in the united states’ industrial sector, is a good example of the board’s use of micro and macro data and how they interrelate. the board receives input to the industrial production indexes from a variety of sources including sample surveys of independent firms, trade organizations, and other agencies. the micro data are weighted and aggregated to generate the macro data series. micro data collection, editing, and storage the board acquires micro data through a variety of methods. the board purchases data from independent vendors for information on sectors such as commercial interest rates, stock and bond market prices, and nonfinancial industries. the board also receives some micro data from other regulatory agencies. finally, the board conducts surveys to collect data within the financial services industry. over 60 surveys of varying frequency are currently collected. some surveys consist of a sample of entities from a population (such as commercial banks). the population is then estimated from the sample and the macro data are produced. relatively few surveys are a periodic census of a population such as the consolidated report of condition and income for a bank, commonly known as the call report. the federal reserve system maintains an ongoing census of the banking industry’s structure by collecting data on structure changes either directly from the bank, the bank holding company, or from the bank’s primary regulator. the federal reserve system’s structure system contains descriptive, geographic, and ownership information. it also identifies events such as mergers between depository institutions and bank holding companies. the structure system includes all bank holding companies, all banks and their branches, all thrifts, and all credit unions in the united states. it also contains some foreign bank and other financial and economic sector data. because so much micro data are collected directly from businesses, the federal reserve system has a large infrastructure to collect, edit, process, and distribute data. although the process varies slightly among surveys, the majority of surveys follow the process outlined in figure 1. the collection process begins with the reporter (usually a financial institution) transmitting the data requested on a form to the responsible federal reserve bank. the transmission processes range from reporters mailing or faxing paper forms to secured electronic data transfers that load the data directly into the board’s editing system, depending on the survey and reporter. iassist quarterly winter 2004 13 the editing system is a data repository designed to quickly and accurately edit massive2 quantities of data. the system is parameter-driven and new data surveys and edits can be added relatively quickly. four types of edits are performed in the editing system; validity, quality, intraseries, and interseries. validity and quality edits confirm that the data within a transmission are consistent and do not violate any standard rules. for example, validity edits ensure that all required fields are not null and that impossible scenarios (such as a negative number of employees) do not occur. quality edits compare the data within the transmission to ensure that the data follow accounting (or other) rules and generally make sense. for example, an accounting rule for a balance sheet is that assets = liabilities + owner’s equity. since rounding can create some slight inequalities tolerances are usually associated with each edit. the tolerances can include percentage differences and/or unit differences. intraseries edits look for consistency between periods and over time. for example, if a bank reported a 1000 percent increase in total assets over one quarter it is likely that extra zeros were erroneously added to the submission and the data should be reviewed for accuracy. intraseries edits also have percent and unit tolerances associated with each edit. since there are some reports that collect similar data but on different frequencies interseries edits are also performed. these are edits that compare similar data items on different surveys for comparable dates. for example, the bank’s weekly report of deposits collects some information that is similar to the quarterly call report which contains some deposit information. the deposit information on the two surveys can be compared to ensure that they are comparable. once the data have been edited they can be stored in the financial data repository (fdr). however, for some surveys there is additional processing that is done to the data such as aggregating data, deriving commonly used data items, and incorporating additional data already housed within the federal reserve system. after the processing is complete the data are stored in the fdr db2 relational database. the banking industry structure data are processed similarly but in a different collection and editing facility because the nature of processing and editing structure data is different from that of financial data. the structure data repository is also in db2 but the collection, transmission, and editing are done on microsoft sql servers. the structure data are broken into different databases to capture attribute (name, location, entity type), relationship (parent, ownership), and merger (acquirer, type of merger) information. the edits figure 1 14 iassist quarterly winter 2004 in this system ensure consistency within the databases and adherence to a predefined set of rules for displaying the data. one example of an edit is that the attribute data for an entity must end if the entity is acquired in a merger. micro data are acquired through a variety of additional methods to meet the special needs of individual surveys. these other methods include contracting with research organizations, direct mailings of surveys, purchasing commercial products, and supplying custom software. the method of data collection and processing used depends on the type of data, frequency, population being surveyed, and other special needs of the survey. how these data are stored and accessed also varies depending on the method of acquisition, user needs, and license restrictions. micro data documentation because the central bank collects data ranging from bank balance sheet values to kilowatt hours generated, the need for centralized and comprehensive documentation became apparent in the early years of data archiving. for the surveys collected and stored via the method described in figure 1, there is an on-line dictionary that defines the characteristics and content of each of the surveys. this dictionary is called the micro data reference manual (mdrm) and contains descriptions of the surveys as well as meta data for each data item stored. the data are organized by survey and consist primarily of financial and structure data. the mdrm documents the labels and values (meta data) associated with each data item in a survey. a web interface is used to access and display the mdrm meta data. for the collection and storage process each survey is given a mnemonic such as edds (the report of deposits). each accounting concept or data item collected is given a number such as 2200 (total deposits). because comparable accounting concepts are collected across various surveys the same number is used for all comparable data items regardless of the survey. combining the survey mnemonic and variable number references a specific data item within a specific survey. the combined mnemonic and number is commonly referred to as the mdrm number. for example, the mdrm number svgl2170 refers to total assets on the thrift report of condition and cusa2170 refers to total assets on the credit union report of condition. within some surveys the same accounting concept may be collected several times but for different populations or periods. in these cases, multiple mnemonics (a.k.a subseries mnemonics) can be used for one survey. an example of this is in the edds survey, total deposits are collected weekly for each day of the week. to identify which day of the week the data apply, the edds mnemonic is broken into edd1 to represent tuesday’s data, edd2 to represent wednesday’s data … edd7 to represent monday’s data. the mdrm has three main components; the reporting forms, the mnemonics information, and the data dictionary. the reporting forms section is a historical pdf library of all the forms and instructions used to collect micro data. the forms are the visual representation of what data should be reported and the instructions are provided to give detailed information about who, how, when, and exactly what to report. the mnemonics section describes each survey, provides a hierarchy of subseries mnemonics, and links the mnemonics to the published reporting form names. the data dictionary is the heart of the mdrm and defines each data item collected on any of the surveys. for each data item the mdrm provides the starting and ending dates it was collected for each survey. it also provides a long caption, a short caption (similar to the long caption but limited to 40 characters), a confidential indicator, a data type (financial, structure, ratio, or derived), and a long description which provides the detailed instructions, history, and idiosyncrasies of an accounting concept between surveys. the web interface displays all surveys associated with a specific data item number or all data items within a survey. each survey also has a glossary which identifies unique information pertinent to the survey such as edd1 represents tuesday data. the documentation of data purchased from vendors or collected through contractors varies between surveys and sources of the data. vendor commercial packages generally have user guides but there is not currently a central documentation facility. macro data collection, storage, and documentation most of the economic macro data used at the board is housed in fame databases which are designed to store large volumes of time series data. as with the micro data, the macro data are collected from a variety of sources in a variety of formats ranging from pdf to delimited files. the sources include board staff’s aggregations of micro data, other government agencies, research organizations or universities, private organizations, and commercial data providers. once the data are received they are processed and renamed, using a board nomenclature, and stored in a temporary database. while in the temporary database, several quality edits are run against the data. the edits look for nulls, excessive variability over time, and noncontiguous date problems. the edit routines run against the data are standard for all series with a few minor exceptions. once the quality of the data is confirmed the series are loaded to the production database. the macro data used for most official forecasting are stored using a hierarchical nomenclature in fame databases. the main fame databases hold either us or international data. iassist quarterly winter 2004 15 to get to a specific series, users can ‘drill down’ through the nomenclature to retrieve the desired time series. for example, the us database includes income data for which the first letter of the nomenclature is ‘y’. within income there are several categories including personal (‘p’) which also contains several categories. within personal income is disposable outlays (‘d.o’) which has the full nomenclature of y.p.d.o. these data can be viewed at this level or they can be further disaggregated and viewed. in the international database the country is denoted by a suffix, such as .uk. the fame database allows for some self-documentation and each series has several attributes that provide descriptive data. the attributes include information regarding where the data originated, where the data are stored in fame, board contacts, periodicity, adjustments to the data, units, unit multiplier, currency, update frequency, and several attributes that describe formulas. the nomenclature shows the relationship to parent series. because hundreds of thousands of macro data series are stored in fame, there is also the need for data documentation. to aid in the documentation and navigation of macro data several web-based tools are available at the board including a source book of nonfinancial economic data. the source book lists the various sources of data, what data are received from each source, data definitions, the data collection process, and adjustments to the data, as well as information specific to the source. challenges and the evolution of data archiving as the need for more and varied data grows, we have encountered new challenges in the collection, storage, access, and documentation of data. twenty years ago, the majority of micro data used at the board were collected and stored in a model similar to figure 1. today, the reliance on market information and data obtained through other sources eclipses the volume of collected data, forcing users to spend more time researching what data are available and learning how to access the data. catalogue one of the greatest challenges is the cataloguing of all data purchased at the board. as in academia, there are often several individuals independently studying different aspects of the same industries. these different studies can be performed in unrelated areas of the board but have similar data needs. for example, when citicorp bank holding company and travelers insurance company merged, the supervision community was interested in insurance company micro data to determine the effect of the merger on the bank holding company’s safety and soundness and the banking industry in general. simultaneously, the economic research departments are interested in the insurance industry’s impact on the economy and financial markets. as with most corporate, academic, or government bureaucracies, budgets and responsibilities are dispersed throughout the agency. as the volume of data purchased has increased, ensuring that similar data are not purchased multiple times has become increasingly difficult. to address this challenge, the board has created a “data and news catalogue” of purchased data. the catalogue contains information for each data purchase including the database, the vendor, internal contacts, the form of access, and license information. it also contains a brief description of the data purchased and links to the vendor’s website or the database if available. the catalogue was originally populated by surveying the budget areas to identify all data expenditures. the catalogue is maintained by soliciting information from the budget areas as well as by having some areas update it directly. an annual review during the budget cycle catches additions or deletions not otherwise captured during the year. the board also maintains several ongoing databases, such as merger adjusted balance sheet data for banks. the time and effort to create and maintain these databases are extensive. to avoid possible duplication of effort or purchasing of data already captured in-house, the catalogue also describes ongoing databases created and maintained by board staff. once users know what data are available, the next step is accessing the data. storage and cross referencing much of the purchased data are stored within vendor software and can only be accessed via the vendor software. oftentimes the data can be exported to a sas dataset or excel file that can be further manipulated and analyzed or combined with other data. many license agreements limit the way data can be stored and the usability of vendor software limits the ability to standardize the storage and documentation of micro data. the volume and complexity of purchased data also make it impractical to try to put all micro data in a standard storage repository with documentation such as the mdrm or fame nomenclature. the uniformity of micro data storage is just beginning to be pursued but is hampered by license agreements and data availability. similar licensing problems are encountered on the macro data level and in some cases licenses need to be negotiated to allow for the loading of data into fame. data consistency over time is another challenge. as users’ needs and the economy change, so do the data. for example, the calculation of goodwill has changed over the last decade as accounting rules and regulatory requirements have changed. in addition to changes over time, there are often different names for the same accounting concepts across industries such as capital vs. net worth. to address this challenge, wherever possible, micro data items with like meanings are given the same numbers in the mdrm 16 iassist quarterly winter 2004 or new data items that can be used over time are derived from the existing data. in addition, slight changes to micro data items that result from changes in the markets or regulatory changes are generally captured through notes to the description rather than giving the item a new number. changes that have a large immediate impact require new numbers. for macro data, the fame nomenclature enables historical series to be link with current series. for example, when the u.s. census bureau changed from using standardized industrial classifications (sic) to north american industrial codes (naic) there was a one-toone relationship between a number of the sic and naic codes. therefore, series for a sic code that had a direct relationship to a naic code were flagged as historical series to the naic series. researchers also have increasing needs for more data. this is particularly prevalent in the macro data series. specifically, there is demand for geocoded data including an increasing demand for geocoded housing data. for example, there is a desire to be able to drill down from us data (1 series) for a specific series to specific regions (4-12 series), states (50 series), msas, and counties. this type of geocoding on a large number of series can have an exponential effect on the volume of data stored. in addition to needing more data, researchers have an increasing need for cross sectional data. as technology improves and sources of data grow, the availability of market data has increased. purchasing data from a vendor also adds flexibility to a research project because the lead time is short and there isn’t an investment in infrastructure, so changing the sources of data is easy. in recent years, evaluating stock and bond market data in conjunction with corporate structure and regulatory accounting data has become increasingly feasible. these different types of data can come in a variety of formats and from various sources. linking the data between sources is challenging because it is difficult to join data from different vendors or in-house databases. to address this challenge, the board is currently evaluating compiling a database of key fields. the board’s entity identifier (id_rssd), the stock market ticker, and the bond cusip number are some of the key fields being included in the date-sensitive database. once complete, this database will allow users to join data from different sources based on the key identifiers and date of the data. technology advances the introduction of extensible markup language (xml) is allowing the data archiving process to evolve in new ways. xml brings the ability to tag individual concepts (text or numbers) with context so data can be searched quickly and accurately. for example, the word ‘bank’ refers to a financial institution or the side of a river. xml enables each word to be given a different tag that will signal to a search engine which meaning the word has in the context in which is it being used.3 xml is valuable as a transmission protocol because it allows for sending and receiving large quantities of data. to further facilitate the transmission of data, several groups are developing transmission standards. the accounting industry is focusing on the xbrl standard which is designed for financial statement data. another transmission standard for statistical data is the statistical data and metadata exchange (sdmx). the board is currently implementing production systems using each of these transmission protocols. xbrl is being used for an interagency project between the three primary u.s. bank regulators to collect bank call report data over the internet. sdmx is being used for the downloading of macro data in a new bulk data download facility available on the board’s website. conclusion the international and domestic economies, financial markets, and economic models continue to become more complex requiring more data to evaluate the growing complexities. the increasing number of sources of data and the increasing volume of data add to the challenges faced by data librarians and other data archivists. advances in technology, such as xml, help to facilitate the resolution of some of these challenges but at the same time create additional technology hurdles. to address the challenges that lie ahead, data archivists need to ensure that they have organized, flexible, and well-documented data archives. meta data, naming conventions and nomenclatures, and other data documentation are becoming more important to describe large and diverse data archives. data quality verification and definitional consistency are necessary to ensure the usability of the data archived. maintaining meta data and complying with nomenclature rules are timeconsuming and tedious tasks, but if you have a large group of users over a long period it will save resources and avoid frustration. finally, centralizing data of various file types can be difficult and time-consuming but ensures that data are not lost and reduces the burden on end users needing to know how to use a variety of data access tools. bibliography bayard, kimberly and morin, norman. “industrial production and capacity utilization: the 2003 annual revision.” federal reserve bulletin winter 2004. 20 december 2004. http://www.federalreserve.gov/pubs/ bulletin/2004/winter04_ip.pdf. cannon, sandra, chief of economic information management. federal reserve board of governors. d.c. personal interview. 3 january 2005. fame time series database. 2003. sungard data iassist quarterly winter 2004 17 management solutions. 20 december 2004. http://www.data.sungard.com/infrastructure/fame/index. htm federal reserve statistical release g.17 industrial production and capacity utilization. 14 december 2004. board of governors of the federal reserve system. 20 december 2004. http://www.federalreserve.gov/releases/g17/current/ default.htm wallison, peter j. “enhanced business reporting gets a start.” the new republic magazine, december 20, 2004, pp on/1 – on/4 endnotes 1 linda f. powell, board of governors of the federal reserve system, washington, d.c. email: linda. powell@frb.gov. 2 data from over 2,500 bank holding companies and over 18,000 depository institutions, each of which reports hundreds of financial statement data items, are processed and stored at the board each quarter. this is in addition to other weekly, quarterly, and annual reports that follow the process outlined in figure 1. 3 p. j. wallison, “enhanced business reporting gets a start,” the new republic, 20 december 2004, p. on/3 38 iassist quarterly end user searching and data: the graduate business resource center experience by rona ostrow' deputy director graduate business resource center baruch college city university of new york introduction the graduate business resource center of baruch college, c.u.n.y. is a technology based facility located in close proximity to the school of business but physically separated from the library. it is an attempt to serve a population of approximately 3000 graduate students and faculty in the school of business and public administration without having an actual graduate level business library. since there are so many of them, and so few of us, we have emphasized enduser applications wherever possible. moreover, we feel that part of our mission is to prepare the graduate students for 'presented at the international association for social science information service and technology (iassist) conference held in washington, d.c., u.s.a. on may 26-29, 1988 the "real" world of business they'll discover upon graduation. we aim, therefore, to make available to them those data sources which they are likely to encoimter within the business community or to which they can request access once employed by a firm. our goals are to use access points that are as user friendly as possible and to teach the students where and how to obtain the data. we do not attempt to analyze the data, but do make available such software programs as lotus 1-2-3 to help them analyze it on their own. of course, many of our constituents, particularly faculty members, have complex data needs beyond those of the average student/enduser. these needs are referred to otir data resources service, headed by professor siman and our computerized information services, headed by professor lowe. since both of my colleagues have already addressed to you, i'd like to concentrate on those areas of our service which do indeed focus on enduser and user-friendly services and which meet the needs of the average data user at this level. many of our constituents need data that is readily available to them once they know it exists. for example, our students often request such data as demographics, market share, market segmentation figures, advertising reach, financials, and usage of materials in production. our role in the gbrc is to publicize the available data, emphasize new user-friendly methods of access, and teach the endusers how to meet their own data requirements. spreading the word publicizmg the available sotirces is a major function of our center. the gbrc publishes newsalerts several times each semester to alert both students and faculty to new data sources. summer 1988 iassist quarterly 39 programs, demonstralions, and services. (see figures #1 and #2). by bringing new datafiles and services to the attention of our constituency, we are trying, in essence, to "create a market" of potential users. typically, our users may not even be aware that such data exists and, if they are, may not realize how they can use it in their research. some newsalerts focus on new datasets available through the data resoiuces service while others aimounce workshops and demonstrations of data available through commercial vendors via online information retrieval. teaching this, in turn, brings us to the second major function of our program. once we have interested our researchers in the available data, we then proceed to introduce them to il we let them try it out for themselves so that they will learn how to access data without the intervention of an intermedian . since enduser searching is becoming increasingly the norm in the business community, we want our students to feel confident that they can access the information the\ need long after they've completed their degrees. in addition to the more or less formal training obtained through our seminars, we also offer point-of-use assistance through access guides (brochures we prepare to guide endusers) and one-on-one help in the person of trained professional consultations to assist each researcher in the selection of the most appropriate data sources available both at baruch and elsewhere. availability fiiully, to ftilfill the third part of our perceived mission, we bring the data to them in as user-friendly a format as possible. this has, imtil recently, meant that we provide the students with after-training access to the dow jones news/retrieval service (which is available to us on a prepaid monthly basis as an educational institution) and prosearch (which provides user-friendly access to both the dialog and brs information services). for budgetary reasons, the latter must be monitored closely and hmiled to research for theses, dissertations, and articles for publication. more recently, the advent of cd rom technology has allowed us to make several of these same data sources available to a much larger audience since they, too, are prepaid. although our facility is still small we are able, through cd roms, to make the disclosure database (including spectrum ownership) and standard and poor's corporations available to our students and faculty along with such bibliographic databases in cd rom format as psvclit abl/inform. and eric. dow jones news/retrieval service one of the most pressing needs for data we have at the gbrc is for company financial information. our students are constantly on the lookout for income statements, balance sheets, key ratios and the like. one of the easiest ways to make this data available to them is through the dow jones news retrieval ser,'ice . we have ananged to prepay for the service and have configured our pcs to automatically dial up, connect, and enter a password through smartcom. through dow jones news , our students have access to disclosure (including extracts from over 10,000 publicly held summer 1988 40 iassist quarterly company's 10-k repon and other sec filings see figure #4), media general (detailed corporate financial information on approximately 4,300 companies and 170 industries). standard and poor's (which provides brief profiles of over 4,500 companies including earnings, dividend and market for the current and past four years). historical quotes (including daily volimie, high, low and close for stock quotes and composites). historical averages , and current quotes, among others. our students make particularly good use of dow jones news' "quicksearch" feature. by simply entering "//quick" and a company's ticker symbol, they can immediately key into a wealth of financial information drawing from multiple dow jones news/retrieval files (see figure #5). data includes current quotes, latest news stories, a financial and market overview, earnings estimates, company profiles, and investment research reports, (see figtire #6). the students may print their dow jones news/retrieval results and/or save the information to a diskette for future use. prosearch among the business, economic and demographic databases most in demand at the gbrc are the pts family of databases (forecasts. times series , and annual reports abstracts as well as the bibliographic database promt excellent for market share information, and mars useful for targeting advertising audiences). in addition, students and faculty make great use of donnelly demographics and cendata for additional demographic information (see figure #7). in order to use prosearch . the researcher merely highlights the desired category and subject by using the up and down keys and the return key. by the way, all of this is done while offiine, so no charges are accumulating for typing time. once the researcher has selected a subject, he or she selects a database from the "catalog cards" screen (see figure #8) and enters the search request on a grid designed to enter fields with the mere touch of a key. the researcher uses boolean connectors in the usual way and. when the request is complete, cormects automatically to dialog or brs by simply hitting the f5 key. although we often send our students to the library to use the pts print sources, their ability to use boolean logic easily through prosearch greatly enhances their chances of finding the exact statistical or demographic table they need. ehalog and brs databases are also available at the gbrc. to make searching available for our endusers, we provide prosearch software. the software is very user-friendly. we do, however, require both students and faculty to attend a short training seminar before beginning to use it. we also provide some back-up support in the form of trained graduate student assistants who also monitor usage (this costs money per connect hour, per record downloaded or printed, etc.). within a very short period of time most graduate students (and even most faculty) become proficient at accessing the databases with minimal assistance. graphic presentations we have also begun to make software available for out students to analyse and present the data they've found. although we do not teach the use of lotus 1-2-3 and other statistical packages, we do have a copy available for knowledgeable students to use. we also provide access to a plotter, laser printers, and a scanner which our students use to create graphic represnetations of their findings (see figure #9). summer 1988 iassist quarterly — 41 cd roms al present, the gbrc provides access to both standard and poor's corporations and disclosure in addition to bibliographic databases in cd rom formal once again, having this prepaid data allows us to make it available to a much wider audience. instead of being limited to online searching for thesis, dissertation, or publication purposes, all our constituents may access this tjusiness and financial data at any time. we hope in the near future to enhance our capabilities through the acquisition of a new cd rom product, lotus one source which contains both the data our students need and lotus 1-2-3 on a single laser disk. another anticipated acquisition is batelle's america 2000 software package which will enable the students to make economic projections based on extensive data included in the package. in sum, there is quite a bit of data in the fields of business and economics which can be made available to the novice user who does not require raw data or anything very sophisticated in the way of data manipulation. the way to get this data to the user is threefold: publicity, teaching, and availability. through our publications, training sessions, and consultations, we at the graduate business resource center try to give our students access to the data they need and, more importantly, the skills to continue to meet their own requirements for information.n summer 1988 42 [assist quarterly illustration 1 baruch cor.lf.gn library g r b c graduate businrss resource c e n t e i? news alert october 13, 1988 telephone: 725-7114 editor; roha oitrow gbrc announces fall programs, services and hours welcome 10 the (alm988 lomeilefl tlie graduale business resource center plans a leiles of woik shops <iiu new services designee! to meet specific research needs workshops online inforniation for endusers learn to access dialog dalabd^es online without the assistance ol a prolessional liitermejlar/ participants receive training, hands-on operlence, and bcivsi to qnllije inlomiallon, pioleisor id* lowe will conduct all seminars in booit) 1 jj4, 3^0 parlt avenue south, yv^rkshopsjre o^eillo al l baruch f^curfvaii^qfaduaiestu^enis, dates andtlmci friday, october 21, 1968 10:u0a.m12:00p m friday, november 4, 1988 10:c(ja.m.ij'ioqp m. thursday. decembers, j988 3:u0pm.5:00 p.m. to replster, please complete the attached application and mail to: professor tela lowe, gbdc, box 2g2. or phono 725-7) 14. fall hours monday-thursday 9 30 a.m. • 9:u0 p in friday 9:30 a in • <1:00 p >ii. saturday 2:00 p in. • g;uu p in the gbftc is not open when llie college is closed on days when there are no ciassl-s, but the colletjb is open, the glutc will close fit <1:uu p.m. the l!!/p'i'iii'ijijj.abyviliijgl^t_5(e5ejl!, bg qlk'ii on salurdjys. cd rom demonstrations the newest innovation in information technology is the database in cornpact optical disk format (cd roms). now, without professional assistance, any lesearcher may access both biblioyraphlc and infcrmational databases iraa ol charge. currently available cd flofvl databases at the gqrc include conipacf discloiute. piycin. abiinform. iric. and standardi poor'i corpo'iatlons. professor ida lowe will demonstrate how cd ftoms can benefit research in a series of progran^s to be givers at the gbnc, room 1224, 360 park avenua south. the procl^a^^s are op^ n to all baruch figiltv and graduate ttujeiut. oaieisnj tfinei gsneritl introduction to cd ftom daiabatet thursday, october 20, 1988 12:30 p.m. -1:30 p.m. dinlness liiforniation on cd rom: abi-lnforn\ andpiyclit thursday, novembers, 1988 12:30 p.m.l:3sv m. social sciences information on cd rom: eric and piyclit thursday, november 17, 1988 12:30p.m.1:30 p m. coiporale direclories on cd ho\f: disclosui» and standard & poor's thursday, december 1, 1988 12.30p.m.l;30p.m. (ove,) summer 1988 iassist quarterly 43 illustration 2 baruch coriegr library g r b c gradual business resource center news alert march 26. 1986 telephone: 725-3301 editors: t. atkins r. ostrow 1985 general social surveys the 198s central socl|il survcyi (css) are hare. produced by the universicy of chlcago'i national opinjon research center (norc), the ccnoral social surveys provide a cross-sectional sample ot the united slates adult population. the' data has been collected almost every yeai" since 1972 (no survey was done in 1979 or 1981). the surveys are based on a 300 question interview which idoniiries respondents' attitudes and opinions oij such issues as ilia family, social mobility, social coutrol, race rclalloni, sexual mores, and nallaaal morale. this high quality data (s used for research and instruction in tnahy fields including sociology, markeling, psychology, and consumer behavior. in |l982 the national opinion research center oversampled the black population, thus providing an excellent dnia set for studying this minority group. this data set is now available in machinereadable formal on computer tape at the cuny/ucc. it is a cuitiulative data file, with each annual turv|e/ contained in a separate subfile. merging all 12 yearly files greatly simplifies the use of the general social surveys for trend analysis of specific questions. subsets of the data may be created for research or instructional purposes. the subsets may be used on the mainframe, either on tape or on disk, or duwnioadi;d for use with microcomputers. special data modules may be prepared for liudeni exercises and classroom work using these files. in order tq access the data, you will need the following information i'or tap* cda 11$; file i is the raw dalafile; file 2 is i (ct of spss control cards; file 3 contains the spss contiol cards for the first half of the data; file a is the spss comrol cards for the second half of the data; and file s is a file of spssx control cards f^ir the data. contents of tape volume c0aii5 //file number dsname recfm lrecl block est. blksiee count feet 1 icpsftoss«'<33.a fb 10 32000 m 317.1 2 icpsr c$s!433 b f8 10 ]i:o j. icfsr.gssjos.c fb 10 3i;o 4. 1cp5r.0ssjo5.0 fb so 3120 ). icpsr.gs2i.i15 e fb 10 si20 summer 1988 44 iassist quarterly niustration 3 id o en o < oq z <t q << uj \— oq jq (x —' co a; to jo a-5 si is iiil 2 o 8 ? 3 9! ^ e p e ll 5 i i s j s ?i5§ " ? s c s • ? c £a i) d p ^ f, "^ " -o ^ ^ o * " ?? j! s * .^ u o o 01 \3 o s "5 «i o c e 2 of^89 e e «. £ 2 ^ j 2|^ 58° 55 e 3 c ?'^ 3 b summer 1988 lassist quarterly 45 illustration 4 figure h: disclosvire i^atabase <st:e<r:nir. cvnershi::) d1sclcsub.z iktshnatiokxi. bosikzss vachines corp ikstitutionxl 1 owne^^s •rmtr 1 2 3 a 5 6 7 a 9 10 11 12 .13. 14 is 16 17 kake wells fxrgo bxnk n.x. morgan j p i co inc bankers trust n y corp college retxre equities bernstein saneord c t co mellon bank corporation michigan state treasttrer new york st cckhon ret. wellington/thorndike ctiase manhattan corp capital guardian trust capital research k mgmt jkt.titance-capital. hgmt state street boston cor? pnc financial corp calif public empl retirk manufacturers hanover tr shares latest qtr filing held change date 9, 5*3,720 -1,205,630 06/30/37 s, 533,000 -271,000 06/30/87 7, 990,373 -253,664 06/30/87 7,,785,100 -124,600 06/3 0/87 5,,773,247 69,055 06/30/87 5,,675,094 -169,164 06/30/8" 5,,496,029 557,000 06/30/87 s,,175,000 -111,000 06/30/8": ,564,513 -709,384 06/30/8: ,311,346 -10,225 06/30/8'; 4 ,068,000 -39,200 03/31/8: ,447,700 30,000 06/30/a: ,321,999 -879,_661 06/30/8: ,295,125 62,003 03/31/8-. ,135,723 -293,795 c6/30/8". ,099,100 -26,500 03/3 1/8-. ,032,505 -177,552 06/30/8summer 1988 46 — iassist quarterly illustration 5 'igure is: "quicksearch." start-up screen. ouicksearch copyright (c) 1987 dow johes & compai^y, inc. an autonated method ^for accessing quotes^ company news, financial data and profile information from eight news/retrieval services press to 1 search by company stock sytubol 2 access a quicksearch help menu or enter as much of the company name as you're sure of and press return ibm *end* press for 1 ibm credit corp. or enter another name or /t for top summer j 988 iassist quarterly 47 illustration 6 figure 16) part of a dow jonea news/ketrieval qulcksearch. doh j01ie3 quick3e/\ncii idm credit conp. piiesa for j'/a 1 curreiix quotes h/a a la'vest hews oh d.iad w/a 3 rihahcial ahd harkkt oveuvii 4 kmiiirnca estimates f n/a 5 cohpahy v3 iiidustry pehformaiici \l/\ 6 income stateheiit3, |lal 6iieets • 7 cohpahy pnorilb 6 ih31l)er tiiadiiic; suhhary h/a » i1ive3theht reaeailcll reports h/a type pllltrr folu)hei) by iteh llohdeha, bkparatbd by commaj, to prjht fltlecteo becriohfl ok u'llt; hepouj.' at k^guiau u3aqb ratu . ' cxamplbi pniirr i,i,» press bel-uru fon ihsti\uctio»is ahd pricihc ihtormatioh. nil disclosure year i»b6 x969 1984 l«t1 •s-yr growth rati rive year summary sales 6^0, 147,000 5^4, 176,000 396,439,000 1^3, 639,000 l|.l,017,000 (i) xlo.t disclosurc ihcoh quarterly report fofll hrr sales cost ok goods 0u033 prokit r&d expenditures bell gen t atjhiii exp inc ber dep k ahort depreciation t amout non-operatino inc iktcrest expense incohb before tax prov for inc taxes minority int income invest gains/losses other incohb net ihc brf ex items ex items ( disc ops net income outetanoind shares 164 110 s4 30 }} k 6tatimbnt 03/31/87 ,993,000 ,689,000 ,303,000 wa ,813,000 ,490,000 ha ha ha ,490,000 ,614, ooo ha lla ha ,876,000 ' 'l" , 876,000 net incohb lis, 118,000 103,033,000 63,691,000 41,670,000 31,374,000 69.4 xbh credit corp eps .41 .46 .4) .40 .39 0.4 ibm credit corp summer 1988 iassist quarterly illustration 7 file 575 donnelley demographics dialog file 575 sample record st. louis citk mo la t ion aa — 453,085 440 198 / ^^^ total popu -^-2.7%—j .0% 423,019 total hous "holds all — 178,048 177 796 172, 327 household population ac — 443,305 430 418—, -2.8% 41 t, 239 average ho jsehold si2 e at -2.5 2.4 -v\-2.6% 2.4 average ho jsehold inc af —$11,712 1980 c $14 860^^r-^ s 1 9 , : 7 9 \ 1964 1989 number percent estimate p ojection total population by age 453,085 100.0% 440,198 423,019 • 5 bt>-—38,447 8.5»—-iil = 6.6%— -cb 8.8% 6-13 bc . -49,814 11.0%^--bm10.7%— -cc 10.9% 14 17 bd =— 30,175 6.7»-^-bij5.7%—-cd 5.3% 18 2« be = —61,143 13.51—-bp= 12.0%— -ce 10. 3% 25 h bf =— 64,774 14. 3«-—-b0= 16.8%— -cf 17. 3% 35 44 bc =— 37,347 8.2«-^-br = 9.6%— -cc 12.5% 45 54 bh =-42,939 9.5«-^-b5= 8.4%— -ch 7.9% 55 64 bj — 46,526 10.7«— -bt = 10. 1%— -cj= 9.0% 65 + bk = -79,920 --*249,293 17. 6i— 100. os -bu= 18.1%241,490-" -ck 18.1% female population by age 230,645 0-5 19,065 7.6» 7.6% 7.8% 6-13 24,568 9.94 9.7% 9.9% 14 17 15,078 6.01 5.2% 4.9% 18 24 32, 362 13. ot 11.2% 9.4% 25 34 33,704 13.5% 16.0% 16.5% 35 44 20,385 6.2% 9.4% 12.0% 45 54 24,043 9.6% 8.5% 8.0% 55 64 $8,068 11 . 3% 10.6% 9.5% 65 » e . 52,020 —^203,792 20.9% 100.0% 21.7% 196,706 . 22.0% male population by ag "tttriti^0-5 19,382 9.5% 9.7% 9.8% 6-13 25,246 12.4% 12.0% 12.2% 14 17 15,097 7.4% 6.4% 5.9% 18 24 28,781 14.1% 13.1% 11 .3% 25 34 31,070 15.2% 17.7% 18.2% 35 44 16,962 8.3% 9.6% 13.0% 45 54 18,896 9. 3% 6.3% 7.8% 55 64 20,458 10.0% 9.4% 8,3% 65 * 27,900 13.7% 13.7% 13.4% median age total popu lat on 31.6 32.4 33.5 median age adult popu lat on 45.9 44.0 42.9 total population 453,085 100.0% 440,196 423,019 white da.—— 242,576 53.5% -di:. 52.2%dj= 49.7% black dl . —— 206,386 45.6%—-[* = 46.8%dk49. 2» othet dc= — -4,123 .9%— -dc 1.0%— dl = 1.1% household income s s 7,499 1 a = 59 999 31 6%t j: 26 4%e5 = 19. s 7,500 s 9,999 cb = -18 776 10 5%fk. 8 6%— et = 6. $10,000 514,999 tc = -30 501 17 1%tl. 15 4%eu= 12. $15,000 $24,999 td= -40 516 22 7%tm = 26 6% ev = 27. $25,000 $34,999 tl= -18 080 10 1%1 m14 4%ew= 20. 535,000 $49,999 f f= -7 382 4 1%k' = 6 0%ex = 10. $50,000 s75.000 1 574,999 rc = -2 i h= 316 873 1 5«^ 1 li: 1 9%ey = 7%lz = 1. 1 . summer 1988 iassisl quarterly 49 illustration 8 figure *8: prosearch selection screens. the database selection screen after you select the high-levd interface, the dacabasc selection screen appears: cat*<7ari«« |>6usinasa caqin««rlng 4 sei sub'^scrs acqulaltions/x«rf adv«rxl«lnq xsaoclacxona accounting: kajtvaao busimtss revtsw accounting: kxsvxilo bosiness axview aceouaciag: ali/ixrorn is l accounting: aai/uffowt ntro abatracts covara all ptiaaaa oc bualnaaa and >anaga»«nt. scraaaaa ganaral inforbatlon appllcaela to many bualnaaaaa and lndustrlaa. including eoapany caaa hlatorlaa, coapatltlva intalllqanca, naw product davalopaant, and daelalon aaking. covara ov«r too primary pu&llcatlona. 1971: monthly updacaa. 229.000 racorda. si.lt/mlnuta $.20/oefllna print s.30/onxina display summer 1988 50 iassisl quarterly illustration 9 co e (0 c oo q c .2 o d) 00 j—" o da o \-^ gl co co q) d o "s^̂ "^^s kwww^ r^3^^^;^3^r^ kwwww 5̂ kwwwwkwwwww^̂ kwwwwww-x iwwwwwwŵ \\\\\\\\\\\\\\\\\\v \\\\\\\\\\\\\n^^^?^^^ ^s^^^^^^^k n h h .r.e^\o\i9r\b'>jj ^15 f p, u k ,35. s »> y •« ^ w <j o •j jc x o ««• o 2 t 3u o o o o 66 « » h i c d x) i • c « i i rh m j • • i3 h p 1 0kni 3 -< < o ) oo • n i <) • <t « 4> ^ .-i -h <h • q o /] o 1 h g y u i( 6 u o 1 i i i i << m <.' o ri 1 summer 75 vol29-4.indd 4 iassist quarterly winter 2005 editor’s notes welcome to the fourth issue of the iassist quarterly vol. 29. the true year 2005 ended some time ago, but iassist is now ready for entering the fourth decade of iassist quarterly. the iq has changed through the years and will hopefully continue to do so in the future. since the beginning, much more iassist communication is now available. besides the printed iq and the yearly iassist conference, we have the iassist mailing list, local iassist meetings, the iassist website for further information (and the iq), and the blog. so browse around at http://iassistdata.org and visit the iassist weblog (blog) iassist communiqué – at http://iassistblog.org. happy surfing. the first article is a paper presented at the iassist 2006 conference in ann arbor, michigan, at the session: “moving beyond data to networked knowledge.” julie lamb, from the department of sociology at university of surrey in the uk, is the author of “disseminating survey information in the networked world: a uk resource.” the paper discusses the development and use of the question bank, an innovative web resource which is used to teach students and researchers about uk social surveys with a focus on large-scale quantitative surveys. currently, the question bank contains the full questionnaires for over 50 surveys produced by agencies such as the office for national statistics and the national centre for social research. the question bank supports reuse of survey questions and the accompanying coding frames, including information about the surveys, and often this is used as a benchmark for new surveys. the second article is a paper that was also presented at the iassist 2006 conference. the paper titled “documenting religion worldwide: decreasing the data deficit,” by brian j. grim, the pew forum on religion and public life and association of religion data archives (arda), and roger finke, pennsylvania state university and association of religion data archives, was presented in the session “compare and contrast: using crossnational data.” the paper is a presentation of the arda with data on 238 different countries and territories and some arda-coded measures. the international social survey programme, world values survey, and the general social survey are among the well-known datasets in the archive. the article not only discusses the archive, but also demonstrates research carried out on arda materials by the authors in constructing and analyzing some religious indexes. i noticed another demonstration implicit in the article. the issp, wvs, and gss also are available from other archives, so we have two typologies of archive materials: 1) the traditional archives that are geographically based (national, regional, university, etc.) and 2) the subject archives, of which the association of religion data archives is an example. some redundancy can be solved with links to central sites, other problems will be more problematic like the version problem of extended metadata being constructed at separate archives. the last article is from the 2006 conference session “the big picture: gis data challenges and solutions.” the paper titled “consideration for information security issues in geospatial information services of local governments” was presented by makoto hanashima from the institute for areal studies in tokyo and the institute of information security in yokohama. the author leads off with how gis has changed from being “geographic information system” to being “geospatial information service.” the gis has serious aspects of it security – mostly because the gis is part of the infrastructure for local government. security is mostly about securing the geospatial information service and protecting the geospatial information as public property, and that the service must not threaten the safety of the public. based on a threat analysis, a set of safeguards for the baseline security of the geospatial information service is selected. the paper relates closely to “iso/iec tr 13335 guidelines for the management of it security (gmits)” and provides an overview of the many points herein. the workflow of the research carried out has the following main points: a) discussion of it security policy regarding the geospatial information service, b) outline risk analysis, c) making of a safeguard catalog, d) making of a baseline security model for geospatial information service, and e) proposal of a baseline security guideline prototype. in this process the paper also contains a list of “specific threats” such as tampering with and forgery of data, illegal copying and distribution of data, and attack by an unauthorized service. the safeguards for “specific threats” of gis will be included in further research. articles for the iassist quarterly are most welcome. articles can be papers from iassist conferences, from other conferences, from local presentations, discussion input, etc. contact the editor via e-mail: kbr@sam.sdu.dk. karsten boye rasmussen, august 2006 iassist quarterly summer 2012 5 iassist quarterly editor’s notes research data in demystified compliance, with “glocal” indexing, and archival development welcome to this volume 36-2 2012 of the iassist quarterly (iq). this editorial is written in october 2013, so you might rightfully wonder what time zone iq is in. yes, we are truly sorry that there have been hiccups in the process. however, the good news is that we have several issues in the pipeline being compiled by guest editors. with the many efforts of these busy people we hope to catch up on the production schedule with these coming special issues. now, however, we have three articles in this issue with investigations emanating from institutions that we term as “data archives” or “data libraries”. natascha schumann and astrid recker from the data archive for the social sciences, at the gesis leibniz institute for the social sciences in cologne, are demystifying some central issues on data archiving in their paper “de-mystifying oais compliance: benefits and challenges of mapping the oais reference model to the gesis data archive”. papers for the iq often evolve from presentations at conferences, as did this one, being updated from the iassist 2013 conference session “beyond bits and bytes: the organizational dimension of digital preservation.” the authors are exploring the use of oais (open archival information system) as an abstract reference model as it is being mapped to the gesis data archive. “welcome to the real world!” the authors focus on how the oais is considered to be a language and as such does not present a concrete solution. oais has since the late 1990s shaped and influenced the digital preservation discourse. the gesis data archive followed in the footsteps of icpsr and the uk data archive when it tested “oais compliance”, when investigating the functions of the archival information system by mapping processes to the oais. the investigation presented in the paper will be of benefit to others working at data archives and similar institutions. by the way, the authors conclude that “oais compliance” is not simply a yes/no question. compliance is complicated! also from gesis – leibniz institute for the social sciences comes the next paper “thesaurus-based indexing of research data in the social sciences: opportunities and difficulties of internationalization efforts” by katrin baum and andreas oskar kempf. internationalization and standardization are areas supported and enhanced by iassist. the authors cite the data documentation initiative (ddi) as an example of an international standard for describing data, facilitating international data exchange. before data can be exchanged it has to be identified. the authors are investigating international indexing which enhances the precision of searches for data materials. the paper highlights some international organizations offering support for indexing, for example through descriptions in the “european language social science thesaurus” (elsst) and the “topic classification” used by the member organizations of cessda (european data archives). international indexing and local indexing both have their pros and cons. the paper is proposing a “glocal” solution that combines and integrates positive contributions from both types of indexing. the last paper “differences among faculty ranks in views on research data management” is research carried out by katherine g. akers and jennifer doty when they worked in e-science and data management at emory university libraries. katherine g. akers now works as a postdoc at the university of michigan libraries. the authors investigated how faculty researchers manage their data during and after their research projects, and their views on data sharing and preservation. the pragmatic purpose is for the libraries to develop research data services tailored to their specific needs. researchers’ age and amount of experience are often thought to be important factors, but other studies have not shown this conclusively. this research found that senior faculty stated more often that sharing their research data requires too much time and effort. i suggest that this gives data archives further incentive to continue to improve the ease of depositing research data. articles for the iassist quarterly are always very welcome. they can be papers from iassist conferences or other conferences and workshops, from local presentations or papers especially written for the iq. authors are permitted “deep links” where you link directly to your paper published in the iq. chairing a conference session with the purpose of aggregating and integrating papers for a special issue iq is also much appreciated as the information reaches many more people than the session participants, and will be readily available on the iassist website at http://www.iassistdata.org. authors are very welcome to take a look at the instructions and layout: http://iassistdata.org/iq/instructions-authors. authors can also contact me via e-mail: kbr@sam.sdu.dk. should you be interested in compiling a special issue for the iq as guest editor(s) i will also be delighted to hear from you. karsten boye rasmussen october 2013 editor iassist quarterly vol 22 no2 16 iassist quarterly data sharing is a disputed norm in scientific affairs (fienberg et al. 1985; weil and hollander 1990; fienberg 1994; mishkin 1995). on the one hand, principal investigators argue that they and their research teams are the most competent analysts of originally collected data and best able to safeguard the data against release of confidential information. they know the details and nuances of the sampling procedures, instrumentation, data reduction, and missing data. they have an investment in the original research that should be repaid by first rights of publication. they also argue that for certain kinds of complex studies, for example, observational research, organizational research, longitudinal research, clinical research, and research involving geo-coded data or administrative records linked to survey data, they are the only or principal safeguard against violations of confidentiality of the data. on the other hand, researchers argue that publicly supported data collections should be available to the public, or at least to competent researchers. data sets can be purged or cleaned of identifying information. competent researchers can do responsible secondary analyses of the data while simultaneously upholding the normative requirements for protection of confidentiality. the investment of public funds in data supercedes ownership rights at least with respect to access to the data, as also do the norms of science as an activity open to and dependent upon the scrutiny and review of other scientists. since 1962, icpsr has been responsible for many of the technical and normative developments in social science data sharing. as an archive that acquires data from many principal investigators, icpsr has had to develop and implement procedures that assure original investigators that the distribution of their data will not compromise the protection of confidentiality. as an archive that distributes data to a wide variety of users, icpsr has had to develop and implement these same procedures to substantially reduce or eliminate the opportunity for secondary users to compromise confidentiality even if they wanted to. over the past 36 years, icpsr has had to respond to new technical challenges in protecting the confidentiality of data, while simultaneously charting a course that satisfies both proponents and opponents of data sharing, both data producers and data users. in this paper, we briefly review the origins of ethical requirements and regulations for the protection of confidentiality of research data and ways that confidentiality can be violated. that discussion sets the stage for a description of the nature and development of icpsr practices to assure confidentiality of research data. these practices have had to take account of both technical developments in the capacity to store, distribute and analyze data and normative developments in the biomedical and social sciences about data sharing. finally, we describe some trends in research that pose yet new problems for protecting confidentiality of research data and some new approaches to protecting confidentiality. ethics and regulations biomedical sources. surprisingly, a review of the foundational documents that raised the consciousness about, and led to federal regulation of, the protection of human subjects in research revealed very little attention to or concern with the privacy of research data and the protection of confidentiality. the nuremberg code (oprr 1993c) addressed informed consent, social benefits of research, avoidance of suffering and injury, risks to subjects not greater than the importance of the problem, and preparations and facilities for protection of subjects against injury, disability and death. but it did not address issues of confidentiality and privacy. beecher’s seminal publications (1966a,1966b) focused primarily on safeguarding the physical health of research subjects, the absence of voluntary participation, and the need for informed consent. the belmont report’s (oprr 1993a) discussion of three basic ethical principles (respect for persons, beneficence, and justice) did not mention safeguarding privacy of research subjects or protecting the confidentiality of data obtained from them. the closest it came was in describing the principle of beneficence as making efforts to secure the well being of persons through minimizing possible harms. of the basic documents, only the helsinki declaration mentioned privacy: “every precaution should be taken to respect the privacy of the subject and to minimize the impact of the study on the … subject.” (oprr 1993b) but it did not extend this discussion of principles to its practical implication for protecting confidentiality. finally, in the current federal protecting confidentiality in archival data resources by christopher s. dunn & erik w. austin * summer 1998 17 regulations governing human subjects protection, confidentiality is mentioned only as an element of content of an informed consent statement. “…in seeking informed consent, the following information shall be provided to each subject: … (5) a statement describing the extent, if any, to which confidentiality of records identifying the subject will be maintained;” [45 cfr 46.116(a)(5)]. it seems likely that these foundational documents largely ignored privacy and consent issues because of their biomedical research origins, their primary concern with protection of the physical health and well being of the subjects, and the (false) assumption that physicians would be the primary personnel conducting biomedical research with people. under these conditions, research information from or about human subjects is equated with information obtained under the privacy and confidentiality of the physician-patient privilege in a clinical relationship. so apparently little if anything was said about privacy and confidentiality in the early biomedical discussions. early social science data collection organizations. a sharp contrast is presented in the early history of the social sciences. eckler, a former director of the census bureau, reported that for the first five censuses (17901830), copies of returns were publicly posted for corrections or additions of missing information and were deposited with local courts (1972:164). he also reported that the sixth census (1840) was the first to instruct assistant marshals (i.e., field enumerators) that they were to “consider all communications made to him in the performance of his duty, relative to the business of the people, as strictly confidential” (eckler 1972:165). eckler speculated that this phrase was introduced into the instructions either to deter or curtail the private use of an increased amount of economic information collected in the 1840 census, or to improve the reliability of reports to enumerators. protecting the confidentiality of data was an important concern for two of the early leaders of social statistics, francis a. walker, superintendent of the 1870 and 1880 censuses, and carroll d. wright, first commissioner of labor beginning in 1885 and later director of the census. up until 1902, the census was a temporary organization brought into existence each decade by legislation and terminated soon after issuing its reports. congresses during the 19th century were indisposed to creating new, permanent federal agencies. walker was appointed superintendent of the ninth census (1870), the plans for which had become embroiled in larger political issues of apportionment of house of representative seats and black suffrage (anderson, 1988:76-81). the 40th congress set aside plans for a more scientific census proposed by then rep. (later president) james a. garfield and the 1870 census proceeded under the 1850 census legislation. in the ensuing decade, walker suggested a number of scientific and operational reforms for the census and a quinquennial census in 1875 (which never came to pass) (wright 1900:58). the 1870 census was the last census that used judicial marshals appointed by the senate to supervise data collection in the states. the tenth census in 1880 used “supervisors of census” appointed by the president and who numbered more than twice as many as the judicial marshals, thereby providing more direct supervision of the actual work of enumeration (wright 1900:59) and centralized planning and control (anderson 1988:99). each enumerator had to make daily reports and submit signed copies of original data schedules. most importantly (for our present concern), the enabling legislation for the 1880 census provided elementary forms of protection of confidentiality of the data. first, the oath of office signed by enumerators required that they “will not disclose any information contained in the schedules, lists or statements obtained by me to any person or persons, except to my superior officers.” (wright 1900:937:section 7 of the act to provide for taking the tenth and subsequent censuses). second, section 12 of the enabling legislation made it a crime to violate the confidentiality of responses: “that any supervisor or enumerator, who, having taken and subscribed the oath required by this act, … shall, without the authority of the superintendent, communicate to any person not authorized to receive the same, any statistics of property or business included in his return, shall be deemed guilty of a misdemeanor, and upon conviction shall forfeit a sum not exceeding five hundred dollars.” (wright 1900:938) in the eleventh census (1890) the language about “any statistics of property or business” was changed to “any information gained by him in the performance of his duties.” (wright 1900:946) as walker’s reforms proceeded (including appointments based on merit rather than patronage), the size of the 1880 census organization grew but it ran out of appropriated funds in 1881. walker resigned in 1881, moving to the presidency of m.i.t. after criticism and buffeting by congress and the popular press, control of the census remnants and reporting finally passed to wright in 1885 and was finally completed in 1888 just before the need for legislation for the eleventh census (1890). the policy language about confidentiality that had undergone modest changes from 1840 through 1890 applied only to data collectors. other census employees, in particular, tabulation clerks, and increasingly in 1880 and 1890, professional staff, were not similarly enjoined. thus, in the law providing for the twelfth census (1900), eckler reported that “confidential treatment of the census records was, for the first time, required of all employees, and penalties for violation were applicable to everyone 18 iassist quarterly (1972:165). similarly, in the law that provided for the 1910 census, eckler reported that “the possibility of disclosure through published reports” was addressed by instructions in the industrial censuses that indicated that publication was to be made in such a way as “not to reveal the report of any establishment” (1972:165). this same provision was not extended to the population and agriculture censuses until 1930, presumably because of the lower risk of identifying people than companies. for the 1920 census, data sharing with other government officials was strictly limited “by the provision that in no case should the information thus furnished be used to the detriment of the person to whom it relates” (eckler 1972:165). things were more informal in the department of labor. plewes (1985:222) reported that carroll wright operationalized the standards for protecting confidentiality of data on a personal basis. he sent telegrams to businessmen, pledging his word “as a government officer that names of your plants and of city and state in which located shall be concealed (plewes 1985:222). plewes suggested that obtaining cooperation for data collection about sensitive topics like working hours and conditions, child labor, and wage practices was the motivating force behind these personal persuasions. these practices eventually became associated with such higher objectives as “integrity, impartiality and independence” (plewes 1985:222). plewes also noted that (as of the date of his remarks, march 1985), the bureau of labor statistics was one of only two federal statistical agencies whose policies of protecting confidentiality have existed without the protection of an agency wide confidentiality statute. the expansion of the federal government has been accompanied by the expansion of its information collection role and activities. as more and more kinds of data have been collected, issues surrounding the confidentiality of and access to government statistics have also increased. in many instances, agency practices have been formalized into statutory protections of confidentiality of statistical data and prevention of compulsory disclosure. for example, title 13 of the united states code governs the activities of the u.s. census bureau. in section 9, requirements for the confidentiality of census data are spelled out. (a) neither the secretary, nor any other officer or employee of the department of commerce or bureau or agency thereof, or local government census liaison, may, …, (1) use the information furnished under the provisions of this title for any purpose other than the statistical purposes for which it is supplied; or (2) make any publication whereby the data furnished by any particular establishment or individual under this title can be identified; or (3) permit anyone other than the sworn officers and employees of the department or bureau or agency thereof to examine the individual reports. (13 usc 9) where individual reports are allowed to be shared with government officials, those records are “immune from legal process, and shall not, without the consent of the individual or establishment concerned, be admitted as evidence or used for any purpose in any action, suit, or other judicial or administrative proceeding” (13 usc 9(a)(3)). microdata from u.s. department of justice supported research also has confidential status and is prohibited from uses in the legal process other than statistical research: …, no officer or employee of the federal government, and no recipient of assistance under the provisions of this chapter shall use or reveal any research or statistical information furnished under this chapter by any person and identifiable to any specific private person for any purpose other than the purpose for which it was obtained in accordance with this chapter. such information and copies thereof shall be immune from legal process, and shall not, without the consent of the person furnishing such information, be admitted as evidence or used for any purpose in any action, suit, or other judicial, legislative, or administrative proceedings. (42 usc 3789g) professional association ethical guidelines. a third source of confidentiality restrictions is the ethical guidelines of professional associations. the post civil war decades of the 19th century and first two of the 20th century brought immense technological development, world changing scientific discoveries in physics and chemistry, major demographic changes in american society, and the development of professions and professional organizations. the american statistical association and the american economics association were front runners in the movement to lobby for a permanent census bureau. these associations were made up of persons who had prior direct experience with the censuses or whose graduate students worked with the census or with the department of labor. thus, it is not surprising that ethical guidelines or codes of professional social science organizations eventually reflected confidentiality policies. the same people who were leaders in the associations were also leaders in the emerging disciplines and professions of the social sciences and social statistics in which confidentiality policies were first introduced. in general, professional associations are concerned with promoting the professionalism (and status) of their work. some essential aspects of professionalism are the ability to control or discipline members at the fringes of respectable practice and the provision of members with resources against outside disciplinary or malpractice actions. associations have developed codes of ethics that educate summer 1998 19 members about allowable practices or ethically suspect practices, and that guide behavior in gray areas. many address the protection of confidentiality of sources or data. the american sociological association requires sociologists to “take reasonable steps to ensure that records, data, or information are preserved in a confidential manner,” and that when confidential records, data or information are transferred to other persons or organizations, “they obtain assurances that the recipients … will employ measures to protect confidentiality at least equal to those originally pledged.” (asa 1997, section 11.08). the american political science association addresses the potential conflict between civic and legal obligations to cooperate with governmental organizations and the “professional duty not to divulge the identity of confidential sources of information or data developed in the course of research.” (apsa 1998:section 6) they are also required to observe federal and university rules and regulations for the protection of human subjects, including protection of confidentiality of data. (apsa 1998: section 34) the american statistical association, founded in 1839, has recently released a new draft publication, ethical guidelines for statistical practice, for comments. the section on ethical responsibilities to research subjects includes the following item: “protect the privacy and confidentiality of research subjects and the data they provide.” (american statistical association 1998: section ii.d.3 at http://www.amstat.org/about/ethics.html) the american association of public opinion research also has a code of professional ethics and practices policy pledging confidentiality. “unless the respondent waives confidentiality for specified uses, we shall hold as privileged and confidential all information that might identify a respondent with his or her responses.” (aapor 1998: section ii.d.2 at http://www.aapor.org/ethics/ principl.shtml) summary. the present emphasis on the biomedical roots of modern human subjects protection regulations and their original implementation in the department of health and human services obscures some important origins of the protection of confidentiality of records and data. foundational documents of ethical principles of biomedical research are largely silent on issues of privacy and confidentiality. in contrast, mid 19th century us census legislation required enumerators to keep information they collected in the course of the census confidential, principally as an instrumental means to promote subject cooperation and truthful response. the practice of maintaining confidentiality of census data was extended to all census employees in 1900. gradually, professional associations adopted policies for the protection of confidentiality. these policies are based not on instrumental values like improving the cooperation of respondents and accuracy of the data but on ethical principles like safeguarding the privacy of individuals and minimizing potential harm to subjects through disclosure of sensitive information to third parties. us census and bureau of labor statistics confidentiality practices initiated in the late 19th and early 20th centuries anticipated two of the four major possibilities for failure to maintain confidentiality. the early statements about treating information as confidential in the census enabling legislation from 1840-1890 and their extension to all census employees in 1900 recognized that individuals with legitimate access to microdata could also behave illegitimately by selling or transferring data to third parties. industrial census guidelines in 1910 and population and agriculture census guidelines in 1930 recognized that individual or microlevel identities could be deduced from macrolevel tabular data with small cell sizes, thereby reflecting the first concerns about statistical disclosure. in the next section we describe four main categories of failure to maintain confidentiality as a preface to describing activities undertaken to protect confidentiality of archival data. ways that confidentiality can be violated there are four major ways that confidentiality can be violated, resulting in the release or deduction of individual identities and/or identifying characteristics: accidental release; malicious release; compulsory release; and statistical disclosure. accidental release may be due to sloppy data management procedures, ignorance or errors on the part of staff, or failure to follow standard procedures. malicious release may be due to theft or unauthorized transfer of data by disgruntled staff or by staff or others seeking financial gain, or through breaches of computer systems security. compulsory release may occur as the result of legal action or court order. statistical disclosure results from logical use or analysis of data to identify cases or events that are infrequent or rare, or unique patterns of characteristics which when associated with data from other sources, lead to subject identification. the value of these categories is not merely descriptive. they also direct attention toward objects or mechanisms for maintaining confidentiality. the idea of accidental release http://www.aapor.org/ethics/principl.shtml http://www.aapor.org/ethics/principl.shtml http://www.aapor.org/ethics/principl.shtml 20 iassist quarterly suggests that confidentiality is preserved by: • educating staff about the need for confidentiality protection procedures; • training and monitoring staff in the application of those procedures; • performing quality control checks on data files that are developed for restricted use or public release; and • maintaining adequate security for confidential information. the central feature of malicious release is the idea that information (and hence data) has value and that there are people, whatever their motives, who may attempt to translate that value into cash or otherwise use the information inappropriately. disgruntled staff, for example, may satisfy a symbolic urge for retaliation or retribution by unauthorized transfer or release of information. regardless of whether the motive is instrumental or symbolic, the inappropriate, illegal behavior can be counteracted by deterrence and punishment. these dynamics suggest that organizations should have and use policies that prohibit the unauthorized use, transfer, or release of data. in the case of public release, even though there is no restriction on who can access the available data, there ought to be use restrictions consistent with the research and educational purposes of the organization. the matter of compulsory release is too complicated and uncertain to be dealt with in an encapsulated discussion here. it is sufficient to note that the ethics of research are not the only requirements that researchers face and that the legal protection accorded the confidentiality of research data is not absolute or uniform across states or in different legal matters. researchers have been ordered to release confidential data. some have complied, others have refused and been penalized, still others have had initial orders overturned or modified on appeal. again, the focus with this type of release seems to be a strong organizational policy against compulsory release that has as its basis the necessity of confidentiality in social research. where possible, such policies should be backed up by regulatory or statutory nondisclosure protections, such as the dhhs certificates of confidentiality or the us department of justice statutes (42 usc 3789g) and regulations (28 cfr 22) prohibiting evidentiary or other non-research uses of justice research data. the topic of statistical disclosure is also too complex to be dealt with in an encapsulated discussion. but fortunately, there is more information available on this topic than on the others. statistical disclosure has been the focus of both professional and academic attention. there are a variety of established methods for preventing disclosure (cox et al. 1985; omb 1994). there is also developmental work in progress for devising and testing new methods (e.g., duncan undated; dutta chowdhury et al. undated). some of these have been discussed at this meeting. but the central feature of this way that confidentiality is preserved is its technical focus on the data themselves. in general then, there are four approaches on which to focus attention for protecting the confidentiality of research data: • education and training of persons who work with data; • data management techniques and statistical procedures that can be applied to data; • organizational policies that mandate confidentiality and data security; and • government regulations and laws that protect the confidentiality of research data. the next section of this paper focuses attention on the first two of these approaches at icpsr. practices at icpsr to assure protection of confidentiality data modifications. these sections borrow heavily from the icpsr guide to social science data preparation and archiving. (material taken from second printing 1997:16-17 is italicized.) two kinds of variables often found in social science data sets present problems that could endanger the confidentiality of research subjects. most familiar are the direct identifiers that may have been obtained in the process of data collection. these include items such as names, addresses (including zip codes), telephone numbers (including exchanges), social security numbers, and other linkable identification numbers such as driver license numbers, certification numbers, etc s. data collectors should remove all such identifiers when preparing public use data sets. if data sets are received with such variables, icpsr will remove them as part of the lowest level of study processing. increasingly, consideration is being given to returning to investigators data sets that are received with direct identifiers. this is because icpsr practice is to preserve originally submitted data that could become the focus of legal action should it be known that icpsr maintains a copy of such a data set. another category of variables can often become problematic depending on the content of the data collection and the nature of the research subjects included in the data summer 1998 21 set. these are indirect identifiers that might be used (in combination or in conjunction with publicly-available information) to identify individual respondents. this category is harder to deal with, since it includes items that are often the focus of or useful for statistical analysis. that is probably why such information was collected in the first place. some examples of these indirect identifiers are detailed geography (e.g., state, county, or census tract of residence), organizations to which the respondent belongs, educational institution from which the respondent graduated (and year of graduation), exact occupations held, place where the respondent grew up, exact dates of events, detailed income, and offices or posts held by the respondent. such indicators should be reviewed by the principal investigator/data collector and a judgment made about the effect of retaining such items upon the confidentiality of the research subjects before depositing the data in a public archive. sometimes, variables usually considered to be indirect identifiers can become direct identifiers depending upon features of the research design. job title or occupational role can directly identify a respondent when there is only one such position in an organization, one such organization in a town (or department in an organization), and the town (or organization) is identified, as well as the date of the data collection. for example, if the police chief, presbyterian minister, high school principal, or any other unique figure in a community or organization identifies their job title or occupational role, and the community or organization is also identified, and the date of the data collection is known, then it is easy to find out exactly who that person was at that time. handling indirect identifiers. if, in the judgment of the principal investigator, a variable might act as an indirect identifier (and thus could be used to compromise the confidentiality of a research subject), the investigator should “treat” that variable when preparing a public use data set. modifications commonly used are: • removal—eliminating the variable from the data set entirely; • bracketing—combining the categories of a variable; • top-coding—grouping the upper range of a variable to eliminate outliers; • collapsing and/or combining variables—merging the concepts embodied in two or more variables by creating a new summary variable. the following example is taken from the icpsr guide to social science data preparation and archiving (1997:17). an example from a national survey of physicians (containing many details of each doctor’s practice patterns, background, and personal characteristics) may help to illustrate each of these categories of treatment of variables to protect confidentiality. variables identifying the school from which the medical degree was obtained and the year graduated should probably be removed entirely, due to the ubiquity of publicly available rosters of college and university graduates. the state of residence of the physician could be bracketed into a new “region” variable (substituting more general geographic categories such as “east,” “south,” “midwest,” and “west.”) the upper end of the range of the “physician’s income” variable could be top-coded (e.g., “$150,000 or more”) to avoid identifying the most highly paid individuals. finally, a series of variables documenting the responding physician’s certification in several medical specialties could be collapsed to a summary indicator (with new categories such as “surgery,” “pediatrics,” “internal medicine,” “two or more specialties,” etc.). icpsr staff consult with principal investigators to help them design or modify a public use data set that maintains (to the maximum degree possible) the confidentiality of respondents. the staff will additionally perform an independent confidentiality review of data sets submitted to the archive and will work with the investigators to resolve any remaining problems of confidentiality. the goal of this cooperative approach is to ensure that all reasonable steps have been taken to protect the confidentiality of research respondents whose information is contained in icpsr’s public use data sets. research trends that pose problems for confidentiality some types of studies include variables that pose unusually difficult or problematic threats to confidentiality but are also difficult to modify because of their central importance to the study. one such study is the multi level study having hierarchical files with linkage variables between files. another type is the study that has exact event dates and birth dates. a third type is the study with geo-coded information. a fourth type is the qualitative narrative interview study. a fifth type, the longitudinal panel study, is not especially problematic when ready for archiving, but the need to maintain linkage and locator identifiers from one round to the next makes the study vulnerable to threats to confidentiality during its operational phases. multi-level studies, where data is collected about places, organizations, households, persons and events, simultaneously, is especially difficult to handle with the usual means of modifying variables. often, information in the multiple levels of files will make it easy to identify individual subjects, but the linkage variables between files are essential to maintain the multi-level value of the study. where identification risks are high because the multiple levels of information make it easy to narrow the focus on individuals, icpsr will consider making the study a 22 iassist quarterly restricted use data set. studies with many precisely dated events and birth dates also pose risks to confidentiality, especially if the event information also might have been publicized in the media or recorded in publicly available administrative records (e.g., court dockets). exact dates in the study information and event characteristics can be matched against media or administrative record data allowing subjects to be easily identified. nevertheless, the exact date information is often useful for various forms of time dependent analyses like survival analysis or event history analysis. removing exact dates reduces the value of the information. once again, the solution may be creating a restricted data set rather than removing information. studies with geo-coded information are also problematic. depending upon the nature of other information in the study and the degree of area resolution, geo-coded studies may make it easy to identify subjects, especially when public information is available. for example, it would be inappropriate, unethical, and potentially dangerous to release a data set with the address locations of rape victims. again, resolving these kinds of problems caused by multiple levels of information is not a simple process of modifying indirect identifiers because of the nature of the study. qualitative narrative interviews are another type of problematic study. the level of detail provided through indepth interviews is extensive and often contains many references to people, places, events, associations, organizations, family relationships, persons not liked at work, and so forth. someone with intimate knowledge of these patterns of information may be able to easily identify the individuals involved. the very richness of the detailed information is simultaneously the value of the study and the threat to confidentiality. original investigators are loathe to restrict the richness of the narratives, yet are unwilling to release such detailed information because of the ease of identifying individuals involved in the scenes. providing access to original indirect identifiers it is rarely the case that variables removed or modified to maintain confidentiality are without value for research purposes. archives and other data providers, therefore, frequently field requests for some form of access to original data values. three of these forms of access that have been utilized will be discussed here: customized data analysis performed by the archive/data provider; private use data sets; and front-end software. the first method of providing access to restricted indirect identifiers retains the data in secure form but permits researchers to design analyses that use those data. customized data analysis (often performed at cost to the researcher) affords the opportunity of obtaining analytic results from restricted variables. typically, researchers will be asked to provide detailed analytic instructions— usually in the form of software commands—and the requested analyses are performed at the archive, with analytic output sent to the requesting party. at icpsr and elsewhere, the output is examined by staff to ensure that the analysis results will not endanger the confidentiality of respondents. delivery of a private-use data set allows original data values to be provided to a researcher, with the requestor explicitly assuming responsibility for maintaining confidentiality of those data. most organizations that provide private-use data sets require a transaction form, replete with both researcher and official signatures certifying that such data will be securely held, to be used only by the requesting party in ways that protect respondent confidentiality. a third mechanism bundles an entire data set in an analytic software package which prevents examination of discrete values/cases while allowing statistical access to all variables. this front-end software alternative usually prevents extracting or downloading of original values on some or all variables. (the national center for education statistics’ data analysis system [das] is one example of such a software-based method of protecting the confidentiality of research subjects. other such front-ends are actively being explored, including at icpsr.) each of the mechanisms described above has advantages as well as drawbacks. none are completely satisfactory to both the research community and the repository/holder of original data. tightest control of original data values is an attraction of the customized data analysis option, but is the least popular with active researchers. it is typically costly (in terms of both time and money), and frequently thwarts the iterative analytic style most common in the social sciences. private-use data sets permit the most researcher control of the analytic process, at the expense of certainty of protection of respondent confidentiality. enforcement of private-use data set provisions agreed to by requestors is difficult to effect, and sanctions against violators of promised assurances would inevitably involve a litigious voyage on mostly-uncharted waters. possibly the most secure yet flexible alternative is the front-end software option. yet from the archive’s standpoint, this is probably the most expensive of the three alternatives; putting data into one of these packages is so time-consuming that it can practicably be utilized on very few data collections. furthermore, it is doubtful that front-end software is wholly impervious to hacking by a skilled and determined violator. finally, the learning of “yet another” software package and its guaranteed limitations raises the bar over which interested researchers must jump to access needed research data. other alternatives for protecting confidentiality yet other mechanisms have been proposed or are being experimented with in the quest for the “ideal” way of summer 1998 23 protecting respondent confidentiality. brief mention will be made of three “positive” alternatives, before we close this section on a draconian note. licensing a researcher to use a data set containing indirect identifiers is a variant on the private-use data set arrangement described above. like it, a licensed use is agreed to after completion of a transaction form. unlike private-use data set agreements, however, most licenses impose an up-front fee in the form of a security bond as surety for maintaining confidentiality. the fee has been known to range from a few hundred to many thousands of dollars. several license mechanisms also require the researcher and her/his institution to assume all legal liability in any instance of breaching confidentiality. needless to say, the popularity of this form of “access-with-assurance” is quite low in the research community (not to mention in the college/university legal offices). a second alternative method is being discussed in more detail elsewhere at this conference, and so will be briefly alluded to here. this is the “perturbing” of original data values to break the certain bond between any given data value and the (possibly identifiable) individual who may have provided the initial information. since the essence of this technique is the altering of original data values, it remains suspect in the minds of several generations of social scientists. these individuals find it difficult to overcome one legacy of their training—getting error out of research data collections—which clashes with the practice of introducing error into a data set (however noble the purpose underlying that introduction). perhaps more promising is the concept of secure data analysis laboratories. in such facilities, original data would be available for data analysis in a controlled setting, precluding such things as making copies of original data, investigating single cases, or transmitting the data offsite. scholars would apply to visit the site to do data analysis in the laboratory under secure conditions. an experiment using this form of access can be found at carnegie mellon university, for its violence research consortium project supported by the national science foundation and the national institute of justice. data from the national crime victimization survey, which have long been distributed without geographic sector information, are available with geographic information at carnegie mellon to the violence consortium members this mechanism represents, for social scientists, a departure from a long-term trend of facilitating the export of research data from an archive or producer site directly to the institution (or desktop!!) of the interested scholar. it should be noted parenthetically that many research materials utilized by both historians and social scientists are available only by visiting the site where the research materials are housed. included among such facilities are traditional archives and other repositories, including some fine social science collections like those of the henry murray center at radcliffe college. undoubtedly more costly for the individual researcher (and perhaps for the archive as well), this mode of access to confidential data may become more common with heightened concern for preserving confidentiality. the search for suitable mechanisms for protecting confidential microdata promises to become a high-stakes venture. at risk is the big kahuna of post-wwii social scientific research practice—readily available, empirical microdata. some in the statistical and social science communities, as well as in government, are beginning to worry about the release of any microdata, with a few even predicting its demise. conclusion the very progress of social science research methodology has made it more difficult to safeguard the confidentiality of the research data. removing direct identifiers is a foundational requirement for public use data sets but that is essentially a trivial task. more difficult tasks involve investigating which variables could be used as indirect identifiers and modifying them without significantly reducing the value of the data collection. careful attention must be paid to interactions among the context of the study, the nature of the sample, and the characteristics of respondents to prevent ordinarily unrevealing information from becoming the pointer to an individual. but many studies today involve complex research designs with multiple levels of data collection, file linkage variables that are crucial to the statistical analysis, sources of information that are intrinsically locational in nature, or detailed descriptions of events or situations that can be crossreferenced in publicly available sources like the media or administrative records. maintaining complete archival files for these kinds of studies may involve other procedures than simply eliminating or modifying variables. procedures used in the past or under development include: • conducting contracted analyses; • creating private use data sets; • developing front end software to limit access to data records; • licensing data use; • introducing noise (known statistical error) into data records; • developing data laboratories in which the data can not be removed from the site. 24 iassist quarterly references american association of public opinion research. code of professional ethics and practices. aapor 1998 at http:// www.aapor.org/ethics/principl.shtml american political science association. a guide to professional ethics in political science. (second edition) washington, dc: american political science association, 1998. american sociological association. code of ethics. washington, dc: american sociological association, 1997. at http://www.asanet.org/ american statistical association. ethical guidelines for statistical practice. american statistical association 1998 at http://www.amstat.org/about/ethics.html anderson, margo j. the american census: a social history. new haven and london: yale university press, 1988 beecher, henry k., m.d. “some guiding principles for clinical investigation,” jama 195:135, march 1966. beecher, henry k., m.d. “ethics and clinical research”, new england journal of medicine 274:1354, june 1966. cox, lawrence, bruce johnson, sarah-kathryn mcdonald, dawn nelson, and violeta vazquez. “confidentiality at the census bureau,” proceedings of the first annual research conference, march 20-23, 1985, pp 199-218. washington, dc: u. s. department of commerce. bureau of the census. 1985. duncan, george, ramayya krishnan, rema padman, phyllis reuther, and stephen roehrig. “cell suppression to limit content based disclosure.” undated reprint from h. j. heinz iii school of public policy and management, carnegie mellon university, pittsburgh, pa 15213 dutta chowdhury, sumit, george t. duncan, ramayya krishnan, stephen f. roehrig, and sumitra mukherjee. “disclosure detection in multivariate categorical databases: an optimization approach.” undated reprint from h. j. heinz iii school of public policy and management, carnegie mellon university, pittsburgh, pa 15213 eckler, a. ross. the bureau of the census. new york: praeger publishers, 1972. fienberg, stephen e. “sharing statistical data in the biomedical and health sciences: ethical, institutional, legal, and professional dimensions,” annual review of public health 15:1-18, 1994. fienberg, stephen e., margaret e. martin, and miron l. straf, editors. sharing research data. washington, dc: national academy press, 1985. inter-university consortium for political and social research. guide to social science data preparation and archiving. (second printing) ann arbor, mi: icpsr, (september), 1997. mishkin, barbara. “urgently needed: policies on access to data by erstwhile collaborators,” science 270:927-928, (november 10), 1995. office of management and budget report on statistical disclosure limitation methodology. washington, dc: u. s. office of management and budget. statistical policy office. federal committee on statistical methodology, (may), 1994. office of protection from research risks. the belmont report: ethical principles and guidelines for the protection of human subjects of research in protecting human research subjects: institutional review board guidebook. appendix 6 (a6:7-14) washington, dc: u.s. department of health and human services. public health service. national institutes of health, 1993a. office of protection from research risks. declaration of helsinki in protecting human research subjects: institutional review board guidebook. appendix 6 (a6:36). washington, dc: u.s. department of health and human services. public health service. national institutes of health, 1993b. office of protection from research risks. the nuremberg code in protecting human research subjects: institutional review board guidebook. appendix 6 (a6:1-2). washington, dc: u.s. department of health and human services. public health service. national institutes of health, 1993c. plewes, thomas. “confidentiality: principles and practice,” proceedings of the first annual research conference, march 20-23, 1985, pp 219-226. washington, dc: u. s. department of commerce. bureau of the census. 1985. weil, vivian, and rachelle hollander. “sharing scientific data ii: normative issues,” irb: a review of human subjects research 12(2):7-8, (march/april), 1990 wright, carroll d. the history and growth of the united states census. washington, dc: u. s. government printing office, 1900. * presented at the annual meeting of the international association for social science information service & technology (iassist), new haven, ct, may 20, 1998. christopher s. dunn director, national archive of criminal justice data icpsr and erik w. austin director, archival development icpsr. http://www.aapor.org/ethics/principl.shtml http://www.aapor.org/ethics/principl.shtml http://www.amstat.org/about/ethics.html vol21.2 34 iassist quarterly 1. introduction the u.s. census bureau (boc) is developing a prototype statistical metadata repository for use with internet data dissemination and automated integrated survey processing tools. the repository will be an electronic catalog of information about survey designs, processing, analyses, and data sets. access will be through the internet and the world wide web. substantial background work was done before work to build the prototype could begin. statistical metadata is the information and documentation needed to describe and use statistical data sets for the lifetime of the data. the efficient, effective, electronic management of metadata greatly increases the usefulness of those data sets, especially for internet data dissemination. statistical metadata can also be used to facilitate survey design, processing, management, and analysis. automated integrated survey processing systems, which create and use this information, will allow statistical agencies to conduct their programs in ways that were not possible before. the repository is being designed based on standards and data models. it is being implemented as a relational database and organized through these standards and models. international, american, and internal census bureau standards are all being brought to bear in the development of the repository. three models have been developed and integrated to form the structure of the repository. the models are the business data model, the data element registry model, and a metamodel. tools for the collection of the metadata and querying the repository are under development. without the cooperation of the survey designers and analysts who create the metadata, the repository will never be populated. general, intuitive, and easy to use tools must be developed to collect the data. conversely, the information in the repository will not be useful if it cannot be retrieved in an easy way. a survey business process model, or table of contents, has been developed for users and analysts to find the type of information they may want to provide. this table of contents is being used as a template in the design of the tools. also, it can be used to design a low level interface for other systems to access and use the repository. substantial benefits should be available to the census bureau when the repository is functional. it organizes the documents, data sets, and variable descriptions of the agency. the repository will allow for comparisons across surveys (data or designs) which previously have not been easily available. finally, the repository will make the public information of the agency fully available from a common source. if other statistical agencies around the world adopt similar approaches, the concept of a “single world-wide statistical agency” on the internet could become reality. this paper will define what statistical metadata is, describe the design of the repository (including the standards and models), describe the tools under development for populating and querying the repository (including the table of contents outline), and discuss the ramifications for the agency of implementing the repository. 2. definitions statistical metadata is descriptive information or documentation about statistical data, i.e. microdata and macrodata. statistical metadata facilitates sharing, querying, and understanding of statistical data over the lifetime of the data. the two types of statistical data (electronic or otherwise) are described as follows (see lenz, 1994): • microdata data on the characteristics of units of a population, such as individuals, households, or establishments, collected by a census, survey, or experiment. • macrodata data derived from microdata by statistics on groups or aggregates, such as counts, means, or frequencies. the extensive nature of statistical metadata lends itself to categorization (see sumpter, 1994) into three components or levels: • systems the information about the physical characteristics of the application’s data set(s), such as the statistical metadata repository: an electronic catalog of survey descriptions at the u.s. census bureau by daniel w. gillman and martin v. appel* summer 1997 35 location, record layout, database schemas, media, size, etc; • applications the information about the application’s products and procedures, such as sample designs, questionnaires, software, variable definitions, edit specifications, etc; • administrative the management information, such as budgets, costs, schedules, etc. the systems, applications, and administrative components help to differentiate the sources and uses of statistical metadata. some authors (see, for example, sundgren, 1991b, 1992, 1993) refer to the applications and administrative components of metadata as metainformation. we chose to use the term metadata because it seems to simplify the discussion. statistical metadata and metadata repositories have two basic purposes (see sundgren, 1991a, 1991b, 1992, 1993): • end-user oriented purpose: to support potential users of statistical information, e.g. through internet data dissemination systems; and • production oriented purpose: to support the planning, design, operation, processing, and evaluation of statistical surveys, e.g. through automated integrated processing systems. a potential end-user of statistical information needs to • identify, • locate, • retrieve, • process, v interpret, and • analyze statistical data that may be relevant for a task that the user has at hand. the production-oriented user’s tasks belong to the following types of activities: • planning/design/maintenance, • implementation/processing/operation, and v evaluation. an input-oriented statistical agency is one where the statistical surveys they conduct or manage are also the natural building blocks of its organization. the boc is currently an example of such a statistical office. an output-oriented statistical agency is one which focuses on meeting the needs of its customers. the boc is striving to become more output-oriented. see sundgren (1991a, 1991b, 1992, 1993) for a more detailed discussion of these ideas. output-oriented database systems relate data from different surveys. they need special software and metadata tools for reconciling data from different sources and for helping the users to interpret and analyze the data. this paper describes the pieces necessary to build those metadata tools. statistical metadata repository (mdr) is a planned repository of statistical metadata and pointers to other metadata (such as documents or images). a proof-ofconcept system has been built (see gillman and appel, 1994), and a series of prototypes are under development. the design, uses, and functionality of the mdr will be discussed in more detail below. 3. statistical metadata repository the mdr is being designed to assist with two new types of tools which are under development at the boc: internet data dissemination ; and automated integrated survey processing systems . these tools correspond to the enduser oriented purpose and production oriented purpose, respectively, of statistical systems. statistical systems are known formally as statistical information systems (sis) (see sundgren, 1991b, 1992, 1993; or gillman, appel, and laplant, 1996). 3.1 purposes the eventual plan for the mdr is that it will contain the metadata for survey designs, processing, analyses, datasets, and related information for all surveys the boc performs. links to the data files, documentation, and images (such as questionnaire forms) will also be stored (see sundgren, et al, 1996; or appel, et al, 1996). this has led to the management of data in a decentralized and non-uniform way. on one hand, there is a need for the survey management to process and manage their data in the most efficient way. on the other, there is a need for data users to be able to find and access data efficiently and effectively. the mdr will facilitate a solution for the data users while allowing the survey data managers to find a smooth transition to standard data management strategies. there are many functions for which the mdr is being designed. primarily, the mdr will be a standard tool for researchers and analysts to locate survey data and metadata. data dictionaries, record layouts, questionnaires, sample designs, and standard errors are examples of information 36 iassist quarterly that will be directly available. links from subject types, e.g., income, race, age, and geography, to data sets will allow users to locate data sets by subject. less obviously, users can compare designs of different surveys and find common information collected by them. the mdr will help facilitate data administration at the boc. many surveys define data elements with the same name but with (slightly) different definitions. an aim of the mdr is to help people manage this problem. if definitions and other attributes of data elements are standardized across surveys, through the use of a data element registry (a subset of mdr), then confusion generated by the differences in meaning will be reduced. naming standards and conventions are also needed to reduce the confusion. the mdr will provide the information necessary for the user to understand the distinctions and similarities among data elements from multiple data sources. the design of the data element registry part of the mdr will be based on a standard, and it will be discussed in more detail in section 3.2.2. many of the purposes for the mdr are associated with both the end-user orientation and production orientation. here we will list the end-user oriented purposes. the typical end-user oriented sis is an internet data dissemination system. some of the major functionality for the mdr in support of this is: • location of data sets by survey name and date or content (e.g. household income); • names, definitions, and related information about data elements and links to the surveys and data sets that use them; • links to documentation describing aspects of survey design, processing, or analysis; • links across documents to identify common themes contained in them; • links to images (e.g. questionnaire forms) that are of interest; • the ability to search the information potential through query languages such as sql. the typical production oriented sis is an automated integrated survey processing system. most of the purposes of the mdr for the end-user oriented systems are common to the production oriented systems as well. often, production oriented sis users will be survey analysts working within the boc (statistical agency). they have and need access to confidential data to which external endusers cannot have access. the additional functionality must support this use, such as: • links to all the data sets produced by the instance of a survey (e.g. current population survey, june 1996); • links to frame, sample, and administrative records files; • links to a management information system; • links to some confidential metadata such as disclosure analysis algorithms. these lists are not meant to be inclusive, but to give a fairly extensive picture of the potential uses for the mdr. 3.2 models the design of the mdr is based on three data models. within the repository, these models have been integrated into one extensive model which covers many aspects of statistical metadata. extensions to the model are planned as new items or needs are identified. the three models represent the major dimensions to the mdr model (see figure 1). they are described briefly here and will be discussed in more detail below: • business data model the model describes the business of the boc surveys. it describes survey designs, processing, analyses, datasets, products, and documents as related to statistical surveys. • data element registry model a data element registry is a mechanism for managing the names, definitions, permissible values, and other attributes of data elements. metadata describing data elements is entered into the registry by a process called registration. expanding the concept of registration to include surveys, products, datasets, and documents, this model handles the needs of registering metadata. • metamodel this model describes application specific areas and other non-business related items such as security, access control, database schemas, record layouts, and time frames. the metamodel provides the repository’s view to itself. the mdr prototype also uses a business process model described below. • table of contents a business process model has also been developed. it is in the form of an outline, or table of contents (toc). the toc describes the processes of a survey from design to data dissemination. the mdr model can also be divided into five functional areas. this view gives a clearer picture of how the integrated model works (see figure 2). summer 1997 37 the functional areas are: • data element registry manages the names, definitions, permissible values, and other attributes of data elements (see above and below). • registration manages the metadata needed to register items for which the repository keeps track: surveys, data elements, documents, datasets, products. this section handles the information types which are common to each of the objects which are registered in the mdr, much like an electronic card catalog system. • metamodel manages the application specific information such as security and access control, search criteria, record layouts, database schemas and access, etc (see above and below). • business data manages information about surveys, including design, processing, and data (see above and below). • documentation manages information about documents. the association of documents to different records within other parts of the model acts as a classification system for the documents. 3.2.1 business data model the business data model (bdm) describes the business (statistical surveys) of the boc. it is composed of entities, attributes, and relationships which describe information that a statistical agency needs to keep about surveys. much of this information is in the form of specifications or procedural documentation. the model supports the storage of metadata as single attributes or as documents. figure 3 is a high level er diagram of the bdm, and see appendix figure 1: overview of integrated model 38 iassist quarterly a for an entity definition list. the bdm describes survey designs, processing, analyses, and datasets. it contains entities for each of the important parts of a survey: universe, frame, sample, questionnaire, etc. the model allows for the organized storage and search for metadata about a survey, and it allows searching for metadata items across surveys. many statistical metadata systems in use today address the metadata needs for a single survey or application, but the bdm addresses the metadata needs for many surveys. an important feature of the bdm is that documentation is handled in a general way. each entity of the model allows for many documents to be attached to a single record. the documents can be distinguished by version, document type (e.g. specification, procedure, memo, etc.), the entity the document is associated with, and the relationships the given record has with other records in the model. this provides a comprehensive classification scheme for documents which helps users search directly for the information they need. coupled with the indexed and key word search provided by most internet search engines, the bdm is a powerful document management paradigm. the model also provides several other features listed below: • maintains a list of all current surveys conducted by the agency; • allows for comparing designs, specifications, or procedures across surveys; • allows for reuse of designs, specifications, or procedures; • provides for assembling complete documentation for a survey. figure 2: repository model overview: functional areas summer 1997 39 3.2.2 data element registry model data elements (or variables) are the fundamental units of data an organization collects, processes, and disseminates. a data element registry (der) is a mechanism for managing data elements in a logical fashion. der’s organize information about data elements, provide access to the information, facilitate standardization, help identify duplicates, and facilitate data sharing. der’s are like data dictionaries in that they contain definitions of data elements. but more than data dictionaries, they contain all the information about individual data elements that an organization requires. data dictionaries are usually associated with single data sets (files or databases), but a der contains information about the data elements for an entire program or organization. the information contained in a der is part of an organization’s metadata. therefore, the registry itself will be part of the mdr. important applications for der’s include sis’s. electronic data dissemination requires easy access to information about data elements. data element names, definitions, and classification schemes will help users in locating and understanding data sets. automated integrated survey processing systems that will include sample and questionnaire design, automated edits and imputation, and coding systems require full descriptions of data elements. designers need to know the definitions of all variables that may be affected by the programs they are using. the der model provides for all the metadata needed to describe data elements. it also provides the entities necessary for registration and standardization of data elements. generalizing the concept of registration (see section 4.2 below) to include documents, datasets, products, and surveys provides a framework for merging the der and the bdm. a consequence of registering the important metadata items in the mdr is that the repository, from the registration point of view, becomes a card catalog of metadata items. the integration must also include linking data elements to each of the entities in the bdm which use them (e.g. frame, sample, survey dataset, question, etc.). an important feature of the der is that data elements are composed of a concept (data element concept) and a representation or value domain (set of permissible values). the power of this is seen as follows: • sets of similar data elements are linked to a shared concept, reducing search time; • every representation associated with a concept (i.e. figure 3: business data model 40 iassist quarterly each data element) can be shown together, increasing flexibility; • all data elements that are represented by a single (reusable) value domain (e.g. sic codes) can be located, assisting administration of a registry; • similar data elements are located through similar concepts, again assisting searches and administration of a registry. see figure 4 for a high level er diagram of the der, and see appendix b for an entity definition list. 3.2.3 metamodel the metamodel is the repository’s view of itself. it contains application specific entities necessary for the functioning of particular sis’s, and information which controls access to metadata in the rest of the repository. the kinds of information the metamodel handles are access control, security, physical location of data, machine addresses, record layouts, database schemas, access procedures, etc. the development of the metamodel has been iterative. no specific metamodel has been built. instead, as new functions are identified, they have been added to the mdr model. the partnerships (see section 3.4) that have been formed with sis developers for using the mdr model have been a rich source for metamodel entities and attributes. as these partnerships continue and the sis’s are further developed, more information is added to the metamodel and to the mdr model. .3.2.4 business process model a table of contents (toc) outline view (see census bureau, 1996) of survey processes has been developed. it was patterned after work done by a boc reinvention lab and at statistics sweden (see rosen and sundgren, 1991). the toc is formally a business process model. it is figure 4: data element registry model summer 1997 41 divided into eight chapters, each detailing a different aspect of survey processing. the chapter names and their descriptions follow below: • content the content refers to the nature of the information that is the subject of the survey, i.e. what the universe is, a description of the data collected, and a description of the resulting products. may contain definitions, and data standardization and coding information. • planning documentation related to the planning and management of the design; the conduct of the survey and the analysis, dissemination and disposition of the data. this includes documentation related to budgeting, manpower, and training. • design the design and specifications for how the survey will be conducted. includes the design of the frame, sample, and questionnaire; and the specifications for edits, coverage, and estimations. • data collection obtaining information from respondents and the conversion of that data into a form which can be processed. • data processing the stage of a project, following collection and receipt of the original material and preceding report-writing, during which the information is entered onto a machine-readable medium (or directly into a computer system) and eventually used to produce tabulations and statistical analyses. • data analysis documentation related to all statistical processes used to analyze the survey results or those used for displaying or presenting the resultant information. • data dissemination the process of making data available to users, electronically or otherwise. electronic data dissemination includes use of the internet or cd-roms. • data any information gathered as the result of a survey or added to a survey form. there are two uses that are being developed for the toc: 1) to be used as a “check list” for users who need to provide metadata or users who want to search metadata from the mdr (see section 4.2); and 2) to serve as a mapping between the mdr and other repositories which need to share metadata (see gillman, appel, and laplant, 1996). in particular, the toc can be used as a means to classify documents from another repository in the mdr. 3.3 standards in this section the applicable standards which have been used to guide the development of the mdr and its associated tools will be described briefly. 3.3.1 data element standards the model for the data element registry portion of mdr is based on the conceptual framework contained in the ansi draft standard, the metamodel for the management of shareable data (mmsd), ansi x3.285. it, in turn, incorporates all the principles described in an emerging international standard, specification and standardization of data elements, iso/iec 11179 (see ansi x3l8, 1996). ansi x3.285 provides a conceptual model for building a data element registry and contains some extensions to the framework described in iso/iec 11179. a complete data dictionary describing all the entities, attributes, and relationships of the conceptual metamodel is included in this document. the mmsd metamodel provides a detailed description of the types of information which should belong to a data element registry. it provides a framework for how data elements are formed and the relationships among the parts. implementing this scheme will provide users the information they need to understand an organization’s data elements. iso/iec 11179 is being developed in six parts. the names of the parts, a short description of each, and the status follow below: • part 1 framework for the specification and standardization of data elements provides an overview of the concepts in the rest of the standard. the current status of this document is committee draft. • part 2 classification of data elements describes how to classify data elements. the current status of this document is working draft. • part 3 basic attributes of data elements defines the basic set of metadata for describing a data elements. this document is an international standard. • part 4 rules and guidelines for the formulation of data definitions specifies rules and guidelines for building definitions of data elements. this document is an international standard. • part 5 naming and identification principles for data elements specifies rules and guidelines for naming and designing non-intelligent identifiers for data elements. this document is an international standard. • part 6 registration of data elements describes the functions and rules that govern a data element registration authority. this document is an international standard. 42 iassist quarterly 3.3.2 survey design and statistical methodology metadata content standard the survey design and statistical methodology metadata content standard (sdsm) (see laplant, et al, 1996; or census bureau, 1997) is a draft statistical metadata content standard for the boc. it will provide a description of the information or documentation about statistical data. the content and design of the standard is based primarily on the bdm. the entities of the bdm specify the content sections of the sdsm. sdsm will provide developers and users of statistical products with a common vocabulary for describing the design processing, analysis, and data sets for censuses and surveys. the sdsm also will serve as a glossary of statistical metadata concepts. broad agreement on the meaning and organization of these concepts will provide the basis for improved communication among the producers and users of economic and demographic statistical data sets. each of the 29 sections in the sdsm consists of a list of entries, some that reference other sections. each entry is a metadata data element. any of these metadata data elements may be used to identify specific instances of metadata. the metadata may be some specific information (such as a number or text) or a url to a file of some type (e.g. documents, gif’s, etc.) the sdsm has been submitted to the formal standards review process of the boc, and is expected to be issued as a boc standard in summer 1997. once this occurs, it is hoped that other statistical agencies will adopt the sdsm or similar standards. 3.3.3 other standards information resource dictionary system (irds) is a standard which addresses the use, control, organization, and documentation of the information resources of an enterprise (see nist, 1989). it is an application of another standard, reference model for data management (rmdm) (see iso, 1995). the organization of the mdr model is based on the organization specified in irds. see graves and gillman (1996) for a more detailed discussion. the federal geographic data committee (fgdc) of the u.s. government has developed a family of metadata standards which addresses the geographic content of data. executive order 12906 has mandated that all u.s. agencies that produce geographic based data use these fgdc standards. most boc data is based on geography, therefore these standards will apply to boc data. government information locator service (gils) (fips192) is an extensible standard which describes a format and specifies the underlying protocol (niso z39.50) for making metadata available on the internet. another executive order (the paperwork reduction act of 1995) mandates that all u.s. agencies create and maintain gils records. this provides a mechanaism for the public to find information about what their government is doing and producing through electronic means. 3.4 partnerships several groups within the boc developing sis’s have agreed to use the mdr structure to support the underlying metadata needs of those systems. a short description of each sis follows below. dads (data access and dissemination system) is the name for the census bureau initiative to develop and implement data access and dissemination focused on the 2000 decennial census and continuous measurement data sets, but with the ability to accommodate other data sets having geographic detail, such as those produced from the economic and agricultural censuses. the main objective of dads is to provide one general (electronic) system for all access to census bureau data. the system will be designed to be fast, flexible, and costefficient. to achieve this, four cross-directorate teams were formed to study and recommend policies or designs for user input, promotion and outreach, pricing for products, copyright or trademark, corporate look and feel, data archiving, metadata and documentation, and coordination of various efforts and activities. dads will attempt to incorporate other work, such as ferret, where it is appropriate. the dads team is following a schedule to produce a new prototype each year, in the month of september, until the full production system is built in 2002. the 1996 prototype was a success, though limited in scope. the mdr model will be used to organize the metadata for the 1997 prototype. ferret (federal electronic research and review extraction tool) (see capps, 1995) is a data extraction tool available on the internet that allows users to find information about monthly demographic survey data using a world wide web browser. users can select microdata items (individual survey question items) which can be used to create custom data queries. in addition, users can select macrodata (aggregated or summarized) tables to get preformatted survey data. results of data queries can be output in sas datasets or ascii files. these results can be viewed on the screen or can be downloaded to a local computer. the sas output allows one to get the results in pie charts, bar charts, or summarized on a u.s. map. the ascii output can be brought into an excel spreadsheet. the ferret system can be divided into four major parts. the first part is the user interface which is via the world wide web. the ferret repository contains metadata such summer 1997 43 as basic variable definitions, keywords, concepts, and other items. the document management system handles the documents which describe the survey design, processing, and analysis. finally, there are two databases handling all the microdata and macrodata. ferret currently handles current population survey data. plans are to add other demographic survey data in the future. work is also underway to make the ferret repository model and the mdr model compatible. this will enable people to work with dads and ferret systems seamlessly. steps (standard economic processing system) (see steps, 1996) is an integrated survey processing system the objective of which is to eliminate redundant processing by combining existing survey systems into one system. the scope of the steps system includes providing the following basic survey processing functions: data review and correction; edits; imputation; outliers; estimation; estimation variance; disclosure analysis; time series; queries (canned/ad hoc); tables (canned/ad hoc); management information; and survey control operations (for scheduling of batch mode processes). it will also provide the following additional functions: generate standard and non-standard mail files for mail-out operations; generate standard telephone files for telephone follow-up operations; maintain standard variable names and flags; maintain standard data structures; allow entry of survey design specifications including edit and imputation parameters as determined by analysts or through automated historical data analysis provide audit trails and backup capabilities; provide access to ssel; and provide access to other economic area surveys and censuses. the above provides a view of the functionality which steps will be designed to provide. implementation details are not yet available. the steps system developers plan to use the mdr as a source of information about variables. product registration is a multi-divisional effort to unify the systems that manage the production, inventory, distribution, and sale of census bureau products. the mdr model will be used to register products, i.e. link products to the variables, surveys, geography, and other items that will enable users to locate them. this work has recently started. 3. metadata management the main aspects of managing metadata are content, storage, collection, registration, retrieval, system integration, and metadata administration. this section will describe how the standards based approach and the proposed design architecture address each of these aspects. 4.1 content and storage content refers to the identification of which metadata will be collected and stored in the mdr, and storage refers to the how, i.e. the physical and logical mechanisms for storing the metadata. much of the paper to this point has been addressing these issues. the prototype mdr is being built using oracle rdbms as its underlying storage mechanism and is based on the models and standards discussed above. the models and standards describe the metadata content and how that content is organized for storage. 4.2 collection metadata collection is recognized as a very difficult problem because of the fundamental changes that the survey design and analyst teams must go through to perform their work. at the boc and other statistical agencies, metadata (mostly documents, often in the form of memos) is created either electronically or on paper for each survey, but it is just beginning to be stored in an organized repository, database, or document management system. asking people to use a new system to capture this metadata and organize it represents a big change. the tools that are created must mimic as closely as possible the working paradigm already in place, such as the use of certain word processors and templates for creating documents. a major problem is that the working paradigm for each survey design and analysis team is different. so, creating common tools will require substantial planning. also, incentives must be found so that the designer/analysts will want to provide the metadata to the mdr. no matter how well designed, tools without an obvious payoff to the user will not be used. management can help with the adoption of metadata collection tools by supporting their use, but the end-users will ultimately decide their fate. 4.3 registration registration is the process of providing the mdr with its knowledge about the metadata, e.g. name, location, type, etc. the general classes of items which need to be registered are data elements, surveys, products, datasets, and documents. registration requires several things: 44 iassist quarterly • all the necessary attributes are specified; • all the necessary links are made (e.g. linking a dataset to all the data elements in its data dictionary); • classifying the registered item. registration tools will have to be designed, probably one for each class of item. the tools will require a template for the user to supply the necessary attributes and make the links to other metadata as needed. appropriate classification structures will need to be accessible through the tool so each item can be classified. useful classification schemes already exist which can be incorporated into registration tools, such as • toc; • themes as specified in the cultural and demographic data metadata draft standard of the federal geographic data committee; • thesauri from statistics canada and the university of essex (u.k.). it will be useful for the boc to build a taxonomy of statistical terms to help with the classification problem. of course, effective classification schemes also help with the search for metadata and for understanding the semantics of data or metadata several prototype metadata collection tools are in place at the boc and other statistical agencies. scbdok (at statistics sweden), document management system (dms in use with ferret at boc), and the commercial document management system pcdoc (for 1997 economic censuses) are all designed or being designed under the framework outlined above. 4.4 retrieval retrieval refers to querying metadata in the mdr. querying will be part of the design of general purpose browsers and of sis’s which work with the mdr. user interfaces for metadata-driven systems will let users query the metadata to locate data or other survey information. query languages such as sql will allow the user to retrieve any metadata which is in the mdr. other search mechanisms such as wais, key word, and hyper-text are available through the internet. this is especially important for documentation databases. the toc view of the sdsm can be used as a check list for categories of metadata. for users wishing to find information about surveys, searching the toc for the appropriate subject (e.g. questionnaire design) will be useful. since the toc process model is designed to be a complete description of survey design, processing, analysis, and data sets, then the toc view will provide users access to all the metadata the boc has about a survey. a prototype metadata browser for the mdr has been built, and browsers are being built for the dads and ferret data dissemination systems. 4.4 system integration in addition to the tools for collecting and querying metadata, the integration of the mdr with other sis’s needs to be seamless. two general possibilities for accomplishing this exist. first, the toc can be used. a mapping exists between the toc and the mdr model, and maps can be built from the toc to the other sis’s by mapping the toc to their metadata models. then, a map will exist from the mdr to each sis, through the toc. the mdr will act as a hub, a central communication link between the different sis’s in use at the boc (see gillman, appel, and laplant, 1996). another solution, probably more effective, is for developers of sis’s to adopt the mdr model for the metadata portion of the sis. if every sis at the boc uses the mdr model, then a distributed metadata repository (each piece based on the same model) will be built. tools designed to search the metadata in one sis will be able to search the metadata in all sis’s. a seamless view of the metadata for the entire agency will result. users who look for boc data in ferret will be able to locate data that is only accessible through dads without having to know which tool to go to first. the actual viewing or downloading of the data will probably require switching tools, but that problem should be minor. 4.5 metadata administration the adoption of the mdr model for storing metadata will require more than supplying information about data elements, surveys, documents, or datasets. metadata administration is the active management of the information about all the agency’s metadata. no function of this type exists at the boc at this time at the agency level. the registration process described in section 4.3, and the der described in section 3.3.2, define generally the information that is required for accurate and complete data administration. the mdr model has expanded the notion of data registration to include metadata. the registration tools discussed above will handle the entering of metadata into the mdr, but there is a human side to metadata administration which must not be lost in the discussion of the mdr. some of these functions are: • determining which data elements have the same meanings as others; summer 1997 45 • determining whether metadata items have been properly classified; • ensuring all necessary information is properly supplied for each registered metadata item; • working with metadata administrators of other agencies to facilitate the sharing of data and metadata; • designing rules for forming metadata definitions. • designing and implementing naming conventions; metadata administration will require a large commitment from the boc, but it will greatly enhance the usefulness of boc data, make the mdr a better tool, and facilitate the sharing and understanding of data and metadata among groups within the boc or with other agencies. 4. prototype a series of prototypes is currently under development. the first version is complete. it implemented a subset of the mdr model and contained some information about some data elements and documents. a browser tool was developed using the toc as a search mechanism for specific types of documents. the browser is a web based tool that uses a combination of basic html, cgi-perl scripts, and java. the second prototype is under development now. it is expected to be complete in july. it will implement the complete mdr model, contain substantially more documents, and use an improved version of the browser. two important functions will be demonstrated: the ability to find metadata across surveys and a tool to register metadata for products. subsequent prototypes will add more functionality each time. usability testing is planned for some of the prototypes. both the registration tools and the browser will require user feedback to ensure that the tools are useful for users. unfortunately at this time, the prototypes cannot be released on the web to the internet. much of the metadata in the mdr is not available for the public, and the security functions for the mdr have not been developed to the point where this information is secure. 5. conclusion this paper has discussed the work at the boc to design and build a prototype statistical metadata repository (mdr) using standards developed by international, national, and u. s. government organizations. detailed data and metadata models have been built and integrated. the integrated model is the basis for the mdr architecture. it provides a structure for storing the metadata which describes survey designs, processing, analyses, and datasets. the model supports the card catalog metaphor for organizing the boc metadata. the mdr will not be an end in itself. instead, it will work in conjunction with internet data dissemination and automated integrated survey processing tools. several examples of both of these tools are under development at the boc. the mdr prototypes must be ready in time to meet the schedules of these other tools. the first mdr prototype has been built and subsequent ones are planned. registration and query tools are being developed, and the prototype mdr is being populated with metadata. increasing interest in using the mdr model for storing metadata for various projects has increased the chance that a seamless distributed metadata repository for the boc can be developed. further research, planning, and work will be necessary to bring this plan to reality. 6. references appel, m. v., gillman, d. w., laplant, w. p. jr., creecy, r. h. (1996), “towards unified metadata systems and practices”, isis-96, bratislava, slovakia, may 21-24, 1996. ansi x3l8 data representations (1996), “iso/iec 11179 part 1 framework for the specification and standardization of data elements, working draft 7”, february 1996. capps, c. (1995), “overview of the technical architecture for ferret”, census bureau internal document, demographic surveys division. census bureau (1997), “statistical design and survey methodology metadata content standard”, draft, census bureau internal document, april, 1997. census bureau (1996), “table of contents for statistical design and survey methodology metadata content standard”, draft, census bureau internal document, july 2, 1996. gillman, d. w. and appel, m. v. (1994), “metadata database development at the census bureau”, presented at the un/ece metis working group meeting, geneva switzerland, november 22-25, 1994. gillman, d. w., appel, m. v., and laplant, w. p. jr. (1996), “design principles for a unified statistical data/ metadata system”, proceedings of ssdbm-8, stockholm, sweden, june 18-20, 1996. graves, r. b. and gillman, d. w. (1996), “standards for management of statistical metadata: a framework for collaboration”, isis-96, bratislava, slovakia, may 21-24, 1996. iso (1995), “reference model for data management”, 46 iassist quarterly iso/iec 10032:1995(e). laplant, w. p. jr., lestina, g. j. jr., gillman, d. w., and appel, m. v. (1996), “proposal for a statistical metadata standard”, census annual research conference, arlington, va., march 18-21, 1996. lenz, h.-j. (1994), “the conceptual schema and external schemata of metadatabases”, proceedings of ssdbm-7, pp160-165, charlottesville, va, september 28-30, 1994. nist (1989), national institute for standards and technology, “information resource dictionary system (irds)”, federal information processing standard (fips) publication 156, april 5, 1989. rosen, b. and sundgren, b. (1991), “documentation for reuse of microdata from the surveys carried out by statistics sweden”, research and development statistics sweden, june 28, 1991. steps (1996), “standard economic processing system document 1: concepts and overview”, internal census bureau document, april 16, 1996. sumpter, r. m. (1994), “white paper on data management”, lawrence livermore national laboratory document, 1994. sundgren, b. (1991a), “towards a unified data and metadata system at the australian bureau of statistics final report, december 2, 1991. sundgren, b. (1991b), “statistical metainformation and metainformation systems”, r&d report statistics sweden, 1991:11. sundgren, b. (1992), “organizing the metainformation systems of a statistical office”, r&d report statistics sweden, 1992:10. sundgren, b. (1993), “guidelines on the design and implementation of statistical metainformation systems”, r&d report statistics sweden, 1993:4. sundgren, b., gillman, d. w., appel, m. v., and laplant, w. p. (1996), “towards a unified data and metadata system at the census bureau”, census annual research conference, arlington, va., march 18-21, 1996. * paper presented at iassist/ifdo ‘97, odense, denmark, may 6-9,1997. summer 1997 47 appendix a: entity definitions for business data model entity name entity definition entity note data element a single unit of data that in a data is a representation of certain context is considered facts, concepts, or indivisible. it cannot be instructions in a form that decomposed into more allows them to be collected, fundamental segments of data organized, processed and stored that have useful meanings in a retrievable form for within the scope of the communication, interpretation, enterprise. or processing by human or automated means. emprise an identifiable effort to this appeared in prior models as (project) generate deliverables not project specific to a single survey instance emprise_dataset a dataset containing either case this appeared in prior models as (project_dataset) level data, aggregation of case project_dataset level data, or statistical manipulations of either. frame a dataset containing all the cases identified for a survey instance based on a survey’s universe definition methodology a structured approach to solve a problem product a finished deliverable of a project or survey instance for external use. program a group of surveys related by a common theme. a program can be made of other programs purchaser an external organization or individual who buys census bureau products question a request for one or more related pieces of information from a case. a question can contain other questions questionnaire an identifiable instrument containing questions for a particular survey instance 48 iassist quarterly sample a dataset containing a subset of a frame for a particular set of survey instances, selected with a specific sampling technique. for a census, the sample incorporates the entire frame. supplied_dataset a dataset acquired from sources outside the bureau of the census. can be case level or aggregated/transformed data. supplier an external organization which provides data to augment the census bureau’s efforts survey an investigation about the characteristics of a given universe survey-dataset a dataset containing either case level data, aggregation of case level data, or statistical manipulations of either the case level or aggregated survey data, for a single survey instance survey-instance an identifiable activity which uses a system(s) to gather and process a set of data items from an identifiable set of cases, for a defined period of time, resulting in one or more specific deliverables system an identifiable process, either fully automated or computer assisted, which implements one or more techniques to produce one or more deliverables. a system can be composed of systems technique an identifiable algorithm which is used to implement all or part of a methodology universe the total defined set of interest to one or more surveys summer 1997 49 appendix b: entity definitions for data element registry model entity name entity definition ________________ ______________________ administered data a generalization for a data registration component element, value domain, data concept, object class or property. classified data a subtype of administered registration component data registration component, all the data components that require classification. conceptual domain the set of possible valid values of a data element expressed without representation. drc name context an association between an administered data registration component and a name context. drc registration a registration authority that authority has registered a particular data registration component. data element a single unit of data that in a certain context is considered indivisible. it cannot be decomposed into more fundamental segments of data that have useful meanings within the scope of the enterprise. data element concept the human perception of a property of an object set, described independently of any particular representation. data registration classification schemes which classification scheme are used to classify registered data. datatype a category used to classify the collection of letters, digits, and/or symbols to depict values of a data element based upon the operations that may be performed on the data element. derivation type an entity used to define different types of derivations. used to normalize the 50 iassist quarterly derivation type attribute associated for derived data elements and data element concepts. derived dec to derived an association that tracks a de mapping derivation mapping at the conceptual level to a derivation mapping at the data element level if such a mapping were to exist. this is not required for all derivation mappings. enumerated vd a list of all permissible values. formula an entity that represents an algorithm to compute values. formulas involve input quantities (data elements) and produce output quantities (data elements). keyword an entity that expresses potential search keywords that users of the registry will use to search for and access data element concepts. name context the system, database, standard document, or other environment in which the logical metadata class functions and the name has meaning. non enumerated vd a range used for specifying the lower limit and the upper limit of permissible values. object class a set of concepts, abstractions, or things in the natural world that can be identified with explicit boundaries and meaning and whose properties and behavior all follow the same rules. organization an accredited agency authorized to declare logical metadata classes as registered. (from earlier definition of registration authority). permissible values allowed values in a value domain property a classification of any feature that humans naturally use to summer 1997 51 distinguish one individual object from another. it is any one of the characteristics of an object class that humans use as a label, quantity or description registration authority the organization authorized to register entries in the registry. representation class a classification of value domains based upon the type of representational form. synonym lists a relationship that captures the fact that two distinct data elements have different names but the same meaning (synonym). value domain a set of permissible values, used to represent a data element. value meaning meaning associated with permissible values in an enumerated domain. iassist quarterly 11 the swedish election studies: studies of 30 years of swedish electoral behavior bv iris alfredson 1 before using data for secondary analysis, the researcher must be aware of the principles used in compiling the material. the aim of this paper is to give a short description of the swedish election studies, and discuss some of the problems which may arise if one is lo study changes over time using these data.. quite a few variables are common lo all ten studies, but are they comparable? there may have been slight changes in question wording, or the coding of the variable may have changed over lime. these are some of the changes thai affect direct comparisons. additional complications arise when making international comparisons, but such problems are outside the scope of this paper. '•presented at the international association for social science information service and technology (iassist) conference held in vancouver, british columbia, canada on mav 19-22, 1987 / hope that this presentation will act as an introduction to the extensive material available on swedish elections. for those whose requirements are more comprehensive, a bibliography of books, reports and papers on the subject is provided. background in conjunction with the local elections of 1954, jorgen westerstahl and bo sarlvik of the department of poliucal science, university of gothenburg, conducted the first scientific field survey of swedish voting behavior. the survey was a local election study conducted in gothenourg and the countryside around boras. the survey was inspired by that conducted by lazarsfeld, berelson and gaudet in eire county. the questions concerned mainly social background, political party sympathies and motives, previous political sympathies, stated and actual vote behavior, reasons for change in sympathies, public reaction to content of media coverage, political sympathies of friends, and knowledge of and poliucal opinions on specific elecuon campaign issues. the 1954 survey was a pilot test for a larger national study. two years later, in conjunction with the parliamentary election of 1956, the first nation wide survey on party choices, participation and poliucal opinions of the swedish electorate was conducted. this was the real start of the swedish elecuon research program. since then, similar surveys have been carried out at all elections. further, studies have been conducted in conjunction with the two referenda that have taken place since then, the referendum on the general supplementary pension scheme (atp) in 1957, and a referendum on nuclear power in 1980. the swedish election studies are, together with those carried out in the united states, norway, summer 1987 12 iassist quarterly france and west germany, the only academic election studies that extend as far back as the 1950s. the swedish parliamentary election studies have been carried out at every election since 1956, making them one of the most comprehensive sources on voting behavior. financing the local elecuon study of 1954 and the national election study of 1956 were funded by the swedish social science and legal research council, the foundation for research in sociology at the university of gothenburg, the swedish broadcasting cooperation and the political parties. since 1960, the election studies have been financed through government grants and carried out as a part or the election statistics program by the ceniral bureau of statistics (staususka central byran). survey design since the mid-1950s, there has been a political behavior research program at the department of poliucal science, university of gothenburg. the elecuon studies have, with the exception of the 1976 study, come into being through a close collaboration between the department of poliucal science at the university of gothenburg and the swedish central bureau of stausucs (scb). the 1976 study was conducted b\ the department of political science in uppsala in collaboration with scb. the survey research center of the central bureau of statistics is responsible for the sampling for the election studies, their permanent interview organization performs the field work, and they also collect additional data from public registers. the research project at the department of poliucal science is responsible for the general planning of the studies, the construction of the questionnaires, and the analysis and presentation of data. the basic principles of the studies, the questionnaires, and the coding scheme were originally designed by bo sarlvik. the surveys were initiated by professor jorgen westerstahl and directed by bo sarlvik (1956, 1960, 1964, 1968, 1970, 1973), olof petersson (1973, 1976), and soren holmberg (1979, 1982, 1985). the sample represents the resident, enfranchised population. the samples for the earlier election studies were drawn from the survey research center's sampling framework which consisted of a nationwide set of primary sampling units which provide the framework for a 'general purpose' two-stage population sample. since 1973, another method has been applied. the samples are now drawn directly from the scb register over the total population (rtb), bymeans of the scb standard program for random samples. respondents not included in the target population because of ineligibility to vote are excluded from the sample. the target population even includes swedes resident abroad, but these are not included in the sample. people above 80 years of age at the time of the study are excluded in order to avoid the difficulties encountered in interviewing very old people : the 1968 survey was an exception, the age limit was 84). the proportion excluded by this age limit is about 3% of the total sample. in 1956, respondents were interviewed twice, once before and once after election day. from the 1960 study onwards, with the exception of the 1970 study, field work has been carried out in two stages. the total sample is split into two subsamples of equal size. one subsample is contacted for personal interviews during the field work stage preceding the election. respondents in this subsample are contacted again after elecuon day through a short mail summer 19s7 iassisl quarterly 13 questionnaire. the primary purpose of this mail questionnaire is to obtain information about the final vote decision of these respondents. the second subsample is contacted for personal interview during the weeks immediately after the election. most of the interviews are held between mid-august and mid-october. the 1970 election study differs from the others because the entire survey was carried out after election day. different techniques were used in 1970: about 1/3 of the sample were interviewed in their homes, and additional 1/3 through telephone interviews, and the remainder received a short mail questionnaire. in the other election surveys, respondents were interviewed in their homes. some busy respondents were interviewed by telephone. respondents interviewed before the election received, immediately after election day, a short mail questionnaire which mainly contained questions on final voting decision. panel, in which half of the 1973 sample was reinterviewed in 1976. the "new" respondents in 1976 were reinterviewed in 1979, and so on. in this way, all respondents are interviewed twice. at present, there are four panels: 1973-1976, 1976-1979, 1979-1982 and 1982-1985. in every survey, a supplementarysample of first-time voters, who have become entitled to vote since the last election, is also drawn. in conjunction with the 1957 referendum, respondents were interviewed three times; this design was also applied to the 1980 referendum. the questionnaire which was sent out to respondents in the 1980 referendum sample was also sent out to respondents belonging to the 1979 election study sample. in this way, two long term panels were created, one for the years 1976-1979-1980 and one for the years 1979-1980-1982; this makes it possible to do analyses over a longer time span. panels in order to facilitate the study of variation and constancy in voting behavior, a panel design has, with one exception, been used. the panel technique used has varied over the years. in 1956, respondents were interviewed twice, once before and once after election day. in 1960, no panel was used. in 1964, a three-stage panel was started, in which the same respondents were interviewed at the elections of 1964, 1968, and 1970. a major problem with a panel extending over such a long period of time is that the sample loss has a tendency to increase at every stage. in order to facilitate the study of individual changes, but at the same time avoid too large a sample loss, a new type of panel was introduced in 1973. this was a kind of "rolling" two-stage sample loss during the first decades in which swedish field surveys were conducted, sample loss rates were low, about 5% 7%. in the 1960s and the beginning of the 1970s, there was a dramatic rise, the sample loss increasing to about 20% before stabilizing at that level. sample loss in the swedish election studies have also followed this pattern: the 1965 sample loss was 5%. during the 1960s it slowly increased (8%, 8% and 12%) until it reached 14% in the 1970 survey. then there was a sudden increase, from 18% in the 1973 survey to 26% in the 1976 survey. in the latest election studies, sample loss has been reduced to less than 20%, mainly through the use of shorter interviews. the main portion of the sample loss was due to refusals to participate, and through shorter interviews with respondents who are unwilling summer 1987 14 iassist quarterly or are pressed for time, sample loss can be reduced. the shortened interviews take about half the time of a normal interview. some respondents answered an extremely short interview, containing only those questions on final vote decision. the framing of the questions the swedish election studies series now extends over a period of thirty years. at present, the series consists of ten surveys conducted in conjunction with parliamentary elections, and two conducted in conjunction with referenda. during this period, a great number of questions have been asked. some questions are repeated in all surveys, which makes it possible to study changes over a period of thirty years, and there are also questions specific to one or two surveys. some surveys have a large number of questions on media, while others (1970 and 1973) don't touch on the subject at all. in the 1973 and 1976 surveys, there were many questions on international politics and events, an area not covered in the other surveys at all. the 1968 election study was almost twice as large as the others, largely because of the goal of the survey which was to cover the system of representation. comparable questions were also asked in an interview survey of members of the lower house of parliament 1 . questions asked on current topics include: questions on public representation on bankboards (1968, 1970), possible swedish membership in the european economic community (1968, 1970, 1973), quesuons on nuclear power (1976, 1979), and wage-earners" investment funds (1979, 1982). : representationsundersokningen, by sarivik et al.) bo however, the central questions in an election study are always those on party preference. therefore, all surveys contain questions on: respondent voting habits, party preferences, and voting in both current and previous elections. voter participation is also checked using electoral registers. information on father's political sympathies is also available in most of the surveys. other recurring questions cover political interest and party identification. social background factors such as information about the respondents' date of birth, sex, marital status, education, occupation and trade union affiliation are also available. information on occupation, education, trade union affiliation and political party membership in the 1970 survey (the third stage of the 1964-1968-1970 panel) was extracted from the 1968 survey, for all respondents except those lost from the 1968 sample and those added in the 1970 supplementary sample. for these respondents, information from 1970 was used. the swedish social science data service is at present compiling a continuity guide to all questions and variables used in the swedish election studies. problems of comparability a series of surveys extending over thirty years could be a gold mine for researchers wishing to study changes over time. such comparisons are not, however, always problem-free; a number of factors make comparison difficult, such as changes in question wording between surveys. in addition, questions may not address the same groups, the variable coding may have changed, the source of background variables may change from registers to interviews, or vice versa. changes in society, such as an increasing proportion of women in the work force, and summer 1987 lassist quarterly 1? new education systems, also make direct comparisons difficult in order to save time and trouble, it is necessary' to study the construction of the variables carefully before starting analysis. the following are problems that have been identified in the swedish election studies. marital status: information on the respondent's marital status has alternately been gathered through questionnaires and from registers. in 1956, the question asked was: "are you married, unmarried, widow/widower, divorced?". in 1960-1976. this information was extracted from a population register. in the most recent surveys, the information has again been collected in the interview. the question in 1979/1982 was worded: "which of the following alternatives best describes your marital status (married or cohabiting, unmarried or divorced, widow/widower)? the advantage of this method is that the data collected are more up-to-date. the number of code categories used has also varied over time; from 1956 to 1968, four categories were used: married, widow/widower, divorced or unmarried. since the end of the 1960s, it became more common that people live together without being married; this is reflected in the election studies. in the 1970 survey, the category 'unmarried but cohabiting' was introduced. since the 1976 survey, the number of categories has been reduced to three: married/cohabiting couple, unmarried or divorced, widow/widower. education: changes have been made to the variables on education over the years, because of changes in society. new school systems have been introduced, and the general level of education has risen dramatically. the coding scheme has also changed: the election studies conducted from 1968 to 1976 used a very detailed coding for education. question wording has remained fairly constant over the years: have you ar.y education above 'folkskol' level? (if yes:) what education do you have?" (1956-1964) "do you have any practical or theoretical education above 'folkskol' level? what education? have you any other practical or theoretic education?" (1968-1973) "what education do you have? have you any other practical or theoretic education?" (1976-1982). the early election studies (1956-1964) have the same education code categories, except that category 2 (folkhogskola, yrkesskola) in 1956 was divided into two categories in the following two surveys. categories in the earlier surveys remain in the more recent surveys (1979-1982), but changes in the school system are easy to track. categories 3 and 4 in the 1960 and 1964 election surveys have been combined into a new category 3. this category includes 'grundskola', a level which did not exist earlier. there is a new category, '4', which includes education in 2-year 'gymnasie' courses; this is also a new level of education, as compared with earlier studies. during the period 1968-1976, a very detailed coding scheme was used for education. the question on education was split with a variable for general basic education, followed by addiuonal variables for other education. education is coded with a three digit cod, of which the first digit denotes the main group, the second the level of education, and the third the type of education (degree). occupation: major problems of comparability occur in the definitions of work and class. there are problems in the classification of married women and students, and a new classification system for occupation has been introduced. the variable 'occupation group' is present in all surveys. in the 1956 survey, it had 12 categories; since 1960, it has had over 30 categories. married, non-working women have been included in different categories over the years: in the 1956 survey, all married women were classified according to husband's occupation. in 1960 and 1964, working married women were coded according to their own occupation, while non-working married women were coded according to husband's occupation. in 1968, 1970 and 1973, a special code was used for these women, and since 1976. they have been classified according to previous occupauon. summer j 987 id iassist quarterly students have also been treated differently over the years. since 1970 they have had a separate code category, but before that, they were coded according to the social status or occupation of the family head. since 1968, a detailed, 3—digit code has been used for occupation. the first two digits denote area of occupation, the third digit status of occupation. place of residence: all surveys, with the exception of the 1970 and 1973 studies, include information on respondent's place of residence. the categories have changed from survey to survey. with the exception of the 1982 survey, it is possible, by combining categories, to extract three comparable categories: large cities (stockholm, goteborg and malmo), other towns, and rural areas and villages. party preference: the question "which party did you vote for?" is asked in all surveys. respondents interviewed before the election, with exception of the 1964 survey, received a questionnaire immediately after the election. respondents in the 1964 pre-election sample were asked "which party do you like best?" information on who actually voted in the election is also available. in most surveys (1956 to 1964 and 1976 to 1982), one variable is used to summarize the contents of the two variables on voting and election participation. respondents who stated that they voted for a certain party, but who, according to the election register, didn't vote, are excluded. with the exception of the 1956 study, all surveys have contained a question on how the respondent voted in earlier elections: "did you vote in the election 19..? (if yes:) which party did you vote for?" similar data can, from the 1956 survey, be extracted by combining the answers to the following questions: "did you vote for the same party at earlier elections?" and for those who answered 'no': "which party did you vote for? (1956-1964) there is only one question: "some people are strongly convinced adherents of their party. others are not so strongly convinced. do you yourself belong to the strongly convinced adherents of your party? in the later surveys, party identification is measured in three questions: "many feel strongly for a particular party, whilst others do not feel the same allegiance towards any of the parties. how do you see yourself, as a liberal or social democrat or moderate or centrist or communist? or don't you have this attitude towoards any of the parties?" the respondents who consider themselves party adherents are asked the following question: "which party do you like best?" and finally, those who named a party were asked: "some people are strongly convinced adherents ...". newspapers: with the exception of the 1970 and 1973 surveys, questions on which newspapers respondents read have been asked. but one must be carefull with these data. in the earlier studies (1956 to 1964), respondents were asked which newspapers they read daily, while in the remaining surveys (1968, 1976-1982), respondents were asked which papers they read regularly. in the 1976 survey, 'regularly' was defined as at least 4 times/week for daily newspapers and at least every second week for weekly papers, while in the 1979-1982 surveys, 'regularly' was defined as at least once/week for all papers. in 1968, only one-half the sample were asked which papers they read regularly. this question was asked in the mail quesuonnaire which was sent to pre-election respondents, and 'regularly' was not defined. the coding of the variable in the earlier surveys makes direct comparison with later surveys impossible. the code categories included a variety of combinations of information on subscriptions, political affiliation of the newspapers, type of newspaper, etc. since 1968, the names of the newspapers have been coded. party identification: all studies include questions on party identification. in the earlier surveys a project to recode the oldest surveys, is currently ongoing at the department of political summer 1987 iassisl quarterly 17 science, gothenburg university. this will, hopefully, solve many of the problems of non-comparability among the surveys. swedish election studies at the swedish: social science data service ((ssd)) all election studies through 1982 and the 1980 referendum survey have been deposited in the swedish social science data service (ssd). with the excepdon of the referendum study, they have all been documented with the aid of the gido-system, which produces a data file and machine-readable codebook in osiris format these osiris-format files can be converted to other formats. the 1980 referendum is currently being documented, and we hope that, in the near future, the 1985 elecdon survey will also be available. it is also possible that the referendum study of 1957 will be documented. the following is a short summary of the swedish election studies currendy available from the swedish social science data service (ssd): swedish election study, 1956 (ssd 0020) principal investigators: jorgen westerstahl and bo sarlvik, department of polidcal science, university of gothenburg. total sample: 1,146 number of respondents: 1,088 sample loss: 58 (4.9%) number of variables: 220 weight: persons born 1876-1885 were sampled using 1/2 probability. these are represented by two cards in the data-file. the total number of respondents is therefore 1.131 (43 duplicates). method: interview in home. panel: respondents were interviewed twice, once before and once after elecdon day. format: osiris, spss-x, machine-readable codebook. swedish election study, 1960 (ssd 0001) principal investigator: bo sarlvik, department of polidcal science, university of gothenburg. total sample: 1,603 number of respondents: 1,466 sample loss: 137 (8.5%) number of variables: 215 method: interview in home. one half of the sample was interviewed before elecdon day, the other half after. pre-elecdon respondents also answered a short mail quesdonnaire after the election, which mainly contained quesdons on final vote decision. panel: none. format: osiris, spss-x, machine-readable codebook. swedish election study. 1964 (ssd 0007) principal investigator: bo sarlvik, department of polidcal science, university of gothenburg. total sample: 3,109 number of respondents: 2,849 sample loss: 260 (8.4%) number of variables: 219 method: interview in home. one-half of the sample was interviewed before elecdon day, the other half after. pre-elecdon respondents also answered a short mail questionnaire after the elecdon, which mainly contained quesdons on final vote decision. panel: the first stage of a three-stage panel study in which the sample was reinterviewed in conjuncdon with the parliamentary elecdons of 1968 and 1970. format: osiris, spss-x, machine-readable codebook. swedish election study, 0039) 1968 (ssd summer 1987 18 iassist quarterly principal investigator: bo sarlvik, department of political science, university of gothenburg. total sample: 3,356 number of respondents: 2,943 sample loss: 413 (12.3%) number of variables: 532 method: interview in home. one-half of the sample was interviewed before election day, the other half after. pre-election respondents also answered a short mail questionnaire after the election, which mainly contained questions on final vote decision. panel: the second stage of the three-stage 1964-1968-1970 panel study. format: osiris, spss-x, machine-readable codebook. swedish election study, 1970 (ssd 0047) principal investigator: bo sarlvik, department of political science, university of gothenburg. total sample: 4,815 interview in home 1,602 telephone interview 1,580 mail questionnaire 1,633 number of respondents: 4,130 interview in home 1,355 telephone interview 1,407 mail questionnaire 1,368 sample loss: 685 (14.2%) interview in home 247 (15.4%) telephone interview 173 (10.9%) mail questionnaire 265 (16.2%) number of variables: 223 method: three types of interviews: in-home interviews, a somewhat shorter telephone interview, and a mail questionnaire with only a small number of questions. all interviews were conducted after the elecuon. panel: the study represents stage three in the 1964-1968-1970 panel. format: osiris, spss-x, machine-readable codebook. swedish election study, 1973 (ssd (hmo) principal investigators: bo sarlvik and olof petersson, department of political science, university of gothenburg. total sample: 3,179 number of respondents: 2,596 sample loss: 583 (18.3%) number of variables: 239 method: interview in home. one-half of the sample was interviewed before elecuon day, the other half after. pre-election respondents also answered a short mail questionnaire after the elecuon, which mainly contained questions on final vote decision. panel: the study represents stage one in the 1973-1976 panel. format: osiris, spss-x, machine-readable codebook. swedish election study, 1976 (ssd 0008) principal investigator: olof petersson, department of political science, university of uppsala. total sample: 3,580 number of respondents: 2,652 sample loss: 928 (25.9%) number of variables: 290 weight: respondents belonging to the 1973-1976 panel who did not respond in 1973 had sample probability halved in 1976; when processing the data, these were weighted by a factor of 2. method: interview in home. one-half of the sample was interviewed before election day, the other half after. pre-election respondents also answered a short mail questionnaire after the election, which mainly contained questions on final vote decision. panel: the study represents stage two in the 1973-1976 panel and stage one in the 1976-1979 panel. format: osiris, spss-x, machine-readable codebook. swedish election study, 1979 (ssd 0089) summer 1987 lassist quarterly 19 principal investigator: soren holmberg, department of political science, university of gothenburg. total sample: 3,498 number of respondents: 2,816 sample loss: 682 (19.5%) number of variables: 303 weight: respondents belonging to the 1976-1979 panel who did not respond in 1976 had sample probability halved in 1979. these are represented by two records in the data file. the number of respondents is therefore 2,905 and the sample loss is 853. method: interview in home. one-half of the sample was interviewed before election day, the other half after. pre-election respondents also answered a short mail questionnaire after the election, which mainly contained questions on final vote decision. panel: the study represents stage two in the 1976-1979 panel and stage one in the 1979-1982 panel. format: osiris,' spss-x, machine-readable codebook. swedish election study, 1982 (ssd 0157) principal investigator: soren holmberg, department of political science, university of gothenburg. total sample: 3,597 number of respondents: 2,943 sample loss: 654 (18.2%) number of variables: 303 weight: respondents belonging to the 1979-1982 panel who did not respond in 1979 had sample probability halved in 1982. these are represented by two records in the data file. the number of respondents is therefore 2,980 and the sample loss is 744. method: interview in home. one-half of the sample was interviewed before election day, the other half after. pre-election respondents also answered a short mail questionnaire after the election, which mainly contained questions on final vote decision. panel: the study represents stage two in the 1979-1982 panel and stage one in the 1982-1985 panel. format: osiris, spss-x, machine-readable codebook. publications a number of books, reports, and papers have been published based on the results of the election studies. the following list does not claim to be comprehensive.n bibliography asp, kent journalisternas inflytande i valkampen. nordicom-information ni. 3-4, 1983. asp, kent, per hedberg, and peter esaiasson. massmedieundersokningen. folkomrostningen 1980. teknisk rapport [ts] goteborg: statsvetenskapliga instituuonen, goteborgs universitet 1984. asp, kent rikspoliiik och kommunalpolitik. valrorelsernas utrvmme i svensk dagspress 1956-1985. (statens offentliga utredningar, 1987:6) stockholm, 1987. asp, kent. the struggle for the agenda party agenda, media agenda and voter agenda in the 1979 swedish election campaign. communication research 10(3):333-355, 1983. asp, kent sveriges radio och 1982 ars valrorelse. [ts] goteborg: statsvetenskapliga institutionen, goteborgs universitet 1983. asp, kent v'aljarna och massmediernas partiskhet in: valiare parti er massmedia . stockholm: liberforlag. 1982. clausen, aage, soren holmberg and l. dehavens. contextual factors in the accuracy of leader perception of constituents' views. summer 1987 20 iassist quarterly journal of politics 45(2):449-472, 1983. esaiasson, peter. partiledarna infor valiarna . (forskningsrapport 1985:4) goteborg: statsvetenskapliga institutionen, goteborgs universitet, 1985. esaiasson, peter and mikael gilljam. v'aljarna och de vilda strejkerna. statsvetenskaplig udskrift 86(3): 197-207, 1983. georgsson, ake. den unge v'aljaren och poliuken. lie. diss., statsvetenskapliga institutionen, goteborgs universitet, 1973. gilljam, mikael and lennart nilsson. svenska folket och den offenuiga sektorn . (rapportserien 1984:2) goteborg: statsvetenskapliga institutionen, goteborgs universitet, 1984. gilljam, mikael and lennart nilsson. svenska folkets asikter om den offentliga sektorn tva forklaringsansatser. statsvetenskaplig udskrift 88(2): 123-139, 1985. gilljam, mikael. v'aljarna och v'arldspoliuken. statsvetenskaplig udskrift 87(2), 1984. granberg, donald and soren holmberg. modeling the relauonships among preference, expectauons and voung behavior . (rapportserien 1983:3) goteborg: statsvetenskapliga insutuuonen, goteborgs universitet, 1983. granberg. donald and soren holmberg. prior behavior, recalled behavior, and the predicuon of subsequent voung in sweden and the united states. human relations 39(2): 135-148, 1986. holmberg, soren. facklig akuvitet bland kvinnor och man. in: jamstalldhetsperspektiv i forskningen . rapport fran ett symposium arrangerat av riksbankens jubileumsfond. april 1980 (rj 1980:4) holmberg, soren, hans nordlof and mikael gilljam. folkomrostningsundersokningen 1980. teknisk rapport . (valundersokningar, rapport 7) stockholm: liberforlag, 1984. holmberg, sorcn, and kent asp. kampen om karnkraften. en bok om valiare, massmedier och folkomrostningen 1980 . (valundersokningar, rapport 6) stockholm: liberforlag, 1984. holmberg, soren. karnkraften och 1980-talets poliuska konflikter. in: moderna uder . rapport fran en symposium om teknik, poliuk och samhallsdebatt arrangerat av riksbankens jubileumsfond, mars 1979. (rj 1979:4) holmberg, soren and olof petersson. marxismen och folkbegreppel [ts] goteborg: statsvetenskapliga insutuuonen, goteborgs universitet, 1972. holmberg, soren and olof petersson. nagra anteckningar om klassbegreppet. [ts] goteborg: statsvetenskapliga insutuuonen, goteborgs universitet, 1971. holmberg, soren. poliuska sakfragor och den representauva demokraun. [ts] goteborg: statsvetenskapliga insutuuonen, goteborgs universitet, 1971. holmberg, soren. 'riksdagen representerar svenska folket'. empiriska studier i representauv demokrau. diss., statsvetenskapliga insutuuonen, goteborgs universitet (lund: studenuitteratur) 1974. holmberg, soren. riksdagsm'an och valiare . stockholm: forum for samhallsdebatt, 1974. holmberg, soren. svenska valiare . (valundersokningar, rapport 4) stockholm: liberforlag, 1981. holmberg, soren. valet 1979. in: allmanna valen 1979 . part 3. (sveriges officiella stausuk) stockholm: staususka centralbyran, 1981. holmberg, soren and hans nordlof. valundersokning 1979. teknisk rapport. (valundersokningar, rapport 5) stockholm: liberforlag, 1981. holmberg, soren and mikael gilljam. valundersokning 1982. teknisk rapport . (valundersokningar, rapport 9) stockholm: liberforlag, 1985. holmberg, soren, mikael gilljam and maria oskarson. valundersokning 1985. teknisk rapport. (valundersokningar, rapport 11) stockholm: liberforlag, 1987. holmberg, soren. valiare i forandring . (valundersokningar, rapport 8) stockholm: liberforlag, 1984. summer 1987 [assist quarterly 21 holmberg, soren and mikael gilljam. valiare och val i sverige . (valundersokningar, rapport 10) stockholm: bonniers [sic], 1987. holmberg, sbren. valjarna och lontagarfonderna. in: valiare partier massmedia . stockholm: liberforlag, 1982. oscarsson, vilgol politiskt deltagande. en jamforande studie av politisk participation i sverige, usa, storbritannien och vasttyskland. lie. diss., statsvetenskapliga institutionen, goteborgs universitet, 1973. oskarson, maria. yrkesoch socialgrappsklassificeringar i valundersokningama 1956-1985. [ts] goteborg: statsvetenskapliga institutionen, goteborgs universitet, 1986. petersson, olof. klassidentifikation. in: valiare partier massmedia . stockholm: liberforlag, 1982. petersson, olof. new trends in the swedish electorate: a focus on the 1976 election, paper presented at the ecpr workshop "social structure and political change". grenoble, april 1978 petersson, olof. social class and electoral change: sweden. 1956-1973. [ts] uppsala: statsvetenskapliga institutionen, uppsala universitet. 1975. petersson, olof. stormaktspolitik i svensk opinion. internationella studier 1975(4). petersson, olof. stormaktspolitik i svensk opinion 1973-1976. internationella studier (1):9-12, 1978. petersson, olof. valundersokningen 1976. teknisk rapport (valundersokningar, rapport 3) stockholm: liberforlag, 1978. petersson, olof and bo sarlvik. valet 1973. in: allmanna valen 1973 . (sveriges officiella statistik) stockholm: statistiska centralbyran, 1975. petersson, olof. valet 1976 in: allmanna valen 1976 . part 3. (sveriges officiella statistik) stockholm: statistiska centralbyran, 1978. petersson, olof. valiarna och valet 1976 . (valundersokningar, rapport 2) stockholm: liberforlag, 1977. petersson, olof. valiarna och varldspolitiken . (statsvetenskapliga foreningen i uppsala, nr. 90) stockholm: norstedt & soners forlag, 1982. petersson, olof. valjarnas okade osakerhel [ts] uppsala: statsvetenskapliga institutionen, uppsala universitet, 1978. petersson, olof. v'anster-hoger dimensionen och variationer i politisk participation. [ts] goteborg: statsvetenskapliga institutionen, goteborgs universitet, 1973. petersson, olof. the 1973 general election in sweden. scandinavian political studies 9:219-228, 1974. sainsbury. diane. class voting and left voting in scandinavia. [ts] stockholm: statsvetenskapliga institutinen. stockholms universitet, 1986. smith, l, aage clausen, and soren holmberg. perceptual accuracy as a function of legislative attitudes and the distribution of constituency views. paper presented at the 1981 annual meeting of the american political science association, new york, september 1981. sarlvik, bengl stabila valjare och partibytare. statsvetenskapli e tidskrift (3):121-151, 1975. sarlvik, bo. det kommunala sambandet och valjaropinionen. en studie rorande valjarnas information och attityder i forfattningsfragor. [ts] goteborg: statsvetenskapliga institutionen, goteborgs universitet, 1965. sarlvik. bo. electoral behavior in the swedish multiparty system, diss., statsvetenskapliga institutionen, goteborgs universitet, 1970. sarlvik, bo. intervjuundersokningen. in: allmanna val. riksdagsmannavalen 1959-1960 . part 2. (sveriges officiella statistik) stockholm, stockholm: statistiska centralbyran, 1961. sarlvik. bo. intervjuundersokningen. in: allmanna val, riksdagsmannavalen 1961-1964 . part 2. (sveriges officiella statistik) stockholm: statistiska centralbvran, 1965. sarlvik, bo. lnteryjuundersokning rorande valet till andra kammaxen 1968. in: allmanna val. riksdagsmannavalen 1965-1968. part 2. summer 1987 22 iassist quarterly (sveriges officiella statislik) stockholm: statistiska centralbyran, 1970. sarlvik, bo. qpinionsbildningen vid folkomrostningen. 1957 . (statens oftentliga utredningar, 1959:10) stockholm, 1959. sarlvik, bo. paitibyten som matt pa avstimd och dimensioner i partisystemet. sociologisk forsknine 5(l):35-80, 1968. sarlvik, bo. party politics and electoral opinion formation: a study of issues in swedish politics 1956-1960. scandinavian political studies vol. 2, 1967. sarlvik, bo. party profiles in the swedish electorate. a collection of data drawn from the 1964 elecuon study. [ts] paper presented at the conference on comparability in voting studies, loch lomond, scotland, julv 1968. sarlvik, bo. political stability and change in the swedish electorate. scandinavian political studies (l):188-222, 1966. sarlvik, bo. politisk rorlighet och stabilitet i valmanskaren. statsvetenskaplig tidskrift 67(4): 185-219, 1964. sarlvik, bo. the relation of partisan orientation to political engagement in the swedish muluparty system. [ts] paper presented at the international conference on comparative bectoral behavior, university of michigan, ann arbor, mich., april 1967. sarlvik. bo and olof petersson. rikspolitik och iokalpolitik i valet 1973 . (valundersokningar, rapport 1) stockholm: liberforlag, 1975. sarlvik, bo. skiljelinjer i valmanskaren, statsvetenskaplie tidskrift 68(2-3): 141-183. 1965. sarlvik, bo. socioeconomic determinants of voting behavior in the swedish electorate. comparative political studies 2(1):9 c>-135, 1969. sarlvik, bo. socioeconomic position, religious behavior, and voting in the swedish electorate application of computerized classification techniques. quality and quantity 4(1):95-116, 1970. sarlvik, bo. socio-economic predictors of voting behavior: research notes from a study of political behavior in sweden. [ts] paper presented at the third international conference of the isa research committee on political sociology, berlin, january 1968. sarlvik, bo. sweden: the social bases of the parties in a developmental perspective, in: richard rose (ed.)/ electoral behavior: a comparative handbook . new york, ny: free press, 1974. sarlvik, bo. the swedish elecuon study in 1968: a note on survey design & draft questionnaire. [ts] paper presented at the conference on comparability' in voting studies, loch lomond, scotland, july 1968. sarlvik. bo. the swedish party system in a developmental perspective. [ts] goteborg: statsvetenskapliga institutionen, gbteborgs universitet, 1971. sarlvik, bo. valet 1970. in: allmanna valen 1970 . part 3. (sveriges officiella statistik) stockholm: statistiska centralbyran, 1973. sarlvik, bo. valrorelse och valjarna. tiden 47(6):334-340, june 1955. sarlvik, bo. voting behavior in shifting 'election winds'. an overview of the swedish elections. 1964-1968. scandinavian political studies 5:241-283, 1970. westerstahl, jorgen, bo sarlvik, and esbjorn janson. an experiment with information pamphlets on civil defense. public opinion quarterly 25:236-248, summer 1961. westerstahl. jorgen, bo sarlvik, and esbjorn janson. fragor rorande civilforsvar och psykologiskt forsvar. [ts] beredskapsnamnden for psvkologiskt forsvar. 1958. westerstahl. jorgen and bo sarlvik. svensk hjalp till mindre utvecklade omr'aden. en attitydundersokning. [ts] centralkommitten for hjalp till mindre utvecklade omraden, 1957. westerstahl, jorgen and bo sarlvik. svensk valrorelse 1954. tva lokala studier. arbetsrapport i. [ts] goteborg: statsvetenskapliga institutionen, goteborgs universitet. 1955. westerstahl. jorgen and bo sarlvik. svensk summer 1987 iassist quarterly — 23 valrorelse 1954. lcke-rostning mediastudier. arbetsrapport ii. [ts] goteborg: statsvetenskapliga institutionen, goteborgs universitet, 1956. westerstahl, jbrgen and bo sarlvik. svensk valrorelse 1956. arbetsrapport i: intervjuundersokningen. [ts] goteborg: statsvetenskapliga institutionen, goteborgs universitet, 1957. summer 1987 vol30-2neu.indd iassist quarterly summer 2006 by by andrea janssen and jeanette bohr* microdata information system missy introduction in recent years, the number of official microdata sets accessible as scientific use files has increased significantly in germany. these microdata are of great interest to both economists and social scientists but are not, however, easy to work with. official microdata are surveyed to meet the data requirements of the german federal statistical office. the data contain special classification types which need to be documented for the user. users from the scientific community require more than superficial descriptions of a particular dataset; they also require detailed information pertaining to every variable in it. such an example of a german information system designed to fulfill the need of researchers, is the microdata information system (missy) presented below1. missy contains metadata or “data about data”, (jacobs 2006) about the german microcensus. missy piloted a project containing the descriptions for two census years 1995 and 1997. the next section introduces the microcensus and describes the functions and benefits of missy. the last section introduces the steps necessary to fully implement the system. questions concerning general rules for documenting and presenting metadata derived from the experiences in the first phase of the project are also addressed. the german microcensus the microcensus is the biggest continuing survey in germany. conducted annually since 1957 by the federal statistical office, it samples one percent of all german households or approximately 820.000 people. the main topics of the microcensus are occupation and qualification, labor markets and household and family structures. every four years the microcensus contains additional questions about health or housing conditions, for example. the large sample size and the broad scope of topics make the microcensus an invaluable data source for different scientific questions of varying complexity. for example, the microcensus enables one to examine higher education among relatively small groups of immigrants, e.g. italians or greeks. the microcensus cannot be accessed by the scientific community in its entirety, but the federal statistical office extracts a 70 per cent subset and provides it to researchers. the microcensus has attributes that are not commonly included in social sciences surveys: the classifications of professions (kldb – klassifikation der berufe) and economic sectors (wz – wirtschaftszweige) are used only in the official statistics and require some explanation to the researcher. another characteristic of the microcensus is its vast quantity of derived variables, the so-called “bandsatzerweiterungen und typisierungen”, whereby the latter are based on different concepts of families and living arrangements. the generation of these variables is not easy to comprehend; again they are only partially accessible to the scientific community. as a result, to work competently and efficiently with the microcensus data, the researcher requires information exceeding what a superficial description of the dataset can provide. for this reason missy was developed. missy missy is a product of the german microdata lab (gml), formerly named the department for microdata at the zuma (centre for survey research and methodology) in mannheim.2 since the 1980s, the gml’s focus has been on the microcensus and as part of this focus it has offered an array of comprehensive services to support use of the files. the gml, in collaboration with the federal statistical office, ensures that the procedures necessary for anonymizing the data to protect confidentiality are in place. all files are checked prior to being released to the scientific community and comprehensive documentation of the files is created. as well, the gml provides support to researchers by offering advice on both the methodology and the content. to facilitate work with, e.g., classifications unique to the microcensus, microdata tools are developed. finally, and of equal importance, the gml organizes user conferences and workshops to promote the advantages of the microcensus data for scientific research and to enable and increase the opportunity for scientific communication among researchers (lüttinger et al. 2004). 6 iassist quarterly summer 2006 figure i: missy missy facilitates research based on the german microcensus. gathering all the necessary metadata incorporating the knowledge of the gml is the first step; this includes official documents of the federal statistical office.the second step is one that connects all the metadata in a way that considers the textual relationships between the data and the enquiries of social scientists and economists. the implementation accomplished by missy is based on the ddi (data documentation initiative) 2.1 standard. missy is an exclusively german system; there are two reasons for it being unilingual. at first, researchers are forbidden from using the data abroad. the second and more important reason is that all documents and descriptions of the microcensus are in german making a knowledge of the german language essential. for demonstration purposes the most important expressions in the following examples have been translated. to classify the metadata type, missy utilizes the categories of sundgren’s dimensions. the metadata for the microcensus includes both pragmatic and semantic as well as syntactic aspects (fischer 2005, sundgren 2003). this means that missy encompasses data answering questions of why (pragmatic aspects), what (semantic aspects) and how (syntactic aspects). in the microcensus documentation, it is helpful to differentiate between general information about the entire study and specific information about the variables. the difference between the elements “study description” and “data description” is found as defined in accordance with the “data documentation initiative” (jacobs and thomas 2006). note that in the middle of the screen there are multiple access points for retrieving specific information. furthermore, there is a brief overview of the function of missy and there are also links to more information about the microcensus and missy. in addition to this “main entrance” for access to specific information, the short list on the left sidebar of the screen under the red header “variableninformationen” (specific information) can be used as well. the second list with the green header “allgemeine informationen” contains general information only about the microcensus: an introduction, questionnaires, codebooks, interviewer guides, frequencies and some tips and recommendations on working with the data. information about classifications used by the federal statistical office or the scientific community is included here. for an easier navigation of the missy pages all specific information has red headers and all general information has green headers. points of access the different points of access were designed to simplify the search for variables while recognizing the varying needs and skills of the users. the first point of access is a list of all variables, subdivided by census year. this is useful when information about a specific variable for a particular year is required and it is preferable that the user already has some knowledge of the structure of the microcensus. an easier way of access, albeit longer, is given by the thematic structure illustrated below: thematic access is appropriate when the researcher is http://www.gesis.org/dauerbeobachtung/gml/missy/ iassist quarterly summer 2006 7 figure ii: thematic structure interested in a specific subject or field of research and wants to know if the microcensus has relevant content. to start with, the researcher may choose from eleven topics leading to the secondary level; at this point there are two links. the first is to publications based upon the microcensus that pertain to specific subjects, “ethnic minorities and migration” being one example (see fig. ii). this makes it easy for the investigator to determine what research might be undertaken or see what research has already been done based on the microcensus. the second link connects to tables containing examples of analyses. again, using “ethnic minorities and migration” as an example, the user will find multiple tables including one which shows a comparison of the graduation rates between the german and turkish population. the tables were created to assist novice data users, e.g. students. the aim is to encourage researchers and future researchers to use official microdata in their analysis. the fastest method for obtaining specific information is via a matrix containing all variables for every year covered by the scientific use files (see fig. iii). the variables names in the matrix cells are linked to specific information about the variables. furthermore, the matrix provides an overview of characteristics surveyed in specific years. because the microcensus is conducted annually, it can be used to address questions requiring a consistent long-term view to observe social change in society. in order to examine longer time periods the researcher requires information about the comparability of the variables over a specific timeframe. the matrix presents the most important changes that have occurred in the variables. there was a significant change to the microcensus questionnaire in 1996; many variables were split into two, marked in the matrix. the changes are indicated and explained in the tool tips. when it is possible to generate comparisons of variables, links to spss-syntax are provided. with these instructions, comparisons between many of the variables, both before and after 1996, can easily be made. specific information: variables documentation for each variable, all available metadata are centralized on a single site. not only are the variable labels included, but also the text of the questions and related notations, if available. there is information about the guiding filters of the questionnaire or what attributes the respondent had to fulfill in order to be asked this special question. the value labels and frequencies give first impressions of the variable’s distribution. in the first lines of the variable description are links to shortcuts to detailed information of comparable variables for other years and for different levels of the thematic structure. these links are marked with red buttons to create visual consistency with the list containing the different possibilities of access to specific information. analogously, the links to general information that could be http://www.gesis.org/dauerbeobachtung/gml/missy/zugang/thematische_gliederung.d.html 8 iassist quarterly summer 2006 figure iii: matrix of interest according to the particular variable are marked with green buttons. with these links, the researcher will be directed to the exact reference in the questionnaire, the codebook or the interviewer guide that contain the information concerning the variable of interest. as stated in the introduction, the microcensus contains a variety of derived variables which cannot be tracked in their composition. information about the generation of these variables is documented for the internal use of the federal statistical office only and cannot be accessed via the internet. missy provides an additional link from the special information about derived variables that point to a site on which the generation of this particular variable is described. if the generation of the variable is based upon a special concept of families, households or living arrangements used by the federal statistical office, another link to a description of these concepts is provided. furthermore, researchers can go to a catalogue that contains definitions of the terms used in the microcensus. below is an example for a description of the variable “type of working hours of the reference person of the family”: this information makes it relatively easy to use even rather complicated variables in an appropriate way. conclusion what conclusions can be drawn about the implementation of an information system for microdata? first of all, the concept of the system requires knowledge and research experience with the particular data. the special characteristics of the data should be understood and adequately documented. secondly, knowing the data should make it possible to connect different kinds of information about the contents and thereby facilitate the search on particular topics. ideally researchers should find not only all information they are looking for but also other helpful information that they may not even know existed. another important point to appreciate after the first http://www.gesis.org/dauerbeobachtung/gml/missy/zugang/variablen_zeitpunkte_matrix.html#3.0.0.0.0 iassist quarterly summer 2006 9 figure iv: variable information implementation of an information system of course is to ensure a maximum of usability. the next step is to have experts and users of microcensus data analyze missy’s performance and make recommendations for improvement. following this, they are plans to extend missy by including all available microcensus scientific use files. two more specialized files will be added: the panel file and the regional file. because of the concept of the microcensus as a rotating panel the panel file would include four years of census microdata. the regional file contains microdata of a very differentiated regional level but on a less differentiated topic level to ensure the necessary confidentiality. to include these new types of data sets in missy adequately, some modifications to the information system will be required. another emphasis will be the extension of the category “tables” that provide an overview of the research possibilities using the microcensus. an exercisebased introduction into working with microcensus data is planned. with this concept the main focus of collecting and providing metadata for datasets will be expanded in a direction with implications for more practical and concrete advice for special problems that arise when working with microcensus data. the result will be that the proportion of metadata with syntactic aspects will increase in missy. references fischer, birgit (2005). metadaten in der amtlichen statistik im internationalen vergleich. berliner handreichungen zur bibliothekswissenschaft, no.45. berlin: institut für bibliothekswissenschaft der humboldt-universität zu http://www.gesis.org/dauerbeobachtung/gml/missy/klassifikationen/konzepte_und_definitionen/konzept.d.html@studiengruppe=mikrozensus&studienuntergruppe=grundfile&studie=mz1997&variable=ef605.html 10 iassist quarterly summer 2006 berlin. jacobs, jim (2006). looking into the future... iassist conference 2006 workshop presentation. http://www. iassistdata.org/conferences/2006/presentations jacobs, jim, and wendy thomas (2006). evolution of data documentation. iassist conference 2006 workshop presentation. http://www.iassistdata.org/conferences/2006/ presentations/ lüttinger, paul, bernhard schimpl-neimanns, georg papastefanou, and heike wirth (2004). the german microdata lab at zuma: services provided to the scientific community. schmollers jahrbuch 124, 455-467 sundgren, bo (2003). developing and implementing statistical metadata systems. http://www.epros.ed.ac.uk/ metanet/deliverables/deliverables.html * andrea janssen and jeanette bohr. contact: andrea janssen, centre for survey research and methodology, zuma p.o. box 12 21 55 68072 mannheim germany. email janssen@zuma-mannheim.de. footnotes figure v: explanation of the generation of a variable iassist quarterly summer 2006 11 1 http://www.gesis.org/dauerbeobachtung/gml/missy/ 2 http://www.gesis.org/en/social_monitoring/gml/index. htm iassvol201 4 iassist quarterly how to make osiris more interoperable while waiting for the ideal system by lennart brantgarde and leo rubinstein1, the concept of interoperability will in this paper be understood as a quality that makes a system flexible over time and adaptable in relation to other systems. it stands for high portability but also for high usability for different purposes. the osiris format was long ago sentenced to death. its deathstruggle has now been streched into an extremely painful decennial process and i would guess that there will be several years more to come before it is over. the reason for this extended process is not so much dependent upon an abundance of virtues of osiris but on the lack of alternatives. osiris is amazingly enough still the only format that easily can reach a high interoperability versus a variety of functions that you have to operate at a modern dataoriented archive. two contrary principles for a data archive are: 1. archive and store what you get and disseminate what you have 2. transform what you get into a format that enables you to disseminate what is wanted in terms of formats by the market. we suppose no one operates exactly on one of these extremes we are located on the continuum in between, some more to the first position, some more to the second. this paper deals with how to achieve a position close to the second while still using osiris as the basic archival format for preservation and documentation. as you know one of the basics of osiris is its oldfashioned fixed format. this makes it easy to program utilities for since every single unit of information always is found at the same spot. one way of making osiris more interoperable is to create utilities. and we would guess that a whole set of utilities have emerged at various data libraries over the years but unfortunately they have never been collected and distributed to the benefit of all. at our site we have therefore been forced into this kind of utility production. we are going to present five such utilities and show how you can get hold of the programs. from osiris type 1 to osiris type 3 in the first place we had to convert an osiris-type-1codebook into an osiris-type-3-codebook. the binaries in the type-1-codebook were of no use in the long run at least if you want full control of the file. we created a program that is called 123. this program comprises two things 1. conversions of binary numbers to decimal codes. 2. conversion of character codes, usually from ebcdic to ascii as far as the binary numbers are concerned a complicating fact is that some data entries in a type-1 codebook are stored binary in 2 bytes and can have values between zero and 65336. the same information in a type 3 codebook has only four digits reserved and they are supposed to be stored decimally. consequently no higher numbers than 9999 can be generated. our solution to this was the gv-convention. this comes into consideration only for numbers above 9999 and consists of a hexadecimal kind of system where zero thru nine and capital a thru capital f have been replaced by the minor letters g thru v where these letters stands for zero thru 15. in other words base 10 has been replaced by base 16 which allow us to take care of all binary numbers without losing any information and at the same time preserve the basic format of type 3. the second problem moving from character code representation to another was solved by using, as a second parameter, a name of a file containing a translation matrix, a table with 256 lines. it is up to the user to define this table or modify existing ones. in such a table there are just two columns where the left one has entries equal to the line number -1 and the right one tells what character should replace the former. value -1 in the right column means that the code is undefined. when encountering such undefined values the program shows the partially converted line and asks for a character to put in. the program can be initiated by the command: 123 codebook [table] where codebook is the stem of name of the codebook file and table the name of the translation table file. the program assumes that infile is codebook.os1 and writes the outfile codebook.os3. if the optional parameter table is left out the program assumes that the existing internal table named ascii 5summer 1996 should be used. making an eyefriendly document out of an osiris-type-3codebook. the original basic fixed formatted osiris codebook file is not very useful for having as a readable document on the bookshelf. you have to get rid of all abundant figures and rearrange the information into a userfriendly textbook. this is done in our environment by a program called kbl the program is initiated by the command: kbl infile outfile [dumpfile] where infile defines your osiris-type-3-codebook-file and outfile defines what file you want to get out. in a unix environment the outfile can be browsed using the unix utility more and printed using lpr . moving an osiris-type-3-codebook into world wide web or auis. a modern data library or a modern archive has to consider desktop browsing of the archival holdings. this is absolutely necessary for an archive that has to deal with a scattered academic community located miles and hours from the office. to our help facilities like www, netscape and mosaic have emerged like benevolent fairies out of the dark. like the original fixedformatted osiris the modern www is carrying an heritage from the childhood of computing. to get a text into good order in a web you have to tag it into html which is a sample of sgml. the next utilityprogram to be presented here tranforms an osiris-typ-3-codebook into an html-tagged document that, put out on the server, not only makes use of bold characters and italics but also automatically creates internal weblinks between the table of contents and the rest of the codebook document. in one stroke a stereotype osiris codebook is turned into a decent modern eyefriendly document and a fullyfledged hypertext document. the program is called htkbl and is started by the command htkbl infile outfile [dumpfile] exactly repeating the moments of the former program. dumpfile is the name of a file to where wrong lines in a codebook (infile) will be copied if any. for the auisenvironment a utility program called atkkbl can be used the same way. getting from osiris to spss and further out so far we have just talked about utilities facilitating the infomative functions of an archive to get osiris ready for printing and make it exposable through modern networks. fundamental to an archive is to disseminate both data and documentation. in order to facilitate distribution of data in a variety of formats we have created a utility that enable us to go automatically from osiris-type-3 to a portable systems file in spss. the program is called otosp and is started thru command otosp codebook setup datafile exportfile where codebook is the name of your osiris codebookfile (type 3), where setup is the name of file including all spss control cards -including value labels , where datafile is the name of your osiris datafile and where exportfile is the name of your outfile created by spss. the program can also be started by its mere name otosp in which case the program will be prompting questions for all other parameters. once in spss there are other conversion programs that can bring you further out into the djungle of statistical packages. how do i get hold of these utilities? all of them are stored under /pub/programs in ftp://oden.ssd.gu.se you are welcome to pick what you want. source codes to 123, otosp, kbl, atkkbl and htkbl are free to be copied , compiled and used. but it is to be noted that these programs were created to suit the needs and requirements at the ssd and so far they have been tested only here. you should pay some attention to readme-files in subdirectories for improving your chances to install and customize the programs for your environment. for those of you who wou ld like to have a training example we have also made available a codebook of the oldfashioned icpsr type, the codebook to almond and verba the civic culture. this exist in a type-1 binary version and is supposed to be converted to a typ-3 ascii version and further on to www and spss. for installing copy a complete subdirectory respectively to your computer read the read-me file modify the make file: change the dir-string to the name of an appropriate directory on your computer (in search path for executable files). run the make file. further questions should be addressed to: leo.rubinstein@ssd.gu.se good luck! 1. this paper has been presented at the css96/iassist conference at minneapolis, university of minnesota, may 12 19, 1996. it describes a set of programs that have been built around a major documentation system called a-side developed by stephen greene for swedish social science data service and presented in stephan greene: a functional approach to documentation and metadata, iassist quaterly vol 19 no 1, spring 1995 pp18. support for this presentation has been received from the swedish institute, stockholm, grant 989/301/40 lassist newsletter vol. 4 nos.3&4 a strategy for archiving government data to meet the needs of the research community dr. jake knoppers senior advisor (information management) public archives of canada background govemmenl administration and survey data as a research resource by far the greatest coheclor and user of information is the government the largest portion of the data collected is for administrative purposes, eg. collection of taxes; distribution of socio-economic benefits; regulation of industry, trade and commerce, etc as a result of these activities, the government builds and maintains enormous stores of information or data banks. this information is collected continuously albeit with some periodicity, eg annual tax returns, monthly filings in addition to the information which is collected pursuant to acts of parliament and initiatives of the government, the various departments also carry out innumerable surveys, or commissijn surveys and opinion polls to be earned out by third parties. altogether, statistics canada has identified 2,352 major data banks so far, including both manual and machine readable. (1) some data relates to individuals while other data refers to larger aggregates such as firms, households and national accounts the number of subjects in each data bank ranges from as few as 200 to as high as 21 ,500,000. as a matter of fact, there is a growing recognition that the various departments of govemnient may have gone too far in their mynad data collection activities the government has responded to these concerns through the "reduction of paperburden" initiative this finds expression in such management tools as forms control, data collection approval mechanisms, and data locator systems. perhaps the desire to reduce paperburden, while meeting the needs of the research community, can be dealt with more effectively if one institutes a new strategy for archiving these government data in a rational and consistent manner. the needs of the research community because of the nature of the research scientist's work, his concern for access to government records is both highly selective and strongly motivated neither the research scientist nor an informed public can rely solely on the summary tables produced by government departments and statistical agencies nor do such tables of aggregated statistics take into consideration the concerns of all scholars. it must be possible for the public and the research community 10 obtain the raw micro-data, conduct independent analyses, and draw their own conclusions since the major portion of the data collected by government for administrative purposes involves privacy considerations, ways must be found to balance the freedom to do research against the individual's (including legal persons') right to pnvacy this principle is stated quite succinctly in a resolution of the social science federation of canada, passed at their may 17, 1979 meeting the resolution reads: the opinions expressed in this paper ; of canada, for which he prepared it : "there are socially significant fields of research for which access to personal records is indispensable. there is, therefore, a need to use personal data held by government agencies, for statistical and research purposes, in order to promote scientific understanding of important contemporary problems. this use of govemmenl data is not incompatible with the need to protect the privacy of individuals. therefore, any federal or provincial laws for the protection of personal privacy or for access to government documents should make a clear distinction between administrative or regulatory uses of personal information, which directly affect a person, and statistical or research uses, which do not, and should explicitly recognize the legitimacy of using personal data for statistical or research purposes accordingly, provisions in these laws should set out the right of researchers to obtain access to personal data under specified conditions, and should specify these conditions, the most important being a wntten undertaking not to reveal data on specific individuals without their express consent those of the author, and are not intended to reflect the position of the public archives a consultant, nor that of any other canadian govemmenl department or agency [60] lassist newsletter vol. 4 nos.3&4 such laws bhould also provide, rn case such access is refused, a right of appeal lo an independent authority, such as an ombudsman, or a court, or preferably both." (2) finally, the data needs of the research community and those of the government as vvcll would be met if it were possible lo develop a longitudinal, rational, and integrated master sampling frame of govemmeni administrative records the creation of such a mechanism and entity would also lead lo a high probability of reduction of paperburden and extraneous survey data collection. the role and responsibility of (he public archives o( canada the broad mandate of the public archives of canada is lo collect public (i.e . government) records, documents and other historical material of every description which reflect canadian society from iis beginnings to the present day the public archives has the overall control responsibility for the "life-cycle" of all records, once their creation has been approved. the components of the "life-cycle" of records include: identification, registration and classification; care and custody; controlled circulation; the esiablishmeni of retention and disposal schedules; and final disposal through destruction or permanent retention records include all diferent kinds of physical storage media for information, eg . paper, electronic, film. it is the responsibility of the dominion archivist to identify and appraise all govemmeni records as to their historical or research value, granting permission for the destruction of those records which are of no archival value and thus not eligible for permanent retention the gathering and preserving of archival records is done with the purpose of making them available to the govemmeni. researcher and the public. the role and responsibility of the public archives to the research community is that the desired micro-data found in govemmeni administrative records and surveys is preserved for use. an advisory council on public records in which the social sciences are represented provides a vehicle whereby the research community can make its wishes known. o) the political and legal environment present legislation and public attitudes one of the rights which has been recognired in law in many countries is the right to information privacy, i.e.. that individuals have the right lo be informed about the storage of personal data about themselves and lo control the collection, use, and dissemination of such information this definition is grounded on the conviction that, in the end. information about individuals belongs to the individuals themselves — indeed, such information is the essence of individuality. while recognizing that an individual divulges information about himself for a purpose — in exchange for a good, service or benefit, or as required under law — this approach holds that the information is nonetheless his, and thai he retains rights with respect to it. in general, with respecl to the operations of government, privacy legislation is designed to enable individuals lo control how, when, and to what extent information about them is communicated to others — especially where such communications are related to the decision-making or administrative processes which affect them personally. consequently, privacy legislation normally contains provisions designed to control national or federal data banks, with controls on collection, storage, dissemination, retention and corrections of personal information. in many countries, the exercise of the rights under privacy is assisted through requirements to list or report on all personal data collection activities the emphasis in many countries is on the surveillance of compuierized rather than manually recorded information much of the pressure for greater openness in govemmeni has been relieved by legislation on privacy, even though this aspect has been treated separately from disclosure of govemmeni information of the more general kind. in canada, privacy legislation finds its expression in part iv of the canadian human rights act it requires that after march i . i98t) the government shall inform data sources who knowingly provide information to a govemment institution, during the course of collection, the purpose and use to which the information will be put (4) this rider applies specifically lo administrative uses of information for decision-making purposes which impact directly on the individual concerned apart from privacy concerns, many other acts of parliament set stringent conditions on access lo their records some acts stale specifically which officers or which other departments can have access to their data most of these access restrictions for administrative purposes are presently also used to deny access for research purposes and in some instances even to deny access to archivists when the latter wish to appraise the historical or research value of records. required legal environment before outlining some of the basic legal principles which parliament should embrace through legislation, the research community would do well to take note of a comment in the decision of the u.s. supreme court in the "kissinger case." (5) the justices quoting the senate report to the federal records act of 1950 highlighted that: "it is well to emphasize that records come into existence, or should do so, not in order to fill filing cabinets or occupy fioor space, or even to satisfy archival needs of this and future generations, [one would assume that this means the needs of the research community] but first of all to serve the administrative and executive purposes of the organization that creates them there is a danger of this simple, self-evident fact being lost for lack of emphasis, ,," (6) lassist newsletter vol. 4 nos.3&4 the stalemenl applies universally to administrative records of all organizations, including those in the public sector changes in the legal environmeni which would assist in meeting the needs of the research community must at the least be compatible with administrative requirements fortunately. a legal environment which stresses efficient and costeffective management of recorded information will also, as we shall see, encourage the making available of administrative data for research purposes. the legal environment required lo allow for access to government administrative and survey data should embody the following concepts; • that no government record may be destroyed or altered in any form without proper authority; • that the government be able to identify, inventory, and describe all their information holdings; • that the government be able to identify and know the information contents of all its data collection activities; • that institutions of government be permitted to disclose government records under their control lo any scholar or research institution for research purposes; such records to be in both identifiable and anonymized microdala and in original and complete form without legal barrier or statutory exemptions. • that the government accepts the principle that the security sensitivity of classitied records declines with the passage of time and that declassification schemes are established and implemented. • that for all government records a "life-cycle" is established and adhered to e . records retention and disposal schedules). • that the research community apart from their rights as individuals under privacy and as individuals or groups under freedom of information be granted the right to advise the government in its decision-making process pertaining to the disposal (i.e. destruction or permanent retention) of government records; and • that the dominion archivist can effectively declare as archival any government record (or a copy), to be kept permanently for historical or research purposes. on the researcher's side a comparable legal environment must also be created. such an environment would be based on the principles of ethics in research which inter alia would also include the nght of privacy and protection of the individual where personal information is involved researchers might want to consult government records that are not normally accessible to the public. one would expect that when access 10 such information is granted for legitimate research purposes, it would be granted only: • if the research subject has consented lo the intended use. if the vested interests of the research subject are not harmed or involved because of the type of information, general public knowledge of ihe information, or the type of data processing involved; or if the researcher agrees in writing not lo reveal, publish or otherwise disclose information which would make it possible to identify any individual person, business, or organization except "\" years after ihe birth of an individual, or when the individual is dead or ".\" years after death, or ".x" years after ihe taking of a census or survey, or "x" years after the receipt of information relating to a business, institution or organization outside of govemmenlthe reason for the "x" years is thai these are policy decisions for the govemmcnt lo make, hopefully in consultation with ihe research community. however, the written undertaking or contract between the researcher and ihe government not to reveal information in any form thai could reasonably be expected lo identify ihe research subject. 1 e,. guaranteeing anonymity, would cover over 91)% of ihe research needs should an individual researcher not agree with such conditions, one would expect him to raise the question of public access (as distinguished from access for research purposes under specified conditions) either through the path open to him under freedom of information legislation or through the consultative channels between the research community and the government the above is an outline of the basic political-legal environment for government records which would ensure that both administrative and research needs can be met the reason for the statements made in this section wil become apparent in the model solution a model solution basic approach the public ai large benefits from legitimate uses of administrative data for research and analysis by government, businesses, non-profit organizations, academics, etc. such data is used to analyze the effectiveness of the delivery of existing socio-economic programs; to study the causes of disease, poverty, crime, or migration; and lo discern trends in society which may be of interest lo ihe nation as a whole, to particular interest groups, or lo the individual researcher. yet ihe public is also concerned about ihe increasing burden of providing the required information ("paperburden") and about the real or imagined possibility of the misuse of the daia provided by them funher. at the moment, the legislation affecting the use of adininisirative data for research purposes is far from clear and on the whole presents a negative approach instead of dealing [62 lassist newsletter vol. 4 nos.3&4 with all these laws and regulations on a casc-hy-casc basis, the model solution proposed here presents a glohal approach which nevertheless should be applicable on the detailed level. the success of the approach taken here depends, however, on the establishment a priori of the following overall requirements for government records • the government must establish the information in its possession (physically located in one of its institutions) over which rr has ownership of the crown in right of canada and the condition of such ownership this will also be a requirement for information falling under fol, privacy and archives legislation, and is of particular relevance to information created as part of joint ventures of the government, e.g., federal-provincial, as well as for information created or collected as a result of the contracting out of work using public funds • no government record may re destroyed or altered in any form without the consent of the dominion archivist or his designate. and improper destruction must call for automatic penalties through the establishment of authorized records retention and disposal schedules, the dominion archivist ensures efficicnl and cost-effective management of government records, destroying those records of no further use or value and selecting for permanent retention those records having historical or research value • the dominion archivist must have the right to declare any government record or a copy thereof an archival record and where a copy is concerned take early possession where efficiency or the nature of the record dictates the greater proportion of administrative data are created as a result of ongoing programs such data ac generically known as case hies, i.e.. nies containing records related to specihc repealable actions, events, persons, organizations, products, objects, etc and which are usually filed and retrieved by name, number or any other systematic identifier for the sake of administrative efficiency and timely delivery of programs the government has in recent years resorted to technologies which help it reduce the "paper-shuffling." information received from individuals in paper form is microfilmed, coded and maintained in higher density storage media a long letter from a program participant is reduced to a single computer change-order slip; old addresses and records of those no longer participating are. in due course, deleted automatically normally when a file such as that of a canada pension plan beneficiary finally reaches the archives, all that might be found inside the file jacket are a few computer change orders, the current address and the current benefits. information on that individual's fifty or sixty years of interaction with government will not be there since modern storage technologies lend themselves quite readily to copies, the historical and research interests would be met if procedures were instituted so that at least a complete microdala sample of such administrative data would be deposited in a continuous fashion with the public archives. in return for authorizing the alteration of records from one storage medium to another, eg, paper to microfilm or machine readable form, the dominion archivist would retjuire a sample to be forwarded to the archives. sampling schemes adopted for such series of case files or data banks should; • be applicable to the paper, micrographic and machine readable components of a data bank; • be consistent with the technology used for processing the information of that data bank; • be cost-effective and administratively implementable, given the human and financial resources at hand; • be statistically acceptable, • allow for longitudinal studies, ie , be periodic and consistent; • allow for comparative studies, i.e., be integrated with other data banks so as to create as high a probability as possible in capturing information on the same data subject from different data banks the last point touches on the problem of record linkage record linkage is the process whereby data from different record systems are linked on a case-by-case basis, on grounds that any of the single record systems is incomplete either with respect to data or required coverage or both record linkage is used to "reconstitute" large portions of (past) populations or to consider "statistically" any large population or sample in a multivariate context. for most large data banks the basic file series is organized by unique identifiers such as the social insurance number, corporate taxation number, or like unique systematic numerical or symbolic identifiers. derivative file series, e.g.. frauds and persecutions, appeals, are organized either as part of a mam file series with colour coding or as a separate file series, often in alphabetical sequence other ca,se files whose existence is the result of voluntary rather than mandatory participation are often organized alphabetically within some subject, geographic region, or industrial or product coding schemes. the approach of the archivist has always been that of sampling. it is estimated that on the whole the public archives presently selects for permanent retention less than 59r of all the paper records. case files can be selected for permanent retention on the basis of any of the following criteria; • a case file may be important for the issues involved, due to the issue itself or the context in which it occurred; i 63 i lassist newsletter vol. 4 nos.3&4 • a case file may be regarded as important for its mfiuerice in the development of principles, precedents, or standards of judgement in such matters as the definition of the jurisdiction, operations, and mandate of the department concerned; • a case file may be regarded as important for its contribution to the development of methods and procedures; • a case file may be regarded as important for its documentary/illustrative value; • a case file may be regarded as important for its research value, which even if minimal for one case file would be high enough to warrant permanent retention if a sufficient number of case files were selected basically, the archivist uses two general kinds of sampling techniques, "judgement sampling" and "probability sampling " in both cases the archivist selects for permanent retention a collection of records on only a few members of the universe. a more common term for judgement sampling is selective retention, while probability sampling or just "sampling" is a technical term for a procedure whereby one selects a number or "sample" of items from a defined "population" of items in such a way that every item in the population has a known chance of being selected. for example, if one decided to sample by terminal digit 5 of the social insurance number one would know that the probability of any individual being in the sample would be 1 out of 10 or lu9f should one instead decide to sample first letter of ti.e surname, for example by the letter "r", the probability would be 1 in 2u or 59t. the exact details of sample scheme for administrative and survey data will be presented at a later date. (7) if one can now assume that the dominion archivist has acquired administative data banks in loio or in sample form, these are some of the basic parameters that might be applied to the release of microdata under controlled conditions: • there must be a legitimate and important research purpose to be served by the process; • the researcher must sign a written contract specifying the degree of detail below which information taken from the microdata may not be disclosed; • al time of the conclusion of the research project or no later than some specified date, the researcher, if granted temporary pos.session of microdata, must either return such data to the public archives or submit an affidavit as to its destruction; • there must be a prohibition against dissemination of the microdata to a third party without the written authorization of the dominion archivist; • the researcher must submit a copy of the publication containing the data derived from the microdata to the public archives and in some cases prior to publication; • significant and mandatory sanctions or penalties for improper disclosure of microdata would have to be founded in law; • the researcher must have available an ombudsman mechanism to deal with confiicls relating to the terms of a research contract before it is signed; where the researcher is allowed to take the microdata to his own facilities for processing and analysis, the facilities should provide a level of security comensurate with that required for the microdata a sample application—the federal government in terms of actual application, we can consider the following (hypothetical) example for a socio-economic program (sep) the program is about to shift from a totally paper mode to a multi-storage media approach, i.e , paper, microform and edp. the program consists of five different levels of file series, namely: title organization benefits & claims by terminal sin digit appeals ditto, colour coded frauds & persecutions alphabetic judicial decisions alphabetic medical files by terminal sin data n of participants 22,()00,1x)0 1,i()(),(kk) :20,()0() 44.0(ki 2.(xk).()ik) [64] lassist newsletter vol. 4 nos.3&4 the adminisiralors of the program propose to destroy all paper records for tile series 1,2 and 5 upon receipt, microfilming the "relevant" portions instead they further propose to destroy all files for file series i and 2 two years after last action, and file series i and 5 five years after last action, while maintaining only summaries of 4, the judicial decisions after some deliberation, the dominion archivist approves the records retention and disposal schedule subject to the following limitations: : administrators are lo depositfor file series i the programn at the public archives, • a 01% sample of all paper records consisting of those participants whose sin number ends in 5555, • a 17c sample of all records microfilmed consisting of those participants whose sin number ends m 555, • a lo'/f sample of all edp records consisting of those participants whose sin number ends in 5. for file series 2, the samples for(a) would be 1%, eg all red colour coded file jackets where the sin number ends in 55 for (b) and (c), no separate samples would be taken as it is assumed that for (c) there would be a code for "appeals" in the edp record for file series 3, a svr sample would be taken of all those whose first letter of the surname starts wiih an "r " (note: 10% of the "r's" would also be found in the level 1 , edp samples this would ensure high linkage probability ) for file series components. 4, the sample would consist of two • those case files for which the judicial decision was of special significance because of person involved, nature of case, precedents, etc., • a 10% sample consisting of the lener "r" (4.94%) and the letters "a" (2.92%) and "n" (1.69%). while the letter "b" for example would give a 10% sample directly such a sample would not be "linkable" to file senes .1 one should therefore approach alphabetic sampling in terms of "common building blocks." in our hypothetical example, the department agrees with this sampling plan, knowing that the public archives will abide by the access and disclosure provisions of the legislation pertaining to this data. some of the data can be made readily available in anonymized form, eg., the edp portion, and the remainder will become available at a later date. this assures a respect for the totality of the "fonds " now this is a single case. if the same approach were applied to cover all the data banks, a national archival sample would be created the matter of a national archival sampling strategy for administrative data would include advice to the dominion archivist from such agencies as statistics canada, the national research councils, and the research community. in order lo maximize the probability of record linkage for longitudinal and comparative studies, a national archival sampling strategy would have to be applied uniformly lo all administrative data. the immediate benefit would be a reduction in survey costs and resulting paperbucden it may well happen that a government institution "a" wishes to use the data collected by another department "b" (administrative or survey data) for research purposes. however, the legal questions involved may take some time to resolve department "a" makes its case to the public archives, which in turn requests department "b" to deposit a copy of the data in question, e.g., a machine readable data set and documentation, in the public archives, which will hold the magnetic tape until the legal differences are resolved. such a mechanism might be of particular interest to government departments, in thai ii provides for a single "neutral" depository of administrative data, thus ensuring that longitudinal series can be created without a sudden hiatus caused by some legal problems of direct transfer from "a" to the public archives would be the neutral repository of these samples, which would be subject to the transfer conditions and prevailing legislation. when demand from the research community warrants, "public use files" would be prepared. in other instances, a research file would be created from the master file. the public archives may even decide to split the master file into two separate components, removing all unique identifiers lo a xref file and substituting a control number instead the research community would peruse the federal government inventory of data banks (required under foi and privacy legislation) and the national archival sample to determine whether the administrative data being archived (and the variables contained therein) met their research needs if not, a researcher or a research project could make its wishes known to the dominion archivist for the permanent retention of certain administrative data, even though legal and financial problems pertaining to access for research purposes may take some time lo resolve. the ombudsman mechanism could be fulfilled through the creation of a special committee of the already existing advisory council on public records consisting of representatives of the learned societies and like representative bodies. references: 1 ) federal inventory annual report, march 1980, prepared by federal inventory data base group, statistics canada. for further information contact mr. dave sally, federal data bank manager, statistics canada, ottawa, canada, k i a 0t6. 2) as cited in j knoppers, "a freedom of information act and the future of social science research", social sciences in canada. 7 (december 1979) 4:9-10 [65] lassist newsletter vol. 4 nos.3&4 3) at present the canadian historical association and the canadian pohtical science association represent the research community on the advisory council on public records. the public archives is giving serious consideration to widening the diversity of representation of the research community 4) the full text of part i v of the canadian human rights act can be found as an appendix to the 1980 index of federal information banks which is available for reference in every post office and canada manpower office across canada. 5) kissinger vs reporters committee for freedom of the press et al.. supreme court of the united states. march 3. 1980. no 78-1088 and 78-1217. the case involved summary or verbatim transcripts of kissinger's conversation notes and telephone notes which he "unlawfully removed" from the state department and al a later date deposited al the library of congress with severe access restriction. in short, the supreme court ruled that the plaintiffs had "no standing" and that only the national archives and records service and/or the state department could pursue kissinger for return of the notes. 6) senate report to the federal records act of 1950. s rep no. 2140. 81st congress. 2nd session, at 4<i950). 7) the distribution of sin numbers by first letter of surnames for r is 4 94%. the author is currently undertaking a project to establish an overall mixed numeric-alphabetic sampling scheme for case files falling under privacy legislation. this includes a study on the distribution of first letters of surnames of language, ethnic and provincial groupings the conduct of user surveys dennis d. mcdonald, ph.d. king research, inc washington. dc i'd like to start by making a few statements in general about user surveys, and then to discuss their goals, their different varieties, and finally some methodological pointers first, there's no such thing as an "end user" second, information isn't a commodity like a bar of soap which can be bought, sold, and priced "over the counter." third, don't believe that users and potential users can tell you pointblank what they really want and need in the way of information services. finally, the more you know about a user before you conduct a user survey, the better off you'll be user survey goals can be classified into the following five categorie:s: * input prior to system development. you may want to find out the needs of a potential user group before you invest a lot of dollars in the development of systems or services. this kind of user survey is difficult to conduct since what people suy they may need and what they actually end up using may be two entirely different things. information about your competition. you may be in the process of developing a system or service, and you may want to assure yourself that related information products and services, i e., your "competition." are not satisfying the same needs. this kind of survey is essentially a form of intelligence gathering, and conducting it will force you to consider how your product or service will supplement what is already available identity of current users. you may have a system in operation, and you may want to find out not only how and why people are using your system, but whether they are satisfied with what your system is supplying this is what most people consider to be a "user survey." potential users. you may have a system which is operating and satisfying the needs of a certain core group of users but you want to expand the system's use you must then seriously think about potential user populations and how to ask 66 ] 8 iassist quarterly 2013 iassist quarterlyiassist quarterly abstract sue a. dodd was active professionally in efforts to define and describe social science data files for library catalogs. her two books on guidelines for cataloging data files and software were important contributions to our understanding of the core elements used for this purpose. numerous reports and documents were produced through her work with iassist as part of the classification action group. this overview discusses a selection of sue dodd’s published works and the following list of references are discussed individually. also included are references for works of significance which relied heavily on sue dodd’s research. keywords: cataloging, bibliographic control, metadata, library catalogs it is important to remember that until the work carried out by dodd and others there were no widely known and systematically organized catalogs, inventories, or bibliographies of data files extant in the u.s. some work had been discussed through the international social science council (issc), largely due to the efforts of stein rokkan, and in the u.s. through the council of social science data archives (adams 2006). a publication calling for bibliographic conventions and standards was written by david nasatir under contract to unesco to study “overcoming the barriers to realizing the fullest utilization of machine-readable social science data.” nasatir called for not only the preparation of bibliographic details, but also for archives to enable variable-level searching across studies and across archives (nasatir, d., 1973, p. 3 and p. 46). the following items make interesting reading; many of the topics discussed are issues the data archive community is still concerned with today. a few details about these early days are illustrative. for example, there is mention of the “data explosion” to describe a situation not unlike the way today’s authors refer to the “data deluge.” and it should be noted that at that time, social science data primarily consisted of surveys, enumerations, public opinion polls, and some administrative records. mainframe computers were the only electronic tools available to the researcher for carrying out statistical analysis. today’s variety and size of file formats, including videos, images, simulations, games, etc., were not yet in the mainstream of what we now think of as research data. still the volume of material even then was notable. in addition to concerns about the amount of data available during the period when sue dodd was active, there was much discussion about the potential role for libraries in data management and calls for action. even so this concept was not new and had been expressed since at least the late 1950’s when, as nasatir mentions, “york lucci and stein rokkan proposed a library centre of survey research data in a project sponsored by the school of library service at columbia university in 1957.” (nasatir d., p 10). and library leaders such as clifton brock, whose 1967 work is quoted by nasatir, state “[s]ocial science data archives have developed entirely outside the scope of library systems,” and that “scholars, organizations, and publications” … [in the “data sector”] … “are wholly outside the world of conventional librarianship” (brock, 1967, p. 305). brock’s social science data files and bibliographic control: contributions sue a. dodd by libbie stephenson1 iassist quarterly 2013 9 iassist quarterly lament is echoed in a speech also described in nasatir’s 1973 work in which ralph bisco in 1967 spoke about “why should university libraries undertake data services when social science data archives are already providing them?”(bisco, 1967). further, the national academy of sciences, wrote nasatir, conducted a study headed by phillip e. converse which concluded that “few research libraries are adequately staffed” … [and] “libraries appear overwhelmed by information revolutions on other fronts” (national research council, 1967, chapter 3). so it was within this environment that sue dodd began her work. the shift in thinking about roles and responsibilities for management of social science data files can be found in sue’s writing in drexel library quarterly, journal of the american society for information science, journal of library automation, library trends, and library resources and technical services. these pieces are illustrative of increasing receptiveness of libraries to focus on the bibliographic aspects of describing and classifying social science data files. there were significant activities within the american library association to develop rules and guidelines for cataloging data files, and the machine readable catalog (marc) format. but it would seem that the report on the conference on cataloging and information services for machine-readable data files: march 28-31, 1978 had an impact that ensured the catalog rules would be implemented more widely in libraries. and sue’s books, cataloging machine-readable data files: a interpretive manual in 1982 and cataloging microcomputer files: a manual of interpretation for aacr2 in 1985, aided even lone catalogers in small libraries and archives to create accurate records for their catalogs of holdings. the following listing contains entries for publications, reports and other documents produced by sue dodd and her colleagues and cover the period between 1977 and 1990. also included are pieces containing substantial reference to sue dodd’s work and her role within iassist. these works include sue dodd’s books on cataloging data files and cataloging software for what were then called micro-computers. there are reports she prepared for the iassist newsletter (now known as the iassist quarterly) in her role as a member of the iassist us classification action group. sue was active in professional organizations within the american library association and presented her work at meetings which were critical for their impact on and advancement of bibliographic identification of data files. there have been a number of developments in both libraries and archives since sue dodd’s works were published. in the library cataloging realm barbara tillett’s functional requirement for bibliographic records (frbr) entity-relationship model provided a new way to think about the links between bibliographic records for all types of materials and is concerned with entities, relationships and attributes (metadata) (tillett, 2004). a new set of guidelines for cataloging will replace the anglo-american cataloging rules; resource description and access (rda) will provide rules and instructions on recording data to reflect attributes and relationships associated with the entities defined in the frbr2. further developments have led to the resource description framework (rdf) model which extends the utility of entity-relationship models such as frbr. and rdf links to metadata schema have been demonstrated. for example, stefan kramer et al. within the ddi alliance, have produced a document on connecting “rdf-described datasets to other related resources … in the semantic web … more specifically, … to leverage the data documentation initiative (ddi) to enable semantic linking of social science data to other data and related resources on the web” (kramer, et al., 2012). work to ensure that datasets used in research are properly cited, to develop item-level, actionable metadata schema, and to link data to published content as well as to other resources has been driven by some of the same practitioners who were guided by sue dodd’s earlier efforts and by many newcomers to the field. one likes to think that efforts to consider the possibilities, build the theory and applications, and develop the tools available today have at least a kernel of kinship with the ideas and directions defined in sue dodd’s work. perhaps it may be said without too much hyperbole that this was her expectation given that she wrote in the iassist newsletter in 1978, “[s]uccess is invariably measured not by what you hope to achieve, but by what you are able to produce. at the same time, the sum total of the lessons learned and the refinements made in the initial stages of any new endeavor becomes the foundation for future successes. the work described here is still developmental but it is designed to be expanded and implemented by other parties.” (dodd, 1978, p. 37). selected bibliography of books, articles and citations (text extracted directly from works is in quotes) dodd, s. a. (1977, january). cataloging machine-readable data files: a first step? drexel library quarterly, 13, 48-69. philadelphia: drexel university. 1977 commentary: “libraries are traditionally well equipped to handle … informational needs, in that they have standardized procedures for maintaining bibliographic control on a multimedia collection of materials. coupled with a recent commitment to automated systems, library procedures offer an important way of dealing with the existing problems of organizing, classifying, and cataloging information on social science data files. this paper will focus on the feasibility and future implications of applying library procedures and standards to mrdf, starting with the most important step- cataloging.” dodd describes the process by which rules and guidelines for cataloging data files were developed through american library association committees and library of congress. the iassist classification group work and projects are outlined. this article includes one of the first calls for national level efforts to “establish a national program of information services … through a shared and cooperative network.” one could easily repeat dodd’s insistence in 1977 that “the time has come to focus our attention on our national and commonly held responsibilities.” dodd, s. a. (1977, february). report of the joint canadian-united states action groups on classification (cag). iassist newsletter, 1(2), pp. 9-11. commentary: “agenda topics: discussion of useful areas of concern and future coordination of tasks; 2) review of cataloging efforts to date … including discussion of the working manual for cataloging machine-readable data files [mrdf] compiled by sue dodd; 3) discussion of the ramifications of mrdf catalog records, such as the national union list of social science data; shared cataloging; cataloging-in-production; the marc ii record as a standard format for storing automated bibliographic records …; 4) practical exercise in applying subject headings and descriptors for several large and uniquely held datasets, with a view towads compiling the beginning of the thesaurus or authority list of social science 10 iassist quarterly 2013 iassist quarterly terms for data files; 6) discussion of the lack of adequate subject headings and sub-headings currently provided by the library of congress for social science data files, with a view toward providing constructive recommendations.” dodd, s. a. (1977, may). cataloging and classification of machinereadable data files: a preliminary repo[rt] on the us iassist classification cataloguing project. iassist newsletter, 1(3), pp. 23-24. retrieved january 2014 from <http://archive.org/stream/ newsletterserial13inte/newsletterserial13inte_djvu.txt> commentary: “the primary emphasis of the us classification action group of iassist has been on establishing standards and on the study of library information systems as they may apply to social science data files. some of the recent developments within the library system which the classification group is examining include: 1) the development of rules and guidelines for cataloguing machine-readable data files (mrdf); 2) the development toward the acceptance of the marc (machine-readable catalog) record format as a universal standard for the automated bibliographic record; 3) the development of networks and on-line information systems which allow for multiple input and immediate retrieval of information; 4) the development of thesauri and controlled vocabularies for social science terms; and, 5) the development towards future considerations of a national union list of available mrdf and their location.” dodd, s. a. (1977, may). report of the joint canadian-united states action groups on classification (c ag). iassist newsletter, 1(3), pp. 8-10. retrieved january 2014 from <http://archive.org/stream/ newsletterserial13inte/newsletterserial13inte_djvu.txt> commentary: “the primary focus of the cag at the canadian working conference centered on the first problem of how to cite properly a social science numerical data file in the published literature… agenda topics : current problems and associated tasks within the mandate of the classification action group include developing (1) examples and guidelines for bibliographic references for social science numerical data files, (2) a “cataloging-in-production” scheme for major producers of social science data files, (3) a more “universally based” classification scheme for social science data files; and, reviewing (4) the cataloging efforts to date and discussing any or all related problems and (5) existing printed thesauri in the social sciences in terms of their future applications to data files.” dodd, s. (1977, fall). the emerging priority in bringing bibliographic control to social science machine-readable data files. iassist newsletter, 1(4), pp. 11-18. retrieved january 2014 from <http:// archive.org/stream/newsletterserial14inte/newsletterserial14inte_ djvu.txt> commentary: “two recent attempts to compile “catalogs” of social science data have encountered the lack of consistency among titles for the same data set. one attempt has been the recent cataloging efforts at the universities of north carolina, wisconsin, princeton, and yale, whereby traditional library cataloging records are created for social science data generated by academic research. the other has been the efforts of the association of public data users (apdu) to compile a directory of publicly available data files which represent primarily government produced data. both groups have experienced the same problem: variance of titles for the same data file. yet, without some control over titles and some mutually agreed upon primary source of title information, there can be no bibliographic control of social science data and none of the related products such as a union list of machine-readable data files. this paper will attempt to offer some suggestions for remedying the situation, including guidelines for transcribing titles; for creating a “title page”; for compiling a bibliographic reference; and for establishing an “authority list” for titles.” dodd, s. (1978, spring). building a bibliographic/marc data base for social science data files in a network environment. iassist newsletter, 2(2), pp. 34-37. retrieved january 2014 from <http:// archive.org/stream/newsletterserial22inte/newsletterseria22inte_ djvu.txt> commentary: “social science numerical and textual data files represent a vast amount of valuable and publicly available information. for example, they are widely used by students, faculty and policy makers engaged in research. not only have such data files had an unprecedented growth in the last decade, but with the advance of small and relatively inexpensive computer terminals, data analysis and computer simulation models have moved into the classroom as legitimate instructional tools. specialized files, often referred to as “educational data packages” have been developed to teach students analytical skills, so as to better understand social and economic phenomena. according to nesvold (1976): “experience with machine-readable ‘laboratory’ materials should be as appropriate to the beginning social science student as is the laboratory for the beginning chemistry student.” unfortunately, many such data resources are not fully utilized because potential users are unaware of the existence and accessibility of social science data files. at the present time, information on usable mdrf is fragmented among varying government agencies, research institutions, and university computing and data centers. among these various agencies, there is no common format for information on the existence of data files, nor is there any standardized structure that would facilitate retrieval of information from many different sources. existing information on computerized files is available to some but not to all. what is needed is a central source of information within the public domain that would provide equal access to all interested users. what is needed is some sort of bibliographic control and national standards for social science files-not unlike that which is available for printed materials.” dodd, s. a. (1978). characteristics of machine-readable data files. report on the conference on cataloging and information services for machine-readable data files: march 28-31, 1978 (pp. 5-8). airlie house, warrenton, va: data use and access laboratories. retrieved january 2014 from <http://books.google.com/books/about/report_on_the_ conference_on_cataloging_a.html?id=bf7gaaaamaaj> commentary: “to coordinate the development and implementation of standards for controlling mrdf, a conference on cataloging and information services was organized by dualabs … and supported by a grant from the national science foundation…the conference brought together at least 55 key persons having an active interest in establishing a framework within which a national program of cataloging and information services could be developed. the specific objectives of the conference were to identify key technical issues requiring resolution prior to implementing a coordinated cataloging effort, define the operational components of a centralized clearinghouse for mrdf cataloging…and … initiate a national program to catalog machine-readable data. sue dodd’s iassist quarterly 2013 11 iassist quarterly presentation was focused on “the act of cataloging mrdf and the related information products.” dodd, s.a. (1978, spring). mrdf cataloging conference. iassist newsletter, 2(2), pp. 49-52. commentary: reports on the national conference on cataloging and information services for machine-readable data files, held on march 29-31, 1978. the conference was an attempt at “establishing a framework within which a national program of cataloging and information services could be developed.” the organizers might be forgiven for feeling these initial steps have yet to bear fruit as far as a national infrastructure. even so, the conference resulted in a call to action for work to be carried out with major data producers in academic institutions and government agencies to use rules and guidelines for building a bibliographic record of data files created. dodd’s preliminary report was followed by a final report on the conference on cataloging and information services for machinereadable data files, march 28-31, 1978 at airlie house, warrenton, virginia. produced by the mrdf conference secretariat, dualabs. arlington, va. 1978. dodd, s.a. (1979, march). bibliographic references for numeric social science data files: suggested guidelines. journal of the american society for information science, 30(2), 77-82. doi:10.1002/ asi.4630300203 commentary: “in the last two decades, private research organizations, government agencies, and foundations have invested heavily in the collection of social science numeric data, contributing to the proliferation of machine-readable data. however, the development of information technology and the ability to produce data have progressed much more rapidly than our capacity to organize, classify, and reference its availability. there is an immediate need for some type of bibliographic control over mrdf, including guidelines on how to create a proper bibliographic reference. the purpose of this article is twofold: (1) to outline some of the information components associated with social science numeric data files, and (2) to provide guidelines, examples, and a uniform vocabulary for the creation of a bibliographic reference.” includes numerous illustrations and examples. dodd hoped that these guidelines “will soon appear in the “authors’ guide” section of social science journals and will eventually be included in such works as the chicago a manual of style and kate l. turabian’s a manual for writers. the ultimate goal would be to pave the way for social science data files to be included in printed bibliographies, end-of-work references, and indexing and abstracting works such as the social science citation index.” dodd, s. a. (1979, march). building an on-line bibliographic/marc resource data base for machine readable data files. journal of library automation, 12(1), 6-21. retrieved january 2014, from <http://eric. ed.gov/?id=ej203534> commentary: “explains how a multipurpose bibliographic/ marc data base of machine-readable data files (mrdf), created according to the marc ii record format, was conceived, and how much an information resource would benefit the general user and professional librarian.” dodd describes the experiences of the social science data library at the institute for research in social science at the university of north carolina, chapel hill to build an online catalog of data files. the article provides an example of early efforts to use the facilities and expertise of academic main frame computing centers, marc format for data files, and simple database design. dodd, s. a. & rowe, j. s. (1980) a model bibliographic information system for machine-readable data files in the humanities and social sciences. in data bases in the humanities and social sciences ed. by j. raben and g. marks. amsterdam/new york: north-holland, 1980, pp. 275-279 commentary: “models for the design of a database containing references to mrdf can be found in existing bibliographic information systems currently used for other types of research material books, journal articles, technical reports, government documents, maps and audio-visual materials. the paper summarizes some recent developments in the area of bibliographic control of mrdf, outlines a model integrated system of data elements and indicates some products and services that could be derived from such a system.” dodd, s. a. (1982). cataloging machine-readable data files: an interpretive manual. chicago: american library association. commentary (from the ala press release for this book): “one of the effects of the information explosion is the proliferation of machinereadable data files (mrdf). in order to assure better bibliographic control over this medium, the council on library resources awarded a grant to sue a. dodd to support her groundbreaking work on a manual for cataloging mrdf. the result of her efforts, cataloging machine-readable data files will be published. . . by the american library association. the first section of dodd’s manual is designed to demystify data files and computer programs and to make mrdf more comprehensible to those who must catalog and store them. section two explicates rules for cataloging mrdf from the ninth chapter of aacr2, offering interpretations and examples and revealing the “how and why” of cataloging. included is an extremely valuable outline of the steps for cataloging the microcomputer programs thaty are now found in many public and school libraries. the third section offers guidance on bringing bibliographic control to computerized files, including a bibliographic citation and a data abstract. the only available work of its kind, cataloging machinereadable data files provides the guidance which data producers, data archivists, and data librarians need to supply consistent bibliographic information for the mrdf they service. newcomers to this rapidly developing field will appreciate the accessible presentation and the glossary of terms included in the manual.” dodd, s. a. (1982, winter). toward integration of catalog records on social science machine-readable data files into existing bibliographic utilities: a commentary. library trends, 30(3), pp. 335-361. retrieved january 2014 from <http://eric.ed.gov/?id=ej286983> commentary: comments on “significant steps that have contributed to current level of bibliographic control of social science machinereadable data files (mrdf). . . and . . . outlines some of the remaining problems to be considered before mrdf can be integrated into existing bibliographic utilities.” twenty references and marc format, catalog entry, and data abstract for an mrdf are appended. the article summarizes and describes the process, groups and people who promoted the management and bibliographic control of data files, beginning in the late 1950’s. dodd includes details on the 12 iassist quarterly 2013 iassist quarterly role iassist played; including testing of a cataloging manual by several archives and reporting on the experience. there is a section describing how dodd funded and wrote her first book cataloging machine-readable data files. development of a “multi-purpose automated cataloging system” at icpsr is outlined as is work carried out to further test cataloging rules for data files. writing in 1982, dodd states “there is no doubt that machine-readable data will play an even greater role in research…[m]ore and more data [will be] needed for government and private research.” “[a] new dimension to the information explosion is now apparent; and with it an increasing demand for access to more and better documented data files.” “communicating the availability of usable data is an insperarable part of research and an integral part of librarianship. in the near future, libraries will have no choice but to become more involved.” dodd, s.a. (1984). characteristics and sources of public opinion polls in the united states. in numeric databases, ed. by ching-chih chen and peter hernon. norwood, nj: ablex publishing co. pp. 153-170. commentary: dodd provides a detailed history and overview on public opinion polling, polling agencies, methods of data collection and problems in analyzing and interpreting poll data. dodd, s. a. (1985). changing aacr2 to accomodate the cataloging of microcomputer software. library resources and technical services, 29(1), 52-65. retrieved january 2014, from <http://eric. ed.gov/?id=ej312387> commentary: examines rules affected by revisions approved by american library association’s joint steering committee for revision of anglo-american cataloging rules (aacr2) and those covered in the new “guidelines for using aacr2 chapter 9 for cataloging microcomputer software.” changes still needed to provide adequate bibliographic control are suggested. thirteen references are included. dodd, s. a., & sandberg-fox, a. m. (1985). cataloging microcomputer files: a manual of interpretation for aacr2. chicago: american library association. commentary: see discussion by paden, 1986 (below) dodd, s. a. (1990). bibliographic references for computer files in the social sciences: a discussion paper. university of north carolina chapel hill. chapel hill: institute for research in social sciences. retrieved january 6, 2014, from ( <http://people.virginia.edu/~pm9k/info/ compref.html> ) commentary: “a recent discussion among the participants of the e-mail “informal list for official representatives of icpsr” centered around citing computer files in references, footnotes, and bibliographies; whether to cite a codebook or file (providing you have both); and a discussion on citing primary or secondary sources. with respect to the last two concerns, there appeared to be adequate response indicating that it is better to cite the file as opposed to the codebook, and that one generally cites primary data sources. however, the first concern required more information and the icpsr or meeting was targeted as the next opportunity for such a discussion. note: this paper was first presented at the icpsr or annual meeting in november 1989, but has been revised for the may-june 1990 iassist meeting in poughkeepsie, n.y. … the march 1981 issue of social forces was the first time that a major social science journal had provided instructions (in the “authors’ guide” section) on how to cite a machine-readable data file (mrdf) -currently referred to as a “computer file.”… it is not possible in this discussion paper to provide anything but brief examples, but more detailed instructions on the components of a bibliographic citation are provided in the jasis article (dodd, 1979) and in part three, chapter 9 of cataloging machine-readable data files (dodd, 1982).” gray, a. s. & dodd, s.a. (1984). the roles of libraries and information centers in providing access to numeric databases. in numeric databases, ed. by ching-chih chen and peter hernon. norwood, nj: ablex publishing co. pp. 247-262. commentary: although this chapter was published in 1984, much of what gray and dodd say is still appropriate to ongoing debate about how libraries can take on the work of maintaining and providing access to data. the authors recognized that this would involve additional financial investment by libraries and they promote “cooperative arrangements which reduce the cost.” “keeping in mind the goal of providing access to the data itself, the library should also solicit information from those facilities which might also provide services for accessing the data.” the authors suggest one option might be to share computing facilities and to provide statistical consulting. discussion on acquistions and collections, cataloging, and public service are well defined and could inform libraries still. the authors advise that libraries “should be prepared to evaluate new roles in providing access.” roistacher, r. c., dodd, s. a., robbin, a., & noble, b. (1980). a style manual for machine-readable data files and their documentation. washington, d.c.: u.s. dept. of justice, bureau of justice statistics. retrieved january 2014 <https://www.ncjrs.gov/pdffiles1/ digitization/62766ncjrs.pdf> commentary: for many years, this manual was the de facto bible in the u.s. one could use to properly and fully document data files. each of the guidelines is illustrated with a rationale and examples. the authors state that “bibliographic identity is provided by six kinds of information; … [that] which identifies … describes the content … classifies … in a set of descriptors or keywords … [and provides] information required to access, analyze, … [and] archive the mrdf.” there is a comprehensive section on the kind of information required to describe each variable in a dataset: wording of questions, variable names, variable labels, explanatory text, code values, category labels, frequency count, and universe definition. these elements were to form the basis of a data dictionary and/or users’ guide we think of as today’s codebook. a chapter includes a detailed checklist. this document played a significant role in standardizing the kind of information required to ensure usability of datasets and impacted data management practices of researchers and archivists alike. works citing or making use of sue dodd’s rsearch. fox, m. j. (1990). descriptive cataloging for archival materials. cataloging and classification quarterly, 11(3-4), pp. 17-34. doi:doi: 10.1300/j104v11n03_02 commentary: this paper describes the significant characteristics of archival materials and of archival methods of description and arrangement. key sections of archives, personal papers, and manuscripts are explicated with particular reference to the ways in which the archival approach to descriptive cataloging reflects the nature of contemporary archival records and practice iassist quarterly 2013 13 iassist quarterlyiassist quarterly while remaining compatible with the style and structure of bibliographically oriented cataloging. the relationship of catalog records to other forms of archival finding aids is explained. kinney, t., & jones, r. (1988). microcomputers, government information, and libraries. government publications review, 15, pp. 147-154. retrieved january 2014, from <http://dx.doi.org/10.1016/02779390(88)90040-4> or < http://www.sciencedirect.com/science/ article/pii/0277939088900404> commentary: an increasing proportion of government information is being disseminated in machine-readable and electronic form. an important subset of this information—information distributed on diskette or optical disk or available online—is accessible using microcomputer technology. this article examines the role that libraries can play in helping their users to locate and use this microcomputer-accessible government information, and the potential of such a role in helping libraries to fulfill their responsibility to provide broad access to government information. kranz, j. (1988). microcomputer software cataloging: the need for consistency. cataloging and classification quarterly, 9(1), pp. 83-96. doi:doi: 10.1300/j104v09n01_09 commentary: “bibliographic records for microcomputer software in the oclc online union catalog are evaluated primarily for the purpose of focusing catalogers’ attention on selected areas in need of more consistent treatment. the degree of cataloging inconsistency evident in these records is examined with respect to the application of rules and prescriptions embodied in aacr2 chapter 9, [and] the ala guidelines for cataloging microcomputer software... a secondary purpose of this quantitative/qualitative study is to provide a general assessment of the overall composition of microcomputer software cataloging …” nasatir, m. (1981). the cataloging and classification of machinereadable data files, part 1: a case for incorporating records of machine-readable data files into the public catalog. cataloging and classification quarterly, 1(1), pp. 23-41. doi:doi: 10.1300/ j104v01n01_03 commentary: “standard catalog entries … constitute primary records by which computer-readable data files should be controlled and accessed. it is appropriate that academic and research institutions would want to record and provide access to files of data in machinereadable form in the public catalog of their libraries where entries already appear for other media …it is high time that the feasibility and desirability of incorporating records of machine-readable data files… in the public catalog … should be explored.” nasatir, m. (1982). the cataloging and classification of machinereadable data files, part ii: a practical application of the developing principles. cataloging and classification quarterly, 1(4), pp. 41-62. doi:doi: 10.1300/j104v01n04_04 commentary: “fortified with aacr1, aacr2 , the final report of the catalog code revision committee, the working manual for cataloging machine-readable data files and documentation from the data library at ucla’s institute for social science research (issr), a test run of cataloging mrdf was undertaken. this article chronicles the difficulties encountered during that practical application of the developing cataloguing principles.” nasatir, m. (1982). the cataloging and classification of machinereadable data files, part iii: subject description of machine-readable data files. cataloging and classification quarterly, 2(3-4). pp. 45-58 doi:doi: 10.1300/j104v02n03_03 commentary: “in 1976, the international association for social science information service and technology (iassist) classification action group participated in a project to test the feasibility of cataloguing and classifying mrdf. after applying the rules for descriptive cataloging presented in sue dodd’s working manual, based on the recommendations drawn up by the ala catalog code revision committee’s sub-committee on rules for cataloging mrdf… the project resulted in three recommendations. “the first is that [library of congress subject headings] lcsh be used for catalog … subject description of mrdf. … the second recommendation … is for people … to follow dodd’s guidelines and to provide as many descriptive terms as are applicable to the study… the third … is for a group representing substantive academic disciplines, government agencies, and catalogers to draw up useful terms at all levels of the hierarchy … to evaluate terms … to submit suggested changes to lcsh … to combine … terms into interdisciplinary thesauri … [and] to coordinate the consistent use of standard terms … “ nasatir, m. (1983). machine-readable data files: the future is here. journal of library administration, 4(1), pp. 1-5. doi:doi: 10.1300/ j111v04n01_01 commentary: this article “argues for the inclusion of a format for machine-readable data files in the existing bibliographic networks” given that a machine-readable cataloging (marc) format has been developed. the author notes that “large academic, research, and special libraries are requesting the capability of … [using] the us/ marc formats in combination with aacr2 [to] define the content of the data elements in a marc record. some of the elements … are a data file description which shows existence and source of data; a detailed abstract which includes the genesis and history of the file so as to link modified files; a keyword structure; physical characteristics of tape, file and software needed; applicability of the data to solving specific problems or analytic needs; and the link between data files and the software created to manage or operate them. these links reveal the presence of accompanying documentation, the bibliographic citation of accompanying documentation, software compatibility, and linkage with other files or programs.” further, nasatir stresses “[t]he more that users from all kinds of institutions contribute to and access a mrdf database, the more its economic viability will be assured … [and that] “communicating the accessibility of usable data [is] an integral part of librarianship.” paden, j. c. (1986). cataloging computer software: a guide for the inexperienced microcomputer user. cataloging and classification quarterly, 7(1). pp 19-33. doi:doi: 10.1300/j104v07n01_03 commentary: “an examination of aacr2 chapter 9 and the cc:da guidelines for using aacr2 for cataloging microcomputer software (chicago: ala, 1984) for catalogers not familiar with microcomputers. includes seven full descriptive cataloging examples of microcomputer software using these recently developed guidelines.” sue dodd and ann m. sandberg-fox’s cataloging microcomputer files: a manual of interpretation for aacr2 is widely quoted. paden states the guidelines for using aacr2’s chapter 9 for mrdf’s were “written in the mid ‘70’s before microprocessors and microcomputers 14 iassist quarterly 2013 iassist quarterlyiassist quarterly were fully developed and available on the scale they are today.” the article is a snapshot of the time and how changes in technology impacted libraries. these machines were so new, and they required a variety of software to be installed by the user in order to operate, unlike the early 21st c. when most machines come pre-equipped. libraries maintained copies of each version of programs that would then be installed as needed. there was confusion on exactly what to catalog, since one program usually contained several files to install and operate. “only a cataloger with considerable experience in computer software could be expected to determine the nature of each file in a program.” the section on how to describe physical characteristics is quite charming, with its directions on how to measure various storage media as well as their containers. wenzel, p. (1988). microcomputer-based access to machine-readable numeric databases. reference services review, 16(1-2), pp. 51-55. doi:(doi: 10.1108/eb049009 commentary: in a volume containing articles on the national archives in the u.s., national technical information service (ntis), national archives of canada, roper center, and an article by margaret hedstrom on state archives, wenzel describes a project, carried out at the data and program library service at the university of wisconsin madison to enhance access to their collection of machine-readable data files. project goals were based on the problems in access described by dodd and others. three levels of access are discussed. the archive level, study level and variable level each serve to identify the organization in which data are housed, provide broad descriptions of individual studies, and to document each variable within an individual dataset. this article documents early attempts by domain-specific data archives to develop and make visible online catalogs of their holdings. wittenborg, k. (1985). machine-readable data files. in patricia a. mcclung (ed.), selection of library materials in the humanities, social sciences and sciences (vol. 1, pp. 375-387). chicago: american library association. commentary: “in the past many libraries have been reluctant to acquire mrdf, as they presented a number of obstacles. to some they seemed prohibitively expensive; to others, they fell outside the library’s purview since they did not appear in bibliographies, abstracting and indexing services, or even databases. in addition, they require a computer and a degree of technical expertise to use. …machine-readable data files present a myriad of collection development challenges.” the author suggests readings by sue dodd and judith rowe, and participation in iassist as key to understanding the research and kind of data used by quantitative social scientists. in describing how libraries would manage data acquisitions, wittenborg addresses the use of bibliographic links between data and documentation with attention toward version control, and level of processing (curation). the article concludes with an admonition: “new information and new tools appear with astonishing frequency and the knowledge and tactics one has taken pains to acquire become outdated with alarming speed. …the selector must simply make a great effort to keep in touch…” references adams, m.o. (2006). “the origins and early years of iassist.” iassist quarterly, fall, pp. 5-14 <http://iassistdata.org/downloads/ iqvol303adams.pdf>. bisco, r. (1967). “the research library and data archives for social research.” speech prepared for the dedication of the graduate research library, university of florida. brock, c. (1967). “bibliography: current state and future trends.” political science. eds. robert b. downs and frances b. jenkins. urbana: university of illinois press. dodd s.a. (1978). “building a bibliographic/marc data base for social science data files in a network environment.” iassist newsletter 2:2, pp. 34-37. kramer, s. et. al. (2012). “using rdf to describe and link social science data to related resources on the web. working paper series no. 1. retrieved march 2014. <http://www.ddialliance.org/system/ files/usingrdftodescribeandlinksocialsciencedatatorelated resourcesontheweb.pdf> nasatir, d. (1973) “data archives for the social sciences: purposes, operations and problems.” reports and papers in the social sciences, no. 26. paris. unesco. national research council/committee on information in the behavioral sciences (1967). communication systems and resources in the behavioral sciences. publication 1575. washington, dc: national academy of sciences. chapter 3. tillett, b. (2004). “what is frbr? a conceptual model for the bibliographic universe.” originally published in technicalities 25:5 (sept/oct 2003). retrieved march 2014. <http://www.loc.gov/cds/ downloads/frbr.pdf> notes 1. libbie stephenson is director of social sciences data archives at university of california, los angeles. libbie@ucla.edu please contact ms. stephenson with any additions or corrections to this list of works 2. joint steering committee for development of rda. retrieved march 2014 <http://www.rda-jsc.org/rdaprospectus.html> 12 iassist quarterly 2014 iassist quarterly iassist quarterlyiassist quarterly abstract data visualization has grown in significance and complexity as the quantity of data and the technology supporting it have developed. understanding and using data visualization is now a core skill that should be incorporated into information literacy goals by librarians and educators. competency in data visualization is also closely related to data literacy and other quantitative literacies. undergraduate students and other general learners should be exposed to the fundamentals of data visualization early in their education. this article proposes that evaluation, critique, and use of data visualization be the initial focus of education, and discusses some starting points for training in these three areas. keywords:data visualization, information literacy, data literacy. introduction data visualization has existed in certain forms for centuries. william playfair (1786) is credited as the first to systematically use tools such as bar charts and line graphs in print to illustrate his arguments. pioneers such as charles minard and charles marey (1878) continued to develop new techniques in the 19th century, with applications to social and natural sciences. advances in statistical analysis entrained refinements in the visual expression and exploration of data, exemplified in the more recent work of william cleveland (1985, 1993) and john tukey (1977). but the contemporary imagination has been captured by the ever more rapidly developing forms of visualization generated by sophisticated and powerful computing techniques applied to large volumes of data. today’s data visualization, as expressed in interactive web graphics such as the enormous variety presented at visual complexity (2014), represents one of the most beautiful and intriguing flowers of “big data”. this new attention to data visualization raises the question of whether it has a place in general education, and most particularly in information literacy. information literacy addresses the elements required for interpreting and understanding the informational content of the contemporary, technologically rich environment for communication, whether scholarly and otherwise. information literacy standards and guidelines have been incorporated into the academic aims of institutions, to be promulgated and assessed by librarians and accrediting bodies. by its relative longstanding and well-developed framework, information literacy offers an appealing model to emulate for data visualization. the association for college and research libraries’ information literacy competency standards for higher education are a leading example of such a framework (acrl, 2014). while information literacy began with a textual focus, the approach has been extended to encompass media literacy, numerical literacy, data literacy, and other literacies as they have emerged. since data visualization is now emergent, and represents a major tool for the communication of complex results from large and often heterogeneous data sources, it is natural to consider data visualization as another type of literacy. the more advanced reaches of data visualization encompass high-performance computing, advanced graphic design, sophisticated studies of the cognitive perception of visual imagery, and other expert research. for examples of the frontiers of research in this area, consider linsen (2012) for medical imaging, marchese and banissi (2013) for humanities applications, and huang (2014) for an overview of human-centered design in visualization. in fact, the field’s rapid development has been recognized as a challenge by educators (owen, 2013). this paper does not attempt to address or survey this vast range of material. rather, it focuses on the data visualization skills and literacies that should form the foundational elements of the knowledge of a generally educated person today. these core elements should retain utility and validity even as the field changes. just as someone trained in information literacy can evaluate and use textual information with greater sophistication than the untrained, going beyond a simple and unreflective understanding of data visualization will improve the communication and analytical skills of students. after all, data visualization is just another way of presenting, interpreting, and data visualization and information literacy by ryan womack1 http://www.visualcomplexity.com/vc/ http://www.ala.org/acrl/standards/informationliteracycompetency http://www.ala.org/acrl/standards/informationliteracycompetency iassist quarterly 2014 13 iassist quarterlyiassist quarterly using information. it is time to bring data visualization into the literacy training offered by librarians and educators. literacies: data, statistical, quantitative, and more while data visualization has not yet entered the literature of library and information science [lis] to a large degree, a body of work has developed on the various literacies appropriate to the numeric side of lis that the international association for social science information services and technology (iassist) represents. gray (2004, p. 24) emphasizes statistical literacy as an important value, and hints at the importance of critical analysis of data graphics, saying ‘we live in the information age with rapid distribution of news and content, where content is often overlooked in favor of images, and at a time when more and more statistics and data products are being made available to a larger and less data-literate audience’. she goes on to argue for an expanded role of libraries in training others in statistical literacy concepts. schield (2004) addresses the differences and overlaps between the concepts of data literacy, statistical literacy, and information literacy. he argues that both statistical literacy and data literacy need to be taught more widely. stephenson and caravello (2007) describe the challenges of implementing a classroom instruction program on data literacy. data information literacy [dil] is a shift of emphasis that has emerged out of the increased attention to research data, whether the data is big or not. while many of the concepts of data information literacy are not new, and are certainly well-known to the social science data community, what is new is the emphasis on the educational mission of academic institutions to train new scholars in this set of skills. wright et. al. (2012) identify 12 categories of training needs for data information literacy, one of which is data visualization. certainly one place for literacy associated with data visualization is under this organizing rubric. carlson et. al. (2013a) suggests that faculty do not necessarily feel they have all of the knowledge necessary to train their students in dil, and welcome assistance from others in the education effort. data information literacy is closely tied to the educational and outreach efforts surrounding research data management, and thus parallels the content of data management training. data management training courses have sprung up at universities around the world, such as the university of minnesota (jeffryes and johnston, 2013) or the university of edinburgh (rice and haywood, 2011). these courses are typically focused on graduate students in the disciplines, teaching them how to handle and present their research findings and data. they are targeted to students who are past their general phase of learning, and who will therefore have many specific ways of handling data visualization that are appropriate and customary to their disciplines. designing general data visualization content at this stage is therefore difficult. besides, the data training agenda is crowded, and there is little time to add training to meet new goals. qin and d’ignazio (2010) mention data visualization as only one of 20 topics in a science data training context. despite these difficulties, a brief reminder of the need for thoughtful and effective data visualization may be appropriate in the graduate context, as well as pointers to places to learn more information. particular disciplines may use data visualization extensively, but the specialized techniques of the discipline are best addressed by specialists, not by generalists like librarians from outside the discipline. in other contexts, terms such as quantitative literacy or numeric literacy are used to describe the skills needed, placing the focus on the mathematical aspects of understanding data. certainly quantitative reasoning is an important part of educational goals and may include making numerical sense out of graphs and charts. other literacies could be adduced, but the purpose of this article is not to provide precise definitions of the boundaries between these interrelated forms of literacy, or to introduce new terminology for data visualization information literacy. instead, data visualization should be thought of as a component that relates to many aspects of literacy as described. skills in data visualization support a range of literacies and should be viewed as complementary to them. data visualization as a basic component of information literacy if data visualization is not to be taught as separate or specialized content, how can it best be integrated into general education goals? while the lis literature has recognized that data visualization has potential significance (thomas, 2012) and is a topic whose time has come (bell, 2010), specific goals for data visualization have not been articulated. a few skills relevant to data visualization are mentioned by interviewees in the data information literacy project, but not in a general education context (carlson, 2013b). this article represents a further step towards defining data visualization’s place in general education on information literacy. although the acrl information literacy competency standards for higher education (acrl, 2014) are undergoing revision, the basic competencies are currently as follows: 1 the information literate student determines the nature and extent of the information needed. 2 the information literate student accesses needed information effectively and efficiently. 3 the information literate student evaluates information and its sources critically and incorporates selected information into his or her knowledge base and value system. 4 the information literate student, individually or as a member of a group, uses information effectively to accomplish a specific purpose. 5 the information literate student understands many of the economic, legal, and social issues surrounding the use of information and accesses and uses information ethically and legally. the first competency can be considered a preliminary stage that defines the research project. this is not to minimize its importance. phetteplace (2012, p. 97) states, ‘it bears repeating: the first step to good data visualization is good data. most of the thought and effort should go into to [sic] collecting and analyzing data; playing with visuals until you find a compelling option is the reward for your due diligence.’ however, most of the initial research definition, data preparation, and other steps do not directly relate to data visualization itself. the second competency dealing with access does have a visualization component, but one that is arguably not an analytical one. because most visualizations will be encountered in the course of general research into a topic, via websites and publications, the information seeker may not need to learn new skills just to tap into data visualizations. while there are many applications of data visualization, such as the mapping of census data, where graphical elements are prominent, and there are some sites that specialize 14 iassist quarterly 2014 iassist quarterly in creating visualizations of preexisting data, effective and efficient access to data visualization is not a universal or high-priority need. initial steps in general education for data visualization should lie elsewhere. however, that does not preclude librarians and educators from building data visualization access tools into their repertoire of resources and guides. the last of the five competencies, relating to ethical and legal considerations, is a general proviso that applies to all information use. data visualization may engender some unique ethical and legal considerations, such as the safeguarding of individually identifiable information in a graph of social network relationships, or whether derived data distilled into images for distribution abides by terms of use for the data. however, these are more likely to arise in specialized contexts, and are not appropriate for introductory educational efforts. the competencies relating to evaluation and use (3 and 4) are the most relevant to data visualization in practice, and the most appropriate for incorporation of data visualization goals into introductory outreach. the acrl standards focus on the intellectual framework required to achieve the competencies, not the use of specific technologies or tools. for data visualization, the focus on the intellectual framework should remain the same, because the tools will change rapidly. the acrl has also developed visual literacy competency standards for higher education (acrl, 2011), but visual literacy as defined in these standards focuses more on the interpretation and use of visual imagery and media outside of the data context. these standards tangentially refer to the context of data visualization in standard four, part 1.f, as follows: “determines the accuracy and reliability of graphical representations of data (e.g., charts, graphs, data models)”. elsewhere, the standards remain focused on images in general. still, there are parallels with the broader acrl information literacy standards and the data visualization concepts discussed in this paper. standard three specifies the ability to “interpret and analyze” visual imagery, which is related to the concept of critique discussed below. standard four of the visual literacy standards deals with evaluation, and standard five deals with use. see hattwig et. al. (2013) for further discussion of the visual literacy standards. evaluation, critique, and use the third acrl information literacy standard requires students to evaluate sources critically, and to incorporate them into their knowledge base. there is enough work in both the evaluation and critique of data visualization resources that considering these elements separately is justified. evaluation, as used here, refers to the basic questions that must be asked of a particular data visualization to establish its quality, accuracy, and reliability. the danger inherent in a visual medium is that the power of the image will overwhelm the substantive content that it represents. when presented with a data visualization, the user should ‘interrogate the image’ and establish the source of the data, the reliability of the source, and the appropriateness of the visualization for the kind of data. if the underlying data is of poor quality, no amount of elegant graphics can compensate for this. understanding the methodology that produced the data is also essential (gray, 2004). concepts such as edward r. tufte’s lie factor (tufte, 2001) can be introduced to provide a framework for systematically checking the level of distortion inherent in an image. the lie factor is computed by dividing the size of the effect shown in the graphic by the actual size of the effect in the data. for example, an increase in magnitude from 11 to 12 can be made to appear as a doubling, if we set the baseline at 10 (+1 vs. +2 over the baseline). students should be introduced to a basic range of visualization types (bar, line, scatterplots, box and whiskers plots, etc.) and learn appropriate uses for each. for example, connecting points into lines to show a time series trend is a good idea, while connecting points in a scatterplot usually has no meaning and can be misleading. a box-and-whisker plot can summarize the variation of a dense dataset. the r package ggplot2 includes a sample dataset of 50,000 diamond prices with related characteristics. for example, in figure 2, the box-and-whisker plot of diamond prices classified by the cut of the diamond shows a considerable number of highpriced outliers beyond the “whiskers”, but a relatively compact central range of prices covering the 25th to 75th percentile of the data within the boxes. iassist quarterly 2014 15 iassist quarterly students should be aware that there are many other methods available to them, recognizing that time constraints may limit what is presented in an introduction to the topic. one resource describing such methods is the periodic table of visualization methods (2014) which displays an extensive and suggestive classification of available techniques, even if the classification is not quite as rigorous as the periodic table of elements. students should learn to evaluate data visualizations that they plan to incorporate into their research, just as they weigh and evaluate textual sources to cite. evaluation answers the fundamental question of whether or not a particular data visualization is sound and reliable to use as a basis for scholarship. critique, in the sense proposed here, is evaluation raised to the next level, and attempts to answer the question of whether or not a particular data visualization is among the best possible in its domain for a particular application. fox and hendler (2011) argue, among other things, that as web technologies have improved the ease of implementing more complex and interactive data visualizations, science should make greater use of these techniques for the masses to explore and interpret data. as research uses more sophisticated visualization techniques, students will need to understand and appreciate these nuances. critique involves comparison among different data visualizations in order to develop understanding of which visualizations exemplify best practices. general principles such as striving for clarity, avoiding clutter, and emphasizing the most relevant data apply to most visualizations. in addition, the best visualizations enable rich understanding of complex datasets with relative ease. the techniques developed to produce these visualizations are both an art and a science, and should be appreciated and emulated by students, who should also learn to be cautious of oversimplification and approaches that sacrifice features of the data in favor of graphical elegance. for example, fisher, dempsey and marousky (1997) show that despite more complex 3d forms being appealing to the eye, simpler 2d graphics were preferred when the task required actual extraction of information from the graphs. new methods, such as tableplots, are also being developed to simplify the visualization of large datasets (tenneke, de jonge and daas, 2013). figure 3 is a tableplot based on the ggplot2 diamond price dataset which provides a snapshot of the relationship among variables in this large dataset. data visualization should remain a tool in service of scientific knowledge, not an end in itself, in spite of the undeniable aesthetic attractiveness of many of the best visualizations. a good discussion of the difference between pretty infographics and high quality and accurate representations of data is beauty is as beauty does (lyons, 2011). senay and ignatius (1999) provide a useful schema of data visualization types and their attributes, along with extensive guidelines for their use at rules and principles of scientific data visualization. their rules are distilled from the works of several pioneers in the field such as cleveland, tukey, and tufte. senay and ignatius emphasize the need for further research, saying ‘to gain better understanding of the effectiveness and expressiveness of various visualization primitives, it is essential that empirical studies of visualization techniques should be undertaken. otherwise, the techniques that are generated may be viewed as meaningless ‘pretty pictures’. ‘ lessons distilled from the work of researchers in the field should be used to transmit some of these concepts to students at the undergraduate level and higher, using relevant examples. use is the third proposed area of focus. use puts the emphasis on putting data visualization into practice. with the menu of visualization types and rules for their use previously introduced, students should have the opportunity to practice doing their own data visualizations. guiding students through introductory examples, working in sandbox environments, and using various demos and examples will lead students through the process of actually developing their own visualizations based on the choices before them. more experience with actual creation of data visualizations will develop skill and wisdom in making good selections, and will reinforce the concepts learned about evaluation and critique. the actual form that the ‘use’ component takes will be determined by current technology and research needs in a particular setting. owen et. al. (2013) classify data visualizations into three areas: 1) scientific or data visualization, in which the data dimensions correspond to physical reality (e.g., remote sensing); 2) information visualization, for multi-dimensional data from a defined field of interest; and 3) visual analytics, which is massive and heterogeneous. these boundaries are fluid and may overlap, but this is one potentially useful schema for types of 16 iassist quarterly 2014 iassist quarterly visualization. each of these types will have its own software and design decisions. owen et. al. go on to address several areas in the creation of visualizations: the user, the design stage, visual presentation, interaction techniques (required for visual analytics), communication, collaboration, evaluation, and displays. not all of these categories are concerned with the analytical, intellectual literacy skills that relate to information literacy, but again this can serve as a template for developing practical examples that allow students to create and use visualizations. as a concluding example from the literature, kelleher and wagener (2011) offer 10 simple guidelines that apply to almost any kind of data visualization. these guidelines are distilled from the theoretical literature in a reliable way, and are very useful for basic training in data visualization. they range from [#1] ‘create the simplest graph that conveys the information you want to convey’ to [#10] ‘select an appropriate color scheme based on the type of data’, and are illustrated with examples of effective and poor practice. could this serve as an elements of style (strunk and white, 1979) for graphics? perhaps. while no one would argue that strunk and white are forever definitive and prescriptive, few would argue against the benefits of attempting to identify and reinforce specific fundamental and useful principles. more importantly, practicing such rules reinforces literacy. such preferred practices have evolved for static graphics, but recommendations for the new world of interactive and dynamic graphics have not yet been distilled into concise and widely accepted principles. as education and literacy efforts for data visualization grow, the body of knowledge describing these best practices will grow in parallel. conclusion as argued here, evaluating, critiquing, and using data visualizations have become an essential literacy, one that is now required to understand and make use of the information products of our datadriven age. by focusing on a limited set of the most fundamental principles of evaluation, critique, and use, data visualization can be incorporated into introductory information literacy efforts targeted at undergraduates and general learners. librarians and other educators should be equipped to instruct in these areas. data librarians and other data professionals are clearly positioned to lead these efforts. while ‘the devil is [still] in the details’ (womack, 2014), and the implementation of actual instruction programs will depend greatly on the preferred technologies and topics appropriate to each institutional environment, helping students make better use of information as presented via data visualization supports the core goals of information literacy. other quantitative, data, and numerical literacies will also benefit from a dose of data visualization. most importantly, students exposed to more sophisticated data visualization training will be better able to understand, not only data visualizations, but the world around them, and to develop the skills to influence their world. references acrl (2011) visual literacy competency standards for higher education. [online] available from: http://www.ala.org/acrl/ standards/visualliteracy. [accessed: 21 september 2014]. acrl (2014) information literacy competency standards for higher education. [online] available from: http://www.ala.org/ acrl/standards/informationliteracycompetency. [accessed: 22 june 2014]. bell, m. (2010) do you see what i see? multimedia and internet @ schools. march/april 2010, pp. 39-41. carlson, j. et. al. (2013a) developing an approach for data management education: a report from the data information literacy project. international journal of digital curation. 8(1), pp. 204–217. http://dx.doi.org/10.2218/ijdc.v8i1.254 carlson, j. et. al. (2013b) [online] findings from the dil interviews: data visualization and representation. available from: http://docs.lib. purdue.edu/dilsymposium/2013/dilcompetency/12/. [accessed: 22 june 2014]. cleveland, w. (1985) the elements of graphing data. wadsworth. cleveland, w. (1993) visualizing data. at&t bell laboratories. “findings from the dil interviews: data visualization and representation” http://docs.lib.purdue.edu/dilsymposium/2013/dilcompetency/ fisher, s., dempsey, j. and marousky r. (1997) data visualization: preference and use of two-dimensional and three-dimensional graphs. social science computer review. 15, p. 256. <http://dx.doi.org/doi:10.1177/089443939701500303> fox, p. and hendler j. (2011) changing the equation on scientific visualization. science. 331, 11 february 2011, pp. 705-708. gray, a. (2004) data and statistical literacy for librarians. iassist quarterly. 28 (2/3), pp. 24-9. hattwig, d. et. al. (2013) visual literacy standards in higher education: new opportunities for libraries and student learning. portal: libraries and the academy. 13(1), pp. 61-89. <http://dx.doi.org/ doi:10.1353/pla.2013.0008> huang, w. (ed.) (2014) handbook of human centric visualization. springer. jeffryes, j. and johnston, l. (2013) an e-learning approach to data information literacy education. 120th asee annual conference and exposition, paper #6956, june 23-26 2013. kelleher, c. and wagener, t. (2011) ten guidelines for effective data visualization in scientific publications. environmental modelling & software. <http://dx.doi.org/doi:10.1016/j.envsoft.2010.12.006> linsen, l. , et. al. (eds.) (2012) visualization in medicine and life sciences ii: progress and new challenges. springer. lyons, r. (2011) beauty is as beauty does. [online] available from: https://libperform.wordpress.com/2011/10/28/beauty-is-as-beautydoes/ [accessed: 24 june 2014]. marchese, f. and banissi e. (eds.) (2013) knowledge visualization currents: from text to art to culture, springer. http://www.ala.org/acrl/standards/visualliteracy http://www.ala.org/acrl/standards/visualliteracy http://www.ala.org/acrl/standards/informationliteracycompetency http://www.ala.org/acrl/standards/informationliteracycompetency http://dx.doi.org/10.2218/ijdc.v8i1.254 http://docs.lib.purdue.edu/dilsymposium/2013/dilcompetency/12 http://docs.lib.purdue.edu/dilsymposium/2013/dilcompetency/12 http://docs.lib.purdue.edu/dilsymposium/2013/dilcompetency http://dx.doi.org/doi https://libperform.wordpress.com/2011/10/28/beauty iassist quarterly 2014 17 iassist quarterly marey, e. (1878) la méthode graphique dans les sciences expérimentales et principalement en physiologie et en médecine. g. masson. owen, g. et. al. (2013) how visualization courses have changed over the past 10 years. ieee computer graphics and applications, july/ august 2013, pp. 14-19. periodic table of visualization methods (2014) [online] available from: http://www.visual-literacy.org/periodic_table/periodic_table.html . [accessed: 24 june 2014]. phetteplace, e. (2012) effectively visualizing library data. reference and user services quarterly. 52(2), pp. 93-97. playfair, w. (1786) commercial and political atlas: representing, by copper-plate charts, the progress of the commerce, revenues, expenditure, and debts of england, during the whole of the eighteenth century republished in the commercial and political atlas and statistical breviary, 2005, cambridge university press. qin, j. and d’ignazio, j. (2010) the central role of metadata in a science data literacy course. journal of library metadata. 10, pp. 188–204. <http://dx.doi.org/10.1080/19386389.2010.506379> rice, r. and haywood, j. (2011) research data management initiatives at university of edinburgh. the international journal of digital curation. 6 (2), pp. 232-244. schield, m. (2004) information literacy, statistical literacy and data literacy. iassist quarterly. 28 (2/3) pp. 6-11. senay, h. and ignatius, e. (1999). [online] rules and principles of scientific data visualization. available at https://www.siggraph.org/ education/materials/hypervis/percept/visrules.htm [accessed: 24 june 2014]. stephenson, e. and caravello, p. (2007) incorporating data literacy into undergraduate information literacy programs in the social sciences: a pilot project. reference services review. 35 (4), pp. 525-540. strunk, w. and white, e. (1979) the elements of style. third edition. macmillan. tennekes, m., de jonge, e. and daas, p. (2013) visualizing and inspecting large datasets with tableplots. journal of data science. 11, pp. 43-58. thomas, l. (2012) think visual. journal of web librarianship. 6, pp. 321–324. <http://dx.doi.org/doi:10.1080/19322909.2012.729388> tufte, e. (2001) the visual display of quantitative information. second edition. graphics press. tukey, j. (1977) exploratory data analysis. addison-wesley. visual complexity (2014) [online] available from: http://www. visualcomplexity.com/vc/. [accessed: 22 june 2014]. womack, r. (2014) data visualization and information literacy, iassist annual conference, toronto, canada, june 5, 2014. 2014. <http:// dx.doi.org/doi:10.7282/t37p8wm1> wright, s. et. al. (2012) a multi-institutional project to develop discipline-specific data literacy instruction for graduate students. libraries faculty and staff presentations. paper 10. http://docs.lib. purdue.edu/lib_fspres/10 notes 1. ryan womack is data librarian at rutgers, the state university of new jersey (new brunswick campus). correspondence may be addressed to rwomack@rutgers.edu or 169 college avenue, new brunswick, nj 08901 [usa]. http://www.visual-literacy.org/periodic_table/periodic_table.html http://dx.doi.org/10.1080/19386389.2010.506379 https://www.siggraph.org/education/materials/hypervis/percept/visrules.htm https://www.siggraph.org/education/materials/hypervis/percept/visrules.htm http://www.visualcomplexity.com/vc http://www.visualcomplexity.com/vc http://docs.lib.purdue.edu/lib_fspres/10 http://docs.lib.purdue.edu/lib_fspres/10 mailto:rwomack@rutgers.edu vol25s.1 iassist quarterly spring 2001 15 mission (multi-agent integration of shared statistical information over the [inter]net) -the data archive perspective by joanne lamb* abstract mission [1] is a multi-national project funded by the european commission. it aims to provide a modular system of software that will enable providers of official statistics to publish their data in a unified, and unifying, framework, and to allow consumers of statistics to access these data in an informed manner with minimum effort. the objective of the project is to develop an integrated set of software modules, which will • allow suppliers of statistics to subscribe to an integrated network of datastores via an interface to their existing data while retaining control over all aspects of access to their data: their level of involvement; the data they supply; the users who can access it; and the level of resources to commit. • allow users to make declarative requests, with a minimum of understanding of statistics, or the domain area, and still retrieve meaningful results from our internal routines or through an interface with external statistical packages. • give the user a range of options for automatic harmonisation of statistical data, with clear indication on the interpretation of the results. • provide audit trails of data manipulation and analysis, so that methods can be retained, re-used and published. • maintain libraries of metadata that can be made available to other users. • provide a flexible architecture that allows third parties to act as independent metadata providers, thus encouraging the free exchange of knowledge. • allow users to build up individual profiles, accessing data and methods most relevant to their needs. • offer a number of independent, interoperable systems that can run on different hardware platforms and access heterogeneous data storage systems. mission is a development of a fourth framework project, addsia [2], but it brings a number of new initiatives to the basic ideas of that project. these are: • the use of agent technology to optimise queries; • the use of the unified modelling language (uml) in designing the system; • the development of the concept of metadata libraries that are independent of data sources, and which provide a middleman service to the user; • tools to enable expert users to develop and share their methodology. the mission project mission is a european union r&d project, numberist1999-10655. the project started in january 2000 and is due to finish in december 2002. we are therefore halfway through the project. mission grew out of the addsia project and has the same partners, who are: university of edinburgh, scotland, uk (coordinator) office for national statistics, uk central statistics office, ireland tilastokeskus (statistics finland), finland university of athens, greece university of ulster, northern ireland, uk desan marktonderzoek bv, the netherlands the objectives of the project can be summarised as follows: we aim to build a software suite that will allow statistical data providers to integrate their publication of data on the web. this software will have a number of features, based on the requirements of data suppliers and data users. for suppliers, the system will allow them to subscribe to an integrated network of datastores via an interface to their existing data. they will be able to retain control over all aspects of access to their existing data. 16 iassist quarterly spring 2001 data users will be able to make requests in a declarative manner, with a minimum of understanding of statistics, or the domain area, and still retrieve meaningful results. the users will be able to tailor their environment, from simple requests to detailed in-depth analysis. they will also be able to build up individual profiles, accessing the data and methods most relevant to their needs. a key objective is to allow methods of data manipulation and analysis to be retained, re-used and published. this will be done using libraries of metadata and of tables, both of which can be developed in the mission system and then published for re-use by other users. we have in fact separated the functions of data providers and metadata providers. this approach gives mission a very flexible architecture, which will allow third parties to act as independent metadata providers. the mission system will be implemented on independent, interoperable systems running on different hardware platforms and accessing heterogeneous data storage systems. the core of the mission approach is the notion of allowing a user to form a query over several datasets, using a centralised statistical engine, which would parse the request and send partial queries to different datasets held in the system. this was the key concept of addsia, which has been further developed in mission. to achieve this we identify three basic concepts: the client, the library and the data server. the client is a piece of software that can be downloaded from a mission site and installed on the user’s machine. it connects to a host mission site, and offers both the user interface and the user’s workspace. the library is the core of the mission system. it is a repository for different types of statistical metadata: access, methodological and contextual. it also contains the statistical processing engine. libraries communicate with each other via agents, and therefore once a user is connected to a host library, he or she potentially has access to all mission libraries. the data server is the unit that gives access to the data. it provides the link from the data to the library. data servers register with libraries, and then register their datasets. the access rights components of the data server enables the data suppliers to specify who can access which parts of their data. only aggregated data is sent to the library. when a request is sent to the data server, the result of that query is computed, and this aggregated result is sent to the library. in this way, sensitive micro data is not sent outside the data supplier’s site, and also the amount of data transferred over the internet is reduced, thus making the retrieval efficient. figure 1 gives the overall picture of the mission system. innovation in mission mission has extended this basic idea in a variety of ways. first, we are using agent technology to enhance the system in a number of places. in the formulation of a query, the user is presented with logical variable names and descriptions, and agents will search a number of libraries to discover datasets containing these logical variables. the agents will then process the metadata for these datasets and can use a number of techniques to optimise the user’s request. these agents first construct the correct query, and then plan the execution of that query. further agents connect with the data figure 1: the mission system client library library library dataserver dataserver dataset dataserver datasets dataserver datasets iassist quarterly spring 2001 17 servers. this use of agents is illustrated in figure 2, where the first part of the diagram shows the query planning stage, and the second part shows the query execution phase. second, we have designed a graphical user interface, which will allow the user to specify and format his query using a table template to build up the query. this graphical specification is translated into the internal query language, and the results of the query populated the table frame that the user has built up. there will also be opportunities for the user to ‘post process’ this query, by graphically specifying the modifications, as shown in figure 3. a third aspect worth mentioning is the development of the mission model of metadata. figure 4 shows the mimamed data model. mimamed stands for micromacrometadata model, and is a model designed to library matching agent negotiation agent covering/costing agent query optimisation agent query planning agent brokering agent brokering agent brokering agent execution tree mameob + other metadata logs other libraries library matching agent negotiation agent covering/costing agent query optimisation agent query planning agent brokering agent brokering agent brokering agent execution tree mameob + other metadata logs other libraries other libraries figure 2: agents in mission var 2 var 1 male female cat1 cat2 total c1 c2 t c1 c2 t course 1 school national course 2 school national course 3 school national total school national merge courses insert variable insert splitting variable: e.g. gender figure 3: post processing a results table 18 iassist quarterly spring 2001 describe all three types of data, and to process the data and metadata simultaneously. in this model, developed by the university of ulster [4], the ‘sunkey table’ is a reference table for a set of particular aggregated data – the result of a query. the summary table is the actual data, all other tables contain metadata, which, when aggregates are combined, will also be processed. in addition, we have demonstrated that this model maps to other common models, such as data warehouse structures, cubes and the cristal [5] model. finally, the mission project is considering how the system will be made available at the end of the project. we have made an in principle commitment to open source [6] publishing of the source code, subject to this being a legal option in the setting of a european union r&d project. process to date the aspects of the system described above will feature in the first prototype, due in september 2001. in february, we successfully tested the connectivity of the system, with a client in ulster linking to a library in edinburgh, which queried three data servers in athens (running on different operating systems and relational databases). we have also completed the data server installation package. in may 2001, we repeated the experiment with a more complete query processing, and are about to complete the library installation package and supply the demonstration sites with these test versions. with the first prototype in september, the user will be able to: • browse metadata i.e. view dimensions of an aggregated dataset; • produce ‘simple’ tables from one source; • produce ‘composite’ tables from different logically homogeneous sources. future plans in the second phase of the project we plan to provide more functionality that will enable the expert user to share methods and publish results via the system. when looking at the requirements of the system for these features, we need to consider three different aspects: • the model for handling and displaying ‘transformations’; • storing mission tables for publishing; • more contextual metadata. for the first point, we will build a metadata model developed round the transformations that we have identified in figure 5. however, practically this will be stored as xml, for two reasons: • first, it is a principle that all metadata held in the client workspace will not require any software except a java environment – so the data cannot be held in a database. • second, we are developing a metadata user interface, which will allow the user to search any xml file. this metadata browsing tool will give us the ability to exploit xml files from different sources. we are particularly interested in the development of the ddi for aggregated data, and also in the standards table dtd. for importing data files, pc-axis, and spss will be the first formats that we support. we are considering which other formats are likely to be in demand, and expect to see a demand for ddi in future. figure 4: the mimamed model maptable reference table categorical attribute sunkey table categorical attribute numerical attribute numerical attribute summary table conversion table attribute level note data dictionary table level note data level note survey table attribute level note attribute label level note maptable reference table categorical attribute sunkey table categorical attribute numerical attribute numerical attribute summary table conversion table attribute level note data dictionary table level note data level note survey table attribute level note attribute label level note maptable reference table categorical attribute sunkey table categorical attribute numerical attribute numerical attribute summary table conversion table attribute level note data dictionary table level note data level note survey table attribute level note attribute label level note maptablemaptable reference tablereference table categorical attribute categorical attribute sunkey tablesunkey table categorical attribute categorical attribute numerical attribute numerical attribute numerical attribute numerical attribute summary tablesummary table conversion tableconversion table attribute level noteattribute level note data dictionarydata dictionary table level notetable level note data level notedata level note survey tablesurvey table attribute level noteattribute level note attribute label level note attribute label level note iassist quarterly spring 2001 19 the concentration for prototype 1 has been on the downloaded client software. users register to the library, and have privileges to access certain datasets according to the data suppliers’ stipulation. thus control of the use of data is left with the supplier. the degree to which data is confidential is also at the suppliers’ discretion, and we will advise caution in this area. while a single request may be easy to monitor, tracking a series of requests that may lead to disclosure is more difficult. in contrast, the public user – i.e. a user who accesses the system via the web, without registration – has no direct access to data. access to metadata is freely available, and also to pre-defined tables which are created through mission. therefore the amount of information accessible to the public depends on the number of tables published in a library. the second prototype is scheduled for march 2002, and the final version for december 2002. implications for data archives we have depicted the mission library as a separate entity from the data server, and we picture the two different modules being run by different organisations that have different expertise. the inspiration for the independent metadata provider scenario came from the way in which social science currently uses quantitative data. typically a researcher will get data from an archive, which will be well documented for secondary analysis. however, after a two or three year project, the findings are published in theoretical papers, and the modified data is destroyed for legal or economic reasons. the researcher may not legally be permitted to keep the modified data, since this would be for a purpose other than that of the original project. alternatively, there may be a financial cost for using the data for another reason. if the researcher cannot keep the modified data for his own purposes, it is even more difficult for another researcher to pick up on this work and continue. it has been observed that the manipulation of data prior to analysis encompasses the hypotheses of the research [7]. it is therefore important to capture not only the algorithm of the transformation, but also the reasoning behind it. if this can be presented to the analyst as a tool for aiding his own work, then the overhead of supplying this metadata is not seen as arduous. once this reasoning has been captured, it is available to give a reasoned method for other to use. the wider context this section places mission in the wider context of statistical information systems research in ces and in relation to other eu initiatives in which we are involved. the ces is participating in two r&d projects and three networking projects. while mission is concerned with data dissemination, the other project, iqml [8] is concerned with data collection. the three networking projects will be described briefly. metanet a network of excellence for harmonising and synthesising the development of statistical metadata started in november 2000, and held its first conference in april 2001. the proceedings from the conference are due at the end of may, and will be available from the website [9]. amrads accompanying measure to r&d in statistics – is a project concerning issues of technology transfer from r&d projects to national statistical institutes, and between national statistical institutes. while the focus of this project is on official statistics, the issues it addresses – the transfer of new ideas and technology between research and services institutes – also has relevance for data archives. the project has six themes led by experts in the area, and ces is responsible for metadata. the first significant event in this project will be the etk/ntts2001 conference in crete in june 2001, where amrads has had significant input to the programme. cosmos, a cluster of systems of metadata for official statistics, will start in september 2001. clusters are specifically for eu fifth framework projects, to encourage the sharing of knowledge during the lifetime of the projects. in cosmos, dohtem tnemmoc noitatupmoc tluserehttcurtsnocotnoisserpxecitemhtiranasesu noitacnurt arofyllausu-edocafodneehtmorfstigideromroenosevomer noitacifissalc desab-elur y=tlusernehtxxxfimrofehtfosnoitcurtsnifoseiresa gnidnab elgnisaotdespallocebdluohsseulavfoegnaratahtnoitacificepsa edoc elbatgnidocsnart stluserlanifdnalaitinifoelbata xelpmoc sdohtemfonoitanibmoca sdohtemfonoitacifissalc:5erugif 20 iassist quarterly spring 2001 lead by ces, we have five projects, whose acronyms are: mission (see above) iqml a software suite and extended mark-up language [xml] standard for intelligent questionnaires faster flexible access to statistics, tables and electronic resources metaware statistical metadata support for data warehouses ipis integration of public information systems and statistical details of the projects can be obtained from the information society technology website [10]. figure 6 shows the relationship between the projects. the two r&d projects are at the bottom of a hierarchy of generalisation. these feed into cosmos, where the objective is to demonstrate interoperability between some components of the five projects. they, and cosmos, feed into metanet, which is looking to build a conceptual framework of statistical metadata. finally, metanet and metadata form one strand in the general investigation of issues concerning technology transfer from r&d in official statistics. conclusions in conclusion, we would like to emphasise the following points. first, mission is an open system, which aims to help users exchange methodologies as well as utilise data from different sources. this gives the opportunity for data archives and data libraries to host metadata sites as well as access to data. we deliberately called these site libraries, since we feel that much of the metadata knowledge is held at the servicing level, rather than at the data provision level. the expertise of librarians and archivists is complementary to that of producers of (official) statistics. we expect to input the metadata from providers automatically, and are keen to utilise as many standards of metadata that it is feasible to handle. finally, mission is also contributing to other activities aimed at getting a shared understanding of statistical metadata needed for the processing, documentation and preservation of statistical data using modern technology. references 1. http://www.epros.ed.ac.uk/mission 2. http://www.ed.ac.uk/~addsia 3. http://www.epros.ed.ac.uk 4. http://www.epros.ed.ac/metanet/conferences/ proceedings.html 5.van bracht, e., de jonge, e. & kaper, e. cristal data objects. an object model for cubic, raw, or intermediate statistical data. statistics netherlands (march, 2000). 6.pardue, h (2000) open source software development: a business model. paper presented at the 31st annual meeting of the decision sciences institute, orlando florida november 1821, 2000. 7 fenelon j-p, grelet y, houzel y (1997) analysing transitions in the labour market through individual longitudinal data: some methodological issues. paper presented to the fourth transitions in youth workshop, dublin 1997. 8. http://www. epros.ed.ac.uk/ iqml 9. http://www.epros.ed.ac.uk/metanet 10. http://www.cordis.lu/ist/projects * paper presented at the iassist/ifdo conference 2001, amsterdam. joanne lamb, centre for educational sociology, university of edinburgh, scotland, uk mission iqml cosmos metanet amrads figure 6: relationship between projects http://www.epros.ed.ac.uk/mission http://www.ed.ac.uk/~addsia http://www.epros.ed.ac.uk http://www.epros.ed.ac/metanet/conferences/proceedings.html http://www.epros.ed.ac/metanet/conferences/proceedings.html http://www. epros.ed.ac.uk/iqml http://www. epros.ed.ac.uk/iqml http://www. epros.ed.ac.uk/iqml http://www.epros.ed.ac.uk/metanet http://www.cordis.lu/ist/projects who owns contract and grant data in the u.k. and who can use it? by cally brown school of library archive and information studies university college london ii marcia taylor i ssrc data archive i" university of essex in 1980, a working group of the social research association published a report on the "terms and conditions of social research funding in britain". (1) issues singled out for special consideration included control over publication and ownership of data and copyright. the ensuant discussion was not as useful as it might have been, however, oecause the differences inherent in contract-funded research and grant-funded research were not always recognised. it therefore seems worthwhile here to start by distinguishing between these two concepts insofar as they relate to the british situation. f in grant-funded research, money is awarded by the commissioning body on a broad understanding as to the results of the research. the study is usually initiated by the researcher rather than by the funding body, and the funder does not normally see itself as the primary user or beneficiary of the results. rather, the intended audience is seen to be other practitioners and theoreticians in the general field of the enquiry and, ultimately, 'the citizen'. ownership of results-including both data collected and interpretations of those data--are usually left in the hands of the researcher. in contract-funded research, money is awarded to the researcher for a specific study defined by the funder. the researcher may be pre-selected by a 'closed-tender' process, or chosen from a group invited to apply for the contract on a competitive 'open-tender' basis or, occasionally, appointed after public advertisement. a customer-contractor relationship is entered into where the commissioner purchases the researcher's services and is the prime user of the research results. the funder usually retains much clearer control over the research process than in the case of grant-awarding bodies, and usually claims rights of ownership over all material created during the activity of the enquiry. in effect, the contracted researcher is paid for his time and expertise but has no rights to the product of his labour. as the social research association report suggests, public funding of social research in the u.k. is organised in two ways: through government departments and quangos, and through more generally oriented independent bodies such as the social science research council (ssrc). increasingly, government departments are using contracts to administer research, whilst the ssrc usually allocates research funds through a grant system. the report also suggests that central government departments are increasing their direct support of social research whilst, at the same time, public funds made available for more general social research dre being diminished. with this picture in mind, we have first looked at the ways in which central government departments directly administer their research funds, paying special consideration to contract practice and how this affects ownership of, and access to, data; and then we have gone on to consider britain's most prolific independent funding agency of social research, the ssrc. the examples used as illustration have generally been based on personal communication. practice in central government departments r ,1 funding conditions the extent of control exercised by government departments over externally conducted research initially depends on whether the funds have been awarded on a contract basis or on a grant basis. practice varies considerable those departments that usually employ a grant system for funding research do not, typically, have a tradition of internal research activity. this may be due to a variety of factors, but in particular, may be due to the nature of the policy area involved. for example, the department of education and science and the department of health and social security are concerned with the school system and the national health service respectively. in both instances, the policy control is decentralised, and in both cases, the department favours a grant system as the most appropriate way of administering research funds. departments with a tradition of performing their own research and which have-or have had--internal research units, are most likely to work on a contract basis. the department of employment, the department of the environment and the home office, for example, all tend to put researchers under contract. access to research results as we have already suggested, contract-funding can give the funder greater control over the research he has sponsored. under common law in the u.k., an employer has the right to claim ownership of all materials created by an employee during the period of employment, so long as this is stipulated in the employment contract. the employee's right to publish no longer applies although under the copyright act ownership of copyright arises from evidence of authorship. for example, in british universities, ownership of copyright material authored by researchers is usually claimed by the university in its employment contract. whether the university subsequently exercises this power of ownership is another question. similarly, government departments can claim ownership of work performed by externally contracted researchers. 10 iii her majesty's stationary office, apparently, advises departments to include a specific clause in contracts making any written results of research subject to crown copyright, and the inclusion of a further clause claiming ownership of materials encoded in machine-readable forms. (ownership of machine-readable data is particularly unclear in english law as this area is not addressed by the 1956 copyright act currently in force.) the inclusion of such ownership clauses in government contracts is, however, left to the discretion of individual civil servants. this means that not only does contract practice vary between departments, but it can also differ within a department. in some contracts, the most restrictive veto on publication is both stated and implemented. in others, the department reserves the right to prior publication--indeed, in one case where such a right was reserved, the sponsoring department used this right to delay publication of results, which meant that the researcher concerned was unable to publish his somewhat contradictory interpretation . the publication veto is, however, usually more liberally interpreted by individual civil servants and although the formal contract appears to be restrictive, the researcher will often receive an accompanying letter of agreement relaxing any publication restrictions stipulated in the contract. it can also be noted that contracts from a few departments place no conditions upon publication apart from requiring acknowledgment of sponsorship and a waiver of departmental responsibility. the variation in contract implementation practice within departments is perhaps best illustrated by two views independently expressed to us concerning the same department. one researcher stated that the department was "the best--very liberal" in its attitudes, whilst the other was emphatic that the department was "a bugger-always gives me trouble". machine-readable data are preserved, if at all, on the initiative of individual government departments or by the social survey division of the office of population censuses and surveys. there is, at present, no public records office machinereadable archive where records collected in this form by government research must be deposited. although in march 1981 a review committee recommended to the lord chancellor that such an archive be established with the greatest possible speed (2), the future development of such an archive remains a matter for conjecture and discussion continue with the ssrc data archive as to the possible form this might take . meanwhile, public access to government-initiated data remains discretionary. typically, data are made available through individual arrangement between researchers and civil servants. more general access is frequently provided through deposit in the ssrc data archive and localised access is sometimes made possible by deposit in one of the smaller data archives established in several british universities. in the absence of legislation, access is often as haphazard as the interpretation of publication rights. some examples may serve to illustrate this. 11 in one type of departmental contract, the data are provided by the department to the contracted researcher. in one such case, the researcher was obliged to sign the official secrets act and was forbidden to allow anyone access to the data who was not specifically named in the contract. he was further required to return the data uncopied. in another case where the department supplied the data, however, the researcher was permitted to mount a copy on his local machine in perpetuity. another contract reported to us stipulated that the data be destroyed or returned to the departments after use but constructed variables could be retained by the researcher. an absolute lack of caution on the part of the department is illustrated by one somewhat incredible case where the researcher was supplied with a sample of highly confidential, individual records taken from a central register. nowhere in the contract was reference made to the preservation of confidentiality or to subsequent use of the data. in the other type of departmental research contract where data are collected by the contracted researcher, there are equally contrasting examples of departmental attitudes towards data access. in one case, a seemingly non-controversial enquiry in the area of medical research, the researcher, himself, was very keen to deposit resulting data in the ssrc data archive. the funding body, however, remained adamant that the data should be withheld, effectively ensuring that no further access could be made to the data by the contracted researcher or secondary analysts. in direct contrast to this, a well established research institute which conducts numerous social surveys of medical care under government contract, deposits data as a matter of course in the data archive. the lack of any legislation defining public rights of access to data may have been a contributory factor in the last example where data were scheduled for deposit in the data archive following a project initiated during the life of one government but, with a change in administration, this decision was reversed and no further access was granted. as we see it, where social research is administered directly by government departments on a contract basis, the government funds the research, specifies the research and has the power to control dissemination of the research results. there is, however, no consistency in practice and ownership and access conditions vary between contracts. current developments there is no automatic right of access by u.k. citizens to public records. access is not governed by any written rule of law but is, as we said earlier, at the discretion of the government and, de facto , of individual civil servants. there is, however, an indication that this situation may be changing and that the government is becoming aware of the need for some sort of coherent policy on public access to government data and--by extension--to government contracted research data. 12 ' in 1980, a review of the government statistical services was carried out under the chairmanship of sir derek rayner. it recommended, among other things, that government departments should seek less costly and more flexible means of enabling interested members of the public to have access to government figures, and that clear rules about the use of data should be published in order to enable more statistical research to be performed outside the civil service. (3) where this recommendation is acted upon, it may well help to increase the public availability of data collected under government contract. an indication of this already happening can be seen by the number of government department approaches made to the ssrc data archive in order to use the archive's facilities for disseminating data and statistical series, practice in the social science research council funding conditions the social science research council (ssrc) was established in 1965 to promote and fund research activities within the social sciences and to provide a continuing overview. the bulk of its research funds are spent on grants. the grant-awarding process is entirely reponsi ve--that is, application is made to the council at the initiative of the individual researcher. proposals are reviewed by a committee appointed for this purpose by the council, consisting of leading scholars in the field, and then put out for independent refereeing. this 'peer group assessment' is considered to be an essential part of the award process. awards are typically made to institutions rather than to individuals, so that applications are subject to further scrutiny by the research committess in applicants' home institutions. access to research results once a grant has been awarded, the researcher is usually left to his own devices to complete the research and submit his report. a copy of this report is usually deposited by the ssrc in the british library lending division. additional publications by the researcher are encouraged by the council, and copyright rests with the investigator. the ssrc aims to assure access to any machine-readable data generated during the study by making it a condition of the grant that a copy of the data be offered to the ssrc data archive for subsequent use by secondary analysts. failure to do so may affect the success of any future grant application made by the researcher. current developments in recent years, responding to both a growing shortage of funds and to public pressure to make research more relevant to policy issues, the ssrc has allocated an increasing proportion of its grant budget to specific research initiatives defined by a specially appointed board of the council. 13 administration of these research initiatives more closely approximates the contracting methods used by government departments, where the procedures are more formalised and supervision is likely to be more stringent. whilst directing research may be an efficient way for the ssrc to administer its restricted resources, this trend towards a research-initiatives policy has led to an emphasis being placed on short-term, ad-hoc and specific policy-oriented research at the expense of more long-term, basic research. in spite of the ssrc's moves towards initiating and directing research, there has been increasing criticism of the council for supporting esoteric and irrelevant studies, culminating in demands in parliament and the press for its closure. (4) however, the rothschild report published in may 1982, outlinging the results of a review of the ssrc's functions and functioning, recommended that the council should not only not be closed, but that it should be asked to return more diligently to its original remit of promoting the future development of social science research, particularly mul tidiscipl inary research that will not only advance the understanding of current issues of public importance, but will also fundamentally question the working of society. (5) the report stressed the importance of 'peer review' of social research, emphasising the need for independence from government departments in research initiation. if the recommendations in this report are implemented, the main thrust of ssrc funding can be expected to return to a grant-awarding system with its more liberal copyright and data access arrangements. it can reasonably be expected that independence from government control of data access will be assured by the council's continued committment to the broadest possible airing of research results. concl usion in the volatile and often contradictory situation we have described, it is difficult to envisage specific suggestions which could be made to guarantee the public availability of government-funded data. the social research association has begun drawing up a list of "desirable and undesirable contract conditions", and it recommends that "steps be taken to secure agreement to such a list". (1) whilst we feel that the first of these tasks is formidible and the second monumental, nevertheless we feel that such an exercise is a necessary pre-requisite to any further action. we also echo the association's view that many of the difficulties surrounding ownership of, and access to, data would be "alleviated if there were greater harmonisation of contract condi tions" . (1 ) we would add, in conclusion, that until data protection regulations and freedom of information legislation are in force in the u.k., little progress can be made towards any consistent or just policies to ensure access to data. 14 notes and references (1) social research association, terms and conditions of social research in britain . (london: sra, 1980) [report of a working group] (2) modern public records: selection and access . (london: hmso 1981) [report of a committee appointed by the lord chancellor. chairman sir duncan wilson, march 1981.] cmnd. 8204 (3) rayner, sir derek, review of government statistical services . (london: hmso 1981") [report to the prime minister, december 1980] (4) see, for example. the guardian 20 may, 1982. (5) an enquiry into the social science research council by lord rothschild . (london: hmso 1982) [presented to parliament in may 1982] cmnd. 8554 1982 annual conference report the lassist 1982 annual conference and workshops were held may 27-30, at the hotel del coronado in san diego, california, u.s.a. according to only moderately biased reports, the conference was considered a huge success by those attending. the site was fantastic, the program participants were well prepared and the hospitality suite closet bar was a different experience. conference attendance highlights include: ninety-five participants over the four day period. good workshop participation. fifty-three percent of those attending the conference attended a workshop. significant canadian, european and australian representation. over 22% of the participants were not from the united states. sporadic participation by the hotel police force in the post mid-night session of the turkey action group. the police comprised about 10% of the group. amazing representation at sunday's business meeting. thirty-seven percent or 35 people appeared. approximately 50% of the conference participants were not previously lassist members. twenty-four or 25% of those individuals subsequently became members. only 44% of the conference participants were from california. 15 beyond the public use file: confidentiality of archival records a case study. the national center for health statistics by cynthia g. fox national archives and records service governments and society need information to plan, execute and evaluate in a rational manner. access to timely and accurate information is the cornerstone of liberty. it i s no accident that orwell's hero in 198^ was an information manager or that henry ford's motto "history is bunk," and therefore subject to change was the watchword of huxley's brave new world . free societies have a need and a right to know. in the united states, for example, the social security act exemplifies the notion that information once considered completely private, that is work record and salary, should be collected about individuals to insure rational execution of social programs. citizens of free societies fund these efforts through taxation and participate in the collection of data by responding to census questionnaires, filing income tax returns, registering births, deaths and marriages, and applying for licenses. in most cases, they do so willingly and often without threat of criminal 1 iabi 1 i ty . this fundamental support for the collection of data, however, is tempered by fear, both real and perceived, of a loss of individual privacy and the creation of the all knowing "big brother." at the u.s. federal level, this fear is translated into legislation and regulation. the privacy act of 197^ requires a government to inform the citizens of the existence of systems of information which may contain data about them. it permits citizens to request the destruction of files which the government has created or collected on them or correct misinformation contained in those files. another piece of legislation, the freedom of information act, permits individuals to obtain access to much of the information the government collects. these two laws attempt to insure that the individual's rights to privacy and access to information are protected. the two laws work in concert. the freedom of information act, which protects the right of access, exempts from disclosure personal information the release of which would clearly constitute an invasion of personal privacy and specifies medical data. the result is that in the united states, information managers at the federal level are at the center of a triangle composed of 1) the need to collect data of a personal nature in order to assure rational planning; 2) the right of the citizens to access the information collected by the federal government; and 3) the need to insure the privacy of the individuals about whom the data is collected. if the confidentiality of the individual cannot be protected, then the ability of the government to collect the accurate information it requires will be impaired. similarly, the right of access to information cannot infringe on the right of the individual to maintain his personal privacy. the dilemma seems less of a problem when discussing machine-readable records. one collects that data needed to insure rational decision making, drops off the personal identifiers, edits or aggregates the data, and releases that new version, "a public use tape" to researchers. unfortunately, edited and suppressed records are not always the answer and "disclosure free" data tapes are not always totally 10 disclosure free. i will attempt to describe the efforts of one federal agency to prevent the unwarranted invasion of personal privacy and the efforts of the national archives to continue to protect the confidentiality of this type of record when transferred. the national center for health statistics, an arm of the u.s. public health service, states that its primary mission is "to develop statistical information on health matters in the united states and to provide that information as quickly and in as useful a form as possible to all who desire it."(l) this mission is conducted within the context of strict controls and guidelines aimed at insuring the confidentiality of individuals and entitites about whom they collect the information. the national center for health statistics or nchs functions under two basic propositions which are that the transfer of personally identifiable data from one custodian to another should occur only in accordance with carefully formulated and widely understood written rules and that there is a basic difference between information collected and used only as statistical evidence and personally identifiable data used directly to affect the rights, benefits, privileges, responsibilities, duties, or proscriptions of individuals. they collect data for statistical and reporting purposes only and they do so in a fashion controlled by regulations, written agreements and signed assurances. there are two laws which permit the center to provide the confidentiality it requires to carry on its work. section 308(d) of the public health service act (^2 u.s.c. 2^2m) provides the basic legal authority for the protection of nchs files. it states that no information obtained in the course of activities undertaken by nchs and its sister agency, the national center for health service research may be used for any purpose other than the purpose for which it was supplied unless authorized by the secretary of health and human services. the law states that information obtained in the course of health statistical activities may not be published or released in another form if the particular establishment or person supplying the information or described in it is identifiable unless this establishment or person has consented. whenever nchs requests information, it specifies to the person or agency supplying the information that the data will be used for a limited purpose or purposes. in most cases, this use is limited to statistical research and reporting. under a second law, the privacy act of is^t, nchs has obtained a "k-^" exemption for i ts stati sti cal systems. this means that nchs does not have to allow the subjects of its data files to have access to the records about themselves in those files. this exception to privacy act requirements is permitted because nchs does not have in its data files any records that are used in any direct way to affect the persons whose records exist in these files, (2) the files are used strictly for statistical and related purposes. the public health service act and the privacy act augment the basis for exemption from the freedom of information act {k u.s.c. 552). subsection (6) of the act specifically exempts personnel and medical files and similar files "the disclosure of which would constitute a clearly unwarranted invasion of personal privacy" and subsection (3) provides that matter "specifically exempted from disclosure by statute" are also excluded from the disclosure requirements. the public health service act provides the necessary statutory restriction to prevent the disclosure of individual information . the center makes every effort to assure that no breach in confidentiality occurs at any stage of the life of their files. in all cases when nchs contracts 11 with any organization outside the center for the collection or use of information that identifies individuals and/or establishments not advised that all information obtained from them will be made public, the contracts are carefully worded to assure compliance with either the privacy act of is?'^ or the public health service act. the contracts contain stipulations to assure confidentiality and physical security of the information and ensure that the contractors' employees abide by the center's stipulations. (3) nchs uses three standard wordings in data collection contracts, depending on the laws governing the particular project from which the information is to be collected. the first alternative is to be used when both the privacy act of is?'* and section 308(d) of the public health service act apply. the project could involve the collection of information about identified individuals which nchs is authorized to carry out. the second alternative is used when section 308(d) of the public health service act applies to the project but the privacy act does not. for example, in the case of a survey of institutions providing health services, the privacy act covers only individuals and does not apply. the third alternative would be used when the privacy act but not the public health service act covers the study. such a situation would be rare but might occur if congress authorized a study not normally conducted by nchs and required that nchs conduct the study. the fourth wording is used in contracts called for when contractors process confidential data. (4) nchs contracts wi th organ izat ions for the collection, processing, and analysis of data because such organizations are specifically equipped and staffed to perform such services effectively, although the contracting organization has no intrinsic rights in the data. these contractors provide such services as extensions of nchs and, as such, they are subject to all the confidentiality strictures which apply to nchs itself, and the contractors' employees are subject to the same privacy act actions as employees of nchs. the center's policy on protection of records also applies. (5) all contractor employees must sign a "nondisclosure statement" as do all nchs employees. in addition to signing the statement which specifies the force of law and punishment for violation, employees agree to maintain the same physical protections that the records are afforded in the center. this includes keeping them locked up in fireproof cabinets or locked rooms at all times when they are not actually being used, keeping them out of sight of persons not authorized to work with the records, limiting duplicates, and transferring them in sealed containers. furthermore, in all statistical programs involving confidential information, records containing identifiers or individuals or establishments should be held to the minimum number required to perform the center's function. identifiers are never carried beyond the original survey or report document when the data are processed, unless there is a legitimate and important reason for doing so. documents containing identifiers are stored in secure areas as soon as possible. nchs releases its statistical data in the form of traditional published tables. it also creates public use versions of its data files by deleting personal identifiers or suppressing selected elements to ensure that the identity of the individual is protected. the availability of machine-readable microdata files enhances the value of research conducted by nchs. it permits other public and private institutions to use the data for purposes other than the rational planning of health related services. the center has determined that the release of the data in machine-readable form is a desirable objective. (6) however, nchs must achieve this goal without compromising the confidentiality of its files. 12 the question of how to achieve both goals raises a wide range of ethical, legal, technical, technological and economic issues. five different classes of constraint must be considered in deciding if microdata can be released and if so how it should be handled. i have previously mentioned the public health service act which provides the legal constraint on release. as i mentioned, the phs act makes it clear that data collected by the center must be processed in such a manner that the identity of no individual is disclosed and that the individual identification will be used only by persons engaged in achieving the purpose for which the information was originally assembled. the identifiers are deleted as soon as possible in the processing sequence. policy and practices in nchs and other general purpose statistical agencies have given strict interpretation to this principle, not only protecting individual records against unauthorized use, but also in adopting tabulation and publication procedures that are designed to make it virtually impossible to isolate, identify, or extract facts in such a way that a specific individual or business establishment can be identified by use of released data. (7) in addition to these legal constraints, there is an ethical concern governing the release of microdata files. on basic human grounds, nchs has an obligation to protect individual respondents against any invasion of their personal privacy and to be absolutely sure that the confidentiality of privileged communications is not breached. in the united states this ethical responsibility is embodied in the previously discussed exemption of medical files from disclosure under the freedom of information act. however, nchs is a federal government agency and as such must make every effort to assure that the citizens of the country have maximum access to bodies of information that the government assembles. while nchs has received a k-k exemption from privacy act disclosure of individual files, the ethical concern embodied in the privacy and freedom of information acts still exists. a government in a free society is and ought to be responsible for informing its citizens about the scope and content of systems of information that it maintains. any release of microdata must be done under public scrutiny providing equal access and equal protection for all involved. the third consideration in the possible release of microdata is a technical constraint. because nchs programs are generally designed to produce estimates for the entire united states, some studies require elaborate scientific sampling procedures. release of the microdata without a complete explanation of the editing, weighting, and ratio adjustments that it has undergone would be close to releasing inaccurate data. the center must determine if the sampling and manipulating can be explained and duplicated by a researcher. technological considerations provide the fourth constraint on the release of data. hardware dependency is an example of this type of consideration. another example would be data in other than a statistical form. in some nchs programs, original data includes such records as xray films, electrocardiograms, paper tapes, tape recordings of speech samples, blood specimens, and photographs. the ability to provide reproductions of these types of information are bound not only by technological difficulties by financial constraints as well. this is the fifth consideration in the release of data, the economic problems. financial means have become very important in determining the nature of the center's statistical output. the center's resources in funds, personnel, and equipment are limited. the highest priority and prime purpose for which nchs surveys are conducted are the prompt production of general purpose statistical tabulations for use by a broad spectrum of consumers. the preparation of public use files takes second place to the production of these tabulations. (8) 13 these five constraints are reflected in the nchs policy statement on release of data for individual elementary units and special tabulations. it reads, "within prevailing ethical, legal, technical, technological and economic restrictions, it is the policy of the national center for health statistics to augment its programs of collection, analysis, and publication of statistical information with procedures for making available, at cost, transcripts of data for individual elementary units-persons or estab! i shments-in a form that will not in any way compromise the confidentiality guaranteed the respondent." micro-data tapes are released after they have been reviewed and approved by the director of the center as conforming to guidelines and conditions set forth in the policy statement. operational considerations and attention to the matter of unit costs mean that for most data the effective format for release for the individual elementary units is a standardized microdata tape transcript. a descriptive catalogue of public use tapes is published. the catalogue is supplemented by a newsletter of updates. each public use data set is governed by a procedure designed specifically for that particular study. in general, however, the data set is composed of a standardized transcript, a computer tape image of the edited, weighted, and adjusted data from which all evidence is deleted which might possibly identify the respondent, the image is arranged in a fixed format and is accompanied by documentation which explains the editing and weighting and permits the use of the tape. any codes which appear on the tape are scrambled or offered in the most general terms. data which has proved faulty is represented by blanks. the center will not modify standardized transcripts or produce special tapes when standardized transcripts have been prepared. finally, each purchaser must sign, as part of the order form, a statement of assurance regarding the use of the data: "the undersigned gives assurance to nchs that individual elementary unit data on the micro-data tapes being ordered will be used solely for statistical research or reporting purposes." the center's reference technical services are handled under contract with the national technical information service (ntis). ntis maintains and sells out of print nchs publications in addition to reference copies of the public use files. however, proprietary responsibility for the data is still in the hands of nchs. when the agency no longer requires the information to perform its function or when data is replaced or superceded the records become eligible for transfer to the national archives. the national archives and records service has been concerned with the protection of confidentiality since its establishment in ibs'*. long before "privacy" became a national issue, nars successfully protected personal, restricted, and classified material from unauthorized disclosure. nars' general restrictions placed a 75-year restriction on files which contained personal and medical information the release of which could be embarrassing to an individual. this principle has been enforced consistently unless prospective researchers assure nars that data is to be used for statistical or reporting purposes only. according to the u.s. code, records transferred to the national archives become the responsibility of the administrator of general services (the parent agency of the national archives). statutory restrictions on their use continue to apply for 30 years from creation or longer if the archivist and the head of the transferring federal agency so advise (9), and other restrictions may be negotiated between the transferring agency and the archivist. ]k micro-data files which are deposited into the national archives are treated in precisely the same fashion. since nchs records are subject to the restriction imposed under the phs act previously discussed, they are to be used for statistical and reporting purposes only, for a period of not less than 30 years. in addition, nars general restrictions, revised february i'*, i983, restrict for 75 years "records containing information about a living individual which reveal details of a highly personal nature. .. incl uding but not limited to information about the physical or mental health or medical or psychiatric care or treatment of the individual .. .not known to have been previously made public." (10) such records, the restrictions go on, may be disclosed to "researchers for the purpose of statistical or quantitative research when such researchers have provided the national archives with adequate written assurance that the record will be used solely as a statistical research or reporting record and that no individually identifiable information will be disclosed." (11) in addition to the guidelines and restrictions which pertain to all archival records, two office of management and budget publications offer guidelines for determining the confidentiality of machine-readable records. 0mb ' s computer security guidelines for implementing the privacy act of 197^ (firs #1 ) , prepared by the national bureau of standard, requires that provisions exist for stripping "records of individual identifiers so that identities cannot be discerned when statistical research or reporting records are disclosed or transferred 1(a) (10) and for ensuring "that an individual's identity cannot be discerned from tabulations or other presentations of statistical data by combining various statistical records or referring to other available information" 1 (a) (11). by requiring that the agency transferring records to the national archives supply a complete and accurate record layout as part of the documentation in hard copy form, the machine readable archives branch is able to determine with relative ease if a file contains information of a personal nature. these precise definitions of potentially restricted data elements along with detailed information relating to the specific assurances given to the respondents make the production of public use tapes in extract or summary form possible. 0mb ' s privacy act implementation guidelines and responsibilities is the second publication which may be used to determine confidentiality of data elements. according to the gui del i nes the elements of personal information may constitute an invasion of personal privacy if retrieved through personal identifiers. this covers a number of areas including: income or census information, private or subjective information about an individual or family, inaccuracies, information about individuals used by a federal agency for making a policy decision, information gathered under an implied or expressed promise of confidence, or any combination of public elements which combines to form too intimate or too detailed a profile. the national archives operates within these guidelines and under its own restrictions and staff guidance and procedures to identify data elements which are subject to protection from disclosure. in addition, the following steps are taken to ensure the physical and intellectual security of restricted files: a. all files are stored in an environmentally controlled vault; b. access to the vault is limited to nars reference and technical personnel and the washington national records center security officer; c. the file and its documentation are stored separately; d. no file is maintained in an office area except when in transit; e. documentation does not accompany files for computer processing; f. all programming is controlled; 15 f. agency restrictions are strictly enforced; h. all restrictions are published; and i. all requests for access to restricted data, even for statistical and reporting purposes, must be approved. in the case of nchs data files, the national archives will accession not only the full master files of unsuppressed microdata but the public use versions as well. the master files meet all of the criteria for restriction for at least 75 years unless an inter-agency agreement reduces that to 72 as has been done for census files. the public use versions created by the agency will be accessioned rather than created by nars because the center's record of protecting the privacy and confidentiality of their data has been outstanding. the national archives and records service wants to preserve the records of the federal government and provide access to any researcher as completely and conveniently as possible without violating the confidentiality of the information or the privacy of the source of that information, nars and nchs share these common goals and are working together to insure that a wealth of health related data is preserved. notes (1) u.s. department of health, education, and welfare, public health service, staff manual on confidentiality nchs , dhew publication no. 1 (phs) 78-12'4i4, (hyattsville, md: july, 197s) 10. (2) ibid ., 3. (3) ibid .. 13. c*) ibid ., 26-37. (5) ibid ., 9. (6) national center for health statistics, policy statement on release of data for individual elementary units and special tabulations , (washington, dc: 1978) 3(7) ibid., 5. (8) ibid. , 6-7. (9) '4'* u.s.c. 210'*. (10) k] c.f.r. part 105-61 . 5302. '(a) (11) ibid ., 3(a) . (12) u.s. department of health, education, and welfare, public health service, the model state health statistics act: a model state law for the collection sharing , and confidentiality of health statistics. (hyattsvi 1 le . md: march, i98o) t. 16 comparative charting of social change in four industrialized societies by simon langlois' institut qiubicois de recherche sur la culture et department de sociologie, universiti laval contemporary analysts of social change have somehow abandoned the ambitious tenet of building a general theory which could predict or explain all the contemporary forms of social change. a search of laws of history, or works on the stages of development (rostow), are no more pertinent r. boudon has proposeid an analysis of the reasons why such a loss of interest for a general theory on social change occurred. according to boudon, social systems are not regulated by general laws which are extensible to all societies. consequently, the search for universal laws must be abandoned and analysis of social change must give up the "monologic" perspective (boudon 1983:104). in its place, he suggests to draw some possible statements (in french, "dnonc6s de possibles"), i.e. relationships which have a certain probability, in a certain location and a certain time, instead of general laws of the form of conditional laws or transitional laws. from the boudon's individualistic approach, there is no more universal laws because sociologist's observed propositions are in part the results of aggregate actions of individuals. almost all general propositions (laws) in social sciences are in fact situated observations which are not universal, and it could be almost possible to pose counter-examples. otherwise, boudon suggests to consider the existing theories of social change as formal models which need to be adapted to every specific and concrete situations. ted caplos also criticized social change studies, more precisely empirical studies, saying that the parallel between social change and technological change is misleading. "instead of being continuous, like scientifictechnological progress, social change seems to be episodic, non-consultative, non-consistent and nonreversible" (caplow, 1988:3). too severe a diagnosis? in fact, these two briefly referred to comments, from different perspectives, reveal or show a growing uneasiness with the traditional ways of studying social change. a new approach is essential, which will make possible to observe and to consider the diversity of ongoing social change, and, more precisely, which facilitate the observation of non-cumulalive and non-consistent change. this can be done with social trend analysis. social trend analysis social trend analysis is sitting between social reporting and system analysis of the global society. social indications "designate statistics that are supposed to have significance for the quality of life, and sets of social indicators designate a social report" (mikalos, 1982:2). there is a normative perspective in constructing social indicators. they intend to measure goal attainment, or to evaluate and compare the respective performance of different countries. by themselves, social indicators do not form a social system. they are built to measure different aspects or dimensions of social life, or to evaluate current public policy and programmes. international comparisons of social indicators are aimed at comparing differential goal attainment from one country to another. countries are in fact located on a continuum. consider, for example, life expectancy in good health after 65. behind this social indicator, there is a clear or precise objective: staying in good health as long as possible. if this indicator increases, one expects to say that the quality of life is increasing; and the quauty of life is supposed to be higher in a society in which this indicator is also higher. "social indicators facilitate establishing social goals and pohcies. for example, a time series comparing life expectancy of several countries shows that life expectancy in the united states if lower than in the united kingdom, canada, japan and sweden. such information on levels attained by other countries shows us what is attainable and stimulated action" (ferris, 1988:609-610). opposed to social reporting, one finds global diagnosis which summarize, under a single heading or a macrotrend, a great lot of particular trends: post-modem society, dependent society, traditional or modem society, etc. a large number of segments are postulated to go in the same direction, and the macro-trend either summarizes these particular or specific trends, or is presented as the cause which otientate them in a certain direction: this approach raises many questions: to what extent all the trends are really convergent? for example, is france really a post-modem society? is it possible to speak of the process of europeanization in contemporary united states (oxford analytica)? assist quarterly trend analysis, in the perspective proposed here is somewhere between the two approaches stated before. a french research team, who have published an article under the pseudonym louis dim, define trends in this way: the trends which have been identified are sometimes in the nature of behaviour which can be isolated and taken as an indicator of a more general movement. for example, the decrease in religious practice or the increase in participation in sport or spwts-related activities. other trends are much more global, they formulate an overall judgement on a social sector or on an aspect of society. they refer to a sociological theory by which the phenomenon is delimited and can be judged. (louis dim, 1985:401) the analysis of trends if often associated with the study of phenomena which are quantifiable ot which have direct relevance to government policies (income, unemployment, population, etc.). some trends are better documented than others, mainly because they are based on data or indicators gathered by a government statistical office, indicators which have, in a way, received the lion's share of analysis. changes in income, voting, fecundity, unemployment, for example, are better laiown that those in division of domestic labour or in use of leisure time. however trends can be analysed other than quantitatively, it is also possible to discem the direction of phenomena with a certain amount of accuracy by using qualitative studies or monographs: strengthening of kinship, increases in new forms of religious practice, etc. it is this broad concept of trends such as these that we are speaking of. the trend is not an indicator nor a statistical series. it is a diagnosis on a social segment, narrowly defined (declining fertility, for example) or much more wide (increasing mobility of daily hfe). in the first case, the trend is in fact close to an indicator, but in the second one, it comes from the convergence of many indicators, of many statistical series. our unit of analysis is a trend , i.e. a series of values representing the incidence of some item of social behavior in a given population at points of time in a consecutive sequence. guidelines, annexed). the trend is a sector-based diagnosis of the changes in and the direction of a phenomenon or of an aspect of social reality, for example, a drop in the birth rate, an increase in disposable income per inhabitant, an increase in poverty, a decrease of the inequality between the sexes, an increase in the importance of kinship, a growth of individualism, a decrease in the practice of rehgion, etc. trend analyses are in fact studies of the present situation in light of the past it must not be confused with prediction of the future, nor with futurology. it is a way of studying changes "en cours". this perspective is important for another reason: the comparative purpose. theoretical framework and method the array of trends to be analysed is extremely thxjad and we did not want to use too narrow a theoretical framework. the choice of trends cannot be firmly fixed since it will be modified somewhat as the analysis goes along, as the french experience has shown. the choice is not based on definite theoretical bases, thus we are not limiting ourselves to the study of marginal trends which may reveal tomorrow's norms, for it is far from certain that this will be the case. the critique of reich's work, the well-known essay on the new culture published in the 60's, by hamilton and wright (1986) is revealing in this respect. it is no longer necessary to limit oneself to known territory, because analysis should also be able to discem new trends, especially outside the sjaere of work, although this remains important this explains why the choice of trends to be studied is based mainly on a group of hypotheses, not a single one, and remains c^n to change. the trends are in fact derived from sociologically driven categories. we have agreed with the other research teams to give priority to behaviour, ritualized situations in institutions and to structures, giving lesser importance to values and social perceptions. the priority given to recently identified trends does not mean that the realm of social perceptions will be completely absent social representation is less "objective", for lack of a better term, than behavior, at least a priori. people's jobs, salaries, levels of education, whether they have a reugious or civil marriage, their actual number of children, are thing which can be observed quite accurately, taking errors of measures into account. aspirations as to salary, the feeung of being deprived, the kind of marriage intended, the number of children desired, satisfaction in a couple's relationship, are all kinds of perceptions and attitudes. measures of these things are not only filled with errors, they are also more unstable. the number of children, salary or level of education are characteristics which are probably more stable than aspirations, for example. as to theory, it seems difficult to omit the domain of social imagery and to only take into account behaviour, or more factual data on individuals. let us not forget that ^tors also give a meaning to their conduct. measuring perceptions, while difficult, thus seems relevant and necessary to an in depth understanding of social phenomena and conduct the analysis of relations among trends is intended to be still more inductive. at this stage of the project we want to hypothesize as little as possible about the relationships. critical analysis of the literature shows that this inductive approach can be very fruitful. we are not starting out with a general theory about global society, such as those of d. bell, h. braverman and others, no matter how attractive and pertinent they may be. we intend rather to work out empirically what the overall interpretation might be, somewhat like the method of oxford analytica in america in perspecuve . an example will illustrate our method. it is already possible to gather, from the analysis of several indicators, some spring 1990 43 overall trends; an increase in individualism, mobility in daily life, a change from hierarchy lo network, etc. these broad trends are the result of observation of a large number of indicators. it is now the analyst's task to interpret the, to discern all the implications, so that the process results in a tentative generalization or theory about all the data. it must be noted that the proposed study of trends will try to work out, as far as possible, the variations in the subgroups such as age, sex, social and cultural group and region. the trends are not consistent and enormous differences exist side by side. this has been shown in sevctal recent studies, bella's work, for example. ttie analysis of trends will not be limited to the study of average or means only. studying trends in four industrialized societies our proposed research project intends to identify the principjj trends with characterize global society by doing secondary analysis of existing data. therefore, the first task will be to trace and synthesize published or available observations and analyses. up to now, nothing original, for such studies yet exist in number, on all the possible objects. for example, canadian social trends publish excellent analysis on different indicators and different trends. the same for publications like british social attitude . social and cultural report (holland), works published by eurobarometers, etc. there are fewer analyses of relationship between trends, and fewer again are the attempts to study systematic relationship between a large set of trends. this is precisely what we are planning to do. the level of the proposed analysis is somewhere between the social reporting and the systemic analysis of a large, global, macro-tend which summarize a very large sample of specific trends. the louis dim team in france took the initiative by inviting other research groups from various countries lo undertake a comparative analysis of trends, adopting as closely as possible the methods which they have worked out over the last several years. three teams have already agreed to participate in the project, one from the usa, led by theodore caplow, a german team led by glatzer and hondrich, and an iqrc team from quebec, led by s. langlois. a preliminary meeting was held in paris in may of 1987 to establish the main goals of the project and to choose which trends which would be observed. the project at hand is now part of a co-operative venture involving several teams from various countries, teams which share the same goal and approach to the study. the quebec team, based at iqrc, has been given the secretarial role in this little group, coordinating the development of this project the area to be covered and the indicators are not chosen for the purpose of verifying a particular theory of social change, nor to illustrate a dominant trend (for example, the increase in individualism in contemporary society). we know as well that many of the indicators of social change which are used, in scientific analyses and in public debate, reflect an era when the majority of the population spent most of its time working and the times when survival was a daily challenge: unemployment, standard of uving based on steady income, poverty, jobs and social position were and are the indicators most relied on in studies of the social structure. we do not deny their relevance, quite the contrary, but it seems to us necessary to develop others which can reveal ongoing social changes (owning a second home, mass media consumption, touristic travel abroad, etc.), indicators which are usually outside the sphere of work which will point out new trends. anything to do with the world of work will, however, still have an important place in the analysis of trends. the final choice of the data to be analysed, of the trends to be examined, was based on hypotheses in existing monographs or put forward in some theories about social change, but also, relied on the observations of researchers and on research going on elsewhere, always taking available date into account a number of 79 trends were identified at the first meeting of the international group in which our team is participating, based on a preliminary list drawn up by the french team. we agree that the area to be covered is vast it can be done, however. the louis dim group's first attempts which is going on in france showed this. basically, it is a question of drawing up a brief synthesis of what is known, in the form of trends, about each of the things on the list each team undertakes an analysis of the trends which describe its own society, while following the common method as closely as possible, basically that suggested in the article by l. dim, 1985, in order to allow for later comparative study. the importance, or the interest, of a comparative study is obvious. it is, among other things, an excellent way to the extend the analysis of causal relations. let us lock at a known example: has development of education promoted an increase in intergeneralional social mobility? there may be a whole new light put on the analysis of relations between these two trends after comparative examination by four societies. this comparative study of different societies poses enormous problems. faced with these problems, there are two attitudes lo lake. we could do nothing, since there are so many difficulties. or we would try lo iron out the problems in order to prepare a trial method, however imperfect it may be at first. this is what we decided upon. in order to smooth out the difficulties, it was agreed that a common grid of theses for the study of trends would be drawn up and that as much as possible, the representative data conceming the overall society would be used and the measures to be used would be clearly elucidated. but, above all, it was agreed what the proposed comparative analysis would bear on the direction of &ends and the relations among the trends themselves. this approach minimizes the problems of comparison to some extent. thus, it is not a rate or a precise measure which is to be compared (rate of unemployment, real income, etc.) but the direction of a trend and its relation to other trends. lassist quarterly the data we will be carrying out a secondary analysis of existing data and a synthesis of published works on a given subject. we will be trying to obtain a set of statistics, standardized as much as possible, starting from 1%1 if possible, or from 1970, so as to have at least about fifteen years of observation in order to discern a trend. where statistics may not be available, we will look for data observed for at least three separate periods of time, again so as to discern trends. this will be done mainly for the secondary analysis of data from surveys and the study of changes in social perceptions. we have already identified about 250 series of statistics with which to characterize the trends we will deal with. others will be added as the project advances, because the research consists precisely of identifying or even constructing such series for later analysis. the list of data is too long to include here. examples are number of automobiles, real income per capita, circulation of daily newspapers, etc. priority will be given to quantitative data which can be compared in a given society and later between societies. preuminary examination indicates that for the majority of trends in the attached list it is possible to fmd out at least one set of statistics. this preliminary data base will be completed by qualitative observations and analyses or by monographs concerning "phenomena portending the future", always working from secondary sources. these observations will clarify the process of some extent scientinc and social significance of the project this project may seem ambitious to some. however, the french experience during the past four or five years proves that it can be done and the results can be fruitful. the scope of the project is broad: the diagnosis of global society and its social changes using secondary analyses. one of the project's interests is to use existing date in order to arrive at a more in-depth analysis. the comparative dimension should also be stressed. despite the difficulties that this presents, we believe that the results will be {woductive. the comparison with other countries will allow us to go ahead with certain interpretations. comparative charting of social change list of trends (quebec, december 1988) 0. context 0.1 demographic trends 0.2 macro-economic trends 0.3 macro-technological trends 1. age groups 1.1 youth 1.2 elders 2. microsocial 2.1 self identification 2.2 kinship networks 2.3 community and neighbour-hood types 2.4 local autonomy 2.5 voluntary associations 2.6 sociabiuty netwwks 3. women 3.1 female roles 3.2 childbearing 3.3 matrimonial models 3.4 women's employment 3.5 reproductive technologies 4. labour market 4.1 unemployment 4.2 skills and occupational levels 4.3 types of employment 4.4 sectors of the labour force 4.5 computerization of work 5. labour and management 5.1 structuring of jobs 5.2 personnel administration 5.3 size and types of enter-prises 6. social stratification 6.1 occupational status 6.2 social mobiuty 6.3 economic inequality 6.4 social inequality 7. social relations 7.1 conflict 7.2 negotiation 7.3 norms of conduct 7.4 authority 7.5 public opinion spring 1990 8. state and service institutions 8.1 educational system i2 health system 8.3 welfare system 8.4 presence of state in society 9. mobilizing institutions 9.1 labour unions 92 religious institutions 9.3 military forces 9.4 political parties 9.5 mass media 10. institutionalization of social forces 10.1 dispute settlement 10.2 institutionalization of labour unions 10.3 social movements 10.4 interest groups 11. ideologies 11.1 political differentiation 1 1.2 confidence in institutions 1 1.3 economic orientations 11.4 radicalism 1 1.5 religious beliefs 14.3 athletics and sports 14.4 cultural activities 15. educational attainment 15.1 general education 15.2 professional education 15.5 continuing education 16. exclusionary phenomena 16.1 immigrants and ethnic mincxities 16.2 crime and punishment 16.3 behavioral and emotional 16.4 poverty 17. attitudes and values 17.1 satisfaction in life domains and in general 1 7.2 perceptions of social problems 17.3 orientations to the future 17.4 values 17.5 national identity 'presented at the ifdo/iassist 89 conference held in jerusalem, israel, may 15-18, 1989 12. household resources 12.1 personal and family income 12.2 informal economy 12.3 personal health and wealth 13. lifestyle 13.1 market goods and services 13.2 mass information 13.3 personal health and disorders beauty practices 13.4 time use 13.5 daily mobility 13.6 household production 13.7 forms of erotic expression 13.8 intoxication 14. leisure 14.1 amount and use of free time 14.2 vacation patterns lassist quarterly the international research group for the comparative charting of social change (club de quebec) guidesines for new members introduction the international research group on the comparative charting of social change in advanced industrial societies, informally known as the club de quebec, was founded in 1986 by a group of sociologists and historians from france, west germany, canada, and the united states, who had been studying social trends in the respective countries. these separate studies attracted a fair amount of .scholarly and popular attention but did not advance our general understanding of contemporary industrial society as much as they should have, for want of a comparative perspective. without systematic international comparisons, it is impossible to know whether trends we discover in national societies are local accidents or features of a larger system. another reason for combining our efforts was that we had been using different methods to delineate social trends and wanted to standardize them to facilitate comparison. the project was initially sponsored by council for european studies, but each national team has found the financial support of its own research operations. the canadian team assumed responsibility for maintaining a central secretariat, at the universitd laval in quebec. meetings have been hosted in rotation by national teams: at charlottesville in april 1986> paris in april 1987, bad homburg in june 1988, quebec in december 1988, and charlottesville again in may 1989. at quebec in december 1988, responding to a request from professor constantin tsoukajas of ijie national social research center in athens, the members voted to admit additional national teams on the following conditions: 1 . persons proposing to establish a national team, and professional personnel subsequently added to a national team, must be individually approved by a vote of the genera] membership of the international research group. 2. national teams are obligated to follow the research methods adopted by the international research group, including the standard format of trend reports, the list of trends for which reports are prepared, and the criteria for acceptable data. 3. trend reports may be written in any language, but each national team has the responsibility of eventually making its reports available for publication in english. 4. members of any national team shall have full access for their own scholarly purposes to the data gathered by other national teams, and shall make their own data reciprocally available. the research problem social change is too large a topic to be manageable without further specification. we are specifically interested in the late twentieth century, the industrialized or partly industrialized nations and the social structures and institutional patterns that characterize the behavior of mass societies, especially those associated with the family, voluntary associations, work, leisure, education, religion, government and politics. our unit of analysis is a trend , i.e. a series of values representing the incidence of some item of social behavior in a given population at points of time in a consecutive sequence. most of our work has been done with time series 10 to 60 years long, ending as recently as possible, and covering such matters as family income, household expenditures, employment and unemployment, working conditions, the informal economy, marriage and divorce, household composition, kin networks, housing, migration, educational achievement, criminality, leisure patterns, health care, social movements, and so forth. in the scholarly literature, a few trends have received the lion's share of attention. economists have looked very closely at trends in economic growth, prices and wages. political scientists have studied twentieth century trends in voting and party affiliation. demographers have scrutinized trends in fertility and mortality. it is no coincidence that these are the areas of social life which lend themselves most readily to quantification and offer the longest time series. but a description of social change that limited itself to trends in economic development, political participation and population would be incomplete indeed. even though quantification is initially more difficult in other institutional sectors, many of the difficulties have been overcome in recent years, and we can anticipate that the quality of data will continue to im{hx)ve. the international research group's standard list of spring 1990 trends and indicators currently includes 77 trends grouped into 18 major categories. it is attached hereto. note that the development of a national profile calls for the preparation of 77 trend reports, each cwresponding to one of the numbered subheadings in the list, from 0.1 demographic trends to 17.5 national identity. the first four national teams have already finished this phase of work, which takes from 12 to 36 months depending upon the manand woman-power available, and they are currently engaged in the more challenging task of comparing their results. we anticipate that new national teams will complete their national profiles at various times during the next five years and we believe that this staggered schedule may be intellectually advantageous. following the list of trends and indicators, you will find, first, the standard format for trend repots, and second, a review of the project's long-term goals by two participants. format for the presentation of trend reports coverage a trend report is i»«pared for each numbered subheading in the list of trends and indicators, e.g. 0. 1 demographic trends, 0.2 macroeconomic trends, 0.3, technological trends. the items listed underneath each subheading constitute a check-list of topics that should be covered in the trend report, but the checklist is not intended to be inclusive. arrangement a trend report normally has four sections: a brief summary of about five lines at the beginning, an explanatory text, a section of tables and figures, and a bibliography. each table or figure is placed on a separate page. factual statements in the explanatory text, as well as all tables and figures, should be fully referenced, but the bibliography is usually much more extensive than would be required for direct referencing alone. if possible, trend reports should be prepared on an ibm or ibm-compatible personal computer, using word perfect for the text labels every page of every trend report is labeled for each identification, as in this example: [ccsc-us, 3/8/88, trend #4,1, draft #3, ac, page #4] interpreted as follows: comparative charting of socialccsc-us change, trend #4.1: draft #3: ac: case page 4: taken from list of trends and indicators, 4.1 unemployment. self-explanatory preparer's initials, a. carrier in this of this trend report criteriafor data 1 . the data were obtained by empirical measurement or observation. 2. the data refer to an entire national society or a representative sample thereof. 3. the data may be expressed as a time series. 4. the time series covers a period of at least ten years, ending in 1983 or later. 5. the time series include measurements or observations for three or more time-intervals, obtained contemporaneously. 6. the data are defined in such a way that the measurements or observations can be replicated in other national societies. 7. the data are defined in such a way that the measurements or observations can be replicated in the same national society in future years. 3/8/88: united states date prepared 48 lassist quarterly vol30-1.indd iassist quarterly spring 2006 by by helena laaksonen, sami borg & janez stebe* setting up acquisition policies for a new data archive introduction newly established or emerging european social science data archives, like the finnish social science data archive (fsd) and the slovene social science data archives (adp), work in national settings where research with secondary data is not deeply rooted. in such a national setting a new service provider must have an active role to prove its value. on the one hand, the archive should relatively quickly supply sufficiently interesting data repository for the scientific community to attract secondary research. on the other hand, it should also be active on the ’demand side‘ by promoting the usefulness of secondary data in research and training. by raising general awareness about the benefits of secondary research and placing more emphasis on national institutional support and regulation, a culture of data sharing arises in a national setting. this data sharing creates an environment for higher quality acquisition. our paper discusses fsd’s and adp’s data acquisition policies in relation to their general archival development strategy. we will attempt to systemize the relevance of policy issues by taking a look at current acquisition practices. we will also share experiences from the very grass root level acquisition practices and try to shed light on who are the actors and how their interests play a role in the acquisition. additionally, we discuss the national institutional support for data sharing. we report the practices of national research funding organizations and present our colleagues’ views on how the current practices are operating. data archives are typical service sector establishments. a quality of their service can be assessed from users’ perspective. how relevant is a service provided by archives to users requirements in type and content. a typical user is a researcher planning secondary analysis; therefore generally the acquisition procedure should be focused on assessing the usefulness of the data for scientific inquiry1. national statistics offices have established a set of criteria for the quality of the statistical data serving users requirements that can also be used in the context of a data archive supply. the criteria are accuracy, timeliness, accessibility, comparability, coherence, and finally, cost and burden associated with the products (eurostat 2000). accuracy means that the data needs to meet certain methodological standards; timeliness that the period between collection and distribution should be short. accessibility to the data is increased by additional processing and distribution channels used by data archives. the added costs and obligations from the archiving process, and to the data providers, are nevertheless minor in comparison to the whole cost of data collection. the provision of micro data also requires additional confidentiality measures to be established to minimize disclosure risks. acquisition and data sharing appropriate data acquisition policies support the principles of data sharing and open access to research data. these principles have also recently been highlighted in the oecd declaration of open access to research data from public funding2. widely adopted and efficiently organized data sharing is expected to provide, among other things, better research and science ethics, better quality of research and learning, and more efficient use of public funding. the benefits just mentioned are well argued and accepted in relevant literature. (hyman 1972; miller 1977; royal statistical society & the uk data archive 2002). however the level of the culture of data sharing is uneven among countries. acquisition policies and archival development of a single data archive must be adjusted to its mandate, mission, the resources of the archive, and to its national operational environment. in finland fsd has had quite a strong position, as it is the first and so far the only national unit whose main aim is to preserve and disseminate research data for secondary use. fsd also has a clear, and too restrictive, mission to archive social science data, and to serve the universities as a national service provider. as the infrastructure for research with secondary data has emerged at the national level only recently, the national culture for data sharing is still rather weak in most science disciplines and organizations. for this reason, our article also takes a look at the national environment and institutional support for data acquisition. 6 iassist quarterly spring 2006 the role of slovene adp in its national setting is similar to the extent that it declares its position as a national social science data archive. this implies certain obligations about the acquisition policy that is implemented in practice. users would expect to find all important studies in a collection. whether or not this is the case, it should be accompanied with the appropriate measures to achieve that goal. after attaining sufficient coverage of narrower fields of social science disciplines (sociology, political science and media studies) in its first years of existence, a need may arise to go beyond those disciplines into some of the areas that rely heavily on the empirical data: economy, education, psychology. one can easily locate the main data producing institutions and individuals in those fields, but to establish contacts and share the same understanding of a data sharing culture is yet another story. common to both the fsd and the adp is a rising awareness that a more systematic acquisition policy is needed after the starting period of abundance of data supply, when a whole set of legacy studies were at the priority of acquisition. now is the time for a systematic evaluation of gaps in a collection and of a more active approach to identify and attract new and important contemporary data sets. after the initial period, a natural claim arises to extend the thematic coverage to disciplines and areas that are partially beyond the limits set at the institutional establishment. developments in scope and coverage until now fsd’s focus has primarily been on archiving quantitative social science data. the present collection includes about 600 datasets, all fully documented with the ddi format. for most datasets, documentation is also available in english3. documentation and processing levels for different types of data were set recently. a new classification and its implementation will hopefully help to adjust future data acquisition policies with the general strategic development of the archive. most typical datasets in fsd’s collection are nationally representative quantitative surveys from barometers of some non-university national organization. their time coverage is usually 10 15 years. now the archive is also trying to acquire more data from educational sciences and, as with the adp, from health sciences, too. acquisition of qualitative data started some years ago, but their share of the collection is not more than five per cent. compared to quantitative surveys, acquisition of qualitative data has proven to be more time consuming and labor consuming. more often than not, the sufficient legal requirements for archiving are not met, as the primary researcher has promised to the persons interviewed that (s)he is the only one who will use the data. many collectors of qualitative data also tend to think that their data are their personal property. therefore, in this area, fsd presently aims to affect the overall development of research guidelines and agreements, in order to move them in a direction that would allow easier access to qualitative data. the depositors in adp, while mostly academic, do include public and private institutions which gather data themselves. more than 400 studies are currently in the collection. over half of them are processed to a highest level that includes complete variable level documentation in an electronic codebook. the study level descriptions are available in english4. initiated with the special project, supported by the ministry of information society in 2002, the adp started collaboration with the statistical office of the republic of slovenia and some commercial marketing institutes engaged in data collection. that experience showed that when there is an explicit agreement to share in efforts to process data sets, and when there is additional financial support that covers part of those efforts, then there is also willingness to contribute. qualitative data are only sporadic in the slovene adp. the archive has not developed special procedures for processing qualitative material. rather, appropriate sections in ddi format have been used to describe the material and to store the actual qualitative evidence in an electronic format (e.g. scanned paper versions). it is one of the fields that calls for additional efforts in acquisition, as this is the type of data that is not readily available due to different research tradition and corresponding notions of the usefulness of secondary research. criteria for evaluating datasets and deciding what to archive basically, the fsd and the adp criteria for evaluation of data is similar to many other european social science data archives. key evaluation criteria are linked to the scope, to sufficient legal conditions, and to the re-use potential of the data in research and teaching. the adp was not using strict criteria for selection until 2005 when it introduced a new ‘inquiry form’ to gain access to rough information about studies that is later on used for selection. reuse potential of a study is difficult to asses. one may observe current use patterns, where comparative and longitudinal national studies predominate. experts from a discipline may assist in selection to evaluate scientific potential in collaboration with the archivists, who in turn may make decisions based on administrative criteria. one may also take notice of users suggestions and evaluations of potential studies to be archived. data depositors can set different type of conditions for the re-use of their data. fortunately, this seldomly happens. in the adp, few of the data sets are under embargo or available only by special permission from the primary researchers. in most cases, the embargo or special iassist quarterly spring 2006 7 conditions apply only for a certain period of time or to users from nonacademic provenance. unlike countries that truly support the idea of open access to research data from public funding, in finland depositing data or offering it to a data archive is not mandatory in any circumstances. similarly, in slovenia collaboration in the whole process from identification to acquisition is entirely voluntary. that means that special effort needs to be invested into the process to negotiate and gain access to the data sets. on the other hand, the datasets deposited and archived are easily and equally available for researchers and university students. for the moment, the basic data services of the fsd are free of charge for all data providers and end-users. generally, the fsd does not pay for the data they acquire. only in very few cases the fsd has paid the costs of preparing data for archive. the adp does not charge for educational purposes nor does it charge researchers who don’t have institutional support. the charges are minimal for other academic or public use purposes, and full for commercial purposes. in exchange for the data provision, the adp offers the data providers free access to their data and to an equivalent amount of data provided by other depositors. they require collaboration in the processing stage, but the archive personnel tries to do most of the editing work themselves and asks only for a proofreading of the final version of documentation. these measures, together with a set of guidelines that describe the goals and procedure of data acquisition for potential depositors, are set up to increase willingness to supply data to a data archive by reducing their burden. thinking of different types of coverage issues, emphasis has mainly been on data with sufficient, and at least national geographical coverage. the fsd has tried to acquire both panel data and time series that allow examination of trends with several cross sections. unfortunately, the supply of panel data has been very moderate. until now, the work load from updating data and metadata (that has been processed once already) has not yet increased. this allows both of the archives to focus on efforts that aim at a rapid increase of datasets in its holdings. the data evaluation has not yet been very selective if a potential dataset has met the formal requirements posed by the archive. the adp aims primarily at collecting theoretically or practically important studies. studies that fill a research gap or have many implications for a wide range of practical research problems, and have long term scholarly value. comparative or continuous research with methodological excellence have top priority. in practice, the archive does not reject the data offered, based on the selection criteria. in cases of occasional studies of low quality, the data and materials are just stored as received, without further processing. in active searching for new studies the adp strives to attain most relevant data, to build high quality collection. the studies that comply with the above mentioned criteria receive the most intensive processing and full electronic documentation. it involves conversion of paper documentation to electronic format and variable level documentation that includes full question texts. from localisation to receiving the data for archiving data acquisition is very labour intensive. researchers seldom contact the archive expressing a desire to deposit their data for wider use. the staff has to be active and persistent. this is probably the case for most data archives in countries where archiving data is voluntary, and maybe even if it is compulsory. in the following, we share some experiences in acquisition work in finland and slovenia. the acquisition process can be divided into three distinctive stages: identifying potential datasets (we call this stage ’localisation‘), negotiating with data creators, and receiving the data and other relevant material. after the last stage, additional information and material are often required to be able to create the metadata and a dataset version suitable for secondary use. to keep track of all the contacts made regarding any particular dataset, an efficient operative database is needed. everyone involved in the acquisition process has to enter information on contacts made into the database. nobody can remember everything, and it may take years from the first contact to the moment when the archive actually receives the material. there are now over a thousand potential datasets recorded in the database of the fsd. in fsd’s new internal operational database5, the archiving process can be followed right from the identification of a potential dataset to the publication of metadata in the archive’s online catalogue. according to the exchange theory one is most likely to respond to an inquiry ‘when perceived costs of doing so are minimised, the rewards are maximised and the expected rewards will be delivered’ (dillman 1983). with the more active acquisition policy in the adp in 2005 an ‘inquiry form’ was set. it is suitable for online collection of evidence about potential new studies, so as to ‘minimise costs’ of providing initial information6. a request for giving information was circulated widely to the general user community and specifically to potential data providers. the ‘rewards’ were emphasised by stressing that providing information to a study’s inventory one will contribute to a user-centred data collection with the more exhaustive coverage of high range studies. a modest initial response to this inquiry shows that only additional, more personalised 8 iassist quarterly spring 2006 communication, in most cases e-mails followed by telephone calls, produce positive responses. the adp aims to get 30 new studies into the archive each year. in the fsd the goal has been around 80–100 studies, including the international surveys. the adp, thanks to the inquiry and personal contacts has been able to fulfil the goal in 2005. the studies that are being processed following this initiative will add to topical variety. almost all of the potential new studies were of highest quality so that no selection was needed except based on the criteria of availability of materials for further processing and distribution. the finnish archive has reached or almost reached the quite ambitious goal in most years. identifying potential contemporary datasets (collected in the 1990s or later) when identifying potential acquisitions, the fsd is looking for data with potential for secondary use. primary sources of information are academic journals and the news media. joining a wide range of e-mail lists has also proven useful, likewise regular browsing through universities’ online publication catalogues and the web sites of research funding bodies. the majority of relevant new data can be traced to these sources. academic literature is less useful, as information published in monographs or in articles of edited collections has usually been published somewhere else previously. in slovenia, main sources of information to identify potential new studies are personal contacts with colleagues that are experts in a field and recommend their own or their close collaborators’ studies. other productive sources for first information are news media, as it is widely adopted practice to have a news conference after the first results of a research project are published. other sources are sicris (slovene current research information system)7 that covers publicly funded research projects since 1998. updates on information gathered through the course of the research project are often missing, including if a project is of empirical nature and if a data set exists at all. the other information system is cobiss8 which is a national record-keeper of all slovenian libraries and includes an upto-date bibliography of registered researchers. institutional home pages are another source of information, useful in particular when one knows what one is looking for and to add reference to initial information, gained from some other sources. the adp also keeps and updates information on availability of some international data sets kept in other national archives or in specialised ad hoc project archives. all in all, work on data acquisition is labour intensive and demanding task already in a stage of identification of datasets. detective work to preserve older data when it comes to tracking down older research material, literature reviews, and particularly the contacts with experienced academics, are important. one can naturally browse through older issues of academic journals to acquire information on data colleted during past decades. however, without personal contacts it is often difficult to discover the present whereabouts of an older dataset. in finland, the data collected prior to the 1990s, even when identified, very seldom end up at the archive. it may be laborious to migrate the data into a present day format. still, this is mostly manageable. what often makes the archiving of older data impossible is that available metadata is not at all adequate. if the archive cannot find out what the variables mean, who the respondents were and how the sample was drawn, the data cannot be used. in slovenia there were some cases when no data set was available any more for some past studies, and the archive tried to preserve at least what was remaining on paper, such as reports and questionnaires. as it has proven difficult to acquire older data, its share in the collections is a minor one. around a half of the archived quantitative data were collected in the 1990s, in both of the archives, and around one third since the year 2000. thus, getting data collected in the 2000s has not been too difficult. contacting potential depositors at the fsd, all personnel are involved in localisation. the main responsibility for getting data into the archive, however, lies on the shoulders of the director and the information officer. they are the ones who usually contact the data creators. the research officer in fsd is a specialist in research ethics and qualitative research. she is mainly responsible for acquisition of qualitative data. at the adp, the personnel consists of only two persons. the initial contacting of the depositors is mainly the director’s duty, but further acquiring and processing of new studies is done by both. acquiring data is a tough job convincing researchers takes a lot of persuasion, whether in finland or slovenia. one has to repeat the same reasoning over and over again, to the same person or maybe to several persons, as it sometimes takes time to find the right person. there are two main stages. first, one needs to reach the ‘ok, i will give you the data’ agreement stage. this does not necessarily get the process much further. encouraging the researcher to act and actually transfer the data into the archive is another story. at the fsd, in one relatively easy negotiation process, where the dataset was deposited within a year of the first contact, altogether 10 contacts were made (e-mails and phone calls). in some cases it has taken only about half a year from the first contact to the arrival of the data. still, iassist quarterly spring 2006 9 several e-mails have to be sent and many phone calls made, before the actual delivery of the data. in most cases, researchers agreed to archive their data at the first contact. it was much more difficult to get them to actually send the material and deposit agreements to the archive. the biggest negotiation challenge is to persuade an individual researcher to archive a one-time study. if the fsd already has established contact with someone in an organisation regularly collecting data, the job may be easier but not necessarily. established contacts must be maintained. generally speaking, it is easier to deal with medium-sized organisations than with large ones. why? a probable reason is that large organisations lack the culture of archiving and data sharing. it is also difficult to find persons in large organisations who would take responsibility for these matters. recently collected datasets, with primary analyses made and published, are easier to acquire than older research material, which may have been put aside somewhere. however, there are a couple of moments when the archive has good chances of success: when researchers are approaching retirement age or when they move office. at these times researchers may be willing to let their beloved research data fall into the hands of the archive staff. at the adp experiences are similar to the extent that it is hard to persuade researchers to move forward. after the first contact and when usual arguments are being exchanged about the purpose and benefits of having the data set offered to the data archive, in most cases researchers are willing to provide their data. what is a main obstacle in that even finding the data files and collecting the documentation may take additional time that is extremely scarce. that adds to an argument that some form of institutional legal obligation accompanied with the financial support would probably help to reduce ‘perceived costs’. in addition, the proper reference to an authority increases mutual trust that is a basis for expecting that ‘rewards will be delivered’. pr and information services an essential part of acquisition, at least for fairly recently founded archives like the adp and the fsd, is to make the archive known, and to promote new ways of thinking within the research community. researchers have all kinds of reasons why they do not want to archive their data for re-use. everyone with experience of data acquisition is familiar with some of the reasons. qualitatively oriented researchers are the toughest cases. they need to be addressed with particular care to make them see data sharing as an essential part of the research process. in making the archive known, standard pr methods are used. however, approaching the social science research community may be tricky. one has to rouse their interest, but not appear too ‘commercial’. one needs to provide information, but as an expert rather than as a salesperson. this is not necessarily easy to carry out. the fsd distributes general information on its services through its own channels, and the personnel write short articles for the publications of other organisations and associations. newsletters on recently published data and other services are sent through fsd’s e-mail list and by post to selected target persons and groups. new generations of students and future researchers are also an important target group, as are university teachers. contacts with key persons, university libraries and research institutes are important. these bodies are informed of the data archive’s activities and collections through all possible channels. in addition, the archive staff visit them, invite them to visit the archive or to attend the conferences and seminars organised by the fsd. conference presentations and posters, articles in scientific journals, lectures for students of different faculties, and other channels of information about the adp’s activities and holdings are more often being conceived as an essential part of its activity. that this is not an easy task shows evidence that in personal contacts the adp personnel still encounter researchers who are unfamiliar with the mission of the adp. most often, the reaction after explaining the basic principles of data preservation and reuse, is exclamation of surprise and of support. the adp plans to integrate into its activities thematically dedicated workshops. these workshops will share the specialised practical knowledge of the eminent researchers, who have long term experiences with the analysis of particular data sets, with the wider community of users. in the end, only intensive contacts and personalised communication have been shown to achieve the desired results. to promote a new culture of data sharing within the research community, it is also necessary to proclaim the virtues of open access, transparency, possibility of replication and validation of research, and other good research practices. these issues must be brought up repeatedly when contacting individual researchers. however, caution is needed to not get the opposite result from the one intended. as mentioned, researchers are a tricky bunch of people. it is often useful to appeal to the researchers’ own interests. what are the advantages for them? a free-of-charge, reliable preservation for their data, and of course, fame and glory, when other people get acquainted with their excellent data, and cite them. however, using the latter argument may sometimes backfire. once the principal researcher told to an fsd employee that the data was ‘so bad’, i.e., of so low quality that he did not want to share it – and refused to archive it on those grounds. on that ground one may claim that willingness to provide the data to a data archive is also a guaranty of its consistency and overall quality, that can 10 iassist quarterly spring 2006 be further tested by secondary analysts. which is another argument that public funding agencies could use to make the data sharing a legal obligation. after all, the data are made on public money and with the collaboration of lay people. support from the funding organisations in europe during the spring of 2005 the fsd turned to 20 european data archives to collect experiences of what kind of support social science funding organisations give to archiving and data sharing. nine data archives sent their answers. a short summary and some examples of practice are provided below. the fsd thanks all contributors9. we asked: 1) what kind of guidelines or regulations the research funding organisations have on the archiving, accessibility and re-use of research data, 2) in what kind of a document do the guidelines appear, and 3) are the guidelines implemented, i.e. are researchers doing what the guidelines say they should do. in all but one country (italy) at minimum one funding body has at least a recommendation to deposit data for archiving. in some cases it was difficult to interpret whether it was a recommendation or a requirement. according to our interpretation, in five countries the funding body or bodies had taken a stronger stance for archiving than mere recommendations. however, even when the funding is given ‘on the condition that data be deposited’ there are seldomly sufficient effective ways of controlling that the demand is fulfilled. the most effective means, a sanction actually, is set in the uk, where the project might not receive part of its funding, if it has not offered its data to the data archive. this applies to funding received from the economic and social research council (esrc). accordingly, the guidelines are implemented at a much higher level than earlier. recommendations or not, very often it seems to be the archives who have to do the policing afterwards. yet, in some countries the funding body itself checks to determine whether or not its policy is followed by the research projects. for example, in germany, national science foundation (dfg) monitors final project reports to check whether data has been deposited and the science council checks in its evaluations of institutes, whether they conform to policy recommendations. in switzerland, the swiss national science foundation (snsf) selects a subset of projects in which they explicitly ask the primary investigator to take contact with sidos to discuss the opportunity of depositing the data for further use and to include information on that contact in the first intermediate report. if the report is missing, a reminder is sent to the researcher. the sidos evaluates the case together with the pi and makes a recommendation. the data on which sidos and pi agree are usually deposited. we were also interested in financial support for research projects to prepare data for archiving and re-use. we asked whether researchers include expenses for preparing data for archiving and re-use in their research proposals to funding bodies. we also asked whether they would get the money if they applied for it. there is only one country where the research projects usually apply for money to prepare data for archiving. that is the uk, where they also have a sanction for not depositing the data. in five countries, researchers almost never or never apply funds for this purpose. in one country they sometimes do apply, and in one they seldom do. yet, seven of our respondents estimated that, if applied for, funding would usually be given to the projects. in three countries (according to our interpretation of the answers) they would usually not get the money. further, we asked whether the archives have tried to influence the research funding organisations’ policies regarding data sharing and archiving, and regarding financial support for preparing data for archiving. in all countries, the archives have been and will have to continue to be more or less active in lobbying for the cause of archiving and data sharing. it is clear that newer archives have a longer way to go to receive full institutional recognition and support from national legislation. the role that the data archives have is similar to well established ‘national heritage’ institutions like national libraries, museums and classical archives. new data archives could make a step further in this direction with the help of examples of good practice in countries where infrastructural role of a data archive is to support high quality scientific production and high quality education. concluding remarks at present, both the fsd and the adp are clearly moving from the first phase of a newly established archive to a second phase. at first, they concentrated on building up the data collection, and could not afford to be very selective. they needed to be active, and contacted a large number of persons and organisations within respective research communities. now the reserve of potential, unidentified datasets is diminishing within the core social science disciplines. but at the same time the data preservation needs of other disciplines are growing. therefore the fsd will, based on its own decisions, expand its coverage to some related research fields, mainly health and educational sciences. including more qualitative data into the holdings clearly is also a very important choice of acquisition during this second phase of development. the adp intends to cover the fields mentioned and in addition intends to explore more intensive collaboration with economists and iassist quarterly spring 2006 11 psychologists. the archives hope to see the third phase of development soon. that would require more institutional support for them, especially from research funding organisations. to ensure a steady flow of data into the archive, new institutional policies are needed. the projects and current policies aim at establishing a research culture where data sharing is considered an inherent part of a research project – to be taken into account already when preparing research proposals. the cost of preparing data for archiving should be included in research applications. researchers are to be advised to include into contracts and communication with sponsors, and human subjects who collaborate, an explicit agreement about data delivery to the data archive for the purpose of secondary analysis. both online and printed material will be produced to promote this view, and all these actions will be linked to the national implementation of the oecd guidelines (oecd 2004). major research funding bodies should be encouraged to place more weight on the issue, and to include in their funding decisions recommendations or requirements for the data to be archived. when implemented, such institutional changes would improve the overall efficiency of acquisition in a strategy, which can set the requirements of a quality and relevance for new studies higher. that can be accomplished if the coverage of the studies accessible for acquisition would be almost complete, that is, if information about studies would be supplied regularly and data ready for processing without further obstacles. researchers already in a planning stage of a project would be advised to include an option for giving data to an archive when negotiating a contract with sponsors, to include initial preparation of data and documentation among the tasks, and communicate with the research subjects about the secondary use. if all this will happen in the near future, our next paper on data acquisition might concentrate more on the issue of designing and improving the collection. the portion of staff time that is currently devoted to initial negotiations about sharing the data could be concentrated on quality of processing and distribution instead. references dillman, don a. (1983): mail and other self-administred questionnaires.” in: rossi, p.h., j.d. wright, a.b. anderson (eds.): handbook of survey research. academic press, inc., san diego. 359-378. eurostat (2000): standard quality report. assessment of the quality in statistics. luxemburg, 4-5/04/2000. www. unece.org/stats/documents/2000/11/metis/crp.3.e.pdf gutmann, m., k. schürer, d. donakowski and h. beedham (2004): the selection, appraisal, and retention of social science data. data science journal, 3, 2004. http://journals. eecs.qub.ac.uk/codata/journal/contents/3_04/3_04pdfs/ ds386.pdf hyman, h.h. (1972): secondary analysis of sample surveys. new york: john wiley & sons, inc. miller, w.e. (1977): “the less obvious functions of archiving survey research data”. in: hofferbert, r.i., j.m. clubb (eds.) (1977): social science data archives. beverly hills: sage. mochmann, e., p. de guchteneire (1988): “data services for the social sciences”. in: l. kiuzadjan, k.t. saelen, g. soloviev: information needs, problems and possibilities. vienna: european coordination centre for research and documentation in the social sciences (vienna centre). http://www.ifdo.org/data/data_archive_workflow01.html oecd (2004): science, technology and innovation for the 21st century. meeting of the oecd committee for scientific and technological policy at ministerial level, 29-30 january 2004 final communique. http://www.oecd. org/document/0,2340,en_2649_34487_25998799_1_1_1_ 1,00.html royal statistical society & the uk data archive (2002) preserving & sharing statistical material. working group on the preservation and sharing of statistical material: information for data producers. colchester: uk data archive, university of essex. sivonen, jouni: (2004) new user interface for managing the archiving process in fsd. paper presented at iassist conference, madison, may 2004. see: http://www. iassistdata.org/conferences/2004/presentations/c3_sivonen. ppt * the article is based on two presentations at the iassist conference, edinburgh, 25 may, 2005 in session a3: enlightened policies: improving collections and acquisitions. the presentations were pulled together and updated in 2006. helena laaksonen & sami borg, finnish social science data archive, janez stebe, social science data archives, university of ljubljana, slovenia. email: helena.laaksonen@uta.fi, sami.borg@uta.fi, janez. stebe@fdv.uni-lj.si endnotes 1 “a general guideline is whether on not the data are usable for future scientific research” (mochman and guchteneire 1988); “the extent to which the data will advance knowledge” (guttman et al. 2004). 2 oecd 2004. after the declaration the member countries have set up a working group to specify possibilities and 12 iassist quarterly spring 2006 practices for the implementation. 3 see more at http://www.fsd.uta.fi/english/data/ 4 see http://www.adp.fdv.uni-lj.si/opisi/ 5 fsd’s operational database was presented at iassist 2004 (sivonen). 6 see http://www.adp.fdv.uni-lj.si/edan/. 7 http://sicris.izum.si/ 8 http://cobiss.izum.si/ 9 the inquiry was sent to 20 data archives in europe, and nine responded. the answers were provided by ekkehard mochmann for za, germany; reto hadorn for sidos, switzerland; susan cadogan for the uk data archive; hans jørgen marker for dda, denmark; janez stebe for adp, slovenia; carlo pisano for adpss-sociodata, italy; iris alfredsson for ssd, sweden; marion wittenberg for steinmetz archive, the netherlands; and gry henriksen for nsd, norway. a more comprehensive overview of the answers was given in the appendix to the actual iassist paper by s. borg & h. laaksonen (2005). it can be attained from the fsd. vol223 spring 1998 17 the history data service – using technology to enhance access by cressida chappell, oscar struijvé, sheila anderson * the history data service (hds) <http:// hds.essex.ac.uk> is funded by the uk joint information systems committee (jisc) <http://www.jisc.ac.uk/> to collect, manage, and encourage re-use of digital resources which result from or support historical research and teaching. the hds is located and integrated in the uk data archive <http:// dawww.essex.ac.uk/> and is the arts and humanities data service (ahds) <http://ahds.ac.uk/> service provider for the historical disciplines. the ahds provides archival, training and other functions to the archaeology, history, performing arts, textual studies and visual arts communities, and consists of five subject-based service providers and a managing executive. the hds collection covers a time period from the late tenth century to the mid twentieth century, and includes a wide range of historical data, which has been transcribed or compiled from original sources. the hds is committed to using technology to improve access to its collection through a programme of work that is essentially needs rather than technology-driven. although this programme of work aims to make effective use of established and state of the art technologies and concepts for data and metadata storage, presentation and delivery, the real emphasis is on recognising and responding to endusers’ needs. the hds has an active and ongoing policy of consulting with actual and potential users. for example, in april 1998 the hds held a workshop <http:// hds.essex.ac.uk/reports/user_needs/final_report01.stm> to explore, assess and prioritise the needs of end-users in the historical community. the central goal of this programme of work is to improve and increase access to historical data by making the location, identification, assessment and use of data easier. the objective is to get more people using and experimenting with historical data, and the hds is aiming to gradually extend its user-base into the less computerliterate and non-computing segments of the historical community by acting as an information provider as well as a data provider. in a very broad sense, the hds is seeking to improve the relationship between end-users and data and it is essential that this is carried out within existing resources. the hds approach to using technology to enhance access is underpinned by a model called testmix, short for telescope, stethoscope and microscope. the principle behind this model is that users want to locate, identify and assess data. the model maps user needs and actions to system functions, and it maps the system functions to the system components or tools required by users. the top level of this model deals with user actions and needs. from left to right we have the need to discover information about data centres, the need to identify and select suitable data, and lastly the need to make detailed assessments of the relevance of data, along with the need to explore data in detail and make sub-selections. the next layer maps user needs and actions to system functions, so we have the provision of information about data centres, the provision of searchable data descriptions, and lastly the provision of data ‘close-ups’ and subsets. the next layer maps the system functions to the system components or supporting tools that are required, so we have gateways and search engines likened to telescopes; catalogues and thesauri likened to stethoscopes; and browsing, subsetting and visualisation tools likened to microscopes. the hds uses testmix as a framework to structure and prioritise development work; to clarify relationships between services and systems; to identify scope for improvement and collaboration; and as a focus point for technical and service activities. it stimulates a coherent approach, and provides simple metaphors for communication. within this framework the hds is implementing a multilevelled strategy to improve and increase access to data. increasing the number of metadata access points the first level involves increasing the number of metadata access points and the hds has two different approaches. one approach uses conventional online catalogues and operational examples include the uk data archive’s information retrieval system, biron <http:// biron.essex.ac.uk/cgi-bin/biron/> and the cessda idc (council of european social science data archives integrated data catalogue) <http://dastar.essex.ac.uk/ cessda/idc/> information about the hds collection is http://hds.essex.ac.uk http://hds.essex.ac.uk http://www.jisc.ac.uk/ http://dawww.essex.ac.uk/ http://dawww.essex.ac.uk/ http://ahds.ac.uk/ http://hds.essex.ac.uk/reports/user_needs/final_report01.stm http://hds.essex.ac.uk/reports/user_needs/final_report01.stm http://biron.essex.ac.uk/cgi-bin/biron/ http://biron.essex.ac.uk/cgi-bin/biron/ http://dastar.essex.ac.uk/cessda/idc/ http://dastar.essex.ac.uk/cessda/idc/ 18 iassist quarterly also being made available through the prototype ahds integrated access gateway <http://prospero.ahds.ac.uk:8080/ahds_live> which is based upon the dublin core and the z39.50 network applications protocol, and which acts as a virtual union catalogue for the collections of the five subject-based service providers. in the coming months information about the hds collection will also be accessible via the cheshire information retrieval system – an sgml-based system which utilises the data documentation initiative (ddi) codebook dtd. the other approach uses a tree-based structure as an alternative way of accessing information about the hds collection. this will allow users to adopt a ‘drill down’ approach to locating data in addition to the more sophisticated search options offered by online catalogues. users will be able to drill down to metadata via three dimensions. a time dimension by centuries, a geographic dimension by countries and administrative subdivisions within the uk, and a subject categories dimension. providing online access to additional information the second level involves providing online access to additional information and will allow users to access to types of information that are not generally found in catalogue records. in particular we are interested in providing users with access to online documentation with the option to preview a sample of data. we believe that this will make it much easier for users to make detailed assessments of the suitability of data and that it is a more efficient way of supplying users with information. these services will initially apply to areas of the hds collection where there is critical mass of related materials, because there is a greater potential to create additional documentation. we are intending to apply this approach to a collection of early twentieth century surveys, a collection of european state finance data, and a collection of electoral poll book data. these services will be freely available to all users and will not require registration. providing a data and documentation ftp service the third level involves providing a data and documentation ftp service and will give registered users online access to the vast majority of the hds collection. we envisage that users who register with the hds will be able to select and download data when they require it in suitable easy-to-use formats such as tab or comma delimited ascii. the main exceptions will be difficult to use data and the minority of hds data which has more http://prospero.ahds.ac.uk:8080/ahds_live winter1998 19 restrictive access conditions. developing online browsing, sub-setting, combining and downloading facilities the fourth level involves the development of online browsing, sub-setting, combining and downloading facilities for major collections of value-added data and will allow registered users to explore fully documented data collections online. the hds has developed this service for a large collection of nineteenth and twentieth century statistics, the great britain historical database (gbhd), and work is now being carried out on developing a similar service for a large collection of individual-level nonanonymised british historical census data which includes the 1881 census for england and wales digitised by the genealogical society of utah and the uk federation of family history societies. the gbhd has been assembled by humphrey southall at queen mary westfield college, london and incorporates demographic statistics, marriage statistics, mortality statistics, employment statistics, trade union statistics, government unemployment statistics, poor law statistics and small debt statistics. the gbhd online system allows users to sub-set this data by geographical area, at present either standard regions and/or counties. the system also allows users to specify the tables to be searched for relevant data and to specify the variables to be included in the result. finally the resulting subset and customised documentation can be viewed and browsed online or downloaded to the user’s workstation by ftp. data is formatted as fixed width ascii with headers, which can be imported into a variety of software packages for further manipulation and analysis. for more information about gbhd online (including details about registering as a user) please see the gbhd online webpages at <http://hds.essex.ac.uk/ gbh.stm> the hds is confident that this multilevelled strategy will encourage and enhance use of and experimentation with the hds collection. we believe that this programme of work will widen and improve access to historical data, and that it will contribute to the creation of an infrastructure which will enable historians and others to explore the full potential offered by digital resources. *paper presented at the 1998 iassist/css conference, yale university, new haven, usa, may 19-22, 1998, by c.chappell, o. struijvé, s. anderson, history data service, uk data archive, university of essex http://hds.essex.ac.uk/gbh.stm http://hds.essex.ac.uk/gbh.stm zbw iassist quarterly fall winter 2012 17 iassist quarterly abstract this paper gives insights into the findings of the nestor working group “preservation policy” which was founded in the beginning of 2012. it is led by two of the nestor partners: the german national library and one of the goportis libraries, the zbw – leibniz information centre for economics. the working group attempts to help institutions involved in digital preservation to develop their own preservation policies. to support this task, the group has created guidelines for the development of an institutional preservation policy which will be published during the first quarter of 2014. shedding light on the policy development process and providing guidance concerning the content and structure of a preservation policy, the guidelines describe what a policy is needed for, which content it could have, which staff members should be involved in the development and how its quality can be ensured. keywords: preservation policy, digital preservation, guidelines, nestor introduction according to the iso standard “audit and certification of trustworthy digital repositories”, a preservation policy is a “[w]ritten statement, authorized by the repository management that describes the approach to be taken by the repository for the preservation of objects accessioned into the repository” (iso 16363, 2012). the standard explains that the policy has to be consistent with the preservation strategic plan. in contrast to the policy, the preservation strategy addresses how the preservation is carried out and therefore focuses on workflows and technical strategies. in practice, the policy and strategy are often (but not necessarily) addressed in the same document which complicates delimiting between the two. the iso standard requires that the preservation strategy match the preservation policy and vice versa. for example, an institution cannot state in the policy that 100% of digital content is preserved if the strategy makes it possible to consider only parts of the entire collection for preservation due to technical obstacles. preservation policies are an essential tool in digital preservation, serving both the purpose of creating trust and offering a formally binding frame of reference for the preservation activities of a given institution. however, although many institutions in germany and all over europe have already begun to engage in digital preservation, only a few have published a preservation policy of some kind (angevaare, 2011, p. 5). thus, as part of the 2011 digcurv survey of training needs, 454 institutions were asked if they engaged in storing digital material. the institutions surveyed were cultural heritage institutions such as libraries, archives, or museums from 44 countries. respondents mostly came from european countries (81.3%), but also from the united states (12.3%), canada (1.5%) and a small percentage (4.7%) from other countries (engelhardt, strathmann and mccadden, 2011, p. 10). more than 75% (n = 437) replied that they were involved in digital curation. an additional 18% (n = 437) stated that they were planning to store digital materials in the future (ibid., p. 15-17).thus, according to the survey, 331 institutions were already engaged in digital curation in 2011 and it is likely that this number has grown over the last two years. how to develop a preservation policy guidelines from the nestor working group by yvonne friese1 18 iassist quarterly fall winter 2012 iassist quarterly but what is the state of preservation policies? two resources serve to support angevaare’s claim: 1. with the aim of surveying “the current state of digital preservation policy planning within cultural heritage organizations” sheldon (2013) collected and compared publicly available policies worldwide and counted 33 documents: 15 from libraries, 16 from archives, and two from museums. it should be noted, however, that sheldon limited her analysis to published policies which are written in english (see 2013, p. 4). therefore her study excludes, for example, the bsb (bavarian state library, 2012) and the dnb (german national library, 2013) preservation policies, as neither has an english translation yet. 2. the scape wiki on published preservation policies (last updated in november 2013) lists 40 institutions with a published policy. the list is not limited to english-language material and includes dutch, german, and danish policies. as the wiki is built collaboratively and receives updates from many authors from different countries, it seems safe to assume that it is fairly comprehensive even though it surely is not complete. there is an overlap of 26 published preservation policies found by sheldon and listed by the scape authors. sheldon includes four policies not listed in the wiki, and there are 14 policies listed in the scape wiki not taken into account by sheldon. hence, 44 institutions with published digital preservation policies are known. although the different scope and design of the surveys used here is not entirely identical, the numbers support angevaare’s perception: the number of institutions actively archiving digital material (at least 331) greatly exceeds the number of institutions with published preservation policies (at least 44). the nestor working group although a preservation policy is such an important part of an organization’s commitment to digital preservation, a certain reluctance to develop and adopt one is understandable. firstly, a transparent policy which can be accessed by users, partners and investors is a big commitment. secondly, it can be quite difficult to determine the level of detail and decide on length and scope of a preservation policy. to support the widespread development and adoption of digital preservation policies, in 2012 a nestor working group was formed to establish guidelines for the creation of a preservation policy for memory institutions such as archives, museums or libraries. its 12 members come from germany and switzerland. the working group is part of nestor (network of expertise in long-term storage and availability of digital resources), the german-language competence network for digital preservation founded in 2003. initially funded by the bmbf (federal ministry of education and research) in two phases (2003-2006 and 20062009), since 2009 nestor is acting as an independent, self-financing network. as of today, it has 16 members, mostly german memory institutions such as libraries, archives and museums (see figure 1). there are more institutions interested in becoming a nestor partner, so the number of partners is likely to grow further. currently, nestor consists of eight working groups for different important digital preservation tasks and topics, e.g. av-media, cost, rights and emulation. the network is also engaged in standardization work and has developed three national standards over the last three years, among others the catalogue of criteria for trustworthy digital archives (din 31644). since 2013, german digital archives have the possibility to receive the nestor seal for trustworthy digital archives which is based on these criteria (nestor, 2013a). in addition, creating guidelines and making international standards accessible to the german-speaking community belongs to the tasks of nestor and its working groups. e.g. a translation of the oais model into german was published in 2012 (nestor, 2013b). nestor also monitors the state of digital preservation and curation in germany and the german-speaking countries, and just recently a baseline study of the digital curation of research data in germany (also available in english, neuroth et al., 2013) was carried out. starting with a review of already existing preservation policies (the national archives, 2009; nlnz, 2011), the working group noted that policies vary considerably in length, depth, and detail. for example, the preservation policy of the national library of new zealand (2009) and archives new zealand (anz) includes parts that due to fast technical changes – would have to be updated quite often and in our opinion should rather be included in the preservation strategy. as no german guidelines on this topic exist, the working group reviewed existing english guidelines on policy development (the national archives, 2011). these were used to decide which parts are important for the german community as well. the resulting german-language guidelines, which will be published in the first quarter of 2014, consist of five main chapters: 1. goals of our guidelines 2. use of a policy 3. development of a policy (motive, responsibility, publication, relation to other related documents and strategic papers) 4. possible content of a policy 5. updates of a policy (policy watch) in the following, an overview of the most important findings of the preservation policy working group is given. it is these findings which form the basis for the content of the guidelines. . goals of our guidelines as already emphasized, a preservation policy is an important element in securing long-term-access to digital objects. digital preservation depends on technical as well as organizational issues, and a policy serves to address these. it demonstrates the figure 1 nestor partners iassist quarterly fall winter 2012 19 iassist quarterly commitment and responsibility of the archiving institution and adds to its trustworthiness. it is to be expected that it will be common practice in a few years for all digital archives to have a published preservation policy, and accordingly the pressure for every institution to get involved in this topic grows. against this background, the work of the nestor group aims to simplify the task of writing a preservation policy and to raise the awareness of the need of a publicly visible policy, especially for a german audience. our guidelines provide a tool box: the institutions using it decide themselves which parts will be relevant for their institutional policy and on this basis create a policy suitable for their needs. thus the guidelines aim to assist in the development of a policy, but they will not dictate any mandatory rules as the needs of the different institutions and digital archives are very heterogeneous. accordingly, the guidelines inform users about the impact, use, typical questions and difficulties of (creating) a policy. they help them to unmask their blind spots and increase awareness of dependencies and consequences. in addition, a generic policy example – abstracted from already existing policies – gives an idea to the users of the guidelines of what a policy might look like. purpose of a policy in general, the purpose of a policy is to show why – and, possibly, how an institution is involved in digital preservation and to define its benefits (the national archives, 2011). it demonstrates that an institution is part of the preservation community and is aware of important standards. more specifically, the purpose of a policy derives from its audience, or, simply: its users. the latter can be internal or external users. for internal users such as staff members, a policy forms an important basis for decisions. it can also serve to mitigate possible financial cuts: affected staff can point to the policy and insist at least on the budget needed for the minimum standards the institution has publicly committed to maintain. for external users, for example the “consumers” of the digital assets, a policy supports the building of trust as it creates security that the assets will remain accessible, citable, and usable for the long term. the same is true for the data producers, who have an interest that their findings serve future users as well. having a published preservation policy means that stakeholders will have transparent information about what the archive does to secure long-term access, as will potential clients who are considering outsourcing the digital preservation of their assets. thus, the institutional preservation policy is likely to be the basis for service level agreements between the archiving institution and any third party (beagrie et al., 2008). finally, a policy can be also useful – or even mandatory – for certification and audit processes and certainly will help acquiring third-party funds. developing and publishing a policy: the why, how, who and where there are multiple motives for starting to develop a policy. an obvious reason would be the beginning of digital preservation activities, but as the findings cited above show, this is rarely the case. in fact, one reason not to adopt a preservation policy early on in the process of building a digital preservation system is that this policy is likely to be revised frequently as the system and the experience grow and the workflows are implemented. for example, the marriot library of the university of utah in salt lake city, usa, revised its policy three times during the last three years. in contrast to the many institutions which conduct a digital archive without having developed a preservation policy yet, the marriot library published the first version of its preservation policy in 2010 – two years before they purchased the digital preservation system they are using today. they deliberately developed a policy so early because they felt it would help them to decide which preservation software to purchase once there would be something suitable available for them. in the case of the marriot library, writing the policy helped to shape the preservation program and to raise awareness about digital preservation plans and actions among the staff members. it is likely that there will be another revision once the digital archive is well established and fully implemented in the library workflows (keller, 2012). in contrast, the german national library (2013) and the bavarian state library (2012) had already been engaged in digital preservation for a number of years before they published their policy. in these cases, it was preferred to set up the policy after the digital archive was established and the full extent of the system was known. furthermore, technical or organizational changes within the institution could be the reason to start the development of a policy: an external evaluation of the institution, or – as mentioned before – an audit or a certification of the digital archive. depending on the organizational structure of the institution, a number of different staff members can be responsible for developing the policy content. possible scenarios are described in the nestor guidelines. in most cases, both members of the management and practitioners are likely to be involved. the development process and later adjustments of the policy will be time-consuming, especially if many staff members need to be involved. if possible, it is therefore recommended to keep the number of involved persons to the necessary minimum. where and how the policy is published is partly dependent on its scope. a policy might contain confidential matters and therefore will only be published within the respective institution. this might concern the whole policy or just certain chapters. the language used in the policy strongly depends on the target group. usually the national language is used and often an additional english translation for an international audience is created. generally, the language used has to be comprehensible for a wider audience and should avoid technical terms. additionally, the policy will most likely refer to other documents or strategic papers. it is recommended, for example, to address technical solutions not in the policy text but in other, related documents, as this content is likely to change very fast. as for the description of ingest workflows and preservation strategies like migration, these are better explained in the preservation strategy plan instead of in the policy (the national archives, 2011, p. 7). it is also highly important to ensure that the policy does not conflict with laws, rules or tasks of the institutions or already existing policies, for example the preservation policy for printed material. due to the relative novelty of the field, digital archives are often still in a development phase or in a very early stage of productive use. 20 iassist quarterly fall winter 2012 iassist quarterly therefore, the status quo of a given archive is often still not stable enough to frame certain principles. as mentioned above, this could be one of the reasons why many institutions seem hesitant to publish a (final) policy. in these cases it is possible to express the status quo of an archive or to create an “aspirational policy” (the national archives, 2011, p. 6), but both possibilities bear a risk of having to revise the policy fairly soon. policy content: the what the areas covered in a policy can vary a lot. analyzing 33 policies in the english language, sheldon (2013, p.6) observes that some are only one page long, whereas others consist of 30 pages or more. from the point of view of the nestor working group it is therefore an important task of our guidelines to give an overview of possible content of a policy and to emphasize the consequences that adding a particular content item will have for future work and the need to update the policy regularly. again, the guidelines refrain from prescribing too much because each institution will have very individual needs and there will be no “one size fits all” solution. the policy content is the main chapter of our guidelines as the possibilities are diverse and multifaceted. therefore, only a selection of possible aspects can be highlighted in this paper. in creating its guidelines for policy content, the working group took into account beagrie’s model of a preservation policy (2008; see table 1), the findings of sheldon’s analysis (see table 2), the practical experience of the members of our working group, and our own analysis of existing policies we consider to be a good example. the working group decided not to include all these criteria in its guidelines because from our point of view a compact policy with a manageable number of topics is easier to develop and to maintain. thus, some of the topics identified by sheldon (e.g. preservation planning, storage, duplication, and backup) might better be placed in a preservation strategy, which addresses more technical topics like preservation planning, storage, duplication and backup, and which will have to be revised more often. it is evident that the objective and the scope of the policy should be embedded in the general strategy of the institution and has to be compliant with its focus, priorities and tasks. it is important to define this objective in time and to address it within the policy. the goals of preservation, e.g. maintaining the usability, authenticity and integrity of the archived digital objects, can be a main part of the policy, as this is the heart of all preservation activities and of particular interest for the target group. a policy can also address how these preservation goals will be reached. furthermore, it is recommended to name the responsible units within the institution, those responsible for the archiving workflows, and the staff member or members responsible for the content and the updates of the policy itself. as there will most likely be some fluctuation in the staff, it is recommended to only point out staff functions rather than including names. from the perspective of the working group, other important topics for a policy are: • the organizational structure of the institution (including secure funding for the future) • mandate of the institution • legal and technical framework • principles of digital curation, e.g. maintaining integrity, authenticity and accessibility • protecting sensitive data from unauthorized access (e.g. medical research data). some points are not mandatory but could be useful depending on the scope of the policy: • purpose and scope of the archived digital material table 1: policy content suggested by beagrie (2008) 1 principle statement (benefits) 2 contextual links (relation to other strategies and documents) 3 preservation objectives 4 identification of content (scope of digital content) 5 procedural accountability (responsibilities) 6 guidance and implementation 7 glossary 8 version control (review of the policy) 1 access and use 2 accessioning and ingest 3 audit 4 bibliography 5 collaboration 6 content scope 7 glossary/terminology 8 mandates 9 metadata or documentation 10 policy/strategy review 11 rights and restriction management 12 preservation planning 13 rights and restriction management 14 roles and responsibilities 15 security management 16 selection/appraisal 17 staff training/education 18 storage, duplication, and backup 19 sustainability planning table  2:  common  policy  content  iden3fied  by  sheldon  (2013) iassist quarterly fall winter 2012 21 iassist quarterly • staff and other resources used for digital preservation (as this can also change over time, a rough estimate might be enough) • preservation activities (information in detail might lead to regular updates). the process and criteria for the selection of digital objects for the archive could also be part of the policy. furthermore, the access to different collections – if the institution is a light archive with user assess – can be an important part of the policy as well. again, however, as digital collections are growing, the policy would possibly have to be extended quite often. thus, if an institution does not want to update the policy regularly, it might be a good decision to deal with this issue in another, related document. the scope of the archive collection could be described in such a document and the description of newly acquired material could then be added to this document to avoid that the policy has to be edited too often. policy watch: updating and evaluation among the members of the working group opinions about whether or not a policy should be changed, and how this should happen, diverge. on the one hand, by revising its policy regularly, an institution can show that it actively watches technology and developments in digital preservation and keeps the preservation policy up to date. on the other hand, a preservation policy should be a commitment for the long term, something the institutional staff, stakeholders and clients can build and rely on. accordingly, if the decision to change the policy is made, it is a matter of trust to make the reasons for updates and changes transparent and to archive the older versions and keep them accessible, for example on the institution’s website. it is possible to indicate the next review date within the policy, as the national archives have done (the national archives, 2009, p. 10). of course, such a review might reveal that there is no need to change the policy. a policy update becomes necessary, if the policy no longer matches the daily work. for example, if an institution had a dark archive and adds an access component, the policy is likely to lack guiding principles for this. it will be necessary to extend the policy in order to cover access to the archived collections, and this will have to happen in a transparent and comprehensible way. if the policy includes detailed technical aspects, there will likely be a need to adjust it quite often to account for technical developments and changes of workflows. one possibility to deal with this issue is to state it in the policy and thus announce it from the beginning. an evaluation of the preservation policy can include the question whether the policy has met its goals. for example, it might be necessary to adjust existing workflows to the policy in this context. this is the best case scenario. an evaluation might also reveal that the reality cannot be adjusted to the policy and the policy has to be changed because certain procedures cannot be implemented. due to the lack of experience in this still relatively new field this possibility cannot be ruled out. again, it is recommended to create transparency in this case by giving comprehensive explanations about the changes. as the example of the marriot library mentioned above shows, a policy can be revised because the first draft of the policy has been published at a very early stage and therefore has to be updated more often and extensively as the implementation of the actual workflows take place. a look into the future the nestor preservation policy working group aims to publish its guidelines in the first quarter of 2014. subsequently, there will be a workshop for practitioners and possibly other follow-up activities. the guidelines will be available as an open access resource (in german; an english translation is currently not planned). the topic of preservation policies will also feature in nestor’s upcoming best practice wiki2. the wiki will supplement the publication of the guidelines and will provide a competence network for the discussion of practical questions and issues. it will also serve as a platform to address yet unresolved or even unknown aspects of drafting and maintaining preservation policies. for example, the majority of the institutional policies sheldon (2013) examined, address the issue of collaboration. currently, this is not part of our guidelines, although some of us participate in a digital preservation consortium. apparently, there are still blind spots to be detected by us and by others! references angevaare, i., 2011. policies for digital preservation. [pdf ] available at: <http://www.digitalpreservationsummit.de/presentations/ angevaare.pdf> [accessed10 december 2013] bavarian state library, 2012. sicherung des in digitaler form vorliegenden wissens für die zukunft – die langzeitarchivierungsstrategie der bayerischen staatsbibliothek. [pdf ] available at: <http:// www.babs-muenchen.de/content/dokumente/2012-11-22_bsb_ preservation_policy.pdf> [accessed10 december 2013] beagrie, n. et al., 2008. digital preservation policies study, part 1: final report october 2008. [pdf ] available at: <http://www.jisc.ac.uk/ media/documents/programmes/preservation/jiscpolicy_p1finalreport.pdf> [accessed10 december 2013] engelhardt, c., strathmann, s. and mccadden, k., 2011. report and analysis of the survey of training needs. [pdf ] available at: <www. digcur-education.org/eng/content/download/3322/45927/file/ report%20and%20analysis%20of%20the%20survey%20of%20 training%20needs.pdf> [accessed10 december 2013] german national library, 2013. langzeitarchivierungs-policy der deutschen nationalbibliothek. [pdf ] available at: <http://d-nb. info/103157140x/34> [accessed10 december 2013] keller, t., 2012. digital preservation policy. [online] available at: <http:// www.lib.utah.edu/collections/digital/digital-preservation.php> [accessed10 december 2013] international organization for standardization, 2012. space data and information transfer systems – audit and certification of trustworthy digital repositories. iso 16363:2012. washington d. c., usa nestor, 2013a. nestor seal for trustworthy digital archives. [online] available at: <http://www.langzeitarchivierung.de/subsites/nestor/ en/nestor-siegel/siegel_node.html> [accessed10 december 2013] nestor, 2013b. referenzmodell für ein offenes archiv-informationssystem. nestor-materialien 16. [pdf ] accessible at: <http:// files.d-nb.de/nestor/materialien/nestor_mat_16-2.pdf> [accessed10 december 2013] neuroth, h., et al., 2013. digital curation of research data. [pdf ] available at: <http://nestor.sub.uni-goettingen.de/bestandsaufnahme/digital_curation.pdf> [accessed 10 december 2013] nlnz, 2011. digital preservation strategy. [pdf ] available at: <http:// archives.govt.nz/sites/default/files/digital_preservation_strategy. pdf> [accessed10 december 2013] the national archives, 2009. preservation policy. [pdf ] available at: <http://www.nationalarchives.gov.uk/documents/tna-corporatepreservation-policy-2009-website-version.pdf> [accessed10 december 2013] 22 iassist quarterly fall winter 2012 iassist quarterly the national archives, 2011. digital preservation policies: guidance for archives. [pdf ] available at: <http://www.nationalarchives.gov.uk/ documents/information-management/digital-preservation-policiesguidance-draft-v4.2.pdf> [accessed10 december 2013] scape wiki, 2013. published preservation policies. [online] available at: <http://wiki.opf-labs.org/display/sp/published+preservation+polic ies> [accessed10 december 2013] sheldon, m., 2013. analysis of current digital preservation policies. archives, libraries and museums. [pdf ] available at: <http://www. digitalpreservation.gov/documents/analysis%20of%20current%20 digital%20preservation%20policies.pdf?loclr=blogsig> [accessed10 december 2013] notes 1. yvonne friese is affiliated at the leibniz information centre for economics in kiel. the main focus of her work is on digital preservation of research paper and digitized material and organizational issues around digital preservation. her contact email is y.friese@zbw. eu. 2. the wiki has already been established as a test and will be made available for public access in the future. vol30-2neu.indd by 12 iassist quarterly summer 2006 by josefi na j. card, tamara kuhn, and thomas wells* user-centered design and innovation in the sociometrics social science electronic data library (ssedl) abstract this paper presents the current state (scientifi c content, formats, platforms, distribution partners) of the sociometrics data archives, collectively known as ssedl, the social science electronic data library. it then peers into the future by describing areas of topical expansion, new target audiences, and new sciencebased resources currently being built around ssedl. usage information is also given. in the twenty years since the fi rst topically focused data archive was established at sociometrics (card 1989, card 1996, card 2000, carley and card, 2000), there has been a tremendous increase in the availability of inexpensive and powerful computing resources (davey et al., 2006), an expansion of federal requirements and incentives for data sharing (melichar, evans, and bachrach, 2002, nih, 2003), and a burgeoning of cost-effective data distribution options such as cd-rom and the internet. these factors have made use of data archives an increasingly attractive option for research and teaching, and have spurred the development of diverse collections of primary research data for conducting secondary research. the data archives at the inter-university consortium for political and social research (icpsr) at the university of michigan are the largest collection of social and behavioral research data. the data archives at sociometrics collectively known as ssedl, the social science electronic data library continue to be an attractive supplement, especially for users interested in public health issues and in use of data for novice researchers or for teaching purposes. the continual addition of new datasets to ssedl makes the resource a rich source of data for those in the public health, medical, nursing, social work, and social science professions. in this article we provide an overview of the current content of ssedl. we describe the features that make ssedl easy to use by novice and expert researchers alike. we end with a look into the future and share plans for upcoming content and user-focused innovations. organization and content the sociometrics social science electronic data library is a premier health and social science resource that is comprised of nine topically-focused data archives. each data archive has exemplary datasets selected by a distinguished scientist expert panel for their scientifi c merit, substantive utility, potential for secondary data analysis, and program or policy relevance. table 1 gives an overview of the contents of ssedl. details on each archive, including a complete list and description of included datasets, can be found at www. socio.com/dataarchives.htm. with some 600 datasets from more than 250 different studies comprising nine topicallyfocused collections, ssedl is a unique source of high quality health and social science data and documentation for researchers, educators, students, and policy analysts. more than eighty percent of the ssedl collection is unique and not available from any other public source (including icpsr). table 1: overview of sociometrics’ data archive collection topically-focused archive studies datasets variables adolescent pregnancy 162 286 80,000 aging 3 22 19,000 child well-being & poverty 12 36 20,000 complementary & alternative medicine 8 17 10,000 contextual 13 29 19,000 disability 19 40 25,000 family 20 122 66,000 hiv / aids / std 19 30 19,000 maternal drug abuse 7 13 5.,000 total 263 595 263,000 iassist quarterly summer 2006 13 user-focused features product packaging each dataset in the collection is made available with a standard set of eight machine-readable data and documentation fi les: (1) the raw data fi le, (2) spss program statements that defi ne each variable in the dataset and provide both variable and value labels, (3) sas program statements, (4) spss data dictionary, (5) spss frequencies, (6) an spss portable fi le, (7) a sas transport fi le, and (8) a user’s guide with standard sections: description of study, description of machine-readable fi les, complete list of variables sorted by their topic and type, frequencies for key variables included in most datasets (e.g., race, gender, marital status, etc.), and results of data completeness and consistency checks conducted by archive staff. this standard packaging and documenting of each dataset in ssedl assists users in familiarizing themselves with the resource. once a data analyst has worked with one ssedl dataset, it is easy for him or her to work with any of the others in the collection. search aids data users are able to identify datasets and variables that meet their needs and specifi c variables of interest via a search mechanism freely available on sociometrics’ web site (www.socio.com/search.htm). analysts can specify whether they want to search the entire ssedl collection, a combination of data archives, or a single data archive. the keyword search utilizes standard boolean search strings and searches key fi elds that include variable labels, value labels, study name, and investigator names. for each variable that the search returns, the display shows the variable label, the value labels, the names of the original investigators, and the study name with a link to additional study information (brief abstract, summary of methodology, number of variables, number of cases, and purchase options). product formats the format of data distribution has changed signifi cantly over the past twenty years, in keeping with technological advances in this period. initially, large datasets were made available on mainframe tape and smaller datasets were made available on diskette. now datasets and accompanying documentation are distributed in user’s choice of cd-rom or internet download. multi-level acquisition options purchasers can acquire data in one of three confi gurations: an individual dataset, a complete topical archive, or the complete ssedl collection (currently nine topical archives). individual datasets can be obtained on cd-rom or downloaded from the sociometrics web site. complete topical archives can be ordered on cdrom at a cost that is signifi cantly less than purchasing each of the datasets individually. the complete ssedl collection is available to universities and other institutions via subscription through thomson gale, the exclusive worldwide distributor of ssedl. all faculty, staff, and students at the subscribing institution are allowed free and unlimited access (via internet download) to all of the several hundred ssedl datasets and data-related materials. additionally, subscribers have immediate download access to new datasets as they become available. this dissemination format meets users’ need for quick access to the data, relieves the burden on data librarians of providing access to the data, and is extraordinarily cost effective. usage report during the past fi ve years, the explosive growth of the internet is refl ected in the increase in usage rates of sociometrics’ web site and data-related web features. as seen in table 2, during the past fi ve years the number of downloads of data-related products, including datasets, has increased by nearly 300% and the number of “hits” to data table 2. internet usage & purchase rates of sociometrics’ data archive collection, by year year 2001 2002 2003 2004 2005 visits to sociometrics website 136,598 145,237 175,984 182,528 253,636 hits to all data archive pages 42,003 70,022 86,770 101,819 210,494 downloads of data and data-related products 10,704 16,325 29,729 34,954 39,672 number of units ordered (non-subscribers 171 222 156 167 193 number of purchasers (non-subscribers 82 105 91 98 138 14 iassist quarterly summer 2006 archive-related web pages has increased by nearly 400%. in the past year alone, the number of hits to all data archive pages has more than doubled. the large increases in hits and visits appear to be a function of an increase in referrals from search engines, with the majority of new visitors being referred from google, yahoo, and aol. as more datasets are added to the archive collections and as internet use continues to increase we anticipate the rates of both web site visits and dataset usage to continue to increase in the coming years. looking toward the future: expanding data archives and new technologies we are continually expanding the content and capabilities of our data archives. as we move toward the future, userfocused innovation will be present in both content and the technology accompanying selection and use of the data. we are in the process of adding three new topically-focused data archives. the first is the data archive of longitudinal studies on childhood problem behaviors, which is being established with funding from the u.s. national institute of mental health. this archive will consist of an online and cd-rom collection of important longitudinal studies on childhood problem behaviors. the archive will include content-rich longitudinal studies that are not currently available for public use, as well as those studies in the public domain that have not received widespread use figure 1. topic and type distribution search matrix interface (partila view) iassist quarterly summer 2006 15 among the research community. the second new archive, the communication disorders data archive, is being established with funding from the u.s. national institute on deafness and other communication disorders. this archive will house state-of-the-art research datasets that address the prevalence and the social, behavioral, and occupational antecedents and consequences of hearing impairment and speech and language disorders. both archives have recently completed the dataset selection stage and each has an archived dataset. the objective of our newest archive, the welfare reform evaluation data archive, is to facilitate access to high quality welfare reform evaluation studies that will enable welfare policy research among a broad pool of scholars and researchers. updates to current archives in addition to creation of new topically-focused data archives, the existing archives in the social science electronic data library are continually being expanded through the addition of new datasets. the u.s. national institute of child health and human development is providing funds for the addition of datasets each year to the data archive on adolescent pregnancy and pregnancy prevention (daappp). recently archived daappp datasets include the national longitudinal study of figure 2. topic and type distribution search results 16 iassist quarterly summer 2006 adolescent health (add health), wave iii, 2001-2002, the public use education data; national survey of family growth, cycle 6, 2002; and the national longitudinal study of adolescent health, wave iii, 2001-2002 (add health). the other archives shown in table 1 are in the process of being updated and prospective datasets for each archive are currently being prepared for review by a scientist expert panel. at the conclusion of this cycle, each archive will have been updated with new datasets. an upcoming user-focused innovation: guided search through a data archive’s topical “areas of richness” a key element of assisting researchers in the use of secondary data is helping them identify the best datasets for their research questions and topics of interest. although users can currently perform web-based keyword searches on variables in each of sociometrics’ archives, it was determined that the search process could be made more productive if users were given a broader overall sense of the areas of topical areas of richness within each of the archives, and then were able to identify specific variables of interest within those topics. as a result we have begun development of a simple, cost-effective search interface to meet that need. figure 3. variable level search output iassist quarterly summer 2006 17 the search interface for each data archive displays matrices of variables by the topic and type distribution for that archive. figure 1 shows the prototype matrix for this new search capability. in figure 1 the “areas of richness” of the complementary and alternative medicine data archive can be seen from the cells with high numbers (many variables of the given topic and type). using a small javascript program that generates a help balloon when called by a mouse-over, each topic and type is clearly defined for the user by simply placing the mouse pointer over the topic and type heading in the matrix. the number of variables in the archive associated with any topic and any corresponding type are displayed in the matrix. each of the numbers in the matrix is a link that calls a pre-populated defined keyword query to the verity system requesting a search for variables containing only the topic and type corresponding to that box of the matrix. for example, clicking on the number 5 in the cell corresponding to the topic = meditation, yoga, and relaxation and type = behavior yields the referenced five “hit” variables. figure 2 gives the first screenful (four) of these variables. as seen in figure 2, the search returns a formatted list of variables matching the search topic and type keyword search criteria. each variable in the result list is displayed with the variable name, variable label, and an excerpt from the html page that corresponds to that variable’s information within the search index. clicking on a hit variable’s name and label then returns a web page displaying the variable name, variable label, value labels, the variable’s topic and type codes, and information about the dataset including the study title and the original investigators. for example, clicking on the first “hit” variable in figure 2, “past year relaxation techniques for pain,” results in information on the metadata associated with this variable (figure 3). finally, clicking on the study title returns complete information about the study and provides links to purchase or download the dataset, if desired. this user-focused search mimics the thinking of the analyst in searching for data that might address his or her topic of concern. first the analyst is advised in advance of the “areas of richness” of a data archive (figure 1). then s/he is systematically guided through the contents of the data archive, through the topics and types by which all of the several hundred thousand variables in ssedl have been indexed. finally, metadata about the variable and acquisition information about the dataset are provided. conclusion during the past two decades the evolution of sociometrics’ data archives has reflected current trends and changes in technology, in this manner expanding the definition and potential usage of data archives. during the next decades we plan to continue the expansion of the collections’ content and capabilities and continue the focus on data quality and user-focused design, search, and dissemination. * this article was presented at the iassist 2006 conference in ann arbor, michigan, at the session “innovations in data dissemination”. the authors josefina j. card, tamara kuhn, and thomas wells are all at sociometrics corporation. correspondence to: dr. josefina j. card, sociometrics corporation, 170 state street, suite 260, los altos, ca 94022, (650) 949-3282 x211, jjcard@socio.com. references card, j.j. (1989, january). facilitating data sharing. asa footnotes. card, j.j. (1996). development of the sociometrics data library on families, aging, substance abuse, and aids. social science computer review, 14, 305-309. card, j.j. (2000). development and dissemination of an electronic library of exemplary social science data. social science computer review, 18, 82-86. carley, m. and card, j.j. (2000). the social science electronic data library: serving the needs of data librarians and users. iassist quarterly, 24, 8-14. davey, m.e., matthews, c.m., moteff, j.d., morgan, d., schacht, w., smith, p.w., morrissey, w.a. (2006). federal research and development funding: fy2007. congressional research service report. library of congress. melichar, l., evans, j., & bachrach, c. (2002). data access and archiving: options for the demographic and behavioral sciences branch. national institute of child health and human development (nichd), august. national institutes of health. (2003). nih data sharing policy and implementation guidance. http://grants1.nih. gov/grants/policy/data_sharing/ the electronic triumvirate: the archives; the data processors; the new york state department of correctional services inmate files. a case study byhughw.shinn' new york state archives and records administration abstract: the new york state archives and recwds administration (sara) was one of the first state archives in the united states to accession electronic records into its holdings and niake them available to the public. sara worked with the new york state education department's (sed) electronic data processing division (edp) to obtain main frame computing services. sara's dependency on sed edp fchdata processing services required the development of a positive relationship with sed edp. this case study examines the relationship between a government archival institution working with electronic records and a centralized data processing unit that is completely unfamiliar with the operations and requirements of a data archive. there are two levels of a successful relationship: formal, which includes agreements on hardware, disk space, training, etc.; and informal, including the development of creative solutions to technical w procedural problems. these levels of interaction were necessary because sara (unlike other sed divisions) developed and executed its own applications rather than use the traditional edp services. sara the new york stale archives and records administration's (sara) program for the archival preservation of and research services for electronic records had its origins in 1988 with the release of the special media records project repcxl: strategic plan for managing and preserving electronic records in new york state government . this plan gave birth to sara's center for electronic records (cer) in 1990, which is to be the focal point of electronic records program development in sara. among its many charges, cer was the unit assigned to tning electronic records transferred from agencies to sara under archival control and to provide reference services for those data files. in this instance, archival control refers to variablelevel descriptions of the data set, explanations of the data set's technical aspects, the arrangement and description of the data set as a records series, and verification of the data and the documentation. in archival terminology, the process of bringing electronic (or paper) records under intellectual and physical control is termed accessiomng . archival administratiod the accessioning procedure for electronic (and paper records) is preceded by the records retention and disposition scheduling process in which agencies estabush minimum retention requirements for records and determine their final disposition: eith^ destruction or transfer to sara. most records are destroyed after they are no longer useful to the agency because they do not have enduring legal, administrative, evidential, ot research value. before electronic records are accessioned by sara, they are appraised for archival value. appraisal involves the examination of electronic records from both a content and a technical point of view. once the records have been determined to be of archival value, the data set's technical aspects are examined to determine whether sara has the skill and equipment to preserve the data. technical appraisal is based on discussions with agency personnel and careful examination of the technical and data documentation of the data set data processing before electronic records accessioning operations could begin in 1990, sara had to obtain computer and technical services from the state department of education (sed) division of electronic data processing (edp). sara initially developed an informal plan that specified cer's requirements and contributions to the joint venture of accessioning electronic records. requirements of the archives cer operates differently from other units in sed with respect to data processing requirements. most units work with a specified type of data and a specific set of data files. conversely, sara collects data from all executive branch agencies and must contend with a wide array of data files, types, and formats. the result is that diagnostic programs must be individually designed for specific data files. generally, sed edp customers have informational products such as reports and publications that must be produced. unfortunately, the user may lack the skills. winter 1992 time, or equipment to accomplish the unit's tasks. typically, at the customa-'s request, edp will carry out an examination of the unit's particular information requirements, conduct a needs assessment, and where appropriate, assist in the development of goals for the automated system. after developing these plans, sed edp designs, produces, and tests the programs that accompush the project's stated objectives. the data processing shop is also reqwnsible for system upgrades and major changes. the customer determines when the data system will run, edp determines how it will operate. this type of arrangement is practical for systems that are totally dependent upon edp for development and support sara contends with data from different systems that contain few (if any) common denominators, designed by personnel with varying skill levels. each data set requires unique levels of handung and effort to document the data fully. in order to provided this level of service, the sara analyst must become a hybrid of programmer and end user, successfully combining attributes from both worlds. it is this requirement for multi-faceted operation that sets sara apart form traditional sed edp customers. one aspect of commonality among the data sets that sara has already accessioned is that they were produced by automated systems that are now largely defunct. electronic record systems are suspended for many reasons including: loss of funding, migration to new equipment, and the completion of a temporary commission's task. when this is the case, the source code is generally missing, and the original programmers have long since departed. under these conditions, sara must become an investigative body to determine the scope and function of aging or superseded electronic recwds systems. these investigations seldom conform to predictable schedules such as annual reviews or a decennial census. rather, they may appear at inopportune moments in an astonishing array of formats. sara is currently working with agencies to replace this practice of sporadic submission of out-dated material with the mwe predictable method of scheduled transfers of current data. cer's role as an investigate and the necessity for multifaceted operations, produced a number of data processing requirements. the services required by cer for screening purposes are as follows: 1 . access to magnetic tape drives on a non-scheduled basis. 2. programs to check the physical condition of magnetic tapes for problems such as parity errors. 3. non-standard training for sed's unisys mainfirame software tools. (the goal is to develop a degree of indqiendence from the edp staff.) 4. training in the use of software packages (such as spss and insyte) that can be used to analyze agency data sets. 5. on-line access to data files. 6. ability to manipulate data files from the programmer's point of view. this includes variable-level operations involving the verification of data values. 7. the ability to carry out data verification procedures by comparing data values with existing documentation. (generally carried out with frequency distributions or dis^^ys of value ranges.) 8. the ability to isolate and examine specific columns and records within a data set from the operating system level. 9. production of copies and extracts of data sets, and a method for distributing the data to customers. armed with these requirements, cer met with representatives of sed edp in search of agreements and assistance. for the most part, the data processing staff were helpful and available for consultation. sed edp's operating procedures sed is an enormous and complex organization. it has an excess of 3,(xx) employees, many of whom use or maintain the department's fifty diverse mainframe applications or its state-wide data network. the size and complexity of the department required that sed edp develop standard operating procedures for most types of data and technical functions. these procedures are as follows: 1 . all requests for services will be directed to an edp staff member in the unit assigned to the customer's organization. 2. customers will not have direct contact with the technical suppwt staff. 3. customers that have programs requiring the use of magnetic tape drives will request that their edp contact run the program and submit a 'run sheet' to the edp operations staff 4. formal training above the level of technical manuals and vender programs is not available for assist quarteriy software packages or higher-level languages. most of these procedures reflect edp's customoservice orientation. in fact, even the lack of formal training relieves the customer of the tedium of writing code and deciphering arcane documentation. sed edp's view is that the customs should simply have to make a request, and the data processing ihx)fessionals will provide the user with the required product it is an efficient method fot managing requests from a large number of customers who require routine services. combined procedures: the first attempt to combine sara's data processing requirements with edp's existing procedures occurred when cer began to examine records for the department of correctional services (docs). the docs inmate under custody data files were the testing vehicle for the electronic records accessioning program at sara. exxs began collecting data on punch cards in 1956. the data were used to produce reports on the numbers and characteristics of inmates undercustodyfix)m 1956-1974. the information contains incarceration data including facility name and date received; crime and sentencing data; detailed demographic data including race, religion, nativity, occupation, and education; criminal history data. each annual file is a reflection of the prison population at a given point in time. these files are unique in that they contain relatively complete inmate-level demographic data on the population of new york state's correctional facilities ^. the initial foray into the unknown and the semi-known began when the first data files from the docs were examined. the technical infwmation did not coincide with the data sets, nor did the data documentation match the actual data values. it became apparent that additional tools were necessary if the proper file formats were to be discovered. the mystery file the mystery file appeared with the second shipment of data files from the department of correctional services. this data file refused to coincide with any of the printed documentation. it was difficult to determine the structure, size, record lengths and other important characteristics of this data set the tools and skills that could be used to augment sara's level of training and experience and bring the mystery file under archival control were distributed throughout sed edp's sections. while individual sections often cooperated with each other, staff in one section had only a limited need to understand the activities in the other sections. in this instance, the single contact system introduced an additional level of bureaucracy into the ditroi. the contact had to transmit sara's requests accurately to the ^propriate edp section, and accurately transmit that section's response. unfortunately, the mystery file stretched the single contact system beyond its limitations, and the file had to be abandoned the mystery file situation demonstrated that cer's data processing requirements were not routine and new procedures had to be de\e\aped by edp to assist cer in its attempt to accession electronic records. sara had the ability to perfotm much of the work normally provided by sed edp. with this in mind, sara requested access to the tools, software, and training that would lead to self-sufficiency. edp's response to sara's requests sed edp's response to sara's requests for services was to set up a liaison system where cer would contact a programmer in the applications section assigned to sara. the liaison would respond to sara's requests for services, or attempt to find someone who could. negotiations between sara and sed edp resulted in the modification of edp's procedures with respect to non-routine operations. 1 . on-line access to data files from a programmer's point of view was always available. 2. sed edp responded rapidly to requests for service. 3. sed edp willingly participated in the development of tools and procedures to solve unforeseen problems. 4. sed edp allowed limited access to the technical support staff. these modifications allowed cer to make use the old, inadequate documentation, examine the data, and determine the tine nature of the docs files. in addition, sara was able p-oduce more useful documentation for the docs inmate under custody files. attempts at the formal level to resolve the issues and problems raised by the mystery file situation were not entirely successful; and it became apparent that an additional and different type of relationship with sed edp was required. fortunately, strong working relaticxiships were developing between cer and several sed edp sections. this led to the development of informal relationships with specific programmers ft-om selected sections of the data processing shop. winter 1992 the basis of the infcmnal relationship was the realization that personal cooperation between specific sed edp programmers and the sara analyst would be a more effective than an organizational approach for solving unique problems. one of the mystery file issues was resolved when the programmers and the analyst devised a method for examining specific columns and records in a data set admittedly, this was one cer's formal data processing requirements. it became an informal item when it could only be operationalized by cooperation between the analyst and the sed edp programmers. the informal relationship was designed to augment the liaison system by developing contacts in other sections of edp. the liaison served as a facilitator, providing the sara analyst with the initial introduction to the appropriate sections. from that point, the analyst and the programmer in that section would work together as needed to resolve the problem at hand. the strength of the informal relationship is its flexibility. the combination of the informal and formal relationships formats results in the creation of a relationship network. this network is based on the premise that formal relationships are effective for repetitive qjerations; while informal relationships are more suited to solving unique problems. there are formal connections between sara and the edp liaison and between the edp liaison and edp's individual sections. through these connections flow procedural recommendations, written procedures, and formal requests for assistance. the formal relationship is particularly useful for contending with repetitive activities such as transferring data files. in addition, it provides a structure for communication between cer and sed edp. infwmal relationships also exist between sara and the edp liaison and among sed edp's various sections. additionally, infcnmal relationships exist between sara and the individual sed edp sections. the infonnal relationship is useful for the resolution of unique problems such as accessioning new types of electronic reccntls. in short, the infonnal relationship is more effective in resolving cutting-edge problems rather than completing routine tasks. informal relationships are legitimized by the existence of formal relationships, and cannot function effectively in the long-term without them. conclusion the flexible aspects of accessioning electronic records require that the archivists and the programmers have the ability to adjust to the variable demands of differing electronic records formats. flexible operations ^>ring from adjustable, non-stagnant relationships. the development of a relationship network between sara and sed edp provided the necessary structure for a flexible electronic records accessioning program. the relationship network combines the aspects the formal and the informal relationship formats. as a result, the electronic records program at sara is flexible, efficient, and has the capability to address technical and logistical problems (e.g. reading data dictionaries and transferring data files). the development of a strong relationship network has provided a stable foundation for cooperation between sara and sed edp. *** a diagram follows (see p. 12 of text) **** 1 presented at the lassist 92 conference held in madison, wisconsin, u.s.a. may 26 29, 1992. 2. new york state archives and records administration, bureau of records analysis and disposition. appraisal report # 87-28n; may, 1987; p. 3. relationship network formal inform/^l: ^ edp section sara r "k ' php ^ 4 t ' ^ ^liasony .^^ edp seclion 6 lassist quarteriy vol241 iassist quarterly winter 1999 19 abstract this paper describes the arl gis literacy project and its role in providing support for continued access to government data which is increasingly distributed only in digital form. in particular, it will address the university of missouri’s (mu) experience in the broad context of the arl gis literacy project goals as well as in comparison to the reported experiences of other participating institutions. it will discuss what mu has produced in terms of gis services and what has been learned about broadening awareness of gis. the mu experience will be examined as an example of the creation of support mechanisms for integration of gis into the digital library environment. introduction when geographic information systems (gis) moved beyond the domain of professional geographers and into “mainstream” technology in the early 1990s, libraries began seeking effective ways to utilize this powerful research tool. a particular focus has been on the ability of gis to deal with government data that is crucial to social science research (as well as many other disciplines) and which is increasingly distributed only in digital form. in the midst of this, the association of research libraries (arl) established the gis literacy project as a support mechanism for libraries interested in learning about and introducing gis into their services. as a graduate student in the university of missouri’s (mu) library and information science program, i became intrigued when i learned in a government information course that gis was being integrated into library services in order to provide access to digital spatial data. i was particularly interested when i learned that mu had been an early participant in the arl project and set out to find out more about the topic through a literature review of library gis services and the arl project, as well as interviews with staff members of mu’s ellis library regarding the institution’s experience with the project and the current state of the library’s gis services. this paper is thus a summary of my initial exploration into the world of library gis services. it provides an overview of the arl gis literacy project and addresses mu’s experience in the broad context of the project goals as well as in comparison to the reported experiences of other participating institutions. and it discusses what mu has produced in terms of gis services, what has been learned about broadening awareness of gis, and mu’s experiences with creating support mechanisms for integration of gis services into the digital library environment. arl gis literacy project the arl gis literacy project was initiated in 1992 as a multi-phased project in partnership with esri (the leading producer of gis software) and other public and private partners. the goals of the project are designed to meet the current needs of libraries and users while addressing the changes libraries are undergoing as they enter the 21st century, and to provide the tools and expertise necessary to insure that digital government information can be used effectively and remain in the public domain. these goals include: • introduction of gis to a variety of libraries (e.g., public, state-based, academic, and university libraries in public and private institutions) to address diverse user information needs; • development of a team of gis professionals in the research library community willing to lend time and expertise to applications, user training, and education programs; • encouragement of connections among federal, state, and local gis users and information; • promotion of research, education, and the public right-to-know through improved access to government information; • initiation of library projects to explore new applications of spatially referenced data and evaluate the introduction of these services in research libraries; and • implementation of programs to allow institutions that have invested in networking capabilities to leverage the sharing of resources via networks. the project seeks to provide a forum for libraries to experiment and engage in gis activities by introducing, educating, and equipping librarians with the skills needed the arl gis literacy project: support for government data services in the digital library by mary french* 20 iassist quarterly spring 2000 to provide access to digital spatial data. in cooperation with gis vendors and foundations, arl organizes training sessions for project participants, sponsors an electronic mail list, and works with government agencies on gis programs and related issues. financial support, data, software, hardware, and expertise to assist in the project goals have been provided by gis vendors and foundations. the project is still active, although the focus has shifted more towards enhancement of programs at institutions with gis services now in place rather than the earlier focus on introduction of these services. occasional training sessions are still provided, as are other types of support for institutions wishing to develop gis services. participant experiences in 1997, arl conducted a survey to determine how, in the years since the project began, participants have organized their delivery of gis (davie et al, 1999). the survey addressed four main categories of gis service: 1) general information about the library’s role in delivering gis, 2) the number, level and academic preparation of other training of staff involved, 3) the amount and kind of equipment, software, and data files that support gis in the library, and 4) the kind of service offered and by whom it is used. seventy-two of the 121 project participants responded to the survey. the following summarizes the survey results: general information • 89% of the responding institutions reported that they provide gis services. • gis services were administered by the library at 83% of the institutions and by academic departments offering gis courses at 70% of the institutions (both the library and academic departments administer gis services at many institutions). • only 5% of the libraries reported having discrete gis units; most library gis services were found to be located in the government documents center (48%) or map center (52%). staffing • at 81% of the libraries with gis services, the services were directed by a librarian with an mls; 54% of those librarians held at least one additional graduate degree. • the most common gis training for respondents was through arl’s gis literacy project; others had received training through gis software providers or gis coursework. infrastructure • 78% of the libraries with gis services utilized esri’s arcview software. • 58% operated their gis on windows95/np platforms, 56% on windows 3.1; the remainder operated on dos, unix, and macintosh platforms. • 61% utilize computer networks for their gis services. • the government printing office depository program provided digital data files used for gis services in 83% of the libraries; 67% supplement those files through purchases (70% of those libraries had funding of less than $2000 for such purchases). service • 53% of the responding libraries offered gis support service 20 hours a week or less, 24% offered more, and three institutions offered no support at all. • the typical number of users of library gis services was about seven per week; students comprised about half of those users, while faculty, staff, businesses, local government, and the general public fairly equally comprised the remainder of the users. also in 1997, arl published transforming libraries: issues and innovations in geographic information systems. this publication presents a number of case reports that provide an overview of experiences and lessons learned in attempting to develop support mechanisms for gis services in libraries. included is a set of questions for library planners to answer in designing or rethinking gis-based services: key questions for planners • what kind of service should we provide? • how will collections be built? • who will staff the gis-based services? • how will we learn – and educate others – about gis? • with whom will we collaborate? • how and where will we store data? • what will it cost? these questions point to the critical themes that emerged from the case reports – themes such as gis service planning, partnering, policy development, staff training and expertise, resource allocation, and user support. following is a summary of three case reports presented in transforming libraries. these cases are selected for discussion because they provide good examples of ways in which institutions have successfully addressed particular issues in their efforts to provide support for gis services. specifically, the university of georgia is noted for its planning efforts, penn state for its extensive partnerships, and north carolina state university for addressing issues of staff training and expertise. planning – university of georgia libraries careful planning proved crucial in the development of gis services at the university of georgia libraries. in 1994, a comprehensive survey commissioned by university administration and issued by a campus-wide committee identified current and future campus gis instructional and research activities, gis software needs, and a host of potential gis services. the results of this survey, which iassist quarterly spring 2000 21 are available on arl’s transforming libraries gis website (http://www.arl.org/transform/gis/) along with the survey developed by the university of georgia, allowed the university’s map library to create a small gis lab and design focused and responsive services which provide patron access to the library’s digital spatial data as well as to spatial data available on the internet. partnerships – the pennsylvania state university libraries penn state university libraries began to plan for gis services in 1995 with a brief analysis of existing resources that revealed a lack of appropriate coordination of gis data and its use. a new mission was drafted to create gis services which would provide for acquisition, organization, and archiving spatially referenced data and make it available to the widest number of users through in-house facilities and the internet. partnership development and/or enhancement of existing relationships with the campus’ geography department, cartography lab, computing center, and a semi-independent research group proved crucial in fulfilling this mission. these partnerships resulted in the creation of an in-house gis center at the university libraries with trained, continued student staffing provided by an internship program with the geography department and hardware and software support provided by the campus computing center. additionally, partnering with the campus cartography lab and a semi-independent research group with gis expertise resulted in the development of a web interface which distributes pennsylvania-based spatial information both within and beyond campus via the internet. training and expertise – north carolina state university libraries in the early 1990s, north carolina state university organized a small team, consisting of librarians and a computer center staff member in liaison with a faculty member with gis expertise, to initiate the libraries’ gis start-up effort. the team quickly identified staff and user training and expertise as key challenges in providing support for gis-based services. to tackle the problem of a staff with little gis expertise, members of the team who acquired training began to offer sessions in basic gis skills to other staff to provide them with an understanding of the scope of the services and enable them to assist in user support. the gis team also developed introductory gis workshops and classes for campus faculty, staff, and students which would enable users to work more independently with the data and software. additionally, a position of librarian for spatial and numeric data services was created to provide a professional staff member with the appropriate experience and proficiencies to carry the responsibility for development and management of the libraries’ gis and other spatial and numeric data resources and services. providing in-house training and expertise for gis services was an important component in the development of significant and successful gis services at north carolina state. the mu experience project participation the university of missouri-columbia was an early participant in the arl gis literacy project, assigning two professional staff members to participate in the program in addition to their other duties. these staff members attended training sessions and the library’s data services center received a computer and software to begin experimentation in developing gis services. staff members found it to be an extremely time-consuming service to offer and time and staffing restraints have prevented the development of these services. the data services center has received only a small number of requests that utilized the gis tools, and has not yet been able to allocate the staff time and other resources to maintain the requisite hardware and software, develop the policies and expertise necessary to support gis services inhouse, and enable them to publicize these tools to potential users. although development of gis services remains in the library’s long-term plans, it is not currently experiencing the demand that would establish development of these services as a high priority. the staff members assigned to the arl project remain aware of gis activities, but are not currently actively participating in the project. access to spatial data the university of missouri-columbia is a federal depository library, thus its government documents center receives digital spatial data that it must provide public access to. the government documents center has a computer that meets the minimum government standards for running gis programs, but has only standard printing capabilities and limited staff expertise and time for supporting the users who wish to utilize and manipulate the data. the center currently only receives about six requests a semester for gis-based data and these are primarily from experienced users to whom the materials can be loaned or who require minimal in-house support. the government documents staff finds they are least able to support the casual user who requires more intensive support to work with spatial data. the staff currently plans to attempt to increase awareness of these data resources in the user community, an effort which hasn’t taken high priority in the past due to support limitations. the government documents center also has a loose agreement with the campus’ geographic resources center (a unit of the department of geography) to assist with gis-related needs that can’t be met by the library. impact of project participation in my introduction, i note that mu’s experience in the arl project would be examined in this paper as an example of the creation of support mechanisms for integration of gis 22 iassist quarterly spring 2000 into the digital library environment. although mu clearly has not yet been able to achieve the well-developed services reported in some of the case reports, a level of support for gis services has resulted from participation in the project. gis services have been introduced into the library, a level of staff experience in using the software and equipment has been achieved, there has been a broadening of awareness of the data and services that can be provided, and government distributed spatial data is accessible. in fact, looking at the status of gis services at mu in comparison to the findings of arl’s 1997 survey suggests mu’s experience may be typical of that of many other participating institutions. the following compares mu’s status with the survey results. (note: mu did not participate in the 1997 survey). general information • mu is providing gis services, as are 89% of the responding institutions. • like many of the responding institutions, gis services are provided through both an academic department and the library. • library gis services are currently offered primarily through the government documents center, like those at 48% of the responding institutions. staffing • mu’s library gis services are supported by librarians with an mls, as are those at 81% of the responding institutions. • like the majority of the responding institutions, library gis training has been provided primarily through the arl project. infrastructure • mu’s library gis services utilize esri’s arcview software, as do 78% of the responding institutions. • mu’s library gis services are operated on a windows platform, as are the majority of responding institutions. • the gpo is the primary source of mu’s digital spatial data files, as is the case at 83% of the responding institutions. service • mu offers less than 20 hours per week of library gis support service, as do many of the responding institutions • mu currently experiences a very low demand for library gis services – only a few requests per semester, as opposed to the seven per week which the survey found typical. future possibilities the achievements mentioned above provide the groundwork on which mu can continue to build their gis services, particularly if they utilize the experiences and suggestions offered by other project participants such as the three case previously summarized. utilizing the university of georgia’s survey might assist in shaping gis service planning efforts. development of in-house workshops and expertise, like the north carolina state example, could assist in addressing support issues. additionally, mu’s existing interdisciplinary resources appear to hold great potential for exploration of partnerships like those that have been implemented to support library gis services at penn state. mu’s geographic resources center, a multidisciplinary applied research and teaching facility for geographic and remote sensing data analysis, already provides some support in meeting gis-related requests received by the library. the missouri spatial data information service, which provides gis and census data about missouri via the internet, is run in close association with the geographic resources center and is another rich resource for creating service and resource partnerships. the mu integrated spatial analysis of environmental systems mission enhancement proposal, sponsored by the school of natural resources, department of geography, and college of engineering, has received administrative funding and support. this proposal seeks to focus mu’s efforts in the geographic information sciences and to enable participation in a global network of research and outreach in the analysis of geographic information integrated across traditional disciplinary boundaries. although this mission enhancement area does not specifically include development of library gis services, library staff do participate in meetings of the mission enhancement area group and the proposal and its support provide groundwork which the library can utilize in its own proposals to obtain support for gis services. conclusion there is a substantial subset of library literature which focuses on gis services, much of which consists of accounts of institutions who have participated in the arl gis literacy project. in addition, there are email lists (e.g., gis4lib@u.washington.edu ), websites (e.g., www.mcmaster.ca/library/maps/gis_libr.htm ), and conferences (e.g, esri’s international conference on gis in education and libraries) which have focused on this topic. having reviewed information from a variety of these resources, i’ve drawn the following conclusions: • libraries seem to achieve great success in developing their gis services when they focus on the particular issues or areas that work best within their larger institutional context for creating support mechanisms for their services. these issues/areas include planning, partnering, policy development, staff training and expertise, resource allocation, and user support, all of which are essential to creation of support mechanisms but, as the case studies above iassist quarterly spring 2000 23 have illustrated, can have varying roles in developing support mechanisms. • the arl project has offered important assistance in developing the support mechanisms which have enabled the creation of successful gis services at many of the participating institutions; however, there are many institutions that have yet to be heard from. only 72 of the project’s 121 participants responded to the 1997 survey and the literature tends to focus on those institutions who have been able to more fully develop their services. gathering a wider range of case reports from those institutions that may still be struggling to implement their gis services will be essential in continuing to find way s to further support mechanisms for these services. • a variety of limitations have thus far prevented mu from achieving the same level of growth in their gis services as some other institutions; however, participation in the project has provided a forum for the library to experiment and engage in gis activities and broaden awareness of the potential this tool may hold for the library’s long-term plans to provide access to and support for digital data resources. references “arl gis literacy project” [ www.arl.org/info/gis/ index.html ]. davie, d. kevin, james fox, and barbara pierce. the arl geographic information systems literacy project. (washington: association of research libraries, spec kit 238, march 1999). deckelbaum, david. “gis in libraries: an overview of concepts and concerns.” issues in science and technology librarianship, no. 21, winter 1999. [ www.library.ucsb.edu/istl/99-winter/article3.html ] gis for libraries listserv [ gis4lib@u.washington.edu ] information technology and libraries, vol. 14, no. 2, june 1995. special issue: making gis a part of library service. journal of academic librarianship, vol. 23, no. 6, november 1997. special issue: geographic information systems (giss) revisited: a symposium. herold, philip. “maps and legends: plotting a course for geographic information systems,” american library association, association of college and research libraries, acrl 1997 national conference papers, partnerships and competition session a07, saturday april 11, 1997. [www.ala.org/acrl/paperhtm/a07.html] mount, jack d., robert change, and patricia j. morris. “planning and developing a gis program in a large academic library.” (1998 esri international user conference proceedings) [www.esri.com/library/userconf/ proc98/proceed/to600/pap551/p551.htm] soete, george j. transforming libraries 2: issues and innovations in geographic information systems. (washington: association of research libraries, spec kit 219, february 1997). [www.arl.org/transform/gis/ gistrans.html] university of missouri-columbia, office of the provost. “mission enhancement – global information (mogaia) funded proposals. global access to the information age: integrated spatial analysis of environmental systems” [web.missouri.edu/~provost/memogaia_abstracts.html] * paper presented at iassist 2000 (chicago, 7-10 june 2000). mary french school of information science & learning technologies university of missouri-columbia email: mef884@mizzou.edu iassvol201 12 iassist quarterly keywords: internet, training, information skills, world wide web information. technology. networks. training. four words which are often thrown together and thrown around quite carelessly. our view is that in whatever combination they occur, technology, to quote janis joplin, always seems to come out on the top. our view is that the present and the future demand that we take training, and training in information skills in particular, very seriously. only then can we deal with the challenges which the advances in networked technology presents. only then can we begin to help others to come to grips with new possibilities and use existing skills and experience to avoid new versions of old mistakes. nicky ferguson writes: let’s start where everyone should start these days on the web. often people have noticeboards in their kitchens i find these quite compulsive reading. they are covered with the essential detritus of individual, or family, life. business cards from the plumber and the piano teacher; appointment cards from the dentist, the clinic and the acupuncturist; receipts from the washing machine repair woman and the milkman; shopping lists, opening hours, bus timetables, parking tickets. sometimes i learn something from these unauthorised perusals (“gosh the ante-natal clinic! congratulations”) but mostly the charm resides in glimpsing the ordinariness, the minutiae of other people’s lives. nosiness, not to put too fine a point on it. i’m not sure that trawling through most web pages is very different. a superficial fascination but often no new information, no new ideas. this is actually fine we don’t advise people to keep their noticeboards covered with a black cloth in case someone else wastes their precious time reading personal trivia. it is up to me to discipline myself not to spend all day browsing the appointment cards. we should be explaining this to children, students, trainees, users call them what you will. it is not the web’s fault that people find it useful for collecting and collating their personal and work trivia and signposts, and sharing that information with their friends. it is up to the browsers, and here i mean the people not the software, to recognise quality when they see it and when they don’t, and to develop strategies to make their work time more productive: and of course it is up to the professional providers of quality information to point people to worthwhile sources and to run quality services. so what are these quality services? how will our students recognise them? equally important how will they recognise those that aren’t worth spending their time on? moreover, since the line between information consumers and information providers will become increasingly blurred, how can we encourage them to make their information available in a neighbourly, useful, socially responsible, creative, even fun way? let’s look at some sites and see what we think of them. let’s say i’m interested in goldfish someone’s told me that i should look at “sharon’s home page”, ok, let’s explore. h’mmit seems sharon likes goldfish too but there’s not really much more about goldfish than that. there seems to be a jane austen archive and there are other things which may or may not interest us, but no real material about fish; still there is a list of other places to look for fishy stuff including dave’s page “really interesting” it says so we’ll go and look there. dave’s page is strong on hyperbole but again it lacks content. still it does have a list of fish sources (which look similar to the ones on sharon’s page, now you come to mention it). there seems to be a heavy metal music archive and there are other things which may or may not interest us, but no real material about fish; still there is a link to the “sofa home page” (i know you’ve all heard of the small orange fish association) so we’ll go and have a look there. well this isn’t quite what i expected this seems to be the home page for my kinswoman finlay ferguson, rabid scot and fish fan, as well as chairwoman of sofa. it does have a list of fish sources (which look similar to the ones on sharon’s page, now you come to mention it). there seems to be a clans and tartans archive and there are other things which may or may not interest us, but no real material about fish; still there is a link which reads “goldfish lovers click here “ now that sounds exactly what we’re after, so we’ll go and have a look there. oh dear... “sharon’s home page” ... back where we started. one could argue that the act of classification itself (“some fish pages i have gathered together”) adds value to the web; but in fact a classification which only points at further pointers obfuscates rather than clarifies. which is all a roundabout way of saying that many web pages seem to be a part of a self-referential, charmed (but not charming) circle, training in the age of digital abundance technology or information? by nicky ferguson and lesly huxley1 , university of bristol 13summer 1996 merely pointing to each other without adding much to the sum of human knowledge or even networked information. of course a classification which also describes, not a mere listing but a descriptive record, meta-data in the jargon, is a different matter. depending of course on its own provenance and reliability, such meta-data does add value to the web and to the resources it describes and it can be amply justified. care should also be taken to point either to resources themselves or, in some cases, to further, fuller or more specialised descriptive lists, but not to bare lists of titles, whether or not they are long or have hyperbolic introductions. what implications does this have for us when we attempt to train others to construct worthwhile web-based resources? what should the resources do? 1. value they should add value in some way, probably providing descriptions of the resources they point to, preferably doing more. 2. classification they should systematically categorise and take advantage of the uniquely “virtual” nature of their medium to crossclassify, so that users can get used to knowing where to look and can find things where they expect to not where we think they should. 3. maintenance resources should be continuously maintained to give currency in addition to reliability; networked information changes so fast that unmaintained lists go off quicker than milk on the doorstep. 4. quality they should not just uncritically dump everything that might be relevant into a huge list. quality judgements should be made and continue to be made so that resources which are set up in a burst of enthusiasm and left to wither and become irrelevant are spotted and deleted. 5. variety of access users should be offered a variety of access methods or interfaces they should have the choice of searching or browsing and preferably have different browsing options. and what are the implications for the searchers, the users (and those who seek to teach them)? there will probably be no “one way” of doing things and no one source or megastore which will satisfy all your information needs. you may expect to call at three or more locations and use different techniques before arriving at your goal. you should take care that you are narrowing the field along the way, not skipping from one unordered list to another. reward the productive search paths by taking a few moments to retrace your steps and add bookmarks, penalise the not-so-charming circles by making a mental note to avoid them in the future. so the tao of webbing will be that there are many paths to enlightenment, you will browse and then search, search and then browse. in searching, how to search? in my experience the information strategies of the average user are limited. many if not most of the postgraduate students i come into contact with have managed to get good degrees without knowing what the three letter word “and” means when it is used as a boolean operator. they will search for “marx and engels” (or often “marx and spencer” but here is not the place to consider the quality of secondary education, spelling or the commodification of culture) and expect, very reasonably if no-one has told them otherwise, that they will find every resource which mentions “marx” and also everything containing “engels”. yet many sites providing search facilities on the web will offer far more sophisticated options which go largely unnoticed or unused. there are two approaches to this problem. the first is the one that the technologists adopt. what we need to do, they tell us, is to make the search engines, the software, the facilities, so clever that users don’t need to know about search strategies, they can just type in their queries in natural language. i’m a big fan of natural language searching, at least i will be when i find a system that works, but i doubt if most users have given sufficient thought to what they are seeking, even to phrase their quest in natural language. perhaps i will be convinced by someone here today who has designed a natural language search engine with artificial intelligence, inbuilt dictionary/thesaurus and pre-search feedback mechanisms so that when a user enters on the search form the word “aids”, before searching the database or sending the robot off to examine the web, our intelligent engine will ask the user “do you mean handy gadgets for disabled people to allow them to operate machinery, pick things up, hear better and things like that, or do you mean the disease or do you mean home helps or do you mean something else i haven’t mentioned here?”. even in that unlikely event i will still maintain that a sophisticated approach to searching will encourage sophisticated thought and that surely sophisticated thought is needed for sophisticated analysis. of course i do not think we should discourage the development of excellent search mechanisms i am in fact involved in a project in which we spend a lot of time discussing exactly what such a mechanism should and should not do for the user but alongside the technological development, we should be encouraging and promoting amongst so-called “ordinary users” an understanding of strategies for finding, retrieving and using information. which brings us to training. you will have guessed by now that i think training should encompass more than the latest technological buzz, more than which buttons to press in version x of software y which will be replaced by software z in a few years, months or weeks. we are after some more general appreciation of ways to use these technologies for real work, even real life. i used to think that it was important to make internet training a totally pleasant and stress free 14 iassist quarterly experience, newcomers tend to bring quite enough stress and anxiety with them when attending a course on such a daunting and overhyped subject. i now think that it is somewhat mischievous and misleading to prepare and engineer such things to gloss over the difficulties and make everything too smooth and easy. training, when related to preparing yourself for other activities such as running a marathon, swimming the channel or even a walking holiday, implies a certain amount of effort, dedication, commitment and practise. you make yourself, or your trainer makes you, do unpleasant things, push yourself, stretch, extend your capacity you will be subjected to nauseating exhortations such as “no pain, no gain”. perhaps we should be taking the “make ‘em sweat” approach a bit more in this area too. of course it is necessary to present beginners with step by step practical exercises outlining every key press, what are known as hand-holding exercises. but if that is all we do, however impressive our evaluation sheets at the end of the day, we are not giving them the confidence to go further on their own. better to suffer a few adverse comments at the end of the training day but produce trainees who, when the hand is taken away will wobble off on their own bicycles and disappear round the corner, not collapse in a heap. so let’s make our poor trainees answer questions, don’t just let them follow the instructions, or wander off on their own. we should, of course, provide reference materials and the equivalent of reading lists, citations catalogues and bookshelves, but as many university lecturers have found, it’s not always best to dole out photocopies as it can encourage the belief common amongst students that the act of clipping a photocopy into a ring binder osmotically transfers information and comprehension of it to the brain. often better to force the unwilling student to search for and actually read the article before deciding whether it is worth copying and archiving in empty cornflake boxes. similarly with exploring the networks next time they will be able to cope better if this time they had to work it out from a sketch map rather than being led by the nose. in summary, i would like to beg, plead and cajole you, as information professionals to share your skills with the horde naive users like myself who are blundering and about to blunder into this huge global virtual library that is creating itself. i am asking you to consider, going out of your institutions and talking to the people who are and will be using the internet. get involved with training initiatives and patiently explain that people have thought about information issues before netscape was installed on their pc. go to the places where they are beginning to use this stuff the school classrooms, the undergraduate internet clubs, the cyber-cafes and worse (yes it’s a dirty job but someone’s got to do it). spread the word about information handling skills, information seeking skills and user-friendly information provision. don’t let the code-writers monopolise the new image of international networked information they will reinvent the wheel if you let them and it will be triangular (but with retractable spokes and flashing lights). what’s more, as the provision and use of networked information explodes, the technology will change at least every couple of years. but information skills will become more relevant, more important, more marketable. be there, or be triangular. lesly huxley writes: the task, when i was appointed sosig documentation and training officer in mid-1995 was twofold: to produce sosig promotional, publicity and reference materials drawing attention to a service which had done ‘some of the hard work’ in searching for quality and relevant social science resources; most academics who tried the web when it first emerged from their university computer experts’ clutches found it a significant time-waster and severely wanting. we wanted to bring them back into the fold, bringing the newcomers with them, to show them that there were ways of locating useful networked information quickly. secondly i was to provide internet workshops at uk universities and colleges of higher education, supported by training materials tailored for social scientists and the particular needs of the site concerned. the target audience comprised mainly newcomers to networking in the social science field and those tasked with training and supporting them. the workshops and materials were not to be set entirely in the ‘press that button’ mould although newcomers would need some precise instruction, the aim was to provide a forum for learning both the tools and techniques and an attitude of enquiry which would allow them to cope with and extract the best and most relevant information for their work not only on the day of the workshop but well into the net future a future with little discernible shape. one difficulty was in reconciling participants’ time constraints and potential technophobia (or netphobia) with the ever-changing, ever-challenging internet environment to which i was trying to introduce them. another was to satisfy the needs of on-site trainers and support staff for materials which could be adopted, adapted and cascaded to others beyond the dozen or so attending the workshops each time. the route from task specification to task completion (not that it will ever really be complete that’s not the way of the networks nor the way of training!) was a challenge to ideas about teaching (training) and learning. initially there was a period of stock-taking, using paper and on-line materials developed during sosig’s formative years. an on-line welcome page was loaded in a browser and bookmarked before participants entered the room and greeted them when they arrived. they were invited to browse through it to gain some www and netscape background, something which seemed to challenge their ideas of what a workshop should be: they expected to be welcomed, introduced, led gently into the topic with a talk or perhaps a demonstration, not allowed to explore on their own. some had dabbled before, thought they knew a lot and were expecting something a lot more sophisticated. others felt lost: how could they get on and experiment when no-one 15summer 1996 had taught them what to do? many simply sat and stared at the text on screen and then, as time went by, at the screen saver. others browsed through the comforting pieces of paper they had been given, hoping for guidance from there, but still reluctant to put fingers to keyboard or mouse. the know-it-alls clicked off into a bravado show of hypertext highjumps which did little to reassure their colleagues. a few of the newcomers caught on and read, understood and followed links, started exploration. the aim had been to avoid giving them sosig “on a plate” but to provide a menu of ingredients with which they could experiment to find the most appropriate mix to prepare themselves for future learning, future exploration, with some guidance, some structure. in the main the challenge presented by this slightly unconventional learning experience failed them. instead of prompting reflection, questioning, experimentation, it engendered resentment amongst the knowing, misunderstandings and misconceptions amongst the beginners: some thought the welcome page was sosig, some were unaware of what they were using a browserto view this information: was it a word processor? a text reader? how had it appeared on their screens? worse still, their ability to come to grips with paper exercises later on and their confidence to proceed further were seriously affected. the unconventional start which should have set a positive note of enquiry for the rest of the workshop instead proved a barrier to learning. the pattern of the workshops and the materials evolved gradually for a time as i tried to improve them with small adjustments, but eventually changed dramatically to take on a more conventional look which could still incorporate prompts for reflection, areas of challenge. a very traditional start of welcome speech, presentation with slides about the internet and demonstration of how to load the browser and the on-line tutorial (a several page and quite complex development on the one-page welcome page) now leads participants gently into the recipe, but the emphasis is placed early on on self-paced learning, exploration, challenge within a structured framework. the on-line ‘slides’ provide more flexibility than their powerpoint predecessors:the levels of experience of each audience varies enormously, from one university to another, department to department and within the same workshop. the hypertext links in the slides provide for a longer, detailed ‘talk’ for newcomers, with the ability to bypass the basics and/or offer off-the-cuff demonstrations for a more experienced set. the presentation can be expanded in almost any direction dictated by the experience and interests of the audience. within the traditional framework of talk and presentations and step-by-step exercises implying apprenticeship, acceptance of information learnt from the expert, lie semi-socratic interventions which stop short of destroying all former knowledge in order to clear the way for future learning: at all stages participants are questioned and challenged, either to think further about what they are doing, to seek information and provide an ‘answer’ and prompted to develop their own questions, their own enquiry. once the expected presentation is over, participants load a browser, enter the url for the on-line tutorial and start exploring. some are hesitant and follow slavishly what is on screen, but for most, the flexibility of the tutorial structure encourages them to set their own agenda, their own pattern for learning. after these initial explorations via the tutorial and a talk about and demonstration of sosig, participants are finally given the comforting pieces of paper many of them still crave. there is generally a collective sigh of relief at this stage: they have had a taste of freedom but they do not yet feel ready for total liberation. the pilot step-by-step exercises have been developed into a workbook and are interspersed with questions demanding reflection and further questioning in turn and to try to interrupt slavish adherence to instructions without understanding. quizzes are provided to reinforce and test the newcomers’ learning and to engage the more experienced. in both cases they are designed to illustrate how sosig and the resources it points to can be used to support teaching and research, to prompt lateral thinking, provoke consideration of searching and browsing strategies. emphasis throughout is on consideration and development of the latter, on the different tools and techniques available, the different thinking that may be required depending on the design and content of the resource. this is followed through in further exercises involving other uk national services and international www search engines such as alta vista, excite etc. participants are encouraged throughout to use their own search terms or subject areas for browsing rather than sticking rigidly to the examples in the exercises. after the initial, traditional ‘presentation’ introducing the workshop and enough background to get them going, the rest of the day and these are full-day sessions is given over largely to participants’ explorations. there is no requirement to use all of the workbook or to follow the sections in any particular order. most are delighted to be able to set their own pace, to follow up their own lines of enquiry within the framework the workbook provides. during the second two thirds of the workshop my role is to respond to questions, interrupt occasionally with comments and demonstrations on issues which arise and offer collective or individual guidance if asked. requirements and the materials to meet them are constantly changing. materials are frequently updated with screen shots and instructions, urls and comment. i have to address participants’ increasing levels of net experience arising from extensions to campus networks and the increasing availability of graphical browsers on academics’ office machines. few sites now want coverage of using telnet to access www resources via lynx, more want an introduction to html authoring. web-based evaluation forms and questions and comments during workshops provide useful feedback in tailoring workshops and materials. information is also collected at the workshops on participants’ previous usage of the internet and world wide web as part of the evaluation of 16 iassist quarterly the effectiveness and usefulness of the gateway and subjectbased services in general, as well as of the training. followup questionnaires and, in some cases, telephone interviews, seek to provide comparative data for usage after the workshops, to be analysed by a consultant employed under a related project. the most common comment on the workshop forms has been the usefulness of ‘protected time’ to explore, of not being forced through at a particular pace and of being allowed to follow up own lines of enquiry. early on in the workshop participants are exposed to ways of recording and saving information found on the web, through bookmarks, saving and copying and, in some cases, using electronic mail. fewer now reach for the pen to write down urls of useful sources they have found. many copy bookmarks to disks to take away as a new starting point. if participants have not started following links or searching for information themselves by the middle of the second session i feel that the workshop has failed. the most successful sessions from my own and from participants’ points of view passed on through evaluation forms are those where questions come thick and fast, where the text and graphics appearing on screens as i roam around the room are those i have never seen before, where participants call colleagues’ attention to resources they have found, sometimes scampering excitedly around the room like children. their enthusiasm for exploration, the discovery that amongst the abundance of networked information sources there are some which could prove really useful, that there are ways of finding and handling information which are not trivial or a waste of time, is a great joy. even more so when they begin to think out loud, follow lines of thought on how they might incorporate some of the resources in their teaching, how they might introduce students to them. from then on the sosig has jumped off the plate, each participant leaves with a handful (or mindful) of ingredients which they can fashion into their own individual recipes for locating, using and perhaps in the future building networked information resources. 1. this paper has been presented at the css96/iassist conference at minneapolis, university of minnesota, may 12 19, 1996. nicky ferguson and lesly huxley, centre for computing in the social sciences, university of bristol 8, woodland rd bristol bs8 1tn e-mail:nicky.ferguson@bristol.ac.uk ccss home page at: http://sosig.ac.uk/ccss/ sosig at: http://sosig.ac.uk/ 34 iassist quarterly spring summer 2009 controlled vocabularies for ddi 3: enhancing machine-actionability introduction controlled vocabularies (cvs) are organized lists of terms used for metadata and information retrieval. they may be short (flat) lists or more complex constructs, containing hierarchical relationships. subject thesauri, such as lcsh (library of congress subject headings) or the multilingual elsst thesaurus used by european data archives, are examples of the more complex constructs, containing broader terms, narrower terms, related terms, synonyms, and scope notes. ideally, the terms in a controlled vocabulary should be exhaustive (covering the whole dimension of the issue), mutually exclusive (no overlaps between terms) and clearly defined (definitions/scope notes given for the meanings of the terms). cvs are often used in specific contexts, and definitions/scope notes clarify and disambiguate the meaning of a term in a particular context as it may differ from the meaning in natural language. controlled vocabularies play a critical role in metadata standards, which have two basic components: 1) semantics – definition of the meaning of metadata elements, and 2) content – declaration of instructions for what and how values should be assigned to elements (chan and zeng, 2006). controlled vocabularies belong to the domain of content as they specify the values allowed in an element or attribute. an extensive set of controlled vocabularies is now being developed for the data documentation initiative (ddi) metadata standard, to be used to describe specific aspects of a dataset across the data life cycle. this paper discusses the advantages of and reasons for using controlled vocabularies; the history of the effort to create controlled vocabularies for ddi; the work of the ddi controlled vocabularies group (cvg) in developing new vocabularies; and a new standards-based system for managing cvs in ddi, as well as other future developments. advantages of controlled vocabularies using controlled vocabularies has several advantages for ddi, many of them related to overcoming difficulties caused by natural language in documentation and information retrieval. control of synonyms. in natural language, there are synonyms that use different terms to refer to the same entity. in a controlled vocabulary, synonyms are no longer a source of concern because the vocabulary defines the preferred term (chu, 2007). control of lexical anomalies. cvs control lexical anomalies by minimizing any superfluous vocabulary or grammatical variations that could potentially create noise in the users’ results set (chamis, 1991; garshol, 2004), e.g., removing leading articles, prepositions, conjunctions, etc., or ensuring consistency (macgregor & mcculloch, 2006). for example, when describing data collection methods, a ddi vocabulary will tell us whether to use ‘self-completed questionnaire’ or ‘self-administered questionnaire’. even in the case of simple issues such as describing the (same) planned frequency of data collection, there may be surprisingly many variations: ‘twice every year’, ‘biannually’, ‘every 6 months’, ‘every six months’, ‘two times a year’. add to that all the possible variations in different languages, and it becomes clear how difficult it often is to produce comparability and how vocabularies can enhance semantic interoperability between organizations and systems. promotion of consistency and efficiency. controlled vocabularies enhance consistency and high-quality metadata not only by providing a single form of the term to be used but also by promoting more consistent use of ddi elements themselves. a cv is a clear indication of the intended content of an element. we must also factor in improved efficiency: persons providing metadata tend to change over time and vocabularies lessen the burden of learning. vocabularies also make metadata production quicker – less time will be spent on trying to figure out the meaning of elements and how this or that entity should be described. not surprisingly, new staff members tend to be fond of vocabularies. clearly defined terminology. definitions/scope notes for terms provided in the ddi vocabularies are another way by taina jääskeläinen, meinhard moschner and joachim wackerow1 iassist quarterly spring summer 2009 35 of improving consistency, comparability and efficiency. the definitions explicate the meaning of a term in the given context, clarifying the difference between, say, a proxy and an informant as a response unit. the controlled vocabularies group members have found that meanings of terms are rarely so clear they seem at first glance. our experience was that more often than not, discussion of a particular vocabulary had to be postponed until we had time to find term definitions from methodology handbooks or other relevant sources in order to know what we were talking about. therefore, we fully expect that the definitions provided for the terms in ddi vocabularies will be an advantage both to metadata production and information retrieval. the controlled vocabularies group also learned during the process that institution-level vocabularies currently used in their own organizations leave a lot to be desired. we found, quite often, that they contain overlapping terms, lack some necessary terms and are too geared toward describing the collections of a particular archive. this makes us confident that the vocabularies suggested for ddi 3 elements will be an improvement to many institution-level vocabularies. promotion of interoperability. it is becoming generally accepted in the information community that interoperability is one of the most important principles in metadata implementation. using the same controlled vocabularies for metadata in different collections enables cross-collection searching. if different vocabularies are used, interoperability can be provided by mappings. interoperability at the repository level with harvested or integrated records from varying sources can be enabled by mapping value strings associated with particular elements (chan and zeng, 2006). if a data organization feels obliged to continue to use an institutional-level and context-specific vocabulary instead of the one recommended by the ddi standard, the organization should provide a terminology mapping to the ddi vocabulary. terminology mappings are intellectually created crosswalks from the terms in one vocabulary to the terms in another, providing a network of equivalent, broader, narrower and related term relationships (mayr and petras, 2008). support for machine-actionability. the structure of cvs also facilitates the use of codes or notation which can then be associated with terms. such notation is mnemonic, predictable, and language-independent (broughton, 2004). in fact, one of the most important advantages of controlled vocabularies is that they facilitate the production figure 1. filtered search enabled by controlled vocabulary 36 iassist quarterly spring summer 2009 of metadata that are not only machine-readable, but also machine-interpretable (ddi 2) and machine-actionable (ddi 3). cvs do not usually replace machine-readable textual descriptions used to give more in-depth information, but since ddi controlled vocabulary terms have codes, they provide precise and unambiguous means for controlling software processes, as in figure 1 below. in this example, a search application looks for all studies conducted by cati. the user sees the human-readable text, which is the definition of the code interview.telephone.cati from the controlled vocabulary modeofcollection, and then retrieves a list of studies to browse. if statistical measures, for instance, are to be extracted and compared in an automated way, resulting in a time series chart or a geographical representation, it has to be guaranteed that identical statistical measures are being compared. in addition, one needs to control for universe, analysis unit, etc. if the matching of datasets is managed by an application, the terms used to describe dataset formats or character sets have to come from a controlled vocabulary. another example of a cv with the potential to control software processes is the iso 3166 alpha-2 country code provided by ddi 3 for identifying countries. the country of a data provider or user may be relevant for access conditions, and an application can use the country codes to control access. similarly, the iso 639-1 alpha-2 standard coding for the most common languages also facilitates machine-actionability. if questionnaires are marked up using the language cv, their selective and separate indexing by retrieval tools can be controlled by software. software can also determine which language versions of metadata are available for which elements in a database (or ddi instances) in order to control text retrieval processes and to offer transparency to end users in multi-language applications. crossing language barriers in documentation and information retrieval. controlled vocabularies form an important class of language tools. they can be used to assist both in manual and automatic translation of metadata (svenonius, 2003). if an editing tool has implemented controlled vocabularies, the tool may be designed to produce automatic translation of some metadata elements from one language to another. this, of course, only applies to the elements or attributes that have cvs. controlled vocabularies can be used to support behind-thescenes query expansion across languages. in addition to element-specified search, they can also be used as tools in free-text search. in fact, multilingual subject thesauri are useful for crossing language barriers even in cases where they have not been used for indexing data. if implemented as behind-the-scene search tools, they enable the user to discover data in different languages when the query term he/she has used is both 1) a thesaurus term and 2) appears in the text of an abstract or in question wording. more precision and recall in information retrieval. controlled vocabularies presuppose less previous knowledge about the content of a resource or a repository, or even a virtual collection of repositories. they may help to bridge the initial gap between the user and the resources he or she needs by focusing the request. an information seeker is more likely to achieve high recall (fraction of the relevant documents that are successfully retrieved) if all entities of the same kind are named in the same way. the seeker would achieve lower precision (fraction of the retrieved documents that are relevant) if terms have multiple meanings in different contexts. there is even more “added value” if the combination of allowed values in relevant elements provides a pre-selection mechanism for potentially comparable results, for example, data resulting from measuring the same concept at a certain point in time and space, for a comparable universe and using a certain methodology. the challenge is to build a system for “bringing like things together and differentiating among them” (svenonius, 2000). search tools can display thesauri and other controlled vocabularies to allow users to improve their query formulation. data portals may display broader, narrower and related terms of the term the user has used in his/her search. this will enable users to see which concepts/terms are related to their areas of interest and maybe give them ideas of other potentially relevant query terms to use. ddi controlled vocabularies project the implementation of controlled vocabularies has been an ongoing topic of discussion across the history of the ddi standard. early versions of the standard, ddi 1 and 2, were expressed in the form of a document type definition (dtd) that contained some controlled vocabularies within it. this was problematic, however, because to change the embedded cvs, the entire specification had to be reissued – clearly not an ideal situation. when ddi 3 was released in 2008 as an xml schema, it was clear that a review of controlled vocabularies was in order, both in terms of content and their relation to the xml schema. accordingly, a controlled vocabularies working group (cvg) was established by the ddi alliance in late 2007 to develop vocabularies for ddi 3 elements and attributes. the technical implementation committee (tic) provided a list of elements and attributes for which vocabularies might be considered. the multilingual and multicultural controlled vocabularies group has members from several different countries. the group was initially chaired by ken miller from the uk data archive, and after his retirement in mid-2009, by taina jääskeläinen from the finnish social science data archive. the group has been meeting via videoconferences approximately every two weeks iassist quarterly spring summer 2009 37 and expects to publish the vocabularies on the ddi alliance web site within the next few months. controlled vocabularies created for ddi at the moment, there already are a number of controlled vocabularies embedded in the ddi 3 schema, including valuetypecodetype (e.g., greater than, less than, equal to, etc.). some elements use well-established external controlled vocabularies. for example, the countrycodetype element uses iso country codes. the cvg has developed the following controlled vocabularies which correspond with related elements or attributes in ddi 3.1; the usage of these controlled vocabularies will be enabled with the next version in the ddi 3 development line. • lifecycleevent • commonality • intendedfrequency • timemethod • modeofdatacollection • responseunit (for survey type data) • aggregationmethods • datatype • softwarepackage • characterset • categorystatistic • summarystatistic • datecalendar • analysisunit • contributorrole • publisherrole • kindofdata (referring to the kind of data disseminated) the ddi alliance will recommend the usage of the cvs for ddi 3, and the vocabularies will be published on the ddi alliance web site. each vocabulary will have its own version number. each entry in a vocabulary has a code and a corresponding term and definition in english (see example in table 1). terms and definitions in other languages can be added as required. it is possible to add region-specific language versions for terms and definitions. while this can make sense in some cases, in general one language version should suffice. this avoids confusion caused by multiple terms in the same language. some of the proposed vocabularies developed are hierarchical, containing broader and narrower terms. the narrower terms may not cover the whole dimension of the broader term, and users are advised to use the broader term if none of the narrower terms is suitable. all vocabularies have an unspecified ‘other’ term, unless there is a clear reason for not including it. users are advised to specify in the documentation what they mean by ‘other’. if the element is of the codevaluetype, it has an ‘othervalue’ attribute that can be used to enter the specific meaning. cessda, for example, is considering capturing the othervalue information in order for the cvg to determine whether there are additional terms that should be added to specific vocabularies. the work on ddi 3 controlled vocabularies is still in progress. the group has revised the original draft vocabularies after receiving comments from the cessda data archives and other data providers. the biggest challenge the cvg has encountered in its work is that, because ddi 3 is as yet not widely used by data organizations, there are few experts on the standard. the group has consulted tic on several occasions. we expect that when the standard becomes more widely used, there may be suggestions for changes to some vocabularies, or requests for vocabularies for new elements/attributes. the qualitative data exchange working group activities may eventually bring additional term suggestions. at the moment, while the cvg has done its best to take qualitative data into account, the vocabularies are somewhat more geared to quantitative data. code term definition median median (mdn) the score value below which (and above which) half of the scores in a distribution fall (50th percentile). validcases valid cases cases with observations considered to be valid, i.e., providing substantial information and to be included for calculation. invalidcases invalid cases cases which are considered and defined as “missing” (e.g., not ascertained, not applicable, etc.) to be excluded from calculation. minimum minimum the lowest valid score in a variable. table 1: extract of the controlled vocabulary for summary statistics 38 iassist quarterly spring summer 2009 if a vocabulary has been created for a ddi 3 element/ attribute that has a corresponding element/attribute in ddi 2, the vocabulary can also be used for ddi 2. this approach is backward compatible. the documentation of the forthcoming new version in the ddi 2 development line will include information on the use of these cvs. the alliance has also published ddi best practices for controlled vocabularies . genericode the ddi cvs will be published in the genericode format, separately from the ddi xml schemas. genericode defines a standard format for defining code lists, also known as enumerations or controlled vocabularies. genericode aims to provide a standard model and xml representation for the contents of a code list. this is an oasis genericode committee specification. the genericode format has a tabular model for code lists. the “rows” are individual entries in a code list, where an entry is a set of one or more codes, plus other metadata, that is associated with a single conceptual entry in the code list. the “columns” are individual (typed) pieces of metadata that can be applied to each entry in a code list. so columns define what kind of data can be in the code list, while rows define what actual data are in the code list (oasis code list representation requirements, 2007). an advantage of using a controlled set of semantic concepts is in localization where the associated documentation for the coded values can include descriptions in different languages, thus not requiring the coded values themselves to be translated, or where translation is desired, the semantic equivalence of values can be described (oasis code list representation tc charter). maintenance and management of ddi controlled vocabularies for the vocabularies to remain up-to-date and viable, they need to be maintained. the cvg will function as the management team, reviewing any suggestions for changes, monitoring the types of terms that have been used for ‘other’ and their documentation, keeping track of different language versions and liaising with the bodies/ persons responsible for them. updated vocabularies will be published with new version numbers. flexible approach for specific needs ddi controlled vocabularies can be customized to meet local requirements or even replaced by institution-specific vocabularies, if needed. however, any extension or change of terms puts interoperability at risk, particularly if data are to be published or harvested in cross-institutional data portals. if extensions or revisions to ddi cvs are made locally, mappings from the more detailed local vocabulary version to the ddi cv are recommended. usage of ddi controlled vocabularies for other applications while the ddi controlled vocabularies have been developed for usage with ddi 3, they can be used by other applications as well. the ddi controlled vocabularies are a separate product of the ddi alliance, published independently of ddi xml schema. cessda plans cessda is planning to establish a european research infrastructure for the social sciences in 2011. all members in this coalition will use a common metadata standard for data documentation. the standard will include mandatory or recommended use of controlled vocabularies in certain ddi elements/attributes, and therefore the adopted vocabularies will be translated into the local language(s) of the member organisations. using the controlled vocabularies will help to facilitate the eventual transformation of ddi 2 documentation to ddi 3. references broughton, v. (2004). essential classification. london: facet publishing. chamis, a.y. (1991). vocabulary control and search strategies in online searching. westport conn.: greenwood publishing group [ref. macgregor & mcculloch, 2006] chan, l.m., and zeng, m.l. (2006). metadata interoperability and standardization – a study of methodology part i, d-lib magazine, vol. 12, no 6. chu, h. (2007). information representation and retrieval in the digital age. american society for information science and technology. asist monograph series. garshol, l.m. (2004). metadata? thesauri? taxonomies? topic maps! making sense of it all. journal of information science, vol. 30, no. 4, 378-391. mayr, p., and petras, v. (2008). building a terminology network for search: the komohe project. paper at the 2008 international conference on dublin core and metadata applications. also available online from http:// dcpapers.dublincore.org/ojs/pubs/article/viewarticle/931. macgregor, g. and mcculloch, e. (2006). 'collaborative tagging as a knowledge organisation and resource discovery tool', library review, vol. 55, no.5, 291–300. oasis code list representation requirements (2007). version 1.0.1, p. 6. retrieved 7 may 2010 from http://www. oasis-open.org/committees/download.php/23844/oasiscode-list-representation-requirements-1.0.1.pdf. oasis code list representation tc charter. retrieved iassist quarterly spring summer 2009 39 7 may 2010 from http://www.oasis-open.org/committees/ codelist/charter.php. svenonius, e. ( 2000). the intellectual foundation of information organization. cambridge, mass.: mit press. svenonius, e. (2003). design of controlled vocabularies. encyclopedia of library and information science, 2nd ed. zeng, m.l., and chan, l.m. (2006). metadata interoperability and standardization – a study of methodology part ii, d-lib magazine, vol. 12, no 6. notes 1. taina jääskeläinen, finnish social science data archive, finland: email taina.jaaskelainen@uta.fi. meinhard moschner, gesis leibniz institute for the social sciences, germany: e-mail meinhard.moschner@gesis. org. joachim wackerow, gesis leibniz institute for the social sciences, germany: e-mail joachim.wackerow@ gesis.org. 2. the alpha-2 standard can be extended to alpha-3 or supplemented by region or script subtags where necessary. 3. ddi best practices documents are available at: http:// www.ddialliance.org/resources/publications/working/ bestpractices 4. the organization for the advancement of structured information standards (oasis) is a global consortium that drives the development, convergence and adoption of e-business and web service standards, web site http://www. oasis-open.org/. 5. the council of european social science data archives, web site http://www.cessda.org/ 34 iassist quarterly educating the data user: online information retrieval in a rapidly changing environment in which cunent accurate information is a key element in staying up-to-date and maintaining an edge in an ever-more-competitive marketplace, access to on-line databases and information services is definitely worthwhile. baruch college library has been offering access to online information since 1981. currently, we subscribe to about twelve online services, which provide access to bibliographic, textual, and numeric databases. in the beginning the service was available only through intermediaries (i.e. trained expert searchers), but as demand has increased, direct access by end users (i.e. those who actually need the information, but who are, in general, inexperienced searchers), has been provided. by ida lowe' baruch college, city university of new york 'presented at the international association for social science information service and technologv (iassist) conference held in washington, d.c., may 26-29, 1988 obstacles to accessing online databases the major obstacle to accessing online databases is that information online is costly. in general, online services charge for every second of time that one is connected, and retrieving information can be a very expensive endeavor, even when the search is performed by an expert, and more so when performed by an end user. some services offer special academic rates. for example, dow jones news/retrieval charges academic institutions a monthly fiat fee for unlimited access to most of their databases. this service offers comprehensive company and industry information, as well as general business information. in addition, major services such as dialog . brs , and orbit offer special training as well as 'after hours' agreements which provide access to most of their databases at gready discounted rates. baruch college library has taken advantage of all these special rates in order to make online access available to end summer 1988 iassist quarterly 35 the second major obstacle is that getting the information is not easy. the nature of the majority of databases makes it difficult to extract information from them. most online databases were originally created to be accessed directly not by the end user of the information, but by intermediaries working on his or her behalf those intermediaries were information professionals, usually librarians, familiar with the particular jargon and structure that apply to information-searching in the library field. as a result, the interface to most online databases was structured for their needs. not surprisingly, a lot more work and a great deal of confusion occurs when end users try to get at such information on their own. for most online databases, the situation comes down to the end user needing to learn the information professional's language and methods if he or she wants direct access to information. users who want or need to conduct their own searches for information have to invest some time and eftort in learning how to use these systems. in the process, they often have to deal with almost as many information retrieval formats as there are databases. most online services strongly recommend training for new users. for example. dialog already has a two-day training session for beginners, and halfto full-day sessions for special categories of databases, such as business and economics. however, for the end user who will need to use the system only once in a while, this training may be too much, and will generally be wasted for lack of practice. there is a movement among online services to make at least some of their information more accessible to users who are not information professionals. dialog information already has two services that provide simplified access: the business connection allows users to easily access detailed information on thousands of corporations and businesses, and the medical cormection provides menu-driven access to information on clinical medicine and medical research. easvnet. a service of telebase systems inc. of narberth, pa., attempts to overcome the potential confusion inherent in trying to interface with so many difterent databases by providing a single user interface that quizzes the user in order to narrow down and define the information he or she is seeking. easvnet then uses that information to determine which databases it should access and how to conduct the search. easvnet serves as a central access point to approximately 900 database and information services. there are hidden disadvantages of which users may not be aware, including restrictions on the nimiber of databases or type of information that can be accessed. also, recent developments of more sophisticated, user friendly interfaces are making end user searching easier. prosearch , from personal bibliographic software, allows the user to access any database on dialog or brs without specific knowledge about the databases or the systems. by using a scheme of menus and submenus, the information seeker can set up a search statement and verify its correcmess before logging on to the system. prosearch provides descripuons of each database structure. all the user has to do is highlight the database fields and type the necessary keywords. prosearch uses this information to construct the search statement users who want access to on-line information must take the lime to learn the best way to get at the information, or at least make sure they have access to someone who can. a final problem users may face is that, even if they are able to easily access the type of information they need, it may not be in a form that is especially useful to them. the more useful services make information available in a form that can easily be incorporated into dbase iii, lotus 1-2-3 or similar formats. summer 1988 36 iassist quarterly the development of databases on cd-rom will contribute substantially to the solution of all three problems enumerated above, i.e. there are no connect charges, the access software can be simple and effective, and the user can have a choice of formats of the information retrieved. one database producer that has successfully taken advamtage of this medium is disclosure, a database containing financial information on public companies. the end user can search disclosure effectively without any knowledge of boolean logic, truncation, field delimiters, etc., and download the information into lotus 1-2-3, or ascii for use with spssx or sas. taking all these obstacles into consideration, a plan to promote the use of information online was developed at baruch college. the first factor to be considered was the audience we wanted to reach: * undergraduate students, the majority of whom are business majors, * graduate students, business and education majors (we subdivide this group further into masters degree candidates, doctoral studems.and research assistants), and * faculty. ai present online information retrieval services are completely subsidized by the college. searches done by professional librarians (intermediaries) are not available to undergraduates. in order to serve our patrons, we developed several approaches: 1. at the undergraduate level, online information retrieval is introduced in the basic library course, which is a three-credit course, part of the core curriculum for all students at the college. a more advance course on online information retrieval is offered for juniors and seniors. the students are introduced to basic concepts such as boolean logic, truncation, proximity operators, field delimiters, and basic search strategy preparation. the systems introduced are dow jones news/retrieval, a menu-driven system, and dialog , a command-driven system. the goal of the course is to provide the student with enough of the basic concepts, so as to enable him or her to learn a new system without too much trouble. they learn that the basic concepts apply to all systems, only the mechanics change. a similar course is offered at the graduate level in the masters in educational technology program. the command-driven system introduced in the latter course is brs, because this is the prevalent system in elementary and secondary schools. online information retrieval workshops are offered to graduate students and faculty. these workshops last four hours, and the participants learn to search dialog and brs using prosearch . anyone who has taken this workshop is then allowed to do their own searching in a special information lab. the lab has five ibm xt's with 1200 baud modems. each computer has prosearch set up with the appropriate passwords, and keeps track of all searches. a trained research assistant supervises the lab and provides basic assistance. if the searcher needs more help, a professional librarian can be consulted. in addition to the general application workshop described above, special subject-oriented workshops are offered to thesis students and research assistants. for example, accounting and finance students can take a workshop on accessing i. p. sharp (a service which contains over 40 million time series of primarily economic and financial data), learn to download time series data to a diskette, and upload them to the city university of new york central computer for use with such statistical summer 1988 iassist quarterly — 37 packages as sas or spss.x. 4. doctoral students are required to take a research methods course in which they are taught to use prosearch for onhne information retrieval, and are expected to use it throughout the semester for all their projects. 5. in order to promote the use of those systems which are simple to use, and require no special training, such as dow jones news/retrieval and databases on cd-rom, we hold biweekly demonstrations for anyone interested. the emphasis in the workshops and demonstrations is on 'hands on' experience. we make sure all participants carry out a search. 6. finally, we distribute to all facultypromotional literature on new products that we acquire or to which we subscribe. baruch college has been oftering these programs for several years. the form and content have changed to keep up with new developments, but the purpose remains the same, i.e. to make ever.' member of the baruch college community "online information literate." i feel quite confident that we have been successful .n summer 1988 vol18172 7spring/summer 1994 background archivists have come somewhat belatedly to the idea that there should be formal standards for description and for the exchange of data about their materials. observing progress made in these fields in north america, british archivists began work on constructing the necessary instruments in 1984. the archival description project was set up at liverpool university, supported by funds from the british library research and development department and the society of archivists. the project team has produced two successive texts of a manual of archival description, affectionately known as mad. the second edition, mad2, published in 1990, was published by gower, and has received a reasonable degree of trialing2 the archival community in britain, however, finds itself in a difficulty as regards the formal adoption of a standard. there is a national council on archives, and a working party of this, chaired by dr. kitching (who is secretary to the royal commission on historical manuscripts), has recommended the adoption of mad2. in a rather similar way, the society of archivists has issued signals of approval, and has asked its professional methodology panel to carry out tests and development work. these measures are somewhat short of a formal endorsement, but they do indicate acceptance at a practical level, and show that there is a will to continue developing the work. the second edition of mad contains rules for the description of a number of special formats, commonly found amongst archives. these are: title deeds (legal documents transferring land) letters and correspondence photographs cartographic archives architectural and engineering plans sound archives film and video archives machine-readable archives this section of mad2 must still be regarded as experimental, and it has not yet received adequate trialing. the principes on which the rules and guidelines are based, are coherent over the whole body of mad2 and will be discussed later in this paper. the mad2 special formats are intended for use in general archves repositories and services, not in specialist institutions. this important restriction should be emphasised. the second set of archival description standards which should be mentioned are the international ones. the lnternational congress on archives held in montreal in september 1992, received the text of two new standards: 1. statement of principles regarding archival description (the madrid principles). since this had been debated by the profession since 1991, this text was adopted. 2. general international standard archival description (isad(g)). this was received as a draft for dissemination and discussion. the first of these texts is now will be available. the second, (isad(g)), is not immediately available as it is in course of publication. it is intended that there should still be discussion of the topics presented, so that the process of maintenance and development may proceed. isad(g) itself has indeed not yet received formal adoption, but since it is in its second draft, and has received considerable discussion all over the world, it must be regarded as being near completion. the other main standard applicable to archives, which ought to be mentioned here concerns data exchange. this is the marc format, an archival application which was developed in the usa in 1984. it has become widely used in north america to allow archival descriptions to appear in the bibliographic databases, rlin and oclc. these databases are not widely available in britain as yet, and the resistance of archivists to bringing in a library standard has been such that up to now marc has been virtually unused for archives. there has indeed been little opportunity for it to be used. this situation appears to be changing, and a version of ukmarc in the archival format (amc) is due to appear in 19933 turning now to the management and use of machinereadable archives, few british archivists have yet had managing machine-readable archives: progress with description and exchange standards. by michael cook 1 archival description project, university of liverpool 8 iassist quarterly much experience. the esrc data archive and the edinburgh data library have been almost alone in the field in this country. the public record office had ambitious plans to establish a data archive department during the mid 198os. these have not progressed as might have been hoped. recently there have been signs of life from this quarter, and we are given to understand that the pro’s computer-readable data archive will be established in 1995, with public access in 19975. it is important to make clear that there is a distincton between machine-readable files and datasets (which are the material administered by the esrc data archive and other similar services) and machine-readable archives. the latter, like archives generally, are materials produced by, and forming part of the activity of, an organisation of some kind (such as a government). archives of any sort are therefore unlikely to be one-time studies, or to have enough individual distinctness to allow them to be treated as discrete objects, comparable with books. archives belong together in aggregations, which owe their character to the administrative system which produced them. some of the consequences of this distinction are discussed further below. the guiding characteristics of mad2 the work both of the archival description project and of the ica’s ad hoc commission on archival description, has shown that certain basic principles underlie all description of archives. international agreement on this, at least as far as traditional records are concerned, is quite remarkable. the description of machine-readable archives, therefore, is likely to require attention to these principles, if only to test their applicability to new materials. the following section attempts to summarise what the basic rules are. 1. levels of arrangement and description. the idea that there are standard levels of arrangement is not new. the concept was first indicated in europe at the start of the 20th century, then clarified in the usa6. it has been rediscovered and republished in different forms ever since. mad2 restates the principle, but also extends it. a table of levels is given which looks at first sight like the hierarchical continuum characteristic of a classification scheme, and numbered like one: 0 repository level: suitable for combined descriptions covering more than one repository. 1 management levels assemblies of archival groups brought together on the basis of some common feature, for the convenience of the repository. e.g official/nonofficial archives, ecclesiastical archives, private papers. subordinate groupings may be numbered using decimals of 1. 2. group or collection level (internationally fonds): the archives of distinct entities. subgroups (functional divisions within the group) are numbered using decimals of 2. 3. series (within britain, termed class): physically related sets of archives. subseries are given decimals of 3. 4. items: the unit of physical handling (volume, file, box). 5. pieces: indivisible components; documents. levels 4 and 5 may be used interchangeably in some cases. the interesting thing about this table is its universality. yet it is unlike a general classification scheme because it is tied to observable external phenomena at three points: fonds (level 2) always relates to the total archival product of a distinct entity (organisation or individual); series (level 3) are always the physically and systematically related product of an administrative activity, sets that belong together because of the way they were created and used; items (level 4) are always the physical units of handling. no level of arrangement is compulsory; though in the madrid principles it is stated that the level of the fonds is “the broadest unit of description”7. therefore, provided that we accept that the three levels above must always be set to correspond to the appropriate physical entities, any or all of the levels of arrangement can be used, above the fonds, or below the item, as convenient. there can be problems in identifying what should constitute a fonds. mad2 advises that administrative or political levels of dependence should be disregarded. thus an overall or umbrella organisation can be the origin of a fonds, but so can organisations which are administratively part of it. an extreme illustration would be that the government of a country could be the source of a fonds (provided that it did actually produce records as such); but so could any of its departments, or even lower subdivisions, sections etc. if any organisation is complete enough in itself to produce its own archives, it can originate a fonds 1 2.the multi-level rule the multi-level rule in mad2 states that archival descriptons should normally embrace more than one 9spring/summer 1994 level of arrangement. this is fully consistent wth the multi-level rule laid down. in isad(g), and in the madrid principles. however, mad2 has a further elaboration of the principle, which has an important use in the context of finding aids. this is the concept of the ‘macro’ and ‘micro’ description. these two terms do not relate to the specific levels of arrangement which are being described, but to the relationship between them. for example, finding aids frequently contain descriptions at fond, series and item levels. in these, the macro-micro relationship has a triple form: fonds description: a macro description governing: series description 1: a micro description in relation to the above, but a macro governing. item descriptions: micro descriptions of items in series 1, governed by the above. series description 2.... etc in the mad2 models, guidelines suggest that these relationships of dependence should be demonstrated to the user by the use of narrower margins, left and right; this assumes a hard-copy finding aid using standard pages. that is a common situation but not the only one. the important thing is that in any given case, the macro and micro descriptions may relate to any level of arrangement: fonds/item; management group/fonds; item/piece, etc. it is therefore a misconception to regard the macro description as peculiar to the ‘higher’ levels of arrangement, and the micro to the ‘lower’ ones. macro descriptions are written from a different standpoint than from micro descriptions. their standpoint is the aggregate (whichever it is). micro descriptions give information specific to each case. in the example above, the fonds description will give information relating to the fonds as a whole (probably including provenance information, but this is a separate issue); it also gives all information common to the series which follow, in order to avoid redundancy the series descriptions which follow have a dual character. in so far as they are micro descriptions, they deal with each series one by one, giving specific information. each serie descripticn then operates as a macro for the items which follow. as macros they give information which relates to the series as a whole, and common data for the items. finally, the items give data specific to each case. this rule has been explained at some 1ength because it makes it immediately clear that, and why, standards originating in library practice are not suitable for archival applications. 3. the data elements table and its structure archival descriptions require data of two different kinds: information about the origin, background, context and provenance of the archive; and information about its content. descriptions must therefore be essentially structured. the project team drew up a list of the data elements that can be found in these descriptions, and drew them together into seven ‘areas’. like the levels, most data elements and areas are optional, and are brought into use only when required for the specific case. mad2 sets out a number of models which govern the way in which descriptions can be set out, using the data elements and areas. these models accommodate the multi-level rule and allow the dependence of micro upon macro descriptions to be demonstrated so as to be easily perceived by users. 4.access points and provenance isad(g) introduces the concept of access points, which should be subject to authority control. access points should be provided for provenance information as well as for data from the contents of documents. work on authority files, sadly lacking in the archive world, is therefore needed. standards for the description machine-readable archives unless it is true that machine-readable archives are quite unlike any other archive, description standards for them should follow the models and rules for archival description, including the basic principles outlined above. section 25 of mad2 deals with this problem. although it may be anomalous to speak of levels of arrangement where the materials can never be physically arranged, it is nevertheless true that there must be levels of description. both the fonds (the archive of a whole organisation) and the series still appear to have a real existence. there is some debate about whether or not machinereadable archives must be treated in a radically diferent way from other archives9. those who concentrate on the media which carry electronic documents, are conscious above all of its evanescence, its lack of objective existence. those who look primarily at the origin, context and purpose of the document will have a much more traditional picture. the fonds will doubtless also contain descriptions of traditional archives, or archives in alternative forms. the series is normally the dataset which can most be regarded as a complete entity for 10 iassist quarterly description and managenient purposes. it most clearly resembles the datasets held by the esrc data archive. mad2 proposes that there should be short descriptions of the entities at these two levels, written into the main finding aids of the repository. when this is done, separate and specialised descriptions of the machine-readable groups and classes can be established, with a linkage between the two systems. this method allows a generalist repository to have a finding aid system which is an effective intellectual control over its total holdings, while at the same time designing a specialist finding aid which is appropriate to technically different materials. the specialist description may itself be multi-level, or it may be a flat file, according to circumstances. it must clearly contain all the metadata required: the technical information needed to record the internal structure of the file and its software dependence. the data elements needed for this are listed in section 25. a final note might be that background, context and provenance information should always be provided, because without it the meaning of the electronic record is lost. indeed this point is conceded by the practice of data archives. this ‘macro’ information, however, does not necessarily have to be held in the detailed, specialised, file which is the direct finding aid to the machinereadable data. it may be held in the main finding aid system of the repository. in future, this main finding aid may of course be itself held in a machine-readable form; or it may be processed so as to enter it into a national index, or into a data entry system. for these, both cataloguing and data exchange standards will be needed. 1. paper presented at iassist 93 in edinburgh. 2. m.cook & k.grant. manua1 of archival description. society of archivists, 1986. [some exemplars of this edition were wronly marked ‘2nd edition’] m.cook & m.procter. manual of archival descriotion 2nd ed. gower, 199o. 3. copies of texts and curent drafts are obtainable from the secretariat of the inteinational councl on archives ad hoc commission on archival description, national archives of canada, ottawa 4. alan hopkinson & m.cook. lnformation from the former at the library and archive, tate gallery, london. 5. alexandra nicol and steven duffield. unpublished paper to seminar on electronic records held at the school of library archives and information studies, universitv college, london, 10 dec 1992. 6. richard lytle. subject retrieval in archives: a comparison of the provenance and content indexing methods. phd, university 0f maryland, 1979. 7. statement of principles regarding archival description, first version, revised, section 2.2. 8.terry cook. treatment of the archival fonds: theorv, method and practice. bureau of canadian archivists, ottawa, 2. it is interesting that the concept was known to, but misunderstood by, the marxist regimes ot eastern europe. they adopted the habit of setting the fonds at too high a level of institutional independence, hence most institutions had to be the originators of sub-fonds. this misuse serves to underline the validity of the concept when used properly. 9. charles dollar. new developments and the implication on information handling’. in information handling in offices and archives, ed. angelika menne-haritz. k.g.saur, 1993 pp56-66. vol192 36 iassist quarterly by household survey standards, the slid database will be large and complex. even with our best efforts to make it approachable, researchers will need to make an “up front” investment of time and effort to come to grips with it. why is this so? number of variables and hierarchical structure perhaps the most fundamental reason is the size of the dataset and its internal relationships. as a rough estimate, there are 500 distinct variables in the full dataset, without taking the time dimension into account. this means that events, spells, variables collected annually and variables collected as many times as applicable are all counted only once — and there are many such variables in the dataset. hierarchical relationships abound in the data. a person can have several employers and information is collected on up to six jobs per year. there may be several work absences from each job. over time, even if a person does not change employers, he or she can have several occupations, wage rates and work schedules. the survey will also yield information at the household and family level. because of the hierarchical nature of the survey content, we are processing the data in a relational database environment and are also proposing to use a relational database for the microdata output.2 time dimension like all longitudinal surveys, slid users will need to grapple with the time dimension. from the time perspective, we can distinguish different types of variables. first, variables like gender, year of birth and ethnic origin, are fixed if an error is detected these variables may be corrected but otherwise they do not change over time. next, there are annual variables, such as weeks worked during the year and investment income. for these variables, the reference period is by definition the calendar year. thus, for a full panel, there will be six observations for each record. there are also cumulative variables, like years of schooling, years of work experience and number of children where, depending on the respondent’s activities or circumstances, the values may or may not require updating each year. finally there are dynamic variables which relate to spells. the duration of a spell may range from a week to several years3. slid’s content includes many variables expressed as spells and, to facilitate analysis, spells that cross the seam between two reference years (for example, an unemployment spell that begins in november and ends the following march) will be disseminating data from longitudinal surveys: issues facing the survey of labour and income dynamics by maryanne webber1 statistics canada i. introduction the survey of labour and income dynamics is one of several new longitudinal household surveys being mounted by statistics canada. like the others, slid is preparing for the release of its first round of microdata. the dissemination of microdata from longitudinal surveys poses several challenges. the purpose of this paper is to outline these challenges and some of the measures being proposed to deal with them. the paper begins with a brief overview of the survey content and design as context, but the main purpose of the paper is to provoke discussion on general dissemination issues, using slid as a case study. the intended audience is research librarians and others who will play a role in the dissemination process. ii. outline of the survey slid is designed to track the experiences of individuals in the labour market, their level and sources of income and changes in family life over a period of six years. the first panel began in 1993, with labour and income information collected from about 31,000 persons aged 16 and over. a second panel will begin in 1996, doubling the sample size. in 1999, when the first panel ends, a third one will begin. this approach of rotating, overlapping panels ensures that the sample remains representative. during the six years, 13 interviews are conducted. a preliminary interview is done when a panel first starts up, to collect background demographic, education and work experience information. one year later, an annual cycle of labour and income interviews begins. every january, information on the person’s labour market activities throughout the previous year is recorded; in may, income sources and amounts for the previous year are collected. a summary list of variables from the survey and a chart depicting the main types of information are presented in appendix. major research areas will range from employment and unemployment dynamics and labour market transitions linked to the life cycle, to job quality, workplace inequality issues, family economic mobility (dealing with shifts in income level), low income dynamics (or flows into and out of poverty), demographic events and the relationship between work and education. researchers are expected to come from many disciplines. iii. database size and complexity: the main challenge 37summer 1995 linked up on the database. in effect, the dataset that will ultimately look like the information for a six-year period was collected retrospectively at the end of the six years, as opposed to being a series of unrelated snapshots. units of analysis another factor that adds to the learning curve — and this again is due to the hierarchical properties of the data — is that there are many possible units of analysis. the person is the basic unit. in addition to being the appropriate unit for many types of research focused on the individual, the person will also generally be used for studies of the family. because family composition can change over time, the definition of family poses some sticky problems in longitudinal research4. one can however define the person as the unit of analysis and develop typologies to characterize the person’s family circumstances over the study period. the person-job is a unit of analysis used with data from labour market surveys with a one-year reference period, like the survey of work history and the labour market activity survey. we expect that researchers will also use the personjob for slid studies. this unit of analysis came about as a way of handling the fact that a person may have several jobs, concurrently or consecutively, during a one-year period. instead of using complex and arbitrary assumptions to select a main job for the year, all jobs are included and weighted using the respondent’s sample weight. sometimes they are further weighted by annual hours worked, so that part-time jobs lasting one month are given less weight than full-year, full-time jobs. some studies will use spells as the unit of analysis. for example, if a person is unemployed for two separate stretches during the study period, the two spells of unemployment will be included, both receiving the respondent’s sample weight. demographic and other characteristics can be treated as attributes of the spell. similarly, researchers may use transitions as a unit of analysis. some transitions can be identified from dynamic variables, when one state ends and another begins. some data users will no doubt want to develop definitions of transitions tailored to a particular study. for example, it should be possible to use slid to study work-to-retirement transitions or job promotions. but since these are complex processes, there is no variable or flag on the database identifying these events. rather, the user will need to look at a range of variables and explicitly define the event of interest. iv. tools to help researchers get started the survey staff are very aware of the challenge data users face in getting started. it is incumbent on us to develop tools and user support strategies that increase data accessibility. what are these tools and strategies? database design because of the size and complexity of the data, a data model was developed. this is a device for structuring the survey content and giving explicit expression to the relationships in the data. the development of the data model was done following two important principles, both of which were intended to aid the data user. first, variables were defined in keeping with the survey’s content objectives, rather than as a simple reflection of the questions and response categories used in data collection. the survey questions are designed to accommodate data collection, and are often not that useful as analytical variables. for example, to collect one content item, there may be several different questions addressed to various subgroups. second, the decision to collect data annually was based on respondent recall and other operational considerations. it was decided that this feature of the data collection operation should be transparent in the output variables (except of course in cases where annual observations make sense from a content point of view). the data for a six-year panel should look like they were collected once covering the full six-year period. these principles required a significant “up front” design and development effort but hopefully they will pay off in downstream benefits to data users who would otherwise have to recreate “seamless” data from a series of snapshots. software to retrieve data from database we are planning to provide a public-use microdata file with front-end software that, at a minimum, allows users to select variables and subpopulations of interest, for specified timeframes. these smaller datasets can then be downloaded into a flat file for further analysis using whatever software the user chooses. there will also be easy ways of producing simple frequency counts from the full dataset, to help users define their study populations. cd-rom the public-use microdata file will be available on a cdrom. this will hopefully increase data accessibility. major reference products there are three types of documentation in the works: technical documentation of the database content and structure; a user handbook; and research papers providing detailed documentation on specific topics. the main slid database is being designed with the technical user documentation — variable names, descriptions, definitions, algorithms for derived variables, code lists and user notes — as an integral part. this documentation is being stored in a relational format, so it is possible to extract parts and produce customized reports. microdata users will be able 38 iassist quarterly to access the documentation electronically as it will be imbedded in the product. a handbook or “friendly” user guide is also being developed. this should be of interest to users of custom tabulations as well as to actual and potential microdata users. after the first few editions, this publication will probably stabilize and enjoy a relatively long shelf-life — perhaps we will re-issue it every six years to coincide with the completion of a panel. finally, slid has a general purpose research paper series. since 1992, we have produced about 15-20 of these reports each year5. we are beginning to use this series as a repository for detailed information on specific variables, for example, the composition of “roll-up” categories for mother tongue and ethnic origin. workshops to get started, some users may be interested in participating in a workshop. we are quite sure that there will be interest in a workshop on the content and structure of the database. we have already been asked by a few groups to do workshops of this type and have agreed. there may also be interest in analytical techniques appropriate for use with these data. sharing information on research in progress throughout the survey development process, decisions and issues have been documented in the quarterly newsletter, dynamics. while there will still be developments to communicate in coming years, we expect that the role and content of dynamics will gradually shift, hopefully becoming a forum for exchange on research underway outside as well as inside statistics canada. it is very beneficial for the survey staff and the agency to be aware of data uses (as well as research not being done because of the lack of a few key variables). short research summaries in dynamics would keep us up to date and could supplement whatever other exchange mechanisms exist among researchers in a particular field. v. confidentiality longitudinal surveys in general face a challenge because the events and transitions that they document — and that are central to their analytical potential — may create risks of disclosing the identity of respondents. moreover, when the first wave is released, it is impossible know what patterns of change over time will be common or rare several years down the road, which means that we may need to reconsider the content of the public-use file as the data from successive waves build up. in slid’s case, there are difficult trade-offs between geography, family information and labour market detail. the data are supposed to meet the needs of researchers in a range of disciplines and to allow analysis of the interactions that exist between labour market behaviour, family circumstances and income. this makes it very difficult to protect confidentiality without “short-changing” any particular user group. the search for solutions is very lively. research is under way on techniques for quantitatively assessing disclosure risk and on alternatives to suppression and collapsing. other statistical agencies are being consulted on their approaches. an attempt is being made to prototype a remote access system, which would allow researchers to write and test their programs off-site and telecommunicate them to us so we could execute them against the full database. we are also investigating the possibility of licensing researchers to use a middle-level file for a specified purpose, following stringent rules regarding access, security and disposal. there is enough concern and energy being devoted to this issue to hope that solutions will emerge. in the meantime, we are defining the content of a public-use microdata file that would be screened using the usual statistics canada procedures. several analytically interesting derived variables are being added to the file to reduce the impact of missing detail. here are a few examples: * several occupation typologies; * the relevant low-income cutoff, or a measure showing family income as a ratio of the relevant lico; * a derived variable showing the link between occupation and major field of study. hopefully, variables such as these will help researchers to proceed with their work even if some of the very detailed information (like 4-digit occupation) is not on the public-use file. we also face a dilemma with respect to family information. on the main base, it is possible to link up family members (and previous family members) but, to provide this capacity on the public-use file, it would be necessary to reduce the amount of labour market information. as a compromise, we are proposing to include a good range of family variables, but only for a subsample of respondents. this means that researchers have access to more variables on the public-use file and, should they require results for the full population, the same program can be re-run against the full data base. these measures will ensure that, even if some variables are missing, the public-use file will still be a rich source of information. vi. computer-assisted interviewing and user documentation although it does not exclusively concern longitudinal surveys, the move to computer-assisted interviewing for 39summer 1995 household surveys at statistics canada is raising some interesting documentation issues. we are finding that efforts to document the questionnaire are proving to be very labourintensive and error-prone. we have been searching for tools and techniques to improve the process and trying to promote some measure of consistency across surveys. a working group was set up recently in the household surveys area to address this issue. it looked at a number of options. one idea was to produce a print image of each screen. however, this would yield very bulky documents and, for surveys with complex branching (like slid), it would be nightmarish to follow flows. also, even with that level of detail, many special features such as hot keys and edits would not automatically be documented. similarly, the idea of producing a diskette with the questionnaire is appealing at first blush but this would not be very meaningful as a “stand-alone” product. the user would need to learn the data collection software. moreover, many survey applications —particularly longitudinal ones — do not start with a blank sheet. there are prefilled items that affect questionnaire flow. without these prefills, one cannot get into various branches of the application. after examining these and other options, the working group found that, at least for the time being, the best approach is to concentrate on producing a good survey codebook. among other advantages, this is an approach where standards or guidelines across surveys are a reasonable goal and where the documentation reflects the data user’s perspective. this means that the user documentation of a questionnaire would begin with the output variables and work backwards, ending with the questions underlying the variables. instead of expecting users to follow complex flows through hundreds of questions, each question or group of questions would have a “universe statement” describing the question’s target population. the group also concluded that different surveys would require different supplementary tools, depending on audience, length, complexity and periodicity. in slid’s case, flow diagrams showing the organization of the survey content at increasingly detailed levels are being developed. v. conclusion once established, longitudinal surveys can be invaluable — but it can take time to become established. in the current fiscal and social policy climate, time is at a premium. new longitudinal surveys cannot afford many years to demonstrate their value. there is therefore a pressing need to support researchers in getting started. in this paper, some of the dissemination measures planned for slid have been reviewed. feedback on current plans will help us to get off to a good start. at the same time, this is a learning experience for survey staff as well as researchers. we fully expect to make adjustments to products and services and therefore hope to sustain a dialogue on enhancements. job information absences from work employer attributes household family information job characteristics jobless person work experience labour market activity patterns income sources assests & debts monthly receipt of ui/wc/sa level of schooling educational activity geography information on person's children demographics disability ethno-cultural income & wealth personal chararcteristics education person labour slid organization of contenet 40 iassist quarterly appendix: overview of slid content partial list of variables i. labour nature and pattern of labour market activities -spells of employment and unemployment (start and end dates, durations) -weekly labour force status -total weeks of employment, unemployment and inactivity by year -multiple jobholding spells -work absence spells work experience -years of full-time and part-time employment -years of experience in full-time, full-year equivalent characteristics of jobless spells -job search during spell -dates of search spells -desire for employment -reason for not looking job characteristics (all characteristics updated each year and dates of changes recorded; collected for up to six jobs per year) -wage -work schedule (hours and type) -benefits -union membership -occupation -supervisory and managerial responsibilities -class of worker -tenure -first date ever worker for this employer -how job was obtained -reason for job separation characteristics of work absences lasting one or more weeks (collected on first and last absence each year, for each employer) -absence dates -reason -paid or unpaid employer attributes -industry -firm size ii. income and wealth personal income -annual information on about 25 income sources -total income -taxes paid -after tax income receipt of compensation (whether benefits were received from each source and, if so, in which months) -unemployment insurance -social assistance -worker’s compensation assets and debts information might be collected once or twice in life of panel on roughly 20 asset and debt categories. iii. education educational activity -enrolled in a credit program, months attended -type of institution -full-time or part-time student -certificates received educational attainment (updated annually) -years of schooling -degrees and diplomas -major field of study iv. personal characteristics demographics -year or birth 41summer 1995 -family events (separation, death, birth) main features of slid objectives 13 interviews over 6 years: -preliminary -6 labour ( jan) -6 income (may) first panel started jan. 1993 31k persons 16 and over second panel starts jan. 1996 results of preliminary interview released (publication) now processing first wave 1. paper presented at iassist 21st annual conference may 9-12, 1995, quebec city, canada. 2. the first wave (including results from the preliminary interview) will, however, be released as a rectangular file. the content has not yet been finalized but our best estimate is that the record length will be about 3000 bytes for a total file length of roughly 90 kb. every year, the dataset grows, i.e., the second year’s file will incorporate and replace the first. 3. the basic time unit used in dynamic variables differs depending on the state being measured. for example, spells of employment and unemployment are measured in weeks, as are work absences. marital states, job tenure and receipt of ui are among the variables measured in months. 4. for a discussion of this issue, see greg duncan and martha hill, “conceptions of longitudinal households: fertile or futile?,” journal of economic and social measurement (1985) vol 13, pp. 361-375. 5. abstracts appear in our quarterly newsletter, dynamics. also, an annual supplement to dynamics presents abstracts for all research papers produced during the year. for major developments and issues, there is generally also a longer write-up in dynamics. -sex -current marital state and date it began -year/age at first marriage -number of children at home -parents’ schooling ethno-cultural -ethnic origin -member of an employment equity designated group -mother tongue -citizenship -country of birth activity limitation -annual information on activity limitations and their impact on working -satisfaction with work information on person’s children -number of children born, raised -year and person’s age when first child born geography and geographic mobility -economic region or cma of current residence -size of community -moved during year -move dates -reason for move -nature of move (full household/household split) household and economic family information (annual summary information at household level, e.g., size, type) -key characteristics of other individuals in household (e.g., age, sex, relationship, income, annual hours worked) -household/family size and type -family income -relevant low-income cutoff bibliographic references for computer files in the social sciences: a discussion paper by sue a. dodd' background a recent discussion among the participants of the e-mail "informal list for official representatives of icpsr" centered around citing computer files in references, footnotes, and bibliographies; whether to cite a codebook or file (providing you have both); and a discussion on citing primary or secondary sources. with respect to the last two concerns, there appeared to be adequate response indicating that it is better to cite the file as opposed to the codebook, and that one generally cites primary data sources. however, the first concern required more information and the icpsr or meeting was targeted as the next opportunity for such a discussion. note: this paper was first presented at the icpsr or annual meeting in november 1989, but has been revised for the may-june 1990 lassist meeting in poughkeepsie, n.y. identifying the problem there is good and bad news. the good news is that researchers arc beginning to cite computer files in the reference sections of social science journals. the bad news is that for every person who does cite his or her data source, another twenty to thirty continue to provide no citations. this means that valuable data sources will not be indexed by bibliographical services such a social science citation index : and more importandy, the next researcher who would like to analyze these data may not have sufficient information to acquire them. despite efforts to provide researchers with examples and information on how to compile a data file citation, the overwhelming majority continue to describe data sources within the text of their articles, but do not follow-up with a citation in the reference section. the march 1981 issue of social forces was the fu-st time that a major social science journal had provided instructions (in the "auuiors'guide" section) on how to cite a machine-readable data file (mrdf) — currently referred to as a "computer file." to see if this effort had any impact on the number of computer file citations that could be visibly detected in the reference sections of social forces . i took a two year eye-readable sample for 1988 and 1989. out of approximately 90 articles describing some form of secondary analysis, there were only 12 computer file citations. one of the citations had its own unique style (see below), but it nevertheless included enough information to gain access to the data. the point being that it is better to err in-the-effort than not to give any information. inter-university consortium for political and social research. 1979. icpsr study 7708 data description. police departments, arrests and crime in the u.s., 1860-1920, principal investigator: erick monkkonen. ann arbor. there were several citations for codebooks, which i assumed to be an indirect way to cite the actual data; and with one or two exceptions, most references to census information came in the form of the gpo printed documents. what this means is that we must renew our eforts to educate and encourage researchers to cite actual data sources. we must also encourage editors and review boards within the various social science disciplines to do likewise. how can we do this and what role can lassist members play? here are some suggestions: — lassist should undertake the task of publishing a small pamphlet or work that would provide sufficient instructions and examples ofbibliographic citations for computers files. such a publication would offer a researcher the luxury of a personal/desk reference sourceeasily retrieved when needed. apparently, the information provided as part of the data acknowledgement form and that given in certain codebooks is not getting the proper attention. this work should refiect the editorial styles of different social science journals including the american journal of sociology , the american sociological review . social science information . government publications review . demographv . and the american political science review . — lassist might also want to sponsor an announcement reminding researchers to cite their data sources. this public announcement might read: don't forget to cite your computer files ... it should be sent to the various social science journals, and space permitting, it is likely that they would run it. anotlier announcement might be joinuy sponsored by several editorial boards e.g., the following edititorial boards endorse the practice of citing social science data sources ... the various associations' newsletters might also be a vehicle for this type of announcement. — individual lassist members should contact editors and discuss the importance of citing computer data sources in references. point out that researchers are obliged to cite machine-readable sources as well as the printed ones. in addition, this practice should be encouraged so that no data source is described within the text of the article without it also appearing in the reference section. — individual members should contact review boards and authors that prepare or oversee "style manuals"— including the chicago a manual of stvle and kate l. turabian's a manual for writers . note: as i was preparing this paper, i discovered a new manual put out by the american poutical science review entitled stvle manual for pohtical science . — individual members should assist icpsr in their efforts to provide quality control over bibliographic descriptions of data files and accompanying documentation. better control over bibliographic elements facilitates the citation process. ttie icpsr staff is making valiant efforts in this regard, but need more guidance and feedback. for example, when iassist members discover any discrepancies between the bibliographic elements on a title page and those presented in a citation on the verso of the title page, then this should be pointed out so that it can be corrected. — iassist members should get more involved with the national and international groups dealing with standards associated with computer publishing, production and access. social science data producers are in the minority compared with computer software producers. without more active involvement and visibility, decisions are made that exclude the needs of social science data users. for example, thens is a national information standards organization (z39) committee— known as the niso committee ff: computer software description — that includes information on bibliographic citationsfor computer software. however, dhere is no similar effort for text or data files. to the style of the respective publication and are adhered to by the author with the exception of the bracketed information designating the computer-readable format. social forces. american political science association, and demography u.s. bureau of the census. 1989. american housing survev. 1985: national core file [computer file]. icpsr ed. ann arbor interuniversity consortium for political and social research. u.s. bureau of the census. 1979. 1979 census of popululation and housing. fourth count population summarv tape [computer file]. washington: u.s. bureau of the census [producer] arlington, va.: dualabs [distributor]. american sociological review gentemann, karen m. 1978 survev of north carolina women [computer file]. chapel hill: institute for research in social science, university of north carolina. verba, sidney and norman h. nie. political 1967 participation in america [computer file]. icpsr ed. ann arbor: inter-university consortium for political and social research. university of north carolina. north carolina 1969 information system [computer file]. chapel hill: institute for research in social science [producer]. it is not possible in this discussion paper to provide anything but brief examples, but more detailed instructions on the components of a bibliographic citation are provided in the jasis article (dodd, 1979) and in part three, chapter 9 of cataloging machine-readable data files (dodd. 19821. using different editorial styles, the following examples of bibliographic citations are given below. in most cases, the computer file in question is considered a "published work" or book equivalent— even though some computer files in the social sciences do not have "sewn or glued bindings" nor are they always boxed and sealed in packages. the composition of the citations are according american f.conomic review elkins, david j., blake, donald e. and johnston, richard, british columbia election studv: 19791980 [computer file], vancouver, b.c.: department of political science, university of british columbia, 1980. social science information studies and political psychology davis, j.a., smith, t.w. and stephenson, c.b. (1978). general social surveys 1972-1978: cumulative data [computer file]. chicago: national opinion research center louis harris and associates (1973). harris 1973 confidence in government. studv no. 2343 [computer file]. new york: l. harris and associates (producer); chapel hill, n.c.: l. harris data center (distributor). electronic journal material a new phenomenon brought about by computer technology is the so called "electronic journal." computers have changed the way that scholarly articles are created. for example, articles may be prepared using computers and word processing programs, then sent to colleagues for review via email and computer networks, and later returned to the author. once completed, they are submitted to a discipline-related electronic journal and/or computer list-server for storage and access on demand. an example of an electronic journal is the public access computer system review (pacs review). this journal focuses on "public access" computer systems that libraries make available for patron use. articles are stored as files on the pacs forum hst-server. the table of contents section of the review is sent to all pacs forum users, who can retrieve articles of interest from the list-server by following the instructions contained in that section. pacs review is published three times a year, has an editorialboard and is copyrighted. it also features special departments and reviews of others works. the first volume and issue appeared in january 1990. because standards and past traditions fall behind technology andthe capability to fjroduce computer-generated works, there are no definitive "guidehnes" for citing an article that appears in an electronic journal. however, common sense plus building on what is currently in place, makes the leap from print to computer-readable amanageable feat the following examples reflect articles that have appeared in pacs review. morgan, james jay. 1990. "expansion and testing of a meridian cd-rom network" [computer file]. houston, texas: public-access computer system review . electronic journal. 1(1) 34^2. (access via email "get morgan prvinl" listserv@uhupvm1 or lib3(a)uhupvml .bitnet] stigleman, sue. 1990. 'text management software" [computer file]. houston, texas: public-access computer sv.stem review . electronic journal. 1(1)522 . (access viaemail "get stiglemaprv 1n l ' listserv@uhupvm1 or lib3@uhupvm 1 .bitnet). on-line databases treated as serials many computer works take the form of true serials or ongoing databases (sometimes called "dynamic databases"). they can be cited in bibliographies just as other types of computer files. the only difference is that they have a beginning date and ending date — provided the serial is complete. for those that are continuing, then only the beginning date is given, followed by a hyphen and blank spaces. university of north carolina. 1989irss catalog of data holdings [computer file]. 3rd ed. chapel hill, n.c.: institute for research in social scienc on-line database. (uirdss@uncvml.bitnet). . 1989ouestions from the louis harris surveys 1958 to the present [computer file]. chapel hill, n.c.: institute for research in social science. on-line database. (uirdss@uncyml.biinet). , 1989ouestions from the usa today polls 1983 to the present [computer file]. chapel hill, n.c.: institute for research in social science, on-line database. (uirdss@uncyml.bitnet). email and computer related items use of electronic mail and networks among social scientist has grown rapidly in recent years, but most are only using a fraction of the power and resources available world-wide. networking in the future will be the way to access and disseminate data resources— especially if the cost remains so low. tapping into these data resources and alerting others to their availability becomes the responsibility of all the email and network users. just as with unpubishcd manuscripts, thesis, dissertations, or letters, there are ways to give credit to authors and provide sufficient information for subsequent access. works created using the computer and later circulated via email and networks would most likely fall into the category ofunpubushed material and more specifically "typescripts." in fact, to coin a new phrase, they would more aptly be 16 assist quarterly called "computerscripts." when citing unpublished or forthcoming computer works such as an email letter, thesis, computer-readable article, etc., be descriptive about the nature of the item and include as much information as is reasonable. jones, paul. 1989. "what is the internet?" academic computing services, university of north carolina at chapel hill. email. (pjones@samba.acs.unc.edu). holland, alccia, ed. 1989. institute for research in social science newsletter (jan.). university ofnorth carolina at chapel hill. computer-readable mimeo. cox-byme, sarah. email letter tolauraann guy dated october 5, 1989. email correspondence. petterson, lynne m. 1990."the impact of international competition, technological adoption and industrial restructuring on the evolving geographic distribution of the contemporary american machine tool industry, "[computer file]. ph.d. diss. university of north carolina at chapel hill. updegrove, daniel a., john a. muffo, john a. dunn, jr. "electronic mail and networks: new tools for institutional research and university planning." [computer file]. air professional file. forthcoming. references dodd, sue a. 1979 "bibliographic references for numeric social science data files: suggested guidelines." journal of the american society for information science . 30:77-82. dodd, sue a. 1982 cataloging machine-readable data files: an interpretive manual . chicago: american library association. pp.169-172 'presented at the iass 1st 90 conference held in poughkeepsie, n.y. may 30 june 2, 1990. institute for research in social scienceuniversity of north carolina chapel hill, n.c. 27599 usdodd@ uncvml.bitnet rev. may 1990 permission to reprint and distribute given only if this attribution is also given. summer 1990 vol221 spring 1998 7 introduction the environmental situation in new jersey is complex. the state is replete with areas tremendous natural beauty: pristine beaches and sand dunes, lush green countryside, historic towns, rolling hills, scenic river valleys and mountain overlooks, lush wetlands, and dense pine forests. by many measures, however, new jersey is the most polluted of the united states’ fifty states, and its environmental heritage is undermined by past and present treatment of the environment. efforts to remediate environmental damage resulting from new jersey’s industrial history and current industrial practices, and attempts to preserve open space through intelligent management of future development, have resulted in a tremendous amount of scientific study within the state. however many, perhaps most, of the resulting research reports are unpublished, elusive, and unavailable for secondary research purposes. this paper will outline the efforts of a partnership in new jersey to make environmental information widely and readily available. in addition to discussing the project’s background and implementation, we will cover in some detail the technological considerations for creating a webbased product that will be used to manage and query diverse information types. this discussion will include: determining system architecture and requirements, database design and programming, design of user interface and graphical design for web accessibility, and need for awareness of and compliance with current information technology standards. these are general issues, which all producers of end-user information systems must address. more specifically, this paper will discuss the technological issues involved in building a database based on a topical and geographical approach to information management. the new jersey environmental information network is neither a digital library nor a traditional catalog/index to information; it is a hybrid of both, and it encompasses all available media and formats important to the study of the environment. the njein contents range from print inventories of species within a geo-region to directory information about local experts to digital models and gis layers. the hybrid nature of the product has given rise to some unusual data management issues. project description: the new jersey environmental information network (njein) is a prototype of a web-based environmental information system for new jersey ecosystems. the authors’ participation in this project is as members of the new jersey ecological research partnership, a group of academic, non-profit, and corporate organizations concerned with making scientific information on the environment available to all potential users. special emphasis is placed on making scientific research available to enable sound decision-making processes within the state. the specific goals of the partnership include: • developing consistent, quality assured data and data dissemination mechanisms; • promoting involvement of the public in data collection, utilization, and education, • promoting mechanisms to facilitate transfer, access and use of environmental information. the njein was envisioned by the partnership as the appropriate mechanism for realizing these goals. working with new jersey’s department of environmental protection and rutgers’ ecopolicy center the authors, representing rutgers university libraries’ scholarly communication center, have received support to develop a first version of the prototype. a conceptual description of the njein would include these essential elements: 1. njein will serve as an electronic clearinghouse of information about the new jersey environment that can be queried by location, topic, or both. 2. it encompasses all topic-relevant information, regardless of physical format. 3. the system provides retrieval of located sources through: a) downloading of data; b) “scan-on-demand” system for non-digital objects (fee-based); the new jersey environmental information network: providing access to new jersey’s environmental literature by ron jantz & linda langschied * 8 iassist quarterly c) web links where available; d) referral to a physical repository where appropriate. 4. the njein will grow as items are entered into the database by data holders/producers, rather than by librarians/information managers. 5. njein’s structure (both intellectual and technological) reflects an emerging concept of “placebased management” which is being embraced by multiple federal and state agencies; it also acknowledges the naturally-occuring approach to environmental studies that scientists and other researchers employ. the importance of place-based information organization is particularly manifested in the prominence of gis usage by environmental researchers. therefore we determined that all items entered into the database, whether digital or nondigital, will be geo-located and thereby retrievable by geographic query of the database. 6. because of the foregoing, and in order to accommodate the gis data layers produced by the new jersey department of environmental protection and others, the njein will utilize the fgdc metatdata standard for descriptive cataloging of items in the database. 7. particular emphasis is given to capturing the abundance of “gray” literature and data that is produced in the state, often through project-specific research done by consultants, and making it available for secondary research. new role of librarians in information systems design: the development of the njein in a library setting illustrates the rapidly changing role of librarians in providing information access to electronic information. a decade ago, electronic information service librarians acted as the intermediary between patrons and remote, commanddriven databases; with the advent of end-user systems, librarians became coaches to hands-on users. new tools and, perhaps even more importantly, new institutional imperatives (in particular, the impetus to create digital libraries) have contributed to reshaping the librarian role yet again, with librarians entering the picture earlier in the process as participants in creation of electronic products. developing the njein, an electronic information tool intended for all levels of research from a diverse user community, requires an application of technology that is sophisticated enough to satisfy expert scientific inquiry, yet friendly enough for the interested citizen. the new jersey ecological research partnership elected to seek librarians, rather than systems programmers, to lead the project’s design and development in the belief that technology is only as good as a deep understanding of the needs of its users. using this interpretation, the field of librarianship does embody numerous characteristics that are essential to assuring quality development of this kind of product. among them are: 1. direct experience in user services: librarians employed in public services and collection development have a demonstrated understanding of what kinds of data users seek, and they strategies they employ in information seeking. 2. experience with information needs of a diverse audience: the project’s primary intended audience is those officials in new jersey (e.g., municipal and county administrators, planning boards) who are charged with local decision-making that could effect the environment. but a corollary goal is to make the same information available to those people who might influence government decision-makers: new jersey’s scientists, students, and citizens. librarians are accustomed to working with a diversity of users, and could anticipate and account for differing levels of user skill in accessing information, particularly in designing a user interface for both entering and searching the database. 3. technological skills: while these skills could have been obtained from other sources, as noted above, the project leaders among the partnership (scientists from the department of environmental protection and ecopolicy center) expressed great concern that systems programmers might not have adequate knowledge of or sensitivity to end-user needs. therefore, their preference was that information specialists would take the lead in bringing the project online. fortunately, recent organizational and technological developments in rutgers university libraries ensured that the library is prepared to undertake the development projects like njein. in particular, the creation of the scholarly communication center (scc) within alexander library, rutgers’ graduate library for social science and humanities research, provided a foundation for digital initiatives. placing the project: the scholarly communication center it was, in fact, the scholarly communication center which drew the original partners to the rutgers libraries. in order to meet the dual requirement of user-orientation design and sophisticated technological application, the partnership sought to join forces with librarians adept in both areas. ultimately, they were referred to the librarians involved with planning rutgers university libraries’ scholarly communication center (scc). the scc, officially opened in october 1998, is a technological research, teaching and learning center. its components include: 1. a teleconference lecture hall, which gives presenters spring 1998 9 access to a plethora of multimedia presentation options, has satellite uplink/downlink capabilities, and interactive distance conferencing technology. locally, this facility allows us to hold meetings among rutgers three distant campuses; it recently enabled journalists in poland to “attend” a discussion on contemporary international journalism with faculty from rutgers school of communication, information, and library studies. 2. two hands-on, multimedia infused computer classrooms. the classrooms are used primarily to deliver hand-on instruction, and one lab has distance education capabilities. 3. the humanities and social science data center, which serves as the hub for reference services, and for research and development activities for the scc. the data center is intended to serve as a testbed site for creation of new information tools. one of its first projects was to bring online the public opinion polls conducted by rutgers prestigious eagleton institute of politics. in making the decision several years ago to raise the $3 million dollars to build the scc, our institution committed itself to a technological future. creating and maintaining this facility necessitated our hiring of new talents. thusfar, we have revised open positions to create new placements for an information technology librarian, data librarian, humanities computing specialists, and system programmers. while rutgers librarians, as members of the faculty, have tremendous autonomy in setting their professional priorities, large projects undertaken by librarians are expected to conform to overall institutional priorities. in evaluating whether to assume the development of the njein as an scc project, the librarians looked to the mission statement of the scc which includes the following goals: to serve as a testbed and demonstration site for the application, development and evaluation of electronic resources by: • providing opportunities for developing electronic resources, multimedia programs and for handling electronic data, text and images. • providing guidance, instruction and training in the development, use and evaluation of electronic resources in all formats. • delivering remotely accessible resources in support of the goals of the educational and research mission of the university. • fostering specialized projects using resources of particular interest to the rutgers community and the state of new jersey. the development of the njein clearly fit the articulated goals of the scc, and the decision to join the partnership was made. the scc would “host” the njein, and its librarians would undertake project development with financial support from the new jersey department of environmental protection. summary our ultimate vision with the njein and similar efforts is to develop focused, domain specific collections that are of specific interest to the rutgers university community and the citizens of new jersey. through these efforts, we hope to impose structure and access methods on a large amount of very useful, but also distributed and uncataloged information. traditional librarian skill sets in describing and organizing information, and in providing user education services, as well as new technological ones, are critical to projects such as njein. nevertheless, many (probably most) libraries will not have all requisite skills in-house, and that is true of this project. developing the njein is an illustration of the need for collaboration and partnership in order to meet shared goals among diverse stakeholders. in this instance, we refer to the environmental scientists, creators and users of data, and librarians/information technology specialists — all are critical contributors to njein, and its success depends on a continuing partnership. what is at stake is the future of decision and policy making within the state. we are very proud to be a part of a process that will help to preserve new jersey’s environmental heritage. designing the njein design methodology. it is appropriate to say a few words about our design methodology. traditional methodologies have generally required a static set of requirements that precede the development phase. this approach frequently assumes that the analyst can somehow anticipate and understand all the complex interactions that might occur in an information retrieval system. our approach has been to embody the requirements in a prototype which allows us to see the interactions and introduce the product at a very early stage to potential customers. hence our requirements document is, in effect, the prototype. this approach has enabled a highly interactive and iterative design process in which we might make several changes to the prototype in one day. we have been able to do this and still keep the prototype up and running while students continue to load data into the njein. many development organizations have found that prototyping is one of the most effective methods for determining system requirements. finally, we have taken the approach that there is a small design team that controls the design of the system and the database. systems cannot 10 iassist quarterly be designed by committee; the design team enters into many discussions about the design in committees, small groups, and one-on-one interactions. the resulting design changes are integrated into the prototype where the team makes final decisions based on user needs, technology available, schedule, and other factors. these approaches have been used in other information retrieval systems (crawford, 1996) and many other similar product developments. although the njein is still in prototype form, our efforts so far have allowed us to learn much about the types of information that will be entered into the database, the user interface, and the definition of the database. njein architecture and platform. to a large degree, the scc is acting as a “technology conduit” into the library; we want to explore various technologies, use the technology in prototypes and hopefully help promulgate stable technology platforms throughout the various rutgers university libraries in many useful applications. one of our objectives is to establish a technology platform that can be used repeatedly for similar types of applications such as providing access to a variety of databases on the web. as our university addresses how it will acquire and deploy digital material (sewell, 1998), we want to have platforms in place that can be used by librarians and others to provide access to electronic resources. basically, a platform is a set of identified components that work well together and remain relatively stable over a period of time. this approach allows developers to become experts and the components to be thoroughly tested so that reliability is increased and learning time is decreased. the primary platform components are illustrated in the discussion of the njein architecture below. the architecture for the njein is relatively simple and is illustrated in figure 1. we have incorporated off-the-shelf software that has been frequently used in other database applications in the scc (e.g., eagleton archive, event scheduling, microforms database) with the objective of standardizing on software components to minimize maintenance and support efforts. three basic web-enabled functions are available: 1) create/ modify the reference database, 2) search/browse the reference database, and 3) retrieve/download the digital document when available (e.g. scanned text document, image, or numeric data). the primary components of the figure 1 – njein architecture spring 1998 11 computing and application platform are nt 4.0 with internet information server, frontpage, cold fusion, and ms access. using frontpage and cold fusion, the process of providing database access on the web is fairly straightforward and can be accomplished simply with sql statements and without writing a complicated script. from an architectural point of view, the important part of this diagram is the role that cold fusion plays in enabling access from the client workstation to the server reference database. the cold fusion server processes database requests through the use of templates. the templates are similar to an html file with a “.cfm” extension and special cf (cold fusion) tags which the server recognizes while ignoring the regular html statements. using these tags and the sql query language, one is able to quite easily design web pages that will create database records or alternatively query the database to retrieve information based on a user request. as a result of a database query, output is passed back to the web server from the cold fusion server in order to create a dynamic web page. as figure 1 indicates, we are running microsoft’s internet information server (iis) which has three major features which make all of this possible: 1) internet database connector (idc) which provides built-in access to odbc databases, 2) security integrated with windows nt, and 3) isapi, a robust high-performance method of communicating with gateway programs (blum, 1997). the reference database. description. the reference database contains descriptive and access information about the document. our subject domain is focused on the environment and is geographically limited to new jersey and surrounding areas (e.g. adjacent states) that might have an impact on the environment of new jersey. the audience for this database is just about anyone in new jersey: students, practitioners, scholars, researchers, new jersey citizens, and decision makers. the medium is primarily print on paper and electronic text and images although other media such as microfilm or video will not be excluded. the document format includes reports, inventories, studies, theses, map images, numeric and gis data. the domain deals with the sources of data and where and how these data are found, selected and acquired. for the njein, as mentioned previously, we will be focusing on what is sometimes referred to as “gray” literature, or literature that has not been catalogued or indexed figure 2 – major fields in reference database 12 iassist quarterly previously. our sources will be primarily universities, institutes, governments at the county and township level, and consultants who do work for new jersey state and local governments. metadata. definition. there are many definitions of metadata and the following operational definition (ng, et al) highlights the flexibility required in an internet environment: metadata is data which characterizes source data, describes their relationships, and supports its discovery and effective use. for our database, we wanted a metadata scheme that was very flexible and would support both intrinsic aspects of the document (e.g. subject, title, etc) and also extrinsic aspects related to administration and non-bibliographic issues such as size or system requirements. further, we knew that many of our documents would have a geospatial component. given these requirements, we have decided to use a subset of the content standards for digital geospatial data (federal geographic data committee) for the definition of the reference database. our reasons for doing this are as follows: • a subset allows us to have a relatively simple database so that users (non-specialists) will be able to enter data into the reference database. • using a subset of the fgdc standard will allow us to map our data to other fgdc databases with the objective of being able to exchange data with other organizations and institutions. • using the fgdc standard will enable us to take advantage of many of the continuing efforts that are producing related tools such as parsers and compilers. • we want to be closely aligned with state and federal efforts related to gis. a 1994 presidential executive order has directed federal geographic information to be described using fgdc (executive order 12906). figure 2 shows some of the key metadata fields in the reference database, organized by major fgdc categories. for example, items in section 6.4.2 of figure 2 identify extrinsic aspects of the data whereas sections 1 and 8 contain intrinsic data items. accessing the archive. as stated earlier, a major objective of this project is to impose structure and provide convenient access to new jersey’s environmental information. although pursuing this objective from a librarian’s perspective, we are departing from traditional approaches that libraries might take in organizing and cataloging information. as in many endeavors to organize information on the web (vellucci, 1997), we do not expect to have exclusive ownership or control of the information resources. our focus is shifting from ownership to providing access and our “collection” process consists of finding environmental information and developing partnerships with those organizations who do produce and own the data. this approach lacks the architectural and user interface simplicity of a single, physical opac. however, we believe the compromise distributes the effort of maintaining the collection while also improving access. as vellucci has pointed out, “it is essential and desirable that the confining parameters that define a collection be expanded to accommodate documents that are not owned and physically housed within the library’s walls.” to accommodate this diverse environment, we have a developed an architecture that has a centralized reference database but may in fact link to other searchable databases, resulting in two tiers of searchable databases. the njein should provide the ability to not only locate a document but to also retrieve a copy of the specific document. although we are in the early stages of the prototype, we plan to put considerable effort into making it possible to actually obtain a document (as opposed to just finding a reference to it). as indicated in figure 1, there are three general types of document formats in the archive. print documents will be located at institutions such as alexander library special collections, the nj state library, and designated partner institutions throughout new jersey. we will enable a user to request a print document electronically through alexander library. frequently requested print documents will be scanned and moved to the digital documents database. digital documents include both scanned print documents as well as other static documents such as map images. this material will be located on the scc server and downloadable to a local workstation through standard web browsers. the third type of archive material is gis/numeric databases. the important distinction here is that this digital information is likely to be processed further by a user. for example, numeric data would likely be analyzed using a statistical tool such as spss and gis shape (.shp) files might be used in a gis tool such as arcview. it is here that the extrinsic aspects of the metadata become very important such as file decompression technique or transfer size (see metadata items in figure 2). we have taken a pragmatic approach to handling this type of gis and numeric data. creating and maintaining gis databases is a time-consuming and complex process. our approach is two-fold and recognizes the complexitiy of the data, our partnerships and also the diverse user population. our townships (murphy, 1997), counties and states are creating and maintaining a wealth of gis data. our spring 1998 13 reference database will have an abbreviated record that contains the essential data about the gis database and points to either the actual document or a search interface so that a prospective user can locate the desired information, understand its content and determine system requirements such as file format and size. further, where possible we will capture the resulting digital map image, index the image, and place it on the scc server. this approach has the advantage of making the map images easily available to many of our customers who are not able to deal with gis data while also having the continuing support of the gis databases reside in the locations where they are created and maintained. figure 3 provides an architectural illustration of this approach. the user interface – basic principles effective user interfaces are extremely difficult to design. the designer has to understand the user of the database, have a good grasp of system design principles, and also be familiar with the subject content. in our project, we will have in the order of 1000’s of records (as opposed to 100,000s) which permits a relatively simple navigation and search structure. we are using the tree structure as shown in figure 4 below which limits flexibility to some degree but has the advantage of a user interface model that is straightforward and readily understood. there are a few simple design principles that we have tried to keep in front of us while working on the prototype: • the user interface should entice and encourage people to want to use the system. • we will not segment the user community by introducing “advanced” search techniques (i.e. design so that everyone can use all search and browse options). • there should be many ways to access the subject content and • screen display should follow two rules of thumb (thomson, 1996): 1) no more than 30% of the screen should be filled with text and 2) the optimum range of options available at any one point is between 5 and 9. opening screen. the opening screen should describe the content, scope, access options, and size of the database (anderson, 1997). given the above rules, a major challenge in the opening screen is to present the essentials of content and access without using a large amount of obscure and inappropriate figure 3 reference database and archive relationship 14 iassist quarterly text. figure 5 provides a representation of the opening screen for the prototype. data entry. as has been discussed previously, a unique aspect of the njein is to allow new jersey citizens and institutions to add data to the collection. to date, librarians and students have been “seeding” the database with relevant documents of all types. ultimately, we would like to see the njein become self-sustaining by establishing partnerships with institutions, local governments, and citizens by which they would enter environmental data as it becomes available. the data entry screens are relatively straightforward and will not be discussed in detail here. for ease of entry, data entry is performed by selecting a document type such as “”thesis” or “report”. this allows us to customize data entry for each type of document and provides a natural context for the user who has the specific type of document in their hands. in many respects, the njein is a hybrid combining aspects of digital libraries and more conventional opacs. as mentioned previously, we will have both digital and print documents available. this objective has led us to provide a minimal level of indexing and cataloging as opposed to a database that contains all digital documents and relies on automatic indexing of the entire document for effective retrieval (witten, et al, 1996). there are two unique aspects to the data entry functions provided in the njea. creating a database record and controlled vocabulary. keeping in mind our objective of enabling users to enter bibliographic data, we have required only a few fields to be entered in each record. for all of our document types, in addition to author or originator, a record also requires entries for title, abstract, primary theme, primary place, and document type. generally, the title can be taken directly from the document without undue difficulty or user confusion. in terms of traditional cataloging tasks, the abstract will be most difficult for the novice user. users will, in all likelihood, struggle to accurately and succinctly describe what the document is about. to assist in helping users describe what the document is about we have required three additional fields: primary theme, primary place, and document type. to enter data from these fields, a user selects from a pre-constructed set of themes places, and types. themes have been selected to describe environmental topics that are specific to the state of new jersey and places include the states in the northeast, new jersey counties, and the bioregions of new jersey. as an example, figure 6 shows the screen for “reports entry”. figure 4 – user interface structure spring 1998 15 our user interface assumes that at least some of the users will be willing to take the time and effort to contribute to the system. thus, users have the option of entering as many additional theme and place descriptors as appropriate. this process adds the user’s knowledge state to the representation process (o’connor, 1996). these user interface concepts are all brought together in the single search interface in which a user can use keywords to search across the title, abstract, theme and place fields. these searches can be further limited by selecting from “pick-lists” of primary themes, primary places, and document type. for example, a user may search for “pollution” and limit the search to primary place=”monmouth county” and document type=”map”. metadata review process the metadata review process is intended to supplement the user process of entering bibliographic data. each time a user enters a record from the web, email is sent to a library coordinator. this person is assigned the task of reviewing the record for quality and completeness. since users are required to enter basic information and can take advantage of the pre-constructed lists, we expect this process to be one of eliminating records that don’t make sense or do not provide adequate information as to how to locate and acquire a specific document. the record will not actually be searchable in the database until the librarian has entered a metadata review data (item 7.2 in figure 2). browsing the reference database browsing is the process of scanning by content or structure and results in an awareness of unexpected or new content and paths in the database (dodd, 1996). in the njein, we have put considerable emphasis on browsing for several reasons. our user population is diverse and scattered geographically throughout the state so that it is difficult to educate the user community about subject terms and cataloguing rules. searching is difficult and, according to researchers, can put undue cognitive load on an uninitiated figure 5 – opening screen 16 iassist quarterly user who has to devise search strategies, determine search terms, and grapple with boolean logic (behesti, et al, 1996). the data shows that between 30 and 45% of all searches starting in an online database are concluded with browsing library shelves. browsing allows the user to become familiar with the database contents and structure without trying to understand the design principles of the information retrieval system. our second reason for placing considerable emphasis on browsing stems from a pragmatic system design approach. browsing lends itself to direct manipulation user interfaces; further our database is in its infancy and sophisticated search capabilities are not yet needed. the browsing approach is to provide as many access paths as possible to the database. these paths are summarized in figure 7: figure 6 – approaches to sample data figure 7– approaches to browsing spring 1998 17 searching after a user has familiarized himself with the database through browsing, he can undertake the more mentally demanding task of searching. in the search process, we have tried to use the knowledge gained through browsing. so, for example, the user can employ the rather straightforward search form as shown below which allows a keyword to be searched across multiple fields (title, abstract, theme, and place). this search process can be further limited by using one of the browse approaches. for example, one might limit the search process to only document types of “map image”. status and conclusion the njein prototype has served, and for the foreseeable future, will continue to serve its purpose well. through its development, the authors have gained critical knowledge and experience that will inform the creation of future products. discovery and entry of scientific information into the database is ongoing, and attendant administrative tasks and processes are in place. these operations will continue, as they are quite fundamental to the project. however, recent developments on the national environmental level have propelled us into the delivery of environmental information on a much grander scale than originally envisioned. njein as we have envisioned it may not exist, but instead merge with a national priority for environmental information management. in doing so, the njein stands to become a prototype for the rest of the nation’s states. we invite our readers to visit the njein website at http://scc01.rutgers.edu/njenvironment and send us feedback. a recent partnership between the new jersey department of environmental protection’s gis division and the federal environmental protection agency has been established to create a national registry of environmental information, the environmental information management system (eims). the authors, along with nj-dep’s gis officials, represent new jersey in this initiative. to date, all other participants come from within the epa’s vast bureaucracy. new jersey is the first, and so far, sole state participant. it is with some regret that in order to participate fully with the national project, we must relinquish some of our own design and administrative control and flexibility. yet we are convinced that uniformity and interoperability across the various eims databases is a worthy aim, intended to take researchers smoothly across the possible points of access to environmental information and so we have become partners again with scientists, researchers, and information managers, this time on a more global scale. we hope to be able to report to iassist again, in not too many years, the status of environmental information systems in new jersey, and in the nation. references anderson, j.d. (1997). indexing for information retrieval: the design of indexes for textual databases. (second figure 8 searching with limits http://scc01.rutgers.edu/njenvironment 18 iassist quarterly working draft). behesti, j., large, v. & bialek, m. (1996). pace: a browsable graphical interface. information technology and libraries, 15, (4), 231 – 240. blum, a. (1997). active x web programming. new york: wiley computer publishing. crawford, w. (1996). developing eureka: rapid access to very large databases. information technology and libraries, 15, (1), 9 – 19. dodd, d.g. (1996). grass-roots cataloging and classification: food for thought from world wide web subject oriented hierarchical lists. library resources and technical services, 40, (3), 275 – 286. federal geographic data committee standards: url: http://fgdc.er.usgs.gov/standards/standards.html. murphy, t. (1997/november). how tewkesbury’s gis program aids sensible development. new jersey municpalities. 72,88. ng, kwong bor, park, s. & burnett, k. (1997). control or mangement: a comparison of the two approaches for establishing metadata schemes in the digital environment. http://www.scils.rutgers.edu/~sypark/asis.html. o’connor, b.c. (1996). explorations in indexing and abstracting: pointing, virtue and power. (chapter 9: pp. 145 – 158). englewood, co: libraries unlimited. thomson, william k. (1996). designing effective user interfaces. proceedings of the seventeenth national online meeting. new york: may 14-16. 385 – 391. sewell, r. (ed.) (1998). a bridge to the future: rutgers digital library initiative (draft document). vellucci, s. (1997). options for organizing electronic resources: the coexistence of metadata. bulletin of the american society for information science, 24, (1), 14 17. witten, i; cunningham, s. & apperly, m. (november/ 1996). the new zealand digital library project. d-lib magazine. http://www.dlib.org/dlib/november96/ newzealand/11witten.html * paper presented at the annual iassist conference, new haven connecticut, may 19-22, 1998. http://fgdc.er.usgs.gov/standards/standards.html. http://www.yorku.ca/org/iassist/ http://www.scils.rutgers.edu/~sypark/asis.html. http://www.dlib.org/dlib/november96/newzealand/11witten.html http://www.dlib.org/dlib/november96/newzealand/11witten.html iassist quarterly 2014/2015 5 iassist quarterly editor’s notes the making of meta-analysis through metadata of the data documentation initiative for semantic web welcome to the fourth issue of volume 38 and the first issue of volume 39 in this double-issue of the iassist quarterly (iq 38:4 & 39:1, 2014 & 2015). this special issue is guest edited by joachim wackerow of gesis – leibniz institute for the social sciences in germany and mary vardigan of icpsr at the university of michigan, usa. they have arranged, participated, herded other ddi experts, as well as produced in several workshops and conferences on the issues of the data documentation initiative (ddi). this special issue on ddi addresses the semantic web. i have included the keyword ‘meta-analysis’ in the header for this ddi-issue as i believe that is where the combination of ddi and the semantic web is going to make one of its big impacts. i expect precise and detailed ddi metadata with strong applications will bring remarkable support for overview and accumulation of results from large amounts of datasets. thanks to the guest editors joachim wackerow and mary vardigan, this issue includes papers from numerous researchers involved in ddi development and use. in the overview paper on semantic web applications (thomas bosch and benjamin zapilko) i came across a term like ‘machine-understandable’ i.e., the machine (and software) is capable of understanding. in these days of the turing-movie ‘the imitation game’ you could say that ‘understanding’ is a huge part of what the turing-test investigates when looking into intelligence. however, i believe that we still must be content to stick with ‘machine-actionable’ – i.e., the machine is capable of taking special actions according to the rising conditions. i will frame the difference as the machine is capable of “if then do” while only humans so far are also capable of acting on “if then don’t”! the second paper (thomas bosch, olof olsson, benjamin zapilko, arofan gregory, and joachim wackerow) shows how the ddi has in twenty years developed into the support of the complete data lifecycle. in several figures the concepts are overviewed graphically and the use cases illustrate several scenarios of support. use cases are also at the center of the paper, including the ontology of the ddi (thomas bosch and brigitte mathiak) and here less database bound graphics are showing representations of the ddi conceptual model with a focus on the linked open data cloud and the web of data. this paper also exemplifies the query language sparql. the linked open data cloud is continued in the paper on linking study descriptions (johann schaible, benjamin zapilko, thomas bosch, and wolfgang zenk-möltgen) describing how study descriptions can be enriched with datasets from the linked open data cloud. the trick is to automatically detect that items within different sources can be successfully linked as they carry the same property. the last paper introduces the reader to xkos, the extended knowledge organization system (franck cotton, daniel w. gillman, yves jaques). the paper gives some examples of statistical classification and includes data harmonization. this paper applies the term ‘machineunderstandable’ once again. semantics is about meaning and meaning is understanding. i still don’t think that the turing-test has been passed although i might have paid insufficient attention – but the ddi and semantic web is going in a promising direction. articles for the iassist quarterly are always very welcome. they can be papers from iassist conferences or other conferences and workshops, from local presentations or papers especially written for the iq. when you are preparing a presentation, give a thought to turning your one-time presentation into a lasting contribution to continuing development. as an author you are permitted ‘deep links’ where you link directly to your paper published in the iq. chairing a conference session with the purpose of aggregating and integrating papers for a special issue iq is also much appreciated as the information reaches many more people than the session participants, and will be readily available on the iassist website at http://www. iassistdata.org. authors are very welcome to take a look at the instructions and layout: http://iassistdata.org/iq/instructions-authors authors can also contact me via e-mail: kbr@sam.sdu.dk. should you be interested in compiling a special issue for the iq as guest editor(s) i will also be delighted to hear from you. karsten boye rasmussen july 2015 editor http://www.iassistdata.org http://www.iassistdata.org http://iassistdata.org/iq/instructions mailto:kbr@sam.sdu.dk 6 iassist quarterly 2014/2015 iassist quarterly ddi and semantic web this issue focuses on the ways in which ddi can play an important role in the linked open data environment and recent accomplishments that move this idea forward into reality. the first paper, “semantic web applications for the social sciences,” presents an overview of several representative applications that use semantic web technologies, highlighting social science applications and their benefits for the domain. the second paper, “ddi-rdf disco – a discovery model for microdata,” describes a data discovery ontology based on ddi that enables users to publish their ddi data and metadata in rdf and link them with many other datasets from the linked open data (lod) cloud. the third paper, “use cases related to an ontology of the data documentation initiative,” provides additional detail related to the data discovery ontology disco, offering several use cases to show its value in the world of linked open data. the fourth paper, “linking study descriptions to the linked open data cloud,” presents ways to enrich a study description with various datasets from the lod cloud by exposing selected elements of the study description in rdf. and finally, the fifth paper, “xkos an rdf vocabulary for describing statistical classifications,” offers a brief description of the extended knowledge organization system (xkos), an extension of skos, and a rationale for why it was developed, showing how the semantics of classification systems in the authors’ own offices are represented more faithfully by extending skos with xkos. joachim wackerow joachim.wackerow@gesis.org mary vardigan vardigan@umich.edu guest editor’s notes vol262 12 iassist quarterly summer 2002 iassist quarterly summer 2002 13 by lucy bell * this paper describes the development of the one-stop census registration service (crs), an online system providing quick and simple user registration for access to all the varied resources from the 1971, 1981, 1991 and 2001 uk decennial censuses. introduction uk higher education institutions, through the uk data archive (ukda) and the universities of manchester, edinburgh and leeds, have been providing access to uk census data for over a quarter of a century. in that time, the number of data products associated with the 1971, 1981 and 1991 censuses has increased dramatically and now totals more than 50. this number will increase again with the release of the results from the 2001 census. such a wealth of resources presents the academic user with not only the opportunity to use many sophisticated datasets but also a requirement to accept and fulfil the terms and conditions of the various licences covering these products. a new single licence for all 1971-2001 census products is now in development. in line with this new licence and in time for the results of the 2001 census, the uk data archive has been funded by the economic and social research council (esrc) and joint information systems committee (jisc) to develop, implement and maintain the one-stop census registration service (http://www.censusregistration.ac.uk) to simplify the process of registering for these data products. this paper reviews the background to the project, explaining the need for such a co-ordinated service in light of the varied actual and potential uses of census data in higher and further education (he/fe) and the ways in which these applications are currently developing. it also reports the findings of a small-scale questionnaire survey of uk he/fe staff, undertaken by the paperʼs author, which sheds light on views of the old and new registration systems. lastly, the paper describes the new registration service in some detail and highlights some of the lessons learned during its development. background the official census in the united kingdom has been taken for 200 years, with 2001 being its bicentenary. it has been let us bring you to your census: recent developments in uk census data provision conducted every ten years since 1801, with just two anomalies: it was omitted in 1941 due to war-time security issues; and a small-scale, experimental, five-year census was attempted in 1966, using a 10 percent sample of the population. the end of the 20th century saw more activity than ever before in disseminating census materials to academia. in the late 1970s the uk data archive became involved, as it started to disseminate the data from 1971. in those days it took nearly ten years to finalise the census data taken at the start of the decade and to publish them. progress in the following years has speeded this up considerably. access for the academic community to these data has also been assisted by a joint esrc and jisc programme which started with the 1991 census. this programme funded a series of academic data services, known as census data support units (dsus), which have been disseminating the data to the uk higher and, more recently, further education communities in the past decade. following the development of these academic services, the data products available to uk he/fe have grown in number. the 1991 programme encouraged the development of derived datasets and tools which could be used with the census data. there are now over 50 datasets, including resources such as deprivation indices derived from the census data, which can be used. these datasets themselves comprise thousands of tables. the geographical breakdown of these data is sophisticated and, for certain data, such as the small area statistics, takes the data to very small output areas. the 1991 esrc/jisc programmeʼs online and offline access services have also been extended to bring in some of the 1971 and 1981 data. this all gives the uk he/fe community unrivalled access to a vast collection of census data. the data essentially break down into the following categories: • area statistics • boundary data 14 iassist quarterly summer 2002 iassist quarterly summer 2002 15 • interaction data (origin-destination data) • microdata (the samples of anonymised records or sars) • other derived datasets these data are distributed by the four highly-respected dsus, funded via the esrc/jisc census programme under its 2001 census budget: • census dissemination unit from mimas (university of manchester, http://census.ac.uk/cdu) for the area statistics for 1981, 1991 and 2001 and soon 1971 • census geography data unit (ukborders) from edina (university of edinburgh, http:// edina.ac.uk/ukborders) for digitised boundary data for 1971, 1981, 1991 and 2001 • census interaction data service (universities of leeds and st andrews, http://census.ac.uk/cids) for origin-destination statistics for 1981, 1991 and 2001 • census microdata unit from the cathie marsh centre for census and survey research (university of manchester, http://www.ccsr.ac.uk/sars) for the 1991 and 2001 samples of anonymised records (sars) the combination of the richness of the data resources and the sophisticated dissemination services make the census data valuable for research, teaching and learning; however, one large obstacle has blocked the userʼs way in past years: registration. because the different census products have evolved over time, almost every one has resulted in a new licence being drawn up between the esrc/jisc and one of the three census offices in the uk. these three administrations organise the censuses for their areas: the office for national statistics administers the census for england and wales; the general register office for scotland administers the scottish census; and the northern ireland statistics and research agency administers the census in northern ireland. to complicate matters further, the 1981 and 1991 boundary data for england and wales and the 1981 boundary data for scotland are owned and, therefore, licensed by three additional bodies: the office of the deputy prime minister; a consortium called ed-line; and the scottish executive respectively. because they have been developed by different organisations, each of these licences is slightly different. even those licences which relate to the same country can vary in content. add to this the fact that four distinctly different services distribute the data, all of which have required users to register separately with them, and a complicated system of access becomes apparent. the original registration procedures to access the census data involved the need to locate, print off, complete and have counter-signed locally, a series of forms, most of which were slightly different in both content and format, in order to be able to use all the census data available. a user approaching the census data in the past may well have had to fill in different forms depending on which census year he or she was interested in, which country was needed and which dsu was being used. in order to reduce the confusion and the administration associated with having to fulfil each clause of each licence, the esrc/jisc, the three census offices and hmso have developed a new tripartite agreement which will eliminate all the old licences, bringing the data from each census from 1971 onwards under the same terms. this will create a level playing field, so that all the data may be treated in the same way. the intention is also for this agreement to roll forward to bring in new censuses, as and when they are taken. the uk data archive has developed an online system which co-ordinates all the registrations for access to the census data, after having won the tender from the esrc/ jisc to supply this. it was intended to dispense with the multiple paper forms and create a smooth, user-friendly, web-based registration system which would knock down many, if not all, of the hurdles in the userʼs path to the data. the census registration service, established at the uk data archive in august 2001, has been developing this system over the course of the past year. the system is now live, meaning that just one online registration now entitles the user to access the data from all the dsus, for all the years, countries and types of data, using a simple, online, athens1-compliant system. consultation one of the first tasks in the crsʼs life was to undertake a series of consultations in order to guarantee that the development of the service satisfied all the stakeholders as far as possible. these stakeholders comprised the four dsus whose registration needs the crs serves, other experts and a sample of the current systemʼs users. three meetings with the dsus were held during the first year of the service. in order to facilitate confidential and speedy discussions about the progress of the service a closed jiscmail2 discussion list was also established in november 2001. the list has been used extensively and has proved to be a vital forum for sharing comments and suggestions. consultation with the future users of the service was considered to be just as important as consultation with the dsus. a series of surveys were prepared and undertaken between november 2001 and march 2002 in order to ascertain the users ̓views of the current registration 14 iassist quarterly summer 2002 iassist quarterly summer 2002 15 procedures which were then in place. this was done for three reasons: 1. to elicit their views about the then current and forthcoming census registration systems 2. to start to advertise the forthcoming changes 3. to discover whether any of the current users selected would be interested in beta testing the new system initially, all mimas and uk data archive class tutors using census data and ukborders-selected class tutors were contacted with an email questionnaire on tuesday 4 december 2001 (appendix a). class tutors form a special set of census data users. under the old system, whole classes of students from one university could register in one batch by signing one licence administered by their tutor. this added to the workload of already busy lecturers whose responsibility it became to ensure that their students were signed up to use the data. the email questionnaire was sent to 48 class tutors; 19 replied. their responses indicated both familiarity and discontent with the old system: • 16 said that they either knew how some, most or all of the system worked • 9 said they found the system either impenetrable or quite difficult to use • 7 found it ʻmiddling ̓in terms of ease-of-use • 10 were dissatisfied • 2 were extremely dissatisfied from their written comments, their main concerns were as follows: • gathering student signatures is very time consuming • the registration process involves too much paperwork • the registration forms tend to be long and complex • individual registration for each database is tiresome • tutor involvement should be minimised or eliminated they came up with the following suggestions: • speedy online self-registration for students • a system that is not too technical but easily explained online • a single form for all census data services when asked which aspects would discourage them from using – or wanting to use – a system, they came up almost unanimously (n=15) with: • tutor involvement in gathering students ̓names this was the key. it was reported that it was difficult to organise signatures for an entire class using hard copies of licences because some from the class will always be absent at the time of signature collection. the class tutors commented on their frustration in having to chase the stragglers. following on from the class tutors ̓questionnaire survey, all site representatives from mimas, ukborders and the uk data archive were emailed the url of an online questionnaire on 7 february 2002 (see appendix b for the text of the questionnaire). the ccsr representatives were alerted to the questionnaire through the sars newsletter. in total, 373 representatives were emailed the details of the online questionnaire. seventy-one replied, yielding a response rate of 19 percent. the results from these respondents mirrored those of the class tutors in that they indicated both an understanding of the system and some dissatisfaction about the way it worked. • 57.75 percent (n=41) were familiar with at least some of the system, with 32.4 percent (n=23) feeling familiar with most or all of it. • the respondents did not seem to find the system easy to use. 38.03 percent (n=27) found it ʻquite difficult ̓or ʻvery difficult ̓to use, 33.8 percent (n=24) found it ʻneither difficult nor easyʼ, with only 9.86 percent (n=7) finding it ʻeasyʼ. none of the respondents ticked the ʻvery easy ̓box. • despite their familiarity with the system, the site representatives were, in the main, dissatisfied with it. 42.25 percent (n=30) ticked the ʻextremely dissatisfied ̓or ʻdissatisfied ̓boxes; 14.08 percent (n=10) were, however, satisfied. no-one ticked the ʻextremely satisfied ̓box. the reasons supplied for dissatisfaction were: • the need for counter-signatures on the licence agreements and the reliance on the postal service to deliver them both slow the process down 16 iassist quarterly summer 2002 iassist quarterly summer 2002 17 • the confusing abundance of forms and the difficulty in identifying which forms are needed for which dataset and, indeed, whether all the correct ones have been found, all cause frustration • the time it takes to register puts off potential users (for students in particular, it was emphasised that the system needs to be speedy) • the need to remember multiple user names for the services and to register separately for each one is burdensome the questionnaire included a series of suggested improvements which the respondents were asked to tick should they desire them. although not all of these additional functions were ticked by each representative, they each received a tick from at least 59% of the questionnaireʼs respondents. the preferred registration functionality, in order of preference, was as follows: • a single interface for all services, 95.77 percent (n=68) • immediate access to the data after registration, 81.69 percent (n=58) • step-by-step online guidance on how to register, 73.24 percent (n=52) • the functionality to jump from one service to another without logging on again, 66.2 percent (n=47) • nothing to sign, 60.56 percent (n=43) • helpdesk facility for registration problems, 59.15 percent (n=42) the results of the class tutors questionnaire were combined with the results of the questionnaire sent to the site representatives and the desires of the respondents incorporated into the system as far as possible. the functionality of the registration system after these consultations and further scenario planning, the final specification for the service was decided. this was then programmed, beta-tested and launched. in more detail, the systemʼs features are as follows: • it is short, simple and straightforward, comprising a wholly online system with no need for paper licences and counter-signatures. • it is athens-authenticated, meaning that users will simply have to input their athens usernames to access the data. this does, of course, mean that users will have to have athens usernames before approaching the system; however, this is a straightforward process. in fact, many users will already have athens identities for use with other resources. • the crs verifies users ̓email addresses using a simple procedure. as soon as the user has completed the online form, a message is sent to the email address they entered. the email contains a url which they must visit. once they have successfully completed this, they will have registered to use the census data. there are two reasons for undertaking this email check. first, it confirms the userʼs email address and, second, it checks that the email goes to the person who completed the registration form. • the system has a facility for users to update their own details; for instance, if an undergraduate becomes a postgraduate, his or her status may be altered accordingly online. • an online feedback area exists, where users may submit enquiries to the crs team. • there is a facility for inputting the details of publications which have arisen from use of the census data. one of the conditions of the end user licence is that the user agrees to inform the crs about these publications. in order to make this easy for the user, the crs has established a centralised area of the web site where this may be completed online. • the system offers registered users the opportunity to sign up for the uk data archive as well, removing the need for them to complete additional online forms. • the site also contains other informative pages, containing news and useful links. obstacles and lessons few obstacles were encountered during the serviceʼs development, but those that did arise were complex. these have now been overcome and the relevant issues addressed. the obstacles essentially fell into two categories: licensing issues and technical issues. tracking down all the previous licences, including those for the digitised boundary data which were owned by different organisations than the rest of the census data, proved time-consuming. additionally, the crs had hoped to be able to bring other census-related datasets into the registration system in time for its launch, but this proved to be impossible due to time constraints. the technical obstacles with which the crs had to grapple related primarily to the different levels of access control that each data support unit required for their data. for example, some units were clear that, should the service 16 iassist quarterly summer 2002 iassist quarterly summer 2002 17 be opened up outside the uk he/fe community, it would not be possible for their data to be disseminated to these additional users. the crs was intended to be a one-stop shop but, because of this access issue, it has also been developed with flexibility in mind. this flexibility was achieved through use of the athens profile system. the athens system allows each of the services to be treated differently; each can be set up as a separate ʻresource ̓in the athens terminology. each user who is granted access to a resource has an electronic ʻprofileʼ, to which information may be written and from which information may be read. the census system has been established so that, once a user registers with the crs, the expiry date of his or her registration is written to the userʼs profile. when the user logs into one of the dsus the athens software reads the profile and assesses whether or not the user is entitled to access the data. this system can be extended to include other information regarding the userʼs status to allow or disallow access to individual dsus. for instance, should a user work outside the he/fe sector, a fact which can be determined from the athens username, information about access rights can be included in the userʼs profile, preventing him or her from using some units and allowing them the use of others. lessons about how not to re-invent the wheel were learned early on in the project. the use of externally established systems of communication and authentication assisted the crs enormously. one of these was athens, while another was the jiscmail closed discussion list. this allowed all parties involved to reflect on changes to the system and to suggest the best solutions. it also resulted in agreement among all four of the dsus. support has also been forthcoming from many quarters. the dsus have encouraged the serviceʼs development and given ideas and advice all the way, as have users in the field. colleagues in the uk data archive have been on hand to discuss various options whenever needed. this has supplied the crs team with a rich and varied body of experts from whom to take advice. the launch and the future the service was launched on 2 august 2002. from this point onwards, all uk he/fe census data users have registered with the crs. previous users have had to re-register with the service for legal reasons. the uk data protection act 1998 means that the crs cannot simply transfer users ̓details from one service to another; however, the bonus that exists in re-registration is that the process should be short, simple and straightforward. most importantly, filling in one online form entitles users in the academic community to use all four of the census data support unit services and the uk data archive. the crs is now entering phase 2 of its life and is gearing up to enhance and improve the service. the developments planned include a search interface to the database and more links. the crs team will also be listening carefully to its users ̓suggestions. if additional changes take place, they could well be a result of comments made by those whom the crs values very highly –the people using the data. footnotes 1 athens (http://www.athens.ac.uk) is a system used throughout uk he/fe which supplies students and staff with just one username and password to access many different databases. it works with the education – and other – sectors and with database producers. he/fe institutions must negotiate access for their user communities to different databases, after which the database producers inform athens which institutionʼs members may use their services. the institutions are then granted access to these databases by athens. athens has over 1.8 million user accounts. all uk he/fe institutions and the research councils have already been granted access to the census services, as these are the audiences for whom the services have been funded. 2 jiscmail (http://www.jiscmail.ac.uk) is a service based on listserv which allows uk academics and support staff to join, create and manage electronic discussion lists. *paper presented at the iassist conference, june 2002, in storrs, ct, usa. lucy bell, uk data archive, lajbell@essex.ac.uk. appendix a: email to class tutors, sent to 48 people via email on tuesday 4th december 2001 dear dr ~ the census registration service at the uk data archive <http://www.data-archive.ac.uk/> is in the process of setting up a one stop shop census registration service, providing a user-friendly and simple registration system for access to all the varied census resources from the 1971, 1981, 1991 and 2001 decennial censuses. this means that, soon, just one registration per person will be required to access all census materials from 1971, no matter where the data are located (mimas, ukborders, centre for censuses and survey research etc.). the service is expected to be ready in the summer of 2002. for more information on the project, please follow the link to the uk data archiveʼs projects web pages <http://www.dataarchive.ac.uk/home/censusrs.asp>. my role is to coordinate this process and one of the first things i am trying to do is to canvas opinions on the ideal design and functionality of the new service. colleagues at mimas mentioned your name as a tutor who uses the class registration procedures to access census materials http://www.athens.ac.uk http://www.jiscmail.ac.uk lajbell@essex.ac.uk http://www.data-archive.ac.uk/ http://www.data-archive.ac.uk/home/censusrs.asp http://www.data-archive.ac.uk/home/censusrs.asp 18 iassist quarterly summer 2002 iassist quarterly summer 2002 19 with your students and suggested that you may have extremely valuable views on how the new system should be developed in this regard. if you had five or ten minutes to answer the questions below and email them back to me, i would be very grateful. indeed, any comments about the current and/or future registration systems in relation to the needs of class tutors would be gratefully received. if at all possible, it would be helpful to have all responses back by friday 21st december. many thanks. lucy bell service coordinator census registration service uk data archive. university of essex wivenhoe park colchester essex co4 3sq tel: 01206 873950 email: lajbell@essex.ac.uk **** questionnaire: access to the census datasets **** 1. on a scale of 1 5, where 5 is the greatest, how familiar do you feel with the current systems of class registration required to access the 1971-1991 census datasets? please mark. 1 __ i donʼt know how the system works at all 2 __ i know only a little about how it works 3 __ i know some of it quite well 4 __ i am fairly sure how most of it works 5 __ i know exactly how it works 2. on a scale of 1 5, where 5 is the easiest, how easy do you find the current census registration systems to use? please mark. 1 __ impenetrable 2 __ quite difficult 3 __ middling 4 __ easy 5 __ very easy 3. on a scale of 1 5, where 5 is extremely satisfied, how satisfied are you with the current registration procedures? please mark. 1 __ extremely dissatisfied 2 __ dissatisfied 3 __ neither satisfied nor dissatisfied 4 __ satisfied 5 __ extremely satisfied please give the reason(s) for your answer to question 3: 4. what would make you, as a class tutor, view the new census registration system as a success? 5. which aspects of an online registration system would discourage you, as a class tutor, from using it? 6. if there are any anonymised suggestions or complaints you have picked up from students using the current census registration system and which you feel you can share, please describe them below. 7. any other comments: if you would like to express an interest in the beta testing of the new service in 2002, please just let me know. many thanks for your time in answering these questions. lucy bell service coordinator census registration service uk data archive university of essex wivenhoe park colchester essex co4 3sq tel: 01206 873950 email: lajbell@essex.ac.uk ************************************************ legal disclaimer: any views expressed by the sender of this message are not necessarily those of the uk data archive. this email and any files transmitted with it are confidential and intended solely for the use of the individual(s) or entity to whom they are addressed. ************************************************ 18 iassist quarterly summer 2002 iassist quarterly summer 2002 19 appendix b: online questionnaire for site representatives census registration service questionnaire many thanks for taking the time to visit this page and complete the questionnaire (below). the results will be stored in a database held at the uk data archive. the information is being gathered to help inform the development of the new one-stop census registration service for uk higher and further education. this new service will provide an integrated, seamless, userfriendly and simple registration system for access to all the varied resources from the 1971, 1981, 1991 and, when they are ready, 2001 decennial censuses. this means that, soon, just one registration per person will be required to access all census materials from 1971, no matter where the data are located (mimas, ukborders, cathie marsh centre for census and survey research, cids or the uk data archive). the service is expected to be ready in the summer of 2002. the answers you give to the questions below will help to determine its final design. 1. how familiar do you feel with the current system of registration required to access the 1971-1991 census datasets? • i donʼt know how the system works at all. • i know only a little about how it works. • i know some of it quite well. • i know, reasonably well, how most of it works. • i know exactly how it works. • i have never used the system (please go to question 4). 2. how easy do you find the current census registration system to use? • very difficult. • quite difficult. • neither difficult nor easy. • easy. • very easy. 3. how satisfied are you with the current census registration procedures for uk higher and further education? • extremely dissatisfied. • dissatisfied. • neither satisfied nor dissatisfied. • satisfied. • extremely satisfied. please give a brief summary of the reason(s) for your answer for question 3: 4. if usage statistics for this service for your institution were available, which level of breakdown would you find most useful? by which category of user? (tick all that apply) • i would not need the results to be broken down by category of user. • educational type (staff, postgraduate, undergraduate etc). • department. • subject area. by which time period? (tick one) • annual figures. • quarterly figures. • figures produced more frequently than quarterly. 5. what would you, as a user, like to see in an ideal census registration system? (tick all that apply) • a single interface for logging on to all services. • nothing to sign. • the functionality to be able to jump from one service to another without logging on again. • immediate access to the data after registration. • helpdesk facility for registration problems. • step-by-step online guidance on how to register. • something else (please specify). 20 iassist quarterly summer 2002 iassist quarterly summer 2002 21 6. additionally, please indicate on a scale of 1-5 where 5 is the most satisfied, your satisfaction with the following aspects of the uk data archive registration system: • the registration process 1 2 3 4 5 • logging in on subsequent visits, once already registered 1 2 3 4 5 • the email instructions 1 2 3 4 5 • resolution of registration problems 1 2 3 4 5 7. any other comments relating to the uk data archive or the census registration systems: your details full name department email address telephone number which service(s) do you represent? � mimas � uk data archive � ukborders � athens � ccsr (sars) � none of these if you would prefer your contact details not to be included in the database of respondents to this survey, � please tick this box. the new service will be beta tested in 2002, please indicate below whether or not you would like to be involved in this process. � yes please � no thank you � not sure, please contact me again later on in 2002. -----------------------------------------------------------© copyright 2002 university of essex. all rights reserved. top -----------------------------------------------------------vol29-1.indd 24 iassist quarterly spring 2005 by chuck humphrey* the preservation of research data in a postmodern culture john curtice, one of the plenary speakers at the 2005 iassist conference in edinburgh, presented time-series evidence showing the emergence of postmodern values. derived from major attitudinal and value surveys since the 1960’s, his research reveals a shift in the locus of personal identity formation from institutions, such as organized religion and political parties, to the “marketplace” of individually established identities. in the postmodern world, everyone supposedly shops for her or his own identity using today’s educational systems to compile a uniquely packaged identity through self-actualization. according to this theory, institutions no longer serve as touchstones in determining individual identities. in his address, professor curtice also showed research findings that temper the postmodernist interpretation of today’s world. specifically, many of today’s values represent an interaction between individuals and institutions. nevertheless, institutions are increasingly under attack by postmodern values. these assaults even pose a threat to national data archives and how these archives function in today’s societies. our national data archives are not immune to fundamental changes within our cultures. after all, data archives preserve one aspect of our cultural heritages, namely, digital evidence of research value. therefore, one would expect major shifts in culture to result correspondingly in changes in the ways in which the record of our cultures are preserved. what aspects of postmodernism present a threat to our national data archives? the very nature of the internet reinforces the image of individualism in today’s culture. everyone can have her or his own domain name and an identity on the internet for a modest fee. web technology, including recent web log (blog) software, has the potential of making everyone a publisher. napster enabled peer-to-peer music distribution, much to the chagrin of the entertainment industry. in canada, a temporarily publication ban was imposed by justice gomery on testimony before a commission investigating possible misuses of public funds. shortly after the ban, this information appeared on the internet from a site in the united states. as a consequence, the chief commissioner partially lifted his earlier ban because the evidence had been widely disseminated on the internet.1 these examples illustrate why the internet is perceived as a great leveler enabling individuals to compete against institutional powers within and outside today’s legal boundaries. from the perspective of data services, the internet enables us to provide access to larger numbers of users than we have been able to support in the past. dissemination is more direct and responsive to on-demand access to data resources using the internet. our profession sees these as admirable qualities of this technology. individuals are empowered to retrieve data directly and quickly to their desktops providing researchers a sense of autonomy. an erroneous corollary of this sense of access-autonomy is the concept that everyone on the internet is her or his own archivist. we see this arising within the discourse around digital repositories and the idea of “self-archiving.” this is an unfortunate choice of words to describe the act of contributing individual works to a database of research publications. a motivating force behind digital repositories has been to increase access to scientific findings more quickly and equitably, sometimes even circumventing traditional channels of peer-reviewed print publications. some commentators have generalized this self-archiving concept to incorporate research data among the digital objects researchers should contribute to digital repositories. self-archiving to me is an oxymoron. the act of archiving involves an institutional commitment to preserve knowledge and culture beyond political and technological changes. in the case of research, data archives represent institutions dedicated to the long-term preservation of data. ideally, data archiving is a process throughout the life cycle of research and involves the full range of contributors to a research project. while sole investigators still contribute to the overall output of research, increasingly research projects are organized around teams, especially research that is inter-disciplinary, comparative, multi-national and large in scale. consequently, the idea of an individual being her or his own data archivist runs counter to the way major research is being performed nationally and iassist quarterly spring 2005 25 internationally today. furthermore, the process of such research engages many stakeholders, including government granting agencies, universities, researchers, data producers, publishers, libraries and data archives. all of these contributors play a role in the life cycle of research data. one function of a data archive is to identify the custodial relationships among these stakeholders throughout the various stages of the data life cycle. the concept of selfarchiving is meaningless in the context of a life-cycle model, diminishing the value of institutions in a postmodern world and accentuating the individualistic attributes of the internet threatens data archives as institutions. research in canada conducted in conjunction with the national data archive consultation demonstrated the need for institutional support of data to ensure its long-term access. a study of 100 funded projects by the social sciences and humanities research council in canada between 1977 and 1980 found data for only three of these projects in 2001.2 the data for all three projects had been deposited with the inter-university consortium for political and social research (icpsr) at the university of michigan.3 this canadian study demonstrates the level of risk that research data face without an institution responsible for their longterm care. just as professor curtice found an interaction between institutional and individual values in his research, the challenges facing data archives consist of similar competing values. consider the example represented by digital repositories. these systems are being built with good intentions to create better access. unfortunately, the preservation commitment, established practices and protocols of digital repositories are less developed and tend to be based on an end-state model of preservation rather than life cycle. they are being constructed in disciplines without roots in archival science and their proponents have initiated discussions using terms that confound issues of preservation. iassist members need to enter this dialogue to clarify key institutional principles about preservation. we need to educate the developers of digital repositories and the stakeholders in the research community that: · long-term access is dependent upon preservation, that is, access models that largely ignoring preservation will at best only provide short-term access; · today’s technology cannot replace the professional skills and knowledge of data archivists; · institutions are necessary to support preservation activities over generations of technology and researchers; · preserving research data requires the involvement and enduring commitment of all major stakeholders in the research community. these principles need to be apparent in the data preservation, documentation and citation standards that our profession develops as well as the services that we provide in our local institutions. furthermore, all disciplines now struggling with new requirements to provide access to research data created through public funds need to by aware of and to embrace these preservation principles. 4 the recent oecd ministerial declaration on this issue is an opportunity for iassist members to educate disciplines both within and outside the social sciences about the principles of preserving research data. what options exist to combat postmodern values threatening our data archives? we must be advocates for data archives individually and collectively through our affiliation with iassist and other professional organizations. internet technology supporting individualism in today’s culture should be used to promote institutional solutions for preserving research data. champions for data archives among the stakeholders in the research community need to be identified and supported. we need to be creative in transforming our national data archives so they remain relevant to the knowledge sector in our societies while ensuring that the functions they fulfill in preserving research data are not sacrificed or diminished. finally, we need to monitor attitudes toward data preservation constantly and to combat indifference toward institutions with mandates to preserve data. * chuck humphrey is head of data library at university of alberta. he was president of iassist 1991-1995. he can be emailed on humphrey@datalib.library.ualberta.ca footnotes 1 allison hanes, “media bans meet the internet age,” the gazette: montreal, april 7, 2005, p. a2. 2 charles humphrey, “preserving research data: a time for action” in preservation of electronic records: new knowledge and decision-making: postprints of a conference, symposium 2003, ottawa: canadian conservation institute, 2005. 3 the icpsr, which was created in the 1960’s, is an institution dedicated to the preservation of social science research data. 4 oecd declaration on access to research data from public funding, adopted on 30 january 2004 in paris. i^ssist newsletter vol.1, no. 1 at the lassist-ipsa panel, martinotti presented an analytic framework for the historical and present uses of process-produced information by governmental, administrative and research institutions and the implications of these data for the european data archives and of the growing demand for policy-oriented data. some of the implications he identified included the linking of different data bases, quality of the data, and the political and institutional nature of the consequences of publicly supported social science research. at the lassist-ipsa panel, erwin k. scheuch provided the session's attendees with an historical overview of the development of information systems for storage and management of large data bases. he described the problems encountered with these information systems, some of the large data bases which have been organized and the problems utilizing them, the relationship between the data collectors who have been primarily the public agencies and other communities such as researchers, and the relationship between the data archive and public agency. [editor's note: his paper will be forthcoming.] further input from action group members will help the group coordinators decide which projects are of most immediate interest, but tentative plans call for a continuation of the study of confidentiality laws; a preliminary pilot survey of available files in order to define the magnitude of this potential resource; a documentation and coding manual specifically for aggregate data; and, a survey of completed resenrch using process-produced data. paul mdller has suggested the following priorities: (1) an overview (survey) of process-produced data for each country of interest; (2) decisions and guidelines on which information should be put into machine-readable form; and, (3) recommendations to public administrators for preservation of materials for social scientists. it is likely that given the enormity of the tasks outlined by mdller and leavitt and other issues identified in the moller memo, these issues might well be handled by other action groups, but these matters are still fluid. this action group will be interacting with quantum, the us association for public data use, and other organizations to address subjects of mutual interest. data organization and management canadagreg morrison, social science data archive, department of sociology, carleton university, ottawa, ontario kis 5e6 europeeric tannenbaum, social science research council survey archive, university of essex, wivenhoe park, p.o. box 23, colchester, essex, england c04 3s0 united stateswilliam gammell, social science data center, university of connecticut, storrs, connecticut 06268 sist newsletter vol.1, no. 1 mandate this group addresses the problems of data base organization and management for effective and efficient analysis purposes. the action group will investigate and evaluate existing procedures for data and documentation preparation and data managei ment software and hardware capabilities. after evaluation the group will recommend : guidelines for preparation procedures and software development. workshops and semi inars will be sponsored for the transfer of information and professional training i in the organization, management, and use of machine-readable data bases. activities and plans the steering committee established this action group in recognition of the felt need of many potential lassist members for a forum to discuss matters relating to software for data management because this topic has received far less attention than data analysis and because other organizations have been established to address research needs. eric tannenbaum prepared a paper for the edinburgh meetings on "data preparation procedures in european archives," which addresses such topics as data file quality, cleaning classifications, data verification procedures, software and hardware availability, and organizational characteristics of the archive staff. this paper is based on the results of a questionnaire sent to the seven major european archives (norwegian social science data services, danish data archives, ssrc survey archive, zentralarchiv for empirische sozialforschung, steinmetzarchives, archivio dati e programmi per le scienze sociali, and the belgian archives for the social sciences. this action group will send a similar questionnaire to other archives and the resultant report on existing facilities and procedures will be of value to both existing and developing archives and libraries. at the iassist session at ipsa, bjfsrn henrichsen, norwegian social sciences data service, and terje sande, institute of sociology, university of bergen, presented a paper on software capabilities for mapping process-produced data. entitled, "computerized statistical mapping," the paper describes the adoption of an automated mapping system, using polyvrt, to deal with data structures, and symap and calform, for describing and shading of maps. discussion after the presentation concerned the interaction between norwegian official (governmental agencies) and the data services and the degree of success that the data services has had in convincing the agencies to utilize map producing systems. agenda for iass|st-|psa panel the following papers were presented at the lassist-ipsa panel, 20 august ' 1976. the panel was chaired by ivor crewe of the social science research council survey research archive, university of essex. carolyn geda , lassist chairperson served as discussant. copies of these papers can be obtained by writing to the authors. 1/1 rasmussen, karsten boye (2021), editor’s notes: iassist is now glocal., iassist quarterly 45(3-4), pp. 1-1. doi https://doi.org/10.29173/iq1025 iassist gone glocal welcome to the special double issue of iassist quarterly 2021 (iq vol. 45(3-4) 2021). iassist is an acronym. you may think that the word is the contraction of the two words 'i assist'. in my mind, you are right! whether the word iassist or the long explanation of seven words came first is the problem of the chicken and the egg. however, it is undisputed that when it is spelled out, the first i in iassist is for international. that has been so from its founding in 1974. having iassist members in usa, canada, and some (west) european countries was for a long time what we myopic westerners considered to be international. it is with great pleasure that iassist quarterly now presents a double issue from a regional workshop in africa. even in 2021, it is only a small number of iassist's members who are from regions not part of the western world. however, having a special issue from the african region is an important contribution to making iassist truly international. the phrase 'think globally, act locally' is a good framing of the compressed word 'glocal'. this special issue was compiled by guest editors winny nekesa akullo and robert stalone buwule, and they were also behind the africa regional workshop that took place at makerere university (kampala, uganda) on january 11 to 13, 2021. winny nekesa akullo works as head, library and documentation centre at public procurement and disposal of public assets authority in uganda, and is the iassist africa regional secretary. robert stalone buwule is senior assistant librarian at kyambogo university, also in uganda. the themes of the workshop addressed a world issue: 'data literacy as a catalyst for achieving sustainable development goals (sdgs)'. thus, global problems were addressed from a local viewpoint. great thanks to winny and robert for their lead in the arrangement of the workshop and extra thanks to them for collecting, editing, and making the papers of the regional workshop available to us all, and for making the regional international and the local global. enjoy the reading! submissions of papers for the iassist quarterly are always very welcome. we welcome input from iassist conferences or other conferences and workshops, from local presentations or papers especially written for the iq. when you are preparing such a presentation, give a thought to turning your one-time presentation into a lasting contribution. doing that after the event also gives you the opportunity of improving your work after feedback. we encourage you to login or create an author profile at https://www.iassistquarterly.com (our open journal system application). we permit authors to have 'deep links' into the iq as well as deposition of the paper in your local repository. chairing a conference session or workshop with the purpose of aggregating and integrating papers for a special issue iq is also much appreciated as the information reaches many more people than the limited number of session participants and will be readily available on the iassist quarterly website at https://www.iassistquarterly.com. authors are very welcome to take a look at the instructions and layout: https://www.iassistquarterly.com/index.php/iassist/about/submissions. authors can also contact me directly via e-mail: kbr@sam.sdu.dk. should you be interested in compiling a special issue for the iq as guest editor(s) i will also be delighted to hear from you. karsten boye rasmussen december 2021 https://doi.org/10.29173/iq1025 https://www.iassistquarterly.com/ https://www.iassistquarterly.com/ https://www.iassistquarterly.com/index.php/iassist/about/submissions mailto:kbr@sam.sdu.dk cnrs 22 iassist quarterly spring 2012 iassist quarterly abstract in europe, national legal frameworks frequently enable research access to official statistical data, also including detailed microdata, but cross-country research remains difficult. accreditation is a central element of the framework for access to data that currently is understood to be a barrier especially for trans-national access. to better understand the nature and causes of the problem, and to devise potential solutions, we have mapped current arrangements across european countries. we identify similarities and differences as well as major gaps and inconsistencies across countries, and we single out best practices and new, example-setting solutions. overall, our key results are encouraging: almost all european countries do provide research access to their microdata, and most of them allow non-national european researchers to access their data, though under varying conditions. some of the gaps that we have identified are relatively easy to fill, notably a widespread lack of online information, and unsystematic translation into other languages. a small set of issues, however, will require negotiation and coordination at higher, policy-making levels: the controversial need for institutional accreditation, homogeneization of terminology, and the possibility to introduce special provisions to facilitate trans-national access. some of these issues are under discussion today and some new solutions are being tested or piloted, so that substantial improvements can be expected in the future.. keywords: research access to data, researcher accreditation, trans-national access, highly detailed microdata, european research area. introduction today’s national legal frameworks frequently include provisions that facilitate researchers’ access to data produced by official statistical systems. however, existing solutions are mostly country-specific and data do not circulate easily across borders, even within the european research area (era) where regulations are very similar and a common framework on data protection applies. difficulties are particularly acute for access to confidential (or highly detailed, as they are sometimes called) microdata. despite progress in individual countries as well as at eurostat level, trans-border access to country-level official microdata is still patchy and especially difficult for highly detailed microdata, thereby strongly penalizing comparative research and research on europe as a whole. accreditation is a central element of the framework for access to data across borders that currently is understood to be a barrier to trans-national access. indeed, national statistical institutes (nsis) and other producers of official statistics maintain and recognize different procedures and practices in researcher accreditation, resulting in inequalities among researchers located in different countries in the era, a great deal of red tape, and a negative default position with respect to granting access across borders. more precisely, accreditation can be defined as the process defining the conditions under which a researcher access to official data and researcher accreditation in europe: existing barriers and a way forward by paola tubaro, marie cros and roxane silberman 1 iassist quarterly spring 2012 23 iassist quarterly willing to use official data can be considered a “fit and proper” person. in the eyes of nsis, it would mean comparability to official statistics staff members and subjection to the same rules and penalties. in practice, accreditation involves three main steps: • defining eligibility criteria: who is a researcher, what is research, what is a research project; • establishing application procedures: how to request access, what documentation and evidence to provide; • setting up a service level, including: designing rules for decision-making (who approves applications, on what basis), managing and monitoring the process, ensuring good governance and transparency. answers to these questions contribute to nsis’ risk management framework, defining the scope for safe research access to official microdata. the problem for researchers and data users is that these answers may differ across countries, depending not only on national legislative frameworks, but also on the internal policies and established practices of each institution. even within the same country or institution, there are discrepancies depending on type of data, mode of access, or status of the applicant. there is also some degree of variation due to fees that researchers sometimes have to pay for data provision, whether in the form of bespoke files or of secure access through hightech facilities. fees do not concern the accreditation process strictly speaking, but depend on the actual costs of services provided and are often set independently of the accreditation-granting authority. be that as it may, ambiguities and incongruities arise especially in less clearly-defined cases, particularly with foreign researchers and joint (typically, cross-border) projects. to better understand the nature and causes of these barriers and inconsistencies, and to devise potential solutions, we have set out to map the current arrangements in the different european countries including eligibility criteria, application procedures, and organisation of the service. we endeavour to detect patterns of similarities across countries, and to identify existing best practices of how to enhance access under relatively simple and straightforward conditions for data users, while still protecting the confidentiality of statistical units. on this basis, we discuss possible approaches for the future, along two main lines. firstly, we identify a set of simple and small-scale solutions, that may be easily transposed to a wide range of countries, and whose widespread adoption may lead to small, but tangible improvements that may make a difference for users even in the short run. secondly, in a long-run perspective, we discuss the extent to which a future common standard for accreditation may be considered as a realistic possibility, though perhaps a distant one, and we outline open questions and issues that require negotiation at policy-making level. our work on accreditation is part of the data without boundaries (dwb) project, funded by the european commission under its 7th framework programme for 2011-15, and aiming to support equal and easy access to official microdata for the era. it aims to map the current situation, identify and promote best practices, devise and pilot new solutions for remaining problems. focus is on trans-national access and on highly detailed microdata. the remainder of this paper outlines how we have undertaken this study (section 2), describes the major results we have obtained so far (section 3), and indicates directions for future development (section 4) methods to answer these questions, we have started with a discovery phase aiming to collect information on current researcher accreditation arrangements in the era (including both the eu and the eea countries). we have retrieved most of the information from public domain and secondary sources (particularly nsis’ websites) and existing literature (particularly tubaro et al. 2009; unece 2007). we have also obtained primary data directly from representatives of eastern european nsis, at a dedicated workshop we organized in bucharest, romania, on 23rd january 2012. we have organised the information into a working document (spreadsheet) for internal use to analyze results and to identify common approaches, existing workable solutions, and areas for improvement. in a subsequent consolidation and analysis phase, we have cross-checked this information for completeness, and have identified and revealed the key messages, particularly a list of best practices. the results presented below are the outcome of both our discovery and consolidation phases. to ensure comparability across countries, we focus on nsis only, leaving aside other public-sector data producers (such as iab in germany, or the bank of italy), and we consider nsi data at all levels of anonymisation, not limiting our analysis to confidential data. we emphasize national rather than european data, whose access is managed by eurostat and follows a specific set of rules and procedures. we focus on practices and procedures rather than legal principles strictly speaking, which are being investigated by another team in dwb. practices are often part of the “tacit” knowledge of nsis and are seldom shared or openly discussed, but have potentially strong and concrete effects, that a comparative study of legislative frameworks alone would be unable to bring to the surface. in general, laws are very similar across european countries and are often rather vague: sometimes entirely silent on matters of access to data for scientific purposes, sometimes authorizing access explicitly but without going into the details of who is a researcher, what is research, and how to establish that eligibility conditions are met. regarding highly detailed data, laws often simply state that security must be ensured. therefore, interpretation and practical implementation are even more crucial elements than legal rules themselves, in affecting actual conditions of access. to reveal similarities beyond apparent discrepancies, we strive to use a common terminology here, despite national-level variations in word choice, and nuances in meaning across countries. for example, we broadly distinguish between highly detailed (or confidential) and less detailed data based on disclosure risk only, disregarding the fact that levels of anonymisation and disclosure control may differ across countries, and that modes of access to data with similar disclosure risk may also diverge widely. our third and final phase involves dissemination, and is still in the making. we aim to develop the working document prepared so far into a searchable tool to be made available online, to facilitate data users’ search for information on accreditation in a comparable manner across countries. a repository of web pages will be built, each describing one nsi, and all linking to a data base of official statistical surveys available to researchers, that is being compiled in another part of dwb. although this tool is still in preparation, and its technical characteristics have yet to be finalized, we outline in the conclusions how it can contribute to making the results of our study more actionable, and to improving access by making accreditation conditions more easily intelligible across borders. 24 iassist quarterly spring 2012 iassist quarterly results on this basis, we have obtained a global picture of accreditation procedures and practices across europe, and we have identified similarities and differences across countries. we present them in the order outlined above eligibility, application procedures, and service and we conclude by examining in greater detail the specificity and additional problems that arise for trans-national access to data. eligibility a first key question is how to define a researcher who are the persons who, by law, can be entitled to have access to datafiles that may not be released to the general public. interestingly, european countries’ answers to this question reveal a great deal of commonality. the legal framework in most countries disallows release of data for commercial use and therefore, requires that all requests to access data are for research/study purposes (figure 1). affiliation to a research or higher education institution is sometimes considered as evidence of nonprofit research purposes and, especially for highly detailed data, it often also acts as an additional safeguard for the data provider. indeed some nsis require data users to be employees of a research institution, so as to involve the responsibility of the institution (and to be able to sue it in case of breach of their terms of use, particularly confidentiality rules). in some countries (germany) employment at a public research institution ensures subjection to the same codes of conduct as official statistics staff, and is therefore considered as a stronger guarantee against possible misconduct. the track record of the researcher (in terms of previous experience with microdata, publications etc.) is only required in some countries for access to highly detailed data (the uk’s “approved researcher” scheme for example), often allowing alternatives: for example students, who by definition have no track record, need instead some formal backing by their supervisors. while requirement of a research or study purpose is widely shared, the need for institutional support is much more controversial. the reason is that it is difficult to establish which institutions are eligible, all the more so as there are a growing number of ambiguous cases: publicprivate partnerships, analyses undertaken by research departments of non-research bodies (oecd for example), multi-institutional research consortia of limited duration, think tanks. another difficulty concerns the relationship between the researcher and the institution, often short-lived owing to people’s career moves as well as increasingly frequent fixed-term employment contracts (of post-docs for example). thus, a future shared system will have to carefully consider these issues and design flexible ways of ensuring institutional support, so as to accommodate for these cases, while not resulting in excessive bureaucracy and burden for data users, research institutions and nsis alike. applications the other major question is how to submit an application. an overwhelming number of nsis require a written application, even for less detailed data; online rather than paper submissions are more and more widely accepted. beyond this basic commonality, only about half of our sample has standard application forms: primarily large european countries (for example france, germany, italy, uk) that often make their forms available online. many smaller countries instead, require figure 1: eligibility criteria for research access in european countries (n = 28), absolute frequency. a country may have more than one (depending on type of data and status of users). figure 2: contents of application forms (main items), absolute frequency (n = 21). a country may have different versions of application forms depending on data types. figure 3: most common conditions in contracts between nsi and researcher / research institution, absolute frequency (n = 20). a country may have more than one depending on types of datafiles and modes of access. iassist quarterly spring 2012 25 iassist quarterly a written letter or email but do not have a standardised form (for example lithuania, poland, slovakia). regarding the contents of applications (figure 2), most countries require a research project, though the expected level of detail may vary. it is usually necessary to include title, composition of the research team, abstract, and a comprehensive list of the data and variables needed. for access to more detailed data, a more complete description of the project is typically required (france for example); in particular to better assess disclosure risk, nsis may ask applicants to indicate what analyses they plan to undertake or what statistical tools they intend to use. other elements vary more widely; for example if the applicant is a team, some countries are content with just one application by the team leader on behalf of the group, while others require each team member to apply separately. another variation concerns signatures: sometimes only the researcher or team sign an application form, sometimes an institutional representative is also required to sign (lithuania). finally, a condition often found in cases in which researchers receive data on their own computers (for example on cd-rom or through a ftp server) is to indicate how they intend to physically protect the data: for example using computers with passwords, keeping them in locked rooms, or avoiding copying data on laptops or portable devices (germany). in most cases (and almost always when data are highly detailed), if an application is approved, researchers are expected to sign a written agreement before actually starting using the data (figure 3). this may take different legal forms, from a end user licence to a contract, which we treat as equivalent for the purposes of this study. the most interesting aspect here is the commonality of key conditions, particularly use for research only, no transfer of data to any third party (or no access outside the approved research team), and confidentiality pledges, whereby the researcher undertakes not to attempt to identify statistical units and not to publish results in forms that may enable re-identification by others. notice that confidentiality pledges are common even when data do not present a very high disclosure risk. other conditions are specific to the type of access requested: for example when researchers receive the data for use on their own computers, they are usually asked to destroy the files at completion of the project; but this condition obviously does not apply when, instead, they use the data on the premises of nsis, or through a remote-access secure server where download is disallowed. in such cases, there may be other specific conditions, for example researchers much have their outputs checked for disclosure by nsi staff before being authorised to retrieve them from the system. service decisions on accreditation applications are overwhelmingly made by nsis themselves, or dedicated internal units within nsis, which also manage applications and monitor the whole process. in some countries, however, data archives managed by the research community take responsibility for applications concerning versions of datafiles that are not highly detailed or confidential: for example, issda does so in ireland. for highly detailed data, some countries have a dedicated authority or commission, such as comité du secret in france. representatives of researchers are occasionally involved in the decision-making process, possibly as members of a scientific council in charge of advising the decision-maker. though these solutions introduce further diversity in the european landscape, they empower researchers by involving them directly in the process, allow sharing costs between nsis and the research community, and usually result in greater openness, improved efficiency, transparency, and cost-effectiveness (see for example beagrie and houghton 2012, for the case of the uk). a major, widespread difficulty that hinders a smoother application process is lack of adequate communication to users and the general public. although all european nsis have a website, and all have an english version of (at least part of ) it, many of them provide only limited or no information about existing data, and about criteria and conditions for accreditation and access. even when this information is available, it is often difficult to locate through standard web search engines; what’s more, navigation within a single nsi’s website is frequently clumsy. these gaps, observed in many national-language websites, are typically exacerbated in their english translations. trans-national access can foreign researchers be eligible to access data too? we restrict our analysis to european researchers, who (regardless of their nationality or country of origin) live and work in one of the eu and eea countries, so that they are subject to very similar legal frameworks on personal data protection. figure 4 shows that in most cases (uk’s “approved researcher” scheme for example), they face the same conditions as national researchers and have to go through the same application procedure. in other cases (applications for highly detailed data in france, for example), they are entitled to the same data as national researchers, but may have to undergo some additional procedure or to produce additional evidence, for instance to prove the trustworthiness of their institution. other countries are stricter and allow foreign researchers to access some types of data files only: in particular germany cannot distribute scientific use files (that is, data at intermediate level of anonymisation, where disclosure risk is rather low, but not inexistent) outside its borders and requires foreign users to come to its research centres to use its data. interestingly, however, these limitations apply only to a small number of countries and most encouragingly, no country completely disallows research access to foreigners. nonetheless, a major bottleneck revolves around the extent to which institutions, not just individual researchers, need official accreditation. this is an obstacle especially for transnational access in that it is more difficult to ascertain the suitability of a foreign than of a national institution; more importantly, nsis fear the technical and legal difficulty, figure 4: conditions of access for foreign researchers (from eu and eea countries), absolute frequency (n = 22). one option for each country. 26 iassist quarterly spring 2012 iassist quarterly as well as the higher costs, of suing a foreign institution in case of breach. they thus often tend to be particularly cautious. other impediments to trans-national research access are subtle, and concern practicalities and procedures rather than general principles. as mentioned above, the english versions of nsis’ websites are less complete than national-language ones, further exacerbating the problem of insufficient information mentioned above. if it is difficult for a national researcher to find out what data are available and how to request them, it is even more challenging for a foreigner. more to the point, conditions for trans-national access are rarely spelled out explicitly, and applications forms, when they exist, do not always have an english version. further, there are nuances and differences in terminology that may make it difficult for a user to understand to what extent two apparently similar national datasets are comparable (in terms of available variables, degree and methods of anonymisation, etc.). finally, the architecture of nsis’ websites differs widely across countries, so that data access and accreditation issues are not always classified under the same headings, in a way that makes it difficult for external users to navigate through the european network of official statistics websites. discussion and conclusions overall, our key results are encouraging: almost all european countries do provide research access to their microdata, and most of them also allow researchers from other countries within the era to access their data, though under varying conditions. indeed the “open data movement”, progress in it, and pressure on nsis to extract maximum value from their data collections, have enabled major steps forward in the last few years. in particular, availability of secure it solutions for access to highly detailed data on the one hand, and increased production of highly anonymised public use files and tabulations for distribution through the internet, have allowed nsis to significantly increase their offer of data of value for research. the approach to dissemination of european nsis becomes more and more researcherfriendly, a significant evolution relative to the past. cross-country differences remain though interestingly, they mostly concern actual practices and concrete procedural aspects rather than general guiding principles. it is primarily because of apparently inconspicuous issues that trans-national access remains difficult: in particular differences in terminology and definitions across countries, and unclear or non-explicit rules for trans-national accreditation. some of the gaps that we have identified are relatively easy to fill: particularly the observed widespread lack (or incompleteness) of online information about accessible data and conditions for access, and uneven availability of english-language translations. these problems constitute a major practical obstacle for users, but can be cheaply and rapidly solved. to achieve this, we have identified a set of best practices in accreditation, that may provide guidance to all nsis on how to make progress even in the very short run: • availability of complete english translations of nsi websites, particularly the pages dedicated to data access and accreditation; • adoption of a more common terminology, use of more similar website structures or of indicators and visual clues that help users to more easily locate information on data access and accreditation; • clarity and completeness of information on both general criteria and any special conditions (in particular for trans-national access, but also for a number of other less clear-cut cases such as students’ access); • clarity and completeness of information on how to apply (including prices, if any, and expected timing); • standard application forms (rather than a more generic request of a written letter) with english translations, ideally downloadable from the web and allowing both online and email submission. the database that we are building as part of dwb, and its future release through the web, are meant to contribute to this process by further improving the capacity of the system to communicate and maintain openness, transparency, and readability. other issues, however, cannot be solved in the short run and require negotiation and discussion at policy-making levels. a first major question, extensively discussed above, is whether institutions need an accreditation procedure together with individual researchers’ accreditation, how to recognise foreign institutions, how to assess unconventional or short-lived institutional partnerships, how to account for short-term employment contracts and researchers’ career moves across institutions (and sometimes across national borders too). a second major question concerns terminology, as much of the observed lack of clarity is due to heterogeneity of definitions and denominations. while adoption of a common terminology across europe may be too ambitious (not least because it would involve appropriate translations into multiple national languages), perhaps a more easily applicable solution would be to create some code to “interpret” the notions defined at national level, and guide users to identifying the closest notions in different countries: for example, to understand the extent to which the “public use files” of a country are really similar (for degree of anonymisation and conditions or modes of access) to those that another country may call “campus files” or “anonymised microdata files”. such a solution would also involve a great deal of preparatory work, but may be easier to implement for individual nsis, without requiring them to radically change their operating modes. a third major question is trans-national access and how to improve it at european level, so as to facilitate comparative research. several options are possible. the least demanding one would require some form of mutual recognition of accreditation decisions, or perhaps some simplification of the application process in country b, if a researcher has already received accreditation for an equivalent dataset in country a. even in this case, though, careful attention will be needed to design procedures that do not discriminate among researchers, keep track of data usage in different countries, and allow smooth and continuous communication between the nsis involved. a more ambitious option would be to devise a common standard for all member states with the same criteria for eligibility, shared application procedures and forms, and possibly even similar organisation of the service level. dwb is actively engaging in discussions about a future possible standard, and an emerging concept from these discussions so far is the ambition of a “schengen area” for researchers. as a long-run goal, this would be achieved first by harmonisation of criteria, conditions, and procedures for granting researchers access to confidential data. once a high degree of harmonisation is achieved, the potential for integration of researcher accreditation opens up. why maintain more than one procedure in the era if all procedures are the same and deliver the same results? one manifestation of a “schengen” for researchers is the idea of a “researcher passport” whose definition, applicability and usefulness, however, are still to be assessed. its creation might be integrated within a “european service centre for official statistics” (esc-os), also proposed within dwb, which would centralise information and handling of procedures iassist quarterly spring 2012 27 iassist quarterly and applications on behalf of participating countries (mack, wolf, esteve and silberman 2012; tubaro, cros, kleiner and silberman 2013). it is yet unclear which of these solutions will be preferred by nsis and other stakeholders (particularly data archives and of course, the research community), all the more so as different countries may have different preferences. another issue to be considered very seriously is the funding of these activities, all the more so as the budgets of most european countries are currently under strong pressure. while some of the changes we have identified can be made at relatively low cost (improvements in information in particular), other changes are more demanding. in truth, accreditation procedures are not in themselves very expensive, and any rationalization is likely to further lower down costs; but an improved and smoother accreditation process may globally increase researchers’ demand for access, generating additional costs related to service provision – that is, preparation, documentation and secure delivery of data. one solution would be to involve the social science research community and share the burden, as mentioned above, particularly through data archives such as those that are part of the council of european social science data archives (cessda). the experience of some countries demonstrates that this solution is not only cost-effective, but may even generate high returns to public investment, as recently shown in the case of the uk (beagrie and houghton 2012). in addition, enhanced coordination within the european statistical system, with the leading role of eurostat, may also contribute to reducing some costs, particularly those related to provision and dissemination of information in a comparable and consistent way across countries. in the years to come, dwb will continue its involvement in discussions in the hope to facilitate an improvement and possibly, a standard european model and a collaborative approach between different stakeholders to facilitate accreditation and access to official statistics. references beagrie c. and houghton j. (2012).economic impact evaluation of the economic and social data service, report for the economic and social research council. available at: http://www.esrc.ac.uk/_images/ esds_economic_impact_evaluation_tcm8-22229.pdf mack a., wolf c., esteve a., and silberman r. (2012) report on concept for and components of european service centre for official statistics. available at: http://www.dwbproject.org/export/sites/default/about/ public_deliveraples/d5_1_european_service_centre_report.pdf tubaro p., cros m., kleiner b., silberman r. (2013). accreditation for transnational research access to official micro-data in europe. proceedings of the new techniques and technologies for statistics (ntts) conference 2013. available at: http://www.cros-portal.eu/sites/ default/files/ntts2013fullpaper_145.pdf tubaro p., silberman r., cros m., cornilleau a., kvalheim v., kiberg d. and farago p. (2009). audit of access mechanisms and official statistics in the european research area, technical report d10.1 for cessda ppp. available at: http://www.cessda.org/project/doc/d10.1_audit_of_ access_mechanisms_and_official_statistics.pdf unece task force directed by d. trewin (2007). managing statistical confidentiality and microdata access. available at: http://www. unece.org/fileadmin/dam/stats/publications/managing.statistical. confidentiality.and.microdata.access.pdf acknowledgment this research has been undertaken as part of the “data without boundaries” project, which has received the financial support of the european union’s seventh framework programme (fp7/2007-2013) under grant agreement n° 262608. we thank members of work package 3 for relevant inputs and useful suggestions notes 1. corresponding author: paola tubaro, university of greenwich, room qm163, park row greenwich, london se10 9ls. email: p.tubaro@ greenwich.ac.uk. paola tubaro, university of greenwich, cnrs marie cros, université de lille i roxane silberman, cnrs réseau quetelet iassvol201 6 iassist quarterly introduction the training of librarians in data related issues has become an educational imperative emanating from the technological advances experienced globally. the dramatic strides in technological developments are causing major expansion in the way that end-users access information. this sudden explosion of information has caused a shift from the print to the electronic mode, which in turn has led to a shift from ownership to access. academic librarians make information accessible for research in academic libraries by placing at the individual’s disposal vast resources of information to satisfy information needs. academic librarians are thus challenged to redefine their service philosophy. a reconsideration of roles and mindsets are the profound changes that will have to be effected. parallel to the developments in technology, the political changes and the unfolding transformation process in south africa necessitates academic librarians to reconsider their information provision strategies to redress imbalances. statements of objectives in this paper i will focus on the following: (i) the impact of technology on academic libraries; (ii) the impact of technology on the client of the academic library, with special reference to the university of the western cape; and (iii) the changing roles that academic librarians are required to play. the abovementioned issues will be contextualized within the framework of the political changes occurring in south africa. clerification of concepts for the purpose of this paper, data is defined as facts, statistics or information that can be analysed. knowledge is defined as familiarity gained by experience. it is also viewed to be a person’s range of information or theoretical or practical understanding. information is viewed as desired items of knowledge. one notes a diversity in the definition of information. this is in view of the fact that it is deemed to be intangible and only encountered operationally by its effects. it is suggested to be derived from data, and can be accessed through print and computerized machinery for utilization (zorckoczy, 1988: 11). training as used in this paper refers to the concern with making the best use of the human resources in an organization by providing them with the appropriate instruction to acquire the necessary skills for their jobs (statt, 1991: 154). impact of technology on academic librarries internationally, the age of computers have certainly impacted on the traditional manner of doing things in a library. the dynamic nature of information generation, management and use, as well as the proliferation of publications, force the library environment to either adapt or die. microcomputers have streamlined the operational workflow of routine functions as well as enhance the online search process. cataloguing and circulation automation are providing more effective service and better control over collections. computer technology effected various ways of accessing information more speedier through networked electronic mail facilities. it now allows librarians to automatically logon to local and remote systems and download search results for later printing. this is causing a major shift from libraries’ provision of information from own collections to that of access to remote regional, national and international collections, databases and networks. advanced technology also gave rise to the virtual library a concept used to denote remote access to the contents and services of libraries and other information resources, combining an on-site collection of current and heavily used material in both print and electronic form with an electronic network which provides access to, and delivery from, external worldwide library and commercial information resources. this is a transformation with dramatic results for research which cannot be ignored. computer technology has indeed made it possible for libraries to establish networks, based on co-operation and resource sharing, with each other and with other information centres. the concept of networking, understood to be the building of contacts among professionals, has been given new meaning by the academic library’s utilisation of technology. here one thinks especially of the internet, which is seen as the network of networks. it brings together people, databases, and networks. through technology eager use is also made of periodical indexes and reference works on cd-roms (compact disc read-only-memory). it is interesting to note that this transformation, brought about by technology, does merit a lot of adjustment. this is demonstrated in library school curricula which, today, include a module on information technology or a complete course on computer science. it is true to say that the increasing complexity and sophistication of information technology requires a high degree of specialized technical the need to train librarians in data related issues by julia dawn paris1, university of the western cape 7summer 1996 knowledge. this requirement has far reaching implications for the future training of academic librarians. academic librarians should learn how to exploit new technology to the benefit of their user’s information needs. they have a responsibility to require the knowledge and skills necessary to use and teach the most efficient information techniques the current technology makes possible. a proactive and flexible manner is also required if academic librarians want to take up the challenge of technology, to interact and communicate effectively with clients who may not only be unfamiliar, but also uncomfortable with information technology. therefore, training and retraining should be high on the strategic priority list of academic libraries. in south africa, academic librarians and end-users experience the convenience of information made easier and speedier by sophisticated machinery. on the other hand, they experience difficulty in the accessing of the information and information overload. information overload occurs when locating too much information on the given topic. this dilemma creates a feeling of being overwhelmed and overloaded and this could result in a situation of frustration and anxiety. if a user is not able to understand the automated system used for an online catalogue at any academic library, or the systems potential to give information, or know how to sift and sort through the plethora of online information, it immediately results in information poverty by creating a barrier between the user and the information needed. hence, information technology becomes a problem when it deprives users of information, especially if the information is to satisfy basic needs as experienced by the majority of people in south africa. south africa is a ‘new kid on the block’ in the development of technology and the application thereof. estimations claim that information technology is developing at an annual rate of 35% in south africa. this is viewed to be very slow in comparison with developed countries where information technology is doubling, if not tripling, at that rate every four years. my major concern is with the extent of the impact of this development of technology, however slow, on the information impoverished people of south africa. the south african academic libraries, and in particular the university of the western cape (uwc) have been influenced by the global information explosion. the access of our users to local, national, international databases and networks, has made it necessary for the librarians to take an objective look to what is happening. i agree with makhubela (1995: 15) who states that “though this situation is not necessarily unique to uwc, it becomes especially critical, given that many students at uwc come from economically deprived and disadvantaged communities”. the realization then, that technology is affecting the way that librarians provide access to information should become a catalyst for change. however, outdated mindsets and curricula make it downright impossible to take up the challenges posed by the shift and increased user expectations. challenges facing librarians in south africa are not only caused by the technological developments, but also by the major political change from an oligarchic, apartheid society to that of a democratic one. great educational inequalities, major illiteracy problems, lack of a reading culture, lack of school libraries are some of the stark realities academic librarians are faced with. the nationalist apartheid government fostered a library and information service characterized by a traditional approach as was followed by many of the developed countries. it was focussed to benefit the educated user community, was literacy based, with collections comprising of predominantly books, and an emphasis on facilities and collections instead of the needs of users. in a nutshell it served the needs of the dominant white culture and class (nepi, 1992: 54). the transformation process now underway calls for the library and information service to also undergo transformation and address the social responsibility of libraries. it calls for a shift in emphasis from collections to needs of users. it also challenges librarians to use information as a source of development and empowerment in the education of all those served. in the academic environment librarians need to learn to facilitate the learning environment of disadvantaged students even if it means teaching them step-by-step how to use the technology in its basic forms e.g. opacs, and later teach them how to access cd-rom databases and other remote online networks through information literacy programmes. the impact of technology on clients industrial progress requires a versatile and skilled labour force. this could be effected by well planned education and training strategies. information needs vary according to levels of education and socio-economic status. it is important to assess what it is that information end-users need and want. to end-users accessibility of information is one of the primary issues of concern in satisfying information needs. no matter how organized the information, it will not realize its value until it is made known and put to effective use. a generalization reflected in international literature is the idea that undergraduates usually labour under time constraints. also, that their information needs are most often met by the collections of their own institution. further, that depth of reference queries are not so complex and that less reference assistance is required. this argument definitely assume a situation where undergraduates have been exposed to adequate education, school libraries, and have a sound economic background. at an institution such as uwc, a lack of factors noted above, combined with problems of language, render undergraduates incapable of having a basic understanding of their research topic under investigation. minimum or no knowledge of basic reference sources, inability to verbalise information needs are other shortcomings encountered by academic librarians. 8 iassist quarterly post-graduate and faculty needs are viewed as extending beyond the resources of their own institution. the research of these clients represent unique contributions to the knowledge base of disciplines. therefore, their demands on reference services are more specific and exhaustive. for them the academic library extends into the resources of other information centres through inter-library lending services. with the advent of online searching on remote databases and networks ‘ the world become their campus’. the appearance of computer accessible information databases have changed the way researchers search for information. some scholars rely almost entirely on computerized sources for current citations to published research in progress. this is creating great expectations toward academic libraries to provide the necessary mechanisms to access these sources. at uwc our post-graduate and faculty needs do not differ as widely as that experienced internationally. however, with post-graduate students the librarians still have the responsibility to teach the most efficient information techniques the current technology in the library makes possible. i wish to reinforce benson’s (1995: 57-69) utterance when he said that “our enthusiasm about the new technologies should not outstrip our ability to provide adequate service to those for whom we acquire the services and may cause us to overlook real needs”. i firmly believe that we have a responsibility to be sensitive to research needs of the entire academic community we serve. to reflect this sensitivity, we need to conduct use and needs studies to ensure acquisition of the databases and networks. changing roles of academic librarians traditional role of librarians previously, librarians emphasised preservation of books, making them accessible through cataloguing and classification, and offering a client advice and guidance service. librarians were only required to know where information could be found rapidly to answer the client’s need or questions. librarians were viewed as custodians of collections of library material instead of proactive individuals. this situation is still prevalent in many academic libraries in developing countries. new roles of librarians in the developed countries, the paradigm shift from traditional to electronic libraries have changed the way librarians execute their duties today. the expectations created by this paradigm shift resulted in the need for librarians to adopt a new vision and a new mission to meet users information needs. they are faced with new issues, and new challenges brought about by the increasing use of information technology in the library environment and the rapid changes occurring in this area, understanding of what librarians duties are and the new attitudes that should result from this change in mindset. there is also the added challenge of helping clients to manage the information overload that came about as a direct result of this new era. clients must be assisted to sort through the wealth of information and make wise decisions about what they need. there is a challenge posed at academic libraries to shift the emphasis from physical documents to individuals with needs; from document delivery to information management and transfer; and from question-answering to problem-solving. librarians are challenged to move beyond quantity to quality. a general view seems to be that although this shift has occurred, librarians continue to be more concerned with delivery of documents and have not started to focus on the delivery of contents or the data and information contained in the documents. my contention is that this is due to the failure to grasp that there is a need for this new vision. there should be a move towards acceptance of the facts that major paradigm shifts are creating new definitions of what librarians roles are to be in the technological library environment. we face expanded demands on our time and skills and are afforded a tremendous opportunity to assist students and faculty to solve their research problems. requisite skills that need to be acquired include: competencies ranging from communication skills to knowledge of database searching techniques; sifting and analysis of information; and the reduction of the amount of information provided to the user. larry benson (1995: 58) postulates that “as attitudes toward the use of automated systems change, so must the roles of academic librarians. they face expanded demand on their time and skills and have a tremendous opportunity to assist students and faculty to solve their research problems”. the academic librarian is called upon to serve in overlapping roles namely as information specialist; teacher/trainer; and as consultant. depending on the library’s priorities, elements of all three roles will enter into the professional librarian’s position. as information specialist, the librarian will be called upon to provide adequate access through adequate resources for optimum usage. as teacher, the librarian will have to teach library, information, and technology literacy skills. this includes the teaching of critical thinking skills to assist clients to become active, independent and confident users of information technology. as consultant the librarian is called upon to participate in academic curriculum design and assessment projects and to provide expertise in the selection, evaluation and use of materials and emerging technologies for the delivery of information and instruction, as well as translating curriculum needs into academic library program goals and objectives. a reluctance to assume this role would deny the faculty and administration the benefit of the valuable insights academic librarians have in these areas. it would also minimalize the importance of the academic library to academic outcomes. this expanded role of the academic librarian, brought about by information technology has great potential for improved educational outcomes. at uwc, the role played by academic librarians has not 9summer 1996 changed very much. the new issues and challenges of technology have not resulted in a radical change of roles yet. in fact, the technology applied in information provision and access have not been exploited to its fullest potential. lack of resources have made it impossible to initiate a process of exposing end-users to networks like the internet. language problems and other cultural barriers add to the existing complexities of users social backgrounds. the predominant roles thus played are primarily those of instructor in library skills, information provider, faculty liaison and collection developer. lack of adequate staffing and specialized skills contribute to a work overload which are some of the factors slowing down the move toward the real role of facilitation and empowerment. recommendations for training strategies: the training of librarians should be understood within the broad framework of human resource development and capacity building within the transformation paradigm of south africa. in order to play a proactive role in facilitating information acces and its proper use, it is imperative to look at the role librarians can and should play. the suggestions posed in this section are not proposed as solutions, but should serve as useful information that could assist in appropriate applications of training strategies. any training or retraining embarked upon by librarians in south africa, should be done in line with the needs of those wanting and needing the services. the bottom line for training should always be improvement and efficiency of services to clients. this is especially seen against the background of cultural diversity within the population as well as their previous exposure or non-exposure to information. reference librarians need to be equipped to become more sensitive to cultural diversity and in this new technological environment develop more effective communication skills to understand the multi-cultural aspects of information seeking behaviour. in order to make a difference, these aspects should be inculcated through the curriculum design at library school entrance level. continuous education efforts to ensure lifelong learning in multi-cultural information provision and the evalution of performance should also become a priority. the majority of users of the historically marginalised institutions come from a background of poverty, illiterate parents, overcrowded homes, lack of proper housing, lack of reading facilities, oral tradition and lack of proper schooling, to mention only a few. for these users information is geared towards coping to survive. librarians need to be equipped to understand this dilemma experienced by the majority of users and to devise strategies to effectively address it so that they become instrumental in helping users. information technology and other mechanisms should be applied to enhance the development process of information seekers in south africa, not impede it. librarians must be equipped to break down the barriers which hinder information access. the unfolding transformation process brought about a new dimension of complexity in the form of all the challenges facing academic librarians. librarians should provide an unthreatening atmosphere in which users could access the necessary information. librarians should also become instrumental in computer aided instruction programme development. training must include elements of basic data analysis, networks, and network search commands, the nature of interfaces, search software of specific social science and other databases. general training in the nature of all databases and networks is imperative. training in end-user instruction, to ensure optimum utilization of these costly resources, is necessary. in order to instruct end-users, librarians themselves need to be trained in how to access, search and interpret different databases. the intention with strategies mentioned should be to use existing resources. the challenge is placed before tertiary institutions and other academic alliances to respond to the training and continuing education needs of the academic librarians by transforming themselves in the way they operate. this statement is made in view of the fact that there is considerable competence and skills existing within the various alliances of the library and information service, but outdated principles regard users as passive consumers of information instead of an active participant in the information process. within the south african situation, there are no simplistic answers. here academic librarians face formidable challenges, but also lots of opportunities to make a difference. training should also address all levels and needs in terms of how to supplement current awareness of information technology, how to do retrospective searches and how to use technology in the social science and other related fields. basic training is required to increase the awareness of people about the possibilities offered by the new technology and teaching them effective ways to make use of the equipment. this is the equivalent of teaching people how to use the computerized catalogues. a second kind would be training needed to become fully familiar with a complex word processing or database software package. if the training is neglected, the full benefits of the facilities will not be realized and the drawbacks will be magnified (boston,g,1994:331-337). the introduction of new electronic services has implications for virtually all library staff. inevitably, as electronic information systems expand, workloads will change and resources, both human and financial, will need to be redirected. successful implementation and user satisfaction in the electronic environment, will only be achieved by a highly trained and skilled library workforce. these will 10 iassist quarterly include increased knowledge of automated systems and electronic communications, as well as the hardware such as workstation networks. furthermore, it is important that information workers continue to be active in influencing the newly emerging national and international standards for electronic information, e.g., standards for bibliographic control and for citing articles. “some electronic network travellers will fearlessly set off on their own. others will still rely on librarians to drive the tour bus” (woodward, 1994: 44). in the uk, the library association is encouraging librarians and information workers to analyze the contents of online information systems, databases and cd-roms in the same way that librarians have critically analyzed the contents of reference books for accuracy and up-to-dateness. the information era, as previously noted, puts before us more information than can be consumed and used. the sifting and analysis of information, and the reduction of the amount provided to the user, is an equally important responsibility that is based upon the traditional skills of information workers. the quality and standards of service provision has become an important factor too. agreed and published standards or guidelines are mechanisms to be used, because they are important factors toward improvement in the provision of services. the professional librarian also has the responsibility to maintain his/her level of knowledge, skills and expertise given the speed of change enhanced by the information technology and telecommunications developments. it is our responsiblity as well as those of employers to ensure that the knowledge base and expertise of information workers is kept up-to-date. it should be made mandatory by library associations. literature suggest a code of ethics or conduct as the machinery to exert pressure, should information workers fail to meet the responsibilities as set out above. in training, librarians should be charged with learning the structure of online libraries, the mechanics of searching, and the relationship between database content and user questions. librarians should be trained not only to focus on specific databases within their areas of specialization, but also to look at general news databases which transcend subject discipline. to understand this magnitude of online library collections, a major conceptual shift in thinking must occur. the ability to search full-text articles makes the service unique, and poses special challenges. the importance of constructing precise search statements, critically selecting words and utilizing appropriate commands to narrow the retrieval are some of the special instructional challenges for reference librarians (adalian & rockman, 1995: 99-113). identifying, locating and using networked information can only be accurate with the help of an intermediary specialized information workers who are the links between the information seekers and the information in the databases. in the networked environment it is easy to become suddenly overwhelmed with a large volume of information. in the absence of user driven filtering of information librarians are called upon to play the role of filtering and interpreting the information for those who choose to use libraries. stakeholders to assist in training the many challenges posed to academic librarians cannot be carried by them alone. what will be required, are the effective partnerships of all major stakeholders within the field of library and information services. training structures for librarians already exist, but should be effectively harnessed for continuing lifelong education. stakeholders who should play a vital role in the education and re-education of academic librarians in south africa are the african library organization (alasa), the inter resource forums, the library and information workers’ organization (liwo) and the south african institute of librarianship. these organisations should pool their resources and professional skills to provide a vehicle for information workers throughout south africa to benefit from a broad range of professional expertise. other stakeholders identified are: the vendors of technological and commercial products by giving basic training in their software; library schools by updating and redesigning curricula as the changes in the information provision environment occur; university and academic library administrators through strategic planning of inservice training programmes in word processing, database-, online network searching and information management skills; and finally government departments who should make the necessary funds available to acquire the technology and training facilities needed to empower and build the capacity of academic librarians. conclusion the training of academic librarians in data related issues is crucial if they want to play the new role of facilitator in the technological environment of the academe. to be able to do so academic librarians have to leave their comfort zones as custodians and become proactive role players in the information provision for development and progress. it is my belief that south africa can learn valuable lessons from the international experience. however, priorities for training need to be constructed in accordance with the prevailing needs and conditions in south africa, and not simply in relation to abstract theories or first world research findings. references adalian, paul t & rockman, ilene f. (1995) issues in implementing and teaching the lexis/nexus services in an undergraduate library. the reference librarian, 48: 99-113. benson, larry d. (1995) scholarly research and reference service in the automated environment. the reference 11summer 1996 librarian, 48: 57-69. boon, j.a. (1992) information and development towards an understanding of the relationship. south african journal of library and information science, 60 (2): 63-74. boston, george (1994) new technology friend or foe. ifla journal, 20 (3): 331-337. bowden, russel (1994) professional responsibilities of librarians and information workers. ifla journal, 20 (2 ): 120-129. chatman, elfreda a. & pendleton, victoria (1995) knowledge gap, information seeking and the poor. the reference librarian, 49/50: 135-145. diaz, karen r. (1994) getting started on the net. the reference librarian, 41/42: 3-24. faries, cindy (1994) reference librarians in the information age: learning from the past to control the future. the reference librarian, 43: 19-28. figueiredo, nice (1992) information as a tool for development. the international information and library review, 24 (3) sept.: 189-201. gerber, p.d; nel, p.s. & van dyk, p.s. (1994) human resource management. halfway house: southern books publishers. hopkins, richard l.(1995) countering information overload: the role of thelibrarian. the reference librarian, 49/50: 305333. in focus (1995) human science research council bulletin, feb/march: 37-43. library and information services. (1992) report of the nepi library and information services research group: a project of the national education co-oordinating committee cape town: oxford university press. liu, mengxiong. (1995) ethnicity and information seeking. the reference librarian, 49/50: 123-133. makhubela, l & koen z. (1995) another angle on access: information literacy and student learning. academic development, 1 (1): 13-19. mccallum, sally h. (1994) connectivity for libraries and information services introduction. ifla journal, 20 (2): 145-146. mchombu, kingo. (1991) which way african librarianship. international library review, 23: 183-200. pask, judith m & snow, carl e. (1995) undergraduate instruction and the internet. library trends, 44 (2) fall: 306317. sisson, loren & pontau, donna (1995) the changing instructional paradigm and emerging technologies: new opportunities for reference librarians and educators. the reference librarian, 49/50: 205-215. statt, david (1991) the concise dictionary of management. london: routledge. summerhill, craig a. (1994) connectivity and navigation: an overview of the global inter-networked information infrastructure. ifla journal, 20 (2): 147-157. sykes, j.b. (ed.) (1982) concise oxford dictionary. oxford: clarendon press. weibel, stuart r. (1995) the world wide web and emerging internet resource discovery standards for scholarly literature. library trends, 43 (4) spring: 627-644. westerman, mary. (1989) computers in academic libraries. health science libraries micros and medicines. catholic library world, 60 (4) jan/feb: 150. woodward, hazel. (1994) the impact of electronic information on serials collection management. ifla journal, 20 (1): 35-45. zorckoczy, p. (1988) information technology. oxford: clarendon press. 1. this paper has been presented at the css96/iassist conference at minneapolis, university of minnesota, may 12 19, 1996. julian dawn paris, subject librarian economic and management science, university of the western cape, private bag x 17, bellville, south africa. iassist quarterly summer 2010 5 iassist quarterly editor’s notes new format new papers walter piovesan – our publication officer – had a biking accident. to show that nothing is so bad that it is not good for something, walter used his recovery time to redesign the iq. i hope you like it. i like it! furthermore, walter is part of the local arrangements committee, along with mary luebbe, in charge of the coming 2011 iassist conference, so he is a busy guy. and i’m happy to say that walter is now fully recovered. welcome to the iassist quarterly volume 34 (2) of 2010. this issue of the iq features the following papers: rein murakas and andu rämmer from the estonian social science data archive (essda) at the university of tartu describe in their paper “social science data archiving and needs of the public sector: the case of estonia” how the archive had a historical background in the empirical research of the soviet union. after estonian independence the creation of the social science data bank acquired momentum and the data archive was established in 1996. the following year the data archive became a member of the european council of social science data archives (cessda). this paper was presented at the iassist 2009 conference in the session “building data archives and user communities: greece, estonia and ethiopia”. from the historical background we move to web 2.0 in a paper by angela hariche, estelle loiseau and philippa lysaght on “wikiprogress and wikigender: a way forward for online collaboration”. the paper was presented at iassist 2010 conference at cornell university, ithaca. the authors are working at the oecd and the paper’s statement is that “collaborative platforms such as wikis along with advances in data visualisation are a way forward for the collection, analysis and dissemination of data across countries and societies”, and their paper “presents two wiki platforms that can support such a vision: wikiprogress and wikigender”. these applications support the oecd’s shift from measuring (only) economic production to measuring well-being and the general progress of societies. the third paper addresses an issue of central importance for most data archives. the question concerns balancing data confidentiality and the legitimate requirements of data users. this is a key problem of the secure data service (sds) at the uk data archive, university of essex. the paper “secure data service: an improved access to disclosive data” by reza afkhami, melanie wright, and mus ahmet was presented at iassist 2010 in the session “secure remote access to restricted data”. the sds will allow researchers remote access to secure servers at the uk data archive. in the paper many related issues in legal, contractual, educational, and technical areas are discussed and examined. on top of this support the sds may potentially open access to previously closed data and thus further facilitate new research. the last article has the title “a user-driven and flexible procedure for data linking”. the authors are cees van der eijk and eliyahu v. sapir from the methods and data institute at the university of nottingham. the paper was presented at iassist 2010 in the panel on “virtual research environments: tools for presenting and storing data”. the data linking relates to research combining several different datasets. the paper defines the relational database operations that can be used as flexible and end-user directed choices. the implementation is developed for the piredeu project in comparative electoral research. the authors are combining traditional survey data with data from party manifestos and state-level data. the paper exemplifies the linking of voter and manifesto studies, and also addresses more complex data linking and merging. articles for the iq are always very welcome. they can be papers from iassist or other conferences, from local presentations or papers especially written for the iq. if you don’t have anything to offer right now, then please prepare yourself for the next iassist conference and start planning for participation in a session there. chairing a conference session with the purpose of aggregating and integrating papers for a special issue iq is much appreciated as the information reaches many more people than the session participants and will be readily available on the iassist website at http://www.iassistdata.org. authors are very welcome to take a look at the description for layout and sending papers to the iq: http://iassistdata. org/iq/instructions-authors authors can also contact me via e-mail: kbr@sam.sdu.dk. should you be interested in compiling a special issue for the iq as guest editor or editors i will also be delighted to hear from you. karsten boye rasmussen march 2011 editor vol282-3.indd 4 iassist quarterly summer/fall 2004 editor’s notes i am pleased to be able to present this special double issue of the iassist quarterly (iq issues vol. 28-2 and 28-3). we have eight articles on the subject of “data literacy”. this subject has been the focus of many iassist efforts and the enhancement of data literacy is the daily work of many iassist members. these members are making a great contribution to the building of competences for evaluation of information. furthermore, i am glad to announce that these articles have been collected by some very active chairpersons: louise corti from the uk data archive at the university of essex and wendy watkins from the library data centre at the carleton university library. the work of compiling articles for the iassist quarterly is not always the isolated effort of the iq-editor. others are very much welcome to do special issues like this one. great thanks to louise and wendy. karsten boye rasmussen, june 2005 guest editor’s notes this special issue of iq brings together a collection of papers centred on the important topic of literacy as it relates to data. as we shall see from the contributors, a number of terms are often used: data literacy, information literacy, statistical literacy, quantitative reasoning. these concepts are different, yet they intertwine. but they all point to one key element: the importance of acquiring skills for evaluating ‘information.’ literacy means having the knowledge, skills and experience to tackle the interpretation of research fi ndings and confront data from an informed and critical persuasion. the eight contributions presented here arise out of iassist annual conferences held over the past two years. the introduction of data in the classroom and the teaching of statistical literacy has become a growing concern for data librarians, and this area now features as an increasingly popular theme at successive iassist events. in 2004, at the madison wisconsin conference, a session on developing statistical literacy: think globally, work locally was held while in 2003 in ottawa, two sessions were devoted to advancing research and data literacy: empowering users and understanding the strength of numbers: statistical literacy. data librarians are at the front line students hungry for data cannot completely rely on their own departments for guidance in more complex data or analytic matters. we have all experienced the fl urry of queries at dissertation time, that go well beyond the basic question of how do i download this dataset or open this codebook? frighteningly basic questions plague our help desks on a daily basis: help, i have downloaded data but have no clue what to do next. what is spss? what is the difference between a variable and value? i have 468 cross tabs, all which have a p-value of .5 what does this mean? data librarians have a role to play in helping facilitate the use of real data in the classroom, and to help support faculty to improve data and statistical literacy. but, for the data librarian, how and where does such accumulation of knowledge and skills take place? all the papers here describe particular ways in which the challenges of improving statistical literacy have been addressed at an institutional or regional level. they provide useful and appealing case studies that have been trialed to help educate new students and lecturers, supervisors, mentors and data support staff who do not feel entirely competent in working with numeric data. in the fi rst paper, milo schield considers, specifi cally, the challenges of teaching data and statistical to students, and how data librarians can play a role in this education. he discusses the meaning and interaction of the terms information, data and statistical literacy and how they all engender critical thinking skills. karen hunt goes on to raise the key challenge, highlighted in all the papers, of how to help students develop information literacy skills. she discusses her own experiences, as the university of winnipeg information literacy coordinator, of attempting to help integrate data literacy into the subject curriculum through collaboration with teachers. what are the best practices for developing data literacy and what aspects can be applied from the information literacy fi eld? again, the importance of training for data librarians in how to both promote and teach data literacy arises as a pivotal matter. the next two papers move on from karen hunt’s local initiative to examine broader issues concerning the training of staff within libraries and data libraries to be better equipped to service data requests and support users requests. wendy watkins discusses the creation of a national peer-to-peer training programme for data librarians in canada, that came about as a result of the canadian’s data liberation initiative (dli) in 1996. the initiative opened a new channel of access to statistics canada’s quantitative and spatial data fi les, for which a national support infrastructure was not already established. her paper elucidates the strategies through which a core set of competencies was established throughout data library services in canada. she points to the challenges of initiating a large-scale training programme, starting from a small base of seasoned data librarians, but goes on to highlight the success of regional networking and the rapid escalation of an enthusiastic data community. anne gray’s paper follows on with another canadian example in which she examines the relevance of data and statistical literacy for librarians. she argues that librarians benefi t from having a healthy interest in data and statistics so they can provide greater assistance to users. the paper provides us with a couple of interesting case studies showing ways of evaluating statisticallyoriented publications from a methodological and analytical perspective. how do we recognise and judge quality of content? anne gray usefully points us to canadian and us standards for documenting the quality of government statistical publications. iassist quarterly summer/fall 2004 5 the next paper by susan czarnocki and anastassia khouri offers us a detailed account of how the mcgill libraries electronic data resources service (edrs) was set up and how it is currently run. the service is part of a grouping of library services that provide access and support for various kinds of digital information in the form of maps, electronic data and digital government information. staff with specialisations in particular kinds of data resources work to support for all research and instructional activities. the paper describes examples of approaches taken with respect to supporting use of data in undergraduate curricula and for graduate students. daniel edelstein and kristi thompson describe the work of the data and statistical services (dss) unit in the library at princeton, in providing both consulting on statistical methods and software support for data library users. the staff provide a practical, problem-oriented and intuitive approach that helps individual social science students feel more confi dent about taking on statistical analysis for their research projects. the paper argues that wellresourced data professionals can work proactively and collaboratively with faculty to enhance the statistical skills of undergraduate students. the last two papers consider actual examples of teaching sometimes conceptually diffi cult statistical topics in context. context can mean embedding the learning old techniques in real-life applied social problems, building on social theory, or by situating the confrontation and manipulation of data in a friendly user interface, with simple metadata to hand. louise corti provides a behind the scenes look at a project undertaken to increase the use of real data sources in the classroom, and in a more ambitious sense, to help improve the data and statistical literacy of those studying social sciences, from school students age 16-19 to postgraduates. in a collaborative effort of a national data archive with teachers, the project created a set of free teaching and learning data and statistics-oriented resources based on the study of crime in society. learning strategies were adopted that encourage the teaching of research methods and statistics within a substantive context. an online set of resources was created with data exploration available though the online data browsing system, nesstar. the paper addresses both the positive and challenges that arose from running the project. in a similar way, aaron shrimplin’s and jen-chien yu’s piece provides a case study of using another brand of innovative web-based tools for data presentation and exploration, sda. through close collaboration with faculty, the local initiative showed both students and teaching staff at miami university how data and data analysis could be accessible to them. shrimplin and corti’s papers both describe projects and activities that offer ‘data confrontation’ through the use of tailored datasets delivered through user-friendly web interfaces, but while situated in substantive reasoning. students are enticed and lecturers quickly appreciate the benefi ts that complimentary e-learning approaches can offer their own typically linear teaching pathways. these papers in this special issue offer us useful practical ideas about how to help students with handling, manipulation and interpretation of data. the ideas described point to collaboration as the key to achieving informed and intuitive appreciation and use of data. starting locally by supporting faculty is a great way to test out some of practical strategies described here. other benefi ts of promoting and improving ways of utilising data in the classroom are also apparent. improved data literacy hopefully means greater usage of our richlystocked national data stores. and, upping our user fi gures can be seen as a welcome added benefi t. all of the authors featuring are or have been dedicated iassist members. we all recognise that iassist has a signifi cant role to play in helping nurture new generations of data service providers who can help build a persistent culture of statistically-literate students, researchers and professionals. and may this wave of commitment to and enthusiasm for the cause of statistical literacy thrive!’!’ louise corti and wendy watkins, may 2005 iassist quarterly vol 22 no2 spring 1998 13 howe and graham (1993) proposed that “the goal for the use of metadata and the development of user interfaces should be nothing less than permitting everyone from the novice to the expert to function independently at a desktop machine.” they identified three problems that needed to be addressed: • metadata must be transportable from platform to platform. • there will be pressure on interface designers to make interfaces ever smarter, as more and more naïve users access metadata. • metadata will vary in quality, depending largely upon whether the research team intended the study to be available for secondary analysis. the purpose of this paper is to reassess where the social science community is with respect to the above issues. throughout this paper, we will be careful to distinguish between studies intended for use in secondary analysis and other studies. one of our themes is that tremendous progress has been made over the past five years with respect to data sets intended for use by secondary analysts. in contrast, very little progress has been made with respect to the problem of making metadata available for the tens of thousands of other studies published each year in the social sciences. transportability of the three issues identified by howe and graham (1993), the greatest amount of progress has been made in terms of the transportability of metadata (and data). this is not to say that the problems have all been resolved, but it is now possible to imagine a future in which transportability is a non-issue. while researchers have been able to routinely transmit error-free data around the world for the past six or seven years, it has only been in the past two years or so that the problem of data-storage has been solved. the university of cincinnati has recently purchased an hp 330fx optical storage jukebox to store its social science data collection. the jukebox, with 330gb of direct online storage, is connected to a windows nt file server that it is also easy to forget the fact that the recent success of java promises that access to metadata via the web can, in theory, be unfettered by operating system or platform differences. interface development the uc system compares very favorably to almost any other data archive in terms of access to secondary data. however, as these kinds of storage devices become more commonplace, there have been no comparable improvements in the quality of user interfaces to access data sets. with few exceptions, user interfaces have not progressed appreciably in the last five years. the university of michigan has developed impressive web sites for analysis and extraction of data from the general social survey and the american national election study. the bureau of economic analysis has marginally improved user access to the regional economic information system (reis) cd-rom with the release of a windows interface. unfortunately, the bureau of the census go/extract combination and the national center for health statistics sets software have remained essentially unchanged over the last several years. icpsr also seems to be staking out a position that is distinct from that of industry. the icpsr data documentation initiative is moving in a direction away from that of software developers such as microsoft and sas (microsoft and sas are both members of the meta data interchange specification initiative). while there have been modest advances in the ways that statistical software packages such as sas and spss permit the analyst to make use of metadata in working with a set of data, packages have by and large remained stagnant in the amount and types of metadata they support. most packages do a very poor job of supporting any type of metadata beyond what can be considered “data definition metadata” (i.e., variable labels, value labels, missing value definitions, etc.), and even with respect to these kinds of metadata, the packages’ capabilities are nearly identical to what was available a decade ago, although more of this information is available in point-and-click interfaces. perhaps most importantly, none of the major packages have global access to data resources: where’s the metadata? by mark a. carrozza & steven r. howe * 14 iassist quarterly produced any revolutionary new tools for capturing metadata that archivists will need for bibliographic purposes or that secondary analysts will need for planning their work. as just one illustration of the type of metadata sorely needed but impossible to capture in these packages is information about skip patterns. on the one hand, it must be acknowledged that software package designers must feel frustrated at the lack of standards for metadata in the user community. on the other hand, both spss and sas did at one time pace the user community in terms of promoting better and better data definition features. variability in metadata quality as just noted, producing metadata for a set of data in 1998 is not remarkably different than in 1968: someone involved in the process of research data management has to do a lot of typing. as a result, the metadata available for a study varies tremendously in quality, ranging from very good for large, government-sponsored efforts such as the census to very poorly for the student who has never been taught the fundamentals of research data management. there is, thus, a sharp distinction at present between the accessibility of data resources designed for secondary analyses and virtually all other ones. data sets collected and prepared for the user community as secondary resources are increasingly available via the web and are slowly becoming more and more accessible to end user as the social science community learns what constitutes a useful interface. as metadata standards become better established and cataloging tools become more sophisticated, we can expect the pace at which these studies are made available to accelerate. ironically, the pace at which we are losing primary research data is probably increasing. more and more research is published, and we would guess that smaller percentages of it are being archived. the future our common goal should be nothing less than to create metadata and user interfaces that allow the community of data users to access and process secondary data. metadata standards, although varied and at times painfully subject specific, have emerged. our most popular data management and analysis applications, however, continue to lag in meeting the needs of the social science community. our recommendation for solving the user interface problem is unchanged from five years ago. we suggest an interactive program shell that allows both the researcher and the end user to: • enter information that documents the bibliographic record of the study, including study title, principal investigator, year, funding source, related studies. • create topical files that detail study information – including topics such as sampling, copies of instruments, relationships between study data sets, calculation of weights and standard errors, definition of terms, documentation of calculated variables or fields, and originating hardware and software platforms. • develop data definition and data manipulation structure – including definition of elements and element formats, complete labeling information, descriptive statistics, and free-field explanatory notes. a well-developed system would allow the researcher or other person responsible for data documentation to either create a default minimal metadata collection that would provide facilitate subsequent file access, or create very detailed documentation with all study specifics stored as part of the metadata collection. we also need to work harder at promulgating research data management standards and encourage professional associations, journal editors and funding agencies to require the archiving of research data. references howe, s. r. & graham, r. g. (1993) meta-data and user interfaces: promise and problems. international association for social science information and service technology quarterly, 17, 4-7. *presented at iassist/css 1998 conference new haven, ct., may, 1998. mark a. carrozza institute for policy research steven r. howe department of psychology university of cincinnati (513) 556-5077 mark.carrozza@uc.edu steven.howe@uc.edu mailto:mark.carrozza@uc.edu mailto:steven.howe@uc.edu summer 1998 15 - -- -- --- ----- - --- --- ---- ----- ----- - ------------ ---- --- ---------- ---the international association for social science information service and technology (iassist) and the canadian association of public data users (capdu) announce their joint 1999 conference, "building bridges, breaking barriers: the future of data in the global network". the conference will be held may 16-21, 1999 on the university of toronto campus in toronto, ontario and will address issues of computing and information services in social science research, teaching, and data management. this is iassist's 25th annual conference, and the ninth capdu conference. http://datalib.library.ualberta.ca/iassist/ http://nexus.sscl.uwo.ca/assoc/capdu/index.html http://www.yorku.ca/org/iassist/ ukda iassist quarterly summer 2010 15 iassist quarterly abstract the secure data service is a secure environment funded by the esrc to provide researcher access to disclosive micro data either from their offices, safe rooms in their institutions or on site at the uk data archive. operation is legally framed by the 2007 statistics act which makes possible access to the confidential data for statistical purposes. this short paper introduces this new uk data archive service and proposed specifications, as well as challenges facing data service providers. we envision that the proposed sds infrastructure will meet the requirements of the data security model. the paper also aims to be an exemplar for a secure remote access practice. keywords: remote access, data security, citrix technology, secure data service 1. introduction disclosure of personal information can be harmful and may result in denial of services, embarrassment and loss of reputation and trust which in turn reduces the response rate and jeopardises the future research. research results based on disclosive data can also cause indirect harm by affecting perceptions about a group to which a person belongs. balancing data confidentiality and legitimate requirements of data users is a key problem of the secure data service (sds). confidentiality of individual information can be protected by restricting the amount of information provided by adjusting the released microdata /tables/ statistical outputs (restricted data), by imposing conditions on access to the data products (restricting access), or by some combination of these. the uk data archive secure data service is a new service to allow controlled restricted access procedures for making more detailed micro data files available to some users (approved/accredited researchers), subject to conditions of eligibility, purpose of use, security procedures, and other factors associated with access to the sds data. building on the success of other secure data enclaves worldwide2, and employing security technologies used by the military and banking sectors, the sds will allow trained researchers to remotely access data held securely on central sds servers at the uk data archive. the aim of the service is to provide approved academics unprecedented access to valuable data for research from their home institutions, with all of the necessary safeguards to ensure that data are held, accessed and handled securely. the sds follows a model which recommends that safe use of data should include safe project, safe people, safe setting and safe output (ritchie, 2006. see figure 1). to achieve this goal, data security depends on a matrix of technical, legal, contractual, and educational factors. the structure of this paper revolves around these factors, demonstrating how sds has set up the necessary infrastructure to meet these requirements. we discuss the legal and contractual responsibilities of the users and their institutions followed by issues such as user education and training prior to data access. the technical features and the system specifications are also examined. we also discuss the kind of disclosive data we aim to support in the sds. finally, the challenges facing the sds operation will be examined. 2. legal and contractual framework users of the sds will be required to be either “ons approved researchers”3 or “esrc accredited researchers.” the first of these is defined by the statistics and registration services act 2007 as “an individual to whom the board has granted access, for the purposes of statistical research, to personal information held by it.”4 no definition currently exists of an “esrc accredited researcher,” but we assume that it will have a similar status to an ons approved researcher. this is a person who has been granted access for the purposes of statistical research to personal information which has been licensed to the esds/uk data archive/5university of essex for dissemination on behalf of a government department or some other data provider. neither of these two types of user will be able to use the sds without appropriate training. mandatory training will allow the uk data archive to ensure that end-users are fully aware of any penalties which they might incur if they cause a breach. we believe that if there is user approval to any penalties for breaches, and that they believe that these penalties are reasonable and necessary, we will avoid the inadvertent disclosure to which social science researchers are most likely to be prone. the 2007 act also allows for increased sharing of data between ons and other departments, subject to agreement by parliament on a case-by-case basis. at the same secure data service: an improved access to disclosive data by reza afkhami, melanie wright, mus ahmet1 16 iassist quarterly summer 2010 iassist quarterly time the act also outlines measures designed to protect the confidentiality of personal information. the act states that a person who discloses personal information “is guilty of an offence and liable — (a) on conviction on indictment, to imprisonment for a term not exceeding two years, or to a fine, or both; (b) on summary conviction, to imprisonment for a term not exceeding twelve months, or to a fine not exceeding the statutory maximum, or both.” the sds will immediately suspend access to the service if it believes that any user is perpetrating or attempting to perpetrate any of the breaches listed in sds security breaches or sds confidentiality agreement. a full investigation will follow. users will be required to complete an on-line form which collects personal and institutional details, information about their proposed data usage, and information which demonstrates their expertise and ability to conduct the research described in a competent and secure manner. they will also be required, if they have not already done so, to agree and sign the standard uk data archive end user license (eul), and also agree/sign any special license conditions which apply to the resource they wish to access. this application would be first checked for accuracy, sense and completeness by uk data archive staff, and then forwarded to the data owners for their access authorisation. once authorised, users would be informed and requested to sign up for appropriate training (if they have not already been trained). upon completion of training, users would be granted permission to access the secure data server, either from their own desktop if the owners of the data they wish to access permit, or from their institution’s secure data access room. if their institution does not have a secure data access room, the user will have the option of negotiating access from another nearby institution’s safe room (the sds will offer ‘matchmaking’ introductions, but the specific arrangements must be under the control of the institution hosting the room, as audit trail security will be their responsibility) or coming to the university of essex, to access the service onsite. 3. education &training researchers are the known weakest links in data security. education coupled with the stricter legislative protections mentioned above, can offer another potentially efficient means of improving confidentiality, as disclosure probability can be decreased without imposing costs on rule-abiding researchers. before becoming an active user of the sds, users will have to attend a mandatory training session which will focus first on the user’s legal and ethical responsibilities within their sds user license agreement, the mechanics of how to use the sds, what they can and cannot do in a remote access setting, and the potential of the collaboratory spaces. the second part will focus on principles of the statistical disclosure control, assessment of outputs, and analysis aspects of the particular datasets in the sds. access to the sds will only be granted after users have attended an sds training session. sds staff will vet data analysis outputs for disclosure issues, to ensure that nothing escapes the secure data setting, which could compromise the data security (safe output). one of the purposes of the training is to give researchers the ability to recognise confidential data and distinguish it from statistical results that are safe to remove from the sds. in effect, the training removes the ‘reasonable belief’ defence for a disclosure. we believe that penalties will only be an effective deterrent if they are known, and it should also be clear that we are more concerned with prevention than punishment. 4. system specifications the technology used by the sds must be secure and the system adheres to the highest standards of quality. the technical model that has emerged is one which shares many similarities with both the ons vml and the norc secure data enclave . it is based around a citrix infrastructure which turns the end user’s computer into a ‘remote terminal’ giving access to data, statistical software, and collaboratory spaces on a central secure server held within the uk data archive. the system is flexible, in that depending upon the wishes of the data custodians, access can be restricted to particular users (safe people) and/or particular locations (safe rooms/machines). it is secure because all data manipulation occurs on the server, which is maintained to very strict security protocols. beyond the general security policy, the secure server itself will be subject to additional security measures and controls. approved researchers will access the proposed sds by using vpn (virtual private network/thin-client) technology, which encrypts the data transmitted between the researcher’s computer and the host network. other components of the vpn technology allow control to be established over which network resources the external researcher can access on the host network. the service will employ a citrix xenapp server farm, which participates on two networks mulcahy. et al, 2008). how the system operates? with this technology, although all applications (spss, stata, etc) and data run on a central server at the uk data archive/sds, the approved researcher still interacts with a full windows graphical user interface. this means that the researcher never has to install any complex applications on his/her remote computer; the only application required by the approved researcher is a web browser. this also means that the uk data archive can prevent the researcher from transferring any data from the data archive to a local computer. for example, citrix can be configured so that data files cannot be downloaded from the remote server to the user’s local pc. similarly, the approved researcher cannot use the “cut and paste” feature in windows to move data from the citrix session into an excel spreadsheet sitting on the local computer. finally, the user is prevented from printing data from a local computer. the approved researcher logs onto the sds system remotely via a web secure (https) browser. all data processing is carried out on a central secure server, which processes all requests centrally and returns information about the results. no data travels over the network, except the statistical results sent from the central server to the remote location by an encrypted email after the final outputs are checked against statistical disclosure controls. figure 1: elements of data security (ritchie, 2006) iassist quarterly summer 2010 17 iassist quarterly figure 2: sds system architecture key features • clients cannot remove data • absolutely no webpage access • clients cannot import data • data transfers are logged • all traffic is encrypted • smart auditing • critical security updates are applied daily 5. benefits this system may pose an inconvenience to the user compared to their accustomed ability to use eul (end user licensesimilar to public use file) data on their desktop with all their favourite local software and networked resources. however, it is a price users will gladly pay for local access to data which they might otherwise have had to travel to ons sites to access, or simply have been unable to access. the sds will benefit users in: • a self-contained secure ‘home away from home’ service with familiar analytical environments; • ability to work in their own private work areas or in shared areas with other approved researchers; • access to enhanced, highly sensitive available data storage in tandem with the related metadata through increased capacity and environmental protection; • possibility of data linkage exercise with using existing data in the uk data archive or other administrative data subject to approval of data owners/custodians; and • collaborative functionality including survey and document library, spss/ stata code library, knowledge repository, disclosure review and technical assistance. 6. data a variety of data may potentially be available to users within the sds. we are in ongoing discussions with the owners of key sensitive data resources about how sds might assist in broadening the use and utility of these important resources, whilst assuring that legal, moral and security requirements are met. the specifications of data candidates include: • more detailed variables from existing esrc-funded data resources; • more previously unavailable detailed variables from government social surveys; • other government data previously only available in onsite enclaves, or previously unavailable to academic researchers altogether; • business data which has commercial sensitivity; • administrative data; the sds may be able to provide a secure environment for data linkage activities to researchers whose home institutions lack the technological wherewithal to offer it (restrictions must be negotiated and approved by the data owners/custodians upon user’s application); and • data previously considered too sensitive or potentially disclosive due to its very nature, such as longitudinal data, medical data, etc. in addition, the service will allow users to bring in less disclosive data from the uk data archive standard eul holdings, upon request or researcher’s own data subject to the standard ingest checks and approval. 7. challenges the two main goals of the secure data service are: maximizing utility of microdata for research purposes, and protecting the confidentiality of individual respondents. access to confidential data is an exception to the non-disclosure rule that must be justified according to the balance of the public good of the research against the risk of a breach of respondent privacy. thus, the sds hopes to maximize data utility while minimizing the disclosure risk, utilizing a strategy that is simple, wellcommunicated and acceptable to users. however this task may be daunting as there is no consensus on the definition of what is safe data and second, even more contentious is what information loss means and how it can be measured. as any effort to implement confidentiality protection is associated with some loss of information. maintenance of confidentiality needs a consistent and coherent approach and we must trust the researchers as no environment is free of risk. for example, how can we prevent against a manual data copying or using photographs for researchers who remotely have access to the disclosive data or even user’s memory. for the majority of researchers, data breach happens for access convenience and not out of a malicious intention and surely remote desktop access to the data would diminish that temptation. however, the possibility of disclosure is always there, the legal framework and training and education may deter users -who are approved researchers afterall to perpetrate any confidentiality breaches. 8. evaluation and monitoring of the outputs careful user vetting and the most secure analysis environment in the world cannot on its own ensure that data are not disclosed. the missing piece of the data security puzzle is not what goes into the secure data system, but what comes out of it. for the service to be able to meet the security guarantees placed upon it by the data guardians, it must offer some form of output screening. if an output has been determined to be disclosive, it will be up to the user to determine the best way to render it safe. sds adheres to european-wide essnet standards on good practice in statistical disclosure control of tabular and other statistical analytical outputs (hundepool, et al 2009). the sds disclosure advisor will divide outputs from sds into two main categories: • safe: very low risk of disclosure – output will be released promptly • unsafe: high risk of disclosure – output will be blocked in its current form and won’t be released; the researcher must produce safe outputs and demonstrate that they are free from the disclosure risks there are several solutions available to protect the information of the sensitive cells: 18 iassist quarterly summer 2010 iassist quarterly • combining categories of the spanning variables (table redesign). larger cells tend to protect the information about the individual contributors better. • suppression of additional (secondary) cells to prevent the recalculation of the sensitive (primary) cells 9. summary the sds is a secure environment funded by esrc to provide researcher access to disclosive micro data either from their offices, safe rooms in their institutions or on site at the ukda. it has two goals: to promote researcher access to sensitive micro data and to protect confidentiality. sds operation is legally framed by the 2007 statistics act, which makes access to confidential data for statistical purposes possible. researcher access to microdata serves the public good both by leveraging existing public investments in data collection, and by ensuring high quality science through the replication of scientific analysis. the sds provides approved/accredited researchers with remote access to microdata using the most secure methods to protect confidentiality. this is achieved by implementing technological security (citrix gateway), applying statistical protections, enforcing legal requirements, and training researchers. the sds also ensures that valuable data are preserved for the long term by documenting the data using ddi compliant metadata standards. in addition, the sds aims to engage the research community in using its shared data space to share information which enables collaboration among geographically dispersed researchers. references hundepool, anco et al.: handbook on statistical disclosure control version 1.1, essnet-project. http://neon.vb.cbs.nl/casc/sdc_handbook. pdf. (2009) mulcahy, timothy m, & john nieszel.: towards a secure data service at the uk data archive. sds consultants’ report. (2008) ritchie, f .: disclosure control of analytical outputs. mimeo: office for national statistics, uk. (2006) ritchie, f.: disclosure control for regression outputs, mimeo : office for national statistics, uk. (2007) ritchie, f.: statistical disclosure control in a research environment. mimeo: office for national statistics, uk. (2007) wright, melanie: case for support – secure data service. (2008) notes 1. this paper was presented at the iassist 2010 in the session “secure remote access to restricted data”. corresponding author: rafkhami@ essex.ac.uk, dr. reza afkhami, senior data and support services officer, secure data service, uk data archive, university of essex, wivenhoe park, colchester, essex co4 3sq, tel: +44 (0) 1206 874968, fax: +44 (0) 1206 872003 2. secure remote access is also developing in denmark, netherlands and sweden. 3. http://www.data-archive.ac.uk/orderingdata/agreements/ arformsandnotes.doc 4. statistics and registration services act 2007 § 39 (5). 5. http://www.esds.ac.uk/aandp/access/licence.asp 6. statistics and registration services act 2007 § 39 (9). vol21.2 12 iassist quarterly abstract in the first part of this paper we discuss and define the concept of a networked digital library. we define it both as a new tool for virtual communities engaged in the production and dissemination of information and knowledge, and also as a potential active member of those communities. in the second part of the paper we present the arquitec project, a trial conceived to assess the defined concept of a networked digital library. arquitec is a work in progress that will result in a prototype of a networked digital library for the portuguese academic and research community. introduction the actual and future impact of the internet in our society is one of the most complex and participated discussions of the moment. an emerging issue of this discussion has been the redefinition of the role of libraries, raising the question of what is a “digital library” in a global networked world. in the first part of the paper we discuss and define the internet as a communication medium and as a meetingplace, a “new land” of opportunities for the virtual communities. this vision is discussed in opposition to a common vision of the internet as just a new distribution medium, in the line of the press, the radio or the television. based on that discussion, we define the concept of a networked digital library. a networked digital library is seen not only as a repository for data and information, with the traditional missions of preservation and dissemination of knowledge, but also as an active partner with the potential to stimulate, support and register the process of creation of that knowledge. in the second part of the paper we describe arquitec, a prototype of a networked digital library for the portuguese academic and research community. arquitec is a joint effort undertaken by inesc (an r&d institute), the portuguese national library and jnict (the portuguese r&d funding agency). the purpose of arquitec is to set up a prototype of a networked digital library for the portuguese research and academic community, which will be used to test the concept and the technology. the vision the net isn’t 30 million people, it’s tens of thousands of overlapping groups ranging from a few people to perhaps a couple of hundred thousand at the largest” (o’rally, 1996). it has been broadly pointed out that the information technology in general, and the internet in particular, has been supporting the existence of virtual communities, defined as communities of individuals sharing common interests, but that are not geographically confined. evident demonstrations of that reality are the existing thousands of news groups and electronic mail lists, dedicated to almost all the cultural, professional and political perspectives. with that reality, the internet can be defined as a new virtual space, like a new dimension of the physical and temporal world. it offers a real meeting-place and a multidimensional communication medium, with a social function in the genealogical line of the traditional squares, market places, coffeehouses (see the success of the cybercafes) and the telephone. this is a deeper and vaster view than merely defining it as a simple one-way broadcasting medium, such as the press, the radio or the tv, since in the internet each one can be an equal player, with the same chances to be active as anyone else. this vision has been already a field of concrete experiences in scientific and academic communities. it was maybe first identified by paul ginsparg and steven harnad, that coined expressions like “skywriting”, “esoteric publishing” and “pre-print continuum” (okerson, 1995; harnad, 1990; harnad, 1991; harnad, 1995). harnad presents an interesting perspective on the evolution of the human communication, with the phases of speech, writing, printing and, now with the internet, skywriting. skywriting is defined as both a new medium and a new model of communication, interactive, independent of the space and more suitable with the human cognitive process. this is a scenario favorable to the raising of esoteric virtual communities that, by using the internet for their natural skywriting and pre-print activities, will be able to work and prosper in the production of their knowledge and memory. a digital library for an academic and research community by josé luis borbinha & josé delgado * summer 1997 13 with this reflection we can now complete the view of the internet as the mean (the “ether”) that can allow the library, now converted in the networked digital library, to go and meet the community. networked digital libraries can be important not only for the geographically defined communities (that have already their traditional communal structures), but even more important for the geographically unbounded communities, where they can play as active members in the process of development and creation of knowledge and memory. table 1 resumes that vision for the networked digital library paradigm. in what we call the traditional library, the subject is the book. its value is “sacred” (otherwise it wouldn’t have been purchased) and it is stored “for ever”. in this scenario authors decide what to write and when to edit the book, while the librarian decides whether to buy it or not. finally, the librarians expect the patrons to came to the library and request the book. it was more or less like that until the middle of this century, when the industrial development changed it. the industrial development reduced printing costs, illiteracy and the physical distances, while at the same time it increased the amount of information produced. it is not possible anymore for an individual to absorb all the knowledge produced by mankind, so it is necessary to specialize. the specialization brought thematic magazines, journals, reports, conference, etc. a new subject emergent from this reality is the “paper”, which represents a new type of knowledge. it is not “sacred” anymore, but still formal, being validated by the credibility of an editor or a review committee. this knowledge is not intended to be valid “for ever”, but to be discussed during a period of time, refined and, in the end, what survives is then sanctified in books (while the journals and conference proceedings are stored in the basement). it is difficult for the traditional library to follow the specialization; so the library itself becomes specialized, with the mission to serve specific communities. usually, those communities control now the library content in their own interest, in the sense of who decides which periodicals to subscribe or what to buy. quoting nicholas negroponte: the real value of a network is more related with community than with information. the information super-highway is more than a shortcut to all the books in the library of the congress. it is creating a completely new global social tissue” (negroponte, 1996). in this scenario the library is requested to perform now a more active role. since the communities are well identified, it is now possible to anticipate their needs and to provide customized services, such as the notification of new issues, advertisement of new publications, etc. the scenario changes again with the arriving of the computer. with the desktop publishing tools and www, everyone becomes a potential publisher. the process acquires speed, and the subject is the idea. with computer networks, electronic mail and news groups, communities intensify their interactions. to produce fast results, ideas are submitted in pre-prints or presented to discussion as position papers in informal workshops. ideas that succeed in this process are then published in journals and promoted in formal conferences. what will be the impact of this new reality in the library world? using electronic mail and www, it is easier for the library to reach the communities and provide new services (such as the announcement of workshops, the arriving of new publications, etc.). by the same reason, it is now easy for the users to interact with the library, not only to access online public access catalog (opac) services but, in an extreme scenario, to contribute also with new kinds of meta-knowledge that can enrich notably the library. examples of such contributions can be the tuning and completing of thesaurus and catalogue (allowing dynamic and collaborative cataloguing), the attachment of annotations and comments to the stored documents (allowing collaborative refereeing, for example), etc. after this discussion, we will finish with our vision and a definition for the concept of a networked digital library: a networked digital library is defined not only as an organized repository of data and information, with the traditional mission of preserving that knowledge, paradigms networked digital library specialized library traditional library subject the book the paper the idea knowledge sacred formal informal memory persistent semi-persistent volatile actors author, librarian community, editor community dissemination very slow fast / slow very fast library role passive active interactive table 1: the library paradigms 14 iassist quarterly but as a system with also the mission to stimulate, support and record the process of its creation. it is now our mission to demonstrate how to turn this vision in reality. arquitec arquitec is a trial to test our vision of a networked digital library that will result in a prototype of a networked digital library for the portuguese academic and research community. the system will be accessible over the internet, through a www interface, and will provide access to different kinds of technical documents (such as papers, reports, theses, dissertations, etc.), in any field of the knowledge. the architecture of the system is distributed, with each participating institution (universities and r&d organizations) managing its own repository (see figure 2). based on that infrastructure, the national library will manage an official repository of digital documents. we intend to use arquitec both as a technology demonstrator and a pilot system to develop, test and consolidate expertise in three identified issues: architectures of distributed digital libraries. procedures for management and access to the information, comprising gathering, classification, searching, retrieval and management library procedures. innovative services for networked digital libraries, to exploit the potential of interaction between the library and the community brought by open networks, such as the internet. concerning the management of the information, the main problems will be the procedures for the remote submission of documents and their classification and search, as well as the creation and management of the official archive. the central archive is a repository at the national library, onto which new documents are automatically copied when they are submitted to the local repositories. dealing with documents from different fields of knowledge rises an important issue related with their classification and search. the key problem here is the possible integration of different metadata structures (required by the different contexts and communities) and the use of thesaurus. we will also explore new services to be provided by the networked digital library, such as a filtering service based on the matching of the user profile and documents classification, an annotation service for documents, a collaborative catalogue and thesaurus, etc. the library collection arquitec will provide support for a three-steps workflow in the production of information, comprising: • informal documents: a class of documents usually called grey literature (such as position papers, drafts, preprints, etc.) often useful only in the short/medium term, since it is expected that they will loose interest or they will give rise to refereed documents. • refereed documents: such as full electronic journals, papers presented in conferences or published in conventional journals, etc. • formal documents: theses, dissertations, official reports, electronic books, etc. the increasing scholarly and scientific activity has resulted in the growth of publications rich in new interdisciplinary perspectives. that kind of contents has been raising serious classification problems for traditional libraries, where collections have been classified with catalogues usually defined by static structures. in order to deal with this dynamic classification problem, our digital library should provide users with an interactive catalog of the documents. as illustrated in figure 1, the catalog will be supported by: • a document index. • a multi-context and multi-lingual thesaurus (also interactive). • the user interactions. catalog index users thesaurus repository figure 1: interactive catalog and thesaurus summer 1997 15 the users are able to contribute to the catalog: • directly, by suggesting new keywords for documents or questioning existing ones. • indirectly, by suggesting new relationships to the thesaurus or questioning existing ones. for the development and interaction between the catalog and the thesaurus, experimental work was done with mcf (gutha, 1996), a recent language for meta-content representation. for the thesaurus structure, the iso-5964 standard was followed (iso, 1995). users and services users can access the digital library in one of two modes: anonymous or identified. in order to register a user, the minimum required information is an electronic mail address. however, the users or the system administration can optionally provide other explicit complementary data, useful for some services (such as academic degrees, expertise fields, etc.). an identified user has a profile, composed of the explicitly provided data and by data implicitly extracted from the history of user interactions with the system. for example, if a user retrieves a document related to a specific subject that is not in their explicit profile, this subject is implicitly added to that user’s profile. pending on explicit confirmation, this new subject will be tagged as a potential interest, which the user can easily change later. user profiles serve three main purposes: • searching: for identified users, the profile is used to rank searching results, highlighting documents that match the profile (but not restricting the access to other documents). • filtering: the profile is also used for an information filtering service, supported by electronic mail and by the www interface, through which users can be notified, for example, of new documents of potential interest. • annotations and catalog tuning: interactive services for document annotation and catalog tuning are also provided. during an interaction with the system, any identified user may contribute also with opinions about document classification, by suggesting new keywords, questioning existing ones or by suggesting changes in thesaurus relationships. these contributions are weighed by explicit parameters of the user’s profile (such as the academic degree, for example), and the results of these actions are disseminated by the electronic mailing lists related to the affected documents and subjects. this service gives users a means to interact with the library, not only to access it as an opac service but, in an extreme scenario, to contribute also with a new kind of meta-knowledge” that can enrich notably the library. it is expected that the major part of the documents in the arquitec digital library will be written in portuguese or english, among other languages. due to that, the ability to deal with more than one language will be vital for indexing and searching in documents (for example, to recognize common roots in compound words). the success of this task is one of the main targets of our project, having in mind not only arquitec but also its potential application to other similar situations. a similar problem arises with the diversity of document formats, since we don’t impose a unique format. we try to support as many formats as possible, which is nice for the authors but problematic for us. the integration of such different document formats and languages was done by the development of filters for the indexing and searching modules, rendering the format of documents transparent for the indexing and search tools. to test solutions for those problems we have been experimenting with publicly available indexing (and searching) tools, such as glimpse1 and smart2. these tools have been integrated with palavroso (barreiro, 1993) and correcto (medeiros, 1995), two successful tools developed by the natural language processing group at inesc for morphologic and orthographic treatment of the portuguese language. archiving and persistence a central archive at the portuguese national library will be maintained, with a copy of the formal or refereed documents, after copyright has been secured from their producers. this archive will automatically harvest the new documents from the local servers, storing and cataloguing them in a central repository. a final requirement is name persistence, especially for the documents archived at the national library. depending on whether they are a serial publication or isolated books, printed documents are usually identified by issn or isbn numbers. however, for digital publications such mechanism doesn’t exist yet. it is usual to register cd-rom publications with issn or isbn numbers, specially if they are related to printed publications (such as the cd-roms distributed with magazines), but for on-line publications this is not of great help. the publication of an on-line document is an almost instantaneous process (it requires basically the time to store and to index it in a ftp or http server), and there is no expedient way to require an isbn or issn number for that document compatible with this workflow. another 16 iassist quarterly important problem raised by on-line publications is that its name, or reference, should not only be an unique reference to identify that object in a specific name space, but should also provide a means to access the document (it must “say” where the object is and how to get it). this is a complex problem, globally known as uri uniform resource identifier, and its solution has been addressed by the w3c world wide web consortium3. at present, the most commonly used form of uri is the url uniform resource locator, but urls have a problem: they are not persistent. if we have a document stored at a server where we need to change the structure of the stored information, the original url of that document can become invalid, and any reference to it will originate an irritating “error: the requested document is not valid on this server”. in order to prevent that, we must ensure persistent names for stored objects, through some form of urn uniform resource name. the problem of naming objects in a digital library was generically addressed in the cstr project (anderson, et. al, 1996). that work was reported in the “kahn/wilensky report”, from which emerged the concept of handle as an urn (kahn & wilensky, 1995). that concept was implemented by oclc in the purl persistent url service. in a few words, the purl service is based on the existence of a highly reliable server, where it is possible to register pairs of purls and related urls. in its structure, a purl is a normal url, with a structure like http://dns of the purl server…/object name.... it has a logical meaning that, when used, implies an access to the purl server that acts as a proxy and automatically translates the logical name to the “physical” url of the object referred to (a task performed by a simple http redirect). a purl service, for all the persistent documents with copies archived at the national library, will be provided in arquitec. for each persistent document a purl is automatically and registered at the central purl server. the global architecture before starting the description of the architecture of our system, we will describe some of the most paradigmatic and related projects already done in the field and whose lessons and results we used for our trial. related work the core project started in 1991, and its purpose was to build a database of scanned journals published by the american chemical society (entlich et. al, 1995). by the end of 1994 they had a database of more than 400,000 pages of full text and graphics (in magnetic tapes and cd-rom). the text was converted to ascii and marked-up with sgml (standard general markup language), the database being accessible with dedicated xwindows interfaces. the other major contributors of this project were the cornell university, oclc, bellcore and chemical abstract service. the users accepted the results of the core project very well, but another conclusion was also that “the task of building and maintaining electronic journal databases remains formidable.” a contemporary and also ambitious initiative was the tulip project, started in march 1991 and concluded in the end of 1995 (elsevier, 1996). it was sponsored by elsevier science, and involved nine universities in the usa (c.m.u., cornell, georgia institute of technology, mit, univ. of california, univ. of michigan, univ. of tennessee, univ. of washington, and virginia polytechnic and state univ.). the main goal of the project was to research and test systems for networked delivery and use of scanned journals. elsevier contributed with the scanned page images, ocr generated text and bibliographic data from 43 engineering and materials science journals. the universities provided solutions to deliver these journals in electronic form to their users. the research focus was on technical issues, user behavior and organizational and economic problems. when the project tulip started, the internet was already a reality, but the web was still in an embryonic state. due to that, the delivery technology was based on dedicated graphical clients for x-windows, ms-windows and apple macintosh, besides alphanumeric clients for mainframe terminals. but soon the maintenance costs were evident, and the project shifted to www technology when its advantages and maturity became recognized. in its final conclusions, the project pointed out that the transition from conventional to digital libraries (defined here as libraries with full digital contents), will take much longer and cost more than commonly thought, mainly due to network bandwidth and storage limitations. however, and as it was also pointed out by the core project, we think that this conclusion can not be dissociated from the approach taken: to scan the original material. for example, it was estimated in tulip that a typical journal issue, with 20 articles and 200 pages, requires approximately 17 mbytes of storage, with 16 mbytes for the scanned pages (in tiff format). by comparison, the ascii information resulting from the ocr process requires only 800 kbytes and the indexing and bibliographic information (in sgml format) requires about 200 kbytes. http://dns summer 1997 17 more pragmatic approaches were taken in a series of projects in the computer science reports area. some of the most representative were ucstri unified computer science technical report index (vanheyningen, 1994), ntrs nasa technical report server (nelson, et. al, 1994), waters wide area technical report service (french, et. al, 1995) and cstr computer science technical reports. a common goal of those projects has been easy installation and maintenance of the server sites and support for heterogeneous collections. the idea has been not only to provide scanned versions of printed documents, but also to take advantage of the fact that today it is normal to produce, in the source, those documents already in digital formats (such as ascii, ms-word, pdf, html, etc.). in april 1995, waters and cstr projects joined efforts and conceived a new service: ncstrl networked computer science technical reports library (davis, 1995). ncstrl is a network of servers providing three kinds of services: repository, indexing and user interface. currently ncstrl is a worldwide service, with repositories installed in over 60 universities and research centers across the world. ndltd, a more recent project in the usa, aims to extend that base to provide a generic national digital library of theses and dissertations (fox, et. al, 1996). dienst and ncstrl inesc has been experimenting with the ncstrl technology since middle 1996. we were impressed by its capabilities as a potential framework for future work, especially its open architecture model and its ability to handle documents in several formats. therefore we decided to use it as the core technology for arquitec. in figure 2 we present the main blocks of that architecture. the dienst technology was the main contribution of project cstr for the ncstrl initiative (davis & lagoze, 1994). the ncstrl architecture is based on a network of dienst servers (referred to as s), each one managing a repository of documents (r) the respective index (i) and user interface (ui). the user interface is implemented in html, provided through an http server (the dienst server is written in perl and its interface to the http server uses cgi). a user can access any server from any user interface, since user searches are always performed in all the indexes. optionally, the repositories can be accessed via lite servers (l), the main contribution of project waters for ncstrl. in this case each site only has to provide a metadata description file (m) and have its documents accessible by ftp or http. the lite server converts that metadata to the dienst format, indexes it, and provides normal dienst interfaces for the users and for the other dienst servers. in the specific case of the ncstrl service, it has only one central server for all the registered lite repositories. b i m r m r s ui ir s ui r i s ui r i l ui i figure 2: the ncstrl architecture. a backup server (b) can maintain a copy of all the indexes, which is useful if one of the servers becomes inaccessible. in that case users will not be able to perform retrievals, but at least they will be able to search and find references to the desired documents. finally, our architecture the architecture of arquitec is distributed, with local nodes managing the local repositories at the universities and research institutes, but all the collections are freely accessible for search from any node. the core of arquitec is based on a modified and extended version of dienst 4.0. the required modifications occurred at the three modules of ncstrl, corresponding to three different tasks of arquitec: replacement of the indexing and searching tool, modification of the repository management and modification of the interface. the original dienst indexing and search tool had to be replaced by a more powerful catalog, as described. the new requirements implied modifications at the ncstrl repository interface level, in order to perform full text indexing of as many document formats as possible (such as ascii, postscript, ms-word, etc.), as well as in different languages. concerning the management of information, the main generic problems were the procedures for submission of the documents, their classification and search, as well as the creation and management of the central archive. the submission of documents can be done remotely, with the user authenticated by username and password (stronger security and authentication issues, for which we recognize the importance, were not addressed for now). the submission process starts by the filling and submission of 18 iassist quarterly registration forms, by www. users will be required to provide the location of the original document at an ftp or http server. after that a confirmation procedure takes place: an electronic mail message is sent to the user and the system waits for a reply. after successful confirmation, the document is then retrieved, registered and added to the catalog. the core of the ncstrl system was also modified in order to allow the automatic management of the official central archive. in practice this means that the central host, at the national library, automatically gathers all new persistent indexed documents into a central repository. that repository is used as an official archive, which is especially important for theses and dissertations. it also serves as a mirror repository to provide global fault tolerance. the ncstrl user interface was modified in order to support all the described requirements, new functions and services. the modifications were done essentially in the submission of documents (that can now be done remotely), as also in the support of the search task. all the interface components were redesigned to support multi-lingual access (portuguese and english in the first release). finally, a directory for the registered users was added to the system. it is a distributed directory based in the x.500 model, with an ldap interface (yeong et. al, 1995). future work and open issues medium term work will be concerned with the integration of other spaces, accessible by new interfaces at lite dienst servers. examples will be interfaces for z39.50 servers4, useful for the integration of opac systems such as the catalogs of conventional libraries, and harvest5 brokers, useful for the support of informal publications and other similar material such as mailing lists, source code, etc. examples of other identified research issues requiring our attention in the medium/long term are: • document structuring: research will be done on using sgml and other alternative solutions for structuring the information objects (a specially interesting issue to be applied not only for the original documents but also to represent the associated annotations); • natural language: trials will be done in the classification and search of documents with natural language techniques, with a special concern for the portuguese language; • authentication and certification authorities: the requirements for authentication and certification authorities, for both the documents and users, will be addressed in medium term; • legal issues: among generic problems, such as how to assign and observe other properties of the documents (such as terms and conditions and other copyright problems), examples of new open interesting problems in this field are the legal implications of the new objects, composed by an original document and a list of annotations (or just the legal implications of an annotation); • long term preservation: how will the official repository survive the evolution of the hardware and software, such as storage technology, operating systems, document formats, viewers, etc.? references anderson, g.; lasher, r.; reich, v. (1996). the computer science technical report (cs-tr) project: a pioneering digital library project viewed from a library perspective. the public-access computer systems review 7, no 2, 1996. available at http://info.lib.uh.edu/pr/v7/n2/ande7n2.html barreiro, a.; pereira, m., j.; santos, d. (1993). linguistic options and criteria in the development of palavroso, a computational system for the morphological description of portuguese (in portuguese). inesc report no. rt/54-93, december 1993. davis, j. r. (1995). creating a networked computer science technical report library. d-lib magazine, september 1995. available at http://www.dlib.org/dlib/ september95/09davis.html davis, j. r.; lagoze, c. (1994). a protocol and server for a distributed digital technical report library. technical report tr94-1418, computer science department, cornell university, 1994. elsevier science (1996). tulip final report. elsevier science edition. available at http://www.elsevier.nl/locate/ tulip. entlich, r.; garson, l.; lesk, m.; normore, l.; olsen, j.; weibel, s. (1995). making a digital library: the chemistry online retrieval experiment. communications of the acm, april 1995, vol. 38, no. 4, 54. fox, e. a.; eaton, j. l.; mcmillan, g.; kipp, n. a.; weiss, l.; arce, e.; guyer, s. (1996). national digital library of theses and dissertations. d-lib magazine, september 1996. available at http://www.dlib.org/dlib/september96/ theses/09fox.html french, j. c.; fox, e. a.; maly, k. (1995). wide area technical report service: technical reports online. communications of the acm, april 1995, vol. 38, no. 4, 45. http://info.lib.uh.edu/pr/v7/n2/ande7n2.html http://www.dlib.org/dlib/september95/09davis.html http://www.dlib.org/dlib/september95/09davis.html http://www.elsevier.nl/locate/tulip. http://www.elsevier.nl/locate/tulip. http://www.dlib.org/dlib/september96/theses/09fox.html http://www.dlib.org/dlib/september96/theses/09fox.html summer 1997 19 gutha, r. v. (1996). meta-content format. apple computer. available at http://mcf.research.apple.com/hs/ mcf.html harnad, s. (1990). scholarly skywriting and the prepublication continuum of scientific inquiry. psychological science 1, 342 343. harnad, s. (1991). post-gutenberg galaxy: the fourth revolution in the means of production of knowledge. public-access computer systems review 2, no 1, 39-53. available at http://info.lib.uh.edu/pr/v2/n1/harnad.2n1. harnad, s. (1995). the postgutemberg galaxy: how to get there form here. the information society 11(4), 285291. iso international organization for standardization (1995). iso-5964: documentation guidelines for the establishment and development of multilingual thesaurus. geneva, 1985. kahn, r.; wilensky, r. (1995). a framework for a distributed digital object services. available at http:// www.cnri.reston.va.us/home/cstr/arch/k-w.html medeiros, j., c.,(1995). processamento morfológico e correccao ortográfica do português. master thesis, instituto superior técnico universidade técnica de lisboa, lisboa. negroponte, n. (1996). ser digital. editorial caminho (portuguese edition of the original title “being digital”, 1995). nelson, m. l.; gottlich, g. l.; bianco, d. j.; paulson, s. p.; binkley, r. l.; kellog, y. d.; beaumont, c. j.; schmunk, r. b.; kurtz, m. j.; accomazzi, a.; syed, o. (1994). the nasa technical report server. internet research: electronic network applications and policy, vol. 5, no 2, 25-36. o’reilly, t. (1996). publishing models for internet commerce. communications of the acm, june 1996, vol. 39, no 6, 79-86. okerson, a. s.; o’donnell, j. (1995). scholarly journals at the crossroads: a subversive proposal for electronic publishing. association of research libraries, june 1995. vanheyningen, m. (1994). the unified computer science technical report index: lessons in indexing diverse resources. second international world wide web conference, www’94 oct. 94, 535-543. yeong, w.; howes, t.; kille, s. (1995). rfc 1777: lightweight directory access protocol. ietf network working group. available at http://www.umich.edu/~rsug/ ldap/doc/rfc/rfc1777.txt. 1 http://glimpse.cs.arizona.edu 2 ftp://ftp.cs.cornell.edu/pub/smart 3 http://www.w3.org/www/addressing/addressing.html 4 http://lcweb.loc.gov/z3950/agency 5 http://harvest.transarc.com * paper presented at iassist/ifdo ‘97, odense, denmark, may 6-9,1997. josé luis borbinha (jose.borbinha@inesc.pt) ist technical superior institute (lisbon technical university) department of electrical and computers engineering josé delgado (jose.delgado@inesc.pt) inesc institute for systems and computer engineering telematics systems and services group http://mcf.research.apple.com/hs/mcf.html http://mcf.research.apple.com/hs/mcf.html http://info.lib.uh.edu/pr/v2/n1/harnad.2n1. http://www.cnri.reston.va.us/home/cstr/arch/k-w.html http://www.cnri.reston.va.us/home/cstr/arch/k-w.html http://www.umich.edu/~rsug/ldap/doc/rfc/rfc1777.txt. http://www.umich.edu/~rsug/ldap/doc/rfc/rfc1777.txt. http://glimpse.cs.arizona.edu ftp://ftp.cs.cornell.edu/pub/smart http://www.w3.org/www/addressing/addressing.html http://lcweb.loc.gov/z3950/agency http://harvest.transarc.com mailto:jose.borbinha@inesc.pt mailto:jose.delgado@inesc.pt vol252 iassist quarterly summer 2001 11 the social research informatics center, tárki founded the social science databank in 1985, which was a consortium of different academic and research institutes. the aim of the founders was to create a service-providing center which besides developing the hungarian empirical sociological research, engaging in research consultancy, and conducting surveys functions as a databank, that serves the establishment of the information basis of the hungarian social researches and to improve its methodological co-ordination. since 1992 tárki has its own survey department. in 1998 the tárki incorporation was founded to conduct also profit oriented opinion polls, while the tárki consortium continues the non-profit data archiving. at present, there are 9 member institutions of the tárki social research informatics center, which are the following: • eötvös loránd university of sciences, • university of szeged; • university of debrecen; •hungarian central statistical office, • budapest university of economic studies, • institute of political sciences of the hungarian academy of sciences; • the sociological research institute of the hungarian academy of sciences; • national institute of vocational training; • high school of nyíregyháza. the databank has been functioning already since 15 years and its main task is to archive and disseminate data and survey documentation, and also the acquisition of data form other research institutions. at present the databank contains more than 450, mostly hungarian related, empirical social data sets in spss format, suitable for secondary analysis. these data are mostly originated from nationwide representative sample surveys. the archived surveys are conducted by tárki and by other hungarian research institutes as well. in the databank the are also some data bases suitable for international comparisons. the users of the databank can choose from various kinds of topics, such as attitudes, family, social deviance, health care, life-styles, values, consumer patterns, occupations, mobility, ethnic and migrant groups, local governments, stratification, poverty, social policy, social relations, social strata, rural society, religion, elections etc. data access categories a. free access and dissemination b. free access for hungarian researchers and research institutes, for others with the owner’s permission only. c. free access for the member institutions of tárki , for others with the owner’s permission only. d. access with the owner’s permission only (for max. 5 years) the databank also sells the data. the price of the dataset depends on whether it is simple or aggregate, on the date the survey was conducted and on the status of the purchaser. see table 1 for the current pricelist of the databank. the databank also operates different thematic databank sections. the aim of creating these thematic sections is to collect and organize those data sets which belong to the same topic. one of the thematic sections is the historical archive, which was built jointly with the hajnal istv·n kˆr (hik), presently it includes 23 economic and socialhistorical databases. (http://www.tarki.hu/t_adat/ index.html) tárki data bank – the case of hungary by ildiko nagy 1 12 iassist quarterly summer 2001 the women’s data archive was established in 2000. it sums up social scientific research projects concerning women and gender issues and thus makes data and publications, a register of researchers and urls easily accessible. (http://www.tarki.hu/adatbank-h/nok/ index.html) the tárki databank operates the new democracies web site, which provides on-line access to the questions and answers from the multinational new democracies barometer database. (http:// rs2.tarki.hu:90/ndb-html/) the databank also publishes and distributes the cd-rom version of the hungarian household panel surveys conducted between 1992 and 1997. since the date of its establishment, tárki has been laying emphasis on forming close relationship with the major significant social science data archives of the world. at present it is the member of three international data organizations: ifdo (international federation of data organizations), cessda (council of european social science data archives), icpsr (inter-university consortium for political and social research), ecpr (european consortium for political research). this membership, due to free exchange of data sets between the member institutions, gives the tárki databank an opportunity to make international databases available for hungarian users. the databank also joined the luxembourg income study project and the european household panel network of ceps/ instead in luxembourg, and therefore it gives free access to lis databases for hungarian researchers. these project leading institutions also offer scholarships for researchers interested in international income comparison. figure 1. functional structure and staff archiving, disseminating data; data acquisition operating tárki's web page international contacts tárki databank 3 persons tárki's documentations, journals, reports andorka library library 1 person hardware and software maintenance tarki online it group 3 persons social research informatics center director elpmis tesatad etagergga tesatad erofebshtnom21-0detcudnocyevrus 052 0001 erofebshtnom84-31detcudnocyevrus 661 666 erofebshtnom94detcudnocyevrus 48 333 iyrogetactnuocsid setutitsnihcraeserngierof snoitaroproccilbupnairagnuh ecirptsilehtfo%66 iiyrogetactnuocsid snoitutitsnihcraeserrebmem-nonnairagnuh stnedutsngierof ecirptsilehtfo%33 iiiyrogetactnuocsid ikrátfosnoitutitsnirebmem stnedutsnairagnuh snoitutitsnihcraeserlanoitanretni eerf table 1 price list of the tárki databank june 2001 (in usd) iassist quarterly summer 2001 13 since 1988 the tárki has been taking part in the international social survey program, therefore all related international datasets are accessible through our databank. further international data sets are the east-european comparative surveys, which are conducted by a research foundation consisting of tárki , a czech and a polish research institution. the central european opinion research group (ceorg) ceorg was founded in 1999 and situated in brussels. a significant part of the revenues of the tárki databank comes from the support of the tárki inc. and scientific projects. since 2001 the hungarian scientific research fund has been one of the main supporter of the databank. the membership fees of the tárki member institutions and the data access fees altogether make up only 10% of the annual revenues. the planned total revenues for 2001 is about 58000eur. the hungarian social researchers and students of high education are the main user of the tarki databank. future plans and problems the tárki databank participates in international social science database projects such as the consortium for household panel studies for european socio-economic research (cher) financed by the european commission. this project lasts for 3 years between 2000 and 2002 and it aims at harmonizing several national panel studies with the european community household panel. the databank is going to develop the present thematic sections: the historical data archive and the women’s data archive in 2001. although the tárki databank is a public archive it is not supported by the government. our main problem is connected to the funding of the databank. unfortunately tarki ‘s homepage is just partly available in english, so many information cannot reachable for researcher from abroad. 1. social research informatics center, tárki databank, h-1112 budapest, budaörsi út 45. phone: 36-1-309-7693, fax: 36-1-309-7666, nagyildi@tarki.hu, http://www.tarki.hu figure 4: users of tárki databank in 2000 figure 3 revenues of the databank – plan for 2001 50% 30% 10% 5 % 5 % tárki inc. scientific projects hungarian scientific research fund own revenues membership fees 44% 32% 19% 3% 2% hungarian students researchers from member institutions other hungarian researchers researchers from other countries non-scientific users vol29-1.indd iassist quarterly spring 2005 by by margaret law* reduce, reuse, recycle: issues in the secondary use of research data introduction “reduce, reuse, recycle”, a phrase familiar to us from the environmental movement, can also be used to reflect on the secondary use of research data. secondary research refers to the use of research data to study a problem that was not the focus of the original data collection. this may be data collected for administrative, health or educational purposes, census data, or data collected as part of a previous study. this secondary analysis may involve the combination of one data set with another, address new questions or use new analytical methods for evaluation (szabo & strang, 1997). both benefits and dangers have been attributed to the secondary use of research data. distinctions are often made between large scale data collections, particularly sample survey data collected at public expense, and smaller bodies of data collected at personal expense. there is general agreement that the first should be shared and made generally available in a “timely” fashion, but little agreement about the second. there is also no sense of agreement about what would constitute timely in this situation (clubb, austin, geda, & traugott, 1985). while the questions surrounding the secondary use of research data have always existed, they have become more pressing with the use of new technologies. new capabilities include easier data sharing, faster and more complex analysis, and the development of large scale data banks. previously, the ability of researchers to communicate was limited by time and distance; now data can be shared globally at the click of a mouse. as we adapt to the electronic environment there are new concerns about confidentiality and the threat of security lapses. the potential for finer data resolution becomes possible with better data collection tools and technological innovations. while the fundamental ethical issues have not changed, the possibilities created by new technologies have brought them to the forefront. just as governments have taken a strong leadership role in developing and supporting good environmental habits, they must be encouraged to develop and support good habits concerning the storage and use of research data. this paper summarizes ethical concerns about the secondary use of data and the arguments for encouraging or facilitating it. it includes some potential solutions and discusses the implications of the increased use of new technology. while it focuses on the canadian regulatory environment, similar issues arise in other countries. concerns about data sharing and data confidentiality affect researchers and data librarians across the world. while each country may have a different regulatory environment and a different research culture, the need to find an appropriate balance between the optimal use of data and the protection of individuals is worldwide. with increasing globalization, and an increase in international research, the development and articulation of appropriate guidelines becomes paramount. data sharing is a fundamental value for iassist, and individual or random decisions about data sharing stand in the way of providing the best support for researchers. by looking at the canadian situation, data librarians may develop and share common messages as part of an overall advocacy plan to support data sharing. there must be limits, of course, to protect respondents, but these must be delineated and managed in a coherent way that not only recognizes their rights, but also those of researchers, and of the taxpayers who frequently fund the research. this advocacy effort must be aimed at regulatory bodies, funding agencies and the researchers themselves in order to change the cultural values around secondary use of research data. in canada, much research involving humans is governed by the three major granting councils, who have developed a shared policy statement to govern all research involving human participants done in canada, and by canadians outside of canada. section c3 of the canadian tri-council policy statement (tri-council policy statement: ethical conduct for research involving humans, 1998) lays out guidelines for research ethics board (reb) approval of research that proposes the secondary use of data. it is clear that reb approval is required if identifying information will be involved, but leaves it to the researcher and the reb to determine exactly what constitutes identifying data. 6 iassist quarterly spring 2005 there is sufficient legislation in most provinces to provide a framework for researchers wishing to use data that has been collected for purposes other than research. for example, the alberta health information act (health information act chapter h-5, 2000) provides direction for researchers’ use of health care data. division 3, ‘disclosure for research purposes’, defines the role of the ethics committee, including consideration of whether the researcher should be required to obtain consent from subjects, implying that there is some choice in this matter. the ethics committee must also decide whether the research is of sufficient importance to outweigh privacy concerns, whether the researcher is qualified and whether there are sufficient safeguards. the picture is not so clear, however, when the data was originally gathered for research purposes. this data is not governed by legislation, and guidelines are open to interpretation. there is considerable discussion about who owns the data and decisions about whether to share data are often made by the original researcher and may depend on a number of personal factors. a requirement for all researchers to consider the potential for secondary use of their data, either by themselves of by others, would contribute to a more orderly use of data with resulting benefits for researchers, subjects and the community. concerns about the secondary use of data concerns about secondary use of data generally focus on the potential for harm to the individual subjects of the research and the lack of informed consent. many writers are passionate about the primacy of informed consent for any type of research involving human subjects. consent applies not only to a particular researcher, but also for an identified purpose. to quote kalman (1994), ‘the requirements to seek an individual’s consent to participate and to provide data for a specific purpose must take precedence.’ since researchers generally are not able to predict potential requests for secondary use of data that they are collecting, they are unable to fully inform subjects of the primary research about potential future uses of data. as this full disclosure of information is one of the requirements of informed consent, it follows that it is not possible to get informed consent for unanticipated uses of data. others argue that if the second researcher were to contact subjects to ask consent to re-use data, the original researcher must first identify the individuals thereby breaching their privacy. this situation could be managed by having the original researcher contact the subjects on behalf of the second researcher. privacy is generally defined as a personal issue, defined by the subject. the subject may have felt comfortable disclosing information to the first researcher because of their relationship or rapport, but secondary research could leave him open to actions of researchers with whom he feels less comfortable (homan, 1992). technology-driven data analysis techniques also create the potential for triangulation of data: the combining of variables that allows identification of specific individuals and organizations even though identifying information was removed from the original data sets. for example, there has been concern that the combination of census data and geographic information can allow the identification of small or unique groups. the use of gis allows for closer identification of geographic data through the availability of differing degrees of granularity (trainor & dougherty, 2000). there is additional concern for vulnerable populations that could be at particular risk if their confidentiality were breached. current north american legislation and the media have raised awareness about profiling issues, and certain populations such as those involved in criminal activities or who are hiv-positive have a high risk of harm if they are identified. in addition, particular forms of data such as oral histories, photographs or diaries cannot be made anonymous because the identification of the respondent is a large part of the value of the data. researchers in these situations often feel that they have given unqualified pledges of confidentiality to participants, leading them to bar access to the material unless participants can be contacted for permission (hedrick, 1985). considerable commitment from the first researcher would be needed to contact the subjects for get permission for them to be approached by a second researcher. ethical practice requires a balancing of benefits and harms when conducting research. while this may be assumed to refer to the benefits and harms that may be experienced by the subjects of the research, it could also be interpreted as requiring a consideration of the potential harm to the original researcher. the ‘design and execution of data collection effort is a creative activity that sometimes involves innovative techniques.’ it seems reasonable to question why a secondary analyst should benefit from someone else’s work, particularly if the second research is a potential scholarly competitor (clubb et al., 1985). high quality data are expensive to collect, organize and store in an accessible form. if the data is to be used by someone else at a later date, additional work and documentation may be required. if this is carried out at the original researcher’s expense, it would seem to create the potential for harm with no counterbalancing benefit. this is particularly true if the original researcher is not cited, as it could have a negative effect on tenure and future funding opportunities (sieber, 1991). some writers have also proposed a negative effect on “good science”, brought about as a result of too much data-sharing iassist quarterly spring 2005 7 (stanley & stanley, 1988). “it is a lot easier, faster, and less costly to obtain someone else’s data than it is to design a study, recruit participants, collect and analyze data” (p. 178). this potentially leads to fewer original data sets, reducing the potential for multiple independent evaluations. methodological problems also arise. for example, rules requiring the original researcher to delete identifying information or other methods of anonymization may prevent the accurate use of data files, or interfere with their appropriate linking with other data, resulting in incorrect associations (fienberg, martin, & straf, 1985). if the original data were not well collected or documented the second researcher may lack information about possible errors, the relationship of the data to the universe of responses, details about the sample, ways in which data were analyzed (sieber, 1991) and the assumptions underlying interpretations (fienberg et al., 1985). any of these could lead to flawed research. in support of secondary use of data the arguments in favor of secondary use of data may not be as straightforward, but should also be considered within the framework of the ethical principles in the tri-council policy statement. the basis of these arguments focuses on the ethical obligations to good science, and to the benefit of the community at large. a fundamental principle of research ethics is respect for human dignity, incorporating both the selection and achievement of morally acceptable ends and the morally acceptable means to those ends. if this is understood to mean that individuals are important and should be treated appropriately, one implication is that researchers should make as much use as possible of the data that is collected in order to reduce the burden on research subjects. this would provide an improved benefit/harm ratio for vulnerable groups who may be at risk from repeated data gathering intrusions into their lives. one might argue that there is, in fact, the potential for greater benefit if research with already collected data provides more opportunities to support these groups. taking steps to ensure that interpretations of data are valid through encouraging multiple methodologies demonstrates respect for research subjects through accurate interpretations of their behaviour. secondary analysis creates an opportunity to establish relationships that were entirely unpredictable at the time of the original data collection (dale, arbor, & proctor, 1988). for example, some of our understanding about the causes of disease has occurred through the secondary analysis of medical records that were not collected with the intention of making such a causal relationship (dale et al., 1988). the ability to link data files and to create families of data creates possibilities that together they can contribute knowledge that none could contribute alone. in essence, this is a situation where the whole is greater than the sum of the parts (johnson & sabourin, 2001). if a second research repeats original calculations to assure accuracy, it is not considered to be a secondary analysis. if, however, it analyzes the data from a different perspective or within a different theoretical framework it allows the findings to be challenged and debated, and creates an opportunity both for further discovery and for a deeper understanding of the interpretations of the data (dale et al., 1988). developing and implementing protocols for data sharing creates potential for testing the generality of research findings, and comparing analyses on different data sets across time or across locations allows us to ‘generalize findings about social phenomena’ (fienberg et al., 1985). good science requires that data be available for scrutiny and reanalysis as part of scientific enquiry (fienberg et al., 1985). the practice of a second researcher reanalyzing data is widespread, although this would seem to pose the same concerns about breach of confidentiality as any other access to data by a second researcher. “it seems reasonable to argue that if one is prepared to publish assertions about the nature of reality based on collected data, then one should equally be prepared to allows others to examine that data to check the validity of the assertions”(reidpath & allotey, 2001). concerns about privacy and confidentiality are the most frequently raised objections to secondary use of research data. some writers believe that it is “likely that this obstacle is cited much more frequently than is warranted” (hedrick, 1985 p.142). attempts to quantify the risk of identification, particularly from anonymized records (marsh et al., 1991) support to some extent the notion that the risk is over-stated. other writers assert that achieving informed consent for secondary research is never truly voluntary as there are pressures on subjects to agree simply because they have already agreed once before, but this appears to not have been adequately investigated. the issue is further confused by a discussion of what is meant when the original researcher states on the consent form that personal data will not be shared. some researchers would interpret this in the most conservative way to mean that none of the data will be shared, ever. a more sensible interpretation might be that data can be shared as long as it is properly anonymized and all identifying characteristics are removed (johnson & sabourin, 2001). while the real question is how the subject interprets it, not the researcher, the subject is likely influenced by the researcher’s position. this is an ethical position that must be resolved before it can be managed through improved methodology. the canadian tri-council policy statement allows for a breach of confidentiality in section 3.3 (c) if the individuals to whom the data refer have not objected to secondary 8 iassist quarterly spring 2005 use. the research ethics board is charged with the responsibility of evaluating the sensitivity of information, seeking consent to used the stored data, and allowing the researcher to propose an appropriate strategy. the reb is directed to pay particular attention to the possibility of “harm or stigma” that might be attached to identification. while this obviously does not preclude the secondary analysis of research data, it clearly does not take a strong position in favor of it. the sharing of research data must also be considered under the ethical principle of balancing harms and benefits. in many situations, the original data collection was paid for through research grants, funded by the taxpayer. this ‘harm’ to the community of taxpayers should be balanced by an appropriate benefit; the most logical way of maximizing that benefit is to ensure that the optimal use is made of all data collected. the allocation of harm and benefit in this case needs to be extended to include all of the participants in the research process, not just the immediate subjects of each piece of research. to quote davey smith (1994), “data paid for by public money are public property.” the additional analysis of data also provides the benefit of increased confidence in the outcomes of research to the larger community. “publishing the findings of research in peer reviewed journals implies a high level of confidence by the authors in the veracity of their interpretation. therefore it stands to reason that researchers should be prepared to share their raw data with other researchers, so that others may enjoy the same level of confidence in the findings” (reidpath & allotey, 2001). the ethical principle of reducing harm can also be viewed as support for the secondary use of data. subsequent use of data already collected reduces the impact on the larger population by involving a smaller number of research subjects and subjecting them to a smaller number of tests. on a more practical level, the use of previous studies can help formulate a good research question and refine the analysis carried out in subsequent studies (davey smith, 1994). good methodology is one of the primary mechanisms for reducing harm. the tri-council policy statement articulates the maximization of benefit as a guiding principle. this strongly supports the secondary use of data as a costeffective and convenient mechanism for the advancement of knowledge. as research money becomes more restricted, increased secondary analysis will allow for ongoing research in situations where new data collection is hampered by lack of resources (szabo & strang, 1997). in situations where research will have a significant impact, for example in influencing public policy, it is essential that data be considered from many directions to reduce the possibility of flawed or weak conclusions. hedrick (1985) states that secondary analysis allows for the “reinforcement of open scientific inquiry” (p.127) by providing for evaluation of research and the opportunity to replicate or reanalyze it using the same or different methods. a critical process will increase public confidence in the value of research and reduce the incidence of faked and inaccurate results. increased public confidence may also benefit the research community by providing support for research funding. the sharing of research data is a logical process that maximizes the benefits of research while reducing much of the potential for harm. many of the anticipated risks and harms can be managed through improved methodologies. once this position is understood and widely shared, those solutions will become part of the research ethos. solutions for anticipated risks a number of writers have proposed solutions for the anticipated risks stated by individuals who are not in favor of secondary use of research data. while the list below is not complete, it demonstrates the breadth and ingenuity of researchers who are committed to good science and maximum benefit to the community. a number of the solutions focus on the requirements for confidentiality from the secondary researcher. for example, clubb, austin et al. (1985) recommend “a form of licensing or swearing in as a condition for access to data with the possibility of legal sanctions and penalties for breaches of confidentiality” (p. 62). the british sociological association, cited in heaton (1998) recommends that researchers consider obtaining consent that at least “covers the possibility of secondary analysis.” a number of approaches to the original consent form have been proposed. in some cases the original consent form includes provision for secondary research with the requirement that the secondary study receives approval from an ethics review committee. at the very least, this raises the question of potential secondary use in the minds of both the researcher and the subject, and allows respondents the opportunity to object should they wish. while this may not strictly meet the requirement for informed consent, it demonstrates an effort to resolve the situation early in the research process. it assumes that the secondary analyst is “bound by the same confidentiality and privacy restrictions as the primary analysts”(szabo & strang, 1997 p.7). better anonymization can be built in by the original researcher as a required part of research ethics approval for gathering data concerning humans. proposed methods include a uniform practice of removing names and substituting numeric codes, removing occasional data values that reflect rare attributes and could allow for identification of specific individuals and organizations, iassist quarterly spring 2005 9 aggregating data in such a way that the performance of identifiable individuals or organizations is not obtainable, and various forms of encryption (clubb et al., 1985; johnson & sabourin, 2001). this requires a better understanding of which identifying items data need to be maintained to keep the data useful while protecting the respondents. while it can be argued that some forms of data such as photographs or diaries have little value without identifying information, these should be regarded as exceptions rather than the norm and general policy should not be based on them. a mathematical solution has been proposed that adds enough uncertainty to statistical analysis to prevent the identification of individuals while not significantly affecting the outcome of the analysis. the process, known as “jittering”, is defined by johnson and sabourin (2001) as “adding a small, normally distributed random value with a mean of zero to all fields that might be used to identify an individual by matching against publicly accessible records.” the real problem may not be a lack of potential solutions, but a reluctance to implement them. this could be encouraged through a number of means outlined by sieber (1991) in support of secondary research: · in appealing to enlightened self interest grant bodies could require willingness to share data as funding criterion. this is already required by some funding bodies (davey smith, 1994). if not a requirement, funding priority could be given to those who create and share important data files and to research which builds on upon existing data files. note that this requirement only works if it is monitored and there is a mechanism for sharing. · to minimize the potential harm to researchers through not having their work adequately cited, the research community could require the implementation of clear and enforced standards for citation of data files. standards for authorship should include identification of the source of data. · in order to reduce researcher fears about secondary use of data, research education should be enhanced to include improved understanding about the advantages, process, and barriers in data sharing. funding agency policy statements could be a stronger advocate for secondary use of research data by including further instructions for the original researchers. it should work from the assumption that data sharing is standard practice unless there are specific reasons for prohibiting secondary analysis. it would then include, for example, the requirement for the secondary analyst to properly recognize the original researcher and a clear stipulation of the conditions under which data sharing is prohibited. this would facilitate good science by removing the potential conflict of interest that occurs when a researcher must decide whether or not to share data. research ethics approval could require a process for coding, storing and providing access to the data in a uniform fashion. this could be accomplished by adding an information professional, such as a librarian or an archivist, to the research team (humphrey et al, 2000). a uniform requirement would mean that the burden of this additional work was evenly spread among researchers and would be considered as part of the original research design. this has primarily been a discussion of the ethical issues surrounding data sharing. there is also the practical consideration of whether researchers are prepared to voluntarily share their data, or whether they have maintained it in a form that allows it to be used by other people. (corti, foster, & thompson, 1995; reidpath & allotey, 2001) at this time, to share or not is still largely an ad hoc decision made by individual researchers. a clearly articulated policy that evaluated the situation on scientific merit and an analysis of harms and benefits would ensure that the ethical principles were the basis for decisionmaking. conclusion the secondary analysis of existing research data provides many exciting opportunities for the development of new knowledge. it can be aligned with the ethical principles of research in many countries by minimizing the respondent burden and maximizing the potential benefits from the data. to make a change in the research culture requires strong advocacy on the part of data librarians, to change the thinking of funding bodies, regulatory agencies and researcher. traditionally we have not required that the potential for data sharing be a part of every research proposal. now technology has provided us the opportunity to ‘build the corpus of knowledge, not through the frenzied winnowing that has characterized our evaluations in the past but through an orderly interlocking of the puzzle pieces contributed by the disparate sub-fields. we have the means, for the first time in our history, to begin putting together the full picture of human behaviour’ (johnson & sabourin, 2001). it is important that we champion the changes needed to accept this challenge, and to advocate for the creation of an ethical basis that requires the development and implementation of strategies to overcome potential barriers to data sharing. 10 iassist quarterly spring 2005 “reduce, reuse, recycle”. this phrase has shaped a generation of behaviors about environmental concerns. governments and funding agencies have promoted changed behaviour through investment in infrastructure, and in policy directions. the same thinking can be used to shape our understanding about ways of reducing the costs and burdens of data collection, increasing the value of research, and maximizing the benefits for everyone involved in the research process. * margaret law, university of alberta, margaret. law@ualberta.ca. with this article margaret law won the iassist strategic plan publication award in 2005. the author wishes to acknowledge the advice received from dr. g.griener, department of philosophy, university of alberta. references clubb, j. m., austin, e. w., geda, c. l., & traugott, m. w. (1985). sharing research data in the social sciences. in s. e. m. m. e. s. m. l. fienberg (editors), sharing research data (pp. 39-88). washington, d.c.: national academy press. corti, l., foster, j., & thompson, p. (1995). archiving qualitative research data. social research update, 10. dale, a., arbor, s., & proctor, m. (1988). doing secondary analysis (contemporary social research series no. 17). london: unwin hyman ltd. davey smith, g. (1994). increasing the accessibility of data. bmj, 308(june 11), 1519-1520. fienberg, s. e., martin, m. e., & straf, m. l. (editors). (1985). sharing research data. washington: national academy press. freedom of information and protection of privacy act, rsa 2000, c. f-25, sec. 42 health information act, rsa 2000, c. h-5, ss. 48 – 56, (2002). ottawa: government of canada. heaton, j. (1998). secondary analysis of qualitative data. social research update, (22). hedrick, t. e. (1985). justifications for and obstacles to data sharing. in sharing research data (pp. 123-147). washington, d.c.: national academy press. homan, r. (1992). the ethics of open methods. the british journal of sociology, 43(3), 321-332. humphrey, c. k., estabrooks, c. a., norris, j. r., smith, j. e., & hesketh, k. l. (2000). archivist on board: contributions to the research team. forum qualitative sozialforschung / forum: qualitative social research, 1(3). johnson, d. h., & sabourin, m. e. (2001). universally accessible databases in the advancement of knowledge from psychological research. international journal of psychology, 36(3), 212-220. kalman, c. j. (1994). increasing the accessibility of data. 309(17 september), 740. marsh, c., skinner, c., arber, s., penhale, b., openshaw, s., hobcraft, j., lievesley, d., & walford, n. (1991). the case for samples of anonymized records from the 1991 census. journal of the royal statistical society. series a (statistics in society), 154(2), 305-340. reidpath, d. d., & allotey, p. a. (2001). data sharing in medical research: an empirical investigation. bioethics, 15(2). sieber, j. (1991). social scientists’ concerns about sharing data. in sharing social science data; advantages and challenges (pp. 141-150). newbury park, california: sage publications, inc. stanley, b., & stanley, m. (1988). data sharing: the primary researcher’s perspective. law and human behavior, 12(2), 173-180. szabo, v., & strang, v. r. (1997). secondary analysis of qualitative data. advances in nursing science, 20(2), 6674. trainor, t., & dougherty, k. (2000). selected issues concerning disclosure avoidance in the context of userdefined geography. statistical journal of the un economic commission for europe, 17(2), 133-139. tri-council policy statement: ethical conduct for research involving humans. medical research council of canada//natural sciences and engineering research council of canada//social sciences and humanities research council of canada. august 1998. vol25.4 4 iassist quarterly winter 2001 iassist quarterly winter 2001 5 editor’s notes the theme of the iassist 2002 conference in storrs, ct was “accelerating access collaboration and dessimination”. the conference consisted of many good (“best ever”) and well attended workshops, streams, discussions, panels, and presentations. many of the powerpoint presentation are already available for viewing at the iassist website www.iassistdata.org. on the web-site you can take a look under multimedia presentations from the iassist 2002 conference. if your presentation is not available you should contact the collector lisa neidert (lisan@umich.edu). papers from the conference are scheduled to appear in this and coming issues of the iassist quarterly. you can contact the editor (kbr@sam.sdu.dk). this issue vol. 25-4 of the iassist quarterly contains three papers from the 2002 conference: september 11th has affected nearly everybody, everywhere. shortly after the acts the national opinion research center (norc) lauched the “national tragedy study”. the intention was to replicate the “kennedy assassination study” carried out forty years earlier. for the session on ”the new frontier for archives” tom w. smith and michael forstrom from national opinion research center and university of chicago had prepared a paper with the title “in praise of data archives: finding and recovering the 1963 kennedy assassination study”. the paper explains that this task cannot be said to have been easy. this is a detailed description of a search for documentation and data for a single study. luckily the needle in the haystack was retrieved, but again technology raised new difficulties, in this metaphor the problem was now the missing sewing machine. the conclusion is as the title says “praise for data archives”, if data and documentation had been properly stored at a data archive the similar task could have been routine and not an experiment in data excavation and the task could have been carried out in confidence of a positive outcome. from the session on “extreme intelligence: pushing expert systems to their limits” robert wozniak from the minnesota population center explores in his paper “emerging from the quagmire: building expert systems technologies for the social sciences” the latest technology for the discovery and access to social science data. utilization of the web has moved the extraction expertise from the professional to the user and leaving the user in a quagmire. by taking advantage of the current technology and by utilizing domain knowledge both the novice and the expert will be assisted. the paper uses nhgis national historical geographic information system and its terabyte of united states summary census data as an example. the system is an expert system and is capable of deducing that a requested variable does not exist at the required geographical level and will expand to the next level to fulfil the request. in the session called “make it faster, make it bigger: developments in data delivery systems” micah altman from harvard university presented the paper “open source software for libraries: from greenstone to virtual data center and beyond”. the paper provides an introduction to open source software (oss) and looks into questions like what are the advantages and disadvantages of oss and which are most useful in the library environment. among the advantages are cost and fast respondance to bugs and thus evolution of the software. the risk of using oss could be the stagnation of software and a lesser degree of usability. for the data library or archive oss can be of significance as it can be recompiled and ported to new hardware and operating systems which is important if the digital objects are requiring specific software. (often the archive will try to store the objects in formats that are readable or transformable into other existing and even coming software packages). the paper also directly reviews some oss packages for library use as well as gives addresses to these and other resources. three papers from three sessions at the recent conference. do have a good read of this issue of the iassist quarterly. karsten boye rasmussen, august 2002 http://www.iassistdata.org mailto:lisan@umich.edu mailto:kbr@sam.sdu.dk complex data, simple tools: an introduction to text retrieval packages* by andrew marchant-shapiro ' departments ofpolitical science and sociology union college, schenectady, n.y. introduction before the advent of personal computers, qualitative researchers could often be found "waste" deep in typed interview transcripts. dealing with these transcripts often meant hours of searching for a particular passage or pattern. for many, uttle has changed — except that the transcripts are now word-processed instead of typewritten. but precisely because they are word-processed, these transcripts open up new possibilities for computeraided retrieval and analysis. this article provides an overview of one class of software programs, text retrieval packages (trps), that can provide significant assistance to qualitative sociologists with minimal investments of both time and money. using a hypothetical text retrieval package, 1 suggest some techniques that sociologists can use to maximize the utility of trps. i outline the basic characteristics of trps, and describe a few commonly available software packages that present variations on the trp theme. the techniques introduced here are not specific to any one system, and may be used to advantage with a wide variety of text retrieval packages. 1 should make clear at the outset that this article is about improving access to textual data, with specific application to qualitative data such as transcripts from unstructured ("conversational") interviews. none of the trps described in the following pages provides for analysis of qualitative data. while such programs are readily available, and are used by a growing number of qualitative analysts, they are not my concern in this article. rather, the programs discussed below are simple tools that provide qualitative researchers with greatly improved access to their complex data. thinking of the interview as a database when a sociologist hears the word "database," he or she is likely to think of a collection of coded data arrayed in rows (cases or 'records') and columns (variables). normally, such databases are manipulated with software programs known as data base management systems (dbms). the dbms makes it possible for the researcher to gain rapid access to specific sections of his or her data. for example, a researcher using dbase iv, a popular dbms program, might want to see the ages of all those individuals in her database who uved in ilunois in 1970. depending on how the data are structured, she might give the following command: list age for "il"$state 1970 the dbms program would first locate those rows of data for which the variable state_1970 had the value "il," and then isolate the variable age in each such record and print its value on the screen. a typical output might look like this: 27 28 19 41 30 note that by using the dbms program, the quantitative researcher has gained great power in interrogating her data. she no longer needs to manually search for each case in which a subject was living in illinois in 1970. this ability to isolate particular records for inspection is one of the reasons that dbms programs have gained popularity with quantitative researchers. on large surveys, such programs are often used to ease cleaning of data and allow for the isolation and closer inspection of outlying cases. the usefulness of similar strategies should not be lost on the qualitative researcher. there are many instances in which it is desirable to move rapidly to a section of an structured or unstructured interview that is marked by the occurrence of one or more key words or phrases. these range from the early days of a research project, when one is exploring the transcripts of recent interviews, to the final stages, when a researcher may need to find one or more quotations to reinforce her point. one may even want to test the notion that two words or phrases verbalizing particular concepts occur only (or most frequently) in conjunction with one another. in a set of interview transcripts that can run to hundreds of pages and millions of words, how can you find the particular passage, or passages? one approach, to be recommended for its economy and simplicity, is to use the search function of your word processing program to look for the text in question. but 36 assist quarterly with a few exceptions (noted below), this limits you to searching for a single word or phrase in a single text file at a time. moreover, more complicated searches (such as searching for all paragraphs of text that do not contain a particular word or phrase or combination of words and phrases) are beyond the capabilities of word processing programs. this is where text retrieval programs come into play. the grandparent of modem trp programs is a widely available program called grep.^ it originated on unix mainframes, and is today available on all computers that run the unix operating system. moreover, various public domain' versions of the grep program are available for most microcomputers, and a limited version of grep, called find, is distributed by microsoft with every copy of ms-exds. grep is a simple program, but extremely powerful. in essence, you give grep a word or phrase to search for, and it compares each line in a data file with that specific word or phrase. lines that match can be counted, printed to the screen, or saved into a new file for further manipulation (alternatively, you can do the same thing for lines that don't match). a researcher might be interested in seeing how often the word 'credibility' appears in a particular interview transcript. the transcript is stored in the file trans017.txt, so our researcher invokes grep this way: grep -n credibility trans017.txt grep scans each line oftrans017.txt for the pattern of letters forming the word credibility. the '-n' in the command causes grep to number the lines in the file as it scans them. when it finds a line containing the pattern credibility, it prints that line to the screen, forming an output uke this: 0023 and that was a serious problem for our credibility 0040 was it a credibility problem? no, credibility was 0215 credibility. plain and simple. if it hadn't of the researcher now knows how often the word api)ears in the transcript, as well as where it appears. since grep works very quickly, a few such searches can give the researcher significant insight into his data in a very short time.* note, however, that there are some important drawbacks to the way that grep handles the data. first of all, what we have retrieved are lines of the file, not sentences. from a computer's point of view, lines are a sensible units to use because it is easy to tell where one line ends and the next begins. the computer treats each line as an independent record. but from a human perspective, unes are not very useful as records: they may contain anything from a few words to a few short sentences, and they do little or nothing to establish the context within which any particular datum is found. grep can rapidly locate all lines of text containing matching patterns; we need programs that can retrieve entire chunks of text (sentences or paragraphs, for example), context and all. at a minimum, we need to locate not only the statement containing the word or phrase for which we are searching, but also the stimulus that evoked that statement. a second problem is that grep is limited to searching for a single word or phrase at a time. while a skilled user can compensate somewhat for this limitation through the use of complex 'regular expressions' or root searches (see below), grep is incapable of searching for phrases that break over lines and cannot examine text chunks for the presence and/or absence of multiple words and phrases. grep is limited to searching for a single word or phrase (a sequence of words); we need programs capable of looking for combinations of words and phrases that may or may not be sequential. modem trps answer both of these needs and more. rather than operate on single lines of text, they can operate on paragraph-sized chunks' and can make use of boolean operators (see below) to allow for a variety of ways to combine search texts. moreover, unlike word processing i»-ograms, trps can search many files with a single command, greatly speeding up the retrieval process. in themselves, these improvements over word processors and grep make trps powerful if simple-minded tools. to get optimum performance from a trp, however, requires more than just aiming the program at a file and telling it to go to work. by adding a modicum of structure to the transcript of an unstmctured (or structured) interview, we can realize significant benefits. structuring a transcript actually involves three different aspects: structuring, in the sense of organizing a conversation into meaningful 'chunks'; identifying concepts, adding keywords to the record that either amplify the content of the conversation or actually represent analytic categories; and problem prevention, a process analogous to the 'cleaning' of quantitative data. structuring the transcript if the chunks of data that we wish to retrieve are larger than a single line, we need to stmcture the text so that the trp we are using understands where a particular chunk begins and ends. how we stmcture chunks (or records) depends on the particular data we are looking at and how we plan to use it consider the unstructured interview. such an interview is a conversation, typically made up of paragraphs; the spring 1991 37 first party speaks, then the second, then the first, and so on. typically, the interviewer asks a question, then the informant replies, as in figure 1. * figure 1 q: were you particularly worried about extremists coming into the movement at that time (1978)? a: not especially, no. not until we got word that the news media had mixed up some of our group in the west with the posse committatus. that cut into our credibility something fierce, and made it very difficult for us to get sympathetic press coverage. the question and answer— sometimes with fouowup questions and answers or other interactions— provide the context within which to understand a particular statement. if we can treat each question-and-answer set as a record, that is, as an independent datum, then we are well on the way to transforming what may be a long (and sometimes rambling) interview into a useful database. within each record, we should have not only the full text of the answer, but the stimulus that evoked that answer. bear in mind, however, that the real challenge is not understanding for ourselves what constitutes a record, but organizing our data in some way so that the computer's notion of what constitutes a record is identical to our own. once we have arrived at a definition of a record that is adequate for our own use, we have to think about the structure of the records as the computer sees them. from a computer's perspective, the most useful form of raw data is a file that consists of text (letters, numbers, punctuation and spaces) and a few special characters, such as carriage returns and line feeds. such a file is known as an ascii (american standard code for information interchange) text file, and uses characters in a standardized fashion.^ to divide a file into records that both the researcher and the computerytrp will understand, use single spacing within paragraphs, and double spacing between paragraphs, as in the following example: xxxxxxxxxxxxxxxxxxxxx xxxxxxxxxxxxxxxxxxxxxx <-record 1 xxxxxxxxxxxxxxxxxxxxxxx xxxxxxxxxxxxxxxxxxxx xxxxxxxxxxxxxxxxxx xxxxxxxxxxxxxxxxxxx <-record 2 xxxxxxxxxxxxxxxxxxxxx xxxxxxxxxxxxxxxxxxxxx <-record 3 xxxxxxxxxxxxxxxxxx x in this way, each paragraph of text becomes a distinct record, and the trp can easily distinguish where one ends and the next begins. records should include, at a minimum, a question and the response it invokes. these should be divided in some way, however, so that we can tell at a glance at which part of the interaction we are looking. one way to divide between the two is to insert a line of hyphens (-) between question and answer. the one way not to divide the question and answer is with a double carriage return; this will make the question and answer appear to the trp as two independent records. identifying concepts not infrequently a conversation has more meanings than would be apparent from the text of the interaction itself. in such instances, we may wish to add still another section to the record— again, divided by a string of hyphens or other special characters— that consists only of keywords or comments pertaining to the interaction. such keywords may simply clarify the meaning of the text of the conversation, or they may be analytical categories you have assigned to the particular record. if you do use keywords, it is a good idea to use some special character to mark them so that the trp can differentiate keywords from the rest of the text. for example, all keywords might be preceded and followed by the '*' character— "'aggression*, *money*, and so forth. this helps to avoid confusion between words that are part of the text per se and others that are introduced by the researcher once the interview has been completed. preventing (and resolving) common problems while we are busily creating the perfect data record, however, we need to be aware of complications that we can introduce that may make searching difficult or impossible. the major problem is one that afoicls all text processing: misspelling and inconsistent spelling. this is a particularly nasty problem if someone other than the interviewer transcribes the interview. computers are powerful but infiexible creatures; if you search for the name kamin, you may find that the name doesn't come up in the database— because it has been entered variously as kemin, kammin, camin and/or camyn. there are fiexible trps (discussed below) that may find one or more of these misspellings through the use of 'fuzzy' search criteria, ' but it is probably best not to rely on technology to fix this problem after the fact. the most straightforward answer to this problem is the spelling checker. these programs, often integrated with word processing programs, scan the text, either while it is being entered or once the file is complete, and locate words that do not match a dictionary file. most spelling checkers have provision for an auxiliary dictionary. 38 assist quarterly which contains words not in the main dictionary but which are used by the writer. this is the place to establish a list of names of persons and organizations, so that misspellings will be detected at once and corrected. you should also supply yourself, or the person transcribing the interview, with a list of names and special terms that appear in the interview(s) with which you are working. as you proofread the transcripts, you can add to this list (and the spelling checker list) as you go along. sometimes, however, a new name or term comes up, or a name is garbled. you should have provisions for alternative spellings, such as the following: after that [mankoff/mankov (0213)] told me that an early step in going through the interviews could be to search for the '[' character, so as to locate and resolve problems. alternatively, you can leave the variant spellings in place, in case context later allows you to interpret the garbled passage. if you are not transcribing the interview yourself, have the person who is insert the digit counter value for the point on the tape where the garble occurs. in this way, even imperfect transcriptions (and there are few perfect ones) will be useable by computer search programs. it may sound like a lot of work to structure text in this way, but it is not really that difficult if the structuring is done when the data are first entered into a word processing program. if the researcher himself or herself is entering the interview, then keywords can even be added at the same time. the additional labor imposed by a simple data structure is a small price to pay for the ease of access that will result. in the next section, we turn to the issue of access in order to get a sense of what can be accomplished. whether coded or not, once the data have been structured, the hard part is over. searching for data if we think of our interview data as now consisting of paragraph-sized records, each record consisting of a question and its associated answer and constituting a context for the statements therein, we are in a position to consider how we would like to specify which records to retrieve. searches may be simple (for one word or phrase) or complex (for various combinations of words and/or phrases). if we have added keywords to the interview data, we may search for these as well. simple searches recall the example of the quantitative researcher. her search began by specifying a subset of possible records — those records that contained the value 'il' in the variable state_70. the qualitative researcher does not have variables to work with in the same sense; instead, he has a chain of verbalized (and perhaps coded) concepts. while these are not consistent from record to record— only a few members of the set of all possible concepts are present in any given record— we can test for the presence or absence of particular words.' consider figure 2. this is the same paragraph shown in figure 1 , but now structured as a record and stored, with other records, in the ascii file intrvw.(x)1: figure 2 q: were you particularly worried about extremists coming into the movement at that time (1978)? a: not especially, no. not until we got word that the news media had mixed up some of our group in the west with the posse committatus. that cut into our credibility something fierce, and made it very difficult for us to get sympathetic press coverage. *extremist* *posse* *perception* 'media** 1978* *west* this record contains a stimulus, a response, and a set of keywords that both overlaps (e.g., *media*) and categorizes (e.g., *perception*) the information contained in the interaction. the record shows the presence of such concepts as extremist, movement, media, credibility, 1978, and so on. on the other hand, it does not contain terms indicating such concepts as electoral politics, formal organization, or legitimacy (to name just a few possibilities). so, if we wanted to retrieve only those records that included a verbalized or keyworded conception of credibility, we might give a hypothetical trp a command like this: list 'credibility' in intrvw.ool the result would be a listing of all of the records in intrvw.ool that contain the term credibility, including, of course, the record shown above. if we wanted to search for all records that contained the concept electoral politics, the trp would not retrieve this record. conversely, if we searched for all records that did not refer to electoral politics, this record would be among those retrieved. but suppose that we wanted to search more broadly— for variations on credibility. suppose that our informant didn't actually use the word credibility, but said something like 'it was hard for us to be credible.' since credible is not the same pattern of letters as credibility, the computer would not have found that record. but we can modify the search in one of two ways so that we are more likely to find appropriate records. we can either search using roots, or we can search using multiple terms. spring 1991 39 if we are searching for variants on a single term, searching for a root can do the job. for example, we might search for the pattern common to both words, i.e., credib. to do a root search, consider all of the similar terms you want to retrieve and search for the common portion of those words. legitimacy, legitimate, and legitimation, as well as variations such as iuegitimaie, can all be retrieved through a common root if you use this approach, though, be careful not to shorten the root too much; if you do, the search results may be useless because large numbers of records containing 'noise' words satisfy the search. some trps allow for a variation on the root search method using wildcards. these are special characters that can be inserted into a search that will match any other character or combination of characters. if '+' matches any single character and '_' matches any group of characters, then 'gr+w' matches words such as grew and grow, and 'im_le' would match anything from stimu/ent to impossible. obviously, wild card searches are also subject to 'noise' problems, and should be undertaken with care. complex searches while searching for a single word or for variations on a single word can be helpful in plowing through long transcripts, it is often more useful and more interesting to be able to choose records based on the presence of two terms, or on the presence of one and the absence of another. boolean operators are ways of specifying logical connections between words and/or phrases. online information services such as lockheed's dialog service make use of these operators, as do more common, pc-based systems such as wilsondisc, and virtually all dbms programs. the basic boolean operators, and, or, and not, can be used singly or in combinations to set exacting criteria that records must meet before the trp will retrieve them. widening the search: logical or a root search works by using a single, less rigorous criterion for matches. in contrast, a multiple term search expands the search pattern by allowing a record to be retrieved if it satisfies one or more elements of a set of criteria. to construct such a set, we use the boolean logical operator or. for example, if we give a command to our trp to: list for 'credibility' or 'credible' in intrvw.ool a record that contains either word will be retrieved. obviously this technique can be expanded so that concepts that may be expressed in a variety of ways can be searched. we might want to search for terms like 'credibihty' or 'legitimacy,' for example. using the logical operator or always widens the search, since a record that satisfies any part of the expression that has been ored together is retrieved. sometimes, however, oring things together gets us more than we want. it is then that we can use another logical operator to tighten our search criteria. narrowing the search: logical and and logical not when we and things together, we are telling the computer to retrieve only records that meet multiple criteria. for example, we could exclude the example record shown in figure 2 from a search by asking our trp for the following: list for 'credibility' and 'organization' in intrvw.ool the record satisfies one criterion but not the other, so it is not retrieved. only those records will be found that contain both words. using and takes care, because it is possible to quickly reduce the number of records that match the search to zero. the utility of and and or is increased by adding the third logical operator, not. not allows the trp to retrieve a record only if a particular term is not present in the record. not is seldom useful alone, but in combination with and and or, it allows for very precise specification of searches. if we wish to find only those records that refer to 'this', but not those that also refer to 'that', then we can search for 'this' and not 'that'. grouping logical operators with parentheses while many searches are easy to specify with one or two logical operators, searches can become quite complex, and it is important to specify the priority in which logical operators act fortunately, most trps allow the use of parentheses, which allow the researcher to specify the order in which the trp evaluates logical relationships. we can develop searches such as ('credibility' or 'legitimacy') and not 'organization'. this particular search would first retrieve the subset of all records in which either 'credibility' or 'legitimacy' were present, and then reject the sub-subset of records which also contained the term 'organization'. if we had instead defined the search 'credibility' or ('legitimacy' and not 'organization'), the trp would first find all records containing 'legitimacy' but not 'organization', and then retrieve as well all records containing 'credibility' regardless of whether or not they included 'organization'. an example of the logical operators' power to differentiate among records may be in order here. consider the following one-line records: 1. then bob told carol and ted. 40 assist ouarteriy 2. but of course alice and carol told bob and ted. 3. alice and ted were outraged at that 4. finally, alice left with ferdinand. below are some search criteria and the numbers of the records that each search would retrieve. these should demonstrate clearly the different behaviors of the various operators. 'bob' or 'carol' or 'ted' or 'alice' (u,3,4) 'bob' and 'carol' and 'ted' and 'auce' (2) 'alice' and not ('bob' or 'carol' or 'ted') (4) 'alice' and ('bob' or 'ferdinand') (2,4) searching using keywords with the use of boolean operators, keywords take on a special significance. they are more than merely additional tags that we can use when our informants use varying terms to discuss a single concept. through the use of and, or, and not, we can examine the relationships that exist between keywords that indicate coded concepts and the content of the conversation itself. recall that keywords are marked with special characters (*). these markers affect searching in particular ways. for example, a search on the term 'legitimacy' will be satisfied whether the term occurs in the text or in the keyword section. but '*legitimacy*' will only be satisfied by the term in the keyword section. by combining keywords and logical operators, we can do searches like this: list for 'media*' and 'credibility' in intrvw.ool this search would find only those records noted and marked by the researcher as having some bearing on media issues, and then only the subset of these records that had verbal and/or keyword relations to credibility. by adding keywords to our records, we begin to approach the same kind of specificity and power in searching that dbms programs afford quantitative researchers. making use of the output the goal of all these manipulations is, of course, to find specific records within a large body of information. what you do with that information once you find it is up to you, but you should be aware that not all trp programs allow you to save the data that you find. some merely allow you to view the records that the program has retrieved. most have provisions for saving some or all of the retrieved records to an ascii file. some, such as golden retriever (reviewed below), take you to the point in your transcript where the match occurred and allow you to save as much or as httle of the surrounding material as you desire. since the usefulness oftrp programs lies in their abiuty to winnow data, as it were, you should probably avoid programs that do not allow you to save output to a new file. for example, seelceasy, one of the programs reviewed below, has no provision for placing the retrieved text into a new file. this limits its usefulness in anything other than exploratory research, since the only way to recotd the results of your search is with pencil and paper (or the print screen key). once you have an output file, you can do several things with it. you can simply include the file, or an edited version, in a paper or article you are working on. or, if the number of records retrieved is large, you may be able to treat the new file as a second-order database— searching more specifically within the file. in any event, you should always take a look at the output file before including it in other documents or doing further searches. computers are wonderful servants, but they take everything — including our mistakes— literally. if the results look strange to you, review your search commands carefully. the difference between ('bob' and 'carol') or ('ted' and not 'alice' and 'bob' and ('carol' or 'ted') and not 'alice' may turn out to be significant. a good way to check the search results is to make certain that a randomly-chosen record within the output actually does satisfy your search request from ideal to real: some inexpensive trp programs up to this point, we have been dealing with a hypothetical trp. none of the programs that i discuss below does exactly what our hypothetical model does. rather, each emphasizes one or more features described above. none of these programs costs more than $50, and most cost significantly less; some are available for the asking. for each program, i give a brief summary and then a spring 1991 description and evaluation of how the program works. these descriptions are summarized in table 1 . i also include information on how to obtain each program. originally, i intended to begin this section with a speed comparison across the programs, and to this end i tested each program using the ascii transcript from a three hour unstructured interview— approximately 22,500 words. the slowest program i tested was a version of grep, which look 70 seconds to go through the file; the other programs all had times of 30 seconds or less. consequently, i have not included a speed comparison. the differences here are negligible. rather than focus on speed in deciding which program might meet your needs, i suggest that you consider the features that particular programs emphasize that might make them most useful in your particular work. one important factor to note is that, while the hypothetical trp described above is controlled through command lines, many trps are menu-driven: you select the actions you want from a list, and the computer does the rest these may be simpler to use for those unfamiliar with computers, or in classroom situations. golden retriever golden retriever, version 4.0, shareware — $39.95. golden retriever is a powerful trp; it can be menu or command driven, and it has the capability to search for multiple-word phrases. it even has an adjustable fuzziness level. that is, you can make close guesses at the spelling of terms you don't quite recall, and golden retriever will often find them. the degree of "fuzziness" golden retriever will allow in a search comes preset to a reasonable level, but you can make the search more or less rigorous through a menu choice. golden rennever can make use of logical operators, as described above, but only in a very limited fashion. if you and words together, for example, they will only match exactly tiie same pattern in the file— 'bob' and 'carol' will only match bob carol; it will not match carol bob or bob alice carol. the words in the file must not only appear in the same order as in the search criterion, but they must also be adjacent golden retiiever's menus are clear and easy to understand. when golden retriever finds a word in a file, it takes you to the appropriate record and highlights the word on the screen; you may then use the cursor keys to choose how much of the surrounding material, if any, to save into an output file. one unusual feature allows golden retriever to run in the background, while you work in your word processor or other text entry program. pressing a special key shifts you into and out of the golden retiiever program, allowing you to search for data while you are working on a report, for example. there is a preview version of golden retriever available, the golden retriever pup. golden retriever pup works exactly the same way as does golden retriever except that it will not read data files on a hard disk, which limits its usefulness considerably. the pup version is available via modem from computer bulletin boards, ot for $10 from the national collegiate software clearinghouse (ncsc), duke university press, 6697 college station, durham, nc, 27708. the full version can be ordered for $39.95 from wesware, 42 epping street, lowell, ma, 01852. grep there are dozens of versions of grep available, most of them in the public domain, posted on computerized systems across the country. if you have access to a modem, this is one way to locate a grep program. if you don't, find a colleague or computer center person who can help you. most microcomputer greps will explain themselves to you if you enter grep or grep ?. grep is a good place to start looking at trp systems because it is simple and cheap; you should be able to obtain a copy for free. most (but not all) versions of grep can save output to a file by appending the command '>', followed by a file name, to the end of the search request hence, grep 'bob' intrvw.txt >save.bob saves the results of the search to the ascii file save.bob. resnoter resnoter, version 1.0, (c) ncsc —$35. resnoter is one of the most technically sophisticated programs i evaluated. it is the only one of the trp programs reviewed here that uses indexing. this means that using resnoter is keyword-intensive; if you want to use this trp, you must insert extensive keywording in your data; resnoter will not search raw text. from the keywords that you supply, resnoter constructs a list of code words and their locations in the database. '" if you ask for 'bob' and 'carol', resnoter need only look at the locations in the 'bob' list and compare them to those in the 'carol' list. it can then jump direcuy to the records that satisfy the request. because of this indexing, all logical operators are available and resnoter is very fast. however, you must remember that if you add a new keyword you will not be able to use it until you have re-indexed the database. indexing does not take long, but it is a step you must not forget in working with resnoter. another drawback is that resnoter is loaded with menus. menus should make life easier for the user, but resnoter' s menus are posi42 lassist ouarteriy lively frightening because their operation is highly inconsistent. still, if you need rapid access to large amounts of data, and if you are willing to insert keywords, resnoter is worth learning. you can save output to a file, and you can choose what portion of the record to save (keywords, raw data, or both). resnoter is one of a series of text retrieval and analysis programs published through and available from the ncsc. search search, version 1.3, public domain— free/$25. search is one of the more complete trp programs reviewed. it uses all three logical operators and, unlike golden retriever, is not limited to 'exact' logical matches. that is. search scans the whole record to see if the logical requirements are matched, not just adjacent sets of words. 'bob' and 'carol' will match not only bob carol but also carol bob and bob alice carol. search supports multiple levels of parentheses and can search for any combination of up to fourteen words and/or phrases. one particularly nice feature, useful with logical or searches, allows search to note, either on the screen or in the output file, which of the logical search terms it matcheid in a given record. you can also have search ask you whether or not to save a given retrieved record to its ouqjut file. an option allows search to work like grep, if you want to look only at line-sized records. search has two relatively minor drawbacks. first, it is mainly command-driven. to use it, you must learn to type in a sequence of commands. for example, if you gave this command: search rvtrvw.ool b =bob&caroi > output.txt search would look for paragraph records containing both 'bob' and 'carol' and saves the results to output.txt. these commands are not hard to learn, but may intimidate a first time user. search provides a second, more limited search mode for beginners, which allows for only and and or operators. in this secondary mode, the trp asks the user for search terms and filenames. still, it is not as friendly as a menu-driven system like golden retriever. a second drawback is that matched words are not highlighted in retrieved records when they are printed to the screen — if you are dealing with large records, this can make it difficult to find the exact point at which the match occiured. search is available free on computer bulletin boards or directly from its author for $25, which includes a subscription to future versions. note however that, if you obtain search from a bulletin board, no donation is expected. for further infonnation, contact eric bohlman, 1921 highland avenue, wilmetle, il, 60091. table 1 program version boolean name logic fuzzy search save output max size per record price golden retriever 4.0 yes(l) yes(2) yes no limit $39.95 (3) grep various no no yes lline firee resnoter 1.0 yes no yes no limit $35.00 search 1.3 yes no yes no limit free/$25.00 seekeasy 5.0 no yes (4) no 2 lines —$30.00 (1) golden retriever uses boolean logic to match only adjacent words. (2) golden retriever allows the user to adjust the level of 'fuzziness.' (3) a "sample" version is available through computer bbss. (4) seekeasy's fuzzy search is not adjustabl spring 1991 43 seekeasy seekeasy, version 5.0, shareware— $30.00. seekeasy is a slightly speedier but much less successful implementation of "fuzzy" searching than golden retriever, with considerably less flexibility. you cannot adjust the "fuzziness" level, and the program is limited to two lines of context around each word or phrase it finds. there is no way to save the results of the search to a file. in its favor, seekeasy is extremely easy to use; type what you are looking for and, if it is in the file, seekeasy will find it unfortunately, since it will retrieve the 100 closest matches (in no particular order), it will find a great deal of material you don't want, and you may have to search through all of that to locate what you asked the program to find for you in the first place. this program would be more useful for keeping an address list than for data searching. seekeasy is available from computer bulletin boards or for $10 from the national collegiate software clearinghouse. the author requests a $30 contribution if you make use of the software, and that also entiues you to updated editions, when they are released. for more information, contact correlation systems, 81 rockinghorse road, rancho palos verdes, ca, 90274. other programs in the course of this section 1 have limited myself to a discussion of public domain and shareware programs, with the exception of resnoter. all of these programs are available for less than $50, and some can be had for free. potential users should be aware, however, that a large number of commercial programs exists designed for similar purposes. these include ask sam, gofer, zylndex, notebook i1+, fyi3000, and the word processing package nota bene, which includes an interface to the fyi3000 text database system. some conventional dbms packages, such as dbase iv, have added features that allow them to cope with large bodies of textual data as well. potential users should also be aware, however, that the price of these programs can range from the moderate to the stratospheric. while i would not discourage anyone from investigating some of these programs, i have not found that the increased costs purchase significant increases in power or sophistication intrps." what the increased costs do buy is support. if you are uncomfortable with, or inexperienced in the use of computers, it may be worth spending some extra money to gain access to software support personnel. for those who have moderate computer experience, however, public domain software and shareware come very close to being the proverbial free lunch, and i would encourage you to investigate those sources first. summing up quantitative researchers, with their relatively simple data, have been the first to benefit from the computer revolution. but the increasing speed and power available through microcomputers makes even the complex textual data of qualitative researchers more accessible. this article has described the ways in which common, inexjiensive, trp systems may be useful in dealing with large quantities of interview data. the approaches described here can be applied as well to other textual data — field notes, fw example, or archival research entered through text scanners. any textual data can be made more useful through the application of simple computerized tools. the availability of these tools does not, however, absolve the analyst of his or her responsibility. trps can only retrieve and display data— they cannot understand what those data mean, and they will wiuingly supply answers to queries whether those queries are motivated by theoretical understanding or conceptual blindness. computers are always increasing in power, but never in intelligence, and it is worth remembering the first principle of data processing— gigo '^— whenever one sits down at a keyboard. always think of the computer as an exacting but unimaginative research assistant, and you will not go far wrong. finally, you should be aware that trp programs are rapidly increasing in power and hexibility and that, by the time you read this, there will probably be new versions available of most of the programs discussed here and a host of new trps as yet undreamed of. to find out about the latest programs, contact your computer center or local users' groups. a little lime spent looking at the available trp programs will be rewarded with a simple but powerful data retrieval tool. * for helpful comments on earlier drafts of this paper, i would hke to thank theresa marchant-shapiro, charles tidmarch, martha muggins, renata tesch, and several referees, who shall remain nameless. ' presented at the lassist 90 conference held in poughkeepsie, n.y. may 30 june 2, 1990. ^ grep stands for general regular expression print. regular expressions are ways of expressing complex patterns of letters and numbers. grep was originally designed to search through lists using these expressions and print the results on a teletype terminal. ' public domain software is a body of programs placed by their authors into free public circulation: the programs can be freely copied, used and given away but cannot be sold for profit. computer hobbyists often trade these programs and they are also available through electronic bulletin boards and services such as compuslassist quarteriy erve. finally, there are companies that sell public domain software through catalogs for a 'copying fee,' which is usually no more than a few dollars per disk. pubhc domain software should be differentiated from shareware, where the author of the software freely distributes his or her programs, but asks for a contribution from those who use them. both public domain and shareware are excellent sources fw useful and unusual programs. note however that, for your own peace of mind, you should carefully test such programs. a'ever test new programs on the machine you use fw stoting interview transcripts and book chapters: programmers sometimes, though rarely, accidentally release programs with bugs in them, and it's best to find out without destroying irreplaceable materials. * grep is an extremely flexible tool, capable of rapidly seeking out particular patterns in your text database. rather than go into details here, however, i refer you to the support personnel at your institution. if you have access to a unix-based computer, however, you should be able to get a comprehensive overview of grep with the following command: man grep this will display the pages of the unix manual dealing with grep on your terminal. these will give you some sense of the power of the program.if you are using an ms-dos based computer, on the other hand, your msdos manual will give you an oudine of how to make use of the find program. ' for computational purposes, and for the purposes of this paper, a paragraph includes all single-spaced text that occurs between sets of double carriage returns. ' all of the quotations in this article are constructs based on a series of unstructured interviews i conducted in 1988-89. ' typically, you have an ascii file if, when you use the ms-dos "type" command to show your file on the screen, lines end without wrapping around from the right to the left, and you see only alphabetic, numeric, and punctuation characters on the screen. unfortunately, most of the more powerful word processors do not create pure ascii files. wordstar and wordperfect, to name two popular word processing programs, include special codes in their files to make printing easier. such codes must be stripped out if the file is to be searched by most trps. some word processing programs solve the problem of these codes with a built-in option to save files in ascii format. in wordperfect, for example, you should save your transcription into a "dos text' file. this will be an ascn version of the file, with all special characters removed. for many other word processors, you will need a special conversion program. for the most part, such programs are available free or at a nominal charge. if your word processing jmxjgram is incapable of writing an ascii file, go to your college or university microcomputer lab or computer center, and explain what you need to do. they should be able to help you find a suitable conversion program. ' fuzzy searching is a term that covers a great deal of ground. in general it means one of two things. if they do not find an exact match, some programs will look for words or phrases that contain many of the same characters in the same order as the search phrase. others will seek wotds or phrases that are phonetically similar to the search phrase. ' we might think of each 'record' as being made up of a chain of dummy variables (words). each word in the record indicates the presence of a characteristic, and the absence of a word indicates the absence of that characteristic. in any given search, therefore, we are trying to discover whether particular dummy variables are present or absent. at the same time, we will be ignoring most of the variables in a particular record— all of the words that do not appear in a search command. '° given that you can only search for words that you have explicitly coded, you may want to think carefully about the amount of work involved in such coding before choosing this trp. its power comes in large part from a great deal of time and preparation on the part of the user. coding 80 pages of interview (the outcome of the three hour interview i used to test programs) is, to say the least, a nontrivial investment of time. " while some commercial program may have slight speed advantages over their public domain competitors, the major constraint on search speed is likely to be the access speed of the hard disk in your computer. since this affects all programs equally, and is the major constraint on data retrieval, retrieval speed should not be given undue weight in deciding between two trps. '^ "garbage in. garbage out" spring 1991 45 vol29-2.indd iassist quarterly summer 2005 by by elizabeth hamilton1 providing context for understanding: insight from research on two canadian health surveys introduction a significant question for data producers, as well as for the iassist community, is whether today’s data documentation preserved in the form of user guides and codebooks has all of the information necessary for the analysis of a survey. this is particularly important as the community wrestles with the adoption and implementation of data documentation initiative (ddi) projects to describe data from many of our national surveys. to examine this question, i used a case study based on major population health data in canada and employed a life-cycle perspective. i found that, with respect to the national population health survey (nphs) and, more recently, the canadian community health survey (cchs), much of the important information and research related to placing the survey data in context is derived from activities that precede extensive data analysis.2 it is the argument of this paper that the work involved in the creation of a data collection, which occurs early in the data life cycle, is integral contextual information and as such should be identified, described and preserved, in addition to the formal data collection itself. background launched in 1994-1995, the national population health survey is significant in that it was statistics canada’s first national longitudinal health survey designed to fill a specific and critical data gap in health information: the determinants of the health of canadians over time.3 the problems related to health data were extreme from the perspective of researchers and policymakers alike. in a summit on health information sources, participants enumerated the problems with existing data sources.4 while numerous statements revealed frustrations with the inadequacy of the state of health information, some critics went further and described the state of health information of the era as being in a deplorable state, suffering from a lack of comparability and with serious gaps in coverage. the nphs was a key component in the new health information infrastructure for canada, and it was the hope of statistics canada that flowing from this survey would be the development of health indicators akin to economic indicators for the country, a theme that has been echoed by many in the health field since the inception of the survey. not surprisingly, the expectations for this survey were enormous. by the spring of 1992, statistics canada had treasury board’s support of the project, and research and consultations on the design and methodology of this new longitudinal health survey were well under way. the primary survey instrument was finished by fall 1992, a time frame that, in the best of circumstances, did not allow for the luxury of repeated revisions. the pressure was considerable. the project managers had to devise an instrument that would fill the data gaps over more than just one survey cycle; they had to create the appropriate content, questions, and scales to collect person-oriented health information and meet the needs of researchers and policymakers over a 20-year period. and, as with many national surveys in canada, the sampling was complex to accommodate political and social realities in the country. additionally, companion surveys to the primary household survey were developed for institutions and for the traditionally under-surveyed northern areas of canada. the process of launching the survey came with guidelines that emanated from the task force on health information. the survey was to be flexible and statistically reliable, with timely release of data. further, the process was to be consultative, allow for supplementary content or sample size, and permit linkage with administrative data. the guiding force came from the project team within statistics canada with oversight by a federal-provincial-territorial committee. the project manager brought together expert groups of six to ten people as needed to decide upon about ten minutes of questions to capture data necessary on mental health, health measurements, and other key content areas.5 the questions were brought to focus groups and the questionnaire was modified for field-testing. the sampling design was drawn up, modified, and re-modified. linkage with administrative data was a key component of the nphs, and in the survey respondents were asked if they would permit information to be shared with provincial health departments, health canada, and human resources development canada (hrdc) for statistical purposes. to 18 iassist quarterly summer 2005 diagram 1 solve the dual problems of confi dentiality and researcher access, the nphs team produced three cross-sectional fi les, plus a share fi le, and established dissemination practices that were designed to meet the needs of most researchers.6 the fi rst article by statistics canada on the survey results came in september 1995,7 the announcement of the release of the public use microdata data fi le came in november 1995,8 and the fi rst graduate thesis using nphs was granted in the spring of 1996.9 the nphs had not been in the fi eld long before new data gaps were identifi ed and addressed through a new biennial cross-sectional survey, the canadian community health survey (cchs). as the national population health survey ceased to produce public use microdata fi les and use of the data was restricted to those who met the statistics canada’s research data centres criteria, the cchs provided personoriented health data at the sub-provincial level (health districts) and accommodated the need for periodic inclusion of special topics and special populations, such as mental health component of cchs cycle 1.2 for both the general population and the canadian forces. methodology at the outset of this research project, the methodology used to discover evidence of nphs use was that of standard literature reviews. established peer-reviewed databases were systematically searched to fi nd proof of data use and knowledge transfer related to the nphs. there was compelling evidence from this research that the data fi le was being used extensively, just as the survey planners had intended. since the inception of the survey, there have been over 60 theses and dissertations written, over 400 articles published in 147 journals around the world, and in subject areas as varied as veterinary science, kinesiology, and cardio-thoracic research. workshops and conference presentations abound, as do examples of the use of the data fi les in teaching health research courses. the research also revealed the use of information related to the early stages of the survey in analytical studies. the research methodology was accordingly expanded to fi t the framework of the survey life cycle (diagram 1). while there are many representations of the life cycle of a survey, this framework contains those broad stages that are familiar to all project managers: the data gap analysis, the planning and administration of the survey, the release of the survey data and results, and fi nally, the evaluation of the survey in terms of its future. search techniques and sources were revised to uncover products generated from all stages of the life cycle.10 findings the research revealed much more than simply articles iassist quarterly summer 2005 19 analyzing the data files for the nphs. statistics canada employees produced approximately 30 articles or reports on various aspects of the survey, dealing with topics as disparate as sample retention, consideration of weighting techniques, evaluation of statistical packages for nphs analysis, the creation of dummy files, and the protection of confidentiality. there was no argument that the survey was successful in the public aspect of knowledge transfer; the database now totals some 800 items relating to the nphs. the follow-up question was why the survey was so successful in meeting its initial goals. discussions with the nphs team yielded more information on the extent of internal documentation, such as training manuals, and interviewer feedback reports that help explain this extraordinary success. a) survey guides the detailed user guide accompanying the public use microdata file reflects the information that statistics canada considered necessary for data analysis (diagram 2). it is notably more complete than most documentation disseminated with data files. the guide for nphs cycle 2 runs to 1060 pages (though admittedly, the inclusion of the ontario health survey questions added bulk to the codebook). nevertheless, does it have all the information that is going to help researchers analyze the survey properly? are the traditional documentation and data files enough to ensure appropriate use of the survey data in future years? though the user manual for this survey is extremely rich, i would argue that it is not enough simply to capture the user guide for preservation. at a micro level, information on the question content and interview training is invaluable for focused research; at a macro level, life-cycle information provides context for understanding (and guiding) critical issues throughout the life of the survey. b) background reports/training materials/administrative documents there is an abundance of material produced across the life course of a survey that can help researchers understand the data. in preparing for the nphs, for example, there were reports from the expert groups on selected areas of the questionnaire, developmental studies on new areas of health, questionnaire focus group reports, treasury board documents, studies on sample selection and methods to be used for estimates, coding concordances between survey cycles, analytical exercises for researchers using different software, and the results of post-field feedback surveys for interviewers. statistics canada does pay special attention to training interviewers working on critical surveys, such as the census of population or surveys requesting sensitive information. the interviewer training materials for cycle 1.2 of the canadian community health survey were particularly thorough because the survey’s content focused on the state of mental health on a national level—a true challenge of the skills and training of statistics canada interviewers. protocols in the manual included the handling of difficult respondents, referrals of distressed individuals to a list of resources, and contact information for mental health professionals for the use of interviewers to allow them to decompress after difficult sessions. in the example presented in figure 1, the cchs cycle 1.2 training manual provides information about handling sensitive questions. is this information important to a researcher? depending on whether this information relates to a question in her or his research area, it could be. one of the significant problems in investigating physical or substance abuse is presenting questions in a way that captures accurate replies without creating a biased social response. why is this useful information? although there are many strategies for handling sensitive questions, it is extremely nphs codebook (cycle 2) table of contents background objectives survey content sample design data collection data processing data quality guidelines for tabulation, analysis & release approximate sampling variability tables weighting file usage questionnaire record layout (general and health) data dictionary derived and grouped variables cv tables list diagram 2 20 iassist quarterly summer 2005 valuable when interpreting the survey results to know the specific instructions given to the interviewers. for instance, if the survey was interrupted and the interview continues with a subsequent denial of a previously positive answer, how was this handled? in another example from the nphs manual, interviewers are told that feedback on a question of social support indicated that the questions were problematic for members of certain cultures. in such cases, interviewers were to put “no response” rather than receiving and recording a negative answer.13 why is this potentially useful information? immigrant studies attempt to identify factors that lead to a sense of belonging and social support. if particular cultures regard social support networks as not being applicable in their lives, it will help our understanding of the functioning of these cultures within and alongside the overall canadian cultural norms. if the text or context of the question has prompted a consistent non-response pattern for a cultural group, that is equally important information. c) methodological testing and documentation on a broader level, the background of the nphs is a compelling one, and understanding the role of the nphs within the revamped health information infrastructure highlights the importance of issues under debate today. two critical problems at this point, for example, are the attrition in the sample and the issues of privacy and data quality concerned with file linkage of the nphs to administrative data. initial funding was provided for a sample size of 12,767 individuals, to be tracked over 20 years. a critical problem with the nphs as it heads into the mid-point of the life of the survey is attrition within the original sample. the research on methods to retain the respondents of the original sample and the decisions on how best to make the survey sample robust are vital to the continued viability of this survey. the research community, for its part, needs to understand the efforts invested toward this end and support those efforts. i would argue that researchers benefit in their analysis by knowing the context of the original sample selection and by understanding the dynamics behind the retention of the survey respondents. in this way, they can frame their questions appropriately and interpret analysis accurately. this knowledge can also to contribute to research on sample selection for large longitudinal surveys. much of this information is not in the current documentation but has been discussed in internal documents produced in conjunction with this survey.14 in supporting the development of the nphs in 1991-1992, the chief statistician cited linkage as a key part of a robust health information system. it was assumed, for example, that certain causes of death in the mortality database might be traced back to nphs respondents, and causal relationships on risk factors could be investigated more thoroughly. in the information leading up to the survey, it was argued that a direct product of linkage would be the improvement in the key determinants of health and a more effective use of health resources. linkage is not a simple issue, though, particularly in a country that has responsibility for health split between the federal and ten dep_q26ee1a during the last 12 months, did experience c happen to you? r.: well, yes, there was one night, we were drinking and i just... q. at this point the respondent stops talking and you see a tear come to her eyes. what do you do? exercise-minimize non-response tactics: sad/upset respondent: stop for a minute. be responsive to the respondent in a supportive way; give the respondent a chance to collect themselves, and help them get on track with the interview. 1.“that must have been very upsetting”.....offer to take a break from the interview. 2. ask if the respondent is able to continue with the interview. 3. if not, offer to continue the interview at a later date. fiqure 1 :training excercise cchs iassist quarterly summer 2005 21 provincial governments—and three territories. the barriers and potential solutions to linkage issues inevitably escape the survey user guide, but are recorded elsewhere. soliciting information on how the survey producer and researchers want to see the data analyzed may help to determine the kinds of information to preserve throughout the survey life cycle. if, as the task force on health information asserted, analysis is a process to increase human perception of the significance of data, it is easy to see the degree of richness these additional documents offer to data analysis and interpretation. ancillary documentation can help frame an issue so that the analysis is more meaningful. consider, for example, the value of knowing the following types of information: • what solutions are there to data linkage issues? is there a response at the researcher or policy level that can alleviate technical and legal obstacles? • why did the wording of a question change over time? are the responses still comparable or did the improvement in the clarity of the question change the distribution for this question? • why this particular question? why not another wording that would give greater precision in the response? • why are respondent numbers declining and how do we maintain the reliability of the results? • why this content? why not broader or more specific content? • what scales were adopted or modified for use in this survey? why? • what can the interviewers tell us about the receptivity to more health questions, of geographical or cultural problems with data gathering at an individual or community level? • how were difficult issues resolved? what was considered—and what was rejected? • what is the relationship between this survey and other surveys of similar subject matter? does this question, or order of questions, intentionally replicate the wording in other surveys? • how has the sample changed (over surveys and over time) at the micro level? who are those who have ceased to participate? do they have lower incomes? those in better health? frequent movers? • what should i know to analyze this survey appropriately? in discussions with the project teams of the nphs and cchs, it is clear that they invested a great deal of talent and energy into the design and administration of these surveys. it is not clear, however, that they always appreciate the value of the supporting documentation. i would argue that the documentation associated with the survey life cycle is too critical to be separated from the data itself. but does it all have to be included in the user guide or relevant equivalent (such as ddi encoding)? not necessarily, but while some documents are inappropriate for public release because of confidentiality and disclosure risks, researchers should know, at a minimum, what has been produced. intellectual and physical control at this point, none of the user guides produced by statistics canada include a complete listing of the documentation associated with a survey’s development. the integrated metadata data base (imdb) has become more inclusive with regard to documentation produced as part of a survey, but it does not come close to recording all the studies, reports, and evaluations produced in the life cycle of the survey.15 with continuous surveys, the personnel associated with the surveys will move on. at best, their expertise and knowledge can only be partially transferred to their successors. ironically, those who least appreciate the usefulness of the evolutionary work involved in a project are those who are responsible for the creation of the survey and for the ancillary documents involved in its creation, production, review, and revision in the first place. another key question is whether all of the documentation has to be physically stored or made accessible with the survey data. i would argue that the physical storage is not as critical as the intellectual linkage between the information produced in conjunction with a survey and the final version of the documentation and data file. this is a critical concept to grasp. not all researchers need all of this information—in fact, it can be a disservice to those just beginning to work with a survey. however, the person investigating asthma will want to know that there is conceptual background for the content, that there is documentation on the variations in questions on asthma or asthma medications from cycle to cycle, and that there is a very good explanation as to why the question on pet ownership was omitted after the first cycle. these examples refer to documentation directly related to the production of a survey, but the existence of other information is of equal value. certainly the detailed survey methodology must be preserved even though the survey creators cannot release much of this information because of confidentiality constraints. similarly, the documents prepared for the public release microdata committee may have similar confidentiality concerns. the need identified at this juncture is simply to know about the existence of this information by recording and linking it intellectually to the survey. failure to understand that intellectual control is the highest priority puts information at risk of being suppressed, lost, or discarded as ephemeral, unimportant, or 22 iassist quarterly summer 2005 dangerous. much like a publicly available report released in conjunction with a royal commission, the enduring studies provide lists of evidence and ensure that the background documents are all captured and archived. while the report itself will be preserved as part of the traditional function of the national library, the archives provide the security and archival techniques to preserve the audio, textual, and electronic forms of a commission’s work. conclusion my research on the use of the nphs and cchs has indicated that the intensive planning involved in the survey has produced an extremely rich data file that has endured through the first half of the project. it has also unearthed a wealth of information that points to reasons why the nphs, in particular, has been so successful in ensuring that the data are used. it is axiomatic that data are only valuable when analyzed and interpreted. analysis and critical assessment of the data are the true mark of success of a survey. as front line data curators, librarians, and archivists, our community champions access as a fundamental principle of democratic societies, a principle that, in practice, encourages ground-breaking investigation and research. it is also a cost-effective and sustainable way to produce an understanding of our society. statistics canada has neither the funding nor the mandate to conduct research, particularly in controversial areas. as iassist members push forward in establishing protocols in the description of survey data for discovery and retrieval through ddi, it is incumbent upon us to think about the parameters of the information to be included in ddi information. though it may not be imperative to capture and encode all the information generated through the life cycle of the survey, it is impossible to make an informed decision about inclusion and exclusion without knowledge of what was produced. on their part, the research community has a mandate to delve into societal problems, and also the infrastructure to support long-term projects, including the type that will eventually produce the results needed to reshape a massive challenge like health care in canada. providing the resources to do this research is necessary; but the knowledge of the documents that exist to help explain and contextualize the variables and the data file is nothing short of critical to the success of this joint endeavour. while data professionals may understand this better than the survey producers, it is worth asking whether archiving the documentation through the life cycle of a survey matters to the rest of the world. i would contend that it does. in a recent senate committee hearing, several senators questioned statistics canada on the interview techniques used, and why there were different numbers used in the testimony of others appearing before the committee. senator lavoie-roux: my questions have already been asked more or less by my colleagues because i was wondering how the data were collected. it is easy to establish how many children complete their primary, secondary or university schooling because diplomas or certificates are granted. to me it was really an important issue to know to what extent your statistics could be trusted—i don’t say that in a negative way—since the methods used for data collection are fairly weak. … with the methods that you use, do you believe that we can trust the data? are you sure that it is accurate?”16 it would be reassuring to say that yes, we know that we can trust the data. we have a complete record of evolution of that question, from concept to analysis. references beaudoin, gabrielle. “national population health survey, cycle 6: respondent relations.” workshop on respondent relations, statistics canada. ottawa. february 23, 2004. beland, y. (2002). “canadian community health survey: methodological overview.“ health reports, 13, 3 (2002): 9-14. belanger, alain, berthelot, j.m., martel, l. ”canadian national population health survey and the calculation of canadian healthy life expectancy indicators for the next decade.” international network of health expectancy, reves 11 conference. london. 1999. bernier, julie et al. ”comparison of the health utility index mark iii (hui iii) and the self-assessed health status in the canadian population.” international society for quality of life research (isoqol) annual conference. amsterdam. 2001. bernier, julie et al. ”handling missing data.” statistical society of canada annual meeting. toronto. 2002. canada. parliament. senate. standing committee on social affairs, science and technology. proceedings: evidence. 36th parliament, first session: issue11, 13 may 1998. canada. task force on health information. health information for canada, 1991: report of the national task force on health information. ottawa: health canada, 1991. canadian institute for health information. health information roadmap: beginning the journey. ottawa: the iassist quarterly summer 2005 23 institute, 1999. canadian institute for health information. national consensus conference on population health indicators final report. ottawa: the institute, 1999. catlin, gary. ”shaping a vision for health statistics.” national committee on vital and health statistics 50th anniversary symposium. washington. 2000. <http://ncvhs. hhs.gov/50thcatlin.pdf > (25 june 2005). catlin, g., & will, p. “the national population health survey: highlights of initial developments.” health reports, 4, 3 (1992): 313-329. conference of deputy ministers of health (canada). federal/provincial/territorial advisory committee on population health. strategies for population health: investing in the health of canadians. [ottawa]: the committee, 1994. dale, a., arber, s. & procter, m. doing secondary analysis. london: unwin hyman, 1988. fobes p. & geran l. “cycle 2 and beyond: preparing and storing longitudinal data of the national population health survey.” statistics canada symposium 98. longitudinal analysis for complex surveys. ottawa. 1998 gentleman, j., & tomiak, m. “the consistency of various high blood pressure indicators based on questionnaire and physical measures data from the canada health survey.” health reports, 4,3 (1992): 293-312. helmer, d. “shooting from the hip or target practice? a comparison of conventional and fugitive search results.” medical library association/canadian health library association annual conference. vancouver. may 2000. helmer, d., et al. bibliography on systematic reviews: locating, understanding and using the evidence. bcohta working document, no. 1. vancouver: office of health technology assessment, 1999. helmer, d., savoie, i., & green, c. “evidence-based practice: extending the search to find material for the systematic review. bulletin of the medical library association, 89, 4 (2001): 346-352. helmer, d., savoie, i., & green, c. “how do various fugitive literature searching methods impact the comprehensiveness of literature uncovered for systematic review?” international conference on grey literature. washington, dc. 1999. mantel, h.j. & nadon, s. “dummy file creation for the remote access program of the national population health survey.” in statistics canada. survey methods section. proceedings. ottawa: statistics canada, 1999. mathieu, patrice. “longitudinal research using the national population health survey.” situating place in health research: putting theory into practice. kingston. 2001. norris, d. “using survey data to address policy issues related to health and social support.” canadian association on gerontology annual scientific and educational meeting. halifax. 1998. perez, claudio e. using the bootstrap technique for nphs analysis. ottawa: health statistics division, statistics canada, n.d. roberts, georgia, et al. “bridging the gap between the theory and practice of analysis of data from complex surveys: some statistics canada experiences.” federal committee on statistical methodology research conference. washington. 1999. savoie, i., & helmer, d. ”extended systematic vs. conventional search methods: weighing the quality of the literature retrieved.” international society for technology assessment in health care annual meeting, philadelphia, pa. 2001. savoie, i., helmer, d., green, c.j., & kazanjian, a. “improving the efficiency of the fugitive search: effectiveness of various methods. international society for technology assessment in health care annual meeting. the hague. 2000. scherer, r.w., & langenberg, p. “full publication of results initially presented in abstracts.” cochrane methodology review, 2 (2001). shields, margot. “proxy reporting in the national population health survey.” health reports, 12, 1 (2001):.21-40 (english), 23-44 (french). singh, m.p. et al. national population health survey: design and issues.” in american statistical association. section on survey research methods. proceedings, 1994. alexandria: american statistical association, 1994. http:// www.amstat.org/sections/srms/proceedings/papers/1994_ 138.pdf> (25 june 2005). statistics canada. canadian community health survey, cycle 1.2: mental health and well-being, project code 6502-0 training guide. ottawa, statistics canada, [n.d.]. statistics canada. information about the national population health survey. (82f0068xie). ottawa: statistics canada, 1999. http://www.statcan.ca:8096/bsolc/ 24 iassist quarterly summer 2005 english/bsolc?catno=82f0068x (accessed 1 july 2005). statistics canada. national population health survey 1996-97 household component: user’s guide for the public use microdata files. ottawa: statistics canada, 1998. statistics canada. national population health survey answers to the exercise to minimize non-response, cycle 6, project code 0104-0. [ottawa: statistics canada, n.d.] statistics canada. national population health survey debriefing questionnaire 6, cycle 6 – quarter 4 – 0104-0. [ottawa: statistics canada, n.d.] statistics canada. national population health survey interviewer’s manual, cycle 6, project code 0104-0. [ottawa: statistics canada, n.d.] statistics canada. national population health survey selfstudy, cycle 6, project code 0104-0. [ottawa: statistics canada, n.d.] statistics canada. national population health survey training guide, cycle 6, project code 0104-0. [ottawa: statistics canada, n.d.] statistics canada. cycle 4 (2000-2001) household component longitudinal documentation. ottawa: health statistics division, statistics canada [unpublished], may 2002. stephens, t. measuring the health of canadians: an agenda for developing health surveys. health reports, 3, 2 (1991): 137-145. stukel, d.m., mohl, c. & tambay, j-l. “weighting for cycle two of statistics canada's national population health survey.” statistical society of canada. proceedings of the survey methods section. ottawa: statistical society of canada, 1997. swain, l., caitlin, g., & beaudet, m.p. the national population health survey – its longitudinal nature. health reports, 10,4 (1999): 69-82. tambay, j-l., & catlin, g. “sample design of the national population health survey.” health reports, 7, 1(1995): 29-38. tambay, j-l. et al. “treatment of nonresponse in cycle two of the national population health survey.” survey methodology. 24 (1998): 147-156 (english), 159-169 (french). tolusso, s. & brisebois, f. nphs data quality: exploring non-sampling errors, methodology branch working paper, hsmd-2003-004e. ottawa: statistics canada, 2003. umphrey, g., kendall, o., & macneill, i.b. “assessing the surveillance capability of canada’s national health surveys.” chronic diseases in canada, 22, 2 (2001): 5056. wolfson, michael c. “toward a system of health statistics” daedalus. 123 (fall 1994): 181-195. yeo, d. “after the first steps: the evolution of a longitudinal survey.” workshop on longitudinal research in social science—a canadian focus. london, on. 1999. yeo, d., mantel, h. & liu, t.p. “bootstrap variance estimation for the national population health survey.” in proceedings of the american statistical association. survey research methods section. alexandria, va: the association, 2000. footnotes 1 elizabeth hamilton is head, government documents, data and maps department at university of new brunswick. contact: hamilton@unb.ca the paper is a revision of material presented at the iassist 2005 conference in edinburgh. 2 the context of a survey has been recognized by authors such as dale, arber and procter (1988) as important in secondary analysis, but has received little attention within the context of identification, documentation and preservation of master files and public use microdata files. 3 for further information on the background of the nphs, see: d. yeo, after the first steps : the evolution of the national population health survey, canadian studies in population, 28; 2 (2001):.377-390. 4 “the lack of standard data definitions, minimum data sets, standard edit rules, as well as the quality control and security procedures is completely unacceptable.” extract from the report of the project team on comparability, in canada. task force on health information. health information for canada, 1991: report of the national task force on health information (ottawa: health canada, 1991), 9. 5 the research process had begun long before the approval of funding, of course. for example, blood pressure indicators had been a subject of study in the early 1990s, and there had been a federal-provincial-territorial committee on mental health before nphs was born. 6 two of the programs that evolved from the desire to promote timely access to data were the highly successful data liberation initiative (dli) launched in 1996, and iassist quarterly summer 2005 25 the research data centres (rdcs). nphs project managers contributed products through all statistics canada’s dissemination channels, including tables in health indicators, published articles (in both free and priced publications), public use microdata files through dli, master data files through the rdcs, and remote job submission and custom tabulations for individualized data tabulations. 7 “health of canadians, 1994” in the daily, 22 september 1995 8 the daily, 21 november 1995 9 paula fletcher, falls among the elderly: risk factors and prevention strategies. phd, university of waterloo, 1996 10 see, for example, d. helmer, i. savoie, & c. green, “evidence-based practice: extending the search to find material for the systematic review. bulletin of the medical library association, 89, 4 (2001): 346-352; and d. helmer, “shooting from the hip or target practice? a comparison of conventional and fugitive search results.” medical library association/canadian health library association annual conference. vancouver. may 2000. 11 statistics canada. national population health survey 1996-97 household component: user’s guide for the public use microdata files,. (ottawa: statistics canada, 1998), 3-5.. 12 statistics canada. canadian community health survey, cycle 1.2: mental health and well-being, project code 6502-0 training guide (ottawa, statistics canada, [n.d.]) 183-184. 13 statistics canada. national population health survey: cycle 6: project code 0104-0: training guide (ottawa: statistics canada, [n.d.]), 124. 14 beaudoin, gabrielle, “national population health survey: cycle 6 respondent relations (ottawa: statistics canada, 2004). 15 the entry for the cchs includes the information that there were studies evaluating cchs results in relation to various other surveys, and a study that compared the profile of respondents agreeing to share their data with other departments was conducted. http://www.statcan.ca/english/ sdds/3226.htm 16 canada. parliament. senate. standing committee on social affairs, science and technology. proceedings: evidence. issue 11, 13 may 1998. vol29-2.indd 10 iassist quarterly summer 2005 by stuart macdonald & luis martinez 1 the local data support landscape in the uk introduction this paper will report on existing data support infrastructures within the uk tertiary education community. the paper will then discuss early methods and traditions of data collection within uk territories. in addition it will focus on the current uk data landscape with particular reference to specialized national data centres which provide access to largescale government surveys, macro socio-economic data, population censuses and spatial data. it will conclude by outlining examples of local data support services with particular reference to the ‘accidental’ nature of the data librarian, their organizational role and areas of expertise in addition to future developments. data in the uk the 'tradition' of data collection within the united kingdom can be traced back to the 7th century 'senchus fer n'alba' in gaelic scotland (translated as tradition/census of the men of alba). the original document is now lost but was translated in the 10th century and is regarded as the earliest native tax assessment known in britain. it is part genealogical record, part inventory of the territories of the descendents of eochiad muin and consists of a list of the numbers of men that the various families of the scoto-irish kingdom of dal riata (centred on modern-day argyll) could provide for their navy. in 1086 the domesday book (an inventory of land use and property in england) was commissioned by william the conqueror for administration purposes. the 17th and 18th centuries saw a number of countries conduct national censuses e.g. 1666 quebec province, 1703 iceland, 1749 sweden and 1790 usa however it was not until 1801 that the first comprehensive uk census was conducted, partly to ascertain the number of men able to fight in napoleonic wars, to be subsequently carried out on a decennial basis. the need for a central statistics office was recognised as early as the 1830’s with the introduction of tables of revenue (a type of statistical year book instituted by george richardson porter from the board of trade). a decision to create a cso was agreed in 1880 but never came to fruition and, indeed, it wasn’t until 1941, with the aim of ensuring coherent statistical information, that the central statistics office was founded by winston churchill. following the advent of mainframe computing the ssrc data bank was established at the university of essex in 1967 (later to become the uk data archive2). 1971 saw the social statistics laboratory at strathclyde university3 being set up in addition to the appearance of the first microprocessor. in 1981 the first ibm desktop pc was appearing in the midst of the developing internet. the first data library based in a uk tertiary education institution was set up at edinburgh university in 1983. in 1992 the world wide web was released by cern (the european organization for nuclear research) and in 1996 the cso merged with the office for population, censuses and surveys to become the office for national statistics. from the abridged account above it is evident that the collection, organisation and analysis of records about people is nothing new. what is new, in relative terms, is the microprocessor and pc in addition to advances in telecommunications and web technologies. and as robin rice suggested in the recent article ‘the internet and democratisation of access to data’ (as part of an online discussion for esrc social sciences online – past, present and future4) “data collection has arguably been changed more by computers than analysis itself, which has been dominated for decades by a few well-known statistical, and qualitative, analysis packages” analysis of large research datasets at the desktop requires a different set of skills to those of data discovery. these skills include tools to make the data usable in addition to a familiarity with the construct of the dataset. thus it is as a result of this march of technology that there have emerged data professionals who not only have the necessary data discovery skills but also provide access to, support and train those wishing to use research and statistical data. iassist quarterly summer 2005 11 government statistical services the office for national statistics (ons) is the government department that provides uk statistical and registration services in addition to planning and conducting the decennial census for england and wales. it is also responsible for producing a wide range of key economic and social statistics which are used by policy makers across government to create evidence-based policies and monitor performance against them. publications include: labour force survey quarterly sample survey of households living at private addresses in great britain. its purpose is to provide information on the uk labour market that can then be used to develop, manage, evaluate and report on labour market policies. family expenditure survey continuous survey of household expenditure and income that has been in existence since 1957. annual samples of around 10,000 households (about 1 in 2000 of all united kingdom households) are selected each year. general household survey continuous national survey of people living in private households. the main aim of the survey is to collect data on a range of core topics, covering household, family and individual information. health survey for england annual surveys about the health of people living in england. it began in 1991 and has been carried out annually since then. a number of core questions are included every year but each year’s survey also has a particular focus on a disease or condition or population group. the general register office for scotland (gros)5 is charged with a similar role in scotland. gros is the department of the devolved scottish administration responsible for the registration of births, marriages, deaths, divorces, and adoptions in scotland as well as conducting and publishing the output from the scottish census. in addition, the scottish executive6 provides relevant and reliable statistical information, analysis and advice that meet the needs of government, business and the people of scotland. they are also responsible for commissioning a number of surveys specific to scotland, these include: scottish household survey scottish crime survey scottish health survey scottish house condition survey the northern ireland statistics and research agency (nisra) and the statistical directorate of the national assembly of wales have similar functions within the other uk territories. uk data centres for tertiary education there are a number of uk national data centres offering the uk tertiary education and research community networked access to a library of data, information and research resources. in the majority of cases services are available free of charge to members of uk tertiary education institutions for academic use, although institutional subscription and end-user registration are required for most services. the uk data archive (ukda), based at the university of essex acts as a repository for the largest collection of digital data in the social sciences and humanities in the uk including the aforementioned economic and sociodemographic survey data, amongst many others. its remit includes data acquisition, preservation, dissemination and promotion of social scientific data. it also houses a major collection of computerised historical material and is responsible for the census registration service which facilitates access to census data resources for uk higher and further education. the ukda is a lead partner in the economic and social data service (esds)7 which was launched in january 2003. esds is a distributed service, based on the collaboration between four key centres of expertise, ukda, mimas, the institute for social and economic research (iser), and the cathie marsh centre for census and survey research (ccsr). it acts as a national data service providing access and support for an extensive collection of quantitative and qualitative datasets for the research, learning and teaching communities. it is organised in five sections : • esds government – dedicated to the government surveys such as the general household survey. • esds longitudinal – supporting the use of the uk longitudinal collection, datasets like british household panel survey or british cohort study. • esds international provides access to a range of international macro-economic datasets from organizations like the un, eurostat, oecd and imf. • esds qualidata – specialist service providing access and support to qualitative datasets. data includes interviews ,focus groups, personal documents and photographs. • esds access and preservation – focuses on the data acquisition, processing, preservation and dissemination. hosted by edinburgh university data library, edina8 12 iassist quarterly summer 2005 was launched as a jisc-funded national data centre in january 1996. services include abstract and index bibliographic databases such as biosis, cab abstracts; spatial data services such as digimap (which provides the uk tertiary education community with access to a range of ordnance survey mapping products), ukborders, euroglobalmap and edina agcensus; multimedia services such as education media online (emol) and the education image gallery (eig). edina are also partners in learning and teaching projects (jorum and nln) in addition to spatial development projects such as emapscholar and the resource discovery tool go-geo!, and infrastructure projects such as shibboleth development and support services. mimas9 (based at manchester computing at the university of manchester) is a jisc and esrc-funded national data centre, and like edina provides uk higher education, further education and research community with networked access to key data and information resources to support teaching, learning and research across a wide range of disciplines. services include bibliographic discovery tools such as copac, zetoc, jstor and the isi web of knowledge; learning materials such as nln; spatial data services such as landmap which provides access to uk satellite data; census data via the census dissemination unit; international databanks (via esds international). the arts and humanities data service (ahds)10 is a jisc and ahrb-funded data centre which aids the discovery, creation and preservation of digital resources in and for research, teaching and learning in the arts and humanities. it currently consists of 5 distributed hubs: the archaeology data service; ahds history; ahds visual arts; ahds literature, languages and linguistics; ahds performing arts. the hubs collect, preserve, catalogue, and distribute digital resources which are relevant to their subject areas, facilitate good practice in their creation and use, and offer some user services united kingdom local data support services academic institutions provide support for data services in a variety of ways11. this is reflected in the diversity of organisational representatives for esds, for example. some work in a data library, a university library, a university computing centre, a central research office or an academic department. however, the data support offered by data libraries goes beyond supporting the national data services. data librarians/managers deal with the management and implementation of such services. among their multiple tasks, data libraries: • act as data repositories, developing and preserving local collections • serve as reference services helping researchers to identify appropriate resources and troubleshooting • educate users to access and handle data resources through teaching and learning activities. although there are common activities, the level of support and areas of expertise varies amongst data support services. below there are examples of four different local data support services in the uk. edinburgh university data library edinburgh university data library (eudl)12 was the first such service in the uk, started in 1983. it was set up as a small group with a sociology lecturer as part time manager, with 1.5 staff (one programmer and one computing assistant). currently two qualified librarians provide the service, with administrative and technical support from edina. the current collection covers large scale government surveys, macroeconomic and financial time series, population and agricultural census data and geospatial resources. eudl specialises in data for scotland, and gis resources due to its relationship with edina. the data library staff actively participates in local training activities in addition to providing a consultancy service, helping with the extraction, merging, matching and customization of data; for time-consuming jobs a fee is charged. university of oxford data library the university of oxford data library13 started in 1988. three people formed the computing and research support unit, with one statistician, one computer/statistical software specialist and one data manager. at present it consists of one data manager, with no dedicated it support and is part of the nuffield college library . the current collection comprises survey micro-datasets from uk and elsewhere, including the large government continuous surveys, and many ad hoc, repeated crosssection and panel/cohort academic surveys. subsets of general household survey and labour force survey variables combined over time have been compiled, and are widely used by researchers. for seventeen years now the university of oxford data library has supported researchers in using quantitative datasets; some of the key support functions have been: • data searching and conversion • negotiations and management of contracts with data providers iassist quarterly summer 2005 13 • storage, protection and access arrangements • questionnaire design, data collection and management methods, knowledge of question wording, coding structures and methods, data cleaning. london school of economics data library the lse data library14 was launched in 1997 to support lse researchers in the task of locating quantitative data. it provides an advisory service to phd students, contract researchers and academics, helping them to locate and access datasets. its collection includes a microdata archive covering large scale government surveys, longitudinal data collections and international opinion polls as well as a wide range of aggregated databases providing worldwide socio-economic indicators from igos such as imf, oecd or eurostat. it also includes spatial data such as uk boundaries and eu and world administrative regions. lastly, financial databases are increasingly becoming an important part of the collection, providing company accounts, indexes and bond data, exchange and interest rates, etc. the data librarian offers direct support for users through the weekly data surgery (one-to-one advice) and information literacy courses, helping to locate, access, and format data. a data laboratory, datalab, is being implemented that will provide gigabit connectivity between pcs in a computer classroom and a dedicated server. the datalab will be used for using datasets for teaching as a first stage, and will be hosting all microdata and managing access and metadata. a high level advisory group for guidance on academic priorities formed by one academic from each department is also in place. the group will identify academic priorities across lse departments and will shape strategic planning for the future. london school of economics rlab data services the london school of economics rlab15 data services started in 1999 providing data support to lse’s research laboratory, a unique institution bringing together leading research centres in economics, finance, industrial relations, social policy and demography. the centerpiece of the collection is an electronic library housing approximately 150gb of data. the data is mainly social survey data, with some financial, geographical and medical data. both macro and micro datasets are held in the library, data from individual countries throughout the world, as well as a wealth of international sources from the us, europe, india, china. data support is part of the rlab it service, there is a team of 6 people, as well as a fulltime data assistant. the data manager has the support of two systems professionals, and two information professionals to help with the web site. rlab’s data manager, tanvi desai, is now involved in the esrc review of international data sources and needs. the aim of the project is to gain an understanding of the opportunities and the obstacles presented by international data resources, enabling us to recommend strategies for improving international research and collaboration.16 data information specialist committee, disc-uk uk data librarians have always been represented by the international association of social science information service & technology [iassist]. this association also represents international data archives, statistical agencies, government departments and non-profit organizations. however, a need existed for much closer collaboration to deal with data issues within and relevant to the uk. in october 2002, a mailing list “digging for data” was set up as a forum to help data librarians, national data centre site representatives’ and any other academic staff, support staff, statistical consultants or students to locate and use quantitative data. it was an informal initiative amongst uk data librarians represented by oxford university, the university of edinburgh and the london school of economics. in september 2003 the data libraries from these universities formed a support group called disc-uk, or data information specialist committee-uk17, holding their first meeting in april 2004. the group meets several times per year to compare issues and solutions arising from their daily work and its aims are the following: • foster understanding between data users and providers • raise awareness of the value of data support in universities • share information and resources among local data support staff the founding members intend to open up their group to others performing similar roles in their universities, though they may not work in dedicated data libraries. esds site representatives were emailed a simple questionnaire to find out the level of data support offered at their institutions. although this elicited only one response from a site representative who had little to do with data support in his institution, it is an area which disc-uk will pursue further. 14 iassist quarterly summer 2005 a much better response has been obtained when individuals have been approached. universities such as glasgow, birbeck college, southampton and others have been contacted and links with those institutions have been established for future collaboration. an interesting fact to emerge from such contact is that in most cases people doing data support in those institutions are subject librarians dealing with other electronic resources such as bibliographic databases. the website for the group has been set up, with links to member sites and a description of the group's aims. through time it is hoped that the site develops into a helpful resource hosting online training materials and links to relevant articles. channels of communication between disc-uk and the national data centres are already in place. esds workshops have been organised in each of the member institutions and several improvements suggested to uk data archive’s administration web interfaces. the next step for the group is to plan and coordinate the necessary resources to act as a broker between member institutions and the national data centres in addition to investigating the viability of running ‘train the trainers’ events. the arrangement of regular meetings with such centres will benefit both, establishing means to provide feedback and develop common strategies for the promotion of the data hosted at the national data centres. accidental data librarian so where are the data librarians coming from? is there a common educational background? what are the skills required to be a data librarian? data librarianship is still an underdeveloped profession in the united kingdom. the uk example shows the variety of fig.1 data information specialists committee (disc-uk) website iassist quarterly summer 2005 15 backgrounds for its data librarian/manager professionals. jane roberts, university of oxford data library, comes from a social science background with a politics degree, and worked as a social policy researcher, from questionnaire design and data collection through to analysis. jane also provided a support service for data management and spss. tanvi desai, lse rlab data service, comes from a science background with a degree in mathematics with european studies. she worked as a research assistant doing data analysis on the key government social surveys. subsequently she managed a number of primary data collection projects. luis martinez, lse data library, comes also from a science background with a degree in mathematics specialised in statistics and computer science with very little data experience. robin rice is data librarian to the university of edinburgh. she has a masters degree in library and information studies from the university of wisconsinmadison, where she worked as a data librarian through the 1990s. stuart macdonald, also from edinburgh university data library has a biosciences background with a degree in biochemistry in addition to a postgraduate qualification in information studies. he has worked for the last six years in the world of data. “future of data library development will depend critically upon the combined efforts and expertise of three sorts of practitioner. these are the reference librarian, the applied statistician and the software engineer.”18 as peter burnhill suggested in his paper “towards the development of data libraries in the uk”, the skills required to deal with data management can be summarised as an amalgamation of three core areas : • computing; to manage the systems that will store the data and metadata and will provide access to it. to be able to deal with data formatting, merging and customization. to be aware of the latest technologies which will make a significant impact in dissemination and curation of data. • statistics; to interpret the data and help users to take decisions on the appropriate data analysis. to support conducting new surveys, helping with questionnaire design, sampling methods and coding structures. • librarianship; required to catalogue and organise resources. to be able to structure the information in such a way that users can make the most out of what is available. to conduct training session to raise awareness and educate data users. armed with such skills data specialists educate themselves through the day to day exposure to the world of data, specialising according to the needs of their institutions, which in turn have strengths and weakness in the disciplines that they cover. conclusion to conclude, new web and telecommunication technologies coupled to value-added metadata has resulted in advances in data discovery and dissemination. however, as we move towards the use of powerful grid technologies with the next generation of internet, and institutional repositories there may well be a shift towards a more devolved data services infrastructure. the volumes of data will increase exponentially, bringing with it the need for appropriate data curation activities. access to data will be better controlled although possibly more and more restricted. thus it is evident that the 21st century data professional will need to evolve as fast as these technological events take place in addition to being actively involved in the dialogue protecting access to data. footnotes 1paper presented at the iassist conference, edinburgh, may 2005, by stuart macdonald (edinburgh university data library) and luis martinez (london school of economics data library). contact: stuart.macdonald@ed. ac.uk and l.martinez@lse.ac.uk. 2 www.data-archive.ac.uk 3 www.strath.ac.uk/departments/socialstats/ 4 ‘the internet and democratisation of access to data’ robin rice (eudl), blog for social science week on sosig, www.sosig.ac.uk/socsciweek/ 5 www.gro-scotland.gov.uk/ 6 www.scotland.gov.uk/home 7 www.esds.ac.uk 8 http://edina/ 9 www.mimas.ac.uk 10 http://ahds.ac.uk/ 11 “providing local data support for academic data libraries.” robin rice (eudl), the data archive bulletin: 16 iassist quarterly summer 2005 8-11. [available: http://datalib.ed.ac.uk/discuk/docs/ rice2000.pdf] 12 http://datalib.ed.ac.uk/ 13 http://www.nuff.ox.ac.uk/projects/datalibrary/ 14 http://www.lse.ac.uk/library/datlib/ 15 http://rlab.lse.ac.uk/data/ 16 esrc review of international data resources and needs: http://rlab.lse.ac.uk/esrcdata/default.asp 17 http://datalib.ed.ac.uk/discuk/ 18 peter burnhill,1985. towards the development of data libraries in the uk. [online] unpublished paper. university of edinburgh data library. available at : http://datalib. ed.ac.uk/discuk/docs/devdatlibuk.pdf vol21.2 4 iassist quarterly preservation, access, and the multinationals by trudy huskamp peterson* nation states have been the dominant political organizations of the twentieth century. nation states have national archives. these archives have been dominant, too: developing archival theory and practice, supporting archival organizations, and defining what it means to be an archives and an archivist. let us now frame a few research questions that might be posed about the concluding years of this century, the century of the nation state: • did the move to invite additional nations to join nato reflect the notions of identity of populations with each other, with nation state security concerns, or with a desire to lock ever greater portions of the european land mass into one military system? • did the recent tumult when renault announced its plan to shut down its auto plant in belgium and move manufacturing to spain’s cheaper labor market affect ford’s subsequent decision to continue producing in germany, even though german firms themselves were fleeing to central and eastern europe? • was there a congruence or incongruence between the crumbling status of the dayton peace accord in bosnia and the efforts to rebuild the infrastructure of bosnia in general and sarajevo in particular? to answer the first of these questions, the one on nato expansion, a researcher will have to have recourse not only to the records of the nation states, but also to the records of international government organizations, nato in particular, but also the european union, the united nations security council, and the organization for security and cooperation in europe. the second of the questions, that addressing labor movements, popular protest, and industrial activity, would require access to the records of the headquarters of the firms in question, the records of the local subsidiaries (in both the gaining and the losing country), and pan-european manufacturing and labor data from international governmental sources. the third question, rebuilding sarajevo in the face of political disintegration, requires access to the records of international philanthropic organizations and other non-governmental organizations, numbering in the dozens of dozens. are the questions important? absolutely. are records required to answer them? of course. are they being preserved? it is difficult to know. what is the likelihood that our future researcher could gain access to this data? in a word: poor. let me briefly examine three related questions. first, what are the forces that have made the records of the multinationals significant? second, what are the factors that make preservation of records a particularly difficult problem with multinationals? and third, what are the current possibilities to gain access to these records? galloping globalization. each of the three types of multinationals—international government organizations, international business, and international philanthropic and other nongovernmental organizations—is experiencing galloping globalization. each affects the other two, directly or indirectly, but each acts autonomously. the international governmental organizations have had an astonishing growth in the second half of the twentieth century. unexpectedly, nation states willingly shrank their own powers, agreeing to multinational control structures. why? recently a team of researchers consisting of a russian, a german, and a u.s. economist argued that “the most important national interests of these states [united states, russia, japan, and the nations of europe] converge much more than they conflict. the real interests that the parties share greatly outweigh the interests that divide them.”1 in other words, ceding power has actually been in the national interest. be that as it may, the outcome has been the creation of permanent structures, from the european commission to the world bank, with a permanent corps of civil servants, unaccountable to any single state, who create records on the most important worldwide issues of our day. while the fears of the anti-un activists in the united states, who see in the united nations a conspiracy to establish a world government and extinguish the nation state, are clearly fantasy, it is true that the permanent bureaucratic structures summer 1997 5 of the international governmental organizations create the same self-preservation mechanisms that surround any bureaucracy, but importantly absent are the counter-veiling pressures of a citizenry to whom the officials are accountable. the second type of multinational structure is that of the businesses and commercial establishments. while these have been international for some functions for centuries (think of the chinese painting porcelain for the european trade), the late twentieth century difference is in the assembly of goods through multiple nations producing components; in the move from international goods or financial markets into international service providers; and from the speed with which information and currency flows. this borderless market, however, still relies on corporate headquarters somewhere on the globe. these headquarters may be the traditional home of the company, where the corporate officers have their offices, or it could be a single officer in a location that gives the most advantageous tax position for the company. once again, however, these companies can set their work practices and employment standards without much accountability to anyone, other than to the owners. in the united states, the clinton administration has attempted to forge an agreement with a number of the major clothing manufacturers that will cover their operations world wide. the agreement is for a code of conduct, that would prohibit child labor, forced labor, and worker abuse; establishes health-and-safety standards; recognizes the right to join a union; limits working hours to 60 a week “except in extraordinary business circumstances”; and insists that workers be paid at least the legal minimum wage or the prevailing industry wage in every country in which agreements are made.2 the problem, of course, is monitoring such an agreement. it is a major step, but it is voluntary, it is limited to the manufacturers in one country, and to one industry in that country. the industry is said to be setting up a policing mechanism, but efforts to introduce transparency—including the access to records—in international business operations are sisyphean tasks. privatization—a world-wide trend—also plays a part in the issues surrounding the records of international business. as governments divest themselves of a particular function, the records of that function vanish from the public sphere into the private. from banking to manufacture of weapons, the public track stops at the corporate door. and when the privatized entity is purchased by a foreign corporation (such as the lightbulb maker tungsram of hungary purchased by general electric of the united states), then the policy of retention and access move from that of a national government to that of the foreign parent. the situation with the major philanthropic and other nongovernmental international organizations is different from either the governmental or business model. these organizations, ranging from greenpeace to the rockefeller foundation to iassist, are accountable to members or to boards of directors or, even, to heirs of the original donor. the records of the activity may be centralized or dispersed among national chapters; the sources of capital or the number of members may be publicized or a closely held secret; the actions of the board may be publicly reported or may be absolutely secret. what is clear is that these organizations float above or rest lightly within a single nation’s structure. preserving the information base all three types of international organizations depend heavily on electronic information transfers to accomplish their work. but all three, too, have piles of papers, photographs, videotapes, architectural drawings, and a panoply of other records types. who decides what to preserve? and, if data can flow electronically, is there any need to move physically other record types—such as paper or video tape or cartographic items—to an archives? the record of the united nations and its components on preserving their records is spotty, at best. the central united nations archives and records service in new york has no authority to control the records policies of the components. further, with the 1000 person cuts in un headquarters recently announced, any administrative positions are shaky, especially in something as little valued as the records preservation activity. on the other hand, some un units have solid records and archives programs, such as the food and agriculture organization or the un high commissioner for refugees. the temporary un units—such as unprofor—are less likely to have a sufficient records policy to ensure that essential documents are preserved. other international governmental bodies, from nato to the european union to the world bank and the international monetary fund, are known to have serious records programs. it is not clear that temporary bodies, created for a limited purpose, often as a result of crisis, are adequately documented—just as national governments often have trouble adequately documenting the records of short-term committees, commissions, and boards. who, for example, is the official secretariat for the documents of the campaign to end female genital mutilation, recently announced as a joint campaign of the world heath organization, the un’s children’s’ fund, and the un population fund? turning to international business, it is difficult to gain any general picture of the state of preservation of records, given the secrecy that surrounds commercial enterprises. when royal dutch shell, for example, was under attack for continuing to do business in nigeria, were the records of the nigerian unit physically transported to the netherlands? was the information reported from nigeria to the headquarters deemed to be sufficient for corporate purposes and the disposition of the records in nigeria left to chance? or was it most expedient—if not downright prudent—to destroy the records in nigeria as soon as possible? 6 iassist quarterly one specific problem for an in-gathering of corporate records in europe is the policy of the european union to ensure that if a country wants to keep the records created in its territory, it can. in the case of the renault controversy, for example, that would mean that belgium could prohibit the renault subsidiary from sending its records to the french headquarters, as could spain. whether or not this has actually happened, it is possible. in at least one example, france prevented the records of a monastic order from being sent to the headquarters of the order in rome. whether any government would think that corporate records are essential for documenting the history of the nation is not clear, but other countries than the european union—notably russia—give themselves the legal right to prevent the export of records of a business registered with the government. the best one can say is that at least in such a case the documents would be preserved, although scattered. the problems with the international ngos are quite similar to those of the multinational businesses, with the exception that there is less likelihood that they would be destroyed to prevent the release of corporate secrets. the truth is that, for many records of international nongovernmental bodies, whether commercial or philanthropic or pressure groups, there is no logical archival home. if one expects the corporate headquarters to hold records of business world-wide, there would be mass archival storage on the cayman islands or in liechtenstein. in countries where the national archives has a mandate to hold the records of industry or of any organization or establishment within the country, the national archives might be a possible place of deposit. in some countries, however, such as the united states, unless the records of the nongovernmental body show the functioning of the government, the national archives does not have the authority to hold them. even if the problem of location could be solved, the problem of international transport is daunting. electronic files can be transported with relative ease (and if they cannot, the matter of carrying diskettes is simple), but operating in many languages, with electronic data recorded in a wide range of fonts, currently represents a major technical problem. paper and videotapes are another matter entirely. shell advertises that it operates in 120 countries. mcdonald’s operates in so many that purchasing power parity can be based on the cost of a big mac—with reasonably sound economic predictability. it is simply not realistic to believe that these corporations will move any significant quantity of records around the world—it would not be economic, and for these businesses that has to be the bottom line. turning to the non-commercial sector, there the funds are usually heavily committed to pursuing the mission of the organization and precious little is willingly spent on administration, for that is just the means to the goal. unless the information itself has value in pursuing the objective of the non-commercial organization (such as documentation of human rights abuses), it is unlikely to command the resources required for preservation. accessing the record if the record of a multinational activity is preserved, what is the likelihood that the researcher can gain access to it, in any reasonable time? again, the answer is discouraging. the international governmental organizations may have a policy or a procedure, but it is often both arduous and not timely. (one outstanding exception is the historical archives of the european union in florence, italy.) the records of the cases at the world court are closed for 100 years. the international monetary fund recently balked at a request for access for official business by an employee of its sister institution the world bank. ironically, the closed policy and restrictive conditions of some of the international organizations spill over into the national practice; a recent attempt to adopt a freedom of information act in latvia, for instance, failed (according to a knowledgeable observer) because latvia hopes to join nato and the opponents of the act argued that an open access law would run counter to nato practices! turning to international corporations, the record is opaque. corporations like coca cola have an archives and an access policy; so does the walt disney corporation, some multinational banks, and others. but the records policies of most are unknown. ngos also probably present a mixed access picture. here there is almost no data about the actual access conditions. let me, instead, give you an example of the archival challenges of the open society archives as an archives of a major international ngo. the open society archives is the archives for the worldwide network of soros foundations. from their headquarters in new york, foundations operate in nearly forty countries world-wide—from mongolia to south africa to estonia to guatemala. in addition, there are major philanthropic activities in the united states. the archives itself is located in budapest, hungary, which serves as the de facto european headquarters for the open society institute (as the soros foundation is officially known). in addition to the records of the foundations themselves, the archives holds by contract the records of the research institute of radio free europe/radio liberty, in which we find almost every language of europe and central asia; the records of the index on censorship, with world-wide languages; and so on. obviously, for us fonts and languages are major issues. but so, too, is what to preserve in what format. because the foundation network is completely networked and dependent upon electronic communications, we are looking hard at the electronic data in the foundations, to see what can be transferred to the archives electronically and thereby provide a basic portrait of the foundation’s summer 1997 7 activities. we also are capturing the electronic traffic broadcast within the foundation network in an innovative electronic storage program; this should give basic outlines, too. we cannot reasonably ship paper or audio-visual records from all points of the globe, yet some records in non-electronic formats are indispensable for providing a picture of what the foundations are achieving. for example, the romanian foundation has for a number of years provided grants to companies to take public opinion polls, thereby providing an unbiased source for evaluating attitudes toward issues. the polls are taken by different organizations, and the result is published in a continuing series of hard copy publications. the data is invaluable for research, but to preserve it we have to preserve the hard copy report. similarly, the videotape of the roma microlending project by the hungarian foundation exists only in video format. books published by foundation grants, documentary films supported by them, and all manner of sound recordings are also part of the legacy. for us, the mission of the open society archives is to document the foundations and also to preserve information on the period of communism and post-communism in europe and to document the movements for human rights world-wide. but with the exception of a few other foundations, we are probably a unique philanthropic organization that is willing to consider the historical importance of its records. what is to be done? the issue of preservation and access in archives of intergovernmental organizations has been repeatedly discussed during the last decade. in 1990 at the international congress of historical sciences in madrid, charles kecskemeti, the secretary general of the international council on archives, argued for an archival policy in major intergovernmental systems. this was followed by a paper to the 1995 international conference of the round table on archives by liisa fagerlund, in which she called for “harmonized standards and procedures” and for exploring the possibility of depositing the archives in co-operating archival repositories (such as national archives) “in sites with major concentrations of united nations system organizations.”3 in many democratic states, a rock of the social order is the principle that citizens have the right to know what the government is doing and has done—the basis of freedom of information legislation. the international monitor group, freedom house, now estimates that 60% of the governments in the world have a democratic form._ if this is true, then it follows that organizations made up of democratic states should, themselves, have democratic management—in the instant case, a principle of preserving important archives and providing access to them in an open and consistent way. pressure by member states is critical to ensuring that a discussion of preservation and access goes forward. the goal is democratization of the data. unlike the obvious pressure path of citizen to national government to international organization, when we turn to international business we have no such levers. we can assume that business will do whatever it believes good business practice to be, without regard for future research and for history unless it suits the corporate purpose. but we really know very little about the actual situation in the fortune 500 companies, not to mention such emerging giants as gazprom or lukoil. here i believe the next steps are (1) to survey the actual preservation and access practices in the fortune 500 companies, giving special attention to the status of the records of offshore subsidiaries; (2) to launch a concerted effort to encourage the companies to use international standards to describe the historic records they do hold and, to the extent possible under corporate guidance, to share that information electronically. only with some survey results in hand will the research community be able to assess the preservation and access needs. the records of international philanthropy are somewhere in between the two. if human rights activism uses the politics of guilt, it should be possible to use that same argument with key philanthropies to have them preserve and make available their records. after all, these organizations are or hope to be change agents, and it is important to them to have a means to measure their effectiveness. records do that. by suggesting these few activities, i am conscious of the admonition of the book of daniel: “many run to and fro and knowledge shall be increased.”4 i also remember mao zedong’s wrong-headed idea, as reported in the famous little red book, that “investigation may be likened to the long months of pregnancy, and solving a problem to the day of birth. to investigate a problem is, indeed, to solve it.” 5 survey for the purpose of surveying, preservation for preservation’s sake or access for access’ sake is not the goal. the goal is, rather, that we take the steps now to ensure that the twenty-first century embarks on the preservation of its history, a history—i am convinced—that will have as a critical component the actions of the multinationals. 1 graham allison, karl kaiser, and sergei karaganov, “towards a new democratic commonwealth,” december 12, 1996 draft, pp. 1-2 (copy in possession of the author). 2 dress code,” the economist, april 19, 1997, pp. 54-55. 3 liisa fagerlund, “status of records of the united nations system” (copy in possession of the author). 4 daniel 12:4. 5 as quoted in paul theroux, riding the iron rooster: by train through china, new york: ivy books, 1989, p.72. * paper presented at iassist/ifdo ‘97, odense, denmark, may 6-9,1997. vol21.2 20 iassist quarterly abstract hyperlinked eurotrends are a new groundbreaking approach to present codebook data on the world wide web. the author created the so-called eurobarometer trends codebook system (eutrecs). nearly all available material about eurobarometers is presented in a clear and userfriendly way. eutrecs gives fast and easy access to all metadata and datasets. it is the first comprehensive attempt to use internet technology to serve the most basic needs of secondary analysis researchers. eutrecs is based on four principles: 1) hyperlinks connect all interrelated variables; 2) a comprehensive search engine gives access to full text retrieval in all eurobarometer codebooks; 3) an index of over 1,000 keywords and a classification scheme of over 400 trend variables present an easy way to browse through eurobarometer questions; 4) datasets and codebooks are available for immediate download. it will greatly enhance the dataservice of the archives. this paper describes the basic features of this system. scenario: what a researcher could think suppose you are a social scientist and you want to know which questions were asked in the eurobarometer surveys. you have access to the internet and you are looking for a database containing question wording of eurobarometer questionnaires. after having visited all www search engines in vain, a friend of yours gives you a hint. you finally get through and a welcome page invites you: “please type in your query”. but how can you know what you are looking for, if you don’t know what is in those eurobarometer studies. let’s assume, you know already what you are looking for. you are familiar with the topics of the eurobarometer, but you have never worked directly with eurobarometer data. now you want to do some analysis about attitudes towards the common currency. you go to that question database, ask for ”currency” and find the items you are looking for. but you are still without data. if your institution is a member of icpsr, you are lucky and get the desired data quick via internet. but otherwise you will have to wait until your request for data has reached your home archive and has been processed there. however, you need the data just at this very moment. it would be better for you if you knew someone who has these data already. send him an e-mail and you will have your data within short time. let’s think further. you have got the data and have run an analysis. but you are puzzled by your findings. you rerun the analysis, the findings remains the same. after a while you will be wondering if there is anything wrong with that question: ”what was the exact wording of question q34_a in eb 40.1?” just two days ago you had received the printed codebook . but unfortunately it’s in your office and you are on weekend some hundred miles away. so if you are lucky and you know a friend who has specialized in european politics and analysis of eurobarometers, go to the telephone and give him a call. but otherwise? let’s think positively. you have taken the codebook with you and resolved your question wording problem. now you have a new idea: it could be very interesting to compare the attitudes towards the common currency between various professions and diverse levels of occupational prestige. you find a variable concerning ”occupation” but you have to construct a prestige scale for that variable. after a few minutes you remember that there was an article concerning this issue some months ago. but who was the author and in which journal has it been published? it would be great if you could use his work and not having to redo the construction work for a second time. the eutrecs answer the eurobarometer trend codebook system gives a solution for all these problems. those who are not familiar with eurobarometers can browse in a keyword and trend index. they first scroll through subjects or concepts. if they find something interesting, they click on a link to that variable and have question text and marginals right in front them. is that question asked more often than once they simply click on another link to follow that item over time. browsing through codebooks that way, you may find an interesting keyword and you are wondering what else has been asked concerning that topic. so you follow the keyword link of that concept, finding yourself in the keyword index and having a list of questions with similar content in front of you. hyperlinked eurotrends by lorenz gräf* summer 1997 21 those researchers who are familiar with eurobarometer studies find a fully searchable codebook database. every single word contained in a codebook can be found. all the internal linking between similar variables has been preserved. from within eutrecs users can download every dataset except those which are currently under embargo. this service is open only for non-commercial use and academic purpose. you have to register and accept online the terms of use agreement. once registered researchers get an username and a password. after having selected one or more datasets they get emailed a transaction id. within minutes you can have the data. the hyperlinked eurobarometer web pages provides space for user communication. in the moment there is a list of nearly 800 working papers all about eurobarometers. if users contribute to that pages there will be much more content in the next future. eutrecs in detail overview eurobarometer are surveys conducted on behalf of the european commission in the member countries of the ec. they focus on topics concerning the eu. normally they are conducted in spring and fall every year. sample size in each country is about 1,000 persons. data are gathered using face to face interviews. besides this standard series, there are several so-called flash eurobarometers. they are conducted by telephone, on a smaller sample size (about 500) and focus mainly on one topic. since the breakdown of the former soviet imperium every year a survey similar to the standard eurobarometer is carried out in the central and eastern european countries. two years after the fieldwork the data is normally available for academic researchers. those who want the data earlier need a special permission of the european commission. the series of eurobarometer surveys started in the early seventies. up to now, there have been conducted more than 70 standard, 50 flash surveys and 10 central and eastern european surveys. the eurobarometer site built for the central archive (za) is based upon the codebooks created by icpsr, ssd and za itself. so far the codebooks are only produced for standard surveys which do not underlie the embargo restrictions. background information about survey characteristics and topics covered is available for all eurobarometer surveys. datasets of all type of eurobarometer surveys (standard eurobarometer, flash eurobarometer and central and eastern eurobarometer) can be downloaded, but this service is restricted to registered users. eutrecs consist of 32,000 html-pages, comprises 1,000 keywords and over 400 trend variables. it is built on basic usability principles and aims to encompass user needs. hyperlinked structure the eurobarometer codebooks 22 iassist quarterly make use of the core web technology. hyperlinks between variables make it easy for the user to explore survey content and to follow his associations meanwhile browsing through the hyper codebook. what does this mean? suppose you are interested in attitudes towards technology. you look for the item technology. some of the items you find deal also with computer. in eutrecs you can follow the link underneath the item computer and find a list of all variables containing this term. by browsing these variables you discover a variable containing information about the diffusion of computer in diverse european countries. you can see that at the beginning of the nineties in the netherlands and great-britain computers were summer 1997 23 present in a third of all households. at the same time in germany and france only every fifth household possessed a computer. now you are interested to know how these figures have changed over time. in other information systems on the web you have to go back to your query result and click on the next item. so if you want to follow a fairly large trend you are forced to go forward and backward some nasty long time. eutrecs shows on every variable output all other variables which contain the same trend. therefore it is easy to follow the same variable over time by only clicking on the variable name. eutrecs is based on the already available ascii codebooks. when creating html pages eutrecs needs a continuity table to link identical or nearly identical questions. this table is built in two ways. first, variables with identical labels will be identified and inserted in the table. the second way consists in preparing this table outside of eutrecs and using the internal codebook information to check the applicability of the pre-given information. as shown in the next figure eutrecs enlarges the ascii codebooks and adds links. fulltext search engine eutrecs has a built-in search engine based on freewais-sf. all relevant codebook content can be searched using a fully featured search engine. it is the same engine that serves the cessda database. what is unique to eutrecs is that it preserves all the internal links during processing in the search engine. so, if a user puts a query to locate all variables containing the term ‘technology’ on the display of results he finds links to similar concepts, links to other variables of the same trend and a link to the complete online codebook enabling him to explore the context of this variable in the original questionnaire. keyword and trend index a web based information system should present as much information as possible in html files. following this way, researchers can use the web site like a book. in eutrecs browsing in codebooks is possible. 24 iassist quarterly comparing eutrecs to books, the single online codebooks are the chapter of the book. like in books, identical or similar concepts are dealt with in different chapters. to facilitate navigation, eutrecs provides not only a free text search engine but also a keyword and a trend index. the keyword index groups variables with similar subjects together. the trend index combines identical or nearly identical questions in different surveys. keywords are extracted from variable labels. doing so eutrecs can take advantage of a quasi-controlled vocabulary and is not bound to the question wording. but it heavily depends on the quality of the variable label wording. in eurobarometers this labeling should be improved. only codebooks produced in the last years contain really good labels but standardization is already on the way. it would be better if we had real descriptors for each variable. but for the time being we have to be content with the material as it is. from the variable labels of the available codebooks we extracted over 8,000 tokens. as a first step1 we implemented a stop-word list and ended up with about 1.000 different keywords. eutrecs present the keywords in alphabetical order entirely on one screen. you can see in the following figure how easily and fast navigation is. with only three clicks the user gets to the desired variable. a comprehensive list of all keywords is available but it is nearly 2 mb big. a series of thematically similar surveys offers great opportunities for time oriented secondary analysis. to foster such analysis researchers ought to know which question was asked more than once and in which surveys. the eutrecs trend index gives the most comprehensive trend register of the eurobarometer surveys. it contains trend information identified at zeus (mannheim), and at za (köln). but unlike zeus we defined trends strictly on a single variable level. in our definition a trend is every single variable that can be traced over time. it is this data that can be compared between different points in time. now what is the difference? at zeus trends are identified, named and differentiated on a concept level. in eutrecs only questions with identical wording or identical subject have got the same name. let’s give an example. zeus presents each item of the question ”which of the following aereas of policy do you think should be decided by the (national) summer 1997 25 government, and which should be decided jointly within the european community?”2 under the label ‘common policy aereas’. whereas in eutrecs we name the concept ‘common policy aereas’ but give the trends the name of the stimulus object, i.e. ‘foreign policy’, ‘currency’ or ‘education’. by doing so we can provide links between all variables which have the same question wording. clicking on these links will guide the user through all instances of the trend. as eutrecs is built out of machine readable codebooks together with each question text full marginal information is displayed. in this version we were not strict on question wording and did not differentiate questions with minor deviations in question wordings. we leave it up to the researcher to decide whether questions were identical or not. until now we have identified over 400 single trend variables. to facilitate finding topics we classified the trends in a new classification scheme. in our opinion it was the best method to group similar topics together and avoid getting lost in a mere alphabetical ordering. this classification scheme includes three levels. the third level describes the concept level and is comparable to zeus trend names. beneath that level we have classified the variable trends which we could identify. so we can distinguish between the concept and a measurement level. in the following figure you can see how trends are presented in eutrecs. download within eutrecs the direct download of spss datasets and of codebook material is possible. bandwidth on the internet is normally small. therefore we divided the big codebooks in smaller parts of about ten or fifteen variables. but for printing purposes many users want to have entire codebooks. for offline use we give access to all codebooks in htmland postscript format. keyword and trend indices can be accessed in one-file lists. access to datasets is given via an entirely online procedure. we tried to find the easiest and fastest way to deliver data over the internet without having to compromise to the vital interests of the archives. so we ended up with a three step procedure. in the first step researchers register at the za in cologne. all they need is a functional e-mail address. after registration they get an account and a password via e-mail. provided with username and password in step 2 they can select datasets. after having accepted the terms of use agreement and having paid a fee (depending on the usage conditions of the archive) they get a transaction id by e-mail. with this id they go directly to the download page and get their material immediately. user forum the web site we intend to create should be as interactive as possible. it should be the most efficient site for all researchers 26 iassist quarterly doing secondary research with eurobarometer surveys. we pursue three aims: first, this site should be a platform for the communication between users and the archive and between users themselves. it would be nice, if the archives use this forum to announce any news concerning eurobarometers (bug reports, announcement of availability and declarations about archive policy). second, we try to stimulate a shared knowledge forum. in this forum every researcher should announce his findings in the eurobarometer data. the site now contains a list of working papers done with eurobarometer material. we would appreciate if researchers uploaded their newest paper for the communication with the scientific community. and we encourage researchers to share their operationalisations and measurement attempts of the data. it would be a really common good to have samples of spss or sas statements there for often needed recoding of eurobarometer variables. third, we will give eutrecs users the necessary software tools to use materials of this site, like ghostview to read ps-files. what comes next eutrecs is going to be developed in a multistage process. in stage one it was important to explore possibilities to present large and complex survey metadata on the web. in the next step we need the help and feed-back of the users to add consistency to the pages. we have to find and correct miscoding, misleading labeling and false classification. it has to be tested how usable the site is and how navigation could be improved. also the gap in the codebook production should be closed as soon as possible. in the last step we want to incorporate in the data itself trend information and other measurement suggestions of the research community. 1 in future versions a better suited dictionary will be used to improve building of keywords. 2 exact question wording (eb33 q30): ”some people believe that certain areas of policy should be decided by the (national) government, while other areas of policy should be decided jointly within the european community. which of the following areas of policy do you think should be decided by the (national) government, and which should be decided jointly within the european community?” • foreign policy towards countries outside the european community • education www: http://infohttpsoc.uni-koeln.de/graef http://solix.wiso.uni-koeln.de/~graef/ * paper presented at iassist/ifdo ‘97, odense, denmark, may 6-9,1997. lorenz gräf, university of cologne, e-mail: lorenz.graef@uni-koeln.de http://infohttpsoc.uni-koeln.de/graef http://solix.wiso.uni-koeln.de/~graef/ mailto:lorenz.graef@uni-koeln.de 1/9 cain, jonathan; cooper, liz; demott, sarah; montgomery, alesia (2019) where is qda hiding? an analysis of the discoverability of qualitative research support on academic library websites, iassist quarterly 43(2), pp. 1-9. doi: https://doi.org/10.29173/iq957 where is qda hiding? an analysis of the discoverability of qualitative research support on academic library websites jonathan cain1, liz cooper2, sarah demott3, alesia montgomery4 abstract this study explores the discoverability of qualitative research support services, using a purposive sample of academic library websites (n=95). these services were hard to find on most of the websites in our sample. in this paper, we outline the site characteristics that make discoverability easy or hard. previous studies on qualitative resources at academic libraries have not addressed this topic. our study fills this gap in the literature. our aim is to provide information that can help libraries to improve the visibility of their resources for qualitative researchers and their students. keywords qualitative research support, discoverability, caqdas, research guides, academic libraries introduction in the world of data librarianship, quantitative data remains sovereign. whenever there is a data science initiative on a u.s. campus, it is the quantitative courses, procedures, and tools that administrators offer as illustrations of data science and data research. when administrators allocate funds for data analysis tools, often, qualitative tools are left by the wayside. however, as data science matures, researchers increasingly strive for holistic design, participatory research, and mixed methods—approaches that include qualitative techniques. thus, qualitative data analysis (qda) has become more prominent in research projects, and qualitative digital tools, methods, and sources have grown in importance for scholars in the humanities, sciences, social sciences, and the professions (engineering, law, medicine). as researchers acknowledge the value of qualitative data, we must ask how libraries are meeting the demand for qualitative research support services. to address this concern, our six-member team— composed of academic librarians, research specialists, social scientists, and their overlap—began a project to evaluate the availability of these services. initially, we intended to compile a web directory of links to library qualitative services, with the aim of informing our colleagues in the iassist qualitative social science & humanities data interest group (qsshdig). iassist is “an international organization of professionals working in and with information technology and data services to support research and teaching.”5 however, we soon found that it was hard to find qualitative services on library websites, even when we knew that particular libraries did offer these services. we asked ourselves: “where is qda hiding?!” four members of our original team shifted our focus to conduct an exploratory study of the discoverability of qualitative services on academic library websites. literature review within the social sciences, there are various definitions of qualitative research and data. in this paper, we define “qualitative research” as the naturalistic, interpretive study of social meanings and processes that uses techniques such as in-depth interviews, observations, and textual analysis. we describe “qualitative data” as the work material of researchers who use qualitative https://doi.org/10.29173/iq957 2/9 cain, jonathan; cooper, liz; demott, sarah; montgomery, alesia (2019) where is qda hiding? an analysis of the discoverability of qualitative research support on academic library websites, iassist quarterly 43(2), pp. 1-9. doi: https://doi.org/10.29173/iq957 methods. qualitative data is often called “unstructured” data—for example, the interview transcripts, field notes, and neighborhood photos of qualitative researchers—in contrast to the tabular, numerical, “structured” data of quantitative researchers. since the late twentieth century, qualitative researchers have increasingly digitized and archived their data for secondary analysis, and social media (e.g., twitter, facebook) have become novel sources of data. in this era of “big data,” both quantitative and qualitative researchers sometimes analyze massive assemblages of unstructured data—the kind of material that was once solely the domain of qualitative researchers—and they sometimes engage in mixed methods research (boyd and crawford, 2012; davidson et al., 2018; delyser and sui, 2013; mcfarland et al., 2016). thus, there is an overlap in the data, data analysis tools, and data management and archiving support services that qualitative and quantitative researchers need. yet even when qualitative and quantitative researchers use the same data, tools, and archives, they still often maintain distinct epistemological assumptions, research terminology, analytical approaches, and data reuse concerns. given the dominance of quantitative research in the social sciences, the challenge for social science librarians is to make their services available and discoverable in ways that respect the unique needs, language, and concerns of qualitative researchers. most literature on data services focuses on the needs of quantitative researchers (e.g., geraci et al., 2012). surveys of data services at academic libraries sometimes contain an item about qualitative data support (tenopir et al., 2017), but few studies focus on qualitative data services (corti, 2000; mannheimer et al., 2018). an excellent example of this latter work is mandy swygart-hobaugh’s (2016) study of the qualitative data services offered by academic libraries. swygart-hobaugh outlines the types of services that could be provided to qualitative researchers across the data life cycle: ● during the first stage—discovery and planning—librarians could help qualitative researchers find information about archived materials for secondary analysis and they could inform researchers about project documentation procedures. ● during the intermediate stages—data collection, processing, and analysis—librarians could provide services such as helping researchers learn to use qualitative data analysis software. ● during the final stages—research dissemination and long-term management—librarians could help qualitative researchers to archive and share their data. after presenting this outline, swygart-hobaugh describes the types of services that academic libraries do provide, based on her analyses of (1) academic librarian job postings in the iassist repository (n=148), (2) an online survey of academic librarians (n=112), and (3) qualitative research guides on academic websites (n=53). she found that the academic libraries in her sample emphasize quantitative data support much more than qualitative data support. most iassist job postings (83 out of 148) mention quantitative expectations (for example, statistical skills) while few (22 out of 148) mention qualitative expectations (for example, knowledge of qualitative analysis software). few social science librarians report providing services specifically targeted to qualitative researchers, and few academic library websites link to qualitative research resources. based on her findings, swygart-hobaugh urges the expansion of qualitative data support services, especially at universities in which a significant number of faculty and graduate students do qualitative research. https://doi.org/10.29173/iq957 3/9 cain, jonathan; cooper, liz; demott, sarah; montgomery, alesia (2019) where is qda hiding? an analysis of the discoverability of qualitative research support on academic library websites, iassist quarterly 43(2), pp. 1-9. doi: https://doi.org/10.29173/iq957 as academic libraries respond to researcher demand by expanding their research support services (kennan et al., 2014), they must carefully plan ways to make their qualitative research services visible. our literature review uncovered no studies that focus on the discoverability of qualitative research services on academic library websites. there are usability studies of academic library websites (comeaux, 2017; ridha, 2018), and there are recommendations for making data discoverable (macmillan, 2014), but there does not appear to be research that specifically focuses on the visibility of qualitative research services on library websites. our study fills this gap in the literature. if qualitative researchers cannot find these services on library websites—or if the language framing these services seems tailored to quantitative researchers--then qualitative researchers may not take advantage of these services, even when they are available. discoverability is also important in supporting a community of practice among those of us who provide services to qualitative researchers: being able to easily find the qualitative resources offered by our colleagues at other institutions may help us to forge ties with and learn from each other. methods for our exploratory study, we selected a diverse sample of u.s. academic libraries (n=95) at public and private institutions that vary in size and resources (see appendix). we used the list of american research libraries (arl) to help us identify institutions.6 however, we did not include all of the arl institutions, and a few institutions in our sample are not arl members. after we chose the institutions for the study, we assigned a set of roughly 20 institutions to each member of our research team. none of us reviewed her or his home institution. we focused our assessment on the degree to which it was easy to find the range of qualitative services that faculty and students might use across the data life cycle on the library websites. as part of this evaluation, we examined qualities such as the location of qualitative research resources and the types of data support services offered (for example, software tutorials and workshops, data management tools). discussion it is hard to find qualitative services on the websites of most libraries in our sample. perhaps many of these libraries do not provide these services, which would be in line with the findings of swygarthobaugh’s (2016) study. yet, in some cases, we eventually found mention of qualitative services offered by the libraries after trial-and-error searches using various search terms. we often found these services by entering the brand names of popular qualitative software, such as nvivo, atlas.ti, dedoose, and maxqda. approximately half of our sample references at least one qualitative software brand. libraries mention, for example, that a brand is available on campus computers and/or that staff are available to provide software training. it is problematic that, for some libraries, the only way to find out about their qualitative services is to enter the name of a proprietary software. their patrons may not know these brand names, and they may not even know much about qualitative software. mention of these brands is sometimes buried on pages that list quantitative software and do not include the word “qualitative.” it is also problematic that, for some libraries, data services that qualitative researchers might need—for instance, help with archiving their research or learning a programming language such as python that could be used for web scraping—can be found by entering the term “quantitative” on the library’s homepage but not “qualitative.” it was much easier to find mention of support across the data lifecycle for quantitative research. when support for qualitative software was mentioned on a library website, it was rarely linked to data management webpages, and these webpages tended not to https://doi.org/10.29173/iq957 4/9 cain, jonathan; cooper, liz; demott, sarah; montgomery, alesia (2019) where is qda hiding? an analysis of the discoverability of qualitative research support on academic library websites, iassist quarterly 43(2), pp. 1-9. doi: https://doi.org/10.29173/iq957 address the unique needs of qualitative researchers. in some cases, information about qualitative services was isolated from other data services pages on the library website—for example, the information might reside on a libguide for a course or discipline. in other cases, we found qualitative resources scattered across institutional silos in and beyond the library, such as campus units that support transcription tools or audio and video recording. it seems likely that these qualitative resources have been developed independently across units, with little collaboration or communication, and thus they have not been showcased as a set. these silos include data services, tech support, digital scholarship, digital humanities, and gis services. in contrast, we found that library websites that make it easy to find qualitative services tend to share the following four characteristics: (1) data services are prominently featured on the library home page. (2) links to data services are integrated into the top level of the library website’s menus. o data services webpages are accessible within 1-2 clicks. o users do not have to search through the site. (3) labels for data services are intuitive and include the terminology of qualitative researchers—novice users of qualitative services do not have to guess terms (e.g., software brand names) for finding resources. (4) there is one integrated website for all data services support. o qualitative and quantitative data support are both explicitly mentioned and supported. o services across the data lifecycle are outlined in one place (e.g., data discovery, data analysis including software workshops, data management). o any related units or services that provide support are linked back to this main page (including any research guides). o the structure of the data services website is based on user needs, not the organizational structure of data services. below we illustrate the qualities of websites that make it easy to find data services on their home pages and that have integrated data services websites that mention qualitative services (see figures 1 and 2). https://doi.org/10.29173/iq957 5/9 cain, jonathan; cooper, liz; demott, sarah; montgomery, alesia (2019) where is qda hiding? an analysis of the discoverability of qualitative research support on academic library websites, iassist quarterly 43(2), pp. 1-9. doi: https://doi.org/10.29173/iq957 figure 1: example library home page: one click to finding qualitative resources figure 2: example integrated data services page it is important to note that qualitative data services are rapidly evolving and that some institutions already may be taking steps to improve the visibility of these services and to coordinate their delivery. over the course of our study, as we revisited websites, we noticed that some libraries with hard-to-find qualitative resources have greatly improved their discoverability: for example, one library has made its qualitative services discoverable by bringing them out from behind a firewall. from the library’s homepage, entering the term “qualitative” now returns a treasure trove of resources, including links to qualitative software, a software guide, staff, a user group, and a relevant database. conclusion currently, finding information about qualitative research support on library websites is difficult. although on many campuses and in many disciplines support for qualitative research is secondary to https://doi.org/10.29173/iq957 6/9 cain, jonathan; cooper, liz; demott, sarah; montgomery, alesia (2019) where is qda hiding? an analysis of the discoverability of qualitative research support on academic library websites, iassist quarterly 43(2), pp. 1-9. doi: https://doi.org/10.29173/iq957 support for quantitative research, as more researchers begin to work with unstructured data and others strive for more holistic design, the demand for qualitative research continues to increase. libraries wanting to provide robust support for data services would do well to consider not only swygart-hobaugh’s 2016 recommendations on types of qualitative support libraries might offer across the data lifecycle, but also to dedicate time and resources to ensuring the discoverability of these services on their websites. references al-qallaf, c.l. & ridha, a. (2018), “a comprehensive analysis of academic library websites: design, navigation, content, services, and web 2.0 tools”, international information & library review, pp. 1–14, available from https://doi.org/10.1080/10572317.2018.1467166 (accessed 29 january 2019). boyd, d. & crawford, k. (2012), “critical questions for big data”, information, communication & society, vol. 15, no. 5, pp. 662–679, available from https://doi.org/10.1080/1369118x.2012.678878 (accessed 12 february 2019). comeaux, d.j. (2017), “web design trends in academic libraries—a longitudinal study”, journal of web librarianship, vol. 11, no. 1, pp. 1–15, available from https://doi.org/10.1080/19322909.2016.1230031 (accessed 29 january 2019). corti, l. (2000), “progress and problems of preserving and providing access to qualitative data for social research—the international picture of an emerging culture”, forum qualitative sozialforschung / forum: qualitative social research, vol. 1, no. 3, available from http://www.qualitative-research.net/index.php/fqs/article/view/1019 (accessed 12 february 2019). davidson, e., edwards, r., jamieson, l. & weller, s. (2018), “big data, qualitative style: a breadthand-depth method for working with large amounts of secondary qualitative data”, quality & quantity, available from http://link.springer.com/10.1007/s11135-018-0757-y (accessed 3 october 2018). delyser, d. & sui, d. (2013), “crossing the qualitative-quantitative divide ii: inventive approaches to big data, mobile methods, and rhythmanalysis”, progress in human geography, vol. 37, no. 2, pp. 293–305, available from http://journals.sagepub.com/doi/10.1177/0309132512444063 (accessed 1 february 2019). geraci, d., humphrey, c. & jacobs, j. (2012), data basics: an introductory text, available from http://3stages.org/class/2012/pdf/data_basics_2012.pdf. kennan, m.a., corrall, s. & afzal, w. (2014), “’making space’ in practice and education: research support services in academic libraries”, library management, vol. 35, no. 8/9, pp. 666–683, available from http://www.emeraldinsight.com/doi/10.1108/lm-03-2014-0037 (accessed 31 january 2019). https://doi.org/10.29173/iq957 https://doi.org/10.1080/10572317.2018.1467166 https://doi.org/10.1080/1369118x.2012.678878 https://doi.org/10.1080/19322909.2016.1230031 http://www.qualitative-research.net/index.php/fqs/article/view/1019 http://link.springer.com/10.1007/s11135-018-0757-y http://journals.sagepub.com/doi/10.1177/0309132512444063 http://3stages.org/class/2012/pdf/data_basics_2012.pdf http://www.emeraldinsight.com/doi/10.1108/lm-03-2014-0037 7/9 cain, jonathan; cooper, liz; demott, sarah; montgomery, alesia (2019) where is qda hiding? an analysis of the discoverability of qualitative research support on academic library websites, iassist quarterly 43(2), pp. 1-9. doi: https://doi.org/10.29173/iq957 macmillan, d. (2014), “data sharing and discovery: what librarians need to know”, the journal of academic librarianship, vol. 40, no. 5, pp. 541–549, available from http://www.sciencedirect.com/science/article/pii/s0099133314000950 (accessed 11 february 2019). mannheimer, s., pienta, a., kirilova, d., elman, c. & wutich, a. (2018), “qualitative data sharing: data repositories and academic libraries as key partners in addressing challenges”, american behavioral scientist, available from https://par.nsf.gov/biblio/10062050-qualitative-data-sharingdata-repositories-academic-libraries-key-partners-addressing-challenges (accessed 12 february 2019). mcfarland, d.a., lewis, k. & goldberg, a. (2016), “sociology in the era of big data: the ascent of forensic social science”, the american sociologist, vol. 47, no. 1, pp. 12–35, available from https://doi.org/10.1007/s12108-015-9291-8 (accessed 1 february 2019). swygart-hobaugh, a. (2016), “qualitative research and data support: the jan brady of social sciences data services?”, in l. m. kellam & thompson, kristi (eds.). databrarianship: the academic data librarian in theory and practice, pp. 153–178, american library association, chicago, il, available from https://scholarworks.gsu.edu/univ_lib_facpub/128. tenopir, c., talja, s., horstmann, w., late, e., hughes, d., pollock, d., schmidt, b., baird, l., sandusky, r.j. & allard, s. (2017), “research data services in european academic research libraries”, liber quarterly, vol. 27, no. 1, pp. 23–44, available from https://www.liberquarterly.eu/article/10.18352/lq.10180 (accessed 1 february 2019). appendix: list of libraries 1. university at buffalo, suny, libraries 2. university of calgary libraries and cultural resources 3. university of california, berkeley library 4. university of california, davis library 5. university of california, irvine libraries 6. ucla library 7. university of california, riverside library 8. uc san diego library 9. university of california, santa barbara libraries 10. case western reserve university libraries 11. the university of chicago library 12. university of cincinnati libraries 13. university of colorado boulder libraries 14. colorado state university libraries 15. columbia university libraries 16. university of connecticut libraries 17. cornell university library 18. dartmouth college library 19. university of delaware library 20. duke university libraries 21. emory university libraries https://doi.org/10.29173/iq957 http://www.sciencedirect.com/science/article/pii/s0099133314000950 https://par.nsf.gov/biblio/10062050-qualitative-data-sharing-data-repositories-academic-libraries-key-partners-addressing-challenges https://par.nsf.gov/biblio/10062050-qualitative-data-sharing-data-repositories-academic-libraries-key-partners-addressing-challenges https://doi.org/10.1007/s12108-015-9291-8 https://scholarworks.gsu.edu/univ_lib_facpub/128 https://www.liberquarterly.eu/article/10.18352/lq.10180 8/9 cain, jonathan; cooper, liz; demott, sarah; montgomery, alesia (2019) where is qda hiding? an analysis of the discoverability of qualitative research support on academic library websites, iassist quarterly 43(2), pp. 1-9. doi: https://doi.org/10.29173/iq957 22. university of florida libraries 23. florida state university libraries 24. the george washington university library 25. georgetown university library 26. university of georgia libraries 27. georgia institute of technology library 28. georgia state university 29. university of guelph library 30. harvard library 31. university of hawai‘i at mānoa library 32. university of houston libraries 33. howard university libraries 34. university of illinois at chicago library 35. university of illinois at urbana-champaign library 36. indiana university libraries bloomington 37. the university of iowa libraries 38. iowa state university library 39. johns hopkins university libraries 40. university of kansas libraries 41. kent state university libraries 42. university of kentucky libraries 43. university of louisville libraries 44. mcgill university library 45. mcmaster university libraries 46. university of minnesota libraries 47. new york public library 48. new york university libraries 49. university of north carolina at chapel hill libraries 50. north carolina state university libraries 51. northwestern university library 52. university of notre dame libraries 53. the ohio state university libraries 54. ohio university libraries 55. university of oklahoma libraries 56. oklahoma state university library 57. university of oregon libraries 58. university of ottawa library 59. university of pennsylvania libraries 60. pennsylvania state university libraries 61. university of pittsburgh libraries 62. princeton university library 63. purdue university libraries 64. queen's university library 65. rice university library 66. university of rochester libraries 67. rutgers university libraries 68. university of saskatchewan library 69. simon fraser university library 70. smithsonian libraries 71. university of south carolina libraries 72. university of southern california libraries https://doi.org/10.29173/iq957 9/9 cain, jonathan; cooper, liz; demott, sarah; montgomery, alesia (2019) where is qda hiding? an analysis of the discoverability of qualitative research support on academic library websites, iassist quarterly 43(2), pp. 1-9. doi: https://doi.org/10.29173/iq957 73. stanford university libraries 74. stony brook university, suny, libraries 75. syracuse university libraries 76. temple university libraries 77. university of tennessee, knoxville, libraries 78. university of texas libraries 79. texas a&m university libraries 80. texas tech university libraries 81. university of toronto libraries 82. tulane university library 83. university of utah library 84. vanderbilt university library 85. university of virginia library 86. virginia tech libraries 87. university of washington libraries 88. washington state university libraries 89. washington university in st. louis libraries 90. university of waterloo library 91. wayne state university libraries 92. western university libraries 93. university of wisconsin–madison libraries 94. yale university library 95. york university libraries endnotes 1 jonathan cain, head, data services & librarian for planning public policy and management university of oregon, jocain@uoregon.edu 2 liz cooper, social sciences librarian, university of new mexico, cooperliz@unm.edu 3 sarah demott, research specialist, data services, new york university, sarah.demott@nyu.edu 4 alesia montgomery, subject specialist for sociology, psychology, & qualitative data stanford university, alesiam@stanford.edu 5 a description of iassist can be found at https://iassistdata.org 6 association of research libraries member list, https://www.arl.org/membership/list-of-arl-members, accessed february 10, 2019 https://doi.org/10.29173/iq957 mailto:jocain@uoregon.edu mailto:cooperliz@unm.edu mailto:sarah.demott@nyu.edu mailto:alesiam@stanford.edu https://iassistdata.org/ https://www.arl.org/membership/list-of-arl-members vol30-1.indd iassist quarterly spring 2006 by by jeanette bohr, andrea janßen & joachim wackerow * problems of comparability in the german microcensus over time and the new ddi version 3.0 introduction the data documentation initiative (ddi) was created as an international documentation specification to improve the access to and the analysis of social science data. the ddi can be seen as a reaction to a growing need for data documentation standards, brought about by the increased diffusion of quantitative data in the social sciences, because “[...] accurate use of data depends on access to comprehensive, accurate documentation” (blank/rasmussen 2004, 310). initiated by the inter-university consortium for political and social research (icpsr), the ongoing development of the standard is carried out by the membership-based ddi alliance. the alliance is developing the ddi specification, which is written in xml.1 at present, a new version of ddi, whose new possibilities have been discussed at the 2006 iassist conference2, is under review. as a contribution to this discussion, the paper will point out an example of use for the new ddi version 3.0, with emphasis on the new grouping model. the german microdata lab at the centre for survey research and methodology in mannheim prepares documentation for official data for the scientific community. the documentation of the german microcensus refers to the ddi standard version 2.1. as an annually-repeated survey, the german microcensus contains some changes over time that must be documented. however, one of the biggest limitations of the second version of ddi is the lack of the capacity for documenting repeated surveys. the ddi 2.1 documentation pertains only to a single survey, without an option for recording changes over time in repeated surveys. the new ddi version 3.0 and the new grouping structure enables the documentation of comparability of variables over time for repeated studies. this paper explains this structure and gives an example for using the german microcensus for the application of this new grouping model. ddi version 3.0 the major change of ddi version 3.0 from preceding versions is the increased scope of metadata which can be captured. with ddi 3.0, it is now possible to describe all aspects of the data life cycle (see thomas 2006a, nelson 2006). the consideration of the entire data life cycle necessitates a means for describing groups of studies as well as the relationships within collections of comparable studies. the new grouping functionality can be used to define a set of comparable studies which can be described in one single instance. the possibility of documenting information about the comparability of several studies represents a major improvement over previous versions of ddi. because the new version is still under review, the new possibilities of documentation with ddi 3.0 have yet to be fully explored. this paper will serve as a contribution to that exploration. the example: the german microcensus the german microcensus is a representative annual population sample containing structural population data of one percent of all households in germany. from every survey, a sub-sample of seventy percent is drawn, which is made available to the scientific community as scientific use files. because of the annual repetition, the broad scope of topics and the large number of interviewees, the scientific use files are suitable for the analysis of social structure and the observation of social change in society. to observe social change with the microcensus, however, it is necessary to have information about the comparability of variables among census years. the contents of the microcensus and the questionnaires are regulated through a law called the “mikrozensusgesetz.” changes over time in the mikrozensusgesetz have complicated the comparability of variables over time. between 1995 and 1996 there was an especially significant change concerning the questionnaire, leading to a variety of changes on different levels. many of the variables of the 1995 census were split into two variables in subsequent years. in addition to this, there were changes in the question routes of the questionnaires, and in many cases the question texts were changed as well. further inconsistencies concern the values of the variables and the value labels. all the information about inconsistency among census years must be adequately documented and administered. for this purpose, the new ddi version 3.0 offers new possibilities. ddi 3.0: grouping model within the ddi version 3.0, variable inconsistency can 14 iassist quarterly spring 2006 be described within the grouping model. it offers the possibility to define any information as a standard on a top level and to capture variations or additions on a lower level. consequently, the application of the new model permits the documentation of coherences and variations among different census years. the main improvement the new version offers is the facility for the inheritance of the common characteristics of studies down the hierarchy tree of the metadata. this means a simplification of ddi instances because the common metadata has to be stated only once at the upper level of the grouping structure. the grouping structure consists of several hierarchical levels: · the group as the top level contains common metadata which is inherited down the hierarchy of the grouping structure. · subgroups can be created on one or more lower levels. in the following example the subgroups contain the information for a period of predominantly consistent census years. · the study unit represents a single study on which all the lower-level modules depend. groups and study units both contain a cluster of modules which describe a study. these modules are: the “concept” which contains mainly the study description; the “datacollection” which includes information about the methodology and the instrument; the “logicalproduct” with most of the material in the data description, and finally “physicaldataproduct” and “physicaldatainstance” which capture the file description (see thomas 2006a). the concept of the inheritance means that classes of specific information always and at any level inherit from their ancestor classes. the specified metadata on the top position of the hierarchy is valid for all studies in this group, unless it is overridden at a lower level. if a piece of information is not valid for a member of the group, the mechanism of local overrides allows for an appropriate replacement of the information on a lower level. the german microcensus in the grouping model structure3 the organization of the census years within the grouping structure requires a precise knowledge of the data. in particular, the definition of subgroups presupposes a thorough familiarity with all consistencies and changes over time. for the following example, the census years from 1989 to 2004 were taken into consideration. on the top level, the german microcensus is defined as the group. this class includes the common metadata, which is shared by all census years such as the general part of the study description or the basic conception. on the group level, the group parameters have to be defined. this is a set of required properties of the survey, which microcensus subgroup 1: 1989-1995 subgroup 2: 1996-2004 group: subgroup: studyunit: 1989 1995… 1996 2004… group parameters figure 1: hierarchical structure of the german microcensus in the grouping model iassist quarterly spring 2006 15 determines the relationship between the study units of a group (see thomas 2006b). the group parameters of the german microcensus are marked in the following table: because of changing legal regulations and a significant break in the survey program, two different periods can be defined: the first period including the years from 1989 to 1995 and the second including the years from 1996 to 2004. an extensive variable consistency exists within these periods. consequently, it proves to be useful to define these two periods as subgroups on a lower level, where most of the variable description is included. as an alternative, one could define the census years from 1996 on as a standard in the top level and the period before as a subgroup. the disadvantage of this approach is, that the structure is less flexible for accommodating any future significant changes. finally, the study units represent the single census years of the german microcensus. example 1: variable inconsistency between subgroups the first example illustrates the possibilities of documenting variable inconsistency between two subgroups. in both periods a variable exists with information about the vocational training certificate for the person who is head of the household. apart from the different variable names and labels, the contents of the variables differ as well: the variable in subgroup 1 contains the latest certificate, while the variable in subgroup 2 contains the highest: because of the different meaning of the two variables, we define their variable-specific properties within the specific subgroup. as a result, the information is shared by all members of the subgroup. the relevant metadata for each subgroup in this case are the variable name and label, the category information and the associated question text. in ddi 3.0, information about the variable and the categories of the variable both have to be defined in the module “logicalproduct.” the variable attributes “name” and “label” are described in the container table 1: group parameters parameter tag description time t0 no formal relationship t1 single occurrence t2 multiple occurrence: regular occurrence: continuing t3 multiple occurrence: regular occurrence: limited time t4 multiple occurrence: irregular occurrence: continuing t5 multiple occurrence: irregular occurrence: limited time instrument i0 no formal relationship i1 single i2 multiple: integrated set of 2 or more instruments used for different subgroups i3 multiple: base with topical changes panel p0 no formal relationship p1 single panel surveyed multiple times p2 single panel surveyed once p3 rolling panel (multiple interviews limited duration) p4 different panel each survey geography g0 no formal relationship g1 single geography surveyed multiple times g2 single geography surveyed once g3 rolling geography (multiple interviews limited duration) g4 different geography each survey data sets d0 no formal relationship d1 single data file from a data collection d2 multiple data products from a single data collection d3 integration of multiple data sets into a single integrated structure d4 multiple data files each from a different data collection subgroup 1: 1989-1995 “latest vocational certificate: head of household” (ef198) subgroup 2: 1996-2004 “highest vocational certificate: head of household” (ef568) 16 iassist quarterly spring 2006 “variablescheme.” the variable categories are described in the container “categoryscheme” which is illustrated in simplified form in figure 3. hence, in both elements the adequate attributes can be described for each subgroup. variables which are identical for all census years (e.g. age or gender) are described on the top-level in the group and are valid for all study units. the question text for the variables is defined in the module “datacollection” on the subgroup level. it is described in group subgroup1 subgroup2 logicalproduct logicalproduct variablescheme categoryscheme variablescheme categoryscheme variable: name: ef198 label: “latest…” variable: name: ef568 label: “highest…” category: label definition logicalproduct i.e. age, gender… category: label definition figure 2: ddi structure example 1: logicalproduct subgroup 2: 1996-2004 study unit: 2003 “type of attended school” (ef72) 1 class 1 to 4 2 class 5 to 10 3 class 11 to 13 (sixth form) 4 vocational school 5 university of applied sciences 6 university 9 non response not applicable “type of attended school” (ef74) 1 class 1 to 4 2 class 5 to 10 3 class 11 to 13 (sixth form) 4 vocational school 5 vocational preparatory school 6 vocational school with middle graduation 7 vocational school with higher graduation 8 technical college, university of cooperative education 9 graduation of college and advanced administrative studies 10 university of applied sciences 11 university 12 phd program 99 non response not applicable iassist quarterly spring 2006 17 the element “questiontext” of the container “instrument.” besides the question text, there is specific information about the questionnaire such as the question number which changes from year to year. this annually-changing information should be defined in the single study units on a lower level. this example illustrates how important the knowledge of the consistencies and changes among the census years is for the correct definition of the information in the hierarchical structure of ddi. example 2: variable inconsistency within subgroups the next example deals with a change on the category level within one subgroup. in subgroup 2 (1996 to 2004) the variable “type of attended school” (ef72) contains seven categories; in 2003 the variable differentiates among more categories regarding vocational school and university. in addition, the variable name has changed to ef74. because of this variation within the subgroup, the shared information about the variable categories on the subgroup level has to be overridden for the specific year. this can be realized by defining the information about the categories at the study unit level for the census year 2003. furthermore, the variable name and the response categories of the question should be defined at the study unit level, because these elements are different from the description in the subgroup, too. group subgroup1 subgroup2 datacollection datacollection instrument questionitem: questiontext instrument questionitem: questiontext datacollection figure 3: ddi structure example 1: datacollection studyunit datacollection … for the description of the categories and the variable name, we need the module “logicalproduct” again. but now we have to use the module on the study unit level. however, the variable label does not need to be specified in the study unit because this information is inherited from the subgroup. as in the previous example, the question text will be stated in the module “datacollection.” if the same variation of the variable exists over several years, for instance in 2003 and the following years, the definition of a new subgroup for this information below subgroup 2 could be useful. furthermore, ddi gives the opportunity to document the information about the comparability of non-identical variables. in case of the given example, a new variable can be created by combining the categories of the variable in 2003, which is comparable to the standard variable in the subgroup. information about the new variable as well as the needed recode job can be stated in the element “derivation” of the container “variablescheme.” the possibility of documenting aspects of comparability marks an important advantage in terms of the documentation of the german microcensus. due to frequent variations on the category level between several years, many variables are comparable but not identical. the possible creation of explicit comparability is important 18 iassist quarterly spring 2006 information for using the data for analyses over time. conclusion all in all, it can be stated that the innovations of the new ddi version will improve the data documentation regarding comparability over time. for the german microcensus as an annually repeated survey with some breaks in the survey program, the grouping structure will simplify the documentation over time. for a group of studies, the migration from version 2.1 to 3.0 may become quite complex, but the possibilities for grouping multiple studies and for overriding information on a lower level offer a more efficient means of documentation, especially in the case of variable inconsistency between periods of time or single years. however, there are some limitations of the grouping model. the efficient use of the grouping mechanism for documenting comparability over time requires exact knowledge of the documented data and needs a lot of preliminary work. every extension of the grouping structure results in a higher branching as well, so that the administration of the ddi structure may become quite complex. moreover the flexibility of the structure is limited once a standard has been defined. as a consequence, the basic structure can not be modified to accommodate potential future changes in the survey program. this studyunit (2003) logicalproduct variablescheme categoryscheme variable: name category: label definition subgroup2 datacollection instrument questionitem: questiontext representation � derivation figure 4: ddi structure example 2 limitation concerns not only ddi, but is inherent in working with comparative standards. nevertheless, overall ddi 3.0 provides an instrument for a better managing and processing of metadata. these possibilities should be used to ensure high quality data documentation which is comparable in an international context. * jeanette bohr, andrea janßen und joachim wackerow, centre for survey, research and methodology (zuma), mannheim, centre for survey research and methodology, zuma, p.o. box 12 21 55 68072 mannheim germany. contact: bohr@zuma-mannheim.de. http://www.gesis. org/en/social_monitoring/gml/index.htm. references blank, grant and karsten boye rasmussen (2004). the data documentation initiative. the value and significance of a worldwide standard. in: social science computer review, vol. 22, no. 3, 307-318. nelson, chris (2006). ddi 3.0. conceptual model. iassist conference 2006 presentation. http://www. iassistdata.org/conferences/2006/presentations/w5_ddi. ppt thomas, wendy (2006a). codebook centric to life-cycle iassist quarterly spring 2006 19 centric: in the beginning .... iassist conference 2006 presentation. http://www.iassistdata.org/conferences/2006/ presentations/w5_thomas_inthebeginning.ppt thomas, wendy (2006b). organizing groups. iassist conference 2006 presentation. http://www.iassistdata.org/ conferences/2006/presentations/w5_thomas_groups.ppt endnotes 1 http://www.icpsr.umich.edu/ddi/org/index.html 2 http://www.iassistdata.org/conferences/2006/ presentations/ 3 in the following, the ddi structure examples are illustrated graphically, not in xml. 18 iassist quarterly 2016 iassist quarterly abstract the data documentation initiative (ddi) specification has gone through significant development in recent years. most health and demographic surveillance system (hdss) researchers in low and middle income countries (lmic) are, however, unclear on how to apply it to their work. this paper sets out considerations that lmic hdss researchers need to make regarding ddi use. we use the kisesa hdss in mwanza tanzania as a prototype. first, we mapped the kisesa hdss data production process to the generic longitudinal business process model (glbpm). next, we used existing glbpm to ddi mapping to guide us on the ddi elements to use. we then explored implementation of ddi using the tools nesstar publisher for the ddi codebook version and colectica designer for the ddi lifecycle version. we found the amounts of metadata entry comparable between nesstar publisher and colectica designer when documenting a study from scratch. the majority of metadata had to be entered manually. automatically extracted metadata amounted to at most 48% in nesstar publisher and 33% in colectica designer. we found colectica designer to have stiffer staff training needs and software costs than nesstar publisher. our study shows that, at least for hdss in lmic, it is unlikely to be the amount of metadata entry that determines the choice between ddi codebook and ddi lifecycle but rather staff training needs and software costs. lmic hdss studies would need to invest in extensive staff training to directly start with ddi lifecycle or they could start with ddi codebook and move to ddi lifecycle later. keywords hdss, open-access, metadata, ddi codebook, ddi lifecycle introduction investigators of hdss studies in lmic are realising the importance of preparing their existing data for open access. these data have been used to produce some of the key results leading to better understanding of hiv/aids among other diseases (ghys, zaba, and prins 2007; hallett et al. 2008; porter and zaba 2004; todd et al. 2007; zaba et al. 2013; ndirangu et al. 2011; streatfield et al. 2014; desai et al. 2014). they have been used to shed light on subsaharan africa mortality patterns (indepth network 2002; sankoh et al. 2014). providing open access will increase accessibility of these data to regional trainee scientists and the wider research community and thus maximise their public health benefit. human science research data documentation has gone through considerable methodological advances in recent years. one of these advances is the development of the data documentation initiative (ddi), a specification that is commonly used for documenting observational survey data (rasmussen and blank 2007; wellcome trust 2014). it uses the extensible markup language (xml) format (w3schools.com 2015) and has two main strands: ddi open-access for existing lmic demographic surveillance data using ddi by chifundo kanjala1, jim todd1, david beckles2, tito castillo3, gareth knight1, baltazar mtenga4, mark urassa4, and basia zaba1 iassist quarterly 2016 19 iassist quarterly codebook, originally called ddi 2, and ddi lifecycle, originally called ddi 3. ddi codebook is the simpler of the two and aims to describe a dataset in terms of its structure, contents and layout – a compilation of facts about a dataset mainly for archiving purposes. it has been used worldwide including in lmic through the international household survey network (ihsn) and the world bank (international household survey network 2013). the ihsn implementation of ddi codebook was done using ddi-compliant software for metadata management called nesstar publisher (digital curation centre 2013). once data have been documented in nesstar publisher, the resulting documentation can be presented in various forms including pdf versions of the codebook and cataloguing of the data in webbased catalogues. a commercial data repository and catalogue created by nesstar called nesstar server could be used. alternatively open source software called national data archive can also catalogue data and ddi-compliant metadata. nada was designed by the world bank and the ihsn to facilitate archiving and sharing their national data (international household survey network 2016). ddi lifecycle was developed from the premise that a dataset is an embodiment of a process that produced it, thus, it uses the data life cycle (figure 1) as its conceptual model. it comprises modules which are packages of metadata each roughly corresponding to a stage in the data life cycle. there is one related to study conceptualisation, another related to data collection, another catering for archiving and so on. ddi codebook metadata are still present in ddi lifecycle and are spread throughout its modular structure. it also captures metadata that describe associations between groups of studies. a number of tools for implementing ddi lifecycle are available. these include colectica designer, questasy (de bruijne and amin 2009; de vet 2013), ddi on rails (hebing 2015), dda ddi editor (jensen 2012) among others produced at the gesis leibniz institute for social sciences in germany (http://www.gesis.org/en/institute/) and the north american metadata technology (http://www.mtna.us/). we used colectica designer because when we started the documentation work it was one of the few available ddi lifecycle tools offering the most flexibility to meet our needs. closely related to the ddi lifecycle is the generic longitudinal business process model (glbpm), which outlines steps taken in the process of producing longitudinal data for social and human sciences. the glbpm is shown in figure 2. mapping an organisation’s data production process to the glbpm can determine what metadata to record at each step of data production since glbpm has been mapped to ddi lifecycle (barkow et al. 2013). the lmic hdss studies have generally used metadata standards at the research network level as shown by the example of the indepth network data repository (indepth network 2013a). to the best of our knowledge, only a few individual hdss studies, among them, the africa centre for population health (africa centre for population health 2015) and african population and health research center (african population and health research center 2015) have used ddi. for sites not using ddi, this has led to the documentation of a small subset of all the data that the studies generate, in many cases, less than 20% of the variables on which a typical hdss collects data. this means that the strengths and limitations of the data are not properly understood by secondary users, making it hard for them to interpret their analyses. to demonstrate the use of the ddi metadata standard to document ‘legacy’ data, we applied it to the existing kisesa hdss data. this task required consideration of the metadata editors to use, the amount of documentation needed when using ddi codebook and ddi lifecycle, staff training needs and approximate software costs. study settings and methods study settings the tazama project within the national institute for medical research, mwanza tanzania runs the kisesa open cohort study. it has been described in detail previously (marston et al. 2012; kishamawe et al. 2015; urassa et al. 2001). the backbone of the kisesa study is its hdss. the population in the study area had grown to over 35,000 by 2014 (kishamawe et al. 2015) from about 19,000 in 1994. follow-up data collection rounds have been done at roughly six-month intervals recording new births, migrations and deaths. in addition, marriages, pregnancies and education are recorded. paper questionnaires were used for data collection until round 25. since round 26, hdss data are collected electronically using portable digital assistants and cspro applications. while the kisesa study runs other studies including figure 1: data life cycle: ddi lifecycle conceptual model 20 iassist quarterly 2016 iassist quarterly cause of death analysis and hiv serological studies, we focus on describing the documentation of the hdss, which provides the sampling frame for all the nested tazama studies. once the hdss documentation is understood, it will be easier to apply the principles to the studies that rely on the hdss. the hdss component is implemented in broadly similar ways across a range of studies (sankoh and byass 2012) so such studies can relate to the kisesa experiences. study methods the data production process involved in the implementation of a typical hdss data collection round in kisesa is illustrated in figure 3. at the top is the data evaluation and analysis phase prior to an hdss round. going clockwise, we have the planning and preparation phase followed by activities related to fieldwork, while the last box shows the steps related to office data processing, storage and dissemination. each step was mapped to its closest equivalent within the glbpm (barkow et al. 2013). we then used the existing mapping from glbpm to ddi lifecycle (barkow et al. 2013) to guide us on the likely ddi metadata elements to use for documenting hdss data. once the mapping exercise was completed, we used nesstar publisher to produce ddi codebook and colectica designer for ddi lifecycle. for nesstar publisher, we used the ihsn metadata template and the step-by-step guide (dupriez and greenwell 2007), while for colectica designer we used the information model provided with the colectica online documentation (colectica 2015b). the actual documentation was done in three overlapping phases: preparation, data documentation, and creation of an internal data catalogue. in the preparation phase, we piloted the use of colectica designer and nesstar publisher. in colectica designer, we created an hdss series as a group within which all the hdss data from the numerous data collection rounds could be documented. for rounds 26 and 27, a study metadata package was created, using guidance provided by the colectica user’s guide (colectica 2015a). we gathered and entered foundational metadata including concepts, affiliated organisations and universes for variables, and added metadata pertaining to studylevel, data collection, data processing, dataset and variables. this pilot showed that the levels of training and finances required to do this work using locally recruited staff were not sustainably available for the project. on the other hand, ddi codebook seemed accessible from both our pilot work and examples from other studies (indepth network 2013b), and its use was agreed. two recent graduates from quantitative backgrounds were recruited and trained in the use of nesstar publisher – this initial training took two weeks. data in the project’s databases that required documentation was identified and relevant details lists of the database tables and locations of the databases on the project’s servers recorded. figure 2: generic longitudinal business process model iassist quarterly 2016 21 iassist quarterly figure 3: steps in the kisesa study hdss data collection round we started the documentation phase by importing the data into nesstar publisher where additional metadata were added. metadata not available in the data files were extracted from questionnaires, ethical clearance documents, funding proposals and other supporting documents and entered manually in nesstar publisher. after documentation, we went on to catalogue the data. finally, the metadata, data, supporting documents and publications based on the data were brought together into the data catalogue. the ddi codebook files were transferred from nesstar publisher to nada and we subsequently configured nada to suit our needs. the design of the catalogue provides for demarcation of collections of the data and their associated documentation – in this case we created a collection dedicated to kisesa hdss data. results mapping kisesa study hdss data production to glbpm the results of mapping one round of the kisesa study hdss data production process onto the glbpm are presented in figure 4. the glbpm steps are shown in square brackets. figure 4: kisesa study hdss documentation in nesstar publisher 22 iassist quarterly 2016 iassist quarterly the activities within the evaluation and analysis phase corresponded to the two glbpm steps research / publish (8) and retrospective evaluation (9). the preparation and planning phase corresponded to the glbpm’s first 3 steps which are evaluate / specify the needs (1), design / redesign (2) and build / rebuild (3). the fieldwork phase corresponded to the design/ redesign (2), the build / rebuild (3) and the collect steps (4). the processing, storage and dissemination phase corresponded to the process / analyse (5), the archive / preserve and curate (6) and the data dissemination/discovery (7) steps. data dissemination is done via the data catalogue at the project offices and through correspondence with the project head for remote access. in line with the two properties of the glbpm that it is not exhaustive and non-linear, not all sub-steps were used in the mapping. of the 53 sub-steps in glbpm, 28 were found to be relevant to the kisesa hdss. the excluded sub-steps fell into 3 broad categories: those that were not supported within the kisesa data management system, those that were not applicable to the hdss round under consideration and those that do not apply to the hdss type of studies. the examples of sub-steps not currently supported include 5.4 – imputing missing data, 7.5 – support for data citation, and 7.6 enhance data discovery, among others. most of the sub-steps in step 1 mainly applied to the initial census and first follow-up round of the hdss and were not frequently revisited in the subsequent rounds of the study. since the hdss involves the entire population within a geographically demarcated area, it does not apply any sampling so the sampling and weighting sub-steps are not applicable. implementation of ddi in nesstar publisher and colectica designer table 1 contains counts of some of the main items involved in the documentation of kisesa study hdss data. this gives an idea of the scale of the documentation involved. table 1: counts of items involved in the documentation of kisesa hdss item quantity hdss data collection rounds 27 data files 38 questionnaires 27 computer assisted interviews 2 paper questionnaires 25 variables5 1216 we had completed data documentation for 27 hdss rounds at the time of writing. starting from the baseline round to round 20, there is one data file per round. rounds 21 onwards have either two or three data files for each round, with one file holding household-level data and the other holding individual household members’ data. in round 26, the questionnaire comprises a hierarchical set of 36 householdlevel questions and 53 individual-level questions, generating 41 and 67 variables respectively in the household and individual data files, including derived and administrative variables. in round 27 there are 54 household-level questions and 104 individual level questions, generating 62 and 118 variables respectively. in developing the metadata repository, we extracted data from ms access and sql server databases into stata 12. within stata, we added notes, variable and value labels as needed. the resulting stata files were then imported into nesstar publisher and colectica designer (only two rounds, 26 and 27 for pilot). some metadata were automatically extracted from the stata files: categories, codes, variable names and labels, data file and variable notes. we compiled counts of the metadata items we considered to be important for hdss. the results are given in table 2. the four broad categories into which we classified the metadata are foundational, study level, data collection-related and datasets metadata. universes were identified both at study and variable levels. in nesstar publisher, even in cases where a number of variables shared the same universe, that universe had to be entered for each variable due to lack of mechanisms for reuse of metadata. this is a limitation of ddi codebook not of nesstar publisher. in contrast, in colectica designer, we entered each unique universe once and referred to that universe each time it applied, which explains why there are many more universes in nesstar publisher than in colectica designer. the ability to reuse metadata across studies meant we only needed 2 additional universes during documentation of round 27, since most of them had been entered in round 26. reuse of metadata also led to the reduction in categories, codes, concepts and organisations that needed to be entered for round 27 for colectica designer. categories and codes were automatically extracted from stata files but the concepts, universes and organisations had to be entered manually. automatically extracted foundational metadata contributed 38 per cent of all the foundational metadata needed in nesstar publisher and about 60 per cent of the foundational metadata in colectica designer. iassist quarterly 2016 23 iassist quarterly table 2: nesstar publisher (np) and colectica designer (colectica) documentation regarding study-level metadata, ddi codebook does not have the concept of grouping studies so we had no counts of metadata items for nesstar publisher in table 2 in the “hdss studies group” row. in colectica designer studies are grouped together in what is called a series. we put the hdss rounds together in an hdss series, documenting each round as a separate study. the amounts of metadata required for hdss at study level are comparable for nesstar publisher and colectica designer. there was little reuse of study-level metadata across studies as most of the metadata provided at study-level are specific to the particular study. the data collection section is the one where a lot more metadata are provided for in ddi lifecycle compared to ddi codebook. methodology description and collection events had similar metadata requirements for both nesstar publisher and colectica designer. however ddi lifecycle provides far more metadata and structure related to instrument description. it was possible for us to build digital versions of hdss paper questionnaires or cspro data entry applications for rounds 26 and 27 from colectica designer. the paper questionnaires that we built were similar to the ones that would have been used during the actual data collection if rounds 26 and 27 had used paper questionnaires. however, the data collection applications for cspro generated by colectica designer did not represent their final state, and more work would need to be done to include loops and skips as there are no inbuilt functions to do these in cspro so they are implemented using user-defined functions. ddi codebook mainly provides textual description and bibliographic information for a questionnaire, thus there are few metadata elements for hdss questionnaire documentation in nesstar publisher. the datasets metadata section is divided into metadata relating to a dataset as a whole and variable-level metadata. this is where we entered most of the metadata in nesstar publisher. in both colectica designer and nesstar publisher, variables within a given data file are linked to their source questions where applicable. the same source questions entered during instrument development are referred to in colectica designer. here we also see comparable amounts of metadata between nesstar publisher and colectica designer in round 26 and due to metadata reuse, fewer items are needed for round 27 in colectica designer, mainly to cater for variables not present in round 26. we distinguished between metadata that editors automatically extracted and those that we manually entered. in round 26, 48% of the metadata were automatically entered from stata files for nesstar publisher and 20% for colectica designer. round 27 had a similar percentage of automatically extracted metadata in nesstar publisher (44%) while in colectica designer automatically extracted metadata went up to 34%. further details on staff training needs and the software costs are shown in table 3. (page 28). table 2: nesstar publisher (np) and colectica designer (colectica) documentation round 026 np colec0ca round 027 np colec0ca founda0onal metadata universes categories codelists concepts organisa2ons automa2cally entered 108 41 41 14 12 81 15 32 32 14 12 62 180 67 67 16 12 134 2 (15 referenced from round 26) 8 (32 referenced from round 26) 8 (32 referenced from round 26) 2 (14 referenced from round 26) all 12 referenced from round 26 16 study-level metadata hdss studies group each hdss round automa2cally entered 41 53 49 41 37 (12 referenced from round 26) data collec0on metadata methodology instrument collec2on events 4 11 5 5 1237 8 4 11 5 5 331 (1233 referenced from round 26) 7 (1 referenced from round 26) data processing apach batch edit programs as external resources/ other materials automa2cally entered datasets metadata dataset variables automa2cally entered 20 2808 1404 20 2160 432 20 4680 2054 20 2090 (1512 referenced from round 26) 720 total automa0cally entered items total number of metadata items 1485 3105 494 3637 2188 5103 736 2212 1 24 iassist quarterly 2016 iassist quarterly nesstar publisher colec1ca designer pages of documenta1on in user manual read by documentalist 80 pages 100 pages pages of training material prepared for metadata entry staff powerpoint presenta1ons – 80 slides, 30 pages handbook. handbook under development 40 pages power point presenta1ons 250 slides self-study 1me and courses taken by documentalist 1 – 2 months ini1al self-study ihsn toolkit, nesstar publisher user’s guide and ddi codebook online documenta1on 1 week ddi lifecycle training and one day ddi / colec1ca workshop 4 6 months ddi lifecycle self-study and prac1cal work in colec1ca designer time taken to train metadata entry staff 2 weeks ini1al, 3 months during work not done cost of metadata prepara1on so?ware nesstar publisher free colec1ca designer monthly license us $65 per seat (logged in user) annual license us $59 per month perpetual license – us $2000 per seat cost of archiving service nada free nesstar server commercial fee not specified on website colec1ca repository us $5000 us $74000 depending on selected op1ons 1 table 3: training materials and software costs the documentalist used a combination of short courses and self-study of online resources to get started with ddi and its metadata editors. knowledge of ddi codebook and nesstar publisher was acquired using the ihsn resources in form of a toolkit comprising sample documentation in nesstar publisher and a step-by-step ddi codebook documentation guide (dupriez and greenwell 2007). in addition, the ddi codebook online documentation on the ddi alliance website6 was used. for colectica designer, the documentalist attended a one-week introduction to ddi lifecycle course, and a one-day introduction to ddi lifecycle and colectica course. in addition, he spent between 4 to 6 months of self-study of ddi lifecycle resources available on the ddi alliance website mainly in the form of ddi lifecycle documentation, conference presentations and working papers. parallel to that, practical activities were also carried out in colectica designer. to prepare metadata entry staff, we spent two weeks on initial nesstar publisher training. it then took 3 months of close supervision to get them comfortably working independently. regarding software costs, nesstar publisher is available for free while colectica is commercial software with pricing at the time of writing as given in table 3. most of the online resources were accessible to the documentalist but difficult to understand for metadata entry staff at our disposal. the documentalist made the online resources that he had accessed available to the metadata entry team with follow up explanations to help them understand the content. discussion we investigated the implementation of ddi on the existing kisesa hdss data. in particular, we paid attention to the identification of the steps involved in the kisesa hdss data production and their relationship to the glbpm, the choice of ddi tools to use, the amount of metadata to be entered, the staff training needed and the software costs involved. we used nesstar publisher and colectica designer as our ddi codebook and ddi lifecycle tools respectively. our first finding is that the number of metadata items that had to be entered in nesstar publisher and in colectica designer were comparable when an hdss round was documented from scratch. documenting a subsequent round reduced the amount of metadata entry drastically in colectica designer due to reuse of metadata from the earlier round. this is supported by the fact that we needed to enter 3105 items in nesstar publisher and 3637 items in colectica designer when we documented round 26 from scratch. round 27 required 5103 in nesstar publisher and 2212 in colectica designer. our second finding is that though the metadata editors automatically extracted some metadata from the stata files we used, we still had to manually enter the majority of the metadata in both colectica designer and nesstar publisher. this is supported by the observation that metadata automatically extracted from stata files for round 26 catered for 48% of the metadata entered in nesstar publisher while it was about 14% for colectica designer. in round 27 it was 43% and 33% respectively. in each case we still had to manually enter more than half of the metadata that we considered necessary. iassist quarterly 2016 25 iassist quarterly our third finding is that more staff training and stiffer financial demands were required to implement colectica designer than nesstar publisher. this is supported by the time taken to get the training done for the staff and the reported software costs. the documentalist spent about 2 months of initial study of ddi codebook, the ihsn toolkit and the nesstar publisher user’s manual before embarking on the preparation of training materials for metadata entry staff. it then took another one to two months to get the training material ready. for comparison, colectica designer took a week of formal training by ddi alliance-affiliated ddi lifecycle developers, an introduction to colectica pre-conference workshop and 4 – 6 months of online ddi lifecycle resources searching and study. concurrent to the self-study, the documentalist was having practical sessions learning the colectica designer software. with respect to costs, nesstar publisher and the nada software are free, whereas colectica designer is commercial and so are the colectica repository and portal (the data and metadata storage system and its web application for cataloguing the data). regarding the mapping of the kisesa hdss data production process, we mapped this process to 28 sub-steps of the glbpm. the glbpm sub-steps we did not use are in one of the three categories: not supported within the kisesa hdss data production process, not suitable for the round of hdss under consideration or not applicable to the hdss type of studies. this mapping helped to describe the kisesa hdss data production process in a standardised and coherent manner. we faced some challenges during the mapping. for some activities, we could not find the exact sub-steps to map them to. the mapping also required input from a wide range of staff involved in the data production process who often could not give immediate response as they needed to first study the glbpm. in those cases we made efforts to gather their understanding of the steps they were responsible for and we centrally mapped their feedback onto the glbpm. this procedure is in contrast to that used by another study that worked on a similar mapping but to a different reference model (ausborn, rotondo, and mulcahy 2014). gathering input from staff on their responsibility and then mapping centrally takes away the need for the concerned staff to understand glbpm. the generic tools for data documentation that we used, arguably among the best currently available, still involve a lot of manual entry of metadata and parsing through free-text documents, in the form of questionnaires, protocols or reports, in search of study-level metadata, involved organisations, the concepts being measured and so on. this requires trained documentation personnel who understand ddi, especially if ddi lifecycle is to be produced, having necessary skills to work out study concepts from proposals, questionnaires and publications. this does not mean that the ddi lifecycle standard is unsuitable, however; it just means that its complexity makes it difficult to use generic tools for most of the steps within the glbpm. in practice, hdss studies clearly do not need to leverage most of the additional features of ddi lifecycle; however, there are some parts of the standard that would be advantageous (referenceability, versioning and comparison, for example). the generic tools seem to be most useful once the ddi content has been created. this seems to suggest that a sensible next step would be to consider development of bespoke software solutions, funds permitting. the bespoke tools would cover the parts of the documentation process that involve manual metadata entry. much of the data dissemination and discovery (step 7 in the glbpm) could be supported by using the generic tools. the question of generic versus bespoke tooling therefore needs to be explored for each of the other process steps in the glbpm. we have only considered two metadata editors but there are other ddi lifecycle editors in development that are free -for example, ddi on rails (hebing 2015), the danish data archive’s ddieditor (jensen 2012) and questasy (de bruijne and amin 2009). it would be worthwhile to carry out a more extensive exploration of the wider range of tools to see if any of the ones we did not consider would offer distinct advantages in the documentation of hdss data. we chose colectica designer over the others as it was arguably the most generic at the time we were starting our documentation work. questasy, which was originally designed for the centerdata at tilburg university in the netherlands, is now being developed further to make it more generic (edwin de vet, scientific programmer at centerdata, personal communication). ddi on rails was not yet available when we started. other hdss studies have taken this route of documenting their existing hdss data using ddi codebook. these include the africa centre (ac) for health and population research in south africa (dr. kobus herbst, personal communication) and the africa population and health research centre (aphrc) in nairobi, kenya (aphrc, 2014). these two studies are larger than the kisesa study, covering populations of 85,000 (tanser et al. 2008) and 65,000 (beguy et al. 2015) respectively compared to kisesa’s 35,000. the ac hdss currently acts as a platform for 5 research programmes; each with its own sub-studies. since its inception in 2002, the aphrc has had more than 15 projects, using its hdss as a platform, compared to 4 sub-studies in kisesa. they are also better resourced in terms of it and programming staff, compared to kisesa. but even with this level of sophistication they have not yet adopted the more advanced technology offered by the ddi lifecycle approach, which has hitherto been used only by studies in more developed countries, such as the midus study in the usa (radler, iverson, and smith 2013), the closer project in the uk (gierl and johnson 2012), statistics denmark (nielsen, iverson, and smith 2013), and statistics new zealand (brown et al. 2012). it would appear that this technology will not be rapidly adopted by hdss in lmic. one important finding, which was not part of the original remit of this investigation, is awareness of how much harder it is to include in the study documentation a questionnaire that has been developed for collecting data on an electronic device rather than on paper. hdss, which moved to electronic data collection using specialist software like cspro, need to be aware that for documentation purposes they need to develop paper versions of the questionnaire for explanatory purposes, or supply the code and its interpretation (e.g., as screen shots) as part of the documentation package. summary in summary, our study shows that at least for a typical african hdss, it is not so much the difference in the amount of metadata to be entered but rather, the staff training requirements and the software costs that producers should consider when deciding between ddi codebook and ddi lifecycle. if available staff expertise is capable of learning and implementing ddi lifecycle, an hdss could directly start 26 iassist quarterly 2016 iassist quarterly with ddi lifecycle; otherwise, they would better start with ddi codebook and then move on to ddi lifecycle at a later stage. the kisesa study is used as an example but the general principles would apply to other african hdss studies. references africa centre for population health. 2015. ‘research data management platform’. http://www.africacentre.ac.za/index.php/data-research-manag ement. african population and health research center. 2015. ‘central data catalog’. http://aphrc.org/catalog/microdata/index.php/catalog. ausborn, scot, julia rotondo, and tim mulcahy. 2014. ‘mapping the general social survey to the generic statistical business process model: norc’s experience’. iassist quarterly, 21. barkow, ingo, william block, jay greenfield, arofan gregory, marcel hebing, larry hoyle, and wolfgang zenk-möltgen. 2013. ‘generic longitudinal business process model ddi working paper series – longitudinal best generic longitudinal business process model’. business, 1–26. beguy, donatien, patricia elung’ata, blessing mberu, clement oduor, marylene wamukoya, bonface nganyi, and alex ezeh. 2015. ‘hdss profile: the nairobi urban health and demographic surveillance system (nuhdss)’. international journal of epidemiology, dyu251. brown, adam, jeremy iverson, dan smith, and sally vermaaten. 2012. ‘powering official statistics at statistics new zealand with ddi-l and colectica: a case study’. in . http://www.eddi-conferences.eu/ocs/index.php/eddi/eddi12/paper/view/42. colectica. 2015a. ‘colectica — colectica 5.1 documentation’. http://docs.colectica.com/. ———. 2015b. ‘colectica information model — colectica 5.1 documentation’. http://docs.colectica.com/introduction/information-model/. de bruijne, marika, and alek amin. 2009. ‘questasy: online survey data dissemination using ddi 3’. iassist quarterly 33 (spring): 10–15. de vet, edwin. 2013. ‘update on questasy, a data dissemation tool based on ddi3’. in eddi13–5th annual european ddi user conference. http:// www.eddi-conferences.eu/ocs/index.php/eddi/eddi13/paper/view/71. desai, meghna, ann m buff, sammy khagayi, peter byass, nyaguara amek, annemieke van eijk, laurence slutsker, john vulule, frank o odhiambo, and penelope a phillips-howard. 2014. ‘age-specific malaria mortality rates in the kemri/cdc health and demographic surveillance system in western kenya, 2003–2010’. digital curation centre. 2013. ‘nesstar | digital curation centre’. june 12. http://www.dcc.ac.uk/resources/external/nesstar. dupriez, olivier, and geoffrey greenwell. 2007. ‘quick reference guide for data archivists’. http://www.ihsn.org/home/node/544. ghys, peter d, basia zaba, and maria prins. 2007. ‘survival and mortality of people infected with hiv in low and middle income countries: results from the extended alpha network.’ aids (london, england) 21 suppl 6 (november): s1–4. doi:10.1097/01.aids.0000299404.99033.bf. gierl, claude, and jon johnson. 2012. ‘70 years of uk birth cohort data into ddi lifecycle?’ in eddi12–4th annual european ddi user conference. http://www.eddi-conferences.eu/ocs/index.php/eddi/eddi12/paper/view/38. hallett, timothy b, basia zaba, jim todd, ben lopman, wambura mwita, sam biraro, simon gregson, j ties boerma, and alpha network. 2008. ‘estimating incidence from prevalence in generalised hiv epidemics: methods and validation’. plos med 5 (4): e80. hebing, marcel. 2015. ‘a metadata-driven approach to panel data management and its application in ddi on rails’. indepth network. 2002. population, health and survival at indepth sites. international development research centre. international household survey network. 2013. ‘mission and objectives | ihsn’. http://www.ihsn.org/home/content/about/objectives. ———. 2016. ‘microdata cataloging tool (nada) | ihsn’. http://www.ihsn.org/home/software/nada. jensen, jannik. 2012. ‘ddieditor’. in eddi12–4th annual european ddi user conference. kishamawe, coleman, raphael isingo, baltazar mtenga, basia zaba, jim todd, benjamin clark, john changalucha, and mark urassa. 2015. ‘health & demographic surveillance system profile: the magu health and demographic surveillance system (magu hdss)’. international journal of epidemiology, dyv188. marston, milly, denna michael, alison wringe, raphael isingo, benjamin d clark, aswile jonas, julius mngara, et al. 2012. ‘the impact of antiretroviral therapy on adult mortality in rural tanzania’. tropical medicine & international health 17 (8): e58—-e65. doi:10.1111/j.1365-3156 .2011.02924.x. ndirangu, james, marie-louise newell, claire thorne, and ruth bland. 2011. ‘treating hiv-infected mothers reduces under 5 years of age mortality rates to levels seen in children of hiv-uninfected mothers in rural south africa.’ antiviral therapy 17 (1): 81–90. nielsen, mogens, jeremy iverson, and dan smith. 2013. ‘standardized quality declarations with ddi, sdmx, and colectica’. in . porter, kholoud, and basia zaba. 2004. ‘the empirical evidence for the impact of hiv on adult mortality in the developing world: data from serological studies.’ aids 18 suppl 2 (suppl 2): s9–s17. radler, barry, jeremy iverson, and dan smith. 2013. ‘applying ddi to a longitudinal study of aging’. in north american data documentation initiative conference (naddi 2013), university of kansas, lawrence, kansas. rasmussen, karsten boye, and grant blank. 2007. ‘the data documentation initiative: a preservation standard for research.’ archival science 7 (1): 55–71. sankoh, osman, and peter byass. 2012. ‘the indepth network: filling vital gaps in global epidemiology’. international journal of epidemiology 41 (3): 579–588. sankoh, osman, david sharrow, kobus herbst, chodziwadziwa whiteson kabudula, nurul alam, shashi kant, henrik ravn, abbas bhuiya, le thi vui, and timotheus darikwa. 2014. ‘the indepth standard population for low-and middle-income countries, 2013’. global health action 7. streatfield, p kim, wasif a khan, abbas bhuiya, syed ma hanifi, nurul alam, eric diboulo, ali sié, et al. 2014. ‘malaria mortality in africa and asia: evidence from indepth health and demographic surveillance system sites’. global health action 7: 10.3402/gha.v7.25369. doi:10.3402/gha.v7 .25369. tanser, frank, victoria hosegood, till bärnighausen, kobus herbst, makandwe nyirenda, william muhwava, colin newell, johannes viljoen, tinofa mutevedzi, and marie-louise newell. 2008. ‘cohort profile: africa centre demographic information system (acdis) and population-based hiv survey’. international journal of epidemiology 37 (5): 956–62. iassist quarterly 2016 27 iassist quarterly todd, jim, judith r glynn, milly marston, tom lutalo, sam biraro, wambura mwita, vinai suriyanon, ram rangsin, kenrad e nelson, and pam sonnenberg. 2007. ‘time from hiv seroconversion to death: a collaborative analysis of eight studies in six low and middle-income countries before highly active antiretroviral therapy’. aids 21: s55–63. urassa, m, j t boerma, r isingo, j ngalula, j ng’weshemi, g mwaluko, and b zaba. 2001. ‘the impact of hiv/aids on mortality and household mobility in rural tanzania.’ aids 15 (15): 2017–2023. w3schools.com. 2015. ‘xml introduction what is xml?’ http://www.w3schools.com/xml/xml_whatis.asp. wellcome trust. 2014. ‘enhancing discoverability of public health and epidemiology research data’. wellcome trust. https://wellcome.ac.uk/sites /default/files/enhancing-discoverability-of-public-health-and-epidemiology-research-data-phrdf-jul14.pdf. zaba, basia, clara calvert, milly marston, raphael isingo, jessica nakiyingi-miiro, tom lutalo, amelia crampin, laura robertson, kobus herbst, and marie-louise newell. 2013. ‘effect of hiv infection on pregnancy-related mortality in sub-saharan africa: secondary analyses of pooled community-based data from the network for analysing longitudinal population-based hiv/aids data on africa (alpha)’. the lancet 381 (9879): 1763–71. notes 1. london school of hygiene and tropical medicine 2. independent it consultant, uk 3. university college london hospitals nhs foundation trust 4. national institute for medical research, mwanza tanzania 5. in this case we are counting the instance of each variable within a data collection round as a distinct variable even though many of the variables remain unchanged across data collection rounds 6. http://www.ddialliance.org/ a computer center survey: what students don't like gary p, fisher hay-hug6ins actuarial consultants philadelphia, pa introduction the idea for a survey of users at the "metro" university computer center grev) out of personal observations made when using the facility in september 1980. service times -both to use a keypunch machine and to receive one's printed program -seemed inordinately and uncomfortably long. cramped and crowded waiting areas made these service delays even more unpleasant. the consultants on duty were sometimes nowhere in sight, sometimes besieged by long lines of of students, and usually, once one got an interview, unfriendly. assuming that these experiences were not unusual, i hypothesized that lengthy service times and inadequate consulting might lead to dissatisfaction among users with the computer facilities. a queueing study would have provided precise waiting times and queue lengths for computer center services, but my concern was the impact these factors have on user satisfaction. i anticipated that those segments of the user population with more limited access to the facilities -commuting students traveling 25 or more minutes from home -would feel the waiting times more intensely. these students have less flexibility than "metro" residents in choosing when to avoid those crowded periods that arise unexpectedly at the center. long-distance commuters, i predicted, should therefore experience longer waits and rate the center's performance lower than their non-commuting counterparts. methodology the questionnaire i devised for the pre-test a two-page, thirty-question self-administered survey which solicited demographic information, ratings of some services, average waits for turnover and keypunches, suggestions for improvement (more keypunches, longer hours, and so on), and user characteristics. these user characteristics included how often and at what times a student used the center and what languages he could program 56 my sample of 40 computer center users consisted mostly of males (67.5%), commuters (72.5%), and either sophomores (37.5%) or pre-juniors (27.5%). exactly half of the students travelled more than 25 minutes to reach the center. the relative absence of freshmen (7.5%) and graduate students (7.5%) from the sample is notable but not inexplicable. few freshmen take computer courses because they must first satisfy curriculum requirements and course prerequisites; many graduate faculty (in business or library science, for example) encourage their students to use the facilities at a computer utility elsewhere. the sample embraces a fairly sophisticated group of programmers: 70% report that they know more than one high-level or assembler language; 90% know fortran; 55% know languages other than pl/1, fortran, or cobol. respondents used the center an average of 3 days and 8.5 hours during the week preceding the survey, which took place the week of february 8, 1981. the 40 questionnaires were completed within two full days (tuesday and wednesday). because most students did not answer questions that demanded a written response, i was unable to collect data on users' majors or the course that required them to use the center. index of satisfaction to measure fully a user's satisfaction with the computer facilities, i devised an index of satisfaction composed of the following responses: response value center hours should not be extended blank cards are usually available consultants are helpful consultants are knowledgeable keypunch instructions are clear and machine is easy to operate card reader instructions are clear and machine is easy to operate operating hours are convenient computer facilities are adequate total 10 scores on the satisfaction index included all values from to 10 with a median of 5 and an interquartile range of 3. individual scores were collapsed to create three degrees of user satisfaction, low (0-4), moderate (5-7), and high (8-10) for crosstabulation and analysis. (table 1 the index of satisfaction encompasses and enlarges upon the user's response to the question of the center's adequacy. comprising a number of sources of user satisfaction — ratings of consultants, readability of machine instructions and ease of machine operation, adequacy of operating hours — the index gives a fuller, more objective picture of the user for 57 crosstabulations with demographic variables. in crosstabulations with service factors (turnover and keypunch waits, consultant ratings), the adequacy rating was principally used, since the satisfaction index included too many items unrelated to the problems under analysis. responses concerning waiting times were excluded from the index because the a priori determination of how long a waiting time must be before it equals 1 point on the index begs the question of this study. results the median score of 5 on the 10 point satisfaction index strongly suggests that services could be improved at the "metro" computer center. 44.7% of the respondents rated the center not adequate, contributing to these low index scores. students were also critical of the center's operating hours and consultants. although 74.450 said that the current operating hours were convenient, 64.7% would still like those hours extended. consultants were rated as more helpful (70% found them helpful) than knowledgeable (62.5%). more positive ratings include: 93.9% thought that keypunch instructions were clear and the machine was easy to operate; 87.5% felt the same way about the card reader; 68.8% reported that an adequate supply of blank cards was usually available. to uncover those factors that might contribute to a user's low satisfaction score or rating of the center as not adequate, i crosstabulated demographic variables, user characteristics, and service factors with user satisfaction; and service factors with the rating of center adequacy. negative ratings of consultants (not helpful, not knowledgeable) and high turnover times (over 15 minutes) were the service factors most strongly associated with a user's rating of the center as not adequate (tables 2-4). the availability and waiting times for a keypunch, conversely, exhibited little or no relationship with the adequacy rating (tables 5-6). all three of the service factors that were strongly associated with center inadequacy were only slightly less related to low user satisfaction (tables 7-9). even though both consultant ratings did account for one point each in the index, one can not discount the strong association between user dissatisfaction and the belief that a consultant is not helpful or knowledgeable. on the other hand, only one item related to keypunching revealed any association with user satisfaction. longer keypunch waits, contrary to expectations, tended to go along with high satisfaction scores (tables 10-11). those students who spent the most hours at the center during the week of february 8 (though not necessarily the most days) reported lower satisfaction scores than other students (table 12). the strong associations of turnover time and the availability of helpful and knowledgeable consultants with user satisfaction and center adequacy suggest that difficulties in the running and debugging of computer programs, rather than their keypunching, greatly contribute to a student's rating of the center's performance. "metro" computer center 58 users appear to be goal-oriented. unproductive moments spent waiting for programs to work prove more unsatisfactory than time spent on routine tasks (waiting for keypunches to become available). and the more hours a student spends at the center running, consulting about, and debugging his programs, the greater his dissatisfaction with services tends to be. independent demographic variables (sex, class, commuter or resident student) had little observed effect on user satisfaction. students who travel 25 or more minutes to reach the center reported ratings and exhibited user habits mostly indistinguishable from other students. contrary to expectations, the association between a student's travel time to the center and satisfaction score was weak and statistically insignificant (table 13). these long-distance commuters did not rate consultants appreciably lower than did other students: 71.4% to 68.8% helpful, 60% to 68.8% knowledgeable. they don't use the center more (7.8 to 8.1 hours per week) or at different hours. ratings of center adequacy did not vary significantly when crosstabulated with travel time (table 14). students who travel 25 or more minutes to the center, however, did tend to wait longer for keypunches to become available (table 15), as expected. but longer waits for keypunches, as has been shown, do not lead to lower ratings of the center. it was in their attitudes toward current operating hours that these students were most distinctive. only 57.9% of them, as opposed to 94.7% of the other students, found these hours convenient. extending these operating hours, however, did not receive significantly more approval from long distance commuters 68.8% to 58.8%. one must conclude from this data that problems incurred in the running and debugging of problems have a greater impact on a traveling student's rating of the center than do his scheduling problems and waits for keypunches. conclusions while most students found that keypunches and cardreaders were easy to operate and that center hours were convenient for them, they were less affirmative in their ratings of consultants and the overall adequacy of the center. 44.7% thought that the facilities were not adequate; this contributed greatly to the median of 5 on the 10 point user satisfaction index. the unavailability of and long waits for keypunches did not, however, tend to lower user satisfaction. rather, long job turnover times and low ratings of consultants strongly associated with an inadequate rating of the facility as well as with low satisfaction scores. it is not the waiting, then, but what you wait for that matters to users. disturbances in the running and debugging of programs generate user dissatisfaction. given this goal-orientation of "metro" computer users, efforts spent on improving consulting services and reducing turnover times should have the greatest impact on improving computer center performance, as measured by its users. employing more dispatchers and consultants would be more ameliorative than purchasing new keypunch machines. 59 table "1. user satisfaction index score frequency percent 4 low 15 40 5 7 medium 12 30 8 10 high _1_2 _^ 40 100 table 2. center adequacy by averaqe turnover time average turnover time (ninutes] adequate 0-15 16 or more yes 71% 36% no 29 64 100% 100% n=17 n=14 table 3. center adequacy by consultant knowledceability consultants are knowledgeable adequate yes no yes 80^^ 18% no 20 82 100% 100% n=20 n=ll table 4. center adequacy by consultant helpfulness consultants are helpful adequate yes no yes 76% 25% no 24 75 100% 100% n=21 n=8 60 table 5. center adequacy by average wait for keypunch average wait (minutes] adequate 0-5 6-10 11-15 yes 75% 36% 50% no 25 64 50 100% 100% 100% n=4 n=ll n=4 table 6. center adequacy by availability of keypunches keypunches are available adequate usually not usually yes 36% 71% no 64 29 100% 100% n=28 n=7 table 7. user satisfaction index by average turnover time average turnover time (minutes) satisfaction 0-15 ^ 16 or more high 47% 25% medium 35 13 low 18 62 100% 100% n=17 n=16 table 8. user satisfaction index by consultant helpfulness consultants are helpful satisfaction yes no high 55% 8% medium 30 25 low 15 67 100% 100% n=21 n=9 61 table 9. user satisfaction index by consultant knowledgeabi lity consultants are knowledgeable satisfaction yes no high 55? 8% medium 30 25 low 15 67 100" 100^,: n=20 n=12 table 10. user satisfaction index by average wait for keypunch average wait for keypunch (minutes) satisfaction 0-5 6-10 11-1! high 0% 36^ 50^ medium 20 27 25 low 80 36 25 loos 100:; loo'i n=5 n=ll n=4 table 11. user satisfaction index by availability of keypunches keypunches are available satisfaction usually not usually high 34"^ 25% medium 32 25 low 34 50 100? 100? ri=29 n=8 62 table 12. user satisfaction index by hours spent in center hours spent in center satisfaction 0-5 6-10 11-20 high 565;; 15% 9% medium 25 31 36 low 19 54 55 loo;;; 100% 100% n=16 n=13 n=n table 13. user satisfaction index by travel time to center travel time to center (minutes) satisfaction 0-25 25+ high 32% 30% medium 42 20 low 25 50 100% 100% n=19 n=20 table 14. center adequacy by travel time to center travel time to center (minutes) adequate 0-25 25+ yes 63% 50% no 37 50 100% 100% n=19 n=18 table 15. average wait for keypunch by travel time to center travel time to center (minutes) average wait 0-25 25+ 0-5 min. 38% 18% 6-10 min. 50 55 11-15 min. 12 27 100% 100% n=8 n=ll 63 notes on the social sciences as producers of technologies richard c. rockwell" social science research council washington^ d.c. the focus of this conference is on computerization--the premier technology of the late twentieth century--and what it has meant for the social sciences. in these remarks i consider the production of new technologies by the social sciences themselves. social science research is the source of a number of technological innovations that have proved their utility and value in the commercial market place. this assertion can be supported by examples and by rhetoric, but not yet by systematic research and quantitative measurement. in this paper i develop approaches to the measurement of social science technologies (sst) and suggest research questions that i would like to see social scientists address in this area,' definitions and examples of social science technologies, a simple dictionary definition of technology is "applied science " a somewhat less broad definition refers to a technical method of achieving a practical purpose. neither definition conveys the sense that technology requires a physical process or hardware, but we easily lapse into the equation of technology with tangible things. that fallacious equation has led to the virtual exclusion of sst from official statistics and research on technological innovation and diffusion, and this exclusion has costs for both the science and the society. * this paper was delivered at the 1981 ifdo/iassist conference in grenoble. the views expressed are not necessarily those of the social science research council ' i draw most of my understanding of technology and of the role of the social sciences in society from that society with which i am most familiar, the united states, and i would therefore like to avail myself of the international character of this conference by learning the applicability of my remarks in other contexts. 49 there are many different sorts of practical knowledge that can be put to use. what distinguishes technology from other applications of knowledge is that technology has its roots in scientific research. almost all of the sciences produce technologies, and i shall show that the social sciences are well-represented in this company. where our technologies perhaps most differ with those produced by physics, chemistry, and biology is that our knowledge often directly competes with knowledge based on tradition, experience, and "common sense." but that competition has been strikingly successful in the last fifty years, perhaps because of the growth of our body of knowledge--facts, generalizations, methods, concepts, and theories. there are many applications of social science other than sst. social commentary and criticism, whether by accredited social scientists or by others, has shaped our present world, but by sst's i do not mean these or more general contributions to knowledge, culture, or enlightenment. nor do i mean the design and evaluation of public programs and public policy, or the data gathering and societal monitoring activities that are intrinsic to the information systems of every advanced scientific-industrial society. i do not mean the study of human factors in technologies from other sciences, or the social impacts of these technologies. nor, finally, do i mean the various techniques that the social sciences have developed in pursuit of their own research aims, such as coding of open-ended questions, computer manipulation of large data sets, and computer-assisted interviewing, except insofar as these techniques have also entered the market place. kenneth prewitt, president of the u.s. social science research council, listed a number of sst's in his annual report for 1979-1980. he writes: "standardized educational and intelligence testing, economic forecasting, political polling, psychotherapy, man-machine system design, programmed language instruction, consumer research, cost-benefit analysis, behavior modification, and demographic projection are examples of, but do not exhaust the list of, multimillion dollar industries in the united states that started as social science discoveries" (prewitt, 1980). i would of course add to this list, while taking away psychotherapy, which i do not understand to have much of a base in social science research. among my additions would be self-paced programmed instruction, life tables for the writing of mortality insurance, quality control circles in japanese automobile factories, behavioral profiles of assassins and airplane hijackers, "total immersion" language instruction, and standardized personnel testing procedures. most of these technologies are rooted in more than one of the social sciences. mathematical and statistical methodologies undergird them, just as they undergird technologies produced by our sister sciences of physics, chemistry, and biology. as with most other technologies, persons using them are now largely responsible for their further development and broader application, and social scientists may or may not make further contributions. but each of these technologies was dependent in its early stages, at least, on a seed of knowledge from social science research: self-paced programmed instruction grew from laboratory studies of learning and reinforcement; standardized personnel testing grew from studies of personality and ability; and quality control circles grew from research on group dynamics, productivity, and morale. each is based in some way 50 on concepts or models that social scientists have developed: life tables, on a special case of stable population theory; econometric forecasting models, on either demand-side or supply-side models of national economies; and behavior modification, on organismic responses to reward and punishment. social science even has its own engineering departments and technical advice-giving programs, many of which use sst: city planning, social work, public finance, organizational administration, communication, land use management, public law, and program evaluation. to depart for a moment from these existing technologies, it is an interesting and possibly profitable endeavor to speculate about the sst's that may be important in the future and the research they would require. a critical problem facing a nuclear society is the disposal of wastes that will remain toxic for tens of thousands of years. the technological solution is not solely to be found in concrete, steel, glass, and deep burial. assuming that the chemical, geological, and structural problems are solved, what is the technology that will communicate across the generations the specific message of danger and ensure that this message is believed and acted on? this technology must work for a longer time than has existed complex human society and for many multiples of the time that any human government has stood. it must carry its message undistorted despite the evolution of language for a period longer than was required for the appearance of the many languages of europe and asia. it must survive what could be massive social change, disruption, and war. even over a forty-year period there appears to have been some deterioration of knowledge about those toxic wastes generated by the manhattan project, which was the world war ii atom bomb program that had sites scattered in half the states of the u.s. (walsh, 1981). if the social sciences have knowledge that bears on this problem, it may be found in anthropology's study of priesthoods and ritual, and in history's study of long-lived institutions such as military tradition and religious authority. we could engage in similar speculations about technologies for a globe that will bear double its current population or a society that will implement the communications capabilities the electronic engineers have readied. my guess is that society will demand much more from sst's in the future than it does today, and that technological development will become a recognized part of the work of the social science disciplines. indicators of social science technologies in most of the advanced scientific-industrial nations of the north there exists a body of official statistics that provide quantitative descriptions of aspects of science and technology. these statistics have proved to be useful in setting national science policies and in assessing strengths and weaknesses in the various scientific disciplines and the technological industries that draw on them. legislators and program administrators look to these statistics for descriptions of the conditions of science and for an accounting of what science brings to a society. scholars use these indicators to pursue their own studies of scientific innovation and technological diffusion (elkana et al , 1978; zuckerman and miller, 1980). 51 missing from these national statistics is evidence of the technologies that the social sciences have developed. investments in testing geological strata for oil deposits are tabulated, but investments in testing students for the potential to learn to program computers are not. investments by computer scientists in developing compiler-level languages are recorded but research by linguists on the sign language of the deaf is not. most industries known to man are accounted for except those built on sst's. official indicators of science and technology in the united states began in 1953, with the publication of data from surveys on the "funding and performance" of research and development. in 1979 the u.s. national science board published science indicators 1978 , the fourth in a series of quantitative assessments of u.s. science and technology. this report concentrates on the "inputs" to science of money and people, classified by field, economic sector, and industry. the "outputs" of science and technology are recorded in patents and published articles. as might be surmised from this characterization of the data employed, both basic research and technological development in the social sciences are inappropriately or inadequately treated. in a review of the 1976 report, the general accounting office (an agency of the u.s. congress) observes that the model inherent in these statistics assumes that "physics, chemistry, biology, and math are the most important sciences, and hence deserve the most complete coverage (1979:20)". one cannot look to this report for quantitative information on sst's. nor can one look to the measurement of scientific and technical activities , the influential "frascati manual" proposed by the organization for economic cooperation and development (1976) as a guide to standard practice for surveys of research and experimental development. the frascati manual generally excludes technologies from its purview, and explicitly excludes sst's from its three categories of basic research, applied research, and experimental development. the manual says the following: "many social scientists perform work in which they bring the established methodologies and facts of the social sciences to bear upon a particular problem, but which cannot be classified as research. the following are examples of work which might come in this category and are not r[esearch] and d[evelopment]: interpretative commentary on the probable economic effects of a change in the tax structure, using existing economic data; forecasting future changes in the pattern of the demand for social services within a given area arising from an altered demographic structure; operations research (or) as a contribution to decision-making, e.g., planning the optimal distribution system for a factory; the use of standard techniques in applied psychology to select and classify industrial and military personnel, students, etc., and to test children with reading or other disabilities" (oecd, 1976:24). "^^ we cannot incorporate sst into official statistics by adopting the 2) nor is any guidance to be found in a book on social forecasting by olaf helmer, to which he gave the title social technology (1966). 52 categories and concepts that are used for other technologies. the industrial categories that have apparently proved to be useful for the measurernent of physical, chemical, and biological technologies will not, i think, prove useful in the design of indicators of sst's. these technologies are crosscutting, put to work in more than one industry. personnel testing oervades all industrial sectors; demographic projections are used by manufacturers of baby food and by electric utilities; and the technology for securing toxic waste will serve both the chemical and nuclear industries. in the place of a classification of technology by industry, i suggest experimenting with a classification in terms of the aspect of social, economic, or political affairs that an sst serves. there is a great variety of suitatdle classifications available in the literature on social indicators. for example, the preliminary social indicators guidelines of the united nations identify twelve areas: learning and educational services; health, health services, and nutrition; time use; public order and safety; and so forth (un, 1978:43-50). the frascati manual offers a classification by particular socio-economic objectives for research and development, which incidentally also includes twelve categories: production and rational use of energy, protection of the environment, transport and telecommunications, defense, and so on (oecd, 1976:43-50). another level of classification could categorirre the unit of society that is directly affected by an sst: individuals, business enterprises, governments, other organizations, relations among organizations, and relations among individuals and organizations. still another level of classification comld be based on the concept of basic human needs, of which johan galtung(1980:66) has provided a working list. in the four major categories of security, welfare, identity, and freedom galtung itemizes the basic needs of individual human beings. these classifications have potential for informing the study of all technologies, not just the sst's produced by the social sciences; but because they begin with social concepts, they may be particularly appropriate for our analyses. statisticians count the patents awarded for physical, chemical, or biological innovations; count the scientists and engineers involved in research and development; and count the money spent by governments and industry on research, development, and adoption of technologies. economists estimate the costs and benefits of specific technological changes, sometimes extending their analyses to estimates of the impact on the length or quality of human lives. are such statistics appropriate as points from which to begin the measurement of sst's? i suspect that they are, with the exception of patent statistics, chiefly because they are reasonable things to study and also because such statistics will permit comparisons among the various technologies. for example, we could compare the costs and benefits of social science and solutions to problems of urban transportation, and the increase in productivity and hardware accounted for by quality control circles may be compared to the increase that automation provides. in parallel to these conventional objective measures, we should collect subjective reports from individuals on their perceptions of how sst's have affected their lives, with particular regard to their attitudes, feelings. 53 happiness, and levels of satisfaction. these subjective indicators have been illuminating in such diverse areas as crime statistics and studies of economic growth, and they probably have much to tell us about personnel testing, behavior modification, quality control circles, and man-machine system design. research on social science technologies research can determine which sectors of society control the use of the various sst's and who benefits (or suffers) from their use. it is my impression that most sst's are in the hands of organizations, not of individuals, and that it is often organizational ends, rather than individuals' needs, that they serve. personnel testing serves an organization's needs for differentiating among candidates for a position, but is rarely used to serve the individual's needs for choosing an occupation. demographic and economic forecasting are useful to governments and large businesses, and are under their control. schools elect methods of instruction, and unions or companies organize quality control circles. with the exception of psychotherapy, the practitioners of sst's usually deliver their services to organizations, not to individuals. it is perhaps not odd that sst's, being rooted in understandings of social life, serve the ends of collectivities such as businesses and governments, and not necessarily those of individuals, but i admit i find this disquieting. we need studies of the effect of sst's on the welfare of entire societies. to what degree does a technology serve a society's goals? in the united states what is known as the "social indicators movement" partly grew out of the perception that one sst, economic accounts and econometric forecasting, was not adequate for the needs of society and was in fact rather dangerous. the social indicators movement proposed to create a competing sst for enabling a society to shape its collective fate. lastly, we need assessments of the quality of the knowledge that underlies an sst. life tables, and the more general stable population models, are the products of a research tradition of unusual strength for the social sciences. demographers have a data base of fairly high quality that contains observations extending over many years for many societies. their equations are based on well-understood, straightforward theory, and years of experience in the insurance industry have shown that life tables can be relied upon. is the research base equally strong for quality control circles, behavioral profiles, and personnel testing? i suspect that there are some technologies in wide use that are based on defective research and partial knowledge, and that not all of these are from physics, chemistry, and biology. there is a degree to which social science remains responsible for the scientific quality of the technologies it creates, just as nuclear physics recognizes its continuing responsibility for nuclear engineering. conclusion social scientists have conducted important research on the technologies produced by the physical, chemical, and biological sciences. their research encompasses the nuclear accident at three mile island, pennsylvania; 54 the social effects of aviation and of the elevator; the jacquard loom; highways and airports; and water irrigation systems. social science theorists such as ogburn, lenski, and hawley give a prominent place to technology. international journals publish scholarly articles on science and technology by economists and sociologists. missing from our own research, from our theories, and from the advice we give to official statisticians is the concept of the social sciences as part of the system of science that produces technologies at work in the economy. it is almost as if we discriminate against applications of our own knowledge, preferring the exotic work of other disciplines. comprehensive statistics and research on science and technology must include the sst's in their domains. the sst's are important to the society, the polity, and the economy, and scholars of science and technology will profit from study of them. references elkana, yehuda, joshua lederberg, robert k. merton, arnold thackray, and harriet zuckerman 1978. toward a metric of science: the advent of science indicators. new york: john wiley and sons_ u.s. general accounting office 1979. science indicators: improvements needed in design, construction, and interpretation, washington: g.a.o. galtung, johan 1980. "the basic needs approach." pp. 55-125 in katrin lederer ed.. human needs: a contribution to the current debate. cambridge, mass.: oegeschlager, gunn and hain. helmer, olaf 1966. social technology. new york: basic books. u.s. national science board 1979. science indicators 1978. washington: u.s. government printing office. organization for economic cooperation and development 1976. the measurement of scientific and technical activities. paris: oecd. prewitt, kenneth 1980. "annual report of the president 1979-80." pp. xiii-xxvii in annual report 1979-1980 of the social science research council. new york: s.s.r.c. united nations 1978. social indicators: preliminary guidelines and illustrative series. new york: united nations. walsh, john 1981. "a manhattan project postscript." science 212:1369-1371. zuckerman, harriet and roberta balstad fliller. science indicators: implications for research and policy. vol. 2 no. 5-6 of scientometrics. amsterdam: elsevier scientific publishing co. 55 vol244 iassist quarterly winter 2000 13 since 1985 the international social survey programme (issp) has conducted annual social surveys in the participating countries covering relevant topics from the social sciences. the issp was founded in the early eighties by four social science research institutes for the purpose of adding an international comparative aspect to the existing national social surveys. the founding members were: the issp agreed to four general principles: 1. jointly develop topical modules dealing with important areas of social science 2. field the modules as a fifteen-minute supplement to the regular national surveys (or a special survey if necessary) 3. include an extensive common core of background variables, and 4. make the data available to the social science community as soon as possible. as of 2001 the number of participants has grown from the original four countries to 38 countries worldwide. the map shows the geographical distribution of the current issp members. the member states are: australia japan austria latvia bangladesh mexico brazil netherlands bulgaria new zealand canada norway chile philippines cyprus poland czech republic portugal denmark russia finland slovakian republic flanders slovenia france spain germany south africa great britain sweden hungary switzerland ireland taiwan israel usa italy venezuela issp datawizard – computer assisted merging and archiving of distributed international comparative data by robert strötgen & rolf uher* yrtnuoc etutitsni yevrus asu cron noinipolanoitan ,retnechcraeser foytisrevinu ogacihc ssg yevruslaicoslareneg ynamreg amuz ,negarfmurüfmurtnez nedohtem nesylanadnu sublla eniemeglla egarfmusgnureklöveb netfahcsnessiwlaizos\ taerg niatirb retneclanoitaneht ecneicslaicosrof dnalaicosremrof( gninnalpytinummoc ,)rpcs–hcraeser nodnol asb laicoshsitirb sedutitta ailartsua sssr foloohcshcraeser ,secneicslaicos nailartsua lanoitan ,ytisrevinu arrebnac sssn laicoslanoitan yevrusecneics 14 iassist quarterly winter 2000 the first issp survey in 1985 was conducted in 6 countries, the original 4 plus italy and austria and had the topic ‘role of government’. the following topics have been fielded since then or are being planned for the near future: replicating topics over time has enriched the international comparative aspect by adding a time-series component. at the 1986 issp meeting in mannheim the zentralarchiv1 was chosen as the ‘archive of the issp’. the tasks of the zentralarchiv are to archive, check and maintain data and documentation of the country-specific studies and distribute it to the scientific community. following the structure of the questionnaire and the set of standard background variables, an international comparative data set is prepared accompanied by extensive and detailed comparative documentation. this documentation, also known as a codebook, includes all information that is necessary to analyze and interpret the data set. in addition to the methodological and technical description of the survey the codebook includes the complete questions and answer categories, the frequency-distributions broken down by countries, and also country-specific details like deviations from the agreed standard, problems with translations of indicators and the like. the first step in creating the international file is the production of a ‘standard setup’. this includes the desired structure of the integrated file, the variable-names, variable-labels, codes, value-labels and the definition of the missing values. the starting point for the production of the ‘standard setup’ is the basic questionnaire of the respective issp module and the set of standard background variables defined for the issp. even though this ‘standard setup’ is being distributed to the participating countries in advance, a number of cases have remarkable deviations from the desired standard, which must be considered in the process of merging the country-data to the integrated file prior to archiving the data. starting with 6 country data sets in 1985, by 1998 the issp had grown to include 28 different data sets, dramatically increasing the workload for preparing the merged file. the amount of work that is necessary to harmonize one country-data-set to the standard can be measured by the amount of different documents that have to be viewed and used in order to understand all the details of the countryspecific conditions. documents needed for processing: * issp basic questionnaire * standard setup * standard background variable * codebooks of earlier issp modules * frequencies of original data * original questionnaire ssecorpgnigrem putesdradnats ssps 1yrtnuoc putes 1yrtnuoc atad 2yrtnuoc putes 2yrtnuoc atad 3yrtnuoc putes 3yrtnuoc atad ���������� � ����� �� � ��� �� � � � �� �� �� �� �� �� �� �� �� �� � � � � � � � � � �� ������ 4002-5891scipot tnemnrevogfoelor* skrowtenlaicos* ytilauqenilaicos* ylimaf* snoitatneirokrow* noigiler* tnemnorivne* ytitnedilanoitan* pihsneziitic* 6991,0991,5891> 1002,6891> 9991,2991,7891> 2002,4991,8891> 7991,9891> 8991,1991> 2002,3991> 3002,5991> 4002> iassist quarterly winter 2000 15 * dictionaries * english documentation * re-coding documentation * frequencies of re-coded data * other sources: internet, isco, statistical yearbooks, etc. the predefined standard is being considered in different degrees of quality. the number of recode-statements to harmonize the country data sets to the final standard varies from 50 to 300, in one case over 1000 recode-statements were needed to merge the file. in all of these cases each recode-statement must be checked and proofed to determine whether the desired results have been attained. the time and resource consuming nature of this process lead to ideas for developing a tool that would efficiently support the harmonizing process. the instrument needed to compare the pre-defined standard with the actual country specific setup with the following processing criteria. the first comparison should be done automatically resulting in a list of all deviations. the deviations in the setup and in the dataset itself are then corrected through individual recoding. the interface for this process must be userfriendly. the system should create a detailed report of all steps in the process. issp data wizard: * maps original with standard setup * provides a comparative view * assigns variables and values t ostandard * re-codes data to standard * alows for individual data-processing * reports processing steps the concept of the issp datawizard is being developed cooperatively by the gesis (german social science infrastructure services) institutes of the central archive for empirical social research (za) in cologne and the german social science information center (iz) in bonn. the first ‘real’ application after the explicit test-phase will be done with the 2000 issp module on ‘environment ii’. in a further stage of development the issp datawizard will be prepared for distributed use so that the issp participants can prepare their data-sets at their home institutions. thus the expectation is high that the quality of the data delivered to the archive will be significantly higher according to the pre-defined standard, thereby facilitating the final steps of merging the files to a common international data-set. implementation of the issp datawizard the issp datawizard is implemented at the german social science information center (iz) as a java/swing application.2 the use of a platform independent application is an advantage, particularly when distributing the tool to issp project partners in different countries. although the data are stored in an oracle database, the tool can be used just as readily on any relational jdbc-capable database server. great importance is attached to providing an ergonomic user interface in order to provide users with a sound level of support, focus their attention on the main aspects of their work and not burden them with the unimportant areas or information. the wob model («tool metaphor based strictly object orientated graphic direct manipulative user interface») is a set of software ergonomic proposals which, in their entirety, were designed to create efficient and «natural» user interfaces (krause 1995). the development of the issp datawizard was orientated on this model. based on the analyses of current work processes carried out cooperatively by za and iz, an attempt has been made to optimize support for the steps involved in processing the issp modules. in this context, the wizard manages modules covering the associated default and national setups, country-specific study descriptions and survey records. (see fig 1 : managing study descriptions with the issp datawizard) an import and export interface provides the capability of reading and writing spss setup and data files. the open xml standard of the data documentation initiative3 is also supported. this permits the use of a flexible and maker-independent data format and also enables a straightforward exchange of data with other tools supporting this standard. the setup that defines variables and values with codes, labels and comments thus matching the questionnaire can be viewed, edited or even re-created using the datawizard. it is also possible to alter the sequence of variables and values. this can be done both for the default questionnaire of a module as well as for country-specific questionnaires that implement this standard. in addition to making entries in the editor, users can also adopt variables and values from other modules, e.g. in the case of permanent demographic variables, where questions are asked that pertain to a module with a subject area covered in a preceding module, or in conjunction with values on scales that are used for several variables. it is also conceivable for partners to copy the entire default setup and make the requisite country-specific changes to the copied version so as to reduce the number of transfer errors, such as scale reversals, etc. (see fig 2 : managing setups with the issp datawizard) the central function of the issp datawizard is to compare country-specific setups with the default setup. to this end, the work carried out ‘intellectually’ at the za has been 16 iassist quarterly winter 2000 analyzed and – in the simple cases – translated into rules wherever possible. an automatic mapping process, for instance, identifies reversed scales, incorrect variable codes or similar problems and marks them as errors. variables and values capable of being assigned beyond doubt are marked «ok», and dubious assignments highlighted for intellectual investigation. in the process of intellectual post-processing, which cannot be avoided, these problem cases are highlighted (e.g. using colour markings and an error browser) whereas unequivocal assignments require no further attention. the editor for post-processing compares a country-specific setup with its allocated default setup and synchronizes the display in order to quickly provide a clear view of the information relevant to a variable. assignments can be made from context-related selection lists, thus reducing the likelihood of making incorrect entries and easing the demands on the user’s attentiveness. this way, it is also possible to adopt values into the default setup that have been «overlooked» for a particular country without having to switch to the part for processing the default setup.(see fig 3 : post-editing setup assignments using the issp datawizard) the database archives both the reference setups of issp partners as well as the edited versions. this means it is possible at any time to access the data submitted and check the adaptations made. in addition, it is possible to document all checks and assignments in a text file that additionally contains a complete concordance between a countryspecific setup and the relevant default setup. this creates transparency in terms of merging and mapping; any suspected error can be reliably investigated. adaptations to the setups (to the structure of the data collected) must also be applied to the data itself. divergent value codes, reversed scales, etc. must – as adapted in the setup – also be corrected in the data. the fig 1.: managing study descriptions with the issp datawizard fig 2.:managing setups with the issp datawizard iassist quarterly winter 2000 17 necessary re-coding of data takes place automatically on the basis of the mapped setup. here too, the database archives the data originally submitted as well as edited version. it is also possible to employ customized rules in order, for example, to examine the consistency of filter queries or to re-code country-specific indices to a default index, provided this is possible without intellectual scrutiny. users are provided with complex tools for creating and applying these rules. (see fig 4 : simple data visualization with the issp datawizard) finally, users have the capability of viewing data in straightforward counts that allow them to carry out simple checks, e.g. what age distributions to expect, etc. here, it is also possible to compare the reference data version with the edited version and thus identify the result of the mapping and re-coding process. however, these capabilities only serve to monitor the success of merging; the actual data analysis is still to be carried out using the normal statistics program packages. whereas the current version of the issp datawizard was primarily developed to support the za in merging international data, future versions will have to be developed towards the work distribution between za and issp project partners. it will be necessary to resolve questions of data exchange, access rights etc. conclusion the issp datawizard represents a tool that reduces the effort involved in merging issp records as well as the possibility of errors by automating simple activities and offering support to users in the necessary process of intellectual post-editing. once the tool has proven its worth in practice, it must be further improved and optimized towards the requirements of the applications for which it is used. a general extension of the methods used beyond the issp context for similar application areas cannot be ruled out. a conceivable area for future development is the addition fig 3.post-editing setup assignments using the issp datawizard fig 4.simple data visualization with the issp datawizard 18 iassist quarterly winter 2000 of further elements of artificial intelligence to the issp datawizard that do not focus solely on the data structure of default setup. for instance, plausibility rules could be introduced that contain the anticipated spread of characteristics and provide warnings where countryspecific data stray beyond the given confidence intervals. this way, the user would be drawn to critical points and could consider taking a closer look at the data and documents. developments of this type will need to be the subject of further concept formulation. literature: j.w. becker, james a. davis, peter ester, and peter p. mohler, eds. (1990), attitudes to inequality and the role of government. rijswijk, the netherlands: sociaal en cultureel planbureau petra beckmann, peter ph. mohler, rolf uher (1991), issp. international social survey programme – basic information on the issp data collection 1985-1994, zuma-arbeitsbericht, nr. 15 alan frizzell and jon h. pammett, eds. (1996), social inequality in canada. ottawa: carleton university press alan frizzell and jon h. pammett, eds. (1997), shades of green. ottawa: carleton university press roger jowell, sharon witherspoon, and lindsay brook, eds. (1989), british social attitudes: special international report. aldershot: gower roger jowell, lindsay brook, and lizanne dowds, eds. (1993), international social attitudes: the 10th bsa report. aldershot: dartmouth publishing jürgen krause (1995), das wob-modell, (izarbeitsbericht 1) bonn: informationszentrum sozialwissenschaften n. to_, p.ph. mohler and brina malnar, eds. (1999), modern society and values, a comparative analysis based on issp project. fss, university of ljubljana und zuma mannheim rolf uher, irene müller (1988), the international social survey programme – issp, in: iassist quarterly, vol. 12, no. 4, s. 3 ff rolf uher (2000), the international social survey programme (issp), in: gert g. wagner et. al. (eds.) schmollers jahrbuch, zeitschrift für wirtschaftsund sozialwissenschaften, 120. jahrgang, heft 4, berlin 2000, s. 663-672 footnotes 1 since 1997 the zentralarchiv is supported in its work by the spanish issp partner1 in addition to the authors, siegfried schomisch, udo riege and max stempfhuber are involved in developing this tool. 2 http://www.icpsr.umich.edu/ddi/ * paper presented at the iassist/ifdo conference, amsterdam 2001. robert strötgen, german social science information centre (iz) bonn, and rolf uher, central archive for empirical social science research (za) at the university of cologne. http://www.icpsr.umich.edu/ddi/ rights of researchers and governments to national records the articles which follow are drawn from the papers presented at the 1982 annual conference and focus on the rights of researchers and governments to national records. both present an overview of the policies and legislation which determines ownership and utilization of information in the united states and great britain. this is the first of a two-part issue on the subject. the fall 1982 newsletter will be devoted to a follow-up article on the u.s. and to contributions from germany and sweden. who owns contract and grant data and who can use it?: a look at the u.s.a. by thomas elton brown national archives and records service united states of america the opinions expressed in this article are solely those of the author and in no way reflect the official position of the national archives and records service. information management policies for machine-readable data include two fundamental questions. the first concerns the disposition of the information. this can involve long-term retention of the data by its creator, destruction whether willful or inadvertent, or transfer to another organization or individual. the decision on the disposition is the responsibility of the person or organization who legally owns the information. the second area of concern is who has access to the information and under what conditions. obviously, the disposition can affect the access. if the data is destroyed, no one has access. since different organizations and individuals can differ widely on access procedures, legal and physical custody can determine whether data is available. however these issues are addressed, the main goal of the information management policies should be the best and most efficient use of the information. in the united states, the ownership and disposition of materials in the hands of federal agencies are regulated by the records disposal act. this legislation defined records as "all books, papers, maps, photographs, machine readable materials or other documentary materials, regardless of physical form or characteristics, made or received by an agency of the united states government under federal law or in connection with the transaction of public business. . ."^ under this act, the agency may not destroy or otherwise dispose of those records without the approval of the national archives and records service. for those computer records which the archives appraises as having continuing value, agencies are required to transfer them to the national archives as soon as they become inactive.^ access to government information is controlled by the country's freedom of information act (foia) as amended in 1974. it requires the prompt release of "agency records" unless they fall under one of nine exemptions. if one of these exemptions applies, the agency must release "any segregable portion of the record". however congress failed to degine "agency records" in the foia. 3 most federal officials assumed that the definition in the records disposal act applied to the foia. in 1978, however, the united states court of appeals for the district of columbia ruled otherwise in poland and skidmore v. central intelligence agency et al . the appeals court pointed out that congress "had ample opportunities" to refer to the records disposal act's definition, but had not done so. since the ruling did not provide an alternative definition, it suggested that the meaning of "agency records" would be decided on the individual facts of each case, 4 two years later, the supreme court in forsham v. harris moved toward using the records disposal act's definition for foia purposes. the court noted that both the records disposal act and another related statute associated the creation or acquisition of materials with the concept of the status of "records" and concluded that this association had significance "in this case". before drawing this conclusion, the court warned "these definitions are not dispositive of the proper interpretation of the congressional use of the word [records] in the f0ia."5 to discuss the impact of the records disposal act and the foia on information created under government grants and contracts, one must distinguish between the two. before 1977, government agencies often used grants and contracts interchangeably for administrative convenience to get work done. in that year, a new law required agencies to discriminate between the two forms of federal funding. grants are intended to support a private organization or individual whose functions have a public or general purpose. in contrast, a contract is the result of a procurement process through which the government buys something for its own use. this can include the purchase of information or services.^ in 1978, the united states house of representatives committee investigated the ownership, maintenance, access, and disposition of information produced under united states government grants and contracts. the committee reported that the national government had no consistent policy or guidelines concerning such data.' for example, the committee asked the executive departments for their policies on ownership, use, and disposition of data assembled by contractors. the response ranged from the department of commerce claiming ownership and the right to control distribution to the department of health, education, and welfare vesting ownership in the contractor of all information including that specified for delivery. ° since this congressional investigation, the supreme court clarified some of the issues regarding grant records. in the previously mentioned forsham decision, a group of researchers had sued under the foia to gain access to raw data in the hands of a grantee. the court denied access: congress undoubtedly sought to expand public rights of access to government information when it enacted the freedom of information act, but that expansion was a finite one. congress limited access to "agency records". . . with due regard for the policies and language of the foia, we conclude that data generated by a privately controlled organization, which has received grant funds from an agency (hereafter a "grantee"), but which has not at any time been obtained by an agency, are not "acency records" accessible under the foia. this decision rested on the fact that the granting agency, the national institutes of health, consistently maintained that the records were not government property and had never received a copy of the raw data with the final reports. after equating creation or acquisition as the "threshold" for records status in this case, the court determined that a private organization had made the records and that no federal agency had ever received them.^ records retained by a grantee seemingly are beyond the scope of the records disposal act for disposition and the foia for access. in this situation, the grantee has almost total control over access and disposition of the information, subject only to the specific provisions of the grant. the implication in this reasoning is that whatever information an agency does receive from a grantee is an "agency record" under both statutes. in light of this decision, researchers may be able to obtain data from the grantee in two fashions. first, some grants contain clauses which give the government agency the right to access the data assembled by the grantee. if the agency exercises this right, then seemingly the disposition and access questions would be governed by the statutes and not the grantee. secondly, several granting agencies concerned with supporting research specify that the data be made available to other researchers. for example, the national endowment for the humanities' guidelines for basic research proposals advise: please provide evidence of other scholar' readiness to make use of your data, if you anticipate such use, and discuss your or your institution's plans to make the data available to other researchers .^'^ in these cases, the grantee appears to retain the authority for determining access and disposition. while forsham clarified some aspects concerning grant data, the decision noted the congressional distinction between grants and contracts. thus the questions about contract data remain unanswered. a clear example of this is to look at the table of contents in the federal procurement regulations in the code of federal regulations . in this thousand-plus-page volume which outlines the regulations which civilian agencies must follow, not one word has been included in the section reserved for "data". 11 instead, each agency has developed its own internal guidelines with varying approaches and degrees of specificity for use in each contract. the department of agriculture candidly reported, "most of the contracts awarded by this department do not contain clauses specifying who owns the data, how it can be used, and the ultimate disposition of the data. "12 whatever the agencies general guidelines are, they are put into the specific clauses of the contract. in almost all contracts which provide services or information to the government, "rights in data" clauses define the mutual rights of the government and the contractor to the information. there are three basic approaches: 1.) all data delivered under the contract is acquired with limited rights; 2.) all data is acquired with unlimited rights, and 3.) specified data is acquired with unlimited rights. implied in the rights-i n-data clauses is the authority of the government to order delivery of the data. a recent development in such rights-in-data clauses is the "deferred ordering and delivery of data" provision. several agencies-department of defense, national aeronautics and space administration, department of education, and department of housing and urban devel opment--are using standard clauses to require delivery of the information created during the contract for two or three years after the termination of the contract. this right to order data extends to the entire federal government, not just the contracting agency.'-^ if a federal agency orders and receives the data from a contractor, then the information is seemingly subject to the records disposal act. this is the position of the national science foundation regarding its research centers operated by contractors. if a center transfers any material to the foundation, the records become government property and subject to the government's records management policies regarding disposi tion .^4 while the foia would probably control access to the delivered data as well, this is less certain. are the records which the contractor retains subject to the records disposal act? to rephrase the question, is a government agency "making" the records when it awards a contract to an organization or individual for the contractor to perform a service or gather information? the question is unresolved. possibly the key to this question is whether the contractor is performing a function which the congress has specifically mandated the agency to perform. this position may receive support in the section of the records disposal act which requires: the head of each federal agency shall make and preserve records containing adequate documentation and proper documentation of the organization, functions, policies, decisions, procedures, and essential transactions of the agency. . .^5 since the questions of disposal and access are separate, does the foia apply to the records retained by the contractor? this too is unanswered. the courts have been reluctant to apply the foia to contractor records. but these cases generally have concerned the housekeeping material incidental to the administration of the contract and not to the raw data collected for the government. these legal questions about grant and contract data will become even more critical as federal agencies increasingly rely on private organizations to perform federal functions. the answers to these questions will define the information management issues about the contract and grant information. as the house committee stated in its previously cited report, "the point of a disposition policy is not for the government to acquire all data from a federal grant or contract, but for the best use to be made of the data." in searching for the key to this "best use", the committee concluded that "no single information management provision would be suitable for all federal contracts or grants." and that "different types of information may require different types of management." indeed this committee hoped that its report would produce discussions and more understanding about grant and contract data.^^ hopefully, those interested in the secondary analysis of data either as users or suppliers will join the discussion to clarify the issues and to find suitable disposition and access policies. references ^44 u.s.c. 3301 . 241 c.f.r. 101-11 .411-9. 35 u.s.c. 552. '^trudy huskamp peterson, "after five years: an assessment of the amended u.s. freedom of information act," the american archivist (spring 1980), pp. 165-166. ^ forsham v. harris , supreme court of the united states, no. 78-1118, march 3, 1980. hereafter cited as forsham . ^federal grant and cooperative agreement act of 1977, public law number 95-224, 92 stat. 3. 'u.s. house of representatives, committee on government operations, information policy issues relating to contractor data and administrative markings , hearings, september 18, 1978, 55th congress, 2nd session, pp. 84-223. hereafter cited as hearings . ^ forsham . tounited states national endowment for the humanities, "the general research program," (no date), p. 7. lul c.f.r. 1-1 . ^^ hearings , p. 88. ^^house report , pp. 13-14. john e. kirsch to g. n. scaboo, april 29, 1981. while granting that materials transferred from the centers to the agency are subject to government records disposition, the national science foundation maintains that all data not transferred to the agency is the porperty of the contractor. a copy of this letter is available from the author. ^^44 u.s.c. 3101 . ^^ house report , pp. 3-4, 18, 22. 20 iassist quarterly 2016 iassist quarterly iassist quarterlyiassist quarterly abstract new york university (nyu) libraries has an extremely high-volume chat reference service. this popularity presents a unique opportunity for gaining insight into library patrons’ conceptualizations of their data reference needs and how these needs are changing. through analysis of three years’ worth of chat transcripts, we began to explore user needs and familiarity related to locating secondary data and statistics, performing data analysis, and using existing data services. ultimately, we focused our analysis on requests for census data. this article discusses, in detail, the methods, preliminary results, limitations, and proposed next steps of our investigation. our final goal is to contribute to the growing body of knowledge about how information needs are conceptualized and articulated, and how this knowledge can be used to improve data reference in an academic library setting. keywords: academic libraries, data reference, grounded theory, virtual reference services, chat transcripts introduction nyu libraries serves the nyu ‘global network university’, the main campus of which is situated in greenwich village, next to washington square park, in lower manhattan. the nyu polytechnic school of engineering is housed nearby, in downtown brooklyn, and nyu has portal campuses in abu dhabi and shanghai, as well as 11 smaller global academic centers where students study away for a semester or year. nyu enrolls approximately 45,000 students (half of whom are undergraduate students), and employs approximately 3,000 teaching faculty. bobst library is the flagship of the nyu libraries’ system, with 12 publicly accessible floors, 6 million volumes, and seating for 3,000. the library’s urban location and proportionately small seating capacity, combined with the area’s above-average commute time and a user community spanning the globe, lead to high demand for nyu libraries’ virtual library services. our chat reference service is extremely busy; we receive approximately 15,000 chat transactions annually, 30-40 a day on average, mostly occurring between the hours of 9am and midnight, new york city local time. the average duration of a chat conversation is 16 minutes. this popularity offers a unique opportunity for gaining insight into library users’ conceptualizations of their data needs and how these needs are changing. through analyzing three years’ worth of chat reference transcripts, we began to explore user needs and familiarity related to locating secondary data and statistics, performing data analysis, and using existing data services, focusing on the way patrons initially ask data questions. while existing scholarship has addressed the theory and practice of data reference (gerhan, 1999; kellam and peter, 2011), very little empirical research to date has qualitatively explored users’ articulations of their data needs (wang, 2013). this project is unique in that it employs transcripts of actual reference transactions, as opposed to user understanding academic patrons’ data needs through virtual reference transcripts: preliminary findings from new york university libraries by margaret smith1, jill conte2, samantha guss3 ... patrons and librarians often use the word ‘data’ casually when discussing databases or information in general. iassist quarterly 2016 21 iassist quarterly surveys (read, 2007), as the basis for analysis. furthermore, such a high-volume chat reference service, which is staffed by data specialists and non-specialists alike, offers an opportunity to assess how the service as a whole handles–and can better handle–data reference. research method: grounded theory because little research to date has been done on how users conceptualize and articulate their data needs, we chose a grounded theory approach, which is an exploratory, iterative methodology. this inductive approach seemed well suited for our purposes, as we did not start out with any particular hypothesis or hypotheses, but we knew that we had a rich data set. in grounded theory, researchers constantly ‘move back and forth’ between data collection and analysis (bryant and charmaz, 2007), resulting coincidentally in data refinement and conceptual categorization that leads to increasingly theoretical insight (payne and payne, 2004; bryant and charmaz, 2007). on our first pass at analyzing the chat transcripts, we used the process of ‘open coding’ and ‘memoing’ (grounded theory institute, 2014) to look for common patterns and to recognize and establish emerging themes. from there, we developed nascent codes and descriptors to start categorizing the data; codes were applied to relevant portions or passages of transcripts, while descriptors were applied to entire transcripts. we used the process of ‘constant comparison’ (grounded theory institute, 2014) to scrutinize and further develop codes and descriptors as we applied them. during this initial phase, we communicated on a regular basis through memos and real-time meetings to discuss observations, to deliberate over the shape of the emerging coding/descriptor schema, and to consider strategies that would better focus the data set. this iterative process of collaborative inquiry–i.e., observation, analysis, deliberation, and refinement–likewise marked each subsequent phase of our investigation, as the data collection and analysis processes described below demonstrate. while we remain in the exploratory stage of our investigation, using a grounded theory approach will allow us over time to move from coding, categorizing, and comparing concepts to building an overarching theory that we can then marry with existing literature on the topic (grounded theory institute, 2014). data collection and analysis due to the iterative nature of grounded theory, most of our data collection and analysis processes were inextricably entwined. initially, we collected three years’ worth of chat reference transcripts, as text files, from libraryh3lp, our chat service provider. we then used two main tools to compile our data: filelocator pro, to retrieve transcripts containing data-related keywords, and textcrawler, to remove system-generated librarian identifiers. to analyze these transcripts, we used dedoose, a web-based application developed to perform mixed-methods analyses in the social sciences. dedoose allowed us to categorize each transcript using controlled descriptors–for example, to indicate whether a transcript should be included or excluded from a sample, and also to apply qualitative codes to excerpts of text within the transcripts. these descriptors and codes could then be cross-tabulated, analyzed, and visualized in various ways. in combination, these tools–filelocator pro, textcrawler, and dedoose–were extremely effective for selecting a sample of transcripts; for protecting the privacy of individuals involved; and for classifying and analyzing the transcripts within a sample. the process of gathering data-related reference transactions, however, was a non-trivial task. even generating a starting search strategy required careful consideration of disambiguation. for example, we quickly realized that a search for the phrase, ‘number of’, would also retrieve results where a librarian or patron mentions the call number of a book. after a few minor tweaks to minimize these mismatches, our search strategy settled on this: data or statistics or stats or gdp or demographics or census or mortality or gis or quantitative or numeric or spss or atlas.ti or atlas or nvivo or qualitative or vivo or “data services” or data.services@nyu.edu or “data service studio” or “data.service@nyu.edu” or stata this limited the number of transcripts substantially, but still retrieved an immense number of transcripts that were not datarelated. for example, patrons and librarians often use the word ‘data’ casually when discussing databases or information in general. additionally, there were quite a few hits where the patron was asking for help locating or accessing a book or article that had one or more of our search terms in its title, yet the resource itself was not data-related (e.g., a quantitative study related to nursing). there were also cases where the physical space of our data services department was referenced, but not in regard to data needs (e.g., complaints of an unruly patron or broken computer in that area). in order to ensure that the sample contained as many data-related results as possible, we read through the transcripts, looking for actual relevance to data, and assigned an inclusion or exclusion descriptor to each one. even so, we ended up with 950 datarelated transcripts from just one year’s worth of transcripts. so we further refined our inclusion/exclusion criteria to omit those data-related transcripts involving ‘known item’ questions, such as a patron asking for help locating a specific financial report that contained data they had found via google. while sometimes these patrons seemed clearly interested in the data that the report contained, it was often difficult to say whether this was definitively the case, or whether they were more interested in the report as a whole. we applied these new descriptors to the sample. at this point, 633 transcripts remained, a large proportion of which still involved questions about specific databases for business and financial information. at a loss for ideas of other wide-sweeping exclusions we could make, we made a first pass at creating descriptors and codes for the transcripts in this sample. we read through them separately, coming up with lists of descriptors/codes that seemed potentially relevant, such as which specific resources were mentioned, the general subject area of the query, how accurate the librarian’s response was (on a numeric scale), and how satisfied the patron seemed (on a numeric scale). we then discussed our experiences as a group and quickly realized the overwhelming effort that would be necessary to apply multiple, quantitative descriptors to a sample of this size. we decided to drop nearly all of the descriptors, and instead, apply codes within the text of each transcript, indicating the presence of different characteristics, like ‘inaccurate answer’ or ‘patron satisfaction’. this was a speedier process, and we were able to make better progress in creating, discussing, and assigning codes. 22 iassist quarterly 2016 iassist quarterly although we were now making more progress, we discovered that the sample did not include as many juicy, in-depth data reference questions that we had hoped to explore. after a few more coderefining group discussions, we introduced a new code that indicated simply which transcripts were compelling. we focused on these transcripts, looking for patterns that might help us come up with a new iteration of our search strategy. in doing this, we were surprised by how many reference questions we received that were explicitly related to united states and international census data, and, conveniently, it seemed like these questions tended to be the more in-depth exchanges that we were after. we completely revised our search strategy, so that it included the terms that were frequently used in these interactions: census or factfinder or “social explorer” or “american community survey” or “fact finder”4 this strategy retrieved 147 results across all three years of transcripts, although, of course, there are some caveats to the ‘meaningfulness’ of this search. for example, it only captures use of the word ‘census’, so sometimes questions are included which merely involve the concept of a census or patrons may ask for known items, other than censuses, that happen to have the word ‘census’ in the title. it also relies on user and librarian understanding of when to consult a census: sometimes the user is wrong, sometimes the librarian is wrong, and our sample includes both of these cases. furthermore, this strategy omits censusrelated questions where the patron’s information need was not sufficiently explored or understood, such that a census would have been an appropriate suggestion on the part of the librarian, but the transaction never got that far. for each transcript in this new sample, we started by examining only the patron’s opening question, unnegotiated in any way by the librarian. we made observations about more easily categorizable and quantifiable aspects, like what time period was requested, as well as more qualitative, nuanced observations on the phrasing used by the patron. as before, we separately compiled lists of our observations; these ended up being extremely similar. where there was no difference in what was observed, we created a corresponding code. where disparity occurred, we discussed potential options and implications until consensus was achieved. we then applied this coding scheme to the transcripts. we were interested in exploring further the qualitative aspect of the users’ questions, potentially using this to develop theories about how the users conceptualized data. in consulting the library and information science literature for other studies on how users formulate information requests, we came across an article that examined reference questions submitted to archives staff via email (duff and johnson, 2001). we expanded the scope of our coding beyond the patron’s initial statement of need, categorizing the overall kinds of information given and wanted by the patron, as duff and johnson had done. preliminary findings below is a quantitative and qualitative snapshot of some of the observations and themes we have been able to extract from the data thus far using the iterative processes of coding and categorization. general observations not all patrons asked for ‘data’ in the data reference questions we identified. in fact, users invoked various terms to describe their data needs. figure 1 breaks down the frequency of language that patrons used to communicate their need for data. roughly one quarter of users did ask explicitly for ‘data’. another quarter of users used alternative language that implied that they were looking for quantitative or numeric information, while a third quarter asked for either ‘information’, ‘statistics’, or ‘stats’. the remaining quarter of users asked for specific publications types that possibly contained data, e.g., journal articles, research reports, or books. some patrons were very specific about temporal and geographic aspects of their data needs, while others were not. in some cases, this information was freely given in their opening statements; in others, such details emerged through a reference interview. overall, 49% of users voiced data needs that included a specific time period; of those, 38% sought historical data or data from a range of years, while 9% sought the ‘most recent’ data available. figure 1 words initially used by patrons to describe their data needs. iassist quarterly 2016 23 iassist quarterly in contrast, only 4% of users indicated a specific time scale (e.g., annual, decadal). 82% of users asked for data from a specific geographic location; of those, 68% sought united states data and 27% sought new york city data. 79% of users described data needs that included a particular geographic scale; of those, 32% sought city-level data, 17% sought country-level data, and 12% sought neighborhood-level data. in many cases, it was difficult to know exactly which geographic scale a patron actually needed unless it was expressed at the most granular level. for example, a user asking for new york city data may have actually needed data on harlem (a neighborhood within new york city), which they may have thought–correctly or incorrectly–would be findable in the city-level data set. the nature of patrons’ data needs also varied across subject area, as figure 2 demonstrates. nearly one third of all the data queries we identified were in reference to demographic data, while roughly one fifth were in relation to business, industry, and marketing data. together, demographic and business data reference questions constituted the bulk of our data set. lastly, 36% of the transcripts we identified showed ‘referral activity’. this means that they had been transferred between different librarians within nyu’s libraryh3lp system, that the librarian had consulted with another librarian during the course of the chat, or that the librarian had given the user another librarian’s contact information for follow-up. this suggests the collaborative nature figure 2 subject breakdown of expressed data need. of data reference as well as demand for specialized data and/or subject expertise in our sample. emerging themes data analysis is still ongoing, but a number of themes have emerged that are worth further exploration. although there are many interesting themes related to patrons’ question topics, librarian responses, and general characteristics of the interactions, the ones described below focus on patron behavior, and specifically on how patrons pose their initial questions to the librarian. the easiest/fastest way the first theme describes when a patron specifies that they are not only looking for data or statistics, but specifically for a faster or more efficient way than they can devise on their own. several examples appear below: patron: i’m wondering what is the most efficient way to find ny census data from 18401940...i just need general numbers/ demographics __ patron: hi i’m trying to figure out how many italians immigrated to the us at the end of the 19th-early 20th century patron: is there an easy way to find this? __ patron: hello, i am trying to locate health statistics for the borough of brooklyn from the census. can you suggest a link? the census is a bit convoluted and i am a bit rushed. by asking the question in this manner, the patron could be implying that they believe they have the ability to find what they are looking for if only they had enough time to do it. along the same lines, they could also be phrasing their question this way to ‘save face’–that is, to make it seem to the librarian like they are more confident about their searching abilities than they really are. the patron could also be admitting that they know that what they need is likely to exist, but know that they lack the skills to find it. 24 iassist quarterly 2016 iassist quarterly ask (for) a librarian instead of asking for help finding data, several patrons instead asked directly for a person who might know the answer to their question. for example, patron: hellois there someone who is great with using the census website? __ patron: hi where would i find someone who knows about gov docs? __ the patrons who asked their questions this way showed a fairly sophisticated understanding of the library’s reference service; that is, they understood the concept of specialist librarians, that many data and statistics questions go beyond the realm of general reference, and that there are librarians on staff who specialize in data and statistics areas. of course, it is difficult to know the patron’s true mindset in phrasing a question like this, but it could be read as either benevolent (indicating to the general reference librarian that it is ok if they do not know how to help with a very specialized question) or impatient (immediately asking for a specialist knowing that communicating with the generalist may not be a good use of time). ‘am i in the right place?’ on the other hand, many patrons began their conversation with the librarian by admitting their inexperience with the reference service model in asking a first question about whether or not the librarian might be able to help them, or verifying what they might expect to receive from the librarian. here are a few examples of this: patron: i’m looking for information regarding united states annual steel production as far back as possible, to present librarian: ok patron: would you be able to help me find that info? perhaps recommend some material __ patron: hey i have to find some figures on topics based on cities, if i were to tell you some of these topics do you think you can give me a hunch on where to start or which databases would be helpful? __ patron: i have a question about citing us census data? patron: (i’m not sure if that’s something you could help me with) interestingly, this patron could potentially have the same spectrum of intentions as the savvier patron who asked for a librarian above. by expressing doubt about whether the librarian can help, they again make it ok for the librarian to say they do not know how to help (or to give basic help or make a referral) and it potentially saves time by making sure they are asking the question to the right person in the right place. authority another common theme arose, relating to the authority and reliability of sources the patron had already found–a theme that will be unsurprising to anyone who does any reference or information literacy instruction. for example: patron: hi, i need an academic source that establishes the years for all living generations. could you help me find a reputable source? __ patron: hello! im looking for demographics on southern brooklyn (birth rates, sex, age population). we are not allowed to use wikipedia as a resource __ patron: i can’t seem to find what i want patron: is indexmundi.com a reliable source? __ patron: i am researching the recents stats of homelessness in nyc patron: how can i find accurate numbers? this theme suggests a more substantial knowledge gap for the patron–the lack of ability to evaluate the reliability and authority of a source–but also the wherewithal to acknowledge this gap and ask for help. it is difficult to tell from most chat transcripts whether these patrons were interested in authority for the sake of an assignment (i.e., their instructor told them they can only use authoritative sources) or for the sake of having reliable data for their own projects or needs, but it is likely that both types are represented. ‘where’ vs. ‘how’ another interesting distinction that emerged was that some patrons ask for ‘where’ to find the data they need while others ask ‘how’ to find it. for example: patron: i was trying to find demographic information from 1980 to 1990 for far rockaway, ny patron: where should i look __ patron: hi, do you know where can i find the total number of college students in specific cities versus: patron: hello, i need to find cities in us where people need to use public transportation a lot patron: do you know how can i find the data? __ patron: i want to find the revenue number of taobao.com, an ecommerce website in china patron: could you show me how to find the numbers? thank you while this could simply be a result of different manners of speaking (rather than something deliberate and worth analyzing), it could also reveal clues into the way different patrons are thinking about their data needs and questions. both patrons seem to assume that the data they need exists, but the one who asks ‘where’ also seems to believe that once they know where to look, then the process of extracting or accessing it and understanding what it means will be easy or at least doable. this patron could be a more experienced data user, or could be overestimating their abilities. the patron who asks ‘how’ is acknowledging that they do not know how to approach searching and possibly also does not know what to do iassist quarterly 2016 25 iassist quarterly once the desired data are found. looking at whether a patron asks ‘where’ or ‘how’ may also tell us something about where the patron is in the research process, for example, if they are looking for data or statistics to support an argument that they have already made, or if they are in a more exploratory stage. unanswerable finally, we will explore a broader, more complex category of patron questions that we have chosen to classify as unanswerable for some reason or another. this does not mean that the question is not legitimate or should not have been asked, only that the way that it was asked makes it impossible to answer at face value. essentially, these questions are ones that require a good reference interview on the part of the librarian, and looking closely at the original phrasing of the question gives us interesting insight into how the patron was thinking about the information need and approaching it for the sake of the librarian. there are several flavors of the unanswerable theme, which are discussed after each example, below. patron: hello i have been searching a statistic for two days and i have been unsuccessful and running out of time :( can you help me? patron: i am trying to find the uninsured rate (for healthcare) in canada and cannot for the life of me find it patron: i know canada has universal health care but i need a solid statistic within the past 5 years of those citizens that are uninsured in this case, the patron is asking for something that they admit should not logically exist: if canada has universal health care, then there should not be any uninsured canadian citizens, and therefore no statistics on the number of such citizens. even so, the patron clearly has an information need; it is reasonable to assume that they are aware of this logical fallacy, so the librarian’s job is to help clarify that need and then help fulfill it. this is, in fact, what happened over the course of this chat conversation. it could be that the patron had spent sufficient time on this project such that when they asked the question, they forgot that the librarian would not have the same context to understand what was meant by this query. the example above also hints at the patron’s challenge in operationalizing concepts into variables that are likely to exist and be available. this was observed many other times too, for example: patron: hello i’m currently working on a project about the changing face of jersey street in new brighton, staten island. how would you advise that i find out the culture of crime in that area from 1950s through now? librarian: hi there librarian: are you looking for books? articles? statistics? patron: stats please in order to find statistics on the ‘culture of crime’ in a certain area, this patron will need to decide how the concept should be defined and measured first. patron: how would i find the specific ethnic breakdown and class breakdown of east los angeles? i need information on that specific region likewise, while there exist some standardized definitions for collecting data on ethnicity (though these can and should be scrutinized), there are no similar standards for data on ‘class’ in the united states. this patron will need to clarify what they mean by ‘class breakdown’ before they can find statistics about it in east los angeles. other users asked for things that were simply unlikely to exist or be available publicly or through library databases, for example: patron: i need statistics for us tomato consumption in 1840s, 1850s...thru 1900. usda stats start in 1886 many reference librarians will read this patron’s statement as a successful search: the patron identified the correct authority most likely to have the data if they exist; however, since that authority does not have them, the answer is that those data almost certainly do not exist. there could be some additional discussion about proxy variables or other creative places to look and maybe this patron’s need could still be satisfied, but the interesting part is the difference in how the patron thought about this question versus the way a professional librarian would. most of the queries that fall into this category also raise questions about what background work the patron has already done and whether they might be better served by looking for books or articles instead of data sources. limitations and next steps we acknowledge, of course, that our approach has limitations. while chat transcripts allow us to look back at reference transactions in a way we never could with in-person reference, we also do not have any feedback about the experience from the patron or the librarian. as a result, it is very difficult to truly know what the patron really wanted, or whether the patron or librarian considered the interaction successful. furthermore, the concept of a successful interaction is complicated by the fact that user satisfaction or dissatisfaction does not necessarily equate to a correct or complete answer from the librarian. for example, is an interaction successful if the librarian determines that the desired data exist, but only in pdf format, and then the user leaves discouraged? or if the librarian gives an answer that is wrong or incomplete but the patron is happy with the answer? additionally, the set of transcripts we extracted may be incomplete, because it is difficult to identify transactions where neither the patron nor the librarian recognized a data need, which may be among the most interesting interactions. there are many additional themes in the chat transcripts in our data set; this investigation is a preliminary exploration of how patrons ask data-related questions. more themes– and their relationships to one another–will be discussed in future publications. a grounded theory approach suggests that the next phase of this project will be to begin exploring the relationships between themes and determining what this data set is a study of (grounded theory institute, 2014). from the themes already uncovered, we have several pressing questions: • are these themes specific to census-related questions? are they even specific to data-related questions? or are they more generalizable to all chat reference? • is there a relationship between any of these themes and the overall success or failure of the reference interaction? 26 iassist quarterly 2016 iassist quarterly • how do these examples fit into established models of ‘question-asking’? once we have built a theory or theories from the data, the final step will be to integrate them into the established literature and articulate how our work moves the conversation forward, possibly adding to a growing body of knowledge about the librarian’s role in supporting the data lifecycle. in addition to the theoretical advantages of understanding our users, there are practical aspects of this inquiry as well. this project gives us a rare opportunity to look closely at some of the problems our users and librarians are having with data in reference transactions and to think about how we can improve our services for the benefit of all. in better understanding the kinds of queries we receive, and the ways data needs are conceptualized and articulated, we hope to build better data research guides for our patrons and improve the training, scripts, and guides available to the librarians staffing the service. one clear way to improve service is to offer training to library staff on how to use open-ended questions during the data reference interview. as evidenced by questions classified within the ‘unanswerable’ theme, users often have an incomplete understanding of how to operationalize concepts into variables that could be found in existing data sources. training that allows staff to practice asking the kinds of open-ended questions that will help users and librarians move toward a shared understanding of what the user needs, and what exists, will translate into more effective data reference interactions. our analysis also shows that users struggle with questions related to the reliability and authority of data sources. this could be communicated efficiently through an online guide showing the who, how, and why of data creation, collection, and distribution, as well as strategies for evaluating sources. making this kind of a convenient takeaway available allows librarians to more easily seize a teaching moment, and enrich and expand the learning experience beyond the immediate data reference interaction. these guides are especially valuable because they make it easier for generalists staffing the service to convey specialized information. these are just two possible ways to improve service based on our preliminary findings. as demand for secondary data grows across academic disciplines, strengthening the data reference piece of a larger reference program that is staffed by specialists and generalists alike ensures the future health and relevance of academic reference services. references bryant, a. and charmaz, k. (2007) the sage handbook of grounded theory. [online] london: sage publications ltd. available from: http://srmo.sagepub.com/view/the-sage-handbook-of-groundedtheory/sage.xml. [accessed: 7th july 2015]. duff, w. and johnson, c. (2001) a virtual expression of need: an analysis of e-mail reference questions. the american archivist. [online] 64 (1). p.43-60. available from: http://americanarchivist.org/doi/ abs/10.17723/aarc.64.1.q711461786663p33. [accessed: 13th july 2015]. gerhan, d. r. (1999) when quantitative analysis lies behind a reference question. reference & user services quarterly. [online] 39 (2). p.166-176. available from: http://www.jstor.org/stable/20863727. [accessed: 7th july 2015]. grounded theory institute. (2014) what is grounded theory? [online] available from: http://www.groundedtheory.com/what-is-gt.aspx. [accessed: 7th july 2015]. kellam, l. and peter, k. (2011) numeric data services and sources for the general reference librarian. [print] oxford, uk: chandos. payne, g. and payne, j. (2004) key concepts in social research. [online] london: sage publications ltd. available from: http://srmo.sagepub. com/view/key-concepts-in-social-research/sage.xml. [accessed: 7th july 2015]. read, e. j. (2007) data services in academic libraries: assessing needs and promoting services. reference & user services quarterly. [online] 46 (3). p.61-75. available from: http://www.jstor.org/ stable/20864696. [accessed: 7th july 2015]. wang, m. (2013) supporting the research process through expanded library data services. program. [online] 47 (3). p.282-303. available from: http://www.emeraldinsight.com/doi/abs/10.1108/prog-042012-0010. [accessed: 7th july 2015]. notes 1. margaret smith is physical sciences librarian at new york university and can be reached by email at margaret.smith@nyu.edu. 2. jill conte is librarian for sociology, psychology, and gender & sexuality studies at new york university and can be reached by email at jill.conte@nyu.edu. 3. samantha guss is social sciences librarian (for political science, international studies, geography & the environment, government information, and data & statistics) at the university of richmond and can be reached by email at sguss@richmond.edu. she was previously data services librarian at new york university from 2009-2014. 4. american factfinder is the united states census bureau’s online tool for accessing data. the american community survey is a demographic survey that complements the united states decennial census. social explorer is a commonly used commercial database that repackages u.s. census and other data. iassist quarterly — 51 reviews reviews reviews reviews by daniel tsang, university of california, irvine. directory of statistical microcomputer software. wayne a. woodward, alan c. elliot, henry l. gray. new york and basel: marcel dekker, 1988. the second edition of this statistical software directory, some two to three years after the first, is a massive 744-page book. based on questionnaires to vendors, over 200 statistical software packages are analyzed. entries contain hardware/software requirements, ordering/price information, available docimientation and phone support, statistical feattu^es supported, graphic output, product history, and occasionally, a listing published reviews. for spss-pc+, there is a citation to a review in the august, 1986 issue of "american statistician", but no review is listed for sas. a useful feature is an appendix listing program capabilities for each program. also useful is the name of a contact person at each vendor, although inevitably, that information will become dated readily. rather surprising is the information given for the number of cunent users for the software package — from unknown to a dozen to thousands. a good source for the harder to find statistical package. cartographic and remote-sensing digital databases in the united kingdom. sarah finch and david rhind. boston spa. wetherby, west yorkshire: british library, 1987. this catalog of mrdf in cartography is pan 6 in the british library information guide series. the subjects covered span not only oceanography or cartography, but also administrative and political divisions of great britain. each entry describes the data file and lists the soiu-ce. it may also list availability of hard copy output also sometimes given is compatability with particular statistical packages, such as sas. in total 257 datasets are identified. of special interest is a short essay on "data archives and libraries," arguing that "if use of digital data and their transfer over telecommunications links become commonplace, then efficient storage of the digital records becomes essential." it notes that back in 1984, the british house of lords select committee on science and technology produced a report on remote sensing and digital mapping. among its 46 recommendations was one calling on the british librar. to preserve a retrospective archive of uk digital maps and remote sensing images. subsequently, the british government decided the british library should be an appropriate place for such an archive, and this book is a preliminary first step to survey the field. 52 — iassisl quarterly the essay concludes with a caution to be aware of "major difficulties" facing a library intent on archiving digital images: cost of acquisition, including the cost of obtaining adequate descriptions; the need for skilled staft, and the need to acquire hardware and software to store, retrieve and plot the data. a complete large-scale topographic map coverage of all of britain would produce an estimated 16 gigabytes of data. the book is a outcome of a project, begun in late 1984 and finished shortly thereafter in march, 1985, to compile an inventory of mrdf on cartography. project director was david rhind, a professor of geography at birkbeck college, university of london; the research was done by sara finch. the mrdf identified in the book are not, however, archived at the british library. in the us and canada, the book is distributed by longwood publishing group inc., 27 s. main sl, wolfeboro nh 03894-2069. unlocking the census storehouse for beginning undergraduates by william bosworth ' political science department, lehman college city university ofnew york as the industrial revolution gained momentum, the individual artisan gave way to organized, hierarchized factories. now, academic computer users seem to be evolving backwards: data processing is swiftly moving from an organized mainframe environment to increasingly powerful pc's in the hands of independent computer users, just like individual artisans. but as we become able to store, manipulate, and display even the most complex forms of data, we can easily become isolated from fellow researchers. manipulating data on one's own machine may have its advantages: for example, we can tailor data sets and programs to the specific needs of oitf students. however, when we deal with data sets as universally necessary as the us census, there is a danger of "reinventing the wheel." this paper discusses specific projects developed for beginning students in an urban setting, in the hope that others can find the work useful for their own academic projects. it is always desirable for the researchers in this field to suggest changes and improvements in their various projects. it really makes no sense for each of us to reinvent the same wheel; we may be independent artisans in our census-oriented research, but with a little cooperation we can perfect different approaches and, by sharing them, change that tiresome wheel into a complete, road-worthy vehicle. this paper concentrates on data derived from the us decennial censuses, since such data has a number of special advantages for all researchers dealing with social questions in the us. census items are uniformly labelled, so a program to access and transform them will work everywhere in the us. no data is more reliable. census data reaches down to the level of city blocks, and, transformed into percentages, can enable us to compare almost any government units with any others, large or small. the census bureau itself provides us with hundreds of cross-tabulations so we can characterize in extraordinary detail the qualities of various age, racial, and economic categories of the population. and though most census data is based on geographic units, the pums file is based on a representative sample of people for each major county. finally, census material is available on tape at least from 1960 onwards, so comparisons through time are facilitated. where tracts and blocks have remained basically the same, such comparisons can be done for very small geographic units. thus, those of us interested in getting undergraduates started in data analysis have in the census data our richest storehouse. but a storehouse with a locked door is of no use. confronted by mountains of census information coming raw from the government, each researcher is tempted to develop his own program to convert the data into usable form. here we see the sinister danger of simultaneously re-inventing the wheel. at this point, we should inform one another what works in oiu" experience so that others can profit from it. first, a note about the specific needs and resources that influence our activities here. lehman college is a public commuter college, 80% of whose students come from one county (the bronx, a borough of new york city). thus there is a built-in student interest in studying this area. the bronx is separated by water and a greenbeli from surrounding counties, so it is easy to identify and analyze through time. it is one of those northeast urban areas that have changed dramatically over the past thirty years, so demographic analysis through time is particularly rewarding. and when the bronx is seen in detail (particularly when we examine each of its 4,132 blocks) we see economic and ethnic differences that allow for many other dimensions of analysis. resources fw data analysis are available from our college and from the city university of new york as a whole (the artisan has his own workshop, but he can also whittle away in the modem factory). we have a powerful university mainframe computer, and individual faculty members can have terminals in their offices. the mainframe provides powerful statistical languages such as spss-x and sas, as well as tape storage and disk space for real-time work. university membership in the inter-university consortium for pohtical and social research (icpsr) enables us to get most census tapes without charge. as a us government depository, our library has most of the technical documentation needed to identify census items on the tapes. new york city's city planning commission has developed a mapping system for computer representation of city features down to the individual block. at the college we have a number of classrooms with networked pc's and a common elementary statistical language (abc). there is also a spring 1991 classroom with unix-based pc's connected to an rt file server. here students can display and manipulate census data directly on maps of the bronx, down to the level of city blocks. for this work we use the histwy machine program developed by prof. david miller at carnegiemellon university. the foregoing inventory of resources shows what we start with. most colleges probably have most if not all of them. many undoubtedly have other resources that facilitate the work we shall describe. if other things work better, we would like to hear about it at this point we present to you, in detail, the story of how we at lehman college unlocked the census storehouse for our beginning undergraduates. i. development of data files for 1980 census data going down to the tract and blockgroup levels there are two files: the complete count stfl and the sample stf3. only the latter has information on income and poverty, education, and occupation. using spss-x on the cuny mainframe computer, we selected the stfl and stf3 files for new york city and suburban counties, identified and labelled each of the original census variables, created new variables for "non-hispanic white" (an ethnic category which the census bureau should have developed itself), created percentages from most of the census variables, changed the order of variables to make the file look more logical (in our judgment), and finally "matched" the stfl and stf3 files to create a single master census file for the bronx and other metropolitan areas that could be analyzed in spss-x. the process involved six successive transformations of data. each of the six programs can be made available to interested researchers. for the 1980 pums file we first rectangularized the file by "nesting," changed the order of variables to make the file more logical, and recoded ages and ethnic background to simplify for easier analysis. this process involved two successive data transformations, which interested researchers may obtain. using the mainframe disks we can quickly generate crosstabulations from the bronx pums files. we have found no pc program that will do so for the largest pums, the sample based on 5% of the population (for the bronx, this sample includes over 58,000 respondents and the datafile on pc would be larger than 5 megs). we have created a pc-usable pums subset by randomly selecting onequarter of the pums cases. the file is still over one meg in size, and works best from the v disk of one of the more powerful pc's. for the tract and block group material we have created abc datafiles on our pc network, so students on their own can do the analyses we shall describe below. this information is also a base for student projects involving maps of the bronx on our unix-based pc network. students can study census items for the 64 bronx "health areas," the 356 bronx census tracts, or the 4,132 bronx city blocks. we must reiterate that all the data displays are based on the transformations we did for stfl plus stf3, and for pums, and the programs for these transformations can be made available to colleagues. n. student projects in introductory classes, a neighborhood study is the first project that introduces students to our computerized programs. students describe their bronx neighborhoods (those who are not bronx residents "adopt" a local neighborhood). they draw a map of the neighborhood, indicating its boundaries, and describing why they chose the boundaries they did. the reason for their choice is generally socio-economic rather than spatial, so the students already have generated assumptions about their neighborhood. next, students are presented with tract maps that approximate as closely as possible the neighborhoods they have described (we must be honest: in this project we try to convince students to tailor their neighborhoods to the boundaries of one or more census tracts). then each student is given a printout of 63 variables, with figures from new york state as a whole and from the bronx as a whole. these are selected from items in the computer program for each of 338 bronx census tracts. they include general population figures as well as items on employment and occupation, education, family structure, income and poverty, and housing. items for 1970 are included as well as 1980 items. students thus see on the printout certain "norms." they then predict what they will find for the census tract or tracts constituting their neighborhood. then they are shown how to use the simple "list" command in abc to retrieve the figures for their local tract, .so they see how accurate their guesses were. great differences will stimulate students to make hypotheses about demographic factors they did not consider or perhaps about change in their neighborhood since 1980. throughout this first project, students are encouraged to use their own personal experience in the neighborhood to supplement the statistics they find. the following items are examples of what the students work with (figures for 1980 unless otherwise specified): once they have done the first project, students will have mastered the abc software package and (we hope) will be interested in exploring their local area in other ways. using the same dataset just described, we next show students how to aggregate items among all bronx tracts using weighted means. we soon develop rather complex questions. for example, we can consider all the tracts (there are 55 of them) where the 1980 population was less than half the 1970 population. we can then get weighted means for these tracts to see any peculiarities. we find, for example, that in these 55 tracts that lost so 22 assist quarterly n.y.state bronx label guess for your nbrhd 1table 1 1 name blpct 13.68 31.82 % blacks in pop, 1980 blpct70 24.30 % blacks in pop, 1970 femployd 44.77 37.58 % females, 16+-, who are employed whcoll 20.40 13.89 n.hsp.white 25+: % 4 + yrs college kdipar 21.68 44.41 kids -18: % in 1-parent homes belowpov 13.09 26.98 % pop, income below poverty level vacant 5.35 4.80 % units that are vacant much population, the number over age 65 actually increased by more than a third. or we can select areas where few blacks are below poverty and compare them to areas where many are below poverty, and see the differences in family structure. we can do the same for hispanics and have a couple of dimensions for interesting speculation (at this point we must remind students that they are working with areas, not with individuals). illustrative tables are given below. those students who are particularly interested in the preceding studies are introduced to a second dataset, based on the public use microdata sample (pums file) from the 1980 census. we use the pums file for the bronx, which, when tailored for the pc, includes over 14,000 individuals, around 1.3% of the bronx population. though we cannot look at areas within a county, the pums file uses individuals as its cases, so we can define specific characteristics of a population and make comparisons through crosstabulations without fear of confusing people with census tracts. with pums, we can find unexpected differences among populations, which may well call for reconsideration of univariate table 2 procedure: datafile: bxcor78 partition: popratio it 50 number of case s passing partition: 55 number of case snot passing partition: 283 variable: oldpcrat (totalpop, % 65+: 1980 compared to 1970) weight: totalpop (total population) n total: 55 n included: 55 n weighted: 110,580 minimum code: 0.00 maximum code 274.75 num. unique codes: 44 mean: 138.902 mode: 183 median: 140. sum: 15,359,836.00 standard deviation: 61.981 variance: 3,841.698 univariate 1 table 3 procedure: datafile: bxcor78 partition: blblopov it 10 number of cases passing partition: 1 20 number of cases not passing partition: 2 1 8 variable: blmakds (black: % fams, no hsbnd, own kids) weight: blackpop (black population) n total: 111 n included: 38 n weighted: 60,138 minimum cod(;: 4 maximum code: 28 num. unique codes: 15 mean: 12.6 mode: 15 median: 13.2 sum: 756,937 standard deviation: 3.8 variance: 14.4 spring 1991 23 univariate table 4 procedure: datafile: bxcor78 partition: blblopov ge 40 number of cases passing partition: 98 number of cases not passing partition: 240 variable: blmakds (black: % fams. no hsbnd, own kids) weight: blackpop (black population) n total: 98 n included: 94 n weighted: 147,976 minimum code: 9 maximum code: 68 num. unique codes: 34 mean: 32.2 mode: 40 median: 32.0 sum: 4,758,267 standard deviation: 7.6 variance: 57.5 1 univariate table 6 procedure: datafile: bxcor78 partition: hsblopov ge 40 number of cases passing partition: 130 number of cases not passing partition: 208 variable: hsmakds (hisp: % fams, no hsbnd, own kids) weight: fflsppop (hispanic population) n total: 130 n included: 130 n weighted: 242,455 minimum code: 5 maximum code: 51 num. unique codes: 34 mean: 33.2 mode: 31 median: 32.4 sum: 8,058,798 standard deviation: 6.9 variance: 47.3 table 5 procedure: datafile: partition: univariate bxcor78 hsblopov it 10 number of cases passing partition: number of cases not passing partition: 82 256 variable: weight: n total: n included: n weighted: minimum code: maximum code: num. unique codes: 17 hsmakds (hisp: % fams, no hsbnd, own kids) hisppop (hispanic population) 82 34 16,143 4 46 mean: mode: median: sum: standard deviation: variance: 10.5 6 8.0 169,706 5.9 34.9 certain social policies. for example, if we concentrate on the three major ethnic groups in the bronx (non-hispanic whites, blacks, and hispanics), we note very significant age differences. we can further divided these groups into those who are native bom and those who are not (in the bronx, the blacks who are not native bom are in their majority jamaicans. dominicans are the largest nonnative hispanic group, while almost all native bom hispanics are puerto ricans). if we do this, we find that age differences are even more magnified: kids under 17 are twice as large a constituent of the native bom black and hispanic groups than of the non-native groups, while those over age 65 actually constitute a majority of the non native white group! the pums tables presented below show only one striking aspect of bronx population groups. with the pums file we can spend hours examining other characteristics and relationships. does income increase with education in the same way for male black heads of household as for male white heads of household? how does the specific ancestry of non-native bom whites differ from the ancestry of native bom whiles? does a larger percentage of bronx residents of albanian origin have air conditioners in their apartments than bronx residents of irish background? if age and marital status are held constant, do hispanics of dominican origin still have higher household incomes than hispanics of puerto rican wigin? do bronx residents who spend over an assist quarterty table 6 procedure: xtables datafile: bxpums partition: citizen eq number of cases passing partition: 11820 number of cases not passing partition: 2736 row: agegroup (age in 1980, categorized) column: simplrac (racial/ethnic categories) n total: 11820 n included: 11696 nh col % white black his panic total n's 4 andunder 4.5 10.5 12.4 9.3 1085 5-12 8.0 16.1 17.0 13.9 1621 13-17 6.9 11.8 13.0 10.7 1251 18-24 12.8 13.0 13.1 13.0 1518 25-29 7.3 8.3 8.7 8.1 950 30-39 10.4 14.7 13.7 13.0 1519 40-49 8.4 9.0 10.3 9.3 1086 50-64 22.5 11.3 8.5 13.8 1619 65 and up 19.3 5.3 3.2 9.0 1047 total n's 100.0 100.0 100.0 100.0 3682 3914 4100 11696 procedure: xtables datafde: bxpums partition: citizen ne number of cases passing partition: 2736 number of cases not passing partition: 11280 row: agegroup (age in 1980, categorized) column: simplrac (racial/ethnic categories) n total: 2736 n included: 2522 nh col % white black his panic total n's 4 andunder : 0.3 1.3 1.6: .9 22 5-12 2.1 5.3 5.8: 3.9 98 13-17 1.8 9.0 8.3: 5.4 137 18-24 3.4 14.1 13.0: 8.8 221 25-29 2.5 11.0 15.5: 8.1 204 30-39 8.3 19.1 19.2: 14.0 353 40^9 8.4 14.7 15.4: 11.9 300 50-64 19.9 16.2 14.2: 17.5 441 65 and up 19.3 9.4 7.0: 29.6 746 total n's 100.0 100.0 100.0 100.0 1195 702 625 2522 hour commuting to work have fewer bedrooms than bronx residents who walk to work? from the ridiculous to the sublime, one cannot predict which of these questions will stimulate an undergraduate. the accompanying maps indicate the final stage in our process of introducing undergraduates to census information. our unix-based pc network includes bronx maps showing three schts of geographic units: the 64 bronx health areas, the 356 bronx census tracts, and the 4,132 city blocks into which the bronx is divided. available data is displayed for each unit (note that income, education, and occupation figures are available only down to the census tract level; for city blocks our data shows age and ethnic divisions, family structure, and housing). the mapping system is extremely flexible. students can change the cutpoint values of an item and see the changes instandy on a new map. new variables can be created from two or more existing ones. the screen can be spht so that two maps (showing the same or different geographic units) can be displayed simultaneously. for greater detail there is a powerful zoom feature. most important, all the manipulations are shown instantly and can be printed. the history machine program, mousebased and insulating students from the horrors of unix, is also very easy to learn. we include here three maps to illustrate each geographic unit we are able to present: the health area and census tract maps have a split screen to show changes in variables through time. the block map is just too detailed to reproduce perfectly without zooming in on one part of the bronx; nonetheless, the complete map reproduced here will give an idea of what we can do with our map program. we continue to enlarge the kinds of data manipulation as well as the census information available to students. we are just beginning to incorporate the items of the 1980 stf4a file into our systems. we shall soon create new units from the existing city block map (police precincts; community school districts, for example) so that the data associated with these units can be compared visually with our census data. and, of course, we are preparing to plug in the 1990 census data as soon as it becomes available. whatever the research potential for all this material, we shall not forget that in the first instance it was designed for use in undergraduate teaching. we stand ready to share the materials we have developed with others who think like us ' presented at the lasslst 90 conference held in poughkeepsie, n.y. may 30 june 2, 1990. spring 1991 assist quarterly spring 1991 s o n s sai<<>» ips o 00 <^ 2 2 to s 5 c c 2 co co e3e3 00 m c 3:200 -"> go ii> i-» »-* i-» »-» l^i ». i<o ho zsz< "z»s o'^oc 11 > tl z 0>>mcbo°2 z assist quarterly + 7 ) -. map of the 4,132 blocks in the bronx showing percent hispanics; lightest to darkest cross-hatchej: population 0-2% 3-9% 10-3956 40-745? 75 to 94% 95% and above (blank areas contain no populadon) spring 1991 29 ¥"^2 = o z lassist quarteriy •^ c ^ s.g-b. 5" f 2; izr h^ ^ co ^ ov v» xs» co 2 ""^ s "^ "^ r" o o m f^ <=^ h'^ :i ~ > ti oo -a c/j "3 ^ p>s ^ c "a > 13 :5 o /-, ooa:o™m° a, co "o 00 ^ <: 2 2 ro h ^ > z^ts o o m>° s ^ pd ^ > tn ^ hz 00 m h a^ spring 1991 vol28-4.indd 18 iassist quarterly winter 2004 the international association for social science information service and technology (iassist) invites your participation in its 32nd annual conference entitled data in a world of networked knowledge on may 22-26, 2006 in ann arbor, michigan. the conference will be preceded by workshops and followed by optional weekend activities in the ann arbor area. details about the conference and the association may be found on the iassist website at: <http://www.iassistdata.org> proposals for papers, sessions and poster/demonstrations should be submitted by 16th january 2006. the 2006 conference theme, data in a world of networked knowledge, highlights the role of empirical data in a society that wishes not only to know itself, but also to build an enduring, interconnected storehouse of knowledge for learning and research. once again iassist offers a time and place to explore, enlighten, and energize the participation of data professionals in the networked information world. we seek submissions of papers, poster/ demonstration sessions, and panel sessions on topics that address the full range of digital data life cycle issues, including those that focus on access, documentation, dissemination, preservation, data use and current empirical research activity. additional topics might also include information and statistical literacy, data confi dentiality and statistical disclosure, geographic information systems (gis) and spatial data, as well as publication, annotation, curation and authentication of networked knowledge assets. for other key topics see previous iassist conferences at <http://www.iassistdata.org/conferences/index.html>. about iassist iassist is an international organization of professionals working in and with information technology and data services to support research and teaching in the social sciences. the organization also explores issues of access, stewardship and the interconnections among social science, behavioral, biological, and health data. typical workplaces include quantitative and qualitative data archives/libraries, statistical agencies, research centers, libraries, academic departments, government departments, and non-profi t organizations. see the iassist website at <http://www.iassistdata.org> for further information. iassist conferences bring together data professionals, data producers, and data analysts from around the world for presentations and workshops covering new and persistent issues relating to access to data, its documentation, and digital preservation, with special emphasis on the social sciences. the social sciences have a long history of data sharing activity which will make the conference of interest to colleagues in disciplines where improving data access practices is on the policy agenda, and where there are clear overlaps with digital curation, data publishing, e-science/ cyberinfrastructure initiatives, and new interdisciplinary collaborations. the iassist quarterly (iq), available online from the iassist website and in print, is another important means of communication for the data community. each year, iq features the papers associated with conference presentations. of special note is the iassist publication award, involving a cash prize for the winning paper. for further details see: <http://www.iassistdata.org/publications/pubaward.html >. the iassist outreach committee accepts applications from data professionals in countries with emerging economies for funding to attend iassist 2006. more information about the outreach committee’s work, including funding criteria and online application form, can be found at <http://www.iassist.ucdavis.edu/> procedure iassist 2006 call for papers iassist quarterly winter 2004 19 iassist 2006 the deadline for paper, session, and poster/demonstration proposals is 16th january 2006. the conference program committee will send notifi cation of the acceptance of proposals on or before 10th february 2006. individual presentation proposals and session proposals are welcome. proposals for complete sessions, typically a panel of three to four presentations within a 90-minute session, should provide information on the focus of the session, the organizer or moderator, and possible participants. the session organizer or moderator will be responsible for securing session participants, some of whom may submit paper proposals independently. all proposals, including proposed title and an abstract (recommended length 150 words), should be submitted using the link on the following website: <http://www.icpsr.umich.edu/iassist/call.html> alternatively, proposals may be sent via email to <iassist06@gmail.com>. in this case, please use a subject heading of “paper proposal your name” or “session proposal your name” replace “your name” with the name of the session organizer. further information on travel and accommodations will be available at links from the iassist ‘06 conference website: <http://www.icpsr.umich.edu/iassist/>. online registration is scheduled to open on 1st february 2006. make plans to come to ann arbor for the iassist 2006 conference may 22-26, 2006 20 iassist quarterly winter 2004 submissions due january 10, 2006 announcement of winner march 1, 2006 award: $250 us and one year membership in iassist http://www.iassistdata.org content focus: education. as an organization, iassist has a history of working to educate its members about matters of common professional interest to the social science data community. traditionally, this education has taken the form of professional development opportunities available in member-initiated and member-taught workshops at the annual iassist conference. yet as the world of social science data grows increasingly complex, staying abreast of new developments in the profession is likely to present an ever increasing challenge for iassist members. as a result, the education committee of iassist is receptive to recommendations for employing new instructional methods and technologies as we strive to meet iassist’s educational mission. at its annual conference in may, 2004, the iassist membership approved a 5-year strategic plan that focuses upon three strategic directions: education, outreach, and advocacy. this paper competition has been established as a means of exploring, articulating, and documenting topical issues related to education. for further information about the strategic plan, see: http://www.iassistdata.org/membership/plan_june2004.pdf iassist seeks papers that address one or more of the issues, principles, and strategies for engagement in the following subject areas: 1. iassist-related educational initiatives 2. professional development and educational opportunities of interest to iassist members 3. educational outreach to research communities, including and beyond the social sciences, to promote data preservation and access. papers should include recommendations for action or suggestions of specifi c projects that can be undertaken (or are underway) to further the education goals of iassist. prospective submitters may wish to review the discussion on the iassist blog about the educational issues arising under the topic of the accidental data librarian (available at http://iassistblog.org/?cat=5). criteria for evaluation: call for papers: 2006 iassist strategic plan publication award (competition is not limited to current iassist members) iassist quarterly winter 2004 21 2006 iassist strategic plan publication award --relevance of the paper to one or all of the themes in the iassist strategic plan (specifi cally strategic direction i: improve and expand the educational component of iassist both internally and externally. --inclusion of specifi c suggestions for action or specifi c projects that will further the education of iassist members or otherwise encourage progress in support of the iassist strategic plan --potential in building a base for future iassist activity --quality of writing -bibliographic content including references to related materials --clarity in presenting issues and viewpoints as outlined above competition details: all papers are to be submitted in english. the winning paper will be announced on the iassist list-serve on or around march 1, 2006. in addition to being designated as the winning paper in the iassist quarterly, the author of the winning paper will receive the monetary award and a one-year membership in iassist and be recognized at the iassist conference in ann arbor. all other submissions meeting the criteria for evaluation will be published in the iassist quarterly (iq) (online and print). papers that are submitted to this strategic plan publication award competition may also be submitted for inclusion in the iassist conference in ann arbor in may 2006. papers must be a minimum of 5 pages in length, including bibliography and graphics as appropriate. we strongly prefer that for publication purposes all documents be submitted in word format and each graphic be submitted as a separate fi le in one of the following formats: .gif .jpg .tif .bmp .png papers must not have copyright limitations; iassist quarterly (iq) rights will apply upon publication. all papers must be submitted by january 10, 2006 to the competition web host: david sheaves <sheaves@vance.irss.unc.edu> questions (not papers, please) may be sent to: <iassist-reviews@mailman.srv.ualberta.ca> competition is not limited to current iassist members. the review committee for the iassist strategic plan publication award will be announced on the iassist website www.iassistdata.org and on the iassist list serve. vol282-3.indd iassist quarterly summer/fall 2004 17 by wendy watkins1, elizabeth hamilton, ernie boyko and chuck humphrey introduction the creation of the data liberation initiative (dli)2 in 1996 opened a new channel of access to quantitative and spatial data fi les, making unprecedented amounts of statistics canada’s data available for scholarly research and teaching through an affordable annual fee. the responsibility for the collection and provision of service for the dli fell primarily upon academic librarians, the majority of whom were neophytes to the exciting world of data within an academic setting. this paper examines the training strategy developed in response to an urgent, canada-wide, need for a new program of data dissemination3. background the dli was established as a result of negotiations between statistics canada, the social science federation of canada’s data liberation working group and academic libraries in 1996. initially presented as a fi ve-year pilot project, an extremely positive evaluation during the fourth year of operation led to the transformation to an ongoing partnership program in 2001. under the dli licence, statistics canada provides participating institutions with access to all of its standard data products, which includes databases, public use microdata fi les and geography fi les.4 as part of this contract, universities agree to make the data available to members of their communities and to guarantee that the data are used only for non-commercial teaching and research purposes. an important element of the dli licence is that each institution must designate a local staff member as the dli contact who serves as the intermediary between statistics canada and his or her local institution. the dli contacts are responsible for providing local access to the collection of dli data products and for ensuring that the licence between their institutions and statistics canada is fully observed. in most cases, the dli contact is an information-professional in an academic library. the need to provide training for dli contacts was identifi ed as an early priority in the dli pilot period. without an established baseline of competencies among dli contacts across all subscribing universities, local services would vary radically across institutions. instead of being facilitators to access, dli contacts could have become a bottleneck in providing local access to dli data. the success of this initiative was highly dependent upon dli contacts developing a basic level of knowledge about the dli collection and acquiring a set of skills to disseminate data products to local patrons. the training challenge while dli data resources greatly increase the potential for empowerment and insight into canadian society, they also create new challenges for the librarians and information professionals who are confronted with the task of organising and supporting the access to this material. prior to dli, there were fewer than a dozen data libraries in the country and these were minimally staffed. for the new staff assigned to administer dli, the challenges included: quickly acquiring skills to manage this new licensed resource; understanding the collection and its use; and developing skills to aid patrons in working with these data resources. the challenges were a bit staggering. in addition to the paucity of experienced data librarians across member dli institutions, many of the dli contacts already bore multiple responsibilities in their libraries and had little time for yet another professional development venture on the job. some came to the task with self-confessed statistical literacy defi cits, and were intimidated by data and its technology. few had resources suffi cient to launch a fullservice data library immediately and were dependent on fi nding colleagues both in and outside of the library to help build a local service model. indeed, one of the training challenges was to develop a curriculum fl exible enough to provide for tremendous disparities in local environments. furthermore, the pre-existing data library community was small, stretched across a continent, and had to operate with two offi cial languages (english and french). these vast distances introduced further challenges. for one thing, training costs were a real barrier unless the program was built in such a way as to encourage those without expertise to attend training sessions. creating a national peer-to-peer training program for data librarians in canada 18 iassist quarterly summer/fall 2004 while the importance of building regional strengths was seen as part of the long-term solution to offering data services across the country, the identifi cation of a common curriculum of essential skills appropriate to the different environments of small, medium and large libraries became an initial priority. this curriculum had to recognize the different starting points for dli contacts within the training program. an initial curricular strategy was designed with these realities in mind. a basic level of data service skills (core competencies) was fi rst articulated as part of the curricular strategy. this training would be considered the entry level for staff supporting dli data and would apply to all participating institutions, regardless of institutional size. more advanced training would build upon this basic level. more recently, the curriculum has been revised to address wider issues of data literacy (see appendix a for the latest version of the curriculum.) all training was conducted from a ‘service’ perspective, that is, from a point of view focusing on the clientele of dli data. the purpose of this training was to prepare data services staff to assist their user community with dli data. this was an integral factor in the design and implementation of the training program. (see appendix b for an overview of the dli training principles and implementation process.) following the drafting of the curriculum and training principles, considerable effort was invested to ready the scene for national training of the type envisaged. the fi rst task was to build the core team of trainers for the fi rst wave of training. with only a few seasoned data librarians at the outset, additional trainers had to be recruited from the existing canadian library community and regional contacts. the maxim for this team became “as one learns, one will teach.” this principle fi ltered down to the current training delivery model. because the initial wave of national training was to be conducted in four regions of the country, a comprehensive manual was developed in both offi cial languages. this served as both a training and future reference resource. this manual was distributed in a massive three-ring binder and included everything from technical notes to in-class and take-home exercises. the trainers, new and experienced alike, grew into a supportive community as they gathered in the spring of 1997 to prepare for the delivery of four regional workshops. this community-building approach worked. the new trainers became better acquainted with the veterans, setting up a safety net for those trying to learn and teach data skills at the same time. all were encouraged and, indeed, energized by the retreat-like experience in developing data competency. following the “training the trainers” session, the team was ready to take the curriculum into the four regions of canada, from coast to coast. central locations within the four regions were selected for training and, in the case of new trainers, a seasoned trainer was on-site to infuse the process with confi dence. the seasoned trainers did not take a lead role. to do so would have undermined the role of the new trainers. but the experienced data colleagues performed trouble-shooting with the new trainers and clarifi ed issues on data fi le verifi cation processes, data parts, and elusive bits and bytes. among the factors contributing to the success of dli training, two equally important conditions applied: one was to develop trainers in the regions who became recognized by their peers as mentors; the second was to get participants to training venues. within six months of the inception of the dli training program, a core team of trainers delivered training in all four regions covering both offi cial languages, and over 80 librarians experienced one of these three-day training sessions in basic data competencies. how were such numbers achieved? travel stipends were built into the budget to cover transportation costs and, equally importantly, library directors were contacted to ensure that individuals assigned with the task of administering the dli project were granted the time to attend. the focus on four regional workshops was critical to the long-term goal of the program of establishing a data community to support the emergence of a successful data culture in canada. the on-going training program from the outset, it was recognized that a single data training experience would not suffi ce to build the type of data competencies required to sustain the dli project. the data collection itself is a dynamic one. furthermore, the technology that supports data use is in constant fl ux. as with many learning endeavours, there is typically a period of confusion at the beginning. confusion tends to be followed by clarity as basic concepts are grasped. from clarity, competence develops allowing a layering of detailed information on the fundamental concepts. finally, confi dence is achieved through practice and with the transmission of skills or information to other colleagues. the training program had to allow for this type of learning experience, as well as the knowledge that, as individual competencies were increasing, so too were local environments changing for these individuals. following the regional training sessions in 1997, the dli external advisory committee (eac) discussed mechanisms for continuing the training. the extension of the training has allowed the curriculum to include topics appropriate to the maturation of the service, such as in-class instruction to students, and working with their colleagues locally to assist patrons with using statistics canada iassist quarterly summer/fall 2004 19 products across the information continuum. one of the most striking outcomes of the regional experiences was the development of a sense of community among the participants as they struggled to understand new concepts and worked through their homework together over a meal. notable too was the continuation of contact between participants and trainers after the workshop. this sharing of expertise is in part a fulfi llment of the mentoring that was intended through the choice of regional trainers. several discussions have taken place on the substitution of alternate teaching methods to the workshops but, while web-based training can increase the knowledge base, the face-to-face workshops keep the community vibrant, informed, and cohesive. in a profession where individual participants may feel locally isolated, establishing a broader sense of a community of data professionals is key to attracting and retaining high quality staff. the eac gave priority in the dli budget to continue annual training workshops through administrative support and stipends for travel for all dli contacts to attend the workshops. another observation following the initial training “boot camp” was the importance of the involvement of the dli section in statistics canada in the training. between training sessions, they become the problem-brokers for dli contacts on the front line of library services. dli contacts have participated in a total of twenty-eight training programs (seven in each of the four regions) and in 2003 met nationally in conjunction with iassist in ottawa. they have not only been exposed to virtually all aspects of data, content, software and service, but have also developed a sense of community through their regional, national and international meetings. this community revolves around the training sessions, mentoring developed at the training sessions, and other support mechanisms such as a newsletter (the dli update), a listserv, and the development of a strong central dli section within statistics canada evaluation and future developments from the discussions on the dli listserv and from the high turnout at each regional workshop, there is strong evidence that the delivery of data services has grown substantially and that we now have aware, committed, and skilful colleagues across the country positioned to assist users with data disseminated through dli. colleagues in the areas of the country with solid data services at the outset have perhaps moved more quickly to a point where training is now being offered by a new generation of data librarians, as well as the experienced data trainers. in addition, as data services become established, library directors are increasingly advertising for people with expertise rather than adding yet another hat to an already fully-employed professional. what are the success factors of dli training? while the above discussion focussed on some of the key principles, other factors contributing to this success can be identifi ed. some of these were through design, while others were by circumstance. focus: a clearly defi ned strategy was held that involved a target audience, a common core curriculum, and a set of priorities relating to a regional focus, bilingual training, and a commitment to public service. involvement: the training came from within the community. statistical software providers did not teach sas, stata or spss. rather, this instruction was provided by colleagues in the fi eld who are sensitive to the starting points of our colleagues and who can explain the role of the software in the broader picture of data services. similarly, discussions on statistics canada products have not come from divisions remote from end users. rather, collection issues have been dealt with by those working on the reference desk who can succinctly answer the questions as to the purpose and use of data fi les associated with a product. interestingly, the defi nition of the community has an elastic nature and is expanding to include as colleagues individuals from statistics canada author divisions who understand the dli training program’s goals. relevance: the window of opportunity was limited, forcing the dli training committee to act expediently rather than with a focus on perfection. launching such a large-scale training program so quickly after the need was identifi ed meant that the direction of subsequent workshops has been toward keeping pace with change. it has given the workshops a vital focus, building upon basic skills with timely and critical knowledge for data librarians. equality and accessibility: the training is based on equitable access. each region is served on an equitable basis and the training is designed for maximum accessibility by the provision of travel support. each participant is treated as a colleague, regardless of her starting point in the curriculum or in the size of her home institution. furthermore, those trained are given the tools to teach others, both colleagues and patrons alike. renewal: the fi rst wave of training was followed by a reassessment, at which point training became part of the dli strategic plan with its own budget component. each region is open to the participation of dli contacts from outside the region and, as importantly, in providing opportunities for trainers to continue to learn from other trainers. following several years of gradual turnover in positions, the second and third “boot camps” for new dli contacts were offered in 2001 and 2004. another “train the trainer” initiative was also offered in 2004 with a focus on leadership and renewal as well as skills. 20 iassist quarterly summer/fall 2004 commitment: a commitment was made by dli to the library directors of member dli institutions to provide training for their staff. in return, the expectation was made that the directors would support their dli contacts in participating in this training. consequently, participants come from institutions with an understanding that training is integral to dli. this commitment to training has been the acceptance of responsibility at the individual and institutional level for the success of the program. after seven years of dli workshops, new challenges are emerging. libraries are entities in constant change and there has been turnover in the dli contact community. new materials and new access channels are opening up and peer-to-peer training is moving down to mainstream reference services. we are making new discoveries about statistical literacy and data service delivery that will be refl ected in changes in the content of the curriculum. and, regrettably, some of the trainers are now approaching stages where they are making life changes because of retirement or relocation. there are also new opportunities. with the core level competencies identifi ed and basic training well underway through the workshops, other learning methods and technological solutions can be added to the training program and used within the dli community. while the regional approach to training was adopted to build strengths in all areas of the country and allowed the training to address common interests of the different communities, there is a need to expand the world of talent and expertise within these regional data communities. the 2003 iassist meeting was held in ottawa and was the venue for canada’s largest dli training event ever. all dli contacts were provided with travel support to attend national and international data workshops. participants were able to choose up to four sessions from a total of twelve, ranging in level of expertise from novice to cuttingedge. since iassist is held in canada once every four years, the plan is to organize a national training event each time iassist meets in canada. this will be done in conjunction with the annual, regional workshops and will ensure that dli participants have an opportunity to meet those from outside their region and the country who share many of the same trials and triumphs that they face. conclusions data liberation began as a pilot project and was not established as an on-going program until after it underwent an offi cial independent program evaluation. the role and importance of training was clearly a key factor in the positive evaluation of the program. it is signifi cant to note that this was underscored by each group of stakeholders approached by the evaluators -contacts, managers, academics and statistics canada personnel. continued investment in training has become a given for canada’s data community. as the community grows in expertise, more and more members are able to take part in the enterprise as trainers, thus ensuring that those leading the sessions are a renewable resource. the success of dli in canada has opened a new chapter in library service and sets the stage for creating a more numerate society. while we have yet to address fully the concerns raised by dr. paul bernard, we are well on the way. concerning such issues, the public must have appropriate knowledge and not only hypothetical access to the data. paradoxically, indeed, contemporary societies offer a wealth of information, but workers and citizens can be totally mystifi ed, surrounded as they are by data whose fl ow and codes they do not master5 the challenge that we faced evolved as technology and information practices changed. the solutions were found within the community itself. the initial talent pool was carefully extended to trainers and then on to the broader community. through the use of peer-to-peer instruction, a core set of competencies has been established throughout data library services in canada. this in turn has produced new strengths as this new generation of data service providers has enriched and been enriched by the wider community through iassist. clearly the next challenge is to expand the peer-to-peer approach to the larger community of data users. iassist quarterly summer/fall 2004 21 appendix a dli curriculum knowledge skills attitudes statistics and data literacy 1. understanding the framework of statistics and data 2. understanding the continuum of access 3. understanding the uniqueness of data as a medium 4. understanding the methods of data collection 1. recognizing a data question 2. interpreting data documentation 1. overcoming anxiety 2. nurturing sharing and open access 3. nurturing preservation content 1. knowing where and how are statistics gathered (stc) 2. knowing the collection (dli) 3. knowing about other data collections 1. finding 2. accessing 3. using 4. sharing 5. re-purposing (creating new data from old) 1. advocating access to, openness to and use of statistics and data 2. being tenacious 3. being creative and bold 22 iassist quarterly summer/fall 2004 tools 1. understanding the variety of options: data, software, output fi le types, post processing 2. selecting appropriate tools 3. developing knowledge of access tools 1. using statistical packages 2. using the web tools (access tools) 3. understanding search tools 1. being open to lifelong learning (re-learning tools) 2. being curious and having the courage to attempt the new 3. being positive towards change services 1. recognizing the variety of options for service models 2. adapting a service model responsive to internal and external changes 3. understanding the user/ audience 4. being aware of funding sources 5. administering the service and dli license 1. conducting an environmental scan 2. writing grant applications 3. interpreting the dli license 1. holding a positive attitude toward service to a large community 2. being a champion for data; one who is proactive about promoting data service 3. being positive about consulting (data reference is a consultation) 4. developing a data culture appendix b dli training principles the following principles serve as guidelines for the overall training program sponsored by the data liberation initiative project. these were developed by the national training committee in december 1996 and subsequently modifi ed by the dli eac education committee in october 2003. 1. training under this program is being conducted specifi cally for (1) dli contacts at participating universities, (2) the staff who will provide services for dli data at these institutions, and (3) statistics canada staff directly involved in the support of dli. 2. training will be provided to all of those eligible under the fi rst principle through a variety of formats, including subsidized workshops that are delivered regionally. iassist quarterly summer/fall 2004 23 3. the fi rst training priority is to establish a basic level of data service skills for new dli contacts. this training shall be considered the entry level required to dli data. more advanced training will build upon previous levels. priorities for advanced levels will be determined by the needs of those supporting data services and by the evolution of dli. 4. training priorities for phase ii will address varying levels of expertise and service within the dli community. special attention will be given to strengthening expertise in regions undergoing changes due to retirement or turnover in experienced dli contacts. special attention will be given to regions without a prior tradition or culture of data use or without a previous foundation in data services. this will entail strategies that identify the training and support of key individuals who will become recognized experts in their region. 5. all training will be conducted from a ‘service’ perspective, that is, from a point of view that focuses on the clientele of dli data. the purpose of this training is to prepare data services staff to assist university clients with dli data. 6. a global curriculum plan will guide the course content that is offered through this program. the dli external advisory committee will be responsible for maintaining this plan and for periodically reviewing its content and direction. 7. training will address concerns appropriate both to small and large institutions. 8. training will be regionally based with regular national and international exposure when the opportunities arise. 9. whenever possible, trainers will be recruited from the existing canadian data library community with the expectation that those who are trained may some day be called upon to train others. this perspective operates on the principle that as one learns, one will teach. 10. outreach will be organized for library directors, the user community, stc survey managers, and other general public to communicate the importance of statistical and data literacy. notes 1 contact: wendy watkins, data centre, carleton university, ottawa, canada k1s 5b6. phone: +1 613 520-2600x8376. internet: www.carleton.ca/~ssdata. email: wwatkins@ccs.carleton.ca 2 watkins, wendy and ernie boyko (1996) “data liberation and academic freedom” government information in canada/ information gouvernementale au canada 3, no. 2 (1996). at [[http://www.usask.ca/library/gic/v3n2/watkins2/watkins2.html 3 adapted from a paper by ernie boyko, elizabeth hamilton, chuck humphrey and wendy watkins, presented to ifla’s 69th conference, berlin, august, 2003. this paper was also presented at the iassist conference held in madison wisconsin in may 2004 in the session on “developing statistical literacy: think globally, work locally.”. 4 examples of fi les included in the dli collection are census cartographic fi les at all levels of geography, postal-code conversion fi les, canadian community health survey, volunteer survey, survey of labour and income dynamics and the collection of canadian general social surveys. a complete list may be found at http://www.statcan.ca/cgi-bin/spider/dli_list. cgi 5 bernard, paul (1992) “data and knowledge: statistics canada and the research community,” society/société, may 1992, p. 22. vol29-4.indd iassist quarterly winter 2005 by by julie lamb* disseminating survey information in the networked world:a uk resource abstract survey researchers are increasingly turning to the www in an effort to find information for the data collection stage of their projects as well as for the more traditional activity of searching for literature and reports. this paper will discuss the development and use of the question bank (qb), an innovative www resource which is used to teach students and researchers about uk social surveys produced by survey agencies such as the office for national statistics and the national centre for social research. the question bank contains the full questionnaires for over 50 social surveys and is continually expanding. these questionnaires enable researchers to take questions that have been used in large scale surveys for use in their own research work, thus ensuring that they do not spend time ‘re-inventing the wheel’. the qb also contains information on social measurement in 22 substantive topic areas, and has numerous resources relating to survey data collection methods. the resource is free to all. introduction this paper is about the question bank (qb), a free uk web based resource which helps to disseminate questionnaire metadata to researchers and students’ wishing to see what has been done before. contrary to what the title of the 2006 iassist conference; ‘data in a world of networked knowledge’ suggests, there is not a huge amount of easily accessible data concerning the network or the knowledge of social science survey researchers, especially those involved with professional survey research. even fewer data exist on how such students and researchers use the knowledge that they acquire through the internet. one survey which has studied exactly how students at college use the internet has concluded that: “internet use is a staple of college students’ educational experience. they use the internet to communicate with professors and classmates, to do research, and to access library materials. for most college students the internet is a functional tool, one that has greatly changed the way they interact with others and with information as they go about their studies.” (jones et al 2002:2) the question is then, how do we as social scientists or information professionals provide students with relevant and useful information in a user friendly web interface? the question bank has attempted to do this since 1996, a time when there were very few online resources available, especially in the survey research world. one notable exception is the uk data archive which has been disseminating data since the 1960s, and has had a web presence since 1995. the qb and the uk data archive now work closely together to ensure that the future for uk survey researchers and secondary analysts is one in which a seamless interaction between survey web services is possible. in a previous edition of iassist quarterly, margaret law used the phrase ‘reduce, reuse, recycle’ (law 2005:5) and discussed issues surrounding the phrase when dealing with data. the qb also recycles, but does so with metadata, mainly questionnaires from major uk probability surveys. the idea behind the qb is that those wishing to write their own survey questions do not have to begin from scratch when there are many questions already in use and the accompanying coding frames and documentation can be easily accessed the question bank: a brief history as stated above, the question bank was developed starting in 1995 with a grant from the uk economic and social research council. the question bank was established as part of the uk centre for applied social surveys1, and sought to contribute to strengthening the quality of uk survey research and in particular to try to improve survey measurement. in addition to providing access to questionnaires, many of which were not readily accessible previously, the question bank attempts to present commentary, written by experts, on the quantitative measurement of different survey variables, as well as to provide access to the harmonisation project of the office for national statistics which seeks to standardise survey measurement across government surveys on a number of variables.2 the approach adopted by the question bank has been 6 iassist quarterly winter 2005 an inductive one, that is, to make a very large amount of material available for easy access on the www, and to invite the user to search this material using a powerful search engine on the qb site. this requires ingenuity and insight on the user’s part in order to bridge the gap between concept and variable, or between the abstract idea and the actual question. in some cases this is easier to achieve than in others. the topic commentary is intended to sensitise the user to some of the issues which can arise, but the amount of material on converting concepts into variables is not extensive. the question bank relies upon the user drawing their own conclusions about the comparability of questions and the most appropriate questions to include in the type of survey which the user is designing. one of the key problems in uk survey research is that the large scale surveys mainly funded by the government (which are used for official statistics) are conducted by survey agencies outside academia. these agencies have large staff numbers and budgets. this means that academics, especially students, are far removed from the real world of survey research. the qb has been developed with the aim of helping three main groups. firstly, researchers devising their own survey questionnaires, by providing easily-accessed illustrations of how the topics with which they are grappling have been handled/measured in professionally designed surveys. another group which the qb aims to help is secondary analysts of survey data, either at the stage at which they are seeking out surveys containing material of interest to them, or at the stage when, having worked with particular survey data sets, they wish to learn more about the underlying survey processes and their likely strengths, weaknesses and limitations. finally, we aim to help teachers and students of survey methods, by providing text and examples on the wording of questions and the construction of questionnaires and examples from large surveys already carried out. the question bank has now grown into a very large resource with over 40,000 pages of pdf questionnaire material in around 4000 files! in 2005 funding of the resource was reorganized by the economic and social research council, and it became a separate entity from the cass model, with greater funding for more staff. this will ensure that the qb continues to grow and that essential changes to the look and feel of the web site can be managed. the web site the qb is a very large site organised in a flat file structure. materials are organised around three main themes/areas of the site: surveys, topics, and resources. i will now go through each of these in turn. surveys the surveys area is the largest part of the qb and houses not only the questionnaires from large scale uk social surveys but also supplementary information about those surveys. this ensures that qb users can see what a real survey questionnaire looks like and can find out how a survey is carried out from start to finish. for example, for each survey the qb has extensive links back to the organization that carried out the survey and to technical and final survey reports that give details about the timelines for the survey, the methods used, sampling techniques, advance letters sent out, coding frames, how the data were analyzed , as well as the final results. it is therefore possible, with the questionnaires and information on the qb and the links to other resources such as the department of health and the office for national statistics, for students to run mini replica projects to really find out how survey research is carried out, and then use other resources such as nesstar3 at the economic and social data service4 to attempt to analyse the actual data from the survey at no cost. currently the qb holds information on 57 uk social survey series and is continuing to expand. this amounts to over 40,000 pages of questionnaire material, all of which is fully searchable using our advanced search engine. the surveys for which questionnaires are included in the question bank are all social surveys using a probability sample design. commercial market research surveys and business surveys directed to organisations are not generally included in the question bank. in most cases the population units that the surveys we focus on are intended to study are either individual persons, or domestic groups such as households or families. these surveys deal with a very wide range of topics which relate to the circumstances, behaviour and attitudes of these units. the qb focuses mainly on large-scale quantitative surveys. most of them have quite long and complex questionnaires that are administered in the field (or by telephone) by trained social survey interviewers, the type of survey that would be too vast for a student to replicate for their own projects. a high proportion of the in-scope surveys have been conducted either by, or for, central government departments. others are major academic surveys. many are repeated continuous or longitudinal surveys which have produced annual series of published results. another main reason for selecting the questionnaires of particular surveys for inclusion in the qb are that these surveys on a national scale are generally treated as benchmarks against which other surveys in the same topic areas can be compared. the criteria for selecting surveys as benchmarks are: • that the survey should have been professionally developed and conducted to a high technical standard; • that it should cover a national reference population; iassist quarterly winter 2005 7 • that it should be a prime current source of information on important social science topics that the qb sets out to cover. given their origin, it can be assumed that the questions reproduced in the qb have also been pilot tested. therefore they are likely, on the whole, to perform better as a means of collecting quantitative information for particular purposes than questions which someone coming fresh to a survey topic, without previous question drafting experience, might devise for themselves. the question bank aims to keep up with the constantly increasing tempo of new questionnaire instruments coming on stream, and the release of survey datasets to the uk data archive. retrospectively, we decided to try to cover the period from 1991 (a census year) onwards, but not to attempt systematic coverage of the period before 1991. however, a number of surveys conducted before 1991 are still used as benchmarks, or exemplify particular innovations in question or data collection design. for these we have made exceptions to our rule, so that the qb contains questionnaires for selected surveys conducted during the 1980s. maintaining this material is not an easy task. one full time content manager works to identify, update and add new material to the site. each questionnaire that we obtain has to be fully marked up in adobe acrobat, bookmarked, and linked to all related material, all of which takes time. however, this task is essential if material is to be easily accessed by qb users. topics the topics section of the qb site is aimed at assisting researchers with social measurement. as oppenhiem states; “the questionnaire has a job to do: its function is measurement” (1992:100). the qb clearly has questionnaires from large scale surveys that researchers can use to measure phenomena. however, the topics area of the qb goes further than this and aims to take 22 key social science topics and provide commentary on how social measurement can be done within them. these 22 topics are: · crime and victimisation · demography · economic activity · education · ethnicity and race · family · gender · geography · health, illness and disability · housing and household amenities · household definition and structure · income, expenditure and wealth · leisure and lifestyles · political behaviour and attitudes · religiosity · social attitudes in general · social capital · social class · social protection and care · travel and transport · voluntary associations · working life each topic has a table detailing surveys that contain questions seeking to measure that phenomenon. for each topic area, we aim to provide a summary account of the main concepts involved and current approaches to measuring those concepts using quantitative methods. further, an aim from the beginning of the qb has been to have specially commissioned commentary written by experts in the field on measurement of that topic. the commentary is intended to help users to understand the conceptual structure of each topic area and the way it is reflected in the structuring of questions. it may also make users aware of other concepts and questioning approaches that may be closely related to the ones that they had in mind when accessing the qb. the commentary includes discussion of available objective evidence on the validity and reliability of the measures produced by questions used. this has been problematic for a number of reasons; the main one is that such experts do not have time spare to write. another problem is the nature of the qb itself. as a free online resource, academics do not get recognition from their peers for writing for the web. having said this, the topics area is very well used and is especially popular with those that are beginning measurement of a new topic and who are looking for further guidance. we are making every effort to ensure that this area is as well populated as possible. an editorial board has been set up to monitor the quality of material in the qb, and to commission commentary from experts in their respective subjects. this will further the aim of building up within the qb site an electronic encyclopaedia of the social survey world. in addition we provide bibliographic references to relevant social science research literature on each topic, and hyperlinks to other internet sites containing relevant information. 8 iassist quarterly winter 2005 resources this section of the site houses a number of different resources which researchers may find useful for their survey research. as well as information about the qb, including teaching materials, user guides and historical documents, this area includes an explanation of the following resources: harmonised question forms the office for national statistics (ons) initiated a programme of work and negotiation that aimed at arriving at a set of variables with question wording harmonised across surveys. the result was a booklet entitled ‘harmonised questions for government social surveys’. these harmonised wordings have now been, or are in process of being, adopted in all major continuous government social surveys and are feeding into work on the new integrated household survey. capi documentation iassist members will be well aware that most contemporary large-scale interview surveys are carried out by the interviewer carrying a portable computer on which the questionnaire resides as a program for computer assisted personal interviewing (capi). this represents a major technical advance in the survey process, but also poses a challenge in making the capi interview intelligible to the layperson. students in particular are often unaware of capi and the implications it has for questionnaire design and administration. the qb team is working to make the way capi surveys work in the field as transparent as possible to qb users. survey link scheme the esrc survey link scheme exists to give academic social scientists the opportunity to acquaint themselves with professional social survey research, carried out by the office for national statistics, the national centre for social research and various market research companies. it thus provides a bridge between academia and the practical world in which professional survey research is carried out. two linked components are offered, each typically taking one day each: ·• one-day workshops, which provide a briefing on a particular survey. the day includes an introduction to survey interviewing in the field, an introduction to capi – computer assisted personal interviewing — and guidance through the capi questionnaire of a particular survey by professional staff from the agency carrying out the fieldwork. these workshops are held at various locations throughout the uk ·• the opportunity to go out with a professional interviewer for one day, and observe a social survey interview in the field. this can be arranged close to where the scheme participant lives, and takes place after attendance at the workshop. the website for the scheme is maintained by the question bank and professor martin bulmer directs both resources. further resources are planned for the qb in the coming few years. these include new fact sheets about the survey process, and further information on how researchers can utilise new technologies in their own small scale research. searching the qb in march 2006 the qb implemented a new advanced function for our apr smartlogik search engine. the search engine was specifically chosen because it searches within pdf documents as well as html pages. the new advanced search makes it possible to search in the following ways (taken from the qb search help pages): · search with all the words use this section if there are several different words that are relevant and you want to find only documents that include every one of these words but not necessarily in any particular order. this reduces the number of results and should narrow the search to more relevant ‘hits’. example: a search for “pollution” will generate over 120 hits while a search for “pollution, health” in this section generates around 35 hits (march 2006), these documents will have elements involving each word although not necessarily in the same sentence or question. · search with any of the words use this section if you have several different words that may be equivalent to each other, maybe as alternative descriptions of the same concept, and you want to find documents that include only one, two, or more, of them. example: search in this section for “marijuana, heroin, cocaine” to find questionnaires about illegal drugs ( a search for the word “drugs” would also find prescription drugs mentioned in health surveys). · search with the exact phrase this search is similar to “search with all the words” above but with the addition that the order of the words in the search enquiry becomes significant. please note that the search engine may not use every single word of the exact phrase to find the hits, it ignores the little words like “of”, “and”, “are” or “with”, and so these may not be highlighted in the results. however, provided the phrase you have searched for is reasonably complicated you should have mostly correct hits. example: search in this section for “people like me have no say in what government does” to find all the surveys that have used this particular question about political efficacy. · search without the words this section should be used together with an all words or iassist quarterly winter 2005 9 an any words search. sometimes a term you want to find has more than one use or meaning, so you would use this section to try to exclude the hits that involve the meanings that you do not want to find. do this by putting in this section some words that would often occur close to the search word when it is being used with the meaning you do not want to apply. example: the word “train” could refer to a form of transport or an education process. to find materials about the form of transport enter “train” in the all words search and “learn, program, vocation, job, college” in the without the words section. you will need to add more terms to the without section to eliminate every unwanted hit but this should convey the idea. · combining sections in the search command it is not possible to combine an exact phrase search with a without the words search. but you may combine an exact phrase search with either an any words or an all words search. example: repeat the “marijuana, heroin, cocaine” any words search from above but combine it with an exact words search for “have you ever tried” and then again with “have you ever taken” and compare the two sets of results. · search by survey another way of eliminating unwanted results is to use the survey titles to select just the surveys that you would expect to be relevant to your search. you may select more than one survey to be included by holding the control key down as you make selections from the scrolling list using your left mouse button (mac users should hold their apple key down while making selections with their mouse button). do not forget to clear these selections by clicking on the all item at the top of the scrolling list before starting a new search. · search by questionnaire year it may be possible to restrict your search further by limiting it to a specific year. however this particular element may also prove to be unreliable in some circumstances due to a lack of consistency in recording survey year numbers. where a survey is identified with single calendar years (eg. 1997) this function should work properly. but where a survey is recorded as “97/98” or “97/8” there is less likelihood of this function working well. in some cases selecting 1997 will generate hits in surveys titled 1997/98 and 1996/97. · search by document type when you have become familiar with the question bank you may find this option more useful to limit the types of material offered to you in search results. all of the original materials prepared by the researchers and survey organisations that can be found on this website are stored and displayed in pdf files (which can only be opened by using adobe acrobat reader). these are mainly questionnaires but also include show cards, advance letters, interviewer instructions and technical report extracts. most of the materials prepared by the question bank staff to help users locate particular surveys or to provide background information on survey methodologies, topic guides, links and bibliographies, are stored and displayed as html files (web pages). interpreting the results of a search each hit result is headed by a document title which should include a good indication of the type of document that has been found and, if it is a questionnaire extract, the particular survey series and year applicable. if your search terms have been found in several different parts of a document you will probably be shown a separate hit for each page involved. on the results list you should be able to see the successful search word highlighted in red with about 50 words to show the context in which it has been found. this should help you to select the hits which are most likely to be useful to you. when you click on a particular hit to follow the link there will be a slight delay while the file is downloaded to your computer, after which you will see the beginning of the document on your screen. there may then be another pause before the display changes again to show the relevant section of the file with the search words highlighted in blue or grey. you may want to scroll up or down the document to read more material in the vicinity of the located result, and you may find it helpful to zoom in or out to see the text more clearly or to see its context on the page. to return to the search results, click on the “back” button on your browser toolbar. the document title that you have just opened should now be displayed in a different colour to that of documents you have not followed-up. maintaining the resource as i have already stated, the qb is a large site which is maintenance intensive. until november 2005 when further funding was received for a team of four staff, the qb was run from the university of surrey, department of sociology with a complement of 2 staff, one of whom was part time. this insufficient funding is a common problem for research projects. however i mention it here because it has become clear, now we have a larger team, just how much we could do with the qb in an ideal world. one of the most problematic areas has been the technical knowledge that we have at times desperately needed. from the beginning the qb has been a social science project rather than an it project. all of its team members, with the exception of one past manager, have been social researchers. it is a constant struggle to keep up with the 10 iassist quarterly winter 2005 latest web developments, with developments in server technology and to fix technical issues. in short, the web development side of the qb has been a very steep learning curve for the team, which we will continue to struggle with whilst attempting to keep the balance between a comprehensive resource on social measurement and an easy to use web resource. the main priorities that we have for the next three years are to keep updating the site, particularly the questionnaires and the topics; to conduct a thorough literature review of capi and possibly conduct further research on this; and to enhance the resource through adding some limited european content. we are also hoping to be able to redesign the web pages with a more modern and user friendly look. the qb is also gaining experience and expertise in questionnaire archiving processes, especially with the data documentation initiative. clearly the materials that the qb houses are produced by external agencies and are largely untouched by us, except for the addition of bookmarks and explanatory material about the survey. organisations with which we work closely, such as the uk data archive, have a great deal more technical expertise than we have in initiatives like the ddi and in techniques for digital preservation of data (and metadata), and it is envisaged that we will be working much more closely with the archive in the future. we are therefore extremely pleased to be involved in this iassist conference and the following ddi meeting, and we welcome comments and discussion with iassist members. conclusion students and researchers are increasingly turning to the web in order to find materials for their work. the qb is a large web resource aimed at social science researchers, students and teachers of research methods and secondary analysts. it contains information about large scale uk social surveys and their questionnaires, which enables users to recycle questions in their own work, thus avoiding having to reinvent the wheel. it also aims to have commentary on social measurement in 22 social science topic areas and various other resources linked to survey research. i have explained that the qb is intensive to maintain, but that with additional funding from the esrc, progress is being made towards making the qb into a useful and userfriendly resource for the future. useful links: apr smartlogik: http://www.aprsmartlogik.com esrc question bank: http://qb.soc.surrey.ac.uk esrc survey link scheme: http://qb.soc.surrey. ac.uk/sls.htm economic and social http://www.esrc.ac.uk research council: economic and social data http://www.esds.ac.uk service: office for national statistics: http://www.statistics. gov.uk national centre for social http://www.natcen. ac.uk research: nesstar: http://www.nesstar.org uk data archive: http://www.data-archive.ac.uk references: jones, steve (2002) ‘the internet goes to college: how students are living in the future with todays technology’ washington, pew internet and american life project available online: http://www.pewinternet.org law, margaret (2005) ‘reduce, reuse, recycle: issues in the secondary use of research data’, iassist quarterly, 29, volume 1 2005 pp 5-10 available online: http://www.iassistdata.org/publications/ iq/iqvol29.html oppenhiem a.n (1992) ‘questionnaire design, interviewing and attitude measurement’ london, continuum acknowledgements: i would like to thank graham hughes, laura hyman and zoe tenger for their comments. * paper presented at the iassist 2006 conference in ann arbor, mich., session: “moving beyond data to networked knowledge”. contact: julie lamb, department of sociology, university of surrey, uk. telephone: +44 (0)1483 683762, e-mail: julie.lamb@surrey.ac.uk footnotes 1 http://www.socstats.soton.ac.uk/cass cass ceased to exist in 2005 when the question bank and the short courses were refunded as separate entities. 2 http://www.statistics.gov.uk/about/data/harmonisation/ default.asp 3 nesstar is an online tool which enables basic data analysis. it is used by the economic and social data service to display key uk social survey data sets such as the general household survey. 4 the uk economic and social data service (esds) works to provide access to uk survey data and to support its use. there are five services within esds each with a remit for data support in the specialist area; esds government, esds longitudinal, esds international, esds qualidata and esds access and preservation. all can be found online: http://www.esds.ac.uk vol192 13summer 1995 data libraries as vending machines; or, what we can learn from arthur dent by laura guy 1, 2 university of wisconsin-madison "...technology causes trouble. as a major agent of change it intrinsically, not accidentally, dislocates and distresses established relationships and forces economic, political or social change3." abstract technological change has had a tremendous impact on how we do our jobs. it not only has affected how we organize and provide access to information, but how our users conduct their research. this change has created new challenges for our profession, not the least of which is wondering if it will make us obsolete by replacing us with knowledge-based systems. the nature of these changes is discussed and we fantasize a bit about the data library vending machine. finally, we look at our users and how we might best continue to provide them with the services that they need. the vending machine analogy when i was a little girl we would visit my grandfather where he worked in one of the state office buildings in st. paul, minnesota. in the basement of this building there was a little cafeteria--a lunch counter. it was staffed by one or two people, and they sold sandwiches, soup and drinks. i'm sure that similar lunch counters existed in office buildings all around the country. but if you go to one of these buildings now you will likely see a bank of vending machines. what happened? although it may be overly simplistic, it appears that the people lost their jobs to technology. what were the reasons? vending machines may be cheaper. they take up less space. they are on duty 24 hours a day, seven days a week. they may be more efficient. they don't require vacation time or a health care plan. they don't complain about work conditions. in the last few years, we have seen vending machine technology advance. some of them can talk. they can take $1 bills and even $5 bills and make change. they dispense hot soup and cold salads. the questions that we are dealing with are as follows: can systems be developed that provide users access to data without the need for data librarians? to what extent can what we do be replaced by vending machines--that is--data vending machines? what are the tasks that can or should be automated or eliminated by technology, and after that happens, will there be anything left for us to do? coping with change i doubt that there is one of us who isn't simply breathless at the speed of the technological change that we are experiencing. today's leading edge technologies quickly become commonplace. it's likely that as professionals our primary task over the next decade will be coping with this change. much of the transformation is evolutionary in nature rather than fundamentally discontinuous: the change builds on itself. the speed and breadth of the transformation we are experiencing creates interesting challenges as well as opportunities. as the old axiom goes: "god protect me from living in interesting times." these are indeed interesting and exiting times for us, brought in part on the following changes: 1. change caused by technological advances in hardware: client/server technology: the architecture of the computer systems we use is changing rapidly. hardware is becoming more "personal" and "portable." personal devices are connected to powerful servers that are part of a distributed information system. storage media capacity: multi-gigabyte local storage capacity is becoming commonplace. advances in disk technology and compression rates facilitate the storage of vast amounts of information on-line and provide interactive access. 14 iassist quarterly network capacity: the development of ethernet and token ring networks enhances connectivity and provides extremely reliable data throughput. today's hardwired networks will become tomorrow's wireless networks, enabling users to connect to anywhere from anywhere. speed capacity: over the last few decades we have seen the doubling of raw cpu power every 18-24 months. today's 20 and 50 mips machines will become tomorrow's 1,000 mips machines. memory capacity: we now have 16-megabit memory chips available, and soon we will see 64-megabit chips. future computers will be able to store hundreds of thousands of typed pages in a computer's main memory and millions of pages on local disk drives. 2. breakthroughs in software: new software architecture: the hardware revolution has enabled the development of easier-to-use software. this software in turn encourages the creation of collaboration tools and work group environments. the interfaces are more powerful, and applications are moving from the stand alone model to intelligent workflow. we see multi-tasking and multi-media capabilities. the virtual environment: graphic techniques and animation are transforming the way we visualize information and complex computations. artificial intelligence modeling and virtual reality coupled with enhanced visualization capabilities allow users to explore and interact with a virtual environment. if you think this sounds a bit farfetched, watch one of the america's cup programs on espn. software developed by silicon graphics incorporates a variety of measures such as wind speed and direction, compass readings, boat location, speed and course information into an astounding real-time graphic display. document-based systems: the concept of "document" has become more complex with the advent of hyperlinked data, text, graphics, video, sound, and so forth. the tools that operate on these documents also have become more sophisticated. information becomes document-based when documents exist separately from the applications that create and operate upon them. object-oriented systems: suited to modeling complex problems and processes, these advanced systems have the ability to self-update and communicate output in a variety of manners (voice, visual, etc.). software tools are more modular and applications more flexible and powerful. 3. we see a paradigm shift from the data processing model to the information processing model as described by ronald weissman of next4. in the new model information becomes content enriched, existing in an environment of "creativity" and "ambiguity" (an example might be data in a spreadsheet as opposed to information on the world wide web; the first is static, one-dimensional, unambiguous and incomplex, the second is dynamic, multi-media, possibly ambiguous and capable of presenting complex subject matter). in more concrete terms, we see documents linked with abstract, index, and bibliographic information, and numeric data linked with meta-data. such documents may be more accessible and, for the general public, perhaps more captivating. 4. we've seen changes in our users: scholarly research methods have evolved into what michelson/rothenburg call "network-mediated scholarship5. scholarly communication and collaboration, as well as the broader research process, have undergone significant transformation. many of the changes our users are experiencing are in large part technology driven. 5. we are experiencing changes in levels of connectivity. for those of us who are "wired" there is an enhanced ability to access, analyze, disseminate and communicate information instantaneously and without regard for distance. 6. there is an increasing amount of information published in electronic form (for example, the growth of government information) and a growing number of formats for electronic records and information such as e-mail, cd-rom, magnetic tape, word processor publication, dial-up services, on-line services, g.i.s., spreadsheets, relational databases, floppy disks, and bulletin board systems. 7. recently there has been a hypermedia revolution and accompanying it the concept of the nonlinear document (and the thought process behind it--which is not that new6!). we see multidimensional data that integrate diverse formats of information. recall the information processing and the data processing models mentioned above. the new paradigm includes a growing complexity of systems, increasingly sophisticated applications, and a plethora of document types. 15summer 1995 it will be essential for us as librarians and archivists to project and assess the importance of these changes. clifford lynch warns that in the 1980's "many research libraries...thought that users' needs for access to online database searching were substantially overstated7." by not meeting the challenge of a transforming environment we risk making our libraries irrelevant and ourselves obsolete. and while we must remain agile, at the same time we need to be wary of developing systems and new services that may be poorly matched to the needs of our users, poorly designed or poorly implemented. decentralization of information resources in the last several decades we have seen the transformation of traditional, paper-based, largely manual information systems into automated electronic systems. the obvious example is the library card catalog, which was transformed in the early 1980's by the development of online public access catalogs (opacs). we have experienced a more recent evolutionary change into a distributed or decentralized information êenvironment. for example, by the late 1980's opacs were beginning to be made available through the internet. later, to obviate the need to learn new user interfaces to access the various opacs, the z39.50 standard was êdeveloped. the standard is based on client/ server technology and greatly facilitates network information access. this evolution towards distributed resources closely follows the evolution of the internet itself. the development of the early arpanet in the late 1960's and early 1970's had a strong economic basis because of êthe great expense of computers; it enabled resource sharing. as technology became cheaper, the need to centralize (cost share) decreased. in the 1980's the financial need for groups to come together to share computing resources decreased. the death of centralized mainframe computing followed closely the advent of minicomputers and micros. increased networking capabilities allowed individuals to link to other individuals for reasons other than cost-sharing and without any regard for physical proximity. in the 1990's the dumb terminal attached to a mainframe computer has been replaced by the stand alone "scholar's workstation," which has powerful cpu, lots of disk space, a cd-rom drive and an internet connection. a familiar manifestation of this decentralization has been the proliferation of information resources on the êinternet. the sharing of information has been greatly facilitated by the development of tools such as anonymous ftp, the wide area information server (wais), and the world wide web. research projects that collect data such as the national survey of families and households (nsfh) now have the technological ability to cheaply and easily become their own access providers. individuals now have the capability to become resource centers. bill goeffe's resources for economists on the internet8 , is a good example of what one person can accomplish using technology no more sophisticated than what can fit on a desktop. information decentralization causes librarians quake in their shoes. and rightfully so: as the anarchistic nature of the internet intrudes into information systems there is a recognized lack of standardization, centralized authority, and access control. there is no institutional control over individuals like bill goeffe, who may in the twinkling of an eye forsake his resource and allow it to lapse. similarly with the nsfh, one might ask what happens to the project's data when the funding ends or the research group disbands? in a decentralized or distributed environment access becomes independent of location. the physical location of a resource has little meaning. in fact, something that virtually appears as a single resource might physically exist in separate parts in disparate locations. from the user's viewpoint it is not really important where the resource lives9. in this new information model our old notions of control require reformation. who is responsible for a resource that consists of multiple copies dispersed among multiple locations? new forms of control may be required for insuring the continuation of important information resources. data libraries as vending machines as librarians and archivists we exist to serve users and preserve information. if we are to be replaced by vending machines it will be because they serve users and preserve information better than we do. they would be cheaper, replacing staff and facilities with computer hardware and software; they would be easier to use, enhancing the scholar's workstation and providing access to needed information from the desktop; they would be faster, allowing for the instantaneous access to any information at any time; they would be decentralized, maximizing connectivity through resource sharing; and, they will be wihout boundaries, providing access to users no matter where they may be. most data users go through a process that consists of four separate steps: 16 iassist quarterly 1. identification of potential data sources 2. determination of the usefulness of the data 3. obtaining access to the data 4. obtaining analyses from the data any data vending machine environment would have to play a role in each of these steps. it would have to assist users possessing varying levels of experience as they make their way through a research sequence that is not necessarily linear; for example, it may not be until the user reaches step 4 that they realize that the data do not êmeet their needs. the system would have to deal with the lack of standardized terminology to describe data and the lack of standards for formatting and storing data. the data library vending machine would have to be smart. it would need to be capable of a multitude of sophisticated activities such as conducting reference interviews, answering questions, conducting searches, and assisting with problems. it would need to be capable of serving a diverse group of users that spans a wide range of abilities, including computer knowledge, communication skills, and research expertise. it would need to be able to instruct as well as assist, and preserve as well as make accessible. who are our users and what do they want? its an obvious assertion that we need to know who our users are and how we can best meet their information needs. to be sure, they are a diverse group, but in general we can talk about them in terms of what they are usually involved in: the research process. this process can be divided into five parts: 1. identification of sources 2. communication with colleagues 3. interpretation and analysis of data 4. dissemination of research findings 5. curriculum development/instruction of next generation the research process has undergone a tremendous amount of transformation. these changes in turn have an êimpact on how we, and others associated with the research process, do our jobs. the research paradigm includes a diverse group of agents: researchers, publishers, computer specialists, vendors, librarians, archivists, and professional organizations. information technology has influenced changes in all of them. in the last several decades end-user computing has become more convenient, cheaper, faster, more powerful, and êeasier to use with sophisticated interfaces and the advent of interactive environments. technology has provided increased connectivity that enhances the research process through the expansion of access and the facilitation of communication. as librarians we are more commonly serving a "remote clientele." users frequently do not need or wish to "come through the door" to receive assistance. this change, although gradual, impacts directly on the way we provide êour services. much of the communication is electronic in nature. the need for remote access to documentation as êwell as data is a part of our users' growing demands and expectations for timely and adequate service. we've seen a transformation in the capabilities of our users. some are very technologically sophisticated and have access to the most powerful computing resources. these people, typically faculty in my environment, prefer to have minimal interaction with the library. they want what they want when they want it, and prefer to require little human assistance. i can imagine this group working well in a "data library vending machine" environment of the type i believe is practical within the constraints of today's technologies: an environment that includes ftp or www accessible data and meta-data, intelligent searching, extraction, and analysis capabilities. 17summer 1995 others users, often but not always students, have varying degrees of computer sophistication and varying access to resources. many have very little experience with ftp or the world wide web. they do not have the computer resources required to, for example, download large files off the internet. many lack a basic understanding of the nature of numeric data. i have difficulty thinking of this second group as a potential vending machine clientele. the front end that would be needed to teach these people what they need to know in order to use any given data set is beyond my imagination. we are tied inextricably to the information needs of our constituency and must always monitor these needs êcarefully. it is critical that we work with our users and the other agents involved in meeting their needs (for example information providers and computer scientists) to continue to push forward information management. the "expert systems" fear expert systems, sometimes called "knowledge-based systems" are a type of computer program that uses the knowledgebased techniques of artificial intelligence. simon hayward10 described them as computer programs that represent knowledge and apply expertise to manipulate that knowledge and to achieve solutions. over the last decade or more the fear of being replaced by expert systems has been sounded in many places. by-in-large, this has not come to happen. certainly we will see, coming out of artificial intelligence, the development of intelligent agent technology that will provide aids for locating, evaluating, analyzing and interpreting information. but, for data librarians the fear of being replaced by sophisticated expert systems that interface data and meta-data with users is a concern that may, for the foreseeable future, be unwarranted. the omniscient, omnipotent and omnipresent data library vending machine described above will probably not exist for a long time. nevertheless, we must consider: to what extent can what we do be replaced by expert systems? alternatively, what is it that we do would we like to see eliminated by technology? what might we do with the extra time that will facilitate our users and enhance our role in the research process? and, can this process be viewed positively rather than as a threat? an informal users' survey: during the months of march and april, 1995, i conducted an informal survey of users of the dpls. i kept track of the complexity of their needs, their level of sophistication, and their access to resources. the question i wanted to answer was the following: which of these users would operate well in a data vending machine environment? i found that 30% of our users would be well-served by the vending machine data library. they are experienced and familiar enough with numeric data that they need very little help to use it; these people just need to know where the data are. as more data are made available publicly on the internet (for example labstat , the psid, and the penn tables) we will have less frequent contact with these users. the other 70% are the users who don't know what data sets are available or which data they want to use. they don't have the experience working with data to understand what they need to do to access it or how statistical software works. there are many who simply like to come in and talk about their ideas and appreciate whatever type of feedback they might get. some users who have little experience using computers are scared and need comforting and hand-holding. some of our users prefer to look at paper-based resources. there are those who don't have a clue as to how to do secondary analysis. the continued growth of cd-rom publishing, and the development of increasingly sophisticated interfaces will change the above percentages. for example, compare an internet-based interface that merely provides a raw data set (like ftp) with an interface that allows people to do sas or spss runs interactively in real-time on a remotely stored data set (such as can be built today on the world wide web). or compare the latter with an on-line data center system that provides all metadata associated with a multitude of data sets, including variable-level information, instruments, methodology discussions, bibliographies, and so on. the percentages will also be changed by the users themselves. over the years we've seen an increased recognition among scholars of the importance of quantitative analysis skills, and a spreading of this recognition to less traditional fields such as history and education. but as the shear numbers of users doing quantitative research increases, they are coming to us more technologically sophisticated because they are exposed to computers at an ever earlier age. the new technologies described above promise to eliminate some of the more tedious parts of our jobs such as tape 18 iassist quarterly rollovers and extracting data for users. this will free us up to deal with important issues such as proper levels of documentation, the structure of the information systems providing data and meta-data, and bibliographic, abstracting and indexing problems. it allows us to devote more time to assisting users, teaching, and developing standards and policies. its likely that in the process we will redefine what it is that we do and what it means to be a "librarian." we need to ask ourselves what does the word "library" mean? is a library just a physical place or might it become more? an environment? a mind-set? a virtual world? or, could it become less? a computer attached to the internet? we must also ask ourselves as user demands and expectations grow, can we provide? so what does arthur dent have to do with all this? when i first started to think about this panel i remembered something that i had read: "after a fairly shaky start to the day, arthur's mind was beginning to reassemble itself from the shellshocked fragments the previous day had left him with. he had found a nutri-matic machine which had provided him with a plastic cup filled with a liquid that was almost, but not quite, entirely unlike tea. the way it functioned was very interesting. when the drink button was pressed it made an instant but highly detailed examination of the subject's taste buds, a spectroscopic analysis of the subject's metabolism and then sent tiny experimental signals down the neural pathways to the taste centers of the subject's brain to see what was likely to go down well. however, no one knew quite why it did this because it invariably delivered a cupful of liquid that was almost, but not quite, entirely unlike tea." the hitchhiker's guide to the galaxy11. there is a real danger that the vending machine data library will almost, but not quite, entirely be inadequate. the human factor in the vending machine data library might the human factor be missed? it is important to keep in mind that in electronic systems the human factor plays a very important role. in the words of one science fiction author, it is the human factor which gives these systems their "heart." thus, it will be humans who must deal with pertinent issues such as standards development, information integrity, accountability, and responsibility; it will be humans who realize the importance of metadata as an essential supplement to standard bibliographic approaches; it will be humans who design the systems, who implement policies and develop the tools and criteria by which the systems will operate; it will be humans who will develop new descriptive systems, finding aids, navigational aids êand informational hooks that are suited to the constantly changing electronic environment and user demands. there are equations where the human factor hasn't been missed, for example, in bowling alleys there once were "pin boys." it is a cold hard fact that if an entity can be adequately replaced by technology under the current êsystem it probably doesn't deserve to survive. interestingly, at the university of wisconsin, human-staffed delicatessens are proliferating. why? the profit motive? the need for human interaction? could it be that we still do need to have that perfect cup of tea? conclusion we live in a time when concepts like "unique" and "multiple" are becoming obscure, as are "library" and "archive." our survival and transformation as a profession depends on how we respond to the changes we are experiencing and the new paradigm in which we find ourselves. it is important that these changes not be perceived as threats but as opportunities, and that we work to turn what might be weaknesses into strengths. as the information infrastructure becomes stable and established, focus will shift to areas that are open to contributions that we can play an important role in making. for example: meta-data engineering for better methods and tools for describing information. development of new descriptive systems and finding aids. development of access tools to facilitate navigation and information retrieval. development of improved user interfaces. development of new governance and control mechanisms over information. 19summer 1995 standards for document and data management dealing with diverse areas such as scanning, text encoding, and storage, and retrieval issues. insuring the continuation and widening of the information infrastructure. teaching colleagues and users. copyright, intellectual property,privacy, and public use issues. promotion of the archival mandate and the protection and preservation of electronic information. we must temper "carpe diem" in an environment of decreasing funds, private sector competition, and the need for us to develop advanced skills and expertise. in light of the changes we are experiencing, it is clear that there are many challenges ahead if we are to remain a viable and useful profession. these challenges will require êinnovation, agility and deep understanding of our environment and the needs of our users. defining these challenges may help us to continue and grow as a profession; meeting them will definitely lead us to live in interesting times. 1 paper presented at iassist 21st annual conference may 9-12, 1995, quebec city, canada. 2. i would like to thank derek zahn for his substantial contributions to this paper. 3. coates, joesph f. "science, technology, and human rights," technological forecasting and social change 40 (1991): 389-391. 4. weissman, ronald e. "archives and the new information architecture of the late 1990's," american archivist 57 (winter 1994): 20-45. 5. michelson, avra and jeff rothenberg. "scholarly communication and information technology: exploring the impact of changes in the research process on archives, " american archivist 55 (spring 1992): 236-315. 6. bush, vannevar. "as we may think," atlantic monthly 176 (july 1945): 101-108. 7. lynch, clifford. "archiving the promise: a proposed startegic agenda for libraries and networked information resources in the 1990's." in netwoks for networkers ii, edited by barbara evans markuson, 52-84. new york: nealschuman, 1983. 8. goeffe, bill. "resources for economists on the internet." version 8.02 <url:http://www.shusu.edu/ftp/economics/ econdata/.www/econfaq/econfaq.html> 9. that is, until it disappears. 10. hayward, simon a. "is a decision tree an expert system?" in proceedings of the british computer society specialist group on expert systems 4th conference, 185-192. cambridge, england: cambridge university press, 1985. 11. adams, douglas. "the hitchhiker's guide to the galaxy." new york: harmony books, 1980 vol30-1.indd 20 iassist quarterly spring 2006 by san cannon and meredith krug* everything but the kitchen sink: building a metadata repository for time series data at the federal reserve board in support of the board’s duty to conduct monetary policy for the united states, the research divisions at the federal reserve board use a variety of time series data for both research and forecasting. the collection, maintenance, and upkeep of more than 50,000 time series from more than sixty sources in a central location are daunting tasks; documenting the metadata for the compilation and use of these data is even more so. we are currently building a comprehensive metadata repository that links three kinds of metadata about our time series: structural metadata, reference metadata, and operational metadata. many of the pieces of this puzzle currently exist in an array of disparate formats: attributes in a proprietary database, html pages on a website, word documents buried on a file server, etc. we are bringing these pieces of information together in a relational database setting to allow users to search and display relevant metadata for a particular series or economic concept. in addition, we are working to make the metadata entries in the repository time sensitive to accommodate the database of contemporaneous “real time” or “snapshot” time series that we are building for future research purposes. this paper will examine the different types of metadata gathered for time series data, and how this information is collected, stored, and made available to staff at the federal reserve board. metadata types the concept of metadata or “data about data” raises questions concerning what users need to know about different pieces of information. we have identified three types of metadata that cover the information we gather. the first and most fundamental type is structural metadata. structural metadata are usually a small set of concrete details that identify and define a particular time series; they can often be expressed as short strings of text that convey the information essential to the meaning of a number. for example, users would have difficulty interpreting the statement that output was 12487.1. what kind of output? how is it measured? when was it measured? these questions may seem basic, but their answers are integral to the correct interpretation of this string of digits as economic information. if a few short strings such as description, country, time period, and unit of measure were defined, and you were told that the gross domestic product of the united states for 2005 was 12,487.1 billion u.s. dollars, the number would now have meaning. further details on the construction of a statistic are better classified as reference metadata1 reference metadata contain more detail about the calculations behind a number and may describe the data collection, sampling methodology, preparation of the estimates from the sample data, treatment of revisions, and the reliability of the estimates. reference metadata are often presented as a document provided by an issuing agency in an article or bulletin, which may eventually appear on its website. researchers often need to understand the more complicated details of how a series is constructed. reference metadata are invaluable when deciding on the appropriate series to use in research, especially if more than one possible measure is available. for example, the department of labor publishes the employment cost index (eci) and the employer cost for employee compensation (ecec), both of which have data on the cost of labor to employers. typical structural metadata on these statistics would indicate that the eci series are indexes with a base year of 2005 = 100, and that the ecec series are measured in current u.s. dollars. the reference metadata, however, would get to the heart of the calculation behind the time series and show that the eci indexes use “fixed weights to control for shifts among occupations and industries” whereas compensation figures from the ecec use “current weights to reflect today’s labor force composition.” a researcher can then decide which construction best represents the economic concept under study. the third type of metadata we collect can be classified as operational metadata. these are the processing instructions that explain exactly what is done with the data once they are received and exactly how it is accomplished. this type of metadata is extremely valuable to the data maintenance staff and facilitates the job of collecting, uploading, and disseminating information. typical operational metadata contain information on the source agency’s publication schedule and procedures, file formats of the data, the location of updating programs and how they work, archiving information, data contact information, and other special considerations. iassist quarterly spring 2006 21 all three types of metadata are closely connected. a change in the reference metadata could have consequences for the structural and operational documentation. for example, when the census bureau changed the sampling universe of permit-issuing places from 19,000 to 20,000 for the new residential construction statistics, the reference metadata explain in detail the change in definition. the structural metadata, however, would need to note that there is an important break in the continuity of the series, and the operational metadata would need to be updated with the operational consequences of processing the newly defined data. where we started the two main production databases used by the research departments of the federal reserve board contain about 62,000 time series (52,000 domestic and 10,000 international). these numbers are inflated by multiple versions of series in different frequencies (monthly, quarterly, etc). removing the frequency dimension, there are about 45,000 unique economic measures in the two databases. our current “method” of storing all three types of metadata for this information is a hodge-podge of legacy systems that differ by metadata type. in the past, different information was stored in different formats on different platforms, depending on where the metadata were originally placed, who placed them there, and who needed access to them. structural metadata the structural metadata for our time series data are stored with the data in a fame database. designed to manage times series, fame software allocates some characteristics when a series is created, such as name, description, database name, first value, and last value. while these native attributes are useful, they do not provide enough information for the data to be uniquely identified and used correctly. so that the data will be more useful to the economists and analysts who use them, we have implemented two additional metadata constructs to convey information about the series: a detailed, hierarchical nomenclature system that restricts the name of the series, and a set of additional attributes to help capture some of the more critical information about each data series. the hierarchical nomenclature system is our key to storing, in a single repository, a vast collection of data series covering a wide range of topics. the system was developed in the 1980s from a skeletal structure purchased from the consulting firm of townsend greenspan, and over the past quarter century board staff have expanded its scope. it allows for the categorization of a variety of macroeconomic concepts and is adaptable enough to accommodate the changing economic landscape. we construct our series names from prefixes, roots, and modifiers with a frequency designation, but only the root and the frequency code are required. there are 62 high-level roots covering broad economic topics such as gross domestic product, consumer expenditures, employment, and price indexes. each of these high-level keys has a “branch” of subcategories with further breakdowns as necessary. the result is 22,000 unique “roots” describing the economic concepts. some roots are succinct: for example, pjp indicates producer price index, where pj is the high-level key indicating the broad concept of price index and the additional p indicates the subcategory of producer prices. such a breakdown might be depicted as: pj: price index pj.p: producer other roots, such as the one for two-year installment loans to consumers (riflpbciplm24), are painfully detailed: r: rate r.i: rate of interest in money and capital markets r.i.f: federal reserve system r.i.f.l: long-term or capital market r.i.f.l.p: private securities r.i.f.l.p.bc: commercial banks r.i.f.l.p.bc.i: consumer installment loans r.i.f.l.p.bc.i.pl: personal loans r.i.f.l.p.bc.i.pl.m24: 24-month loan once the root is established, a series name can have up to four “modifiers”—special codes that add more information and are preceded by an underscore. modifiers describe concepts such as base year, data source, product designation, industry (naics and sic definitions), region of the united states, country, currency, import designation, export designation, commodity group, occupation (soc categories), asset class, age group, and market category2. modifiers can also be used to identify a series as seasonally adjusted or not seasonally adjusted, break adjusted or merger adjusted or to indicate that a series contains seasonal factors. in addition, we can identify series that are simple calculations from other series, such as different types of percentage change. a straightforward example with only two modifiers is the monthly producer price index for electric power. as shown here, it has two modifiers listed after the root: one indicating the bureau of labor statistics commodity code and one indicating that the series is not seasonally adjusted: pjp_g054_n.m: pj: price index pj.p: producer _g: bls commodity group 05: fuels and related products and power 22 iassist quarterly spring 2006 05.4: electric power _n: not seasonally adjusted m: monthly finally, a series must have a frequency code at the end of the series name, preceded by a dot. this producer price index ends with “.m,” indicating that it is a monthly series. other examples of frequency codes are .a (annual), .q (quarterly), .b (business daily), and .wf (weekly friday). the meticulous identification of the economic concept with our nomenclature, however, is not sufficient structural information for our researchers and analysts. our second construct is a set of additional attributes stored in the proprietary database format with the data. these attributes provide needed structural metadata not included with the native attributes allotted by fame when a series is created.3 currently there are fifteen additional attribute categories. some attributes are administrative, such as when the series was updated (to the second), who has permission to make changes, what group is responsible for the series, and who to contact with questions. other attributes are more substantive, such as strings containing the source agency and publication, the table and line number on which the data are presented, restricted value attributes for units, unit multiplier and currency, an annual rate boolean, and the frequency of updates. for example, the additional attributes for the producer price index for electric power would be agency: department of labor / bureau of labor statistics publication: producer price index release table: ppi release table 3 line_number: 39 units: index: 1982 = 100 unit_mult: one currency: annual_rate: no section: eim data_group: ppi:price contact: san cannon updaters: meredith krug, other eim updaters update_series: pjp_g054_n.m update_freq monthly time_stamp: 20060418083440.00 reference metadata many times knowing the nomenclature or the discrete list of attributes for a series is not enough information to make a decision about the appropriateness of a series for a particular use. for detailed methodology or reference metadata, economists in the research divisions can refer to the source agency’s documentation. in addition, we maintain a data sourcebook—a set of web pages created from the documentation compiled by a research economist on staff. this information, gathered mainly from the source agency’s information and from conversations with staff members at different agencies, started life as word processing documents and were then converted to static html pages. these pages outline the more detailed methodological information for most of the major statistical releases from which we gather data. typical topics covered include the source, useful links, principal data provided, history, concepts and definitions, sample design, data collection, revision information, reliability information, seasonal adjustment details, and typical uses. for some statistical releases, the metadata are fairly straightforward. the entry for new residential sales data, for example, has only three definitions of interest: “sale,” “houses for sale,” and “sales price.” other series require as many as ten concepts. for example, the sourcebook entry for the current population survey, from which we obtain the household estimates for the employment situation release, explains the meanings behind “civilian noninstitutional population,” “employment,” “unemployment,” “civilian labor force,” “unemployment rate,” “duration of unemployment,” “reason for unemployment,” “not in the labor force,” and “discouraged workers.” these details are important for distinguishing between various definitions of economic concepts and for understanding the properties of some series. in these detailed pages economists will find notations, for example, about the difference between “all persons” and “all employees” measures for productivity and cost data and about the types of adjustments the bureau of labor statistics makes for that distinction in the payroll data. operational metadata clean structural and reference metadata are vital to a data user’s understanding of a time series, but more information is required for the data maintenance staff to properly maintain the actual time series. we have two operational metadata tools to keep track of the statistical releases that need to be updated to the production databases and outline the procedures for updating them. the first and most essential is our data release calendar, which contains all the statistical releases that we follow and keeps track of exactly what date they will be published—daily, weekly, monthly, quarterly, or annually. like the data sourcebook, this calendar was converted to static html from a word processing document. our second tool, the data documentation system, stores the procedures that we follow for each statistical release. each page was originally created as a static document in word and html. the files varied in structure and organization, depending on who initially wrote them. a few pieces of information were common across pages, iassist quarterly spring 2006 23 such as who had updating responsibility for the data and some external contact information. most content, however, varied according to what each author thought was important. some documentation had secondary pages with additional information, such as screen captures and cryptic instructions to “hit control-p” twice. the only consistent element was a loose description of how to incorporate the data into the database on publication day, but even that was frequently out of date. building a new system the three types of metadata are of interest to different sets of users: structural metadata are important to all the users of a time series; reference metadata are important to the smaller group of analysts doing more detailed research; operational metadata are important only to the people managing the data. as the audience size for each type of metadata slowly decreases, our influence over the metadata as data managers increases. we have little say in what structural metadata are specified; the decisions are usually made by committee, and programming for the additional attributes is done by other staff members. we have slightly more to say about the reference metadata, as the decisions are usually made by the analyst or economist most familiar with the data. finally, we have complete control over the specification, collection, and storage of the operational metadata for the data we maintain. operational metadata in working on a system to tie all three types of metadata together, we started with the operational metadata. several years ago, we moved our first operational tool, the data release calendar, from static html pages to a relational database. each statistical release is its own entry in the calendar and provides links to an interface for editing the entry. as publication dates are known, releases can be easily added to, moved within, or deleted from the calendar. the interface also provides information such as the individual responsible for that update; a link to the statistical release on the source agency’s website; and a link to the relevant page in our second tool, the data documentation system. after moving the calendar information to the relational database, we improved the data documentation. first, we identified the key operational metadata categories for maintaining our data that apply to all statistical releases. then we stored each release and its standardized information in the relational database and created a web interface. the main data documentation page is an alphabetical listing of all the statistical releases from which we retrieve data—currently around 100 releases and growing steadily. linked to each release entry is a secondary page containing all the operational metadata categories from the database. these categories explain how we obtain all the information we need and what we do with the information once we have it. the fields include publication schedule, the releasing agency and contact information, the location of the press release, and how to retrieve the data. we also detail how our update programs work, what other programs are run after the data are updated, and who is notified of their availability. we maintain additional documentation about what happens when there is a revision or when our usual update methods are unavailable. an additional section of notes allows the user to document important issues, lessons learned, or information gathered in phone conversations with agency contacts. this new data documentation system is similar to a “restricted wiki,” allowing only the data maintainers to update it and providing simple formatting options for different types of entries. for each statistical release, an edit button allows a user to edit the text fields and make selections from drop-down boxes and radio buttons for each defined category. behind the interface is an sql database table with rows indicating the statistical data product and columns indicating the various fields. editing the appropriate field on the appropriate page is a simple task that updates the database so that the new information can be rendered on the presentation page immediately. the transition to this system from the word documentturned-static html page was lengthy and painful but the payoffs have been tremendous. we now have a clear, concise, and consistent framework in which to store our operational metadata, which is easy to edit in a timely manner. less energy is expended in deciphering and maintaining the documentation. the new system also provides an organized way for all data maintainers to communicate with each other and has made training new data managers much easier. reference metadata our next step is to work on the other type of metadata over which we have some influence. we are currently planning to alleviate some of the pain of editing html code by hand by transferring the data sourcebook from static html pages to a table in the sql database. this system is in its very early stages but is similar in structure to the data documentation system. on the main page, each row is a statistical release and has an edit button to enable the “restricted wiki” interface that accesses the database table containing all the reference metadata. each entry links to a secondary page with standardized categories for the reference metadata with text fields that can be updated. currently, the fields we have identified are: data source, principal data provided, concepts and definitions, history, data collection, preparation of estimates, revisions, reliability, seasonal adjustments, uses, secondary uses, additional resources, miscellaneous, and references. there is the possibility for much more variation in the 24 iassist quarterly spring 2006 reference metadata categories than in the operational metadata categories. for example, every statistical release will have a source, but, as noted above, the number of concepts and definitions will vary. we are still working on the best way to handle such variations. regardless, our hope is that the relevant analysts and the data maintenance staff will be able to easily update the information in the system in a simple and timely fashion. structural metadata because the structural metadata are so closely tied to the actual data, it is more practical to have the primary storage location in the database with the time series. to complete our relational database repository, we simply copy the structural information to a table in the sql database. those who are not interested in using the data in fame often want to know how to find what they want and get it out of fame. we have built a web interface with a search facility that researchers can use to find time series in the database based on the structural metadata stored in the relational database. they are then able to extract the data from the database and use it somewhere else. the challenge of time the next step is to incorporate the element of time into our integrated metadata. neither data nor metadata are static, and changes to the former should be documented by changes to the latter. we are currently building a library of “vintage” or “real-time” data. similar efforts are underway at the federal reserve banks of st. louis and philadelphia, but instead of re-creating history as those projects are doing, we are building a collection for future use.4 every day we store a snapshot of the production databases in their native format. so far we have stored about one year of databases, but in a few years we will be able to look back and identify exactly what information the research staff had for any given business day. capturing snapshots of the data over a period of years, however, may not allow for a complete understanding of the changes to those data over time. the structural metadata stored in the database with the time series should reflect the correct contemporary information, but for many purposes that information will not be enough. changes to the components of an aggregate series may not be reflected in the mnemonic or in any of the stored fame attributes. if the source agency changes their method of seasonally adjusting data or redefines the population it is sampling, the changes would be noted in the reference metadata but not in the structural metadata. to understand the full meaning of a time series, it is essential to know the precise definitions and data construction for a given time period. for full use of the rich metadata resource we are working toward, the metadata storage must be able to depict contemporary information. to capture it, we need to add time-sensitivity to the relational database. we need to be able to capture when edits to various fields in the metadata are made so that we will know the period for which the information in that edit was valid. then researchers will be able to see the reference metadata as they were on any particular date and thus make the most of the vintage data available to them. as before, we are starting with the operational metadata storage and will then work on the reference metadata storage. we are currently building a time component into our data documentation system. we maintain a detailed archive in which data files and programs are automatically stored every time a data file is processed. the archive enables the data managers to restore data in case of a technical problem or to track down the source of an error. being able to reproduce the data-updating instructions for any publication date allows us to take full advantage of the archive we have built; we will have complete instructions on what to do with the various files that are stored with date and time stamps. next, we plan to work on adding a time component to our data sourcebook. final thoughts the goal for this project is to have a gateway to data and metadata that is informative, simple to use, and time sensitive. the biggest challenges we face are those of information retrieval and interface design. we may have a wealth of useful information in our database, but if users are presented with too much information in an incoherent way, the collection is worthless. we need to build a search facility that will allow users to find the information they need for a given point in time, regardless of the type of metadata. after the pertinent information has been retrieved, we need to present it clearly and concisely to the user. once these hurdles have been cleared, the metadata repository for the research staff at the federal reserve board will be an invaluable research tool. * paper presented in the session “the essential role of metadata in resource discovery” at the iassist 2006 in ann arbor by san cannon and meredith krug, federal reserve board. contact: san cannon, scannon@frb. gov. chief, economic information management, federal reserve board, washington dc 20551. endnotes 1 the distinction between structural and reference metadata, as well as the terminology, was adopted from the sdmx initiative’s metadata common vocabulary. see statistical data and metadata exchange initiative (draft, march 2006), “sdmx content oriented guidelines: metadata common vocabulary,” sdmx, www.sdmx. org/news/document.aspx?id=146&nid=67 (accessed june 26, 2006). iassist quarterly spring 2006 25 2 details on the north american industrial classification system (naics) and its relationship to the standard industrial classification (sic) system can be found on the census bureau website at www.census.gov/epcd/ www/naics.html. details on the standard occupational classification (soc) system can be found on the bureau of labor statistics website at www.bls.gov/soc/soc_majo.htm. 3 the “native” fame attributes include information about the object type; measurement information that affects frequency conversions; index information including type of index and first and last value indicators; and creation and update dates. 4 for information on the saint louis project, see katrina steirholz (2005), “economic data as snapshots in time,” iassist quarterly, vol. 29 (2) (summer), p. 5. to access data and documentation from the philadelphia system, see the federal reserve bank of philadelphia’s website at www.philadelphiafed.org/econ/forecast/readow.html. 5 iassist quarterly winter/spring 2010 of harmonization. mary tao provides a timely look at data on the credit crisis. she provides starting points for when a researcher asks for data and/or statistical information on mortgages, credit cards, financial institutions, and other financial crisis-related materials. both fee-based and free resources are included. kristi thompson then turns to the developing world as she looks at survey data on international development. she focuses on survey data for developing countries. she looks at some of the different groups that conduct surveys and make the microdata available. she provides helpful comparisons of the major sources, their sample sizes, geographic and temporal coverage, and access issues. the issue concludes with michele hayslett and lynda kellam focusing on the major changes that will occur with the 2010 united states census and american community survey. a series of examples illustrate the challenges researchers face with comparability. the authors and i hope that future conferences and issues of this journal will have a renewed emphasis on the actual data. there are many subject areas to explore and sharing makes us all more knowledgeable. thanks to kristi thompson for taking my ongoing call to arms at conferences and taking the lead to make it materialize into two panels that formed the basis for this double issue. thanks to these authors for answering the call to arms to talk and to reshape their presentations into full articles about data reference. thanks again to kristi for all of her organizational work in coordinating the original speakers into panels and for serving as a speaker, author, and helping to edit this issue. bobray bordelon, princeton university library guest editor’s notes in recent years, iassist conferences have tended to focus primarily on data documentation, infrastructure, technology, and other technical aspects. all of these topics are crucial to the process. however, the core is the actual data and understanding the subject matter. with the largest number of participants being specialists who help researchers find appropriate data, it was alarming that so little emphasis was placed on a deeper understanding of the subject content and how researchers use the data. a call to arms to focus on the actual data resulted in two sessions collectively titled “data reference in depth” at the 2009 iassist conference, “social data and social networking: connecting social science communities across the globe”. this special issue contains papers that emerged from these two sessions. the presentations have been expanded to further elaborate on data reference. in most research settings, data librarians are required to be fluent in a range of sources dependent on the needs of their researchers. the data librarian needs to be prepared to move easily between often diverse subject areas. the data librarian has to transition from answering a question on population to trade to health to terrorism to religion and everything in between typically across all geographic boundaries and languages. this special issue gives readers a sampling of a data librarian's "day in the life" by highlighting a few major subject areas and their primary data sources. readers will see a few of the particular reference challenges by getting an introduction to some key sources and particular challenges that crop up in a sampling of different areas. in addition, the question “how does providing access to data, as a unique format, affect library reference services?” will be answered. kristin partlo examined how the data reference interview is especially crucial in the provision of research assistance, from the viewpoint of working with undergraduates at a small liberal arts college. the lessons learned have a bearing on research institutions of all sizes and provides a model for getting to the heart of what is really needed. of particular importance is how to balance providing reference with instructing the researcher to learn from the experience. partlo shares the data reference worksheet she devised and provides clear examples. the issue then moves on to major resources available in five distinct subject areas. amy west provides a detailed analysis of the sources and interfaces for international trade, commodities, prices and production. walter giesbrecht follows with a guide to labor looking at both national and international sources as well as the problems acquisition and use of electronic records in the national archives of sweden by magnus geber,' national archives ofsweden, stockholm, sweden this paper genctally reflects the experiences made at the national archives of sweden concerning the management of electronic records and describes how we try to solve the new archival problems that have been caused by these media. 1 will stress what seems to be special for the swedish development and possible reasons for that. before describing the situation in sweden i will present some figures about the general situation concerning electronic records at national archives in the world. since the archival congress in september 1992 1 have been trying to gather information about this matter, mainly by distributing a questionnaire. the table below probably contains most of the countries where national archives have received electronic records. the questions might have been interpreted in different ways so the numbers specified shall not be looked upon as exactl as seen above the national archives of sweden has of today received more than 10 000 magnetic tapes containing electronic records. that would be about 1 terabyte of information totally (1600 and 6250 bpi). this means that although sweden is a small country the amount of electronic recwds delivered to the national archives is amongst the highest in the world. what might then be the explanation of the amount of delivered electronic records in sweden? two main reasons are to be seen, the relatively early and high degree of computwisation among swedish governmental agencies and the existence of the swedish data protection act this is also related to an old and maybe quite unique swedish tradition of keeping a lot of information about the citizens as governmental personal records. swedish governmental agencies expanded especially during the 60s which among other things led to an early introduction of computer systems. many agencies introduced large administrative systems. with large amounts of the governmental information made machine readable the national archives became involved in making disposal decisions concerning some of these systems already during the 70s. this was natural as this information in its traditional form on p^)er constituted basic series of archival records, frequently used by researchers. the other reason was the data protection act. it was promulgated in 1973 to secure the personal integrity in computer systems, governmental or private. the act specifies that all computer systems containing personal table 1. electronic records at national archives use (annual frequence)number of received tapes/cassettes files/datasets datamedia paper(occasions/pieces) canada usa sweden 9000 2 520 10 100 12 000 11646 371 40/1300 360 25/25 danmark norway finland 1200 265 289 985 2 5/10 1/many germany france 1092 2 500 3 500 switzerland 304 527 italy uk netherlands (is said to have) (plans to have 1995) (is having a project on ihe matter) winteir1992 7 information of a certain degree of sensitivity, must be permitted by the data inspection board, and have a limited time to exist. this fact originally prevented the national archives from receiving these electronic records, since the data protection act didn't support giving permits for archival reasons. but with the change of the act 1982, the data inspection board could decide that the electronic records from these systems were to be delivered to the national archives instead of being destroyed. no permission was needed for the national archives. disposal decision were also to be made after consultations with the national archives. still it meant that the disposal concerning a big part of the governmental information wasn't decided by the national archives. however, normally the view of the national archives is accepted by the data inspection board. altogether this has led to a high amount of transfers to the national archives. when a system is ended or parts of the information is getting to old to be stored in the system a decision has to be made. the information must be destroyed or transferred to the national archives. transfers of electronic records had actually started already during the 70s. at that time the records mainly came from different governmental committees which had fuiished their activities, having all their archival records transferred. after the above mentioned change of the data protection act the transfers increased. the records mainly came from large govermental administrative systems like systems for unemployment, insurance, governmental accounting or £^)plying for university studies. the largest part of the information has come, and still comes, from the systems for taxation where tapes are delivered annually from 25 regional agencies. there is also a certain amount of transfers from research projects, the universities being a part of the governmental sector in sweden. a transfer to the national archives offers a possibility for the researcher to have the electronic records preserved with the personal identification even when a study is finished. if there is a need for a followup in the future the records may be requested from the national archives on the condition that a new permit from the data inspection board is given. another main part of the transfers, the foremost being from the taxation systems, is the records from statistics sweden. statistics sweden has of course a large amount of computer systems with personal information. this is regulated by the data protection act. after long discussions between statistics sweden and the data inspection board also involving the national archives, the swedish government decided that a large part of the these records that had reached a certain age were to be transferred to the national archives. this was done in 1989. at the same time the national archives received extra economical resources for this matter. the transfer was partly made voluntary by statistics sweden to avoid the costs caused by clause 10 in the data protection act (that clause enables every person having their personal data registered in a computer system to have an outprint of all that information once a year. but the electronic records transferred to the national archives are excluded from this plight.) the transferred records are still frequently used by statistic sweden as i will mention later. for a long time there was no special staff for the electronic records at the national archives. but following increased funding a section dealing with modem media was created in the beginning of the 80s. today there is an edp-section as a part of the division for technical matters. this division also includes micrognq)hy, conservation and book binding. the edp-section deals with the internal edp use and with the devek^)ment of archival appucations like computerised inventories apart from managing transferred electronic records. for many years there was no computer equipment at the national archives. all use and copying had to be done through service bureaus. some years ago we bought a unix-computer. but we have had problems finding suitable software for our needs. today we are examining the possibility to use a dos-system with special software and different tape-drives, to be able to convert both different physical and logical formats. (the influense of finding such a solution came from visits to the national archives in usa and canada last autumn.) our tapes are stored in a climate archive. they are rotated, rewinded and then copied after 10 years. all tapes are transferred and stored in 2 copies, one being kept in an archives far north of stockholm. when electronic records are transferred the national archives sets certain requirements. this is possible on the basis of the swedish archives act which gives the national archives the right to regulate the management of electronic records considered governmental archival records. the requirements are mainly specified in acccrdaike with different standards. the aim is to get the electronic records in a form as undependent of original hardware and software as possible. we accept ascii and ebcdic but the numerical information may not be in any packed formatthe files shall be in fixed format and may only contain one record-type. the idea is to have a structure corresponding to and also directly importable to a relational database. this means for example that variable files from a system with a hierarchical structure have to be converted before transfer. we have old transfers with files not satisfying matching our requirements. these files will be converted when we make a new generation of storage copies. (it could mean that a variable file containing 5 record-types would be converted to 5 fixed files.) the transferrwl tapes shall lassist quarterly contain files labelled with iso or ibm labels. today we only accept 9-channal spool tapes, 1600 or 62s0 bpi. but in the near future we plan to accept other physical formats. probably we will choose two types for internal longtime storage, for example 3480 and dat. we will possiwy also accept other types of data-media f« transfo' and convert them as long as our software can handle them and we can charge an extra conversion fee. there are also certain regulations about the form of the documentation of the electronic record. generally the documentation is kept on papo'. but we would like getting code tables machine-readable, like an extra table in a relational database. to make it easi^ to import the files to a relational database system, we wouldn't mind getting the record descriptions machine readable. in the future we probably also would like to get the descriptions of how to create a certain archival record firom a number of storage filesaables expressed as standardised sqlcommands. it is important to have in mind that the national archives is collecting public records in machine readable form, which may come from complex administrative systems. the actual record does not have to be similar to what is one data file. the national archives also takes part in differents fields of the standardisation work. primarily, in the standardisation of archival techniques such as terminology, paper, book-binding and storage conditions. but because of our work with electronic records we also take part in the standardisation work on it, linked with iso/ec jtc 1. this has been a good way of getting information in this field. but because of lack of time and technical knowledge the possibilities to contribute to the work and influence the standardisation work has been less then we would have wished. finally i would like to mention a little about the use of the electronic recwds transferred to the national archives. the basis for this is the freedom of information act which is one of the fundamental laws (constitutions) of sweden. it stipulates that all governmental information is open to the public, with the restrictions specified in the act on security (privacy act). this also includes electronic recwds. however the use of these in the national archives has been low until the transfers from statistics sweden were made. the reason is probably that the users do not know about, lack the technical knowledge or are not interested in modem records. today the main user is statistics sweden having need of their fomer records, and doing it so much that there is a transport from the national archives twice a week. but the use from others, mostly research institutions, also has increased in recent years. the researchers using the electronic records are mostly in the field of medicine and may also ask for some material originally from statistics sweden. all this is done in machine-readable form. it will thctefrae require a permit from the data inspection board if the recoixls shall contain personal information. 99% of the electronic records transferred to the national archives contains personal information. however it would be possible to get some special records without a pctmit if the personal information was excluded. there is an example of ssd having received electronic records from the national archives in that way. up to may 1993 it was very rare that someone wanted information as an out-print. but that month we got the national register of private boats as a transfer. the responsible governmental agency had to end it after a political decision. just now we receive about 25 questions a week concerning these records. mostly it has been the police, the navy or the public asking about the owners of lost boats. the information are distributed by mail, fax or telephone. today we have no possibilies for the public and the researchers to do on-line work with the electronic records. but we plan to have it in the future. there is actually a co-operation between the edp-divisions/ sections of the national archives in the nordic countries (denmark, finland, iceland, norway and sweden). that has resulted in a common project called team (availibility of electronic archival material). within the project the nwwegians are constructing a relational data base system to import data from the storage files and then using some suitable software to present the data. through this system we hope that in the future the public will be able to get direct access to and fkint-outs from our electronic records, naturally under the restrictions that are set by the act on secrecy and the data protection act 1 paper p-esenied at iassist/ifdo'93 conference, edinburgh, scotland. magnus geber, national archives of sweden, p.o. box 12541, s-10229 stockholm, sweden ph +46-8 737 6486 fax +46-8 73736474 2 the figures are taken from a questionnaire answered from september 1992 to may 1993. the numbers describe the total holdings and the annual use and form of distribution of electronic records to researchers and the public winter 1992 1/15 hagman, j.c. and bussell, h. (2022) going qual in: towards methodologically inclusive data work in academic libraries, iassist quarterly 46(2), pp. 1-15. doi: https://doi.org/10.29173/iq1022 going qual in: towards methodologically inclusive data work in academic libraries jessica hagman1 and hilary bussell2 abstract data literacy and research data services are a growing part of the work of academic libraries. data in this context is often presumed to mean only numeric data or statistics, leaving open the question of what role qualitative research plays in services and programming for research data and data literacy. in this paper, we report on the results of interviews with academic librarians about their understanding of data literacy, qualitative research, and academic library infrastructure around qualitative research. from the interviews, we propose a model of data literacy that incorporates both interpretive and instrumental elements. we conclude with suggestions for incorporating qualitative data and analysis methods into academic library programming and services around data literacy and research data. keywords data literacy, research data services, qualitative data, qualitative research, methodology introduction many academic libraries offer instruction on data literacy and research data services to help researchers collect, analyze, and manage their research data. as social sciences librarians who work frequently with researchers using a wide range of methodological approaches, we have found that work around data literacy and research data services programming often seems based on the assumption that data is inherently quantitative. this observation is supported by recent research by cain et al. (2019), pearce et al. (2019), and swygart-hobaugh (2016) on the difficulty of locating support for qualitative research in academic libraries. in this study we draw on in-depth interviews with academic librarians to examine perceptions of how data is defined in data literacy and research data services work to better understand the existing support for qualitative research and to identify spaces for developing greater methodological inclusivity. data literacy services and programming that are based on the presumption of quantitative data and post-positivist research paradigms could be failing to address the needs of researchers and students who work with qualitative or mixed data and methods of analysis, which ultimately limits the methodological and epistemological inclusiveness of data-related work in academic libraries. specifically, we ask: 1. how do academic librarians define qualitative research? 2. do academic librarians define data literacy in a way that is inclusive of qualitative data and methods of analysis? 3. how do academic librarians perceive their library’s support for qualitative research data literacy instruction and research data services? drawing on the results of these interviews, we propose a model of data literacy and research data services work that is both interpretive (related to evaluating others’ use of data) and instrumental (focused on developing skills for using data) and conclude with suggestions for incorporating support for qualitative research across data-related work in academic libraries in pursuit of greater methodological and epistemological inclusivity. while we focus here on the type of data and https://doi.org/10.29173/iq1022 2/15 hagman, j.c. and bussell, h. (2022) going qual in: towards methodologically inclusive data work in academic libraries, iassist quarterly 46(2), pp. 1-15. doi: https://doi.org/10.29173/iq1022 associated methods of analysis in qualitative research, we see these questions as linked to the broader issue of types of research valued in academic libraries and higher education. literature review data literacy and research data services in the past decade, a number of academic libraries have turned their attention towards teaching data literacy and offering research data services, driven, at least in part, by the sense that the growing amount of data available online requires new skills (corrall, 2012; throgmorton, norlander and palmer, 2019; burress, mann and neville, 2020). within library and information science, the growth in data availability and access has been seen as an opportunity for library staff to deploy their specific skills and expertise (shields, 2004; carlson et al., 2011), with the library as an ideal site of instruction around the use of data (fontichiaro et al., 2017; dai, 2019). data literacy may be embedded into individual course instruction or student projects (macmillan, 2015; beauchamp and murray, 2016; widener and slater reese, 2016). or, data literacy may serve as the basis for ongoing library programming aimed at researchers or community members interested in using existing data for analysis, or collecting and managing their own research data outputs (hogenboom, phillips and hensley, 2011; okamoto, 2017; schöpfel, prost and malleret, 2018; willaert et al., 2019). definitions of data literacy often draw on conceptions of data as solely quantitative, such as the definition offered by dechman and syms (2014), who point to the lack of training in quantitative methods among those who now have access to datasets available online. similarly hogenboom, phillips, & hensley (2011) describe a 'shift to quantitative research methods in social sciences' and define data literacy as 'the ability to read and interpret data, to think critically about statistics and to use statistics as evidence' (p. 410). this trend seems to be continuing, as a recently published title from the american library association is titled data literacy in academic libraries: teaching critical thinking with numbers (bauder, 2021). this focus on quantitative data is not universal, however. some definitions are agnostic about the nature of data in data literacy, such as prado and marzal (2013) who define data literacy as 'the component of information literacy that enables individuals to access, interpret, critically assess, manage, handle and ethically use data' (p. 126). similarly, dai (2019) takes a broad view of data literacy as 'critical thinking applied to evaluating data sources' (p. 2). dai also explicitly positions statistical literacy as a 'companion' to data literacy (p. 2). some authors take care to be inclusive about the types of data that may be relevant to data literacy, such as deahl (2014), who proposes a definition of data literacy as 'the ability to understand, find, collect, interpret, visualize, and support arguments using quantitative and qualitative data' (p. 41). qualitative research we use the phrase qualitative research to refer to any research that makes use of non-numeric data, with recognition of the difficulty of offering a succinct, yet inclusive definition that accounts for the wide range of research that is labelled qualitative (guest, namey and mitchell, 2013; aspers and corte, 2019). small (2021) offers a useful reminder that when we speak of qualitative research, we can be referring to the method of data collection, format of the data, and the analysis approach, with no requirement that all three be used in a single project. ultimately, the choice of data, collection approach, and analysis strategies rests on the researcher’s study design and methodology, which are embedded with theoretical and epistemological assumptions about how new knowledge should be generated (crotty, 1998; staller, 2013). qualitative research is frequently contrasted with investigations using quantitative data, but any type of data could be used within a particular theoretical and epistemological frame, or paradigm (crotty, 1998). https://doi.org/10.29173/iq1022 3/15 hagman, j.c. and bussell, h. (2022) going qual in: towards methodologically inclusive data work in academic libraries, iassist quarterly 46(2), pp. 1-15. doi: https://doi.org/10.29173/iq1022 qualitative research is often used in critical approaches to research, which use a variety of theoretical frames to explore the construction, maintenance and deconstruction of systems of inequality such as intersectional analyses (esposito & evans-winters, 2022), critical disability studies (minich, 2016), and the use of indigeous methodologies (lilley ,2018). these paradigms recognize that the researcher is inherently embedded in the process of data collection and analysis, and ultimately unable to develop objective observations of social worlds. research in these paradigms may even be used to challenge the idea of objectivity altogether and ‘master narratives of knowledge’ (nadar, 2014, p. 23). by contrast, quantitative research is frequently linked to positivist or post-positivist research that aims to develop ostensibly objective and often generalizable claims about the nature of the world (crotty, 1998; williams, 2000). we should be clear, however, that qualitative data and analysis methods can, however, be used in positivist or post-positivist research, particularly when researchers seek to identify causal relationships between research concepts or identifying trends in concepts that cannot be easily quantified (su, 2018). the diversity of paradigms in which research can be conducted means that there is no single way to evaluate the quality of new knowledge claims. qualitative research, however, is sometimes critiqued for not following the same criteria for rigor as work using quantitative data in positivist or post-positivist paradigms (anfara, brown and mangione, 2002; nadar, 2014). such messages can lead to researchers viewing their qualitative work positioning them as outsiders within their own discipline (benton et al., 2012; roger et al., 2018). dempsey (2018), for example, surveyed authors of qualitative works in library and information science and found that some believed peer reviewers to be unprepared for evaluating qualitative analyses and that qualitative research would be unfairly dismissed for lacking sufficient sample size or lack of predictive power when undergoing peer review. researchers who use qualitative data and analysis methods have articulated criteria on which such work can be assessed. tracy (2010) has defined eight criteria for assessing qualitative research, with the recognition that different research paradigms and methodologies place value on different criteria. likewise, bhattacharya (2017) outlines factors that contribute to assessments of rigor in qualitative work, such as the 'alignment of epistemology, theoretical frameworks, methodology, and methods, data analysis, and representation' or the use of multiple types of data (p. 23). transparency about the choices that have guided the research design process and researcher reflexivity have also been identified as important elements of rigor for studies using qualitative data or analysis outside of positivist and post-positivist paradigms (anfara, brown and mangione, 2002; pillow, 2003; davidson, thompson and harris, 2017). academic library support for qualitative research there is currently limited research on library services specifically for qualitative work. swygarthobaugh (2016) has questioned whether qualitative research is the 'jan brady' of data services, always receiving less attention than research using quantitative data, based on an analysis of data-related job ads, a survey of data services librarians, and a review of libguides on qualitative research. a more recent survey of library websites found that information on support for qualitative work is difficult to locate and can often only be discovered by searching for the names of specific, proprietary qualitative data analysis software programs (cain et al., 2019). the few existing needs analyses indicate that qualitative researchers would benefit from a research data infrastructure that supports their work throughout the research cycle, from preparing irb applications and research design to analyzing data and writing research reports in ways that meets the expectations of readers who may not work in the same research paradigm (downing et al., 2019). even when research data services are available for qualitative work, individual researchers are often https://doi.org/10.29173/iq1022 4/15 hagman, j.c. and bussell, h. (2022) going qual in: towards methodologically inclusive data work in academic libraries, iassist quarterly 46(2), pp. 1-15. doi: https://doi.org/10.29173/iq1022 unaware of such services and see the library primarily as a site of collections access rather than for active methodological learning (pearce et al., 2019). for those providing data-related services in libraries, it may not be clear what role the library or individual librarians can play in the work of qualitative researchers. librarians interviewed by downing et al. (2019) noted that they did not see addressing questions of research design as appropriate for their role. similarly, some data services librarians surveyed by swygart-hobaugh (2016) were unclear on how they could offer support for qualitative researchers given that qualitative data is not often reused, indicating a conception of the library’s role as primarily for data location rather than data analysis. in the areas where qualitative support is specifically discussed, the focus is largely on instruction around the use of qualitative data analysis software, such as swygart-hobaugh’s description of developing an nvivo workshop in collaboration with disciplinary faculty (swygart-hobaugh, 2019). similarly, røddesnes, faber, and jensen (2019) have written about the process of developing nvivo workshops in their library. hagman (2021) has proposed a model of workshop development that is centered on the qualitative data analysis strategies used by qualitative researchers, even when offering instruction on a specific software tool. thielen and hess (2017) provide one exception to this trend, as they recount offering instruction on research data management practices in the context of a graduate course in qualitative research methods. methods participants and data collection in this study, we draw on in-depth, semi-structured interviews with academic librarians about their experiences with data literacy, research data services, and qualitative research in their work. participants were recruited via our personal social media accounts, relevant email lists, and targeted outreach to librarians named as the owners of libguides about qualitative research. potential participants indicated their interest using a qualtrics form. we also used a snowball sampling technique in which we asked participants to suggest additional names for recruitment, and we followed up with suggested contacts (while not revealing the name of the participant who suggested the contact). we conducted interviews during october and november 2020. the participants were 13 academic librarians based in the united states. of this group, 11 worked at research universities, while two were based at liberal arts colleges. most of the participants (11) had subject specialist responsibilities and five had data services roles within their library. the researchers took turns facilitating the interviews, which were conducted via zoom. the zoom live transcription feature provided a rough transcript of each interview, from which we created a corrected transcript while re-viewing the recording of the interview session. each participant was given the option to pick a pseudonym or to have one chosen for them. the interview guide is listed in appendix i. we developed the interview questions based on our reading of the existing literature around data literacy and library support for qualitative research, as well as our own experiences with data literacy and qualitative work on our campuses. data analysis we followed deterding and waters’ (2021) approach for the collaborative analysis of interview data, first reading through the interviews and coding responses to each interview question, or what deterding and waters call index coding. throughout the index coding process, we noted potential concepts that were relevant to our research questions and wrote memos about our initial perceptions https://doi.org/10.29173/iq1022 5/15 hagman, j.c. and bussell, h. (2022) going qual in: towards methodologically inclusive data work in academic libraries, iassist quarterly 46(2), pp. 1-15. doi: https://doi.org/10.29173/iq1022 of the data. we drew on these memos in developing the second stage of the analysis process, which involved coding relevant parts of the transcripts using analytic codes. in some cases, index codes overlap with analytic codes, as in when we explicitly asked participants to offer a definition of data literacy, though we often found pertinent information beyond the scope of a specific interview question. the analytic codes we explored to answer the questions for this study included: 1. definitions of data literacy 2. definitions of qualitative research 3. actual support for qualitative research at the participant’s library 4. ideal support for qualitative research at the participant’s library 5. participant’s perceptions of others’ attitudes towards qualitative work we used maxqda20 for our analysis to iteratively code the data and made use of the summary feature to explore the data within each analytic code and develop an analysis of recurrig ideas and relationships between elements. for example, we were interested in the ways in which participants defined data literacy, both when explicitly asked and in other parts of the interview. by using the summary grid feature, we were able to move through all the answers coded with 'defining data literacy', for example, and pull out the descriptions of the elements that participants said are embedded in data literacy and ultimately develop the interpretive and instrumental approaches to data literacy described in our findings. in the results and discussion below we weave together data from across these analytic codes as we propose answers to our research questions and consider the implications of our analysis. interview findings defining qualitative research we asked the participants to define qualitative research and found that their responses mirrored the complexity of existing definitions in the research literature (guest, namey and mitchell, 2013; aspers and corte, 2019; small, 2021). their answers included examples of methodologies and types of data that they believed to be qualitative, actions that they considered to be part of qualitative research processes, and characteristics of qualitative research. three participants explicitly noted the difficulty of defining qualitative research, including penelope who feared that she’d get the definition 'wrong' even though she had conducted her own qualitative work, and clarence who saw qualitative and quantitative research to be closely related, often ‘mov[ing] back and forth a lot.’ in discussing methods and data types that make up qualitative research, participants frequently pointed to text as a form of data and described the collection of data through interviews, focus groups, and open-ended survey questions. overall, participants described a wide range of methodologies, including grounded theory, ethnography, interviews, and case studies. two participants drew on the concept of 'stories' in defining qualitative work, including jane who referred to qualitative research as examining 'stories, the human side of any question.' jane continued in this vein when she noted her own interest in qualitative work, as she sees qualitative data as having 'more depth, impact, nuance' in comparison to quantitative research. jane’s conception of qualitative work as deeper than quantitative mirrors irene’s and penelope’s descriptions of qualitative methods as allowing more exploration of the topic under study, in contrast with research using quantitative data. comparisons to quantitative work came up frequently among our participants who described qualitative work as 'so much more interesting' (siobhan), 'more subjective' (clarence), 'more interpretive' and 'more interactional' (anya), and 'more about developing, exploring, confirming, understanding themes' (pheobe). in describing their conceptions https://doi.org/10.29173/iq1022 6/15 hagman, j.c. and bussell, h. (2022) going qual in: towards methodologically inclusive data work in academic libraries, iassist quarterly 46(2), pp. 1-15. doi: https://doi.org/10.29173/iq1022 of qualitative research, participants pointed to actions that they saw as central to this in-depth work, including close reading and deep engagement with data. they also described qualitative research as being interested in identifying themes and patterns. defining data literacy our second research question asked how participants defined data literacy. we explicitly asked participants to define this term, but also identified implicit definitions throughout our conversations with participants. participants’ definitions included elements that we are categorizing as interpretive and instrumental. interpretive elements of data literacy definitions emphasize understanding presentations of data by others, while instrumental elements focus on skill-building for an individual's own use of data. we use the term elements here, because most of the participants offered definitions that included both interpretive and instrumental aspects, an indication of the complexity of defining a concept like data literacy. we can see the interpretive elements of data literacy in participants’ focus on identifying the context of data. understanding context refers to examining the provenance of a particular dataset or presentation of data, as well as understanding the general historical, social, and disciplinary landscapes in which different types of data are created. for example, david’s definition of data literacy included: 'data is contextual...it’s a bit connected to the researcher who collects it and presents it.' for david, data literacy instruction would involve teaching students to 'reflect on why this data exists and its purpose;' a process he related to the research as inquiry frame in the acrl framework. arthur echoed this understanding when he defined data literacy as 'thinking more about the context in which that data was collected, the context in which it was curated, and the potential ethical ramifications of that sort of surrounding context of the data set.' similarly, irene pointed to the importance of knowing the ‘history’ of a data set to identify any potential biases. beyond the idea of identifying context, we can also see interpretive definitions that focus on identifying and avoiding the use of information that has been improperly manipulated. cathy saw data literacy as an important skill for students to learn because 'as you know, statistics lie. you can make them say whatever you want. so, data literacy would be that aspect of letting the data be the data.' like cathy, david expressed concern about the misuse of statistics and said, 'you hear lots of stuff about people lying with statistics and data,' particularly in the context of social media posts in an election year. in a similar vein, tracy related data literacy to trusting interpretations of data and learning to 'see what the original underlying data looks like so i can actually trust this statement.' instrumental elements of data literacy refer to helping patrons learn to 'work with data' (anya, jane, and tracy all used this phrase) for their own purposes, usually in the context of academic research. examples of activities that fall into this category come from across the research cycle beginning with developing questions that can be answered with data, as we can see in sally’s definition of data literacy as 'a facility and the ability to ask data related questions…' participants also frequently mentioned the importance of data management skills, including making decisions about ethically and securely storing data from human subjects. ethics also played an important role in conversations about building skills in presenting data. phoebe, for example, wanted to explore presenting data in 'effective, clear, authentic, like honest kinds of ways' while still making presentations of data 'visually interesting.' cathy, who emphasized the importance of interpreting data to avoid being misled by statistics, also described data literacy as the ability to 'present data in an ethical, truthful way.' throughout our interviews, we found that participants shared a belief that other people view data literacy as solely linked to numeric data, particularly when we asked them whether their definitions of data literacy were in sync with their colleagues’ understanding of the concept. veronica saw https://doi.org/10.29173/iq1022 7/15 hagman, j.c. and bussell, h. (2022) going qual in: towards methodologically inclusive data work in academic libraries, iassist quarterly 46(2), pp. 1-15. doi: https://doi.org/10.29173/iq1022 colleagues as tending to ‘jump to quantitative’ in conversations about data work within the library. clarence said that her colleagues see data as only numeric content, a 'traditional' but 'limited' view, 'considering how today’s individual’s work. we move between numbers and other symbolic meaning.' likewise, tracy saw subject liaisons with which she collaborates with as focusing narrowly on coding, using big data or r (the programming language) in discussions of data literacy, in contrast with her own department which is more broadly focused on data education. data literacy work in support of qualitative research our final research question asks how participants perceive their library’s data literacy work in support of qualitative research. we asked participants both to describe their existing resources and services for qualitative research and to share what they might like to do if limits on time and other resources were eased. the range of resources and services available for qualitative researchers varied greatly, with some participants describing groups who support qualitative work and others indicating that there is almost no campus or library infrastructure specific to qualitative research. participants also pointed to barriers that they perceive limits attention to qualitative research on campus, particularly the campus and disciplinary valuing of research using quantitative data rather than the use of qualitative data. access to qualitative data analysis software and instruction on use of these tools was the most frequently mentioned way that participants’ libraries address data literacy for qualitative work and support qualitative researchers. david, for example, told us that his campus had recently 'invested' in nvivo, and jane said that she taught workshops and offered consultations on the use of this proprietary qualitative data analysis software. in cathy’s case, she believed she was the only person on campus who could advise on the use of qualitative data analysis software. cathy expected to retire a few months after our interview, meaning that the campus would be without support for the use of software for qualitative analysis. by contrast, sally described a five-person 'qualitative user group' that works on issues related to qualitative research, providing workshops and individual consultations, on the use of both nvivo and the open-source tool taguette. like sally, veronica, arthur, and irene all described more robust programming that supports qualitative research. veronica partners with disciplinary faculty to teach qualitative data analysis software in a way that emphasizes features of the software and the ways it can be used within qualitative methodologies. she sees the outcomes of these workshops helping students learn the logic of the tool and countering the mistaken notion that qualitative data analysis software has a 'magic button' that can output research results without extensive analysis from the researcher. similarly, arthur’s library provides tutorials on multiple software programs for qualitative data analysis, as well as instruction around collecting data. interestingly, arthur also described trying to 'slip in' aspects of data literacy into tutorials on qualitative data analysis tools, which are ostensibly focused on the use of the software, as data literacy is 'not necessarily one of the direct research goals or learning goals.' irene also mentioned that she seeks to integrate data literacy concepts into software tutorials, out of recognition that the concept of data literacy is linked to the library, and not seen as a goal of most patrons. many participants, including those whose libraries offer little to no support for qualitative research as well as those with robust programming in this area, pointed to the value, attention, and resources given to research using numeric data on their campuses, in contrast to the lack of attention given to qualitative work. tracy, for example, believed that qualitative research was 'ignored' in favor of the 'shiny toy' of open data for stem disciplines. while jane teaches workshops on nvivo, she thought https://doi.org/10.29173/iq1022 8/15 hagman, j.c. and bussell, h. (2022) going qual in: towards methodologically inclusive data work in academic libraries, iassist quarterly 46(2), pp. 1-15. doi: https://doi.org/10.29173/iq1022 that her flagship state university should offer more resources for qualitative researchers and better coordinate what services are available, a situation she attributed to both the lack of state funding and an under-valuing of qualitative work. even sally, whose library has a group of staff supporting qualitative work, noted that qualitative research can be seen as 'niche' work that doesn’t quite fit within the support offerings for computational methods in research centers around campus. discussion and recommendations in this research project, we have explored how practicing academic librarians understand qualitative research, define data literacy, and perceive their library’s data literacy work and data-related services as addressing the needs of qualitative researchers. analysis of these interviews indicates that while our participants see data literacy as theoretically inclusive of qualitative work and would ideally like to see services developed that support qualitative research, they identify barriers to the full development of a data literacy and research support infrastructure that addresses qualitative work. in exploring participants’ definitions of data literacy, we find both interpretive and instrumental elements, often from the same participant. while this distinction points to the complexity of data literacy as a concept, we contend that approaching data literacy with these two elements in mind means we can draw out more specific ways to build robust infrastructures that are inclusive of the wide variety of methodologies and epistemological approaches used on our campuses. interpretive elements of data literacy emphasize the importance of learning to understand the context in which presentations of data were created in order to avoid being manipulated by improper uses of data. while the participants do not use the term 'interpretive' we see the focus on understanding the context of research and evaluating presentations of data as interest in the rigor of research processes. there is, however, no single way to evaluate the rigor of research, particularly the diversity of research using qualitative data and analysis (tracy, 2010; staller, 2013; bhattacharya, 2017). viewing data literacy as a tool for avoiding being duped by ill-gotten statistics limits the scope and potential of this concept. we contend that data literacy instruction should, explicitly and intentionally, address broader questions of research rigor with recognition of the multiple ways in which we can evaluate any claims to new knowledge. instruction should include examples of research that make sense of qualitative data in a variety of methodologies and paradigms. given how often our participants pointed to others’ understanding of data as solely quantitative and the under-valuing of qualitative research, we recognize the potential difficulties of countering such discourses. maintaining the status quo, however, means capitulating to narrow conceptions of research and ignoring the wealth of knowledge that draws from paradigms that value deep and reflexive exploration of concepts and experiences. in addition to interpretive elements of data literacy, we also identified instrumental elements, or the resources, services, and instruction provided for researchers who are collecting, analyzing, managing, and presenting their own data. participants frequently noted that they believe qualitative research is less valued on their campuses, in the disciplines they work with, and even among colleagues in the library and in the library and information science field. while additional research is needed to understand whether patrons share this perception, this study encourages us to think about how our assumptions about what it means to do research and work with data are embedded in our services and instruction. further research may also consider the impact of broader cultural understandings of how research data is used to understand the world, including in popular media. whether the limited infrastructure in support of qualitative research stems from a lack of resources or uneven valuing of research approaches, we believe that services for researchers can be offered and https://doi.org/10.29173/iq1022 9/15 hagman, j.c. and bussell, h. (2022) going qual in: towards methodologically inclusive data work in academic libraries, iassist quarterly 46(2), pp. 1-15. doi: https://doi.org/10.29173/iq1022 presented in ways that are intentionally inclusive of diverse research paradigms. for some libraries, inclusive research infrastructure may require the investment in software and other tools for collecting and analyzing data. library staff members may need additional training or the resources to explore what it means to conduct qualitative research to provide these services, since mls graduates working in academic libraries often report that their program did not prepare them to conduct original research (kennedy and brancolini, 2018). but the expenditure of additional resources may not be sufficient to develop a more inclusive approach to data literacy and research data services; services may need to be framed in new ways. given the difficulty of identifying resources for qualitative work on academic library websites (cain et al., 2019) it may be that researchers using qualitative data may not see themselves as the likely customers for research data services provided by academic libraries. or, they may share the idea that qualitative research materials do not necessarily constitute the types of data that requires management. to counter this assumption will require targeted outreach to qualitative researchers and updating the language used to describe data services programming to include examples of qualitative data and analysis approaches. such outreach does not necessarily require expertise in specific methodologies. instead, we would argue that librarians’ skills in working across disciplinary boundaries in public services roles positions them well to develop an understanding of the diversity of research conducted qualitatively and the ongoing campus and disciplinary conversations about what constitutes rigor in academic research. exploring these discourses may be a fruitful area of collaboration for those working in data services and subject specialists who can bring knowledge of the types of research conducted within and across academic disciplines. conclusion academic libraries are a vital part of the research processes, through both the provision of existing information as well as through resources and services that support researchers in developing new knowledge. research across our campuses is conducted within a wide array of epistemological paradigms and using many different methodologies. too often, conversations around data within academic libraries emphasize numerical data, which ultimately limits the potential audience for these services to those who see their own work in library communications. this imbalance is more than just a matter of fairness to library patrons. academic libraries must consider how their approaches to data literacy and research data services can serve to limit or expand the notion of what counts as research, and even who can bring their research knowledge to the scholarly conversation. developing services that are ostensibly open to all library patrons but ultimately only serve those whose research uses only one type of data sends a message about what kinds of research are valued and worth supporting. by critically examining the explicit and implicit messages we share about knowledge creation processes, academic libraries have the opportunity to consider their broader role in higher education and even wider social systems (honma & chu, 2018). while we focus here on the use of qualitative data and methods broadly, we recognize that further research is needed to explore the complexity of the scholarship and scholars in this broad category. future work to understand how to support the work of critical scholarship will be of particular value given the power of these frameworks to explain and develop responses to the persistent and systemic inequalities in our social systems (esposito & venus-williams, 2022; guyen, 2022; minich, 2016). the challenges of our collective future require that we embrace diverse approaches to building new knowledge. academic libraries can be better partners in the process of discovery and innovation by https://doi.org/10.29173/iq1022 10/15 hagman, j.c. and bussell, h. (2022) going qual in: towards methodologically inclusive data work in academic libraries, iassist quarterly 46(2), pp. 1-15. doi: https://doi.org/10.29173/iq1022 recognizing and affirming the value of diverse approaches to research and moving toward a more inclusive data work. references anfara, v.a., brown, k.m. and mangione, t.l. (2002) ‘qualitative analysis on stage: making the research process more public’, educational researcher, 31(7), pp. 28–38. doi: http://www.jstor.org/stable/3594403. aspers, p. and corte, u. (2019) ‘what is qualitative in qualitative research’, qualitative sociology, 42(2), pp. 139–160. doi: https://doi.org/10.1007/s11133-019-9413-7. bauder, j. (ed.) (2021) data literacy in academic libraries: teaching critical thinking with numbers. chicago: ala editions. beauchamp, a. and murray, c. (2016) ‘teaching foundational data skills in the library’, in kellam, l.m. and thompson, k. (eds) databrarianship : the academic data librarian in theory and practice. chicago, il: association of college and research libraries, pp. 81–92. benton, a.d. et al. (2012) ‘of quant jocks and qual outsiders: doctoral student narratives on the quest for training in qualitative research’, qualitative social work, 11(3), pp. 232–248. doi: https://doi.org/10.1177/1473325011400934. bhattacharya, k. (2017) fundamentals of qualitative research. new york: routledge. available at: https://www.taylorfrancis.com/books/9781351865982/chapters/10.4324/9781315231747-2 (accessed: 12 august 2021). burress, t., mann, e. and neville, t. (2020) ‘exploring data literacy via a librarian-faculty learning community: a case study’, the journal of academic librarianship, 46(1). doi: https://doi.org/10.1016/j.acalib.2019.102076. cain, j. et al. (2019) ‘where is qda hiding? an analysis of the discoverability of qualitative research support on academic library websites’, iassist quarterly, 43(2), pp. 1–9. doi: https://doi.org/10.29173/iq957. carlson, j. et al. (2011) ‘determining data information literacy needs: a study of students and research faculty’, portal: libraries and the academy, 11(2), pp. 629–657. doi: https://doi.org/10.1353/pla.2011.0022. corrall, s. (2012) ‘roles and responsibilities: libraries, librarians and data’, in pryor, g. (ed.) managing research data. facet publishing, pp. 105–133. crotty, m. (1998) the foundations of social research: meaning and perspective in the research process. thousand oaks, ca: sage. dai, y. (2019) ‘how many ways can we teach data literacy?’, iassist quarterly, 43(4), pp. 1–11. doi: https://doi.org/10.29173/iq963. https://doi.org/10.29173/iq1022 http://www.jstor.org/stable/3594403 https://doi.org/10.1007/s11133-019-9413-7 https://doi.org/10.1177/1473325011400934 https://www.taylorfrancis.com/books/9781351865982/chapters/10.4324/9781315231747-2 https://doi.org/10.1016/j.acalib.2019.102076 https://doi.org/10.29173/iq957 https://doi.org/10.1353/pla.2011.0022 https://doi.org/10.29173/iq963 11/15 hagman, j.c. and bussell, h. (2022) going qual in: towards methodologically inclusive data work in academic libraries, iassist quarterly 46(2), pp. 1-15. doi: https://doi.org/10.29173/iq1022 davidson, j., thompson, s. and harris, a. (2017) ‘qualitative data analysis software practices in complex research teams: troubling the assumptions about transparency and portability’, qualitative inquiry, 23(10), pp. 779–788. doi: https://doi.org/10.1177/1077800417731082. deahl, e.s. (2014) better the data you know: developing youth data literacy in schools and informal learning environments. ma thesis. massachusetts institute of technology. available at: http://www.ssrn.com/abstract=2445621 (accessed: 28 october 2019). dechman, m.k. and syms, l.r. (2014) ‘working together to maximize the utilization of open data across social science and professional disciplines’, behavioral & social sciences librarian, 33(4), pp. 188–207. doi: https://doi.org/10.1080/01639269.2014.964617. dempsey, p.r. (2018) ‘how lis scholars conceptualize rigor in qualitative data’, portal: libraries and the academy, 18(2), pp. 363–390. doi: https://doi.org/10.1353/pla.2018.0020. deterding, n.m. and waters, m.c. (2021) ‘flexible coding of in-depth interviews: a twenty-firstcentury approach’, sociological methods & research, 50(2), pp. 708–739. doi: https://doi.org/10.1177/0049124118799377. downing, k. et al. (2019) ‘capturing the narrative: understanding qualitative researchers’ needs and potential library roles’, in mueller, d.m. (ed.) recasting the narrative: the proceedings of the acrl 2019 conference. association of college and research libraries, cleveland, oh: association of college & research libraries, pp. 163–175. esposito, j. and evans-winters, v.e. (2022) introduction to intersectional qualitative research. thousand oaks, ca: sage. fontichiaro, k. et al. (eds) (2017) data literacy in the real world: conversations & case studies. michigan publishing. guest, g., namey, e.e. and mitchell, m.l. (2013) collecting qualitative data: a field manual for applied research. london: sage. guyan, k. (2022) queer data: using gender, sex and sexuality data for action. bloomsbury academic. hagman, j. (2021) ‘centering analysis strategies and open tools for qualitative data analysis’, in mueller, d.m. (ed.) proceedings of the association of college & research libraries conference. association of college & research libraries, chicago: association of college & research libraries, pp. 394–401. hogenboom, k., phillips, c.m.h. and hensley, m. (2011) ‘show me the data! partnering with instructors to teach data literacy’, in mueller, d.m. (ed.) declaration of interdependence: the proceedings of the acrl 2011 conference. association of college & research libraries, chicago: association of college & research libraries, pp. 410–417. honma, t. m., & chu, c. m. (2018). positionality, epistemology, and new paradigms for lis: a critical dialog with clara m. chu. in r. l. chou & a. pho (eds.), pushing the margins: women of color and intersectionality in lis (pp. 447–465). library juice press. https://doi.org/10.29173/iq1022 https://doi.org/10.1177/1077800417731082 http://www.ssrn.com/abstract=2445621 https://doi.org/10.1080/01639269.2014.964617 https://doi.org/10.1353/pla.2018.0020 https://doi.org/10.1177/0049124118799377 12/15 hagman, j.c. and bussell, h. (2022) going qual in: towards methodologically inclusive data work in academic libraries, iassist quarterly 46(2), pp. 1-15. doi: https://doi.org/10.29173/iq1022 kennedy, m.r. and brancolini, k.r. (2018) ‘academic librarian research: an update to a survey of attitudes, involvement, and perceived capabilities’, college & research libraries, 79(6), pp. 822–851. doi: https://doi.org/10.5860/crl.79.6.822. macmillan, d. (2015) ‘developing data literacy competencies to enhance faculty collaborations’, liber quarterly: the journal of european research libraries, 24(3), pp. 140–160. doi: https://doi.org/10.18352/lq.9868. minich, j.a. (2016) ‘enabling whom? critical disability studies now’, lateral, 5(1). doi: https://doi.org/10.25158/l5.1.9. nadar, s. (2014) ‘“stories are data with soul” – lessons from black feminist epistemology’, agenda, 28(1), pp. 18–28. doi: https://doi.org/10.1080/10130950.2014.871838. okamoto, k. (2017) ‘introducing open government data’, reference librarian, 58(2), pp. 111–123. doi: https://doi.org/10.1080/02763877.2016.1199005. pearce, a. et al. (2019) ‘qualifying for services: investigating the unmet needs of qualitative researchers’, in baughman, s. et al. (eds) proceedings of the 2018 library assessment conference: building effective, sustainable, practical assessment. library assessment conference—building effective, sustainable, practical assessment, association of research libraries, pp. 321–333. doi: https://doi.org/10.29242/lac.2018.29. pillow, w. (2003) ‘confession, catharsis, or cure? rethinking the uses of reflexivity as methodological power in qualitative research’, international journal of qualitative studies in education, 16(2), pp. 175–196. doi: https://doi.org/10.1080/0951839032000060635. prado, j. and marzal, m.á. (2013) ‘incorporating data literacy into information literacy programs: core competencies and contents’, libri: international journal of libraries & information services, 63(2), pp. 123–134. doi: https://doi.org/10.1515/libri-2013-0010. røddesnes, s., faber, h.c. and jensen, m.r. (2019) ‘nvivo courses in the library: working to create the library services of tomorrow’, nordic journal of information literacy in higher education, 11(1), pp. 27–38. doi: https://doi.org/10.15845/noril.v11i1.2762. roger, k. et al. (2018) ‘exploring identity: what we do as qualitative researchers’, the qualitative report, 23(3), pp. 532–546. doi: https://doi.org/10.46743/2160-3715/2018.2923. schöpfel, j., prost, h. and malleret, c. (2018) ‘research and development in the field of research data and dissertations. the d4humanities project at the university of lille (france)’, grey journal (tgj), 14, pp. 30–36. shields, m. (2004) ‘information literacy, statistical literacy and data literacy’, iassist quarterly, 28(2– 3), pp. 6–11. doi: https://doi.org/10.29173/iq790. small, m.l. (2021) ‘what is “qualitative” in qualitative research? why the answer does not matter but the question is important’, qualitative sociology, 44, pp. 567–574. doi: https://doi.org/10.1007/s11133-021-09501-3. https://doi.org/10.29173/iq1022 https://doi.org/10.5860/crl.79.6.822 https://doi.org/10.18352/lq.9868 https://doi.org/10.25158/l5.1.9 https://doi.org/10.1080/10130950.2014.871838 https://doi.org/10.1080/02763877.2016.1199005 https://doi.org/10.29242/lac.2018.29 https://doi.org/10.1080/0951839032000060635 https://doi.org/10.1515/libri-2013-0010 https://doi.org/10.15845/noril.v11i1.2762 https://doi.org/10.46743/2160-3715/2018.2923 https://doi.org/10.29173/iq790 https://doi.org/10.1007/s11133-021-09501-3 13/15 hagman, j.c. and bussell, h. (2022) going qual in: towards methodologically inclusive data work in academic libraries, iassist quarterly 46(2), pp. 1-15. doi: https://doi.org/10.29173/iq1022 staller, k.m. (2013) ‘epistemological boot camp: the politics of science and what every qualitative researcher needs to know to survive in the academy’, qualitative social work: research and practice, 12(4), pp. 395–413. doi: https://doi.org/10.1177/1473325012450483. su, n. (2018) ‘positivist qualitative methods’, in cassell, c., cunliffe, a., and grandy, g., the sage handbook of qualitative business and management research: history and traditions. london: sage, pp. 17–31. doi: https://doi.org/10.4135/9781526430212.n2. swygart-hobaugh, m. (2016) ‘qualitative research and data support: the jan brady of social sciences data services?’, in kellam, l. and thompson, k. (eds) databrarianship: the academic data librarian in theory and practice. chicago: association of college & research libraries, pp. 153–178. swygart-hobaugh, m. (2019) ‘bringing method to the madness: an example of integrating social science qualitative research methods into nvivo data analysis software training’, iassist quarterly, 43(2), pp. 1–16. doi: https://doi.org/10.29173/iq956. thielen, j. and hess, a.n. (2017) ‘advancing research data management in the social sciences: implementing instruction for education graduate students into a doctoral curriculum’, behavioral & social sciences librarian, 36(1), pp. 16–30. doi: https://doi.org/10.1080/01639269.2017.1387739. throgmorton, k., norlander, b. and palmer, c. (2019) ‘open data literacy and the library’, alki, 35(2), pp. 27–29. tracy, s.j. (2010) ‘qualitative quality: eight “big-tent” criteria for excellent qualitative research’, qualitative inquiry, 16(10), pp. 837–851. doi: https://doi.org/10.1177/1077800410383121. widener, j.m. and slater reese, j. (2016) ‘mapping an american college town: integrating archival resources and research in an introductory gis course’, journal of map & geography libraries, 12(3), pp. 238–257. doi: https://doi.org/10.1080/15420353.2016.1195783. willaert, t. et al. (2019) ‘research data management and the evolutions of scholarship: policy, infrastructure and data literacy at ku leuven’, liber quarterly: the journal of european research libraries, 29(1), pp. 1–19. doi: https://doi.org/10.18352/lq.10272. williams, m. (2000). interpretivism and generalisation. sociology, 34(2), 209–224. https://doi.org/10.1177/s0038038500000146. https://doi.org/10.29173/iq1022 https://doi.org/10.1177/1473325012450483 https://doi.org/10.4135/9781526430212.n2 https://doi.org/10.29173/iq956 https://doi.org/10.1080/01639269.2017.1387739 https://doi.org/10.1177/1077800410383121 https://doi.org/10.1080/15420353.2016.1195783 https://doi.org/10.18352/lq.10272 https://doi.org/10.1177/s0038038500000146 14/15 hagman, j.c. and bussell, h. (2022) going qual in: towards methodologically inclusive data work in academic libraries, iassist quarterly 46(2), pp. 1-15. doi: https://doi.org/10.29173/iq1022 appendix i 1. please describe the major roles and responsibilities of your current position. a. how long have you worked in this position? 2. what sort of work around data literacy (if any) happens at your library? a. potential follow up probes: i. who works on data literacy? does this work come out of one department? or is it something addressed across the library? ii. ask for details on specific programs, instructional approaches, etc to get a full idea of what exactly the library is doing. iii. how long has your library been involved with data literacy? has the work changed over time? b. what is your role in addressing data literacy at your library? i. how did you come to this role? (e.g. part of your job when you were hired, or something you took on) c. does the library work with anyone else on campus to address data literacy? i. if yes, which campus partners? how well do you think this collaboration works? ii. are there any campus partners you would like to see more work with in terms of data literacy? d. does any of your library’s work around data literacy explicitly address using data ethically? e. is there anything else you’d like to see your library do to address data literacy, perhaps if there were fewer limitations on time and resources? f. are there other libraries doing work around data literacy that you admire? what are they doing that you think works well? 3. how would you define data literacy? a. do you feel like your definition fits with your colleagues’ definitions? b. what do you think are the most important elements of data literacy? c. has your understanding of data literacy changed at all over time? 4. what is your understanding of what it means to do qualitative research? 5. does your library have any services or programming specifically aimed at qualitative researchers? 6. have you done any research work of your own that you consider to be qualitative? a. if yes, follow up probes: ask for types of methods and data for recent projects. b. what was it like to manage the data for your qualitative research projects? c. what were ethical considerations around the use of qualitative data that you encountered in your own research? 7. do you work with researchers who use qualitative methods in your individual role? (for example, as a subject specialist). a. in what capacity do you work with qualitative researchers? what type of researcher are those researchers doing (e.g. frequently used methods, disciplines). b. how much is working with qualitative researchers part of your job? c. how did you come to this role? (e.g. part of your job when you were hired, or something you took on). 8. how do you or your library address data literacy for qualitative researchers and if so, how do you do that? 9. is there anything you’d like to do to address data literacy for qualitative researchers, in an ideal scenario. 10. have you worked with any researchers (or conducted research yourself) on projects that involved re-using data? https://doi.org/10.29173/iq1022 15/15 hagman, j.c. and bussell, h. (2022) going qual in: towards methodologically inclusive data work in academic libraries, iassist quarterly 46(2), pp. 1-15. doi: https://doi.org/10.29173/iq1022 a. if yes, could you talk about what that process was like for you as a researcher? b. and/or what did you perceive that process to be like for the researchers you were working with? 11. have you ever talked about research data management practices with library patrons? a. if yes, have you discussed research data management practices with qualitative researchers? b. if yes, what did you talk about? 12. do you think there are differences in how research data management is addressed, or should be addressed for those conducting qualitative research, compared to other types of research? 13. is there anything else you think we could consider as we talk about data literacy for qualitative research? 14. is there anyone else you think we should reach out to for this study? endnotes 1 jessica hagman is social sciences research librarian, university of illinois at urbana-champaign, email: jhagman@illinois.edu 2 hilary bussell is associate professor and head of the humanities and social sciences librarians at ohio state university libraries. email: bussell.21@osu.edu https://doi.org/10.29173/iq1022 mailto:jhagman@illinois.edu mailto:bussell.21@osu.edu iassist quarterly spring summer 2009 23 by metadata creation, transformation and discovery for social science data management: the dames project infrastructure abstract this paper discusses the use of metadata, underpinned by ddi (data documentation initiative), to support social science data management. social science data management refers broadly to the discovery, preparation, and manipulation of social science data for the purposes of research and analysis. typical tasks include recoding variables within a dataset, and linking data from different sources. a description is given of the dames project (data management through e-social science), a uk project which is building resources and services to support quantitative social science data management activities. dames provides generic facilities for performing (and recording) operations on data. specific resources include support for analysis through micro-simulation, and support for access to specialist data on occupations, educational qualifications, measures of ethnicity and immigration, social care, and mental health. the dames project tools and services can generate, use, transform, and search metadata that describe social science datasets (including microdata from social survey datasets and aggregate-level macrodata). on dames, these metadata are described by various standards including ddi version 2, ddi version 3, jsdl (job submission definition language), and the purpose-designed jfdl (job flow definition language). the paper describes how dames uses metadata with a range of resources that are integrated with a job execution infrastructure, a web portal, and a tool for data fusion. 1. introduction this paper discusses the use of metadata within the infrastructural provisions of the uk-based dames project (data management through e-social science, www.dames. org.uk). the remit of this project is described, along with the requirements for data services and the role of metadata in this work. a review [1] has been conducted of existing social science metadata standards of relevance to the dames project. this concludes that ddi (data documentation initiative [12]) is the obvious standard with which to coordinate the use of social science metadata. the authors also observed that the transition from ddi version 2 to ddi version 3 supported many additional metadata structures which in principle would assist in the documentation of social science datasets relevant to dames. this paper reports on progress in implementing metadata tools within dames, outlining prototype services and the means by which they exploit ddi and other metadata formats. 1.1 the dames project approach to data management in social science data management is taken to mean the discovery, preparation, and manipulation of social science data for the purposes of research and analysis. more particularly, this refers to social survey research. here, data management is typically undertaken by applied researchers themselves (e.g., a secondary survey data analyst), and/or by a data distributor (e.g., a survey collection agency). typical data management tasks include recoding variables within a dataset, harmonising and standardising variables, and linking data from different sources in order to enhance analysis. data management tasks are highly significant in social science research. they account for a very substantial proportion of the time spent undertaking research. they have the potential to fundamentally change the conclusions from an analysis. they are also critical to the replicability (or otherwise) of the research workflow (e.g., [6]). there is nevertheless relatively little methodological literature on the topic of data management (at least, compared to more voluminous methodological material on collection and analysis of social science data). existing resources include instruction materials on good practices for data management in the context of specific relevant software packages (e.g., [9, 10]). some internet sites and publications offer advice on dealing with specific data resources, variables, and measures (e.g., the uk survey resources network, surveynet.ac.uk). capacity-building activities also offer training and advice (e.g., [11]). our observation is that most such materials are ad hoc, in the sense that they deal with specific problems and requirements, and are difficult for non-specialist researchers to engage with. the dames philosophy is that progress would be made in social science research if jesse m. blum, guy c. warner, simon b. jones, paul s. lambert, alison s. f. dawson, koon leai larry tan, kenneth j. turner1 24 iassist quarterly spring summer 2009 data management aspects of the research process could be followed more systematically by applied researchers. dames is working to make this engagement much easier than is presently common. dames is therefore creating tools and services for management of social survey data. the facilities should be available for general use by applied researchers, and they should support and promote good practice in data management operations. good practice includes providing clear documentation for data management tasks. for example, reusable traces should be defined for the data preparation commands used in a mainstream software package (e.g., [10]). good practice also includes making use of previous research efforts in the field. for example, there should be consideration of the ways in which outcome measures have been operationalised in previous studies, e.g., recoding or numeric standardisation, and implementing those existing operationalisations within new studies. the work on dames is contributing online infrastructural services that are accessible to most social science researchers, and that address the challenges of documentation and engagement with data management in previous research. these goals are being pursued through three strategies: 1. developing web sites and portal systems (using liferay [13]) that offer easy-to-use points of entry into dames services 2. developing online services for researchers to deposit and search heterogeneous data management information resources (dmirs) 3. developing online services that allow researchers to undertake relatively challenging data management tasks (such as matching complex data files) and to document these structured metadata records are central to delivering these strategies for two reasons. first, they are essential to providing adequate, searchable records for the complex, heterogeneous data management information resources relevant to (2). dmirs can take many forms such as quantitative data tables (e.g., aggregate statistics on occupational titles), command files for specific statistical packages, and unformatted notes and documentation. second, structured metadata records can provide the point of connection between the dmirs in (2) and the services needed for (1) and (3). the web site and portal framework can use structured metadata to selectively display suitable information about data resources. the documentation of a data management task itself requires metadata about the component datasets used. the challenge of generating effective online services in dames is increased by the very large volume of data management information resources that are relevant and interesting in social science research. as examples, dames is supporting specialist research domains including the study of occupations, educational qualifications, ethnicity and immigration, social care, and mental health inequalities (www.dames.org.uk/themes. html). the dames infrastructure thus provides access to a selected range of data management information resources. the principal challenges are: • to provide (and document) repeatable data management activities • to provide accessible guidance for using complex, heterogeneous data resources • to provide accessible guidance for otherwise neglected data management techniques (e.g., standardisation of variables, matching data files) • to facilitate and encourage data management activities of a higher standard than is currently common a data curation tool has been developed to collect and maintain metadata about information resources. a data fusion tool has been developed to support data management tasks such as linking data files by exploiting relevant metadata. these are discussed later, focusing on the use of structured metadata. 1.2 the geode project approach to metadata the geode project (grid-enabled occupational data environment, www.geode.stir.ac.uk) was a precursor to dames, focusing exclusively on helping social science researchers obtain better access to occupational information resources. the dames project is continuing this service, whilst generalising it to support other forms of data management information resources. geode collected a diverse range of occupational information resources, and created metadata descriptions for them. this was done by first asking a data depositor to complete a small online form describing the data. this information was stored in a simple xml metadata format. the information could then be supplemented by a much more substantial set of metadata, in ddi 2.1 format, which was added manually to the pool of occupational metadata. a subset of the ddi 2.1 tags was used to document these records, though more tags could be accommodated. the tags are defined by the geode-m schema (www.geode. stir.ac.uk/geode_m_curation.html). the principal structural ddi 2.1 tags are all used (<codebook>, <docdscr>, <stdydscr>, <filedscr>, <datadscr> and <othermat>), as well as most of their immediate sub-elements (such as iassist quarterly spring summer 2009 25 <citation>, <docsrc>, <stdyinfo>, <vargrp> and so on). however, many of the more detailed or application-specific sub-elements were not needed in geode -<topcclas>, <westbl> and <qstn> to name a few. geode-m was used to perform occupational matching on micro-datasets. this links values of semantically equivalent variables with the associated occupational information resource resulting in new classification variable values (e.g., for another occupational scheme), thereby helping researchers to prepare their analyses. references [7] and [8] describe the process of adding metadata to occupational resources and the use of metadata within geode services. dames is extending the work of geode by generating services for several new specialist information resources [14]. dames is also developing a number of new resources that feature significant data management requirements. reference [1] observed that the geode metadata model is likely to be too restrictive to apply readily to other domains. a more generalised approach is needed to support the multiple and highly varied domains in dames that require a wider range of data management tasks. automated metadata construction is desirable given the large volume of datasets. it is also desirable to support the curation of more complex data formats, given the wide range of data management information resources. 2. infrastructure overview as motivated in the preceding section, the dames infrastructure centres upon supporting access to a selected range of data management information and data manipulation resources. each of these resources is described using searchable metadata for improving resource discovery by social scientists. this section outlines the architecture and operation of the dames infrastructure. to facilitate accessibility to social science researchers, dames provides a grid-based infrastructure accessed through an online portal, with a familiar and comfortable mode of working for the researcher. through the portal, the researcher is able to use portlets customised for discovering and accessing specialised information resources [14], and for carrying out data management activities such as data curation and data fusion. figure 1 shows an overview of the dames infrastructure. the infrastructure provides the user-facing portlets with service-oriented access to resources stored at (potentially) physically distributed locations: • datasets may be held in external repositories or uploaded to the dames file store, with curation information entered into the metadata store. • the metadata store is provided by an exist xml database [15]. this is populated with curation information concerning explicitly uploaded datasets and remote datasets. more importantly, in the context of this paper, the metadata store holds details of datasets resulting from data management activities (identities of input datasets, processing activities carried out, the rationale for the activity). this use of metadata is discussed in more detail in the following sections. • the file store is supported by irods [16]. this provides transparent access to distributed file store through an advanced network and server. the file store is populated with uploaded datasets, and also the results of data management activities such as data fusion and standardisation of variables. • the search services allow discovery of curated datasets in both the file store and external locations, through access to the metadata store. this aspect of the dames infrastructure is not yet fully developed; it is planned to offer advanced searches based on semantic web concepts. • the compute resources would typically be in a condor pool [17] to which tasks are submitted for carrying out requested data management activities. these requests are expressed in jfdl (job flow definition language [5]), an extended version of jsdl (job submission definition language [3]) described in the following sections. the job flows themselves become part of the metadata documenting the resultant datasets. • enact fusion is one example of the data manipulation services in the dames infrastructure. these services receive task descriptions created by the researcher through the portal for searching and task customisation. this results in fetching relevant datasets from local and/or external resources through the file access service, submitting jobs to compute resources for execution, returning result datasets to the file store, and curating information in the metadata store as appropriate. the following scenario illustrates common usages of the dames infrastructure. a social science researcher wishes to fuse scottish household survey data with privately collected study data on internet usage. having registered with the uk data archive and obtained the appropriate end user license from them, the researcher uses the data curation and data fusion tools available via the dames portal to upload the data and generate a derived dataset. the metadata about this derived dataset is then made public through the portal. another researcher now searches the portal for data related to the scottish household survey, 26 iassist quarterly spring summer 2009 and finds the metadata file related to this derived dataset. realising that this would be suitable, the researcher obtains a licence and is then able to access the derived dataset. the derived dataset might then be re-purposed, using the dames tools to fuse it with data from a third study. it is important to note that the dames infrastructure provides only tools and support to the researcher, and that it is the responsibility of the researcher to establish the scientific merit of the data fusion. 3. metadata curation and generation the dames infrastructure is still in development at the time of writing. an infrastructure has been implemented to support ddi 3. this is also supported by a grid computing standard called jsdl (job submission description language [3]). a number of metadata standards were reviewed in [1], with the conclusion that ddi version 3 was the most appropriate one to use for curating datasets and resources, and for describing datasets generated or modified by the dames tools. many resources that are currently available for use in dames (and in particular resources available from geode [2]) have ddi version 2 metadata figure 1. the dames infrastructure – key use is made of ddi in the components marked ✩ figure 2. the first stage of data curation using the dames infrastructure iassist quarterly spring summer 2009 27 records. it is therefore necessary to support this version as well. however, the following focuses on metadata curated and generated using version 3. 3.1 data entry and curation the dames services distinguish private datasets, public datasets, and other resources (such as classification schemes and processing instructions). users may populate their own private storage space with data and resources and not have them exposed to other users. the dames services can process private data, public data, or a combination of the two. all datasets resulting from the use of services are considered private. they are therefore available only to the user that generated the data. users may propose making data and resources public (uploaded or generated through tools). however, since uploading data and resources to the dames infrastructure does not require or generate any metadata, the users must first perform curation. curation is handled through a web portal wizard. it takes users through various stages to generate ddi version 3 metadata that describe the items being curated. users can upload existing metadata, which are used to automate much of the curation process. curated data and resources can be proposed for publication to the dames management team. the management team then inspects the items and the metadata, and can make the metadata public or suggest improvements to the users. figure 2 illustrates the first stage of curating a new data resource. metadata records are made public, but not the actual data and resources. this is done to avoid making generated datasets public that are derived from datasets which the user does not have authority to publish. as explained below, the metadata for derived datasets refer to the original datasets and their metadata. the dames services access the originals on behalf of the users and are therefore limited to the access permissions of the users. if the users do not have access to the originals, they are limited to inspecting the metadata of derived datasets in order to determine how to obtain appropriate permissions. the possibility of automatically inferring access rights, particularly for derived datasets, is being investigated. in principle it would be possible to export the ddi metadata records, although this facility is currently not available through the dames services. 3.2 tool inputs and outputs whilst ddi metadata are appropriate for describing resources, an instruction language is also needed to describe composable tasks that services can use in a grid environment. jsdl was chosen owing to its widespread use in existing grid systems, and the need for dames services to run on largescale existing grids. however, using jsdl raised two key problems: how to handle job flows, and how to transform jsdl records into meaningful social science metadata. the first problem with jsdl is that does not support relationships between jobs. this is important because dames services require interactions between numerous jobs to complete their processing. for example, a service for fusing two datasets accessed from different external databases might require jobs for staging in each dataset, imputing variables, mapping variables, and fusing the data. some of these jobs could run in parallel, while others might depend on the results of earlier jobs. the mechanisms offered by condor [4] were considered, but these are hampered by lack of support for data flows. the result is that data cannot be staged in effectively from various external data sources on behalf of users. it is also hard to specify how to effectively stage out the resulting datasets and metadata. the solution adopted was to describe service job flows by combining the purpose-designed jfdl (job flow definition language [5]) with jsdl. jsdl is used to describe service provision issues and data staging. in addition, jfdl is used to describe the relationships between datasets and the transformational tasks needed to realise dames services. the jsdl/jfdl fragment in listing 1 shows job flow and data flow for an example service (omitting the obvious closing tags here). listing 1. an example jsdl/jfdl fragment <jsdl:jobdefinition xmlns:jsdl=”…” xmlns:jfdl=”...”> <jsdl:jobdescription> <jsdl:jobidentification> ... <jsdl:application> <jsdl:applicationname>dames::fusion ... <jfdl:jobflow> <jfdl:dataset file=”gamma1.csv” id=”acsv”/> ... <jfdl:method submitfile=”submit.imputea” id=”imputea”/> <jfdl:job id=”j1”> <jfdl:datasequence> ... <jfdl:method ref=”imputea”/> <jfdl:job id=”j3”> <jfdl:parent_job ref=”j1”/> <jfdl:parent_job ref=”j2”/> ... <jfdl:method ref=”fusion”/> <jsdl:datastaging> <jsdl:filename>gamma1.csv <jsdl:creationflag>overwrite <jsdl:source/> <jsdl:datastaging> <jsdl:filename>output.csv ... <jsdl:target> <jsdl:uri>irods:/home/myfusion/ad7a835325cd2bd d9f549c0271274796/f.csv ... 28 iassist quarterly spring summer 2009 the second problem with using jsdl (and jfdl) to describe dames tasks is how to transform these records into meaningful social science metadata. when dames services succeed in outputting datasets created from jsdl/ jfdl inputs, the jsdl/jfdl instructions are metadata describing how the output datasets were created. however, these records lack descriptive qualities such as the purpose of the investigation, the investigator, and a logical view of the input and output datasets. the records are also not in a standard format for sharing by social scientists. to solve these problems the dames team has developed a solution that uses xslt (extensible stylesheet language transformations) to transform jsdl/jfdl instances, along with ddi 3 instances describing user and job profiles, into ddi 3 dataset records. following translation, the ddi metadata are stored in the metadata database for future searches by researchers. in addition, the jsdl/jfdl file is stored so that the job can be re-run if necessary. the jsdl/ jfdl fragment in listing 1 results in the ddi 3 fragment shown in listing 2 (again, omitting closing tags here). the dames team has developed a data fusion service to test the infrastructure. social science researchers can use this service to specify two datasets to be grouped by common variables (imputing data if necessary). a wizard was developed as a web portlet to generate the jsdl/ jfdl file and submit it to the service. the service stages in the datasets and processing algorithms on behalf of the user, runs the jobs, generates the metadata, and makes the output dataset and metadata available to the researcher. additionally, researchers can inspect the status of failed jobs to determine the causes of failure. figure 3 illustrates the metadata cycle implied by the data fusion service: searching the metadata discovers datasets for processing. the outcome of processing is curated figure 3. the dames metadata cycle processing ddi3 search <ns1:ddiinstance xmlns:ns1=”ddi:instance:3_0_cr3” ...> ... <g:group time=”t0” instrument=”i0” panel=”p0” geography=”g0” dataset=”d3” language=”l0”> <r:citation> ... <g:purpose>... ... <a:archive> <r:name>dames data archive ... <g:concepts> <!-a conceptual component describing the data that are being fused is given for each mapped and imputed variable --> ... <g:datacollection> <d:datacollection> <d:processingevent isidentifiable=”true” id=”[fusion method id]” isderived=”false”> <d:coding isidentifiable=”true” id=”[fusion method submit file]”> <d:generationinstruction> <d:sourcevariable ...> <r:urn… <d:mnemonic>donor dataset <d:sourcevariable ...> <r:urn>… <d:mnemonic>recipient dataset ... <r:command> <r:commandfile formallanguage=”[language such as spss or stata]”> ... <r:uri>[uri for the command file] ... <g:logicalproduct> … <l:categoryschemereference ... uri=”[uri to category scheme for dataset]”> ... <g:studyunit> <!-donor dataset --> ... <g:studyunit> <!-recipient dataset --> ... <!-a comparison is added for each mapped donor & recipient variable --> <cm:comparison> ... <cm:variablemap> <cm:sourceschemereference isexternal=”true” uri=”[uri to donordatasetvariable]” /> <cm:targetschemereference isexternal=”true” uri=”[uri to recipientdatasetvariable]” /> ... listing 2. an example ddi 3 fragment iassist quarterly spring summer 2009 29 along with its own metadata, ready for re-discovery in subsequent searches. the data fusion tool outputs datasets described by metadata records that make use of the grouping and processing event description capabilities of ddi 3 (as shown in listing 2). the metadata files group the corresponding input datasets as referenced study units, along with references to command files containing the processing instructions. they also contain citation, grouping purpose and information about the fusion tool, along with concepts and logical products composed from references to the input datasets’ metadata. the principal structural ddi 3 tags currently used by the dames infrastructure are <ddiinstance>, <citation> and <group>, with detailed metadata held within <group> elements using <purpose>, <concepts>, <datacollection> <logicalproduct> and <studyunit> tags. 4. discussion and conclusion this paper has presented the dames infrastructure which is being developed to support the provision of data access and processing services for social science researchers. we have implemented prototype tools based on the infrastructure for resource curation and data fusion. the curation tool generates standardised metadata about social science data and resources to operationalise and fuse the data. the data fusion tool uses existing metadata to perform data linkage tasks and generates metadata for its output datasets. the dames infrastructure integrates these tools in a framework that manages depositing, accessing, and processing data. prototype and early versions of the services are accessible through the dames web site at www.dames.org.uk/resources.html. this paper has shown that ddi 3 features are useful to the curation and fusion services. the paper has also shown that ddi 3 can be integrated with approaches making use of additional metadata schemas purposed for service execution such as the jfdl for job flows description and jsdl for job submission description. acknowledgements the dames project is funded by the uk economic and social research council (grant res 149-25-1066). the authors thank their colleagues in the national e-science centre at the university of glasgow for their fruitful collaboration. notes 1 jesse m. blum, research assistant, computing science and mathematics, email jmb@cs.stir.ac.uk. guy c. warner, research assistant, computing science and mathematics, email gcw@cs.stir.ac.uk. simon b. jones, lecturer, computing science and mathematics, email sbj@ cs.stir.ac.uk. paul s. lambert, lecturer, applied social science, email paul.lambert@stir.ac.uk. alison s. f. dawson, research fellow, applied social science, email a.s.f.dawson@stir.ac.uk. koon leai larry tan, research assistant, computing science and mathematics, email klt@cs.stir.ac.uk. kenneth j. turner, professor, computing science and mathematics, email kjt@cs.stir.ac.uk. university of stirling, stirling, fk9 4la, uk references 1. j. m. blum and k. j. turner. the dames metadata approach, technical report csm-177, department of computing science and mathematics, university of stirling, dec. 2008, issn 1460-9673. 2. p.s. lambert. an illustrative guide: using geode to link data from soc-2000 to ns-sec and other occupationbased social classifications, edition 1.1, geode project technical paper number 2, university of stirling, 2007. 3. a. anjomshoaa et al. job submission description language (jsdl) specification version 1.0, global grid forum, nov. 2005 4. d. thain, t. tannenbaum and m. livny. distributed computing in practice: the condor experience, concurrency and computation: practice and experience, 17(2–4):323–356, feb.–apr. 2005. 5. j. blum and g. warner. job flow definition language (jfdl 1.0), available from www.cs.stir.ac.uk/~jmb/dames/ schemas, accessed oct. 2009. 6. a. dale. quality issues with survey research, int. j. of social research methodology, 9(2):143–158, 2006. 7. p. s. lambert and k. l. l. tan. instructions for using the geode portal, edition 1.1, geode project technical paper no. 1, university of stirling, available from www. geode.stir.ac.uk, 2007. 8. p. s. lambert, k. l. l. tan, k. j. turner, v. gayle, k. prandy and r. o. sinnott. data curation standards and social science occupational information resources, int. j. of digital curation, 2(1), 73-91, 2007. 9. r. levesque and spss inc. programming and data management for spss statistics 17.0, spss inc., chicago, 2008. 10. j. s. long. the workflow of data analysis using stata, crc press, boca raton, 2009. 11. v. van den eynden, l. corti, m. woollard and l. bishop. managing and sharing data: a best practice guide for researchers, uk data archive, colchester, 2009. 12. m. vardigan, p. heus and w. thomas. data documentation initiative: towards a standard for the social sciences, int. j. of digital curation, 3(1):107–113, 2008. 30 iassist quarterly spring summer 2009 13. liferay. open source enterprise portal with web content management system, collaboration and social networking – liferay, available from www.liferay.com/web/guest/ products/portal, accessed oct. 2009. 14. p. s. lambert, v. gayle, k. l. l. tan, j. m. blum, a. bowes, s. jones, k. j. turner, g. warner, r. o. sinnott and e. bihagen. grid enabled specialist data environments: forward planning for the ge*de services for specialist data on occupations, educational qualifications, and ethnicity, technical paper 2008-1, university of stirling, available from www.dames.org.uk, 2008. 15. exist. an open source database management system built using xml technology, available from www.exist-db. org, accessed oct. 2009. 16. irods. data grids, digital libraries, persistent archives, and real-time data systems, available from www.irods.org, accessed oct. 2009. 17. condor. high throughput computing, university of wisconsin, available from www.cs.wisc.edu/condor, accessed oct. 2009. vol19.1 22 iassist quarterly abstract: a report from an enquete. data professionals have given their views to questions on a future codebook format (“should it be sgml?”, “should it be supported by vendors?”, etc.). this forms a description of “what we want”. on the other hand archives have described their holdings with regard to levels of machine-readable documentation. this section focuses on the present and on the actual data thus: “what we have”. the paper presents the key figures from questionnaires sent out by “the iassist codebook action group”. maybe we are moving in a direction demanding less knowledge about special formats from our users? hinting to the conference theme the subtitle could be “access for other than partners?”. background information history at the iassist conference in edinburgh in may 1993 several sessions were centred around documentation, and some specifically with documentation at the variable level (codebooks) as opposed to documentation at the study level (study descriptions). the sessions concerned with codebooks were: “roundtable on codebooks” (wednesday, may 12) “poster session on production of codebooks” (thursday, may 13), and the “codebook session” (friday, may 14). i chaired the first and last of these sessions and the dda contributed to the “poster session”. the willingness to present papers at the “codebook session” as well as the many arguments and the eagerness of discussions all pointed out that many professionals were interested in working in the area of codebook documentation. the papers0 at the “codebook session” contributed to many discussions at the conference, and after the conference the paper by the icpsr director richard rockwell created many e-mail discussions. at the business meeting at the conference an action group was formed and named: “codebook documentation of social science data”. i was appointed the chairman or co-ordinator of this action group and during 1993 was joined by lennart brantgärde and bill bradley as formal members of the action group. my intentions with the action group was broadcast on the iassist listserver in late may 1993 as follows: during the last years there has been some confusion concerning who are making which changes to the osiris codebook. the osiris codebook format has for more than twenty years been the de-facto standard for archives around the world. the most obvious reasons for this standard are that the osiris codebooks can store full text, are input to retrieval systems, and that the codebooks are easily converted to other formats (sas and spss). i propose that the task of the working group is to remedy this confusion by: 1) identifying the tasks and areas that cannot be handled by the osiris codebook in its present form. two examples: a) many archives are looking for a feature of presenting tabulations in the codebooks not just frequencies. b) a less rigid print out of codebooks not limited by the original card image input format. it is important to note that the first is a problem of storing a type of information in the codebook which does not fit the present format. the second concerns the utilisation of the codebook, and could be changed without changing the codebook format (by flowing the text, and using a different font). 2) identifying the number of data sets at archives all over the world. the data sets should be grouped by the level of documentation: a) full text machine readable codebooks (what format?); b) abbreviated machine readable documentation (osiris dictionary / sas / spss ?); c) no machine readable documentation (paper / scanned information). data sets that are originally stored at other archives (icpsr etc.) are to be counted only at the original archive. another question would be whether the archive is producing machine readable documentation at all. the rationale behind the second identification is that archives who have produced and stored machine readable documentation (e.g. osiris codebooks) will be able to convert these to the new format (missing the special features of the documentation what we have and what we want: report of an enquete of data archives and their staff by karsten boye rasmussen1 danish data archives 23fall/winter 1994 new format). there will not be invented a new codebook format which automatically documents what has not been documented. the archives who now uses the same documentation format will be able to share software for making the conversion to the new format. i must emphasize, that it is not my intention that the codebook working group should produce a complete proposal for “this is how a codebook should look like”. we have no way to enforce a new standard other than waiting for people and organisations to realise its superiority to older standards. the task of the working group is to identify and structure the problems: these are common problems; these are problems with historical data; these are problems with complex data; etc.i think one feature of the new codebook format can be revealed now: it has to be so flexible that new features do not require a new format. another reason for not putting forth a new codebook standard is that it is my belief that documentation at both the study description level and the codebook level (and levels in between) has to be integrated. my plans for the actions group were not fully met. one of the things that things that changed was not surprisingly the time schedule. i had planned to report to the iassist conference in 1994, but due to other obligations this was postponed till 1995. in the meantime several activities took place. in europe a cessda seminar with the title “variable level documentation” took place at the ssd in göteborg in august 19931; the participants list was not restricted to europeans. in august 1994 a cessda seminar was held in grenoble2, this time the theme was “networking and internet?”, again icpsr willingly sent a participant so the perspective was broadened and more global than just european. although the term “codebook” was not part of the seminar title the searching and availability of codebooks on internet and therefore also the format of codebooks -were discussed intensively during the seminar. in october 1994 icpsr announced a commitment in the development of an sgml dtd for codebooks. this issue was addressed again in mid february 1995 when icpsr announced an international committee. the questionnaires in the summer of 1994 i formulated two questionnaires. one questionnaire was individual and gave the opportunity to present personal views on present and future codebook formats. the other questionnaire intended to count the number of studies at different archives and to show the distribution of different levels of documentation. in late autumn i received comments and advices concerning the questionnaire from a group consisting of charles k. humphrey (data library of university of alberta), lennart brantgärde (swedish social science data archive in göteborg), and william bradley (canadian health and welfare, ottawa). the population and the return rate finally on 29 november 1994 the two questionnaires were sent to the listservers consisting of the iassist membership (178 recipients), official representatives of the icpsr (195 recipients), and to the small ifdo list (27 recipients). a number of individuals were on more than one of these lists. the guess is that individuals from around 200 institutions were given the opportunity to answer the questionnaires. all recipients had furthermore the opportunity to forward a copy of the questionnaires to other individuals whom might be interested. a pull for the return of the questionnaires were made shortly after the announced deadline on january 15th 1995. several questionnaires were received thereafter and the last questionnaire was received at the end of february. at that time the number of received individual questionnaires had reached 50, and the number of institutional questionnaires were totalling 20. these numbers do not hold evidence of the representativeness or the missing representativeness of the collected data. the investigation was never intended to be representative. however it is a fair assumption that the individuals that took the time to answer the questionnaire are individuals that have an interest in the development of documentation formats. and it is a fact that the archives that have returned the institutional questionnaire are principally amongst the national social science archives. all though it is disappointing that not all ifdo members made the effort of answering the institutional questionnaire. the individual questionnaire as shown in the “appendix 1” the individual questionnaire is mainly about 25 different assertions about codebook documentation that individuals rate with their level of agreement on a five level scale from “strongly agree” to “strongly disagree”. the middle category was labelled “indifferent, don’t know” and this presented a problem to some individuals as these two answers are not completely identical. 24 iassist quarterly the methodology of exposing individuals to a battery of items is well know. looking back a feature that limited the amount of points which each individual could distribute would have been effective. it is much easier to answer that a lot of factors are important than actual to rank the factors and point out that “these are most important”. on the other hand the distribution via e-mail called for both a very simple layout and a simple question structure. this led to the concentration on the battery of items without any filtering structure or deepening sub-questions. in “appendix 2” the distribution of the answers in these 25 variables is shown together with information about the mean and the number of non-missing answers. the mean was simply calculated by appointing the values 1 through 5 to the five answer categories: 1 strongly agree 2 agree 3 indifferent, don’t know 4 disagree 5 strongly disagree the treatment of the items calculating the mean implies that the items are viewed as belonging to an interval scale. lots of arguments can be put forth in favour of or against this decision. in this context i find that “keep it simple” will suffice as legitimisation of this manoeuvre, as a more appropriate measure like “mode” will lose information. if you are interested in viewing the actual distribution of the answer categories for each of the items you should take a look at “appendix 2”. as the direction of the assertions differs a high mean on one variable and a low mean on another does not necessarily imply that these two variables do not support each other. the mean can range from 1 to 5. in order to compare two variables of different direction you should keep in mind that a mean of for instance 4.2 is equally strong as a mean of 1.8. means around the number 3 implies that the community has no fixed strong feelings about this particular assertion. in the following the themes of the 25 assertions have been boiled down to a few headings. the need for standards i do not intend to enter a long philosophical discussion about standards. i am sure that we are all aware of the benefits of standards. it is equally true that we all spend time getting from one standard to another, e.g. getting from one analysis package to another. the common sense meaning of standard in this context is “as a rule”. we expect to be able to connect our electronic equipment when we move to a new house, the plugs are supposed to be “standard”. but having been to iassist conferences we know that this is not true when moving between countries. thus the less common sense and the more sophisticated the more standards are available. or to put it more precise: the many standards are a painful fact of life. so it is no surprise that the lead question among the assertions “1. there is no great need for standardization of codebooks” receives a very strong disagreement value (mean 4.2). there is a need for standards, and we can continue the search for the standards. the ultimate alternative to the codebook namely “no codebook” is considered a very bad solution. “2. a data user should be content with a study description and photocopies of relevant pages from the questionnaire”, with a mean of 4.3 this item receives the strongest disagreement of all 25 items. one of the currently used and widespread standards is osiris3 and this format is drawn into the discussion in “5. there is a great need for more structured information than is available in the osiris codebook format” that shows a weak agreement (mean 2.6). half the individuals answer “indifferent, don’t know” and maybe they answer the latter simply because they don’t know the osiris format and its structure. other formats and these are more commercial and more available formats are also given the same weak agreement (mean 2.6) in “10. let us stick with commercially supported formats for social science data (e.g. sas and spss)”, but it should be noted from the distribution that there is a higher variance in this question. the support of standards standards live because they have supporters. the supporters do not have to be very loudspeaking as the history of osiris shows. in the 70’ies osiris was rather widespread as a social science analysis package, but spss4 and sas5 gained momentum and osiris was abandoned by the researchers. but many archives continued to utilise the osiris documentational format, and still osiris is used by many archives. often the researcher that receives archive materials does not know that the sas or spss setup actually is an automated product created from an osiris codebook. when standards 25fall/winter 1994 depend upon supporters in order to get a strong standard you would naturally want strong supporters. two items express the need for commercial support: “18. a new format should be supported by the analysis software industry (e.g. sas and spss)” and “19. a new documentation format should be supported by major document software and applications (word, wordperfect, www)”. they both receive high agreement (mean 1.8). but does this mean that if we can not persuade sas, spss, word (microsoft), wordperfect (novell) to support a new codebook format then we will have to give up? no, not in my opinion! it would be nice with support from sas and spss, but the history of data conversion shows6 that the packages always almost do the job. that leaves you with problems concerning the character set and especially about missing data. till now it has been much easier to write documentation conversion software that will reduce and format the documentation to the levels supported by sas and spss. another item indirectly addresses the commercial support: “21. a new documentation format will not be of any interest unless the data producers directly produce their documentation in this format”. here we are not only demanding software to follow the standards, but humans and maybe persons we know are asked to follow suit. the mean for item 21 drops to 2.5. then it is getting really hot: “22. a new documentation format is only interesting if all archives abide by the new standard”. the answer is close to indifferent (mean 2.7), but now we have moved from vendors to other people and finally we are trapped ourselves. the support of standards is a nice feature, and the less of one’s own work that is involved, the nicer the feature gets. support from word processing companies will depend upon the new format. but if the new format is going to be directly connected to the formats of the word wide web (html7) there is no doubt that the format is going to be supported8 at least when the word processor is used as a viewer. however we have to be careful here. html is moving and changing, new versions and new features are being introduced. the safe ground to build upon is sgml where definitions can be made. furthermore a codebook defined as a document type in sgml and marked up accordingly is very easily converted or reduced to any html level. labels in the codebook new formats for codebooks or not, everything will not change overnight. we are going to continue to support analysis packages that can only handle limited amounts of text. but there is harsh disagreement that users or archives should be content with this level of information. “7. a documentation format with short labels for variables and values is sufficient for the user” (mean 4.0) and “8. a documentation format with short labels for variables and values is sufficient for the archive where a study is deposited” (mean 4.2). when the question is directed specifically towards the length of the variable labels it is no great surprise that a longer label is preferred to the shorter: “14. variable labels composed of 24 characters is sufficient” is slightly disagreeable (mean 3.4) whereas the next item is found slightly agreeable (mean 2.7): “15. variable labels composed of 40 characters is sufficient”9. the question of course is what the labels are “sufficient” for? i interpret the results as specific for the labels, and not that the codebook documentation should consist of nothing more than the variable labels. oldies but goodies of the codebook information about the marginals is considered one of the main benefits of the codebook. “3. codebooks should contain marginal frequencies that will enable the users to check the data they have received” is heavily agreed upon (mean 1.6). new fashions are not considered that important: “4. codebooks should contain cross-tabulations so the user has more information about how to analyse the data” only makes it slightly above the threshold of indifference (mean 2.8). most of the items about the layout and printing of the codebook receives the same low level of attention. “9. the presentation of a printed codebook is very important” (mean 2.5); “24. there is little need for a printed codebook if a machine readable codebook exists” (mean 2.8); and “25. i prefer to browse data documentation in files on my own computer” (mean 2.6). new possibilities of the codebook if a new codebook format is to be developed we could be talking about transferring information, or we could use the opportunity to expand the current format with new possibilities. some possibilities are mentioned amongst the assertions, and the highest score of all items is received by “20. a new documentation format should include the study description” (mean 1.4). once again this demonstrates the methodological problem of having unlimited resources (points) when filling out the questionnaire. we must conclude from the questionnaire that there is a craving for including the study description in the codebook. from our practical lives we must conclude that we have to solve one thing at a time. redesign of the codebook has to consider but not necessarily to solve the implementation of the study description. 26 iassist quarterly i assume that most people are aware of the new possibilities but these are not considered very important for the codebook. the answer to “11. a documentation format should be able to incorporate pictures and sound” is indifferent (mean 2.9). the item on scanned images “23. i would be content to receive scanned images of the questionnaires” receives a slightly less favourable score, but that is due to poor verbalisation on my behalf. i believe that receiving only scanned images will not satisfy the user, but the combination of the text based codebook with the actual questionnaire as presented through scanned images will be of interest to most users. efforts of internationalisation have been mentioned at several meetings on documentation. one of the main obstacles in reading other languages than ones own. the introduction of english labels is not greeted with much enthusiasm: “6. english variable labels should be used with all studies regardless of the language of the original study” receives only a mean between agreement and indifference (mean 2.5). with the syndrome of unlimited resources you would have expected this to be easily agreed upon. furthermore people from english speaking countries (uk, usa, canada) seem to be only a fraction more inclined towards agreement. the institutional questionnaire the intention of the institutional questionnaire is to determine: a) the number of datasets amongst the social science data archives. b) the number of individual datasets (not held as a copy from another archive). c) the distribution of used archive formats. datasets in social science archives one of the big subjects has been to make a clear definition and to answer the question about the presumed unit of analysis “what is a study?”. as archives would tend to see their importance depending on the magnitude of studies, and especially as compared with other archives, the study unit seemed to be of crucial importance. this led to some correspondence over the issue as it was argued that some studies consist of many datasets, have a complex scheme, and are distributed on several tapes, other studies have only few variables and a couple of hundred cases. however i found that the introduction of a “true” unit of analysis could not succeed within the limited time period for returning the questionnaire. it should be noted that the individual questionnaire demanded only the person to make up his mind. in the institutional questionnaire more questions are asked based upon the unit of analysis. even though the unit of analysis (the study) is not the same at all archives, it was more important that the individual archive had the easiest possibility of answering the questions; it was hard enough anyway. at each archive they had to do some kind of stocktaking and to place the studies within the categories demanded by the questionnaire. categorisations and fundamentals of archives the archives that have answered the questionnaire are not many. 20 in numbers. one returned questionnaire was afterwards disregarded because both a staff member and the director of the archive had returned a questionnaire. 19 is now the basis. the reason for some people to fill in the individual questionnaire as the only person from their archive, and yet not fill in the institutional questionnaire must be that the institutional questionnaire involved much more work in order to be filled out. therefore i want to express my sincere thanks to the persons who took the time to fill out the time-consuming institutional questionnaire. when preparing to discuss what kind of codebook documentation format should be used in the near future it is interesting to know whether an archive produces documentation. the answer of 7 archives to the question “does your institution produce original machine readable documentation for studies in your archive?” was “no we are only storing”. most of these archives are university data archives in north america, but the english national data archive was also a member of this group. if you do not produce documentation you will end up having all kinds of documentation formats. the interest of these archives in a new documentation format must be for the utilisation of the splendours of a new format, more than how a new format will act as a standard and solve some of our current documentational problems. of the 19 archives 13 archives have english as their native language. 27fall/winter 1994 paper documentation vs. machine-readable documentation first the respondent from the archive is asked about the number of studies with only paper documentation (q1_1) secondly about the number of studies with some kind (any kind) of machine-readable documentation (q1_2). the 19 archives share among them more than 27,000 studies, of these close to 11,000 have some kind of machine-readable documentation. the ratio of machine-readable documentation compared to all studies is thus 0.4. but this ratio hides great differences amongst the archives, some archives go as high as 0.9 others below 0.2 (the mr ratio). it is important to note, that the majority of studies at data archives do not have any machine readable documentation at all. even when the machine readable text information is very limited (e.g. only describes variable labels and few categories) it can be used as input for software that automatically will convert the information to a new documentation format. this is a qualitative difference compared to having no machine readable text information. the studies without machine readable documentation will demand much more work to process to a new documentation format. own holdings vs. deposition from source archive in q2 the question is put about how many of the studies are actually stored and available from another archive (a source archive). the number of studies totals to more than 11,000. as many of these archives report their main source archive to be icpsr, and as icpsr is among the 19 answering archives we can decide to be bold and just subtract this from the total, leaving us with around 16,000 unique studies. the ratio of own studies varies from 1.00 (all studies are own studies) to ratios close to zero (no studies that do not come from a source archive). the scattergram shows the distribution between the ratio of own studies and the ratio of machine-readable studies in the figure below. types of mr then five questions about levels of machine readable documentation are asked. the respondent is asked to distribute the studies counted in q1_2 having machine-readable documentation into five separate categories where the lowest rank is scanned images and the highest rank is a full codebook. the five categories are expected to sum to the total of machinereadable studies, but only half of the archives manage their stock data in such a way as to make the balance right. the sum of the studies in the five categories in q3 totals to 12,627, whereas the total of machine readable studies in q1_2 is 10,946. scanning as mr documentation scanning (q3_1) of questionnaire pages as the highest level of machine-readable documentation is to close to 100 pct. only 28 iassist quarterly found at the amsterdam based steinmetz archive. they have earlier shared their research and experience in scanning and made us aware of choosing tiff-4 as the scanning format10. it is obvious that scanning is used at other archives both for security reasons as well as an easy deliverable especially over internet but most studies will have some kind of higher level documentation as well. one of the important things to remember when talking about scanning is that the scanned images should be referenced / pointed to / tagged from the character document that constitutes the codebook. plain text as mr documentation after some editing11 the distribution of question q3_2 with the category “unstructured and untagged text (text from ocr, questionnaire from wordperfect, etc.)” ends with the following result: at first it is surprising that the largest figure is given without any information about the format. on the other hand these formats are without importance because the formats are not directly related to the variables in the codebook or structured into elements of the variables. this category contains cases where the information is more like a stream of text for instance the non-edited result of ocr. when “ascii” is mentioned this is not a reference to the irrelevant character format, “ascii” here means “plain ascii” which again means “no formatting information”. dictionary as mr documentation the category in q3_3 “machine readable dictionaries (information only about variable locations, labels, missing data in sas, spss, osiris or other format.)” means basic dictionary information on the variable level without information about the categories. in common database systems this is the dictionary information documenting the single fields of the database. however the introduction of information about missing data values and meanings is a social science product. a few formats are not widely used. for instance the israeli archive mentions that they store their catalogue information in their aleph system, while the data is stored and documented in sas. personally i do not find the differences in these systems interesting. as the dictionary level is defined in the question q3_3 all these system will be very much alike. at the same time within a single product for instance the much used spss system the differences between different forms of spss can be as challenging to the user as receiving different products. spss can mean: export files, system files (what platform), setups (what version) etc. 4,520 files are deliverable with the lowest level of machine readable variable level information. dictionary+ as mr documentation this level is defined as the dictionary information plus information on values: q3_4. “machine readable codebook documentation (as 3. above but with the addition of explanation of the coding categories such as value labels in spss or user formats in sas)”. this is a restricted codebook format that does not have a lot of codebook levels and elements and does not support unlimited amounts of text. two interesting facts: ddms is the system developed at canadian health and welfare and the system is being used at different canadian archives12. the nsd stat is the format belonging to the statistical package developed by the norwegian archive13. format frequency 1. scanning 505 2. text 1943 3. dict 4520 4. dict+ 3656 5. dict + codebook 2003 total 12627 text format frequency ascii 532 wordperfect 299 other 1112 total 1943 dictionary format frequency dbaseiv 20 osiris 248 sas 266 sps s 3550 other 436 total 4520 29fall/winter 1994 again the spss format is most often mentioned. and again my same views concerning formats between packages and within packages apply. notice that the total number of studies archived drops when we demand value label information compared to the dict-level before. dictionary codebook as mr documentation in q3_5 the level is defined as “machine readable codebook documentation (as 4. above but including all questionnaire text and other information like in the osiris format codebook).” this can give some difficulties if the respondent does not know the osiris format. some are still mentioning sas and spss, even though these packages cannot go above the level of dict+, but these packages are then combined with other text. others mention different cataloguing systems and searching capabilities. all in all a conservative guess of the magnitude of fully documented studies can be as low as 1256 (ddms plus osiris). if we find this figure disappointingly low then on the positive side we can now remember that some archives have not answered the institutional questionnaire. however these 2,003 studies plus the 3,656 studies belonging to the dict+ level are the studies with machine readable documentation available. these are the studies that the user will be able to analyse without consulting other material except of cause the study description. these 5,659 studies should optimistically be the studies that can be automatically converted to a new codebook format. if we are going to divide the task between us, i will personally choose to make the conversion of the osiris codebooks because of their very simple format! conclusion future of the codebook? codebook of the future! we have seen that “what we want” is “everything”. “what we have” is pessimisticly close to nothing (1300 mr documented datasets). we noticed that this high level of documentation should preferably be produced by “somebody else”. we know that only a small percentage of the studies are fully documented, and these can easily be converted to a new codebook format. we know that more than half of the studies have no machine readable documentation at all. we can conclude that we are facing a daunting big job. however: if a new format can combine all (or almost all) our needs for structural elements in the documentation, the archivist will save time. i am very much looking forward to the work on a new format in the icpsr “sgml codebook committee”. we have a great opportunity in using the tools that are being offered. because the archives immediately will develop tools for processing the documentation for use in many surroundings like cdrom, hypertext, internet, etc. the possibilities of the new format will also be of interest to data producers. laura guy14 talks about the archivist pleading with producers to “.. please .. document it properly?”. with these modern add-ons there will be great benefits of doing proper machine readable documentation. dict+ format frequency ddms 137 nsd stat 100 osiris 268 sas 116 spss 2138 other 897 total 3656 dict format frequency ddms 45 osiris 1211 sas 37 spss 37 other 673 total 2003 30 iassist quarterly paper presented at the iassist conference in quebec city, may 1995 1 report from ssd cessda seminar “variable level documentation”, göteborg 1993. 2 the report and paper collection from the grenoble meeting is not yet published. 3 the osiris format is documented in osiris iii, volume i, isr 1973, univ. of michigan. 4 ”spss for windows, release 6.0”, 1993, chicago, spss inc. 5 sas has meters of manuals, i’ll mention 4 cm: “sas language: reference, verions 6”, 1990, cary, sas institute inc. 6 poster session at iassist 92 and “converting data” in dda-nyt 62, summer 1992. 7 ”hyper text markup language” is a document type definition (dtd) made in sgml (standard generalized markup language). html is used as the document format in www. a very usable sgml book is: “practical sgml” by eric van herwijnen, 2. ed., 1994, kluwer, dordrecht. 8 in march 1995 microsoft announced their html extension available for word (but so far only for the us version 6.1). 9 as a curiosity i can mention that a cross tabulation of the two items (14 and 15) shows that some persons find 24 character labels agreeable but find 40 character labels disagreeable. 10 ”exchange of scanned documentation between social scientists and data archives: establishing an image file format and method of transfer”.repke de vries and cor van der meer in iassist quarterly vol.16, number 1/2. 11 apart from summarizing i have taken the liberty to evenly distribute figures connected to more than one format. if an archive gave the number 200 and mentioned the formats ascii and wordperfect both categories received 100. 12 william bradley mentioned in note 1 is the leader of the group developing ddms. it should be noted too that ddms is a full codebook system, not a restricted system. 13 the documentation format is only mentioned at the norwegian archive. but the use of the nsd stat pc-package is much more widespread 14 ”the need for revised data documentation standards: new solutions for old problems”, laura guy in iassist quarterly vol. 17 num. 3/4. 31fall/winter 1994 appendix 1: the questionnaires in the following the questionnaires as e-mail to the listservers. part 1: data documentation preferences codebook documentation of social science data: an iassist action group in this survey we are interested in your individual opinionsabout improving data documentation and the formats used with data documentation. the information obtained through this survey will become part of a report from an action group under the international association of social science information service and technology (iassist) about the current practices of social science data documentation and proposed standards for codebooks. your thoughtful completion of this questionnaire is appreciated. this first part is an individual questionnaire. if your institution stores or archives social science data we ask you to fill out the second part with information about your institution and archival format. you are kindly invited to complete this questionnaire and return it to karsten boye rasmussen, dansk data arkiv, islandsgade 10, dk-5000 odense c., denmark by mail, fax (+45 66113060) or e-mail (kb@dda.dk). if you are using e-mail please observe that you are not responding to the list server but directly to kb@dda.dk. please return this questionnaire before the 15th of january 1995. q1. below is a series of statements about data documentation. please indicate for each statement the degree to which you agree or disagree with its content. use the following five-point scale: 1 strongly agree 2 agree 3 indifferent, don’t know 4 disagree 5 strongly disagree __ there is no great need for standardization of codebooks. __ a data user should be content with a study description and photocopies of relevant pages from the questionnaire. __ codebooks should contain marginal frequencies that will enable the users to check the data they have received. __ codebooks should contain cross-tabulations so the user has more information about how to analyze the data. __ there is a great need for more structured information than is available in the osiris codebook format. __ english variable labels should be used with all studies regardless of the language of the original study. __ a documentation format with short labels for variables and values is sufficient for the user. __ a documentation format with short labels for variables and values is sufficient for the archive where a study is deposited. 32 iassist quarterly __ the presentation of a printed codebook is very important. __ let us stick with commercially supported formats for social science data (e.g. sas and spss). __ a documentation format should be able to incorporate pictures and sound. __ missing data should always be coded as numeric values. __ coding of data fields using alphabetical and special characters should be discouraged. __ variable labels composed of 24 characters is sufficient. __ variable labels composed of 40 characters is sufficient. __ changing to a new documentation format would be very difficult to implement at our institution. __ a new documentation format should be a specialized implementation of a general document format (e.g. sgml). __ a new format should be supported by the analysis software industry (e.g. sas and spss). __ a new documentation format should be supported by major document software and applications (word, wordperfect, www) __ a new documentation format should include the study description. __ a new documentation format will not be of any interest unless the data producers directly produce their documentation in this format. __ a new documentation format is only interesting if all archives abide by the new standard. __ i would be content to receive scanned images of the questionnaires. __ there is little need for a printed codebook if a machine readable codebook exists. __ i prefer to browse data documentation in files on my own computer. q2. your name : _______________________________ your position: _______________________________ institution: _______________________________ country: _______________________________ q3. are you a member of iassist ( ) yes ( ) no q4. please comment further on any aspect of data documentation that is a concern to you. 33fall/winter 1994 part 2: institutional data documentation codebook documentation of social science data: an iassist action group if your institution stores or archives social science data we are interested in information about your institution and archival format. you are kindly invited to complete this questionnaire and return it to karsten boye rasmussen, dansk data arkiv, islandsgade 10, dk-5000 odense c., denmark by mail, fax (+45 66113060) or e-mail (kb@dda.dk). please return this questionnaire before the 15th of january 1995. this questionnaire is to be completed from an institutional perspective. individual opinions on social science data documentation are to be expressed in the first questionnaire. several individuals from the same institution can fill out the individual questionnaire, but only one person needs to answer the institutional questionnaire. the information obtained through this survey will become part of a report from an iassist action group about the current practices of social science data documentation and proposed standards for codebooks. your thoughtful completion of this questionnaire is appreciated. your name: _____________________________ your position: _____________________________ name of your institution: _____________________________ _____________________________ country: _____________________________ q1. we are interested in the variety of documentation that accompanies social science data and the number of studies with each type of documentation. please indicate the total number of studies available with only paper documentation and those with some form of machine readable documentation. a study should only be counted once. number of studies archived with: _________ 1. only paper documentation (no machine readable documentation) _________ 2. some form of machine readable documentation (may also be available in print) q2. how many of your studies have you received from another archiving institution? please specify number of studies: _________ studies from “source” data archive q3. of the total number of studies with some form of machine readable documentation, please indicate how many studies are available in the following formats. please report a study only once at the highest documentation level (1=low 5=high). number of studies in machine readable format consisting of: _________ 1. scanned images please specify the most commonly used format: 34 iassist quarterly _________ 2. unstructured and untagged text (text from ocr, questionnaire from wordperfect, etc.) please specify the most commonly used format: _________ 3. machine readable dictionaries (information only about variable locations, labels, missing data in sas, spss, osiris or other format.) please specify the most commonly used format: _________ 4. machine readable codebook documentation (as 3. above but with the addition of explanation of the coding categories such as value labels in spss or user formats in sas). please specify the most commonly used format: _________ 5. machine readable codebook documentation (as 4. above but including all questionnaire text and other information like in the osiris format codebook). please specify the most commonly used format: q4. does your institution produce original machine readable documentation for studies in your archive? ( ) no we are only storing ( ) yes if “yes” please describe the process and format used for preparing machine readable documentation. q5. does your institution have a policy about the format used for documentation at the variable level? ( ) no ( ) yes if “yes” please describe the format used for describing variables. 35fall/winter 1994 appendix 2: distribution and mean of the individual questionnaire strongly agree indiffe disstrongly mean n nonagree -rent, agree disagree missing don’t know 1. there is no great need for standardization 0 4 3 18 24 4.2 49 of codebooks 2. a data user should be content with a study 1 3 2 14 30 4.3 50 description and photocopies of relevant pages from the questionnaire 3. codebooks should contain marginal 23 25 1 1 0 1.6 50 frequencies that will enable the users to check the data they have received 4. codebooks should contain cross4 17 16 8 5 2.8 50 tabulations so the user has more information about how to analyze the data 5. there is a great need for more structured 6 12 23 5 1 2.6 47 information than is available in the osiris codebook format 6. english variable labels should be used with 5 25 10 6 4 2.5 50 all studies regardless of the language of the original study 7. a documentation format with short labels 2 4 4 22 18 4.0 50 for variables and values is sufficient for the user 8. a documentation format with short labels 2 3 4 14 27 4.2 50 for variables and values is sufficient for the archive where a study is deposited 9. the presentation of a printed codebook is 10 21 5 10 4 2.5 50 very important 10. let us stick with commercially supported 11 12 13 6 6 2.6 48 formats for social science data (e.g. sas and spss) 11. a documentation format should be able to 5 9 22 10 4 2.9 50 incorporate pictures and sound 12. missing data should always be coded as 15 12 14 3 6 2.4 50 numeric values 13. coding of datafields using alphabetical 18 11 8 9 4 2.4 50 and special characters should be discouraged 14. variable labels composed of 24 1 11 12 15 10 3.4 49 characters is sufficient 15. variable labels composed of 40 7 16 12 9 5 2.7 49 characters is sufficient 16. changing to a new documentation format 6 4 17 17 5 3.2 49 would be very difficult to implement at our institution 17. a new documentation format should be a 12 17 14 4 1 2.2 48 specialized implementation of a general document format (e.g. sgml) 18. a new format should be supported by the 19 23 6 2 0 1.8 50 analysis software industry (e.g.sasandspss) 19. a new documentation format should be 18 25 6 1 0 1.8 50 supported by major document software and applications (word, wordperfect, www) 20. a new documentation format should 29 19 2 0 0 1.4 50 include the study description 21. a new documentation format will not be of 12 14 9 14 1 2.5 50 any interest unless the data producers directly produce their documentation in this format 22. a new documentation format is only 7 14 16 11 2 2.7 50 interesting if all archives abide by the new standard 23 i would be content to receive scanned 3 10 10 16 11 3.4 50 images of the questionnaires 24. there is little need for a printed codebook 9 17 1 17 6 2.8 50 if a machine readable codebook exists 25. i prefer to browse data documentation 8 16 10 12 2 2.6 48 in files on my own computer iassist quarterly 2014/2015 47 iassist quarterly xkos an rdf vocabulary for describing statistical classifications by franck cotton, daniel w. gillman, yves jaques1 introduction this paper contains a brief description of the extended knowledge organization system (xkos) and a rationale for why it was developed. in particular, there is a focus on describing statistical classifications with xkos. for statistical data, statistical classifications are essential for categorizing complex domains, such as industries or occupations; presenting dimensions on which to aggregate data, such as in tables or time series; providing the means to stratify populations; and supplying survey respondents with standard response choices. xkos is an extension of the simple knowledge organization system (skos)2 applicable to the needs of statistical offices and social science data users. as we show in this paper, some limitations in skos leave it inadequate to the task of describing statistical classifications. xkos is designed to fill these gaps. skos was published in 2009 as a world wide web consortium (w3c)3 recommendation, and in the same year was extended in another vocabulary named skos-xl. this was to better meet the needs of multilingual thesauri. the purpose of skos is to provide a representation for knowledge organization systems, of which statistical classifications and thesauri are examples, in a machine-understandable way within the framework of the semantic web4. therefore, skosencoded statistical classifications are appropriate for use within the linked open data (lod)5 community. lod is a set of recommendations for building the semantic web, described by tim berners-lee in 2006,6 and has been taken up by a wide variety of communities including biodiversity, environment, statistics, gis, libraries, archives, and museums. its promise to provide crosswalks across domains and types of data is especially attractive to the growing “open access” and “open data” movements that in the social science data community are beginning to force change to the business-as-usual practice of considering each dataset part of its own closed world. implementing the lod recommendations provides new abilities to find, understand, and combine data on similar or otherwise related domains by organizing and linking data and metadata. lod adds value to disparate, difficult to link datasets by employing frameworks such as the resource description framework (rdf).7 rdf, described further in the the purpose of skos is to provide a representation for knowledge organization systems 48 iassist quarterly 2014/2015 iassist quarterly resource description framework section, is a w3c standard used for organizing and linking data. links are used to navigate and find related data and metadata; therefore the technique, among other features, provides an easy to leverage mechanism for building mash-ups (data from multiple sources). implications of using lod for data harmonization were initially explored in a paper by gillman (20108), which includes references to work to mash-up crime, traffic, workplace safety, and natural disaster risk data to create a livability index for us cities. even though the cited work did not employ lod per se, the ideas are very similar to lod recommendations, and the reader is encouraged to understand the example. moreover, the example shows that to do lod right in the statistical framework is not at all straightforward. however, as a growing collection of new tools and many applications have been built with lod, there is an expanding community of interest in employing the technology, and many benefits are promised.9 the statistical data community needs to be paying attention to these developments. along with xkos, other rdf developments that affect the statistical data community have taken place. the data cube vocabulary built through cooperation between lod experts and sdmx technical experts has produced a rendition of sdmx for lod10 which is already in wide use by major initiatives such as data.gov.uk.11 similar work is planned for ddi, and the workshops held at schloβ dagstuhl12 in germany on semantic statistics for social, behavioural, and economic sciences: leveraging the ddi model for the linked data web in september 201113 and october 201214 were devoted to the topic. in particular, this is where xkos was first developed. the original skos is used widely in lod applications, as seen in the skos implementation report.15 as a result, a group was formed at the dagstuhl workshops (in 2011 and 2012) to look at the suitability of using skos in the statistical data community for lod work. as will be described in the skos / what is missing section, skos was found to have shortcomings, so the group looked to address the issues of how to extend skos to meet the needs of the statistical data community. several extensions were deemed important enough for inclusion under a new initiative, xkos, with the intention of submitting this as a w3c editor’s draft. fortunately, the entire design and culture of rdf is based on a spirit of re-use and extension, so extending skos is technically easy. the results of the workshops and subsequent output are reported here. in this paper, we provide introductory remarks to set the stage for discussion, provide a short primer on rdf, describe skos in general and the limitations to statistical classifications embedded in the design in some detail, and lay out the extensions to skos that form the xkos specification. in particular, we show how the semantics of classification systems in our own offices are represented more faithfully by extending skos with xkos. resource description framework this section gives a brief primer on rdf, a w3c standard that facilitates the exchange of structured data on the internet. based on a simple subject-predicate-object model commonly referred to as “triples,” it allows for a generic, standardized structuring of resources that can be used to model and disseminate everything from taxonomies to statistical observations to metadata records. the model used by rdf is also commonly referred to as a “graph model” consisting of “nodes” (which are vertices) and “edges” or “arcs.” see the figure 1 below for an example. the rdf model, which by itself contains only the barest set of classes (subjects and objects) and properties (predicates), is extended using rdf schema,16 another fairly limited set of classes and properties that together with rdf form the foundation of the framework which can then be endlessly extended and specialized as needed. each extension is known as a vocabulary, which is bounded by a namespace. namespaces allow implementers to specify the set of classes and properties that belong to a vocabulary and give a strong assurance of uniqueness even in the open waters of the world wide web (www). this is a concept that will be familiar to those who know xml schemas. the other very important aspect of rdf is that as with its namespaces, all of its classes and properties are also uniquely identified using the underpinning naming mechanism of the internet, the uri17 (uniform resource identifier). in the same way that all web pages are uniquely identified by a uri (web pages actually use the url,18 a subset of the uri specification), all rdf classes and properties are uniquely identified by a uri. in practice this enables a powerful, standardized method for uniquely identifying information of all kinds with great certainty that the information will remain unique not only within the closed context of an internal database, but also across the www. as mentioned before, each vocabulary uses a namespace to scope its set of classes and properties. this namespace is known by a uri, and by common convention the unique identifiers for the classes and properties are appended to this common namespace uri with an intervening hash or forward slash. for example, the commonly used friend of a friend (foaf)19 vocabulary, designed to link instances of people and information, uses the common namespace http://xmlns.com/foaf/0.1/. all of the foaf classes and properties are then appended to this namespace, e.g., the foaf class person is uniquely identified by its uri as http://xmlns.com/ foaf/0.1/person. just as in an xml schema, one can define a namespace prefix to act as a shortcut for the entire namespace. thus in a group of foaf statements (written in xml syntax) one will commonly find a statement such as xmlns:foaf=http://xmlns.com/foaf/0.1/. this simply means that once this foaf shortcut has been defined, one can now refer to the uri that uniquely identifies the class foaf person more compactly as foaf:person. one of the other important aspects of rdf is that it does not rely on a particular syntax for its expression. thus, there are a handful of interchangeable syntaxes that can and are used depending on a variety of requirements that one may have such as brevity or readability. this paper uses the popular turtle (terse rdf triple language20) syntax, prized for its readability. returning to the foaf example, here is how one might make the simple triple statement that one of the authors of this paper is a thing known as a person (with a web page to provide an identifier for the actual person): < http://aims.fao.org/community/profiles/yjaques> <http://www.w3.org/1999/02/22-rdf-syntax-ns#type> <http://xmlns.com/foaf/0.1/person>. iassist quarterly 2014/2015 49 iassist quarterly so to recap, we have a subject “yves jaques”, a “type” predicate (defined in rdfs), and an object foaf:person. to put it in another way, “yves jaques” is an instance of the class “person”. in rdf “type” gets used so often that turtle lets you simply use “a” for convenience: < http://aims.fao.org/community/profiles/yves-jaques> a <http://xmlns.com/foaf/0.1/person>. let’s say we want to make our statement a little shorter. we can define namespace prefixes one time and then use the shortcut for all the other triples in our graph: @prefix foaf: <http://xmlns.com/foaf/0.1/> . @prefix rdf: <http://www.w3.org/1999/02/22-rdf-syntax-ns#> . @prefix aims: <http://aims.fao.org/community/profiles/> . so with those shortcuts defined, we can now write the same statement as (putting the triple on a single line this time): aims:yves-jaques rdf:type foaf:person . or using the turtle shortcut for rdf:type: aims:yves-jaques a foaf:person . let’s say we want to put a few triples together so we can say a little bit more: aims:yves-jaques a foaf:person ; foaf:name “yves jaques” . so here we are seeing the short-hand turtle notation for two sets of triples. in words, these triples are “the yves-jaques aims profile web page is a person.” “the person is named yves jaques.” this illustrates another feature of rdf. the triples may be linked together to tell a story. the object in the first triple is then used as the subject in the next (possibly many) triple(s). to think about what rdf looks like graphically, here is a nice diagram courtesy of marek obitko.21 the round-cornered boxes are classes or instances of classes (subjects/objects), the arrows are properties (predicates), and the square boxes are literals. literals are typically used to represent simple numeric values, dates, or labels. literals can also have a datatype, a powerful mechanism to enforce restrictions on permissible values: and here is the corresponding turtle (note the use of the empty namespace shortcut): @prefix : <http://www.example.org/~joe/contact.rdf#> . @prefix foaf: <http://xmlns.com/foaf/0.1/> . @prefix rdf: <http://www.w3.org/1999/02/22-rdf-syntax-ns#> . :joesmith a foaf:person ; foaf:givenname “joe” ; foaf:family_name “smith” ; foaf:homepage <http://www.example.org/~joe/> ; foaf:mbox <mailto:joe.smith@example.org> . to briefly recap, rdf is a framework that is designed to organize structured data about resources and their relationships over the internet in a standard way. it is designed from the ground-up to be endlessly extensible and able to maintain the uniqueness of the things it represents even in the radically decentralized www. skos what is missing skos provides a means for representing knowledge organization systems using rdf, and this makes the use of skos immediately applicable to lod and the semantic web. so, skos is important for figure 1: rdf graph 50 iassist quarterly 2014/2015 iassist quarterly organizations that wish to use lod and employ classifications and code sets. it is beyond the scope of this paper to provide a detailed description of skos. we direct the interested reader to the skos website (see end note 2). however, skos contains the following basic ideas, whose definitions we paraphrase here: • concept scheme – any knowledge organization system (including statistical classifications and code sets) • concept – any abstract idea or unit of thought • definition – formal statement conveying the meaning of a concept • label – lexical representation for a concept, may be preferred or alternate; provides means to communicate the concept • notation – a symbolic notation for the concept (such as a code) that is typically data-typed • semantic relation – broad category for relations between concepts, such as broader than, narrower than, and related to (these relations can include relations to concepts found in other concept schemes) the basic ideas listed above are the minimum required to describe a classification scheme. we can account for the scheme itself (concept scheme), all its underlying concepts (or categories as they are often called in statistics) with concept, what each concept means (definition), the labels and codes associated with a category (label / notation), and relationships between a concept and its parent and between and concept and all of its children (semanticrelation). so, is anything missing that is needed for statistics? skos is based on the now withdrawn standard iso 2788 guidelines for the establishment and development of monolingual thesauri. this standard describes three basic kinds of relations between concepts: generic, partitive, and instantiation. the generic relation refers to a generic / specific situation, such as between family and genus/species in the biological classification of living things. for instance, all homo sapiens are mammals. the partitive relation refers to a part / whole situation, such as between an automobile and a steering wheel. instantiation is the relation between a kind and an instance, such as each of the authors of this paper are instances of the class of people. both the generic and partitive relations are used in statistical classifications, but instantiation is not. interestingly, the generic and partitive relations are not provided in skos, only the more generic broader than and narrower than, which are often referred to in more technical settings as superordinate and sub-ordinate, respectively. both the generic and partitive relations are specializations of broader than / narrower than. in the skos primer,22 this simplification is acknowledged by the following: “not covered in basic skos is the distinction between types of hierarchical relations: for example, instance-class and partwhole relationships. the interested reader is referred to section 4.7, which describes how to create specializations of semantic relations to deal with this issue.” these more specialized relations were included in the past in skos, but they are now deprecated. xkos, in part, is the effort to put them back. skos also specifies the possibility of an association relation between concepts, but this is not made any more detailed. it is possible to specialize associations somewhat, and that is done in xkos through sequential, temporal, and causal relations, none of which are in skos. the sequential relation refers to ideas where one is the antecedent of the other, either temporally or spatially. an example is the relationship between production and consumption. the specialized temporal relation is based on time. an example is the relationship between spring and summer. finally, the causal relation relates cause and effect, such as the detonation of a hydrogen bomb and nuclear fall-out. upon inspection of some classification schemes in the statistical offices of the authors, some of these relations are needed. there is also a structural deficiency in skos; there is no satisfactory way to represent the idea of levels in concept schemes. levels in statistical classifications are used to identify aggregation levels in reported statistics, which provide producers a consistent way to report their data or provide a way to reduce the threat of disclosures. therefore, xkos also needs to account for levels in concept schemes. examples below are some examples that illustrate the need for the extensions we have identified above: 1. the us standard occupational classification system (soc – 2012). take, for example 27-2000 – entertainers and performers, sports and related workers 27-2040 – musicians, singers, and related workers 27-2042 – musicians and singers the appropriate relation between 27-2000 and 27-2040 is generic, i.e. musicians, singers and related workers is a specialization of entertainers and performers, sports and related workers. the same relation is found between 27-2040 and 27-2042, i.e., musicians and singers is a specialization of musicians, singers and related workers. so, the generic relation is needed to specify the semantics of the us soc. 2. the us occupational injury and illness classification24 (oiics – 2012). occupational injury and illness is a four-facet classification: nature, body part, source, and event. in the body part facet, for example 3 – trunk 31 – chest 313 – heart 315 – lungs 32 – back, including spine, spinal cord 321 – thoracic 322 – lumbar going from broad to lower detail in this snippet of the body part classification illustrates the partitive relation. the chest and back are parts of the trunk. the heart and lungs are part of the chest. finally, the thoracic and lumbar regions are part of the back and spine. note that it would not be proper to use the generic relation here. therefore, the partitive relation is needed to specify the semantics of the us oiics. iassist quarterly 2014/2015 51 iassist quarterly 3. the us american time use survey — activity coding lexicons,25 last updated in 2011. the classification is a hierarchy, but some activity categories depend on what has occurred before. for instance, 04 – caring for & helping non-household members 0402 – caring for & helping non-household children 040204 – arts & crafts with nonhousehold children 040212 – dropping off/picking up nonhousehold children dropping off non-household children is a sequential activity related to having supervised arts-and-crafts activities (or some other activity in the 04 group) previously. so, there are associations between some pairs of activities within this classification. in this case, the sequential or possibly the temporal relation is needed to convey the additional semantics that some activities depend on the triggering of other prior activities. xkos we move now to a description of the xkos vocabulary. as already mentioned, just as skos-xl extends skos for the needs of multilingual thesauri, xkos extends skos for the needs of statistical classifications. it does so in two main directions. first, it defines a number of terms that allow the representation of statistical classifications with their structure and textual properties, as well as the relations between classifications. second, it refines skos semantic properties to allow the use of more specific relations between concepts. those specific relations can be used for the representation of classifications or for any other case where skos is employed. classifications for the representation of statistical classifications, xkos borrows from the neuchâtel model,26 which is a de facto standard created by a group of statistical institutes and maintained in the united nations economic commission for europe’s common metadata framework.27 xkos is not a complete translation of the model, though. in particular, the notion of a classification index is not supported. there are other areas where minor differences exist between the xkos and neuchâtel model approaches: these will be described below. to begin with the classification itself, we distinguish within xkos the notion of classification and that of classification scheme. a classification is a set of classification schemes that share a wellknown name, for example, the european statistical classification of economic activities (nace) or the international standard industrial classification (isic). typically, a classification scheme will be a major version of a given classification. for example, nace is a classification, and each version of nace (the original 1970 version, the 1990 nace rev. 1, the 2003 nace rev. 1.1, and the 2008 nace rev. 2) are classification schemes belonging to this classification. the neuchâtel model also defines the classification variant, which is an adaptation of a classification version to a certain context or usage. in a variant, items can be split, aggregated, added, or suppressed relative to the standard structure of the base version. a variant can also be represented as an xkos classification scheme, albeit of a particular type. xkos does not create its own object classes to represent classifications, classification schemes, and classification items, but directly uses classes already defined in skos. classification items will be represented as instances of skos:concept, with normal skos properties for codes, labels, etc. a classification scheme will simply be a skos:conceptscheme, which is defined as an aggregation of concepts and semantic relationships between those concepts. a classification itself will also be a skos:concept, which can in turn be included in concept schemes representing classification families (e.g., “occupational classifications”, “activities classifications”, etc.). however, xkos defines a set of properties that can be used to link classifications and classification schemes. for example, xkos:belongsto allows one to attach a classification scheme to its classification, and xkos:follows or its sub-property xkos:supersedes can link classification schemes representing successive versions of a classification. xkos also provides a set of properties that indicate how a classification covers its field (e.g., exhaustively, without overlap, both). the field itself would be a skos concept that can be taken from a well-known thesaurus such as eurovoc28 or the library of congress subject headings.29 of course, existing standard rdf properties are available to capture versioning information, textual documentation, etc. examples of these are the dublin core30 dcterms:valid property, or the radion31 radion:version property. also, skos:note can be used to record documentation or other descriptive resources relative to classifications and schemes. in keeping with the rdf spirit of re-use, the existing classes and properties of broadly supported vocabularies are used wherever possible. the main purpose of a classification is to classify the entities that belong to or operate in the field that it covers. in linked data terms, classification results in the creation of an rdf triple where the subject is the resource representing the entity and the object is the concept representing the classification item. xkos defines a generic property, xkos:classifiedunder, that can be used in such statements, but classification criteria are often quite complex: for example, the same enterprise could be classified in different items of a classification of activities, depending on the rules that are used to measure its main economic activity. thus, it is expected that xkos:classifiedunder will be specialized for use in specific contexts. another important notion in the classifications terminology is the notion of level. many statistical classifications, especially those that are international standards, are organized in embedded levels. for example, the isic rev. 4 has four levels: the top is composed of 21 sections that cover broad economic sectors, and there are three more levels that go into greater and greater detail: divisions, groups, and classes. in skos terms, classification levels are just collections or at most ordered collections of concepts, but their hierarchical organization within a classification scheme gives them extra characteristics not covered by skos. thus, xkos defines a dedicated subclass of skos:collection to represent them, which is the xkos:classificationlevel. the levels or instances of xkos:classificationlevel, are structured as an rdf list, starting with the most aggregated, and the list is attached to the classification scheme by the xkos:levels property. an xkos:depth property can be used to express the distance of a given level from the (abstract) root node of the level hierarchy, and an xkos:organizedby property 52 iassist quarterly 2014/2015 iassist quarterly can be used to record the generic name of the items of a given level (e.g., “section”, “division”, etc.). the structure of a classification scheme can be described using the usual skos properties. more precisely: • skos:inscheme (or the more specific sub-property skos:topconceptof if the items belong to the most aggregated level) links the classification items to the classification scheme • skos:member connects the classification level to the items that it contains • skos:broader and skos:narrower represent the hierarchical relations between the classification items in this last case, the more precise sub-properties defined by xkos to express partitive or generic relations between concepts (see below) may be used instead of skos:narrower or skos:broader. figure 2 illustrates a simple abstract case of the usage of skos properties to represent the structure of a classification scheme. textual properties good classifications usually come with a fair amount of textual material, generally organized as notes attached to the classification items or to the scheme itself. these notes typically explain the content of a given classification item by describing what should be classified under this item and what should go elsewhere. for example, here is an excerpt from the official publication of nace:33 we see that the explanatory notes have a defined structure: they first describe what is included in the item, then what is excluded. for the inclusions, a distinction is made between what is evidently included (sometimes called “central content” or “core content”), and what is “also” included, by convention or experts’ decisions, even if it does not result obviously from the item’s label. for the exclusions, the note often refers explicitly to the item(s) where the content should in fact be classified. it is perfectly satisfactory to represent explanatory notes with skos generic notes (skos:note) or scope notes (skos:scopenote), but it can be useful to be able to easily distinguish between the different types of note. for this purpose, xkos introduces four sub-properties of skos:scopenote, which are represented in the figure 3 below. in the case of the nace class 46.34 cited before, we would have three rdf triples to represent the explanatory notes with predicates, respectively, xkos:corecontentnote, xkos:additionalcontentnote and xkos:exclusionnote. skos does not specify which type the objects of these triples should be, nor does xkos. as a side note, eurovoc uses an interesting mechanism that allows the representation of the notes as xhtml fragments, thereby opening the possibility of rendering the references to other items as html links. correspondences between classifications different classification schemes can cover the same classification, the same field, or even fields that are different but semantically related. this induces semantic relations between the classification items that belong to these schemes. a simple example of this is given by two successive major versions of a classification: some items may remain unchanged in the new version, but others will disappear, merge, be created, etc. more complicated n to m correspondences between items of the two versions are frequent. a much more complex example of relations between classifications or classification schemes is given by the international system of economic classifications maintained by figure 2: structure of a classification scheme a rdf:list a xkos:classificationlevel a xkos:classificationlevel (level 1) (level 2) a skos:concept (item 11) a skos:concept (item 21) a skos:concept (item 22) a skos:conceptscheme (classification scheme) xkos:levels skos:member skos:member skos:broader sk os :to pc on ce pt o f sk os :in s ch em e skos:ins chem e iassist quarterly 2014/2015 53 iassist quarterly the united nations statistical division. the european view of this system is well described in the online publication of the nace rev. 2 (op. cit., chapter 1.1). the economic classifications forming this system are linked either by a common structure which gets more detailed as one goes from the international to the european to the national levels, or by semantic correspondences between the economic fields covered: activities, products, and goods (e.g., activities create products). here again, the high-level links established between classifications result in more fine-grained correspondences between items: a given activity will create one or more specific products. thus, there are different types of correspondences between classifications, schemes, or items: • between classifications on the same field, for example, north american and european activities classifications • between different linked fields, for example, classifications of activities and products • historical correspondences, for example, sic to naics • versioning of items over time within a given classification scheme since classification items are represented as skos concepts, we could use the usual skos associative properties to represent figure 3: xkos note properties figure 4: concept association example 54 iassist quarterly 2014/2015 iassist quarterly correspondences between them. however, this simple approach has some limitations: • as mentioned above, relations between items in correspondences are often n to m, whereas skos properties relate one unique concept to another unique concept. it is always possible to decompose an n to m relation into several 1 to 1 relations, but it is better to have a global vision of a given correspondence. we also want to be able to represent 0 to n relations, for example, when an item is created or disappears in a new version of a classification. • more globally, we want to be able to group all the fine-grained item associations that compose a given high-level relation between two classification schemes, such as the ones that exist in the international system of economic classifications. such a collection of item associations is called a correspondence table, conversion table, or concordance. • lastly, it is often useful to be able to attach additional information (for example, notes) to item associations, for example, to describe what proportion of the different items are linked in the association. for these reasons, xkos defines the xkos:conceptassociation class that can be used to represent correspondences between classification items where the skos properties are not sufficient. each xkos:conceptassociation may have input or source skos:concept(s) and output or target skos:concept(s). the complete collection of such associations for all the concepts in two skos concept schemes forms a correspondence and is expressed as an instance of the xkos:correspondence class. the xkos:madeof property is used to link the xkos:correspondence to its xkos:conceptassociation components. to those familiar with entityrelationship diagrams, what xkos does is to take the skos:related relationship (property) and “decompose” it into its own entity (class) to solve the n to m relationship problem as well as to be able to add additional properties to the relationship. figure 4 illustrates a simple example of a concept association: three classification items are re-combined into two. the xkos:conceptassociation is similar to the correspondence item in the neuchâtel model, but it can describe in a single instance the relationship of any number of source concepts to any number of target concepts rather than expressing the association through a set of pair-wise relations. the xkos concept association can also represent the item change class of the neuchâtel model. however, in this version, xkos does not define any properties or sub-classes for xkos:correspondence and xkos:conceptassociation for modeling the different types of correspondences that we described above, nor can xkos describe the typology of item changes detailed in the neuchâtel model (annex 3). these may be added in a future version. semantic properties semantic properties constitute the second direction in which xkos extends skos. concept schemes are not just lists of concepts: as the skos primer puts it (section 2.3), “the meaning of a concept is defined not just by the natural-language words in its labels but also by links to other concepts in the vocabulary.” skos intentionally defines few properties, but introduces the fundamental distinction between hierarchical and associative relations. in both these categories, xkos creates more precise properties which are described below. the reader can refer to the figure provided in annex 1 to find a panoptic view of skos and xkos properties. hierarchical properties skos defines several hierarchical properties, but the most used are skos:broader and skos:narrower, which are each other’s inverse. these are the two properties that are refined in xkos. a concept is broader than another one if it encompasses a wider portion of the field covered by the concept scheme, and thus includes the scope of the narrower concept. note that the skos:broader property has the narrower concept for the subject and the broader one for the object, for example, “car” “broader” “vehicle”; and the skos:narrower property has the broader concept for the subject and the narrower one for the object, for example, “green” “narrower” “olive”. as we made clear in the previous sections, it is important, at least for statistical purposes, to represent generic and partitive relations between concepts. xkos therefore defines two couples of inverse properties: xkos:specializes and xkos:generalizes on the one hand, xkos:ispartof and xkos:haspart on the other. all are sub-properties of skos:broader and skos:narrower, but the terminology is a bit tricky here: xkos:specializes goes from the more specific concept to the more generic one, and thus is a sub-property of skos:broader. similarly, xkos:haspart is a sub-property of skos:narrower. for example, head ispartof body and chest haspart heart. associative properties in terms of associative properties, skos defines the very general skos:related, and a set of mapping properties (skos:closematch, skos:exactmatch, etc.) intended for establishing links between concepts of different schemes. xkos proposes a hierarchy of skos:related sub-properties that convey more precise semantics. this hierarchy is organized in three branches. the xkos:disjoint property forms a branch of its own. in some circumstances, it is useful to explicitly state that two given concepts do not overlap (for example, private company and non-profit organization in the class-of-work classification of the us current population survey), especially when it has not been specified that the scheme covered its field without overlap (see a.1 in figure 4 above). the second line of xkos associative properties is dedicated to causal relationships. this class of link between concepts is frequently encountered (physics, biology, history, law, etc.). the generic xkos:causal is further subdivided into xkos:causes and xkos:causedby, so that the direction of the causality can be expressed. the last branch of properties is the most populated and deals with sequential relationships; it is represented on figure 5 below. the top node of this branch is xkos:sequential, a refinement of skos:related that just indicates that two concepts in a scheme are in a sequential relationship, for example, notes in a musical scale. below are xkos:succeeds and xkos:precedes that can be used when the sequence has a known order between the concepts. a third sub-property of xkos:sequential is xkos:temporal, which can be used when the sequence is of a temporal nature (i.e., events in time.). xkos:temporal itself is the parent of xkos:before and xkos:after. it was found useful to add two more precise sub-properties of xkos:precedes and xkos:succeeds, namely xkos:previous and xkos:next. previous and next imply that there is no intermediary concept iassist quarterly 2014/2015 55 iassist quarterly between two sequentially linked concepts. these two properties are of course not transitive, although their parents are. conclusion in this paper, we laid out the general rationale and purpose for why xkos was developed. we explained the basic extensions to skos that were identified as needed to describe statistical classifications in the lod domain, and we gave examples from the statistical community to justify our choices. some unresolved issues were also discussed. finally, we gave a rationale for the importance of skos and xkos, appealing to the burgeoning lod community of practice, the use of rdf, and the growth of the semantic web in general. it is interesting that some of the extensions (generic and partitive relations) were originally included in skos. given the amount of discussion in the lod and semantic web communities about semantics and precision, it is even more remarkable that these specific relations were left out. on the other hand, there was a clear desire by the skos designers to make building semantic web applications as simple as possible. since skos is the simple knowledge organization system, this design choice begins to make sense. yet, we have also seen that the worlds of thesauri and classifications are often too complex to model in skos. thus, the vocabulary was quickly extended with skos-xl to handle the need to treat labels, not as literals, but as actual class instances (a process sometimes referred to as reification) that could participate in relationships with other instances and have properties of their own. while skos-xl extends skos for the particular needs of the multilingual thesaurus community, xkos adds the extensions that are desirable to meet the requirements of the statistical community. figure 5: xkos sequential properties skos is a very popular specification, and we hope the xkos extensions will simply serve to increase its adoption. the proof of whether xkos is useful will be found when statistical offices implement it. this work is already underway. however, xkos is still a work in progress, and unresolved issues remain. we hope the users of xkos will offer help with these issues, provide comments to the authors on the effectiveness of xkos, and give guidance as to what other areas should be extended as we prepare to submit the standard as a w3c editor’s draft. acknowledgements the authors wish to thank the organizers of the dagstuhl workshops – richard cyganiak, arofan gregory, wendy thomas, and joachim wackerow – for their support and encouragement in developing the xkos ideas. the authors also wish to thank the participants not already mentioned in the xkos development group: thomas bosch, rob grim, and jannik jensen. notes 1. franck cotton, institut national de la statistique et des études économiques. daniel w. gillman, us bureau of labor statistics, yves jaques, food and agriculture organization of the united nations. the opinions in this paper are those of the authors only and do not necessarily reflect the policies and programs of the institut national de la statistique et des études économiques, the us bureau of labor statistics, or the food and agriculture organization of the united nations. 2. http://www.w3.org/2004/02/skos 3. http://www.w3.org 4. http://en.wikipedia.org/wiki/semantic_web and http:// semanticweb.com/tag/tetherless-world-constellation 5. http://linkeddata.org 6. http://www.w3.org/designissues/linkeddata.html 56 iassist quarterly 2014/2015 iassist quarterly 7. http://www.w3.org/rdf 8. http://www.unece.org/fileadmin/dam/stats/documents/ece/ces/ ge.40/2010/wp.4.e.pdf 9. http://linkeddata.org 10. http://publishing-statistical-data.googlecode.com/svn/trunk/specs/ src/main/html/cube.html 11. http://data.gov.uk 12. http://www.dagstuhl.de 13. http://www.dagstuhl.de/en/program/calendar/ evhp/?semnr=11372 14. http://www.dagstuhl.de/en/program/calendar/ evhp/?semnr=12422 15. http://www.w3.org/2006/07/swd/skos/reference/20090315/ implementation.html 16. http://www.w3.org/tr/rdf-schema/ 17. http://tools.ietf.org/html/rfc3986 18. http://www.ietf.org/rfc/rfc1738.txt 19. http://xmlns.com/foaf/spec/ 20. http://www.w3.org/teamsubmission/turtle/ 21. http://www.obitko.com/tutorials/ontologies-semantic-web/rdfgraph-and-syntax.html 22. http://www.w3.org/tr/skos-primer/ 23. http://www.bls.gov/soc/ 24. http://www.bls.gov/iif/oshoiics.htm 25. http://www.bls.gov/tus/lexicons.htm 26. http://www1.unece.org/stat/platform/pages/viewpage. action?pageid=14319930 27. http://www1.unece.org/stat/platform/display/metis/ the+common+metadata+framework 28. http://eurovoc.europa.eu/ 29. http://id.loc.gov/authorities/subjects.html 30. http://dublincore.org/documents/dcmi-terms/ 31. http://www.w3.org/ns/radion 32. http://unstats.un.org/unsd/cr/registry/isic-4.asp 33. http://epp.eurostat.ec.europa.eu/cache/ity_offpub/ks-ra-07015/en/ks-ra-07-015-en.pdf iassist quarterly 2014/2015 57 iassist quarterly annex 1 skos and xkos properties relating concepts note: skos properties are in the two upper boxes, xkos in the two lower. annex 1 skos and xkos properties relating concepts note: skos properties are in the two upper boxes, xkos in the two lower. 58 iassist quarterly 2014/2015 iassist quarterly iassist 2016 will take place in bergen, norway, hosted by the norwegian social science data services. for any questions please contact: heidi.tvedt@nsd.uib.no 59 iassist quarterly 2014/2015 online application iassist member ($50.00 (usd)) subscription period: 1 year, on: july 1st automatic renewal: no please fill in the information our online form the application is in usd, however, we do accept canadian dollars, euro, and british pounds as well. the membership rates in all currencies as well as the regional treasurers who manage them are listed on the treasurers page the international association for social science information service and technology (iassist) is an international association of individuals who are engaged in the acquistion, processing, maintenance, and distribution of machine readable text and/or numeric social science data. the membership includes information system specialists, data base librarians or administrators, archivists, researchers, programmers, and managers. their range of interests encompases hard copy as well as machine readable data paid-up members enjoy voting rights benefit from reduced fees for attendance at regional and international conferences sponsored by iassist. join today by filling in our online application: http://www.iaassistdata.info/ iassist international association for social science information service and technology a s s o c i at i o n i n t e r n at i o n a l e pour les services et techniques d’information en sciences sociales http://www.iaassistdata.info iassia quarterly 19 public data use: a view from the telecommunications industry in the united states by a. dianne schmidley, staff manager, bell atlantic bell atlantic is one of eight u.s. telecommunications firms resulting from the breakup of the at&t owned bell system on january 1, 1984, the largest divestitiu-e and reorganization in corporate history. bell atlantic owns an assortment of companies engaged in various aspects of providing telecommunications services and products. these companies can be divided into two categories which we call the "enterprises group" and the "network services group." the enterprises group provides services and products to a variety of geographic locations in the united states and canada and it is a relatively unregulated entity. activities of the network services group are concentrated in the seven politica] jurisdictions of washington d.c., delaware, maryland, new jersey, pennsylvania, virginia, and west virginia. management services, incorporated (msi) and the operating telephone companies are part of the network services group. the economic analysis district (ead), my organization, is located in the business planning and financial management department of the msi. the primary occupation of staff members in the ead is the provision of internal consulting support to the network services group, which is dominated by the concerns of the operating telephone companies: the chesapeake and potomac telephone companies of washington, d.c., maryland, virginia, and west virginia, new jersey bell telephone company, bell of pennsylvania, and diamond state telephone company in delaware. these concerns can be divided into four functional areas: regulatory, personnel, facilities plaiming. and marketing. because the network services group is almost wholly comprised of regulated telephone companies, "regulatory issues" are the most important concern of the ead. state regulatory agencies, often called public utility commissions, determine, through pricing decisions, who will bear the burden of the rates the companies charge to recoup operating expenses and guarantee the investors in bell atlantic stock a competitive rate of return on their investment dollar. demographic and economic analyses of the size, distribution and composition of each company's market provide the basis for determining the effects of various pricing configurations. summer 1986 20 iassist quarterly analytical work undertaken by bell atlantic economists and demographers has been made more complicated by the divestiture, since the breakup of the bell system hterally led to the breakup of telephone-served geography. the seven jurisdictions we serve contain 19 "local access and transport areas" or latas, which do not correspond to any other political or statistical entity, although they are associated with metropolitan statistical areas (msas) in many cases. these latas, or large exchange areas, obtain their external communication links from the interexchange carriers (lecs), telecommunications companies engaged in long distance calling services which the local telephone companies (such as those owned by bell atiantic) are constrained from offering, owing to federal regulatory restrictions. relations between the lecs and the local telephone companies are regulated by federal communications commission (fcc) rulings, legislative requirements established by the u.s. congress, and executive branch decisions made through entities such as the justice department and the federal coiut system. in order to comply with the various rulings and legislative mandates, and sometimes question their logic. bell atiantic must have knowledge of the size, distribution and composition of the market within and between the latas. in addition to the regulatory function, another important activity of the companies the had supports is the personnel function. whether we are addressing equal employment opportunities issues, employment site location studies, force planning or how to strategically locate our work crews relative to population growth and migration chum, we turn to data available in the public domain to answer questions about the size, distribution and composition of the local labor force and the telephone-served population. our third major area of support for the operating telephone companies involves the "facilities planning" function. the telephone companies operate networks, which consist of central offices containing switching systems (large computers), miles of cable, and microwave towers. we are constantiy concerned with plant capacity and demand for our services which translates into changes in demand for central office switching and transmission capabihty. population and economic forecasts based on public data make the forecasting of demand possible and enhance oiu ability to plan efficientiy and effectively. although we are heavily regulated by the government, we do have many marketing concerns, and the market we serve constitutes the fourth area of functional responsibility for the ead. some of the more familiar marketing efforts the telephone companies engage in include the distribution of the white and yellow pages directories, and the provision of operator services such as call completion and information retrieval. in addition, we offer products such as business to business directories, and services such as cable television access. we serve the government at the national, state and local level. we serve large industries, such as the steel mills in west virginia and pennsylvania, and small enterprises such as a savings and loan company in maryland. we serve the elderly, the handicapped, homeowners, travellers, and, a new customer since the divestiture, interexchange carries [sic]. knowledge of the consitutents of this market is derived from public data coupled with our own internal surveys. what are the kinds of pubuc data used by bell atiantic? generally, we use as much of the demographic and economic data as we can obtain from the federal, state and local governments, whether it comes from censuses, surveys or administrative records, but our most important source of demographic or socio-economic information is the 1980 u.s. census of population and housing. these data are available in many forms: published, on microfiche, and on magnetic tapes. the problem is that there is more data than we can handle, so we have implemented an online summer 1986 iassist qvuirterly 21 demographic data retrieval system to assist us. as i mentioned earlier. bell atlantic has a imique problem. the divestiture left the company with odd service areas called latas. in order to provide information to our companies, the ead modifies public data from the economic and demographic censuses and surveys to make it conform to the geographic area bell atlantic serves. prior to the divestiture, the operating telephone companies were concerned with the same geographical areas they serve today. the imit of concern, however, was the wire center area or central office district (cod) as it is sometimes called. to complicate matters further, local exchange areas (smaller and different from the latas described above) were also a concern. fifty years ago, all three entities were represented by the same geographic area, corresponding to a community or settlement technological change, which allowed the newer central offices to serve more than one of the old wire center areas, population change, and concessions to consumers with regard to their calling access led to an erosion of this one-to-one conespondence. as a result, the telephone companies not only have served and continue to serve areas unlike any other known pohticjil or statistical geographical areas, they serve a number of entities that do not correspond to one another. because of the continuous need to determine the demographic/economic characteristics of telephone service areas, in order to address the functional areas described above, the requirement for tailored public data arose long before the divestiture. the key to tailoring the demographic and economic data used to develop construction plans, engage in force planning and answer the questions of the regulators is census geography. census tract and block group information, aggregated to user described areas, is the rosetta stone of managers engaged in economic and demographic analysis. with the advent, in 1970, of the first fully automated census, the laborious task of aggregating census tract and block group information by hand became, mercifully, obsolete. today, there are three major methodological approaches underlying the automated demographic data retrieval systems which provide information for user defined geography. 1. federal information processing codes (ftps) are assigned to every pohtical and statistical entity in the united states. this means that all political and statistical geographic units, such as states, counties, msas, and census tracts, have unique identification codes. in the automated system, the user can retrieve information associated with these codes. this approach is efficient if the user is seeking information for a list of states, coimties, or municipalities. the first attempts to aggregate data for user described areas, such as wire center areas, were based on combinations of block groups/census tracts, and relied on this mechanism. when thousands of geographic units were involved, however, (the old bell system had 10,000 wire center areas) this particular approach proved to be extremely time consuming, even after automation. 2. the assignment of geo-coordinates (latitudinal/longitudinal coordinate points) to census data provided the basis for a major breakthrough in the automation of demographic data retrieval. every census block in the united states received, in 1970, a centroid assignment of a unique set of coordinate points. the centroid is the geographic or population center of a block area; there are variations in the way these assignments are made, but discussion of this topic is beyond the scope of this paper. in 1970, point assignments developed by the u.s. bureau of the census, were listed in the master enumeration districts list (meds) and in 1980, census bureau point assignments were listed in summer 1986 22 iassist quarterly the master area file reference list (marf). in 1990, they will probably be found in the topographically integrated geographic referencing and encoding system (tiger). in themselves, the centroid assignments are useless for solving the problem of demograhically describing user defined areas. software linking the coordinate assignments and user described boundaries of study areas, which have been transcribed into binary code, are needed to complete the demographic data retrieval operation. to date, most of the software for this type of application is owned by non-governmental sources, and licensing arrangements must be purchased in order to make use of the private sector product before the divestiture, the operating telephone companies in the jimsdictions now served by bell atlantic had transcribed their wire center area boundaries into binary coded polygon files. census data based on the meds and marf assignments could be aggregated to produce demographic profiles of the user described areas. since latas are aggregations of wire center areas, all that had to be done after the creation of the latas was to aggregate the wire center polygons into lata polygons. as was mentioned earlier, latas axe also aggregations of the smaller exchange areas, and/or cod areas. at the lata level, however, the difference between the three telephone entities (wire center areas, central office districts, and exchanges) disappears. thus, aggregating the wire center areas to equate to latas does not cause discrepancies. the final result is census data tailored to our lata areas. 3. the third type of geographic linking system available for tailoring govenmicnt produced demographic data to meet user defined needs is the geo-based files/dual independent map encoding or gbf/dime process. briefly, this process makes possible the matching of census address records for urbanized areas with user records. in the case of the telephone company. these are customer records. customer records processed through the gbf/dime program can be linked, at the census tract level, with specific socio-economic characteristics. this process was used to provide the washington, d.c. public utility commission with information concerning hnks between telephone availability and characteristics of the inhabitants of areas under study. this process is more limited than the other two, however, since the gbf/dime files are only available for urbanized areas. at bell atlantic, economic data available from the government for political or statistical areas are disaggregated into user defined telephone service areas through the use of population weights derived from the centroid point assignment process described above, or the feps code process. since economic data are available from the government for the whole counties contained in the latas, disaggregation only occurs in the case of spht counties. the census tract components of counties are assigned to their respective latas using the procedures outlined in 1. and 2. above. the census profiles developed through the use of the geographic unking systems provide the basis for developing time series data and forecasts of population through iterative proportional fitting schemes, when linked to historic and forecast information for the aggregates of counties which correspond to the latas. economic data, in turn can be derived through the use of the time series and forecast versions of the population weights. thanks to the geographic linking processes developed jointly by the government and private sector firms with software capabilities. bell atlantic is able to address problems in the major corporate functional areas outlined earlier in this paper, utilizing public data as they relate to our odd geographic areas, n summer 1986 lassist newsletter, vol. 2, no. 1 (winter 1978) editorial comment t year news ilea ms. of w firm and gan form begi publ tor tion cere tlie year his i of pu letter fion, al ice iscons found also p for t ation nning icatio and th have desi high s by al ssue b blicati the under bobbin in, ha ation f rovided he dis within of th howe e forma changed re to tandard ice. egm on firs the of diso or t a m semi lass e se ver, t o upho s se s the of the t year edito the un n, pr he pub uch ne nation istcond both f the it is id and t in t se ias of rshi iver ovid lica eded of with year the publ my ex he f cond sist pubp of sity ed a tion orinthe of ediicasintend irst on in f o the d here, produ ditio at th a tex the netwo scrip the p are i ing pm m netic delig enter es of vo newsl pu5ir as we est t exten time tion go ou reaso tide net wo from sist year, prese vide tion e of rmat ouble t ces a n of e sa t edi kentu rk ( t) is ublic ntrod apers achin tape hted ed us pecia lume etter caeio 11 as o ias ds t probl :o s ; unt i i h ;--al :ks 5aper ;onf e the fro col he d pub a ma me t ting cky cms use atio ucin or e re "to ing wa n of pro sist hat ems ubmi il t ave 1 q and s gi renc ke h in info ough ican obv m th umn oubl lica gazi ime sy educ eait d to n, g is oth adab in f have scri with t wa s mo sub vidi mem tren this t ar his sel eali ne ven e in ope this rmat t c t im lous e fi style e co tion ne or cons stem ation or a ente one i tha er c ie f act , ^pap pt co the s cle ving stant ng ne bers. d. imp! tide issue ected ng wi twork at t febr that issu ion oncer porta ait rst ill lumn in jou erve aval al nd r an nnov t of ontr orm i ers mman fere year ustr la the rnal s pa labl comp sate d f o atio ace ibut on woul air ds. nces is ated yout traand per . e on uter rloo rmat n we eptions magd be eady fourth ar th t o w a r ive ar ws of this becau icit i cou fo sever th co ing--d he 197 uary o the ar e wil and st ning a nee t issue at the d the tides interissue se of nvitaid not r that al armputer erived 8 lasf this tides 1 proimulan area o our ted ome f th of signif enterprise the term "network," as used by those of us who use computers, may refer to communications networks, computer networks, or both. the papers in this issue are directed primarily toward computer networks although many allusions are made to the communications networks on which the computers rely. what are computer networks? one common 8 uterne or users work . main-s and/or genera munica withou networ munica ly use net wo more via so the c ite, remo liy it tions t com ks can tions d defin rk is comput me com omputer or h te com may be ne two puters, not exi network itio the ers muni s ma ost puti sa rks b st n of a follow accesse cations y consis comput ng syst id that can e ut comp without cominq: d by nett of ers, ems. comxist uter comi atte char port esta netw the on t the sour espe comra borm eval soci base n the f mpt to acterist unities blishmen orks. impact he socia problems ca shar ciaiiy f unity, an . uates th al struc d scient ollow expl ics, prov t and alic of co 1 sci and ing or t is di and e eco ture if ic ing pa ore s proble ided use e robb mputer ence d poten taroug he soc scusse richar nomic in t commun ges we ome of ms, an through of com in exp netwo ata lib tials o h netw iai sc d by lo d roist climat he net ity. will the d opthe uter ores rking rar y. f reorks, ience raine acher e and workwhile att the inclusio articles, that the ne means of com members arou suit, the n to provide s reports and around the ution to t newsletter w aeear—if, pecially lik are attempti letter, plea loolcing rorw newsletter i empting t n of mor we have wsietter municafio nd the wo ewsletter pace for for news world. hese seg ill be gr for any r e, or dis ng to do se write ard to wo n 1978. not is ns f rid. wil act and your ment eatl easo like with or c rkin ove ubst for a p or 1 as 1 co ion note co s o y ap n, y , w the all. g wi toward antive gotten rimary assist a rentinue group s from ntribf the preciou eshat we newst am th the substantive papers and book reviews will be accepted throughout the year and scheduled for publication as space becomes available. material for "news and notes" and "action group reports" should conform to the following deadlines for the june, september, and december issues: may 15, 1978 (june) august 15, 1978 (september) november 15 (december) material received after these deadlines will be held for the next issue, providing the notices are not time dependent. your help in observing these deadlines will be appreciated. t-w-h. 2 lassist newsletter vol. 4 nos.3&4 president's report on 1981 iassist—ifdo conference alice robbin this is a report on a planning meeting for the 1981 lassist—ifdo conference that i attended in my capacity as president of lassist. the report summarizes agreements on the objectives of the conference, its contents, audiences, and lasslst's role and responsibilities on october 6lh and 7th, 1 attended the meeting in paris, france my trip was funded by the national science foundation as part of its program to help international organizations sponsor and plan professional society meetings the interests of lassist correspond rather well with the objectives of the foundation's travel grants and, in particular, with the french-american program of the international program division in attendance was the executive planning committee, composed of french members, f. bon and b. bouet of the university of grenoble, j -p grem y of the university descartes (and also a member of lassist), and myself. g. martinotti, president of ifdo, was unable to attend the meeting was a delight, and we worked steadily with few interruptions for hours on end, all of us committed to completing a long agenda of items and each recognizing that we had very little time to exchange information. i am most grateful for the careful pre-meeting planning and enthusiasm that the french members of the executive committee demonstrated and the good humor that reigned throughout the long days. here are the particulars: title: the impact of computerization on social science research: data services and technological developments/ l'impact de rinformatique sur les recherches en science sociales: les banqucs de donnees et les developpements technologiques date: september 14-18, 1981 place: university of grenoble, grenoble, france conference organizers: lassist, ifdo, laboratoire d'lnformalique pour les sciences de i' homme (lish)/ centre national de la recherche scientifique (cnrs), and the university of grenoble working languages: english and french (sessions will have both french and english speaking professionals who will act as infonnal intermediaries.) registration fee: $.^5 us. or 150 ff. preliminary schedule: • end of october: 'call for papers' (by this time all lassist members should have received a copy.) • middle of february 1981: intention to submit a paper should be submitted to the french planning commitee. • beginning of april 1981: abstract of paper submitted to the french planning committee (no more than 500 words). begirning of may 1981: notification of acceptance of papei • end of june—beginning of july 1981: preliminary agenda for the conference this schedule should allow ample time for lassist members to decide whether they want to attend the grenoble meetings and to apply for institutional and/or extramural funding support. (and also to review our french) conference agenda: plenary session: title: computers and information: their societal impact new types of research will include the processing of large ecological files and survey, textual, historical, and biographical data, complex data (historical, biographical, time series, genealogical, topological, and textual): their creation, structure, administration, and preservation, and problems ansing from their exchange; secondary analysis; computer mapping, and problems of aggregation and disaggregation new institutions will include development of data banks and their perspectives; politics of creating and disseminating data, longterm storage, sociological approaches to the data worid (producers, users, services); information systems about data and machine readable files. [57] lassist newsletter vol. 4 nos.3&4 new tools will include networks, shared and distributed data bases, on-line data bases; microand minicomputers; data base management systems for complex data bases and perspectives on the application of artificial intelligence and knowledge representation; compatibility (through translation. pivots languages), and scienlific information retrieval ' relations between data producers and researchers will include data description and documentation, user needs for information and data, economics of data services; information on organisational sources for public and commercial data; political, organizational, and technical problems associated with process-produced data (public/administrative records), public opinion polls; their quality, effects on political decision-making, access conditions; anonymization of data (and deductive disclosure). as you can see this is a broad methodological, technical, research and policy agenda for the conference. it permits wide participation by members of lassist, who can define their own contributions within the proposed topics the breadth of subjects is indicative of the wide range of interests by members of lassist. ifdo, and the french social scientific community cartographic, and the forth on satellite data) topic 2: the organization and management of data services. this is viewed as a working session, lasting either a half or a whole day (in conjunction with topic ?i). during which time at least one individual will discuss the political, economic, and administrative problems associated with the organization and management of data services; budgeting; personnel recruitment and training (but not to be confused with the next topic); building a collection; user services; relationships of personnel to users, software and computer; maintenance and perservalion; user needs, linkages to other information services, etc. topic 3: the formation or a professional data archivist and librarian. this is viewed as a working session, lasting either a half or a whole day (in conjunction with topic 2). during which time at least one individual will discuss problems in recruitment and training of data services' personnel, based on previous experience, and knowledge necessary to provide data services hands-on experience in the various aspects of understanding data and providing data services will be provided. workshop agenda: a one-day series of workshops devoted to three topics of substantial interest to lassist members and the french social research community and government agencies has been proposed the workshops will be the responsibility of lassist and its members lassist members should immediately signal their interest in participating in a workshop, submit an abstract of the objectives of their contribution, and as extensive as possible an outline of the proposed discussion (including the estimated length of time required for the presentation) copies of the abstract and outline should be sent to lassist 1481 workshop chair. lame ruus. and members of the program committee (robbin. mcmanus. rowe. von brunken. gavrel, and schrik. addresses found below). lassist will attempt to secure travel and lodging suppori for those individuals selected to participate in these workshops (at the same time, individuals should attempt to locale matching funds) the program committee will develop an extensive outline of the contents of each of the workshops, to be forwarded to various funding agencies this outline should be completed by the middle of december; therefore, lassist members wishing to participate m these workshops should contact lassist program committee members and the president as soon as possible. topic i: the assay/evaluation of survey, ecological, cartographic, and satellite data. this is viewed as a working session (read "technical"'). lasting an entire day. during which lime four individuals will present the conceplu.il and methodological activities required to evaluate the quality of data (one person on survey data, the second on ecological, the third on exposition/demc nstration during the conference: three areas of interest have been identified, and others are sought. build the exposition around minicomputing and social research: display of materials created and used by lassist and ifdo members, with demonstrations during the conference • invite data base vendors (e.g.. viewdata) to demonstrate their wares to the social science community invite the network people of euronet and regional and international network organizations (e g . chronos data bank of the eec) to demonstrate on-line access and relneval funding the cost of attending this conference should not be minimized in an effort to assist social and information scientists' pariicipation in the grenoble conference, a variety of funding strategies are suggested. first, as president of lassist, i have made application to the u.s. national science foundation for a group travel grant on behalf of us social and information scientists if awarded, the grant would provide i."^ u.s. scientists with air travel between their home institution and the conference (an average grant of $884) the award will be competed for and is open to all us scientists. selection criteria include a good distribution between young scientists whose professional development will profit by attendance and experienced scientists who are knowledgeable abt>ut social science information lassist newsletter vol. 4 nos.3&4 infras(ruclure development and can conlnbute lo cooperative social science research activity; good geographic and institutional distribution, and a tnix of conlnbuted papers, invited papers, and w9rkshop organizers the selection committee is composed of social and information scientists who have long been involved in efforts to develop social science data services and in organizing access lo machinereadable data. the committee is chaired by dr joseph w. duncan, director of the us office of federal statistical policy and standards should the grant be awarded, lassist members and other social scientists will be immediately informed through their professional journals instructions will be supplied on to whom and in what form your proposed paper should be described the grant proposal will be submitted by january 1 , 1981 1 hope that a decision will be forthcoming by april isl and that the selection committee will make its final decision by the end of may you are very much encouraged lo respond immediately to the 'call for papers' (directly to frederic bon) i will be kept informed by the french organizers about the lassist members' and other social and information scientists' interest in participating in the grenoble conference. should you have any questions, please feel free to call me second, lasslst members should make every effort to secure institutional or extramural funding support if they wish to attend this conference national research funding agencies (e.g.. social science research councils and the like) have limited amounts of funding for individual travel grants. for example, the u.s. national science foundation has a limited budget for individual awards to individuals who have been invited to organize a special session or lecture at a plenary session of an international scientific meeting other national funding bodies have funds for individuals who will give papers/communications in sessions at international scientific meetings. since there is ample time to make special requests for institutional support, this strategy should be employed. third, an effort will be made to interest data processing and minicomputer manufacturers to assist in the conference. all ideas on this front are solicited from the members of lassist please contact me if you have any thoughts on this matter. lassist is an iniernaiional organization it was organized lo refiecl the needs of professionals involved in providing data services. also included in its mandate is the objective of providing the expertise from among us members to assist in developing data services in countries where this need has been articulated. the french scholarly community has requested our assistance in the planning and in the success of this conference it behooves each of us to diffuse the knowledge and experience we have gained in providing data services. this international conference will allow us to do just that. 1 look eagerly for your enthusiasm and commitment to making this conference a success. lassist 1981 workshop chair lalne ruus, data library. computing centre. university of bntish columbia, 2075 wesbrook place, vancouver, british columbia, canada v6t iw5 program committee members (lassist) 1 . alice robbin. president. lassist. oata and program library service. university of wisconsin—madison, 4452 social science building, madison, wisconsin 53706 (tel. no.: (608) 262-7962) 2. sue gavrel, machine readable archives, public archives of canada. 395 wellington street. ottawa. ontario kia 0n4 canada 3. judith rowe. computer center. princeton university. 87 prospect avenue. princeton. new jersey 08544. usa 4. henk schnk. slelnmetz archives. herengracht 410-412, 1017 bx amsterdam, the netherlands 5. nancy mcmanus, social science research council, 1755 massachusetts ave., n.w. washington, dc. 20036 usa 6. erika von brunken, medical information center, karollnska institutel, s-104 01 stockholm. sweden [59] iassist quarterly iassist quarterly 2013 5 abstract in recent years, the level of detail in confidential data made available to social scientists has increased dramatically. one particularly important growth area has been ensuring that research outputs do not present any residual disclosure risk. traditionally this has been managed by specifying rules for researchers (the ’rules-based’ model), but it is increasingly recognized that a ‘principles-based’ approach is both more secure and more cost-effective. the principles-based approach requires a higher level of expertise from those managing access to data, and places the subjective assessment of risk at the forefront of decision-making; these two factors make data managers uncomfortable. in addition, knowledge of this approach is concentrated amongst a relatively small community, whereas the rules-based approach has dominated for half a century; data managers may not be aware of an alternative perspective. this paper reviews the arguments for the two different approaches. they are not mutually exclusive: both take simple rules as a starting point, but the rules-based approach also finishes there. this has advantages in some circumstances, but the value of the principles-based approach increases with the sensitivity of the data and the scope of researchers to innovate. the paper considers how the two approaches can be implemented. although the principles-based model requires greater initial investment by both researchers and those managing access to data, this can bring substantial auxiliary benefits to the latter. the paper therefore concludes that a principles-based approach is generally preferable, and it is essential for the remote research data centres which dominate access solutions for the most sensitive data. keywords data access, data security, statistical disclosure control, principles-based, output sdc, researcher management . acknowledgements we are grateful for comments by members of the administrative data research network; don webber and richard welpton; and the anonymous referees, who suggested the inclusion of the summary table in section 4 introduction since the early 2000s, social scientists have seen an explosion in data availability. the most important has been the increasing research access to highly confidential but high utility data (see for example trewin et al., (2007) or elliot and purdam, (2015) for a discussion of this trend). this has been made possible by the development of secure data access solutions. principlesversus rulesbased output statistical disclosure control in remote access environments by felix ritchie1 and mark elliot2 the main advantages of a rulesbased model is the certainty and lack of ambiguity 6 iassist quarterly 2015 iassist quarterly these allow the managers of such facilities to control the data access process with a great deal of security, but researchers are no longer necessarily restricted to physically visiting the facility manager’s site; instead, they increasingly do so through remote access systems. some organisations have invested in secure remote job submission, where the researcher submits code to a server holding the data and receives back statistical results; for some data this is the most appropriate model, but for many types of data the dominant model is the remote research data centre (rrdc), where researchers access data through ‘thin clients’ (that is, where all the processing is done by the server holding the data, and the researcher controls the research through a web browser, for example). these allow researchers almost complete freedom to work with the data (other than removing them from the facility) and so are popular with researchers; they are also popular with data access managers, as such systems still allow a high level of control to be applied without the need for continual surveillance. a key element of that control is checking to ensure that statistical outputs do not present any residual disclosure risk: a researcher will use the data to produce statistical outputs, and it is possible that those outputs could inadvertently breach the confidentiality of the underlying data (for example, by revealing that only one person in a small area has a particular illness). a critical point here is that when a researcher publishes output they are effectively moving data (for output is still data) from a highly secure setting to a completely insecure one. the change in data environment, means that the status of the data can, in principle, change from non-personal to personal (see mackey and elliot, 2013, for a review ). hence, disclosure checking of output is a vital part of the governance of secure data access systems. traditionally, ‘output statistical disclosure control’ (osdc) has been managed by specifying rules for researchers to follow; for example, a requirement for all table cells to have at least three observations contributing to that cell. if the table meets the rules it can be released; if not, not. this ethos is reinforced by a half-century of statistical research focused on making tabular outputs safe (hundepool et al., 2012). however, in the last ten years it has become clear that simple models for tables have limited value in modern research environments, and increasingly common to discuss ‘output sdc’ as a separate research field3. the complexity of outputs led some data managers (e.g. ritchie, 2007) to argue that the rules-based approach was both unsafe and inefficient; instead, a ‘principles-based’ approach could be both more secure and more cost-effective. the principles-based approach uses rules to first-approximate a decision to release or not; but all preliminary decisions are subject to review and change if the researcher or the person responsbile for approving the release of output can make the case. the key is that both parties agree on the aims of the sdc process – and one of those aims can be to use the resources of the facility efficiently. the principles-based approach does not accept that outputs can be definitively classified as safe or not, only that the balance of probability says so. finally, the principles-based approach acknowledges the value of research output in any decision. these differences may seem subtle, but they have profound implications for the way the facility and the researchers are managed. the principles-based approach requires a higher level of expertise from the managers of research facilities, who must have both technical knowledge and an understanding of the research environment and researcher. in addition, it places the subjective assessment of risk at the forefront of decision-making. these two factors often make facility managers uncomfortable, as such organisations are typically risk-averse (ritchie, 2014a). in addition, knowledge of the principles-based approach is concentrated amongst a relatively small community, whereas the rules-based model has been the dominant approach for half a century. hence, facility managers may not be aware that there is an alternative perspective; if they are, they may not appreciate the subtleties of the principles-based approach, preferring instead a simpler model of data security (ritchie and welpton, 2014). this paper aims to help facility managers make decisions about sdc, using the current best understanding of the pros and cons of each of the two process methodologies. for explanatory purposes, it uses the example of deciding on an approach to sdc for a remote rdc. this is because this is the case in which the difference between the two is starkest, and this paper demonstrates that the value of the principles-based approach increases with the sensitivity of the data and with the degree of freedom that researchers have to innovate. our conclusion is that for remote rdcs the advantages of principles-based sdc are clear. beyond this, there are also lessons for other environments. the two approaches are not mutually exclusive: both take simple rules as a starting point, but the rulesbased approach also finishes there. this has advantages in some circumstances (for example, rules-based is more appropriate for european statistical system outputs; eurostat, 2014), but as the principles-based approach is the generalisation of the rules-based approach, facility managers would do well to consider both. the paper notes that, although the principles-based model requires greater initial investment by both the facility managers and the researchers, the necessary training can bring substantial auxiliary benefits to the facility manager. put simply, the necessity to train researchers gives the facility manager an opportunity to encourage other positive behaviours, leading to increased ‘legitimacy’ and improved researcher behaviour (ritchie and welpton, 2014). again, this is not always feasible, which is why rules-based modelling is sometimes more appropriate. the aim of this paper is to show when these benefits can be realised. the next section considers the definition, benefits and costs of the rules-based approach, whilst section three evaluates the principlesbased approach. both consider the evidence for claims made. this is particularly important for the principles-based model, which brings risk-assessment to the fore. section four summarises the dicussion, and concludes that the principles-based approach is essential for rdcs, whether remote or not. section five reviews implementation issues; section six concludes4. some definitions are necessary for the paper: an ‘output’ in the context of this paper is any statistical product arising from use of the data, intended for distribution beyond the technical confines of the research facility. an output may be a table, graph, regression model, frequency count, survival function etc., or it may be a paper containing multiple statistical outputs. a ‘commentary’ in the context of this paper refers to any discussion about the analysis produced by the researcher. this includes written and verbal discussion. iassist quarterly 2015 7 iassist quarterly ‘statistical disclosure control’ (sdc) means techniques to ensure that a statistical product or data does not breach confidentiality guidelines. ‘input sdc’ (often just referred to as sdc) is concerned with protection of data before researchers have access to it, and is the subject of a different paper. ‘output-based sdc’ (osdc) is only concerned with statistics for distribution, and is the focus of this paper. ‘disclosive’ is used in this paper as a short-hand for ‘output which should not be made generally available as it retains a non-negligible residual risk that an individual population unit could be identified within them’; there might be variations in outputs between what is strictly unlawful and what is unwise and undesirable. for the purposes of exposition, we assume that disclosure equals a breach of confidentiality, with legal consequences and/or ethical implications. the discussion below sits within the ‘five safes’5 framework (desai et al. , 2014; see camden, 2014, or sullivan, 2011, for examples of use), which is a way of identifying sources of risk in data access: • safe projects – whether the data use is lawful • safe people – whether the researchers can be trusted to hold and use the data appropriately • safe settings – whether the manner of accessing the data offers protection • safe data – whether there is any inherent protection in the data • safe outputs – whether the outputs from the research pose a disclosure risk the final criterion recognises that, however well-intentioned and competent the researcher is, accidents can happen – either through ignorance or through complexity. this paper only considers the ‘safe outputs’; that is, it is based on the assumption that data access is lawful and that researchers are not deliberately trying to misuse breach data confidentiality. how the researchers access the data is not relevant to this discussion. inherent protection in the data is mostly irrelevant to this discussion of general principles and procedures, (although it does have some bearing on training issues and rules-based approaches, to be discussed later). therefore, this paper does not place any limits on the data used to generate these outputs. rules-based osdc how it works rules-based osdc consist of the application of a set of rules to determine whether an output should be released or not. example rules might be • “a table may only be released if there are at least three observations for each cell” • “a regression may be released if not based entirely on categorical data” • “a herfindahl index of over 0.3 should only be released as ‘over 0.3’” • “variance-covariance matrices x’x may not be released” these are hard rules; they are expected to be applied consistently. this generates certainty in what output is acceptable, and allows machine-based sdc to be applied rules-based osdc rules-based osdc is simple and transparent. it is popular with data owners as it reflects the guidelines used to create official statistics, which are largely tabular. it is also necessary for the increasing number of automated systems which allow researchers to produce tables and analysis on the fly. the main advantages of a rules-based model is the certainty and lack of ambiguity. this allows untrained (or partially trained) staff to clear outputs, and requires no training on the part of the researcher. researchers can be given all the information they need in the form of hand-outs. this, for example, is how researchers at the eurostat safe centre have been advised. criticisms of rules based osdc there are three main disadvantages with the rules-based approach. all three are a consequence of the inevitable trade-off between confidentiality and efficiency problems. consider devising a rule for the number of observations that have to be in a cell for it to be released: • the confidentiality problem: a low limit increases the probability of disclosive cells being published • the efficiency problem: a high limit increases the probability of non-disclosive findings not being published first, no rule can guarantee non-disclosure, and the application of strict rules can provide a false sense of security. for example, in some cases no amount of units in a cell prevents that cell breaching confidentiality in some way. hence, the rules based approach may under-protect the data in some cases second, a rules-based approach tends to over-protect the data. protection will tend to dominate user value: the rule is the only protection against disclosure, and hence has to be much stricter than if other factors (such as the specific data being tabulated) are taken into account. as well as being inefficient, this can also create credibility problems when expert users are asked to follow rules which do not make sense to them. third, rules cannot cover all conditions; new rules need to be devised and agreed as new possibilities occur. a proliferation of rules is possible and this can in turn lead to contradictions. consider defining dominance rule for table cells of n units (that is, determining whether one or two units contribute so much to a cell total that the cell can be considered, for all practical purposes, to only include those units). rules that have been put forward include: • the top unit does not account for more than w% of the total of the bottom n-2 records • the top unit does not account for more than x% of the cell total • the top two units do not account for more than y% of the cell total • the herfindahl index does not exceed z% each rule identifies a potentially problematic distribution of data, but the rules will not necessarily agree. requiring all four rules to be met is over restrictive. finally, unless the person clearing the output knows how the output was created, it is not possible solely from the output to determine whether the rules have been met or not. 8 iassist quarterly 2015 iassist quarterly in summary, the simplicity of rules-based osdc is also its main limitation, particularly when dealing with an expert user base when rules-based osdc can suffer from credibility problems. the blunt instrument of rules-based osdc can under-protect in some cases, but is more likely to over-protect. this can cause frustration in researchers, which in turn is one of the factors associated with confidentiality breaches (desai and ritchie, 2010). principles-based osdc (pbosdc) how it works principles-based osdc is characterised by: • researchers and output checkers both trained in sdc • rules-of-thumb rather than hard rules • freedom to approve any output in principle • no duty to release any output • responsibility for producing good output resting with the researcher • output checkers considering the value of the output • output checkers considering resource constraints pbosdc starts from the same perspective as a rules-based model: a set of rules exist to guide output approval. the difference is that these now become rules-of-thumb, rather than hard rules. they are there to guide the output-checker, but do not necessarily need to be followed. the output checker has complete freedom to exercise discretion, both to release an output and to decide not to release it. clearance for release thus becomes a negotiation between researcher and approver. but an unrestricted negotiation is inefficient and so the rules-of-thumb provide the starting point. it is important that both parties recognise the costs and the benefits of checking an output for clearance. both parties want outputs to be processed quickly. the researchers want their outputs to be cleared; approvers want to be satisfied that the outputs are non-disclosive. a rejected output imposes costs on both parties; both parties therefore want to avoid this. the key to pobsdc is that the person best able to assess whether an output should be released is the person who created it. the researcher knows whether the data was sampled, whether there are any dominance issues, how many observations there are in a cell, and so on. most importantly, the researcher knows the value of the output to his or her research. if the researcher knows the broad criteria on which output is checked he/she can ensure that (1) the output meets those criteria (2) the output checker has the information necessary to come to the same conclusion quickly and easily. if the output does not meet the prima facie conditions but is of high importance to the researcher, the researcher can try to persuade the output checker that this case is a valid exception to the rules-of-thumb. as this is likely to involve effort on both parties (and the output checker is under no obligation to concede the arguments), the researcher will have to consider whether the output is worth it. the final part of the system is that the outputs can be rejected not just on the basis of disclosiveness, but on the basis of whether they are a good use of the output checker’s time. it is irrelevant how much time the output checker actually has; the point of this rule is to allow the output checker to reward compliant behaviour and punish the malcontents by setting up appropriate incentives (welpton and ritchie, 2011, ritchie and welpton, 2012). consider a (male) researcher wanting to produce a number of outputs which are (to his mind) non-disclosive, but nevertheless breach the rules-of-thumb. he has three alternatives a. send the outputs through and hope for the best b. not send the outputs through c. raise the issue with the output checker and try to identify a solution which works for both parties option (a) is often a good strategy as a one-off; output checkers should be tolerant of researchers periodically sending through output which requires more work to check. however, as a repeated strategy it risks aggravating the output checker who can delay or stop checking future outputs, particularly if the researcher shows a failure to understand the principles of clearance. option (b) is common in practice; researchers working in restricted environment typically produce a large amount of outputs, not all of which is wanted. as a general rule, pbosdc aligns with good statistical practice; for example, in discouraging very low cell counts. option (c) is the most interesting one, and frequently used in restricted facilities using pbosdc. researchers learn that a ‘no surprises’ policy appeals to output checkers, who also want an efficient clearance process. this can lead to inventive solutions which work for both parties; for example, validating program code rather than outputs. hence, researchers are incentivised to produce good output, and are encouraged to talk to output checkers before problems arise. output checkers are encouraged to promote good practice amongst researchers, and to listen to user perspectives on the value of certain outputs. to be most efficient, a pbosdc process should adopt the ‘safe’/’unsafe’ statistics dichotomy (ritchie, 2008; brandt et al., 2010; ritchie, 2014b). a ‘safe’ statistic is something which has very little or no inherent disclosure risk, such as regression coefficients. an ‘unsafe’ statistic is one which presumed to present a disclosure risk unless proved otherwise; for example, a simple tabulation. ‘unsafe’ statistics, being ex ante disclosive, are much more time consuming to check. under pbosdc, researchers can expect that ‘safe’ statistics will be cleared unless the output checker can demonstrate that it is disclosive; by definition, there will be almost no instances where this is the case. in contrast, ‘unsafe’ statistics will not be cleared unless the researcher can demonstrate that there is no disclosure risk. ‘unsafe’ statistics passed for clearance impose a cost on the researcher related to the output checker’s cost of clearance. the researcher should therefore be incentivised to concentrate on producing outputs made up of ‘safe’ statistics; and because the researcher is aware of how the value of specific outputs, there is an incentive to focus on important results and not large quantities of ‘nice to have’ information. this dynamic needs to be emphasised in the service access training. finally, pbosdc ideally attempts to raise awareness of speculating about the identity of records when commenting on results. for example, although a researcher may produce good statistical outputs, he or she might draw attention to the presence of a particular outlier which affects results. this cannot be dealt with by iassist quarterly 2015 9 iassist quarterly a rules-based approach, but only by ensuring that researchers are aware of the risks posed by incautious discussions of data quality. in summary, the aim of pbosdc is to create an atmosphere where clearance is seen as the joint responsibility of researchers and output checkers. the common interest encourages the production of safe and useful outputs cleared by an efficient process. however, for this to happen effectively, both need to understand the incentives of the other party, as well as the principles of sdc. hence, there is a need for training of researchers and of output checkers. this is a major difference compared to the rulesbased approach. advantages of principles based osdc the direct advantage of pbosdc is that it should allow both more and safer outputs than the rules-based approach. consider again the potential errors noted above arising from a threshold rule, the confidentiality problem (too low a limit —> disclosive outputs) and the efficiency problem (too high a limit—>safe outputs rejected). under pbosdc, there is no conflict: a high limit, likely to remove almost all disclosure risk, is set as the initial rule of thumb. if a researcher feels that this is too high in a particular case, he or she can make that argument, confident that the arbitrary rule of thumb will be replaced by a review of the circumstances in the specific case. this works because most researchers requiring access to detailed data want it for analytical purposes, which are more likely to be safe statistics. when unsafe statistics are the main focus of output (for example, the oecd commissioned work on high-growth companies which required very many tabulations from the vml), output checkers and researchers have an incentive to work together to agree in advance the range of permissible outputs. the indirect advantage of pbosdc is the ability to develop a culture of confidentiality awareness amongst researchers. this has positive feedback effects: a demonstrably educated and trustworthy researcher base can have systems designed to reflect that knowledge and trust, and a more appropriate working environment encourages positive behaviour from researchers (desai and ritchie, 2010; ritchie and welpton, 2014). this has been the pattern of development in uk research data centres over the past decade. criticisms of principles based osdc the three main criticisms of pbosdc are uncertainty, inconsistency, and resource requirements. pbosdc, by design, introduces uncertainty into the process, as the whole ethos of pbosdc is that decisions are taken in specific contexts. training and ongoing engagement are therefore required to build trust. if several output checkers are reviewing outputs, they may make different decisions; it is also possible that the same output checker will make apparently inconsistent decisions in different circumstances. this is because the checkers are making decisions responding to the particular circumstance of output, data and researcher and not just a rule about an output. to ameliorate this good training (of both researchers and output checkers) ensures that there will be agreement on the principles. an important part of the process is that is that there is no rules-based yes/no answer; therefore output checkers should be aware that any decision necessarily has an element of subjective judgement and may need to be justified. it has been argued that the need for output checkers rather than automatic processes (or checking by staff with limited expertise) increase the costs of a facility, as does the explicit allowance for researchers to challenge decisions. while this has not been the case to date in facilities where pbosdc is fully implemented, it is clear that some expenditure on training for staff and researchers is necessary; when one facility failed to train its new staff, there was an immediate impact on clearance rates and quality. pbosdc in practice in practice, none of the criticisms suggested in the previous section have proved significant. the evidence for this comes from the eight years of pbosdc at the virtual microdata laboratory (vml) at the uk office for national statistics, where the process was developed, as well as more recent experience at sds and hmrc data lab, also in the uk6 . uncertainty does exist. however, there should be no uncertainty about the process or the criteria for deciding whether to release an output or not. the purpose of researcher training is to help researchers understand the uncertainty and manage it. researchers have made mistakes, and while these are mostly oversights on the part of the researchers, a small number are due to researchers not understanding the principles of sdc. in general these were dealt with by discussion with the researchers, explaining the error. recidivism rates were negligible, but a very small number of vml researchers were asked to re-attend training (less than five people in seven years, out of over eight hundred trained researchers). although there are no figures, the number of requests for output refused at the vml was believed to be around 5%. inconsistency exists but there is little evidence to date of it being significant. over seven years around twenty output checkers were employed at ons; periodic checks were carried out, and although output checkers showed some slight variance, there was no practical difference in outcomes. it was discovered that one output checker was seen as ‘softer’ by some researchers, but even those outputs were well within the safety margin. in one case, a researcher’s output was seen by five output checkers (including the ‘soft’ checker), all of whom independently gave the same opinion. so whilst differences do exist, in practice these have no notable impact. if the research facility is not centrally run, or if several facilities are trying to co-ordinate sdc policies, there is another level at which variance may happen: between centres. there may be some cultural variability and the potential for local ethos to develop should be acknowledged and monitored. a common training framework where the principles of sdc are gone into in some depth, practice sharing between centres, “test” submissions and cross-centre case study reviews can help to mitigate this. the reason inconsistency and uncertainty have not been major problems to date is because of the built-in safety margins. as noted above, ignoring the efficiency problem means that the rules of thumb to address the confidentiality problem can be made much stricter; a confidentiality breach therefore requires a considerable error by both researchers and output checkers. 10 iassist quarterly 2015 iassist quarterly this margin of error is also at the heart of the efficiency of the process. output checkers can clear large amounts of output because they have confidence in the extra-safe rules-of-thumb for “unsafe statistics”, and because “safe statistics” require little scrutiny. at its peak the vml was dealing with 2500 clearance requests a year, or roughly ten every working day; one person was allocated to be output checker for that day, and the target was that this should take no more than half an hour, a target generally achieved. as a ‘clearance request’ typically consisted of several regressions, a couple of descriptive tables (and in one case over fifty graphs), and/or a log file, the actual number of statistical outputs being checked was much more than ten per day. this work rate was maintained by emphasising the researcher’s role in clearance. for example, the vml output checkers had an informal policy of clearing the easiest outputs first; this was communicated to researchers at training sessions. this gave researchers an incentive to build up a reputation of producing ‘good’ (i.e. easy to check) outputs. finally, there is evidence that the training sessions (and the relationship built up between researchers and output checkers) do foster a culture of confidentiality awareness, with examples of researchers being self-policing. it is unlikely that academics using the vml, sds and hmrc data lab would see themselves as being particularly sdc aware, as this is their only experience of it7. nevertheless, a comparison with users of facilities in other countries shows that the uk academia is much better informed. this is why the uk model of researcher training was adopted largely unchanged by eurostat as recommended best practice rules-based principles-based complex no generally not – rules-of-thumb usually sufficient flexible no yes transparent yes yes, if output checkers record factors leading to exceptional judgment consistent yes generally, but scope for minor variations secure limited by efficiency yes able to handle lower and higher risk cases efficiency limited by security yes tailored to circumstances risk management yes/no model explicit ’balance-of-risks’ model sensitive to context no yes sensitive to user needs no yes suitable for automation yes only as initial gatekeeper – humans are final arbiters requires researcher training no yes requires researcher engagement no yes requires co-ordination of output checkers no yes cost low high initial cost, low or negative ongoing costs 1 (brandt et al., 2010), and why researcher training should be seen as an investment rather than an expense. rules-based versus principles-based osdc the various aspects of the two approaches can be summarised as follows. although described above as alternatives, there is a relationship between rules-based and principles-based output sdc. the rules act as the starting point for pbosdc output checking (see brandt et al., 2010). however, there are two crucial differences: • in pbosdc, the rules are ‘rules-of-thumb’ – explicitly ad hoc, and amenable to adjustment, up or down, depending on circumstances. • these rules of thumb can then be more restrictive as prima facie efficiency has a low priority in the setting of the default values. for ‘safe statistics’ the ‘hard rules’ and ‘rules-of-thumb’ are the same; by definition, ‘safe statistics’ are those which are amenable to the identification of simple yes/no cases (ritchie, 2008). ultimately, pbosdc takes rules-based sdc as a good first-order approximation, but gives expertise and experience the final decision. for this reason it is the recommended approach for the rdcs. in situations where facility managers have less opportunity to manage researchers’ activities and outputs, pbosdc may be harder to implement, as the active engagement of researchers is essential. implementing pbosdc conceptual development the most detailed expression of pbosdc is the eurostat-approved guidelines of brandt et al. (2010). these were derived almost entirely from the vml rules, with the exception of the dominance rules. since the publication of the eurostat guidelines there have been a number of minor developments. for example, on regression, several authors (e.g. reznek and riggs, 2005; ronning, 2011; bleninger et al., 2011) have investigated options for deliberately creating misleading regression results, and us and australian works have studied regression models in remoteexecution systems. ritchie (2012, 2014b) incorporates most of these results, but new queries appear (for example, on the tabulation of binary variables, and how single-observation categories are treated in regressions) in addition, not all uk work (for example, on variance-covariance matrices) was adopted as it was felt to be too obscure for a general document. none of these are major challenges but they highlight the need for updating the current state of knowledge and communicating this to facility managers and researchers. when the vml guide to sdc was the only source and was owned by the vml team, updating was straightforward. one of the negative consequences of the wider use of the pbosdc approach is that there is no clear mechanism for maintaining and disseminating knowledge. iassist quarterly 2015 11 iassist quarterly facilities wishing to implement pbosdc may therefore need to collaborate with other facilities to ensure that a consistent approach can be developed and applied8. it has been suggested that the theoretical basis for pbosdc and the safe/unsafe statistics model does not provide sufficient reassurance to potential data suppliers. in addition, theoretical models may miss some plausible outcomes (that is, those likely to occur in practice) arising from, for example, naïve researchers. one solution is to set up an ‘ethical hacking’ mechanism to probe both genuine and fake outputs to test what information could be acquired. a second would be to have periodic audits of outputs from the facility by a competent body, perhaps another facility. this has the added advantage of encouraging consistency across facilities. need for training researcher training is essential for pbosdc. untrained researchers cannot be expected to know the basis of sdc, and the training itself should be used to develop an affinity between the researchers and the facility. there are several pbosdc training manuals/presentations which can be drawn upon and updated. this training could be part of any user accreditation process. there is already significant training experience in the uk and other countries, and therefore any facility would not need to start from scratch. however, as one of the key elements is to encourage researcher engagement, training needs to be sensitive to the needs and indeed interests of the researcher community; what may appeal to mexican researchers may be different from the expectation of norwegians. training also needs to be sensitive to the data. much of the extant training focuses on social and business survey data, which are of more limited risk, but administrative data brings additional problems (data may be a census; health data is typically more prone to outliers and retains its sensitivity over time). this is particularly important if tables are likely to comprise a lot of the outputs. finally, there is the need to train and update the knowledge of output checkers. periodic peer review has proved useful in the past. discussion forums allow knowledge to be shared, discussed and updated across facilities. if made available to researchers, these would also have value for researchers, although not necessarily in the same detail and with a full range of views expressed. one option would be to set up a discussion forum for facility staff, with a summary/faq page for researchers. trusted researchers one decision to be made is whether some researchers could have different rules applied. either particular types of people (e.g. full professors) or those who have built up reputations with the output checkers could have fewer checks imposed on them. the argument is that the burden for output checking is unnecessary, and creates ill-will amongst senior researchers who have proven their expertise. some facilities do use such differentiated models. in practice, these arguments do not hold. seniority is no guarantee of good practice. indeed the opposite can be true; experience shows that junior staff are more enthusiastic adopters of safe practices. similarly, familiarity may make mistakes less likely but does not eliminate them. as mistakes are far and away the most likely reason for failed clearances, long-term usage does not seem sufficient cause on its own to remove checks. these arguments are also based on the assumption that all outputs are equal. as should be clear from the discussion above, over time researchers should develop a sense of what is and is not allowed, and tailor their outputs accordingly. in other words, the pbosdc encourages the development of expertise in output assessment by all parties, and thus efficient exchanges. there is no need to create an artificial group of ‘good’ researchers; besides, creating ‘classes’ of users could discourage knowledge and experience sharing. malevolent researchers any output disclosure control system will have additional difficulties with a user who deliberately sets out to breach the system. pbosdc assumes that researchers are well-intentioned and interested in generating good statistical outputs. it may be possible to spot unauthorised outputs, but in practice a user set on breaching procedures would be able to disguise inappropriate outputs (for example by burying discovered data in complex model output). it should be noted that all the data used in current training programmes for controlled environments is wholly invented, yet plausible. the success of pbosdc therefore depends upon the ‘safe people’/’safe project’ dimensions of the five safes model. that said, at present, there are no known examples of academic researchers maliciously breaching confidentiality rules. there are numerous examples of well-intentioned researchers making mistakes, and a smaller number of cases of researchers deliberately ignoring procedures to make life easier for themselves (but again, without intending to disclose confidential data). setting up a system which would stand a chance of picking up malicious attacks therefore has no realistic prospect of success, and would increase greatly costs and clearance time. if a facility believed that malicious attack was a significant risk, then examining user logs and codes would be a more useful place to look for malpractice – and would also be necessary for forensic evidence in the event of a disciplinary event. in summary, pbosdc cannot stop ill-intentioned researchers; it is designed to deal with cases of accidental rule breaking and errors. if malicious attack is felt to be a problem, then this should be tackled at the ‘people’ level9. multi-stage clearance the uk vml operated a two stage-clearance procedure: ‘intermediate’ outputs could be released to researchers who could work on the further analysis at their home institution. ‘final clearance’ was given when papers were ready for general distribution. the uk secure data service only allowed ‘final clearance’: papers are fully prepared within the sds virtual facility. these two choices reflect physical differences. the vml involved travel to a specific location; it was not thought to be a good use of restricted facilities to have researchers editing papers, and the opportunity to discuss results with co-researchers was limited. in the sds these two issues are less important. note however that vml applied full pbosdc at the intermediate stage; the assumption was that once an output had left the vml’s control it could end up in the wild however well-intentioned the researcher. this assumption turned out to be correct, even if unsanctioned releases were rare. the vml’s ‘final clearance’ stage imposed a level of super-checking on outputs – asking researchers to limit outputs to the minimum necessary (compared to the 12 iassist quarterly 2015 iassist quarterly intermediate stage were multiple variants might be tried). as this reflected the research stages (exploratory work, produce many outputs, refine, and then publish) the model worked. the two-stage model did introduce an extra layer of administration, as now both intermediate and final clearances had to be checked and recorded. however, it did speed up the intermediate level clearance, as vml staff knew that if they made a mistake they had a ‘second chance’ to address it. this increased the already-wide margin for error in the vml pbosdc procedures. given that researchers were only bound by vml procedures (and not required by law) not to publish intermediate outputs, the 100% co-operation (allowing for mistakes) can be seen as a positive reflection of how researchers respond to appropriate training. however, the vml did make researchers aware that failure to follow procedures would affect the way that future access to data was viewed. this may have had more of an effect, and may be of relevance to a facility choosing to adopt two-stage clearance. allowing intermediate output out implies increasing risk unless adequate mitigation is in place. relying 100% on trust of researchers without any bounds implies allowing unmeasurable variations in risk and therefore is not sufficient to provide that mitigation. appropriate risk mitigation implies processes such as: • secondary licensing agreements. • minimum required security arrangements for intermediate output. • specified lists of persons who will have access to the intermediate outputs. • occasional random audits. summary this paper recommends pbosdc be adopted for use in rdcs, whether physical or remote. the reasons for this are twofold (i) principles based sdc can produce output that is of higher quality at the same or lower level of risk; and (ii) the opportunity of building a relationship with researchers can generate multiple benefits. in short, pbosdc is both safer and more efficient that rules-based approaches and it encourages the development of a culture of expertise in confidentiality. there are already guides and training programmes boasting several years’ experience of pbosdc. a facility wanting to adopt pbosdc can build upon these, perhaps tailoring them more to its particular researcher group and data. pbosdc needs to be integrated into a training programme; it assumes that researchers are well-intentioned (if liable to make occasional mistakes). there also needs to be a mechanism to ensure consistency across checkers (and possibility sites). the facility may also want to invest some resources in ethical hacking to provide extra reassurance to data owners. finally, the facility needs to determine whether it wants a oneor two-stage clearance process. while pbosdc at the point the output leaves the restricted facility can be the same, the perceived and actual security differs. for non-rdc environments, the case for pbosdc is less clear. for example, if researchers’ only sensible interaction with the facility manager is receiving a partially anonymised file on cd plus guidelines for publishing safe statistics, then a rules-based approach may be simpler. however we would recommend that even a discussion of rules should be placed in the context of the principles of sdc: in general, the support officer is more welcomed than the policeman. references bleninger p., drechsler j., and ronning g. (2011) remote data access and the risk of disclosure from linear regression”, stat. and op. res. trans. special issue: privacy in statistical databases, pp 7-24 http:// www.idescat.cat/sort/sortspecial2011/dataprivacy.1.bleninger-etal. pdf brandt m., franconi l., guerke c., hundepool a., lucarelli m. , mol j., ritchie f., seri g. and welpton r. (2010) guidelines for the checking of output based on microdata research, final report of essnet subgroup on output sdc, eurostat http://neon.vb.cbs.nl/casc/essnet/ guidelines_on_outputchecking.pdf camden m. (2014) “confidentiality for integrated data” in work session on statistical data confidentiality 2013; eurostat. http://www. unece.org/fileadmin/dam/stats/documents/ece/ces/ge.46/2013/ topic_3_nz.pdf desai t. and ritchie f. (2010) “effective researcher management”, in work session on statistical data confidentiality 2009; eurostat http:// www.unece.org/stats/documents/ece/ces/ge.46/2009/wp.15.e.pdf duncan, g., elliot, m. j. and salazar, j. j. (2011) statistical confidentiality. springer, new york elliot m.j. and purdam k. (2015) ‘the changing social data landscape’ in halfpenny, p. and procter, r. (eds.) innovation in digital research methods. sage eurostat (2014) treatment of statistical confidentiality, manual and exercises, luxembourg: eurostat hundepool, a., domingo-ferrer, j., franconi, l., giessing, s., nordholt, e. s., spicer, k., and de wolf, p. p. (2012) statistical disclosure control. chichester, uk: john wiley & sons. mackey e. and elliot, m.j. (2013) ‘understanding the data environment’ xrds 20(1); 37-39. http://xrds.acm.org/article.cfm?aid=2508973 ritchie f. (2007) statistical disclosure control in a research environment, mimeo, office for national statistics; available as wiserd data resources paper no. 6 http://www.wiserd.ac.uk/wp-content/ uploads/2011/12/wiserd_wdr_006.pdf ritchie f. (2008) “disclosure detection in research environments in practice”, in work session on statistical data confidentiality 2007; eurostat; pp399-406 http://epp.eurostat.ec.europa.eu/portal/page/ portal/conferences/documents/unece_es_work_session_statistical_ data_conf/topic%203-wp.37%20sp%20ritchie.pdf ritchie f. (2012) output-based disclosure control for regressions”. working papers in economics no. 1209. university of the west of england, bristol. http://www2.uwe.ac.uk/faculties/bbs/bus/ research/economics2012/1209.pdf ritchie f. (2014a) “resistance to change in government: risk, inertia, and incentives”, working papers in economics no. 1412. university of the west of england, bristol. http://www2.uwe.ac.uk/faculties/bbs/bus/ research/economics%20papers%202014/1412.pdf ritchie f. (2014b) “operationalising ‘safe statistics’: the case of linear regression”, working paper no. 1410, department of economics, university of the west of england, bristol. http://www2.uwe.ac.uk/ faculties/bbs/bus/research/economics%20papers%202014/1410. pdf ritchie f. and welpton r. (2012) “sharing risks, sharing benefits: data as a public good”, in work session on statistical data confidentiality 2011; eurostat http://www.unece.org/fileadmin/dam/stats/ documents/ece/ces/ge.46/2011/presentations/21_ritchie-welpton. pdf ritchie f. and welpton r. (2014) “understanding the human factor in data access”, working papers in economics no. 1413. university of iassist quarterly 2015 13 iassist quarterly the west of england, bristol. http://www2.uwe.ac.uk/faculties/bbs/ bus/research/economics%20papers%202014/1413.pdf ronning g. (2011) disclosure risk from interactions and saturated models in remote access, iaw discussion papers no. 72, june http:// www.iaw.edu/repec/iaw/pdf/iaw_dp_72.pdf sullivan f. (2011) the scottish health informatics programme, presentation to health statistics user group. http://www.rss.org.uk/ uploadedfiles/userfiles/files/frank-sullivan-linkage.ppt trewin d., andersen a., beridze t., biggeri l., fellegi i., toczynski t., (2007) managing statistical confidentiality and microdata access: principles and guidelines of good practice; geneva, unece /ces. welpton r. and ritchie f. (2011) “incentive compatibility in data security”, presentation to iassist 2012, vancouver http://www. iassistdata.org/downloads/2012/2012_a2_welpton_etal.pdf notes 1. corresponding author: felix ritchie, bristol business school, university of the west of england bristol, coldharbour lane, bristol bs16 1qy. email: felix.ritchie@uwe.ac.uk. 2. mark elliot, school of social sciences and data research institute, university of manchester, oxford road, manchester m13 9pl. email: mark.elliot@manchester.ac.uk 3. there is a large literature on statistical disclosure risk assessment and control methods. we do not discuss this in detail as here we are talking about two top-level process methodologies rather than the specifics of individual technical methods. we would direct the reader who is interested in the detail to recent comprehensive field reviews (duncan et al. 2011 and hundepool et al. 2012) 4. a more extended discussion of this argument is available in ritchie (2007). 5.. it is important to stress that “safe” is used here not in its absolute postpositive sense (free from danger or risk) but in its relative sense (the degree to which a solution affords security or protection from risk); see ritchie (2014b, section 5). 6. pbosdc is also used at other restricted facilities in mexico, germany and the netherlands, as well as informally in other countries. 7. an exception is medical sciences where data protection has had a much higher profile, and researchers in all countries tend to have a greater awareness of confidentiality issues. 8. in january 2015, all the uk rrdcs initiated a working group on sdc; one of the group’s functions is to determine how guidelines can be effectively maintained, distributed and updated. 9. note that rules-based osdc is even more susceptible to malicious attacks, as the yes/no approval process means that anything that looks acceptable will be approved without further checking. instructions for authors of the iassist quarterly 1/2 schwartz, ofira & hayslett, michele (2024) assessing needs and developing solutions, iassist quarterly 48(2), pp. 1-2. doi: https://doi.org/10.29173/iq1121 the creative commons-attribution-noncommercial license 4.0 international applies to all works published by iassist quarterly. authors will retain copyright of the work and full publishing rights. editors’ notes: assessing needs and developing solutions welcome to the second issue of iassist quarterly for 2024, iq 48(2). it was wonderful to meet so many old and new colleagues at the best iassist ever in halifax. it was really inspiring to learn about all the great work that is being done by members of this community. for those of you who presented, please consider turning your conference presentation or poster into a paper and submitting it to iq. this will allow you to share your expertise with a wider audience. if you were not able to attend the conference, you may have missed the announcement about the winner of the iassist conference paper competition. this year’s winner is the paper “how are we fair-ing? creating a fair self-assessment checklist for data repositories” by lauren phegley and lynda kellam. in the paper the authors describe a project in which a data repository’s staff wanted to gauge how well they were enabling fair principles. a small team from penn libraries found that much of the literature about fair was from the perspective of data creators, so they developed a fair principles self-assessment tool for repository teams. we look forward to publishing this paper in a future issue. we would like to take this opportunity to encourage you to look ahead to submitting your papers for next year’s paper competition. in addition to bragging rights, the award incudes free registration for the first author to the following year’s iassist conference. the four papers included in this issue of iq introduce tools developed in several institutions, representing a wide geographic diversity, to assess and resolve operational challenges. in the article titled ”research analysis: a world data system and canadian coretrustseal cohort needs assessment,” lee, gonzalez, payne and goins describe how they designed a method to identify the needs and challenges faced by members of the world data system (wds) and canadian coretrustseal pilot. they also describe the assessment tool they developed and the overarching challenges and goals identified through the usage of this tool. based on their findings, they provide recommendations on how best to assist the wds members and the cohort of canadian data repositories. constanzo and cooper, in their article ”developing institutional research data management strategies in canada: setting the foundation for stronger partnerships and collaborations,” describe national surveys developed by research intelligence expert group (rieg) to gauge institutions’ readiness for developing an institutional rdm strategy required by the government of canada’s tri-agency. the first survey was conducted in 2019 and a follow-up survey in 2022 in order to assess the progress of institutions in creating their institutional strategies and identifying additional challenges. the authors and report the findings and recommendations from their study and share their survey instruments. https://doi.org/10.29173/iq1121 https://creativecommons.org/licenses/by-nc/4.0/ 2/2 schwartz, ofira & hayslett, michele (2024) assessing needs and developing solutions, iassist quarterly 48(2), pp. 1-2. doi: https://doi.org/10.29173/iq1121 in ”enhancing fair compliance: a controlled vocabulary for mapping social sciences survey variables,” authors bach and klas introduce the gesis controlled vocabulary (cv) for variables in social sciences research data. this cv is designed to enhance semantic interoperability across various organizations and systems, and facilitates harmonization across different study waves. this endeavor aligns with the fair data principles, and aims to foster a more integrated and accessible research landscape. obasola and usman in their article ”digitising old yoruba newspapers at kenneth dike library,” describe in detail the digitisation of a collection of old yoruba newspapers stored at kenneth dike library in ibadan, nigeria. the project was undertaken in order to preserve this historical and delicate material, which includes rich details of local history. in addition to providing a detailed workflow, the authors share lessons learned. we hope you enjoy reading and wish you a productive summer. ofira schwartz and michele hayslett, june 2024 https://doi.org/10.29173/iq1121 vol264 8 iassist quarterly winter 2002 iassist quarterly winter 2002 9 by anne sofie fink* lately it has been considered to create a stronger web presence for the danish data archives (dda). in order to do so an analysis of the dda̓ s current web presence has been made and new strategy for web presence has been suggested. in this article i will try to outline these two steps. the new strategy is based on an ambition of creating a web presence that covers all services carried out by the data archive in a way that supports internal work processes and integrate these with related work processes performed by external parties. the article will describe the implications for web presence that follows and suggest that web content should be structured around three resource fields named data production, data archiving and data usage. the article is based on a presentation at the iassist conference 2002. about the dda and data archiving the dda was established in 1973 as a national data bank for quantitative research carried out primarily in the social sciences. as such the dda collects, preserves and disseminates machine readable research data. in 1993 the dda became an independent unit in the danish state archives. at present the archive has 15 full-time employees. traditionally data archives are associated with the social sciences. however, the dda has never been seen as a resource for the social sciences exclusively. work areas for the dda are defined by the potential for exploiting a core competence in preservation of empirical data of a certain structure e.g. a survey conducted as a study within medical science is received just as well as a social science survey is.1 as colleagues working in data archives know well, it is no easy job to retrieve data from the research community and great insistence needs to be exercised by the archivists. in an ideal world data archiving would be an integrated step within any research project. in reality researchers seldom regard data archiving as part of their research project and this often becomes an unexpected burden when the research project is finished. when data material arrives, data and documentation are converted to an archival format, which secures technical preservation for the future. according to priority the data materials are processed, which implies standardisation and various check of the material. during this process it is often necessary to request information from the researchers in order to make the documentation as complete as possible. dissemination of data material is carried out by providing search catalogues on the internet that let users search and select materials on their own. some material will be on-line accessible within nesstar2 whereas others will need to be ordered from the data archive. obligations as data archive as the national data archives for social, medical science and history, the dda has an obligation to act as an active partner in the danish research environment centred on empirical research. the archiveʼs contact and services to external parties is in this respect largely dependent on the web. among other things this means that www-based facilities for interacting with the research community continuously must be adapted and developed. although danish researchers and students will be perfectly able to use a web site in english, one aspect of this obligation is in my view to supply dda̓ s internet service in the national language, danish. to national producers and users of data, a web site in danish will call attention to the dda as an active player in the national research environment. our current web presence it is frequently pointed out that a lot of organisational web sites are internally focused in the way that the sites are much more about describing the organisation, than about offering information relevant to outside stakeholders. the dda web site is currently no exception. the web pages are used as if they constituted an information folder. they are static, there are no links to other sites and visitors are met with long ʻdead ̓texts to point out a few examples. besides the web site the dda supplies: a search catalogue a danish research portal on the internet 8 iassist quarterly winter 2002 iassist quarterly winter 2002 9 on the archiveʼs holdings, a search catalogue which is part of the integrated data catalogue (idc) and a nesstarbased catalogue. this article will not make an evaluation of these search facilities, but – for obvious reasons – they are exclusively devoted to search and location of data sets. figure 1 below shows the relation between the data archive and its stakeholders and web presence. figure 1 an interpretation of this figure shows that the web site and search catalogues mirror to a simplified production line where dissemination of data is the only activity given visibility on the web. some important insights can be gained from making an elaborate interpretation of the figure 1. � actors there are a variety of actors present. but these actors may in real life perform different roles – the producer may become a user of data, the producer may become disseminator of data, the user may become a producer, etc. � flow from the model linearity is presupposed, but more arrows could be made since both data sets and communication may be seen to move back and forth among actors and sometimes jump the actor first in line. � universalism the model which is used seems universalistic in scope, but any research environment will be unique in many ways e.g. due to the unique national context and it should be possible to mirror this on the web. � limited transparency web site and search catalogues cut into the ʻtravel ̓performed by a data material by focusing only on dissemination. however, all parts of the travel are relevant not only to the data archive but also to external parties. the web service should support all activities performed by the archive, not to create a duplicate existence to the dda on the web, but to create a bunch of complementary activities implying great synergy effects to be gained. another perspective on web presence the interpretation of figure 1 suggests viewing the data archive not just as collector and disseminator of data but also as an intermediary agent between data producers and data users (and disseminators of knowledge products). as intermediary the data archive would enable flows of communication and data sets among actors who will be taking on interchangeable roles of data producers, figure 2 10 iassist quarterly winter 2002 iassist quarterly winter 2002 11 data archive and data users. therefore a network model seems more fertile in showing relations between data users, data archive and data producers, see figure 2 below. as roles are interchangeable i suggest giving up viewing them as separate entities to be targeted and instead complementing the network on the web by three resource fields – a field for data production, a field for data archival issues and a field for data usage. in this network the data archive would act as the organiser of information content leading users into a virtual space structured around the three fields. this would create a broad coverage of subjects related to use, storage and production of data sets not as closed rooms but linked in a way that supports the visitorʼs work where – by way of examples – issues concerning secondary data will be intervened with data production and preservation. this would be to supply an integrated information gateway that supports non-linear work processes and communication flows among actors. by constructing the web site around three resource fields, the web service will incorporate great flexibility. the strategy would be to supply a bundle of content related to each field, which under goes continuous construction and adaptation according to visitors ̓ needs and demands. thereby the service will be tailored to the national research environment the data archive is part as an ever-changing reflection and support ʻorganismʼ. with this broader perspective on services provided on the web, the data archive should not only be seen as a service towards data users but just as well to data producers and data archivists since the service will be addressed both externally and internally. external services to data producers will have the aim of opening up the black box the archive has been so far. for instance when a data material is handed over to the archive, the data producer is no longer part of or aware of the work processes taking place within the archive. what will be done in this respect is to turn the archive inside out. closely linked to this ʻturning inside out ̓is that the services should include services supplied to the data archivists themselves e.g. search facilities for administrative purposes, standards for documentation and networks for corporation to name some. the data archive as organiser will not define what is relevant for users of the service/visitors of the portal instead this will be put up to them to suggest and add content. what the archive will provide is management of structuring the content and where relevant supply of information content. in this way the same navigation rules will be place on internal and external users of the fields. a field for data usage search of data sets search of complementary kinds of information: personal contacts, research products methods and techniques: educational material, discussions, resources, … issues concerning secondary data analysis: information about legislation, discussions, … networking: support for personal contacts, mailing lists… etc. figure 3 a field for data production best practices documents and discussions interchange of experiences with data collection, documentation, analysis, storage, practise vs. theory etc. issues concerning secondary data analysis information about legislation, discussions, … publications based on empirical data information about software packages etc. a field for data archival information on publication of data sets and metadata standards for data processing information about software products data resources (archives) and their services ad hoc activities e.g. eu-funded projects discussions of standards, archival formats, new initiatives etc. issues concerning secondary data analysis information about legislation, discussions, … etc. 10 iassist quarterly winter 2002 iassist quarterly winter 2002 11 in order to conceptualise this kind of expanded web service supplied and managed by the data archive the word web portal seems more appropriate. through this research portal data producers will be able to check issues to be aware of when carrying out a survey, data archivists will be reminded about procedures for data processing, users will find data sets across europe to throw light on a problem, etc. content of the resource fields in figure 3 a selection of ideas for content for the three different resource fields is listed. as visualised by the figure the three fields share some aspects whereas others relate predominantly to one of them. to clarify the argument for the resource field view i will comment on a few items related to each of the fields. search for data sets and more to a resource field for data usage obviously there must be links to facilities for searching for data. an excellent example for this is the nesstar system, which facilitates search, data and metadata browsing, on-line access to data analysis, graphic presentations and downloads. ideally there should be no limits to the search engines offered. the challenge is to provide users with an efficient gateway to this a vast amount of engines. one feature would be to make users ̓evaluation of different engines available to all of the users. users searching for data will seldom have an interest limited to the data and metadata provided. they will often have an interest in complementary kinds of information such as articles based on certain data sets, experience in analysing the specific data set, etc. (as is acknowledged by the madiera project3), too. but moving in the opposite direction ʻback ̓to the data archive and producers is also relevant (see figure 1). for instance the user might need to know about the principles for data processing used on a material or might want to know if it is possible to retrieve additional documentation by approaching the data producer personally. in this way the portal will acknowledge that access to data is only a part of what the data users need and the portal will support this. in this respect a resource field for data usage will in some respects be intervened with the two other fields, although the data user will experience this as integrated discoveries within a search process. data processing on the surface a field for data processing seems to be relevant only to data archivists – a view which some times is carried forward by data archivists – but it is actually an issue that concerns all users. to give an example, users will often need to become familiar with standards for data documentation, since they will influence the quality and/or sophistication of the data analysis the data material will support. to carry the example further, in order to accommodate this, standards for documentation have to be acceptable to the data producers and the data archive alike. the implication of this is that the archive needs to gain support and acceptance for its documentation accepted by the data producers in order to make them willing to supply the necessary information (and ideally without the data archivist having to make explicit demands for it). the field for data archiving could provide standards for data documentation applicable to data production and data archiving simultaneously. although users seem to have a great hunger for documentation, inflexible demands for elaborate documentation put on data producers might lead them to refuse to offer their data to the archive. transparency and communication mediated by the research portal are expected to lead to greater understanding among actors. data production as it is stressed several times now by way of example the expectation is that a field for data production is not a source only relevant to data producers, primarily researchers, but also to other visitors of the portal. for instance legal issues concerning conduction of a survey is something the researcher must be informed about. but also to the data archivists certain legal issues are important to be aware of when accepting to store, process and disseminate survey data. to users of the data it will also be important to be aware of the legal restriction a survey is conducted under in order to integrate this background information into the analysis process and conclusions. personal contacts an often-celebrated feature of the internet is the possibility of connecting people across space and time; the portal must obviously support this. discussion forums and mailing lists are aspects to take advantage of in this respect. as an archive concerned with adapting to visitors ̓needs and demands, virtual forums may be used to pose and discuss questions e.g. discussions of standards, archival formats, initiatives, etc. the support for personal contacts and discussions obviously reaches across the different fields and will act as a tool for defining and providing information content for the portal. by elaborating these few examples the argument for a portal structured around resource fields has been made. hopefully, the goal of supporting work processes and communication that does not fit in with a traditional strict division of usage, storage and production has been detected. to conclude; whether one is conducting surveys, collecting data, storing data, performing secondary analysis there are interrelated issues to be concerned with. a portal structured around the three resource fields will be able to represent an issue as a whole but offer multiple entries to the visitors based on his/her immediate preferences. 12 iassist quarterly winter 2002 iassist quarterly winter 2002 13 managing a portal in order to supply a portal in line with needs and expectations of the three target communities – users, archivists and producers – it is essential not only to collect feedback, but also to collect it in a way so that is readily available for evaluation and implementation. the feedback then needs both to be structured and to be discussed openly. this means that feedback will be collected and put up on the web site to prompt discussion among users from all communities. for a small data archive the workload involved in constructing and running a research portal might seem overwhelming. in order to meet with the challenge put on the archive, two fundamental principles should guide the work. first of all, external parties must be involved in the work and teamwork between external and internal parties will be essential. external parties ̓involvement will admit access to knowledge, insights and artefacts, which is not held by the data archive. the second principle is implementation of a modular structure of content. it should be possible to add, change, remove, and adapt content step by step and thereby incrementally construct – and de-construct – the portal content. in that way the portal should never be seen as complete or finished. what is essential however, is that strict structural guide lines are implemented in order to make sure that content are accessible in a well-organised and systematic way. this is also to make sure that the portal stay ʻmanageable ̓to the data archive as the key actor, otherwise the objective of supplying content with a high degree of usability will slip out of hand. “accelerating access” 2002ʼs theme for the iasssist conference was “accelerating access” and this article was in line with this theme. however, my goal has in some ways been broader than that. with the resource field approach the data archive is expected to open up to the outside by supplying a wellstructured information gateway with access to a broad range of topics concerning production, storage and use of data. on a micro level the portal will offer a multipurpose tool for supporting people working with empirical data. on a macro level the portal is supposed to raise awareness and use of secondary data as well as data archiving in the social sciences and other scientific disciplines – potentially also to other areas such as consultancy work and public accounts. last but not least it is our hope that the portal will give greater visibility to the dda as a competent and experienced actor on the national research scene and thereby support the activities carried out by the archive. footnotes 1 from 1996-2002 the dda received a separate funding for registration and archiving of data from the medical sciences. this initiative was named eras http:// www.sa.dk/dda/eras/english/eras/default.htm. 2 nesstar – networked social science tools and resources – www.nesstar.org – is a web-based system for searching, browsing, analysing, presenting and downloading statistical research data. the dda is one among several data archives distributing data through nesstar. 3 the madiera – multilingual access to data infrastructures of the european research area – project picks up where other nesstar-centred projects have stopped. the project objective is to create a multilingual search engine that gives users access to a broad range of data materials and related materials. * paper presented at the iassist conference, may 2002, in storrs, ct, usa. anne sofie fink, researcher, danish data archives, asf@dda.dk. http://www.sa.dk/dda/eras/english/eras/default.htm http://www.sa.dk/dda/eras/english/eras/default.htm http://www.nesstar.org mailto:asf@dda.dk 1/11 steeves, vicky; rampin, rémi; chirigati, fernando (2020) reproducibility, preservation, and access to research with reprozip and reproserver, iassist quarterly 44(1-2), pp. 1-11. doi: https://doi.org/10.29173/iq969 reproducibility, preservation, and access to research with reprozip and reproserver vicky steeves1, rémi rampin2, fernando chirigati3 abstract the adoption of reproducibility remains low, despite incentives becoming increasingly common in different domains, conferences, and journals. the truth is, reproducibility is technically difficult to achieve due to the complexities of computational environments. to address these technical challenges, we created reprozip, an open-source tool that automatically packs research along with all the necessary information to reproduce it, including data files, software, os version, and environment variables. everything is then bundled into an rpz file, which users can use to reproduce the work with reprozip and a suitable unpacker (e.g.: using vagrant or docker). the rpz file is general and contains rich metadata: more unpackers can be added as needed, better guaranteeing long-term preservation. however, installing the unpackers can still be burdensome for secondary users of reprozip bundles. in this paper, we will discuss how reprozip and our new tool, reproserver, can be used together to facilitate access to well-preserved, reproducible work. reproserver is a web application that allows users to upload or provide a link to a reprozip bundle, and then interact with/reproduce the contents from the comfort of their browser. users are then provided a persistent link to the unpacked work on reproserver which they can share with reviewers or colleagues. keywords research reproducibility, data management, data librarianship, digital preservation, access, reuse 1. introduction libraries, museums, and archives are tasked with saving the world's knowledge and culture, and keeping it re-usable throughout time. with the advent of born-digital works, attaining this goal becomes more complicated. we see digital stewardship as the long-term active management of digital materials and their supporting information including, but not limited to, metadata, documentation, and provenance. the goal is preservation and unencumbered access in the longterm. the proliferation of bespoke, unique processes for conducting research and the technical requirements to access the outcomes are ever-evolving, often taking place in proprietary environments that have been difficult for digital stewards to preserve (chassanoff and altman, 2019). digital preservation practitioners are constantly looking for improved methods to reliably preserve the works of their communities, as well as more efficient and scalable means of providing access to materials. in the case of those working with research communities, this often entails verifying that research data can be opened in the future by forward-migrating formats, and ensuring that there is enough documentation about these datasets to be useful to other researchers (johnston, 2014). this has since evolved to include preserving software alongside data, especially the project-specific software often developed by researchers for their work (rios et al., 2017). similarly, capturing and https://doi.org/10.29173/iq969 2/11 steeves, vicky; rampin, rémi; chirigati, fernando (2020) reproducibility, preservation, and access to research with reprozip and reproserver, iassist quarterly 44(1-2), pp. 1-11. doi: https://doi.org/10.29173/iq969 describing these software dependencies and scripting workflows have been a growing necessity for researchers seeking to make their work reproducible. reproducible research implies that when another researcher tries to rerun the same analysis—using the same data and code as the original creator— they will achieve the same result (goodman et al., 2016). this is extremely difficult, however. dependencies of a research process are not always known in full detail by the original researcher, much less the digital steward tasked with the long-term safety of the work (marwick, 2015). archivists and librarians who endeavor to accession and preserve reproducible research materials therefore must have a corresponding toolkit up to the task. but why is it so important for reproducibility and archival access to access materials in the original computational environment? given how much modern research practices rely on unique toolkits, the output from any analysis is highly dependent on the actual software in which the research happens (pawlik et al., 2019). results of research change depending on versions of software, on operating system changes, and other custom configurations (gronenschild et al., 2012). ensuring that the research process includes documentation about dependencies is important for the integrity of the scholarly record. likewise in archival science, the idea of authenticity of access to materials remains an important facet of both analog and digital work. authenticity refers to the idea that materials should be viewed/rendered exactly as they would have at the time of their creation, and has led to large-scale efforts in emulation services within digital archives (espenschied and lialina, n.d.). just as web archivists often need an emulated browser to view warc files (i.e. webrecorder player4), research archivists often need emulated computational environment to engage with original research materials and to confirm they are reproducible (rhizome, 2015). we envision an ecosystem of open infrastructure that makes reproducibility easy for both the originating researcher and those seeking to view, reproduce, and extend the original work. ideally, a researcher should have access toa self-contained, distributable, and reproducible bundle of their work using as few commands as possible. other researchers/reviewers should be able to reproduce and reuse experiments with a few mouse clicks, and without having to deal with all the complex chains of dependencies required to run the corresponding computational processes. to that end, we have developed two open-source tools: reprozip5 and reproserver6. reprozip is a tool that can automatically and transparently capture a computational experiment, creating a single, distributable bundle, without the researcher having to specify all of the required dependencies. other researchers can then use reprozip to automatically unpack this bundle and set up the environment in order to reproduce the research, even if the operating systems are different (chirigati et al., 2016). our second tool, reproserver, is a web application that provides easier access to reprozip bundles: instead of having to install different software together with reprozip to set up and reproduce a bundle, reproserver allows researchers to reproduce and reuse other people’s research from the comfort of their browser (rampin et al., 2018a). there is existing work in providing authentic access to materials at scale, though with varied results for reproducibility. binder7 is one such system. binder can provide access to r scripts and jupyter notebooks directly from a url, doi, or gitlab or github repository, allowing users to interact with these notebooks in a browser-based live environment. however, this is a solution specific to jupyter https://doi.org/10.29173/iq969 https://webrecorder.io/ https://webrecorder.io/ https://webrecorder.io/ https://webrecorder.io/ https://www.reprozip.org/ https://www.reprozip.org/ https://www.reprozip.org/ https://www.reprozip.org/ https://www.reprozip.org/ https://server.reprozip.org/ https://server.reprozip.org/ https://server.reprozip.org/ https://server.reprozip.org/ https://server.reprozip.org/ https://server.reprozip.org/ https://server.reprozip.org/ https://mybinder.org/ https://mybinder.org/ https://mybinder.org/ https://mybinder.org/ 3/11 steeves, vicky; rampin, rémi; chirigati, fernando (2020) reproducibility, preservation, and access to research with reprozip and reproserver, iassist quarterly 44(1-2), pp. 1-11. doi: https://doi.org/10.29173/iq969 notebooks and r scripts, and users must provide a file that describes the dependencies for the notebooks, i.e., users must manually capture such dependencies, which can be a burdensome task given there are chains of dependencies that can be unknowable to the user without proper documentation (jupyter et al., 2018). emulation as a service infrastructure8 (eaasi) is another system that provides access to operating systems in-browser and at scale. eaasi excels at emulation and provides the user access to older systems. that said, users’ emulated environment is a sandbox in which to do work, and while it can certainly help in efforts to recreate research projects from the past, users still must manually declare their environment (cochrane, 2018). this presents similar problems as with binder. we posit that core digital stewardship goals intersect with computational reproducibility, and that applying tools for computational reproducibility in an archival context can help ensure that materials preserved as key parts of the scholarly record can be verified, reused, reproduced, and accessed with integrity in the future. in this paper, we will describe how reprozip and reproserver can be used in tandem to: 1) capture a research process/workflow with all its dependencies and components in a preservation-ready format, the reprozip bundle, and 2) give patrons easy access to view and execute the contents of those bundles in the original computational environments, via their browser, with reproserver. 2. preservation with reprozip 2.1 reprozip infrastructure reprozip is an open-source tool that automatically captures all the dependencies of a research product or application originally run in a linux environment, and creates a single, distributable bundle that can be used to reproduce the entire experiment in another environment (e.g., on linux, windows, or macos) (chirigati et al., 2016). the tool works in two steps: 1. packing. given a research application that runs on a linux os, reprozip automatically and transparently traces all of the system calls related to the execution of that application. it captures all dependencies at the os level, including software, data files, databases, libraries, environment variables, parameters, and os and hardware information. using this information (which can optionally be customized by the user), reprozip creates a compendium for it: an rpz file containing the dependencies. the packing step must be done via the command line; future work includes a gui and support for non-linux environments. unpacking. given an rpz file, other users can use reprozip to set up the application in their environment, even if their os is different than the one used for the creation of the experiment. users can choose among different unpackers (e.g.: vagrant or docker for an isolated reproduction). reprozip then automatically sets up the environment and allows users to rerun the application. the unpacking step can be done either via the command line or using the provided gui. for instance, suppose that a user, alice, has a python script named prediction.py that represents her research. the script predicts the values of handwritten digits from an input image (using a variety of software packages) and outputs a new image with predictions. to pack this script with reprozip, she https://doi.org/10.29173/iq969 https://mybinder.org/ https://mybinder.org/ https://mybinder.org/ https://mybinder.org/ https://www.softwarepreservationnetwork.org/eaasi/ https://www.softwarepreservationnetwork.org/eaasi/ https://www.softwarepreservationnetwork.org/eaasi/ https://www.softwarepreservationnetwork.org/eaasi/ https://www.softwarepreservationnetwork.org/eaasi/ https://www.softwarepreservationnetwork.org/eaasi/ https://www.softwarepreservationnetwork.org/eaasi/ 4/11 steeves, vicky; rampin, rémi; chirigati, fernando (2020) reproducibility, preservation, and access to research with reprozip and reproserver, iassist quarterly 44(1-2), pp. 1-11. doi: https://doi.org/10.29173/iq969 first prepends reprozip trace to the execution of her script (via the command line): reprozip trace python prediction.py. reprozip then detects all the necessary dependencies in order to rerun this script in the future. she then can use reprozip pack to create an rpz bundle for it. if she shares this bundle with another researcher, bob, he can use the reprounzip component to set up and reproduce the script. because he works on macos and the script was originally created on a linux os, bob chooses the unpacker reprounzip-docker to let reprozip automatically set up the experiment using docker. reprozip takes care of all the unpacking details, and bob is able to reproduce the experiment even though he is not familiar with docker. provided that reproziphas access to the original command-line execution, it can trace and reproduce a variety of research products, including python/r/julia/any language scripts, applications based on compiled source code (e.g.: jar files, c/c++ binaries), jupyter notebooks, interactive applications, gui applications, and more involved scenarios such as client-server applications (e.g.: databases). reprozip has supported research in a variety of disciplines, including neuroscience, physics, computer vision, computational humanities, data science, data visualization, and data journalism. some users have contributed explanations for their research with steps to create the reprozip bundle and to access the contents, collected on the reprozip examples website9. reprozip is currently being used by reproman10, a tool that helps create and manage computing environments in the neuroimaging domain, and by corr11, a web platform from the national institute of standards and technology (nist) that stores and manages data from computational and experimental materials science. reprozip has also been used to support reproducibility evaluation for research papers published in different venues, including: the information systems journal12; acm sigmod13; and conferences that follow the artifact evaluation process guidelines14. 2.2 archival benefits of reprozip among the many benefits of reprozip, digital preservation is the most relevant. using reprozip, digital archivists and librarians can meticulously trace computing environments in which research takes place, from data files and applications to complicated chains of software dependencies. this process allows patrons to reproduce environments over time (steeves et al., 2018). the rpz file generated by reprozip is a preservation-ready object for several reasons: • flexibility. the file is generalized, completely agnostic to the unpacking technology being used: it can be rerun using many different unpacker plugins through reprozip. in addition, thanks to reprozip’s plugin model, unpackers can be added and removed as systems become more or less popular or usable, allowing long-term preservation. for instance, if docker becomes obsolete, then a reprozip unpacker can be written for another containerization system and the reprozip packages remain usable. completeness. the rpz file contains all the necessary files to reproduce and reuse the packed research. in addition, it provides a configuration file that is machine-readable and lists all the information about the computational environment and dependencies necessary to recreate it if in the future there are no containers or virtual machines, the archivist/librarian can still use the robust technical and administrative metadata from the bundle. https://doi.org/10.29173/iq969 https://examples.reprozip.org/ https://examples.reprozip.org/ https://examples.reprozip.org/ https://github.com/repronim/reproman https://github.com/repronim/reproman https://github.com/repronim/reproman https://www.nist.gov/programs-projects/cloud-reproducible-records https://www.nist.gov/programs-projects/cloud-reproducible-records https://www.nist.gov/programs-projects/cloud-reproducible-records https://www.nist.gov/programs-projects/cloud-reproducible-records https://www.nist.gov/programs-projects/cloud-reproducible-records https://www.journals.elsevier.com/information-systems/ https://www.journals.elsevier.com/information-systems/ https://www.journals.elsevier.com/information-systems/ https://www.journals.elsevier.com/information-systems/ http://db-reproducibility.seas.harvard.edu/ https://www.artifact-eval.org/guidelines.html https://www.artifact-eval.org/guidelines.html 5/11 steeves, vicky; rampin, rémi; chirigati, fernando (2020) reproducibility, preservation, and access to research with reprozip and reproserver, iassist quarterly 44(1-2), pp. 1-11. doi: https://doi.org/10.29173/iq969 bundle size. because reprozip captures only the necessary files for reproduction (i.e., the minimal set of files), the rpz file is often very small compared to the entire computational environment in which the research runs. this makes it easier to store and share these bundles. as an example of how reprozip can be used to streamline digital preservation, the tool was recently used as the core component in a recent project, saving data journalism15. this institute of museum and library service16 (imls) funded endeavor is aimed at preserving interactive news applications, i.e., dynamic websites that allows readers to fully interact with the stories (e.g.: dollar for docs17). due to its technological complexity, dynamic web content cannot currently be fully archived or preserved by libraries, newsrooms, or cultural institutions. as part of the project, we built a prototype, called reprozip-web18, that leverages reprozip and webrecorder to correctly preserve this type of application (boss et al., 2019). preliminary tests with the prototype showcased its usefulness in preserving and replaying news applications; a demonstration video is available online19. while reprozip significantly reduces the barrier to reproducing and preserving different applications, to unpack and rerun rpz bundles, users must still download the reprounzip component and their unpacker of choice (e.g.: reprounzip-docker to use docker20, and reprounzip-vagrant to use vagrant21), together with the software to be used by the unpacker (e.g.: docker or vagrant and virtualbox22). even if this is only required once, having to download and set up these tools can prove to be a heavy burden (and intrusive). therefore, there is still a barrier to accessing rpz files, as one might not have the required software to run reprozip and rerun the research. to address this limitation, we started implementing a web application called reproserver, detailed next. 3. reproserver — where preservation meets access 3.1 reproserver infrastructure reproserver is a cloud-native application designed to run reprozip research bundles from the browser. it builds on one of reprozip’s unpacker plugins, reprounzip-docker, and provides a web interface through which users can interact with bundled research and software in their original environment (rampin et al., 2018b). the system is functional and we maintain a public deployment at https://server.reprozip.org/. when provided an rpz file, either via direct upload from the user’s computer or via a link to a supported data repository, reproserver builds a container of the preserved environment and allows the user to run the program with the option to use different parameters or input data. the program is then executed in a cluster, and the results are shown to the user with all the same interactions that they might have with the reprounzip component on their desktop, including the ability to download the output files and upload their own input files. https://doi.org/10.29173/iq969 https://savingjournalism.reprozip.org/ https://savingjournalism.reprozip.org/ https://savingjournalism.reprozip.org/ https://savingjournalism.reprozip.org/ https://www.imls.gov/ https://www.imls.gov/ https://www.imls.gov/ https://www.imls.gov/ https://www.imls.gov/ https://projects.propublica.org/docdollars/ https://projects.propublica.org/docdollars/ https://github.com/reprozip-news-apps/reprozip-web https://github.com/reprozip-news-apps/reprozip-web https://github.com/reprozip-news-apps/reprozip-web https://github.com/reprozip-news-apps/reprozip-web https://github.com/reprozip-news-apps/reprozip-web https://github.com/reprozip-news-apps/reprozip-web https://www.docker.com/ https://www.docker.com/ https://www.vagrantup.com/ https://www.vagrantup.com/ https://www.virtualbox.org/ https://www.virtualbox.org/ https://www.virtualbox.org/ https://www.virtualbox.org/ https://www.virtualbox.org/ https://server.reprozip.org/ 6/11 steeves, vicky; rampin, rémi; chirigati, fernando (2020) reproducibility, preservation, and access to research with reprozip and reproserver, iassist quarterly 44(1-2), pp. 1-11. doi: https://doi.org/10.29173/iq969 figure 1. overall architecture of reproserver. reproserver is built on docker and kubernetes, which are common, well-supported technologies for the deployment of this type of applications. kubernetes is an open-source container orchestration system, which can be deployed on many infrastructures, and it is supported by most cloud providers (e.g.: google cloud, amazon web services, microsoft azure, digital ocean). it features automatic scaling, allowing for efficient medium-scale deployments accessible to the public. reproserver was designed with collaboration and interoperability in mind. we rely on existing thirdparty data repositories and we currently support the open science framework23, zenodo24, and figshare25. reproserver can use a reference to the file in the repository (e.g. file url), which appears in the url of the environment in reproserver. each execution of the unpacked research also gets a url that can be sent to others to show the results and retrieve output files. in addition to batch experiments (i.e., non-interactive applications), reproserver also supports interactive web-based experiments (e.g.: websites and jupyter notebooks), providing the user with a persistent url where they can access the program (which can be shared with others) while it is running. this url is shareable with others so that they can easily access an unpacked work without having to rebuild it every time. figure 2. the workflow of using reproserver to unpack and interact with an rpz file hosted on the osf. https://doi.org/10.29173/iq969 https://osf.io/ https://osf.io/ https://osf.io/ https://osf.io/ https://figshare.com/ https://figshare.com/ https://figshare.com/ https://figshare.com/ https://figshare.com/ https://figshare.com/ https://figshare.com/ https://figshare.com/ https://figshare.com/ 7/11 steeves, vicky; rampin, rémi; chirigati, fernando (2020) reproducibility, preservation, and access to research with reprozip and reproserver, iassist quarterly 44(1-2), pp. 1-11. doi: https://doi.org/10.29173/iq969 for instance, suppose that sarah is a reviewer on alice’s paper, and wants to look at her entire research process to verify the claims within the paper. she can use reproserver to reproduce alice’s work, without having to install any additional software on her computer. she chooses to upload the rpz bundle to reproserver and immediately reruns the prediction script, which returns consistent results. given alice’s research is reproducible, she feels comfortable signing off on a positive review. sarah then decides to try the script with a few other images she often uses in her research because she finds the research that alice has done is novel. she goes back to reproserver, changes the input file of the execution, and reruns the script again for different image files. for one of sarah’s files, the prediction seems to be wrong. she then shares the link to the unpacked experiment with alice so she can see and interact with these results also through reproserver, and hopefully improve her model. sarah would have neither been able to do that without the ease of access and use that goes into reproserver, nor if alice had not had been a great steward of her research and packed it for reproducibility. 3.2 archival access with reproserver researchers and digital stewards alike can use reprozip to create a compendium of work that is easily shareable, citable (if deposited in an institutional or subject-specific repository), and that they and the community-at-large can use. accompanying that, reproserver can act as a virtual, interactive reading room for those interested in exploring preserved reproducible research, for purposes such as studying the history of science (comparing workflows across time in a given niche, for instance), verifying results of research and building on it with in-browser interactions, and surely more ways that we have not envisioned — patrons are infinitely creative when approaching collections in libraries/archives. for the librarian/archivist, the workflow for preserving these complex research processes and objects becomes easier to manage. they would straightforwardly trace and pack the given research with reprozip, ingest into a repository or secure storage layer, and provide a button on the record page that links to reproserver for easy access to patrons to allow them to fully reproduce the work. for archival researchers, interacting with primary source materials as authentically as possible is extremely important for understanding the complexities (kim and mennerich, 2015). if libraries/archives provide the primary source materials (the rpz bundle) in the collections via the web browser, archival researchers can view and interact with materials in their original environment, with detailed metadata and provenance information, which ensures high-fidelity, with respect to original functionality and utility. research bundles like reprozip’s rpz files can be hard to read at first blush -in fact, there have been guides written such as how to read a research compendium to educate even computationallyadvanced patrons in understanding the materials captured in these packages (nüst et al., 2018). by designing a simple interface with which to access and interact with these files, reproserver can be an onramp for patrons who might want to explore different methods of research, but lack some of the computational literacy to fully understand how to read the bundle or use it on their own computer. https://doi.org/10.29173/iq969 8/11 steeves, vicky; rampin, rémi; chirigati, fernando (2020) reproducibility, preservation, and access to research with reprozip and reproserver, iassist quarterly 44(1-2), pp. 1-11. doi: https://doi.org/10.29173/iq969 4. future work to build on our current work with reproserver, we would like to support interacting with more varied types of research processes from the browser. right now, we only allow uploading new inputs, but we would like to, for instance, give users the chance to interact with the packed terminal applications via a terminal right on the webpage, or conversely interact with gui applications from their browser. this would give greater freedom to those patrons who want more in-depth investigations of the unpacked works. for patrons that use big data, we are also looking into integrating with high-performance computing (hpc) clusters and cloud providers directly, to allow the reproduction of research processes that are distributed, use large data, or need to be run on specific hardware (for example, specific gpus for artificial intelligence research that might exist on one hpc cluster but not the current public reproserver deployment). this means that patrons with credentials to specialized computing environments would be able to bridge reproserver and their custom compute cluster, and use their own compute to reproduce work via reproserver. we also want to support a wider range of data repositories, especially those being supported and maintained by librarians and archivists at their institutions. reproserver can currently obtain files from zenodo, osf, and figshare, but it can be similarly integrated with more software by writing adapters for their apis. we want to support all the repositories where researchers are likely to deposit their software, including the self-hostable options that might be used by institutions, such as dataverse26, dspace27, and samvera28. finally, a long-term goal of reprozip is the support of more operating systems. we are investigating ways to capture and reproduce research on macos and microsoft windows since those operating systems are prevalent in many research domains. we are also looking into creating linux environments with the same dependencies that were used on the author’s machine (for example, python/r/ruby packages are usually available for all operating systems), to alleviate the need to run virtual machines with alternative oses (at substantial cost). 5. conclusion ultimately, in a time when the initiatives of libraries and archives in particular are being cannibalized by for-profit and corporate entities, it’s more important than ever to build out and support open infrastructure that can empower institutions to safely archive the work of their communities, and provide user-friendly means of accessing that work authentically. reprozip and reproserver were built with this in mind. there is no vendor lock-in, and no need to work within or port to commercial platforms for reproducibility. any institution could instantiate reproserver (with minimal computing infrastructure) and allow their users to reproduce the work held safely in their repositories. likewise, reproserver adds to the open-source ecosystem for authentically accessing research in the cloud. reprozip and reproserver increase open-source support for reproducibility in the long-term and ensure that researchers and the archivists/librarians they trust to steward their work will be able to safely and accurately maintain access to reproducible research materials. https://doi.org/10.29173/iq969 https://dataverse.org/ https://dataverse.org/ https://dataverse.org/ https://dataverse.org/ https://dataverse.org/ https://dataverse.org/ 9/11 steeves, vicky; rampin, rémi; chirigati, fernando (2020) reproducibility, preservation, and access to research with reprozip and reproserver, iassist quarterly 44(1-2), pp. 1-11. doi: https://doi.org/10.29173/iq969 ultimately, reprozip and reproserver improves research by lowering the barriers to reproducibility, and the way in which the lis community can preserve and make discoverable these research compendia. reproserver builds on this important work by making access to these materials seamless; through an easily accessible web interface that allows for multiple ways of providing an rpz file (including integrations with repositories), multiple ways of interacting with the unpacked work, and open-sourcing this for use by anyone/any institution. acknowledgements we would like to acknowledge dr. juliana freire, the principal investigator of the reprozip project, for her support in continuing to build reprozip and now reproserver. we would also like to thank genevieve milliken for providing feedback on this paper prior to submission. we would also like to acknowledge the support from the gordon and betty moore foundation as well as the alfred p. sloan foundation. the moore-sloan data science environment was vital to the development of reprozip. references boss, k.e., steeves, v., rampin, r., chirigati, f., hoffman, b., 2019. saving data journalism: using reprozip-web to capture dynamic websites for future reuse (preprint). lis scholarship archive. https://doi.org/10.31229/osf.io/khtdr chassanoff, a., altman, m., 2019. curation as “interoperability with the future”: preserving scholarly research software in academic libraries. j. assoc. inf. sci. technol. 0. https://doi.org/10.1002/asi.24244 chirigati, f., rampin, r., shasha, d., freire, j., 2016. reprozip: computational reproducibility with ease, in: proceedings of the 2016 international conference on management of data, sigmod ’16. acm, new york, ny, usa, pp. 2085–2088. https://doi.org/10.1145/2882903.2899401 cochrane, e., 2018. emulation as a service infrastructure (eaasi). https://doi.org/none espenschied, d., lialina, o., n.d. authenticity/access | one terabyte of kilobyte age. url https://blog.geocities.institute/archives/3214 (accessed 11.1.19). goodman, s.n., fanelli, d., ioannidis, j.p.a., 2016. what does research reproducibility mean? sci. transl. med. 8, 341ps12-341ps12. https://doi.org/10.1126/scitranslmed.aaf5027 gronenschild, e.h.b.m., habets, p., jacobs, h.i.l., mengelers, r., rozendaal, n., van os, j., marcelis, m., 2012. the effects of freesurfer version, workstation type, and macintosh operating system version on anatomical volume and cortical thickness measurements. plos one 7, e38234. https://doi.org/10.1371/journal.pone.0038234 johnston, l.r., 2014. a workflow model for curating research data in the university of minnesota libraries: report from the 2013 data curation pilot (report). university digital of minnesota conservancy. jupyter, p., bussonnier, m., forde, j., freeman, j., granger, b., head, t., holdgraf, c., kelley, k., nalvarte, g., osheroff, a., pacer, m., panda, y., perez, f., ragan-kelley, b., willing, c., 2018. binder 2.0 reproducible, interactive, sharable environments for science at scale. proc. 17th python sci. conf. 113–120. https://doi.org/10.25080/majora-4af1f417-011 https://doi.org/10.29173/iq969 https://doi.org/10.31229/osf.io/khtdr https://doi.org/10.1002/asi.24244 https://doi.org/10.1145/2882903.2899401 https://blog.geocities.institute/archives/3214 https://doi.org/10.1126/scitranslmed.aaf5027 https://doi.org/10.1371/journal.pone.0038234 https://doi.org/10.25080/majora-4af1f417-011 10/11 steeves, vicky; rampin, rémi; chirigati, fernando (2020) reproducibility, preservation, and access to research with reprozip and reproserver, iassist quarterly 44(1-2), pp. 1-11. doi: https://doi.org/10.29173/iq969 kim, j., mennerich, d., 2015. jeremy blake’s time-based paintings: a case study – electronic media review four. marwick, b., 2015. how computers broke science – and what we can do to fix it [www document]. the conversation. url http://theconversation.com/how-computers-broke-science-andwhat-we-can-do-to-fix-it-49938 (accessed 3.27.17). nüst, d., boettiger, c., marwick, b., 2018. how to read a research compendium. arxiv180609525 cs. pawlik, m., hütter, t., kocher, d., mann, w., augsten, n., 2019. a link is not enough – reproducibility of data. datenbank-spektrum. https://doi.org/10.1007/s13222-019-00317-8 rampin, r., chirigati, f., steeves, v., freire, j., 2018a. reproserver: making reproducibility easier and less intensive. arxiv180801406 cs. rampin, r., chirigati, f., steeves, v., freire, j., 2018b. reproserver: making reproducibility easier and less intensive. arxiv180801406 cs. rhizome, 2015. cyberspace, the old-fashioned way [www document]. rhizome. url http://rhizome.org/editorial/2015/nov/30/oldweb-today/ (accessed 11.1.19). rios, f., contaxis, n., almas, b., jabloner, p., kelly, h., 2017. exploring curation-ready software: use cases. url http://www.softwarepreservationnetwork.org/exploring-curation-readysoftware-use-cases/ (accessed 7.8.17). steeves, v., rampin, r., chirigati, f., 2018. using reprozip for reproducibility and library services. iassist q. 42, 14–14. https://doi.org/10.29173/iq18 end-notes 1 vicky steeves is the librarian for research data management and reproducibility, a dual appointment between new york university division of libraries and nyu center for data science. she can be reached by email <vicky.steeves@nyu.edu> and her orcid is 0000-0003-4298-168x. 2 rémi rampin is a research engineer at the tandon school of engineering at new york university. he can be reached by email <remi.rampin@nyu.edu> and his orcid is 0000-0002-0524-2282. 3 fernando chirigati is a postdoctoral researcher at the tandon school of engineering at new york university. he can be reached by email <fchirigati@nyu.edu> and his orcid is 0000-0002-95665835. 4 https://webrecorder.io/ 5 https://reprozip.org 6 https://server.reprozip.org 7 https://mybinder.org/ 8 https://www.softwarepreservationnetwork.org/eaasi/ 9 https://examples.reprozip.org/ 10 https://github.com/repronim/reproman 11 https://www.nist.gov/programs-projects/cloud-reproducible-records 12 https://www.journals.elsevier.com/information-systems/ 13 http://db-reproducibility.seas.harvard.edu/ 14 https://www.artifact-eval.org/guidelines.html 15 https://savingjournalism.reprozip.org/ 16 https://www.imls.gov/ 17 https://projects.propublica.org/docdollars/ https://doi.org/10.29173/iq969 http://theconversation.com/how-computers-broke-science-and-what-we-can-do-to-fix-it-49938 http://theconversation.com/how-computers-broke-science-and-what-we-can-do-to-fix-it-49938 https://doi.org/10.1007/s13222-019-00317-8 http://rhizome.org/editorial/2015/nov/30/oldweb-today/ http://www.softwarepreservationnetwork.org/exploring-curation-ready-software-use-cases/ http://www.softwarepreservationnetwork.org/exploring-curation-ready-software-use-cases/ https://doi.org/10.29173/iq18 mailto:vicky.steeves@nyu.edu https://orcid.org/0000-0003-4298-168x mailto:remi.rampin@nyu.edu https://orcid.org/0000-0002-0524-2282 mailto:fchirigati@nyu.edu https://orcid.org/0000-0002-9566-5835 https://orcid.org/0000-0002-9566-5835 https://webrecorder.io/ https://reprozip.org/ https://server.reprozip.org/ https://mybinder.org/ https://www.softwarepreservationnetwork.org/eaasi/ https://examples.reprozip.org/ https://github.com/repronim/reproman https://www.nist.gov/programs-projects/cloud-reproducible-records https://www.journals.elsevier.com/information-systems/ http://db-reproducibility.seas.harvard.edu/ https://www.artifact-eval.org/guidelines.html https://savingjournalism.reprozip.org/ https://www.imls.gov/ https://projects.propublica.org/docdollars/ 11/11 steeves, vicky; rampin, rémi; chirigati, fernando (2020) reproducibility, preservation, and access to research with reprozip and reproserver, iassist quarterly 44(1-2), pp. 1-11. doi: https://doi.org/10.29173/iq969 18 https://github.com/reprozip-news-apps/reprozip-web 19 https://www.youtube.com/watch?v=_sfvmzknfo&list=pljgz3v4gfxpwc7xc9ukiygy8a8rlgcvv5 20 https://www.docker.com/ 21 https://www.vagrantup.com/ 22 https://www.virtualbox.org/ 23 https://osf.io/ 24 https://zenodo.org/ 25 https://figshare.com/ 26 https://dataverse.org/ 27 https://www.dspace.com/en/inc/home.cfm 28 https://samvera.org/ https://doi.org/10.29173/iq969 https://github.com/reprozip-news-apps/reprozip-web https://www.youtube.com/watch?v=_sfvmz-knfo&list=pljgz3v4gfxpwc7xc9ukiygy8a8rlgcvv5 https://www.youtube.com/watch?v=_sfvmz-knfo&list=pljgz3v4gfxpwc7xc9ukiygy8a8rlgcvv5 https://www.docker.com/ https://www.vagrantup.com/ https://www.virtualbox.org/ https://osf.io/ https://zenodo.org/ https://figshare.com/ https://dataverse.org/ https://www.dspace.com/en/inc/home.cfm https://samvera.org/ 1/14 kong, ningning nicole (2018) one store has all? the backend story of managing geospatial information toward an easy discovery, iassist quarterly 42 (4), pp. 1-14. doi: https://doi.org/10.29173/iq927 one store has all? – the backend story of managing geospatial information toward an easy discovery ningning nicole kong1 abstract geospatial data includes many formats, varying from historical paper maps, to digital information collected by various sensors. many libraries have started the efforts to build a geospatial data portal to connect users with the various information. for example, geoblacklight and opengeoportal are two open-source projects that initiated from academic institutions which have been adopted by many universities and libraries for geospatial data discovery. while several recent studies have focused on the metadata, usability and data collection perspectives of geospatial data portals, not many have explored the backend stories about data management to support the data discovery platform. the objective of this paper is to provide a summary about geospatial data management strategies involved in the geospatial data portal development by reviewing related projects. these data management strategies include managing the historical paper maps, scanned maps, aerial photos, research generated geospatial information, and web map services. this paper focuses on the data organization, storage, cyberinfrastructure configuration, preservation and sharing perspectives of these efforts with the goal to provide a range of options or best management practices for information managers when curating geospatial data in their own institutions. keywords geospatial information, data management, geoportal introduction in united states, academic libraries began to provide geospatial datasets to their users from the 1990s when the census materials were given to depository libraries as tiger (topologically integrated geographic encoding and referencing) files (gabaldón and repplinger, 2006). the way how these datasets were managed and distributed by libraries depends on how they were acquired. in the early years, as these government datasets were distributed as part of the depository program, they arrived libraries as cd-roms (abbott and argentati, 1995). the datasets can be managed by libraries from their cataloging systems as same as any other electronic resources. as the arl geographic information systems (gis) literacy program introduced more librarians gis skills, library’s gis service expanded and reached more information users (association of research libraries, 1999). since then, academic libraries started to include more geospatial datasets into their collection development. the management and access to these geospatial data was one of the key challenges that librarians face. many academic libraries have purchased computer hardware and software to store and manage these datasets, and provided users data access via library visits (boissé and larsgaard, 1995). https://doi.org/10.29173/iq927 2/14 kong, ningning nicole (2018) one store has all? the backend story of managing geospatial information toward an easy discovery, iassist quarterly 42 (4), pp. 1-14. doi: https://doi.org/10.29173/iq927 during the past decade, as technology evolves and the library’s geospatial data collection expands, libraries began to explore new ways to provide access to geospatial information. several reasons have contributed to this information access change. first of all, library’s geospatial data collection has expanded to include a great variety of datasets, from the government datasets to proprietary datasets, including generic gis files, scanned or georeferenced maps, regional collected aerial photos, as well as research data collection (longstreth et al., 1995; bennett and nicholson, 2007; newton, miller and bracke, 2010). since these data are scattered in different server spaces, it is difficult for geospatial information users to find their interested datasets. there is a need to have an inclusive geospatial information catalog serving for an easy data discovery. secondly, the datasets previously distributed from the depository program became available online as public domain data. often these data were difficult to find and difficult to use, so libraries began to provide links and general guides of using these datasets instead of directly managing them (morris, 2006). finally, the development of gis industry has made it possible to distribute maps and geospatial datasets via web services instead of on-site visits. as a logical response, many academic libraries started to develop geodata portal to facilitate federated search across different data provider services. examples include the inside idaho project at university of idaho, the scholars geoportal project at ontario council of university libraries, and the open geoportal federation project initiated at tufts university (kenyon, godfrey and eckwright, 2012; florance et al., 2015; trimble et al., 2015). geoportals provide a gateway to discover and access geographic information. with an effective design, users can discover their interested geospatial datasets from multiple data providers by a map, as well as a keyword or faceted search. geoportals can improve data access and sharing between libraries, public domain platforms, as well as any other data providers, and greatly foster geospatial data discovery and usage. however, academic libraries are facing many challenges in creating the geoportal. these challenges include how to choose and create the cyberinfrastructure for the geoportal development; how to effectively organize the geospatial metadata and develop the data catalog to support the geoportal; how to manage the library’s datasets; and the portal usability design. many studies have explored different aspects of these challenges, such as the geospatial data catalog, metadata schema, portal federation, and usability (kollen et al., 2013; hardy and durante, 2014; florance et al., 2015; blake et al., 2017). among these challenges, my research interest is particularly focused on how to manage the ever increasing datasets in the library’s environment, so that the information can be well organized and preserved to serve as a reliable node for the geoportal. in this paper, i collected data management information from two different sources, including existing geoportal systems and opengeoportal participating institutions, in order to learn about the data management practices in different cases. with this information, i hope this paper will be a reference to help librarians to understand the current status, technology and major challenges about managing geospatial data. background geospatial data refers to the wide variety of datasets that have a geographic component, and that can typically be viewed as representing a portion of the earth’s surface in some way. different with other electronic resources in library’s collections, geospatial data are notoriously difficult to manage due to the following reasons. first, geospatial data doesn’t have a uniform data model. it includes both vector and raster datasets and might reside in complex, multi-file objects. many file formats https://doi.org/10.29173/iq927 3/14 kong, ningning nicole (2018) one store has all? the backend story of managing geospatial information toward an easy discovery, iassist quarterly 42 (4), pp. 1-14. doi: https://doi.org/10.29173/iq927 used to store this information are proprietary and are linked directly to the version of the program in which they were designed. especially in recent years, many of these data are being stored in relational geodatabases requiring sophisticated storage and archiving schemes (janée, 2009). second, geospatial data are often quite large with datasets commonly having gigabyte granularities and with some datasets growing by terabytes per day. the geospatial data collection efforts are ongoing and long-lived programs. due to the improved sensor technologies, geospatial data are increasing exponentially during the last decade (chuck killpack, 2011). the gis data libraries are challenged by the increased requirements to collect, curate, and make available more data and services than ever (goldberg et al., 2014). third, geospatial metadata may be voluminous because the associated technology and how it has changed over time need to be documented, which is difficult to find or not included with other information about the file itself (erwin and sweetkindsinger, 2009). geoportal provides an effective way to exchange and sharing the complicated geospatial datasets between people and organizations. academic libraries have started their efforts to adopt the geoportal technology into their own discovery tools, such as the development of the alexandria digital library project (smith and frew, 1995) and the g-portal project developed at singapore (lim et al., 2002). in recent years, these efforts were boosted by the open source development concept, which allows more institutions to participate and contribute into the development of a single geoportal, such as the open geoportal federation and the geoblacklight community (florance et al., 2015; battista et al., 2017). geoportal is developed and implemented using three distributed components: a discovery interface or portal connecting to the metadata catalog; a set of web services which comply with existing standards for describing, accessing and exchanging digital data; and a data management system which provides a managed environment for both raster and vector geographic contents as shown in figure 1 (maguire and longley, 2005; tait, 2005). the first component is related to geospatial metadata. there are existing metadata standards such as iso 19115 and fgdc (federal geographic data committee) content standard to describe geospatial information. although challenges exist when document metadata using these standards and transform metadata records between the standards and library’s generic metadata formats, many studies have addressed these issues and provided best management practices or crosswalk (nogueras-iso, zarazaga-soria and muromedrano, 2005; batcheller, 2008; hardy and durante, 2014). starting from 2014, a github repository, opengeometadata, emerged as an online space for libraries to share their geospatial metadata in an open, standards-agnostic way and has become an essential piece of infrastructure for building cross-institutional catalogs (battista et al., 2017; opengeometadata, 2018). it greatly fostered the collaboration between individual libraries. the majority of opengeometadata represents geospatial data objects held within the respective institution’s repository as well as records extracted from government open data portals. the second component is about standard web services. standards specify communication protocols between data servers that provide geospatial datasets. they are used to ensure interoperability of different datasets (strain, rajabifard and williamson, 2006). several technical standards defined by the open geospatial consortium (ogc) and the world wide web consortium (w3c) play an important role in the dissemination of geospatial data (bocher and martin, 2012). for example, ogc has specified geospatial data delivery standards including web mapping service (wms), web feature service (wfs) and its transactional equivalent (wfs-t), and the web coverage service (wcs) (ogc, 2013). specifications developed by the w3c for data dissemination include html, xml, svg, soap, wsdl, etc. libraries should consider to adopt these standards as much as possible when designing their geospatial data management systems so that data could be shared in an interoperable environment. https://doi.org/10.29173/iq927 4/14 kong, ningning nicole (2018) one store has all? the backend story of managing geospatial information toward an easy discovery, iassist quarterly 42 (4), pp. 1-14. doi: https://doi.org/10.29173/iq927 the third component of geoportal, which will be further discussed in this article, is a data management system. a geoportal is only as good as the information that is published through the site. thus, it is essential to keep the datasets reliable, appropriate, and professionally maintained (tait, 2005). in order to effectively share the datasets, we need to design the spatial data management infrastructure compatible with the existing standards. many researches in the geographic information science has explored the efficiency of geospatial web services, which will provide great insights for building the data management system (yang et al., 2006; li et al., 2015). academic libraries has also explored to manage geospatial data via their institutional data repositories (newton, miller and bracke, 2010; durante and hardy, 2015). we will explore the data management options in various projects and analyze the capabilities offered by each option. geoportals have been existing outside of library’s domain ever since 1990s (tait, 2005). government organizations or multi-institution projects have been using geoportal as a way to facilitate data access and sharing for various reasons. initiated from government organizations, the united states national spatial data infrastructure (nsdi) released the geospatial one-stop (gos) geoportal in 2003, which aims to promote coordination and alignment of geospatial data collection and maintenance among all levels of government. in europe, the inspire geoportal intends to be europe’s internet access point for geospatial data discovery, access and services (crompvoets and wachowicz, 2004). to facilitate multi-organizational project data sharing, different geoportals have been implemented for marine administration, forest management, countryside development, and biodiversity research (askew et al., 2005; strain, rajabifard and williamson, 2006; flemons et al., 2007; mcinerney et al., 2012). examine the data management solutions in these projects as well as in the current library’s practices would provide great insights for librarians to manage the expanding geospatial datasets. client internet portal interface web server + metadata catalog data exchange service map service or file transfer protocols data management system client client figure 1. components of a geoportal infrastructure. https://doi.org/10.29173/iq927 5/14 kong, ningning nicole (2018) one store has all? the backend story of managing geospatial information toward an easy discovery, iassist quarterly 42 (4), pp. 1-14. doi: https://doi.org/10.29173/iq927 methods in order to learn about the data management strategies in different geoportal projects, i have reviewed articles about geoportal project development. the data management portion of these projects has been analyzed and compared in order to generate useful references for librarians. since different articles introduced their projects from different aspects, the data management solutions were not introduced in consistent details and procedures. some of these papers focused on the planning stage of the project, while others focused on the geoportal outcome and assessment. i integrated the information extracted from planning concerns and assessment results, then summarized as the considerations that librarians should take in designing their geospatial data management system. i also listed possible options librarians could take, followed by specific examples in non-library and academic library settings. in addition, i reviewed the metadata records contributed from opengeoportal community, and analyzed the online data link properties managed by different contributors in order to learn about the data management practices in each participating institution. with this information, i hope this paper will help readers to understand the current status, technology, and major challenges about managing geospatial data in academic libraries. results designing the spatial data management strategies in the many existing geoportal development projects, two of them focused or introduced the planning stage of their projects. the needs assessment is essential as the first step to design a geospatial data management system. the purpose of such an assessment is to help understand the program goal, geospatial data properties, existing systems, expected gis functionalities, etc., so that the data management system could be designed around these information. smith et al. (2015) documented a needs assessment process that applied to one of the geospatial programs in national park service, which provided a comprehensive and logical workflow for the data system administrators and developers. in library setting, shell australia technical library has also conducted a very detailed initial assessment when designing the data management system, which includes the details about file size, data lifecycle, file formats, and management needs. there are several considerations in designing the data management solutions for geoportal. first, how data will be organized, maintained and updated in the system? the data management portion of many geoportal projects intends to create a centralized system to store, describe, and manage the datasets from multiple sources. planning for such a system requires to set up common practices for data formats, data organization rules, metadata schema and content standard. in addition, the different license agreements for various datasets need to be considered and access policy need to be designed in order to disseminate the data to appropriate user groups. some projects also require the geospatial datasets to be maintained or updated on time to ensure the users can access the latest datasets. in those cases, a centralized enterprise infrastructure with version control is needed. the second consideration is about the expected functionality for the data service, such as data downloading, visualization, and additional mapping functions. the most common function for many geoportal projects is downloading the datasets and simple metadata. some portals were designed https://doi.org/10.29173/iq927 6/14 kong, ningning nicole (2018) one store has all? the backend story of managing geospatial information toward an easy discovery, iassist quarterly 42 (4), pp. 1-14. doi: https://doi.org/10.29173/iq927 with a map preview function either as a thumbnail image or an interactive map so that users can have an idea about the datasets before downloading (schwarb et al., 2011; trimble et al., 2015). as the geographic analysis expands from its primary domain of geoscience to a broad range of disciplines, the new user groups need to access geospatial data collections in the way that reflects the current environment of web-based research and publishing activities (durante and hardy, 2015). so, web-based spatial visualization became a common requirement in the recent year’s geoportal development. the spanish national spatial data infrastructure has monitored the different types of visits on their geoportal. according to their reports between 2009 to 2012, the data visualization functionality was the most visits, which composed more than 80% of their yearly visits. data downloading service was the second most visit type. in addition to data visualization and downloading, some geoportals also offer gis functions such as location or attribute based query within datasets, data format conversion, projection conversion, etc. (smith et al., 2015; trimble et al., 2015). these additional requirements indicate that web-based map service is a must-to-have function in their portal, which requires a relational database management system (dbms) and a spatial database engine to extend the functionality of the datasets and operate geospatial industry standards. finally, long-term preservation is another consideration for a data management system. although it is not a frequently mentioned topic in most of the prevailing geoportal projects, it was studied in library-centered data management systems and the long-standing data programs such as the datasets managed by nasa (u.s. national aeronautics and space administration) (sweetkind, larsgaard and erwin, 2006; durante and hardy, 2015; khayat and kempler, 2015; beaujardière, 2016). the long-term preservation requires a well-designed technical architecture which is usually independent with the routine geospatial data management system. most of these projects have designated data repository as their preservation solution, and set up a set of preservation protocols, including format registry (to ensure the data format could be used for a long-term), metadata documentation, rights management and contracts, as well as collection management. spatial data management options based on the current technology, there are many options that librarians could take in order to develop their data management system. delivery of spatial data over the internet can be realized in various ways ranging from data transmit via emails to ogc standard web services (crompvoets et al., 2004). while a simple file structure could serve for the purpose, many organizations have chosen a spatial dbms because it offers extended visualization, mapping, and gis capabilities. the major proprietary options of spatial dbms include oracle 10g spatial and the esri enterprise geodatabase which needs to be implemented on top of an enterprise dbms choice, such as microsoft sql. these options provide most of the ogc standard web services and have been widely adopted in many u.s. government agencies and big organizations. the largest user base in free and open source software (foss) market is postgis, which adds spatial data types and analysis functions to the foss database postgresql (bocher and martin, 2012). comparisons with proprietary spatial dbms show that postgis is a comparable alternative when considering functionality, robustness, support and price. other foss spatial dbms include mysql spatial extension which adds spatial support to mysql, and the spatialite project for sqlite. https://doi.org/10.29173/iq927 7/14 kong, ningning nicole (2018) one store has all? the backend story of managing geospatial information toward an easy discovery, iassist quarterly 42 (4), pp. 1-14. doi: https://doi.org/10.29173/iq927 data management solutions in non-library geospatial projects in the non-library setting, there are two major kinds of geospatial data management approaches – file management system and spatial dbms. the file management system is implemented by a metadata inventory which organizes the detailed information about each dataset in a centralized database. the file could be accessed by ftp, odbc, or http protocols. this approach is usually adopted in two extreme cases, either the datasets are simple and consistent enough to manage (ganor, 2017) or the datasets are distributed on multiple servers in very different formats which is impossible to centralize them in one unified system (tsontos and kiefer, 2002; johnson, 2017). spatial dbms is used in most of the geospatial projects because of the additional spatial functionalities, especially the scientific visualization capability. one major kind of implementation is the foss option which includes postgis and geoserver. examples of this kind of implementation include the forest data portal project at european countries (mcinerney et al., 2012) and the tioga project which intends to provide data management for scientific visualization as part of the global change research project (stonebraker et al., 1993). the other major implementation is the adoption of esri products including arcgis server and sql database as discussed by many u.s. and european geospatial portal projects (maguire and longley, 2005; porcal-gonzalo, 2015). there are cases that more customized data management system needs to be developed in order to fit the projects’ needs, especially in the case of handing big data. in one of the nasa’s remote sensing data preservation projects, a customized system was developed based on flexible extensible digital object repository architecture (fedora), which is a digital repository addressing the data management aspects of university and institutional libraries (khayat and kempler, 2015). noaa managed its big data projects by exploring cloud-based infrastructure options including providers from amazon, google, ibm, microsoft and open cloud consortium. noaa not only used these cloud service as storage spaces, but also explored additional functionalities from these platforms such as the unified access framework that improved the discoverability and accessibility of regularly gridded data and the dataset identifier project which assign dois to archived datasets (beaujardière, 2016). data management solutions in library geospatial data projects the data management system in academic libraries varies a lot, ranging from cd-roms, library maintained workstations to enterprise database. serving for the purpose of geoportal, libraries mainly use spatial dbms or institutional data repository as the options to manage their geospatial data collection. more than a dozen of academic libraries have implemented opengeoportal project in their institutions (opengeoportal community, 2017). opengeoportal provides data visualization and download functions for its users. serving for spatial visualization purpose, libraries have the options to use either postgis and geoserver technology stack, or esri enterprise. or, libraries can have the option of not providing the map preview. for the data download function, libraries can choose to provide data via geoserver’s processing function, or ftp/http file transfer protocols, or through other data transfer format. libraries can also build in their physical map collection into the system and provide the call number from the portal. in the next section, metadata shared from participating institutions will be further analyzed to review the data management options from each institution. https://doi.org/10.29173/iq927 8/14 kong, ningning nicole (2018) one store has all? the backend story of managing geospatial information toward an easy discovery, iassist quarterly 42 (4), pp. 1-14. doi: https://doi.org/10.29173/iq927 the inside idaho project uses sql database for data management and arcgis server for web map visualization (kenyon, godfrey and eckwright, 2012). in the florida geographic data library project, data quality assurance/quality control (qaqc) was a major concern in the project design. spatial dbms, more specifically, arcgis products and oracle database were selected because the qaqc process is relatively complex and the dbms provides an excellent environment to store, manage and relate information about the qaqc process (goodison, guillaume thomas and palmer, 2016). drawing upon the increased needs of data management in academic libraries, digital repository is an excellent option for geospatial data management. there are several advantages of choosing data repositories for the geoportal purpose. first of all, many academic libraries already have digital repositories in their institution which accept geospatial datasets, so there is no additional requirement to create a new data infrastructure. second, digital repository could include a wide range of dataset formats, which fits the requirements of geospatial data especially for collections from multiple sources. third, digital repository manages dataset as distinct object which includes metadata, access policy and long-term preservation services. however, in order to visualize the spatial data from the portal, additional visualization function has to be developed and additional infrastructure such as a gis server has to be added. examples of this kind of option include the geohydra project at stanford university and the geoportal developed by the ontario council of university libraries in canada (durante and hardy, 2015; trimble et al., 2015). spatial data management in opengeoportal community the participating institutions in the opengeoportal community shared their metadata database via the portal. i analyzed all the metadata records with an online link to understand the data management strategies in each institution. by the time we harvest the metadata for this analysis, there are 44,603 records with online link information. among them, 75% records are geospatial information shared publicly, and the other 25% are restricted information only available to the specific institution or group of users. the participating institutions and their shared records are shown in figure 2. figure 2. opengeoportal participating institutions and the number of records in their collections. 0 2000 4000 6000 8000 10000 12000 opengeoportal participating institutions https://doi.org/10.29173/iq927 9/14 kong, ningning nicole (2018) one store has all? the backend story of managing geospatial information toward an easy discovery, iassist quarterly 42 (4), pp. 1-14. doi: https://doi.org/10.29173/iq927 figure 3. different types data registered in the opengeoportal community. figure 3 shows the data type of all the records discussed above . overall, about half of the records are vector files saved as polygon, line or point type. ten percent of the records are raster files, and another ten percent are scanned maps. twenty-five percent are defined as paper maps. five percent are undefined library records with an established web link. it must be pointed out that due to the limitations on metadata content standard, institutions might define their data type very differently. for example, “paper map” in harvard university refers to those georeferenced paper maps which are available online as a web map; while in mit literally refers to the library’s physical paper map collection with metadata information existing in library’s record system. in order to learn about the data management practices in each institution, i further analyzed the preview and download link for each vector and raster datasets (including the scanned or paper map records which refer to georeferenced datasets) managed by different institutions. figure 4 shows the type of primary data links from each institution. the majority of the institution chose to use ogc standard web services, wms, wcs, or wfs to disseminate their datasets online. wms is especially the most popular choice and became the dominate type in many universities such as harvard, mit, stanford and university of arizona. most of the institutions using these ogc standards manage their data using geoserver and postgis database. other than that, pasda (pennsylvania spatial data access), purdue, and a small portion of records from university of minnesota use web services offered by arcgis server to manage their geospatial datasets, including esri web map services, web feature services, and image services. a big portion of the records managed by united nation’s food and agriculture organization (fao) are managed by geonetwork. geonetwork is a catalog application to manage spatially referenced resources, and was started in 2001 as a spatial data catalogue system for the fao. without a spatial database engine, geospatial data can still be managed and transferred in a zip file format for users to download. the downloadable zip file via ftp https://doi.org/10.29173/iq927 10/14 kong, ningning nicole (2018) one store has all? the backend story of managing geospatial information toward an easy discovery, iassist quarterly 42 (4), pp. 1-14. doi: https://doi.org/10.29173/iq927 figure 3. data management strategies managed by different institutions. or http(s) is also a commonly used method within the community and it is at least used by five institutions. discussion managing geospatial data is never an easy task. a well-designed data management system along with good metadata practices are essential for academic libraries to manage the great variety of geospatial datasets. visualizing the spatial data using a web interface will be the trend for general information users, as geospatial information is expanding across disciplines (kong, zhang and stonebraker, 2014). this requires libraries to serve these datasets as web map services. to fulfil this need, libraries will need to maintain the geospatial datasets beyond their original formats and start to build spatial dbms and spatial database engines such as geoserver or esri arcgis engerprise to implement these services. in order to provide these map services in a stable, immediate, and up-todate fashion, libraries will need a lot of management efforts to explore and compare different factors contributing to the map services, such as data formats, map service properties, and database versioning so that the data are managed with minimal storage space requirements while being delivered to the web browsers in an interoperable format with a satisfying speed (kong, 2015). many libraries started to use digital repository as a platform to manage their geospatial datasets. digital repositories is an excellent platform to centralize the datasets. however, many academic libraries are still in the early stage of employing data repository in their institutions and the collection policy might not include all the library’s geospatial datasets which were usually collected from different sources. more development needs to be made in order to combine the gis server capability with the digital repository so that the geospatial datasets can be well maintained and disseminated. 0% 10% 20% 30% 40% 50% 60% 70% 80% 90% 100% data management systems arcgisrest zip wcs wfs wms geonetwork figure 4. data management strategies managed by different institutions. https://doi.org/10.29173/iq927 11/14 kong, ningning nicole (2018) one store has all? the backend story of managing geospatial information toward an easy discovery, iassist quarterly 42 (4), pp. 1-14. doi: https://doi.org/10.29173/iq927 finally, it will need a long-term discussion and effort between librarians in order to preserve geospatial datasets. although many agencies provide the current collected geospatial data, it is not in their responsibility to provide the previous decades data, even though these data could provide very important information for researchers. what is the mechanism to preserve these big datasets will be a challenging topic for libraries with increased data collection efforts and increased sensor technology. references abbott, l. t. and argentati, c. d. (1995) ‘gis: a new component of public services’, the journal of academic librarianship, 21(4), pp. 251–256. askew, d., evans, s., matthews, r. and swanton, p. (2005) ‘magic: a geoportal for the english countryside’, computers, environment and urban systems, 29(1 spec.iss.), pp. 71–85. doi: 10.1016/j.compenvurbsys.2004.05.013. association of research libraries (1999) the arl geographic information systems literacy project, spec kit 238. washington, dc. available at: http://www.istl.org/06-fall/refereed.html (accessed: 1 may 2015). batcheller, j. k. (2008) ‘automating geospatial metadata generation-an integrated data management and documentation approach’, computers and geosciences, 34(4), pp. 387–398. doi: 10.1016/j.cageo.2007.04.001. battista, a., majewicz, k., balogh, s. and hardy, d. (2017) ‘consortial geospatial data collection: toward standards and processes for shared geoblacklight metadata’, journal of library metadata. routledge, 17(3–4), pp. 183–200. doi: 10.1080/19386389.2018.1443414. beaujardière, j. d. la (2016) ‘noaa environmental data management’, journal of map & geography libraries, 12(1), pp. 5–27. doi: 10.1080/15420353.2015.1087446. bennett, t. b. and nicholson, s. w. (2007) ‘research libraries: connecting users to numeric and spatial resources’, social science computer review, 25(3), pp. 302–318. doi: 10.1177/0894439306294466. blake, m., majewicz, k., tickner, a. and lam, j. (2017) ‘usability analysis of the big ten academic alliance geoportal: findings and recommendations for improvement of the user experience’, code4lib, (38). bocher, e. and martin, j. (2012) geospatial free and open source software in the 21st century, springer. doi: 10.1007/978-3-642-10595-1. boissé, j. a. and larsgaard, m. (1995) ‘gis in academic libraries: a managerial perspective’, the journal of academic librarianship, 21(4), pp. 288–291. chuck killpack (2011) ‘geospatial content – big data, bigger opportunity’, geospatial world, 1(9), pp. 18–26. crompvoets, j., bregt, a., rajabifard, a. and williamson, i. (2004) ‘assessing the worldwide developments of national spatial data clearinghouses’, international journal of geographical information science, 18(7), pp. 665–689. doi: 10.1080/13658810410001702030. https://doi.org/10.29173/iq927 12/14 kong, ningning nicole (2018) one store has all? the backend story of managing geospatial information toward an easy discovery, iassist quarterly 42 (4), pp. 1-14. doi: https://doi.org/10.29173/iq927 crompvoets, j. and wachowicz, m. (2004) ‘impact assessment of the inspire geo-portal’, proc. of the 10th ec …, (june), pp. 23–25. available at: http://www.ec-gis.org/workshops/10ecgis/papers/23june_crompvoets.pdf. durante, k. and hardy, d. (2015) ‘discovery, management, and preservation of geospatial data using hydra’, journal of map and geography libraries, 11(2), pp. 123–154. doi: 10.1080/15420353.2015.1041630. erwin, t. and sweetkind-singer, j. (2009) ‘the national geospatial digital archive: a collaborative project to archive geospatial data’, journal of map & geography libraries, 6(1), pp. 6–25. doi: 10.1080/15420350903432440. flemons, p., guralnick, r., krieger, j., ranipeta, a. and neufeld, d. (2007) ‘a web-based gis tool for exploring the world’s biodiversity: the global biodiversity information facility mapping and analysis portal application (gbif-mapa)’, ecological informatics, 2(1), pp. 49–60. doi: 10.1016/j.ecoinf.2007.03.004. florance, p., mcgee, m., barnett, c. and mcdonald, s. (2015) ‘the open geoportal federation’, journal of map & geography libraries, 11(3), pp. 376–394. gabaldón, c. and repplinger, j. (2006) ‘gis and the academic library : a survey of libraries offering gis services in two consortia’, issues in science and technology librarianship, 48, pp. 1–8. doi: 10.5062/f4qj7f8r. ganor, t. (2017) ‘an integrated spatial search engine for maps and aerial photographs on a google maps api platform’, journal of map and geography libraries. taylor & francis, 13(2), pp. 175–197. doi: 10.1080/15420353.2016.1277574. goldberg, d., olivares, m., li, z. and klein, a. g. (2014) ‘maps & gis data libraries in the era of big data and cloud computing’, journal of map & geography libraries, 10(1), pp. 100–122. doi: 10.1080/15420353.2014.893944. goodison, c., guillaume thomas, a. and palmer, s. (2016) ‘the florida geographic data library: lessons learned and workflows for geospatial data management’, journal of map & geography libraries, 12(1), pp. 73–99. doi: 10.1080/15420353.2015.1038861. hardy, d. and durante, k. (2014) ‘a metadata schema for geospatial resource discovery use cases’, code4lib, (25). janée, g. (2009) ‘preserving geospatial data: the national geospatial digital archive’s approach’, in archiving 2009 final program and proceedings. arlington, va: society for imaging science and technology, pp. 25–29. available at: http://www.ingentaconnect.com/content/ist/ac/2009/00002009/00000001/art00007 (accessed: 2 april 2018). johnson, v. (2017) ‘leveraging technical library expertise for big data management’, journal of the australian library and information association. routledge, 158, pp. 1–16. doi: 10.1080/24750158.2017.1356982. kenyon, j., godfrey, b. and eckwright, g. z. (2012) ‘geospatial data curation at the university of idaho’, journal of web librarianship, 6(4), pp. 251–262. doi: 10.1080/19322909.2012.729983. khayat, m. and kempler, s. j. (2015) ‘life cycle management considerations of remotely sensed geospatial data and documentation for long term preservation’, journal of map & geography https://doi.org/10.29173/iq927 13/14 kong, ningning nicole (2018) one store has all? the backend story of managing geospatial information toward an easy discovery, iassist quarterly 42 (4), pp. 1-14. doi: https://doi.org/10.29173/iq927 libraries, 11(3), pp. 271–288. doi: 10.1080/15420353.2015.1072122. kollen, c., dietz, c., suh, j. and lee, a. (2013) ‘geospatial data catalogs: approaches by academic libraries’, journal of map and geography libraries, 9(3), pp. 276–295. doi: 10.1080/15420353.2013.820161. kong, n. (2015) ‘exploring best management practices for geospatial data in academic libraries’, journal of map & geography libraries, 11(september), pp. 207–225. doi: 10.1080/15420353.2015.1043170. kong, n., zhang, t. and stonebraker, i. (2014) ‘common metrics for web-based mapping applications in academic libraries’, online information review, 38(7), pp. 918–935. doi: 10.1108/oir-06-20140140. li, w., song, m., zhou, b., cao, k. and gao, s. (2015) ‘performance improvement techniques for geospatial web services in a cyberinfrastructure environment a case study with a disaster management portal’, computers, environment and urban systems. elsevier ltd, 54, pp. 314–325. doi: 10.1016/j.compenvurbsys.2015.04.003. longstreth, k., librarian, m., library, m. and arbor, a. (1995) ‘gis collection development, staffing, and training’, (july), pp. 267–274. maguire, d. j. and longley, p. a. (2005) ‘the emergence of geoportals and their role in spatial data infrastructures’, computers, environment and urban systems, 29(1), pp. 3–14. doi: 10.1016/j.compenvurbsys.2004.05.012. mcinerney, d., bastin, l., diaz, l., figueiredo, c., barredo, j. i. and san-miguel ayanz, j. (2012) ‘developing a forest data portal to support multi-scale decision making’, ieee journal of selected topics in applied earth observations and remote sensing, 5(6), pp. 1692–1999. doi: 10.1109/jstars.2012.2194136. morris, s. p. (2006) ‘geospatial web services and geoarchiving : new opportunities and challenges in geographic information services brief overview of digital geospatial data services in’, geographic information systems and libraries, 55(2), pp. 285–303. newton, m. p., miller, c. c. and bracke, m. s. (2010) ‘librarian roles in institutional repository data set collecting: outcomes of a research library task force’, collection management, 36(1), pp. 53– 67. doi: 10.1080/01462679.2011.530546. nogueras-iso, j., zarazaga-soria, f. j. and muro-medrano, p. p. r. (2005) geographic information metadata for spatial data infrastructures, springer. doi: 10.1007/3-540-27508-8. ogc (2013) ogc standards, open geospatial consortium. available at: http://www.opengeospatial.org/standards/is (accessed: 9 october 2013). opengeometadata (2018). available at: https://github.com/opengeometadata. opengeoportal community (2017) open geoportal. available at: http://opengeoportal.org/ (accessed: 26 march 2017). porcal-gonzalo, m. c. (2015) ‘a strategy for the management, preservation, and reutilization of geographical information based on the lifecycle of geospatial data: an assessment and a proposal based on experiences from spain and europe’, journal of map and geography libraries, 11(3), pp. 289–329. doi: 10.1080/15420353.2015.1064054. https://doi.org/10.29173/iq927 14/14 kong, ningning nicole (2018) one store has all? the backend story of managing geospatial information toward an easy discovery, iassist quarterly 42 (4), pp. 1-14. doi: https://doi.org/10.29173/iq927 schwarb, m., acuña, d., konzelmann, t., rohrer, m., salzmann, n., serpa lopez, b. and silvestre, e. (2011) ‘a data portal for regional climatic trend analysis in a peruvian high andes region’, advances in science and research, 6, pp. 219–226. doi: 10.5194/asr-6-219-2011. smith, j. w., slocumb, w. s., smith, c. and matney, j. (2015) ‘a needs-assessment process for designing geospatial data management systems within federal agencies’, journal of map and geography libraries, 11(2), pp. 226–244. doi: 10.1080/15420353.2015.1048035. stonebraker, m., chen, j., nathan, n., paxson, c. and wu, j. (1993) ‘tioga : providing data management support for scienti c visualization applications 1 introduction 2 the tioga programming paradigm’, vldb, pp. 1–12. strain, l., rajabifard, a. and williamson, i. (2006) ‘marine administration and spatial data infrastructure’, marine policy, 30(4), pp. 431–441. doi: 10.1016/j.marpol.2005.03.005. sweetkind, j., larsgaard, m. l. and erwin, t. (2006) ‘digital preservation of geospatial data’, library trends, 55(2), pp. 304–314. doi: 10.1353/lib.2006.0065. tait, m. g. (2005) ‘implementing geoportals: applications of distributed gis’, computers, environment and urban systems, 29(1 spec.iss.), pp. 33–47. doi: 10.1016/j.compenvurbsys.2004.05.011. trimble, l., woods, c., berish, f., jakubek, d. and simpkin, s. (2015) ‘collaborative approaches to the management of geospatial data collections in canadian academic libraries: a historical case study.’, journal of map & geography libraries, 11(3), pp. 330–358. doi: 10.1080/15420353.2015.1043067. tsontos, v. m. and kiefer, d. a. (2002) ‘the gulf of maine biogeographical information system project: developing a spatial data management framework in support of obis’, oceanologica acta, 25(5), pp. 199–206. doi: 10.1016/s0399-1784(02)01209-4. yang, c. p., cao, y., evans, j., kafatos, m. and bambacus, m. (2006) ‘spatial web portal for building spatial data infrastructure’, geographic information sciences, 12(1), pp. 38–43. doi: 10.1080/10824000609480616. end-notes 1 nicole kong is an assistant professor and gis specialist at purdue university libraries. she can be reached by email: kongn@purdue.edu https://doi.org/10.29173/iq927 1/17 lafia, sara, million, a.j., and hemphill, libby (2024) exploratory and directed search strategies at a social science data archive, iassist quarterly 48(1), pp. 1-17. doi: https://doi.org/10.29173/iq1087 the creative commons-attribution-noncommercial license 4.0 international applies to all works published by iassist quarterly. authors will retain copyright of the work and full publishing rights. exploratory and directed search strategies at a social science data archive sara lafia1, a.j. million2, libby hemphill3 abstract researchers need to be able to find, access, and use data to participate in open science. to understand how users search for research data, we analyzed textual queries issued at a large social science data archive, the inter-university consortium for political and social research (icpsr). we collected unique user queries from 988,475 user search sessions over four years (2012-16). overall, we found that only 30% of site visitors entered search terms into the icpsr website. we analyzed search strategies within these sessions by extending existing dataset search taxonomies to classify a subset of the 1,554 most popular queries. we identified five categories of commonly-issued queries: keyword-based (e.g., date, place, topic); name (e.g., study, series); identifier (e.g., study, series); author (e.g., institutional, individual); and type (e.g., file, format). while the dominant search strategy used short keywords to explore topics, directed searches for known items using study and series names were also common. we further distinguished exploratory browsing from directed search queries based on their page views, refinements, search depth, duration, and length. directed queries were longer (i.e., they had more words), while sessions with exploratory queries had more refinements and associated page views. by comparing search interactions at icpsr to other natural language interactions in similar web search contexts, we conclude that dataset search at icpsr is underutilized. we envision how alternative search paradigms, such as those enabled by recommender systems, can enhance dataset search. keywords research data, information search, query log analysis, user behavior, web analytics introduction data sharing in the social sciences allows researchers to build upon the work of others. funders require awardees to share their data to increase scientific efficiency, enhance research transparency, and promote fair access, among other benefits (national research council et al., 1985). however, data sharing does not guarantee discoverability or reuse by others. data curation activities, such as creating descriptive metadata and documentation, promote data findability, accessibility, interoperability, and reuse (levenstein & lyle, 2018). large-scale data archives, such as the inter-university consortium for political and social research (icpsr), support long-term data preservation and provide data curation services to enhance the quality of deposited data (akmon et al., 2020). data archives also offer search and discovery tools for data retrieval (pienta et al., 2018). prior research has studied the impact of https://doi.org/10.29173/iq1087 https://paperpile.com/c/jjxwah/sgixn https://paperpile.com/c/jjxwah/klmws https://paperpile.com/c/jjxwah/x8owq https://paperpile.com/c/jjxwah/l6ln6 https://creativecommons.org/licenses/by-nc/4.0/ 2/17 lafia, sara, million, a.j., and hemphill, libby (2024) exploratory and directed search strategies at a social science data archive, iassist quarterly 48(1), pp. 1-17. doi: https://doi.org/10.29173/iq1087 curation and archiving decisions on data reuse (he & han, 2017; hemphill et al., 2021); however, less is known about intermediate data discovery steps, such as the specific sequences of actions that users take when seeking data (lafia et al., 2023) and disciplinary search strategies for finding research data (gregory et al., 2020; kacprzak et al., 2017). search systems facilitate information discovery and retrieval in several ways. in particular, academic search tasks are often exploratory and complex. they emphasize learning and discovery, are often illdefined, multi-perspective, and require browsing to support learning alongside search (r. w. white, 2016). academic search tasks require support for the user as they “learn” a knowledge domain (h. d. white et al., 2004). ideally, “context-driven discovery” allows users to learn as they search and gain proficiency with a given subject (solomon, 2002). approaches, such as the visualizations of scientific terms (e.g., in maps), balance designer-initiated (global) and user-driven (local) conceptualizations (börner et al., 2003). other design considerations, such as search facets, can guide users to understand possible kinds of interactions within a system (hearst, 2009). well-designed systems overcome human-system communication's “vocabulary problem” (furnas et al., 1987) by aligning user concepts with system specifications. this is important for supporting interdisciplinary research, where various disciplinary terms may describe similar phenomena (institute of medicine et al., 2005) across multiple levels of expertise (hembrooke et al., 2005). importantly, search systems must also balance exploratory and directed search tasks by allowing users to retrieve known items (r. w. white, 2016). search interfaces are often evaluated based on their support of user search strategies, including monitoring, file structure, search formulation, term, and idea tactics (wilson et al., 2009). our prior work found that users follow direct, orienting, and scenic search paths while navigating dataset searches at a large-scale social science data archive (lafia et al., 2023). approaches proposed to increase the accessibility of archival collections include introducing novel finding aids that support federated queries across collections and adding context to boost search relevancy within collections (renspie et al., 2015). archives and repositories can develop responsive systems that encourage dataset discovery and reuse by learning from how users search for data. to understand how prospective users search for curated social science research data, we analyzed 1,554 unique user queries issued across 988,475 user search sessions spanning four years (2012-16) at icpsr. we asked: 1) what are the most common features of queries issued at a large-scale social science data archive?; and 2) what strategies do prospective users employ to search for research data? based on our analysis, we discuss opportunities for improving data discovery and eventual reuse by supporting exploratory and directed search strategies. background query or transaction logs provide a foundation for analyzing human information behavior (hib). while hib models provide a theoretical basis for representing user search behavior (bates, 1989; marchionini, 1997; meho & tibbo, 2003), query logs offer detailed insights into search strategies that users employ in their everyday lives (jiang et al., 2013). taxonomies bridge search log analysis and theoretical models by describing high-level patterns in users’ observed search behavior. for instance, https://doi.org/10.29173/iq1087 https://paperpile.com/c/jjxwah/bgjtr+n8nok https://paperpile.com/c/jjxwah/qhfx9 https://paperpile.com/c/jjxwah/hvoi4+bboky https://paperpile.com/c/jjxwah/19vhm https://paperpile.com/c/jjxwah/19vhm https://paperpile.com/c/jjxwah/vx6id https://paperpile.com/c/jjxwah/vx6id https://paperpile.com/c/jjxwah/ko8kp https://paperpile.com/c/jjxwah/eyrta https://paperpile.com/c/jjxwah/rnteu https://paperpile.com/c/jjxwah/gezio https://paperpile.com/c/jjxwah/hbxqz https://paperpile.com/c/jjxwah/h3zva https://paperpile.com/c/jjxwah/19vhm https://paperpile.com/c/jjxwah/txeww https://paperpile.com/c/jjxwah/qhfx9 https://paperpile.com/c/jjxwah/dw8dv https://paperpile.com/c/jjxwah/5zosp+aquwu+fq4f0 https://paperpile.com/c/jjxwah/5zosp+aquwu+fq4f0 https://paperpile.com/c/jjxwah/jz5l9 3/17 lafia, sara, million, a.j., and hemphill, libby (2024) exploratory and directed search strategies at a social science data archive, iassist quarterly 48(1), pp. 1-17. doi: https://doi.org/10.29173/iq1087 log analysis has been used to summarize the intent behind commercial web searches as navigational, informational, and transactional (broder, 2002). query log analyses have been applied to study commercial search engines (kumar & tomkins, 2010; silverstein et al., 1999), digital libraries (carevic et al., 2020; jones et al., 2000), and data portals (degbelo, 2020; kacprzak et al., 2017). query log analysis can be used to enhance clickthrough search performance (joachims, 2002), infer users’ information needs by analyzing search topics (abebe et al., 2018), and appraise gaps in collections by identifying failed searches (pienta et al., 2018). analyses can be constrained (e.g., to a single day of searches issued on a given portal) (herskovic et al., 2007) or cover longer durations to study changing user behaviors (e.g., characterize search as a learning process) (eickhoff et al., 2014). information behavior can also be inferred from users’ responses to search systems. for example, during query refinement or reformulation, users modify their search queries to retrieve more relevant results; query modification feedback can be explicitly provided by the user (e.g., clicks within a query) or implicitly derived by the system (e.g., semantic document similarity mining) (baeza-yates & ribeironeto, 2011, chapter 5). prior work has identified unique considerations for designing dataset retrieval systems (wang et al., 2021). for example, while systems index datasets as discrete objects, users may want to perform interactions such as combination and subsetting (chapman et al., 2019). leading dataset search systems, such as google’s dataset search, rely on original, high-quality metadata for indexing (brickley et al., 2019). other dataset search systems, such as government data portals, encourage users to explore and browse for data rather than issue known-item searches (kacprzak et al., 2017). users’ dataset search strategies also vary across domains; for example, social scientists tend to trace publication references and explore survey data banks more than earth scientists and astronomers, who follow “bounded” strategies (e.g., searching by journal, location, and time) to find data (gregory et al., 2019). social scientists need descriptive metadata to support their search needs; these include contextual information about prior data use (e.g., evidenced in publication citations) (faniel et al., 2019). however, existing systems do not tend to include explicit, contextual information about how data have been reused by others or curated, for example, in search indexes (sun & khoo, 2017). generally, users’ information needs are often far more detailed and expressive than the dataset search queries that they issue (papenmeier et al., 2021). in this study, we analyze query logs to develop a baseline understanding of users’ expressed information needs and search behaviors when seeking social science data. methods we analyzed user search queries at the inter-university consortium for political and social research (icpsr), a large social science data archive. specifically, we used google analytics (ga) to track user queries issued through the icpsr website’s search box (i.e., “site searches”) across research metadata, variables, data-related publications, and documentation about icpsr. icpsr holdings include over 250,000 data files in 10,000 public use studies and 295 series. ga is set to omit searches performed by icpsr staff based on ip addresses. we recognize that google analytics collects far more data than https://doi.org/10.29173/iq1087 https://paperpile.com/c/jjxwah/7dcls https://paperpile.com/c/jjxwah/jnhtn+qidm4 https://paperpile.com/c/jjxwah/jnhtn+qidm4 https://paperpile.com/c/jjxwah/ifimo+jq2ks https://paperpile.com/c/jjxwah/hvoi4+guifp https://paperpile.com/c/jjxwah/vzn7u https://paperpile.com/c/jjxwah/54gyz https://paperpile.com/c/jjxwah/54gyz https://paperpile.com/c/jjxwah/l6ln6 https://paperpile.com/c/jjxwah/pzuiw https://paperpile.com/c/jjxwah/gbcvd https://paperpile.com/c/jjxwah/fust1/?locator_label=chapter&locator=5 https://paperpile.com/c/jjxwah/fust1/?locator_label=chapter&locator=5 https://paperpile.com/c/jjxwah/4tffh https://paperpile.com/c/jjxwah/btwdd https://paperpile.com/c/jjxwah/wmrpp https://paperpile.com/c/jjxwah/hvoi4 https://paperpile.com/c/jjxwah/hvoi4 https://paperpile.com/c/jjxwah/7nkex https://paperpile.com/c/jjxwah/7nkex https://paperpile.com/c/jjxwah/zhtcn https://paperpile.com/c/jjxwah/zhtcn https://paperpile.com/c/jjxwah/aifsl https://paperpile.com/c/jjxwah/1ouxe 4/17 lafia, sara, million, a.j., and hemphill, libby (2024) exploratory and directed search strategies at a social science data archive, iassist quarterly 48(1), pp. 1-17. doi: https://doi.org/10.29173/iq1087 we present here, and in doing so presents privacy risks for icpsr site visitors. we do not control the google analytics settings at icpsr, but we mitigate these risks in our study by minimizing the data used here to only variables of interest and not those that could identify individual site users. using older data also helps mitigate risks to individuals – for instance, someone searching “crime” in the data we analyzed may no longer be connected to that research topic. we only considered the 30% of sessions (988,475/3,434,937) that included site search interactions. from these sessions, we collected all unique user queries issued across user search sessions from 9/1/2012-9/1/2016. we selected the period for our analysis based on the stability of icpsr’s website design and the consistency of available ga data; site changes to ga since 2016 made more recent data challenging to analyze. data processing we processed website queries using open refine, a data-cleaning tool. we removed whitespace, normalized text to lowercase, removed punctuation, transformed plural to singular forms, checked spelling, and clustered similar query strings. this approach matched queries that contained the same words in different orders (“crime mental illness”, “mental illness crime”) and deduplicated nearly identical queries by merging them into a single entry. we did not, however, merge name variants or synonyms (“national longitudinal study of adolescent health”, “add health”, “nls”) since these reflected diverse search strategies. by applying these rules, we merged a total of 900 queries. query classification to classify queries, we first aligned and extended existing categories of data-specific queries proposed by kacprzak et al. (2017) and pienta et al. (2018). we selected these categories based on their relevance to the task of describing and classifying data-related queries from search logs. a summary of the categories and the rules used to code the icpsr queries is provided in table 1. the prior analysis by pienta (2018) found that users relied on exploratory keywords – indicating subjects, locations, and timeframes – along with directed terms corresponding to known items – such as studies, series, and author names – when searching for datasets. we used these categories (keyword; name; number; author; and format) to code the 1,554 most popular queries in our sample, which were present in more than 57% of all search sessions (562,723/988,475) and which users searched for more than 100 times across all user sessions in our sample. two authors agreed on a schema based on the taxonomies mentioned above. then two authors coded a subsample together until they reached agreement. the first author then coded the remaining items. all coding was conducted by the first author. one of thecategories, listed in table 1, was then assigned to each query. most queries that contained multiple categories (e.g., “chinese household income 2002”) referred to study or series names; however, in ambiguous cases (e.g., “english second language in texas”), the category with more words or that appeared first in the query string was assigned. to interpret the coded queries, they were grouped into one of two search task categories: exploratory, corresponding to searches using keywords or formats, or directed, corresponding to searches by name or author (r. w. white & roth, 2009). exploratory searches facilitate browsing for unknown items, whereas directed searches indicate a specific item that the user is seeking. table 1. query classification scheme and alignment with prior categories https://doi.org/10.29173/iq1087 https://paperpile.com/c/jjxwah/hvoi4 https://paperpile.com/c/jjxwah/l6ln6 https://paperpile.com/c/jjxwah/l6ln6 https://paperpile.com/c/jjxwah/5zosp+suulf+3osgi 5/17 lafia, sara, million, a.j., and hemphill, libby (2024) exploratory and directed search strategies at a social science data archive, iassist quarterly 48(1), pp. 1-17. doi: https://doi.org/10.29173/iq1087 category rules related category from pienta et al. (2018) related category from kacprzak et al. (2017) keyword place, date, topic (exploratory) includes a geographic place name, time, or concept. keyword or phrase (e.g., “diabetes”) location (name of city, town, geographical area) time frame (years, months, weekday) format type (exploratory) uses the name of a known file format or analysis method. file and dataset type (.csv, .pdf, html, table) name study, series; number study, series (directed) uses a number in icpsr’s study or series number range (not a year or other identifier). study name (e.g., “icpsr 2896”) numbers named serial collection (e.g., “nsduh”) abbreviations (acronyms from controlled list or manually verified) author institutional, individual (directed) includes an author’s full or last name, or uses the name of an organization. author/principal investigator name (e.g., “lillard”) feature selection to characterize groups of queries classified in our analysis, we selected features from google analytics described in table 2. we chose these query-level features based on prior findings by kathuria et al. (2010), who defined query intent using query-level features, such as query length and reformulation strategy. we also based our feature selections on work by sharifpour et al. (2022), who proposed distinct user groups by performing hierarchical clustering on query logs (2022). we selected these categories based on their relevance to differentiating user behavior and profiling users based on their web queries. the features we selected (google analytics, 2023) were: results page views per search (i.e., the number of items a user looked at after searching); percent search refinements (i.e., the share of sessions where a user adjusted or reformulated their search); average search depth (i.e., number of pages clicked on following a search); time after search (i.e., amount of time spent in the session after a search); and query length (e.g., number of words in the query). table 2. features extracted from google analytics to characterize queries feature definition from google analytics related category from kathuria et al. related category from sharifpour et al. https://doi.org/10.29173/iq1087 https://paperpile.com/c/jjxwah/l6ln6 https://paperpile.com/c/jjxwah/hvoi4 https://paperpile.com/c/jjxwah/nz5h0 https://paperpile.com/c/jjxwah/olo2i https://paperpile.com/c/jjxwah/olo2i https://paperpile.com/c/jjxwah/2dlk3 https://paperpile.com/c/jjxwah/2dlk3 https://paperpile.com/c/jjxwah/2dlk3 6/17 lafia, sara, million, a.j., and hemphill, libby (2024) exploratory and directed search strategies at a social science data archive, iassist quarterly 48(1), pp. 1-17. doi: https://doi.org/10.29173/iq1087 (kathuria et al., 2010) (sharifpour et al., 2022) results page views per search views of search result pages divided by total unique searches results viewed page views percent search refinements repeated searches using another term divided by views of search result pages query reformulation average search depth average number of pages viewed after performing a search total click-throughs time after search amount of time in seconds users spend on site after performing a search total time spent query length number of terms contained in a particular query number of query terms number of unique query terms results users searched with short, unique phrases to characterize the queries, we measured their lengths and checked if they contained interrogative terms (e.g., “who”). on average, queries were shorter than two words, meaning that most users entered a single word or phrase. exploratory searches, which facilitated browsing and were not directed to retrieve known items (r. w. white & roth, 2009), were shorter on average than directed searches for known items, such as the names of social science studies (1.5 words versus 2.7 words). in terms of query formulation, only four queries contained one or more interrogative keywords proposed by bendersky and croft (2009) suggesting a question (e.g., the word “do” indicates the question in: “'do children of asian immigrants speak english in the home more often than children of latino immigrants?'”). few popular search terms were shared across users; instead, users searched with distinct query terms, resulting in a long-tailed distribution (figure 1). for example, the most popular query in our sample (icpsr study number “21600”) was issued 10,148 times, while many more queries (“surveillance”, “infertility”, “religious attitudes”) were only issued 100 times each. the figure also shows that there were only 0-50 query terms that were used more than 1,000 times. https://doi.org/10.29173/iq1087 https://paperpile.com/c/jjxwah/nz5h0 https://paperpile.com/c/jjxwah/olo2i https://paperpile.com/c/jjxwah/olo2i https://paperpile.com/c/jjxwah/5zosp+suulf+3osgi https://paperpile.com/c/jjxwah/iwvwd 7/17 lafia, sara, million, a.j., and hemphill, libby (2024) exploratory and directed search strategies at a social science data archive, iassist quarterly 48(1), pp. 1-17. doi: https://doi.org/10.29173/iq1087 figure 1. histogram with fixed size bins (bins=50) indicating the number of unique query terms (xaxis) and the frequency with which they were searched (y-axis) figure 2. treemap of labeled queries shows that search by topic and name were most common searches were dominated by topics and names we classified more than 66% (1,030/1,554) of the queries as “keyword (topic)”, meaning that the user entered one or more social science subject terms into the site search box. the second largest category of queries used part of or the full “name (series, study)” of a social science study or series. searches by “number (series, study)”, “author (institutional, individual)”, and “format (type)” were the least common kinds s (figure 2). https://doi.org/10.29173/iq1087 8/17 lafia, sara, million, a.j., and hemphill, libby (2024) exploratory and directed search strategies at a social science data archive, iassist quarterly 48(1), pp. 1-17. doi: https://doi.org/10.29173/iq1087 exploratory searches included more refinements and page views most queries were exploratory (73%), which included keyword and format-based searches, while directed queries (27%) included study and series names, numbers, and authors. we summarized the distributions for each feature across the exploratory and directed query groups (figure 3). we observed that sessions had a similar search duration (in seconds) and search depth (by page views) across query types. however, directed queries tended to be longer than exploratory ones. sessions with exploratory queries included more refinements, meaning that users edited and re-issued search terms more often; exploratory sessions also included more result page views than their directed search counterparts, suggesting that they enabled more browsing and navigation behaviors. figure 3. enhanced boxplots show analytics features of exploratory and directed queries discussion by analyzing query logs from icpsr, we determined that data searches use exploratory and directed queries to find research data. our findings align with prior studies of information seeking, which differentiate between exploratory and directed search tasks (bates, 1989; marchionini, 2006; r. w. white & roth, 2009). keyword-based queries that use dates, places, or topics to search suggest that https://doi.org/10.29173/iq1087 https://paperpile.com/c/jjxwah/5zosp+suulf+3osgi https://paperpile.com/c/jjxwah/5zosp+suulf+3osgi 9/17 lafia, sara, million, a.j., and hemphill, libby (2024) exploratory and directed search strategies at a social science data archive, iassist quarterly 48(1), pp. 1-17. doi: https://doi.org/10.29173/iq1087 users do not have known items in mind. at the same time, searches for particular study or series names, numbers, and authors are better characterized as “information lookup” tasks, in which users expect to retrieve a specific item (buckland, 1979). while users who issued exploratory queries were able to navigate icpsr’s website and take additional actions, such as expanding and refining their searches, users may benefit from more explicit support for query reformation; icpsr offers search facets, but their integration with users’ queries could be enhanced (hearst, 2006). by pointing users to semantically related resources, query reformulation would be helpful to prevent users from exiting the website after issuing searches that have few or no search results (pienta et al., 2018). in addition, the popularity of known-item searches suggests that there may be a need for additional functions, such as search “bookmarks”, to help users store and take shortcuts to retrieve their previous queries and search results (aula et al., 2005). one question we are unable to answer with our data is whether researchers’ search strategies at icpsr are successful – i.e., do they find the data they need? we were not able to compare icpsr queries with searches at other archives (e.g., gesis, roper) or with aggregators that search across archives (e.g., google dataset search). the exploratory strategies evident in the icpsr query terms indicate that users do not always know what they are looking for. tools (e.g., metasearches) and training (e.g., teaching researchers to search multiple archives’ holdings) to help searches find data, wherever it resides, would likely be useful. the brevity of queries, and the lack of questions entered in icpsr’s site search, may indicate room for improvement to the user experience. kacprzak et al. (2017) for example, found that most queries entered into open government data portals were a single word in length. shorter queries may indicate that users are not confident in the capabilities of search engines to interpret intent in complex prompts and return relevant results (jansen & spink, 2006). user behavior reflected in icpsr’s search logs suggests that site search is generally treated as an entry point for data exploration. as dataset search matures, queries may start to resemble other kinds of web-based search interactions, such as more complex, natural language queries issued to commercial search engines (taghavi et al., 2012). comparative analysis of user search behavior would help to establish the prevalence of exploratory search across other curated data repositories that offer text-based search against metadata with text descriptions and controlled vocabularies. given that icpsr shares many core features with other archives (e.g., faceted search, variable indexing, controlled metadata), the user behaviors we identify in the present study are likely generalizable. given that keyword-based exploration by topic is the dominant category of search interaction at icpsr (i.e., 66% of the queries we coded), we plan to extend our analysis by investigating the relationships between the topics of user queries and search results. we are interested in exploring the potential to support query expansion with word embeddings. developing a finer set of categories such as social science methods (“factor analysis”), populations of interest (“homeless youth”), and historical events (“hurricane katrina”) would allow for detailed search refinements. the prevalence of specific words and phrases may help predict query intent. prior studies of search behavior at icpsr found evidence of the stability of search topics across time (pienta et al., 2018). thus, detecting significant shifts in topic popularity may also be informative for supporting data search and discovery. https://doi.org/10.29173/iq1087 https://paperpile.com/c/jjxwah/2zteo https://paperpile.com/c/jjxwah/nyw8k https://paperpile.com/c/jjxwah/l6ln6 https://paperpile.com/c/jjxwah/zcs0p https://paperpile.com/c/jjxwah/hvoi4 https://paperpile.com/c/jjxwah/ge0tt https://paperpile.com/c/jjxwah/zjtcq https://paperpile.com/c/jjxwah/l6ln6 10/17 lafia, sara, million, a.j., and hemphill, libby (2024) exploratory and directed search strategies at a social science data archive, iassist quarterly 48(1), pp. 1-17. doi: https://doi.org/10.29173/iq1087 in terms of our study’s limitations, we were restricted to using queries issued between 2012 and 2016. we selected this time period based on a number of changes made to icpsr’s analytics collection process. while our findings are still relevant based on the stability of icpsr’s website and catalog design, we recognize that user search behavior may have shifted in more recent years in response to new search technologies and an increasing emphasis on interdisciplinary research. we also focused on the relationship between the most popular queries and features identified in prior studies. this meant that less popular queries, which may not be well-supported by icpsr’s search system, were omitted from our analysis. to code the queries, we developed a scheme where a single label neatly described most queries; however, we encountered queries that would be better represented by multiple labels (e.g., “chicago homicide” includes place and topic keywords). in some cases, it was also unclear whether a search (e.g., “are you happy”) referred to a known variable label or was an exploratory topic. we note that we are also limited in the inferences we can draw from queries alone, which indicate how users approach search, but do not describe what exactly users are evaluating or their internal cognitive states. we plan to triangulate the present analysis with information about types of users and their narratives about search experiences drawn from interviews that we are conducting to develop search personas for recommendation systems. conclusion by charting the sequences of actions users take to discover research data (lafia et al., 2023), and describing directed and exploratory data search strategies, we are better positioned to propose responsive search tools that support research data discovery and encourage data reuse. analysis of search logs at icpsr shows the prevalence of exploratory search behavior within data archives. improving search and discovery tools for data exploration also supports users with additional needs, such as learning and gaining expertise in a new research domain. while current search methods support exploratory browsing and known-item retrieval for research data to varying degrees, the ability to explore semantically related datasets still needs to be improved. for example, current search processes help users identify data with relevant keywords, but do not help users find similar data, such as those with descriptions indicating subjects, geographies, or methods in common. in addition, site search at icpsr is underutilized, and the ways that users query the system are limited. directed searches were longer and more descriptive than exploratory searches. compared to directed searches, exploratory searches required users to expend more effort to refine their queries and review results. future work will explore approaches, such as recommender systems and aggregators, that balance search efficiency with data exploration to support the serendipitous discovery of research data available in archives, such as icpsr. acknowledgments we thank aalap doshi, lara cooper, sai sandeep reddy bedadala, and the user experience engineering team at icpsr. https://doi.org/10.29173/iq1087 https://paperpile.com/c/jjxwah/qhfx9 11/17 lafia, sara, million, a.j., and hemphill, libby (2024) exploratory and directed search strategies at a social science data archive, iassist quarterly 48(1), pp. 1-17. doi: https://doi.org/10.29173/iq1087 award information this material is based upon work supported by the national science foundation under grant 2121789. data availability the code and data for this project are available in a github repository (https://github.com/icpsr/query-analysis). https://doi.org/10.29173/iq1087 https://github.com/icpsr/query-analysis 12/17 lafia, sara, million, a.j., and hemphill, libby (2024) exploratory and directed search strategies at a social science data archive, iassist quarterly 48(1), pp. 1-17. doi: https://doi.org/10.29173/iq1087 references abebe, r., hill, s., vaughan, j. w., small, p. m., & andrew schwartz, h. (2018). using search queries to understand health information needs in africa. in arxiv [cs.cy]. arxiv. https://doi.org/10.48550/arxiv.1806.05740 akmon, d., lafia, s., thomer, a., hemphill, l., pienta, a., yakel, e., bleckley, d., & tyler, a. (2020). measuring and improving the efficacy of curation activities in data archives. https://hdl.handle.net/2027.42/163501 aula, a., jhaveri, n., & käki, m. (2005). information search and re-access strategies of experienced web users. proceedings of the 14th international conference on world wide web, 583–592. https://dl.acm.org/doi/10.1145/1060745.1060831 baeza-yates, r., & ribeiro-neto, b. (2011). modern information retrieval: the concepts and technology behind search. addison wesley: edinburgh. bates, m.j. (1989), "the design of browsing and berrypicking techniques for the online search interface", online review, vol. 13 no. 5, pp. 407-424. https://doi.org/10.1108/eb024320 bendersky, m., & croft, w. b. (2009). analysis of long queries in a large scale search log. proceedings of the 2009 workshop on web search click data, 8–14. https://doi.org/10.1145/1507509.1507511 börner, k., chen, c., & boyack, k. w. (2003). visualizing knowledge domains. annual review of information science and technology, 37(1), 179–255. https://doi.org/10.1002/aris.1440370106 brickley, d., burgess, m., & noy, n. (2019). google dataset search: building a search engine for datasets in an open web ecosystem. the world wide web conference on www ’19, 1365– 1375. https://doi.org/10.1145/3308558.3313685 broder, a. (2002). a taxonomy of web search. sigir forum, 36(2), 3–10. https://doi.org/10.1145/792550.792552 buckland, m. k. (1979). on types of search and the allocation of library resources. journal of the american society for information science. american society for information science, 30(3), 143– 147. https://doi.org/10.1002/asi.4630300305 carevic, z., roy, d., & mayr, p. (2020). characteristics of dataset retrieval sessions: experiences from https://doi.org/10.29173/iq1087 http://paperpile.com/b/jjxwah/54gyz http://paperpile.com/b/jjxwah/54gyz http://paperpile.com/b/jjxwah/54gyz http://paperpile.com/b/jjxwah/54gyz https://doi.org/10.48550/arxiv.1806.05740 http://paperpile.com/b/jjxwah/x8owq http://paperpile.com/b/jjxwah/x8owq http://paperpile.com/b/jjxwah/x8owq http://paperpile.com/b/jjxwah/x8owq http://paperpile.com/b/jjxwah/x8owq https://hdl.handle.net/2027.42/163501 http://paperpile.com/b/jjxwah/zcs0p http://paperpile.com/b/jjxwah/zcs0p http://paperpile.com/b/jjxwah/zcs0p file:///c:/users/homepc/documents/iassist%20quarterly%20processing/march%202024%20issue/,%20583–592.%20https:/dl.acm.org/doi/10.1145/1060745.1060831 file:///c:/users/homepc/documents/iassist%20quarterly%20processing/march%202024%20issue/,%20583–592.%20https:/dl.acm.org/doi/10.1145/1060745.1060831 http://paperpile.com/b/jjxwah/fust1 http://paperpile.com/b/jjxwah/fust1 http://paperpile.com/b/jjxwah/fust1 http://paperpile.com/b/jjxwah/fust1 https://doi.org/10.1108/eb024320 http://paperpile.com/b/jjxwah/iwvwd http://paperpile.com/b/jjxwah/iwvwd http://paperpile.com/b/jjxwah/iwvwd http://paperpile.com/b/jjxwah/iwvwd https://doi.org/10.1145/1507509.1507511 http://paperpile.com/b/jjxwah/eyrta http://paperpile.com/b/jjxwah/eyrta http://paperpile.com/b/jjxwah/eyrta http://paperpile.com/b/jjxwah/eyrta http://paperpile.com/b/jjxwah/eyrta http://paperpile.com/b/jjxwah/eyrta https://doi.org/10.1002/aris.1440370106 http://paperpile.com/b/jjxwah/wmrpp http://paperpile.com/b/jjxwah/wmrpp http://paperpile.com/b/jjxwah/wmrpp http://paperpile.com/b/jjxwah/wmrpp http://paperpile.com/b/jjxwah/wmrpp https://doi.org/10.1145/3308558.3313685 http://paperpile.com/b/jjxwah/7dcls http://paperpile.com/b/jjxwah/7dcls http://paperpile.com/b/jjxwah/7dcls http://paperpile.com/b/jjxwah/7dcls http://paperpile.com/b/jjxwah/7dcls https://doi.org/10.1145/792550.792552 http://paperpile.com/b/jjxwah/2zteo http://paperpile.com/b/jjxwah/2zteo http://paperpile.com/b/jjxwah/2zteo http://paperpile.com/b/jjxwah/2zteo http://paperpile.com/b/jjxwah/2zteo file:///c:/users/homepc/documents/iassist%20quarterly%20processing/march%202024%20issue/(3),%20143–147 file:///c:/users/homepc/documents/iassist%20quarterly%20processing/march%202024%20issue/(3),%20143–147 https://doi.org/10.1002/asi.4630300305 http://paperpile.com/b/jjxwah/jq2ks 13/17 lafia, sara, million, a.j., and hemphill, libby (2024) exploratory and directed search strategies at a social science data archive, iassist quarterly 48(1), pp. 1-17. doi: https://doi.org/10.29173/iq1087 a real-life digital library. digital libraries for open knowledge, 185–193. https://doi.org/10.1007/978-3-030-54956-5_14 chapman, a., simperl, e., koesten, l., konstantinidis, g., ibáñez, l.-d., kacprzak, e., & groth, p. (2019). dataset search: a survey. the vldb journal: very large data bases: a publication of the vldb endowment. https://doi.org/10.1007/s00778-019-00564-x degbelo, a. (2020). open data user needs: a preliminary synthesis. companion proceedings of the web conference 2020, 834–839. https://doi.org/10.1145/3366424.3386586 eickhoff, c., teevan, j., white, r., & dumais, s. (2014). lessons from the journey: a query log analysis of within-session learning. proceedings of the 7th acm international conference on web search and data mining, 223–232. https://doi.org/10.1145/2556195.2556217 faniel, i. m., frank, r. d., & yakel, e. (2019). context from the data reuser’s point of view. journal of documentation, 75(6), 1274–1297. https://doi.org/10.1108/jd-08-2018-0133 furnas, g. w., landauer, t. k., gomez, l. m., & dumais, s. t. (1987). the vocabulary problem in humansystem communication. communications of the acm, 30(11), 964–971. https://doi.org/10.1145/32206.32212 google analytics. (2023). https://support.google.com/analytics/ gregory, k., groth, p., cousijn, h., scharnhorst, a., & wyatt, s. (2019). searching data: a review of observational data retrieval practices in selected disciplines. journal of the association for information science and technology, 70(5), 419–432. https://doi.org/10.1002/asi.24165 gregory, k., groth, p., scharnhorst, a., & wyatt, s. (2020). lost or found? discovering data needed for research. harvard data science review. hearst, m. (2006). design recommendations for hierarchical faceted search interfaces. acm sigir workshop on faceted search, 1–5. https://flamenco.berkeley.edu/papers/faceted-workshop06.pdf hearst, m. (2009). search user interfaces. cambridge university press. he, l., & han, z. (2017). do usage counts of scientific data make sense? an investigation of the dryad repository. library hi tech, 35(2), 332–342. https://doi.org/10.1108/lht-12-2016-0158 https://doi.org/10.29173/iq1087 http://paperpile.com/b/jjxwah/jq2ks http://paperpile.com/b/jjxwah/jq2ks file:///c:/users/homepc/documents/iassist%20quarterly%20processing/march%202024%20issue/,%20185–193 file:///c:/users/homepc/documents/iassist%20quarterly%20processing/march%202024%20issue/,%20185–193 https://doi.org/10.1007/978-3-030-54956-5_14 https://doi.org/10.1007/s00778-019-00564-x https://doi.org/10.1145/3366424.3386586 https://doi.org/10.1145/2556195.2556217 https://doi.org/10.1108/jd-08-2018-0133 https://doi.org/10.1145/32206.32212 https://support.google.com/analytics/ https://doi.org/10.1002/asi.24165 https://flamenco.berkeley.edu/papers/faceted-workshop06.pdf https://doi.org/10.1108/lht-12-2016-0158 14/17 lafia, sara, million, a.j., and hemphill, libby (2024) exploratory and directed search strategies at a social science data archive, iassist quarterly 48(1), pp. 1-17. doi: https://doi.org/10.29173/iq1087 hembrooke, h. a., granka, l. a., & gay, g. k. (2005). the effects of expertise and feedback on search term selection and subsequent learning. journal of the american society for information science and technology. https://doi.org/10.1002/asi.20180 hemphill, l., pienta, a., lafia, s., akmon, d., & bleckley, d. (2021). how do properties of data, their curation, and their funding relate to reuse? journal of the american society for information science and technology, 73(10), 1432–1444. https://doi.org/10.1002/asi.24646 herskovic, j. r., tanaka, l. y., hersh, w., & bernstam, e. v. (2007). a day in the life of pubmed: analysis of a typical day’s query log. journal of the american medical informatics association: jamia, 14(2), 212–220. https://doi.org/10.1197/jamia.m2191 icpsr thesaurus. (2023). https://www.icpsr.umich.edu/web/icpsr/thesaurus institute of medicine, national academy of engineering, national academy of sciences, committee on science, engineering, and public policy, & committee on facilitating interdisciplinary research. (2005). facilitating interdisciplinary research. national academies press. jansen, b. j., & spink, a. (2006). how are we searching the world wide web? a comparison of nine search engine transaction logs. information processing & management, 42(1), 248–263. https://doi.org/10.1016/j.ipm.2004.10.007 jiang, d., pei, j., & li, h. (2013). mining search and browse logs for web search: a survey. acm trans. intell. syst. technol., 4(4), 1–37. https://doi.org/10.1145/2508037.2508038 joachims, t. (2002). optimizing search engines using clickthrough data. proceedings of the eighth acm sigkdd international conference on knowledge discovery and data mining, 133–142. https://doi.org/10.1145/775047.775067 jones, s., cunningham, s. j., mcnab, r., & boddie, s. (2000). a transaction log analysis of a digital library. international journal on digital libraries, 3(2), 152–169. https://doi.org/10.1007/s007999900022 kacprzak, e., koesten, l. m., ibáñez, l.-d., simperl, e., & tennison, j. (2017). a query log analysis of dataset search. in lecture notes in computer science (pp. 429–436). springer international publishing. https://doi.org/10.1007/978-3-319-60131-1_29 https://doi.org/10.29173/iq1087 https://doi.org/10.1002/asi.20180 https://doi.org/10.1002/asi.24646 https://doi.org/10.1197/jamia.m2191 https://www.icpsr.umich.edu/web/icpsr/thesaurus https://doi.org/10.1016/j.ipm.2004.10.007 https://doi.org/10.1145/2508037.2508038 https://doi.org/10.1145/775047.775067 https://doi.org/10.1007/s007999900022 https://doi.org/10.1007/978-3-319-60131-1_29 15/17 lafia, sara, million, a.j., and hemphill, libby (2024) exploratory and directed search strategies at a social science data archive, iassist quarterly 48(1), pp. 1-17. doi: https://doi.org/10.29173/iq1087 kathuria, a., jansen, b. j., hafernik, c., & spink, a. (2010). classifying the user intent of web queries using k‐means clustering. internet research, 20(5), 563–581. https://doi.org/10.1108/10662241011084112 kumar, r., & tomkins, a. (2010). a characterization of online browsing behavior. proceedings of the 19th international conference on world wide web, 561–570. https://doi.org/10.1145/1772690.1772748 lafia, s., million, a. j., & hemphill, l. (2023). direct, orienting, and scenic paths: how users navigate search in a research data archive. proceedings of the acm on human information interaction and retrieval (chiir). levenstein, m. c., & lyle, j. a. (2018). data: sharing is caring. advances in methods and practices in psychological science, 1(1), 95–103. marchionini, g. (1997). information seeking in electronic environments. cambridge university press. marchionini, g. (2006). exploratory search: from finding to understanding. communications of the acm, 49(4), 41. https://doi.org/10.1145/1121949.1121979 meho, l. i., & tibbo, h. r. (2003). modeling the information‐seeking behavior of social scientists: ellis’s study revisited. journal of the american society for information science and technology. https://asistdl.onlinelibrary.wiley.com/doi/abs/10.1002/asi.10244 national research council, division of behavioral and social sciences and education, commission on behavioral and social sciences and education, & committee on national statistics. (1985). sharing research data. national academies press. papenmeier, a., krämer, t., friedrich, t., hienert, d., & kern, d. (2021). genuine information needs of social scientists looking for data. proceedings of the association for information science and technology, 58(1), 292–302. https://doi.org/10.1002/pra2.457 pienta, a. m., akmon, d., noble, j., hoelter, l., & jekielek, s. (2018). a data-driven approach to appraisal and selection at a domain data repository. international journal of digital curation, 12(2). https://doi.org/10.2218/ijdc.v12i2.500 renspie, m., shepard, l., & childress, e. (2015). making archival and special collections more accessible. oclc research. https://doi.org/10.29173/iq1087 https://doi.org/10.1108/10662241011084112 https://doi.org/10.1145/1772690.1772748 https://doi.org/10.1145/1121949.1121979 https://asistdl.onlinelibrary.wiley.com/doi/abs/10.1002/asi.10244 https://doi.org/10.1002/pra2.457 https://doi.org/10.2218/ijdc.v12i2.500 16/17 lafia, sara, million, a.j., and hemphill, libby (2024) exploratory and directed search strategies at a social science data archive, iassist quarterly 48(1), pp. 1-17. doi: https://doi.org/10.29173/iq1087 sharifpour, r., wu, m., & zhang, x. (2022). large-scale analysis of query logs to profile users for dataset search. journal of documentation, 79(1), 66–85. https://doi.org/10.1108/jd-12-2021-0245 silverstein, c., marais, h., henzinger, m., & moricz, m. (1999). analysis of a very large web search engine query log. sigir forum, 33(1), 6–12. https://doi.org/10.1145/331403.331405 solomon, p. (2002). discovering information in context. annual review of information science and technology, 36(1), 229–264. https://doi.org/10.1002/aris.1440360106 sun, g., & khoo, c. s. g. (2017). social science research data curation: issues of reuse. libellarium: journal for the research of writing, books, and cultural heritage institutions, 9(2). https://doi.org/10.15291/libellarium.v9i2.291 taghavi, m., patel, a., schmidt, n., wills, c., & tew, y. (2012). an analysis of web proxy logs with query distribution pattern approach for search engines. computer standards & interfaces, 34(1), 162–170. https://doi.org/10.1016/j.csi.2011.07.001 wang, x., duan, q., & liang, m. (2021). understanding the process of data reuse: an extensive review. journal of the association for information science and technology, 72(9), 1161–1182. https://doi.org/10.1002/asi.24483 white, h. d., lin, x., buzydlowski, j. w., & chen, c. (2004). user-controlled mapping of significant literatures. proceedings of the national academy of sciences, 101(supplement 1), 5297–5302. https://doi.org/10.1073/pnas.0307630100 white, r. w. (2016). exploration, complexity, and discovery. in interactions with search systems (pp. 201–230). cambridge university press. https://doi.org/10.1017/cbo9781139525305.009 white, r. w., & roth, r. a. (2009). exploratory search: beyond the query-response paradigm. synthesis lectures on information concepts retrieval and services, 1(1), 1–98. https://doi.org/10.2200/s00174ed1v01y200901icr003 wilson, m. l., schraefel, m. c., & white, r. w. (2009). evaluating advanced search interfaces using established information-seeking models. journal of the american society for information science and technology, 60(7), 1407–1422. https://doi.org/10.1002/asi.21080 https://doi.org/10.29173/iq1087 https://doi.org/10.1108/jd-12-2021-0245 https://doi.org/10.1145/331403.331405 https://doi.org/10.1002/aris.1440360106 https://doi.org/10.15291/libellarium.v9i2.291 https://doi.org/10.1016/j.csi.2011.07.001 https://doi.org/10.1002/asi.24483 https://doi.org/10.1073/pnas.0307630100 https://doi.org/10.1017/cbo9781139525305.009 https://doi.org/10.2200/s00174ed1v01y200901icr003 https://doi.org/10.1002/asi.21080 17/17 lafia, sara, million, a.j., and hemphill, libby (2024) exploratory and directed search strategies at a social science data archive, iassist quarterly 48(1), pp. 1-17. doi: https://doi.org/10.29173/iq1087 wu, m., psomopoulos, f., khalsa, s. j., & de waard, a. (2019). data discovery paradigms: user requirements and recommendations for data repositories. data science journal, 18. https://doi.org/10.5334/dsj-2019-003 zhang, g., wang, j., liu, j., & pan, y. (2021). relationship between the metadata and relevance criteria of scientific data. data science journal, 20(1), 5. https://doi.org/10.5334/dsj-2021-005 endnotes 1 sara lafia, icpsr, university of michigan 2 a.j. million, icpsr, university of michigan 3 libby hemphill, icpsr, university of michigan and umsi, university of michigan https://doi.org/10.29173/iq1087 https://doi.org/10.5334/dsj-2019-003 https://doi.org/10.5334/dsj-2021-005 iassist quarterly 2010 / 2011 67 iassist quarterly abstract within this paper the situation of qualitative longitudinal (ql) research and archiving systems in lithuania is presented along with an overview of existing research cultural issues and trends. it is highly important to stress that qualitative longitudinal research in this country is experiencing a new prominence as for many years the quantitative tradition was predominant in the scientific community. finally a few statements will be made on possible constructive ways to develop further existing resources and infrastructure for qualitative archiving and for ql research and resources. keywords: qualitative longitudinal research, archiving system, social science data archiving, qualitative resources management, lithuania. introduction fsince the 1990s, there have been numerous highlevel qualitative research studies made by lithuanian scientists, especially in the social sciences of sociology, education and political sociology (e.g., value studies, regional and urban studies, migration issues, attitudes towards careers among school teachers, elite studies, attitudes towards eu integration etc.). however, according to krupavicius and gaidys (2009), the situation in eastern european countries is almost completely different to that of western europe, since social science data archives, as a necessary element of science infrastructure, are almost non-existent or are still in the early phases of development. in many cases in eastern europe empirical data for social sciences are still available through various widely dispersed institutions or through individual contacts2. this difference is certainly not due to the dearth of empirical social science or empirical data sets, but rather because of a lack of legal and institutional arrangements and funding capacities to promote a regular process of social data archiving as well as access to this data by a broad community of social scientists within different countries and from abroad according to clear and transparent rules. lithuania is not an exception in this. moreover, krupavicius and gaidys (2009) go on to note that lithuania is lagging behind such countries as slovenia, the czech republic, hungary, estonia and romania, which already possess the basic structures of social science data archives. this slower development in the area of social science data archives in eastern europe, and for lithuania in particular, needs to be considered in terms of a few major variables in order to achieve change and formulate adequate solutions or policy decisions. there are few official documents signed by government representatives that support a social science data archiving and sharing policy. however, various european projects sustain this whole process and help to further develop its infrastructure. there exploring qualitative longitudinal research and qualitative resources the lithuanian case by jurate butviliene and tomas butvilas1 68 iassist quarterly 2010 / 2011 iassist quarterly is the lithuanian humanities and social science data archive (lida) that was initiated in july 2006 as a two year national level project with objectives for “storage and administration of empirical data and information for lithuanian humanities and social science data”. the project has been supported by the eu european social fund. the project has been implemented by kaunas university of technology policy and public administration institute in partnership with vilnius university, institute for social research, the republic of lithuania ministry of education and science. the project was completed in september 20083. there are some official documents regarding national policies on data archiving and sharing, i.e. resolution no. 1389 dated november 22, 1996 of the government of lithuania “regarding the order of distribution of legal deposit copies of publications and other documents to libraries” (žin, 2006, no. 136-5170)4. this document assures the following activities take place: i) performs control of free legal deposit copy delivery to the national archive of published documents; ii) stores and preserves documents published in lithuania and other documents with national content; and iii) stores and preserves web documents in the archive of electronic resources. less than a decade ago on-line resources were non-existent in lithuanian libraries. lithuanian participation in eifl.net5 opened access to affordable on-line resources for lithuanian researchers, students, and the general public6. expenditure for on-line resources is included in annual statistics and the evaluation of libraries (especially the academic ones) also takes these figures into consideration. as eifl.net grows to meet the evolving challenges of electronic resource acquisition and management, so the lithuanian research library consortium member libraries’ expands its range of activities in partnership. the aim of this paper is to present the situation regarding qualitative longitudinal research and archiving systems in lithuania and also to discuss possible means to sustain and develop this quite new phenomenon in the national and local social science culture. two main methods were chosen: i) scientific literature analysis, evaluation, and interpretation and ii) law acts and normative documents analysis, comparison, and their dissemination in written format. lithuanian qualitative archiving infrastructure and main empirical research institutions as glosiene (n.d)7 states, the landscape of continuing professional development (cpd) for luxembourg income study (lis)8 is rather heterogeneous in lithuania. there are several important institutions in the network. the main state-supported institution for cpd of public librarians as well as for museum cultural center workers is the cultural ministry that offers training courses only for one segment of the lis community – public libraries; academic, school and special libraries are not offered cpd courses. the second institution is the martynas mazvydas national library of lithuania9 (nll). it is a responsibility of the nll to provide both support and training to different libraries in different fields, public ones first of all but also to the school libraries to a certain extent. the center for librarianship which is a part of nll offers lectures and seminars on the actual topics of libraries’ activities and their modernization but they are organized as a part of methodological support activities and do not fit into the concept of cpd precisely. the lithuanian central state archive10 has been a member of the international federation of television archives (fiat/ifta) since 200411 . the institution’s audiovisual holdings consist of film, sound and video recordings as well as photo documents. the division of image and sound is the main repository of cinema heritage in lithuania and holds a total of 7.612 titles: lithuanian chronicles from 1920-1940, chronicles of the second world years, diverse lithuanian newsreels and sketches from the post-war period, lithuanian feature films, documentaries of independent film studios and individual creators, starting from 1991. the lithuanian music libraries network consists of over 150 different libraries possessing music stocks12 . libraries belonging to research library system the national library, vilnius university library, library of lithuanian academy of science and library of lithuanian academy of music and theatre have the most important holdings of music documents. the national library is the leading library in the country and the music department of the national library provides professional assistance to the network of music libraries in order to introduce music specification into the field of library science and activity. the music document stocks are formed by integrating specialized and universal library functions and holdings contain all forms of music documents such as printed music, manuscripts, audio video materials, books, serials etc. the main sources of funding of those institutions are as follows: i) the lithuanian government (i.e. lithuanian ministry of education and science and the state fund of science and studies); ii) eu project funds; iii) private funds (rarely). a detailed map of main empirical research institutions in lithuania is given in appendix 1 of this paper. qualitative longitudinal (ql) data as chenail (1992) emphasizes, much of qualitative research is dominated by research traditions from education, sociology, and anthropology. the researchers from these fields favor such methods as ethnography, participant observation, and naturalistic inquiry. in addition to these popular methods, qualitative research can also include methods from fields such as communication (i.e. discourse analysis or conversation analysis), literature (i.e. narrative analysis or figurative language analysis). valantiejus (2005) adds to this that contrary to the popular definition of qualitative research as the new mode of cognition, we can see qualitative method as having at its essence the concepts of hermeneutics, pragmatism, and radical micro-sociology. thus it becomes more important to analyze ql data collecting and archiving issues and consequently to consider strengthening this position in lithuanian social science data research. development planning or steps to consider although in recent years lithuania has made obvious progress in the development of the information society, much still has to be improved in order to achieve an inclusive information society in which everyone can participate on equal terms13 . among the principal targets there are widespread installation of broadband access, the development of e-content and e-skills, as well as motivating inhabitants to take up new e-services and measures aimed at enabling them to do so. another important area for attention to accelerate the participation of target groups at risk of exclusion. thus existing organisations such as iassist and cessda could certainly strengthen lithuanian present institutions and research centres that mainly concentrate on furthering their substantive and methodological programmes. also support from international agencies and funders would be much of help for the lithuanian department of statistics that collects and disseminates important national information via different channels. according to krupavicius and gaidys’ (2009), the international dimension of the development of the lithuanian social science and humanities data archive (lida), mentioned earlier, is especially important for: i) enabling free access to existing lithuanian and international empirical data and ii) sharing expertise obtained through cooperative agreements with similar academic service organizations abroad. such iassist quarterly 2010 / 2011 69 iassist quarterly international agreements help to attract more material and intellectual investments from all domestic institutions. moreover, the capacity of lida to facilitate access by lithuanian scholars to international empirical data holdings and research development know-how is seen as a way of building confidence and consensus among domestic lithuanian institutions for the process of forming and expanding a national data archive. thus participation of lida in collaborative international projects could help to create a resource base for linking qualitative and quantitative research data, even though there are still differing opinions as to the importance of these two research strategies in the social sciences (bryman, 2008; denzin, 2008 et al.). finally, international cooperation is almost a precondition in order to obtain sufficient funding from national and international sources for the further development of lida as well as for qualitative data archiving. references bryman, a. (2008). social research methods. oxford: oxford university press. chenail, r. (1992). ‘qualitative research: central tendencies and ranges’. the qualitative report. 1 (4). [online]. available at: http:// www.nova.edu/ssss/qr/qr1-4/tendencies.html denzin, n. and lincoln, y.(2008). the landscape of qualitative research. los angeles (ca.): sage krupavicius, a and gaidys, v. (2009). empirical social research in lithuania. [online]. available at: http://www.ceesocialscience.net/ archive/empirical/lithuania/report1.html valantiejus, a. (2005). ‘the rules of radical microsociology’. sociologija: mintis ir veiksmas. 1(15). [online]. available at: http://www.ku.lt/ sociologija/en/issue.php?uid=15 notes 1. jurate butviliene, phd student of sociology lithuanian centre for social researches, lithuania e-mail: jurate.terepaite@gmail.com tomas butvilas, phd associate professor at vilnius university, education department, lithuania e-mail: tomas.butvilas@fsf.vu.lt 2. similar positions were publicly shared during the bremen workshop (2009) from other countries representatives, e.g., poland, finland, switzerland, and czech republic etc. 3. more at: http://www.lidata.eu/en/page.php?page=apie_projekta 4. for further information see: http://www.lnb.lv/lv/bibliotekariem/ konferencu-materiali/20070420/regina_varniene.pdf 5. eifl.net is a not for profit organisation that supports and advocates for the wide availability of electronic resources by library users in transitional and developing countries. its core activities are negotiating affordable subscriptions on a multi-country consortial basis, supporting national library consortia and maintaining a global knowledge sharing and capacity building network in related areas, such as open access publishing, intellectual property rights, open source software for libraries and the creation of institutional repositories of local content [taken from: http://www.eifl.net/cps/sections/ about]. 6. see at: http://www.ifla.org/iv/ifla74/papers/148-banionyte-en.pdf 7. see at: http://tltc.shu.edu/doclib/data/doc/glosiene.1176321015.pdf 8. the luxembourg income study (lis) is a cross-national data archive and a research institute located in luxembourg [taken from: http:// www.lisproject.org/]. 9. more at: http://www.lnb.lt/lnb/selectlanguage.do?language=en 10. more at: http://www.filmarchives-online.eu/partners/ lithuanian-central-state-archive 11. further information at: http://www.filmarchives-online.eu/partners/ lithuanian-central-state-archive/view?set_language=en 12. see at: http://www.iaml.info/activities/lithuania/2006/report 13. see at: https://countryprofiles.wikispaces.com/lithuania 70 iassist quarterly 2010 / 2011 iassist quarterly appendix 1 main empirical research institutions in lithuania institution data collections since/ remarks lithuanian academy of sciences institute of philosophy, sociology, lithuanian academy of sciences late 1960s; sociological and demographic research data lithuanian academy of sciences institute of economics, lithuanian academy of sciences since the 1970s lithuanian academy of sciences institute of social studies a successor institution to the institute of philosophy, sociology, lithuanian academy of sciences; since april 1, 2002 universities kaunas university of technology since the 1970s; sociological and political research universities university of klaipėda since the 1990s; sociological empirical research universities university of law (present mykolas romeris university) since the mid-1990s; sociological and political empirical researchuniversities university of šiauliai since the late 1990s; sociological empirical research and educational sciences universities university of vilnius since the 1970s; sociological and political research universities vytautas magnus university since the 1990s; sociological empirical research statistical offices department of statistics   other government and non-government institutions lithuanian bank  other government and non-government institutions ministry of finance   other government and non-government institutions ministry of health care   other government and non-government institutions central electoral committee since 1992 other government and non-government institutions lithuanian free market institute since 1990 international organizations european bank for reconstruction and development (ebrd), international monetary fund, transparency international, population activities unit of the united nations economic commission for europe, international organization of migration etc. various years private institutions of public opinion and market research baltic surveys since 1992private institutions of public opinion and market research vilmorus since 1993 private institutions of public opinion and market research social information center since 1993 source: krupavicius, a., gaidys, v. (2009). empirical social research in lithuania. interactive: http://www.ceesocialscience.net/archive/empirical/lithuania/report1.html by 22 iassist quarterly winter/spring 2010 by mary tao1 financial crisis data resources: a brief guide the u.s. housing boom was brought to a halt by the subprime mortgage crisis in 2007. as the housing market hit a bad patch, it affected other areas as well. the situation had evolved into the credit crisis by 2008. banks’ balance sheets took a few hits, leading to a liquidity crisis. the united states was not the only country hit hard by events; greece and iceland, among others, saw their share of pain and suffering. ireland is the latest country in the news, with the government stepping in to save two major irish banks. here’s an example of how closely intertwined the fortunes of global companies are. the troubles of an american company, lehman brothers, meant that it was unable to repay the money due to its creditors, including depfa bank, an irish firm. as a result, depfa was unable to pay its own creditors, eventually causing major problems for its parent company, hypo real estate group, based in germany. a lot has been written about the financial crisis and a lot more are yet to come. listed here are just a few resources to use as starting points. banking statistics crsp-frb link • allows researchers to compare companies over time, taking into account mergers and acquisitions. the dataset matches regulatory entity codes to crsp permcos for publicly traded banks and bank holding companies from january 1990 to december 2007. http://www.newyorkfed.org/research/banking_research/ datasets.html assets and liabilities of commercial banks in the united states h.8 • this weekly release provides an estimated aggregate balance sheet for all commercial banks in the united states including u.s. branches and agencies of foreign banks. http://www.federalreserve.gov/releases/h8/about.htm reports of condition and income (call reports) • this database allows one to obtain financial and structural information for most fdic-insured institution and. compare banks with its peer group the data goes as far back as march 31, 2001 and can be downloaded as (pdf), semicolon delimited format (sdf), or extensible business reporting language (xbrl) format. bulk data is also available for all reporters in the form of tab delimited format or (xbrl) format. https://cdr.ffiec.gov/public/ quarterly summary of banking statistics • quarterly synopsis of balance sheet and income statement developments for all u. s. commercial banks. the banks are broken into two categories: those held by the 8 largest domestic bank holding companies (bhcs) jpmorgan chase, bank of america, citigroup, wells fargo, pnc financial services group, us bancorp, bank of new york mellon, and suntrust and all other commercial banks. http://www.newyorkfed.org/research/banking_research/ quarterly_summary.html quarterly summary of banking statisitcis :3rd quarter 2009 23 iassist quarterly winter/spring 2010 fee-based resources include snl financial and bloomberg (wdci writedowns and credit losses versus capital raised ) housing statistics federal housing finance agency (fhfa) • housing market indicators http://www.fhfa.gov/default.aspx?page=66 u.s. bureau of the census • datasets for decennial census housing files 19402000, housing vacancies and homeownerships, housing starts and building permits. http://www.census.gov/cgi-bin/briefroom/briefrm u.s. department of housing and urban development • access to the original datasets including the american housing survey, as well as microdata from research initiatives on topics such as housing discrimination, the hud-insured multifamily housing stock, and the public housing population. http://www.huduser.org/portal/datasets/pdrdatas.html icpsr $ • offers american housing survey datasets, census of population and housing as well as the underlying data cited in research publications. http://www.icpsr.umich.edu/icpsrweb/icpsr/ mortgage statistics mortgage market statistical annual $ • print and cd (xls format). available data include monthly average mortgage rates calculated for different mortgage terms, monthly new home sales inventories, house prices, mortgage markets, total securities issuance volume, and more. u.s. credit conditions: mortgages • contains maps of delinquent mortgages around the country by state and county and uses publicly available data on mortgage delinquencies and foreclosures. http://data.newyorkfed.org/creditconditions/ other useful sources include fee-based sources such as first american corelogic loanperformance, absnet, equifax, haver analytics, and bloomberg. mortgage-backed securities statistics (mbs) fannie mae • monthly reporting data on fannie mae’s mbs http://www.fanniemae.com/mbs/data/index. jhtml?p=mortgage-backed+securities&s=monthly+repor ting+data freddie mac • a variety of data relating to mortgage securities http://www.freddiemac.com/mbs/html/security_data.html inside mortgage finance $ • publisher of mortgage market statistical annual and various newsletters containing the latest data on mbs. credit cards statistics major providers of credit card statistics tend to be feebased sources such as equifax, bloomberg, and snl financial. snl financial $ • information on the major credit card issuers (american express, bank of america, jpmorgan chase, citigroup, discover) ○ credit card delinquency data 30+ day delinquencies; chart; spreadsheet of underlying data federal reserve liquidity facilities the following programs were authorized by the board of governors of the federal reserve system under section 13(3) of the federal reserve act to provide credit and liquidity during a time of financial stress. commercial paper funding facility (cpff) http://www.newyorkfed.org/markets/cpff.html primary dealer credit facility (pdcf) http://www.newyorkfed.org/markets/pdcf.html money market investor funding facility (mmiff) http://www.newyorkfed.org/markets/mmiff.html term asset-backed securities loan facility (talf) http://www.newyorkfed.org/markets/talf.html term securities lending facility (tslf) http://www.newyorkfed.org/markets/tslf.html forms of fed lending chart http://www.newyorkfed.org/markets/forms_of_ fed_lending.pdf other financial data ● credit ratings (of companies and countries) ○ moody's, standard & poor's, fitch $ 24 iassist quarterly winter/spring 2010 ■ moody's corporate default and recovery rates, 1920-2009 ● www.moodys.com/corporate_ default_and_recovery_rates_02_10 ● free registration for some reports; a subscription is needed to acquire most datasets ■ bloomberg, snl financial $ ratt historical ratings trends 2000-2010 • credit risk – data on the financial stability of nations ○ moody's, standard & poor's, fitch, bloomberg $ ○ financial soundness indicators (fsis) imf source that measures the capital adequacy of deposit takers. meaning does a certain country’s financial institutions have enough capital at hand to withstand shocks to their balance sheets? further reading ashcraft, adam, paul goldsmith-pinkham, and james vickery. 2010. “mbs ratings and the mortgage credit boom.” federal reserve bank of new york staff report no. 449. blundell-wignall, adrian, paul atkinson and se hoon lee. 2008. “the current financial crisis: causes and policy issues.” financial market trends, 1-21. gramlich, edward m. 2007. “booms and busts: the case of subprime mortgages.” federal reserve bank of kansas city economic review, fourth quarter:105-113. http://www.kansascityfed.org/publicat/econrev/ pdf/4q07gramlich.pdf jarvis, jonathan. 2008. “the crisis of credit visualized.” retrieved june 1, 2010 from http://vimeo.com/3261363 stackhouse, julie. 2010. the financial crisis: what happened? ( audio version: http://www.stlouisfed.org/ education_resources/awordontheeconomy/player.html text:http://www.stlouisfed.org/education_resources/ awordontheeconomy/presentation.pdf ) notes: 1 mary tao, federal reserve bank of new york, contact: mary.tao@ny.frb.org the views expressed are those of the author and do not necessarily reflect the position of the federal reserve bank of new york or the federal reserve system. vol233 4 iassist quarterly winter 1999 automated preservation of electronic records: a case study of the archival preservation system by fynnette eaton* the national archives and records administration has had a program for accessioning, describing, preserving, and providing reference service to the electronic records (machine readable) records) created by federal agencies and transferred to the national archives for almost thirty years. although there have been many changes in the name of the office, its basic mission has remained the same: to preserve and make available those records created by federal agencies in electronic format that the national archives has determined to have value beyond the short-term need of the originating agency. most people think of the national archives as the keeper of the constitution and the declaration of independence. even the most experienced researchers are largely unaware of the growing number of files in electronic format. since the creation of the center for electronic records in october 1988, the number of files transferred has literally skyrocketed. in 1988 the archives received 150 files from federal agencies. in fiscal year 1991, the number was 1500, ten times as many in three years. the numbers jumped again in 1992 to 8730 files. unfortunately, the center became involved in a resource-draining court case, which forced tom brown and his staff to reduce their efforts in accessioning new files, with a resultant decrease in accessions in fy94 and fy95 to 843 files and 1590 files respectively. nevertheless this is a vast increase compared to earlier years. currently, the center has accessioned about 23,000 files produced by over 100 bureaus, departments, and other components of executive branch agencies and their contractors. these files range from the american soldier surveys of world war ii to records of the 1980 and 1990 decennial censuses. these files include education data illustrating the variety of education programs of the federal government; health and social science data incorporating both biomedical and sociological information and efforts to measure the effectiveness of a variety of social programs; international data including import-export statistics and usia-sponsored surveys. the represented military data ranges from prisoner of war records for world war ii and the korean conflict, and casualty records for the korean and vietnam conflicts, to a large collection of data files resulting from the use of computers for military operations, management, and research dating from the 1960’s especially during combat in southeast asia. clearly, as the size of our holdings grew, with the tremendous increase in transfers of files, the center recognized the need to develop new methods for accessioning and preserving these new files. my paper will discuss the development and implementation of the archival preservation system (which i will refer to as aps). another system, the archival electronic records inspection and control system or aeric was also developed in the early 1990’s to provide automated validation of electronic files and to build a base of descriptive data drawn from the data elements of these files. but that is another paper, by a different author. as i stated at the beginning of this paper, the center has been involved with the various archival activities associated with electronic records for more than twenty years. but it was only in the late 1980’s that the number of files being transferred overwhelmed the staff, requiring reexamination of the methods used to process these files. permit me to give you a brief overview of how the staff used the resources that were available at the time to perform the basic preservation work at the national archives. during the 1970’s all computer processing required by the machine readable archives division was performed at service bureaus. the division had an ibm 029 keypunch, which the staff used to punch cards for the programs to copy or dump tapes. the program card deck was wrapped in a rubber band along with a sheet of instructions to the service bureau indicating which tape volumes were to be used for input and output and any other special instructions. the card decks, instructions, input tapes, and blank output tapes were boxed and sent by courier to the service bureau. the service bureau staffs usually ran the jobs at night with a twenty-four to forty-eight hour turnaround. the division preservation and reference staff checked the jobs, labeled the tapes, assigned the output tapes location numbers and took the tapes to the storage area in the washington national records center in suitland, maryland. iassist quarterly winter 1999 5 during this period the original agency tapes was kept as the master tape and the nars-created tape became the reference copy to be used as input tapes to make copies for researchers. in 1975/1976 the transfer of dod files written in nips (national military command system information processing system) software required special handling. gsa made available to nars a copy of the nips software at the dc share computer facility. generally the staff transported the tapes and card decks to dc share. because of the faster turnaround time, staff began to use dc share to run other jobs as well. in 1981 the machine readable archives division acquired on loan a dumb terminal so that staff could run some jobs for the fbi appraisal task force (another court case that severely strained nars resources) at the national institutes of health (nih). the division took advantage of the access to this computer center by performing some of its normal copy and dump jobs at nih. because of organizational changes within the national archives and records service, the machine readable archives division became the machine readable branch of the special archives division in 1982. plans to procure a minicomputer for its own use did not materialize and the branch lost the terminal with access to the nih computer center. by october 1982 all requests for preservation and reference work for computer files was submitted to another office for processing. although reference work continued, only a few preservation jobs were successfully completed after this transfer. in 1984 the branch acquired a decwriter terminal that the motion pictures branch was surplusing to use for its work on the catalog of holdings. beginning in january 1985, the branch received authorization to use the terminal to access the national institutes of health computer center, as it assumed responsibility for preservation copying of its accessioned files. within a few months the branch obtained additional terminals for submitting jobs to the nih computer center. the computer center at nih served as the computer resources for all of the branch’s requirements. this acknowledgment that the machine readable branch should perform the work is borne out by the few statistics i could locate in the files. during the period fiscal year 82 through 84 less than 100 tapes were copied. the machine readable branch staff copied over 200 tapes in six months of 1985. another report indicates that in fy 87 104 reels had been copied by may. yet these numbers were far too low, once the center began its accelerated accessioning program. the preservation work, which required making two copies of each file offered by a federal agency, was performed using the mainframe computers at the national institutes of health computer center in bethesda, maryland. this mainframe computer center used ibm machines, so the types of outputs that we could produce were limited to the options that were available to us at that center. in addition, the file formats that we could accept for transfer were limited to those formats that we could process at this site. our requirements, which were published in the code of federal regulations, stated that agencies were to transfer permanent computer files in a hardware and software independent format. specifically the files were to be written on half-inch magnetic tape in ebcdic or ascii, without internal control characters on 7 or 9 track open-reel magnetic tape recorded at 800, 1600 or 6250 bytes per inch or on 3480 cartridges and blocked not higher than 32,000 bytes. these requirements clearly reflect the use of a mainframe computer in creation of our preservation copies of these files. even though we had control over the preservation copying of our files, the center was at a disadvantage. we had to relinquish physical control of both the agency tapes and the blank tapes or cartridges that we would use to make the preservation copies on when they were sent to the nih computer center to be mounted on tape drives there, but that was the only real choice we had. we were able to streamline some of our copying procedures, but we often had to wait in line for access to tape drives because we were dependent on modems and phone lines to connect to the nih computer center which had thousands of users. we determined that although the staff time required to copy and compare files created at nih was between 2 1/4 and 4 hours, the time that elapsed from when we prepared the tapes to be transported to the computer center and their return after successful copying, was literally one week. one of the first priorities enunciated by the director, ken thibodeau, when he joined the center in the winter of 1988/89, was to reengineer this process by developing an in-house capability for making preservation copies of electronic files sent to the center by federal agencies. by examining the processes associated with producing preservation copies of electronic files, the staff developed a statement of work for prospective contractors that defined the basic requirements and outlined additional features we would like to see developed over the life of the contract. although we defined the processes based upon our current practice, we also recognized the opportunity to streamline some of the work that had almost developed a life of its own. the development of the statement of work took about a year. the national archives issued a request for procurement in march 1992 and selected as the successful bidder, muller media conversions of new york city in may 1992. although in the statement of work we defined what the processes should be, we did not define how they should be accomplished. 6 iassist quarterly winter 1999 there were four objectives with this contract. first, we wanted to retain control of the media and perform the work in-house, streamlining and saving valuable time and therefore increasing productivity. second, we wanted to be able to handle a wider variety of file formats from agencies and produce standardized output. third, we wanted to capture automatically from the processing of the files the technical attributes of the file, such as the logical record length, blocksize, character code and media. in effect, we wanted to automate the collection of technical description during the processing, rather than entering this information into a separate database (tapes). fourth, we sought to increase the types of media we could accept from agencies and, as well, increase the types of media on which we could output records. muller media began developing the software which was the largest component of the contract. the estimated costs for the archival preservation system (aps) included the cpu unit, a 66 mhz 486 ibm value point running on os2, with two 9 track overland tape drives, 2 overland cartridge tape drives, cables, and bar code apparatus at a cost of $203,165. the software development, $129,500, was more than 63% of the contract cost. it had been our intention to develop a system, operate it for a certain period of time, and if it performed as we expected, to purchase additional systems as money permitted to increase our efficiency in copying files. we had one system, but i had six programmers; so we anticipated purchasing additional systems. however, we did not anticipate purchasing additional systems even before the software was developed, but life overtook plans. on january 19, 1989--the last day of the reagan administration--scott armstrong, among others, filled freedom of information act (foia) requests for information stored on the computer system in the offices of the president from its date of installation in 1985 until the end of the reagan administration. they sued the government, including the national archives, asking for the court to declare many of the materials on the system to be federal and presidential records. until the issues could be resolved, the court ordered the government not to destroy or alter any of the systems’ backup computer tapes since they contained the only extant copies of some of the information on the system. the lawsuit carried on throughout the bush administration. on the day after the election, when bush was defeated, the plaintiffs extended the lawsuit to include materials residing on the computer systems in the bush white house as well. on january 6, 1993 -two weeks before bush formally left office -the court ruled that some materials on the white house computer systems were federal records and presidential records and the court directed the government, specifically, the archivist of the united states, to “take all necessary steps to preserve, without erasure, all electronic federal records generated” by the white house agencies. this court order was not to be taken lightly as events showed. the archives took physical custody of all of the tapes from the white house as the bush administration was leaving and the clinton administration took office. in late may, the court ruled that the archives had not complied with his order to preserve the tapes, found the archives in a state of civil contempt, and levied fines which would amount to $2.5 million if the archives did not come into compliance within thirty days. when the ruling also referred to ”increases in . . . sanctions reserved . . . for any further noncompliance. . .” it meant the threat of possible jail time was real. while the fines were lifted, upon appeal, it was apparent that the threat of fines and possible imprisonment was very real. since the national archives had acquired custody of the computers files and since the records in question were electronic, it was inevitable that the center for electronic records would become involved in this case. this in fact happened in march 1993 when the acting archivist, trudy huskamp peterson, transferred the responsibility for preserving these computer backup tapes from the office of presidential libraries, which had had physical custody of these materials from the time they left the white house until late may, to the center for electronic records within the office of special and regional archives. while the aps had been conceptualized to expedite routine preservation processing, we had to use it first in the very non-routine processing of the materials from the white house. responding to the crisis situation, the aps contractor muller media conversions compiled enough software so that the staff with the requisite clearances were able to make duplicate copies of files in most cases. we encountered problems in copying some of the files and noted any problems. but we complied with the court orders and have successfully avoided legal sanctions. if the center for electronic records had not previously developed the concept of the archival preservation system and did not have the aps on order, nara would have been unable to comply with the court orders. unfortunately the work required by what became known as the armstrong v. executive office of the president case overtook the development of the aps system for the next full year (1994). refinements to the software were made so that problems that we initially encountered were either alleviated or at least better documented by the system. over the next two years, more than 5900 tapes and/or cartridges were copied, using one or more versions of the archival preservation system software. i can state with pride that the center successfully copied more than 99.998% of the media transferred from the white house. out of 5906 items, only twenty nine unique items had a data error. although it had been anticipated that the full development iassist quarterly winter 1999 7 of the aps would take approximately 150 days, the requirements of the court case overshadowed the development of the full system. nonetheless, the staff devoted many hours to developing the data elements for the catalog database, which was the only part of the system that had not been well defined in the contract. basing the database on a preexisting system that was used to track the technical attributes of electronic files, known as tapes, the staff sought to include the essential elements from the tapes database, to capture preservation activities that were previously recorded in a second database (preslog) and to capture information about the individual files as the aps processed the new files. these long staff meetings paid off because the database reflected the needs of the branch in capturing the information necessary for tracking new media, the progress of preservation work, and the technical attributes of files processed on this system. as i said, the contract was not completed as soon as we had anticipated, because we had to make adjustments to the aps system to enable the center to meet the requirements of the court order for preserving the backup files. there were problems that had to be overcome in processing these backup tapes from a variety of computer systems at the white house. in many cases the tapes had been written over, since they were used for weekly backups; so when we attempted to copy the files, there was information beyond the tape mark which meant we did not get a normal end of file mark to cease the copying operation. we were not able to determine what some of these problems were until we obtained greater functionality with the aps system. with the bush cartridges our greatest problem was the fact that the backup utility had also used compression, so the aps system had to be modified so it could make duplicate copies of compressed files. under our normal operating procedures, data in compressed formats do not conform to our transfer requirements. without aps the national archives could not possibly have met the requirements set by the court, duplicating all of the backup tapes, thus ensuring preservation of whatever is found on these backup tapes. but the cost was the delay in full implementation of the aps for the “normal” processing for which this system was designed and purchased. in fact, the center only accepted the system as meeting the basic requirements as outlined in the statement of work in spring 1994. we are still working with a system that clearly is evolving. currently we are moving the catalog database to a network, which will make the information available to the center staff and increase the functionality of the system by being able to use any number of drives for performing copy jobs. unfortunately for us most of the additional development in the aps system up until last september had been in refinements to address additional problems encountered in copying the white house system backup tapes. does that mean that aps is not meeting the needs of the center in making preservation copies of electronic records? absolutely not. we have been processing “normal” files on aps since the fall of 1994, whenever we were not processing files associated with the profs case. we have copied over 1,375 accessioned files using this system. it has greatly increased our ability to handle a wide variety of file formats, which we could not handle previously. one example is the 1990 decennial census public use sample files that the bureau of the census has been transferring to us for more than two years. although they are in a format that is hardware and software independent, user labels and the blocking factor used with these records, prevent us from being able to process these tapes at the nih computer center. perhaps even more importantly, the aps provides the center with a mechanism to accept files on wider variety of media. in using the nih mainframe, we had two choices: records on 9 track tape or 3480-class cartridge. we have both of those options with aps, but we have ordered a cd-rom drive to be installed, so that we can begin to copy scheduled electronic files transferred to us on cd-roms. we also want to have the ability to copy files from diskettes which was never possible previously. and, there are other forms of media which we might want to be able to access, which will be possible, by attaching drives to the system. for the classified profs files, for example, we had to install both a 4mm and 8mm drive, because some of the files were recorded on that type of media. perhaps even more importantly, the center anticipates using the archival preservation system to make copies of files to fill reference requests. we have always used the national institutes of health computer center to make reference copies of electronic files. we have had to limit the choices of output to 9 track tape or 3480 cartridge. but most users now use personal computers or are attached to networks. in many cases researchers do not have access to tape drives. so, one of our goals is to use the aps to make copies of reference requests and to output some files on diskettes when appropriate, and possibly other media such as cd-rom as well. we have just received the funding to purchase the system with a cd-r drive attached to the system. again, it will mean that we will not lose physical control over our records. the tapes and/or cartridges will not be exposed to poor environmental conditions, while they are in transit between our vaults and the external computer center, and we can tailor the output to meet the needs of our customer base. has the archival preservation system been a success? it literally saved us from a contempt ruling in a contentious court case. we are employing it to make preservation copies of files. our current objective is to improve the functionality of the catalog database, by moving it from 8 iassist quarterly winter 1999 a faircom server to oracle on a risc6000 computer and integrating the two systems that are currently used to make copies of accessioned files. and, recognizing the utility of this system, we want to secure another system to make copies of files, and possibly extract of files, to fill reference requests. do i regret that its development was delayed by the court case? yes, but there is a silver lining. the aps system was able to deal with nonconforming media and software dependent files. the success we had in overcoming the problems posed by the backup computer tapes has given us greater confidence in being able to use the aps system beyond the narrow confines for which it was developed. the center will be able to modify this system to meet the demands posed by the newer media being employed by federal agencies. the aps is a viable system for preserving information into the twenty-first century. * fynnette eaton smithsonian, institution archives. [this paper was presented at the may 1996 iassist meeting in minneapolis, mn and reflects the situation at the national archives through 1997. the author left nara in may 1997 to join the smithsonian institution archives, where she is currently employed.] iassist quarterly winter 1999 9 editor's notes 42 2 1/2 rasmussen, karsten boye (2018) editor’s notes: digital curation after digital extraction for data sharing, iassist quarterly 42 (3), pp. 1-2. doi https://doi.org/10.29173/iq944 editor's notes digital curation after digital extraction for data sharing welcome to the third issue of volume 42 of the iassist quarterly (iq 42:3, 2018). the iassist quarterly presents in this issue three papers from geographically widespread countries. we call iassist ‘international’, so i am happy to present papers from three continents in this issue with papers from zimbabwe, italy and canada. the paper 'the state of preparedness for digital curation and preservation: a case study of a developing country academic library' is by phillip ndhlovu, who works as the institutional repository librarian and liaison librarian, and thomas matingwina, who is a lecturer at the department of library and information service at the national university of science and technology (nust) in bulawayo, zimbabwe. modern day libraries have vast amounts of digital content and the authors noted that because these collections require very different management than the traditional paperbased materials, the new materials’ longevity is endangered. their study assessed the state of preparedness of the nust library for digital curation and preservation, including the assessment of awareness, competencies, technology infrastructure, digital disaster preparedness, and challenges to digital curation and preservation. they found a lack of policies, lack of expertise by library staff, and lack of funding. you might conclude that investigating your own organization and reaching the very well known conclusion that 'we need more money!' is not so surprising. however, you have to take note that the jeff rothenberg statement from 1995 that 'digital information lasts forever – or five years, whichever comes first' has not yet sunk in with politicians and administrators, who will immediately associate the term 'digital' with 'saving money'. this study shows them why this is not a valid connotation. it is a study of a single institution, and as the authors note it cannot be generalized even to other academic libraries in zimbabwe. however, other libraries also outside zimbabwe have here a good guide for making their own assessment of the digital preparedness of their institution. the second paper was as was the paper above presented at the iassist conference in 2018 and is also about the transition from media known for thousands of years to new media and digital forms. peter peller presented the paper 'from paper map to geospatial vector layer: demystifying the process'. he is the director of the spatial and numeric data services unit at libraries and cultural resources at the university of calgary in canada. the conversion of raster images of maps to vector data is analogous to ocr technologies extracting words from scanned print documents. thereby the map information becomes more accessible, and usable in geographic information systems (gis). an illustrative example is that historical geospatial information can be overlaid in google earth. the description of the entire process incorporates examples of the various techniques, including different types of editing. furthermore, descriptions of the software used in selected studies are listed in the appendix. it is mentioned that 'paper texture and ink spread' can be responsible for introducing noise and errors, so remember to keep the old maps. this is because what is considered noise in one context might become the subject for https://doi.org/10.29173/iq944 2/2 rasmussen, karsten boye (2018) editor’s notes: digital curation after digital extraction for data sharing, iassist quarterly 42 (3), pp. 1-2. doi https://doi.org/10.29173/iq944 interesting future research. in addition the software for extracting information will most certainly improve. for once both the author and we at iassist quarterly have been quite fast. the data for the third paper was collected in late 2017 and the results are presented here only a year later. in october 2017 a message appeared on the iassist mail list with the start of the sentence 'i would share the data but...' it quickly generated many ways of completing that sentence. flavio bonifacio who works at metis ricerche srl in torino, italy quickly launched a questionnaire sent to members of the mail list and to others from similar communities of interested individuals. the questionnaire was an extension of an earlier one concerning scientists' reuse and sharing of data. the paper includes many tabulations and models showing the background as well as the data sharing attitudes found in the survey. a respondent typology is developed based upon the level of propensity for sharing data and the level of experiencing problems in data sharing into a 2-by-2 table consisting of 'irreducible reluctant', 'reducible reluctant', 'problematic follower', and 'premium follower'. in the nordic countries we tend to have the impression that certain services are publicly available and for free. this impression is plainly superficial because we nordic people also know very well that 'there is no such thing as a free lunch'! all services must be paid for in one way or another. if you have many services that carry no direct cost, it is probably because you and others paid for them beforehand through taxation. because of cuts in the public economy one of the things flavio bonifacio wanted to investigate was the question 'is there a market for selling data-sharing services?' the results imply that 'reducible reluctants' can be a target for services that reduce the problems of that group. submissions of papers for the iassist quarterly are always very welcome. we welcome input from iassist conferences or other conferences and workshops, from local presentations or papers especially written for the iq. when you are preparing such a presentation, give a thought to turning your one-time presentation into a lasting contribution. doing that after the event also gives you the opportunity of improving your work after feedback. we encourage you to login or create an author login to https://www.iassistquarterly.com (our open journal system application). we permit authors 'deep links' into the iq as well as deposition of the paper in your local repository. chairing a conference session with the purpose of aggregating and integrating papers for a special issue iq is also much appreciated as the information reaches many more people than the limited number of session participants and will be readily available on the iassist quarterly website at https://www.iassistquarterly.com. authors are very welcome to take a look at the instructions and layout: https://www.iassistquarterly.com/index.php/iassist/about/submissions authors can also contact me directly via e-mail: kbr@sam.sdu.dk. should you be interested in compiling a special issue for the iq as guest editor(s) i will also be delighted to hear from you. karsten boye rasmussen november 2018 https://doi.org/10.29173/iq944 https://www.iassistquarterly.com/ https://www.iassistquarterly.com/ https://www.iassistquarterly.com/index.php/iassist/about/submissions mailto:kbr@sam.sdu.dk 1/7 kabatangare, tumuhairwe goretti (2021) data literacy integration into development agenda. a catalyst to achieving the sustainable development goals (sdgs), iassist quarterly 45(3-4), pp. 1-7. doi: https://doi.org/10.29173/iq1003 data literacy integration into development agenda. a catalyst to achieving the sustainable development goals (sdgs) tumuhairwe goretti kabatangare 1 abstract the ‘fourth industrial revolution’ (4ir) era characterized by ‘information communication technology’ (ict) based data literacy with respect to research data collection, documentation, preservation, intellectual protection/control and dissemination is a functional catalyst. it enables the realization of the sdgs of the united nations (un) global agenda. the objectives of this study were to identify the role of data literacy in catalyzing the achievement of the sustainable development goals (sdgs); identify challenges faced; and provide recommendations to the challenges faced. the study, employed a desk bound literature review research design, conceptualized that ict based digital data literacy can catalyze an enabling of the realization of the sdgs of the united nations (un) global agenda. according to the literature reviewed, global government ‘ministries, departments and agencies (mdas) are managing voluminous (big) digital data to support strategic decision making, policy implementation and operational optimization towards realizing the sdgs’. this however, requires effective competence (literacy) in digital data analytics to facilitate sdgs based data processing to enable global government mdas to accurately utilize data for policy implementation and decision making towards effectively realizing the sdgs. the study findings recommend a scaling up of digital data literacy and internet infrastructure development as well as power accessibility especially in developing countries among others. keywords data literacy, sdgs, catalyst, 4ir, integration 1.0 introduction 1.1 background scientific research data literacy describes a growing global trend of skills required to support research data services (qin and d’ignazio, 2010; carlson et. al., 2013; schneider, 2013; koltay, 2015). data (information) literacy is often referred to as science data literacy and described as the ability to understand, use and manage data (qin & d’ignazio, 2010). stakeholders pertinent to the global development agenda often rely on the skills of data (information) literate professionals (experts) for their information needs to search for, find, curate, organize, cite, and use data (information). according to zins (2017) with the explosion of data and data-intensive research, the need to preserve and curate data for long-term access, reproduce research and comply with funding agency policies has become a critical component of the research lifecycle. these critical global sustainable development prerequisite needs present to information professionals and any and all pertinent stakeholders, several unique opportunities in fitting traditional old age (now obsolete) information literacy into the new digitally inclined 21st century ‘fourth/4th industrial revolution’ (4ir) era (acrl, 2013). in addition to subsequent inherent challenges in establishment, acquisition and development of new skill/knowledge/competency sets/ models and technical/logistical infrastructural as well as financial support as noted by acrl (2013). taking the initiative to impart africa with new 4ir based skills and knowledge in data (information) literacy management could not have come at a more opportune moment as africa embarks on collecting, organizing, analyzing, monitoring, and presenting data around sdgs and targets as highlighted by the (united nations general assembly, 2015). as data-intensive (big data) research increases around sdgs and targets with required compliance to mandates from funding agencies to make datasets available, https://doi.org/10.29173/iq1003 2/7 kabatangare, tumuhairwe goretti (2021) data literacy integration into development agenda. a catalyst to achieving the sustainable development goals (sdgs), iassist quarterly 45(3-4), pp. 1-7. doi: https://doi.org/10.29173/iq1003 data-driven decision making to inform policy and practice grows; expanding the responsibilities for researchers, information scientists and development policy professionals (experts) to support sdgs’ data management services. didham and ofei-manu (2015) affirmed that data literacy (education) is a critical key to the global integrated framework of sdgs that serves as an important means of implementation for sustainable human development due to the number of positive benefits it brings across the development goals and targets. additionally, quality basic education is a necessary formation for learning throughout life in a complex and rapidly changing world (bokova in unesco, as cited in didham & ofei-manu, 2015:96). quality basic education is about how and what is learned and its influence on personal and collective choices for sustainability (ibid). this allows every human being to acquire knowledge, skills, attitudes and values necessary to shape a sustainable future as pristinely articulated on by mckeown et al. (2002) and nascimen (2016). 1.2 problem statement the ‘4th industrial revolution’ (4ir) based data literacy (education) in today’s 21st century digital world was by majority consensus approved, from the literature reviewed, as key in catalyzing the achievement of the global sustainable development goals (sdgs). but sadly however, the literature reviewed was also telling of many african developing countries’ challenge to harness this potential. a critical problem exists in africa’s lower compliance degree to current data literacy integration mandates into development agenda as mandated by sdgs’ funding agencies relative to the developed world. current scientific sdgs funding agencies’ mandated research data management applications’ utilization, a critical factor in catalyzing the achievement of the global sdgs is especially low on the african continent with many development challenges subsequently poorly understood and managed, a problem that the study attempted to examine. 1.3 aim the aim of the study was to examine the integration of data literacy into the development as a catalyst in achieving the sustainable development goals. 1.4 objectives the objectives of the study were to: i. identify the strategic significant value/role of data literacy (education) in catalyzing the achievement of the sustainable development goals (sdgs) within the 4th industrial revolution (4ir) era. ii. identify the inherent challenges faced. iii. identify the solution-based recommendations to overcome the identified inherent challenges faced. 1.5 research questions i. what is the strategic significant value/role of data literacy (education) (dependent variable) in catalyzing the achievement of the sustainable development goals (sdgs) within the 4th industrial revolution (4ir) era)? ii. what is the inherent challenge faced? iii. what is the solution-based recommendation (remedial action) to overcome the inherent challenge faced? https://doi.org/10.29173/iq1003 3/7 kabatangare, tumuhairwe goretti (2021) data literacy integration into development agenda. a catalyst to achieving the sustainable development goals (sdgs), iassist quarterly 45(3-4), pp. 1-7. doi: https://doi.org/10.29173/iq1003 1.6 significance manda and backhouse (2017) envisioned an onset of 4th industrial revolution (4ir) based data literacy disruptive change to the labor market in their projected increased demand for data literate labor. their key indicator for data literacy catalyzed sustainable development goals (sdgs) achievement of literacy quality and how effectively it was used for development (personal, social, physical, cognitive, moral, psychological and emotional). the study strived to trigger an increased supply of a data literate labor force in africa to increase the continent’s labor market force’s compliance to sdgs funding agencies’ mandated research data management applications’ utilization, a critical factor in catalyzing the achievement of the global sdgs. the study outcome bares significant implications for information professionals, academia, government ministries, departments and agencies (mdas), development policy makers, development partners, the civic community and the labor market navigating the highly dynamic 21st century world of work within the 4th industrial revolution (4ir) era today. 1.7 conceptual framework the study conceptualized ‘data literacy’ (dependent variable) as being a significant catalyst in the ‘achievement of the sustainable development goals (sdgs) in the 4th industrial revolution (4ir)’ (independent variable) and the ‘inherent challenges’ (intervening variable) as an impediment to both the dependent and independent variables. the study further conceptualized that the dependent, independent and intervening variables affected stakeholders pertinent to sustainable development. schematic illustration of conceptual framework 2.0 literature review 2.1 role of data literacy in achieving sdgs the united nations educational, scientific and cultural organization (unesco) (2018) asserts that achieving sustainable development requires a change in the way we think and act. consequently, a transition to sustainable lifestyle, consumption and production patterns which is enhanced through education for all. the 2030 incheon declaration on education recognizes the important role of education as a catalyst in achieving sdgs (incheon declaration, 2015 as cited in unesco, 2018, p.34) adding that education empowers people of all ages to take personal responsibility for creating a sustainable future through data literacy. dependent variable independent variable data literacy as an sdgs achievement catalyst in the 4ir sdgs achievement in the 4ir intervening variable stakeholders african industry, academia, government mdas, policy makers, development partners, labor market inherent challenges faced https://doi.org/10.29173/iq1003 4/7 kabatangare, tumuhairwe goretti (2021) data literacy integration into development agenda. a catalyst to achieving the sustainable development goals (sdgs), iassist quarterly 45(3-4), pp. 1-7. doi: https://doi.org/10.29173/iq1003 libraries, especially academic and research libraries are increasingly taking on the task of supporting their community of users in data collection, analysis, management, and preservation an area of charge that came to be known as research data management (tenopir and et al, 2014). through the sdgs library guide portal a one-stop shop was established for resources on the sdgs, which enable researchers to locate key information easily. the portal has, among other contents, key internet sites, open-access reports and statistics, open education resources, books, journal articles among others (nhamo &malan, 2021). namhila and niskala (2013) noted that access to information was an essential strategy in achieving the sdgs and that information professionals and centers like librarians and libraries were not only key government partners but were already contributing towards achievement of the 17 sdgs by supporting their implementation in regards to providing access to (data) information. the national library of uganda (nlu) provides information communication technology (ict) training to female farmers in weather forecast, crop price and online market access in their local languages, increasing their economic wellbeing (ifla, 2015). in addition, makerere university has been at the forefront in providing health information services to health professionals and local communities in uganda (musoke, 2014). thus, education in any form provides transferable skills relevant to collaborative problem solving. this means that, access to basic education emphasizes socio-emotional skills for achieving positive life outcomes and reducing educational and social disparities that would have hindered sustainable development goals (sdgs) (world education forum, as cited in unesco, 2018). 2.2 challenges data literacy faces the curriculum development change process which passes through different subsequent steps that demand effective and efficient monitoring and control of each process (nevenglosky, 2019). accordingly, curriculum development is not a simple process that can be accomplished without multi-stakeholder collaboration and this negatively affects the implementation of data literacy in achieving sustainable development goals in uganda (sintayehu et al, 2017). the main challenge facing the successful achievement of the united nations agenda 2030 based sdgs in the 4ir 21st century digital world is to manage the data revolution in support of sustainable development; prioritizing the broadening and deepening of production; dissemination and use of statistical data and identification of those population groups which are most vulnerable and making governments more accountable to their citizens (eele, 2015). although ideas on their diverse participation in the curriculum development process that ensure its most effective and efficient execution have emerged in research papers published in scientific journals and presented at conferences (veselin et al., 2014). curriculum development remains a complex and iterative process with a great number of activities involving many stakeholders whose roles are most commonly associated with evaluation of complete activities. 2.3 solutions according to abbott et al. (2008), there is a need to explore the role that the states can play in setting future policy and developing and mandating licensure and certification requirements for all educators on data literacy as a catalyst to achieving sdgs in addressing the following concerns: i. will states require education institutions to offer data literacy programs? ii. will education institutions be held accountable to show evidence of graduate data literacy? iii. how will graduate data literacy be measured? https://doi.org/10.29173/iq1003 5/7 kabatangare, tumuhairwe goretti (2021) data literacy integration into development agenda. a catalyst to achieving the sustainable development goals (sdgs), iassist quarterly 45(3-4), pp. 1-7. doi: https://doi.org/10.29173/iq1003 iv. how will such requirements stimulate change throughout the education system and impact practice on the development of sdgs? therefore, there is a clear need for a scientific and comprehensive survey and inventory of the existence of courses and the extent to which data-driven practices are integrated into existing courses (mandinach et al., 2011). additionally, it is necessary to understand how states are dealing with the issue. researchers and policy implementers need to know what accreditation and licensure requirements around data literacy have been issued by the states, how those requirements are being implemented, and how institutions of higher education are responding. the information from such a comprehensive survey and inventory will provide fundamental data from which policy organizations and education schools can build and respond. currently, there is only anecdotal information that is insufficient and even inappropriate, given the policy and practice emphasis on the importance of data-driven practice and this affects the achievement of the sdgs (amunga, 2011). inclusion and equity in and through education and training are vital to ensuring a transformative education agenda and the right to safe quality education and learning throughout life based on the principles of non-discrimination, gender equality and equal opportunity for all must be ensured to improve on data literacy to achieve the sdgs (banisa, 2015). commitments are needed to support lifelong learning opportunities for all to ensure necessary competencies for personal development, decent work and sustainable development, with attention to climate change, adaptation and mitigation (efa global monitoring report, 2015: 294). education institutions must provide children, youth and adult learners with the competences to be active citizens in democratic and sustainable societies. this includes efforts to promote education for sustainable development and sustainable lifestyles, democracy and human rights, gender equality, age-appropriate comprehensive sexuality education, physical education and sports, education in native language, peace and non-violence (bradley, f. 2016). 3.0 discussion and conclusion given that data literacy is still a major challenge in africa, greater efforts are needed to eradicate illiteracy through formal and non-formal education and training and ensure equitable access to digital literacy, as well as media and information literacy as a continuum of proficiency levels within a lifelong learning perspective. various extant and emergent forms of academia (education), state (public) and industry (private), sector cooperation (collaboration) on data (information) literacy promotion with special emphasis on curriculum development must take center stage moving forward. concrete approaches to support developing countries in the area of specific instruments and requirements provisions for data literacy and data ecosystems capacity building and development should take on a stronger efficient service delivery, transparency, and accountability narrative. the success of the sdgs should not be measured by a head count of cities with better data, but by how many cities have used data to solve the problems their citizens face. data governance and data literacy are indispensable for managing data quality, and thus by their overarching nature, making use of them is a prerequisite of effective and efficient data services as a catalyst in achieving sustainable development goals (association of college and research libraries, 2015). it is important for information professionals to acquire and provide effective data literacy (education) skills and services irrespective of the fact that these competencies extend beyond the knowledge and skills of typical information professionals. paying attention to the sdgs is an important step towards making all audiences accept the information professionals’ mission to provide data literacy skills and services to their full satisfaction as opined by acrl (2000). the success of the 4th industrial https://doi.org/10.29173/iq1003 6/7 kabatangare, tumuhairwe goretti (2021) data literacy integration into development agenda. a catalyst to achieving the sustainable development goals (sdgs), iassist quarterly 45(3-4), pp. 1-7. doi: https://doi.org/10.29173/iq1003 revolution depends on leadership from all sectors working together to leverage the opportunities and address the challenges of the 4th industrial revolution in achieving the sustainable development goals. political leadership, for example, is responsible for developing and implementing an enabling environment for digital transformation and innovation. hence information centers are coping with this new and emerging arena of data services through professional development and re-skilling efforts as a few are able to hire specialized staff as noted by christensen-dalsgaard et al. (2012). references acrl (2013). intersections of scholarly communication and information literacy: creating strategic collaborations for a changing academic environment. working group on intersections of scholarly communication and information literacy. chicago, il: association of college and research libraries. available online at: http://acrl.ala.org/intersections/ acrl (2000). information literacy competency standards for higher education. chicago, il: association of college and research libraries. available at: http://www.ala.org/ala/mgrps/divs/acrl/standards/standards.pdf (accessed 2 may 2016). association of college and research libraries (2015). framework for information literacy for higher education. chicago, il: association of college and research libraries. available online at http://www.ala.org/acrl/standards/informationliteracycompetency abbott, d. v. (2008). a functionality framework for educational organizations: achieving accountability at scale. in e. b. mandinach & m. honey (eds.), data-driven school improvement: linking data and learning (pp. 257–276). new york, ny: teachers college press. amunga, h. a. (2011). information literacy in the 21st century universities: the kenyan experience. banisa, d. (2015). access to information central to the post-2015 development agenda. international federation of library associations & institutions (ifla). boeren, ellen (2019). "understanding sustainable development goal (sdg) 4 on “quality education” from micro, meso and macro perspectives." international review of education 65.2 (2019): 277-294. bradley, f. (2016). a world with universal literacy: the role of libraries and access to information in the un 2030 agenda. ifla journal, 42, 118-125. christensen-dalsgaard, b., van den berg, m., grim, r., jansen, d., pollard, t. & annikki r. (2012). ten recommendations for libraries to get started with research data management: final report of the liber working group on e-science/research data management. available online at https://libereurope.eu/wp-content/uploads/2020/11/the-research-data-group-2012-v7final.pdf easton, j. q. (2009, july). using data systems to drive school improvement. keynote address at the stats-dc 2009 conference, bethesda, md. easton, j. q. (2010, july). helping states and districts swim in an ocean of new data. keynote address at educate with data: stats-dc 2010, bethesda, md. eele, g. (2015). building statistical capacity the challenges, paris21 discussion paper no. 7, july 2015, www.paris21.org/sites/default/files/paris21-discussionpaper7-statisticalcapacity.pdf efa global monitoring report. (2011). education counts: towards the millennium development goals. paris. retrieved from http://unesdoc.unesco.org/images/0019/001902/190214e.pdf heritage, m., & yeagley, r. (2005). data use and school improvement: challenges and prospects. in j. l. herman & e. haertel (eds.), uses and misuses of data for educational accountability and improvement: 104th yearbook of the national society for the study of education, part 2 (pp. 320–339). malden, ma: blackwell. https://doi.org/10.29173/iq1003 http://acrl.ala.org/intersections/ http://www.ala.org/ala/mgrps/divs/acrl/standards/standards.pdf http://www.ala.org/acrl/standards/informationliteracycompetency www.paris21.org/sites/default/files/paris21-discussionpaper7-statisticalcapacity.pdf http://unesdoc.unesco.org/images/0019/001902/190214e.pdf 7/7 kabatangare, tumuhairwe goretti (2021) data literacy integration into development agenda. a catalyst to achieving the sustainable development goals (sdgs), iassist quarterly 45(3-4), pp. 1-7. doi: https://doi.org/10.29173/iq1003 nhamo g., malan m. (2021) role of libraries in promoting the sdgs: a focus on the university of south africa. in: nhamo g., togo m., dube k. (eds) sustainable development goals for society vol. 1. sustainable development goals series. springer, cham. https://doi.org/10.1007/978-3-030-70948 manda, m.i. & backhouse, j. (2017). digital transformation for inclusive growth in south africa. challenges and opportunities in the 4th industrial revolution. 2nd african conference on information science and technology, cape town, south africa. tenopir, c., sandusky, r. j., allard, s., & birch, b. (2014). research data management services in academic research libraries and perceptions of librarians. library and information science research, 36(2), 84-90. united nations. general assembly (2015, october 21). transforming our world: the 2030. agenda for sustainable development. retrieved from http://www.un.org/en/ga/search/view_doc.asp?symbol=a/res/70/1&lang=e zins, m. (2007). conceptual approaches for defining data, information, and knowledge. journal of the american society for information science and technology, 58(4), 479–493. endnotes 1 tumuhairwe goretti kabatangare (mrs.) is a librarian i at kabale university of southwestern uganda and can be reached by email: tgkabatangare@kab.ac.ug or gkabatangare@gmail.com https://doi.org/10.29173/iq1003 https://doi.org/10.1007/978-3-030-70948 http://www.un.org/en/ga/search/view_doc.asp?symbol=a/res/70/1&lang=e mailto:tgkabatangare@kab.ac.ug gkabatangare@gmail.com ;sist newsletter vol.1, no. 1 editor's note this is the first issue of the lassist newsletter . it represents an international cooperative effort on the part of individuals managing, operating or utilizing machine-readable data archives, data libraries and data services. the newsletter has as its primary intent the dissemination and exchange of information on significant developments in these information centers. four times a year this newsletter will report on activities related to the production, acquisition, preservation, processing, distribution, and utilization of machine-readable data carried out by its members and others in the international social science community. your contributions and suggestions for topics of interest are encouraged and welcomed. in the end, the success of the newsletter as a communications mechanism depends on you. i chairperson's report a series of lassist meetings were held on august 16-20, 1976, in conjunction with the international political science association world congress in edinburgh, scotland. these meetings were held to finalize the constitution, approve the operational structure of the association, review the progress of the regional sectors, project future activities, and formally establish lassist. the steering comnittee members present were per nielsen, alice robbin, judith rowe, joseph bonmariage, cees middendorp, guido martinotti, bjjjrn henrichsen, thomas madron, ekkehard mochmann, and carolyn geda. a draft constitution was reviewed and rewritten based on lassist experiences of the past year and a half. the individuals involved with writing the constitution were allen barton, columbia university, united states; joseph bonmariage, university of louvain, belgium; carolyn geda, university of michigan, united states; bjjsrn henrichsen, university of bergen, norway; thomas madron, western kentucky university, united states; guido martinotti, university of milan, italy; ivor crewe, university of essex, england; cees middendorp, steinmetzarchives, netherlands; ekkehard mochmann, university of cologne. federal republic of germany; per nielsen, danish data archives, denmark; alice robbin, university of wisconsin, united states; judith rowe, princeton university, united states; and, marcia taylor, university of essex, england. we recognize that in a dynamically growing organization the constitution will be subject to change. your comments on the constitution will be welcome. membership dues were established on the basis of the information provided by the various regional representatives. although the association has been created on the basis of individual membership, the dues structure has been developed to accomodate the need for institutional affiliation. complete fee information is provided on the last page of this newsletter . sist newsletter vol.1, no. 1 an lassist newsletter will be produced quarterly. it will be the official vehicle for communication and will systematically report on the activities of the association. we would like to take this opportunity to urge you to make contributions to the newsletter as well as suggestions of topics that you would like to see addressed. this first issue is being edited by alice robbin who has graciously accepted the responsibility to compile the existing reports on lassist activities. a list of steering committee members is included on the cover of this newsletter . changes occur as previous members move to different institutions or countries and as the association expands to warrant the creation of new regions of affiliation, thereby necessitating the appointment of a regional secretariat. the regional secretariats have made a great deal of progress the past year. the united states, canadian and european secretariats have produced several mailings to the membership. as a result. action group chairpersons have been designated for each of these regions. the canadian secretariat activities are now handled by sharon chappie of the data clearing house for the social sciences. the asia-african secretariat has also changed. naresh nijhawan of the indian council of social science research will assume this responsibility. activities in this region have been delayed temporarily, but will resume when mr. nijhawan returns to the council at the end of november after a leave of absence. krzysztof ostrowski of the institute of philosophy and sociology in warsaw has volunteered to act as the east european secretariat. reports from the united states, canada, and west european secretariats are included in this newsletter . information on latin american, east european, and asia-african activities will appear in future newsletters . the mandate and definition of the action groups were revised somewhat at the edinburgh meetings. these revisions were made in response to the need for the united states and west european action groups to focus on slightly different problem areas. these changes are fully captured by the descriptions provided in the action group section of the newsletter . papers were presented at an lassist panel chaired by ivor crewe at the ipsa world congress on the theme of the contribution of data archives to the empirical study of public policy. carolyn geda acted as discussant. copies of the papers are available from the authors. further activities of lassist include a north american working conference during february 1977, an international meeting in ann arbor, michigan during august 1977, and an international meeting in uppsala, sweden in conjunction with the international sociological association world congress in august 1978. 4. flexible ddi storage 1/12 hopt, oliver; klas, claus-peter; mühlbauer, alexander (2018) flexible ddi storage, iassist quarterly 42 (2), pp. 1-12. doi: https://doi.org/10.29173/iq923 flexible ddi storage oliver hopt, claus-peter klas, alexander mühlbauer1 abstract the current usage of ddi is heterogeneous. it varies over different versions of ddi, different grouping, and unequal interpretation of elements. therefore, provider of services based on ddi implement complex database models for each developed application, resulting in high costs and application specific and non-reusable models. this paper shows a way to model the binding of ddi to applications in a way that it works independent of most version changes and interpretative differences in a standard like ddi without continuous reimplementation. based on our ddi-flatdb approach, shown first at eddi 2015 & 2016, we present a complete implementation along the use case of a web-based questionnaire editor including sustainable solution for version management and efficient handling of large ddi structures. the user interface is adaptable to different usage scenarios to come. the application supports ddi-lifecycle from the first question draft and hands over structured (meta-) data to survey institutes and data archives and supports the collaborative questionnaire development for the pre-election cross-section of the german longitudinal election study from early 2017 on. keywords questionnaire development, storage, compatibility, ddi, software architecture introduction ddi is the best known and widely used standard for describing the social sciences data lifecycle from the conceptual start of a “study” to the support of reusing data found in data archives. but the definition of a standard format (together with the implicit data model underneath) is not enough to actually support all the steps in this lifecycle. there is also an enormous need for software systems to gather the metadata to be filled into the standard and to provide it to the scientific and general community. making use of ddi in software systems is often a large investment on the implementation side and also on the standardization side. for several software architectural approaches this investment will last until the next version of the standard, until it has to be renewed. this paper presents an storage architecture to work independent of most version changes and interpretative differences in a standard like ddi without continuous reimplementation. problem description the expressive strength of ddi on one hand is a disadvantage on the technical side when it comes to implementing software with and against ddi. the enormous complexity is a matter of cost within classical development, where each class to be supported becomes also a class within the data model. even those bits from the standard that just bring structure and understandability to the metadata, but have no software-technical necessity have to be implemented. https://doi.org/10.29173/iq923 2/12 hopt, oliver; klas, claus-peter; mühlbauer, alexander (2018) flexible ddi storage, iassist quarterly 42 (2), pp. 1-12. doi: https://doi.org/10.29173/iq923 concerning the evolution of ddi, we see a constant extension of complexity together with the extension of expressiveness. but the underlying principles of data modeling change on a large scale especially between ddi codebook and ddi lifecycle. a good example is the documentation of questions, which are placed as simple <qstntxt> within variables (<var>) in ddi 2.1. in 3.0 they moved into an own package with widely increased details as <questionitem> or <multiplequestionitem>. and even within the different minor version of ddi lifecycle, the documentation of questions is not 100% compatible. <multiplequestionitem> in 3.0 and 3.1 became <questiongrid> and items within those were inline <questionitem> and became a set of <category> within one <griddimension>. a more detailed discussion on the technical issues concerning meeting the complexity of and diversity within ddi could be found in [amin et al. 2011]. to make it even harder for developers, the standard is not always interpreted in the same way by all of its users due to its semantic complexity. this is even the case within institutions using ddi, where not all people have the same understanding of how to use certain fields. in the following sections of the paper we will first describe our way of storing ddi xml files across versions and across different interpretations which we call dialects. second, we will describe a portal application to develop and manage questionnaires in, both from the usage perspective and some technical details. finally, we give a summary and outlook on the next steps for our portal and extended use cases for the ddi-flatdb. making use of ddi-flatdb the main idea behind the ddi-flatdb architecture is to store the ddi xml files not in a database, but to store them on the hard disk and keep them managed in a version control system, such as git. this will eliminate the necessity to develop an entity relationship model for each ddi dialect, ddi version and custom ddi semantics. the file is stored as is. in addition, the version control system provides basic revision and versioning of the ddi files. in order to enable efficient access to the elements within the ddi files, a customizable split mechanism is provided, which splits each xml file into identifiable xml elements, according to the functional requirements of the application. these elements are then stored in a simple database, currently with one table, which holds the main keys and ids of the element along with its xml content. this database is intended as “proxy only” to provide efficient read and write access and should be configured strictly to functional and application specific requirements. the xml elements within the database, respective in the file storage, are accessible via a restful api, depicted in figure 1: ddi-flatdb implementation. https://doi.org/10.29173/iq923 3/12 hopt, oliver; klas, claus-peter; mühlbauer, alexander (2018) flexible ddi storage, iassist quarterly 42 (2), pp. 1-12. doi: https://doi.org/10.29173/iq923 figure 1: ddi-flatdb implementation the main functionalities provided by the api are ● upload/import and split an ddi xml file with its attached split configuration, ● get element lists (e.g. all studyunits), ● get, set and delete elements. upload and split ddi xml document the upload, split and versioning process is implemented as depicted in figure 2: upload, split and versioning process. the initial step is to upload the ddi xml document incl. the application specific split configuration via rest post. the original document is stored within the flatdb incl. automatic revisioning. the next step is to apply the split configuration, inserting the elements from the split xml document into the flatdb. now it is possible to use the elements within an application. through a scheduled poll or event driven service the document is finally stored in the file system and version controlled via git. https://doi.org/10.29173/iq923 4/12 hopt, oliver; klas, claus-peter; mühlbauer, alexander (2018) flexible ddi storage, iassist quarterly 42 (2), pp. 1-12. doi: https://doi.org/10.29173/iq923 figure 2: upload, split and versioning process example for a simple split configuration in the following a small example for a split configuration is given to split the “studyunit” into the elements “archive” and “othermaterial”, given a ddi3.1 xml document. the main entity in the file is to describe the split paths via xpath expressions. this line also defines the type of the element. the first line in the split configuration gives each element “archive” the type “ddiinstance.studyunit.archive” and is split at “/ddiinstance/studyunit/archive”. in addition the identifier path gives xpath location relative to the elements xpath. the split configuration is used to split and integrate the xml, so the order of the “paths” is important to recreate a well-formed and validated xml file. access elements in ddi-flatdb the api provides efficient access to lists of ddi element types (e.g. all studyunits) or single elements (e.g. one questionconstruct within a questionnaire) for read, write or delete actions. in addition all revisions of a document or element can be loaded. fetch lists of elements: <host>/getelementlist/<type> fetch single element: <host>/getelement/<elementid> fetch revision of element: <host>/getelementrevisions/<elementid>/<since> ddii ddiinstance.studyunit.archive.path = /ddiinstance/studyunit/archive ddiinstance.studyunit.archive.identifierpath = ./@id ddiinstance.studyunit.othermaterial.path = /ddiinstance/studyunit/othermaterial ddiinstance.studyunit.othermaterial.identifierpath = ./@id ddiinstance.studyunit.path = /ddiinstance/studyunit ddiinstance.studyunit.identifierpath = ./@id ddiinstance.path = /ddiinstance ddiinstance.identifierpath = ./@id nstance.studyunit.archive.path = /ddiinstance/studyunit/archive ddiinstance.studyunit.archive.identifierpath = ./@id ddiinstance.studyunit.othermaterial.path = /ddiinstance/studyunit/othermaterial ddiinstance.studyunit.othermaterial.identifierpath = ./@id ddiinstance.studyunit.path = /ddiinstance/studyunit ddiinstance.studyunit.identifierpath = ./@id figure 3: example split configuration https://doi.org/10.29173/iq923 5/12 hopt, oliver; klas, claus-peter; mühlbauer, alexander (2018) flexible ddi storage, iassist quarterly 42 (2), pp. 1-12. doi: https://doi.org/10.29173/iq923 the <elementid> corresponds to the actual identifier within the ddi file. the <type> of the element is derived from the split configuration. connecting the ddi-flatdb to the user interface implementation as described before, the split configuration defines the elements to be accessible from the ddiflatdb. in a user interface the mapping from xml elements to java objects is needed. as we wanted to be flexible and again follow the functional dependencies, we decided to implement the necessary objects as simple java beans. to be independent of the ddi dialect, we provide for each element a configuration, how and where to read and write the values. the following code snippet gives a simple example, how this is done for the property label from the entity questionitem: the path to the label value is given by an xpath expression. as the label can be given by multiple languages, the phrase %lang% is replaced by the current content language set in the application context. in addition the key “ldep” (language dependent) is set for processing. thus ddi-flatdb can handle different language versions. complex entities, such as multiple publications with own attributes can also be instantiated, but need a more complex configuration. within the java bean, the configuration of “questionitem.label” fills the values of java object attribute label, using java reflection and the setter and getter pattern. in the next section we describe our use case project, the questionnaire and study editor, depending heavily on the ddi-flatdb approach to also be used in the future, when new ddi versions come up. in the best case, we only need to adapt the above described configuration files, instead of remodeling the database. questionnaire and study editor following one of the main goals, implementations around the ddi-flatdb are driven by their use cases, the first step towards the questionnaire and study editor was to collect requirements from the actual users of this tool. initially it was projected to support the german longitudinal election study (gles) que questionitem.label.ldep=true questionitem.label.multi=false questionitem.label.complex=false gesisquestionnaire32.questionitem.label.path=/d:questionitem/d:questionitemname /r:string[@xml:lang='%lang%'] stionitem.label.ldep=true questionitem.label.multi=false questionitem.label.complex=false gesisquestionnaire32.questionitem.label.path=/d:questionitem/d:questionitemname /r:string[@xml:lang='%lang%'] figure 4: property file for entity parsing https://doi.org/10.29173/iq923 6/12 hopt, oliver; klas, claus-peter; mühlbauer, alexander (2018) flexible ddi storage, iassist quarterly 42 (2), pp. 1-12. doi: https://doi.org/10.29173/iq923 with management facilities for their questionnaires for the german federal election in 2017. the resulting workflow, depicted in figure 5: questionnaire workflow looks very similar to the general process of creating questionnaires [blumenberg et al. 2015]. figure 5: questionnaire workflow caused by the need for a new study (1) the researchers want to either start with a complete or partial copy of a questionnaire from a previous wave (2) or with a blank questionnaire to be filled with new questions (3). now, the actual collaborative and iterative process starts to fine tune the instrument (4), including editing, discussing and to rate existing questions as well as the creation of new questions both from scratch and as import from the collection of questions already existing inside the question database. after finalizing the definition of the new questionnaire (5), it should no longer be edited by the researchers in general. only a study administrator is then able to do error corrections on the existing content. this study administrator will also perform a delivery to a survey institute as a set of possible exports (6). these exports include a customizable questionnaire document and an unfilled file of any statistical software including the defined variables. if survey institutes will be able to process ddi in the future, the application will be able to hand over such files instantly as it is directly relying on ddi https://doi.org/10.29173/iq923 7/12 hopt, oliver; klas, claus-peter; mühlbauer, alexander (2018) flexible ddi storage, iassist quarterly 42 (2), pp. 1-12. doi: https://doi.org/10.29173/iq923 files. along with the definition of the questionnaires, we included a study editor to document the study itself. the numbering of these steps (1-6) will reappear later in this paper to associate these steps with the actual implemented functionality. the current mapping into ddi is based on version 3.2. this includes both a split configuration on the level of the ddi-flatdb and a mapping onto the used java entities. these entities and the resulting user interfaces are modeled along the data model behind ddi lifecycle. the consequence of this approach is that some splits may not work in the same manner in other versions of ddi. the used structure for items within question grids for example cannot be mapped even into ddi 3.1 without using the same xml snippet to fill several different entities. to avoid dirty reads, legacy data could be viewed within the described portal but not edited. for reuse in new questionnaires the entity classes could be used to convert copies to ddi 3.2. the portal currently contains 13 individual studies with 30 questionnaires containing about 4000 questions only for the german election study alone. portal components the current editor application contains only functionality that needs authorization. therefore, every visit of the portal will lead first to a login form. after authentication, the user is directed to a questionnaire/study selection view (see figure 6), containing all studies with respect to his/her group affiliation and according to the access rights, a set of management functions. these consist of creating a new study (3) or clone an existing study (2) and create a new questionnaire (3) or clone an existing questionnaire (2). the distinction between studies and questionnaire is necessary due to the requirement within gles to have a series of questionnaires included in one data set. after choosing a questionnaire or study, the top menu can be used to either edit the study level documentation or the questionnaire content (4). https://doi.org/10.29173/iq923 8/12 hopt, oliver; klas, claus-peter; mühlbauer, alexander (2018) flexible ddi storage, iassist quarterly 42 (2), pp. 1-12. doi: https://doi.org/10.29173/iq923 figure 6:study and questionnaire selection view the portal offers two slightly different views for editing questions. the first one offers a list of all questions contained in a questionnaire sortable by different criteria. one of these consists of personal tags to be associated with questions to mark them as of individual interest or with importance (rating). it is also possible to filter within each of these criteria. the second view, show in figure 7, offers a hierarchical list of logical blocks (ddi:sequences) and questions in the straight forward flow of the questionnaire. within this view, it is not possible to sort, because the order of the overview table is used as the arranging device of the questionnaire. researchers can drag and drop or move single questions or entire blocks of questions through the questionnaire flow, so that the overall questionnaire can be rearranged (4). both views contain functionalities for creating and editing questions (4) and for exporting certain questionnaire elements or even the entire questionnaire (6). the tree-table form for logical blocks (sequences) is covering not much more than label and description to be stored within <sequence>. the <controlconstructreference> elements included are managed by the tree view on the left side of the main screen. the center of this main view always contains a preview of the selected element. in case of a question grid (or item battery), the preview will give an impression, how the element would look like in the final questionnaire. simple question look similar but the preview of a sequence just consists of its definition part and a list of all contained elements. https://doi.org/10.29173/iq923 9/12 hopt, oliver; klas, claus-peter; mühlbauer, alexander (2018) flexible ddi storage, iassist quarterly 42 (2), pp. 1-12. doi: https://doi.org/10.29173/iq923 figure 7: question view the editing of questions is described along the form for question grids (or item batteries) as it is more or less an extension of the field set available for simple question items (see figure 8). concerning the mapping into ddi 3, this form consists of several parts. ● the first part contains the essential question elements like name, title and question text (a).this information is stored within <questiongrid> (or <questionitem>). ● the items of the grid (b) are defined individually for every question grid. as a result of the user requirements, no reuse of item sets is wanted. this results in a pair of <codelist> and <categoryscheme> referenced from one <griddimension>. ● incoming filter definitions (c) are stored within universe. as this is referenced from the <questionconstruct> that maps the grid (or question) into the questionnaire flow, reuse of these definitions is possible. the user requirements contained the request to not manage this information independently but to integrate it into individual questions. ● answer categories, concepts and interviewer instructions (d) are also referenced from <questiongrid> (<questionitem>) and <questionconstruct>. they are accessed on right side of the form (e) both for selecting the used instances and for managing/editing. ○ answer categories are stored as pairs of <codelist> and <categoryscheme> each. ○ interviewer instructions go straight forward into <interviewinstruction>. ○ concepts are managed as <concept> within a <conceptscheme> that is placed outside of the questionnaire file, so that the concepts are consistent throughout the entire gles study series. ● variables (f) can also be edited, but as this portal targets to a state before an actual data set is produced/survey is conducted, the variables are only defined by name and label, e.g. for hand over to the survey institution in an statistical file. https://doi.org/10.29173/iq923 10/12 hopt, oliver; klas, claus-peter; mühlbauer, alexander (2018) flexible ddi storage, iassist quarterly 42 (2), pp. 1-12. doi: https://doi.org/10.29173/iq923 figure 8: question edit view “external” components following the idea of micro services, functionality that covers information surrounding the ddi structure is stored in separate services. our portal currently makes use of services for: 1. user authentication to support all tools within the same tool chain with a consistent user management, this is externalized. 2. tagging of ddi elements for several reasons, private and public tagging can support the productivity within large (meta-) data collections. we developed a service that associates tags with identifiers so that the tags could always be linked to the resources behind the identifiers. 3. export to word this service will evolve to a general export mechanism based on word templates to be filled with the content from ddi elements. currently it supports questionnaires with a fixed template. summary and outlook the application of the ddi-flatdb architecture described in this paper is not only implemented with the questionnaire editor, but also is used within our automatic codebook report generation service and applied to the study editor (see figure 9 blue boxes). in addition we will adopt further editor tools, like the variable documentation. and besides the editor tools, we plan to provide search and browse portals on the documented studies. https://doi.org/10.29173/iq923 11/12 hopt, oliver; klas, claus-peter; mühlbauer, alexander (2018) flexible ddi storage, iassist quarterly 42 (2), pp. 1-12. doi: https://doi.org/10.29173/iq923 figure 9: overview of services based on ddi-flatdb we have achieved up to this point a complete application from storage to a user interface fully based on ddi (within one year development time), with the following features: ● collaborative and web-based, ● not fixed to a specific version or dialects of ddi, ● includes revision and version handling, ● includes user management and access rights. for the future versions of the ddi-flatdb an extension is planned to explicitly instantiate the links within ddi, e.g. external links between studies and study waves as well as internal links between questions and variables to be published as linked open data. this extension will provide faster access to entities and provide search and browse functionality. for the questionnaire editor, we will go live in april and will extend the functionality further to meet the user requirements, like more collaborative functions such as a discussion platform. the ddi-flatdb will be published as open source. https://doi.org/10.29173/iq923 12/12 hopt, oliver; klas, claus-peter; mühlbauer, alexander (2018) flexible ddi storage, iassist quarterly 42 (2), pp. 1-12. doi: https://doi.org/10.29173/iq923 references amin, alerk, barkow, ingo, kramer, stefan, schiller, david, williams, jeremy. 2011. “representing and utilizing ddi in relational databases”. ddi working paper series (other topics) ddi alliance. http://dx.doi.org/10.3886/ddiothertopics02 blumenberg, manuela s., wolfgang zenk-möltgen and claus-peter klas. 2015. "implementing ddi-lifecycle for data collection within the german gles project." eddi15 – 7th annual european ddi user conference, kopenhagen. klas, claus-peter, oliver hopt, wolfgang zenk-möltgen and alexander mühlbauer. 2015. "ddiflat-db – a lightweight framework for heterogeneous ddi sources" eddi16 – 8th annual european ddi user conference, cologne. klas, claus-peter, oliver hopt, wolfgang zenk-möltgen and alexander mühlbauer. 2015. "ddi flatdb: efficient access to ddi." eddi15 – 7th annual european ddi user conference, kopenhagen. 1 oliver hopt, claus-peter klas, alexander mühlbauer are all affiliated with gesis leibniz-institute for the social sciences in germany. oliver hopt are contact author at oliver.hopt@gesis.org. https://doi.org/10.29173/iq923 http://dx.doi.org/10.3886/ddiothertopics02 mailto:oliver.hopt@gesis.org 6 iassist quarterly summer 2007 linda f. powell and andrew boettcher*1 modernizing financial data collection with xbrl abstract in 2003, three u.s. banking regulatory agencies combined resources to revolutionize the collection, editing, storage, and dissemination of commercial bank reports of income and condition. the regulatory agencies relied heavily on web-based technology and the extensible business reporting language (xbrl) transmission protocol. this paper will review the creation of an interagency data collection and dissemination facility. it will focus on the business problem that needed to be solved, the evolution of the technology that enabled the project, what xbrl is, and why was it selected as the transmission protocol. the paper will also review the challenges and benefits associated with using a standard transmission protocol versus creating a customized xml transmission facility. introduction in 2003, u. s. bank regulators, the federal deposit insurance corporation (fdic), board of governors of the federal reserve system (frs), and office of the comptroller of the currency (occ) combined resources to revolutionize the collection, editing, storage, and dissemination of the commercial bank report of income and condition (call report). the system development process took approximately two years and was completed in march 2005. the new process utilizes internet technology and xbrl protocols. the benefits of the new process include easier distribution of reporting requirements, fewer revisions to reported data, and easier dissemination of data to the market. the call report is a regulator-specified report which collects primarily financial statement data. this means that the data collected are standardized across the industry and predefined by the banking regulators. the financial data conform to u.s. generally accepted accounting principles. however, the data are not synonymous with data that are collected from other market regulators such as the securities and exchange commission (sec). many of the data variables collected are comparable to those collected by other regulators or published in the financial statements, but additional variables are uniquely defined for and used specifically by the banking regulators. the call report is filed by all u.s. insured commercial banks and state chartered savings banks, which currently includes over 7,700 institutions. each bank is required to file a quarterly report, which contains over 2,000 variables, to its regulator. although all banks must file a call report not all banks are required to file the exact same set of data. for example, banks with foreign operations are required to provide details about their foreign activities that are not applicable to banks that operate solely in the domestic realm. the criteria for which variables a bank must report are detailed in the instructions. historical data collection model the collection of call report data has evolved over the years from a manual collection process to the current internet submission process. figure 1 shows the regulators’ data collection, editing, and distribution model in use before the 2003 call report modernization project. the historical model required all banks to file their call report electronically but involved multiple regulators maintaining multiple copies of the call report data and, individually contacting banks regarding any questions or modifications to that data. the historical model included the following data flows: 1. reporting form changes are communicated to banks and software vendors via pdf and word files. 2. each regulator and bank modifies its internal applications to receive or send data. 3. banks compile the required data and send a standardized formatted file to a central collection vendor. 4. the collection vendor compiles the banks’ filings and transmits the data to the federal reserve system (frs). 5. the federal reserve parses the data and transmits the data to the fdic. 6. each banking regulator (frs, fdic, occ) maintains its own copy of the most current data. iassist quarterly summer 2007 7 7. the frs and fdic edit the data for their banks and call the banks to resolve data anomalies and process revisions. as revisions are processed, the regulators transmit revisions amongst one another. 8. non-confidential data are released to the public in various forms from the regulators business problems with the historical collection model the historical business reporting model, outlined in figure 1, presents some obvious and not-so-obvious risks and challenges. the challenges that needed to be addressed fell into three categories: 1) multiple collection and storage sites, 2) difficulties for the industry in implementing changes to the data collection requirements, and 3) improvements to data quality. multiple data collection and storage sites coupled with the ability of the banks to revise their data introduces the risk of inconsistencies in the data maintained and published by the various regulators at any one time. the inconsistencies were the result of timing differences, human error, or communication breakdowns between organizations. additionally, redundancies in overhead existed across the regulatory agencies. separate databases and editing applications were maintained by each regulator. although the regulators collaborated and combined efforts where feasible, such as in communicating with the banks, activities such as database administration, edit identification, and metadata maintenance were performed separately at each agency. a larger but less obvious challenge was the need for multiple organizations (regulators, banks, and software vendors) to maintain metadata to support electronic filing. early in the call report modernization process, industry participants (banks and software vendors) identified the need for quality metadata that could be electronically communicated and consumed. necessary metadata include the variables collected, variable definitions, and data edits (business rules that verify the accuracy of the data). revisions to the reporting structure, including adding and removing variables as well as definitional changes and clarifications, can occur multiple times a year. under the historical model, the regulators, banks, and software vendors had to each make revisions to their applications based on changes that were communicated in hard copy, text, or pdf documents. electronic distribution of metadata was appealing because it can be easily altered to reflect reporting requirements. additional business needs included making the editing process more efficient and improving the data quality for the industry. as noted above in the historical model, all data editing was done after the data were received by the regulator. after editing the data the regulator would call the bank to determine if data revisions were necessary or to gather explanations for data anomalies. it was much more difficult to resolve data anomalies and edit failures after the data were submitted to the regulators than at the time the data were compiled by the banks. for example, if a total variable does not equal the sum of the component variables, it is much easier for the bank staff to determine the reason for the difference and fix the data or note the cause of the difference when compiling the data than several weeks later after the data have been submitted to the regulator. 8 iassist quarterly summer 2007 evolution of technology enabling xbrl collecting, distributing, and sharing data have always been very labor and time consuming and are often further complicated by incompatible software and the need to provide volumes of text that explain the data and define the data file layout. although it is currently hard to imagine professional communication before the development of the internet, call report data were collected and disseminated long before common acceptance and use of the internet, which occurred around 1995. in the late 1990s the innovation of xml (extensible markup language) enabled web developers to move from static html documents to interactive web pages. in other words, beyond simply displaying preset reports, xml supports data exchange and machine-actionable processing over the internet. unlike html, xml also enables descriptions of the data to be associated with the data and embedded in the web applications. these descriptive data are commonly referred to as metadata. shortly after xml was introduced, business users began to harness the power of xml and apply standardized data descriptions and naming conventions. the compilation of naming conventions and metadata is referred to as taxonomy. since the creation of xml, several transmission protocols were created and have become industry standards. xbrl (for financial data), sdmx (for time series statistical data), and ddi (for social science data) are three internationally recognized data exchange formats capitalizing on the xml format. specifically, xbrl marries new internet-based technology across dissimilar computer systems, with business rules and practices. xbrl can be thought of as a set of accounting standards coupled with information technology (it) standards that simplifies the exchange of data. this marriage of accounting and it is accomplished through the taxonomy, which is an integral part of the xbrl standards. xbrl taxonomies provide definitions for tags that are identified in the xml schema. within the taxonomy, there are a number of required data elements that provide information about the data characteristics. xbrl examples include the currency of monetary variables, if the variable is from the balance sheet or income statement, and if the variable is a debit or credit. using the xbrl taxonomy facilitates solving the business problems of distributing data collection requirements. the new collection model the collection model that resulted from the call modernization project is outlined in figure 2. in the new model, xbrl supports the collection, editing, and distribution of call report data. each quarter, the regulatory agencies create one taxonomy that includes all of the variables to be collected, metadata, and edit calculations. the taxonomy is downloaded by banks or bank vendors and is used to programmatically update their internal reporting systems. the data are then compiled by the banks and submitted to a central data repository (cdr). as noted above, before the implementation of xbrl, every time the call report was revised each agency manually updated its collection, editing, and data storage systems. in addition, each bank or bank vendor manually updated its individual regulatory reporting system. with the use of xbrl, the taxonomy can be programmatically imported into vendor software, while all three agencies use one system to collect and edit call data. xbrl also makes it possible for edits to be distributed to the banks because the taxonomy contains data validation routines. thus, each bank edits its data before submitting to the cdr, which then reedits the data upon receipt. if any edits fail or anomalies are not adequately explained, the report submission is automatically rejected and the bank is notified immediately. this pre-editing routine reduces the number of follow-up phone calls the agencies must make to analyze questionable data, thereby reducing the burden on both the banks and the regulators. an added benefit of the xbrl taxonomy is the ability to use the metadata to create data presentations. the presentation metadata arrange the call data into a human readable format. for the call reports, it organizes the data into balance sheet and income statement formats. once the presentations are defined, web reports can be created dynamically from the metadata, enabling the cdr to easily publish call data to the general public. although the agencies have made the data available to the public on the internet for years, the process to maintain the web reports was time consuming and the reports were only published in bulk when all of the banks’ data were finalized. under the new xbrl model, each bank’s data are published three business days after being submitted to the cdr. the cdr also has a web service that allows the public to retrieve all call report data. this service is an efficient method to disseminate more timely banking industry data to the public and allows public users to repackage and easily disseminate data2 the creation of an internet-based data collection and dissemination facility is a tremendous undertaking. the u.s. banking agencies had two advantages in place that helped facilitate the new xbrl collection model: well structured data and variable level metadata. the call report data are well structured because comparable information is collected from all entities with well defined rules regarding which banks must report what data. although some of the rules are quite involved, they are structured enough to define programmatically. the variable-level metadata predated the cdr and are housed in a metadata warehouse know as the micro data iassist quarterly summer 2007 9 reference manual (mdrm) at the federal reserve board. the mdrm was originally built as a tool for end users who access the data. however, its well structured naming convention and nomenclature are well suited for the data collection process. although the mdrm did not contain all of the xbrl metadata variables (which have since been added manually), it did provide variable naming conventions, variable history, and detailed descriptions of every variable. pros & cons of using an industry-wide protocol vs. a custom built web interface all data exchange formats provide structure for sending and receiving data. the xbrl protocol includes metadata for each data variable, such as if the variable is from the balance sheet or income statement, what currency it is denominated in, and if the variable is an integer, monetary value, ratio, or text item. the detailed metadata remove any doubt about the meaning of data variables. for example, within xbrl, the variable “total deposits” can be defined in the metadata so there is no confusion whether the deposit is a liability (such as for a bank) or an asset (such as for an individual with a savings account). clearly, an internet based collection system could have been built without using a data exchange protocol. there is even an argument that not using an industry-wide protocol might have been easier to build since having standards requires adherence to its ensuing rules and guidelines. for example, one hurdle encountered by using xbrl is that call report data are reported in thousands of dollars while the xbrl standard requires that data be transmitted in single dollars. to meet the xbrl standard, additional steps are taken to convert thousands of dollars to dollars and then back to thousands of dollars for the reporters and end users who expect the data to be in thousands of dollars. in spite of the added hurdles, the benefits of using the industry-wide protocol outweigh the challenges of meeting the requirements of the protocol. having the standard ensures that any user of the data can access and understand the information as long as he or she has basic knowledge about the standard’s data elements and definitions. one of the earliest examples of the benefits of standardization is the standard rail gauge. the standard rail gauge defined the distance between the rails on railroad tracks. before the standard rail gauge was adopted in the 1870s trains could not travel across the country and cargo had to be unloaded and reloaded to a new train each time the rail gages changed. adopting the standard rail gauge allowed trains to travel on any track providing tremendous economic benefits3. when data are transmitted under different formats or standards the data need to be transformed each time the data are transmitted. just as the rail gauge facilitated the shipping of freight over long distances, transmission protocols allow for an easier flow of data along the information highway. 10 iassist quarterly summer 2007 where xbrl is going there is a growing body of research on the benefits of xbrl. the published research falls into three main perspectives: financial statement consumers, financial statement creators, and financial statement auditors. financial theory suggests that more frequent, reliable, and readable financial statement reports will result in a healthier marketplace. the first published studies focused on market participants’ decision making. recently, research is moving toward xbrl’s impact on reporting costs and regulatory enforcement. financial theory suggests that a market with more informed participants would price risky assets closer to the assets’ true economic value. market participants use reported financial statements to assess a company’s financial strength and adjust their investment in the company accordingly. however, there is a prohibitive search and learning cost to understanding financial statements. as a result, risky assets are priced by fewer investors or worse, misguided investors. while all statements retain core structure (i.e. equity = assets liability), line items and footnote disclosures can differ significantly from industry to industry and even company to company. therefore, early research focused on the financial statement consumers and whether tagged financial data improves transparency and results in a healthier marketplace. one of the first xbrl studies (hodge, kennedy, maines 2002)4 tested investors’ judgments regarding the reporting method of stock option compensation data. stock option compensation is allowed to be reported either on the face of the financial statement or disclosed in the footnotes. investors were shown tagged financial statements to see if xbrl reduces the difference between the financial statement recognition and the footnote disclosure reporting methods. the authors found that without tagged data, the influence of footnote disclosures diminished, resulting in investors making less informed decisions. when using tagged data, investment decisions did not change between reporting methods, suggesting that the xbrl format equalizes recognition of the footnote and the financial statement disclosures. furthermore, investors reported trusting tagged statements more than the untagged data. while the authors found xbrl tagged data eliminated the financial sophistication required to understanding financial statements, xbrl only had added benefits for investors who used the available search technology. the authors found that almost 50% of the survey participants did not use the search technology to the extent that xbrl enabled. new research is examining xbrl’s impact on financial statement creators. the current statement filing process is expensive and time consuming, requiring audited statements to be filed annually and unaudited statements to be filed quarterly. an xbrl-based reporting system offers the potential of more frequent, potentially continuous, financial reporting. a paper by hunton, wright, and wright (2003) argues that there is an optimal reporting frequency somewhere between quarterly and continuous. the authors claim that, on appearance, continuous reporting would align fair market value with asset prices resulting in reduced price volatility but that psychological factors could undo and further exacerbate price volatility. while further research is needed to determine the ideal frequency for financial statement reports, xbrl can report any time frequency desired. successful xbrl implementations, growing international use, and positive academic research influenced the security and exchange commission (sec) and the chartered financial analyst institute to publicly support adoption of the xbrl standard in the united states. both the sec commissioner and cfa president are promoting xbrl reporting standards and its importance to the industry. with these inherent benefits of xbrl, government regulators worldwide are either already using or pursuing the use of xbrl. other significant adoptions of xbrl standards include, but are not limited to: 1) the australian prudential regulation authority (apra) to collect data from australia’s super funds, insurers, and banks, 2) the u.k inland revenue services, 3) the bank of japan to gather data from financial institutions in february of 2006. 4) japan’s financial services agency’s launch of the electronic disclosure for investors’ network (edinet) system in march of 2008, 5) the netherlands’ data collection process for corporate tax, business financial, and business statistical data in 2007. 6) united states generally accepted accounting practices (gaap) and the international accounting standards committee (iasc) taxonomies. there is a sustained need for faster, better, and easier data retrieval in the financial industry. new technology has opened doors for financial reporting that were unimaginable as recently as the early 1990s. although the need and desirability of financial data reporting on a continuous basis is debatable, the current technological environment of xbrl facilitates that option. regardless of the optimal frequency, governments and regulators have an ongoing desire to improve the efficiency of data collection and consumption. in addition to helping the government regulators, improved data flows may also help to improve market efficiencies, as they rely on current and accurate data. using the evolving technology while enforcing compliance with transmission protocol standards may take our information highway to destinations that we are currently unable to envision. bibliography australian prudential regulation authority. “statistics project – data submission – xbrl,” apra, http://www. apra.gov.au/statistics/xbrl.cfm (accessed april 2008). iassist quarterly summer 2007 11 bnet business network. “business services industry,” bnet, http://xbrl.squarespace.com/japan-edinet-reporting/ (accessed april 2008). edgaronline. “home page and various content pages,” edgaronline, http://www.edgar-online.com/ (accessed april 2008). hodge, frank, kennedy, jane jollineau, and maines, laureen (2002). “recognition versus disclosure in financial statements: does search-facilitating technology improve transparency?,” november, 2002. hunton, james, wright, arnie, wright, sally (2003). “the supply and demand for continuous reporting,” publications of the univesiteit van amsterdam (netherlands). pricewaterhousecoopers llc. roohani, saeed j, eds. (2003) trust and data assurances in capital markets: the role of technology solutions. pricewaterhousecoopers llc. securities and exchange commission. “recent publications – various publications,” sec. http://www.sec. gov/ (accessed april 2008). xbrl international home page. “extensible business reporting language,” xbrl, www.xbrl.org/home/ (accessed may 2008). xtech conference (2007). “large scale xbrl introduction for dutch egovernment presentation”, xtech, http://2007. xtech.org/public/schedule/detail/202 (accessed may 2008). * linda f. powell and andrew boettcher, board of governors of the federal reserve system, washington, d.c. this was presented at the iassist conference in stanford 2008. contact: linda.powell@frb.gov. fotenotes 1 the opinions are of the authors and not the federal reserve board. 2 the taxonomy and data are publicly available at cdr.ffiec. gov/public/default.aspx 3. encarta 97 encyclopedia, “railroads”. cd. 4. ”recognition versus disclosure in financial statements: d3oes search-facilitating technology improve transparency?” hodge, frank; kennedy, jane jollineau; maines, laureen a november, 2002 vol29-4.indd 16 iassist quarterly winter 2005 by makoto hanashima* consideration for information security issues in geospatial information services of local governments abstract geographic information system (gis) is one of the important elements of the it infrastructure of local governments. in recent years, in the consequence of development of web-based gis technology, the definition of gis has been changing to “geospatial information service”. for a local government, the geospatial information service is becoming an important it system in the field of information service or in the field of information disclosure for its residents. since all of the governmental geospatial information is public property, it should not only be utilized but be protected. however, there is no standard or guideline for the information security regarding the geospatial information service in japan. therefore, a study from the stance of it security is required in order to secure the geospatial information service and to protect the geospatial information as public property. iso/iec tr 13335 gmits that is one of iso standards for the it security management is a useful framework for discussing regarding it security policy of an information service. this paper aims to consider regarding it security requirements for the geospatial information service and analyze the specific threats to them according to the framework of gmits. based on the threat analysis, a set of safeguards for the baseline security of the geospatial information service is selected. in addition, technical issues of the safeguards to clarify the feasibility of them are discussed. introduction geographic information system (gis) is one of the important factors that constitute social it infrastructure today. furthermore, the system that carries out interoperability of the geospatial information data through the internet is becoming feasible. in this situation, gis has been changed to the terminology meaning geospatial information service. gis in local governments is the main application field of gis. introduction of gis in japanese local governments shows rapid progress ignited by the kobe earthquake in 1995. according to a survey in 2004 by the national spatial data infrastructure promoting association (nsdipa), the introductory rate of gis in the allprefectures agency is 100%. moreover, in 40% of the cities, towns and villages, gis introduction has been completed. the main purpose of gis in local governments is management of geospatial information in regard of administration and decision support. in addition, utilization of gis for information disclosure and the public information service has become important year by year. since the distribution of the geospatial information service can contribute to the improvement of the welfare of residents, it is a desirable usage of public resource. on the other hand, standardization of the specification of web map service (wms) has been carried out by open geospatial consortium (ogc) inc., and the technical infrastructure of geospatial information service has been developed year by year. not only the service of conventional “end to end” but also the service chaining by combining two or more geospatial information services are becoming possible. it is expected that the circumstances of the geospatial information service in local governments also change a lot by such technology. in some local governments, the operation of “wide-area integrated gis” which has multi-service interoperability of geospatial information has already started. in such condition while a user’s convenience and the quality of service improve greatly, the possibility of information security problem also increases. the service that a local government provides must not threaten the residents’ safety. in order to prevent illegal usage of the geospatial information that is public property, the countermeasure based on suitable information security is indispensable. however, in japan, those countermeasures are being entrusted to each local government. therefore, in consideration of the circumstances of geospatial information services of local governments, to discuss those requirements for the information security systematically is required the meaning and the purpose of research in order to make the information service on the internet secure, many technologies and standards already exist and iassist quarterly winter 2005 17 are used in various fields. thus, it is possible to secure the geospatial information service by technology theoretically. all the same, the question i have to consider here is whether or not to discuss about the information security of geospatial information service. generally, in order to implement suitable it security management, it is necessary to define clearly it security requirement peculiar to the system or field. even if it is using the same web-based system technology, a security requirement is not the same in a medical information system and in gis. it will be very natural to set up a suitable and logical security requirement in consideration of the characteristics of each it system. nevertheless, i think that such discussion has not been fully held in the field of geospatial information service. probably, there may be an opinion that the discussion about the it security requirements should be held at the process of system design of an actual system. such opinion will be reasonable if it is from the stance focused on implementation. it will not matter if sufficient security in that approach is ensured. however, since such bottom-up approach is not necessarily systematic, in many cases, there is a risk of the defect of security requirements. therefore, i need the systematic and logical discussion regarding the requirements for an information security of geospatial information service. as a framework for discussing such issues, there are the information security related standards and guidelines of iso. i think the most appropriate standard is iso/iec tr 13335 guidelines for the management of it security (gmits). gmits provides a systematic framework for it security management. being based on the concept of gmits, we would like to clarify the meaning of discussing it security of geospatial information service. “the part 4” of gmits has described the selection process of the safeguard of it system for threats. according to the description of that part, when setting up the security level in an it system, there are two approaches. one is the approach based on a “detailed risk analysis.” a detailed risk analysis estimates risks based on detailed evaluation of the information property that is a target for protection, the threat evaluation to the information assets, and vulnerability evaluation of it system. therefore, it is possible to select the safeguard optimized to the target it system. on the other hand, since this approach needs high-level technical knowledge and a great effort, it requires many resources. thus, it is suitable for the system that needs an advanced information security. oppositely, depending on the type of system, it may become surplus specification. another process is “baseline approach”. baseline approach selects a safeguard (baseline safeguard) so that the minimum-security level (baseline security) decided for each type of it system may be satisfied. because this approach can be implemented in minimum time and effort for a risk analysis or for selection of safeguards, for a system that does not need a high security level, its cost benefit is far good. instead, this approach depends on the adequacy of baseline security. when the standard baseline security already verified in a system of the same kind does not exist, it cannot be implemented. actually, the peculiar baseline security guideline is already specified by a certain kind of application fields (medical service, financial information system, etc.). the application in such fields can implement a minimum safeguard without a detailed risk analysis. however, as far as i know, in japan, such a guideline regarding geospatial information service of local governments does not exist. the fact that there is no guideline of baseline security suggests that there can be three cases in the approach of it security of geospatial information service of local governments. those cases are as follows (1) it is based on the detailed risk analysis. (2) it is based on informal approach. (3) it is based on the baseline safeguard of other information service. since (1) requires high costs as i mentioned previously, it is difficult to carry out in a standard local government. therefore, it is thought that (2) or (3) will mainly be chosen. in that case, following problems may arise by choosing these approaches. a. when the interoperability of geospatial information service is carried out, complicated processing is required because of the differences in the security implementation between services. b. redundant investments to the security countermeasures that may not be so effective will continue in many local governments. c. some specific risks of geospatial information service may remain without consideration. in order to prevent these problems, a systematic and logical discussion will be required regarding baseline security for geospatial information service. if the specific requirements for the information security of geospatial information service become clear, the guideline that contains these requirements in baseline security can be proposed. even if there is no specific security requirement, certain criteria about baseline security should be shown. 18 iassist quarterly winter 2005 from the above reasons, this research aims at contributing to the making of a practical guideline by discussing the baseline security of the geospatial information service in a local government, and proposing the prototype. the workflow of research in this research, the framework of gmits is referred to, in order to discuss the baseline security of geospatial information service. fig. 1 shows the relation between the framework of the security management of gmits and this research. as shown figure 1, the target of this research is a proposal of a baseline security guideline for the geospatial information service with consideration for the institutional restrictions of a local government and the actual condition of business environment in japan. for that purpose, not only gmits but also other standards may be applied. the reason for referring to these standards is to provide the systematic conceptual framework, but not to create the proposal based on a specific standard. (of course, there is a possibility of designing a proposal based on the standard because of discussion.) now, let us explain the workflow of our research. a. discussion of it security policy regarding the geospatial information service there are two levels of the security policy in it security management. a high-level it security policy is for the whole organization, and the security policy of individual it system is created based on it. in this research, the security policy of the geospatial information service is discussed as one of the individual it security policies, assuming that overall it security policy of the local government is already created. based on the governmental guideline, i interpret the security issues regarding the distribution of the geospatial information, and extract the security requirement of an it figure 1 iassist quarterly winter 2005 19 figure 2 system. b. outline risk analysis gmits recommends evaluating the importance and influence of an it system, before considering a detailed safeguard. (see fig. 2) that process is called “outline risk analysis”. when the result of “outline risk analysis” indicates that the requirements of it security are high, a detailed risk analysis needs to be implemented. (for example, a national-defense information system, the system regarding high-energy waste, etc.) in that case, i do not need to discuss about that kind of system, since the it security is implemented depending on the specification of the system. on the other hand, when the result of “outline risk analysis” shows that it is not necessary to implement detailed risk analysis, an appropriate baseline approach should be considered. therefore, the process of “outline risk analysis” should be clarified. the plan of this research is to build up a simple evaluation model of “outline risk analysis” based on a minimum risk assessment for the geospatial information service in local governments. c. making of a safeguard catalog the requirements for selecting the safeguard to risks of it security become clear according to process b. by cataloguing systematically the safeguard that suits the requirements, it becomes easy to define it security requirements. d. making of a baseline security model for geospatial information service on the basis of the safeguard catalog, a set of the safeguards is selected to construct a baseline security model for the geospatial information service. the baseline security model is defined in order to fill the security requirements of the geospatial information service. e. proposal of a baseline security guideline prototype the last step of the research is to propose a baseline 20 iassist quarterly winter 2005 security guideline for local governments. this paper shows the result of the discussion regarding the above-mentioned a and b. from now on, the research in this paper will be developed further and it is necessary to make it more practical. future research is due to show the prototype of a guideline. the security policy regarding geospatial information service at present, there is no ordinance that refers concretely to the security policy of geospatial information service in japan. the only document, “the guideline regarding distribution of governmental geographic information”, has been released by the “related ministry liaison conference for gis” in 2003. this guideline is a de facto guideline of the geospatial information service by the public institution in japan. the most important point of this guideline is recognizing that the properties of geospatial information are different from those of other administration information. it also points out that the guideline is necessary in order to promote the distribution and disclosure of geospatial information. the main points of the guideline are summarized as follows. (1) governmental geospatial information is public property which all the people can enjoy conveniently. therefore, it should not only be used inside an institution, but be positively provided for people. (2) although geospatial information is one of the administration information, smooth distribution may be unable to be promoted only by being based on the existing government ordinance, because of the specific properties of geospatial information. those specific properties are as follows. a. since most of geospatial information is created by association of the work of various authors, the provider cannot avoid various rights management in regard of copyright. b. in the case of almost all administration information, governmental duty is achieved by the disclosure of information, but for geospatial information it is important that a user can reproduce or process it after its distribution. (3) therefore, the guideline regarding information service based on the properties of geospatial information is necessary. (4) the basic policy regarding the distribution of geospatial information that the government owns is defined as follows. a. in principle, the government shall not set a restriction in usage of geospatial information that is distributed freely through the internet as much as possible. b. from a viewpoint of ensuring governmental accountability, whereabouts, distribution propriety, distribution process, distribution conditions, etc. shall be indicated on the internet. c. the government shall make sure that geospatial information may not correspond to the nondisclosure information defined by “freedom of information act” of japan as much as possible. furthermore, in 2004, the liaison conference released the “collection of q&a” as the guide of actual operation. the handling of the copyright of governmental geospatial information and the interpretation of the government regarding the issues for actual operation are described in this “q&a.” in addition, this guideline describes, “although judgment regarding geospatial information is fundamentally entrusted to a local government, it is possible to provide geographic information according to the distribution guideline.” when a local government considers its distribution of geospatial information, these two documents are considered to become a source of a basic security policy. therefore, i will be able to extract the security policy requirements for geospatial information service of a local government from this guideline. the requirements are as follows. a. protection of geospatial information regarding privacy. b. ensuring confidentiality of nondisclosure geospatial information. c. ensuring integrity and authenticity of geospatial information. d. management of the access privilege of geospatial information. e. prevention from violation of the copyright of geospatial information. f. maintenance of accountability of local government for geospatial information. g. ensuring availability of geospatial information service. threats and damages over geospatial information service next, i would like to consider the threats over these requirements, and the damage caused by them. first, threats are divided into two classifications. one iassist quarterly winter 2005 21 is “typical threat” that is enumerated in gmits part 3 annex c “list of possible threat types”. another one is “specific threat” that is peculiar to the geospatial information service. “typical threat” and “specific threat” are not independent of each other. there is an intersection area as shown in fig. 3. in order to clarify actual risks, some “typical threat” should be redefined as “specific threat”. based on our consideration about “specific threat” regarding gis of local governments, i have enumerated them as follows. a. protection of geospatial information regarding privacy this requirement is endangered by leakage, theft, failure, loss, tampering, forgery, etc. of privacy information. the possible damages in this case are as follows. a. risk of infringing on individual (corporate) security and profits arises. b. the legal liability based on a related statute is prosecuted. c. the evaluation to administration falls. d. the promotion of utilization of service is obstructed. “specific threat” which causes such damages includes: (1) exposure of the privacy information by a connected referencability (a vulnerability that is able to guess the nondisclosure information in a certain database by indirect reference of the information in other databases.) (2) tampering and forgery of data (3) illegal copy of data b. ensuring the confidentiality of nondisclosure geospatial information this requirement is endangered by leakage, theft, failure, tampering and forgery of nondisclosure information. the possible damages in this case are as follows. a. the threat on the security by leakage of a social infrastructure or national-defense-related nondisclosure information is caused. b. the threat to safety of residents is caused. “specific threat” which causes such damages includes: figure 3 22 iassist quarterly winter 2005 (1) exposure of the nondisclosure information by a connected referencability (2) tampering and forgery of data (3) illegal copy of data (4) data error c. ensuring integrity and authenticity of geospatial information this requirement is endangered by forgery, tampering and defect of geospatial information. the possible damages in this case are as follows. a. there is a possibility that all of the services using the geospatial information will be influenced by the defect of information. b. there is a risk that tampered digital map or geospatial information is abused as an official document. c. trouble or crime by the tampered geospatial information is caused. d. the reliance on geospatial information falls. «specific threat» which causes such damages includes: (1) tampering and forgery of data (2) illegal copy of data (3) data error (4) masquerading of a source d. management of the access privilege of geospatial information this requirement is endangered by fraudulent procurement or setting error of the access privilege of geospatial information. the possible damages in this case are as follows. a. the serious threat over the security of entire information system is caused. b. the services that use geospatial information are disturbed. c. the reliance on geospatial information service falls. «specific threat» which causes such damages includes: (1) setting error of an access privilege (2) the attack by unauthorized service (3) the attack to web application e. prevention from violation of the copyright of geospatial information this requirement is endangered by violation or masquerade of copyright of geospatial information. the possible damages in this case are as follows. a. risk of infringing on the profits of an author or a user is caused. b. the legal liability relevant to protection of copyright is prosecuted. «specific threat» which causes such damages includes: (1) setting error of an access privilege (2) masquerading of an author or a source (3) tampering and forgery of data (4) illegal copy and distribution of data f. maintenance of accountability of local governments for geospatial information this requirement is endangered by failure, disturbance, and loss and tampering of audit information. the possible damages in this case are as follows. a. risk of overlooking an information crime and the illegal use of information is caused. b. possibility that spoils the accountability of administration is caused. c. the legal liability regarding business administration of local government is prosecuted. «specific threat» which causes such damages includes: (1) tampering and deletion of an audit log (2) data error (3) setting error of an access privilege (4) tampering and forgery of data (5) illegal copy and distribution of data g. ensuring the availability of geospatial information service this requirement is endangered by failure of system, iassist quarterly winter 2005 23 hardware, network, etc. the possible damages in this case are as follows. a. it interferes with the correspondence in the time of disaster and emergency. b. it becomes impossible to maintain the business continuity of a local government. c. it has a bad influence on the service of other local governments. d. the reliance on geospatial information service falls. «specific threat» which causes such damage includes: (1) failure of interoperability between systems (2) the attack by unauthorized service (3) the attack to web application (4) tampering and forgery of data (5) data error 6. “specific threat” to geospatial information service since «typical threat» has been described by gmits, let us explain regarding «specific threat» concretely here. (1) tampering and forgery of data it is difficult to distinguish the tampered geospatial information data. for example, when the polygon data of a land partition is changed or a non-existing block is forged, the first people that see the map are unable to detect the tampering with the geospatial data. the geospatial information data of a local government can be used as an official document, if the criterion is satisfied. therefore, the general user will not suspect the authenticity of the map. when a tampered map data is provided from the gis service of a local government, a very dangerous situation may be caused. (2) illegal copy and distribution of data although geospatial information data is public property, it cannot be necessarily reproduced unconditionally. copyright is generated even if the geospatial information has been made by public institutions. since the utilization conditions of geospatial data are sure to be specified by the government, all the reproductions that do not follow the conditions are illegal copies. when one map image is the result of editing a set of geospatial data, copyright exists for all the data. therefore, in such reproduction of a map image, it is necessary to obtain agreement of all the copyright holders. if a large amount of the illegal copy that disregarded the utilization conditions is distributed, there is a possibility to cause the trouble of copyright infringement. (3) the attack by unauthorized service interoperability technology realizes cooperation of several web services, but it also enables malicious web service attacks to gis without a user’s awareness. a web service hijacked by the attacker can be injected with tampered geospatial information data which can execute damages to the service. (4) the attack to web application usually, a geospatial information service is implemented as a web application. although various measures are taken about it security of a web application, it is not easy to build a web application without vulnerability. when a web application is attacked, it causes execution of malicious codes, hijack of a system, tampering of a web site, and so on. (5) masquerading of the author or the source on the utilization conditions of geospatial information, the designation of the copyright holders or the source is required in many cases. however, when there is no way of attesting the authenticity of the author, a malicious third party may misrepresent the author, and the user cannot know it. it is not necessarily authentic information just because the copyright notice is indicated. therefore, it is possible that copyright notice is abused in order to camouflage the information tampered by the third party. (6) setting error of the access privilege generally, geospatial information consists of the data of many features. in order to manage nondisclosure information, it is necessary to set up the access privilege of each feature, and it is thought that very complicated management is performed. if the setup of an access privilege has an error, it causes leakage of nondisclosure information immediately and causes various security problems. (7) exposure of confidential information by a connected referencability although geospatial information itself is not confidential information, a certain kind of geospatial information may expose confidential information, when it is used combining other information. a safeguard cannot be implemented individually in each case, since it is very difficult to detect and prevent a connected referen24 iassist quarterly winter 2005 cability. if the exposure of nondisclosure information occurs, there is a risk that a legal liability is prosecuted. (8) data error it is very difficult to discover the error of geospatial information. although a geometric check, a numerical check of attribute information and so on by a computer are possible, there are many errors undetectable with those checks. for the service that an independent local government provides, its influence will stop in that region, but when interoperability of geospatial information is being performed, there is a risk that erroneous data is diffused in many services. (9) tampering and deleting audit log information the audit information, such as audit log files is indispensable, because accountability is an important requirement for a local government. when such audit information is lost or is tampered, a serious threat over the safety of the whole it security is caused. moreover, since accountability of a local government is lost, there is a risk that a legal liability is prosecuted. (10) failure of interoperability between systems when geospatial information service is operating based on the interoperability of web services, a failure of one service will affect other services immediately. for example, if the service that provides the background map stops, others cannot continue their services. therefore, such failure of interoperability between systems is a serious threat to geospatial information service. the safeguards to «specific threat» following consideration above, i discuss what kind of safeguard is applicable to each «specific threat». the results of the discussion are shown as follows. (1) tampering and forgery of data a. the access control to data b. authentication of the authenticity of the data based on digital signature c. tampering-proof data generation (2) an illegal copy and distribution of data a. authentication of the authenticity of the data based on digital signature. b. authentication of the data provider by digital signature (3) the attack by an unauthorized service a. mutual authentication by the security framework of a web service b. mutual authentication in an application level c. reinforcement of the detection function of an unauthorized service (4) the attack to a web application a. reinforcement of robustness of a web application b. reinforcement of attack detection function c. using rich client (5) masquerading of an author or a source a. authentication by digital signature of an author or a source b. using digital watermarking (6) setting error of an access privilege a. application of an access-control model b. using an access-control framework (7) exposure of the confidential information by a connected referencability. a. the limitation of the resolution by metadata (8) data error a. early distribution of error information b. audit of the updating log of data (9) tampering and deletion of an audit log information a. reinforcement of a logging system (10) the failure of interoperability between systems a. implementation of the error-tracking function of a web service summary of the safeguards the above-mentioned list of safeguards includes those whose feasibility has not been examined. that is, they can iassist quarterly winter 2005 25 be sorted into the following two categories. a. the safeguards that can be implemented by existing technologies (1) implementation of ws-security technology web service-related safeguards are almost feasible by applying the technology of ws-security in which standardization is promoted by organization for the advancement of structured information standards (oasis). (2) application of a secure data transmission protocol regarding the data transmission of «end to end», it is secured by using secure sockets layer (ssl). however, in the case of the data transmission over more than one service, other technique is required. (3) application of an access model based on access models, such as role based access control (rbac) model, the access control that is able to ensure consistency of the whole system is feasible. b. the safeguards that need further discussion (1) improvement in the robustness of a web application since it is depending on the implementation of an individual service, the systematic process that defends the attack on web applications has not been established yet. the discussion regarding construction of a robust web application system should cover not only the software skill but also the aspect of operation. (2) realization of the traceability of a web service component if the interoperability between web services becomes complicated, it will be difficult to detect the failure of a web component. the function to trace the web service that is malfunctioning in a service chain and to perform suitable failure processing is necessary. (3) implementation of the digital signature and authentication protocol in the open geospacial consortium (ogc)’s open architecture. there is no description regarding security in the open architecture specification of ogc, such as web map service. how to implement security requirements is depending on each application system. however, regarding a signature and authentication of geospatial data, i consider that it will be more suitable to build in the framework of gis technology. thus, more technological consideration is required about this issue. as mentioned above, regarding the safeguard to «specific threat», a continuous discussion is necessary. conclusion in this paper, i clarify meaning of discussing it security of the geospatial information service in local governments at the first, and then specified the workflow of research. then, i consider regarding the framework of baseline security for geospatial information service of local governments, and suggest that the baseline approach of gimits was an appropriate framework. furthermore, based on the governmental guideline, i show the it security requirements for geospatial information service of local governments, and discuss regarding the threat over those security requirements. finally, i enumerate the possible safeguards to «specific threat» of geospatial information service, and consider regarding their technical issues. although this paper is the first step in my research, i think that i can show the meaning and the importance of the baseline security of geospatial information service. from now on, the following issues will be discussed in the research. (1) a conceptual model for «outline risk analysis» (2) revalidation of «specific threat» (3) verification of safeguards against «specific threat» (4) abstract specification of the baseline-safeguards for geospatial information since it is thought that the image of a specific system is required to progress the discussion, i am going to define the geospatial information service as a virtual workbench. further consideration regarding the safeguards for «specific threat» of gis will be carried out with the virtual workbench. references belussi,a,et al. “an authorization model for geographical maps” in proc. gis’04. nov. 12-13. 2004. downs,r & lenhardt,c. “privacy and confidentiality issues with spatial data” iassist 2003. iso. “geographic information web map server interface” 26 iassist quarterly winter 2005 iso/dis 19128. 2004. iso. “iso/iec tr 13335 guideline for the management of it security” jis handbook 2005: 74-415. iso. “iso/iec15408:2005 common criteria for information technology security evaluation version 2.3” iso. 2005. it strategy committee. “e-japan strategyⅱ” ministry of internal affairs and communications (mic). japan. 2003. http://www.kantei.go.jp/jp/singi/it2/ japan standards association. “jis handbook – information security 2005” japan standards association. 2005 joshi,j, et al. “digital government security infrastructure design challenges” ieee computer. 2001. ministry of internal affairs and communications. “manual for implemetation and operation of integrated gis” 2004. nsdipa. “annual survey of gis in local governments” http://www.gisportal.jp/case/lo_case/h16.html oasis. “wss soap message security (ws-security 2004)” oasis open. 2004. oasis. “xacml profile for role based access control (rbac)” committee draft 01. oasis open. feb. 2004. ogc. “ opengis reference model” ogc 03-040. open geospatial consortium inc. 2003. ogc. “opengis web map server cookbook” ogc 03050r1. open geospatial consortium inc. 2004. related ministry liaison conference for gis. “q & a: the guideline regarding distribution of governmental geographic information” 2004. related ministry liaison conference for gis. “the guideline regarding distribution of governmental geographic information” 2003. taylor,k & murty,j. “implementing role based access control for federated information systems on the web” australasian information security workshop 2003 (aisw2003) watanabe, kozo. “practical guide for integrated gis in local government” nikkan kogyo shimbun inc. 2003. * makoto hanashima is a senior researcher at the institute for areal studies, foundation (ias), tokyo and the institute of information security (iisec), yokohama. email: mhana@ias.or.jp. the paper was presented at the iassist 2006 conference in ann arbor, michigan, usa in the session, «the big picture: gis data challenges and solutions». vol252 iassist quarterly summer 2001 19 in 2000 the russian public opinion research center (vciom) initiated a project on compiling a national sociological data archive supported by the ford foundation. the aim of the project is to work out the content, the organizational and financial principles of the formation and further functioning of the national sociological data archive on the basis of a restricted number of vciom surveys as well as of other research institutes carrying out representative sociological surveys. beginning from the mid and late eighties, the need for a national archive has been discussed by sociologists more than once. however, it was only by the late 90s that real conditions had been created for the project to become viable. first, at present there are many research companies conducting national and international representative surveys which are of scientific value for a wide circle of researchers. this situation posed new questions related to the long-term data retention, their standardization and availability for researchers. of course, each research company solves itself these problems dealing with their own databases and their possible users. however, as vciom’s experience shows, possibilities of the solution of these problems within each separate company are limited. second, besides the existing companies, there are more and more researchers and research teams who conduct surveys supported by grants from national and foreign funds. the data obtained in these surveys are used by the researchers themselves and are unavailable for other interested users, which is contrary to the very nature of non-commercial support of research. third, due to the development of sociological education in russia for at least 10 years, there is an increasing circle of potential consumers of sociological data, who need access to such data for secondary analysis in the teaching process, for preparing graduation papers, dissertations, articles and monographs. the facts mentioned above make the sociologist’s society move towards compiling a national archive. however, we realize that this task cannot be solved overnight. it will take at least 2-3 years to create the archive, as well as the effort and goodwill of many interested organizations, first of all those which are prepared to deposit their data in the archive for storage and dissemination and also those which are prepared to give financial, material and organizational support. 1. the expected results of the pilot project are as follows: 2. data files formed on the basis of the data provided by vciom and other organizations which agreed to take part in establishing the archive; 3. the information-retrieval system which will enable the users to find the necessary information; 4. organizational, technological and financial principles of the archive functioning which will make the archive a social institution (designproject); the duration of the pilot project is one year – from january to december 2001. the project manager is dr. l. khakhulina (vciom), coordinator and executive manager – dr. l. kosova (vciom), consultant – dr. a. kryshtanovsky (higher scholl of economics). it is planned to hold an international seminar to present the results of the pilot project of establishing the national archive and to discuss its main technological and organizational principles. the principal characteristic of the pilot stage of the formation of the sociological archive is the fact that vciom, the coordinator of the work, plans to do it together with other interested companies which carry out sociological surveys. in the course of negotiations with the leaders of the most well-known and highly qualified institutes such as a.oslon (public opinion foundation), e.bashkirova (romir), l.drobizheva, n.rostegayeva, (institute of sociology) m.gorshkov, n.tichonova (russian independent institute of social and nationalities problems) showed interest in the participation in the setting up of the national archive. working together with these organizations it is planned, first, to form data archives which will constitute the “core” of the future archive, second, to agree on the principles of cooperation of the archive and the depositors owners of the information, on the one hand, and of the archive and its users, on the other hand; these principles will be the basis for the functioning of the national archive. third, and it is probably the most important item, to form the board of on the pilot project “sociological archive” by ludmilla khakhulina & larisa kosova 1 20 iassist quarterly summer 2001 trustees (experts) of the program “sociological archive”, which will define the main directions of the archive’s work from the scientific and organizational points of view. upon completion of the pilot project the work on the sociological archive formation will be continued. it means that the general design-project of the archive, the sociological data files and the information-retrieval system upon completion of the pilot project will be handed over, with all the legal formalities observed, to the program “sociological archive” functioning within the framework of the independent institute “social policy”. priorities in the selection of information at the pilot stage. in theory, the sociological data archive can and should contain any data obtained in the course of sociological surveys if they meet certain requirements of the archive. besides, the archive may contain also texts based on these data (articles, monographs, etc.). however, since the time of conducting the pilot survey is limited, we restricted ourselves at this stage to prepare the data files which are the most “popular” from the users’ point of view (assessed by the number of request sent to the vciom), namely: 1. russian data of the last 2-3 years obtained on a national sample 2. data from comparative international surveys in which russia took part 3. electoral surveys 4. sociological trends. requirements to the data submitted these are the same requirements as they are set in other archives of europe and the usa. following these rules we get the possibility to become part of the “archive” community. naturally, it will require additional work in some of the companies to fulfill the requirements. we have worked out the outline of a document based on the general requirements and our experience – “guidelines for the depositor” which describes the content and volume of work to be done for transferring the data to the archive. if a company agrees to participate in this work, it receives this document to prepare the data for the archive. relations with the depositors an important issue in the relations between the archive and the depositor is the copyright for the information. we proceed from the general provisions which are as follows: the archive is a social institution established on the basis of a voluntary agreement of information owners for the purpose of storing and disseminating this information. it means that the proprietor of the information is its “producer”, while the archive only receives the copyright for the information administrator to perform certain functions, namely: • storing the information received in the established order, having full responsibility for its physical integrity • dissemination of the information among the users for secondary analysis according to the provisions discussed with and approved by the producer (acting as the depositor) and the general rules of all the archives. from the legal point of view the above stated means that the archive and the depositor conclude a contract (agreement) on transferring the copyright for the information, which stipulates the rights and duties of each party. this contract can stipulate all the conditions related to the transfer of the data to the archive. the body of the contract states that the organization, a potential depositor, agrees to prepare its data and the relevant documentation in accordance with the “guidelines for the depositor”. in its turn, vciom as the coordinator of the project is obligated to ensure the physical integrity of the prepared data and documentation until they are transferred to the archive (within the framework of the institute “social policy”). if an organization prefers not to transfer the data to the project coordinator (vciom, in this case) it will submit only the required documentation for the research (study description, questionnaires, methodological reports, etc.), which will be entered in the data-retrieval system of the future archive. thus at the initial stage we proceed from the possibility that the archive could be based on a distributed storage of data files. in other words, a certain part, preferably the larger part, is stored in the archive, while the rest of the data are kept in the organizations – depositors, which agree to make the data available for the user in the required quality, format and time. another problem in the relations with the potential depositors is the stimulation to deposit their data in the archive. at present, we suggest that the following stimuli should be used. the founders of the archive, on the one hand, have the possibility to present their surveys and their companies to the interested public and to have an open access to the data of other companies in the archive and, on the other hand, they can form the quality standard for surveys which can be accepted by the archive for storage and dissemination. the practice of relations between the archive and its users within the framework of this pilot project is not elaborated. the vciom as the project coordinator does not take upon itself the task of disseminating the information (data) of other companies. first of all because the legal procedure of transferring the copyright for information has not been worked out. we just generally assume that, as is iassist quarterly summer 2001 21 the practice of all other archives, information is made available for the academic community free of charge, especially under the current circumstances of very modest financial resources of both academic institutes, universities and researchers, professors and students. as it is the practice of the european archives, the principle of free data provision does not extend to commercial organizations (consulting companies, advertising and pr agencies, marketing organizations, etc.), which may apply to the archive for certain data. the model of relations with such organizations is planned to be different this question, however, has not been considered from the practical point of view so far. at present, a data file is being formed, the informationretrieval system is being designed, and the interaction with potential depositors is being worked out. 1. contact: vciom, kazakova 16, moscow 103064, russia, phone: +7 095 2654772. ludmilla khakhulina: lkhahul@wciom.ru, larisa kosova: lkos@wciom.ru 1/4 nyangoma, judith (2021), data dissemination by uganda bureau of statistics, iassist quarterly 45(3-4), pp. 1-4. doi: https://doi.org/10.29173/iq991 data dissemination by uganda bureau of statistics nyangoma judith1 abstract data plays a big role in educating the population on various issues that contribute to development. one of the major activities conducted after data collection is, dissemination of the data to different stakeholders. uganda bureau of statistics disseminates data to its users through a number of channels. this paper discusses each method in detail and how it's used during this process. the major channel of sharing data with users is through dissemination workshops and the website. other channels used for dissemination include the library and resource centre, social media and physical delivery to stakeholders in district public libraries. having the above-mentioned channels of data dissemination in place, has helped ubos remain the centre of excellence in dissemination of data to users, countrywide and in africa. keywords data, dissemination, uganda bureau of statistics, uganda introduction uganda bureau of statistics (ubos) is the principal data collecting, processing, analysing and disseminating agency responsible for coordinating and supervising the national statistical system. formerly, ubos was known as the statistics department under ministry of finance, planning and economic development. it was later on, transformed into a semi-autonomous body by the uganda bureau of statistics act no. 12, 1998. the decision to establish the bureau arose from the need for an efficient and user-responsive agency that would meet the growing demand for statistics on social, economic and political developments in the country. ubos is, therefore, coordinating the development and maintenance of a national statistical system which will ensure collection, analysis and dissemination of integrated, reliable and timely statistical information (ubos, 2020). ubos is mandated to develop and maintain an integrated, coherent and reliable national statistical system (nss). the bureau, therefore, has the dual role of producing and disseminating quality statistical information as well as coordinating, monitoring and supervising the nss. in totality, the bureau produces key statistics to support and inform the national and international results based management (rbm) development agenda. the bureau under one of its divisions (communication and public relations), disseminates data after having sensitised the public of its intended surveys, carrying out field interviews and final analysis of data. lots of leaflets concerning the survey are given away to the public. the media is also involved to ensure the public is aware of what is going on in regards to a specific survey. the end-result is the published reports which are well distributed to the public and and some african countries. objectives the objectives of the study were to: i. find out the type of data disseminated by uganda bureau of statistics. ii. find out the channels used in data dissemination. iii. find out the challenges faced in data dissemination. https://doi.org/10.29173/iq991 2/4 nyangoma, judith (2021), data dissemination by uganda bureau of statistics, iassist quarterly 45(3-4), pp. 1-4. doi: https://doi.org/10.29173/iq991 type of data disseminated the uganda bureau of statistics disseminates socio and macro-economic statistics. the type of data disseminated by ubos includes, among others, population and social statistics, households, agriculture, business establishments, macro-economic statistics i.e., gdp, cross border trade, consumer price index, among others. methods of data dissemination the methods of data dissemination are dissemination workshops, hand-delivery, social media, library and resource centre, word-by-mouth and online delivery. channels ubos uses to disseminate data word of mouth channel of communication wherever we go by use of taxis, motorcycles or on foot, we as ubos staff tend to find out if the public has heard about us. at least, they are familiar with the population census and household surveys. they, particularly, participate in responding to the questionnaires. then, we go ahead to tell them the information that we need to be shared with colleagues or their next clients. the same approach is applied to taxi drivers and passengers. by this channel of dissemination, a wide coverage of cities or districts is done, hence fulfilling our aim of data dissemination. while talking to someone in a taxi or any other type of vehicle, a message is heard from the horse's mouth (ubos) and easily trusted and at the same time, clearly understood. that way, this person passes on the same message to another. the challenge with the above channel would be an incentive. most of these people tend to ask for it. the question, these days is, how are they to benefit from passing on the message on ubos' behalf? however, it's an effective method to pass on data through all those people who use public means or those who tend to have walks. along the way, you get to walk besides someone, hence initiating a talk. we intend to make the word-of-mouth channel of communication become official because of its effectiveness. we ensure to get the mobile numbers of the persons/motorcyclists and passengers that we talk to, in order to acquire feedback from them. the social media channel the social media that is normally used for data dissemination is whatsapp. we also use the email addresses and hence, collect as many as possible. majority of people no longer make calls to friends, officemates or relatives but find using the whatsapp and twitter social media as the easiest and flexible means of communication. this has also extended to the delivery of data to the users and especially the policy-makers. many groups for such a purpose are now in place, for example, the ’ubosstaff’ which has effectively helps to deliver information to staff, especially on the covid 19 updates through a newly created newsletter that we circulate to friends and relatives, ministries, departments and agencies, who also circulate the information. there is another group in place, named baseline education survey on which we continuously deliver related data. websites and emails do a similar purpose but not as effective as whatsapp and twitter, from which there is instant information dissemination and feedback. the uganda bureau of statistics website is https://www.ubos.org. it's a necessity for all institutions to make official, the use of whatsapp and twitter as an effective method of information delivery. hand-delivery channel hand-delivery of data or information is a good channel especially to the districts in the country, that don't have access to internet. here, you get the chance to discuss further with the officials and get https://doi.org/10.29173/iq991 https://www.ubos.org/ 3/4 nyangoma, judith (2021), data dissemination by uganda bureau of statistics, iassist quarterly 45(3-4), pp. 1-4. doi: https://doi.org/10.29173/iq991 feedback if needed immediately. brochures and publications by ubos are delivered to them by the help of vehicles, accompanied by an official letter and officer. it would be a good move if the motorcycles are also involved in the delivery of this data. this channel of hand-delivery has been continuously used world-wide before the internet came on board, and we still use it. as uganda bureau of statistics, this is handy because we always have different ongoing surveys, almost in the same period due to deadline requirements from government and donors. hence, fuel and transport are always available for data dissemination, without necessarily using only the dissemination budget. brochures and newsletter channel use of brochures and newsletters remains a channel for data dissemination. these provide summarised long-term data. these can as well be scanned and placed on any social media of choice, distributed through word-of-mouth, motorcycles and vehicles. generally, a lot of hand-delivery is done for data dissemination because of its being handy. the brochures and newsletters normally have between one and six pages. after reading information therein, an individual is free to pass it on to another person or the family for awareness as well as information sharing with neighbours. the library and resource centre the ubos library is open to both the public and staff from monday to friday, 8:00am-5:00pm. that means, it's able to entertain library users from within and outside. all ubos publications including those current and past, are found in our library. we ensure all the several displays that we have in the library can have our different products and also have some for take-away. we normally host between 200 and 300 guests per month. in the library, there is always someone to respond to questions and in event that the request is complicated, we have statisticians and directors from different directorates to give a hand. so, the library users go away with instant feedback from our staff. the library is privileged to host students, researchers and policy-makers like the legislators, ministers, executive directors, heads of ministries, departments and agencies. they are always excited to receive our data. all libraries with displays for their products, would be doing their organizations/institutions a great favour since the users of this information take it away for further reading and give-away. challenges ubos faces in disseminating data the challenge faced by ubos is delayed finalisation of the statistical products, i.e., reports, due to delayed receipt of funds from donors and government. failure to print them in time denies current data to the public. timely delivery of products is affected by delay in processing funds to disseminate the information. fuel for vehicles, allowances for officers to disseminate information in the districts and internet bundles are all affected by the delayed funds. the challenges to data dissemination include but not limited to: limited funding hence limited surveys conducted, people don’t value statistics when they start relating it to politics, motorcyclists request for a bribe in order to deliver or pass on information, lack of internet all over the country hence some intended persons are left out, if it is the only option to be used. https://doi.org/10.29173/iq991 4/4 nyangoma, judith (2021), data dissemination by uganda bureau of statistics, iassist quarterly 45(3-4), pp. 1-4. doi: https://doi.org/10.29173/iq991 recommendation to improve data dissemination there is need to have the whatsapp and twitter social media officially recognized in delivering and disseminating information to the targeted people or audience. the channels of word of mouth and motorcyclist sharing of information needs to be officially recognized because of their wide dissemination. strategies to improve data dissemination at ubos the strategic plan is to use data to make better and faster decisions to as many organisations and individuals, as much as possible. this calls for continued collection of online addresses so that it is posted accordingly. presenting our findings in a simple terms so that even the common person can apprehend it, is paramount for the benefit of all intended persons, ministries, departments and agencies. the continued use of various channels of communication would allow wider information dissemination, countrywide and to africa. there is need to widely disseminate the dissemination policy to stakeholders who are involved in data production so that they know what to expect or be done. let us make official the use of social media and public means as channels of data delivery because they are fast. finally, targeting the right people to whom information can be disseminated is important. for example, reports should be sent to institutions other than individuals after dissemination workshops. that way, the institutions would disseminate it as required. conclusion data dissemination is a worthwhile exercise when the beneficiaries utilise the data provided, and later on, continue to share it. references 1. the uganda bureau of statistics library and resource centre 2. the uganda bureau of statistics act no.12 of 1998 3. the uganda bureau of statistics' communication and public relations manual 2018 endnotes 1judith nyangoma is an information assistant at uganda bureau of statistics (ubos) and can be reached by email: joykuse@gmail.com https://doi.org/10.29173/iq991 mailto:joykuse@gmail.com 1/13 wiltshire, d., lichtwardt, b., & bishop, l. (2024) building human networks to drive forward innovations in international data access: introducing the international secure data facility professionals network (isdfpn), iassist quarterly 48(3), pp. 1-11. doi: https://doi.org/10.29173/iq1097 the creative commons-attribution-noncommercial license 4.0 international applies to all works published by iassist quarterly. authors will retain copyright of the work and full publishing rights. building human networks to drive forward innovations in international data access: introducing the international secure data facility professionals network (isdfpn) deborah wiltshire1, beate lichtwardt2, and libby bishop3 abstract the international secure data facility professionals network (isdfpn) was set up in 2022 as part of the social sciences and humanities open cloud project (sshoc) to bring together international colleagues working in or towards trusted research environments (tres), to share expertise and experiences, and to spark collaboration as well as develop new ideas. while various international networks and collaborations exist currently, these aim to improve the international data infrastructure landscape, and mainly focus on establishing connections between tres. however, there is also a forum needed for those working in these tres, which provides a platform for regular knowledge exchange as well as opportunities to drive innovation. isdfpn is not a network for developing infrastructure, rather it is a place to share experience, expertise, and ideas, which is not yet formally available internationally. as the fast-changing secure data landscape evolves, this network will be a vital resource for collaborative work towards finding solutions for shared and emerging problems experienced by tre staff. where tres are at different stages in their development, such a forum is vital as services seek to learn from each other. the isdfpn held its first virtual meeting on 30 march 2022, and, although the sshoc project ended in april 2022, the group has continued to meet twice a year, co-chaired by the uk data service, and gesis leibniz institute for the social sciences. this paper traces the isdfpn from its origins and highlights both its aims and objectives as well as its activities so far. keywords trusted research environments, tres, secure data facilities, international data access, secure access data, secure data, controlled data, professional networks introduction enabling safe, efficient, and impactful research on societal challenges is crucial to facilitate evidencebased policy development and effecting positive change for daily lives and livelihoods within and across countries. in order to provide the foundation for such essential research, increasingly more trusted research environments (tres) are enabling access to very detailed, sensitive microdata.4 it was the national statistical institutes (nsis) that took the lead in the early 2000s in setting up physical safe rooms to provide on-site access for researchers to their sensitive, potentially disclosive microdata. in the years that followed other entities followed suit in establishing safe rooms for onsite access to sensitive data, including several national data archives such as the uk data archive and https://doi.org/10.29173/iq1097 https://creativecommons.org/licenses/by-nc/4.0/ 2/13 wiltshire, d., lichtwardt, b., & bishop, l. (2024) building human networks to drive forward innovations in international data access: introducing the international secure data facility professionals network (isdfpn), iassist quarterly 48(3), pp. 1-11. doi: https://doi.org/10.29173/iq1097 microdata online access (mona) – statistics sweden, and research institutes such as the leibniz institute for educational trajectories (lifbi). on-site access has several disadvantages for researchers though, e.g. the expense of traveling to sometimes quite distant physical locations and the associated time constraints of short research visits (bishop et al, 2022). by 2011, work was well underway to move towards offering remote access to microdata, either via remote desktop systems (e.g. uk data archive) or via remote job execution (e.g. institute for labor market and occupational research of the federal employment agency research data center (iabfdz)). starting in 2013, the uk data archive has successfully enabled remote desktop access to its controlled (secure access) data via the ukds securelab, playing a longstanding part in the uk data provision landscape. this service broke new ground by removing the need to travel to a fixed location for accessing detailed microdata, and ‘secure remote access to data in the uk was born’ (welpton, 2021; scott, lichtwardt & woods, 2022). ukds securelab, initially called the secure data service (sds), was the first national secure data facility in the uk providing accredited researchers with secure remote access to controlled data, including a wealth of linked longitudinal datasets. following the ukds securelab provision and best practice, other tres in the uk – for instance, the office for national statistics’ (ons) secure research service (srs) – moved towards offering remote access. ukds securelab has therefore ‘provided the blueprint’ (welpton, 2021) for secure remote data provision in the uk and beyond. in germany, selected research data centres (rdcs) also enable remote access. for example, the research data center at the leibniz institute for educational trajectories (lifbi) prepares and disseminates survey data from the german national educational panel study. more sensitive versions of the data are made available via their remote desktop system ‘remoteneps’, whilst the most sensitive data remains accessible only via their onsite data security room in bamberg. gesis leibniz institute for the social sciences has provided access to sensitive data via its secure data center safe room, a physical data enclave, since 2013. in 2020, a new programme of work began to develop a new remote desktop access system. this will remove the need for researchers to travel, often some distance, to the safe room to access these data. it has long been recognised that enabling cross-border international access to sensitive data would be a positive move forwards in the international data access landscape and already in the 2000s there were a few initiatives to enable safe room remote desktop access across international borders (bishop et al, 2022; woollard et al, 2021). the holy grail was to find ways to enable researchers to visit a safe room in one country to remotely access sensitive data held by a tre in another country. a recent review of the secure data access landscape showed that progress in this area was initially quite slow (bishop, 2021), but one of the early success stories was the 2004 opening of an access point at the inter-university consortium for political and social research (icpsr) in the us to the iab-fdz in germany and the subsequent opening of access points in numerous safe room locations at tres both within germany and internationally5. in 2018, the international data access network (idan) was founded6 as a collaboration between six european research data centres, including casd (france), cbs (the netherlands), gesis (germany), iab (germany), the uk data service (uk), and ons (uk), forming a network to facilitate research use of secure access data via reciprocal provision of safe room remote desktop access. the most recent development was the setting up of a bilateral remote access connection between the ukds securelab and the secure data center at gesis, cologne as part of the social sciences and humanities open cloud (sshoc) project, a large eu collaboration of 47 organisations7 (wiltshire, voronin & lichtwardt, 2022; lichtwardt et al, 2022; lichtwardt & wiltshire, https://doi.org/10.29173/iq1097 3/13 wiltshire, d., lichtwardt, b., & bishop, l. (2024) building human networks to drive forward innovations in international data access: introducing the international secure data facility professionals network (isdfpn), iassist quarterly 48(3), pp. 1-11. doi: https://doi.org/10.29173/iq1097 2023). this marked a significant step forward in the drive towards opening international access to sensitive data. due to limited funding, legal barriers and other challenges, most of the infrastructure development for remote access has thus far occurred in a bottom-up fashion, mostly on an individual institutional level or else with small-scale bilateral collaborations. too often nsis have little incentive to invest in such collaborative infrastructure as this is usually not in their mandates, and national data archives generally have other priorities and limited resources to invest in the development of international data access infrastructure. nonetheless, there is a growing list of secure data facilities and corresponding demand from researchers, driving improvements. infrastructures must, of course, provide essential ‘plumbing’ hardware, software, platforms, resources, and so on. however, without adequate human support such as a full complement of staff, skills, training, etc., too often infrastructures are built but not adopted, embraced but not established, or started but not sustained. by nature of their specialist niche, tre staff tend to be widely distributed, sometimes it is only a single individual within a large institution. therefore, improving the human network landscape, especially internationally, is essential. even for well-established national tres, a substantial amount of work and expertise is required to enable safe room remote desktop access across borders. with the now rapidly evolving international secure access data landscape, professionals working in tres need a more structured mechanism for exchanging knowledge, learning from each other, as well as having access to and contributing to a one stop shop for relevant materials and resources (resource platform). within the tre sector, dedicated, experienced staff grapple with the challenges of facilitating access to and keeping safe a growing range of sensitive and potentially disclosive data, even more so with newly emerging challenges of new international sensitive data access possibilities. this work comes with a unique set of tasks that must be conducted within an often-complex legal governance landscape. it has long been recognised within the sector that little formal training or support is available to provide guidance and assistance for these roles. for new tre staff, training is primarily ‘on the job’ with support provided only by their colleagues. this approach can work in larger tres where new recruits have access to more experienced colleagues but falls short in smaller tres where there may be just one or two full-term team members. it also falls short in helping tre staff develop their skills and advance their careers. as the number of tres, and consequently the number of those working in tre settings, grew, a move started to change this ‘on the job’ approach to active support, with the belief that providing more substantial and sustained human support would also do a great deal to advance and support the developments of national as well as international secure data access. the sdap group an example of a national safe data access professionals network there are several examples of national safe data access professional networks both nationally and interationally (across europe and beyond). rdc-net, for example, aims to bring together participating tre partners from across germany8. multi-nation projects such as idan and sshoc bring together the international community. however, their primary focus lies on practical outcomes such as developing infrastructure or establishing cross-organisational connections rather than on providing support and development for the practitioners. https://doi.org/10.29173/iq1097 4/13 wiltshire, d., lichtwardt, b., & bishop, l. (2024) building human networks to drive forward innovations in international data access: introducing the international secure data facility professionals network (isdfpn), iassist quarterly 48(3), pp. 1-11. doi: https://doi.org/10.29173/iq1097 in the uk, a long-established network, the safe data access professionals (sdap) group which started as an informal gathering in 2011, provides training and support via a forum for data stewards working in tres across the uk9. early members came primarily from the consortium of organisations tasked with making office for national statistics data and esrc-funded longitudinal studies data available. the group's original aim was to support professionals working within this very specialised sector. throughout the evolution of the group, it has become more formalised, and its membership has grown substantially but its primary aim remains professionalising the work of data stewards working in tres and giving them a forum where they can exchange experiences and expertise with peers from across the sector. the group currently has around 50 members from across the uk. the group meets quarterly, with all activities and future planning being overseen by a steering committee. participation in this group has brought many benefits to its members, not least in providing a forum where they can talk about their work and where they can seek expert advice from others familiar with the regulatory landscape within which they work. the sdap group not only provides vital support to tre data stewards, but through its members it has produced several important deliverables focused in three main areas: researcher training, staff skills and competencies, and statistical disclosure control (sdc), that are widely utilised across the sector. 1. researcher training: in the uk it is often a mandatory requirement for researchers to receive some form of training prior to gaining approval to access secure data. even where this is not mandatory, many tres are now looking at developing their own researcher training. sdap members from cancer research uk and the health foundation developed a set of canonical training materials that can be freely downloaded and adapted by other services as needed10. this approach was adopted later by the sshoc project team who developed a new canonical set of training materials for the european tre sector (wiltshire, 2021; wiltshire, 2024). 2. staff skills and competencies: the area of staff skills and development is perceived as an important area to address as staff often come into the sector more by accident than design, and whilst they gain considerable skills through their work, the role of the secure data steward has not been professionalised. since 2016, the sdap group have been working to develop a competency framework11. the competency framework sets out the skills required for staff working in secure data facilities and can aid staff development as a way of setting objectives, identifying strengths and areas for improvement, performance management, and preparing for future roles within the sector. it also designed to assist tres with the process of recruiting new staff. 3. statistical disclosure control: in 2017, the sdap group began reviewing the available guidance on statistical disclosure control. the review concluded that the existing guides were by then some years old, and whilst the primary theoretical principles of sdc have changed little, new methodologies and data types had emerged. the result of this review was the publication by the sdap group of the sdc handbook in 2019, followed by its translation into spanish in 202012. the sdc handbook was widely applauded and is considered the main ‘go to’ guide for tre data stewards responsible for carrying out statistical disclosure control. the sdap group have their own website where these resources and others including presentation slides from previous meetings and events are made freely available to anyone interested in tres and sensitive data access and use13. the international secure data facility professionals network (isdfpn)with the expansion of secure data access possibilities across international borders via projects and collaborations such as idan and sshoc, it was recognised that a network would be highly beneficial https://doi.org/10.29173/iq1097 5/13 wiltshire, d., lichtwardt, b., & bishop, l. (2024) building human networks to drive forward innovations in international data access: introducing the international secure data facility professionals network (isdfpn), iassist quarterly 48(3), pp. 1-11. doi: https://doi.org/10.29173/iq1097 to those tasked with developing and running these internationally focused tres. secure data facility professionals working within these services are stepping into new, uncharted territory and as such, there is an emerging need to provide a space for secure data professionals internationally to meet one another, to exchange knowledge, and discuss pertinent issues arising from these new connections. until 2022 there was no formal intenational forum for secure data professionals that allowed sharing of experiences and expertise. similar to the the uk, there was no clear career path to train data stewards to work in tres. as more international bilateral connections are built, the establishment of an international forum became an urgent priority. the international secure data facility professionals network (isdfpn) was set up in 2022 as part of the social sciences and humanities open cloud project (sshoc) with the aim of bringing together international colleagues working in or towards trusted research environments (tres) and to provide support and collaborative opportunities to the international tre community (lichtwardt, wiltshire & bishop, 2022). the network is open to different disciplines including the social sciences, the health sector and humanities. isdfpn is unique in discussing issues not only in relation to quantitative secure data but also actively working on, for the first time, generating options for making qualitative sensitive data available in via secure data facilities in the future. the isdfpn group started as a formal structure. the international secure data facility professionals network (isdfpn) was set up as part of the social sciences and humanities open cloud (sshoc) project, which ran from january 2019 until april 2022. the uk data service has been leading the sshoc deliverable to setup and establish this network as a member of work package 5 (‘innovations in data access’), task 5.4 ‘remote access to sensitive data’. since the conclusion of the sshoc project, the network is stewarded through a collaboration between the uk data service (ukds secure lab) and gesis leibniz institute for the social sciences (secure data centre). prior to the first isdfpn meeting terms of reference (tor) were drafted so they are ready for discussion and comments during the first meeting14. the timing and frequency of meetings (steering group meetings, topical member meetings) were outlined in the tor, and the chair and secretariat were named. the tor also outlines the objectives and deliverables of the group, which are illustrated below. table 1 objectives and aims of the isdfpn objectives deliverables set the strategic direction for the isdfpn establish a steering group for isdfpn oversee the achievement of deliverables agree on an annual calendar of meetings and events for isdfpn members and the wider tre community establish isdfpn as an ongoing forum within an international context identify strategic needs and set up work streams with associated projects15 run topic-based networking and knowledge exchange events for both isdfpn members and the wider tre community agree and oversee the communications and digital strategy ensure collaboration with other professional groups where appropriate agree on a basic action plan (annual planner including all meetings, events, and work strands) https://doi.org/10.29173/iq1097 6/13 wiltshire, d., lichtwardt, b., & bishop, l. (2024) building human networks to drive forward innovations in international data access: introducing the international secure data facility professionals network (isdfpn), iassist quarterly 48(3), pp. 1-11. doi: https://doi.org/10.29173/iq1097 foster continued collaboration amongst isdfpn members and the wider tre community develop a community code of conduct for isdfpn members and for all external public facing forums e.g., events, social media platforms further work streams will be established as the group evolves and grows. key goals and deliverables may also change from year to year in response to changes in the international secure data access landscape. meetings are held bi-annually, and, for reasons of inclusivity, online. there are two presentations at each meeting, exploring and discussing issues related to quantitative and qualitative secure data. the isdfpn group held its inaugural meeting on 30 march 2022. the meeting was attended by twenty-two people from 13 different institutions, and 5 countries. a number of other people expressed interest in joining the network, even though they were unable to join the first meeting. the inaugural isdfpn meeting in the inaugural meeting attendees were asked to brainstorm ideas about current and future staff skills and staff training needs in tres. the following four staff skills and training questions had been placed in a collaborative document: • are there skills missing at this point in time? • what training is needed? does this exist, and, if so, where? • what skills might the tre professional of 2032 need? • other thoughts/comments? missing skills at this current time identified by participants included machine learning and ai; statistical software programming; anonymisation and pseudonymisation tools; synthetic data techniques; talking to data holders; workflows for ingesting data; research reproducibility; and knowledge regarding non-tabular data. when asked to look into the future and anticipate what skills the tre professional of 2032 might need, participants listed confidence, awareness of ethical complexities in outputs, and automation of low-level outputs as key priorities. in terms of training needs, the group acknowledged that some training resources exist already such as the output checker course offered by the dragon project, the safe researcher training (srt) provided by many uk tres, and fair data stewardship training. in addition, many data archives run annual summer schools which include courses on research data management. but it was felt that an overview of which tres offered remote access to sensitive data sources and of what training was currently on offer would be a useful exercise, and that certificates or other formal recognition for training participation would be an important step forward. having a system for formal recording or recognising training activities led to an expressed desire to see the status and pay of staff involved in key tre tasks such as statistical disclosure control (sdc) increase, and the lack of role professionalisation tackled, especially given the major legal implications of these tasks. further comments suggested a diversification of roles might be required within tres to reflect the increasing complexity of services involved. the second brainstorming exercise focused on gathering future topics that would be relevant to and of interest to the group as a precursor towards identifying possible work strands for the first few years https://doi.org/10.29173/iq1097 7/13 wiltshire, d., lichtwardt, b., & bishop, l. (2024) building human networks to drive forward innovations in international data access: introducing the international secure data facility professionals network (isdfpn), iassist quarterly 48(3), pp. 1-11. doi: https://doi.org/10.29173/iq1097 of the network and towards building a calendar of topics for the network meetings. this sparked a lively and engaged conversation with a long wishlist of topics that included among other things: an overview of existing tres/tre infrastructure, including provision of remote access systems an overview services involved in isdfpn and future directions technical infrastructures for secure data automatization of secure data facility tasks (e.g., output checking; self-administration platforms for users) public engagement and involvement legal challenges such as how gdpr supports remote secure data access authentication and authorisation procedures qualitative data in secure data facilities resources library or knowledge bank. whilst the list gave the network plenty to work with, it was clear from the discussions that the group wished to start with an inventory of what is available at present, and further discussions about what direction the group should take. this starting point was followed closely, according to the group, by topics where there is a need for consensus on best practice, such as technical infrastructure and solutions, legal questions and solutions to manage and scale up the daily workloads. events so far since the first meeting in march 2022, and after the end of sshoc, we had three more meetings/events all of which sparked great interest and very fruitful discussions. the topics presented in each of the meetings are detailed in the table below. table 2 isdfpn meetings themes and dates meeting date presentation 1 presentation 2 1 30.03.2022 international secure data facility professionals network (isdfpn) mind the skills gap: creating capacity for data access: compentancy framework 2 07.09.2022 certifiying reproducibility with confidential data: forst results from the french cascad/casd cooperation data sharing with rdc qualiservice 3 08.03.2023 enabling access to cenfidential qualitative data through data enclaves building a safe researcher accrediatation scheme 4 06.09.2023 researcher passport: a digital user credential for assessing restricted data introducing the safe points 5 17.04.2024 sane: an off-the-shelve, data holderagnostic tre output disclosure control for qualitative data in trusted research environments: current state and next steps https://doi.org/10.29173/iq1097 8/13 wiltshire, d., lichtwardt, b., & bishop, l. (2024) building human networks to drive forward innovations in international data access: introducing the international secure data facility professionals network (isdfpn), iassist quarterly 48(3), pp. 1-11. doi: https://doi.org/10.29173/iq1097 all of the topics are along the themes that attendees have mentioned in the initial brainstorming exercises. the presentations and discussions have been proven to be extremely valuable and have already led to new connections being made. meanwhile, in our fourth meeting, we have repeated the second part of our initial exercise and asked members to identify three topics that they would like to see covered in our isdfpn meetings in 2024. this way we ensure we know what the expectations of the group are, and are able to look for inspiration on these topics worldwide. equally, we are looking for new developments to share with the group, e.g., the development of the safepod network16 to offer also safepoints, a portable tre infrastructure system that could be shipped and installed anywhere in the world. we are evaluating what that could mean for the international community and trying to develp standards for adoption of technical solutions, enabling remote access much quicker, at a higher-standard. outlook the international tre community’s response to the first and all subsequent isdfpn meetings clearly demonstrates the need and desire for such a network. whilst the sshoc project was the original driver for setting up the network, the network has continued following the official end of sshoc in april 2022, under the joint stewardship of the uk data service and gesis-leibniz institute for the social sciences. the first deliverable, to appoint a steering group, has been completed and the group now meets bianually to drive the network forward and to oversee work on the remaining deliverables. steering group members come from across the two lead organisations and contribute their time on a voluntary basis. post-sshoc the network received no formal funding, so the secretariat has been provided from existing resources at the uk data service. a new eu funded project, eosc-entrust17, aimed at creating a european network of trusted research environments, will now provide support for the network until early 2027. this support will hopefully enable the network to increase the regularity of its meetings to 3-4 times per year, and to consider having a dedicated website. since its inaugural meeting in march 2022, the network has held whole group meetings twice a year, with two guest presentations each time covering topics such as proposals for a safe researcher accreditation system in germany, qualitative secure data sharing explorations , enabling reproducibility of findings based on secure access data, and other new developments such as the opportunities a safe researcher passport type scheme currently being proposed in several countries could offer18. whilst the network is still in its infancy, it has made a strong start in building constructive and collaborative connections across the global tre community that are vital for advancing work to facilitate international access to sensitive data. projects like idan and sshoc have highlighted the importance of having consensus over minimum requirements, comparable standards and procedures across tres for agreeing and implementing bilateral data access solutions. whilst idan, eoscentrust and other projects focus on infrastructure development and setting up connections, isdfpn will play a key role by providing a forum where discussions can occur, tres can share and exchange knowledge and drive new developments, such as qualitative data in secure data environments, while bridging the gap between different disciplines. the isdfpn chairs have been active in promoting the network through presentations at international conferences and through other networking activities and these activities play a key role in recruiting new members globally. membership in the network continues to grow; as of january 2024 it had 33 individual members from 25 organisations across 9 countries19. the network’s communication is managed via a dedicated jisc mailing list. isdfpn welcomes new members. joining is free, and open to all those who are involved in safe data facilities around the world. https://doi.org/10.29173/iq1097 9/13 wiltshire, d., lichtwardt, b., & bishop, l. (2024) building human networks to drive forward innovations in international data access: introducing the international secure data facility professionals network (isdfpn), iassist quarterly 48(3), pp. 1-11. doi: https://doi.org/10.29173/iq1097 references bishop, l. (2021). ’ms28 assessment of existing platforms (1.0)’. zenodo. https://doi.org/10.5281/zenodo.5914390 (accessed 28/07/2023). bishop, l., broeder, d., van den heuvel, d., kleiner b., lichtwardt, b., wiltshire, d. & voronin, y. (2022). ’d5.10 white paper on remote access to sensitive data in the social sciences and humanities: 2021 and beyond (1.0)’. zenodo. https://doi.org/10.5281/zenodo.6719121 (accessed 28/07/2023). international data access network (2023) idan – international data access network. available at: https://idan.network/ (accessed 29/09/2023). lichtwardt, b and wiltshire, d. (2023). ’crossing borders without leaving – sharing secure data internationally’. ukds data impact blog, 20 june 2023. https://blog.ukdataservice.ac.uk/sharingsecure-data/ (accessed 28/09/2023). lichtwardt, b., wiltshire, d. and bishop, l. (2022) ‘d5.12 international secure data facility professionals network (isdfpn)’. zenodo. https://doi.org/10.5281/zenodo.6583379 (accessed 03/07/2023). lichtwardt, b., woollard, m., wiltshire, d. & bishop, l. (2022). d5.11 eran pilot: setting up a secure remote connection between two trusted research environments (1,0). zenodo. https://doi.org/10.5281/zenodo.6676393 (accessed 15/08/2023). konsortswd (2023) secure data access point network – rdcnet. available at: rdcnet https://www.konsortswd.de/en/konsortswd/the-consortium/services/rdcnet/ (accessed 29/09/2023). scott, j., lichtwardt, b. and woods, c. (2022) ’uk data service securelab: pioneers in enabling safe data-driven research for over a decade’ ukds data impact blog, 12 january 2022. available at: https://blog.ukdataservice.ac.uk/securelab-ten-year-anniversary/ (accessed 03/09/2023). scott, j and woods, c. (2020) ’statistical disclosure control handbook now available in spanish. ukds data impact blog.’ ukds data impact blog, 23 july 2020. available at: https://blog.ukdataservice.ac.uk/statistical-disclosure-control-handbook-spanish/ (accessed 05/04/2022). safe data access professionals (2023) safe data access professionals: home. available at: https://securedatagroup.org/ (accessed 29/09/2023). welpton, r. (2021) ’celebrating 10 years of secure remote access in the uk’. ukds data impact blog, 12 october 2021. available at: https://blog.ukdataservice.ac.uk/ten-years-secure-remote-access/ (accessed 29/09/2023). wiltshire, d. (2021) ’d5.20 training materials of workshop for secure data facility professionals (v1.0)’. zenodo. https://doi.org/10.5281/zenodo.5638596 (accessed 03/07/2023). wiltshire, d. (2024). developing canonical ‘safe researcher’ training materials for trusted research environments. iassist quarterly 48 (1). https://doi.org/10.29173/iq1093/. https://doi.org/10.29173/iq1097 https://doi.org/10.5281/zenodo.5914390 https://doi.org/10.5281/zenodo.6719121 https://idan.network/ https://blog.ukdataservice.ac.uk/sharing-secure-data/ https://blog.ukdataservice.ac.uk/sharing-secure-data/ https://doi.org/10.5281/zenodo.6583379 https://doi.org/10.5281/zenodo.6676393 https://www.konsortswd.de/en/konsortswd/the-consortium/services/rdcnet/ https://blog.ukdataservice.ac.uk/securelab-ten-year-anniversary/ https://blog.ukdataservice.ac.uk/statistical-disclosure-control-handbook-spanish/ https://securedatagroup.org/ https://blog.ukdataservice.ac.uk/ten-years-secure-remote-access/ https://doi.org/10.5281/zenodo.5638596 https://doi.org/10.29173/iq1093 10/13 wiltshire, d., lichtwardt, b., & bishop, l. (2024) building human networks to drive forward innovations in international data access: introducing the international secure data facility professionals network (isdfpn), iassist quarterly 48(3), pp. 1-11. doi: https://doi.org/10.29173/iq1097 wiltshire, d, voronin, y. and lichtwardt, b (2022) ‘m29 tested connections between partners with live data and researcher projects (v1.0)’. zenodo. https://zenodo.org/record/7684215#.y_3pnxbmjpz (accessed 03/07/2023). woollard, m., lichtwardt, b,. bishop, l and müller, d. (2021) ’d5.9 framework and contract for international data use agreements on remote access to confidential data (v1.0)’. zenodo. https://doi.org/10.5281/zenodo.4534286 (accessed 03/07/2023). https://doi.org/10.29173/iq1097 https://zenodo.org/record/7684215#.y_3pnxbmjpz https://doi.org/10.5281/zenodo.4534286 11/13 wiltshire, d., lichtwardt, b., & bishop, l. (2024) building human networks to drive forward innovations in international data access: introducing the international secure data facility professionals network (isdfpn), iassist quarterly 48(3), pp. 1-11. doi: https://doi.org/10.29173/iq1097 appendix 1: isdfpn terms of reference https://doi.org/10.29173/iq1097 12/13 wiltshire, d., lichtwardt, b., & bishop, l. (2024) building human networks to drive forward innovations in international data access: introducing the international secure data facility professionals network (isdfpn), iassist quarterly 48(3), pp. 1-11. doi: https://doi.org/10.29173/iq1097 https://doi.org/10.29173/iq1097 13/13 wiltshire, d., lichtwardt, b., & bishop, l. (2024) building human networks to drive forward innovations in international data access: introducing the international secure data facility professionals network (isdfpn), iassist quarterly 48(3), pp. 1-11. doi: https://doi.org/10.29173/iq1097 1 deborah wiltshire (corresponding author); gesis-leibniz institute for the social sciences, unter sachsenhausen 6, 50667 cologne, germany; deborah.wiltshire@gesis.org; orcid 0000-0001-65332426. 2 beate lichtwardt; uk data service, university of essex, wivenhoe park, colchester, essex, co4 3sq, united kingdom; blicht@essex.ac.uk; orcid 0009-0006-5304-5634. 3 libby bishop; gesis-leibniz institute for the social sciences, unter sachsenhausen 6, 50667 cologne, germany; elizabethlea.bishop@gesis.org 4 increasingly, personal/confidential, and sensitive data are made available through secure data facilities which can be also referred to as secure access facilities/ secure research facilities/ safe settings/ or trusted research environments (tres). examples for these are a) research data centres (rdcs), e.g., the iab fdz, rdc lifbi, rdc soep, b) datalabs, such as the ukds securelab, hmrc datalab, ons srs, justice data lab etc., and c) data safe havens, to name just a few. 5 https://fdz.iab.de/en/about-us/appointment-locations-and-fdz-online-calendar/ [accessed 05/04/2022] 6 https://idan.network [accessed 05/04/2022] 7 https://sshopencloud.eu [accessed 08/04/2024] 8 https://www.konsortswd.de/en/konsortswd/the-consortium/services/rdcnet/ [accessed 05/04/2022] 9 https://securedatagroup.org/ [accessed 03.08.2023] 10 https://securedatagroup.org/training2/ [accessed 05/04/2022] 11 https://securedatagroup.files.wordpress.com/2018/07/sdap_competency_framework-01_00.pdf accessed 05/04/2022] 12 https://securedatagroup.org/sdc-handbook . some of the authors discussed the reception of the sdc handbook in this blog published on the ukds website. scott, j and woods, c. (2020) ’statistical disclosure control handbook now available in spanish. ukds data impact blog.’ ukds data impact blog, 23 july 2020. available at: https://blog.ukdataservice.ac.uk/statistical-disclosure-controlhandbook-spanish/ (accessed 05/04/2022). 13 https://securedatagroup.org/events/ [accessed 05/04/2022] 14 please see appendix 1 for the complete tor draft. 15 projects within these work strands may be led by a steering group member, or by an isdfpn group member, with involvement from a steering group member. 16 https://safepodnetwork.ac.uk/ provides portable safe settings for secure data access across the uk [accessed 08/04/2024] 17 home | european network of trusted research environments (eosc-entrust) project [accessed 10/04/2024] 18 a safe researcher passport scheme would provide an electronic record of a researchers affliation and training status that could be used by tres and data access organisations to carry out authenication and authorisation checks. one example is the icpsr researcher passport: https://radius.icpsr.umich.edu/radius/passport/static/about [accessed 08/04/2024]. 19 if you would like to join the network, offer a talk or request more information regarding the network, please email isdfpn@ukdataservice.ac.uk. https://doi.org/10.29173/iq1097 mailto:deborah.wiltshire@gesis.org mailto:blicht@essex.ac.uk mailto:elizabethlea.bishop@gesis.org https://fdz.iab.de/en/about-us/appointment-locations-and-fdz-online-calendar/ https://idan.network/ https://sshopencloud.eu/ https://www.konsortswd.de/en/konsortswd/the-consortium/services/rdcnet/ https://securedatagroup.org/ https://securedatagroup.org/training2/ https://securedatagroup.files.wordpress.com/2018/07/sdap_competency_framework-01_00.pdf https://securedatagroup.org/sdc-handbook https://blog.ukdataservice.ac.uk/statistical-disclosure-control-handbook-spanish/ https://blog.ukdataservice.ac.uk/statistical-disclosure-control-handbook-spanish/ https://securedatagroup.org/events/ https://safepodnetwork.ac.uk/ https://eosc-entrust.eu/ https://radius.icpsr.umich.edu/radius/passport/static/about mailto:isdfpn@ukdataservice.ac.uk a profile of data preservation activities in university data libraries and archives alice robbin data and program library service university of wisconsin-madison i. introduction for the last two decades, funding priorities have dictated allocation of resources to national centers as principal sources of archival data for research and teaching. however, the importance of local centers in universities for preserving, disseminating, and describing computer-readable data cannot be understated. these local, campus-based libraries and archives play a critical role in the information transfer system for the social scientific community, both within and outside their institutions. even though technological developments make it possible to receive and transmit data from great distances (thereby obviating in principle the need for being a local repository), current economic and political realities have constrained efficient use of the modern computer technology. these realities suggest that distributed data centers at the local level will continue to play a major role in the transmission of information. it is worthwhile briefly to describe the important role these local data libraries have played in the last fifteen years, and to suggest how they will continue to participate in social inquiry. a number of the centers were established before their national governments established machinereadable archives divisions within the national archives. as a result they became de facto repositories for federally-produced data. for many studies obtained from outside their institutions, no adequate archival facility existed elsewhere, and the centers became by default permanent repositories. in other cases, although these centers had not been designated repositories for federally-produced machine-readable data, federal agencies turned to them for assistance in retrieving data files which the agencies produced but could no longer retrieve. the data centers also acted as a transfer agent and depository for data files produced by foreign governmental agencies and research institutes. for other studies which could theoretically be obtained elsewhere, some data centers assumed archival responsibility on the grounds that the supplier could not adequately preserve and maintain the valuable resource, or that it was economical in time and money for the university data center to maintain an on-site copy of the data. there are of course other reasons for preserving data and for supporting a system of distributed data centers within a national and international context. a growing literature (cf. clubb, hofferbert, miller, rokkan, boruch and wortman, nesvold) makes cogent arguments for supporting the national data centers, data libraries, and data laboratories for social scientific research and teaching. the underlying philosophy of preservation and access holds that transfer of the data collections from private research organizations, governmental agencies, and foundations to these data centers has greatly magnified the return on the original private and public investment in data gathering, and has encouraged and facilitated social scientific inquiry. the data archive participates in the processes of innovation, dissemination of scientific results, and information transfer. it acts as a scientific laboratory which encourages the sharing of data, multi disciplinary exploitation of evidence, and "multiple and complex analytical applications" (hofferbert and clubb, p. 383). the data center makes a pedagogical contribution by allowing the student to participate directly in empirical scientific inquiry, developing problem-solving modes of behavior like those of students in the natural sciences. less obviously, the data archive plays a role as an agent for administrative and technical assessment of information transfer activities and mechanisms. it offers administrators and researchers the opportunity to assess, in a rigorous and analytical fashion, the technical, administrative, economic, and policy issues related to standards of data quality, documentation, access, and diffusion. collegial behavior is facilitated. common access encourages standardization and "commonality of research among widely separated scholars" (miller, p. 411). finally, a democratic society such as ours is committed to public access to the products of research and to knowledge-producing modes of behavior, and it is through the data centers that this access is facilitated. rockwell argues that in the 1980s these data centers will assume greater importance for the social sciences than they have in the past, because of such factors as the increasing cost of gathering data and conducting surveys. we can expect that in the 1980s social scientists will capitalize on resources such as those found in these local data cen ers. for example, the increasing number uf time series of replicated data is becoming a major resource for cohort and panel analysis. rockwell suggests that it is unlikely that we will see many new surveys mounted during the coming decade and that "from the perspective of social indicators research, the resources of these data centers are important precisely to preserve long time series of data." he gives the following reasons: "more generally, cumulative social science research demands ample opportunity to return to the same data bases for repeated inquiries. the journals increasingly reflect the field's recognition of the importance of good social measurement: standardized measures available on a repeated basis in a time series data base." some of the data centers now face critical problems preserving data on magnetic tape, the principal medium of long-term storage, because of changes in computer technology, aging of the collection, and magnetic tape deterioration. the data and program library service at the university of wisconsin, for example, has found that a growing number of tapes, even ones guaranteed for 15 or 20 years, are developing non-recoverable read-write errors. these errors are seriously affecting the quality of the data preserved. the problems of physical deterioration are not limited to elderly reels of tape. dpls has encountered quality control and deterioration problems with tapes purchased two and three years ago from a manufacturer with whom dpls has dealt for some years. in addition, dpls has found, as have others, that the computing center's tape drives have influenced the condition of dpls magnetic tapes, and have been responsible for parity error problems. in response to these and other problems, including changes in computer technology, dpls instituted a minimal tape maintenance program three years ago to convert data sets written on older tapes. when dpls began calculating the costs of a full-scale tape maintenance program, it became obvious that the magnitude of its collection, the staff time required to rectify the problems, and the computing and capital equipment expenditures required were beyond dpls's limited resources. it was at this point that the dpls staff began an investigation into current and potential mass storage media developments and decided to conduct a survey of data centers to ascertain how their staffs are handling data stored on magnetic tape. dpls thought that the literature might provide some insights into developing its own program of tape maintenance and would also provide the data library and archival community with information on the status of data preservation. the following discussion is a report on technical problems related to preserving machinereadable records, and on the findings from the small survey. although some may believe that data stored on magnetic media can be treated like a book left on a shelf to gather dust, that in fact is not the case. problems discussed in part ii of this paper suggest the need for accelerating the development of long-term archival storage media now in the experimental stage, particularly because of the increasing generation of statistical data. the problems facing the north american data libraries and archives, and how their staffs are coping with the preservation of data stored on magnetic tape, are taken up in part iii. the survey results suggest that data library and archive staffs recognize that preserving their valuable resources is necessary to ensure continuing support to the research and teaching community. all have either developed formal tape maintenance programs or are aware of the need to develop better practices for preserving their collections. ii. technical problems associated with preserving machine-readable records among the major problems faced by the librarian and archivist who deal with computerized records is that current magnetic storage medial and most of the mass storage devices2 now in various stages of development fail to meet archival storage requirements— that is to say, the preservation of digital data for a very long period. problems associated with permanent preservation of data include the physical size of data, machine independence and media standardization, reliability of the storage medium, the medium's sensitivity to environmental conditions, lifetime maintenance and cost of the medium, accessibility of the information, and cost of duplication. volz, dollar, and geller elaborate on these problems from the archivist's perspective, and their comments bear repeating. to convey the problem of physical size of data, volz presents this example: a typical book contains about 10 to the 7th bits. (i interpret him to mean bytes or characters rather than bits. --author) encyclopedia brit annica contains about 10 to the 9th bits. such volumes can be readily stored in today's technology. for example, britannica would roughly fit on a single ibm 2220 disk pack. however, the problems are not storage of a single volume of text, but rather large collections of such volumes. the dpls collection contains many such "volumes." for example, one data file in the collection contains approximately 545 million characters. and although most files do not approach this size, the dpls collection contains more than 6000 data files. in the last year, data files which fill two or more 2400foot magnetic tape reels have become the norm. we expect that with the 1980 u.s. census of population and housing , average data files will be stored on multiple reels of tape. increasingly as scholars turn to administrative records for research, data collections will require adequate storage devices. with the exception of magnetic tape, which offers compatibility when written on different tape drives if the same density and character codes are used and if utility software is available to translate the character codes (dollar, p. 29), all other magnetic storage media (to my knowledge) are machine-dependent and non-standard. (for example, cassette tapes produced by the sykes corporation cannot be used on ibm equipment.) without machine independence and media standardization, archival storage becomes a very nearly unsolvable problem in compatibility. machine dependence also affects preservation of data in another way: system files written on one machine cannot be transferred to another computer system (e.g., spss system files produced and transferred between computers of the same manufacturer may not be readable). preservation is also affected by machine obsolescence. the rapidly changing computer technology has resulted in removal of equipment used for the initial creation and copying of the data. thus, the data archives created in the 1960s, when magnetic tape was read and written on 7-track tape drives, have found that their computing centers have replaced their equipment, and that their data cannot be read on the new equipment. the result is reduced access to their collections and increased preservation costs, b. cause all their data must be converted to meet the specification of the new equipment. another archival concern is how long the media retain a reliable image of the data. volz notes that "due to the relatively short span of time over which really large mass memory devices have been in use, only limited empirical data is available." dollar comments (p. 29): permanent preservation of digital data requires storing the records in a mode in which under normal conditions the recording signal will not degrade and the medium will not deteriorate to the point that data recovery is impossible... this means a non-erasable mass storage capability which is not vulnerable to irreparable loss of archival records through human carelessness or system malfunctions. geller, manager of the magnetic media group at the national bureau of standards, elaborates (pp. 37-38): "experimental evidence has shown that failures to extract the information from magnetic media are aimost always attributable to the physical deterioration of the media rather than to the deterioration of the data." although there are now estimates of an archival lifetime of 10 years for magnetic tape, there is really no way to simulate the reliability of a magnetic tape as a storage medium for a long period of time. most archival data, our records indicate, are accessed infrequently (every few years) or not at all. thus, without a regular and frequent maintenance program, tape deterioration is not apparent until the data are requested, accessed, and copied. archives in existence since the 1960s face aging problems associated with the quality of the magnetic storage medium. tapes produced before 1972 cannot be reused because of the deterioration resulting from the poorer magnetic tape. archives whose holdings date from before 1972 may need to replace large parts of their magnetic tape collections. 3 environmental and handling conditions affect the lifetime of the magnetically encoded data and necessitate expensive environmental controls to prevent adverse forces and "debilitating humidity and temperature conditions" (dollar, p. 29) from affecting the recording signal and from impairing the storage medium. the magnetic tape on which all data libraries store their collections is well known to be susceptible to environmental conditions. for example, a 2400-foot tape will "try to change its length by approximately one foot for every 10 degree change in temperature or 10 percent change in humidity," volz states. he continues: friction of the tape wound upon a reel tends to prevent these changes in length from taking place, resulting in high pressures on the tape and perhaps some permanent changes to the tape. occasionally some slippage may occur resulting in flaking of the oxide from the surface of the tape, which not only may itself lose information, but creates debris which will interfere with the reading of other bits. lifetime maintenance and media costs are considerable. to ensure accessibility to the stored information, tapes must be duplicated so that the archive always retains a reliable image of the data (i.e., so that unrecoverable errors on one file do not result in irretrievable loss of data and a file's integrity is assured by maintaining a second copy. for an archive of record, costs of maintenance and preservation can be significant: if the original data file has gone through several data processing activ ties (updates, corrections) over time, all the file iterations must be maintained. data files stored on magnetic tape must be "rolled over" (i.e., copied) at least once every two or three years. this entails a considerable allocation of resources for an archive. new magnetic tape must be purchased and computer time must be "bought for copying; staff time must be available for carrying out the maintenance program—preparing the software, documenting the procedures, evaluating results of magnetic tape quality, and completing the administrative records to document output onto the new storage medium. to the extent that documenting and administrative record-keeping can be automated, human resource savings can be significant, since it is the record-keeping activities which are labor-intensive. what this discussion suggests is that the social scientific research activity requires adequate funding to maintain necessary supporting facilities. the laboratory of the social scientist requires modern and reliable mass storage equipment for long-term preservation of the materials used for scientific discovery. maintenance, while perhaps more visible in a natural sciences laboratory or a traditional library (where there are devices for controlling humidity, facilities for rebinding books, and programs for the security and physical protection of the collection), is a necessary condition for social scientific activities. hi. the survey between january and march 1980, dpls conducted a mailout-mailback survey on tape maintenance activities in data centers (libraries and archives) located in north america. three sources of information were used to identify these centers. the list of data centers provided in ss data: a newsletter of archival acquisitions was supplemented by a review of all the catalogues of data holdings at dpls and of the dpls administrative correspondence files. with the exception of one survey research data archive (which was identified in ss data ), survey research institutes were systematically excluded, as were governmental archives (e.g., the u.s. national archives and records service and the public archives of canada), national and international repositories such as the inter-university consortium for political and social research and the roper public opinion research center, and federal agencies, such as the u.s. bureau of the census and statistics canada, which disseminate data to the social research community. icpsr member institutions which provide access to icpsr data through a departmental faculty member were also excluded because most of those departments would not qualify as a data center or library/archive in any rigorous way, particularly because control over their materials is lacking and because they play a minimal or nonexistent role in information dissemination about data for other than the consortium's holdings. 4 (this statement is of course offered without any hard evidence, and needs verification.) the questionnaire was sent to 37 organizations, of which 34 had responded by the end of march 1980. after reviewing the completed questionnaires, four data centers were deleted from the final sample. either most of the items in the questionnaire were not relevant to their organization, or their holdings were so specialized that the information we sought could not be utilized in our analysis, or they were not a university-affiliated organization. only one university-affiliated data center did not respond. the final sample on which our analysis is based is 30 university data libraries and archives. 5 because we cannot say with any assurance that our original list constituted the universe of data libraries and archives in north america, our review of tape maintenance activities offers no tests of statistical significance. rather, our intention here is to describe current data library maintenance activities and to present a profile of these activities in a select group of data centers. we need to probe more deedlv into the state of these data organizations to understand how they are structured, what activities they carry out, and their influence on social scientific activity at their institutions. these are all important questions for which we have little or no information. but certainly these questions are worth pursuing, for they add another dimension to what we know of how organizations charged with information transfer participate in the knowledge flow process. the very high response rate and the enthusiasm with which people responded is evidence that these staffs do want more information about the problems of their colleagues and how they are coping with current economic and political realities. a. a profile of north american data centers in an effort to reduce respondents' reporting burden, questions about their organizational structure, activities other than tape maintenance, funding, and collection were kept to a minimum. we wanted to know when the center was established and the estimated size of the data and tape libraries. we posited that an early establishment date and a large collection would lead to inadequate levels of funding for purchase of magnetic tapes and for maintenance activities. it would lead also to dissatisfaction with the quality of the maintenance program. we were interested in knowing whether there were any differences in maintenance activities if a data center were an independent department or affiliated with another department, library, research organization, or computing center. we wanted to know from where the data collection was derived, that is, its original sources; its estimated growth in data files and magnetic tapes over the next five years; how the data were used; and whether the staff had noted any changes in the number of requests and in types of files requested in the last two years. on the basis of our services at dpls, we have noted an increasing tendency toward use of government-produced data and toward larger and more complex files requiring at least several reels of magnetic tape. until rather recently, dpls served as a research support facility, and undergraduate class projects have constituted no more than 15 to 20 percent of our use. we wondered how different or similar the situation was at other data centers. we thought that if staffs were noting changes in the number of data files and types of data being requested, this could signal the growing complexity of the data being used by members of their institutions, and of increasing demands being placed on the library staff. neither the questions nor the responses permit us to infer what is happening at the local level, although we can make some educated guesses. concerning the computer facility available to the data center, we wanted to know what computer is primarily used for most activities. we wanted to know how the data center stores its data and the current storage mode on magnetic tape. we then turned our attention to whether the organization was encountering or anticipated tape storage problems and whether the staff had investigated any ways other than magnetic tape for transfer and long-term storage. our last set of questions concerned the burden of changes made to the computing center, and the adequacy of financing to preserve the integrity of the collection and carry out maintenance activities. figure 1 shows that 17 of 27 data centers, or 63 percent, were established between 1966 and 1972.6 these years correspond to a period when universities and external funding agencies provided increased financial support to the social sciences. between 1973 and 1977, we see a decline in the number of data centers being established; but in 1977 we once again see an increase. figure 2 shows that more than half of the data centers (n=17) have collections of between 100 and 699 data files. the size of the data collection appears to have little relationship to when the center was founded. see table 1, why is unknown; it may have something to do with the size of the user community, as well as the resources available for collection building. figure 3 shows the size of the magnetic tape collection. here we see that 18 of 30 centers have fewer than 399 magnetic tapes. figure 1. date of establishment. figure 3. number of reels. 1963 1 1964 1 1965 1966 1 1967 3 1968 4 1969 1970 2 1971 3 1972 4 1973 2 1974 1 1975 1 1976 1 1977 3 27 figure 2, estimated size of data collection. 100-299 6 ****** 300-499 6 ****** 500-699 5 ***** 700-899 900-1099 2 + * 1100-1299 2 ** 1300-1499 1 * 1500-1699 2 ** 1700-1899 1900-2099 2100-2299 1 * 2300-2499 1 * 2500-2699 2700+ 3 *** -100-199 9 ********* 200-399 9 ********* 400-599 4 **** 600-799 2 ** 800-999 1000+ 6 30 ****** 29 table 1. date established by size of collection. size of collection (number of reels] 100-899 900-2999 3000+ 19731968-72 -1967 16 date established 10 when we look more closely at the relationship between the number of data files and number of magnetic reels, we see some connection, but at the same time we see that storage conditions vary. several data centers utilize modern storage technology to pack large amounts of data on a small number of tapes, while others store their data at much lower densities. see table 2. we asked the staffs to estimate the growth in the number of data files and magnetic tapes per year over the next few years. for those centers which supplied this information, estimated growth in the number of files per year was the following: 32 percent (n=7) estimated between 10 and 30 files; 64 percent (n=14), 31 to 75 files (a large spread); and four percent (n=l), 150 or more. estimated growth in the number of magnetic tapes, as expected with the advance in storage technology, was 43 percent (n=9), between two and 25; 48 percent (n=10), between 30 and 80; and nine percent (n=2), between 100 and 125. the next set of questions dealt with the sources of the data in the collection, how the collection was used, and whether there have been changes in the types of requests made to the staff. by far the largest source of data is the inter-university consortium for political and social research (icpsr). of 29 data centers reporting sources of data, 45 percent (n=13) report having up to 59 percent of the collection from icpsr, while 55 percent report between 60 and 100 percent icpsr materials. the average was about 65 percent. surprisingly, acquisitions from the federal government are very low; 79 percent (n=23) report between and 25 percent of their holdings from federal sources. because social scientists are increasingly consuming federally-produced data, we expected a greater percentage of the centers' collections to be from the government. of course, it is quite possible that the centers are obtaining federal data from icpsr and then reporting icpsr rather than the government as the supplier. not surprisingly, the private sector accounts for an insignificant percentage of collections: 78 percent of the centers report between and 5 percent. the center's own institution, the roper center, and other distributors make up only a small fraction of the remaining suppliers: 83 percent report no more than 15 percent from their own institutions; 83 percent have between and five percent from the roper center; and 79 percent get from to 10 percent from other sources, primarily international and intergovernmental. data centers have typically been the product of research activity at an academic institution. as new generations of graduate students trained in quantitative methods and data handling enter the teaching profession, quantitative methods and the use of the computer are introduced into the classrooms. considering that data handling was the purview of sophisticated graduate students during the middle and late sixties, we should expect a third generation of former graduate students now to be faculty members and a data center to be responsive to their teaching needs. we therefore expect that instructional use of the data center, as a laboratory for scientific activity, will constitute a significant part of its over-all use. table 3 shows the use of the collection for teaching, research, and other (primarily policy) activities. here we see that research is indeed the principal reason for the use of the data center (mean=65 percent), but that instructional use does represent a significant activity (mean=33 percent), we wondered whether staffs had noted any changes in the number of requests and types of data files requested during the last two years. we asked whether there were any increases in the number of requests, whether the files were structurally more complex, and whether files were requiring more than one or two reels table 2. size of data collection (number of data files) by size of magnetic tape collection (number of tapes). number of tapes -100-399 400-999 1000+ 1700+ 2 1 2 5 size of data 700-1699 7 1 8 collection 100-699 8 5 2 15 17 6 5 28 table 3. percentage of collection used for teaching and research. instruction research other % n 0-20 13 (45*) 25-50 11 (38%) 60-90 _5 (17%) 29 mean = 33% ; n 10 40 6 (21%) 50 80 16 (55%) 90 100 7 29 (24%) 5 33 100 65'. n 26 1 1 j_ 29 (90%) ( 3%) ( 3%) ( 3%) table 4. changes in the types of requests and data files over past two years. increase in number of requests files more complex files require more than 1 or 2 reels of magnetic tape increase in file requests and data files more complex increase in file requests and files require more reels increase in requests and files more complex and require more reels files more complex and require more reels no changes noted not ascertained 5 (17%) 4 (13%) 2 ( 7%) 1 ( 3%) 1 ( 3%) 9 (30%) 3 (10%) 2 ( 7%) _3 (10%) 30 12 of magnetic tape. the increase in the number of requests and growing complexity of the data were the two largest single categories of changes noted during the last two years; 30 percent of the respondents noted all three changes. the next series of questions reports on media storage of the collection, current and anticipated storage problems, and investigation of other storage media. as table 5 indicates, magnetic tape is the medium of storage for the the data centers, with 29 of 30 centers storing between 95 and 100 percent of their data on magnetic tape. the current storage modes appear to be ebcdic, nine channel, 1600 bpi, althouqh there are still data centers (almost a third) which store their data in seven channel, even parity, 556 bpi, and an increasing number of centers which are moving to ebcdic, nine channel, 6250 bpi. 7 more than half the data centers said that they had no tape storage problem now (n=16, 53 percent), while 47 percent (n :: 14) reported problems. when asked what kinds of storage problems they were encountering, almost half gave lack of space as the principal one. table 6 indicates a pattern to the tape storage problem: too many tapes, lack of space (usually associated with onsite storage rather than off-site), leading to off-site (or remote) storage as a necessity. ° data centers located in computing centers and affiliated with libraries indicated no problems, whereas those which were independent or affiliated with research organizations appear to be encountering storage problems. in response to the question about whether they anticipate a storage problem in the future, 62 percent (n=18) responded yes, and 39 percent (n=ll), no. of the 18 who anticipate problems, the need for storage (whether onor offsite), storage costs, and the large collection were cited as major problems. 9 we wondered whether any data centers had investigated ways other than magnetic tape for transfer and long-term storage of their data collection. 37 percent (n=ll) had, whereas 63 percent (n=19) had not. for those who had investigated other media, off-line disk was cited by five, video disk by two, microfilm by one, and computer by two. finally, we were interested in knowing whether those who had noted a tape storage problem had also investigated other media for long-term preservation. as our table 7 indicates, 57 percent (n=8) of the 14 responding that there were tape storage problems had not done any investigating, while 31 percent of those indicating no tape storage problem had investigated other ways of storing data. the last set of questions explores the impact of changes in computer technology, of increased requests for data files, and of the adequacy of funding for tape purchase and maintenance activities. we asked whether the computing center had made or planned to make changes which have affected or would affect the way in which the data center stored its data. almost half (47 percent) noted that the computing center had made changes but the changes did not affect the way the data were stored; however 33 percent did note that seven channel tape drives were being phased out and that the data center must convert its data collection. in response to the question about whether changes in the types of requests being made for data had placed a burden on the library in terms of available budgetary resources to maintain the physical integrity of the collection, 73 percent (n=22) responded no, while 27 percent 13 table 5. media storage of the collection. cards tapes disk % n % n % n 22 (73%) 75 1 ( 3%) 25 (83%) ] 2 ( 7%) 95 5 (17%) 1 1 ( 3%) 2 3 (10%) 98 2 ( 7%) 3 1 ( 3%) 5 2 ( 7%) 99 3 (10%) 5 2 ( 7%) 25 1 30 ( 3%) 100 19 30 (63%) 40 1 30 ( 3%) table 6. tape storage problems. too many tapes storage charges not enough space lack of environmental controls unused files combined with active files 1 ( 7%) backup charges prohibitive remote storage necessary 2 (14%) age j. ( 7% ) 14 first problem second problem 3 (21%) 1 ( 7%) 1 (33%) 6 (43%) 1 (33%) es 1 ( 7%) 1 (33%) table 7. whether tape storage problem exists by whether data center has investigated other storage media investigated other storage media yes no 14 tape storage problems yes no 6 (43%) 8 (57%) 5 (31%) 11 (69%) 11 19 16 30 14 (n=8) said yes. while 90 percent (n=26) claimed adequate funding for tape, only 66 percent (n=19) said they could afford a tape maintenance program. for the 10 centers wtiich said that funds were inadequate, major reasons cited were that there was not enough staff (30 percent) or that both funding and staffing were inadequate (30 percent). 10 the data centers affiliated with a computing center and a research institute have no difficulty in supporting a tape maintenance program, while more than half the centers which are independent departments or affiliated with a teaching department cite inadequate funds to maintain such a program. our last question on financing asked where financial support came from. the data library budget is the source for maintenance activities for 41 percent (n=12); computing centers account for 17 percent (n=5); a mix of library and computing center for 10 percent (n=3); the data library budget and ad hoc requests for maintenance funding, 10 percent (n=3). the remaining percentage was divided among ad hoc requests, other, no support provided, and a mix of data library, computing center, and ad hoc funding. b. tape maintenance activities in this section we examine the quality of tape maintenance activities, degree of satisfaction with the data center's program, and whether there is any difference in the quality of activities between those who are satisfied and those who are dissatisfied with their program. our concern here is with the set of activities to preserve data on magnetic tape. according to the literature, tape maintenance involves creating back-ups, controlling the movement of the magnetic medium from abrupt environmental changes, maintaining environmental controls ( j < emperature and humidity) in the storage area(s), monitoring these controls periodically to observe changes, having access to an off-site facility for storage, and maintaining a record keeping system for effective control and administration of the tape library. what we observe in tables 8 and 9 is that data centers can clearly be given high marks for protecting their data by maintaining back-ups of every data file and controlling the movement of the medium; but their monitoring of environmental controls, providing off-site storage for the data, and maintaining complete evaluation histories of the magnetic tapes are not as good as they should be. considering that data centers are transferring data from supplier to data center and data center to computing facility, fully 66 percent are not letting their magnetic tapes sit for at least 24 hours before mounting them. this can result in too much stress on the medium, cracking, and data destruction. while 70 percent say they have environmental controls in the storage area(s), only 33 percent say they monitor the controls. unless data centers can guarantee full protection (against loss) of their master and backup copies, off-site storage of at least one copy is a requisite for preservation. yet only 57 percent (n=17) say they have access to off-site storage. sixty percent of the data centers say they maintain adequate procedures for recording status of each magnetic tape; yet further examination of their responses indicates that this is not so: fully one-third do not appear to be recording the results of their periodic review of their archival and working tapes. although 77 percent state that periodic review is carried out, cleaning and testing is carried out only by 55 percent, and certification and precision rewinding by around 23 percent; however, a number of the data centers (particu15 table 8. tape maintenance activities. "many data and tape libraries have a set of activities to preserve their data on magnetic tape. check as appropriate those carried out by your library or archive." yes no total a. maintain back-ups of every data file b. control movement of medium c. magnetic tape sits for 24 hours d. environmental controls in storage area(s) e. monitoring of environmental controls f. off-site facility for storage g. record history of status of each mag tape gl. age g2. manufacturer g3. size g4. certification evaluation history other (tape contents) periodic review of archival and working tapes hi. cleaning h2. testing (evaluation) h3. certification h4. precision rewinding h5. other (roll -over, reading) g5. g6. 29 1 30 24 6 30 10 20 30 21 9 30 10 20 30 17 13 30 18 12 30 14 4 18 8 10 18 12 6 18 9 9 18 6 12 18 6 12 18 23 7 30 12 10 22* 12 10 22 4 18 22 5 17 22 17 5 22 *1 not ascertained. table 9. contents of the record keeping system. "do you have a record keeping system (either manual or automated)' check as appropriate." no totalyes a. identifies each reel b. identifies reel's contents c. provides location (and movement) of the reel d. describes status 30 30 30 30 22 8 30 17 13 30 table 10. periodic tape cleaning. yearly every two years every 5 years or more not at all /very seldom other ( ad hoc basis) na 5 4 3 9 6 3 35 16 larly the ibm community) are using special software to scan their tapes before the data are used. as table ]o shows, only 30 percent (n=9) are regularly cleaning their tapes (yearly or every two years is wnat is recommended). some 47 percent do not clean their tape collection at all or on an ad hoc basis. only three data centers indicated that they neither had developed nor had access to special software to evaluate the physical integrity of their magnetic tapes; thus, maintenance responsibilities (or the lack thereof) are not explained by inaccessibility of software. and although almost half say they must pay for cleaning and evaluating services supplied by their computing services, only a few data centers, as we described earlier, have indicated a funding problem. rather, the explanation probably lies in the availability of staff to carry out maintenance activities. almost two-thirds of the sample stated that the library staff is responsible for tape maintenance. a more in-depth analysis of the level of responsibilities and demands placed on the data center staffs would give us more information about this aspect of the tape maintenance problem. the next two questions deal with the level of satisfaction with the data center's tape maintenance activities. in response to the question, "are you satisfied with the things you do to protect your collection?" 53 percent (n=16) said "yes," and 47 percent (n=14) said "no." probing further into the "no" responses, we asked, "if not satisfied, would you do any of the following?" clearly, respondents are aware that they need to improve present practices of tape maintenance: tape quality must be monitored on a regular basis. somewhat less than half responded that they must upgrade their record keeping practices. only a small percentage attend to the need to establish environmental controls. it may very well be that they believe that they have less direct control over site environmental conditions and that, therefore, any attempts to influence the quality of these conditions would be fruitless. on the othe*hand, they may feel that environmental controls are already satisfactory and that this is not an issue in their tape maintenance practices. we wondered whether there were any differences between those respondents who said they were satisfied with their present practices and those who said they were not, with respect to the activities each group is carrying out. table 12 looks at all respondents who said they carry out maintenance activities. comparing satisfied data center staffs to the dissatisfied ones (both conducting good maintenance practices), we see little difference in the absolute numbers in each group except in two areas: more satisfied than dissatisfied staff members control the movement of magnetic tape and record the status of each magnetic tape in the collection. in sum, it might be suggested that the degree of satisfaction with one's maintanance practices lies in the quality of record keeping. c. needs the last question in our survey asked respondents whether a document on minimal standards for tape maintenance of an archival data collection would be useful to them. with only two exceptions, the response was positive. we also asked them what they would like to see in such a document. the responses are described in table 13. indeed, the greatest interest lies in report forms for record keeping in the tape library and a bibliography of the state-of-theart research on archival storage (23 of 28 respondents). next are procedures for protecting the magnetic tapes that undergo environmental changes, and inventory control procedures (17 and 18 of 28, respectively). the responses 17 are consistent with behavior reported by the data centers' staffs and known to dpls: record keeping is always a lower priority in an organization which has user services as its primary goal. record keeping is neglected because it takes time and is transparent to the user. staff is usually inadequate to support quality record keeping (which also includes inventory control). table 11. satisfaction with tape maintenance activities. "if not satisfied with present maintenance practices, would you do any of the following?" maintain back-ups of master files monitor tape quality on regular basis (evaluation, certification) develop complete records on status of every tape in collection establish environmental controls establish off-site facility for tape storage *1 not ascertained. table 12. satisfaction/dissatisfaction with tape maintenance practices by those who conduct tape maintenance activities. satisfied dissatisfied total let mag tape sit for 24+ hours control movement of mag tape establish environmental controls establish off-site facility conduct periodic monitoring record tape history and status carry out periodic cleaning table 13. contents of a document on tape maintenance. yes no total es no 12 total 1 13* 8 5 13 6 2 5 7 11 8 13 13 13 4 4 8 13 10 23 10 9 19 8 8 16 5 4 9 10 7 17 11 10 21 procedures for protecting mag tapes which undergo environmental changes 17 11 28 procedures for maintaining environmental controls in storage facility 13 15 28 report forms for managing the tape library 23 5 28 inventory control 18 10 28 bibliography of state-of-the-art research on archival storage 23 5 28 also, many believe that record keeping takes one away from the data, which are the raison d'etre of the center. the reality, however, is that without good record keeping practices, good user services cannot be provided and the collection is placed in jeopardy. iv. concluding remarks during the coming decade, precisely at a time when managers and administrators have come to recognize the importance of organizations wfiich preserve, maintain, and disseminate statistical and other data, university-affiliated data centers will be faced with limited funds to maintain their collections. obviously the economic realities call for creative technical and administrative solutions to the costly problem of data preservation. one solution is the development of storage devices which ensure long-term and stable preservation, to reduce the cost of yearly or bi-yearly file roll-over. devices are now in the experimental stage, and prototypes offer hope that effective media will be available at reasonable cost within ten years. another solution is better administrative practices to reduce the labor-intensive activity of record keeping. the computer and data base management software offer an opportunity to become more efficient and cost-effective—that is, to employ labor-saving devices for maintaining records of the data and tape libraries, inventory control, and retrieval and updating of information for periodic review of the status of the data and storage medium. nevertheless, both the new storage devices and use of data base management systems are initially costly items: tape drives which permit writing of data at a density of 6250 bpi may involve outlays of anywhere between $125,000 and $150,000--perhaps beyond the means of all but the largest computing centers. development of the data bases for record keeping systems will involve a sophisticated programming staff familiar with data base management; for although most computing centers already provide some type of data base management software, it is usually not designed with administrative record keeping in mind. it requires software interfaces, necessitates substantial data base investment and data entry personnel, and incurs continuing operational and maintenance costs. unless the administrative data base can be designed with multiple users in mind, developmental costs will have to be borne by the data center itself. unless the data center has unlimited free computing and programming assistance, use of labor-saving devices such as a data base management system will not occur. the best strategy for a university-affiliated data center with limited funds is to investigate the potential user market for such automated administrative record keeping systems and to convince this user community of the need to develop good administrative practices to maintain their data. in the meantime, however, data centers would be well advised to upgrade their present practices of tape maintenance to preserve access to their collections. inadequate attention is being given to the importance of environmental controls and the need for monitoring these controls on a regular basis. better protection for the collection through off-site storage of the master files is required. there appears to be too much reliance on the data center's computer center for carrying out basic maintenance. computing centers are primarily involved in throughput operations and not preservation activities. tape cleaning and evaluation need to be performed regularly and must be followed by adequate record keeping of the evaluation. since the majority of the data 19 centers do not hold large collections, periodic cleaning and evaluating of their tape libraries should not prove too time-consuming to be carried out within the constraints of their budgets. since all data centers expect growth in their collections, and particularly since they have probably underestimated this growth, they would be well advised to activate a program of good tape maintenance. because the focus of this survey was narrow, the data do not provide us with insights into the organizational problems of the data centers, staff allocations, demands for services, and the current budgetary situation, all of which probably influence the quality of the maintenance practices. the increasing reliance on statistical evidence for research, policy, and program planning, and the influence of libraries generally in the information transfer process, suggest that further examination of the data center would be useful. this small survey of tape maintenance practices should be followed by a more extensive survey of the data centers, to reveal more fully how they facilitate the flow of information and contribute to intellectual inquiry. current national funding priorities promote too centralized, too structured, and too hierarchical use of data repositories. this policy risks paralysis of the larger system and denies the pluralistic nature of information needs and services. local data centers are important contributors in a pluralistic system. their efforts in the areas of dissemination and maintenance of valuable archival data resources need to be fostered. endnotes 1. i mean conventional mass storage media, such as magnetic tape and disc; existing new storage systems, such as the ibm photostore and sdc tbm ii; extensions of current magnetic tape technology, such as the calcomp automatic tape library, ibm 3850, cdc 385000 system, precision instruments (omex) system 190. 2. potential mass storage developments (based on other technologies) include the omex vidicon system, video disk, direct digital film-based storaae, holographic storage, electron beam memories. volz comments that "there will be a number of new mass storage devices to reach the marketplace over the next two to three years. however, the immediate concern of the developers of these devices is to achieve a high recording density of large system capacity without extensive consideration for longevity of the media. this means that while some of the techniques do have some potential for archival purposes, many of the first applications are likely to be for large volumes of data which are nonarchival, that is, data which can be safely discarded after a few years. once adequate recording densities and access time are achieved, attention will be more focused toward the archival properties of the media existing mass storage devices are not truly adequate for the archival (sic) of large collections of data and the most imminent new technologies will probably also not be acceptable. a really good solution for large data collection archival (sic) is still a number of years down the road and good higher level software support is still further away." for a description of holographic storage, see maugh. the public archives of canada have been investigating the technology of recording data on special discs by exposure to focussed laser light--the video disk. 20 locke states that this "technology has recently reached the point of sufficient technical maturity such that it should be seriously considered as a basis for the storage of archival materials." he goes on to say that "laser recording provides the only economical basis for large-scale, machine-readable storage. in addition, data recorded by laser are expected to exhibit longer lifetimes and better security than data stored by any other information storage process. what is more, archival materials converted to digitally coded laser records can be preserved forever without any degradation whatsoever by the simple process of periodic replication protected by error-correction coding." in conversations with the author, harold naugler, director of the machine readable archives of canada, noted that it would probably be some years before production versions of these recording devices are on the market and have been tested. volz comments that the "hardware to perform both recording and playback is projected to be in the vicinity of $200,000 for a trillion bit storage. however, devices that only read are expected to be available for just a few thousand dollars. commercial marketing of the device is probably two years away." the high cost of the device certainly puts it beyond the capital equipment acquisition budget of every data archive and probably most computer centers. these devices also must have a write capability to be of any utility to the archive which has as a major function the dissemination of its collection. 3. another problem may be recording technique. prior to 1600and 6250-bpi with phase and block-coded recording, the recording techniques for data did not have the capability of correcting for errors caused by minor flows in the magnetic survace of the tape. in addition, newer tapes have a much smoother surface than those manufactured in the late 1960s. as a result, there is less wear on the read-write heads, less wear on the tape, and less likelihood of debris accumulating on the tape, according to volz. 4. as will be described later, a few of the identified data centers disseminate only icpsr data. it also turns out that disqualifying icpsr member institutions in our sample resulted in eliminating a few data centers that could have participated in our survey. 5. twelve are classified as independent departments or organizations; three, affiliated with a teaching department; ten, affiliated with a research organization; three, affiliated with a computer center; and two, affiliated with a library. 6. the histogram excludes one data center established in 1941 and two centers which could not supply this information. 7. at this point it is useful to describe the computer hardware at these data centers: ibm accounts for 47% (n=14), amdahl, 7% (n=2), cdc, 27% (n=8), dec 10, 3% (n=l), dex-vax, 3% (n=l), itel as6, 7% (n=2), xerox, 3% (n=l), and univac, 3% (n=l). '8. the question about the type of tape storage problem was left open-ended; as a result, totals exceed number of centers which responded that they had problems. 9. almost 30% of the data centers noted that they are storing their master files on-site, and 44% both onand off-site. these high statistics are cause for some degree of concern for long-term data preservation. 21 10. respondents were asked to check any of the following which applied: collection too large, financial support inadequate, or not enough staff. references boruch, r.f. and wortman, p.m. "an illustrative project on secondary j^jysis. secondary analysis. san francisco: jossey-bass, inc., 1978, dollar, cm. "problems of magnetic recording in archival storaqe " diqest of pagers. spring compcon 1977, the 14th ieeec society internationa l conference, san francisco, 28 february-3 march 1977, 28-30. geller, s.b. "layaway, standby and reactivation procedures for computer magnetic media." (n.d.) hofferbert, r.i and clubb, j.m. "introduction." america n behavioral scientist, 19(4), 381-386. " locke, j.w. "videodisc pilot project progress report: phase i (8sep78 to 31mar79)." prepared july 1979 for the public archives of canada (mimeo). maugh, t.h. "holographic file: an industry on the verge of birth." science , miller, k'.e. "the less obvious functions of archiving survey research data." american behavioral scientist , 19(4), 409-418. nesvold, b.a. "instructional applications of data archive resources." american behavioral scientist , 19(4), 455-466. "~ robbi'n, a. "technical guidelines for preparing and documenting statistical data. _ in boruch, r.f., wortman, p.m., and cordray, d. (eds.), secondary analysis; policy and practice in applied social research . san f rancisco'jossey-bass, inc. (forthcoming) rockwell, r. personal communication, 20 december 1979. rokkan s. "data services in western europe: reflections on variations in the conditions of academic institution-building." american behavioral scientist , 19(4), 443-454. li-hjvolz, r.v. "computer based mass storage technology." prepared for the conference on archival management of machine-readable records, ann arbor michigan, 7-10 february 1979. 22 1/34 nesvijevskaia, anna (2021) databook: a standardised framework for dynamic documentation of algorithm design during data science projects, iassist quarterly 45(2), pp. 1-34. doi: https://doi.org/10.29173/iq989 databook: a standardised framework for dynamic documentation of algorithm design during data science projects anna nesvijevskaia1 abstract this paper proposes a standard documentary framework called databook for data science projects. this proposal is the result of five years of action-research on multiple projects in several sectors of activity in france and a confrontation of standard theoretical processes of data science, such as crisp_dm, with the reality of the field. the minimalist and flexible structure of the databook prototype, described and illustrated in this paper, has revealed its operationality on more than a hundred projects and has been recognised by various stakeholders as an excellent facilitator of human data mediation, especially for multi-skilled projects. beyond its proven benefits for project efficiency, this framework, conceived as a frontier object, can be applied more broadly to data project portfolio management and data value, governance and quality. by surpassing the computational aspect of the models, the databook is an answer to the issues of interpretability and auditability of algorithms. keywords data science, artificial intelligence, documentation, reproducibility, algorithm transparency, project process, fair, human data mediation introduction the proliferation of data science projects has accelerated knowledge discovery and generated new algorithms for commercial use. it has also produced massive amounts of exploratory data and metadata. yet exploration is still poorly equipped to understand the processed data in terms of meaning, utility and value. on one hand, recent data science platforms tend to structure only the technical aspects of the data engineering pipeline, data linkage and algorithmic libraries. this technicity erects a barrier to the understanding of the data by all project stakeholders, especially in complex multi-skilled projects. on the other hand, traditional master data management tools, which handle this type of metadata on the key records of an organisation, are generally incomplete and unable to absorb all the data created during data science projects, including the final algorithmic model. the lack of standards for the capitalisation of this data leads to difficulties in replicating results, a lack of transparency and efficiency of the arbitrations made during these very dynamic projects, and a loss of resources during the data understanding and qualification phases of subsequent data science projects (portfolio management). these limitations are critical in the context of increasing european regulation and growing acculturation of business decision-makers to algorithms. both trends require a shared and facilitated data understanding that goes beyond technical and mathematical measures. this paper proposes a standard documentation framework, called databook. it has been conceived for data science projects in which algorithms are designed. the emergence of the databook is guided by (1) the theoretical and practical limitations of standard data science processes. this gap has been (2) compensated in the field by a databook prototype with a unique structure. the prototype was (3) tested and confirmed as efficient for several purposes and stakeholders: this paper proposes its evolution to the first standard algorithm design documentation framework. https://doi.org/10.29173/iq989 2/34 nesvijevskaia, anna (2021) databook: a standardised framework for dynamic documentation of algorithm design during data science projects, iassist quarterly 45(2), pp. 1-34. doi: https://doi.org/10.29173/iq989 1. standard process in data science projects: from theory to field reality to begin with, we will consider the theoretical processes of data science projects and their results, the main limits of these processes identified in the field, such as the lack of documentation, and the first attempts to fill the documentation gaps in practice. 1.1 overview of the standard process in data science projects data science projects aim to build an algorithmic model for a specific practical purpose (a usage). the model consists of an input data, a finite sequence of well-defined operations, and an output en terms of analytical result. the choice of a model depends on the problem to be solved, such as a phenomenon prediction or its correlation to root causes. the solution is usually an articulation of several algorithms chosen among thousands of possibilities. the algorithms are applied to data selected for the project from an expanding number of available sources: the data are then assumed to contain in past observations an insightful signal that is key to the problem, and the algorithm is, therefore, a means of revealing this signal. the uncertainty of these projects is highly substantial because the presence and usefulness of the signal in the data must be explored, and sometimes the emergence of a signal predates to formulation of the need. as recent technological advances have had an impact on the entire data chain value (bertino et al. 2011; miller & mork 2013), the cost of these exploration projects has reduced and opened up new horizons for possible usages in all sectors (manyika et al. 2011; mayer-schönberger & cukier 2013). the business needs covered by these projects are currently very diverse, most of them being assimilated to knowledge generation or decision-making acceleration. in both cases, the sense and value creation by the algorithmic model depends on a broader usage device (brynjolfsson et al. 2011; provost & fawcett 2013) that includes a purpose, a context, a decision-making process, a user community, an interface or a workflow and many other elements that are impossible to standardise. despite this variety in terms of usages, algorithms and exploitable data, data science projects are composed of a similar sequence of activities described in the widespread use of data mining for knowledge discovery in databases (fayyad et al. 1996; piatetsky-shapiro 1994). commercial actors, researchers and companies leading these projects have attempted to standardise these activities: the most successful attempt is the cross-industry standard process for data mining, or crisp_dm (chapman 1999; shearer 2000; wirth & hipp 2000) that resulted from a convergence of reference processes and their confrontation in the field by a mixed consortium funded by the european union. this process model breaks down the project life cycle into six phases: business understanding, data understanding, data preparation, modelling, evaluation and deployment. it captures the complexity of data exploration by identifying the main iterations between these phases and remains neutral in terms of usage, tools and data. since the suspension of the consortium, several proposals to improve the standard process have remained pending: to expand the number of use cases, to map and describe the activities and their results in more detail or to link the process to different project management methods. as the most stable and widely used process in data science (camiciotti & racca 2015; provost & fawcett 2013), crisp_dm was defined as a reference in the course of an action-research which was conducted on seven different projects from 2014 to 2017 (nesvijevskaia 2019). the objective of this thorough qualitative multiple cases study was to understand why big data, as a myth-bearing sociotechnical phenomenon (boyd & crawford 2012) reflected in companies by the implementation of the https://doi.org/10.29173/iq989 3/34 nesvijevskaia, anna (2021) databook: a standardised framework for dynamic documentation of algorithm design during data science projects, iassist quarterly 45(2), pp. 1-34. doi: https://doi.org/10.29173/iq989 first data science projects, did not generate the expected value. the relevance of the crisp_dm framework was confirmed, as were its expected limitations. the iterative nature of the main tasks was revealed as a regular and beneficial overlap between the six phases. the framework also found to be too focused on the algorithmic model, with a risk of uprooting the project results from the practitioner’s activity (nesvijevskaia 2017). it explains potential project failure in terms of result exploitation. indeed, the process delays the anticipation of the usage, data inclusion/exclusion criteria but also the co-construction of the results restitution. this delay creates a risk of inadequate expectations, project costs drifts for production launch, errors in analytical strategy, but also a lack of capitalisation throughout the project. this confrontation between the reference standard process and the field leads to the building of a global data project device called brizo_ds2 which includes an adjusted crisp_dm model. the reference outputs of each phase of the adjusted model are mapped to the process critical path and the process documentation (figure 1). figure 1 output mapping of the adjusted crisp_dm the mapping of the reference outputs in the figure above is based on the following principles. the critical path is composed of intermediate analytical outputs: raw data lead to selected data which are structured to feed algorithmic models, and the best models are selected to generate knowledge or decision-making. some or all of the critical path outputs may be automated. these computational https://doi.org/10.29173/iq989 4/34 nesvijevskaia, anna (2021) databook: a standardised framework for dynamic documentation of algorithm design during data science projects, iassist quarterly 45(2), pp. 1-34. doi: https://doi.org/10.29173/iq989 objects are specific to data science projects as components of the finally used algorithm. they are materialized by the code of the algorithm and are accessible for the coding project team members. all intermediate outputs can be versioned: the first version allows the launch of the next phase of the project; the intermediate versioning explains the overlap between phases and the progressive optimisations of the algorithm during the project; and the final version corresponds to a component of the algorithm used for exploitation. by isolation of this critical path, the documentation is composed of all the other outputs that can be shared between stakeholders in a tangible or intangible format. they can be generated before the execution of a phase (anticipation), during the phase or once the phase is finished. the numerous and non-mandatory possible documentation outputs (see the most common ones in figure 8 in appendices) can be classified into three main categories: critical path analytical outputs documentation: all the outputs in this category describe and qualify the intermediate analytical results and guide the convergence on the final optimum algorithm. usages: these outputs describe the operational conditions for the final analytical result and knowledge activation (this activation can take place before the finalisation of the algorithm). mediation milestones: all outputs in this category trace the arbitrations realized during the project and guide the project management. the three categories above remain interdependent: the progressive design of analytical outputs feeds the project management; the project management makes decisions considering the value generation through usages; the usages emerge from analytical exploration and impose constraints and priorities. once the outputs of each phase of the process are classified, their nature and production methods can be analysed by comparing practices and reviewing the state-of-the-art practices. 1.2 main limits of the documentation outputs the critical path has been broadly supported by the development of analytical tools (data engineering platforms, algorithmic libraries…) and the skills of freshly and progressively professionalised data scientists (davenport & patil 2012). it has therefore been increasingly productive. however, the actors implied in field projects can be more diverse and most of them are not supposed to open a data science application, read a line of code (even a well-commented code) or juggle technical and mathematical concepts. in small-size projects, the most representative skills are usually divided into business skills and data skills, for instance when a data scientist works with a decision-maker: bridging the gap between these skills’ carriers is still testified as insufficient and critical (austin et al. 2021). in more complex projects observed in the field, more individuals are implied. data skills can be carried for instance by machine learners, data engineers, data stewards or data analysts. business skills are devised into strategic, analytical and operational. the last type of skills is often carried by users’ representatives (for instance product owners) or knowledge managers, but they are very dependent on the expected usage. some skills are dual, such as business intelligence skills. the most complex project teams are composed of members with mixed skills and with different levels of maturity: they require a strong mediation through project management (nesvijevskaia 2019). besides skills complexity, team members can be https://doi.org/10.29173/iq989 5/34 nesvijevskaia, anna (2021) databook: a standardised framework for dynamic documentation of algorithm design during data science projects, iassist quarterly 45(2), pp. 1-34. doi: https://doi.org/10.29173/iq989 confronted with difficulties related to data complexity (numerous sources, numerous extractions including erroneous ones, lack of meaning sharing by all actors, complex treatments, progressive identification of bias...). in the field, this complexity and high uncertainty require successive arbitrations implying a diversity of stakeholders. to achieve efficient arbitrations, heterogeneous stakeholders urgently need a common intelligible framework and semantics of raw and processed data at each stage of the project. however, reference documentation outputs are little mentioned in research work on data process, lacking anthropocentric anchoring. the principle of the first databook, as a dynamic documentation device, was guided by this urgent need and by a broader interdisciplinary approach to data quality. it had to take into account both computational (berti-equille 2012; wang 1998) and cognitive (arruabarrena et al. 2019; broudoux & scopsi 2011; cottin & nesme 2017; odeh & chartron 2016) aspects of transforming data into useful information by reducing uncertainties (mayère 1990) in a given economic context (doucet 2010). these historical approaches generally apply to the master data management (loshin 2010), where data governance issues are more largely focused: its objective is to ‘increase business performance (by adjusting the value of the data) and reduce the costs associated with the processing and management of master data’ (mariko 2016). the inspiration also came from digital knowledge media engineering, seeking to establish standard attributes of knowledge elements (zacklad et al. 2007). it also implied to follow the processes of capitalisation, sharing, knowledge creation, learning, selection and evaluation of useful information (ermine 2003). however, the amount of data explored, created and discarded during time-limited data science projects did not allow for a full comprehensive data quality process, and the issues of a single exploratory project were not as significant to the investment as the meticulous processing of master data. the iterative exploration process required a flexible and dynamic data quality and knowledge sharing device that was difficult to transpose from mdm. it also had to be more practical for the project and the knowledge management needs common in consultancy practices. in opposition to the theoretical limitations and field pressure, a first databook was imagined and tested in real-life situations. 2. the databook: prototype structure the databook prototype is a generic documentation output specific to data science projects. it describes all the algorithm components and the decisions that occurred during the algorithm design. hence, its structure reproduces the critical path outputs documentation and encapsulates the dynamic links between the algorithm, its usages and its design process punctuated by the mediation milestones. this structure is materialised in a single, common and shareable excel file: each spreadsheet of this file represents a module. all the prototype modules are listed in figure 2 and classified following the documentation categories listed in section 2. these modules can be supplemented with less specific documentation objects in different formats, such as powerpoint reports, data visualisation interfaces or tools required by the usage or by the project management. https://doi.org/10.29173/iq989 6/34 nesvijevskaia, anna (2021) databook: a standardised framework for dynamic documentation of algorithm design during data science projects, iassist quarterly 45(2), pp. 1-34. doi: https://doi.org/10.29173/iq989 figure 2 databook prototype modules the following sections present in detail the different databook modules and their flexible building mechanism which relies on a clear distinction between core structure and metadata structure. 2.1 core structure of the databook the core structure is the skeleton of the algorithm, usages and mediation milestones documentation. it is introduced in a guide (module 0), which is a manual presenting the ten following modules grouped into three categories detailed in the following sections. 2.1.1 documentation of the analytical outputs of the project the structure of modules from 1 to 5 replicates the critical path of intermediate analytical outputs of each phase of a standard data science process, excluding the deployment phase. an intermediate analytical output is defined as a data object (for example, a table) composed of elements (for example, variables in the table). the breakdown of data objects into more detailed elements remains specific to each project, but it always results in a structured list of homogeneous items. this list is usually presented like a hierarchical directory. each data object or element in this list can then be completed with attributes, or metadata. these attributes result from the addition of several descriptive criteria that will drive the choice to keep or abandon an element for the following phase. this decision is a qualification traced through a status of the data object resulting from its judgement based on different criteria. these metadata (criteria and statuses) are also usually organised thematically or/and hierarchically: the choice of this structure remains specific to each project and will be presented in section 2.2. the matrix representation of the structured list of homogeneous items associated with metadata fits perfectly formats such as excel. for instance, the module 2 (source data) aims the qualification of all the data that must be explored: an illustration of this module completed in a real-world project is presented in appendices in figure 9. another illustration can be found in figure 10 for the module 3 (model structure), with the qualification of all new generated variables and their selection in a context of multiple algorithm development. the other 3 modules follow the same matrix structure. databook prototype modules 0 databook guide mediation milestones a project roadmap b method of data inclusion/exclusion c exploration report critical path analytical outputs documentation 1 perimeter 2 source data 3 model structure 4 analytical results 5 functional results usages 6a usage roadmap 6b expérience return https://doi.org/10.29173/iq989 7/34 nesvijevskaia, anna (2021) databook: a standardised framework for dynamic documentation of algorithm design during data science projects, iassist quarterly 45(2), pp. 1-34. doi: https://doi.org/10.29173/iq989 2.1.2 documentation of usages and knowledge the last deployment phase is often restricted in the literature to the usages directly aimed by the project. however, the databook framework includes the documentation of both direct usages (technical and operational aspects) and knowledge generated throughout the project. it splits knowledge into two types: knowledge that can be potentially transformed into a business lever (indirect usage) and knowledge that can be useful for further data science projects (data project experience). direct usages are immediately operational levers which have been decided upon for deployment. each direct usage is associated with deployment actions that can be described in terms of purpose, modus operandi, associated version of the solution to deploy, expected benefits, deadlines, responsibilities, key indicators to monitor and so on. in the case of an automatized algorithm, actions include the pipeline automatization tasks, and sometimes the interface development specifications. indirect usages are potential levers with remaining uncertainties to investigate after the project. they result from knowledge that still requires concrete actions to be transformed into levers. both types of usages require a usage roadmap (module 6a): it contains a structured list of actions, usually broken down into a list of tasks and associated with metadata. as with the previous modules, the metadata includes the descriptive criteria and the status of each action, corresponding the decision to activate it or not. the matrix representation of this roadmap is very appropriate as a basis to feed other formats, more commonly presented to deciders (for instance, a report, a monitoring interface or the last version of the application). an illustration of this module is presented in appendices in figure 11. data project experience covers all the qualitative feedback in the form of knowledge capitalisation useful for further data projects (module 6b). usually this experience feedback is tacit, intangible or orally shared, but it can also be documented, especially when the knowledge must be shared with stakeholders outside the project. for example, if team members judged that cleaning up the data in a particular table was not necessary for the project but had an intuition that it would be of great value to other projects or existing usages, this intuition can be capitalised upon. another example can be a good coding practice capitalisation, or a business concept explored and finally judged as not of interest. this knowledge can potentially save significant time in future. faced with the variety of possible knowledge that can arise from the experience of a project, the databook prototype stops at a proposal to incrementally draw up a list of insights by application domain without seeking to structure the qualification of these ideas. in each project context, these ideas can then be shared with the appropriate stakeholders in the most suitable format. 2.1.3 documentation of the milestones of the project the milestones are usually the tip of the iceberg for the project management and provide essential elements for arbitrations throughout the project. as project decision facilitators, these milestones are usually more convenient to present with storytelling components, including texts, graphs and other project management best-in-class practices. however, they also necessarily include elements that must be fed with structured data and metadata issued from the other modules presented above. this category includes three modules graded a, b and c. the module a is a project roadmap with project advancement statistics based on statuses of elements qualified at each phase. it is illustrated in appendices (see figure 12). each time the databook is versioned, the module a represents a https://doi.org/10.29173/iq989 8/34 nesvijevskaia, anna (2021) databook: a standardised framework for dynamic documentation of algorithm design during data science projects, iassist quarterly 45(2), pp. 1-34. doi: https://doi.org/10.29173/iq989 photography of the version. the module b is a data inclusion/exclusion methodology synthetizing the structure of descriptive criteria and their expected impact on the qualification statuses. it represents the project decision rationale traced through metadata and results from a more complex mechanism described in section error! reference source not found.. the module c refers to exploration reports: i t is very specific to each project. if the project has only one exploration report, it can be directly integrated in this module. however, usually a project generates several reports and each report is realized in its own format such as a data visualization or a presentation. in this case, the module c lists the different reports, their versions, associated decision milestones (for instance, the date of a project committee) and key elements and decisions. this module forms then a bridge between the decision milestones and earlier databook versions as well as with other possible project documents. 2.1.4 core structure synthesis each of the ten modules remains adaptable to the complexity of different projects thanks to a flexible database composed of custom lists, data objects, elements and associated metadata. the complete databook core structure is presented below in figure 3, with an illustration of the most common documented elements observed in the field for each module. figure 3 databook core structure and its modules, illustrated with the most common elements and data objects perimeter source data model structure analytical results functional results usage roadmap phenomenon individuals drivers time perimeter eval. criteria key indicators documents sources ∟bases ∟tables ∟variables ∟values aggregations cleaned data derived data learning / test / validation matrix filters types of models ∟models ∟models’ parameters existing results models’ results ∟adjustments to the model results actions remaining uncertainties databook guide project roadmap method of data inclusion/exclusion exploration report 0 a c 1 2 3 4 5 6a exp erie n ce re tu rn b summary, description of the used modules and of their purposes, reading grid work remaining to be done for the project and synthesis of statuses at a given date for modules 1 to 6a for each data object in modules 1 to 6a, descriptive criteria and methods of statuses qualification (metadata structure) list of reports (data visualizations, presentations…), their characteristics, versions and associated decision milestones metadata metadata metadata metadata metadata metadata legend illustration of main elements and data objects of each module and of their typical structure. modules 1 to 6a are completed with metadata which is piloted, rationalised and shared through modules a to c. module namesdatabook introduction documentation of the mediation milestones of the project documentation of analytical outputs of the project documentation of usages and knowledge insights by application dom ain / by skill 6b https://doi.org/10.29173/iq989 9/34 nesvijevskaia, anna (2021) databook: a standardised framework for dynamic documentation of algorithm design during data science projects, iassist quarterly 45(2), pp. 1-34. doi: https://doi.org/10.29173/iq989 2.2 metadata structure as presented above, each data object is described in terms of criteria and qualified with a status during the project. this documentation treatment is recorded through metadata. but, unlike the analytical treatments carried on the data throughout the critical path realization, the metadata treatment does not follow a sequential dynamic. indeed, as seen in figure 1, documentation can occur before the analytical work to anticipate it, during or after critical path completion. this dynamic is closely linked to the nature of the uncertainties reduced at each stage of the project: the skills required to anticipate the risks of each phase and to carry out the associated treatments are usually the same. the next sections reproduce the standard process and concentrate on the skills associated to each phase. it shows how those skills are implied in the documentation beyond their intervention in the critical path. 2.2.1 business understanding the data project must be anchored in a given business context in order to lead to a result that will be in line with the business strategy. this anchoring can be reflected in strategic criteria, business priorities and confidence in the relevance of each data object bearing real-life concepts in a business context. business understanding metadata also includes all the regulatory constraints such as gdpr or discrimination rules that can lead to some data or model exclusion despite their statistical significance. the documentation in terms of business understanding requires the skills of strategic management and business analysis. 2.2.2 data understanding while source data understanding is part of the critical path, it does not stop here. all data objects produced during the project need semantics, units, names and other metadata that will make sense to all stakeholders in various communities. as observed in the field, this is one of the most used metadata in the databook, facilitating co-construction and appropriation of all intermediate outputs. the level of detail can vary from a name of a data element to a definition or an in-depth explanation of its generation process. it must be aligned with the level of maturity of the stakeholders, their usual vocabulary, language, shortcuts and so on. if the databook is to be shared outside the project team, these semantics can be completed with translations, comments and other facilitators. this semantic metadata can be structured and used as a dictionary or a repository, for instance in exploration reports or data visualisation, to complete technical structured nomenclatures of explored data with meaning. this documentation requires the skills of data stewardship and data analysis. 2.2.3 data preparation this documentation is predominantly technical and describes the data engineering issues in order to anticipate the exploration and the exploitation pipelines. the metadata for this purpose is composed of the function of data elements in the pipeline structure (keys, filters, place in query structures, filling controls, duplication controls…), of their formats and volumes with associated calculation time, and of other technical criteria. the technical documentation can lead to choose a given tool, language or even model or usage device and interface, and sometimes the exploration pipeline will differ from the exploitation pipeline. this documentation supports the anticipation of the usage deployment (controls, automations…) and requires the skills of data engineering and data analysis. https://doi.org/10.29173/iq989 10/34 nesvijevskaia, anna (2021) databook: a standardised framework for dynamic documentation of algorithm design during data science projects, iassist quarterly 45(2), pp. 1-34. doi: https://doi.org/10.29173/iq989 2.2.4 modelling much more mathematical, this documentation links data objects to algorithmic models and evolves during the project as the model is anticipated, realised, calibrated, benchmarked and adapted to the usage. modelling metadata corresponds to the analytical uncertainties and signal detection in the data. it includes metrics such as minimums, maximums, standard deviation of a given variable, modality distributions and a varied set of algorithm-specific metrics, such as parameters, hyperparameters or statistical evaluation criteria calculation. it also anticipates the re-learning mechanism and its monitoring if it is needed for the algorithm exploitation. this documentation requires the skills of machine learning and data analysis. 2.2.5 evaluation this documentation is paramount to understand the consistency of each data object in terms of its contribution to the key performance indicators of the expected usages. naturally, this documentation concerns the translation of analytical results into business value: the statistical evaluation criteria (for example, the false positive rate for a churn prediction algorithm) must be translated into business evaluation criteria (for example, full-time equivalent or operational cost). this translation goes beyond semantics and includes the value calculation methods and its intelligible representation. interpretability, rapidity, user appropriation or maintainability of a model can also be considered as performance indicators to judge a model for a given usage. this means that evaluation metadata can be both quantitative and qualitative. the best practice is to imagine the performance indicators first and then derive the statistical evaluation criteria from the business criteria, even for exploratory projects aimed at generating original knowledge. in this case, the initial performance indicators are specified as they are developed. however, this type of documentation applies not only to analytical results but also to all the other data objects. for instance, interpretability can be judged for source data or newly generated variables in terms of consistency between the definition and the perceived meaning. value calculation rules and orders of magnitude can also be anticipated since the beginning and progressively controlled. for instance, the control of the consistence of a customer database used in the project needs a comparison between the volume of the database lines and the number of customers usually measured by the company in existing reporting. these consistency controls are particularly useful when heterogeneous stakeholders from different parts of a company must make project decisions on a common objective basis. consistency controls occur for intermediate and final data objects: they explain a significant number of iterations because they help to detect errors. they include not only obtained, but also expected metadata: for instance, a usage value can be judged through the delta versus an expected value. this documentation requires the skills of business intelligence, but also both data analysis and business analysis. 2.2.6 deployment this documentation corresponds to the usage anticipation throughout the project, whether they are defined from the start or gradually emerging. indeed, this usage anticipation can lead to operational priorities or exclusions despite the business, technical or mathematical importance of certain data objects. for example, if the usage is a real-time decision-making, but a data source is collected only on a monthly basis, this data source may be eliminated from the project despite its strong predictive power. another recurring example is the volume of explored data: big volume can be perfectly usable https://doi.org/10.29173/iq989 11/34 nesvijevskaia, anna (2021) databook: a standardised framework for dynamic documentation of algorithm design during data science projects, iassist quarterly 45(2), pp. 1-34. doi: https://doi.org/10.29173/iq989 for the exploration but inappropriate for usage exploitation. this qualification is also notably critical for exploitation of personal data. the usage anticipation is necessary to avoid the risk of producing interesting but unexploitable results. deployment documentation describes the operational constraints linked to the usage exploitation and strongly impacts the qualification of all the previous data objects. indeed, their criteria must be adequately judged to determine the status. operational criteria are very variable from a project to another. they depend on the stakeholders responsible for activating the project results and piloting the usage. deployment requires the skills of product ownership for direct usages (this skill must be specified for each field of application) and/or knowledge engineering for indirect usages, usually completed with skills of business analysis. 2.2.7 project management as the analysis work progresses, data elements are treated at different phases of the critical path: these treatments are documented with associated metadata. this analytical work can be done phase by phase, but also iteratively or through phases’ overlaps thanks to the versioning of intermediate analytic outputs, as described in figure 1. a version then corresponds to a set of data elements that will evolve. for example, a machine learner can start working on a model with incomplete data, in order to test the first assumptions, without waiting for a complete dataset to be prepared by a data engineer, who is waiting for the last data extractions. in another situation, a business stakeholder may perceive a meaningful variable without knowing if this variable exists in a source. this variable will then be added to the databook as being perceived as useful, but not directly confirmed as treatable. these advancement tactics cannot be followed if only completed analytical treatments are documented. however, they are perfectly followed when metadata are generated as soon as the treatment anticipation occurs. these advances are synthesised by a status of each data element in terms of qualification. the most common statuses are to be started, in progress, included or excluded, and can be adjusted for multiple uses in the same project (for example, a source can be included for algorithm a and excluded from algorithm b). the statuses are essential for sticking to the design dynamic and represent the pivot between phases: they vary with the versioning (databook is then versioned in parallel) and can be supplemented with workload estimates or difficulties anticipated. this qualification thus requires the skills of project management, but also business analysis and data analysis to guarantee the understanding of each activity stakes and associated qualification methods. 2.2.8 metadata structure synthesis in the databook, metadata can be generated independently form the analytical critical path as it aims both the anticipation and the reduction of different uncertainties of the project. each phase of the project represents a type of uncertainty and involves a specific skill to reduce it. this skill must also be involved in the anticipation. this documentation mechanism generates rich criteria guiding the projects decisions traced through the statuses. therefore, this framework is a boundary object than can be successfully adapted to different points of view for all the skills’ carriers engaged in the project and robust enough to maintain identity between them (star & griesemer 1989). thanks to this structure, skills’ carriers participate using their documentation capacity in the main arbitrations leading to the construction of project results. https://doi.org/10.29173/iq989 12/34 nesvijevskaia, anna (2021) databook: a standardised framework for dynamic documentation of algorithm design during data science projects, iassist quarterly 45(2), pp. 1-34. doi: https://doi.org/10.29173/iq989 given the diversity of possible metadata, figure 4 reports a non-exhaustive list of the most commonly used criteria. this illustration represents the databook modular core structure (vertical) enriched with typical metadata classified by type of uncertainty reduced for each phase of the project (horizontal). for each concrete project, this structure can be documented in the module c, i.e. data inclusion/exclusion method. all metadata are then listed with associated modules and skills (usually, skills are represented by an individual skills’ carrier) and for each criteria the status establishment method is mentioned. for instance, if the criteria ‘personal data’ is flagged as ‘yes’ for a variable, it will have to be flagged as ‘excluded’ in the status. this flexible mechanism linking the core structure and the metadata structure of the databook results from an iterative confrontation between theorical and practical requirements. figure 4 illustration of the most typical metadata used in the databook modules 3. the databook: from a prototype to a standard documentation framework in the next section, field feedback on the prototype is discussed in order to stabilise the databook as a generic documentation framework and propose its main reading grids. 3.1 benefits perception the imagined prototype appeared in the field to be a complete, autonomous and dynamic object, logically linked to the stakeholders’ documentation needs. proposed as an excel file with basic core and metadata structures, it was progressively filled up by the project teams on seven projects. the structure of the prototype appeared flexible enough to be adapted to meet the urgent needs and priorities of stakeholders, and if was successfully judged as operational. besides its usability, the databook prototype revealed its significant effectiveness in stimulating cooperation, enlightening the business understanding data understanding data preparation modelling evaluation deployment project management 1 perimeter pertinence, business priority, consistence as business concepts sense of each key metric temporal perimeter, filters, keys… utility in the model construction translation of a key metric into data (calculation rules) importance in decision making, activability, historization possibilities… difficulties to reach exploitable data, status 2 source data confidence in the meaning, regulatory constraints… sense of each source data element, units, cognitive bias… accessibility, integrity, completeness… precision, bias, statistical metrics, nature of values… consistency between the perceived sense and the source data exploitability of the volumes, formats, techniques and resources needed for exploitation… difficulties to prepare the source data, status 3 model structure confidence in the construction process, regulatory constraints… sense of created data, units, nomenclatures, cognitive bias… description of treatment (keys, filtering, gross use, derivation, rules, matrix cutting…) models impacted, weight of signal for each data element… consistency between the perceived sense and the new data exploitability of created data, resources needed for deployment and maintenance… difficulties of the modeling, status 4 analytical results level of model and result generation process understanding sense of each model benchmarked (name, family…) calculation time and resources… statistical evaluation criteria : performance and confidence levels… consistency between the model process and its explanation frequency or thresholds of learning, maintainability… difficulties of result evaluation or monitoring (tools, competences…), status 5 functional results business criteria evaluation, confidence in benefits potential… sense of the retreated results (name, nature, units…) and of evaluation metrics added indicators, format changes… retreatments issued from business appropriation (adjustments…) translation between business and statistical criteria (calculation rules, qualitative…) exploitation performance target and its confidence level… difficulties of deployment preparation of direct and indirect usages, status 6 usages usage priorization, decision logs and criteria (benefits, uncertainties…) sense of each usage and their indicators nature of treatments of integration, automatization, controls, securisation... further model treatments needed for exploitation (self-learning parameters…) level of contribution of a usage to benefits… exploitation & maintenance modalities estimation of benefits and remaining uncertainties needing further investment, status https://doi.org/10.29173/iq989 13/34 nesvijevskaia, anna (2021) databook: a standardised framework for dynamic documentation of algorithm design during data science projects, iassist quarterly 45(2), pp. 1-34. doi: https://doi.org/10.29173/iq989 arbitrations made during the project, and guiding the appropriation of the results by the stakeholders. different beneficiaries highlighted through qualitative feedback that the device was used for two main purposes: project efficiency and data documentation efficiency. the project efficiency was improved through the facilitation of the mediation milestones (meaning shearing, progress visualisation, decisions traceability…), the knowledge capitalisation useful for further data projects and the identification of indirect usages requiring further analytical investigation. the impact on the project results quality was also highlighted by the users and by business decisionmakers: the traceability of all the analytical components of the algorithm and project decisions was perceived as a quality and solution auditability guarantee. moreover, the recorded indirect usage ideas inspired not only further analytical iterations but also business offer and process evolution. several data and business project team members also reported a change in posture: at the beginning of the project, they perceived filling in the databook as an additional workload of little use, but as the project progressed, they realised that the device was indispensable and generated time savings at each iteration. the documentation efficiency was appreciated not only by the team members but also stakeholders outside the project. for instance, data governance managers were very interested in semantic and other metadata generated during the project and the identification of referent data owners. data protection officers kept and reused the personal data identification to control the discrimination drifts and the usage purposes. financial managers were also very interested in the possibility to retrace the value generated by the usage and link it to the different data sources: this opened a new field for exploring patrimonial, finance and accounting concepts around data assets development. finally, it managers reused the databook to size the technical resources required to explore, deploy and maintain other business applications. this feedback highlighted that the purpose of the databook largely exceeded the project efficiency and aimed data project portfolio management and globally the data quality and value management. 3.2 operational limitations the prototype also showed its weaknesses, the main one being its excel format which is not very practical for several modules. indeed, exploration report, functional results and usage roadmap modules had to be completed with more specific formats such as data visualisation tools or powerpoint reports. the analytical results module was also completed outside the excel file when it was too specific to the benchmarked algorithms. the project roadmap module was simplified as much as possible in order to be coupled with more appropriate project management formats. if the prototype was to evolve to a more sophisticated tool, it would be necessary to handle other types of formats for these modules or favour the compatibility or interoperability with other dedicated tools. the remaining modules have been adopted in excel format and adapted to each project. for more complex projects that implied a production of several articulated algorithms, the prototype’s modules have been multiplied to serve better the skeleton of the algorithm articulation. most of the time, one or more modules was left empty: this selection revealed the adaptation to the specific needs and resources of each project. finally, the module with the data exclusion/inclusion method, predefined with fixed metadata, was redefined orally by the stakeholders and applied mostly through column creation in modules 1 to 6a. some mandatory regulatory criteria were dropped, such as the personal https://doi.org/10.29173/iq989 14/34 nesvijevskaia, anna (2021) databook: a standardised framework for dynamic documentation of algorithm design during data science projects, iassist quarterly 45(2), pp. 1-34. doi: https://doi.org/10.29173/iq989 data flag for anonymisation for projects without personal of data. given the field feedback, the databook structure is confirmed as flexible enough for different projects but attempts to make it more rigid (fixed metadata, mandatory modules…) are qualified as inoperative in practice. module or metadata aborts were explained by the fact that time investment priorities were set at the small scale of each project, and rarely at the scale of a project portfolio or the company. they were also partly explained by organizational and human factors. indeed, as a collaborative tool for a multi-skilled data science project team, the databook has raised several organisational issues. when a data science projects remains restricted to a small team, for example with one data scientist and one business decision-maker, the data scientist is often expected to produce all the documentation alone. moreover, documentation production can be refused by some team members who do not see its benefits for their own technical tasks. finally, data science is still a young profession and stakeholders often lack acculturation and experience: the anticipation of uncertainties is clearly difficult and time-consuming without experienced skill-carriers. in these circumstances, the definition of responsibilities can remain a weak point. in theory, the operational application of the databook requires a clear prior distinction between skills, individuals and responsibilities. skills are needed to produce qualification metadata, as presented in section 0: they are essential in the databook construction. individuals can carry one or more skills, and each of their skill can be tainted with different maturity level. the maturity level is key for anticipation: an inexperienced team member is usually able to document his production only a posteriori or execute a qualification procedure only if it has been predefined by a more experienced skill carrier. responsibilities are defined according the project specificities and individual skill range. usually, this definition is realized by the project manager for the duration of the project. but the distinction of these three concepts and their articulation is often more confusing in practice. surprisingly, the databook appears as a good communication facilitator that can be used for responsibilities clarification. 3.3 pilot evaluation since the prototype first tests, the databook prototype was freely accessible to several data science teams in a leading french data science company called quinten. its appropriation continued on data science projects between 2017 and 2020, i.e. more than a hundred of projects in health, perfume, insurance, banking, media and industry sectors. several completed databooks have been reused from one project to another, mainly for projects for one given company, using the same data sources or with similar usages and data objects. the pilot phase main qualitative feedback confirms the precedent advantages and limitations: it is summarised in the appendices (figure 13). this confrontation between theory and various fields reality still must be considered as potentially biased by the data science practice of one company. quinten is characterised by its own values, business offer and managerial practices. most of the usages produced through the company’s artificial intelligence projects are aimed at human users in highly regulated domains, thus transparency of the analytical work remains a priority. the following proposal is an unprecedented attempt to standardise the most useful documentation principles and functionalities. the databook, as a flexible framework should then be tested in other contexts. https://doi.org/10.29173/iq989 15/34 nesvijevskaia, anna (2021) databook: a standardised framework for dynamic documentation of algorithm design during data science projects, iassist quarterly 45(2), pp. 1-34. doi: https://doi.org/10.29173/iq989 3.4 the databook as a standard documentation framework the databook is founded on the principle of distinction between the realization of analytical outputs on the critical path of algorithm design and the production of the documentation. both dynamics imply similar skills but involve them at different stages and with different purposes. the databook as a documentation framework provides an opportunity to share the story of how data is transformed into useful information during a collaborative data science project. it can be used as a dynamic device for capitalising on knowledge, a material object that helps to gradually retrace the memory of the project and to give transparency to the resulting algorithmic model. the databook guarantees and respects by design the fair principles by guiding the construction of findable, accessible, interoperable and reusable data and metadata, as far as the business context allows the sharing of sensible data. for the data science projects, it provides the same advantages as the use of the fair principles in crossinstitutional projects (hansen et al. 2019): project planning, navigation through project changes, information sharing and data sharing outside the group. from a more operational point of view, the databook structure, illustrated in appendices (figures 14 to 24), provides three clear reading grids for each different purpose. 3.4.1 cross-skill data object qualification the databook can be used by all project actors to qualify one given data object and determine together its further treatment. it is the most basic use, achievable independently for each module from 1 to 6. this cross-skill use promotes convergence towards the most relevant result and stimulates the productivity of the project team. the convergence is accelerated by more convenient representations of data, such as graphs, explanations and other reports listed in module c. the convergence is also improved by module b, representing the collective arbitration procedure. in parallel, the dynamics of the convergence are monitored in module a through the statuses. in this situation, the project in less piloted through iterative or overlapped phases than by the progress of the qualification of one given data object. the databook guide in module 0 can facilitate this reading grid by presenting in priority core structure elements (horizontal in figure 4). for this application, the databook should be read starting from the module containing the data object and then viewing the project management modules. for example, figure 5 illustrates how to use the databook when the complete project team needs to qualify together the analytical results. https://doi.org/10.29173/iq989 16/34 nesvijevskaia, anna (2021) databook: a standardised framework for dynamic documentation of algorithm design during data science projects, iassist quarterly 45(2), pp. 1-34. doi: https://doi.org/10.29173/iq989 figure 5 databook reading grid for cross-skill qualification of a data object 3.4.2 skill capitalisation the databook can also be used by one given skill carrier who wishes to gain experience on all the data objects qualified and processed during the project. in this case, his attention will only focus on the qualification criteria that confer his skills on modules 1 to 6. for example, the databook can be used by a strategic manager to understand how the components of the algorithm are related to the business priorities, by a data steward who is interested in the semantics of all elements processed during the project or by a data engineer to optimise the exploration pipeline. the capitalisation is also very interesting for project management. indeed, the difficulties solved throughout the project constitute a rise in maturity: the estimation of the benefits and remaining uncertainties on the usages at the end of the project can lead to a broader roadmap than a technical deployment of the result. all the knowledge capitalised can be qualitatively consolidated in module 6b in the form of a return on experience from each of the skills’ carriers involved. this capitalisation is essential to generate new ideas of usages and gain in productivity for further data projects (data projects portfolio management) requiring similar data objects or skills. the databook guide in module 0 can facilitate this reading grid by presenting in priority metadata structure elements (vertical in figure 4). for this application, the databook should be read starting from each module from 1 to 6 containing metadata of the skill of interest, and then list the generated knowledge in module 6b, as presented in figure 6. https://doi.org/10.29173/iq989 17/34 nesvijevskaia, anna (2021) databook: a standardised framework for dynamic documentation of algorithm design during data science projects, iassist quarterly 45(2), pp. 1-34. doi: https://doi.org/10.29173/iq989 figure 6 databook reading grid for skill capitalisation from all data objects 3.4.3 final algorithm understanding the databook traces of all the data objects that compose the final algorithm. this trace is the documentation of the critical path: all the data elements notified as included are contained in the final versions of each intermediate output. this documentation is key for many purposes. first, it is a detailed specification for usage deployment. then, it makes the algorithm auditable, including for external actors. and, finally, this traceability gives the possibility to propagate the value generated by the usage back on data sources: this inverse value cascade is hence an original tool for evaluating data assets from both patrimonial and operational views. all these purposes follow the same reading grid illustrated in figure 7. for example, within the framework of an audit of the final algorithm, the investigation will consist in going back from the operational usage to the mathematical algorithm, then to the data which feeds this algorithm, then to a set of source data. these source data will then point to key business concepts chosen in the construction of the algorithm. as the audit is interested only in used data elements included in the algorithm, his reading is focused on the diagonal in figure 4 as soon as we consider that the final status of each type of data object is qualified by the skill carrier that produced the data object. if the auditor of the algorithm must go further to understand the design process, he may also be interested in project management metadata. this reading can be facilitated by the guide, especially by filtering the entire databook only on elements with a status included or by zooming on more detailed project management metadata. https://doi.org/10.29173/iq989 18/34 nesvijevskaia, anna (2021) databook: a standardised framework for dynamic documentation of algorithm design during data science projects, iassist quarterly 45(2), pp. 1-34. doi: https://doi.org/10.29173/iq989 figure 7 databook reading grid for the final algorithm understanding these different reading grids, not exhaustively described above, can fit the priorities chosen for each project or purpose. this flexibility remains one of the advantages of such a documentation framework and justifies databook definition as a new boundary object, like a portolan chart for navigating through the algorithm’s metadata. conclusion the databook emerges from the urgent documentation needs of data project stakeholders in the field and from interdisciplinary concepts inspiring the gap filling in the standard state-of-the-art data science process, crisp_dm. its structure is based on one main principle: in these exploratory algorithm design projects, each phase realization needs specific skills and all these skills are required to progressively adjust the entire process by producing dynamic documentation. documentation is then the result of both anticipatory and informative qualification work, and the documentation process generates a faster convergence on the best project results. the databook traces this dynamic and constitutes a boundary object for all the stakeholders. as the responsibilities and skills of a data scientist are still poorly and heterogeneously defined, databook description sheds some light not only on the essential skills but also on their mobilisation mechanism. as a prototype, it is decanted and confirmed as a very efficient human-data mediation facilitator. it can still be improved in terms of ergonomics, but its simple and flexible core structure completed with free metadata structure remains compatible with classical tools and independent of the data project purpose and pace. outside of https://doi.org/10.29173/iq989 19/34 nesvijevskaia, anna (2021) databook: a standardised framework for dynamic documentation of algorithm design during data science projects, iassist quarterly 45(2), pp. 1-34. doi: https://doi.org/10.29173/iq989 projects, the databook is useful for the development of a company's data assets and the governance of data quality. besides these theorical and operational considerations, one of the main benefits of the databook remains its contribution to algorithm transparency in business companies. this lack of transparency is too often reduced to the algorithm learning process, especially for the deep learning. however, an algorithm is much more than that: the databook can reveal all the objects that constitute it from end to end, but also the human choices than have driven its progressive design. the fear of this lack of transparency, crystalized after several scandals in the last years, lead to a search for a french and european position that has so far been unsuccessful, and to the affirmation of founding principles such as loyalty and vigilance by the cnil in 2017 (falque-pierrotin et al. 2017). these principles are reflected in a set of recommendations, such as ethics training of all implied actors, the mediation between users to make algorithms more understandable, the subordination of algorithms to human freedom and to the general interest from the design phase, or the creation of a national algorithm audit platform. capitalisation on french assets such as cultural values oriented towards people and ethics or the quality of training in engineering sciences and mathematics, is already mobilised, as presented by inria's annual report and its publications in 2017. while the contributions to the algorithm documentation framework remain limited and mainly oriented towards the control of external algorithms3, this proposal offers the possibility for all algorithm designers to achieve transparency of their own algorithms. acknowledgement i would like to gratefully thank professor ghislaine chartron for her determined and caring supervision during my thesis years and her valuable feedback on this paper. many thanks to quinten’s team that made possible these data science projects’ observations and the confrontation of the databook prototype to real life. i am also very thankful to catherine lesperance for the careful and thorough revision of this article in english. references arruabarrena, b., kembellec, g., & chartron, g. (2019, march). data littératie & shs : développer des compétences pour l’analyse des données, presented at the codata data value chain, val d’europe. austin, r. d., joshi, m. p., su, n., & sundaram, a. k. (2021). why so many data science projects fail to deliver. mit sloan management review, spring 2021. retrieved from https://sloanreview.mit.edu/article/why-so-many-data-science-projects-fail-to-deliver/ berti-equille, l. (2012). la qualité et la gouvernance des données : au service de la performance des entreprises, paris; cachan: hermes science publications. bertino, e., bernstein, p., agrawal, d., … widom, j. (2011). challenges and opportunities with big data, cyber center publications. boyd, d., & crawford, k. (2012). critical questions for big data. information, communication & society, 15(5), 662–679. broudoux, é., & scopsi, c. (2011). introduction. études de communication, (36), 9–22. https://doi.org/10.29173/iq989 https://sloanreview.mit.edu/article/why-so-many-data-science-projects-fail-to-deliver/ 20/34 nesvijevskaia, anna (2021) databook: a standardised framework for dynamic documentation of algorithm design during data science projects, iassist quarterly 45(2), pp. 1-34. doi: https://doi.org/10.29173/iq989 brynjolfsson, e., hitt, l. m., & kim, h. h. (2011). strength in numbers: how does data-driven decisionmaking affect firm performance? (ssrn scholarly paper no. id 1819486), rochester, ny: social science research network. camiciotti, l., & racca, c. (2015). creare valore con i big data. gli strumenti, i processi, le applicazioni pratiche, 1 edizione, milano: edizioni lswr. chapman, p. (1999). the crisp-dm user guide, presented at the brussels sig meeting, ncr systems engineering copenhagen. cottin, m., & nesme, m.-f. (2017). la qualité : variations autour d’une notion essentielle, quality: variations on an essential notion. i2d – information, données & documents, 53(4), 28–29. davenport, t. h., & patil, d. j. (2012). data scientist: the sexiest job of the 21st century. harvard business review, 90(5), 70–76. doucet, c. (2010). la qualité, paris: presses universitaires de france. ermine, j.-l. (2003). la gestion des connaissances, hermes lavoisier. falque-pierrotin, i., mahjoubi, m., & villani, c. (2017). comment permettre à l’homme de garder la main ? rapport sur les enjeux éthiques des algorithmes et de l’intelligence artificielle, cnil. retrieved from https://www.cnil.fr/sites/default/files/atoms/files/cnil_rapport_garder_la_main_web.pdf fayyad, u., piatetsky-shapiro, g., & smyth, p. (1996). the kdd process for extracting useful knowledge from volumes of data. communications of the acm, 39(11), 27–34. hansen, z. n., kruse, f., & thestrup, j. b. (2019). managing data in cross-institutional projects. iassist quarterly, 43(3), 1–10. loshin, d. (2010). master data management, morgan kaufmann. manyika, j., chui, m., brown, b., … hung byers, a. (2011). big data: the next frontier for innovation, competition, and productivity, mckinsey global institute. retrieved from https://www.mckinsey.com/business-functions/mckinsey-digital/our-insights/big-data-the-nextfrontier-for-innovation mariko, d. (2016). le master data management (mdm) et la qualité des données de l’entreprise : synergies digitales et collaboratives, intd-cnam. mayère, a. (1990). pour une économie de l’information, c.n.r.s. editions. doi: doi.org/10.3917/cnrs.mayer.1990.01 mayer-schönberger, v., & cukier, k. (2013). big data: a revolution that will transform how we live, work, and think, houghton mifflin harcourt. miller, h. g., & mork, p. (2013). from data to decisions: a value chain for big data. it professional, 15(1), 57–59. https://doi.org/10.29173/iq989 https://www.cnil.fr/sites/default/files/atoms/files/cnil_rapport_garder_la_main_web.pdf https://www.mckinsey.com/business-functions/mckinsey-digital/our-insights/big-data-the-next-frontier-for-innovation https://www.mckinsey.com/business-functions/mckinsey-digital/our-insights/big-data-the-next-frontier-for-innovation file:///c:/users/oschwart/stokes/iq/iq45_2/doi.org/10.3917/cnrs.mayer.1990.01 21/34 nesvijevskaia, anna (2021) databook: a standardised framework for dynamic documentation of algorithm design during data science projects, iassist quarterly 45(2), pp. 1-34. doi: https://doi.org/10.29173/iq989 nesvijevskaia, a. (2017). value creating through data science projects : insufficiency of the standart workflows. nesvijevskaia, a. (2019, october 18). phénomène big data en entreprise : processus projet, génération de valeur et médiation homme-données (thesis), paris, cnam. retrieved from http://www.theses.fr/2019cnam1247 odeh, s., & chartron, g. (2016). acteurs et économie des métadonnées du livre en france : analyse et avenir. documentation et bibliothèques, 62(1), 21–32. piatetsky-shapiro, g. (1994). an overview of knowledge discovery in databases: recent progress and challenges. in rough sets, fuzzy sets and knowledge discovery, springer, london, pp. 1–10. provost, f., & fawcett, t. (2013). data science and its relationship to big data and data-driven decision making. big data, 1(1), 51–59. shearer, c. (2000, fall). the crisp-dm model : the new blueprint for data mining. journal of data warehousing, pp. 13–22, pp. 13–22. star, s. l., & griesemer, j. r. (1989). institutional ecology, `translations’ and boundary objects: amateurs and professionals in berkeley’s museum of vertebrate zoology, 1907-39. social studies of science, 19(3), 387–420. wang, r. y. (1998). a product perspective on total data quality management. commun. acm, 41(2), 58–65. wirth, r., & hipp, j. (2000). crisp-dm: towards a standard process model for data mining. in proceedings of the fourth international conference on the practical application of knowledge discovery and data mining, pp. 29–39. zacklad, m., cahier, j.-p., bénel, a., zaher, l., lejeune, c., & zhou, c. (2007). hypertopic: une métasémiotique et un protocole pour le web socio-sémantique. in actes des 18eme journées francophones d’ingénierie des connaissances, francky trichet, p. 13. https://doi.org/10.29173/iq989 http://www.theses.fr/2019cnam1247 22/34 nesvijevskaia, anna (2021) databook: a standardised framework for dynamic documentation of algorithm design during data science projects, iassist quarterly 45(2), pp. 1-34. doi: https://doi.org/10.29173/iq989 appendices 1. reference outputs mapping figure 8 reference outputs of the adjusted crisp_dm model and their classification the figure 8 presents a mapping between the main outputs of the adjusted crisp_dm and the types of outputs presented in this paper. the crisp_dm adjustment consists in rearranging its outputs in order to isolate for each phase its main intermediate analytical output (in blue), determining the dominant skill for this output and grouping the documentation outputs (in grey and red) around this dominant skill. the dominant skills are, in activity order: strategy management, data stewardship, data engineering, machine learning, business intelligence and product ownership/knowledge management. the documentation outputs can be produced collectively or with the help of business analysis and data analysis skills. finally, the project management skill remain transversal to produce associated mediation milestones outputs (in green). individual skills’ carriers and project roles are not considered in this mapping. phases of adjusted crisp_dm model reference outputs critical path analytical outputs critical path analytical outputs documentation mediation milestones terminology project plan background risks and contingencies costs and benefits business objectives business success criteria assessment of tools and techniques (initial and final) requirements, assumptions, and constraints data science success criteria data science goals inventory of resources raw data data collection report data description report data exploration report exploration report data quality report data set description data set selected data rationale for inclusion / exclusion data inc./exc. method data cleaning report data treatment description derivated attributes generated records merged data reformated data modeling technique modeling assumptions test design model description analytical results benchmark (model assessments results) parameter setting (initial and revised) model assessment of data mining results w.r.t. business success criteria approved models description review of process approved models selected models list of possible actions & decisions restitution format final presentation deployment plan monitoring and maintenance plan experience documentation functional results description knowledge project roadmap usage roadmap & knowledge capitalization automated model structured data models perimeter description source data description model structure description analytical results description business understanding data preparation modelling evaluation deployment data understanding https://doi.org/10.29173/iq989 23/34 nesvijevskaia, anna (2021) databook: a standardised framework for dynamic documentation of algorithm design during data science projects, iassist quarterly 45(2), pp. 1-34. doi: https://doi.org/10.29173/iq989 2. examples of databook modules in this appendix, four illustrations of databook modules present its typical applications in artificial intelligence projects. these illustrations are extracted from complete project databooks in excel format and anonymised. for confidentiality reasons, communication of a complete databook prototype with qualified data objects is avoided. however, further real-world illustrations and practical details emerging from the testing phase in france can be found in appendix 11 of the multiple-case study (nesvijevskaia, 2019, appendix 11). figure 9 illustration of the module 2 from a project on compliance the first illustration (see figure 9 above) is extracted from a databook adapted to the context of a french branch of leading insurance company which worked on the detection of contracts with noncompliance risks in order to optimise the control process. this module aims the qualification of variables used for the detection model (module 5). the different components of this module include (from left to right): technical information about the sources of variables needed by the data engineer: reception batch, name of the table and database join key for the table. semantics of the variables: code and meaning of the variable (if the name does not exist in the database, or does not make sense, it is manually added in the databook). type of variable needed to choose the structuring methods (here, left unqualified). business priority of the exploration of each variable (exclusion of a variables by a business decision-maker, based on his perception of the compliance control process). flag of personal data to exclude (discussed with the dpo of the company). status of the variable: the excluded variables are eliminated from the final algorithms. “test forge”: conclusion of a custom analytical qualification based on a mathematical method of elimination of variables with no signal or too much noise (realized by the machine learner). usage of the variable: function of the variable in the learning matrix, qualified by the data engineer. the columns colours represent the different skill-carriers that produced the qualification. statut des variables exclue variables exclues de l'analyse incluse variables incluses dans l'analyse : voire nature d'usage nouvel extract attendu variables / tables nécessitant un extract complémentaire ? en cours d'analyse tableau des données reçues n° de lot table / file jointure(s) variable explication type de variable (az) (2 = texte 1 = num.) priorité métier donnee nominative statut test forge brute dérivée clé autre 1 sinmre15 nopol topcrac ? ? non non exclue non vide 1 sinmre15 nopol topasstr ? ? non non exclue non vide 1 sinmre15 nopol chargecie charge ? non non exclue ok ? 1 sinmre15 nopol regcie règlement ? non non exclue ok ? 1 sinmre15 nopol reccie recours encaissés ? non non exclue ok ? 1 sinmre15 nopol rapcie restant à payer ? non non exclue ok ? 1 ipfmre15 nopol dtmjne date de mise a jour du segment (aqqq) ? non non exclue ok 1 ipfmre15 nopol nopol numero de police (ancien) x(14) ? 1 non incluse ok x x 1 ipfmre15 nopol noint numero de l intermediaire ? 1 non incluse ok x 1 ipfmre15 nopol cdpole code pole ? 1 non incluse ok x 1 ipfmre15 nopol cmarch code marche ? 1 non incluse ok x 1 ipfmre15 nopol cdreg code region ? non non exclue ok 1 ipfmre15 nopol cdprod code produit ? 1 non incluse ok x 1 ipfmre15 nopol csegt code segment ? 1 non exclue non modalité fixe 1 ipfmre15 nopol cssegt code sous segment ? non non exclue ok 1 ipfmre15 nopol dtresilp date de resiliation police ? 1 non incluse ok x x 1 ipfmre15 nopol dttramvt date traitement dernier mvt du traite ? non non exclue ok 1 ipfmre15 nopol noaveder dernier numero d'avenant ? non non exclue ok 1 ipfmre15 nopol nopolori no de police compagnie precedente ? non non exclue ok 1 ipfmre15 nopol nocie numero de compagnie ? non non exclue non modalité fixe informations sur les fichiers reçus informations sur les variables reçues nature de l'usage de la variable https://doi.org/10.29173/iq989 24/34 nesvijevskaia, anna (2021) databook: a standardised framework for dynamic documentation of algorithm design during data science projects, iassist quarterly 45(2), pp. 1-34. doi: https://doi.org/10.29173/iq989 figure 10 illustration of the module 3 from a project on health insurance churn the second illustration (see figure 10 above) is extracted from a databook of a health insurance churn project aiming the generation of two different algorithmic models (prescriptive profile generation and predictive scoring approach). in this project, more than 200 variables were collected and transformed into 500 new variables: these new variables are documented in module 6. it is composed of (from left to right): position of the variable in the new table, code and type of the variable (discrete or continuous): these criteria are critical for the machine learner to use the variables in a learning matrix. the nomenclatures for the variable codes have been defined specifically for the project in order to maintain homogeneity with existing nomenclature methods (data analysis skills). semantics of the new variables: all the variables are organised by groups with homogeneous meaning (customer characteristics, contract characteristics, trends in past claims…) and then described one by one. these semantics correspond to a new data dictionary. type of structuring method to create the variable: closely linked to the data engineering pipeline, this qualification gives the possibility to see at a glance if the variable is identical to the source variable or if it is issued from a more complex treatment. the nomenclature of types of treatments have been defined specifically for the project context. the status of each variable is here split into two columns, each one corresponding to one of the two algorithms (predictive and prescriptive) used in the following step. each status is here binary, representing the inclusion (yes) or exclusion (no) of the variable for each algorithm. the final columns correspond to two specific data treatments for the models: business rules association and mathematical quantiles generation. each one is specifically documented in a complementary module. this module was produced entirely by the data scientist and controlled by the business expert. nomenclature des variables : nature des variables tx_q_xxxx_xxxx : taux ? = variables en attente dt_q_xxxx_xxxx : dates i ou k = dérivée première (source : variables covea) bc_q_xxxx_xxxx : booléenne (0 ou 1) ds = dérivée seconde (source : une variable dérivée première) cd_q_xxxx_xxxx : code dz = dérivée complexe (sources multiples covea et dérivées quinten) mt_q_xxxx_xxxx : montant dm = variable à découper par modalité (_x = modalités) nb_q_xxxx_xxxx : nombre voir onglet "index construction variables" pour plus de détails nu_q_xxxx_xxxx : numéro lb_q_xxxx_xxxx : libellé mm_q_xxxx_xxxx : mois (de 1 à 12) aa_q_xxxx_xxxx : année tableau des variables finales #col intitule_variable type de variable groupe description natureusage prédiction usage préscription spécificité des règles quantilisation 1 lb_q_vs discret churn variable de sortie : churner oui/non (voir onglet périmètre) dz oui oui 2 nu_affa discret affaire santé numéro d'affaire santé k non non 3 cd_type_affa discret affaire santé code type affaire santé k non non 4 cd_adhe discret souscripteur code adhérent i non non 5 cd_cr discret affaire santé code centre de responsabilité i non non 6 dt_effe_affa date affaire santé date de début d'affaire i non non 7 dt_fin_affa date technique date de fin de l'affaire (par défaut) k non non 8 dt_start date technique date de début de décompte des prestations santé k non non 9 dt_stop date technique date de fin de décompte des prestations santé k non non 10 dt_sais_evnm date technique date de churn (si non churner : "none") k non non 11 dt_effe_evnm_sour date churn renseignée pour les churners, correspond à la date de churn i non non 12 cd_moti_rslt_cont discret churn code du motif de résiliation du contrat santé i non non 13 bc_annu discret technique annulation de la résiliation k non non 14 nu_pcp_ede discret souscripteur numéro de souscripteur associé à l'affaire santé k non non 15 ssaa continue technique année de l'extraction des données k non non 16 nu_mois discret technique mois de l'extraction des données k non non 17 cd_type_evnm_ede discret technique evènement s28 = résiliation sens covea, "none"=non résiliation au sens coveak non non 18 nb_q_rnvl_affa continue affaire santé nombre de renouvellements, ie ancienneté de l'affaire santé (churner = date fin date début contrat // non churner = 31-12-2014 date début contrat) en annéesdz oui oui qt_6_ 19 cd_marc discret affaire santé code marché du souscripteur de l'affaire i oui oui oui 20 cd_rgim_asrc_sour_01 discret affaire santé existance d'un bénéficiaire 1 (regime general,volontaire,pers) i oui oui 21 cd_rgim_asrc_sour_02 discret affaire santé existance d'un bénéficiaire 2 (exploitants agricoles (amexa)) i oui oui 22 cd_rgim_asrc_sour_03 discret affaire santé existance d'un bénéficiaire 3 (profession independante (ampi), soit tns)i oui oui 23 cd_rgim_asrc_sour_60 discret affaire santé existance d'un bénéficiaire 60 (regime local alsace-moselle) i oui oui https://doi.org/10.29173/iq989 25/34 nesvijevskaia, anna (2021) databook: a standardised framework for dynamic documentation of algorithm design during data science projects, iassist quarterly 45(2), pp. 1-34. doi: https://doi.org/10.29173/iq989 figure 11 illustration of a usage roadmap presented in power point format and based on the module 6a the third illustration (see figure 11 above) is an example of representation of the usage roadmap in a graphical power point format. this graph represents eight actions that need to be validated before deployment after the project. each action is qualified in terms of: number and name of each action two types of categories of the actions: the colours represent the type of skills necessary for the deployment, and the red circles represent a qualitative priority judgement of the actions estimated value creation (vertical axis) and complexity (horizontal axis) this ergonomic representation is usually completed by a planning as soon as actions are validated. it is based on the module 6a of the databook (usage roadmap): the module is a matrix structure in excel. unsurprisingly, usually this excel file contains not only the qualified actions, presented above, but also more detailed tasks, associated with constraints, responsibilities, dates and charge estimation (time and budget). the charge is usually qualified by the different skills’ carriers of the project. the underlying excel illustration remains confidential. https://doi.org/10.29173/iq989 26/34 nesvijevskaia, anna (2021) databook: a standardised framework for dynamic documentation of algorithm design during data science projects, iassist quarterly 45(2), pp. 1-34. doi: https://doi.org/10.29173/iq989 figure 12 illustration of a synthesis of statuses by module used for a given project milestone the last illustration (see figure 12 above) is an example of synthesis of statuses issued from all the modules from 1 to 6a, presented in module a (project roadmap). for each module 1 to 6a, data elements to be qualified are counted and distributed by type of status. for each type of status and module, a mean workload of qualification estimated: this gives the possibility to translate the remaining qualification workload in man-days. this charge anticipation method can be applied when all the elements and mean charges are correctly anticipated but also for more agile project management when elements emerge progressively and are associated with individual short-term charges. it is then compatible with typical backlog-burning monitoring tools. https://doi.org/10.29173/iq989 27/34 nesvijevskaia, anna (2021) databook: a standardised framework for dynamic documentation of algorithm design during data science projects, iassist quarterly 45(2), pp. 1-34. doi: https://doi.org/10.29173/iq989 3. pilot feedback figure 13 feedback on the implementation of the databook in the field https://doi.org/10.29173/iq989 28/34 nesvijevskaia, anna (2021) databook: a standardised framework for dynamic documentation of algorithm design during data science projects, iassist quarterly 45(2), pp. 1-34. doi: https://doi.org/10.29173/iq989 4. databook: excel format the figures 14 to 24 illustrate each excel sheet of a databook with documentation of the first iterations of an imaginary churn prediction project. figure 14 – module 0: guide figure 15 – module a: project roadmap https://doi.org/10.29173/iq989 29/34 nesvijevskaia, anna (2021) databook: a standardised framework for dynamic documentation of algorithm design during data science projects, iassist quarterly 45(2), pp. 1-34. doi: https://doi.org/10.29173/iq989 figure 16 – module b: method of data inclusion/exclusion https://doi.org/10.29173/iq989 30/34 nesvijevskaia, anna (2021) databook: a standardised framework for dynamic documentation of algorithm design during data science projects, iassist quarterly 45(2), pp. 1-34. doi: https://doi.org/10.29173/iq989 figure 17 – module c: exploration reports figure 18 – module 1: perimeter © anna nesvijevskaia databook version 2020 exploration reports n° document name n° document part itemid version number version date document type comment document storage link milestone name milestone date milestone place participants of the milestone main decisions 1 iq quarterly paper on databook 0 1_0 final 02/04/2021 paper in open access databook description https://iassistquarterly.co m/ conference iassist 16/09/2021 gothenburg, suede researchers use the databook in data science projects 1 iq quarterly paper on databook 1 appendix figure 9 1_1 module 2 illustration fill your own module 2 1 iq quarterly paper on databook 2 appendix figure 10 1_2 module 3 illustration fill your own module 3 1 iq quarterly paper on databook 3 appendix figure 11 1_3 module 6a illustration fill your own module 6a 2 big data phenomenon… 0 2_0 final 2019 thesis pdf http://www.theses.fr/2019 cnam1247 defense of thesis 18/10/2019 paris, france jury & fans write an paper on the databook 2 big data phenomenon… 1 chapter 2.3.1 2_1 first databook concepts write an paper on the databook 3 project x databook v1 0 3_0 initial 01/01/2022 excel presentation of the structure in my computer here : xxx project committee n°1 05/01/2022 company x headquarters project manager, sponsor, xxx continue to use the databook 4 project x intermediary exploration report 0 4_0 final 01/01/2022 ppt fists exploration conclusions sent my mail : xxx project committee n°1 05/01/2022 company x headquarters project manager, sponsor, xxx exclude all personal data 5 project x dashboard 0 5_0 version 1 01/01/2022 powerbi details of the exploration in my computer here : xxx project committee n°2 01/02/2022 company x headquarters project manager, sponsor, xxx include semantics in the dashboard 5 project x dashboard 1 graph1 5_1 version 1 01/01/2022 powerbi name of graph 1 project committee n°2 01/02/2022 company x headquarters project manager, sponsor, xxx 6 project x final exploration report 0 6_0 final 01/02/2022 ppt all exploration conclusions sent my mail : xxx project committee n°3 01/03/2022 company x headquarters project manager, sponsor, xxx deploy the results 5 project x dashboard 2 graph2 5_2 version 2 01/02/2022 powerbi name of graph 2 project committee n°3 01/03/2022 company x headquarters project manager, sponsor, xxx include semantics in the dashboard … … … … … … … … … … … … … … … © anna nesvijevskaia databook version 2020 exploration perimeter n° scope nature n° business concept n° business concept itemid business concept description data source indentified data preparation comment impact on modellisation order of magnitude operational comment status 1 phenomenon of interest 1 active churn 0 1_1_0 resiliation of all the contracts on a customer active demand crm date of the last resiliation demand active churn is the variable to predict 15% per year retain active customers with high probability of churn included 1 phenomenon of interest 2 passive churn 1 non-conformity 1_2_1 resiliation of all the contracts by the compagny litigation base identify customers with litigation exclude passive churn from active churn to predict 2% per year never retain a "bad customer" included 1 phenomenon of interest 2 passive churn 2 passive churn 1_2_2 resiliation of all the contracts resulting from customer death impossible to source identify deceased customers exclude passive churn from active churn to predict 1% per year avoid to call deceased customers excluded 2 individuals (items to analyse) 1 private clients 0 2_1_0 private clients are prioritary to retain crm filter only on "pri" customers 300000 private, 20000 professional concentrate the calls on btoc vendors in progress 3 drivers 1 client characteristics 0 3_1_0 crm to be started 3 drivers 2 contracts characteristics 0 3_2_0 crm to be started 3 drivers 3 historical commercial contacts 0 3_3_0 crm to be started 4 time perimeter 1 prediction horizon 0 4_1_0 crm to be started 4 time perimeter 2 historical clients 0 4_2_0 all customers active in 2019 crm in progress 4 time perimeter 3 last contacts 0 4_3_0 consider only recent contacts (3 months) crm censure contacts on their date in progress 5 eval. criteria 1 target precision 0 5_1_0 crm to be started 5 eval. criteria 2 target interpretability 0 5_2_0 crm to be started 5 eval. criteria 3 time saving for vendors 0 5_3_0 crm to be started 6 key indicators 1 turnover 0 6_1_0 crm to be started 6 key indicators 2 fte 0 6_2_0 annual report to be started 6 key indicators 3 salary costs 0 6_3_0 rh database to be started … … … … … … … … … … … … … … https://doi.org/10.29173/iq989 31/34 nesvijevskaia, anna (2021) databook: a standardised framework for dynamic documentation of algorithm design during data science projects, iassist quarterly 45(2), pp. 1-34. doi: https://doi.org/10.29173/iq989 figure 19 – module 2: source data figure 20 – module 3: model structure © anna nesvijevskaia databook version 2020 source data n° source / document n° base n° table n° variable n° values itemid business priority description compliance volume completeness uniqueness accuracy operational comment status 1 annual report 0 0 0 0 1_0_0_0_0 2 ok to be started 2 crm 1 clients 1 clients 1 id_client 0 2_1_1_1_0 1 unique client identification anonymization needed 346453 lines yes yes yes to show on the screen included 2 crm 1 clients 1 clients 2 name_client 0 2_1_1_2_0 1 nominative data excluded 2 crm 1 clients 1 clients 3 postal_code 0 2_1_1_3_0 1 ok in progress 2 crm 1 clients 1 clients 4 address 0 2_1_1_4_0 1 identifying data excluded 2 crm 1 clients 1 clients 5 gender 1 m 2_1_1_5_1 1 ok no to be cleaned (m=mr=mister) in progress 2 crm 1 clients 1 clients 5 gender 2 mr 2_1_1_5_2 1 ok no to be cleaned (m=mr=mister) in progress 2 crm 1 clients 1 clients 5 gender 3 mrs 2_1_1_5_3 1 ok no to be cleaned (mrs=miss) in progress 2 crm 1 clients 1 clients 5 gender 4 miss 2_1_1_5_4 1 ok no to be cleaned (m=mr=mister) in progress 2 crm 1 clients 1 clients 5 gender 5 mister 2_1_1_5_5 1 ok no to be cleaned (m=mr=mister) in progress 2 crm 1 clients 1 clients 6 segment 1 pro 2_1_1_6_1 2 ok 23004 lines yes in progress 2 crm 1 clients 1 clients 6 segment 2 pri 2_1_1_6_2 1 ok 323449 lines yes in progress 2 crm 2 contracts 1 contracts 1 id_contract 0 2_2_1_1_0 1 anonymization needed in progress 2 crm 2 contracts 1 contracts 2 id_client 0 2_2_1_2_0 1 anonymization needed in progress 2 crm 2 contracts 1 contracts 3 type of contract 0 2_2_1_3_0 1 ok in progress 2 crm 2 contracts 1 contracts 4 start_date 0 2_2_1_4_0 1 ok in progress 2 crm 2 contracts 1 contracts 5 resiliation_date 0 2_2_1_5_0 1 ok in progress 2 crm 3 contacts 0 0 0 2_3_0_0_0 to be started 3 litigation applicaiton 1 litigation base 0 litigation list 0 id_client 0 3_1_0_0_0 1 anonymization needed in progress 3 litigation applicaiton 1 litigation base 0 litigation list 0 date of litigation 0 3_1_0_0_0 1 ok in progress 3 litigation applicaiton 2 reporting base 0 0 0 3_2_0_0_0 to be started 4 hr database 0 0 0 0 4_0_0_0_0 1 nominative data excluded 5 geographical table 0 0 0 0 5_0_0_0_0 1 postal codes description ok in progress … … … … … … … … … … … … … … … … … … © anna nesvijevskaia databook version 2020 model structure n° agregated table n° variable n° values itemid business priority description business concept unit type of treatment volume base separation order of magnitude accuracy operational comment status 1 main table 0 0 1_0_0 1 to run on a monthly basis in progress 1 main table 1 id_client 0 1_1_0 1 unique client identification number client filtered on "pri" and active contracts in time scope 323449 lines 60% learning 30% validation yes to show on the screen with associated client name included 1 main table 2 active_churn 1 yes 1_2_1 1 resiliation of all the contracts on a customer active demand churn flag: resiliation date in time scope + client without litigation 46900 lines 14,8% vs 15% yes in progress 1 main table 2 active_churn 2 no 1_2_2 1 active client or resiliation after litigation churn flag: others 276549 lines in progress 1 main table 3 postal_code 0 1_3_0 1 client characteristics raw in progress 1 main table 4 department 0 1_4_0 1 client characteristics derivation in progress 1 main table 5 region 0 1_5_0 1 client characteristics derivation in progress 1 main table 6 border area 0 1_6_0 1 client characteristics derivation in progress 1 main table 7 gender 1 male 1_7_1 1 client characteristics nb cleaning in progress 1 main table 7 gender 2 female 1_7_2 1 client characteristics nb cleaning in progress 1 main table 8 contract_a 0 1_8_0 1 contract characteristics nb flag derivation in progress 1 main table 9 contract_b 0 1_9_0 1 contract characteristics flag derivation in progress 1 main table 10 contract_c 0 1_10_0 1 contract characteristics flag derivation in progress 1 main table 11 last_contract_purchase 0 1_11_0 1 number of months after the last substription contract characteristics nb months date derivation in progress 1 main table 12 client_seniority 0 1_12_0 1 number of months after the first substription contract characteristics nb months date derivation in progress 1 main table 13 number_contacts 0 1_13_0 1 contacts characteristics nb aggregation to be started 1 main table 14 last_contact 0 1_14_0 1 number of days after the last contact contacts characteristics nb days date derivation to be started 1 main table 15 calls 0 1_15_0 2 contacts characteristics nb flag derivation to be started 1 main table 16 mails 0 1_16_0 2 contacts characteristics nb flag derivation to be started 1 main table 17 meetings 0 1_17_0 2 contacts characteristics nb flag derivation to be started … … … … … … … … … … … … … … … … … … https://doi.org/10.29173/iq989 32/34 nesvijevskaia, anna (2021) databook: a standardised framework for dynamic documentation of algorithm design during data science projects, iassist quarterly 45(2), pp. 1-34. doi: https://doi.org/10.29173/iq989 figure 21 – module 4: analytical results figure 22 – module 5: functional results figure 23 – module 6a: usage roadmap © anna nesvijevskaia databook version 2020 analytical results n° type of model n° model n° model parameters itemid business priority description auc volume lift coverage model interpretability operational comment status 1 scoring 1 xgboost 0 default 1_1_0 2 see final report 0.7821 29481 2,5 69% difficult never show a score to a vendor in progress 1 scoring 2 random forest 1 default 1_2_1 2 see final report 0.7048 32019 1,9 72% mean never show a score to a vendor excluded 1 scoring 2 random forest 2 4 trees, max depth 6 1_2_2 2 see final report 0.7648 30245 2,4 70% mean never show a score to a vendor in progress 1 subgroup identification 1 qfinder 1 order 2, zscore, lift>1,5 1_1_1 1 see final report na 30294 2,2 65% easy associate a lever to each profile excluded 1 subgroup identification 1 qfinder 2 order 3, zscore, lift>2 1_1_2 1 see final report na 14239 3,6 43% easy associate a lever to each profile in progress … … … … … … … … … … … … … … … … © anna nesvijevskaia databook version 2020 functional results n° type of result n° result itemid business priority description mean nb of calls per week target precision time saving result interpretability success rate pilot qualitative feedback turnover saved status 1 existing 0 1_0 3 global retention plan 35 18% 0% yes 25% vendors can not handle more than 35 calls per week in progress 1 existing 1 profile a 1_1 3 recent clients (<2 months) 20 15% yes 26% in progress 1 existing 1 profile b 1_1 3 parisian ancient clients 15 22% yes 24% in progress 2 scoring 2 score flag 1 2_2 2 clients with high churn risk (>75%) 21 36% difficult 20% vendors do not know why thei have to call excluded 1 subgroup identification 1 profile 1 1_1 1 recent customers (<4 months) with only 1 contract 15 38% yes 35% it works when vendors propose a satisfaction interview and cross-sell if satisfied in progress 1 subgroup identification 2 profile 2 1_2 1 in progress 1 subgroup identification 3 profile 3 1_3 1 in progress 1 subgroup identification 4 profile 4 1_4 1 in progress 1 subgroup identification 5 profile 5 1_5 1 in progress … … … … … … … … … … … … … © anna nesvijevskaia databook version 2020 usage roadmap n° action n° task itemid business priority data ressources need time needed risk cost estimated benefits remaining uncertainties status 1 analytic solution deployment 0 1_0 1 see databook final version 10 days low low high delayed 2 monitoring deployment 1 add profile monitoring 2_1 1 see monitoring interface in powerbi 5 days low low medium validated 2 monitoring deployment 2 add vendor performance monitoring 2_2 1 see monitoring interface in powerbi 7 days low low high managerial levers ? validated 3 maintainance anticipation 1 once per year, look for new profiles 3_1 1 see databook final version 10 days medium low low churn stability delayed 3 maintainance anticipation 2 create a hotline for vendors 3_2 1 60 days / year low high low train marketing team ? excluded 4 communication 2 explain the profiles to vendors 4_2 1 30 days medium medium high delayed … … … … … … … … … … … … https://doi.org/10.29173/iq989 33/34 nesvijevskaia, anna (2021) databook: a standardised framework for dynamic documentation of algorithm design during data science projects, iassist quarterly 45(2), pp. 1-34. doi: https://doi.org/10.29173/iq989 figure 24 – module 6b: return of experience © anna nesvijevskaia databook version 2020 return of experience n° idea domain n° idea itemid interested stakeholders remaining uncertainties status 1 project management 1 better define who fill which item in the databook 1_1 data scientists team organisation for the next project ? to be shared 2 data analysis 1 change the crm interface in order to choose only between "mr" & "mrs" for each customer 2_1 it / crm team is it technically possible ? to be shared … … … … … … … … https://doi.org/10.29173/iq989 34/34 nesvijevskaia, anna (2021) databook: a standardised framework for dynamic documentation of algorithm design during data science projects, iassist quarterly 45(2), pp. 1-34. doi: https://doi.org/10.29173/iq989 endnotes 1 anna nesvijevskaia is doctor of the conservatoire national des arts et métiers in science of information and communication and associate researcher at the laboratory dicen ile de france. she is also partner at quinten, expert firm in artificial intelligence, and can be reached by email: anna.nesvijevskaia@gmail.com (version: may 2021) 2 brizo_ds is a model of data project device fundamentally orientated towards value generation through the exploitation of usages, including knowledge capitalisation. it is intended to reduce the uncertainties inherent in these exploratory projects and is transferable to the scale of enterprise data project portfolio management. the device includes an adjusted crism_dm model, completed and reorganised in a gantt chart to facilitate project management. beyond the initials of the three reference indicators coordinating the trade-offs during the project between benefits, resources and incertitudes, the name of the model is inspired by the greek goddess brizo, bearer of prophetic dreams and protector of sailors: the art of predicting the future through dreams is indeed a major asset for a data science project, by nature exploratory in an uncertain environment, and aiming to arrive ‘safely’, i.e. on a value-generating usage, whether anticipated or not. 3 https://www.inria.fr/fr/pour-une-regulation-des-algorithmes https://doi.org/10.29173/iq989 https://www.inria.fr/fr/pour-une-regulation-des-algorithmes 1/10 hettne, kristina maria; proppert, ricarda; nab, linda; rojas-saunero, l. paloma; gawehns, daniela (2020) reprohacknl 2019: how libraries can promote research reproducibility through community engagement, iassist quarterly 44(1-2), pp. 1-10. doi: https://doi.org/10.29173/iq977 reprohacknl 2019: how libraries can promote research reproducibility through community engagement kristina maria hettne1, ricarda proppert, linda nab, l. paloma rojas-saunero, daniela gawehns abstract university libraries play a crucial role in moving towards open science, contributing to more transparent, reproducible and reusable research. the center for digital scholarship (cds) at leiden university (lu) library is a scholarly lab that promotes open science literacy among leiden’s scholars by two complementary strategies: existing top-down structures are used to provide training and services, while bottom-up initiatives from the research community are actively supported by offering the cds’s expertise and facilities. an example of how bottom-up initiatives can blossom with the help of library structures such as the cds is reprohack. reprohack – a reproducibility hackathon – is a grass-root initiative by young scholars with the goal of improving research reproducibility in three ways. first, hackathon attendees learn about reproducibility tools and challenges by reproducing published results and providing feedback to authors on their attempt. second, authors can nominate their work and receive feedback on their reproducibility efforts. third, the collaborative atmosphere helps building a community interested in making their own research reproducible. a first reprohack in the netherlands took place on november 30th, 2019, co-organised by the cds at the lu library with 44 participants from the fields of psychology, engineering, biomedicine, and computer science. for 19 papers, 24 feedback forms were returned and five papers were reported as successfully reproduced. besides the researchers’ learning experience, the event led to recommendations on how to enhance research reproducibility. the reprohack format therefore provides an opportunity for libraries to improve scientific reproducibility through community engagement. keywords open science, reproducibility, hackathon, grassroot initiative, community engagement introduction university libraries around the world play a crucial role in open science, contributing to more transparent, reproducible and reusable research by offering training and services to their researchers. the center for digital scholarship (cds) at leiden university (lu), the netherlands, is a scholarly lab located in the lu library. the cds employs two complementary strategies to improve open data literacy among leiden’s scholars: existing top-down structures are used to provide ’training2’, such as a brief training on data publishing, and a more extensive data fairification workshop, where researchers get hands on experience with making their data findable, accessible, interoperable and reusable (fair) (wilkinson, et al., 2016) according to recent recommendations (jacobsen, et al., 2019). complementary to such training, bottom-up initiatives are actively supported by offering the cds’s expertise and facilities. examples of such bottom-up initiatives initiated by leiden’s research community are data management workshops addressing institute-specific data challenges, for https://doi.org/10.29173/iq977 https://www.library.universiteitleiden.nl/about-us/centre-for-digital-scholarship/workshops 2/10 hettne, kristina maria; proppert, ricarda; nab, linda; rojas-saunero, l. paloma; gawehns, daniela (2020) reprohacknl 2019: how libraries can promote research reproducibility through community engagement, iassist quarterly 44(1-2), pp. 1-10. doi: https://doi.org/10.29173/iq977 example, ‘pilot projects3’ on making book history data fair and an ongoing initiative developing ‘guidelines4’ on how to share medical resonance imaging data. while many libraries make an effort to improve data access, few have focussed on reproducibility as another crucial aspect of open science literacy. reproducibility can be defined as ‘data and code being available to fully rerun the analysis’ (the turing way community, et al., 2019) and has become a widely discussed issue over the past decade (e.g., for an early discussion on the reproducibility crisis in psychological science, see (pashler & wagenmakers, 2012). in 2016, a nature survey (baker, 2016) listed factors contributing to irreproducible research. the top factors, such as intense competition and time pressure, can be tackled by scientific organizations adjusting their policies and incentive structures in a top-down approach. however, factors such as ‘methods, code unavailable’, ‘raw data not available from original lab’ and ‘insufficient peer review’ may be more effectively tackled in a bottom-up fashion by enabling the scientific community in their practices, and provide a low hanging fruit for librarians to pick. one initiative that addresses these factors is ‘reprohack5’. reprohack is a grass-root initiative by young scholars with the goal of improving research reproducibility in three ways: first, hackathon attendees learn about reproducibility tools and challenges by reproducing published results and providing feedback to authors on their attempt. second, authors can nominate their own work and receive feedback on their reproducibility efforts. third, the collaborative atmosphere at the event helps build an interdisciplinary community among researchers interested in making their own research reproducible. the first reprohack in the netherlands took place on november 30th, 2019, co-organised by the csd at the lu library (see ‘website6’ and ‘blog7’). in this paper, we will reflect on the organisation and output of the reprohack format, and the role libraries can play in the process. finally, we will provide recommendations for reproducibility practices based on participants’ feedback and provide resources by which support centers such as libraries can organize their own reprohack together with researchers. event setting the event was hosted at the lu library and library staff was actively involved in the organisation. by providing a location unrelated to a specific department or research field, the organizers hoped to attract a diverse group of attendants. the organizing team itself had diverse backgrounds and consisted of two researchers from university medical centers, one researcher at a computer science institute, a student at the social science faculty and a library staff member. information about the event was spread via a twitter account, newsletters, public agendas and the organizers’ individual channels, such as at field-specific meetings, workshops or by word of mouth. as the meeting rooms were offered free of charge, the organizers had spare funds to provide lunch and drinks for the participants and print stickers as well as posters. the full day event was held on a saturday. fourtyfour participants attended the event and brought in their experiences from the fields of psychology, engineering, biomedicine, and computer science. https://doi.org/10.29173/iq977 https://digitalscholarshipleiden.nl/articles/durable-access-to-book-historical-data https://github.com/dorienhuijser/decisiontreemridata https://github.com/reprohack/reprohack-hq https://reprohacknl.github.io/reprohack/ https://www.software.ac.uk/blog/2020-01-15-reproducibility-hackathon-netherlands-aftermath 3/10 hettne, kristina maria; proppert, ricarda; nab, linda; rojas-saunero, l. paloma; gawehns, daniela (2020) reprohacknl 2019: how libraries can promote research reproducibility through community engagement, iassist quarterly 44(1-2), pp. 1-10. doi: https://doi.org/10.29173/iq977 event structure two speakers framed the event, the first one introducing participants to current developments on tools for reproducible research and the second one putting reproducibility into the broader context of open science. in between, participants worked on ‘hacking’ scientific papers by testing and troubleshooting the papers’ reproducibility and documenting what hurdles they encountered. the event was spontaneously extended with an improvised hands-on tutorial on how to make computational code citable. during finalizing drinks and bites, participants further discussed their experiences of the day and got more opportunity to build their network. the schedule of the day can be found in box 1. box 1. schedule of the reprohacknl day. feedback on the papers attendees reported their findings on the attempt of reproducing the chosen papers using a standardized feedback form (table 1). forms were returned by email to all authors. empty cell question answer type 1 name of participant text field 2 which paper did you attempt to reproduce? text field 3 did you manage to reproduce it? yes/no/almost 10:00 coffee and tea 10:30 welcome 10:40 ‘reproducibility: the value in the practice’ by dr. anna krystalli 11:30 forming groups and start hackathon 12:30 lunch buffet 14:30 ‘a vision for open science beyond the reproducibility crisis’ by dr. john boy 15:00 continue hacking 16:30 drinks and bites https://doi.org/10.29173/iq977 4/10 hettne, kristina maria; proppert, ricarda; nab, linda; rojas-saunero, l. paloma; gawehns, daniela (2020) reprohacknl 2019: how libraries can promote research reproducibility through community engagement, iassist quarterly 44(1-2), pp. 1-10. doi: https://doi.org/10.29173/iq977 4 on a scale from 1 to 10, how much of the paper did you manage to reproduce? scale 1 (none of it) to 10 (all of it) 5 briefly, describe the procedure followed/tools used to reproduce it text field 6 what were the positive features of this approach? text field 7 any other comments/suggestions on the reproducibility approach? text field 8 how well was the material documented? scale 1 (hard to navigate) to 10 (very well) 9 how could the documentation be improved? text field 10 what did you like about the documentation? text field 11 after attempting to reproduce, how familiar do you feel with code and method used in the paper? scale 1 (not at all familiar) to 10 (fully walked through the analysis code) 12 any suggestions on how the analysis could be more transparent? text field 13 rate the project on the reusability of the material scale 1 (not reusable) to 10 (easily reusable) 14 are material clearly covered by a permissive enough license to build on? checkbox permissive license for data included checkbox permissive license for code included 15 any suggestions on how the project could be more reusable? text field 16 any final comments? text field 17 contact email text field table 1. feedback form to be filled in by participants. thirty-one papers on different topics were available for reproducibility checks during the event. the papers were not curated by the organisation. the only inclusion criteria for papers were that authors agreed to their use during the event, and that associated code and data were available online. out of the 31 available papers, 19 papers were freely chosen by participants to work on during the event and 24 feedback forms were provided to their authors. five of the 19 papers were reported as successfully https://doi.org/10.29173/iq977 5/10 hettne, kristina maria; proppert, ricarda; nab, linda; rojas-saunero, l. paloma; gawehns, daniela (2020) reprohacknl 2019: how libraries can promote research reproducibility through community engagement, iassist quarterly 44(1-2), pp. 1-10. doi: https://doi.org/10.29173/iq977 reproduced (answer ‘yes’ to question 3) and six were reported as almost reproduced (answer ‘almost’ to question 3). feedback form scores are shown in figure 1. the organisation made an effort to frame the event as a learning opportunity rather than a performance assessment, hoping to make this an accessible event also for researchers with less experience with reproducible workflows. despite this communication strategy, it remained a major challenge to convince authors to submit their work for the event. the top three reasons why contacted researchers were hesitant to share their work for reproducibility checks were: ‘my code is not nice enough to be shared’, ‘i work with sensitive data that cannot be shared’, ‘i don't want to think about this work i did two years back anymore’. these responses speak to the organisers’ impression that, unfortunately, openly sharing code and data remains an uncomfortable thing to do within the current academic climate. issues related to reproducibility reproducibility (the turing way community, et al., 2019) is the extent to which analysis results can be reproduced given the code and the exact same data as in the original analysis. reproducibility scores ranged from 1 (not reproducible) to 10 (fully reproducible), and 12 feedback forms reported a score higher than 5 (question 4). when participants were asked to rate the documentation and annotation of code, 21 feedback forms reported a score higher than 5 (question 10). thus, there were more documentation scores higher than 5, than reproducibility scores higher than 5. in agreement with these findings, participants reported a number of issues other than documentation in the openly phrased feedback questions (question 5, question 7 and question 12), which caused them to be unable to reproduce the paper: unavailability of data (locked behind a paywall, not available due to privacy issues, time-outs during download), faulty code, and dependence of analyses on proprietary or platform-specific software. various recommendations for improving code documentation (question 9) were provided. first, it was suggested to include a codebook, as often used in survey research, which provides information about the data structure by explaining the survey structure and aids data interpretation and transparency. a second recommendation was adding a readme text file, which explains the nuances and context of a unique data collection and has been deposited to a data repository. further, suggestions included better explanation of how the different parts of the code correspond to results in the paper, more and clearer comments in the code itself, performing a check for typos, and providing orientation on time needed for the code to run. issues related to reusability reusability (the turing way community, et al., 2019) is the extent to which the code and provided data can be used in future studies and how easily the material can be built upon by other researchers. the reported scores for reusability varied between 1 (not reusable) and 10 (easily reusable), and 15 feedback forms reported a score higher than 5 (question 13). less than half (10) of the feedback forms reported the existence of a permissive license for either data or code (question 14). when participants were asked how familiar they were with the code and method after trying to reproduce, answers varied between 1 and 10, with 12 reporting a score higher than 5 (question 11). when asked to provide https://doi.org/10.29173/iq977 6/10 hettne, kristina maria; proppert, ricarda; nab, linda; rojas-saunero, l. paloma; gawehns, daniela (2020) reprohacknl 2019: how libraries can promote research reproducibility through community engagement, iassist quarterly 44(1-2), pp. 1-10. doi: https://doi.org/10.29173/iq977 suggestions to further improve the reusability, adding a permissive license was the most frequent answer (6 times) (question 15). figure 1. distribution of feedback form scores for likert-type questions 4 (a), 8 (b), 11 (c) and 13 (d). comparison across teams three papers were reproduced by more than one team each, which allows for comparison of scores between teams. we compared the feedback form scores from the questions that could be quantatively compared (questions 4, 8, 11, and 13), see table 2. we noted that scores for question 4 and question 13 varied the most (between 1 and 9), but also the answers to question 8 and question 11 varied, between 4 and 8, and between 1 and 7, respectively. the difference in scores for the third paper (paper c in table 2) might be related to the coding experience of the participants. from the text answers we could extract that one of the teams was not able to get the code to run at all and spend a lot of time on setting up the required software, while another team could solve r dependency issues more readily and work through the project. https://doi.org/10.29173/iq977 7/10 hettne, kristina maria; proppert, ricarda; nab, linda; rojas-saunero, l. paloma; gawehns, daniela (2020) reprohacknl 2019: how libraries can promote research reproducibility through community engagement, iassist quarterly 44(1-2), pp. 1-10. doi: https://doi.org/10.29173/iq977 question answers from feedback forms: paper a paper b paper c (4) on a scale from 1 to 10, how much of the paper did you manage to reproduce? 5/9/8 1/1 7/1 (8) how well was the material documented? 8/6/7 4/5 7/6 (11) after attempting to reproduce, how familiar do you feel with code and method used in the paper? 7/7/6 1/3 7/1 (13) rate the project on the reusability of the material 9/8/8 4/5 7/1 table 2. feedback form scores per team for papers that were reproduced by more than one team. read as score from team 1/ team 2/ team 3. language used in feedback forms we noted that the language used in feedback forms was friendly and constructive, reflecting the interpersonal atmosphere at the event. to illustrate, figure 2 shows a word analysis of question 15, containing more positive (blue words) than negative words (green words). figure 2. word analysis of question 15, larger letters reflect higher frequency of the used positive (blue) and negative (green) words. https://doi.org/10.29173/iq977 8/10 hettne, kristina maria; proppert, ricarda; nab, linda; rojas-saunero, l. paloma; gawehns, daniela (2020) reprohacknl 2019: how libraries can promote research reproducibility through community engagement, iassist quarterly 44(1-2), pp. 1-10. doi: https://doi.org/10.29173/iq977 discussion at reprohack, participants could learn about both good and bad practices for writing a reproducible paper and gain experience in using tools like ’github8’ for code development and ’zenodo9’ for depositing code. making data and code available, using interoperable and open software, and documenting code well, are often advocated as the main drivers (baker, 2016) for reproducibility. participants could experience the nuances of these drivers, in that data availability is also about packaging the data so that it is easy and fast to download, and that code availability is also about using platform independent, freely available coding software, and attaching a codebook or contextual explanations about the data collection process. we believe that by participating in a reprohack, researchers become more aware of potential barriers to reproducibility. in addition, we believe that the hands-on, community-driven format of reprohack also helps to normalize making mistakes, and therefore normalizing peer-checking and collaboration. lessons learned can be communicated to others but also help researchers in writing more reproducible papers themselves. finally, participants are encouraged to reuse the reprohack format and re-run the event to their home institutions. potential reprohack organizers might want to discuss with each other which day of the week they want to pick for the event. in the netherlands, a saturday attracted young scholars from across the country. choosing a saturday instead of a weekday might also be easier to organize as meeting rooms are free (if buildings are open) and participants do not have any work duties on that day. the last point is an issue that we are very much aware of. to make reproducibility and open science part of mainstream research, events like these need to be treated like any other training or workshop by supervisors. we suggest listening closely to what the local open science community prefers and to cater to their needs. if library training programs are a fixed part of graduate school curricula, a reprohack could become part of top-down training and hence be held during working hours. there is also a danger of only getting early career researchers already actively involved in open science communities. this challenge can be addressed by organizing smaller and regular reprohacks, for example within a research group, in addition to or replacing regular presentations of intermediate or final results. this will normalize reproducibility practices and create a bigger audience for larger and interdisciplinary events. finally, we recommend picking a location that offers wifi, power outlets, and space for participants to break out into smaller working groups or have lunch, and found the library especially suitable while not burdening the limited funding available to the organisers. there are more community initiatives to promote reproducible research, such as the collaborative book project the turing way (the turing way community, et al., 2019), the ‘cure-tier workshop10’ aimed at librarians, archivists and information professionals, and the ‘codecheck11’ workflow and guidelines for checking code behind research papers. reprohack distinguishes itself from these activities by providing a live setting in which researchers with different levels of experience with computationally heavy research and reproducible analysis can learn together and from each other. the feedback forms provided at reprohack guide participants on what to expect from papers in terms of reproducibility and reusability; participants can incorporate these ideas in their own work. in addition, feedback forms are a possitive reinforcement to authors who shared their work. a topic for future research is to identify factors contributing to how reproducibility is assessed and how much the reviewer's own coding experience impact the reproducibility score. https://doi.org/10.29173/iq977 https://github.com/ https://zenodo.org/ https://www.projecttier.org/fellowships-and-workshops/cure-tier-workshop/#about-the-cure-tier-workshop https://codecheck.org.uk/ 9/10 hettne, kristina maria; proppert, ricarda; nab, linda; rojas-saunero, l. paloma; gawehns, daniela (2020) reprohacknl 2019: how libraries can promote research reproducibility through community engagement, iassist quarterly 44(1-2), pp. 1-10. doi: https://doi.org/10.29173/iq977 conclusion based on the feedback forms from participants, a list of top 10 tips to enhance the reproducibility and reusability of a paper can be created (box 2). box 2. top 10 tips from reprohack participants to enhance the reproducibility and reusability of a paper. libraries have always been a space for people to meet and exchange ideas. in a sense, they offer a ‘neutral’ space outside research institutes for researchers to focus on sub-parts of work. reprohacks and other grassroots initiatives need exactly that: a place to meet, work, think and discuss. libraries are connected within the faculties and they can use their network to reach researchers throughout the university. for this first reprohack in the netherlands, the cds contributed greatly by offering its infrastructure and enhancing the organisers’ outreach through posters, flyers and its twitter account, but also informing their faculty liaisons to give them an opportunity to spread the word via their channels as well. next time, more directed advertisements can be made by informing participants taking part in other workshops organized by the cds. reprohack sparked discussions at other dutch universities around organizing their own reprohack. this process is facilitated by information on how to run a reprohack provided in a ‘github repository5’ and there are plans to organize webinars for those interested in running their own reprohack. references baker, m., 2016. 1,500 scientists lift the lid on reproducibility. nature, 533(7604), p. 452–454. • package data so that it is easy and fast to download • provide non-platform specific code that is written using an open software • use a codebook explaining the data structure • include a readme text file to explain the context of data collection • comment the code generously • perform a typo check • report on time needed to run the code • explain which parts of the code corresponds to which results in the paper • attach a permissive license to code • attach a permissive license to data https://doi.org/10.29173/iq977 https://github.com/reprohack/reprohack-hq 10/10 hettne, kristina maria; proppert, ricarda; nab, linda; rojas-saunero, l. paloma; gawehns, daniela (2020) reprohacknl 2019: how libraries can promote research reproducibility through community engagement, iassist quarterly 44(1-2), pp. 1-10. doi: https://doi.org/10.29173/iq977 jacobsen, a. et al., 2019. fair principles: interpretations and implementation considerations. data intelligence, 2(1-2). pashler, h. & wagenmakers, e., 2012. editors’ introduction to the special section on replicability in psychological science: a crisis of confidence?. perspectives on psychological science, 7(6), p. 528– 530. the turing way community, et al., 2019. the turing way: a handbook for reproducible data science. https://doi.org/10.5281/zenodo.3233853 wilkinson, m. d. et al., 2016. the fair guiding principles for scientific data management and stewardship. scientific data, 3(1), p. 160018. 1 contact author: kristina m. hettne; leiden university libraries, centre for digital scholarschip; witte singel 27, 2311 bg, leiden, the netherlands; k.m.hettne@library.leidenuniv.nl. affiliation and address for authors: ricarda proppert; leiden university, faculty of social and behavioural sciences, wassenaarseweg 52, 2333 ak leiden, the netherlands; linda nab; leiden university medical center, department of clinical epidemiology, albinusdreef 2, 2333 za leiden, the netherlands; l. paloma rojas-saunero; erasmus mc, epidemiology department, doctor molewaterplein 40, 3015 gd, rotterdam, the netherlands; daniela gawehns; leiden university, leiden institute of advanced computer science, niels bohrweg 1, 2333 ca leiden, the netherlands. 2 https://www.library.universiteitleiden.nl/research-and-publishing/centre-for-digitalscholarship/workshops 3 https://digitalscholarshipleiden.nl/articles/durable-access-to-book-historical-data 4 https://github.com/dorienhuijser/decisiontreemridata 5 https://github.com/reprohack/reprohack-hq 6 https://reprohacknl.github.io/reprohack/ 7 https://www.software.ac.uk/blog/2020-01-15-reproducibility-hackathon-netherlands-aftermath 8 https://github.com 9 https://zenodo.org 10 https://www.projecttier.org/fellowships-and-workshops/cure-tier-workshop/#about-the-curetier-workshop 11 https://codecheck.org.uk https://doi.org/10.29173/iq977 https://doi.org/10.5281/zenodo.3233853 mailto:k.m.hettne@library.leidenuniv.nl https://www.library.universiteitleiden.nl/research-and-publishing/centre-for-digital-scholarship/workshops https://www.library.universiteitleiden.nl/research-and-publishing/centre-for-digital-scholarship/workshops https://digitalscholarshipleiden.nl/articles/durable-access-to-book-historical-data https://github.com/dorienhuijser/decisiontreemridata https://github.com/reprohack/reprohack-hq https://reprohacknl.github.io/reprohack/ https://www.software.ac.uk/blog/2020-01-15-reproducibility-hackathon-netherlands-aftermath https://github.com/ https://zenodo.org/ https://www.projecttier.org/fellowships-and-workshops/cure-tier-workshop/#about-the-cure-tier-workshop https://www.projecttier.org/fellowships-and-workshops/cure-tier-workshop/#about-the-cure-tier-workshop https://codecheck.org.uk/ vol282-3.indd 12 iassist quarterly summer/fall 2004 introduction the university of winnipeg information literacy program operates on the premise that students develop information literacy skills and knowledge best when opportunities for learning are integrated into the subject curriculum. this paper discusses the results of attempting to integrate data literacy into the subject curriculum in the same way. while discovering what are the best practices for developing data literacy, the paper considers what aspects can be applied from the information literacy fi eld. and, what is unique to learning how to discover, manipulate and interpret numeric data?2 in the last two days i’ve spent several hours helping students doing a geography assignment that i helped create. as a librarian in a small undergraduate library i also take my turn covering our chat virtual reference service. the chat conversation this evening starts with “i’m hoping this is karen working” and ends with “what time are you here till tonight”. as the assignment deadline approaches another student emails me her spreadsheet to help diagnose a problem. and, yes i can meet with you tomorrow afternoon. the challenges of integrating data literacy into the curriculum in an undergraduate institution by karen hunt 1 figure 1 transcript of chat virtual reference session with student working on a geography assignmen iassist quarterly summer/fall 2004 13 university of winnipeg information literacy meets data literacy one of the advantages of being a librarian in a smaller undergraduate library (university of winnipeg has about 8000 students and seven librarians) is that by necessity i wear multiple hats. my main position is coordinator of the information literacy program, but i am also the designated “data librarian.” with an undergraduate background in geography and quantitative methods it is relatively easy to work with a faculty member in the geography department to include fi nding and manipulating numeric data in a fi rst-year course. integrated information literacy the fi eld of information literacy is becoming well established and there is general agreement that students learn best when information literacy is integrated throughout the curriculum.3 when discussing information literacy with faculty and administration it is helpful to refer to established defi nitions, reports, research and best practices. these include the information literacy competency standards for higher education (endorsed by the association of college and research libraries (acrl) in 2000) and characteristics of programs of information literacy that illustrate best practices: a guideline also published by acrl.4 there is some tension in the fi eld between stand-alone courses in information literacy and a more integrated approach. my colleagues in australia also talk of “embedded” information literacy, which they argue is a better approach than “integrated” information literacy.5 there is also a tension between the service role of a librarian and the teaching role of a librarian-teacher. there is also a tension between recognizing that librarians do not “own” information literacy and the need to fi nd a “home” for information literacy. all this is interesting to me because i see similar discussions occurring in the fi eld of “data literacy”. for example, according to the acrl best practices document, the goals and objectives of information literacy programs should: • articulate the integration of information literacy across the curriculum; • accommodate student growth in skills and understanding throughout the college years; • apply to all learners, regardless of delivery system or location; • refl ect the desired outcomes of preparing students for their academic pursuits and for effective lifelong learning. substitute “data literacy” for “information literacy” and are we on the way to defi ning the best practices for data literacy programs? it is also interesting to note that the writing across the curriculum (wac) movement discusses similar issues.6 human geography case study when the instructor for the human geography course and i fi rst discussed collaboration, he was interested in a term assignment that would get students “into the data,” that was scalable to a large class, and was easy to mark. briefl y, the assignment involves the students using the united nations human development indicators, constructing their own hdi for provinces and a territory in canada, and constructing their own population pyramids for two small towns in canada.7 the assignment includes fi nding and retrieving data, manipulating data using msexcel and understanding how indexes are created. students must also demonstrate an understanding of the relationship between quality of life, population age distributions and migration, and how these factors are distributed geographically. librarian’s involvementlibrarian’s involvement • developed the assignment, initially for a small class (fall 2003) • held labs for the class • available for tutoring and questions instructor’s involvementinstructor’s involvement • created the shape of the assignment • fielded questions from students • lectures on development and population distribution • ultimately responsible table 1: roles in the human geography term assignmenttable 1 14 iassist quarterly summer/fall 2004 there are several things i have learned from our collaboration, which roughly fall into two categories: how this collaboration is the same as other more traditional information literacy assignments and how it is very different! data literacy same as information literacy finding the best methods for ensuring students learn is the same whether we’re discussing information literacy, writing across the curriculum or data literacy (at the 20000 foot level at least). students learn best when the curriculum is relevant and builds on previously learned skills and knowledge, involves opportunities for making connections and practicing, and includes appropriate scaffolds. proponents of data literacy should borrow heavily from information literacy and learning theory. data literacy different than information literacy in the practical implementation (or on the ground) data literacy is quite different than traditional information literacy. toolbox for numbers is more complicated in a typical course-integrated approach to information literacy i can assume the students know how to read, use a web browser and word processor. at this time i can’t assume that students know how to use a spreadsheet or even remember that one can turn a positive number into a negative number by multiplying by negative one. this assignment didn’t even require the use of any statistical package. developing and supporting assignments takes more time and expertise if i develop an assignment that requires english students to compare the process of fi nding a journal article to fi nding a book, i can assume my colleagues can also help students with the assignment. until the training of reference librarians includes more “data literacy,” supporting data assignments will be onerous for those that can provide help. staff training will be critical to make a “data literacy across the curriculum” program scalable. it also seems to be more complicated to develop data assignments, however that may refl ect my own lack of practice in this area. as chat transcript (figure 1) demonstrates, support for this assignment relied on one person. recommendations terminology and defi nition the information literacy fi eld may not agree on the precise defi nition of information literacy, but most people use the term “information literacy” rather than “library instruction” or “information fl uency.” in what i have been calling “data literacy” we have statistical literacy, quantitative reasoning or quantitative literacy, numeracy and data literacy all roughly meaning the same thing. 8 until we start talking the same language amongst ourselves it will be diffi cult to convince our stakeholders that we should be revising curriculum and programs. articulated learning outcomes and best practices in information literacy we have the acrl competency standards (and similar documents in australia and elsewhere). the data literacy fi eld is more fragmented. for example, those discussing “quantitative literacy” and “statistical literacy” often leave out the fi nding and evaluating of existing statistics and data. in my own view, what needs to be done is: • get the players talking to each other; • decide on one term and agree upon a defi nition of "it"; • codify data literacy learning outcomes; • endorse and promote the standards; • develop opportunities for data librarians to learn how to integrate the outcomes; • articulate best practices for data literacy programs. training the success of data literacy will depend on how well we train data librarians about teaching; reference librarians and support staff about data and referrals; and faculty and administrators about why data literacy is imperative. iassist quarterly summer/fall 2004 15 conclusion: finding a home for data literacy the success of the information literacy programs is due, in part, to the library organizations and librarians who have taken a leadership role. the success of data literacy may also depend on organizations such as iassist providing leadership and fi nding or creating a home for data literacy, while always considering the real success will be measured in the strength of the relationships among all partners and the depth of our students’ understanding. notes 1 contact: karen hunt, information literacy coordinator, university of winnipeg library, 4c01a, 515 portage avenue, winnipeg, manitoba, canada r3b 2e9. phone: +1 204 786-9940. internet: http://scholar.uwinnipeg.ca/khunt/ email: k.hunt@uwinnipeg.ca 2 this paper is a rough translation of one delivered at the 2004 iassist conference in madison, wisconsin. the online version, complete with references and links, can be found at http://scholar.uwinnipeg.ca/khunt/iassist2004/ 3 see for example, rockman, ilene f. (2004). integrating information literacy into the higher education curriculum : practical models for transformation.1st ed. san francisco: jossey-bass. 4 american library association (2004) information literacy competency standards for higher education. available at: http://www.ala.org/ala/acrl/acrlstandards/informationliteracycompetency.htm; american library association (2004), characteristics of programs of information literacy that illustrate best practices: a guideline. available at: http://www. ala.org/ala/acrl/acrlstandards/characteristics.htm 5 bundy, a (2004) (ed.). australian and new zealand information litercaylitercayliteracy framework: principles, standards and literacy framework: principles, standards and practice (2nd ed.). adelaide: australian and new zealand institute for information literacy. available at: http://www.anziil. org/resources/info%20lit%202nd%20edition.pdf 6 see http://mendota.english.wisc.edu/~wac/ 7 united nations human development indicators. see http://hdr.undp.org/statistics/data/ 8 see for example schield, milo (2004). statistical literacy and liberal education at augsburg college. peer review summer issue, assoc. of american colleges and universities. available at: http://www.augsburg.edu/ppages/~schield/. 16 iassist quarterly summer/fall 2004 vol30-4.indd 10 iassist quarterly winter 2006 technology of data: collection, communication, access and preservation the 34th international association for social science information services and technology (iassist) annual conference will be held at the stanford university, palo alto, california, usa, may 27-30, 2008. this year's conference, technology of data: collection, communication, access and preservation, examines the role of technology and tools in various aspects of the data life cycle. the theme of this conference addresses how technology can affect aspects of data stewardship throughout the data lifecycle. the methods and media by which data are collected, shared, analyzed and saved are ever-changing, from punch cards and legal pads to online-surveys and tag clouds. there has been an explosion of data sources and topics; vast changes in compilation and dissemination methods; increasing awareness about access and associated licensing and privacy issues; and growing concern about the safeguarding and protection of valuable data resources for future use. the 2008 conference is an opportunity to discuss the role of technology – past, present, and future – in all of these arenas. we seek submissions of papers, poster/demonstration sessions, and panel sessions on the following topics: � issues and techniques for preserving “old” data as well as information “born digital” � methods, technology and questions surrounding data dissemination, including best practices and innovations � archival and preservation challenges presented by new processes � metadata � innovation in the use of data for teaching and research � the legal issues surrounding new technologies � changes in resource discovery methods � data services in virtual spaces � providing services to users with different degrees of technical “savvy” � tools and spaces for research collaboration papers on other topics related to the conference theme will also be considered. the deadline for paper, session, and poster/demonstration proposals is december 17, 2007. the conference program committee will send notification of the acceptance of proposals by february 8, 2008. individual presentation proposals and session proposals are welcome. proposals for complete sessions, typically a panel of three to four presentations within a 90-minute session, should provide information on the focus of the session, the organizer or moderator, and possible participants. the session organizer will be responsible for securing session participants. organizers as well as panel participants are also welcome to submit additional paper proposals but please note that conference program committee may need to limit the number of presentations per person. proposals for papers, sessions, and poster/demonstrations should include the proposed title and an abstract no longer than 200 words. longer abstracts will be returned to be shortened before being considered. please note that all presenters are required to register and pay the registration fee for the conference. registration for individual days will be available. proposals can be submitted via email to: iassist08@gmail.com a conference website with on-line submission form will be available shortly. a separate call for workshops is also forthcoming. 1/16 kirby, jasmine s. (2018) how not to create a digital media scholarship platform: the history of the sophie 2.0 project, iassist quarterly 42 (4), pp. 1-16. doi: https://doi.org/10.29173/iq926 how not to create a digital media scholarship platform: the history of the sophie 2.0 project jasmine s. kirby1 abstract since the mid-2000s digital platforms have emerged to take advantage of the capabilities of new technology to incorporate media content, tell nonlinear stories, and reinvent the book for the 21st century. sophie 1.0, from the university of southern california, the institute for the future of the book (ifb), and computer scientists based in europe, was an attempt to create a multimedia editing, reading, and publishing platform. sophie 2.0 was an international collaboration between the university of southern california and astea solutions in bulgaria to rewrite sophie 1.0 in the java programming language. this research will explore how the sophie 2.0 project was unable to become a viable and well-maintained open source product despite receiving over a million dollars in funding from the mellon foundation. problems included the technological difficulty of creating an easy-to-use but completely customizable open source multimedia e-publishing platform, which was also compounded by competing visions over what this project was to be. stakeholders did not demand a deliverable that actually worked. funders seemed willing to overlook weaknesses in early releases for a more encompassing, if impractical, project. the computer scientists wanted to add the most features possible, while the ifb and usc institute for multimedia literacy focused on creating a product based on the values of a future they hoped to create. understanding what went wrong with sophie 2.0 can help us understand how to create better digital media scholarship tools and to start much-needed discussions about failure in the digital humanities. keywords sophie, digital humanities, failure, usc, institute for the future of the book, astea solutions, introduction we don’t talk enough about failure in the digital humanities. this is a problem because as information professionals we are expected to find, understand, and explain digital tools for scholarship that we want our users to be able to access for eternity. in the field of librarianship and in higher education, there is a particular emphasis on supporting free and open source platforms, since these are thought to embody the values of openness, transparency, and continued access that we promote and aim to achieve. however, it can be very hard to tell which open source projects are going to succeed and which will flop. studying failed projects can give us guidance on what to look for when choosing digital humanities tools for our own research and the research of our library users. while some of these failures are products of their unique time and place, they also speak to many dangers in software development that apply to other projects. although it received over a million dollars from the mellon foundation and others, sophie 2.0, an update on the institute for the future of the book’s sophie 1.0 multimedia e-book reading and authoring platform from the university of southern california and https://doi.org/10.29173/iq926 2/16 kirby, jasmine s. (2018) how not to create a digital media scholarship platform: the history of the sophie 2.0 project, iassist quarterly 42 (4), pp. 1-16. doi: https://doi.org/10.29173/iq926 astea solutions never became a viable digital humanities and media scholarship platform. this paper will explore factors that contributed to the failure of the sophie 2.0 project. situational context there is limited scholarship available about failed digital humanities projects and the failure of the sophie project in particular. this research will draw from “whatever happened to project bamboo?” an article written in 2014, where quinn dombroski, a member of the staff of an institution participating in the digital infrastructure initiative project bamboo, discusses the issues that resulted in the failure of that project. there is not much critical information about the sophie project in particular. in 2010, dan visel, from the institute for the future of the book (ifb), who was in charge of software development for sophie, interviewed the institute’s founder, bob stein, discussed stein’s influential role in the history of computers, and briefly mentioned the sophie project. there is a 2011 interview of bob stein that while focused on his accomplishments and those of institute for the future of the book, does include a section on problems with sophie 1.0 and sophie 2.0. both of these published interviews help us understand a lot of the context and thinking behind the sophie project and what it was building on, blaming its failure on too much ambition and a lack of funding, but do not go into enough detail into the reasons that the sophie project failed. beyond these interviews, there is so little objective information available about sophie and its end that there is still confusion as to whether it is a viable platform. for example, as recently as january 2017, there was an article in the magazine computers in libraries about digital humanities still promoting sophie 2.0, although the author, nancy k. herther, an academic librarian, acknowledged the lack of updates, but mentioned looking forward to what the project is said to bring (herther, 2017). i want to emphasize the importance of seeing computer applications in the greater context of history of media and history of science and technology to contrast how narratives about new software tend to focus on how they are unprecedented and unique to the 21st century. i build on ballatore and natale’s work on the cultural implications of what they call ‘the myth of the death of the book’ and agree that “such prophecies, however, are revealing of the way societies regard media as vehicles for change – precisely because they are embedded in the idea of the future.” (ballatore and natale, 2016). i agree with evgeny morozov, in his work to save everything click here that the internet needs to be studied not as a “mcluhanesque ‘medium’” or the bringer of a unique epoch in human history and instead placed in a greater context (morozov, 2013). another article that inspired the way this research is framed is “listening to pictures” by katie day good, a media scholar. this article discusses the history of radio photologues, a combined radio program and photogravure section in the chicago daily news, and explores the context and significance of this unique media product. similarly, i hope to place sophie in a longer history of mixed media forms. i want to illustrate the important lessons that digital humanists can learn from looking past our cultural myths about media and instead examine the actual complicated messy history of how new media emerge and change over time. a brief history of the future of reading the history of sophie starts in 1981, with publisher bob stein encountering “electronic text that might be readable” and an early demonstration of video embedded in hypertext and “…realized at that point https://doi.org/10.29173/iq926 3/16 kirby, jasmine s. (2018) how not to create a digital media scholarship platform: the history of the sophie 2.0 project, iassist quarterly 42 (4), pp. 1-16. doi: https://doi.org/10.29173/iq926 that the book of the future wasn’t going to be limited to text and figures; we were going to be able to have audio and video on the page”(stein, 2008). sophie 2.0 is but a chapter in the long and troubled history of e-books and e-book platforms. while the idea of the e-book can be found earlier in computer history, in their article about the history of the idea of e-books resulting in the end of the print book, ballatore and natale describe how “the development of the actual idea of the e-book is principally attributed to andries van dam, who coined the term working on a hypertext system in 1967, and michael hart, who founded project gutenberg in 1971” and describe how early attempts at e-books were hindered by the size and limits of computers of the 60s and 70s (ballatore and natale, 2016). in his article “a call to embrace social reading in higher education” business professor matthew dean describes how “in a hard-copy version of a book, one can highlight sentences, annotate in the margins, bookmark important pages, keep the book in a revered spot on a bookshelf, loan it to friends, discuss your favorite parts with others who have read the book, etc.” (dean, 2016). e-book creators face the challenge of either replicating the beloved features of a book in a digital environment or creating something better. bob stein figures prominently in both the history of e-books and the story of sophie. not many other people had the vision to see past the limits of the technology at the time and believe in the possibilities that new technology offered for changing the way we learn. before even the original sophie, bob stein would try to build a multimedia editing platform of one form or another at least twice. while running the voyager multimedia electronic publishing company, bob stein oversaw the creation of the expanded books toolkit, a software that allowed educators to create and edit their own editions of books on floppy disks (rüger et al., 2008). stein is described in a 1996 profile of his voyager company as “the most far-out digital publishing visionary in the new world or the least effective businessman alive – or both” and that “many people who work for stein mention his tremendous intellectual passion and enthusiasm and an almost equal number cite his short attention span and complete disregard for detail” (virshup, 1996). that being said, his laserdisc film collection offering a second audio track of commentary was the first of its kind and changed the way we study film, and voyager offered a variety of multimedia educational products including, most famously, the interactive primary source collection who built america cd-rom (visel, 2018). stein of course was not the only one interested creating a platform allowing users to create multimedia products in the 1980s and 1990s. in 1987 apple created the hypercard application that allowed users to create documents that “could contain images, sounds, and movies; the author could add controls via a simple language called hyperscript” however, “apple killed the project in 2000…” (stein and visel, 2010). the story of hypercard shows that similar programs to allow non-programmers to create interactive multimedia documents had been tried before in the closed source world with limited success. moreover, ballatore and natale in their discussion of the history of e-books indicate how early attempts at e-readers during the 1990s were considered a failed technology and did not survive the dot-com burst (ballatore and natale, 2016). while of course the technology was much more limited, it does show that multimedia e-books and e-book creators were never popular. stein also led the night kitchen company that created tk3, the predecessor software to sophie that was a closed source product (rüger et al., 2008). tk3 was probably the most successful of all of stein’s efforts to create a multimedia editor and reading platform. for example, according to visel, while no one ever put their thesis in sophie 1.0 or sophie 2.0., virginia kuhn, a media scholar, did her thesis in https://doi.org/10.29173/iq926 4/16 kirby, jasmine s. (2018) how not to create a digital media scholarship platform: the history of the sophie 2.0 project, iassist quarterly 42 (4), pp. 1-16. doi: https://doi.org/10.29173/iq926 tk3 and got it accepted which was a major development for e-books since it was the first born digital doctorate (visel, 2018). it is clear that there was some acceptance of the use of tk3 for scholarship purposes. in an interview he did with visel, stein describes the end of tk3 with his refusal to get rid of the macintosh version in order to get funding from microsoft to market the software and then ultimately abandoning the project (stein and visel, 2010). the end of tk3 speaks to stein’s dedication to creating software available for users on all types of platforms and wanting to reach the broadest audience possible. telling off the largest software company in the world at the time also demonstrates that stein had a specific vision of how software should work and be available to people. this also is part of a larger pattern of stein not letting things like profit or feasibility get in the way of his vision. despite the previous limited adoption of multimedia e-book creators, stein was given the opportunity to create an open source version of tk3. specifically, “the mellon foundation approached some of the tk3 team and asked them to build a new multimedia authoring program which would extend tk3 by enabling time-based events and make it able to live on the network. that became sophie” (rüger et al., 2008). in their conference paper “sophie: the future of reading” the creators of what would become sophie 1.0 describe their project as “with sophie we are tackling the long standing issues as keeping documents and their media accessible for a long time (the 200 year problem) and making electronic books living documents that capture and reflect the readers’ interactions and comments (the annotating problem)” (rüger et al., 2008). using funding acquired from the mellon and the macarthur foundations, stein founded the institute for the future of the book in 2004, a think tank that continued his longstanding affiliation with usc but was based in new york (stein and visel, 2010). why did nonprofits, most notably the mellon foundation, ignore the unpopularity and unprofitability of earlier efforts to create multimedia e-book platforms and support this venture? issues inherited from the original sophie the sophie project was funded by the andrew mellon foundation as part of their research in information technology program (mellon/rit). this program sponsored projects that would create “community-source software” and “service-oriented architecture” where universities would work together to develop shared platforms that would reduce the amount of redundant software applications and be open source so people wouldn’t be stuck with certain vendors (fuchs, 2008). in other words, creating big applications that serve lots of functions and can be customized to the unique needs of different institutions to replace lots of small applications (fuchs, 2008). in her article on project bamboo, another failed initiative sponsored by this grant, dombroski argues that the main purpose of this program was to make digital tools that everyone can use instead of everyone building their own tools for their own projects (dombrowski, 2014). this goal of creating large platforms to combine overlapping software needs explains why a grant program would be interested in tk3 with its ambitious goal of allowing for multimedia e-books and previous use for digital scholarship. in a 2007 conference paper from the forum for the future of higher education, visel describes the advantages of sophie with: “more and more students are taught to make presentations with powerpoint, a limiting program that, as edward tufte has pointed out, encourages gimmicky special effects at the expense of coherent thinking…sophie treats all media equally: if adding a slide show would be helpful to the primarily written report, the student can add the slide show to the page it is intended to illustrate without having to switch from word to powerpoint” (stein and visel, 2007). this https://doi.org/10.29173/iq926 5/16 kirby, jasmine s. (2018) how not to create a digital media scholarship platform: the history of the sophie 2.0 project, iassist quarterly 42 (4), pp. 1-16. doi: https://doi.org/10.29173/iq926 is the idea of creating tools around the work that universities are doing rather than being locked into the functionality provided by proprietary software such as word and powerpoint that are more oriented towards the needs of corporations than universities. stein and visel also write “sophie’s aim is to democratize the world of multimedia by making it possible for individuals and small nonprofits to express themselves via compelling multimedia books” (stein and visel, 2007). moreover, an open source version of tk3 would also help mellon’s overall mission of keeping the humanities relevant in the 21st century by letting scholars incorporate multimedia. in interviews about what happened, over-ambition and lack of funding are blamed for the failure of sophie and other digital humanities initiatives. this issue of creating and releasing ambitious and difficult to complete software was compounded by the restructuring of the mellon foundation in december 2009, when rit became part of scholarly communications, resulting in a new group of people to work with, different goals, and less money available (dombrowski, 2014). all of this was made even worse by the 2008 financial crisis and a shifting of priorities at universities away from futuristic long-term projects to software that addressed local needs (dombrowski, 2014). we now know that the most successful open source projects have paid staff working to keep them up to date, which was not the case with sophie which had money to create essentially a demo product but could not get funding to keep it going (visel, 2018). it's also true that there was a dramatic staffing change between the first and second editions of sophie. it’s likely that sophie was facing similar issues to the case of project bamboo where “these staffing changes led to a loss of organization memory, which had particularly negative consequences for the message and tone of the project’s communication with scholarly communities” (dombrowski, 2014). it is still possible to access forum postings from 2008 on slashdot where an alleged former programmer from sophie 1.0 argues that their work has been stolen by usc and astea solutions (“how to kill an open source project with new funding,” 2008). in addition, coordinating a project between usc in los angeles, the institute for the future of the book in new york, and astea solutions in bulgaria with the technology at the time was not the easiest feat to accomplish. a sophie user described that the geographic distance meant that it could be very hard to get tech support; and that they imagined getting support would be impossible if you were not friends with one of the people behind the project who could contact the programmers in bulgaria directly.2 on the website for sophie saved in the internet archive, even the most updated version from june 19th 2015, the link for technical support is a mail:to to daniel visel (“sophie,” 2015). sophie 2.0, the update of sophie 1.0 in java, would’ve been difficult regardless but on top of everything else sophie 1.0 was a very unique software with many incompatible parts and ideas. to start, due to the influence of computing pioneer allen kay and a desire to deal with the problem of digital preservation, the original sophie was written in smalltalk (visel, 2018). however, smalltalk was a very academic programming language that had mostly been used by the swiss banking system, which meant that there was no video player or text editing (visel, 2018). smalltalk had the advantages of being device independent, for example, you can still run smalltalk programs from the 1970s (visel, 2018). people want to be able to access their academic work in the future. for example, although it was accepted and serves as part of the basis of her academic career, to look at virginia kuhn’s thesis in tk3 these days you’d need to use a mac emulator running on a mac (visel, 2018). this is unacceptable for widespread use in research and writing contexts where people need to be able to https://doi.org/10.29173/iq926 6/16 kirby, jasmine s. (2018) how not to create a digital media scholarship platform: the history of the sophie 2.0 project, iassist quarterly 42 (4), pp. 1-16. doi: https://doi.org/10.29173/iq926 access stable versions of their scholarship for years as they build up their careers. creating a multimedia editing platform from a programming language mostly used for banking is a very difficult task by itself. smalltalk was so obscure that everything had to be created from scratch and the programmers couldn’t build on other people’s code (visel, 2018). sophie 1.0 was supposed to be built to last. what’s more, this meant that not only were creators of sophie 2.0 unable to use any of the original code but they also could not use any of the workflows that came out of the sophie 1.0 project since it was a very different kind of project. visel summarized the issues that came from using smalltalk for sophie 1.0 with, “it was a really good idea but a really hard idea” (visel, 2018). looking at sophie and its various iterations it becomes increasingly clear that the focus was on building the tool with the most features and not something that actually could be sustained or even worked. the computer scientists who programmed sophie 1.0 described how “[e]ven though a system may work correctly, it may still fail in the field of user experience and usability if it does not embrace suitable concepts to implement and to offer the possible very large number of expected features”(holz et al., 2009). in a 2011 interview, in a discussion on “sophie and software development” stein mentions that “as a publisher, i learned to live with the “get-it-right-the-first-time” reality of print, but it’s a completely wrong model for software development in the era of the digital network, where the goal is to get out a good-enough first version and then iterate and improve as fast and as often as you can” (gold, 2011). in that way, sophie faced many of the same issues as project bamboo. “however, the infrastructure was architected in such a way that made it difficult to complete and release standalone components that could be tested and used while other parts were incomplete…the extensive development time required for infrastructure components without successfully fulfilled real needs…” (dombrowski, 2014). it is clear that the creators of sophie wanted to build a complete product, which also explains why this project took such a long time to create anything and needed so much money. the innovation and uniqueness aspect of sophie 1.0 also went into the user interface and experience. in a review of the sophie 1.0 alpha release, early adopter, tech blogger james bridle, writes “it’s clearly inspired by existing rich media applications such as flash, but it’s [sic.] target users – the technologically unskilled – don’t use such applications. how are they supposed to get their heads around concepts such as ‘flows,’ ‘timelines’ and different server versions? and if they do get that, why aren’t they using the existing apps? it’s all very disappointing, and i think if:book3 [sic.] know it, which is why they haven’t supported or trumpeted this release in any way” (bridle, 2007). even as early as the alpha release, before the global recession limited the possible support from mellon and academic institutions, there were issues with the design of the software itself. having to build everything from the ground up also likely resulted in a user interface that one sophie 1.0 user described as “outdated” and “like running an emulator.”4 visel, who wrote the documentation, described the software as “deeply, deeply confusing.” on the sophie 2.0 developer site, there is still a page that reads “sophie wishlist (this is a list based on discussion with bob stein and the institute for the future of the book’s sophie users as well as looking through mantis feature requests).” items on the wish list, which were things that the developers openly admitted probably wouldn’t be fixed, included not having a usable windows file format, no way to delete embedded books, and “sophie 1.0 doesn’t handle saves well when the book has been moved to another location – this is something that people tend to do a lot (esp. on macs, i think) and there have been a lot of crashed books because of this. this needs to be handled more gracefully” (visel, 2008). instead of focusing on a less ambitious but functional project, https://doi.org/10.29173/iq926 7/16 kirby, jasmine s. (2018) how not to create a digital media scholarship platform: the history of the sophie 2.0 project, iassist quarterly 42 (4), pp. 1-16. doi: https://doi.org/10.29173/iq926 sophie 2.0 traded the lessons and benefits of smalltalk in order to build a product that would be able to be maintained by a dedicated open source community. and then, it got worse in a forum post on slashdot, a user claiming to be elizabeth daley of the university of southern california and principal investigator of the sophie 2.0 project, explained that part of the reason that sophie never took off was that institutions could not properly support a software written in smalltalk, and that changing the language to something more commonly used would help reach the goal of creating a community to help develop it, as was required to get another grant from mellon (“how to kill an open source project with new funding”, 2008). changing the language it was coded in to java did not change the fact that this was an open source software that did not particularly benefit open source users. sophie was a multimedia editing platform meant for people who could not code, such as publishers and academics. this discussion on slashdot was one of the few examples of any outreach towards programmers and information technology specialists for building the open source community to maintain the software. it is unclear who was ultimately supposed to keep the software up to date. the targeted audience of people who cannot code are the same people who are unable to fix or perhaps even articulate issues that inevitably come up in such a complicated piece of software in an environment of constant and rapid technological change. visel mentioned at the time that the mellon foundation seemed to believe that if a digital technology was released as an open source software a community would build up around it to update it and add new features (visel, 2018). while there have been some historical examples of this, such as the linux operating system, these tend to be technical software aimed at technical users who have some coding experience already. moreover, looking at the sophie 2.0 developer site, it would be very hard to actually get involved in this project if a person wasn’t in bulgaria and working for or in some way affiliated with astea solutions since it appears that potential contributors had to get approval from the main team to do anything (“lpandeff”, 2009). even if contributors were interested in joining this project, the developer’s site is confusing, and it would be difficult coordinating time zones for contributors who were not based in europe, such as it people affiliated with usc. what’s more, sophie 2.0 maintained the sophie 1.0 team’s practice of releasing unfinished alpha software to unsuspecting users. a sophie 2.0 user described how “it was super unstable, crashed, and no clear rhyme or reason to why it kept crashing… in the process of building the sophie book things would be gone and not retrievable not ever getting to the point of being finished.”5 the sophie 2.0 team never released a reliable deliverable; it was always up to the user to imagine what could have been while dealing with the reality of the limitations of the software in front of them. instead of creating a smaller project that could do one thing well, they created a project that could kind of do a variety of things but was buggy, prone to breaking, and rarely up to date. this is neither appealing to the intended end users who are people who do not know how to code or the intended open source community who can’t see what the overall vision of the project was supposed to be based on the code that was released. https://doi.org/10.29173/iq926 8/16 kirby, jasmine s. (2018) how not to create a digital media scholarship platform: the history of the sophie 2.0 project, iassist quarterly 42 (4), pp. 1-16. doi: https://doi.org/10.29173/iq926 who was the audience? another aspect that played a major role in its downfall was that sophie lacked a clear target audience. like many projects relying exclusively on grant funding, the target audience for sophie changed based on who they were talking to. for example, a conference paper about sophie 1.0 from the forum on the future of higher education mentions “while sophie can be used in many settings it is aimed squarely at the world of education” (stein and visel, 2007). there is a promotional video put out by usc iml focused on using sophie 2.0 for journalism (artsj09, 2009). the grant proposals and papers list even more possible uses. for example, smartbook, a project proposing to expand the functionality of sophie 2.0, received funding from the bulgarian national science foundation, and argued that an expanded sophie 2.0 could indirectly benefit the way that science was published and discussed (koychev et al., 2013).there was very little that sophie could not be used for in some way or another. one exception of course comes from the transliteracies journal which pointed to the fact that sophie was not being used for narrative storytelling and only for educational purposes as another weakness of the platform (hudson, 2008). being exclusively for education and not entertainment makes it similar to the radio photologues of the daily news, where the ultimate problem of that platform was not that it was worse than other forms of media during the 1920s but that what people used it for didn’t use narrative structures (good, 2017). it is true that in the mid-2000s mainstream publishing companies were experimenting with releasing mixed media books. the idea was to reach audiences who were more used to, as one publisher explained, “three-minute youtube videos and using social networks,” with experiments such as including videos in electronic books that could be read online or on apple devices or a website where readers could discuss the events in a book and possibly have their comments incorporated into later books (rich, 2009). however, it seems sophie did not reach out to these markets in any meaningful way until a report from 2011. specifically, the research into potential users from the bulgarian software developers focused primarily on reaching out to publishing companies. this research was done as a requirement of a mellon grant and they did marketing research on publishing companies with a survey that didn’t have a high response rate (“sophie 2.0: from projects to publishing initiative two, part one marketing analysis results,” 2011). building digital humanities tools that tell narratives is important since that is the dominant form of how we relate to each other and how we currently communicate information. in trying to build a platform that was usable to everyone they created a product that was not particularly useful to anyone. was sophie ahead of its time? sophie users often lament that the software was ahead of its time. but the real problem was that it was built for a future that would never be. looking at conference presentations and interviews given by some of the minds behind sophie 1.0 and sophie 2.0 it becomes very clear that they, like many people during the first decade of the 2000s, subscribed to a narrative of technological progress and the idea of computing technology completely changing the way we live our lives. for example, a paper by the bulgarian computer scientists who took over the sophie 2.0 project begins with “the information technologies available today have made possible the advent of the e-book that overcomes a number of weaknesses of the classic scroll described so well by socrates 24 centuries ago” (koychev et al., 2013). the scientists see themselves as solving a problem that books themselves were not able https://doi.org/10.29173/iq926 9/16 kirby, jasmine s. (2018) how not to create a digital media scholarship platform: the history of the sophie 2.0 project, iassist quarterly 42 (4), pp. 1-16. doi: https://doi.org/10.29173/iq926 to solve and ushering in a new era in the way that they believed print books did. the scientists building on the sophie project see themselves as part of a long tradition of technological disruption and changing the world to a better place. in their article “a pedagogy for original synners” the authors, scholars from the institute for multimedia literacy describe sophie and similar efforts with “the limited range and noncommercial aspirations of such programs place emphasis on developing conceptual sophistication rather than final polish. we believe that this emphasis on process over product may allow students to pursue more experimental, concept-driven creative and critical production” (anderson and balsamo, 2008). this is a way of saying that they were less concerned about providing students with tools that actually worked than having tools that would encourage students to do a certain kind of creative work. the main users of sophie were educators preparing their students for a certain future they imagined. for example, the “original synners” pedagogy hoped to “address the learning needs of the born digital generation. 1. open… 2. hybrid… 3. media rich…” (anderson and balsamo, 2008). the futuristic stance of the software is also demonstrated in the rare test uses for the software, k-12 and undergraduate writing courses. for example, when describing using a sophie 1.0 book for teaching an ap spanish course, private school teacher sol b. gaitán describes how, “as a teacher of children and adolescents, i firmly believe i have the moral obligation to prepare them for the world they will be part of as adults” (gaitán, 2011). this also explains the way that this software dealt with issues of intellectual property, which is to say for the most part, it didn’t .perhaps this is because under the fair use doctrine in the us, copyrighted materials can be incorporated in educational materials under certain conditions. it is also possible that the neglect of intellectual property law could be what barnett describes “[a]s part of the evolutionary history of e-books, the proliferation of pirated texts as digital files in the 1990s and 2000s created a network and market for the creation and consumption of digital texts. these were frequently circulated as microsoft word files, copied and pasted from ocr scans or typed by fans.”(barnett, 2015). but more importantly, visel described in an interview how there was an expectation during the 1990s and 2000s that with the web information would be free and available online (visel, 2018). this perhaps explains why in the paper for sophie 1.0 the creators mention “sophie to date does not deal with implement or enforce any drm6v related technologies, possibly making some local media resource unavailable for use within sophie” (rüger et al., 2008). both sophie 1.0 and sophie 2.0 in their documentation and description deal very little with issues of intellectual property, despite being a software that is concerned with helping people create and preserve multimedia content. and it’s not that people weren’t thinking about intellectual property issues during the time. for example, a professor who used sophie 2.0 for teaching a computers and writing course mentioned how learning about copyright and dealing with the fact that most film is still protected by copyright was a major part of an assignment of creating a digital edition in sophie 2.0 (bjork, 2012). it’s clear that the few actual test case uses of sophie 2.0 were dealing with copyright, even if the software itself did not. as a software that dealt with both writing and multimedia, there was the constant question of whether sophie was an aggregator or an authoring platform or both. “the sophie server will provide a repository for all sophie books that exist on a given network and will allow users to search sophie books already created, as well as publish new books on it. the repository will also serve as a rich source https://doi.org/10.29173/iq926 10/16 kirby, jasmine s. (2018) how not to create a digital media scholarship platform: the history of the sophie 2.0 project, iassist quarterly 42 (4), pp. 1-16. doi: https://doi.org/10.29173/iq926 of reusable content” (stein and visel, 2007). this reflects another idea from the time period that all new works would be combinations of old works and all merging into one big work as information became available and free online (lanier, 2010). sophie is an authoring, reading, and publishing platform for multimedia documents. while the creators were caught up in the view of a new culture, users on the ground were just confused. in bridle’s review of the alpha release of sophie 1.0, there is a comment from someone identifying themselves as a developer from the sophie 1.0 team who writes “james, the reason bob says ‘assemble’ and i say ‘edit’ is a philosophical difference in how we see sophie…however, you look at it, the simple facts are that sophie does (within the limits of bugs and available developer time) provide an editing tool for structured text, with searching, spellchecking, undo/redo, markup, annotation (for both author and reader, independently sharable by groups working on the same book) *as well as* assembly tools for other content formats. comment by tim rowledge – april 12, 2007 @8:44 pm [sic.]” (bridle, 2007). to which james bridle, the author replied, “tim – what i am suggesting is that ‘philosophical difference’ is not helping end users figure out what they are supposed to do with this software” (bridle, 2007). this confusion over sophie’s purpose continued into sophie 2.0 where “usc’s support for sophie has also included funding a week-long workshop for scholars in may 2008, during which sophie’s affordances were tested in practice in tandem with discussions regarding the ways in which sophie transforms the traditional acts of scholarly reading and writing. organized and led by the iml, the workshop raised several key conceptual issues. one of these centered on the tension between understanding sophie as a compositional environment that sparks new forms of writing, in opposition to imagining sophie to be an aggregator, or a space for gathering and displaying various texts and media objects” (visel, 2008). the focus of the 2008 workshop is how sophie 2.0 could be used to change scholarship, not how sophie 2.0 fit with the scholarly practices at the time. it is clear that what the software was actually for was a question that remained unanswered throughout its entire history. a revolution in reading before going into e-books, stein worked as a leftist publisher for many years (stein and visel, 2010). the influence of leftist political thought played a role in his leadership both of his own companies and of the institute for the future of the book. sophie 2.0 continued the leftist visions of bob stein in particular with the way that the authorship feature is set up. specifically, in a comparison of the strengths and weaknesses of various e-book platforms, one complaint about the authorship feature of sophie 2.0 was “it is not possible to assign roles and planning activities when writing a book. all users can write at the same time, any place, and there is no way to block a section to avoid conflicts when writing nor it’s [sic.] possible to know which section of the book was written by a particular user” (ochoa et al., 2013). this is a structure of authorship without hierarchies where all authors are considered equal in both level of importance and in contribution regardless of amount. this complex rethinking is an idea of authorship that was promoted by the institute for the future of the book. specifically, “‘an old-school author,’ says stein, ‘is somebody whose commitment it is to engage with subject matter on behalf of future readers. a new school author is somebody whose commitment is to engage with readers in the context of subject matter… authors are about to learn what musicians have already learned, which is, they’re going to get paid to show up, whether it’s at a speaking gig at https://doi.org/10.29173/iq926 11/16 kirby, jasmine s. (2018) how not to create a digital media scholarship platform: the history of the sophie 2.0 project, iassist quarterly 42 (4), pp. 1-16. doi: https://doi.org/10.29173/iq926 a university or on a page of their book’” (moyer, 2009). although this was not the economic model that the world of writing operated on during the mid-2000s this was clearly the model that the sophie 1.0 and sophie 2.0 platforms were based on. moreover, in the “what if” article describing the institute for the future of the book, the author discusses how “adherents of the network book, though, such as online wired magazine’s kevin kelly, suggest a luminous future for reading as the number of available books rises with the creation of what is often called the ‘universal library’: ‘plans like google’s [to digitize out-of-copyright and out-of-print books] will allow all the books in the world to become a single liquid fabric of interconnected words and ideas” (moyer, 2009). sophie was built for a world where all information is free and available online. for example, in sophie 2.0 “there is no statistical information to know how many users uploaded, downloaded, or liked a book” (ochoa et al., 2013). part of why there is so little information available about sophie 1.0 and sophie 2.0 was that there were no mechanisms in the platforms themselves for recording such information. this would have made it harder to justify their use by anyone since it would be hard to prove that the software was ever being used. there is a little bit of a ranking feature for individual users on the spinoff virtualbookclub platform with sophieserver with choices of “good,” “ok,” and “bad” and public annotations (hirschfeld et al., 2008). however, this was only on virtualbookclub and the people who actually used sophie 2.0, such as the people reviewing e-book platforms for open educational content were unaware of this by the time they reviewed and these “good” “ok” and “bad” rankings do not tell us very much about who is using these books (ochoa et al., 2013). looking at stein’s writings from the institute for the future of the book, he clearly felt a sense of responsibility for bringing about an inevitable techno-utopian vision of the world. for example, stein describes an institute for the future of the book experiment with commenting and annotation software with “many of the earlier reviewers said the same thing: it is no longer the author speaking, it is now the book speaking” (stein, 2008). this is an erasure of the individual from the experience of text and a promotion of a world where the group overrules the individual. the increased focus on social features of the platform and experiments at the institute for the future of the book in developing social networking capabilities show this dedication to a collective knowledge future where the individual was less important than the whole. after a talk in 2008 for serials librarians, stein was asked about authors who may not want to have reader comments in their work or to change their work based on reader comments, and replied “i think that is a valid question you ask, but i would argue that over time what is going to emerge is a form of expression where artists will take for granted and want an intersection with the reader. the role of the reader and author is going to morph or merge in some way. i’m all for people doing things the way they want to do them, but i think things are going to change in the direction that i’m talking about” (stein, 2008). here is an example where stein does not really give a satisfactory answer to a valid question but rather believes that the culture will inevitably change because of the change in technology. therefore, the job of the creators of sophie was to create the technology that would bring about the culture of the future rather than to build the technology for the culture that existed at the time. sophie was more than just a software for creating interactive multimedia e-books that readers could annotate, it also was the responsibility of the institute for the future of the book to wrestle with important questions of “what does it mean to be human in the age of the digital network” (stein, 2008). during this same talk he warned “in fact, if we just cling to the past, then the techies who actually don’t think about things are going to invent a future for us that we’re really not going to like” (stein, 2008). this is the idea of the responsibility of librarians https://doi.org/10.29173/iq926 12/16 kirby, jasmine s. (2018) how not to create a digital media scholarship platform: the history of the sophie 2.0 project, iassist quarterly 42 (4), pp. 1-16. doi: https://doi.org/10.29173/iq926 and scholars to actively choose and shape what technology they use to share their work. therefore, the marketing strategy for this product seemed to be that in a world where people have many options for how they express their work, people will choose the forms that bring about the vision of the world they want to create. this raises the question of whether sophie was ever supposed to work or was it just a way to promote the values espoused by the institute for the future of the book and the institute for multimedia literacy at the university of southern california. conclusion both sophie and sophie 2.0 fit into a greater history of using new technological advancements as a form of outreach and a way to showcase a world people want to exist. in her article describing the history of the chicago daily news radio and photogravure hybrid programming, good discusses how “first, it is not clear exactly what kind of benefit these ‘supplemental’ extensions of the newspaper brought to the daily news in terms of sales or readership. the daily news approached both the citywide lecture series and the radio photologues as a public service that would boost the paper’s image as a progressive, civic-minded source of information and cultural uplift”(good, 2017). on the sophie 2.0 “about page” usc’s role in the sophie 2.0 project was described as “providing project oversight and evangelism to the academic community” (visel, 2008). the choice of the word “evangelism” is telling in illuminating the heart of the purpose of this project. sophie was promoting a certain vision of the future. sophie embodies the values of free and open source, multimedia document creation, and new forms of authorship and intellectual property that the creators believed were going to be the future. unfortunately, they did not create a product that worked in its present time. it is not the job of librarians and digital humanists to use software we hope will work because it aligns with values we find important, it is our job to recommend and contribute to digital tools that won’t eat our users’ homework. although information about sophie 2.0 and the developer site can still be accessed today, there is little word of what happened to the project and its history. you basically have to call dan visel and ask what happened. scalar is probably the closest thing we have to a platform for creating nonlinear online books where users can incorporate and annotate multimedia content, or what usc and the institute for the future of the book had hoped to accomplish with sophie 2.0. although everyone i talked to insisted that these platforms were not at all related and have different histories, scalar also received funding from the mellon foundation and is based at the university of southern california. tara mcpherson, who edited the volume that the “a pedagogy for original synners,” the article that argues for releasing incomplete user-unfriendly open source software, is currently a pi for scalar. while scalar is much better to use on many aspects, such as being designed around copyright and having very helpful tech support, i worry about the sustainability and digital preservation issues with this software, especially as more of us use it for scholarly purposes. moreover, as a user of open source projects for digital humanities, i’m all too familiar with the issues of buggy software and lost work. while certainly not as extreme as in the case of sophie 1.0 and sophie 2.0, when paired with not publicly discussing https://doi.org/10.29173/iq926 13/16 kirby, jasmine s. (2018) how not to create a digital media scholarship platform: the history of the sophie 2.0 project, iassist quarterly 42 (4), pp. 1-16. doi: https://doi.org/10.29173/iq926 the issues that led to sophie 2.0’s failure, makes me question what, if anything, we as digital humanists have learned from our mistakes. takeaways, questions to ask when evaluating a digital scholarship tool • has this been updated? when was the most recent update? are updates regular or sporadic? • how is this project funded? is this grant funded? what happened to other recipients of these grants? • how well does the software work? has the software ever worked? is this software supposed to work? • where can i get tech support? how fast is the response for questions about glitches? are users told to fix glitches themselves? • was this software created accounting for intellectual property laws and other legal issues faced by users? • does this software claim to be meant for nontechnical users? is there documentation? is there a glossary for software-specific terms? • are there reviews? demo projects? are demo projects created by ordinary users or institutional groups with advanced it resources? references anderson, s. and balsamo, a. (2008), “a pedagogy for original synners”, in mcpherson, t. (ed.), digital youth, innovation, and the unexpected, mit press, cambridge, mass, pp. 241–259. artsj09. (2009), sophie presented by holly willis, available at: https://www.youtube.com/watch?v=kltcdfftiga (accessed 16 march 2018). ballatore, a. and natale, s. (2016), “e-readers and the death of the book: or, new media and the myth of the disappearing medium”, new media & society, vol. 18 no. 10, pp. 2379–2394. barnett, t. (2015), “platforms for social reading: the material book’s return”, scholarly & research communication, vol. 6 no. 4, pp. 1–23. bjork, o. (2012), “digital humanities and the first-year writing course”, in hirsch, b.d. (ed.), digital humanities pedagogy: practices, principles and politics, open book publishers, cambridge, uk, pp. 97–119. bridle, j. (2007), “sophie’s choice (a partial review)”, booktwo, 10 april, available at: http://booktwo.org/note-book/sophies-choice-a-partial-review/ (accessed 15 march 2018). dean, m.d. (2016), “a call to embrace social reading in higher education”, innovations in education & teaching international, vol. 53 no. 3, pp. 296–305. dombrowski, q. (2014), “what ever happened to project bamboo?”, literary and linguistic computing, vol. 29 no. 3, pp. 326–339. https://doi.org/10.29173/iq926 https://www.youtube.com/watch?v=kltcdfftiga http://booktwo.org/note-book/sophies-choice-a-partial-review/ 14/16 kirby, jasmine s. (2018) how not to create a digital media scholarship platform: the history of the sophie 2.0 project, iassist quarterly 42 (4), pp. 1-16. doi: https://doi.org/10.29173/iq926 fuchs, i.h. (2008), “challenges and opportunities of open source in higher education”, in katz, r.n. (ed.), the tower and the cloud: higher education in the age of cloud computing, educause, boulder, co, pp. 150–157. gaitán, s.b. (2011), “children of the screen: teaching spanish with commentpress”, learning through digital media, pp. 47–55. gold, m.k. (2011), “‘becoming book-like: bob stein and the future of the book’ (interview)”, kairos, vol. 15 no. 2, available at: http://technorhetoric.net/15.2/interviews/ (accessed 11 march 2018). good, k.d. (2017), “listening to pictures: converging media histories and the multimedia newspaper”, journalism studies, vol. 18 no. 6, pp. 691–709. herther, n.k. (2017), “top tools for digital humanities research”, computers in libraries, february, vol. 37 no. 1, p. 28. hirschfeld, r., haupt, m., rüger, m., brünn, p., esterluß, r., holz, n., knebel, k., et al. (2008), “sophieserver: the future of reading”, proceedings of the sixth international conference on creating, connecting and collaborating through computing (c5 2008), ieee computer society, washington, dc, usa, pp. 29–35. holz, n., hirschfeld, r., lincke, j., haupt, m. and rüger, m. (2009), “sophie tools and materials in multimedia book creation”, proceedings of the 2009 seventh international conference on creating, connecting and collaborating through computing, ieee computer society, washington, dc, usa, pp. 20–26. “how to kill an open source project with new funding”. (2008), slashdot, 3 october, available at: https://ask.slashdot.org/story/08/10/03/1547256/slashdot.sourceforge.net (accessed 6 february 2018). hudson, r. (2008), sophie, research report, transliteracies: research in the technological, social, and cultural practices of online reading, available at: http://transliteracies.english.ucsb.edu/post/research-project/researchclearinghouseindividual/sophie-2 (accessed 6 february 2018). koychev, i., dicheva, d. and nikolov, r. (2013), “smartbook: semantics inside”, serdica journal of computing, vol. 4 no. 2, pp. 263–278. lanier, j. (2010), you are not a gadget: a manifesto, 1st ed., alfred a. knopf, new york. “lpandeff”. (2009), “development_overview – sophie 2.0”, wiki:development_overview, 15 october, available at: http://sophie2.org/trac/wiki/development_overview (accessed 19 march 2018). morozov, e. (2013), to save everything, click here: the folly of technological solutionism, first edition, public affairs, new york. https://doi.org/10.29173/iq926 http://technorhetoric.net/15.2/interviews/ https://ask.slashdot.org/story/08/10/03/1547256/slashdot.sourceforge.net http://transliteracies.english.ucsb.edu/post/research-project/research-clearinghouseindividual/sophie-2 http://transliteracies.english.ucsb.edu/post/research-project/research-clearinghouseindividual/sophie-2 http://sophie2.org/trac/wiki/development_overview 15/16 kirby, jasmine s. (2018) how not to create a digital media scholarship platform: the history of the sophie 2.0 project, iassist quarterly 42 (4), pp. 1-16. doi: https://doi.org/10.29173/iq926 moyer, s. (2009), “what if?” humanities, august, vol. 30 no. 4, available at: https://www.neh.gov/humanities/2009/julyaugust/feature/what-if (accessed 8 march 2018). ochoa, x., casali, a., deco, c., gerling, v., frango, i., fager, j., carrillo, g., et al. (2013), “analysis of existing technological platforms for the collaborative production of open textbooks”, presented at the edmedia: world conference on educational media and technology, association for the advancement of computing in education (aace), pp. 1106–1115. rich, m. (2009), “curling up with hybrid books, videos included”, the new york times, 1 october, p. a1(l). rüger, m., stein, b. and visel, d. (2008), “sophie the future of reading”, proceedings of the sixth international conference on creating, connecting and collaborating through computing (c5 2008), ieee computer society, washington, dc, usa, pp. 13–20. “sophie”: (2015), 19 june, available at: https://web.archive.org/web/20150619042103/http://www.sophieproject.org/ (accessed 19 march 2018). stein, b. (2008), “the evolution of reading and writing in the networked era”, transcribed by creech, a. the serials librarian, vol. 54 no. 1–2, pp. 43–55. stein, b. and visel, d. (2010), “mao, king kong, and the future of the book”, triple canopy, no. 9, available at: https://www.canopycanopycanopy.com/contents/mao__king_kong__and_the_future_of_th e_book (accessed 11 march 2018). stein, r. and visel, d. (2007), “sophie and the future of reading and writing”, forum futures 2007, presented at the 2006 aspen symposium, forum for the future of higher education, cambridge, mass, pp. 57–60. virshup, a. (1996), “the teachings of bob stein”, wired, 1 july, available at: https://www.wired.com/1996/07/stein/ (accessed 13 march 2018). visel, d. (2008), “aboutpage – sophie 2.0”, 15 december, available at: http://sophie2.org/trac/wiki/aboutpage (accessed 15 march 2018). visel, d. (2018), interview with author, 7 february. end-notes 1 jasmine s. kirby is a subject liaison librarian for psychology and human development and family studies at the iowa state university of science and technology in ames, iowa. https://doi.org/10.29173/iq926 https://www.neh.gov/humanities/2009/julyaugust/feature/what-if https://web.archive.org/web/20150619042103/http:/www.sophieproject.org/ https://www.canopycanopycanopy.com/contents/mao__king_kong__and_the_future_of_the_book https://www.canopycanopycanopy.com/contents/mao__king_kong__and_the_future_of_the_book https://www.wired.com/1996/07/stein/ http://sophie2.org/trac/wiki/aboutpage 16/16 kirby, jasmine s. (2018) how not to create a digital media scholarship platform: the history of the sophie 2.0 project, iassist quarterly 42 (4), pp. 1-16. doi: https://doi.org/10.29173/iq926 2 interview with sophie user who wished to remain anonymous conducted march 1, 2018 3 institute for the future of the book 4 interview with sophie user who wished to remain anonymous conducted march 1, 2018 5 interview with sophie user who wished to remain anonymous conducted march 1, 2018. 6 drm is an acronym for digital rights management https://doi.org/10.29173/iq926 vol21.4bacp fall 1997 11 ilses general overview this document describes the ilses (integrated library and survey-data extraction service) project. ilses: introduction general information ilses is a project of a number of dutch, german, french, and irish institutes. ilses has been accepted by the european commission under the fourth framework (telematics applications of common interest, telematics for libraries) programme as project lb-4050/c and will as such run from september 1996 to september 1999. project goals the ilses project aims to develop a service that enables individuals users to access and retrieve documentary information and empirical data related to large-scale surveys such as the binannual eurobarometer surveys. ilses is designed to serve both end-users and contentproviders of socio-economic information. end-users for end-users ilses facilitates the integrated access and retrieval of two different kinds of information: 1. documentary information as is commonly available from libraries, and 2. empirical data as archived by data archives. ilses allows end-users to extend literature research and searches with a focused access of empirical data which can or have been used in empirical research of the kind reported in the literature they review. this allows them to extend classic literature research with their own original empirical analyses of relevant data. ilses of course also offers the other route: directing dataanalysts searching for specific empirical data to be used in secondary analysis to literature in which results and outcomes of previous and similar analyses have been reported. content-providers for content-providers ilses offers tools and procedures for the normalization, cataloguing and controlled distribution of (distributed) holdings of the documentary and data resources mentioned above. such content-providers are the database administrating staffs of librarian and dataarchival institutions. ilses will enable them to drastically increase and improve the utilization of their information resources while at the same time reducing their support burden per information request or retrieval. general design ilses will be designed as an open system which can be applied to different kinds of library and data holdings. in this project, however, a pilot-application will be focused on socio-economic information as collected by large scale surveys, and on the associated literature. ilses is based on integrated relational databases of metainformation pertaining to both library and data-archive holdings, both of which are typically distributed over many different institutions. in order to productively connect such holdings with each other and with end-users, ilses provides a wide-area network interface utilizing internet and supporting browsing and retrieval tools such as www. ilses partners the following institutes are partners in the ilses project: iec progamma, groningen swidoc, amsterdam university of amsterdam zentral archiv, cologne the following institutes are associate partners in the ilses project, working in association with the university of amsterdam: trinity college, dublin cidsp, grenoble iec progamma ilses (integrated library and survey-data extraction service) project by david a. schweizer* 12 iassist quarterly introduction the interuniversity expertise center progamma is a not_for_profit cooperation of eight dutch, british, and belgian universities, established in 1989. the center promotes the development, quality, and use of computer applications for the social and behavioral sciences. progamma is recognized by the netherlands organization for scientific research (nwo) and supported by the dutch government. role in the ilses project progamma is the coordinating partner in the ilses project. the coordinating partner is responsible for communication with the european commission. progamma is also responsible for the management and financial administration of the project. progamma is also responsible for the development of many of the ilses software tools swidoc introduction the social science information and documentation centre (swidoc) of the royal netherlands academy of arts and sciences, promotes and facilitates the exchange and efficient use of information in social science research. role in the ilses project swidoc will deliver the necessary expertise in the library field for the development and evaluation of the ilses bibliographic data. swidoc will also function as content provider of bibliographic data, and will as such function as a test site for the content-provider tools of lib-ilses. university of amsterdam introduction the university of amsterdam, more specifically the faculty of political, social and cultural sciences (pscw) is since long one of the most prominent locations of advanced academic empirical social research in the netherlands. the faculty contains strong and productive nuclei of methodological expertise, and of innovative large_scale survey research. the combined expertise is particularly strong in the area of electoral research, public opinion studies, multi_level research and methodology of comparative research, all of which are of eminent importance to the ilses project. role in the ilses project the university of amsterdam will function as the main end-user representative. they will research user demands prior to the creation of the ilses tools, coordinate end-user testing at the cidsp and trinity college, and evaluate the results. throughout the project, the university of amsterdam will provide end-user feedback for the developing ilses tools. zentral archiv introduction the central archive for empirical social research at the university of cologne, germany (za) is a resource, research, and teaching centre for national and international comparative research. the za is responsible for archiving, processing and distributing the eurobarometer data, in cooperation with the inter-university consortium for political and social research (icpsr) and the swedish social science data service (ssd). within the framework of the international federation of data organisations for the social sciences (ifdo) and cessda, za has access to the data contained in social science data archives worldwide. role in the ilses project the central archive, on the background of rich research and data management experience and acknowledged competence in integrating international comparative data sets, will help develop the content-provider ilses tools to access survey-data and guarantee the consideration of international standards for data documentation. trinity college introduction trinity college in dublin, is one of the most important academic institutions in the republic of ireland. it is renowned for its library and documentary infrastructure, and for its sound empirical social research. the department of politics, which will be the most involved in the organization of the user validation workshops of the ilses project, is unusually strongly embedded in european social research and collaboration networks, owing to which it is particularly well placed for contributing to such validation studies. this department can be regarded as a prototypical environment for many of the end_users of ilses. role in the ilses project trinity college will serve as end-user test site for the ilses tools. the college will also help disseminate ilses through the organisation of workshops. trinity college is an associated contractor, associated to the university of amsterdam. cidsp introduction the centre d’information des données sociopolitiques, as intermediaries and suppliers, represent a rapidly developing new type of interactive data and documentation resource and interactive service providers. their experience with international organisations such as cessda will enable the input of international best practice and standardization to the ilses project. winter 1997 13 role in the ilses project cidsp will function as end-user test site for the ilses tools. cidsp will also help disseminate ilses through the organisation of workshops. cidsp is an associated contractor, associated to the university of amsterdam. ilses tools in the ilses project the following tools will be developed: administrator-ilses the technical heart of the system lib-ilses for library content-providers dat-ilses for content providers with data archives e-ilses for end-users on standalone machines net-ilses for end-users on the internet administrator-ilses administrator-ilses will be the technical heart of the system. it consists of the datadictionary, containing all metadata about available archive data and bibliographic data, plus the interfaces that access this metadata. administrator-ilses will consists of a formal specification, plus a set of (ms-windows-) dlls that implements this specification. content-providers or other interested parties seeking to expand upon the ilses service may get access to administrator-ilses. the specification of administrator-ilses will be developed by the university of amsterdam, in consultation with the other project members. the implementation of the software modules will be done by iec progamma. according to the ilses planning, the first version of administrator-ilses will be ready in march 1997. lib-ilses lib-ilses will enable the library content-provider to produce and maintain the metadata needed for the ilses end-users to access bibliographic information through net-ilses and e-ilses. lib-ilses will contain of a set of tools to set up and maintain the metadata and will contain also connections to give on-line access to existing bibliographic and documentary databases and their interrelations. lib-ilses will consist of a set of mainly pc-based programs, that will be available for library contentproviders. the first version of lib-ilses will be developed by iec progamma, in consultation with the swidoc, the central archive and the university of amsterdam. the lib-ilses tools will then be installed at swidoc to be integrated with existing library systems. at least 250 new documentary records will be compiled at swidoc and included into ilses. according to the ilses planning, the first version of libilses will be available april 1998. dat-ilses dat-ilses will enable the data-archive content-provider to produce and maintain the metadata needed for the ilses end-users to access statistical data through net-ilses and e-ilses. dat-ilses will contain of a set of tools to set up and maintain the metadata and will contain also connections to give on-line access to existing statistical data archives. end-users will also be able to download selected sets of statistical data. dat-ilses will consist of a set of pcand unix(tm)based programs, that will be available for data-archive content-providers. the first version of dat-ilses will be developed by iec progamma, in consultation with the central archive, the university of amsterdam, and swidoc. the dat-ilses tools will then be be installed at the central archive and integrated with the local retrieval system. a test base for selected eurobarometers will be provided. according to the ilses planning, the first version of datilses will be available november 1997. e-ilses e-ilses is the tool for the end-user to access networked information resources, as services to be provided by both lib-ilses and dat-ilses. however, it should also be capable of operating stand-alone. this is a strict requirement from the research-user who wants to be able to use e-ilses stand-alone as informationand data extractor for distributed spss study files and documentation. this implies that some functionality of the lib-ilses and dat-ilses will also be incorporated into the ilses user tool. when operating stand-alone e-ilses can use any pcbased common database format, including those available through odbc. when operating in client-server mode any relational database capable of understanding sql can be used and manipulated. 14 iassist quarterly through e-ilses, the user may: examine content of archives and retrieve references to articles relating to the holdings combine the contents of several holdings to set up a new study retrieve the selected data e-ilses will be developed as a stand-alone pc-program, running under windows 3.x or windows 95/nt. e-ilses will be developed by iec progamma, in consultation with the university of amsterdam, the central archive, and swidoc. according to the ilses planning, the first version of eilses will be available in june 1997. net-ilses next to the stand-alone version for end-users, e-ilses, a www-based service for end-users will also be developed. this service will offer (part of) the functionality of eilses, and possibly more. as the technology of the internet is very rapidly progressing, the form net-ilses will eventually take, and the relation with e-ilses, will continually be re-evaluated in order to follow technology and end-user demands. net-ilses will be developed by iec progamma, in consultation with the central archive, the university of amsterdam, and swidoc. net-ilses will be developed in the form of one or more html-pages with corresponding code, tools, applets etc. according to the ilses planning, the first version of netilses will be available august 1998. morerinformation more information about ilses can be found at the ilses home page, at url http://www.gamma.rug.nl/ilses there is a mailing list that will keep interested people updated about the progress of the project, and encourages discussions about ilses. it can be subscribed to via the ilses home page, or by sending a mail message to listserv@nic.surfnet.nl, with the message text subscribe ilses-l yournamegoeshere (put your name in the location yournamegoeshere.) you can also contact the project coordinator at iec progamma: david a. schweizer, iec progamma, p.o. box 841, 9700 av groningen, the netherlands, tel: +31 50 363 6900, fax: +31 50 363 6687, e-mail: gamma.post@gamma.rug.nl * paper presented at iassist/ifdo ‘97, odense, denmark, may 6-9,1997. http://www.gamma.rug.nl/ilses mailto:listserv@nic.surfnet.nl mailto:gamma.post@gamma.rug.nl 1/26 fehrmann, paul & mamolen, megan (2020) methods reporting that supports reader confidence for systematic reviews in psychology: assessing the reproducibility of electronic searches and first-level screening decisions, iassist quarterly 44(1-2), pp. 1-26. doi: https://doi.org/10.29173/iq968 methods reporting that supports reader confidence for systematic reviews in psychology: assessing the reproducibility of electronic searches and first-level screening decisions paul fehrmann1 and megan mamolen2 abstract recent discussions and research in psychology show a significant emphasis on reproducibility. concerns for reproducibility pertain to methods as well as results. we evaluated the reporting of the electronic search methods used for systematic reviews (sr) published in psychology. such reports are key for determining the reproducibility of electronic searches. the use of sr has been increasing in psychology, and we report on the status of reporting of electronic searches in recent sr in psychology. in all, we used 12 checklist items to evaluate reporting for electronic strategies. kappa results for most of the items developed from evidence-based recommendations, ranged from fair to almost perfect. data for a stringent ‘prisma’ type of recommended reporting showed that only one of the 25 randomly selected psychology sr from 2009-2012 reported recommended information for all items in the set, and none of the 25 psychology sr from 2014-2016 did so. results for a second less stringent set found that only 36% of the psychology sr reported basic information that supports confidence in the reproducibility of electronic searches. using those two sets of checklist items found similar results for psychology sr published in 2017. moreover, reporting was also very infrequent for a third supplemental set of ‘confidence items’. fuller and clearer recommended reporting of the electronic searches used in sr would provide a stronger basis for confidence in the reproducibility of searches. that reporting, in turn, would strengthen reader confidence more generally in the results and conclusions reached in sr in psychology. keywords systematic reviews, psychology, reproducibility, electronic searches, reporting, confidence 1. introduction 1.1 background recent discussions have shown a significant emphasis on the reproducibility of research in psychology (pashler and wagenmakers, 2012; yong, 2013; cooper and vandenbos, 2013; novotney, 2014; open science collaboration, 2015; gilmore, diaz, wyble, & yarkoni, 2017). systematic reviews are viewed by many as a top level synthesis of research (paul & leibovici, 2014; kisely et al., 2015; ng & benedetto, 2016), and a special emphasis on reproducibility and systematic reviews is evident in the recent psychological bulletin focus on the topic ‘replication and reproducibility: questions asked and answered via research synthesis’. concerns for reproducibility pertain to methods as well https://doi.org/10.29173/iq968 2/26 fehrmann, paul & mamolen, megan (2020) methods reporting that supports reader confidence for systematic reviews in psychology: assessing the reproducibility of electronic searches and first-level screening decisions, iassist quarterly 44(1-2), pp. 1-26. doi: https://doi.org/10.29173/iq968 as results, and the current paper contributes to discussions about the methods that are used for systematic reviews. 1.2 rationale for our reproducibility research briefly stated, our research has been motivated by the following. first, in general, confidence in the methods used for systematic reviews provides a basis for confidence in the conclusions reached in systematic reviews. second, and more specifically, electronic searches provide a crucial part of the methods used to identify data and build the evidence base that is then used for conclusions reached in systematic reviews. third, the reproducibility of electronic searches has been a core recommendation in guidance on systematic review methods. fourth, full and transparent reporting of electronic search methods is key for reader confidence in the reproducibility of electronic searches. fifth, important reporting details also can support reader confidence in the reproducibility of first level screening of electronic searches. and, sixth, confidence in the reproducibility of searching and screening then supports the confidence that readers can have in the conclusions that are reached in systematic reviews (cooper, 2017; vazire, 2017: golder et al., 2013; pashler & wagenmakers, 2012). meta-research has recently cautioned that ’…even when research findings are reported, they can be undermined by a lack of transparency about how they were generated (hardwicke et al., 2020), and the target of our study has been what is described by goodman and others as ‘methods reproducibility’. they argue that methods reproducibility ‘…refers to the provision of enough detail about study procedures and data so the same procedures could, in theory or in actuality, be exactly repeated (goodman, fanelli, & ioannidis, 2016, p. 2). in response, to develop a picture of the basis readers have for confidence in the results of systematic reviews in psychology, we have been evaluating the reporting of electronic searches used for systematic reviews in psychology. 1.3 systematic reviews, reproducibility, electronic searches, data management, reporting, and psychology systematic reviews (sr) have been widely recognized both as methods and as products that involve rigorous, transparent processes for identifying, analyzing, and synthesizing data from primary studies to draw conclusions relevant to a research topic (cooper, 2017, p. 10; gough & oliver, 2012). sr are also given special recognition in a recent national academy of sciences (nas) report ‘reproducibility and replicability in science’ (committee on reproducibility and replicability in science et al., 2019), particularly in a chapter entitled ‘confidence’ which has this conclusion ”multiple channels of evidence from a variety of studies provide a robust means for gaining confidence in scientific knowledge over time’ (committee on reproducibility and replicability in science et al., 2019, p. 155). importantly, in that nas chapter, sr are noted as a significant for gaining confidence in science. sr by design are ’the ensemble of research activities involved in identifying, retrieving, evaluating, https://doi.org/10.29173/iq968 https://www.nap.edu/read/25303/chapter/10 3/26 fehrmann, paul & mamolen, megan (2020) methods reporting that supports reader confidence for systematic reviews in psychology: assessing the reproducibility of electronic searches and first-level screening decisions, iassist quarterly 44(1-2), pp. 1-26. doi: https://doi.org/10.29173/iq968 synthesizing, interpreting, and contextualizing the available evidence from studies on a particular topic’; and as such they address ‘…the central question of how the results of studies relate to each other, what factors may be contributing to variability across studies, and how study results coalesce or not in developing the knowledge network for a particular science domain’ (committee on reproducibility and replicability in science et al., 2019, p. 144). mirroring the nas discussion, a recent introduction from the joanna briggs association explains that sr are specifically designed to achieve results and reach conclusions based on analyzing and synthesizing “all” of the evidence or data relevant to a question (aromataris & munn, 2019). of course an important methological question for readers of sr is ’how did you get that data ?’. sr typically use a range of strategies to find resources that have the data that is extracted, analyzed and synthesized, including use of databases, grey literature, scanning reference lists of key articles, hand searching of journals, web sites, and contacting experts (kugley et al., 2017). however, although different strategies are used for the comprehensive searching that is often stressed as a characteristic of sr, electronic searches have been noted as providing ‘…the largest portion of the evidence base for systematic reviews’ (sampson et al., 2009, p. 944), and a standard expectation is that the electronic searches used will be reproducible (centre for reviews and dissemination, 2009; liberati et al., 2009; lefebvre et al., 2019; agency for healthcare research and quality, 2014; kugley et al., 2017). a follow up question for readers of sr is how to determine if electronic searches used in sr are reproducible. readers might use what is reported about those searches in an attempt to rerun the searches (e.g., see ali & usman, 2018). or, researchers might give their electronic search strategy report to a peer to have them execute the search, and then report the peer review results for readers (mcgowan et al., 2016; faggion, 2019). of course, most readers of sr, including clinicians and policy makers, do not have the time or resources to attempt a rerun. although an alternative is to see if information recommended for confidence in the possibility of reproducing the electronic search is actually reported, research indicates that greater transparency is needed to support that confidence (campbell et al., 2019). as stressed recently, without ‘clear signals’ of practices that increase the ’trustworthiness of scholarly work’, readers are challenged when ‘ascertaining confidence’ they might have in that scholarly work. a recommendation is that such signals would include information that confirms ‘adherance to field-specific reporting requirements’ (jamieson, mcnutt, kiermer, & sever, 2019). internationally accepted guidance on reporting sr search steps has been available for many years in the prisma statement. as the authors of that guidance stated “…systematic reviews should be reported fully and transparently to allow readers to assess the strengths and weaknesses of the investigation” (liberati et al., 2009)3. importantly, that guidance includes descriptions and examples of what should be reported of sr electronic search methods, and the current study drew on that methods guidance to address our research objectives. with respect to the field of psychology, as indicated in table 1., there is clear evidence of growing use of sr. https://doi.org/10.29173/iq968 https://wiki.joannabriggs.org/display/manual/1.1+introduction+to+jbi+systematic+reviews http://www.prisma-statement.org/ 4/26 fehrmann, paul & mamolen, megan (2020) methods reporting that supports reader confidence for systematic reviews in psychology: assessing the reproducibility of electronic searches and first-level screening decisions, iassist quarterly 44(1-2), pp. 1-26. doi: https://doi.org/10.29173/iq968 even though recent work has reiterated that electronic searches are insufficient for confidence in the comprehensiveness of searches (delaney & tamás, 2018), readers of sr still do have a key support for confidence in sr results if they see reporting that helps to ensure the reproducibility of electronic searches used. actually, determining reproducibility is dependent on what is found in this report (niederstadt & droste, 2010; rader et al., 2014; mullins, 2014; atkinson et al., 2015; schalken & rietbergen, 2017), and in our study of methods we evaluate the reporting of electronic search methods that are used for sr in psychology. to date we have not found this kind of assessment of sr methods in psychology4. 1.4 definitions and abbreviations other recent work has pointed to variation in how researchers define and assess ‘reproducible searches’ (sayre & riegelman, 2018; koffel & rethlefsen, 2016; ali & usman, 2018). table 2 contains stipulated definitions for this paper, including definitions for the reproducibility of electronic searches and for confidence in the reproducibility of electronic searches. additionally, in this paper we use the abbreviations that are indicated. for example, ‘search’ means electronic search. https://doi.org/10.29173/iq968 5/26 fehrmann, paul & mamolen, megan (2020) methods reporting that supports reader confidence for systematic reviews in psychology: assessing the reproducibility of electronic searches and first-level screening decisions, iassist quarterly 44(1-2), pp. 1-26. doi: https://doi.org/10.29173/iq968 https://doi.org/10.29173/iq968 6/26 fehrmann, paul & mamolen, megan (2020) methods reporting that supports reader confidence for systematic reviews in psychology: assessing the reproducibility of electronic searches and first-level screening decisions, iassist quarterly 44(1-2), pp. 1-26. doi: https://doi.org/10.29173/iq968 in addition to the stated definitions for the terms ‘electronic resource’ and ‘electronic search’, for this paper these terms refer to a range of approaches for finding the information sources that are eventually selected for use in sr.electronic searches can involve using discipline-specific commercial resources (e.g., ebsco’s psycinfo), public resources (e.g., pubmed), web search engines (e.g. google scholar), systematic review library databases (e.g., campbell library; cochrane library), electronic grey literature resources (e.g., proquest dissertations and theses), scholarly ‘cited reference’ databases (e.g., scopus, web of science), and a range of other possibilities. also, in our definitions for electronic resource, electronic search result, and electronic search report we refer to data. the extracting or selection of data from resources found using electronic searches is basic to sr, and in our definitions the term ‘data’ refers to the quantitative or qualitative information selected and extracted from multiple sources and then analyzed and synthesized to address research questions in sr. 1.5 research objectives 1. we sought to compare reporting of electronic searches in psychology sr to the reporting of electronic searches in sr that are known to be completed with rigorous methods for reporting. 2. most if not all research on reporting of sr has looked at reporting of individual items (individual steps) used for electronic searches. assuming that higher levels of reproducibility are supported as more of a set of search steps are reported as recommended, we looked to document the extent that psychology sr provide reporting of electronic searches according to a set of widely accepted recommendations for what should be reported. as detailed in our methods section, we called this a prisma set. data and discussion our assessment of a ”non-prisma” set are available in supplementary files on this paper’s osf site5 (hereinafter ’the osf site’). 3. there is information that is not needed to execute and see results for what we defined as a reproduced electronic search, but which nevertheless can support reader confidence in the reproducibility of searches used. we looked to document reporting for this kind of information in campbell and psychology sr. 2. methods 2.1 checklist items used for this study amstar, prisma, and press are three major resources providing guidelines to support the design, execution, and evaluation of sr, and we discuss those resources and the evaluation of reproducibility in supplementary files on the osf site. other authors have also created checklists or reporting guidelines relevant to evaluating electronic searches (e.g., booth, 2006; yoshii et al., 2009; maggio et al., 2011; atkinson et al., 2015). editors have also been urged to use checklist items to check reporting in order to ’protect the reporting process’ and ‘to signal the trustworthiness of science’ (jamieson, mcnutt, kiermer, & sever, 2019). https://doi.org/10.29173/iq968 https://osf.io/g6x4k/ 7/26 fehrmann, paul & mamolen, megan (2020) methods reporting that supports reader confidence for systematic reviews in psychology: assessing the reproducibility of electronic searches and first-level screening decisions, iassist quarterly 44(1-2), pp. 1-26. doi: https://doi.org/10.29173/iq968 for this study we used twelve items drawn from a set of 36 checklist items created during a search evaluation project pursued over a number of years. these items focusing on reproducibility were based on evidence based recommendations in publications and manuals of major sr organizations (e.g., centre for reviews and dissemination, 2009; liberati et al., 2009; kugley et al., 2017; lefebvre et al., 2011; agency for healthcare research and quality, 2014). additional explanation for the eight basic and four additional confidence checklist items used is found in sections 2.3 and 2.4 below.6 2.2 identification and selection of sr for this study psycarticles and the campbell library were used to identify sr for this study. psycarticles is a resource for identifying ‘peer-reviewed publications of the american psychological association (apa) and affiliated journals’ that cover ‘the science of psychology and behavior’ (psycarticles, n.d.). the campbell library is also a resource that consists of peer-reviewed sr publications. we used the reporting for campbell collaboration sr (campbell sr) as a model or standard of comparison for assessing the psychology sr. similar comparisons have been reported in the health sciences (sampson et al., 2008; yoshii et al., 2009; popovich et al., 2012; golder et al., 2013). the searches used with psycarticles and with the campbell library are presented in table 3. https://doi.org/10.29173/iq968 8/26 fehrmann, paul & mamolen, megan (2020) methods reporting that supports reader confidence for systematic reviews in psychology: assessing the reproducibility of electronic searches and first-level screening decisions, iassist quarterly 44(1-2), pp. 1-26. doi: https://doi.org/10.29173/iq968 the 29 hits retrieved in psycarticles in december, 2012 were sr as defined by apa (apa databases methodology field values, n.d.), and for this study 25 articles were randomly selected to represent sr in psychology for the timeframe of 2009-2012. each of those 25 sr contained electronic searches. the 29 hits from psycarticles are listed in appendix 1, with 25 included for this study marked with asterisks7. the 69 hits noted in the updated psycarticles search in april 2016 were assumed to be https://doi.org/10.29173/iq968 9/26 fehrmann, paul & mamolen, megan (2020) methods reporting that supports reader confidence for systematic reviews in psychology: assessing the reproducibility of electronic searches and first-level screening decisions, iassist quarterly 44(1-2), pp. 1-26. doi: https://doi.org/10.29173/iq968 sr in psychology. this search was completed by the first author who examined all 69 to see if an electronic search was reported. ten of the set of 69 articles did not report electronic searches, and a random sample of 25 was selected from the remaining 59 to represent sr in psychology for the timeframe of 2014-2016 (as of date of search). appendix 3 lists all 69 of the second set of sr from psycarticles; and there we indicate both the ten which did not provide electronic search reports as well as the 25 randomly selected sr that we used. similarly, 25 campbell library articles were randomly selected from the 42 initial search results. the 42 campbell collaboration results are in appendix 2, with 25 sr marked that were included for this study. in the remainder of this paper the campbell library systematic reviews may be referred to as ‘campbell sr’ and the phrase ‘psych sr’ may be used to refer to the psycarticles assessed. a search update was used in 2017 to collect data for the second research objective; and appendix 4 shows the 45 sr articles identified along with indication of the 18 sr that were assessed for the current study. the 18 sr assessed were those papers that explicitly discussed prisma by name or cited and listed prisma in their references. like other recent studies (leclercq et al., 2019; page et al., 2020), we took that discussion or citing of prisma to indicate that the authors of those sr had seen prisma as a guide for their reporting of electronic searches used. 2.2. data collection for the psych sr, reports were in the article’s method section, in appendixes, or in supplemental files. for the campbell sr, the search reports were located either in a section entitled ‘search methods for identification of studies’ or in an appendix. we used qualtrics (qualtrics xm experience management software, n.d.) to create a data entry tool for checklist scores for the 93 sr assessed. before collecting data for this study, the authors pilot tested the 36 checklist items mentioned above using sr from psychology journals and sr from the campbell library. the sr used for pilot testing were not used as sources of data for addressing our research topics. prior to reaching consensus scores, search reports for the sets of sr published from 2009-2016 were evaluated and scored independently by the authors. checklist items ask about electronic search elements that are documented in electronic search reports. items are scored yes (y), provisional yes (ps), not sure (ns), no (n), or na. y for an item means that the evaluator believes that what is noted in that item is clearly reported. n means the opposite of y. ps means the evaluator feels confident they can guess what was done, and ns means the opposite of ps. na for an item means that the evaluator believes that what is noted in that item is not applicable. https://doi.org/10.29173/iq968 10/26 fehrmann, paul & mamolen, megan (2020) methods reporting that supports reader confidence for systematic reviews in psychology: assessing the reproducibility of electronic searches and first-level screening decisions, iassist quarterly 44(1-2), pp. 1-26. doi: https://doi.org/10.29173/iq968 figure 1 shows a copy of the first checklist item as it appeared on the data entry tool. as shown, two levels of assessment (a. and b.) were used for that search report item. a comment box allowed for qualitative observations used for our discussions as we reached consensus-scoring decisions for items. 2.3 data analysis we used and report below on three approaches to assess the electronic search reports in sr drawn from the psychology literature. information and data for one additional approach is on the osf site. campbell collaboration comparison. first, we evaluated campbell sr to determine report frequencies for eight basic, individual search elements. sr in the campbell library, like those completed under guidance of the cochrane collaboration, are completed using guidance that increases the possibilities for higher quality reporting of searches (kugley et al., 2017). moreover, because the cochrane collaboration has been a leader for guidance and high quality sr, studies in the health sciences have compared the quality of reporting in cochrane sr to that found in noncochrane sr (moher et al., 2007; sampson et al., 2008; popovich et al., 2012). mirroring the health studies research, then, we compared reports in sr published in the psychology literature to those sr from the campbell library. we used fisher exact tests to evaluate the significance of difference in proportions between the report frequencies for the psych and campbell sr. prisma set. as a key approach to assessing psych sr for 2009-2012, 2014-2016, and 2017, a set of basic search element reporting recommendations that are found in the prisma statement were used. the prisma guidance urges the reporting of the ‘full electronic search strategy’ for at least one electronic resource used (liberati et al., 2009). above we noted our evaluation of the frequencies of reporting for individual recommended search elements related to reproducibility. however, when evaluating a specific sr paper, readers are likely to be concerned with reporting for sets of search elements relevant to that sr. to that end, we used the search elements listed below to see the extent to which a prisma type of full electronic search strategy has been reported in psychology sr. we also evaluated the campbell sr. to date we have not found this kind of assessment of sr in psychology. https://doi.org/10.29173/iq968 11/26 fehrmann, paul & mamolen, megan (2020) methods reporting that supports reader confidence for systematic reviews in psychology: assessing the reproducibility of electronic searches and first-level screening decisions, iassist quarterly 44(1-2), pp. 1-26. doi: https://doi.org/10.29173/iq968 below we briefy explain our choice of search elements. we then explain how we used those elements. our selection of four elements can be viewed as representing what is recommended in the prisma statement as a report of the ‘full electronic search strategy’ (liberati et al., 2009). these elements also reflect relevent items in the american psychological association’s ‘meta-analysis reporting standards’ (apa publications and communications board working group on journal article reporting standards, 2008).8 moreover, this set corresponds to an overlap of elements provided in two papers which used extensive methods for identifying key elements, including consultation with groups of searching experts. see table 3 in mullins et al. (2014), and table 1 in rader et al. (2014). this list also overlaps with search reproducibility elements noted in other papers (sampson, et al., 2008; yoshii et al., 2009; atkinson et al., 2015; meert et al., 2017; koffel & rethlefsen, 2016), and corresponds to recommendations in the campbell collaboration guidelines (kugley et al., 2017). the 1-4 order of the list below tracks the order of percentages previously reported for the number of sr reporting those information elements (mullins et al., 2014, see their table 5). those percentages are listed here in parentheses. 1. the names of all electronic resources used (94% of 102 sr reporting). 2. the publication time frames of articles to be included (34% reporting). 3. copies of search strategies for any electronic resource used (13% reporting for at least one or all databases used). 4. the vendor names of any electronic resource used (8% reporting). we assumed that higher levels of reproducibility are supported as more of a set of recommended reporting criteria are met, and we evaluated electronic searches using the following sequence. first, in spss we identified those sr which were given a consensus yes score for item 1 just above. we then identified those sr that were given a consensus yes score for both items 1 and item 2. next we identified sets with positive scores for 1, 2, and 3. finally, we identified a set with positive scores for 1-4. by using the basic percentage order indicated to sequence our analysis we looked to prevent premature elimination of those sr which had not reported the names of vendors but which had provided copies of search strategies of every database used. confidence items. during our current research project, we developed and used four items that go beyond the basic report information typically required by readers if they want to actually reproduce an electronic search. that is, as one reads and evaluates a systematic review (sr), if rerunning an electronic search is not feasible then these are supplemental report elements that may enhance confidence in search reproducibility and confidence in the potential use of searches. two supplemental search report elements involve reporting the final number of hits for each database search, and reporting the final number of hits combined across the databases used. this kind of reporting can be found in sr that use a prisma type flowchart (see the prisma statement at prisma-statement.org; also see the ‘flow diagram’ in gensby et al., 2012). https://doi.org/10.29173/iq968 12/26 fehrmann, paul & mamolen, megan (2020) methods reporting that supports reader confidence for systematic reviews in psychology: assessing the reproducibility of electronic searches and first-level screening decisions, iassist quarterly 44(1-2), pp. 1-26. doi: https://doi.org/10.29173/iq968 a third confidence element involves reporting the use of two researchers for inclusion/exclusion decisions when viewing the title and abstracts (mcdonagh et al., 2008) and also reporting inter-rater agreement results (e.g., kappa) for those inclusion/exclusion decisions (liberati et al, 2009). as indications of potentially reduced selection errors, good reporting for either of these would support reader confidence in the possibility of reproducing the electronic search first level inclusion choices. that is, there could be support for increased confidence that the original use of electronic search results (‘hits’) could be reproduced. a fourth confidence element would be reporting information that identifies hits actually chosen for potential inclusion as a result of evaluating the title and abstract. this is prior to evaluating the full text. if a reader reproduced an electronic search process and saw that the items they choose while screening the titles and abstracts match those identified by the original researcher, then that reader can have increased confidence that they are mirroring the original researcher’s use of inclusion criteria with the electronic search hits. actually, without rerunning a search, just seeing this information reported would support reader confidence in the possibility of reproducing the selection process as well as the searches used. to briefly summarize, this kind of supplemental information in search reports can support the assumption of reproducibility for the search and selection process used for sr. as this paper was prepared, except for an increase in use of prisma type flowcharts showing the number of hits for searches, the reporting of information for these elements has been very infrequent. additionally, with a recent exception (schalken & rietbergen, 2017), looking for reporting of this information also does not seem to be a part of assessing reports of searches (e.g., atkinson et al., 2015; booth, 2006; golder et al., 2008; golder et al., 2013; maggio et al., 2011; mullins et al., 2013; niederstadt and droste, 2011; tunis et al., 2013; yoshii et al., 2009; koffel and rethlefsen, 2016; meert et al., 2016). as a contribution to discussions of the reproducibility of searches and confidence in the use of search results, we looked at reporting for these four supplemental elements across our set of 93 sr from campbell and psycarticles. 3. results and discussion 3.1 inter-rater agreement and potential use for basic checklist items the disccussion and table below provide a look at data for the eight basic checklist items we developed as well as for our use of those items to address our first two research questions. section 3.4 presents similar discussion for our four confidence items. similar to the item assessment reported for the development of amstar (shea et al., 2009), we looked at inter-rater agreement for individual checklist items. some important studies that have evaluated reports of search strategies have not included inter-rater agreement results for the elements or items used to evaluate search reporting (shea et al., 2002; sampson et al., 2006; moher et al., 2007; yoshii et al., 2009; golder et al., 2013; mullins et al., 2014). other recent related papers have provided this kind of checklist information (willis and quigley, 2011; fehrmann and thomas, 2011; popovich et al., 2012; pieper et al., 2015; meert et al., 2016). that said, the current paper is the first to report this kind of reliability information for individual items on a checklist specifically https://doi.org/10.29173/iq968 13/26 fehrmann, paul & mamolen, megan (2020) methods reporting that supports reader confidence for systematic reviews in psychology: assessing the reproducibility of electronic searches and first-level screening decisions, iassist quarterly 44(1-2), pp. 1-26. doi: https://doi.org/10.29173/iq968 developed and worded to assess the reporting that supports the reproducibility of electronic searches. kappa coefficients were used to assess inter-rater agreement, and, following landis and koch (1977), those coefficients were interpreted with these categories: below chance considered poor; 0.01 to 0.20 slight agreement; 0.21 to 0.40 fair agreement; 0.41 to 0.60 moderate agreement; 0.61 to 0.80 substantial agreement; and 0.81 to 1 almost perfect agreement. the asterisks *** in the table 4 indicate those items where we scored all sr articles as ‘no’ or ‘yes’ and so kappa could not be calculated. the asterisks indicate 100 percent agreement. as shown in table 4, our testing with eight basic multi-level items showed them overall to have ‘fair’ to ‘almost perfect’ inter-rater scores (kappa). our findings do indicate challenges (e.g., item 6 kappa results for the campbell sr), and more training with our items might have improved our kappa results. others found repeated refining of and training with their ‘keyword’ item improved the kappa to .64 for an item similar to item 6 in our table 4 (meert et al., 2016). similarly, others reported a kappa ‘range’ of .66 – 1 (willis and quigley, 2011), using items from the prisma statement (liberati et al., 2009), indicating that the prisma items relevant to reproducibility (6, 7, 8) had kappas at or above the lower end of that range. https://doi.org/10.29173/iq968 14/26 fehrmann, paul & mamolen, megan (2020) methods reporting that supports reader confidence for systematic reviews in psychology: assessing the reproducibility of electronic searches and first-level screening decisions, iassist quarterly 44(1-2), pp. 1-26. doi: https://doi.org/10.29173/iq968 it does seem, then, that the current set of checklist items could usefully serve others as part of an approach to evaluating search reports to determine reproducibility. the use of items specifically worded to assess search report elements provides a consistent framework for assessing searches. this is similar to the specific items found in amstar9 and the press tool10. that said, just as others have invited researchers to use their checklists (shea et al., 2009; tong et al., 2012; meader et al, 2014), additional work with the checklist items we used could have value for research for assessing electronic searches, and for determining and possibly extending the value of these items. also promising for checklists that researchers, readers, and editors can use is the related work underway both on the overall prisma update (“prisma,” n.d.; page et al., 2020), and on the prisma-s11 extension that focuses more generally on the fuller set of searches done for sr (including electronic searches). 3.2 the status of recommended search reporting in psychology sr published from 2009-2017 table 4 presents data on the reporting in psychology sr for widely recommended search elements, and here we focus on results for our checklist item 1. more detailed comments on the other checklist items are available on the osf site. it appears unanimous in sr guidelines that a copy of the full electronic search is desirable and even expected for at least one of the electronic search sources used (liberati, et al., 2009; higgins and green, 2011; kugley et al., 2017), and item 1. was used to assess the reporting or provision of what we call copies of the electronic search strategy. for this study we looked for the kind of copy that is possible using a search resource function for saving or printing a search history, or for a copy and pasted representation of the search steps. a typical printout may well have the terms used (free text and thesaurus terms), and show how terms were combined, the sequence of entering terms and their combination, the use of adjacency, and the use of truncation, along with other details. this checklist item is different from our other checklist items, because such copies can potentially include information for many of those other items. in fact, depending on the resource used, and the search run, either now or in the future such copies might include all of the information indicated in the other checklist items that would be relevant for providing a ‘full electronic search strategy’ for that resource. providing such copies could be a straightforward way to efficiently and accurately show most if not all of the steps actually used with a given resource. our results showed that 14 of the 25 campbell sr provided such a copy for at least one electronic source used (see item 1.a., column 5). these articles are identified in appendix 2. only one of our set of 25 psych sr from 2009-2012 provided a copy for at least one electronic source used (see item 1.a., column 3, and article 24 in appendix 1). two of our psych sr sample for 2014-2016 provided this in their reports, and these are identified in appendix 3. four of our select set of 18 psych sr for 2017 provided copies for at least one resource used, and are indicated in appendix 4. frequency results for copies of search strategies for every electronic source used (item 1.b) were similar (psych sr, 1/25; campbell, 13/25).12 we did not find any in our 2016 psych sr sample that https://doi.org/10.29173/iq968 https://amstar.ca/ https://www.cadth.ca/resources/finding-evidence/press https://doi.org/10.17605/osf.io/ygn9w 15/26 fehrmann, paul & mamolen, megan (2020) methods reporting that supports reader confidence for systematic reviews in psychology: assessing the reproducibility of electronic searches and first-level screening decisions, iassist quarterly 44(1-2), pp. 1-26. doi: https://doi.org/10.29173/iq968 reported copies for all searches used, and 2 out of 18 psych sr assessed in 2017 provided copies for all resources used. assuming that having copies of search strategies supports confidence in reproducibility as well as facilitating the actual reproducing of electronic searches, the results above suggest that there should be more frequent provision of copies of search strategies for every electronic source used. others have similarly argued recently for this expanded reporting for electronic searches (shokraneh, 2019). however, even without such copies provided for every electronic resource listed in an sr, using the campbell sr results as a comparison that indicates what might generally be possible and expected, in keeping with the prisma guidance, the psychology sr could and should more frequently provide a copy for at least one of the electronic resources used. sr in psychology routinely use psycinfo, and such copies have been possible with psycinfo available from vendors such as ebsco, firstsearch, ovid, and with psycarticles from the american psychological association. although our results and interpretation will benefit from additional verification, there are concerns that approximately 90% of the 68 psychology sr we assessed did not provide copies for any of the searches used. assessment of sr published more recently than those we looked at could show improvements in this kind of reporting. as an alternative to providing copies as we defined them is not possible, researchers often list and report information for the individual search strategy elements that they used. we evaluated sr that used this approach to reporting, and we found that only 38% of those psychology sr reported basic information that supports confidence in the reproducibility of electronic searches . that assessment and data are discussed on the osf site (see ‘non-prisma’ in supplementary files). in the past the space limits for journals have been a challenge to detailed reporting, and key guidelines have recognized such challenges, even while calling for fuller reporting (liberati et al., 2009). recent signs, such as the use of supplemental online files, indicate that the online environment will help to reduce the impact of space limits (cooper & vandenbos, 2013; lebel et al., 2013; atkinson et al., 2015), and the journal archives of scientific psychology, and the open science framework13 are two venues that support provision of needed reporting information. however, in addition to reducing space constraints, the need for fuller reporting to support reproducibility will be addressed significantly if what we have described as copies are required as a part of all electronic searches in sr. 3.3 a prisma set. the reporting of recommended information for recommended sets of search elements. as our second approach to evaluating psych sr, and reporting for searches, the results in table 5 are for a set of elements/items that may be viewed as a fairly stringent prisma type of reporting https://doi.org/10.29173/iq968 https://osf.io/ https://osf.io/ 16/26 fehrmann, paul & mamolen, megan (2020) methods reporting that supports reader confidence for systematic reviews in psychology: assessing the reproducibility of electronic searches and first-level screening decisions, iassist quarterly 44(1-2), pp. 1-26. doi: https://doi.org/10.29173/iq968 pertaining to reproducibility. these results show what we found as we assessed our samples of psych sr and campbell sr. looking at item 3 in table 5, our findings show that, for the randomly selected 25 campbell sr that we examined, only 14 provided recommended information for the first three of our prisma set of recommended search report elements. in other words, forty-four percent did not provide this prisma search set information. the requirement of seeing vendor information reduced that number to 10 of 25 reporting desired information for our prisma set of reporting information elements. in comparison, the data show that the reporting for a prisma set in the 25 randomly selected psych sr for 2009-2012 is much lower. only 1 of those 25 psych sr provided information for that set of prisma elements. the data for the randomly selected set of psych sr for a 2014-2016 publication time frame showed that low reporting for that prisma set continued, and data for the 2017 set of sr was only slightly better. overall, our assessment shows about half of the campbell sr reporting for this prisma set, and reporting this information in the 68 sr from psychology is considerably lower. if electronic search reproducibility is viewed as dependent on reporting that is equivalent to what we call our prisma set, our findings suggest that readers would not have a strong basis for confidently assuming that the electronic searches in psychology sr are reproducible. this is a concern to address in the current studies and discussions of reproducibility in psychology. https://doi.org/10.29173/iq968 17/26 fehrmann, paul & mamolen, megan (2020) methods reporting that supports reader confidence for systematic reviews in psychology: assessing the reproducibility of electronic searches and first-level screening decisions, iassist quarterly 44(1-2), pp. 1-26. doi: https://doi.org/10.29173/iq968 3.4 confidence items there are search report elements/items that can provide added support for confidence in the reproducibility of searches, as well as for confidence in the use of those searches. the results in table 6 present kappa and report frequency data for these search elements based on our assessment of sr from psychology and the campbell library for the years 2009-2016, as well as for psychology sr for 2017. our kappa results suggest challenges for assessing some of these confidence elements with our items; and, again, more training could give better agreement results. additionally, in comparison to our item 2, other studies using the related but more general amstar item for ’duplicate study https://doi.org/10.29173/iq968 18/26 fehrmann, paul & mamolen, megan (2020) methods reporting that supports reader confidence for systematic reviews in psychology: assessing the reproducibility of electronic searches and first-level screening decisions, iassist quarterly 44(1-2), pp. 1-26. doi: https://doi.org/10.29173/iq968 selection and data extraction’ found kappas of .93 (popovich et al., 2012) and .77 (pieper et al., 2015). relative to our third research question, our results also show low frequencies for clear reporting. frequency of reporting for this item 1 was low, although reporting such numbers has been recommended by many including the prisma group (liberati et al., 2009). reporting information for our item 2 is encouraged, or even expected, for some projects (e.g. major health, psychology, or policy topics). this recommendation is seen in the iom’s finding what works in health care: standards for systematic reviews (committee on standards for systematic reviews of comparative effectiveness research, board on health care services, & institute of medicine, 2011), and is reiterated by the cochrane collaboration in the mecir standards for the reporting of new reviews of interventions (mecir manual, n.d.).the reporting of inter-rater agreement at the point of assessing the title and abstract (item 3) also is encouraged in the prisma guidelines for study selection. our findings show some of this reporting for selection at the point of screening title and abstracts. reporting for our item 4, the identification of items chosen or included at the title and abstract screening was infrequent, though it was evident in some sr. while space considerations could account for that finding, this would be information that supports readers who wish to be confident that they are reproducing the initial selection choices made as title and abstracts are screened. information for each of these confidence items is easily reported, and, going forward, such reporting will support confidence in conclusions of sr. that confidence in sr conclusions is important for all readers including other researchers, clinicians, and those who develop policies or share sr research information with the public. summary and conclusions recent discussions and research in psychology show a significant emphasis on reproducibility. concerns pertain to methods as well as results, and this paper contributes to discussions about the methods that are used for systematic reviews. we specifically examined the reporting of electronic searches used for sr in psychology. such reports are key for determining the reproducibility of electronic searches. confidence in the reproducibility of electronic searches can also impact the confidence that readers have in the overall results or conclusions of systematic reviews. in this paper we first discuss systematic reviews, reproducibility, electronic searches, transparent reporting, and the increased use of sr in psychology. based on evidence-based recommendations, we developed and used 12 checklist items to evaluate electronic search reporting that supports reproducibility. item kappa results ranged from fair to almost perfect. then, mirroring comparisons of reporting in cochrane sr to that found in non-cochrane sr, using those checklist items we compared reports in sr published in the psychology literature to those sr from the campbell library. reporting of basic recommended electronic search step information that supports reproducibility was seen significantly less in psychology sr. additionallly, we found that 90% of the 68 psychology sr that we assessed did not provide what we defined as a copy of a full search strategy for any of the electronic search resources used. moreover, https://doi.org/10.29173/iq968 19/26 fehrmann, paul & mamolen, megan (2020) methods reporting that supports reader confidence for systematic reviews in psychology: assessing the reproducibility of electronic searches and first-level screening decisions, iassist quarterly 44(1-2), pp. 1-26. doi: https://doi.org/10.29173/iq968 assuming that higher levels of reproducibility are supported as more of a set of recommended reporting criteria are met, we used a set of checklist items to represent a ’prisma’ type of recommended reporting. we found that only one of the 25 randomly selected psychology sr from 2009-2012 reported recommended information for all items in the set, and none of the 25 psychology sr from 2014-2016 did so. furthermore, although the set of 18 psych sr from 2017 was used because each referred to the prisma statement, only 3 reported information for all in our ’prisma set’ of search report elements. we also looked at reporting for we view as ‘confidence items’ that can be a part of reporting of electronic searches in sr. items covered reporting for the number of hits for every electronic search, the number of hits for all electronic searches combined, the use of two or more researches for independent title/abstract screening, the inter-rater agreement for two or more researches for independent title/abstract screening, and the identification of items selected for possible use at the end of title/abstract screening. about half of the 68 psych sr we assessed reported the number of hits combined across all electronic searches; and reporting for the other search items was very low. based on our findings, we had six general conclusions. 1. electronic search reporting in published sr in psychology shows that improvements should be made that support confidence in the reproducibility of electronic searches used. 2. as shown in our assessment data for the sr from the campbell collaboration, it does seem possible to report what we called our prisma set for every electronic resource used. 3. reporting for what we called a prisma set should be seen more in sr published in psychology. 4. it does seem that reporting in sr could include more reporting for all of what we called confidence items. this information supports reader confidence not only in the searches run but also in the use of search results. 5. findings from the current study could serve as a baseline for this kind of reporting in sr that are published in psychology. 6. the research checklist we developed and used had inter-rater agreement that suggests it might serve as a resource for those concerned with the reproducibility of electronic searches. that checklist, or some version, might be used for evaluating electronic searches, for research on the items themselves and/or for research on the reproducibility of electronics searches in sr. improving the reporting of electronic search strategies so that readers can be confident in the reproducibility of such searches can be challenging. however, we believe the improvements that we describe for sr are possible. and, going forward, improvements in the reporting of electronic searches used for sr can serve as ’clear signals’ that provide a stronger basis for confidence in the results and conclusions of sr in psychology. https://doi.org/10.29173/iq968 20/26 fehrmann, paul & mamolen, megan (2020) methods reporting that supports reader confidence for systematic reviews in psychology: assessing the reproducibility of electronic searches and first-level screening decisions, iassist quarterly 44(1-2), pp. 1-26. doi: https://doi.org/10.29173/iq968 references agency for healthcare research and quality. (2014). methods guide for effectiveness and comparative effectiveness reviews. agency for healthcare research and quality, retrieved from https://effectivehealthcare.ahrq.gov/sites/default/files/pdf/cer-methods-guide_overview.pdf ali, n. b., & usman, m. (2018). reliability of search in systematic reviews: towards a quality assessment framework for the automated-search strategy. information and software technology, 99, 133–147. https://doi.org/10.1016/j.infsof.2018.02.002 apa databases methodology field values. (n.d.). retrieved october 30, 2019, from https://www.apa.org website: https://www.apa.org/pubs/databases/training/method-values apa publications and communications board working group on journal article reporting standards. (2008). reporting standards for research in psychology: why do we need them? what might they be? american psychologist, 63(9), 839–851. https://doi.org/10.1037/0003-066x.63.9.839 aromataris, e., & munn, z. (2019). chapter 1: jbi systematic reviews jbi reviewer’s manual jbi global wiki. in jbi reviewer’s manual. retrieved june 29, 2020, from https://wiki.joannabriggs.org/display/manual/chapter+1%3a+jbi+systematic+reviews atkinson, k. m., koenka, a. c., sanchez, c. e., moshontz, h., & cooper, h. (2015). reporting standards for literature searches and report inclusion criteria: making research syntheses more transparent and easy to replicate. research synthesis methods, 6(1), 87–95. https://doi.org/10.1002/jrsm.1127 booth, a. (2006). brimful of starlite : toward standards for reporting literature searches. journal of the medical library association, (4), 421. campbell, m., katikireddi, s. v., sowden, a., & thomson, h. (2019). lack of transparency in reporting narrative synthesis of quantitative data: a methodological assessment of systematic reviews. journal of clinical epidemiology, 105, 1–9. https://doi.org/10.1016/j.jclinepi.2018.08.019 centre for reviews and dissemination. (2009). systematic reviews. crd’s guidance for undertaking reviews in health care. retrieved june 29, 2020, from https://www.york.ac.uk/crd/guidance/ committee on reproducibility and replicability in science, board on behavioral, cognitive, and sensory sciences, committee on national statistics, division of behavioral and social sciences and education, nuclear and radiation studies board, division on earth and life studies, … national academies of sciences, engineering, and medicine. (2019). confidence. in reproducibility and replicability in science. https://doi.org/10.17226/25303 committee on standards for systematic reviews of comparative effectiveness research, board on health care services, & institute of medicine. (2011). finding what works in health care: standards for systematic reviews (j. eden, l. levit, a. berg, & s. morton, eds.). national academies press. https://doi.org/10.17226/13059 https://doi.org/10.29173/iq968 https://effectivehealthcare.ahrq.gov/sites/default/files/pdf/cer-methods-guide_overview.pdf https://doi.org/10.1016/j.infsof.2018.02.002 https://www.apa.org/pubs/databases/training/method-values https://doi.org/10.1037/0003-066x.63.9.839 https://wiki.joannabriggs.org/display/manual/chapter+1%3a+jbi+systematic+reviews https://doi.org/10.1002/jrsm.1127 https://doi.org/10.1016/j.jclinepi.2018.08.019 https://www.york.ac.uk/crd/guidance/ https://doi.org/10.17226/25303 https://doi.org/10.17226/13059 21/26 fehrmann, paul & mamolen, megan (2020) methods reporting that supports reader confidence for systematic reviews in psychology: assessing the reproducibility of electronic searches and first-level screening decisions, iassist quarterly 44(1-2), pp. 1-26. doi: https://doi.org/10.29173/iq968 cooper, h. m. (2017). research synthesis and meta-analysis: a step-by-step approach (fifth edition). thousand oaks, california: sage publications, inc. cooper, h., & vandenbos, g. r. (2013). archives of scientific psychology: a new journal for a new era. archives of scientific psychology, 1(1), 1–6. https://doi.org/10.1037/arc0000001 delaney, a., & tamás, p. a. (2018). searching for evidence or approval? a commentary on database search in systematic reviews and alternative information retrieval methodologies. research synthesis methods, 9(1), 124–131. https://doi.org/10.1002/jrsm.1282 faggion, c. m. (2019). should a systematic review be tested for reproducibility before its publication? journal of clinical epidemiology, 110, 96. https://doi.org/10.1016/j.jclinepi.2019.02.008 fehrmann, p., & thomas, j. (2011). comprehensive computer searches and reporting in systematic reviews: computer searches and reporting in reviews. research synthesis methods, 2(1), 15–32. https://doi.org/10.1002/jrsm.31 gensby, u., lund, t., kowalski, k., saidj, m., jørgensen, a. k., filges, t., … labriola, m. (2012). workplace disability management programs promoting return to work: a systematic review. campbell systematic reviews, 8(1). https://doi.org/10.4073/csr.2012.17 gilmore, r. o., diaz, m. t., wyble, b. a., & yarkoni, t. (2017). progress toward openness, transparency, and reproducibility in cognitive neuroscience: openness in cognitive neuroscience. annals of the new york academy of sciences, 1396(1), 5–18. https://doi.org/10.1111/nyas.13325 golder, s., loke, y. k., & zorzela, l. (2013). some improvements are apparent in identifying adverse effects in systematic reviews from 1994 to 2011. journal of clinical epidemiology, 66(3), 253–260. https://doi.org/10.1016/j.jclinepi.2012.09.013 golder, s., loke, y., & mcintosh, h. m. (2008). poor reporting and inadequate searches were apparent in systematic reviews of adverse effects. journal of clinical epidemiology, 61(5), 440–448. https://doi.org/10.1016/j.jclinepi.2007.06.005 goodman, s. n., fanelli, d., & ioannidis, j. p. a. (2016). what does research reproducibility mean? science translational medicine, 8(341), 341ps12-341ps12. https://doi.org/10.1126/scitranslmed.aaf5027 gough, d., thomas, j., & oliver, s. (2012). clarifying differences between review designs and methods. systematic reviews, 1(1), 28. https://doi.org/10.1186/2046-4053-1-28 hardwicke, t. e., serghiou, s., janiaud, p., danchev, v., crüwell, s., goodman, s. n., & ioannidis, j. p. a. (2020). calibrating the scientific ecosystem through meta-research. annual review of statistics and its application, 7(1), null. https://doi.org/10.1146/annurev-statistics-031219-041104 higgins, j. p. t., & green, s. (2011). cochrane handbook for systematic reviews of interventions. john wiley & sons. https://doi.org/10.29173/iq968 https://doi.org/10.1037/arc0000001 https://doi.org/10.1002/jrsm.1282 https://doi.org/10.1016/j.jclinepi.2019.02.008 https://doi.org/10.1002/jrsm.31 https://doi.org/10.4073/csr.2012.17 https://doi.org/10.1111/nyas.13325 https://doi.org/10.1016/j.jclinepi.2012.09.013 https://doi.org/10.1016/j.jclinepi.2007.06.005 https://doi.org/10.1126/scitranslmed.aaf5027 https://doi.org/10.1186/2046-4053-1-28 https://doi.org/10.1146/annurev-statistics-031219-041104 22/26 fehrmann, paul & mamolen, megan (2020) methods reporting that supports reader confidence for systematic reviews in psychology: assessing the reproducibility of electronic searches and first-level screening decisions, iassist quarterly 44(1-2), pp. 1-26. doi: https://doi.org/10.29173/iq968 jamieson, k. h., mcnutt, m., kiermer, v., & sever, r. (2019). signaling the trustworthiness of science. proceedings of the national academy of sciences, 116(39), 19231–19236. https://doi.org/10.1073/pnas.1913039116 kisely, s., chang, a., crowe, j., galletly, c., jenkins, p., loi, s., … macfarlane, s. (2015). getting started in research: systematic reviews and meta-analyses. australasian psychiatry, 23(1), 16–21. https://doi.org/10.1177/1039856214562077 koffel, j. b., & rethlefsen, m. l. (2016). reproducibility of search strategies is poor in systematic reviews published in high-impact pediatrics, cardiology and surgery journals: a cross-sectional study. plos one, 11(9), e0163309. https://doi.org/10.1371/journal.pone.0163309 kugley, s., wade, a., thomas, j., mahood, q., jørgensen, a. k., hammerstrøm, k., & sathe, n. (2017). searching for studies: a guide to information retrieval for campbell systematic reviews. campbell systematic reviews, 13(1), 1–73. https://doi.org/10.4073/cmg.2016.1 landis, j. r., & koch, g. g. (1977). the measurement of observer agreement for categorical data. biometrics, 33(1), 159–174. https://doi.org/10.2307/2529310 lebel, e. p., borsboom, d., giner-sorolla, r., hasselman, f., peters, k. r., ratliff, k. a., & smith, c. t. (2013). psychdisclosure.org: grassroots support for reforming reporting standards in psychology. perspectives on psychological science, 8(4), 424–432. https://doi.org/10.1177/1745691613491437 leclercq, v., beaudart, c., ajamieh, s., rabenda, v., tirelli, e., & bruyère, o. (2019). meta-analyses indexed in psycinfo had a better completeness of reporting when they mention prisma. journal of clinical epidemiology, 115, 46–54. https://doi.org/10.1016/j.jclinepi.2019.06.014 lefebvre, c., glanville, j., briscoe, s., littlewood, a., marshall, c., metzendorf, m.-i., … wieland, l. s. (2019). searching for and selecting studies. in cochrane handbook for systematic reviews of interventions (pp. 67–107). https://doi.org/10.1002/9781119536604.ch4 lefebvre, c., manheimer, e., & glanville, j. (2011). chapter 6: searching for studies. in j. higgins & s. green (eds.), cochrane handbook for systematic reviews of interventions version 5.1.0. retrieved june 29, 2020, from www.handbook.cochrane.org liberati, a., altman, d. g., tetzlaff, j., mulrow, c., gøtzsche, p. c., ioannidis, j. p. a., … moher, d. (2009). the prisma statement for reporting systematic reviews and meta-analyses of studies that evaluate health care interventions: explanation and elaboration. plos medicine, 6(7), e1000100. https://doi.org/10.1371/journal.pmed.1000100 maggio, l. a., tannery, n. h., & kanter, s. l. (2011). reproducibility of literature search reporting in medical education reviews: academic medicine, 86(8), 1049–1054. https://doi.org/10.1097/acm.0b013e31822221e7 mcdonagh, m., peterson, k., raina, p., chang, s., & shekelle, p. (2008). avoiding bias in selecting studies. in methods guide for effectiveness and comparative effectiveness reviews. agency for https://doi.org/10.29173/iq968 https://doi.org/10.1073/pnas.1913039116 https://doi.org/10.1177/1039856214562077 https://doi.org/10.1371/journal.pone.0163309 https://doi.org/10.4073/cmg.2016.1 https://doi.org/10.2307/2529310 https://doi.org/10.1177/1745691613491437 https://doi.org/10.1016/j.jclinepi.2019.06.014 https://doi.org/10.1002/9781119536604.ch4 https://www.handbook.cochrane.org/ https://doi.org/10.1371/journal.pmed.1000100 https://doi.org/10.1097/acm.0b013e31822221e7 23/26 fehrmann, paul & mamolen, megan (2020) methods reporting that supports reader confidence for systematic reviews in psychology: assessing the reproducibility of electronic searches and first-level screening decisions, iassist quarterly 44(1-2), pp. 1-26. doi: https://doi.org/10.29173/iq968 healthcare research and quality (us). retrieved june 29, 2020, from http://www.ncbi.nlm.nih.gov/books/nbk126701/ mcgowan, j., sampson, m., salzwedel, d. m., cogo, e., foerster, v., & lefebvre, c. (2016). press peer review of electronic search strategies: 2015 guideline statement. journal of clinical epidemiology, 75, 40–46. https://doi.org/10.1016/j.jclinepi.2016.01.021 meader, n., king, k., llewellyn, a., norman, g., brown, j., rodgers, m., … stewart, g. (2014). a checklist designed to aid consistency and reproducibility of grade assessments: development and pilot validation. systematic reviews, 3(1), 82. https://doi.org/10.1186/2046-4053-3-82 meert, mlis, d., torabi, mlis, n., & costella, dds, msc, mlis, j. (2017). impact of librarians on reporting of the literature searching component of pediatric systematic reviews. journal of the medical library association, 104(4), 267–277. https://doi.org/10.5195/jmla.2016.139 mecir manual. (n.d.). retrieved january 2, 2020, from https://community.cochrane.org/mecirmanual moher, d., tetzlaff, j., tricco, a. c., sampson, m., & altman, d. g. (2007). epidemiology and reporting characteristics of systematic reviews. plos medicine, 4(3), e78. https://doi.org/10.1371/journal.pmed.0040078 mullins, m. m., deluca, j. b., crepaz, n., & lyles, c. m. (2014). reporting quality of search methods in systematic reviews of hiv behavioral interventions (2000-2010): are the searches clearly explained, systematic and reproducible? research synthesis methods, 5(2), 116–130. https://doi.org/10.1002/jrsm.1098 ng, c., & benedetto, u. (2016). evidence hierarchy. in g. biondi-zoccai (ed.), umbrella reviews (pp. 11–19). https://doi.org/10.1007/978-3-319-25655-9_2 niederstadt, c., & droste, s. (2010). reporting and presenting information retrieval processes: the need for optimizing common practice in health technology assessment. international journal of technology assessment in health care, 26(4), 450–457. https://doi.org/10.1017/s0266462310001066 novotney, a. (2014). reproducing results. monitor on psychology, 45(8). retrieved june 29, 2020, from https://www.apa.org/monitor/2014/09/results open science collaboration. (2015). estimating the reproducibility of psychological science. science, 349(6251), aac4716–aac4716. https://doi.org/10.1126/science.aac4716 page, m. j., mckenzie, j. e., bossuyt, p. m., boutron, i., hoffmann, t., mulrow, c. d., … moher, d. (2020). mapping of reporting guidance for systematic reviews and meta-analyses generated a comprehensive item bank for future reporting guidelines. journal of clinical epidemiology, 118, 60– 68. https://doi.org/10.1016/j.jclinepi.2019.11.010 https://doi.org/10.29173/iq968 http://www.ncbi.nlm.nih.gov/books/nbk126701/ https://doi.org/10.1016/j.jclinepi.2016.01.021 https://doi.org/10.1186/2046-4053-3-82 https://doi.org/10.5195/jmla.2016.139 https://community.cochrane.org/mecir-manual https://community.cochrane.org/mecir-manual https://doi.org/10.1371/journal.pmed.0040078 https://doi.org/10.1002/jrsm.1098 https://doi.org/10.1007/978-3-319-25655-9_2 https://doi.org/10.1017/s0266462310001066 https://www.apa.org/monitor/2014/09/results https://doi.org/10.1126/science.aac4716 https://doi.org/10.1016/j.jclinepi.2019.11.010 24/26 fehrmann, paul & mamolen, megan (2020) methods reporting that supports reader confidence for systematic reviews in psychology: assessing the reproducibility of electronic searches and first-level screening decisions, iassist quarterly 44(1-2), pp. 1-26. doi: https://doi.org/10.29173/iq968 pashler, h., & wagenmakers, e. (2012). editors’ introduction to the special section on replicability in psychological science: a crisis of confidence? perspectives on psychological science, 7(6), 528–530. https://doi.org/10.1177/1745691612465253 paul, m., & leibovici, l. (2014). systematic review or meta-analysis? their place in the evidence hierarchy. clinical microbiology and infection, 20(2), 97–100. https://doi.org/10.1111/14690691.12489 pieper, d., buechter, r. b., li, l., prediger, b., & eikermann, m. (2015). systematic review found amstar, but not r(evised)-amstar, to have good measurement properties. journal of clinical epidemiology, 68(5), 574–583. https://doi.org/10.1016/j.jclinepi.2014.12.009 popovich, i., windsor, b., jordan, v., showell, m., shea, b., & farquhar, c. m. (2012). methodological quality of systematic reviews in subfertility: a comparison of two different approaches. plos one, 7(12), e50403. https://doi.org/10.1371/journal.pone.0050403 prisma. (n.d.). retrieved october 31, 2019, from http://www.prisma-statement.org/ psycarticles. (n.d.). accessed june 29, 2020, from https://www.apa.org/pubs/databases/psycarticles/index qualtrics xm experience management software. (n.d.). retrieved october 31, 2019, from qualtrics website: https://www.qualtrics.com/ rader, t., mann, m., stansfield, c., cooper, c., & sampson, m. (2014). methods for documenting systematic review searches: a discussion of common issues. research synthesis methods, 5(2), 98– 115. https://doi.org/10.1002/jrsm.1097 sampson, m., & mcgowan, j. (2006). errors in search strategies were identified by type and frequency. journal of clinical epidemiology, 59(10), 1057.e1-1057.e9. https://doi.org/10.1016/j.jclinepi.2006.01.007 sampson, m., mcgowan, j., tetzlaff, j., cogo, e., & moher, d. (2008). no consensus exists on search reporting methods for systematic reviews. journal of clinical epidemiology, 61(8), 748–754. https://doi.org/10.1016/j.jclinepi.2007.10.009 sampson, m., mcgowan, j., cogo, e., grimshaw, j., moher, d., & lefebvre, c. (2009). an evidencebased practice guideline for the peer review of electronic search strategies. journal of clinical epidemiology, 62(9), 944–952. https://doi.org/10.1016/j.jclinepi.2008.10.012 sayre, f., & riegelman, a. (2018). the reproducibility crisis and academic libraries. college & research libraries, 79(1), 2–9. https://doi.org/10.5860/crl.79.1.2 schalken, n., & rietbergen, c. (2017). the reporting quality of systematic reviews and metaanalyses in industrial and organizational psychology: a systematic review. frontiers in psychology, 8, 1395. https://doi.org/10.3389/fpsyg.2017.01395 https://doi.org/10.29173/iq968 https://doi.org/10.1177/1745691612465253 https://doi.org/10.1111/1469-0691.12489 https://doi.org/10.1111/1469-0691.12489 https://doi.org/10.1016/j.jclinepi.2014.12.009 https://doi.org/10.1371/journal.pone.0050403 http://www.prisma-statement.org/ https://www.apa.org/pubs/databases/psycarticles/index https://www.qualtrics.com/ https://doi.org/10.1002/jrsm.1097 https://doi.org/10.1016/j.jclinepi.2006.01.007 https://doi.org/10.1016/j.jclinepi.2007.10.009 https://doi.org/10.1016/j.jclinepi.2008.10.012 https://doi.org/10.5860/crl.79.1.2 https://doi.org/10.3389/fpsyg.2017.01395 25/26 fehrmann, paul & mamolen, megan (2020) methods reporting that supports reader confidence for systematic reviews in psychology: assessing the reproducibility of electronic searches and first-level screening decisions, iassist quarterly 44(1-2), pp. 1-26. doi: https://doi.org/10.29173/iq968 shea, b., moher, d., graham, i., pham, b., & tugwell, p. (2002). a comparison of the quality of cochrane reviews and systematic reviews published in paper-based journals. evaluation & the health professions, 25(1), 116–129. shea, b. j., hamel, c., wells, g. a., bouter, l. m., kristjansson, e., grimshaw, j., … boers, m. (2009). amstar is a reliable and valid measurement tool to assess the methodological quality of systematic reviews. journal of clinical epidemiology, 62(10), 1013–1020. https://doi.org/10.1016/j.jclinepi.2008.10.009 tong, a., flemming, k., mcinnes, e., oliver, s., & craig, j. (2012). enhancing transparency in reporting the synthesis of qualitative research: entreq. bmc medical research methodology, 12(1), 181. https://doi.org/10.1186/1471-2288-12-181 tunis, a. s., mcinnes, m. d. f., hanna, r., & esmail, k. (2013). association of study quality with completeness of reporting: have completeness of reporting and quality of systematic reviews and metaanalyses in major radiology journals changed since publication of the prisma statement? european journal of radiology, 269(2), 413–426. vazire, s. (2017). quality uncertainty erodes trust in science. collabra: psychology, 3(1), 1. https://doi.org/10.1525/collabra.74 yong, e. (2013). psychologists strike a blow for reproducibility. nature, nature.2013.14232. https://doi.org/10.1038/nature.2013.14232 yoshii, a., plaut, d. a., mcgraw, k. a., anderson, m. j., & wellik, k. e. (2009). analysis of the reporting of search strategies in cochrane systematic reviews. journal of the medical library association : jmla, 97(1), 21–29. https://doi.org/10.3163/1536-5050.97.1.004 williams, m., bagwell, j., & nahm zozus, m. (2017). data management plans: the missing perspective. journal of biomedical informatics, 71, 130–142. https://doi.org/10.1016/j.jbi.2017.05.004 endnotes 1 paul fehrman, ma, mls, is a research and instruction services librarian with the university libraries at kent state university, pfehrman@kent.edu 2 megan mamolen, phd, mlis, is phd, mlis is a reference, instruction, e-resources librarian at lakeland community college, mmamolen1@lakelandcc.edu 3 with respect to transparent reporting of searches to identify resources with data for sr, recent related discussion of data management planning also argues for clear documentation of such ’upstream activites that determine data quality’ (williams et al., 2017). https://doi.org/10.29173/iq968 https://doi.org/10.1016/j.jclinepi.2008.10.009 https://doi.org/10.1186/1471-2288-12-181 https://doi.org/10.1525/collabra.74 https://doi.org/10.1038/nature.2013.14232 https://doi.org/10.3163/1536-5050.97.1.004 https://doi.org/10.1016/j.jbi.2017.05.004 mailto:pfehrman@kent.edu mailto:mmamolen1@lakelandcc.edu 26/26 fehrmann, paul & mamolen, megan (2020) methods reporting that supports reader confidence for systematic reviews in psychology: assessing the reproducibility of electronic searches and first-level screening decisions, iassist quarterly 44(1-2), pp. 1-26. doi: https://doi.org/10.29173/iq968 4 we want to express appreciation to matthew cox, mlis. his contribution was significant for the work on reproducibility presented at the annual campbell collaboration colloquium, held may 21-23, 2013 in chicago. that poster is on the osf site noted just below. 5 osf site for this paper https://osf.io/g6x4k this site includes supplementary materials noted in the paper. 6 in the spring of 2016, we learned of work on ‘prisma-search’. the goal of that project has been to develop an extension to the prisma statement to guide the reporting of components critical to a reproducible search. information has been available in ‘reporting guidelines under development’ on the equator network website (www.equator-network.org/). moreover, in the spring of 2019, the authors of prisma-s shared that checklist and explanation documents were available for review. information for prisma-s has been available in ‘reporting guidelines under development’ on the equator network website and more recently on the open science framework https://doi.org/10.17605/osf.io/ygn9w. we anticipate that when the extension work is complete it will provide additional support for our selection of items to assess electronic search reproducibility. 7 the four appendixes noted in this paper, as well as additional background information about the checklists used for this study, are available on the paper’s osf site noted above. 8 as an example of a paper using the apa guidance see this paper and supplemental file published in the archives of scientific psychology. youngstrom, e. a., genzlinger, j. e., egerton, g. a., & van meter, a. r. (2015). multivariate meta-analysis of the discriminative validity of caregiver, youth, and teacher rating scales for pediatric bipolar disorder: mother knows best about mania. archives of scientific psychology, 3(1), 112–137. https://doi.org/10.1037/arc0000024 9 https://amstar.ca/ 10 https://www.cadth.ca/resources/finding-evidence/press 11 https://doi.org/10.17605/osf.io/ygn9w 12 of the 14 campbell sr in item 1. a., one did provide a full copy of at least one electronic search, but not for all of the electronic resources used (see 11. in appendix 2). 13 https://osf.io https://doi.org/10.29173/iq968 https://osf.io/g6x4k https://doi.org/10.17605/osf.io/ygn9w https://doi.org/10.1037/arc0000024 https://amstar.ca/ https://www.cadth.ca/resources/finding-evidence/press https://doi.org/10.17605/osf.io/ygn9w https://osf.io/ vol30-4.indd iassist quarterly winter 2006 by by jinfang niu* reward and punishment mechanism for research data sharing abstract many funding agencies require grantees to deposit their data into an archive after they finish their research projects. the archive processes and disseminates the data for public use. these deposited data sets are public goods that benefit users and society. however, under voluntary contribution, public goods tend to be under-provided. for normal public goods, the contributors benefit from their own contributions as much as free-riders. contributors are not harmed by their contributions. in the data sharing case, data producers make efforts to prepare the data for deposit, but the benefit of the data preparation largely goes to secondary users. in addition, data producers are at risk of being harmed by the misuse and misinterpretation of data by unqualified users, or by being charged with misconduct. that makes freeriding even more attractive. to motivate data producers to prepare and share data, there must be some incentive mechanisms. in this paper, i built a simple mathematic model to analyze the effects of punishment and reward. hopefully it will help policy makers decide on incentive mechanisms for data sharing. background many funding agencies require grantees to deposit data into an archive after they finish their research projects, such as the national institute of justice (nij), the national institutes of health (nih) in the united states, and the medical research council (mrc) and the economic and social research council (esrc) in the united kingdom. the archive processes and disseminates data for public use. data sharing benefits society in many ways. it saves funding and avoids repeated data collecting efforts, allows the verification and replication of research findings, facilitates scientific openness, deters scientific misconduct, and supports communication and progress before deposit, data depositors need to prepare their data according to the requirements of data archives. the purpose of the preparation is to help secondary use of data and protect the privacy of human research subjects. data preparation includes three kinds of work: preparing data, creating documentation and processing confidential information. data preparation includes checking the integrity1 and consistency of data, careful naming of variables and choice of variable labels that will be easy for secondary users to understand, organizing the variables such as grouping them to enable secondary analysts to get an overview of the data quickly, etc. (icpsr, 2005). data documentation provides metadata about the data sets and research projects, such as the principal investigator of the project, when and where the data were collected, the methodology and procedures used to collect the data, details about codes, definitions of variables, frequencies, and the like (nih, 2003). even data collection instruments, such as questionnaires and interview guides are required parts of documentation. documentation is indispensable for the searching, managing, preserving and re-using of data. in other words, without adequate documentation, secondary users of a data set will not be able to find the data, nor will they be able to interpret and analyze the data. as a result, the goal of data sharing will not be achieved. in addition, insufficient documentation might lead to the misuse of data or incorrect conclusions. to protect confidential information in data, all direct identifiers, such as names, addresses, telephone numbers, and social security numbers, have to be removed. in addition, indirect identifiers and other information that could lead to “deductive disclosure” of participants’ identities should also be removed or processed before the data are made public. data preparation involves a lot of work, and a fair amount of it is done only for secondary users. for example, data producers do not have to process the confidential information if they keep the data for their own use. many data producers do document data for their own use. however, documentation created for the producer’s own use are informal and biased toward short-term needs. to share with others, data producers have to take extra effort to shape the “public face” of their documentation (markus, 2001). in addition, data producers should take the main responsibility for preparing data. zimmerman (2003) found that both secondary data users and data managers (intermediaries or data archivists) agree that no one understands the data better than the scientists who gathered them, and that it is the data producers who must document data. 12 iassist quarterly winter 2006 when publicly funded research data are disseminated to the public through the website of a data archive, no one is excluded from using them, and one individual’s use of the data and documentation does not reduce the amount available for other people. those data sets are public goods by definition (mas-colell, et al., 1995). since no one is excluded from the online data archive whether or not they have deposited data, as with other public goods people have strong incentives to free ride in preparing and depositing data. from the game theory perspective, free-riding in voluntary contribution to public goods tends to be the dominant strategy in a non-cooperative game (bergstrom, et al., 1986; cornes & sandler, 1986). for normal public goods, the contributors benefit from their own contributions in the same way as free-riders, and they are not harmed by their contributions. for example, once a bridge is built, the contributors and free riders get the same benefit. however, in the data sharing case, the depositor of a data set does not benefit from the data he deposited in the same way as secondary users. a data depositor is unlikely to use his own data deposited into a data archive, either because he has used it before, or because he keeps his own data for future use. people mostly benefit from others’ contributions. the benefit of depositors’ effort in data preparation largely goes to the users. in addition, data producers are at risk of being harmed by the misuse and misinterpretation of data by unqualified users, or by being charged with misconduct. that makes free riding even more attractive. to change this situation, there must be some incentive mechanisms to motivate researchers to prepare and deposit data. the incentive mechanisms that some funding agencies have implemented focus on punishment for non-compliance with data sharing requirements, and pay less attention to rewards. according to the policy of the national institutes of health (nih), in the case of noncompliance (depending on its severity and duration), nih can take various actions to protect the federal government’s interests. in some instances, for example, the nih may make data sharing an explicit term and condition of subsequent awards (nih, 2003). under the policy of the esrc, “the final payment of an award will be withheld until data has been deposited in accordance with the requirements. the requirements of the data sharing policy are now a condition of esrc research funding.” (esrc, 2000). the data sharing policies of nih and esrc do not mandate that users cite the data they use, and they are against the idea that data producers require co-authorship as a condition for sharing the data. nih explicitly stated that they do not offer rewards for doing well in data-sharing. data from a survey2 of the grantees of a funding agency showed that some grantees expect rewards for data deposit. for example, one grantee said he would be more likely to deposit data if there were some sort of acknowledgment that he had deposited data, such as a certificate. some other grantees claimed that some sort of punishment would make them more likely to deposit data, for example, if data deposit were mandatory to receive new funding from nij, or a prerequisite for publishing a paper derived from the data. one grantee was strongly against a punishment mechanism. he said: “do you really want a system where archiving data prevents people from publishing or from doing new work? this would be a triumph of bureaucracy over common sense. if the funding agency becomes obsessed with bureaucratic requirements, they will drive away talented researchers.” i believe that either punishments or rewards would provide incentives for data producers to take more effort in data preparation. but when decide the punishment or reward mechanisms, the level of the punishment and reward should be carefully chosen. otherwise, unintended consequences might occur. to help illustrate this, i have built a very simple mathematical model the model three parties are involved in the model. they are the data producer, the data user and the funding agency. in reality, there are many data producers and users. to make the problem simple, i only consider one data producer and one data user. the data producer has total fund p. he chooses θ and e to spend on research and data preparation respectively (p = θ + e). he benefits ω(θ) from spending θ on research. i assume that ω(θ) is concave, differentiable and ω(θ) >0, meaning that the more the data producer spends on research, the more he will benefit, but the increase rate of the benefit decreases3. see figure 1 for the graph of ω(θ). i also assume that the data producer always tries to maximize his benefit when making decisions. i analyzed and compared the social benefits generated from three scenarios: no reward & no punishment, punishment only and reward only. social benefit is defined as the sum of the benefit gained by the data producer, the user and the funding agency in each scenario. figure 1 iassist quarterly winter 2006 13 scenario 1: no reward & no punishment in the no reward & no punishment scenario, the data producer does not benefit from spending effort on data preparation, and he loses nothing if does not spend any effort on data preparation. to state this formally, the utility function of the data producer is ω(θ), 0 < θ ≤p. since the more the data depositor spends on research, the more he benefits, the data depositor would spend e = 0 on data preparation to maximize his utility. his maximized utility is ω(p). for the user, since the data depositor did not spend any effort on data preparation for deposit there is no data to use, so the user’s utility is 0. the social benefit is the sum of the benefit of the depositor and the user: ω(p). scenario 2: punishment only in this scenario, the data producer will be punished if the effort he spends on data preparation is lower than a threshold. to state this formally, the benefit function is: if e ≥ e’, the data depositor’s benefit is ω(θ), if e < e’, the data depositor’s benefit is: ω(θ) – f, (f > 0). f is a fine that the data depositor has to pay to the funding agency if he is punished. e’ is the threshold for punishment. in this case, to maximize his benefit, the data producer needs to compare the highest possible benefit he could get if he passes the threshold versus if he does not. mathematically, he needs to maximize a benefit function of two parts, and pick the one that is larger. when e ≥ e’, the data depositor’s benefit is ω(θ) = ω(p-e). since ω'(θ) > 0, to maximize ω(θ), we need to minimize e, the smallest value of e is e’, and the maximized benefit of the data depositor is ω(p e’). when e < e’, the data depositor’s benefit is ω(θ) – f = ω(p-e) – f, again we need to minimize e to maximize the data depositor’s benefit. the smallest value of e is 0, so the maximized utility of the data depositor is ω(p) – f. now compare (p e’) and ω(p) – f. if ω(p e’) > ω(p) – f <=> f > ω(p) ω(p e’), the function is maximized at e = e’, which means that the data producer will benefit more by passing the threshold. so the data producer would choose to pass the threshold to avoid punishment. there are two explanations for this. first, if the threshold (e’) is fixed, this means that the punishment is severe enough (f is big enough) to make f > ω(p) ω(p e’). second, if the punishment level is fixed (keep f constant), this means the threshold is easy to meet (e’ is low), so the data depositor would like to meet the threshold to avoid the punishment. if ω(p e’) = ω(p) – f, the data depositor is indifferent between preparing data for deposit and getting punished. if ω(p e’) < ω(p) – f <=> f < ω(p) ω(p e’), the data producer benefits more from being punished than from preparing and depositing data. to maximize his benefit, the data depositor will choose to spend nothing on data preparation and be punished. there are two explanations for this. first, if the threshold is fixed, it means the punishment is not severe enough to deter non-compliance behaviors. second, if the punishment level is fixed, it means the threshold e’ is too costly to meet, so the data depositor would rather be punished than meet the threshold. based on the analysis above, we can see that the data depositor’s benefit in this scenario is max [ω(p e’), ω(p) – f]. for the user, when the data depositor would prefer to be punished than deposit data (ω(p e’) < ω(p) – f), there is no data to use. so the user’s benefit is 0. if max [ω(p e’), ω(p) – f] = ω(p) – f, the data depositor loses f, but the funding agency gets f4. the social benefit is the sum of the benefits of the data depositors, the users and the funding agency. so the social benefit = ω(p) – f +f +0 = ω(p). this is equal to the social benefit in the no punishment & no reward scenario. we can see that too weak a punishment or too high a standard for data preparation is not effective. the data depositor is punished, yet there is no gain in social benefits. this actually confirms the findings of existing literature that punishment is effective only when it is relatively harsh (trevino & ball, 1992). when the data depositor chooses to meet the threshold for data preparation, there is data available to use. but deposited data sets are not always used. in reality, there are various reasons. for example, a user does not use a data set because it does not fit his research purpose, or because the documentations of the data is not sufficient. here, i assume that the probability that the data is used depends on the fund that the data producer spends on data preparation. the more fund the data producer spends on data preparation, the more likely the data is used by the user. to state this formally, there is a probability π(e’) (π'( e) > 0, π( 0) =0) that the user will use the data. if he uses the data, the user will benefit v, so the user’s expected benefit of using data is v * π(e’). the social benefit is: ω(p e’) + v * π(e’). remember the social benefit in the no punishment & no reward is ω(p). so when ω(p) < ω(p e’) + v * π (e’), it means that an appropriate punishment and a carefully selected threshold causes higher social benefit than no punishment & no reward. when ω(p) > ω(p e’) + v * π(e’), it means the reverse. scenario 3: reward only in this scenario, when the deposited data is used, the producer of the data gets a reward r, and the user of the data set benefits v. the deposited data has a probability π(e) (π'( e) > 0, π (0) = 0) of being used. so the depositor’s expected benefit from the reward is r * π(e), the user’s expected benefit v*π(e). the producer’s total benefit is 14 iassist quarterly winter 2006 the expected benefit from the reward plus the benefit from doing research: ω(θ) + r * π(e) <=> ω(p-e) + r * π(e). to maximize the benefit of the data producer, we need to check the first order condition of ω(p-e) + r * π(e). if there is a value of “e” which makes [ω(p-e) + r * π(e)]’ =0 <=> r * π’(e) = ω’(p-e), then the benefit of a data producer is maximized when the marginal benefit of spending an additional amount of funding on research is equal to the product of reward and the marginal probability of being used. if the reward “r” is so big that no matter how small the marginal probability of the data being used (π'(e)) is, the product (r * π'(e)) is always greater than the marginal benefit of doing research [r * π'(e) > ω'(p-e)], it means that the function ω(p-e) + r * π(e) is monotonically increasing in the interval e ∈ [0, p]. in this case, utility is maximized when e = p, which means that to maximize his benefit, the data depositor should spend all funding available on data preparation. if the reward is so small that no matter how big the marginal probability of the data being used, the product (r * π'(e)) is always smaller than the marginal benefit of doing research [r * π'(e) < ω'(p-e)], the function ω(p-e) + r * π(e) is monotonically decreasing in the interval e ∈ [0,p]. here the benefit is maximized when e = 0, which means that to maximize his benefit, the data depositor should spend all funding on research. then the data producer will not deposit data and there is no data to use. in this case, the data producer is not rewarded because he did not deposit data. his benefit is ω(p). the user does not benfit because there is no data to use. the funding agency does not need to pay any reward to the producer. the social benefit = sum of benefit (producer, user and funding agency) = ω(p). it is exactly the same as the case with no reward & no punishment. in this case, the small reward is not effective at all. this confirms the findings of other literature that rewards should be of sufficient value, as rewards of insufficient value are the same as no reward at all (buhler, 1992). neither of these two cases are what we want. so we need to be careful not to make the reward too big or too small. suppose the data depositor’s benefit is maximized at e = e*, and the data user’s benefit is v*π (e*). the data depositor’s benefit is ω(p-e*) + r * π(e*), but the r * π(e*) is from the funding agency. in other words, the funding agency loses r * π(e*) in rewarding the data depositor. so social benefit = sum of the benefit of (depositor, user and funding agency) = ω(p-e*) + v*π( e*). ω(p) is the special point for ω(p-e*) + v*π( e*) where e = 0. e* is the maximized point, so ω(p) cannot be greater than ω(p-e*) + v*π( e*). so reward causes at least as much social benefit as no reward and no punishment. but we need to find an appropriate reward to make sure that the data depositor does not choose e = 0 or e = p. this simple model reveals the importance of choosing an appropriate level of punishment and reward, and an appropriate threshold for punishment. it does not deal with specific kinds of punishment or reward. for example, we do not consider whether we should punish non-compliers by withholding 10% of their final grant, or by factoring the quality of deposited data into consideration of future grants. i propose the following reward mechanism for data sharing policies: make the citation of data sets or the acknowledgement of data providers a mandatory requirement of publishing, the violation of which is treated in the same way as using but not citing published papers. treat the citation of data the same as the citation of published papers in the performance evaluation of researchers. as a complement for the model, here is a qualitative analysis of the punishment and reward mechanisms. effective punishments force all data producers without plausible excuses to prepare and deposit data, which would make all data collected under public funding accessible to the public. this gives users chances to verify the research findings of data producers, which would deter scientific fraud and misconduct. on the other hand, not all data sets will be used heavily (niu and hedstrom, 2007). under the punishment scenario, even if the data is very unlikely to be used in the future the data producer still needs to prepare and deposit data to avoid punishment. also, the archive needs to process, disseminate and preserve the data. enforcing uniform strong punishment on all data sets would cause the waste of resources. unlike the coercive and uniform nature of punishments, rewards are inductive and selective. rather than forcing researchers, rewards induce researchers to prepare and deposit data. researchers who expect their data to be used by other people will be motivated to do better in data preparation. data depositors who do not expect their data to be used will not prepare and deposit data, which may be a good choice. in this case, not all federally funded data sets will be made available to the public. the chance to verify some research is lost. also, data producers decide their effort in data preparation based on the expected future use of their data, which might be hard to anticipate. acknowledgments this research is funded by the national science foundation, award # iis 0456022 as part of the project entitled “incentives for data producers to create archiveready data sets.” i gratefully thank linhong chen’s help and instructions in building the model. references bergstrom, t. c., blume, l. & varian, h. r. (1986). on iassist quarterly winter 2006 15 the private provision of public goods. journal of public economics, february 1986, 29(1), pp. 25– 49. buhler, patricia m. (1992). the keys to shaping behavior. supervision v. 53 (jan. ‘92) pp. 18-20. cornes, r. c. & sandler, t. (1986). the theory of externalities, public goods, and club goods. cambridge: cambridge university press. esrc (economic and social research council), (2000). economic and social research council data policy. http:// www.esrcsocietytoday.ac.uk/esrcinfocentre/images/ datapolicy2000_tcm6-12051.pdf icpsr (inter-university consortium for political and social research). (2005). guide to social science data preparation and archiving: best practice throughout the data life cycle. http://www.icpsr.umich.edu/access/dataprep.pdf markus, m. l. (2001) toward a theory of knowledge reuse: type of knowledge reuse situations and factors in reuse success. journal of management information systems. 18(1), pp. 57-93. nih (national institute of justice). (2003). nih data sharing policy and implementation guidance. http:// grants2.nih.gov/grants/policy/data_sharing/data_sharing_ guidance.htm niu, j. & hedstrom, m. (2007). streamlining the “producer/archive” interface: mechanisms to reduce delays in ingest and release of social science data. digccurr 2007. april 18-20, chapel hill, nc, usa. trevino, l. k. & ball, g. a. (1992). “the social implications of punishing unethical behavior: observers’ cognitive and affective reactions.” journal of management, 18(4), pp. 751-768. zimmerman, a. (2003). data sharing and secondary use of scientific data: experiences of ecologists. unpublished dissertation, information and library studies, university of michigan, ann arbor. endnotes: 1. integrity means no wild codes or impossible values. for example, a respondent has 99 rather than 9 children. consistency means the variable values are consistent, for example, a respondent doesn’t work but reports earnings. 2. that survey was done in 2006 by the team of the nsf project “incentives for data producers to create archiveready data sets.” 3. p: the total fund available for the research project. θ: the amount of fund the data producer spent on research. e: the amount of fund the data producer spent on data preparation. ω'(θ): the first derivative of ω(θ). 4. the fine is paid by the data depositor to the funding agency. so when the data producer pays f to the funding agency, the data depositor loses f, and the funding agency gets f. *jinfang niu is a phd candidate at the school of information, university of michigan. she has a master degree in library science and 3 years working experiences in tsinghua university library, china. her research interests include data sharing, metadata, digital preservation and digital libraries. the paper was presented at the iassist 2007 conference in may in montreal, canada. contact: niujf@umich.edu. 4 iassist quarterly winter/spring 2010 editor’s notes welcome to this double issue of the iassist quarterly: volume 33 (number 4, 2009) and volume 34 (number 1, 2010). the iassist quarterly (iq) has now published several double issues focusing on a particular theme. often the issues have arisen from one or more sessions at an iassist conference or from workshops and other professional gatherings. furthermore, the focused issues are often used in workshops and serve as a welcome collection of relevant and related papers for presentation and discussion. the special issues of the iq normally have guest editors. on this occasion we are happy to have bobray bordelon from princeton university library as our guest editor. the special theme is: the subject content and how researchers use the data. bobray comments that we spend much effort concentrating on various technical aspects of data management and procurement; this special issue however looks into the actual subject matter of the data. as editor in chief of the iq i know that encouraging authors to turn their presentations into papers and pulling them together for publication can be a long-winded process with many repetitive steps. so much greater the contentment when the finished product surfaces. special thanks to bobray bordelon and also thanks to kristin partlo, amy west, walter w. giesbrecht, mary tao, kristi thompson, michele hayslett and lynda kellam for the compilation of this iq publication. articles for the iq are always very welcome. they can be papers from iassist conferences, other conferences, from local presentations or papers especially written for the iq. if you don't have anything to offer right now, then please prepare yourself for the next iassist conference and start planning for participation in a session there. chairing a conference session with the purpose of aggregating and integrating papers for a special issue iq is much appreciated as the information reaches many more people than the session participants and will be readily available on the iassist website at http://www.iassistdata.org. authors are very welcome to contact me via e-mail: kbr@ sam.sdu.dk, and should you be interested in compiling special issues for the iq as guest editor(s) i’d be delighted to hear from you. karsten boye rasmussen december 2010 . editor’s notes 1 1/8 manuel, kevin, orlandini, rosa and cooper, alexandra. (2022) who is counted? ethno-racial and indigenous identities in the census of canada, 1871-2021, iassist quarterly 46(4), pp. 1-8. doi: https://doi.org/10.29173/iq1016 who is counted? ethno-racial and indigenous identities in the census of canada, 1871-2021 kevin manuel1, rosa orlandini2, alexandra cooper3 abstract finding data on race, racialized populations, and anti-racism in canada can be a complex process when conducting research. one source of data is the census of canada which has been collecting sociodemographic data since 1871. however, the collection of racial, ethnic, or indigenous data has changed throughout the years and from census to census. in response to the need for more support in finding ethno-racial and indigenous data, the ontario council of university libraries’ ontario data community has created an online guide to provide guidance, in part, about the terminology used for indigenous and racialized identities over time in the census. in this article, the modifications to how ethno-racial origin questions have been asked, and the ongoing changes to sociocultural perceptions impacting the census are reviewed. keywords canada, census, data, ethnicity, indigenous, race, racialization introduction a common question that a data librarian or library professional at a university in canada will receive from a researcher is ‘where is the data on race?’ or ‘where is the data about indigenous peoples?’ the answer will often end up being a complex explanation about the history of the canadian census and its transformation overtime. unlike the united states census, the canadian census has not always asked respondents about racial identity but instead asked about ethnicity, origin, or country of birth. in the summer of 2020, a working group of data librarians and professionals from the ontario council of university libraries’ ontario data community volunteered to form a working group to examine the context of indigenous and racialized data over time in the census and other data sources in canada. their goal was to create an online research guide4 that could be used by information professionals and researchers who are seeking out data about indigenous and racialized people or groups. the research guide addresses the following: a) provides a curated list of datasets that include ethnicity and race variables which can be used to facilitate research on racialized people or groups in canada; b) describes how has the terminology about indigenous and racialized groups or identities has changed in the census of population since 1871; c) describes how to use census terminology to help find data about racialized people or groups outside of the census; https://doi.org/10.29173/iq1016 https://learn.scholarsportal.info/featured/data-on-racialized-populations/ 2 2/8 manuel, kevin, orlandini, rosa and cooper, alexandra. (2022) who is counted? ethno-racial and indigenous identities in the census of canada, 1871-2021, iassist quarterly 46(4), pp. 1-8. doi: https://doi.org/10.29173/iq1016 d) describes how to find contemporary and historical data, including census data, that contains data about racialized peoples or groups. in this article, the authors will focus on the census of canada and the modifications to how ethno-racial origin questions were asked. these questions reflect the ongoing perception of sociocultural attitudes, from rigid colonial thought to more recent ideas of diversity and inclusion. but even in the 2021 census of canada, a critical lens needs to be applied, as there are still potential opportunities to improve how societies across canada are represented in the official data. a note on terminology due to the historical nature of the data sources referred to in this article, terminology may include language that is problematic and/or offensive to contemporary users. specifically, vocabulary used to refer to racial, indigenous, ethnic, religious, and cultural groups, is specific to the time period when the data was collected and does not reflect the attitudes and viewpoints of contemporary society. history of the census of canada the history and present of canada is rooted in its colonial past. canada is a constitutional hereditary monarchy that is a federation of ten provinces and three territories. the head of state is queen elizabeth ii of the united kingdom and canada is part of the commonwealth of nations, formerly known as the british empire. some aspects of the country’s affairs, such as the economy and indigenous relations, are centrally administered through the federal government which is located in the capital ottawa, ontario. one of the agencies of the federal government in canada is its national statistical agency statistics canada (formerly the bureau of dominion statistics until 1971) which is responsible for the census. first conducted in 1871, the census of canada provides a snapshot of the people living in canada, collecting socioeconomic data to help inform public policy, decide parliamentary representation, and direct funding to resources across the country. initially run every ten years, the quinquennial census was introduced in 1956. throughout its history, the census has continued to evolve and change reflecting canada’s political and social transformations. race and ethnicity historically in the census of canada the origins of the census of canada are tied directly to colonialism. the first census in north america was conducted in 1666 in new france, now quebec to identify who was living in its colonial territory claims. following british expansion after they annexed new france in 1760, census activity was limited to individual british north american colonies. the 1867 british north america act (constitution act) created canada by uniting the provinces of ontario, quebec, nova scotia and new brunswick. the act also provided that a census of canada should be conducted every ten years starting in 1871 and new provinces and territories joined the confederation over time. the population at the time in 1871 was about 3.6 million and by 1911 the population had doubled to over 7.2 million, mostly through immigration of european settlers (urquhart, 1993). the census evolved to monitor the status of these new settlers but at the same time ignored and overlooked the indigenous and racialized populations. the embedded racist https://doi.org/10.29173/iq1016 3 3/8 manuel, kevin, orlandini, rosa and cooper, alexandra. (2022) who is counted? ethno-racial and indigenous identities in the census of canada, 1871-2021, iassist quarterly 46(4), pp. 1-8. doi: https://doi.org/10.29173/iq1016 attitudes of the colonial administration perpetuated the hegemony of erasure and assimilation of those marginalized. currently, statistics canada (2015) defines ethnic origin as ‘the ethnic or cultural origins of the person's ancestors... an ancestor is usually more distant than a grandparent.” but from 1871 to 1891, the term ‘place of origin’ was used to indicate where a respondent was born and what their race was. in 1901, the term ‘racial origin’ was introduced and remained in use until 1941. unlike the united states census which continued to use ‘racial origin’ after 1941, the census of canada removed any reference to ‘race’ and used ‘ethnic origin’ in 1946 or ‘origin’ in 1951 to define a respondent’s race. in the decades to follow there was a shift away from the term race and in 1961, 1971, and 1981 ‘ethnic or cultural group’ was used in the census. a problem for researchers in this period of the census is the lack of distinction between different ethnic, racial, and cultural groups recorded. some racial groups would self-identify through their country of origin, unable to record their actual race. as such: “when asked about their ethnic origins for the 1981 census, jamaican-descended respondents reported that they were british, while those of haitian descent identified as french. some indians (from india) identified as “status indian”—a category intended to enumerate the aboriginal population—because they believed the question inquired about their ethnic origin and immigration status” (thompson, 2020). in 1986, statistics canada introduced changes to the census that would help address the issue of a respondent being able to properly self-identify their ethnic or racial origin. respondents were now able to enter multiple ethnic or cultural groups in addition to stating which country(ies) they were a citizen of. the 1986 census is also the first time that a question about indigenous identity was asked as separate from a racial or ethnic origin question (this will be discussed further in the section on indigenous representation in the census). for example, the following was asked in the 1986 census asked: “to which ethnic or cultural group(s) do you or did your ancestors belong? mark or specify as many as possible. groups include: french, english, irish, scottish, german, italian, ukrainian, dutch, chinese, jewish, polish, black, inuit, north american indian, métis. other groups were written in by the individual and/or census taker.” (statistics canada, 1990). further changes were to come in the 1990s that would attempt to distinguish more specific ethno-racial data in the census. visible minority in the census of canada visible minority identity came about as a result of “close to 200 local, regional and national ethnic groups were contacted and asked to specify their data needs… most agreed that data on visible minorities are needed, but felt that a race question could be sensitive and controversial, even for visible minority groups” (statistics canada, 1990). introduced in 1996, the visible minority question allows a respondent to selfidentify as such based in the employment equity act and is defined as: ““persons, other than aboriginal peoples, who are non-caucasian in race or non-white in colour”. the visible minority population consists mainly of the following groups: south asian, chinese, https://doi.org/10.29173/iq1016 4 4/8 manuel, kevin, orlandini, rosa and cooper, alexandra. (2022) who is counted? ethno-racial and indigenous identities in the census of canada, 1871-2021, iassist quarterly 46(4), pp. 1-8. doi: https://doi.org/10.29173/iq1016 black, filipino, latin american, arab, southeast asian, west asian, korean and japanese” (statistics canada, 2021). modern versions of the census of canada do not include indigenous people within the visible minority classification. recently some researchers have begun to refer to the census classification of visible minorities as racialized peoples in order to better identify racialized groups in the data. accordingly, some organizations in canada have also responded to the needs to support racialized peoples. the ontario human rights commission defines racialized as “people from marginalized creed groups including ethnic origin, colour, ancestry, place of origin and citizenship... (and) is a generalized term to refer to people who are not indigenous or white” and as such it recognizes “...the unique specific historical experiences of indigenous peoples and considered separately from those of other racialized people” (2019). the ontario human rights commission elaborates that racialized includes people who identify with the statistics canada categories of south asian, chinese, black, filipino, latin american, arab, southeast asian, west asian, korean, japanese, or more than one of these categories.” so, in this instance, racialized is not race like that in the united states census, but based on ethnic categories from statistics canada. researchers trying to compare race and ethnicity between canada and the united states often encounter difficulty as the ethno-racial data is collected differently in the two countries. it is worth noting that there is no international standard on how each countries’ census questions are asked and categorized. often, international data comparison requires what is referred to as harmonization whereby data is organized into similar categories for analysis. as indigenous identity is not a visible minority category in the census of canada, it is important to elaborate on the history of how indigenous peoples have been represented in the census. indigenous representation in the census of canada indian act it is important to recognize that the census of canada’s categories for indigenous peoples is closely tied to the colonial legislation that was created in the past. the 1876 indian act was introduced to give the federal government wide sweeping powers over indigenous peoples in canada. it attempted to generalize a varied population of peoples, defining them by their “indian status” and assimilating them into nonindigenous society. “indian status” was defined as “any male person of indian blood reputed to belong to a particular band”, as well as “any child of such person” and to “any woman who is or was lawfully married to such person” (parrott, 2020). the indian act also defines how an individual can lose status, for example until the 1986 amendment to the indian act, women with indian status who married someone without status lost their status rights (parrott, 2020). the act does not directly reference non-status first nations, inuit or métis, but does define the structure of indigenous political structures, governance, cultural practices and education. the act has been gradually amended since 1951 by removing many discriminatory sections. however, the indian act still exists, and throughout its history the regulations, policies, and ideologies resulting from the indian act have defined and restricted indigenous peoples; and https://doi.org/10.29173/iq1016 5 5/8 manuel, kevin, orlandini, rosa and cooper, alexandra. (2022) who is counted? ethno-racial and indigenous identities in the census of canada, 1871-2021, iassist quarterly 46(4), pp. 1-8. doi: https://doi.org/10.29173/iq1016 these same policies and ideologies are reflected in the census and define how indigenous-identity questions are asked in the census. indigenous identities and terminology in history of the census in the 1870 census of manitoba and 1871 census of canada, and until 1891, indigenous identity was categorized under 'origin’ as ‘indian’. between 1901 and the 1931 census, an ‘indian’ was defined as an aboriginal person whose origin or race was ‘indian’ on the mother’s side, while in the 1941 and 1951 census, this was categorized under ‘indian’ or ‘eskimo’ through the father’s side. the 1951 census further defined: “for person’s of mixed and indian parentage, the origin recorded will be as follows: (a) for those living on indian reserves, the origin will be recorded as “native indian”; (b) for those not on reserves the origin will be determined through the line of the father” (dominion bureau of statistics, 1941). the 1951 amendment to the indian act replaced the concept of first nations ancestry with status through registration which is reflected in the 1961 census: “if a person reports “native indian” ask an additional question: is your name on any indian band membership in canada” and “note that “treaty indians” should be marked “band member”; if a person is of mixed white and indian parentage: a) consider those living on indian reserves as “indian” and determine band status [as outlined above] b) for those not on reserves, determine the ethnic or cultural group through the line of the father” (dominion bureau of statistics, 1961). used until 1981, ‘native indian’ was replaced in the 1986 census by the term ‘aboriginal’ with the following identifiers: first nations, métis and inuit. first nations is a term used to describe indigenous peoples in what is now canada, who are not métis nor inuit. it came into common usage in the 1970s and 1980s and has replaced the terms ‘indian’, ‘native indian’ and ‘native american’, although ‘indian’ is still used as the legal term (kestler, l et. al, 2009). inuit and métis prior to 1981, inuit were enumerated as ‘indian’ until the 1931 census. from 1941 to 1971, census enumerators entered ‘eskimo’, to describe an inuk (singular for inuit). inuit appeared on the census form in 1981. today, the métis are one of the three recognized indigenous peoples in canada alongside first nations and inuit. however, this was not always the case in the eyes of the government of canada as witnessed in an 1885 house of commons speech by prime minister sir. john a. macdonald. he stated, in reference to the métis, “if they are indians, they go with the tribe; if they are half-breeds, they are whites” (house of commons debates, 1885). the census reflected this attitude by categorizing métis in the ‘origin’ variable as ‘half-breed’ in the 1870 census of manitoba and the 1871 census of canada. in 1881, métis were enumerated as either ‘white’ or ‘indian’ and by 1901, métis are considered ‘half-breed’ again, but the term was not used again until 1941. finally, it is not until the 1981 census that a respondent can selfidentify as métis as one of the choices in the ethnic or cultural group question. from this point forward, métis is always included in the census. https://doi.org/10.29173/iq1016 6 6/8 manuel, kevin, orlandini, rosa and cooper, alexandra. (2022) who is counted? ethno-racial and indigenous identities in the census of canada, 1871-2021, iassist quarterly 46(4), pp. 1-8. doi: https://doi.org/10.29173/iq1016 the census in the 21st century in 2021, statistics canada replaced the term ‘aboriginal’ with ‘indigenous’ as the collective term for first nations, métis and inuit (statistics canada, 2021b). however, it is worth noting that the terminology is not changed retroactively so when looking at historical census data the terms native indian and aboriginal remain. in recent years of the census, some indigenous communities have not participated in the census. as they are considered sovereign territories, they choose not to have federal government census enumeration and rather collect the data themselves to support their own communities. the legacy of colonialism has created mistrust between indigenous communities and the various levels of government in canada, especially at the federal level. 92.5% of indigenous communities participated in census 2016 up from 89.9% in 2011 (grant, 2016). as such, fourteen first nations communities did not give permission to census 2016 enumerators (statistics canada, 2016). furthermore, “many urban indigenous people tend not to participate in the census due to factors such as poverty and its associated lack of a fixed address, mobility between communities and historical distrust of government and colonial policies' ' (yfile, 2019). these indigenous communities that did not participate in the census would collect demographic data themselves for their own decision making. over time, the census of canada has shifted and evolved in its classifications and categories of ethno-racial and indigenous identity as it responds to changes in the sociocultural constructs of canada. conclusion what has become evident in working on this project of identifying indigenous and racialized data in the census of canada is that there have been points in which certain groups of people have been excluded and ignored. the colonial mindset rendered those it considered not important to progress irrelevant. the terminology about ethno-racial identity over the past one hundred and fifty years in the census of canada has changed tremendously. in the early censuses of the 19th century, the information collected was based on where you were born and this fixed your identity to that place. but “in the mid-19th century, science and the scientific community served to legitimize society’s racist views” (smithsonian, 2021) so between 1901 and 1941, racial identity did become a part of the census. consequently, the census became an instrument to intensify the classification, segregation, and assimilation of the non-white elements of canadian society. there have been improvements to the canadian census in recent years. in addition to the greater recognition of what indigenous peoples and visible minorities bring to canadian society, the 2021 census added a question about gender identity in addition to sex of the respondent. however, sexual orientation was not included as statistics canada still considers this a ‘sensitive topic.’ the federal government is often seen as a monolith slow to change, but it is at the point it appears to be going in the right direction in terms of greater inclusivity with the census. however, it will only be evident in the next census in 2026 if there are new measures in place to better identify racialized, indigenous, and marginalized groups in canada. https://doi.org/10.29173/iq1016 7 7/8 manuel, kevin, orlandini, rosa and cooper, alexandra. (2022) who is counted? ethno-racial and indigenous identities in the census of canada, 1871-2021, iassist quarterly 46(4), pp. 1-8. doi: https://doi.org/10.29173/iq1016 references dominion bureau of statistics (1941) ninth census of canada enumerator manual. ottawa. dominion bureau of statistics. available at: https://publications.gc.ca/collections/collection_2017/statcan/cs98-1941i-eng.pdf) dominion bureau of statistics (1961) 1961 census of canada: enumeration manual. ottawa. dominion bureau of statistics. available at: https://publications.gc.ca/collections/collection_2017/statcan/cs98-1961-i-2-eng.pdf grant, t. (2016) ‘rise in census participation from indigenous communities in canada’, the globe and mail, 6 november. available at: https://www.theglobeandmail.com/news/national/rise-in-censusparticipation-from-indigenous-communities-in-canada/article32694558/ dominion of canada (1885) house of commons debates, 5th parliament, 3rd session: vol. 4. july 6, page 3113. available at: https://parl.canadiana.ca/view/oop.debates_hoc0503_04/557?r=0&s=3 kestler, l., crey, k, and hansen, e. (2009) indigenousfoundations.arts.ubc.ca: terminology. available at https://indigenousfoundations.arts.ubc.ca/terminology/ (accessed: 31 august 2021) ontario human rights commission (2021) under suspicion: research and consultation report on racial profiling in ontario. toronto: government of ontario. available at: http://www.ohrc.on.ca/en/under-suspicion-research-and-consultation-report-racial-profilingontario/1-introduction parrott, zach (2020) ‘indian act’, canadian encyclopedia. available at: https://www.thecanadianencyclopedia.ca/en/article/indian-act smithsonian, national museum of african american history and culture (2021) historical foundations of race. available at: https://nmaahc.si.edu/learn/talking-about-race/topics/historical-foundationsrace statistics canada (1990) 1991 census content development final report. ottawa. statistics canada statistics canada (1990) general review of the 1986 census. ottawa. statistics canada statistics canada (2015). ethnic origin of person. available at: https://www23.statcan.gc.ca/imdb/p3var.pl?function=dec&id=103475 statistics canada (2016) incompletely enumerated indian reserves and indian settlements. available at: https://www12.statcan.gc.ca/census-recensement/2016/ref/98-304/app-ann1-2-eng.cfm statistics canada (2021a) 355 years and counting. available at: https://census.gc.ca/about-apropos/355years-355-ans-eng.htm statistics canada (2021b) statistics on indigenous peoples. available at: https://www.statcan.gc.ca/eng/subjects-start/indigenous_peoples https://doi.org/10.29173/iq1016 https://publications.gc.ca/collections/collection_2017/statcan/cs98-1941i-eng.pdf https://publications.gc.ca/collections/collection_2017/statcan/cs98-1961-i-2-eng.pdf https://www.theglobeandmail.com/news/national/rise-in-census-participation-from-indigenous-communities-in-canada/article32694558/ https://www.theglobeandmail.com/news/national/rise-in-census-participation-from-indigenous-communities-in-canada/article32694558/ https://parl.canadiana.ca/view/oop.debates_hoc0503_04/557?r=0&s=3 https://indigenousfoundations.arts.ubc.ca/terminology/ https://indigenousfoundations.arts.ubc.ca/terminology/ https://indigenousfoundations.arts.ubc.ca/terminology/ http://www.ohrc.on.ca/en/under-suspicion-research-and-consultation-report-racial-profiling-ontario/1-introduction http://www.ohrc.on.ca/en/under-suspicion-research-and-consultation-report-racial-profiling-ontario/1-introduction https://www.thecanadianencyclopedia.ca/en/article/indian-act https://nmaahc.si.edu/learn/talking-about-race/topics/historical-foundations-race https://nmaahc.si.edu/learn/talking-about-race/topics/historical-foundations-race https://www23.statcan.gc.ca/imdb/p3var.pl?function=dec&id=103475 https://www12.statcan.gc.ca/census-recensement/2016/ref/98-304/app-ann1-2-eng.cfm https://census.gc.ca/about-apropos/355-years-355-ans-eng.htm https://census.gc.ca/about-apropos/355-years-355-ans-eng.htm https://www.statcan.gc.ca/eng/subjects-start/indigenous_peoples 8 8/8 manuel, kevin, orlandini, rosa and cooper, alexandra. (2022) who is counted? ethno-racial and indigenous identities in the census of canada, 1871-2021, iassist quarterly 46(4), pp. 1-8. doi: https://doi.org/10.29173/iq1016 thompson, d. (2020) “race, the canadian census, and the interactive political development”, studies in american political development, vol. 34, no. 4, pp. 44-70. thompson rivers university (2021) canadian census. history of the census. available at: https://libguides.tru.ca/censuscanada/history urquhart, m. et al. (1993) historical statistics of canada. ottawa. statistics canada. available at: http://epe.lac-bac.gc.ca/100/200/301/statcan/historical_statistics_can-e/sectiona/toc.htm yfile: york university’s news (2019) ‘toronto has twice as many urban indigenous people than previously believed.’ available at: https://yfile.news.yorku.ca/2019/05/02/toronto-has-twice-as-many-urbanindigenous-people-than-previously-believed/ endnotes 1 kevin manuel works at ryerson university in toronto, ontario and has been a data librarian for nearly 15 years. he has a ba in anthropology, a ma in sociology and a masters in library and information science. he also provides research support for black studies, lgbtq+ studies and is involved with the library’s indigenous engagement plan. kevin can be reached at kevin.manuel@ryerson.ca 2 rosa orlandini is a data services librarian at york university, toronto, ontario. she has a bsc in geography and a masters in library information studies. she provides research support for environmental studies, geography, geospatial information, and social science data and statistics. rosa can be reached at rorlan@yorku.ca 3 alexandra cooper is the data service coordinator at queen’s university, kingston, ontario, providing data and research data management support to researchers and students for about 20 years. she has a ba in religion and culture and canadian history. alexandra can be reached at coopera@queensu.ca 4 https://learn.scholarsportal.info/featured/data-on-racialized-populations/ https://doi.org/10.29173/iq1016 https://libguides.tru.ca/censuscanada/history http://epe.lac-bac.gc.ca/100/200/301/statcan/historical_statistics_can-e/sectiona/toc.htm https://yfile.news.yorku.ca/2019/05/02/toronto-has-twice-as-many-urban-indigenous-people-than-previously-believed/ https://yfile.news.yorku.ca/2019/05/02/toronto-has-twice-as-many-urban-indigenous-people-than-previously-believed/ mailto:kevin.manuel@ryerson.ca mailto:rorlan@yorku.ca mailto:coopera@queensu.ca https://learn.scholarsportal.info/featured/data-on-racialized-populations/ vol28-4.indd iassist quarterly winter 2004 5 by by oliver watteler 1 the iassist wiki – a copyleft solution for the 4th decade “all together now....“ the beatles introduction entering its fourth decade, iassist has set out to plan the future. according to the strategic plan for the period until 2009, the three main strategic directions are education, outreach and advocacy.2 what do these three points have in common? they are dealing with the distribution of information. iassist currently publish a magazine, maintain a website and run an e-mail list, which includes logs that are searchable, thereby realizing one of its prime objectives as described in article iii of the constitution that it wants to “foster international exchange and dissemination of information regarding substantive and technical developments related to social science machinereadable data.” there is no question concerning the importance of information. thus, donald waters wrote in 2002 when talking about the relationship between researchers and archives: “there are many dimensions to the good to be achieved, but two of them merit special mentioning. on the one hand, there is the joining together by scholars and the agents of education—universities, libraries, scholarly societies, and publishers—in serving the common interest of future scholarship by keeping good, or preserving, the digital resources now being created. on the other hand, there is the research and learning thereby made possible, which are the indelible marks of a good scholar. in other words, good archives make good scholars.”3 if one continues on this line of thought it is logical that good archivists and good data librarans are the basis for good archives and that they are well trained specialists that keep up their knowledge e.g. through participating in the activities of iassist. one possible way of sharing information is of course the internet. there are numerous information gateways on the web, for example: the virtual training suite4, which was presented by heather dawson at the iassist conference in 2001, the social science information gateway5, both being part of the resource discovery network in the united kingdom, or the vascoda6 portal in germany. and iassist itself which revised its site only recently. most of the gateways use state of the art technology like “harvesters” sifting through the enormous wealth of information on specific topics that can be found on the web. they are also encountering sites maintained by individuals and institutions that hold bits of information and hyperlinks of particular interest, which means that they are relying on existing content. but what if you want to foster information exchange on an international level? link again to the same sites as so many other pages have before? the wikipedia offers an interesting concept of a website maintained through the joint effort of a dedicated community, much like the members of iassist. wikipedia – what is it? most members of iassist will probably know the wikipedia.7 in its own words: “wikipedia is a web-based, multi-lingual, ‘copyleft’ encyclopedia designed to be read and changed by anyone. it is collaboratively edited and maintained by thousands of users via the wiki software […] and it is hosted and supported by the non-profit wikimedia foundation.”8 the wikipedia allows anybody who is registered to “boldly” edit whatever content she or he is interested in. a wiki generally means a web site comprised of the perpetual collective work of many authors. similar to a web log in structure and logic, a wiki allows anyone, using a web browser, to edit, delete or modify content that has been placed on the web site including the work of other authors. the wikipedia offers style guides, codes of conduct and several guidelines on how to proceed when working on an entry. the overall work is being loosely coordinated via the “community portal” including information on possible contributions (e.g. a task list) and basically everything you need in order to participate.9 furthermore, the wikipedia foundation encourages everybody to translate entries into 6 iassist quarterly winter 2004 their respective language, which leads e.g. to 180,000 german or 90,000 japanese equivalents of some of the 450,000 english articles. what does a wikipedia article look like? an article usually consists of: a) the entry itself, possibly extended by sections such as “see also” “external links” and “further reading”, b) four index-tabs above the entry, allowing the user to get to the discussion list, the editing view and the history log c) the toolbox offering links to a list of other entries that link to the entry in question (“what links here”), to a list of recent changes made in the wikipedia as a whole (“related changes”), and to a list of special pages e.g. pages that offer help to the user. articles that concern ambiguous topics are listed on so called “disambiguation pages”. if you are e.g. looking for “data” you are being told that a “datum is a statement accepted at face value (a “given”)” and some things about the etymology and the usage of the term. the disambiguation page then also tells you, among other options, that data is “a fictional android character in the star trek universe”. what are the pros and cons of the wikipedia? of course, an open source, freely editable pool of information also raises doubts and criticism. some of the cons should be mentioned here in order to reflect the experiences that have been made with the wikipedia. if you go through the pages criticising the wikipedia but which are part of this platform and thereby make the development progress more transparent you will find several severe points of negative assessment. the entry “wikipedia: why wikipedia is not so great”10 talks for example about the lack of accuracy, completeness and npovness (npov = neutral point of view) of several articles. the author argues that the credibility of sources “can be dubious because of the anonymous nature of the wiki”, that too many people merely add so called stub (short articles with little content marked to be expanded) and do not care about expanding existing articles, and that some participants refrain from maintaining a neutral point of view. the complete freedom of editing articles leads to bad style, nonsense content and arbitrary changes of articles in order to support individual points of view. this freedom can even end in what is called “edit wars”. this term refers to conflicts between competing content providers, a rather odd thing for an information system that was actually meant to serve a beneficial goal. “articles are sometimes copied virtually verbatim from other sources infringing on (international) copyright, particularly when no credit is given.” the wikipedia thus obviously has a serious copyright problem. those are just some of the arguments speaking out against “vandals” and “geeks” that seem to have a strong impact on the quality of the wikipedia. they uncover the immediate problem of completely open systems: they are democratic, but seemingly out of control. one of the harshest critics of the wikipedia concept was a former editor of the internet version of the encyclopaedia britannica, robert mchenry. in an online article he complains very much about the same details as mentioned above, but also admits that planning an encyclopaedia from the top is an almost unattainable aim:”i know, to begin with, that it can’t be done in any thoroughgoing way. the job is just too big. professional reviewers content themselves with some statistics -so many articles, so many of those newly added, so many index entries, so many pictures, and so forth -and a quick look at a short list of representative topics.”11 on the pro side the openness of the system is what makes it very appealing to users. a forum founded on a database system that allows easy access which obviously results in a decent pool of information. more than 168,000 registered users show that there are obviously quite a number of people having an active interest in the wiki-kind of information. and the collaborative work has produced some remarkable results. iassist and information sharing i see the objective of reinforcing the international information exchange, as mentioned above, as being at the core of an organization that obviously has a very positive public image. the notion that there is an entire community of people out there that may have encountered the same problems that i have can be a relief. no bureaucratic framework is slowing down the work and the people active within iassist are highly motivated. like many other voluntary organizations iassist wants to get more people to become involved and would like to see increased activities concerning training and education. an iassist wiki could be an opportunity for people to contribute when they could and in the way they like. as the e-mail list shows, people are seeking contact and the large world of data archives and users suddenly becomes very small. this is a good starting point for further activities. a wiki encyclopaedia could serve as a framework for information on the topics that most members are interested in: data preservation, data editing, data access and dissemination, iassist quarterly winter 2004 7 data confidentiality, data sources, and statistical and methodological issues. this follows very much the framework of iassist’s action groups and the organization of its website. the presentations from the iassist conferences could be utilized as quarries of information for the wiki knowledge base. or better still it could be a repository for the information presented at the meetings i.e. prepare your slides and turn them into wiki content at the same time. information bits that show up in e-mail conversations could be directly transferred into the data-base and scattered references could be bundled up in one place. there should certainly be some rules in order to avoid the problems with the wikipedia as described above: first of all, editing should be restricted to registered members of iassist. since all of them are professional using the organization for day-to-day problems on the job, there should be no complications with “geeks” that try to wreck the system. secondly, all entries should be in proper english. thirdly, the recommendations of the wikipedia on how to deal with your information should be taken over in order to make life easy for potential authors. the system should only be responsible for the rendering and formatting. and, last but not least, an open-minded notion toward changes and corrections of entries should be advocated, so that the harmony within iassist is not be disturbed by any “edit wars”. but i honestly believe that this is no problem within iassist. conclusion both the wikipedia and iassist live on the voluntary input from a dedicated community committed to sharing knowledge. this is a point they have in common, though the communities differ tremendously in size. the collaborative effort that ends up in well filled pools of information is probably the most appealing thing about the wikipedia. it is obviously this aspect of the wikipedia that draws a lot of interest toward it, which is an effect that could be utilized by iassist for its plans to reach out to the interested public. as in many organizations few people do a lot of the work, but there is always potential for small contributions that could add up to a solid base. a wiki kind of information system could be a bank for this input. with regards to the scope, an iassist wiki should be between an encyclopaedia britannica style knowledge base and an open forum for volunteers to contribute 24/7. to avoid the problems of the wikipedia a certain regulative framework should be established that facilitates the dissemination of helpful and well considered information. of course, it will take funding and maintenance to run the system, but the idea of wiki could carry the iassist spirit one step further into the virtual world. references members of the iassist strategic plan action group (ed.), iassist strategic plan, 2004-2009, 2004: <http:// www.iassistdata.org/membership/plan_june2004.pdf> the state of digital preservation: an international perspective. conference proceedings, washington dc 2002 (= clir report no.83): <http://www.clir.org/pubs/reports/ pub107/pub107.pdf> social science information gateway: <http://www.sosig. ac.uk/>. vascoda: <http://www.vascoda.de/>. virtual training suite: <http://www.vts.rdn.ac.uk/>. wikipedia: <http://en.wikipedia.org/wiki/wikipedia>. endnotes 1 oliver watteler, gesis – zentralarchiv für empirische sozialforschung an der universität zu köln, cologne (frg). e-mail: watteler@za.uni-koeln.de 2 members of the iassist strategic plan action group (ed.), iassist strategic plan, 2004-2009, 2004 <http:// www.iassistdata.org/membership/plan_june2004.pdf> 3 the state of digital preservation: an international perspective. conference proceedings, washington dc 2002 clir report no.83), p.91 <http://www.clir.org/pubs/ reports/pub107/pub107.pdf> 4 see <http://www.vts.rdn.ac.uk/>. 5 see <http://www.sosig.ac.uk/>. 6 see <http://www.vascoda.de/>. 7 for an overview of the wikipedia, see <http:// en.wikipedia.org/wiki/wikipedia>. 8 copyleft is a concept contrasting the idea behind copyrights. for details see article “copyleft” at wikidepdia. 9 on the technical side everything is open source and the development can be followed on the sourceforge platform. wikipedia uses php and the mysql dbms. 10 see <http://en.wikipedia.org/wiki/wikipedia:why_ wikipedia_is_not_so_great>. 11 see <http://www.techcentralstation.com/111504a.html> 1/11 dai, yun (2019) how many ways can we teach data literacy?, iassist quarterly 43(4), pp. 1-11. doi: https://doi.org/10.29173/iq963 how many ways can we teach data literacy? yun dai1 abstract academic libraries are ideally positioned to teach data literacy. what is ‘data literacy’ in the first place? is it the new information literacy? will the ways we teach information literacy limit imaginative ways to teach data literacy? with those questions in mind, the library of new york university shanghai has explored multiple ways to teach data literacy to undergraduate students through university events, ‘for-class’ instruction and workshops, and online casebooks. (1) we initiated the yearlong series of events titled ‘lying with data’, inviting faculty across disciplines to each address one core data literacy question that students of data science may misunderstand (2) we offered workshops and in-class instruction that are up-to-date with the latest technology and that fit with the curriculum. (3) we created online casebooks on various topics in the data lifecycle, tackling user needs at different levels. essential to our teaching activities are two core values: ‘let the quality speak for itself’, and ‘outreach by teaching’. keywords: data literacy, information literacy, instruction, data services, academic library introduction academic libraries are ideally positioned to teach data literacy. what is ‘data literacy’ in the first place? is it the new information literacy? will the ways we teach information literacy limit imaginative ways to teach data literacy? with those questions in mind, the library of new york university shanghai2 has explored multiple ways to teach data literacy through university events, ‘for-class’ instruction and workshops, and online casebooks. our library's data services unit is structured in the larger ‘data services ecosystem’ of the university, in which we have found our niche as data literacy advocates. our programming complements course offerings and formal scholarly activities, seeking to bridge the gap between those with rising interest but little knowledge in data and the seasoned data practitioners. we also work with universitywide platforms to promote literacy programs to the more general audience. while the data services program started to take shape since 2017, we have been broadly experimenting with forms of data literacy instruction. some of these instructional strategies, such as in-class instruction and workshops, are common practice in academic libraries, while others may be less conventional. specifically, we initiated the yearlong series of events titled ‘lying with data’, in which we invited faculty across disciplines to each address one core data literacy question that students may misunderstand we also offered for-class instruction and workshops that are up-to-date with the latest technology and that fit with the curriculum. outside of the classroom, we created online casebooks on various topics in the data lifecycle, tackling user needs at different levels. essential to our teaching activities are two core values: ‘let the quality speak for itself’, and ‘outreach by teaching’. in the following sections, i share the ideas that we experimented with, the rationale behind them, and the lessons we learned from these experiences. the bulk of the paper focuses on the event form of data literacy activities, which is believed to be the format less studied. https://doi.org/10.29173/iq963 2/11 dai, yun (2019) how many ways can we teach data literacy?, iassist quarterly 43(4), pp. 1-11. doi: https://doi.org/10.29173/iq963 revisiting data literacy instruction in 2004, iassist quarterly dedicated a whole double issue (vol. 28-2 and 28-3) to the discussion of data literacy and library instruction of data literacy. this was fifteen years ago, long before big data became a big deal; yet the questions asked then are still as relevant, or maybe even more so, today. one question that we are still asking is ‘what is data literacy’? another is ‘what is the role of a librarian or a data services support team at an academic library regarding service models and challenges in data literacy education and advocacy’. to the first question, several approaches to data literacy instruction arise from the body of literature on data literacy instruction in higher education. the first approach aligns closely with information literacy instruction. across the years, data literacy has been discussed with reference to the different versions of acrl information literacy competency standards for higher education, the most recent one of which is published in 2016 (acrl, 2016). data literacy, and its variations (data information literacy, research data literacy and science data literacy in terms of terminology), has been viewed by the library community as a subset, or extension, of information literacy in the time of high performance computing and an increasingly networked environment (schield, 2004; stephenson & caravello, 2007; carlson, fosmire, miller, & nelson, 2011; prado & marzal, 2013; womack, 2014; maybee & zilinski, 2015; shorish, 2015; frank & pharo, 2016). from this perspective, data literacy is, first of all, critical thinking applied to evaluating data sources and formats, and interpreting and communicating findings. as a companion to, or component of, data literacy, statistical literacy is the ability to evaluate statistical information as evidence; this involves the ability to understand summary statistics and graphs, and gauge the presentation of statistics, including the potential misuse of numbers (schield, 2004; gray, 2004; stephenson & caravello, 2007; womack, 2014). occasionally, it goes beyond discovering and evaluating datasets to include operationalizing research questions into measurable variables (beauchamp and murray, 2016). the second approach builds upon information literacy instruction and centers around data management and curation. core competencies proposed span knowledge and skills in databases, data acquisition and collection, documentation and standards, data management, curation, reuse, preservation, and sharing (qin and d’ignazio, 2010; schneider, 2013; carlson et al., 2011; carlson, johnston, westra, & nichols, 2013; mooney, collie, nicholson, & sosulski, 2014; shorish, 2015). the third approach encompasses the whole data lifecycle, ranging from data discovery and access, data interoperability and manipulation, data analysis, databases, data visualization, metadata, data preservation, data sharing and reuse, and best practices and ethics (fosmire and miller, 2008; prado & marzal, 2013). with the first question in mind, it is time that we re-envision the library’s role in data literacy instruction. it is generally agreed that the library should lead efforts in critical thinking, statistical literacy, information literacy, and data visualization literacy (gray, 2004; schield, 2004; womack, 2014). beyond that, it is the faculty’s role to teach statistical theory within the discipline, while the library’s complementary role can be found in offering practical knowledge to the patrons (thompson & edelstein, 2004), and instructing ethics and preservation of data (shorish, 2015). data literacy instruction https://doi.org/10.29173/iq963 3/11 dai, yun (2019) how many ways can we teach data literacy?, iassist quarterly 43(4), pp. 1-11. doi: https://doi.org/10.29173/iq963 is believed to be better delivered where it is integrated into the subject curriculum (hunt, 2004; stephenson & caravello, 2007; carlson et al., 2011; maybee & zilinski, 2015) and research projects (carlson et al., 2013), and where it is combined with the efforts of the strategic partners on campus (hogenboom, phillips, & hensley, 2011). while information literacy serves as the foundation of the data literacy program, i would argue that data literacy instruction at the library may not be limited within the context of information literacy. for instance, when teaching statistical literacy, data services specialists may introduce students to more sophisticated documents than summary statistics. we may also include the layer of computational skills, which many have, and perhaps some of the methodology component. in the age of technology advancement and big data, data visualization, web-interfacing technology supporting data needs, or even machine learning basics may have already become part of everyday language. in the section below, i will share a few of our experiments and regular instruction methods that seek to meet new challenges in the curriculum and beyond. teaching data literacy in three ways 1. campus wide dialogue we initiated the yearlong series of events titled ‘lying with data’, which builds upon our relationship with the center for data science and artificial intelligence and committee on critical inquiry by creating a platform to address core data literacy questions. the invited faculty tapped into the courses they have taught previously or their current research projects, which reduced the effort for faculty to prepare their sessions. topics included how to view graphs skeptically, getting the truth out of opinion surveys, understanding causal inference beyond correlations, and how to evaluate the appeal of statistics in advertisements. other proposed topics included blind spots in machine learning and fraudulent reporting on pharmaceutical trials as scientific misconduct. the plan was to cover a wide range of issues arising from various fields and schools of research, certainly not limited to data science or statistics. together with trivia, we have hosted four of the six proposed topics. this event evolved to be a campus wide dialogue once the vice chancellor volunteered to give a lecture, with support from two deans of the academic departments. but more importantly, this event was a collaboration with two partners on campus, the center for data science and artificial intelligence and the committee on critical inquiry. the committee was itself an initiative from the provost office. with its information literacy subcommittee, the committee directs students to participate in interactive discussions and hands-on activities to develop skills to evaluate information. the center is a research hub with groups of interdisciplinary scholars conducting the most advanced research with quantitative and computational methods. the collaboration also shows how the library’s data services program stands in the middle ground of the ‘data spectrum’ in relation to the scholarly activities carried out by the research institutes and the general information literacy efforts on campus. in this section, i will summarize these sessions, investigating the contents presented, approaches adopted, and messages delivered. each session potentially has offered many ways to review, but below i https://doi.org/10.29173/iq963 4/11 dai, yun (2019) how many ways can we teach data literacy?, iassist quarterly 43(4), pp. 1-11. doi: https://doi.org/10.29173/iq963 highlight where it concerns teaching data literacy, situating it in the larger ‘data ecosystem’ within the university. i have given more attention to some cases than others that deserve closer scrutiny. advocacy with graphs. this session analyzed three cases that relied on visual displays to show how advocates sometimes used graphs to persuade their audience to perceive patterns that did not really exist. the essence of this talk rested in where the speaker contrasted and compared alternative ways to array elements of a graph. the cases were examined both in their current forms as well as the counterfactual scenarios that could have left drastically different impressions on the readers. the first case is the 1987 gotti case (raab, 1987), where the speaker presented the graph ‘criminal activity of government informants’ that john gotti’s attorneys used to show the jury that the reputation of the witnesses were questionable. the graph arranged the columns with names of seven witnesses and rows with their long list of criminal records, where the most outstanding record was placed in the middle to catch people’s attention. but the graph could have been organized in a different way to reduce the visual prominence of some parts of it, such as aligning some crimes at the bottom. all those adjustments would make the audience think that some witnesses might be actually believable and therefore influence the jury’s decision. figure 1 criminal activity of government informants. reprinted from envisioning information (6th ed., p. 31), by e. r. tufte, 1998, cheshire, ct: graphics press. copyright 1990 by edward rolf tufte. in the next two cases, the speaker mimicked situations in political campaigns where advocates for democrats and republicans use the same sets of data to advance their own agendas. a group of tricks to engage emotions of the readers have been presented and inspected on a graph of the median family income of the u.s. (19472017). these tricks include: censoring time series to flatten the trend and to create an image of stability, or conversely, decreasing the units to increase the fluctuations; adding irrelevant reference data above or below the line being discussed to induce negative or positive moods in the audience; intentionally cutting off parts of the vertical axis to move the line up and down to different levels; selecting data points to shift the slopes in order to strengthen a message; not adjusting https://doi.org/10.29173/iq963 5/11 dai, yun (2019) how many ways can we teach data literacy?, iassist quarterly 43(4), pp. 1-11. doi: https://doi.org/10.29173/iq963 for inflation to distort the real trend. therefore, advocates always have a compelling story to tell by manipulating the level, trend and volatility of this line graph: ‘when we are in power, things get better; when they are in power, things get worse’ (lehman, 2018). following the analysis, the speaker shared some suggestions for how to make a graph fair but effective, and for how to view a graph with a critical eye. the core message, ‘every chart, every graph, represents a set of decisions about how to represent the information’ (lehman, 2018), was consistently conveyed in each case. getting the truth out of opinion survey. this session delved into a specific phase of quantitative social sciences research, data collection, through evaluating data sources and sampling schemes of survey design and implementation. drawing on the experience of collecting data for opinion studies in hong kong and shanghai, the speaker assessed how biases and errors in sampling could affect sample representativeness and data quality. technical notions, such as sampling frame, population specification, sampling errors, questionnaire and fieldwork design, and social desirability, were integrated into the delivery of the talk with small examples full of details for each concept. the key takeaway was to differentiate between the problems that should be avoided (e.g. measurement error) in survey data collection and the ones that could not be completely resolved (e.g. sampling error ), which could therefore be accommodated (wu, 2018). how to make causal inference out of casual correlation. this session looked into another phase in social sciences research, analyzing and interpreting data. the empirical challenge arises in making causal inference when ‘the fact that two trends seem to fluctuate in tandem does not prove that they are meaningfully related to one another’ (lu, 2018). the entire session was designed to deal with this challenge. the session started with an introduction of typical logical fallacies when making causal inference, including omitted variables, survivorship bias, and reverse causality. it then proceeded with concrete examples from the speaker’s own research area in finance, where the speaker illustrated how econometric tools, such as regression discontinuity design and difference-in-differences, could be leveraged in causal inference beyond correlation. the delivery style of this session resembled an academic seminar where the speaker presented his/her research, followed by a discussion among the group of audience with questions, comments, and debates. stop being fooled by advertisements. this session touched upon a real-world problem outside academic research. it demonstrated how companies’ marketing strategies capitalize on the statistical appeal of scientific evidence featuring numbers to influence consumers’ perception and purchases. in the past, companies have fabricated data, used biased polling and meaningless reference groups, cherry-picked data, and created misleading visualizations to appeal to customers. this session was constructed in a way that the speaker engaged the audience with a myriad of advertisements across industries and made connections with topics of the preceding lectures on sampling, causality versus correlation, and misleading graphs. a student posed an interesting question to the speaker: if as a marketing expert, could he, himself, effectively avoid the traps? the speaker replied by stating that even when conscious of the tricks of the trade, one could still be misled. however, ‘knowing and recognizing our cognitive biases will help prevent us from being fooled by https://doi.org/10.29173/iq963 6/11 dai, yun (2019) how many ways can we teach data literacy?, iassist quarterly 43(4), pp. 1-11. doi: https://doi.org/10.29173/iq963 misleading ads’, since ‘individuals tend to embrace information that supports their wishes and reject information that contradicts them.’ (yan, 2018) trivia. this session invited students to test their knowledge and critical thinking on topics discussed in the ‘lying with data’ series with game-based quizzes. students earned points of one to three during the quiz depending on how difficult the question was. unlike a lecture, which did not leave much room to exhibit technical details of a graph or a problem, the trivia session was devised to evaluate the granularities of a problem. students were encouraged to speak freely for each question, noting what went wrong and how the situation could have been improved or avoided. following the competition, the library staff reviewed all the questions and offered solutions to the puzzles. unfortunately, the series only presented quantitative social science perspectives, though we attempted a session on machine learning and another on critical examination of scientific misconduct. these sessions were later cancelled by the speakers due to conflicts with their schedules. materials from the series will be recycled and reused in two ways. the social sciences capstone project class will utilize the video recordings as course materials, since each of the topics are relevant to the course objectives. the series will also grow to be an online mini course consisting of pieces of short videos under the same theme, reproduced for a wider audience outside the university walls. moreover, how we marketed this event may also be worth noting. we created a mascot, the ‘lying with data’ pinocchio, for promoting the event throughout the year (see figure 1). pinocchio, a cartoon character whose nose grows when he lies, was selected for its common association with lying to reinforce the title of the series. we built several three-dimensional posters of human size that stood at different locations of the campus building (see figure 2). a student worker custom designed a limited number of metallic bookmarks in the shape of the mascot (see figure 3). students were encouraged to post selfies with the mascot on their social media, in order to generate buzz and publicity about the mascot and the event. figure 2. ‘lying with data’ mascot. illustration by keyin wu. https://doi.org/10.29173/iq963 7/11 dai, yun (2019) how many ways can we teach data literacy?, iassist quarterly 43(4), pp. 1-11. doi: https://doi.org/10.29173/iq963 figure 3 ‘lying with data’ poster. created by keyin wu. figure 4 ‘lying with data’ bookmark. designed by keyin wu. 2. for-class instruction it is noted that data literacy is best taught when integrated into the subject curriculum (hunt, 2004; stephenson & caravello, 2007; carlson et al., 2011; maybee & zilinski, 2015). we offer such instruction, both up-to-date with the latest technology and that fit with the curriculum. part of this program consists of stand-alone workshops outside classrooms; part of it are custom sessions designed for a course. we used to refer to the later kind as ‘in-class’ instruction, but here i am replacing the term with ‘for-class’ instruction to include those not happening physically in the classrooms but created for a class. one type of such for-class instruction is recorded videos. we have a case where our staff recorded eight videos on data scraping with python for the business analytics class (konagai, 2018). these are short video tutorials produced at our studio of the library. https://doi.org/10.29173/iq963 8/11 dai, yun (2019) how many ways can we teach data literacy?, iassist quarterly 43(4), pp. 1-11. doi: https://doi.org/10.29173/iq963 videos were distributed to the class as recommended resources. the benefit of this practice is that the contents can be recycled for future classes, and that students can view the materials at their own pace. but the risk also exists, especially when the tools are updating themselves quickly and thus the tutorials can be soon outdated. conventional in-class instruction sessions are delivered by request, most frequently coming from digital humanities, urban design and occasionally the business department. for instance, the gis specialist has been asked to deliver workshops on spatial analysis and geoprocessing. for technical sessions, a typical way is to give out a handout that contains step-by-step instructions of using a tool at the beginning of the class (luo, 2019). the majority of class time is devoted to the practices of problem solving, where students consult the handouts to complete a task. before the practices, conceptual questions are first explained, and student readiness are assessed with quizzes. for non-technical topics, a more visual display of the contents is desired to gather students’ attention. for instance, a session on the topic of storytelling with data visualization used the platform story maps to showcase the narrative revelation of data, rather than using powerpoint slides. the media, story maps, was itself part of the delivery in a session on visualizing gis data. 3. online casebooks in addition to the events and the for-class instructions, we have created online casebooks that are hosted on the data resources and services website (dai, 2019). the website has been set up as complementary resources to the pre-scheduled library workshops on various topics in the data lifecycle from data discovery (of chinese datasets), data preprocessing, data analysis, to data visualization. for each tool, we created tutorials targeted at users of various needs. for instance, topics under ‘coding smartly’ are tricks and tips with examples from real projects to solve particular problems. topics under ‘cases’ are tutorials that demonstrate the full workflow of a project with technical details. the remaining contents are introductory tutorials for a certain step in a project, such as exploratory analysis, transforming variables, handling strings, dates and times etc. after a few trials, we found that library workshops, even in-class instruction, were better suited for introductory topics. several attempts at more advanced, or intermediate workshops did not work out as expected in terms of attendance and student reception of the contents. this may be due to the fact that our main student body are undergraduate students. besides, students more often learn more advanced materials through self-exploration rather than teaching. therefore, intermediate and more advanced topics have been moved to the online space. discussion 1. how literate should we be there is always the question, how ‘literate’ we, librarians and technologists, should be in order to teach others data literacy. part of this question concerns the tools. the question puts more pressure on professionals nowadays as the field and industry of data science are advancing daily; this anxiety and need to upgrade ourselves constantly is real. at a new institution, the stress can double when there is no room for old technology to ‘die out’. yet service providers may have been trained with a small number of tools, while students arrive every year with experience using numerous emerging tools. https://doi.org/10.29173/iq963 9/11 dai, yun (2019) how many ways can we teach data literacy?, iassist quarterly 43(4), pp. 1-11. doi: https://doi.org/10.29173/iq963 the other part of this question concerns support on methodological or conceptual aspects in software assistance. some libraries refrain from offering such support, while others do not. perhaps an intermediate step would be asking ourselves to get familiar with the methodological or conceptual knowledge of a domain even though we may not offer direct support in advising or teaching. we should be able to join the conversation alongside data users, if not data scientists, to generate more meaningful communication and collaboration. 2. marketing essential to our teaching activities are two core values: ‘let the quality speak for itself’, and ‘outreach by teaching’. in a way, everything we do is about marketing. in fact, the ‘lying with data’ events probably have said much more about who we are than any brochure could possibly do; the data resources and services website presents to the university community the many facets of our services and capacity. word of mouth from faculty advocates carry us further than advertising. acknowledgement huge thanks go to caitlin mackenzie mannion, xiaojing zu, adrian hodge, edward junhao lim, jennifer anne wood stubbs, and fan luo for your valuable and detailed feedback to the paper. reference acrl (2016). framework for information literacy for higher education. retrieved from http://www.ala.org/acrl/sites/ala.org.acrl/files/content/issues/infolit/framework_ilhe.pdf beauchamp, a., & murray, c. (2016). teaching foundational data skills in the library. in kellam, l. & thompson, k. (eds.), databrarianship: the academic data librarian in theory and practice (pp. 81-92). chicago, il: acrl. carlson, j., fosmire, m., miller, c. c., & nelson, m. s. (2011). determining data information literacy needs: a study of students and research faculty. portal: libraries and the academy, 11(2), 629–657. <http://dx.doi.org/10.1353/pla.2011.0022> carlson, j., johnston, l., westra, b., & nichols, m. (2013). developing an approach for data management education: a report from the data information literacy project. international journal of digital curation, 8(1), 204-217. <http://dx.doi.org/10.2218/ijdc.v8i1.254> dai, y. (2019). data resources and services. retrieved from https://shanghai.hosting.nyu.edu/data/ fosmire, m., & miller, c. (2008). creating a culture of data integration and interoperability: librarians and earth science faculty collaborate on a geoinformatics course. retrieved from https://docs.lib.purdue.edu/cgi/viewcontent.cgi?article=1850&context=iatul frank, e. p., & pharo, n. (2016). academic librarians in data information literacy instruction: a case study in meteorology. college & research libraries, 77(4), 536-552. <http://dx.doi.org/10.5860/crl.77.4.536> https://doi.org/10.29173/iq963 http://dx.doi.org/10.1353/pla.2011.0022 http://dx.doi.org/10.2218/ijdc.v8i1.254 https://shanghai.hosting.nyu.edu/data/ https://docs.lib.purdue.edu/cgi/viewcontent.cgi?article=1850&context=iatul http://dx.doi.org/10.5860/crl.77.4.536 10/11 dai, yun (2019) how many ways can we teach data literacy?, iassist quarterly 43(4), pp. 1-11. doi: https://doi.org/10.29173/iq963 gray, a. s. (2004). data and statistical literacy for librarians. iassist quarterly, 28(2/3), 24-29. <http://dx.doi.org/10.29173/iq793> hogenboom, k., phillips, c. m. h., & hensley, m. (2011). show me the data! partnering with instructors to teach data literacy. retrieved from http://www.ala.org/acrl/sites/ala.org.acrl/files/content/conferences/confsandpreconfs/national/2011/ papers/show_me_the_data.pdf hunt, k. (2004). the challenges of integrating data literacy into the curriculum in an undergraduate institution. iassist quarterly, 28(2/3), 12-15. <http://dx.doi.org/10.29173/iq791> konagai n. data scraping with python. retrieved from nyu stream nyu community private access. lehman, j. (2018). advocacy with graphs. lecture presented at the ‘lying with data’ events. lu, y. (2018). how to make causal inference out of casual correlation. lecture presented at the ‘lying with data’ events. luo f. (2019). spatial analysis and geoprocessing. retrieved from http://bit.ly/2uhizr8 maybee, c., & zilinski, l. (2015). data informed learning: a next phase data literacy framework for higher education. in proceedings of the 78th asis&t annual meeting: information science with impact: research in and for the community (p. 108). american society for information science. <http://dx.doi.org/10.1002/pra2.2015.1450520100108> mooney, h., collie, w. a., nicholson, s., & sosulski, m. r. (2014). collaborative approaches to undergraduate research training: information literacy and data management. advances in social work, 15(2), 368-389. <http://dx.doi.org/10.18060/15089> prado, j. c., & marzal, m. á. (2013). incorporating data literacy into information literacy programs: core competencies and contents. libri, 63(2), 123–134. qin, j., & d'ignazio, j. (2010). lessons learned from a two-year experience in science data literacy education. retrieved from http://docs.lib.purdue.edu/iatul2010/conf/day2/5/ raab, s. (1987, march 14). a weakness in gotti case; major u.s. witnesses viewed as unreliable. the new york times, retrieved from https://www.nytimes.com/1987/03/14/nyregion/a-weakness-in-gotticase-major-us-witnesses-viewed-as-unreliable.html schield, m. (2004). information literacy, statistical literacy and data literacy. in iassist quarterly, 28(23), 6-11. schneider, r. (2013). research data literacy. in european conference on information literacy (pp. 134140). springer, cham. <http://dx.doi.org/10.1007/978-3-319-03919-0_16> https://doi.org/10.29173/iq963 http://dx.doi.org/10.29173/iq793 http://www.ala.org/acrl/sites/ala.org.acrl/files/content/conferences/confsandpreconfs/national/2011/papers/show_me_the_data.pdf http://www.ala.org/acrl/sites/ala.org.acrl/files/content/conferences/confsandpreconfs/national/2011/papers/show_me_the_data.pdf http://dx.doi.org/10.29173/iq791 http://bit.ly/2uhizr8 http://dx.doi.org/10.1002/pra2.2015.1450520100108 http://dx.doi.org/10.18060/15089 http://docs.lib.purdue.edu/iatul2010/conf/day2/5/ https://www.nytimes.com/1987/03/14/nyregion/a-weakness-in-gotti-case-major-us-witnesses-viewed-as-unreliable.html https://www.nytimes.com/1987/03/14/nyregion/a-weakness-in-gotti-case-major-us-witnesses-viewed-as-unreliable.html http://dx.doi.org/10.1007/978-3-319-03919-0_16 11/11 dai, yun (2019) how many ways can we teach data literacy?, iassist quarterly 43(4), pp. 1-11. doi: https://doi.org/10.29173/iq963 shorish, y. (2015). data information literacy and undergraduates: a critical competency. college & undergraduate libraries, 22(1), 97–106. <http://dx.doi.org/10.1080/10691316.2015.1001246> stephenson, e., & schifter caravello, p. (2007). incorporating data literacy into undergraduate information literacy programs in the social sciences: a pilot project. reference services review, 35(4), 525–540. <http://dx.doi.org/10.1108/00907320710838354> thompson, k. a., & edelstein, d. m. (2004). a reference model for providing statistical consulting services in an academic library setting. iassist quarterly, 28(2), 35-38. <http://dx.doi.org/10.29173/iq795> tufte, e. r. (1998). envisioning information (6th ed.). cheshire, ct: graphics press. womack, r. (2014). data visualization and information literacy. iassist quarterly, 38(1), 12–17. <http://dx.doi.org/10.29173/iq619> wu, x. (2018). getting truth out of opinion survey. lecture presented at the ‘lying with data’ events. yan, d. (2019). stop being fooled by advertisements. lecture presented at the ‘lying with data’ events. 1 contact: yun dai, nyu shanghai library. email: yun.dai@nyu.edu 2 nyu shanghai is a portal campus in the nyu global network and a joint venture between nyu and east china normal university. founded in 2012, it is a campus with around 1,200 enrolled undergraduate students from more than 70 countries and more than 200 full-time international faculty members. https://doi.org/10.29173/iq963 http://dx.doi.org/10.1080/10691316.2015.1001246 http://dx.doi.org/10.1108/00907320710838354 http://dx.doi.org/10.29173/iq795 http://dx.doi.org/10.29173/iq619 mailto:yun.dai@nyu.edu the creative commons-attribution-noncommercial license 4.0 international applies to all works published by iassist quarterly. authors will retain copyright of the work and full publishing rights. guest editors’ notes robert stalone buwule1 and winny nekesa akullo2 we are excited to note that iassist’s africa chapter has continued to grow bigger and stronger. after a successful first iassist africa regional workshop in uganda during january 2021, a second iassist africa regional workshop took place in ibadan, nigeria in west africa october 4th through october 7th, 2022. we are delighted to share with you the papers in this issue, most of which were presented at the second iassist africa regional workshop at the university of ibadan, nigeria. the first paper unpacks the application of emerging technologies for research support in academic libraries in the modern era. the authors are dr. sophia v. adeyeye and taofeek abiodun oladokun who explain how emerging technologies offer innovative ways of supporting research activities. these emerging technologies provide tools and resources that streamline the research process and ensure proper visibility for the research outputs of academic libraries’ clients. the article explores various areas where academic libraries can apply emerging technologies such as data mining, data management, artificial intelligence, library automation and scholarly communication, among others. the article further highlights the setbacks academic libraries in nigeria are facing in the application of emerging technologies such as lack of infrastructure, librarians’ skills, and negative attitude towards change. the second article, authored by ms. akinyoola oladoyin grace, is titled ”knowledge and perception of librarians towards cloud-based technology in academic libraries in southwest, nigeria”. in this paper the author reveals that librarians in academic libraries in southwest nigeria are familiar with and use cloudbased technologies. however, the librarians seem to have a negative attitude towards the use of these technologies. therefore, there is a need for a staff development program that would enable the librarians to keep pace with the latest technologies. such a program could be funded by government and executed through seminars, conferences, and workshops so as to enhance the librarians’ skills with cloud-based technologies. the third paper presents the preservation of election data and security in the fourth industrial revolution. election malfeasance and violence have been experienced in nigerian political systems since 1959. in this paper sunday tunmibi and wole olatokun explore how the world’s gradual move into the fourth industrial revolution (4ir) could be harnessed to ensure the preparation of free and fair elections. the paper suggests specific 4ir technological solutions to electoral data security and preservation challenges. it also suggests policies to serve as catalysts for the independent national electoral commission. https://creativecommons.org/licenses/by-nc/4.0/ suggested policies for 4ir technologies relate to artificial intelligence, big data, internet of things, robotics, block chain, cloud computing and 3-d printing. these are the 4ir technologies that dictate the pace of activities in all walks of life including security and managing a free and fair national election. in the fourth paper ologbosere oluwatosin abiodun discusses the significance of data literacy in the era of big data. it emphasizes the role of big data as a fundamental building block of truth focusing on the emergence of data literacy. data literacy is a crucial subset of information literacy necessary for navigating the virtual landscape. essential data literacy skills that are needed for navigating the dynamic twenty-firstcentury environment are highlighted. integrating data literacy into higher educational programs, particularly in libraries, is stressed for relevance in meaningful information resource utilization. this paper shows how the integration of data literacy in higher education emphasizes the critical role of data literacy in the context of economic growth, development and informed decision-making and fosters sustainable development. in the fifth paper titled ”data protection and right to privacy legislation in kenya”, author andrew matoke mankome articulates how the parliament of kenya enacted the data protection legislation in november 2019. this new law guaranteed the right to privacy as a fundamental right. data protection and citizens’ right to privacy is now a topical concern in kenya and around the world. this paper expounds how the new law comes at a time when data security and privacy concerns are prevalent and lack of them result in loss of reputation and identity; safety concerns; legal penalties; and compensation for damages or loss of business. this is mainly because of the increasing globalization, cross-border transactions, internet penetration, and the use of social media and digital platforms among citizens, private institutions, and governments. this paper reviews the crucial provisions of the data protection law which covers regulated actions seeking compliance by data controllers and processors under the stewardship of the office of the data protection commissioner ('odpc'). a comparative analysis of the practice in other jurisdictions is also provided. retracted 10/2024: see retraction notice. 1 dr. robert stalone buwule is the university librarian of mbarara university of science and technology in uganda (rbuwule@must.ac.ug) 2 ms. winny nekesa akullo is the head of records and information management at the national social security fund of uganda (winny.nekesa@yahoo.com) https://iassistquarterly.com/index.php/iassist/article/view/1080 instructions for authors of the iassist quarterly 1/4 watkins, trevor and cain, jonathan o. (2022) systematic racism in data practices, iassist quarterly 46(4), pp. 1-4. doi: https://doi.org/do.be/doo systemic racism in data practices trevor watkins1 jonathan o. cain2 positionality statement as we begin to discuss this issue, its origins, and its importance in contemporary society, i wanted to acknowledge my positionality and the role that it may play in the formation of this issue. jonathan o. cain is an african-american male working in the lis field. before moving into administration, i taught data and digital literacy and worked on developing programs that focused on improving access to these critical skills at zero cost to learners. it is important to acknowledge my positionality and the lens through which i see the data science field. trevor watkins is an african american male working in the lis field at an academic institution in an academic library. i teach critical data literacy workshops and engage in diversity and bipoc-related digital projects with faculty, students, and the broader academic community across the country. i am also a researcher and practitioner in artificial intelligence (ai) and data science. the global pandemic, its impacts, and why it matters we first met in august 2020 to discuss the possibilities of this special issue about five months into the pandemic. we spent a good chunk of that meeting getting to know each other and, most importantly, discussed the toll the pandemic placed on our communities and us. it is probably safe to say that many of you, at some point, were uncertain of the future. like most people worldwide, we lost family and friends or knew of people who succumbed to covid-19 and other illnesses that weren't treated because the focus shifted to covid-19. we get it. at one point, covid-19 killed over three thousand people per day (centers for disease control and prevention (cdc), 2022). according to data from the cdc, 90% of the 385,676 people who died between march and december 2020 had covid-19 listed as the underlying cause of death on their death certificate. the murders of ahmaud arbery in february, breonna taylor in march, and george floyd in may 2020 sparked civic unrest across the united states (us) and protests across the globe in solidarity against racial injustice. when we announced this special issue and initiated a call for papers, we didn't get much of a response initially. we expected and acknowledged that it would probably take some time before we received inquiries or proposals about the issue, the intent to submit, or any submissions. like many of you, we are still picking up the pieces from 2020 and dealing with the aftermath of covid19. the pandemic may be over now, depending on whom you ask, but the emotional scars are still there and may remain so for quite some time. patience was the one quality we all had throughout this process, which is why we can present this publication today. data and liberatory technology liberatory technology. this is a concept that invited contemplation as we sat down to record our reflections on this special issue. in drawing together scholars, educators, and practitioners to address the issue of data and its relationship to race, ethnicity, and representation, we, as coeditors, were making a statement about the importance of data, the material impact that this seemingly abstract and ethereal object can and does have on individual and community lives. and thinking about that https://doi.org/do.be/doo 2/4 watkins, trevor and cain, jonathan o. (2022) systematic racism in data practices, iassist quarterly 46(4), pp. 1-4. doi: https://doi.org/do.be/doo impact brought liberatory technology to the front of our minds. the definition of liberator technology offered by the ida b. wells just data lab intrigues us and invites us to grapple with that topic. they defined liberatory as something that "supports the increased freedom and wellbeing of marginalized people, especially black people outside of capitalism and settler colonial power structures" and technology as "a tool used to accomplish a task." and as we contemplate this set of definitions, we are left to question whether data can be a liberatory technology or not. (liberatory technology and digital marronage, n.d.) in liberation technology: black protest in the age of franklin, richard s. newman draws parallels with the asserting ownership and mastery of new communication technologies and black liberation activities. reflecting on the transformative nature of print technology, he writes, "if the marquis de condorcet was right in 1793 that print had unshackled europe from medieval modes of thought and action, then it is also true that print was perhaps the first technology to liberate blacks from the servile images that had long haunted their existence in western culture." and draws a 19th-century example of how it expressly connects to black lives post-emancipation noting "w. e. b. du bois certainly thought that black history and print history worked in tandem. wherever one found newspapers in the postcivil war south, he observed, one found some form of black freedom" (richard s. newman, 2009, p. 175). he even notes how scholars note that black activists embraced other communication technologies like photography "to reshape the image of african americans in nineteenth-century culture." (richard s. newman, 2009, p. 175) we have no shortage of examples of how data and data-driven technologies fail to support the "increased freedom and wellbeing of marginalized people outside of capitalism and settler colonial power structures." in 2016, propublica published machine bias, a report that looks at risk assessment technologies used in arraignment and sentencing. they report that "the formula was particularly likely to falsely flag black defendants as future, wrongly labeling them this way at almost twice the rate as white defendants" and "white defendants were mislabeled as low risk more often than black defendants" (julia angwin, 2016). a 2021 article, fairness in criminal justice risk assessments: the state of the art, in their analysis, noted, "the false negative rate is much higher for whites so that violent white offenders are more likely than violent black offenders to be incorrectly classified as nonviolent. the false positive rate is much higher for blacks so that nonviolent black offenders are more likely than nonviolent white offenders to be incorrectly classified as violent. both error rates mistakenly inflate the relative representation of blacks predicted to be violent. such differences can support claims of racial injustice. in this application, the trade-off between two different kinds of fairness has real bite." (berk et al., 2021, p. 33) these are just a few examples of how these technological developments, on their own merits, fail to meet the definition offered by the authors of the "liberatory technology and digital marronage" zine from the ida b. wells just data labs. reflecting on the technological path illustrated by newman, the work of ownership and mastery of the tool provides the potential for it to be liberatory. through this lens, the work of the just data lab is exemplary for this meditation; it draws a direct line from technology, education, mastery, and liberatory technology. https://doi.org/do.be/doo 3/4 watkins, trevor and cain, jonathan o. (2022) systematic racism in data practices, iassist quarterly 46(4), pp. 1-4. doi: https://doi.org/do.be/doo data in higher education data literacy education is an area that has been a focus of our careers in librarianship. it's a space where we saw the libraries' ability to make a meaningful impact. data has had a tremendous impact on college campuses, from how research is conducted to the pressures colleges feel from stakeholder groups: students, governments, funders, donors, and employers to prepare students with the data and technology skills to gain employment in the knowledge economy. as colleges and universities have turned (with varying degrees of success) to meet the needs of these communities, a myriad of explorations on the importance of the representation of these marginalized communities in these systems—to combat and dismantle the harmful practices that we see embedded in the systems that drive society and the potentially debilitating consequences they produce. that is partly why the works in this special issue are so important at this moment in time. these scholars and scholar-practitioners are engaging with these issues that drive the opaque structures surrounding us. and hopefully, their work can give us another perspective on how to engage with these structures and transform them to support liberatory practices. the entries in this issue we have some fantastic articles for you to read in this issue. we open with an article by kevin manuel, rosa orlandini, and alexandra cooper, who discuss how the collection process of racial, ethnic, and indigenous data has evolved in the canadian census since 1871, the erasure of minorities and indigenous citizens from those censuses, and the work to restore and accurately identify and categorize racialized groups. in the next article, leigh phan, stephanie labou, erin foster, and ibraheem ali present a model for data ethics instruction for non-experts by designing and implementing two data ethics workshops. they make important points about the failure of academia to incorporate the ethical use of data in course curriculums and digital literacy training and demonstrate how academic libraries have become an essential resource for the academic community. their workshop structure can be modeled for any academic library that endeavors to provide a similar service to its community. in the third article, natasha johnson, megan sapp nelson, and katherine yngve, interrogate the collective and local purposes of institutional data collection and its impact on student belongingness and propose a framework based on data feminism that centers the student as a person rather than a commodity. finally, our closing article from thema monroe-white focuses on marginalized and underrepresented people in the data science field. the author proposes that racially relevant and responsive teaching is necessary to recruit more people from these groups and diversify the field. she discusses how the ladson-billings model of cultural relevant pedagogy has been applied and is beneficial to stem curriculums, and how a liberatory data science curriculum could promote a student's voice and sense of belonging. conclusion we want to thank all those involved in producing this special issue. we want to thank the authors first. their patience, dedication, and perseverance throughout this process were much appreciated. the https://doi.org/do.be/doo 4/4 watkins, trevor and cain, jonathan o. (2022) systematic racism in data practices, iassist quarterly 46(4), pp. 1-4. doi: https://doi.org/do.be/doo reviewers provided timely, very detailed, and thorough feedback. we would be remised if we didn't acknowledge their hard work and labor. we would like to thank the iq editorial team, michele hayslett and karsten boye rasmussen, for working with us over the last two years, and ofira schwartzsoicher, for helping us get to the finish line. trevor watkins jonathan o. cain references berk, r., heidari, h., jabbari, s., kearns, m., & roth, a. (2021). fairness in criminal justice risk assessments: the state of the art. sociological methods & research, 50(1), 3–44. https://doi.org/10.1177/0049124118782533 flipsnack. (n.d.). liberatory technology zine. flipsnack. retrieved december 17, 2022, from https://www.flipsnack.com/ebc8cd77c6f/liberatory-technology-zine.html liberatory technology and digital marronage. (n.d.). ida b. wells just data lab. retrieved december 17, 2022, from https://www.thejustdatalab.com/tools-1/liberatorytechnology-and-digital-marronage mattu, j. a., jeff larson,lauren kirchner,surya. (n.d.). machine bias. propublica. retrieved december 17, 2022, from https://www.propublica.org/article/machine-bias-risk-assessmentsin-criminal-sentencing richard s. newman. (2009). liberation technology: black printed protest in the age of franklin. early american studies: an interdisciplinary journal, 8(1), 173–198. https://doi.org/10.1353/eam.0.0033 endnotes 1 trevor watkins is the teaching and outreach librarian at george mason university libraries. he can be reached at twatkin8@gmu.edu. 2 jonathan o. cain is the associate university librarian for research and learning at columbia university libraries and can be reached at joc2122@columbia.edu. https://doi.org/do.be/doo https://doi.org/10.1177/0049124118782533 https://www.flipsnack.com/ebc8cd77c6f/liberatory-technology-zine.html https://www.thejustdatalab.com/tools-1/liberatory-technology-and-digital-marronage https://www.thejustdatalab.com/tools-1/liberatory-technology-and-digital-marronage https://www.propublica.org/article/machine-bias-risk-assessments-in-criminal-sentencing https://www.propublica.org/article/machine-bias-risk-assessments-in-criminal-sentencing https://doi.org/10.1353/eam.0.0033 mailto:twatkin8@gmu.edu mailto:joc2122@columbia.edu iassvol201 17winter 1996 social science data services during the last five years of the millennium : developments in the delivery and support of data services for academic research in europe and north america. introduction, aims and background in general, data librarians are supported by researchers and computer staff in their view that demands for data services are likely to increase. even where there have been technology related savings2, other technology based tasks have arisen to add to the number of tasks performed by data support services for example, the management and/or construction and maintenance of web interfaces to data and associated documentation and literature. the case for investing in data support services may seem clear to members of organisations such as iassist, css and cause. and in the more recent reports produced by research/teaching support funding bodies such as the uk’s esrc and jisc there has been a marked increase in references to the data support environment. the main aim of this paper is to see if there is some empirical basis for the claim that investment in the continued development of data support services is worthwhile. the establishment of this claim will provide a sound basis from which to present the likely development scenarios of academic data services up to 2000. the background to this study is associated with observations of a number of trends in institutional policy in the broad areas of social science support and administration in european and north american universities and associated research centres. these trends include: • increasing funding pressures on researchers and research supervisors to speed-up submission rates for example, from 1998, the uk’s economic and social research council, (esrc) will only fund phds at institutions where 60% or more students submit their phds within 4 years (currently, this is set at 50%)3; • the development of institutional measures to ensure researchers (and teachers) have necessary data resources and appropriate information systems (is) infrastructure about 70% of us academic institutions claim to have is strategies4; and • changes in is infrastructure which have affected the resource demarcation between hitherto autonomous entities for example, the integration of some or all of audio-visual, computer, data, library, network and telecommunications services5. to this can be added societal changes, such as the rise of so-called meritocratic practices such that “good policy” requires that position/status and associated resourcing have some empirical basis, as in evidence-based planning requirements6. the first set of facts gathered in this study relates to the current and projected growth rates in empirical research. whichever way these are indicated, this growth is dramatic. the results of three methods of assessing empirical trends are summarised as follows: • article-content analysis shows a consistent growth in the proportion of empirically based journal articles (oswald, 1992; figlio, 1994; stigler, 1995; and platt, 1996)7; • data access enhancements, particularly those associated with networking and interfacing, continue to speed-up the process of acquiring data and associated bibliographic references for example, biron, bls, ibss, icpsr and many local developments such as the data subsetting services at the nber, ssdc and cepis8; and • it enhancements (storage and processing) have enabled major increases in productivity recorded in studies of empirical research outputs (cep/lse) and business productivity measurement (mit’s ccs)9. the fulbright study is the basis for the second part of this investigation. it captures data support experiences from three perspectives: research, data services and it/computer support. on the issue of efficiency of research and in-house data support, views are summarised as follows: • the researcher-teacher view (28/30) is that data support (local and central) is an essential component of an efficient research environment but, according to some (10/30), this may not be for ever; by adam lubanski1, information systems manager esrc centre for economic performance 18 iassist quarterly interviewees researcher data support it support total interviewed 30 27 26 female1 6 18 5 male 24 9 21 job 9 sras (senior research assistant/officer 8 data librarians 3 data archivists 7 it assistants 13 professors 3 data consultants 14 it managers 8 directors (i.e. directors of research centers 2 info. managers 4 res. managers 7 data managers 5 it directors research center = 8 (data center = 7) 8 (of 400 fte) 5 (of 100 fte) 8 (of 20 fte) 7 (of 22 fte) 8 (of 14 fte + it) 7 (of 15 fte) university = 20 (17 research) 17 (of ? fte) 12 (of 25 fte) 11 (of ? fte) published 27 (ibss) 15 (cause/effect, iassist, ibss) 8+ (cause/effect iassist, ibss) cited 23 (isi) 90? citations 12 (cause/effect, iassist approx. not counted (size/scale average) research center data center university research university teaching 80 fte 50 fte 5,000s+?f 6,000s+?f 2.5 fte 2 fte 2 fte 1 fte 2 fte 3 fte 30 fte (lse) 30 fte (lse) • the data support service/person view (25/27) is that data and information services are experiencing a major upturn in demand from both research and teaching activities (4 respondents also cited an increase in administration demands for data advice); and • the majority it/computer support service/person view (17/26) is that the acquisition of information and data support skills has become vital to their career prospects others felt that networking and teaching support together with some integration of audio-visual support appeared a more fruitful path. what seems obvious to the stakeholders, however, may not be fully recognised by the funders and planners. ultimately, in the long run, data support funding will be determined by economic criteria. the bad news at the present time, is that many data services are not well placed within the order of things to ensure that their strong economic arguments are well represented. the good news is the data. background data on data support services the selection of ninety or so interviewees, split about equally by the above types, was based on publication-citation methods (gutman library, february 1995). briefly, this method adopted the following research sequence: data was captured mainly during 30-40 minute interviews (during some thirty visits to north american research institutions, march to may 1995); additional data was gathered from preliminary internet searches and follow-up email to interviewees typically, clarification of interview notes. supplementary environmental evidence was gleaned from institutional policy documents as they related to data services/ support, and collected during the study. these included institutional responses from lse, esrc, national and state archives, data suppliers (e.g. the bls) and a sample of north american universities. the sample population: 1. an empirical basis for the prominence of women in computer-based data support is reported in anderson, r.e., 1987. ? denotes that figures were not noted at time of interview. (all figures are for economic/social sciences. it support includes networking and systems staff.) 19winter 1996 of seventeen research universities visited, virtually all (16) had a level of local data support far higher than that found (informally) in the uk. the one university which did not claim to have any formal arrangements for data support did, however, provide a very competent it and computing advisory service together with a catalogued tape library facility. basic advice about dataset management tasks was given by a program advisory team which referred detailed queries to an “analyst programmer with experience of databases”. of the sixteen research universities claiming to provide a “resourced data service”, four were classified as having basic data support10 ; nine were classified as having intermediate data support11; and three fitted the classification of full data support12. • of all forty-eight respondents interviewed at the 17 research universities, 45 expressed unprovoked favourable opinions of data support services although nearly all said this was an under resourced area (and were actively lobbying for better funding); only two researchers said that their level of data support was adequate for local needs; and • about half of the data support services/libraries were managed by the library service a trend which was generally welcomed, but opposed by 4 respondents (3c and 1d i.e. three computer staff and one data support person) who favoured independence. of seven research universities with large independently funded economic/social research centers i.e. similarly configured to lse and cep: • all seven had data support facilities (often named data libraries) both centrally based typically, managed by the library (1) or it services divisions (2), sometimes independent (4) and devolved in the research centers themselves (data from interviews held at harvard/nber, princeton/opr, cornell/ciser, syracuse/cpr, wisconsin/ssc, ohio/nlsy and michigan/icpsr); • both central and devolved models of data support appeared to function and coordinate well (according to interviewees), and were associated with high levels of researcher (and support staff) satisfaction; and • all seven researchers (all experienced professors) interviewed had recently visited research universities in the uk (typically, the lse and one or two others), and expressed some dismay at the poor level of data and it support facilities for researchers although conventional uk academic library facilities were rated highly. teaching universities provided data services through the library and it/computer services. the vp of one university described an “innovative plan” for creating information (and data) support teams attached to academic departments and managed by the library service. each team would comprise a “subject librarian”, a “computer/network adviser” and an “avgraphics-teaching resources manager”. overall, although a few data staff had major reservations, this sort of reorganisation the clio model was expected to be a feature of the is future in teaching institutions. of nine such support staff interviewed, all looked forward to re-defined jobs, some with enthusiasm (5) others with apprehension (4). resources and costs associated with data support facilities (it infrastructure) all respondents seemed aware of the major time-savings enabled by technological advances. in particular, researchers were keen to cite benchmarks for various modelling and statistical tasks. the following examples are typical of a dozen or so proffered. these were provided by an industrial relations researcher (lse and dti) and a trade/productivity research economist at esrc’s cep: level of data basic data intermediate full data support service data service service research 4 9 3 universities (17) research 1 3 4 centers (8) 20 iassist quarterly * sun will run two identical jobs in less than twice this time (actually, 51 secs). the following times correspond to the running time of the same gauss program solving a non-linear equation system for a grid of points. * unix times cannot be guaranteed if multi-user. typical hardware platforms found in the research centers included several unix boxes (hp, ibm and sun) and about one 486/pentium per fte researcher (excluding part-time postgraduates who typically shared pooled 386/486 facilities). • novell (stable), nt (expanding) and unix (stable) network servers were typical 486/pentium servers (1 gb to 5gb) and unix cluster (5 to 50gb) were typical storage capacities • a typical mid-sized research center had 52 dos/windows pcs, 10 macs (classic), 3 unix, 1 vms and 2 novell servers supporting about 200 postgraduate students and 40 fte research staff • use of campus-wide email (with approx. allocation of 1mb space per user) was typical in research centers as opposed to own email installation • most common/popular packages were wordperfect/word, netscape/mosaic, elm/pine/cc:mail, gauss, excel/123/ quattropro, sas/spss/stata little evidence of the use of programming languages such as fortran and c. most interviewees reported major changes in the pattern of it/computer support. for example, the central program advisory service, still prevalent in many uk universities, had all but disappeared in the us institutions in the fulbright study. typically, programmers had been relocated and redesignated as departmental or cluster it support staff. in the department, experienced programmers were often expected to provide a wide range of skills, covering software and hardware installation as well as teaching support duties. some common responses to this major structural change were as follows: the majority of “ex analyst-programmers” experienced what they saw as a “deskilling process” a minority were optimistic about the challenge/value of learning new skills. the majority of this group expressed disquiet over “cost recovery” policies, and some expected this to lead to their extinction. some data staff felt that lone researchers in particular had lost a valuable resource, the program advisor. nearly all data staff said they now found it necessary to provide some programming support for basic data management tasks typically, sas, spss and stata. just over half of all data staff interviewed (i.e. 14 of 27) appeared familiar with one or other of these programs most of these said they had always seen data management programming as part of their remit, although they also reported less demand for detailed program advice. year stata v3 machine cost (new) time (approx.) 1985 pc-xt $=£1,400 6,000 secs 1990 386dx20 $=£2,000 500 secs 1993 486dx33 $=£2,000 90 secs 1995 pentium90 $=£2,000 23 secs 1996 pentiumpro200 $=£3,000 7 secs 1996 sun model20-71 $=£7,000 30 secs* year gauss machine cost (new) time (approx.) 1993 486dx66 $=£2,200 68 secs 1995 pentium90 $=£2,000 19 secs 1996 dell latitude $=£2,600 17 secs (laptop p120) 1996 1996 pentiumpro $=£3,000 7 secs 1996 sun model20-60 $=£7,000 20 secs* 21winter 1996 in turn, nearly all it support staff (23 of 26) reported concern that their skills needed to be upgraded (19), or had already been upgraded (4), to cope with new information and data management tasks. many computer staff reported a major decline in the demand for their programming skills, and some said they had stopped all programming activity “many years ago”. the following it/computer issues were cited frequently: • around one third of data staff stressed that “computer skills were a basic requirement for data support staff” (10d=10 citations) • some it support staff (below managers) were concerned that it managers were not offering appropriate/relevant training for it staff, particularly data skills (5c) • the use of public pc facilities by students was often 100% with queuing at peak times, indicating that demand exceeded supply although some universities experienced a decrease in use of public facilities, as students stayed off-site (long journeys, bad weather, good support for modem links or local networking, etc. encouraged purchase of laptops) the 1997-2000 it outlook as predicted by richard rockwell (iassist, 1993), more powerful personal desktop pcworkstations (running unix and windows nt/95), have continued to enable researchers to process large-scale datasets extracted from local and wide area networks. all respondents in the fulbright study expected continued performance improvements in desktop processing and data management, enabling further gains in research output. there continues to be general optimism about the contribution of it hardware and software advances and their contribution to greater research productivity. the demand for it support of remote laptop and home computing (distance learning) is expected to increase, stimulated by the growth in quality teaching software. in the light of so much concern expressed about reorganisation of support services, we may expect a professional review of all research support services. it, library and data support staff will take the initiative (data piloted) and produce a more useroriented information service. data consultancy/services (local data support) according to researchers and data staff, the delivery of data support has become far more proactive. all data staff provided examples of “going out there” to find datasets, to advise on the best use of the data service (and other larger data facilities) and to help construct enhanced services through interlinked web pages. library based data staff were most enthusiastic about the contribution of cd-rom based datasets; some others, particularly experienced data support staff, seemed sceptical about the ultimate value of this form of data dissemination. researchers, closely followed by everyone else, were perplexed by the management and demarcation cd-rom data (typically supported by the library service) and data on other media (typically found in data centers). this was generally put down to some form of historical determinism, and there was little evidence of plans for change in this respect! researchers and data staff reported that internet type enhancements to data services had become expensive to maintain. expectations were high, following the early lead taken (voluntarily) by data staff in constructing useable interfaces to datasets. web weavers reported time costs between 2 hours and two days per week for basic to comprehensive coverage of data services. much of this work had been undertaken without additional funding. data and research staff had become de facto web advisors. invariably, data support staff expressed “grave concerns” about data security and quality particularly, in environments of decreased it support services. the following data issues were cited frequently: • researchers in research groups/centers appeared less interested in programming support although lone researchers still 22 iassist quarterly needed help (4d+1r=5 citations) • researchers seemed more concerned about quality of data accessibility, particularly with respect to speed (21r+10d+6c=31 citations) • data staff and some it staff were troubled by the ease of passing on large-scale undocumented (or poorly documented) datasets (22d+12c=34 citations) • a small number of experienced researchers were concerned about a possible decline in the quality of data analysis due to the trend to increase in accessibility (2r) • european data could be difficult to locate, and often impossible to acquire (8d+9r= 17 citations) • the majority of researchers prefer to download data directly to their personal machines, using their own data checking skills (18r) • some data bureau seem reluctant to develop user services so icpsr (good at data checking) were playing an essential role (5r) • the majority of researchers preferred to download entire file rather make “front-end decisions” even in the case of very large datasets (21r) • researchers were keen to support the central university data repository saying good local availability was important to research productivity (23r) • about half of the experienced researchers interviewed said they liked to send their research assistants to the it/data center the other half tended to seek assistance directly from it and data staff as appropriate the 1997-2000 data outlook general expectation of increased researcher self-sufficiency (with network infrastructure and data support) in programming/computing tasks. alongside this, more effort in enabling access to datasets through the internet. later rather than sooner (evidence suggests) someone brave will pull cd-rom data together with other data media. expectation, in the long run, of large investment in distributed data services via web/internet economists/accountants will work with data staff, network/communications staff and higher education planners to produce properly resourced infrastructure. in the mean time, data staff will continue to produce prototype web data servers without proper funding, and to experiment with the linking of datasets, documentation and bibliographic information. many data staff will change from being de facto web advisors to de jure information managers. data archives and data services • a significant number of data support staff (and others) were concerned about small-scale institutions in particular, their inability to manage and afford big datasets (7d+5r+4c) • even in larger institutions, data staff said that if funding problems persist, central archives such as icpsr would become still more important (3d) • a few it support staff and researchers stated their preference for getting data directly from central large-scale/national archiving (essex, icpsr, roper, etc.) which might assure the quality and security of important datasets (4c+2r) • some experienced researchers appeared keen to get data direct from source, and to bypass both local and national data services and archives (6r) • researchers and data staff based in specialist research centers expected to play a major role as data resource centers, claiming the “full set of research, data and computing expertise” at the necessary level of expertise to advise specialist research projects (6 of 7r + 7 of 7d + 5 of 7c) the 1997-2000 data archive outlook there was some expectation of devolution of large-scale data archives major research universities and research centers will negotiate to get data associated with their specialism direct from source. specialist research centers will work with major archives to distribute datasets, associated materials and expert advice to high level research projects. data archives will continue to distribute datasets to the majority of non-specialist institutions, and to provide some further “one-stop-shop” support for institutions unable to resource a local data service. data archives will combine with national data services and social science information gateways to lead the management and coordination of specialised data services and 23winter 1996 associated expertise based at universities and research centers. at an international level, they will plan and manage the network of specialist data servers, and they will jointly work towards making national datasets statistically comparable. structural/organisational trends whilst the majority of researchers and data staff supported developments in the integration of data services and libraries, some had major reservations citing loss of autonomy, deskilling and reduced service as likely outcomes. it support devolution and cost recovery continued to alarm support staff and over-exercise is managers. it appears that every institution has completed or is considering a major reshaping of teaching and research support services. most researchers (17) supported the development of “one-stop-shopping” i.e. the integration of it and data services although some (4), typically experienced, researchers questioned whether extending the ranges of skills might dilute the expertise. one experienced researcher and a few (4) data staff viewed integration plans as “cosmetic”, and counter-productive in that experienced data staff were likely to be lost or become disaffected in the transition. integration of data and library catalogues was fully supported. about half of data staff reported that datasets and library records “have been or are being fully integrated”; a third said it was “being planned”; and the others said they “expected integration of catalogues to happen soon”. most researcher-teachers (18 of 21 interviewed) supported “the trend” to deliver research (project) based courses to undergraduates using “real datasets”. some “research-led teachers” discussed the need for a more effective information systems structure to enable appropriate support for courses which required a range of data and information inputs together with more advanced information management and processing techniques cf. courses which employ artificial intelligence methods13. in this scenario the popularity of the one-stop-shop was very evident. network and communications remain central services albeit with evidence of growth in the number of local servers. the move towards full integration of voice and network services continues reaching over 50% in research universities14. about half of the institutions visited charged for ethernet (or token ring) connections, and some added a rental charge $130 for network installation (+ $5 additional annual rental in some) was typical. there were reports of an increase in (hitherto flatish) demand for remote computing which it support staff expected to further stretch their reduced (in real terms) resources. teachers, students and researchers are already expecting computer advice from remote locations i.e. from home, conference locations, etc). the 1997-2000 structural outlook the continued devolution of large-scale central it services seems likely, although a few of the very successful central systems should be able to construct a professional/economic case. professional groups such as iassist and css will cooperate to produce documentation of “models of successful research support systems”. joint work on teaching support will enable teachers to deliver remote/distant learning courses using real datasets extracted from central archives (for general/introductory courses) and specialist data servers (for advanced courses). problem areas data staff were seriously concerned over a number of data security issues. all experienced staff (23 of 27) said they had initiated (8 of 23) or were initiating (10 of 23) or would/should initiate (5 of 23) procedures for data checking in light of bad dataset transfers. large scale data transfers using ftp were commonly cited as error prone15, and bad windows transfers (via file manager, particularly from cd-rom) and tape backups were also reported. the following problems were cited frequently: 24 iassist quarterly • proliferation of forms (8d) • a few staff mentioned the importance of getting away from the “format statement”, which was seen as problematic for researchers and time consuming for support staff (3d+1c+2r)) • variable extraction via web interfaces was generally supported but there were some fears that speedy extraction would mean misuse of data, particularly if “data alerts” were not built into the system (5d) • researchers complained about the work overload at icpsr which meant they had to plan for up to eight weeks delay from data order (via icpsr and local data library) one researcher said she advised colleagues to “order data on the expectation that you may need it!” (6r) the 1997-2000 problems outlook expect “contents to check contents”, i.e. auto check for data consistency. data users to feedback errors though speedy “feedback system”. overdue replacement of paper forms by electronic forms. data staff will make a professional case for greater investment in the integration of metadata with datasets, and experienced researchers will advise web data server designers on the attachment of appropriate data documentation to subsets. success factors and performance indicators associated with data support all researchers interviewed showed a keen awareness of research technologies and their contribution to research productivity. about half said they were sceptical of windows style guis, but all research respondents said that productivity gains from advances in operating systems (for example, multi-tasking and large memory management) had made a major contribution to their (and others) empirical research. about a third (9 of 30) volunteered detailed benchmark figures consistent with those cited earlier in this paper. most respondents said that the quality of output was higher due to both it and data support (roughly equally, when prompted). two respondents (senior/experienced researchers) said their own research benefitted very little from local data services and a great deal from it support -system programmers advice. in one case, local data services did not feature at his institution he tended to use highly skilled systems programmers to assist with data management tasks. six researchers argued that their research could not be undertaken properly without assistance from local “highly skilled data support staff”. as might be expected, productivity issues cited by data staff invariably reflected the content of the recent iassist newsletter/journal coverage. the importance of documentation featured in all interviews. content analysis of respondence to an open-ended question on “what matters most?” shows the following recent articles to be representative of the range of issues cited: for example, general issues covered in rasmussen, 1995, on-line codebooks by sheih 1995, quality and accessibility by beedham 1995, production by winstanley and standards by greene 1995. it/computer support staff were much more likely than researchers and data staff to mention shortfall in training, both in terms of their own needs and the requirements of end-users. data staff were most concerned about getting additional resources for new developments such as the delivery of subsetting services and related documentation. the fulbright study showed a strong association between high levels of local data support and good performance16. nearly all researchers were keen to empathise with the data services view that local data services are a key factor in the production of highly cited research publications. unprompted, over half of all respondents expected data support to be a major component of developing teaching methods, particularly new and redesigned undergraduate courses. there are also strong a priori grounds for associating data support with good research performance. the evidence for growth in quantity and quality of empirical research is very strong, and it is also clear that academics are rewarded for research performance measured by publications and citations. the variables that most distinguished the academics in the sample who had been promoted from those who had not included rate of publication in refereed journals, level of citation, research grants applied for and obtained and the number of phd students under a person’s supervision. likelihood of promotion was correlated negatively with self-reported commitment to teaching17. 25winter 1996 the 1997-2000 resource outlook expect the organisation of local research support to be investigated more rigorously with a view to expansion in light of its proven contribution to research and teaching productivity. some economic conclusions one thing is for certain in this study: researchers, data services managers and it staff all feel the funding future to be uncertain. while they may have clear visions of the what the direction of research support services ought to be, they are nervous about policy-making. the bad news is that this concern is well-founded. most researchers and virtually all research support staff (outside the library) are badly placed (in the order of things) to make a big impact on resource policy. and, as virtually every interviewee in the study has mentioned, failure to compete professionally for the appropriate level of resources will not enable their data utopia to become reality or even virtual reality. the quite good news is that they do have lots of real data on the productivity benefits of data support to enable the construction of a strong case for further investment. they also have the skills to disseminate this evidence. what is required is a framework for evaluating the contribution, and for this they may need to find time to review the small but growing literature in information economics. recent work on the economics of the internet and on the contribution of it to business may provide some clues as to what to look for and what to measure in the context of research inputs. as in other areas, the returns to investment in research support services can be measured in terms of productivity changes, performance and consumer benefits. a number of recent research papers claim that investment in it is associated with increased productivity, increased consumer benefits but unchanged business performance18. according to brynjolfsson (1993), these results are compatible with conventional economic theory, i.e. “... firms are making the it investments necessary to maintain competitive parity but are not able to gain competitive advantage”. productivity gains to business and benefits to consumers due to investment in it have been found to be strong. however, the impact of it on business performance seems to be slight, sometimes negative. it appears from a stream of it literature on business performance (1989-1993, cited by hitt) that firms are unable to increase their profits through it investment; indeed, while it may be creating enormous value, it may simultaneously be intensifying competition and enabling entry, and thus lowering prices. the really good news for data support services is that their contribution to empirical research is truly, widely and deeply recognised. it is time to invest some of the energy and enthusiasm of research and teaching support services into the production of an empirically based case for expansion. references url references: biron: esrc’s data archive system for data searching on-line at essex university: http://dawww.essex.ac.uk/biron.html ccs: mit’s center for coordination science at mit: http://ccs.mit.edu/ccswp190.html cep: esrc’s centre for economic performance based at the lse,: http://cep.lse.ac.uk/ ibss: esrc/jisc’s lse based international bibliography of the social sciences at bids, via http://www.niss.ac.uk/ or http://www.lse.ac.uk/ restricted access 26 iassist quarterly icpsr: inter-university consortium for political and social research: http://www.icpsr.umich.edu/icpsr_homepage.html jisc (1996): policy for jisc dataset services provision at niss, via: http://www.niss.ac.uk/it/jiscdatapol.html nber: national bureau for economic research pwt data subsetting service: http://nber.harvard.edu/pwt56.html ssdc: university of california at sd’s social science data center,: http://ssdc.ucsd.edu/ restricted access to some dataset services bibliographic notes: anderson, r.e. and coover, e.r., 1976, “academic social research organisations and computerization”, social science information 15 (4/5): 741-754. anderson, r.e., 1987, “females surpass males in computer problem solving”, journal of educational computing research, vol.3(1). brynjolfsson, e., 1993, “the productivity paradox of information technology: review and assessment.”, centre for coordination science, mit. cave, m., hanney, s., kogan, m. and trevett, g., 1989, the use of performance indicators in higher education, jessica kingsley, london. centre for economic performance, 1993, unpublished report to esrc datasets policy committee, prepared by layard, r., lubanski, a. and wadsworth, j. on behalf of the centre for economic performance, lse. figlio, d., “trends in the publication of empirical economics”, journal of economic perspectives, 8 (summer 1994): 179-87. gaspar, j. and galeser, e., 1996, preliminary draft, communications technology and the future of cities, stanford, harvard university and nber. hitt, l. and brynjolfsson, e., 1995, “productivity without profit? three measures of information technology’s value.”, mis quarterly. (see ccs under url references above). munson, j.r., richter, r.l. and zastrocky, m.r., 1994, cause id 1994 profile. oswald, a. (1991) “progress and microeconomic data”, the economic journal, 101 (january 1991), 75-80. over, r., 1993, “correlates of career advancement in australian universities.”, higher education, vol.26, no.3, pp.313-329 platt, j., 1996, “has funding made a difference to research methods?”, sociological research online, vol.1, no.1,<http:// www.socresonline/1/1/5.html>. ramsden, p. (1994) “describing and explaining research productivity”, higher education, v.28, n2, p 207-226. ruus, l.g.m., 1990, “planning a data service facility”, data library service, university of toronto, 30/11/90 (internal document). 1.paper presented at the annual meetings of iassist, may 15, 1996, minneapolis 2 for example, in some areas the requirement for help with acquisition of data documentation has been reduced due to investment in on-line services. 3 esrc annual report, 1994-95, december 1995. 27winter 1996 4 from cause id survey 1994. in the uk, the esrc and jisc have recently produced their respective policies for dataset services see url references. 5 from cause id survey 1994, supplemented by evidence from my own study see fulbright report in url references. 6 verified from ncds studies at university of sussex, 1994. 7 refer to bibliographic notes 8 refer to url references 9 refer to bibliographic notes hitt and brynjolfsson, 1995. 10 as defined by laine ruus’ ‘planning a data service facility’, in ruus (1990). 11 op. cit. 12 op. cit. 13 for example, richard freeman’s new economics course at harvard university. 14 cause id 1994 profile. 15 a cep researcher reported 5% failure rate (identified in data checking) in ftp transfers of some 60 data files of between 5 and 25 megabytes. 16 as indicated by publication and citation rates. see kogan, m., et al, 1991, in bibliographic notes below. 17 see over, r., 1993, in bib. 18 hitt and brynjolfsson, 1995. microsoft word 49-1-golden-final (1).docx 1/12 golden, madison (2025) adaptive data governance for research data management, iassist quarterly 49(1), pp. 1-12. doi: https://doi.org/10.29173/iq1128 the creative commons-attribution-noncommercial license 4.0 international applies to all works published by iassist quarterly. authors will retain copyright of the work and full publishing rights. adaptive data governance for research data management madison golden1 abstract the field of research data management librarianship has grown significantly in past years but continues to face the challenges of knowledge gaps, frequent changes to policy and guidance, and the complexity and context that comes from data that varies both in type and format. as a research data librarian, i face these issues on a daily basis and have adopted an adaptive approach that combines multiple styles to balance the individual needs of researchers while complying with policies and best practices. this approach was adopted from my past experience in data governance at a corporation in which we faced the same core challenges. incorporating the four styles of data governance as laid out by gartner provides a framework for librarians and data governance specialists alike to prioritize competing needs and guide researchers through the data lifecycle. the benefits of this approach include increased flexibility in data management practices, continuous improvement of services and resources, efficiency, and empowerment of researchers and related stakeholders. keywords research data management, data governance, adaptive data governance, data librarianship introduction while research data management support in academic libraries is becoming an essential service, work remains to be done to balance priorities and offer the most impactful services possible with the resources available. policies and standards emerging from various organizations, training gaps, and the complexity and context of academic research datasets create a challenging landscape for researchers and librarians alike (sheikh et al, 2023). these challenges are best met with a flexible and institution-specific approach that can balance individual researcher needs with policies and best practices. prior to becoming a research data librarian, i had three years of experience in data analytics and data governance at a large insurance corporation. during this time, analysts and data engineers encountered many of the same challenges as researchers; complying with policies, managing sensitive data, and documenting data workflows. the data governance team was created to address these issues and included ten people. we primarily collaborated with data engineering teams to apply security, classify data, and standardize metadata for datasets across the organization. additionally, we worked with data analysts to understand the data they were subject matter experts in, as well as, assisting them in finding the appropriate data. we trained all groups working with data on the company’s data policies and the data catalog we managed for the company. members of the team 2/12 golden, madison (2025) adaptive data governance for research data management, iassist quarterly 49(1), pp. 1-12. doi: https://doi.org/10.29173/iq1128 also coordinated a data stewardship program to expand our reach and receive continuous input from across the company. throughout these efforts, my manager employed the adaptive data governance approach to balance the fast-paced and goal-oriented approach of the company’s leadership with the longer term goal of achieving proper data management. this paper proposes how gartner’s adaptive data governance model can be mapped onto the field of research data management librarianship. gartner is a research and consulting firm located in the united states that focuses on business and technology. according to their website, gartner specializes in ”actionable, objective insights”, along with ”expert guidance and tools” (gartner, 2025). they have been in business for over 40 years and work with businesses in nearly 90 countries and territories (gartner, 2025). this approach will assist data librarians and those in related positions to balance policy and best practices with the individual needs of researchers and departments, as well as develop services and resources as data management maturity grows on campus. there are several existing guides and frameworks for developing research data management services. the digital curation center developed the research infrastructure self-evaluation (rise) framework in 2017 to “facilitate rdm service planning and development at the institutional level” (rans and whyte, 2017, p.3). oclc research also developed a three-part research data management service guide in 2018 covering how to understand local needs, identifying incentives for various services, and whether to create or buy identified services (bryant, 2018). additional institution-specific case studies and frameworks regarding the development of research data management services were also developed. examples include oxford university, central washington university, and the university of toronto (chiarelli et al, 2022; fu et al, 2022; perrier & barnes, 2018). while these guides and frameworks address similar issues of institution-level needs assessment, service maturity, and the costs versus benefits of various services; the model presented in this paper offers a new framework for categorizing and evolving services over time by utilizing multiple styles of governance. this paper will rely on my experiences as a research data librarian at the university of utah. as such, generalizability of this work is limited and requires further research and application at other institutions. for context, the university of utah is an r1 research university with approximately 26,827 undergraduate students and 8,409 graduate students as of 2023. there are 1,592 full-time tenure-line faculty and 1,863 full-time career-line faculty. in the 2023 fiscal year, the institution received $768 million in research funding (university of utah, 2024). prior to august of 2023, there was a brief gap in research data support from the library which was filled by myself and one other research data librarian faculty member. these hires and subsequent reinstatement and expansion of research data services were a direct result of the ostp memo of 2022 and an increased focus on open science across campus. in addition to the library, research data infrastructure comes from the center of high performance computing, guidance from the vice president of research office, and data science initiatives. methodology at its core, this is a conceptual article which utilizes theory adaptation to expand upon gartner’s adaptive data governance model, wherein four styles of data governance are used simultaneously to meet variable needs. because of this, the paper relies on gartner research’s methodology. gartner states that they “offer a full range of research methods such as in-depth proprietary studies, peer and industry best practices, trend analysis and quantitative modeling” (gartner, 2024). their proprietary 3/12 golden, madison (2025) adaptive data governance for research data management, iassist quarterly 49(1), pp. 1-12. doi: https://doi.org/10.29173/iq1128 methodology was developed and utilized by their global experts numbering over 2,000. in regards to their objectivity, gartner cites their strict code of conduct and guidelines to ensure neither the company nor its employees are in a position to benefit from any one company, industry, or technology (gartner, 2024). in addition, they were ranked by forbes as #92 on their america’s best companies list for 2025 (forbes, 2024). theory adaptation as a methodology “develop[s] contribution by revising extant knowledge—that is, by introducing alternative frames of reference to propose a novel perspective on an extant conceptualization” (jaakkola, 2020, p. 23). according to jaakkola, this necessitates identifying a theory of interest, adjusting or expanding the original theory’s scope, and justifying why the shift is needed, including why the selected theory is the best fit (2020). in keeping with these requirements, i will first identify and explain the theory of interest, which will be gartner’s adaptive data governance framework. next, i will justify why the shift is needed by identifying shared core challenges of data governance and research data management, as well as how these challenges are addressed by the framework. finally, the shared core challenges along with the mapping of research data management services and their iterative development into the framework will be used to argue the efficacy of this theory adaptation. as with all conceptual articles, the work has not been empirically proven “but rather build[s] on theories and concepts that are developed and tested through empirical research” (jaakkola, 2020, p. 19). as such, this paper will also include limitations and a call for further evaluation and critique. data governance overview data management association international (dama) defines data governance as “the planning, oversight, and control over management of data and the use of data and data-related sources” (dama, 2017), typically within companies and government organizations. this includes a wide range of activities, including data architecture, data modeling and design, data storage and operations, data security, data quality, metadata, data warehousing, reference and master data, document and content management, and data integration and interoperability (dama, 2017). it also includes collaborating with data scientists, data analysts, and business leaders to coordinate efforts and align goals. in a fast-paced business environment, particularly those motivated by growth, data governance can often be seen as a block to progress rather than an enabler of it. this is due to a tendency in traditional governance towards reactivity and a top-down approach. however, if data is not properly organized, managed, and protected, there will be larger complications down the line whether they be from legal action resulting from not adhering to policies, data breaches, or consistent duplication of effort and poor resource management (abraham et al, 2019). adaptive data governance is an approach to data governance created to alleviate some of the issues with traditional governance styles described above. at the heart of the approach is agility, which allows data governance to accomplish its goals while encouraging innovation and growth (gulzar and kopcho, 2024). adaptive governance is achieved by combining multiple styles of governance, which will be outlined in the following sections. the use of multiple styles also allows for data governance teams to handle more complexity and disparate needs across an organization (rama, 2013). in addition, this approach is complementary to the rapid changes and development in organizations, enabling efficiency rather than preventing it. 4/12 golden, madison (2025) adaptive data governance for research data management, iassist quarterly 49(1), pp. 1-12. doi: https://doi.org/10.29173/iq1128 gartner, a leader in data and it management research, defined an adaptive data governance framework that incorporates four governance styles; control, outcome, agility, and autonomous (gartner). each style builds upon the last in maturity and enables an organization to handle increasing complexity. control is at the heart of the model and closely resembles the traditional approach to data governance in which compliance with policies and rules guides all work (gulzar and kopcho, 2019). in data governance at a corporation, this would take the form of data policies such as sensitive data policies and data access policies, along with government regulations and company or industry standards. the outcomes style of governance still incorporates the control of the previous style, but introduces analytics as a way to balance priorities and make informed decisions (gulzar and kopcho, 2019). here, standards and rules may be altered so business performance goals may be met while still adhering to policies and regulations. the agility style introduces increased flexibility by distributing empowerment to dispersed groups in an organization to accelerate decision making by placing power in the hands of subject matter experts (gulzar and kopcho, 2019). this stands in stark contrast to control where decisions are made from the top down and policies are created and managed at the upper levels of management. in this bottom-up approach, decisions can be made that balance individual teams’ needs and policies to move away from the ‘one size fits all’ approach. the final governance style in the model is autonomy, including distributed authority from the autonomy style, along with input from ai and other automated tools (gulzar and kopcho, 2019). the use of ai and algorithms increases the ability for complex decision-making by incorporating a variety of factors and real time data while still adhering to policies. the model is not designed to be unidirectional. rather, the goal is to situate services and tools for given data management activities within the governance style that best meets the needs of an individual organization at a given point in time. the governance style employed for a given activity, such as monitoring data quality can progress from one style to another as needs and overall data maturity of an organization change. for example, an organization may use the agile approach to monitor data quality wherein responsibility is dispersed to departments within a large organization. as the organization’s data strategy matures, they may pivot to the autonomous style of governance by utilizing an algorithm to monitor data quality instead. on the other hand, the introduction of novel policies such as the ostp memo in 2022 could necessitate moving data management practices that were previously governed with an outcomes or agile style to the control style because more oversight is needed to comply. at the university of utah, this was the case with data sharing practices. therefore, success of implementing the model depends on whether the organization is able to accurately understand the level of data maturity across the organization at a given time and what governance style is needed for data-related activities. shared core challenges: research data management and data governance research data is defined as “ the outcome of experiments or observations that validate research findings, and can take a variety of forms including numerical output (quantitative data), qualitative data, documentation, images, audio, and video” (national library of medicine, 2022). research data is similar to industry or government data in that it comes from a variety of sources, contains a combination of sensitive and non-sensitive data, and is used for a variety of purposes. it is unique in that there are additional variations in the data types and formats, including video, audio, code, and simulations. additionally, policies and processes come from a variety of sources such as government, funding agencies, and institutions. this means internal regulations, or lab-specific regulations and 5/12 golden, madison (2025) adaptive data governance for research data management, iassist quarterly 49(1), pp. 1-12. doi: https://doi.org/10.29173/iq1128 processes, tend to be looser and more disparate, even between research labs within the same discipline (reichmann et al, 2020). data management, similar to research data, is a very broad concept that encompasses "data management planning, documenting your data, organizing data, improving analysis procedures, securing sensitive data properly, having adequate storage and backups during a project, taking care of your data after a project, sharing data effectively, and finding data for reuse in a new project," (briney, 2015, p.7). the scope of data governance is slightly larger than data management, as data management activities are seen as a component of data governance. however, each works to achieve similar goals where proper data management improves usability of data while managing risks such as security breaches, data loss, and poor data quality. several core challenges are shared by research data management and data governance that can be improved with the adaptive approach. complexity and context create challenges due to a variety of data sources and types. variation in these areas creates difficulty in developing policies and practices that fit across data types. internal and external policies introduce complexity as they may have conflicting requirements. disparate standards and processes can be a barrier when introducing new policies or best practices, and can also create challenges for instruction development for a wide audience. additionally, understanding all of these differences across disciplines is unfeasible for many organizations. another shared challenge is a lack of awareness among data creators and users of policies, standards, and best practices. lack of awareness arises from training gaps, either within a research group or across an institution. there are also multiple sources of guidance from professional organizations, groups within institutions, and government and funding agencies. lastly, it is difficult to reach all applicable audiences, especially in a decentralized organization where data management is not seen as essential. finally, constant developments and changes across the data landscape contribute to the above challenges of lack of awareness and handling complexity and context. updates to policies and regulations may go unseen by affected parties unaware of what applies to them. this includes the development of data sharing requirements, either from states like california issuing the california consumer protection act, or from research funders and publishers. in addition, the introduction of new tools, systems, and methods may make current policies and instruction inadequate. adaptive data governance for research data management in order to explain how the adaptive data governance model, as laid out by gartner, can be applied to research data management, each of the four governance styles described earlier will be recontextualized for research data management within an academic institution. as described in the introduction, this will primarily stem from the library perspective and examples from the university of utah. 6/12 golden, madison (2025) adaptive data governance for research data management, iassist quarterly 49(1), pp. 1-12. doi: https://doi.org/10.29173/iq1128 figure 1: conceptual map of linear relationship between service modality complexity and maturity needed to achieve each adaptive governance style control in data governance, the control style is defined by policies and adopts a passive, compliance-based approach (judah, 2023). to re-define this style for research data management; the control style of research data management seeks to mitigate risks by monitoring compliance to applicable research data policies and standards arising from funding organizations, publishers, government, and institutions. a core drawback to this approach is strict adherence to policies that are often not defined by those in the library or other academic offices supporting researchers. however, these policies and laws create the jumping off point for service development and resource acquisition. clear requirements for research data management in the form of policies create clear service priorities and an opportunity to perform outreach to researchers. additionally, control establishes the need for a compliance and policymaking body within an organization, which can serve as the basis for further infrastructure and support staff for researchers. the research office serves as the primary policymaking and compliance group on campus. however, in contrast to a company, institutions have more disparate compliance processes built in. for example, grant offices assist researchers in reporting on compliance to their funders and researchers are also accountable to comply with policies on data sharing set forth by publishers. regardless of this slight decentralization, the main mechanism driving data management in this context is compliance with policies and regulations created in a top-down manner. at the university of utah, examples include assisting researchers developing data management plans, selecting appropriate repositories for data sharing, and advising on infrastructure development to support these requirements. in terms of benefits, interfacing with researchers under this style has the dual benefit of incentivizing researchers to seek support from the library to comply with policies regardless of their origin and gives the library insight into key support needs of researchers. assisting researchers complete data management plans (dmps), selecting appropriate repositories, and storing data securely opened conversations on common needs and questions. some examples included difficulty describing data and generating metadata, difficulty selecting a repository or knowing what characteristics to look for, and confusion over where to seek guidance on research data management. these conversations prioritized service and resource development in the following governance styles and created educational opportunities that matured many researchers’ understanding of research data management. 7/12 golden, madison (2025) adaptive data governance for research data management, iassist quarterly 49(1), pp. 1-12. doi: https://doi.org/10.29173/iq1128 outcomes the outcomes style relies on balancing risk and performance to meet the policies outlined in the control style while still prioritizing efficiency and business goals (judah, 2023). in the context of research data management, this style supports research data management through the development of services and resources across multiple modalities, informed by the needs and gaps of an institution’s researchers as well as relevant standards and requirements. utilizing learnings from the previous section, or from a more structured approach such as a survey, serves as the foundation for service and resource development within this governance style. understanding certain departments’ level of need for services, as well as the modalities most requested, can guide the need for additional support teams. additionally, resources and standards aimed at meeting the goals of data policies should be introduced. for example, the fair principles assist researchers in meeting data sharing policies while improving the findability, accessibility, interoperability, and reusability of those shared datasets. at the university of utah, there was a clear value to and need for developing an institutional data repository to assist researchers in meeting data sharing requirements quickly and freely. this gives researchers greater flexibility when depositing data and we are able to adapt different aspects of the deposit process and standards to meet their needs. for example, researchers fill out a qualtrics form when requesting to deposit which we use to generate a readme on their behalf. usually, one section of the form is writing out a codebook. defining variables in a codebook is essential to future users understanding a dataset, which is complimented by methodological information included in a readme document. however, we have had several researchers with over 100 variables who we allowed to upload a codebook separately which is referenced in the readme file. remaining open to these small adjustments is essential for research data management as the amount and format of data (and metadata) vary widely. for this and the following styles, it is important to focus as much on what an institution’s researchers don’t need as what they do need. a key part of the outcomes style is recognizing what efforts and resources are not yet necessary and would take away from needed support. on the flip side, analysis of current offerings may show a gap in services. for example, the university of utah’s research data repository recently increased its retention period from five to ten years due to longer retention requirements in federal funding and publisher policies. going forward, analytics on downloads and citations metrics will be used to make deaccessioning decisions at the end of that period. other examples include deciding whether or not to hire additional staff for data curation support, paying for additional storage space for data deposits, offering coding or software support, and providing in person or on demand workshops. given the ample online resources from academic institutions, professional organizations, and other groups, investigating free and already available content should be top of mind. agility the agility governance style emphasizes placing decision making power in the hands of subject matter experts to allow decisions to be made efficiently. it also prioritizes on demand resources and services to support individuals and teams. in the context of research data management, agility empowers researchers and supporting staff to perform research data management using self service tools that align with their individual needs. as shown in the model, additional data management maturity is 8/12 golden, madison (2025) adaptive data governance for research data management, iassist quarterly 49(1), pp. 1-12. doi: https://doi.org/10.29173/iq1128 needed at an institution to enable localized decision making and the use of self service tools. some challenges referenced earlier, particularly addressing training gaps, will need to be addressed by the following styles prior to introducing agile approaches. at the university of utah, this is the stage when research data management responsibilities were dispersed across the university between the libraries, research office, it, and researchers and research labs. structurally, the library provides access to training and resources while the research office handles policy making and compliance, it handles data security and classification policies along with data storage for in-progress research, and researchers along with grant officers are accountable for compliance and reporting for any policies affecting their work. in this system, various groups are able to decide what gaps on campus they have capacity to fill and how those resources are made available. collaboration and pooling of resources can also be done to complete larger projects, such as the center for high performance computing offering additional storage space that the library-managed repository cannot accommodate. the library specifically offers several self-service tools to support agility and empower researchers who are more comfortable with the data management process. dmptool is an online tool that allows users to access data management templates and guidance based on the institution they belong to or their funder. libguides are another on demand resource with comprehensive information on the research data lifecycle and reusing research data, including a repository selection tool. while the ultimate goal is to provide self service resources when possible, it is important to consider areas where this may not be the best fit. for example, we have noticed many researchers are unfamiliar with building readmes and generating metadata using controlled vocabularies or standards. as such, it makes sense to retain control of the deposit process in our institutional data repository rather than a self-deposit model. this is a great example of not having the required level of maturity on campus for an agile or autonomous style for this service. it is also an example of how a blended style approach meets institution-specific needs. autonomy autonomy requires the highest level of maturity but is able to handle the most complexity by combining the power of automated tools and agile decision making. this style allows practitioners to manage research data via individuals and automated tools attuned to the researcher’s needs while complying with relevant policies and regulations. many tools and services that fit within this approach are novel and not yet widely used. examples include machine readable data management and sharing plans that will allow easier compliance monitoring after the grant cycle, additional automation for our repositories deposit process which would greatly speed up the process and reduce data-entry type tasks, and dataset curation or metadata creation done by ai to allow for more efficiency and a more discipline specific approach than we are able to provide currently. thus far, the university of utah has not reached the automation level. implementing autonomy style tools will require resources and technical expertise in addition to data management knowledge across campus. utilizing use cases from data management savvy researchers and possibly the assistance of grant funding to develop services will be key mechanisms for implementing this style. conclusion the primary purpose of this model is to guide research data librarians and related practitioners at research institutions on the development and prioritization of research data management services 9/12 golden, madison (2025) adaptive data governance for research data management, iassist quarterly 49(1), pp. 1-12. doi: https://doi.org/10.29173/iq1128 and resources over time. the model may also be useful to adjust services and resources as the landscape of research data management changes due to increased data literacy at a given institution, the introduction of new policies, and/or an increased commitment to open science. while existing case studies, guides, and frameworks provide helpful guidance in research data management service development, this adaptive data management approach provides a novel strategy to blend multiple approaches and scale services up and down over time. therefore, it provides the benefits of flexibility, continuous improvement, and efficiency and empowerment. flexibility is a core benefit of the adaptive approach to research data management. combining multiple styles facilitates institutions meeting the needs of as many researchers as possible. practically, this can look like offering training and informational content in multiple modalities such as videos, text, consultations, and live presentations and workshops. it also takes the form of accommodating individual needs in various processes such as depositing in an institutional repository. as mentioned in a previous section, research data varies even more widely than industry data with fewer standard practices within disciplines. this variation requires a flexible approach to assisting researchers in meeting policies and requirements rather than one size fits all processes and resources. the adaptive approach allows for that flexibility while maximizing the support provided through a mixture of internal and external services, information offered in multiple modalities, and prioritization based on the most pressing needs. continuous improvement is achieved through the adaptive approach by having a gradually maturing model built in. as training gaps are closed and more individualized services are requested across campus, more agile and autonomous management styles can be implemented. ensuring basic policies and requirements are met as a first priority creates an opportunity to educate researchers on a host of other topics including standards like fair, the use of metadata schemas and standardized vocabularies, how to select a data repository, and how to handle sensitive data. that knowledge can be built upon over time resulting in the use of self service tools like dmptool, freeing up time and resources to be put to developing additional tools or automating processes. receiving regular input and concerns from researchers also results in agility and re-prioritization over time. finally, efficiency and empowerment are encouraged through self service resources, automation wherever possible, and dispersed accountability. managing research data through a compliancebased mindset is often reactionary and can hamper innovation. moving away from this style to decentralize over time as expertise with data management grows across key groups such as grant offices, research administrators, and researchers themselves reduces the need for command and control efforts. investing resources across each style gives researchers more options for managing their data and aligns with the ‘teach a man to fish’ ethos common in libraries and academic institutions. the power of the adaptive data management approach comes from the centering of the researcher and striving to mold a set of practices and policies around their knowledge and the resources available. as stated at the beginning of this paper, this synopsis of the adaptive data governance approach and its application to research data management arise from my experience in both fields as well as the referenced articles, which are primarily from gartner. adapting research data management tools, services, and resources exists along two axes. the first is the chosen tools, services, and resources offered to support activities across the research data lifecycle. the second is the level at which those tools, services, and resources are offered. examples include synchronous or 10/12 golden, madison (2025) adaptive data governance for research data management, iassist quarterly 49(1), pp. 1-12. doi: https://doi.org/10.29173/iq1128 asynchronous, virtual or in person, self service or mediated, and informational or hands-on. as these decisions are further complicated by resource limitations and staff expertise, this adaptive model reimagined for research data management can assist in guiding and prioritization over time. additional research on how institutions with varying characteristics situate services and resources within the model shared in this paper would be instrumental in validating and understanding this application of adaptive data governance to research data management. limitations the primary limitations of this paper are its conceptual nature, reliance on gartner’s proprietary methodology, and use of examples derived from only one university. a conceptual article is inherently non-empirical, and therefore does not meet the requirements of a paper based on empirical research. further application by other institutions would be necessary to fully evaluate the effectiveness of this theory adaptation. secondly, while gartner claims to have high standards for independence and objectivity in their research, those methodologies are proprietary and were therefore not fully evaluated in this paper. finally, the examples and experiences used to adapt gartner’s adaptive data governance framework were based on one university. generalizability of this work relies on future application, evaluation, and critique by other institutions. references abraham, r., schneider, j., & vom brocke, j. (2019). ‘data governance: a conceptual framework, structured review, and research agenda’. international journal of information management, 49, pp. 424-438. https://doi.org/10.1016/j.ijinfomgt.2019.07.008 briney, kristin. (2015). data management for researchers : organize, maintain and share your data for research success. pelagic publishing. bryant, r., faniel, i., & lavoie, b. (2018). research data management: planning guide. oclc research. https://www.oclc.org/research/areas/research-collections/rdm/guide.html chiarelli, a., beagrie, n., boon, l., mallalieu, r., johnson, r., may, a.w., & wilson, r. (2022). to protect and to serve: developing a road map for research data management services. insights: the uksg journal, 35, pp. 4. https://doi.org/10.1629/uksg.566 dama. earley, s., & henderson, d., sebastian-coleman, l (eds.). (2017). ‘the dama guide to the data management body of knowledge (dama-dm bok)’. bradley beach, nj: technics publications, llc. forbes. (2024). gartner. forbes. https://www.forbes.com/companies/gartner/ fu, p., blackson, m., & valentino, m. (2022). developing research data management services in a regional comprehensive university: the case of central washington university. ifla journal, 49(6), pp. 443-451. https://doi.org/10.1177/03400352221116923 gartner. (2024). research and advisory overview. gartner. https://www.gartner.com/en/research/methodologies 11/12 golden, madison (2025) adaptive data governance for research data management, iassist quarterly 49(1), pp. 1-12. doi: https://doi.org/10.29173/iq1128 gartner. (2024). independence and objectivity. gartner. https://www.gartner.com/en/research/methodologies/independence-and-objectivity gartner. (2025). about. gartner. https://www.gartner.com/en/about gulzar, r. and kopcho, j. (2019). ‘succeed with digital business through adaptive governance’. gartner. (available at https://www.gartner.com/document/3975569?ref=authbottomrec&refval=4276799) this gartner report is archived and is used to provide historical context only. gulzar, r. and kopcho, j. (2024). ‘define principles for adaptive governance to quickly respond to change’. gartner. (available at https://www.gartner.com/document/4276799) jaakkola, e. (2020). designing conceptual articles: four approaches. ams review, 10, pp. 18-26. https://doi.org/10.1007/s13162-020-00161-0 judah, s. (2023). ‘2024 strategic roadmap for data and analytics governance’. gartner. (available at https://www.gartner.com/document/5028631?ref=solrall&refval=413150173&) gartner is a trademark of gartner inc. and/or its affiliates. the national library of medicine. (2022) ’research data management’. national library of medicine. https://www.nnlm.gov/guides/data-glossary/research-data-management. perrier, l. & barnes, l. (2018). developing research data management services and support for researchers: a mixed methods study. partnership: the canadian journal of library and information practice and research, 13(1). https://doi.org/10.21083/partnership.v13i1.4115 rama, d. (2013). ’adaptive data governance: the at-ease change management approach’. gartner. pp.164–189. https://doi.org/10.1201/b15034-12. rans, j. & whyte, a. (2017). ‘using rise, the research infrastructure self-evaluation framework’ v.1.1 edinburgh: digital curation centre. available online: www.dcc.ac.uk/guidance/how-guides reichmann, s., klebel, t., ilire hasani-mavriqi and ross-hellauer, t. (2020). ‘between administration and research understanding data management practices in a mid-sized technical university’. socarxiv (osf preprints). https://doi.org/10.31235/osf.io/75ac6 sheikh, a., malik, a., & adnan, r. (2023). ‘evolution of research data management in academic libraries: a review of the literature’. information development. https://doi.org/10.1177/02666669231157405 university of utah | university analytics and institutional reporting. (2024). fast facts 2024 [infographic]. data.utah.edu. https://data.utah.edu/wpcontent/uploads/sites/61/2024/11/fast-facts-2024-final-11.5.24.pdf 12/12 golden, madison (2025) adaptive data governance for research data management, iassist quarterly 49(1), pp. 1-12. doi: https://doi.org/10.29173/iq1128 endnotes 1 madison golden, assistant research data librarian, university of utah. madison.golden@utah.edu https://orcid.org/0009-0004-4993-3503 1/7 akmon, dharma and jekielek, susan (2019) restricting data’s use: a spectrum of concerns in need of flexible approaches, iassist quarterly 43(3), pp. 1-7. doi: https://doi.org/10.29173/iq941 restricting data’s use: a spectrum of concerns in need of flexible approaches1 dharma akmon2, susan jekielek3 abstract as researchers consider making their data available to others, they are concerned with the responsible use of data. as a result, they often seek to place restrictions on secondary use. the research connections archive at icpsr makes available the datasets of dozens of studies related to childcare and early education. of the 103 studies archived to date, 20 have some restrictions on access. while icpsr’s data access systems were designed primarily to accommodate public use data (i.e. data without disclosure concerns) and potentially disclosive data, our interactions with depositors reveal a more nuanced range of the needs for restricting use. some data present a relatively low risk of threatening participants’ confidentiality, yet the data producers still want to monitor who is accessing the data and how they plan to use them. other studies contain data with such a high risk of disclosure that their use must be restricted to a virtual data enclave. still other studies rest on agreements with participants that require continuing oversight of secondary use by data producers, funders, and participants. this paper describes data producers’ range of needs to restrict data access and discusses how systems can better accommodate these needs. keywords data archives, restricted access systems, privacy, confidentiality introduction responsible stewardship of data requires the ability to restrict access when data could identify individuals and potentially cause harm through their disclosure. at the same time, access restrictions by definition limit data’s use, so data archives must also apply restrictions judiciously, ensuring restrictions match the level of disclosure risk. icpsr4, a membership-based social science data archive, currently offers three main dissemination options for sensitive data: secure download, virtual data enclave (vde), and physical data enclave. each of these dissemination options imposes a stringent application process, but differ in where the data are accessed and what is required to access the data. icpsr’s stepwise series of dissemination options have been devised to serve sensitive data that vary in their probability of disclosing individuals and the potential severity of harm were individuals’ information disclosed. in this paper, we discuss icpsr’s options for restricting access to data using three examples from one of its topical archives: childcare and early education research connections5. these examples demonstrate some of the reasons for restricting access to data. in doing so, they also highlight the ways in which the need for restricting use does not always align with the systems and tools we have implemented for accessing the data. after describing the examples, we discuss design implications for https://doi.org/10.29173/iq941 https://www.icpsr.umich.edu/icpsrweb/ https://www.researchconnections.org/childcare/welcome https://www.researchconnections.org/childcare/welcome 2/7 akmon, dharma and jekielek, susan (2019) restricting data’s use: a spectrum of concerns in need of flexible approaches, iassist quarterly 43(3), pp. 1-7. doi: https://doi.org/10.29173/iq941 restricted access systems that will better serve the needs of the data, researchers, and study participants. icpsr and restricted access options icpsr, founded in 1962, is a consortium based at the university of michigan of over 760 member institutions around the world to archive and make social and behavioral science data available to researchers. icpsr holds over 10,000 data collections, approximately 1,670 of which have at least one dataset with restrictions on its use. in restricting access to data, icpsr aims to protect the confidentiality of study participants and ensure that the benefits of research outweigh the potential harms to individuals. information such as criminal activity, antisocial activity, and medical conditions could cause harm were it associated to particular individuals, who could be ascertained through direct or indirect identifiers. direct identifiers include name, phone number, social security number, and location, while indirect identifiers include information that can be used to identify a subject when combined with other information (for example, gender, birth date, geographic indicator and other descriptors). the only way to ensure 100% confidentiality protection is by blocking access to data. as data stewards, we must balance the tradeoff between the value of making the data available for others to use in new research and the responsibility to protect the confidentiality of study participants. the vast majority of data archived with icpsr are public use files: these are data that icpsr, often along with the data producers, have assessed to present very little, if any, risk of harm and/or probability of disclosure. a secondary user accesses these data through a simple web download: she logs into her icpsr mydata account, agrees to terms of use by checking a box, and after doing so, the data download immediately to her local machine. for sensitive data, icpsr staff use several approaches to maintain confidentiality: they modify data to reduce the chance of reidentification (e.g. through masking data); physically isolate the data and use secure technologies to provide access; train researchers in the responsible use of the data; and require researchers to agree to particular guidelines of use through restricted use agreements. typically, the agreement includes a responsible use statement, a research plan, institutional review board (irb) sign-off for the plan, a data protection plan, behavior rules (e.g. the researcher will not attempt to identify individuals and will not share the data with others), a security pledge, and an institutional signature on the agreement to use the data. data with moderate risk of harm and probability of disclosure (in other words, data with some risk of reidentification, disclosure, and non-trivial risk to study subjects), are generally offered via icpsr secure download. the secure download option requires a researcher to submit an irb-approved research proposal; a data use agreement signed by her institution; an agreement that the data will only be used within a very particular computing configuration (for example, on a stand-alone machine that is not connected to the internet); and an affidavit of the data’s destruction at the conclusion of the use period (generally one year, but the researcher can apply for extensions to the agreement). once approved, researchers receive an encrypted file via email, which they download to the secure location specified in the application. for higher risk data—that is, data that might be reidentified and also cover sensitive topics where disclosure could harm the study participants—icpsr offers both a virtual and physical data enclave. https://doi.org/10.29173/iq941 3/7 akmon, dharma and jekielek, susan (2019) restricting data’s use: a spectrum of concerns in need of flexible approaches, iassist quarterly 43(3), pp. 1-7. doi: https://doi.org/10.29173/iq941 each of these enclaves require a similar set of application materials as for secure download, however, the data can only be accessed and analyzed from within a highly restricted environment. in the case of the virtual data enclave (vde), researchers access the data through a virtual machine they launch from their own computer but that operates on a remote server. the virtual machine is completely isolated from the user's physical computer, restricting the user from downloading files or parts of files to their physical computer. the virtual machine is also restricted in its external access, preventing users from emailing, copying, or otherwise moving files outside of the secure environment, either accidentally or intentionally. to receive final output from the vde, the data must be vetted by icpsr staff. furthermore, icpsr has the capability to shut off access to the data when the terms of the data use agreement end and in the rare case of a user violating the terms of the data use agreement. data that are deposited in the physical enclave can only be accessed on site in ann arbor, michigan. the machines in the physical enclave are not connected to the internet, and an icpsr staff member is present at all times when a researcher is using the enclave. icpsr staff conduct a disclosure review of all files that the researcher wants to use after leaving the enclave. three examples of restricted-use data icpsr provides data dissemination for more than 20 federal and non-governmental sponsors via topical archives on topics such as addiction and hiv, aging, arts and culture, and criminal justice. icpsr’s childcare and early education research connections archive (hereby referred to as “research connections”) is funded through the office of planning, research and evaluation, administration for children and families (opre), u.s. department of health and human services and curates and provides access to over 300 studies. as an archive at icpsr, research connections has at its disposal the restricted-use access options described above and works closely with both the sponsor and data producers to ensure that confidentiality is maintained while facilitating the broadest use of data possible. working closely with study stakeholders has given the research connections project team the opportunity to understand the myriad concerns at play when making sensitive data available to secondary users. for the purposes of this paper, we discuss three studies from research connections that demonstrate the nuanced needs of restricting access to data. the head start family and child experiences survey the first study we discuss is the head start family and child experiences survey6, also known as “faces.” the faces study provides longitudinal data on a periodic basis on the characteristics, experiences, and outcomes of head start children and families as well as the characteristics of the head start programs that serve them. faces also provides information on the relationship among family and program characteristics and outcomes. several cohorts of faces have been fielded since 1997, and, through research connections, icpsr has curated and provided access to data collected from six cohorts of this study. https://doi.org/10.29173/iq941 https://www.researchconnections.org/childcare/series/236 4/7 akmon, dharma and jekielek, susan (2019) restricting data’s use: a spectrum of concerns in need of flexible approaches, iassist quarterly 43(3), pp. 1-7. doi: https://doi.org/10.29173/iq941 as study sponsors and data producers work with research connections to make their data available to secondary users, they are concerned not only with the potential disclosure of individual study participants, but also with the disclosure of particular centers (e.g. a specific head start center), since such information could potentially cause harm to the center and the families it serves. for that reason, icpsr works to ensure that both personal and center characteristics are non-disclosive. with all waves of faces, the direct and indirect identifiers have been removed from the data made available through research connections, leaving virtually no chance that individuals or centers could be reidentified. these protective actions might suggest treating this study as a public-use study, and that simple download, where nothing more than logging in and agreeing to icpsr’s standard terms of use is required, would be the most appropriate dissemination method. yet, because of the moderate risk of harm to the centers should they be reidentified, the sponsor and data provider were keen to place additional requirements on researchers that want to access and use the data. they wanted to both lightly screen potential researchers to ensure that they are using the data for legitimate research or public policy purposes and create a means of addressing data misuse should it occur. icpsr’s simple download currently provides no means to meet those goals, but the standard secure download does. however, no one involved felt the disclosure of the data was sufficiently risky to warrant additional barriers to researchers for using the data, including irb sign-off, a signed agreement from the researcher’s institution, and strict technology prerequisites. data collected from the earliest faces cohorts predated icpsr’s online restricted data application system, and data access was administered via a paper application form. the advent of icpsr’s online restricted data access system brought some challenges for icpsr to consider for the dissemination of faces. because simple download did not satisfy the dissemination needs of the study, and the new on-line application system placed more barriers than were necessary to responsibly disseminate the data, we developed a hybrid approach. for faces, we continue to gather the required researcher information through the original paper form that researchers download from the study’s homepage, and then—once we approve the researcher’s access—we deliver those data via a download link emailed to the user. once the researcher gets the data, the data are subject to the same rules as public-use data: the researcher is not to redistribute the data or attempt to identify individuals or organizations and must properly cite the data in any publications they make from using the data, but they are not required to destroy the data upon completion of their work with them. the head start impact study the second example demonstrating the need for flexible technical approaches comes from the head start impact study7 (hsis), which is a nationally representative sample of head start programs and over 4,500 children. as with faces, the disclosure of centers presents moderate risk of harm to the head start centers included in the study. while all of the study’s direct identifiers have been removed, some indirect identifiers are available to enhance the analytic utility of the data. the concerns with disclosure were, therefore, greater than they were for faces, and the data producers and sponsor agreed that the restrictions associated with the secure download option were most appropriate. however, they were highly concerned with making the most sensitive of the study’s files available through secure download. the center analysis file contains a compilation of publically available population and household characteristics data for each center’s local community at the county and https://doi.org/10.29173/iq941 https://www.researchconnections.org/childcare/studies/36968 https://www.researchconnections.org/childcare/studies/36968 5/7 akmon, dharma and jekielek, susan (2019) restricting data’s use: a spectrum of concerns in need of flexible approaches, iassist quarterly 43(3), pp. 1-7. doi: https://doi.org/10.29173/iq941 census tract level. even though centers are not directly disclosed in the file, because the information in the file is unique for communities, the centers could conceivably be identified through triangulation with public data sources. as a result, the data producers and study sponsor had serious concerns with allowing researchers to download the file, even with the strict technology requirements of secure download. in fact, the data producers were only willing to make this file available through the vde, where researchers must confine their work with and analysis of the data. only vetted, final output can be removed from the vde, further reducing the disclosure risk for this file. in some ways, the needs of hsis fit very well with icpsr’s options for restricting access. however, the inclusion of a single file requiring vde dissemination within a study where every other file could be offered via secure download somewhat challenged technical systems that do not easily allow us to specify a different means of dissemination for a single file within a restricted use study. therefore, we had to create what essentially became two different studies at icpsr: 1) the head start impact study8, where files could be applied for and accessed via secure download and 2) the head start impact study with center analysis file9, where researchers could apply and access all of the studies files within a vde. this means that whether or not a researcher needs to work with the center analysis file dictates which version of the study she should apply to access. american indian/alaska native head start family and child experiences survey the last example we discuss to better understand restricted data access needs is the american indian/alaska native head start family and child experiences survey10 (ai/an faces). ai/an faces stands as the first ever national study on ai/an head start children and families. as with the faces study discussed earlier, there is a moderate risk of harm to centers if disclosed. the risk of disclosure is somewhat higher due to narrow focus of the study on a particular population, however direct identifiers and other indirect identifiers were removed. what makes this study unique and adds nuance to how we should think about restricting access is the study’s focus on a specific community and the associated agreements that the sponsor and data producer have with the tribal communities that participated. specifically, researchers’ permissions to use of the data depends on their commitment to use the data with consideration for the tribal centers in which the data were gathered. this includes both training and a researcher review process that is far outside the scope of icpsr’s internal restricted use application review and requires reviewers with expertise in working with tribal communities. the risk of disclosure and harm associated with disseminating ai/an faces best fit with the requirements and application process for icpsr’s secure download dissemination option. in addition, for the aforementioned reasons, we had to meet the study’s need for third-party review of applications, which none of icpsr’s restricted use options were designed to do. we created a workflow that leveraged both the study’s homepage and secure download application system, resulting in what we might think of as “secure download plus.” researchers must follow two separate application processes that serve different functions and come together in the final submission package to icpsr. on the ai/an faces homepage, interested researchers access the ai/an faces application guide11, which includes instructions and guidelines for submitting an application package to the thirdparty panel (the ai/an faces data committee) as well as links to required trainings to use the data. https://doi.org/10.29173/iq941 https://www.researchconnections.org/childcare/studies/29462 https://www.researchconnections.org/childcare/studies/36968 https://www.researchconnections.org/childcare/studies/36968 https://www.researchconnections.org/childcare/studies/36804 https://www.researchconnections.org/childcare/studies/36804 https://www.researchconnections.org/files/childcare/pdf/aian_icpsr_guide.pdf https://www.researchconnections.org/files/childcare/pdf/aian_icpsr_guide.pdf 6/7 akmon, dharma and jekielek, susan (2019) restricting data’s use: a spectrum of concerns in need of flexible approaches, iassist quarterly 43(3), pp. 1-7. doi: https://doi.org/10.29173/iq941 all research plans are reviewed by the external ai/an faces data committee, which is comprised of individuals with expertise in conducting research with tribal communities as well as representatives from native american head start programs or tribal community representatives. the committee reviews each research plan to evaluate whether the research plan demonstrates a sound understanding of the study design and sample and the proposed research questions are able to be answered by the data; assess the plan for disseminating findings; and evaluate the expertise and experience of the research team in working with tribal communities. at the same time, researchers follow the icpsr restricted use application process. in the end, the researcher submits an application to icpsr that includes an irb approval/exemption notification letter, the signed ai/an faces data committee12 notification letter, a signed acknowledgement that the researcher has read best practices for working with ai/an faces data13, and a signed data use agreement. after everything has been approved, researchers receive the encrypted files via email, as they would with icpsr’s standard secure download process. conclusions the three studies we described reveal nuance in the dissemination needs of sensitive data. specifically, these examples show that flexibility is needed around the following three key areas: the information collected from would-be users as a prerequisite to data delivery; the differing degrees of sensitivity of individual files within the same study; and who to include in the review and approval of applicants to use the restricted data. like faces, some studies represent enough disclosure risk to warrant collecting additional information about applicants and screening them prior to disseminating the data. as our workaround shows, building in the capability to specify for particular studies which information about a researcher must be collected (e.g. full name, affiliation, research questions, contact information) without the more stringent requirements of irb sign-off, institutional legal agreements, and strict technology setups would help appropriately protect study participants without placing unnecessary obstacles to accessing and using the data. the faces study requires an icpsr staff member to review and approve applicants, but we can also imagine scenarios where it is sufficient to collect additional information about secondary users (i.e. not null) without requiring staff review and approve the researcher’s access to the data, suggesting another area where flexibility is needed. as the head start impact study demonstrates, individual datasets within a study can represent different levels of disclosure risk, and so our systems also need to allow us to specify dissemination restrictions at the file level. while icpsr’s current systems allow for a mix of restricted-use and publicuse files within the same study, they do not easily allow for the restricted files within the same study to have varied dissemination methods. customizable settings that allow staff to specify different access requirements for different restricted-use files promise to improve the user experience for researchers who currently must determine which version of the study they need and then are directed to completely different application systems (one for secure download; another for virtual data enclave) depending on the files of interest. https://doi.org/10.29173/iq941 https://www.researchconnections.org/files/childcare/pdf/aian_faces_best_practices_for_researchers.pdf 7/7 akmon, dharma and jekielek, susan (2019) restricting data’s use: a spectrum of concerns in need of flexible approaches, iassist quarterly 43(3), pp. 1-7. doi: https://doi.org/10.29173/iq941 our third example demonstrates that some studies require more specialized vetting of users than can be provided by repository staff on their own, particularly when consent agreements define a particular review process as a prerequisite to data access. studies such as ai/an faces demand not only irb sign-off, institutional agreements, and strict technology configurations, but also a commitment that all interested researchers complete specialized training and have their research plans approved by a panel of individuals with expertise specific to the study. icpsr’s current system met this need with an existing application portal that allows for the inclusion of additional documents in the submitted application package. interested researchers follow two application processes—icpsr and the ai/an faces panel—that come together in the end with a single application package to icpsr. finally, all three examples demonstrate the importance of building in documentation and review of individual case studies to support continuous system improvement. while icpsr’s systems are flexible enough to provide solutions for each of the cases described, the details of each case can be used to inform future system development that facilitates researcher access to and analysis of secondary data while maximizing protection of data. in this way, data repositories more effectively balance the need to protect study participant confidentiality with facilitating the broadest data reuse possible and bolster their credibility as responsible stewards of research data. 1 this paper was originally presented at iassist 2018, montreal, canada. 2 dharma akmon is assistant research scientist and director of project management and user support at the inter-university consortium for political and social research icpsr and can be reached by email: dharmrae@umich.edu 3 susan jekielek is assistant research scientist and director, education archives at the interuniversity consortium for political and social research and can be reached by email: jekielek@umich.edu 4 https://www.icpsr.umich.edu/icpsrweb/ 5 https://www.researchconnections.org/childcare/welcome 6 https://www.researchconnections.org/childcare/series/236 7 https://www.researchconnections.org/childcare/studies/36968 8 https://www.researchconnections.org/childcare/studies/29462 9 https://www.researchconnections.org/childcare/studies/36968 10 https://www.researchconnections.org/childcare/studies/36804 11 https://www.researchconnections.org/files/childcare/pdf/aian_icpsr_guide.pdf 12 the ai/an faces 2015 data committee is comprised of individuals with expertise in conducting research with tribal communities and representatives from region xi ai/an head start programs. 13https://www.researchconnections.org/files/childcare/pdf/aian_faces_best_practices_for_resear chers.pdf https://doi.org/10.29173/iq941 mailto:dharmrae@umich.edu mailto:jekielek@umich.edu https://www.icpsr.umich.edu/icpsrweb/ https://www.researchconnections.org/childcare/welcome https://www.researchconnections.org/childcare/series/236 https://www.researchconnections.org/childcare/studies/36968 https://www.researchconnections.org/childcare/studies/29462 https://www.researchconnections.org/childcare/studies/36968 https://www.researchconnections.org/childcare/studies/36804 https://www.researchconnections.org/files/childcare/pdf/aian_icpsr_guide.pdf https://www.researchconnections.org/files/childcare/pdf/aian_faces_best_practices_for_researchers.pdf https://www.researchconnections.org/files/childcare/pdf/aian_faces_best_practices_for_researchers.pdf report on possibilities for a photograph database by adam engst' consultant 901 dryden road #88 ithaca. ny 14850 introduction as gould colman explained it to me, the transfer of the bibliographic information on the archival photograph collection to a computer database is currently open to a large number of possibilities. while there are advantages and disadvantages to programs on a number of machines, there are no external forces currently requiring a certain solution. i was retained to research the possibilities and present my findings, either recommending hardware and software combinations to test or recommending that the archives wait several years before repeating this process. 1 have gone through several steps to come up with this report first, i tried to determine precisely the needs and desires of the archives. second, i researched the software possibilities on three different hardware platforms which are available to the archives, the macintosh series, the ibm pc line, and a general category of mainframe. third and finally, i weighed the advantages and disadvantages of various combinations, adding in my knowledge about the cornell community, the state of the software industry, and the history of some companies in particular. most of the information below comes from my notes on telephone conversations held with representatives of the various companies. questions since the archives is not locked into using any specific program or computer, i was left to figure out what sort of a system would best fit the needs of the department as i see it, there are some requirements placed on any system by the size of the data, the nature of the data, the use to which the data is put, and the cost of the hardware, software, and programming time. size and speed gould told me that the archives currently has between 50,000 and 100,000 photographs. assuming one record in the database for each photograph, the system must be able to handle 100,000 records with decent searching speed. this was the first question i asked of the various database companies. their replies must be taken at face value though, since the only way to really test the each system is to put 100,000 representative records into each and do some searches. keywords since a small number of photographs are the end result of any search through the database, the system must be able to handle a relatively large number of keywords, or else researchers will have difficulty narrowing down their searches. a number of databases require programming convolutions to be able to deal with a field containing an unknown, but potentially large, number of keywords. such a hmitation does not rule out a database, it simply downgrades it in terms of ease of setup and programming. in addition, selecting the keywords is an extremely important task which must be thought out carefully. graphical information bibliographic information is useful for providing a brief description of each photograph and locating it within a collection, but for a researcher who is trying to find a certain photograph, possibly out of hundreds of similar ones, bibliographic information will not allow that researcher to select a certain photograph with surety. a graphical method of describing the photographs would decrease the amount of time it would take a researcher to find the right photograph for the use he or she has in mind. the main possibility is displaying images on a videodisc. few of the database systems can directly control a videodisc player. there are serious cost drawbacks to displaying visual information though, so inability to control a videodisc does not disqualify a database. initially, scanning the images and storing them on a cd-rom would seem to be feasible because each cdrom can hold 600 megabytes of information. unfortunately, cd-rom is not feasible because a scanned photograph has an average file size of 300k, which, when multiplied by 100,000 photographs, would force you to use close to 50 cd-roms holding 600 megabytes each. in comparison, a single videodisc can hold 108,000 images. costs the cost of the software are minimal in comparison to the costs of transferring the images to videodisc, although the purchase of expensive mastering equipment can reduce the overall costs. another cost which cannot 46 lassist quarleriy be ignored is the cost of programming and setup with whatever software is decided on. as a result, ease of programming does play a financial role in the final decision as well. so the questions that i asked each database company were as follows: •can your program handle 100,000 records with a fast search speed for a single record, say under 10 seconds as a worst case scenario? •can your program handle unlimited length text fields in an index (to retain searching speed) or is there a simple way around the program's inability to do so? •is there any way for your program to access images stored on a standard videodisc player? •how hard would it be to set up your program with a simple interface for researchers who may be inexperienced with computers? i also tried to get a feel for each company—how easy they would be to work with, how much help they would be if we needed any technical support, and whether or not they would still be in business in several years. these are intangibles, but potentially useful pieces of information. a note before i get into the details. i've tried to write this so no technical knowledge is required to understand it i'm sure that in some places i have failed because there is simply no other way to talk about certain features and actions of computers. in those places, i've included a footnote or tried to explain the term i use within the text. if at any time, you are confused reading this, please call me, and i will attempt to clear up the source of the confusion. hardware the companies with whom i spoke have database programs that run on the macintosh series of microcomputers, the ibm pc line of microcomputers, and (in the cases of oracle and notis) almost all minicomputers and mainframes. the archives currently has several ibm pcs and clones and will be getdng several macintosh se/30s shortly. in addition, i gather that the department has access to the mainframe resources of the library and the university. so existing hardware does not bias the decision. with a few exceptions, all of the software packages i researched can handle the large size of the database without a loss in searching speed. obviously, the minimum (and preferred) hardware configuration in each case does vary slightly, although some packages run fine on less powerful machines, which is a bonus since it will reduce the costs. in general, and like all generalizations this one is not to be trusted completely, the macintosh will be the easiest to set up and for both researchers and staff members to use. ibm pc clones have the advantage of being in the majority, although powerful systems are not really much cheaper than macintosh systems. microcomputers have the advantage (and disadvantage) of local control—if something goes wrong with the computer you can have it fixed quickly if necessary, whereas you must wait for another department to respond to your problem with a mainframe. on the other hand, if something with a microcomputer fails, you must deal with it, unlike with a mainframe, which will have a staff to deal with problems. mainframes often suffer from poor interfaces as well, although there are ways of avoiding the poor interfaces. using a mac and hypercard with a videodisc is simple, while using a pc with a videodisc requires a special device driver, which is a small program which allows the computer to control the videodisc. such programs are available, often from the videodisc maker, and there are also programmers who could write a custom device driver if necessary. general hardware conclusions based on my experiences with the various types of computers and my knowledge of the cornell user community, i recommend using a macintosh. macs are predominantly easier to work with in the setup phase, and they are far easier for inexperienced users to work with. since the entire point of this project is to provide easy access to information, i think that the interface is one of the most important parts of the system, and better interfaces can be created on the macintosh. in any case, my research covers all three platforms, and i hope that my recommendation of a hardware platform is bom out by the software possibilities on the macintosh. in addition, the macintosh database companies were far more knowledgeable about controlling videodiscs, which is why i often have more information on the macintosh databases. macintosh software company: istdesk systems program: istteam hardware: mac price: $795 their powerful relational database, called istteam, is compatible with hypercard and can store up to 255 characters in each field. offhand, 255 characters doesn't sound like it would necessarily be enough for our keywords, but spring 1991 47 perhaps their literature will shed more light on the subject. otherwise, istteam is certainly a possibility because it should be fast enough and can control a videodisc through hypercard, although it is much less well-known than either 4th dimension or omnis 5. company: acius program: 4th dimension hardware: mac price: $695 4d can handle 100,000 records with no problems, but it would have trouble with indexing. without indexing, the search speed slows tremendously, but 4d cannot index its unlimited length text fields, which we would use for holding keywords. so setting up the keywords would be a litue tricky in 4d. there are a number of different ways around for this problem, but they would require a bit more work programming. in the first french version, there was some kind of external command which could control a videodisc, although they may not still exist in the current version. there is a demo database, called minifans, which we could look at if 4d turned out to be a likely candidate. against 4d, i've heard that it is one of the slower databases for the mac, which is a problem for this project a test of its speed would definitely be needed before i could recommend it any farther. acius is one of the major database companies for the mac, but i have been unable to get through to them at all, which may indicate mediocre customer support while this is not a complete argument against using 4d, it doesn't bode well for future support needs. i can't really recommend them unless i can get through to talk to them. company: biyth software program: omnis 5 hardware: mac and pc price: $695 omnis 5 from blyth software certainly has the power to deal with 100,000 records, and it can be extended to do even more such as control a videodisc, although the representative didn't think such an external command had been written so far. alternately, either blyth could do it for us fw free if it was small and fairly easy or an independent programmer might be willing to write such a thing for a fee. omnis can index variable length text fields, so it would have no problem with a field containing a variable number of keywords. if the software to control a videodisc was difficult to write or acquire in other ways for omnis, it can work with hypercard so that hypercard uses the omnis database while acting as a front-end. however, omnis can also create simple interfaces easily, so it should not be necessary to link the two together on that account. omnis's language is supposedly english-like and easier than most database programming languages. an advantage of omnis over any of the hypercard extensions is that omnis is a full-fiedged database, and as such, can generate reports and display multiple windows, which would be good for displaying a number of records which met the search criteria. omnis needs a minimum of 1 megabyte of memory and is happier with a fast machine and more memory. overall, i was quite impressed with the possibilities of using omnis 5, since it seems to meet all the requirements and be fairly easy to work with in addition. the representatives have been extremely knowledgeable and responsive, unlike some of the other companies, such as acius. company: fox software program: foxbase plus and foxbase/mac hardware: pc/mac price: $395/$495 foxbase can have up to 254 characters in text fields, and can search on unhmited length text fields, but they aren't indexed which slows the search. however, fox claims that foxbase can handle up to 1 billion records, and that it is the fastest of all the mac databases by a great deal (some 30 times faster than 4d), although some of the pc databases come close in speed. reportedly, a new version of foxbase/mac can use hypercard's external commands (such as the ones to control a videodisc player) directly, which would be a major point in its favor. foxbase is generally accepted to be better than dbase ill-ton the pc and the mac version is cwrespondingly good, if not better since the mac version can handle unlimited length text fields. foxbase cannot control a videodisc, although it could work with a cd-rom. despite foxbase's speed and file compatibility between machines, i think it is somewhat too umited in this situation because of its inabiuty to index unlimited length text fields and its inabihty to control a videodisc. in addition, if it crashes for any reason, it will often corrupt the entire database rather than just losing the last record entered. this is a serious problem because you can never predict crashes. company: odesta corp. program: double helix 11 hardware: mac assist quarterly price: $395 double helix cannot link to hypercard and cannot control a videodisc, although there is a version that runs on vax mainframes (for about $5000) which would help the speed and storage problems. despite the fact that double helix can handle 100,000 records and has a simple method of programming, i doubt that this program is a real answer. double helix simply doesn't have enough to recommend it over any of the other major databases except its idiosyncratic programming environment, which may be easier than most hypercard extensions all of the following jjroducts require hypercard, or one of two hypercard clones, supercard or plus, which provide the same basic features as hypercard but with significant extensions. if a hypercard system is decided on, it would be well worth the time to investigate creating the database in supercard or plus rather than in hypercard itself. the various extensions hsted below may or may not work with supercard or plus, although there is a good chance that they will. the areas in which supercard and plus go beyond the capabilities of hypercard include reporting, graphics, multiple windows, and color. whether or not these features are worth moving away from hypercard is another question entirely, and one that need only be asked if the archives decides to go with hypercard rather than one of the fullfledged databases. company: answer software program: hybase (under hypercard) hardware: mac price: $150 size and speed are not problems, since hybase can handle up to 2 billion records and can usually find a single one in about 5 seconds. all fields are unlimited in size, or at least very large. the company claimed that it is not difficult but that some programming experience is helpful. answer software could set up the database for us if necessary. however, gregory crane at harvard said that he used hybase on project perseus for a while and found it very difficult to work with. dealing with the people at answer software was rather difficult and based on gregory crane's advice, i don't think that hybase is a good possibility. it suffers from difficult set up, which is unnecessary for this project company: discovery systems program: hypersearch (under hypercard) hardware: mac price: $99 i have not yet received any information from discovery systems regarding their hypersearch package, so i cannot make any specific statements for or against it. however, library of congress is using it in their american memwy project, and they seemed pleased with its speed and ease of use. company: knowledgeset corp. program: hyperkrs (under hypercard) hardware: mac price: $195 hyperkrs works completely within hypercard so it would be simple to design the database. nothing else need be done in terms of setup except for generating the index, which is fairly slow, but only needs to be done once. a mac plus is all that is required for searching. hyperkrs was designed for cd-rom, which accounts for its speed. i tested the demo software they sent me and i wasn't remarkably impressed. i had trouble finding anything, mostly because i was unfamiliar with the information for which i was searching. the speed was good but not great, but my macintosh is not that fast, which is certainly an issue with this program. on the whole, hyperkrs sounds like it may be the simplest of all the hypercard extensions to set up initially. after that, i have no real numbers to compare its speed with hyperhit or xearch. company: novasoft engineering group program: gridfile (under hypercard) hardware: mac price: $195 novasoft has a sample application called clipfile for gridfile which is being used right now to access pictures on clip art cd-roms. we would have to do the indexing and setup ourselves, which would be difficult without the aid of a relational database expert. gridfile is extremely fast, though, and is able to search any database for a unique record in 3 disk reads (certainly under 1 second). the cons of gridfile include the fact that it requires 2 megabytes of memory and a fast hard disk; it does not provide as good data packing as some other databases, which makes the file larger, it slows down on smaller databases in comparison to the others; it doesn't support split files over two or more hard disks; and it would be hard to set up. the pros of gridfile are that it is blindingly fast (faster even than some mainframe databases) and that it uses hypercard as a front-end, which can then control a videodisc. spring 1991 on the whole, i think gridfile is very powerful, but possibly too difficult to work with. there are other programs which provide similar speeds, but are easier to work with and require less hardware. company: sortstream international program: hyperhit (under hypercard) hardware: mac price: $195 steve hannaford, the technical support representative for hyperhit said that hyperhit has extremely fast searching speed (~1 sec) and can handle unlimited length fields with no problems. the information does not need to be textual—it could be pictures or sounds. the setup is not trivial but not that hard, and the hyperhit system is entirely contained in external commands that work within hypercard. steve didn't think it would be hard for someone without database training to use. the data file is external to the hypercard stack which would control the videodisc and thus requires only a small amount of space. another advantage to the external data file is that there could be a number of different interfaces since the hypercard stack does not have the data embedded in it. search time is usually under 1 second with a mac plus, and it would drop with a faster machine and hard disk. there is httle speed degradation when the file size increases (3 hundredths of a second when going from 1000 records to 10,000 records). in fact, the videodisc might be the bottleneck, depending on how fast it can find each frame. steve hannaford was very helpful and said that he wasn't getting many calls as the technical support person for hyperhit, which could mean that people aren't having any problems worth calling about. the advantages of hyperhit are that it is extremely fast despite what sort of machine it runs on, it supposedly isn't difficult to set up (although gregory crane will be testing it for project perseus soon and will have an opinion on its ease of use), and it will allow simple interfaces and videodisc access through hypercard. overall, hyperhit sounds like a good possibiuty. company: the voyager company program: videostacks hardware: mac price: $99.95 videostacks is a set of external commands to control a videodisc for a number of different videodisc players. there is a possibility that some of the external commands would be available from apple free of charge or they might be distributed with the videodisc itself. some sort of videodisc drivers will be necessary. company: xiphias program: xearch hardware: mac price: $??? xearch is an external command for searching in hypercard which xiphias uses in their cd-rombased product. time line of history. it sounds like it would be fast enough and they do have a licensing agreement, although i don't yet have the details. initially xearch sounds like it could be quite useful, although 1 don't have a sense of how easy or fast it is in comparison to hyperhit or hyperkrs. general hypercard software conclusions i think that all of the various packages mentioned above will probably provide hypercard with the searching speed necessary to use the database. the main distinction then, lies in the ease with which each is set up. hyperkrs and xearch are probably the easiest, with hyperhit, hybase, and gridfile lining up in increasing order of difficulty. more specific research and testing would need to be done to determine speed and ease of use in order to choose between the various extensions. hyperhit may be the best compromise between speed and difficulty. general macintosh software conclusions i am of two minds in this category. i think that hypercard is a wonderful program (not to mention the fact that it is free with all macs), and it will become integrated into the macintosh hardware and system software in the next few years, making it even stronger and faster. on the other hand, it really is not a database and does not provide the features that a full-fledged database provides, such as reporting and fast searching. it may be necessary to use hypercard in some fashion to facilitate access to a videodisc, which lends strength the cases of those database products that can link to hypercard, such as 4d, omnis 5, and istteam. on the other hand, the structure of the proposed database is very simple and does not really require the full power of a relational database. in the final consideration, i think i would currently recommend omnis 5 because of its power, flexibility, and ability to link to hypercard, not to mention the quality of the customer support, with which i was very pleased. pc software company: ashton-tate program: dbase iv/dbase mac lassist quarteriy hardware: pc/mac price: $795/$49s neither dbase iv nor dbase mac have any internal way of controlling a videodisc, and the representative didn't know of any external ways either, although he thought one might be possible. in addition, neither can index on variable length text fields, which would slow them down a great deal for this purpose. add these problems to the fact that ashton-tate is undergoing major problems as a company and has publicly announced that they will not be upgrading dbase mac at all, and you get a company to stay away from. company: borland international program: paradox/reflex plus hardware: ibm pc/mac price: $725/$279 the borland representative didn't think that either paradox, the more powerful pc program, or reflex plus, a decent macintosh database, could control a videodisc. it took several phone calls and some time on hold to get that much information, so i didn't pursue it farther. however, tim at turquoise filma'ideo productions (one of the mastering services) said that he was thinking about re-writing his custom database in paradox because it was fast and fairly easy to work with. he also said that paradox runs on a number of machines and is probably file compatible with reflex. as a result, paradox sounds like the best of the pc databases that i've looked into. using paradox would require some additional device driver to control the videodisc, but such a program might be available from a number of sources, including turquoise productions. paradox won most of the speed tests i saw in the course of my research, so i would recommend it over the other pc databases based on what i currently know. company: dataease international program: dataease hardware: ibm pc price: $700 dataease cannot control a videodisc, although it supposedly can interface with three scanners for using pictures. however the storage of those scanned images would be ridiculous and dealing with graphics on the pc is more difficult than on the mac. dataease does have long text fields which can be searched, although the representative wasn't sure about whether or not they were indexed, which is a major concern. i have a demo disk from them which may answer the indexing question, although i see nothing special about dataease otherwise. company: image concepts program: c-quest hardware: pc or unix mainframe price: $6000 or $25000 c-quest is a proprietary system for storing photographic information and controlling a videodisc. it has been around for several years, but doesn't seem to have a devoted following. clif nickerson of image concepts was somewhat helpful, although his system is designed more for a stock photograph collection than a histotical research collection. the main evidence of this is the way it uses synonyms of keywords, a method which allows the user to search on "stream" and get "brook" and "river" and "run" and "creek". unfortunately this is not nearly as useful with proper names of people and places, since they tend to be specific. the only use i can think of it is use modifiers, so you could have frank rhodes walking, talking, shaking hands, or making a speech, and search on the action involved. i don't know if that is too much trouble to set up and key in or not c-quest runs under unix mainframes as well as pc clones. under the unix system, c-quest can display 18 pictures at once; on the pc it can only display one at a time. its speed is dependent on the number of subjects used in the search, but clif said something about speeds of under 1 second, which he said was faster than the mainframe database ingres (and he thought than oracle). the c-quest interface is menu-driven and not particularly good. it does not have a simple interface for researchers to use, although one is being proposed. image concepts will change, add, or remove fields from the menus for a nominal fee, which is not as good as setting it up oneself. c-quest is not cheap, by any means, at $6000 for the pc version of the software and $250(x) for the unix version. clif said that the easiest way of getting images on disc were to buy a video camera and a writable videodisc, at which point you could do it all inhouse. i suspect that quahty wouldn't be as good, although there is no way to know without trying. he recommended making a 35mm film image in case the resolution of the monitors increased enough to make it worthwhile to re-master a videodisc. c-quest has an impressive list of clients, although 1 suspect that is from being the only game in town for 4 years, since no one else does this on the pc at all. evidently, the library of congress system is slower than c-quest, although that doesn't really spring 1991 mean much without more details. c-quest can control videodiscs from a number of companies, such as sony, pioneer, philips, and panasonic. overall, i find their system to be somewhat clumsy, expensive, and not really suited to the needs of the archives. the archives photographs have specific subjects without synonyms and need a very simple interface for researchers. while $6000 is not truly expensive in relation to the cost of mastering the videodisc, i think it is quite a bit more than you would pay for any other system. it might require less setup initially, although it is still a generic program that would require some customization. i cannot recommend it, especially since i heard from another consultant that the version of c-quest at the united nations was actually quite slow. company: microrim program: rbase for dos hardware: ibm pc price: $725 rbase will handle an unlimited number of records but the representative didn't know if it could handle an unlimited length text field. i didn't want to hold any longer to find out if it might be able to control a videodisc since the representative didn't think so. i see no reason to specifically recommend rbase. company: symantec program: q&a hardware: ibm pc price: $??? q&a does not have variable length text fields, but it can have large ones which are indexed so that the search speed doesn't suffer. the speed isn't good, though, at 15-20 seconds average, partly because q&a is not a high-powered relational database. there is no videodisc access, although the representative thought that an external program might work. considering the.speed problem, i can't recommend looking any further at q&a. general pc software conclusions i think most of the major pc database programs will handle the textual part of the database without trouble. however, it seems as though it will be more difficult to link the textual information in a pc database to the frame numbers of a videodisc. these device drivers do not seem to be readily available or supported by the database companies. for instance, the representative of ashton-tate knew nothing about linking to a videodisc, yet supposedly the library of medicine is using dbase ill-t-. however, a device driver might come with the videodisc player. i also feel that it will be more difficult within these pc programs to create a foolproof interface for researchers who are inexperienced with computers. i missed at least two major databases for the pc, revelation and nutshell, because i was unable to find phone numbers for them. however, i am not particularly wcxried that they are the perfect database because none of the other pc database companies had much of an idea what a videodisc even was, much less if their program could control it. the pc database companies were also much harder to reach on the telephone and much less willing to talk. if someone else turns out to be using revelation or nutshell, it would be worth checking them out. otherwise, i think paradox will be the best on the pc side. mainframe software company: oracle program: oracle (runs under a hypercard front end on the mac) hardware: mac/pc/minicomputers/mainframes price: variable depending on version—from $299 to $1299 oracle for the macintosh is a port of the most popular database program in the world. it retains complete compatibility with all other oracle databases on all other machines, which is a plus if this data will be shared with other people. in addition, a macintosh running the hypercard front-end to oracle can use any oracle database on any machine. because hypercard is the front end to the actual database, oracle can control a videodisc through hypercard. should the archives wish, they could probably find a mainframe or minicomputer on which they could use oracle. oracle's advantages are speed, portability, and ease of use with hypercard, although it might be a bit more expensive than the archives would want initially. there is a developer's version for the mac for $299, which would allow us to test its capabihties (with the only limitation being that this version cannot link to other oracle databases on other machines). the main disadvantages to oracle are that it is potentially more expensive (although the macintosh version is quite cheap) than other databases, and that it may simply be too complicated for the relatively simple database information we have. i gather that setting up a database in oracle is not all that easy. in addition, i've heard that oracle for the macintosh is not that fast and occasionally does strange things to data files. oracle as a company is excellent, with toll free support and guaranteed stability. they are the 52 assist quarterly largest database company in the world and the third largest software company in the world. company: northwestern university program: notis hardware: ibm mainframe price: free notts has a number of advantages, although it also spotts major several disadvantages. notts is currently installed and running in the library, so there are no added software or hardware costs to the system other than a macintosh and videodisc from which people can search in the archives. tt is relatively fast and can certainly handle another 100,000 records in its database. notts has the advantage of being accessible from anywhere on campus, but researchers may not use it unless they can also see the videodisc images because it is difficult to search for photographs based solely on bibliographic information. it has been in use at cornell for some time now, so many people are familiar with its interface, although its interface is also one of its main disadvantages. searching and moving between the various results of a search in notts is difficult and completely not intuitive. its other main disadvantages include the fact that it would be very difficult to link it to a videodisc, if it is possible at all, and the problem of portability of data since notis does not have the abiuty to export its information to another program, something which all of the microcomputer databases can do and which is very important for future expansions or modifications. it might be possible to sidestep notts's poot interface with a hypercard interface currently being worked on at mann library. in addition, there is a commercial product that will be available soon from texas a&m and apple, called macnotis, which also provides a better interface to notis. because both of these fm-oducts use hypercard, it is theoretically possible to have the hypercard interface control a videodisc while using the information from the notts database. howard curtis of mann library thought that this was possible, although extremely clumsy and prone to break whenever either notts or hypercard changed much. other people thought that it would be an unworkable situation even if it was theoretically possible. howard also said that it might be possible, though difficult, to program hypercard to download records from notts to the mac, which would allow the records to be used by microcomputer databases. notts is an easy solution because it requires no new hardware or software, but putting the records into notts removes them from a certain level of accessibility. tt would be difficult and clumsy to attach a videodisc to a macintosh running one of the hypercard interfaces, if it is indeed possible at all. more research would need to be done to determine tjie reality of such a setup. even worse, it would be hard to transfer those files to any microcomputer system. however, there are some ways of moving from microcomputer databases to a format which notts can read, which points towards putting the records into a microcomputer database first, and then, if there is interest, transferring a copy to notts. as much as notts seems like the simplest solution, i don't feel comfalable recommending it given the possibility for videodisc access and the inaccessibility of the data once it is in notts. i reauze that notts data can be shared by other mainframe cataloguing databases, but they don't (on the whole) provide the kind of features that microcomputer databases do. being able to move data between systems is important, and customized mainframe databases are a blockade to such a move. videodisc mastering services company: image premastering services program: videodisc services hardware: na price: variable image i'remastering services claims they are known for having the highest image quality for still frame transfers. some of their main clients have been the united nations, the library of congress, the mayo cunic, and the american college of radiology. they can handle absolutely any original—for the library of congress they laid down 30,000 glass plate negatives without cracking any. of course, slides are the cheapest method, and run anywhere from 550 to $1.35 per slide. other original media are correspondingly more expensive, although presumably the cost goes down with quantity. in addition, there are some basic initial costs which cannot be avoided. these costs total $3 1 50, although that is minor compared to the cost of transferring 100,000 photos to the disc. $2000 for the master disc $500 fw a check disc $150 fw the videotape $150 fw a duplicate/backup, kept at their site $350 as a basic setup fee they claim that stokes is mainly a slide copy service and makes a 35 mm film negative, which is a second-generation picture of the original. if the spring 1991 image is from a print, then the videodisc image is third-generation picture and suffers correspondingly in quality. stokes does colorcorrection, so the colors may be bright, but they are likely to be inaccurate. stokes is also generally cheaper because everything is automated in their process. image, on the other hand, is specifically dedicated to mastering videodiscs and they have patented technology for the process. they use a 2 foot lens over a 12 foot optical bench, which gives them two advantages. 1) the light comes in at a perfect 90° angle, which gives much better edge definition to the image. 2) they use an aerial image transfer, which somehow projects the image so that there is no film grain in the resulting videodisc image. it also allows them to easily perform custom sizing. the image representative recommended the mac and said that a military project used hypercard with oracle. he had heard something about 4d, but didn't know of anyone who was using it. he didn't recommend using the pc at all because the hardware is more expensive and is harder to set up the software to interface easily with the videodisc. company: stokes mastering services program: videodisc services hardware: na price: variable i spoke with john stokes and jim couch of stokes mastering service. in regard to costs, stokes estimated that a basic image transfer of positive images would be somewhere between $2 and $3 per image, although his estimate for complete costs (ie. in-house handling and database work) was closer to $4.50 per image. the work is done onsite and includes a person to come and do it with stokes's somewhat specialized equipment there is not much difference between stokes's doing the work and it being done in-house except for the fact that he claimed they had higher quality control, which is fairly likely. i suspect this is somewhat cheaper than image premastering's prices, although not by as much as i had originally thought they have a number of projects going on, the most notable of which is for the library of medicine, whose cost was about $2.40 per image. that price is slightly inaccurate because it was a test run in some ways and the library got two sets of negatives and slides for each of 70,000 images. the perschi to talk to at the library of medicine is lucy kiester, phone number 301-496-5962. the library of medicine is using a pc with dbase iii+ for their database. stokes claimed that videodisc access was incredibly simple with any database and that you didn't need a custom driver, but he did admit that you had to write some software. the library of congress is also working with stokes, which is curious since the representative at image premastering said that the library of congress was working with them. perhaps there are two different departments? in any case, stokes claimed that the library of congress is using some in-house computer system rather than an off-the-shelf software package. stokes said something about how that was their policy. this is not necessarily true since the american memory project is using a macintosh and hypercard. one advantage of stokes's method is that you can get negatives of each image as well, which allows you to reduce handling of the original images by making additional negatives. the library of medicine uses these negatives for public access to avoid giving out their originals. interestingly enough, stokes said that more people are doing the imaging first, then the database work, partly because stokes can uansfer the images to videodisc faster than the database can be set up. i was unsure about the real reasons for this, but they could be determine by talking to some of the people stokes referred me to. as far as hardware goes, stokes sounded like he doesn't really know very much about the mac. he claimed that the mac had no advantage over the pc in ease of use if the software was designed well, although i disagree with that rather strongly. based on a number of years of working with novices on both systems, the pc is less intuitive and clumsier than the mac when it comes to user interfaces. in any case, stokes has developed a database package under informix (which is oracle compatible, or so he said) which runs on a number of different machines. if we decide to use oracle, we could buy the oracle package and then stokes would provide us with his custom database for only the cost of support. this is curious because he could very easily create a stand-alone oracle database and then just sell that (or give it away if he wanted) without the customer having to buy their own copy. stokes wouldn't really comment on image premastering except to give me the name and 54 assist quarterly number of bill perry (202-857-7537) at national geographic, which did independent tests of both stokes and image premastering (and one other, actually whose name steves did not mention). in addition, stokes claimed that their quality has improved since then. evidently, the american college of radiology had 1 l"xl4" x-rays which stokes claimed were optimized for the image premastering system and those came out better. as far as quality goes, any transfer from a positive image will lose quality in the transfer process, much as copying a tape or videotape loses quality. the contrast of a videodisc is usually around 20 to 25 with a maximum of 45, whereas a transparency is about 1(xx) and a slide about 250. thus a great deal of contrast is lost when going to videodisc in any case. the transfer to a negative reduces this contrast lost by spreading out the contrast rather than clipping it, aluiough it also puts it through several generations of imaging. stokes is aiming at a contrast factor of two to three times better than high definition video, which is as good as a monitor will get in the near future. he said that it is very difficult to match the original exacuy, and that matching the original better is their main task right now. company: turquoise film/video productions program: videodisc services hardware: na price: variable turquoise said that they can provide anything up to a turn-key system. their background is in motion picture processing, and they moved from that to providing software and hardware as well. their database is an in-house one currently but it can import and export to a number of other formats. they are thinking about re-writing in paradox, which can also run on a number of different machines. they charge an average of $2 per image, including hardware, and they can shoot either in st. louis or on-site. they use a special motion picture film to get better quality images, but i have no sense how the quality of their images compares to the quality of either stokes's or image premastering's images. it doesn't seem that there is anything remarkable about turquoise in relation to the other two mastering services, but should the archives decide to have a videodisc mastered externally to cornell, it would be a good idea to talk more specifically to all three companies. other cornell projects i spoke with anne camell about the project in the university photography department, and she said that they are having someone in publications develop an inhouse program. this p-ogram will catalog and store all of the information on their photographs, but it will also provide billing , usage, and reporting capabilities. they didn't think one of the commercial programs could provide all that, something which i doubt, given the power of some of these databases. they are looking at videodisc in the near future, for much the same reason as the archives, perhaps in the next year or so. she didn't seem to have a wonderful grasp on what is entailed with the entire technology, since she didn't know about the problem with file sizes for scanned images, and she didn't know how many images could be stored on a videodisc. i also spoke with dave watkins, who is the head of media services and in doing so found my way back to the original stokes project mentioned in the memos to and from chris pelkie. they are ciurenuy selecting slides, negatives, and prints for a free sample videodisc to be supplied by stokes. partly to test the quality of stokes's service, they are trying to assemble a number of different types of images for inclusion. if the archives wishes participate in this project to see how it works out, they should contact dave watkins. he is looking for up to three hundred of the most difficult type of images. the deadline for submission to this project is january 1st, 1990, which is fast approaching. other departments that may be interested in this project include the department of entomology, which has 30,000 slides that they wish to use for diagnostic work, obviating the need to go to the slide collection itself. these slides must be of the highest quality because of their use and the fact that the disc will be sold to other universities. plant pathology and veterinary medicine may wish to do similar things. evidenuy the hotel school has a videodisc of wine labels and the vet school and the law school are still looking into the possibilities of some sort of image database on a videodisc. if the archives wishes to look more closely at the specifics of a videodisc system in future, i strongly recommend that the department provide a number of photographs to dave watkins for this test project the project will provide a videodisc which can be shown to potential donors and with which we can test the pros and cons of various software packages. such an opportunity should not be passed up lightly! i met with margaret webster, who runs the architecture school's slide library. she has for some time been planning a videodisc project to keep track of the 350,000 slides in the library. tliey hired an outside consultant, nancy humphries of etech, to research the possibilities and provide a system. margaret has been putting records into the database and will be setting up the pilot project after the slide library moves in january spring 1991 55 of 1990. they will be using a pc database called btricve/xtrieve (with which i'm not familiar because it was never compared in the literature with the more wellknown databases) along with a specialized graphics board in the pc that will allow them to manipulate the images electronically. manipulation of images is something which i did not explore particularly because of the expense involved, but is certainly a possibility fw the archives. the question that must be answered to justify the cost of such a board is the use to which these photographs are being put. if the photographs are ending up in publications designed and executed on a personal computer, or the images must be frequently manipulated, then a graphics board makes sense. however, if the publications in which these photographs appear use traditional methods of production, then additional 35mm negatives would be mwe useful. margaret webster knows more than most people on campus about videodisc systems because their system has been in the research phase for close to two years now. the final person from cornell with whom i spoke was mike oltz from the interactive multimedia group. he gave me some bits of information that may be useful. he thought that the architecture school and the history of art department were looking into something similar and were probably working along different lines. as it turns out. history of art has dropped their project entirely, whereas the architecture project is the closest to reality of any of the ones i've heard of. the exception to this is the medical school, which has a system for pathology training using a macintosh pseudo-database called guide. they started out with no funding at all and now have close to five million dollars of computer equipment, which does lend hope to the archives getting funding for a videodisc. the interactive multimedia group has a sample videodisc from image premastering services, which is currendy lost, but mike will try to find it and let me know via email. he mentioned that revlon, the makeup people, are also doing something like this. if we wanted to learn more we should talk to the advertising photography department there is no reason to assume that any of the other systems are better than a system the archives could come up with, although the ability to perform information transfer might be useful in the future to avoid re-keying records. similarly, standard information in each record would help in translating the records from another system when other departments wished to archive various photographs. database information can be shared between pcs and macs without too much trouble, so the specific machines used by different departments should not really matter, although the ability to transfer between the various databases matters. cornell coordination one problem that has come up time and time again in my research is that there are a number of cornell departments working completely independently on similar videodisc projects. while such a lack of communication is not unusual at cornell, it is regrettable, particularly in a field such as this where the information really is fairly finite. it would be extremely useful if there could be a single person who would, if nothing else, have copies of all the various pieces of information collected by the different departments. that way, whenever anyone was thinking about starting such a project, the information would be more or less at hand and would include names of the people at cornell who are good resources. this person would merely disseminate information and would refrain from making any recommendations as far as hardware or software go in order to avoid the politics. the interactive multimedia group would seem to be a logical group to coordinate or various videodisc information, but because they exist completely on soft money for specific projects, they are not set up to handle any sort of coordination. geri gay said that media services was one place coordination could come from, and some part of cit services would be another. people to talk to in cit include larry fresinski, donna tatro, and if all else fails, stuart lynn. another area in which the various departments could pool resources would be in setting up facilities at cornell for transferring images to videodisc. i gather that there are close to a million images at cornell that could be put on videodisc if the process was cheaper and easier. dave watkins in media services is the person to talk to about such a project margaret webster in the architecture slide library would also be very interested. other non-cornell projects i've found names of people at other institutions who have done something along these lines or are thinking about it talking to them might help the final decision because you can get an opinion from someone in a similar position. i did not get more detailed information from these people since it is often easier in this situation to use academic channels for sharing information, and the archives ah-eady has contacts in some of these institutions, whereas i would be going in cold. so it doesn't make sense for me to talk to everyone immediately unless it seems that they have something important to offer to the decision-making process right now. that step can come if and when the archives decides on a specific system or type of system. if i know of a way of contacting the jjeople below, i've mentioned it most of this information comes via electronic mail, so i can ask for additional contact information if desired. my apologies for the lack of organization, but no method proved 56 assist quarleriy itself better than formatting for maximum readability. elizabeth wood mentions joint project between the emergency medicine and radiology l5epartment of los angeles county and the university of southern california medical center that had their mastering done by image premastering. elizabeth h. wood computer services librarian norris medical library university of southern california ewood%phad.hsc.usc.edu@ usc.edu the av department of hombake library at the university of maryland is developing a videodisc in conjunction with the national agricultural library. 1 wonder if this is the fwestry service collection mentioned above. david austin mentions two other projects. first, andrew eskind at the eastman house is working on something to do with a videodisc. second, jim sheldon at the mit media lab is working on a videodisc of edweard (sic) muybridge motion pictures in conjunction with the addison gallery of american art, phillips academy, andover, ma, 01810. finally, david says "also, make sure you check the sn/g: report on data processing projects in art (1988). it is not yet on-line but available in hard copy, maybe even at cornell. it is a list of projects registered with the scuola normale superiore, pisa, italy and the getty art history information program, los angeles." david austin u29716@uicvm jim sheldon jls@ media-lab.media.mit.edu the aviador (avery videodisc index of architectural drawings on rlin) project at the avery architectural and fine arts library at columbia sounds very similar lo what the archives might want to do. in addition, rlg is working on a way of hnking a videodisc to an rlin terminal, which would be very interesting. janet parks sent me a copy of their literature on the videodisc system. janet parks curator of drawings avery architectural and fine arts library columbia university new york, ny 10027 212-854-6738 jane kleiner mentions several videodisc projects. one of which is the emperor i collection done by ching chi chen at simmons. it is quite sophisticated and includes sound as well. she thinks mit has an architectural collection on videodisc and adds that the national agricultural library has a collection of historical photographs from the forestry service on videodisc. jane kleiner notjpk@lsuvm lennie stovel mentions that there would be more information in the library of congress* literature on their prints and photographs division's videodiscs, although he does not give a specific contact. lennie stovel library systems analyst research libraries group bl.mds<a)rlg.bitnet david finkelstein at stanford university academic information resources says that they are currently digitizing a large slide collection, which will eventually reside on videodisc. they are currently using hypercard because of its ease of use but are looking into more powerful database programs such as ingress, which is in use but is not necessarily well-liked at mit's athena project. david finkelstein academic information resources stanford university davef@jessica.stanford.edu steve cisler from apple computer has available a technical report on basic videodisc production. it is called "multimedia production: a set of three reports", and includes: "casual multimedia production", "videodisc basics", and "videodisc production of the visual almanac". it was done by apple's multi-media lab for the production of the forthcoming visual almanac. it is written for the non-technical person and includes mastering costs, sources of replicators, techniques. steve mentioned that some people at the visual resources association fell that the image quality was not high enough for scholars. he also said that there is a new method of distributing videodisc images with a certain type of network called broadtalk. steve can be contacted for more information on broadtalk, and he will send a cqjy of the videodisc report to anyone who sends him a request on university letterhead along with a selfaddressed mailing label. i have the report and recommend it highly for anyone who is actually starting on the specifics of producing videodisc. steve cisler apple library spring 1991 57 10381 bandley drive ms: 8c cupertino, ca 95104 sac@apple.com in response to another question, steve cisler gave the address of several replicators for videodiscs which i have yet to contact. these are as follows. crawford communications 506 plasters ave atlanta, ga 30324 404-876-8722 pioneer communications 1058 e 230th st. carson, ca 90745 3m optical recording 223-5s 3m center st. paul. mn 55144 612-733-2142 cynthia read-miller at the ford museum is using a videodisc and microcomputer catalog set up by a company called argus. bernard littau at uc davis is putting together a radiology learning system for the veterinary school. he warns about several problems with videodisc production. first, it is extremely expensive to master the first disk from videotape. it requires a great deal of staff and equipment time just to make the videotape, and then the frame numbers on the resulting videodisc must be matched with the photograph database. second, he feels that videodiscs are more suited to video sequences since that was what they were originally designed for, and, videodiscs have limited resolution for displaying still images in comparison to a digital image stored on cdrom and displayed on a computer monitor. unfortunately, as subsequent conversations with bernard proved, 100,000 photographs is simply too many to put in cdrom format because it would require 30 or more cdrom discs. third, he said that when they did the videodisc, they were forced to use three frames for each image by the mechanics of the process of transferring the images to videotape. as such, they ended up using three frames of the videodisc for each image, reducing the storage capacity by three. using three frames per image had the advantage of safety if one or two of the images were bad for some reason or other. i wonder if this problem appears if a commercial mastering service does the work since bernard's project was done in-house, i believe. bernard littau vm radiological sciences school of veterinary medicine, university of california davis, ca 95616 916-752-0184 internet: vmrad@ucdavis.edu bitnet: vmrad@ucdavis there is an integrated image database package with its own programmers application language developed by pcm, inc. the package is called pc album and runs on an ibm-pc. pcm, inc. 8330 boone blvd. suite 430 vienna, va 22180 703-356-1600 or 800-654-5845 ernst robl recommends an expensive system called inmagic. it is sold by a company of the same name in cambridge, ma. the los angeles pubuc library uses it to catalog their extensive photograph collection and speaks very highly of it the version which runs on ibm-pc type microcomputers is $1000, and there is a version which runs on vax mainframes as well. ernst says that inmagic allows a considerable amount of individual configurations and handles variable length data well. it can accept data from other sources, which is good for compatibility reasons. in fact, the la public library has staff members do the cataloguing on laptop computers in the stacks rather than bring the collection out to a terminal. ernst has served a couple of terms as chair of the picture division of the special libraries association and has authored an introductory book on picture librarianship, organizing your photographs [amphoto,1986]. in connection with the above, he has visited a large variety of institutional and commercial picture collections. (the picture division no longer exists as an individual entity with sla, but its interests have been taken over by several other divisions.) his book points out some general issues to consider in the cataloging of photos, although the sections on computers are fairly basic because of its audience. ernest h. robl systems specialist (tandem system manager), library systems 027 perkins library, duke university durham, nc 27706 (919) 684-6269 w; (919) 286-3845 ehr@ecsvax russell grau mentions a project he worked on with a company called laser recording systems. the project consisted of taking images, scanning them onto a worm drive, and then accessing the images via bibliographic information stored in a database. the whole 58 lassist quarterly thing ran on ibm-pc type microcomputers. the person russell worked with was named tom cwsten, but he may not be there any more. laser recording systems, inc. 270 sparta ave. sparta, new jersey 07871 201-729-3055 russell grau 916-920-9092 gordon fair mentions that oracle for macintosh can work with supercard as well as hypercard. unfortunately, supercard is much slower than hypercard and a project like this does not require supercard's color and animation abilities. gordon fair gf07-i-(2)andrew.cmu.edu there is a package called videodisc showmaker that will allow you create a database of entries with keywords and then search over the fields in the database. it is intended for a substitution for a slide projector in classes that require many images (it was originally designed for graphic arts education). it is a collection of stacks for the novice hypercard/macintosh user; it interacts with the videodisc players using videodisc drivers from apple; and it can handle large databases of images (around 2(xx)). the current version of showmaker uses the hypercard find function and is not very fast, but it does what it is supposed to. it is going to be released in december through a company called ztek, which supplies interactive-video software. if interested, get in touch with them or with the professor in charge, mark sanders. my only problem with videodisc showmaker is that it is definitely not fast enough for 100,000 images. some sort of hypercard extension software would be required to increase the search speed. mark sanders msanders@vtvml.cc.vledu msanders@vtvml (703) 231-6480 bob samson at the university of texas at arlington might be setting up a system using the series 2000 laser-optic filing system from tab products. i don't know much about this jwoject, but i gather that the system is a digital system, so i don't know how they are getting enough storage space for 350,(xx) photographs even though it comes with either a 5.25 or 12 inch optical disk. the system also includes a computer (no indication of what kind), a scanner, a high-resolution monitor for viewing the images, and possibly a laser printer for creating hard copy. he would be using this to store 350,000 images from the photographic archives of a local newspaper. since many of the photographs are quite old, he wishes to avoid physical contact when possible. bob samson university of texas at arlington b366rcs@utarlvm1 817-273-3000 lucy kiester (phone: 301-496-5%2) at the national library of medicine is finishing up a videodisc project and used stokes mastering service to transfer her images to disc. bill perry (phone: 202-857-7537) at the national geographic society has also done some wwk with stokes mastering service. mike segel recommends using the informix database (which runs on many different microcomputer and mainframe systems) because it allows you to store blobs (binary large objects) in the data base. i don't know if storing images as blobs would take up less space, but if it didn't the space requirements would be prohibitive. mike segel segel(a)quanta.eng.ohio-state.edu ed heath is an intern at the library of congress and is working on the american memory project. he sent me quite a bit of information on american memory. the project uses a pioneer laservision player and the machine is controlled by a macintosh iix using hypercard and discovery systems' hypersearch. the photographs were mastered onto the videodisc before the project started for another reason so ed didn't know too much about the specifics. american memory deals with keywords by using free text searches with a "visual materials" thesaurus developed by the library of congress. otherwise it is the american memory setup also includes a cd-rom. ed heath special projects university computing george mason university fairfax, va 22030 (703)323-2941 • eheath@gmuvax lloyd davidson tells of an article in byte magazine (january 1988—("a better way to compress images", byte 13/1, 215-218, 220-223)) in which a method using fractal geometry achieves graphic compression at ratios of over 10,(xx) to 1. such compression ratios would easily allow a cd-rom to store a great many images and would make them far more feasible for spring 1991 extremely large image collections. he also mentions a second article about the same researchers in the november 4, 1989 issue of the new scientist. p.40. lloyd a. davidson seeley g. mudd library for science and engineering northwestern university evanston, il 60208 l_davidson@nuacc.acns.nwu.edu overall conclusions and comments taking everything i currently know into consideration, i would recommend using a macintosh se/30 with 2 megabytes ofram and at least an 80 megabyte hard disk. that will satisfy any of the programs and leave plenty of room for expansion. as far as the programs go, i currently recommend omnis 5 with hypercard to provide videodisc access if no external routine for this are easily available. of course, testing would be necessary before a final decision. no matter what software is used, there will be a fair amount of programming time necessary to set it up and get it running. in addition, the time it will take to enter 100,000 records into the database will be considerable. i cannot make any recommendations as to the mastering services because i do not have enough hard evidence to work with. ideally, cornell would set up its own facility for transferring images to videodisc. since transferring images to a videodisc is so expensive, i can only recommend that the archives search for donors. in the meantime, deciding on a database and starting to enter the data would be useful whether or not a videodisc is ever monetarily feasible. once the data is entered into a microcomputer database, it would not be too difficult to move it to another system, should a standard appear or merely a better method of working with a videodisc. i see no reason to wait on starting the database for this reason, and the cost of transferring images to videodisc will not drop much in the future, if at all. if cornell set up a facility to transfer images to videodisc, it might be cheaper, although one never knows. an important procedure for the moment is to think about the format of the database. the fields of bibliographic information are set, but some thought must be given to the keywords. the problem is with keywords because there are simply too many different possibilities, since everyone thinks different keywords are important. the architecture school has a thesaurus, which helps, but they will still need to add some keywords and ignore others. three to four levels of hierarchical keywords ( ie. post 1905 people professors professor kaplan) are probably as detailed as you want to go at first, since it is too easy to come up with keywords which only make sense to the cataloguer after four levels in the hierarchy. a good way to figure out a system is to find a picture and then work backwards so you see what steps you went through to find it. hopefully this will be solved easily by simply using the categories of information that are already set up for the current system. looking at the architecture school's system might be helpful in determining the number and type of keywords. some more questions to keep in mind while designing such a system include who will be using the system, how will they be using it, and what is the end result going to be? answering these questions before the database is designed will help in the design stage to make sure that the database is really set up correctly for the purposes at hand. as a caveat, let me merely mention the problem of copyright. i assume that cornell owns the copyright on the photographs in the archives, but if not, it is technically a breach of copyright to transfer these images to a videodisc. further questions i have left a number of questions unasked in my inquiries because of the preliminary nature of the investigation. most of these deal with the specifics of the videodisc access, since that is the great unknown in the whole project. the questions into which i have yet to delve are as follows (along with my current opinions). •is the videodisc access feasible soon or farther in the future? my feeling is that the videodisc access will be very nice once it is set up, but it will be expensive to master the disc. assuming the prices quoted by the mastering services, a videodisc of 100,000 images could easily run over $300,000, if not more. unless a munificent source of funds appears, i suspect that the videodisc will simply be too expensive for the moment i don't think prices will drop much in the future, because the imaging and material handling work involved will remain more or less the same. in the event that several hundred thousand dollars should become available, the database should be able to support a videodisc. otherwise, the files would have to be exported and imported into another p-ogram, a procedure which can be difficult and time consuming. •what would be the best videodisc players for the archives's purposes? i didn't check into this at all, although i know there are a number of models that would work with either a macintosh or a pc. all the prices that i've seen are in the $2500 range. the writable videodisc which you can 60 assist quarteriy record to directly is quite a bit more expensive, at $13000 to$15000. •would the archives wish to use an outside mastering service or set up an in-house mastering service? the outside services are likely to be more expensive, although they would also probably give higher quality results. if there is enough interest. media services might set up a videodisc mastering system at some point, which would be ideal. their system might be somewhat cheaper and would certainly be closer. in any event, price and image quality seem to be directly related, so the nicer your pictures look, the more it will cost if in-house mastering is deemed unfeasible, then some testing between the three mastering services would be in wder. in addition to these specific questions, i'm afraid that this report brings forth more questions yet that can only be answered by careful thought on the part of the archives. i've made recommendations and hopefully discovered sources of information that will help answer these questions, but much more work will need to be done before a project like this becomes reality. for instance, it took margaret webster almost two years to start her pilot project if i can be of assistance at any later point in the project, please feel free to call me. appendix addresses and phone numbers istdesk systems 7 industrial park rd. medway, ma 02053 508-533-2203 800-522-2286 makers of istfile and related programs. acius 20300 stevens creek blvd cupertino, ca 95014 408-252^w44 makers of 4th dimension. they are very hard to reach. answer software 20045 stevens creek blvd. suite ie cupertino, ca 95014 408-253-7515 makers of hybase, a hypercard database extension. ashton-tate 20101 hamilton ave. terrance,ca 95052 213-329-9989 213-329-8000 makers of dbase iv and dbase mac blyth software 3655 campus dr. san mateo, ca 94403 415-571-0222 makers of omnis 5. i spoke with jennifer blome. borland international 1800 green hius road scotts valley, ca 95066-0001 408^38-8400 makers of paradox and reflex plus anne camell 1 159 comstock hall cornell university ithaca, ny 14853 255-7675 anne works in the university photography department. howard curtis mann lithary information technology section cornell university ithaca, ny 14853 255-9570 database international 7 cambridge ave. trumbull, ct 06611-9983 203-374-8000 makers of database discovery systems 7001 discovery blvd. dublin, oh 43017 614-76m197 makers of hypersearch, a hypercard database extension. ducsoft 238 columbus ave sandusky, oh 44870 419-626-6797 makers of applications and routines for 4th dimension, but nothing for videodiscs. fox software 27493 holiday lane perrysburg, oh 43551 419-874-0162 makers of foxbase-iand foxbase/mac spring 1991 61 image concepts p.o. box 211 west boylston, ma 01583 508^81-6882 clif nickerson, marketing manager. note that the phone number is different from the old literature. image concepts make c-quest, a videodisc program for the pc and unix boxes. image premastering services 1781 prior avenue north st. paul, mn 55113 612-644-7802 a videodisc mastering service. interactive media center geri gay or mike oltz cornell university ithaca, ny 14853 255-5530 knowledgeset corp. 888 viua st, suite 500 mountain view, ca 94041 415-968-9888 makers of hyperkrs and hyperlndexer, hypercard extensions. microrim 3925 159th ave ne redmond, wa 98073-9722 206-885-2000 makers of rbase. novasoft engineering group 2343 ridgewood ave. edgewater.fl 32032 904-423-5189 makers of gridfile, a hypercard database extension. odesta coip. 4084 commercial ave. northbrook, il 60062 312-498-8852 312-498-5615 makers of double helix n oracle corp. 20 davis drive belmont, ca 94002 800-345-3267 makers of oracle database software. spoke with a robert silverberg, exl 2019 softstream international 19 white chapel drive mount laurel, nj 08504 800-262-6610 609-866-1187 marketing company for hyperhit, a hypercard database extension softstream—steve hannaford 19 white chapel drive. mount laurel, nj 08054 215-543-5194 technical support representative for hyperhit. stokes mastering service austin, tx 512^58-2201 a videodisc mastering service. symantec corp. 10201 torre ave. cupertino, ca 95014 408-253-9600 makers of q&a turquoise filnvvideo productions st. louis, missouri 63088 314-843-1998 a videodisc mastering service. voyager company 239 manning ave. los angeles, ca 90025 800446-2001 makers of videostacks, a set of videodisc drivers and other software. dave watkins media services b-27mvr cornell university ithaca, ny 14853 255-5431 dave is the head of media services and is working with the videodisc project. margaret webster architecture slide librarian b-30 sibley dome cornell university ithaca, ny 14853 255-3300 xiphias 12464 washington blvd. marina del rey, ca 90292 213-841-2790 makers of xearch, a hypercard searching extension 62 assist quarterly a videodisc is an optical disc which is read by a laser. the images are analog, which means essentially that they are stored as a snapshot consisting of shades of gray or color, rather than being divided into individual dots which can be either on or off, which is how an image would be stored in digital format cd-rom stands for compact disk read only memory. it is a digital fwmat, which means that any pictures are made up of individual dots which can be either on or off, black or white. as a result, pictures take up a great deal of space on a cd-rom, so much space that a project this size would not be feasible. mainframe storage systems would be required to store so many photographs. there are two types of databases, relational databases and flat-file databases. flat-file databases work just like a file cabinet in that each record is stored separately. relational databases can share information between files, so you would not need to duplicate information if you had a database of addresses and a database of phone numbers because the two files could share the person's name. in addition, relational databases tend to be faster and more powerful. hypercard is a program described as a "software erector set" by its author. it is free with every macintosh and allows non-programmers to create sophisticated programs, called stacks. hypercard works on the metaphor of a stack of note cards, although it has a great deal of easily-accessed power which seems unrelated to a stack of cards. hypeicard is not a database, but it is an information manager and manipulator. an external command is a small program that can be inserted into another program to give the second program additional functionality. they are extremely common with hypercard and provide numerous ways of enhancing hypercard. a front-end is what you see and work with, whereas the back-end is the part which actually does the work. for instance, the front-end of a washing machine is the control panel where you set the type of wash and the amount of time. the back-end is the drum and vibration mechanism which actually washes the clothes. you have to be able to use the front-end, but you don't have to know how the back-end works to get your clothes clean. i don't quite understand their technology and am merely trying to repeat it verbatim in hope that someone more well versed in the photographic arts will understand. i don't know what the units in question are since stokes didn't mention them. a board is a piece of hardware which plugs into a slot in the computer and provides some sort of added functionality. add-on boards usually fulfill a specific need which most people do not care about, which is why such functionality is not built in to the computer itself. the ibm pc clones and the macintosh n hne can easily accept such boards. worm stands for write once, read many. essentially, a worm drive is just like a cd-rom drive except for the fact that the user can write to the drive once. it is very useful for archiving information because it cannot be erased afterwards. ' prepared for the department of manuscripts and university archives at cornell university olin library, ithaca, ny 14850. distributed with the permission of the department of manuscripts and university archives december i4th, 1989. paper presented at the lassist 90 conference held in poughkeepsie, n.y. may 30 june 2, 1990. spring 1991 63 1/1 rasmussen, karsten boye (2019) editor’s notes: the interest group on qualitative data sums up and continues, iassist quarterly 43 (2), pp. 1-1. doi https://doi.org/10.29173/iq961 editor's notes: the interest group on qualitative data sums up and continues welcome to the second issue of volume 43 of the iassist quarterly (iq 43:2, 2019). with joy and pride the many people behind each issue of the iq are here presenting a special issue. iassist has several interest groups of members committed to selected important areas under the umbrella of iassist. be aware that you could become a member of an interest group (see: https://iassistdata.org/about/committees.html#interest). if an interest area that you find important is not presently on this list, you are invited to start campaigning for the formation of a new interest group. the interest groups discuss and document their area and often arrange sessions at the iassist conferences. more formalization and continued documentation of the group’s work are presented in conference papers and papers published here in the iq. this issue of the iq is dedicated to papers on qualitative data presented by members of the group named ‘qualitative social science & humanities data interest group’ (qsshdig) and related practitioners. lynda kellam from the cornell institute for social & economic research and mandy swygart-hobaugh of george state university end their leadership of the group with this special issue. lynda kellam and celia emmelhainz (qualitative research librarian at the university of california berkeley) are guest editors of this issue and their introduction to the issue is following this page. i want to express my great thanks from the iq to lynda and celia for taking the job of compiling a special issue. support for qualitative data is important and a growing area. i trust you as readers will find valuable information and excellent advice in the papers of the many authors that are committed to improving the use and value of qualitative data. submissions of papers for the iassist quarterly are always very welcome. we welcome input from iassist conferences or other conferences and workshops, from local presentations or papers especially written for the iq. when you are preparing such a presentation, give a thought to turning your one-time presentation into a lasting contribution. doing that after the event also gives you the opportunity of improving your work after feedback. we encourage you to login or create an author login to https://www.iassistquarterly.com (our open journal system application). we permit authors 'deep links' into the iq as well as deposition of the paper in your local repository. chairing a conference session with the purpose of aggregating and integrating papers for a special issue iq is also much appreciated as the information reaches many more people than the limited number of session participants and will be readily available on the iassist quarterly website at https://www.iassistquarterly.com. authors are very welcome to take a look at the instructions and layout: https://www.iassistquarterly.com/index.php/iassist/about/submissions authors can also contact me directly via e-mail: kbr@sam.sdu.dk. should you be interested in compiling a special issue for the iq as guest editor(s) i will also be delighted to hear from you. karsten boye rasmussen june 2019 https://doi.org/10.29173/iq961 https://iassistdata.org/about/committees.html#interest https://www.iassistquarterly.com/index.php/iassist/about/submissions mailto:kbr@sam.sdu.dk 1/13 majawa, felix and hall, ralph, p. (2021), establishment of data centre at mzuzu university: a survey of anticipations and aspirations of key project stakeholders, iassist quarterly 45(3-4), pp. 1-13. doi: https://doi.org/10.29173/iq999 establishment of data centre at mzuzu university: a survey of anticipations and aspirations of key project stakeholders felix majawa1 and ralph p. hall2 abstract mzuzu university lost its library as a result of a fire that took place on december 18, 2015. in response, the university established two processes to ensure the library services were not interrupted. the first process was to restore information services within six months by creating an interim library. the second was to design a new library in collaboration with virginia tech’s school of architecture and design in the united states. a total of three conceptual designs were developed, from which mzuzu university selected a final design. one key aspect of each conceptual design was a dedicated space for a data centre. the initial concept was that the data centre would support research activities at the university, within malawi, and with international partners outside malawi, such as virginia tech. this paper captures the anticipations and aspirations of the key stakeholders involved with the library design project at mzuzu university in malawi and virginia tech in the usa. data were captured by a survey that was shared via email with 29 stakeholders. a total of 10 responded at mzuzu university, and 12 responded at virginia tech. a key finding from the survey was the need to create clear plans for each aspect of the project to ensure the effective implementation of the data centre. critical aspects to the project include staffing, equipment procurement, the management of the data centre, data literacy programming, and the long-term sustainability of the data centre. developing a policy/process to guide the operations of the data centre was also found to be critical. the library construction began in february 2021 and is expected to end in february 2023. having a clear plan for how the data centre could be operationalized will be essential to ensuring the centre is successful. the data centre will be a new facility for the university and this paper is a first step towards shaping the requirements of, and potential for, this new facility. keyword data centre, data literacy, data management, mzuzu university, malawi, virginia tech 1. introduction an increase in research by universities and other institutions associated with research has resulted in the production of large amounts of data (fox, 2013), advancing the need for mechanisms to generate, store, process, and use such data. one mechanism has been the development of data centres, which can be independent, providing data services to organisations on a commercial basis, or attached to specific institutions such as universities, providing data-related support functions to promote research endeavours, learning, etc. data has been defined as the representation of facts, figures, and concepts in a manner suitable for processing, interpretation or communication by human, or automated means while the term “data centre” has been defined “as a centralized repository, either physical or virtual, for the storage, management, and dissemination of data and information organized around a particular body of knowledge or pertaining to a particular business” (techtarget, 2012). however, telecity group (2011) defines “data centres as buildings that have within them electrical and mechanical infrastructure that creates an environment in which computing and telecommunications equipment can run without interruption.” the former definition emphasizes the use of a data centre while the later emphasizes the technical equipment that enables a data centre to perform its functions. the former definition is more suitable for the mzuzu university data centre as it focuses more on service provision and the need for data literacy. https://doi.org/10.29173/iq999 2/13 majawa, felix and hall, ralph, p. (2021), establishment of data centre at mzuzu university: a survey of anticipations and aspirations of key project stakeholders, iassist quarterly 45(3-4), pp. 1-13. doi: https://doi.org/10.29173/iq999 data literacy means “the ability to understand and use data effectively to inform decisions” (mandinach and gummer, 2013). further, calzada et al. (2013) state that data literacy enables individuals to access, interpret, critically assess, manage, handle, and ethically use data. also, carlson et al. (2011) and rin (2011) outline the necessary data literacy competencies as follows: discovery and acquisition of data; data management; data conversion and interoperability (dealing with the risks and potential loss or corruption of information caused by changing data formats); metadata; data curation and re-use; data preservation; data analysis; data visualization; and ethics, including the citation of data. mzuzu university (mzuni) was established by an act of parliament in 1997. it has the following six faculties: faculty of education; faculty of environmental sciences; faculty of health sciences; faculty of humanities and social sciences; and faculty of science, technology, and innovation. the university has about ten thousand students. the university library is mandated by the university act of 1997 as an integral part of the university’s teaching, learning, and research. the library was established with the mission “to provide up to date and relevant information resources; promote effective utilization of those resources; and facilitate rapid access to information held within and in remote places through convectional and electronic means.” progress in pursuing goals for achieving improved library services were hampered by the fire that destroyed the library on december 18, 2015, resulting in the loss of all its collections and ict equipment. specifically, 53,000 books, 68 desktop computers, 403 reading chairs, 62 reading tables, 111 shelves, three heavy duty photocopiers, eight printers, and other countless valuable items. the total value of items damaged was mk 5,891,214,532 (approximately $7,854,952) (chawinga and majawa, 2018). upon learning of the fire, virginia tech, a long-standing partner of mzuni, stepped forward to assist with the design of a new library. the school of architecture and design provided support by developing three conceptual designs, from which the executive leadership at mzuni selected one design for further development. the provision of a data centre is a critical component of the final design. 2. literature review in recent years, data centres have become an essential, if almost a surreptitious, element of business and social life throughout the modern world (jones et al., 2013). the critical functions associated with data management such as processing, storage, management, and exchange are performed in data centres, hence data centres have become the driving hub of the economy, and in some ways, of society (ibid.). for a data centre to operate, it needs specialized equipment. kumar (2020) outlines some of the equipment needed for a data centre to function, such as “rack server, san (storage area networking) & nas (network attached storage) devices, network, security, backup devices, additional services of equipment which help to support customers through noc (network operation center) & soc (security operation center) operations.” similarly, jones et al. (2013) identified rack cabinets, servers, routers, switchers, and air-conditioners as necessary equipment for a data centre to function effectively. data centres require people with specialised skills and knowledge to manage and operate them. snipes (2018) stated that “on a broader scale, the concepts of big data, data-driven decision-making and data literacy have become an important part of life of librarians.” similarly, cox and pinfield (2014) argue that academic library services are well positioned to play an important role in research data management (rdm). koltay and hungary (2015) reasoned that in the same way as libraries have traditionally facilitated access to documents, they now need to facilitate access to data. however, for this to happen, new expertise is needed from experts in ict, statistics/data analytics, data visualization, etc., expertise that is not normally available among library scientists. https://doi.org/10.29173/iq999 3/13 majawa, felix and hall, ralph, p. (2021), establishment of data centre at mzuzu university: a survey of anticipations and aspirations of key project stakeholders, iassist quarterly 45(3-4), pp. 1-13. doi: https://doi.org/10.29173/iq999 data literacy also plays a critical role in ensuring usability for a data centre. the concept of data literacy refers to “the ability to transform information into actionable instructional knowledge and practices by collecting, analyzing, and interpreting all types of data” (gummer and mandinach, 2015). this definition reveals that data literacy is central to the provision of data centre services. data literacy incorporates aspects of statistical literacy, assessment literacy, pedagogical knowledge, and data-driven decision making (henderson and corry, 2020). carlson et al. (2011) and rin (2011) outline the following range of data literacy competencies: discovery and acquisition of data; data management; data conversion and interoperability (dealing with the risks and potential loss or corruption of information caused by changing data formats); metadata; data curation and re-use; data preservation; data analysis; data visualization; and ethics, including the citation of data. these aspects are considered to be critical for the effective operation of a data centre. 3. problem statement the increasing use of it services has been accompanied by an increase in the need to process, manage, preserve, and use data in support of decision making. data centres came into being when a large number of companies required rapid and constant internet connectivity and dedicated data storage facilities (jones et al., 2013). the idea of establishing a data centre at mzuzu university developed through discussions between mzuzu university and virginia tech. the type of data centre being considered for mzuzu university is unique in malawi. given the potential and scope of a new data centre, it was important to identify the expectations of the key stakeholders at mzuzu university and virginia tech. the findings from this paper are intended to support the development of a proposal for a data centre that aligns with the ambitions of the stakeholders and provides guidance on the equipment, staffing, and manage structures needed for mzuzu university to run a state-of-the-art data centre. 4. objectives of the study the aim of this study was to investigate stakeholders’ expectations for the establishment of a data centre at mzuzu university. the specific objectives of the study were as follows: (a) to identify the staffing and equipment needed to run the data centre; (b) to establish the functions that the data centre would perform; (c) to determine the need for data literacy in carrying out the functions of the data centre; and (d) to identify the beneficiaries of the services provided by the data centre. 5. methodology data for this study were collected using a semi-structured questionnaire that was distributed through email to assistant librarians, senior assistant librarians, deans, faculty, and directors of research and studies at mzuzu university and virginia tech. in total, 15 stakeholders at mzuzu university and 14 at virginia tech were sent the survey. a total of 10 responded at mzuzu university and 12 responded at virginia tech, implying a response rate of 67% for mzuzu university and 86% for virginia tech, and an overall of response rate of 76%. data analysis was undertaken using spss. 6. findings and discussion 6.1 definition of data centre the first area of interest was to understand how respondents understood the concept of data centre. a definition by techtarget (2012) – who defined a data centre as “a centralized repository, either physical or virtual, for the storage, management, and dissemination of data and information organized around a particular body of knowledge or pertaining to a particular business” – was presented in the questionnaire and respondents were asked whether they agreed or not with this description. of the https://doi.org/10.29173/iq999 4/13 majawa, felix and hall, ralph, p. (2021), establishment of data centre at mzuzu university: a survey of anticipations and aspirations of key project stakeholders, iassist quarterly 45(3-4), pp. 1-13. doi: https://doi.org/10.29173/iq999 22 respondents, 15 agreed with this definition while 8 disagreed (figure 1). one respondent was noncommittal and neither agreed nor disagreed with the definition, stating “the promise of the data centre as a physical place should provide access to data and information relative to a broad base of knowledge as well as targeting the knowledge clusters of the partnering institutions. it should be widereaching and possess the capability to change, re-orient, and grow. the data centre’s mission should specify the desired extent of this.” another respondent, despite agreeing with this definition, said “i agree because it covers the different aspects of data (storage, management and dissemination). however, it does not cover data generation.” one of the respondents who disagreed with the definition stated: “the data center needs to be conceived in a much broader sense. first, the center should do much more than simply curate and share data. it needs to process, organize, and add value to the information. i also think the center should not limit itself to a specific body of knowledge. if this happens, it will greatly reduce the potential for future research/funding. i believe the center needs to be designed around a purpose/mission that is focused on unlocking the potential of data to advance positive social and environmental change.” similarly, another responded pointed out that “it would be best not to limit the data center to a single focus. though data on environmental science might be one area of concentration, the facility should accommodate new alternative research as it develops along with the university curriculum.” finally, one responded observed that the issue of information security should be part of the definition. the definition by techtarget (2012) proved to be a useful way for stakeholders (respondents) to reflect on the concept of a data centre. however, the definition was limited in the sense that it emphasised the collection and management of data in a particular body of knowledge. another limitation of the definition raised by a respondent is that it “does not capture the transformation/computation of data that is so important in the modern world. i would find it hard to classify google or microsoft’s data centers as data centers under the above definition.” this observation, combined with the multidisciplinary nature of mzuzu university and virginia tech, indicates that the data centre should be able to accommodate a broad array of research and adapt to new research opportnuities as they arise. another respondent remarked, “this definition is limited to a data-centric perspective, and a more forward thinking and sustainable definition would include the organizational and policy aspects of a data center. if these critical components of a data center are not included, then the perspective may become data-centric, [losing …] the human component, which is then at risk [of becoming …] an afterthought.” in this regard, one of the respondents suggests that the definition should include the following statement: the data centre should be “organized around one or more particular bodies of knowledge or pertaining to one or more businesses.’ fig 1: agreement with the provided definition of a data centre 15 8 agree disagree https://doi.org/10.29173/iq999 5/13 majawa, felix and hall, ralph, p. (2021), establishment of data centre at mzuzu university: a survey of anticipations and aspirations of key project stakeholders, iassist quarterly 45(3-4), pp. 1-13. doi: https://doi.org/10.29173/iq999 6.2 staffing the data centre the survey asked stakeholders about the type of staff needed to run/manage a data centre. as indicated in figure 2, most respondents agreed that librarians should be part of the team managing the data centre. one rationale for this is that librarians are specialised in acquisition, processing, and dissemination of information and these skills could certainly be applied to the management of data. this view is supported by snipes (2018) who states that “on a broader scale, the concepts of big data, data-driven decision-making and data literacy have become an important part of life of librarians. as information professionals, they need to understand data issues and how to integrate them into their work.” ict experts were also identified as critical for managing the technical aspects of the data centre. researchers are both generators and consumers of data and are therefore key staholders of the project. however, one question raised by a respondent is whether it will be neccesary to hire a researcher for the data centre. however, another respondent stated that researchers should not be employed as members of staff, but that the data centre staff should be able to work with researchers and support their data/analysis needs. statistical experts could play an important role in supporting the analysis of data. one respondent refered to statisticians as “experts in qualitative and quantitative data analysis.” another respondent highlighted the need to have a data center project manager, but did not specify the skills or qualifications needed for such a position. the need for “marketing/communications for effective dissemination” of research, was also highlighted by one respondent. this role could as well be performed by data literacy experts. a similar role could also be played by librarians by expanding their information literacy activities. another respondent pointed out that “the center should develop a system where advanced graduate students can work in the center to develop and use their skills and support services provided to other researchers/students.” while another respondent indicated that there was need to have computational scientists, who could also be system administrators and engineers. another skillset identified was the need for “data visualization experts, artists, designers […] to help get the ‘data’ out into usable formats.” domain/subject experts and transdisciplinary knowledge management experts were also considered as important. https://doi.org/10.29173/iq999 6/13 majawa, felix and hall, ralph, p. (2021), establishment of data centre at mzuzu university: a survey of anticipations and aspirations of key project stakeholders, iassist quarterly 45(3-4), pp. 1-13. doi: https://doi.org/10.29173/iq999 what this feedback reveals is the broad range of skills that are needed to support a successful data centre. fig 2: recommended staff for the data centre 6.3 equipment for the data centre the survey asked respondents about the types of equipment needed to operate the data centre. all respondents agreed that equipment such as public server (available in the public cloud), computers, routers, switches (1-gbps ethernet switch for local connectivity and connectivity to the campus network), wireless access points for local user access, rack cabinets, air conditions, storage systems, and water and smoke detectors were all considered to be important (figure 3). jones et al. (2013) highlight rack cabinets, servers, routers, switchers, and air-conditions as necessary equipment for a data centre to function. similarly, lam et al. (2020) indicated that the computing resources for data centres include servers, storage and access devices, communications equipment (routers and switches), databases, and software applications. however, some respondents added that there was need to have printers, genset for power back up, large smart screens/projectors, networking equipment, security cameras, a load balancer (to direct incoming traffic to various application nodes/servers), a telecommunication platform, a data center environmental monitoring system (e.g., sensaphone), a fire containment system that will not destroy the hardware during a fire event, and scanners for the conversion of nondigital data (pictures, notes, maps, etc.) to a digital form. one of the respondents observed that “storage, possibly multiple levels: some amount of high speed on-line primary storage, slower but much higher capacity near-line backup storage, off-line storage (generally tape) and off-site backup storage. a means for off-site backup of critical data and a plan for disaster recovery/continuity of operations is very important. if the data is important enough to collect, it is important enough to protect.” as kumar (2020) highlights the importance of it equipment power utilization: “it includes power load utilization of it equipment, i.e., rack server, san (storage area networking) & nas (network attached storage) devices, network, security, backup devices, additional services of equipment which help to support customers through noc (network operation center) & soc (security operation center) operations.” 19 20 14 17 11 ict experts librarians researchers statistical experts other https://doi.org/10.29173/iq999 7/13 majawa, felix and hall, ralph, p. (2021), establishment of data centre at mzuzu university: a survey of anticipations and aspirations of key project stakeholders, iassist quarterly 45(3-4), pp. 1-13. doi: https://doi.org/10.29173/iq999 finally, one respondent commented that “all of this equipment will be necessary. the degree of investment will be dependent on the mission. if the goal is to serve university students the investment will be modest. however, if the data center is to be a regional, national, or multi-country resource, outside support will be necessary.” this comment highlights the critical need to develop a mission, vision, and goals for the data centre that clearly outline the scope of work that the centre will support. fig 3: equipment for the data centre 6.4 functions of the data centre the development of a clear set functions (or capabilities) of the data centre will be essential to securing funding for research and building the reputation of the centre both nationally and internationally. as shown in figure 4, respondents agreed that the data centre should perform a wide range of functions associated with the acquisition, processing, and management of data. similarly, jones et al. (2013) observed that critical functions associated with data management such as processing, storage, management, and exchange are performed in data centres, hence they have become the driving hub of the economy, and in some ways, of society. similarly, laughton (2012) describes the open archive information system (oais) functionalities of a data centre such as ingest, archival storage, data management, preservation planning, and provision of access. another function mentioned by the respondents was the importance of data ‘synthesis,’ for new horizons that develop from the intersection of seemingly disparate or different pieces of information. however, some respondents added that functions such as data generation, ingestion, and retrieval should be shaped by the organizational structure and mission of the centre. 0 5 10 15 20 https://doi.org/10.29173/iq999 8/13 majawa, felix and hall, ralph, p. (2021), establishment of data centre at mzuzu university: a survey of anticipations and aspirations of key project stakeholders, iassist quarterly 45(3-4), pp. 1-13. doi: https://doi.org/10.29173/iq999 fig 4: functions of the data centre 6.5 data literacy topics necessary for the beneficiaries of the data centre all respondents agreed that the data literacy topics presented in figure 5 are important. in addition, one respondent indicated that data security is also an important topic for users of the data centre. jones et al. (2013) also argue that the constantly evolving concern of data security should be included in data literacy programmes offered by data centres. they further state that “there are many potential threats to data centres ranging from intruders physically trying to gain access to the facility and ultimately to the servers to cyberattacks and hackers trying to gain access to the network and to the data stored on the servers” (ibid.). also, creation of data and metadata, long term backup and storage, offsite mirror, intellectual property and licensing, and ethics, including the citation of data, is very important since data carries risk to be used in misinformation and unethical applications. in addition, mandinach and gummer (2013) enumerate data literacy skills that include knowing how to identify, collect, organize, analyze, summarize, and prioritize data. in this case mandinach and gummer are emphasizing what people should be able to do after going through a data literacy programme. 0 5 10 15 20 25 other data dissemination data virtualisation data acquisition data literacy data processing data storage https://doi.org/10.29173/iq999 9/13 majawa, felix and hall, ralph, p. (2021), establishment of data centre at mzuzu university: a survey of anticipations and aspirations of key project stakeholders, iassist quarterly 45(3-4), pp. 1-13. doi: https://doi.org/10.29173/iq999 fig 5: data literacy topics 6.6 beneficiaries of the data centre the survey also sought to identify the potential beneficiaries of the data centre. the most commonly cited beneficiaries were researchers, lecturers, students, and communities surrounding mzuzu university, and institutions in and outside malawi. in addition, the respondents also mentioned ngo’s, private companies, government institutions, and research and grant partners. one stakeholder remarked that “all of these could have access but with ‘tracking’ to make sure information is not misused.” similarly, henderson et al. (2014) stated that primarily, the beneficiaries of data services are the undergraduate and graduate students as well as researchers. in relation to providing services to communities surrounding mzuzu university, one respondent observed that such action would need to be based on “a policy decision. access to the [data centre …] services will give the community greater opportunities and increase the general well-being of the population, besides providing talent for working at mzuzu university.” this statement reveals that there is need to develop a data centre policy to guide the operations of the centre. another respondent alluded to the fact that the communities should be able to use the centre when they collaborate with staff/faculty at the university. one respondent indicated that “private companies represent a special case regarding revenue and expenses, though any group outside the university will need an agreed contract.” however, the respondent cautioned that there is need to make sure that all intellectual property created by the center is carefully managed. another respondent further observed that extending services to institutions in and outside malawi will be essential for generating financial resources for the center. the above feedback reveals the need to identify the costs of providing services and the need for a clear plan to manage the revenue generated by the data centre. while it is anticipated that the data centre will not be focused on profit making, it will need to generate sufficient funds to offset the costs of providing services. jones et al. (2013) highlight that “in the face of rising costs and seemingly ever more sophisticated technological innovation, many companies have not been in a position to develop and manage their own data centres. there is need therefore to find ways of sustaining the functionalities of the data centre.” 0 5 10 15 20 25 https://doi.org/10.29173/iq999 10/13 majawa, felix and hall, ralph, p. (2021), establishment of data centre at mzuzu university: a survey of anticipations and aspirations of key project stakeholders, iassist quarterly 45(3-4), pp. 1-13. doi: https://doi.org/10.29173/iq999 fig 6: beneficiaries of data centre services 6.7 institutional relationship the final section of the survey focused on understanding potential institutional arrangements between mzuzu university and the data centre. a majority of respondents (13 out of 22) indicated the data centre should be located under the leadership of the library. this position is supported by snipes (2018) who states that “on a broader scale, the concepts of big data, data-driven decision-making and data literacy have become an important part of life [for …] librarians.” similarly, cox and pinfield (2014) stated that several commentators have proposed that academic library services are well positioned to play an important role in research data management (rdm). however, survey respondents also mentioned that the data center should serve the university and be advised by a board or panel of key actors/stakeholders. such a board/panel could act as an important policy decision making body for the data centre, while allowing the data centre to be managed by library leadership. however, six of the respondents opted for the centre to be under the directorate of research, while three opted for the centre to be independent reporting to the vice chancellor. however, one respondent observed that “being independent directly to vc is risky unless monetization is [a] primary goal.” similarly, other respondents had different opinions regarding the position of the centre in the university. “why not a panel with representatives of all the entities? this will keep it from becoming political. rather than centralized authority, perhaps an ‘oversight committee’ or ‘work group’ comprised of important stakeholders, […] librarian, director of research, vc, members of key academic departments, technical professionals, etc.” 0 5 10 15 20 25 other (specify) institutions in and outside malawi people from the community surrounding mzuzu university researchers lecturers students https://doi.org/10.29173/iq999 11/13 majawa, felix and hall, ralph, p. (2021), establishment of data centre at mzuzu university: a survey of anticipations and aspirations of key project stakeholders, iassist quarterly 45(3-4), pp. 1-13. doi: https://doi.org/10.29173/iq999 fig 7: potential institutional relationships 7. conclusion and recommendations this exploratory study reveals a number of important issues that need further consideration with regards to establishing a data centre at mzuzu university. these issues range from defining what is meant by a data center, to considerations of the expertise needed to manage the centre, what the centre’s core functions should be and what this means for equipment, what data literacy programs might need to accompany the launch of the centre, how the centre will be positioned institutionally within the university, and who the target beneficiaries will be. the key observations from this study are as follows: • defining the data centre: while there are different perceptions of the concept of a data centre, adopting a broader view of the scope of the data centre would position it to engage in a wide range of research and educational opportunities. these opportunities would also enable the centre to diversify its potential revenue, creating a more robust financial model. • data centre expertise: the operation of the data centre would require people with different types of expertise, ranging from ict, statistics, data management, etc. library staff with skills relating to the curation of data and communication were also identified as important. in addition, the role of post graduate students in supporting the services offered by the data centre was highlighted as way to enable them to use and develop their skills. • equipment: a broad range of equipment needs were identified that included servers, computers, routers, switches, rack cabinets, air conditioners, and water and smoke detectors. however, equipment for off-site backup of critical data and a plan for disaster recovery were also deemed important tools for the centre. • data centre services/functions: several functions of the data centre were revealed through the study in relation to the generation, processing/analysis, visualization, storage, and dissemination of data. the role of a data literacy programme was viewed as essential for enabling data centre users to fully benefit from its services. one recommended subject to include in such a programme is data security. • data centre beneficiaries: the stakeholders revealed different groups of people who could benefit from the centre. however, there is need to first establish a policy that clearly outlines who can use the centre and what the related costs may be, so revenue generated can offset some of the costs of running the data centre. under the library leadership, 13 under the directorate of research, 6 independent reporting to vc, 3 other, 8 under the library leadership under the directorate of research independent reporting to vc https://doi.org/10.29173/iq999 12/13 majawa, felix and hall, ralph, p. (2021), establishment of data centre at mzuzu university: a survey of anticipations and aspirations of key project stakeholders, iassist quarterly 45(3-4), pp. 1-13. doi: https://doi.org/10.29173/iq999 • institutional location of the data centre: given the physical location of the data centre and the alignment between the services it could provide and those offered by mzuzu university’s library, it is recommended that the data centre be positioned under the library. the above points reveal a broad range of issues that need serious consideration if mzuzu university and its partners are to realise the full potential of the data centre. while construction begins on the new library, mzuzu university has a unique window of opportunity to develop a clear plan for creating a state-of-the-art data centre. it is recommended that a team of administrators, faculty, and staff at mzuzu university and virginia tech is convened with the charge of developing this plan. an important aspect of this effort would be to identify the initial areas of expertise at mzuzu university and virginia tech, where faculty could come together around joint research projects that leverage the data centre’s services/capabilities. the identification of these synergistic research opportunities should also help shape the type of services provided by the data centre. on the other hand, mzuzu university needs to consider collaborating with other universities in this project as the data centre could also support research data services for other institutions. one such institution is the kamuzu university of health sciences, which could benefit from sharing the services provided by the data centre. however, currently the project is being undertaken in close collaboration with virginia tech, thereby providing an opportunity to share data in the future when the project is realised. references carlson, j., fosmire, m., miller, c.c., and nelson, m.s. (2011), “determining data information literacy needs: a study of students and research faculty”, portal: libraries and the academy, vol. 11 no. 2, pp. 629-657 calzada prado, j. and marzal, m.á. (2013), “incorporating data literacy into information literacy programs: core competencies and contents”, libri, vol. 63 no. 2, pp. 123-134. chawinga, w. and majawa, f. (2018), “an assessment of mzuzu university library after a fire disaster”, african journal of library, archives and information science. vol. 28 no.2, pp. 183-194. cox, a.m. and pinfield, s. (2014), “research data management and libraries: current activities and future priorities”, journal of librarianship and information science vol. 46 no. 4, pp. 299–316. fox, r. (2013),"the art and science of data curation", oclc systems & services: international digital library perspectives, vol. 29 no. 4, pp. 195 – 199. gummer, e. and mandinach, e. (2015), “building a conceptual framework for data literacy”, teachers college record, vol. 117 no. 4, pp. 1-22. henderson, m., raboin, r., shorish, y., and van tuyl, s. (2014) “research data management on a shoestring budget”, bulletin of the association for information science and technology, vol. 40 no. 6. jones, p., hillier, d., and comfort, d. (2013), “data centres in the uk: property and planning issues”, property management, vol. 31 no. 2, pp. 103-114. koltay, t. and hungary, j. (2015), “data literacy: in search of a name and identity”, journal of documentation. vol. 71 (2) 401-415. https://doi.org/10.29173/iq999 13/13 majawa, felix and hall, ralph, p. (2021), establishment of data centre at mzuzu university: a survey of anticipations and aspirations of key project stakeholders, iassist quarterly 45(3-4), pp. 1-13. doi: https://doi.org/10.29173/iq999 kumar, r., khatri, s.k., and diván, m.j. (2020), “efficiency, measurement of data centers: an elucidative review”, journal of discrete mathematical sciences and cryptography, vol 23 no. 1, pp. 221-236. lam, p., lai, d., leung, c., and yang, w. (2020), “data centers as the backbone of smart cities: principal considerations for the study of facility costs and benefits”, facilities, vol. 39 no. 1/2, pp. 80-95. laughton, p. (2012), “oais functional model conformance test: a proposed measurement”, program: electronic library and information systems, vol. 46 no. 3, pp. 308-320. mandinach, e.b. and gummer, e.s. (2013), “a systemic view of implementing data literacy in educator preparation”, educational researcher, vol. 42 no 1, pp. 30-37. rin (2011), “the role of research supervisors in information literacy, research information network”, london, available at: www.rin.ac.uk/system/files/attachments/research_supervisors_report_for_screen. pdf snipes, g. (2018), “everyone’s a data librarian now”, journal of new librarianship, vol. 3 no. 1, pp. 28-31. techtarget (2012), “what is a data center”, available at: http://searchdatacenter.techtarget.com/definition/data-center. telecity group (2011), “annual report and accounts 2011”, available at: www.telecitygroup.com/annual-reports/telecitygroup_plc_annual_report_and _accounts_2011.pdf endnotes 1 felix majawa, librarian, mzuzu university, malawi. email address: majawa.f@mzuni.ac.mw 2 ralph p. hall, associate professor and associate director of the school of public and international affairs (spia) at virginia tech, usa. email address: rphall@vt.edu https://doi.org/10.29173/iq999 http://www.rin.ac.uk/system/files/attachments/research_supervisors_report_for_screen.pdf http://www.rin.ac.uk/system/files/attachments/research_supervisors_report_for_screen.pdf http://searchdatacenter.techtarget.com/definition/data-center http://www.telecitygroup.com/annual-reports/telecitygroup_plc_annual_report_and http://www.telecitygroup.com/annual-reports/telecitygroup_plc_annual_report_and mailto:majawa.f@mzuni.ac.mw mailto:rphall@vt.edu 1/13 thompson, kristi and sullivan, carolyn (2020) mathematics, risk, and messy survey data, iassist quarterly 44(4), pp. 1-13. doi: https://doi.org/10.29173/iq979 mathematics, risk, and messy survey data kristi thompson and carolyn sullivan1 abstract research funder mandates, such as those from the u.s. national science foundation (2011), the canadian tri-agency (social sciences and humanities research council, 2018), and the uk economic and social research council (2018) now often include requirements for data curation, including where possible data sharing in an approved archive. data curators need to be prepared for the potential that researchers who have not previously shared data will need assistance with cleaning and depositing datasets so that they can meet these requirements and maintain funding. data deidentification or anonymization is a major ethical concern in cases where survey data is to be shared, and one which data professionals may find themselves ill-equipped to deal with. this article is intended to provide an accessible and practical introduction to the theory and concepts behind data anonymization and risk assessment, will describe a couple of case studies that demonstrate how these methods were carried out on actual datasets requiring anonymization, and discuss some of the difficulties encountered. much of the literature dealing with statistical risk assessment of anonymized data is abstract and aimed at computer scientists and mathematicians, while material aimed at practitioners often does not consider more recent developments in the theory of data anonymization. we hope that this article will help bridge this gap. keywords anonymization, deidentification, sensitive data, survey data introduction as a result of an open government mandate in 2014 (government of canada, 2014), many datasets and surveys were released on canada’s government open data portal, giving researchers access to a trove of previously unavailable survey data. among these datasets were a set of surveys that had been conducted by firms on behalf of health canada. these surveys came to the attention of a group of canadian data librarians in part because the files were in several cases released in formats that were difficult to use and without adequate documentation. in other cases, documents were released that discussed surveys that had not been made available. the efforts of this group to track down documentation and datasets and obtain or build easier-to-use survey files have been described in a previous article (thompson, 2018). as a result of these efforts, one of the co-authors of this article was entrusted with a small collection of health canada datasets, which needed additional work beyond the group’s now standard practice of documenting and formatting. among the issues were that the datasets had not been fully assessed for anonymization, though obvious identifiers like name and address had been removed. the first coauthor, a data librarian with a computer science background, spent some time familiarizing herself with the theory behind data anonymization and statistical disclosure risk assessment in order to deal with this unexpected data issue. the second coauthor, a master of library science student who also has a computer science background, became involved later when work on the data became a part of her internship in data librarianship with the first co-author. this paper will provide a background to the field of data https://doi.org/10.29173/iq979 2/13 thompson, kristi and sullivan, carolyn (2020) mathematics, risk, and messy survey data, iassist quarterly 44(4), pp. 1-13. doi: https://doi.org/10.29173/iq979 anonymization and explain the work we did to ensure that two particularly difficult datasets had been anonymized. based on our experiences, we will describe a set of processes for reviewing messy survey data to make sure it has been properly anonymized. data anonymization background the first step in data anonymization is the removal of all direct personal identifiers data elements that can be directly linked to a specific individual such as names, telephone numbers, social media identifiers, and so on. this step is obvious and inexperienced data curators may assume it is sufficient. however, demographic variables can also pose risk. a common example might be occupation and geography; if the dataset includes name of town and occupation, then the only doctor in a small town is at risk of being identified. add additional variables such as ethnicity, country of origin for immigrants, gender, and family structure, and even doctors in larger cities might be at risk of reidentification. these variables persistent demographic characteristics of people that might be used to discover their identities are known as quasi-identifiers. the problem of reidentification with quasi-identifiers becomes more acute when you consider not just the unlikely possibility of the hypothetical doctors’ neighbours reading through the dataset and recognizing them, but the unfortunately plausible possibility that an intruder might use external information from directories and other public sources to attempt to reidentify people maliciously, for fun, or for profit. a variable should only be considered a quasi-identifier if an intruder could plausibly match that variable to information from another source. some variables may be used to derive other quasiidentifiers; for example, community size could be combined with a broader geographic grouping to guess the precise community someone lives in. remaining variables in the dataset are nonidentifying variables and will include opinions and ratings, temporary measures such as recent food consumption or exercise, and other research questions. risk is created when there is the potential for an intruder to link external information to identifiers or quasi-identifiers in a dataset to gain additional information about individuals. the question then becomes, how do you decide whether there are people in a dataset who are at risk of being reidentified, and how do you keep this from happening? in the example above, an obvious approach might be to remove the quasi-identifiers of ‘town’ and ‘occupation’ from the dataset. but there might still be some unusual combination of ethnicity, family structure, age and other demographic variables that could lead to the reidentification of an individual. removing all quasi-identifiers would remove risk from the dataset but is usually overkill; a variable such as gender or age on its own does not pose significant risk. and demographic variables are of great utility to researchers. every variable removed from the dataset decreases the dataset’s research value. rather than removing variables completely, a common practice is to group values into broad categories; a variable such as occupation might be grouped into categories of professional and nonprofessional occupations or according to some other scheme, age might be grouped into 10 year categories, and similar measures may be taken with other variables. but the question remains: how can we be sure we have done enough to ensure that participants remain anonymous? https://doi.org/10.29173/iq979 3/13 thompson, kristi and sullivan, carolyn (2020) mathematics, risk, and messy survey data, iassist quarterly 44(4), pp. 1-13. doi: https://doi.org/10.29173/iq979 introduction to k-anonymity a common approach is k-anonymity, one type of what is called data analytic risk assessment (dara) or statistical disclosure risk assessment (eliot, mckay et al, 2015). k-anonymity was first developed by computer scientists sweeney and samarati (1998) and has formed the basis of formal data anonymization efforts since then. ayala-rivera, mcdonagh et al (2014) describe it as a ‘fundamental principle of privacy’ and ‘widely discussed and adopted in a variety of domains’. simi et al call it the ‘primary model proposed for microdata anonymization’ and ‘the base from which further expansions have been developed.’ the concept behind k-anonymity is relatively straightforward: it should not be possible to isolate fewer than k individual cases in your dataset based on any combination of identifying variables, where k is an integer set by the researcher, typically 5. that is, a record cannot be distinguished from k-1 other records. a minimum of 3 is commonly suggested for k; in practice a value of 5 is often used. according to el emam and dankar (2008) data custodians should ‘select a value of k commensurate with the re-identification probability they are willing to tolerate—a threshold risk.’ they also note that ‘it is uncommon for data custodians to use values of k above 5 .‘ an example may help illustrate this concept. imagine your survey has four demographic variables: marital status, age group, gender, and ethnic group. if an individual in the dataset is white, married, over 65, and female, then for the data to have k-anonymity with k=5, there must be at least four other individuals in the dataset with the same set of characteristics. this also must be true for every other individual in the dataset; each person must have at least four data twins. even if an intruder knew that an individual was in the dataset and was able to match their characteristics against the data, they would not be able to tell which of the five cases was the target individual. a set of data twins, or cases with the same values on all potentially identifying variables, is called an equivalence class. an individual in the dataset who does not have any data twins an equivalence class of one is called a sample unique. this individual is at risk of being re-identified. if the dataset is a complete sample of a small population (for example, employees at a particular company) then this sample unique will also be a population unique, and an intruder looking at this dataset will be able to definitively identify this person. even if the dataset is not a complete sample, it is still possible that this person may be a population unique or have some combination of rare characteristics that makes it easy to narrow down their identity to a small number of candidates. k-anonymity is not always sufficient k-anonymity is a useful technique for limiting the possibility of exposure of the identity of the individuals in a data set. however, it may not always be sufficient to prohibit attribute disclosure, as many articles have noted. consider the above dataset, hypothetically a complete sample of all employees at a particular company. by grouping and categorizing variables, we have achieved kanonymity with a k of 5 the smallest equivalence class in the dataset contains at least five respondents. none of the individuals in the dataset can be identified with any certainty. however, imagine that all the respondents in an equivalence class answered a question the same way for example, all the employees at our company with a specific combination of characteristics answered ’yes‘ to a question about organizing a union. if a manager at this company looked at the data, that https://doi.org/10.29173/iq979 4/13 thompson, kristi and sullivan, carolyn (2020) mathematics, risk, and messy survey data, iassist quarterly 44(4), pp. 1-13. doi: https://doi.org/10.29173/iq979 manager would now know an attribute (interest in joining a union) about everyone in that equivalence class. k-anonymity is not sufficient to prevent attribute disclosure. respondents to surveys are generally told that their responses will be kept confidential, not merely that no one will know which line of data contains their specific answers; a k-anonymous dataset may not fulfill that promise. variations on k-anonymity, such as p-anonymity and l-diversity, have been developed in an effort to deal with the attribute disclosure issue. however, as domingo-ferrer and torra (2008) note in their review of k-anonymity and its variants, ‘neither k-anonymity nor its enhancements … are entirely successful in ensuring that no privacy leakage occurs while keeping a reasonable data utility level.’ a brief explanation of one of the variants of l-diversity will serve to illustrate this problem. a data set is said to satisfy distinct l-diversity if, for each group of records sharing a combination of demographic attributes (an equivalence class), there are at least l different values for each confidential variable. in our example workplace dataset, every group of data twins would need to include both yes and no responses to the union question. however, the effort needed to check for and recode to ensure this by hand, assuming that, as in many datasets, there are dozens of responses that need to be kept confidential, is daunting. in addition, one can easily imagine a scenario where nearly every attribute needed to be manipulated, grouped or partially suppressed in some way, or alternatively every equivalence class needed to be enlarged. as we will see later in this paper, a relatively small set of quasi-identifying variables can easily lead to a very large number of potential equivalence classes in a dataset. multiply that number by the set of attribute responses that would need to be considered and you will have a sense of the scope of the problem. at the end of all that manipulation, the resulting dataset would probably have very limited analytic value. the manipulation would have the result of destroying close relationships between the quasi-identifying demographic variables and the remaining variables in the dataset and determining accurately the correlation between demographic variables and attributes is a key part of research. k-anonymity is not always necessary until now we have been considering an example dataset which surveys an entire population a set of workers at a particular place of employment. it is very difficult to ensure confidentiality of responses in datasets of this nature. however, many datasets do not survey entire populations but are instead sample surveys. imagine that only one in 10 of the workers were surveyed. even if a worker belonged to an equivalence class of one, it would not be clear from the dataset whether that worker might have ‘data twins’ outside the dataset a sample unique might or might not be a population unique, although someone with perfect knowledge of the population being sampled might still be able to determine that individual’s identity in the case that they were a true population unique. sampling can also protect against attribute disclosure, since members of an equivalence class are likely to have data twins outside the dataset whose potential responses to confidential questions are unknown. according to eliot, mckay et al (2015), ‘sampling is one of the most powerful tools in the toolbox. the key point is that it creates uncertainty that any given population unit is even in the data at all.’ the larger the population from which a sample is drawn, the less likely an intruder is to have perfect knowledge of the population, and the less likely it is that a sample https://doi.org/10.29173/iq979 5/13 thompson, kristi and sullivan, carolyn (2020) mathematics, risk, and messy survey data, iassist quarterly 44(4), pp. 1-13. doi: https://doi.org/10.29173/iq979 unique will also be a population unique. we argue that attribute disclosure in the absence of identity disclosure is not a genuine concern in the case of a sample being drawn from a large population. if an equivalence class in the dataset can be assumed to have many co-equivalents in the general population being sampled whose non-identifying attributes or opinions are unknown, then membership in an equivalence class cannot be said to reveal confidential attributes or opinions2. national anti-drug strategy survey the national anti-drug strategy (nads) survey series is a series of surveys conducted by the environics research group on behalf of the government of canada during the years 2008 – 2012 (environics research group, 2012). an initial baseline survey assessed the behaviours and attitudes of teens aged 13-15 years towards drug use through responses collected from both teens and parents of teens. after this baseline data was collected, health canada launched an anti-drug media campaign. subsequent surveys followed a large proportion of the original respondents as well as new participants to gauge the potential impact of different aspects of the campaign. in this paper we are concerned with deidentification of the baseline survey, which had 1502 respondents. this dataset is highly sensitive; not due to the population under examination (this is a survey of the general population, and while respondents were adolescents at the time, all participants are now adults), but due to the subject. survey responses that confirmed acquisition and use of illegal drugs by participants and related to their comfort level in discussing these issues with family members could still impact these individuals and their relationships. anonymization of this dataset did not require removal of direct identifiers, as none were provided in the version of the dataset made available to the authors. potential quasi-identifiers noted within the dataset included subject age, gender, region of residence, household composition (i.e. number of parents, presence of older siblings), aboriginal status and visible minority status. based on factors including lists of key quasi-identifiers to check for as well as concepts such as attribute persistence (household composition as an adolescent will not persist into adulthood) we decided to focus on age, sex, geographic region, visible minority status and aboriginal status. some necessary context: in canadian government data, aboriginal status is considered an ethnopolitical category. survey respondents may be aboriginal or visible minority but are only considered both if they are a mix of aboriginal and some other visible minority group, such as asian. visible minority status is coded as a binary variable; respondents who are not caucasian or aboriginal are considered a visible minority. in this dataset age had three categories as respondents ranged from 13 to 15, and region had seven categories: one category for canada's four small atlantic provinces and six representing each of canada’s remaining six provinces. documentation informed us that respondents from the three northern territories were grouped with the geographically nearest province. the remaining variables were binary. simple multiplication would suggest a total of 168 possible equivalence classes, but since the aboriginal and visible minority statuses are effectively mutually exclusive the actual total is 126 possible equivalence classes. if these were distributed equally across the dataset, we would expect each equivalence class to contain about 12 cases. however, as is usual in practice, they are https://doi.org/10.29173/iq979 6/13 thompson, kristi and sullivan, carolyn (2020) mathematics, risk, and messy survey data, iassist quarterly 44(4), pp. 1-13. doi: https://doi.org/10.29173/iq979 not distributed equally across the dataset. some classes are much larger than others. the aboriginal group, for example, made up less than 5 percent of the sample, and the regions are similarly unbalanced with the largest two making up over half the sample. when the equivalence classes were calculated (see appendix for code samples in stata and r) we found that our dataset had 21 equivalence classes with only a single member, and a total of 42 equivalence classes with less than 5 members. clearly, even looking at a relatively small number of variables with only a few categories each, in a fairly large dataset, it is very difficult to produce a dataset that satisfies k-anonymity let alone any more stringent criteria. we were able to achieve k-anonymity by deleting the region variable; on the remaining four variables there were no equivalence classes smaller than 5. but how risky would it have been to include the region variable in this survey? this survey is a sample of a much larger population, the population of people aged 13 to 15 in canada in 2008. recall that if an equivalence class in the dataset can be assumed to have many co-equivalents in the general population, then membership in an equivalence class cannot be said to reveal confidential attributes or opinions. but how does one know if this is a safe assumption? the first co-author decided to investigate further. the census of canada 20163 (statistics canada, 2019) is a complete sample of the population of canada; as such it includes the complete population of people who were age 13 to 15 in 2009, minus any intervening deaths or emigrations. a public use sample is made available for academic researchers to download; although this is a subset of the full data, this sample includes a weight variable that allows it to be weighted back to represent the full population. i decided to use this to determine the size of the general population equivalence classes for our sample unique cases. it was straightforward to create a version of the census of canada dataset that matched the variables and population of the nads. the census of canada also includes questions on gender, visible minority status, and aboriginal identity. the latter two variables needed to be recoded to binary variables, but this could be accomplished unambiguously. the province / territory variable was similarly recoded to match the nads region variable, with the atlantic provinces being grouped and the populations of yukon, nunavut and the northwest territories each being assigned to the nearest province. the one variable which posed a slight difficulty was the age variable. nads respondents would have been age 20 to 22 in 2016; the census microdata file divided age into groups including a category for 20 24. i created an artificial variable for ages 20 to 22 by randomly assigning one fifth of the 20 to 24 year old respondents to each of the categories of 20, 21 and 22. the remaining two fifths of the 20 to 24 year old respondents, representing the group that was 23 and 24 years old, were dropped from the dataset along with other respondents outside the targeted age range. this left a nationally representative sample that matched the population from which the nads was drawn that could be used to estimate the population risk of the variables under consideration. when i calculated the equivalence classes using the weighted canadian census of 2016, the smallest equivalence class was estimated to have 370 cases, with the next smallest containing 518, and the remaining 214 equivalence classes being considerably larger. each sample unique in our survey was estimated to have a minimum of 369 data twins in the general population. even though we did not come close to achieving k-anonymity with this set of variables, the sampling factor means that even https://doi.org/10.29173/iq979 7/13 thompson, kristi and sullivan, carolyn (2020) mathematics, risk, and messy survey data, iassist quarterly 44(4), pp. 1-13. doi: https://doi.org/10.29173/iq979 with the region variable this dataset is low risk, with a population reidentification risk of at most 1/370 for the riskiest case, rather than the sample estimated risk of 1. the approach outlined here is straightforward to implement and may provide a good option for estimating population risk in appropriate cases. this approach will only work if the sample is drawn from a known population, there exists a survey or census that can be weighted to that population, and that survey contains the correct set of demographic variables. in sample surveys of some clearly defined subset of the general population, a national census may be an appropriate choice, although other large, nationally representative surveys might be used. surveys of smaller and more distinct populations that do not approximate a national sample of the general population or a national sample of a readily defined subset of the national population are inherently of higher risk and cannot have their population risk threshold estimated using a national sample. in cases of relatively small samples of a defined population it is reasonable to enforce kanonymity with a k of 5, while relying on the sampling factor to deal with the concern of revealing opinions or attributes within equivalence classes. if a survey is a complete sample of some defined population (e.g. every person attending a school) or a large fraction of that population then the dataset is of very high risk and should only be shared with extreme caution regarding presence of quasi-identifiers, if it contains no sensitive information, or if the survey respondents were not promised confidentiality. drinking water quality survey in addition to anonymity testing using the type of statistical disclosure risk assessment exemplified by k-anonymity, a researcher may also attempt to check the sensitivity of a dataset to identity disclosure through penetration testing, as the second co-author will demonstrate in our next example. a series of surveys completed by ekos research for health canada in collaboration with indian and northern affairs canada during 2007, 2009, and 2011 assessed water quality on first nations reserves and rural communities (ekos research associates, 2011). here we will consider data from the 2007 survey. demographic data collected from survey participants included their year of birth/age, a binary gender, linguistic preference (english/french), aboriginal status, whether they lived on or off reserve for at least six months of the year, the number of persons in their household, number and age by category of dependent children, number of seniors and vulnerable adults in their household, and whether their home was used as a daycare facility. ekos also collected information on their place of residence, including whether a drinking or boil water advisory was currently in effect for their region, how many times their community had been under a drinking or boil water advisory during the past five years, forward sortation area (fsa) (a grouping of postal codes often used for releasing census information in canada), province of residence, rural/urban status, their distance from the nearest city, and the population of their community. as with the nads survey, special attention must be taken to prevent re-identification of survey participants, though for different reasons even beyond the usual promise of confidentiality given to survey respondents. in the nads dataset, full anonymization was particularly important because of https://doi.org/10.29173/iq979 8/13 thompson, kristi and sullivan, carolyn (2020) mathematics, risk, and messy survey data, iassist quarterly 44(4), pp. 1-13. doi: https://doi.org/10.29173/iq979 the sensitivity of the topic (illegal drugs); here, anonymization is crucial due to the population being considered, as first nations individuals and communities have been systematically marginalized within canada. as may be anticipated, these data in combination can be used as key variables by a data intruder to reidentify a survey participant. as the location of first nations reserves in canada are publicly known and there is often only a single reserve within a given fsa, a data intruder can easily identify the community of origin for participants who identified as living on-reserve. in the cases for which there are multiple reserves within a given fsa, the data intruder may still discriminate between candidates for the community of origin by comparing the reported population of the community against the 2006 census records, or juxtaposing reported water advisories against those listed in old news documents, data gathered by civic action groups like water today or the government of canada’s list of long-term water advisories (2020). even if the data intruder is unable to choose between multiple candidates for a community of origin at this phase though, the populations of first nations reserves are sufficiently small that individuals demographically unique to the sample will likely be unique to the population of a fsa, especially as the population of individuals living on reserve in a fsa may only number in the tens or hundreds. this data in its original form then threatens the anonymity of survey participants as individuals. it is also worth noting that even if an individual is not sample-unique or population-unique, if all demographically similar individuals can be located to the same reserve, or if all the candidate reserves for these individuals fall under the same tribal governance, locating the participants’ community of origin may still pose a problem. historically, data on first nations communities has influenced perceptions of them as a group, which is one reason why the research principles of ocap4, standing for ownership, control, access and possession were developed (first nations information governance centre, 2014); ocap understands ownership of data as the collective responsibility of the first nations people from whom it originates. anonymizing this survey data for public release should then ensure the community of origin for this data cannot be established. a naive attempt to anonymize this survey may assume removal of the fsa to be sufficient. as we will demonstrate, this strategy underestimates the ability of data intruders to use publicly available information and simple programming skills to recover the fsa from still-included information on distance to the nearest city through data linkage. by way of proof, my co-author and i attempted an experiment. she presented me with a dataset from which the fsas had been suppressed. i constructed a database containing the name, province, population, forward sortation area, location, and distance to the nearest town with population over 15,000 of every first nations reserve in canada using resources available to anyone, regardless of academic or governmental affiliations. information on name, province, and population of first nations reserves were scraped from wikipedia pages using python and the beautifulsoup code library (richardson, 2020). given this information, i then used the government of canada’s geolocation webservices (2018) to find the latitude and longitude of these communities. geonames.org’s findnearbyplacenamejson and findnearbypostalcode functions (geonames team, 2020) were used to discover the postal code of each first nations reserve, and their straight-line distance to the nearest place with population equal to or greater than 15000 (this being the most appropriate filter accessible through the app). having created this database, i then wrote a programming script that would compare the distanceto-the-nearest-city reported by each respondent who had identified as living on reserve, to the https://doi.org/10.29173/iq979 9/13 thompson, kristi and sullivan, carolyn (2020) mathematics, risk, and messy survey data, iassist quarterly 44(4), pp. 1-13. doi: https://doi.org/10.29173/iq979 distance-to-the-nearest-city for each reserve in that respondent’s province within my database. as the distance-to-the-nearest-city recorded in my database was a straight-line distance, while that reported by survey participants was an estimation, an exact match could not be expected. i decided then to create a list of candidate reserves for each survey participant based on whether their estimated distance-to-the-nearest-city came within a given margin of error to the distance-to-thenearest-city recorded in my database. for 98 of the participants living on-reserve, only a single reserve appeared possible for their location. when my co-author compared my ‘guesses’ for the fsa of these individuals to the information present in the non-deidentified dataset, 25% of them were correct. while the correct guesses only amounted to 24 individuals out of 98 supposed correct guesses, out of 1114 individuals surveyed, this experiment demonstrates the risk presented by a data intruder with only simple programming skills and access to public information. accuracy of the constructed database and its distance-to-the-nearest-city could be improved through use of geographic information systems (gis) software, which allows measurements and calculations to be made using digital maps. for example, the population size considered to be a city by the ekos survey was not always consistent with the results returned to me using the geocoding app, which could only limit ‘cities’ by populations over 5000, 10000, or 15000. an early iteration of the code, using the definition of a city as an area populated with over 5000 people was of low accuracy in predicting participant locations. it is expected that accuracy would improve if a data intruder could select cities of a population closer to that used by the ekos survey. a wide margin of error had to be used to account for the low accuracy of comparing a straight-line distance-to-the-nearest-city within the constructed database to the estimated distance-to-the-nearest-city by road. gis layers, such as dmti route maps, could be used to find the shortest distance by transportation to the nearest city. this would enable us to decrease our allowed margin-of-error, decrease the number of candidate reserves possible for each respondent living on reserve, and increase the number of respondents for which a single location can be positively identified. based on the results of this penetration test, the variable ‘distance to nearest city’ will be dropped from any publicly accessible version of this dataset. discussion the authors’ experiences should help illuminate the complexity of determining if a survey dataset has been successfully deidentified and provide some guidance into deidentifying other survey datasets. as a practical approach, steps for deidentification of a dataset might include, first, the removal of all direct identifiers. the set of risky quasi-identifiers to be preferentially retained needs to be identified next. frequency tables can be used to identify small categories on these quasiidentifiers and determine appropriate groupings. (‘small’ is relative and will depend on the size of the sample and the size of the population from which a sample was drawn. as a first pass, groups smaller than 5% of the population might be considered.) bivariate tables of the grouped quasiidentifiers can be used next to identify variables that produce small groups. (software such as the program amnesia or the r package sdcmicro can help automate this process but are beyond the scope of this paper.) the data custodian may wish to consider suppressing individual values rather https://doi.org/10.29173/iq979 10/13 thompson, kristi and sullivan, carolyn (2020) mathematics, risk, and messy survey data, iassist quarterly 44(4), pp. 1-13. doi: https://doi.org/10.29173/iq979 than regrouping at this stage. for example, in a unique case of a respondent being married and under age 16, the custodian might delete the response to the marriage question instead of regrouping the otherwise non-risky variables of ‘age group’ and ‘marital status’. larger groups of variables can be iteratively investigated to locate potentially small groupings, until the data custodian comes up with a final set of equivalence classes based on the full list of modified risky variables. if the dataset has achieved k-anonymity with an appropriate value of k (usually 3 or 5) the dataset may be considered provisionally safe. if unacceptably small equivalence classes remain in the dataset and the data custodian would prefer not to drop or regroup variables any further, at this stage the population k-values of the equivalence classes can be checked using an appropriate large national population-weighted dataset, if one can be located. if the dataset has a low population reidentification risk, current variables may be retained, otherwise they will need to be dropped, suppressed or grouped further. as a final step, variables that relate to geography in any way should be treated with extreme caution. as we have demonstrated, non-obvious geographic variables such as distance from nearest city, combined with contextual information such as survey respondents living on a reservation, can be used to pinpoint geographic location with surprising precision. other geography-adjacent variables that might need to be considered in relation to contextual survey information might include community size and presence or lack of resources such as a major hospital or public airport in a community. as penetration testing is likely to be beyond what is practical as a part of routine data deidentification, the data custodian should be proactive in considering whether there is a strong analytic interest in retaining such variables and dropping them if there is not. references ayala-rivera, v., mcdonagh, p., cerqueus, t. and murphy, l. (2014). ‘a systematic comparison and evaluation of k-anonymization algorithms for practitioners’. transactions on data privacy, 7(3), pp.337-370. available online: http://hdl.handle.net/10197/9109 canadian institutes of health research (cihr), the natural sciences and engineering research council of canada (nserc), and the social sciences and humanities research council of canada (sshrc) (2018). ‘draft tri-agency research data management policy for consultation’. available at https://www.science.gc.ca/eic/site/063.nsf/eng/h_97610.html domingo-ferrer, j. and torra, v. (2008). ‘a critique of k-anonymity and some of its enhancements’. third international conference on availability, reliability and security, ieee, pp. 990-993. available at https://doi.org/10.1109/ares.2008.97 ekos research associates inc. (2009). ‘water quality on-reserve quantitative research’. available at http://www.ekospolitics.com/articles/0559.pdf ekos research associates inc. (2011). ‘perceptions of drinking water quality in first nations communities and general population’. available at http://www.ekospolitics.com/articles/015-11.pdf https://doi.org/10.29173/iq979 http://hdl.handle.net/10197/9109 https://www.science.gc.ca/eic/site/063.nsf/eng/h_97610.html https://doi.org/10.1109/ares.2008.97 http://www.ekospolitics.com/articles/0559.pdf http://www.ekospolitics.com/articles/015-11.pdf 11/13 thompson, kristi and sullivan, carolyn (2020) mathematics, risk, and messy survey data, iassist quarterly 44(4), pp. 1-13. doi: https://doi.org/10.29173/iq979 elliot, m., mackey, e., o'hara, k. and tudor, c. (2016). the anonymisation decision-making framework, manchester: ukan. available at https://ukanon.net/wp-content/uploads/2015/05/theanonymisation-decision-making-framework.pdf environics research group (2012). ‘2012 national anti‐drug strategy (nads) youth advertising recall and tracking survey’. available at https://epe.lac-bac.gc.ca/100/200/301/pwgsc-tpsgc/poref/health/2012/068-11/report.pdf first nations information governance centre (2014). ‘ownership, control, access and possession (ocap™): the path to first nations information governance’. available at https://fnigc.ca/sites/default/files/docs/ocap_path_to_fn_information_governance_en_final.pdf geonames team (2018). ‘geonames’. available at https://www.geonames.org/export/webservices.html government of canada | gouvernement du canada. (2014). canada's action plan on open government 2014-16. available at https://open.canada.ca/en/content/canadas-action-plan-opengovernment-2014-16 government of canada | gouvernement du canada. (2018). ‘geolocation service’. available at https://www.nrcan.gc.ca/earth-sciences/geography/topographic-information/webservices/geolocation-service/17304 government of canada | gouvernment du canada (2020). ‘ending long-term drinking water advisories’, government of canada | gouvernement du canada. available at https://www.sacisc.gc.ca/eng/1506514143353/1533317130660 richardson, l. (2020). ‘beautiful soup 4.9.1’. available at https://www.crummy.com/software/beautifulsoup samarati, p. and sweeney, l. (1998). ‘protecting privacy when disclosing information: k-anonymity and its enforcement through generalization and suppression’. proceedings of the ieee symposium on research in security and privacy, oakland, ca. available at https://dataprivacylab.org/dataprivacy/projects/kanonymity/paper3.pdf simi, m s; sankara, nayaki and sudheep elayidom, m. (2017). ‘an extensive study on data anonymization algorithms based on k-anonymity’. iop conf. series: materials science and engineering, 225. available at https://doi.org/10.1088/1757-899x/225/1/012279 statistics canada (2019). ‘2016 census of population [canada] public use microdata file (pumf): individuals file [public use microdata file]’, ottawa, ontario: statistics canada [producer and distributor]. accessed through odesi, http://odesi.ca thompson, k., (2018). ‘documentation as data rescue: restoring a collection of canadian health survey files’. against the grain, 29(6), p.12. available at https://docs.lib.purdue.edu/cgi/viewcontent.cgi?article=7876&context=atg https://doi.org/10.29173/iq979 https://ukanon.net/wp-content/uploads/2015/05/the-anonymisation-decision-making-framework.pdf https://ukanon.net/wp-content/uploads/2015/05/the-anonymisation-decision-making-framework.pdf https://epe.lac-bac.gc.ca/100/200/301/pwgsc-tpsgc/por-ef/health/2012/068-11/report.pdf https://epe.lac-bac.gc.ca/100/200/301/pwgsc-tpsgc/por-ef/health/2012/068-11/report.pdf https://fnigc.ca/sites/default/files/docs/ocap_path_to_fn_information_governance_en_final.pdf https://www.geonames.org/export/web-services.html https://www.geonames.org/export/web-services.html https://open.canada.ca/en/content/canadas-action-plan-open-government-2014-16 https://open.canada.ca/en/content/canadas-action-plan-open-government-2014-16 https://www.nrcan.gc.ca/earth-sciences/geography/topographic-information/web-services/geolocation-service/17304 https://www.nrcan.gc.ca/earth-sciences/geography/topographic-information/web-services/geolocation-service/17304 https://www.sac-isc.gc.ca/eng/1506514143353/1533317130660 https://www.sac-isc.gc.ca/eng/1506514143353/1533317130660 https://www.crummy.com/software/beautifulsoup https://dataprivacylab.org/dataprivacy/projects/kanonymity/paper3.pdf https://doi.org/10.1088/1757-899x/225/1/012279 http://odesi.ca/ https://docs.lib.purdue.edu/cgi/viewcontent.cgi?article=7876&context=atg 12/13 thompson, kristi and sullivan, carolyn (2020) mathematics, risk, and messy survey data, iassist quarterly 44(4), pp. 1-13. doi: https://doi.org/10.29173/iq979 u.k. economic and social research council (2018). ‘research data policy’. available at https://esrc.ukri.org/funding/guidance-for-grant-holders/research-data-policy/ u.s. national science foundation (2011). ‘dissemination and sharing of research results’. available at https://www.nsf.gov/bfa/dias/policy/dmp.jsp 1kristi thompson is the research data management librarian at western university and can be reached by email: kthom67@uwo.ca. carolyn sullivan is a student in the faculty of information and media studies at western university. 2 inferential disclosure determining from a dataset that having a particular combination of characteristics makes an individual more likely to possess some attribute is occasionally mentioned in the literature. as forming inferences is the point of most research this is not something that can be eliminated in the general case. more generally the issue of community stigmatization should be considered as a part of the general ethical review of datasets, but this is not precisely a deidentification problem and is beyond the scope of this article. 3 a long form census of canada was not conducted in 2011 as it would usually have been, so 2016 was the closest available census occurring after the 2009 survey. 4 the history of research on first nations peoples in canada is complex and deeply problematic. as an official government survey this data on water collection did not fall under ocap, but we felt the principle of first nation community rights to privacy should apply when we prepared the public version for deposit. https://doi.org/10.29173/iq979 https://esrc.ukri.org/funding/guidance-for-grant-holders/research-data-policy/ https://www.nsf.gov/bfa/dias/policy/dmp.jsp mailto:kthom67@uwo.ca 13/13 thompson, kristi and sullivan, carolyn (2020) mathematics, risk, and messy survey data, iassist quarterly 44(4), pp. 1-13. doi: https://doi.org/10.29173/iq979 appendix: code for finding equivalence classes in stata and r -stata - * stata code for checking k-anonymity * kristi thompson, may 2020 * create the equivalence groups egen equivalence_group= group(var1 var2 var3 var4 var5) * create a variable to count cases in each equivalence group sort equivalence_group by equivalence_group: gen equivalence_size =_n * list the id numbers of equivalence groups containing 3 or fewer cases tab equivalence_group if equivalence_size < 3, sort * list the values of the quasi-identifiers for each small equivalence class. e.g. if 1 list var1 var2 var3 var4 var5 if equivalence_group == num --r - # r code for checking k-anonymity # carolyn sullivan, may 2020 # install plyr, a useful data manipulation package. install.packages("plyr") # load the library. library('plyr') datafile <" location of the data file csv format " # read the csv file. df <read.csv (datafile) # figure out what equivalence classes there are, and how many cases in each equivalence class. dfunique <ddply(df, .(var1, var2, var3, var4, var5), nrow) dfunique <dfunique[order(dfunique$v1),] view(dfunique) https://doi.org/10.29173/iq979 iassvol194 25winter 1995 background over the past fifteen years, there have been several collaborative studies of the archival value of scientific records in the united states. between 1978 and 1983, representatives of the history of science society, the society of american archivists, the society for the history of technology, and the association of records managers and administrators worked together on the joint committee for the archives of science and technology (jcast), assessing the state of documentation on research and development, the dissemination of ideas, technology transfer, and professional education in science and technology2. a self-acknowledged sequel to the jcast project, was the collaboration of joan haas, helen samuels and barbara simmons at the massachusetts institute of technology (mit) which resulted in the publication of appraising the records of modern science and technology: a guide in 19853. beginning in 1989, the center for the history of physics of the american institute of physics (aip) inaugurated a long term study of fields of physics and related sciences where multi-institutional collaborations are prominent4. while the three projects were undertaken in different organizational contexts, with varying focus and goals, they share a primary concern with the records of research and development activities as resources for historical research. in these endeavors, the potential long-term value of the records of science and technology for further research in these fields themselves has been recognized, but not explored in depth. generally, consideration of enduring value for science has focused on the data records generated in research and development activities. in fact, jcast declared, “the first consideration regarding retention of data must be the needs scientists themselves have for these records.” both the jcast and mit publications recognized that the actual retention of data for science is usually in the hands of the scientists themselves or of specialized scientific data centers, although archivists may occasionally face the necessity of deciding on the retention of scientific data for scientific purposes5. the aip project took into account the future needs of physicists. it identified categories of records which should be retained by scientific laboratories and science libraries, but did not articulate criteria for identifying records with continuing value for science. both the jcast report and the mit guide drew attention to the distinction between observational and experimental data, and suggested that long-term value is more often found in observational data than in experimental results. the argument which supports this generalization is that experiments are repeatable, while observational data of ten relate to unique or rare events or sequences of events. with respect to scientific data, the aip has concluded that, in high energy physics at least, very little data should be preserved for long periods, and then for purposes of exhibit, rather than scientific research6. scientific records in the national archives of the united states the national archives and records administration (nara) of the united states has been involved in appraising and preserving the records of science and technology since its inception. the types of scientific records in the national archives include project case files, technical reports, laboratory notebooks, drawings and specifications, maps, charts, graphs, aerial photography, motion pictures, sound recordings, and digital data files. the subjects reflect the broad range of scientific and technical activities in which the government of the united states has engaged, including astronomical, geological and meteorological observations, land and stream classifications, patents, weights and measures, nuclear energy, mineral deposits, weapons systems, aircraft and spacecraft, epidemiology and biometry, entomology, and many other subjects. it is probably impossible to categorize in general terms the reasons why such scientific and technical records have been accessioned into the national archives. however, among other factors, nara has been concerned with the continuing value of these records for science and technology themselves7. concern with the long term scientific value of records derives from a crucial provision of u.s. law, the definition of a federal records. this definition, articulated in title 44 of the united states code, states that federal records are preserved or appropriate for preservation either as evidence or “because of the informational value of the data in them.” as jcast recognized, the primary informational value of scientific data is for scientists. in appraising the records of science, then, nara has an obligation to consider their continuing value to scientists. the most obvious means of exploring this value is to consult with scientists, as was the case in all three f the earlier projects i mentioned. in fact, the reports of all three projects recommended this practice as standard. in the case of research funded by the u.s. governpreserving scientific information on the physical universe by kenneth thibodeau 1 26 iassist quarterly ment, there are especially important reasons to include the perspective of specialists: (1) the functions of federal agencies in sponsoring research, (2) the role of researchers outside of the government in the life-cycle of the data, and (3) the knowledge of the provenance and life-cycle of the data that these scientists have. (1) agency functions: commonly, federal records are documents used by agencies in the exercise of mission functions. the internal revenue service, for example, collects tax returns as instruments critical to its function of collecting taxes. often, however, the science agencies do not use the data that result from the research they sponsor; rather, their function is to sponsor research. in many cases where the funding agency requires the researchers to deliver the resultant data to the government, the primary, if not sole, purpose for this is to make the data available to yet other researchers outside of the government, in many cases for long periods of time. (2) role of outside researchers: a large proportion of the scientific data generated by the federal government is created as a result of the initiatives of investigators outside of the government. outside scientists of ten originate the research proposals and are responsible for the organization and conduct of the research. through peer review of proposals, other members of the research community have a decisive role in determining what data are collected and how they are collected and organized. even in cases where research is conducted by scientists employed by an agency, it is not uncommon to have government laboratories reviewed by peer groups composed of outside scientists. (3) the life-cycle of scientific data: in many cases, scientific data are in the custody of researchers outside of the government for a major part of their life-cycle. these researchers collect the data, calibrate and refine it, and analyze it, definitively shaping the record. in many cases, the records are not transferred to government custody until the records cease to be active. in other cases, the records although owned by the government remain in the custody of outside researchers throughout their life-cycle. the relevant framework of appraising scientific data sets, thus, is not defined by the business activities or the need for corporate memory of the sponsoring agency, but by the research community. seeking the input of scientists in the appraisal of the data recognizes that the roles and the actions of academic researchers are at least as important as the functions of the agency that funded the research or launched the satellite. since 1990, nara has sponsored two important efforts to obtain the advice of subject matter experts on the retention of data. the first study, undertaken by the national academy of public administration (napa), focused on major federal databases used in support of mission activities. this study had a twofold purpose: first, to identify these databases and, second, to recommend what data should be preserved in the national archives8. the napa project included the review of some scientific and technical data, notably in the areas of natural resources, the environment, and health. however, large collections of scientific data were intentionally excluded from consideration in the napa project because nara felt that such a large and complex area as “big science” merited separate attention. scientific data are the focus of the second recent study sponsored by nara. this project, inaugurated in 1992, was undertaken by the national academy of sciences’ national research council (nrc). the nrc study was divided into five subject areas: (1) space sciences; (2) physical, chemical and materials sciences; (3) earth sciences, (4) atmospheric sciences, and (5) ocean sciences. panels of experts were organized to develop recommendations for the preservation of records in each of these areas of research. a steering committee oversaw the work of the panels and formulated generalized recommendations and criteria for the retention of scientific records based on the work of the panels. the nrc project gave the national archives the opportunity to interact with the records creators and to engage them in a dialogue on the long term value of the data for secondary use. it was hoped that the nrc project would serve to raise the addressing the complete potential life-cycle of the records during the development and performance of research projects. the final report: preserving scientific data on our physical universe. a new strategy for archiving the nation’s scientific information resources<end underline>, moves towards fulfilling that hope9. the report, which was completed in march of this year, makes several sweeping recommendations which can be grouped under the twin headings of retention and responsibilities. retention recommendation: “as a general rule, all observational data that are nonredundant, useful, and documented well enough for most primary uses should be permanently maintained. laboratory data sets are candidates for long-term preservation if there is no realistic chance of repeating the experiment, or if the cost and intellectual effort required to collect and validate the data were so great that long-term retention is clearly justified10.” the report makes two procedural suggestions related to appraisal. the first is that each program or project should have a data management plan established at the origin and governing the entire life-cycle of the data: “planning activities at the point of data origin must include long-term data management and archiving.11” the second is that 27winter 1995 appraisal itself is a multifaceted, continuing process: “formal appraisals should be kept to minimum, appraisals should be performed according to the data management plan established for each project12.” the first step in this process would be an interdisciplinary consensus regarding broad classes of data: “all stakeholders... should be represented in the broad, overarching decisions regarding each class of data13.” “scientists, information technology professionals, data managers, librarians, and archivists must unify their expertise in the establishment of a coherent strategy for end-to-end data and information management14.” principal investigators and program managers would then appraise the long term value of individual data sets: “the appraisal of individual data sets ... should be seen as an ongoing, informal process associated with the active research use of the data, and therefore should be performed by the most knowledgeable about the particular data.... in some cases, they may need to involve an archivist or information resources manager to help with issues of long-term retention.15” finally, the judgements of the primary users would be supplemented with some sort of peer review. the purpose of this review would not be to appraise the data, but to determine if they are as purported and if they are adequately documented. several options for peer review are identified, ranging from “a formal peer review to certify integrity and completeness,” to “documented evidence of the use of the data set in publications in peer-reviewed journals,” or evidence from expert users that the data set “is as described in the documentation.” responsibilities recommendation: “as a general principle, data collected by an agency should remain with that agency indefinitely.16” “collection” in this context refer to data collected under agency sponsorship, through contracts and grants, as well as data created within the agency. the proposal is for scientific data to be held for the long term in distributed archives, typically in discipline oriented data centers, such as the national space science data center in nasa and the national geophysical and solar-terrestrial data center in noaa. the novelty in the approach advocated by the report is a proposal for coordination of the activities of such distributed archives “the federal government should create a national scientific information resource federation — an evolutionary and collaborative network of scientific and technical data centers and archives....17” this recommendation to establish the federation suggests action by the clinton administration’s information infrastructure task force, the national science and technology council, and/or the office of management and budget to initiate the federation. the report recommends that either an independent commission or an agency with an established mandate in both the physical sciences and information technology, such as the national science foundation, provide executive support for the federation. however, the organization is to be true federation; i.e., a collaboration among equals, not a top-down activity directed by the government. hypothetical profile of the life cycle of a data set {in an effort to understand the recommendations of the nrc report, i have constructed the following hypothetical profile of the life-cycle of a data set in accordance with these recommendations. this profile has been reviewed by both the project director and the chair of the steering committee. they both agreed that it accurately reflects the intent of the recommendations.} when a research project is proposed, the principle investigators, with appropriate collaboration by the program manger in the funding agency, draft a data management plan. in evaluating the long term value of the data, the principals consult with nara to learn of any recommendations that may have been made by diverse groups of stakeholders about the enduring value of data in the relevant class. the plan is completed no later than project initiation. following generally accepted criteria, the plan assumes that most observational data produced by the project will be subject to long term retention. laboratory or engineering data sets, however, are candidates for long-term preservation only if there is no realistic chance of repeating the work, of if the cost and intellectual effort required to collect and validate the data would be so great that long term retention is clearly justified. other experimental laboratory data or engineering data generally will not need to be retained after completion of the project. the provisions of the plan conform to well established standards for information technology and documentation, as endorsed by the nsir federation. nara maintains liaison with the sponsoring science agency and, periodically or when necessary, consults with the agency and the investigators to remind them of their responsibilities for long term retention, management, and access. during the conduct of the research, the investigators and managers will occasionally and informally consider whether the data actually collected merits retention and also whether the history of the project gives rise to any special requirements for documentation. at the end of the project, the data is deposited in a data center or field archives, as specified in the data management plan. this repository is designated, or operated, by the lead agency in the subject area, where staff have appropriate expertise and close ties with the relevant researcher community. the repository has mechanisms for access to the data by individuals beyond the primary users. for the data set to be accepted into the data center of distributed archives, it must undergo some form of peer review to ensure that it adequately meets the standards of uniqueness and accessibility. along with the data, metadata essential for others to use the data is transferred to the repository. when any required services related to retention or access to the data are available elsewhere in the federation at lower costs than would be incurred by the organization with primary responsibility, those alternative services are used when feasible. from the time the data set becomes available for use outside of the project, it is identified in a hierarchical information locator system. the data should be available for remote access and/or file transfer, ideally as an extension of the locator system. nara is informed of the existence and location of the data. nara monitors the preservation and accessibility of the data over time and, acting in an advisory capacity, helps the custodians with any problems in these areas. if the custodians can no longer meet the needs of the user community, or if the data is no longer in regular use, the data should be considered for transfer to some other federal science agency or, as a last resort, to the national archives. the nrc report has been received too recently for nara to have been able to take a position concerning its recommendations. however, as an individual archivist, i would like to offer some observations. they are entirely my own; they do not represent the position of the national archives not, as far as i know, of any other individual in the national archives. preserving scientific data on our physical universe purports to offer “a new strategy for archiving the nation’s scientific information resources.” there are certainly new elements in this strategy. one is the articulation of the general principle that data collected in the observational sciences should be preserved permanently. another is the creation of the nsir federation, conceived as a collaboration facilitated, but not directed, by the federal government. a major innovation entailed by the recommendations is that the scientific community would have to recognize data management, data retention and data access as valuable activities by scientists, and would have to adjust the culture of science to include rewards for these activities, on a par with the publication of research results. these things would be significant changes, but they would be changes in the scientific community. from an archival perspective, the impact would be different. as the report recognizes, very little scientific data has been deposited in the national archives, and there is no grounds for expecting this to change. the assertion that entities within the scientific community should be responsible for the long-term retention of scientific data sets and for access to them is not only consonant with the status quo, but also consistent with archival concepts, as articulated by the three projects i described at the start of this paper. furthermore, the national archives does not play a forceful role in the management, retention, or access to scientific data. the report, in fact, argues that such a role would be counter-productive. it suggests that the appropriate role for nara would be that of consultant and collaborator on archival preservation and access issues. on empirical grounds, one might say that this is role that nara has been playing in the domain of scientific data. thus, one might argue that, from an archival perspective, the recommendations of the nrc report reduce, by and large, to a confirmation of the status quo. references 1 the views stated in this paper are those of the author and do not represent the position of the national archives and redords administration. (paper presented at iassist95 may 1995 quebec city, quebec, canada.) 2 clark a. elliott, editor. understanding progress as process. documentation of the history of post-war science and technology in the united states. final report of the joint committee for the archives of science and technology. chicago. society of american archivists, 1983. 3 joan k. haas, helen willa samuels and barbara trippel simmons. appraising the records of modern science and technology: a guide. cambridge, ma. massachusetts institute of technology, 1985. 4 joan warnow-blewett and spencer weart. aip study of multi-institutional collaborations, phase i: high-energy physics. report no. 1: summary of project activities and findings. project recommendations. new york: american institute of physics, 1992. 5 elliott, p. 33-34. haas et al., pp. 60-61. 6 joan warnow-blewett, lynn maloney, and roxanne nilan. aip study of multi-institutional collaborations, phase i: high-energy physics. report no. 2: documenting collaborations in high-energy physics. new york: american institute of physics, 1992. pp. 75-76, 89-91. 7 trudy huskamp peterson. presentation on the national archives and the records of science for the national academy of science/national research council study on the long term retention of scientific and technical records. plenary session, july 7, 1993. 8 national academy of public administration. the archives of the future: archival strategies for the treatment of electronic databases. a report for the national archives and records administration. washington, d.c. national academy of public administration. 1991. 9 national research council. preserving scientific data on our physical universe. a new strategy for archiving the nation’s scientific information resources. commission on physical sciences, mathematics, and applications. washington d.c. national academy press. 1995. 10 ibid. p. 4. the report argues that technological developments make it possible to save everything and to provide access to it. however, the report recognizes that data management activities in general, not just preservation, are chronically underfunded and that they are at best a secondary concern in scientific culture. it does not adequately address how these difficulties can be overcome. 11 ibid. p. 50 12 ibid. p. 40 13 ibid. 14 ibid. p. 50 15 ibid. p. 4 16 ibid. p. 56 17 ibid. p. 51 1/2 rasmussen, karsten boye (2019) editor’s notes: standardization and certification save us from the frustrations of the greek drama, iassist quarterly 43 (1), pp. 1-2. doi https://doi.org/10.29173/iq953 editor's notes standardization and certification save us from the frustrations of the greek drama welcome to the first issue of volume 43 of the iassist quarterly (iq 43:1, 2019). the iassist quarterly presents in this issue three papers illustrated in the title above. chronologically we start from an early beginning. no, not with turing, we time travel further back and experience ancient greece. in this submission the greek drama delivers the form, while data librarians deliver the content on data sharing. and it makes you a proud iassister to know that altruism is the rationale behind data sharing. the drama continues in the second submission when librarians get frustrated because they suddenly find themselves first as data librarians and second as frustrated data librarians because ends do not meet when the librarians have difficulties servicing the data needs of their users, in combination with the users having unrealistic expectations. finally, the third article is about standardization and certification that makes the librarians more secure that they are on the right track when building a tdr (trustworthy digital repository). enjoy the reading. the first article is different from most articles. there is a first for everything! not often are we at iq offered a greek drama. and here is one on data sharing. the article needed the layout of a play so even the typeface of this contribution is different. the paper / play is called 'an epic journey in sharing: the story of a young researcher’s journey to share her data and the information professionals who tried to help’. the authors are sebastian karcher and sophia lafferty-hess at duke university libraries. the reason for using greek drama as a template is that form can help us think differently 'out of the box’! the play demonstrates the positive intention of data sharing, and by sharing contributing to something larger. the article references other researchers showing that scholarly altruism is a driving force for data sharers. no matter the good intentions of the protagonist, she finds herself locked in a situation where she is not able to take identifiable data with her when leaving the institution. and leaving the university is what undergraduates do. without the identification, it is impossible to obtain re-consent from participants. yes, it does look murky but there is even a happy ending in the epilogue. the second article is about librarianship, and how that task is not always easy. 'frustrations and roadblocks in data reference librarianship’ is by alicia kubas and jenny mcburney who work at the university of minnesota libraries. like many others, they have observed that many librarians find themselves as 'accidental data librarians'. that this brings frustration can be seen in the results of a survey they carried out. the methodology is explained, and descriptive statistics bring insight to what librarians do as well as to the frustrations and roadblocks they experience. let us start with the good news: some librarians are never frustrated with data questions. the bad news is that only 3% fall into that category. on the other hand, 83% mention 'managing patron expectations’ among their biggest frustrations. it sounds as if matching of expectations should be a course at library school. maybe it is already, and users with high expectations simply do not understand the complexity of the work involved. fortunately, some frustrations can be lessened by experience, but there are https://doi.org/10.29173/iq953 2/2 rasmussen, karsten boye (2019) editor’s notes: standardization and certification save us from the frustrations of the greek drama, iassist quarterly 43 (1), pp. 1-2. doi https://doi.org/10.29173/iq953 others – called roadblocks, e.g. paywalls or lack of geographic coverage – that all librarians meet. among the comments after the survey was that data persist as a difficult source type for librarians to support. the questionnaire developed and used by kubas and mcburney is found in an appendix. the last article in this issue raises sustainability as an important issue for long term data preservation, and the concept forms part of the title of the submission 'coretrustseal: from academic collaboration to sustainable services'. the paper is from an international group of authors comprising hervé l'hours, mari kleemola, and lisa de leeuw from uk, finland and the netherlands. the seal is a certification for repositories curating data. the last sentence in the abstract sums up the content of the paper: 'as well as providing a historical narrative and current and future perspectives, the coretrustseal experience offers lessons for those involved in developing standards and best practices or seeking to develop cooperative and community-driven efforts bridging data curation activities across academic disciplines, governmental and private sectors'. in order to attain coretrustseal tdr certification and become a trustworthy digital repository (tdr), the repository has to fulfil 16 requirements and the coretrustseal foundation maintains these requirements and the audit procedures. the certification draws on preservation standards and models as found in open archival information systems and in the catalogues of iso and din standards. the authors emphasize that the coretrustseal is founded on and developed in a spirit of openness and community. the paper's sharing of the experience follows that spirit. submissions of papers for the iassist quarterly are always very welcome. we welcome input from iassist conferences or other conferences and workshops, from local presentations or papers especially written for the iq. when you are preparing such a presentation, give a thought to turning your one-time presentation into a lasting contribution. doing that after the event also gives you the opportunity of improving your work after feedback. we encourage you to login or create an author login to https://www.iassistquarterly.com (our open journal system application). we permit authors 'deep links' into the iq as well as deposition of the paper in your local repository. chairing a conference session with the purpose of aggregating and integrating papers for a special issue iq is also much appreciated as the information reaches many more people than the limited number of session participants and will be readily available on the iassist quarterly website at https://www.iassistquarterly.com. authors are very welcome to take a look at the instructions and layout: https://www.iassistquarterly.com/index.php/iassist/about/submissions authors can also contact me directly via e-mail: kbr@sam.sdu.dk. should you be interested in compiling a special issue for the iq as guest editor(s) i will also be delighted to hear from you. karsten boye rasmussen may 2019 https://doi.org/10.29173/iq953 https://www.iassistquarterly.com/index.php/iassist/about/submissions vol29-4.indd iassist quarterly winter 2005 by by brian j. grim & roger finke* documenting religion worldwide: decreasing the data deficit abstract religion’s prominence in national and international affairs makes the availability of empirical measures on religion a pressing concern for researchers, policy makers, and data archivists. unfortunately, good international religious data are scarce. this paper describes the expanded mission of the association of religion data archives (www.thearda.com) to archive and develop data on religion worldwide. the arda archives data on 238 different countries and territories including arda-coded measures for 195 countries from the us state department’s annual international religious freedom reports. this paper highlights three important indexes developed by the arda that contribute to a better understanding of religion and religious tensions around the globe. these data are freely accessible online and available for download despite the increasingly obvious role of religion in national development and global relations, international data on religion are few, scattered, and often difficult to access. a recent review concludes that the study of world politics runs the “risk of stagnation” unless data are updated (bremer, regan, & clark, 2003). data sets which are available, such as the minorities at risk [coding] project, have limited measures that are specifically focused on religion as a category distinct from ethnicity. other crossnational studies, such as the world values study (wvs) and the international social survey programme (issp) with multiple measures on religion, have few data that specifically relate to such important topics as religion and conflict. such a topic is not of trivial interest. a recent pew research center poll of the u.s. public (2005) found that three-quarters of the public consider that religion either has a great deal (40 percent) or a fair amount (35 percent) to do with most wars and conflicts in the world today. to address the data deficiency on worldwide religion, the association of religion data archives (www.thearda. com) is pursing an international religion data initiative to archive key religion indicators for 238 different countries and territories. the arda, housed at the pennsylvania state university, is unique among the major academic data archives in that it focuses on the specific topic of religion. this offers several advantages. first, it archives data collections that might be excluded from more general or regional archives. second, it assists users in finding quality data that are embedded in larger social scientific studies, such as the issp, the wvs, and the general social survey (gss). third, it provides more support information and metadata that are customized for users interested in the topic of religion. and, fourth, online software and web features are customized for the topic at hand. currently, the arda’s international online collection includes several different types of data on religion including the world christian encyclopedia’s data on adherents of major religious traditions (barrett, kurian, & johnson, 2001). religion variables from the issp survey and four waves of the world values survey will soon be highlighted and available for online analysis at the arda. the collection also has other important international measures, including: freedom house’s indexes on political rights, civil liberties, religious freedom, and freedom of the press; basic economic measures such as purchasing power parity (ppp), income inequality (gini index), and the heritage foundation’s index of economic freedom; and key united nations indexes such as the human development index (hdi) and the gender-related development index (gdi). providing these measures in one site and in one downloadable data set improves access to such data, increases the use of the data, and allows comparisons between countries and regions of the world. unique to the arda collection are coded measures on the religious situation in 195 different countries based on the u.s. state department’s international religious freedom reports (2003). (for details on the coding and inter-rater reliability statistics see grim & finke, 2006; also see grim, finke, harris, meyers, & van eerden, 2006.) these annual reports to congress were mandated by the 1998 international religious freedom act and are one of the most comprehensive global treatments of social and political factors related to religion. though the report’s title rightly indicates a focus on religious freedom, the reports detail many issues related to religion including situations where religious actors are either perpetrators or victims of conflict and/or violence. each u.s. embassy 12 iassist quarterly winter 2005 follows a common set of guidelines and training is given to embassy and state department personnel who investigate the situation and prepare the reports (see u.s. state department 2001-2005). the reports cover the following standard reporting fields for each country: religious demography, legal/policy issues, restrictions of religious freedom, abuses of religious freedom, persecution by terrorist organizations, forced conversions, improvements in respect for religious freedom, and the us government’s actions. the u.s. state department has been compiling such annual reports since 1997. in 2001, they took on the reporting format shown in figure 1. one of the measures coded by grim and finke (2006) was whether tensions related to religion are present in the country (see table 1). it may come as no surprise that every country in south asia3 has tensions related to religion. it is sobering to see that, in countries with populations of at least two million, the presence of religious tension is the norm (84 percent). for those in the west, these religious tensions may be less present, given that tensions based on religion are reported in only 64 percent of the countries, much below the world norm. but, as the pew poll cited above reveals, these tensions are nonetheless felt. knowing that tensions are present is useful information, but it is even more useful to have measures that can help understand the government and social factors related to these tensions. the arda provides indexes for three important dimensions of religion in society that are related to these tensions: government regulation of religion index (gri), government favoritism of religion index (gfi), and social regulation of religion index (sri). each of these measures is significantly correlated with the presence of religious tensions.4 the first index (gri) is based on a set of questions related to the government regulation of religion. this is the most visible form of regulation for outsiders and the one that receives the most attention from theory and research. grim & finke (2006, p. 7) define government regulation as the restrictions placed on the practice, profession, or selection of religion by the official laws, policies, or administrative actions of the state. the index is built upon measures that are not just limited to the formal laws of the state. although the vast majority of countries promise religious freedom in their constitution, they often support administrative sanctions or open hostilities toward specific groups. moreover, minority groups can face battles for zoning approvals, tax-exemption status, and public corporation status. thus, restrictions against religions can come in the form of blatant laws against their existence or more subtle policy restrictions that limit their operations. actions of the state, however, are also supportive of religion. indeed, many countries openly favor select religions. grim & finke’s figure 1: international religious freedom report format (for each country) introductory overview [untitled section] 1. religious demography 2. status of religious freedom a. legal/policy framework b. restrictions on religious freedom c. abuses of religious freedom † d. forced religious conversion * e. improvements in respect for religious freedom ‡ 3. societal attitudes 4. u.s. government policy † section is absent for countries with no reported abuses. * beginning in 2004, a section on “persecution by terrorist organizations” was added. ‡ section is only present when improvements since the last report have been made. table 1: tensions related to religion tensions related to religion countries with populations greater than 2 million; n = africa 83% 35 east asia & pacific 80% 20 europe & eurasia 88% 42 near east & north africa 94% 17 south asia 100% 7 western hemisphere 68% 22 all countries > 2 million pop. 84% 143 iassist quarterly winter 2005 13 second index (gfi) is a measure of government favoritism of religion, which they define as subsidies, privileges, support, or favorable sanctions provided by the state to a select religion or a small group of religions (2006, p. 8). this favoritism can come in many forms. like government regulation, subsidies can be constitutional guarantees or they can result from the more capricious actions of administrative offices. the most obvious are specific constitutional privileges and the financial subsidies directly supporting religious institutions. less obvious are the supports of state institutions and administrations for such things as the teaching of religion in state supported schools and subsidy of service institutions run by religious groups. religious regulation is more than just the laws, policies, and administrative actions of a government. the third index (sri) provides a measure of the social regulation of religion that moves beyond the realm of the state. grim & finke (2006, p. 8) define social regulation as the restrictions placed on the practice, profession, or selection of religion by other religious groups, associations, or the culture at large. this form of regulation might be tolerated or even encouraged by the state, but is not formally sanctioned or implemented by government action. social regulation can be extremely subtle, arising through the pervasive norms and culture of the larger society, or it can include blatant acts of persecution by militia groups. often, though not always, this form of regulation is a product of religion. religion itself can regulate other religions. like other groups, religions seek to gain advantage by forming cartels and alliances that can regulate the culture and give the group a competitive advantage. regulation and favoritism of religion within countries is widespread. of the 143 countries with populations of two million or more, 76 percent have some form of government regulation of religion, 89 percent have some sort of government favoritism of religion, and 83 percent have some level of social regulation of religion. interestingly, the countries of the western hemisphere are more similar to the mean of all other countries with populations of two million or more in terms of government favoritism (4.7 compared to 4.8), but less similar when considering government regulation (1.4 versus 3.5) (see table 2). it is also interesting to note that the mean for social regulation of religion is highest in south asia (8.4) and the near east and north africa (8.3), both regions of the world experiencing high levels of conflict today. our intent here is not to provide an explanatory model of religion and conflict; rather, we have briefly highlighted examples of the types of data currently available at the arda on worldwide religious tensions, government regulation of religion, government funding of religion, and social regulation of religion. these are data researchers can use to better study religious tension and conflict in the world. these new measures are just one part of the arda’s new initiative to generate, assemble, and disseminate international data on religion. one of the chief aims is to make such data available in ways that reach multiple audiences. a user-friendly website for each country helps to accomplish this aim (see figure 2). the religious regulation and favoritism measures, in addition to other information, are clearly visible on the country summary page. intriguing statistics are also found, such data showing that saudi arabia is only estimated to be 92.7 percent muslim, due to the large number of foreign workers resident in the country who come from many different faiths, including at one time, the lead author of table 2: government regulation and favoritism, and social regulation of religion extent to which countries regulate and/or favor certain religions mean level* of: countries with populations greater than 2 million; n = government regulation of religion (gri) government favoritism of religion (gfi) social regulation of religion (sri) africa 3.0 3.2 3.4 35 east asia & pacific 4.4 3.9 3.4 20 europe & eurasia 3.1 5.3 3.7 42 near east & north africa 6.2 7.5 8.3 17 south asia 6.3 6.4 8.4 7 western hemisphere 1.4 4.7 2.0 22 all countries > 2 million pop. 3.5 4.8 4.1 143 * range 0-10, with 10 being the highest level of regulation or favoritism 14 iassist quarterly winter 2005 this paper. tabbed pages for each country take the visitor to additional information on adherents, religious freedom, and socio-economic conditions in each country. additional tabs will be added, including one planned for public opinion surveys related to religion done in the country, e.g., highlighting key variables from the world values survey and the international social survey programme (issp). the web site also provides easy links to many downloadable data sets and their sources. in the years ahead, the arda will continue to expand its collection of data specifically related to religion around the world and make these data freely available to the general public. additional online software tools for creating reports, graphs, and maps will also be developed as well as other online resources including press releases, figure 2: web shot of a country page summary reports, tutorials, information software tools, and user support. finally, the arda is producing a series of publications addressing the technical issues of creating indexes from the data (e.g., grim & finke, 2006) and initiating debate on the key substantive issues (e.g., religion’s relationship to peace, war, health, and democracy). international and interdisciplinary in scope, the international initiative’s new data, press releases, online resources, and publications also aim to contribute to the agenda for future research and provide solid resources for informed public discourse on international religion. refrences barrett, d.b. kurian, g.t., & johnson, t.m. (2001). world christian encyclopedia: a comparative survey of churches and religions in the modern world, 2nd ed. iassist quarterly winter 2005 15 new york: oxford university press. bremer, s., regan, p.m., & clark, d.h. (2003). building a science of world politics: emerging methodologies and the study of conflict. the journal of conflict resolution 47: 3–12. grim, b.j., & finke, r.. (2006). international religion indexes: government regulation, government favoritism, and social regulation of religion. interdisciplinary journal of research on religion 2, article 1. internet www. religjournal.com. grim, b.j., finke, r., harris, j., van eerden, j., & meyers, c. (2006). measuring international socio-religious values and conflict by coding u.s. state department reports. paper presented at the american association of public opinion researchers (aapor) annual conference, montreal, may 18-21, 2006. pew research center. (2005). views of muslim-americans hold steady after london bombings. internet http:// pewforum.org/docs/index.php?docid=89. u.s. state department. 2003. 2003 international religious freedom report. internet http://www.state.gov/g/drl/rls/irf/. * paper presented at the iassist 2006 conference in ann arbor in the session “compare and contrast: using cross-national data”. brian j. grim, the pew forum on religion and public life & association of religion data archives. roger finke, pennsylvania state university & association of religion data archives. direct correspondence to brian j. grim, bgrim@pewforum.org. supported by a grant from the john templeton foundation. the opinions expressed in this article are those of the authors and do not necessarily reflect the views of the foundation. the data reported in this article were downloaded from the association of religion data archives, www.thearda.com, and were collected by roger finke and brian j. grim. footnotes 1 major us university-based archives include the interuniversity consortium for political and social research (icpsr) at the university of michigan, the roper center for public opinion research at the university of connecticut, the howard w. odum institute for research in social science at the university of north carolina, the henry a. murray research archive and the harvardmit data center (hmdc), both members of the institute for quantitative social science at harvard university. other us university-based archives include the cultural policy and the arts national data archive (cpanda) at princeton university, uc data archive and technical assistance (uc data) at uc-berkeley, the center for international earth science information network (ciesin) at columbia university, the social sciences data archive (ssda) at ucla, the data and program library services (dpls) at the university of wisconsin-madison, and the national data archive on child abuse and neglect (ndacan) cornell university. major european archives include the central archive for empirical social research at the university cologne (which archives major surveys such as the issp and evs). the council of european social science data archives (cessda) provides a map of and links to the 21 major european archives, as well as links to five archives outside of north america and europe (see map and links at http://www.nsd.uib.no/cessda/europe. html). 2 the one country missing from this data set is the united states, due to the state department’s mission being limited to reporting on countries outside of the usa. 3 afghanistan, bangladesh, bhutan, india, nepal, pakistan, and sri lanka 4 the pearson correlations with religious tensions are: gri .279 (p = .001, two-tailed), gfi .188 (p = .025, two-tailed), and sri .463 (p < .001, two-tailed). all three of these indexes and the questions used to construct them, as well as the variable on religious tensions, are downloadable. they are also displayed online at an arda web page developed for each country. a more detailed description of the indexes and how they were created can be found in grim & finke (2006). iassist quarterly iassist quarterly 2013 7 abstract in recent years, semantic web technologies have matured and have made their way into various domains – for example, bioinformatics and egovernment -where they are used in different applications to provide value-added services for users. in this paper, we present an overview of several representative applications that use semantic web technologies and show the potential of these technologies in the linked data world. we then present existing semantic web applications specifically for the social sciences and highlight their impact on this domain. integrating the semantic web vision into the social sciences results in clear benefits, which we identify and discuss. keywords semantic web, linked data, semantic web applications, social sciences . introduction the corporate landscape is moving to the semantic web in a big way. major companies like adobe, oracle, ibm, hp, software ag, ge, northrop grumman, altova, and microsoft offer (or plan to offer) semantic web tools or systems. others like novartis, pfizer, and telefónica are using the semantic web (or are considering using it) as part of their own operations. active participants in w3c semantic web related groups include hp, agfa, sri international, fair isaac corp., oracle, boeing, ibm, chevron, siemens, nokia, pfizer, and eli lilly. in addition to the corporate sector, we see major communities such as digital libraries, defense sector, egovernment, the energy sector, financial services, health care, the oil and gas industry, and the life sciences adopting the technologies. for social science researchers, the semantic web and linked data hold great promise as gregory and vardigan (2010) illustrate in detail. the adoption of semantic technologies has the potential to make the discovery of data and metadata in the web more efficient, to enhance the reuse of social science metadata, and to decrease the technical barriers to employing data in research. through linked data for the social sciences, users can easily discover the existence of data and can determine how the metadata are structured and whether the data are suitable for their interest. therefore, it is necessary that the data is well-documented and that quality and provenance are explicit. this will enable the identification of complex relationships such as whether datasets are comparable or whether there are relationships to other versions of the same dataset. since data collections published as linked data can easily be linked with each other, the integration and merging of heterogeneous datasets is also facilitated. in this paper, we provide an overview of tools and applications that apply semantic web technologies and linked data. we also present social science projects and applications that make use of semantic technologies, and we discuss the benefits of applying these technologies to the domain of the social sciences. semantic web applications the number of applications utilizing semantic web technologies and linked data is increasing and illustrates the potential of these technologies for semantic web applications for the social sciences by thomas bosch1 and benjamin zapilko2 the number of applications utilizing semantic web technologies and linked data is increasing and illustrates the potential of these technologies for various domains 8 iassist quarterly 2014/2015 iassist quarterly various domains and heterogeneous communities. in this section of the paper we provide illustrative examples. hcls demo the w3c health care and life sciences interest group (hcls) have developed demonstrations of the usage of semantic web technologies (w3c 2008a), which herman (2011) has described. the demonstrations are designed to show the hcls community how the semantic web can be used and to show the semantic web community how this semantic web technology can be useful in this specific application area. the core of the hcls demo is the access and the integration of public datasets via the semantic web. one of the hcls demonstrations is the allen brain atlas, which is currently available only through an html interface. mouse brains are cut in slices and stained for the presence of gene expression: 20,000 genes, 400,000 images at high resolution. the allen brain atlas is ‘mashed up’ with google maps. google maps allows the user to upload his or her own ‘maps’, i.e., the uris of bitmap images. the user then gets the navigation of large bitmaps for free. the goal of the demonstration is to find the right images in the atlas data to really provide navigable data. how is it done? what happens is that the query on the atlas data is based on scraping the html structures, extracting the uri, and using sparql to combine these into more complex queries that result in the right image uris being fed into the google service. via rdf, one gets a standard query sparql interface for free, providing much power to the end user. if the original authors of the brain atlas had stored the data in a public mysql database, one could have achieved the same user interface, but the fact is that they did not. rdf and sparql allow one to produce the relevant images easily without interfering with the original data in any way, using standard, off-the-shelf tools. the demo shows that sometimes rdf and sparql help in a very simple way -semantic web applications are not necessarily very complex. nasa grove and schain (2008) describe how to find the right experts at nasa using semantic web technologies. the nasa tool is an expertise locater for nearly 70,000 nasa civil servants. the tool uses rdf integration techniques across geographically distributed databases, data sources, and web services. the authors use internal ontologies/vocabularies to describe the knowledge areas, and a combination of the rdf data and these ontologies to search through the (integrated) databases for specific knowledge expertise. the dump results from a faceted browser developed by the company to view the data on experts. semantic mediawiki semantic mediawiki (2012) is a module of the mediawiki (2012) software and it extends wikis with ideas from the semantic web discipline. semantic mediawiki enables the user to make facts available for machines, thus making it easier for humans to search and reuse information. articles, such as an article about the iassist conference in 2014, can be annotated semantically using the rdf format. in this way, rdf triples can be built and stored in the wiki code directly. relations between subjects and objects can be defined semantically. the iassist conference in 2014 would be the subject of the rdf triples stated on the article. it could be specified that the iassist conference has participants, presentations, publications, key note speeches (e.g., relation ‘haskeynotespeech’). authors can also write articles for each relation and for each object to explain the semantics. in the context of this example, authors can write articles about the meaning of the relations to key note speeches and also about each key note speech. on the site about a specific key note speech, the key note speaker could write an article about the content of the presentation. humans as well as machines can query all the information included in the wiki sites. people can query information about other articles when they edit new articles. then the query results appear directly in the wiki site. these articles are updated automatically when the dependent sites are updated. as a consequence, overview articles are always up-to-date and consistent with the detail sites. visitors of semantic mediawikis can download semantically enriched information directly in rdf. a large number of general-purpose rdf tools and specialized external programs can now reuse and process the rdf data in an easy and standardized way. the data and metadata exchange in rdf enables the combination of information from different sources like wikis. semantic mediawiki is free software under the gpl license. more than 150 public wikis use the semantic mediawiki extension. in particular, semantic annotations in wikis (e.g., lexwiki, concepthub-wiki) are adopted in medical and biology sciences to create biomedical terminologies and ontologies collaboratively. web repositories and server systems in this section we present two popular systems for accessing, maintaining, and publishing data on the web that utilize semantic web technologies. fedora repository project the fedora (flexible extensible digital object repository architecture) repository project (2013) is an open source software system originally developed by researchers at cornell university. the underlying architecture enables the storage, management, and access of digital content. fedora allows for expressing digital objects and relationships among them by assigning so-called ‘behaviors’ (i.e., services) to them. in addition to a core repository service with well-defined apis, fedora includes services for searching, oai-pmh, messaging, and administrative clients, to name a few. regarding semantic web technologies fedora supports interaction with rdf data, since the repository software can be connected with rdf triple stores. data stored in a triple store can be accessed and used by every service of fedora. there are various scenarios and domains dealing with digital content where fedora is applied. it can be used for “digital collections, e-research, digital libraries, archives, digital preservation, institutional repositories, open access publishing, document management, digital asset management, and more” (fedora, 2013). virtuoso virtuoso (2013) is a data server that can be applied to various ways of storing, maintaining, and accessing different kinds of data. the core of virtuoso is an object-relational sql database. for serving dynamic web pages, virtuoso provides a flexible built-in web server, which can process pages written in virtuoso’s own web language (vsp) and other standard languages like php or asp. virtuoso also provides functionalities for maintaining and managing the published web pages like versioning, automatic metadata extraction, and full text searching. it is also available as an open source edition at http://www.openlinksw.com/wiki/main/. iassist quarterly 2014/2015 9 iassist quarterly with respect to semantic technologies virtuoso currently enables the storing and querying of rdf data in its database. since this is currently a sql database, a translation of sparql queries into sql is supported in order to query the rdf data and to provide rdf as output format as well. there are plans to extend the access and storage capabilities of connected databases, which would also enable particular technologies like inferencing on rdf data. thesaurus management tools thesauri and classifications are commonly used instruments for describing and annotating metadata about various kinds of documents. they have a long tradition in libraries and archives. in the following paragraphs we describe two thesaurus management tools that use semantic web technologies and datasets. poolparty thesaurus server poolparty thesaurus server (ppt) (2013) is a software platform that enables the management and maintenance of complex knowledge models such as taxonomies, thesauri, controlled vocabularies, and similar data. the metadata are fully organized and modeled using w3c´s semantic web standards rdf and skos, since skos focuses particularly on knowledge organization systems. managing this data inside ppt is enhanced by text mining functionalities and linked data mapping technologies. in addition to processing rdf data, the apis provided by ppt are also based on semantic technologies, the sparql standard of w3c. this allows also an integration of the maintained knowledge models in other systems, e.g., cms, erp-systems, or wikis. by using complex semantic web based approaches like text corpus analysis, entity extraction, linked data enrichment, and skos thesaurus management, it is possible to build, maintain, and publish large and complex knowledge models based on rdf data. vocbench vocbench (2013) is a vocabulary editing and workflow tool developed by the food and agriculture organization (fao). the web-based application enables the transformation of multilingual knowledge organization systems like thesauri, authority lists, and glossaries into skos/rdf concept schemes. thus, traditionally maintained thesauri can easily be used in semantic web applications. besides the transformation, vocbench also allows for managing, maintaining, and editing the data. this includes collaborative editing as well as validation and quality assurance tasks. vocbench is an open source project and based on protegé. currently, vocbench is used to manage several datasets held at fao like the agrovoc thesaurus (2013), the biotechnology glossary (2013), and bibliographic metadata used in fao. for future releases fao plans to include a native interface for skos and skos-xl and configurable support for hosting of different triple store technologies. also, they plan to support generic owl ontologies. vivo project the vivo project (2013) is an open source semantic web application, which was originally developed and implemented at cornell university. the application maintains profiles of researchers and organizations. these profiles can be populated with additional information like activities or interests. through extensive search and browse capabilities, it is possible to discover information across institutions and disciplines. although the vivo software is installed locally, the different installations worldwide are connected with each other in a network, which also enables integrated searching and browsing across the information of all connected installations. the information that can be discovered can be used in different contexts, e.g., in visualizations or in applications like vivo searchlight, which allows to search for vivo profiles based on textual information from any web page. the open source project is available at http://vivo.sourceforge.net. the structured data in vivo is represented in rdf using the vivo ontology (mitchell et al., 2011). this ontology focuses on describing researchers and networks of researchers across organizations and disciplines. it also covers researchers’ teaching activities, their expertise, their research, and which service activities they provide. the ontology has been developed inside the vivo project. semantic web applications for the social sciences figure 1. sofiswiki – search for research institutions 10 iassist quarterly 2014/2015 iassist quarterly so far, there are just a few semantic web applications for the social sciences. we present them in this section in detail. for each semantic web application for the social sciences, we show individual benefits for users in the social sciences community regarding semantic web supported functionalities. a preview of these social science applications: • sofiswiki provides research institutions the possibility to publish information about their research institution, their research activities, and their research projects • the microdata information system (missy) is an information system to document german and european studies at the variable and the study level. • nesstar enables publishing of a huge amount of statistical data and metadata using semantic web technologies. • colectica is a fast way to design, document, and publish survey research using open data standards. sofiswiki are you conducting a social science research project? are you writing a thesis or doing advanced work on a social science topic? if so, then you may want to create your project in sofiswiki and make your research work transparent for the scientific community. figure 2. sofiswiki – research institution and associated research projects gesis has developed sofiswiki (2012), the first semantic mediawiki in the social science community. sofiswiki informs researchers about research activities and projects as well as research institutions in the german-speaking social sciences. sofiswiki contains all entries from the last ten years of sofis (2012), a central database hosted by gesis that delivers information about over 50,500 research projects. using sofiswiki, research institutions like universities can enter and represent information about their research projects and their institutions in convenient forms. other researchers can get an overview of the wiki content and can also search for and find this information. the collected project information is also available using the social science portal sowiport (2012). as part of future work, sofiswiki could be extended to support the documentation of research activities, projects, and institutions all over the world. figure 1 shows the graphical user interface of sofiswiki to enable a search for research institutions by institution name, country and location, content orientation, and institution type. figure 2 shows the result of the above query for the department ‘monitoring society and social change’ of the research institution iassist quarterly 2014/2015 11 iassist quarterly ‘gesis’. department metadata such as address, homepage, and contact person as well as the associated research projects are delivered. benefits for the social sciences community research institutions like universities can manage and publish information about their research institutions and their research projects using this tool, and researchers can get an overview of research projects, institutions, and activities. internally, the metadata are represented using rdf to ease the integration of the metadata items. rdf is also used to integrate sofiswiki with diverse other portals like sowiport and sofis. the microdata information system (missy) general description of missy the microdata information system (missy) (gesis 2012) maintains the largest household survey in europe – the german microcensus, which provides statistics about the general population in germany, including the employment market (occupation, professional education, income, legal insurance). missy consists of approximately 500 variables and questions and captures data for 25 years, since 1973. figure 3 shows the graphical user interface of missy. in the detailed view of the variable, the user gets details about the variable “gender” including associated question, values, value labels, and absolute and relative frequencies. missy has two parts: • missy web, the end-user front-end • missy editor for the metadata documentation, which is the back-end several use cases are covered by missy: • thematic classification: variables by thematic classification and year • variables by year • generated variables by year • details of variables with statistics • variable-time matrix: variables by thematic classification and year (selectable) • questionnaire catalogue in the third generation of missy, additional surveys such as eu-silc (european union statistics on income and living conditions), eu-lfs (european union labour force survey), and evs (european values study) will be integrated. the missy editor will be implemented as a web application. in future, it should also be possible to browse variables by survey and by country. missy and linked data the missy-specific data model is based on the ddi-rdf discovery vocabulary – disco (bosch et al. 2012) for the following reasons: • disco contains the most salient components of both ddicodebook and ddi-lifecycle for data discovery (as a figure 3. missy – graphical user interface 12 iassist quarterly 2014/2015 iassist quarterly consequence, not all of the over 830 xml elements of ddi 3.1 are covered) • disco serves as a first step in developing a model-driven ddi specification to document microdata within the social, behavioral, and economic sciences • disco will be officially published and publicly available • more than 20 experts from the statistics and the linked data community of eight different countries have contributed to the development of disco in three workshops and additional working groups as not every requirement within the missy context is covered by the ddi ontology, an individual data model is defined on top of the common abstract data model. other software projects intended to document studies at the study and variable level using ddi-l can also be based on disco and can reuse existing code, which is made available on a github repository (missy 3 2012). to show how this works, an example is provided below. the skos (simple knowledge organization system) (w3c 2009), whose purpose is to define hierarchies of concepts, is reused in the disco ontology to a large extent (see figure 4). for instance, codes, categories, ddi concepts, and study subjects are represented in disco as skos:concepts. the next figure visualizes skos:concepts within missy. the skos:concept used in the ddi discovery vocabulary is extended by the concept defined in missy. using the property ‘skos:preflabel’, category labels can be stored in an rdf format. the datatype of the category labels is specified as a normal string. one requirement in missy is to store category labels in different languages such as english, german, and french. thus, we have defined the type ‘multilingual’ in missy. in order to represent category labels, the property ‘skos:preflabel’ of the datatype ‘multilingual’ is used and therefore the initial common abstract data model is extended. in a well-defined software architecture, the application itself does not need to know how the data are stored. the application just needs to know the api, i.e., methods that are provided to access and store objects. these methods may be abstracted away from the actual implementation. an actual implementation or strategy can just be a matter of configuration. a strategy is an implementation of the actual type of persistence or physical storage, e.g., ddi-l-xml, ddi-rdf, xml-db, or relational-db. the persistence api defines the persistence functionality for model components regardless of the actual type of physical persistence. several components implement the persistence functionality defined in the persistence api with respect to the usage of relational dbs, ddi-xml, and ddi-rdf. one concrete implementation of the persistence api is ddi-rdf, the rdf representation of the developed ddi ontology. missy will offer several export formats – one of them will be ddi-rdf. we will implement two additional concepts providing rdf data. first, missy websites will be annotated semantically using rdfa (w3c 2012), which are generic annotations in xhtml documents. thus, machines can crawl the missy websites in order to import exactly the information needed for further processing, since now machines ‘know’ the meanings of the provided information. rdfa metadata should and will be provided according to the ddi ontology disco and according to the schema.org vocabulary (schema. org 2012). launched in 2011, schema.org is an initiative from bing, google, and yahoo to provide a vocabulary (a collection of concepts and their properties) to be used by web masters to markup web content in ways recognized by major search providers. search engines will rely on this markup to improve the display of search results, making it easier for people to find the right pages they search for. the second way of exporting semantic information to the social science community is to build a sparql endpoint. rdf triples are stored in a triple store in parallel to the xml documents containing both the data and metadata of multiple studies offered by missy. benefits for the social sciences community other software projects documenting studies on the study and variable level using ddi-l can reuse existing github repository code. missy will provide multiple export formats (e.g., ddi-rdf). ddi data as well as metadata can be published in the linked open data cloud. missy websites will be annotated using rdfa according to ddi-rdf and schema.org. as a consequence, search engines will improve their query results. by writing sparql queries, ddi data and metadata can be accessed from sparql endpoints. nesstar nesstar (norwegian social science data services 2012) is a semantic web application for documenting both statistical data and metadata. nesstar can be seen as an extraordinarily easy to understand sw application, as it does not specify sophisticated ontologies, does not use advanced rdf features such as reification, and does not use logical inference. in 1998, the european union funded the research and development project called ‘networked social science tools and resources,’ abbreviated as nesstar. the eu project with the name ‘faster’ (faster 2012) followed the goals associated with the nesstar project. assini (2002) gives a rather detailed description of nesstar. the aim of nesstar is to make a huge quantity of statistical data and metadata accessible using semantic web technologies. before the implementation of nesstar, statistical data as well as metadata was typically only available in a human-readable and understandable form and not additionally in a machineunderstandable form that could be further processed by class  concept-­‐2 ddidiscovery::skos:concept -­‐   skos:definition    :rdf:langstring -­‐   skos:notation    :rdfs:literal -­‐   skos:preflabel    :rdf:langstring concept +   skos:preflabel    :rdf:langstring 0..* skos:narrower 0..*0..* skos:broader 0..* figure 4. missy – abstract and individual data model iassist quarterly 2014/2015 13 iassist quarterly computer programs. nesstar was intended to revolutionize the way people access statistical information, bringing the advantages of instant access to the world of statistical data dissemination. on the nesstar website (nesstar 2013), a list of nesstar catalogues (e.g., surveys, tables) is provided by the nesstar’s demo server. figure 5, for example, shows information such as the associated question, the values and the categories, summary statistics, interviewer instructions, and the total responses about the variable gender of a demo survey. conceptual model of nesstar the nesstar object model is defined in rdfs. about 15 classes represent the key domain-specific concepts within the statistical domain like studies, data files, variables, indicators, and tables. the conceptual model also includes relationships between domainspecific concepts: studies, for example, may contain cubes having one or more dimensions. additionally, 10 domain independent support classes are part of the object model of statistical data and metadata. the class server, for instance, represents the server where the metadata objects are hosted. it provides basic administrative functionality such as file transfer, server reboot, and server shutdown. starting from a server and by recursively traversing the objects’ relationships, applications can reach all the server’s objects. the domain independent support class catalog groups metadata objects. instances of catalogs can be browsed and you can get a list of all the metadata objects which are included in the catalog. there is also the possibility to search for particular objects. many research studies contain sensitive information that cannot be made available without restrictions. within nesstar, access control policies can be defined in order to follow a security model. to implement this, classes such as user, role (e.g., administrator, final user, data publisher), and agreements (e.g., ‘i agree to use this data only for noncommercial research purposes’) are specified. nesstar is based on lightweight, object-oriented web middleware named neoom nesstar object oriented middleware. neoom is based on web and semantic web standards like html, http, rdf, and rdfs. neoom is a set of guidelines on how to use web as well as semantic web technologies to build distributed object-oriented systems. the neoom guidelines are described extensively in assini (2001b) and very briefly in assini (2001a). rdfs does not provide a way to describe the behavior, the operations of statistical objects (e.g., queries, statistical operations, file transfers, tabulate, and frequency). how to specify the operations formally? in the neoom object model, specific methods (e.g., login) are defined as sub-classes of the method class. concrete method invocations are then instances of the method class. behavioral view on nesstar according to the nesstar conceptual model, data publishers make their statistical data and metadata available on the web as objects. these objects are represented by rdf resources according to the nesstar object model. each data publisher runs its own server, which is an instance of the class server. nesstar figure 5. nesstar – variable gender 14 iassist quarterly 2014/2015 iassist quarterly servers host the maintained objects. nesstar servers provide www resources such as html pages and images as well as statistical objects. nesstar is fully distributed and each server is totally independent and integrated. users have the possibility to access statistical objects remotely by simply typing objects’ urls. soap (w3c 2007) is used for remote object-oriented calls. similar to using search engines like google, users can search for remote statistical objects: they could for example type the search term ‘find all variables about political orientation’. in nesstar, there are different kinds of user access possibilities: nesstar explorer, figure 7. colectica – questions figure 6. colectica – variables iassist quarterly 2014/2015 15 iassist quarterly nesstar light explorer, nesstar publisher, and object browser. nesstar explorer is similar to a common web browser. users can enter objects’ urls and www resources are displayed as they would be displayed in a web browser. nesstar publisher is a tool for editing metadata, for validating, and for publishing. the object browser’s purpose is to test and to administer statistical objects. benefits for the social sciences community nesstar enables users to publish a huge amount of statistical data and metadata using semantic web technologies. now statistical data and metadata is not only available in a human-readable and understandable form but also in a machine-understandable form which can be further processed. colectica rdf services colectica is a fast way to design, document, and publish survey research using open data standards. the colectica platform provides features for statistical agencies, survey research groups, public opinion researchers, data archivists, and other data intensive operations. colectica can increase the expressiveness and longevity of the data collected through standards-based metadata documentation (colectica 2012c). ddi-l allows for reuse and harmonization of metadata items through the use of referencing. with the colectica 4.0 repository addin, the relationships between metadata items are indexed (colectica 2012b), which makes it possible to execute queries on these relationships. figure 6 displays the documentation of variables using colectica designer. you can specify variable names, variables labels, variable descriptions, the response unit, associated concepts and universes, and the representation of the variable (e.g., numeric, textual, or coded representation). figure 7 shows how to document questions using colectica designer. for questions you can specify question names, question text, the question scheme, and the response domain (e.g., numeric, textual, or coded). the colectica rdf model is created by hand based on the colectica ddi-l model. each description of a ddi metadata item is stored as a named graph. the rdf services architecture can be deployed with colectica repository. all ddi metadata items, which are stored and versioned in the repository, are also stored in rdf (colectica 2012a). several external vocabularies such as rdf, rdfs, simple dublin core (dc), the dcmi metadata terms (dcterms), owl, xsd, and foaf are reused (colectica 2012b). sparql (w3c 2008b) is a query language created for searching rdf data and is standardized by the w3c. it allows for searching based on the relationships and literal data stored in an rdf graph or store. sparql can be used to construct very precise questions about ddi metadata items referencing multiple metadata items. there are two deployment scenarios which can be distinguished in colectica: internal rdf stores and external rdf stores. using colectica repository’s internal rdf store, sparql 1.0 as well as the draft version 1.1 of sparql are supported. the sparql update functionality is disabled in order to maintain consistency with the versioned ddi metadata items in the repository. for colectica repository, it is also possible to replicate the rdf to external already existing rdf stores, which is the second deployment scenario (colectica 2012a). one can query ddi-l as rdf either using a web service from colectica repository or using a sparql endpoint on colectica web. in addition, each ddi-l metadata item, which is stored in the colectica repository, can be downloaded as an rdf dump (colectica 2012a). one example of such a sparql query in the statistical domain could be: which studies has dan smith the software developer of colectica authored since the beginning of 2010 (colectica 2012a)? prefix ddi: <urn:ddirdf:> prefix ddit: <urn:ddirdf:type:> prefix dc: <http://purl org/dc/elements/1.1/> select ?study where { ?study a ddit:studyunit; dc:date ?creation_date; dc:creator <http://dan.smith.name/who#dan>. filter (xsd:datetime(?creation_date) > “2010-01-01 00:00:00”^^xsd:datetime ) . } order by ?study another example of a sparql query would be: how many times has a variable been reused across multiple datasets (colectica 2012a)? prefix ddi: <urn:ddirdf:> prefix ddit: <urn:ddirdf:type:> prefix dc: <http://purl.org/dc/elements/1.1/> select ?variable count (?parent) as c where { ?variable a ddit:variable ; ?parent ddi:hasvariable ?variable . ?parent a ddit:dataset. } group by ?variable dan smith’s website (colectica 2012b) also offers several examples of ddi-rdf serializations. a further sparql example from the ddi-l us 2010 census sample file is also provided. as part of future work, predicates will be updated when the official and community adopted ddi discovery vocabulary is available (colectica 2012a). benefits for the social sciences community ddi-l data and metadata can be queried as rdf using the colectica repository web service or the colectica web sparql endpoint. ddi-l metadata items, stored in the colectica repository, can also be downloaded as rdf dumps. conclusion and future work we have presented several representative applications that apply semantic web technologies to a high degree. while semantic technologies and linked data have yet not been widely used in the social sciences, we have identified initial applications exclusively developed for this domain. the impact of semantic web and linked data are exposed in these applications. additional potentials and benefits for an adaption of semantic technologies for scientific purposes can easily be identified. we have shown individual benefits for users of the social sciences community regarding semantic web functionalities. references agrovoc 2013 agrovoc thesaurus, viewed 6 may 2015, < http:// aims.fao.org/vest-registry/vocabularies/agrovoc-multilingualagricultural-thesaurus > assini, p 2001a ‘objectifying the web the ‘light’ way: an rdf-based framework for the description of web objects’, proceedings of the international world wide web conference, hong kong, 01 mai 2001, tenth international world wide web conference. 16 iassist quarterly 2014/2015 iassist quarterly assini, p 2001b neoom: a web and object oriented middleware system, [online], available: http://www.nesstar.org/sdk/neoom.pdf [6 may 2015]. assini, p 2002, ‘a semantic web application for statistical data and metadata’, proceedings of the international world wide web conference, hawaii, 07 may 2002, 11th international world wide web conference. biotechnology glossary 2013 biotechnology glossary, viewed 6 may 2015, <http://www.fao.org/biotech/biotech-glossary/en/> bosch, t, cyganiak, r, wackerow j & zapilko b 2012 ‘leveraging the ddi model for linked statistical data in the social, behavioural, and economic sciences’, proceedings of the international conference on dublin core and metadata applications, kuching, 03 september 2012, international conference on dublin core and metadata applications, pp46 55. colectica 2012a accessing ddi 3 as linked data: colectica rdf services, viewed 6 may 2015, <http://www.iassistdata.org/conferences/2012/presentation/3326 >. colectica 2012b ddi 3 meets rdf and sparql with colectica repository, viewed 6 may 2015, <http://dan.smith.name/2011/10/ ddi-3-meets-rdf-and-sparql-with-colectica-repository/>. colectica 2012c colectica website, viewed 6 may 2015, <http://www. colectica.com/>. faster 2012 faster, viewed 6 may 2015, <http://fasterproject.eu/>. fedora 2013 fedora repository project, viewed 6 may 2015, <http:// fedora-commons.org/>. gesis 2012 the microdata information system (missy), viewed 6 may 2015, < http://www.gesis.org/missy/>. gregory, a & vardigan, m 2010 the web of linked data: realizing the potential for the social sciences, viewed 6 may 2015, < http://odaf. org/papers/201010_gregory_arofan_186.pdf> grove, m & schain, a 2008 case study: pops — nasa’s expertise location service powered by semantic web technologies, viewed 6 may 2015, <http://www.w3.org/2001/sw/sweo/public/usecases/ nasa/>. herman, i 2011 semantic web adoption and applications, viewed 6 may 2015, <http://www.w3.org/people/ivan/corepresentations/ applications/>. mediawiki 2012 mediawiki.org, viewed 6 max 2015, <http://www. mediawiki.org/wiki/mediawiki>. mitchel, s, chen, s, ahmed, m, lowe, b, marks, p, rejack, n, corsonrikert, j, he, b, ding, y 2011 ‘the vivo ontology: enabling networking of scientists’ acm webscience conference, koblenz missy 3 2012 missy 3 project, viewed 6 may 2015, <https://github. com/missy-project>. nesstar 2013, welcome to nesstar’s demo server, viewed 6 may 2015, < http://nesstar-demo.nsd.uib.no/webview/>. norwegian social science data services 2012 nesstar, viewed 6 may 2015, <http://www.nesstar.com/>. ppt 2013 poolparty thesaurus server, viewed 6 may 2015, < http:// www.poolparty.biz/> schema.org 2012 schema.org, viewed 6 may 2015, <http://schema. org/>. semantic mediawiki 2012 semantic mediawiki, viewed 6 may 2015, <http://semantic-mediawiki.org>. sofis 2012 sofis social science research information system, viewed 6 may 2015, <http://www.gesis.org/en/services/research/ sofis-social-science-research-information-system/>. sofiswiki 2012 sofiswiki, viewed 6 may 2015, <http://www.gesis.org/ sofiswiki/hauptseite>. sowiport 2012 sowiport, viewed 6 may 2015, < http://sowiport.gesis. org/ >. virtuoso 2013 virtuoso universal server, viewed 6 may 2015, <http:// virtuoso.openlinksw.com/> vivo 2013 vivo project, viewed 6 may 2015, <http://www.vivoweb. org/> vocbench 2013 vocbench, viewed 6 may 2015, < http://aims.fao.org/ vest-registry/tools/vocbench-2> w3c 2007 soap version 1.2 part 0: primer (second edition) w3c recommendation 27 april 2007, viewed 6 may 2015, <http://www. w3.org/tr/2007/rec-soap12-part0-20070427/>. w3c 2008a hcls/banff2007demo, viewed 6 may 2015, <http://www. w3.org/wiki/hcls/banff2007demo>. w3c 2008b, sparql query language for rdf, viewed 6 may 2015, <http://www.w3.org/tr/2008/rec-rdf-sparql-query-20080115/>. w3c 2009 skos simple knowledge organization system namespace document html variant 18 august 2009 recommendation edition, viewed 6 may 2015, <http://www. w3.org/2009/08/skos-reference/skos.html>. w3c 2012 rdfa 1.1 primer rich structured data markup for web documents w3c working group note 07 june 2012, viewed 6 may 2015, <http://www.w3.org/tr/2012/ note-rdfa-primer-20120607/>. notes 1.thomas bosch gesis leibniz institute for the social sciences, mannheim, germany e-mail: thomas.bosch@gesis.org 2. benjamin zapilko gesis leibniz institute for the social sciences, köln, germany e-mail: benjamin.zapilko@gesis.org 1/18 kubas, alicia & mcburney, jenny (2019) frustrations and roadblocks in data reference librarianship, iassist quarterly 43 (1), pp. 1-18. doi: https://doi.org/10.29173/iq939 frustrations and roadblocks in data reference librarianship alicia kubas1 jenny mcburney2 abstract as data skills are incorporated into academic curriculum and data becomes more widely available and used in everyday life, many librarians find themselves serving as 'accidental' data librarians in their subject areas, a phenomenon widely discussed at conferences (including iassist), in publications, and in the field at large. due to this evolving landscape and growing data need, it is increasingly important for librarians to be familiar with data resources and able to answer secondary data reference questions. to learn more about this area of librarianship, this study uses survey responses from librarians who answer data questions to explore the challenges and frustrations that arise from data reference questions and interactions. our key findings reveal that frustrations are ever present in data reference regardless of how much experience a librarian has, and many frustrations arise due to factors such as patron expectations, subject-specific and data-related jargon, and data formats and accessibility, some of which are beyond the data librarian’s control. keywords library, data reference, library reference services, data services, data literacy, data librarian introduction questions about locating hard-to-find data, including international data, historical and time series data, and microdata, are increasingly important and frequent at many types of educational institutions and libraries. data reference is a growing area of expertise that is becoming a vital baseline skill for library professionals across subject areas, not just for data librarians, as more teaching faculty adopt data as a regular component of assignments and as data librarians assist patrons of all types in locating secondary data for their research pursuits. in order to learn more about the landscape of secondary data reference and how librarians can better support researchers and users in this area, data-focused librarians at the university of minnesota developed a survey targeting library staff across different library settings and geographic locations who work with data-related questions. the survey focused on the types of data questions received and from whom, how librarians tackle these questions, frustrations and roadblocks they experience, and opportunities for increasing expertise in this area of librarianship. https://doi.org/10.29173/iq939 2/18 kubas, alicia & mcburney, jenny (2019) frustrations and roadblocks in data reference librarianship, iassist quarterly 43 (1), pp. 1-18. doi: https://doi.org/10.29173/iq939 while many themes and issues were identified from the survey results, one important theme that emerged was the frustrations librarians and library staff face as they attempt to answer data reference questions and become more skilled in finding and accessing datasets. this paper focuses on finding out and discussing which frustrations were influenced by other variables, namely the volume of data questions received, a librarian’s number of years in the field, and the type of data questions received in terms of time (time series, historical, current, etc.). our key finding is that frustration is not impacted by years of experience, which we did not expect. rather, data persists as a difficult source type for librarians to support, regardless of years of experience, due to issues such as format challenges, jargon that varies across disciplines, difficulty in locating certain types of data, and management of patron expectations. in addition, a number of the frustrations encountered by data librarians are out of their control, so experience does not eliminate them entirely, even though more experienced librarians may have determined better strategies for dealing with these roadblocks. methods during the fall of 2017, we developed an online survey targeting librarians who answer any secondary data questions in their work. the survey was created in qualtrics and received irb approval through the university of minnesota. the survey was sent to various national, international, and local listservs and mailing lists including: • iassist (international association for social science information services & technology) • govdoc (library professionals working with government information) • rusa (american librarian association (ala)-reference & user services association) • rss (ala-reference services section) • brass (ala-business reference and services section) • nmrt (ala-new members round table) • uls (association of college & research libraries university libraries section) • minnesota library association (mla) • mla public library division • mla academic and research libraries division • lib-minndocs (federal depository library coordinators in mn, mi, and sd) • research and learning division at university of minnesota-twin cities additionally, some librarians who received the survey via an email list further shared it more broadly on their personal twitter accounts. https://doi.org/10.29173/iq939 3/18 kubas, alicia & mcburney, jenny (2019) frustrations and roadblocks in data reference librarianship, iassist quarterly 43 (1), pp. 1-18. doi: https://doi.org/10.29173/iq939 the survey was open from november 27 to december 22, 2017, and a reminder email was sent to each list halfway through the open period. the quantitative data was exported from qualtrics into a csv file. data clean-up was done in google sheets and the cleaned dataset was analyzed in spss. we used dichotomous variables and therefore utilized chi-squared tests in our analysis. we considered our analysis to be statistically significant if the p-value was less than 0.05. questions that had a 'select all that apply' option were broken down into individual responses, and dichotomous variables were created from the absence or presence of the response. for example, one question had nine possible answers we considered for the analysis and therefore we created nine dichotomous variables out of that question. three hundred sixty respondents began the survey. only responses where respondents continued past the demographics section and also indicated that they answered at least one data question on average per month were included in the data analysis, which amounted to 278 total responses. in the survey, we defined data as follows: in this context, 'data' refers to existing datasets, statistics, and data points that users are trying to find, access, or cite. we are not referring to data collection, management, curation, or analysis. respondent demographics of the 278 respondents retained in this data analysis, 87% (n=187) were from the united states, and 8% (n=21) were from canada. four percent (n=10) of respondents were from other countries, which included south africa, the united kingdom, germany, slovenia, switzerland, and zimbabwe. two percent (n=6) did not respond to this question. of the 241 respondents from the united states, 12% (n=32) were from the western census region, 39% (n=109) were from the midwest, 19% (n=53) were from the south, and 17% (n=47) were from the northeast. the higher number of respondents in the midwest is likely due to our access to local listservs and organizations, including the minnesota library association and the federal depositories in the university of minnesota region of the federal depository library program. the respondents worked at a wide range of library types. the majority (82%, n=228) worked at a university or college library, whereas 7% (n=19) were from public libraries, 3% (n=8) worked at community, technical, or tribal college libraries, and 8% (n=23) were from other types of libraries, including but not limited to state, law, or special libraries. additionally, respondents held a wide variety of roles at their libraries, with most people holding multiple roles. the most common duties included reference emails and consultations, instruction, and liaison or subject specialist roles. many people also staffed a reference or information desk, performed https://doi.org/10.29173/iq939 4/18 kubas, alicia & mcburney, jenny (2019) frustrations and roadblocks in data reference librarianship, iassist quarterly 43 (1), pp. 1-18. doi: https://doi.org/10.29173/iq939 collection development duties, or were a functional specialist in areas such as data services or government information (table 1). table 1 most common duties of respondents which duties are part of your position? n respondents out of 278 possible respondents % out of 278 possible respondents reference/research emails & consultations 216 78% instruction 203 73% liaison/subject specialist 194 70% reference/information desk 171 62% collection development 167 60% functional specialist (data services, government info, copyright, etc.) 131 47% administration 47 17% acquisitions 40 14% e-resources 40 14% web development 32 12% metadata & cataloging 28 10% other 19 7% archives 13 5% results to discover what frustrations librarians experience when answering data reference questions, we included a multiple-choice (select all that apply) question and a free-response question. the first question, 'what are your biggest frustrations or roadblocks when answering data questions?', allowed respondents to check as many of the answers that applied to them. the possible answers (table 2), were chosen based on common themes in current literature and our personal experiences. the second question, 'what other frustrations or roadblocks do you experience when answering data questions?', was left open-ended since we knew the response choices would not encompass all possible answers. this paper focuses on the initial quantitative (categorical) question, but in the future we will examine the open-ended question as well. only the first nine response options for the first question were considered during analysis, and the 'other' and 'na' responses were not included. in our survey, all eight response options were coded as binary variables which we refer to as frustrations throughout this paper. some of these frustrations are roadblocks, meaning that they are external factors that are beyond the control of the data librarian. these roadblocks are encompassed within our larger https://doi.org/10.29173/iq939 5/18 kubas, alicia & mcburney, jenny (2019) frustrations and roadblocks in data reference librarianship, iassist quarterly 43 (1), pp. 1-18. doi: https://doi.org/10.29173/iq939 conceptualization of frustrations, which are any issues a data librarian faces including those that can be improved with experience or training. each of the nine provided response options were selected by at least 25% of the possible respondents, and a large majority of respondents (83%, n=231) indicated that ‘managing patron expectations’ was one of their frustrations (table 2). table 2 – most common frustrations/roadblocks what are your biggest frustrations or roadblocks when answering data questions? n respondents out of 278 possible respondents % out of 278 possible respondents managing patron expectations 231 83% data resources are hard to navigate 127 46% lack of geographic coverage 121 44% lack of time series and/or historical data 115 41% can’t access data due to paywalls 103 37% jargon or vocabulary 94 34% data questions are time consuming 90 32% don’t know where to start looking 71 26% data not in an easily accessible format 70 25% other 22 8% n/a i’m never frustrated with data questions! 9 3% when considering these frustrations, we wanted to see if other variables from our survey were related to the frustrations a librarian experienced. we examined the relationship between the nine frustrations reported and three other multiple choice questions: volume of data questions, years of experience in the field, and type of data questions received in terms of time (historical, time series, current, etc.). frustrations and volume of data questions of the 278 individuals who answered at least one data reference question per month, 22% answered 10 or more questions (n=61), whereas just over 50% answered 1-4 questions (n=141). we did not ask those who reported answering 10 or more questions per month the exact number they received on average per month and therefore re-coded the 10+ answers as 10 for statistical analysis. this means that our statistically significant differences are very likely more significant than they appear in our current data results. https://doi.org/10.29173/iq939 6/18 kubas, alicia & mcburney, jenny (2019) frustrations and roadblocks in data reference librarianship, iassist quarterly 43 (1), pp. 1-18. doi: https://doi.org/10.29173/iq939 figure 1 – % volume of questions received one of our areas of interest was if librarians who answer more questions are more or less likely to have certain frustrations versus librarians who answer fewer data questions. additionally, we wondered if having more or fewer frustrations on average would depend on the number of questions received. we found a statistically significant relationship when examining the number of data questions one receives by one’s self-reported frustration with jargon or vocabulary. overall, data librarians’ likelihood of being frustrated by jargon consistently declined as their number of questions increased. those who receive approximately five or more questions on average per month are less likely to report being frustrated by jargon or vocabulary related to the question or topic than those who answer four or fewer questions (chi-sq.=18.986, df=9, p=0.025). in other words, it is statistically significant that librarians who receive approximately 4 or fewer data questions per month are more likely to report being frustrated by jargon or discipline-specific vocabulary. finally, there was no statistically significant difference in having more or fewer frustrations overall based on the number of questions received. thus, librarians report the same number of frustrations across the board regardless of the average number of questions received, so frustrations and roadblocks do not cease once a librarian starts regularly working on data questions. https://doi.org/10.29173/iq939 7/18 kubas, alicia & mcburney, jenny (2019) frustrations and roadblocks in data reference librarianship, iassist quarterly 43 (1), pp. 1-18. doi: https://doi.org/10.29173/iq939 frustrations and years of experience in the field the respondents also represented a wide range of experience levels (figure 2). early-career librarians with fewer than five years of experience represented 12.6% of respondents (n=35), whereas 26.6% were very experienced with at least 25 years of experience in libraries (n=74). figure 2 – length of time working in libraries we expected that librarians who have spent more years in the field answering secondary data reference questions would report fewer frustrations on average or be less likely to have certain frustrations than those librarians who are less experienced. surprisingly, we found no statistically significant difference in the number of frustrations reported by librarians with more experience answering data questions versus those with less experience. this suggests that data librarians, regardless of how long they have been doing this kind of work, will always be frustrated in some way with the process of doing data reference. this result is similar to what we found when looking at the relationship between frustrations and volume of data questions. frustrations and finding historical and time series data high percentages of respondents reported that they answer questions about three time-based categories of data: historical data, current data, and time-series data (table 3). https://doi.org/10.29173/iq939 8/18 kubas, alicia & mcburney, jenny (2019) frustrations and roadblocks in data reference librarianship, iassist quarterly 43 (1), pp. 1-18. doi: https://doi.org/10.29173/iq939 table 3 – most common types of questions received what kinds of data questions do you receive? (time category) n respondents who chose this frustration out of 278 possible respondents % out of 278 possible respondents historical data 220 79% current data 246 88% time-series (e.g. a 20-year span of the same data over time) 186 67% librarians who report that they receive data questions where the patron is looking for historical data or data reported over time expressed significant variation in their frustrations. librarians who receive historical data questions and/or time series questions are more likely to be frustrated with patron expectations around data availability (chi-sq.=14.826, df=1, p<0.001; chi-sq.=25.050, df=1, p=0.000) 23_5 x 12_1 and 23_5 x 12_3). both types of data come with other frustrations as well: historical data questions inspire frustration with data formats that are not easily accessible (chi-sq.=7.929, df=1, p=0.005), and finding answers to time series data questions can be extremely time consuming (chisq.=4.636, df=1, p=0.031). discussion librarians are increasingly called upon to work with data, even though many lack specific training and expertise. some of the frustrations we will discuss are likely due to increasing pressure to provide data services coupled with the unique demands of supporting data and the experience--often minimal--of the librarians called upon to provide that support. these data librarians continue to play vital roles in the research lifecycle and utilize specific subject and resource expertise to locate secondary data and statistics for patrons. this is a specialized skillset that is honed with experience answering questions and time spent in the field, even though data librarians are already fulfilling many other roles in their institutions. many authors have described the trend toward data services in libraries. in databrarianship: the academic data librarian in theory and practice, kellam and thompson describe major changes in data services over time, from libraries offering data files on magnetic tape in the ’80s and ’90s, to assisting in accessing data via the internet in the mid-2000s, to the more recent interest in data information literacy (p. 2). data services are increasingly part of library services and the librarian’s job scope, and data reference is becoming especially common for librarians who are not specifically 'data librarians.' in databrarianship, bobray bordelon’s helpful chapter 'data reference: strategies for subject librarians' points out that as part of a subject librarian’s expertise in a particular field, they will need to be aware of major data sources and types of data and methodologies common to their subject (p. 36). in numeric data services and sources for the general reference librarian, lynda m. kellam and katharin peter contend that at many institutions, the social sciences or business librarian becomes the de facto data https://doi.org/10.29173/iq939 9/18 kubas, alicia & mcburney, jenny (2019) frustrations and roadblocks in data reference librarianship, iassist quarterly 43 (1), pp. 1-18. doi: https://doi.org/10.29173/iq939 librarian due to the nature of requests the library staff receive, and that this can become a problem if that librarian does not have the necessary knowledge or opportunities for training in data reference (p. 3). they aptly observe that the 'accidental' data librarian is quite common (p. 151). some of the major frustrations we identified in our survey are also reflected in the literature. in 'the pedagogical data reference interview,' kristin partlo describes how students’ expectations of finding data are sometimes very different from reality, and that students’ misunderstandings can lead to challenging reference consultations. in our survey, of the possible frustrations or roadblocks library staff may encounter when answering data questions, by far the most highly selected was 'managing patron expectations,' with 83% of respondents indicating that this was a major frustration. one particular area of frustration can be understanding jargon and vocabulary that is used when talking about data or statistics or when fielding questions in specific subject areas. as rice and southall note in the data librarian’s handbook, 'data-related queries tend to be more difficult, more time consuming, and sometimes even bewildering to the non-subject expert trying to provide assistance to the experienced researcher' (p. 40). they also note that acronyms, jargon, and other vocabulary should be addressed with the user during the reference interview (p. 43). we found this to be true in our survey results as well, finding that librarians who field fewer questions on average struggle with jargon and vocabulary as a frustration. essentially, the less experience a librarian has working with secondary data inquiries, the more time she needs to spend figuring out verbiage related to how data is presented and discussed in a variety of disciplines that have data needs. this extra time required to simply understand the question may lead to more frustration. for example, receiving a data question from an economics student trying to find seasonally adjusted data on the purchasing power of a particular country could cause frustration because the librarian may be unfamiliar with what exactly purchasing power is or how it is measured, or what it means for data to be seasonally adjusted. it is likely that data librarians will always be frustrated in some way working with data, regardless of experience level, sometimes due to factors beyond their control. data continues to be a specialized source type that requires managing user expectations, a great deal of on-the-spot instruction, and navigating a non-standardized discovery environment, which contributes to the level of difficulty and frustration when working with these questions. kellam and peter (2011) note that for many data librarians 'true learning comes from daily work with data-related questions,' and learning on the job is especially typical for data librarians since most come from a variety of educational backgrounds (p. 153). they interviewed a variety of data librarians about their jobs and heard from several who explained that data reference questions are very complex and time-consuming, and that the more questions a librarian answers, the more she knows about finding data and fielding questions (p. 155). bauder (2014) further underscores this idea by noting that 'for librarians accustomed to the relatively organized world of books and journal articles, trying to find data can be a frustrating experience' (p. 11). while examples like this in the literature suggest that librarians with more experience would find answering questions https://doi.org/10.29173/iq939 10/18 kubas, alicia & mcburney, jenny (2019) frustrations and roadblocks in data reference librarianship, iassist quarterly 43 (1), pp. 1-18. doi: https://doi.org/10.29173/iq939 easier and report fewer frustrations overall, our survey found that there were no statistically significant differences in the number of types of frustrations between librarians with more experience answering data questions and those with less experience, based on our measure of experience which encompasses time in the field and volume of data questions. this may mean that data librarians, regardless of how long they have been doing this kind of work and the volume of questions they receive, will always be frustrated in some way with the process of doing data reference. it is possible that some frustrations could be reduced with experience (i.e. not knowing where to start looking), whereas others, that we call roadblocks, are out of the librarian’s control (i.e. paywalls). over time a librarian could learn better ways of dealing with some frustrations, such as certain strategies for managing patron expectations or knowing where to start looking for specific data, but experience will not help overcome roadblocks such as lack of geographic coverage of data or a paywall restriction. in other words, more experienced librarians could better anticipate these challenges and have strategies for dealing with these types of roadblocks, but this does not eliminate them entirely. these factors may improve naturally over time because more content could be digitized or organizations and agencies could release more data that has not been publicly available, but this is once again beyond the control of the data librarian. while current time series data has become easier to locate due to improved online availability, finding historical data has remained difficult due to various issues including geographic, methodological, and sampling technique changes over time. another major issue that comes into play is format changes over time, which we found to be true from our survey where librarians expressed frustration with data formats when looking for historical data. older data is often available only in print, which means that for data analysis purposes, the data will need to be manually entered, likely causing a disruption in patron expectations. furthermore, older data is often exclusively available in print or, particularly in the case of government-produced data, in obsolete formats like floppy disks and cdroms, where an emulator is needed to run software to access the files or particular hardware is required. this is particularly frustrating to the librarian who may need to take quite a bit of time to look at the data in a less accessible format, but also it may be frustrating for the patron since this is likely not what they were hoping for or expecting. additionally, finding historical data can be frustrating for patrons and librarians when the dataset does not go back far enough to the desired time period. kellam and peter (2011) note that this historical data is not possible to find if the dataset wasn’t collected during the requested years (p. 73). due to these various restrictions, data librarians will often need to work with their patrons to help them understand the limitations of historical or time series data. our survey has some limitations that could affect our results. first, while we targeted a number of listservs that we knew included library staff who answer data reference questions, not all data librarians may be on listservs, and we may have missed other major listservs that we are not currently aware of that would have broadened our range of participants. there could be certain types of librarians who answer data reference questions but for whatever reason do not make use of one of the listservs we included. furthermore, the survey was shared informally by participants with other librarians that they knew through channels such as email and twitter, thereby increasing the audience but not in a way that https://doi.org/10.29173/iq939 11/18 kubas, alicia & mcburney, jenny (2019) frustrations and roadblocks in data reference librarianship, iassist quarterly 43 (1), pp. 1-18. doi: https://doi.org/10.29173/iq939 we can track. generally, we do not know the complete size of the target audience or how many potential respondents we reached with our survey distribution, so we cannot calculate the overall response rate or estimate the nonresponse bias. we also had a much greater number of responses from the midwest region of the united states, since that is where we are located and we have access to networks of librarians in this area. we further have very low representation from other countries, again due to our access to networks. these various limitations to our survey participation could lead to regional bias in our results. in our analysis of our survey of librarians who answer data questions, we found that there are many frustrations that they experience in their roles as data librarians. one of our key findings is that frustrations will always exist regardless of a librarian’s years of experience. many factors contribute to the experience of frustration, particularly the popularity of historical and time series data questions. while answering many questions exposes the data librarian to a wide variety of jargon and vocabulary, years of experience in answering data questions does not eliminate the tendency to be frustrated. there will always be aspects of data reference that librarians have little to no control over that lead to frustration, but there are some strategies librarians could employ to lessen and better cope with their frustrations and roadblocks and successfully contribute to data services within the library. more experienced data librarians can share their expertise with colleagues, especially subject liaisons and ‘accidental’ data librarians who may have less experience with data. librarians at all levels can also seek out opportunities to expand their repertoire and depth of knowledge of data sources. they can also advocate for administrative support in training, professional development funding, increased staffing in data services, and opportunities to contribute to strategic planning and discussion at larger structural levels. librarians can also develop strategies for improving data literacy among users, including targeting courses that incorporate data components, being prepared to cover basic data concepts during brief reference interviews, and aligning user expectations with the realities of data reference and discovery. in recognition of this challenging area of librarianship, future directions of this research will include analysis of data librarians’ free text responses of suggested strategies for dealing with tough questions as well as more focused looks into the answers of academic librarians. we are particularly interested in the responses of academic librarians as they made up the majority of respondents and we are interested in learning about the experiences of other librarians in similar roles to ourselves. acknowledgements we thank andrew kubas for his critical role in advising on data analysis and methodology and running analysis in spss. we also thank danya leebaw, amy riegelman, carl mcburney, and the peer reviewers for their feedback and critique. https://doi.org/10.29173/iq939 12/18 kubas, alicia & mcburney, jenny (2019) frustrations and roadblocks in data reference librarianship, iassist quarterly 43 (1), pp. 1-18. doi: https://doi.org/10.29173/iq939 references bauder, j. (2014) the reference guide to data sources. chicago: ala editions. bordelon, b. (2016) ‘data reference: strategies for subject librarians’ in kellam, l. and thompson, k. (eds.) databrarianship: the academic data librarian in theory and practice. chicago: association for college and research libraries, pp. 35-49. kellam, l. and thompson, k. (eds.) (2016) databrarianship: the academic data librarian in theory and practice. chicago: association for college and research libraries. kellam, l. and peter, k. (2011) numeric data services and sources for the general reference librarian. oxford: chandos publishing. partlo, k. (2010) ‘the pedagogical data reference interview’, iassist quarterly, 33(4), pp. 6-10. available at https://iassistquarterly.com/index.php/iassist/article/view/884/876 rice, r. and southall, j. (2016) the data librarian’s handbook. london: facet publishing. https://doi.org/10.29173/iq939 https://iassistquarterly.com/index.php/iassist/article/view/884/876 13/18 kubas, alicia & mcburney, jenny (2019) frustrations and roadblocks in data reference librarianship, iassist quarterly 43 (1), pp. 1-18. doi: https://doi.org/10.29173/iq939 appendix: copy of survey questionnaire answering secondary data questions in a library setting welcome! the goal of this survey is to collect information from librarians and library staff who assist users in locating secondary datasets and statistics. we hope to collect data that is rich enough to inform how this growing area of librarianship is changing and how librarians can support researchers in finding secondary data. this research study is being conducted by alicia kubas and jenny mcburney at the university of minnesota libraries. it should take less than 10 minutes to complete. if you have questions about the survey, please contact alicia kubas (akubas@umn.edu). your participation is completely voluntary. all individual responses will be confidential, and you can stop taking the survey at any time. continuing the survey indicates that you consent to participate. this study has received irb exemption from the university of minnesota. demographics at what kind of library do you work? • university or college (1) • public (2) • community/technical college (3) • tribal college (4) • state (5) • law (6) • special (7) • other (please describe) (8) ________________________________________________ where is your library located? ▼ united states of america (187) ... zimbabwe (1357) display this question: if list of countries = united states of america: in which census region of the us is your library located? • united states west (1) • united states midwest (2) • united states south (3) • united states northeast (4) display this question: if list of countries = united states of america https://doi.org/10.29173/iq939 14/18 kubas, alicia & mcburney, jenny (2019) frustrations and roadblocks in data reference librarianship, iassist quarterly 43 (1), pp. 1-18. doi: https://doi.org/10.29173/iq939 are you a federal depository library program coordinator? • yes (1) • no (2) which duties are part of your position? [check all that apply] • liaison/subject specialist (1) • functional specialist (data services, government info, copyright, etc.) please specify: (2) ________________________________________________ • reference/information desk (3) • reference/research emails & consultations (4) • administration (5) • instruction (6) • archives (7) • collection development (8) • metadata & cataloging (9) • acquisitions (10) • e-resources (11) • web development (12) • other (please describe) (13) ________________________________________________ how long have you been working in libraries? • 0-4 years (1) • 5-9 years (2) • 10-14 years (3) • 15-19 years (4) • 20-24 years (5) • 25+ years (6) data questions for this part of the survey, we are asking about how you answer data-related questions. in this context, 'data' refers to existing datasets, statistics, and data points that users are trying to find, access, or cite. we are not referring to data collection, management, curation, or analysis. on average, how many data questions do you receive per month? • 0 (1) • 1 (2) • 2 (3) • 3 (4) • 4 (5) • 5 (6) • 5 (7) • 6 (8) • 7 (9) • 8 (10) • 9 (11) • 10+ (12) https://doi.org/10.29173/iq939 15/18 kubas, alicia & mcburney, jenny (2019) frustrations and roadblocks in data reference librarianship, iassist quarterly 43 (1), pp. 1-18. doi: https://doi.org/10.29173/iq939 skip to: end of survey if on average, how many data questions do you receive per month? = 0 on average, how much time do you spend on each question? • less than half an hour (1) • half an hour to an hour (2) • between 1-2 hours (3) • more than 2 hours (4) what kinds of data questions do you receive? [check all that apply in the 4 categories below] geography: [check all that apply] • local or state related (1) • regional (e.g. midwest) (2) • national (3) • international (4) • other (please describe) (5) ________________________________________________ time: [check all that apply] • historical data (1) • current data (2) • time-series (e.g. a 20-year span of the same data over time) (3) miscellaneous: [check all that apply] • microdata (1) • spatial data (2) • study data data already collected for research that is now available for reuse (3) • other (please describe) (4) ________________________________________________ topical areas: [check all that apply] • business/economics (1) • agriculture (2) • politics (3) • education (4) • tourism/culture (5) • health (6) • demographic (7) • environment (8) • psychology (9) • science (10) • weather (11) • astronomy (12) • other (please describe) (13) ________________________________________________ what are your strategies for answering these data questions? [check all that apply] • google it to get background information (1) https://doi.org/10.29173/iq939 16/18 kubas, alicia & mcburney, jenny (2019) frustrations and roadblocks in data reference librarianship, iassist quarterly 43 (1), pp. 1-18. doi: https://doi.org/10.29173/iq939 • consult reference tools like wikipedia or free or subscription encyclopedias (2) • ask the patron follow-up questions/reference interview (3) • use personal knowledge of sources (4) • consult research guides (5) • consult with colleagues at your institution (6) • consult with colleagues outside of your institution (listserv, call or email, etc.) (7) • citation pearl growing (using a citation or other piece of information to find additional information) (8) • other (please describe) (9) ________________________________________________ what types of resources do you use to find data? [check all that apply] • library subscription databases (1) • freely available databases (2) • government websites/portals (3) • data repositories (icpsr, etc.) (4) • ngos or intergovernmental organizations (un, eu, worldbank, etc.) (5) • analog data (data in physical form such as print, cassette, cd-rom, etc.) (6) • other (please describe) (7) ________________________________________________ how successful are you at finding the data for which you are looking? • always find it (1) • usually find it (2) • occasionally find it (3) • rarely find it (4) • never find it (5) do you teach data literacy topics (e.g. methodology, authority, sample size, public availability, proprietary information, etc.)? [check all that apply] • yes; as part of a college course (1) • yes; as part of a graduate course (2) • yes; as part of a college or university freestanding workshop (3) • yes; as part of a non-college/university freestanding workshop (at a public library or special library, etc.) (4) • yes; for your colleagues (5) • yes; created tutorials or learning objects (6) • no (7) how do you stay current on developing your personal knowledge of data sources? [check all that apply] • conferences (1) • webinars (2) • blogs or websites (3) • scholarly publications (4) • other (please describe) (5) ________________________________________________ demographics of patrons https://doi.org/10.29173/iq939 17/18 kubas, alicia & mcburney, jenny (2019) frustrations and roadblocks in data reference librarianship, iassist quarterly 43 (1), pp. 1-18. doi: https://doi.org/10.29173/iq939 from whom do you receive data questions? [check all that apply] • community members (1) • k-12 students (2) • undergraduate students (3) • graduate students (4) • faculty (5) • other librarians and colleagues (6) • other (please describe) (7) ________________________________________________ considering data literacy skills (e.g. understanding methodology, authority, sample size, public availability, proprietary information, etc.), on a scale of 1 5, 1 being very weak data literacy skills and 5 being very strong data literacy skills, how data literate are your typical users? • 1 (very weak) (1) • 2 (2) • 3 (3) • 4 (4) • 5 (very strong) (5) challenges and opportunities what are your biggest frustrations or roadblocks when answering data questions? [check all that apply] • don’t know where to start looking (1) • can’t access data due to paywalls (2) • data resources are hard to navigate (3) • jargon or vocabulary related to the question or topic (not knowing what seasonally-adjusted data is, etc.) (4) • managing patron expectations around what data actually exists or its availability (5) • lack of geographic coverage (other countries, local data, etc.) (6) • lack of time series and/or historical data (data from 1952, data from 1970s to present, etc.) (7) • data not in an easily accessible format (in print, on cd-rom, microfiche, etc.) (8) • data questions are time consuming (9) • n/a i’m never frustrated with data questions! (10) • other (you will be able to describe in the next question) (11) what other frustrations or roadblocks do you experience when answering data questions? ________________________________________________________________ what tips, tricks, or strategies for answering tough data questions do you have to share with other library staff? ________________________________________________________________ what data source do you use most often? ________________________________________________________________ https://doi.org/10.29173/iq939 18/18 kubas, alicia & mcburney, jenny (2019) frustrations and roadblocks in data reference librarianship, iassist quarterly 43 (1), pp. 1-18. doi: https://doi.org/10.29173/iq939 end-notes 1 alicia kubas is the government publications and data librarian as well as the regional depository coordinator at the university of minnesota libraries and can be reached by email at akubas@umn.edu. 2 jenny mcburney is the economics librarian, liaison to the institute for advanced study, and the research services coordinator for social sciences and professional programs at the university of minnesota libraries and can be reached by email at jmcburne@umn.edu. https://doi.org/10.29173/iq939 mailto:jmcburne@umn.edu three reasons for the underutilizath of social science data services in the information age alice robbin data and program library service university of wisconsin-madison societies everywhere are being affected by the new information technologies. in addition, they have become increasingly dependent on statistical information for making important public policy decisions. social science data services, a direct result of the new technologies, have been established to provide easier access to computerized statistical information. one would expect therefore to have seen over the last 15 years a great many data services established throughout institutions of higher learning and government. yet these data services are few and the ones that exist, underutilized. there are obviously many reasons for their underutilization. today, i will address three reasons which contribute to the current situation. poor quality data impede good decision making and research. lack of coordination and planning of the statistical information system make it very difficult to produce, locate and retrieve data. new information technologies are modifying our societies. but social scientists are not directing enough attention to how society is being altered and we lack appropriate models and data. my concluding remarks suggest a number of ways that social scientists can contribute to improving the current situation. social science data archives and services, like their predecessor libraries and archives of print documents and film, represent one component of a society's institutional memory. the underlying philosophy of preservation and access holds that transfer of the data collections from their producers to these data centers greatly increases the return on the original public and private investment. this paper was delivered at the 1981 ifdo/iassist conference in grenoble. the author gratefully acknowledges helpful comments on earlier versions from thomas flory, nancy mcmanus, and richard c. roistacher. 17 most of these centers were established before their national archives created machine-readable divisions. although these centers have not been designated official repositories for government records, governments have turned to them for assistance in retrieving government data files. as recent experiences in several countries demonstrate, more government data producers are delegating archival responsibilities to university data repositories in recognition that government cannot preserve and maintain its own records. as society's problems have grown more complex, statistics have become more important to effective decisionmaking. not only do policymakers face increasingly complex issues, but many problems now interact with one another (12,136). the resources of data centers, for holding historical collections of data and for generating new ones, are essential if national policy decisions are to be made in a more rational manner. existing administrative records systems, used for secondary analysis or linked to new data collection activities, provide a means for responding efficiently to new policy questions. data services are also important for cumulative social science research activities. common access creates a "commonality of research among widely separated scholars" (9,411). the data archive acts as a scientific laboratory which encourages the sharing of data, multidisciplinary exploitation of evidence, and "multiple and complex analytic applications" (5,393). the data center makes a pedagogical contri lution by allowing the student to participate in scientific inquiry, developing problem-solving techniques and behavior like those of students in the natural sciences. a recently completed study of factors influencing the sharing of computer-based resources for higher education and research shows a direct connection between utilization and sharing. it suggests that the "seemingly indirect attempts to broaden 'computer literacy' and computer use might have systemic effects on the level and nature of computer-based sharing" (8,4.44). less obviously, the data archive plays a role as an agent for assessment of information transfer activities. it offers administrators and researchers the opportunity to assess the technical, administrative, economic, and policy issues related to standards of data quality, documentation, access, and distribution. nevertheless, 33 years after the roper center at williams college in massachusetts and 20 years after the establishment of the steinmetz archives in amsterdam, the zentralarchiv fur empirische sozialforschung at the university of cologne, and the inter-university consortium for political and social research at the university of michigan, no more than 50-odd data services exist throughout the world, almost all university based. national governments have been slow to accept the idea that data services play an important role in information policy development. information technologies and services produced and offered by the privatefor-profit sector are beginning to dominate access channels. why are there now so few social science data archives and why do they appear to be underutilized? that they have been is due to a wide array of reasons. rather than providing an inventory of these reasons, i will address the complex and interdependent issues of the quality of statistical data, factors responsible for the lack of coordination of data resources, and the need to make social science more relevant to policy choices. most of my remarks have been stated in one form or another during the last five years in many countries. i address the creation of statistical data and administrative records produced by government because it is a major provider of the data resources which social scientists use. and i expect that in the future, government's influence on statistical data production will determine even more how the social scientific community conducts itself. my remarks about the role of social science in an information age have been influenced by recent political events, in which many questions have been raised about the relevance of social science. i believe that relevance implies and requires philosophical reflection. relevance requires use of theoretical perspectives about human and social interests. relevance requires new models which integrate our natural and social worlds with scientific and technological discoveries. my recommendations for improving data quality, planning, and coordination should be understood as two aspects of the larger philosophical and moral dilermas which we confront. thus, the last part of my address reflects on some of the questions social scientists must seek to answer as they confront social changes which are the result of new information technologies. ii. problems of data quality and of coordination and planning a. data quality dissatisfaction with the quality of data is widespread throughout the scientific community and government, although enormous strides have been made to improve measurement, david r. lidd, jr., director of the office of standard reference data at the u.s. national bureau of standards, recently wrote "that a considerable amount of information in such archives is erroneous". he cited almost 200 reported measurements of the heat conductivity of copper--"a range of values so great that most of the data are clearly off the mark" (12). publications of social and science indicators, on which many projections in the united states are based, contain obvious statistical errors--obvious, that is, once the data are examined--and inadequate information on sectors of the society which we know are undergoing rapid changes. these errors are due in part to inadequate sampling frames and improper methodological tools applied to data gathering and analysis. 19 three factors that influence the quality of statistical and other data and their analytic potential are demand (or user requirements), supply (or the resources of the system), and structural or environmental conditions. user requirements . a recently published white paper on the u.s. statistical system notes that "the complexity and urgency of issues facing policymakers often leads them to demand more data and more timely data, with little regard for quality" (11,164). policymakers tend to be uncritical about the quality of the data they use; social scientists only somewhat less so. the immediate demands for completing the administrative function, a budgetary horizon of one to two years, and legislative demands for information for modifying policy impede the necessary gestation period for designing and gathering data. political ends influence the quality of data. "... some of the most important statistics are held hostage to political ends by their visible and direct use in politically important decisions which allocate [national] resources" ( 4,204 ). resources for maintenance and improvement of quality . at least in the u.s., there has been no thorough government-wide review of classification standards for statisticians for about three decades. professional training in data handling is received (or not received, as the case may be) on the job, with little influence by non-governmental sources of expertise. the social science community, which has discovered many useful tools for improving data quality, has little opportunity for interaction with the governmental data producer and statistician. this interaction is not encouraged by government and the university organization nor by attitudes of the government administrator or academician. civil servants' opportunities for career development and participation in conferences such as this one are limited. the white paper offers other explanations. budgets for statistical programs and projects do not include resources for internal and/or external measurement of quality. funds are seldom provided for methodological research to improve quality, except where there are clear indications of serious deficiencies. such deficiencies may not become obvious until the effects of poor policy decisions are felt. political bodies are then moved to apply remedies (which rarely reflect the underlying systemic problem). little attention is given to the basic design of surveys, evaluation studies, program experiments, and data bases developed for policy analysis. competitive procurement activities (contracts, for example) seldom receive adequate technical review, and selection panels often lack the technical skills to make an informed judgment (11 ). a 1978 study by the u.s. general accounting office of federallysponsored attitude and opinion surveys found serious technical flaws which limited the usefulness of the results in all five surveys which were reviewed in detail. the gao concluded that "better guidance and controls were needed to improve federal surveys of attitudes and opinions" (11,162). another study, sponsored by the american statistical association 20 and funded by the u.s. national science foundation, evaluated 26 sample surveys conducted in 1975 and found that 15 of the 26 surveys had serious technical flaws. all but two of the 26 federally sponsored surveys were conducted under contract by universities or other private survey research organizations (1 ). structural factors affecting quality . increasingly, statistical services are being procured from outside the government under contract. agencies often have funds to acquire these statistical services, but no budget to develop staff and inhouse organs to build services and decide on technical specifications and selection. operations which include data collection by other units of government are notoriously difficult to monitor and to standardize. for example, a large portion of the data collection activities conducted under the auspices of the intergovernmental cooperative health statistics system program in the u.s. is being eliminated; quality control was cited as a major factor in this decision (13). producers outside government are typically unaware of the uses to which their data will be put, or of the utility of the data they provide or of the administrative needs of an agency. analysts are often unaware of important limitations of data because technical standards of data description have not been instituted by government agencies. restrictions on interagency sharing often result in the lack of comparability in data produced by different agencies. such restrictions sometimes result in failure to fully exploit expensive data bases. although policy may require linkages of materials gathered in several agencies and from several records series, legal procedural, and operational mechanisms to provide linkage are few and far between(2). b. lack of coordination and planning poor information management practices applied to statistical and administrative records and the internal organization of bureaucracy are in part responsible for difficulties in accessing records. these problems have led "to a growing incidence of overlap, duplication, mismatch and gaps in data and analysis, and increasingly complex problems of access by users and statistical agencies to various federal data" (11,143). nora and mine give three examples of this kind of compartmentalized development in france. hospitals have developed systems for billing medical expenditures and hospital -stay expenditures without collaborating with social security. within social security itself, compartmentalization into three branches, each with its own data processing centers, has led to manual retrieval of data produced by the computers of the other branches. as a result of the present departmental separation [they write before various reorganizations within the mitterand government], the direction ge'ne'rale 21 des impots and the direction de 1 ' amenagement foncier et de i'urbanisme (land development and urban affairs) has each established a land use data bank, the former for tax purposes, the latter for development purposes. the legal definitions and the types of information differ. nevertheless, there are broad common areas, but nobody worries about them. in addtion to the waste, the establishment of these two data banks prolongs administrative isolation. strengthened by this investment, both administrations are prepared to resist attempts at rapprochement (10,115). within the u.s. government, the federal trade commission in its quarterly financial reports asks for data which are available in quarterly filings with the securities and exchange commission. and there are currently three duplicate mortgagee interest surveys (11,149). in wisconsin, the department of public instruction refuses to turn over computerized records that the department of revenue needs for statistical analyses and modeling. the department of revenue is forced to collect this information manually if it is to perform its work in a timely way. the application of data processing technologies has been uneven throughout government, and as nora and mine note, although "penetration has been extremely rapid," it has "taken place in uneven ways, strengthening barriers, immobilizing the structures that it penetrates for a long time" (10,112). they note that in the majority of cases, each department acquires data processing capabilities without worrying about the possible difficulties that its plan may cause elsewhere, and especially without measuring the "synergistic" effects that better coordination with other departments might have produced (10,112). the high rate of change in administrative data processing has resulted in a phenomenon that could be called input without throughput. delays in the implementation of data base management systems, complications in electronic data entry systems, pressures to maintain routine adminstration in the face of high staff turnover in data processing, and the imposition of computer technology on organizations designed for manual systems have created serious bottlenecks in routine administration. procurement policies emphasize centralization and are costly and a serious impediment to acquiring the most economical and efficient technology available. little attention is given to identifying areas where decentralization of the information system would improve an agency's capabilities. on the other hand, administrators have few possibilities and little incentive to improve coordination because statutes delimit an agency's mission. even when research access to identifiable information is not in question, attention has not been given to maintenance and preservation of machine-readable records. constraints on administrative activity 22 tend to reduce incentives for "backward" looks, those that would require that records be maintained and preserved. the resulting costs can be very high. for example, efforts now underway to create public use samples from microfilmed versions of the 1940 and 1950 u.s. censuses of population are to cost $8 million. much of that information was on punch cards at one time. records managers and archivists do not usually participate in decisions about retaining and destroying computerized records. as a result, computerized records are not integrated into records management practices. records managers leave decisions about retention to those with programnatic responsibility and concern themselves with managing paper and microfilm records. records and computer centers see themselves as repositories for magnetic tape, with responsibility for decisions about tape maintenance left in the hands of an agency. individual analysts retain information on the contents of files for which they have programmatic responsibility. data processors are often the only persons knowledgeable as to format and physical attributes of computerized records. documentation for mrr may not exist or may be scattered among the various agency personnel responsible for the different aspects of mrr. valuable data are routinely erased and the tapes are reused when tape shortages occur, often without prior systematic review. iii. society and the new information technologies the emerging information technologies are already altering the nature of our society and affecting existing political, economic, and social institutions and values. data processing is accelerating production, with less but more effective work and jobs very different from those imposed by industrial life. this change has already begun: a great decrease in the labor force in the primary and secondary sectors, an increase in the services, and above all, a multiplication of activities in which information is the raw material (10,126). already, computerization of formerly manually performed tasks is rendering the semi-skilled and unskilled worker unemployable. robots are beginning to replace humans, performing certain tasks more efficiently and increasing industrial productivity. however, not only the unskilled or low-skilled are being replaced. the introduction of automation is affecting highly skilled technical workers. for example, although more than 12,000 air traffic controllers walked off their jobs in the united states, air traffic was only partially reduced because computers assisted in air traffic decision making. in the opinion of some, computers were used as a strike-breaking tool (3). the federal aviation administration hopes within 10 years to have computerized en route air control to such an extent that at least 50% fewer controllers will be needed and those that will be needed will be computer managers(6) 23 economic changes will be accompanied by a change in the structure of organizations and by fluctuations in attitudes toward work. as numerous examples have demonstrated, the new technologies related to automation and data processing can flourish in small as well as large organizations. the psychological and social bonds that were created by the work place and that fostered worker solidarity will weaken as automation enforces isolation. monetary and other rewards will go increasingly to those who have the means to produce and manipulate the technology, creating new elite structures and placing political decision-making in the hands of technicians. as duncan mcrae has noted, the "risk of technocracy lies in the possibility of uncontrolled power held by an elite and devoted to special values and interests rather than to the general welfare" (7,45-46), iv. recommendations in what ways can social scientists contribute to improving the present environment of the information system? the information system in which statistical data production and analysis take place is highly complex and dependent on new technologies. it requires expertise from many disciplines and specializations. it requires modifications in the institutional framework in order to cope effectively with societal change and to anticipate unexpected policy and political demands. the social scientist and policymaker have many common interests. they have a great deal to gain by cooperating, to improve the quality of data, coordination and planning, and access to computerized records. governments must use available expertise "in data collection and analysis activities, starting at the design stage, and continuing through to evaluation of how results are used" (11,166). social scientists can contribute through methodological research in measurement of errors to improving collection methods and to improving the presentation of information about methodology structure and other limitations of the data products and analyses. the results of methodological research must then be widely disseminated so that they can be evaluated, criticized, and competing methods proposed if necessary. we must be concerned with creating an integrated output and with producing cross-cutting analyses over a wide range of issues. social scientists can assist in substantive integration activities, by developing standard concepts, definitions, classifications, survey frames, and procedures, and by monitoring and promoting their utilization by government and by the private sector. social scientists can assist in developing a "consistent conceptual framework or model based on behavioral relationships in various disciplines" (11,172). there needs to be increased use of administrative records to produce statistics and to respond to public policy questions. public use samples should be drawn from administrative records. administrators should be 24 made to produce public use files and to coordinate record linkage and analyses. through their activities, social scientists can promote record linking at the microlevel and demonstrate ways in which the data's analytic potential can be enhanced. (it is important to note, by way of illustration, that social scientists and government officials in germany have been meeting to discuss the creation of public use samples. this meeting should be emulated by other countries.) some of the problems of use of social science methodological and policy research can be traced to the fact that researchers are not part of the policy formation activities of government. if social researchers are to play a greater role in social policy formation and are to increase utilization of their research, there must be a higher rate of communication between researchers and policymakers. this communication is more successful if social scientists participate in internal organizational decisions (14). social scientists must make a concerted effort to involve themselves in these decisions. involvement in the internal decision-making process will indirectly improve the quality of civil servants' activities and directly improve utilization of their research and policy recommendations. with administrators and policymakers, social researchers can assess research needs and examine the relation of the statistical system to research activities outside the government. they can apply their training in organizational theory and public administration to improving information management activities in government. indeed, some of these very activities are already underway in italy, norway, germany, the unitec. states, and great britain. closer ties between data producers and analysts will result in data that are more relevant to policy issues and will also improve the quality of both data and analyses. producers of data will have more direct feedback on quality from major users of data . . .users will come to have a better understanding of the operational problems of collecting and processing data, and will design and perform their analyses with a better understanding of the limitations of the data (11,168). what should be the role of social science in an information age? this is a much more difficult question than the one which asks what knowledge should be applied and how? let me identify only a few salient public policy issues that form part of an agenda for information technologies-related social research and training. (1) society will require a decreasing amount of work. will work as a value lose its importance? how will the remaining work be distributed? what educational and job training programs will be needed, ones that are more compatible with the requirements of the post-industrial and infonnation age? if the number of hours of leisure time is 25 increased, what social and psychological changes will occur; what changes will be necessary? (2) new organizational structures are evolving and, increasingly, innovation takes place and new products develop in small units. what should be the role of the state in reorganizing the production structures? how do we design tax policies and write administrative regulations to provide incentives for industrial and university research and development, to foster innovation and risk-taking in the highly productive information technologies? if basic research outside industry is a prerequisite for innovation and continuing productivity, are the existing models of research in a more decentralized fashion, along the lines of the u.s. model, or research in the colbertist tradition any longer relevant; or is some mix more appropriate to optimize available resources and to encourage innovation? (3) critical shortages of trained scientific and technical personnel are beginning to be felt. in what ways can we improve the quality of our science and social indicators to reflect the current situation? how can we estimate the impact of these shortages on the economy and on a nation's productivity? what roles should the state and the private sector play in ameliorating these conditions? if university budgets continue to experience serious erosion, how will a nation's productivity and general welfare be affected? yet, if attention is turned only to reducing these shortages, do we risk neglecting the education of the "well -informed citizen" who is necessary for democratic control of technical decisions? "^o we thus accelerate the creation of a society which is, to quote shils, "victim of the parochial preoccupations of specialized technical experts"? [in mcrae). if we emphasize scientific knowledge to the detriment of valuative discourse will we neglect the education of both the. scientist and the consumer of technology? (4) the design industry and regulatory arms of the state have been preoccupied with hardware systems, with minimal consideration of human factors and a disregard for worker participation. the accident at three mile island nuclear power facility on march 28, 1979, dramatically illustrates the failure to integrate the reactor operator into the system. the kemeny commission pointed to the mutual isolation of the operator and equipment in the highly complex sociotechnical system as a root cause of the accident (15, 57). the social scientist malcolm brooks observed that the events were a direct function of the electro-mechanical system design and detail (15, 58). in what ways can we improve the man-machine interface in order to reduce isolation and alienation? if it is necessary to modify the work environment, in what ways? are our theories of participatory democracy relevant to the emergence of new environments based on information technologies? (is the model of industrial democracy relevant in a postindustrial information society?) can the new information technologies and new sources of knowledge enhance autonomy and responsibility, make possible mastery of the natural and social world, and emancipate rather than imprison us? 26 (5) instrumental reason has spread to many areas of social life and there is an increasing tendency to define practical problems as technical issues. will technocratic domination erode the institutional framework of society? what value system will it dictate? will the technical values of efficiency and economy dominate the selection of means for realizing social goals? (6) the ability to communicate has always been the purview of the educated and dominant classes. will standardization of access vocabularies affect language and syntax and authority structures? if language will be of a different nature, simplified, to reduce communication costs, will we then sacrifice part of the content? what will occur when the essential meaning of messages related to daily life becomes available to anybody? will new communication structures create more open and accountable authority structures? do they offer the potential of transforming the state into one more easily supervised by the "public"? (7) the cultural model of a society also depends on its memory, control of which largely conditions the hierarchy of power. will access to infinitely greater sources of information entail basic social changes and affect the social structures by modifying the procedures for acquiring knowledge? (10, 313). how will data banks restructure knowledge? how much social control will be exercised by the producers of data banks? to understand the nature and direction of technological change demands a vigorous and sustained program of social research related to information technologies. tie frameworks of the social science disciplines and social thought can help us in orienting our discourse and directing it to problems of action and choice. new information bases and new knowledge can improve political choices in an increasingly technological society. they can assist social groups to transform society, to use new resources effectively and to their benefit, and to create control mechanisms for the new information order. this effort requires engaging and appropriating competing traditions of philosophy and social thought, new philosophical approaches and different methodologies, and creativity and innovation unfettered by the narrow confines of the empirical sciences. references (1) bailer, b.a. and lanphier, cm. (1978) development of survey methods to assess survey practices . washington, d.c.: american statistical association. (2) david, m. and robbin, a. (1981) the great rift between administrative records and knowledge created through secondary analysis. review of public data use , (forthcoming) (3) faa withheld air control system until strike: aspin. capital times, august 24, 1981, 23. 27 (4) federal statistical system project staff. (1980) improving the federal statistical system: report of the president's reorganization project for the federal statistical system. statistical reporter , may, 1980, 197-212. (5) hofferbert, r.i. and clubb, j.m. (1976) introduction. american behavioral scientist . 19 (4), 381-386. (6) kolata, g.b. (1981) faa plans to automate air traffic control. science . 213, 21 august, 845-46. (7) mcrae, d. (1976) the social function of social science . new haven, ct; yale university press. (8) mebane. d.d. and mebane, r.m. (1981) factors influencing the sharing of computer-based resources for higher education and research . princeton, nj: educom. (9) miller, w.e. (1976) the less obvious functions of archiving survey research data. american behavioral scientist , 19(4). 409-418. (10) nora, s. and mine, a. (1976) the computerization of society |l' informatisation de la society . cambridge, ma: mit press .translation). (11) the president's reorganization project for the federal statistical system. (1981) improving th-^ federal statistical system: issues and options. statistical re jorter , february 1981, 133-221. ("white paper"! (12) sullivan, w. (1981) data services map the way in labyrinths of information. the new york times , august 23, 1981. (13) u.s. department of health and human services. (1980) directions for the '80s . final report of the panel to evaluate the cooperative health statistics system . washington, d.c. national center for health statistics. (dhhs publication no. (phs) 801204). (14) van de vail, m. and bolas, c.a. (1981) external vs. internal social policy researchers. knowledge: creation, diffusion, utilization , 2(4), 461-482. (15) wolf, c.p. (1979) the accident at three mile island: social science perspectives. items, 33(3/4), december 1979, 56-61. 28 1/7 ologbosere, oluwatosin abiodun (2023) data literacy and higher education in the 21st century, iassist quarterly 47(3-4), pp. 1-8. doi: https://doi.org/10.29173/iq1082 the creative commons-attribution-noncommercial license 4.0 international applies to all works published by iassist quarterly. authors will retain copyright of the work and full publishing rights. data literacy and higher education in the 21st century oluwatosin abiodun ologbosere 1 abstract this abstract discusses the significance of data in the era of big data, emphasizing its role as a fundamental building block of truths. the concept of datafication, the transformation of various aspects of life into digital data, is explored, focusing on the emergence of data literacy as a crucial subset of information literacy necessary for navigating the virtual landscape. the write-up underscores the skills essential for data literacy, highlighting the role of data science, authentic context, and quantitative reasoning. it emphasizes the importance of data literacy programs covering data analysis techniques, real-world applications, critical thinking in data, and data ethics. the mention of metaskills as higher-order abilities crucial for navigating the dynamic twenty-first-century environment adds depth to the discussion—the abstract delves into the distinction between data literacy and information literacy, emphasizing their complementary nature. integrating data literacy into educational programs, particularly in libraries, is stressed for relevance in meaningful information resource utilization. the context extends to higher education in nigeria, where the role of institutions in developing a knowledge economy and human capital is explored. the abstract underlines the global importance of higher education for sustainable development and emphasizes the critical role of data literacy in this context. challenges faced by nigeria, including research productivity and data literacy, are discussed, highlighting the need for skills like critical thinking and data comprehension in the twenty-first century. the abstract concludes by advocating for incorporating data literacy education at all levels in nigeria's educational system to foster growth, development, and informed decisionmaking. keywords data, higher education, 21st century introduction data has long been a driving force in science, and it is now doing the same in vocational education, where success is frequently determined by the ability to comprehend information. nigeria's data literacy is gradually gaining importance as the country embraces digital transformation; there is a growing awareness of the need for individuals and organizations to understand, interpret, and use data effectively (deja, rak, & bell, 2021). data literacy is a foundational skill that empowers individuals in higher education to navigate the evolving landscape, make informed decisions, foster innovation, and prepare students for the challenges of the 21st century. furthermore, in an era of increasing interdisciplinary research and technological advancement, data literacy is essential for researchers in order to be able to collect, analyze, and interpret data to conduct meaningful research and contribute https://doi.org/10.29173/iq1082 https://creativecommons.org/licenses/by-nc/4.0/ 2/7 ologbosere, oluwatosin abiodun (2023) data literacy and higher education in the 21st century, iassist quarterly 47(3-4), pp. 1-8. doi: https://doi.org/10.29173/iq1082 innovatively. the role of data in academic institutions is undergoing a profound evolution driven by technological advancements changing educational paradigms, therefore making any institutions that prioritize data literacy better positioned to adapt to change, improve educational outcomes, and contribute to the overall advancement of society. however, challenges such as limited access to quality education and technology infrastructure still impact the widespread development of data literacy ( raffaghelli, manca, stewart, prinsloo, & sangrà, 2020). data literacy is the ability of an individual to identify, work with, analyze, and communicate data effectively. in the rapidly evolving landscape of higher education in the 21st century, a critical gap exists among students and educators in data literacy skills. this deficiency needs to improve the ability of academic institutions to harness the full potential of data-driven decision-making, improve students' learning outcomes and promote institutional advancements. the data literacy landscape in higher education could be characterized by a proactive approach to integrating data skills into academic programs, a growing recognition of its importance, efforts to integrate data literacy into curricula, and increased demand for data skills among students. in the age of big data, the sheer volume of digital resources is overwhelming. data are perceived as the fundamental building blocks of truths. data collection, processing analysis and use for decision-making form the foundation of self-efficacy, which gives people control over their lives and careers (carlson, johnston, westra & nichols, 2013). collecting, deploying and analyzing data for decision-making is referred to as datafication. datafication transforms various aspects of life, activities, and information into digital data. in the modern digital age, many activities and phenomena are being quantified, recorded, and analyzed as data. this process involves converting analog information into a digital format that computer systems can store, process, and analyze. data, the smallest quantitative unit of knowledge, are now used to interpret reality objectively. data literacy is regarded as a subset of information literacy (information). people may require skills such as data literacy and specialized knowledge to be at ease and competent in virtual information. this may also apply to academics and postgraduate students, where data literacy skills may benefit professional advancement (mandinach & gummer, 2016). overview of data literacy data literacy is defined as the ability to use critical thinking to draw valuable conclusions from data, make sense of abstractions, and apply analysis results (shreiner & dykes, 2021). comprehending abstractions becomes critical in becoming data literate because data is meaningless unless linked to indicate a relationship between concepts; this is because the ability to make sense of multiple independent concepts and infer the specific link between them is similar to computational thinking. this results from careful thought and critical analysis, improving the individual's skills and producing more insightful interpretations of the data (davenport & patil, 2012; elder & paul, 2020). data science, including data collection, calculation, analysis and interpretation, and communication, are examples of primary data literacy skills. for students entering the labor market, these skills could be measured with a test at the end of the course. authentic context and quantitative reasoning are essential skills related to intermediate and advanced data literacy. data collection, calculation, analysis and interpretation, and communication skills are critical for teaching students, academics, https://doi.org/10.29173/iq1082 3/7 ologbosere, oluwatosin abiodun (2023) data literacy and higher education in the 21st century, iassist quarterly 47(3-4), pp. 1-8. doi: https://doi.org/10.29173/iq1082 and other tertiary institution staff members how to critically evaluate the utility of data models as part of teaching basic data literacy skills (mckendrick, 2015). understanding data enables the use of data insights to make predictions. as a result, data literacy programs should assist people in improving their understanding of patterns and their ability to overcome obstacles as they arise (klidas & hanegan, 2022). data literacy programs involve content such as data analysis techniques, real-world applications, critical thinking in data and data ethics and privacy, among others. when properly designed and implemented, programs like these will reinforce data literacy in society while promoting the development of other abilities known as meta-skills (liquete, 2012). meta-skills refers to higher-order abilities that enable individuals to acquire, adapt, and apply various specific skills. examples are critical thinking, problem-solving, communication, and learning how to learn. a meta-skilled individual can navigate a dynamic and evolving work environment. (kumar, kumar, & lochab,2022). meta-skills are critical for increased participation in the rapidly changing twenty-first century. however, they are not fundamental but build on the foundation of data literacy and pre-existing abilities. carlson and johnston (2015) provided the foundational skills for data information literacy based on a literature review written in their book after evaluating students' performance that "the high level of interest in these competencies are divided into twelve (12) major themes: introduction to databases and data formats, data discovery and acquisition, data management and organization, data conversion and interoperability, quality assurance, metadata, data curation and reuse, cultures of practice, data preservation, data analysis, data visualization, and ethics, including data citation. sub-themes and specific lessons must be learned to gain mastery and competence in data information literacy. furthermore, after evaluating students' performance in their book, "the high level of interest in fundamental topics, such as data formats and an introduction to databases, indicate the relative need for preparation in the core technological skills required to work in an e-research environment. students lack the fundamental technological skills required to function in a data-driven society, implying that this tendency exists in other societies carlson & johnston, 2015). it is widely agreed that a data-literate person—someone who can comprehend and evaluate knowledge derived from accurate data or facts—should be able to apply mathematical ideas to a realworld problem related to his or her area of expertise and demands. it skills, the third component of data literacy, are intended to simplify data analysis and synthesis by displaying facts in a virtual form. as a result, a person with basic data literacy skills should be able to find and use appropriate it resources for his or her needs first and foremost (gebre, 2018). however, teaching it skills is complicated because users typically do not investigate database technologies until they genuinely need to collect and use data (maybee & zilinski,2015). as a result, demanding that people learn these abilities in an entirely simulated educational environment, in the abstract, is difficult. more often than not, scholars distinguish between data literacy and information literacy (palsdottir, 2021). data literacy is related to information literacy (shields, 2005; koltay, 2016). the quality of information is critical in the data literacy-information literacy relationship. the level of trust placed in the source, that is, the level of transparency, is determined by how quickly one can recognize the veracity of facts, how quickly one can grasp their logic, how valuable the data are, and how closely they adhere to the standards (koltay, 2016). to evaluate source material in data literacy, knowledge of primarily (but not exclusively) quantitative approaches to data model construction, as a structure https://doi.org/10.29173/iq1082 4/7 ologbosere, oluwatosin abiodun (2023) data literacy and higher education in the 21st century, iassist quarterly 47(3-4), pp. 1-8. doi: https://doi.org/10.29173/iq1082 describing a set of intentionally gathered facts, is required. in this sense, information literacy is a broader concept, with qualitative analytical methods much more frequently used to confirm the reliability of information (shrestha, 2018). to enable data literacy in practice, it is paramount to emphasize the importance of information literacy-like critical thinking about information resources. nowadays, data literacy is just as important as information literacy. they complement each other well and unmistakably contribute to libraries' educational mission of encouraging the meaningful use of information resources to create knowledge and invent new things. as a result, it is well justified for inclusion in library educational programs. academics and librarians can use data literacy skills to design learning programs that integrate with the faculty and demonstrate the fruit of their labor to avoid a situation where the results of their academic work are indiscernible (augood, 2019). higher education in nigeria bernett (2017) sees higher educational institutions as distinct from others in terms of research, defining higher education in terms of the levels and functions of the educational experience offered. higher education has been widely acknowledged as a critical tool for developing a knowledge economy and human capital worldwide ( adepoju & okotoni 2018). according to peretomode (2018), higher education is the facilitator, bedrock, powerhouse, and driving force for a nation's socioeconomic solid, political, cultural, healthier, and industrial development, as higher education institutions are increasingly recognized as wealth and human capital-producing industries. higher education is essential for all developing countries if they are to prosper in a global economy where knowledge has become a critical competitive advantage. the quality of knowledge generated in higher education institutions is critical to national competitiveness. countries can achieve sustainable development by improving the skills of their human capital through higher-level training. higher-level human resources training has been recognized as a primary tool for national development on a global scale. such high-level educational provision enables citizens to acquire skills and techniques that can be applied to increase human productivity (thom-otuya & inko-mariah, 2016). according to the federal ministry of education (2004) section 8 (59), the goals of higher education in nigeria are as follows: contribution to national development through high-level workforce training development and instillation of appropriate values for individual and societal survival, individuals' intellectual capabilities to understand and appreciate their local and external environments are being developed, acquisition of physical and intellectual skills that will allow the individual to be a selfsufficient and valuable member of society. scholarship and community service for national unity and national and international understanding and interaction are promoted and encouraged. nigeria has a population of approximately 154 million people. a growing population necessitates expanding higher education to meet the quality challenges in higher education in nigeria in the twenty-first century. data literacy and the 21st century in nigeria the world is changing, and to remain relevant, one must change as well. the rate of change in the world is accelerating, and inventions are appearing at an unprecedented rate. nigeria, a developing country with an active and prosperous population, cannot catch up in this race. the country's ability to advance as it should has been hampered by low research productivity and data literacy levels. combing through the literature, countries with higher data literacy rates develop faster. (pingali, aiyar, https://doi.org/10.29173/iq1082 5/7 ologbosere, oluwatosin abiodun (2023) data literacy and higher education in the 21st century, iassist quarterly 47(3-4), pp. 1-8. doi: https://doi.org/10.29173/iq1082 abraham, & rahman, 2019). nigeria has one of the largest populations of young people, with a better chance of providing a unique opportunity to build a productive society. this has not been the case, as several hindrances have slowed the country's growth and development. being on the cutting edge of technological and informational advancement is critical. in the twenty-first century, some talents are in demand and will provide a better future with appropriate applications. as a result, citizens of the twenty-first century must be capable of dealing with issues such as critical thinking, data comprehension, and making data-driven decisions (chinien & boutin, 2011; wanner, 2015). data literacy education in post-primary (secondary) education in europe is gaining increased recognition as societies become more data-driven. the emphasis on data literacy reflects the growing importance of understanding, interpreting, and critically evaluating information in various forms. (fontichiaro, & johnston, 2020). this, however, is different with nigeria. nigerian scholars need to be exposed early enough to the rudiments of data literacy. data literacy is one of the skills every nigerian in the twenty-first century must have. aside from the individual, having these abilities will allow nigeria to enjoy advancements and consistent growth and development. data literacy education has been shown to improve students' study habits and learning skills because these skills are fundamental to developing other literacies (information, statistical, digital, media, computational, and visual), known as meta or trans-literacy. furthermore, studies have shown that higher-education students can better manage higher-order thinking and provide more practical answers (mackey and jacobson, 2011; vahey et al.,2012). this will affect the research output of these societies, allowing their students to more effectively validate and generate research with a high impact factor and broad applicability (frau-meigs, 2012; hattwig, bussert, medaille, & burgess, 2013). consequently, to achieve this result, these abilities must be introduced into society and the educational system when they can pique individuals' interest and lead to their success. conclusion scholarly work focusing on conceptualizing information literacy to improve learning in higher education may shed light on the emergence of data literacy in nigeria. for higher education institutions to reach their lofty goals, in the 21st century, students, academics, and other staff must be able to analyze and manipulate data to make informed decisions. data literacy skills must be taught at all levels of education in nigeria, particularly at higher learning institutions. references adepoju t.,& okotoni c.(2018) higher education, knowledge economy and sustainable development in nigeria, journal of education and practice 9, no 18,issn 2222-288x augood, d. c. (2019). excavating an occluded genre: creating visibility of disciplinary values and goals in prompts (doctoral dissertation, california state university, sacramento). bernett r (2017) higher education: a critical business. buckinggham: the society for research university press. carlson, j., johnston, l., westra, b., & nichols, m. (2013). developing an approach for data management education: a report from the data information literacy project. the international journal of digital curation, 8(1), 204-217. https://www.doi.org/10.2218/ijdc.v8i1.254. https://doi.org/10.29173/iq1082 https://www.doi.org/10.2218/ijdc.v8i1.254 6/7 ologbosere, oluwatosin abiodun (2023) data literacy and higher education in the 21st century, iassist quarterly 47(3-4), pp. 1-8. doi: https://doi.org/10.29173/iq1082 carlson, j., & johnston, l. (2015). data information literacy: librarians, data, and the education of a new generation of researchers. purdue university press. chinien, c., & boutin, f. (2011). defining essential digital skills in the canadian workplace: final report. retrieved from http://www.nald.ca/library/research/digi_es_can_wor kplace/digi_es_can_workplace.pdf davenport, t., & patil, d. (2012). data scientist: the sexiest job of the 21st century. retrieved from harvard business review: https://hbr.org/2012/10/data-scientist-the -sexiest-job-of-the-21stcentury/ar/1 deja, m., rak, d., & bell, b. (2021). digital transformation readiness: perspectives on academia and library outcomes in information literacy. the journal of academic librarianship, 47, 102403. elder, l., & paul, r. (2020). critical thinking: tools for taking charge of your learning and your life. foundation for critical thinking. fontichiaro, k., & johnston, m. p. (2020). rapid shifts in educators' perceptions of data literacy priorities. journal of media literacy education, 12(3), 75-87. frau-meigs, d. (2012). transliteracy as the new research horizon for media and information literacy. media studies, 3(6), 14-27. gebre, e. h. (2018). young adults’ understanding and use of data: insights for fostering secondary school students’ data literacy. canadian journal of science, mathematics and technology education, 18(4), 330-341. klidas, a., & hanegan, k. (2022). data literacy in practice: a complete guide to data literacy and making smarter decisions with data through intelligent actions. packt publishing ltd. koltay, t. (2016). data governance, data literacy and the management of data quality. ifla journal, 42(4), 303-312. kumar, p., kumar, s., & lochab, a. (2022). impact of individual personality traits on organizational commitment of it professionals in india: the moderating role of protean career. south asian journal of management, 29(1). liquete, v. (2012). can one speak of an “information transliteracy”? international conference: media and information literacy for knowledge societies. moscow, russia. retrieved from https://hal.archives-ouvertes.fr/hal00841948 . mackey, t. p., & jacobson, t. e. (2011). reframing information literacy as a metaliteracy. college & research libraries, 72 (1), 62-78. doi:10.5860/crl-76r1. mandinach, e. b., & gummer, e. s. (2016). data literacy for educators: making it count in teacher preparation and practice. teachers college press. maybee, c., & zilinski, l. (2015). data informed learning: a next phase data literacy framework for higher education. proceedings of the association for information science and technology, 52(1), 1-4. mckendrick, j. (2015). data driven and digitally savvy: the rise of the new marketing organization. forbes insights. retrieved from https://www.turn.com/livingbreathing/asset s/089259_datadriven_and_digitally_savvy_the_rise_o f_the_new_marketing_organization.pdf https://doi.org/10.29173/iq1082 http://www.nald.ca/library/research/digi_es_can_wor%20kplace/digi_es_can_workplace.pdf https://hbr.org/2012/10/data-scientist-the%20-sexiest-job-of-the-21st-century/ar/1 https://hbr.org/2012/10/data-scientist-the%20-sexiest-job-of-the-21st-century/ar/1 https://hal.archives-ouvertes.fr/hal00841948 https://www.turn.com/livingbreathing/asset%20s/089259_data-driven_and_digitally_savvy_the_rise_o%20f_the_new_marketing_organization.pdf https://www.turn.com/livingbreathing/asset%20s/089259_data-driven_and_digitally_savvy_the_rise_o%20f_the_new_marketing_organization.pdf 7/7 ologbosere, oluwatosin abiodun (2023) data literacy and higher education in the 21st century, iassist quarterly 47(3-4), pp. 1-8. doi: https://doi.org/10.29173/iq1082 palsdottir, a. (2021). data literacy and management of research data–a prerequisite for the sharing of research data. aslib journal of information management, 73(2), 322-341. peretomode v.f (2018) what is higher in higher education. benin-city: justice jecko press and publishers ltd. pingali, p., aiyar, a., abraham, m., & rahman, a. (2019). transforming food systems for a rising india (p. 368). springer nature. raffaghelli, j. e., manca, s., stewart, b., prinsloo, p., & sangrà, a. (2020). supporting the development of critical data literacies in higher education: building blocks for fair data cultures in society. international journal of educational technology in higher education, 17, 1-22. shields, m. (2005). information literacy, statistical literacy, data literacy. iassist quarterly, 28(2-3), 66. shreiner, t. l., & dykes, b. m. (2021). visualizing the teaching of data visualizations in social studies: a study of teachers’ data literacy practices, beliefs, and knowledge. theory & research in social education, 49(2), 262-306. shrestha, b. (2018). information literacy at the workplace: digital literacy skills required by employees at the workplace. thom-otuya, b. e., & inko-tariah, d. c. (2016). quality education for national development: the nigerian experience. african educational research journal, 4(3), 101-108. vahey, p., rafanan, k., patton, c., swan, k., van’t hooft, m., kratcoski, a., & stanford, t. (2012). a cross-disciplinary approach to teaching data literacy and proportionality. educational studies in mathematics, 81, 179-205. doi:10.1 007/s10649-012-9392-z. wanner, a. (2015). data literacy instruction in academic libraries: best practices for librarians. archival and information studies student journal, 1, 1-17. retrieved from http://ojs.library.ubc.ca/index.php/seealso/article/view/186 335. world bank (2004) improving tertiary education in nigeria for development. washington d.c. i 1 ologbosere oluwatosin abiodun, lead city university, ibadan, nigeria, department of information management, can be reached at: ologbosere.oluwatosin@lcu.edu.ng. https://doi.org/10.29173/iq1082 http://ojs.library.ubc.ca/index.php/seealso/article/view/186%20335 mailto:ologbosere.oluwatosin@lcu.edu.ng vol244 6 iassist quarterly winter 2000 harmonising methods of disseminating urban heritage by alejandro delgado-gómez* introduction in cities with a long history, we can often find a high level of dispersion of the cultural heritage, as well as methods of preserving, describing and disseminating it. this paper explores a process to reach the harmonisation of documentation and retrieval of cultural heritage at a local level. although this process is valid in any context, we will focus our example on the most complex area that we have found: city planning. cartagena is a medium size city, with a long history starting from the carthaginian period. the city owns several collections connected to planning and landscape, although, as a sample, we will concentrate this presentation on the development of the city carried out between 1875 and 1934, known, at a local level, as the “ensanche.” during this period there was an effort to modernise and enrich the city with a strong modernist orientation, coincident with a period of industrialisation, promoted by foreign investors. in addition, nowadays the city is remodelling the old ensanche, in such a way that the council and private organisations are generating a significant volume of active records. the first remodelling, on the other hand, generated files, architectural drawings, archival documents, administrative regulations, as well as ancillary products, like bibliographic essays, photographs, paintings, etc. of course, the different quarters and buildings are the main product of that effort. therefore, we can discriminate, conventionally, two kinds of collections, “static” and “dynamic”: 1. dynamic records, still active, related to urban reforms in progress, as well as to the retrieval of carthaginian, roman and modernist architectural heritage. we are interested, at the moment, in this last period. 2. static different cultural collections, including archival files, plans, drawings, photographs, etc.; museological and bibliographic collections, ancient serials, some samples of paintings and other fine arts, etc. additionally, not all of this documentation is owned by the local government, but also by individuals, foundations, private companies, etc. of course, these collections have not been preserved and described in a consistent way over the years. as an obvious example, the active documentation is managed through the organisation of the information programs, and the closed documentation through more static databases. since cartagena is a growing city, oriented towards tourism and heritage retrieval, and, at the same time, with a population strongly involved in his cultural environment, one of the priorities of the council is the documentation, preservation, restoration and, mainly, dissemination of the history in a consistent way. the solution we show uses, as a pretext, a physical and digital exhibition of the urban history of the city harmonising in such a way preservation, description and dissemination of the above mentioned materials. we must notice, however, the fact that the cultural heritage of the city, as well as the urban planning, has been managed rather poorly for years. this implies a constraint in the harmonisation process. the first constraint is the chaotic situation that we found among the records. the second constraint is determining a way tore-arrange in some conventional way the information. these two tasks have to succeed before moving to more sophisticated techniques and procedures. contents and organisation of the information we have to work with two quite different kinds of documents, data repositories and, as a consequence, information: 1. the urban planning department is developing a “new city” and generating a great deal of new information. it is using a conventional computer supported co-operative workflow tool, based on microsoft products: visual basic, access and so on. however, because the department is divided into several offices, some of them are still using obsolete programs and tools, for instance, wordperfect 5.1 or 486 processors. because most of the staff doesn’t have strong computer skills, we cannot remove suddenly an iassist quarterly winter 2000 7 old, familiar system to implement a more adequate system. on the other hand, we are developing our work in co-operation with the data processing centre, but we are not responsible for the urban planning department. therefore, we have to work, simultaneously, in four different steps: a) the oldest databases. we have to get a correct migration from these to our archival system, and this implies the use of an intermediary program, friendly for the civil servants, but terribly annoying and certainly useless for us. b) the cscw. since this is a more updated system, we are using it, at the moment, like an intermediary step, with two aims: on the one hand, to train the staff in new uses of the technology; on the other, to migrate newly created records, in such a way we can minimise the “disturbing” effect of the oldest databases. c) development of a modeling system. we need to develop a system that is capable of structuring the information that the urban planning department is generating. the users do not want to know the modeling system. there needs to be an interface to the system. in such a way, we will migrate data according to our interests, avoiding a negative reaction on behalf of the department. to reach this end, we are using the well-known idef techniques, specifically idef0 and idef5, such as developed by kbsi. 8 iassist quarterly winter 2000 d) interface. we will develop a mask with the appearance of a “windows-based program” but actually independent of any platform. this seemingly strange mixture of procedures, techniques and tools allows us to allow for a structured migration of information with a minimal impact on the staff that input the information. it allows us to reconcile and harmonise this information with other databases. 2.the councillor of culture office, on the other hand, is promoting the retrieval, arrangement and dissemination of cultural heritage, in all its different facets, that is to say: museums, libraries, archives, buildings, fine arts, etc. since all of these are betterconsolidated areas, the situation is not as problematic as with the urban planning department. these institutions are using relational databases to describe their items, although not the same model of databases: access, fox pro, dbase, etc. it will be easier to implement changes here. these staff are skilled professionals who understand the concept of harmonisation and have strong computer skills. however, in spite of the co-operative staff, we have an additional problem. because of the relationships between urban planning and historical and cultural information there is a permanent flow and re-flow of documents and information. we are modeling cultural institutions in a similar way, in order to reconcile procedures and techniques. our aim is not to reach one homogeneous database, but to allow every professional to manage his or her own databases. however, all will follow similar procedures, which will make the retrieval of information easier for the professional, the intermediary and the enduser. all of this means that we have to deal at least with the following “instances”, and their relationships: a) active records, managed according three different means: old databases, based on ms-dos platforms and dbf files; recent databases, based on windows platforms and mdb files; and a mask to replace the former databases. b) non-active records and files, consisting of: -archival materials, since 1245 containing a highly dispersed range of information. -bibliographic materials, dealing mainly with the history of the urban development of the city. -ancient serials, newspapers and other newspaper materials, also with a high level of dispersion, as they reported the first modernist development of the city day-by-day. -cartographic materials and other plans, drawings and projects. -ancillary materials, such as photographs, paintings and other fine and decorative arts, etc. -buildings and other architectural items, ranging from squares, markets, façades, fountains, to parts of buildings such as doors or windows, which are protected by the regulations about cultural heritage. iassist quarterly winter 2000 9 since most of these materials are hosted by dissimilar kinds of institutions their records are collected and maintained differently. some of them are registered and described by means of word processors others by means of different relational databases. the highest level of homogeneity is being reached by public institutions, because of an agreement between the archives department, the data processing centre and a private company. in this way, we are solving at least one of the problems. both the archives, the libraries and the public museums are using only one programming language visual fox pro and only one associated kind of relational database. in addition, they are describing their materials according to a single structure the iso 2709 standard, making in this way the interchange of information easier. each type of repository uses the most adequate description standards: usmarc formats, isad(g)2, cidoc standards. this issue is irrelevant, since our interest is to obtain a homogeneous structure to interchange information, and an indexing and classification system capable of retrieving those information from disperse points. we are aware of the technological poverty of this solution. on the one hand, we think that, at least, it is realistic, and allow us to put a bit of order into the chaos. on the other, this solution allows us to use a distributed database, instead of disperse and heterogeneous data repositories. visual fox pro, merged with some other proprietary applications, allows us to manage the documents and the information in a robust and reasonably flexible way. of course, all of us hope this will be, also, a temporary solution. at an internal level, we are modelling, as we said, this information, following the idef techniques, and, at the same time, migrating data to conventional html files, in order to get a more homogeneous display. with regards to some other associated problems, the archives department is signing agreements with private institutions hosting materials, to manage them. 10 iassist quarterly winter 2000 finally, and since, as we said, this situation is provisional and cannot be sustained for a long time, we have been planning a technological process, currently in progress, to ensure a persistent harmonisation of materials and associated data and data repositories, through the use of modelling techniques and metadata languages. technological steps the following is a list of all the technological steps in this process. some of them are finished while others are still in progress. 1. digitisation of materials not digitised yet, or associated documentation in the case of archaeological and architectural items most of the significant materials for public institutions buildings, architectural items and drawings, archival documents regarding the first urban reforms have been digitised. this is not the case for private institutions. therefore, one priority is to digitise these materials. however, we have finished a complete union catalogue of the architectural heritage. 2. analysis of materials, in order to develop a model of classification and indexing, allowing a refined retrieval in a subsequent step. since one of our main interests is a sophisticated retrieval of the information, oriented to users’ needs, not strictly to contents, a detailed definition of the ontology, as well as its entities, attributes and elements is a sine qua non requirement. with regards to thematic indexation, we are using the oecd macrothesaurus, as it is simple, and allows quite a correct retrieval, taking into consideration contents are not our primary interest. a classification according to the users’ needs is more complicated, as it implies a market analysis and the use of statistical and psychological devices. at the moment, we are using a conventional solution, perhaps too easy, but useful: we are classifying the items according to some of the udc auxiliary tables, basically, that for people. in such a way we can retrieve information according to the users’ age, skills, education or planned use of the information. obviously, a combined search, by indexing terms and users’ classification is also possible. a more refined classification will have to wait. with regards to the analysis of the end-users, we have added a module to create and control statistical data. at the moment it is quite simple, but we are finishing a new and more sophisticated module, and a model of survey to be incorporated next academic year. iassist quarterly winter 2000 11 3. development of a generic relational database this will allow the professionals to enter any kind of basic data. at a second level will allow them to add any kind of relevant information. as we said above, we are able to customise these databases, by creating profiles according to the different kinds of professionals’ needs. 4. conversion to a generic metadata language. this was quite a problematic issue, basically because of the current “inflation” of specialised metadata languages. after reviewing carefully the state-of-the-art, we had to make a decision between two options: -to use a simple, basic, language, to interchange information. the most obvious example is dublin core; but we realised elements in dublin core were clearly insufficient to accommodate an exhaustive description, necessary in some cases. -to use specific metadata languages for each type of data repository. for instance, chio for museums, ead for archives, marc-dtd for libraries, tei for publishing departments, and so on; but this option was, simply, unmanageable, and, in addition, we ran the risk of returning to the initial chaotic situation. we took into consideration the use of only one specific language, in different contexts -maybe ilses or ead, but, even so, both of them are difficult to reconcile with active records description, that requires quite a qualitatively different method, such as, for instance, an adaptation of gils or ddi. finally, we refused these options, and chose an easier one. since cultural heritage databases are using a standardised language, marc, easy to convert, and even the active records are going to finish their lifecycle in the archives, we are working, simply, with xml and associated, and displaying the information by means of html and associated. 5. design of a search engine capable of filtering the information, not according to its contents, but according to the users’ interests. at the moment we have to work with conventional tools: the ifla guidelines to display information through opacs, and the ansi/niso z39.71 standard to display holdings. we can customise also this search engine, according to different users’ needs, as well as the printed reports. 12 iassist quarterly winter 2000 6. develop satisfactory “visual” outputs. we are dealing with a thematic digital library, with a large visual component. thus, we must also work on “visual” outputs, both to satisfy the end-users and the involved professionals, using ourselves vrml technologies, and allowing them to use authoring tools. at the moment, we have to be simple. each item has a description and associated with it, one or more “contextual files”, depending on their relevance: static images, sound, video or text. 7. if the mentioned steps are developed correctly and at the moment that is the situationwe will be able to elaborate any kind of digital or physical output website, intranet, kiosk, opac, dvd, webtv, printed materials...using consistent and conventional technologies. anyway, outputs are not a problem for us, if we can harmonise different databases in a distributed virtual database. in fact, even although we will have to fight still against the chaos for a long while, we are developing the website, the intranet and the opacs, based on the current achievements. these outputs, at the moment in progress, will replace the current tools at the users’ disposal: a poor website, a rather unfriendly opac, and a tv channel lacking information. we hope we will be able to put into operation some of the outputs by summer, and the project will be finished in no more than one year. * paper prepared for the iassist/ifdo conference, amsterdam 2001. alejandro delgado-gómez, cartagena city council-archives, publications, libraries and information science department. phone: 0034 968 128855, fax: 0034 968 128856. archivo@ayto-cartagena.es 26 iassist quarterly 2015 iassist quarterly improving the quality of digital preservation using metrics by mari kleemola1 abstract the finnish social science data archive (fsd) is dedicated to preserving digital research data in the social sciences and to providing good quality services to the research community. since the research data landscape is perpetually changing, digital repositories need to stay alert and be ready to review their policies and procedures, continuously. in fsd’s case, it is critical that our operations and procedures are up-to-date and consistent with relevant standards and best practices, and that our stakeholders trust us. in this paper we outline fsd’s venture into the world of digital preservation standards and assessments. we started by exploring the key standard for long term preservation of digital data, the open archival information system (oais). we then proceeded to conduct a self-assessment within the framework of the audit and certification of trustworthy digital repositories (tdr) checklist, followed by the cessda trust process which resulted in fsd’s successful application for the data seal of approval certification. the process has been fruitful and we can safely say that in the case of using metrics to improve quality, it is the journey that matters as much, if not more, than the destination.2 keywords: digital preservation, certification, oais, trusted digital repositories, assessment, dsa the finnish social science data archive fsd the finnish social science data archive (fsd) is a national resource centre that promotes open access to research data as well as transparency, accumulation and efficient reuse of scientific research data. fsd was established in 1999 to archive and disseminate digital, quantitative data for social science research, teaching and learning. over the years, it has broadened its services to include qualitative data archiving as well as provision of guidance on research ethics and data management. in many ways the turn of the century proved to be an excellent time to set up a national data archive. in the late 1990s the internet was already a daily working tool for social scientists, and standards like the data documentation initiative were emerging. most importantly, the international social science data archiving community was well established and networked, having become a part of the research scene in the 1950s and 1960s. the various social science data archives were actually the very first institutions to handle and preserve digital material (doorn and tjalsma 2007). our new archive was privileged to learn from the practices and experiences of the pioneers of digital data archiving. early on we realised that preservation, management and dissemination of datasets as well as building high-quality knowledge-based services needed to be done in a transparent way, at the most appropriate time possible and without over-burdening researchers. it was also clear that for the outcome to be successful, many organisational issues and pieces needed to be in place, including policies, procedures and sustainable resources. consequently, we paid attention to documenting our processes and procedures from the iassist quarterly 2015 27 iassist quarterly very beginning: our first internal handbook was created during the first two years of operation and our first archives formation plan, or ams, emerged in 2003. the ams is fsd’s highest-ranking document concerning provision of data services. it is based on guidelines set by the national archives of finland and it describes tasks and processes, selection criteria, preservation periods and forms, confidentiality and legislative issues, responsibilities, and data systems and security. nowadays, fsd is an acknowledged and active national centre of expertise in the areas of preserving and providing access to digital research data. fsd’s core user community consists of researchers, teachers and students from finland and abroad. most fsd’s services, such as the data management guide, are openly and freely available via our website and in 2014 the number of successful web page requests reached 1.2 million. access to fsd’s data holdings is provided via the aila data portal. in june 2015, aila contained 1200 datasets and had 1300 registered users. fsd is funded by the ministry of education and culture and operates as a separate unit at the university of tampere. fsd is finland’s service provider for the pan-european research infrastructure, cessda3. measuring trustworthiness of digital preservation hedstrom (1998) defines digital preservation as ‘the planning, resource allocation, and application of preservation methods and technologies necessary to ensure that digital information of continuing value remains accessible and usable’. digital preservation is therefore an active, and in the optimal case, even a pro-active process that does not include extended periods of inactivity. digital preservation is also something that needs to be done presently for the future; we need to think about a time period long enough to be concerned with the impacts of changing technologies or a changing user community (ccsds 2009; giaretta 2011). the demands these definitions make are powerful and illustrate well the difficulties in providing long-term access to digital data. preserving digital data is a challenge. for the outcome to be successful, many organisational and practical issues need to be in place. key aspects in demonstrating trustworthiness include: transparency, documentation, adequacy and information security. in addition, one has to keep in mind that one needs to evaluate trust into the future. (dobratz et al. 2010; giaretta 2011.) for quite some time it has been evident that methods are needed to assess the trustworthiness of a digital repository. in 1996, the commission on preservation and access and the research libraries group called for a certification programme for repositories claiming to serve an archival function, and since then several stakeholders have explored certification issues (see dobratz et al. 2010). the need to be able to test a repository’s claims about digital preservation was also one of the key drivers of the open archival information system (oais) reference model (giaretta 2011, 461). in the european data archive world assessment and trust issues came up roughly ten years ago. in 2006, the council of european social science data archives (cessda) research infrastructure was identified as an existing pan-european ri recommended for a major upgrade by the european strategy forum on research infrastructures roadmap. as a direct result of this, cessda launched the preparatory phase project (ppp) in 2008. the project recommended, amongst other things, that adherence to standards should be a part of cessda eric’s membership criteria since common use of standards is necessary for compatibility. as possible tools, the project brought forward the oais model in particular, and the dsa criteria. (dusa et al. 2010.) the oais model (iso14721:2003) is the international key standard for archival systems and for organisations engaged in long-term preservation of digital data (see, for example, spence 2006; lavoie 2004; giaretta 2011). prior to the cessda ppp project, the oais model had been explored by the uk data archive (beedham et al. 2005) and the icpsr (vardigan and whiteman 2007). a finnish version of the standard was published in march 2010 by the finnish standards association sfs4. the oais model introduced important concepts and paved the way to a certification standard for digital repositories. oais conformance is necessary for trustworthiness but not sufficient since the oais model is very general and does not cover, for example, financial aspects. metrics were therefore needed. the trustworthy repositories audit and certification checklist (trac) was released in 2007, and its revised version, entitled the trusted digital repository (tdr) checklist, in 2011. the checklist derives from the oais, and the iso 16363:2012 standard5 is based on it. the first edition of the data seal of approval, dsa, emerged in 2008. it was initially developed for use in the netherlands, but already in 2009 an international dsa board was established. the dsa contains 16 guidelines, and it is the first step in the three-level european framework for audit and certification of digital repositories6 that was developed in 2010. other certification initiatives include, for example, the nestor catalogue of criteria (din 31644) and the drambora toolkit (for an overview of key guidelines and frameworks, see kvalheim et al. 2013). all these developments prompted us to take a critical look at our functions at fsd. we wanted to find out if they were adequate and up-to-date and hoped to find ways to rationalise our processes. we also wished to demonstrate that we can be trusted by our key stakeholders: data depositors, users and funders. furthermore, an important goal was to ascertain that fsd would able to fulfill the cessda membership criteria. to achieve all this, we decided to use suitable standards and metrics to assess our systems, policies and procedures. step one: start by oais our first step was to familiarise ourselves with the oais model in 2010. an oais compliant organisation has to support the information model of the standard and fulfil the six minimum requirements. thus, the organisation must: • negotiate for and accept appropriate information from information producers, • obtain sufficient control of the information provided to the level needed to ensure long-term preservation; • determine, either by itself or in conjunction with other parties, which communities should become the designated community and, therefore, should be able to understand the information provided; • ensure that the information to be preserved is independently understandable to the designated community; • follow documented policies and procedures which ensure that the information is preserved against all reasonable contingencies, and which enable the information to be disseminated as 28 iassist quarterly 2015 iassist quarterly authenticated copies of the original, or as traceable to the original; and • make the preserved information available to the designated community. (ccsds 2009.) fulfilling these general requirements proved straightforward. fsd negotiates with researchers when they are depositing their data, and fsd’s designated community is stated in the regulations as social science researchers, teachers and students. fsd’s archives formation plan (ams) contains specific information on fsd’s tasks (including selection criteria, archival process and data protection practices) and the internal manual contains detailed practical instructions. archived research data are processed and described in compliance with international standards and formats, and the research community is able to find information about the data on the fsd website. the oais standard also contains a functional model consisting of six main entities: ingest, archival storage, data management, administration, preservation planning, and access. additionally, an organisation has to provide common services, such as various technical support services. exploring these functional entities in detail provided us with a new perspective about our processes. all the entities were identifiable in fsd’s operations and luckily we were not able to find any serious defects. this in turn resulted in a strengthened conception that we had been doing the right things in the right way. however, it became clear that our processes could be further clarified and improved and that we should, for example, collect and store more extensive preservation information and to better structure it. the oais compliance was expected for many reasons. first, fsd has been established to function as a repository for digital data, and also fsd’s operational model has been based on the examples of well-established social science data archives. all in all, fsd emphasises functions slightly differently in comparison to the oais model. while the oais describes ingest only briefly and practically omits acquisition, they are central processes in the fsd. on the other hand, the oais presentation of administration and preservation planning is more complicated than fsd’s procedures, mainly because fsd is still a relatively small archive. as fsd grows, it must be prepared to adopt more sophisticated administrative processes. our experiences about oais were very similar to those reported by other data archives, for example by the uk data archive (beedham et al. 2005) and the icpsr (vardigan and whiteman 2007). also gesis has carried out a mapping of the oais functional model to gesis’s operations (schumann and recker 2012). step two: tdr checklist the results of our oais exercise were encouraging. however, the oais model is very high-level and thus did not provide enough concrete details to support the assessment of our day-to-day practices, which we thought would help us ensure the quality of our practices and processes and consequently the quality of our data services. therefore, in spring 2012, we continued our assessment exercise with the help of the tdr checklist. this decision was influenced by the uk data archive’s draft audit against the emerging iso 16363 standard (see woollard, 2011). at the time of our self-assessment, the tdr checklist was at its final stages of approval as iso 16363 and was available as ccsds 652.0m-1 magenta book (ccsds 2011). we used both the magenta book and the preliminary version of the tdr checklist in excel format provided by the center for research libraries. the checklist is designed to cover all the aspects necessary to demonstrate that an organisation can be trusted by its stakeholders. the main sections include organisational infrastructure, digital object management and infrastructure, and security risk management. for each of the 100+ tdr clauses, we identified and described the written evidence and assessed our compliance using the scale of 0 (not compliant) to 3 (fully compliant). we found fsd to be well compliant (3) or almost compliant (2) with 66% of the clauses and not compliant (0) or only somewhat compliant (1) with 10% of the clauses. six clauses were thought to be out of scope for fsd. in 18% of the clauses we could not make a straightforward assessment because we did not understand them completely. this is probably at least partly due to the fact that in many cases fsd relies on a manual process and a set of systems whereas the tdr checklist describes a larger and more automated system. our estimation is that fsd would be not compliant (0) or only somewhat compliant (1) with most of the “unclear” criteria. however, not all criteria have the same importance or level of risk. examples of fsd’s full compliance: • mission statement exists • collection policy exists • preservation policies exists • short and long term business planning in place • appropriate deposit agreements exist • minimum information requirements specified • minimum descriptive information captured examples of fsd’s non-compliance: • no formal succession plan • no commitment to regular self-assessment or external certification • no documented process for testing for an understanding of the aip content information • no systematic analysis of security risk factors the good news was that we found no major failures or risks. however, the analysis revealed several weak points. for example, early on in the process it became evident that fsd needed to clarify several processes and to especially improve the documentation that describes the technical infrastructure and security risk management. as a result of the self-audit, fsd made an action plan for various changes and improvements. some minor adjustments were made immediately and major changes became subject to discussions and have been – or will be – implemented gradually. our motivation for the self-assessment was to improve our practices and to plan for the development of our processes. for this, the tdr checklist worked very well, although the learning curve was rather steep and understanding what constitutes adequate evidence was sometimes difficult. the self-assessment was also rather time-consuming and therefore we decided not to go too deep into the clauses and evidence. since the goal of this tdr self-assessment was not a certification we did not analyze in detail which of the non-conformities would be acceptable and which were not. our checklist assessment is also in draft state (for example, some texts are in english, some in finnish) and while it is sufficient for internal use, it has not been published. so even though the self-assessment provided us with a lot of insight into iassist quarterly 2015 29 iassist quarterly our processes and helped to recognise and correct deficiencies, due to the lack of transparency it is not very useful for building trust with our key stakeholders. step three: cessda trust process in 2013, cessda archives carried out a two-phase self-assessment exercise also known as the “cessda trust process” that aimed to progress the cessda archives in the area of trusted digital repository status. as cessda is on its way to becoming an european research infrastructure (eric), all its service providers need to meet the membership criteria. in the cessda trust process, the data seal of approval (dsa) was selected as a reference point because it is the base-level in the european framework for audit and certification of digital repositories. the 16 guidelines of the dsa allow for verifying quality aspects concerning the creation, storage, use and reuse of digital data, and they can be seen as a minimum set distilled from proposals like nestor, drambora and tdr. the guidelines focus on three stakeholder groups: the data producers, the data repositories, and the data consumers. a repository is designated a trusted digital repository if it complies to the ten guidelines for repositories and if it enables data producers and data consumers to comply with their three guidelines (data seal of approval 2014). at the beginning of the cessda trust process, the dsa criteria were mapped to the cessda member obligations. after the mapping, all the european data archives undertook a dsa selfaudit. the process was coordinated by an expert panel consisting of four people, representing the british uk data archive, the dutch dans, the german gesis and the fsd. the expert panel counselled the archives during the self-audit, reviewed all self-assessments and provided feedback and a gap analysis. all in all, european data archives appeared to have good practices in terms of, for instance, data reusability. long-term preservation of data is also well managed, although the documentation of practices was inadequate across many data archives. during the cessda trust process, fsd gained valuable knowledge about the dsa as well as of certification in general. discussions with other data archives clarified the meaning of the guidelines and helped to understand and identify the relevant evidence ie. the documentation that was needed to demonstrate that we meet the dsa criteria. feedback from the self-assessment and reviews led us to further improve our documentation. step four: dsa application the cessda trust process confirmed that fsd was in a good position to apply for the dsa. in may 2014 we implemented major changes in our data services that took us from paper-based data ordering to an advanced online data download system, so we decided to postpone our dsa application until august 2014. this ensured that we had time to revise both our internal and external documentation to reflect the changed service model. before we sent our application, we also translated some key documents, like our archives formation plan (ams), into english. our actual dsa application process was very straightforward and not time-consuming. it took only a day. this somewhat contrasts with the gesis’ experience (schumann 2012) and reflects the fact that we already had many pieces in place following the oais, tdr and cessda trust exercises. for example, we had all the necessary documentation readily available to support our assertions, most of it in both finnish and english. it is also worth mentioning that fsd had perceived good documentation as a cornerstone of its operations from the beginning so our documentation is the result of years of collective, continuous, and innovative work by fsd experts. our documentation consists of three main components: our website, the archives formation plan (ams) and the internal manual. fsd’s website contains detailed information on archived data, archiving and ordering data, and on managing research data. the archives formation plan includes, for example, our preservation plan, ingest criteria, and data protection practices, as well as the internal manual contains very detailed practical guidelines for all processes. the extensive and well-maintained documentation ensures, for example, that if any two data managers would process the same data according to the instructions, the resulting archival information packages would be substantially similar. on september 23rd, 2014 we were awarded the 2013 dsa certificate7, and are planning to renew it regularly since we view it as a tool that can be used in preparing for the changes and challenges we will inevitably face. fsd was the first finnish organisation to acquire the dsa and we have received several inquiries about our certification. we have been very pleased that, for its own part, fsd’s trust process has raised awareness about trust and long-term preservation of digital data in finland. conclusion over the last few years, fsd has successfully used metrics to improve the quality of its processes and operations. as was to be expected, fsd conforms to the oais model. the tdr exercise confirmed that fsd is doing things in the right way, the cessda trust process allowed us to compare fsd with other archives, and the dsa certification increased trust with stakeholders. the use of models and metrics to assess our procedures and policies have raised our awareness about the challenges of digital preservation, revealed existing and possible problems and weaknesses as well as strengths, steered and initiated minor and major changes in our operations, and resulted in improved documentation. as a consequence, many of our processes are now better and more efficient or, they will be better – some of the bigger changes will take time to implement. we are also able to better manage risks, provide more trustworthy services for the research community, and demonstrate fsd’s trustworthiness to our stakeholders. in addition, we are in a good position to meet the cessda membership criteria. our journey into the world of standards and measurement has so far taken four years, and we see metrics such as the tdr and dsa as tools to be used continually for improving processes and services. the actual intensive work related to oais, tdr and dsa has required in total about six person months’ worth of effort which we see as resources well spent. we are also considering taking the next step in the european framework for audit and certification, which is a structured, externally reviewed and publicly available self-audit based on iso 16363 or din 31644. the research landscape and thus the research data landscape is changing rapidly. as a digital repository we need to stay alert, be ready to review our policies and procedures methodically and critically, and build and share our competence continuously. 30 iassist quarterly 2015 iassist quarterly funding note in order to be successful in any project, commitment from management as well as sufficient resources are needed. the work described in this paper has been part of the fsd upgrade and veric projects8, both funded by the academy of finland and both aimed at strengthening the finnish national service provision as part of the cessda eric process. references beedham h, missen j, palmer m and ruusalepp r (2005). assessment of ukda and tna compliance with oais and mets standards. joint information systems committee (jisc), united kingdom. available at: http://www.esds.ac.uk/news/publications/oaismets.pdf [accessed 4.6.2015] ccsds (2011). audit and certification of trustworthy digital repositories. magenta book. issue 1. september 2011. available at: ccsds 652.0-m-1. http://public.ccsds.org/publications/ archive/652x0m1.pdf [accessed 10.4.2012] ccsds (2009). reference model for an open archival information system (oais) draft recommended standard. ccsds 650.0-p-1.1 (pink book), issue 1.1, august 2009. available at: http://public.ccsds. org/sites/cwe/rids/lists/ccsds%206500p11/attachments/650x0p11. pdf [accessed 17.2.2010] center for research libraries (2012). tdr checklist in excel format (preliminary version). available at: http://www.crl.edu/sites/default/ files/attachments/pages/tdr_checklist_self_audit.xls [accessed 10.4.2012] data seal of approval (2014). dsa overview article. available at: http:// datasealofapproval.org/media/filer_public/2014/10/03/20141003_ dsa_overview_defweb.pdf [accessed 4.6.2015] dobratz, susanne, peter rödig, uwe m. borghoff, björn rätzke and astrid schoger (2010). the use of quality management standards in trustworthy digital archives. international journal of digital curation 2010, vol. 5, no. 1, pp. 46-63. doi:10.2218/ijdc.v5i1.143 doorn, peter and heiko tjalsma (2007). introduction: archiving research data. archival science 7(1), 1-20. doi:10.1007/s10502-007-9054-6 dusa a, krejčí j, štebe j, fábián z, hegedus p, hausstein b (2010). wp6 final report: strenghtening the cessda ri (d6.1). cessda ppp. available at: http://ppp.cessda.net/doc/wp6_final_report.pdf [accessed 4.6.2015] giaretta, david (2011). advanced digital preservation. heidelberg: springer-verlag berlin. hedstrom, margaret (1998). digital preservation: a time bomb for digital libraries. computers and the humanities 31(3): 189-202. doi 10.1023/a:1000676723815 kvalheim v, kiberg d, kvamme t, balster e, de bruijne m, wijnant a, recker a, lenkiewicz p, widdop s, and wloka b (2013). roadmap for preservation and curation in the ssh. dasish report d4.1. available at: http://dasish.eu/publications/projectreports/d4.1_-_roadmap_ for_preservation_and_curation_in_the_ssh.pdf/ [accessed 4.6.2015] schumann, natascha and recker, astrid (2012). de-mystifying oais compliance: benefits and challenges of mapping the oais reference model to the gesis data archive. iassist quarterly summer 2012, 6-11. available at: http://www.iassistdata.org/downloads/iqvol36_2_ recker_1.pdf [accessed 4.6.2015] schumann, natascha (2012). tried and trusted. experiences with certification processes at the gesis data archive. iassist quarterly fall winter 2012, 23-27. available at: http://www.iassistdata.org/ downloads/iqvol36_34_schumann_0.pdf [accessed 4.6.2015] vardigan, mary and whiteman, cole (2007). icpsr meets oais: applying the oais reference model to the social science archive context. archival science 7:73-87. available at: http://deepblue.lib.umich.edu/ bitstream/2027.42/60440/1/vardigan.whiteman.applying%20oais. pdf [accessed 4.6.201] woollard, matthew (2011). standards-based approach to preservation planning. presentation at the aligning national approaches to digital preservation conference, tallinn, estonia, 23-25 may 2011. available at: http://www.data-archive.ac.uk/media/276922/ woollard_20110524.pdf [accessed 4.6.2015] notes 1. mari kleemola is information services manager at the finnish social science data archive. she can be reached by email: mari.kleemola@ uta.fi. 2. this paper is an updated version of a presentation given at iassist 2012: kleemola, mari (2012). improving operations using standards and metrics: self-assessment of long-term preservation practices at fsd. a presentation at the 38th annual iassist conference, washington dc, june 06, 2012. available at: http://www.iassistdata. org/conferences/2012/presentation/3328 3. consortium of european social science data archives, http://www. cessda.net 4. the finnish version is called sfs 5972 viitemalli pitkäaikaissäilytysarkistolle. 5. ccsds 652.0-m-1 -audit and certification of trustworthy digital repositories (september 2011) contains the final draft standard submitted to iso for review and approval and is freely available from the ccsds website as a recommend practice document: http:// public.ccsds.org/publications/archive/652x0m1.pdf [3.6.2015] 6. http://www.trusteddigitalrepository.eu/ 7. https://assessment.datasealofapproval.org/assessment_109/seal/ pdf/ [4.6.2015] 8. more information about fsd’s projects: http://www.fsd.uta.fi/en/ news/projects.html 1/20 hayslett, michele & jansen, matthew (2022) factors contributing to repository success in recruiting data deposits, iassist quarterly 46(2), pp. 1-20. doi: https://doi.org/10.29173/iq1037 factors contributing to repository success in recruiting data deposits michele hayslett & matthew jansen1 abstract what factors make data repositories successful in recruiting research data deposits from scholars? while quite a few studies outline researchers’ data management needs and how repositories can meet those needs, few have assessed the success of various approaches. this study examines infrastructure for accepting data into repositories and identifies factors influential in recruiting data deposits. keywords data repositories, data deposits, deposit recruiting, best practices, marketing introduction throughout the early 2000s, librarians have worked to build repository infrastructure, deposit workflows, access features, and support services to meet the needs of their audiences. needs assessments and evaluation of research practices often informed the creation of repository features and services, and descriptions of these design efforts abound in the literature. however, few postdeposit assessments have been published to identify the most successful recruitment practices for datasets. moreover, the content of many institutional repositories (irs) is focused on textual deposits such as pre-prints and publications. fewer repositories enable the preservation of datasets, and the literature yields very few assessments of data repositories. many questions can be asked about the work to recruit dataset deposits in repositories. what features and services have proven to be the most useful and attractive to researchers? staffing, both overall and specifically related to data deposits, may logically have significant positive correlations with larger numbers of depositors, but does choice of marketing media likewise have a high association? this project addresses the over-arching research question: what factors are most associated with larger numbers of data depositors? literature review as the field of data management has grown, a great deal of the literature has detailed the overall benefits of open data and more specifically the advantages an individual scholar would accrue from sharing their research. many articles also focus on ascertaining investigators’ needs and how to design repositories to meet those needs. these two pools of research overlap and offer repositories formative information with which to plan recruiting strategies for research materials generally, including data. little is available in the literature, though, about the next step in the process: evaluation of how successful those early needs assessments and system designs have been, and what other factors are positively correlated with larger numbers of depositors. a variety of searches of the international journal of digital curation between october and november 2021 (‘evaluation success’; ‘marketing’; ‘effects deposit rate’; ‘factors affecting deposit rates’) yielded only one result related to the evaluation of repositories broadly. mchugh et al. (2008), describes the digital repository audit method based on risk assessment (drambora), a flexible framework for self-audit to assess ‘demonstrable, and not just inferred, success (p. 135)’ but does not actually discuss results of any specific assessment. while self-audit can provide valuable information for repositories and its results can be kept private, wider sharing of assessment metrics may yield valuable information for the broader community. this study seeks to address this gap. https://doi.org/10.29173/iq1037 2/20 hayslett, michele & jansen, matthew (2022) factors contributing to repository success in recruiting data deposits, iassist quarterly 46(2), pp. 1-20. doi: https://doi.org/10.29173/iq1037 data sharing and incentives many articles have reported on the benefits of sharing data, both as a public good that improves reproducibility and the overall quality of research, and for individual researchers, raising the profile of their work and increasing their impact (gardner et al., 2003; national research council, 1985; pienta, alter & lyle, 2010; piwowar, day & fridsma, 2007). however, numerous surveys have found that while researchers often declared willingness to share their data, many barriers obstruct them from actually sharing (fecher, friesike & hebing, 2015; gardner et al., 2003; lowenberg, 2017; national research council, 1985; tenopir et al., 2011; wallis, rolando & borgman, 2013). where compliance serves as an incentive for data deposit, researchers may sometimes only encounter a funder’s compliance requirement when they apply for grant funding. faniel and connaway (2018) note that several librarians in their survey found their data management services being contacted just prior to grant proposal deadlines. this matches the authors’ own experience with many researchers seeking help writing data management plans only days before a proposal deadline. irs may be better positioned than some other repositories to assist researchers in this last-minute way, with low barriers to deposit. it might then follow that researchers find irs more important for preserving data than articles. bryant, lavoie, and malpas (2017) note, ‘given the extensive network of discipline-, consortialand national-scale [research data management (rdm)] services, many institutions have scoped their local rdm service bundles to be complementary to, rather than parallel with, these external options’ (p. 30). hudson-vitale et al. (2017) found research libraries that provide data curation services viewed providing a persistent identifier as the most important data curation activity overall, but it remains to be seen if researchers agree with this. the authors have encountered quite a few depositors who did want to deposit their data specifically to obtain the persistent identifier assigned by the ir to insert in an upcoming publication based on those data. but what are repositories broadly (not just irs) experiencing? what other services are repositories finding to be in demand? does offering more advanced services like file or code review, encryption, and direct deposit via electronic lab notebooks (elns) correspond with having more depositors? needs assessments and repository design quite a few needs assessments and pilot projects for repositories have been reported, covering everything from the benefits for depositors to repository technical infrastructure to metadata creation services (abrams et al., 2014; burton & treloar, 2009; hudson-vitale et al., 2017; mattern, jeng, he, lyon, & brenner, 2015). many note special challenges associated with archiving data. as early as 2008, salo observed in her ‘repository as a roach motel’ article that the ‘build it and they will come’ approach to repository collection development with which many institutions started had not been successful and argued that different incentives than commonly cited ones like preservation and heightened impact were needed. plale et al. (2013) noted difficulties publications-based repositories can encounter in archiving data and explored how those challenges can be addressed through requirements, policy, and architecture. minor et al. (2014) describe how pilot ingests of datasets directed the design and implementation of the uc san diego library’s research data curation program. borgman et al. (2016) explored cloudbased services as a data management solution for ‘long-tail’ research projects, that is, smaller projects with few resources for documenting and preserving their data. while these services were found to be useful for some basic tasks, they were insufficient for more complex needs such as development of specialized data tools and long-term preservation, needs that can be addressed in repositories. peer and green (2012) describe how the push to make more research open access is reflected in yale’s efforts to host an open access repository on open-source software with the goal of supporting research replication as well as re-use and instruction. few of these case studies have published followhttps://doi.org/10.29173/iq1037 3/20 hayslett, michele & jansen, matthew (2022) factors contributing to repository success in recruiting data deposits, iassist quarterly 46(2), pp. 1-20. doi: https://doi.org/10.29173/iq1037 up reports on their particular results (although peer and green do note positive early feedback and outcomes from users), but there are some studies of outcomes generally. tillman (2017) surveyed self-deposit rates at 55 u.s. irs and concluded, ‘…the short answer to the question ‘is our faculty depositing?’ is ‘not really,’ or the even more straightforward ‘no.’…everyone is trying. few are succeeding’ (p.13). but tillman also evaluated other variables’ correlation with high deposit rates: the age of the ir; the software on which the repository is built; and the outreach methods of the ir. (she cites several reasons why the software would be important, both from the standpoint of a faculty member’s willingness to use [and re-use] the ir, and as an indicator of the administration’s investment in the ir.) her results are particularly revealing of the difficulties repositories face in achieving success. with regard to the age of the ir, she notes: it appears to take a minimum of two years on average for repositories to have even a 50% likelihood of getting at least a single self-deposit per month, with a much greater likelihood of success after five years. this time period allows for the ir to become an established entity on campus and for responsible parties to do a variety of outreach and build new strategies after failures in the > 2-year and 2–5 year ranges. however, as 66.67% of repositories that had existed for at least five years still had self-deposit rates of 20 items or fewer and a full 25% had 0 average monthly self-deposits, the comparatively positive correlation of success and age should not be considered a guarantee. age correlates even more strongly with failure. (p. 14) the results about software seemed to indicate it was important to high self-deposit rates but not conclusively why that was so: institutions reporting higher rates of self-deposit are more likely to report the use of either a homegrown system or one that involves high levels of developer time and engagement. this factor may indicate that the institution is investing heavily in the repository, implying that it allots more staff time and effort for other activities that promote deposit. it may also indicate that the user interface for deposit in turnkey models does not promote self-deposit. (p. 14) finally, she notes that reports on irs’ outreach are also correlated ‘strongly with both a strong deposit profile and the lack thereof. no conclusions, therefore, can be drawn about either in general (p. 14).’ tillman did not differentiate between publications and data deposits, though, and focused on selfdeposit rates. this study focuses on data, in whatever way they are deposited. marketing over time, as research norms change, more researchers will begin to deposit their data if such practices become the accepted cultural norm. outreach is one way for libraries to reinforce such a cultural shift. bryant, lavoie, and malpas (2018) note: there is an evangelistic aspect to…[educational data management] outreach—although researchers may not be ready to deposit data at the time the outreach occurs, they are at least made aware that data management services are in place to support them when needed…successful outreach program, can over time, cultivate the demand that will help establish [rdm] as a critical piece of scholarly infrastructure (p. 16). indeed, many universities, especially in europe, are targeting graduate students, seeding cultural change in the next generation of researchers. the university of groningen (2013) requires its doctoral students to make their data ‘available for further research,’ although exemptions are possible for ‘compelling reasons’ (p. 13). to propose a similar directive be implemented at the delft university of https://doi.org/10.29173/iq1037 4/20 hayslett, michele & jansen, matthew (2022) factors contributing to repository success in recruiting data deposits, iassist quarterly 46(2), pp. 1-20. doi: https://doi.org/10.29173/iq1037 technology (tu delft), dunning (2017) compiled policies from groningen and seven other universities in the u.k. and the netherlands: utrecht, leiden, twente, bristol, southampton, bath, and manchester. most of these stated their policies as an expectation, without making deposit mandatory. tu delft decided their doctoral students ‘…starting from 1 jan 2019 will have to share data unless they have a compelling reason not to’ (a. dunning, personal communication, may 9, 2019). their policy states the expectation of this covering ‘all data and code underlying completed phd theses,’ and that they be ‘appropriately documented and accessible for at least 10 years from the end of the research project …’ (p. 7, tu delft, 2018). many academic repositories’ outreach efforts also aim to change faculty attitudes, of course. otto (2016) describes the outreach message rutgers used with the ‘…primary objective…to fully inform faculty so that they were motivated to make the open access choice …’ (p. 11). she goes so far as to quote thomas jefferson in his description of the declaration of independence, that their objective was ‘to present [to the ‘tribunal of the world’] ‘the common sense of the subject, in terms so plain and firm as to command their assent.’ (jefferson, 1825)’ (p. 11). she concedes, though, ‘…evidence of its own efficacy remains, for the most part, unavoidably anecdotal’ (p. 4). several projects are assembling lists of european institutions that have established policies, and many of these apply to all researchers, not only graduate students. laurence horton with the london school of economics collaborated with the u.k.’s digital curation centre to compile the 'overview of uk institution rdm policies' web site (m. donnelly, personal communication, june 28, 2019). as of june 21, 2022, it included 86 institutions plus one that was drafting a policy (although none are dated past 2016, this number has changed during the writing of this paper and some policies are undated). kerstin helbig at humboldtuniversität zu berlin noted their web site listing with links to dozens of german institutions’ policies in a message to the research-dataman [listserv, a research data management] list (k. helbig, personal communication, june 28, 2019). and fairsharing.org, a manually curated educational resource that describes and captures the relationships between standards, databases and data policies, indicated it would ‘soon be inviting submissions to [its] data policy registry for these types of institutional research data policies’ (p. mcquilton, personal communication, june 28, 2019)—it included 155 such policies as of june 21, 2022; however it does not appear to allow filtering by type of organization, e.g., ir, disciplinary repository, publisher, professional or research association, or other. few examples in the literature have evaluated social media as a marketing method to reach potential depositors, though. boulton (2020) reviewed the literature around institutional repositories and engagement and found that, broadly, the repository community focused on improving systems and structures to make use of their repositories easier rather than on evaluating direct outreach to their audiences. given that void, that paper turned to examining the social media practices of irs but did not detail exactly how, nor how many irs were examined, merely noting, ‘engagement through social media channels such as facebook and twitter appears common across institutional repositories …’ (p. 1). this practitioner’s aim was to go beyond using a single social media channel and describe instead a coordinated campaign using multiple media at griffith university in australia. the case study analyzed the traffic driven by two blog posts (covering the research backstories and results of 15 deposited articles) that were promoted in tweets from the library’s twitter account and concluded that ‘…social media and blog posts could be used by the library to increase engagement with an external audience and to drive traffic to the repository’ (p. 4). in other words, the particular aim of this experiment was to increase research impact rather than to increase depositors but the same strategy could potentially achieve both ends. certainly, the griffith repository succeeded in highlighting the articles: during the two months of the campaign, ‘... the 15 featured articles were accessed close to 500 times, this being approximately 60% of their combined total for the previous six months …’ (p 3). https://doi.org/10.29173/iq1037 https://www.dcc.ac.uk/guidance/policy/institutional-data-policies https://fairsharing.org/search?fairsharingregistry=policy 5/20 hayslett, michele & jansen, matthew (2022) factors contributing to repository success in recruiting data deposits, iassist quarterly 46(2), pp. 1-20. doi: https://doi.org/10.29173/iq1037 lafferty-hess et al. (2018) discuss how their thought exercise of conceptually grouping the dcn’s data curation activities helped them not only identify appropriate services for their respective institutions given different available resources, but also to consider communication strategies to make their services clear to researchers, and strategies for measuring their success. it is the authors’ hope that the current study will also aid repositories in these important ways. methodology the population under examination in this study was north american data repositories, whether open or not, whether independent or based in an institution. (note: respondents were not asked to identify the type of repository in which they worked. also, the terms “researchers” and “faculty” are used interchangeably in this paper.) the focus was on those that accept datasets. while brief information was invited from repositories that do not, few non-data repositories participated. the appeal for survey participation (shown in appendix a) was sent to repository professionals via three email lists heavily populated by data curation professionals: the membership list of the international association for social science information services and technology (iassist); that of the u.s.-based research data access and preservation (rdap) association; and datacure, a list established by geographically scattered participants in the digital curation curriculum (digccurr, pronounced “dij-seeker”) at the school of information and library science (sils) at unc-chapel hill after returning to their home institutions, and whose membership has grown rapidly since. no organization is currently maintaining a directory of all repositories (opendoar offers a directory only of open access repositories) so this approach was deemed best. the initial appeal was sent to all three lists in august 2019, with two follow up reminders sent at the threeand five-week marks. to give an idea of the audience reached by the appeal, list membership at the time of the survey stood as follows: datacure – 231; iassist – 539; and rdap – 557, for an estimated total of 1,327 invitations. however, there is some overlap among the three organizations’ membership, and iassist includes members from outside of north america, so the actual number of eligible participants reached is unknown. data were collected by an online qualtrics survey, which was pretested by several volunteers from the 2019 rdap conference. the consent form was presented as the survey’s first page (see appendix b for the consent form and appendix c for the full survey instrument). the survey solicitation and the survey instrument specified that one response per repository was requested, and the survey instrument was structured to allow response over multiple sessions to encourage input from multiple individuals as necessary to construct a full picture of each repository’s characteristics. response was requested within a six-week window, by september 30. the survey instrument was left available for some weeks after that deadline in case of late responses; the last response was recorded on october 10. the number of respondents was small, impacting what conclusions may be drawn from these results: 31 submissions overall, with two who did not accept data and one who accepted data but did not provide their number of depositors. because of non-responses in number of depositors and other variables of interest, the sample size varied between 26 and 28 (see appendix e for specific results by hypothesis). still, the data were sufficient to identify multiple trends. survey respondents were invited to indicate their willingness to participate in follow-up interviews but very few did so; consequently, plans for follow-up interviews were abandoned. results eight hypotheses were tested for this paper: 1. repositories with the highest staffing will have a significant positive correlation with larger numbers of depositors. https://doi.org/10.29173/iq1037 6/20 hayslett, michele & jansen, matthew (2022) factors contributing to repository success in recruiting data deposits, iassist quarterly 46(2), pp. 1-20. doi: https://doi.org/10.29173/iq1037 2. larger numbers of depositors will be significantly correlated with offering advanced curation services. 3. repositories with larger numbers of depositors will be those which have depositors referred by faculty versus referrals from other places. 4. larger numbers of depositors will be significantly correlated with using social media to promote the repository. 5. larger numbers of depositors will not be significantly positively correlated with infrastructure except where repository ingest is linked to electronic lab notebooks (elns). 6. larger numbers of depositors will be significantly positively correlated with repositories that encrypt data. 7. larger numbers of depositors will be significantly positively correlated with having more staff dedicated specifically to data deposits. 8. larger numbers of depositors will be significantly positively correlated with older repositories. all tests were performed at a 95% confidence level. because of the small response rate, the authors focused on measuring relationships between just two variables at a time. this approach enabled rejection of a hypothesis but not assertion of causation nor indication of direction, i.e., which variable might have caused the other. the number of depositors as tested for each hypothesis does not follow a normal distribution and has large outliers, so a two-sided asymptotic wilcoxon ranked sum (aka mann-whitney) test was used for categorical data and a spearman test for numeric data. statistical testing was performed using r version 4.0.5 on a x86_64-w64-mingw32/x64 (64-bit) platform running windows 10 x64 (build 19043). the r code and markdown files for the analyses are described and linked in appendix d. due to the uncertainty in how many the survey invitation actually reached, it was not possible to calculate response rates. 1. repositories with the highest staffing will have a significant positive correlation with larger numbers of depositors. analysis supported this hypothesis. spearman correlation tests were performed (using the midranks method to deal with ties) to examine relationships among numeric data and to control for large outliers. two scatterplots are included in the markdown (.rmd) file to compare the lack of trend lines among data points when outliers were included, with the spearman results which enabled exclusion of those extreme values and a more detailed view of the data. having a larger staff was significantly correlated with having a larger number of depositors (ρ =0.43, z = 2.25, pvalue = 0.02). because the correlation coefficient, rho (ρ), is positive, the relationship is positive: when one variable goes up, the other goes up. 2. larger numbers of depositors will be significantly correlated with offering advanced curation services. analysis supported this hypothesis. the survey asked respondents to indicate which of the following services each offers: minting digital object identifiers (dois); basic data curation tasks (e.g., checking data files, reading documentation, assignment of keywords/subject headings, running checksums); more staff-intensive data curation tasks (e.g., verification of file organization, file format normalization or migration over time, code checking, or replication); inserting an internal link between repository records for a deposited dataset and a deposited article based on those data; inserting a link from the repository record for a deposited dataset to a different repository’s record for a deposited article; data encryption; scanning for personally identifiable information; and other services (a write-in category). write-in answers for the other category were excluded from analysis. the remaining services were grouped into basic versus advanced services offered as shown in table 1 below. in most cases repositories offering advanced services provided https://doi.org/10.29173/iq1037 7/20 hayslett, michele & jansen, matthew (2022) factors contributing to repository success in recruiting data deposits, iassist quarterly 46(2), pp. 1-20. doi: https://doi.org/10.29173/iq1037 some if not all of the basic services as well, so testing focused on whether or not a repository offered advanced services. a wilcoxon-mann-whitney test was employed to handle tie values and outliers in the data, since it is based on ranks instead of the raw values and p values were computed using mid-ranks to break ties. offering advanced curation services had a significant association with having a larger number of depositors (z = -2.06, p-value = 0.04). repositories offering advanced curation services had 12.5 more median depositors than those not offering such services. table 1. groupings of curation services basic advanced minting dois basic data curation tasks inserting an internal link between repository records for a deposited dataset and a deposited article based on those data inserting a link from the repository record for a deposited dataset to a different repository’s record for a deposited article more staff-intensive data curation tasks data encryption scanning for personally identifiable information 3. repositories with larger numbers of depositors will be those which have depositors referred by faculty vs referrals from other places. this hypothesis was undetermined. with only three of 28 repositories without referrals from faculty, conclusions drawn from statistical testing seemed unrepresentative, but the fact that so many repositories had referrals from faculty is notable. 4. larger numbers of depositors will be significantly correlated with using social media to promote the repository. analysis supported this hypothesis. using social media had a significant association with having a larger number of depositors (z = -2.40, p-value = 0.02). those repositories using social media had on average 22 more median depositors than those that those not using it. 5. larger numbers of depositors will not be significantly positively correlated with infrastructure except where repository ingest is linked to electronic lab notebooks (elns). this hypothesis was undetermined. with only two of 28 repositories linking to elns, conclusions drawn from statistical testing seemed unrepresentative, but the fact that so few repositories offered such integration is notable. https://doi.org/10.29173/iq1037 8/20 hayslett, michele & jansen, matthew (2022) factors contributing to repository success in recruiting data deposits, iassist quarterly 46(2), pp. 1-20. doi: https://doi.org/10.29173/iq1037 6. larger numbers of depositors will be significantly positively correlated with repositories that encrypt data. this hypothesis was undetermined. with only four of 27 repositories offering encryption, conclusions drawn from statistical testing seemed unrepresentative, but the fact that so few repositories offered this service is notable. 7. larger numbers of depositors will be significantly positively correlated with having more staff dedicated specifically to data deposits. analysis supported this hypothesis. an asymptotic spearman correlation test found a significant correlation between the number of depositors and the number of dedicated data staff: those repositories with more data staff also have larger numbers of depositors (ρ= 0.71, z = 3.64, pvalue = [<0.01]). 8. larger numbers of depositors will be significantly positively correlated with older repositories. this hypothesis was unsupported. an asymptotic spearman correlation test found repository age was not significantly correlated with numbers of depositors (ρ= 0.03, z = 0.16, p-value = 0.87). limitations the small sample size (n varied between 26 and 28) and non-response bias are the greatest limitations of these analyses. stating the researchers’ intention to deposit the data publicly may have discouraged more from participating, but perhaps many repositories chose not to respond which were similar to each other but different from those that did participate. promoting the survey on several relevant email lists was intended to procure responses from a wide variety of repositories, and the outliers in the data seem to indicate this was successful. possible analyses were limited to two-way tests, however; results of regressions would have been questionable with such a small dataset but in a larger study could control for the effect of multiple factors at once. there is also a time-variant issue: the current study was conducted only at one point in time. if the study were administered year after year, changes in the number of depositors could be better related to changing repository characteristics. the current study was also limited by what is generally known about the community of data repositories broadly. since no complete directory of data repositories exists, targeted outreach and calculation of a response rate are impossible. discussion data were so imbalanced for hypotheses three, five, and six as to make statistical analysis inappropriate—for a valid analysis, the smaller group should be big enough to be plausibly representative. these results still offer interesting insights, though. for hypothesis three regarding faculty referral, 25 of 28 responding repositories received referrals from faculty. the result is notable because so many repositories are doing well in this aspect of service. the importance of building trust with and offering good service to customers, as well as the importance of early adopters, make good service a critical factor in the success of repositories. commercial retail research shows the importance of word-of-mouth advertising. businesswire (2011) quoted an american express survey: consumers will tell others about their customer service experiences, both good and bad, with the bad news reaching more ears. americans say they tell an average of nine people about good experiences, and nearly twice as many (16 people) about poor ones – making every individual service interaction important for businesses. (p. 2) https://doi.org/10.29173/iq1037 9/20 hayslett, michele & jansen, matthew (2022) factors contributing to repository success in recruiting data deposits, iassist quarterly 46(2), pp. 1-20. doi: https://doi.org/10.29173/iq1037 rogers’ (2003) classic research into the spread of new ideas identified early adopters as key to persuading many more people to adopt a given technology. this more recent edition notes that, although the internet has increased the speed of adoption of new technologies exponentially, early adopters are still key to the process. taken together, these studies suggest that ensuring early repository users have a good experience to share with colleagues is important. the insights generated by hypotheses five and six are similar: out of 26 respondents, only two linked ingest to electronic lab notebooks (hypothesis five), and of 28 respondents, only four offered encryption (hypothesis six). both services require extensive resources but may be services for which demand will grow in the future. macdonald and macneil (2015) point out the efficiencies to be gained by linking elns to repository ingest. describing researchers’ reaction to how this streamlined the (mandated) process of depositing data, they noted: when an initial, limited trial of [the new eln] was rolled out to ten labs … researchers from no less than nine of the labs reported that it was the ability to use [the eln] in conjunction with the … repository that was of most benefit. (p. 170) respondents to a survey by lagzian, et al. (2015) ranked fourth most important out of 46 factors the statement, ‘the ir is intuitive and easy to use.’ making data deposit easier and seamless with systems researchers already use (or that can otherwise save them time) may be key to boosting repositories’ success. in addition, offering encryption could increase deposits by substantially broadening the types of data a repository could accept, particularly in the u.s. in light of the data management requirement the national institutes of health will be implementing in january 2023. this mandate may result in a substantial increase in the number of research studies to be archived which contain personally identifiable information, health data, and/or student data, all of which require a high degree of protection. other results of the current study are unsurprising, especially since this analysis offers no indication of directionality: it is perhaps predictable that repositories with the largest number of staff have the highest deposit rates as hypothesis 1 supposed. it may well be that having a larger staff means one or more of those employees is able to focus on promotion and outreach, so that the large staff directly results in more deposits. likewise, offering higher levels of curation service (hypothesis 2) and having staff dedicated to data deposits (hypothesis 7) may indicate a larger staff that can, again, afford to devote time to those services and/or devote a person to promoting the repository, perhaps directly resulting in a larger number of deposits. on the other hand, this study assumed that a larger staff would be an indicator of a repository’s success. however, some researchers suggest a small repository staff working either with a supportive data community or backed by a larger team of reference/subject librarians can also be successful (akers & green, 2014; bailey, 2005; bell, foster et al., 2005; springer & cooper, 2020; ruediger et al., 2022). finally, the analysis not supporting hypothesis eight was unsurprising, that the age of a repository would correlate with more depositors, confirming tillman’s (2017) findings that older repositories do not automatically have more depositors, although that might seem counter-intuitive on the surface. the result of the final hypothesis, number four, relating use of social media to more depositors, was perhaps not unexpected, but the fact that only half of respondents reported using social media for outreach was surprising given its direct access to patrons and economy over print marketing. perhaps boulton’s (2020) observation of repositories’ widespread use of social media to promote their contents rather than their services explains this. repository staff may simply be more likely to promote their depositors than themselves. nevertheless, although chugh, grose and macht (2021) found that not all academics use social media, of those who do, a major purpose is for communication. they point out that o'keeffe (2019) documented ‘academics’ perception that twitter is a useful tool to assist with https://doi.org/10.29173/iq1037 10/20 hayslett, michele & jansen, matthew (2022) factors contributing to repository success in recruiting data deposits, iassist quarterly 46(2), pp. 1-20. doi: https://doi.org/10.29173/iq1037 informal academic development and learning, in particular learning about academic knowledge and practices’ (p. 990 [emphasis added]). referring to the previously mentioned importance of word-ofmouth advertising, it follows that repositories can benefit from having a presence in channels faculty are using for related purposes. conclusion the early twenty-first century has been a time of consciousness-raising with researchers about the importance of data management and data re-use, characterized by the proliferation of irs in particular but also repositories more broadly. the ‘if-you-build-it-they-will-come’ approach was unsuccessful, and many repositories reconsidered their approach. those that have survived and those that accept data deposits need to manage resources carefully and find the most efficient platforms, features and services to entice data deposit. this study extends the work of projects like hudson-vitale et al.’s data curation spec kit (2017), to help repository staff understand which factors are associated with greater numbers of depositors. future researchers may want to explore further the circumstances in which repositories with smaller staff sizes are successful, and specifically the role of librarians as partners in the repository and the part they play in the success of repositories. they may also want to consider whether the small number of respondents in this study may have been a result of the authors’ stated intention to deposit the data openly. responses to the penultimate survey question about comfort with various models of sharing the data (i.e., with different audiences) indicated discomfort even among some who chose to participate, and the number willing to share the data did not vary much regardless of the audience with whom the data were to be shared. table 2 below shows the distribution of answers. (also, the last item in appendix e, the detailed statistical test results, provides more detail about the pattern of respondents’ answers.) table 2. answer patterns for opinions on sharing data no maybe yes open to all 14 9 5 open only to researchers 11 9 7 open only to data curation researchers 12 10 5 the final question of the survey elicited reasons for the respondents’ opinions on sharing data. many indicated discomfort, either their own or that of their institutional leaders, with disclosing detailed information due to uncertainty about their own authority to release information publicly or (in particular) concern about making budget information public. overall recommendations from this study are: • further research should explore whether a stated intent to deposit (even deidentified) data deters participants; how repositories with small staff sizes achieve success; and what role reference/subject librarians play in making repositories successful. https://doi.org/10.29173/iq1037 11/20 hayslett, michele & jansen, matthew (2022) factors contributing to repository success in recruiting data deposits, iassist quarterly 46(2), pp. 1-20. doi: https://doi.org/10.29173/iq1037 • more repositories may want to connect with audiences through social media (while this study did not determine directionality, future research could also test this); • repositories may want to explore offering more advanced curation services such as checking code, linking to elns (or other university systems such as those that track grants), and/or offering encryption; and • repositories will want to continue to develop good relationships with researchers. finally, more repositories may want to evaluate whether their original structures and services are in fact meeting their audiences’ needs and publish those evaluations. only with more publicly available data will the repository community be able to benefit from past experience and more efficiently and effectively target services to their researchers. references abrams, s., cruse, p., strasser, c., willet, p., boushey, g., kochi, j., laurance, m., rizk-jackson, a. (2014). datashare: empowering researcher data curation. international journal of digital curation, 9(1), 110–118. https://doi.org/10.2218/ijdc.v9i1.305 akers, k.g. and green, j.a. (2014). towards a symbiotic relationship between academic libraries and disciplinary data repositories: a dryad and university of michigan case study. international journal of digital curation (9)1, 119–131. https://doi.org/10.2218/ijdc.v9i1.306 association of american universities-association of public & land-grant universities public access working group. (2017, november 29). aau-aplu public access working group report and recommendations. washington, d.c. https://www.aau.edu/sites/default/files/aau-files/keyissues/intellectual-property/public-open-access/aau-aplu-public-access-working-groupreport.pdf bailey, c. w. (2005). the role of reference librarians in institutional repositories. reference services review, 33(3), 259-267. https://doi.org/10.1108/00907320510611294 bell, s., foster, n.f., & gibbons, s. (2005). reference librarians and the success of institutional repositories. reference services review, 33(3), 283-290. https://doi.org/10.1108/00907320510611311 borgman, c. l., golshan, m. s., sands, a. e., wallis, j. c., cummings, r. l., darch, p., & randies, b. m. (2016). data management in the long tail: science, software, and service. international journal of digital curation, 11(1), 128–149. https://doi.org/10.2218/ijdc.v11i1.428 boulton, s. (2020). social engagement and institutional repositories: a case study. insights, 33, 1-9. https://doi.org/10.1629/uksg.504 bryant, r., lavoie, b. and malpas, c. (2017). part 2: scoping the university rdm service bundle. in the realities of research data management. dublin, oh: oclc research. https://doi.org/10.25333/c3z039 bryant, r., lavoie, b., & malpas, c. (2018). incentives for building university rdm services. the realities of research data management, part 3. dublin, oh: oclc research. https://doi.org/10.25333/c3s62f https://doi.org/10.29173/iq1037 https://doi.org/10.2218/ijdc.v9i1.305 https://doi.org/10.2218/ijdc.v9i1.306 https://www.aau.edu/sites/default/files/aau-files/key-issues/intellectual-property/public-open-access/aau-aplu-public-access-working-group-report.pdf https://www.aau.edu/sites/default/files/aau-files/key-issues/intellectual-property/public-open-access/aau-aplu-public-access-working-group-report.pdf https://www.aau.edu/sites/default/files/aau-files/key-issues/intellectual-property/public-open-access/aau-aplu-public-access-working-group-report.pdf https://doi.org/10.1108/00907320510611294 https://doi.org/10.1108/00907320510611311 https://doi.org/10.2218/ijdc.v11i1.428 https://doi.org/10.1629/uksg.504 https://doi.org/10.25333/c3z039 https://doi.org/10.25333/c3s62f 12/20 hayslett, michele & jansen, matthew (2022) factors contributing to repository success in recruiting data deposits, iassist quarterly 46(2), pp. 1-20. doi: https://doi.org/10.29173/iq1037 bryant, r. and faniel, i.m. works in progress webinar: identifying and acting on incentives when planning rdm services. oclc research webinar, november 13, 2018, https://www.oclc.org/research/events/2018/111318-incentives-when-planning-rdmservices.html burris, b. (2009). institutional repositories and faculty participation: encouraging deposits by advancing personal goals. public services quarterly, 5(1), 69-79. https://doi.org/10.1080/15228950802634212 burton, a., and treloar, a. (2009). designing for discovery and re-use: the ‘ands data sharing verbs’ approach to service decomposition. international journal of digital curation, 4(3), 44– 56. https://doi.org/10.2218/ijdc.v4i3.124 businesswire. (may 3, 2011). good service is good business: american consumers willing to spend more with companies that get service right, according to american express survey. https://www.businesswire.com/news/home/20110503005753/en/good-service-is-goodbusiness-american-consumers-willing-to-spend-more-with-companies-that-get-serviceright-according-to-american-express-survey chorus. (2019). how university of denver librarians used chorus institution dashboards in conjunction with their own internal data to help monitor public accessibility to the university’s publicly funded research. [brochure]. author. https://www.chorusaccess.org/wpcontent/uploads/ud-success-story-final-011518-3.pdf chugh, r., grose, r. & macht, s.a. (2021). social media usage by higher education academics: a scoping review of the literature. education and information technologies 26(1) 983–999. https://doi.org/10.1007/s10639-020-10288-z data curation network (2018). checklist of curated steps. https://docs.google.com/document/d/1rwt2obxooejrrfmvo9vakl4h41cl33zm5yyny3hbpz 8/edit#heading=h.ir351ea2236s davis, p., and connolly, m. (2007). evaluating the reasons for non-use of cornell university's installation of dspace. d-lib magazine, 13(3/4). http://www.dlib.org/dlib/march07/davis/03davis.html dunning, a. (2017, december 22). phd policies for research data in netherlands and uk. [powerpoint presentation]. open working web site. https://openworking.wordpress.com/2017/12/22/phd-policies-for-research-data-innetherlands-and-uk/ faniel, i.m., & connaway, l.s. (2018). librarians' perspectives on the factors influencing research data management programs. college & research libraries, 79(1), 100-119. https://doi.org/10.5860/crl.79.1.100 fecher, b., friesike, s., & hebing, m. (2015). what drives academic data sharing? plos one, 10(2), e0118053. https://doi.org/10.1371/journal.pone.0118053 ferguson, l. (2014, november 3). how and why researchers share data (and why they don't). https://www.wiley.com/network/researchers/licensing-and-open-access/how-and-whyresearchers-share-data-and-why-they-dont https://doi.org/10.29173/iq1037 https://www.oclc.org/research/events/2018/111318-incentives-when-planning-rdm-services.html https://www.oclc.org/research/events/2018/111318-incentives-when-planning-rdm-services.html https://doi.org/10.1080/15228950802634212 https://doi.org/10.2218/ijdc.v4i3.124 https://www.businesswire.com/news/home/20110503005753/en/good-service-is-good-business-american-consumers-willing-to-spend-more-with-companies-that-get-service-right-according-to-american-express-survey https://www.businesswire.com/news/home/20110503005753/en/good-service-is-good-business-american-consumers-willing-to-spend-more-with-companies-that-get-service-right-according-to-american-express-survey https://www.businesswire.com/news/home/20110503005753/en/good-service-is-good-business-american-consumers-willing-to-spend-more-with-companies-that-get-service-right-according-to-american-express-survey https://www.chorusaccess.org/wp-content/uploads/ud-success-story-final-011518-3.pdf https://www.chorusaccess.org/wp-content/uploads/ud-success-story-final-011518-3.pdf https://doi.org/10.1007/s10639-020-10288-z https://docs.google.com/document/d/1rwt2obxooejrrfmvo9vakl4h41cl33zm5yyny3hbpz8/edit#heading=h.ir351ea2236s https://docs.google.com/document/d/1rwt2obxooejrrfmvo9vakl4h41cl33zm5yyny3hbpz8/edit#heading=h.ir351ea2236s http://www.dlib.org/dlib/march07/davis/03davis.html https://openworking.wordpress.com/2017/12/22/phd-policies-for-research-data-in-netherlands-and-uk/ https://openworking.wordpress.com/2017/12/22/phd-policies-for-research-data-in-netherlands-and-uk/ https://doi.org/10.5860/crl.79.1.100 https://doi.org/10.1371/journal.pone.0118053 https://www.wiley.com/network/researchers/licensing-and-open-access/how-and-why-researchers-share-data-and-why-they-dont https://www.wiley.com/network/researchers/licensing-and-open-access/how-and-why-researchers-share-data-and-why-they-dont 13/20 hayslett, michele & jansen, matthew (2022) factors contributing to repository success in recruiting data deposits, iassist quarterly 46(2), pp. 1-20. doi: https://doi.org/10.29173/iq1037 gardner, d., w.toga, a., ascoli, g. a., beatty, jackson t., brinkley, j. f., dale, a. m., … wong, s. t. c. (2003). towards effective and rewarding data sharing. neuroinformatics, 1(3), 289–296. https://doi.org/10.1385/ni:1:3:289 hayslett, m. & jansen, m. (2022). factors contributing to repository success in recruiting data deposits (de-identified repository survey dataset in comma separated format), [computer file]. https://doi.org/10.17615/a4hj-4w45 hayslett, m. & jansen, m. (2022). factors contributing to repository success in recruiting data deposits (repository survey r script and markdown files), [computer files]. https://doi.org/10.17615/s6ps-qx27 helbig, k. (2019). data policies: institutionelle policies. https://www.forschungsdaten.org/index.php/data_policies#institutionelle_policies hudson-vitale, c., imker, h., johnston, l., carlson, j., kozlowski, w., olendorf, r., & stewart, c. association of research libraries,. (2017). spec kit 354: data curation. horton, laurence and data curation centre. (2016). overview of uk institution rdm policies. http://www.dcc.ac.uk/resources/policy-and-legal/institutional-data-policies jaradeh, m.y., auer, s., prinz, m., kovtun, v., kismihók, g., & stocker, m. (2019, april 29). open research knowledge graph: towards machine actionability in scholarly communication. https://arxiv.org/pdf/1901.10816.pdf jefferson, t. (1825, may 8). letter to henry lee. http://www.nlnrac.org/american/declaration-ofindependence/primary-source-documents/jefferson-to-lee lafferty-hess, s., rudder, j., downey, m., ivey, s., & darragh, j. (2018, may 30). conceptualizing data curation activities within two academic libraries. lis scholarship archive works. https://doi.org/10.31229/osf.io/zj5pq lagzian, f., abrizah, a., & wee, m. c. (2015). critical success factors for institutional repositories implementation. the electronic library, 33(2), 196–209. https://doi.org/10.1108/el-04-20130058 lowenberg, d. (2017, december 18). where’s the adoption? shifting the focus of data publishing in 2018 [blog post]. https://medium.com/@uc3cdl/wheres-the-adoption-shifting-the-focus-ofdata-publishing-in-2018-8506f80371cd lynch, c.a. (2003, february). institutional repositories: essential infrastructure for scholarship in the digital age. arl: a bimonthly report. http://old.arl.org/resources/pubs/br/br226/br226ir.shtml macdonald, s. and macneil, r. (2015). service integration to enhance research data management: rspace electronic laboratory notebook case study. international journal of digital curation, 10(1), 163-172. https://doi.org/10.2218/ijdc.v10i1.354 mchugh, a., ross, s., innocenti, p., ruusalepp, r., and hofman, h. bringing self-assessment home: repository profiling and key lines of enquiry within drambora. international journal of digital curation, 3(2), 130-142. https://doi.org/10.2218/ijdc.v3i2.64 https://doi.org/10.29173/iq1037 https://doi.org/10.1385/ni:1:3:289 https://doi.org/10.17615/a4hj-4w45 https://doi.org/10.17615/s6ps-qx27 https://www.forschungsdaten.org/index.php/data_policies#institutionelle_policies http://www.dcc.ac.uk/resources/policy-and-legal/institutional-data-policies https://arxiv.org/pdf/1901.10816.pdf http://www.nlnrac.org/american/declaration-of-independence/primary-source-documents/jefferson-to-lee http://www.nlnrac.org/american/declaration-of-independence/primary-source-documents/jefferson-to-lee https://doi.org/10.31229/osf.io/zj5pq https://doi.org/10.1108/el-04-2013-0058 https://doi.org/10.1108/el-04-2013-0058 https://medium.com/@uc3cdl/wheres-the-adoption-shifting-the-focus-of-data-publishing-in-2018-8506f80371cd https://medium.com/@uc3cdl/wheres-the-adoption-shifting-the-focus-of-data-publishing-in-2018-8506f80371cd http://old.arl.org/resources/pubs/br/br226/br226ir.shtml https://doi.org/10.2218/ijdc.v10i1.354 https://doi.org/10.2218/ijdc.v3i2.64 14/20 hayslett, michele & jansen, matthew (2022) factors contributing to repository success in recruiting data deposits, iassist quarterly 46(2), pp. 1-20. doi: https://doi.org/10.29173/iq1037 mattern, e., jeng, w., he, d., lyon, l., & brenner, a. (2015). using participatory design and visual narrative inquiry to investigate researchers’ data challenges and recommendations for library research data services. program, 49(4), 408–423. https://doi.org/10.1108/prog-01-20150012 mayo, c., vision, t. j., & hull, e. a. (2016). the location of the citation: changing practices in how publications cite original data in the dryad digital repository. international journal of digital curation, 11(1), 150-155. https://doi.org/10.2218/ijdc.v11i1.400 meishar-tal, h., & pieterse, e. (2017). why do academics use academic social networking sites? the international review of research in open and distributed learning, 18(1). https://doi.org/10.19173/irrodl.v18i1.2643 minor, d., critchlow, m., hutt, a., fleming, d., bergstrom, m. l., & sutton, d. (2014). research data curation pilots: lessons learned. international journal of digital curation, 9(1), 220–230. https://doi.org/10.2218/ijdc.v9i1.313 national research council. (1985). sharing research data. washington, d.c.: national academies press. o'keeffe, m. (2019). academic twitter and professional learning: myths and realities. international journal for academic development, 24(1), 35–46. https://doi.org/10.1080/1360144x.2018.1520109 otto, j.j., (2016). a resonant message: aligning scholar values and open access objectives in oa policy outreach to faculty and graduate students. journal of librarianship and scholarly communication. 4, p. ep2152. http://doi.org/10.7710/2162-3309.2152 paul, s. (2012). institutional repositories: benefits and incentives. international information and library review, 44(4), 194–201. https://doi.org/10.1080/10572317.2012.10762932 peer, l. and green, a. (2012). building an open data repository for a specialized research community: process, challenges and lessons. international journal of digital curation, 7(1), 151–162. http://dx.doi.org/10.2218/ijdc.v7i1.222 pienta, a. m., alter, g. c., & lyle, j. a. (2010). the enduring value of social science research: the use and reuse of primary research data. in the organisation, economics and policy of scientific research. torino, italy. https://doi.org/https://deepblue.lib.umich.edu/handle/2027.42/78307 plale, b., mcdonald, r. h., chandrasekar, k., kouper, i., konkiel, s., hedstrom, m. l., … kumar, p. (2013). sead virtual archive: building a federation of institutional repositories for long-term data preservation in sustainability science. international journal of digital curation, 8(2), 172– 180. https://doi.org/10.2218/ijdc.v8i2.281 rogers, e. m. (2003). diffusion of innovations. (5th ed.). new york: free press. ruediger, d., macdougall, r., cooper, d., carlson, j., herndon, j., & johnston, l. (2022, august 9). leveraging data communities to advance open science: findings from an incubation workshop series. https://doi.org/10.18665/sr.317145 https://doi.org/10.29173/iq1037 https://doi.org/10.1108/prog-01-2015-0012 https://doi.org/10.1108/prog-01-2015-0012 https://doi.org/10.2218/ijdc.v11i1.400 https://doi.org/10.19173/irrodl.v18i1.2643 https://doi.org/10.2218/ijdc.v9i1.313 https://doi.org/10.1080/1360144x.2018.1520109 http://doi.org/10.7710/2162-3309.2152 https://doi.org/10.1080/10572317.2012.10762932 http://dx.doi.org/10.2218/ijdc.v7i1.222 https://doi.org/https:/deepblue.lib.umich.edu/handle/2027.42/78307 https://doi.org/10.2218/ijdc.v8i2.281 https://doi.org/10.18665/sr.317145 15/20 hayslett, michele & jansen, matthew (2022) factors contributing to repository success in recruiting data deposits, iassist quarterly 46(2), pp. 1-20. doi: https://doi.org/10.29173/iq1037 salo, d. (2008). innkeeper at the roach motel. library trends; fall library & information science abstracts, 57(2), 98-123. https://muse.jhu.edu/article/262026/pdf savage, c. j., & vickers, a. j. (2009). empirical study of data sharing by authors publishing in plos journals. plos one, 4(9), e7078. https://doi.org/10.1371/journal.pone.0007078 springer, r. and cooper, d. (2020). data communities: empowering researcher-driven data sharing in the sciences. international journal of digital curation, 15(1). http://dx.doi.org/10.2218/ijdc.v15i1.695 tenopir, c., allard, s., douglass, k., aydinoglu, a. u., wu, l., read, e., … frame, m. (2011). data sharing by scientists: practices and perceptions. plos one, 6(6), e21101. https://doi.org/10.1371/journal.pone.0021101 tillman, r.k. (2017). where are we now? survey on rates of faculty self-deposit in institutional repositories. journal of librarianship and scholarly communication, 5 (general issue), ep2203. https://doi.org/10.7710/2162-3309.2203 tu delft (2018, june 26). tu delft research data framework policy. https://zenodo.org/record/2573160/files/tu%20delft%20research%20data%20framework% 20policy.pdf?download=1 university of groningen (2013, september 1). university of groningen phd regulations. section 4.1.5. https://www.rug.nl/about-us/organization/rules-andregulations/onderzoek/promotiereglement-14-en.pdf wallis, j. c., rolando, e., & borgman, c. l. (2013). if we share data, will anyone use them? data sharing and reuse in the long tail of science and technology. plos one, 8(7), e67332. https://doi.org/10.1371/journal.pone.0067332 https://doi.org/10.29173/iq1037 https://muse.jhu.edu/article/262026/pdf https://doi.org/10.1371/journal.pone.0007078 http://dx.doi.org/10.2218/ijdc.v15i1.695 https://doi.org/10.1371/journal.pone.0021101 https://doi.org/10.7710/2162-3309.2203 https://zenodo.org/record/2573160/files/tu%20delft%20research%20data%20framework%20policy.pdf?download=1 https://zenodo.org/record/2573160/files/tu%20delft%20research%20data%20framework%20policy.pdf?download=1 https://www.rug.nl/about-us/organization/rules-and-regulations/onderzoek/promotiereglement-14-en.pdf https://www.rug.nl/about-us/organization/rules-and-regulations/onderzoek/promotiereglement-14-en.pdf https://doi.org/10.1371/journal.pone.0067332 16/20 hayslett, michele & jansen, matthew (2022) factors contributing to repository success in recruiting data deposits, iassist quarterly 46(2), pp. 1-20. doi: https://doi.org/10.29173/iq1037 appendix a: email solicitation message a word version of the invitation to participate is available in the carolina digital repository at https://doi.org/10.17615/bj43-pw90. appendix b: consent form a word version of the consent form is available in the carolina digital repository at https://doi.org/10.17615/q0yy-8h57. appendix c: survey instrument a word version of the survey instrument is available in the carolina digital repository at https://doi.org/10.17615/7vn5-1g28. appendix d: code used in data analysis the code with which data analysis was performed is available in the carolina digital repository at https://doi.org/10.17615/s6ps-qx27 and on github at https://github.com/unc-libraries-data/reposurvey. • the r script with which data analyses were performed is create-analysis-data.r, and • the markdown file with which tables and diagrams were created is analysis.rmd. https://doi.org/10.29173/iq1037 https://doi.org/10.17615/bj43-pw90 https://doi.org/10.17615/q0yy-8h57 https://doi.org/10.17615/7vn5-1g28 https://doi.org/10.17615/s6ps-qx27 https://github.com/unc-libraries-data/repo-survey https://github.com/unc-libraries-data/repo-survey 17/20 hayslett, michele & jansen, matthew (2022) factors contributing to repository success in recruiting data deposits, iassist quarterly 46(2), pp. 1-20. doi: https://doi.org/10.29173/iq1037 appendix e: detailed results of statistical tests by hypothesis note: hypotheses marked with an asterisk (*) were not tested. tables of these data are presented to show their imbalance. hypothesis 1: repositories with the highest staffing will have a significant positive correlation with larger numbers of depositors. (n=28) asymptotic spearman correlation test z = 2.2466, p-value = 0.02467 alternative hypothesis: true rho is not equal to 0 ρ = 0.4323595 staff n median depositors 0.2 1 1.0 0.3 1 4.0 0.5 1 0.0 1.0 6 20.0 1.2 2 6.5 1.5 3 2.0 1.7 1 1.0 2.0 2 4.5 3.0 1 20.0 3.5 1 468.0 4.0 2 30.5 5.0 1 0.0 6.0 1 25.0 9.0 1 22.0 12.0 1 25.0 17.0 1 45.0 25.0 1 10.0 110.0 1 360.0 hypothesis 2. larger numbers of depositors will be significantly correlated with offering advanced curation services. (n=26) offer advanced curation services count median depositors no 10 4.0 yes 16 16.5 na 2 16.5 asymptotic wilcoxon-mann-whitney test z = -2.6309, p-value = 0.008516 alternative hypothesis: true mu is not equal to 0 https://doi.org/10.29173/iq1037 18/20 hayslett, michele & jansen, matthew (2022) factors contributing to repository success in recruiting data deposits, iassist quarterly 46(2), pp. 1-20. doi: https://doi.org/10.29173/iq1037 hypothesis 3.* repositories with larger numbers of depositors will be those which have depositors referred by faculty vs referrals from many places. (n=28) receive referrals from faculty count no 3 yes 25 statistical testing is not supported by results with this level of category imbalance but the fact that so many repositories had referrals from faculty is notable. hypothesis 4. larger numbers of depositors will be significantly correlated with using social media to promote the repository. (n=28) use social media count no 15 yes 13 asymptotic wilcoxon-mann-whitney test z = -2.4026, p-value = 0.01628 alternative hypothesis: true mu is not equal to 0 hypothesis 5.* larger numbers of depositors will not be significantly positively correlated with infrastructure except where repository ingest is linked to electronic lab notebooks (elns). (n=28) eln linked count no 26 yes 2 statistical testing is not supported by results with this level of category imbalance but the fact that so few repositories linked with elns is notable. hypothesis 6.* larger numbers of depositors will be significantly positively correlated with repositories that encrypt data. (n=27) offer encryption count no 23 yes 4 na 1 statistical testing is not supported by results with this level of category imbalance but the fact that so few repositories offered encryption is notable. https://doi.org/10.29173/iq1037 19/20 hayslett, michele & jansen, matthew (2022) factors contributing to repository success in recruiting data deposits, iassist quarterly 46(2), pp. 1-20. doi: https://doi.org/10.29173/iq1037 hypothesis 7. larger numbers of depositors will be significantly positively correlated with having dedicated staff for data deposits. (n=27) have data staff count median depositors false 4 0 true 23 11 na 1 468 asymptotic spearman correlation test z = 3.6407, p-value = 0.0002719 alternative hypothesis: true rho is not equal to 0 ρ = 0.71 hypothesis 8. larger numbers of depositors will be significantly positively correlated with older repositories. (n=28) age in yrs count median depositors 1 1 33 2 4 24.5 3 4 1 4 4 13 5 3 8 6 3 7 7 3 20 8 1 0 10 1 0 12 1 3 30 1 10 57 1 360 asymptotic wilcoxon-mann-whitney test z = 0.16139, p-value = 0.8718 alternative hypothesis: true rho is not equal to 0 ρ = 0.03 https://doi.org/10.29173/iq1037 20/20 hayslett, michele & jansen, matthew (2022) factors contributing to repository success in recruiting data deposits, iassist quarterly 46(2), pp. 1-20. doi: https://doi.org/10.29173/iq1037 opinions about sharing data, detailed answer patterns the table below displays how the 29 respondents answered on their feelings about sharing with different audiences. all who answered in each pattern open to anyone open only to researchers open only to data curation/repository researchers total answers maybe maybe maybe 7 maybe maybe yes 2 no maybe maybe 1 no no maybe 1 no no no 11 no yes yes 1 yes yes yes 4 yes (blank) (blank) 1 (blank) (blank) (blank) 1 endnotes 1 michele hayslett is the librarian for numeric data services and data management and matthew jansen is the data analysis librarian in the university libraries at the university of north carolina at chapel hill. questions may be sent by email to the lead author: michele_hayslett@unc.edu. https://doi.org/10.29173/iq1037 mailto:michele_hayslett@unc.edu 16 iassist quarterly winter/spring 2010 by walter w. giesbrecht 1 data reference in depth: sources of international labour data abstract most countries provide access to national labour data on the internet, but finding it can be a frustrating exercise, especially if the country in question does not provide the information in a language understood by the researcher. foreign labour data are also often sought when attempting to make international comparisons of particular labour statistics. this article reviews internet-accessible multi-country compilations of labour data that provide access in multiple languages, with attention given to those that permit international comparisons to be made. keywords: labour data, labor data, comparative data, international data while data and statistics on labour and employment for wide variety of countries are now readily available on the internet, finding statistics for a particular country and/ or on a particular parameter for multiple countries can be a challenge. even more frustrating can be trying to find statistics that are comparable across countries. in this article, i will describe a number of sites that provide easy access to both national labour statistics and international and comparative labour statistics. international labour organization (ilo) <http://www.ilo.org> the ilo should be the first stop for labour data beyond one's own borders. it is the hub for labour statistics within the un system. the ilo is "... devoted to advancing opportunities for women and men to obtain decent and productive work in conditions of freedom, equity, security and human dignity. its main aims are to promote rights at work, encourage decent employment opportunities, enhance social protection and strengthen dialogue in handling work-related issues."2 the ilo provides, in english, french and spanish, official core labour statistics and estimates for over 200 countries since 1969. besides the data, it also provides considerable metadata, definitions and methodological descriptions of main national statistical sources. the data provided use mostly national definitions, except in limited cases; the difficulties this poses will be discussed at the end of this paper. on the ilo home page, a link to ‘statistics and databases’ is given on the right side. this page shows a list of databases, both of statistics and literature, with brief descriptions. the ones that are primarily data-related are listed under the ‘statistics’ subheading: 1. laborsta database of labour statistics 2. key indicators of the labour market (kilm) 3. statistical information and monitoring programme on child labour (ipec-simpoc) 4. labour force surveys plus, separate from these four sites, but devoted primarily to data that are more current, is 5. ilo global job crisis observatory laborsta <http://laborsta.ilo.org/> laborsta covers official core labour statistics and estimates for over 200 countries since 1969. also provides methodological descriptions of main national statistical sources. [3] 3 at the time of writing, there were 224 entries for countries, of which some are subnational areas included for historical reasons (e.g., the four entries for germany: germany as it is now, the former frg and gdr, and the 5 new länder plus what was east berlin), or overseas territories (e.g., st. pierre & miquelon, an overseas territory of france). the first page offers the user a choice of statistics for 11 general topics, in most cases available in both annual and monthly frequencies: • total and economically active population • employment • unemployment 17 iassist quarterly winter/spring 2010 • hours of work • wages • labour cost • consumer price indices • occupational injuries • strikes and lockouts • household income and expenditure • international labour migration following the link to ‘employment', for example, leads to six links to data sources, some of which go to other databases. the hover text for each link lets the user know how far back the data go, at least as an upper limit. there is also a ‘by topic’ link on the left-hand menu bar that leads to a page that lists all the available topics. the user can also see what is available by country, or by publication (for those familiar with the range of topics covered by ilo yearbook of labour statistics, ilo bulletin of labour statistics, and ilo october inquiry). ultimately, the data can either be viewed online, or downloaded to excel. the online version has links to relevant metadata embedded in it; the excel spreadsheet has limited metadata. the user needs to look at the ‘view data’ page to have a clear idea of what the data represent; if nothing else, the spreadsheet should contain links to the necessary metadata. on the laborsta home page (and on the left menu on all subsequent pages) is a set of metadata links: • definitions of terminology used (e.g., economically active population, employment, etc.) in each case, an essay is provided with the definitions used, potential sources of data on the given concept, plus alerts as to potential problems with interpretation; • classifications detailed information on the various international standard classification schemes used: o international classification by status in employment o international standard classification of education o international standard classification of occupations o international standard industrial classification of all economic activities o system of national accounts 1993 • sources and methods information on the scope of the statistics, their definitions and the methods used by the national statistical services in establishing the data published.4 this is the full text of the ten-volume series sources and methods: labour statistics 5 perusal of these metadata links is critical to a full understanding of the data and their limitations for making comparisons. the laborsta homepage also contains links to short term indicators of the labour market; these are drawn from official national statistical sources, based on national definitions, and are not seasonally adjusted. they contain monthly data for the past twelve months, quarterly data for the past five quarters, and annual data for the past three years. the data are available by topic or by country; in addition, complete country profiles and selected series by country are available in pdf or xls formats. key indicators of the labour market (kilm) <http://www.ilo.org/empelm/what/lang--en/ wcms_114240> available in print, online and as a standalone software package, the kilm has been published every two years since 1999. currently in its sixth edition, the kilm • is a comprehensive database of country-level data on 20 key indicators of the labour market from 1980 to the latest available year; • is a source of the latest ilo world and regional estimates of employment and unemployment indictors. a training tool on development and use of labour market indicators; • highlights of current labour market trends; • provides analyses of key issues in the labour market6 the online version, known as kilmnet (still in beta) offers 32 tables in six groupings: • participation in the world of work (3 tables) • employment indicators (11 tables) • unemployment indicators (6 tables) • educational attainment (2 tables) • wages and labour costs (7 tables) • performance and poverty indicators (3 tables) the interface is similar in some ways to that of the world 18 iassist quarterly winter/spring 2010 development indicators 7, except that all the data parameter selections can be made on the same screen the differences between the kilm and laborsta are addressed in the document "guide to understanding the kilm" 8. basically, laborsta (and its print equivalent, the yearbook of labour statistics) are the best source of nationally-reported labour statistics, whereas kilm supplements these data from other sources when those other sources are considered more accurate or more complete, and offer better international comparability. the kilm is not restricted to using national data as reported; it makes efforts to use indicator series that are more comparable across time and geographies. the kilm offers three comparable or harmonized series: labour force participation rates, employment-to-population ratios, and the inactivity rate. other series have been made as comparable as possible; anomalies in definition and methodologies are clearly indicated in the table notes. in other words, when national data are desired, the user should start with laborsta; when comparisons between countries are to be made, kilm should be the starting point statistical information and monitoring programme on child labour (ipec-simpoc) <http://www.ilo.org/ipec/childlabourstatisticssimpoc/ lang--en/index.htm> this is the statistical arm of the international programme for the elimination of child labour (ipec). it claims to offer: • specific questionnaires for child labour surveys • manuals and training kits on how to carry out child labour data collection in households, schools and at the workplace • guidance on how to properly process and analyse the collected information • micro datasets and survey reports from around the world • research on critical statistical issues • regular trend reports microdata sets for 30 countries (as of 2010.09.26) are available under the left menu option "surveys"; the data are often available for multiple years, and are provided in ascii, spss or stata, along with metadata labour force surveys <http://www.ilo.org/dyn/lfsurvey/lfsurvey.home> this page lists links to the websites for the labour force surveys (defined as “a standard household-based survey of work-related statistics") of all the countries in the ilo. it does not provide access to the microdata, and is inconsistent as to what it does link. using examples from countries the author of this paper is currently familiar with, • in the case of canada, it links to the press release associated with the most current release of the labour force survey; it does not mention the survey of labour and income dynamics; • for the united states, it links to the "current labor statistics" page from the monthly labor review; it provides access to labour data from both the current population survey and current employment statistics; • for australia, it simply provides a link to the "statistics by topic: labour" page on the australian bureau of statistics website, with no direct access to the household, income and labour dynamics in australia (hilda) survey in each case, however, the ilo does provide access to a page describing the methodology of the main labour force survey of each country. also, the user gets a link to the country’s relevant website -this in itself can be valuable, as it enables one to follow up with the country itself on any questions, as well as potentially providing access to reports on data unavailable from the ilo. global statistics on the labour market <http://www.ilo.org/pls/apex/f?p=109:11:0> part of the ilo's ilo global job crisis observatory 19 iassist quarterly winter/spring 2010 website 9, this site offers "the latest national data for indicators which have been selected for their ability to reflect recent and short term changes." data are available by topic or country, and are not seasonally adjusted or otherwise altered by the ilo. updates of indicators and associated publications are usually monthly, but can be more or less frequent. other data sites the ilo is not the only site from which one can get compiled labour data from a variety of countries. the following section of this paper describes some others. international labor comparisons (us) <http://www.bls.gov/data/#international> this site provides labour data for selected countries using a variety of indicators, all of which are adjusted to u.s. concepts 10 unless otherwise noted. data series are presented as: a group of most requested series; either a one-screen or multi-screen data search process; a series of static tables (in html, pdf or xls formats). some series, such as the supplementary tables comparing manufacturing productivity and unit labour cost trends, go back to 1950; most, however, start sometime in the 1990s. oecd labour statistics <http://www.oecd.org/> this section of the article is divided into two parts: for oecd ilibrary (formerly sourceoecd) subscribers and those who do not subscribe. non-subscribers to oecd ilibrary finding labour and employment statistics on the oecd site can be a little confusing, as there are three different pages that provide related information. (1) from the above address, you can reach the oecd statistics portal <http://www.oecd.org/statsportal/0,3352, en_2825_293564_1_1_1_1_1,00.html> this has a section on labour: <http://www.oecd.org/topicstatsportal/0,3398, en_2825_495670_1_1_1_1_1,00.html> the labour portion of their portal has two sections: "labour statistics" (13 indicators) and "unemployment statistics" (6 indicators), along with links to various reports, definitions and other metadata. the links for each indicator can take one to a publicly accessible (i.e., to nonsubscribers) portion of oecd.stat, or one of the oecd's many other labourand employment-related databases. much of the data are inaccessible to non-subscribers. (2) the oecd also offers access to their content by general subject area; via this route, there is a section on "employment" <http://www.oecd.org/topic/0,3373, en_2649_37457_1_1_1_1_37457,00.html> clicking on 'statistics" in the menu bar on this page leads one to a long list of discrete statistical publications on employment, many of which, but not all, are also under the "unemployment statistics" heading on the oecd statistics portal page on labour. (3) the oecd also has a directorate for employment, labour and social affairs <http://www.oecd.org/department/0,3355, en_2649_33729_1_1_1_1_1,00.html> the link to "statistics" within this section leads one to a page that is actually more helpful, in many ways, that the other two: a set of links to relevant pages of statistics are given at the top of the page. the relevant one is to employment <http://www.oecd.org/els/employment/data> • the employment database offers current statistics for international comparisons and trends over time. • key employment statistics has summary tables for oecd countries with indicators on labour market outcomes and policies and how they compare with the oecd average. in both cases, links to sources and relevant metadata are readily accessible. while the written reports may be unique to the oecd, the data are derived from national labour force surveys. given this, you may find it easier to use laborsta. subscribers to oecd ilibrary <http://www.oecd-ilibrary.org/> for subscribers, the quest is much simpler. from the oecd ilibrary home page, follow the link to "statistics" at the top of the page, and then find “oecd employment and labour market statistics" in the list of available databases. the abstract states that this database includes a range of annual labour market statistics and indicators from 1960 broken down by sex and age as well as information about part-time and short-time workers, job tenure, hours worked, unemployment duration, trade union, employment protection legislation, minimum wages, labour market programmes for oecd countries and non-member economies. the data are more accessible, and many preconfigured tables are available. related oecd publications are clearly 20 iassist quarterly winter/spring 2010 indicated. a nice feature of the new ilibrary is a link, for each available dataset, to a page that provides a citation and links to download the citation information to a variety of citation managers other related sites a source that can be generally useful in finding or interpreting labour data is how to find labour statistics 11. the focus is naturally on the ilo’s resources, but it does offer a selection of resources available to the public from other international organizations and partner institutions. resources, including a variety of metadata, are listed by topic as well as either globally or regionally. problems of comparing data from different countries as alluded to earlier, problems can occur when attempting to compare data across geographies. in the ilo definitions of terminology, one frequently encounters phrases such as “national definitions of [desired criterion] may differ from the recommended international standard definition” or “national practices vary between countries” or “the comparability of the data is hampered by the differences between countries and even within a country”. a concrete example: in canada, a part-time worker is one who usually works less than 30 hours per week; in australia and the u.s, the criterion is less than 35 hours or more per week, and in the eu, it is whatever the individual being asked considers part-time work. in addition, in terms of classifying one’s status as employed vs. employer, most countries classify managers and directors of incorporated enterprises as employees, while in some others they are classified as employers. one wonders if comparability is possible at all. efforts are being made to create labour data that can be compared across national boundaries. as previously mentioned, the kilm database incorporates efforts to create harmonized variables that are comparable between countries. other efforts include the crossnational equivalent file (cnef project based at cornell university12 , which includes “equivalently defined variables” for data from the u.k., australia, korea, u.s., switzerland, canada and germany. the comparative perspectives database, an offshoot of the gender & work database 13, is a project that is attempting to create a variety of cross-tabulations on the subject of precarious employment using data from the u.s., canada, australia and the countries of the european union. these efforts, as well as others, will make the task of comparability much simpler, albeit not as comprehensively as one might like. making variables have equivalent definitions is not necessarily an ideal solution either. the legal and regulatory frameworks surrounding labour and employment are dependant on national definitions; creating equivalent definitions removes the data from its national context, thereby, in some ways, making the data less relevant to the people and environment it represents. it is an issue one has to keep in mind, depending on the reason for comparing the data. the important thing to remember, when attempting to compare statistics on similar parameters from different countries, is to examine the metadata, especially the definitions, closely to make sure apples are being compared to apples, not apples to pineapples, as it were. just because something is called the same thing does not mean it is the same thing. ultimately, it is the responsibility of the user of the data/statistics to ensure that their use is appropriate, but the data professional can, at least, alert her/him to the issues. references 1 author contact information: walter w. giesbrecht, 203e scott library,york university, toronto, on, canada m3j 1p3 (email: walterg@yorku.ca) 2 “about the ilo”, international labour organization, accessed august 18, 2010, http://www.ilo.org/global/ about_the_ilo/lang--en/index.htm. 3 “statistics and databases – what we do”, international labour organization, accessed august 18, 2010, http:// www.ilo.org/global/what_we_do/statistics/lang--en/index. htm. 4 “sources and methods: labour statistics”, international labour organization, accessed september 25, 2010, http:// laborsta.ilo.org/applv8/data/ssme.html. 5 “sources and methods in labour statistics”, international labour organization, accessed september 27, 2010, http://www.ilo.org/stat/publications/sources/. 6 “key indicators of the labour market (kilm), sixth edition “,international labour organization, accessed september 26, 2010, http://www.ilo.org/empelm/what/ pubs/lang--en/wcms_114060/index.htm 21 iassist quarterly winter/spring 2010 7 “world development indicators “,world bank, accessed september 26, 2010, http://data.worldbank.org/datacatalog/world-development-indicators. 8 “guide to understanding the kilm”, international labour organization, accessed september 26, 2010, http://kilm.ilo.org/kilmnetbeta/pdf/guide%20to%20 understanding%20the%20kilmen-2009.pdf. 9 “ilo global job crisis observatory”, international labour organization, accessed september 26, 2010, http:// www.ilo.org/pls/apex/f?p=109:1:0. 10 “international comparisons of annual labor force statistics, adjusted to u.s. concepts, 10 countries, 19702009: introduction”, bureau of labor statistics, accessed september 26, 2010, http://www.bls.gov/fls/flscomparelf/ notes.htm#introduction. 11 “how to find labour statistics “,international labour organization, accessed september 26, 2010, http://www.ilo. org/public/english/support/lib/resource/subject/labourstat. htm. 12 “cornell university user package for the crossnational equivalent file (cnef), 1970-2008 “, department of policy analysis and management, cornell university, accessed september 27, 2010, http://www. human.cornell.edu/pam/research/centers-programs/germanpanel/cnef.cfm. 13 “gender & work database”, accessed september 27, 2010, http://www.genderwork.ca. 1/10 beeken, jeannine (2020) a recommendation to the ssh community: take a linguist on board, iassist quarterly 45(1), pp. 1-10. doi: https://doi.org/10.29173/iq992 a recommendation to the ssh community: take a linguist on board jeannine beeken1 abstract in this paper we address how natural language processing (nlp) approaches and language technology can contribute to data services in different ways; from providing social science users with new approaches and tools to explore oral and textual data, to enhancing the search, findability and retrieval of data sources. by using linguistic approaches we are able to process data, for example using automated speech recognition (asr) and named entity recognizers (ner), extract key concepts and terms, and improve search strategies. we provide examples of how computational linguistics contribute to and facilitate the mining and analysis of oral or textual material, for example (transcribed) interviews or oral histories, and show how free open source (os) tools can be used very easily to gain a quick overview of the key features of text, which can be further exploited as useful metadata. keywords natural language processing (nlp), language technology, data and metadata services and infrastructure, social sciences and humanities (ssh) introduction introducing (more) linguistics into social science research is no simple feat. moreover, the fact that both oral and written material has to be considered presents another complicating factor. in the following sections, we investigate how (computational) linguistics and tools can contribute to social science research by providing new approaches and tools and by enhancing findability and retrieval, taking into account the fact that these principles promote machine-actionability. we will especially pay attention to findability (online searchable and discoverable) and interoperability (using for example standards and schema, controlled vocabularies, keywords, thesauri or ontologies) of metadata and data. a basic suite of linguistic tools and methods in this section, we introduce a possible suite of different linguistic tools and methods which could, or rather, should be added to the standard package of services (for external users) and infrastructure (internal metadata and data managers) at archives. 1. nlp tools which optimize search, findability and retrieval, are spell checkers and correctors, stopword excluders, autosuggest functionality based on one or more thesauri, clustering of keywords and their synonyms (for example ’war’ and ’armed conflict’), priority lists of abbreviations/acronyms as used in social science research (for example ’als, ehs, closer, gus’) and language-specific stemmers (searching for ’tax’ finds correctly studies about ’tax, taxes, taxation’, but not about ’taxi’ or ’taxis’). 2. automated speech recognition (asr) tools can partially take over manual transcription, while separating oral language from, for example, silences. they are able to distinguish spoken natural language from surrounding noise and can also recognize and distinguish https://doi.org/10.29173/iq992 2/10 beeken, jeannine (2020) a recommendation to the ssh community: take a linguist on board, iassist quarterly 45(1), pp. 1-10. doi: https://doi.org/10.29173/iq992 between different speakers taking part in, for example, an interview. most importantly, they convert spoken language into written language or text, using word segmentation, whereby the written text has been aligned with the spoken fragments. 3. named entity recognizers (ner) identify and classify named entities, such as person names, organisations, geospatial terminology etc. they can simplify anonymization and assist checks on disclosure and de-identification. this feature could be introduced as part of, for example, a self-depositing portal or an ingesting infrastructure. 4. information extraction (ie) tools detect keywords, thus assisting indexing and pre-populating the relevant metadata fields. also, this feature could be useful when self-depositing or when adding keywords during the ingesting process. moreover, matching or aligning the extracted keywords with thesaurus terms or any other (standardized) controlled vocabulary of topics, for example, would improve findability and retrieval. 5. concordances (kwic/keywords in context) and correlations in a (group of) texts can easily be generated. this feature helps detect possible and unexpected clusters and patterns, for example between ’schools’ and ’knives’, which leads to new research questions and insights. basic linguistic challenges when developing tools and services for finding and retrieving archived oral and written data, linguists face three main challenges: 1. disambiguation or separation of, for example, homonyms such as ‘book’ (as in ‘book a holiday’ or ‘reading a book’), ‘bank’ (as in ‘a bank of a river’, ‘a savings bank’, ‘a bank of snow’), ‘current’ (as in ‘my current job’, ‘ocean currents’, ‘a current flowing to a lamp’) 2. clustering or grouping of, for example, synonyms (such as ‘current, contemporary, presentday, present, ongoing’), words sharing the same stem (such as ‘nurse, nurses, nursing’), coreferences (such as ‘i voted for sammy, since she is my sister and the current chairperson of the board’), multiword expressions and idioms (such as ‘black money’, ‘black widow’, ‘black humour’, ‘with respect to’, ‘giving the cold shoulder’, ‘being all ears’). 3. reduction or control, for example removing ‘meaningless’ stopwords such as ‘was, be, you, me, to, for, or, if, when’ (excluding ‘not, n’t’) while keeping ‘meaningful’ keywords and domain dependent or domain specific terminology, for example ‘was’ or ’was’ meaning ‘wealth and assets survey’. as said, introducing more linguistics into social science research is no simple feat. for example, some of the challenges mentioned appear to contradict each other—separation, but at the same time also grouping— and the fact that both oral and written data have to be taken into account presents another complicating factor. for example, how do asr tools (speech to text) distinguish between homophones such as ‘be’ and ‘bee’, ‘meet’ and ‘meat’, ‘friar’ and ‘fryer’, ‘nun’ and ‘none’, ‘grease’ and ‘greece’, ‘knead’ and ‘kneed’ or ‘need’. improving search, findability and interoperability by using language-specific tools search, findability and retrieval can be optimized by implementing well-known nlp tools such as spell checkers or spell correctors, taking into account caseand diacritic-insensitivity. other tools that are commonly used within the wide domain of search engine optimization (seo) provide autosuggestions based on a (domain specific) thesaurus or a controlled list of keywords and their synonyms, which improves findability to a great extent. for example, in the uk data service data https://doi.org/10.29173/iq992 3/10 beeken, jeannine (2020) a recommendation to the ssh community: take a linguist on board, iassist quarterly 45(1), pp. 1-10. doi: https://doi.org/10.29173/iq992 catalogue, searching for ’health’ generates the autosuggested list ‘health and well-being’, ’health behaviour’, ’health care’ and a list of thesaurus (in this case hasset) keywords ’public health risks’, ’men’s health’. these keywords form an opening to clusters of (quasi-)synonyms (i.e., different form, similar meaning), for example, ’war’ and ’armed conflict’, or ’energy prices’ and ’energy tariffs’, ’fuel prices’. next, an extended list of acronyms (for example ’closer’) and abbreviations (for example ’ehs’) as used in social science research can be implemented as part of the search algorithm, giving a higher scoring priority to domain-specific terminology and conceptual meaning than to common language meaning. for example, ’gus’, ’closer’, ’dots’, ’als’ or ’ehs have a different meaning in social science research (uk data service data catalogue) than in everyday language. meant is ’growing up in scotland’; ’cohort and longitudinal studies enhancement resources’; ’imf direction of trade statistics’; ’active lives survey’;’ english housing survey’; but not ’global university systems’ or a person’s name; ’more close’; a type of punctuation marks (namely ’...’); ’amyotrophic lateral sclerosis’; ’environment, health and safety’. when searching for well-known abbreviations and acronyms of datasets/studies and series, for example ’qlfs’, ’closer’, ’gbhd’ or ’was’, the search engine searches for both the acronym/abbreviation and the full term of the study or dataset/series in the relevant metadata fields. figure 1. shows the results for studies, when searching for ’closer’, sorted by ’relevance’. figure 1. moreover, implementing a standardized list of stopwords as an extendable blacklist prevents the retrieval of too many and unwanted results. it therefore improves recall/precision to a high extent. common stopwords are, for example, ’a, the, by, your, was, she, be, here’. an average list contains between 150 and 250 stopwords. an example: searching the ukds data catalogue for ’was’ only retrieves results connected to the ’wealth and assets survey,’ and excludes all instances of ’was’ as https://doi.org/10.29173/iq992 4/10 beeken, jeannine (2020) a recommendation to the ssh community: take a linguist on board, iassist quarterly 45(1), pp. 1-10. doi: https://doi.org/10.29173/iq992 used in, for example, ’it was’. consequently, the system retrieves correctly 21 studies in ukds data catalogue (see figure 2.), whereas without using a list of stopwords, the number of results would be 5000+ due to the appearance of the verb ’was’ in for example the abstracts of studies. figure 2. findability and retrieval are also improved by implementing language-specific stemmers and lemmatizers. these tools automatically search for all formal (form) and semantic (meaning) variants of a search term, i.e. all the terms that share the same stem or are variants of the canonical form as found in a dictionary. for example, a search for ’nurse’ equals a search for ’nurse, nurses, nursing’, but does not consider or retrieve studies for ’nurseries’. lemmatizers help with searches for ’good, better, best’, where not all variants share the same stem. consequently, searches for singular or plural terms and for gerunds yield the same number of results, for example, ’tax’ or ’taxes’; ’nurse’, ’nurses’ and ’nursing’. the sorting order of the results, however, will correspond with the specific search term. when searching for ’tax’, the results containing the singular form will be higher up in the ranking than the results with the plural form. when searching for ’taxes’, the results containing the plural form will appear first (see figures 3. and 4.). https://doi.org/10.29173/iq992 5/10 beeken, jeannine (2020) a recommendation to the ssh community: take a linguist on board, iassist quarterly 45(1), pp. 1-10. doi: https://doi.org/10.29173/iq992 figure 3. figure 4. https://doi.org/10.29173/iq992 6/10 beeken, jeannine (2020) a recommendation to the ssh community: take a linguist on board, iassist quarterly 45(1), pp. 1-10. doi: https://doi.org/10.29173/iq992 important tools such as electronic thesauri or ontologies and other types of controlled vocabularies and terminology provide an excellent extension and solution, for example, when a search for ’cat’ returns studies about ’felines’ (its broader concept), since the term ’cat’ was not used for indexing but ’felines’ was. improving findability and interoperability by implementing language-independent algorithms language-independent algorithms optimize the overall search to a great extent. an instruction such as ‘search for (“x y”) or (x and y)’, applied to, for example, ’energy prices’ searches automatically for both ’energy prices’ (i.e. the words must be adjacent, but can vary in order) and ’energy’ and ’prices’ (the words are not adjacent and can be in any order). the results list will display all studies for which the search terms are found at least once in the metadata record, more specifically, in the metadata fields in which the search is carried out. often, when results are found, they are displayed in a specific order according to, for example, ’most recently released’ or ‘relevance’. it is important to know which logic and algorithm sits behind the concept ’relevance’. at uk data service ‘relevance’ is facilitated by using a score resulting from hits within a small, relevant set of available or populated metadata fields (see table 1.), combined with a relative boosting weight (indicated in brackets) and a relative, proportional weight (a hit in a title of 3 words scores higher than a hit in a title of 10 words). in case of the same score, the results are ordered in descending order of version date. i.e. the most recent or newest first. title (50) country study number (15) geographical coverage abstract (10) spatial unit alternative title (10) town / village topic (10) other geography primary investigator (5) sampling procedure keyword (2) population data collector time period depositor time dimension sponsor kind of data grant number data type data producer language of study description series number language of study documentation subtitle data access tool type of key dataset table 1. metadata fields that are searched (with boost weight 50-1) for relevance ranking https://doi.org/10.29173/iq992 7/10 beeken, jeannine (2020) a recommendation to the ssh community: take a linguist on board, iassist quarterly 45(1), pp. 1-10. doi: https://doi.org/10.29173/iq992 improving text mining by adopting tools based on computational linguistics in this section, we focus on how linguistics, and more specifically computational linguistics, may contribute to and facilitate the mining and analysis of spoken and written data/material, i.e. qualitative data. for example, (pre-)processing interviews or oral histories for text-mining, i.e. preparing data for analysis and interpretation, may include the following steps and technology: • converting spoken data into written text data; parsing in order to group synonyms and multi-word expressions, sentence splitting; filtering by using stopword lists; discovering of patterns using frequency lists, concordances and correlations. • using automated speech recognition (asr) tools, which are able to distinguish spoken natural language (sound) from surrounding noise. they are able to translate or convert spoken language into written language (spelling, text), which can be aligned with the spoken fragments. the transcriptions can also include, for example, repetitions, incomplete sentences or onomatopoeia (e.g. ‘mmm’, ‘pfff’). • adopting complementary or advanced technology, which is able to distinguish different speakers in interviews, i.e. speaker diarisation, or to tag positive/negative emotions, for example ‘awful’ (negative) vs. ‘awfully nice’ (positive). multimodal technologies can also be used to investigate the importance of silences and role-taking in social interaction, tone and pitch, facial expressions and body language in audio-visual material. • enhanced speech-to-text transcription tools correctly assign capital letters useful for the recognition of named entities and abbreviations -, sentence splitting and punctuation useful for specifying questions or exclamations -, etc. as we will demonstrate below, a wide range of natural language processing (nlp) tools can indeed improve and offer help with the meaningful ‘human’ understanding or interpretation of texts. basic nlp can be adopted to enhance the quality of ‘machine’ generated frequency lists (which are currently not or only partially based on meaning/semantics), the creation of word clouds and term/keyword extraction. advanced nlp, such as syntactic parsers, help to disambiguate homophones like ‘friar, fryer’; ‘none, nun’; ‘knead, kneed, need’; ‘cense, cents, scents, sense’. they also detect co-referencing, identifying for example whether something refers to the same person or not, as in, for example, ‘warren arrived early this evening. the presidential candidate was accompanied by her daughter’. this type of information is important when mining and analysing texts using frequency lists (based on both meaning and form) or investigating and focusing on one and the same person. it also contributes to the reduction of disclosure risk, for example, re de-identification (direct) and anonymization (indirect). another type of nlp tools concerns information extraction. named entity recognizers (ner), for example, recognize and classify named entities, such as person names, organisations, locations or geospatial terminology, percentages, quantities, dates etc. the automatically generated and produced lists can be very useful for social science research, since they can assist with both simplifying anonymization and checking on or controlling disclosure and de-identification. https://doi.org/10.29173/iq992 8/10 beeken, jeannine (2020) a recommendation to the ssh community: take a linguist on board, iassist quarterly 45(1), pp. 1-10. doi: https://doi.org/10.29173/iq992 using extraction tools for the detection of keywords and controlled terminology, in this case social science terminology and jargon, may also improve the quality of human text mining and analysis, including (semi-automatic) indexing. as mentioned before, as an example, the relevant meaning of ‘was’ as used in the uk data service data catalogue is ‘wealth and assets survey’, and not the verb ‘was’ (which it would be for linguistic research concerning auxiliary verbs or passive constructions); ‘closer’ stands for a range of longitudinal studies; its meaning in everyday language ‘more close’ is rather irrelevant in this respect. as previously mentioned, keywords can be used to identify and describe the content of a study; they also improve information retrieval in terms of precision and recall, where precision is the result of dividing the number of true positives by the sum of all positives, and recall is the result of dividing the number of true positives by the sum of true positives and false negatives. (nb the ‘keyword’ metadata field has boosting factor 2, see above) the following examples illustrate keyword extraction and auto-summarisation applied to a text example from the uk data archive website (abstract copyright uk data service and data collection copyright owner): the commercial victimisation survey (cvs) provides a source of information on crime and crime-related issues as they affect businesses in england and wales. it provides additional detail on the extent of crime to be used alongside the other main sources of information on crime. these are the crime survey for england and wales (csew) (formerly the british crime survey), which covers crimes against private individuals and households, and the police recorded crime statistics, which cover crimes reported to the police. in common with the csew, the cvs also includes crimes that are not reported to the police. the police recorded crime data tables are available from the gov.uk website. the cvs was conducted in 1994, 2002, 2012, 2013, 2014, 2015, 2016 and 2017 (at present, the archive only holds data from 2002 onwards) and the survey has been commissioned to run in 2018. further information on the cvs, with links to findings by year, can also be found on the gov.uk crimes against businesses webpage. keywords: crime, information, england, cvs, csew, wales (http://keywordextraction.net/keyword-extractor) summary: the commercial victimisation survey (cvs) provides a source of information on crime and crime-related issues as they affect businesses in england and wales. these are the crime survey for england and wales (csew) (formerly the british crime survey), which covers crimes against private individuals and households, and the police recorded crime statistics, which cover crimes reported to the police. (https://summarygenerator.com/) a lot of nlp tools also generate concordances (kwic/keywords in context) and correlations in a text or group of texts, i.e. text corpus. this helps the human user to detect possible links and (unexpected) patterns, for example between ‘school’, ‘learning’, ‘teachers’ and ‘knives’. below is an example produced by sketchengine’s ‘concordance’ functionality, when searching for the word ‘travel’ in an interview with a black immigrant to the uk. the human user can easily detect that interviewee 1 travelled before, but interviewee 2 travelled for the first time abroad by boat and that she/he disembarked in southampton. an important fact here is that the answers from respondent 1 https://doi.org/10.29173/iq992 http://discover.ukdataservice.ac.uk/series/?sn=200009 https://www.gov.uk/government/publications/police-recorded-crime-open-data-tables https://www.gov.uk/government/publications/police-recorded-crime-open-data-tables https://www.gov.uk/government/collections/crime-against-businesses http://keywordextraction.net/keyword-extractor https://summarygenerator.com/ 9/10 beeken, jeannine (2020) a recommendation to the ssh community: take a linguist on board, iassist quarterly 45(1), pp. 1-10. doi: https://doi.org/10.29173/iq992 and respondent 2 have been identified and separated (see figure 5), a process similar to speaker diarisation in audio recordings. figure 5. the voyant os online tool, for example, offers a ‘correlations’ functionality. correlations are words or terminology that often appear in each other’s neighbourhood. searching for ’legal’ in a text about medicinal cannabis in california informs the human user that both the laws for the use of cannabis and the attitudes towards it have changed, since it became legal in the 1990s (see figure 6.). figure 6. concluding remark all of the tools and services described above are available for different languages. it is certainly worth considering to add these linguistics-based services and tools to the standard package of https://doi.org/10.29173/iq992 10/10 beeken, jeannine (2020) a recommendation to the ssh community: take a linguist on board, iassist quarterly 45(1), pp. 1-10. doi: https://doi.org/10.29173/iq992 services and infrastructure offered by archives to social science and humanities researchers, because they • promote and support multidisciplinary research and cooperation. • facilitate interoperability between research approaches and methods, technology and tools. • increase awareness of a wide variety of language technology tools which may assist or improve ssh research. • illustrate and demonstrate the potential and benefit of computational social science. • result in a better user experience with search, retrieval, extraction and analysis tools and create a better understanding, and therefore openness to unknown or lesser-known technology. references adam kilgarriff, vít baisa, jan bušta, miloš jakubíček, vojtěch kovář, jan michelfeit, pavel rychlý, vít suchomel. the sketch engine: ten years on. lexicography, 1: 7-36, 2014. [bibtex] [download pdf]. http://www.sketchengine.eu home office, crime and policing analysis unit. (2020). commercial victimisation survey, 2017. [data collection]. uk data service. sn: 8352, http://doi.org/10.5255/ukda-sn-8352-1 sinclair, stéfan and geoffrey rockwell, 2016. voyant tools. web. http://voyant-tools.org/. end-notes 1 jeannine beeken is senior metadata and ontologies officer at the uk data service, university of essex, uk. she can be reached by email: jeannine.beeken@essex.ac.uk https://doi.org/10.29173/iq992 javascript:void(0); https://www.sketchengine.eu/wp-content/uploads/the_sketch_engine_2014.pdf http://www.sketchengine.eu/ http://doi.org/10.5255/ukda-sn-8352-1 http://voyant-tools.org/ mailto:jeannine.beeken@essex.ac.uk 1/2 schwartz, ofira & hayslett, michele (2024), developing systems to encourage fair and secure research data, iassist quarterly 48(1), pp. 1-2. doi: https://doi.org/10.29173/iq1115 the creative commons-attribution-noncommercial license 4.0 international applies to all works published by iassist quarterly. authors will retain copyright of the work and full publishing rights. editors’ notes: developing systems to encourage fair and secure research data welcome to first issue of iassist quarterly for 2024; this is volume 48 of the journal (iq 48(1) 2024). the three papers in this issue represent different aspects of researchers’ use of data archives. experts from three major data archives (icpsr, ipums, and gesis) share their experiences developing systems to make research resources in their archives more findable, accessible, and usable. lafia, million, and hemphill in their article “exploratory and directed search strategies at a social science data archive” investigate search strategies of users of a large social science data archive. in an effort to understand how users search for curated research data, the authors analyze data queries issued through the inter-university consortium for political and social research (icpsr) website’s search box. the study is meant to inform archives and repositories of their users’ experience and help develop systems that encourage dataset findability, accessibility and reuse. the authors identify two types of data searches, exploratory and directed searches. they find that while users that issued exploratory queries were able to navigate the icpsr website successfully and refine their searches, they may benefit from more explicit support for query reformation. the authors suggest ways development of tools and training could improve users’ experience. in the article “stewarding our resources: building a sustainable ipums archival document access system” daina magnuson describes ipums experience building a web interface to support exploration and dissemination of archival materials associated with ipums international (ipums-i). ancillary materials, including thousands of unique pieces of census and survey documentation were obtained by ipums-i during data acquisition efforts. this rich source material was curated and preserved by archival staff. magnuson describes in details the development of a system that would make these resources findable, searchable and downloadable to internal users as well as ipums researchers. strategies describe may inform other data archives considering developing similar tools. deborah wiltshire in her article “developing canonical ‘safe researcher’ training materials for trusted research environment” describes the development of training materials for researchers applying to access restricted data through trusted research environment (tre). the increased use of a virtual desktop environment, as opposed to physical safe rooms, offers researchers easier and more flexible access to restricted data; however, it offers fewer physical safeguards. attending a mandatory ‘safe researcher’ training is one of the principles of the five safe framework1 developed by the uk’s office for national statistics to bridge this gap and assure the safe use of sensitive data. the process described in the article could be easily adapted by other tres according to their needs and settings (e.g., in-person vs. online training). https://doi.org/10.29173/iq1115 https://creativecommons.org/licenses/by-nc/4.0/ 2/2 schwartz, ofira & hayslett, michele (2024), developing systems to encourage fair and secure research data, iassist quarterly 48(1), pp. 1-2. doi: https://doi.org/10.29173/iq1115 in the spirit of making information findable, accessible and reusable, we’re happy to share that the open journal system platform from the university of alberta has enabled author linking to orcid. this means that iq authors are now able to connect their orcid ids to their iq publications, so that their orcid profiles will automatically update with those citations as their articles are published. for existing iq account holders, log into your account at the upper right of the iq page, and go to profile > public. then click “create or connect your orcid id” (just below the homepage url field). for details, see this video walkthrough of profile authentication. new iq users will automatically be prompted to link (or create) their orcid ids during iq account creation. authors are also encouraged to note their orcid id link with their affiliation and contact information in an endnote when submitting a manuscript for publication. we hope to see many of you at the 49th annual iassist conference in halifax, nova scotia, canada, may 28-31, 2024! ofira schwartz and michele hayslett, march 2024 1 desai, t., ritchie, f. and welpton r. (2016) ‘five safes: designing data access for research’. economics working paper series 1601, pp. 1-27. available at: https://www2.uwe.ac.uk/faculties/bbs/documents/1601.pdf https://doi.org/10.29173/iq1115 https://nam12.safelinks.protection.outlook.com/?url=https%3a%2f%2fiassistquarterly.com%2f&data=05%7c02%7coschwart%40princeton.edu%7cb9306cab5b654a82c5e908dc44560a19%7c2ff601167431425db5af077d7791bda4%7c0%7c0%7c638460383536191103%7cunknown%7ctwfpbgzsb3d8eyjwijoimc4wljawmdailcjqijoiv2lumziilcjbtii6ik1hawwilcjxvci6mn0%3d%7c0%7c%7c%7c&sdata=l4xykzmqq88kydhyupguys59a1x9ecrumnxyyrnhk5c%3d&reserved=0 https://nam12.safelinks.protection.outlook.com/?url=https%3a%2f%2fvimeo.com%2f374415404&data=05%7c02%7coschwart%40princeton.edu%7cb9306cab5b654a82c5e908dc44560a19%7c2ff601167431425db5af077d7791bda4%7c0%7c0%7c638460383536201232%7cunknown%7ctwfpbgzsb3d8eyjwijoimc4wljawmdailcjqijoiv2lumziilcjbtii6ik1hawwilcjxvci6mn0%3d%7c0%7c%7c%7c&sdata=6t9u%2bc8f3a5aebo0mzrvffjzszjssb83twkkrj8ub6g%3d&reserved=0 https://nam12.safelinks.protection.outlook.com/?url=https%3a%2f%2fvimeo.com%2f374415404&data=05%7c02%7coschwart%40princeton.edu%7cb9306cab5b654a82c5e908dc44560a19%7c2ff601167431425db5af077d7791bda4%7c0%7c0%7c638460383536201232%7cunknown%7ctwfpbgzsb3d8eyjwijoimc4wljawmdailcjqijoiv2lumziilcjbtii6ik1hawwilcjxvci6mn0%3d%7c0%7c%7c%7c&sdata=6t9u%2bc8f3a5aebo0mzrvffjzszjssb83twkkrj8ub6g%3d&reserved=0 https://www2.uwe.ac.uk/faculties/bbs/documents/1601.pdf by 31 iassist quarterly winter/spring 2010 the american community survey: benefits and challenges abstract in the united states' decennial census, all persons living in the us are asked to fill out a short form asking basic questions such as age, race, and number of people living in a housing unit. in addition to the short form, starting in 1960 a sample of housing units were asked to fill out a long form with both the basic demographic questions plus questions about socioeconomic topics, such as education, income, housing characteristics and more. in 2010 the united states will conduct its constitutionally mandated census of the population, but a major change will occur. the long form will no longer be distributed and in its place will be the american community survey (acs). this article discusses the development of the survey and its benefits and challenges. the acs will provide researchers and policymakers more timely information of the characteristics of areas. nevertheless, there are still some questions and concerns about how to use the data and challenges for the implementation of the survey. keywords: population, demographics, census, socioeconomics every ten years the united states is required by its constitution to conduct a census of the population. article 1, section 2 of the constitution of the united states maintains that: representatives and direct taxes shall be apportioned among the several states which may be included within this union, according to their respective numbers...the actual enumeration shall be made within three years after the first meeting of the congress of the united states, and within every subsequent term of ten years, in such manner as they shall by law direct.2 in the united states’ decennial census, all persons living in the us are asked to fill out a short form asking basic questions such as age, race, and number of people living in a housing unit. these demographic data are used for the apportionment of congressional seats. in addition by michele hayslett and lynda kellam1 to the short form, starting in 1960 a sample of housing units were asked to fill out a long form with both the basic demographic questions plus questions about socioeconomic topics, such as education, income, housing characteristics and more. although this sample survey is not constitutionally mandated, it serves an essential function for policy-makers and planners. in 2010 the united states will conduct its constitutionally mandated census of the population, but a major change will occur. the long form will no longer be distributed and in its place will be the american community survey (acs). this article will discuss the development of the survey and its benefits and challenges. the acs will provide researchers and policymakers more timely information of the characteristics of areas. nevertheless, there are still some questions and concerns about how to use the data and challenges for the implementation of the survey development and design of the american community survey efforts to create the american community survey began in 1996 when the survey was launched at four test sites. with the 2000 census, a test form of the acs was conducted as the census 2000 supplement survey (c2ss) and was launched in 1,200 counties. the purpose was to test “the feasibility of collecting acs statistics in a decennial census year.” (herman, 2008) full nationwide implementation of the acs began in 2005 except for group quarters data which began in 2006. the acs is a self-enumeration survey with questionnaires sent by mail to chosen survey households. enumerators conduct follow up telephone calls and visits to addresses that have not mailed in their questionnaires. approximately 250,000 addresses receive a questionnaire each month totaling about 3 million households each year, resulting in a sample size of approximately one in eight households. the costs of conducting a monthly survey prevent an increase in the sample size to match the census long form sample size. because of the smaller sample size and because the sample is accumulated progressively over time, the release of the data is tiered based on the size of geographic areas (mather, rivers, jacobsen, 2005). hence estimates for geographies 32 iassist quarterly winter/spring 2010 with larger populations (more than 65,000 people) can be calculated on the sample accumulated within just one year, but estimates for geographies with smaller populations must wait to be calculated until three years or five years of data have been collected. acs data prior to 2005 are available for geographies with 250,000 people or more and are considered test data. in 2006, the census bureau published acs data collected in 2005 for geographies with at least 65,000 people and data for these large geographies with over 65,000 people will be available on an annual basis. for a geographic area with between 20,000 and 65,000 people, three year estimates first became available in december 2008 using data collected from 2005 to 2007. in 2009, the three year estimates for 2007 through 2009 were released. for a geographic area with fewer than 20,000 people, a five year estimate will be required. the first five year estimates for the period 2005-2009 will begin to be released in late 2010 (see figure 1). as with the long-form sample in the decennial census, the acs is sample survey data and will have margins of error and confidence intervals. the census bureau maintains that the estimates are within the range of a 90% confidence interval. for example, in the 2005-2007 three-year estimate the population of greensboro, nc is 237,423 with the margin of error of +/-2,958. this statement tells us that the census bureau is 90% certain that the population of greensboro is between 234,465 and 240,381. another defining characteristic of the acs is the collection of data over a period of time. this is in direct contrast to the data collection for the decennial census. whereas the decennial census has a reference point of april 1 for determining residency, the acs’s reference period for residency varies depending on the month in which the specific household receives the questionnaire. this has numerous effects on understanding data related to specific figure 1: acs release dates reference periods especially employment, income, and school enrollment. although the acs replaces the long form in the conduct of the census, these differences related to reference periods affect the comparability of acs data to decennial long-form data. benefits the census bureau developed the acs in response to users' demands for more timely data. although the supreme court determined that only 100% data can be used for apportionment of congressional seats, planners and policy makers needed more frequent data releases to make better decisions and determine whether programs were successful and working as intended. thus, the immediate benefit of acs data is that it is collected every year and released the following year. businesses, government agencies at all levels and the public will no longer have to wait ten years to find out how the country and local communities have grown and whether planning and public policy is meeting people's needs. no longer will they have to wait two to three years after the decennial census is taken for data on income, education and housing characteristics to be released. moreover, because the survey is run every year, the data provide a way to track rapidly changing community trends and the opportunity to change data collection to respond to current events, including natural disasters like hurricanes and forest fires, and economic crises. the census bureau believes that, despite having a smaller sample size, the acs will actually provide more accurate data than the decennial long form for two reasons. first, because the acs is being run constantly, a professional staff has been hired on a permanent basis to work in local areas. instead of having to hire a huge number of temporary, non-professional staff who have to be trained in a very short period of time, this permanent staff will gain deeper experience and local knowledge over time that will improve data collection. for instance, issues such as reaching non-english speaking groups will become easier to address since these long-term staff will either be members of those communities themselves or able to develop relationships with leaders in those communities. this is also the reason the census bureau cites for the acs saving money over the decennial long form, that it is more cost effective to maintain a smaller collection and processing staff throughout the decade than to hire and train a much larger number of workers once a decade3. second, the non-response follow-up procedures for acs are more extensive than those of the decennial long form, including telephone contacts as well as in-person visits (u.s. census bureau 2008b, 82). as an example, “a comparison between acs and census 2000 data for the bronx showed that while the census 2000 had a higher initial mail response rate than the acs, it was less effective than the acs during 33 iassist quarterly winter/spring 2010 follow-up phases, when information is collected from nonrespondents” (u.s. census bureau 2008c, 8). the acs is also capable of producing some data that the decennial census was not. the decennial long form asked people to answer questions based on their "usual residence" defined as "the place where the person lives and sleeps most of the time" (u.s. census bureau 2007, c-1). if someone received a form at an address where they did not live most of the time, the form would indicate they should only fill out a form for their usual address. consequently the decennial census had no mechanism for counting temporary populations like people who live in florida in the winter months or people who live in the northern states in the summer. the acs, however, counts people at their "current residence," defined as "everyone who is currently living or staying at a sample address…except for those staying there for…less than two consecutive months" (u.s. census bureau 2009a, 6-1). moreover, the counting goes on year-round instead of on one day, so the acs is able to account for temporary residents regardless of season. areas that have significant seasonal migrant worker populations will also notice higher acs counts versus the decennial long form figures since the year-round data collection will better account for such groups. some researchers have stated concern about the comparability of school enrollment data since the acs will collect data in the summer months when children are not in school (gage 2006, 247). however, at least as far back as 2005, the questionnaires have been worded to ask whether children have been enrolled "in the last three months." consequently, the time of year when a respondent receives the survey should not matter for this variable. overall, researchers are beginning to appreciate the advantages in the acs data over the decennial long form data. in a 2006 study, gage found, after graphing multiple variables for two california counties: in most cases, even when statistical tests identified differences [from decennial long form data] as significant, the acs data generally appeared useful and usable. simply observing a statistically significant difference provides no guidance as to which data are better….for practical purposes it appears that most of the acs data could, on an annual basis, be used in place of the census data and should provide a more current measurement, especially as the census count ages and remains static throughout the decade (247). however, challenges still abound, particularly for new users. challenges and how to meet them the degree of difference between the methodology of the acs and the decennial long form survey results in a number of notable challenges for data users who want to do time series analysis. essentially, the two surveys are not comparable. the simple cost of running the survey every year results in a significant compromise: the sample size of the acs is decidedly smaller than that of the decennial long form. griffin and waite from the census bureau argue that the “estimates of sampling error for the five-year acs estimates will be about one-third higher than those from decennial census estimates” (2006, 216), but they maintain that “this is acceptable given the reduction of bias due to timeliness and the potential for reductions in nonsampling errors because of factors such as the use of automated instruments and experienced interviewers” (2006, 216). it must be emphasized that the census bureau's goal with the acs is not to produce a population count but rather to produce an estimate of the characteristics of the population. that is, this data will be less useful than the census long form for noting the absolute numbers of the population but very useful for studying trends over time. this is an important distinction because, with its smaller figure 2: ranking table – percent of people with a disability (note option on left to view as a chart) 34 iassist quarterly winter/spring 2010 sample size, the acs is not very useful for pin-pointing exact numbers. the smaller sample size is also the reason the census bureau is publishing the confidence intervals (cis) for each estimate with the acs, to demonstrate the accuracy of each figure. while the numbers which appeared in the decennial long form were also estimates, the sample size was sufficient that the bureau didn't feel the need to emphasize the cis. unfortunately one of the results of this was users came to see the long form numbers as actual counts rather than the calculated figures they really were. with the acs, particularly for smaller geographies and smaller groups of population (by race or income, etc.), the cis can be quite large despite targeted over-sampling to off-set this problem. hence, it's more important for the bureau to highlight them and explain what they mean. the state ranking tables offer a visual display of the cis that is very helpful once one understands how to interpret them (see figure 2). figure 2 shows a ranking table for percent of people in each state with a disability. you will see on the left side of the page there is an option to view as a chart. figure 3 shows the chart. the red dots represent the estimates while the blue lines bracketing each dot represent the cis. while the estimate will always be shown in these figure 3: ranking chart – percent of people with a disability (note overlap of cis) charts as the center of the ci, technically the definition of a confidence interval is that the true value may be anywhere within the ci. so when a user looks at the chart, anywhere the estimates' cis overlap, technically the actual values for those states might be the same. so the order of those rankings might be considered ties, or even be reversed. this illustrates why the data should be used for tracing trends, not as absolute numbers. another effect of the very different methodology is that many acs variables are not comparable to decennial longform ones, even when they have the same name. novice users will almost certainly be tempted to make direct comparisons without realizing they are trying to compare apples and oranges instead of apples to apples. for example, while the decennial census is taken on a single date, the acs is a rolling survey, with responses collected every month of the year. consequently rather than try to ask about respondents' income "last year" as the decennial does, respondents will be asked to provide their income during the twelve months prior to the date they receive the survey. this is likely to be a challenge for respondents to even answer. on the decennial census date, april 1st, most u.s. respondents are working on or have finished their federal income tax returns (due april 15th) and can easily cite their previous calendar year's income. citing the previous twelve months' income for the acs, however, will require some figuring, especially if the given twelve-month period encompasses a change in rate of pay, or commission income that varies from month to month. users of the final estimates are likely to think "last year's income" is essentially the same as the "income of the last twelve months" without realizing the significance of the different reference periods involved. the acs's reference period is also an example of why this data should be used for trend analysis rather than point-in-time exact figures. related to the rolling nature of the survey, another feature of the acs methodology that will complicate use is that the survey draws on data from multiple years to accumulate a sample size large enough to create estimates for smaller geographies. geographies with populations between 20,000 and 65,000 will have estimates based on the average of three years of data, while geographies with populations less than 20,000 will 35 iassist quarterly winter/spring 2010 have estimates based on the average of five years of data. because the distinction is based on population totals, large cities will have single-year estimates while small towns will have threeor five-year estimates—and the three different levels are not comparable to each other. instead, in addition to the one-year estimates, the census bureau is making available averaged-year estimates for larger geographies that should be used for comparisons to smaller geographies. for example, to compare the state of north carolina and the city of charlotte, one can use one-year estimates for both since the population of each exceeds 65,000 people. however if one were researching the city of kannapolis, its population was 36,699 in the 2000 decennial census. because this falls between 65,000 and 20,000 people, the acs will only provide estimates based on the average of three years of data in order to have enough respondents in the sample to create accurate estimates for the size place it is. in this case, to compare kannapolis with the state, one would need to use north carolina's three-year averaged data instead of its one-year estimate. likewise to compare north carolina and mount airy, a town of 8,460 in the decennial census, one would need to use the state's fiveyear averaged data since mount airy will only have acs estimates based on the average of five years of data. data for the smallest geographies, all those with less than 20,000 people (including all census tracts and block groups), have not yet been released. the american community survey began full-scale data production in 2005 (with the exception of group quarters data which was added in 20064 ), so until it has had five full years of data collection, the pool of respondents will not be large enough to create the estimates for the smallest geographies. with an extra year for processing time, the census bureau will not release five-year averaged estimates until close to the end of the 2010 calendar year. consequently for a while yet data users will be frustrated when trying to find data on small places or rural areas. however, this issue will disappear entirely once the first five-year estimates are released since five-year estimates will be available every year thereafter. knowing which estimate to use for larger geographies and how to explain the use of different figures in context when writing a grant proposal, for instance, will be a particularly difficult issue for novice users. another issue related to sample size is the suppression of data. in the decennial census, the census bureau employs thresholds below which data for very specific occupations or population groups will not be published in order to protect confidentiality. because of the bureau’s confidence in the sample size based on five years’ worth of data, the acs does not employ such thresholds. instead, staff tests for the statistical reliability of the oneand three-year estimates and suppress tables when at least 50 percent of the included estimates (that is, cells within the table) fail the coefficient of variation test. the bureau states that the five-year estimates will not be tested at all since the sample size based on five years of data will ensure viable estimates. an example of this might be detailed race breakdowns in a rural state, especially ones that tend to be more homogenous racially. for example, montana might figure 4: base table, b02003. race universe: total population 36 iassist quarterly winter/spring 2010 be more likely to be suppressed for this variable than north carolina. one way the census bureau handles this is by the production of base versus compressed tables. base tables provide all the detail users are used to seeing in the decennial long form data. but for a table that is likely to be suppressed because more than half of its cells fail the statistical test, the bureau may produce a compressed, or c, table for the same subject. figures 4 and 5 demonstrate the difference between the two. figure 4 is the base (or b) table for race from the 2008 1-year estimates. it was necessary to run this report for the country as a whole—even the most populous states were suppressed. this is understandable when one considers how detailed the categories are for "population of two races," with fifteen different race combinations including, for instance, one for those respondents who indicated they were both american indian/alaska native and native hawaiian/other pacific islander. in figure 5, the c table figure 5: compressed table, c02003. race universe: total population for the same variable, you can see that the two or more race categories have been severely compressed to the four most commonly chosen categories and one titled "all other two race combinations." here it was possible to generate data for states at both ends of the population spectrum as well as for the nation. users familiar with the p(opulation) and h(ousing) tables of the decennial data will easily translate to the acs system of labeling tables b(ase) or c(ompressed) in the title, as noted in these figures. another method the bureau recommends5 for ameliorating estimates with very large margins of error (moes) is to combine several geographies or several variable categories. the method is straightforward: one simply sums the geographies or categories to create a larger “sample.” of course, it can only be used for straight summed data like population, race, sex, etc.; it cannot be used with calculated values such as medians. then to calculate the new moe for this new “estimate,” one squares each original moe 2005 2006 2007 2008 2009 2010 single year estimates 20.0 21.2 23.3 28.6 32.6 35.1 3-year estimates (2005-2007) 21.521.521.5 3-year estimates (2006-2008) 24.824.824.8 3-year estimates (2007-2009) 28.628.628.6 5-year estimates (2005-2009) 25.925.925.925.925.9 3-year estimates (2008-2010) 32.232.232.2 5-year estimates (2006-2010) 28.928.928.928.928.9 table 1. estimates for a geography with more than 65,000 people 37 iassist quarterly winter/spring 2010 and adds them together, then takes the square root of that sum, or √ (moe2 + moe2 + moe2). of course, this method needs to be used with some care. reliable data will not result from combining geographically distant geographies. geographies at the same summary level (e.g., tracts combined with tracts or counties combined with counties) that border one another and have similar characteristics to the one under examination are to be strongly preferred. to return to the issue of estimates based on the averaged data of several years, we can better understand the effect of averaging several years of data by considering two 2005 2006 2007 2008 2009 2010 single year estimates 20.0 21.2 23.3 28.6 32.6 35.1 3-year estimates (2005-2007) 21.521.521.5 3-year estimates (2006-2008) 24.824.824.8 3-year estimates (2007-2009) 28.628.628.6 5-year estimates (2005-2009) 25.925.925.925.925.9 3-year estimates (2008-2010) 32.232.232.2 5-year estimates (2006-2010) 28.928.928.928.928.9 2005 2006 2007 2008 2009 2010 3-year estimates (2005-2007) 21.521.521.5 3-year estimates (2006-2008) 24.824.824.8 3-year estimates (2007-2009) 28.628.628.6 5-year estimates (2005-2009) 25.925.925.925.925.9 3-year estimates (2008-2010) 32.232.232.2 5-year estimates (2006-2010) 28.928.928.928.928.9 2005 2006 2007 2008 2009 2010 5-year estimates (2005-2009) 25.925.925.925.925.9 5-year estimates (2006-2010) 28.928.928.928.928.9 table 2. estimates for a geography with more than 20,000 and less than 65,000 people table 3. estimates for a geography with less than 20,000 people 2005 2006 2007 2008 2009 2010 single year estimates 20.0 21.2 23.3 28.6 32.6 35.1 3-year estimates (2005-2007) 21.521.521.5 3-year estimates (2006-2008) 24.824.824.8 3-year estimates (2007-2009) 28.628.628.6 5-year estimates (2005-2009) 25.925.925.925.925.9 3-year estimates (2008-2010) 32.232.232.2 5-year estimates (2006-2010) 28.928.928.928.928.9 2005 2006 2007 2008 2009 2010 3-year estimates (2005-2007) 21.521.521.5 3-year estimates (2006-2008) 24.824.824.8 3-year estimates (2007-2009) 28.628.628.6 5-year estimates (2005-2009) 25.925.925.925.925.9 3-year estimates (2008-2010) 32.232.232.2 5-year estimates (2006-2010) 28.928.928.928.928.9 2005 2006 2007 2008 2009 2010 5-year estimates (2005-2009) 25.925.925.925.925.9 5-year estimates (2006-2010) 28.928.928.928.928.9 figure 6: example 1: item with year-to-year increases (foreign-born population) examples presented by deborah griffin and her colleagues at the state data center/business and industry data center annual national training conference in 2004. in the first example, values are steadily increasing over time—the percentage of foreign-born population might be such a variable. tables 1 through 3 show hypothetical values for such a variable. (for simplicity's sake in this example, the values across the geographies are the same, although this would probably seldom be true in actuality.) consider that an average is a measure of the middle. consequently when one averages data to create an estimate, the estimate will tend more toward the middle of the figures averaged. figure 6 shows how the averaged estimates' values will follow the trend of the annual estimates in cases of steady figure 7: example 2: item with year-to-year increases and decreases (home ownership rates) 38 iassist quarterly winter/spring 2010 increases but will tend to lag slightly behind it. this would also be the case with steady decreases in values. however, what happens when the trend fluctuates? figure 7 shows such an example, a hypothetical view of home ownership rates. here you can see that the averaged estimates tend toward the middle of the varying values, describing a trend that smoothes the highs and lows to a flatter line. this is perhaps the biggest disadvantage of the acs methodology and it is not so much an issue of accuracy as of precision. researchers familiar with the decennial long-form data will miss that survey's ability to provide (essentially) one-year estimates for all geographies. some users of the acs data have even indicated that one should not use overlapping estimates; in other words if one uses a 2004-2006 threeyear estimate, it would be better to wait for the 2007-2009 data with which to compare the same geography. however, the census bureau would again assert that the acs data is best used to understand trends, and that the 2005-2007 three-year estimate will provide an update to the 2004-2006 one, even if much of the pool of respondents remains the same. geographic boundaries of the most recent year in multiyear averaged estimates apply. to do this, acs staff re-create the earlier years' estimates with the current year's geographic boundaries in order to include respondents for the new geography for all years of the average. also, dollar values for earlier years of an average are inflation-adjusted to the most recent year. (griffin, et al., 2004) in a year that a small-sized geography crosses the threshold to the next size (i.e., from less than 20,000 to between 20,000 and 65,000) it will begin to have three-year averaged estimates produced as well as five-year estimates. likewise, in a year when a medium-sized geography crosses the threshold to the large size (i.e., from between 20,000 and 65,000 to over 65,000) it will begin to have single-year estimates produced as well as threeand five-year. the reverse is also true. if a geography loses population and drops below the threshold, it will lose the estimates of the larger category—a place dropping below 65,000 would lose the single-year estimates and a place dropping below 20,000 would lose the three-year estimates. the future how specifically the surveys are able to describe a community has always been at the forefront of the challenges the census bureau faces. protecting confidentiality is of paramount importance, punishable by fines and imprisonment. yet the american public demands the smallest level of geography possible for both political and economic planning reasons. officials at the bureau recognize that striving for this level of detail is costly. at a hearing of the congressional joint economic committee, former census bureau directors louis kincannon and kenneth prewitt both testified that even for the decennial census, data at the block level is unnecessary for the purposes of redistricting and dropping the smaller geographic levels would significantly cut costs. (2009, timestamp 77:15) with follow-up questioning, prewitt stated that data at the census tract level would provide sufficient detail (2009, timestamp 98:18). conclusions the best preparation for understanding a community's acs figures is to know the community very well. local knowledge will help researchers identify when the acs data are incorrect or insufficient. where researchers are not familiar with local communities, they must carefully attend to the moes and decide when the data are sufficient to the research purpose at hand and when they are not. for novice users, guidance on using the acs is critical. librarians need to be on-hand in academic and public libraries to assist users with both navigating the american factfinder interface and understanding the acs data. the census bureau is fully aware of how difficult the acs is to use, particularly for novices, and it works constantly to make tools available to assist with it, including extensive technical documentation, guidance on making comparisons between different editions of acs data, and compass handbooks customized for different audiences. there is also an e-tutorial to assist novice acs data users. this suite of tools is available on the acs's how to use the data web site at http://www.census.gov/acs. while learning to use the acs will take some effort, it is imperative to do so. the long form on the decennial census will not return and the acs will remain the best data available. references gage, l. (2006). comparison of census 2000 and american community survey 1999-2001 estimates: san francisco and tulare counties, california. population research and policy review, 25(3), 243-256. griffin, d., hubble, d., love, s., & mcginn, l. (2004). american community survey technical training. presented at the state data center/business and industry data center annual national training conference, september 21, 2004. griffin, d. and waite, p. (2006). american community survey overview and the role of external evaluations. population research and policy review 25(3), 201-223. herman, e. (2008). the american community survey: an introduction to the basics. government information quarterly 25(3), 504-519. mather, m., rivers, k. & jacobsen, l. (2005) “the american community survey,” population bulletin 60(3), 1 united states census bureau. (2006). american community survey 2006 subject definitions. washington: 39 iassist quarterly winter/spring 2010 u.s. census bureau. http://www.census.gov/acs/www/ downloads/2006/usedata/2006%20acs%20subject%20 definitions.pdf united states census bureau. (2007). 2000 census of population and housing, summary file 3 technical documentation, appendix c. data collection and processing procedures. washington: u.s. census bureau. http://www.census.gov/prod/cen2000/doc/sf3.pdf. united states census bureau. (2008a). the american community survey (acs) mail questionnaire from 2005 to 2009. washington: u.s. census bureau. http://www. census.gov/acs/www/downloads/acs%20mail%20 questionnaire%20(2009).pdf. united states census bureau. (2008b). 2007 american community survey technical documentation, appendix c. data collection and processing procedures. washington: u.s. census bureau. http://www2.census.gov/ acs2007_1yr/summaryfile/acs_2007_sf_tech_doc.pdf. united states census bureau. (2008c). a compass for understanding and using american community survey data: what general data users need to know. washington: u.s. census bureau. http://www.census.gov/acs/www/ downloads/acsgeneralhandbook.pdf. united states census bureau. (2009a). chapter 6: survey rules, concepts, and definitions. in acs design and methodology. washington: u.s. census bureau. http:// www.census.gov/acs/www/sbasics/desgn_meth.htm. united states census bureau. (2009b). toolkit for members of the house and senate, sample news release. washington: u.s. census bureau. http://2010.census.gov/ partners/pdf/congressnewsrelease.doc. united states congress, joint economic committee. (2009). the federal statistical system in the 21st century: the role of the census bureau. [video of hearing] washington: u.s. congress, joint economic committee. http://jec.senate.gov/index.cfm?fuseaction=hearings. hearingscalendar&contentrecord_id=89ad0feb-50568059-7698-744705750411®ion_id=&issue_id=. notes: 1 authors are michele hayslett, university of north carolina at chapel hill, michele_hayslett@unc.edu, and lynda kellam, university of north carolina at greensboro, lmkellam@uncg.edu. 2 national archives and records administration. the constitution of the united states. the charters of freedom: a new world is at hand. <http://www.archives.gov/exhibits/ charters/constitution_transcript.html>. accessed october 19, 2010. 3 personal communication (by telephone and email) with bob coats, north carolina's liaison to the governor for the census on august 6, 2009. the bureau has described decennial census-taking as the "largest peace-time mobilization of personnel in u.s. history" (u.s. census bureau 2009b, 2). 4 users should be aware that acs data prior to 2006 does not include the population in group quarters. from the acs 2006 subject definitions: "this change in universe may affect the distribution of characteristics in areas where a significant proportion of the population lives in group quarters." (1) see united state census bureau. 2006. acs 2006 subject definitions. <http://www. census.gov/acs/www/downloads/data_documentation/ subjectdefinitions/2006_acssubjectdefinitions.pdf>. accessed october 19, 2010. 5 personal communication with kelly karres at the north carolina state data center annual meeting, raleigh, nc, august 23, 2010. 6 http://www.census.gov/acs/www/guidance_for_data_ users/e_tutorial/ 1/26 antognoli, erin; avila, regina; sears, jonathan; christiansen, leighton; tieman, jessica; hart, jacquelyn (2020) reproducibility literature analysis a federal information professional perspective, iassist quarterly 44(1-2), pp. 1-26. doi: https://doi.org/10.29173/iq967 reproducibility literature analysis a federal information professional perspective erin antognoli1, regina avila2, jonathan sears3, leighton christiansen4, jessica tieman5, jacquelyn hart6 abstract this article examines a cross-section of literature and other resources to reveal common reproducibility issues faced by stakeholders regardless of subject area or focus. we identify a variety of issues named as reproducibility barriers, the solutions to such barriers, and reflect on how researchers and information professionals can act to address the ‘reproducibility crisis.’ the finished products of this work include an annotated list of 122 published resources and a primer that identifies and defines key concepts from the resources that contribute to the crisis. keywords reproducibility, reproducibility crisis, replicability, research data, landscape analysis, culture shift introduction over the last number of years, the terms ‘reproducibility’ and ‘replicability’ have left the realm of science and become widely discussed in mainstream media. books, blogs, and television news shows have talked about the emergence of a ‘reproducibility crisis’ which brings the validity of scientific research into question. numerous studies and reports identifying the status of research reproducibility reveal problems at all levels of study across multiple research disciplines. even published research from prominent journals and institutions suffer from reproducibility issues (weir, 2015). while some question the idea that the issues surrounding reproducibility constitute a ‘crisis’ (baker, 2016c), evidence points to a widespread difficulty to reproduce published scientific results. as data managers and information professionals in u.s. federal libraries working in a variety of disciplines and backgrounds, we understand the difficulties surrounding this topic. in order to address these concerns, we formed a team and embarked on a project to identify the ‘crisis’ and the ongoing challenge it creates for ourselves and our stakeholders. our operating definitions of ‘reproducibility’ and ‘replication’ were as follows: ● reproducibility measures whether a study or experiment can be reproduced in its entirety. to achieve adequate reproducibility, studies implement measures to support verification of research, including, for example, sharing data and methods. no single factor or method alone achieves reproducibility in a study, and likewise, many factors can result in a study with poor reproducibility (munafò et al., 2017). https://doi.org/10.29173/iq967 2/26 antognoli, erin; avila, regina; sears, jonathan; christiansen, leighton; tieman, jessica; hart, jacquelyn (2020) reproducibility literature analysis a federal information professional perspective, iassist quarterly 44(1-2), pp. 1-26. doi: https://doi.org/10.29173/iq967 ● replication is the attempt to recreate the conditions believed sufficient for obtaining a previously observed finding and is the means of establishing reproducibility of a finding with new data (open science collaboration, 2015). it should be noted that formal definitions of the two terms ‘reproducibility’ and ‘replicability’ were offered in a report by the national academies of sciences, engineering, and medicine (2019). the distinct differences between the two definitions did not play a key part in this project. the academies’ definitions were published after most of our own investigation was complete, and the terms were deemed somewhat interchangeable throughout this exercise. to fully understand the reproducibility crisis, we must understand the multitude of contributing factors influencing reproducibility (open science collaboration, 2015). we gathered a number of resources and began studying. our findings resulted in two products discussed in this paper: an annotated list of 122 resources reviewed to understand the ‘crisis,’ and a primer that lists the top issues and solutions defined within the resources. these items provided our team a common language and understanding of the problem which we can use in our profession moving forward. project origins our journey started with a proposal brought up within cendi7, a u.s. federal scientific and technical information managers group. cendi is ‘a volunteer-powered membership organization that serves the federal information community that is, all those who create, manage, aggregate, organize, and provide access to federally-funded data and publications’ within federal scientific and technical information agencies. its member organizations represent a cross-section of federal data and publication stakeholders—including libraries, data centers, aggregators, information technology developers, and content management providers. cendi’s mission is to ‘increase the impact of federally funded science and technology by improving the management and dissemination of data and information’ (cendi, 2019). cendi is home to a small number of working groups, including the data curation discussion group (dcdg). while principle cendi members are the managers of federal scientific and technological information libraries, the dcdg members are, in the main, hands-on data management and curation staff within the libraries and home agencies. the dcdg’s goal is ‘to collaborate across agencies, employ data curation best practices, tools, and workflows, promote efficiencies and consistency, work through challenges, and avoid ‘reinvention.’’ (christiansen, 2017). in late 2017, cendi leadership proposed that dcdg develop tools to assess and address the ‘reproducibility crisis’ with possible outcomes being: ● developing or populating a website with content that puts reproducibility challenges in context and identifies both real issues and spurious concerns ● sharing information about approaches to reproducibility among the cendi members ● disseminating information about best practices within the respective agencies, based on consensus findings in cendi and/or noted elsewhere https://doi.org/10.29173/iq967 3/26 antognoli, erin; avila, regina; sears, jonathan; christiansen, leighton; tieman, jessica; hart, jacquelyn (2020) reproducibility literature analysis a federal information professional perspective, iassist quarterly 44(1-2), pp. 1-26. doi: https://doi.org/10.29173/iq967 project goals and methodology the dcdg began the project over the first half of 2018 as an effort to familiarize cendi members and interested parties on the topics and issues surrounding reproducibility. a subset of dcdg members formed a team that collected and annotated resources on various aspects of scientific reproducibility and replicability. the resources the team reviewed were primarily articles appearing in top results from a google scholar search, as well as resources cited within those initial results, dating from 2005 to the present. a number of other relevant resources such as books, presentations, and websites were also included. the team divided the list of resources and each member was assigned as a reader who offered an annotation or abstract of the resource. they also identified the top three issues and/or solutions shared within each. these issues and solutions populated separate spreadsheets and given definitions based on the literature. the goals of this exercise were to: 1. identify the variety of issues named as barriers to reproducibility 2. identify solutions to such barriers 3. reflect on how researchers, information professionals, and librarians perceive the ‘reproducibility crisis’ although this was not a formal analysis of the literature, or a refined scientific experiment, this information gathering exercise served to inform our group about the reproducibility crisis. the result of dcdg’s work included here is an annotated list of 122 resources which were reviewed, and a primer which identifies and defines key terms that surfaced in this exercise. these terms are a snapshot of our interpretation of the resources when it was undertaken in early 2018. when reviewing this list and the readers’ annotations, a formal rubric was not developed or required for participation. nor did members attempt to agree on definitions or classification. as each member brought their own professional background to bear on their assessment of the themes in each resource, we attempted to achieve a group understanding of a very broad issue with each participant contributing their own perspective. reporting these results is our desire to share our findings without any attempt to filter understanding of each participant. additionally, while most resources in this list acknowledged a reproducibility problem on some level, we did not differentiate between ‘pro-crisis’ or ‘no-crisis’ authors. group members read through each resource, identifying the primary problems or issues raised, as well as any solutions proposed. the final products the dcdg produced (as of october 2019) from this effort are publicly available at https://doi.org/10.18434/m32150 they include: ● reproducibility resources: an annotated list of resources that were evaluated, with annotations penned by dcdg members, with resource citations ● issues: a list of terms deemed as ‘issues’ or problems relating to reproducibility, with definitions https://doi.org/10.29173/iq967 https://tinyurl.com/dcdg-repro-2018 https://doi.org/10.18434/m32150 4/26 antognoli, erin; avila, regina; sears, jonathan; christiansen, leighton; tieman, jessica; hart, jacquelyn (2020) reproducibility literature analysis a federal information professional perspective, iassist quarterly 44(1-2), pp. 1-26. doi: https://doi.org/10.29173/iq967 ● solutions: a list of terms deemed as ‘solutions’ to the reproducibility issues, with definitions ● metrics: tallies of how often each term was selected as a primary theme of the resource, the resource dates, the scientific disciplines represented, and the types of resources reviewed resource metrics source information for this project consists of a variety of material types from numerous research disciplines. in total, we reviewed 122 resources relating to reproducibility in research data. to provide perspective on our source material, we include metrics for the resources selected for this project. the bulk of our research was derived from scholarly journal articles (59 percent), but we also included other sources such as books (4.1percent), presentations (0.8 percent), and websites (7.4 percent). figure 1: resource types reviewed by the group varied, but consisted overwhelmingly of journal articles. most of the resources, about 93 percent, discussed reproducibility and pointed to or offered solutions and best practices. a few, roughly 7 percent, pointed out reproducibility issues but did not speculate on causes and/or offered no solutions. https://doi.org/10.29173/iq967 5/26 antognoli, erin; avila, regina; sears, jonathan; christiansen, leighton; tieman, jessica; hart, jacquelyn (2020) reproducibility literature analysis a federal information professional perspective, iassist quarterly 44(1-2), pp. 1-26. doi: https://doi.org/10.29173/iq967 the compiled resources represent a broad sampling of items written for, or about, researchers within a variety of disciplines. the majority of the resources address science in general, but the total number reflects an impressive breadth of disciplines. they range from biomedicine to economics to psychology. in all, the materials target over 20 different disciplines, all discussing or analyzing the issue of reproducibility from their respective viewpoints. while numerous, this collection does not represent an exhaustive list of resources. therefore, this compilation of resources, while annotated, does not constitute a proper ‘literature review’ in the typical sense. there was no comprehensive selection of papers from any single discipline, nor were all scientific disciplines represented. this analysis did not intend to favor any one discipline over another, as we intended to gather more general information about reproducibility wherever the topic appeared in various resources. figure 2: disciplines represented in the reviewed resources varied greatly over 20 are represented in this exercise. the resources reviewed date from 2005 to the present. over 70 percent of these resources were published between 2014 to 2017. the lack of articles for 2018 occurred because the group compiled most of these resources in early 2018, at the start of the project. https://doi.org/10.29173/iq967 6/26 antognoli, erin; avila, regina; sears, jonathan; christiansen, leighton; tieman, jessica; hart, jacquelyn (2020) reproducibility literature analysis a federal information professional perspective, iassist quarterly 44(1-2), pp. 1-26. doi: https://doi.org/10.29173/iq967 figure 3. charts showing the number of resources published in years 2005-2018. the top chart shows the numbers from the dcdg resource list. the bottom charts publication counts from nexis.com and web of science that contained ‘reproducibility’ or ‘replicability’ in the title, headline, or lead paragraph. the overall increase in publications on the topic from 2014-2017 is similar both in the dcdg resources list, and publications indexed in the other two databases. while the resources reviewed comprise only a fraction of what was published during that time period, they do appear to be representative of that period. results from other databases reflect a similar https://doi.org/10.29173/iq967 7/26 antognoli, erin; avila, regina; sears, jonathan; christiansen, leighton; tieman, jessica; hart, jacquelyn (2020) reproducibility literature analysis a federal information professional perspective, iassist quarterly 44(1-2), pp. 1-26. doi: https://doi.org/10.29173/iq967 increase in publications on the topic of reproducibility during that same period. a search of nexis.com reveals the number of news items mentioning ‘reproducibility’ or ‘replicability’ in headlines or lead paragraphs follow the same peak, from 2014 to 2017. results from the science database web of science (https://apps.webofknowledge.com/) show a steady increase in the publications indexed from 20052018, with the highest being the latter years, 2017 and 2018. only one publication was added to our collection long after the others were gathered: the aforementioned report by the national academies of science, which provides recommendations to improve reproducibility and replicability in science. variety of issues and solutions many resources centered on the general problem of reproducibility and replicability, while others focused more on specific causes of the reproducibility crisis. of the 57 issues identified, the most common problems or issues discussed in our cross-section of resources include: replicability (in 23 resources); reproducibility (20); reproducibility crisis (20); bias (10); cherry picking (10); publish or perish (10); data sharing (8); data quality (7), and, researcher misconduct (6). the remaining 48 primary topic issues presented in five or fewer of the reviewed resources. figure 4: the nine (9) issues most frequently selected as a primary topic in the reviewed resources. solutions presented in the resources typically fell into three categories: best practices/standards; transparency/sharing; and, culture. solutions in the best practices category included general https://doi.org/10.29173/iq967 https://apps.webofknowledge.com/ 8/26 antognoli, erin; avila, regina; sears, jonathan; christiansen, leighton; tieman, jessica; hart, jacquelyn (2020) reproducibility literature analysis a federal information professional perspective, iassist quarterly 44(1-2), pp. 1-26. doi: https://doi.org/10.29173/iq967 recommendations on how to conduct data management at different stages of the research. solutions in the transparency category covered topics about sharing methods and data. the culture category comprised a broader look at research community behaviors and proposed methods for turning best practices into regular practice. many of the terms and definitions overlap, and many issues were also selected as solutions. the scientific research community is diverse and complex. ergo, the issues often overlap and require a comprehensive view when considering solutions. of the 55 solutions identified, the most common solutions discussed in our resource list were: experimental design (in 16 resources); transparency (16); code sharing (13); data sharing (13); replication studies (13); publication policy (12); standards (12); best practices (11); methods sharing (11); training (11); and, quality assurance (11). the remaining 44 solutions were selected as a primary topic in ten or fewer of the reviewed resources. figure 5: the eleven (11) most frequent solutions selected as a primary topic in the reviewed resources. primer of terminology and findings reproducibility issues or challenges this exercise revealed many concepts surrounding research reproducibility, though several ideas appeared much more frequently across the board. this section highlights some of the terms derived from the reviewed materials and most commonly identified as reproducibility issues or challenges. https://doi.org/10.29173/iq967 9/26 antognoli, erin; avila, regina; sears, jonathan; christiansen, leighton; tieman, jessica; hart, jacquelyn (2020) reproducibility literature analysis a federal information professional perspective, iassist quarterly 44(1-2), pp. 1-26. doi: https://doi.org/10.29173/iq967 many of the reported issues relate to or feed off one another. in the terms defined below we note other terms from the list that are related, where applicable. for a full list of the selected issues and their definitions, review the related data file at https://doi.org/10.18434/m32150. replicability replication is the attempt to recreate the conditions believed sufficient for obtaining a previously observed finding and is the means of establishing reproducibility of a finding with new data (open science collaboration, 2015). [related terms: reproducibility; repeatability] reproducibility reproducibility measures whether a study or experiment can be reproduced in its entirety. to achieve adequate reproducibility, studies implement measures to support verification of research, including, for example, sharing data and methods. no single factor or method alone achieves reproducibility in a study, and likewise, many factors can result in a study with poor reproducibility (munafò et al., 2017). [related terms: replicability; repeatability] reproducibility crisis the reproducibility crisis is defined as widespread failure to replicate the results of experiments and studies (weir, 2015). while many acknowledge the problem and see a need for research and experimental reform, many people debate the reproducibility problem as exaggerated (baker, 2016c). bias bias includes prejudice in favor of or against one thing, person, or group compared with another, usually in a way considered unfair. two common types of bias in research studies are confirmation bias and hindsight bias. confirmation bias promotes the tendency to focus on evidence that is in line with our expectations or favored explanation. hindsight bias is the tendency to see an event as having been predictable only after it has occurred (munafò et al., 2017). [related terms: cherry picking; replication studies (solution)] cherry picking cherry picking data includes suppressing evidence, or the fallacy of incomplete evidence by pointing to individual cases or data that seem to confirm a particular position or statistical significance, while ignoring a significant portion of related cases or data that may contradict that position (baker, 2016a). [related terms: bias; p-hacking] publish or perish ‘publish or perish’ is a phrase coined to describe the pressure in academia to rapidly and continuously publish academic work to sustain or further one’s career. frequent publication is one of the few methods at scholars’ disposal to demonstrate academic talent. the desire or need to publish work at a near-constant rate can lead to problems in reproducibility, and can lead to issues concerning selective reporting, also known as ‘cherry picking’ (baker, 2016a). [related terms: incentives; publication policy (solution); culture shift (solution)] data sharing definition is included alongside ‘code sharing’ in the solutions section below, as it appeared as both an issue and solution in this analysis. https://doi.org/10.29173/iq967 https://doi.org/10.18434/m32150 10/26 antognoli, erin; avila, regina; sears, jonathan; christiansen, leighton; tieman, jessica; hart, jacquelyn (2020) reproducibility literature analysis a federal information professional perspective, iassist quarterly 44(1-2), pp. 1-26. doi: https://doi.org/10.29173/iq967 data quality many definitions and factors determine data quality. however, data ‘fit for their intended uses in operations, decision making and planning’ are generally considered high quality. incorrect or incomplete data used to influence decision-making minimizes accuracy and strategic advantage (redman, 2008). bad or low-quality data can result from many avenues, including negative cultural influences such as ‘publish or perish’ as well as researcher misconduct, bias, p-hacking, or cherry picking data, among others. [related terms: validation; trust; quality assurance (solution)] researcher misconduct ‘the national science foundation (2001) defined scientific misconduct as fabrication, falsification, or plagiarism in proposing, performing, or reviewing research or in reporting research results. such misconduct is committed intentionally, knowingly, or in disregard of accepted practices. fabrication of data involves totally inventing a data set, while falsification refers to manipulation of equipment or changing data such that the research is not accurately represented in the research report’ (stroebe, postmes and spears, 2012). [related terms: trust; data quality] potential solutions to the reproducibility crisis while many potential solutions to reproducibility issues appeared throughout the reviewed resources, here we highlight some of the terms appearing the most frequently, or which we felt were important. as with the reproducibility issues discussed above, many of these reported solutions intersect. it is worthwhile to reiterate that many of the terms and ideas encountered during our research appeared as both issues and solutions—such as publication policy, incentives, training, data sharing, and data quality. for example, lack of ‘data sharing’ is cited as a hindrance to reproducibility. others cite ‘data sharing’ as a potential solution. again, where applicable, we append other terms to these definitions that are related to those we highlight here. for a full list of the selected solutions and their definitions, review the related data file available at https://doi.org/10.18434/m32150. experimental design experimental design aims to describe or explain the variation of information under conditions that are hypothesized to reflect the variation, and use this knowledge to collect more accurate data (acevesbueno et al., 2017). a framework for a systematic process to guide researchers and reviewers in assessing, documenting, and mitigating the sources of uncertainty in a study enhance comparability and reproducibility (plant et al., 2018). experimental design features should enhance, or facilitate inference about, the reproducibility and generalizability of the expected results (würbel, 2017). [related terms: pre-registration of results; case study] transparency simply put, transparency means ‘provable to the outside’ (bartling and fecher, 2015). transparency is the basis of open science, which refers to the process of making the content and process of producing evidence and claims clear and accessible to others. transparency is a scientific ideal, and adding ‘open’ should therefore be redundant (munafò et al., 2017). [related terms: open review; open science; data sharing; code sharing; methods sharing; trust (issue)] https://doi.org/10.29173/iq967 https://doi.org/10.18434/m32150 11/26 antognoli, erin; avila, regina; sears, jonathan; christiansen, leighton; tieman, jessica; hart, jacquelyn (2020) reproducibility literature analysis a federal information professional perspective, iassist quarterly 44(1-2), pp. 1-26. doi: https://doi.org/10.29173/iq967 data sharing, code sharing studies and research that implement data sharing support verification of research and conduct alternative analysis. sharing data in public repositories offers field-wide advantages in terms of accountability, data longevity, efficiency and quality (peng, dominici and zeger, 2006). likewise, code sharing provides discovery and access to the details of computational analysis including programming code and data (gezelter, 2015). these terms appear as both issues and solutions, since lack of sharing creates reproducibility issues. [related terms: data discovery index; methods sharing; open science; transparency] replication studies somewhat related to experimental design, replication is a term referring to the repetition of a research study, generally with different situations and different subjects, to determine if the basic findings of the original study can be applied to other participants and circumstances (gezelter, 2015). replication studies also help identify potential biases in the original study and serve as a basis for confirming or disconfirming prior findings (spector, johnson and young, 2014) (camerer et al., 2016). [related terms: collaborative replication; replication files; reproducible research standard (rrs); bias (issue)] publication policy journals have power to enforce transparency and reproducibility through their review and publication policies. this could help establish and enforce best practices. for instance, journals could require authors to register reports in advance so that the study protocol and analysis plan is locked in place before data collection even begins, and scientists should be encouraged to store methods, data, and code in repositories to help other groups reproduce experiments. one source suggested 5 to 10 percent of research funding should be spent on replication studies, and journals should devote more space to replication studies and null results (mcnutt, 2014) (begley and ioannidis, 2015). [related terms: data citation; funding agency requirements; incentives (issue); open review; public access; publish or perish (issue)] standards standards comprise the fundamental reference for a system of weights and measures, against which all other measuring devices are compared. standards contribute to improved research practices and promote positive change (capes-davis and neve, 2016). [related terms: best practices; metrological standards; reporting guidelines; reproducible research standard (rrs)] training open science, the movement to make scientific products and processes accessible to, and reusable by all, relies on culture and knowledge as much as it does on technologies and services. convincing researchers of the benefits of changing their practices, and equipping them with the skills and knowledge needed to do so can happen through training and education. a recommendation from citizen science states that “training, in particular, has been shown elsewhere to enhance accuracy and credibility (freitag, 2016, kosmala, 2016). [related terms: best practices; standards; trust (issue)] https://doi.org/10.29173/iq967 https://esajournals.onlinelibrary.wiley.com/doi/full/10.1002/bes2.1336#bes21336-bib-0011 https://esajournals.onlinelibrary.wiley.com/doi/full/10.1002/bes2.1336#bes21336-bib-0018 12/26 antognoli, erin; avila, regina; sears, jonathan; christiansen, leighton; tieman, jessica; hart, jacquelyn (2020) reproducibility literature analysis a federal information professional perspective, iassist quarterly 44(1-2), pp. 1-26. doi: https://doi.org/10.29173/iq967 quality assurance quality is an encompassing term comprising utility, objectivity, and integrity (national institute of standards and technology, 2009). whether in a laboratory setting or defining a quality system, quality assurance is the “system of activities whose purpose is to provide to the producer or user of a product or a service the assurance that it meets defined standards of quality with a stated level of confidence” (taylor, 1987). implementing quality assurance may involve a variety of checks for data completeness, validity, consistency, precision, and accuracy, among other aspects of the data (wiggins et al., 2011). research units often take an ad-hoc approach to methods and workflow, but standardizing operations and following certified protocols increases confidence in research results (baker, 2016b). [related terms: standards; best practices; data quality (issue); trust (issue); validation (issue)] incentives rewards for publishing, often tied to showcasing certain results, comprise the biggest challenge to widespread adoption of open data. conversely, well-conceived incentives may also provide solutions for increased reproducibility. if journals in particular regulate and highlight incentives for research practices promoting reproducibility, researchers will more widely adopt these positive practices (gezelter, 2015) (begley and ioannidis, 2015). [related terms: funding agency requirements; publication policy; publish or perish (issue)] culture shift an overarching theme with regard to reproducibility solutions boils down to a culture shift throughout the research endeavor and data gathering. culture shift encompasses changing beliefs, behaviors, and outcomes. industries must broadly address their practices at all stages of data collection, processing, publication, dissemination, and preservation to make reproducibility commonplace (baker, 2016b). opportunities for future study the resource list and primer presented here are an introduction to the landscape of the reproducibility crisis. the terms are defined broadly and remain at surface-level with regard to the topics described. while the final resource list contains annotations with key terms and definitions, more targeted research will uncover more nuances of specific problems and solutions. subject-specific analysis of reproducibility, as well as further and more honed examination of any of the issues and/or solutions may produce additional understanding. the work of this group is only the beginning for dcdg and other data and information managers. numerous opportunities for future research studies remain. such work could underpin the formation of other resources for those pursuing study of this topic, and for those who wish to improve reproducibility as a means to increase confidence in science within their own organizations. this collection of annotated resources spans the past fourteen years, with the bulk published in the last decade when digital methods have been the norm for scientific research. while digital data and modern computing and modeling practices certainly may cause their own unique reproducibility issues, an analysis of research practice in earlier literature may reveal more clarity into the scope and depth of these problems (bastian, 2016). as new research data insights, trends, and studies emerge, this resource list and primer should be updated to reflect the latest information that pertains to the reproducibility https://doi.org/10.29173/iq967 13/26 antognoli, erin; avila, regina; sears, jonathan; christiansen, leighton; tieman, jessica; hart, jacquelyn (2020) reproducibility literature analysis a federal information professional perspective, iassist quarterly 44(1-2), pp. 1-26. doi: https://doi.org/10.29173/iq967 crisis. in addition to adding more literature to the existing compilation of resources, performing more indepth exercises with these resources—such as textual analysis—may uncover new insights that can accurately inform training or education strategies to increase reproducibility. conclusion the issues surrounding reproducibility present great challenges. comprehending the numerous, nuanced causes within this exercise sometimes seemed insurmountable. while potentially overwhelming, our group made strides to understand the issues, root causes, and history of the reproducibility crisis in order to create a guide for ourselves. many organizations and individuals who recognize research reproducibility as an issue may not currently have the knowledge or resources to effect significant change. however, in summarizing a portion of the available literature on this topic, it is our hope that information professionals now have a better starting point to begin incremental change in promoting reproducible research. ultimately, we concluded that because the reproducibility crisis stemmed from such a wide variety of causes, stakeholders must take a multi-pronged approach to tackling the problem. a culture shift across all branches of research must occur to reverse the distrust this crisis has engendered. in moving forward, we must also realize our limitations. as data managers and librarians, many reproducibility issues stem from actions that occur prior to or after our typical involvement with the research. we recognize that we can facilitate reproducibility from within our own roles. given our positions as information professionals reflecting on the nature and scope of these reproducibility issues, our objective should involve education, training, and building awareness. for example, as data managers, we may not write data management plans, but we do share information and guidelines about how to write, curate, and archive them. we do not generate the data that results from research, but we can assist with organizing data, finding documentation standards for data, creating proper metadata, and assist in building and managing trusted repositories for the data. we do not format or publish the data, but we can share best practices for fair data, thereby contributing to research data that is findable, accessible, interoperable, and reusable (wilkinson, 2016). even from our set positions we can take concrete steps to aid and influence activities that move toward the goal of reproducibility. increasing education and awareness within our fields helps the cause. common past practice may have seen librarians contributing to the scientific endeavor in very limited ways, such as assisting with initial literature access and reviews, or as cataloging and preserving reported scientific results. however, the evolution of modern scientific research has, as discussed above, opened up a number of roles for library, information, and data professionals throughout the entire scientific research lifecycle. our participation can positively impact scientific reproducibility and replicability. first, we must understand the issues, and this paper is one contribution to develop that understanding. next we should apply that comprehension to aid our colleagues across research disciplines. https://doi.org/10.29173/iq967 14/26 antognoli, erin; avila, regina; sears, jonathan; christiansen, leighton; tieman, jessica; hart, jacquelyn (2020) reproducibility literature analysis a federal information professional perspective, iassist quarterly 44(1-2), pp. 1-26. doi: https://doi.org/10.29173/iq967 references aceves-bueno, e., adeleye, a., feraud, m., huang, y., tao, m., yang, y. and anderson, s. (2017). the accuracy of citizen science data: a quantitative review. the bulletin of the ecological society of america, 98(4), pp.278-290. https://doi.org/10.1002/bes2.1336 baker, m. (2016a). 1,500 scientists lift the lid on reproducibility. nature, 533(7604), pp.452-454. https://doi.org/10.1038/533452a baker, m. (2016b). how quality control could save your science. nature, 529(7587), pp.456-458. https://doi.org/10.1038/529456a baker, m. (2016c). psychology’s reproducibility problem is exaggerated – say psychologists. nature. https://doi.org/10.1038/nature.2016.19498 bartling, s. and fecher, b. (2015). could blockchain provide the technical fix to solve science’s reproducibility crisis?. lse impact of social sciences. [online] available at: http://eprints.lse.ac.uk/67354/1/could_blockchain_provide_technical_fix.pdf [accessed 10 sep. 2019]. bastian, h. (2016). reproducibility crisis timeline—milestones in tackling research reliability. [online] phys.org. available at: https://phys.org/news/2016-12-crisis-timelinemilestones-tackling-reliability.html [accessed 10 sep. 2019]. begley, c. and ioannidis, j. (2015). reproducibility in science. circulation research, 116(1), pp.116-126. https://doi.org/10.1161/circresaha.114.303819 camerer, c., dreber, a., forsell, e., ho, t., huber, j., johannesson, m., kirchler, m., almenberg, j., altmejd, a., chan, t., heikensten, e., holzmeister, f., imai, t., isaksson, s., nave, g., pfeiffer, t., razen, m. and wu, h. (2016). evaluating replicability of laboratory experiments in economics. science, 351(6280), pp.1433-1436. https://doi.org/10.1126/science.aaf0918 capes-davis, a. and neve, r. (2016). authentication: a standard problem or a problem of standards? plos biology, 14(6), p.e1002477. https://doi.org/10.1371/journal.pbio.1002477 cendi.gov. (2019). cendi. [online] available at: https://cendi.gov/ [accessed 10 sep. 2019]. christiansen, l. (2017). expanding the u.s. federal data curation community: year 01 at the national transportation library. research data alliance tenth plenary. montreal, quebec, canada. https://doi.org/10.21949/1504432 crane, h. (2017). why “redefining statistical significance” will not improve reproducibility and could make the replication crisis worse. [online] arxiv.org. available at: http://arxiv.org/abs/1711.07801 [accessed 10 sep. 2019]. freitag, a., meyer, r. and whiteman, l. (2016). strategies employed by citizen science programs to increase the credibility of their data. citizen science: theory and practice., 1(1), p.2. https://doi.org/10.5334/cstp.6 https://doi.org/10.29173/iq967 https://doi.org/10.1002/bes2.1336 https://doi.org/10.1038/533452a https://doi.org/10.1038/529456a https://doi.org/10.1038/nature.2016.19498 http://eprints.lse.ac.uk/67354/1/could_blockchain_provide_technical_fix.pdf https://phys.org/news/2016-12-crisis-timelinemilestones-tackling-reliability.html https://doi.org/10.1161/circresaha.114.303819 https://doi.org/10.1126/science.aaf0918 https://doi.org/10.1371/journal.pbio.1002477 https://cendi.gov/ https://doi.org/10.21949/1504432 http://arxiv.org/abs/1711.07801 http://dx.doi.org/10.5334/cstp.6 15/26 antognoli, erin; avila, regina; sears, jonathan; christiansen, leighton; tieman, jessica; hart, jacquelyn (2020) reproducibility literature analysis a federal information professional perspective, iassist quarterly 44(1-2), pp. 1-26. doi: https://doi.org/10.29173/iq967 gezelter, j. (2015). open source and open data should be standard practices. the journal of physical chemistry letters, 6(7), pp.1168-1169. https://doi.org/10.1021/acs.jpclett.5b00285 kosmala, m., wiggins, a., swanson, a. and simmons, b. (2016). assessing data quality in citizen science. frontiers in ecology and the environment, 14(10), pp.551-560. https://doi.org/10.1002/fee.1436 mcnutt, m. (2014). journals unite for reproducibility. science, 346(6210), pp.679-679. https://doi.org/10.1126/science.aaa1724 munafò, m., nosek, b., bishop, d., button, k., chambers, c., percie du sert, n., simonsohn, u., wagenmakers, e., ware, j. and ioannidis, j. (2017). a manifesto for reproducible science. nature human behaviour, 1(1). https://doi.org/10.1038/s41562-016-0021 national academies of sciences, engineering, and medicine (2019). reproducibility and replicability in science. washington, dc: the national academies press. https://doi.org/10.17226/25303 national institute of standards and technology (2009). nist information quality standards. [online] available at: https://www.nist.gov/nist-information-quality-standards [accessed 18 oct. 2019] open science collaboration (2015). estimating the reproducibility of psychological science. science, 349(6251), aac4716. https://doi.org/10.1126/science.aac4716 peng, r., dominici, f. and zeger, s. (2006). reproducible epidemiologic research. american journal of epidemiology, 163(9), pp.783-789. https://doi.org/10.1093/aje/kwj093 plant, a., becker, c., hanisch, r., boisvert, r., possolo, a. and elliott, j. (2018). how measurement science can improve confidence in research results. plos biology, 16(4), p.e2004299. https://doi.org/10.1371/journal.pbio.2004299 redman, t. (2008). data driven: profiting from your most important business asset. boston: harvard business review press. spector, j., johnson, t. and young, p. (2014). an editorial on replication studies and scaling up efforts. educational technology research and development, 63(1), pp.1-4. https://doi.org/10.1007/s11423-0149364-3 stroebe, w., postmes, t. and spears, r. (2012). scientific misconduct and the myth of self-correction in science. perspectives on psychological science, 7(6), pp.670-688. https://doi.org/10.1177/1745691612460687 taylor, j. (1987). quality assurance of chemical measurements. chelsea: lewis publ. weir, k. (2015). a reproducibility crisis?. monitor on psychology, [online] 46(9), p.39. available at: https://www.apa.org/monitor/2015/10/share-reproducibility [accessed 10 sep. 2019]. wiggins, a., newman, g., stevenson, r. and crowston, k. (2011). mechanisms for data quality and validation in citizen science. 2011 ieee seventh international conference on e-science workshops. https://doi.org/10.1109/esciencew.2011.27 https://doi.org/10.29173/iq967 https://doi.org/10.1021/acs.jpclett.5b00285 https://doi.org/10.1002/fee.1436 https://doi.org/10.1126/science.aaa1724 https://doi.org/10.1038/s41562-016-0021 https://doi.org/10.17226/25303 https://www.nist.gov/nist-information-quality-standards https://doi.org/10.1126/science.aac4716 https://doi.org/10.1093/aje/kwj093 https://doi.org/10.1371/journal.pbio.2004299 https://doi.org/10.1371/journal.pbio.2004299 https://doi.org/10.1007/s11423-014-9364-3 https://doi.org/10.1007/s11423-014-9364-3 https://doi.org/10.1177/1745691612460687 https://doi.org/10.1177/1745691612460687 https://www.apa.org/monitor/2015/10/share-reproducibility https://doi.org/10.1109/esciencew.2011.27 16/26 antognoli, erin; avila, regina; sears, jonathan; christiansen, leighton; tieman, jessica; hart, jacquelyn (2020) reproducibility literature analysis a federal information professional perspective, iassist quarterly 44(1-2), pp. 1-26. doi: https://doi.org/10.29173/iq967 wilkinson, m.d., dumontier, m., aalbersberg, i.j., appleton, g., axton, m., baak, a., blomberg, n., boiten, j.w., da silva santos, l.b., bourne, p.e. and bouwman, j., 2016. the fair guiding principles for scientific data management and stewardship. scientific data, 3. https://doi.org/10.1038/sdata.2016.18 würbel, h. (2017). more than 3rs: the importance of scientific validity for harm-benefit analysis of animal research. lab animal, 46(4), pp.164-166. https://doi.org/10.1038/laban.1220 https://doi.org/10.29173/iq967 https://doi.org/10.1038/sdata.2016.18 https://doi.org/10.1038/laban.1220 17/26 antognoli, erin; avila, regina; sears, jonathan; christiansen, leighton; tieman, jessica; hart, jacquelyn (2020) reproducibility literature analysis a federal information professional perspective, iassist quarterly 44(1-2), pp. 1-26. doi: https://doi.org/10.29173/iq967 appendix a list of resources evaluated aceves-bueno, e., adeleye, a., feraud, m., huang, y., tao, m., yang, y. and anderson, s. (2017). the accuracy of citizen science data: a quantitative review. the bulletin of the ecological society of america, 98(4), pp.278-290. https://doi.org/10.1002/bes2.1336 aguinis, h., cascio, w. and ramani, r. (2017). science’s reproducibility and replicability crisis: international business is not immune. journal of international business studies, 48(6), pp.653-663. https://doi.org/10.1057/s41267-017-0081-0 akpan, n. (2017). why bad science is plaguing health research—and how to fix it. [online] pbs newshour. available at: https://www.pbs.org/newshour/science/why-bad-science-is-plaguing-healthresearch-rigor-mortis-richard-harris [accessed 10 sep. 2019]. anderson, c., anderson, j., assen, m., attridge, p., attwood, a., axt, j., babel, m., bahník, š., baranski, e. and barnett-cowan, m. (2019). reproducibility project: psychology. [online] osf. available at: https://osf.io/ezcuj/ [accessed 10 sep. 2019]. https://doi.org/10.17605/osf.io/ezcuj aschwanden, c. (2015). science isn’t broken. (online). fivethirtyeight. available at: https://fivethirtyeight.com/features/science-isnt-broken/ [accessed 10 september 2019] asendorpf, j., conner, m., de fruyt, f., de houwer, j., denissen, j., fiedler, k., fiedler, s., funder, d., kliegl, r., nosek, b., perugini, m., roberts, b., schmitt, m., van aken, m., weber, h. and wicherts, j. (2013). recommendations for increasing replicability in psychology. european journal of personality, 27(2), pp.108-119. https://doi.org/10.1002/per.1919 baker, m. (2015). over half of psychology studies fail reproducibility test. nature. https://doi.org/10.1038/nature.2015.18248 baker, m. (2016). 1,500 scientists lift the lid on reproducibility. nature, 533(7604), pp.452-454. https://doi.org/10.1038/533452a baker, m. (2016). how quality control could save your science. nature, 529(7587), pp.456-458. https://doi.org/10.1038/529456a baker, m. (2016). muddled meanings hamper efforts to fix reproducibility crisis. nature. https://doi.org/10.1038/nature.2016.20076 baker, m. (2016). psychology’s reproducibility problem is exaggerated – say psychologists. nature. https://doi.org/10.1038/nature.2016.19498 bal, l. (2015). is science broken? the reproducibility crisis. [blog] on biology. available at: https://blogs.biomedcentral.com/on-biology/2015/03/20/is-science-broken-a-reproducibility-crisis/ [accessed 10 september 2019]. https://doi.org/10.29173/iq967 https://doi.org/10.1002/bes2.1336 https://doi.org/10.1057/s41267-017-0081-0 https://www.pbs.org/newshour/science/why-bad-science-is-plaguing-health-research-rigor-mortis-richard-harris https://www.pbs.org/newshour/science/why-bad-science-is-plaguing-health-research-rigor-mortis-richard-harris https://osf.io/ezcuj/ https://doi.org/10.17605/osf.io/ezcuj https://fivethirtyeight.com/features/science-isnt-broken/ https://doi.org/10.1002/per.1919 https://doi.org/10.1038/nature.2015.18248 https://doi.org/10.1038/533452a https://doi.org/10.1038/529456a https://doi.org/10.1038/nature.2016.20076 https://doi.org/10.1038/nature.2016.19498 https://blogs.biomedcentral.com/on-biology/2015/03/20/is-science-broken-a-reproducibility-crisis/ 18/26 antognoli, erin; avila, regina; sears, jonathan; christiansen, leighton; tieman, jessica; hart, jacquelyn (2020) reproducibility literature analysis a federal information professional perspective, iassist quarterly 44(1-2), pp. 1-26. doi: https://doi.org/10.29173/iq967 bartling, s. and fecher, b. (2015). could blockchain provide the technical fix to solve science’s reproducibility crisis?. lse impact of social sciences. [online] available at: http://eprints.lse.ac.uk/67354/1/could_blockchain_provide_technical_fix.pdf [accessed 13 aug. 2019]. bastian, h. (2016). reproducibility crisis timeline—milestones in tackling research reliability. [online] phys.org. available at: https://phys.org/news/2016-12-crisis-timelinemilestones-tackling-reliability.html [accessed 10 sep. 2019]. bauer, h. (2015). how medical practice has gone wrong: causes of the lack-of-reproducibility crisis in medical research. journal of controversies in biomedical research, 1(1). https://doi.org/10.15586/jcbmr.2015.8 begley, c. (2013). six red flags for suspect work. nature, 497(7450), pp.433-434. https://doi.org/10.1038/497433a begley, c. and ioannidis, j. (2015). reproducibility in science. circulation research, 116(1), pp.116-126. https://doi.org/10.1161/circresaha.114.303819 bergh, d., sharp, b., aguinis, h. and li, m. (2017). is there a credibility crisis in strategic management research? evidence on the reproducibility of study findings. strategic organization, 15(3), pp.423-436. https://doi.org/10.1177/1476127017701076 bissell, m. (2013). reproducibility: the risks of the replication drive. nature, 503(7476), pp.333-334. https://doi.org/10.1038/503333a bollen, k., cacioppo, j., kaplan, r., krosnick, j., olds, j. and dean, h. (2015). social, behavioral, and economic sciences perspectives on robust and reliable science. [online] nsf.gov. available at: https://www.nsf.gov/sbe/ac_materials/sbe_robust_and_reliable_research_report.pdf [accessed 10 sep. 2019]. boos, d. and stefanski, l. (2011). p-value precision and reproducibility. the american statistician, 65(4), pp.213-221. https://doi.org/10.1198/tas.2011.10129 borgman, c. (2019). research data, reproducibility, and curation. digital social research: a forum for policy and practice — oxford internet institute. [online] oii.ox.ac.uk. available at: http://www.oii.ox.ac.uk/events/?id=487 [accessed 10 sep. 2019]. camerer, c., dreber, a., forsell, e., ho, t., huber, j., johannesson, m., kirchler, m., almenberg, j., altmejd, a., chan, t., heikensten, e., holzmeister, f., imai, t., isaksson, s., nave, g., pfeiffer, t., razen, m. and wu, h. (2016). evaluating replicability of laboratory experiments in economics. science, 351(6280), pp.1433-1436. https://doi.org/10.1126/science.aaf0918 camerer, c., dreber, a., holzmeister, f., ho, t., huber, j., johannesson, m., kirchler, m., nave, g., nosek, b., pfeiffer, t., altmejd, a., buttrick, n., chan, t., chen, y., forsell, e., gampa, a., heikensten, e., hummer, l., imai, t., isaksson, s., manfredi, d., rose, j., wagenmakers, e. and wu, h. (2018). evaluating the replicability of social science experiments in nature and science between 2010 and 2015. nature human behaviour, 2(9), pp.637-644. https://doi.org/10.1038/s41562-018-0399-z https://doi.org/10.29173/iq967 http://eprints.lse.ac.uk/67354/1/could_blockchain_provide_technical_fix.pdf https://phys.org/news/2016-12-crisis-timelinemilestones-tackling-reliability.html https://doi.org/10.15586/jcbmr.2015.8 https://doi.org/10.1038/497433a https://doi.org/10.1161/circresaha.114.303819 https://doi.org/10.1177/1476127017701076 https://doi.org/10.1038/503333a https://www.nsf.gov/sbe/ac_materials/sbe_robust_and_reliable_research_report.pdf https://doi.org/10.1198/tas.2011.10129 http://www.oii.ox.ac.uk/events/?id=487 https://doi.org/10.1126/science.aaf0918 https://doi.org/10.1038/s41562-018-0399-z 19/26 antognoli, erin; avila, regina; sears, jonathan; christiansen, leighton; tieman, jessica; hart, jacquelyn (2020) reproducibility literature analysis a federal information professional perspective, iassist quarterly 44(1-2), pp. 1-26. doi: https://doi.org/10.29173/iq967 capes-davis, a. and neve, r. (2016). authentication: a standard problem or a problem of standards?. plos biology, 14(6), p.e1002477. https://doi.org/10.1371/journal.pbio.1002477 chang, a. c., and li, p. (2015). is economics research replicable? sixty published papers from thirteen journals say ”usually not”. finance and economics discussion series 2015-083. washington: board of governors of the federal reserve system. https://doi.org/10.17016/feds.2015.083 collins, f. and tabak, l. (2014). policy: nih plans to enhance reproducibility. nature, 505(7485), pp.612613. https://doi.org/10.1038/505612a cook, b. (2014). a call for examining replication and bias in special education research. remedial and special education, 35(4), pp.233-246. https://doi.org/10.1177/0741932514528995 couchman, j. (2013). peer review and reproducibility. crisis or time for course correction?. journal of histochemistry & cytochemistry, 62(1), pp.9-10. https://doi.org/10.1369/0022155413513462 crane, h. (2017). why “redefining statistical significance” will not improve reproducibility and could make the replication crisis worse. [online] arxiv.org. available at: https://arxiv.org/abs/1711.07801 [accessed 10 sep. 2019]. dafoe, a. (2013). science deserves better: the imperative to share complete replication files. ps: political science & politics, 47(01), pp.60-66. https://doi.org/10.1017/s104909651300173x donoho, d. (2010). an invitation to reproducible computational research. biostatistics, 11(3), pp.385388. https://doi.org/10.1093/biostatistics/kxq028 donoho, d., maleki, a., rahman, i., shahram, m. and stodden, v. (2009). reproducible research in computational harmonic analysis. computing in science & engineering, 11(1), pp.8-18. https://doi.org/10.1109/mcse.2009.15 elman, c. and kapiszewski, d. (2013). data access and research transparency in the qualitative tradition. ps: political science & politics, 47(01), pp.43-47. https://doi.org/10.1017/s1049096513001777 epa press office (2018). epa administrator pruitt proposes rule to strengthen science used in epa regulations | us epa. [online] us epa. available at: https://www.epa.gov/newsreleases/epaadministrator-pruitt-proposes-rule-strengthen-science-used-epa-regulations [accessed 10 sep. 2019]. estimating the reproducibility of psychological science. (2015). science, 349(6251), pp.aac4716-aac4716. https://doi.org/10.1126/science.aac4716 evanschitzky, h. and armstrong, j. (2013). research with in-built replications: comment and further suggestions for replication research. journal of business research, 66(9), pp.1406-1408. https://doi.org/10.1016/j.jbusres.2012.05.006 fidler, f. & gordon, a. (2013). science is in a reproducibility crisis: how do we resolve it? [online] phys.org. available at: https://phys.org/news/2013-09-science-crisis.html [accessed 10 sep. 2019]. https://doi.org/10.29173/iq967 https://doi.org/10.1371/journal.pbio.1002477 https://doi.org/10.17016/feds.2015.083 https://doi.org/10.1038/505612a https://doi.org/10.1177/0741932514528995 https://doi.org/10.1369/0022155413513462 https://arxiv.org/abs/1711.07801 https://doi.org/10.1017/s104909651300173x https://doi.org/10.1093/biostatistics/kxq028 https://doi.org/10.1109/mcse.2009.15 https://doi.org/10.1017/s1049096513001777 https://www.epa.gov/newsreleases/epa-administrator-pruitt-proposes-rule-strengthen-science-used-epa-regulations https://www.epa.gov/newsreleases/epa-administrator-pruitt-proposes-rule-strengthen-science-used-epa-regulations https://doi.org/10.1126/science.aac4716 https://doi.org/10.1016/j.jbusres.2012.05.006 https://phys.org/news/2013-09-science-crisis.html 20/26 antognoli, erin; avila, regina; sears, jonathan; christiansen, leighton; tieman, jessica; hart, jacquelyn (2020) reproducibility literature analysis a federal information professional perspective, iassist quarterly 44(1-2), pp. 1-26. doi: https://doi.org/10.29173/iq967 firestein, s. (2016). op-ed: why failure to replicate findings can actually be good for science. [online] los angeles times. available at: http://www.latimes.com/opinion/op-ed/la-oe-0214-firestein-sciencereplication-failure-20160214-story.html [accessed 10 sep. 2019]. fitzjohn, r., pennell, m., amy zanne, a., & cornwell, w. (2014). reproducible research is still a challenge. [blog] ropensci. available at: https://ropensci.org/blog/2014/06/09/reproducibility/ fomel, s. and claerbout, j. (2009). guest editors’ introduction: reproducible research. computing in science & engineering, 11(1), pp.5-7. https://doi.org/10.1109/mcse.2009.14 gezelter, j. (2015). open source and open data should be standard practices. the journal of physical chemistry letters, 6(7), pp.1168-1169. https://doi.org/10.1021/acs.jpclett.5b00285 goecks, j., nekrutenko, a., taylor, j. and galaxy team, t. (2010). galaxy: a comprehensive approach for supporting accessible, reproducible, and transparent computational research in the life sciences. genome biology, 11(8), p.r86. https://doi.org/10.1186/gb-2010-11-8-r86 goodman, s., fanelli, d. and ioannidis, j. (2016). what does research reproducibility mean?. science translational medicine, 8(341), pp.341ps12-341ps12. https://doi.org/10.1126/scitranslmed.aaf5027 greenland, s., senn, s., rothman, k., carlin, j., poole, c., goodman, s. and altman, d. (2016). statistical tests, p values, confidence intervals, and power: a guide to misinterpretations. european journal of epidemiology, 31(4), pp.337-350. https://doi.org/10.1007/s10654-016-0149-3 harris, r. (2017). rigor mortis. 1st ed. basic books, pp.1-279. hoffman, j. (2016). archive computer code with raw data. nature, 534(7607), pp.326-326. https://doi.org/10.1038/534326d höller, y., uhl, a., bathke, a., thomschewski, a., butz, k., nardone, r., fell, j. and trinka, e. (2017). reliability of eeg measures of interaction: a paradigm shift is needed to fight the reproducibility crisis. frontiers in human neuroscience, 11. https://doi.org/10.3389/fnhum.2017.00441 hossenfelder, s. (2017). science needs reason to be trusted. nature physics, 13(4), pp.316-317. https://doi.org/10.1038/nphys4079 ioannidis, j. (2005). why most published research findings are false. plos medicine, 2(8), p.e124. https://doi.org/10.1371/journal.pmed.0020124 ioannidis, j. (2008). why most discovered true associations are inflated. epidemiology, 19(5), pp.640648. https://doi.org/10.1097/ede.0b013e31818131e7 ioannidis, j. (2012). why science is not necessarily self-correcting. perspectives on psychological science, 7(6), pp.645-654. https://doi.org/10.1177/1745691612464056 ioannidis, j. (2014). how to make more published research true. plos medicine, 11(10), p.e1001747. https://doi.org/10.1371/journal.pmed.1001747 https://doi.org/10.29173/iq967 http://www.latimes.com/opinion/op-ed/la-oe-0214-firestein-science-replication-failure-20160214-story.html http://www.latimes.com/opinion/op-ed/la-oe-0214-firestein-science-replication-failure-20160214-story.html https://ropensci.org/blog/2014/06/09/reproducibility/ https://doi.org/10.1109/mcse.2009.14 https://doi.org/10.1021/acs.jpclett.5b00285 https://doi.org/10.1186/gb-2010-11-8-r86 https://doi.org/10.1126/scitranslmed.aaf5027 https://doi.org/10.1007/s10654-016-0149-3 https://doi.org/10.1038/534326d https://doi.org/10.3389/fnhum.2017.00441 https://doi.org/10.1038/nphys4079 https://doi.org/10.1371/journal.pmed.0020124 https://doi.org/10.1097/ede.0b013e31818131e7 https://doi.org/10.1177/1745691612464056 https://doi.org/10.1371/journal.pmed.1001747 21/26 antognoli, erin; avila, regina; sears, jonathan; christiansen, leighton; tieman, jessica; hart, jacquelyn (2020) reproducibility literature analysis a federal information professional perspective, iassist quarterly 44(1-2), pp. 1-26. doi: https://doi.org/10.29173/iq967 iorns, e. and chong, c. (2014). new forms of checks and balances are needed to improve research integrity. f1000research, 3, p.119. https://doi.org/10.12688/f1000research.3714.1 ishiyama, j. (2013). replication, research transparency, and journal publications: individualism, community models, and the future of replication studies. ps: political science & politics, 47(01), pp.7883. https://doi.org/10.1017/s1049096513001765 jarvis, m. and williams, m. (2016). irreproducibility in preclinical biomedical research: perceptions, uncertainties, and knowledge gaps. trends in pharmacological sciences, 37(4), pp.290-302. https://doi.org/10.1016/j.tips.2015.12.001 jasny, b., wigginton, n., mcnutt, m., bubela, t., buck, s., cook-deegan, r., gardner, t., hanson, b., hustad, c., kiermer, v., lazer, d., lupia, a., manrai, a., mcconnell, l., noonan, k., phimister, e., simon, b., strandburg, k., summers, z. and watts, d. (2017). fostering reproducibility in industry-academia research. science, 357(6353), pp.759-761. https://doi.org/10.1126/science.aan4906 jones, a. and kemp, a. (2016). why is so much research dodgy? blame the research excellence framework. [online] the guardian. available at: http://www.theguardian.com/higher-educationnetwork/2016/oct/17/why-is-so-much-research-dodgy-blame-the-research-excellence-framework [accessed 10 sep. 2019]. jordan, p. (2014). give young scientists a level playing field. science, 346(6210), pp.711-711. https://doi.org/10.1126/science.346.6210.711-a kappenman, e. and keil, a. (2016). introduction to the special issue on recentering science: replication, robustness, and reproducibility in psychophysiology. psychophysiology, 54(1), pp.3-5. https://doi.org/10.1111/psyp.12787 lawton, j. (2016). reproducibility and replicability of science and thoracic surgery. the journal of thoracic and cardiovascular surgery, 152(6), pp.1489-1491. https://doi.org/10.1016/j.jtcvs.2016.08.044 leek, j. and peng, r. (2015). opinion: reproducible research can still be wrong: adopting a prevention approach: fig. 1. proceedings of the national academy of sciences, 112(6), pp.1645-1646. https://doi.org/10.1073/pnas.1421412111 leveque, r. (2009). python tools for reproducible research on hyperbolic problems. computing in science & engineering, 11(1), pp.19-27. https://doi.org/10.1109/mcse.2009.13 lupia, a. and alter, g. (2013). data access and research transparency in the quantitative tradition. ps: political science & politics, 47(01), pp.54-59. https://doi.org/10.1017/s1049096513001728 makel, m. and plucker, j. (2014). facts are more important than novelty. educational researcher, 43(6), pp.304-316. https://doi.org/10.3102/0013189x14545513 maniadis, z., tufano, f. and list, j. (2017). to replicate or not to replicate? exploring reproducibility in economics through the lens of a model and a pilot study. the economic journal, 127(605), pp.f209f235. https://doi.org/10.1111/ecoj.12527 https://doi.org/10.29173/iq967 https://doi.org/10.12688/f1000research.3714.1 https://doi.org/10.1017/s1049096513001765 https://doi.org/10.1016/j.tips.2015.12.001 https://doi.org/10.1126/science.aan4906 http://www.theguardian.com/higher-education-network/2016/oct/17/why-is-so-much-research-dodgy-blame-the-research-excellence-framework http://www.theguardian.com/higher-education-network/2016/oct/17/why-is-so-much-research-dodgy-blame-the-research-excellence-framework https://doi.org/10.1126/science.346.6210.711-a https://doi.org/10.1111/psyp.12787 https://doi.org/10.1016/j.jtcvs.2016.08.044 https://doi.org/10.1073/pnas.1421412111 https://doi.org/10.1109/mcse.2009.13 https://doi.org/10.1017/s1049096513001728 https://doi.org/10.3102/0013189x14545513 https://doi.org/10.1111/ecoj.12527 22/26 antognoli, erin; avila, regina; sears, jonathan; christiansen, leighton; tieman, jessica; hart, jacquelyn (2020) reproducibility literature analysis a federal information professional perspective, iassist quarterly 44(1-2), pp. 1-26. doi: https://doi.org/10.29173/iq967 mcdermott, r. (2013). research transparency and data archiving for experiments. ps: political science & politics, 47(01), pp.67-71. https://doi.org/10.1017/s1049096513001741 mclellan, m., brannon, p., campa, a., daley-laursen, s., kannan, g., olsen, n., taylor, r. and thilmany, d. (2016). reproducibility and rigor in ree’s portfolio of research. technical report. 9pp. https://doi.org/10.13140/rg.2.2.24369.17761 mcnutt, m. (2014). journals unite for reproducibility. science, 346(6210), pp.679-679. https://doi.org/10.1126/science.aaa1724 mcnutt, m. (2014). reproducibility. science, 343(6168), pp.229-229. https://doi.org/10.1126/science.1250475 monroe, d. (2015). when data is not enough. communications of the acm, 58(12), pp.12-14. https://doi.org/10.1145/2833138 morrell, k. and lucas, j. (2012). the replication problem and its implications for policy studies. critical policy studies, 6(2), pp.182-200. https://doi.org/10.1080/19460171.2012.689738 mullane, k. and williams, m. (2017). enhancing reproducibility: failures from reproducibility initiatives underline core challenges. biochemical pharmacology, 138, pp.7-18. https://doi.org/10.1016/j.bcp.2017.04.008 mullane, k., curtis, m. and williams, m. (2018). reproducibility in biomedical research. research in the biomedical sciences, pp.1-66. https://doi.org/10.1016/b978-0-12-804725-5.00001-x munafò, m. (2017). metascience: reproducibility blues. nature, 543(7647), pp.619-620. https://doi.org/10.1038/543619a munafò, m., nosek, b., bishop, d., button, k., chambers, c., percie du sert, n., simonsohn, u., wagenmakers, e., ware, j. and ioannidis, j. (2017). a manifesto for reproducible science. nature human behaviour, 1(1). https://doi.org/10.1038/s41562-016-0021 national academies of sciences, engineering, and medicine (2016). statistical challenges in assessing and fostering the reproducibility of scientific results: summary of a workshop. washington, d.c.: national academies press. https://doi.org/10.17226/21915 national academies of sciences, engineering, and medicine (2017). fostering integrity in research. washington, dc: the national academies press. 326pp. isbn 978-0-309-39125-2. https://doi.org/10.17226/21896 national academies of sciences, engineering, and medicine (2019). reproducibility and replicability in science. washington, dc: the national academies press. https://doi.org/10.17226/25303 national institutes of health (nih) (2019). rigor and reproducibility. [online] national institutes of health (nih). available at: https://www.nih.gov/research-training/rigor-reproducibility [accessed 10 sep. 2019]. https://doi.org/10.29173/iq967 https://doi.org/10.1017/s1049096513001741 https://doi.org/10.13140/rg.2.2.24369.17761 https://doi.org/10.1126/science.aaa1724 https://doi.org/10.1126/science.1250475 https://doi.org/10.1145/2833138 https://doi.org/10.1080/19460171.2012.689738 https://doi.org/10.1016/j.bcp.2017.04.008 https://doi.org/10.1016/b978-0-12-804725-5.00001-x https://doi.org/10.1038/543619a https://doi.org/10.1038/s41562-016-0021 https://doi.org/10.17226/21915 https://doi.org/10.17226/21896 https://doi.org/10.17226/25303 https://www.nih.gov/research-training/rigor-reproducibility 23/26 antognoli, erin; avila, regina; sears, jonathan; christiansen, leighton; tieman, jessica; hart, jacquelyn (2020) reproducibility literature analysis a federal information professional perspective, iassist quarterly 44(1-2), pp. 1-26. doi: https://doi.org/10.29173/iq967 neuroskeptic (2015). psychology should aim for 100% reproducibility [blog] neuroskeptic available at: http://blogs.discovermagazine.com/neuroskeptic/2015/09/07/100-percentreproducibility/#.wqbyo2rwzaq. [accessed 10 september 2019] pellizzari, e., lohr, k., blatecky, a. and creel, d. (2017). reproducibility. research triangle park, nc: rti press. https://doi.org/10.3768/rtipress.2017.bk.0020.1708 peng, r. (2015). the reproducibility crisis in science: a statistical counterattack. significance, 12(3), pp.30-32. https://doi.org/10.1111/j.1740-9713.2015.00827.x peng, r. and eckel, s. (2009). distributed reproducible research using cached computations. computing in science & engineering, 11(1), pp.28-34. https://doi.org/10.1109/mcse.2009.6 peng, r., dominici, f. and zeger, s. (2006). reproducible epidemiologic research. american journal of epidemiology, 163(9), pp.783-789. https://doi.org/10.1093/aje/kwj093 plant, a., becker, c., hanisch, r., boisvert, r., possolo, a. and elliott, j. (2018). how measurement science can improve confidence in research results. plos biology, 16(4), p.e2004299. https://doi.org/10.1371/journal.pbio.2004299 plant, a., locascio, l., may, w. and gallagher, p. (2014). improved reproducibility by assuring confidence in measurements in biomedical research. nature methods, 11(9), pp.895-898. https://doi.org/10.1038/nmeth.3076 pröll, s. and rauber, a. (2014). a scalable framework for dynamic data citation of arbitrary structured data. proceedings of 3rd international conference on data management technologies and applications. https://doi.org/10.5220/0004991802230230 rauber, a. and pröll, s. (2015). scalable dynamic data citation rda-wg-dc position paper. [online] research data alliance. available at: https://www.rd-alliance.org/groups/data-citationwg/wiki/scalable-dynamic-data-citation-rda-wg-dc-position-paper.html [accessed 10 sep. 2019]. resnik, d. and shamoo, a. (2016). reproducibility and research integrity. accountability in research, 24(2), pp.116-123. https://doi.org/10.1080/08989621.2016.1257387 richardson, d., kwan, m., alter, g. and mckendry, j. (2015). replication of scientific research: addressing geoprivacy, confidentiality, and data sharing challenges in geospatial research. annals of gis, 21(2), pp.101-110. https://doi.org/10.1080/19475683.2015.1027792 sarewitz, d. (2015). reproducibility will not cure what ails science. nature, 525(7568), pp.159-159. https://doi.org/10.1038/525159a sarewitz, d. (2015). reproducibility will not cure what ails science. nature, 525(7568), pp.159-159. https://doi.org/10.1038/525159a scannell, j. and bosley, j. (2016). when quality beats quantity: decision theory, drug discovery, and the reproducibility crisis. plos one, 11(2), p.e0147215. https://doi.org/10.1371/journal.pone.0147215 https://doi.org/10.29173/iq967 http://blogs.discovermagazine.com/neuroskeptic/2015/09/07/100-percent-reproducibility/#.wqbyo2rwzaq http://blogs.discovermagazine.com/neuroskeptic/2015/09/07/100-percent-reproducibility/#.wqbyo2rwzaq https://doi.org/10.3768/rtipress.2017.bk.0020.1708 https://doi.org/10.1111/j.1740-9713.2015.00827.x https://doi.org/10.1109/mcse.2009.6 https://doi.org/10.1093/aje/kwj093 https://doi.org/10.1371/journal.pbio.2004299 https://doi.org/10.1038/nmeth.3076 https://doi.org/10.5220/0004991802230230 https://www.rd-alliance.org/groups/data-citation-wg/wiki/scalable-dynamic-data-citation-rda-wg-dc-position-paper.html https://www.rd-alliance.org/groups/data-citation-wg/wiki/scalable-dynamic-data-citation-rda-wg-dc-position-paper.html https://doi.org/10.1080/08989621.2016.1257387 https://doi.org/10.1080/19475683.2015.1027792 https://doi.org/10.1038/525159a https://doi.org/10.1038/525159a https://doi.org/10.1371/journal.pone.0147215 24/26 antognoli, erin; avila, regina; sears, jonathan; christiansen, leighton; tieman, jessica; hart, jacquelyn (2020) reproducibility literature analysis a federal information professional perspective, iassist quarterly 44(1-2), pp. 1-26. doi: https://doi.org/10.29173/iq967 schulz, j., cookson, m. and hausmann, l. (2016). the impact of fraudulent and irreproducible data to the translational research crisis solutions and implementation. journal of neurochemistry, 139, pp.253270. https://doi.org/10.1111/jnc.13844 sené, m., gilmore, i. and janssen, j. (2017). metrology is key to reproducing results. nature, 547(7664), pp.397-399. https://doi.org/10.1038/547397a shaw, s. and d’intino, j. (2017). evidence-based practice and the reproducibility crisis in psychology. national association of school psychologists, 45(5). available at: https://www.mcgill.ca/connectionslab/files/connectionslab/ebp_nasp_cq.pdf [accessed 10 sep. 2019]. spector, j., johnson, t. and young, p. (2014). an editorial on replication studies and scaling up efforts. educational technology research and development, 63(1), pp.1-4. https://doi.org/10.1007/s11423-0149364-3 stanford university (2019). meta-research innovation center at stanford | metrics. [online] metrics.stanford.edu. available at: https://metrics.stanford.edu/ [accessed 10 sep. 2019]. steward, o. (2016). a rhumba of “r’s”: replication, reproducibility, rigor, robustness: what does a failure to replicate mean?. eneuro, 3(4), pp.eneuro.0072-16.2016. https://doi.org/10.1523/eneuro.0072-16.2016 stodden, v. (2009). the legal framework for reproducible scientific research: licensing and copyright. computing in science & engineering, 11(1), pp.35-40. https://doi.org/10.1109/mcse.2009.19 stodden, v. (2009). the reproducible research standard: reducing legal barriers to scientific knowledge and innovation. presented at: communia conference 2009: global science & economics of knowledge-sharing institutions, torino, italy, june 2009. [powerpoint] stodden, v. (2014). 2014 : what scientific idea is ready for retirement? reproducibility. [online] edge.org available at: https://www.edge.org/response-detail/25340 stodden, v. (2015). reproducing statistical results. annual review of statistics and its application, 2(1), pp.1-19. https://doi.org/10.1146/annurev-statistics-010814-020127 stodden, v., bailey, d., borwein, j., leveque, r., rider, w. and stein, w. (2019). setting the default to reproducible. [online] stodden.net. available at: http://stodden.net/icerm_report.pdf [accessed 10 sep. 2019]. stodden, v., borwein, j. m., & bailey, d. h. (2013). “setting the default to reproducible” in computational science research. siam news, 46(05). stodden, v., guo, p. and ma, z. (2013). toward reproducible computational research: an empirical analysis of data and code policy adoption by journals. plos one, 8(6), p.e67111. https://doi.org/10.1371/journal.pone.0067111 stodden, v., leisch, f. and peng, r. (2014). implementing reproducible research. boca raton, florida: crc press. https://doi.org/10.29173/iq967 https://doi.org/10.1111/jnc.13844 https://doi.org/10.1038/547397a https://www.mcgill.ca/connectionslab/files/connectionslab/ebp_nasp_cq.pdf https://doi.org/10.1007/s11423-014-9364-3 https://doi.org/10.1007/s11423-014-9364-3 https://metrics.stanford.edu/ https://doi.org/10.1523/eneuro.0072-16.2016 https://doi.org/10.1109/mcse.2009.19 https://www.edge.org/response-detail/25340 https://doi.org/10.1146/annurev-statistics-010814-020127 http://stodden.net/icerm_report.pdf https://doi.org/10.1371/journal.pone.0067111 25/26 antognoli, erin; avila, regina; sears, jonathan; christiansen, leighton; tieman, jessica; hart, jacquelyn (2020) reproducibility literature analysis a federal information professional perspective, iassist quarterly 44(1-2), pp. 1-26. doi: https://doi.org/10.29173/iq967 stodden, victoria, enabling reproducible research: open licensing for scientific innovation (march 3, 2009). international journal of communications law and policy, forthcoming. available at ssrn: https://ssrn.com/abstract=1362040 stroebe, w. and strack, f. (2014). the alleged crisis and the illusion of exact replication. perspectives on psychological science, 9(1), pp.59-71. https://doi.org/10.1177/1745691613514450 the national academy of sciences arthur m. sackler colloquium on reproducibility of research: issues and proposed remedies (2017). reproducibility of research: issues and proposed remedies youtube. [online] youtube. available at: http://www.youtube.com/playlist?list=plgjm1x3xqek0ferdgkcyvybh8tkubtwfv [accessed 10 sep. 2019]. the national science foundation (2016). dear colleague letter: encouraging reproducibility in computing and communications research. [online] nsf.gov. available at: https://www.nsf.gov/pubs/2017/nsf17022/nsf17022.jsp [accessed 10 sep. 2019]. träger, u. (2019). going beyond impact factors—reforming scientific publishing to value integrity. [online] phys.org. available at: https://phys.org/news/2016-08-impact-factorsreforming-scientificpublishing.html [accessed 10 sep. 2019]. uncles, m. and kwok, s. (2013). designing research with in-built differentiated replication. journal of business research, 66(9), pp.1398-1405. https://doi.org/10.1016/j.jbusres.2012.05.005 vedantam, s. and penman, m. (2016). when great minds think unalike: inside science’s ‘replication crisis’. [online] npr.org. available at: https://www.npr.org/2016/05/24/477921050/when-great-mindsthink-unlike-inside-sciences-replication-crisis [accessed 10 sep. 2019]. voelkl, b. and würbel, h. (2016). reproducibility crisis: are we ignoring reaction norms?. trends in pharmacological sciences, 37(7), pp.509-510. https://doi.org/10.1016/j.tips.2016.05.003 warren, m. (2018). make replication studies ‘a normal and essential part of science,’ dutch science academy says. science. https://doi.org/10.1126/science.aat0224 weir, k. (2015). a reproducibility crisis? monitor on psychology: american psychological association, 46(9), p.39 wiggins, a., newman, g., stevenson, r. and crowston, k. (2011). mechanisms for data quality and validation in citizen science. 2011 ieee seventh international conference on e-science workshops. https://doi.org/10.1109/esciencew.2011.27 würbel, h. (2017). more than 3rs: the importance of scientific validity for harm-benefit analysis of animal research. lab animal, 46, pp.164–166 https://doi.org/10.1038/laban.1220 yousefi, m. and dougherty, e. (2012). performance reproducibility index for classification. bioinformatics, 28(21), pp.2824-2833. https://doi.org/10.1093/bioinformatics/bts509 https://doi.org/10.29173/iq967 https://ssrn.com/abstract=1362040 https://doi.org/10.1177/1745691613514450 http://www.youtube.com/playlist?list=plgjm1x3xqek0ferdgkcyvybh8tkubtwfv https://www.nsf.gov/pubs/2017/nsf17022/nsf17022.jsp https://phys.org/news/2016-08-impact-factorsreforming-scientific-publishing.html https://phys.org/news/2016-08-impact-factorsreforming-scientific-publishing.html https://doi.org/10.1016/j.jbusres.2012.05.005 https://www.npr.org/2016/05/24/477921050/when-great-minds-think-unlike-inside-sciences-replication-crisis https://www.npr.org/2016/05/24/477921050/when-great-minds-think-unlike-inside-sciences-replication-crisis https://doi.org/10.1016/j.tips.2016.05.003 https://doi.org/10.1126/science.aat0224 https://doi.org/10.1109/esciencew.2011.27 https://doi.org/10.1038/laban.1220 https://doi.org/10.1093/bioinformatics/bts509 26/26 antognoli, erin; avila, regina; sears, jonathan; christiansen, leighton; tieman, jessica; hart, jacquelyn (2020) reproducibility literature analysis a federal information professional perspective, iassist quarterly 44(1-2), pp. 1-26. doi: https://doi.org/10.29173/iq967 endnotes 1 erin antognoli https://orcid.org/0000-0003-0569-0808 is the metadata librarian and data curator, at the national agricultural library, united states department of agriculture, and can be reached by email: erin.antognoli@usda.gov 2 regina avila https://orcid.org/0000-0002-4340-2558 is the digital services librarian, at the national institute of standards & technology, and can be reached by email: regina.avila@nist.gov 3 jonathan sears https://orcid.org/0000-0002-9045-713x is the data scientist, lac group, on assignment at the national agricultural library, united states department of agriculture. 4 leighton l. christiansen https://orcid.org/0000-0002-0543-4268 is the data curator at the national transportation library, in the bureau of transportation statistics, at the united states department of transportation. 5 jessica tieman https://orcid.org/0000-0002-9547-0448 is the digital preservation librarian, at the u.s. government publishing office. 6 jacquelyn hart https://orcid.org/0000-0002-5408-6853 is the acquisitions & cataloging librarian, canada & oceania section, at the library of congress. 7 the name cendi was derived from original membership from departments of commerce, energy, nasa, and the defense information managers group. current membership includes several other federal agencies. https://doi.org/10.29173/iq967 https://orcid.org/0000-0003-0569-0808 https://orcid.org/0000-0003-0569-0808 mailto:erin.antognoli@gmail.com https://orcid.org/0000-0002-4340-2558 https://orcid.org/0000-0002-4340-2558 mailto:regina.avila@nist.gov https://orcid.org/0000-0002-9045-713x https://orcid.org/0000-0002-9045-713x https://orcid.org/0000-0002-0543-4268 https://orcid.org/0000-0002-9547-0448 https://orcid.org/0000-0002-9547-0448 https://orcid.org/0000-0002-5408-6853 vol274.indd by 16 iassist quarterly winter 2003 by angela dale * research access to microdata: an attempt to provide a context background national statistical institutes have an obligation to compile statistics that provide the information required by government. in the uk, following the 1980 review by sir derek rayner, the remit of the government statistical service was restricted to meet the specific needs of government departments rather than the broader needs of the business community, local government and academia. however, the launch of national statistics in june 2000 involved an explicit commitment to meet the needs of a broader range of users that included the general public. the framework document (june 2000) that accompanied the launch set out the governmentʼs commitment to providing a “statistical service that is open and responsive to societyʼs needs and the public agenda: better and more reliable official statistics that command public confidence.” under the aims and objectives of national statistics1, in section 3, the third bullet point lists: to provide researchers, analysts and other customers with a statistical service that assists their work and studies; however, statistical offices have to tread a careful balance between providing the data needed by all sections of society and maintaining the confidence of the general public who supply most of the data. the experience of some other countries shows that if the public lose confidence in the national statistical office then the process of data collection will be undermined and may not recover. for example, germany has not taken a full population census since the census planned for 1983 had to be postponed until 1987 because of public concern over proposals to use census returns to update the local population registers. the netherlands has not taken a census since 1971, following a significant level of refusal in the 1971 census and poor test results in 1979. data in the public domain in the past the general public only had access to government statistics through reports in local libraries. however, in recent years greater dissemination by statistical offices, largely through the opportunities offered by the web, have brought statistical information into the homes of a large sector of the population and into the offices of voluntary organisations, schools and other locally based organisations. particularly through the development of neighbourhood statistics (which includes data from the 2001 census, surveys and also administrative sources), there is now readily accessible information about the places where people live and work. in addition, there is also unrestricted on-line access to reports on the social and economic conditions of the population and the tabulations that underpin them – for example, the living in britain report produced annually by the office for national statistics (ons). there is, therefore, a developing reciprocal relationship between the population that provides the data and the statistical office which collects and compiles that data. for the first time the average person in the street, or student in school (as well as businesses and local authorities) is able to obtain recent and high quality data from the uk statistical offices without charge. the very high rate of hits on the ons web-site, and neighbourhood statistics in particular, suggests that the public are, indeed, accessing these data. this development should be an important step towards retaining and increasing public acceptance of the conduct of the census and government surveys. however, the very fact that these data are public and easily available means that they must not reveal any identifiable information, either now or at some unforeseen time in the future. but it is next to impossible to predict what technologies or techniques may become available in the future that could lead to the identification of individuals and what motivations there may be for using them. therefore the balance between providing a public service by making data easily available and ensuring the confidentiality of the data is very difficult to get right. an additional and little-researched factor is the impact of public perceptions. a wrong belief that people can be identified in government statistics may be as damaging to public confidence as the reality – and, for many people, the two may not be distinguished. what evidence is available [1] suggests that people are unsure about the extent to which information they supply in a census is passed to iassist quarterly winter 2003 17 other government departments and this is also the case in the usa and australia. there is also confusion about the source of information used in direct marketing and whether or not it comes from government data sources. protecting confidentiality it is widely accepted that geographical detail is a key factor in identifying individuals. in small geographical areas (e.g. the 2001 census output areas with about 125 households) residents are likely to have good knowledge of the characteristics of their neighbours. in this size of area there may only be one woman aged 45 who is living in privately rented accommodation or only one man of black caribbean origin who works in education. to ensure that such an individual cannot be identified, much less detail on characteristics such as industry, occupation, age and ethnic group can be provided for small geographical areas than for larger areas. this is reflected in census outputs, when tables at local district level have more detail than those at the level of output area. in addition, ons have added protection to tables from the 2001 census that have small cell sizes. cells containing 0, 1, 2 or 3 respondents have been changed to 0 or 3. for many members of the public and many researchers, information about local areas is what is required. where information on national or regional social and demographic characteristics is needed then tables are available on a range of topics. the role of academic research in the social sciences however, for many researchers these publicly available data sources provide only a first port of call. academic research needs to go beyond published reports and pre-prepared tables to conduct original research using microdata (that is, individual records for individuals and households). academic social research has a vital role to play in understanding social change. it can provide methodologically rigorous analysis of issues that are of fundamental importance: for example the household composition of the ageing population, migration patterns, ethnic diversity, regional differentiation and much more. academic analysis can go beyond the descriptive to seek explanation and to test hypotheses. multivariate analysis is needed that includes all variables of importance to the outcome of interest. furthermore, these variables need to be derived in a way appropriate to the analysis. for example appropriate age groupings will vary by whether one is analysing labour market activity or family formation. bespoke classifications need to be developed that are specific to a particular analysis – for example measures of exclusion based on information about all household members. existing classifications or indicators need to be subject to challenge and to re-working based on different definitions. at the heart of scientific research is the requirement that results are published and open to challenge. the ability to replicate analyses is fundamental to good scientific practice. it is also essential that data collected at public expense is used as extensively as possible, consistent with the undertakings given to the respondents. in this spirit, the results of research should be available in an accessible and reader-friendly form as well as through publications in scientific journals. analysis of microdata files from the 1991 uk census has had a major research impact, including analyses of unemployment that allow both individual and area-level characteristics to be included [2] and analysis of ethnic differences in womenʼs employment over the life course [3]. a summary of this research is available from the ccsr web site (www.ccsr.ac.uk/sars/findings). however, there is an increased risk of identification with microdata by comparison with pre-defined tables, and this is recognised in the procedures used to ensure that confidentiality is protected. the first protection is that microdata files represent only a sample of the population. therefore there is only a small chance – perhaps 2 or 3 in 100 that an individual will be included. in addition, care is taken over the amount of detail that can be released and geographical detail is always heavily restricted. finding the appropriate balance requires careful assessment of the risk of data disclosure. but it also requires recognition that absolute safety jeopardises any significant research activity. therefore the risk of not supporting research also has to be considered. it is also worth noting that, where breaches of confidentiality have occurred, (see above) these have not been associated with research use of data. safety: a double balancing act we can define two interacting dimensions when considering access to data the level of safety associated with the dataset; and the level of safety associated with the access setting. level of safety associated with dataset this will depend heavily on the degree of detail in the data; the proportion of the population in the sample; the ease of identifying the data either through matching or spontaneous recognition. thus a microdata file with a low level of risk may be a sample with very restricted individual detail and little geographical information. level of risk will also vary with the extent to which disclosure protection methods (e.g. perturbation or data swapping) have been used on the data. level of safety associated with access setting this will range from access confined to a safe setting within the statistical office – at one extreme – to unrestricted access where data is distributed to users with few if any conditions of use. 18 iassist quarterly winter 2003 the two dimensions interact so that, at one extreme, if the data are judged to be entirely safe, then the access arrangements can be very open. this is exemplified by the public use microdata files produced by the us bureau of the census, which can be downloaded without restriction from the web-site of the us bureau of the census. these files are samples – 1% and 5% where the amount of both individual detail and geographical information has been heavily restricted to preserve confidentiality. by contrast, if the data are very detailed and/or contain information that could be used to identify someone, then greater safety needs to be built into the access conditions. an example is the ons longitudinal study that contains data with a great deal of individual and geographical detail, from the census and from vital events, but where access is highly restricted and only available within a secure setting inside ons. we have, therefore, a continuum from safe data to safe setting – with all protection built into the data in the former and all protection built into the setting in the latter. research and safety public use microdata files are of considerable value because they can be readily used anywhere at any time. access is quick and easy and these kinds of data are ideal for teaching, where students need to interact with data. however, datasets that are safe enough to need no restrictions will usually lack some of the detail required by researchers. for example, in safe data variables such as occupation or ethnic group may be very broadly banded and thus may not provide the distinction required for some analysis purposes. a lack of geographical detail may also hamper research into the respective effects of individual characteristics and local labour markets. some bias may also have been introduced into the data through perturbation or suppression, in order to ensure that unusual individuals or households cannot be recognised. these are all concerns which have been addressed in the development of microdata samples from the 2001 census. at the other end of the spectrum, secure in-house access, e.g. within ons, where researcher credentials are screened, all data is available under strictly controlled conditions and all outputs are carefully checked, can allow access to much more detailed data. in this kind of safe setting the analyst may be able to access detailed geographical information on place of residence or place of work, or data that is very sensitive – for example information on cause of death, cancer registration or, in the case of business surveys, information on business performance. however, in-house safe settings are expensive to set up and run and also difficult for researchers who have to travel long distances and spend considerable time away from home. finding the middle ground the two extremes of safe data and safe setting both have disadvantages for conducting research. we therefore need to explore a range of options that lie between these polar opposites and that can allow researchers access to data that is of sufficient detail and quality to meet research needs while also retaining the level of confidentiality required by the national statistical institutes. fundamental to this middle ground is the need to recognise that researchers have no interest in breaching confidentiality. research is concerned with establishing statistically significant differences between social and demographic groups, not with attempting to identify individuals. researchers do, however, have a very strong interest in promoting good practice and respect for research data. the safeguards set out below provide varying degrees of protection and can be used singly or together to increase data protection beyond that required for public use files. they should therefore allow a concomitant increase in detail in the data. the role of institutional controls research is conducted in recognised institutions (one definition of a research institution is recognition to administer research grants). these institutions can be asked to accept responsibility for research data used by their staff. this control was used in the uk with dissemination of the samples of anonymised records from the 1991 census. institutions where staff or students wanted to use the data were asked to identify a responsible person who actively managed data access. microdata under licence statistics netherlands provides access to microdata for research purposes under license. researchers in the uk who wish to use microdata from the data archive are required to agree to a confidentiality undertaking. however, this could be extended to provide a more explicit and binding contract between the researcher and the statistical office. this would include use of the data for a fixed length of time and a requirement to return all copies of the data after that time. a safe setting on-site in canada and the usa, statistical offices are increasingly setting up secure data centres for analysis of microdata files. these represent safe settings that, for the researchers who happen to be located nearby, can provide access to the most detailed microdata. whilst these settings can provide very safe conditions, they are expensive to run and privilege those able to use the facility. nonetheless, it is possible to imagine a situation where most universities could support a safe room that would allow access to relatively detailed microdata. there are established iassist quarterly winter 2003 19 procedures for access controls to prevent data being removed from the room. this should be far enough along the safe setting spectrum to allow access to much more detailed microdata than that released as public use files. increased use of technological developments there are a growing number of examples of remote safe-settings where microdata files are held on a secure server that may be located in a statistical office or any other safe location. access to the data can be indirect – as with the luxembourg income study, where the researcher submits a request to run an analysis; the request is physically downloaded and moved across a firewall to a secure server holding the data. the results, which are controlled to prevent disclosure, are then returned by email and the researcher has no access to the actual microdata. alternatively, researchers may be able to interrogate data files through the use of additional controls such as a password authorisation system backed up by a license agreement and registered ip addresses for authorised computers. increasingly, the grid and associated middleware allow imaginative solutions that can maximise research use whilst retaining confidentiality. conclusions in a time of increased concern over data security there is a growing need to explore all possible ways in which data collected at public expense can be fully analysed, while at the same time ensuring the confidentiality of the respondents. in the spirit of ensuring that there is some payback to the public who provide responses to censuses and surveys, there is a strong argument that accessible research findings should be posted on national statistics web-sites. by doing so, we would make the value of research based on government data more apparent to all. [1]fieldhouse, e. and gould, m. i. (1998) “ethnic minority unemployment and local labour market conditions in great britain,” environment & planning a 30, no.5, 833-53. [2]framework for national statistics, june 2000, http:// www.statistics.gov.uk/about_ns/downloads/framedoc1.pdf [3]holdsworth, c. and dale, a. (1997) “ethnic differences in womenʼs employment,” work, employment and society 11, 435-57. [4]marsh, c. (1993) “privacy, confidentiality and anonymity in the 1991 census” in dale, a. and marsh, c. (eds) the 1991 census userʼs guide, london: hms notes 1 section 3 of the ons framework document http://www.statistics.gov.uk/about_ns/downloads/ framedoc1.pdf an earlier version of this paper was published in significance, issue 1, of the royal statistical society we are grateful for permission to reproduce it here i am grateful to chuck humphrey, university of alberta and joris nobel, statistics netherlands for helpful comments and suggestions. * angela dale, ccsr, university of manchester. contact: angela.dale@man.ac.uk http://www.statistics.gov.uk/about_ns/downloads/framedoc1.pdf http://www.statistics.gov.uk/about_ns/downloads/framedoc1.pdf http://www.statistics.gov.uk/about_ns/downloads/framedoc1.pdf http://www.statistics.gov.uk/about_ns/downloads/framedoc1.pdf 1/10 hansen, zaza nadja lee; kruse, filip; thestrup, jesper boserup (2019) managing data in cross-institutional projects, iassist quarterly 43(3), pp. 1-10. doi: https://doi.org/10.29173/iq950 managing data in cross-institutional projects zaza nadja lee hansen1, filip kruse2, jesper boserup thestrup3 abstract this paper provides guidelines for data management professionals and researchers on how fair data usage can help improve the planning, execution and overall success of a cross-institutional project. cases from danish cross-institutional projects are detailed to illustrate this point – as well as the lessons learnt with implementing fair data principles in such projects. key learnings from this paper are: • using fair data principles in cross-institutional projects can help manage the data used in the project in terms of knowledge sharing, access rights, use of templates, metadata and further sharing the data after the project has ended. • to benefit the most from using fair data in a cross-institutional project it should be considered and planned for early in the project process. • if fair is not considered early in the project process problems can arise such as a lot of time spent on converting formats, obtaining permissions and assigning metadata. • it is necessary for researchers and research projects to have infrastructure and other services in place which support fair data usage. keywords fair, data management, cross-institutional projects 1. introduction in order for research institutions to be competitive and to transfer knowledge into social and commercial gain there is an increasing need to share research data between multiple actors. however, when research data has to be shared across institutional boundaries challenges arise due to project members being used to different data management approaches. fair data management is a requirement in many research projects, including horizon 2020 applications (european commission, 2018). the fair data principles (force11, 2018) state that data has to be findable, accessible, interoperable, and re-usable; in other words that data should be as open as possible. however, some data cannot be made fully open – like healthcare data. however, it could be open to some degree. hence, in the projects discussed in this paper, in particular the fair across project, the participants were told to follow fair principles and make data as open as possible – but as closed as needed. in other words, given a good reason – for example that it would be a breach of the general data projection regulation (gdpr), it is possible to make more data (semi-) open as openness is then seen on a scale and not as an either/or decision which for security https://doi.org/10.29173/iq950 2/10 hansen, zaza nadja lee; kruse, filip; thestrup, jesper boserup (2019) managing data in cross-institutional projects, iassist quarterly 43(3), pp. 1-10. doi: https://doi.org/10.29173/iq950 reasons can make sure stakeholders go for the “safe” option and then say no to completely open data. by implementing the fair principles data can more easily be shared in a cross-institutional project, thus making it more likely that project deliverables are done on time and give the expected outcome. this can be done by ensuring that all data related to the project follows some specific guidelines, standards and formats, agreed upon at the project start. examples include where and how to store and share data, who is responsible for updating which data, which data format to use, including how and where to share metadata. this paper will describe how the fair principles can be applied to research data with a specific focus on cross-institutional research projects. furthermore, a series of adaptation cases for usage of fair data principles at danish universities will be presented. 2. the relevance of fair data in cross-institutional projects 2.1. the fair across project in 2018 a working group was created under deic – danish e-infrastructure cooperation with representatives from aalborg university, copenhagen university, the royal library, technical university of denmark, copenhagen business school and the danish national archives. this working group investigated how fair data principles could be used to encourage researchers to share data across data types, disciplines and institutions. the working group interviewed researchers at several of the 8 danish universities in order to create guidelines and material on the fair principles to help aid researchers share data when taking part in cross-institutional projects. 2.2. key findings for relevance of fair data in cross-institutional projects fair can be used to contribute towards the success of a cross-institutional project in three main ways: 1. project planning of cross-institutional projects 2. navigating project changes in cross-institutional projects 3. sharing information in cross-institutional projects 4. sharing data outside the project group the primary focus areas within the fair data principles are making data findable and accessible so everyone in the project has access to the same data at the same time. the key findings from the project can be seen in table 1. https://doi.org/10.29173/iq950 3/10 hansen, zaza nadja lee; kruse, filip; thestrup, jesper boserup (2019) managing data in cross-institutional projects, iassist quarterly 43(3), pp. 1-10. doi: https://doi.org/10.29173/iq950 table 1: summary of benefits of fair data usage in cross-institutional projects how fair data usage can help expected benefits project planning of crossinstitutional projects • define responsibilities for data management among the project partners, e.g. in a data management plan. • agree on common standards / conventions for collecting, storing and documenting the data. • ensure that all data from the project are shared on a common, secure platform. • you save time introducing new members to the project. • you avoid data loss when someone leaves the project. • you enable easy reuse of data generated so far. navigating project changes in crossinstitutional projects • document methods and decisions in a systematic manner, e.g. using common templates • use common standards for collecting, storing and documenting data, e.g. on file names, file formats, table contents, etc. • store all data in a structured way on a shared and secure system that is accessible for the project members. • you can extend your analysis by finding new suitable datasets more quickly. • you can reuse existing datasets more easily for e.g. new analysis. • you minimize the risk of wasting time and money on duplicate work. sharing information in crossinstitutional projects • make agreements on how research data are managed and shared between research partners. • establish clear guidelines on how to document data in a consistent form, e.g. as templates for tables or notebooks. • use a secure shared storage system with controlled access management for all data. • you avoid misunderstandings regarding access to and use of the data. • you foster collaboration to make best use of the data, speed up the processes and reduce administration. • you improve the quality of the data and thus the scientific output from the project. sharing data outside the project group • agree on a procedure for when data or findings are considered “complete”/”valid” and thus ready to be shared outside the project working group. • establish clear guidelines on the criteria for data to be fully open, semi-open or closed, keeping in mind • you make as much data open as possible – without fear of security issues/gdpr concerns • you foster further use of the data beyond the lifetime of the given project https://doi.org/10.29173/iq950 4/10 hansen, zaza nadja lee; kruse, filip; thestrup, jesper boserup (2019) managing data in cross-institutional projects, iassist quarterly 43(3), pp. 1-10. doi: https://doi.org/10.29173/iq950 that the goal is to make as much data as open as possible. • include an it security expert or gdpr expert or dpo as needed in creating the above guidelines. • document the decisions for data openness in the project • agree on what platforms/media data should be released on and a person responsible for this task the findings were validated through interviews, workshops and a conference with data management professionals and researchers. 3. fair implementation cases 3.1 a brief outline of the project data management in practice the project data management in practice (dmip)4 aimed at providing danish researchers with an operational infrastructure for research data management. the general objective was to establish a danish infrastructure setup with services covering all aspects of the lifecycle of research data: from application and initial planning, through discovering and selecting data and finally to the dissemination and sharing of results and data. further, the setup should include facilities for training and education in research date management. researchers’ needs and demands should form the basis of the services and close connection to and cooperation with research projects already in progress was an integral part of the project setup hence the “in practice”. finally, the project should explore the role of research libraries regarding research data management. the project organization can be described as a hybrid middle path between a purely case-based project with individual institutions each working on their own sub-projects, and a thematic project with institutions working within one or more broad themes. the structure chosen contained six themes: data management planning; data capture, storage and documentation; data identification, citation and discovery; select and deposit for long-term preservation; training and marketing toolkits; and sustainability. each of the participating institutions worked on specific cases, such as ongoing research projects, well-defined data collections etc., all cases covering the entire data lifecycle. each case should relate to the themes to be able to draw conclusions on both a more case-specific and a more general theme-specific level. the cases spanned the main academic fields of humanities, social sciences, science and technology. for an overview of the project see kruse & thestrup (2018). below we shall explore selected cases and analyse which challenges fair data principles posed in the practical data management process, how these challenges were met and to what extent. one important lesson from the process is that fair data principles have to be an integral part of the data management plan and the project plan right from the outset. https://doi.org/10.29173/iq950 5/10 hansen, zaza nadja lee; kruse, filip; thestrup, jesper boserup (2019) managing data in cross-institutional projects, iassist quarterly 43(3), pp. 1-10. doi: https://doi.org/10.29173/iq950 3.2 identifying case-specific examples of fair data principles 3.2.1. the larm case the aim of the larm case was to enable researchers and students to use radio programs as research objects by establishing a digital archive with an associated research infrastructure with relevant tools (see hansen et al. 2018, pp. 12-15) the royal danish library’s larm, the sound archive for radio media, offers researchers the larm.fm ( https://www.larm.fm/) platform as the workspace for research projects using radio data. this platform enables researchers to work with the material, add metadata, annotate and structure selections of items etc. an important aim of the project was to provide facilities to overview the new data, a solution for long-term preservation of data (location and format) and to ensure future reuse of data for other research projects. among the results of the larm case, the following is of special interest: the development of the royal danish library’s dmp template (see hansen, (2018b, 26-355 and 54-566),, which is a central element in the danish version of dmponline (https://dmponline.deic.dk/), preparations for a future data harvesting facility, and technical decisions on data formats etc. the larm.fm platform requires the researcher to confirm that all data she enters into the system such as new metadata, annotations etc. are shareable under a cco license. data entered is not issued with an author identification, which is a drawback, the data have an internal id, while the metadata for the program includes a ‘doms-id’, facilitating location in mediestream.dk7. data sharing as an element of data identification, citation and discovery is not possible, however, for this data, as some of the media data is sensitive or protected by copyright. the annotations for the programs and the metadata for the annotations often also contain personal information or program descriptions from e.g. new agencies. this has led to the contingency of designing of a data repository with restricted access, which would still allow for reuse of data for research purposes, library controlled access repository (lcar). some of the challenges outlined above result from the data collections being large and inhomogeneous. in practical research, it will be possible for the researcher to select, download and share smaller data sets, after ensuring that the data is not protected or restricted in use, thus meeting several fair criteria. the royal danish library’s library open access repository (loar, https://loar.kb.dk/) will be well suited for this task. 3.2.2. the netlab case the aim of the netlab case was to use the royal danish library’s webarchive, netarkivet.dk, to map the historical development of the danish web (see hansen et al. 2018, pp. 16-20). to achieve this aim it was necessary to develop an infrastructure with provisions for corpus creation, tools, workspace, storage etc. the royal danish library’s webarchive, netarkivet.dk, poses much the same problems for the researchers as larm. as material and data harvested from the web contain personal information or copyrighted material, the archive is only accessible to researchers who have requested and been granted special permission to use the collection for specific research purposes. analogous to access to the archive, sharing of data generated through research in the archive is prohibited if it contains https://doi.org/10.29173/iq950 https://www.larm.fm/ https://dmponline.deic.dk/ https://loar.kb.dk/ 6/10 hansen, zaza nadja lee; kruse, filip; thestrup, jesper boserup (2019) managing data in cross-institutional projects, iassist quarterly 43(3), pp. 1-10. doi: https://doi.org/10.29173/iq950 personal information or material otherwise restricted. deposit of such data in lcar will still be an option. regarding data identification and discovery the loar and lcar repositories mentioned above supplies data with datacite dois thus making data searchable through datacite’s services. by the use of oai-pmh metadata is available for harvesting or searching through the danish data archive https://www.sa.dk/en/ both larm and the danish netarchive ask the question of ownership of data. the royal danish library owns data in the netarchive and collective or individual rights owners own data in larm. research data created from these sources can be the product of individual or group efforts. in order to clarify legal issues of ownership and copyright to research data a legal framework was drawn up, the model agreement on data management in cooperative research projects8. based on the two cases outlined above we may conclude firstly that to make data fair it is necessary to take these principles into account at the earliest possible stage of the research process. all the solutions were developed ex post as answers to challenges occurring in the course of ongoing research projects. secondly, that while access to research data for obvious reasons can be restricted without necessarily preventing its reuse care should be taken to make metadata findable and accessible for external users and non-researchers. these external users should for example be able to search metadata and should be given information on how to obtain access to data if they find data, which is not public accessible. 3.2.3. the kierkegaard case in the kierkegaard case (hansen et. al. 2018a, pp. 21-25) the project participants were involved in a process keeping data available for future researchers. søren kierkegaard’s collected writings were republished from 1997 to 2009 in 55 volumes. as part of this publication an electronic version was made available on www.sks.dk. the online version is based on a format called kierkegaard normalformat (kn1). this format was developed as part of the publication process. the goal of this dmip case to preserve the data in a format, which would allow researchers to annotate and collaborate with the text and ensure that the data can be reused in 50 years (hansen et. al. 2018, p. 23). this goal was achieved by reformatting the data into tei (text encoding initiative9). tei is an international standard, which allows long-term preservation and gives researchers the possibility of collaboration. about 80 % could be reformatted automatically. the remaining 20% had to be manipulated by a computer scientist in order to ensure that all data could be saved in the tei format (hansen et al. 2018, p. 23). in order to share the data the project group tested a dataverse10 server. dataverse is a software, which makes it possible to upload and share data. the data is now available if a researcher requests access. but the data is stored on a platform called erda11 which offers no direct access to the data for the general public. since the data is not shared the datasets are not issued with a doi and metadata has not been shared (hansen et al. 2018a, p. 24). https://doi.org/10.29173/iq950 https://www.sa.dk/en/ http://www.sks.dk/ 7/10 hansen, zaza nadja lee; kruse, filip; thestrup, jesper boserup (2019) managing data in cross-institutional projects, iassist quarterly 43(3), pp. 1-10. doi: https://doi.org/10.29173/iq950 the case underlines that fair as a principle need to be taken into consideration early in a given process in order to ensure that data can be reused. it must be planned to store the data in an open format, in an early phase of the research project. the case shows that it can be necessary to allocate considerable resources to make data fair, especially in a case like this where the data format has been designed years ago. the case also shows that the data must be stored within an infrastructure that can add adequate and open metadata to make data findable. 3.2.4. the dtu wind energy and dtu space cases two of the science case of the dmip project concerned data from dtu wind energy (http://www.vindenergi.dtu.dk/english) and dtu space (http://www.space.dtu.dk/english) . regarding the wind energy project, the project group should “document, catalogue and archive datasets to make them available wherever possible”. the staff worked at the dtu bibliometrics and data management office (hansen et al, 2018, p. 38). the data was gathered over a period of at least 20 years and contains information, which can be used by other researchers, governments and industry. the group realized that the datasets consist of very different types of data such as measurements, experiments and models. the data was stored on different types of media and without sufficient metadata (hansen et al., 2018, p. 37). the project group worked with several topics. evaluation and usage of an existing data management plan (dmp) template was used to define criteria for repositories and longtime preservation. then the group compiled and shared two datasets via zenodo12 (hansen et al. 2018, pp. 40-42). based on the case the dmip report concludes that a dmp describing the data and how to store and share is easy to create via the tool dmponline, but requires a good template for any given dmp. zenodo uses the datacite metadata schema and offers the possibility of exporting the metadata via oai-pmh. this allows for harvesting and transfer of metadata to other data catalogues. the use of a wide array of data types and formats stresses the importance of choosing a metadata standard suitable for data discovery now and in the future.. the case showed that fair principles have to be part of the research project as early as possible in order to ensure that data can be shared together with relevant metadata. the data from the dtu space case contained information on the magnetic field of earth collected by 17 ground stations. the stations are located in the south atlantic, greenland and denmark. the data is transferred from the stations to servers on dtu space. the goals in this dmip case was to “increase visibility of the group’s research”, to give access to data and make the data citable (hansen et al. 2018a, 43-44). as part of the process, the researchers produced several versions of a dmp. the final dmp was used to design the work of the dmip project case. an infrastructure to share the data was described, but as of 1. april 2019, it is not yet in operation, but the existing infrastructure is supplemented with necessary guidelines etc. the researchers have begun uploading their data to the dtu data repository (http://data.dtu.dk) adding dois, thus improving data discovery. the case gave the project members improved insight in how to plan a workflow and share data, which can be used when other research projects are planned. regarding fair principles, the case demonstrated that https://doi.org/10.29173/iq950 http://www.vindenergi.dtu.dk/english http://data.dtu.dk/ 8/10 hansen, zaza nadja lee; kruse, filip; thestrup, jesper boserup (2019) managing data in cross-institutional projects, iassist quarterly 43(3), pp. 1-10. doi: https://doi.org/10.29173/iq950 issues such as infrastructure and metadata must be taken into consideration in order to be able to share data. 3.2.5. the kepler case since 2008 the kepler mission (https://www.nasa.gov/mission_pages/kepler/main/index.html) collected data on stars and extrasolar planets orbiting the stars. this has generated over 100 tb of data, which must be shared and preserved for future research. the dmip project explored how to ensure longtime preservation and sufficient metadata. the case showed that researchers have to consider issues like formats, metadata and infrastructure as a whole, and thus also the fair principles, very early in the research project. to share and longtime preserve 110 tb of data requires knowhow on formats and infrastructure in order to ensure the desired outcomes for example the suggested solution to ensure longtime preservation was to store data in a format called bagit13, where all information regarding a given star would be stored as individual datasets with a unique doi per dataset. 3.3 services established as a result of the dmip project several services were established in order to provide researchers with access to data infrastructures: the danish version of dmponline and the open access repository loar (library open access repository). loar offers 5 years preservation of up to 10 gb of research data free of charge for researchers from danish universities. preservation for longer periods or larger data sets is possible for a fee. researchers are expected to share the data using creative commons licenses. thus loar facilitates reuse of data. a restricted access repository, lcar (library restricted access repository) is developed but as of april 15. 2019 awaits decision for launching. both of the latter services use the free open source software dspace (https://duraspace.org/dspace/) for the building of repositories and are similar in structure and terms of service, except for terms of access. further, in order to create a common legal framework a model agreement14 on handling data was designed (hansen et al. 2018a, p. 8). as can be seen from the cases mentioned above, fair has to be integrated in the planning of any given research project in order to ensure access to data with a minimum of work. this requirement relates closely to data formats, metadata, and legal constraints for access, but also access to infrastructure and general legal frameworks. 4. discussion implementation of the fair data principles can help a cross-disciplinary project to deliver better results and solutions with more efficient use of project resources. fair data principles support and facilitate knowledge discovery and sharing as roles and responsibilities are clearly indicated and data is delivered according to standards and issued with the necessary metadata markings. however, we saw that it is vital that the project participants agree to use fair principles from the start of the project and that the infrastructure is in place as it is very time-consuming and costly to add this later. the problem we repeatedly encountered during the dmip project in relation to making data from the cases in the project fair was that data can be sensitive, contains private or personal information or is protected by copyright. such ethical and legal issues can to some extent be addressed from the https://doi.org/10.29173/iq950 https://www.nasa.gov/mission_pages/kepler/main/index.html https://duraspace.org/dspace/ 9/10 hansen, zaza nadja lee; kruse, filip; thestrup, jesper boserup (2019) managing data in cross-institutional projects, iassist quarterly 43(3), pp. 1-10. doi: https://doi.org/10.29173/iq950 start. but not if the use of this type of data occurs during the project. temporary solutions could be necessary to consider: anonymization, access to limited parts of the data, access to metadata only, etc. this could be helpful to researchers interested in the project. the use of open access repositories such as the royal danish library’s loar could also encourage the researcher as data owner to incorporate fair principles also in the process of choosing repository. again, this stresses the importance of issue of infrastructure. 5. conclusion this paper has provided guidelines for data management professionals and researchers on how fair data principles can help improve the planning, execution and overall success of a cross-disciplinary project and have shown cases of how fair data principles can be used and the benefits and challenges with doing so. in conclusion, the implementation of fair data principles can be a useful tool to help ensure the success of a cross-disciplinary project and is a great tool to aid in knowledge sharing as well among the project participants as with external parties. however, it is important to decide upon using fair principles before or at least as early on in the project as possible to lessen cost and minimize difficulties in their use and also to ensure that infrastructure and services are in fact in place to support this type of knowledge sharing between the institutions participating in the project. acknowledgements thanks to anne sofie fink kjeldgaard, section head at the danish national achieves, for input, feedback and general support. reference list force11. fair data principles, https://www.force11.org/group/fairgroup/fairprinciples. accessed 31/08/2018. hansen, k. k. et. al. (2018a). data management in practice results and evaluation. copenhagen: deff. doi: 10.7146/aul.243.174. http://ebooks.au.dk/index.php/aul/catalog/book/243 hansen, k. k. et. al. (2018b). data management in practice supplementary files. copenhagen: deff. doi: 10.7146/aul.244.175. http://ebooks.au.dk/index.php/aul/catalog/book/244 european commission (2016). h2020 programme: guidelines on fair data management in horizon 2020. http://ec.europa.eu/research/participants/data/ref/h2020/grants_manual/hi/oa_pilot/h2020-hi-oadata-mgt_en.pdf . accessed 04/15/2019 https://doi.org/10.29173/iq950 https://www.force11.org/group/fairgroup/fairprinciples http://ebooks.au.dk/index.php/aul/catalog/book/243 http://ebooks.au.dk/index.php/aul/catalog/book/244 http://ec.europa.eu/research/participants/data/ref/h2020/grants_manual/hi/oa_pilot/h2020-hi-oa-data-mgt_en.pdf http://ec.europa.eu/research/participants/data/ref/h2020/grants_manual/hi/oa_pilot/h2020-hi-oa-data-mgt_en.pdf 10/10 hansen, zaza nadja lee; kruse, filip; thestrup, jesper boserup (2019) managing data in cross-institutional projects, iassist quarterly 43(3), pp. 1-10. doi: https://doi.org/10.29173/iq950 kruse, f. and j. b. thestrup (2018). data management in practice – knowing and walking the path. ercim news, no. 114: 40-41. https://ercim-news.ercim.eu/en114/r-i/data-management-in-practiceknowing-and-walking-the-path. accessed 04/15/2019 1 project manager at the danish national archives 2 senior advisor at the royal danish library 3 communications officer at the royal danish library 4 the project partners: ruc roskilde university, kb the royal library (merged in 2017 with the state and university library as the royal danish library), dda danish data archive, dtic – dtu library, technical information center of denmark, sb state and university library, now the royal danish library, aub aalborg university library, sub university library of southern denmark. deff, denmark’s electronic research library and the participating institutions funded the project evenly. the project period was march 2015 june 2017, final report january 2018: hansen et al.: data management in practice, results and evaluation, available at: http://ebooks.au.dk/index.php/aul/catalog/book/243 accessed 04/15/2019. 5 in danish 6 in danish 7 mediestream.dk provides online access to danish cultural heritage, newspaper, radio and television programs etc. doms (digital object management system) is an in-house custom-built system for preservation of cultural heritage metadata. 8 http://www.au.dk/samarbejde/erhvervssamarbejde/samarbejde-med-forskere/modelaftale-forsamarbejde-om-forskningsdata/ the model agreement has been approved by aarhus university. 9 http://www.tei-c.org/index.xml, accessed 31/08/2018. 10 see more on the software here: https://dataverse.org/. accessed 31/08/2018. 11 copenhagen university’s electronic research data archive (erda) 12 http://doi.org/10.5281/zenodo.160136 and http://doi.org/10.5281/zenodo.161966. both accessed 04/15/2019. 13 bagit is a file packaging format designed for storing and transferring digital content. http://www.digitalpreservation.gov/series/challenge/data-transfer-tools.html. accessed 04/10/2019 14 https://www.deic.dk/da/news/2017-08-15/modelaftale. only in danish. accessed 04/15/2019. https://doi.org/10.29173/iq950 https://ercim-news.ercim.eu/en114/r-i/data-management-in-practice-knowing-and-walking-the-path https://ercim-news.ercim.eu/en114/r-i/data-management-in-practice-knowing-and-walking-the-path http://ebooks.au.dk/index.php/aul/catalog/book/243 http://www.au.dk/samarbejde/erhvervssamarbejde/samarbejde-med-forskere/modelaftale-for-samarbejde-om-forskningsdata/ http://www.au.dk/samarbejde/erhvervssamarbejde/samarbejde-med-forskere/modelaftale-for-samarbejde-om-forskningsdata/ http://www.tei-c.org/index.xml https://dataverse.org/ http://doi.org/10.5281/zenodo.160136 http://doi.org/10.5281/zenodo.161966 http://www.digitalpreservation.gov/series/challenge/data-transfer-tools.html https://www.deic.dk/da/news/2017-08-15/modelaftale vol28-4.indd by 8 iassist quarterly winter 2004 by chiu-chuang (lu) chou 1 my three-day encounter with argus: a report for iassist have you heard of the story about argus in greek mythology? he was a giant with one hundred eyes. hera, zeus’s wife, assigned him to guard zeus’s mistress io. eager to get io back, zeus sent hermes to kill the giant. argus was then transformed into a peacock. argus is also the name for a practical tool to guard data. it has been developed by the computational aspects of statistical confidentiality (casc) project, which is part of the fifth framework of the european union. argus was developed under windows nt and runs under windows versions from windows 95. it intends to “modify unsafe data in such a way that safe (enough) data emerge, with minimum information loss,”2 so that data producers can safely release the data to researchers and the public. argus has two components in achieving the statistical disclosure control (sdc). µ-argus is for safeguarding microdata, while τ-argus is designed to make tabular data safer. as a data librarian at a large research-oriented university, i am well aware of the confidentiality issues related to social science data sets. since the data and program library service, where i work, only houses and disseminates public use data sets, i am interested in what the principal investigators need to do before they can safely release their data. in this report, i will summarize what i have learned in a three-day workshop called statistical disclosure control for data confidentiality. this workshop was hosted by the center for demography of health and aging (cdha) from november 10-12, 2004 at the pyle center in the university of wisconsin madison. anco johannes hundepool, eric schulte nordholt, and peter paul de wolf, three specialists from statistics netherlands, were the instructors. their lectures and exercises covered the mathematical aspects of statistical disclosure control and the application of these methods in the argus software. participants came from the national bureau of economic research, the bureau of labor statistics, national center for health statistics, and research centers at the university of wisconsin, the university of michigan and the university of pennsylvania. why is statistical disclosure control important? statistical disclosure control has attracted much attention recently because complex statistical analysis can easily be done on powerful pcs. in addition, the ease of linking files from different data sources presents a real threat of re-identifying business entities or individuals from public use microdata and tabular data. to balance the stricter legal regulations and increased data needs from policy makers, several statistical agencies3 in the european union have taken on the challenges in designing better sdc methods and building a new tool to balance the need for data and the need for confidentiality protection. the computational aspects of statistical confidentiality (casc, http://neon. vb.cbs.nl/casc/) project is the result of this collaboration from 2001 to 2003. the casc project comprises not only statistical theories and methods, but also argus software development. statistical agencies and other data collectors have always removed direct identifying variables, like names, addresses, and social security numbers, before they released their data. however, such conventional practices during data processing are no longer adequate to protect respondents in the current computing world. rare combinations of indirect/non-sensitive identifiers can re-identify certain respondents in microdata. reducing these disclosure risks is very important before the release of microdata. data producers can apply various statistical methods and risk models to make their data safe using statistical packages like sas, spss and stata. however, it is very time consuming to do the global recoding and case swapping with any existing statistical packages. meanwhile, the data producers need to document any changes they have made to the original data to meet their sdc criteria. as one can imagine, it is a substantial task to produce a safe data file. to address the needs of sdc in producing public-use data and to build an efficient tool to apply sdc was the main goal of the casc project. argus is the sdc application derived from the collaboration of many casc researchers. it was first written in borland c++ and then converted to visual c++ with its user interface written in visual basic.4 iassist quarterly winter 2004 9 μ-argus: a sdc tool for creating safer microdata files the current version of µ-argus can read ascii data files in fixed format, free format with a defined separator or free format with variable names in the first line. users can provide a metadata description file or use µ-argus to specify the metadata interactively. value lists of the variables can be supplied as external files or entered as metadata attributes. µ-argus will identify the records at risk by checking frequency tables of combinations of identifying variables. low frequencies are considered a risk of re-identification. after the metadata and data file are read in to µ-argus, users can specify the set of tables manually or use one of the two basic rules used in statistics netherlands for producing microdata files for researchers and for public use files. when users are satisfied with the tables, they press the button “calculate tables” and µargus will calculate the frequency tables automatically. after the tables are calculated, users can start disclosure control in µ-argus. you will select those variables that are identified as posing dangers of re-identification and apply various methods, such as recoding variables, suppressing values, perturbing values, applying top and bottom coding, adding noise, masking, pram (postrandomisation) and micro-aggregation to bring them to a safe level. if the result is a file with too much information loss, users can easily go back to the original file and apply a different risk model to reduce the level of information loss. any changes made to the original data are documented for future reference, so users can examine the log files and see what risk models have been applied and how the data have been changed. when finally a safe file is generated, it can be output as an ascii file with an accompanying metadata file. τ-argus: an sdc tool to publish safe tabular data it is a misconception that aggregated data such as tabular data is safe. there are risks of disclosure if aggregated tables are not constructed with statistical disclosure control methods. in general the cells in a table should not be “too small” to disclose confidential information. all statistical agencies probably have their own rules for the minimum safe values for cells in their released tabular data. in addition, they need to detect when the information in aggregated tables is not just statistics but has the risk of group disclosure. τ-argus, like its twin µ-argus, has sdc methods built in to facilitate evaluation of tabular data to see if they are safe. τ-argus includes many sensitivity measures, such as a minimum number rule (threshold rule), an (n, k) dominance rule, a p% rule and a p/q rule (prior-posterior rule) to check the disclosure risks for magnitude tables. when a large number of sensitive cells are present in a table, data producers can use the table redesign feature in τ-argus to combine rows and columns to eliminate the sensitive cells. other methods such as suppression or rounding techniques can also make those cells safe. when a series of tables are created from the same microdata source, τ-argus can effectively perform sdc on all these tables, so they can all be protected in one single session instead of several sessions of sdc. τ-argus has four output options for writing out safe tables. they are csv-format, csv for pivot table, text file with code-value, or intermediate format. argus software and users’ manuals are freely available from the casc web site, http://neon.vb.cbs.nl/casc/. users can click on µ-argus and τ-argus links on the side menu to download the most current version of this tool. legal issues and practices pertaining to the netherlands eric schulte nordholt gave a report on the legal issues related to sdc in the netherlands and the european union. he covered the ethical codes in several statistical organizations, laws and statutes in the netherlands and how they affect the way statistics netherlands distributes their microdata and tabular data. even with well-implemented sdc, certain sensitive data sources are still unsafe to be released as public-use microdata or tabular data. however, researchers and policy makers need access to the data to conduct their studies. to balance the protection of confidentiality and the need for sensitive data, statistics netherlands has set up the center for the research of economic microdata (cerem, http://www.cbs.nl/en/service/research/cerem/). this on-site data center provides researchers access to enterprise data in the netherlands. researchers need to comply with a set of strict rules before they can use the restricted data in the center. the on-site room is equipped with a stand-alone pc without e-mail, internet, or any external drives. all the prospective publications will be screened for safety. in 2002 the center for policy studies was established to provide ministries with optimal statistical information. in addition to an on-site data center, researchers can submit their scripts remotely to be executed. first the job is run on test datasets and errors are corrected. the final script is then executed on real data and the results are sent to the researchers. at the start of their projects, researchers have to take an online course on sdc offered by the center for policy studies. any subsets created in the on-site data center or obtained via remote execution have to go through a safety check. it is labor intensive but necessary. to make this job easier, a new tool, ρ-argus has been developed and is being tested now. current user base of argus software since argus is a fairly new tool for statistical disclosure control (sdc), it is mainly used in the national/state statistical offices among those countries involved in the european union’s the computational aspects of statistical 10 iassist quarterly winter 2004 confidentiality (casc) project. however, anco johannes hundepool, eric schulte nordholt, and peter paul de wolf, the three specialists at statistics netherlands have conducted several argus workshops to promote this tool and its sdc methodology. their latest workshop was given on april 13 and 14, 2005 in sydney australia at the 55th session of the international statistical institute. it is likely that they will give their workshop in future iassist annual conference. statistical disclosure limitation practices in the u.s. statistical agencies unlike the netherlands, the u.s. federal statistical system is not centralized and is comprised of over 70 agencies according to an office of management and budget (omb) report5. so how do these agencies protect the confidentiality of data that they collect? what sdc methods are used by federal statistical agencies before they disseminate their public use microdata and tabular data? the interagency federal committee on statistical methodology (fcsm) was established in 1975 to recommend standards for statistical methodology to be followed by federal statistical agencies. fcsm investigates problems which affect the quality of federal statistical data, as well as makes suggestions for improving statistical methodology in federal agencies. the fcsm has about twenty members. this network of federal agency personnel has focused primarily on data quality. it has published a confidentiality and data access committee (cdac) checklist (http://www.fcsm.gov/committees/ cdac/checklist_799.doc). this list consists of a series of questions that can assist an agency’s disclosure review board to determine the suitability of releasing either public use microdata files or tables. please note that fcsm uses the term statistical disclosure limitation (sdl) or statistical disclosure restriction (sdr), not statistical disclosure control, when it discusses different statistical disclosure techniques. another important document is fcsm’s statistical policy working paper # 22 (spwp # 22): report on statistical disclosure limitation methodology (http://www.fcsm.gov/working-papers/ spwp22.html). it provides 12 recommendations to improve disclosure limitation practices. these two documents are the viable foundation for sdl practices in federal statistical agencies. similar to statistics netherlands, u.s. federal statistical agencies have developed their own procedures to provide researchers access to their sensitive data. these procedures can be classified into three categories: on-site research centers, remote access, and data use agreements or licenses.6 three examples follow. in 2003, the census bureau’s center for economic studies developed and opened several research data centers (rdcs) around the country. the rdcs provide a secure census bureau environment where researchers may have limited access to confidential economic and demographic microdata, with appropriate safeguards to protect data confidentiality. researchers need to submit their proposals to the rdcs first. after their research projects are approved, they will pay for the costs associated with the work, such as computer charges. at each rdc site, standalone workstations without removable media and network connections are set up in a secured and locked room. all the researchers’ materials will be inspected before they are removed from the rdc. disclosure reviews are performed on the researchers’ output. the national center for health statistics (nchs) has a remote access system for researchers in addition to an rdc in their headquarters in hyattsville, maryland. after their proposals are approved, researchers can submit their work electronically to staff at the nchs’ rdc. all submitted programs are reviewed for non-allowed commands, such as proc tabulate or proc iml in sas. all output goes through sdl review before they are sent back to researchers. in the third case, the national center for education statistics (nces) licenses their restricted data to researchers for them to use at their home institutions. in their formal letter of data request, researchers need to specify their research scope and the time period for the loan of the restricted files. they also need to provide a security plan compiled with nces’s requirements. each data user of the restricted files is required to sign an affidavit of nondisclosure. nces conducts unannounced, unscheduled inspections of the licensee’s site to assess compliance with the provisions of the license, security procedures, and the licensee’s submitted security plan. any violation subjects the licensee to immediate revocation of the license by nces, or a report of the violation to the u.s. attorney. the restricted data files need to be returned to nces upon the completion of the project. the confidentiality and data access committee (cdac) web site (http://www.fcsm.gov/committees/cdac/cdac.html) has a link to resources for confidentiality and data access. it lists many important papers and reports for people who are interested in the topic. epilogue the workshop announcement had proclaimed: “workshop materials and presentations will be most accessible to those with graduate training in statistical methods (e.g., econometrics, demographic methods) and researchers experienced in the quantitative analysis of panel or longitudinal survey data.” i was a bit concerned about how accessible those materials would be to a data librarian, like me, who has no formal training in either statistics or research methods. so with a curious mind, i went to the workshop and sat through all three days of lectures and iassist quarterly winter 2004 11 exercises even though i did not understand any of the intimidating mathematical formulas. i was very impressed by how intuitive the argus interface is. users with the appropriate statistics background can easily make informed choices among built-in sdc methods and create a safe file for distribution. lacking any statistical training, i am not in a position to appraise the built-in risk models and sdc methods in argus. yet, this workshop convinced me of the importance of sdc. my plan is to spread the sdc messages and share argus on my campus. i hope that you will find my report on this workshop useful. to learn more about sdc and argus, please visit the casc web site, http://neon.vb.cbs.nl/casc/. it has links to many research papers relevant to the development of argus software and the sdc theories and methods that are applied in argus. endnotes 1 chiu-chuang (lu) chou, senior special librarian, data and program library service, university of wisconsin madison, united states. email: cchou2@wisc.edu. 2 leon willenborg and ton de waal, elements of statistical disclosure control, lecture notes in statistics 105 (new york: springer-verlag, 2001). 3 casc project team includes statistics netherlands, istituto nationale di statistica (italy), university of plymouth (uk), office for national statistics (uk), university of southampton (uk), the victoria university of manchester (uk), statistisches bundesamt (germany), university la laguna (spain), institut d’estadistica de catalunya (spain), institut national de estadisica, tu ilmenau (germany), institut d’investigacio intelligencia artificial-csic (spain), universitat rovira i virgili (spain) and universitat politecnica de catalunya (spain). 4 anco hundepool, “the argus-software”(paper presentation, un-ece/eurostat worksession, luxembourg, april 7-9, 2003).. 5 u.s. office of management and budget, “statistical programs of the united states government: fiscal year 2004,” http://www.whitehouse.gov/omb/inforeg/04statprog. pdf (accessed on february 10, 2005) 6 virginia a. de wolf, “issues in accessing and sharing confidential survey and social science data,” data science journal 2, (2003). sist newsletter vol.1, no. 1 some of the respondents to the membership campaign letter recommended the creation of other action groups, such as the hardware/software technology action group (data organization and management) and a cross-file indexing action group. other suggestions included a group to interface between information data bases and potential users and a group to examine the relationship between bibliographic data bases, statistical data bases, and lassist. initial response to the lassist mailing indicates that there is considerable enthusiasm for such an organization in canada. it will be up to the canadian members of lassist to decide whether their needs can best be served by merging with the us organization in a north american lassist movement or by working closely with the americans in matters of conmon interest, but continuing to maintain a separate secretariat and a separate voice at the international level. west european secretariat report per nielsen danish data archives in europe, the lassist membership campaign in september, 1975, resulted in an immediate re'ponse from more than 80 interested individuals — and names are continuously cominj in. in the spring of 1976, a number of european social science data archives agreed to provide information dissemination facilities for lassist. consequently, all individuals having indicated an interest in lassist will receive mailings according to the following geographical division, where bracketed numbers indicate the approximate "membership" size in may, 1976: greece, italy, and spain: archivio dati e programmi per le scienze sociali (adpss), milan (7); belgium, france, and french-speaking switzerland: belgian archives for the social sciences (bass), louvain-la-neuve (8); denmark: danish data archives (dda), copenhagen (7); eire, israel, and the united kingdom: ssrc survey archive, essex (21); finland, norway, and sweden: norwegian social science data services (nsd), bergen (13); the netherlands: steinmetzarchief , amsterdam (7); austria, the federal republic of germany, and german-speaking switzerland: zentralarchiv (za), cologne (14). until recently, the za has also disseminated information to half a dozen potential members in eastern europe; these people (and, we hope a lot more to be brought in) will now be served by dr. ostrowski of the polish academy of sciences, warsaw. prior to the edinburgh meetings an overall mailing was carried out, including information on the action groups in europe. within a few months, we hope that the action groups have defined the priority of tasks so that workshops and other substantive activities accomplishing the objectives of lassist can be scheduled for 1977. instructions for authors of the iassist quarterly 1/8 monroe-white, thema (2022) emancipating data science for black and indigenous students via liberatory datasets and curricula, iassist quarterly 46(4), pp. 1-8. doi: https://doi.org/10.29173/iq1007 emancipating data science for black and indigenous students via liberatory datasets and curricula thema monroe-white1 abstract despite findings highlighting the severe underrepresentation of women and minoritized groups in data science, most scholarly research has focused on new methodologies, tools, and algorithms as opposed to who data scientists are or how they learn their craft. this paper proposes that increased representation in data science can be achieved via advancing the curation of datasets and pedagogies that empower black, indigenous, and other minoritized people of color to enter the field. this work contributes to our understanding of the obstacles facing minoritized students in the classroom and solutions to mitigate their marginalization. keywords emancipation, data harms, data science, liberatory pedagogy, curricula introduction the data science profession is currently dominated by white and asian males (harnham report, 2021; duranton, et al, 2020). in fact, just 27% of data scientists, defined as individuals who code, collaborate, and communicate by transforming data into insights using techniques in statistics, analytics, and machine learning (ho et al, 2019) in the u.s. identify as women; 6% as latinx and 3% as black, (harnham report, 2021) despite representing 51%, 19% and 12% of the population respectively (us census, 2020). this pattern is particularly concerning, as according to the u.s. department of labor, data science occupations are expected to rise + 25.9% and outpace projections of every other computer occupation (+ 12.7%) or occupations overall (+ 5.2%) (rieley, 2018). furthermore, despite the rapid growth in data talent brahm et al, 2019), the supply of data workers is not expected to keep up with demand (miller and hughes, 2017). therefore, strengthening workforce capacity in data science, defined as the ecosystem dedicated to the systematic collection, management, analysis, visualization, explanation, and preservation of structured and unstructured data (marshall and grier, 2019) will require attracting more black, indigenous, and other marginalized people of color into the data workforce. background race/ethnicity and gender workforce disparities are not new. the marginalization and exclusion of black, latinx and indigenous/native people remain a problem in the science, technology, engineering/computer science and mathematics (stem) fields. according to the 2018 survey of earned doctorates, latinx students were awarded 6.7% of stem doctoral degrees, black students earned 4.9%, and american indian/alaska native students earned just 0.2 % (nsf, 2019). in computing, women’s overall share of u.s. undergraduate computing degrees has dropped from 28.5% in 1995 to 18.1% in 2014; and the trends for women of color in computing are more alarming, as their completion rates either flat-lined at 1.75% (in the case of latinx women) or dropped from 5.10% to 2.61% in the case of black women over that same timeframe (payton and berki, 2019). as of 2016, black and latinx women combined made up just 4.4% of u.s. undergraduate computing degrees (nsf, 2017). this race/ethnicity and gender underrepresentation of minoritized groups in stem has been attributed to pervasive problems in recruitment (i.e., motivations to enter) and retention (i.e., intentions to remain) of minoritized students in higher education (fox, 2009; drury, siy and cheryan 2011; smith-doerr, alegria and sacco 2017). data science, like its core disciplinary predecessors (computer science, information systems, mathematics, and statistics), suffers from a critical lack of https://doi.org/10.29173/iq1007 2/8 monroe-white, thema (2022) emancipating data science for black and indigenous students via liberatory datasets and curricula, iassist quarterly 46(4), pp. 1-8. doi: https://doi.org/10.29173/iq1007 race/ethnicity and gender diversity. therefore, meeting the nation’s need for a data workforce that is reflective of its populace requires new ways of educating and training data professionals from minoritized backgrounds. furthermore, if these patterns persist, data harms (preliminarily defined as the adverse effects caused by uses of data that may impair, injure, or set back a person, entity, or society’s interests) (redden, brand and terzieva 2020) resulting from a homogeneous (i.e., white male) data workforce will reinforce as opposed to resolve structural racist practices that continue to marginalize black and indigenous people in the u.s. increasing the u.s. data workforce is likely to remain a top priority over the coming decades. this creates a unique opportunity for non-stem and minoritized students to find career paths in data work (especially data visualization and data journalism, which align nicely with topics of interest to humanities students (manovich, 2015). given that most minoritized students graduate in non-stem majors (see figure 1), this population represents a large untapped resource for the future data science workforce. through exposure to data professionals and use of datasets that reflect the unique historical and varied experiences of black, indigenous, and other minoritized people of color, we can facilitate inclusive growth in the u.s. data workforce. therefore, this paper aims to offer strategies for identifying more affirming and relevant datasets and exemplary data scientists that cater to the wants and needs of non-stem minoritized populations. facilitating a racial equity orientation in data science the unprecedented influence of data science tools and technologies on our social institutions (benjamin, 2019) coupled with a pervasive lack of diversity in the sciences and data science in particular, limits the range of perspectives (page, 2008) needed to address the sociotechnical complexities of “biased” (i.e., harms associated with the deployment of models that are trained on datasets that reflect broader structural inequalities) datasets and analytical processes. at present, outcomes of data science workflows lead to a profound impact on the public and underrepresented and minoritized groups in particular. biased datasets and analytical processes have furthered inequities in judicial systems, i.e., recidivism risk (o'neil, 2016; flores, bechtel and lowenkamp, 2016; dressel and farid, 2018; angwin et al, 2016) search engine outputs (noble, 2018), and facial recognition (buolamwini and gebru, 2018) among others. joy buolamwini and timnit gebru’s research exposed data harms caused by facial recognition software that disproportionately 14.4% (n=66143) 85.6% n = fig 1. image by author. sankey diagrams of 2017 bachelor’s degree completions by discipline of undergraduate minoritized students overall (left) and in stem (right) (source: u.s. department of education nces ipeds database.) https://doi.org/10.29173/iq1007 3/8 monroe-white, thema (2022) emancipating data science for black and indigenous students via liberatory datasets and curricula, iassist quarterly 46(4), pp. 1-8. doi: https://doi.org/10.29173/iq1007 misclassified darker-skinned female faces (34.7% error rate) at a rate 43 times higher than that of lighter-skinned males (.8% error rate). systems like these are used in concert with massive lawenforcement databases with images of over 117 million u.s. adults (over half the entire u.s. population) for police and government surveillance, leading to the false identification of innocent suspects (garvie, bedoya and frankle, 2019). in addition to economic and data harms justifications for increasing the diversity of the data workforce, a smaller yet growing segment of the academic literature explores the substantive contributions of minoritized groups to the u.s. workforce. for example, evidence suggests that minoritized students of color (1) care more about using their work to assist others (miller et al, 2000), (2) are more likely to endorse communal goals (seymour and hewitt, 1994), (3) emphasize collectivist values, (smith et al, 2014) and (4) address issues of social justice (allen et al, 2015). the equity ethic framework, defined as a “principled concern for social justice and for the well-being of people who are suffering from various inequities,” (mcgee and bentley, 2017) places these motivating forces at the center of the underrepresented minoritized stem student experience. specifically, an equity ethic is understood as a psychological attribute that is characterized by the degree to which an individual’s social justice concerns are developed and ultimately acted upon (naphan‐kingery et al, 2019). this line of research centers on actions stemming from a social justice orientation, which entails helping others for the purpose of reducing social inequalities. findings suggest that students from historically marginalized backgrounds in engineering and computer science are “likely to develop an equity ethic because they are likely to experience oppression and discrimination and to recognize inequity and social suffering in similarly situated groups” (ibid., p. 3). in data work, an equity ethic would express itself in the form of proactively using statistics and machine learning to address social problems (i.e., human trafficking, algorithmic bias, etc.) that disproportionately affect marginalized communities. the author refers to these data scientists with an equity ethic orientation as emancipatory (monroewhite, 2021). emancipatory data science is defined as data work that frees members of marginalized communities from being the ‘object’ to the ‘subject’ of data science framings and where decisions regarding why, how, what, when, and where data are collected, managed, analyzed, interpreted and communicated are maintained by and for members of minoritized, marginalized and vulnerable communities (monroe-white, 2021). emancipatory data science matters, because for members of marginalized, minoritized, and vulnerable communities, who provides the service (e.g., medicine, education, finance, housing, etc.) affects how the service is delivered (gershenson et al, 2018; alsan, garrick and graziani, 2019; steenbarger, 2020). the resulting impact is that having same race doctors and teachers lead to greater standards of care and improved educational outcomes for members of these groups. knowing this; however, requires a reflexive perspective that until recently was missing from the data science academic literature. as an applied social science, majoritarian data science (and data scientists) tend to neglect the study of their own social conditions and normative behaviors (merton, 1949) in favor of examining the data of others. however, as more institutions develop data science programs to meet growing student and industry demand, the opportunity to address inequities earlier as informed by empirical evidence becomes even more pressing (kelly, 2005). furthermore, unless intentionally and actively corrected, these data science programs will continue to reinforce systemic and structural biases that have disproportionately marginalized students from these fields (monroewhite, marshall and contreras-palacios, 2021; 2022). by leveraging lessons from culturally-relevant pedagogy, data curators (i.e., those responsible for creating, organizing and maintaining data sets) and data science instructors can identify and provide increased access to 1) datasets established about black and indigenous peoples and 2) tools that facilitate minoritized student adoption and interaction. ultimately, the aim of these efforts would be to make data science more inclusive and affirming for non-stem (i.e., business, humanities, education etc.) members of black and indigenous communities https://doi.org/10.29173/iq1007 4/8 monroe-white, thema (2022) emancipating data science for black and indigenous students via liberatory datasets and curricula, iassist quarterly 46(4), pp. 1-8. doi: https://doi.org/10.29173/iq1007 thereby increasing capacity within the u.s. data workforce, and minimizing data harms caused by antiblack, anti-indigenous datasets and algorithmic models utilizing what zuberi and bonilla-silva term “white logic” (2008). that is, the context in which white supremacy has defined the techniques and processes employed to analyze data as well as the reasoning used by researchers to understand society. in the world of data science, this requires a critical reassessment of the logic behind and implications of data science processes. liberatory data science pedagogical framework there is an increasing need for racially relevant and responsive teaching in university settings. the ladson-billings model of culturally relevant pedagogy has been applied to stem courses to promote a more inclusive culture for minoritized students (larnell, et al, 2016; johnson and elliott, 2020). one component of this three-part model requires cultural competence or creating a classroom culture where students feel they can be themselves. being culturally competent; however, often involves changes to the curriculum in ways that incorporate students’ cultural knowledge, as in a data analytics course where students analyze and visualize data related to racial justice. this helps “students recognize and honor their own cultural beliefs and practices while acquiring access to the wider culture” (ladson-billings, 2008) learning modules would teach students fundamental and in-demand data science software (i.e., r, python, sql, tableau etc.) and skills using culturally relevant datasets (see figures 2a & 2b) that are widely accessible by faculty and students alike. a liberatory data science curriculum would include prepared datasets and guidance on analytical processes (i.e., data acquisition, preprocessing, modeling, visualization, and interpretation) in a way that promotes student voice and sense of belonging, highlights the work of historical and contemporary data scientists from non-dominant cultures, and encourage students to contribute their own cultural knowledge through class assignments and activities. for example, the following black and indigenous resources are publicly available and readily accessible to data science instructors. 1. native land: https://native-land.ca/ 2. racial terror lynchings: https://lynchinginamerica.eji.org/explore 3. ida b. wells just data lab: https://www.thejustdatalab.com/ 4. slave voyages: https://www.slavevoyages.org/ learning modules that leverage datasets and data visualization resources like these can be made available to a broad audience via open access platforms (e.g., github) (banerjee, 2015) and serve as training tools for data science educators by leveraging best practices in culturally relevant instructional design. racial justice would also be of particular importance to members of minoritized groups and other marginalized communities, including those pertaining to racial activism (i.e., location of black lives matter protests) and racial injustice (i.e., anti-black police shootings), economic empowerment (i.e., historical significance of black wall street) and economic disparities (i.e., white flight and gentrification), and public health (e.g., covid-19 rates within urm communities) (benjamin, 2019). for example, a black undergraduate history, anthropology, theology or linguistics student, after taking an introductory data science course employing a liberatory data science framework, may as an endof-semester project choose to create an interactive map of the transatlantic slave trade to examine survival rates by ship, country of origin and destination (see figures 2a and 2b). this would be particularly meaningful as w.e.b. dubois (the first formerly trained black sociologist in the united states) created such a map by hand working with students at the atlanta university center in 1900. https://doi.org/10.29173/iq1007 https://native-land.ca/ https://lynchinginamerica.eji.org/explore https://www.thejustdatalab.com/ https://www.slavevoyages.org/ 5/8 monroe-white, thema (2022) emancipating data science for black and indigenous students via liberatory datasets and curricula, iassist quarterly 46(4), pp. 1-8. doi: https://doi.org/10.29173/iq1007 adding additional levels of interactivity would allow the user to; for example, identify particularly treacherous travel routes, and offer data-driven explanations for the variety of cultural, religious and linguistic innovations produced by descendants of survivors throughout the america’s including: 1) the gullah-geechee whose creolized language, religion and culture is practiced in the coastal areas and sea islands of the u.s can trace its origins to sierra-leone and benin; 2) haitian creole (kreyòl) which combines elements of the french language with central-african (i.e., kongolese), west-african (i.e., akan twi) religious and cultural practices, and; 3) palenque, which combines kongolese and spanish and is spoken by afro-colombian descendants of escaped slaves (e.g., maroons or cimarrónes) in the pacific and northwest regions of colombia, south america. adding racially relevant and global dimensions to data science educational curricula is both empowering and intellectually liberating for minoritized students, as it humanizes the gruesome realities while simultaneously respecting the diasporic experiences of descendants. students in the liberal arts (i.e., history, anthropology, english, etc.), are then motivated to learn data science by exploring topics of personal relevance and academic interest. by making data accessible, analyzable, and interpretable to non-stem students and leveraging datasets to teach data science (and close cousins such as data journalism, data mining, data visualization, etc.) programs can provide non-stem minoritized students with an opportunity to explore phenomena that are academically relevant and personally meaningful, thereby exposing and attracting a broader, more diverse segment of the population data careers (jackson et al, 2019). conclusion infusing black and indigenous history and liberatory pedagogy into data science education empowers educators to create more affirming and inclusive pathways into the field for minoritized scholars. this could ultimately lead to the mitigation of data harms by having a more diverse and inclusive data science workforce capable of identifying and challenging biased datasets and algorithms predeployment. as providers, curators and instructors of data best practices and systems, we are responsible for preparing members of the data science workforce to intelligently contend with the socio-technical complexities of their work, create liberatory data science pedagogy and curricula (castillo-montoya, abreu and abad, 2019; johnson and elliott, 2020) and advocate for the use of data to empower black, indigenous and marginalized people of color. fig 2a “the slave route” (source: unesco 2006) fig 2b “the georgia negro” (source: w.e.b. dubois, 1900) https://doi.org/10.29173/iq1007 6/8 monroe-white, thema (2022) emancipating data science for black and indigenous students via liberatory datasets and curricula, iassist quarterly 46(4), pp. 1-8. doi: https://doi.org/10.29173/iq1007 references allen, m. j., muragishi, g. a., smith, j. l., thoman, d. b., brown, e. r. 2015. to grab and hold: cultivating communal goals to overcome cultural and structural barriers in first-generation college student’s science interest. transl issues psychol sci, 1, pp. 331–341. alsan, m., garrick, o., and graziani, g. 2019. does diversity matter for health? experimental evidence from oakland. american economic review, 109(12), pp. 4071-4111. angwin, j., larson, j., mattu, s., and kirchner, l. 2016. machine bias: there’s software used across the country to predict future criminals and it’s biased against blacks. https://www.propublica.org/article/machine-bias-risk-assessments-in-criminal-sentencing banerjee, s. 2015. citizen data science for social good: case studies and vignettes from recent projects. doi: 10.13140/rg. 2.1. 1846.6002 benjamin, r. (2019). race after technology: abolitionist tools for the new jim code. polity. isbn: 9781509526390. brahm, c., sheth, a., sinha, v., dai, j. 2019. advanced analytics talent will double. it’s still not enough. bain and company. https://www.bain.com/insights/advanced-analytics-talent-willdouble-its-still-not-enough-snapchart/#:~:text=thanks%20to%20a%20remarkably%20rapid,occur%20primarily%20outside% 20the%20us. buolamwini, j., and gebru, t. 2018, january. gender shades: intersectional accuracy disparities in commercial gender classification. in conference on fairness, accountability and transparency. pp. 77-91. pmlr. https://proceedings.mlr.press/v81/buolamwini18a.html castillo-montoya, m., abreu, j., and abad, a. 2019. racially liberatory pedagogy: a black lives matter approach to education. international journal of qualitative studies in education, 32(9), pp. 1125-1145. https://doi.org/10.1080/09518398.2019.1645904 dressel, j., and farid, h. 2018. the accuracy, fairness, and limits of predicting recidivism. science advances, 4(1), eaao5580. doi: 10.1126/sciadv.aao5580 drury, b. j., siy, j. o., and cheryan, s. 2011. when do female role models benefit women? the importance of differentiating recruitment from retention in stem. psychological inquiry, 22(4), pp. 265-269. https://doi.org/10.1080/1047840x.2011.620935 duranton, j., erlenbach, j., brégé, c., danziger, j., gallego, a., and pauly, m., 2020. what’s keeping women out of data science? boston consulting group. https://www.bcg.com/publications/2020/what-keeps-women-out-data-science.aspx flores, a. w., bechtel, k., and lowenkamp, c. t. 2016. false positives, false negatives, and false analyses: a rejoinder to machine bias: there's software used across the country to predict future criminals. and it's biased against blacks. fed. probation, 80(38). https://www.uscourts.gov/federal-probation-journal/2016/09/false-positives-falsenegatives-and-false-analyses-rejoinder fox, m. f., sonnert, g., and nikiforova, i. 2009. successful programs for undergraduate women in science and engineering: adapting versus adopting the institutional environment. research in higher education, 50(4), pp. 333-353. https://rdcu.be/cuwrd garvie, c., bedoya, a. m., and frankle, j. 2019. the perpetual line-up. unregulated police face recognition in america. georgetown law center on privacy and technology. https://www.perpetuallineup.org/ gershenson, s., hart, c. m., hyman, j., lindsay, c., and papageorge, n. w. 2018. the long-run impacts of same-race teachers (no. w25254). national bureau of economic research. harnham report. 2021. “usa diversity in data and analytics: a review of diversity within the data and analytics industry in 2019” https://www.harnham.com/harnham-data-analyticsdiversity-report. accessed: september 4, 2022 https://doi.org/10.29173/iq1007 https://www.propublica.org/article/machine-bias-risk-assessments-in-criminal-sentencing https://www.bain.com/insights/advanced-analytics-talent-will-double-its-still-not-enough-snap-chart/#:~:text=thanks%20to%20a%20remarkably%20rapid,occur%20primarily%20outside%20the%20us https://www.bain.com/insights/advanced-analytics-talent-will-double-its-still-not-enough-snap-chart/#:~:text=thanks%20to%20a%20remarkably%20rapid,occur%20primarily%20outside%20the%20us https://www.bain.com/insights/advanced-analytics-talent-will-double-its-still-not-enough-snap-chart/#:~:text=thanks%20to%20a%20remarkably%20rapid,occur%20primarily%20outside%20the%20us https://www.bain.com/insights/advanced-analytics-talent-will-double-its-still-not-enough-snap-chart/#:~:text=thanks%20to%20a%20remarkably%20rapid,occur%20primarily%20outside%20the%20us https://proceedings.mlr.press/v81/buolamwini18a.html https://doi.org/10.1080/09518398.2019.1645904 https://doi.org/10.1126/sciadv.aao5580 https://psycnet-apa-org.ucheck.berry.edu/doi/10.1080/1047840x.2011.620935 https://www.bcg.com/publications/2020/what-keeps-women-out-data-science.aspx https://www.uscourts.gov/federal-probation-journal/2016/09/false-positives-false-negatives-and-false-analyses-rejoinder https://www.uscourts.gov/federal-probation-journal/2016/09/false-positives-false-negatives-and-false-analyses-rejoinder https://rdcu.be/cuwrd https://www.perpetuallineup.org/ https://www.harnham.com/harnham-data-analytics-diversity-report https://www.harnham.com/harnham-data-analytics-diversity-report 7/8 monroe-white, thema (2022) emancipating data science for black and indigenous students via liberatory datasets and curricula, iassist quarterly 46(4), pp. 1-8. doi: https://doi.org/10.29173/iq1007 ho, a., nguyen, a., pafford, j.l. and slater, r., 2019. a data science approach to defining a data scientist. smu data science review, 2(3), pp. 4. https://scholar.smu.edu/datasciencereview/vol2/iss3/4 jackson, l. f., kuhlman, c., jackson, f. l., and fox, k. 2019. including vulnerable populations in the assessment of data from vulnerable populations. frontiers in big data, 2(19). 10.3389/fdata.2019.00019 johnson, a., and elliott, s. 2020. culturally relevant pedagogy: a model to guide cultural transformation in stem departments. journal of microbiology and biology education, 21(1), 21.1.35. https://doi.org/10.1128/jmbe.v21i1.2097 kelly, p. j. 2005. as america becomes more diverse: the impact of state higher education inequality. national center for higher education management systems (nchems). https://eric.ed.gov/?id=ed512586 ladson-billings, g. 2008. yes, but how do we do it?: practicing culturally relevant pedagogy. city kids, city schools: more reports from the front row, pp. 162–177. larnell, g. v., bullock, e. c., & jett, c. c. (2016). rethinking teaching and learning mathematics for social justice from a critical race perspective. journal of education, 196(1), 19-29. https://doi.org/10.1177/002205741619600104 manovich, l. 2015. data science and digital art history. international journal for digital art history, (1). https://doi.org/10.11588/dah.2015.1.21631 marshall, b., and geier, s. 2019. “targeted curricular innovations in data science,” 2019 ieee frontiers in education conference (fie), 2019, pp. 1-8, doi: 10.1109/fie43999.2019.9028491 mcgee, e. o., and bentley, l. c. 2017. the equity ethic: black and latinx college students reengineering their stem careers toward justice. american journal of education, 124(1), pp. 1–36. merton, r. k. 1949. the role of applied social science in the formation of policy: a research memorandum. philosophy of science, 16(3), pp. 161-181. https://www.jstor.org/stable/185512 miller, p. h., rosser, s. v., benigno, j. p., zieseniss, m. 2000. a desire to help others: goals of high achieving female science undergraduates. women's studies quarterly, 28(1–2), pp. 128–142. https://www.jstor.org/stable/40004449 miller, s., and hughes, d. 2017. the quant crunch: how the demand for data science skills is disrupting the job market. burning glass technologies. https://www.bhef.com/publications/quant-crunch-how-demand-data-science-skillsdisrupting-job-market monarrez, t., & washington, k. (2020). racial and ethnic segregation within colleges. research report. urban institute. https://www.urban.org/sites/default/files/publication/103279/racial-and-ethnicsegregation-within-colleges.pdf monroe-white, t.; marshall, b.; and contreras-palacios, h. (august, 2022), “social exclusion in data science: a critical exploration of disparate representation in higher education” (2022). amcis 2022 proceedings. 9. https://aisel.aisnet.org/amcis2022/sig_si/sig_si/9 monroe-white, t. (2021, june). emancipatory data science: a liberatory framework for mitigating data harms and fostering social transformation. in proceedings of the 2021 on computers and people research conference (pp. 23-30). https://doi.org/10.1145/3458026.3462161 monroe-white, t., marshall, b., & contreras-palacios, h. (2021, february). waking up to marginalization: public value failures in artificial intelligence and data science in aaai 2021 workshop on diversity in artificial intelligence. proceedings of machine learning research. https://proceedings.mlr.press/v142/monroe-white21a.html https://doi.org/10.29173/iq1007 https://scholar.smu.edu/datasciencereview/vol2/iss3/4 https://doi.org/10.3389%2ffdata.2019.00019 https://doi.org/10.1128/jmbe.v21i1.2097 https://eric.ed.gov/?id=ed512586 https://doi.org/10.1177/002205741619600104 https://doi.org/10.11588/dah.2015.1.21631 https://www.jstor.org/stable/185512 https://www.jstor.org/stable/40004449 https://www.bhef.com/publications/quant-crunch-how-demand-data-science-skills-disrupting-job-market https://www.bhef.com/publications/quant-crunch-how-demand-data-science-skills-disrupting-job-market https://www.urban.org/sites/default/files/publication/103279/racial-and-ethnic-segregation-within-colleges.pdf https://www.urban.org/sites/default/files/publication/103279/racial-and-ethnic-segregation-within-colleges.pdf https://aisel.aisnet.org/amcis2022/sig_si/sig_si/9 https://doi.org/10.1145/3458026.3462161 https://proceedings.mlr.press/v142/monroe-white21a.html 8/8 monroe-white, thema (2022) emancipating data science for black and indigenous students via liberatory datasets and curricula, iassist quarterly 46(4), pp. 1-8. doi: https://doi.org/10.29173/iq1007 monroe-white, t. and marshall, b. 2019. "data science intelligence: mitigating public value failures using pair principles" proceedings of the 2019 pre-icis sigdsa symposium. 4. https://aisel.aisnet.org/sigdsa2019/4 naphan‐kingery, d. e., miles, m., brockman, a., mckane, r., botchway, p., and mcgee, e. 2019. investigation of an equity ethic in engineering and computing doctoral students. journal of engineering education, 108(3), pp. 337-354. https://doi.org/10.1002/jee.20284 national science foundation. 2017. women, minorities, and persons with disabilities in science and engineering. special report nsf 17-310. national center for science and engineering statistics. https://www.nsf.gov/statistics/2017/nsf17310/ national science foundation. 2019. doctorate recipients from u.s. universities: 2018. special report nsf 20-301. national center for science and engineering statistics. https://ncses.nsf.gov/pubs/nsf20301/ noble, s. u. 2018. algorithms of oppression: how search engines reinforce racism. nyu press. o'neil, c. 2016. weapons of math destruction: how big data increases inequality and threatens democracy. broadway books. page, s. e. 2008. the difference: how the power of diversity creates better groups, firms, schools, and societies-new edition. princeton university press. payton, f. c., and berki, e. 2019. countering the negative image of women in computing. communications of the acm, 62(5), pp. 56-63. doi: 10.1145/3319422 redden, j., brand, j., and terzieva, v. 2020. data harm record. https://datajusticelab.org/dataharm-record/ rieley, m. 2018. “big data adds up to opportunities in math careers,” beyond the numbers: employment and unemployment. 7(8) u.s. bureau of labor statistics, june 2018. https://www.bls.gov/opub/btn/volume-7/big-data-adds-up.htm seymour, e. and hewitt, n.m. (1997) talking about leaving: why undergraduates leave the sciences. westview press, boulder. smith-doerr, l., alegria, s. n., and sacco, t. 2017. how diversity matters in the us science and engineering workforce: a critical review considering integration in teams, fields, and organizational contexts. engaging science, technology, and society, 3, pp. 139-153. https://doi.org/10.17351/ests2017.142 smith, j. l., cech, e., metz, a., huntoon, m., moyer, c. 2014. giving back or giving up: native american student experiences in science and engineering. cultural diversity ethnic minor psychol, 20, pp. 413–429. https://doi.org/10.1037/a0036945 steenbarger, b. 2020. why diversity matters in the world of finance. forbes. https://www.forbes.com/sites/brettsteenbarger/2020/06/15/why-diversity-matters-in-theworld-of-finance/?sh=1ba5671c7913#215041397913 u.s. census bureau. 2020. current population survey. available at https://www.census.gov/en.html endnotes 1 thema monroe-white is an assistant professor of data analytics in the campbell school of business at berry college and can be reached by email at: tmonroewhite@berry.edu https://doi.org/10.29173/iq1007 https://aisel.aisnet.org/sigdsa2019/4 https://doi.org/10.1002/jee.20284 https://www.nsf.gov/statistics/2017/nsf17310/ https://ncses.nsf.gov/pubs/nsf20301/ https://datajusticelab.org/data-harm-record/ https://datajusticelab.org/data-harm-record/ https://www.bls.gov/opub/btn/volume-7/big-data-adds-up.htm https://doi.org/10.17351/ests2017.142 https://psycnet.apa.org/doi/10.1037/a0036945 https://www.forbes.com/sites/brettsteenbarger/2020/06/15/why-diversity-matters-in-the-world-of-finance/?sh=1ba5671c7913#215041397913 https://www.forbes.com/sites/brettsteenbarger/2020/06/15/why-diversity-matters-in-the-world-of-finance/?sh=1ba5671c7913#215041397913 https://www.census.gov/en.html mailto:tmonroewhite@berry.edu 6 iassist quarterly 2016 / vol 40 no 4 iassist quarterly iassist quarterlyiassist quarterly more data, less process? the applicability of mplp to research data by sophia lafferty hess1, thu-mai christian2 what constitutes data quality has much to do with users’ needs and preferences for discovering, accessing, interpreting, and using data. abstract in their seminal piece, “more product, less process: revamping traditional archival processing,” greene and meissner (2005) ask archivists to reconsider the amount of processing devoted to collections and instead commit to the more product, less process (mplp) ‘golden minimum.’ however, the article does not specifically consider the application of the mplp approach to digital data. data repositories often apply standardized workflows and procedures when ingesting data to ensure that the data are discoverable, accessible, and usable over the long-term; however, such pipeline processes can be time consuming and costly. in this paper, we will apply the principles and concepts outlined in mplp to the archiving of digital research data. mplp provides a useful lens to discuss questions related to data quality, usability, preservation, and access: what is the ‘golden minimum’ for archiving digital data? what unique properties of data affect the ideal level of processing? what level of processing is necessary to serve our patrons most effectively? these queries will contribute to the discussion surrounding how data repositories can develop sustainable service models that support the increasing data management needs of the research community while also ensuring data remain discoverable and useable for the long-term.. keywords data curation, data reuse, data quality, data service introduction while meissner and greene’s (2005) seminal article, more product, less process: revamping traditional archiving processing, was written with traditional archives in mind, the authors’ appeal for a critical assessment and recalibration of archival processing is no less relevant to digital data archives. soon after the article was published, the new mplp doctrine became the cause célèbre for much of the archival community, which the authors took to task for its proclivities toward all-or-nothing archival processing that had exacerbated growing backlogs of unprocessed and therefore inaccessible materials. mplp reinstated user access as the highest of archive priorities, which absolved archivists of the minutiae of item-level arrangement, description, and preservation. for data archives, however, user access is very much tied to the minutiae. the usability of data is inextricable from its specific context: the research question the data were intended to answer, the instruments used to collect the vol 40 no 4 / iassist quarterly 2016 7 iassist quarterly data, the software programs executed to manipulate and analyze the data, and the methods employed for data collection and analysis (borgman, 2015). data are complex, and the archival processing—or what we equate to data curation—required to make them available and usable for researchers has been informed by uncompromising standards of quality. achieving these quality standards requires data archives to complete a laundry list of skill-intensive, labor-intensive, and time-intensive data curation tasks including normalizing file formats, mitigating confidentiality risks, checking for and correcting data errors, generating and enhancing descriptive metadata, assembling contextual documents, recording checksums, defining undefined variable and value codes, reconciling discrepancies between datasets and codebooks, and so on (peer, green, & stephenson, 2014). not so different from the backlog situation in traditional archives, compromises in data curation processes are inevitable given the nature of tightly resourced environments in which many data archives operate. with increasing demand from funding agencies, journal publishers, and research communities for public access to quality data, data archives (and institutional repositories that are becoming ad hoc data archives) must examine the economies of curating archival data collections to the highest degree of quality at scale while keeping user needs at the forefront of curation approaches. the application of mplp to data curation raises several essential questions about what compromises are allowable, if not inevitable, that will enable data archives to remain solvent while continuing to serve the needs of the user community. in this paper, we discuss the odum institute data archive’s application of the principal concepts of mplp to data curation as part of an exercise to assess the scalability of data archive services as demand for them increases. mplp offers a methodology that enabled us to not only reconsider, but also reaffirm the parameters of the standard data curation processes we use to provide access to quality data. data quality standards one of meissner and greene’s main criticisms of archival processing that obliged them to formulate their mplp approach is “...the persistent failure of archivists to agree in any broad way on the important components of records processing and the labor inputs necessary to achieve them” (p. 209). this criticism cannot be wholly directed at data archivists, who have achieved some consensus on the important components of data curation as demonstrated in documented best practices that have received wide acceptance among data archives (icpsr, 2012; digital curation centre, n.d.). these best practices prescribe specific data curation actions that support data quality standards. what constitutes data quality has much to do with users’ needs and preferences for discovering, accessing, interpreting, and using data. based on results from a study of users’ perceptions of data quality, wang and strong (1996) identified four dimensions of data quality: 1) intrinsic, referring to the accuracy and credibility of the data; 2) contextual, or the relevancy of the data to the user’s goals; 3) representational, relating to the ability to interpret and use the data; and 4) accessibility, or the ability to obtain the data. a more recent study conducted by faniel, kriesberg, and yakel (2015) to determine the factors that elicit social science researchers’ satisfaction with data reuse found that users associate data quality with attributes that align with wang and strong’s quality dimensions. they include completeness (contextual), accessibility (accessibility), ease of operation (representational), and credibility (intrinsic) of the data. these aspects of data quality are often summed up in the notion of data being ‘independently understandable’ to their intended users (cssds, 2012; king, 1995; lee, 2010; peer, green, & stephenson, 2014). this is the quality standard to which research data are being held, particularly those subject to data management and sharing mandates that have grown in popularity among funders and journals (e.g., national endowment of the humanities, 2012; national institutes of health, 2003; national science foundation, 2010; nature, 2015; plos, 2014; science, n.d.). to some degree, this standard also acknowledges the recent scrutiny of published scientific studies, a concerning number of which were reported to have failed to meet the reproducibility benchmark of scientific integrity (chang & li, 2015; freedman et al., 2015; open science collaboration, 2015). in response, the scientific community has called for greater research transparency, which carries the presumption that data underlying published findings are not only shared, but shared in professional data archives that have the expertise and infrastructure to ensure that data are independently understandable to the research community (da-rt, 2015; center for open science, 2015). data archives have long accepted the charge from the scientific community to meet data quality standards to support research transparency. for years, data archives have instituted baseline protocols for acquiring data submissions, preparing data materials for repository ingest, and providing access to usable dataset files based on archival standards for trustworthy repositories. the reference model for an open archival system (oais) is something of a magna carta of archival standards, informing the processing approaches of many data archives (lee, 2010). oais provides a framework of high-level concepts for understanding the requirements for long-term preservation and access of materials. fundamental to oais is the concept of the ‘designated community,’ which is defined as an “identified group of potential consumers who should be able to understand a particular set of information” (cssds, 2012, p. 1-11). in accordance with oais, data archives are responsible for giving access to materials that meet the ‘independently understandable’ criterion for data quality. meeting this criterion requires that data packages held in data archives include sufficient information for users to apprehend the content, context, and structure of the data, as well as information regarding the unique identity, original source, and allowable uses of the data. what this has meant in practice for the odum institute is that, even beyond the various automatic ingest processes executed by archival system technologies, the data archivist is responsible for performing an assortment of critical data curation tasks. table 1 provides a complete illustration of the odum institute data curation pipeline. this skill-, time-, and resource-intensive data curation is similar to peer, green, and stephenson’s (2014) data quality review adopted by yale university’s institution for social and policy studies (isps) and the data curation pipeline employed at the inter-university consortium 8 iassist quarterly 2016 / vol 40 no 4 iassist quarterly for political and social research (icpsr) (vardigan, 2007). this high-level, or maximal, data curation approach involves an exhaustive list of processing actions that are possible to execute only in small-scale operations as in the case of isps, or for well-resourced operations such as icpsr. for data archives for which neither category applies, maximal data curation may not be feasible and/or sustainable even though our users require it. here we arrive at an impasse where we need to confront problems of expectation management and resource management as we weigh user requirements for data quality against data archive capabilities. more data, less process? resolving the problems of expectation management and resource management is at the core of mplp, which “...can help archivists make decisions about balancing resources so as to accomplish their larger ends and achieve economies in doing so...” (meissner & greene, 2010, p. 176). mplp petitions archivists to pursue the ‘good enough,’ or ‘golden minimum,’ in archival processing work, which gives permission to archivists to spend the minimum amount of effort necessary to serve users’ needs. anything beyond the minimum must have “clearly demonstrable business reasons” (p. 240). however, what is ‘good enough’ for traditional archives may not be ‘good enough’ for data archives. to determine what level of processing is considered ‘good enough,’ mplp directs archivists to examine three primary task areas: arrangement, description, and preservation. in traditional archives, arrangement refers to the organization of files into physical and intellectual collections in order to preserve the context of the files’ creation as well as the order of the files as they were created. description provides detailed information about the context, characteristics, and content of the materials to allow users to discover them and evaluate their relevance. preservation deals with the long-term maintenance and protection of materials (society of american archivists, n.d.). though the materials mplp refers to differ from data objects, these archival processing activities do have their equivalents in data curation. data curation pipeline • review the dataset file package to ensure all components necessary to describe and interpret the data are present (i.e., codebook, instruments, reports, etc.) • build the document set (i.e., construct codebooks, locate external documents) • review data for confiden@ality risks • review data for errors (i.e., wild or out-of-range codes, missing or inconsistent variables, undefined missing values) • perform data cleaning opera@ons to anonymize data, correct data errors and inconsistencies, and standardize missing values • assign a persistent iden@fier (i.e., doi) • apply standard vocabulary • generate standard ddi metadata to include methodological informa@on and links to associated publica@ons • add full variable and value label text to dataset • normalize files to nonproprietary, sojware-agnos@c preserva@on formats • generate deriva@ve files for widely-used sojware plakorms • review the replica@on data materials for completeness (i.e., readme file, code file, etc.) • review the code for inclusion of commands and comments required for execu@on • execute code and compare results to the tables and figures in the manuscript • link the replica@on dataset to the published ar@cle table 1. odum ins@tute data cura@on pipeline st an da rd c u ra ti o n re pl ic at io n v er fi ca ti o n 1 vol 40 no 4 / iassist quarterly 2016 9 iassist quarterly arrangement in mplp, finding the ‘good enough’ in arrangement tends towards deliberations over re-labeling and re-foldering archival materials, and whether or not doing so for individual objects is necessary to fulfill the intended purpose of arrangement. according to meissner and greene, as well as other foundational texts on archival practices, arrangement is a way to organize materials both physically and intellectually in a way that preserves their context (society of american archivist, n.d.). mplp disputes the meticulousness with which some archivists organize and apply labels to individual objects. rather than impulsively engaging in such “overzealous housekeeping, writ large” (p. 241), mplp insists that the archivists discharge themselves of such object-level physical arrangement, which contributes little to users’ understanding of the context. greene and meissner wrote: “if a user is given an understanding of the whole and the structure and identity of its meaningful parts, then the vagaries that occur within a folder will not prove daunting, and probably not even confusing” (p. 241). in applying mplp recommendations to arrangement of data collections, primary focus is on the ‘understanding of the whole,’ which, for data, is an understanding of their context. this context is contained in codebooks that define each variable and value code; documented data collection instruments such as interview or survey protocols; methodology reports containing comprehensive information on data collection, cleaning and analysis procedures; links to related research products including publications citing the data; and the programming code used to execute data analysis. arrangement of data is ensuring the presence of these materials and the sufficiency of the information contained in these materials so that the data are ‘independently understandable.’ where any of these documents do not exist, we might construct them from scratch, a cumbersome practice of stitching together information from the data producer, related publications, or any other sources that offer useful clues about the context of the data. data archivists might also insist on performing a meticulous variable-by-variable check of the dataset file to identify and correct errors and inconsistencies. mplp questions how much of this attention and diligence to contextual materials is necessary for users’ understanding of the whole. reviewing datasets and correcting coding errors is as, if not more, tedious as shuffling documents among folders. assembling and copyediting supplemental documents for a dataset is not such a far cry from the meticulous practice of re-labeling folders. instead, processing approaches for archival arrangement for data should keep focus on the goal of ‘understanding of the whole’ and determine what the fundamental requirements are for achieving that goal. what is most critical for understanding data is having the information necessary to decipher cryptic variable names and undefined value codes. data archives may need to reconsider the benefits to users of providing additional and/or enhanced supplementary materials and performing variable-by-variable checks against the amount of resources the archive has to commit to these practices. description as is the case for any type of archival material, the primary purpose of description is to assist users in discovering and accessing materials of interest. mplp suggests that archivists provide enough information to afford users ‘decent access’ without expending extra effort on composing lengthy descriptions of an individual object or its context. mplp discourages verbosity in description, which is considered gratuitous and does not necessarily lend itself to an increase in users’ understanding of the materials or their location. description as it is performed in the data curation pipeline involves applying standard vocabularies, generating metadata, enhancing variable labels, and assigning a persistent identifier for the data. the generation of standardized data documentation initiative (ddi) metadata for data discovery is extended to include both methodological and contextual details extracted from the document set. while this provides robust metadata for search and discovery, mplp asks us to consider whether it is perhaps ‘good enough’ to provide basic discovery metadata without taking the time to incorporate these methodological details such as sample size, weighting procedures, and other contextual information that is available within supplementary documents. greene and meissner make the point that as archivists it is not our job to do the research for our patrons, and efficiencies could potentially be gained by minimizing the generation of metadata. generating metadata and assigning a persistent identifier also underlie the creation of a stable data citation. data citation is an essential practice for not only ensuring data producers receive appropriate attribution but also providing persistent access to data, documentation, and code. the joint declaration of data citation principles (2014) communicates the importance of data citation as a scholarly practice and provides information on the purpose, function, and attributes of data citations. these principles highlight the role data archives play in the creation of data citations by generating metadata and assigning persistent identifiers and reaffirm this as an essential curation practice. the other key description task within our pipeline is the enhancement of variable level metadata. for social science survey data, this often takes the form of adding the complete question text to variables. this variable level description allows for much more detailed and comprehensive discovery, examination, and analysis of the data within the repository platform. however, the mplp model would suggest against this ‘item level’ description as a processing benchmark and would instead suggest archivists focus on describing the materials as a whole. understanding how researchers interact with and use repository metadata would help us understand what is ‘good enough’ for description. repositories have employed usability testing to inform interface design and the expansion of platform functionalities (gibbs et al., 2013), an extension of these types of studies could increase our understanding of what metadata fields are most useful to researchers and provide additional evidence for how best to serve users’ access needs. preservation because our focus is on the work of the archivist, a discussion of archive systems technology that are required to effectively preserve data is beyond the scope of this paper. while much of archival preservation actions take place within technological systems, there are 10 iassist quarterly 2016 / vol 40 no 4 iassist quarterly some preservation tasks that archivists perform. since digital materials are far more fragile than analog materials (rothenberg, 1999), we concede that mplp’s ‘good enough’ does not apply as readily to preservation of digital data. however, a consideration of mplp in our examination of activities in the data curation pipeline--file normalization and optimization--that support long-term preservation allows us to identify potential efficiencies. normalizing files into open or preferred file formats allows files to remain accessible and protects against obsolescence. without normalizing files, data may become unreadable and therefore unusable. a paramount requirement when serving users is ensuring digital material remain accessible into the future; therefore, normalization can be seen as an essential processing practice. in some cases, multiple different derivative copies of a dataset may also be created to allow expanded access to the data. while this increases the dissemination of the data and facilitates reuse, one file normalized into a non-proprietary file format would serve basic user requirements. although the researcher would then have to read the file into his or her preferred software and variable-level metadata stored within the software package would be lost, as long as that contextual information is available within accompanying documentation then researchers would still be able to fundamentally understand the data. the original file in the proprietary format may also be made available alongside the preservation copy. perhaps ‘good enough’ is creating a single non-proprietary version of the data file. another possible option includes shifting the burden to the data depositor and only accepting certain file formats for inclusion within the repository. for instance, guideline two of the data seal of approval (2013) states that “the data producer provides the data in formats recommended by the data repository” (p. 12). this guideline shows how a dsa-certified data repository at a minimum must provide recommendations for appropriate file formats but the onus may be placed upon the data producer to comply. another more automated solution can be seen in certain repository software platforms, such as the dataverse, that generate a derivative preservation copy for certain data file types upon ingest (crosas, 2011). an expansion of these types of system functionalities could also lessen the processing burden. good enough” for data curation a reconsideration of primary data curation activities has helped to identify those activities that are essential to ensuring access to quality data. for each of the three processing task areas, there are some activities that, if not performed, will likely make it impossible for users to discover, interpret, and use the data. we offer this ‘minimal curation’ model as a point of reference for which to engage the data archives community in a discussion on the necessity of intensive data curation processes for supporting data quality and reuse, particularly for tightly resourced environments. in the mplp-based ‘minimal curation’ model (see table 2), arrangement is reduced to the single task of data file package review, description requires only metadata generation and persistent identification, and preservation is limited to file normalization. arrangement most important to arrangement is ensuring that necessary documentation is included within the data package so that users can understand the context of the data. whether this documentation takes on the form of a codebook or survey instrument, or some other format, at a minimum documentation should define variables and values and provide some indication of the research methodology and process, for which links to external publications may be sufficient. no longer part of arrangement in the minimal data curation scheme is the variable-by-variable review of the data to identify and remedy errors, discrepancies, and/or sensitive information in the data. by scaling back on the comprehensive data review to this degree, the archive may no longer be able to guarantee the quality of the minutiae of every dataset. in some cases, variables and values may be left undefined, missing values inconsistently or incorrectly coded, and sensitive variables in the dataset might be awaiting unauthorized disclosure. certainly, these compromises have the potential to impact overall usability; however, the goal of ‘minimal curation’ is to ensure that enough information is present for users to understand and interpret the data as a whole. the users still have contextual information available to them so as to assess the overall credibility of the data and to determine whether or not the data are relevant to them. the presence of variable and value definitions in codebooks enables users to make necessary corrections in the data. it is also “minimal” data curation pipeline arrangement description preservation • review the dataset file package to ensure all components necessary to describe and interpret the data are present (i.e., codebook, instruments, reports, etc.) • assign a persistent iden>fier (i.e., doi) • generate basic descrip>ve standard ddi metadata for discovery and access • normalize files to nonproprietary, sogwareagnos>c preserva>on formats table 2. proposed minimal data cura>on pipeline 1 vol 40 no 4 / iassist quarterly 2016 11 iassist quarterly not unreasonable to set policies that make data producers responsible for removing sensitive information in their data files and users responsible for reporting the presence of sensitive data. should the data archiving and research community determine that comprehensive variable level review is the ‘golden minimum’ for data, then we must also provide “clearly demonstrable business reasons” that dictate this additional task and take into account the additional resources that will be required as a result. description the minimal curation model reduces the amount of metadata generated for a given dataset. instead of generating extensive ddi metadata and enhancing the variable labels for question-text search queries, the archive would simply generate enough descriptive metadata to allow users to discover and access the data and understand the general scope and topic of the data. this descriptive metadata would also include a persistent identifier and all the information required for a standard data citation. while this strategy may compromise some of the discovery and online analysis potential for a dataset, it would still serve basic discovery and accessibility requirements. description to support understanding of the content and context of the data would be left to information contained in supplementary documentation. preservation in regard to preservation, the archive would continue to normalize files into a non-proprietary, system-agnostic file format, as we believe this is necessary to ensure that users are able to properly render the data into the future even as hardware and software systems become obsolete. rather than producing several different file derivatives, normalization is limited to a single file format that users may convert for use in various software platforms. although we did not discuss other preservation activities such as generating and recording checksums, performing fixity checks, and migrating digital content, these are archival processes that are essential for long-term access and reusability and therefore cannot be compromised. however, these preservation activities are performed by archival platform systems and have little effect on archivist-led data curation processes. what we have identified as a minimal data curation pipeline is neither an endorsement of a new standard of data curation, nor of mplp itself. ‘minimal curation’ is not suggested as an alternative to maximal curation. maximal curation supports the sharing of the highest quality data that gives greater assurances that data will be reusable into the foreseeable future. ‘minimal curation’ is presented to address conflicting priorities in a search for efficiency gains. discussion the outcome of the mplp exercise of reconsidering data curation processes is a recognition and greater appreciation of our commitment to providing access to quality data for our user community. in our examination of each of our current data curation activities, we were able to reaffirm the value of our practices to our users and their specific needs and expectations for quality data. as mplp predicted, this exercise reminded us that “choices can be uncomfortable” (p. 233) when attempting to find efficiencies in our current practices, all of which we deem indispensable to our users. but we have little choice but to do so as we anticipate an increase in demand for data curation services. mplp forced us to think about how each task in our data curation pipeline contributes to our goals. in doing so, we also reaffirmed the necessity and non-negotiable nature of some tasks that must be performed regardless of their intensity. the search for ‘good enough’ for data has again left us in a quandary since in many ways meeting the requirements for reuse requires labor-intensive data quality review processes. several data repositories have implemented a variety of resource management strategies for addressing the challenge of providing high quality data access with limited support. for example, the uk data archive has developed different levels of data curation to most effectively respond to users’ varying needs (uk data archive, 2013). a key aspect of this is clear communication of the data curation tasks that will (and will not) be performed as part of a program of expectation management that distinguishes roles, rights, and responsibilities of the data producer, data user, and data archive. shifting responsibility for certain data curation tasks from the data archive to the data producer and data user assumes that data producers and users have an understanding of data quality requirements and the tasks required to meet those requirements, which, unfortunately, is not always the case. to address these challenges, information professionals have produced a proliferation of educational materials and programs to teach researchers strategies for effectively managing their data with eventual data archiving and sharing in mind. while online education programs (such as mantra and the research data management and sharing mooc) have the potential to reach researchers worldwide, the impact of such education programs is not immediate and does not necessarily guarantee that data meet the standard of being ‘independently understandable.’ tools that facilitate and provide additional functionalities to streamline data curation processes also present opportunities for efficiency gains. likewise, tools, such as the open science framework, that help moderate and structure research workflows with an end goal of archiving and sharing data have the potential to assist researchers in creating data packages that meet data reuse requirements. however, even with these tools, certain tasks will continue to require manual data curation processes. ultimately, determining which data curation processes are essential for archiving and sharing of data that meet certain quality standards requires further research. this research will provide the empirical evidence and rationale for data archives’ roles in curating data for 12 iassist quarterly 2016 / vol 40 no 4 iassist quarterly reuse in accordance with the needs of our designated community. although previous studies have already clearly demonstrated the importance of contextual information (faniel & jacobsen 2010; faniel et al., 2012), additional research is needed to investigate: 1) the designated community’s expectations of the archive’s role in providing quality data; 2) how variable level reviews affect reuse; 3) how the presentation of contextual information affects use; 4) how users interact with contextual and variable level metadata; and 5) how specific data curation tasks performed by archives directly impact the data quality and satisfaction criteria discussed within the literature. by expanding our knowledge on the connections between user needs and data curation processes, we will be better equipped to determine what is ‘good enough’ for data. likewise, we will be able to substantiate the necessity of data curation, whether it be maximal or minimal, for informed reuse and make clear that simply making data available does not automatically equate to data that are useable. we will then be able to expand and build upon initiatives advocating for the development of sustainable funding models for data archives (icpsr, 2013). conclusion performing the conceptual exercise of applying mplp in many ways raised more questions than it answered. mplp reaffirmed our belief that a certain amount of processing is necessary to adequately meet users’ needs. mplp also brought to light some gaps in our knowledge about data use that prevents us from truly determining the minimum amount of processing needed. future research will help us build better understanding of the connection between user needs and data curation processes. in many ways, the exercise suggests that ‘good enough’ for data still sets the bar pretty high, and building sustainable models to fund data curation will require the data archiving community to articulate the amount of skills, time, and labor that are non-negotiable when a high level of data quality is expected. essentially, this exercise boils down to a quotation from clifford lynch: “it is clear that an enormous imbalance exists between the resources currently available to fund these efforts and the potentially almost infinite demands of a fully realized data stewardship program; a key strategy in managing this imbalance is the effective use of the specific policy goals, such as data reuse, as shaping and prioritizing mechanisms in shaping an overall stewardship effort” (lynch, 2013, p. 408). with the growth of data sharing mandates and the increasing focus on research transparency, data archives will play an essential role. however, questions still remain as to how we can best support these needs in a sustainable way that results in data that meet the requirements for reuse. references borgman, c.l. (2015) big data, little data, no data: scholarship in the networked world. cambridge, massachusetts: the mit press. chang, a.c., li, p. (2015) is economics research replicable? sixty published papers from thirteen journals say “usually not.” finance and economics discussion series 2015-083. washington dc: board of governors of the federal reserve system. doi:10.17016/feds.2015.083 center for open science. (2015) transparency and openness promotion (top) guidelines. available from: https://cos.io/top/#signatories consultative committee for space data systems. (2012) reference model for an open archival information system (oais) (magenta book no. 650.0-m-2). washington, dc: national aeronautics space agency. crosas, m. (2011) the dataverse network®: an open-source application for sharing, discovering and preserving data. d-lib magazine 17 (1-2). doi:10.1045/january2011-crosas da-rt. (2015) the journal editors’ transparency statement (jets). available from: http://www.dartstatement.org/#!blank/c22sl data citation synthesis group. (2014) joint declaration of data citation principles. martone m. (ed.) force 11, san diego ca. available from: / datacitation data seal of approval. (2013) data seal of approval guidelines (v.2). available from: http://datasealofapproval.org/media/filer_public/2013/09/27/ guidelines_2014-2015.pdf digital curation centre (dcc). (n.d) curation reference manual. available from: http://www.dcc.ac.uk/resources/curation-reference-manual faniel, i.m., jacobsen, t.e. (2010) reusing scientific data: how earthquake engineering researchers assess the reusability of colleagues’ data. computer supported cooperative work (cscw), 19 (3-4), p. 355–375. doi:10.1007/s10606-010-9117-8 faniel, i.m., kriesberg, a., yakel, e. (2012) data reuse and sensemaking among novice social scientists. proceedings of the american society for information science and technology, 49 (1), p. 1–10. doi:10.1002/meet.14504901068 faniel, i.m., kriesberg, a., yakel, e. (2015) social scientists’ satisfaction with data reuse. journal of the association for information science and technology. doi:10.1002/asi.23480 freedman, l.p., cockburn, i.m., simcoe, t.s. (2015) the economics of reproducibility in preclinical research. plos biol 13 (6): e1002165. doi:10.1371/ journal.pbio.1002165 gibbs, e., lin, l., quigley, e. (2013) dataverse usability evaluation: final report. available from: http://dataverse.org/files/dataverseorg/files/ dataverse_usability_report-participant_omitted.pdf?m=1458571553 greene, m.a., meissner, d. (2005) more product, less process: revamping traditional archival processing. the american archivist, 68 (2), p. 208-263. inter-university consortium for political and social research (icpsr). (2012). guide to social science data preparation and archiving (5th ed.). ann arbor, mi: icpsr. available from: http://www.icpsr.umich.edu/files/deposit/dataprep.pdf inter-university consortium for political and social research (icpsr). (2013, june 24-25) sustaining domain repositories for digital data: a call for change from an interdisciplinary working group of domain repositories. available from: http://www.icpsr.umich.edu/files/icpsr/pdf/ domainrepositoriescta16sep2013.pdf king, g. (1995) replication, replication. ps: political science & politics, 28 (3), p. 444–452. doi:10.2307/420301 lee, c.a. (2010) open archival information system (oais) reference model. in encyclopedia of library and information sciences. taylor & francis, p. 4020–4030. vol 40 no 4 / iassist quarterly 2016 13 iassist quarterly lynch, c. (2014) the next generation of challenges in the curation of scholarly data. in j. m. ray (ed.), research data management: practical strategies for information professionals. west lafayette, indiana: purdue university press. available from: http://www.cni.org/wp-content/ uploads/2013/10/research-data-mgt-ch19-lynch-oct-29-2013.pdf meissner, d., greene, m.a. (2010) more application while less appreciation: the adopters and antagonists of mplp. journal of archival organization, 8 (3-4), p. 174–226. doi:10.1080/15332748.2010.554069 national endowment for the humanities (neh). (2012) data management plans for neh office of digital humanitites proposals and awards. washington dc: national endowment for the humanities. available from: http://www.neh.gov/files/grants/data_management_plans_2015.pdf national institutes of health (nih). (2003) final nih statement on sharing research data (no. not-od-03-032). bethesda, md: national institutes of health. national science foundation (nsf). (2010) dissemination and sharing of research results. arlington, va: national science foundation. available from: https://www.nsf.gov/bfa/dias/policy/dmpfaqs.jsp#1 nature. (2013) availability of data, material and methods policy. available from: http://www.nature.com/authors/policies/availability.html open science collaboration. (2015) estimating the reproducibility of psychological science. science, 349 (6251), aac4716. doi:10.1126/science. aac4716 peer, l., green, a., stephenson, e. (2014) committing to data quality review. international journal of digital curation, 9 (1), p. 263–291. doi:10.2218/ ijdc.v9i1.317 plos. (2014) data availability policy. available from: http://journals.plos.org/plosone/s/data-availability rothenberg, j. (1999) ensuring the longevity of digital information. washington, dc: council on library and information resources. science. (n.d.) editorial policies: data deposition. available from: http://www.sciencemag.org/authors/science-editorial-policies society of american archivists. (n.d.) glossary of archival and records terminology: preservation. available from: http://www2.archivists.org/ glossary/terms/p/preservation#.vxknkfkrj9m society of american archivists. (n.d.) glossary of archival and records terminology: arrangement. available from: http://www2.archivists.org/ glossary/terms/a/arrangement#.vxkn2_krj9m uk data archive. (2015) data ingest processing standards. available from: http://www.data-archive.ac.uk/media/54782/cd079-dataingestprocessin gstandards_08_00w.pdf wang, r.y., strong, d.m. (1996) beyond accuracy: what data quality means to data consumers. journal of management information systems, 12, p. 5–33. doi:10.1080/07421222.1996.11518099 notes 1. sophia lafferty-hess is the research data manager at the odum institute for research in social science (228 davis library, cb# 3355, university of north carolina at chapel hill), slaffer@email.unc.edu 2. thu-mai christian is the assistant director of archives at the odum institute for research in social science (228 davis library, cb# 3355, university of north carolina at chapel hill), thumai@email.unc.edu vol30-4.indd iassist quarterly winter 2006 by by rachael e. barlow* mashing maps introduction this is the story of a class at a small, liberal-arts college. the class attempted to do something meaningful with data for people living in the community nearby. the college is trinity college. the nearby community is hartford, connecticut. the students in the class created five “google maps mashups.” one group of students mapped food resources in hartford, everything from community gardens to grocery stores to food pantries (see figure 1). another group mapped houses in disrepair in hartford’s south end, along with the contact information of the absentee owners whose negligence had caused such deterioration (see figure 2). the question: why should those in the data world care about this rather small project at a rather small school? as a sociologist, i am inspired to bring in a little sociological perspective. susan leigh star and james r. griesemar talk about boundary objects, “objects which are both plastic enough to adapt to local needs and the constraints of the several parties employing them, yet robust enough to maintain a common identity across sites.”1 sociologists who study science often talk in terms of upstream and downstream processes. upstream processes refer to what happens before the point at which a technological innovation is considered “done” (ready for the marketplace, for consumption, etc.). downstream processes refer to what happens after this pivotal, and as sociologists will note, socially-constructed point.2 i argue that mashups, although not objects in the ordinary sense (since they are digital, not material), are boundary objects. in the case of the trinity-college class, mashups facilitated the cooperation of those inside the college’s walls (faculty and students) with those outside (community members). these insiders and outsiders of academe approached such mashups while entertaining very different ideas about mashups what were good for. yet despite these two groups’ different perspectives and intentions, the mashups, “plastic enough to adapt to local needs,” had the potential to appease both groups. upstream processes: making the maps but to buy such an argument about the mashups’ status as boundary objects, the reader needs to learn a little about the upstream processes of the mashups created by the students in this course. long before the creation of these mashups, a trinitycollege professor and i composed a set of goals for the course we were going to co-teach. we wanted the students to learn something about cities generally and about hartford specifically. we wanted students to develop technical skills for managing data, but also communication, networking, and problemsolving skills. and we wanted students to participate in the construction of the knowledge they gained in the course by working with and doing something for the local community. perhaps now is the time for a quick definition. in the web 2.0 world, a mashup refers to the product one creates when mixing together the dynamic elements of preexisting websites. as rich gibson and schuyler erle put it, “in music, when you create a new song by taking the melody from one song and the lyrics from another, it is called a mashup. a lot of times things go poorly, but now and then the results are stunning.3 the same, they explain, is true for web mashups. remixing websites might produce something silly or extraneous, or something significant and revealing. hence, we titled our course invisible cities for a reason: so that students could experiment with rearranging data in ways that would allow them to reveal something about the city of hartford that was otherwise invisible. mashups seemed like a timely, if not faddish, means to this end. the whole idea of a mashup is to take what already exists, stir it around and create something not yet seen.4 months prior to the beginning of the semester, the professor and i asked the leaders of two community organizations whether they would like to work with our class. the leaders agreed, but not for the reasons that mattered to us. after all, why would these organizations care about what privileged trinity students learned? instead, the leaders of these groups were primarily concerned with furthering their own organizational goals: informing and mobilizing residents in the south end of hartford. 6 iassist quarterly winter 2006 the course began and the students learned that they were expected to work with and for community groups. they also learned that their work would take the form of maps: collecting data for them, arranging that data in spreadsheets, and “mashing” the data with google maps using the plethora of online tools that make such mashing easy. during the early part of the course, we trained students on the array of skills they would need to make all of this possible: they needed to know what a google mashup was and how to make one, what a spreadsheet of data looked like and how to manage it, and what concerns hartford residents and how to work with them. a month into the semester, we arranged for the students to meet with the community-organizations’ leaders to exchange ideas about what kind of data would make sense for a map mashup. this meeting allowed us to hear what issues had relevance for these organizations. the list was long. they imagined maps that plotted, among other things: banks with free checking (for residents who otherwise cannot afford a checking account), known places where buses idled (emitting pollution and sickening kids), voting stations throughout the city (that were otherwise not well advertised), the location of advocacy groups (like themselves), and places for city residents to access the internet (to access these maps and the larger web). after the meeting, when the class reconvened, the students faced the task of deciding what maps they could make, wanted to make, and mattered most. two of the maps that “won” in this contest are the two i mentioned at the beginning of this article. i mentioned these two because they represent different ways students went about mashing. the students who created the food resources map inherited an excel spreadsheet from a thirdparty organization in hartford that had already collected its own data on locations to access food in the city. the students did not have to collect data anew; instead, their work involved rearranging the excel file so that it was readable by the online tool (called “zeemaps”) they had chosen to make their mashup. but the students who created the mashup of abandoned properties did have to collect data. one of the organizations with which we collaborated had given this latter group of students an initial list of properties it had identified as problematic. but the students had to locate a large amount of additional data on their own: data from hartford’s assessor that confirmed whether the houses had recently changed ownership, data from the connecticut secretary of state about whose names were behind many of the limited liability corporations listed in place of owners for some of these properties, and data from a recent “city scan” project that described the character of these properties’ blight (broken windows, lawns in need of mowing). finally, the students turned to the phonebook to get owners’ contact information, since the organization that requested the map wanted community members to call these owners and ask them to clean up their properties. downstream processes: after the maps but let us put aside for a moment the upstream processes that led to the creation of these mashups and examine their downstream processes, what occurred after the mashups were online, accessible to the public. the mashups, once created, were both opportunity-makers and pressurecookers. the attention they received from the local media and the college administration allowed those of us who taught the course and the students enrolled in it to receive more accolades than any of us perhaps deserved. the mashups made the college “look good” in front of a statewide audience that often has perceived trinity as disengaged from its urban environs. but the attention simultaneously put pressure on the professor and me to make more maps. in the wake of the mashups’ online publication, other local organizations were soon knocking on our doors, as were other trinity faculty asking us if we could help them enter the map-mashup world. in other words, the mashups never were finished, even when we pretended that they were by putting them online for official consumption. the attention also put pressure on the students and community groups to keep the original maps updated. as the world the data described changed, the maps needed to change, too. furthermore, almost immediately, a few individuals wrote emails complaining that the mashups contained misinformation. responding to these issues was a challenge. although mashup technologies make updating easy, fact-checking takes time and people, both of which—at least inside the ivory tower—tend to disappear quite quickly once a semester ends and winter/summer break begins. finally, i personally questioned whether all five maps matched equally well the desires of the organizations with which we worked. the fact that the organizations perceived some maps as more useful than others was reflected in which maps required further refinement after the semester’s conclusion. while we are still working with one organization on one map a year and a half after we started, some of the maps we never touched again. i suspect this was not because these latter maps were perfect: they were not. instead, because the class did not perfectly read the organizations’ and community’s needs, these latter maps found no constituency to care about them, discover their faults, and ask for a better product concluding remarks time to take stock. i propose that there are a few things we can learn from our foray into map mashing. mashups allowed our students to stand in a somewhat strange, but useful position relative to the local organizations they were trying to “serve” (a term i use with some trepidation). mashups were easy: for the students to make and for the organizations to imagine and to use. in mashing, students were not so much producers as they were interpreters, iassist quarterly winter 2006 7 since they were not creating a product from scratch as much as they were using online, freely-available tools to provide organizations with a new perspective on what these organizations, in another form, already knew. for a small, liberal-arts college that lacks the resources larger schools often have, mashing allowed the students to do something with the community they probably could not have done any other way in the tight timeframe of an academic semester. their newfound role as data massagers was a suitable one, given the institutional constraints. as edward maloney has noted, “what makes mash-ups interesting from a teaching and learning perspective is that they permit people with very little technical know-how to manage knowledge online, modeling solutions for others to see, collaborate on, and use in new ways.5 furthermore, as boundary objects, mashups succeeded not only in mixing up online content and tools, but also people, in this case: the students, local organizations, faculty, administration, and media that participated in the project’s upstream and downstream processes. in this way, mashups were a means to “open data,” the idea behind the most recent iassist conference, where i first presented this paper. now, as gibson and earle point out, mashing does not always have stunning results. as my reportage of the downstream processes makes clear, the mashing of people was not always as one might have hoped. however, this is all part of the mashing gamble: that the benefits of open data will outweigh the risks. acknowledgements special thanks to dan lloyd, my professor-collaborator on this project, david tatem, for helping with some of the technical aspects, and vincent boisselle, for giving me permission to get involved. references becker, howard, robert r. faulkner, and barbara kirshenblatt-gimblett, eds. art from start to finish: jazz, painting, writing, and other improvisations. chicago, il: university of chicago press, 2006. gibson, rich, and schuyler erle. google maps hacks: tips & tools for geographic searching and remixing. cambridge, ma: o’reilly, 2006. maloney, edward j. “what web 2.0 can teach us about learning.” the chronicle of higher education 53, no. 18 (2007): b26. oudshoorn, nelly, and trevor pinch, eds. how users matter: the construction of users and technologies. cambridge, ma: mit press, 2003. star, susan leigh, and james r. griesemar. “institutional ecology, ‘translations’ and boundary objects: amateurs and professionals in berkeley’s museum of vertebrate zoology, 1907 39.” social studies of science 19, no. 3 (1989): 387 420. endnotes 1. susan leigh star and james r. griesemar, “institutional ecology, ‘translations’ and boundary objects: amateurs and professionals in berkeley’s museum of vertebrate zoology, 1907 39,” social studies of science 19, no. 3 (1989): 393. 2. howard becker, robert r. faulkner, and barbara kirshenblatt-gimblett, eds., art from start to finish: jazz, painting, writing, and other improvisations (chicago, il: university of chicago press, 2006), nelly oudshoorn and trevor pinch, eds., how users matter: the construction of users and technologies (cambridge, ma: mit press, 2003). 3. rich gibson and schuyler erle, google maps hacks: tips & tools for geographic searching and remixing (cambridge, ma: o’reilly, 2006), 67. 4. note here there are many different kinds of mashups, not just ones involving google maps. mashups might not even involve maps at all! 5. edward j. maloney, “what web 2.0 can teach us about learning,” the chronicle of higher education 53, no. 18 (2007): b26. * rachel e. barlow presented this at the iassist 2007 conference in montreal as “maps that mash: daring, dangerous, or dumb?”. contact information: rachel e. barlow, trinity college library, raether library and information technology center, trinity college, 300 summit street, hartford, ct 06106. rachael. barlow@trincoll.edu, +1 (860) 297 – 4114 8 iassist quarterly winter 2006 iassist quarterly winter 2006 9 1/2 schwartz, ofira & hayslett, michele (2025) perspectives on data: management, access, and education across institutions, iassist quarterly 49(4), pp. 1-2. doi: https://doi.org/10.29173/iq1187 the creative commons-attribution-noncommercial license 4.0 international applies to all works published by iassist quarterly. authors will retain copyright of the work and full publishing rights. editors’ note: perspectives on data: management, access, and education across institutions dear iassisters, welcome to iassist quarterly, vol. 49 no. 4. iq’s editors and editorial board are continuously working on developing policies to provide a clear, consistent, and equitable experience for authors, reviewers, and readers. authors who are considering submitting a manuscript are encouraged to download and read over the author template if they haven’t published with us recently. one recent update is a request that as part of the review process, authors include a separate letter, apart from their manuscript, explaining how the reviewers’ comments were addressed. this letter offers authors an opportunity to clarify why certain comments were not fully implemented or were addressed differently than recommended. as many of you may know, the qualitative social science & humanities data interest group (qsshdig) has been working on a special issue highlighting the challenges of data sharing in the context of qualitative research. we are looking forward to publishing this issue in the spring of 2026. since this will be a double issue, the release date may vary from our regular publishing schedule; however, we’ll announce it debut on the list as usual. the current issue, iq 49(4), brings together diverse perspectives on the evolving landscape of research data management, access, and data literacy within academic and research contexts. the four featured articles explore critical challenges and opportunities in supporting data-driven scholarship. collectively, these contributions underscore the importance of cross-campus collaboration and shared understanding to advance data stewardship and literacy. opening the issue is the winning submission of the 2025 iassist conference paper competition, titled “assessing data management and sharing plans: the “state of play” at duke and opportunities for crosscampus collaborations.” authors sophia lafferty-hess, william krenzer, jenny ariansen and jennifer darragh, present key findings from a data management and sharing plans (dmsp) assessment project jointly undertaken by two research support groups at their institution. they further discuss how data management specialists can use this cross-campus collaboration model for ongoing education, training, and resource development. in their article “in the data steward’s shoes: an autoethnographic exploration of everyday challenges,” authors auriane marmier, stefan stepanovic, and tobias mettler use an autoethnographic approach to provide a practice-based perspective on the role of data stewards in academic institutions and the challenges they face. the authors seek to initiate a discussion about the current positioning of data https://creativecommons.org/licenses/by-nc/4.0/ 2/2 schwartz, ofira & hayslett, michele (2025) perspectives on data: management, access, and education across institutions, iassist quarterly 49(4), pp. 1-2. doi: https://doi.org/10.29173/iq1187 stewards within academic institutions in order to understand how data stewardship can better support both regulatory compliance and research innovation. graeme campbell, katie cuyler and alex guindon offer us an overview of the current landscape of the canadian census portals. in their article “assessing the landscape for discovery and access to historical canadian census data,” the authors argue that fragmented and inconsistent access to canadian census data serve as a barrier to research and emphasize the need for a single comprehensive access point for canadian census data. the paper “conception of data literacy in statistics education literature” presents results from a scoping review of data literacy articles within the field of statistics education. the concept of data literacy is interpreted quite differently by librarians and statisticians. authors julia bauder and libby cave review the landscape of data literacy education in statistics, offering librarians and other information professionals a map for coordinating their data literacy work with disciplinary faculty, and contributing to data literacy education. wishing you a joyous holiday season and a prosperous new year! ofira schwartz and michele hayslett, december 2025 editors’ note: perspectives on data: management, access, and education across institutions opening the issue is the winning submission of the 2025 iassist conference paper competition, titled “assessing data management and sharing plans: the “state of play” at duke and opportunities for cross-campus collaborations.” authors sophia lafferty-... 22 iassist quarterly 2014 iassist quarterlyiassist quarterly abstract this paper presents findings on existing potentials for the establishment of social sciences digital data archives in bosnia and herzegovina, croatia, and serbia. findings are based on a standardized survey that was conducted in all three countries on representative sample of social sciences researchers, with a 63% average rate of completed questionnaires. results of the survey show that the potential for establishment of digital data archives in all considered countries are large, regarding the scope of data produced, as well as positive attitude of researchers toward data sharing and benefits of data archives. also, results point to lack of knowledge in dealing with metadata as the main obstacle for data sharing, which further implies that existence of national data archives, with staff trained to underpin researchers’ efforts in process of data documentation and preservation, should be very beneficial for the future development of social sciences in these countries. keywords: data archive, data services, social science, bosnia and herzegovina, croatia, serbia introduction today, knowledge is one of the most important sources of global economy growth (wb, 2012) and a key driver of a company’s value (bock, zmud, kim & lee, 2005). transmission of information, which is widely expanded as a result of internet development and the possibility of collection, storage and dissemination of that information, has given the new possibilities for the use of knowledge. this is particularly important having in mind that knowledge has the characteristic of growing in the process of dissemination (arzberger et.al., 2011). research has indicated (ukda, 2002; corti, et al., 2011) that there has been a sharp increase in collecting data that has been used in studies of economic, political and other social issues, over the last decades. regarding the fact that the process of collecting primary data is the most expensive and time consuming phase in the research process; establishment of national digital data archives for research data in social sciences and their integration into the standardized system for data sharing on the international level is considered as a cost savings solution (bradić-martinović, zdravković, 2012). data collection was always a part of scientific research process. natural sciences use data from experiments, while social sciences use various methods of primary data collection, such as questionnaires, interviews, focus groups etc. the collected data are often necessary to keep in order to check the results or to be used for further research. the phenomena of digital data have changed the way that data are collected and preserved. for that purpose, many countries established data centers or data archives. some data archives have a very long history, like the ucla social science data archive, us (formed in 1961), and the united kingdom data archive, ukda, uk (formed in 1967). in the western balkans region, none of the countries have digital archives in social science. thanks to the successful implementation of the fp7 serscida project (support for establishment of national/ regional social sciences data archives) bosnia and herzegovina, croatia, and serbia have an opportunity to establish data services in the social sciences. it is a strategic project, designed to support cooperation and knowledge exchange between those eu countries that researchers’ interest in data service in bosnia and herzegovina, croatia, and serbia by aleksandra bradić-martinović1 and aleksandar zdravković2 iassist quarterly 2014 23 iassist quarterly are members of the council of european social sciences data archives (cessda) and the western balkan countries in the field of social science data archiving. the project addresses the existing potential for use of information-communication technologies for the benefit of scientific research and exchange of knowledge, as laid down in the call for proposals topic. it aims to produce tangible results and improve the capacities for exchange of knowledge and data collected through research in social sciences between the european countries and western balkan countries involved.3 the results presented in this paper are outcomes of work package 2 of serscida project analysis of existing potentials for the establishment of social sciences digital data archive, presented at the iassist 2013 conference in cologne, germany. methodology and sample the first step of analysis was to develop appropriate methodology by local partners with assistance of cessda partners. to fulfill that aim we designed an online questionnaire for researchers which considered both their experience of documentation, re-use, and disseminating of research data; and also which type of statistical/ analytical software packages, methodology and data they used primarily in their research. accession to the questionnaire did not imply any restriction (academic network users, for instance) and no registration were needed. the survey has five parts which covered: 1) characteristics of respondents; 2) producing data; 3) methods of data gathering; 4) archiving practices and preferences; and 5) use of data and secondary analysis. it was conducted during june and july, 2012. bosnia and herzegovina had 139 completed questionnaires out of 225; croatia had 186 completed questionnaires out of 307 and serbia had 322 completed questionnaires out of 493. the database of potential respondents was made on the basis of extensive gathering of researchers’ contact addresses either from relevant government institutions or individual websites of research institutions involved in social sciences. the average rate of completed responses for all three countries was 63%. hereafter we selected the key questions that will shed light of the situation on this topic in bosnia and herzegovina, croatia, and serbia. results of the survey characteristics of respondents the first few questions in our survey were aimed to make us more familiar with the basic characteristics of our respondents regarding their principal activity and research discipline. the results on principal activities are slightly different between countries. in bosnia and herzegovina (bih) over a half of respondents were undergraduate students, doctoral students, or teaching assistants and researchers or professors and in croatia and serbia about 80% of respondents were doctorial students or teaching assistants and researchers or professors. there are also some differences between countries regarding research discipline. according to the respondents’ answers within the context of research discipline in bih major researchers were in law science (21%), sociology (13%) and economic (12%), while other discipline like psychology, education science and teacher training, political science and journalism are below 10%. in croatia most of the researchers are economics 12% sociology 13% psychology 8% education   science  and   teacher   training   8% political  science 7% journalism 7% law 21% other 24% bosnia  and  herzegovina economics 16% sociology 19% psychology 28% education   science  and   teacher   training   14% political   science 4% journalism 2% law 1% other 16% croatia economics 30% sociology 9% psychology 9% education   science  and   teacher  training   11% political  science 6% journalism 3% law 10% other 22% serbia figure 1 principal activity of researchers in bih, croatia, and serbia 24 iassist quarterly 2014 iassist quarterly in the field of psychology (28%), sociology (19%), economics (16%) and education science and teacher training (14%). in serbia most of the researchers are in the field of economics (30%), while other disciplines, education science and teacher training (11%), law (10%), sociology and psychology (9%) have a much smaller share. differences in the structure of the respondents might be biased by the willingness of respondents to cooperate on collegial solidary basis, in regard to the primary discipline in which institutions, that sent questionnaires to researchers, are engaged4. ( see fiqure 1) producing data the second part of survey was dedicated to questions related to production of data and research activity within the past five years. the first question attempted to determine how many datasets were produced during that period. in each country over 50% of ] figure 2 type of stored data in bih, croatia, and serbia iassist quarterly 2014 25 iassist quarterly researchers confirmed that they produced five or more datasets during the past 5 years and based on this we can conclude that there is enough research potential in our countries, and that our researchers produce a substantial amount of datasets. in all countries the largest numbers of researchers have produced between 6 and 10 datasets (bih 19%, croatia and serbia 24%). but there is a slight difference between the countries in the case of a subsequent frequency; in bih 13% of researchers produced 11-20 datasets and 12% produced 21 and more. in croatia 19% of researchers produced 5 datasets, while only 8% produced 10-20. in serbia the situation is similar to croatia because 15% of researchers produced 5 datasets. insight into the amount of research figure 3 current access to data vs. ideal level of access in bih, croatia, and serbia (darker columns present current situation and lighter present opinions about ideal situation) 26 iassist quarterly 2014 iassist quarterly conducted was very important because it showed that there is a sufficient number of datasets in all three countries and that the establishment of the data archive is justified, considered this criterion. additionally, we determined that the number of datasets generated has a growing trend. the results of all three countries are almost identical. within last five years researchers completed 45% of all datasets in 2012, 36% in 2011 and the rest in 2010 and earlier. methods of data gathering questions about applied data collection methods and the financial sources for projects were an essential part of the third segment of our survey. the question about applied data collection had openended answers with several offered examples (online questionnaire, structured interview, focus groups, experiment, etc.) so we received different results among the countries. multiple answers were allowed and percentage share for each answer is defined with respect to total number of respondents (such approach is also applied in the archiving practices and preferences section). in bih most researchers used either questionnaires (32%) or interviews (47%) in data collection in the last five years. in croatia, the dominant method was surveys (53%) and quantitative (70%) or qualitative (33%) questionnaires (focus groups and interviews). questionnaire (49%) was the dominant method of data collection in serbia. results for the second question, about financing of research was very interesting to us. in croatia and serbia approximately 40% of all research was financed by public funding through national science funding bodies, while in bih most of the research (40%) was financed through international funds/projects while only 7% had the support of national funding. international funding is also provided in the other two countries, but with a much smaller share compared to bih. in croatia, 17% of research was funded that way, and in serbia 26%. the rest of the projects had been funded by institutions which conducted projects, publicly funded from other sources, private sector, and other. the method of financing research in croatia and serbia is not the best possible, because the funds for the science are allocated from the budget of these countries. therefore, these resources are often insufficient, particularly for the social sciences. archiving practices and preferences the most important and interesting segment of the survey was the fourth part about existing archiving practices and preferences of researchers in three countries. answers to the first question about type of stored data were very similar in all countries. most of the researchers keep the data in the raw form, as data prepared for analysis (with transformations, created index, and recorded), or as cleaned data. but the main obstacle for further use is the absence of well-documented data with metadata. detailed answers are presented in the figure 2. answers to the question regarding where researcher keep data stored are also consistent between countries. the dominant number of researchers keep the data in their own computers (in average over 50%) or several copies in different computers (in average over 40%), and only a few of them keep the data in some form of institutional repository (approximately 3%). our opinion is that the obtained result is not good for three reasons. first one is that researchers usually do not have proper procedures for backup, so the risk of losing the data is very large. also, the other researchers in most cases do not have access to the data kept on personal computers and finally absence of well documented data completely prevents the reuse of these data, because it is impossible to find them. during the design of our survey we assumed the answer on the previous question and the related problems. therefore, we allowed respondents to provide comparative answers to questions about their current level of access to data and what would be the ideal level of access, according to their opinion. in figure 3 we present comparative values obtained on two questions. it is very interesting to see that in all three countries current access to the data is dominantly limited to the research team; however most researchers think that data should be publicly available (open access) or at least available to the broader scientific community. these results are very encouraging for us regarding establishment of digital data archives in our countries. the last question in this segment was intended to provide insight on the willingness of researchers to provide research data to an archive, if the data would be safely preserved and access regulated. the great majority of researchers in all three countries want to provide research data to archive if the data would be safe with regulated access because 45% of them (in average) answered with yes, certainly and 40% answered with yes, probably. these responses are very encouraging, because they indicate a justification for the establishment of digital archives in bih, croatia, and serbia, as well as a positive attitude within the research community about the future deposits and secondary use of data. nevertheless, we are fully aware that the number of positive answers is probably higher than the real disposition of overall population of social scientists, primarily due to self-selection of survey participants based on their interest in the subject of research. nonethelss, a data-archival institution would obviously address the reported existing needs and help in overcoming the current issues identified with respect to safe archiving and enabling of access to research data. use of data and secondary analysis the last part of our survey was conducted to understand the practices of researchers in the field of secondary analysis and use of data. more than half of the respondents (bih 75%, croatia 51%, serbia 64%), have stated that the sharing of research data is very important in their discipline and only 2% (on average) find it not very important. the answers to the question “would your scientific work benefit if you had better access to research data produced locally or internationally?” were expected. we offered two modalities for this question (for local and international research data). in bosnia and serbia the great majority of researchers stated “yes, considerably” as the predominant answer for both types of research data, while croatian researchers considered that their scientific work would benefit more considerably from better access to international than local research data. detailed responses are presented in the figure 4. we anticipated these answers due to the existing problems in funding of scientific work as the researchers are always faced with a lack of data. the final question assured us of the great possibility for establishing data archives in bih, croatia, and serbia, because over 50% of researchers consider it as very useful and only 1% as not useful at all. but regarding these answers we must be very careful. the researchers in these countries do not have appropriate knowledge about data archives, just general information about them or iassist quarterly 2014 27 iassist quarterly absence of any knowledge. we assume that the situation will change by raising awareness of this issue in the near future. conclusion in this paper we report findings of a survey conducted in bosnia and herzegovina, croatia, and serbia based on standardized questionnaires for all three countries on a representative sample, aimed to shed light on existing potentials for the establishment of social sciences digital data archives. we analyzed numerous single issues grouped into four general topics comprised by the survey—production of primary data, methods of data gathering, archiving practices and preferences and use of data (secondary analysis), with particular interests in scope and quality of data production, documentation, and dissemination. similarity in the structure of responses among countries allowed us to easily generalize our conclusions across general topics, despite expected variations in responses on the level of particular issues. results of the survey related to issues of data production are quite promising, as researchers in social sciences in all three countries have been considerably active in recent years producing between 6 and 10 datasets, and what is more important, that scope of data production has a growing trend. however, the situation is not so bright when comes to the issues of data documentation and especially data dissemination. most researchers keep data either in raw form or partially prepared for further analysis, but without appropriate documentation with metadata in accordance to international standards. the absence of appropriate documentation is identified as the main obstacle for further use of data in secondary analysis. therefore, it is not surprising that researchers mostly keep data in their own computers (in average over figure 4 better accesses to data as benefit for scientific work 28 iassist quarterly 2014 iassist quarterly 50%) or several copies in different computers with an approach limited to research teams. it is encouraging that researchers are willing to share their data and are also aware of the benefits that centralized data archive could bring about. thus, we can conclude that potentials for establishment of digital data archives in considered countries are large, regarding the scope of data produced in social sciences and the positive attitude of researchers toward data sharing and benefits of data archives. in addition, existence of national data archives, with staff trained in assisting researchers in the process of data documentation and preservation according to international standards, will remove key obstacle for data sharing and further use of data in secondary analysis. . references 1. arzberger, p., schroeder, p., beaulieu, a., bowker, g., casey, k., laaksonen, l., moorman, d., uhlir, p., wouters, p. (2004). promoting access to public research data for scientific , economic, and social development. data science journal, 3 (november), pp.135-152. 2. bock, g., zmund, r., kim, y., lee, j. (2005). behavioral intention formation knowledge sharing: examining the roles of extrinsic motivators, social-psychological forces, and organizational climate. management information systems quarterly, 29(1), pp. 87-111. 3. bradić-martinović, a., zdravković, a. (2012). integration of western balkan countries into the european system of digital data archives in social sciences: case of serbia. review of applied socio-economic research. vol. 4, issue 2, p.p. 32-41 4. corti, l., van den eynden, v., bishop, l., morgan-brett, b. (2011). managing and sharing data. uk data archive, colcehester, essex. 5. ukda, (2002). preserving and sharing statistical material, the royal statistical society & the uk data archive, university of essex. 6. world bank, knowledge for development – k4d, http://web. worldbank.org/wbsite/external/wbi/wbiprograms/kfdlp/0,,co ntentmdk:20269026~menupk:461205~pagepk:64156158~pipk:641 52884~thesitepk:461198,00.html (last visit 05 may 2014) notes 1. aleksandra bradić-martinović, phd is research fellow in the institute of economic sciences in belgrade, serbia in the fields of business information systems, e-banking and data management. she is also engaged in teaching as an associate professor in belgrade banking academy in the same fields and can be reached by email: abmartinovic@ien.bg.ac.rs. 2. aleksandar zdravković; ma is research associate in the institute of economic sciences in belgrade, serbia in the field of econometrics and macroeconomics. he is also a phd student at economic faculty university of ljubljana and can be reached by email: aleksandar. zdravkovic@ien.bg.ac.rs. 3. the more detailed information about serscida project can be found at www.serscida.eu. 4. survey was conducted by three institutions, local participants in serscida project. in bih it was human right centre, university of sarajevo. in croatia it was faculty of humanities and social sciences, university of zagreb and in serbia it was institute of economic sciences, belgrade. 38 iassist quarterly 2014/2015 iassist quarterly linking study descriptions to the linked open data cloud by johann schaible1, benjamin zapilko2, thomas bosch3, and wolfgang zenk-möltgen4 abstract the gesis data catalogue contains the study descriptions for all archived studies at gesis, currently more than 5000 datasets mainly from survey research in the social sciences. these descriptions include information about primary researchers, research topics and objects, used methods, and the resulting dataset, which is mainly used for archiving and retrieval in order to serve secondary researchers. for this purpose the existing metadata can be enriched with further information about the study investigators, involved affiliations, collection dates, content, and more from other sources like dbpedia or the name authority file of the german national library. in recent years the paradigm of linked open data (lod) encouraged various research organizations to expose their data to the web according to semantic web standards. this has increased the number of available data sources and the feasibility of their reuse. in this paper, we present ways to enrich a study description with various datasets from the lod cloud. to accomplish this, we expose selected elements of the study description in rdf (resource description framework) by applying commonly used vocabularies. this optimizes the interoperability to other rdf datasets and the discovery of links to them. for link detection we use silk, a framework for discovering relationships between data items within different lod sources. once links are detected, the study description is linked to adequate entities of external datasets and therefore holds additional information for the user, e.g. further metadata on the principal investigator of a study. keywords: semantic web, linked open data, data transformation, rdf, link discovery, metadata introduction the linked open data (lod) cloud5 comprises data from diverse domains. various best practices and principles (bizer et al. 2009) guide a data publisher in modeling and publishing data as linked data. to use semantic web technologies such as rdf6 and sparql7 and to include links to external data providers are two essential points in the guidelines, as this leads to better discovery of information by linked data applications and users (heath and bizer 2011). the gesis data catalogue (dbk)8 comprises study descriptions for all archived studies at gesis. it contains metadata about each study, such as the primary researchers, research topics and objects, used methods, etc., which is archived to serve as an information pool for secondary researchers. thus, the visibility of such a dataset is an important aspect. to publish this metadata as linked open data would increase the visibility because external data providers can set links to particular linked data sources. this way, secondary users are able to discover the data from multiple points of access. furthermore, the existing metadata can be enriched with additional information from other external data providers. for example, the gesis data catalogue can be enriched with additional information about the study investigators, involved organizations, collection dates, content, and more from external sources like to publish metadata as linked open data would increase the visibility of the data iassist quarterly 2014/2015 39 iassist quarterly dbpedia9 or the name authority file (gnd)10 of the german national library (also named pnd). note that the publication of the metadata as lod is intended, not the publication of the quantitative dataset. in terms of computer science both are data and could be published as lod. but the quantiative datasets can only be ordered or downloaded by agreeing to the usage regulations of the gesis data archive. however, the metadata of the data catalogue is freely available and was modeled as lod in this paper. please note, that the lod representation of the data catalogue has not been published yet, and the links provided in various examples are as yet hypothetical. in this article, we describe the modeling and the publishing of a dataset as linked open data and the procedure for how to interlink this resulting linked dataset to external data sources. hereby, we especially focus on the difficulties in producing linked open data. our dataset is an excerpt from the gesis data catalogue comprising specific metadata about social science studies. this metadata is stored as xml flat files. the mapping to existing rdf vocabularies is done manually. to transform it into rdf, we use plain xslt scripts. we use the link discovery tool silk11 to detect links from the rdf representation of the gesis data catalogue to external data sources. we discuss our observations on the benefits of the described approach to publish data. in detail, we inspect whether we gain any efficiency in handling of the data, whether we gain new information from external data providers, and what is possible with such a dataset stored in rdf in contrast to xml. we provide answers to these questions with respect to the effort and difficulty in producing such linked open data. the article is structured as follows: in section 2, we describe the gesis data catalogue in detail. furthermore, we illustrate what metadata it contains and which data elements we used for our excerpt. in section 3, we demonstrate the transformation of the xml data into rdf. this also includes the choice of the existing vocabularies as well as the mappings to terms from these vocabularies. section 4 provides an insight into the link discovery framework silk. we present how silk can be used to detect links to external datasets containing information on the same resources. we present the results of our work in section 5 and describe the advantages and the disadvantages of publishing data as linked open data. in section 6, we conclude our work and give an outlook to future work. the gesis data catalogue (dbk) the gesis data catalogue (dbk) comprises the study descriptions from all archived studies and empirical primary data mainly from survey research and historical social research which are published on the gesis homepage by the application dbksearch. it is possible to search within the study descriptions by using a simple or advanced search. the simple search is carried out in all or selected fields, whereas the advanced search combines more search terms in different fields. the management of this metadata is implemented by the dbkedit application that also handles internal metadata and workflows. the gesis data archive uses the data catalogue also to publish the metadata in other portals and systems, such as zacat12, the cessda data portal13, sowiport14, and the data registration agency da|ra, which again is linked to the metadata store of datacite16. the applications dbkedit and dbksearch are also available as an open-source for other providers under the name dbkfree17. the list of structured data which describes a dataset of the archive and makes it easier to find is defined by the metadata schema of the data catalogue (zenk-möltgen and habbel 2012). since the establishment of the central archive for empirical social research 50 years ago (now part of gesis), the metadata schema as a system for study description has always been refined in the context of the cooperation of the international archives and is continuously being developed and adapted to new standards (mochmann 1979, bauske 1992, and bauske 2000). the metadata schema contains a number of mandatory core elements which have to exist for the creation of a new study description. furthermore, optional metadata elements can be used to describe the data more precisely. for some elements other applicable standards are used, e.g., iso standards for dates or geographic locations. the dbk metadata schema is compatible with the codebook and lifecycle standards of the data documentation initiative18 (ddi) and can be exported into the ddi2 and ddi3 xml formats. moreover, it is compatible with the metadata schema of the gesis agency for data registration da|ra and datacite (hausstein et al. 2011). in addition to the datacite metadata schema, the dbk metadata contains specific social science information which supports retrieval and especially allows for a methodological comprehensive description of research data. currently, other social science data archives like the icpsr19 in the u.s., dda20 in denmark, nsd21 in norway, and the ukda22 in the united kingdom use similar study descriptions for their holdings. to enrich the study descriptions with additional information using semantic web technologies, it is possible to publish the data catalogue as linked open data. for this the data catalogue xml files have to be transformed into rdf. we used the ddi codebook xml format and extracted some entities from the dbk which seem to be most promising with respect to finding additional information for the studies. for example, “title,” “author,” and “abstract” are such important entities, but “caseqnty” (number of variables in the data file) is not. following is the entire list of the selected important entities for a study description, and figure 1 displays a pseudo-xml of the structure of the entities. • title statement: the title statement contains a mandatory element “title” and an optional list of elements named “alternative title“. alternative titles can also be of the type project title, original title, or subtitle. • responsibility statement: the responsibility statement contains the repeatable element “authoring entity” with an “affiliation” of the authoring entity as an attribute. this element contains the principal investigators that should be cited for the creation of the study. their institution is named in the affiliation attribute. sometimes institutions are named directly as the principal investigator. • production and distribution statement: the production statement comprises the elements “producer” and “distributor”. the distribution statement currently contains the name of the gesis data archive with its abbreviation and website url as attributes. the element “funding agency” is currently not used by the dbk in the ddi study descriptions. • study info: in the entity study info there is a list of topic classifications for the study from the za-category system and a detailed thematic description of all the variables in the dataset in the “abstract” element. both elements are available in german and english, but for some study descriptions there is still a lack 40 iassist quarterly 2014/2015 iassist quarterly of translations into english for the abstract. in this section there is also the list of “geographic coverage,“ which contains country and region names from the iso-format and additional free text, and a description of the “universe“ that the data applies to (both language dependent). • data collection: in this entity there is the list of collection dates in iso-format under the element “time method”. in addition, there are the elements “data collector”, “sampling procedure”, and “collection mode” (all language dependent) which describe the methodology of the data collection process. • data access: data access comprises a section “data set availability” which contains the element “access place” for describing the location of the access place and an uri of the place as attribute, and the “availability status” of the study which is described in english and german. • other study material: in the entity “other study material“ there are the elements “related material” containing data and document files that may be downloaded with name and url, “related publications” with the full citation, and “other references” with further remarks that may contain notes to the study (language dependent). converting the data catalogue xml into rdf to convert xml data into rdf two steps have to be passed: the mapping and the technical conversion. while the latter step can be solved by writing and executing scripts like xsl transformations, the mapping of xml elements to rdf properties and classes requires expert knowledge for the domain of the data as well as for semantic web vocabularies. that is because on the one hand the data must be converted correctly to rdf without losses or changes in its semantics. on the other hand interoperability with other data expressed in rdf and semantic web applications has to be ensured. as described in bizer et al. (2009) and heath and bizer (2011), it has become best practice to reuse properties and classes of existing and popular semantic web vocabularies as much as possible. but the search for the most adequate properties and classes for representing the semantics of the source xml data can be a time-consuming task, especially if there are several potential suitable rdf vocabularies or if the data is not fully covered by them. the search is complicated since the number of rdf vocabularies has increased massively during recent years. hence it requires expert knowledge for deciding which vocabularies should be used for representing the data. there are several typical decisions that have to be made when defining a mapping of metadata entries to properties and classes of rdf vocabularies. some of them depend on the trade-off between a semantically rich expressiveness of the resulted rdf data and an intensive reuse of existing and popular vocabularies. one has to decide consistently for the full mapping and especially for particular data elements whether a correct and full semantic expressiveness of the data or a technical interoperability with other linked data sources is of higher relevance. this influences directly the amount of used vocabularies and whether the definition of stdydscr [study description] citation titlstmt [title statement] titl [title] alttitl [alternative title] rspstmt [responsibility statement] authenty [authoring entity] @affiliation prodstmt [production statement] producer fundag [funding agency] diststmt [distribution statement] distrbtr [distributor] @abbr [abbreviation] @uri stdyinfo subject [language depended] topcclas [category; language depended] abstract [language depended] sumdscr colldate [collection date] universe [language depended] method datacoll [data collection] timemeth [language depended] datacollector [language depended] sampproc [sampling procedure; language depended] collmode [collection mode; language depended] dataaccs setavail [availability statement] accsplac [access place] @id @uri usestmt [usability statement; how to use the study?] contact othrstdymat [other study material] relstdy [related study] relpubl [related publication] othrefs [other references; further remarks] figure 1: the extracted entities from the dbk as pseudo-xml iassist quarterly 2014/2015 41 iassist quarterly an own vocabulary becomes necessary. if the preservation of the semantic meaning of every data element is the highest goal for a conversion, then it is very likely that not all elements can be represented by existing rdf vocabularies and it is necessary to define individual classes and properties in their own vocabulary. the following examples present cases where these considerations are of importance: • in some cases there is more than one adequate property or class to represent a particular data element. for instance, there are several properties for describing elements of the xml data e.g., title or date. these properties are typically part of different vocabularies like dublin core23 or particular bibliographic vocabularies. one has to decide which property or class of which vocabulary to use for the representation of a particular data element. • there may be a loss of semantics when mapping a data element to a property of a popular vocabulary instead of mapping it to a property of a less popular vocabulary, which represents the semantics of the element more precisely. for instance, the data element describing a particular time (e.g., the time period observed in a study) is not represented adequately by the general date property from the dublin core elements vocabulary instead of a more precise property of a lesser known vocabulary. • two data elements with the same data type, but a slightly different semantic meaning, e.g., starting date of a survey and modification date of a dataset, can lose their meaning if they are represented by the same property (again, e.g., the date property from the dublin core elements vocabulary). such data elements should be represented in rdf by different properties in order to keep the semantic difference between them. additionally, it has to be decided whether data elements should be represented as resources or as properties. a resource is represented with an uri and is in a general sense a “thing”. every resource has properties, which we define as literal values describing the resource. this design decision has to be made carefully, because only resources can be linked to other resources of the linked open data cloud. the instances of properties are commonly expressed as plain literals and cannot be enriched by further information and links. for example, if the principal investigator of a study were modeled as a literal value, it would not be able to interlink this property with an external dataset containing information about persons. on the other side, if the principal investigator were modeled as a resource, it can be interlinked with another resource from an external data source. the structural difference between a resource and a property is defined in the structure of an expression in rdf, as it is a collection of triples, each consisting of a subject, a predicate, and an object. the subject is in most cases an rdf uri that references a resource. the object is usually either an rdf uri that also references a resource or a literal value describing the subject. the predicate is also an rdf uri that links the subject to the object. for example, the resource “study” is the subject. it has the object “principal investigator”, which is also a resource, and the object “study title” that is denoted as a literal. the predicate “hastitle” links the resource “study” to the object “study title” containing a literal value, and the predicate “hasprincipalinvestigator” links the resource “study” to a resource “principal investigator”. these two expressions are considered to be triples. for the conversion of study descriptions to rdf in order to detect links we decided to reuse existing vocabularies, but as few of them as possible. by choosing popular vocabularies we allow for high interoperability with other datasets of the lod cloud. this was also the reason we did not define our own properties and classes, although some data elements cannot be covered to the same full semantic extent in rdf as in their original xml representation. the choice of reusable vocabularies that can express the dbk entities in the best possible way was based on the description of the vocabulary and its human-readable documentation. as most appropriate vocabularies, we have identified the ddi-rdf discovery vocabulary (disco)24 , the dublin core vocabulary (dcterms), as well as the semantic web for research communities vocabulary (swrc)25 . the disco vocabulary covers many ddi2 elements that are used in the data catalogue ddi2 xml export. however, all of the terms from the disco vocabulary that were considered as appropriate mapping are reused classes and properties from the dublin core vocabulary. thus, it is more convenient to use the classes and property from dublin core directly. the swrc vocabulary is widely used to model entities of research communities such as persons, organizations, and bibliographic metadata on publications, which suits our purpose very well. as mentioned earlier, the first step to transform the data catalogue xml files into rdf is to map the various entities to the classes and properties from the vocabularies we have identified as most appropriate. table 1 shows the possible mappings of all entities dcterms swrc title dcterms:1tle  (*) swrc:1tle alterna1ve  title dcterms:alterna1ve  (*) authoring  en1ty dcterms:creator  (*) swrc:author affilia1on swrc:affilia1on    (*) producer dcterms:agent  (*) distributor dcterms:publisher  (*) category dcterms:subject  (*) abstract dcterms:abstract  (*) swrc:abstract universe dcterms:coverage  (*) time  method dcterms:date  (*) swrc:startdate
 swrc:enddate data  collector dcterms:contributor  (*) sample  procedure dcterms: 
 accrualmethod  (*) collec1on  mode dcterms: 
 accrualmethod  (*) access  place dcterms:loca1on  (*) related  publica1on dcterms:rela1on  (*) other  references swrc:note  (*) table 1: mapping of the data catalogue entities to terms from the different vocabularies. the vocabulary terms marked with a “(*)” are the one that were chosen to be used 42 iassist quarterly 2014/2015 iassist quarterly from the dbk excerpt to the terms from the different vocabularies. we finally mapped the entities in the left column to the terms that are followed by an asterisk (*). the mapping was done manually. this way it was likely to preserve as much of the semantic richness of the data as possible. the technical process of the conversion can be conducted by different scripting languages. since the source data is xml and rdf can also be serialized in xml, it seems likely to use xsl transformations. hereby, we extracted the entities from the xml we intended to express in rdf and defined an xslt script, where we specified how the entities should be transformed. figure 2 provides an example that shows how we have transformed the title entity of an xml file into an rdf representation re-using the dublin core property dcterms:title. we can see in figure 2 that the xml element provides the information about how an entity is encoded. we use this information to make an xslt script and generate an rdf property. we first identify the entity “title”, which is marked purple in the xml. it has a language attribute that is marked orange and a value, which is marked blue. in the xslt we define a new element with the name “dcterms:title” that has a new attribute with the name “xml:lang” and the value “en”. additionally, the value from the xml element is extracted using xpath from the path “titlestmt/ title/”. this results in a new property dcterms:title in rdf that has a language attribute and the value from the xml. this procedure has to be done for every entity in the data catalogue xml. it is very important to note that the example in figure 2 does not display an entire and valid rdf representation, as it only a single rdf property, without a subject to complete the triple. discovering links to external data sources the rationale for publishing data as linked open data is to increase its visibility and make it easier for secondary users to consume the data, but also to gather information from other data providers who published their data as linked open data. to achieve the latter, we have to identify external data sources that might hold noteworthy data; second, we have to discover links to equivalent resources; and third, we have to include the links in our rdf representation. the search for external data sources containing further information for the data catalogue’s study descriptions was performed manually, since currently there is no satisfactory way of searching lod instances automatically. the data hub26 linked open data group provided an appropriate set of data sources for this, as it contains all datasets included in the lod cloud. the first candidate is the integrated name authority file (gnd). it originates from the german library community and contains a broad range of elements to describe authorities in detail. this way it aims to solve the name ambiguity problem. another candidate that might comprise data for enriching the study descriptions is dbpedia. it contains structured information that was extracted from wikipedia, i.e., the information boxes on the top right corner of many wikipedia pages. the data comprises information on persons, places, organizations and more. to discover links to instances from these two external data sources, there are so-called “link discovery tools”. one of these tools is silk – a link discovery framework. it detects relationships between items within different linked open data sources based on various comparison methods that are applied on literal properties of all items. the included comparison methods cover typical similarity measures like levenshtein distance, jaccard similarity coefficient, or even geographical distance. figure 3 displays the general workflow of this procedure, where the relationship is defined as owl:sameas and the comparison method is an absolute string equality measure. if the value of “property 1” in the initial dataset is equal to the value of “property 1” in the external dataset, the value of “property 2” in the initial dataset is equal to the value of “property 2” in the external dataset, and the value of “property 3” in the initial dataset is equal to the value of “property 3” in the external dataset, the both resources are considered to be related to each other in the meaning of owl:sameas (note http://www.w3.org/tr/owlref/#sameas-def). this relatedness is expressed by a value, which is computed out of the applied similarity measures. as a benefit the figure 2: the xsl transformation of the entity “title” iassist quarterly 2014/2015 43 iassist quarterly properties “property a” and “property b” in the external dataset can now be gathered as additional information. to guide the user through the process of creating link specification for such relationships, silk provides the “silk workbench”. the user has to go through three basic steps: (1) specify the data sources and the linking tasks, (2) define explicit linkage rules, and (3) evaluate the correctness of the discovered links. in the following, we will describe the link discovery procedure along an example study description from the data catalogue. for the first step, silk allows the user to specify several data sources by either providing the sparql endpoint of the data source or its rdf dump that has to be downloaded and stored on the local machine. figure 4 shows the data catalogue and the gnd data sources (named pnd) that are specified as rdf dumps. after defining the data sources, it is essential to specify a linking task. the user can also denote an output file, where all results can be saved, but this has to be done for every linking task. a linking task describes what kind of relationship shall be found between two data sources. therefore, the user has to declare the source dataset, the target dataset, and the link type. figure 5 illustrates that for our work we have chosen the data catalogue as the source dataset and the gnd name authority file as the target dataset. the link type is set to owl:sameas, as we intend to find equal resources. this is the most common approach to find the resource within an external data source. data represented in rdf is structured as a graph. the user can add source and target restrictions that specify the node in the rdf graph from which silk starts to compare the property values. this can be very helpful if the data is very big or if specific concepts should not be part of the comparison. if no restrictions are provided, silk starts at the root node. having defined the linking task along with two data sources, the user comes to the second step and has to define linkage rules that specify how two literal values have to be compared. hereby, silk displays a set of all properties used in both data sources the user has specified in the previous step. the user chooses the properties he intends to compare. to accomplish this task, silk provides an intuitive drag and drop mechanism. every literal value can also be transformed, e.g., by capitalizing or extracting all numerical values, in order to avoid miss matches due to different encoding schemes. also, it is possible to select different comparators. for example, the user can choose the comparator that utilizes the levenshtein distance. this way it is possible to deal with spelling mistakes. figure 6 illustrates the creation of such a linkage rule. it is shown that for our purpose we selected the property swrc:name from the data catalogue and the gnd:preferrednamefortheperson from the gnd name authority file. each value of these properties is transformed to lower case. then each value of swrc:name is compared to each value of gnd:preferrednamefortheperson by applying the levenshtein distance. the user can also specify other options for the comparator to make the comparison even more precise. furthermore it is possible to compare several properties with each other. for example, the user could also compare the values of the properties foaf:birthday and gnd:dateofbirth. this allows the user to define linkage rules such as “only if the names are the same and the birthday dates are the same, then the resources should have an owl:sameas relationship”. figure 3: general link discovery procedure with owl:sameas as defined relationship 44 iassist quarterly 2014/2015 iassist quarterly the third step comprises the evaluation of the links that silk has detected between the two specified data sources. as an output, silk shows the compared values and to which percentage it considers the resources to be related. figure 7 displays such an output. it is displayed that the comparison is a levenshtein distance transformed on the input properties swrc:name and gnd:preferrednamefortheperson. the values “tomka, miklós” and “tomka, miklós” are considered to be a 100% match. therefore the resources containing these properties are considered to be related in the meaning of owl:sameas. as a result, the detected link can be included in the initial dataset of our study description and thus enriches it with additional information. according to figure 7 this would be the link to http://d-nb.info/gnd/134232240 (the person tomka, miklós). the entire procedure including all the three steps that were explained in this section has to be done for every concept which is intended to be enriched with additional information. for example, the data catalogue comprises descriptions of topical categories of the studies. these are mostly very general terms such as “political attitude” that can be linked to similar terms from dbpedia or various thesauri like the gesis thesoz (zapilko et al. 2012). results based on the entities that we have extracted from the data catalogue, the rdf modeling decisions, and the chosen external data sources we intended to link to, silk was able to detect links to enrich the data on various entities. the name authority file of the german national library provided a lot of additional information on persons who contributed to a study. unfortunately, we were not able to gather further information from dbpedia on the topic category of a study, as silk did not return any links. the same applies to the specification of the data of a study. we intended to link it to an extraction of time events from wikipedia that is published as lod (hienert and luciano 2012). however, no links were detected, as the dates in the dbk data are encoded as a timespan (“january 2003 to december 2003”), whereas the dates from the extracted time events are encoded as a point of time (“2003”). another challenging task was the disambiguation of a person, as the set of the first name, the last name, and the affiliation is simply not unique to certainly identify a person. we did not set the levenshtein distance very low in order to link resources despite spelling mistakes. hence, the evaluation of the discovered links took longer than intended to ensure the disambiguation of the persons, and sometimes it was simply impossible. the topic category of a study in the data catalogue is described with terms from a controlled vocabulary27. in rdf the category was first described as a resource that had the terms from the controlled vocabulary as a property. we designed a linkage rule in silk, which compared the term from the controlled vocabulary figure 4: the definition of the data catalogue and the gnd data sources as rdf dumps in silk figure 5: defining an owl:sameas link type between the data sources data catalogue and the gnd iassist quarterly 2014/2015 45 iassist quarterly with the labels of articles in dbpedia. for example the category “income” was supposed to be linked to the dbpedia data of the wikipedia article about “income”. silk did not find any links, though. this was due to several reasons. first, there are a lot of articles in wikipedia that do not have a structured information box. therefore, there is no dbpedia entry for such articles. second, some categories are described with multiple terms, like “legal system, legislation, law”. dbpedia on the other hand does not describe entities with multiple terms. therefore, silk will not find any links between resources that probably describe the same thing, but the comparison of their properties fails due to syntactical difference. to bypass these problems, we have mapped the categories of the study descriptions to concepts of the thesaurus for the social sciences (thesoz). thesoz has been already published as linked data (zapilko et al. 2012). this way we were able to gather additional information from the thesoz such as the translations of the categories in german and french language as well as the hierarchical structure of the categories. for further information we specified the linking rules in silk to detect links between the concepts of the thesoz and other thesauri such as eurovoc28 . silk detected these links without any problem providing further information about the categories. the properties “abstract” and “other references” of a study description were not as helpful for discovering links as we had intended. in order to use the information within these entities, some natural language processing (nlp) algorithms have to be applied to extract keywords and perform a link discover using those keywords. however, this is not part of this work, but can be strongly considered as future work. besides discovering links for the authority entities, topic categories, and the date of a study, the properties “title”, “alternative title”, “producer”, and “publisher” were modeled to help to link the study another instance of itself from an external data source. unfortunately, such a data source was not found on the linked open data cloud. the remaining properties “universe”, “data collector”, “sample procedure”, “collection mode”, “access place”, “related publication”, and “other references” have not been used yet for link discovery and remain as future work. conclusion and discussion in this work, we have demonstrated how semantic web technologies can be used to link study descriptions to external data sources and enrich them with additional information on various entities such as contributors and the categories of the study. we first extracted the entities, which we intended to enrich with further information and several other entities, which seemed to be most promising to help the link discovery process. we transformed the representation of the study description from xml to rdf using xslt scripts. hereby, we provided detailed information on the difficulties of such a transformation, especially the mapping of entities to classes and properties from existing vocabularies. we illustrated the workflow of the linking process with the link discovery framework silk along with an example and provided the results of our work. for publishing data as linked open data, one has to have good knowledge of rdf as well as the principles and best practices of the modeling and publishing process. “it is especially important to understand whether information should be published as a resource or as a literal, if the intention is to interlink the data with external data sources. in linked open data only resources can be linked together via link types like the owl:sameas statement. therefore, if the intention is to gather additional information on a specific entity such as the principal investigator, it has to be modeled as a resource containing properties that describe the resource such as “first name” and “last name”. for disambiguation purposes, it is strongly advised to use unique identification characteristics such as an isbn number for books, or orcid29 for researchers. another possibility to disambiguate entities is to use several identification characteristics such as “first name”, “last name”, “birthplace”, and “date of birth”. if the aim is to reuse existing vocabularies to express the data, it is important to know which figure 6: definition of the linkage rule to compare two property values with the levenshtein distance figure 7: the result from the comparison 46 iassist quarterly 2014/2015 iassist quarterly vocabularies will fit the best. resources are modeled as classes, so it is important to investigate several vocabularies to determine if they provide classes that can represent entities as resources in a semantically correct way. the same applies to entities which are intended to be modeled as properties. if the data publisher does not know such vocabularies, the search for them might result in a lot of effort. to help the data publisher to find appropriate terms from existing vocabularies, there are vocabulary search engines like lov30 or swoogle31 , or novel concepts that recommend classes and properties during the modeling process (schaible et al. 2013). to link to external data sources, one has to discover such data sources in the first place, for example, by searching a repository like the data hub. the next step is to understand the structure of the external datasets and locate the concepts of interest for linking and their properties for comparison. in the beginning, this might be time consuming, as datasets are generally modeled differently. however, this is a crucial step because it is necessary to specify the linkage rules in link detection tools like silk. the setup of datasets, linking task, and linkage rules in silk is straightforward. nevertheless, several problems did occur, due to the complexity of the specifications of comparison methods and the not very detailed documentation. data from the domain of the social sciences are not very widespread in the linked open data cloud. to find additional information on such type of data is very hard. once the lod cloud gets populated with datasets covering social science studies with detail about their contributors, it will be a lot easier to link the gesis data catalogue to these data sources and thereby enrich its study descriptions with additional information. one example for such a domain would be the publications of scientific papers in the area of the semantic web, as was discussed by schaible and mayr (2012). references bauske, f. (1992), ‘europäische informationsbasis über datensätze in cessda-archiven‘, za-information, vol. 31, pp. 109-111. bauske, f. (2000) ‘das studienbeschreibungsschema des zentralarchivs‘, za-information, vol. 47, pp. 73-80. bizer, c., heath t. & berners-lee, t. (2009), ‘linked data-the story so far’, international journal on semantic web and information systems, vol. 4, no. 2, pp. 1–22. brank, j., grobelnik, m. & mladenić, d. (2005), ‘a survey of ontology evaluation techniques’, proceedings of the conference on data mining and data warehouses (sikdd). hausstein, b., zenk-möltgen, w., wilde, a. & schleinstein, n. (2011), ‘da|ra metadatenschema version 1.0.‘, gesis working papers 2011/14, doi:10.4232/10.mdsdoc.1.0. heath, t. & bizer, c. (2011), ‘linked data: evolving the web into a global data space’, synthesis lectures on the semantic web: theory and technology, vol. 1, no. 1, pp. 1–136. hienert, d., luciano, f. (2012), ‘extraction of historical events from wikipedia’, proceedings of the first international workshop on knowledge discovery and data mining meets linked open data (know@lod 2012). mochmann, e. (1979), ‘bericht über die iassist konferenz in ottawa‘, za-information, vol. 4, pp. 24-27. schaible, j., gottron t., scheglmann s. & scherp a. (2013), “lover: support for modeling data using linked open vocabularies”, proceedings of the joint edbt/icdt 2013 workshops (edbt ‘13), acm, new york, ny, usa, 89-92, doi=10.1145/2457317.2457332, http://doi. acm.org/10.1145/2457317.2457332. schaible, j., mayr, p. (2012): ”discovering links for metadata enrichment on computer science papers”, gesis-technical reports, 2012/10, köln: gesis. volz, j., bizer, c., gaedke, m. & kobilarov, g. (2009), ‘discovering and maintaining links on the web of data’, proceedings of the international semantic web conference (iswc), pp. 650-665. zapilko, b., schaible, j., mayr, p. & mathiak, b. (2012), “thesoz: a skos representation of the thesaurus for the social sciences”, semantic web: interoperabilty, usability,applicability, doi: 10.3233/ sw-2012-0081. zenk-möltgen, w. & habbel, n. (2012), ‘der gesis datenbestandskatalog und sein metadatenschema‘, version 1.8, gesis technical reports 2012/01. notes 1. johann schaible research associate and ph.d. student at gesis. unter sachsenhausen 6-8, 50667 köln, germany. email: johann. schaible@gesis.org 2. benjamin zapilko research associate and ph.d. student at gesis. unter sachsenhausen 6-8, 50667 köln, germany. email: benjamin. zapilko@gesis.org 3. thomas bosch research associate and ph.d. student at gesis. b2, 1, 68159 mannheim, germany. email: thomas.bosch@gesis.org 4. wolfgang zenk-möltgen team leader and project manager at gesis. unter sachsenhausen 6-8, 50667 köln, germany. email: wolfgang. zenk-moeltgen@gesis.org 5. http://lod-cloud.net/ 6. http://www.w3.org/rdf/ 7. http://www.w3.org/tr/rdf-sparql-query/ 8. https://dbk.gesis.org/dbksearch/ 9. http://dbpedia.org/about 10. https://wiki.d-nb.de/display/lds 11. http://wifo5-03.informatik.uni-mannheim.de/bizer/silk/ 12. http://zacat.gesis.org 13. http://cessda.net/data-catalogue 14. http://www.sowiport.de 15. http://www.da-ra.de/ 16. http://www.datacite.org/ 17. https://dbk.gesis.org/dbkfree2.0/ 18. http://www.ddialliance.org/ 19. http://www.icpsr.umich.edu 20. http://samfund.dda.dk/dda/default-en.asp 21. http://www.nsd.uib.no/nsd/english/index.html 22. http://data-archive.ac.uk/ 23. http://dublincore.org/documents/dcmi-terms/ 24. http://rdf-vocabulary.ddialliance.org/discovery 25. http://ontoware.org/swrc/ 26. http://datahub.io/group/lodcloud 27. https://dbk.gesis.org/dbksearch/categories.htm 28. http://eurovoc.europa.eu/drupal/?q=node 29. http://orcid.org/ 30. http://lov.okfn.org/dataset/lov/ 31. http://swoogle.umbc.edu/ irssist newsletter, vol. 2, no. 1 (winter 1978) news and notes under the revised format of the newsletter this section will include notices on a wide variety of topics. first, of course, will be organizational information dealing with iftssist. other standard features will consist of upcoming conferences and meetings; reports of organizations of interest to lassist members; and educational and/or research opportunities. the editor would welcome any additional suggestions concerning this section and will include information members would like to have available or make available to others. the absolute has not been realized. the current format allows new members of iassist to belong to an aclassist news tion group immediately upon entry into the organization. this format chairperson^ report, february does not readily facilitate prog151x7 "t'gtb ress on the designated products. sharon henry, canadian secretariat, the second north american iashas been asred to review the action sist conference took place on febgroup format and provide some alruary 8-11, 1978, at the carson ternative approaches. individuals inn, itasca, illinois. "state of interested in this activity should the art; perspectives" was the address comments to her. theme of the conference. the members of the program committee, tony during the business meeting in falsetto, public archives of canitasca a motion was made and passed ada; sheldon laube, c. m. leinwond to create the position of arcniassociates; richard roistacher. vist/recorder for lassist. this oniversity of illinois; and patrick position would be responsible for bova, national opinion research the preservation of the records of center, are to be commended for lassist including all reports, patheir efforts which made this meetpers, and other designated docuing a success. richard roistacher, ments. an appointment to this poswith the assistance of barbara noition will be made in the near ble. university of illinois, did an future, exceptionally fine job of nandiing the local arrangements. james daperhaps of greatest immediate vis, harvard university socioloorganizational significance of all eist, was the guest speaker at the the actions taken in itasca was the anguet. creation of a nominations and elections committee to handle the electhe conference was organized tion procedures for members of the around seven formal panels and seslassist steering committee. the sions of lassist action groups. committee is composed of elliott panel topics included the followavedon, university of waterloo; ing: documentation: privacy versus nancy carmichael, u. s. social scifreedom of information; software ence research council; and ekkehard and analysis techniques for nonmochmann, university of cologne, rectangular files; alternative orthe committee met several times ganizational arrangements for data during the itasca conference and access; network environments; acthrough those meetings it became quisition and preservation; and clear that some revisions to the networking products and services. constitution were necessary. the the papers presented during the appropriate revisions and election panels will be made available in procedures are presented in a sepathe form of published proceedings rate insert to this newsletter. and a selection of the papers will special attention was given €o geobe featured in the lassist neusletgraphical distribution of the msmter, now under the editorship of bership to assure adequate reprethomas wm. madron, western kentucky sentation. university. seventy-five individuals attended the meeting. a list carolyn geda, chairperson of participants is available through judith rowe, u. s. secre1979 north american lassist one of the developments at the conterence itasca conference was some discussion concerning the role of the acthe 1979 north american lassist tion groups, which in turn is stimconference is scheduled for ottawa, ulating further review. orginally canada, during may, 1979. as plans the action groups developed with for the meeting are developed, !:ur"products" as the objectives, but ther information will be published 25 lassist newsletter, vol. 2, no. 1 (winter 1978) in the newsletter. in the associations. u. s. delegates meantime, stiaron "henry, canadian should contact: secretariat, can furnish material ,„,.... concerning the planning of the congroup travel. dniimited, inc. fprpnrp 1025 connecticut avenue, nwlelence. washington, d. c. 20036 telephone: (202) 659-9555 iassist euro£e canadian delegates should contact: the iassist conference to be ms. jan buchanan, manager held in conjunction with the interconvention services national sociological association p. lawson travel congress at uppsala, sweden, will suite 1415-2 take place on wednesday and thurscarlton street day, august 16-17, 1978. action toronto, ontario m5b 1 k2 group meetings are scheduled for the afternoons and the panels are a special mailing to lassist memscheduled for the evenings, bers is scheduled concerning the 8:30-11:00 p.m. conference. three panels are scheduled and include the following: "issues in comparative data and research," elorganizational reports liott avedon, department of recreation, university of waterloo, misist/social science information waterloo, canada, chair; "researcn problems associated with complex the following report was redata bases," john devries, departceived from fred riggs, a member of ment of sociology, carleton univerlassist and unisist, describing the sity, ottawa, ontario, canada kis november, 1977 pans meeting (see 5b6, chair: and "privary versus newsletter, 1, 4, 41-42, for furfreeedom of information," guido eker information concerning oni' ' ' ' -----itte--martinotti, archivio dati e prosist) of the ad hoc committee or grammi per le scienze sociali, via social science information. g. cantoni 4, 20144 milan, italy, chair. individuals interested m the purpose of the ad hoc cornpresenting a paper should contact mittee is to advise the unesco dithe appropriate chairperson. pavision and unisist on how best to pers may also be presented at ac— link the interests of the social tion group meetings. abstracts for science community with those of the action group papers shouia be subnatural science and engineering mitted to the appropriate action communites for whom unisist was group coordinator or regional secorijinaly designed. two north retariat, americans were invited to attend this meeting: professor jerome the european lassist conference clubb, executive director of the is being held in conjunction with inter-university consortium for pothe meeting of the international litical and social research (ann sociological assocation and in orarbor) and executive secretary of der to attend the world congress of the social science history associasociology the registration; and ms. sharon henry, execution/reservation form found elsetive director of the candaian where in this newsletter must be clearinghouse for social science completed and includ'e advance paydata and secretary for the canadian ment of registration fees. the fee chapter of iassist. before april 30, 1978, is $65 (u.s.) for isa members and $80 for non-members. fees cover access to all congress sessions and exhibits. regional database the printed congress program, the cartograpey printed book or abstracts of congress papers, and a list of partica number of european research ipants. workers active in the analysis of regional problems met in bergenall accommodations in uppsala norway, november 7-3, 1977, with will be reserved for isa particithe representatives of data servpants. therefore, in order to seices and centers of cartography to cure reservations, the accomodadiscuss the possibility for joint tions section of the reservation european action to link databases form must be completed. a copy of and facilities for computer mapthis form shouia also be sent to ping. the meeting was organized by your regional lassist secretariat. the norwegian social science data isa is not organizing any charters services, financed by the norwegian to uppsala but is leaving this to research council and sponsored by the national sociological 27 i^ssist newsletter, vol. 2, no. 1 (winter 1978) the social sciences committee of research. the training session is the european science foundation. directed toward graduate students and junior faculty. those interthere was broad agreement on ested in attending should apply begoals and objectives thought imporfore hay 1, 1978 to: rant by the group and the attendees agreed to seek support for the orbjorn henrichsen ganization of two pilot projects executive director during 1978-79. papers on the pinsd lot projects will be circulated bechr istiesgate 15-19 fore the next meeting of the group, n-5014 bergen-univ. possibly in the summer of 1978. bergen, norway educational and research ssrc (britain) visitinq fellowship,~ oppoittu'ritie^ uhzll micro data collectio n methods in the survey archive invites apeconomics plications to its visiting fellowship program for 1978-79, from soan introductory course in licial scientists interested in brary management of numerical maundertaking either substantive or chine readable data files designed methodological research based on to meet the interests and present the archive's holdings. while diand future needs of librarians, inrect monetary compensation is iimformation specialists, and social ited, computer resources and office scientists, will be held hay space are provided by the fellow30-june 16, 1978, at the university ship. applications, with a deadof wisconsin-madison. line of march 31, 1978, and a curriculum vitae should he addressed course objectives are to into: crease awareness and knowledge of machine readable data through expothe director sure to protessionals working with ssrc survey archive large data bases, data base manageuniversity of essex ment, social science research, and wivenhoe park data library and archive organizacolchester, essex tion and management; to instruct in the latest techniques tor dealing with this medium; and to provide practical experience within a real icpsr summer program data library and computer environment. the 1978 inter-university consortium for political and social advance registration is required research has announced its summer and should be completed prior to program. at least three sessions april 15, 1978. additional informshould be of special interest to ation and application forms may be lassist members: acquired from: 1. workshop on management, alice bobbin or al schubert library control and use data s computation center of computer readable inuu52 social sciance building formation. uw-hadison hadison, wi 53706 2. small computer system phone: (608) 262-7962 hardware and software. 3. database management for complex social and hisnorweqian social s cienc e data torical data. services in addition, there will be two spethe norwegian social science cial workshops dealing with crimidata services will organize a nai justice data, one of which will training course to test a set of be directed toward data processing teaching materials based on survey and data management problems in the data. the course will be held from field. for further information june 18 through the 24th, 1978. contact: the course will be based on a cross-national data package develsummer program oped within the program of the inicpsr ternational social science council, p. 0. box 1248 as well as a set of norwegian surann arbor, michigan 48106 vey data prepared within the natelephone: (313) 764-2570 tional program of electoral 28 lassist newsletter, vol. 2, no. 1 (winter 1978) position announcehenis programmer/analyst the roper center at yale university has acnounced an opening for a computer programmer/analyst. a degree is preferred but not required. the position requires a person with several years or experience. a resume along with a letter of intent should be sent as soon as possible to: donald r. deluca p. 0. box 1732, yale station yale university new haven, connecticut 06520 an equal opportunity employer. ^^ „!21!s,. registration^^ "•^r' form^y sociology^^^ ^ 9th world congraat ol sociology m'^. c/o heso congrom s«rvlce ..^"s.___ s-105 !4slool.holm,s..a.n p,„„ ^^ ^ ^„ '.'™' m, 1 lim. °"" '""" .„.„, ..»,., :™~,'"«':t'™ p','"™ ™zc°' t.ie. .,.,»„..««. «cco««odaiion t„,« sr s::r s™ .-««««.«, r"«"o'°.'»n""m.'oi h„,l,^,,.u00>.» ?" ™''.m "'-i't^* s,.o,m«o<,™..u„„» ««.,««,» »5,ot.».~ o.teof...,,.l d.ie of de,.„„.e o<,™.-,,...ct.~™.™. ..„„„.. "?.'",';';.'.,., »,„„„.„„. .,.„„,.. „„.,,„„, recistbation .,,.« jt.";™ srv./™ "„„, i™ ~™s" x™„~ ""'"" .s.™.™.. sk. 21s 5.. joo 5.. .05 ~0u ,.., ,», ,.„ «o..„~.„ sk. ,00 >. itl ., ,n r.«ot'™"' °" '°" s"",'™,. s., « .. ,so »:"";».,, s>. ,s .. wo '"'•'""»—»" s.. 8i >. wo ««,.d.»,.., .0,5.. ..00 ,.,„,„. .0«lfu,s.., m ^ww"^' "" '*°' e:.re:::r;rr" "'"•'-" »imli^=^^"= 0.,. ^ -"• 1 ^^~\\ ^^ j. u^^sist ^-—-^ membership application hake halllng address ikstltinional affiliatiom telephone nuhber membership fees for calendar year 1977 individual: regular sis student ss institutional (two individual meoibershipsl: s35 charter individual menbership (three years}siod institutional subscription: sjs paytwnt enclosed (amount) .assist, send payment to:malie check or money order payable e!e°ir/t"z:r'"""'" following action groups: data archive registry data orlianizfltion and management data archive developheni classification data acquisitiof: documentation process-produced data membership in the lassist includes a subscrip tipn to the lassisi newsletter and a bership^affords the opportunity t. participat i applied for membership in lassist dues paid (arount) date 29 1/32 miller, o’hanlon, & sanni-anibire (2024) adventures in data literacy: when the gap you were trying to identify turns out to be a chasm., iassist quarterly 49(3), pp. 1-32). https://doi.org.10.29173/iq1138 the creative commons-attribution-noncommercial license 4.0 international applies to all works published by iassist quarterly. authors will retain copyright of the work and full publishing rights. adventures in data visualization support assessment: when the gap you were trying to identify turns out to be a chasm meg milleri, grace o’hanlon,ii & hafizat sanni-anibireiii abstract in an era where post-secondary students are seen as digital natives and novel knowledge mobilization is becoming an expected part of scholarly discourse, this paper synthesizes insights from multiple surveys about this topic. this research was conducted in 2020 and 2022 with participants from programs across the university of manitoba (a canadian public research university of around 30,000 students). this paper aims to illuminate the campus landscape and assess library support and resources for research visualization; additionally, the authors also explore challenges and potential pathways for improvement. keywords data literacy, data visualization, gis, knowledge mobilization, academic libraries background academic researchers' traditional knowledge mobilization (km) activities include journal articles, conference presentations and reports. as technology has evolved, km practices are beginning to shift towards communicating with a more diverse audience beyond academic. novel communication methods are being proposed to do this (weller, 2011), with the idea of data storytelling emerging. dykes defines data storytelling as being "more than just creating visually-appealing data charts. data storytelling is a structured approach for communicating data insights, and it involves a combination of three key elements: data, visuals, and narrative" (dykes, 2016, para. 2). in the not-so-distant past, data visualization was considered a specialist field, requiring certification to use the tools. as technology has become more embedded in our daily lives, there is an expectation for researchers to integrate it into their professional practice (stevens, 2016; weller, 2011). based on the author’s experience at various institutions and conversations with peers, many institutions including the author’s own offer few training opportunities outside of computer science programs and often do not acknowledge these as emerging practices. the university of manitoba is a research institution which, as of 2024, had over 30,000 students, split between the smaller health sciences campus and main campus (at the south end of the city) located in winnipeg, manitoba, canada. geographic information systems (gis) and data visualization research support is an area of growth in academic libraries (chin roemer & kern, 2019; neville & crampsie, 2019; pagowsky & mcelroy, 2016; saba & shearer, 2018). the university of manitoba libraries defined providing gis and data visualization support in its strategic mandate. one of the ways to enact this mandate was with the newly created gis & data visualization librarian role, which the author, miller, https://doi.org.10.29173/iq1138 2/32 miller, o’hanlon, & sanni-anibire (2024) adventures in data literacy: when the gap you were trying to identify turns out to be a chasm., iassist quarterly 49(3), pp. 1-32). https://doi.org.10.29173/iq1138 was the first to occupy in 2019. in this role, she assists all faculty and graduate students at the university of manitoba in these functional areas. unlike many gis and data librarian roles at other institutions, the gis & data visualization librarian at um has no subject area or liaison responsibilities. duties are split between consultation work, teaching, planning and systems development. miller’s previous career as a gis professional provides an applied lens for how she navigates this role. to try and work efficiently and to be able to better direct her efforts, the author decided to formally identify gaps researchers were experiencing in their research visualization practices by conducting multiple surveys to assess the library's supports and resources in this area. context after having spent a year in their new position, miller commenced an exploratory project in collaboration with the subject librarian for earth and environmental resources (grace romund), seeking to identify the researchers using novel knowledge mobilization methods in their work, framing it around the emerging trend of 'data storytelling' (this study will be referred to as the data storytelling project in future references). a literature review revealed two significant themes: knowledge mobilization effectiveness in specific fields and data storytelling as an emerging trend, especially in journalism. while existing literature helps discuss the area within a specific context, it does not address the question of whether researchers generally adopt these new trends in communicating their research and how they do that. our study sought to identify novel research mobilization method use patterns and support structure needs in order to understand how libraries could support this user group better. while revealing some interesting trends (implication of online learning, desire to push beyond traditional knowledge mobilization methods, factors that influence selection of data visualization resources and more), and measurable outcomes, the results did not encompass every facet of the original question the authors sought to address. in 2022, miller with the assistance of sanni-anibire undertook a second study further exploring another area identified in the data storytelling project: learning support. this was proposed as a sister project and would echo the four subject area breakdowns used in the previous study and assess how miller’s implementation of github pages as a platform to host training materials worked for users. this platform was adopted as lockdowns from the covid-19 pandemic forced a move to online instruction. some studies exist in the literature where the aim is to evaluate pedagogical techniques in software instruction (al hashlamoun & daouk, 2020; rickles et al., 2017), identify the challenges of creating and using online learning objects (acosta et al., 2018; diaz, 2018), record user experiences from nontraditional programs adopting data visualization tools (harmon & gross, 2010; henshaw & meinke, 2018) or document the rise of open tools (neville & crampsie, 2019; pugachev, 2019). in terms of library instruction traditional library instruction is approached as a series of one-off sessions (chin roemer & kern, 2019). this model, however, conflicts with pedagogical approaches to software instruction suggested in the literature, where a scaffolded method using practical/local examples is accepted as the appropriate way to engage with the student (al hashlamoun & daouk, 2020; fouh et al., 2012; rickles et al., 2017). miller's approach to instruction is to structure materials as an academic curriculum with a valiant attempt to incorporate epistemology in support. their background in geographic information systems, where they became accustomed to using open tools such as github to keep materials as accessible, interoperable and as organized as possible informed their selection of github as a platform. the author https://doi.org.10.29173/iq1138 3/32 miller, o’hanlon, & sanni-anibire (2024) adventures in data literacy: when the gap you were trying to identify turns out to be a chasm., iassist quarterly 49(3), pp. 1-32). https://doi.org.10.29173/iq1138 has also incorporated some of this pedagogy into their practice as a librarian by implementing the workflows discussed by kaitlin newson (2017). the second study (referred to as the github assessment project in future references) sought to assess using github pages in data visualization instruction and to identify gaps in library support for novel knowledge mobilization (km) and examine them through the lens of library instruction grounded with the student voice. in both of these studies the author was trying to determine: • who is doing visualization work on campus? • what supports are lacking? • how can libraries help? methods data storytelling project this exploratory study sought to identify academic research translation patterns, methods and needs in novel knowledge mobilization support that libraries could adopt to better support this user group. to do this, the university of manitoba employees with an active research portfolio were recruited and asked to complete a web-survey asking questions about their demographics/user groups and practices (see appendix a). github pages assessment project this exploratory study sought to assess the effectiveness of github pages as a platform for data visualization instruction materials and if their content provided the support that university of manitoba learners required. workshop participants over the course of the spring 2022 semester were recruited to complete a web survey to answer questions around user values and needs (see appendix b). shared in both of these exploratory cases, surveys offered participants opportunities to answer structured and unstructured questions. participants were grouped under the following program categories: health sciences, science and technology, arts and humanities, and social sciences. grounded theory (charmaz, 2014) (discussed below) was then used to analyze the results to identify themes and trends. to achieve this outcome, data analysis was iterative; researchers considered the emerging patterns as they worked through the data cleaning process (charmaz, 2014). initial impressions of the data were noted while reading through the response. after that, the researchers coded the data in two phases: initial coding and focused coding. coding was carried out inductively, meaning that the researcher let the data determine the themes. using open coding, the response to each survey question was considered— which allowed the researcher to immerse themselves in the data (charmaz, 2014) -and assigned a word or phrase that best described the responses. these words and phrases served as the initial codes. during focused coding, researchers compared the initial codes with one another and grouped similar ideas together to form the themes. finally, these derived themes were compared against the entire data set to check that they captured the essence of the participant responses. https://doi.org.10.29173/iq1138 4/32 miller, o’hanlon, & sanni-anibire (2024) adventures in data literacy: when the gap you were trying to identify turns out to be a chasm., iassist quarterly 49(3), pp. 1-32). https://doi.org.10.29173/iq1138 findings data storytelling project this project's survey contained questions grouped into the following sections: demographics, data, software, knowledge mobilization, and storytelling. one hundred thirty-three people responded to the survey which was sent out to all (1264) campus members with an active research portfolio. 59% (n=79) of participants were faculty and librarians, 34% (n=45) were grad students and nine respondents (7%) were staff. the responses from faculty members were similar to those of students, with additional themes emerging from specific subject areas. figure 1 (below) depicts the breakdown of the study population by subject area clusters. figure 1: study population by subject area clusters regarding subject areas, the largest group of respondents were from science and technology (33%), with health sciences making up the second largest group. in terms of the types of data being collected, different subject clusters tended to different types of data. below are the subject clusters with the highest usage data type highlighted. arts & humanities n=26 science & tech n=44 health science n=29 social sciences n=18 qualitative 35% 16% 17% 17% both 58% 30% 34% 50% quantitative 7% 52% 48% 33% table 1 types of data collected by different subject clusters of the respondents who used supplementary data sets to augment their work, 9% of respondents said they always use the most current data possible, 33% expressed a preference for using mostly current data, 45% (the largest group) responded they use a mix of current and historical data, and finally 11% of respondents said they mostly use historical data to supplement their work. in terms of how study participants initially learned and stayed up to date with the software and methods used to visualize their data, some interesting trends were revealed. the majority of respondents arts and humanities, 20% health sciences, 22% libraries, 2% science and technology, 33% social sciences, 14% missing, 9% https://doi.org.10.29173/iq1138 5/32 miller, o’hanlon, & sanni-anibire (2024) adventures in data literacy: when the gap you were trying to identify turns out to be a chasm., iassist quarterly 49(3), pp. 1-32). https://doi.org.10.29173/iq1138 reported using formal training to receive initial training and shifting to self-learning for additional training. it is interesting to note that while students seem to depend more heavily on peers and work experience for ongoing learning (16% of respondents), only 4% of faculty followed this trend. figure 2: initial software data and training methods by participant group (percentages) figure 3: software and data training methods by participant group (percentages) most respondents indicated they consider their audience when they are creating their outputs. a higher percentage of faculty than students reported that they always or frequently write to the room. interestingly, 6% of faculty and 13% of student respondents indicated that they never considered the needs of their audience. 0 10 20 30 40 50 60 formal training mixed self-taught work experience/colleagues not applicable initial training students staff faculty 0 10 20 30 40 50 60 formal training mixed self-taught work experience/colleagues not applicable ongoing training students staff faculty https://doi.org.10.29173/iq1138 6/32 miller, o’hanlon, & sanni-anibire (2024) adventures in data literacy: when the gap you were trying to identify turns out to be a chasm., iassist quarterly 49(3), pp. 1-32). https://doi.org.10.29173/iq1138 figure 4: responses to the question: do you consider your audience when presenting your results? only 77 respondents (57%) answered the questions about knowledge mobilization aspirations. the overwhelming majority (62%) indicated that they were hoping to incorporate novel knowledge mobilization methods (i.e., social media, blogs, infographics, data visualization, apps) in the future. among the faculty in health sciences field, 85% expressed interest in pursuing novel km methods, while in science & technology field, 38% say they focused on traditional km strategies (e.g., writing for books or journals), and 42% said they use a mix of traditional and novel km methods. constant comparisons of participants' qualitative responses revealed these key themes: • desire to push beyond traditional km • need for support • tension with key terms used in the study 1. desire to push beyond traditional km: respondents believed that novel knowledge mobilization methods provided opportunities to share their work, with one researcher, noting it allows them to "communicate in a way that is interesting, but not oversimplified." a health sciences faculty member also stated that: "i find that the more i work to communicate our research to the general public and non-experts, the better i become at telling the story in academic (e.g., journal) formats, too, because i am constantly being forced to think, what does my research really mean, why is it relevant, at its core? i also feel that we have a social responsibility to communicate our results to the public, which has always supported our research financially in some ways. and it's our responsibility to find the right way to communicate with this audience." while many respondents wished to be able to engage with other audience than the academy, a faculty member from science & technology noted that: "the issue with mobilization with websites or social media or podcasts is that they are far less valued in the publish-or-perish model, aren't as valuable for promotion, tenure, or grant success. so, spending the time to learn how to use these alternative platforms and to put your research out there in these different formats never seems worth it it's just not valued in the academic markers for success." 0% 10% 20% 30% 40% 50% always frequently sometimes infrequently never audience consideration faculty staff students https://doi.org.10.29173/iq1138 7/32 miller, o’hanlon, & sanni-anibire (2024) adventures in data literacy: when the gap you were trying to identify turns out to be a chasm., iassist quarterly 49(3), pp. 1-32). https://doi.org.10.29173/iq1138 2. need for support participants from all faculty subject groups and ranks noted a need for support. one social sciences researcher noted that time was their major constraint by saying, "general help with km would be appreciated. it feels like researchers have to do more and more management, budgeting, analyzing, public outreach...it's a lot!". others noted that they felt creative outputs were not their strong suit, with a student noting, "i would need the help of an artist or another type of thinker to make these efforts more efficient and productive." another gap that was identified was the lack of skills and formal training opportunities, with a science student reporting "i wish i had better tools or knowledge/ability to make nice figures to visually support my story." a faculty member from the health sciences commented that: "data visualization, storytelling and effective communication will keep us relevant and help our research have a greater impact. it will also help it have greater uptake to our intended audience. some universities have entire credited courses on this for grad students. i'm not aware of this at the uofm and think it would be very useful." 3. tension with terms the biggest surprise for the authors was an underlying tension that a few researchers expressed regarding the terms 'data' and 'storytelling', and 'novel knowledge mobilization'. this did not seem to link back to participant ranks. some notable comments were: two different articulations of this tension were expressed in the science & technology group: 1. "[…] words have meaning, stop trying to make yourself sound better with flowery language. science is not the place for storytellers". 2. data is a very loaded word/concept to which interdisciplinary and humanities scholars have very different relationships than social sciences and scientific scholars have. from the arts & humanities group: "[i] see data as constructed and understanding as evolving and so, i don't want to signal a fixed interpretation, which i think the term data tends to convey" from the social sciences: "sounds like a nice way to say you will propagandize the results." from an unidentified subject grouping: "knowledge mobilization activities is a term that doesn't make sense outside of your group. i can't answer this because […] i don't even know what the hell you're asking" although these comments have been highlighted, the vast majority of study participants had a positive outlook on data/ data storytelling and novel scholarly communication methods as a way to share their work with others. github pages assessment project this project's survey contained questions grouped into the following sections: demographics/ learner profiles, resource preference, learning object structure assessment, and learning object content assessment. https://doi.org.10.29173/iq1138 8/32 miller, o’hanlon, & sanni-anibire (2024) adventures in data literacy: when the gap you were trying to identify turns out to be a chasm., iassist quarterly 49(3), pp. 1-32). https://doi.org.10.29173/iq1138 nineteen people responded to the survey which was sent out to the ninety-eight learners who had attended workshops on topics covering data visualization theory, gis, infographics, data cleaning, network visualization and dashboards. one person answered two questions only and was excluded from the analysis. 83% (n=15) of participants were graduate students, 11% were faculty members (n=2), and one respondent was a postdoctoral researcher (6%). the responses from faculty members were overall consistent with students' responses. the majority of respondents indicated no experience (figure 4) when asked about their perceived expertise levels of session content. figure 5: study population by levels of expertise in terms of subject areas (29%) were from environment, earth, and resources (traditional geography users), 4 (24%) were from health sciences; the other participants were spread across several disciplines including: arts, education, science, social work, engineering, agriculture and food sciences. in terms of general discussion about learning resources for gis and data visualization the following responses were shared by participants: when it came to how session participants preferred to learn a new visualization skill, two categories were highlighted either an in person or virtual class with an instructor present or via online step-bystep documentation (table 2). preferred learning method responses in-person/ virtual class 39% (n=7) online step-by-step documents 33% (n=6) instruction through an experienced mentor 11% (n=2) vendor training 6% (n=1) youtube videos 6% (n=1) table 2: preferred learning methods of study participants. 83% of respondents also indicated that they were likely to reuse/return to resources once discovered (always, n=7; frequently, n=8). 78% (n=14) respondents also noted that they were more likely to attend training if it was done using an open resource (now commonly referred to as oers). no experience 61% beginner 28% intermediate 11% expert 0 https://doi.org.10.29173/iq1138 9/32 miller, o’hanlon, & sanni-anibire (2024) adventures in data literacy: when the gap you were trying to identify turns out to be a chasm., iassist quarterly 49(3), pp. 1-32). https://doi.org.10.29173/iq1138 when asked what they were looking for when signing up/ attending sessions respondents identified that they were split between looking for introductory and advanced training (table 3) this makes sense considering that 89% of respondents self-identified as having no experience or beginner. respondent goals responses introductory training 41% (n=7) intermediate training 24% (n=4) advanced training 35% (n=6) table 3: respondent goals in signing up for a session a likert scale question was used to ask participants about what the author should focus on when creating new contentfrom deeper content on existing topics, to more basic content on additional topics/ software. learners identified that there was a strong preference for new materials to focus on building depth (scaffolded materials) (table 4). learner priority responses depth 17% (n=3) mostly depth 39% (n=7) breadth and depth 33% (n=6) mostly breadth 11% (n=2) breadth 0 table 4: learner priority breakdown for training materials finally, respondents were asked if they though their responses to the questionnaire would have been different if sessions had been in person and we had not had to switch to online learning because of pandemic lockdowns (table 5). respondent response responses would have been different 33% (n=6) might have been different 39% (n=7) would not have changed 22% (n=4) unsure 6% (n=1) table 5: impact of training format on responses some reasons participants thought their responses would have been different include that "i would not be able to pay attention in person but can take breaks as needed virtually," or that "in person learning is always different." some believed that the responses may have been different, responding that "it could be, but i believe that pandemic changed lifestyles, and we should get used to having virtual learning." constant comparisons of participants' qualitative responses revealed these key themes: • implications of the move to online learning resources on learner experience • factors that influence the selection/reuse of data visualization resources • pros and cons of learning objects • content versus skills https://doi.org.10.29173/iq1138 10/32 miller, o’hanlon, & sanni-anibire (2024) adventures in data literacy: when the gap you were trying to identify turns out to be a chasm., iassist quarterly 49(3), pp. 1-32). https://doi.org.10.29173/iq1138 1. implications of online learning: respondents believed that the availability of asynchronous online learning tools (github pages site) provided opportunities for enhancing their knowledge of technological tools, provided flexibility, and encouraged independence. for example, one noted, "i think it broadened my familiarity with new forms of technology [visualization tools] that i was familiar with but not using regularly, and now assist in teaching others to use (part of my work function)." online learning objects also provided more flexibility for learners as another respondent noted: "i've discovered that i learn a lot better when i have the flexibility to go for walks or multitask with mindless work. i have more time for thinking and therefore for problem solving now that i have more flexibility with my schedule." the availability of the github pages resource has also encouraged learners to "problem solv[e] on my own" leading people to take more responsibility for their learning. faculty noted that in general online learning encouraged them to seek different ways to communicate effectively with their students. on the administrative side, a respondent noted that: "the pandemic helped the university facilitate online course participation, which helped students take courses that they couldn't because of the time limitation." however, some respondents felt that online learning had disadvantages, including the lack of a sense of community that accompanied virtual learning and the challenges it poses for atypical learners. one respondent wrote "i have a sensory processing disorder…i prefer engaging with physical material and physical spaces". although the workshops took place virtually due to the university's move to online learning, it is important to consider the positive and negative implications virtual learning has on learners and their ability to engage with data visualization instruction. 2. factors that influence selection/reuse of data visualization resources: the availability and accessibility of learning resources, reviews, and user-friendly features were some of the factors that respondents considered when choosing data visualization resources in general. the availability of tutorials, troubleshooting resources, and positive feedback from previous users, including colleagues, influenced respondents' decision to choose a particular resource. regarding specific workshops delivered by the library, respondents liked that the library research visualization content were presented in a website with slides and pdf. one participant's comment echoes other participants' observations: "the website organizes pieces of it better. i liked the slides being incorporated into the page. i suspect this is more accessible." most participants did not prefer the workshop resource being presented as either a webpage or pdf but thought incorporating both was effective. respondents also thought that the data visualization instruction being open source and accessible was of great benefit as it allowed them to reinforce their learning by referencing the material multiple times. one person explained, "it is impossible to learn everything the first time. i always have to go back and read them again. the[m] being available is critical for me." the github pages site also served as a reference for inspiration for researchers (who did not attend workshops) seeking ideas to visualize research data. https://doi.org.10.29173/iq1138 11/32 miller, o’hanlon, & sanni-anibire (2024) adventures in data literacy: when the gap you were trying to identify turns out to be a chasm., iassist quarterly 49(3), pp. 1-32). https://doi.org.10.29173/iq1138 figure 6: screenshots of github pages training resource. left is a page including embedded slides, right is a page of an exercise walkthrough. 3. pros and cons of github pages site: although some found the information presented at the workshops to be basic (it was an introductory workshop), respondents were unanimous in their observation that the learning objects allowed for information to be presented clearly and gradually. respondents indicated that they found the information to be "useful and useable." they found the instructions easy to follow, and "helpful" and found that the github pages site was "intuitive" and had "accurate headings for navigation". some of the cons, however, included the look and feel of the resource, the lack of adequate practice exercises, the duration of most workshops (1 hour was too short) and the online mode of delivery. respondents suggested the "infusion of colourful elements" and a more attractive interface. respondents also thought that having more practice exercises, access to more workshop data, and increasing the workshop duration would be beneficial. finally, some participants expressed a preference for in-person workshops. 4. content vs skills: respondents thought that learning the skills for data visualization was most important regardless of whether local (geographically) or field-specific data was used. some expressed that using local data was appropriate except when the data being used influenced the kind of visualization skill that could be learned. one respondent explained that: "the subject matter is not necessarily relevant, however the subject often influences the type of data available or being used. this in turn will affect the examples and methods being shown. therefore, data typical in my field of study will be more useful than an example with data that is not." although they found the workshops helpful, many respondents expressed the need for data visualization instruction using data specific to their field of study/research area, including qualitative data visualization. respondents were also given the opportunity to suggest improvements to the learning objects. feedback here could be grouped into four major categories: https://doi.org.10.29173/iq1138 12/32 miller, o’hanlon, & sanni-anibire (2024) adventures in data literacy: when the gap you were trying to identify turns out to be a chasm., iassist quarterly 49(3), pp. 1-32). https://doi.org.10.29173/iq1138 • communication • multi-modal delivery • using relatable data • access to resources respondents thought the github pages site was helpful and valuable and suggested better communication about the availability of workshops to students and faculty. they suggested contacting students through listservs and limiting the number of attendees to facilitate deeper engagement. also, while respondents were pleased with the workshop delivery, they suggested that "a combo of web, power point, and in-person packages will deliver more." many respondents wanted more in-person workshops and opportunities to follow up with data visualization experts later (this service is already available via the library future training may emphasize the availability of additional or follow-up supports). respondents suggested that the workshop data be discipline/field-specific to make the workshop more relatable and applicable to their research work. respondents also requested workshops on data visualization of qualitative data as quantitative data is "often really challenging to do in a clear and concise way" finally, respondents suggested that data used in the workshops should be available to participants for future reference and to reinforce their learning. discussion this paper was born out of one of the author’s starting a new job (a newly created role for the institution) and wondering where they could fit into the campus data literacy landscape. the initial study was conceived to answer the question of where researchers were in their knowledge mobilization practices and what type of support they were seeking. during conversations with faculty and students, they expressed frustration with the lack of support for digital communication tools on campus. the author, prioritized creating supports for campus users by working with their internal learning and instruction team. this included consultation hours, lab drop-ins, workshops integrated into graduate studies programming, class integration, and embedding into lab groups. on top of this forward-facing work, they also took on roles in systems development and software and data license management to build a base in libraries. the outcome of both of these surveys was that key user groups and needs were identified, and the critical connection between learner and content creator identified by ain et al. (2016) would be created. below are the current trends in data literacy at um. users users from all program groupings (health sciences, science & technology, social sciences and arts & humanities) are interested in using digital tools to improve their communication. they recognize that different audiences will be able to understand the information shared by using innovative communication methods. however, many of our research participants, including faculty, staff or students, experience barriers to accomplishing this due to lack of tools, training and time. https://doi.org.10.29173/iq1138 13/32 miller, o’hanlon, & sanni-anibire (2024) adventures in data literacy: when the gap you were trying to identify turns out to be a chasm., iassist quarterly 49(3), pp. 1-32). https://doi.org.10.29173/iq1138 ongoing training in these tools is usually self-directed, with original exposure in a class, workshop put on by libraries, or other formal training environment. users expressed interest in using sample data where they can 'see themselves'. making online training materials available allowed users to work through them on their own time and at their own pace outside formal training sessions. learners also expressed frustration at not knowing what type of support was available on campus in terms of software, data, and services. many researchers whom the author had met via consultation bookings assumed they were alone in their tool use or lack of training. users with intermediate levels of expertise in tools, often had experience for a previous workplace or educational institution. librarian implemented supports in response to this feedback, the author adjusted the supports they were providing, they will be discussed below. resources: 1. creation of libguides gis & geovisualization, and data visualization provide listings of data sets that researches could integrate into their work, as well as links to different tools and resources with notes to what is supported by the institution. these were in the top three used subject guides at the university of manitoba in the last calendar year. 2. integrated gis analysis environment (gis hub) researchers did not know which visualization tools were available and found the process to access the gis tools too convoluted to bother with. the author streamlined the process of esri (gis software for mapping and spatial analytics) license management on campus by implementing saml authentication and developing an enterprise instance that integrates with arcgis online, desktop and mobile applications. this authentication integration, built trust with the central it department as well as other researchers on campus. the system also acts as an open and proprietary data repository highlighting authoritative data that users can use in their work. managing users as named accounts allows um to restrict access to active university users easily. from the implementation date to the time of writing, the population of researchers using the gis hub has grown from less than one hundred to over nine hundred active users. 3. workshop and teaching content hosted in github pages the move to online learning during the pandemic lockdown provided impetus for the author to create an open online repository of their training materials integrating slides, walk-throughs, and in some cases, recordings. while it took a lot of work to create this repository, the feedback overall has been very positive. drupal (used by um) would be an alternative platform that provides a more polished experience, but the author values the ability to host and share content openly in github. programming survey responses indicated a level of disconnect among users which was previously unknown to the author. to resolve this tension, i changed how i talk about and describe the data visualization support i offered. https://doi.org.10.29173/iq1138 https://www.esri.com/en-us/home 14/32 miller, o’hanlon, & sanni-anibire (2024) adventures in data literacy: when the gap you were trying to identify turns out to be a chasm., iassist quarterly 49(3), pp. 1-32). https://doi.org.10.29173/iq1138 1. teaching and workshops hearing how desired skill sets were less tied to specific programs and tension around certain terms prompted workshops to be rebranded. "data dashboards for beginners" shifted to "powerbi: a gentle introduction" and most recently, "integrating word-clouds into research" has proven more popular than its predecessor, "data cleaning for with openrefine," which is the same session with the same description. this approach is also taken when describing how to use a tool during workshops. for example, if a menu item is labelled as 'data', then describing its intended use to the users is reframed to "this is where you will click to find the file you want to visualize" instead of "this is where you import your data." 2. gis days the previous map librarian ran a day of gis programming, bringing in speakers from local governments and organizations and highlighting large projects with a gis component at the institution. as author's role is broader, they took a different approach to moving this initiative forward. instead, um is partnering with western university to offer a week of researcher-focused virtual programming from presenters at all stages of gis expertise. in 2023, two in-person panels were also run, one of students and the other of faculty from different areas, to discuss their experiences in gis at um. this was very well received and prompted the resurrection of the data-viz drop-in. 3. data viz drop-in um is quite a siloed institution, internally and externally. while there are pockets of interdisciplinary work going on, there are not many opportunities for researchers to come together across departments and learn from one another. the author had run a drop-in session out of a library lab for a short time before the lockdowns and did not prioritize it once the campus opened back up; however, one of the students articulated its value during a gis day panel. the drop-in session was restarted in the winter semester (2024) as it gives students from different programs an opportunity to become aware of peers with similar interests and have the opportunity to follow up with one another and share expertise. this peer-assisted learning method has been identified in the literature (al hashlamoun & daouk, 2020; harding & engelbrecht, 2015) as an effective way for student learners to build their skills. a lack in digital tool support across campus is not a gap one person in libraries can or should be expected to fill. all this information gathered has allowed the author to better articulate the overwhelming demand and maxed-out capacity narrative driving their professional life for the past four years. in consulting with their unit coordinator, new boundaries have been laid, and many of the conversations and meetings with faculty have shifted to one of hiring priorities. during faculty meetings, the author advocates that if a program promotes these novel knowledge mobilization techniques within the data visualization sphere to students, academic departments cannot depend on libraries alone to troubleshoot, advise and teach – they need to hire their own experts. and slowly it is beginning to happen in the last six months, new professors in agriculture and architecture have been hired, both of whom have expertise in the digital realms of their field, and teaching release time has been provided to a linguistics department member to build a new course on gis methods. the author writes letters of support for researchers applying for grants to hire ras with the required technical skills and helps them craft these job descriptions. while this approach results in fewer reference stats being recorded by the librarian, it feels much more sustainable and has started a broader conversation on campus that dovetails with other data service questions, especially those related to research data management. https://doi.org.10.29173/iq1138 15/32 miller, o’hanlon, & sanni-anibire (2024) adventures in data literacy: when the gap you were trying to identify turns out to be a chasm., iassist quarterly 49(3), pp. 1-32). https://doi.org.10.29173/iq1138 future while these findings are specific to the author's institution, they offer questions a data services provider could ask themselves. discussion of data visualization as an emergent trend has been the focus in the literature, and this instead focuses on what is being done to support users in the library and will advance the practice of engaging with and instructing this group of users. references acosta, m. l., sisley, a., ross, j., brailsford, i., bhargava, a., jacobs, r., & anstice, n. (2018). student acceptance of e-learning methods in the laboratory class in optometry. plos one, 13(12), e0209004. https://doi.org/10.1371/journal.pone.0209004 ain, q., aslam, m., muhammad, s., awan, s., pervez, m. t., naveed, n., basit, a., & qadri, s. (2016). a technique to increase the usability of e-learning websites. pakistan journal of science, 68(2), 164-169. al hashlamoun, n., & daouk, l. (2020). information technology teachers’ perceptions of the benefits and efficacy of using online communities of practice when teaching computer skills classes. education and information technologies, 25(6), 5753–5770. https://doi.org/10.1007/s10639020-10242-z burton, m., & lyon, l. (2017). data science in libraries. bulletin of the american society for information science and technology, 43(4), 33-35. https://doi.org/10.1002/bul2.2017.1720430409 charmaz, k. (2014). constructing grounded theory (2nd ed.). sage. chin roemer, r., & kern, v. (eds.). (2019). the culture of digital scholarship in academic libraries american library association. diaz, c. (2018, june 7). jekyll and institutional repositories. northwestern university research and data repository. https://arch.library.northwestern.edu/concern/generic_works/6q182k274 dykes, b. (2016). data storytelling: the essential data science skill everyone needs. forbes. https://www.forbes.com/sites/brentdykes/2016/03/31/data-storytelling-the-essential-datascience-skill-everyone-needs/ fouh, e., akbar, m., & shaffer, c. a. (2012). the role of visualization in computer science education. computers in the schools, 29(1–2), 95–117. https://doi.org/10.1080/07380569.2012.651422 harding, a., & engelbrecht, j. (2015). personal learning network clusters: a comparison between mathematics and computer science students. educational technology & society, 18(3), 173–184. https://www.jstor.org/stable/jeductechsoci.18.3.173 harmon, j. e., & gross, a. g. (2010). the craft of scientific communication. university of chicago press. henshaw, a. l., & meinke, s. r. (2018). data analysis and data visualization as active learning in political science. journal of political science education, 14(4), 423–439. https://doi.org/10.1080/15512169.2017.1419875 herther, n. k. (2019). library carpentry: a toolkit for researchers. information today, 36(3), 16–18. neville, t., & crampsie, c. (2019). from journal selection to open access: practices among academic librarian scholars. portal: libraries and the academy, 19(4), 591–613. https://doi.org/10.1353/pla.2019.0037 https://doi.org.10.29173/iq1138 https://doi.org/10.1371/journal.pone.0209004 https://doi.org/10.1007/s10639-020-10242-z https://doi.org/10.1007/s10639-020-10242-z https://doi.org/10.1002/bul2.2017.1720430409 https://arch.library.northwestern.edu/concern/generic_works/6q182k274 https://www.forbes.com/sites/brentdykes/2016/03/31/data-storytelling-the-essential-data-science-skill-everyone-needs/ https://www.forbes.com/sites/brentdykes/2016/03/31/data-storytelling-the-essential-data-science-skill-everyone-needs/ https://doi.org/10.1080/07380569.2012.651422 https://www.jstor.org/stable/jeductechsoci.18.3.173 https://doi.org/10.1080/15512169.2017.1419875 https://doi.org/10.1353/pla.2019.0037 16/32 miller, o’hanlon, & sanni-anibire (2024) adventures in data literacy: when the gap you were trying to identify turns out to be a chasm., iassist quarterly 49(3), pp. 1-32). https://doi.org.10.29173/iq1138 newson, k. (2017). tools and workflows for collaborating on static website projects. the code4lib journal, 38. https://journal.code4lib.org/articles/12779 pagowsky, n., & mcelroy, k. (2016). critical library pedagogy handbook: essays and workbook activities association of college and research libraries. pugachev, s. (2019). what are “the carpentries” and what are they doing in the library? portal: libraries and the academy, 19(2), 209–214. https://doi.org/10.1353/pla.2019.0011 rickles, p., ellul, c., & haklay, m. (2017). a suggested framework and guidelines for learning gis in interdisciplinary research. geo: geography and environment, 4(2), e00046. https://doi.org/10.1002/geo2.46 saba, f., & shearer, r. l. (2018). transactional distance and adaptive learning: planning for the future of higher education. routledge. https://doi.org/10.4324/9780203731819 stevens, h. (2016). [review of the book big data, little data, no data: scholarship in the networked world, by christine l. borgman]. technology and culture, 57(3), 706–708. https://doi.org/10.1353/tech.2016.0099 weller, m. (2011). the digital scholar: how technology is transforming scholarly practice. bloomsbury academic. https://doi.org/10.5040/9781849666275 https://doi.org.10.29173/iq1138 https://journal.code4lib.org/articles/12779 https://doi.org/10.1353/pla.2019.0011 https://doi.org/10.1002/geo2.46 https://doi.org/10.4324/9780203731819 https://doi.org/10.1353/tech.2016.0099 https://doi.org/10.5040/9781849666275 17/32 miller, o’hanlon, & sanni-anibire (2024) adventures in data literacy: when the gap you were trying to identify turns out to be a chasm., iassist quarterly 49(3), pp. 1-32). https://doi.org.10.29173/iq1138 appendix a: data storytelling survey consent form: user clicks “i consent” at bottom of letter and the following survey opens: page 1: demographics: the following section looks at who you are within the institution: type question option: 1 dropdown please select your primary faculty: • faculty of agricultural and food sciences • faculty of architecture • school of art • faculty of arts • h. asper school of business • faculty of education • price faculty of engineering • clayton h. riddell faculty of environment, earth and resources • extended education • faculty of graduate studies • libraries • rady faculty of health sciences • school of dental hygiene • dr. gerald niznick college of dentistry • max rady college of medicine • college of nursing • college of pharmacy • college of rehabilitation sciences • faculty of kinesiology and recreation management • faculty of law • desautels faculty of music • faculty of science • faculty of social work • university 1 • university administrative units https://doi.org.10.29173/iq1138 18/32 miller, o’hanlon, & sanni-anibire (2024) adventures in data literacy: when the gap you were trying to identify turns out to be a chasm., iassist quarterly 49(3), pp. 1-32). https://doi.org.10.29173/iq1138 2 radio buttons please select your position: radio buttons for question 2 asking for participant’s position on campus. options include: instructor i, insturctor ii, senior instructor, lecturer, assistant professor, associate professor, professor, archivist, general librarian, assistant librarian, associate librarian, lbirarian, phd candidate, master’s cadidate, researcher other. 3 radio buttons + text line are you cross appointed? if yes, with what other faculty? yes/ no + line that appears if ‘yes’ is selected 4 radio buttons + text line do you have professional affiliation with any other centres/institutes? if yes, with what centre/institute? yes/ no + line that appears if ‘yes’ is selected page 2: data: in thinking about your research data type question option: 5 check boxes what type of data do you typically collect? check boxes for always qualitative, mostly qualitative, qualitative and quantitative, mostly qualitative, always quantitative. 6 multi-line open text are there secondary data sets that you regularly use to supplement the data you collect (crop inventory, census etc.)? please include the dataset name and producer. https://doi.org.10.29173/iq1138 19/32 miller, o’hanlon, & sanni-anibire (2024) adventures in data literacy: when the gap you were trying to identify turns out to be a chasm., iassist quarterly 49(3), pp. 1-32). https://doi.org.10.29173/iq1138 7 check boxes how old are the data sets you regularly use to supplement your data? check boxes for age of data sets: historical, mostly historical, current and historical, mostly current, always most current available. 8 ranking looking at the data analytics lifecycle below, rank from most to least how much time you spend in each section of the cycle. randomly sorted list of elements from the data analytics lifecycle: maintain, process, analyze, communicate, capture. https://doi.org.10.29173/iq1138 20/32 miller, o’hanlon, & sanni-anibire (2024) adventures in data literacy: when the gap you were trying to identify turns out to be a chasm., iassist quarterly 49(3), pp. 1-32). https://doi.org.10.29173/iq1138 page 3: software: in thinking about the software that you/your lab group use for your analysis type question option: 9 multi-line open text what software or coding language(s) do you use for analysis? 10 multi-line open text where did you initially learn to use those tools? 11 multi-line open text where do you get your ongoing training and/or support for using these tools? 12 radio buttons + text line does the software you use significantly alter the format of your original data? (eg: csv file to cartographic output) if yes, how? yes/ no + line that appears if ‘yes’ is selected 13 radio buttons + text line do you use different software than you use for analysis to create your data visualizations? if yes, what software and why? yes/ no + line that appears if ‘yes’ is selected page 4: knowledge mobilization sshrc defines knowledge mobilization as "moving knowledge into active service for the broadest possible common good." type question option: 14 check boxes + text line what knowledge mobilization activities do you do? check boxes for: conference posters, conference presentations, journal articles, trade publication articles, podcast, blog/social media, professional website, and other. *randomly sorted each time, if ‘other’ is selected , text input opens for user to type https://doi.org.10.29173/iq1138 21/32 miller, o’hanlon, & sanni-anibire (2024) adventures in data literacy: when the gap you were trying to identify turns out to be a chasm., iassist quarterly 49(3), pp. 1-32). https://doi.org.10.29173/iq1138 15 check boxes + text line who is the audience for these knowledge mobilization activities? check boxes for: industry, students, government agencies, academics across various fields, general public, academics in your field, other. randomly sorted each time, if ‘other’ is selected , text input opens for user to type 16 check boxes do you create different data products for different audiences? check boxes for always, frequently, sometimes, infrequently, never 17 multi-line open text are there knowledge mobilization activities that you do not currently do, but you have plans to or would like to in the future? if yes, please describe. page 5: data storytelling: keeping dyck's definition of data storytelling in mind ("a structured approach for communicating data insights, and it involves a combination of three key elements: data, visuals, and narrative."); answer the following section for yourself. type question option: 18 multi-line open text what is data storytelling to you? 19 multi-line open text do you consider yourself a data storyteller? if yes, why? 20 multi-line open text do you have any additional comments about data storytelling? submit https://doi.org.10.29173/iq1138 22/32 miller, o’hanlon, & sanni-anibire (2024) adventures in data literacy: when the gap you were trying to identify turns out to be a chasm., iassist quarterly 49(3), pp. 1-32). https://doi.org.10.29173/iq1138 page 6: exit + option to submit contact info to be part of focus groups in a future study (stored in separate form) https://doi.org.10.29173/iq1138 23/32 miller, o’hanlon, & sanni-anibire (2024) adventures in data literacy: when the gap you were trying to identify turns out to be a chasm., iassist quarterly 49(3), pp. 1-32). https://doi.org.10.29173/iq1138 appendix b: github pages survey https://doi.org.10.29173/iq1138 24/32 miller, o’hanlon, & sanni-anibire (2024) adventures in data literacy: when the gap you were trying to identify turns out to be a chasm., iassist quarterly 49(3), pp. 1-32). https://doi.org.10.29173/iq1138 https://doi.org.10.29173/iq1138 25/32 miller, o’hanlon, & sanni-anibire (2024) adventures in data literacy: when the gap you were trying to identify turns out to be a chasm., iassist quarterly 49(3), pp. 1-32). https://doi.org.10.29173/iq1138 https://doi.org.10.29173/iq1138 26/32 miller, o’hanlon, & sanni-anibire (2024) adventures in data literacy: when the gap you were trying to identify turns out to be a chasm., iassist quarterly 49(3), pp. 1-32). https://doi.org.10.29173/iq1138 https://doi.org.10.29173/iq1138 27/32 miller, o’hanlon, & sanni-anibire (2024) adventures in data literacy: when the gap you were trying to identify turns out to be a chasm., iassist quarterly 49(3), pp. 1-32). https://doi.org.10.29173/iq1138 https://doi.org.10.29173/iq1138 28/32 miller, o’hanlon, & sanni-anibire (2024) adventures in data literacy: when the gap you were trying to identify turns out to be a chasm., iassist quarterly 49(3), pp. 1-32). https://doi.org.10.29173/iq1138 https://doi.org.10.29173/iq1138 29/32 miller, o’hanlon, & sanni-anibire (2024) adventures in data literacy: when the gap you were trying to identify turns out to be a chasm., iassist quarterly 49(3), pp. 1-32). https://doi.org.10.29173/iq1138 https://doi.org.10.29173/iq1138 30/32 miller, o’hanlon, & sanni-anibire (2024) adventures in data literacy: when the gap you were trying to identify turns out to be a chasm., iassist quarterly 49(3), pp. 1-32). https://doi.org.10.29173/iq1138 https://doi.org.10.29173/iq1138 31/32 miller, o’hanlon, & sanni-anibire (2024) adventures in data literacy: when the gap you were trying to identify turns out to be a chasm., iassist quarterly 49(3), pp. 1-32). https://doi.org.10.29173/iq1138 https://doi.org.10.29173/iq1138 32/32 miller, o’hanlon, & sanni-anibire (2024) adventures in data literacy: when the gap you were trying to identify turns out to be a chasm., iassist quarterly 49(3), pp. 1-32). https://doi.org.10.29173/iq1138 end notes 1 meg miller is the gis & research visualization librarian at the university of manitoba, she can be reached by email at meg.miller@umanitoba.ca ii grace o’hanlon is an associate librarian at the university of manitoba libraries. iii hafizat sanni-anibire is a second-year phd student in the faculty of education at the university of manitoba. https://doi.org.10.29173/iq1138 mailto:meg.miller@umanitoba.ca gesis 18 iassist quarterly winter 2011 iassist quarterly abstract what do researchers need from archives? what do archives need from researchers? these questions cover two types of researchers that encounter data archives: those who create the data (data creators) and those who re-use it (data re-users). these groups have different needs and archives mediate between them. the role of an archive for creators and re-users is to support them in producing quality data, metadata and documentation and to facilitate wide and multipurpose data dissemination. by supporting multipurpose reuse, to the fullest extent possible, archives help realize the value of public investment in academic research. this paper discusses the optimization of research data management training and support for research data creators, and data dissemination and long-term preservation for social science data archives. it outlines the gesis plan to create a research data management and archive training centre for the european research area, to cater to both data supply and data demand. the training centre will look to ensure excellence in the creation and long-term preservation of reusable data in the european research area, contribute to promoting and to the adoption of standards in research data management, and promote data availability and reuse. finally, the centre will provide and coordinate training on technologies and tools used by data professionals. keywords: : archives, research data management, incentives, sharing, training. introduction2 social science data archives connect two primary audiences. one is data creators––those who bring social science data into being. in this category, we place principal investigators of studies as well as researchers who work in data collection procedures. the other audience is data re-users. here we mean researchers who either use data they themselves created some time ago or use data created by others to examine social phenomena. the ligaments connecting these audiences are data archives: organizations that facilitate data ingest and dissemination. by accepting data into their catalogue for preservation and reuse, then furnishing the research community with that data, the archives establish a connection between the two audiences. however, it is a dynamic relationship fashioned by two forces: a movement towards data sharing for reuse and a set of resistances to data reuse. in this paper, we discuss these forces and we highlight actions to promote data sharing and reuse. the basis of our perspective is a supply and demand model of data archives and thus the basis of our proposals are for both audiences. we focus on attempts to introduce practical policy suggestions to facilitate an easier relationship between creators, archives, and re-users primarily within the cessda-eric consortium of european social science data archives. the data sharing movement the contemporary movement towards data sharing for reuse is a trend enabled and assisted by technological innovation. the means by which one can share data and collaborate on research have become cheaper and easier to utilize. negating the barriers towards reuse and collaboration posed by time, distance, cost, and logistics are developments in instantaneous means of communication, large capacity data transfer, cheaper digital storage costs, and the power of data analysis software packages. today we can do more research with more data in less time and at less cost. indeed the range, scope and potential applications of data created, available, and analyzed can reach such a size that it may even purposing your survey: archives as a market regulator, or how can archives connect supply and demand? by laurence horton, alexia katsanidou1 iassist quarterly winter 2011 19 iassist quarterly challenge the primacy of the experimental hypothesis approach in doing social research (anderson, 2008). in recognition of these phenomena, the european commission commissioned a report on how to best direct this changing data environment towards scientific and economic innovation. its high level expert group on scientific data envisioned …a scientific e-infrastructure that supports seamless access, use, re-use, and trust of data. in a sense [...] the data themselves become the infrastructure – a valuable asset, on which science, technology, the economy and society can advance (european union, 2010 p.4) the belief that technology is changing patterns of research and publications has a normative basis in the argument that publicly funded data is a public good and that funders can maximize the value of research they support with a requirement that data be shared to the fullest extent possible. this argument is based on the position of the organization for economic co-operation and development (oecd) that publically funded research data should as far as possible be openly available to the research community for re-analysis, repurposing, and long-term preservation (oecd, 2007). the riding the wave report (european union, 2010) echoed an expectation of transparency in data creation. an expectation that the methods of generating and manipulating data be clear so data is comprehensible to others outside of, and remains comprehensible to as time passes, the original data creators themselves. in addition, there is an acceptance as the norm in good scientific research that findings be based on data that is available (where legally and ethically possible) for independent verification, analysis, and reuse. this is a movement accelerated by a requirement of some academic journals that publication of articles is dependent on the authors’ making available the underlying data if it is not already accessible. we find an example of this trend in dryad. dryad is an open data repository for articles published in the natural sciences and lists a number of journals as partners for which it either holds, or works with, to preserve and disseminate data (dryad, 2011). an additional example is european data watch explained (edawax) (european data watch extended, 2011) this german-based project examines the absence of incentives in economics for the replication of results and data reuse with the intention of creating a publication data archive. a similar project for political science, but with narrower focus is the gesis data infrastructure team’s data policy availability project. this project empirically investigates data policies of all top academic journals in political science, analyses their content and finally proposes policy guidelines. in an era of tight pressures on public spending, the political attraction of these arguments is clear. the european commission has committed itself to an open data policy that it estimates would provide an extra €40 billion a year to the eu economy. “taxpayers have already paid for this information, the least we can do is give it back to those who want to use it in new ways...” stated commission vice president neelie kroes. “your data is worth more if you give it away” (european commission, 2011a) she added. however, the ec policy is tied to public sector data, not publicly funded academic research data which remains exempt (european commission, 2011b). yet this too can be, and is, considered a public investment to be shared thereby maximizing its value. we find examples of this belief in the emergence of policies that mandate data sharing be addressed as an aspect of proposals seeking public funding. the united states national institutes of health (nih) enforced a data sharing policy in 2003, with a requirement for funding applications to include a plan for data sharing (national institutes of health, 2003). the national science foundation (nsf) followed in early 2011 by adopting a similar requirement to produce a data management plan for sharing (national science foundation, 2011). in the american environment it is often institutions that provide a preservation and dissemination service. examples include university of california-san diego (2010), university of illinois at urbana-champaign (2005), cornell university (2005), massachusetts institute of technology (2005), and university of rochester (2008). however, these approaches have been institution-specific rather than national infrastructure tools as nih and nsf aside, the united states lacks the regional, national and supranational level funding regime of european countries such as the united kingdom and germany. similar developments have occurred in europe. in may 2011, research councils uk––the strategic partnership agency of the united kingdom’s seven main research councils––published a set of common principles on data policy intended to provide an overarching framework for individual council policies on data reuse. the principals include an explicit statement that: publicly funded research data are a public good, produced in the public interest, which should be made openly available with as few restrictions as possible in a timely and responsible manner that does not harm intellectual property. (research councils uk, 2011) uk councils may vary in the specifics of data, but this principal holds across the field. the economic and social research council (esrc), natural environment research council (nerc) and the british academy all mandate research data be offered to data centers. in the case of the esrc (economic and social data service, 2011) and nerc (natural environment research council, 2011), through council funded data centers. other uk funders expect or encourage data sharing but do not mandate places of deposit. the engineering and physical sciences research council (epsrc) has introduced a policy (from may 2015) mandating that institutions ensure well documented data is preserved and available for a minimum of 10 years from last request for access by a third party (engineering and physical sciences research council, 2011). from an institutional perspective, university of edinburgh, followed by the university of hertfordshire (2011), became the first uk universities to adopt an institutional research data management policy. this included, in edinburgh’s case, a commitment that: research data management plans must ensure that research data are available for access and re-use where appropriate and under appropriate safeguards (university of edinburgh, 2010). in germany, the main publically funded research organizations have adopted a set of principles for the handling of research data. this 2010 agreement does not take as strong a tone as its rcuk equivalent; however, it does support long-term preservation and the “principle” of open access to research data, as well as the development of subject-specific requirements, standards, and metadata to facilitate interdisciplinary research and supporting infrastructure (alliance of german science organisations, 2010). these principals drew, in part, from an earlier set of proposals submitted by the german research council (dfg) that encourage researchers to take into account data management issues. reinforcement of this invitation is by guidelines promoting data sharing for experts on review panels. 20 iassist quarterly winter 2011 iassist quarterly the dfg raise the issue of data management and demand secure preservation and visibility for those data publically funded and used for publications, but limit this demand to a ten-year period (deutsche forschungsgemeinschaft, 1998). since then, greater effort has occurred to promote effective and consistent data management but not explicitly formulated in an official publication of the german research council. thus, the causes of a movement towards data preservation and sharing are clear: technology and financial benefit. furthermore, the demand is there. two of the largest data archives, the uk data archive (ukda) as part of the economic and social data service (esds) and in the united states the inter-university consortium for social and political research (icpsr) (2011a), have both seen significant increases in orders for data they hold since offering online access to data (economic and social data service, various). a similar phenomenon is apparent in the gesis leibniz-institute for social science’s user statistics––specifically for eurobarometer data, for which the number of datasets distributed has jumped between 2005 and 2009 (gesis leibniz institute for the social sciences, 2010). resistances to data reuse however, let us look at the supply side in the social sciences. here there are still obstacles that prevent data sharing. primary limitations are those placed by law and ethics. neither data archives nor funding agencies believe in sharing all data with everyone, or even within the academic community. the policies and recommendations presented above recognize, as we do, that there has to be protection of intellectual property, professional credit, and critically––moral and ethical protection of research participants. however, alongside these recognized limitations there are additional resistances to data sharing. opposition remains to the idea of sharing research data. this phenomenon in the social sciences can draw on a range of arguments. low-level (researcher-level) ignorance as to why others would want to use their data. this was a reason cited by a small number of researchers interviewed for the ukda’s data management planning for esrc centres and programmes (uk data archive, 2010 pp. 17-21). it is not resistances to data reuse itself, but an inability to imagine that the type of data generated would be of interest to anyone else. we can overcome this problem through more interaction within the scientific community and open presentation of opportunities for data sharing. additional to the ignorance of researchers about potential reuse of their data, there are also epistemological concerns. these cover congruence, reflexivity, and context. essentially, data creators holding this objection claim understanding and value of data can only exist in the specific context of their creation. they are concerned that their data, abstracted from the methodologies and ontologies adopted at the time of creation cannot adapt into a different research project. these problems of course need proper consideration particularly where the reflexive relationship between researcher and participant is critical to understanding the data, but given appropriate documentation, they should not prevent future reuse3 a clear problem is the lack of incentives to share data. as long as the main metric of career progression remains publications and citations of publications, data sharing will be a secondary concern. however, data creation requires the investment of a lot of scientific effort and expertise. reusing an existing dataset builds on the scientific work of other researchers who should be not only acknowledged, but also credited for their achievements. widespread recognition and implementation of a system for acknowledgement of data citations as an indication of research quality and establishing them as equivalent to publication citations would remove a reservation against data sharing. data creators often have concerns as to the ethics of reuse concerning research participants. specifically, a concern of compromised anonymity and confidentiality of participants emerges when disseminating data to other researchers. there are ways to anonymize data but some data are extremely sensitive and easily trackable. thus, researchers can be reluctant to share on principle of protecting their participants’ anonymity. we propose that the character and structure of the current social science research environment determines attitudes to reuse. outside of large-scale surveys, the concept of data reuse is not dispositional. there is still no established culture of archiving, sharing and reuse. the environment described above is situational. a strong situational determinist research environment should not only coerce researchers into creating reusable data, but also give them confidence to do so, thereby creating a researcher disposition towards creating reusable data. using the colloquial metaphor that seems to be prevalent in research data management discussions, the current situation is mostly sticks and few carrots, and we need more carrots. promoting reuse: cognition vs. emotion there is a case to be made, and has been made by funding councils and institutions, that data management and reuse be addressed as a mandatory requirement in any funding application. the reasoned argument for data management stands clear: it is fundamental to transparent, high quality sustainable data generation. therefore, in psychological terms, data management for reuse is a “cold”, cognitive task – an intellectually conscious, controlled process based on explicit learning (kahneman, 2003). however, often the resistance to reuse draws not so much on logic, but sources that are more emotive. drawing on movements within political psychology, what we feel should not happen is to dismiss emotive impulses. we believe that emotions should be brought into the discussion between data creators and re-users. this is predicated on the belief that emotional responses are great motivators. emotions can be harnessed to aid decisions, for example, the emotion to care. ambition, incentives, professional acclimation can all be connected with data sharing and help researchers reach their decision to share. researchers make an effort in collecting and working with data, and therefore they should develop an affective relationship with them. they are their intellectual creators and they should be given reason and tools to present them to the community in the same way they do with publications. if we can tie good research data management and data sharing into recognized career advancement, we can bring with it esteem of peers not just for the publications but the data underpinning publications. if we can instill professional pride in replication and peer scrutiny of data creation like the academic community has instilled in journal publications, then by sharing data researchers will be a more ”important” with wider recognition than those who do not share because they will help advance the state of their discipline. those who chose not to, however, will have another emotion to mange – fear: the fear of professional irrelevance (king 1995, p.445). for without emotions such as care, or fear, what incentive––and as we have suggested, incentives are currently lacking––is there to think of the consequences of actions? through iassist quarterly winter 2011 21 iassist quarterly this, we could hope to see a dispositional environment towards data sharing emerge. support for data sharing procedures is an important factor in facilitating sharing as lack of awareness can be a serious obstacle. while resources exist to support data creators in generating reusable data, they are often not discipline-specific. for example, the first versions of the digital curation centre’s (dcc) data management planning tool (digital curation centre, 2011) or the australian national data service’s data management planning advice (australian national data service, 2011) offer detailed but generic support. although discipline-specific focuses are emerging, promoted in part through programs like jisc’s managing research data (joint information systems commission, 2009), as most are either generic tools or pure data management projects, these resources do not occupy the brokerage positions that data archives can assume. the ”brokerage” role of data archives – the supply and demand model the responsibility of a broker is as a third-person facilitator to bring ”sellers” and ”buyers” together. we can therefore think of the brokerage role for an archive in terms of facilitating the ”buying” (acquisition) and ”selling” (dissemination) of data between data creator and data re-user. archives know their ”market” for data, and have established relations with creators ”sellers” and re-users ”buyers”, they are institutions that talk to both communities from acquisition to dissemination via ingest. consequently, they become important regulators of this data market. they regulate the inflow and the quality of data on the supply side by encouraging data creators to share, leading the move to professional credit for sharing by making data citation possible and advising and supporting data creators on avoiding unnecessary obstacles to creating shareable data. however, they also regulate the output of data towards the demand side by disseminating them, increasing their visibility, and providing a service for responsible reuse of data. to highlight four cases, the ukda (2011), the icpsr (2011b), in the netherlands the dans (data archiving and networked services, 2011a) and the iqda (irish qualitative data archive, 2010) are national archives that have produced resources to aid data creators as well as providing data and dissemination support. however, archives do not only regulate supply and demand. through division of labor and specialization, they also add value to the data life cycle. archives undertake tasks that enhance data quality and data survival in an uncertain technological world. though not exhaustively, data archives provide long-term preservation of data with a strategy to ensure readability as file formats and technologies change. in addition, archives add value to data through structured metadata, catalogue records, and harmonization with comparative data collections. archives develop networks for secure and easier access of data for reuse. nevertheless, to provide high quality data, archives must adopt modern technologies and standards, ensure cooperation between same-discipline archives across countries, and promote dialogue with archives operating in other disciplines. through systematic interaction, archives can be the critical ligament that facilitates data sharing. incentives the role of the archive is to build incentives for both audiences to adopt best practices when dealing with data. from the supply side, it is important to increase the cognitive and emotional incentives for data sharing. we have already stated the important enticement for creators in making data available for reuse is their publications record, as their rewards and career advancements depend on that. the first step is then to make data citable. to do so, we need to provide the infrastructure and technology that allow the efficient referencing of data files. the most commonly used form of identifier is the digital object identifier (doi®) system (international doi foundation, 2011). these persistent identifiers are codes that connect a digital object such as a dataset, with accompanying metadata that includes author names, year of data collection and other important information of relevance. dois digitally identify journal articles, thus researchers are already familiar with their basic uses and functions. by having a doi allocated to a dataset, the researcher can be sure that by using that specific doi they refer to the same dataset. therefore, referencing a dataset within the publication used to create it becomes effective. a reader of this publication can then identify the very same dataset with no alterations and replicate the analysis. this ensures research quality and the primary investigator is acknowledged. gesis is a data archive that has a project providing persistent identifiers for data files in its collection. one example, hosted by gesis, is the da|ra project (gesis leibniz institute for the social sciences, 2011) gesis’s registration agency for social science research data. the da|ra infrastructure lays foundations for permanent identification, storage, and localizing to create citable research data. initiated in 2010 with a pilot phase, on entering 2012 the project is now in an upgrade phase. an expansion phase from 2013 to 2014 will centre on the development of useful services like user statistics, citation indexing, peer review possibilities for data, and registering of other data formats. another project of note is the effort by dans (data archiving and networked services, 2011b) to produce a competitive alternative to the doi. the dutch archive is involved in the design and implementation of a persistent identifier (pi) infrastructure in cooperation with the infrastructure-oriented surffoundation [sic], and koninklijke bibliotheek (national library of the netherlands). this collaboration seeks to establish a mechanism called the national resolver that would translate the pi into the current url of the object. in highlighting the projects and arguments we have presented thus far, it is our main goal to encourage researchers to take pride in their data creation activity, not just the outputs, and to invest time in making it reusable and archivable. we also aim to encourage researchers to value the work of other researchers who collect data, and to acknowledge this process as important and equal to other publication activities. to do that we focus on a new innovative data management training facility which we are involved in developing at gesis: the archiving and data management training and information center (gesis, 2012) a concept for training a new development that builds on the supply and demand model is the gesis plan to create a research data management and archive training centre for the cessda-eric european area. this area is inclusive of data archives in twenty european nations (cessda, 2011) the training centre will provide a central reference point for european researchers and archives, containing original resources and links to significant external resources, with the aim to ensure excellence in the creation and long-term preservation of reusable data, contribute to promoting the adoption of standards in research data management, and to advance data availability and reuse. the centre will also provide and coordinate training on technologies and tools used by data professionals. 22 iassist quarterly winter 2011 iassist quarterly by networking, and through surveys of demand for training needs, we are identifying themes and developing training concepts through potential collaborations with expert instructors. these concepts will be the basis on which courses are developed. the idea is to build resources around them using mixed and matched smaller thematic units depending on the needs of each specific course. our website will hold resources created by us, links to external resources, and will host information on the training center’s consulting activities. the virtual centre of competence will allow for consultation on best practice in research data management and archiving, including personal development and the promotion of skills training, provide information on our training activities, and offer structured teaching and self-learning materials. specifically, the centre will support data creators in implementing international standards of metadata and documentation. information for data creators about the importance and uses of persistent identifiers and will be given, plus advice on ethics and consent, details on issues of data ownership, and an overview of archiving software systems. the main support for data reuse is through the training of data archive staff to provide quality user support and to deal with increased volume of support requests. in addition, there will be information for archive professionals about new projects, new technologies, data discovery, and dissemination tools. furthermore, the presentation of projects on data harmonization will enable archive professionals to add to the value of data for their users and create an online user community engaged in task of harmonization. finally, and perhaps most importantly, the training centre looks to support other archives, libraries and repositories in ensuring state of the art data-related functions and in keeping up with the constant development of new technologies. this feature is not only useful for institutions either in a formative stage or that are not specialized social science resources, it is essential for all institutions operating in the data world to keep up with innovation, establish clear workflows, and strive for internationally accepted standards. the centre seeks to bring together the best examples and expert individuals to provide training. training will not only have the traditional form of workshops. it will be an active form of community building and incentive development through all communication channels provided to us by the new technologies. the core of our training concept is to negate all the reasons outlined in this text that allow researchers to sit on their data without sharing, and this can only be done with systematic incentive building. this training centre is only one way to augment the incentives of data sharing by bringing the subject closer to researchers’ hearts. however, the other driving factors mentioned and analyzed in this paper have to be pushed forward in order to ensure the emotive connection of researchers to sharing data, and to establish it an integral part of the scientific contribution. in the world of data, the imperative to share is clear. we have enough sticks; it is time to cultivate the carrots. references alliance of german science organisations (2010) “principles for the handling of research data” http://www.allianzinitiative.de/en/ core_activities/research_data/principles/ anderson, c. (2008) “the end of theory: the data deluge makes the scientific method obsolete”, wired magazine, 16(8) http://www. wired.com/science/discoveries/magazine/16-07/pb_theory australian national data service (2011) “data management planning” http://ands.org.au/guides/data-management-planning-awareness. html bishop, l. (2009) “ethical sharing and reuse of qualitative data”, australian journal of social issues, 44(3), pp.261-267 cessda: council for european social science data archives (2011) “member organisations” http://www.cessda.org/about/members/ cornell university (2005) “registry of digital collections” http://rdc. library.cornell.edu/search/index.php?mode=browse&type=collect ion data archiving and networked services (2011a) “data management plan” http://www.dans.knaw.nl/en/content/categorieen/diensten/ data-management-plan data archiving and networked services (2011b), “persistent identifiers” http://www.dans.knaw.nl/en/content/categorieen/diensten/ persistent-identifiers deutsche forschungsgemeinschaft (1998) “proposals for safeguarding good scientific practice wiley-vch” http://www.dfg.de/download/ pdf/dfg_im_profil/reden_stellungnahmen/download/empfehlung_ wiss_praxis_0198.pdf digital curation centre (2011) “dmp online” https://dmponline.dcc. ac.uk/ dryad (2011) “dryad partners” http://datadryad.org/partners economic and social data service (2011) “about the economic and social data service” http://www.esds.ac.uk/about/about.asp economic and social data service (various) “annual reports” http:// www.esds.ac.uk/news/publications.asp engineering and physical sciences research council (2011) “implementing the delivery plan” http://www.epsrc.ac.uk/plans/ implementingdeliveryplan/pages/default.aspx european commission (12 december 2011a) “digital agenda: turning government data into gold” http://europa.eu/rapid/pressreleasesaction.do?reference=ip/11/1524&format=html&aged=0&language =en&guilanguage=en european commission (12 december 2011b) “digital agenda: commission’s open data strategy, questions & answers” http:// europa.eu/rapid/pressreleasesaction.do?reference=memo/11/891& format=html&aged=0&language=en&guilanguage=en european data watch extended (2011) “about edawax” http://www. edawax.de/about/ european union (2010) “riding the wave: how europe can gain from the rising tide of scientific data. final report of the high level expert group on scientific data” http://cordis.europa.eu/fp7/ict/e-infrastructure/docs/hlg-sdi-report.pdf gesis leibniz institute for the social sciences (2010) “eurobarometer data service – international data sets distributed via archive networks, 2005-2009” http://www.gesis.org/fileadmin/upload/ dienstleistung/daten/umfragedaten/eurobarometer/contacts/ eb-user-stat_archives_2005-2009_v2.pdf gesis leibniz institute for the social sciences (2011) “über da|ra” http:// www.gesis.org/dara/home/ueber-dara/ gesis leibniz institute for the social sciences (2012) “archiving and data management training and information center” http://www.gesis.org/ archive-and-data-management-training-and-information-centre inter-university consortium for political and social research (2011a) “icpsr usage statistics” http://www.icpsr.umich.edu/icpsrweb/icpsr/ curation/usage.jsp iassist quarterly winter 2011 23 iassist quarterly inter-university consortium for political and social research (2011b) “guidelines for effective data management plans” http://www.icpsr. umich.edu/icpsrweb/icpsr/dmp/index.jsp international doi foundation (2011) “welcome to the doi® system” http://www.doi.org/ irish qualitative data archive (2010) “preparing qualitative data for archiving” http://www.iqda.ie/content/ preparing-qualitative-data-archiving joint information systems commission (2009) “managing research data (jiscmrd)” http://www.jisc.ac.uk/whatwedo/programmes/mrd. aspx kahneman, d. (2003) “a perspective on judgment and choice: mapping bounded rationality” american psychologist, 58(9), pp.697‐720. doi: 10.1037/0003‐066x.58.9.697 king, g. (1995) “replication, replication” ps: political science & politics, 28, pp.444-452 massachusetts institute of technology, “dspace@mit” http://dspace. mit.edu/ national institutes of health (2003) “final nih statement on sharing research data” http://grants.nih.gov/grants/guide/notice-files/ not-od-03-032.html national science foundation (2011) “nsf data management plan requirements” http://www.nsf.gov/eng/general/dmp.jsp natural environment research council (2011) “data centres” http:// www.nerc.ac.uk/research/sites/data/ organisation for economic co-operation and development (2007) “oecd principles and guidelines for access to research data from public funding” http://www.oecd.org/dataoecd/9/61/38500813.pdf research councils uk (2011) “rcuk common principles on data policy” http://www.rcuk.ac.uk/research/pages/datapolicy.aspx the engineering and physical sciences research council (2011) “epsrc policy framework on research data” http://www.epsrc.ac.uk/about/ standards/researchdata/pages/impact.aspx uk data archive (2010) “data management practices in the social sciences” http://www.data-archive.ac.uk/media/203597/datamanagement_socialsciences.pdf pp.17-21 uk data archive (2011) “create and manage data” http://www.dataarchive.ac.uk/create-manage university of california-san diego (2010) “digital library program” http://libraries.ucsd.edu/about/digital-library/index.html university of edinburgh (2010) “research data management policy” http://www.ed.ac.uk/schools-departments/information-services/ about/policies-and-regulations/research-data-policy university of hertfordshire (2011), “data management policy” http:// sitem.herts.ac.uk/secreg/upr/im12.htm university of illinois at urbana-champaign (2005) “ideals” http://www. ideals.illinois.edu/ university of rochester (2008) “ur research” https://urresearch.rochester.edu/home.action notes 1. gesis-leibniz institute for the social sciences, unter sachsenhausen 6-8, 50667 cologne, germany, tel. +49-221-47694 494. email: laurence.horton@gesis.org and alexia.katsanidou@gesis.org 2. all online literature references available 29 march 2012. 3. a persuasive case for qualitative data reuse is made by bishop (2009) 14 iassist quarterly 2015 iassist quarterly iassist quarterlyiassist quarterly abstract the continuous analysis of many cameras (cam2)2 project is a research project at purdue university for big data and visual analytics. cam2 has the ability to collect over 60,000 publicly accessible video feeds from many regions around the world. these video feeds were originally collected for improving the scalability of image processing algorithms and are now becoming of interest to ecologists, city planners, and environmentalists. with cam2’s ability to acquire millions of images or many hours of videos per day, collecting these large quantities of data raises questions about data management. the data sources have heterogeneous policies for data use, and some sources have no policies. separate agreements had to be negotiated between each video stream source and the data collector. in this paper, we propose to compare data use policies that are attached to the video streams and study their implications for open access. the need for common points of legal guidance for webcam stream users and publishers is demonstrated through this analysis of usage agreements. keywords: video, cctv, public access, open data, re-use, privacy, webcams. introduction thousands of network cameras have been deployed by many organizations for different purposes. for example, many departments of transportation (dot) deploy cameras along highways or congested streets (city of new york, 2015). national parks deploy cameras showing the views from visitor centers. some zoos deploy cameras showing animals’ activities. graham et al. (graham, riordan, yuen, estrin, & rundel, 2010) used geo-located cameras for a plant phenology monitoring system. the national park service of the united states deploys cameras observing air quality (national park service). nearly 20,000 cameras from a single site (weather underground) allow users to see weather worldwide. another site has more than 40,000 cameras (online promotion ag) watching tourist attractions. the data are available to the public through the internet: anyone connected to the internet can see the data (image or video) from these cameras. the data may provide insightful information about our world, such as traffic congestion, air quality, weather, etc. an important value proposition for this data may lie in the ability to extract and compare information from multiple sources. even though the data are publicly available, it is not easy using the data for scientific research due to accessibility challenges: there is no single site where the data from disparate sources are available. in the united states, the department of data sharing and re-use policies for webcam video feeds from international sources by line c. pouchard, megan sapp nelson, and yung-hsiang lu1 we have been building a system that allows large-scale analysis of image and video iassist quarterly 2015 15 iassist quarterly transportation of each state has its own website showing the traffic cameras. different websites have different data formats and require different ways to retrieve the data. more importantly, different organizations have different policies and restrictions about how data may be used. this paper describes the challenges of using the globally available camera data for scientific research. in order to facilitate research using the data from global camera networks, we have been building a system that allows largescale analysis of image and video (kaseb et al, 2015a; kaseb et al, 2015b; hacker et al, 2014). cam2 (continuous analysis of many cameras,) is a research tool for solving some of problems mentioned above: it is a single site through which the data from many heterogeneous cameras can be retrieved and analyzed. the cloud-based computing engine can handle large quantities of data. an event-driven programming interface offers the flexibility to execute diverse programs analyzing the data. users execute the analysis, using either their own scripts or with scripts provided by the system on the computing engine at the back-end of cam2, and download the analysis results. in most cases, the data are intended to be viewed by human eyes through web browsers. cam2 uses only publicly available data -no password is required to access the individual feeds, only to login into cam2. a big challenge in constructing cam2 is to obtain the permissions granting usage of the data for scientific research. it is possible to write computer programs to retrieve the data from the sources’ web sites but, as a courtesy, we requested explicit permissions from the data owners. in remote sensing and video stream applications, data arrive at very high frequency. they exemplify the volume and velocity characteristics of big data, characteristics that have implications for infrastructure and privacy policies. volume refers to the amount of storage needed to accommodate the data, and velocity to the speed at which the data is produced and transferred through networks (jagadish et al., 2014). issues of data retention, sharing and re-use arise for big data that are scantily addressed in the legislation on privacy. individual policies sometimes address these issues but each data owner has a different policy and set of restrictions. thus, we research the terms that data owners use to articulate access and re-use of their data in those policies and how these policies in turn affect re-use of the data for scientific research. this paper reports our results from the analysis of the policies we obtained from disparate data sources. we frame the discussion with examples of data privacy law in the us and eu where a comprehensive, unified framework is given to member states by the european council directive. we perform qualitative analysis on the 15 policies we obtained and present the results in this paper. the discussion shows a comparison of the policies with a focus on terms of use. the cam2 project was designed to test and improve image analysis algorithms in real time, but other researchers, such as environmental scientists, city planners, and ecologists are becoming interested in the data as a resource for additional scientific analyses. we are illustrating the implications of those policies for other researchers interested in re-using the data. one of our findings is that these policies tend to be sparse in terms of re-use, so we are also trying to understand the implications of these gaps for researchers interested in re-using the data. literature review we analyzed the uk’s data protection policy as applied to cctv for insight into areas that may be addressed by a cctv usage policy. this policy was selected due to the widespread adoption in the uk and the multiple iterations that this policy has gone through in response to appeals to the european court of human rights. the framework for the uk data protection policy is provided by the european union data protection directive of 1995 (european council, 1995) and enforceable through the eu member states since 2004. the 1995 eu data protection directive and its amendments provide three important items with regard to the video streams we are interested in: 1 iimages or voice are considered personal data (article 29, working party); 2 an all-encompassing definition of processing is provided, including “collection, recording, organization, storage, adaptation or alteration, retrieval, consultation, use, disclosure by transmission, dissemination or otherwise making available, alignment or combination, blocking, erasure or destruction” (article 2(b)); 3 when data is transferred to another country, member states must ensure that this country affords equal protection to individual privacy. the eu data protection framework decision applies to crossborder exchanges of personal data processed in the framework of police and judicial cooperation in criminal matters but does not apply to domestic data (european digital rights, 2009). the uk cctv code of practice describes the many factors that a cctv operator must consider and the rules that must be followed before the implementation of a cctv system (information commissioner of the uk, 2014, 2014a). operators should conduct a privacy impact assessment and ensure that effective administration with decision-making power about the storage, possible encryption, retention and use of the images is put into place. images should be retained in a secure location for no longer than strictly necessary as prescribed by the purpose of recording them, accessible only by authorized personnel in a controlled location, and erased once the defined period of retention is reached. any operator putting images on the internet must consider the possible disclosure of individuals’ personal data and proceed accordingly. it is also recommended that operators remain in control of the information and conduct periodic audits to review the many provisions of the code of practice. another important aspect of these practices is the emphasis on making the use fit the purpose: operators are to ensure that the images are used according to the purpose for which the system was put in place (often surveillance). but the surveillance and video systems in place in public places evolve rapidly with new technology and a greater acceptance by the public. examples of emerging technology which impacts the application of regulations include ubiquitous computing and wearable imaging devices, automatic licence (or number) plate recognition, unmanned aerial vehicles, and cameras equipped with direct access to the internet. these new technologies and the increased ability to link information challenges the protections afforded by the eu data directive and the framework put in place by member states. 16 iassist quarterly 2015 iassist quarterly as a consequence of a more interconnected society, re-use of information for purposes completely different than the ones for which the information was captured is becoming common (coudert, 2009). the principle of collecting and retaining data specifically for the purpose described at inception of the system is being eroded. safeguards put into place are not respected or impractical in the context of interconnected video networks. the uk government itself recognized that the lack of definitions of data sharing hampered communication between recovery efforts in the aftermath of the 2005 london bombings (coudert, 2009). some scholars argue that the protections are toothless, the regulatory bodies lack the resources to enforce them and rely upon the goodwill and cooperation of those they are regulating (gras, 2004). nonetheless, in the uk, a legal and regulatory framework exists that is relatively homogeneous, and affords some amount of protection. there is no equivalent data and privacy protection framework in the us at the federal level but instead what many call a ”patchwork” of regulators and regulations, common law, federal legislation, the us constitution, state law and certain state constitutions. privacy is regulated primarily by industry on a sector-by-sector basis and us regulation of the private sector is minimal (levin and nicholson, 2005). the major components to privacy law as it applies to data include the fair information practice principles, the bulk of which were adopted in the early 70s and 80s. the federal trade commission issued a series of streamlined principles at the beginning of the 21st century, focused on online privacy, including the notions of notice and consent for personal information collection. notice and consent cover the requirements to inform consumers that private information is collected in the course of providing a service and to give them the choice of accepting that their information may be used for other purposes. one observes a trend that these principles have been weakened to focus on the procedures of giving notice and consent rather than on substance (strandburg, 2014). the us is more concerned with protecting the privacy of its citizens from government than regulating industry practices. at the same time, industries are pushing back from any regulation and privacy protection is more the result of market constraints and unpredictable common law (levin and nicholson, 2005). as an example of the piecemeal approach, many regulatory acts issued by various regulatory bodies currently govern privacy and could be applied to data (see appendix 2). privacy protection in public places, which is where video surveillance cameras operate, does not exit, following a supreme court and numerous federal court rulings that there are no privacy expectations in a public place (slobogin, 2002). data use is largely unregulated except in a few exceptions (health care, credit reporting) and re-use of big data renders meaningless the notions of notice and consent due to the complexity and numerous possibilities of aggregating the data (ohm, 2014). figure 1: formal policies and coding distribution iassist quarterly 2015 17 iassist quarterly analysis we analyzed the contents of fifteen usage agreements with a variety of different entities, both international and united states, government and business entities. as several of these agreements specify that the terms shall not be made public, no entity will be identified here, other than as a representative of a class (government entity, business entity, us or international). the agreements were analyzed with nvivo using a coding structure that was developed based upon the terms present in the sample of usage agreements. the coding schema developed in this project is included in this article as appendix 1. the nodes were identified based upon the terms that were present in the collected agreements and the context in which those terms were presented (term creator; creator classification; technical guidelines and specific information that we sought as researchers including data sharing and recording or duplication). as a basic classification scheme we divided formal from ad hoc policies. by formal policies, we mean that contractual agreements were signed between the researcher at purdue university or a university representative and a legal representative of a data providing entity. these policies tend to be longer and contain figure 2: ad hoc policies and coding distribution many disclaimers protecting the provider entity from any liability (costs, disputes, responsibility for the actions of a third party) from the use of the feeds. they also tend to assert that the policy itself is protected by copyright and in one case not to be made public without express written permission from the provider. ad hoc policies are those created by the manager of a data stream. these tend to be focused upon the technical aspects of accessing the data stream, are not formally ratified, and provide little guidance about how the stream may be used. we had 10 ad hoc policies and 5 formal ones. figure 1 shows the terms present in the formal policies and the total number of codes recorded for those terms. for instance, the code “use restriction” is assigned 12 times over 3 separate policies. branding is assigned 7 times over 5 policies. the ad hoc policies are much less formal in content and form. often they are in the form of an email to the researcher granting the researcher the right to use the video feeds. they are often technical in content, explaining how to access the feeds, providing apis, and sometimes expressing concerns or providing limitations about the download rate so as not to slow down their systems. because of the technical content of these emails, it often 18 iassist quarterly 2015 iassist quarterly appears that they might have been written by the developers or maintainers of the system. figure 2 shows the distribution of coding across the ad hoc policies documents and the total number of times a particular code is encountered. discussion the policies were divided into two primary classifications, formal and ad hoc. formal policies were developed by law entities as contracts specifying usage guidelines. these required signatories for both the user and the data stream provider to sign a legally binding agreement prior to accessing the data stream. these legal agreements included many usage restrictions and guidelines in common, including branding or marketing, attribution expectations, re-use guidelines, data retention restrictions, and access termination. ad hoc policies were created on the fly by the data stream provider. they did not require formal signatures. generally they included minimal guidelines for use and focused on technical guidelines and explicit permission for the end user to access the data stream. in the case of the ad hoc policies, the permission was given as a person-to-person communication in the form of an email or other written communication. the ad hoc policies were focused upon granting permission to use the web streams to dr. lu. use restrictions were imposed in seven different agreements, the five formal policies, and two ad hoc policies. there were several different types of restrictions. the most common restriction was a technical restriction. to protect servers from being overwhelmed, the rate of download was restricted, generally either in time between downloads and file size that should be downloaded and occasionally both. the download rate restrictions included are in the table below. for the researcher, each of these separate restrictions has to be built into the automated retrieval system. as shown above, a “universal” rate of download could be one picture per hour, as all other exceptions fall within that rate. however, the researcher would be missing a large quantity of data, as demonstrated by the allowable download rate of one image per second. therefore the cam2 system needs multiple use categories built in so that examples of ,me limit or file size limit a picture will not be captured more than once every two minutes allowed one picture per hour per camera allowed one 320 x 240 jpeg per second no camera will be accessed more than once every five minutes no more than a cumula,ve 24 hours of images that are no more than one week old. table 1 1 different cameras can be accessed at different rates. this leads to a significantly more complicated download script. additional restrictions are focused on attribution and branding. for both government and business entities, attribution of the original data stream was requested. however, there was no general agreement in how that attribution should be carried out. in some cases, a hyperlink on the cam2 website was requested. in others, a logo should be placed on the cam2 website where it is visible during the search of camera data streams. others generally requested that the camera owner be identified but did not specify where or how. in other cases, an academic citation to the source was requested on publications or websites. again, this broad range of requests for attribution makes it difficult for the researcher to efficiently meet these terms. in general, the “universal” policy may be to represent the data owner with logo and hyperlink on the screen for all cameras, and cite all data providers in academic articles where that data was used. however, unless a script is run that identifies the data owner for all cameras used in a given simulation, the sheer number of cameras (60,000 and growing) that may be included in the analysis make this unwieldy. data sharing was the primary provision of interest to the researchers. perhaps not surprisingly, this was not on the radar of most of the entities. it did appear in four formally developed policies. in those cases, one entity indicated that data sharing was not required. two entities expressly prohibited data sharing and reuse. the fourth entity allowed reuse only in the case of broadcast journalism. this is obviously discouraging to researchers who may want to continue to develop a video analysis system to create truly reproducible research results on specified sets of camera and image data. in the case of big data, reproducibility is an ongoing area of concern and research. further developments in this aspect of big data research may help to deal with some of the issues presented by these data sharing provisions. there were also restrictions on the ability to retain, copy or duplicate footage. one policy explicitly prevented retaining more than 5 consecutive minutes of images, and no more than 24 hours of footage that is a week old or more. but this policy does not prohibit data sharing explicitly. another restricts all right to copy or duplicate in any way the video feeds unless the user is a media outlet. these restrictions on retention and copying place an additional burden on the possibility of re-use of the data by other researchers. the policies provided to the researcher were for the most part unable to provide the researcher with the rights, permission, or guidance needed to extend use of the data streams to other research projects. data sharing and re-use was built into only one ad hoc policy and two formal policies. in those cases, data sharing and re-use was explicitly allowed to broadcasters only in one formal policy, explicitly disallowed in another formal policy, and is specifically left open to the discretion of the end user in the ad hoc policy. there is no similar language between the three policies. while these video streams present many interesting possibilities for large scale analysis having to do with climate, economics, and engineering, the policies that mention data sharing for the most part restrict it. many more policies do not mention sharing one way or another. even if the researcher took the most liberal interpretation of that hole in the policy, iassist quarterly 2015 19 iassist quarterly it leaves the researcher with the problem of having to build restrictions within the cam2 system for which camera feeds can be re-used by other researchers. another solution is to not use cam2 data feeds that explicitly prohibit re-use. additionally, recording and duplication are prohibited by two formal usage policies. this means that specific considerations for feeds that cannot be re-used beyond the cam2 purpose would have to be built in, as well. in those cases, the relative value of video streams that cannot be reused may need to be reconsidered and these feeds may need to be dropped from the project, because the real value of the cam2 system goes beyond the simultaneous analysis of data feeds for improving image processing algorithms. if those streams cannot be included for future scientific use, are these video streams actually useful in this context? the policies add value to the video streams. those policies that do have specific information about the re-use of data streams and duplication of images are giving the researcher the chance to do extensive new science. the lack of policy indicates a video stream that may end as a liability to the research project due to lack of clear guidance on appropriate duplication, re-use, and preservation actions that can be taken. suggested actions we propose that an agency such as niso, the national information standards organization, should create a template usage agreement that highlights terms that encourage scientific reuse of video data. the template may include language giving permission for the creation of derivative works and information regarding appropriate storage and retention as well as preservation (including protection of privacy and metadata to preserve quality and identify date and time of creation.) this template could then be offered by the researcher to the video data producers where no formal policy already exists for that entity. the template could also be adopted by local and state government agencies as the starting point for their formally adopted policies, thereby ensuring that the relevant terms needed for scientific reuse are included in these legally binding agreements. we propose the following as key components of this template: data provider identification one problem that we identified was that in many ad hoc policies the data owner was not identified clearly. this would help with both proper attribution and accountability. download rate many different download rates may be technically feasible due to the differences in camera and server specifications but for ease of use in planning and building data analysis systems, we propose that a specified rate of download be identified by niso for a generic usage agreement, with specific higher or lower rates of download negotiated on a project-by-project basis between the data owner and data user. file size as in download rates, many different file sizes may be technically feasible due to differences in the specifications of the equipment used by both the data owners and the data users. we suggest that the template include language that guides the clear negotiation of file size. statement of re-use that allows for general scientific investigation a key finding of our investigation was that scientific or academic re-use was sometimes prohibited partially or in full. a statement that the data stream can be re-used for scientific or academic research generally would enable multidisciplinary research investigations on the same data set. additionally, a statement that the data streams may be used to create derivative data sets would help further scientific research, as subsets of collated data sets could then be created to accompany publications. privacy a statement governing appropriate use of the data set regarding individuals’ privacy should be included in all terms of video data re-use. we believe that the default should be that individuals’ privacy will not be infringed upon by the re-use of the streaming video data. for projects that seek to use video to develop face identification algorithms and similar technologies, these terms should be negotiated on a project-by-project basis. quality control if data producers are worried about data streams being re-distributed in a misleading way, a date and time stamp, metadata regarding the source of the original data stream, as well as a branding icon may be required on still images or video clips. those requirements should be spelled out in the terms of use. attribution a suggested attribution including an academic citation that is generated as part of the usage agreement would go a long way towards ensuring that data users have an easy, consistent way to refer to the data stream that they are accessing. as part of this, data providers may consider instituting a persistent url for the website hosting their streaming data. additionally, the date or time stamp may be referenced in the attribution as well. retention and preservation data streams may only be valuable if analysis can be performed over extended collections of data representing days, months or years worth of data. suggested language may include negotiable terms for the storage and preservation of streaming data for use in longitudinal data analysis systems. accountability the template should include the responsibilities of the data user to the data provider whether that be proper attribution, reports back to the data provider on how the data is used, or assurance that quality control measures have been put in place. these terms should also be negotiated and included in the agreement. conclusion this paper reported our experience requesting and obtaining permissions using publicly available video data for scientific research. even though the data are already publicly available on the internet, the heterogeneous sources of data and the terms of use specified or missing create many administrative difficulties. 20 iassist quarterly 2015 iassist quarterly there is no consistent policy among different states or cities of the same country. several changes could be made at the regional and national levels to facilitate the re-use of data. to start with, we recommend that the government accountability office or other similar agency develop a single video data re-use policy that is applicable to all agencies. on the international level, central governments should establish a common set of rules governing video data. among different countries, an international standard could be established for the appropriate terms that enhance scientific reuse to be included in usage policies. ideally, a legal framework should be created that will protect both video data producers and end users. this requires a culture change that encourages additional legal guidelines governing the privacy and data management practices of public and private companies and government agencies. legal scholars are currently calling for this sort of legal framework but some are more concerned with citizens’ privacy rights to be protected from government action (slobogin, 2002; greer, 2012) than with sharing scientific data. with the advent of big data, other legal and privacy scholars are lately calling for a comprehensive framework and national discussion about the scope and foundational concepts of privacy in the context of big data re-use (lane, stodden, et al, 2014). if templates and policies similar to these were adopted, many more scientific uses of video data streaming on the internet could be embarked upon, ranging from environmental and ecological studies to economic assessments based upon the movement patterns of individuals showing up on cameras in a variety of public places. until these policies are specified, the legal liability of the researcher that presumes too much is too great to enable these more advanced, longitudinal studies of streaming video data. references city of new york. (2015). nycdot: real time traffic. retrieved from http://nyctmc.org european council. (1995). directive 95/46/ec of the european parliament and of the council of 24 october 1995 on the protection of individuals with regard to the processing of personal data and on the free movement of such data. (official journal of the european union, no. l 281/31). brussels: european council. european digital rights. (2009). data protection framework decision adopted. retrieved from http://history.edri.org/edri-gram/ number7.3/data-protection-framework-decision graham, e. a., riordan, e. c., yuen, e. m., estrin, d., & rundel, p. w. (2010). public internet‐connected cameras used as a cross‐continental ground‐based plant phenology monitoring system. global change biology, 16(11), 3014-3023. greer, o. j. (2012). no cause of action: video surveillance in new york city. michigan telecommunications and technology law review, 18, 589-626. hacker, t. j., & lu, y.-h. (2014). an instructional cloud-based testbed for image and video analytics. proceedings of the 2014 ieee 6th international conference on cloud computing technology and science (cloudcom). singapore: ieee. information commissioner of the uk. (2014). cctv code of practice draft for consultation. retrieved from https://ico.org.uk/media/ about-the-ico/consultations/2044/draft-cctv-cop.pdf information commissioner of the uk. (2014a). in the picture: a data protection code of practice for surveillance cameras and personnel information. retrieved from https://ico.org.uk/media/fororganisations/documents/1542/cctv-code-of-practice.pdf jagadish, h., gehrke, j., labrinidis, a., papakonstantinou, y., patel, j. m., ramakrishnan, r., & shahabi, c. (2014). big data and its technical challenges. communications of the acm, 57(7), 86-94. kaseb, a. s., berry, e., rozolis, e., mcnulty, k., bontrager, s., koh, y., . . . & delp, e. j. (2015a). an interactive web-based system for largescale analysis of distributed cameras. proceedings of imaging and multimedia analytics in a web and mobile world conference, san francisco: spie. kaseb, a. s., chen, w., gingade, g., & lu, y.-h. (2015b). worldview and route planning using live public cameras. proceedings of the imaging and multimedia analytics in a web and mobile world, san francisco: spie. lane, j., stodden, v., bender, s., & nissenbaum, h. (2014). privacy, big data, and the public good: frameworks for engagement. new york: cambridge university press. levin, a., & nicholson, m. j. (2005). privacy law in the united states, the eu and canada: the allure of the middle ground. university of ottawa law and technology journal, 2(2), 357-395. national park service. (2015). nps explore nature. [data file]. retrieved from http://www.nature.nps.gov/air/webcams/ ohm, p. (2014). changing the rules: general principles for data use and analysis. in j. lane, v. stodden, s. bender & h. nissenbaum (eds.), privacy, big data, and the public good: frameworks for engagement. new york: cambridge university press. online promotion ag. (2015). webcams.travel. retrieved from http:// www.webcams.travel/ scherr, c. (2007). government surveillance in context, for e-mails, location, and video: you better watch out, you better not frown, new video surveillance techniques are already in town (and other public spaces). i\sp: a journal of law and policy for the information society, 499(3). slobogin, c. (2002). public privacy: camera surveillance of public places and the right to anonymity. mississippi law journal, 72, 1-87. strandburg, k. (2014). monitoring, datafication and consent: legal approaches to privacy in a big data context. in j. lane, v. stodden, s. bender & h. nissenbaum (eds.), privacy, big data, and the public good: frameworks for engagement. new york: cambridge university press. weather underground. (2015). weather webcams. retrieved from http://www.wunderground.com/webcams/ wright, d., friedewald, m., gutwirth, s., langheinrich, m., mordini, e., bellanova, r., & bigo, d. (2010). sorting out smart surveillance. computer law & security review, 26(4), 343-354. iassist quarterly 2015 21 iassist quarterly table 2: list of nodes and instances of coding references node sources references ad hoc policy ad hoc policies are those created by the manager of a data stream. these tend to be focused upon the technical aspects of accessing the data stream 10 13 formal policy contractual agreements were signed between the researcher at purdue university or a university representaeve and a legal representaeve of a data providing enety 8 8 sell photos some webcams only provide sell photos 5 7 videos other webcmas provide video footage of the area they cover 4 5 alribueon meneoning the data provider as the source of the data 9 10 branding displaying a visual ideneficaeon of the source (logo or image) 9 10 copyright claiming ownership of the data 1 1 costs financial transaceons regarding the feeds and equipment to obtain them 4 7 data sharing involvement of a third party, including making the feeds accessible 5 11 download rate frequency at which the data is downloaded, expressed in frames per unit of eme, volume, or bandwidth usage 5 7 explicit permission for research to use stated permission to use the data in a policy. purpose of the use is someemes restricted. a login password is someemes provided. 14 14 privacy specificaeon that the cam2 system renders personal ideneficaeon impossible 2 3 public domain specificaeon that the data provided to cam2 is already in the public domain 2 2 quality assurance claims regarding the accuracy, currency or completeness of the data feeds or accessibility of cameras 7 8 recording or duplicaeon storing or copying webcam feeds. 2 2 report back request by data provider to report progress and outcomes to the data source 1 2 reteneon explicit permission to retain a certain amount of data for a specific period of eme 1 5 rights the non-exclusive permission to use the data 6 8 technical instruceons descripeon of the technical mecanisms for accessing webcam feeds at the provider’s site 10 14 terminaeon the ability for the provider to end the provision of feeds without penalty or prior noece 8 15 use restriceons any reduceon of the rights of the cam2 project with regard to the use of the webcam feeds. 7 16 1 appendix 1 22 iassist quarterly 2015 iassist quarterly appendix 2 table 3: examples of the us regulatory bodies who have issued acts that affect privacy. federal trade commission fair credit repor2ng act consumer repor2ng agencies must maintain accurate records and can forward records to anyone with a legi2mate interest. 1970 department of jus2ce privacy act regulates the use of data by government agencies. 1974 federal communica2on commission cable communica2ons policy act cable companies are not allowed to collect or share personal informa2on without individual consent. 1984 federal communica2on commission enforces video privacy protec2on act video stores cannot disclose their customers’ rental history. 1988 department of health and human services health insurance portability and accountability act protects pa2ents’ health informa2on from being released to poten2al employers. 1996 federal trade commission financial moderniza2on act financial ins2tu2ons must have and share a privacy policy by which customers can decline sharing their personal informa2on with third par2es. 1999 state of california online privacy protec2on act one of the most comprehensive laws. websites’ privacy policies must be highly visible and customers must be informed of third party use of their data. 2003 1 iassist quarterly 2015 23 iassist quarterly notes 1. line c. pouchard (pouchard@purdue.edu) is assistant professor, and computational science information specialist at purdue university libraries. she specializes in big data. megan sapp nelson (mrsapp@purdue.edu) is associate professor of library sciences and engineering librarian at purdue university libraries. yung-hsian lu (yunglu@purdue.edu) is associate professor of computer and electrical engineering at purdue university and acm distinguished scientist. this paper is not substantially different from the paper presented at the 2015 iassist. 2. https://cam2.ecn.purdue.edu/ microsoft word 49-1-pival-final-mc (1).docx 1/8 pival, paul r. (2025). support for computer-assisted qualitative data analysis software in arl libraries, iassist quarterly 49(1), pp. 1-8. doi: https://doi.org/10.29173/iq1144 the creative commons-attribution-noncommercial license 4.0 international applies to all works published by iassist quarterly. authors will retain copyright of the work and full publishing rights. support for computer-assisted qualitative data analysis software in arl libraries paul r. pival1 abstract academic libraries are filled with niche support services that are unique to their primary clientele. while rarely taught to librarians in a formal academic setting, the support of computer-aided qualitative data analysis software is one such niche that appears in academic libraries across north america. but how common is it, and at what level is support offered amongst members of the association of research libraries? this paper attempts to answer this question. keywords caqdas, qda, academic libraries, qualitative data introduction qualitative data unveils the rich tapestry of human experiences, capturing the emotions, stories, and nuanced insights that numbers alone can never tell. academic libraries are uniquely positioned to support scholars using qualitative data, as they operate at the crossroads of disciplines, bringing together diverse users, resources, and services that reflect the multifaceted ways in which knowledge is discovered, shared, and applied. qualitative data, generally, can be described as the non-numerical information that researchers collect through interviews, focus groups, and observations (qualitative data, nnlm, n.d.). this type of data can be crucial for understanding complex social phenomena, and over the past several decades the software available for analysing them have become increasingly sophisticated. i believe an academic library is a logical choice to offer support for computer-assisted qualitative data analysis software (caqdas) for the following reasons: 1. research support: academic libraries are already central hubs for research support on campus. they often assist with various aspects of the research process, making qualitative data analysis (qda) software support a natural extension of their services. 2. interdisciplinary nature: qda software is used across multiple disciplines in social sciences and humanities. libraries serve all departments, making them well-positioned to support diverse user needs. 3. technology infrastructure: academic libraries often have robust it infrastructure and staff familiar with various software tools, enabling them to host and support qda software effectively. 4. expertise in information management: librarians are skilled in organizing and managing data, which aligns well with the data handling aspects of qda software. 2/8 pival, paul r. (2025). support for computer-assisted qualitative data analysis software in arl libraries, iassist quarterly 49(1), pp. 1-8. doi: https://doi.org/10.29173/iq1144 5. neutral space: libraries provide a neutral, cross-disciplinary environment where researchers from different departments can collaborate and share knowledge about qda methods. 6. existing training programs: many libraries already offer workshops and training sessions on various research tools, making it easy to incorporate qda software instruction. in 2022 liz cooper published an article describing how she came to develop a qualitative data analysis support structure at her university library and argues that this should be a common role for academic libraries (cooper, 2022). her article described my journey so accurately that i felt compelled to examine the website of every member of the association of research libraries (arl) to learn how prevalent this role and service is. were we outliers, or two among many? because this is a role that i spend a fair amount of time on, i was specifically interested in learning whether other academic libraries support the use of computer-assisted qualitative data analysis software (caqdas) on their campuses. this is an exploratory study seeking to determine how many members of the association of research libraries offer support, and at what level, for caqdas. literature review several papers have explored one aspect or another of qualitative data support within academic libraries. in a groundbreaking book chapter, mandy swygart-hobaugh analysed job postings for social sciences data librarians in north america, surveyed social science librarians, and examined 53 research guides describing qualitative data support services offered by social science librarians specifically in the united states. one of the primary goals of that research was to compare the numbers and roles of librarians in supporting qualitative and quantitative data. as expected from earlier research on the subject, quantitative data support services were far more prevalent than qualitative. of the 53 research guides that were identified, “eighteen (34.0%) were general qualitative research guides with no discipline specified, fifteen (28.3%) were created for specific qualitative methods courses, eleven (20.7%) were dedicated computer assisted qualitative data analysis software (caqdas) guides, eight (15.1%) explicitly targeted individual or multiple disciplines, and one (1.9%) was a data management guide with recommendations on managing/sharing qualitative as well as quantitative data” (p.170). “among the eleven dedicated computer assisted qualitative data analysis software (caqdas) guides, five (45.5% of 11) indicated that librarians were available for consultations or training workshops on the software, while the remaining six (54.5% of 11) did not explicitly indicate as such (p.171).” one of the major conclusions of her work is that “both the survey results and the online research guides suggest that very few librarians are offering caqdas support” (p. 172). she concludes with recommendations for four key areas in which social sciences librarians might expand their qualitative data support services, including qualitative data analysis, qualitative data discovery, qualitative data for teaching and learning, and qualitative data management and sharing (swygart-hobaugh, 2016). my research focusses on the first of these, with an attempt to determine how many association of research library members are offering support for qualitative data analysis, and more specifically, which services they offer. 3/8 pival, paul r. (2025). support for computer-assisted qualitative data analysis software in arl libraries, iassist quarterly 49(1), pp. 1-8. doi: https://doi.org/10.29173/iq1144 recognizing that this knowledge is not typically taught as part of the curriculum leading to a library degree, catherine hansman describes strategies and methods for teaching qualitative research methodology to librarians and novice researchers, a necessary step along the path of being able to support caqdas within the academic library. she also takes time to examine various research methodologies that should be introduced to those seeking to understand the nature of qualitative research and discusses which learning theories prove to be most effective in teaching the discipline (hansman, 2015). librarians who already offer support for qda have noted that it’s often difficult for students to learn where to go to receive support. cain et al. explored “the discoverability of qualitative research support services” on a sample of academic library websites (p. 1). in essence, this was a usability study to determine now how many libraries were offering services, but of the ones that were, what are the characteristics that make discoverability easy or hard. for this study, the authors sampled 95 libraries at public and private institutions in north america. they worked primarily from the list of the association of research libraries, but did not include them all, and added ‘a few’ libraries that are not members of arl. the authors determined, “it is hard to find qualitative services on the websites of most libraries in our sample.” perhaps many of these libraries do not provide these services, which would be in line with the findings of swygart-hobaugh’s (2016) study. yet, in some cases, we eventually found mention of qualitative services offered by the libraries after trial-and-error searches using various search terms. we often found these services by entering the brand names of popular qualitative software, such as nvivo, atlas.ti, dedoose, and maxqda. approximately half of our sample references at least one qualitative software brand. libraries mention, for example, that a brand is available on campus computers and/or that staff are available to provide software training. (p. 3) (cain et al., 2019). these findings informed my approach when searching the websites of all the arl libraries. quantitative data appears to be a more familiar subject to many libraries, especially those serving in roles of data librarianship, including research data services. hagman and bussell interviewed 13 “academic librarians about their understanding of data literacy, qualitative research, and academic library infrastructure around qualitative research” (p.1). as did all the others in this section, they note the nearly ubiquitous nature of support within libraries for quantitative data, while support for qualitative data and analysis is much less frequent. they conclude that, “academic libraries must consider how their approaches to data literacy and research data services can serve to limit or expand the notion of what counts as research, and even who can bring their research knowledge to the scholarly conversation. developing services that are ostensibly open to all library patrons but ultimately only serve those whose research uses only one type of data sends a message about what kinds of research are valued and worth supporting” (p.9) (hagman & bussell, 2022). caqdas is certainly taught in workshops within libraries, as described by røddesnes et al. (2019), (cooper, 2022), and (kang & sinn, 2024). the first two discuss the development of caqdas workshops at two specific institutions, while the third offers a survey of technology workshops offered by 43 american libraries, a small number of which include caqdas. michalovich describes introductory nvivo training workshops and consultations provided by one arl member library in canada (michalovich, 2022). 4/8 pival, paul r. (2025). support for computer-assisted qualitative data analysis software in arl libraries, iassist quarterly 49(1), pp. 1-8. doi: https://doi.org/10.29173/iq1144 none of these reports, though, provide a clear picture of how many libraries in north america offer such support. the aforementioned seed article by cooper (2022) chronicles both the birth of qualitative data support within academic libraries, and, through personal narrative, makes the case for why the library is a natural fit for supporting qualitative data. she also provides examples of different levels of support that can be offered and makes suggestions for how and where librarians can bootstrap their knowledge to be able to begin offering support for qualitative data analysis. methods i visited the websites of all 128 association of research library (arl) member libraries between august 19-23, 2024 (list of arl members — association of research libraries, n.d.). as discussed by cain et al. (2019), it is not often easy to determine whether caqdas support is offered within a particular library setting. they found that approximately half of the sites in their study mentioned a qda product by name, and there was no ‘universal’ acronym used to discuss the concept of caqdas on the sites. in addition, that study found that often, “information about qualitative services was isolated from other data services pages on the library website—for example, the information might reside on a libguide for a course or discipline” (cain et al. , 2019, p. 4). forewarned by these findings, i employed several search strategies to gather results. when available, i used a ‘search this site’ choice on the library’s home page to search for any of the following terms, (nvivo, qda, ‘qualitative data’, atlas.ti, maxqda, dedoose). several sites did not appear to include results from springshare's libguides platform (libguides content management and curation platform for libraries, n.d.) within their results, so i also examined each site for the existence of a ‘research guides’ or ‘library guides’ section, and, if found, those pages were also searched for the same terms. finally, if a ‘search this site’ option was not available, i used google’s site: operator to search for each of those terms against a specific site (e.g. maxqda site:library.university.edu). when any result was found, i made a closer inspection of the result to determine whether the library offered support for the use of qualitative data analysis software. if so, at what level, or if they were simply acknowledging the existence of these tools, either generally, or elsewhere on campus. cooper (2022) suggests libraries can provide tiered levels of support for qualitative data analysis, depending both on the level of expertise of the supporting librarian, and the level of need perceived at the institution. summarizing her tiers,  tier 1: provides a bibliography of key methods books and articles at your library.  tier 2: research guide that discusses different qualitative research methodologies and includes tips about searching in library databases for articles that utilize qualitative methods.  tier 3: includes software-specific documentation.  tier 4: offers instruction and /or consultations on the use of qualitative data analysis software. with these tiers in mind, i also categorized each page i found by tier. an example search workflow follows. 5/8 pival, paul r. (2025). support for computer-assisted qualitative data analysis software in arl libraries, iassist quarterly 49(1), pp. 1-8. doi: https://doi.org/10.29173/iq1144 brown university library (https://library.brown.edu/) offers a search box on their home page. making sure to choose the option to search the library website (not the default), i searched for the phrase, “qualitative data”, which yielded many results. had there been none, i would have moved on to the other terms in my search strategy. one of the results led directly to a libguide for nvivo (wilkinson saldaña, n.d.). upon visiting that page, i learned that brown offers a site license for the software, and that a team of librarians at brown is available for consultations to help “evaluate nvivo / qualitative coding for your project, discuss automation features, etc.”. there are links to resources to learn more about how to use the software, including traditional literature licensed by the campus, and a brownspecific workshop called “nvivo essential training” has the following description: “this workday learning course explores how to leverage nvivo for collecting, organizing, and analyzing nonnumerical research data, such as images and text. the course covers key terminology, importing documents, using nodes (containers for nvivo data), coding to organize the data, tools to analyze data, and exporting a summary of your coding structure to word and excel for sharing with others.” finally, i looked back at the initial website search results and saw that there was a separate libguide that discusses qualitative methodology more generally (qualitative research resources for research rigor & transparency library guides at brown university, n.d.). taking all of these resources into account, i scored the library at brown university as offering tier 4 support. results of the 128 members of the association of research libraries, 43, or 33%, offer some form of support for qualitative data analysis (qda). while this is a smaller number than the 53 research guides that had been identified by swygart-hobaugh in her 2016 study, her sample came from academic libraries of all sizes across north america. of the 43 arl libraries that include information about qda, all of them offer some form of web-based support, most often in the form of a libguide. nineteen libraries offer tier 4 support for qualitative data analysis, which includes either workshops, consultations or one-on-one training sessions. eighteen of the libraries offer tier 3 support, which consists of general information about qualitative data analysis, as well as at least some softwarespecific guidance. three libraries each offer tier 2 or tier 1 support of qda. one library offers regular workshops on the use of nvivo, but otherwise provides no web-based research guide support. see figure 1. 6/8 pival, paul r. (2025). support for computer-assisted qualitative data analysis software in arl libraries, iassist quarterly 49(1), pp. 1-8. doi: https://doi.org/10.29173/iq1144 figure 1 arl libraries offering qda support nvivo is by far the most frequently supported caqdas title, being mentioned by 37 libraries. atlas.ti, maxqda, and dedoose all appear to be supported by at least one library. taguette, an open-source tool, is mentioned by 4 libraries. these frequencies fall in line with the literature in other fields using caqdas (woods et al., 2016), (paulus, 2023), (o’kane & smith, 2022). discussion as pointed out by swygart-hobaugh in her research (2016), “a recurring theme was that data support services should be guided by the local needs of the institution’s researchers, and thus the primary focus for data support should be either quantitative or qualitative, depending on the predominant need” (p. 167). the fact that caqdas support is offered at only one third of arl libraries may be evidence that this type of support is not required at every institution, though that seems difficult to believe for the schools that make up arl membership. or it may be evidence of a continued lack of expertise amongst librarians, as suggested by hansman (2015). many of the campuses in which arl libraries reside do offer support of qda software through their information technology, or other departments. however, since i am interested explicitly in library support, i have not included those campuses in the numbers above. it may be useful for future studies to examine caqdas support across the entire institution, rather than limit to library support. the fact that of the libraries that do support qualitative data analysis, a large majority of them offer support at a high level suggests that there is a strong level of expertise within arl libraries. for those at institutions currently lacking this expertise, there is a solid core of librarians who could teach or be tapped for peer learning opportunities. tier 1 7% tier 2 7% tier 3 42% tier 4 44% arl libraries offering qda support 7/8 pival, paul r. (2025). support for computer-assisted qualitative data analysis software in arl libraries, iassist quarterly 49(1), pp. 1-8. doi: https://doi.org/10.29173/iq1144 conclusions to revisit the original question of learning whether librarians supporting computer-assisted qualitative data analysis software are outliers, the data suggest that, at least amongst arl-member libraries, we are somewhat, but not entirely. librarians wishing to offer these types of services have colleagues upon whom they can draw or could explore other departments on campus to steer students towards. future studies could expand the scope of this survey to other geographies or types of libraries. perhaps support of caqdas is more prevalent outside of north american research libraries. perhaps it clusters in the libraries of universities supporting specific disciplines. or, perhaps the above results are accurately indicative of yet another niche service offered by your friendly neighborhood librarian. references cain, j., cooper, l., demott, s., & montgomery, a. (2019). where is qda hiding? an analysis of the discoverability of qualitative research support on academic library websites. iassist quarterly, 43(2), 1–9. https://doi.org/10.29173/iq957 cooper, l. (2022). creating library services to support qualitative researchers. qualitative and quantitative methods in libraries, 11(3), 573-585. hagman, j., & bussell, h. (2022). going qual in: towards methodologically inclusive data work in academic libraries. iassist quarterly, 46(2), 1–15. https://doi.org/10.29173/iq1022 hansman, c. a. (2015). training librarians as qualitative researchers: developing skills and knowledge. the reference librarian, 56(4), 274–294. https://doi.org/10.1080/02763877.2015.1057683 kang, g., & sinn, d. (2024). technology education in academic libraries: an analysis of library workshops. the journal of academic librarianship, 50(2), 102856. https://doi.org/10.1016/j.acalib.2024.102856 libguides—content management and curation platform for libraries. (n.d.). retrieved october 30, 2024, from https://www.springshare.com/libguides/ list of arl members—association of research libraries. (n.d.). retrieved october 30, 2024, from https://www.arl.org/list-of-arl-members/ michalovich, a. (2022). graduate students’ modes of engagement in computer-assisted qualitative data analysis. international journal of social research methodology, 25(2), 247–260. https://doi.org/10.1080/13645579.2021.1879359 o’kane, p. m., & smith, a. d. (2022). computer aided/assisted qualitative data analysis software in management and organizational research. academy of management proceedings, 2022(1), 10636. https://doi.org/10.5465/ambpp.2022.10636abstract 8/8 pival, paul r. (2025). support for computer-assisted qualitative data analysis software in arl libraries, iassist quarterly 49(1), pp. 1-8. doi: https://doi.org/10.29173/iq1144 paulus, t. m. (2023). using qualitative data analysis software to support digital research workflows. human resource development review, 22(1), 139–148. https://doi.org/10.1177/15344843221138381 qualitative data | nnlm. (n.d.). retrieved october 30, 2024, from https://www.nnlm.gov/guides/data-glossary/qualitative-data qualitative research—resources for research rigor & transparency—library guides at brown university. (n.d.). retrieved february 8, 2025, from https://libguides.brown.edu/c.php?g=553177&p=8156537 røddesnes, s., faber, h. c., & jensen, m. r. (2019). nvivo courses in the library. nordic journal of information literacy in higher education, 11(1), 27–38. https://doi.org/10.15845/noril.v11i1.2762 swygart-hobaugh, m. (2016).qualitative research and data support: the jan brady of social sciences data service?, in kellam, l. and thompson, k. (eds) databrarianship: the academic data librarian in theory and practice (pp. 153–178). association of college and research libraries (acrl). wilkinson saldaña, c. (n.d.). home—nvivo—library guides at brown university. retrieved february 8, 2025, from https://libguides.brown.edu/nvivo woods, m., paulus, t., atkins, d. p., & macklin, r. (2016). advancing qualitative research using qualitative data analysis software (qdas)? reviewing potential versus practice in published studies using atlas.ti and nvivo, 1994–2013. social science computer review, 34(5), 597–617. https://doi.org/10.1177/0894439315596311 endnotes 1 paul pival is director of emerging technologies at the university of calgary. he can be reached at: ppival@ucalgary.ca 1/27 breuer, johannes; al baghal, tarek; sloan, luke; bishop, libby; kondyli, dimitra; linardis, apostolos (2021) informed consent for linking survey and social media data, iassist quarterly 45(1), pp. 1-27. doi: https://doi.org/10.29173/iq988 informed consent for linking survey and social media data differences between platforms and data types johannes breuer1, tarek al baghal2, luke sloan3, libby bishop4, dimitra kondyli5, apostolos linardis6 abstract linking social media data with survey data is a way to combine the unique strengths and address some of the respective limitations of these two data types. as such, linked data can be quite disclosive and potentially sensitive, it is important that researchers obtain informed consent from the individuals whose data are being linked. when formulating appropriate informed consent, there are several things that researchers need to take into account. besides legal and ethical questions, key considerations are the differences between platforms and data types. depending on what type of social media data is collected, how the data are collected, and from which platform(s), different points need to be addressed in the informed consent. in this paper, we present three case studies in which survey data were linked with data from 1) twitter, 2) facebook, and 3) linkedin and discuss how the specific features of the platforms and data collection methods were covered in the informed consent. we compare the key attributes of these platforms that are relevant for the formulation of informed consent and also discuss scenarios of social media data collection and linking in which obtaining informed consent is not necessary. by presenting the specific case studies as well as general considerations, this paper is meant to provide guidance on informed consent for linked survey and social media data for both researchers and archivists working with this type of data. keywords informed consent, social media, surveys, data linkage 1. introduction social media data have been a popular subject of study in the social sciences (as well as various other scientific disciplines) for quite some time as they have become a part of everyday life for many people and are used for a variety of activities that are of interest to social scientists, such as communication, information seeking, news consumption, and relationship management. much of the research on the use and effects of social media data in the (quantitative) social sciences is based on survey data. when studying the use of media, however, several studies have shown that selfreports can be unreliable due to issues of social desirability or difficulties in recalling instances or patterns of usage (araujo et al., 2017; prior, 2009; scharkow, 2016). a way to assess social media use more reliably is to use data obtained directly from the platforms. the types of data available depend on the platform, and they can be collected in different ways (see the following section). notably, social media data can not only be used to study social media usage itself but also to investigate a variety of other topics, such as political communication or the formation and expression of opinions. https://doi.org/10.29173/iq988 2/27 breuer, johannes; al baghal, tarek; sloan, luke; bishop, libby; kondyli, dimitra; linardis, apostolos (2021) informed consent for linking survey and social media data, iassist quarterly 45(1), pp. 1-27. doi: https://doi.org/10.29173/iq988 while social media data have several advantages compared to survey data, they also have certain limitations. two important ones, especially for social-scientific research, are that they often lack indepth explicit information about the individuals, e.g., regarding their socio-demographic attributes or attitudes, as well as relevant outcome variables, such as voting or purchasing behaviour or offline forms of civic engagement. to combine the unique advantages and deal with their respective limitations, data from surveys and social media can be linked (stier et al., 2020). the linkage of surveys and social media data holds great potential and can be used to study a large variety of subjects (for a few examples, see the special issue ‘integrating survey data and digital trace data’ of the journal social science computer review).7 if researchers want to link surveys and social media data, there are several things they need to consider and address. one key issue is that of informed consent. as linked survey and social media data can be quite extensive, disclosive, and potentially also sensitive, obtaining informed consent is an important step in the process. while there can also be other legal bases for collecting and processing social media data for research, from an ethical perspective, obtaining informed consent is the preferable option for linking surveys and social media data (menchen-trevino, 2018). in this paper, we will discuss what researchers need to consider with regard to informed consent when they link surveys with social media data. following some general considerations, we present experiences and solutions from three case studies in which survey data were linked with data from 1) twitter, 2) facebook, and 3) linkedin. we will compare these different cases and highlight similarities as well as differences between the platforms that are relevant for obtaining informed consent from participants. we also discuss cases in which obtaining informed consent is not required. more broadly, we discuss what to consider when ingesting such data into repositories and provide guidance on what researchers should pay attention to with regard to informed consent if they want to link surveys and social media data and subsequently archive them via a data repository. accordingly, the considerations and suggestions in this paper are mostly targeted at researchers but are also relevant for staff at data archives who want to archive linked survey and social media data. 2. linking surveys and social media data there are two important factors that determine how survey data can be linked with social media data: 1) the type of social media data, and 2) the way(s) in which they were collected. social media data can come from a wide range of platforms with very different purposes and attributes. in addition, the same platform can provide various types of data. data from the platforms can include textual data (tweets, posts, comments, etc.), audio-visual material (images, video, etc.), network data (connections between users or content), or user profile information (name, location, occupation, etc.). similar to the types of data, the ways in which they are collected or acquired can also vary (see breuer et al., 2020). the most widely used approach is that researchers collect social media data themselves via application programming interfaces (apis) provided by the platforms or web scraping. however, they can also acquire data by entering into direct cooperation with the platforms or purchasing data from data resellers or market research companies. as an alternative to acquiring data via the platforms, researchers can also directly collaborate with users to collect social media data (see halavais, 2019; we will discuss this option in more detail in a later section). finally, it is also possible to reuse existing collections of social media data that have been created by other researchers and made available through data repositories or some other service. importantly, the https://doi.org/10.29173/iq988 https://journals.sagepub.com/toc/ssce/38/5 3/27 breuer, johannes; al baghal, tarek; sloan, luke; bishop, libby; kondyli, dimitra; linardis, apostolos (2021) informed consent for linking survey and social media data, iassist quarterly 45(1), pp. 1-27. doi: https://doi.org/10.29173/iq988 type of data and how they are acquired affects how they can be linked. for example, the terms of service (tos) of a platform api or contractual agreements with the platforms or data resellers may place restrictions on how the data can be used. there are different ways in which social media data can be linked with survey data (see stier et al., 2020). depending on the type of social media data and how they are acquired, they can be linked with survey data on the individual level or an aggregate level, and they can be collected together for the same units of observation (ex-ante linking) or separately and linked subsequently (ex-post linking that uses existing survey and/or social media datasets). within these types of data linking, different research designs are possible; for example, in the case of individual-level ex-ante linking, researchers can start with the survey and ask respondents to share or allow the collection of their social media data. likewise, they can also first collect social media data and then invite users whose data they have collected to participate in a survey. in both cases, informed consent for collecting or using people’s social media data and linking it with the survey data can be obtained as part of the survey. the type of social media data that is collected, as well as the way in which it is supposed to be linked to survey data, determine what the informed consent needs to look like. in general, the informed consent for linking survey data with social media data needs to be in accordance with relevant local legal regulations, such as the general data protection regulation (gdpr) in europe, and should satisfy relevant ethical standards as defined by institutional review boards (irb), ethics committees, or the ethical guidelines of scholarly societies. gdpr requires a legal basis for processing personal data and, when linking data, informed consent is the standard. notably, when survey and social media data are linked, at least during data collection, identities are known, so the data are always personal (and sometimes also sensitive, e.g., when they include information about religious or political beliefs) and, thus, gdpr applies. under gdpr, consent needs to be voluntary, informed, unambiguous, specific, and a clear affirmative action. ‘passive’ consent, for example, the use of pre-ticked boxes, is not acceptable. consent forms need to be in language suitable for the intended audience. participants should be informed about: how any personal data collected about them will be used, stored, processed, transferred, who the data controller is (and their contact details), the legal grounds and purpose of the processing, any recipients of the personal data, the period of retention and their rights (including that they can complain to the supervisory authority; see, e.g., the ukds gdpr guidelines8). these consent requirements are similar, but not identical, to the ethical requirements of many ethical review bodies. an ethical review will typically also require addressing additional issues, such as the participation of children or vulnerable people. finally, the formulation of the informed consent also needs to take into account the characteristics of the social media platform and the specific type(s) of data that should be linked with survey data. in the following section, we will focus on individual-level linking that starts with the survey. however, many of the considerations regarding the platform attributes and what implications these have for obtaining informed consent are also applicable to other kinds of social media data and linking approaches. https://doi.org/10.29173/iq988 https://www.ukdataservice.ac.uk/manage-data/legal-ethical/gdpr-in-research.aspx 4/27 breuer, johannes; al baghal, tarek; sloan, luke; bishop, libby; kondyli, dimitra; linardis, apostolos (2021) informed consent for linking survey and social media data, iassist quarterly 45(1), pp. 1-27. doi: https://doi.org/10.29173/iq988 3. differences among social media platforms and data types that are relevant for informed consent there is a tendency to treat social media data (and platforms) as homogenous, and this extends into the literature on survey and social media data linkage. the assumption is that there are universal rules and protocols that can be applied to ensure informed consent for data linkage, but this is true only to a limited extent. the platforms have different purposes, the data are structured differently, the data are collected in different ways, and ascertaining a unique identifier for a respondent on a platform is simple in some cases and complicated in others. there are also complications concerning what is actually considered public, as some platforms allow anyone (logged in or not) to view data, others require a researcher to have an account and to log in, and some may even require there to be a link (e.g., following, friendship, connection) between the respondent and researcher before any data can be viewed. beyond the technical questions of visibility and data access, users also have different expectations about how private specific types of information are on different platforms. given the potential disconnect between users’ views on the privacy and sensitivity of data and levels or ways of accessing platforms, it may be that, even though a respondent might consider their data to be ‘public’ and be happy to share it, the technological attributes of a platform can make that data difficult to access. to demonstrate this and discuss what it means for obtaining informed consent, we discuss three case studies of survey and social media platform data linkage covering three platforms: twitter, linkedin, and facebook. 3.1. case study 1: twitter the first case study on twitter data is based on two studies, one from the uk and one from germany. the uk study detailed in al baghal et al. (2019) draws upon three representative surveys of the british adult population: the british social attitudes survey 2015, the understanding society innovation panel 2017, and the natcen panel 2017. the design of the german study was different from the uk study in several regards. the participants in this study came from a non-probability web-tracking panel in which participants have agreed to have their browsing behavior tracked. the panel is maintained by a professional market research company. for a project with a methodological interest in questions of data linking and a substantive interest in online news consumption, researchers purchased access to the web tracking data for one year. the participants of this panel were invited to different online surveys. in the first of these, those who reported having a personal twitter account were asked for consent to link their twitter data to their survey responses. public or private? if we consider social media platforms to sit on a continuum with ‘public’ at one end and ‘private’ at the other, then twitter is quite firmly at the ‘public’ end. notwithstanding debates about the ‘imagined audience’ of a tweet (marwick and boyd 2010), twitter is a broadcast medium through which tweets can be viewed by anyone. they are visible via search engines and can be viewed without having to log in to the site. users can select to mark their tweets as protected, which means that their tweets are visible only to their followers (and users have to approve who these followers are), but these options are made clear to users meaning that the public/private nature of a tweet is well defined on this platform. https://doi.org/10.29173/iq988 5/27 breuer, johannes; al baghal, tarek; sloan, luke; bishop, libby; kondyli, dimitra; linardis, apostolos (2021) informed consent for linking survey and social media data, iassist quarterly 45(1), pp. 1-27. doi: https://doi.org/10.29173/iq988 specifics of informed consent sloan et al. (2020) discuss their procedure for gaining informed consent for survey and twitter data linkage. they identify five areas from singleton and wadsworth (2006) which need to be addressed: (a) why the data is being collected; (b) what will be done with it; (c) what is being collected; (d) secure data storage; and (e) maintaining anonymity. accordingly, they developed the following consent statement: (a) as social media plays an increasing role in society, we would like to know who uses twitter, and how people use it. (b) we are also interested in being able to add people’s, and specifically your, (c) answers to this survey to publicly available information from your twitter account such as your profile information, tweets in the past and in future, and information about how you use your account. (d) your twitter information will be treated as confidential and given the same protections as your interview data. (e) your twitter username, and any information that would allow you to be identified, will not be published without your explicit permission. sloan et al. (2020, p. 65) any consent statement needs to address the specific types of data that a platform generates, using terminology that users will understand. in the extract above, the statement mentions tweets, profile information, and information about how the platform is used. these three broad areas simplify the complexity that underlies twitter data. notably, when extracting data from the twitter api (see below), a single tweet can have over 150 attributes associated with it, covering everything from the content of the tweet itself to the number of followers the user has and various measures of geographical location. it is also not possible to explain the complexity of the analysis that this linked data will be subjected to. sloan et al. (2020) acknowledge that there is a compromise here between complete information and the need to provide a practical and comprehensible explanation that enables participants to make an informed decision. further information was provided in a series of help screens that participants could access if needed, covering: what information will you collect from my twitter account? what will the information be used for? who will be able to access the information? what will you do to keep my information safe? what if i change my mind? the language of the consent statement in the german study was based on the one developed by sloan et al. (2020). the text was translated into german and slightly adapted to reflect the design and purpose of the study. still, the wording is very similar to that used by sloan et al. (2020). what https://doi.org/10.29173/iq988 6/27 breuer, johannes; al baghal, tarek; sloan, luke; bishop, libby; kondyli, dimitra; linardis, apostolos (2021) informed consent for linking survey and social media data, iassist quarterly 45(1), pp. 1-27. doi: https://doi.org/10.29173/iq988 was different in this study compared to sloan et al. (2020) was that the more detailed information about the data collection and handling was not provided via additional info screens but on a separate website that was linked in the consent statement in the online survey. the full text that was presented on that website is included as an appendix for this paper. accordingly, the consent statement in the german study was the following (note: we translated the german text into english, trying to be as literal as possible with our translation): since social media play an increasingly important role in society, we would like to know who uses twitter and how people use twitter. we are also interested in combining the answers from people, and also your responses from the survey with publicly available information from your twitter account. would you be willing to provide us with your twitter username for this research project so that we can link your twitter data with your responses from this survey for scientific purposes? of course, your data will be treated confidentially and not used for commercial purposes. your twitter name will not be mentioned in any publication and all twitter data will be protected by us with the same care as the data from the survey. you can find more information on how we process the data here [link to website with information]. another feature of the consent statement for twitter and survey data linkage is the need to specify that consent is being given to collect both historic and future data. as a microblogging platform, twitter is not static, and the platform encourages frequent interaction with other users and continuous production of content what edwards et al. (2013) describe as locomotive. because of the fast turnover of information, it is important that respondents are given a cue to consider their past behaviour and published content on the platform. some users will have tweets going back years, and, unlike a biographical platform such as linkedin where users are encouraged to keep their profiles current, twitter users are unlikely to monitor or regulate their past activity. unique identifiers unique identifiers are essential for the data linkage process as they allow the researcher to identify an individual user on a given social media platform in an unambiguous manner. while this may seem obvious, there are two related issues : 1) researchers must know what this identifier should be, and 2) when working with the unique identifiers, measures to protect participant privacy need to be taken. for twitter data, the question of what the unique identifier should be is easy to answer. twitter usernames are unique to each user, and the user can specify what this username should be. the username is often referred to as a twitter handle, and they are the mechanism through which people tweet each other (a mention), and can be used as an alternative to a phone number or email address when logging into the site. it is reasonable to expect a survey respondent to know what their username is, although recall ability and accuracy will be determined by how heavily they use the platform and when they last logged in. https://doi.org/10.29173/iq988 7/27 breuer, johannes; al baghal, tarek; sloan, luke; bishop, libby; kondyli, dimitra; linardis, apostolos (2021) informed consent for linking survey and social media data, iassist quarterly 45(1), pp. 1-27. doi: https://doi.org/10.29173/iq988 al baghal et al. (2020) detail the questions used in the same group of studies, as discussed by sloan et al. (2020) above. the version used in the understanding society innovation panel 2017 is as follows: what is your twitter username (e.g. @usociety)? soft check: twitter username does not begin with ‘@’ or contain spaces ‘please check and amend. twitter usernames should begin with an @ character and should not contain any spaces.’ the use of the @ symbol on twitter is the universal standard for addressing a user. therefore, having an @ in the prompt for their username further clarifies what is required and should be understood by any twitter user. the further check of ensuring there are no spaces is intended to avoid respondents confusing their username with their twitter name (which is normally the actual name of a user). again, in the german study, the language was quite similar: please enter your twitter username (e.g. @gesis.org) into the free-text field. my twitter username is: @__________ the instructions, as well as the additional soft check in the uk study, illustrate what can go wrong when linking survey and twitter data via the username. people may misspell their usernames or even (intentionally or unintentionally) provide a handle that is not theirs. this happened in both of the studies that this case study is based on, meaning that the linkage of survey and twitter data failed in these cases. to minimize data loss due to typos or the provision of a wrong username, one solution can be to have participants follow and/or send a direct message to a twitter account created by researchers for the purpose of the study. while it is helpful to remind respondents of the expected format, at least in the german study, some respondents may have been confused by the @ symbol as they entered their email address instead of their username. this confusion is even more understandable when considering that an email address is what many users use to log into their twitter account. another consideration is that usernames can change. hence, if a substantial amount of time passes between obtaining informed consent and the username and collecting the data, the username may have changed. it is also possible that accounts are deleted in the meantime. one way to address the issue of changing user names is to obtain the user id based on the username via the twitter api. unlike the username, the user id is persistent. to increase data privacy, twitter usernames should only be used as unique identifiers when necessary. for the linking process, using a unique generic id is preferable. in addition, the full survey and twitter data should be kept separate. sloan et al. (2020) present a workflow that ensures that there is no linked dataset that contains the full survey data and the full twitter data. the german study went further and included an explicit reference to this on the website containing the extended information on the collection and use of the twitter data (the full text can be found in appendix a): https://doi.org/10.29173/iq988 8/27 breuer, johannes; al baghal, tarek; sloan, luke; bishop, libby; kondyli, dimitra; linardis, apostolos (2021) informed consent for linking survey and social media data, iassist quarterly 45(1), pp. 1-27. doi: https://doi.org/10.29173/iq988 only information that is no longer personally identifiable (e.g., how often you tweet, how often you address political issues on twitter, etc.) is linked to the survey data. data access except when the username has changed, a researcher can easily identify an individual user profile through searching on the twitter website, using a search engine, or via the twitter apis. the latter method is widely used by researchers, and there are all manner of tools developed for researchers that draw on the apis, allowing researchers to access historical data (the rest api) or current data (the stream api). when collecting data via the twitter api based on usernames, data for protected accounts cannot be collected via the api. a potential alternative method for collecting twitter data that also allows accessing data for protected accounts is to have participants export their personal twitter data archive (which is an option available via the twitter account settings) and share it with the researchers. these data are not limited by the limitations of the api, which restricts the amount of historical data that can be accessed. however, this method of data donation (which we will discuss again for the next case study) means more effort for the participants and requires a safe solution for transferring the data to the researchers. rights to the data another consideration that needs to be made is the question of who has the rights to the data. what is important to note here is that none of the authors of this paper are lawyers, so what we say here as well in the corresponding sections for the other two case studies are our personal views based on our experience as researchers and/or data archives personnel and should not be taken as legal advice. there are other sources that provide legal opinions on matters related to the use of social media data. one such example is the expert opinion included in the report on ‘big data in social, behavioural, and economic sciences’ by the ratswd [german data forum] (2020). while its focus is on web scraping, it also includes a short section on the ‘binding effect of the twitter api terms of use’. also, while the expert opinion was written for the german case, it includes several sections discussing eu law, including the gdpr. in general, if social media data are collected via apis, their terms of service (tos) are an important thing to consider when assessing what can be done with the data. notably, tos can be somewhat open to interpretation, especially for the case of academic research. while this is also not based upon legal expertise, a blog post by justin littman provides a good breakdown of ‘twitter’s developer policies for researchers, archivists, and librarians’.9 one aspect on which the twitter api tos and developer policies place restrictions is the sharing of the data. notably, even when twitter data are linked with survey data and informed consent is obtained via the survey, the twitter data collected via the api are observed through a platform owned by a commercial company rather than directly provided by the individuals (as would be the case in a data donation scenario; see the next case study). this means that platform tos and developer policies need to be considered by researchers and archivists when deciding how the data can be used and shared. https://doi.org/10.29173/iq988 https://medium.com/on-archivy/twitters-developer-policies-for-researchers-archivists-and-librarians-63e9ba0433b2 https://medium.com/on-archivy/twitters-developer-policies-for-researchers-archivists-and-librarians-63e9ba0433b2 9/27 breuer, johannes; al baghal, tarek; sloan, luke; bishop, libby; kondyli, dimitra; linardis, apostolos (2021) informed consent for linking survey and social media data, iassist quarterly 45(1), pp. 1-27. doi: https://doi.org/10.29173/iq988 data sharing the twitter developer policies state that data accessed via the twitter apis cannot be shared in full with third parties. most importantly, one of the requirements is that only the tweet ids can be shared (not the tweet text or the associated metadata). hence, if researchers archive twitter data, they typically only archive tweet ids (see kinder-kurlanda et al., 2017). in their faq10, the uk data service also lists this as a requirement for depositing twitter data. if other researchers want to reuse the data, they need to collect the tweets again based on the list of tweet ids; a process called rehydration. of course, tweets and accounts can be deleted. thus, while the use of tweet ids and rehydration respects the users’ ‘right to be forgotten’, it reduces the reproducibility of research findings. an alternative to sharing tweet ids is to only share derived data. of course, while this increases privacy protection, this option somewhat limits the reproducibility of findings based on such data as well as their potential reuse value. only sharing derived data is the solution employed by the german study, which is described in the extended information on the collection and processing of the data: in accordance with the general terms and conditions of twitter, we will not publish the data or pass it on to third parties. only features derived from the data without any personal reference may be shared with other scientists under certain circumstances (e.g., which topics you are particularly interested in, how active you are on twitter). we will never pass on information to third parties by which you can be directly personally identified. twitter also limits the number of tweet ids (and user ids) that can be shared but makes an exception for academic research: academic researchers are permitted to distribute an unlimited number of tweet ids and/or user ids if they are doing so on behalf of an academic institution and for the sole purpose of non-commercial research. for example, you are permitted to share an unlimited number of tweet ids for the purpose of enabling peer review or validation of your research. (twitter, 2020) the openness of twitter in supporting academic studies is significant, and such allowances demonstrate an understanding of the needs of the research community by addressing issues concerning transparency and replication. 3.2. case study 2: facebook data the second case study is based on the german project described in the previous case study. in the second online survey within that project, respondents who reported having a personal facebook account were asked to install and use a browser plugin that collects public posts (as well as some metadata on them, such as the number of likes and other reactions they have received) from the users’ personal facebook feeds. hence, what was collected was not content produced by a user but content by other sources (e.g., media outlets or other organizations) that the user is exposed to. the browser plugin was available for the desktop version of the chrome and firefox browsers and could be installed via the official plugin stores. a detailed description of the plugin and its use can be found in haim and nienierza (2019). https://doi.org/10.29173/iq988 https://www.ukdataservice.ac.uk/help/faq/deposit.aspx#socialmedia 10/27 breuer, johannes; al baghal, tarek; sloan, luke; bishop, libby; kondyli, dimitra; linardis, apostolos (2021) informed consent for linking survey and social media data, iassist quarterly 45(1), pp. 1-27. doi: https://doi.org/10.29173/iq988 public or private? coming back to the hypothetical continuum between private and public for social media data, data from facebook is more on the private end of this spectrum. while twitter is generally meant and used for public communication, facebook is more often used for personal communication. on the technical side, unless a user profile is public which, unlike twitter, is not the default case their status update and profile information can only be seen by their facebook friends. although the data that the browser plugin collected public posts from a user’s news feed can be considered less sensitive than posts made by the users themselves, they are private in the sense that only the users can access their personal facebook news feed. specifics of informed consent given that people generally consider facebook data to be private and sensitive and, because the installation and use of the browser plugin required more effort than the provision of the twitter handle in the first survey, the consent statement for the facebook data was a bit more detailed: for many people, facebook is an important source of information. as you probably know, the display of news items on facebook is highly personalized. since facebook provides virtually no information about this, it is unclear how this selection is made. as independent scientific researchers, we are interested in how the personalized display of messages on facebook works. to this end, we cooperate with researchers who have developed a browser plugin (for firefox and chrome) that collects public posts in the news feed of individual users. we would like to link the data we already have from the survey and web tracking with data on the public posts in your facebook news feed. would you be willing to install this browser plugin? the plugin only records posts from your news feed that have actually been publicly shared on facebook and can, therefore, be seen by any facebook user. private posts, such as status updates from friends or private messages, are not recorded. login codes and passwords are also not recorded. in addition, you can view the data collected from your news feed at any time and delete it if necessary. you can find more detailed information on data protection for the browser plugin here [link to a website with information]. similar to the consent statement in the online survey, the information presented on the linked website was also a bit more extensive (see appendix b). unique identifiers as facebook user names are not unique (in most cases, people use their real names for their facebook profiles) and because the data were not collected via the facebook api (see the following section on this issue), a different unique identifier was needed to link the facebook data with the survey responses. for that reason, participants were asked to generate a six-digit code in the survey: first letter of mother’s first name, first letter of father’s first name, first letter of own first name, day from date of birth, last letter of own hair colour, last letter of own eye colour. to create the link, https://doi.org/10.29173/iq988 11/27 breuer, johannes; al baghal, tarek; sloan, luke; bishop, libby; kondyli, dimitra; linardis, apostolos (2021) informed consent for linking survey and social media data, iassist quarterly 45(1), pp. 1-27. doi: https://doi.org/10.29173/iq988 participants had to enter the code again as part of the installation process for the browser plugin. data access as described at the beginning of this subsection, a browser plugin that the participants had to install was used to collect the facebook data. the plugin only collects current data, so it is not possible to access historical data. users can also deactivate the plugin and delete data that has been collected with the browser plugin. the use of the browser plugin was necessary in this study for two reasons: 1) it is the only way to directly capture exposure to content on facebook via the news feed, and 2) data access via the facebook api has essentially become unavailable to academic researchers as a consequence of the cambridge analytica scandal. as platform providers can substantially alter or even completely close apis at any time, some researchers have argued that research with social media data may be facing an ‘apicalypse’ (bruns, 2019) or entering a ‘post-api age’ (freelon, 2018). asking users to install and use a browser plugin to collect facebook data is one way of partnering with users to address this issue (see halavais, 2019). another option is a data donation model in which users export parts of their personal facebook data archives and share them with researchers (see thorson et al., 2019 for an example). mancosu and vegetti (2020) have also suggested a web scraping routine for collecting public facebook data. rights to the data while privacy is less of a concern for public facebook posts that cannot be directly associated with the user in whose news feed they appeared, a legal issue that needs to be considered for these data is copyright. as many of the public posts in users’ news feeds come from media outlets or companies, many of them are protected by copyright. data sharing the fact that many of the captured posts are likely protected by copyright means that the full raw data cannot be easily shared. to increase data privacy, the survey data should only be linked with data derived from the posts, such as counts of different types of posts. while the users from whose feeds the posts were collected cannot be directly identified from these data, the issue of copyright, as well as the fact that identification of users cannot be ruled out completely, means that the raw data cannot be shared freely. for those reasons, the part on data access and sharing in the extended information document read as follows: the anonymized (aggregated) linked data, which includes your survey responses and web tracking data as well as information on public posts from your facebook news feed, is used for scientific purposes only. commercial use of the data is excluded. access for third parties to the complete linked data will only be possible in a special secure environment. 3.3. case study 3: linkedin data the study on consent to link survey responses and linkedin data is being conducted during the fourteenth wave of the understanding society innovation panel (ip). to the best of our knowledge, https://doi.org/10.29173/iq988 12/27 breuer, johannes; al baghal, tarek; sloan, luke; bishop, libby; kondyli, dimitra; linardis, apostolos (2021) informed consent for linking survey and social media data, iassist quarterly 45(1), pp. 1-27. doi: https://doi.org/10.29173/iq988 this is the first survey asking for linkedin linkage consent. understanding society has a focus on measuring labour market activity, and linkedin focuses on employment and businesses, being used largely as a professional networking site. in terms of scale, a recent survey in the uk by regulator ofcom (2019) found that 16% of uk internet users used linkedin; however, its employment focus means users are mostly a subset of the population who are or would like to be economically active. about half of the uk population (based on understanding society data) is employed, suggesting that linkedin coverage of its target population could be higher than twitter is for the population it targets (~25% of internet users in the uk use twitter). linkedin is what edwards et al. (2013) call punctiform – ‘[it] capture[s] the structure of social relations at particular moments and [is] therefore ‘punctiform’ in providing a snapshot of these relations.’ interestingly, edwards et al. (2013) originally classified all social media as locomotive, and defined social media data, by definition, as not being punctiform; but, when comparing the information turnover and purpose of linkedin with a microblogging platform, such as twitter, it is, indeed, quite static by comparison. linkedin, as a biographical profile site, does fit the description of being a snapshot of a user’s career status. public or private? on the private-public spectrum, linkedin is perhaps the most public of all social media sites. the main purpose of the site is to network professionally, including looking for new business and employment opportunities. having a private profile would naturally limit that objective. moreover, this public nature of the profile has been recognized on a legal basis. recent us litigation determined that such scraping was indeed legal (woollcott, 2019; also see mancosu & vegetti, 2020), partly based on the understanding that linkedin profile data is owned by the users and that user profiles are public for the purpose of being accessed by others. specifics of informed consent given that if the profile is made public, it can be accessed and data scraped directly, the initial need to obtain consent is for ethical considerations. as the linkedin project grew out of the uk project on twitter, the specifics of informed consent are based almost entirely on that project. there was a focus on the same five areas addressed with twitter and facebook for informed consent, and the language was similarly based on that developed by sloan et al. (2020). the main changes were on being more linkedin-specific, including a focus on employment and education content. accordingly, the language for the linkedin consent is as follows: we would like to know who uses linkedin, and how people use it. we are also interested in being able to link the information people have provided for this study to publicly available information from their linkedin accounts, such as their employment or education history, their connections, or information about their employer. information collected from your linkedin account will be treated as confidential and protected in the same way as your interview data. any linkedin information that would allow you to be identified will not be published. https://doi.org/10.29173/iq988 13/27 breuer, johannes; al baghal, tarek; sloan, luke; bishop, libby; kondyli, dimitra; linardis, apostolos (2021) informed consent for linking survey and social media data, iassist quarterly 45(1), pp. 1-27. doi: https://doi.org/10.29173/iq988 are you willing to tell me the name of your personal linkedin account and for your linkedin information to be linked with the information you have provided for this study? the additional help text included with this question provides information regarding what is being asked and ensures greater informed consent. this includes information on what data will be collected and why, who will have data access, and data security procedures. again, this is largely based on the wording developed for the twitter study. besides changing the focus to linkedin information, the main difference with what was provided when asking to link twitter data is the inclusion of a statement about gdpr, which came into effect after the uk twitter study. full wording for these help links is included in appendix c. unique identifiers equally important when asking consent, however, is the need for additional data to be collected from the respondent to identify the correct linkedin profile to link to survey responses. unlike usernames on twitter, linkedin user ids are largely not chosen by (and unknown to) individuals. when a user signs up, the site assigns a user id based on the person’s first and last name with an alphanumeric string appended (e.g., first-last-81341b34). these can be customized by users, but many users do not. after obtaining consent to link the data, survey questions can ask for this id (as would be the case in twitter linkage), but most respondents will not be able to provide an answer. rather, for most respondents, additional questions need to be asked to identify the correct linkedin profile from which to scrape data and link to survey responses. these can only viably be asked after consent has been obtained. it is possible to employ programming scripts (written, e.g., in python or r) to search for profiles automatically using linkedin’s search functionality. to limit search returns, and to correctly identify the respondent’s profile, additional information about the linkedin profile needs to be collected. this information needs to include, at a minimum, the name the respondent has on their linkedin profile, but more information is needed to limit returns to the most likely matches to the respondent. another obvious identifier would be an employer listed on linkedin. however, the ability of these two fields to be limiting may be lacking, depending on the uniqueness of the name and employer combination. for example, ‘bill gates microsoft’ returns only one profile. however, ‘tom smith tesco’ returns 126 profiles. additional questions about the profile should therefore be included but should be focused on what is likely included on profiles for most while avoiding overburdening respondents, especially given that all of the information requested is personal identifiers. an initial set of possible questions are included in the fourteenth wave of the understanding society innovation panel (ip). in addition to profile id (if known) and the respondent’s name and most recent place of work listed on the profile, consenting respondents are asked for their profile job title, location, and most recent place of education listed. data access collecting user data from linkedin to link to survey data is not as simple as it is from twitter, as access to linkedin apis is largely closed to research. however, collecting linkedin data also does not https://doi.org/10.29173/iq988 14/27 breuer, johannes; al baghal, tarek; sloan, luke; bishop, libby; kondyli, dimitra; linardis, apostolos (2021) informed consent for linking survey and social media data, iassist quarterly 45(1), pp. 1-27. doi: https://doi.org/10.29173/iq988 require an additional plugin and user login, as in the case of facebook. rather, researchers can collect linkedin data directly from the website using established data scraping techniques (haag, 2020) in programming languages such as python or r. the set of identifiers provided by respondents is used in the linkedin search function, which returns a set of one or more profiles. given the lack of easily obtainable unique user identifiers and the need to scrape web pages, there may be multiple returns on search results using the set of identifiers collected after the initial consent question. to make these matches, we propose two methods. the first, deterministic linkage, requires exact matches on identifiers. these can include cases where the respondent knows their linkedin id or where the identifiers provided yield only one return. however, in some instances, an exact match may not be possible, for example, due to entry errors or where multiple returns exist on the set of identifiers used. therefore, in these cases, we utilise probabilistic linkage methods that identify likely matches with a quantified level of uncertainty. probabilistic linkage involves linking data based on statistical techniques that calculate from nonunique identifier sets the likelihood of links between records in each data source being correct, given the other links possible between records (sayers et al. 2016; doidge & harron 2019). for each sample member, the most likely link (determined from linkage weights computed for each considered possible link) should be included in the linked dataset, although a similarity threshold below which ‘best’ links are considered incorrect is often applied to reduce linkage errors in the dataset. hence, such methods are particularly useful for linking records when, due to entry errors and other sources of differences (for example, the university of essex will not match with essex university when deterministic linkage methods are used), identifiers for given subjects may be mismatched between data sources. rights to the data given that the data is being scraped directly from websites and not through linkedin’s api, considerations regarding the tos of the api do not need to be factored in. further, a legal precedent suggests that data on public profiles is open to all. however, that legal case was in the united states and may not hold if challenged in other contexts. additionally, linkedin posts may contain copyrighted material that needs to be considered in data collection and curation. data sharing again, unlike the twitter project, since linkedin data is not collected via the api, the situation in regards to data sharing is less clear. also, unlike the twitter project, the work on linkedin has not focused on plans for archiving or comprehensive data sharing. the focus, rather, has been on the data collection and linkage of linkedin and survey data; the amount of work and programming required is non-trivial, in part due to lack of access to the linkedin api. future expansions linking linkedin and survey data will place more efforts on ways to ensure efficient data sharing. however, some data sharing is planned to generate processes and possible next stages for work. as noted above, we explain to respondents who will have access to what data from this process. we note that data from survey answers and linkedin information will be made available to researchers if they are able to present a strong scientific case to ensure that the information is used responsibly https://doi.org/10.29173/iq988 15/27 breuer, johannes; al baghal, tarek; sloan, luke; bishop, libby; kondyli, dimitra; linardis, apostolos (2021) informed consent for linking survey and social media data, iassist quarterly 45(1), pp. 1-27. doi: https://doi.org/10.29173/iq988 and securely. we will also generate summary information from linkedin accounts, which would not allow identification and will have the same access controls as survey answers, which will be accessible by other researchers. 4. platform attributes relevant for informed consent as the case studies in the previous section have illustrated, social media platforms have specific attributes that are relevant for the formulation of informed consent. some of these features are the same or similar across platforms, whereas others differ. based on the case studies we have presented, the key similarities are: all three platforms offer different types of data that vary with respect to their (perceived) privacy and sensitivity, and all of them have complex data structures (whether extracted via api or scraping) that are too complicated to communicate to a lay audience, which means that informed consent will always be a compromise between a simplistic explanation of what the data is and how it will be used versus what the data actually are and how they will be actually used. despite these similarities, as discussed in the previous sections, there also are some clear differences between the platforms. table 1 presents the differences between the platforms we have considered that need to be taken into account for creating informed consent statements and providing appropriate information to participants. in contrast to the description of the case studies in the previous sections, this table focuses on the platforms and their attributes rather than specific methods of data collection. hence, while one of the comparison categories (unique identifiers) is the same, the others are different here. https://doi.org/10.29173/iq988 16/27 breuer, johannes; al baghal, tarek; sloan, luke; bishop, libby; kondyli, dimitra; linardis, apostolos (2021) informed consent for linking survey and social media data, iassist quarterly 45(1), pp. 1-27. doi: https://doi.org/10.29173/iq988 table 1. differences between twitter, facebook, and linkedin that are relevant for the formulation of informed consent twitter facebook linkedin private/public ● twitter is mostly used for public communication ● if user accounts are not protected, much of the data is publicly visible ● facebook data is generally considered more private ● unless users have public profiles ( most do not), their activities and full profile information are only visible to logged-in users with whom they are connected ● linkedin is used mostly for public professional networking and job search ● a us court decided public accounts are public-domain data, as the expectation is access by others dynamic nature of the content ● twitter content is dynamic and changing ● it is important to request access to historic and future data to get a fuller picture for individual users ● facebook content is highly dynamic and changing ● whether researchers can access historical or future data depends on the data collection method ● linkedin data is less dynamic and volatile as users build a profile that is reasonably stable ● it is not necessary to explicitly ask for historic data from linkedin users because the ‘live’ data is by definition historic unique identifiers ● user names are unique and can be used to link the data, but user names can change ● user ids are stable and can be accessed via the api with a list of usernames ● while there are user ids, these are usually not known to users ● other identifiers need to be used to link the data ● a unique alphanumeric id is assigned by the site, which can be customized ● it is unlikely for users to know their linkedin id, so there is a need to rely on other profile identifiers and to employ probabilistic linkage https://doi.org/10.29173/iq988 17/27 breuer, johannes; al baghal, tarek; sloan, luke; bishop, libby; kondyli, dimitra; linardis, apostolos (2021) informed consent for linking survey and social media data, iassist quarterly 45(1), pp. 1-27. doi: https://doi.org/10.29173/iq988 while we have covered three platforms that differ in several important regards in our case studies, there are many other types of social media data that can be linked with survey data. some of these types of data have properties with substantial implications for informed consent. to illustrate this, we will briefly discuss two such categories in the following section: aggregated social media data and social media data for figures of public interest. 5. data from persons of public interest and aggregated data the focus of the case studies presented in the previous section was on individual-level data for normal users of the platforms. however, beyond those presented in the case studies above and differing in several important regards, there are other types of users and forms of social media data that can be linked with survey data and also have implications for the issue of informed consent. the first type that we want to discuss here are social media data from figures of public interest or institutions. such data are often collected in the context of elections. for example, social media data collections for politicians and other relevant public actors (parties, public authorities, etc.) for the german federal elections in 2013 (kaczmirek and mayr, 2015) and 2017 (stier et al., 2018) have been published via the gesis data archive.11 as the politicians are figures of public interest, at least when they use their professional social media accounts, it is not necessary to obtain their informed consent. while the data can be considered personal, what is important to also keep in mind in this context is that informed consent is only one of the possible legal bases for processing such data according to gdpr. another one is a task carried out in the public interest, which is certainly something researchers can claim when studying the social media activities of politicians or other public actors in the context of elections. also, if the data are generated by institutions, such as public authorities, they are also typically not personal data. these criteria are also important for questions regarding the publication of social media data. for example, the decision flow chart for the publication of twitter communications by williams, burnap, and sloan (2017) suggests that tweets by organisations and public figures can generally be published. the second type of data is aggregated social media data from public figures that is published through other means than completed data collections available for download via a repository. the collection of social media data around federal elections in germany has since been converted into an ongoing project with the gesis social media monitoring.12 instead of providing completed collections for specific elections, this platform offers aggregated data for user-defined periods of time, topics, or types of actors. importantly, aggregated social media data can also be linked with individual-level survey data. in that case, there would be no one-to-one matching but a one-tomany-linking. examples could be to link survey data to data on the volume or sentiment of tweets about a specific topic for a certain region and period of time. of course, if aggregated data is used, it is not possible to gather informed consent for the linking from the individuals whose data was used to create the aggregate values. a service that is similar to the gesis social media monitoring in several regards is the social web observatory.13 the social web observatory is an initiative aiming to help researchers, mainly from the social sciences and digital humanities, to investigate information diffusion in the social web. the project aims to monitor various sources of information, such as websites and the most popular social https://doi.org/10.29173/iq988 https://www.gesis.org/en/services/finding-and-accessing-data http://mediamonitoring.gesis.org/ https://socialwebobservatory.iit.demokritos.gr/#/about https://socialwebobservatory.iit.demokritos.gr/#/about https://socialwebobservatory.iit.demokritos.gr/#/about 18/27 breuer, johannes; al baghal, tarek; sloan, luke; bishop, libby; kondyli, dimitra; linardis, apostolos (2021) informed consent for linking survey and social media data, iassist quarterly 45(1), pp. 1-27. doi: https://doi.org/10.29173/iq988 media platforms (facebook, instagram, twitter). users can gather data about different entities, such as politicians or other public actors, by using a wide variety of sources, such as keywords, hashtags, monitoring of websites. the material retrieved through a keyword search can be analyzed based on parameters that allow the extraction of indicators, such as the emergence of trends, emotions, attitudes about a phenomenon, event, or product (tsekouras et al., 2020). similar to the gesis social media monitoring, the data can also be aggregated over different time periods. as part of an informal collaboration between the clarin: el14 and sodanet15 infrastructures, members of the ekke / sodanet research team have set up entities to follow the campaign of political parties and candidates for both municipal and national elections in greece between may and july 2019 by providing information about their official facebook or/and twitter accounts, wikipedia pages, and relevant keywords. again, similar to the gesis social media monitoring, users cannot extract raw data from the social web observatory. instead, processed or aggregated data, such as the number of articles, comments, or tweets or information about the domains containing the articles and comments are provided. cases in which only aggregated data are used and shared are the second type of social media data collection that does not require informed consent from individuals. besides the social web observatory and the gesis social media monitoring, which are geared towards social scientists, there also are other continuous social media collections. one example of those is tweetskb16 (fafalios et al., 2018), which is a “corpus of anonymized data for a large collection of annotated tweets” that includes “metadata information about the tweets as well as extracted entities, sentiments, hashtags and user mentions” (description on the tweetskb website). all of the services presented here are data sources that can serve as alternatives to data collections via web scraping, apis, or data donation, as presented in the case studies. while researchers have no direct control over the actual data collection, these services can provide comprehensive data that can also be linked with survey data with the added benefit that the linking, in this case, does not require researchers to obtain informed consent from the individuals whose data are included in these collections. 6. conclusion the three case studies discussed in this paper provide examples of how informed consent for social media and survey data linkage can be obtained. however, there are clear differences in what information needs to be given to participants, depending on the platforms in use. social media platforms are not homogenous in the way that they are used by individuals, the purposes they serve, or the manner in which they are structured and interacted with, both by content creators and the wider public. accordingly, it is no surprise that it is difficult to provide concrete guidance on informed consent that can be applied to all platforms and types of data. this is further exacerbated by the fact that platforms can change or disappear, and new ones emerge. however, despite the fact that providing general solutions for informed consent for linking surveys and social media data is not possible, the cases and aspects we have discussed should serve as guiding points for researchers and archivists working with such data. it is worth noting that the informed consent process detailed for the twitter case study has been adopted and modified for later projects indicating that there is value in adapting the work of others. https://doi.org/10.29173/iq988 https://www.clarin.gr/en https://www.clarin.gr/en http://www.sodanet.gr/ http://www.sodanet.gr/ https://data.gesis.org/tweetskb/ 19/27 breuer, johannes; al baghal, tarek; sloan, luke; bishop, libby; kondyli, dimitra; linardis, apostolos (2021) informed consent for linking survey and social media data, iassist quarterly 45(1), pp. 1-27. doi: https://doi.org/10.29173/iq988 based on what we presented in the paper, some of the general recommendations for informed consent for linking surveys and social media data are to take into account and address what types of social media data are collected and by what means, how private and sensitive they are, how exactly they will be linked to the survey data, how they are stored and can be accessed, and whether current, future, or historic data are required and collected. acknowledgment the work of johannes breuer, libby bishop, dimitra kondyli, and apostolos linardis on this paper was funded by the consortium of european social science data archives (cessda) as part of the wp2020 project ‘new data types’. the work of luke sloan and tarek al baghal on this paper is associated with the funded esrc project ‘understanding [online/offline] society: linking surveys with twitter data’ (es/s015175/1). johannes breuer wants to thank pascal siegers and sebastian stier for their assistance in writing the informed consent (including the extended data privacy information) for the german study which case study 2 and parts of case study 1 in this paper are based on. references araujo, t. et al. (2017) ‘how much time do you spend online? understanding and improving the accuracy of self-reported measures of internet use’, communication methods and measures, 11(3), pp. 173–190. doi: https://doi.org/10.1080/19312458.2017.1317337. breuer, j., bishop, l. and kinder-kurlanda, k. (2020) ‘the practical and ethical challenges in acquiring and sharing digital trace data: negotiating public-private partnerships’, new media & society, 22(11), pp. 2058–2080. doi: https://doi.org/10.1177/1461444820924622. bruns, a. (2019) ‘after the “apicalypse”: social media platforms and their fight against critical scholarly research’, information, communication & society, 22(11), pp. 1544–1566. doi: https://doi.org/10.1080/1369118x.2019.1637447. doidge, j. c. and harron, k. (2018) ‘demystifying probabilistic linkage’, international journal of population data science, 3(1). doi: https://doi.org/10.23889/ijpds.v3i1.410. edwards, a. et al. (2013) ‘digital social research, social media and the sociological imagination: surrogacy, augmentation and re-orientation’, international journal of social research methodology, 16(3), pp. 245–260. doi: https://doi.org/10.1080/13645579.2013.774185. fafalios, p. et al. (2018) ‘tweetskb: a public and large-scale rdf corpus of annotated tweets’, in gangemi, a. et al. (eds) the semantic web. cham: springer international publishing, pp. 177–190. freelon, d. (2018) ‘computational research in the post-api age’, political communication, 35(4), pp. 665–668. doi: https://doi.org/10.1080/10584609.2018.1477506. german data forum (ratswd) (2020) ‘big data in social, behavioural, and economic sciences: data access and research data management’, ratswd output paper series. doi: https://doi.org/10.17620/02671.52. https://doi.org/10.29173/iq988 https://doi.org/10.1080/19312458.2017.1317337 https://doi.org/10.1080/19312458.2017.1317337 https://doi.org/10.1177/1461444820924622 https://doi.org/10.1177/1461444820924622 https://doi.org/10.1080/1369118x.2019.1637447 https://doi.org/10.1080/1369118x.2019.1637447 https://doi.org/10.23889/ijpds.v3i1.410 https://doi.org/10.23889/ijpds.v3i1.410 https://doi.org/10.1080/13645579.2013.774185 https://doi.org/10.1080/10584609.2018.1477506 https://doi.org/10.1080/10584609.2018.1477506 https://doi.org/10.17620/02671.52 https://doi.org/10.17620/02671.52 20/27 breuer, johannes; al baghal, tarek; sloan, luke; bishop, libby; kondyli, dimitra; linardis, apostolos (2021) informed consent for linking survey and social media data, iassist quarterly 45(1), pp. 1-27. doi: https://doi.org/10.29173/iq988 haag, f. (2020). ‘linkedin scraping with python’, medium, 28 february. available at: https://medium.com/federicohaag/linkedin-scraping-with-python-d8d14519602d (accessed: 28th may 2020) haim, m. and nienierza, a. (2019) ‘computational observation: challenges and opportunities of automated observation within algorithmically curated media environments using a browser plugin’, computational communication research, 1(1), pp. 79–102. doi: https://doi.org/10.5117/ccr2019.1.004.haim. halavais, a. (2019) ‘overcoming terms of service: a proposal for ethical distributed research’, information, communication & society, 22(11), pp. 1567–1581. doi: https://doi.org/10.1080/1369118x.2019.1627386. kaczmirek, l. and mayr, p. (2015). ‘german bundestag elections 2013: twitter usage by electoral candidates’. gesis data archive, cologne, za5973 data file version 1.0.0. doi: https://doi.org/10.4232/1.12319. kinder-kurlanda, k. et al. (2017) ‘archiving information from geotagged tweets to promote reproducibility and comparability in social media research’, big data & society, 4(2), p. 205395171773633. doi: https://doi.org/10.1177/2053951717736336. mancosu, m. and vegetti, f. (2020) ‘what you can scrape and what is right to scrape: a proposal for a tool to collect public facebook data’, social media + society, 6(3), advance online publication. doi: https://doi.org/10.1177/2056305120940703. marwick, a. e. and boyd, danah (2011) ‘i tweet honestly, i tweet passionately: twitter users, context collapse, and the imagined audience’, new media & society, 13(1), pp. 114–133. doi: https://doi.org/10.1177/1461444810365313. menchen-trevino, e. (2018). ‘digital trace data and social research: a proactive research ethics‘ in foucault welles, b. and gonzález-bailón, s. (eds.) the oxford handbook of networked communication. oxford: oxford university press, pp. 519–538. prior, m. (2009) ‘the immensely inflated news audience: assessing bias in self-reported news exposure’, public opinion quarterly, 73(1), pp. 130–143. doi: https://doi.org/10.1093/poq/nfp002. sayers, a. et al. (2016) ‘probabilistic record linkage’, international journal of epidemiology, 45(3), pp. 954–964. doi: https://doi.org/10.1093/ije/dyv322. scharkow, m. (2016) ‘the accuracy of self-reported internet use—a validation study using client log data’, communication methods and measures, 10(1), pp. 13–27. doi: https://doi.org/10.1080/19312458.2015.1118446. sloan, l. et al. (2020) ‘linking survey and twitter data: informed consent, disclosure, security, and archiving’, journal of empirical research on human research ethics, 15(1–2), pp. 63–76. doi: https://doi.org/10.1177/1556264619853447. https://doi.org/10.29173/iq988 https://medium.com/federicohaag/linkedin-scraping-with-python-d8d14519602d https://doi.org/10.5117/ccr2019.1.004.haim https://doi.org/10.5117/ccr2019.1.004.haim https://doi.org/10.1080/1369118x.2019.1627386 https://doi.org/10.1080/1369118x.2019.1627386 https://doi.org/10.4232/1.12319 https://doi.org/10.4232/1.12319 https://doi.org/10.1177/2053951717736336 https://doi.org/10.1177/2053951717736336 https://doi.org/10.1177/2056305120940703 https://doi.org/10.1177/2056305120940703 https://doi.org/10.1177/1461444810365313 https://doi.org/10.1093/poq/nfp002 https://doi.org/10.1093/poq/nfp002 https://doi.org/10.1093/ije/dyv322 https://doi.org/10.1093/ije/dyv322 https://doi.org/10.1080/19312458.2015.1118446 https://doi.org/10.1080/19312458.2015.1118446 https://doi.org/10.1080/19312458.2015.1118446 https://doi.org/10.1177/1556264619853447 https://doi.org/10.1177/1556264619853447 https://doi.org/10.1177/1556264619853447 21/27 breuer, johannes; al baghal, tarek; sloan, luke; bishop, libby; kondyli, dimitra; linardis, apostolos (2021) informed consent for linking survey and social media data, iassist quarterly 45(1), pp. 1-27. doi: https://doi.org/10.29173/iq988 stier, s et al. (2018). ‘social media monitoring for the german federal election 2017’, gesis data archive, cologne, za6926 data file version 1.0.0. doi: https://doi.org/10.4232/1.12992. stier, s. et al. (2020) ‘integrating survey data and digital trace data: key issues in developing an emerging field’, social science computer review, 38(5), pp. 503–516. doi: https://doi.org/10.1177/0894439319843669. thorson, k. et al. (2019) ‘algorithmic inference, political interest, and exposure to news and politics on facebook’, information, communication & society, advance online publication. doi: https://doi.org/10.1080/1369118x.2019.1642934. tsekouras, l. et al. (2020) ‘social web observatory: a platform and method for gathering knowledge on entities from different textual sources’, in proceedings of the 12th language resources and evaluation conference. marseille, france: european language resources association, pp. 2000–2008. available at: https://www.aclweb.org/anthology/2020.lrec-1.246 (accessed: 12th november 2020). twitter (2020) developer agreement and policy. available at: https://developer.twitter.com/en/developer-terms/agreement-and-policy (accessed: 12th november 2020). williams, m. l., burnap, p. and sloan, l. (2017) ‘towards an ethical framework for publishing twitter data in social research: taking into account users’ views, online context and algorithmic estimation’, sociology, 51(6), pp. 1149–1168. doi: https://doi.org/10.1177/0038038517708140. woollacott, e. (2019) ‘linkedin data scraping ruled legal.’ forbes, 10 september 2019. available at: https://www.forbes.com/sites/emmawoollacott/2019/09/10/linkedin-data-scraping-ruledlegal/#30bdd8311b54 (accessed: 26th october2020) endnotes 1 johannes breuer is a senior researcher at gesis – leibniz institute for the social sciences in germany and can be reached via email: johannes.breuer@gesis.org 2 tarek al baghal is senior research fellow and associate director of understanding society, questionnaire design, essex university uk and can be contacted at talbag@essex.ac.uk 3 luke sloan is deputy director of the social data science lab and professor at the school of social sciences, cardiff university uk. he can reached via email at sloanls@cardiff.ac.uk 4 libby bishop the coordinator for international data infrastructures in the data archive at gesis-leibniz institute for social sciences in germany and can be reached at elizabethlea.bishop@gesis.org 5 dimitra kondyli is a senior researcher at national centre for social research (ekke) – institute of social research in greece and can be reached via email: dkondyli@ekke.gr 6 apostolos linardis is a senior researcher at national centre for social research (ekke) – institute of social research in greece and can be reached via email: alinardis@ekke.gr 7 https://journals.sagepub.com/toc/ssce/38/5 https://doi.org/10.29173/iq988 https://doi.org/10.4232/1.12992 https://doi.org/10.1177/0894439319843669 https://doi.org/10.1177/0894439319843669 https://doi.org/10.1177/0894439319843669 https://doi.org/10.1080/1369118x.2019.1642934 https://doi.org/10.1080/1369118x.2019.1642934 https://doi.org/10.1080/1369118x.2019.1642934 https://www.aclweb.org/anthology/2020.lrec-1.246 https://www.aclweb.org/anthology/2020.lrec-1.246 https://developer.twitter.com/en/developer-terms/agreement-and-policy https://doi.org/10.1177/0038038517708140 https://doi.org/10.1177/0038038517708140 https://www.forbes.com/sites/emmawoollacott/2019/09/10/linkedin-data-scraping-ruled-legal/#30bdd8311b54 https://www.forbes.com/sites/emmawoollacott/2019/09/10/linkedin-data-scraping-ruled-legal/#30bdd8311b54 mailto:johannes.breuer@gesis.org mailto:talbag@essex.ac.uk mailto:sloanls@cardiff.ac.uk mailto:elizabethlea.bishop@gesis.org mailto:dkondyli@ekke.gr mailto:alinardis@ekke.gr https://journals.sagepub.com/toc/ssce/38/5 22/27 breuer, johannes; al baghal, tarek; sloan, luke; bishop, libby; kondyli, dimitra; linardis, apostolos (2021) informed consent for linking survey and social media data, iassist quarterly 45(1), pp. 1-27. doi: https://doi.org/10.29173/iq988 8 https://www.ukdataservice.ac.uk/manage-data/legal-ethical/gdpr-in-research.aspx 9 https://medium.com/on-archivy/twitters-developer-policies-for-researchers-archivists-and-librarians63e9ba0433b2 10 https://www.ukdataservice.ac.uk/help/faq/deposit.aspx#socialmedia 11 https://www.gesis.org/en/services/finding-and-accessing-data 12 http://mediamonitoring.gesis.org/ 13 https://socialwebobservatory.iit.demokritos.gr/#/about 14 https://www.clarin.gr/en 15 https://www.sodanet.gr/ 16 https://data.gesis.org/tweetskb/ https://doi.org/10.29173/iq988 https://www.ukdataservice.ac.uk/manage-data/legal-ethical/gdpr-in-research.aspx https://medium.com/on-archivy/twitters-developer-policies-for-researchers-archivists-and-librarians-63e9ba0433b2 https://medium.com/on-archivy/twitters-developer-policies-for-researchers-archivists-and-librarians-63e9ba0433b2 https://www.ukdataservice.ac.uk/help/faq/deposit.aspx#socialmedia https://www.gesis.org/en/services/finding-and-accessing-data http://mediamonitoring.gesis.org/ https://socialwebobservatory.iit.demokritos.gr/#/about https://www.clarin.gr/en https://www.sodanet.gr/ https://data.gesis.org/tweetskb/ appendix 1/27 breuer, johannes; al baghal, tarek; sloan, luke; bishop, libby; kondyli, dimitra; linardis, apostolos (2021) informed consent for linking survey and social media data, iassist quarterly 45(1), pp. 1-27. doi: https://doi.org/10.29173/iq988 appendix a website text with extended information on twitter data [project/study name] data protection information: twitter data your twitter data are collected by [name + address of institution] below you will find all information about our data collection that is relevant to you. you can contact us at the above address or via the email address [project email address] if you need more information about our research project. what information is collected about my twitter account? we will only collect information about your twitter account that is publicly available. this includes information about your account (such as your profile description, who you follow and who is following you), the content of your tweets (including text, pictures, videos, and links), and background information about your tweets (e.g. when you tweeted, what kind of device you used for it or provided you have enabled this feature the location from where you posted). we will collect information about your past tweets and will regularly update this information with current tweets for the duration of our study. what is this information used for? we use the data exclusively for scientific research. linking your twitter data with the survey data allows us to better understand your activities on the internet and your opinions. with additional data from social media we can... ● better understand who uses twitter and for what purposes. ● investigate whether twitter contains scientifically relevant information and how good the quality of this information is. ● identify topics that people are concerned about but which are not part of our surveys. ● gather information in addition to that from the survey to capture attitudes and opinions of the population. ● test assumptions about the relationship between the use of social media and political attitudes and behavior. what do you do to protect my personal information? all information is stored and used in accordance with the general data protection regulation (gdpr). since the information from twitter is publicly available, it is impossible to completely anonymize the collected data. only information that is no longer personally identifiable (e.g. how often you twitter, how often you address political issues, etc.) is linked to the survey data. in accordance with the general terms and conditions of twitter, we will not publish the data or pass it on to third parties. only features derived from the data without any personal reference may be shared with other scientists under certain circumstances (e.g. which topics you are particularly https://doi.org/10.29173/iq988 appendix 2/27 breuer, johannes; al baghal, tarek; sloan, luke; bishop, libby; kondyli, dimitra; linardis, apostolos (2021) informed consent for linking survey and social media data, iassist quarterly 45(1), pp. 1-27. doi: https://doi.org/10.29173/iq988 interested in, how active you are on twitter). we will never pass on information to third parties by which you can be directly personally identified. who will have access to the data? the anonymized linked data, which includes both your survey responses and your twitter information, will be used for scientific social research purposes only. commercial use of the data is excluded. access to the complete linked data will only be possible in a special secure environment. your rights you can withdraw your consent to the collection of your twitter data at any time. to do so, just send an email to [email address for the project] or a written letter to [name + address of the institute] please note that your twitter username must be mentioned in the email or letter, otherwise we cannot correctly assign your data for deletion. with regard to your personal data, you can make use of the following rights at any time: right of access to information right of rectification right to deletion (“right to be forgotten”) right to limit processing right to data transferability you also have a right of appeal to a data protection supervisory authority. contact person with all general questions and requests concerning data protection at [name of institution] you can contact: [name + address of data protection officer] https://doi.org/10.29173/iq988 appendix 3/27 breuer, johannes; al baghal, tarek; sloan, luke; bishop, libby; kondyli, dimitra; linardis, apostolos (2021) informed consent for linking survey and social media data, iassist quarterly 45(1), pp. 1-27. doi: https://doi.org/10.29173/iq988 appendix b website text with extended information on facebook data [project/study name] data protection information: facebook data your facebook data are collected by [name + address of institution note: the browser plugin used in the study was created and maintained by an external collaborator whose contact details were provided here] your data will be transmitted for analysis to [name + address of institution running the study/project] below you will find all information about our data collection that is relevant to you. you can contact us at the above address or via the email address [project email address] if you need more information about our research project. what information is collected about my facebook account? only posts from your facebook news feed that have been publicly shared are collected. private posts, such as status updates from friends, are not collected. the following data is collected: ● the author of the public post in your news feed, ● date and time when the post was created, ● if applicable, the person or page who publicly shared that post on facebook, ● contained text, contained image or video file, contained links, ● number of reactions (e.g. likes) and number of comments to the post, and ● position of the post within the news feed. personal login information, such as email address, login codes and passwords, are also not collected. although only public posts from your news feed are collected, we cannot exclude the possibility that the data collected may still contain personal information (for example, if one of your facebook friends posts publicly and tags you or others in these public posts). we anonymize such information or delete it before the data are analyzed. what is this information used for? we use the data exclusively for scientific research. combining the data on public posts in your facebook news feed with survey and web tracking data enables us to better understand your activities on the internet and your opinions. with additional data from facebook we can... ● better understand who gets exposed to which news on facebook. ● investigate whether the facebook news feed contains scientifically relevant information and how good the quality of this information is. ● identify issues that people may be concerned about but which are not part of our surveys ● test assumptions about the relationship between the use of social media and political attitudes and behavior. https://doi.org/10.29173/iq988 appendix 4/27 breuer, johannes; al baghal, tarek; sloan, luke; bishop, libby; kondyli, dimitra; linardis, apostolos (2021) informed consent for linking survey and social media data, iassist quarterly 45(1), pp. 1-27. doi: https://doi.org/10.29173/iq988 what do you do to protect my personal information? all information is stored and used in accordance with the eu general data protection regulation (eu-gdpr). the collected data are encrypted and transmitted to research servers, all of which are located in germany. in addition, you have the possibility at any time to view all data collected about you via the page [website for the browser plugin] after entering your personal identification (which you generate yourself in the questionnaire and the browser plugin). through that website, it is also possible for you to delete your facebook data. if you do not want the public posts from your facebook news feed to be collected, you can also deactivate the plugin. by simply clicking on the respective symbol (in the upper right corner of your browser) you can deactivate and activate the plugin. only information that is no longer personally identifiable is linked to the survey and web tracking data (e.g. how often you have seen news from a particular provider in your facebook news feed). who will have access to the data? the anonymised (aggregated) linked data, which includes your answers from the survey and web tracking data as well as information on public posts from your facebook news feed, will only be used for scientific research. commercial use of the data is excluded. access for third parties to the complete linked data will only be possible in a special secure environment. your rights you can withdraw your consent to the collection of your twitter data at any time. to do so, just send an email to [email address for the project] or a written letter to [name + address of the institute] please note that your twitter username must be mentioned in the email or letter, otherwise we cannot correctly assign your data for deletion. with regard to your personal data, you can make use of the following rights at any time: right of access to information right of rectification right to deletion (“right to be forgotten”) right to limit processing right to data transferability you also have a right of appeal to a data protection supervisory authority. contact person with all general questions and requests concerning data protection at [name of institution] you can contact: [name + address of data protection officer] https://doi.org/10.29173/iq988 appendix 5/27 breuer, johannes; al baghal, tarek; sloan, luke; bishop, libby; kondyli, dimitra; linardis, apostolos (2021) informed consent for linking survey and social media data, iassist quarterly 45(1), pp. 1-27. doi: https://doi.org/10.29173/iq988 appendix c linkedin additional help links and text what information will you collect from my linkedin account? we will only collect information from your linkedin account that you have made publicly available. this may include information from your profile (for example your work or education history and your connections), the profiles of your connections (such as information about your employer), and posts you have made (including text, images, videos and web links). we will update this information. this information will be collected and stored for as long as they are useful for research purposes. you can withdraw your consent at any time. if you do so, we will not collect any more of your linkedin data and will make no further links. however, previously collected data which has had your identifiers removed will be kept. what will the information be used for? the information will be used for social research purposes only. adding your linkedin information and your survey answers will allow researchers from universities, charities and government to better understand your experiences, such as with work and education. for example, using information from your linkedin account, researchers can start to: * understand who uses linkedin and how they use it * see what linkedin information can tell us about people and their work * collect information about things we don’t ask in our survey * understand what happens between waves of the survey who will be able to access the information? datasets which include both your survey answers and linkedin information will be made available for social research purposes only. researchers who want to use your detailed linkedin information must apply to access it and present a strong scientific case to ensure that the information is used responsibly and securely. summary information from your linkedin account which would not allow you to be identified will have the same access controls as your survey answers. at no point will any information that would allow you to be identified be made available to the public without your express permission what will you do to keep my information safe? all information we collect will be held in accordance with current data protection legislation (gdpr). to keep your information safe, researchers will only be able to access the matched survey answers and detailed linkedin information in a secure environment set up to protect this type of data. only approved researchers who have gone through special training may access this information, and they will have to apply to do so. summary information from your linkedin account which you cannot be identified from will have the same level of protection as your other survey answers. https://doi.org/10.29173/iq988 24 iassist quarterly 2015 iassist quarterly research data repositories: review of current features, gap analysis, and recommendations for minimum requirements by claire c. austin1,2,3, susan brown1,4, nancy fong5, chuck humphrey1,6, amber leahey1,7, peter webster1,8 abstract data sharing is increasingly recognized as integral to scientific research and publishing. this requires informed and thoughtful preparation from initial research planning to collection of data/metadata, interoperability, deposit in data repositories, and curation. research data canada (rdc) is a collaborative, non-government organization that promotes access to and preservation of canadian research data. the rdc standards and interoperability committee (rdcsinc) surveyed 32 canadian and international online data platforms for storage, data transfer, curation activities, preservation, access, and sharing features. we developed a checklist to compare criteria and features between platforms. the survey revealed a heterogeneity of features and services across platforms, non-standardized use of terms, uneven compliance with relevant standards, and a paucity of certified data repositories. recommendations for online digital infrastructure development to meet evolving researcher and end-user needs centre around persistent identification and citation of datasets, data reliability, version control, metadata, data sharing, privacy controls, long-term preservation of data, and certification of data repositories. we identified a need in canada for investment in an integrated, comprehensive national digital infrastructure for research data. keywords: data sharing, data publication, digital repository, data interoperability, data standards, data deposit, digital infrastructure. introduction research data sharing is increasingly recognized as an essential component of scholarly and scientific research. increased sharing improves the ability to reproduce results, replicate findings, and generate new knowledge (parr and cummings, 2005; hernan and wilcox, 2009; peng, 2011; poisot et al., 2013; stodden et al., 2014, 2015). although some disciplines (e.g., astronomy) have a long established practice of sharing and citing scientific data sets (socha, 2013), a very large number of researchers are still very reluctant to do so. perceived risks in data sharing sometimes put forth by researchers, such as damage to the researcher’s reputation, misinterpretation of the data, or misappropriation of the data (socha, 2013), all immediately disappear the moment the data are properly managed and documented. some surveys have found that approximately half of researchers share data (alsheikh-ali et al., 2011; vines et al, 2014). however, this most probably follows publication of results in peer-reviewed journals, often years after the data were originally collected, and data sharing does not necessarily mean the data are useable by another researcher. the usability of shared data relates to best practices in data management, data structure, interoperability, metadata, licensing, and iassist quarterly 2015 25 iassist quarterly accessibility (jones et al., 2006; peer and green, 2014). in canada, increasing public access to scientific research data will help drive innovation and discovery across the broader scientific community, as well as implementation of better data management practices (government of canada, 2014). a major source of research funding in canada is tri-council plus (tc3+): the social science and humanities research council (sshrc), the natural sciences and engineering research council (nserc), the canadian institutes of health research (cihr), and the canada foundation for innovation (cfi). tri-council plus has stated that, “the potential of data-intensive research is progressively and rapidly outstripping our ability to manage and to grow the digital ecosystem to meet 21st century needs” (government of canada, 2013). in an effort to establish a greater culture of data stewardship, the canadian granting councils agreed to promote and develop appropriate data management systems and capabilities, in line with existing data and best practices globally. in early 2015, the canadian research councils formulated a harmonized open access policy that requires all peer-reviewed journal publications funded by one of the three granting agencies to be made freely available online by depositing the manuscript(s) in an online repository within 12 months of publication (government of canada, 2015). cihr-funded researchers are also required to deposit their research data into a relevant disciplinary repository immediately after publication of research results, and they must retain original data sets for a minimum of five years. this is enormous progress, but it also begs some important questions. why are original datasets required to be kept for only five years? why are nsercand sshrc-funded researchers not also required to deposit their research data in a digital repository? when and how will tri-council provide incentive to researchers and reward them for data publication, elevating the practice to a first-class research output on par with traditional forms of journal publication and thereby lead the way for needed change in the academic reward system? would it not benefit the researcher, the broader scientific community, and the common good if data publication were to precede journal publication, even? data publication should be peer reviewed as rigorously as journal articles in the academic and scientific literature, and data should be openly shared in curated data repositories. data are the foundation of everything else that follows, and researchers must receive credit for producing reliable data (costello, 2009; atici et al, 2013; kratz and strasser, 2014). we recognize that principal investigators (p.i.’s) have a primary responsibility in data management and data publication (see endnotes 1&2). it must also be emphasized that there needs to be a robust digital infrastructure in place to support proper data management and to ensure that data are preserved in a useable form for people other than the creators of the data – whether or not the p.i.’s care about this, although they should. credible data publication requires effective data management and a robust digital infrastructure (bloom et al., 2015). is such an infrastructure currently in place so that governments and funding agencies can take that next step in requiring robust data management plans and deposit of research data in data repositories? this is the question that the present paper seeks to answer, at least in part. methods the research data canada (rdc) standards and interoperability committee (sinc) surveyed canadian and international online data platforms to identify currently implemented standards, requirements, and features related to the management and sharing of research data across a variety of academic disciplines. this work was done in parallel with the development of, ’guidelines for the deposit and preservation of research data in canada’ (research data canada, 2015a). the categories for assessment used in the present work were developed from community guidelines and digital preservation literature (see references section). online data platforms that were publicly accessible via the world wide web and that allowed data upload were included in the survey. the survey was performed during the period october 2014-february 2015. the first phase focused on a group of large, established, general platforms (specifically dryad, figshare, dataverse, icpsr, pangaea). although it is a metadata platform not a data repository, datacite was also included. publicly available information, including upload and submission instructions, data requirements, recommended metadata and file naming conventions, data sharing and deposit policies, user guidelines and documents, data dissemination formats, persistent identifiers, and stated data preservation activities were reviewed. in some cases, online platforms restricted user access and did not have openly available documentation regarding metadata and data submission requirements. in those cases, we created a user account and password and attempted to load a sample dataset into the data platform for the purposes of the review. in the second phase, a total of 32 online platforms were surveyed for the following: deposit and submission, storage, description, curation, preservation and archiving, dissemination policies and features, collaboration options, and open access (table 1). these included platforms in the biological & life sciences, social sciences (economics, sociology, political science, etc.), medical & life sciences, earth & environmental sciences, one from physics, and one from astronomy. there were 19 platforms -covering multiple disciplines in the same general domain area (e.g. medical sciences, social sciences). comparison of the 32 online platforms was a challenge due to the heterogeneity of features and the non-standardized use of terms. platform features and data criteria to be surveyed were developed based on ’data seal of approval’ guidelines (data seal of approval board, 2013), ’trustworthy repositories audit & certification’ (trac) criteria and checklists (center for research libraries, 2007), and an initial survey of features observed in the selected online platforms. these features and data criteria were compiled in a checklist that was used as a tool to compare features and requirements across platforms (table 2). the use of the checklist to identify majority practice (i.e., >50% across platforms) with respect to any feature or data criteria was still exceedingly difficult. therefore, for the summary results we used a lower threshold, arbitrarily set at 40%, as a more informative indicator of relatively common practice with respect to the inclusion or exclusion of the features/criteria. results summary results from our survey of the 32 online platforms are found in table 3. detailed results can be viewed online in the, ’repository requirements features review spreadsheet’ found in the rdc-sinc dataverse repository (research data canada, 2015b). 26 iassist quarterly 2015 iassist quarterly subject areas we found that a large number of platforms surveyed handled a variety of data and were multidisciplinary in scope. however, the majority identified with a particular domain or area of study (e.g. earth and environmental sciences, social sciences, medical and life science, etc.). online platforms surveyed often had strong government and academic affiliations, with nearly 41% (13 out of 32) being supported directly by government. an additional seven were ngo’s, six were institutional (academic), three were corporate or commercial, and the remainder the affiliation was unclear. metadata as to the kinds of features the platforms supported for metadata and description of datasets, we noted that they generally recognized depositors as being central to the data publication process. datasets and metadata uploaded to these platforms often contained information concerning authors, publishers, subject matter, dates of collection, abstract etc. support for metadata ingestion and creation was a feature that we looked at particularly closely. the majority of platforms surveyed, 69% (22 out of 32), used some kind of local or custom metadata profile or schema for description and documentation of datasets. nearly 38% of platforms surveyed (12 out of 32) supported or were mapped to a standard metadata set for resource description, e.g., dublin core (dc) or datacite. additional support was noted among some of the platforms for discipline specific standards such as the fgdc and/or iso 19115 for geographic information (7 out of 32 platforms), or the data documentation initiative (ddi) (6 out of the 32 platforms). while many platforms used standardized metadata, a number of the major platforms used non-standardized, internally devised metadata schemas which could not be crosssearched and that were not interoperable with any other system or resource. the granularity of metadata varied significantly across platforms, and we would note that there was very limited support for dataset or file-level metadata descriptions across platforms, however this was not fully captured in this survey. persistent identifiers typically, the platforms surveyed ensured that uploaded datasets were assigned a unique or persistent identifier (e.g. uri, pid) for proper online identification and access. however, they varied in their approach to the use of persistent identifiers, with some providing a resolvable url to the dataset’s associated metadata. approximately half (17 out of 32) of the platforms surveyed supported the digital object identifier (doi) standard for persistent identification of datasets. other persistent identifier standards that were used included dspace handles and urns (6%, or 2 out of 32), with the majority using a local or some unknown unique identification system. typically, identifiers were assigned at the level of metadata description for the dataset or study. concerning the ease of data citation, we found that close to 63% of the platforms (20 out of 32) provided a direct data citation and/or some other mechanism to cite stored data. version control we found that version control, although an important issue, was still an unresolved problem in most repostitories. more than two thirds of the platforms, (22 out of 32), allowed depositors to edit files after they had been uploaded. about 40% (13 out of 32), offered a standard version control system, or version statement. approximately 72% (23 out of 32) of the platforms provided time stamping of uploaded files. time stamping appears to be the most common practice applied to identify changed files, but this does not constitute version control. only one platform offered a systematic and persistent method for identifying versions of datasets (universal numeric fingerprint (unf)). ownership and data reuse approximately three quarters of the platforms surveyed (24 out of 32) associated a creative commons or other open license with the datasets. the majority also supported other data use licenses – often customized to the specific platform – but not meeting any standards. these included restricted licences where ownership rights were retained and that defined limited terms of use for datasets. provision for access to data with restrictions was noted in close to 84% of the platforms (27 out of 32). nearly 66% of the platforms (21 out of 32) published a specific policy on data sharing, terms of use, and ownership. three provided no information concerning terms of use of shared or downloaded data. fees and access from the outset, we’ve assumed that data should be made available online and shared for free, when there were no legal or ethical reasons not to do so. nearly all of the platforms surveyed offered some form of open, free, or anonymous access to data. two thirds of them (23 out of 32) also offered free data deposit. 25% (8 out of 32) sought some form of payment or funding from some or all data depositors for services such as data publication, including preparation, curation or preservation. in assessing for open access amongst the platforms, nearly all of the platforms provided some public information concerning access to data and their terms of use. when data were provided openly, it was not always provided strictly anonymously. most of the platforms surveyed, 78% (25 out of 32) offered some form of authentication whereby users needed to “sign in” in some way to gain access. more work is needed to understand the kind of restrictions applied and the reasons for them, especially as this relates to open access. data usage with regard to tracking data usage, approximately half (15 out of 32) of the platforms surveyed indicated that they offered download or other usage statistics to demonstrate access to and reuse of datasets. the remainder provided no information related to usage. dataset curation and publication in general, the platforms surveyed offered data providers some level of support for dataset publication, although these activities varied greatly between platforms and across disciplines. approximately two thirds (23 out of 32) indicated that they offered some sort of data curation service, including metadata support, or review of the data, prior to publication. generally, few platforms provided detailed explanation about curation services. for those that did state that there was some data curation activity, the detail and extent of the curation services provided were vague or unclear. interoperability in general, in terms of standards for the effective access and exchange of data and metadata, we note that support for open and interoperable standards is not widespread. only 34% of the iassist quarterly 2015 27 iassist quarterly platforms surveyed (11 out of 32) supported open archives initiative (oai) protocols, such as the oai-pmh protocol for the open exchange and harvesting of data and metadata. however, nearly two-thirds (19 out of 32) offered alternative access to data and metadata through some form of application programming interface (api) for online access and exchange. sixteen of the platforms surveyed supported either xml or json format for export and exchange. preservation we were able to extract very few details from the information provided concerning preservation. nonetheless, nearly 56% (18 out of 32) indicated that they did offer long-term storage and preservation of data and had a preservation policy and practices statement. additionally, 44% (14 out of 32) indicated that the platform set-up included multiple redundancy and backup for files. fewer than 13% (4 out of 32) indicated the use of standard file transfer and copy systems such as ‘lockss’ (stanford university, 2015) or a closed system using lockss technology (clockss, 2015). certification only 20% (6 out of the 32) of the platforms surveyed were certified under some form of community assessment or certification body such as ’world data system (wds)’ or ’data seal of approval’ (data seal of approval board, 2013). with only two of the platforms providing information concerning their succession plans, we note that statements and policies concerning plans for data after the online platform ceases to exist were virtually non-existent. discussion increased data sharing and greater openness of scientific research requires robust data infrastructure and sound data and metadata management practices. the present survey is a broad overview of the current features of canadian and international repositories and data sharing platforms. this work is not a comprehensive list of available online data platforms or data repository requirements and features, nor is it a replacement for repository assessment or accreditation. it has, however, identified areas where action is needed to develop the necessary national digital infrastructure in canada to support researchers with management, sharing, and preservation of research data. the checklist and findings may also assist further study and development of best practices. moving forward, investment is needed to develop an integrated, comprehensive digital infrastructure and to improve data sharing and reuse of research data in canada. the initiative funded under the european union’s horizon 2020 research and innovation programme, resulting in eudat (2015), is a good example. for data sharing to be effective, data must be reliable, usable, easily discoverable, accessible, and stored in a persistent manner for the long-term. most importantly, datasets must be considered legitimate research outputs and be appropriately acknowledged for their value in promotion, tenure, and funding decisions to the same degree as are other peer-reviewed publications. the emergence of data journals publishing peer-reviewed scholarly and scientific datasets is a step in this direction. however, this needs to be accompanied by a significant culture change in the academic community in order to become a reality. new data journals, and increasingly, traditional journals, recommend or use existing digital online platform infrastructures (figshare blog, 2015). their data policies vary in terms of standards, compliance enforcement, and data review (stodden et al., 2013; peer and green, 2015). data sharing is frequently a ‘self-deposit’ model, whereby the publisher recommends a list of online data platforms that may or may not perform quality control or review of the data deposited (nature publishing group, 2015). the journal plos one, for example, recommends 76 repositories which have been grouped into one of 11 categories: unstructured and/or large data; sequencing; omics; structural databases; neuroscience; model organisms; taxonomic and species diversity; biomedical sciences; biochemistry; physical sciences; and, social sciences (plos one, 2015). support for standard metadata is highly variable between repositories and data sharing platforms. more than two thirds (23) of the online platforms surveyed provided some support for metadata creation (i.e., guidelines, templates, review etc.), but most large ones still left metadata quality control largely in the hands of the data providers. metadata are the backbone of any dataset and ongoing quality control of metadata is as important as the data. metadata are vital in ensuring that the data are correctly understood and can be effectively used. given the importance of quality control, it is noteworthy that the majority of the platforms surveyed did not address this issue. data curation is the activity of managing and promoting the use of data from the point of creation to ensure that the data are fit for contemporary purpose and available for discovery and reuse (research data canada, 2014). for dynamic datasets this may mean continuous enrichment or updating to maintain fitness for purpose. higher levels of curation also involve links with annotation and with other published materials. one third of the platforms surveyed provided no information concerning data curation, and the remainder provided only vague or unclear information. additional work is needed to understand the curation process used by different online platforms in much greater detail, to understand what is meant by curation in each case, how data selection, retention and quality control decisions are made and what processes are in place. in the development of research data management services and support, the primary focus has been at the institutional level. this often coincides with the need to develop institutional online repositories such as those that now exist at harvard university, hong kong university of science and technology, john hopkins university, monash university, and purdue university (wong, 2009). beyond the needs of repository managers and organizations who are primarily interested in digital preservation, few resources are available for researchers, survey managers, granting agencies, publishers, librarians, or archivists to assess the suitability of online platforms for research data deposit and sharing (humphrey, 2015; guindon, 2014). however, there do exist excellent repository assessment and best practice guidelines, such as the ’trusted repository audit checklist (trac)’, ’trustworthy digital repository checklist (tdr)’, ’digital repository audit method based on risk assessment –drambora (digital curation centre, 2009)’, and the ’data seal of approval (data seal of approval board, 2013)’. these can be used to adopt and use best practices in the selection of repositories for data deposit, and a lot can be said about the benefits of certification especially for the selection of repositories by researchers. researchers will also benefit from the ‘repository platforms for research data interest group’ that gathers and 28 iassist quarterly 2015 iassist quarterly analyzes research data use cases in the context of repository platform requirements (research data alliance, 2015). privacy and confidentiality of personal or sensitive information is an issue which has not been addressed by research data repositories, in part due to the lack of appropriate tools and methods. harvard university (2015) is developing ’data tags’, a promising new tool which will help researchers to share and use sensitive data in a standardized and responsible manner. clearly, costs associated with data and metadata infrastructure and curation services are considerable, and these will increase with the success and growth of each repository. datacite canada, for example, is providing its doi minting service for free to non-profit organizations until march 31, 2016. this business model is currently under review. research data management is, in fact, a transdisciplinary field. one of the challenges in this endeavour is finding a common language to overcome domain-specific methods and terminology. research data canada has developed a living glossary of more than 500 terms and definitions to help researchers and others better communicate and understand the various aspects of research data management, including the sharing and perservation of data. the consortia advancing standards in research administration information (casrai) has made this glossary freely available over the internet on a semantic mediawiki that includes a discussion page for each term (research data canada, 2015c). conclusion although principal investigators are ultimately responsible for the integrity of the data upon which their research findings are based, few have the knowledge, time, or resources to implement state-ofthe-art data management practices or evaluate online data storage and sharing options (guindon, 2014). the results of the present survey suggest that there is still a great deal of work to be done to ensure that online data platforms meet minimum standards for reliable curation and sharing of data. we believe that canada’s tri-council is wise in being cautious about what it requires from researchers in terms of data management and online deposit of research data until a robust national digital infrastructure, including supported data management, is established in canada. academic libraries and archives already have experience with client service and with storage of a vast array of file types: audio, images, software code, and datasets. a logical next step for improving digital infrastructure in canada would be the expansion of existing library and archive services in the development of a national data infrastructure, including institutional repositories, with complementary data management consultation services to support researchers (wong, 2009). we also recommend that best practices for data management and the systematic use of data repositories be incorporated into the curriculum at university undergraduate and graduate levels in the humanities, business, sciences, engineering, computer science, mathematics and statistics, and medical sciences, to begin to building capacity and skills in this area. author statement all authors contributed equally to the writing of this paper. all authors declare no conflict of interest. opinions expressed in this paper are those of the authors and do not necessarily reflect the policies of the organizations with which they are affiliated. feedback readers are invited to provide feedback to the authors at the following link: https://www.surveymonkey.com/s/ rdc-sinc_iassist2015_paper_feedback funding this work was supported by research data canada (rdc). references alsheikh-ali, a., qureshi, w., al-mallah, m., and ioannidis, j. p. a. (2011). public availability of published research data in high-impact journals. plos one, 6(9) doi: http://dx.doi.org/10.1371/journal.pone.0024357 atici, l., kansa, s. w., lev-tov, j., and kansa, e. c. (2013). other people’s data: a demonstration of the imperative of publishing primary data. journal of archaeological method and theory, 20(4), 663-681. doi: http://dx.doi.org/10.1007/s10816-012-9132-9 bloom, t., dallmeier-tiessen, s., murphy, f., austin, c.c., whyte, a., tedds, j., nurnberger, a., raymond, l., stockhause, m., and vardigan, m. (2015). workflows for research data publishing: models and key components (submitted version). international journal on digital libraries research data publishing special issue. 27 pages, june 30, 2015. doi: 10.5281/zenodo.20308 clockss. (2015). controlled lockss. http://www.clockss.org/clockss/ home costello, m. j. (2009). motivating online publication of data. bioscience, 59(5), 418-427. center for research libraries (february 2007). trac – trustworthy repositories audit & certification: criteria and checklist v 1.0. center for research libraries & online computer library center. http://www. crl.edu/sites/default/files/d6/attachments/pages/trac_0.pdf digital curation centre. (2009). digital repository audit method based on risk assessment – drambora. http://www.dcc.ac.uk/resources/ repository-audit-and-assessment/drambora data seal of approval board (july 2013). repositories – data seal of approval guidelines, v2. http://datasealofapproval.org/media/ filer_public/2013/09/27/guidelines_2014-2015.pdf eudat (2015). research data services, expertise & technology solutions. http://www.eudat.eu figshare blog (2015). the rise of the ’data journal’. macmillan publishers. http://figshare.com/blog/the_rise_of_the_data_journal_/149?utm_ source=users+%2b+advisors&utm_campaign=7d388f3383figshare_integrates_with_projects8_9_2013&utm_ medium=email&utm_term=0_e5f7149158-7d388f3383-97189037 government of canada. (2013). capitalizing on big data: toward a policy framework for advancing digital scholarship in canada. tri-council plus. http://www.sshrc-crsh.gc.ca/about-au_sujet/ publications/digital_scholarship_consultation_e.pdf accessed 2015-04-27. government of canada. (2014). canada’s action plan on open government 2014-2016. http://open.canada.ca/en/content/ canadas-action-plan-open-government-2014-16 accessed 2015-04-27. government of canada. (2015). tri-agency open access policy on publications. http://www.science.gc.ca/default. asp?lang=en&n=f6765465-1 guindon, a. (2014). research data management at concordia university: a survey of current practices. feliciter 60( 2), p15. harvard university (2015). privacy tools project data tags. harvard school of engineering and applied sciences. http://privacytools.seas. harvard.edu/datatags hernan, m. a. & wilcox, a. j. (2009). epidemiology, data sharing, and the challenge of scientific replication. epidemiology, 20(2), iassist quarterly 2015 29 iassist quarterly 167-168. doi: http://journals.lww.com/epidem/fulltext/2009/03000/ epidemiology,_data_sharing,_and_the_challenge_of.3.aspx humphrey, c. (2015). preserving research data in canada the long tale of data. chuck humphrey blog. http:// preservingresearchdataincanada.net accessed 2015-04-27. jones, m. b., schildhauer, m. p., reichman, o. j., and bowers, s. (2006). the new bioinformatics: integrating ecological data from the gene to the biosphere. annual review of ecology, evolution and systematics, 37, 519-544. doi:10.1146/annurev. ecolsys.37.091305.110031 kratz, john; strasser, carly (2014). data publication consensus and controversies, v3. http://f1000research.com/articles/3-94/v3 nature publishing group (2015). questionnaire to assist with scientific data repository evaluation.. http://www.nature.com/uploads/ ckeditor/attachments/1301/scidata_respository_evaluation_ march2015.docx parr, c. s., and cummings, m. (2005). data sharing in ecology and evolution. trends in ecology & evolution, 20(7), 362 363. http:// www.sciencedirect.com/science/article/pii/s0169534705001308 peer, l., green, a. (2014). committing to data quality review. presented at the 9th international digital curation conference. http://isps. yale.edu/sites/default/files/files/commitingtodataqualityreview_ idcc14-preprint.pdf peer, l., green, a. (2015). research data review is gaining ground. political science replication blog. https://politicalsciencereplication. wordpress.com/2015/03/26/guest-post-research-data-review-isgaining-ground-by-l-peer-and-a-green/ peng, roger d. (2011). reproducible research in computational science. science, 334(6060), 1226-1227. poisot, t., mounce, r. and d. gravel. 2013. moving toward a sustainable ecological science: don’t let data go to waste! ideas in ecology and evolution, vol 2, no. 6. http://library.queensu.ca/ojs/index.php/iee/ article/view/4632/4992 plos one. (2015). data availability and recommended repositories. http://journals.plos.org/plosone/s/ data-availability#loc-recommended-repositories research data alliance. (2015). repository platforms for research data working group. https://rd-alliance.org/ group/repository-platforms-research-data/case-statement/ repository-platforms-research-data-case research data canada (2014). rdc glossary of terms and definitions v1.0. standards and interoperability committee. http://www.rdc-drc. ca/glossary research data canada (2015a). guidelines for the deposit of research data in canada. standards and interoperability committee. http:// www.rdc-drc.ca/wp-content/uploads/guidelines-for-deposit-ofresearch-data-in-canada-2015.pdf research data canada (2015b). research data repository requirements and features review. standards and interoperability committee. http://dataverse.scholarsportal.info/dvn/dv/rdcsinc research data canada (2015c). rdc glossary of terms and definitions v2.0. standards and interoperability committee. http://dictionary. casrai.org/category:research_data_domain socha, y.m. (ed.) (2013). out of cite, out of mind: the current state of practice, policy, and technology for the citation of data. codataicsto task group on data citation standards and practices. data science journal, volume 12. https://www.jstage.jst.go.jp/article/ dsj/12/0/12_osom13-043/_pdf accessed 2015-04-27. stanford university (2015). lots of copies keep stuff safe. http://www. lockss.org stodden, victoria; guo p, ma z (2013) toward reproducible computational research: an empirical analysis of data and code policy adoption by journals. plos one 8(6): e67111. doi:10.1371/ journal.pone.0067111 stodden, victoria; leisch, friedrich; peng , roger d. (2014). implementing reproducible research. crc press. isbn 9781466561595. stodden, victoria; miguez, sheila; seiler, jennifer (2015). researchcompendia.org: cyberinfrastructure for reproducibility and collaboration in computational science. computing in science & engineering, 17(1), 12-19. vines, t. h., albert, a.y.k., andrew, r. l., de´barre, f., bock, d.g., franklin, m. t., . . . rennison, d. j. (2014). the availability of research data declines rapidly with article age. current biology 24(1), 94–97. wong, g. (2009). exploring research data hosting at the hkust institutional repository. serials review, 35(3), 125–132. definitions 1. rdc defines ’data management’ as, “the activities of data policies, data planning, data element standardization, information management control, data synchronization, data sharing, and database development, including practices and projects that acquire, control, protect, deliver and enhance the value of data and information” (research data canada, 2014, 2015c). rdc views the principal investigator (p.i.) as having responsibility in this area, his or her role being defined as the person who “has a research leadership role and is the point of contact for a project or partnership that applies the scientific method, historical method, or other research methodology for the advancement of knowledge resulting in independent, objective, high quality, traceable, and reproducible results. the p.i. has primary responsibility for the intellectual direction and integrity of the research or research-related activity, including data production, findings and results, and ensures ethical conduct in all aspects of the research process including but not limited to the treatment of human and animal subjects, conflicts of interest, data acquisition, sharing and ownership, publication practices, responsible authorship, and collaborative research and reporting. while various tasks may be delegated to team members, some of whom may have greater expertise in specific areas, the p.i. is familiar with the various technical and scientific aspects of a project and how they fit together, is able to identify and remediate gaps, and ensure communication within the team and with users of the research data and results” (research data canada, 2014, 2015c). 2. rdc uses the following terms and definitions relevant to the deposit and preservation of research data (research data canada, 2014, 2015c): ’data centre’ a facility providing it services, such as servers, massive storage, and network connectivity. ’data repository’ an archival service providing the long-term care for digital objects with research value. the standard for such repositories is the open archival information system reference model (iso 14721:2003). ’repository’ repositories preserve, manage, and provide access to many types of digital materials in a variety of formats. materials in online repositories are curated to enable search, discovery, and reuse. there must be sufficient control for the digital material to be authentic, reliable, accessible and usable on a continuing basis. ’trusted digital repository (tdr)’ a repository whose mission is to provide its designated community with reliable, long-term access to managed digital resources.” please see the glossary for definitions of other related terms (research data canada, 2014, 2015c). 30 iassist quarterly 2015 iassist quarterly table 1. online data platforms surveyed 3tu.datacentrum h0p://datacentrum.3tu.nl/en/home/ icpsr h0ps://www.icpsr.umich.edu/icpsrweb/landing.jsp arcgis online h0p://doc.arcgis.com/en/arcgis-online/share-maps/ share-items.htm immport h0p://www.immport.org/immport-open/public/home/home archaeology data service h0p://archaeologydataservice.ac.uk/ iris h0p://www.iris.edu/hq/ b.c. conservakon data centre h0p://www.env.gov.bc.ca/cdc/ journal of applied econometrics data archive (queen’s university) h0p://qed.econ.queensu.ca/jae/ barcode of life data systems (bold) h0p://www.boldsystems.org/ labarchives h0p://www.labarchives.com/ biolinc (biologic specimen and data repository informakon coordinakng center) h0ps://biolincc.nhlbi.nih.gov/home/ nakonal snow and ice data centre h0p://nsidc.org canadian astronomy data centre (canfar) h0p://www.canfar.phys.uvic.ca/canfar/ nesstar* (<odesi>, cessda) h0p://www.nesstar.com h0p://www.cessda.net cern open data portal h0p://opendata.cern.ch/?ln=en ocean networks canada h0p://www.oceannetworks.ca/informakon ckan* h0p://ckan.org openaire / zenodo repository h0p://www.zenodo.org/ datacite h0ps://www.datacite.org opencontext.org h0p://opencontext.org/ dataverse* (ocul dataverse, harvard) h0p://dataverse.org openicpsr h0ps://www.openicpsr.org/ dryad h0p://datadryad.org/ pangaea h0p://www.pangaea.de/ easy (dans) h0ps://easy.dans.knaw.nl/ui/home polar data catalogue h0ps://www.polardata.ca/ figshare h0p://figshare.com scratch pads h0p://scratchpads.eu/ flowrepository h0ps://flowrepository.org/ sda h0p://sda.berkeley.edu/archive.htm geoss portal h0p://www.geoportal.org/web/guest/geo_home_stp uk data archive / reshare (e-prints) h0p://www.data-archive.ac.uk/home * data repository soeware 1 iassist quarterly 2015 31 iassist quarterly table 2. features checklist used to compare 32 online data platforms category sub-category detailed features hardware & infrastructure server (server resources, pla/orms etc.) ● cloud (i.e. amazon s3) ● dspace ● local ibm server and storage, vmware esxi virtualized redundant server farm ● other cost ● free to access, download, and deposit data ● free to access, but contribulon suggested or required for deposit, i.e. funding structure for access/ deposit beyond the threshold ● publishing charge $ ● formal agreement with research/monitoring program for funding to support archiving and serving the datasets size ● size of repository (number of files, datasets) descrip;on domain ● mulldisciplinary ● earth & environmental science ● medical & life sciences ● social sciences (economics, sociology, polilcal science, etc.) ● physics ● biological and life sciences preserva;on redundancy ● mullple redundant copies ● clockss geographically and geopolilcally distributed network of redundant archive nodes persistent idenlfiers ● doi (specify where possible) ● dspace handle (hdl) ● other persistent ids ● other unique resource idenlfiers (i.e. uris) (not persistent) ● ezid registralon management or other persistent idenlfier registralon persistent data deposit ● long term preservalon of data curalon ● data curalon (specify where possible) privacy & security security ● authenlcalon mechanisms ● dislnclon between public and private data archiving author idenlfier ● orcid id ● scopus id ● digital author idenlfier timestamping and version control ● timestamped upon upload ● data can be edited following upload ● version statement ● universal numeric fingerprint (unf) citalon and references ● citalon provided (specify format) 1 32 iassist quarterly 2015 iassist quarterly submission data types accepted (list exceplons per repo) ● datasets ● metadata (supported upload of exchange formats (xml)) ● computer code ● other files ● figures ● audio (mp3, wav) ● video (mpg: mpeg2 for pal, vlc, mp4: aac, mpeg-4 for hdtv) ● photo (lff, jpeg) ● file sets ● formaded documents (pdf(a), odf, ascii) ● geospalal (kml/kmz, web map service/context, georss, gml) ● raster/matrix ● vector ● most kinds of data (text, spreadsheets, video, photographs, soeware code, compressed archives of mullple files, non-data files) ● publicalons (papers, posters, presentalon) ● compressed (zip) size (storage allocated, upload limits etc.) ● specify size metadata data submission (where applicable) ● metadata ● other (local schema, discipline specific) ● digital resource descriplon (dublin core, datacite, marc21) ● geospalal metadata (iso 19139, fgdc, iso 19115, inspire, etc.) ● health (nih cde etc.) ● ddi (ddi v2, lifecycle etc.) ● controlled language terminology ● readme file (data descriplon, definilons of column & row headings, data codes including missing data, units, data processing steps, contact info, etc.) support ● support for data preparalon and quality control ● formal review and approval of submided metadata and data before availability online access & sharing online access ● data available for free and open download (no registralon, must anonymously "agree" to terms of use) web services ● api for harveslng & search access, proprietary (rest or soap) ● oai pmh harveslng and search access, oai-pmh exchange format ● other web service exchange metadata ● exchange formats (xml, json) 2 iassist quarterly 2015 33 iassist quarterly license ● crealve commons license (adribulon or zero) ● government license ● other license (open) ● other license (restricted) linkages ● linkage between data and publicalon, and / or citalon indexes collabora;on mullple user collaboralon ● collaboralon (project workspace for mullple users) policy mandate ● under what authority does the repository operate (i.e. governing enlty) guidelines ● terminology or glossary of terms data sharing policy ● data rights and usage statement (data for use) ● data sharing policy available (data for deposit) data deposit policy ● terms & condilons by which data are ingested into the repository data ownership policy a repository's statement about ownership of the data it ingests formats policy ● the digital formats accepted by the repository and whether of formats is performed preservalon policy ● a repository's statement about its preservalon praclces succession plan ● aclons to be taken in the event that the repository is closed administra;on tracking ● counts views & downloads tabular data view data ● tabular data view, map view etc. file conversion formats ● file conversion oplons i.e. formats, projeclon etc. download ● download oplons cer;fica;on status trusted repository status ● icsu world data system ● data seal of approval 3 34 iassist quarterly 2015 iassist quarterly table 3. summary of detailed features found in 32 online data platforms surveyed * features and data criteria yes no not available other cloud (i.e. amazon s3) 7 8 17 free to access, download data, and deposit data 23 7 0 2 free to access, but contribuaon suggested or required for deposit (i.e. funding structure for access / deposit beyond the threshold) 13 19 publishing charge $ 8 24 formal agreement with research/monitoring program for funding to support archiving and serving their resulang datasets 7 23 2 size of repository (number of files, datasets) note 1 muladisciplinary 20 12 earth & environmental science 21 11 medical & life sciences 15 17 social sciences (economics, sociology, poliacal science, etc.) 17 15 biological and life sciences 17 14 physics 1 31 mulaple redundant copies 14 10 7 1 lockss/clockss geographically and geopoliacally distributed network of redundant archive nodes 4 18 10 doi (specify where possible) 17 14 1 dspace handle (hdl) 1 30 1 other persistent idenafiers (urns; purls) 1 2 other unique resource idenafiers (i.e. ids) (not persistent) 13 17 2 ezid registraaon management or other persistent idenafier registraaon 3 28 1 long term preservaaon of data 18 9 4 1 data curaaon (specify where possible) 23 7 2 authenacaaon mechanisms 26 3 3 1 iassist quarterly 2015 35 iassist quarterly disancaon between public and private data 28 3 1 orcid id 6 24 2 scopus id 1 27 4 digital author idenafier 1 30 1 timestamped upon upload 23 7 2 data can be edited once uploaded 22 3 7 version statement 13 12 7 universal numeric fingerprint (unf) 1 29 2 citaaon provided (specify format) 20 12 datasets 31 1 metadata (supported upload of exchange formats (xml)) 12 11 9 computer code 19 6 7 other files 22 4 6 figures 16 8 8 audio (mp3, wav) 13 9 10 video (mpg: mpeg2 for pal, vlc, mp4: aac, mpeg-4 for hdtv) 12 10 10 publicaaons (papers, posters, presentaaon) 21 4 7 file sets ("mulaple related") 22 3 7 compressed (zip) 22 3 7 most kinds of data (text, spreadsheets, video, photographs, soiware code, compressed archives of mulaple files, non-data files) 21 5 6 photo (aff, jpeg) 20 4 8 formaled documents (pdf(a), odf, ascii) 22 4 6 geospaaal (kml/kmz, web map service/context, georss, gml) 19 4 9 raster/matrix 17 6 9 vector 19 4 9 specify size 19 metadata other (local schema, discipline specific) 22 7 3 2 36 iassist quarterly 2015 iassist quarterly metadata digital resource descripaon (dublin core or datacite) 11 20 metadata geospaaal metadata (iso 19139, fgdc, iso 19115, inspire, etc.) 7 25 metadata health (nih cde etc.) 2 30 metadata ddi (ddi v2, lifecycle etc.) 6 26 metadata (controlled language terminology) 8 21 3 readme file (data descripaon, definiaons of column headings & row labels, data codes including missing data, units, data processing steps, contact info) 14 14 4 support for data prep and quality control 23 7 2 formal review and approval of submiled metadata and data before availability online 20 9 3 data available for free and open download (no registraaon required, must anonymously "agree" to terms of use) 22 9 1 api for harvesang & search access, proprietory (rest or soap) api 16 14 2 oai pmh harvesang and search access, oai-pmh exchange format 11 19 2 other web service 10 20 2 exchange formats (xml, json) 17 13 2 creaave commons (alribuaon or zero) 19 11 2 open government license 9 21 2 other license (open) 20 11 1 other license (restricted) 19 11 2 linkage between data and publicaaon, and / or citaaon indexes 18 10 4 collaboraaon (project workspace for mulaple users) 15 15 2 3 iassist quarterly 2015 37 iassist quarterly under what authority does the repository operate (i.e. governing enaty) not e 1 terminology or glossary of terms 10 19 2 data rights and usage statement (data for use) 14 17 1 data sharing policy available (data for deposit) 21 10 1 terms and condiaons by which data are ingested into the repository 21 7 4 a repository's statement about the ownership of the data it ingests 19 10 3 the digital formats accepted by the repository and whether normalizaaon of formats is performed 18 9 4 a repository's statement about its preservaaon pracaces 16 15 1 acaons to be taken in the event that the repository is closed 2 26 4 counts views & downloads 15 14 3 tabular data view, map view etc. 18 12 2 file conversion opaons i.e. formats, projecaon etc. 9 20 3 download available 27 3 2 cerafied as a trusted repository? (i.e. data seal of approval) 7 24 1 creaave commons (alribuaon or zero) 19 11 2 open government license 9 21 2 other license (open) 20 11 1 other license (restricted) 19 11 2 linkage between data and publicaaon, and / or citaaon indexes 18 10 4 collaboraaon (project workspace for mulaple users) 15 15 2 under what authority does the repository operate (i.e. governing enaty) not e 1 terminology or glossary of terms 10 19 2 4 38 iassist quarterly 2015 iassist quarterly notes 1. research data canada, standards and interoperability committee, contact: amber.leahey@utoronto.ca 2. environment canada 3. carleton university 4. university of guelph 5. university of toronto 6. university of alberta 7. scholars portal 8. saint mary’s university data rights and usage statement (data for use) 14 17 1 data sharing policy available (data for deposit) 21 10 1 terms and condiaons by which data are ingested into the repository 21 7 4 a repository's statement about the ownership of the data it ingests 19 10 3 * the use of the checklist to idenafy majority pracace (i.e. >50% across plasorms) with respect to any feature or data criteria was sall exceedingly difficult. therefore, a lower threshold, arbitrarily set at 40%, was used as a more informaave indicator of relaavely common pracace with respect to their implementaaon (or not). relaavely common pracace across plasorms is idenafied by the cells highlighted in colour in the table. note1: see repository spreadsheet (research data canada, 2015b). note 2: the majority indicated up to 2gb upload (remote submission). 5 1/12 de, suparna; moss, harry; johnson, jon; li, jenny; pereira, haeron and jabbari, sanaz (2022). engineering a machine learning pipeline for automating metadata extraction from longitudinal survey questionnaires, iassist quarterly 46(1), pp. 1-12. doi: https://doi.org/10.29173/iq1023 engineering a machine learning pipeline for automating metadata extraction from longitudinal survey questionnaires suparna de1, harry moss2, jon johnson3, jenny li3, haeron pereira4, sanaz jabbari2 abstract data documentation initiative-lifecycle (ddi-l) introduced a robust metadata model to support the capture of questionnaire content and flow, and encouraged through support for versioning and provenancing, objects such as basedon for the reuse of existing question items. however, the dearth of questionnaire banks including both question text and response domains has meant that an ecosystem to support the development of ddi ready computer assisted interviewing (cai) tools has been limited. archives hold the information in pdfs associated with surveys but extracting that in an efficient manner into ddi-lifecycle is a significant challenge. while closer discovery has been championing the provision of high-quality questionnaire metadata in ddi-lifecycle, this has primarily been done manually. more automated methods need to be explored to ensure scalable metadata annotation and uplift. this paper presents initial results in engineering a machine learning (ml) pipeline to automate the extraction of questions from survey questionnaires as pdfs. using closer discovery as a ‘training and test dataset’, a number of machine learning approaches have been explored to classify parsed text from questionnaires to be output as valid ddi items for inclusion in a ddi-l compliant repository. the developed ml pipeline adopts a continuous build and integrate approach, with processes in place to keep track of various combinations of the structured ddi-l input metadata, ml models and model parameters against the defined evaluation metrics, thus enabling reproducibility and comparative analysis of the experiments. tangible outputs include a map of the various metadata and model parameters with the corresponding evaluation metrics’ values, which enable model tuning as well as transparent management of data and experiments. keywords automated metadata extraction, longitudinal surveys, machine learning, model provenance, hyperparameter tuning, ddi lifecycle introduction ddi-lifecycle (ddi-l) (ddi alliance, 2014) introduced a robust metadata model to support the capture of questionnaire content and flow and encouraged through support for versioning and provenancing objects such as ‘basedon’ for the reuse of existing question items. however, the dearth of questionnaire banks including both question text and response domains has meant that an ecosystem to support the development of ddi ready computer assisted interviewing (cai) tools is limited. archives hold the information in pdfs associated with surveys but extracting that in an efficient manner into ddi-lifecycle is a significant challenge. survey specification and development tools for standards-compliant questionnaire development using ddi-l require scalable and effective methods to enable automation of cai. with a range of social sciences and biomedical domains’ longitudinal studies forming part of closer discovery (https://discovery.closer.ac.uk), it offers a rich collation of questionnaire construct definitions and measurement approaches employed over a period, as well as scope for cross-study research. due to the increased volume of data, with more questionnaires getting added to closer discovery, the ease, efficiency, and robustness of metadata extraction of question items are crucial considerations in https://doi.org/10.29173/iq1023 https://discovery.closer.ac.uk/ 2/12 de, suparna; moss, harry; johnson, jon; li, jenny; pereira, haeron and jabbari, sanaz (2022). engineering a machine learning pipeline for automating metadata extraction from longitudinal survey questionnaires, iassist quarterly 46(1), pp. 1-12. doi: https://doi.org/10.29173/iq1023 questionnaire processing, and ultimately, scaling to provide a high-quality question bank hosting survey questions for reuse by studies and data collection agencies. use of question banks would have both a utility for discovery and through the accurate reuse of questions between studies, encouraging study and analyses reproducibility. closer has been annotating metadata from the study questionnaires in ddi-l. however, much of this has been done manually or semi-manually, making the extraction of structured metadata for data management purposes burdensome. to move away from manual processing of questionnaires and enable efficiencies in the survey process, automated methods are needed for questionnaire item metadata extraction. computational methods such as machine learning (ml) techniques, especially, supervised learning algorithms are an intuitive candidate approach for automating the extraction of valid ddi items from the survey questionnaires in pdf format that form part of closer discovery. the existing processed and marked-up (in xml) questionnaires form the training and validation dataset for applying supervised ml models. the extraction of the questionnaire items can be modelled as a text classification problem, distinguishing the questions, responses and instructions etc. as specific categories. thus, this paper adopts a supervised ml approach for automated questionnaire item extraction from pdf survey questionnaires for inclusion in a ddi-l compliant repository. given a set of inputs and labelled outputs, supervised ml algorithms are geared towards allowing a model to learn over time, adjusting to minimise the error through a loss function. this necessitates a continuous build and integrate approach, with the different combinations of input data, feature engineering methods, model parameters and their resultant outputs being attached to an ml pipeline. this also requires capturing the various combinations experimented with, and the corresponding outputs, as metadata attached to the experiments, to ensure reproducibility, comparative analysis, and provenance of the pipelines. however, tracking the data and model transformations manually is time consuming and error prone. existing initiatives such as ibm’s prov-ml schema (souza, azevedo, et al., 2019), that incorporate ml model aspects into the w3c prov-o recommendation (w3c, 2013), is a promising development in this regard. the open-source provlake (souza, mattoso, et al., 2019) python library enables collection of provenance data related to function calls, with input arguments and output values captured. it provides a data tracker api which can be integrated into the machine learning workflow source code. thus, this paper also showcases the integration of the abstraction of model parameters through pipelines (through data version control (dvc) (kuprieiev et al., 2021)) and automating the process of attaching metadata related to each model experiment (through provlake). the process is illustrated with an implementation of the naïve bayes ml model which is an instantiation of a probabilistic classification algorithm. tracking of model parameters, input data and output metrics is implemented in the hyperparameter tuning of the naïve bayes model where the dvc pipeline execution outputs a set of evaluation matrices, each corresponding to an individual variation in the hyperparameters. the different evaluation metrics for each parameter variation are then captured in a structured format, as a provlake log file. other metadata recorded are the runtime information such as start time and end times of each data transformation execution. longitudinal survey questionnaires dataset dataset source the dataset source was closer discovery (https://discovery.closer.ac.uk). the content is generated by a collaboration between closer and its partner studies (johnson, 2021) and contains metadata in ddi lifecycle 3.2 for the questionnaires, datasets, study and data collection level information and a set of topics which are assigned to each question and resultant variable. https://doi.org/10.29173/iq1023 https://discovery.closer.ac.uk/ 3/12 de, suparna; moss, harry; johnson, jon; li, jenny; pereira, haeron and jabbari, sanaz (2022). engineering a machine learning pipeline for automating metadata extraction from longitudinal survey questionnaires, iassist quarterly 46(1), pp. 1-12. doi: https://doi.org/10.29173/iq1023 training dataset preparation the training dataset was extracted using python 3 (van rossum & drake, 2009) code (li, 2021) from closer discovery utilising the colectica repository rest api (colectica, 2021). the dataset included the ddi-l 3.2 items: ● question item, question text (questionname, questionliteral), ● response domain items (codedomain, textdomain, numericdomain, and datetimedomain) ● conditionals (ifconditional, loopwhile) ● interviewer instructions (instructiontext) ● statements (literal) ● associated urns for all of the above the urns allow the tracking and subsequent analysis of predicted values from different models. the extracted train and test dataset was output as a tab-separated-value (tsv) file, to serve as input into the machine learning program. the entire process is illustrated in figure 1 below. https://doi.org/10.29173/iq1023 4/12 de, suparna; moss, harry; johnson, jon; li, jenny; pereira, haeron and jabbari, sanaz (2022). engineering a machine learning pipeline for automating metadata extraction from longitudinal survey questionnaires, iassist quarterly 46(1), pp. 1-12. doi: https://doi.org/10.29173/iq1023 figure 1: a schematic view of the training dataset generation from the closer discovery metadata store and model training. dataset description the dataset used as input into the machine learning models contains 187,105 rows, with the following item types (and their respective distributions), as shown in table 1 below. https://doi.org/10.29173/iq1023 5/12 de, suparna; moss, harry; johnson, jon; li, jenny; pereira, haeron and jabbari, sanaz (2022). engineering a machine learning pipeline for automating metadata extraction from longitudinal survey questionnaires, iassist quarterly 46(1), pp. 1-12. doi: https://doi.org/10.29173/iq1023 table 1: distribution of items in training dataset item type count questionname 40,546 question literal 40,545 interviewer instruction 3,051 statement 7,551 response domain codelist 85,547 response domain datetime 486 response domain text 613 conditional 8,414 loop 352 total 187,105 closer ml pipeline git (git, 2021) is used for version control of the underlying code used to pre-process input data, generate features for training from the data and for model training and evaluation. additionally, text outputs of experiments and basic plots are versioned with git. in this structure, each broad model family occupies a branch, with individual experiments represented by a directory containing output files following model training and evaluation. the ml pipeline relies upon dvc to perform both dataset versioning and experiment tracking. in this context, an experiment refers to the model training process: from the choice of input parameters to the performance of the trained model on validation data according to several metrics of interest, such as accuracy, precision, recall, f1-score and the area under the receiver operating characteristic curve, referred to here as auc score and roc curve. datasets versioned by dvc are referenced in the versioncontrolled codebase, managed by git, and transferred via secure shell (ssh) to remote storage on the university college london (ucl) research data storage service (ucl, 2021). experimental setup, ml model and code versioning dvc is used to version the state of the input data as it evolves and associate that with a specific git commit hash. dvc is also used within this work to transfer the versioned data over ssh to remote storage, from which it is accessible to other authorised users on the general-purpose highperformance computing (hpc) cluster at ucl. model training is performed using a node on the ucl hpc cluster with a 36 core intel(r) xeon(r) gold 6140 cpu @ 2.30ghz and attached nvidia v100 gpu. a local environment with one amd ryzen 5 3600 cpu and one nvidia rtx 3070 gpu was additionally used for development purposes. https://doi.org/10.29173/iq1023 6/12 de, suparna; moss, harry; johnson, jon; li, jenny; pereira, haeron and jabbari, sanaz (2022). engineering a machine learning pipeline for automating metadata extraction from longitudinal survey questionnaires, iassist quarterly 46(1), pp. 1-12. doi: https://doi.org/10.29173/iq1023 datasets and trained models are stored, under dvc versioning, on the ucl research data storage service (rdss). the rdss is a petabyte-scale storage facility intended for research data storage in ongoing projects, includes data backup and is accessible via the hpc cluster described previously. dataset versioning is handled by dvc, which calculates a 32-character md5 hash for each file within a directory and uses a json format file to record the relative locations of files within directories. the resulting hashes are stored in dvc configuration (.dvc) files, which are added to git commits in order to pair input datasets to the relevant version of the codebase. an overview of the model, dataset and code versioning is provided below in figure 2. figure 2: a schematic view of the relationship between git and dvc version control and the local code repository. all code, documentation, input parameters, and dvc configuration files are versioncontrolled via git and stored in a remote repository. input datasets and trained models are versioned by dvc and are stored in a separate remote repository. the codebase is entirely python-based, relying on the pytorch and scikit-learn machine learning libraries, and is broadly structured as follows, with some differences occurring between different model choices. in the top level of the repository, steering code exists to launch the relevant workflow via function calls and relies upon the correct setting of values in the parameters yaml (yaml ain't markup language) file. the yaml file dictates the type of dataset to process, model type, output directory and file names and a range of hyperparameters for model training. https://doi.org/10.29173/iq1023 7/12 de, suparna; moss, harry; johnson, jon; li, jenny; pereira, haeron and jabbari, sanaz (2022). engineering a machine learning pipeline for automating metadata extraction from longitudinal survey questionnaires, iassist quarterly 46(1), pp. 1-12. doi: https://doi.org/10.29173/iq1023 hyperparameter tuning is concerned with choosing a set of optimal parameters which define the model architecture. in contrast with model parameters, hyperparameters cannot be directly trained from the data. the performance of a model can significantly vary according to the hyperparameters’ values. since finding the ideal hyperparameter can be time-consuming, search algorithms such as grid search and random search are typically employed. the key distinction between these two methods is the requirement to test all parameters. randomsearchcv works with a few 'random' combinations (though typically with a set number of iterations) out of all the available combinations, whereas gridsearchcv scans all possible combinations. both gridsearchcv and randomsearchcv feature are functions available in scikit-learn’s model selection package and fit the model to the training set by looping through predefined hyperparameters. both functions use the cross-validation method which divides the train data further into two parts the train data and validation data to test the model for all possible combinations of the values given in the dictionary. once the correct workflow and initial arguments are provided, raw data consisting of text content and an item type category is read and stored as a pandas (reback et al. 2021) dataframe object. input raw data is cleaned by removing underscores, renaming labels via regular expression matching and the removal of entries containing null entries for either text content or item type label. following data cleaning, output directories for the experiment are created and the cleaned data is saved. cleaned data is then passed as an argument to a function specific to each model type. at this stage, the model family is determined, and features are generated from the cleaned dataset in a specific manner for that model type. data is split into ‘train’, ‘validation’ and ‘test’ sets, with a ratio defined in the input yaml. text data is then tokenized and encoded, a process that converts individual words in the dataset into integer indices from a vocabulary used by the model. the data is further converted into a relevant format in the case of neural network-based models before commencing model training, using the model type defined in the input parameters yaml file. following model training for a userdefined number of iterations, model performance is evaluated against the validation dataset. output metrics and plots from model evaluation are saved to the output directory of the model along with a record of the input parameters used to obtain results. model runs are initiated via the dvc ‘project’ feature. input parameters are provided by a yaml file (params.yaml) in a nested format, with the top level representing a ‘stage’ of the dvc project. by providing the dvc api with a project name, python file dependencies, expected output files, a file defining input parameters and a python script to run, dvc configuration files are produced following model training that records all of the above information and associates it with the project name, enabling reproducibility of results. in addition, when output metrics are tracked under dvc, running the same ‘experiment’ with differing parameters will display the effect on performance. the results of model training and inference are stored either via git on a remote github repository or in the case of trained models, via dvc on the ucl rdss. this project utilises the git branch structure to separate different model families within the project, with separate output directories denoting different experiments. a pytest (krekel et al. 2004) test suite is provided with the code to ensure changes to the code are error-free and can reproduce expected results after model training on a small test dataset. changes to the code in the remote repository are performed using a continuous integration service via github, which sets up a typical environment with which to run code tests and notifies the user if changes to the code will introduce breaking changes. in this way, it is also possible to check if changes to the code which at first may appear to be error-free in fact introduce changes to model behaviour. model experiments an example of the entire ml pipeline for the case of the multinomial naïve bayes classifier follows, with an overview given below in figure 3. multinomial naïve bayes is a supervised learning https://doi.org/10.29173/iq1023 8/12 de, suparna; moss, harry; johnson, jon; li, jenny; pereira, haeron and jabbari, sanaz (2022). engineering a machine learning pipeline for automating metadata extraction from longitudinal survey questionnaires, iassist quarterly 46(1), pp. 1-12. doi: https://doi.org/10.29173/iq1023 classification method that can be applied for categorical text data analysis. it is based on the assumption that each feature being classified is independent of all others and works by calculating the probability of each tag for any given text input, with the tag with highest probability forming the output. figure 3: a schematic view of the stages of arbitrary model training from data ingest and cleaning to output of model results and their storage under version control via git and dvc. input parameters, the dataset of choice and the ‘multinomialnaïvebayes’ model are selected in the yaml file. steering code is called which reads the parameters yaml file, performs data cleaning and sets up the directory structure to store model outputs. the cleaned data and parameters are passed as arguments to a specific naïve bayes classifier function, which handles feature generation, transformation of each text item into term frequency–inverse document frequency (tfidf) vector representations and splits the data into training and test subsets. multinomial naïve bayes classifiers are used in a one-vs-rest strategy, fitting one classifier per class. performance is assessed during model training using a k-folds cross validation approach, where ‘k’ is the number of ‘folds’. this approach defines a different subset of the training data as validation data in each ‘fold’ which is then used to assess the loss of the trained model. after validating model performance over all k-folds, the model is https://doi.org/10.29173/iq1023 9/12 de, suparna; moss, harry; johnson, jon; li, jenny; pereira, haeron and jabbari, sanaz (2022). engineering a machine learning pipeline for automating metadata extraction from longitudinal survey questionnaires, iassist quarterly 46(1), pp. 1-12. doi: https://doi.org/10.29173/iq1023 then trained on the entire training dataset and its performance evaluated with the held-back test dataset. model metrics such as the accuracy, precision, recall, f1-score, auc score, roc curve plot and confusion matrix plot are saved to the output directory. metrics scores are saved as nested json files, a per-class report and confusion matrix are saved as csv files while roc curve and confusion matrix plots are saved as svg files. the trained model is exported as a python joblib file and added to dvc version control, before manually being sent over ssh by the user to the rdss. all other outputs are version controlled via git and pushed to the remote github repository. hyperparameter tuning is performed for the parameter var_smoothing which is a parameter added to the distribution’s variance in the gaussian naïve bayes model using gridsearchcv. since the naïve bayes algorithm, with its gaussian distribution assumption essentially gives more weights to the samples closer to the distribution mean, var_smoothing adds a user-defined value to the distribution’s variance to account for more samples that are further away from the mean. the algorithm cross validates the model using each value for the parameter and outputs the best parameter combination with the help of the best_params_ built-in function. the accuracy, precision, recall and f1 score are calculated for each hyperparameter adjustment, and the ideal value is calculated. provenance tracking is included in the hyperparameter tuning code, which keeps track of the parameter list and its associated evaluation metrics. this information is encoded in a provlake standard json format and is saved in a prov ml log file. the provml log file is stored in the dvc directory with a standard naming pattern of ‘prov’ followed by the workflow name and the workflow execution start time. the prov log file includes runtime details, input, and output data values in a standard format for every function in a workflow; additionally, we can keep separate prov log files for separate workflows, making analyzing data variation after each function execution much easier. for example, the log details of hyperparameter tuning function includes information such as unique id for that function, start time, generated time, end time, function name, the status of the run, input parameters (list of hyperparameter value used), and output parameters (list of evaluation metrics for each hyperparameter variation). the structured format of the provenance log file is illustrated in figure 4. https://doi.org/10.29173/iq1023 10/12 de, suparna; moss, harry; johnson, jon; li, jenny; pereira, haeron and jabbari, sanaz (2022). engineering a machine learning pipeline for automating metadata extraction from longitudinal survey questionnaires, iassist quarterly 46(1), pp. 1-12. doi: https://doi.org/10.29173/iq1023 figure 4: a snippet of the input and output data values as captured in a provlake json log file. each row of the evaluation metrics’ values map to the corresponding var_smoothing hyperparameter value. the prov log file can be examined to gain a general understanding of the major functions used, the data variation they produce, as well as their particular runtime and fundamental run characteristics. in summary, these prov log files assist in keeping track of data and can be shared across team members to gain a thorough understanding of data flow. conclusions source code versioning, through tools such as git, is well established in the software engineering community. but there are additional challenges in machine learning and data science, which require data version control as well as managing changes to models and datasets. towards meeting this challenge, this paper has showcased a reproducible ml model training and execution method, which also generates logging metadata in a structured format, enabling tracking of various combinations of input data, model features and hyperparameter tuning with the obtained output values. the developed ml pipeline is applied to automate the extraction of data and metadata from longitudinal survey questionnaires through a supervised machine learning pipeline approach. the pipeline approach employs git for code versioning, dvc for model and data files versioning and links these proprietary methods in ml model versioning to an open provenance standard (provlake). https://doi.org/10.29173/iq1023 11/12 de, suparna; moss, harry; johnson, jon; li, jenny; pereira, haeron and jabbari, sanaz (2022). engineering a machine learning pipeline for automating metadata extraction from longitudinal survey questionnaires, iassist quarterly 46(1), pp. 1-12. doi: https://doi.org/10.29173/iq1023 in addition to the aggregate measures, the rich provenance structure of the ddi-lifecycle schema, allow the analysis of prediction of specific item types (e.g. question text) through the urns linked to each input, which are inked to the specific questionnaire, study or if tagged to an ontology from which it originated, to give insights into where the predictions are more or less robust. these insights can be used to inform improvements in the models being used, and the potential need for more training data of specific types, the effects of different hyperparameter tuning approaches across both the whole training dataset and within specific subsets. this further enables future comparative analyses, transparent data management, effective execution of experiments as well as provenance of the ml experiment settings. funding machine learning to enhance metadata in cohort studies. science and technology facilities council st/s003916/1 automating capturing structured content from questionnaires. esrc es/k000357/1 references colectica (2021). colectica. available at: https://www.colectica.com. accessed 4 oct. 2021. ddi alliance (2014). ddi lifecycle 3.2 [computer software]. available at https://ddialliance.org/specification/ddi-lifecycle/3.2/. accessed 4 oct. 2021. git (2021). git – fast-version-control. available at: https://git-scm.com. accessed 23 nov. 2021. krekel, h. et al. (2004). pytest 6.2.2. computer software kuprieiev, r. et al. (2021). dvc: data version control git for data & models (2.3.0) doi:10.5281/zenodo.4892897 johnson, j. (2021). managing multiple uk longitudinal studies. [online] zenodo. available at: https://doi.org/10.5281/zenodo.3775272. accessed 4 october 2021. li, j. (2021). python interface to the colectica api (version 1.0). computer software. https://github.com/closer-cohorts/colectica_api. accessed 4 october 2021. reback, j. et al. (2021). pandas-dev/pandas: pandas 1.2.1 (v1.2.1). zenodo. doi:10.5281/zenodo.4452601 rogers, f.b. (1963). medical subject headings. bulletin of the medical library association, [online] 51, pp.114–116. available at: https://pubmed.ncbi.nlm.nih.gov/13982385/. accessed 4 oct. 2021. ucl (2021). research data storage service (research software repository). available at https://www.ucl.ac.uk/isd/services/research-it/research-data-storage-service. accessed 14 oct. 2021. souza, r., azevedo, l., lourenço, v., soares, e. f. d. s., thiago, r., brandão, r., netto, m. a. s. (2019). provenance data in the machine learning lifecycle in computational science and https://doi.org/10.29173/iq1023 https://www.colectica.com/ https://ddialliance.org/specification/ddi-lifecycle/3.2/ https://git-scm.com/ https://doi.org/10.5281/zenodo.4892897 https://doi.org/10.5281/zenodo.3775272 https://github.com/closer-cohorts/colectica_api https://doi.org/10.5281/zenodo.4452601 https://doi.org/10.5281/zenodo.4452601 https://pubmed.ncbi.nlm.nih.gov/13982385/ https://www.ucl.ac.uk/isd/services/research-it/research-data-storage-service 12/12 de, suparna; moss, harry; johnson, jon; li, jenny; pereira, haeron and jabbari, sanaz (2022). engineering a machine learning pipeline for automating metadata extraction from longitudinal survey questionnaires, iassist quarterly 46(1), pp. 1-12. doi: https://doi.org/10.29173/iq1023 engineering, in works 2019 workflows in support of large-scale science co-located with sc 2019 acm/ieee international conference for high performance computing, networking, storage, and analysis, denver, usa. souza, r., mattoso, m., azevedo, l., thiago, r., soares, e. f. d. s., santos, m. n. d., valduriez, p. (2019). efficient runtime capture of multiworkflow data using provenance, in escience 2019 15th international escience conference, san diego, united states. van rossum, g. & drake, f.l. (2009). python 3 reference manual, scotts valley, ca: createspace. w3c (2013). prov-o: the prov ontology. w3c recommendation. [online]. available at: http://www.w3.org/tr/prov-o/. accessed 14 oct. 2021. endnotes 1 suparna de is a lecturer in computer science at the university of surrey and can be reached by email: s.de@surrey.ac.uk. (version: december 2021) 2 harry moss and sanaz jabbari are with the centre for advanced research computing, ucl. 3 jon johnson and jenny li are at closer, ucl social research institute. 4 haeron pereira is in the department of computer science at the university of surrey. https://doi.org/10.29173/iq1023 http://www.w3.org/tr/prov-o/ mailto:s.de@surrey.ac.uk failure as the treatment for transforming complexity to complicatedness 1/2 rasmussen, karsten boye (2018) editor’s notes: failure as the treatment for transforming complexity to complicatedness, iassist quarterly 42 (4), pp. 1-2. doi https://doi.org/10.29173/iq949 editor's notes failure as the treatment for transforming complexity to complicatedness welcome to the fourth issue of volume 42 of the iassist quarterly (iq 42:4, 2018). the iassist quarterly presents in this issue three papers. when you know how, cycling is easy. however, data for cycling infrastructure appears to be a messiness of complications, stakeholders and data producers. the exemplary lesson is that whatever your research area there are often many views and types of data possible for your research. and the fuller view does not make your research easier, but it does make it better. the term geospatial data covers many different types of data, and as such presents problems for building access points or portals for these data. the second paper also brings experiences with complicated data, now with a focus on data management and curation. i would say that the third paper on software development in digital humanities is also about complicatedness, but this time the complicatedness was not overcome. maybe here complexity is a better choice of word than complicatedness. in my book things are complex until we have solved how to deal with them; after that they are only complicated. the word failure is even among the keywords selected for this entry. again: read and learn. you might learn more from failure than from success. i find that sir winston churchill is always at hand to keep up the good spirit: ‘success consists of going from failure to failure without loss of enthusiasm’. from canada comes the paper ‘cycling infrastructure in the ottawa-gatineau area: a complex assemblage of data’ that some readers might have seen in the form of a poster at the iassist 2018 conference in montreal. the authors are sylvie lafortune, social sciences librarian at carleton university in ottawa, and joël rivard, geography and gis librarian at the university of ottawa. the article is a commendable example of how to encompass and illuminate an area of research not only though data but also by including the data producers and stakeholders, and the relationships between them. the article is based upon a study conducted in 2017-2018 that explored the data story behind the cycling infrastructure in ottawa, canada’s capital city; or to be precise, the infrastructure of the cycling network of over 1,000 km which spans both sides of the ontario and quebec provincial boundary known as the ottawa-gatineau national capital region. the municipalities invest in cycling infrastructure including expanded and improved bike lanes and paths, traffic calming measures, parking facilities, bike-transit integration, bike sharing and training programs to promote cycling and increased cycling safety. the research included many types of data among which were data from telephone interviews concerning ‘who, where, why, when, and how’ in an origin-destination survey, data generated by mobile apps tracking fitness activities, collision data, and bike counters placed in the area. the study shows how a narrow subject topic such as cycling infrastructure is embedded in complicated data and many relationships. ningning nicole kong is the author of ‘one store has all? – the backend story of managing geospatial information toward an easy discovery’. many libraries are handling geographical information and my shortened version of the abstract from the article promises: geoblacklight and opengeoportal are two open-source projects that initiated from academic institutions, which have been adopted by many universities and libraries for geospatial data discovery. the paper provides a summary of geospatial data management strategies by reviewing related projects, and focuses on best management practices when curating geospatial data. the paper starts with a historical https://doi.org/10.29173/iq949 2/2 rasmussen, karsten boye (2018) editor’s notes: failure as the treatment for transforming complexity to complicatedness, iassist quarterly 42 (4), pp. 1-2. doi https://doi.org/10.29173/iq949 introduction to geospatial datasets in academic libraries in the united states and also presents the complicatedness involved in geospatial data. the paper mentions geoportals and related projects in both the united states and europe with a focus on opengeoportal. nicole kong is an assistant professor and gis specialist at purdue university libraries. sophie 1.0 was an attempt to create a multimedia editing, reading, and publishing platform. based at the university of southern california with national and international collaboration, sophie 2.0 was a project to rewrite sophie 1.0 in the java programming language. the author jasmine s. kirby gives the rationale for the article ‘how not to create a digital media scholarship platform: the history of the sophie 2.0 project’ in the sentence: ‘understanding what went wrong with sophie 2.0 can help us understand how to create better digital media scholarship tools’. for the first time we now have failure among the keywords used for a paper in iq. the institute of the future of the book (ifb) was a central collaborator in the development of the sophie versions. the ifb describes itself as a thinkand-do tank and it is doing many projects. the kirby paper gives us a brief insight into the future of reading, starting from basic e-books in the 1960s. when you read through the article you will note caveats like lack of focus on usability and changing of the underneath software language. the article ends with good questions for evaluating digital scholarship tools. submissions of papers for the iassist quarterly are always very welcome. we welcome input from iassist conferences or other conferences and workshops, from local presentations or papers especially written for the iq. when you are preparing such a presentation, give a thought to turning your one-time presentation into a lasting contribution. doing that after the event also gives you the opportunity of improving your work after feedback. we encourage you to login or create an author login to https://www.iassistquarterly.com (our open journal system application). we permit authors 'deep links' into the iq as well as deposition of the paper in your local repository. chairing a conference session with the purpose of aggregating and integrating papers for a special issue iq is also much appreciated as the information reaches many more people than the limited number of session participants and will be readily available on the iassist quarterly website at https://www.iassistquarterly.com. authors are very welcome to take a look at the instructions and layout: https://www.iassistquarterly.com/index.php/iassist/about/submissions authors can also contact me directly via e-mail: kbr@sam.sdu.dk. should you be interested in compiling a special issue for the iq as guest editor(s) i will also be delighted to hear from you. karsten boye rasmussen february 2019 https://doi.org/10.29173/iq949 https://www.iassistquarterly.com/index.php/iassist/about/submissions 1/2 akullo, winny nekesa & buwule, robert stalone (2021), guest editors' notes, iassist quarterly 45(3-4), pp. 1-2. doi: https://doi.org/10.29173/iq1026 guest editors' notes this special issue has nine papers selected from the africa regional workshop at makerere university (kampala, uganda) on january 11th to 13th 2021. the first two papers relate to research data management (rdm). the first one analyses the authorship, volume, visibility, and quality of publications on rdm in sub-saharan africa. the analysis was done using bibliometrics focusing on rdm publications from, and on, sub-saharan africa which are currently indexed in google scholar. the second article presents available open rdm resources for different data practitioners, particularly researchers and librarians at the university of dodoma, in tanzania. some of the rdm resources discussed in this paper are data management plan (dmp) and a data repository available for researchers to freely archive and share their research data with the local and international communities. the third paper highlights the data-sharing attitudes and behaviors of african data curators and data management experts. the paper compares data from an earlier study and analyses the new findings between the data sharing attitudes and behaviors between africans and non-africans. the fourth paper articulates the data literacy integration agenda and how it can catalyze the achievement of sustainable development goals. the paper unpacks the role of data literacy in catalyzing the achievement of the sustainable development goals (sdgs), challenges faced, and suggests recommendations to the challenges. it is however sad to note here that the author of this paper recently passed on 15th december 2021. may the good lord accord gorreti an eternal rest. the fifth paper discourses the establishment of a data center at mzuzu university library in malawi after the unfortunate fire outbreak of 2015 that destroyed the whole library. interesting models are drawn in the paper like; the six-month process of restoring an interim library and the designing & construction of the new library in collaboration with the virginia technological school of architecture & design in the united states. the sixth paper goes further to examine the growth and development of institutional repositories in the east african countries of kenya, tanzania and uganda. the paper contextualizes and discusses in detail the drivers and barriers to the development of institutional repositories in east africa such as: policy formulation, financial support, training, infrastructure, open access awareness among others. the seventh paper focuses on the learning outcomes in literacy and numeracy in uganda in the light of maternal education. in this paper, deeper analysis was conducted on the data mined from the uwezo assessment data to show the effect of the mothers’ education on the numeracy and literacy learning outcomes among children in uganda. the eighth paper illuminates the opportunities and risks of sharing agricultural research data in tanzania. stimulating themes on sharing of research data are developed and discussed in this paper such as: research collaboration, transparency, accuracy, funding, policy, institutional, and government support among others. finally, the ninth and last paper narrates the data dissemination process at the uganda bureau of statistics (ubos). the paper presents in detail the methods, channels of data sharing such as: https://doi.org/10.29173/iq 2/2 akullo, winny nekesa & buwule, robert stalone (2021), guest editors' notes, iassist quarterly 45(3-4), pp. 1-2. doi: https://doi.org/10.29173/iq1026 workshops, websites, libraries, resource centres, social media, and the physical delivery of print resources to the ubos partners and clients. winny nekesa akullo and robert stalone buwule https://doi.org/10.29173/iq vol223 4 iassist quarterly introduction one of the dominant features of the overall socio-economic development since the end of the second world war and especially exhilarating one during the last decades has been the process of european integration. the process which started as a relatively humble goal of preventing for the future any new devastating wars in europe gradually has become one of the dominant factors not only of the whole development in europe but to some extent also worldwide and one of the best examples of the gradual regionalization and globalization of the contemporary world which has become to some extent a model emulated all over the world by various regional and sub-regional clusters of countries. in the next parts of this paper we will deal in more details with some specific features of this process in the context of the contemporary development trends in the european integration vis-a-vis the processes of the eu enlargement to the countries of central and eastern europe and especially with the role of the modern information technologies in the support of these processes. present status of the european union development and the challenges of its enlargement to central and eastern europe during the last over fifty years of its existence, the european union (eu) has been passing through the development process which could be characterized by at least two main characteristics. they have been as follows: a) gradual and steady enlargement of the eu b) systematic increase and development of the common institutions, legislation, various rules and regulations of common policies and all of them leading to further accelerated development but also strengthening of the eu as a whole and its member states individually. in case of necessity, the development of individual relatively weaker member states has been supported through various programs and funds of particular common policies and funds (agriculture, cohesion, infrastructure, etc.) a) in this respect, the eu has developed from its original six members in the mid of 1950s (france, belgium, the netherlands, luxembourg, italy and germany) through gradual enlargements to the existing 15 members union. in addition to the original six members, the present eu includes another nine members who joined the union later on as the united kingdom, ireland, denmark in 1970s, greece, spain, portugal in 1980s and austria, finland, sweden in 1990s. in view of this its gradual development, the eu has become one of the strongest economic powers in the world and in many respects even the most strongest at all. this strength has been increasing not only by each new member but also by the synergic effect of their mutual cooperation and integration b) all above processes of mutual cooperation and integration in no way should be understood as a simple consequence of the gradual process of enlargement of the eu covering at present with few exemptions (norway, switzerland) the most advanced part of the continent. at least as important as the enlargement itself and to some extent even more important has been the gradual development of the institutions, legislation and various other common policies, funds, rules and regulations of the eu which in their mutual interaction have further substantially contributed to the acceleration of the overall development of the union. in this respect we could come to the conclusion that in many ways the eu has gradually developed the unique system of international institutions, legislation, “law”, etc. which of course has to be fully respected not only by all member states but also by all partners outside the eu and in particular by all candidates for the future membership. in this respect the most important institutions of the eu which have been overseeing the overall development of the union and in this respect also being institutions producing the particular rules and regulations of the common policies have been: the european commission global access and local support to the processes of european integration in central and eastern europe through global networking by dusan soltes * winter1998 5 the european council the european parliament the european court of justice these basic four institutions of the eu are further supported by a large number of various other, more specialized and/or technically oriented institutions, agencies, etc. altogether, it represents a staff of more than 20,000 international civil servants serving in three main hubs of the eu i.e. brussels, luxembourg, strassburg and having their official representatives and missions all over the world. of these over 20,000 staff, more than 16,000 have been working directly at the european commission at brussels as the most important and powerful executive arm of the eu providing and executing the day-to-day functioning of the eu internally but also externally and sometimes unjustly referred as “bureaucrats” or “eurobureaucracy” due to the enormous number of various legislation, rules and regulations they have been permanently producing for every aspect of the community life and activities. if in view of the above we would try at least very briefly to assess the challenges of the current process of enlargement of the eu to the central and eastern europe we have to take into account several important aspects. all candidate countries for this forthcoming enlargement are former socialist countries which emerged in the end of the 1980s as new future democracies and market economies. since that time on both sides i.e. the eu as well as these new democracies it was clear that the future development of the eu has to proceed towards its further enlargement by these ten new candidate countries if the europe would like to eliminate any negative consequences of its further and/or continuing division. however, due to the completely different political and socio-economic development in both parts of the europe i.e. in the eu itself on the one side and in the candidate countries on the other, the process of unification and integration of the whole europe has been much more complex, time consuming, fund demanding, etc. than it had been assumed in the end of 1980s and beginning of 1990s when the necessary political preconditions for this kind of integration has been created. among various other differences as e.g. much lower level of the overall development in the candidate countries (only about 6% of the gdp of the eu while by the population it is more than 1/3 of its current population), the main problem is the necessity to prepare the candidate countries for their future in the eu in the terms of legislation and all various community rules and regulation. what in general has been named as a process of approximation and harmonization of legislation of the candidate countries to the so-called “acquis communnautaire” of the eu or its community legislation in its broadest sense. role of global networking in the process of the future enlargement of the european union in order to understand the complexity of the above processes of the approximation and harmonization of the legislation of the candidate countries with the totally different legislation from their former “socialist” past with the “acquis” we have to realize that the later one represents an enormous volume of various legislative acts which in general consists of two main parts: primary legislation all basic treaties by which the eu and before that the european communities have been established i.e. the paris treaty, two rome treaties, maastricht and amsterdam treaties respectively represent the basic part of the “acquis” to which every new member has to access in full and without any exemptions. that of course requires from them to have also the suitable technical tools and means for such an access secondary legislation an enormous amount of various legislative acts, norms, regulations and standards permanently enacted by the eu for the day-to-day running of the eu and mostly initiated by the european commission. in this case again the candidate countries have to adopt and implement all of them in their national legislation. but in difference to the primary legislation, in this case with the possibility of some flexibility in adoption regarding some of them. what on one hand makes the process of harmonization and approximation more flexible but on the other hand it is also more challenging and more complex in seeking the most optimal ways of the particular adaptation according to the national needs and priorities. but in any case it again requires an efficient and direct access to the particular sources of the secondary legislation. in view of the above it is clear that the whole process is very closely related to the efficient utilization of the modern contemporary information and communication technologies and their global networking. it has been a historical coincidence that the processes of european integration in the central and eastern europe have started and been proceeding in the environment of the acceleration of the global networking and direct access to the remote data bases and at the same time of a possibility to use networking also for the local support to these processes. as we will demonstrate in the following parts of this paper, some activities in this respect would almost be impossible if this particular “global” access and “local” support through the contemporary networking technologies would be not existing. in this connection we have to realize that during its over fifty year existence and mainly due to its relatively extensive institutional framework (in particular regarding 6 iassist quarterly the european commission), the eu has up to now accumulated an enormous amount of “acquis” and related and/or derived or supporting information, documents, acts, etc. according to the latest available edition it represents a list of: 51 specialized data bases of what 46 have been on-line data bases and 5 cd-roms 4 world wide web servers in addition there have been a long list and ever growing number of other specialized on-line and offline data bases, new editions of cd-roms and www pages e.g. in the taiex office of the european commission as an office created in 1995 for providing technical assistance to the candidate countries in approximation of legislation according to the white book but also in mastering the latest information and communications technologies. in addition we have to realize that the amount of these data sources and data bases has been further increased by the fact that in accordance with the community rules all of them are either in all 11 official languages of the eu (every member country’s language is an official language of the eu) or in some of them (mostly english, french, german) with at least annotations in all other. the list of the eu data bases is as follows: abel document delivery of official journal l and c series agrep agricultural research projects of the eu apc commission preparatory acts bach harmonized company acts ccl-train common command language training database celex community legislation, case-law, preparatory acts, parliamentary questions, national provisions implementing directives comext intra and extra eu trade (also on cd rom) cordis research and development information service (also cd rom) cordis rtd acronyms research and technology development cordis rtd-com documents commission’s initiatives cordis rtd-contacts contact points of cordis database cordis rtd-eoi expressions of interest cordis rtd-news latest news on rtd cordis rtd-partners partner search service cordis rtd-programmes eu-funded research programmes cordis rtd-projects details of projects cordis rtd-publications abstracts of eu publications cordis rtd-results information on results and prototypes ecdin environmental chemical data (eu, usa, japan) echo news echo news for users eclas european commission library ecu european currency unit emire european employment and industrial relations epistel european parliament press information system epoque european parliament on-line query system eurohistar european historical archives euristote academic research on european integration eurocron general european union statistics eurodicautom directory of terminology eurofarm cd-rom statistics on agricultural holdings eurolib-per collective catalogue of periodicals eurostat cd-rom electronic statistical yearbook of the eu winter1998 7 htcor-db high-temperature database htm-db high temperature materials database i&t magazine industry, telecoms and information market i’m guide information market guide info 92 european internal market and its social dimension iuclid classification and evaluation of existing substances new cronos macroeconomic statistical database oil weekly oil bulletin ovide information service of the european parliament panorama cd-rom panorama of eu industry rapid up-to-date information on eu activities regio regional statistics rem radioactivity environmental monitoring scad community documentation access system sesame energy technology research projects ted tenders electronic daily thesauri structured vocabularies tide technology initiative for disabled and elderly people although not all of the above databases have the same importance for the functioning of the eu, its member states and/or candidate countries it is evident that their proper utilization can in many aspects contribute to the better knowledge and/or communication with the eu and its institutions. especially important it is in the case of the candidate countries as on the bases of some of the above databases they can directly be preparing for the challenges of the future membership. in this respect one of the most important databases is the celex which contains “interinstitutional” documentation system for community law. not only by this its orientation but even more by its content it is the most valuable source of information on “acquis” as in its subsystems it contains information on: legislation (primary (treaties), secondary, supplementary) commission proposals european parliament resolutions economic and social committee opinions court of auditors opinions judgments and orders opinions of advocate-general written parliamentary questions oral questions questions at questions time parliament documents this database is on-line but available against payment only and at present it contains about 200,000 entries with a growth rate of about 10,000 entries per year. the biggest advantage of the celex is the fact that it contains not only all kinds of references and accompanying information but also the official full text of all basic legal documents including primary and secondary legislation. it is evident that such an amount of information could not be effectively handled without possibilities for a global access through a global network. this is the only way how this, the most important source on the community legislation i.e. the back-bone of the whole eu can be made to be directly accessible from the member countries of the eu as well as candidate countries of the central and eastern europe. in this connection we have to realize that this global access is needed not only from the point of view of government institutions but also by any business entity and/or a person being in need to proceed and/or familiarize with the community legislation. an another important aspect of the above databases from the users point of view is the fact that many of them (directly operated by the ec) are available free of charge while some other are available only on the commercial basis. global access and local support to the processes of european integration in slovakia as a candidate country the above system of the eu databases is physically located mostly at the european commission at brussels or in some other institutions of the eu or have been operated by an official database operator under the special arrangements with the european commission. in order to be accessible 8 iassist quarterly also from the other end e.g. the candidate countries it is necessary to make the necessary arrangements and preparations also in the particular country. in the specific conditions of the slovak republic these preparations have started already in 1995 when the european agreement on association with the eu entered into force and it has become clear that it will be necessary to create all necessary preconditions for developing an efficient computer network which would enable all government institutions an access not only to the above eu databases but also it would create conditions for the local support and cooperation between individual government institutions in implementation of the europe agreement and one of its most important task i.e. approximation and harmonization of legislation. on the basis of the technical assistance from the eu, the particular computer network has been established under the phare program in the form of the computer network linking together all government departments into a computer network. this specialized computer network for the processes of european integration has the following main features: it has been operated as a specialized network within the broader “govnet” i.e. the governmental network system it consists of two hubs served by two servers. one being indicated as “cs” i.e. a central server for the network of 28 computers distributed to all central organs and departments of the state administration. the other one “iap” i.e. a central server for a network of 12 computers at the institute of approximation of laws (ial) as a specialized institution of the government for methodical and technical coordination of the specific processes of the approximation and harmonization of the legislation of the slovak republic with the legislation of the eu in order to reduce operational costs of an on-line access to the particular data bases in brussels, they have been physically directly available also in the ial in the form of cd roms with monthly updates for local access and utilization the whole system of govnet enables through the internet facilities a direct connection to the outside world especially to the institutions of the european union and in particular to the european commission in brussels as well as to all other institutions of the eu and/or also other partner international as e.g. oecd in paris, etc. the main contributions in addition to the above global access to the information sources of the eu are mainly in the area of the local support to the processes of european integration through the particular computer network. as all the government departments have been linked together to the particular “intranet” its main benefits are as follows: -direct access and support to all various specialized databases as they have been created or will be created in support of processes of european integration direct support to the processes of approximation and harmonization of laws. especially it is beneficial in case of the legislative acts having “cross-departmental” character i.e. the responsibility for their approximation lies with several departments. in such cases one of them has been acting as a coordinating “gazetteer” department for all other departments participating in the adaptation process. in particular in these cases the function of the global network as well as a local support is inevitable as only the network can create a working environment for on-line cooperation and joint activities on the same “piece” of legislation at the same time by several departments another important contribution of the global access and local support of the network is in the enormous amount of activities related the processes of translation of the “acquis” from the official language of the eu (the slovak republic in this case uses english version of the legislation) to the slovak language and then after its harmonization and approximation back being translated to english for the needs of notification and monitoring by the eu. the whole translation itself without the global access and local support of the network would be almost impossible if we take into account the size of the “acquis” and still relatively low level of english comprehension in the countries of ceec and slovakia in particular as a new country which before its independence had only very limited opportunities for international relations, etc. in view of this, an important role has been played by the computerized thesauri, key words, glossaries, etc. as prepared by the european commission for the needs of the candidate countries another important aspect of the local support to the approximation and harmonization of laws has been in the area of the global access and local support to the winter1998 9 processes of securing proper interpretation of the approximated, harmonized legislation which has to have exactly the same meaning as the official version of the same legislation of the eu. the importance of providing this “common” interpretation of the national legislation with the “acquis” has also been directly supported by the network as the whole process in full utilizes its functions and services. the particular “certificate on the compatibility” when issued by the ial as the official certification authority of the government of the slovak republic is again available at the network for future use in order to prevent any uncertainties in this respect the whole process of the harmonization, approximation of the legislation after being officially certified by the ial has to be notified to the european commission and in particular to its specialized taiex (technical assistance information exchange) office in brussels which provides not only technical assistance to the candidate countries in the harmonization of legislation but is also administering the particular data base on the progress achieved in the whole process. the particular data base monitors the development and progress achieved in the harmonization of legislation in all ten candidate countries (estonia, lithuania, latvia, poland, slovakia, czech, hungary, slovenia, bulgaria, romania) on the basis of so-called “harmonogramme” which traces the whole process of harmonization and serves directly to the evaluation of the particular candidate countries and their process of “approximation” to the eu in this single most important area i.e. legislation. of course that the global networking plays a very important role also in many other related areas of the european integration, in particular regarding e.g. the processes of preparation and in-service training of the civil servants for their new tasks vis-a-vis their new responsibilities in the processes of european integration, approximation of laws, implementation of europe agreement, etc. the particular network in this respect serves as an indispensable source of particular information but at the same time also as a set of tools and means for in service distance education and learning the same function has been played also in relation to the university education of the students specializing in european integration as e.g. at the faculty of management of the comenius university in bratislava as well as at other faculties of the same university e.g. at the faculty of law but also at some other faculties not only in slovakia but also in other candidate countries of ceec. conclusion also from these few examples it is quite evident that the role of the global networking and its functions in global access and local support to the processes of european integration in the countries of the central and eastern europe are inevitable. they make the whole process of integration much more efficient and in many cases it is not at all any exaggeration if we say that without these functions some activities would be almost impossible to carry out within any reasonable time period. otherwise, the process of the european integration and unification would be even more time consuming and hard-to-be completed if we realize that in spite of these enormous support from the contemporary information technologies the gap between the countries of the eu and the candidate countries is still rather widening than narrowing. but this issue is already beyond the scope of this paper. references: europe agreement on association between the european communities and the slovak republic, slovak chamber of trade and industries, bratislava 1994 preparation of associated countries of central and eastern europe for integration to the internal market of the union, volume 1 and 2, the commission of the european communities, brussels 1995 agenda 2000 volume i communications: for a stronger and wider union, the commission of the european communities, strassburg 1997 european union database directory, office for official publications of the european communities, ecsc-eceaec brussels, luxembourg 1995 soltes, d.: enlargement of the european union: some general trend but also current specifics and peculiarities, ceps liechtenstein conference’97, brussels, vaduz 1997 *paper presented at the 1998 iassist/css conference, yale university, new haven, usa, may 19-22, 1998. dusan soltes, faculty of management, comenius university, odbojarov str. 10, bratislava, slovakia. ph:4217-5666 702 fax:421-7-5666 703 email:dusan.soltes@mail.fm.uniba.sk mailto:dusan.soltes@mail.fm.uniba.sk 1/2 rasmussen, karsten boye (2019) editor’s notes: as open as possible and as closed as needed, iassist quarterly 43(3), pp. 1-2. doi https://doi.org/10.29173/iq965 editor's notes: as open as possible and as closed as needed welcome to the third issue of volume 43 of the iassist quarterly (iq 43:3, 2019). yes, we are open! open data is good. just a click away. downloadable 24/7 for everybody. an open government would make the decisionmakers’ data open to the public and the opposition. as an example, communal data on bicycle paths could be open, so more navigation apps would flourish and embed the information in maps, which could suggest more safe bicycle routes. however, as demonstrated by all three articles in this iq issue, very often research data include information that requires restrictions concerning data access. the second paper states that data should be ‘as open as possible and as closed as needed’. this phrase originates from a european union horizon 2020 project called the open research data pilot, in ‘guidelines on fair data management in horizon 2020’ (july 2016). some data need to be closed and not freely available. so once more it shows that a simple solution of total openness and one-size-fits-all is not possible. we have to deal with more complicated schemes depending on the content of data. luckily, experienced people at data institutions are capable of producing adapted solutions. the first article ‘restricting data’s use: a spectrum of concerns in need of flexible approaches’ describes how data producers have legitimate needs for restricting data access for users. this understanding is quite important as some users might have an automatic objection towards all restrictions on use of data. the authors dharma akmon and susan jekielek are at icpsr at the university of michigan. icpsr has been a u.s. research archive since 1962, so they have much practice in long-term storage of digital information. from a short-term perspective you might think that their primary task is to get the data in use and thus would be opposed to any kind of access restrictions. however, both producers and custodians of data are very well aware of their responsibility for determining restrictions and access. the caveat concerns the potential harm through disclosure, often exemplified by personal data of identifiable individuals. the article explains how dissemination options differ in where data are accessed and what is required for access. if you are new to iassist, the article also gives an excellent short introduction to icpsr and how this institution guards itself and its users against the hazards of data sharing. in the second article ‘managing data in cross-institutional projects’, the reader gains insight into how fair data usage benefits a cross-institutional project. the starting point for the authors zaza nadja lee hansen, filip kruse, and jesper boserup thestrup – is the fair principles that data should be: findable, accessible, interoperable, and re-useable. the authors state that this implies that the data should be as open as possible. however, as expressed in the icpsr article above, data should at the same time be as closed as needed. within the eu, the mention of gdpr (general data protection regulation) will always catch the attention of the economical responsible at any institution because data breaches can now be very severely fined. the authors share their experience with implementation of the fair principles with data from several cross-institutional projects. the key is to ensure that from the beginning there is agreement on following the specific guidelines, standards and formats throughout the project. the issues to agree on are, among other things, storage and sharing of data and metadata, responsibilities for updating data, and deciding which data format to use. the benefits of fair data usage are summarized, and the article also describes the crossinstitutional projects. the authors work as a senior consultant/project manager at the danish national archives, senior advisor at the royal danish library, and communications officer at the royal danish library. the cross-institutional projects mentioned here stretch from kierkegaard’s writings to wind energy. https://doi.org/10.29173/iq965 2/2 rasmussen, karsten boye (2019) editor’s notes: as open as possible and as closed as needed, iassist quarterly 43(3), pp. 1-2. doi https://doi.org/10.29173/iq965 while this issue started by mentioning that icpsr was founded in 1962, we end with a more recent addition to the archive world, established at qatar university’s social and economic survey research institute (sesri) in 2017. the paper ‘data archiving for dissemination within a gulf nation’ addresses the experience of this new institution in an environment of cultural and political sensitivity. with a positive view you can regard the benefits as expanding. the start is that archive staff get experience concerning policies for data selection, restrictions, security and metadata. this generates benefits and expands to the broader group of research staff where awareness and improvements relate to issues like design, collection and documentation of studies. furthermore, data sharing can be seen as expanding in the middle east and north africa region and generating a general improvement in the relevance and credibility of statistics generated in the region. again, the fair principles of findable, accessible, interoperable, and re-useable are gaining momentum and being adopted by government offices and data collection agencies. in the article, the story of sesri at qatar university is described ahead of sections concerning data sharing culture and challenges as well as issues of staff recruitment, architecture and workflow. many of the observations and considerations in the article will be of value to staff at both older and infant archives. the authors of the paper are the senior researcher and lead archivist at the archive of the qatar university brian w. mandikiana, and lois timms-ferrara and marc maynard – ceo and director of technology at data independence (connecticut, usa). submissions of papers for the iassist quarterly are always very welcome. we welcome input from iassist conferences or other conferences and workshops, from local presentations or papers especially written for the iq. when you are preparing such a presentation, give a thought to turning your one-time presentation into a lasting contribution. doing that after the event also gives you the opportunity of improving your work after feedback. we encourage you to login or create an author login to https://www.iassistquarterly.com (our open journal system application). we permit authors 'deep links' into the iq as well as deposition of the paper in your local repository. chairing a conference session with the purpose of aggregating and integrating papers for a special issue iq is also much appreciated as the information reaches many more people than the limited number of session participants and will be readily available on the iassist quarterly website at https://www.iassistquarterly.com. authors are very welcome to take a look at the instructions and layout: https://www.iassistquarterly.com/index.php/iassist/about/submissions authors can also contact me directly via e-mail: kbr@sam.sdu.dk. should you be interested in compiling a special issue for the iq as guest editor(s) i will also be delighted to hear from you. karsten boye rasmussen september 2019 https://doi.org/10.29173/iq965 https://www.iassistquarterly.com/index.php/iassist/about/submissions mailto:kbr@sam.sdu.dk instructions for authors of the iassist quarterly 1/8 lafortune,sylvie and rivard, joël (2018) cycling infrastructure in the ottawa-gatineau area: a complex assemblage of data, iassist quarterly 42 (4), pp. 1-8. doi: https://doi.org/10.29173/iq937 cycling infrastructure in the ottawa-gatineau area: a complex assemblage of data sylvie lafortune1, joël rivard2 abstract the ottawa-gatineau national capital region (canada) has a well developed and well used cycling network of over 1,000 km which spans both sides of the ontario and quebec provincial boundary. the purpose of this study is to map out the complex data landscape behind the cycling infrastructure in the national capital region (ncr), which is largely based on inter-jurisdictional cooperation and partnerships with cycling advocacy groups. the questions we try to answer are: what data are collected for cycling infrastructure and activities? who are the data producers and stakeholders? what are the relationships amongst the various data producers and stakeholders? the study reveals that the complexity of the cycling data landscape in the ncr is due to the complexity of the relationships between the various data producers and stakeholders. keywords cycling infrastructure data, cycling advocacy, cycling data stakeholders, cycling data producers, national capital region (canada), active transportation data introduction this article is based on the results of an exploratory study conducted in 2017-2018 and presented as a poster at the iassist 2018 conference in montreal, canada. the theme of the conference was ‘once upon a data point: sustaining our data storytellers’ and it provided us with a great opportunity to explore the data story behind the cycling infrastructure in ottawa, canada’s capital city. we felt that our topic was particularly timely as we had noticed a heightened interest in active transportation research over a number of years. this became apparent with an increase in requests for cycling data which, we might add, are frequently difficult to obtain. as a result, the overall goal of our study was to get a better understanding of the data collected to build and maintain the cycling network in the national capital region (ncr). we chose this geographic region for the following two reasons: 1. the ncr is an interesting and possibly unique location to examine because it is situated across two cities (ottawa and gatineau) as well as two canadian provinces (ontario and quebec). it is also managed by a federal commission which is described below. these five bodies are governed in different ways and follow their own processes for collecting, managing and sharing active transportation data. 2. most of the requests we get for cycling data are generally limited to the ncr, which is where both of our universities are based and where our users are conducting research. https://doi.org/10.29173/iq937 2/8 lafortune,sylvie and rivard, joël (2018) cycling infrastructure in the ottawa-gatineau area: a complex assemblage of data, iassist quarterly 42 (4), pp. 1-8. doi: https://doi.org/10.29173/iq937 background and purpose of the study the national capital region has a well developed and well used cycling network of over 1,000 km of cycling routes which spans both sides of the ontario and quebec provincial boundary. the cycling network started in the 1980s when the national capital commission (federal government) secured the majority of industrialized waterfront lands as public land to create public green spaces, acquired a vast area in the gatineau hills to create a federal park, and established the 203-square-kilometre national capital greenbelt around ottawa.3 over the past two decades the ottawa-gatineau area has seen continued growth in cycling infrastructure but according to citizens for safe cycling, an active advocacy group in ottawa, there has been ’decades of under-investment in active transportation which means that there is a lot of catching up to do’.4 in the past few years, governments in canada have followed the sustainable transportation development trend and shifted funding in this direction. as a result, municipalities such as ottawa have made strategic decisions to further invest in cycling infrastructure. these include expanded and improved bike lanes and paths, traffic calming measures, parking facilities, bike-transit integration, bike sharing and training programs to promote cycling and increased cycling safety. the purpose of this exploratory research is to gain a better understanding of the data landscape behind the cycling infrastructure in the ottawa-gatineau area which is largely based on interjurisdictional cooperation and partnerships with cycling advocacy groups. the questions we try to answer are: 1. what data are collected for cycling infrastructure and activities? 2. who are the data producers and stakeholders? 3. what are the relationships amongst the various data producers and stakeholders? methodology given the scope of this study, we decided to conduct a website content analysis. we began with a general web search on cycling infrastructure in the ottawa-gatineau area. this led to the web sites and planning documents issued by the governments that have jurisdiction over this region: the city of ottawa, the ville de gatineau and the national capital commission. from these webpages, other data stakeholders such as the trans committee, bike ottawa, action vélo outaouais and velogo emerged and were further explored. to validate information found on these web pages and documents, certain individuals were contacted to verify the accuracy of the information. we then compiled the types of cycling data produced and which organization was producing the data. finally, we attempted to establish the relationships between the various data producers and the stakeholders. https://doi.org/10.29173/iq937 3/8 lafortune,sylvie and rivard, joël (2018) cycling infrastructure in the ottawa-gatineau area: a complex assemblage of data, iassist quarterly 42 (4), pp. 1-8. doi: https://doi.org/10.29173/iq937 findings a. data collected the following briefly describes the various data which contribute to the cycling infrastructure in the ottawa-gatineau area, including how they are collected and who collects them. public consultations public consultations are regulatory means of getting feedback from the general public about cycling and its infrastructure in the national capital area. these include in-person and online consultations to help organizations draft reports and plan for future directions. target audiences for public consultations on cycling include residents from both ottawa and gatineau and are conducted by various organizations such as the city of ottawa, the city of gatineau and the national capital commission (ncc). rapport sur l’état du vélo à gatineau en 2015 a comprehensive report on cycling in the ville de gatineau in the province of quebec. this is produced by vélo québec which draws from a province-wide survey of cycling practices in quebec, with a sample of 400 respondents for the ville de gatineau. the report also includes an analysis of the 2005 and 2011 origin-destination surveys of the ottawa-gatineau agglomeration. origin-destination survey the origin-destination (o-d) survey, held every five years, examines the “who, where, why, when, and how” of transportation trips made by residents of the national capital region (ncr) resulting in extensive, up-to-date information on current daily trip patterns of area residents.the survey is conducted through voluntary, confidential telephone interviews over a 12-week period by a team hired by the trans committee, the organization responsible for administering the survey. the survey is conducted at the beginning of fall because during this period trip patterns are usually more stable than at other times of the year. results from the 2005 and 2011 survey are available on the trans committee webpage. the next survey will take place once the first stage of implementation of the light-rail has been implemented at the city of ottawa. user-generated active transportation data data generated by users when they sign up to use a mobile app that publicly tracks bicycle rides, runs and other fitness activities. the data can be uploaded and become part of anonymized datasets which are licensed to city planning groups. strava metro is one of the leading companies which collect user-generated active transportation data. the city of ottawa, the ville de gatineau and the ncc share a strava metro account. cartographic data these include geospatial data, maps (paper and digital) as well as interactive maps that have been created by various local organizations.the cartographic products are used to illustrate the historical, current and future cycling infrastructure in the ncr. quality of facilities measure currently, a small-scale preliminary ‘quality of facilities measure’ study is being conducted in the city of ottawa. it uses the level of traffic stress (lts) methodology to determine the actual and perceived level of safety of the cycling infrastructure. it uses road characteristics such as vehicle https://doi.org/10.29173/iq937 4/8 lafortune,sylvie and rivard, joël (2018) cycling infrastructure in the ottawa-gatineau area: a complex assemblage of data, iassist quarterly 42 (4), pp. 1-8. doi: https://doi.org/10.29173/iq937 speed, number of vehicle lanes, and the presence of parking to determine the quality for a particular segment. collision data the annual collisions report provides data on all reported collisions, including bicycles, on roads within the jurisdiction of the city of ottawa. it should be noted that during our research (spring 2018), this data was only found on the city of ottawa’s open data portal. when searching the city of gatineau’s open data portal, the data was not available. no further steps were taken to confirm whether or not this data existed for the city of gatineau. post-infrastructure implementation survey these are online surveys conducted to get a better understanding of the travel behavior of cyclists once the cycling infrastructure is established. these surveys are part of a larger consultation process conducted through meetings or email. bike counters infrastructure that collects the number of times a bicycle crosses the counter (both directions summed unless otherwise noted) at various locations in the ncr. b. organizations data producers national capital commission (ncc) the national capital commission is a federal crown corporation created by canada’s parliament in 1959 under the national capital act. the ncc is subject to the accountability regime set out in part x of the financial administration act. it reports to parliament through the minister designated as minister responsible for the national capital act. the ncc is the main federal urban planner in canada’s capital region. in this role, the ncc works in collaboration with stakeholders to enhance the natural and cultural character of the capital. the ncc manages 236 kilometres of the pathways in the ottawa-gatineau region, which extend from gatineau park, through ottawa and into the greenbelt. city of ottawa the city of ottawa’s land area covers 2,792 km2 with a population of 934,243 in 2016. it operates under the ontario municipal act and the city of ottawa act, both overseen by the ontario ministry of municipal affairs. in 2013, the city added an extensive cycling plan to the building a liveable ottawa 2031 report, which was a city-wide review of land use, transportation and infrastructure policies, launched in 2012. in 2015, the city owned and maintained a cycling network of 700 km. according to bike ottawa, the city of ottawa has budgeted approximately $25 m in 2018 for cyclingrelated projects. ville de gatineau the ville de gatineau spans a territory of 343 km2 with a population of 276,245 in 2016. it operates under the cities and towns act and the charter of ville de gatineau, both overseen by the ministère des affaires municipales et de l’occupation du territoire du québec. the city is currently creating its first cycling network master plan which is expected to be adopted in 2018-2019. this master plan will update the 2013 plan de déplacements durables. in 2015, the city offered a cycling network of 269 km. according to the 2018 ville de gatineau budget, $7.5m will be available for the https://doi.org/10.29173/iq937 5/8 lafortune,sylvie and rivard, joël (2018) cycling infrastructure in the ottawa-gatineau area: a complex assemblage of data, iassist quarterly 42 (4), pp. 1-8. doi: https://doi.org/10.29173/iq937 development of the cycling network and an extra $1.4m per year afterwards. additionally, the ville de gatineau has planned to spend $470k annually for the cost of maintenance of the cycling network. trans committee the trans committee was established in 1979 to coordinate efforts between the major transportation planning agencies of the national capital region. the committee is a neutral forum for the exchange of information on technical guidelines and best practices. in addition, it manages transportation studies and collects data for transportation planning. the six members of the committee span all three levels of government. they include the national capital commission, the ministère des transports, de la mobilité durable et de l'électrification des transports du québec, the ministry of transportation of ontario, ville de gatineau, the city of ottawa, and the société de transport de l’outaouais. funding responsibilities are shared by the six member agencies. the proportion of contributions may vary for some projects. stakeholders bike ottawa bike ottawa (also known as citizens for safe cycling) is a cycling advocacy group established in 1984 and based in ottawa. according to its website, this organization promotes “cycling as a safe, fun, and environmentally friendly form of transportation.” the organization has a strong volunteer base involved in writing letters to city councillors, participating in city consultations, gathering and analyzing data from statistic canada, the ottawa cycling plan, bike counters, transport canada as well as weather canada to promote safe cycling and cycling as active transportation. action vélo outaouais action vélo outaouais is a cycling advocacy group based in gatineau. according to its website, the focus of this group is on the planning and development of a safe cycling network in the outaouais region (including the ville de gatineau). action vélo outaouais partners with the ville de gatineau by networking with cycling groups in the gatineau area to gather feedback on cycling projects and by submitting briefs and reports to the city. it partners closely with vélo québec, a long-standing provincial cycling advocacy group which produces detailed cycling reports at the regional level in the province. the organization is also involved in promoting active transportation, recreational cycling and cyclotourism, as well as finalizing the development of the route verte (cycling infrastructure throughout the province of québec). bicycle-sharing program a service in which bicycles are made available for shared use to individuals on a very short term basis, for a fee. velogo is the official bicycle-sharing program for the ottawa-gatineau area and it uses “smart-bikes” which come equipped with real-time gps, gsm, rfid and nfc technologies. this program is a public-private partnership between cyclehop, the city of ottawa, the national capital commission and the ville de gatineau. cyclehop shares its velogo gps data with its partners. c. relationships an important insight offered by this study is that the complexity of the cycling data landscape in this area is largely due to the complexity of the relationships between the various data producers and https://doi.org/10.29173/iq937 https://en.wikipedia.org/wiki/bicycles 6/8 lafortune,sylvie and rivard, joël (2018) cycling infrastructure in the ottawa-gatineau area: a complex assemblage of data, iassist quarterly 42 (4), pp. 1-8. doi: https://doi.org/10.29173/iq937 stakeholders. most relationships were outlined above, but to get a deeper understanding of the actual collaboration between organizations, we created the diagram below. relationships amongst the producers and stakeholders of data further observations the national capital region (ncr) is a unique geographic area in canada, where two provinces, two municipalities, and one federal organization must work together to develop a safe and sustainable cycling infrastructure for the residents and visitors who travel within and between both municipalities and federal parks. the cycling data landscape is further complicated by an increasing number of stakeholders such as cycling activist groups and bike sharing companies who have an obvious interest in promoting and challenging transportation projects which involve cycling. finally, an additional challenge in the ncr is that the infrastructure must be implemented in either one or both of canada’s official languages. however, we found that the intricate nature of the region is somewhat counterbalanced by a longstanding formal culture of data sharing through the trans committee as well as an informal one between the various stakeholders (users and data producers). in an era of growing active transportation and open government, how these relationships evolve could be further explored. however, and not surprisingly, we noted that each of the three jurisdictions in the ncr operates independently and information on their cycling planning is not always publicly available in the same manner. https://doi.org/10.29173/iq937 7/8 lafortune,sylvie and rivard, joël (2018) cycling infrastructure in the ottawa-gatineau area: a complex assemblage of data, iassist quarterly 42 (4), pp. 1-8. doi: https://doi.org/10.29173/iq937 further studies on how the ncr stakeholders use and combine the various cycling datasets, more specifically the user-generated active transportation data, to improve infrastructure should also be considered. finally, further investigation on how infrastructure decisions are actually arrived at would require interviews with the data producers and stakeholders. references bike ottawa (2018) ‘bike ottawa annual report 2018’. accessed 20 april 2018 https://drive.google.com/file/d/1bzuzvxbht2sld-6ekvblvq4ate6hbkis/view city of ottawa (2013) ‘ottawa cycling plan’. ottawa: city of ottawa. accessed 22 november 2017 https://documents.ottawa.ca/sites/documents.ottawa.ca/files/documents/ocp2013_report_en.pdf national capital commission (2018) ‘capital pathway strategic plan’. ottawa: national capital commission. accessed 15 february 2018 http://ncc-ccn.gc.ca/our-plans/capital-pathway-strategic-plan o-d survey. (2018) ‘trans committee’. ottawa: trans committee. accessed 15 february 2018 http://www.ncr-trans-rcn.ca/surveys/o-d-survey/ ville de gatineau (2018) ‘gatineau, ville vélo’. gatineau: ville de gatineau. accessed 12 november 2017 https://www.gatineau.ca/portail/default.aspx?p=transport_voirie/velo ville de gatineau (2018) ‘plan directeur du réseau cyclable’. gatineau: ville de gatineau. accessed 12 november 2017 https://www.gatineauvillevelo.ca/ vélo québec (2015) ‘l’etat du vélo à gatineau en 2015’. accessed october 25 2017 https://www.gatineau.ca/docs/transport_voirie/velo/etat_velo_2015.fr-ca.pdf https://doi.org/10.29173/iq937 https://drive.google.com/file/d/1bzuzvxbht2sld-6ekvblvq4ate6hbkis/view https://documents.ottawa.ca/sites/documents.ottawa.ca/files/documents/ocp2013_report_en.pdf http://ncc-ccn.gc.ca/our-plans/capital-pathway-strategic-plan http://www.ncr-trans-rcn.ca/surveys/o-d-survey/ https://www.gatineau.ca/portail/default.aspx?p=transport_voirie/velo https://www.gatineauvillevelo.ca/ https://www.gatineau.ca/docs/transport_voirie/velo/etat_velo_2015.fr-ca.pdf 8/8 lafortune,sylvie and rivard, joël (2018) cycling infrastructure in the ottawa-gatineau area: a complex assemblage of data, iassist quarterly 42 (4), pp. 1-8. doi: https://doi.org/10.29173/iq937 end-notes 1 sylvie lafortune is a social sciences librarian at carleton university, ottawa, canada. her email address is: sylvie.lafortune@carleton.ca. 2 joel rivard is the geography and gis librarian at the university of ottawa, canada. his email address is: joel.rivard@uottawa.ca 3 canada. national capital commission. a legacy to build on, accessed november 11, 2017, http://capital2067.ca/legacy/ 4 citizens for safe cycling, 2017 ottawa report on bicycling, accessed november 11, 2017, https://bikeottawa.ca/images/cycling_reports/2017_ottawa_report_bicycling_final.pdf https://doi.org/10.29173/iq937 http://capital2067.ca/legacy/ https://bikeottawa.ca/images/cycling_reports/2017_ottawa_report_bicycling_final.pdf lassist newsletter, vol. 3, no. 1 (winter 1979) peace research and information systems carl beck university of pittsburgh in a recent article i wrote that cal in a way that will be attended "the need for information retrieval to by academic and policy colsystems in the social sciences is leagues they must be able to docuboth real and apparent, but given ment their position effectively, the ability of many researchers to their ability to be effective is gather idiosyncratic research supgoing to depend upon their ability port, the need is not perceived as to manipulate a basically uncontacute. the evidence that is not reliable information environment, perceived as acute is that few effective access to information is scholars are willing to amend their required if theory buffeted about usual behaviors to participate in by information is to be the basis the building of an effective inforfor conclusions reached as a result mation base which could be shared of peace research, by social scientists. with the seeming inability of anyone to stem unfortunately, we live in an the information explosion and with unstructured information environfew positive steps having been ment particularly in the social taken recently to improve access, sciences. it seems beyond anyone's the information environment in control, even unesco's. we are which we operate has continued to dependent upon information strucdeteriorate. unfortunately, that tore to organize and facilitate deterioration is still not peraccess to that information environceived by most researchers as ment. yet every existing informaacute, and the literature is still tion structure provides only parconcerned more with technology and tial access for a number of reasons conceptualization than on results. that i will discuss later. there have been some important contributhe impact of a less than a-ietions; the work of alan newcomb; quate information environment is the work of the international poliparticularly felt by those engaged tical science association and the in a multi-disciplinary study such work of my own university center as peace research. in peace for international studies in the research there are no established development of the united states authorities and no established and political science information serformalized schools of thought, vice are examples. if this report therefore, the peace researcher stresses the characteristics of must engage in a widely ranging united states political science information search behavior in information service it is only order to touch all of the relevant because i am better acquainted with information bases. peace researchthe ins and outs of that service, ers are also challenged by being in but, i believe, that even those who a field which is policy related. have participated in these activipeace research is by its very ties would recognize the partial nature bound to intrude upon nature of their contribution. as national sensitivities, ideological more data are collected and dissensitiv ities , and/or academic sencussed the information environment sitivities. if peace researchers are to go beyond the merely polemi3 lassist newsletter, vol. 3, no. 1 (winter 1979) expands and the ability of an the international studies associaindividual researcher to control tion, one of whose thematic seethat environment contracts. partions is a peace studies section; tial access to information has many and those associated with my role debilitating consequences. it as director of the university cenleads to waste of research ter for international studies, a resources; it strengthens academic multi-disciplinary center concerned imperialism and orthodoxy; it reinwith internationalizing the teachforces tendencies toward the reifiing, research and service facilication of authorities; it even ties and programs at the university reinforces global divisions such as of pittsburgh. there exists in the north-south division. it also these cases a symbiotic relationmakes it very difficult for new ship between information systems fields of inquirty to develop, or and substantive research. we have to sustain themselves. i believe found that the more effort we put that the problem of information into developing interactive inforaccess is endemic in the social mation systems the more effort we sciences and particularly acute for will able to devote to the analytipersons engaged in peace research. cal characteristics of research rathern than the bibliographic. in the remarks in this paper are doing so we learned a great dela focused on problems and possibiliabout the conceptual, technical, ties for the integration of docuand organizational problems conmentation and information relevant fronting the development of an for peace and conflict studies into information system and the dissemisocial sciences information systems nation of information from one such and for the use of such systems by center to another, peace researchers. there is an additional emphasis upon the delivthe frustrations stem from a ery of such information to individecade of committee work in such duals in developing nations, altill fated activities as the inforhough i tend to believe that the mation retrieval committee of the problem confronting scholars in any council of social science data field that is multi-disciplinary. archives, and other international social scientific and policy and national committee's whose related is probably no greater in efforts here produced many reports, developing nations than the same workshop statements, guidelines, problem is in most developed but whose contributions to anything nations -in fact, it may even be operative that improves structured easier at least to initiate fields access to information has been very of inquiry in those nations in that limited. these experiences have existing educational institutional reinforced a very strong prejudice structures are less formalized. which you will see running throughout this paper: to concentrate much of what i have suggested only on the thesaurus, the informastems from three positive experition program, the information ences and from one very strong retrieval system, and to believe at frustration stemming from other the same time that this constitutes experiences. the positive sets of a direct and positive contribution experience are those associated to improving information access is with the development of the united to engage in an illusionn. the states political science informaillusion is satisfying and, indeed, tion system, those associated with the development of improved concepmy role as executive director of lassist newsletter, vol. 3, no. 1 (winter 1979) tualizations and models is important and will shape the future, but we can't risk the present for the future. since this type of prejudice has, in the past, been described often as american pragmatism i might as well plead guilty to it. however, just for the record it should no longer be considered as american pragmaticism as american research support agencies such as the national science foundation, the office of education, the ford foundation, etc., have made it abundantly clear that they believe the opeaational dimensions of system design and development to be beneath them. someone has been very successful in presuading them that the "american" concern with getting a system running is not as respectable as the european concern with informatics. i believe we need both; we can deliver the something that works and that meets information needs with very little effort and very little funding. indeed, i think one could easily document that any investment in information systems would soon be recouped. the often discussed formula is that 25% of a research budget goes into bibliographic searching. my guess is that 25% is too low (committee on scientific technical information, 1970, 1976a, 1976b) in a discussion of how to get relevant information into information systems appropriate for peace research we have to begin with some truisms. i will not dwell on the truisms of what constitutes an effective information system. we all know that an information system to be useful and to be used must meet the needs of its users. in the conceptualization and design of an information service the nature of the inforamtion to be made available and the nature of the information behaviors of the users to be served must be analyzed or if that is not possible at least inferred from known information structures and known behaviors with space left in the conceptualization and design for unknowns. therefore, we have to start with the question of what constitutes peach research. david singer suggests there are three major streams to peace research: (1) the pure science school who justify their research on grounds of intellectual curiosity; (2) the applied science school that believe that their research will end human suffering; and (3) the radical critique school. singer suggests that three major substantive issues shape the field: strategic deterrence, social reform, and economic development (singer, 1976). wh a t a r cepts that we approac ing a na research a terms and is used to political national abstracts abstracts , the major e some o mark pe h this q rrow p nd then concepts search sc ience pol it and we get terms . f the ace re uest io rof i le see wh as t the un do c urn e ical ps a view major consearch? if n by buildof peace at types of hat profile ited states nts, interscience ycholog ical of some of agreements, including such items as alliances, contracts, blocs, fronts, state agreements, private agreements, international organi zations; conflict, including such items as aggression, agitation, economic conflict, ideological conflict, social conflict, war ; crisis management, including such items as conflict management. lassist newsletter, vol. 3, no. 1 (winter 1979) diplomacy, strategy, bargaining, peace keeping, arms control; 4. weaponary, including such items as arms, missies, military characteristics; 5. structures, including structural prerequisites for world order, peace keeping forces, international law; 6. peace science, including the search for behavioral and structural predictors of peace, correlates of war, events and even analysis; 7. psychologyical conditions, ranging from aggression to consciousness (international peace research newsletter, 7(3)); and 8. special methodologies, such as fights, games, simulations . and of course with all of these the array of methodologies, concepts and analytical modes that can be used to discuss them. peace r will want very wide b tainly enc social scie ogy as wel some peace concepts s extracts f clinical ps apy will pr additionall be concern impacts and able to mon ized transa esearchers as to be able to ody of 1 iteratu ompassing all nces and social 1. if we add researchers con uch as consc rom the liter yc ho logy and ps obably also be y researchers ed with direc they will wa itor events and ctions as we 11. a group search a re, cerof the psycholto that cern with iousness , ature of yc botherrelevant . wi 11 al so t policy nt to be spec ialthe inform at i is characterized information n researchers is i believe the si ment constitute difficult probl peace research or in getting pe mation into ope systems so that will use such sy cussing sources this paper let traditional fas parameters of th peace research i on environment that by the aggregated eeds of peace therefore immense. ze of this environs the single most em in developing information systems ace research inforrating information peace researchers stems. before disas a major item of us discuss in more hion some of the e development of a nformation system. terminological control. information systems vary in the way in whch they search textual information from full text searching to some system of terminological control in which both indexers and users are acquainted with a codebook of terms. weak information systems rest upon a set of descriptors; strong information systems have a structured thesaurus. with testing and refinement the political science thesaurus will meet all of the terminological requirements of peace research information systems in english. such testing is now underway at the university center for international studies. a revised thesaurus, along with appropriate changes in the computer files, will be available sometime in 1980. the limitation is that this thesaurus is only in english and our conceptulassist newsletter, vol. 3, no. 1 (winter 1979) alizations of language and linguistic structure are not developed enough to allow for direct translation from a thesaurus in one language to another. however, even with the sloppiness that will be built into direct translation we ought to be experimenting with it by running a translated thesaurus against existing information data bases. the fact that most computer readable information services are in english, and even more limiting, drawn from english language sources, limits the possibility of building peace research from many different perspectives and traditions (sartori, e_t a_l . , 1975). levels of information. the design of an information system or the design of a format for machine readable descriptions of textual information always involves tradeoffs. i believe that one area that we cannot allow to be weak and partial is in the number of levels of information that should be included. these include a document number, author, contributors, title, source, abstract, tables, figures and charts, cited authors, subject descriptor (developed from the thesaurus), geographic descriptors (developed from the thesaurus) , and significant proper names. i wish that we had added one more, headings in the article, and i wish that we had included the title of the citation as well as the author. however, since we have been operating on no funds except the subsidy that the university of pittsburgh is providing to cover our deficit, it is perhaps astonishing that we have come as far as we have. we made the strong levels of information decision in the belief that by so doing no one would ever have to go back to these documentary representations of scholarly articles to add information, no matter who the clients might be for the information service. i believe we were right. we were also right in not allowing authors to write their own abstracts. our experience shows that authors tend to abstract the article they wish they had written rather than the article they wrote . info syst ter soci of h d isc info syst the esta term at appr info info wo ul "str info syst rmat em . bee al r ours ussi rmat ems sy blis "st 1 ea oach rmat rmat d ong " rmat ems ion si nee ame a esearch have ng and ion in the stem hed. rong" t st one to the ion pr ion sys use in r ion when t retrieval the computool of , millions been spent desig ning retr ieval hope that could be i used the o describe systems levels of oblem in tems. i the term egards to retr ieval hose syslassist newsletter, vol. 3, no. 1 (winter 1979) terns are easily understood, easily manipulated and fully flexible. its internal nature should be left to those who have to mesh hardware and software requirements. recon is still an excellent model. so are dialog (lockheed) , orbit (sdc) , stairs (brs), wise (a new york consortium), trial (northwestern university) , and pirets (university of pittsburgh) . the design of such systems and changing of the necessary programs is now a matter of routine. the swapping of data and information bases is much more important a process in getting information to be used than is the swapping of programs. we have run uspsd tapes on dialog, on orbit, on trial, and on pirets, and on homegrown programs for small ibm computers, large cdc computers, etc. we are always glad to discuss with any institute interested in swapping information bases, mutual participation in the design and implementatin of a retrieval system for access to the swapped materials. dissemination, what do we know about the format requirements of users of a computer-readable information system? actually, we do not know very much. we feel that users do not like to read computer output. i know i do not. we feel that some scholars want traditionally structured information and that other scholars want information sources that they can manipulate. a strong information system will then produce multiple products, a weak one will produce its product in only one or two formats. in uspsis we produce the following: an annual publication of the entire contents of the yearly file; a set of derivative publications in which certain documents which clusters around a theme are pulled from the file and printed on an annual basis (ethnic studies, strategic studies, intercul tur al studies, russian and east european studies, asian studies, and if we could expand our sources -see below — peace research) ; individualized searches in which clients may write or call and discuss their information needs from which we will generate a custom designed profile and run it against all of our holdings (currently about 600,000 entries); tape leasing in which an institute may lease the tapes receiving a new one each quarter; tape leasing to commercial systems which we do through both lockheed and sdc, and soon may be even microfiche and microfilm. all of these outputs seem easy to generate. what is not easy is to know which will please cl ients . lassist newsletter, vol. 3, no. 1 (winter 1979) 5. sources. here is the weakness of all existing answers to the question of how can we help peace researchers meet their information needs. even the literature of the united states, which is the best covered, is badly covered. we can demonstrate this by asking what sets of sources are potentially of interest of peace researchers and how much of a particular set can they access on a sustained basis. please note that i have raised the question, not in terms of articles in specific journals or specific occasional paper series because i believe it is wasteful and inefficient in terms of financing and in terms of meeting the needs of the diverse clientele which constitute peace research to select journals for inclusion in a peace research system. the computer should do the selection: one person's peace and conflict studies literature may not be another's. in the interest of helping researchers we should be able to tell them what materials are covered by an information system. if article are the level of inclusion in a system the client will never really know whether or not a journal has been fully covered by his or her search. when a decision is made to include a journal in a system all articles in that journal should be included so that a user of the system does not have about full cov so that late abstractor/inde not have to go fill in hole data/informatio can be more ea ped than created . it noted that i a low level conce swapping rathe higher level co as networking believe that s possible now networking requ o rgan i za t ional nological devel be ef f ic ient . require much ping , as i d i in the conclusi paper , can be d a fret erage and r another xer does back and s . al so n files sily swapnetworks should be m using a pt such as r than a ncept such because i wapping is and that ires both and techopment to it doesn' t but swapscussed it on of this one today. a) journals: it is a safe statement that through sociological abstracts, psychological abstracts, united states political science documents, the journal literature published in the u.s. relevant to peace research is handled and handled well. clients who wish to search this literature can do so in published format, machine manipulable format, and in some cases tapes can be purchased or exchanged. through the work of alan newcomb and international political science abstracts some of the international literature can also be searched. the english abstracts of lassist newsletter, vol. 3, no. 1 (winter 1979) international political science abstracts can be searched by computer through the united states political science information service. the most effective organizational arrangement for improving the coverage of journal literature is by the development of national or regional equivalents to the work of any of the above. certainly, this is the stated goal of at least one of the above information services and discussions are now under way with a number of institutions on how to bring this objective to fruition. b) books. it is a fair statement that no existing service adequately handles access to books for peace researchers. in the future when the journal of economic literature goes into a computer based information system some progress will have been made. a major issue with this is how to describe the contents of a book adequately to help the client determine whether a study is relevant to the client's needs. those of us who believe in strong information services argue that books should be analyzed on a chapter basis in as structured a format as journal articles. unfortunately, at this moment no one provides the resources even to do the experimentation necessary to determine whether the cost is worth the effort. experiments on how to describe books for the purpose of inclusion in machine readable information systems are critically important. perhaps the journal of economic li teratur-i' s approach of a critical descriptive review is adequate. one service will be exploring soon the possibility of some subvention from book publishers for inclusion of their products in the service. c) convention papers. given the expansive nature of peace research interests, the lack of definition of the field, and given the wide range of authorities -particularly of informational authorities the life of information in peace research is very short. much of the information exists in ephemera or in fugitive files. three major examples of this type of information are convention papers, peace research conference proceedings, and occasional or working papers. eric is about the only 10 lassist newsletter, vol. 3, no. 1 (winter 1979) d) system that handles convention papers and even there it does not do it particularly well. most associations no longer even list convention papers. the only one that i know that does so on a continuous basis and makes them available for purchase after the convention is the international studies association. yet convention papers are not hard to organize. often they can be grouped around a panel which makes the collection of such papers relatively easy. a hopeful sign is the work being done at the carleton school of international affairs under the leadership of jane baumont. she is collecting the titles of convention papers, computerizing them and doing searches using key world in context (kwic) procedures. it is at least a start. conf ere ings . no. 3 tional ne wslet confere to peac 1 isted . rather ever , interes many o f pa pe r s part o tion en we can mater ia nee proceedin vol xvii, of the internapeace research ter , f i f teen nces important e research are this seems typical . howi t wi 1 1 be ting to see how the conference ever become f the informavironment which access. these is, i bel ieve , are of particular interest to a peace researcher because of the nature of the field as discussed above . e) occasional papers and working papers. at times some of us have argued that publication is a symbolic act. by the time an article is published it may be known throughout the invisible college network by those who have their antenna out regarding the work of the author. a media of information that is important to the peace researcher is working papers of institutions. many institutes have working paper series. john fletcher has shown what can be done in regards to letting the world know about economic working papers. this fall, economic working papers will begin computerization and may be available through commercial information systems. unfortunately, economic working papers contains neither abstracts nor a structured terminological control system but mr . fletcher is building a terminological control system inductively. his efforts deserve support. f) research directory including research in 11 lassist newsletter, vol. 3, no. 1 (winter 1979) technological forecasting and social science information services as ever greater demands are current issue of the lassist newsplaced on various resources, letter while a second questionnaire government, industry, and education will appear in the summer issue. a around the world are attempting to report of the findings will be prebetter understand future technologsented in the fall issue, ical developments so that better and more coherent planning can take the delphi method is designed to place. as with other elements of elicit from a group of experts -technological societies, those peoin this case members of lassist — pie engaged in the delivery of informed opinions regarding the information services must have some future. it does this by first idea of the future in order to meet seeking from the experts suggesdemands placed upon them by the end tions about what developments user. one of the problems with the (problems, events, techniques, delivery of such services is that equipment) will develop. the time we meet both technological and conframe you should consider is the ceptual problems. period from 1980 through 2000. the first questionnaire to be presented the technological problems are is, therefore, essentially a blank bound up in developments of compusheet of paper, ter and communicat ins devices, whiole the conceptual problems using all the imagination and relate to the issue of what data information at your disposal we shall be saved and archived. altwould like you to make some suggeshough the need for forecasting is tions about what events and/or evident, the techniques for proproblems you believe will confront ceeding with such forecasts is not the delivery of information serat all apparent. one of the privices over the next twenty years, mary methods used to forecast and from these open-ended responses assess developments in technology will be developed a set of items is the delphi method. delphi is which will be presented in the secuseful under conditions where hard ond questionnaire -you will then trend-line data are not available be asked to rate each of the items or where typical trend models are in terms of the likelihood that an thought to lack descriptive power event will take place, an estimated for future developments. date by which the event will take place, and the desirability that in order to assess future needs the event take place. from the set and problems concerning the delivof responses to the second quesery of machine readable informationnaire an analysis will be tion, with this issue we are begindeveloped which will present the ning a delphi study of the future findings of the study, developments in information technology and service. the first questionnaire appears in the insert 1 lassist newsletter, vol. 3, no. 1 (winter 1979) we are embarking on this study naire by air mail, in an effort to further involve the members of lassist in developing please note goals and objectives for the future of information services and techin order to prepare the second nology. in this insert space is questionnaire we need as much time provided for your responses. when as possible. it is imperative that you have completed as much of the you complete and return your quesquestionnaire as possible, please tionnaire as quickly as possible, return it to: we will structure the second questionnaire using those responses thomas wm. madron received by may 15, 1979. your academic computing and research help and cooperation will be services greatly appreciated. this is an western kentucky university opportunity to provide all of us bowling green, ky 42101 usa with some information which should be useful in planning at our home if you live outside the united inst i tut ionas as well as for lasstates, please return the questionsist. thanks. with respect to each category noted below, please answer the following question: "what conditions will be facing those of us who deliver information services in the 1980s and 1990s?" provide as many responses as you believe desirable in the form of short, declarative sentences . problems for information services: 1. 4. user demands on information services: insert lassist newsletter, vol. 3, no. 1 (winter 1979) 1, 3. 4. new techniques for delivering information services 1, 3. equipment needs and opportunities: 1. insert 3 lassist newsletter, vol. 3, no. 1 (winter 1979) other events or conditions; 1. thank you for your participation insert 4 lassist newsletter, tol. 3, no. 1 (winter 1979) progress. peace researchers need to know who is doing what in the field and what research is being undertaken. directories of research in progress and biographical directories are important information bases for peace researchers. newsletters such as international peace research contains invaluable information on who is doing what and what institutes are producing. but a more systematic and more easily updated system is required. g) policy documents. aside from the new york times information bank what is there? yet peace researchers need to know about transactions, ublic affairs, pronouncements by individual political actors. a model for this type of policy related information system in scorpio at the u.s. library of congress, but scorpio is not public and until it is it is of little use to peace researchers. h) specific data sets. in vol. xvii. no. 3 of the international peace newsletter it is suggested that peace researchers should pay attention to such critical matters as food production. yet how do we find data that we can manipulate for analytic purposes on food production? what is needed is an information system on data files available for secondary analysis. the characteristics of such a system are discussed by paul e. peters in sigsoc bulletin, vol 6, no. 2-3, in 1974, experimental work on developing such files was done by the university center for international studies as early as 1970. however, despite many international conferences on this subject, there is no service available. 6. swapping and networking. the information environment of the peace researcher is immense and continuously growing probably at an expanding rate. the structure for accessing the environment in a continuous controlled manner are inadequate. what can be done to improve this situation? there is always room for improvement in both the technology and terminology dimensions of an information service, but the major impediment is neither conceptual nor technical, it is organizational. we have had our sights set on the system rather than on facilitating system interchange. if through inter-institutional agreements research institutes would sector out the information environment and take on res12 lassist newsletter, vol. 3, no. 1 (winter 1979) ponsibility for building machine-readable files of a sector with agreed upon standards on delivery time, on levels of information, and on terminological control, and then on conditions for swapping, swapping of files could become routine rather than the unusual. uspsis would be glad, as i am sure many research institutes or services would be glad, to swap files, unesco could be of great assistance by providing seed money and organizational backing for such swaps. this many modal model would come closer to satisfying the needs of peace researchers than any one publication, or any one tape, or any one service. as the swaps became networks an infra-structure for peace researchers would develop, and perhaps even a demonstration model for international collaboration could emerge . references committee on scientific technical information (1970). "director of federally supported information analysis centers." cosati-70-1. pb 189 300. springfield, va: national technical information services. -{1976a) . "progress of the united states government in scientific and technical communications." pb 180 867. springfield, va: national technical information services. -(1976b), "proceedings of the forum on federally supported information analysis centers." november 7-8, 1976. sponsored by cosati panel 6, pb 177 051. springfield, va : national technical information services. fred w. , tower of sartori, giovanni, riggs and teune, henry. babel : on the definition and analysis of concepts in the social sciences . international stud ies association, pittsburgh, 1975. singer, j. david. "an assessment of peace research." interna tional security . 1(1), summer 1976, 118-37. 13 by 11 iassist quarterly winter/spring 2010 by amy west 1 sources for international trade, prices, production, and consumption abstract lengthy time-series of international trade, prices, production and consumption of commodities are becoming increasingly easy to find and use via the internet. major international organizations see openly available data sources as key elements of their service mission. at the same time, international organizations maintain use policies increasingly out of sync with the features that their technological innovations offer. in contrast, even though the united states government offers a wide range of international data with no use limits, user interfaces are often significantly less well developed. in addition to international and u. s. government sources, one entirely privately produced resource, lexisnexis statistical insight, is also included and notable for falling somewhere in between these two poles. this article describes the general data availability and stated use policies for the databases comtrade, oecd ilibrary, unctad commodity price statistics, minerals yearbook, faostat, psd online, and statistical insight with respect to data on commodity trade, prices, production, and consumption keywords:trade; commodities; minerals; agriculture; introduction thanks to the internet, librarians and researchers have the best kind of problem these days: sifting through the vast quantities of increasingly freely available statistical timeseries for just the statistics needed for their research. while this paper will not attempt to provide an exhaustive catalog of every database, indicator, country or year of coverage in each of the sources discussed, it will provide some broad outlines for each database’s trade, prices, production, and consumption data. the databases covered are • united nations (un) comtrade • organisation for economic cooperation and development’s oecd ilibrary • united nations conference on trade and development (unctad) commodity price statistics • united states geological survey minerals yearbook • food and agriculture organization (fao) faostat • united states department of agriculture (usda) production, supply and distribution online (psd online) • lexisnexis statistical insight sources covered have at least five decades of data and one or more of the following indicators: • bilateral commodity trade • international prices • international production • international consumption these and many more links can all be found in the author’s del.icio.us account at http://www.delicious.com/umdatalib. this paper will focus on the parts of these databases that relate to bilateral commodity trade, prices, production and consumption, but each contains time-series for many additional topics. where database providers make such features available, the paper will also cover copyright statements, citation tools, alert services, and visualization options. terminology note: while data and statistics are different things (specifically, data are what one uses to create statistics), in practice, the two terms are routinely used interchangeably. most of the databases in this paper use “data” to describe their content and i will follow the same convention. un comtrade the most comprehensive source of international bilateral commodity trade data is the united nations comtrade database at http://comtrade.un.org/db/default.aspx. 12 iassist quarterly winter/spring 2010 comtrade contains data from 170 reporting countries and covers 1962 to the present as completely as possible. comtrade data processing practices include conversion of commodity values from national currency to us dollars using reporter-supplied exchange rates, conversion of reported quantities to metric units, and conversion from the reporter-supplied classification into all versions of the harmonized system (hs), all versions of the standard international trade classification (sitc), and the broad economic categories (bec) classification. the united nations says of its received data, “the data are permanently stored in the un comtrade database server.” the comtrade “about” page (http://unstats.un.org/unsd/ tradekb/knowledgebase/what-is-un-comtrade) contains the statement “browser warning: un comtrade works best with internet explorer web browser. in other browsers, some pages might not work properly. “ mac users can not use internet explorer unless they have a windows emulator, but they should have no trouble using comtrade with firefox. on the other hand, users of google chrome and safari may find some limited functionality. anyone may view comtrade data and anyone may download up to 50,000 records per user per session. if needed, a person or institution may subscribe for additional services including unlimited downloads. visualization options are very limited. comtrade includes some automatically generated pie charts and a map based explorer. however, at the time of writing, the “explorer with map” remains broken in all platforms/browsers checked. expansive user access to comtrade stands in contrast to the stated policies on re-use and dissemination of comtrade data. there are two slightly different copyright statements plus a separate policy on use and re-dissemination. on the “about” page (http://unstats.un.org/unsd/tradekb/ knowledgebase/what-is-un-comtrade) is this statement: “the unsd holds the copyright to comtrade data. all rights reserved. users cannot disseminate comtrade raw data without the express permission of the unsd.“ however, comtrade provides no information on what “comtrade raw data” would be on this or any other pages. the footer of every comtrade page contains this copyright statement: “no part of this material in either its printed or electronic format may be reproduced or transmitted in any form or by any means, electronic, mechanical, photocopying, recording or otherwise, without the prior permission of the copyright owner.” comtrade has a third description of rights and permissions at (http:// comtrade.un.org/db/help/policyonuuseandredissemin ation_11aug2010.pdf) which states that re-dissemination of comtrade data within an institution does not require permission. so, a student including data or tables in a paper can do so without requesting permission provided the paper only goes to the professor. the comtrade policy on use also states that use of a “few tables or graphs in newspaper articles, journals, other magazines or books” doesn’t require permission. in an email from staff in the un publications board and exhibits committee, the author was told “therefore, whenever someone, after downloading some available data, wishes to reproduce it (say, in an article or as photocopies to be used in a class) or forward it (say, to other users, subscribers, include it in a database, etc), that someone needs a permission.” 2 to recap, the un says that using data in papers within a single institution or using a few tables or graphs is fine to do without seeking permission. actually reproducing the data does require permission and depending on the quantity reproduced, a fee as well. what is dissonant about this use policy is that one of the best things about comtrade is the ease with which a user may access the data for precisely these kinds of uses. because librarians ought to direct their users in the proper use of the resources to which they send their users, the author anticipates that in the future she will recommend that, regardless of whether users download data or not, when they cite their sources they should point to the table on the comtrade website that they generated when downloading data for analysis. in pointing to comtrade itself, the user is not re-disseminating data and therefore does not need to seek permission. because there are so many caveats that come with using data from so many different sources, comtrade takes the unusual step of inserting a “readme” screen before the display of a table. (http://comtrade.un.org/db/help/ ureadmefirst.aspx) much like click-wrap licenses for software purchased online, to get past the screen and to the data, users must check a box indicating that they have read the information provided describing data coverage and limitations before they can see the table that they have just constructed. the readme page also appears when a user saves the url to a particular table. thus, if a table is cited in a paper and another user (a reviewer or professor for example) goes to see the table, then that user will also have to indicate that she has read the form. finally, for users who would like to develop their own interface to the comtrade data, the united nations provides an api. some aspects of the comtrade database are freely available via the api, but others are limited to institutions with a license to the database. details are available at http:// comtrade.un.org/ws/. oecd ilibrary the oecd has two versions of its bilateral commodity trade database. sourceoecd is the older version and is in the process of being phased out. its replacement is the oecd ilibrary, officially launched on 20/9/10. 3 13 iassist quarterly winter/spring 2010 the oecd is a 33-nation member organization established in 1961. as the oecd’s publications serve the needs of member organizations, the oecd ilibrary has a different scope from comtrade. the reporting countries are the member nations, although for the newest members little data may be found so far. oecd ilibrary also reports data from a few non-member entities including the eu-15 extra eu4 and taiwan (labeled as chinese taipei and for which data are hard to find). however, each member reports all of the countries with which it trades, so while the reporters are limited to the oecd membership, the database is still quite large. the overall time period covered is 1960-present. oecd ilibrary classifies its data according to the hs and sitc systems. oecd ilibrary may also include volume and value of trade for price calculation if provided by the reporting countries. while some oecd ilibrary content is freely available, full access is fee-based. the cost can be quite high. however, the oecd ilibrary also includes all of the oecd’s textual materials as well as their statistical databases, so subscribers get a lot of content for the cost. several very useful features in the oecd ilibrary make it even more worthwhile for institutions that can afford to subscribe. while providing citations and links to citation management software is pretty standard now in journal databases, similar features are much less common in statistical databases. oecd ilibrary provides citations to whole datasets, links to the major citation management software and will provide citations down to the table level for tables in textual publications. oecd ilibrary uses digital object identifiers (doi) as part of its classification structure. in doing so, they are adhering to emerging international standards for data citation. oecd ilibrary also provides the markup needed to use browser-based citation tools like zotero. rss feeds that alert subscribers to new data as added are another feature of use to researchers studying long-term trends . the oecd also claims copyright to its materials and has a separate terms and conditions document at http:// www.oecd-ilibrary.org/about/terms for subscribers to the oecd ilibrary. however, in each case acceptable uses are clearly defined and for most academic purposes, the oecd maintains a liberal use policy. they summarize acceptable use this way: “oecd encourages the use of its content (textual, statistical, and multimedia). you can copy, download or print content for your own use, and you can include excerpts from oecd publications, databases and multimedia products in your own documents, presentations, blogs, websites and teaching materials, provided that suitable acknowledgment of oecd as source and copyright owner is given. when available, you should use the "cite as" tool within the oecd ilibrary. when the "cite as" tool is not available, you should cite the title of the material, © oecd, publication year (if available) and page number or url (uniform resource locator) as applicable.” 5 whereas comtrade has the greatest volume of trade data, oecd ilibrary provides the most researcher-friendly interface by integrating citation into the database and providing alert services via rss feeds on a database-bydatabase or topic-by-topic basis. the oecd ilibrary also has a more clearly worded statement on acceptable use of their content than does comtrade. unctad commodity price statistics the unctad commodity price statistics at http://www. unctad.org/templates/page.asp?intitemid=1889&lang=1 contains commodity prices and price indices for the period 1960-present. the focus here is not on comprehensive lists of commodities, but instead on raw materials, especially mineral commodities. prices are in us dollars per ton unless indicated otherwise. the commodity price statistics works in multiple browsers and operating systems. the web version of this database is freely available. the unctad commodity price statistics terms and conditions document generally “grants permission to users to visit the site and to download and copy the information, documents and materials (collectively, “materials”) from the site for the user’s personal, non-commercial use”. 6 for researchers looking for prices on minerals and raw materials, this is an excellent choice. usgs minerals yearbook the usgs minerals yearbook at http://minerals.usgs.gov/ minerals/pubs/myb.html provides price and production data either by country with reports from the mid-90spresent or by mineral with statistics from 1900-present. it is freely available online and, as a united states government publication, is not subject to copyright. the web version of the minerals yearbook remains essentially like its print predecessor with a few format modifications. the minerals yearbook retains the organization of the print with volume one covering minerals, volume two covering domestic prices and production and volume three covering international prices and production. most of the content is provided as pdf files or microsoft excel files. there is no single database to search across nor are there any built-in citation tools, or apis. the usgs does provide an rss feed of new publications on minerals that supports researchers with ongoing needs for data. individual tables may include recommended citations. while the minerals yearbook is not very fancy, it is free, easy to find and use, regularly updated, not subject to usage restrictions of any kind, and backed up with print editions in case the site were to become unavailable. faostat the fao database faostat contains a wealth of agricultural and resource data. it is the primary source for comprehensive (or as near to comprehensive as can be) 14 iassist quarterly winter/spring 2010 international • production data (1961-present) • trade (1961-present )quantity, unit value, value • prices (1966-present) in local currency, standard local currency, us dollars • food supply (1961-present) quantity produced, producer price, area harvested, yield per hectare faostat units are metric and values are expressed in us dollars unless noted otherwise. production, price, and supply data can be converted to hs classification from fao codes, but so far, trade data are only available by fao code. faostat is free to use in its entirety as of july 2010, but for researchers in need of very large batch downloads, a free registration is required. faostat says of its copyright “the data of the faostat database shown on this internet site are copyrighted by the food and agriculture organization of the united nations and are provided for your internal use only. they may not be re-disseminated in any form without written permission of the fao statistics division.” 7 there are no separate terms of use. a slightly more positive expression of rights from the fao is “all rights reserved. fao encourages the reproduction and dissemination of material published on this web site. non-commercial uses will be authorized free of charge, upon request.” 8 faostat does not have an rss feed for the whole database or parts thereof. no citation tools are built in either. indeed, all fao says about citation is “all references to faostat data will have to be mentioned with the proper url and the access date.” 9 unfortunately, faostat does not support persistent urls for individual tables, so presumably the proper url is the url for the overall dataset such as tradestat rather than us imports of shelled almonds for 2006 by importing country. faostat targets data users with needs for large amounts of data, but it also provides simple graphical representations of trend data for users with casual data needs. usda production, supply and distribution online (psd online) the usda psd online database at http://www.fas.usda. gov/psdonline/psdhome.aspx contains trade, prices, supply, demand, and consumption data for agricultural commodities from 1960-present. the united states is the only reporter, but bilateral data between the us and other countries are available where they exist. like nearly all us governmental databases, psd online is free. there are also no usage restrictions since this database is excluded from copyright like the minerals yearbook by virtue of being a us government work. it appears to work in all browsers and platforms. in firefox, there are some display quirks, but they do not appear to affect the database’s function or its return of data. users may create free accounts in the database and thereby save their queries for re-use. users may also download datasets for each commodity for analysis in their preferred software. psd online also offers a data availability lookup at http://www.fas.usda.gov/ psdonline/psdavailability.aspx that, if the user finds it first, can save them the effort of constructing a query only to get no results. lexisnexis statistical insight lexisnexis statistical insight is very different from the other resources mentioned here. it is probably closest in many ways to a traditional journal database. it does contain some tables in excel or *.gif, but its value lies in abstracts of print publications from an exceptionally wide range of sources. even when the data are not directly available in statistical insight, it still helps users to figure out what kinds of data might exist and the most likely sources for that data. lexisnexis claims that statistical insight includes data from 1801-present. statistical insight includes all kinds of statistical publications including those with prices, production, and consumption. at the university of minnesota, statistical insight helps to fill a gap in the libraries collection of market research. market research is very, very expensive and the university of minnesota is unfortunately not able to purchase market research reports directly. however, statistical insight can identify sources of data from trade associations that can cover similar territory to market research. thus, it helps us serve researchers that otherwise would have to look elsewhere for such data. statistical insight is less expensive than market research reports, but it is still very expensive, especially if an institution licenses access to additional modules that incorporate more data into the database. no part of it is freely available. while there are limits on claims to copyright of factual material in us law, license agreements can overcome those limits. lexisnexis licenses are typically fairly restrictive. on the other hand, to the extent that statistical insight serves to identify data sources rather than as a source of data itself, rights issues seem a little less important in this case. conclusion not only are more data sources appearing online, but they are doing so with the explicit goal of making more data available to the greatest number of users. data sources that may have launched with short time series find that users always want more and the data source producers strive to meet those desires as openly as possible. many sources provide researcher-friendly features such as update alerts via rss feed and built-in citation tools. at the same time, data source rights statements do seem to be out of sync with the spirit of expanded open access. in an environment in which data source producers enable features like mass 15 iassist quarterly winter/spring 2010 downloads, the expectation that permission must be acquired on a case-by-case basis in order to use those downloads seems odd. however, these are minor quibbles. the increased quantity of data, increased ease of access and improved interfaces far outweigh any drawbacks identified in this paper. notes: 1 amy west, data services librarian. 10 wilson library, university of minnesota, 309 19th avenue south, minneapolis, mn 55455-0414. (612) 625-6368. westx045@umn.edu. 2 morteo, r., 2010. question about the copyright statement for un comtrade. 3 schoenfeld, a., 2010. official launch of oecd ilibrary. 4 glossary:eu enlargements statistics explained. european commission eurostat. available at: http:// epp.eurostat.ec.europa.eu/statistics_explained/index.php/ glossary:eu-15 [accessed october 1, 2010]. 5 oecd ilibrary: copyright and permissions. available at: http://www.oecd-ilibrary.org/about/copyright [accessed september 30, 2010]. 6 unctad.org >> terms & conditions. available at: http://www.unctad.org/templates/page. asp?intitemid=2146&lang=1 [accessed september 30, 2010]. 7 faostat. available at: http://faostat.fao.org/site/567/ default.aspx#ancor [accessed september 30, 2010]. 8 fao: copyright. available at: http://www.fao.org/corp/ copyright/en/ [accessed september 30, 2010]. 9 faostat. available at: http://faostat.fao.org/site/567/ default.aspx#ancor [accessed september 30, 2010]. establishing an australian social science data archive: progress and plans roger 6. jones social science data archives australian national university progress up to 1974, there appears to have been little interest in australia in accessing secondary data. the department of political science in the research school of social sciences at the australian national university held a category c membership of icpr from 1965 but access was largely limited to members of that department and certainly only to members of the university. the only data that was generally available at that time was the 1966 census data distributed by the australian bureau of statistics. during 1974, four events occurred which provided the stimulus for a wider debate on the need for an australian data archive: i) a second department, the department of political science at the university of melbourne, became a member of icpr; ii) icpr decided to reclassify category c schools in australia (and elsewhere) to category b institutions for 1975-76 with a consequent increase in membership charges from $2000 to $3500 per year and further increases to follow. almost simultaneously they proposed a new arrangement under which any number of australian institutions could form a joint organization for australia at a total annual subscription of $4000, subject to one of them acting as a clearing-house for the whole group; iii) the survey research centre was established at the australian national university; iv) don debats, senior lecturer in american studies and politics at flinders university, presented a paper to the academy of social sciences recommending the establishment of an australian data archive. without any one of these factors, it seems doubtful whether any progress would have been made for some time towards establishing an archive. debats had approached the national library two years earlier with a suggestion that they take out a national membership of icpr but his was a lone voice and it was felt that it would be hard to justify offering a new service, and one which would be a totally new departure in the type of material offered, when there was apparently no demand for it. at that time, the and was more concerned that any national membership should not adversely affect their own arrangements than with encouraging wider access. the proposed increase in charges and the presence of at least one other institution to share these charges and maintain them at the previously acceptable level provided the impetus needed for a national membership to be considered. the newly established survey research centre had as one of its objectives the collection of information on survey data that could be made available for secondary analysis and was seen as the logical location for the national clearing-house. as it was, following some preliminary investigation of possible alternatives and canvassing of the level of interest, a meeting was arranged for 16 february 1976 at the anu and representatives of thirteen institutions attended. eleven of these expressed an interest in joining an association of research and teaching institutions formed to take up a national membership of icpsr (as icpr was now called). this was taken out in may 1976 under the name of the australian consortium for social and political research incorporated (acspri). secondary objectives of this organization were i) to collect and disseminate information relating to machinereadable social science data; and ii) to investigate the desirability and feasibility of establishing an archive of australian social science data in australia or elsewhere and, if it is found desirable and feasible, to facilitate the establishment of such an archive. over the last six years, acspri has grown from the two previous icpsr members to nineteen member institutions at present, including twelve of the nineteen universities in australia. each member pays an initial joining fee of $150 and all share equally in the costs of icpsr membership. thirteen acspri nominees have attended icpsr summer training programs. data exchange agreements have been established with the roper center and the ssrc survey archive and agreement to redistribute data acquired from the data and program library service has been obtained. a newsletter is produced twice yearly and distributed through representatives in member institutions and to other interested bodies. over the years, the number of orders for secondary data has remained stable at the level of about 9 a year, although the number of data sets distributed grew from 21 in 1976 to a peak of 85 in 1979 before dropping back to only 27 in 1980. in general, the pattern has been that a member will place a large order for data sets soon after joining, and then orders will be for one or two data sets only. table 1. acspri membership and level of use no. of members no. of no. of year at 31 dec orders data sets 1976 9 7 21 1977 10 8 46 1978 13 10 69 1979 16 7 85 1980 19 9 27 although acspri has been successful in providing australian researchers with access to overseas data, it has been far less successful on its home ground. the anu survey research centre was relied on to undertake any data location and acquisition procedures but found that this was generally impossible due to its other commitments. as a result very few australian data sets have been acquired to date. excluding australian census data, only three australian data collections are available through icpsr and only a further 24 data collections are available from acspri. if this situation had continued for much longer, i believe that membership of acspri would have started to decline, probably quite rapidly. already one of the founding members has dropped out because of lack of interest within the institution. for the great majority of academics, researchers or teachers, local data relating to local characteristics and issues is surely preferable to overseas data. in order to flourish, an archive must substitute for or add to the researchers' data collection activities, as well as provide new opportunities for data analysis, and these possibilities are more obvious with local data. this situation now has a good chance of being rescued following the recent decision of the anu to replace its survey research centre with the social science data archives. the archives will have a staff of six initially and should be fully operational early next year. in preparation for this, some preliminary investigations have been undertaken. in particular, sources of information on survey work in australia have been examined and procedures to follow in acquiring, documenting, advertising and distributing data sets have been considered. the results of these deliberations and some of the questions they raised are presented below. locating survey data through published sources 1. government collections "the australian bureau of statistics is the official statistical organization for the federal and state governments. its main function is to collect statistical information from a wide variety of social and economic areas and to compile statistics and disseminate them to interested users both within the government and the community in general. the abs publishes currently almost 1900 statistical publications either monthly, quarterly, half-yearly, annually or irregularly under approximately 700 different titles." (abs catalogue of publications). the abs is the major data collection agency in australia. however to date, the bureau has taken a very strict line on confidentiality of respondents and has been unwilling to release data in machine readable form in general and certainly not individual record data, de-identified of course. data from the australian censuses of 1966, 1971 and 1976, with 1981 in a few years time, has been made available on magnetic tape aggregated at least to census collector's district level (an average size of 200 dwellings). from the 1976 census in particular. matrix tapes containing counts of individuals or dwellings in cells of multidimensional tables were also made available, although in a format which required a considerable programming effort by the user to read and produce meaningful output. these data tapes are already held by the archives. however, such important studies as the 1974 general social survey, the 1977-78 australian health survey, the 1974-75 and 1975-76 household expenditure surveys, the monthly labour force surveys and many others are inaccessible and likely to remain so. attitudes are changing however and there is some possibility of a sample of individual records from the 1981 census being available for public use. in addition, the abs will under certain conditions and when resources allow, conduct some analyses of individual record data on behalf of researchers. apart from the abs, there are many other government agencies at the federal and state level who undertake data collection activities, and these agencies are generally more willing to make the data available to academic researchers. until recently, information on these data collections was not widely available in any systematic form. however, statistical co-ordination bodies have recently been established by the coimonwealth and state governments and each of these, with the exception of tasmania, has compiled a register of statistical collections undertaken by the various departments and authorities of their respective governments. entries are generally organized under the abs program code or department, and include the title, frequency, time period covered, availability and a contact officer. at the present time each of these bodies uses a different data collection instrument and publishes its information in a different form, but there is some discussion of a unified approach for the future, provided that cuts in staff and available funds allow the continuation of these projects. 2. opinion polls in the period 1941-1971, only one organization roy morgan research centre pty. ltd. conducted regular surveys of public opinion in australia on an interstate basis. two further polling organizations australian nationwide opinion polls (anop) and irving saulwick and associates entered the field during 1971, and mcnair anderson associates pty. ltd. began regular polling in 1973. the recent publication "australian opinion polls 1941-1977" compiled by the university of sydney's sample survey centre provides a subject classification and keywords index to the questions included in the polls conducted by these four organizations up to 1977. data from about half of the 190 surveys conducted by the roy morgan research centre before 1968 are deposited with the roper center and can be made available to australian researchers through a data exchange agreement between roper and acspri. irving saulwick and associates' "age poll" is conducted in association with the political science department at the university of melbourne and permission has been given for these data to be made generally available two years after the completion of fieldwork. negotiations are currently underway with the other three polling organizations to try to establish similar agreements. 3. academic collections in a large and sparsely populated country like australia it is very expensive to build and maintain a national fieldforce of interviewers for use in ad hoc surveys. as a result, any national surveys and the great majority of large regional surveys requiring personal interviews are contracted out to commercial market research agencies for the fieldwork. the only alternatives for large scale survey work are mail self-completion or other self-completion approaches such as surveys of school children conducted under supervision in the classroom. the vast majority of survey work conducted from the academic sector is however based on small samples from small geographic areas. information on the data collection activities undertaken by the academic sector is scattered through a whole range of publications such as annual reports of departments and institutions, reports of the granting bodies who provide funding for much of this research and the journals in which the results of the research appear. the need to provide some form of central register to these activities has been recognized in recent years and some progress has been made in this direction. in 1975 the social welfare comnission produced the first edition of the social welfare research bulletin, which sought to provide a concise listing of social welfare research throughout australia. subsequently, the department of social security took over production of this bulletin and published updated versions in 1977 and 1981. unfortunately, the latest edition is to be the last. a number of other government departments provide bibliographic services on the areas of their particular interest. for example, the department of education maintains a directory of researchers and research in education; the institute of criminology scans publications for australian or australian-related criminological information and aims to collect copies of all publications relating to australian criminology; the department of employment and youth affairs library compiles quarterly bibliographies on a number of topics. however, the entries in these sources are generally limited to author, title and publication, and are thus rarely useful as information sources for the location of machine-readable data files. the survey research centre undertook two projects in an effort to provide more information on academic survey activity. the publication "australian social surveys: journal extracts 1974-78" is based on a search of thirty australian social science journals published in 197478 for articles reporting the use of survey data. approximately 600 entries are organized under subject headings and include author, title, journal reference and, where available from the article, the geographical coverage, date, population, and sample of the survey. the second project, the "inventory of australian surveys", was designed to provide more detailed information on survey work and used a mail questionnaire approach. heads of social science departments in universities and colleges of advanced education were requested to give names and addresses of staff and postgraduate researchers who had conducted surveys from that department since 1970. individual researchers were then contacted by mail and requested to give a detailed description of their work on an inventory questionnaire. details of some 700 surveys are currently held on a computer file. a comparison of the survey references attained in these two projects showed that both approaches suffer from undercoverage. using details of the publications provided in the inventory responses, a brief analysis of written items resulting from these surveys was carried out. based on 617 entries, it was found that about one-third (210) of the surveys had not yet been reported at all, while about one-quarter (145) had resulted in journal articles. of this latter group, at least 107 had published in australian journals although only 69 were covered by the thirty journals selected for the journal extracts. to have located all of these references from a journal search would have required a doubling of the australian journals covered, and inclusion of some overseas journals. table 2 provides details of the types of written reports used. table 2. written reports of surveys included in the inventory. no. of surveys no written items reported 210 journal articles australian journal 107 overseas journals only 19 others only journals not checked, could be either _[9 145 books and monographs 63 academic departments or institutional reports 83 government and other reports 72 published conference proceedings 24 unpublished conference and seminar papers 27 theses 95 n.b. each type of written item reported counted for each survey. on the grounds that there is clearly some time lag between the conduct of the survey and the appearance of a written report, an examination of written output by date of completion of fieldwork was made. surveys resulting in theses, and those in which the dates of fieldwork or type of output was not specified were excluded. as expected, a higher proportion of surveys conducted before 1975 resulted in journal articles, but this was still only 41 percent of all these surveys, and 30 per cent appear not to be written up at least 4 years later (table 3). we have not as yet made any qualitative judgements about the merit of these surveys and it may of course be the case that it would not pay the archive to be too concerned about such work. table 3. written output by date of completion of fieldwork. per cent of surveys completion written up in written up in not of fieldwork journal article other form written up n 29 32 29 16 before 1975 41 1975-76 28 1977-78 21 after 1978 6 30 132 33 130 50 184 78 32 on the other hand, a journal search has some advantages in terms of coverage over the survey approach, due largely to the problems of nonresponse. from a sample of 103 survey reports in the journal extracts we found 9 with no address given and 13 in departments which had not been surveyed. of the 81 remaining, 50 were not reported in the returns from departments, although 14 of these were conducted by researchers who were included in the inventory for different studies. of the 31 studies reported on the department returns, 22 summaries were returned by principal investigators. plans while information on the data collected by government bodies, market researchers, academic researchers and other social science research bodies has improved considerably in australia over recent years, there is a need to co-ordinate these activities and, if possible, establish a uniform approach. the concept of the data clearing house for the social sciences in canada is i believe appropriate for australia, although canada has the advantage of a well established network of data archives. the data clearing house can thus concentrate all its resources on the provision of information services, for which (in 1975) it employed a full-time staff of six professionals, engaged in the developmental and service activities of the program. the broad objectives of the canadian data clearing house are: "1. the preparation of an index of quantitative social science data holdings that exist in machine-readable form and are to be found in canadian universities, as well as in non-profit research agencies and other bodies conducting social science research; 2. the collection from federal and provincial government departments of a continuing description of their holdings and the performance of a liaison role between individual scholars and government departments; 3. the provision of information in response to individual inquiries, referring the inquirer to the source but not attempting to provide the inquirer with the actual data; and 4. the provision of technical information necessary for the more effective use of the data." bulletin, data clearing house for the social sciences, ottawa, nov. 1975. while our objectives have yet to be formulated and agreed, the provision of information about available data is surely necessary for determining a sensible acquisition policy. it would build on the work described above, although there is clearly a need to modify the information collection procedures used in our previous inventory work. a 10 balance must also be found between resources allocated to this activity and resources allocated to data acquisition, processing and dissemination, since we, unlike the data clearing house, will be attempting to provide the inquirer with the actual data . with this modification, the objectives stated above provide the basis for our planning at this time. inventory plans in recent years, there has been a strong movement among data archivists towards standardized documentation and increased bibliographic control of machine-readable data files (mrdf) with the hope that, ultimately, international union listings of available data may be produced. with this in mind, we felt that a new information system should be compatible with overseas developments where practicable. although we have produced our own system for two bibliographies of australian surveys, it is relatively unsophisticated and inexpensive to abandon at this stage. for all practical purposes, we are able to start from scratch. in looking for a suitable description scheme, we required a form which could provide output in the form of a bibliographic citation; a title page; a full description of the study methodology, and content, and associated publications for inclusion in the codebook; and a more compact description for inclusion in a published inventory or catalogue of data holdings. appropriate indices would also need to be generated by machine from the entry. the study description scheme developed at the danish data archives appears, with some reservations, to satisfy these needs. the study description form is essentially a more detailed version of the questionnaire used in our previous inventory work and consequently we are familiar with its style. the questionnaire used there was designed as an instrument which would be completed by the researcher, and returned to us for almost direct processing. in theory, our intervention would be minimal; in practice, it was not. to some extent it may have been due to faulty design, but returns required a significant amount of editing to give consistency, resulting in an untidy copy being sent for processing and thus more editing on the computer. it is therefore anticipated that the description for each study will be completed in-house, and will be based on published reports and other descriptive materials requested from the investigator. based on a very limited trial with three studies, we found only a few problems in completing the sd form in this way, although it is not entirely suitable for our purposes. some sections in part 2, analysis conditions, and part 3, reanalysis conditions, will be omitted, and part 5, variables included, will be compiled as listings of background variables and main variables/topics rather than use the categorized responses provided. the main reason for the latter is that we do not plan to implement a subject classification scheme immediately (preferring to wait for some recommended standard) and will use the main variables/topics as the basis for a keyboard index. as recommended 11 by users of the sd scheme, section 101 will be used to include the necessary elements of a bibliographic citation when these are not already included elsewhere, although an additional section has been added for details of the producer of the mrdf to provide a producer statement. at this stage therefore it seems likely that the sd scheme will be adopted by the australian ssda. there is however one reservation in our minds about adopting this scheme. at present, use of, and interest in using, the sd scheme predominates in the european archives, with only one archive on the north american continent, the leisure studies data bank in waterloo, using it. clearly a standardized system has to be widely adopted to be a standard. given the reported interest in establishing such a standard, we wonder whether the sd scheme is being generally considered outside europe, particularly in the united states; and if not, why not? comments from conference participants on this topic will be very much appreciated. having chosen what we consider to be a suitable study description format, we are still faced with the problem of locating studies for inclusion in the inventory. as indicated by our previous experiences described above, providing reasonable coverage of the academic and other research agencies conducting social science research may be difficult. again looking to the data clearing house model, a national network of designated correspondents and technical co-ordinators may be the answer and will certainly be tried. the ssrc survey archive also has a network of archive representatives covering university and polytechnic social science departments to publicize the archive's services and acquisitions and to simplify request procedures. the basis for such a network is already established by the nineteen acspri representatives, one for each member institution, and efforts will be made to expand and develop this network. compilation of the inventory is seen to require three stages of information collection. firstly, a record will be kept of current research and completed research comprising little more than names and addresses of principal investigators to be contacted, acquired through the network of representatives, reports of grant agencies and other information sources described earlier. essentially, a mailing system for recording details of correspondence between the archive and investigators. secondly, information on completed studies will be compiled from available publications and documentation supplied by the researcher. this will form the basic material, for deciding whether or not the data should be acquired. prerequisites for inclusion of information at this stage is that the data is extant, in machine-readable form, and that the researcher is willing to make the data available to secondary users, perhaps conditionally, at some future date. complete descriptions of data sets will only be made for studies acquired by the archives and available from the archives for secondary users. 12 acquisition plans acquisition policy will generally be determined by reference to the users advisory committee which is being established for the archives. members of the connittee will be drawn largely from the social science departments of the university which include demography, sociology, political science, economics, economic history, law, history, urban research and statistics. in addition, at least one representative of acspri will be on the committee. materials gathered in the course of compiling the inventory will be presented to the committee at regular, probably quarterly, meetings for a decision on the priority to be given to acquiring the data. highest priority will be given to acquisition in response to specific requests, which may be for a specified data set or sets, or for data relating to a specific topic. if the data are not already held by the archive this will clearly involve some delay, but every effort will be made to minimize this. in the longer term, as the holdings of the archive increase, the frequency of such requests should diminish. the third basis for data acquisition relies on the attitudes of research funding agencies in australia towards data archiving. the major funding bodies support a great deal of the primary data collection activity, particularly that of academic researchers, and should be supportive of an activity which will encourage wider use of these resources. grant applications for additional funds to support the salaries and activities of additional staff will be made, which, if successful, will allow the archive to develop more quickly. the australian research grants committte, the major source of academic research funding, agreed three years ago to include in its advice to applicants a request that social science data arising from funding projects be deposited with acspri, but this has achieved little to date. many overseas bodies make the deposit of such data a condition of grant, but this has so far been resisted by the argc. the department of health has this year provided funding to support the establishment of an archive of survey data on drug use in australia and this project is underway. there are i believe major advantages in focusing data acquisition on specific substantive areas where funding is largely centered on a single agency. the problems of locating suitable data can be overcome through reference to the agency's records, and the agency's involvement may act as an inducement to researchers to deposit their data. a substantial collection of related data provides greater opportunity for secondary analysis, and the agency supporting the creation of an archive will surely want to encourage use of the resource which in turn would encourage further support of the archive from the agency. the archives' users advisory committee will also decide the level of data cleaning to be carried out on data acquired. on receipt of the data by the archive, a minimum level of range checking will be done, and where necessary, multi -punch data converted to single-punch. more detailed checking of the data, error corrections, and creation of a codebook 13 by the archive will only be undertaken on studies thought to warrant the effort and expense. on a general point of inter-archival co-operation, it would surely benefit all archives, and new archives in particular, to have information readily available on the types of data set most often requested. the ssrc survey archive provided us with a list of their 25 most heavily demanded data sets which they concluded "demonstrates that national and cross-national rather than local surveys, and longitudinal panel and time series rather than one-off surveys, attract the heaviest use." i feel sure we would all like to know whether this is a general conclusion or one which is perhaps a result of the particular holdings of the archive at the time. british election studies and family expenditure surveys form a significant part of their list, but does this reflect the substantive topics of interest or the quality of the survey work or some other factor? many established archives will surely have conducted user surveys and it is important that the result of these surveys be widely available to all archives. dissemination plans to date, formal advertising of acspri services has been done through distribution of the acspri newsletter. editions of the newsletter are produced in march and september and distributed by the acspri representative largely within their own member institutions. my intention in establishing the newsletter was to carry reports of research and teaching applications of seconc iry data from contributors, but unfortunately no such contributions have been received over the two years of publication. icpsr provides acspri with seven copies of codebooks for all class 1 data sets and these are distributed to codebook centers located around australia, one to each state. each acspri representative receives a copy of the icpsr guide to resources and services and information mailings, and researchers wishing to consult codebooks can borrow them from the nearest codebook center. of course this places researchers at any but the seven institutions with a codebook collection at some disadvantage, but the cost of establishing more of these centers would be considerable. seven points of access to the codebooks is nevertheless clearly preferable to only one. with the establishment of the social science data archives, the primary task will be to provide information on and access to australian data as opposed to data from overseas archives, and to broaden the interest in secondary use of this data. as reported above, attempts will be made to extend the network of representatives down to departmental level as opposed to the current institutional level, and to include more institutions in the network. the principal output from the archive will be derived from the study summaries compiled for the inventory, since it will contain details 14 of many more studies than the archive has in its holdings. for studies which have not been acquired, entries will exclude specific details of the principal investigator to avoid the possibility of unsolicited direct approaches. copies of the inventory will be distributed free of charge to department representatives and be made available to libraries and individual researchers on a subscription basis. for studies held by the archive, documentation will be distributed to acspri member institutions free of charge, but otherwise sold at cost. the newsletter will continue as the main publicity medium, being distributed free through the local representatives. data requests will be charged on a fee-for-service basis. summary there are a number of alternate ways to establish and develop a data archive and we are faced with choosing one of them. essentially i see a data archive as a consumer-oriented marketing activity with the academic social scientist as the primary consumer, the archivist as the marketing manager and data sets as the primary product. the product is not manufactured by the archive but is picked up second-hand from other sources. the archivist has the job of locating suitable products and deciding which to acquire, and whether or not it is worth cleaning up what is acquired before making it available to the consumer. the problems facing our marketing manager are: what data sets to acquire and in what quantities? where to acquire the data sets? which data sets should be cleaned? what promotion activities should be undertaken? with the object of maximizing the consumer awareness and use of the product subject to the constraints of the limited resources available. the marketing manager realizes of course the need for information on which to base these decisions and, being the manager, delegates responsibility to his market researcher. she (in this case) carries out a literature search and, since this is a new product on the australian market, contacts similar marketing operations overseas requesting relevant information. unfortunately, neither source proves very fruitful. the marketing manager is thus placed in something of a dilemma, and decides to take a cautious attitude. there seems little point in filling the warehouse with materials which may never be sold this would simply be doing something for the sake of it. on the other hand, it may be that by filling a warehouse with goods, chosen because they are readily available, and having a good advertising campaign, enough interest could be generated to clear a lot of it even if it was junk for the most part. on balance though, he feels that the consumer market he wants to attract is fairly discerning and that, although they may initially be attracted to the warehouse, their disappointment with the available product will discourage any future interest. 15 taking this view, the manager decides that first priority should be given to establishing a good network of contacts among the producers, creating an information source on the availability of goods of interest. the producers themselves are of course interested in the activities of fellow producers and it is felt that their co-operation would be gained by offering them the results of the information collection in exchange for their involvement, in the way that estate agents pool information on houses for sale in multi-list schemes. the producers here are also the most likely consumers and the information system will both assist them in planning any new product and encourage their interest in the products of others. acquisition during this initial phase will not be substantial, being concentrated on satisfying customer orders, which are also unlikely to be substantial, and pieces of particular merit selected by a board of expert advisors. these special pieces will be used as the center-piece in promotion activities, and seminars and workshops will be devised around them. in the longer term, obtaining input to the information system should become less demanding of the archive's staff allowing redeployment of resources to promotion, cleaning, new acquisitions and distribution activities. with what is essentially a new product on the australian academic market, promotion must be given high priority in order to attract new customers and to keep old customers up to date with new products. information gained from the network of producers and consumers and orders placed during the initial ph se will provide a guide to customer requirements, allowing effective p.anning of and control over future acquisition and cleaning activities. labor statistics for sale on tape ntis, the national technical information service, has available on magnetic tape statistics from the bureau of labor statistics (bls) of the u.s. department of labor, the labstat database includes: 1) manpower information such as labor force characteristics, employment hours and earnings, nationally and by smsa, unemployment data by smsa and labor turnover, 2) the consumer price index , producer price index , and export & import price indexes , and 3) imports statistics, value by industry. each series is updated monthly and is available as a demand item or by subscription. for pricing and ordering information contact: stuart weisman product manager (703) 487-4807 16 1/14 chigwada, j & chiware, e (2024) future models and architecture of data repositories in african universities, iassist quarterly 48(3), pp. 1-14. doi: https://doi.org/10.29173/iq1099 the creative commons-attribution-noncommercial license 4.0 international applies to all works published by iassist quarterly. authors will retain copyright of the work and full publishing rights. future models and architecture of data repositories in african universities josiline chigwada1 and elisha chiware2 abstract research data repositories as part of research infrastructures are being developed and are important tools and components that help to store, preserve, and allow for the re-use of data. as the technologies, networks, and systems that the data repositories are built upon are advancing, this study explores the future models and architectures that african universities can follow to have reliable and sustainable systems for the preservation of research data. a scoping review was done to focus on the future shape of data repositories based on past experiences of the last 10 years of research institutions in establishing data repositories. this study was done to gauge the communities’ responses to the architecture of existing platforms to prepare other institutions planning to establish digital research data repositories. articles were retrieved from scopus, web of science, and dimensions databases using relevant keywords. the content analysis approach was used to establish the requirements for establishing digital research data repositories to develop a framework that can be utilised by other research institutions to develop their repositories. the framework would be handy in providing a roadmap for research institutions that want to establish research data services in africa enhancing the future of research infrastructure in african universities. keywords research data management, research data services, research data repositories, data repository models, data repository architecture. introduction the pace of research data repositories’ development in african universities has been slow compared to other continents, especially those in the global north (patterton et al., 2018). various factors can contribute to this slow progress including lack of financial resources; inadequate research infrastructures, lack of open science and research data management policies and frameworks, unstable electricity grids, and poor internet connectivity as well as limited skills (chiware and mathe, 2015; chiware and becker, 2018). in the last two decades, african universities and other knowledge production centers have developed and implemented digital repositories to showcase their research outputs including collections of electronic theses and dissertations (etds), research articles, conference proceedings, technical reports, and book chapters. however, most of these repositories, built on platforms like dspace, cannot host research datasets. over the last decade, several studies have emerged on how african institutions can develop research data management services. the majority of the writings are based on reviewing existing research data management practices and thereafter proposing frameworks on how new services can be developed and implemented. several authors however recognise the existing challenges and have often questioned these common narratives that seem to ignore the reality of resource constraints in african research institutions. for instance, abebe et al. (2021) argue that these narratives often overlook power imbalances and there is a need for solutions that are grounded in the african context. https://doi.org/10.29173/iq1099 https://creativecommons.org/licenses/by-nc/4.0/ 2/14 chigwada, j & chiware, e (2024) future models and architecture of data repositories in african universities, iassist quarterly 48(3), pp. 1-14. doi: https://doi.org/10.29173/iq1099 data repositories infrastructure in africa the registry of research data repositories (re3data.org) currently lists disciplinary-specific and general research data repositories in several african countries. the majority of the listed repositories are based in south african institutions, followed by six in kenya, three in burkina faso, two in ghana and benin, and single repositories in cameroon, egypt, ethiopia, ivory coast, malawi, namibia, niger, senegal, sudan and tunisia. there are no other recorded research data repositories in any of the other african countries. the types of data repositories found in these african countries are at two levels: disciplinary (subject) specific and general repositories which are mostly found in south african university libraries. disciplinary-specific data repositories include datafirst, a research data service dedicated to providing open-access data from south africa and other african countries. it also promotes high-quality research through the provision of essential open research data infrastructure for discovering and accessing data and skills development. another existing project is the h3abionet (https://www.h3abionet.org/) which was established to develop bioinformatics capacity in africa and specifically to enable genomics data analysis by h3africa researchers across the continent. h3abionet is developing human capacity through training and support for data analysis, and facilitating access to informatics infrastructure by developing or providing access to pipelines and tools for human, microbiome, and pathogen genomic data analysis. the other discipline-specific continental data repositories include; africa rice dataverse (https://www.re3data.org/repository/r3d100011251), afdb statistical data portal (https://www.ruforum.org/directory/afdb-statistical-data-portal), roceeh out of africa database (https://www.re3data.org/repository/r3d100013419), square kilometre array (ska) telescope (https://www.sarao.ac.za/about/the-project/), africa health research institute data repository (https://data.ahri.org/index.php/home), and the west african vegetation (http://westafricanvegetation.senckenberg.de/menu/home.aspx). these repositories seek to enhance accessibility and research data utilisation across disciplines. there are eighteen south african institutions, mainly university libraries and research councils’ facilities that currently host data repositories of a general nature. the majority of these platforms run on proprietary platforms, especially on figshare which was acquired through a national consortium arrangement to enable more uptake among interested institutions. as these data repositories grow there are increasing calls for their integration with existing platforms that host other research outputs (especially those on dspace) (mehnert et al., 2019). there are also calls for institutions to go through the processes of certification to ensure that the repositories are internationally recognised as holding trusted data deposits. trust in repositories will ensure maximisation of research outputs as well as facilitate collaboration and sharing of existing research outputs (mehnert et al., 2019). trusted repositories usually have demonstrated high levels of trustworthiness, through certifications like the core trust seal and membership of the international science council’s world data system (wds) (lin et al., 2020). they will have implemented robust policies, procedures, and technology infrastructures to ensure data quality, security, and preservation. another important aspect relates to the culture of data sharing, which is the widespread adoption of open data practices, where researchers and organizations share data freely and willingly. the culture of sharing also encourages collaboration, reuse, and building upon existing research as well as fostering community norms and values around data sharing. as research data management continues to grow and with support from funders, journal publishers, and the need to respond to international and national calls for coordinated data sharing mechanisms among and beyond research teams and enable more transparency as well as provide impetus to more accelerated scientific discoveries in developing countries and in africa in particular, it is important to find models and solutions for data infrastructure architectures that can be adopted at minimal cost. the goal of this study was to explore how new models can be developed to encourage the uptake of data repositories in african institutions through minimal costs, as well as, ensuring sustainability. https://doi.org/10.29173/iq1099 https://www.re3data.org/ https://www.h3abionet.org/ https://www.re3data.org/repository/r3d100011251 https://www.ruforum.org/directory/afdb-statistical-data-portal https://www.re3data.org/repository/r3d100013419 https://www.sarao.ac.za/about/the-project/ file:///c:/users/oschwart/stokes/iq/revisions/(https:/data.ahri.org/index.php/home http://westafricanvegetation.senckenberg.de/menu/home.aspx 3/14 chigwada, j & chiware, e (2024) future models and architecture of data repositories in african universities, iassist quarterly 48(3), pp. 1-14. doi: https://doi.org/10.29173/iq1099 scope the pace of research data repositories development on the african continent has been limited due to several challenges, including, limited funding, lack of policy framework, skills shortage and limited development within research infrastructures. some of the well-resourced african countries have started using proprietary platforms to manage research data. some of the existing institutional repositories like dspace have no or very limited data management functionalities, thereby limiting the ability of most african universities to use open-source platforms to store and manage research data. for african universities to fully participate in the global open science agenda their scholarly outputs, including data, must be properly managed through data repositories that can be easily accessed. the aim of this paper was therefore to review past experiences and frame future models and define the architecture of data repositories that are more suitable for african university universities. objectives the objectives of this study were: 1. to identify the requirements for establishing advanced data repositories. 2. to define successes and challenges in establishing and managing data repositories. 3. to develop a framework to be utilised when developing research data repository infrastructures. methodology a scoping review was selected due to the complex nature of the topic and the wide range of information sources that might be available for the study since the issues of establishing research data repositories are topical. the review was guided by levac’s (levac et al., 2010) scoping review methodology, which is an improvement of arksey and o’malley’s (2005). the 5 stage methodological framework guided this study through identifying the research question, searching for relevant studies, selecting studies, charting the data, and collating, summarising, and reporting the results. the sixth, optional, stage of consulting with stakeholders to inform or validate study findings will be done as a way of developing the research paper. stage 1: involved the development of a research question. the following questions guided this review: 1) what are the requirements for establishing advanced research data repositories?; 2) what are the successes and challenges faced in establishing and managing research data repositories? stage 2: involved identifying relevant studies. peer-reviewed articles, book chapters, and conference proceedings were retrieved from scopus, web of science, and dimensions. the following search terms were used: “establish research data repository”, “research data repository requirements”, “research data repository and academic library”, “research data repository and research institutions”, and “research data librarian experiences”. stage 3: involved article selection using the inclusion and exclusion criteria. the preferred reporting items for systematic reviews and meta-analyses extension for scoping reviews (prisma-scr) flowchart of article identification, screening, and extraction was used as shown in figure 1. the selection was initially based on the titles, keywords, and abstracts. stage 4: involved data charting and extraction, where the selected articles were subjected to further screening and were retrieved from their databases, and each full-text article was examined. the data was documented on an excel spreadsheet, and two reviewers reviewed the full-text articles to come up with relevant articles for the study. eligible studies met the following criteria: 1) published between https://doi.org/10.29173/iq1099 4/14 chigwada, j & chiware, e (2024) future models and architecture of data repositories in african universities, iassist quarterly 48(3), pp. 1-14. doi: https://doi.org/10.29173/iq1099 2013 and 2023; 2) written in english; and 3) included a discussion of the establishment of research data repositories in research institutions. stage 5: involved collating, summarising, and reporting the results. following a thorough reading of the articles, the authors used the research questions as a guide to identifying themes. all authors drafted and approved the report. figure 1: prisma flowchart findings and discussion during the study, 3,501 articles were retrieved, and 50 articles were considered for this study after screening. details for the screening, exclusion, and inclusion process can be found in figure 1. https://doi.org/10.29173/iq1099 5/14 chigwada, j & chiware, e (2024) future models and architecture of data repositories in african universities, iassist quarterly 48(3), pp. 1-14. doi: https://doi.org/10.29173/iq1099 requirements for establishing research data repositories repositories can be categorised into three groups based on the scope of content they collect and manage. the three groups are: domain repositories, discipline repositories, and institutional repositories and these differ according to the scope of content they collect and manage (lee and stvilia, 2017). this study was focused on institutional research data repositories. patel (2016) developed a three-tier conceptual framework that is aimed at providing guidelines to address research data management issues at an institutional level. the first tier deals with data management, which involves developing institutional policy for data sharing, changing the mindset of researchers, data collection from researchers, copyright and data licensing, cross refer data to methodologies, data classification, data anonymization, data description and identification, data organisation, and an interoperability framework for data. the second tier is about data storage and hosting and covers the selection of file formats, data generated by private-public collaborations, data hosting services, independent data contributions, liability for hosted data, data security, data hosting software, and data backup. the third tier is about data usage and looks at access to data, copyright, data licensing, and rights in derivative works. a useful working guide for higher education institutions planning to start research data management services was developed by jones et al. (2013). nie et al. (2021) indicated that the implementation of research data management services included project kickoff, needs assessment, partnerships establishment, software investigation and selection, software customisation, and data curation services and training. cox et al. (2017) and cox et al. (2019) developed a research data management landscape maturity model with four levels spanning from none, basic, developing, and extensive. level zero deals with audits and surveys to solicit information concerning service and support, level one is the compliance stage with research data management governance boards and research data management policy, and level two deals with capacity-building and reengineering looking at skills, roles, and structures. the activities at levels one and two overlap and they deal with research data management training, data literacy, and advisory services (awareness of data archives, publication, citation storage, data management planning tools, and rights or intellectual property). level three deals with stewardship where there are cultural acceptance and embedded practices looking at data repositories, technical support (selection, catalogue, curation, preservation, metadata), data analysis or visualisation, and research data management shared services. the model was revised and the new model retained the concept of four levels where level one was changed to compliance, level two stewardship, and level three transformation. the major change was in skills where there is the transition of existing skills on level one, reskilling of existing staff on level two, and new skills acquisition on level three (cox et al., 2019). chigwada et al. (2019) proposed a framework for establishing research data management services in zimbabwe which consists of strategies, policies, guidelines, processes, technologies, and services. mushi et al. (2020) developed a planned implementation strategy that can be used by a university to establish research data management services. it includes four phases which include strategy, policy, procedures and infrastructure (phase 1), awareness creation, skills development and repository content development (phase 2), management of active data (phase 3), and data selection and preservation (phase 4) as shown in figure 2. knight (2015) also stated that it is important to determine the research data management requirements within the institutions so that the service would support the evolving needs of researchers. issues such as funding, institutional data management infrastructure, research data management policies and procedures, a research data management website, and expertise should be considered when establishing research data management services at an institution. dora and kumar (2015) pointed out the factors that should be considered in designing and developing research data management services, they include understanding the needs of the various stakeholders, adopting standard recommendations, choosing the software (developing https://doi.org/10.29173/iq1099 6/14 chigwada, j & chiware, e (2024) future models and architecture of data repositories in african universities, iassist quarterly 48(3), pp. 1-14. doi: https://doi.org/10.29173/iq1099 one or adopting an existing commercial or open source software), reviewing the it infrastructure and then developing institutional guidelines. they suggested that institutions should choose from databank, ckan, dataverse, figshare, dryad, or harvard datacerse network (nie et al., 2021; dora and kumar, 2015). figure 2: rdm implementation phases (source: adapted from mushi et al., 2020). experiences of universities in establishing research data repositories the articles reviewed demonstrated the variety of experiences at universities from around the world. knight (2015) noted that they identified various stakeholders that affect research data management and strengthened the institution’s policy framework to address the needs of researchers. as a result, they came up with best practices that can be followed when offering research data management services. these include creating, managing, and sharing research data by contractual, legislative, regulatory, ethical, and other relevant requirements; creating a data management plan for all research projects that capture data; and registering all research data created, no matter where it is hosted. it was noted that some researchers were not willing to share their research data, although they wanted to use research data produced by others (bangani and moyo, 2019). chiware and becker (2018) found out that research institutions in southern africa were offering various services such as support with data management plans, reskilling librarians, reference to highperformance computing centres, dedicated web pages, and advice on data preservation. at the cape peninsula university of technology, research data management services were developed as an eresearch information and communication infrastructure which included several components such as infrastructure development, information flow and management, communication with researchers, development of tools related to the full research life cycle and the means to store, curate, and retrieve data as well as the training of researchers (chiware and mathe, 2015). they emphasised the need for a national e-research infrastructure that would enable the preservation of research data. the university of hong kong utilised the research data stewardship framework which covered policy and https://doi.org/10.29173/iq1099 7/14 chigwada, j & chiware, e (2024) future models and architecture of data repositories in african universities, iassist quarterly 48(3), pp. 1-14. doi: https://doi.org/10.29173/iq1099 procedure settings for research data planning, the establishment of research data infrastructure, data curation services, and online resources and guidelines. john hopkins university developed a new model of data management services involving storage, archiving, preservation, and curation layers (shen and varvel, 2013). cox et al. (2014), naume (2014), searle et al., (2015), perrier and barnes (2018) and martin-melon et al. (2023) stated that librarians had been playing a leading role in the establishment of research data management services and libraries had been leading in policy development, as supported by cox and pinfield (2014). it was noted that libraries were offering advisory, support, and training services rather than technical services, although there were some indications that librarians were upskilling to be able to remain relevant in the new research data management landscape. they added that librarians were not doing the research data management services alone but involved other key stakeholders from the it services department and research support offices, legal office, including the researchers themselves (akers et al., 2014, cox and pinfield, 2014, cox et al., 2017). davidson et al. (2014) pointed out the activities that were done in supporting research data services by the digital curation centre which include understanding funding bodies’ policies, working with individual uk universities to scope rdm and data sharing challenges and opportunities, fostering rdm skills development, supporting data management planning, facilitating data discovery, and assessing research data management costs and benefits. challenges faced when establishing research data repositories institutions encounter several challenges in managing research data, as stated by patel (2016), chigwada (2022), chiware (2020), masenya (2021), chigwada et al. (2017), chiparausha and chigwada (2019), patterton et al. (2018), tang and hu (2019), al-jaradat, (2021), ashiq et al. (2021), huang et al. (2021), ran et al. (2021), m’kulama et al. (2022), chiware and becker (2018), koopman and de jager (2016), chiware and mathe (2015), raju (2014), mohammed and ibrahim (2019), and nhendodzashe and pasipamire (2017). it was noted that the challenges go beyond the institutional level but also include national challenges, as stated by schopfel and rebouillat (2022) and knight (2015). the challenges that were pointed out include lack of storage space on institutional networks, limited computing power and cloud computing accessibility, poor state of research infrastructure, lack of government commitment to fund research data services, lack of clear policy guidelines, uncertainty of software tools to use, uncertainty on documentation standards to apply, security issues, interoperability issues, lack of skills and absence of research data management in some library schools, persistent brain drain, lack of awareness of rdm, and poorly resourced academic and research libraries as shown in table 1. akers et al. (2014) indicated that the challenges that were faced by the eight us universities they studied included difficulties reaching out to researchers for assistance with research data management services and seeking funding for the human resources needed and infrastructure. cox et al. (2017) indicated that libraries play a leading role in offering research data management services but are facing challenges such as low levels of engagement by key stakeholders, uncertainty on the technical infrastructure required, and funding issues. the findings from the study noted a lack of recognition of the need for tackling research data management at the institutional level, difficulties in getting institutional buy-in from the senior management, convincing some academics of the importance and worth of research data management services and some did not get support from the library management within the department (cox et al., 2014). the future of research data repositories the national institute of health (2023) pointed out the desirable characteristics for all data repositories which include unique persistent identifiers, long-term sustainability, metadata, curation and quality assurance, free and easy access, broad and measured reuse, clear use guidelines, security https://doi.org/10.29173/iq1099 8/14 chigwada, j & chiware, e (2024) future models and architecture of data repositories in african universities, iassist quarterly 48(3), pp. 1-14. doi: https://doi.org/10.29173/iq1099 and integrity, confidentiality, common format, provenance, and retention policy. in addition, the confederation of open access repositories (coar) (2022) stated the framework of good practice in repositories, which includes discoverability through the use of metadata standards; harvesting of metadata using oai-pmh; assigning of persistent identifiers, and registration on the registry of repositories; access including limiting access to sensitive research data; reuse through licencing information in the metadata record; integrity and authenticity to prevent unathorised manipulation of resources; quality assurance in line with the policies and procedures; preservation with a digital preservation and business continuity plan; sustainability and governance in terms of managing and funding the repository; and other documentation that provide the scope of the materials that are accepted in the repository. to be able to assess the usage of the research data repository, it should be able to show the usage metrics and citation count (downs et al., 2023). schopfel and rebouillat (2022) stated that the international best practice which includes registration in the re3data directory, certification through the coretrustseal certificate, the world data system, or dta seal of approval should be followed. a national approach to research data management was also suggested by keller (2015) and patterton et al. (2019). the issues of good practices and good standards were emphasised by trippel and zinn (2021). to develop institutional research data repositories that meet international standards, universities in africa should work on the proposed framework in figure 3. figure 3: proposed framework for developing data repository infrastructure. limitations of the study the study only considered published literature, which was easier to retrieve. there is a need to consider grey literature as well, in terms of unpublished reports that document the experiences of data librarians in establishing and maintaining research data repositories in research institutions. this would be done by incorporating the sixth (optional) stage of the methodological framework to consult data librarians as a follow-up to this proposed framework as a way of getting consumer and stakeholder involvement to get additional references and insights beyond those in the literature. https://doi.org/10.29173/iq1099 9/14 chigwada, j & chiware, e (2024) future models and architecture of data repositories in african universities, iassist quarterly 48(3), pp. 1-14. doi: https://doi.org/10.29173/iq1099 conclusion research data infrastructures across research domains, institutions, national boundaries, and beyond continue to grow as the need for good data management practices and sharing is now internationally recognised. the future of african data repositories depends on the development of sustainable platforms that have all the features of internationally trusted repositories, which are secure and driven by clear use guidelines and ensure integrity and confidentiality. the issue of costs is important and collaborative approaches in open source-based development are the only sustainable route to ensure the long-term curation and preservation of african-generated research data outputs. continuous skills development especially among university librarians, research offices, and central information technology services is important as the technological landscape is always in a state of constant change. international pressure, especially from donors, funders, and publishers is likely to drive speedy development and uptake of data repositories across institutions in africa. table 1: challenges faced when establishing rdm services challenge authors lack of storage space on institutional networks chiware and becker (2018), knight (2015), masenya (2021), patterton et al. (2018), tang and hu (2019) limited computing power and cloud computing accessibility knight (2015), patterton et al. (2018), poor state of research infrastructure chigwada et al, (2017), chigwada (2022), chiparausha and chigwada (2019), chiware (2020), cox et al. (2019), huang et al. (2021), mohammed and ibrahim (2019), patterton et al. (2018), tang and hu (2019) lack of government commitment to fund research data services chiware (2020) lack of clear policy guidelines al-jaradat, (2021), ashiq et al. (2021), chigwada et al, (2017), chigwada (2022), chiparausha and chigwada (2019), chiware and becker (2018), huang et al. (2021), masenya (2021), m’kulama et al. (2022), mohammed and ibrahim (2019), nhendodzashe and pasipamire (2017), ran et al. (2021), lack of mandate/ rewards chiware and becker (2018), cox et al. (2019), huang et al. (2021), masenya (2021), lack of institutional buy-in from senior management ashiq et al. (2021), chigwada et al, (2017), chiware and becker (2018), cox et al., 2014; cox et al. (2019), mohammed and ibrahim (2019), tang and hu (2019) uncertainty on documentation standards to apply knight (2015), tang and hu (2019) uncertainty of software tools to use and technical infrastructure required cox et al. (2017), knight (2015), mohammed and ibrahim (2019), security issues al-jaradat, (2021), chigwada et al, (2017), chiparausha and chigwada (2019), chigwada (2022), cox et al. (2019), huang et al. (2021), knight (2015), koopman and de jager (2016), patel (2016), patterton et al. (2018), interoperability issues knight (2015), mohammed and ibrahim (2019), https://doi.org/10.29173/iq1099 10/14 chigwada, j & chiware, e (2024) future models and architecture of data repositories in african universities, iassist quarterly 48(3), pp. 1-14. doi: https://doi.org/10.29173/iq1099 lack of skills al-jaradat, (2021), ashiq et al. (2021), chigwada et al, (2017), chiparausha and chigwada (2019), chigwada (2022), chiware and mathe (2015), cox et al. (2019, huang et al. (2021), masenya (2021), m’kulama et al. (2022), mohammed and ibrahim (2019), nhendodzashe and pasipamire (2017), patterton et al. (2018), raju (2014), ran et al. (2021), tang and hu (2019) absence of research data management in library schools raju (2014) persistent brain drain chiware (2020) poorly resourced academic and research libraries chigwada et al, (2017), cox et al. (2019), huang et al. (2021), funding ashiq et al. (2021), akers et al. (2014), chigwada et al, (2017), chiparausha and chigwada (2019), chiware and mathe (2015), chiware and becker (2018), cox et al. (2017), cox et al. (2019), huang et al. (2021), masenya (2021), mohammed and ibrahim (2019), patterton et al. (2018), tang and hu (2019) researchers not willing to partner ashiq et al. (2021), akers et al. (2014), chigwada et al, (2017), chiware and becker (2018), cox et al., 2014, huang et al. (2021), patel (2016), patterton et al. (2018), tang and hu (2019) low level of engagement by stakeholders ashiq et al. (2021), chigwada et al, (2017), cox et al. (2017), cox et al. (2019), huang et al. (2021), patel (2016), patterton et al. (2018), ran et al. (2021), tang and hu (2019) lack of recognition for tackling rdm at the institutional level ashiq et al. (2021), chigwada et al, (2017), cox et al., 2014, cox et al. (2019), huang et al. (2021), patel (2016), tang andhu (2019) lack of awareness ashiq et al. (2021), chigwada (2022), huang et al. (2021), patel (2016), tang and hu (2019) references abebe, r., aruleba, k., birhane, a., kingsley, s., obaido, g., remy, s. l., and sadagopan, s. (2021) ”narratives and counternarratives on data sharing in africa” in conference on fairness, accountability, and transparency (facct ’21), march 3–10, 2021, virtual event, canada. acm, new york, ny, usa, 12 pages. https://doi.org/10.1145/3442188.3445897. akers, k. g., sferdean, f. c., nicholls, n. h., and green, j. a. (2014) ”building support for research data management: biographies of eight research universities”, international journal of digital curation, 9(2), pp. 171–191. doi:10.2218/ ijdc.v9i2.327. al-jaradat, om. (2021) ”research data management (rdm) in jordanian public university libraries: present status, challenges and future perspectives”, the journal of academic librarianship, 47 (5) 102378. https://doi.org/10.1016/j.acalib.2021.102378. https://doi.org/10.29173/iq1099 https://doi.org/10.1145/3442188.3445897 https://doi.org/10.1016/j.acalib.2021.102378 11/14 chigwada, j & chiware, e (2024) future models and architecture of data repositories in african universities, iassist quarterly 48(3), pp. 1-14. doi: https://doi.org/10.29173/iq1099 arksey, h and o'malley, l. (2005) ”scoping studies: towards a methodological framework”, international journal of social research methodology, 8 (1), pp. 19-32, doi: 10.1080/1364557032000119616. ashiq, m., saleem, qua and asim, m. (2021) ”the perception of library and information science (lis) professionals about research data management services in university libraries of pakistan”. libri, 71 (3), pp. 239-249. https://doi.org/10.1515/libri-2020-0098. bangani, s. and moyo, m. (2019) “data sharing practices among researchers at south african universities”, data science journal, 18 (28), pp. 1–14. doi: https://doi.org/10.5334/dsj-2019-028. chigwada, jp. (2022) "management and maintenance of research data by researchers in zimbabwe", global knowledge, memory and communication, 71 (4/5), pp. 193-207. https://doi.org/10.1108/gkmc-06-2020-0079. chigwada, j. p., hwalima, t., and kwangwa, n. (2019) ”a proposed framework for research data management services in research institutions in zimbabwe” in r. bhardwaj, and p. banks (eds.), research data access and management in modern libraries. hershey, pa: igi global. doi:10.4018/978-1-5225-8437-7.ch002, (pp. 29-53). chigwada, j., chiparausha, b. and kasiroori, j. (2017) “research data management in research institutions in zimbabwe”, data science journal, 16 (31), pp. 1–9, doi: https://doi.org/10.5334/dsj2017-031. chiparausha, b., and chigwada, j. p. (2019) ”accessibility of research data at academic institutions in zimbabwe” in r. bhardwaj, and p. banks (eds.), research data access and management in modern libraries. hershey, pa: igi global. doi:10.4018/978-1-5225-8437-7.ch004. (pp. 81-89). chiware, e. (2020) ”open research data in african academic and research libraries: a literature analysis”, library management, 41 (6/7,) pp. 383-399. doi 10.1108/lm-02-2020-0027. chiware, e and becker, da. (2018) ”research data management services in southern africa: a readiness survey of academic and research libraries”, afr. j. lib. arch. & inf. sc., 28, (1), pp. 1-16. chiware, e and mathe, z. (2015) ”academic libraries’ role in research data management services: a south african perspective”, sa jnl libs & info sci, 81(2), pp. 1-10. doi:10.7553/81-2-1563. coar. (2022) coar community framework for good practices in repositories, version 2 july 19. available at: http://www.coar-repositories.org/ (accessed 10 september 2023). cox, am., kennan, ma., lyon, l., pinfield, s and sbaffi, l. (2019) ”maturing research data services and the transformation of academic libraries”, journal of documentation, 75 (6), pp. 1432-1462. 00220418 doi 10.1108/jd-12-2018-0211. cox, am., kennan, ma., lyon, l., and pinfield, s. (2017) ”developments in research data management in academic libraries: towards an understanding of research data service maturity”, journal of the association for information science and technology, 68(9), pp. 2182–2200. cox, a.m., and pinfield, s. (2014) ”research data management and libraries: current activities and future priorities”, journal of librarianship and information science, 46, pp. 299–316. http://doi.org/10.1177/ 096100061349254. https://doi.org/10.29173/iq1099 https://doi.org/10.1515/libri-2020-0098 https://doi.org/10.5334/dsj-2019-028 https://doi.org/10.1108/gkmc-06-2020-0079 https://doi.org/10.5334/dsj-2017-031 https://doi.org/10.5334/dsj-2017-031 http://www.coar-repositories.org/ http://doi.org/10.1177/%20096100061349254 12/14 chigwada, j & chiware, e (2024) future models and architecture of data repositories in african universities, iassist quarterly 48(3), pp. 1-14. doi: https://doi.org/10.29173/iq1099 davidson, j., jones, s., molloy, l and kejser, ub. (2014) ”emerging good practice in managing research data and research information within uk universities”, procedia computer science, 33, pp. 215 – 222. dora, m and kumar, ha. (2015) ”managing research data in academic institutions: role of libraries” in 10 th international caliber-2015 hp university and iias, shimla, himachal pradesh, india march 1214, inflibnet centre, gandhinagar, gujarat, india, pp 484-495. downs, rr, urquidi díaz, a, xu, q, wang, j, chambodut, a, liu, c, flower, s and payne, k. (2023) ”harvestable metadata services development: analysis of use cases from the world data system”, data science journal, 22 (20), pp. 1–20. doi: https://doi.org/10.5334/dsj 2023-020. huang, y., cox, a. and sbaffi, l. (2021) ”research data management policy and practice in chinese university libraries”, journal of the association for information science and technology, 72 (4), pp. 493506. https://doi.org/10.1002/asi.24413. jones, s., pryor, g. and whyte,a. (2013) how to develop research data management services a guide for heis. dcc how-to guides. edinburgh: digital curation centre. available at: http://www.dcc.ac.uk/resources/how-guides/how-develop-rdmservices#sthash.kfiimtyy.dpuf. (accessed 10 september 2023). keller, a. (2015) ”research support in australian university libraries: an outsider view”, australian academic & research libraries, 46 (2), pp. 73-85, doi: 10.1080/00048623.2015.1009528. knight, g. (2015) ”building a research data management service for the london school of hygiene & tropical medicine”, program: electronic library and information systems, 49 (4), pp. 424-439. doi 10.1108/prog-01-2015-0011. koopman, mm. and de jager, k. (2016) ”archiving south african digital research data: how ready are we?”, s afr j sci., 112 (7/8), pp. 1-7. http://dx.doi.org/10.17159/sajs.2016/20150316. lee, dj. and stvilia, b. (2017) ”practices of research data curation in institutional repositories: a qualitative view from repository staff”, plos one, 12 (3). https://doi.org/10.1371/journal.pone.0173987. levac, d., colquhoun, h. and o’brien kk. (2010) ”scoping studies: advancing the methodology”, implement sci. 5 (69). doi:10.1186/1748-5908-5-69. lin, d., crabtree, j., dillo, i. et al. (2020). ”the trust principles for digital repositories”. scintific data. 7, 144 (2020). https://doi.org/10.1038/s41597-020-0486-7 martin-melon, r., hernandez-perez, t. and martinez-gardama. (2023) ”research data services (rds) in spanish academic libraries”, the journal of academic librarianship. 49 (2023) 102732. https://doi.org/10.1016/j.acalib.2023.102732 masenya, tm. (2021) "research data management practices and services in south african academic libraries", library philosophy and practice (e-journal), 6311. https://digitalcommons.unl.edu/libphilprac/6311. m’kulama, acm., zulu, z, chewe, p and mwiinga, tm. (2022) "preparedness for open science through research data management at the university of zambia in covid-19 and post-covid eras", library philosophy and practice (e-journal). 7268. https://digitalcommons.unl.edu/libphilprac/7268. https://doi.org/10.29173/iq1099 https://doi.org/10.1002/asi.24413 http://www.dcc.ac.uk/resources/how-guides/how-develop-rdmservices#sthash.kfiimtyy.dpuf http://dx.doi.org/10.17159/sajs.2016/20150316 https://doi.org/10.1371/journal.pone.0173987 https://doi.org/10.1038/s41597-020-0486-7 https://doi.org/10.1016/j.acalib.2023.102732 https://digitalcommons.unl.edu/libphilprac/6311 https://digitalcommons.unl.edu/libphilprac/7268 13/14 chigwada, j & chiware, e (2024) future models and architecture of data repositories in african universities, iassist quarterly 48(3), pp. 1-14. doi: https://doi.org/10.29173/iq1099 mohammed, ms. and ibrahim, r. (2019) ”challenges and practices of research data management in selected iraq universities”, desidoc journal of library & information technology, 39 (6), pp. 308-314. doi : 10.14429/djlit.39.6.14443. mushi, ge, pienaar, h. and van deventer, m. (2020) ”identifying and implementing relevant research data management services for the library at the university of dodoma, tanzania”, data science journal, 19 (1), pp. 1–9. doi: https://doi.org/10.5334/dsj-2020-001. national institute of health. (2023) selecting a data repository. available at: https://sharing.nih.gov/data-management-and-sharing-policy/sharing-scientific-data/selecting-adata-repository. (accessed 10 september 2023). naum, a. (2014) ”research data storage and management: library staff participation in showcasing research data at the university of adelaide”, the australian library journal, 63 (1), pp. 35-44. doi: 10.1080/00049670.2014.890019. nhendodzashe, n. and pasipamire, n. (2017) “research data management services: are academic libraries in zimbabwe ready? the case of university of zimbabwe library”, ifla satellite meeting, wroclaw. available at: https://library.ifla.org/id/eprint/1728/1/s06-nhendodzashe-en.pdf. (accessed 10 september 2023). nie, h., luo, p.c. and fu, p. (2021) ”research data management implementation at peking university library: foster and promote open science and open data”, data intelligence 3(1), pp. 189-204. doi: 10.1162/dint_a_00088. patel, d. (2016) “research data management: a conceptual framework”, library review, 65 (4-5), pp. 226-241. http://dx.doi.org/10.1108/lr-01-2016-0001. patterton, l., bothma, t.j.d. and van deventer, m.j. (2018) “from planning to practice: an action plan for the implementation of research data management services in resource-constrained institutions”, south african journal of library and information sciences, 84 (2), pp. 14-26. doi:10.7553/84-2-1761. perrier, l and barnes, l. (2018) ”developing research data management services and support for researchers: a mixed methods study”, partnership: the canadian journal of library and information practice and research, 13, (1). http://dx.doi.org/10.21083/partnership.v12i2.4115. raju, j. (2014) ”knowledge and skills for the digital era academic library”, the journal of academic librarianship, 40 (2), pp. 163-170. https://doi.org/10.1016/j.acalib.2014.02.007. ran, c., yang, l and hu, l. (2021) ”revisit the implementation status of research data management in chinese academia”, the journal of academic librarianship, 47 (2021) 102350. https://doi.org/10.1016/j.acalib.2021.102350. schopfel, j. and rebouillat, v. (2022) ”the landscape of research data repositories in france”, in schopfel, j. and rebouillat, v. (eds) research data sharing and valorization: developments, tendencies, models, pp31-48. https://doi.org/10.1002/9781394163410.ch2. searle, s., wolski, m., simons, n. and richardson, j. (2015) "librarians as partners in research data service development at griffith university", program: electronic library and information systems, 49 (4), pp. 440 – 460. http://dx.doi.org/10.1108/prog-02-2015-0013. https://doi.org/10.29173/iq1099 https://doi.org/10.5334/dsj-2020-001 file:///c:/users/chiwaree/appdata/local/microsoft/windows/inetcache/content.outlook/1cuczfa3/ file:///c:/users/chiwaree/appdata/local/microsoft/windows/inetcache/content.outlook/1cuczfa3/ https://sharing.nih.gov/data-management-and-sharing-policy/sharing-scientific-data/selecting-a-data-repository https://sharing.nih.gov/data-management-and-sharing-policy/sharing-scientific-data/selecting-a-data-repository https://library.ifla.org/id/eprint/1728/1/s06-nhendodzashe-en.pdf http://dx.doi.org/10.1108/lr-01-2016-0001 http://dx.doi.org/10.21083/partnership.v12i2.4115 https://doi.org/10.1016/j.acalib.2014.02.007 https://doi.org/10.1016/j.acalib.2021.102350 https://doi.org/10.1002/9781394163410.ch2 http://dx.doi.org/10.1108/prog-02-2015-0013 14/14 chigwada, j & chiware, e (2024) future models and architecture of data repositories in african universities, iassist quarterly 48(3), pp. 1-14. doi: https://doi.org/10.29173/iq1099 shen, y. and varvel, ve. (2013) ”developing data management services at the johns hopkins university” the journal of academic librarianship, 39 (6), pp. 552-557. https://doi.org/10.1016/j.acalib.2013.06.002. tang, r. and hu, z. (2019) “providing research data management (rdm) services in libraries: preparedness, roles, challenges, and training for rdm practice”, data and information management, 3 (2), pp. 84-101. https://doi.org/10.2478/dim-2019-0009. trippel, t. and zinn, c. (2021) ”lessons learned: on the challenges of migrating a research data repository from a research institution to a university library”, lang resources & evaluation, 55, pp. 91–207. https://doi.org/10.1007/s10579-019-09474-4. xiao, s., ng, t.y. and yang, t.t. (2022) "research data stewardship at the university of hong kong", library management, 43 (1/2) ,pp. 128-147. https://0-doi-org.oasis.unisa.ac.za/10.1108/lm-09-20210079. endnotes 1 josiline chigwada is a postdoctoral fellow at the university of south africa, school of interdisciplinary research and graduate studies. she can be reached by email: chigwaj@unisa.ac.za. 2 elisha chiware is the director at the cape peninsula university of technology libraries. he can be reached by email: chiwaree@cput.ac.za https://doi.org/10.29173/iq1099 https://doi.org/10.1016/j.acalib.2013.06.002 https://doi.org/10.2478/dim-2019-0009 https://doi.org/10.1007/s10579-019-09474-4 https://0-doi-org.oasis.unisa.ac.za/10.1108/lm-09-2021-0079 https://0-doi-org.oasis.unisa.ac.za/10.1108/lm-09-2021-0079 mailto:chiwaree@cput.ac.za vol28-4.indd 4 iassist quarterly winter 2004 editor’s notes welcome to the fourth issue of the iassist quarterly vol. 28. this issue brings us an article on wikipedia. i have to excuse the production time of the iq. when this article was received the concept of wikipedia was not so widespread as it rightfully is now. however the idea has proven durable. oliver watteler from the data archive (zentralarchiv) in cologne in germany – part of gesis – presents us for the fundamentals of wiki: the wikipedia is a web-based encyclopedia. it is a free and copyleft product to be used, read, updated and improved by anyone. the “copyleft” concept is closely connected to open source and freeware, which highlights the free “as in free speech, not as in free beer”. but you can read much more on this at the http://www.wikipedia.org. oliver watteler is proposing a wikipedea for iassist members with coverage of typical iassist topics. as he writes it: “wiki could carry the iassist spirit one step further into the virtual world”. remember to visit a part of the virtuality in iassist: the new iassist weblog (blog) iassist communiqué – at http://iassistblog.org. chiu-chuang (lu) chou is a senior special librarian at the data and program library service (dpls) at university of wisconsin madison, united states. she participated in a three-day workshop called statistical disclosure control for data confidentiality with the software argus. the argus software intention is to “modify unsafe data in such a way that safe (enough) data emerge, with minimum information loss”. as lu chou explains the statistical disclosure control is mainly concerned by obtaining a balance between the need for data and the need for confidentiality. the software was developed as part of the european commission’s fifth framework programme. this was addressing work towards a “user friendly information society” and statistical information is part of that society. lu chou also looks at the us federal statistical system and how statistical disclosure control is obtained there. at the iassist conference in edinburgh 2005 there was a session on “new insights in providing data services: a variety of evidence”. some of the evidence in the form of a case was the presentation on “data archiving at the us central bank”. the article was presented and is written by linda powell, who works at the board of governors of the federal reserve system where vast quantities of data are “consumed”. the article demonstrates and discusses several challenges in connection with the diverse pool of data, user access from several platforms as well as the growth of data archiving. the last also includes the dynamic challenge that data consistency over time means dealing with a changing subject, as some concepts and their calculation change over time. among the technological advances linda powell explicitly mentions xml that is used in several ways, including the xbrl standard. please visit at the iassist website on www.iassistdata. org. you can now find information on previous and coming conferences. the 2006 conference will take place at ann arbor, michigan (23-26 may). you will also find information on the publication award for 2006. furthermore, on the iassist website you will find access to the articles of the iassist quarterly as pdf-files. we hope you will enjoy all of this. papers for the iassist quarterly are most welcome. papers can be from iassist conferences, from other conferences, from local presentation, etc. contact the editor via e-mail: kbr@sam.sdu.dk. karsten boye rasmussen, october 2005 vol30-4.indd 4 iassist quarterly winter 2006 editor’s notes welcome to the fourth issue of the iassist quarterly, vol. 30. the 34th iassist annual conference will take place at stanford university, palo alto, california, usa, may 2730, 2008 with the theme “technology of data: collection, communication, access and preservation”. the conference will examine the role of technology and tools in various aspects of the data life cycle. see the call for papers in this issue of the iq. at the iassist 2007 conference in montreal the best conference ever one of the presentations in the session “data services mash-ups: maps, research and everything!” was rachael barlow from trinity college in hartford, ct presenting “maps that mash: daring, dangerous, or dumb?”. the article is now called “mashing maps”. barlow introduced a class at trinity college a small, liberal-arts college to the facilities of google maps mash-ups; the idea of the class was to attempt to do something meaningful with data for people living in the local community hartford, connecticut. the students in the class created several mash-ups, for example one group of students mapped food resources in hartford, everything from community gardens to grocery stores to food pantries. some maps are shown in the article. as a sociologist rachael barlow investigated the effects of the trinity college class, where the production of the mash-ups made a cooperative connection between the faculty and students inside the college’s walls and the general community members outside those walls. she found that the connection improved the image of the college and also that the creation of the mash-ups created a demand for similar work from other local organizations. barlow concludes that “mashups succeeded not only in mixing up online content and tools, but also people; in this case the students, local organizations, faculty, administration, and media that participated in the project’s upstream and downstream processes”. those were certainly very successful mix-ups. at the same 2007 conference – as you remember “the best ever” in the session “data beyond numbers: using data creatively for research”, jinfang niu from university of michigan presented what is now her paper on “reward and punishment mechanism for research data sharing”. the author is defining the problems in the concept of data as a public good: “in the data sharing case, data producers make efforts to prepare the data for deposit, but the benefit of the data preparation largely goes to secondary users. in addition, data producers are at risk of being harmed by the misuse and misinterpretation of data by unqualified users, or by being charged with misconduct. that makes freeriding even more attractive. to motivate data producers to prepare and share data, there must be some incentive mechanisms”. to add rigidity to the analysis jinfang niu presents some mathematical modeling of the issue. (a very bold reviewer accepted the challenge of the mathematics.) as many of the iq readers are working in data libraries and data archives, it could be easy simply to demand that all research data should be made available for the public. and this can be enforced by not releasing the last portion of grant money until that requirement is met, as in “trust is good, but control is better”. however, you could argue that the backside is that data that were better forgotten are now carried on indefinitely in archives. however, don’t expect that the mathematics can make the definitive judgment concerning whether to keep or to reject/scrap data. but do expect to get enlightened on the issues of reward and punishment, including a proposal to enforce citation of datasets an issue iassist has promoted since the foundation of the organization. in the same session janet stamatel from university at albany presented some very interesting thoughts on “the importance of data visualization in data literacy”. however, approached by one of the other presenters at the same session (this editor), janet stamatel preferred to work a bit more on the issue before publishing her presented thoughts on data visualization. so we as readers can look forward to that. in our talks it turned out that janet stamatel also had presented at the 2006 iassist conference presently the next-best ever and she agreed to publish her paper as an article in the iq with the title: “an overview of publicly available quantitative cross-national crime data”. the article is a specialist article that “reviews the content, data collection methods, geographic and temporal coverage, and accessibility of three main sources of publicly available, quantitative cross-national crime data, with a particular emphasis on recent changes with respect to data availability. these sources are the international police organization (interpol), the united nations crime surveys, and the european sourcebook”. if your subject area is not in crime data, then the article can be read as an important contribution to the issue of globalization, that also emphasizes the problems of validity in cross-national data when the author remarks “researchers have noted considerable inconsistencies across the sources”. remember to have a look at the website http://iassistdata. org and the iassist blog the iassist communiqué – at http://iassistblog.org. articles for the iassist quarterly are very welcome. articles can be papers from iassist conferences, from other conferences, from local presentations, discussion input, etc. contact the editor via e-mail: kbr@sam.sdu.dk. karsten boye rasmussen, november 2007 1/3 rasmussen, karsten boye (2023) editor’s notes: yes, we are international, iassist quarterly 47(2), pp. 1-3. doi https://doi.org/10.29173/iq1091 the creative commons-attribution-noncommercial license 4.0 international applies to all works published by iassist quarterly. authors will retain copyright of the work and full publishing rights. editor's notes: yes, we are international welcome to the second issue of iassist quarterly for the year 2023 iq vol. 47(2). i am very happy with the 'international' in iassist. it is important to learn from outside your own center. in this issue we have a focus on the united states and some african countries with a special focus on south africa. the first article investigates libguides across the many states of the united states. the second article is centered on one of the data resources often found in the libguides pages, but the data itself is about all of the united states. in the third article we shift to the african continent and the described project has a base in south africa with a connection to the united kingdom still part of europe although not of the eu and with research being conducted in several african countries. we can't promise to cover the whole world in each iq issue – but this issue is quite international. the first article is 'taking count: a computational analysis of data resources on academic libguides in the u.s.'. cody hennesy, alicia kubas and jenny mcburney have undertaken the task of collecting links to data and statistical resources from over 10,000 libguide pages at 123 r1 research institutions in the united states. the libguides platform has become the universal resource discovery platform in academic libraries in the u.s. libguides not only support researchers, they also help librarians in orientation among the many resources. the authors reach the conclusion that freely available resources from u.s. government agencies are the most widely used. resources requiring paid licenses or memberships (like icpsr) are also frequent. the analysis suggest traditional licensed statistical resources are more likely to be shared than complex microdata resources. data cleaning of the nearly 200,000 links from the 10,000 guide pages was an essential part of the analysis. the authors cite the data scientist joke that 90% of the work is data cleaning, and they find that the actual number for the cleaning and normalization in this analysis was even larger, performed through python and openrefine. the data process included accessing the libguide pages based on the keywords of 'data' and 'statistic' and then extracting the content links. the links were then cleaned, filtered and further normalized. the data cleaning showed a high degree of inconsistency and dead links, leading the authors to suggest a more centralized management of data resources. the most frequently found links to resources are through icpsr and data.gov, and a table with the 20 most common resources shows that even the most uncommon resource among these 20 are included in more than 73% of the institutions. this demonstrates a high consistency across the institutions. however, the authors remark that they believe that the very few institutions that didn’t include a link to the popular data.gov would benefit from having information about this resource available for their researchers. cody hennesy and jenny mcburney are the journalism & digital media librarian and a social sciences librarian at the university of minnesota, twin cities, and alicia kubas is a librarian at the u.s. government publishing office. https://doi.org/10.29173/iq1091 https://data.gov/ https://creativecommons.org/licenses/by-nc/4.0/ 2/3 rasmussen, karsten boye (2023) editor’s notes: yes, we are international, iassist quarterly 47(2), pp. 1-3. doi https://doi.org/10.29173/iq1091 the second article concerns metadata from ipums projects at the institute for social research and data innovation (isrdi) at the university of minnesota (note, these are among the central sources of data libguides, mentioned several times in the first article). the authors are diana l. magnuson, curator and historian at the institute for social research and data innovation, and wendy l. thomas, now retired curator from the same institution. the title is 'expanding our perspective: building a sustainable metadata culture'. the article describes the learning obtained by isrdi through the submission of an application for certification to the core trust seal (cts). when applying for certification the institution must document that it follows the standards and guidelines for the certification. in the case of the cts as in many other cases of certification the building of a portfolio of documentation of procedures makes the applicant more self-aware of its history, as well as of the routines delivering the final products. the conclusion is also that the certification process has led to a better internal understanding at the isrdi that can support future development as well as preserve the work done. ipums has over the last thirty years created the world’s largest accessible database of census microdata starting with the 1880 historical census project that has been extended in both time directions and now covering more than a hundred years. naturally, processing of data has changed over the years and keeping track of the documentation proved difficult. the decision to use digital object identifiers (dois) led to a persistency and uniqueness that supported the users. this also had internal benefits as references and publications were more easily trackable and the preservation work more accurate and complete for each product version. among the figures of the article, you will find the workflow using the open archival information system (oais) model as well as the ipums business process model. the third article concerns the dilemma of personal data protection versus the benefit of using data for life improvement. the title of the submission is 'data management instruments to protect the personal information of children and adolescents in sub-saharan africa' and concerns health research in this group. on the one hand the researchers naturally must follow the data regulations as they appear in the protection of personal information (popi) act in south africa and the general data protection regulation (gdpr) in the european union, and with special attention to high-risk and vulnerable groups such as children and adolescents. on the other hand, these vulnerable groups are also at risk from a health viewpoint, especially from infectious diseases like infantile paralysis, measles and pneumococci. research and data collected from children has contributed to the development of vaccines, which has led to a dramatic reduction in child mortality and improvements in the quality of life. the project described is a large-scale one that involves many countries and many researchers, making governance and data management crucial to achieving data availability and data security. the article discusses the strategies and instruments used, and addresses the many considerations from both ethical sides and when building a data management plan and decisions on sharing data. the authors behind the article are lucas hertzog, jenny chen-charles, camille wittesaele, kristen de graaf, raylene titus, jane kelly, nontokozo langwenya, lauren baerecke, boladé hamed banougnin, wylene saal, john southall, lucie cluver, and elona toska. many of these are affiliated to the centre for social science research at the university of cape town in south africa and some are connected to the university of oxford. it is important to mention that in addition to the central participation from south africa and the uk, the project is based on partnerships with researchers in zambia, malawi, nigeria, lesotho, tanzania, and kenya. https://doi.org/10.29173/iq1091 3/3 rasmussen, karsten boye (2023) editor’s notes: yes, we are international, iassist quarterly 47(2), pp. 1-3. doi https://doi.org/10.29173/iq1091 submissions of papers for the iassist quarterly are always very welcome. we welcome input from iassist conferences or other conferences and workshops, from local presentations or papers especially written for the iq. when you are preparing such a presentation, give a thought to turning your one-time presentation into a lasting contribution. doing that after the event also gives you the opportunity of improving your work after feedback. we encourage you to login or create an author profile at https://www.iassistquarterly.com (our open journal system application). we permit authors to have 'deep links' into the iq as well as deposition of the paper in your local repository. chairing a conference session or workshop with the purpose of aggregating and integrating papers for a special issue iq is also much appreciated as the information reaches many more people than the limited number of session participants and will be readily available on the iassist quarterly website at https://www.iassistquarterly.com. should you be interested in compiling a special issue for the iq as guest editor(s) you can also contact the iq. take a look at the instructions, layout, and contact at: https://www.iassistquarterly.com/index.php/iassist/about/submissions on a personal note, i have since 1997 been the editor of the iassist quarterly. all good things must end. new people will take over and improve the journal. i find there have been many improvements in the iq during my tenure. special thanks to my good friends walter and jane for their work on the journal. for many years, walter piovesan helped with layout and production, and he established contact with the open journal system staff before retiring from the iq editorial team. jane roberts turned my danglish into english in my iq editorials. i am very happy to quit now, especially because you iassisters will have very competent replacements in michele hayslett and ofira schwartz. they have already for long worked behind the scenes at iq, and have also edited the recent special issue on systemic racism. the iq is in good hands. karsten boye rasmussen june 2023 https://doi.org/10.29173/iq1091 https://www.iassistquarterly.com/ https://www.iassistquarterly.com/ https://www.iassistquarterly.com/index.php/iassist/about/submissions 1/14 kwanya, tom (2021) publishing trends on research data management in sub-saharan africa: a bibliometric analysis, iassist quarterly 45(3-4), pp. 1-14. doi: https://doi.org/10.29173/iq996 publishing trends on research data management in sub-saharan africa: a bibliometrics analysis tom kwanya1 abstract research data management (rdm) is the all-encompassing term used to describe the processes and activities related to the creation, storage, security, preservation, retrieval, reuse and sharing of research data. as is often the case, researchers from sub-saharan africa are lagging behind their counterparts in developed countries in embracing best practices in research data management. one of the factors to which this slow pace of adoption of research data management could be attributed, is inadequate research on the subject. this paper analyses the authorship, volume, visibility and quality of publications on research data management in sub-saharan africa. the analysis was done using bibliometrics. the units of analysis were publications on research data management from, and on, sub-saharan africa which are currently indexed in google scholar. this index was chosen because it is free and is reputed for its liberal selection criteria which does not favour, or discriminate, any discipline or geographic region. data from google scholar was retrieved using harzing’s “publish or perish” software and analysed using nees jan van eck’s vosviewer software. the findings of the study revealed that authorship collaboration, visibility, quality, and quantity of scholarly publications on research data management in sub-saharan africa is low when compared to developed countries like the united states of america, the united kingdom, canada and australia. keywords research data management, rdm, bibliometrics, informetrics, sub-saharan africa, publishing trends 1. introduction schöpfel et al. (2018) assert that research data is multifaceted and dynamic. this makes it easier to describe than define. however, this paper adopts the definition of research data provided by ray (2014) as any qualitative or quantitative evidence collected, observed, generated or created through, and for, scientific research. this data is typically used to interpret, describe or understand the phenomena under study. the data enables researchers to make conclusions which validate or generate knowledge about the phenomena being studied. briney (2015) argues that research data consists of facts and statistics which are collected for analysis and reference in regard to a specific research project. according to ng’eno and mutula (2018) research data is a valuable and unique resource which is irreplaceable and costly to reproduce. therefore, research data should be managed adequately to preserve its value while enhancing its use and usability. according to ng’eno and mutula (2018), rdm is the all-encompassing term applied to refer to the processes used in data creation, storage, security, preservation, retrieval, re-use and sharing. adika and kwanya (2020, p. 447) argue that “rdm also encompasses activities and processes aimed at enhancing the preservation, security, visibility and ethical use of research data. these activities and processes are related to the creation, organisation, structuring, naming, backing up, storage, conservation, and sharing of research data as well as all actions that guarantee the security of research data”. whyte and tedds (2011) opine that the purpose of rdm is to facilitate effective verification of research data and enable new research to be anchored on existing research. https://doi.org/10.29173/iq996 2/14 kwanya, tom (2021) publishing trends on research data management in sub-saharan africa: a bibliometric analysis, iassist quarterly 45(3-4), pp. 1-14. doi: https://doi.org/10.29173/iq996 effective rdm practices yield several benefits. for instance, the effective management of research data provides a sound and reliable foundation upon which to anchor future research and thereby advancing scholarship. thus, effective rdm ensures the continued existence of valid data upon which current and future research can be founded. effective rdm also enables reuse of research data generated, thereby saving costs. rdm also saves the researchers’ time since loss of data or duplication of efforts by recreating existing data are avoided. briney (2015) avers that access and use of existing research data enables researchers to use existing research to generate new knowledge in a cost-effective manner. the place of research data management in the modern research ecosystem is increasingly prominent. indeed, many research funding agencies and publishers have instituted comprehensive rdm requirements. patterton et al. (2018) explain that researchers who apply for funding from the national research foundation (nrf) in south africa are required to demonstrate how they plan to disseminate the data generated by their proposed research projects to other researchers. one recommended strategy of ensuring wide access to research data is archiving it in publicly accessible repositories. similar trends are being applied by funding agencies in the united states of america, australia, canada, and the united kingdom (kahn et al., 2014). it has also been observed that reputable journals such as nature require authors to avail data anchoring their manuscripts for validation (nature, 2014). koopman (2015) asserts that research data is often required by peer reviewers to enable them to verify the findings and prevent scholarly fraud. due to the growing significance attached to research data, effective rdm literacy is now an essential skill amongst researchers. according to adika and kwanya (2020, p. 461), rdm literacy encompasses the capacity of researchers to effectively “plan for, search, find, organise, store, secure, share research data competently”. the authors recommend a training for researchers on effective rdm. besides basic data management, such training, they suggest, may also include how to use digital platforms for sharing research such as institutional repositories, use of credible databases of research material, as well as citation and reference management for scholarly purposes. 2. literature review according to koopman (2015, p. 1), data is “the currency of academic research”. denny et al. (2015, p. 294) assert that researchers describe data as “the lifeblood of their work”. this is due to the fact that research data is intricately intertwined with research performance, and excellence, which influence research networks and support. on their part, chawinga and zinn (2019) opine that scholarly advancement is propelled by research data. the growing acceptance of the significance of data in research is fanning the increased production of data from research projects. several scholars also opine that the rapid growth of research data is catalysed by the ubiquity of digital technologies, devices and networks which make creation, processing and preservation of research data easier (ajibaje & mutula, 2020; asher, 2012; kuo & kusiak, 2019; kwanya et al., 2014; neubert & trischler, 2021). consequently, research data is increasingly being produced in vast volumes and diverse formats (kibeet al., 2020). the existing abundance of research data holds great potential for the advancement of science and innovation (denny et al., 2015). however, kibe et al. (2020) explain that the value of the abundant data cannot be unlocked without effective rdm in terms of access and analytics. according to koopman (2015), a large portion of the existing research data is imperceptible and does not make any meaningful contribution to scientific research and development. therefore, many scholars conclude that access to research data should be enhanced as a means of optimising its potential (adika & kwanya, 2020; kuo & kusiak, 2019; lucasdominguez et al., 2021; vlahou et al., 2021). https://doi.org/10.29173/iq996 3/14 kwanya, tom (2021) publishing trends on research data management in sub-saharan africa: a bibliometric analysis, iassist quarterly 45(3-4), pp. 1-14. doi: https://doi.org/10.29173/iq996 the visibility of existing research data can be enhanced by executing a number of strategies. one of these strategies is to sensitise researchers to understand that by sharing their research data, their research becomes more visible and thereby attracts more citations (van noorden, 2014). however, chawinga and zinn (2019) argue that publicising the benefits of sharing research data alone may not yield the desired results. koopman (2015) opines that some researchers invest immense resources in generating, creating or accumulating research data. thus, they hold the data sets as valuable resources which they would not easily share with anyone else. this data is so valuable to their career that they guard it jealously. also, according to denny et al. (2015), researchers are reluctant to share their data because of the fear that data may be misunderstood and applied in situations which jeopardise the integrity of the original research. nonetheless, anane-sarpong et al. (2018) underscore the inevitability of research data sharing and urge researchers to make adequate preparations for it. several factors hinder research data sharing. these include “lack of time and data misappropriation at the individual level; inadequate data sharing training, absence of compensation and unfavourable internal policies at the institutional level; as well as weak policies, inadequate ethical and legal norms, lack of data infrastructure and interoperability issues at the international level” (chiwanga & zinn, 2019, p. 109). denny et al. (2015) also explain that there is a feeling amongst some researchers, especially in developing economies, that research data sharing may in some cases lead to neo-colonialist behaviour of foreign researchers taking away valuable resources from their territories. anane-sarpong et al. (2018) explain that risks and fears of researchers in under-resourced regions, like sub-saharan africa are less-reported. these include “risks faced by under-resourced scientists and institutions which are slower in translating data produced into new knowledge; absence of harmonised guidelines and structures to help address the risks and institute fairness in data-sharing rewards; and inadequate confidence in available protective safeguards including guidelines” (p. 404). despite all the challenges hindering effective sharing of research data, patterton et al. (2018) argue that their ubuntu2 spirit, researchers in sub-saharan africa exhibit a general willingness to share data with other researchers. they suggest that this positive attitude can be harnessed to inspire researchers to disseminate research data within local research communities and thereafter, gaining the confidence to share it outside. chawinga and zinn (2019, p. 404) suggest that research data sharing can be enhanced by “recognising researchers who share data through data citations, acknowledgement and incentives; investing in infrastructure, conducting training and advocacy programmes; as well as formulating stringent and fair policies for data sharing”. anane-sarpong et al. (2018) also suggest that commodifying research so as to facilitate fees for shared data may motivate researchers in developing economies to share their data. the authors add that this arrangement may provide additional resources for research in developing economies. according to patterton et al. (2018), the behaviour of researchers is essentially similar globally. however, ng’eno and mutula (2018) conducted an analysis of diverse perspectives of rdm in the united kingdom, australia, canada, united states of america, south africa and kenya and concluded that african researchers are lagging behind their contemporaries in the developed world in adopting research data management tenets. pisani et al. (2016) identify the reasons why scholars in africa lag behind the rest of the world in rdm. the reasons include fear to lose control of their shared research data; inadequate incentives for sharing research data; as well as lack of effective technological capacity and infrastructure relevant to rdm. adika and kwanya (2020) also argue that inadequate rdm literacy among researchers in sub-saharan africa constrains their capacity to effectively share research data. in spite of these challenges, there are some efforts amongst researchers in africa to share data. for instance, in south https://doi.org/10.29173/iq996 4/14 kwanya, tom (2021) publishing trends on research data management in sub-saharan africa: a bibliometric analysis, iassist quarterly 45(3-4), pp. 1-14. doi: https://doi.org/10.29173/iq996 africa, denny et al. (2015) found that rdm practices in the country were ad hoc and less formal. nonetheless, the authors report that there were efforts to institutionalise and enforce rdm policies by scholarly and research funding organisations. townsend (2021), with a focus on legal frameworks for sharing health research data in south africa, holds a similar view and suggests new legal mechanisms to regulate and strengthen data sharing in and out of the country. the situation is more or less similar in kenya where ng’eno and mutula (2018) acknowledged that initial efforts are being made but pointed out the need to strengthen institutional rdm capacities and invigorate resource mobilisation for research data sharing, preservation and reuse. 3. rationale and methodology of study the literature reviewed reveals that there is increased acceptance of data-driven research. consequently, there is increased production of research data. this calls for effective rdm. as is often the case in other scholarly metrics, researchers from sub-saharan africa are lagging behind their counterparts in developed countries in embracing best practices in managing research data. one of the factors to which this slow pace in research data management could be attributed, is inadequate research on the subject. whereas several studies, as indicated in the literature review, have been conducted on diverse aspects of rdm in sub-saharan africa, no study was found which has investigated the publication trends on rdm. the purpose of this paper is to analyse the quality, quantity, visibility and authorship of publications on research data management in sub-saharan africa as a means of bridging the gap in literature on this subject. norton (2001) explains that bibliometrics is an approach in the measurement of information. kwanya et al. (2021) explain that in this approach, metadata of publications, including author, publication date, publication channel, and citations are used to assess the quantity, quality and visibility of research output. over the years, bibliometrics has been traditionally linked to quantitative measurement of scholarly materials as a means of quantifying research productivity and excellence. wormell (2001) opines that bibliometrics has been widely applied to assess the production of research output. the advantages of using bibliometrics in research are numerous. however, its capacity to quantify research productivity and excellence has been acclaimed. it is also reputed to be an objective approach to examining knowledge exchange among scholars (dayu, 2012). nonetheless, neuhaus and damiel (2008) as well as kwanya et al. (2021) acknowledge the demerits of bibliometrics. paradoxically, one of the major demerits is linked to its quantitative focus which some scholars view as constricting qualitative perspectives to issues under research (peng & luo, 2021). there is also a view that bibliometrics may be prone to manipulations on issues such as citations (kwanya et al., 2021). these demerits notwithstanding, bibliometrics offered the best mechanism for conducting this study since it can unravel issues of research which other approaches may fail to detect. furthermore, it can also enable researchers to examine how knowledge is created and shared in scholarly communities. bibliometrics approaches were used to analyse publications on research data management from, and on, sub-saharan africa which are currently indexed in google scholar. the index was chosen because it is free and is reputed to have liberal selection criteria which do not favour, or discriminate, any discipline or geographic region. the publications were identified using harzing’s “publish or perish” software. the search was conducted using two key phrases which were “research data management” and “sub-saharan africa”. a total of 184 publications were retrieved. https://doi.org/10.29173/iq996 5/14 kwanya, tom (2021) publishing trends on research data management in sub-saharan africa: a bibliometric analysis, iassist quarterly 45(3-4), pp. 1-14. doi: https://doi.org/10.29173/iq996 4. findings and discussions the findings are structured according to the key themes of the objectives of the study. these are quantity, quality, visibility, and authorship of research publications on rdm in sub-saharan africa. 4.1 quantity of research publications on rdm in sub-saharan africa the findings indicate that of the 184 publications retrieved, the latest were published in 2020 while the oldest was published in 1985. this implies that rdm has been a research issue in sub-saharan africa for about 35 years. the publication trend over the years is as indicated in figure 1. it is evident from the data that the number of publications on rdm grew exponentially from 2006. this indicates a growing interest in the subject over the period. the largest number of publications was put out between 2016 and 2020. the highest number of publications per year was 25, attained in 2018. this was followed by 24 publications in 2019 and 2017 as well as 23 in 2020. only 12 publications were produced in 2016. a quick search on google scholar using the harzing’s software yields more than 1,000 publications on rdm from the united states of america, australia and the united kingdom. for the united states of america, for instance, the first publication on rdm was registered in 1941 which is nearly fifty years before the first paper on the subject in sub-saharan africa. comparatively, in 2020 alone, 929 publications on rdm were produced in the united states of america. similarly, a total of 369 and 330 publications on the subject were produced in the united kingdom and australia, respectively, in the same period. therefore, it can be deduced from the findings above that the quantity of publications on rdm in sub-saharan africa is low. the publication trend revealed above generally follows the overall trends in the production of knowledge by sub-saharan africa when compared to other economies. according to siyanbola et al. (2016), the level of production and use of scientific knowledge by countries in sub-saharan africa is low compared to the rest of the world. they add that the gap between sub-saharan africa and other regions in knowledge production and use has persisted in spite of myriad strategies being executed to bridge it. according to tijssen (2007), the impact of the research out of africa is significantly below the world’s average. it can, therefore, be concluded that the production of scientific publications on research data management in sub-saharan africa, just like in other subject areas, is lower than that of the rest of the world. this can be attributed to many factors key of which are inadequate research funding and infrastructure. using vosviewer, the keywords in the titles and abstracts of the retrieved papers were identified and visualised. the colour coding is used to distinguish the nodes (keywords) in the visual presentation. the connecting lines indicate linkages between the keywords. therefore, connected keywords imply that they appeared together in either the titles (figure 2) or abstracts (figure 3) of the identified papers. from the figures, it is evident that the most prominent themes covered include research data management, university, research, data, data management, and south africa. from this, it can be concluded that the publications cover a range of topics in research data management. the prominence of south africa (as seen in figure 3) implies that most of the publications are produced in or about south africa. the possible higher production rate of publications on research data management by south africa may be attributed to the fact that it is a bigger economy with greater advancement in education, science and technology than other countries in the region. similarly, the presence of “university” as a key term in both the titles and abstracts indicates that most of the studies were focused on or conducted in universities. this is expected because universities are critical institutions in terms of research data production and use. https://doi.org/10.29173/iq996 6/14 kwanya, tom (2021) publishing trends on research data management in sub-saharan africa: a bibliometric analysis, iassist quarterly 45(3-4), pp. 1-14. doi: https://doi.org/10.29173/iq996 figure 1: publishing trends on rdm in sub-saharan africa figure 2: keywords in titles of the publications 4 2 1 4 20 45 108 0 20 40 60 80 100 120 1985-1990 1991-1995 1996-2000 2001-2005 2006-2010 2011-2015 2016-2020 n u m b e r o f p ap e r p u b lis h e d year brackets of publication https://doi.org/10.29173/iq996 7/14 kwanya, tom (2021) publishing trends on research data management in sub-saharan africa: a bibliometric analysis, iassist quarterly 45(3-4), pp. 1-14. doi: https://doi.org/10.29173/iq996 figure 3: keywords in abstracts of the publications 4.2 quality of research publications on rdm in sub-saharan africa in the context of this paper, quality of the publications was assessed based on the citations they received. the rationale for this is the assumption that high quality papers are used and cited more than those of low quality. indeed, this paper acknowledges the fact that many other factors influence the citability of papers. these include length of time of publication, number of authors, and topic of content, among others. in the context of this paper, however, quality was only assessed in terms of the number of citations the papers/publications attracted. it was observed that 91 of the 184 publications have not been cited at all. this is nearly half of all the publications. of the 93 publications which have been cited, only 24 attained at least 10 citations. these are listed on table 1. the majority (43) have been cited once, twice, or thrice. of the remaining publications, six have been cited four times; five have been cited five times; three have been cited six times; two have been cited seven times; three have been cited eight times; while two have been cited nine times. the total number of citations for all the publications is 887. therefore, the citation analysis leads to a conclusion that the quality of the publications on rdm in sub-saharan africa is also low. the low number of citations of the publications implies that sub-saharan africa perspectives to rdm are barely heard. these findings confirm the assertion by stewart (2015) that citations of scientific publications from subsaharan africa are lower than those from developed economies. tijssen (2007) acknowledges that the low number of citations limits the impact of scientific research from sub-saharan africa. the region is https://doi.org/10.29173/iq996 8/14 kwanya, tom (2021) publishing trends on research data management in sub-saharan africa: a bibliometric analysis, iassist quarterly 45(3-4), pp. 1-14. doi: https://doi.org/10.29173/iq996 therefore a net “importer” of scientific products and does not influence research agenda on any thematic areas of research, including research data management. 4.3 visibility of research publications on rdm in sub-saharan africa in the context of this paper, visibility refers to the extent to which researchers can identify, access and use a scholarly publication. according to miguel et al. (2011), many factors determine the visibility of a research publication. they point out, however, that of these factors, the channel in which the research is published plays a pivotal role. thus, ale ibrahim et al. (2014) opine that high impact channels of research publication expose research published therein more and gives them a greater possibility to be identified, accessed and cited. similarly, miguel et al. (2011) assert that since the full-text copies of research published through open access platforms are readily downloadable, such publications are more visible than those published in subscription-based channels. therefore, impediments to their usability are minimal. thus, they are likely to attract more use and citations than their counterparts published in subscription-based channels. an analysis of the channels of publication of the works revealed that a large majority was published in subscription channels such as journals and books. only 71 out of the 184 works were published in openly accessible channels such as institutional repositories and library web sites. recognising the fact that many researchers in sub-saharan africa, and other developing economies, have limited access to subscriptionbased publication channels, they are less likely to access or use the publications. this limitation on the visibility of the publications is one of the factors contributing to the low citation of the publications. the low visibility of research products from sub-saharan africa can also be attributed to the fact that most researchers do not promote their publications. in the age of social media and other open platforms, most researchers have embraced social networking sites to promote and enhance the reach of their publications. in africa, however, neylon et al. (2014) observed less involvement of researchers in social networking sites. they argue that by neglecting social media, researchers in africa lose the opportunity to market their research output directly to their potential users. 4.4 authorship of research publications on rdm in sub-saharan africa most (118) of the works were published by more than one author. this demonstrates a high level of collaboration in terms of co-authorship of the publications. available evidence argues that co-authored research is typically of a higher quality than singly-authored works (hilmer & hilmer, 2005; hart, 2007; andrade et al., 2010; bidault & hildebrand, 2014). according to franceschet and costantini (2010), it is not easy to avoid collaboration and co-authorship in the current scholarly communication landscape. besides, bidault and hildebrand (2014) explain that co-authorship is a strategic instrument for mentoring junior academics by experienced researchers. the authors who co-authored more than one publication were j van wyk (6), h pienaar (6), d hoffmeister (4), m van deventer (4), c curdt (4), wd chawinga (4), bk avuglah (3), s kralisch (3), f zander (3), l lotter (3), bv cendon (2), d nicholson (2), er chiware (2), fg almeida (2), fj abduldayan (2), g coetzer (2), gl coetzer (2); j davidson (2), l horton (2), l jacobs (2), mb macanda (2), mw rammutloa (2), n nhendodzashe (2), p zibani (2), r botha (2), s zinn (3), sm mutula (2), t kramm (2), u lang (2), and v van den eynden (2). figure 4 shows the social networks created through co-authorship of the publications on rdm in subsaharan africa. it shows a number of loosely connected and less dense social networks around h pienaar, https://doi.org/10.29173/iq996 9/14 kwanya, tom (2021) publishing trends on research data management in sub-saharan africa: a bibliometric analysis, iassist quarterly 45(3-4), pp. 1-14. doi: https://doi.org/10.29173/iq996 j van wyk, c curdt and f zander. this finding indicates that the authors have collaborated less with authors with similar interests. figure 4: social networks of co-authors 5. conclusion the value of research data to scientific innovation and development has grown in the recent past. consequently, there is an over-production of research data. researchers can benefit greatly by accessing and reusing existing data. the biggest hindrance to the effective use of research data is inadequate sharing of the data by their creators or collectors. these challenges can be overcome by scientists embracing effective research data management. researchers from most regions have embraced the concept of research data management. however, sub-saharan africa is lagging behind the rest of the world in research data management. research on research data management can contribute greatly to the understanding and adoption of the practice. evidence from this study reveals that the quantity, quality, visibility and scholarly collaboration on research data management in sub-saharan africa is low. therefore, sub-saharan perspectives on research data management are less voiced. 6. recommendations based on the findings of the study, the author recommends as follows: 1. there is need to sensitise researchers in sub-saharan africa about the concept of research data management as a means of stimulating them to adopt the concept. this can be done through appropriate training and publicity programmes. https://doi.org/10.29173/iq996 10/14 kwanya, tom (2021) publishing trends on research data management in sub-saharan africa: a bibliometric analysis, iassist quarterly 45(3-4), pp. 1-14. doi: https://doi.org/10.29173/iq996 2. universities, research institutions and research funders in sub-saharan countries should develop policy guidelines which promote and perpetuate research data sharing and preservation. the policies should cover important elements of research data management. 3. researchers are encouraged to collaborate with each other and create social networks which can be used to strengthen scholarly work. increased collaboration will help the scholars to overcome some of the challenges such as lack of adequate resources. this would be achieved through pooling of resources. 4. researchers are encouraged to publish their work in less restricted channels such as open access platforms and institutional repositories. this will likely increase the reach of the publications thereby enhancing their use and propagation. 5. research institutions are encouraged to develop adequate infrastructure for research data management. the infrastructure may include ict networks, laboratories and libraries. there can be no meaningful research without the requisite infrastructure. https://doi.org/10.29173/iq996 11/14 kwanya, tom (2021) publishing trends on research data management in sub-saharan africa: a bibliometric analysis, iassist quarterly 45(3-4), pp. 1-14. doi: https://doi.org/10.29173/iq996 table 1: citations analysis showing the articles with at least 10 citations cites authors title year 82 rj calantone, sk vickery introduction to the special topic forum: using archival and secondary data sources in supply chain management research 2010 61 lm delserone at the watershed: preparing for research data management and stewardship at the university of minnesota libraries 2008 55 t tuti, m bitok, c paton, b makone innovating to enhance clinical data management using non-commercial and open source solutions across a multi-center network supporting inpatient pediatric care and research in kenya 2016 52 de winickoff, k saha, gd graff opening stem cell research and development: a policy proposal for the management of data, intellectual property, and ethics 2009 48 ma pirog data will drive innovation in public policy and management research in the next decade 2014 48 hh tsai knowledge management vs. data mining: research trend, forecast and citation approach 2013 40 e chiware, z mathe academic libraries' role in research data management services: a south african perspective 2015 36 hh tsai research trends analysis by comparing data mining and customer relationship management through bibliometric methodology 2011 32 m kahn, r higgs, j davidson, s jones research data management in south africa: how we shape up 2014 32 qj groom, p desmet, s vanderhoeven, t adriaens the importance of open data for invasive alien species research, policy and management 2015 22 ta reynolds, m bisanzo, d dworkis research priorities for data collection and management within global acute and emergency care systems 2013 19 c willmes, d kürner, g bareth building research data management infrastructure using open source software 2014 18 m van deventer, h pienaar research data management in a developing country: a personal journey 2015 18 f cloete data analysis in qualitative public administration and management research 2007 16 au aydinoglu, g dogan, z taskin research data management in turkey: perceptions and practices 2017 14 am elsayed, ei saleh research data management and sharing among researchers in arab universities: an exploratory study 2018 14 s friedhoff, c meier zu verl, c pietsch, c meyer social research data: documentation, management, and technical implementation within the sfb 882 2013 14 ds bullock, m boerngen, h tao, b maxwell the data‐intensive farm management project: changing agronomic research through on‐ farm precision experimentation 2019 13 fds choi international data sources for empirical research in financial management 1988 12 o maduka, g akpan, s maleghemi using android and open data kit technology in data management for research in resourcelimited settings in the niger delta region of nigeria: cross 2017 11 er chiware, da becker research data management services in southern africa: a readiness survey of academic and research libraries 2018 10 c neylon building a culture of data sharing: policy design and implementation for research data management in development research 2017 10 f huettmann on the relevance and moral impediment of digital data management, data sharing, and public open access and open source code in (tropical) research: the rio convention revisited towards mega science and best professional research practice 2015 10 ri leihy, ga duffy, e nortje, sl chown high resolution temperature data for ecological research and management on the southern ocean islands 2018 https://doi.org/10.29173/iq996 12/14 kwanya, tom (2021) publishing trends on research data management in sub-saharan africa: a bibliometric analysis, iassist quarterly 45(3-4), pp. 1-14. doi: https://doi.org/10.29173/iq996 references adika, f. o., & kwanya, t. (2020). research data management literacy amongst lecturers at strathmore university, kenya. library management, 41(6/7), 447-466. ajibade, p., & mutula, s. m. (2020). big data research outputs in the library and information science: south african's contribution using bibliometric study of knowledge production. african journal of library, archives & information science, 30(1), 49-60. ale ebrahim, n., salehi, h., embi, m. a., habibi, f., gholizadeh, h., & motahar, s. m. (2014). visibility and citation impact. international education studies, 7(4), 120-125. anane-sarpong, e., wangmo, t., ward, c.l., sankoh, o., tanner, m. & elger, b.s., (2018). you cannot collect data using your own resources and put it on open access: perspectives from africa about public health data sharing. developing world bioethics, 18, 394–405. andrade, h. b., de los reyes lopez, e., & martín, t. b. (2009). dimensions of scientific collaboration and its contribution to the academic research groups' scientific quality. research evaluation, 18(4), 301-311. bidault, f., & hildebrand, t. (2014). the distribution of partnership returns: evidence from coauthorships in economics journals. research policy, 43(6), 1002-1013. briney, k. (2015). data management for researchers: organize, maintain and share your data for research success. exeter: pelagic publishing. chawinga, w.d. & zinn, s. (2019). global perspectives of research data sharing: a systematic literature review. library & information science research, 41, 109-122. dayu, j. (2012) bibliometric analysis tutorial. retrieved october 18, 2020 from https://www.slideshare.net/dayu_jin/bibliometric-analysis-tutorial-by-dayu-jin denny, s.g., silaigwana, b., wassenaar, d., bull, s. & parker, m. (2015). developing ethical practices for public health research data sharing in south africa: the views and experiences from a diverse sample of research stakeholders. journal of empirical research on human research ethics, 10, 290–301. franceschet, m., & costantini, a. (2010). the effect of scholar collaboration on impact and quality of academic papers. journal of informetrics, 4(4), 540-553. hart, r. l. (2007). collaboration and article quality in the literature of academic librarianship. the journal of academic librarianship, 33(2), 190-195. hilmer, c. e., & hilmer, m. j. (2005). how do journal quality, co-authorship, and author order affect agricultural economists' salaries?. american journal of agricultural economics, 87(2), 509-523. jahnke, l.m. & asher, a. (2012). the problem of data: data management and curation practices among university researchers. the problem of data, 3-31. kahn, m., higgs, r., davidson, j. & jones, s. (2014). research data management in south africa: how we shape up. australian academic & research libraries, 45, 296-308. kibe, l., kwanya, t. & owano, a. (2020). relationship between big data analytics and organisational performance of the technical university of kenya and strathmore university in kenya. global knowledge memory and communication. https://doi.org/10.1108/gkmc-04-2019-0052 koopman, m. m. (2015). data archiving, management initiatives and expertise in the biological sciences department, university of cape town (master's thesis, university of cape town). kuo, y. h., & kusiak, a. (2019). from data to big data in production research: the past and future trends. international journal of production research, 57(15-16), 4828-4853. kwanya, t., kogos, a. c., kibe, l. w., ogolla, e. o., & onsare, c. (2021). cyber-bullying research in kenya: a meta-analysis. global knowledge, memory and communication. https://doi.org/10.1108/gkmc-08-2020-0124 https://doi.org/10.29173/iq996 https://www.slideshare.net/dayu_jin/bibliometric-analysis-tutorial-by-dayu-jin https://doi.org/10.1108/gkmc-04-2019-0052 https://doi.org/10.1108/gkmc-08-2020-0124 13/14 kwanya, tom (2021) publishing trends on research data management in sub-saharan africa: a bibliometric analysis, iassist quarterly 45(3-4), pp. 1-14. doi: https://doi.org/10.29173/iq996 kwanya, t., stilwell, c., & underwood, p. (2014). mainstreaming grey literature in research library collections in kenya. libri, 64(2), 134-143. lucas-dominguez, r., alonso-arroyo, a., vidal-infer, a., & aleixandre-benavent, r. (2021). the sharing of research data facing the covid-19 pandemic. scientometrics, 126(6), 4975-4990. miguel, s., chinchilla‐rodriguez, z., & de moya‐anegón, f. (2011). open access and scopus: a new approach to scientific visibility from the standpoint of access. journal of the american society for information science and technology, 62(6), 1130-1145. nature. (2014). scientific data: data policies. retrieved 24 may, 2020 from http://www.nature.com/sdata/data-policies neubert, c., & trischler, r. (2021). “pocketing” research data? ethnographic data production as material theorizing. journal of contemporary ethnography, 50(1), 99-119. neuhaus, c., & daniel, h. d. (2008). data sources for performing citation analysis: an overview. journal of documentation, 64(2), 193-210. neylon, c., willmers, m. & king, t. 2014. illustrating impact: applying altmetrics to southern african research. https://open.uct.ac.za/handle/11427/2316 ng’eno, e., & mutula, s. (2018). research data management (rdm) in agricultural research institutes: a literature review. inkanyiso: journal of humanities and social sciences, 10(1), 28-50. norton, m. j. (2001). introductory concepts in information science. information today. patterton, l., bothma, t.j., & van deventer, m.j. (2018). from planning to practice: an action plan for the implementation of research data management services in resource-constrained institutions. south african journal of libraries and information science, 84(2), 14–26. peng, x., & luo, z. (2021). a review of q-rung orthopair fuzzy information: bibliometrics and future directions. artificial intelligence review, 1-70. pisani, e., aaby, p., breugelmans, j.g., carr, d., groves, t., helinski, m. & mboup, s. (2016). beyond open data: realising the health benefits of sharing data. bmj 355 i5295. ray, j. m. (ed.). (2014). research data management: practical strategies for information professionals. purdue university press. schöpfel, j., ferrant, c., andré, f. & fabre, r. (2018). research data management in the french national research center (cnrs). data technologies and applications, 52(2), 248-265. siyanbola, w., adeyeye, a., olaopa, o., & hassan, o. (2016). science, technology and innovation indicators in policy-making: the nigerian experience. palgrave communications, 2(1), 1-9. stewart, r. 2015. a theory of change for capacity building for the use of research evidence by decision makers in southern africa. evidence & policy, 11(1): 547-57. tijssen, r. j. w. 2007. africa’s contribution to the worldwide research literature: new analytical perspectives, trends, and performance indicators. scientometrics, 71(2): 303-327. townsend, b. (2021). the lawful sharing of health research data in south africa and beyond. information & communications technology law, 1-18. van noorden, r. (2014). confusion over open-data rules. nature, 515, 478. vlahou, a., hallinan, d., apweiler, r., argiles, a., beige, j., benigni, a., ... & vanholder, r. (2021). data sharing under the general data protection regulation: time to harmonize law and research ethics?. hypertension, 77(4), 1029-1035. whyte, a., & tedds, j. (2011). making the case for research data management in dcc briefing papers. edinburgh: digital curation centre. wormell, i. (2001, september). informetrics for informed decision making. in a paper presented at swedish-lithuanian seminar on information management research issues on (pp. 21-22). https://doi.org/10.29173/iq996 http://www.nature.com/sdata/data-policies https://open.uct.ac.za/handle/11427/2316 14/14 kwanya, tom (2021) publishing trends on research data management in sub-saharan africa: a bibliometric analysis, iassist quarterly 45(3-4), pp. 1-14. doi: https://doi.org/10.29173/iq996 endnotes 1 tom kwanya is professor of knowledge management in the department of information and knowledge management at the technical university of kenya. he can be reached on tkwanya@tukenya.ac.ke. 2 african socialist philosophy of taking care of each other. https://doi.org/10.29173/iq996 mailto:tkwanya@tukenya.ac.ke vol282-3.indd 60 iassist quarterly summer/fall 2004 submissions due january 10, 2006 announcement of winner march 1, 2006 award: $250 us and one year membership in iassist http://www.iassistdata.org content focus: education. as an organization, iassist has a history of working to educate its members about matters of common professional interest to the social science data community. traditionally, this education has taken the form of professional development opportunities available in member-initiated and member-taught workshops at the annual iassist conference. yet as the world of social science data grows increasingly complex, staying abreast of new developments in the profession is likely to present an ever increasing challenge for iassist members. as a result, the education committee of iassist is receptive to recommendations for employing new instructional methods and technologies as we strive to meet iassist’s educational mission. at its annual conference in may, 2004, the iassist membership approved a 5-year strategic plan that focuses upon three strategic directions: education, outreach, and advocacy. this paper competition has been established as a means of exploring, articulating, and documenting topical issues related to education. for further information about the strategic plan, see: http://www.iassistdata.org/membership/plan_june2004.pdf iassist seeks papers that address one or more of the issues, principles, and strategies for engagement in the following subject areas: 1. iassist-related educational initiatives 2. professional development and educational opportunities of interest to iassist members 3. educational outreach to research communities, including and beyond the social sciences, to promote data preservation and access. papers should include recommendations for action or suggestions of specifi c projects that can be undertaken (or are underway) to further the education goals of iassist. prospective submitters may wish to review the discussion on the iassist blog about the educational issues arising under the topic of the accidental data librarian (available at http://iassistblog.org/?cat=5). call for papers: 2006 iassist strategic plan publication award (competition is not limited to current iassist members) iassist quarterly summer/fall 2004 61 2006 iassist strategic plan publication award criteria for evaluation: --relevance of the paper to one or all of the themes in the iassist strategic plan (specifi cally strategic direction i: improve and expand the educational component of iassist both internally and externally. --inclusion of specifi c suggestions for action or specifi c projects that will further the education of iassist members or otherwise encourage progress in support of the iassist strategic planiassist members or otherwise encourage progress in support of the iassist strategic planiassist --potential in building a base for future iassist activity --quality of writing -bibliographic content including references to related materials --clarity in presenting issues and viewpoints as outlined above competition details: all papers are to be submitted in english. the winning paper will be announced on the iassist list-serve on or around march 1, 2006. in addition to being designated as the winning paper in the iassist quarterly, the author of the winning paper will receive the monetary award and a one-year membership in iassist and be recognized at the iassist conference in ann arbor. all other submissions meeting the criteria for evaluation will be published in the iassist quarterly (iq) (online and print). papers that are submitted to this strategic plan publication award competition may also be submitted for inclusion in the iassist conference in ann arbor in may 2006. papers must be a minimum of 5 pages in length, including bibliography and graphics as appropriate. we strongly prefer that for publication purposes all documents be submitted in word format and each graphic be submitted as a separate fi le in one of the following formats: .gif .jpg .tif .bmp .png papers must not have copyright limitations; iassist quarterly (iq) rights will apply upon publication. all papers must be submitted by january 10, 2006 to the competition web host: david sheaves <sheaves@vance.irss.unc.edu> questions (not papers, please) may be sent to: <iassist-reviews@mailman.srv.ualberta.ca> competition is not limited to current iassist members. the review committee for the iassist strategic plan publication award will be announced on the iassist website www.iassistdata.org and on the iassist list serve. 1/19 majewicz, karen; martindale, jaime; kernik, melinda (2022). open geospatial data: a comparison of data cultures in local government, iassist quarterly 46(1), pp. 1-19. doi: https://doi.org/10.29173/iq1013 open geospatial data: a comparison of data cultures in local government karen majewicz1, jaime martindale2, melinda kernik3 abstract public geospatial data (geodata) is created at all levels of government, including federal, state, and local (county and municipal). local governments, in particular, are critical sources of geodata because they produce foundational datasets, such as parcels, road centerlines, address points, land use, and elevation. these datasets are sought after by other public agencies for aggregation into state and national frameworks, by researchers for analysis, and by cartographers to serve as base map layers. despite the importance of this data, policies about whether it is free and open to the public vary from place to place. as a result, some regions offer hundreds of free and open datasets to the public, while their neighbors may have zero, preferring to restrict them due to privacy, economic, or legal concerns. minnesota relies on an approach that allows counties to choose for themselves if their geodata is free and open. by contrast, its neighboring state of wisconsin has passed legislation requiring that specific foundational geospatial datasets created by counties must be freely available to the public. this paper compares the implications and outcomes of these diverging data cultures. keywords geospatial data, geodata, gis, open data, local government, minnesota, wisconsin acknowledgment thank you to yijing zhou for research assistance and cartography.4 1. introduction two of the authors, both of whom work at the university of minnesota’s john r. borchert map library, heard the following exasperated question at a recent geographic information science (gis) professional conference in minnesota: ‘why am i able to download a free and open statewide parcel dataset that includes every county in wisconsin, but not in minnesota?’ there was a noticeable whiff of envy in the room. ‘free and open data’ across minnesota’s state and local government has been the stated top priority for the assembled community for the past four years, while wisconsin cleared the hurdle seemingly overnight. what differences in the communities of practice in wisconsin and minnesota led to this uneven landscape? to answer this question, we partnered with our colleague at the university of wisconsin-madison’s robinson map library to delve into each state’s gis history, programs, organizations, and legislation to construct a comparison of our respective open geospatial data (geodata) landscapes. our case studies revealed that (1) legislation, (2) funding models, (3) workflows for contributing to the state’s primary geodata platform, and (4) the involvement of libraries are key differences in the divergent outcomes. 2. overview of open geodata 2.1 qualifications for the purposes of this article, geodata encompasses vector (points, lines, polygons), raster (i.e., lidar dems, orthoimagery, landcover), and database files that contain spatial information. to qualify as ‘open geodata,’ we have identified three criteria. first, the data should have an open license or open status. this means that users cannot be required to sign a license, sharing the data cannot be restricted, and the data https://doi.org/do.be/doo https://doi.org/10.29173/iq1013 2/19 majewicz, karen; martindale, jaime; kernik, melinda (2022). open geospatial data: a comparison of data cultures in local government, iassist quarterly 46(1), pp. 1-19. doi: https://doi.org/10.29173/iq1013 does not contain confidential or private information. second, the data should be accessible for free, without even minimal charges. third, the data should be downloadable as discrete layers and not just viewable from inside an online web map or database application. for context, in figure 1, we have compared our criteria against two other models: the open knowledge foundation’s (okf) ‘open definition 2.1 of open works’ and daniel sui’s article, ‘opportunities and impediments for open gis’ (open knowledge foundation, n.d.; sui, 2014). we chose these models because the okf criteria are widely cited for general open data, while sui’s is specific to geodata. required criteria our model okf open definition 2.1 sui (2014) open license or status x x x free x [reasonable fee] x downloadable x [recommended] open, non-proprietary format x usable: features quality data and metadata x figure 1. a comparison of open data qualification criteria models some aspects of the compared models are more lenient than ours. the okf model allows data providers to charge a fee. although they do mitigate the severity of this with the phrase ‘no more than a reasonable one-time reproduction cost,’ this practice has the potential to function as a kind of loophole. for example, our minnesota case study will observe that, although all government data is defined as ‘public’ by state law, counties and municipalities are still allowed to charge fees to cover the cost of assembling and sharing their geodata. consequently, this practice has effectively prevented geodata from being findable and accessible across a large part of the state. another tolerance that we disagree with is that neither the sui nor the okf models specify that datasets must be downloadable. unfortunately, this is a common constraint on geodata due to the prevalence of web maps, which enable users to view and interact with a preselected set of layers, but typically prohibit dataset downloads. maas (2019) describes this kind of data as ‘captive,’ meaning it cannot be analyzed or mapped outside of the application’s restricted scope. despite this limitation, government agencies have been publishing web maps ever since the national atlas of canada went online in 1994 (kramers, 2008), and there is a perception that they fulfill the ethos of open geodata. we concede that web maps have played a role in increasing public interest in the utility and value of geodata, as users can consume geospatial information without needing to be well-versed in the complexities of gis formats, structures, or technology. however, there is now a wide array of new and emerging user-friendly mapping tools that https://doi.org/do.be/doo https://doi.org/10.29173/iq1013 3/19 majewicz, karen; martindale, jaime; kernik, melinda (2022). open geospatial data: a comparison of data cultures in local government, iassist quarterly 46(1), pp. 1-19. doi: https://doi.org/10.29173/iq1013 have reduced barriers to the public’s ability to collect, merge, and analyze geospatial content from different sources. as a result, a closed web map excessively limits what users can do with its data. other aspects of the compared models feature stricter criteria than ours, and we acknowledge that these represent worthy goals. however, the current nature of gis technology would make those criteria challenging to meet. for example, okf specifies that to be considered “open,” datasets should be available in an open-source format. we agree that open-source formats are ideal but contend that in practice the situation is complicated. some of the most commonly used geospatial file formats are proprietary but still can be read and edited within open-source software. for example, shapefiles are “proprietary but open,” with a technical specification published in 1998 (library of congress, 2021). file geodatabases are also proprietary but can be used within open-source software using communitydeveloped plug-ins. because of this, we felt requiring open file formats would be too restrictive and chose not to include it as one of the criteria in our review. a second condition that we have chosen to permit is data with minimal metadata. sui proposes that open data must be well described and understandable. when assessing whether a dataset’s documentation is sufficient, it is relevant to note that many geodata delivery platforms do not support an intuitive metadata workflow. publishing to an online portal often involves an automated transformation from a comprehensive metadata standard to a reduced set of core fields, thereby limiting the granularity and possibly corrupting the integrity of the original metadata. we see this as essentially a technological issue that should not disqualify items from being considered open. 2.2 sources public geodata is created at all levels of government. federal agencies issue the most public geodata. some of the most well-known examples are satellite imagery from the united states geological survey, real-time weather services from the national oceanic and atmospheric division, and demographic information from the census bureau. most u.s. states maintain foundational geospatial layers, such as transportation networks, elevation, hydrography, aerial imagery, and cadastral information. regional organizations may produce unique sets of resources, such as watershed district boundaries or regional transit systems. counties are typically responsible for maintaining records of tax parcels, address points, and roads. municipalities will generally provide important society data, for instance, city services, neighborhood boundaries, local transit, and community centers. open geodata is provided through a variety of platforms and technologies. the simplest method, typically utilized by smaller organizations, is to publish datasets as direct downloads hosted on an ftp server or static web page. larger cities or regions may choose to use dedicated data search portals, such as the open-source comprehensive knowledge archive network (ckan)5 or the proprietary socrata.6 these applications are designed for general data and can incorporate tabular data, databases, and spatial formats. organizations with a sizable amount of geodata may opt for a dedicated geospatial portal application, such as arcgis hub7 or geoblacklight.8 these specialized applications feature integrated map searches and previews of geospatial web services, which allow users to examine and query the data from within the portal interface without necessitating downloading the data and opening it in a desktop gis application. 2.3 availability and barriers the public availability of open geodata depends upon the administration level that provides it. in the united states, all federally produced geodata (except sensitive data restricted for privacy or security) has been open since 2009 (blatt, 2016). however, policies about the openness of state and local government https://doi.org/do.be/doo https://doi.org/10.29173/iq1013 https://ckan.org/ https://www.tylertech.com/products/socrata https://hub.arcgis.com/ https://geoblacklight.org/ 4/19 majewicz, karen; martindale, jaime; kernik, melinda (2022). open geospatial data: a comparison of data cultures in local government, iassist quarterly 46(1), pp. 1-19. doi: https://doi.org/10.29173/iq1013 data vary from place to place. many states have open data initiatives, and the majority maintain an online clearinghouse that provides geodata produced by state agencies (national states geographic information council, 2019). however, most counties and municipalities are not required to comply with either federal rules or state initiatives for open data. as a result, some regions offer hundreds of free open data layers to the public, while their neighbors may have zero. even when the data is legally declared ‘public,’ a common scenario is that it is not free or accessible online. in those cases, a user must place a data request with the organization and pay a fee. a gis professional then manually prepares and shares the datasets via a hard drive or a file transfer. many studies have investigated why governments may choose not to share their data online freely. johnson et al. (2017) questioned the purported benefits of open data and argued that it is unduly costly. for instance, gis staff would need to implement technology platforms, and they could be subject to an increased workload to maintain and regularly update data. on the other hand, some research disproves the idea that governments would lose revenue. joffe (2003) and maas (2013) contended that embracing open data saves organizations money in the long run by reducing staff workload, as they do not need to fill as many specialized data requests. tombs (2005) described the inclination to keep geodata restricted for public safety and security, but he argued that this practice conflicts with the citizenry’s right to public data access and free speech. overall, the reasons offered against open geodata can be characterized as apprehension about the potential for negative consequences. wirtz et al. (2016) identified general risk aversion among public servants as the main barrier. this assessment aligns with a 2016 survey of gis staff in 59 minnesota counties that revealed four top issues of concern (minnesota geospatial advisory council outreach committee, 2016): 1. the potential loss of revenue from the sale of geospatial data 2. legal liability 3. ‘bad actors’ misusing the data 4. privacy and security concerns when local governments have the authority to choose whether or not to make their data free and open, many will err on the side of caution. unfortunately, the resulting lack of contiguous availability thwarts worthy data aggregation efforts and results in increased costs as organizations that need statewide or regional data must either purchase it or recreate it. 2.4 how local open geodata supports the national landscape geodata produced by local governments may be foremost intended for use within that administration’s local domain. however, many aspects of our environment (e.g., climate and pollution) or infrastructure (e.g., transportation networks) do not terminate at administrative borders. county data layers can be collected, stitched together into statewide layers, and subsequently combined for national frameworks. when geographically adjacent datasets are merged, their value is enhanced by serving expanded areas to inform higher decision-making organizations. this concept has been promoted by national organizations for several projects over the years, including the national spatial data infrastructure (nsdi), the national parcel database, and next generation 9-1-1. in 1994, the clinton administration tasked the federal geographic data committee (fgdc) with the advancement of the national spatial data infrastructure (nsdi) (federal geographic data committee, 1994). the main outcome of the nsdi was to be a set of ‘framework’ data layers that would form the core https://doi.org/do.be/doo https://doi.org/10.29173/iq1013 5/19 majewicz, karen; martindale, jaime; kernik, melinda (2022). open geospatial data: a comparison of data cultures in local government, iassist quarterly 46(1), pp. 1-19. doi: https://doi.org/10.29173/iq1013 of the infrastructure (federal geographic data committee, 1997). tulloch and fuld (2001) analyzed a late 1990s survey of county-level data producers that revealed several challenges to this project. of the respondents, 24% did not create any of the framework layers, and even fewer (approximately 10%) maintained any metadata for the layers. furthermore, data sharing policies were ambiguous or absent. harvey and tulloch (2006) followed up five years later to report that the nsdi had improved the standardization and sharing of federally produced data. however, it was still hampered by participation from local governments. the scholarship on local participation towards the nsdi has fallen off in recent years, but the program received a symbolic boost with the passage of the geospatial data act in 2018. this act was intended to facilitate the nsdi but unfortunately provided no avenues for funding gis departments. interviews with the national states geographic information council leaders indicate that state gis infrastructures are simply not coordinated enough to participate in the nsdi and likely will never be unless federal funding is provided (wood, 2020). an essential part of the nsdi would be a national layer of parcels (sometimes referred to as ‘tax parcels’). parcels are land records that define ownership and boundaries, and they are utilized for many purposes, including land use studies, zoning, taxes, and base maps. except for federal and state-owned lands, individual counties are responsible for creating and maintaining all parcels. having each county create these records independently has led to wide variations between the formats, attributes, and quality. although merging these records would be a massive undertaking, a standardized dataset of all the parcels in the country would have many applications, from facilitating land transfers, to assessing public health needs, to coordinating disaster relief. to get a sense of the difficulty of aggregating all the parcels in the country, consider the relative lack of progress despite long-standing promotion efforts. for example, the national research council issued a guidebook in 1980, in which they provided a template for land records to be digitized, standardized, and combined (national research council, 1980). this process is known as land records modernization. the council followed up twenty-seven years later with a report that lamented how much more work was still needed to create a national layer (national research council, 2007). this goal was reinvigorated in 2010 when the us department of housing and urban development (hud) began a national parcel database project. hud spent a few years evaluating parcel records from over 100 counties. they discovered that the datasets did not have comprehensive metadata and the data models were so incongruous that standardizing them would be complicated and expensive. they further noted that the scope of merely pursuing data-sharing agreements with the counties was daunting. (u.s. department of housing and urban development, 2013). this project’s current status is unclear, but it appears to be no longer active (hud librarian, personal communication, march 11, 2020). a more recent data aggregation effort that may have a higher chance of success is the next generation 91-1 project, a national initiative to improve emergency services to rural areas by creating a complete national gis framework of road centerlines, address points, and administrative boundaries (national emergency number association, 2020). without accurate gis data in rural areas, emergency responders are unable to navigate to their destinations efficiently. unlike other aggregation projects, counties may be more motivated to participate in this initiative, as they will be the direct recipients of benefits that improve the safety of their residents. however, it suffers from the same challenges of coordinating and providing local governments with the resources needed to collect, standardize, and share their data (kemp, 2017). https://doi.org/do.be/doo https://doi.org/10.29173/iq1013 https://www.zotero.org/google-docs/?5ubeey https://www.zotero.org/google-docs/?f9epkt 6/19 majewicz, karen; martindale, jaime; kernik, melinda (2022). open geospatial data: a comparison of data cultures in local government, iassist quarterly 46(1), pp. 1-19. doi: https://doi.org/10.29173/iq1013 3. case studies the following case studies show that the two neighboring states of minnesota and wisconsin share several similarities in their open geodata ecosystems. they both have a long history of supporting gis technology, backed by prominent universities with nationally renowned geography departments and map libraries. they both also have well-supported geodata platforms that can incorporate resources from state, county, and city agencies, as well as nonprofit, business, and educational organizations. however, their stories diverge when it comes to their efforts around local geodata aggregation and availability. 3.1 case study i: minnesota the geospatial community in minnesota has long supported a climate of innovativeness and collaboration that has resulted in a well-established spatial data infrastructure, particularly for state agencies and the twin cities metropolitan region. minnesota was an early hotbed for gis development and coordinated data management endeavors. the minnesota land management information system (mlmis) at the university of minnesota began in 1967 and was one of the first geographic information systems in the world. its mission was to inform land use decisions by maintaining a framework of 19 data layers that could be digitally combined, analyzed, and mapped (university of minnesota, 1976). mlmis is distinct from other pioneering gis projects in that it continued for over a decade and is a direct ancestor to the official state geospatial agency in minnesota today. in the late 1970s, mlmis was transferred to the minnesota state planning agency as the land management information center (lmic). lmic represented an evolution from a research project into a government-run data services center (warnecke, 1992), and it operated for over 30 years. lmic also hosted a search portal, the minnesota geospatial data clearinghouse, that federated gis data from multiple sources. during lmic’s time, other organizations continued developing their own gis programs. several state agencies, such as the pollution control agency and the department of transportation, developed in-house strategies for creating and managing their own gis data. the minnesota department of natural resources (dnr) even built its own open data clearinghouse, the dnr data deli.9 metrogis was established in 1996 as a regional initiative serving the twin cities metropolitan area, and it also maintained its own open data portal for many years, the metrogis datafinder.10 in order to facilitate sharing these collections of open data, the state adopted a custom metadata guideline in 1998. the minnesota geospatial metadata guidelines (mgmg) is a streamlined version of the federal geographic data committee’s content standard for digital geospatial metadata (fgdc). endorsing this guideline was a progressive and forward-thinking step, as most states do not have an official geospatial metadata profile to this day. (minnesota governor’s council on geographic information, 1998). as geospatial technology was flourishing in minnesota, gis professionals began to become concerned about a lack of centralization. lmic was constrained to being an on-demand service organization and, although it acted as the ‘unofficial statewide geospatial coordinator,’ it did not have the power to implement a statewide infrastructure (arbeit et al., 2004; terner et al., 2009). the governor’s council on geographic information was established in 1991 to fill this gap by advising state agencies on gis activities and data sharing. one of the council’s final initiatives was to create a plan for a more authoritative state geospatial agency run by a geospatial information officer (minnesota governor’s council on geographic information, 2009). lmic was then reorganized as the minnesota geospatial information office (mngeo), https://doi.org/do.be/doo https://doi.org/10.29173/iq1013 https://web.archive.org/web/19991013033555/http:/deli.dnr.state.mn.us/ https://web.archive.org/web/19981203091211/http:/www.datafinder.org/ 7/19 majewicz, karen; martindale, jaime; kernik, melinda (2022). open geospatial data: a comparison of data cultures in local government, iassist quarterly 46(1), pp. 1-19. doi: https://doi.org/10.29173/iq1013 which operates today as the official state gis coordinating agency. this action also dissolved the governor’s council on geographic information, which was replaced by the minnesota geospatial advisory council (gac). minnesota has legislation defining public data as open, but it does not require that it must be free. the minnesota data practices act, enacted in 1974, designates all data produced by government entities as open, with the exception of confidential or otherwise non-public information. (minnesota government data practices act, 1974). this act was established before the age of digital data and was designed for people to visit a local government record keeper and inspect physical sheets of data at no charge (maas, 2019). since 1974, the act has been updated and amended many times, often to address privacy issues and to clarify what types of data should be kept confidential. a notable update occurred in 1990, granting counties and municipalities the right to 'charge a reasonable fee for the information in addition to the costs of making and certifying the copies' for digital data (minnesota government data practices act, 1990). this update resulted from the enormous expense counties and cities were shouldering to implement computer systems and technicians to collect, transform, and deliver digital data. maas (2019) notes that the eligible data is still technically public, but it is not guaranteed to be free. this distinction is evidenced by a lack of statewide foundational datasets aggregated from county layers, such as parcels or address points. mngeo does collect these layers from every county for internal use for projects like next generation 9-1-1. while it can share the datasets with other government agencies, it does not make them freely open to the general public because of licensing agreements with the counties. the gac is the state’s most prominent champion of open geodata. this is evidenced by their annual list of top priorities, which is generated by weighing a variety of factors, including community votes and the likelihood of success. the ‘promotion of free and open data’ has been at the top of this annual list for each of the past four years (2018-2021). this priority has also shown itself in the gac’s committees and workgroups: the gac outreach committee has made free and open data the main focus of their recent activities, and a newly formed workgroup is exploring strategies for increasing the number of counties with free and open parcel data. in 2021, the gac evolved further on this issue and upgraded the promotion of free and open data from a ‘priority’ to a ‘guiding principle’ that all committees should incorporate into their work. one of the most successful manifestations of the gac’s open data advocacy has been the development of the minnesota geospatial commons11 (‘commons’), a collectively managed state platform for open geodata that, as of february 2021, contains 900 resources contributed by 45 different organizations. when the commons went online in 2015, it replaced multiple state and regional portals, including the aforementioned minnesota geospatial data clearinghouse, dnr data deli, and metrogis datafinder. the commons accepts data from any public organization, but the primary contributors thus far are state departments and agencies. resources in the commons are well-documented because they must be described with the state metadata guidelines, mgmg. the commons uses a self-service model whereby each contributor has full management over their resources. the state’s advocacy of open government has not brought about a culture of open data in all parts of the state. as of the most recent update in september 2021, only 45 out of 87 minnesota counties offer downloadable geodata for free (minnesota geospatial information office, 2021). furthermore, only ten of these counties have taken advantage of the commons as a platform to deliver their resources to a broader audience. https://doi.org/do.be/doo https://doi.org/10.29173/iq1013 https://gisdata.mn.gov/ 8/19 majewicz, karen; martindale, jaime; kernik, melinda (2022). open geospatial data: a comparison of data cultures in local government, iassist quarterly 46(1), pp. 1-19. doi: https://doi.org/10.29173/iq1013 although the availability of county-level open geodata across the entire state paints a patchy picture, the situation in the twin cities metropolitan area is quite different. under the coordination of metrogis, each of the seven counties in the metropolitan area has declared open geodata policies. they uniformly share datasets for parcels, road centerlines, address points, parks, and trails & bikeways. the success of metrogis can be attributed to the region’s history of cooperation and regional policymaking. although participation in metrogis is voluntary, it is administered and financially supported by the metropolitan council. the council was established in 1967 and is still one of the only regional government entities in the country with the power to create policy and provide services, including public transit, wastewater treatment, and land use planning for a multi-county region. consequently, these counties can rely on longestablished networks of working together and complying with decisions made as a group. metrogis also advocates for open geodata across the rest of the state and has maintained a web page of open data resources12 since 2013. the john r. borchert map library at the university of minnesota has spearheaded several projects contributing to minnesota’s open data landscape. it developed and hosts one of the most widely used resources in the minnesota geospatial community, the minnesota historical aerial photographs online (mhapo)13 website, which provides discovery and access to aerial images dating back as far as 1923. this site features a map interface for finding over 100,000 images that were contributed from a variety of sources, including library holdings, the dnr, and the city of minneapolis (mcauliffe et al., 2017). mhapo was awarded the minnesota governor’s geospatial commendation in 2018, which was accompanied by numerous testimonials of its usefulness (minnesota geospatial advisory council, 2018). the borchert map library is also the project lead for the big ten academic alliance (btaa) geoportal.14 this is a collaboration of thirteen universities in ten states to aggregate metadata records for geospatial resources and provides access to them through a collective geoportal. since the btaa geoportal indexes metadata from state, regional, county, and municipal geodata portals, it fills a gap in minnesota’s open data landscape. although the commons does provide links to externally hosted county portals, the btaa geoportal takes it a step further to index each dataset layer and enrich the metadata with normalized place names, subjects, categories, and dates. lastly, several staff members from the borchert map library are leading efforts to implement a statewide archive for all public geodata. this has taken the form of multiple workgroups made up of members from the library; state, county, and municipal government; nonprofit and commercial sectors; and the minnesota historical society. this project has been many years in the making (dyke et al., 2016) and has a wide swath of support across the geospatial community, ranking third on the gac’s list of priorities for 2020. 3.2 case study ii: wisconsin wisconsin’s dedication to geodata creation across all levels of government has been a long-standing tradition for over 30 years, during which time there has been a steady series of changes in how geodata has been made available to the public. the wisconsin land records committee was established in 1985 to pave a path forward for modernizing land records. this committee developed the wisconsin land information program (wlip) to address several needs, including creating standardized data guidelines, reducing inefficient duplications of effort, saving public money, and keeping up with technology advancements, such as geographic information systems (wisconsin land records committee, 1987). the wlip was officially established by legislation in 1989 and continues to be an active program under the wisconsin department of administration (doa) today. https://doi.org/do.be/doo https://doi.org/10.29173/iq1013 https://www.metrogis.org/projects/free-open-data.aspx https://www.metrogis.org/projects/free-open-data.aspx https://apps.lib.umn.edu/mhapo/ https://apps.lib.umn.edu/mhapo/ https://geo.btaa.org/ 9/19 majewicz, karen; martindale, jaime; kernik, melinda (2022). open geospatial data: a comparison of data cultures in local government, iassist quarterly 46(1), pp. 1-19. doi: https://doi.org/10.29173/iq1013 the wlip has always specified that county participation is voluntary. however, its financial incentives are strong enough that all counties in the state eventually chose to opt into the program. the wlip provides funding to participants in the form of grants and allows them to keep a portion of the fees the state charges on real estate transactions (wisconsin land information board, 1991). in return, each county is required to operate a land information office and create specific geospatial datasets, known as foundational elements. since the counties must share certain foundational elements with the state, the wlip began to advocate for counties to make this data freely available online to the public as well. many of the land information offices across the state created websites with map viewers where the public could view valuable information. however, public access to the raw data files was not assured. some counties restricted access to their geodata by charging fees or setting up licenses. these barriers created an environment that made it difficult for consumers to actually obtain public geodata. academic organizations were one of the entities in wisconsin that could negotiate access to county geodata. beginning in 2005, the university of wisconsin-madison’s arthur h. robinson map library began acquiring local geodata directly from counties for use in academic research and teaching. at that time, over half of wisconsin’s 72 counties required the university to sign formal licenses or data sharing agreements indicating the data would only be used by uw-madison users for academic purposes. other counties agreed to share the data with the university without signing formal agreements, but the general understanding was that the data would be only used for academic purposes. this collection process was sporadic, as the library only requested and archived county geodata files when students or researchers specifically requested them. while this enabled continual growth of the data archive, holdings became inconsistent and unpredictable through time. as a result, the robinson map library changed its county data collection process in 2012 to request a comprehensive set of data layers across all counties at the same time each year. this list of data layers was standardized to encourage broad participation. based on previous experience, more favorable and timely responses were garnered when a specific set of layers was requested, as opposed to a catch-all ‘give us what you have’ request. for many years, academic users needed to visit the library in person to obtain the data on cds, dvds, or portable hard drives. by 2012, advances in cloud-based file-sharing services enabled users to simply download the data directly from the internet. users no longer needed to physically go to the library, because they could submit email requests to access the content at any time. however, the individual requests became too frequent to handle efficiently, and the data archive quickly reached a critical mass of temporally significant content. in response to this growth, the robinson map library made the pivotal decision to develop an online geoportal to serve as the discovery platform for all geodata in the archive. the library collaborated with the wisconsin state cartographer’s office and launched geodata@wisconsin15 in 2014. initially, the only users who were able to download resources from the geoportal were affiliates of uw-madison. this restriction allowed the library to remain in compliance with data-sharing agreements that were still in place for nearly 40 counties. however, once users around the state became aware of the new geoportal, the library was inundated with geodata access requests from students and researchers at other wisconsin campuses. the library staff surveyed the 72 land information officers in each county and found that 70 of them were willing to share their geodata more widely, as long as it continued to be for academic purposes only. the library then changed the geoportal’s authentication protocols to allow access for users affiliated with any university of wisconsin system campus. to avoid having to programmatically address https://doi.org/do.be/doo https://doi.org/10.29173/iq1013 http://geodata.wisc.edu/ 10/19 majewicz, karen; martindale, jaime; kernik, melinda (2022). open geospatial data: a comparison of data cultures in local government, iassist quarterly 46(1), pp. 1-19. doi: https://doi.org/10.29173/iq1013 multiple levels of access authentication, resources from the two counties that did not approve of broader access were simply removed from the geoportal. at the same time, the broader open data movement was taking hold around the country, and there was a sense that the culture was changing in wisconsin as well. generational differences became apparent, and staff turnover in county land information offices resulted in different mindsets. data storage and web hosting services became less expensive and easier to use, and the momentum grew as more and more counties began posting downloadable datasets online. proponents across the state pointed out that the administrative costs of charging fees exceeded the revenue and that the increased user base that comes with free data could translate into increased economic activity. in this changing environment, a major development occurred: the passage of the statewide parcel map initiative. while this initiative alone might not be viewed as the only catalyst for establishing open geodata in wisconsin, the totality of the environment in which it was adopted and carried out is certainly marked by that spirit. the statewide parcel map initiative was established by wisconsin act 20, the biennial state of wisconsin budget for 2013-2015. this act includes statutory directives for a multi-faceted, multi-year collaborative effort of the department of administration (doa) and local governments to coordinate the development of a statewide digital parcel map. it requires counties to submit parcel datasets online in a standardized format and provides additional grant funding administered by the wlip for counties to improve their parcel mapping (wisc. stat., § 59.72). in addition to the parcel information explicitly called for by state law, doa broadened their geodata collection scope in 2017 to include other common foundational datasets, such as address points, street centerlines, land use, zoning, rights of way, and more. the collection of other layers beyond parcels was not made inevitable by the passage of act 20. the expansion was largely motivated by an effort to create a mutually beneficial data collection process for doa and the robinson map library, taking into account the process previously established by the library. doa’s inclusion of additional datasets was driven by the desire to create efficiencies, synergize, and assist where possible to help the library get closer to 100% compliance with their annual data request. the robinson map library plays a central role in creating geospatial metadata required for documenting data collected each year. doa now makes geodata requests to the counties and directs them to send their datasets along with basic metadata to the library. library staff then create fully valid iso 19139 metadata and publish the datasets on geodata@wisconsin, which has since been fully opened to the general public. now, any visitor to the site (not just academic affiliates) can browse and download public geodata. this change in wisconsin’s open data landscape has been bumpy at times, as not everyone in the geospatial community was in support of the initiative at the start. in a small number of instances, counties that initially balked at doa’s request for other geodata layers beyond parcels eventually acceded to the request. sometimes this involved the land information officer working on getting the county’s official policy changed in cases where it was necessary to terminate local policies requiring signed license agreements and fees for the acquisition of data. in one case, doa representatives went before a county land information council and successfully made the case for sharing the other layers over the objections of a county gis staff person. for data not created with wlip grant funding, the requirement for sharing may not be as direct, but wisconsin’s public records laws provide an additional basis for doa to request, collect, and make the data open. unless specifically exempted by federal law or statute, the requested datasets are assumed to be public records under state statute 19.31 and are therefore to be made available upon request (wisc. stat., § 19.31). in cases where a particular county was not sharing its data, https://doi.org/do.be/doo https://doi.org/10.29173/iq1013 https://geodata.wisc.edu/ 11/19 majewicz, karen; martindale, jaime; kernik, melinda (2022). open geospatial data: a comparison of data cultures in local government, iassist quarterly 46(1), pp. 1-19. doi: https://doi.org/10.29173/iq1013 doa officials have simply asked the county why the wisconsin public records law does not apply to the requested records in question. this is often sufficient to bring the county on board, particularly after pointing out that courts around the country have routinely ruled in favor of open data when statutes are challenged (sierra club v. s.c (county of orange), 2013; wiredata inc. v. village of sussex, 2008.). although great strides have been made in gaining open access to much county vector data, there has been less willingness to share raster data in a minority of counties. in 2018-2019, doa began requesting lidar elevation datasets from individual counties. three counties have denied the request because they charge a significant sum of money for the data and do not want the data available for free. doa has decided to defer aggressively pursuing the lidar data from these three counties for now and focus on gathering data from other willing counties (herreid & veselenak, personal communication, wisconsin land information program (wlip) and act 20 email questionnaire responses, feb 24, 2020). despite these challenges, a large amount of wisconsin’s county geodata has now become open data in practice. this can be generally attributed to leadership and support from members of the community, legislation, grant funding, and calls for increased transparency in government operations at all levels. 3.3 commonalities and points of departure one way of assessing a state’s open data landscape is to tally the number of counties that are actively publishing it, either through their own hosted portal or by contributing directly to a state clearinghouse. from this perspective, the digital landscape in minnesota and wisconsin is similar. an examination of each state’s public list of county-level gis websites reveals that roughly half of the counties in each state selfpublish open geodata. (minnesota geospatial information office, 2021; wisconsin land information program, 2021).16 another quantitative method for assessing the open data landscape is to focus on the geographic availability of specific dataset themes. from this perspective, the two states are much more divergent. this evolution can be seen by examining the availability of county parcel datasets over time. at one time, minnesota had more open parcel datasets than wisconsin, but the situation has since flipped. figure 2 illustrates minnesota’s early presence in online open geodata by showing the availability of county parcel datasets in 2005. in that year, eight counties in minnesota were publishing parcel datasets seven in the twin cities area through metrogis, along with the pioneering clay county on the western border (which began the practice all the way back in 1999). there is not a reliable comparison to wisconsin during this period, because the availability of parcel datasets fluctuated depending upon policies in each county at the time. ten years later, the open data landscape in these states paints a different picture. figure 3 displays which counties in minnesota and wisconsin published parcel data as open geodata in the year 2015. while six additional counties in minnesota had joined the open geodata movement, wisconsin now had full coverage of this data for every single county. https://doi.org/do.be/doo https://doi.org/10.29173/iq1013 12/19 majewicz, karen; martindale, jaime; kernik, melinda (2022). open geospatial data: a comparison of data cultures in local government, iassist quarterly 46(1), pp. 1-19. doi: https://doi.org/10.29173/iq1013 figure 2: a map of minnesota showing which counties published parcel data as open geodata in 2005. figure 3: a map of minnesota and wisconsin showing which counties published parcel data as open geodata in 2015. https://doi.org/do.be/doo https://doi.org/10.29173/iq1013 13/19 majewicz, karen; martindale, jaime; kernik, melinda (2022). open geospatial data: a comparison of data cultures in local government, iassist quarterly 46(1), pp. 1-19. doi: https://doi.org/10.29173/iq1013 we have identified four areas that we believe have had the greatest impact on the differing landscapes of open geodata between minnesota and wisconsin. 1. legislation the two states have different approaches to how they regulate the public availability of county geodata. both have long-standing statutes that define government data as open to the public. however, subsequent qualifications to this legislation have had substantial impacts on the open data landscape. in minnesota, the data practices act is not fully enforced, and an addendum allows counties to charge a fee to cover the cost of packaging and sharing their geodata. in wisconsin, the language in the public records law provided justification for a budget act that funds mandatory collection of parcel datasets. it also promoted the notion that more county-level geodata could become free and open. 2. funding dedicated funding for local gis departments makes a big difference for open data, particularly in rural counties. one effective funding mechanism is a recorder’s fee attached to each real estate transaction. in minnesota, counties can optionally use this fee to fund their gis work, but many choose to spend the funds another way. in wisconsin, every county participating in the wlip is required to have a land information office that is funded through retained fees and grants from the program. counties retain a portion of a dedicated $15 real estate document recording fee to fund their land information work, while the remainder is allocated to the state’s land information fund. this fund provides base budget grants to counties that see fewer real estate transactions and generate less than $100,000 per year in retained fees. however, the counties must participate in the program to submit foundational datasets to the state, or else they may not be eligible to receive grant funding. county geodata produced with wlip funding and submitted to the annual call for data is open and publicly accessible. 3. workflows another discrepancy between the states can be seen in how local governments participate in their state’s central geodata platform. minnesota uses a self-service model for contributions to the state geodata platform. although a state agency administers the commons platform itself, the content is fully the responsibility of the contributors. counties need to set up a local node on a file-sharing application, write their own metadata, and upload it bundled with their datasets to the commons; these are all tasks that require a fair amount of staff time to perform. the commons also requires valid metadata that conforms to the state guidelines. without validation, the submission will not go through. this keeps the quality of the data in the commons very high but has the effect of preventing some counties from participating. out of the 45 counties that have open geodata, only ten contribute to the commons. in wisconsin, workflows are more centralized. all counties send specified datasets to the state cartographer’s office and the robinson map library. staff at the state cartographer’s office process the tax parcel data for the creation of the statewide layer, while library staff write full standards metadata for all the incoming datasets and publish them to a geoportal. the geoportal is developed and maintained by the state cartographer’s office and map library, both units at the university of wisconsin-madison. 4. library involvement currently, both states have some level of academic library involvement in open data workflows and discussions. in minnesota, the borchert map library is an active participant in the state’s efforts around open data. it is one of the contributors to the commons, and it hosts one of the most used geospatial access points in the form of a historical aerial photograph finder. more recently, the library has begun to make plans for archiving open geodata in the same way it has done for public domain maps for decades. in wisconsin, the robinson map library has been involved in collecting, archiving, and disseminating https://doi.org/do.be/doo https://doi.org/10.29173/iq1013 14/19 majewicz, karen; martindale, jaime; kernik, melinda (2022). open geospatial data: a comparison of data cultures in local government, iassist quarterly 46(1), pp. 1-19. doi: https://doi.org/10.29173/iq1013 geospatial data since 2005. the library’s process of curating geospatial collections for academic research (out of necessity for users) evolved from user-specific acquisitions to a consistent annual collection of county geospatial data for all of wisconsin. over time, this focused effort became a more formal process with goals for broader access expanding into long-term preservation of the data as well. with the library’s annual data acquisition process in place, it made sense to couple it with the statewide parcel initiative beginning in 2017. doing so means less of a burden for county data providers who only need to respond to a single data request each year. an added benefit to the library’s formal role in wisconsin’s open data acquisition process is the creation of standards-based geospatial metadata. both descriptive and discovery metadata are created by library staff and student assistants with guidelines in place that make the records accurate and consistent. student assistants have always been a significant part of the geospatial metadata workflow. a primary goal of the library is to hire and train students in relevant educational programs. students obtain worthwhile training, education, and applied work experience in data management and documentation. 3.4 additional observations we speculate that minnesota's early flourishing in gis technology could have actually impeded their later open geodata efforts. for example, minnesota was one of the first states in the nation to create geodata on a statewide scale and one of the first to deliver it via open data portals. students and researchers could access open geodata through multiple portals as far back as the 1990s. meanwhile, students and researchers in wisconsin continued to face significant challenges in obtaining geodata without the assistance of the university negotiating on their behalf. this prompted the robinson map library to build a geodata archive years before the borchert map library began investigating a similar project. minnesota, with its plethora of voluntary open geodata, has not had a comparable collection program that can be easily converted into an archive. interestingly, in late 2020, the university of minnesota’s u-spatial program17 began offering limited access to parcel data for every county in the state using a model similar to what the robinson map library started doing in 2005. through this arrangement, students and researchers may request authorized access to parcel data but must agree to use it for research purposes only and not share it. another example of early adoption impeding later progress is the status of minnesota’s metadata guidelines and accompanying tools. mgmg is deeply enmeshed in the documentation and workflows for state agencies and the commons. this is evidenced in the longevity of the primary mgmg authoring tool, known as the minnesota metadata editor (mme). in the intervening years since mgmg and mme were developed, the international standards organization released a new geospatial metadata standard, the iso 191xx series, and arcgis for desktop became the most widely used tool for creating it. in response to these developments, a gac metadata workgroup analyzed mgmg’s compatibility with iso and the arcgis authoring tools. the workgroup concluded in 2017 that ‘there are not yet sufficient business needs to migrate mgmg to be fully compliant with iso’ (minnesota geospatial advisory council, 2017). although the workgroup identified techniques for using arcgis, mme remains the most reliable tool for generating valid mgmg. this is a point of frustration for data creators because mme is an outdated windows-only application that relies upon microsoft access, a deprecated program. in contrast, wisconsin has been able to be more nimble about technological adoption for open data as its efforts have been more recent. the robinson map library uses arcgis pro to create metadata, which can export to either the fgdc or iso standard. they also have been able to take advantage of a more modern interface, geoblacklight, which incorporates geospatial web service previews into item view pages. minnesota had thoroughly developed a state geodata platform before arcgis hub or geoblacklight had https://doi.org/do.be/doo https://doi.org/10.29173/iq1013 https://research.umn.edu/units/uspatial/resources/spatial-data https://research.umn.edu/units/uspatial/resources/spatial-data 15/19 majewicz, karen; martindale, jaime; kernik, melinda (2022). open geospatial data: a comparison of data cultures in local government, iassist quarterly 46(1), pp. 1-19. doi: https://doi.org/10.29173/iq1013 matured as technology options, and it remains invested in using ckan, a technology designed for general purpose data. our examination of the open geodata landscapes in each state indicates that if a government entity values open data, it should look to wisconsin as a model. however, wisconsin's success story may not be wellknown outside of their state. it was not evident in the most recent survey of the national states geographic information council (nsgic) (2019). nsgic conducts surveys every two years to summarize and evaluate the geospatial maturity of each state. although free and open data is one component of the scoring metric, it is not a significant focus of the assessment. the 2019 nsgic survey gave minnesota a grade of 'a' for statewide geospatial coordination. wisconsin received a 'd.' this result is puzzling, as wisconsin is arguably well-coordinated in terms of statewide geospatial activities. for example, wisconsin is one of the only states to require that every county establish a land information office and council, and a dedicated state agency distributes funds to each county. however, the survey did not pose questions related to these aspects. the second area where wisconsin lost many points was whether or not an official state clearinghouse existed. although the university of wisconsin-madison maintains a large geoportal that is at least as comprehensive as any across the country, it was not represented in the nsgic survey as an official state clearinghouse. on the whole, the survey's language often matched minnesota's structure but did not reward wisconsin. 4. conclusion when will the conference attendee we described in our introduction be able to freely download a parcel dataset for every county in minnesota? as of 2021, minnesota continues to gradually increase the number of counties offering open geodata, with a pattern of several new ones signing on every year. although minnesota has many enthusiastic open data supporters that are making real progress, it seems unlikely that the state will attain full open coverage of foundational layers like parcels without adopting one or more of wisconsin's strategies. based upon our case study of wisconsin, we can predict that a few of the remaining minnesota counties will only embrace open geodata if they are mandated to do so while receiving centralized support on multiple fronts. they need dedicated funding for staff positions and technical support for metadata services along with an easy-to-use centralized platform. the state government plays a role by passing legislation and providing a financial incentive, while the libraries are well-suited to play an essential role in resource discovery, metadata, and preservation. references arbeit, d. et al. (2004) a foundation for coordinated gis: minnesota’s spatial data infrastructure. minnesota governor’s council on geographic information. available at: https://www.mngeo.state.mn.us/msdi/mn_iplan_consolidation_final_04oct04.pdf (accessed: 19 february 2020). blatt, a. j. (2016) ‘open-access geospatial data: promise and potential’, journal of map & geography libraries, 12(2), pp. 216–222. doi: https://doi.org/10.1080/15420353.2015.1125405. dyke, k. r. et al. (2016) ‘placing data in the land of 10,000 lakes: navigating the history and future of geospatial data production, stewardship, and archiving in minnesota’, journal of map & geography libraries, 12(1), pp. 52–72. doi: https://doi.org/10.1080/15420353.2015.1073655. https://doi.org/do.be/doo https://doi.org/10.29173/iq1013 https://www.mngeo.state.mn.us/msdi/mn_iplan_consolidation_final_04oct04.pdf https://doi.org/10.1080/15420353.2015.1125405 https://doi.org/10.1080/15420353.2015.1073655 16/19 majewicz, karen; martindale, jaime; kernik, melinda (2022). open geospatial data: a comparison of data cultures in local government, iassist quarterly 46(1), pp. 1-19. doi: https://doi.org/10.29173/iq1013 federal geographic data committee (1994) the 1994 plan for the national spatial data infrastructure. reston, virginia. available at: https://www.fgdc.gov/policyandplanning/nsdi%20strategy%201994.pdf. federal geographic data committee (1997) a strategy for the national spatial data infrastructure. reston, va. available at: https://hdl.handle.net/2027/umn.31951d01539752a. harvey, f. and tulloch, d. (2006) ‘local‐government data sharing: evaluating the foundations of spatial data infrastructures’, international journal of geographical information science, 20(7), pp. 743– 768. doi: https://doi.org/10.1080/13658810600661607. joffe, b. (2003) ‘10 ways to support your gis without selling data’. available at: https://umaine.edu/computingcoursematerials/wpcontent/uploads/sites/511/2017/02/tenways.pdf. johnson, p. a. et al. (2017) ‘the cost(s) of geospatial open data’, transactions in gis, 21(3), pp. 434– 445. doi: https://doi.org/10.1111/tgis.12283. kemp, b. (2017) ‘implementation of ng9-1-1 in rural america–the counties of southern illinois: experience and opportunities’, ieee communications magazine, 55(1), pp. 152–158. doi: https://doi.org/10.1109/mcom.2017.1600457cm. kramers, e. r. (2008) ‘interaction with maps on the internet – a user centred design approach for the atlas of canada’, the cartographic journal, 45(2), pp. 98–107. doi: https://doi.org/10.1179/174327708x305094. library of congress (2021) sustainability of digital formats: planning for library of congress collections. available at: https://www.loc.gov/preservation/digital/formats/fdd/fdd000280.shtml (accessed: 22 september 2021). maas, g. (2013) research & reference documents. white paper. minnesota: metrogis. available at: https://www.metrogis.org/getmedia/a766418c-87a5-4e73-a87e187022b04d50/metrogis_free_open_data_research_resources3.pdf.aspx. maas, g. (2019) questions, answers, concepts, and resources for practitioners. white paper version 6.3. minnesota: metrogis. available at: https://www.metrogis.org/getmedia/e6a25fbe-89cd-43dba80d-fa32fc5ea287/free_open_data_version_6_3.pdf.aspx. mcauliffe, c. p., lage, k. and mattke, r. (2017) ‘access to online historical aerial photography collections: past practice, present state, and future opportunities’, journal of map & geography libraries, 13(2), pp. 198–221. doi: https://doi.org/10.1080/15420353.2017.1334252. minnesota geospatial advisory council (2018) ‘minnesota governor’s geospatial commendation award nomination’. available at: https://www.mngeo.state.mn.us/awards/gov_commendations/govgeospatialcommendationn omination_umn_borchertmaplibrary_2018.pdf. https://doi.org/do.be/doo https://doi.org/10.29173/iq1013 https://www.fgdc.gov/policyandplanning/nsdi%20strategy%201994.pdf https://hdl.handle.net/2027/umn.31951d01539752a https://doi.org/10.1080/13658810600661607 https://umaine.edu/computingcoursematerials/wp-content/uploads/sites/511/2017/02/tenways.pdf https://umaine.edu/computingcoursematerials/wp-content/uploads/sites/511/2017/02/tenways.pdf https://doi.org/10.1111/tgis.12283 https://doi.org/10.1109/mcom.2017.1600457cm https://doi.org/10.1179/174327708x305094 https://www.loc.gov/preservation/digital/formats/fdd/fdd000280.shtml https://www.metrogis.org/getmedia/a766418c-87a5-4e73-a87e-187022b04d50/metrogis_free_open_data_research_resources3.pdf.aspx https://www.metrogis.org/getmedia/a766418c-87a5-4e73-a87e-187022b04d50/metrogis_free_open_data_research_resources3.pdf.aspx https://www.metrogis.org/getmedia/e6a25fbe-89cd-43db-a80d-fa32fc5ea287/free_open_data_version_6_3.pdf.aspx https://www.metrogis.org/getmedia/e6a25fbe-89cd-43db-a80d-fa32fc5ea287/free_open_data_version_6_3.pdf.aspx https://doi.org/10.1080/15420353.2017.1334252 https://www.mngeo.state.mn.us/awards/gov_commendations/govgeospatialcommendationnomination_umn_borchertmaplibrary_2018.pdf https://www.mngeo.state.mn.us/awards/gov_commendations/govgeospatialcommendationnomination_umn_borchertmaplibrary_2018.pdf 17/19 majewicz, karen; martindale, jaime; kernik, melinda (2022). open geospatial data: a comparison of data cultures in local government, iassist quarterly 46(1), pp. 1-19. doi: https://doi.org/10.29173/iq1013 minnesota geospatial advisory council (2017) metadata workgroup final report. minnesota geospatial advisory council. available at: https://www.mngeo.state.mn.us/workgroup/metadata/metadata_workgroup_final_report_201 7.pdf. minnesota geospatial advisory council outreach committee (2016) free and open public geospatial data: 2016 outreach survey results and report. minnesota geospatial advisory council. available at: https://www.mngeo.state.mn.us/committee/outreach/opendatasurveyfindingsreport.pdf. minnesota geospatial information office (2021) ‘status of free and open public geospatial data from minnesota counties’. available at: https://gisdata.mn.gov/dataset/bdry-mn-county-open-datastatus. minnesota government data practices act (1974) minnesota statutes § 13. available at: https://www.revisor.mn.gov/statutes/cite/13 (accessed: 19 february 2020). minnesota government data practices act (1990) minnesota statutes § 13.03(d. available at: https://www.revisor.mn.gov/statutes/cite/13 (accessed: 19 february 2020). minnesota governor’s council on geographic information (1998) minnesota geographic metadata guidelines. available at: https://mn.gov/mnit/government/policies/geo/mn-geographicmetadata.jsp. minnesota governor’s council on geographic information (2009) annual report. available at: https://www.lrl.mn.gov/docs/2010/other/100728.pdf. national emergency number association (2020) nena standard for ng9-1-1 gis data model. available at: https://cdn.ymaws.com/www.nena.org/resource/resmgr/standards/nena-sta-006.1.12020_ng9-1-.pdf (accessed: 6 march 2020). national research council (1980) need for a multipurpose cadastre. doi: https://doi.org/10.17226/10989. national research council (2007) national land parcel data: a vision for the future. washington, d.c.: national academies press. national states geographic information council (2019) 2019 geospatial maturity assessment state report cards. available at: https://nsgic.memberclicks.net/assets/2019gmarawresults/2019gmareportcards/gma%20st ate%20report%20cards.pdf. open knowledge foundation (no date) open definition 2.1 open definition defining open in open data, open content and open knowledge. available at: https://opendefinition.org/od/2.1/en/ (accessed: 12 march 2020). sierra club v. s.c (county of orange), no. s194708 (california supreme court 2013). https://doi.org/do.be/doo https://doi.org/10.29173/iq1013 https://www.mngeo.state.mn.us/committee/outreach/opendatasurveyfindingsreport.pdf https://www.mngeo.state.mn.us/committee/outreach/opendatasurveyfindingsreport.pdf https://www.mngeo.state.mn.us/workgroup/metadata/metadata_workgroup_final_report_2017.pdf https://www.mngeo.state.mn.us/workgroup/metadata/metadata_workgroup_final_report_2017.pdf https://www.mngeo.state.mn.us/committee/outreach/opendatasurveyfindingsreport.pdf https://gisdata.mn.gov/dataset/bdry-mn-county-open-data-status https://gisdata.mn.gov/dataset/bdry-mn-county-open-data-status https://www.revisor.mn.gov/statutes/cite/13 https://www.revisor.mn.gov/statutes/cite/13 https://mn.gov/mnit/government/policies/geo/mn-geographic-metadata.jsp https://mn.gov/mnit/government/policies/geo/mn-geographic-metadata.jsp https://www.lrl.mn.gov/docs/2010/other/100728.pdf https://cdn.ymaws.com/www.nena.org/resource/resmgr/standards/nena-sta-006.1.1-2020_ng9-1-.pdf https://cdn.ymaws.com/www.nena.org/resource/resmgr/standards/nena-sta-006.1.1-2020_ng9-1-.pdf https://doi.org/10.17226/10989 https://nsgic.memberclicks.net/assets/2019gmarawresults/2019gmareportcards/gma%20state%20report%20cards.pdf https://nsgic.memberclicks.net/assets/2019gmarawresults/2019gmareportcards/gma%20state%20report%20cards.pdf https://opendefinition.org/od/2.1/en/ 18/19 majewicz, karen; martindale, jaime; kernik, melinda (2022). open geospatial data: a comparison of data cultures in local government, iassist quarterly 46(1), pp. 1-19. doi: https://doi.org/10.29173/iq1013 sui, d. (2014) ‘opportunities and impediments for open gis’, transactions in gis, 18(1), pp. 1–24. doi: https://doi.org/10.1111/tgis.12075. terner, m., buck, a., and applied geographics, inc. (2009) a program for transformed gis in the state of minnesota: program design & implementation plan. available at: https://www.mngeo.state.mn.us/msdi/dte/programdesign_finalfeb09_v21.pdf (accessed: 19 february 2020). tombs, r. b. (2005) ‘policy review: blocking public geospatial data access is not only a homeland security risk’, journal of the urban & regional information systems association, 16(2), pp. 49–51. tulloch, d. and fuld, j. (2001) ‘exploring county-level production of framework data: analysis of the national framework data survey’, urisa journal. available at: https://www.researchgate.net/publication/239984711_exploring_countylevel_production_of_framework_data_analysis_of_the_national_framework_data_survey. university of minnesota. center for urban and regional affairs (1976) overview of the minnesota land management information system. [minn.]: university of minnesota center for urban and regional affairs. u.s. department of housing and urban development (2013) the feasibility of developing a national parcel database: county data records project final report. available at: https://www.huduser.gov/portal/publications/pdf/feasibility_nat_db.pdf (accessed: 6 march 2020). warnecke, l. (1992) ‘minnesota’, in state geographic information activities compendium. lexington, ky: council of state governments, pp. 212–245. wiredata, inc. v. village of sussex, 751 nw 2d 736 (wisconsin supreme court 2008). wirtz, b. w. et al. (2016) ‘resistance of public personnel to open government: a cognitive theory view of implementation barriers towards open government data’, public management review, 18(9), pp. 1335–1364. doi: https://doi.org/10.1080/14719037.2015.1103889. wisconsin land information board (1991) recommendations and requirements for county-wide plans for land records modernization. available at: https://doa.wi.gov/dir/1991_plan_instructions.pdf. wisconsin land information program (2021) ‘county contacts and websites’. available at: https://doa.wi.gov/dir/county_contacts.pdf. wisconsin land records committee (1987) final report of the wisconsin land records committee: modernizing wisconsin’s land records. madison, wisconsin. available at: https://doa.wi.gov/dir/1987_report_land_records_committee.pdf. wisconsin statutes § 19.31 (1981) available at: https://docs.legis.wisconsin.gov/statutes/statutes/19/ii/31 (accessed: 10 june 2021). https://doi.org/do.be/doo https://doi.org/10.29173/iq1013 https://doi.org/10.1111/tgis.12075 https://www.mngeo.state.mn.us/msdi/dte/programdesign_finalfeb09_v21.pdf https://www.researchgate.net/publication/239984711_exploring_county-level_production_of_framework_data_analysis_of_the_national_framework_data_survey https://www.researchgate.net/publication/239984711_exploring_county-level_production_of_framework_data_analysis_of_the_national_framework_data_survey https://www.huduser.gov/portal/publications/pdf/feasibility_nat_db.pdf https://doi.org/10.1080/14719037.2015.1103889 https://doa.wi.gov/dir/1991_plan_instructions.pdf https://doa.wi.gov/dir/county_contacts.pdf https://doa.wi.gov/dir/1987_report_land_records_committee.pdf https://docs.legis.wisconsin.gov/statutes/statutes/19/ii/31 19/19 majewicz, karen; martindale, jaime; kernik, melinda (2022). open geospatial data: a comparison of data cultures in local government, iassist quarterly 46(1), pp. 1-19. doi: https://doi.org/10.29173/iq1013 wisconsin statutes § 59.72 (2013) available at: https://docs.legis.wisconsin.gov/statutes/statutes/59/vii/72 (accessed: 10 june 2021). wood, c. (2020) states need to improve for national gis infrastructure to work, report says. statescoop. available at: https://statescoop.com/nsgic-geospatial-maturity-assessment-state-gis-reportcard/ (accessed: 6 march 2020). endnotes 1 karen majewicz is the geospatial project manager and metadata coordinator at the john r. borchert map library, university of minnesota (majew030@umn.edu) 2 jaime martindale is the map & geospatial data librarian at the arthur h. robinson map library, university of wisconsin-madison (jmartindale@wisc.edu) 3 melinda kernik is the spatial data analyst and curator at the john r. borchert map library, university of minnesota (kerni016@umn.edu) 4 yijing zhou is the gis & metadata programming intern for the btaa geoportal, university of minnesota 5 ckan, https://ckan.org 6 socrata, https://www.tylertech.com/products/socrata 7 arcgis hub, https://hub.arcgis.com 8 geoblacklight, https://geoblacklight.org 9 minnesota department of natural resources gis data deli, https://web.archive.org/web/19991013033555/http://deli.dnr.state.mn.us/ 10 metrogis data finder, https://web.archive.org/web/19981203091211/http://www.datafinder.org/ 11 minnesota geospatial commons, https://gisdata.mn.gov 12 metro gis free and open data, https://www.metrogis.org/projects/free-open-data.aspx 13 minnesota historical aerial photography (mhapo), https://apps.lib.umn.edu/mhapo 14 btaa geoportal, https://geo.btaa.org 15 geodata@wisconsin, https://geodata.wisc.edu 16 these lists are frequently updated. consult the latest version of the status of free and open geospatial data from minnesota, https://gisdata.mn.gov/dataset/bdry-mn-county-open-data-status and county contacts, https://doa.wi.gov/dir/county_contacts.pdf, for the most up-to-date information. 17 u-spatial, university of minnesota, https://research.umn.edu/units/uspatial/resources/spatial-data https://doi.org/do.be/doo https://doi.org/10.29173/iq1013 https://docs.legis.wisconsin.gov/statutes/statutes/59/vii/72 https://statescoop.com/nsgic-geospatial-maturity-assessment-state-gis-report-card/ https://statescoop.com/nsgic-geospatial-maturity-assessment-state-gis-report-card/ https://ckan.org/ https://www.tylertech.com/products/socrata https://hub.arcgis.com/ https://geoblacklight.org/ https://gisdata.mn.gov/ https://www.metrogis.org/projects/free-open-data.aspx https://apps.lib.umn.edu/mhapo/ https://geo.btaa.org/ https://geodata.wisc.edu/ https://gisdata.mn.gov/dataset/bdry-mn-county-open-data-status https://doa.wi.gov/dir/county_contacts.pdf 1/22 zenk-möltgen, wolfgang (2025) the role of fair principles in high-quality research data documentation, iassist quarterly 49(2), pp. 1-22. doi: https://doi.org/10.29173/iq1119 the creative commons-attribution-noncommercial license 4.0 international applies to all works published by iassist quarterly. authors will retain copyright of the work and full publishing rights. the role of fair principles in high-quality research data documentation: looking at national election studies wolfgang zenk-möltgen1 abstract the fair principles as a framework for evaluating and improving open science and research data management have gained much attention over the last years. by defining a set of properties that indicates good practice for making data findable, accessible, interoperable, and reusable (fair), a quality measurement is created, which can be applied to diverse research outputs, including research data. there are some software tools available to help with the assessment, with the f-uji tool being the most prominent of them. it uses a set of metrics which defines tests for each of the fair components, and it creates an overall assessment score. the article examines differences between manually and automatically assessing fair principles, shows that there are significantly different results by using national election studies as examples. an evaluation of progress is done by comparing the automatically assessed fairness scores of the datasets from 2018 with those of 2024, showing that there is only a very slight yet not significant difference. specific measures which have improved the fairness scores are described by the example of the politbarometer 2022 dataset at the gesis data archive. the article highlights the role of archives in securing a high level of data and metadata quality and technically sound implementation of the fair principles to help researchers benefit from getting the most of their valuable research data. keywords fair principles, data documentation, research data management, f-uji test, research transparency introduction good data documentation is indispensable for working with research data. it is especially important when using data that was collected or gathered by other researchers; knowledge about the data collection procedures, applied concepts, data cleaning and transformation steps, and decisions by principal investigators along the way is necessary to fully understand the findings and implications. in recent years, the movement for more research transparency has grown considerably, partly due to discovery of errors and scientific misconduct (christensen et al., 2019; freese & peterson, 2017). open science and requests for a transparent process of the scientific endeavor led to several improvements with the availability and re-usability of research data. social science data archives such as the icpsr in the u.s., gesis – leibniz institute for the social sciences in germany, or the uk data archive have existed for decades. in addition, research data centers as well as more generic repositories for research data have emerged in recent years (e.g., zenodo2, figshare3, dataverse4, and https://doi.org/10.29173/iq1119 2/22 zenk-möltgen, wolfgang (2025) the role of fair principles in high-quality research data documentation, iassist quarterly 49(2), pp. 1-22. doi: https://doi.org/10.29173/iq1119 the open science framework osf5). all these services support the documentation of research data, however in different degrees of granularity. for all of these services, it is clear that well-done data documentation and provision of research data is costly and needs considerable resources (perry & netscher, 2022). but cases of fraud in more than only a few disciplines (christensen et al., 2019) and acknowledgement for the value of data have led to considerable improvements in data curation, documentation, and access. increasingly also research funders develop guidelines and regulations that support data archiving, metadata documentation, and data sharing (american economic association, 2024; deutsche gesellschaft für soziologie, 2019;u.s. national science foundation, 2018; wissenschaftsrat, 2020). especially in the political science domain, there is a collective understanding and expectation for journals to have a policy that requires authors to deposit the data underlying their article findings at a trusted archive or data repository and make it available for independent scrutiny (da-rt, 2015). it was found that data availability is much higher for articles published in journals with a data sharing policy in place (key, 2016; zenk-möltgen et al., 2018). some policies also require independent verification of the results and sometimes provide staff at the journal to conduct this work. literature review the fair principles have been developed as ‘guiding principles for findable, accessible, interoperable, and re-usable data publishing‘6 within the force11 scholarly community initiative (wilkinson et al., 2016). with the fairsfair (fostering fair data practices in europe) project (devaraju et al., 2022), the fair guiding principles were established in practice as an open standard for evaluating the findability, accessibility, inter-operability, and re-producibility of research data. evaluation of datasets against these criteria have initially been done manually (bishop & hank, 2018; eder & jedinger, 2019; guillot et al., 2023; maxwell et al., 2021). most evaluations in the domains of political science or sociology, but also across disciplines (stall et al., 2019), came to the conclusion that more needs to be done to make the used research data findable, accessible, interoperable, and reusable (betancort cabrera et al., 2020). efforts have also been undertaken to balance requirements of openness with privacy requirements that regularly exist in the social sciences (borgesius, et al., 2016). based on these developments, automated testing tools were developed to assess the ‘fairness’ of a given dataset. for an overview of manual and automatic solutions for evaluation, see the fairassist7 list. the most prominent automatic tool is the f-uji tool (devaraju & huber, 2021) which can be used as a stand-alone implementation or via the provided website8. it is mainly dependent on persistent identifiers (or at least urls) for the datasets and produces detailed test metrics as well as an overall fairness score. a comparison of different automatic fair assessment tools came to the conclusion that there are significant differences in the design, implementation, and documentation of the evaluation metrics for the tools (sun, et al, 2022). assessments for several domain specific research datasets have been done in previous research (alaterä et al., 2022; petrosyan et al., 2023; sofi-mahmudi & raittio, 2022). however, the automatic assessment has also been criticized as not being able to evaluate the quality of metadata or if metadata is ‘rich’ and domain-specific enough to enable reusability (musen et al., 2022). this criticism is certainly valid for a comparison between automatic assessments to a review of metadata conducted by an expert in the fields looking at a small number of datasets. but the automatic procedures allow https://doi.org/10.29173/iq1119 3/22 zenk-möltgen, wolfgang (2025) the role of fair principles in high-quality research data documentation, iassist quarterly 49(2), pp. 1-22. doi: https://doi.org/10.29173/iq1119 the evaluation of a large number of datasets in a standardized way. and given that the metrics were developed within a broad framework of stakeholders and with large support of the scientific community, it can be assumed that they represent at least some basic common understanding of documentation quality (devaraju et al., 2021). however, it remains unclear if progress has happened with research data transparency over the years. to evaluate this question, automatic measurements which are using clearly defined tests might be suitable, even if they cannot perform a qualitative evaluation or determine whether humans are satisfied with the level of data documentation and availability. the aim of this article is therefore to find out if a valid evaluation of fairness can be done by an automatic assessment, if a comparison of these assessments over time can show improvement, and how data archives and repositories can contribute to this. research design using well-known research data seems highly appropriate for a comparison between manual and automatic assessments and for evaluation of progress over time. therefore, the national election studies used by eder and jedinger (2019) are employed for the current paper as an example, since they represent a selection of highly relevant and well-curated social science datasets. to illustrate improvement measures, a single dataset from the gesis data archive is selected from the same area of election studies, namely the politbarometer 2022 (forschungsgruppe wahlen, mannheim, 2023). based on workflows for using the ddi-codebook and ddi-lifecycle metadata standards, gesis has provided good quality documentation for archived datasets for many years (akdeniz and zenkmöltgen, 2017; perry et al., 2019; zenk-möltgen, 2012; zenk-möltgen, 2023). recent work has shown that improvements can be made when using the fair criteria and automatic assessments like the fuji tool (saldanha bach et al., 2023). in addition, it needs to be discussed which of these improvements are simply technical and which do really contribute to higher quality of documentation and thus contribute to more transparent science. given these considerations, this paper will look at election data as an example for the social science domain and will focus on the following research questions: • q1: are there differences between the evaluation of fair criteria with automated tools as compared to manual procedures? • q2: has the fairness of research data changed considerably over the six-year period between 2018 to 2024? • q3: can data archives contribute to transparent science by implementing measures to increase fair scores for research data? fairness evaluation for an assessment of the fair criteria, eder and jedinger (2019) look at eighteen large-scale election studies from western democracies, which cover at least two elections, sample the whole voting population and are mainly conducted for academic purposes (see table 1). they operationalize several criteria within each of the four sections, and present tables with explanations for each of the fair scores, using zero if a criterion is not fulfilled, and one as fulfilled (sometimes also using 0.5 for partly https://doi.org/10.29173/iq1119 4/22 zenk-möltgen, wolfgang (2025) the role of fair principles in high-quality research data documentation, iassist quarterly 49(2), pp. 1-22. doi: https://doi.org/10.29173/iq1119 fulfilled). for all the national election studies, they provide the percent of studies fulfilling a score, but also as a summary index for each of the four criteria (summing up all operationalized scores of an area) (eder and jedinger, 2019, tbls. 8–11). however, they do not calculate an overall fairness score by combining the four indexes. id study abbreviation 1 american national election studies (icpsr) anes 2 australian election study aes 3 austrian national election study autnes 4 belgian national election study bnes 5 british election study (uk data archive) bes 6 canadian election study ces 7 danish national election study dnes 8 dutch parliamentary election study dpes 9 estonian national election study enes 10 finnish national election study fnes 12 german longitudinal election study gles 13 hellenic national election studies elnes 14 italian national election study itanes 16 icelandic national election study icenes 17 new zealand election study nzes 18 norwegian election studies nes 21 swedish national election studies snes 22 swiss electoral studies selects table 1: list of national election studies the overall results from this study in 2018 show that findability of the national election studies is rather good, with an overall mean for the findability index of 5.3 (sd=1.88, min=2.0, max=8.0) (eder & jedinger, 2019, p. 661). assessing the accessibility, eder and jedinger find: ‘most studies perform well (…). however, there seems to be room for improvement in regard to providing information on variables omitted (…) and on how to access variables in cases in which country-specific rules (…). additionally, some studies could benefit from providing more extensive reports’ (eder & jedinger, 2019, p. 663). they calculate a mean for the accessibility index of 5.44 (sd=1.79, min=3.0, max=9.0). for interoperability, the authors state that ‘the provision of metadata is very good, with very little need for improvement’, and note a mean for the interoperability index of 11.94 (sd=1.51, min=7.0, max=13.0) (eder & jedinger, 2019, p. 664). the authors conclude that reusability is ‘quite satisfying’ with a mean for the reusability index of 4.31 (sd=0.84, min=2.0, max=5.0) (eder & jedinger, 2019, p. 665). method and data to answer the first research question ‘is the evaluation of fair criteria with automated tools an alternative to manual procedures?’, i look at the same national election studies mentioned in the study by eder and jedinger (2018). the eder and jedinger study used national election studies (see table 1) that were available through 2016. in a next step i searched for the persistent identifier (in this case a doi, digital object identifier) for the latest data availability and performed a f-uji test for each of these studies. the f-uji tool was used for the automated tests because it is among the most prominent tools, it provides detailed documentation for the test results, interprets the persistent identifiers as identifiers for the data (sun et al., 2022), and because it has been found to correspond https://doi.org/10.29173/iq1119 5/22 zenk-möltgen, wolfgang (2025) the role of fair principles in high-quality research data documentation, iassist quarterly 49(2), pp. 1-22. doi: https://doi.org/10.29173/iq1119 well with manual assessments (gehlen et al., 2022). a comparison was then done to show if the manually created results reported to summarize the situation in 2018 match these of the automated evaluation of the f-uji tool for the 2018 datasets. using the scores provided by eder and jedinger (2019), a sum for an overall fairness score was computed in addition to the single scores for each of the four criteria. to compare the results of the manual assessments to the results of the conducted f-uji tests, the mean percentage of the maximum possible value was calculated for each election study. the frequencies for ten quantiles of the percentages were then plotted and compared to allow an overview of the distribution of f-uji percentages. in addition, two-sided t-tests allowed for assessing the significance of the mean differences. the shapiro-wilk test for normal distribution of the differences was checked as well. because summing up four criteria with each one using a different maximum value leads to some bias towards the criteria with higher maximum values, the same procedure was applied to produce a fairness index using equal weight for each of the four criteria. in that way, differences between manual and automatic assessments of the same datasets were analyzed. to answer the second research question ‘has the fairness of research data changed considerably over the six-year period between 2018 to 2024?’, i updated the list of studies with newer rounds available until 2024. for seventeen of the eighteen election studies, a newer dataset was found, the exception being the estonian national election study (enes). for four countries (i.e., belgium, italy, norway, sweden), module 5 of the comparative study for electoral system (cses) was used for the newer waves. the f-uji tests were repeated for all available newer datasets (seventeen of eighteen), and scores for the single f.a.i.r. criteria (i will use the term f.a.i.r. to indicate the single scores for findable, accessible, interoperable, and re-usable) as well as for the combined fairness score and equally weighted fairness score were recorded. to get the data for the comparison to 2024, the list of available persistent identifiers (dois) was researched for the newer datasets. again, the f-uji tests were calculated, and results were saved for analysis. a comparison of values for each of the four f.a.i.r. components, the overall fairness value as well as the equal weight fairness value was performed in the same way. this involved again plotting the distribution of frequencies for ten quantiles, this time for the automatic f-uji tests for the 2018 datasets and the tests for the 2024 datasets. also, t-tests were again used to find significant differences in the mean percentage values, and shapiro-wilk tests were conducted for assessing normal distribution. to answer the third research question ‘can data archives contribute to transparent science by implementing measures to increase fair scores for research data?’ i use another example dataset, the politbarometer 2022 study (forschungsgruppe wahlen, mannheim, 2023). the politbarometer survey series has been conducted since 1977 for the german tv network zdf, and this dataset contains aggregated annual data from 1977 through 2022. the gesis data archive creates aggregate datasets, documents the data, provides data access, and performs long-term archiving for this series, and also for many others (schumann & mauer, 2013; recker et al., 2017). data and metadata can be accessed (hienert et al., 2019) at the gesis search webpage9 as well as through the gesis politbarometer project webpage10. https://doi.org/10.29173/iq1119 6/22 zenk-möltgen, wolfgang (2025) the role of fair principles in high-quality research data documentation, iassist quarterly 49(2), pp. 1-22. doi: https://doi.org/10.29173/iq1119 since the f-uji tool assessment relies on machine-readable links for data and metadata, results are very much dependent on the availability of standardized metadata. since gesis creates standardized documentation using the ddi standards11 and controlled vocabularies12 for archived studies (akdeniz & zenk-möltgen, 2017), it was expected that the f-uji assessment score will be high. it turned out that f-uji scores were not as high as expected. to address this issue, gesis has been working on several levels to improve the technical implementation of the standards that are used by the f-uji tool for assessment (saldanha bach et al., 2023). by comparing the f-uji tool assessment of this example politbarometer dataset from an initial test on 3rd november 2023 with a later test on 29th april 2024, the effects of technical improvements, as well as effects of standardized metadata can be seen. this uses the list of metrics as applied by the f-uji tool and describes the changes involved to improve the evaluation. the data and scripts for all analyses have been deposited at a trusted repository and are made available for secondary research (zenk-möltgen, 2024). the analysis results and figures were produced using stata 18. results comparing manual and automatic fairness scores a persistent identifier was found for the respective last wave for four of the eighteen election studies investigated, even if in the original study by eder and jedinger recorded no persistent identifier (danish, estonian, icelandic, and norwegian election studies). this is already an improvement and will also help with the f-uji test that needs this persistent identifier for evaluation. only the italian and the canadian election studies do not provide a persistent identifier. the results from the f-uji tool assessment of the national election studies in 2018 (see table 2, autom. 2018) show that the mean values differ quite a lot: findability is quite high (5.78 of 7), and also interoperability is high (2.44 of 4), but accessibility is quite low (1.31 of 3) as well as reusability (3.94 of 10). the overall fairness score mean is 11.97 of 24, showing that the election studies get, on average, only half of the scores that they could get. using the equally weighted value for the overall fairness score, a mean of 13.60 of 24 is only slightly better. we can say that the datasets of election studies conducted until 2018 get only mediocre fairness values from the f-uji assessment. https://doi.org/10.29173/iq1119 7/22 zenk-möltgen, wolfgang (2025) the role of fair principles in high-quality research data documentation, iassist quarterly 49(2), pp. 1-22. doi: https://doi.org/10.29173/iq1119 national election studies fscore ascore iscore rscore fairness score fairness score (equal weight) manual 2018 mean 5.33 5.44 11.94 4.31 27.03 26.92 stddev 1.88 1.79 1.51 .84 4.48 4.65 max 8 10 13 5 36 36 n 18 18 18 18 18 18 autom. 2018 mean 5.78 1.31 2.44 3.94 11.97 13.60 stddev 0.75 0.54 1.15 1.06 5.17 3.37 max 7 3 4 10 24 24 n 16 16 16 16 18 16 table 2: comparing manual and automatic scores for the f.a.i.r. and overall fairness of the national election studies (values for ‘manual 2018’ calculated from eder and jedinger 2018) comparing the distribution of values between the 2018 manual and 2018 automatic assessment can best be done by using percentage values because the maximum values differ for each assessment. the results show (see figure 1) that findability is higher for the automatic assessment. automatic assessed scores for accessibility are rather low, and interoperability is rated at mixed levels for the national election studies. several studies get much lower scores for interoperability with the automatic assessment than with the manual one. reusability is also at quite a low level compared to the manual assessment. figure 1: comparing the f.a.i.r percent scores from manual and automatic assessment for the national election studies 2018 we can see in figure 1 that only for findability does the automatic assessment by the f-uji tool (lilac series) yield higher values than the manual assessments by eder and jedinger (red series). for accessibility, and very clearly also for interoperability and reusability, values for most election studies are lower with the automatic assessment than with the manual assessment. https://doi.org/10.29173/iq1119 8/22 zenk-möltgen, wolfgang (2025) the role of fair principles in high-quality research data documentation, iassist quarterly 49(2), pp. 1-22. doi: https://doi.org/10.29173/iq1119 the overall fairness scores for the automatic evaluation (see figure 2) are medium and lower compared to the manual assessed values. giving equal weight to the four criteria does not change this picture very much. this indicates that overall fairness values will yield lower scores with the automatic f-uji tool assessment than when using a manual approach. figure 2: comparing the overall fairness percent scores from manual and automatic assessment for the national election studies 2018 to see if the differences between the manual results and the automatic results of the mean values are significant, two-tailed t-tests were conducted (see table 3, comparison a-b). since the assumption of normal distribution is required for that, a shapiro-wilk test was performed for each comparison. national election studies f-% a-% i-% r-% fairness % fairness % (equal weight) a manual 2018 mean .695 .544 .947 .875 .769 .765 stddev .233 .167 .067 .165 .105 .113 n 16 16 16 16 16 16 b autom. 2018 mean .826 .438 .609 .394 .561 .567 stddev .107 .181 .288 .106 .123 .140 comparison a-b m diff -.131 .106 .338 .481+ .208 .199 t -1.904 1.575 4.425 8.211 4.429 3.759 df 15 15 15 15 15 15 p 0.076 0.136 0.001 <0.001 0.001 0.002 table 3: comparing manual and automatic f.a.i.r. and overall fairness mean percentage values and t-tests ( + shapiro–wilk test is significant p<0.01 and indicates that the data is not normally distributed for this comparison) the results for the findability of the 2018 datasets with (a) manual evaluation (m=0.695, sd=0.233) and (b) automatic evaluation (m=0.826, sd=0.107) show that findability is higher for the automatic assessment, t(15)=-1.904, p=0.076, significant at the 0.10-level. results for accessibility, interoperability, and reusability show lower values for the automatic assessment by the f-uji tool than with the manual assessment (see table 3, comparison a-b), however, the difference for accessibility is not significant. the higher values for mean interoperability and mean reusability with the manual assessment are significant at the 0.01-level. also, the overall percentage scores for fairness differ significantly for the (a) manual assessment (m=0.769, sd=0.105) from the (b) automatic assessment https://doi.org/10.29173/iq1119 9/22 zenk-möltgen, wolfgang (2025) the role of fair principles in high-quality research data documentation, iassist quarterly 49(2), pp. 1-22. doi: https://doi.org/10.29173/iq1119 (m=0.561, sd=0.123), t(15)=4.429, p=0.001. the same is true for the equal weight fairness percentage. these results suggest that the answer to research question q1 ‘are there differences between the evaluation of fair criteria with automated tools compared to manual procedures?’ is that in this case, the automatic evaluation with the f-uji tool gives lower scores than the manual evaluation, except for the criterium of findability, where we see somewhat higher scores for the automatic evaluation. comparing automatic fairness scores for 2018 and 2024 datasets the results from the f-uji tool assessment of the more recent national election studies in 2024 (see table 4, autom. 2024) show that the mean values also differ a lot: findability is again quite high (5.91 of 7), and also interoperability is high (2.76 of 4), but accessibility is again quite low (1.47 of 3) as well as reusability (4.41 of 10). the overall fairness score mean is somewhat better with 13.75 of 24, showing that the newer election studies get, on average, more than half of the scores that they could get. using the equally weighted value for the overall fairness score, a mean of 14.80 of 24 is again slightly better. all these values are slightly higher than the automatic assessed ones from 2018. we can say that the newer election studies datasets get moderate percentage scores, and to a certain degree better fairness values from the f-uji assessment tool than the 2018 datasets. national election studies fscore ascore iscore rscore fairness score fairness score (equal weight) autom. 2018 mean 5.78 1.31 2.44 3.94 11.97 13.60 stddev 0.75 0.54 1.15 1.06 5.17 3.37 max 7 3 4 10 24 24 n 16 16 16 16 18 16 autom. 2024 mean 5.91 1.47 2.76 4.41 13.75 14.80 stddev 0.85 0.57 1.15 1.62 4.66 3.42 max 7 3 4 10 24 24 n 17 17 17 17 18 17 table 4: comparing 2018 and 2024 scores from the f-uji assessment of the national election studies comparing the distribution of values between the automatic assessments by the f-uji tool for the 2018 and 2024 datasets can again be done by using percentage values. the results show (see figure 3) that findability is nearly the same for both assessments (except that one study seems to be less findable in 2024). assessment scores for accessibility have slightly improved, and interoperability and reusability are at higher levels for the newer election studies. we can see in figure 3 that for all separate indices of findability, accessibility, interoperability, and reusability the assessments for the newer datasets (yellow series) yield higher values than for the older datasets (green series). for accessibility, interoperability, and reusability, this result is much clearer than for findability. https://doi.org/10.29173/iq1119 10/22 zenk-möltgen, wolfgang (2025) the role of fair principles in high-quality research data documentation, iassist quarterly 49(2), pp. 1-22. doi: https://doi.org/10.29173/iq1119 figure 3: comparing the f.a.i.r percent scores from automatic assessment for the national election studies 2018 and 2024 the overall fairness scores for the evaluation of the newer datasets from up to 2024 (see figure 4) are still moderate but somewhat higher compared to the older assessed datasets. giving equal weight to the four criteria of the fairness index does again show the same picture. this makes clear that there is a slight improvement of the overall fairness values for the national election studies that can be shown by conducting the automatic f-uji tool assessment. figure 4: comparing the overall fairness percent scores from automatic assessment for the national election studies 2018 and 2024 differences between the mean results of the older and newer national election studies were again evaluated for significance with a two-tailed t-test (see table 5, comparison c-d), and the shapiro-wilk test for the assumption of normal distribution. the mean comparison results for all indicators show no significant differences. even if means for the newer datasets are slightly higher for all the single (i.e., findability, accessibility, interoperability, and reusability) indicators, none of these differences https://doi.org/10.29173/iq1119 11/22 zenk-möltgen, wolfgang (2025) the role of fair principles in high-quality research data documentation, iassist quarterly 49(2), pp. 1-22. doi: https://doi.org/10.29173/iq1119 reach a statistically significant level (see table 5, comparison c-d). results for the overall indicators of fairness and equal weight fairness percentage are consistent. national election studies f-% a-% i-% r-% fairness % fairness % (equal weight) c autom. 2018 mean .824 .433 .583 .393 .556 .558 stddev .111 .187 .278 .110 .126 .141 n 15 15 15 15 15 15 d autom. 2024 mean .833 .478 .667 .420 .589 .599 stddev .123 .198 .294 .142 .130 .141 comparison c-d m diff -.010 -.044 -.083+ -.027 -.033 -.041 t -0.397 -0.745 -1.581 -0.745 -1.240 -1.401 df 14 14 14 14 14 14 p 0.698 0.469 0.136 0.469 0.235 0.183 table 5: comparing 2018 and 2024 f.a.i.r. and overall fairness mean percentage values and t-tests ( + shapiro–wilk test is significant p<0.01 and indicates that the data is not normally distributed for this comparison) these findings indicate that the answer to the second research question ‘has the fairness of research data changed considerably over the six-year period between 2018 to 2024?’ is that six year later, there is a rather small improvement in the fairness of national election. improving fairness scores research question q3 ‘can data archives contribute to transparent science by implementing measures to increase fair scores for research data?’ requires examining measures for improving the fairness scores and a discussion of the contribution of these measures to transparent science. the politbarometer 2022 study (forschungsgruppe wahlen, mannheim, 2023) is used to illustrate some of the implemented measures. this study is well suited because it has a comprehensive documentation, both in english and german, and is well known and used by many political scientists. table 6 lists for each section of the fair criteria the defined metrics and the number of technical tests performed by the f-uji tool (see table 6, blue section). for sixteen of the seventeen metrics, tests are included. the table also lists the scores that the f-uji test for the politbarometer 2022 study achieved initially on november 3rd, 2023, before implementing dedicated measures to support the fairness assessment (see table 6, orange section). after several technical adaptations and changes in the metadata were implemented, a final f-uji test for the politbarometer 2022 study was conducted on april 29th, 2024, and those results are listed also (see table 6, yellow section). for an easier overview, metrics that improved are indicated by an x in the column titled ‘improvement’. overall scores for each of the f.a.i.r criteria and the fairness scores and percent values are shown in the ‘total’ columns and rows. https://doi.org/10.29173/iq1119 12/22 zenk-möltgen, wolfgang (2025) the role of fair principles in high-quality research data documentation, iassist quarterly 49(2), pp. 1-22. doi: https://doi.org/10.29173/iq1119 table 6: f-uji test comparison on nov 3rd, 2023, and apr 29th, 2024 for politbarometer 2022, https://doi.org/10.4232/1.14103 https://doi.org/10.29173/iq1119 https://doi.org/10.4232/1.14103 13/22 zenk-möltgen, wolfgang (2024) the role of fair principles in high-quality research data documentation, iassist quarterly xx(y), pp. 1-5. doi: https://doi.org/do.be/doo altogether, fairness scores were improved from 12.5 to 18 points out of a maximum of 24, resulting in an increase from 52.1% to 75% (see table 6). this was accomplished by changing data and metadata for six of the seventeen metrics. these six improvements are very generic and do not apply only to this specific study, as will become clear in the following section. the improved fairness score for the politbarometer 2022 dataset moved the study from the ‘moderate’ score category to the ‘advanced’ fairness score category. the findability score for the politbarometer 2022 dataset was improved from 5 to 6 with this measure, leading to the evaluation category ‘advanced’ instead of ‘moderate’. the accessibility score for the dataset was increased from 1.5 to 3, leading to the evaluation category of ‘advanced’ instead of ‘initial’. the interoperability score for this dataset was amended from 2 to 3, leading to an evaluation of ‘advanced’ instead of ‘moderate’. the reusability score for the politbarometer 2022 dataset increased from 4 to 6 with these measures, leading to the evaluation category ‘moderate’ instead of ‘initial’. one general improvement was needed to allow the f-uji tool to access the metadata: because the gesis search webpage is a single page application that uses javascript for asynchronously loading content after the page has already been delivered to the client (ajax), the f-uji tool initially could not extract the schema.org metadata that is provided at each doi landing page. for that reason, each landing page was modified so that it contains a fair signposting13 ‘described-by’-link for retrieving the schema.org metadata in json-ld format, which is supported by the f-uji tool (and other tools). the metadata of the example politbarometer 2022 dataset is displayed in figure 5. subsequent improvements for the f-uji tool assessment were dependent on this first change in the architecture of the gesis search webpage. https://doi.org/do.be/doo 14/22 zenk-möltgen, wolfgang (2024) the role of fair principles in high-quality research data documentation, iassist quarterly xx(y), pp. 1-5. doi: https://doi.org/do.be/doo figure 5: metadata in json-ld format for politbarometer 2022, https://doi.org/10.4232/1.14103 findability in the area of findability, f-uji test scores were improved for: fsf-f3-01m-1 metadata contains data content related information (file name, size, type). fsf-f3-01m-2 metadata contains a pid or url which indicates the location of the downloadable data content. since data downloads are already part of the gesis search webpage, the information about downloadable files, their name, format, size, and url were available. this metadata was included into the json-ld metadata, allowing the f-uji tool to extract this information. having the name, size and type specified led to improvement in the score of the first metric. providing the url for accessing the dataset resulted in the second improvement, even when the download is possible only for registered users. the improvement in the findability score is relevant for all studies with downloadable data files which are the majority of studies on the gesis data archive. accessibility https://doi.org/do.be/doo https://doi.org/10.4232/1.14103 15/22 zenk-möltgen, wolfgang (2024) the role of fair principles in high-quality research data documentation, iassist quarterly xx(y), pp. 1-5. doi: https://doi.org/do.be/doo in the accessibility category the f-uji test scores were improved for: fsf-a1-01m-2 data access information is machine readable. fsf-a1-03d-1 metadata includes a resolvable link to data based on standardized web communication protocols. the first improvement for data access information consists of specifying the json-ld fields ‘conditionsofaccess’ (free text) and ‘isaccessibleforfree’ (controlled vocabulary). this allowed the fuji tool to extract the data access type. providing this information was easy since this metadata is available for all studies at the gesis data archive. the second improvement was solved by the implementation of the findability measure of providing a url to the downloadable dataset (see above). a url starting with ‘https’ is recognized by the f-uji tool as a standard web-protocol and therefore this improvement affects all studies with provided data download links. interoperability in the interoperability category, f-uji test scores were improved for: fsf-i3-01m-1 related resources are explicitly mentioned in metadata. to achieve this improvement rdf fair signposting link headers were added into the html source of the landing page. this technique helps to identify data resources that are machine-actionable on the web, and it allows to specify metadata independently from a specific fair assessment tool. this change was implemented for all studies archived at the gesis data archive. as a result, the f-uji tool assessments were improved. reusability for the reusability category, f-uji test scores were enhanced for: fsf-r1-01md-2 verifiable data descriptors (file info, measured variables or observation types) are specified in metadata. fsf-r1.3-02d-1 the format of a data file given in the metadata is listed in the long-term file formats, open file formats or scientific file formats controlled list. the measure for the first improvement is again due to the findability improvement (see above): providing the data download url also improves this score for interoperability. again, studies with available data downloads benefit from this improvement. the second improvement was possible because the gesis data archive provides most data files as spss and/or stata files. the formats for downloadable data files (provided already for the findability improvement) need to be specified as mime types, and both ‘application/x-stata-dta’ (stata) and ‘application/x-spss-sav’ (spss) formats are recognized as scientific community standards by the f-uji tool. this improvement applies to nearly all studies with downloadable data files, given they have an spss and/or stata file available. https://doi.org/do.be/doo 16/22 zenk-möltgen, wolfgang (2024) the role of fair principles in high-quality research data documentation, iassist quarterly xx(y), pp. 1-5. doi: https://doi.org/do.be/doo summary overall, fairness score for the politbarometer 2022 data that were used as a case study, was improved from 12.5 to 18 points out of a maximum of 24, resulting in an increase from 52.1% to 75% (see table 6). it becomes clear from the example dataset and the explanations that especially the first adaptations for findability are important for allowing the f-uji tool to better assess the research data. for accessibility and reusability, the same measure is relevant and increases the f-uji score of the evaluated resource also for these sections. the increase in fairness score was a result of several modifications: providing the downloadable research data files in standardized form with their name, format as mime type, size, and url, adding the access conditions in machine readable form, and including rdf fair signposting link headers. further modifications were discussed, but not implemented, e.g., using a data license that is part of spdx14 would have had an improvement effect for the accessibility score fsf-a1-01m-3 (‘data access information is indicated by (not machine readable) standard terms’). existing usage regulations15 for the gesis data archive do not allow the re-distribution of data by researchers themselves, and therefore cannot be used with any of the creative commons or other open licenses. however, these regulations allow the use of data for research purposes without any further restrictions (depending on access class) and should therefore still be considered as enabling transparent research practices. other modifications are still under evaluation, e.g., implementing a formal representation of prov-o metadata that would improve metric fsf-r1.2-01m-2 (‘metadata contains provenance information using formal provenance ontologies (prov-o)’). discussion and conclusions it has been shown that the evaluation of fair criteria can be done with automated tools as an alternative to manual procedures. however, the resulting fairness levels are not comparable, and the differences for the mean assessment between both methods are statistically significant (except for accessibility). in the case of using the f-uji tool for an automatic evaluation of the national election studies, the resulting values are higher for findability, and lower for accessibility, interoperability, and reusability than with the manual assessment. regarding the change that might have happened during the last six years, the f-uji tool was used for an automated fairness assessment of the national election studies data. it has been demonstrated that there was not much change with the level of fairness from the datasets available in 2018 and those available in 2024. slight improvements have been found, and for single studies there might have even been considerable improvements – especially in the criteria of accessibility, interoperability, and reusability. overall, there were no significant differences in the average level of fairness scores of the newer datasets compared to the older ones. potentially, changes in the data documentation or website updates between 2018 and 2024 might have resulted in improvements for single studies that were not detected by this comparison (because the metrics were all assessed in 2024). however, this kind of improvement would have contributed to the difference between manual and automatic evaluation examined in the first step. the higher findability score identified in this comparison may explain some of these findings. https://doi.org/do.be/doo 17/22 zenk-möltgen, wolfgang (2024) the role of fair principles in high-quality research data documentation, iassist quarterly xx(y), pp. 1-5. doi: https://doi.org/do.be/doo the study has several limitations that should be mentioned. firstly, the evaluation used a selection of very prominent and widely recognized national election studies which are carefully curated and documented. results might look different when including studies that are not as prominent, have fewer resources for data curation and documentation, or are not survey data at all. secondly, several metrics rely on technical or metadata standards that are currently not commonly used, such as fair signposting or spdx licenses. with an uptake of those standards for web resources, research data might also benefit. custom implementations of those standards require technical and metadata developments, and even large institutions may need years to realize this. with the example of the politbarometer 2022 dataset, the paper describes several modifications that led to improved fairness scores as measured by the f-uji tool. this demonstrates that small technical changes can have a substantial impact. it may be worth mentioning that the implementation of these changes did not require a lot of work by the archive team. given the significant improvement in fairness scores, it should be evident that the investment in providing metadata in a suitable technical format is worth the effort. but how do these technical improvements contribute to more transparent research practices, as asked in research question three? christensen, freese and miguel (2019, p. 12) link the question ‘what is ethical research?’ back to the foundations laid by robert k. merton in 1942 (merton, 1942). they write that ‘openness, integrity, and transparency are at the very heart of merton’s influential articulation of scientific research norms’ (christensen et al., 2019, p. 20). however, they also describe the gap between current research practices and the scientific ideal and formulate recommendations for research practices in four areas: reporting standards, replication, data sharing, and reproducible workflow (christensen et al., 2019, p. 141ff). except for reporting standards, those areas are also relevant when performing fairness tests: replication is only possible with data sharing, which in turn means that the research data is fair: findable, accessible, interoperable, and reusable. reproducible workflows become possible when fair data is the reality. thus, improving the fairness of research data is an essential foundation for ethical research practices. a certain degree of fairness can also be achieved by using unstandardized methods. making research data available upon request or on a website, describing it in a report, and using a commonly used data format may already be a first step for some researchers. but given the developments in data science and big data, the use of social media data to analyze social phenomena, and given the increasing variety, volume, and velocity of change, known as the ‘three v’ (kockum & dacre, 2021), a machineactionable approach is needed (jensen et al., 2019; weller & kinder-kurlanda, 2021; weller & strohmaier, 2014). further on are developments of artificial intelligence in the social sciences that rely largely on machine-accessible data and are increasingly being used (grossmann et al., 2023). supporting researchers working in such environments requires a research infrastructure that delivers fair data services. fairness indicators provide a standardized way to assess the quality of documentation in this respect. the f-uji tool for assessing the fairness of research data makes it possible to document the fairness of research data in a transparent way. as shown in this article, the basis for fairness evaluations in an automated procedure is usage of open standards that are applied according to principles of open science. the openness can be seen when the f-uji tool performs the assessment based on data and https://doi.org/do.be/doo 18/22 zenk-möltgen, wolfgang (2024) the role of fair principles in high-quality research data documentation, iassist quarterly xx(y), pp. 1-5. doi: https://doi.org/do.be/doo metadata that is freely available on the web. anyone can access the basis for evaluation. the integrity is represented by several checks of the f-uji tool which are validations of the provided metadata, e.g., fsf-r1-01md-3 (‘data content matches file type and size specified in metadata’), fsf-r1-01md-4 (‘data content matches measured variables or observation types specified in metadata’) or fsf-f201m-3 (‘core descriptive metadata is available’ checks for defined metadata fields). the transparency of the process is enabled by the availability of the f-uji tool itself as an open source on github under the mit open-source license16 , making criteria for evaluation visible, and allowing users to scrutinize each single metric and test included. data archives and other research data centers can contribute considerably to the level of fairness of the research data they curate and disseminate, as has been shown by the example dataset. however, there is more to ethical research practices than implementing measures to achieve the highest score of research data fairness. additional things to consider include for example reporting standards and transparency in methods for creating scientific results. archives and other data curating institutions can provide scientific data with the highest possible fairness levels, and thus can provide one piece in the mosaic of the scientific ideal. acknowledgements i would like to thank my colleagues at gesis who have made this research possible. christina eder and alexander jedinger have initiated this investigation with their work from 2018. daniel hienert, zhang yudong, janete saldanha bach, brigitte mathiak, and farah karim have all worked on major items for the practical and systematic implementation, and without this my work would not have been possible. and i thank the anonymous reviewers, the editors, christina eder and alexander jedinger, but especially libby bishop for many helpful comments on the manuscript. references akdeniz, e. and zenk-möltgen, w. (2017). ‘ddi-lifecycle at the data archive: the metadata schema for documentation in different software tools’, gesis papers, 2017(18). https://doi.org/10.21241/ssoar.52487 alaterä, t., kleemola, m., ala-lahti, h. and jerlehag, b. (2022). d4.5 report on completed fair data standard adoption and certifications of data repositories in the region. https://doi.org/10.5281/zenodo.7303538 american economic association (2024). data and code availability policy. https://www.aeaweb.org/journals/data/data-code-policy betancort cabrera, n., bongartz, e.c., dörrenbächer, n., goebel, j., kaluza, h. and siegers, p. (2020). white paper on implementing the fair principles for data in the social, behavioural, and economic eciences’, ratswd working paper series. https://doi.org/10.17620/02671.60 bishop, b. w., & hank, c. (2018). measuring fair principles to inform fitness for use. international journal of digital curation, 13(1), 35–46. https://doi.org/10.2218/ijdc.v13i1.630 borgesius, f.z., gray, j. and van eechoud, m. (2016). open data, privacy, and fair information principles: towards a balancing framework’, berkeley technology law journal, 30(3), 20732131. https://doi.org/10.15779/z389s18 . https://doi.org/do.be/doo https://doi.org/10.21241/ssoar.52487 https://doi.org/10.5281/zenodo.7303538 https://www.aeaweb.org/journals/data/data-code-policy https://doi.org/10.17620/02671.60 https://doi.org/10.2218/ijdc.v13i1.630 https://doi.org/10.15779/z389s18 19/22 zenk-möltgen, wolfgang (2024) the role of fair principles in high-quality research data documentation, iassist quarterly xx(y), pp. 1-5. doi: https://doi.org/do.be/doo christensen, g.s., freese, j. and miguel, e. (2019). transparent and reproducible social science research: how to do open science. university of california press. da-rt (2015). ‘data access & research transparency’. https://www.dartstatement.org/about deutsche gesellschaft für soziologie (2019). ‘bereitstellung und nachnutzung von forschungsdaten in der soziologie’. https://soziologie.de/aktuell/stellungnahmen/news/bereitstellung-undnachnutzung-von-forschungsdaten-in-der-soziologie devaraju, a., & huber, r. (2021). an automated solution for measuring the progress toward fair research data. patterns, 2(11), 100370. https://doi.org/10.1016/j.patter.2021.100370 devaraju, a., huber, r., mokrane, m., herterich, p., cepinskas, l., de vries, j., l’hours, h., davidson, j. and white, a. (2022). fairsfair data object assessment metrics. https://doi.org/10.5281/zenodo.6461229 devaraju, a., mokrane, m., cepinskas, l., huber, r., herterich, p., de vries, j., akerman, v., l’hours, h., davidson, j., & diepenbroek, m. (2021). from conceptualization to implementation: fair assessment of research data objects. data science journal, 20, 4. https://doi.org/10.5334/dsj-2021-004 eder, c. and jedinger, a. (2018). fair national election studies: how well are we doing? (gesis sdn10.7802-1761) [data set]. gesis. https://doi.org/10.7802/1761 eder, c., & jedinger, a. (2019). fair national election studies: how well are we doing? european political science, 18(4), 651–668. https://doi.org/10.1057/s41304-018-0194-3 forschungsgruppe wahlen, mannheim (2023). politbarometer 2022 (cumulated data set). (gesis, za7970; version 1.0.0) [data set]. gesis. https://doi.org/10.4232/1.14103 freese, j., & peterson, d. (2017). replication in social science. annual review of sociology, 43(1), 147–165. https://doi.org/10.1146/annurev-soc-060116-053450 gehlen, k. p., höck, h., fast, a., heydebreck, d., lammert, a., & thiemann, h. (2022). recommendations for discipline-specific fairness evaluation derived from applying an ensemble of evaluation tools. data science journal, 21, 7. https://doi.org/10.5334/dsj-2022007 grossmann, i., feinberg, m., parker, d. c., christakis, n. a., tetlock, p. e., & cunningham, w. a. (2023). ai and the transformation of social science research. science, 380(6650), 1108–1109. https://doi.org/10.1126/science.adi1778 guillot, p., bøgsted, m., & vesteghem, c. (2023). fair sharing of health data: a systematic review of applicable solutions. health and technology, 13(6), 869–882. https://doi.org/10.1007/s12553-023-00789-5 hienert, d., kern, d., boland, k., zapilko, b., & mutschke, p. (2019). a digital library for research data and related information in the social sciences. 2019 acm/ieee joint conference on digital libraries (jcdl), 148–157. https://doi.org/10.1109/jcdl.2019.00030 https://doi.org/do.be/doo https://www.dartstatement.org/about https://soziologie.de/aktuell/stellungnahmen/news/bereitstellung-und-nachnutzung-von-forschungsdaten-in-der-soziologie https://soziologie.de/aktuell/stellungnahmen/news/bereitstellung-und-nachnutzung-von-forschungsdaten-in-der-soziologie https://doi.org/10.1016/j.patter.2021.100370 https://doi.org/10.5281/zenodo.6461229 https://doi.org/10.5334/dsj-2021-004 https://doi.org/10.7802/1761 https://doi.org/10.1057/s41304-018-0194-3 https://doi.org/10.4232/1.14103 https://doi.org/10.1146/annurev-soc-060116-053450 https://doi.org/10.5334/dsj-2022-007 https://doi.org/10.5334/dsj-2022-007 https://doi.org/10.1126/science.adi1778 https://doi.org/10.1007/s12553-023-00789-5 https://doi.org/10.1109/jcdl.2019.00030 20/22 zenk-möltgen, wolfgang (2024) the role of fair principles in high-quality research data documentation, iassist quarterly xx(y), pp. 1-5. doi: https://doi.org/do.be/doo jensen, u., netscher, s. and weller, k. (eds) (2019). forschungsdatenmanagement sozialwissenschaftlicher umfragedaten: grundlagen und praktische lösungen für den umgang mit quantitativen forschungsdaten. verlag barbara budrich. key, e. m. (2016). how are we doing? data access and replication in political science. political science and politics, 49(2), 268–272. https://doi.org/10.1017/s1049096516000184 kockum, f., & dacre, n. (2021). project management volume, velocity, variety: a big data dynamics approach. advanced project management, 21(1). https://doi.org/10.2139/ssrn.3813838 maxwell, l., shreedhar, p., dauga, d., mcquilton, p., terry, r., denisiuk, a., molnar-gabor, f., saxena, a. and sansone, s.-a. (2021). fair, ethical, and coordinated data sharing for covid19 response: a review of covid-19 data sharing platforms and registries. https://doi.org/10.21203/rs.3.rs-1045632/v1 merton, r.k. (1942). a note on science and democracy. journal of legal and political sociology, 1(1), 115–126. musen, m. a., o’connor, m. j., schultes, e., martínez-romero, m., hardi, j., & graybeal, j. (2022). modeling community standards for metadata as templates makes data fair. scientific data, 9(1), 696. https://doi.org/10.1038/s41597-022-01815-3 perry, a., & netscher, s. (2022). measuring the time spent on data curation. journal of documentation, 78(7), 282–304. https://doi.org/10.1108/jd-08-2021-0167 perry, a., watteler, o., zenk-möltgen, w. and gregory, a. (2019, may 27-31). how can research projects benefit from standardized metadata like ddi? [conference presentation]. iassist 2019, sydney, australia. https://doi.org/10.5281/zenodo.3612730 petrosyan, l., aleixandre-benavent, r., peset, f., valderrama-zurián, j. c., ferrer-sapena, a., & sixtocostoya, a. (2023). fair degree assessment in agriculture datasets using the f-uji tool. ecological informatics, 76, 102126. https://doi.org/10.1016/j.ecoinf.2023.102126 recker, j., zenk-möltgen, w., & mauer, r. (2017). applications of research data management at gesis data archive for the social sciences. in j. b. thestrup & f. kruse (eds.), research data management—a european perspective (pp. 119–146). de gruyter. https://doi.org/10.1515/9783110365634-008 saldanha bach, j., klas, c.-p., mathiak, b., yudong zhang and mutschke, p. (2023). fairness assessment: a comparison of the rda model and the f-uji automated tool report. https://doi.org/10.5281/zenodo.8308902 schumann, n. and mauer, r. (2013). the gesis data archive for the social sciences: a widely recognised data archive on its way. international journal of digital curation, 8(2), 215–222 https://doi.org/10.2218/ijdc.v8i2.285 sofi-mahmudi, a. and raittio, e. (2022). transparency of covid-19 related research in dental journals. frontiers in oral health, 3, p. 871033. available at: https://doi.org/10.3389/froh.2022.871033 https://doi.org/do.be/doo https://doi.org/10.1017/s1049096516000184 https://doi.org/10.2139/ssrn.3813838 https://doi.org/10.21203/rs.3.rs-1045632/v1 https://doi.org/10.1038/s41597-022-01815-3 https://doi.org/10.1108/jd-08-2021-0167 https://doi.org/10.5281/zenodo.3612730 https://doi.org/10.1016/j.ecoinf.2023.102126 https://doi.org/10.1515/9783110365634-008 https://doi.org/10.5281/zenodo.8308902 https://doi.org/10.2218/ijdc.v8i2.285 https://doi.org/10.3389/froh.2022.871033 21/22 zenk-möltgen, wolfgang (2024) the role of fair principles in high-quality research data documentation, iassist quarterly xx(y), pp. 1-5. doi: https://doi.org/do.be/doo stall, s., yarmey, l., cutcher-gershenfeld, j., hanson, b., lehnert, k., nosek, b., parsons, m., robinson, e. and wyborn, l. (2019). make scientific data fair. nature, 570(7759), 27–29. https://doi.org/10.1038/d41586-019-01720-7 sun, c., emonet, v. and dumontier, m. (2022). a comprehensive comparison of automated fairness evaluation tools. in k. wolstencroft, a. splendiani, m.s. marshall, c. baker, a. waagmeester, m. roos, r. vos, r. fijten, & l.j. castro (eds.) 13th international conference on semantic web applications and tools for health care and life sciences, 44-53. https://ceur-ws.org/vol3127/paper-6.pdf. u.s. national science foundation (2018). data management guidance for sbe directorate proposals and awards. https://new.nsf.gov/sbe/data-management weller, k. and kinder-kurlanda, k. (2021). uncovering the challenges in collection, sharing and documentation: the hidden data of social media research?, proceedings of the international aaai conference on web and social media, 9(4), 28–37. https://doi.org/10.1609/icwsm.v9i4.14687 weller, k. and strohmaier, m. (2014). social media in academia: how the social web is changing academic practice and becoming a new source for research data, it information technology, 56(5), 203–206. https://doi.org/10.1515/itit-2014-9002 wilkinson, m.d., dumontier, m., aalbersberg, ij.j., appleton, g., axton, m., baak, a., blomberg, n., boiten, j.-w., da silva santos, l.b., bourne, p.e., bouwman, j., brookes, a.j., clark, t., crosas, m., dillo, i., dumon, o., edmunds, s., evelo, c.t., finkers, r., gonzalez-beltran, a., gray, a.j.g., groth, p., goble, c., grethe, j.s., heringa, j., ’t hoen, p.a.c., hooft, r., kuhn, t., kok, r., kok, j., lusher, s.j., martone, m.e., mons, a., packer, a.l., persson, b., rocca-serra, p., roos, m., van schaik, r., sansone, s.-a., schultes, e., sengstag, t., slater, t., strawn, g., swertz, m.a., thompson, m., van der lei, j., van mulligen, e., velterop, j., waagmeester, a., wittenburg, p., wolstencroft, k., zhao, j. and mons, b. (2016). the fair guiding principles for scientific data management and stewardship’, scientific data, 3(1), 160018. https://doi.org/10.1038/sdata.2016.18 wissenschaftsrat (2020). ‘zum wandel in den wissenschaften durch datenintensive forschung’. https://www.wissenschaftsrat.de/download/2020/8667-20.pdf?__blob=publicationfile&v=5 zenk-möltgen, w. (2012). ‘metadaten und die data documentation initiative (ddi)’, in r. altenhöner and c. oellers (eds.) langzeitarchivierung von forschungsdaten: standards und disziplinspezifische lösungen (pp. 111–126). scivero verl zenk-möltgen, w. (2023, november 27-29) implementing colectica at the gesis data archive [conference presentation].eddi2023, ljubljana, slovenia. https://doi.org/10.5281/zenodo.10257202 zenk-möltgen, w., akdeniz, e., katsanidou, a., naßhoven, v. and balaban, e. (2018.) factors influencing the data sharing behavior of researchers in sociology and political science, journal of documentation, 74(5), 1053–1073. https://doi.org/10.1108/jd-09-2017-0126 zenk-möltgen, w. (2024). replication data and code for: the role of fair principles in high-quality research data documentation: looking at national election studies, (gesis sdn-10.78022798; version 1.0.0) [data set]. gesis. https://doi.org/10.7802/2798 https://doi.org/do.be/doo https://doi.org/10.1038/d41586-019-01720-7 https://ceur-ws.org/vol-3127/paper-6.pdf https://ceur-ws.org/vol-3127/paper-6.pdf https://new.nsf.gov/sbe/data-management https://doi.org/10.1609/icwsm.v9i4.14687 https://doi.org/10.1515/itit-2014-9002 https://doi.org/10.1038/sdata.2016.18 https://www.wissenschaftsrat.de/download/2020/8667-20.pdf?__blob=publicationfile&v=5 https://doi.org/10.5281/zenodo.10257202 https://doi.org/10.1108/jd-09-2017-0126 https://doi.org/10.7802/2798 22/22 zenk-möltgen, wolfgang (2024) the role of fair principles in high-quality research data documentation, iassist quarterly xx(y), pp. 1-5. doi: https://doi.org/do.be/doo endnotes 1 wolfgang zenk-möltgen works at gesis – leibniz institute for the social sciences, germany. he can be reached by email: wolfgang.zenk-moeltgen@gesis.org. https://orcid.org/0000-0002-2158-3941 2 https://zenodo.org/ 3 https://figshare.com/ 4 https://dataverse.harvard.edu/ 5 https://osf.io/ 6 https://force11.org/info/guiding-principles-for-findable-accessible-interoperable-and-reusable-data-publishing-version-b1-0/ 7 https://fairassist.org/#!/ 8 https://www.f-uji.net/ 9 https://search.gesis.org/ 10 https://www.gesis.org/en/elections/politbarometer 11 https://ddialliance.org/ 12 https://ddialliance.org/controlled-vocabularies 13 https://signposting.org/fair/ 14 https://spdx.dev/about/overview/ 15 https://www.gesis.org/fileadmin/user_upload/usage_regulations.pdf 16 https://github.com/pangaea-data-publisher/fuji https://doi.org/do.be/doo mailto:wolfgang.zenk-moeltgen@gesis.org https://orcid.org/0000-0002-2158-3941 https://zenodo.org/ https://figshare.com/ https://dataverse.harvard.edu/ https://osf.io/ https://force11.org/info/guiding-principles-for-findable-accessible-interoperable-and-re-usable-data-publishing-version-b1-0/ https://force11.org/info/guiding-principles-for-findable-accessible-interoperable-and-re-usable-data-publishing-version-b1-0/ https://fairassist.org/#!/ https://www.f-uji.net/ https://search.gesis.org/ https://www.gesis.org/en/elections/politbarometer https://ddialliance.org/ https://ddialliance.org/controlled-vocabularies https://signposting.org/fair/ https://spdx.dev/about/overview/ https://www.gesis.org/fileadmin/user_upload/usage_regulations.pdf https://github.com/pangaea-data-publisher/fuji the role of fair principles in high-quality research data documentation: looking at national election studies abstract keywords introduction literature review research design fairness evaluation method and data results comparing manual and automatic fairness scores comparing automatic fairness scores for 2018 and 2024 datasets improving fairness scores discussion and conclusions acknowledgements references documenting data for secondary analysis : the primary producer's role and responsibility by bridget winstanley ' esrc data archive, university ofessex, ux background and acknowledgements this paper is based on a session and a round table lunch discussion on the same theme which took place at the lassist conference held in edinburgh in may 1993. the session was convened by sue dodd and bridget winstanley. papers by laura guy, joanne lamb, marcia taylor, paul child and kevin schurer, as well as the numerous participants at the round table discussion have all contributed to the ideas presented here as have the members of a european committee on documentation guidelines, set up by the esrc data archive earlier this year. this committee has representation from the esrc data archive, the office of population censuses and surveys, social and community planning research, the british household panel survey and other areas of the british academic research sectot and the steinmetz archive in the netherlands. the need for guidelines we start from the basic premise that the person or persons best placed to document their data are the primary producers of those data. it is axiomatic that their knowledge of the data must be more complete than anyone else's. yet many primary producers ofdata are reluctant to create documentation of a standard whichgoes beyond their immediate needs for their own analysis of thedata. the reasons for this reluctance, when it occurs, are obvious. the creation of documentation of a substantively and physically high standard is time-consuming and expensive. thereare apparently few incentives to producing such documentation. the culture of data sharing is still largely in a state of infancy even after at least a quarter century of data archiving. and additionally, it is not always apparent to the primary producer what the secondaiv analyst requires in the way ofdocumentation. the primary producers who fail to document their data to an acceptable standard must be balanced by some shining examples ofgood practice in this field. some of the most recent of these include the british household panel survey's two volume user manual(l) and the u.k. employment department's user guide to the quarterly labour force survey (2), both of which are available as machine-readable text files at a much lower cost to the user, as well as appearing in printed paper form. north america can show many examples of data which are well documented by the producer for public use, including the general social surveys produced at the national opinion research center (3). a recent pubucation by the steinmetz archive in the netherlands documentsa dataset put together from a time series of nipo polls (4) with thoroughness and consideration for secondary users of thedata. despite these fine achievements, and many others, by individual research projects, there is much more that data archivists and librarians can do to promote good documentation by f)rimary producers of data. the arguments for doing so encompass both the promotion of good practice and necessity arising from financial and economic constraints facing disseminators. we have already stated what we take to be self-evident, that primary producers are capable of producing the best documentation because of the familiarity with the data. the further imperatives for persuading primary producers that they have a role and a respxmisibility towards the documentation of their own data lies in the decreasing resources and increasing material coming into data libraries and archives. many can no longer afford to create documentation for all (indeed any) of the datasets which they distribute and in any case the upgrading of pxmr documentation after the original project is over is frequently painful and unsuccessful: memories have dimmed and in many cases the wiginal investigators have disp)ersed. yet datasets which are inadequately documented are of no use at all to the secondary users to whom the data are being disuibuted. a further important incentive to the pjroduction of good documentation was described by wj. bradley at the lassist/ifdo 93 conference (5). the ^wnsors of major data collection exercises, typically government depiartments and other pralicy-maidng organisations, expect more for their money than data. they expject information. according to bradley, policy advisors are often quite desp)erate for timely, relevant information. given their wide-ranging and often unpredictable requirements, advisors and decision makers are a prime target audience fox easy, responsive secondary data analysis services that integrate and draw upon the broadest possible base resources. bradley and his colleagues have created software which demonstrates how good documentation, when standardised and 10 lassist quarterly structured, can integrate and front-end rapid and easy access to the data resources that have been documented in this way (6). they also describe how such documentation can actually serve to facilitate the creation of information and knowledge products which in turn can be integrated fw re-use in information retrieval the development of documentation guidelines, togethct with associated methods of standardisation, are keys to the knowledge delivery process. strategies for improved user documentation there are several lanes in the highway which leads towards the ultimate goal of improved documentation by producers of data. we need to convince data funders of the economic arguments in favour of improvements in the standard of documentation. we need to convince data producers of the value of good documentation to the organisation of their own research, as well as of the recognition of their work which will come from their wotk being re-used and acknowledged. we need to convince secondary users to afford this recognition to imimary producers. finally, we need to provide support to primary producers by developing and distributing guidelines on the production of documentation. the case to be made to the funders of data is, as indicated in the previous paragraph, primarily an economic one. many funding bodies are indeed aware of the wastefulness of funding projects with major data-collection components without ensuring that the data are made available for further research beyond its primary research aims. in many cases they are aware, too, that a majw constraint on the re-use of data is the lack of adequate documentation. there is sometimes a perception, however, that the disseminating agency, usually a data archive or data library, will document the data, so the jhxxlucer does not need to move beyond minimal standards. we must make the case that producers are better placed than archivists to create documentation of a high standard for their own data and that it is more cost effective for them to do so. a certain amount of data processing and standardisation will always be necessary in the archive or data library, but the better the incoming documentation, the better the outgoing data and documentation. funders are in a powerful position to provide incentives in the form of additional funds for documentation procedures within the original project funding as well as penalties in the form of blacklisting for those who do not document their data adequately. the judgement as to whether the data are adequately documented for secondary research will probably be the archive's and for this reason we need minimum standards in the form of guidelines. data archives and librarians will rely largely on funding bodies to provide the penalties for inadequate documentation. but they have a major role to play in persuading their depositors or donors of the incentives for providing high quality documentation. above all, the case has to be made for making their data widely usable. why should they care? because usage can be reported back to funding bodies as an argument for more funding; because when data are well-documented there is no need for the constant answering of queries from secondary users; and because usage will bring citation and recognition. here we, the data librarians and archivists, have a task ahead to ensure that use of data which leads to publication also leads to the citation of the dataset. the rules of citation for datasets are well established (see dodd (7)) but we can do more to ensure that they are observed. a scan of examples reveals also that there needs to be clarification on whether the documentation or the data, or both, are being cited. of the examples given above, only the general social survey's documentation (3) gives guidance on both the citation of data with documentation and the documentation alone, although the esrc data archive's citation guidance does make it clear that the citation shown is for data with documentation. the other two cases assume citation for documentation only. guidance on citation should be included in all documentation, editors of journals should be approached to try to ensure their co-operation, and a constant stream of reminders published in newsletters and bulletins. citation has its own rewards in the form of easier identification of data sources for those reading the citation, but also, of course, it ensures the recognition of the achievement of the producer of that dataset in making it publicly available. but citation can only take place when the dataset has a bibliographic identity conferred upon it by its documentation. guidelines are required to show producers how to document their data in a way which will ensure this. existing and future guidelines guidelines already exist for creating the necessary elements for documentation. two us examples are carolyn geda's data preparation manual (8) and richard roistacher et al a style manual for machine-readable data files and their documentation (9). other examples are the u.s. bureau of justice statistics' technical standards fw machine-readable data (10) and patrick collins and jane l. powers the preparation of data sets for analysis and dissemination : technical standards fw machinereadable data (11). excellent as they are, the earlier of these manuals are out of date and need revision while the latest (collins and powers) although providing a attractive introduction to the subject, focuses on the practices required by a particular archive (the national data archive on child abuse and neglect at cornell university) and is consequently short on general detail. a new comprehensive set of guidelines, covering both winter 1992 11 optima] and minimal standards, taking into account new media, new formats, new data collection techniques and a new archival environment, is urgently required. these should include a recognition of the fact that many social scientists are using and creating textual data, or mixed numeric and textual data in their research. it is important that new guidelines should recognise too, the considerable work already undertaken in the humanities and not to duplicate that work. the work of the text encoding initiative should be brought to the attention of social scientists in a way which will be aj^ropriale to their needs. although the guidelines should deal with substance and content, format should not be forgotten. for many primary producers and the archives or data libraries which will be disseminating their data and documentation, the most convenient fchmat in which to produce documentation will be machine-readable. in addition to providing a cheap and convenient means of disseminating documentation on the same medium as the data, machinereadable documentation opens the way to better information systems, allowing the prospective user to examine and compare documentation online before deciding on the ajpropriateness of a particular dataset for his or her particular research. once we have agreed on both optimal and minimal standards fch* documentation we need to think about how to get them accepted. if they have been developed in consultation with data producers and if they are attractive and easy to use, this will be easier. a printed paper version is indispensable but we must also develop software applications of the guidelines. work in this area has already begun, notably by w.j. bradley and his colleagues in the social environment group of health and welfare canada. their work on ddms (6), a pc-based package for managing social science dictionaries and documentation takes into account the data elements recommended by roistacher and provides an easy way to manage data as well as ensuring that these data will be welldocumented. such easy-to-use software in the hands of data producers will be an incentive to the production of complete documentation. the further work by bradley, hum and khosla on dais (data and information sharing) (12) shows how easy, end-user access to data can be provided by documentation that has been structured and standardised via ddms. this system provides a vital incentive to the funders of data who are themselves able, via this system, quickly to locate relevant data items from a broad array of datasets and generate their own analyses using software of their own choice. other work on codebook software has been carried out by the swedish social science data service and further work on codebook production is under way as an lassist action group led by karslen boye rasmussen of the danish data archives. while recognising the contribution this will make to the sharing of data through data archives, this paper, because it is concerned only with the primary jmxxlucer's role and responsibility, does not aspire to enter into the current debate, conducted largely through the lassist listserver, on the desirability of replacing osiris as a codebook tool. it is vital, however, that before we undertake the publicity and training required for the acceptance of software [hoducts, we are agreed on the substance of the guidelines for the documentation of data. conclusion penalties, incentives and support all depend upon the existence of guidelines for documentation. funding bodies have to be persuaded (as many already are) that the provision of funds for research projects to collect data at great expense without making provision for the widct use of these data is intolerably wasteful. for some, such as large governmental organisations, good documentation is essential for sharing within their own organisations, and all that is required is some guidance on how to do it in a way which has a broader application outside their own spheres. other types of funding organisations, who have traditionally seen a single report as the end product of their sponswship, need to be made aware of how much further their money will go if many reports and analyses for different purposes and by different researchers can result from their investment their role with regard to the documentation of datasets which they have funded should be to withhold further funding if the data are not sufficiently documented for further research (stick) and to provide an element of funding sufficient to ensure that the data are documented (carrot). primary researchers have to be persuaded (as many already are) that the creation of a dataset which can be used by others is worthy of recognition, acknowledgement and citation in the course of scientific research and public policy planning. secondary researchers, those making public policy, and the editors of journals should be persuaded to provide the recognition, acknowledgement and citation. the wider use of data and the recognition of the primary producers is dependent on the quality of the documentation which accompanies the data. the quality of the documentation will depend on the guidelines which we, the data librarians and archivists whose task it is to facilitate the flow between primary and secondary researchers, can provide to primary producers. 1 paper presented at iassist/ifdo'93 conference, edinburgh, scotland. (1) taylor, marcia freed (ed.) (1992) the british household panel survey user manual. 2v. colchester university of essex. 12 lassist quarterly (2) great britain. office of population censuses and surveys. social survey division (1992) quarterly labour fwce survey user guide. colchester. esrc data archive [distributor]. (3) davis, james allan and smith, tom w. (1991) general social surveys, 1972-1991 : cumulative codebook. chicago: national opinion research center. (4) eisinga, rob and albert felling (1992) confessional and electoral alignments in the netherlands, 1%2-1992 : documentation of social background variables of 1,067 national surveys conducted by nipo from 1%2 to 1992. amsterdam: steinmetz archive. (5) bradley, wj., diguer, j. and euis, r.k£., methods for producing interchangeable data dictionaries and documentation. paper presented at lasslst '90, poughkeepsie, nj. social environment information health and welfare canada, 1990. (6) bradley, wj., ruus, l., ellis, r.k.e. and diguer, j., ddms : a pc-based package fw managing social science data dictionaries and documentation : reference manual. 9th draft ed. social environment information health and welfare canada, 1991. — ddms [computer files]. social environment information health and welfare canada, june 1991. (7) dodd, sue a. bibliographic references for computer files in the social sciences : a discussion paper. lassist quarterly, v.l4, no. 2, summer 1990. (8) geda, carolyn data preparation manual. icpsr, 1980. (9) roistacher, richard a style manual for machinereadable data files and their documentation. urbana; university of illinois, 1978. (10) u.s. bureau of justice statistics technical standards for machine-readable data (11) collins, patrick and jane l. powers the preparation of data sets for analysis and dissemination : technical standards for machine-readable data. ithaca: national data archive on child abuse and neglect, 1991. (12) bradley, wj., hum, j. and khosla, p., metadata matters : standardising metadata for improved management and deuvery in national information systems. paper presented at lassist/ifdo '93, edinburgh. social environment information health and welfare canada, 1993. winter 1992 13 editors’ notes: much new research, and advances for the iq 1/2 schwartz, ofira & hayslett, michele (2023) editor’s notes: much new research, and advances for the iq. iassist quarterly 47(3-4), pp. 1-2. doi https://doi.org/10.29173/iq1100. the creative commons-attribution-noncommercial license 4.0 international applies to all works published by iassist quarterly. authors will retain copyright of the work and full publishing rights. editors’ notes: much new research, and advances for the iq welcome to this special double issue of iassist quarterly for the year 2023, iq vol. 47(3-4). we are delighted to close out the year by offering the second special issue of the iassist quarterly to showcase articles from the africa workshop. articles in this issue were presented in the second africa workshop which was held in ibadan, nigeria, in october 2022. guest editors winny nekesa and robert stalone buwule have again expertly steered the editorial process to bring us this research. while their guest editors’ notes describe the included articles, we would like to use this space to share a number of announcements about administrative work on the journal. first, please join us in welcoming four new editorial board members for a four-year term: robert stalone buwule, mbarara university of science and technology, uganda winny nekesa, national social security fund, uganda deborah wiltshire, gesis, germany, and ryan womack, rutgers university, united states with these appointments, we achieve two goals. first, we stagger the terms of service of board members so that only half will roll off the board at any one time, ensuring continuity of knowledge moving forward. second, we better diversify geographic representation on the board to reflect iassist’s international membership. deborah and ryan bring perspective from the iassist administrative committee to the board. robert brings experience as an iq guest editor. winny brings experience both from the ac and as a guest editor. over the coming year, iq editorial staff and board members will be exploring a variety of changes to the journal, many of which were proposed by you, the membership. we’ll keep you informed as we make decisions on various of those suggestions. several advances that we have already accomplished are to behind-the-scenes processes but may directly benefit authors who publish with us as well you, our readers. working retrospectively to the last issue, 47(2), as well as for all issues going forward, the editorial staff have published the reference lists of all articles as metadata. this complies with i4oc, a standard that asks for citations to be structured, separable, shareable, and freely accessible (to both human and automated harvesting), resulting in citations that are index-able and searchable. citation-tracking services like crossref also require this. the end result is that people searching for any of the sources listed in our articles will find our articles, which over time may result in greater research impact for our authors. reference linking will also expose articles to new tools, such as openalex, an open citation https://doi.org/10.29173/iq1100 https://iassistquarterly.com/index.php/iassist/about/iqeditorialboard https://i4oc.org/ https://openalex.org/ https://creativecommons.org/licenses/by-nc/4.0/ 2/2 schwartz, ofira & hayslett, michele (2023) editor’s notes: much new research, and advances for the iq. iassist quarterly 47(3-4), pp. 1-2. doi https://doi.org/10.29173/iq1100. metrics tool that can help measure impact. we thank our managing editor, phillip ndhlovu, for his effort in effecting this change. the second change was made by the library open publishing and open education staff at the university of alberta, whose work supports the open journal system platform on which the journal is hosted. their efforts not only keep journal production flowing smoothly, they work continually to improve the technical systems to uphold and improve open access to our content. in this case, they have implemented a research organization registry (ror) feature to allow authors to select their organizational affiliations from the list of organizations in ror. this will not only speed the information input authors must complete during submission, but also standardize it to be represented consistently within the journal, and make it clear and accurate for sharing in external systems such as doaj and crossref. finally, we want to mention a new content feature premiering in this issue. following the receipt of a letter to the editor (to our knowledge the first ever), we’ve added a new section to the journal’s infrastructure to accommodate such conversations. we hope you will enjoy reading this commentary which extends the implications of the hertzog, et al. article in 47(2) to a different type of personally identifiable data, dna. we invite you to take advantage of the option to use this feature in future to correspond with our articles. we wish you all the best for whichever end-of-year holidays you celebrate! we look forward to showcasing your work through the iq in the coming new year, both in the research you submit for publication and in implementing your ideas for evolving the journal’s content and production. ofira schwartz and michele hayslett, december 2023 https://doi.org/10.29173/iq1100 https://ror.org/ 1/15 greer, rebecca. & curty, renata g. (2022), investigating teaching practices in quantitative and computational social sciences: a case study. iassist quarterly 46(3), pp. 1-15. doi: https://doi.org/10.29173/iq1039 investigating teaching practices in quantitative and computational social sciences: a case study rebecca greer1 and renata g. curty2 abstract data education is gaining traction across disciplines and degree levels in higher education. teaching data skills in the social sciences in today's data-driven world is vital for preparing the next generation of dataliterate and critical social scientists. the ability to identify, assess, analyze, and communicate well and responsibly with data is key for scholars and professionals to navigate dynamic and expansive information ecosystems. this paradigm shift demands instructors to adapt their curricula and pedagogy to advance students’ computational and statistical knowledge. this paper presents some of the findings from a local report of a larger national project which explored pedagogical techniques and instructional support needs for teaching undergraduates with quantitative data in the social sciences. results revealed that the core learning goal of instructors is to develop students' critical thinking skills with data, including the conceptual understanding of the research methods employed in the field; the ability to critically evaluate research methodologies, findings, and data sets; and prowess using quantitative and computational tools and technologies. a recurring theme across interviews was students’ fear of math and technology and the challenges these fears pose to data-related instruction. instructors value participation in a community of practice and are eager for more institutional support to advance their computational skills. based on these findings, we suggest avenues for academic libraries to further develop services, activities, and partnerships to aid data instruction efforts in the social sciences. keywords data literacy; statistical literacy; computational literacy; social sciences; data pedagogy background and motivation a thriving number of initiatives have emerged in the last decades as a response to the growing need to equip students with foundational data skills to succeed in our data-driven world. carmi et al. (2020) articulate that data literacy (dl) is critical to achieving data citizenship, combining data thinking, doing, and participation. data thinking entails the critical understanding and evaluation of the data. data doing refers to how individuals engage with data and manipulate it in a more practical way, whereas data participation refers to proactively implementing transformative approaches with data to anticipate problems, and form solutions. libraries play an essential role in helping to prepare the next generation of data literate and critical citizens capable of navigating the data landscape and lifecycle, including the means to identify and assess data sources, extract insights and patterns from data, and communicate them effectively and responsibly. libraries are at the intersection of information exchange and commonly aid in the efforts of researchers, whether they be emerging or established scholars, to engage data to construct new knowledge. henderson and corry (2021) describe dl as a broader term that encapsulates multiple ‘aspects of statistical and assessment literacy, pedagogical knowledge and data-driven decision-making under one umbrella’. for this paper, we concur with their views and define dl as one’s ability to confidently and effectively engage with data. meaning that one should be able to collect data, identify and select existing data sources, read and process data in various formats, derive meaningful insights and accurate conclusions from data, as well as report and share data deliverables in a trustworthy and ethical way. we https://doi.org/10.29173/iq1039 2/15 greer, rebecca. & curty, renata g. (2022), investigating teaching practices in quantitative and computational social sciences: a case study. iassist quarterly 46(3), pp. 1-15. doi: https://doi.org/10.29173/iq1039 also integrated oecd’s (2019) understanding that data literacy encompasses both the technical and social aspects of data, and may include more overarching activities related to quality data management such as data documentation, curation, and citation. risdale et al. (2015) discuss data literacy skills and competencies in higher education and academic research. the authors propose a set of core skills and competencies for dl characterized by five dimensions: 1. conceptual framework: general knowledge and understanding of data, including use and application. 2. data collection: skills and knowledge related to quality data discovery and collection from multiple educational sources. 3. data management: skills related to data organization, preservation, manipulation, curation, and security. 4. data evaluation: skills related to data analysis, presentation, interpretation, and decisionmaking. 5. data application: knowledge and skills needed to share and cite data, evaluate decisions using data, and work with data ethically. through an environmental scan and analysis of current educational data literacy competence frameworks (edl-cfs) and courses, papamitsiou et al. (2021) extended these five dimensions into seven data-related core competence pillars for dl: 1) data location, access, and collection; 2) data comprehension; 3) data interpretation and transformation; 4) data use, application, and act on; 5) data analysis; 6) data evaluation and, 7) data management. these pillars generally map with the more holistic view of the data lifecycle phases used more extensively by research and academic libraries to design services and instruction to support data-related work in academic settings. data instruction opens up a wide array of opportunities to innovate teaching and learning while preparing globally competitive graduates to deal with real-world-oriented tasks. nonetheless, this more authentic and experiential educational environment demands strong partnerships in curriculum planning, continuous professional development efforts, and computational support and lab infrastructure. while the science, technology, engineering, and mathematics (stem) fields have been notoriously at the forefront of data education, more recently, the social sciences have begun to entertain more quantitative and computational approaches to address pressing contemporary social issues. along these lines, conte et al. (2012) emphasize the role of social scientists in exploring complex social systems through quantitative and computational lenses, given the growing integration of technology in our everyday lives and the massive volume of data generated from social interactions. this exploration involves multidisciplinary approaches to better understand the emergence of behavioral patterns in societies, their relationships with one another, and how individuals, groups, and communities interact with their environments both online and offline. in addition to valuing the contextual and ethical aspects ingrained in research processes methodologies, there has been a growing trend in the social sciences to incorporate data literacy and more statistical knowledge and approaches to automate work with data, including data cleaning, processing, interpretation, and visualization to course curricula (stephenson & caravello, 2007). however, such pedagogical transformation does not come without some challenges. these transitions are heavily https://doi.org/10.29173/iq1039 3/15 greer, rebecca. & curty, renata g. (2022), investigating teaching practices in quantitative and computational social sciences: a case study. iassist quarterly 46(3), pp. 1-15. doi: https://doi.org/10.29173/iq1039 dependent on disciplinary and departmental traditions, as well as existing institutional support to facilitate conditions instructors can rely on. thus, we followed an exploratory approach to investigate teaching practices and the factors that may hinder or promote quantitative and computational social sciences instruction at the university of california, santa barbara (ucsb). through our research, we seek to understand not only how the university has been responding to this trend in education and adapting its pedagogy, but, more importantly, to identify what kind of support students and faculty need from the library and other campus partners to advance their data skills. this paper presents the findings of our local study as a part of a larger national project with ithaka s+r involving 19 other academic institutions. the study's goals were: 1) explore pedagogical techniques and support needs in teaching undergraduates with data and 2) provide actionable recommendations for stakeholders within and outside the library to inform new services, and practices to advance data instruction in the field. based on our case study, we hope to showcase opportunities to refine and enhance library services, activities, and partnerships to support computational and quantitative data literacy skills in the social sciences that can be relevant to other academic institutions. research methods our local study3 was carried out in coordination with ithaka s+r and followed the guidelines and research design defined for the national project 'teaching undergraduates with quantitative data in the social sciences’. we followed a qualitative and exploratory approach to understand the current practices of faculty teaching with data. the identification and recruitment of potential participants took into account the selection criteria pre-established by ithaka s+r: a) instructors of courses within the social sciences, considering the field as broadly defined, and making the best judgment in cases the discipline intersects with other fields; b) instructors who teach undergraduate courses or courses where most of the students are at the undergraduate level; c) instructors of any rank, including adjuncts and graduate students; as long as they were listed as instructors of record of the selected courses; d) instructors who teach courses where students engage with quantitative/computational data. a total of 22 instructors were invited to the study, and 10 consented to participation. interviews were conducted between september 2020 and january 2021 and followed a semi-structured interview guide with questions on how students are directed to obtain and engage with data in the course curricula in tandem with instructors’ professional development, and training needs to teach with data. due to covid and the campus shutdown, all interviews were conducted remotely over zoom and were audio-recorded for transcription purposes. interviews produced approximately 12 hours of audio recording. de-identified transcripts and metadata are available through dryad4. we performed coding on maxqda 2020 using a mixed-method approach through the combination of both deductive (top-down) and inductive (bottom-up) strategies. we started with an initial code tree that echoed the main topics present in the interview prompts. this initial coding scheme evolved and was refined as we engaged more closely with the data through iterative rounds of readings and review, which helped us to identify, tag, and rearrange themes that emerged from our conversations with faculty. findings and discussion through our recruitment efforts, we were able to interview 10 individuals from six departments within the social sciences. three of our interviewees were from anthropology, two were from sociology, two from communication, and one interviewee each from economics, global studies, and psychology & brain sciences. predominantly, these individuals held a professorship with three individuals identifying their https://doi.org/10.29173/iq1039 4/15 greer, rebecca. & curty, renata g. (2022), investigating teaching practices in quantitative and computational social sciences: a case study. iassist quarterly 46(3), pp. 1-15. doi: https://doi.org/10.29173/iq1039 rank as assistant or associate. the remaining two interviewees self-identified as lecturers. of the courses taught by these individuals, there was an even split between upper and lower-division courses, with the majority of them being large-enrollment classes that are accompanied by a lab often facilitated with the help of a teaching assistant. the analysis of instructors’ narratives allowed us to organize findings into four main categories, as described in the following sections. expected student learning outcomes and ways students engage with data the desire to develop critical thinking skills and advance students’ data literacy was consistently expressed across interviews. faculty views of critical thinking can be widely defined as the ability for students to actively, and whenever possible, autonomously respond to problem-solving situations. more importantly, instructors desire students to understand the necessary procedural steps of working with data while constructing meaning from the data thoroughly and accurately based on a specific question or problem. most faculty described that their classes are designed around statistical tests students perform when they actively apply learned concepts to assess statistical reports. the instructor's objective is for students to identify flaws in analyses and misleading findings. the following excerpt reflects the need for these skills and how that might manifest for students: ‘what is that? what does that even mean?’ [...] so, those are the critical thinking skills i want them to have to be able to assess right away, if you know, some representation of data that someone is putting out there is problematic. and usually, you can tell it's problematic, just, by the way, they have graphed it, there are ways to graph things to make the pattern look less clear and to make a pattern that's not there look like an actual pattern. (ucsb 1). for most interviewees, teaching students how to correctly interpret data and identify inaccurate statistical findings is key to preparing students to be critical consumers of data. ‘you don't want to just hear something and then take it in without being critical, you need to be a critical consumer to understand if you should believe what you’re being told’ (ucsb 6). similarly, ucsb 9 highlights that students are required to analyze published findings, which leads ‘into a critical discussion of what are the accurate statements you can make based on these data, then which statements misinterpreting correlation as causation’, given that this is a common misconception in statistics. some instructors also described the importance of their courses as means to increase students’ professional skills that align with the job market: ‘i kind of go through some general examples like that, kind of hitting some of the careers that i know that our majors tend to gravitate towards’ (ucsb 6). learning goals interviews surfaced three main learning goals that reflect instructors' experiences teaching with data in the social sciences. conceptual understanding refers to an integrated view of theories, methods, and concepts, their possible applications, interconnectedness, and scenarios where they can be applied. it also reflects students’ abilities to articulate reasonable questions which can be answered statistically. for example, a student learning about group comparisons should be capable of understanding the underlying theory and basic principles behind the most widely used statistical techniques for that purpose (e.g., t-test, f-test, anova, manova), their main differences and relationships, as well as the specific assumptions (e.g., sample size, distribution) they must consider. https://doi.org/10.29173/iq1039 5/15 greer, rebecca. & curty, renata g. (2022), investigating teaching practices in quantitative and computational social sciences: a case study. iassist quarterly 46(3), pp. 1-15. doi: https://doi.org/10.29173/iq1039 critical evaluation represents one’s ability to holistically understand and make an informed assessment of the methods and approaches followed by others, understand the meaning behind the outputs, and evaluate the validity and reliability of the assumptions or conclusions based on data. this learning goal also includes one's ability to identify limitations of collected data and identify potential ethical concerns. critical evaluators can form a plan of action based on their conceptual understanding of disciplinary knowledge in tandem with their ability to identify issues or gaps in the data to synthesize meaning. following the same example above, a student with such skills should be able to evaluate if a given test meets the required assumptions and is correctly employed to analyze the data to answer a specific research question in a particular scenario while being capable of understanding the analytical outputs presented to them. relatedly, critical evaluators would be able to target deviations from an original research question and decide which would be the most appropriate test to answer that question and produce meaningful and effective reports. it thus translates from a general hypothetical scenario to a realworld solution. working with data and/or tools comprises students’ prowess to engage directly with data sets, identify and select existing data sources, gather, manage, and manipulate data, as well as operate (at least at a basic level) tools that can help them to automate analyses and create visualizations to convey meaning. this entails the application of concepts and evaluation to perform hands-on problem-solving beyond analytical reasoning. to satisfy this learning goal, students should confidently work with the dataset and perform their chosen statistical test using a tool such as r, excel, google sheets, stata, or spss to produce meaningful outputs. as illustrated in figure 1, these learning goals are complementary to each other and might play a more or less important role depending on the specificity of the course. some faculty expressed that they dedicate most of their courses to explain concepts and basic statistical principles. others focus more on the evaluation of published studies and statistical reports. in contrast, others emphasize more hands-on practice with tools, and some try to balance all of these learning goals simultaneously. these goals complement one another, the visual depicted below shows the amalgamation of instructor responses. figure 1 learning goals https://doi.org/10.29173/iq1039 6/15 greer, rebecca. & curty, renata g. (2022), investigating teaching practices in quantitative and computational social sciences: a case study. iassist quarterly 46(3), pp. 1-15. doi: https://doi.org/10.29173/iq1039 expected learning goals were developed based on the clustering of skills interviewees desire to equip students with within their courses, as detailed in table 1: table 1 expected learning goals and skills observed learning goals skills definition 'students should be able to…' conceptual understanding develop hypotheses identify relevant questions that could be asked and answered with statistical data. ground stats into the discipline articulate potential applications as well as the advantages of statistics in the context of their field. master key statistical concepts identify variable types, units of analyses, and measurements critical evaluation identify patterns spot and observe trends and correlations in data sets. extract meaning from data read ‘beyond the numbers’ and extract relevant associations from the data. marry concepts and procedures make informed decisions about the best approaches to explore and analyze the data. correctly interpret outputs successfully evaluate and explain statistical analyses and their results. write reports produce statistical reports and effectively communicate findings while following recommended styles and conventions. working with data/tools locate and access data search, identify and access available data sources. basic coding feel more comfortable with tools that require some coding and writing basic scripts to automate statistical analyses. perform basic analysis run statistical tests which are more common to their field. create visualizations create meaningful graphical representations to represent findings. use new tools know how to use different software and statistical packages that could help them to more easily and efficiently work with data. test hypotheses perform hypothesis testing and verify possible correlations and relationships. evidence of learning goals in instructional praxis the goal to establish conceptual understanding is described by ucsb 4 as they relay the importance of introducing students to the hypothesis development: i try to show them the older hypotheses, the gaps in the hypotheses, and you know how those gaps are illustrated by certain examples from the collated data [...] that is to get them to think about, this is still a growing and developing field, perhaps they have some contributions to make and kind of trying to get them excited about it. https://doi.org/10.29173/iq1039 7/15 greer, rebecca. & curty, renata g. (2022), investigating teaching practices in quantitative and computational social sciences: a case study. iassist quarterly 46(3), pp. 1-15. doi: https://doi.org/10.29173/iq1039 instructors strive for students to gain familiarity with potential applications of stats to answer research questions in their field while challenging them to exercise their logical reasoning. the following excerpt exemplifies attempts by an instructor to foster a welcoming environment for students to apprehend key statistical concepts while navigating the provenance and methods behind the data: whenever i introduce a variable, first of all, they have to understand, they’re obliged to understand every aspect of the definition of the units in which it's measured and the real-world process by which somebody arrived at that number. [...] the question is, what units is it measured in, what does it capture, who made the number, who invented the number, what are the components of the calculation that were imputed, who imputed them (ucsb 8). relatedly, ucsb 2 emphasized that their classes allow students to exercise their ability to ‘think synthetically with some of the data and take in data sources from a bunch of different places that may not necessarily have obvious connections or sometimes have very obvious connections’. the importance of students being capable of critically evaluating the data and effectively communicating their insights and inferences was expressed by some instructors who assign statistical reports as deliverables for their classes: the other thing that i think is important is the writing process. […] they write a research paper, they introduce a problem, they review the literature on that problem [...] to address a gap, and then present their data and methods. and then, present their results, and then come back to the original question and say ‘what does this tell us about this question?’, or ‘how does this help us move the literature forward?’. so i think that's a very useful skill to have and to be able to sort of communicating […] others what you've discovered through your analysis. (ucsb 4). ucsb 8 also sees the process of producing their own statistical reports as an opportunity for students to become more data literate and critical of other people’s work: i really want students to be able to not just calculate statistics and values but to understand what those values mean on the other side. [...] we give them a data set and they work their way through statistical analyses and then they report those analyses in apa style. [...] how to report those values to someone else because i think that helps them with actually reading empirical papers. so, when they practice writing a results section, i think they are better equipped to actually read the results section of the paper. the ability for students to search, find, and access relevant existing data sources was highlighted by a few faculty as a part of the skillset of their courses. ucsb 5, for example, described that they usually talk about ‘how one might obtain data sets as well [and] some of the technicalities around [this process].’ such as, ‘finding data online using web scraper technology or using api technology to download data sets’. we observed, however, that instructors in this study most often supplied the data sets for class assignments, and only a few had students generate the data themselves. instructors usually chose publicly available and previously de-identified data such as national statistics, datasets that do not involve human subjects, or even develop assignments around dummy data. to provide student access to the data sets, interviewees often relied on the institution’s learning management system, gauchospace (moodle). https://doi.org/10.29173/iq1039 8/15 greer, rebecca. & curty, renata g. (2022), investigating teaching practices in quantitative and computational social sciences: a case study. iassist quarterly 46(3), pp. 1-15. doi: https://doi.org/10.29173/iq1039 most faculty indicated that at least some part of their course workload covers some basic functionalities of the statistical packages and software, which can help students to compute statistics more easily. these demonstrations are usually provided during lab sessions, and in some cases, are complemented by stepby-step guidelines provided to students on how to use the tool. microsoft excel was the most commonly used tool by instructors, and in some cases, interviewees expressed the necessity to move from microsoft excel to google sheets to ease issues with access to paid software. other software or tools referenced in the order of prevalence by interviewees were: r, spss, stata, q-gis, and eviews. four of the 10 interviewees combined excel with one of the other aforementioned programs for students to analyze or interpret data. beyond time in the classroom or lab, some interviewees described strategies for promoting extracurricular learning as means to help students to advance their data-related skills. ucsb 7, for example, created a series of optional short videos and demos that students can watch at their own pace while working on class assignments. ucsb 5 often recommends youtube videos and khan academy courses, and ucsb 2 referred to their syllabus, which includes ‘all kinds of external resources that are both associated with the university and wiki pages on the internet that are good for learning gis and where people ask questions to learn how to do things.’ a few faculty noted that internships and research projects with undergraduate students helped them become more well-versed in quantitative approaches and computational tools. however, they recognize that only a tiny percentage of students have participated in such activities. main challenges of teaching with data instructors’ perceived challenges were mostly connected to learners’ math and tech anxiety and helping them overcome these limitations without producing cognitive overload. most interviewees acknowledged students’ self-professed fears and obstacles students have had to overcome concerning their readiness to engage with statistics and affiliated tools to perform computations and produce outputs. not only do instructors believe that students’ entry knowledge is often limited, but they also recognize that the field still offers students few opportunities to develop statistical skills to gain confidence in using tools to automate their work, and that their classes are not able to fulfill all existing deficits. in addition, the vast majority of the interviewees see social sciences at ucsb as less inclined to positivist traditions. some interviewees expressed that their classes are the only opportunity for undergraduate students to interact more closely with quantitative data. the lack of options added to the fact that most interviewees signaled that students choose their majors based on their predisposition to soft sciences, which poses some challenges to instructors responsible for introductory courses on quantitative data. i think the thing that’s most relevant to classes involved in data analysis is just students’ fear of math. and that’s going to vary across the divisions, right? i mean, i’m sure engineering students have a lot less fear of math than humanities and social sciences students do. (ucsb 6). ucsb 5 mentioned that students often make comments such as: ‘what, it involves math? logic? no, i don’t want to do that.’ similarly, ucsb 9 mentioned that students are often quite ‘upset about having to touch numbers and having to work with numbers in the first place’, and later expounded: [...] to actually teach them anything about how to do quantitative analysis in any serious way is very difficult, because they're not seeing any of it in any other classes, many of them are taking classes where they're actually actively discouraged from dealing with quantitative social science. they just don't have any training, many of them are math phobic. they have chosen [redacted, https://doi.org/10.29173/iq1039 9/15 greer, rebecca. & curty, renata g. (2022), investigating teaching practices in quantitative and computational social sciences: a case study. iassist quarterly 46(3), pp. 1-15. doi: https://doi.org/10.29173/iq1039 name of the course] because they are either afraid of mathematics or somehow, have some issue with it (ucsb 9). the fear of dealing with tech or the lack of digital dexterity, even among those courses which introduce basic software, such as excel and google sheets, with no coding to perform statistical analysis, was also observed by some of the instructors. for example, one interviewee expressed the expectation that students should be more comfortable handling computer systems, but, unfortunately, that is not always true: ‘and this is something which has sort of surprised me because i just assumed that over time, more and more people would be computer savvy, and it does not seem to be going that way.’ (ucsb3). on the same note, ucsb 8 stated: in fact, i’ve been quite surprised, for example even working with excel some students have a difficult time and don't seem to have much experience in working with excel or uploading something to r for example. i mean it requires a little bit of coding knowledge but the process is kind of similar to uploading a photo to an email. (ucsb 8). besides the general understanding that most students in the social sciences are not as keen on learning statistical and computational approaches to analyze quantitative data, some faculty stressed the challenge of balancing statistical with computational instruction in one course. ‘the whole point of the class is for them to learn how to use the software in addition to how to use the data that they use in the software’ (ucsb2). some faculty expressed concern about the amount of time spent teaching students how to navigate and operate statistical tools and how that can take away their ability to focus on the lessons’ content. this concern was presented as a justification for their choice to work with more basic computational tools. we believe that this may, at least partially, influence instructors to lower their expectations for students within their classes. […] we have them produce histograms, line graphs, you can do box plots, within spss, it has its own kind of data visualization. nothing fancy. and the class, like i said, we keep the class very, very simple and this is not an advanced class. this is a class for students who have an absolutely terrible fear of math and a fear of quantitative data and so we keep it very simple. [...] again, these are students who are not strong enough, resent having to take anything related to math or data analysis and i don't want to overwhelm them [...] (ucsb 6). findings also show that most faculty have been adopting measures to mitigate such challenges, by minimizing obstacles with tools and employing ones that are easier to operate while working with vetted data sets where they have better control, but the data still demonstrates some of the basic statistical concepts and applications. while we understand that this approach seeks to accommodate students’ fears, we believe it reduces opportunities for students to engage more critically with the data to apply learned concepts and tools to different contexts. on the one hand, having more control can help students better acclimate to statistics; on the other hand, it can narrow opportunities for students to practice their problem-solving and to translate general skills more autonomously to specific contexts. instructors’ training and resource sharing instructional training and resource sharing varied among interviewees. aside from their graduate education training, most faculty rely on professional development opportunities, such as academic https://doi.org/10.29173/iq1039 10/15 greer, rebecca. & curty, renata g. (2022), investigating teaching practices in quantitative and computational social sciences: a case study. iassist quarterly 46(3), pp. 1-15. doi: https://doi.org/10.29173/iq1039 conferences and workshops, including carpentry workshops offered at ucsb library to advance their teaching with data skills. the second most common method is through self-discovery, such as reading books and related literature on the topic or program, watching online video tutorials, and following trends with large technology companies such as google llc, meta inc. (facebook inc.), and microsoft corporations. another less referred method observed was through interor cross-departmental collaborations with other faculty. in these few instances, colleagues were often consulted as a resource for independent research to fine-tune their techniques. in other cases, they consulted with colleagues for their expertise as an affiliated resource either for instruction or their research. i’ve started to do some computational work or work that requires computational analyses with big data and i don’t have the skill set to do that. i partner with computer scientists. so, i collaborate. i find people who have the data analysis skills that i don't. (ucsb 6). the majority of interviewees said they were either willing to share or have shared their instructional resources with students and/or fellow instructors. however, only half of the interviewees have used shared resources to develop course materials. one respondent, in particular, expressed that they would not be able to use shared resources or conversely share their own instructional materials due to the nature of their class being the only one of its kind in the department. any resource sharing that could be done inter-departmentally would not be usable for their needs. [...] you know that that's something i've entirely had to figure out on my own. and, you know, i get a lot of practice with that, because nobody else in my department teaches quantitative anything. so i've had to do a lot of it, and i just sort of learned through trial and error […] (ucsb9). instructors described the benefits of engaging in a community of practice through which they can both reuse and share instructional materials. yet, respondents often cited difficulty engaging in professional growth and training in this area due to a lack of time. anecdotal events were referenced, such as a summer institute where instructors were prompted to reflect and retool pedagogical techniques as an opportunity to interrogate their teaching practice and reimagine their approach to measuring student learning. however, examples like these were limited as a robust, integrative opportunity to retool their instructional practices. types of support needed the interactions we had with instructors confirmed our underlying assumption that there is a growing movement or desire toward more computational and quantitative-oriented social sciences. some faculty acknowledged student attentiveness and interest when courses were oriented around digital topics. well, i think that there is a trend in the field that is sort of broadly one, where people are interested in anything digitally, i sort of attach the word digital to it, and people are all of a sudden, more attracted to it [...] because i think they're interested in developing translational skills. (ucsb 2). conversely, most of the interviewees recognized that their departments are not ready yet to fully accommodate this trend, being unable to fulfill this demand since they are more heavily focused on qualitative research presently. during the interviews, most participants identified themselves as one of a few in their departments who conduct statistical research and engage in quantitative data-related https://doi.org/10.29173/iq1039 11/15 greer, rebecca. & curty, renata g. (2022), investigating teaching practices in quantitative and computational social sciences: a case study. iassist quarterly 46(3), pp. 1-15. doi: https://doi.org/10.29173/iq1039 instruction. some interviewees also emphasized how this qualitative orientation affects social sciences degrees at ucsb to advance this direction: faculty members are very, very few quantitative and that has been also a problem because we need more. in order to be able to be stronger in quantitative methods in [redacted department name] here, and to be able to get more students that are quantitative, that want to do quantitative work, you need more faculty that does quantitative work. because if not, you are signaling that this is a qualitative department. (ucsb 7). we also asked interviewees to describe the types of instructional support they need and receive to teach with quantitative data. as mentioned previously, most of the classes discussed were large enrollment classes with associated labs facilitated by teaching assistants. these teaching assistants were often referred to as invaluable to the course as they are the ones who align the critical intersections of theory and methodology with the available technologies. in our courses they lead sections and so [...] in my classes that do data analysis they go over the homework problems with students and help them understand because students do data analysis by hand there and they lead the lab sessions. so, they are the ones teaching spss and showing them how to obtain the outputs to then interpret. (ucsb6). while these teaching assistants structure the practical use and application of the technologies for students, they also, at times, allow instructors of record to learn new data tools. for example, ucsb1 expressed that some students who have become teaching assistants were able to surpass the instructor's expertise with more advanced tools. [...] they've sent me all of the material [...] all the books that they think are the best ones, they've sent me the youtube videos that walk you through it. so i have gotten all these resources from two students who have previously taken my quantitative class [...]. so basically, i've taught them, they have gone beyond me, and now they're teaching me and that is the way learning should work […] it was not uncommon for instructors to express a desire for additional support with learning new programs and further support with technical aspects of the course, especially in a lab setting. for instance, interviewees expressed an interest in having a single, dedicated space on campus to host training on programs and technologies on teaching with data. this suggestion was often affiliated with comments where instructors expressed a lack of time or financial resources to pursue these interests independently. as such, they would value a structural intervention by their department, program, or the university as a whole to centralize professional development in this area. in summary, our conversations with faculty surfaced the need for additional services and resources to support teaching with data in the social sciences on campus more broadly. the small-scale nature of this project, and the fact that not all relevant majors at ucsb were represented in our sample prevent us from generalizing results. we also acknowledge that drawing comparisons between pedagogical face-to-face approaches and instructors’ strategies to accommodate classes to the virtual setting in the context of the pandemic was beyond the scope of this project, but could offer valuable insights for a future study. despite these limitations, we believe our exploratory investigation of quantitative and computational data https://doi.org/10.29173/iq1039 12/15 greer, rebecca. & curty, renata g. (2022), investigating teaching practices in quantitative and computational social sciences: a case study. iassist quarterly 46(3), pp. 1-15. doi: https://doi.org/10.29173/iq1039 teaching practices signals potential directions to address current challenges and limitations preventing data literacy from moving forward in the field of social sciences at ucsb. envisioning ways the library can leverage data instruction considering how academic libraries can actively intervene with the challenges experienced by instructors teaching with quantitative and computational data at ucsb, the co-authors surfaced an important distinction between those who teach credit-bearing courses and library programs and services. instructors-of-record are bound by a curriculum limited by a quarter system which is 10 weeks in length. library programs and services often follow the ebbs and flows of the quarter when responding to instructional and research needs. yet, libraries are not bound in the same way by these conditions. library personnel think expansively about how researchers at the institution engage with the continuum of the data lifecycle and necessary skills needed to work with data with these temporal conditions in mind. library personnel must also remain accountable to the broad spectrum of academic skills and life experiences held by the various users they serve. ucsb is a minority-serving institution and has an expansive portfolio of departments and programs that align with social science research practices. as expressed at the beginning of this paper, to engage in pedagogical transformation we must account for the disciplinary and departmental traditions while examining the institutional supports already in place, including our own. our recommendations reflect possible avenues to support a variety of stakeholders to advance data instruction in the social sciences at ucsb, which could be influential for other academic libraries. in these recommendations, we recognize that this work will likely involve multidisciplinary approaches that can be disruptive to existing patterns in institutional and community contexts and learning modalities. while library staff are contacted with practical research-based questions regularly, it is commonly noted that library staff need further support and opportunities to learn techniques when working with computational and quantitative data (usova & laws, 2021; zaidane & koizumi, 2019). our findings did not directly surface a need for librarians or library staff to be involved in course curriculum development. it was commonly referenced that instructors lack sufficient time to engage in professional development activities to advance these skills. hence, there is an opportunity for library staff to serve as consultants and deliver training to instructors and students at critical points of need. this requires staff to gain both practical experiences with using computational tools paired with a conceptual understanding of methodologies commonly applied in the social sciences if they are not familiar already. currently, ucsb library provides a variety of information sessions and training on tools that can support statistical analysis. however, many of the interviewees were not familiar with these services. library staff are encouraged to upskill and attend training in these areas, but there are no formal requirements nor structured communications to intentionally foster a broader community of practice in the library with statistical tools. efforts can be made to improve communication and outreach not only to instructors but to groups, such as subject librarian liaisons, to express the utility of upskilling in these areas to support our academic community further. active encouragement to engage in internal professional development can be led by distinctive groups, such as the research data services (rds) and the interdisciplinary research collaboratory (irc) at ucsb library. these groups are best positioned to work with subject liaisons to identify current needs and tailor communications to departments when training or workshops are made available. a partnership such as this would provide novel opportunities for library staff to engage their departments and better apprehend how their liaison role can further support data literacy initiatives within the social sciences. https://doi.org/10.29173/iq1039 13/15 greer, rebecca. & curty, renata g. (2022), investigating teaching practices in quantitative and computational social sciences: a case study. iassist quarterly 46(3), pp. 1-15. doi: https://doi.org/10.29173/iq1039 a common expression shared by many instructors was that they perceived students’ fears often interfered with their ability to readily engage with the course content and that the limited time they have with students is usually not enough to overcome such fears and fulfill pre-existing knowledge gaps. extracurricular activities could help mitigate these problems. there are many studies (e.g., carlson, 2015) on ways to support successful extracurricular activities following the ‘learning by doing’ approach through hackathons, bootcamps, research projects, internships, and alike. as acknowledged by interviewees, students are still offered very limited opportunities to engage in similar activities, and the length of their course prevents them from planning these. at ucsb library, some experiential and immersive learning activities around data are being considered for adoption through the newly launched rds workshop series, with planned sessions where attendees will be invited to bring their own data to learn new skills, and have an opportunity to showcase their project deliverables. in an effort to address potential issues of equity and access with immersive learning experiences with data literacy, the library is advancing projects to accommodate remote and asynchronous learning through modularized digital materials. currently, a partnership is being established between rds and teaching & learning (t&l) at ucsb library to integrate these materials into course curricula. rds holds expertise in the different stages of the research data lifecycle, including data gathering, cleaning/wrangling, analysis, documentation, archiving, and preservation. t&l holds expertise with foundational information literacy skills, pedagogy, instructional design, and educational technologies, and often provides instruction to lower-division undergraduate students. to help to allay student fears, rds and t&l propose to create a foundational data literacy instructional module that can be readily deployed within the campus learning management system (lms) and used asynchronously in credit-bearing courses. preand post-assessments can be integrated with the module to measure student gains for both the instructor and the liaison librarians who may work with the class. this instructional delivery method has been successful when working with other entry-level courses, such as in the writing program at ucsb. based on this prior experience, t&l and rds will engage relevant stakeholders in the design process of this module to gather their input on the curricula and design elements, including intended learning outcomes and required assessment criteria. once deployed, the asynchronous module can be paired with online discussion forums within the course lms and drop-in office hours to connect students with library personnel who can further facilitate their research using quantitative data. further support for instructors can also be realized by modifying existing training and services at ucsb library to better tailor to instructor needs. carpentry workshops are offered continuously within the library’s irc in partnership with rds and other campus volunteer instructors. currently, these workshops are made available to all campus affiliates to gain facility with foundational data and computational skills. however, instructors noted that lack of time is a key barrier to participating in similar professional development activities. while working with liaison librarians to craft communications for those who teach computational skills in the social sciences, these messages could also incorporate surveys to gauge interest and availability to attend existing programming. a possible outcome from these efforts may be faculty-only workshops to promote and support the instructor's existing community of practice in a modality that best aligns with their scheduling needs. additional support for instructors may come in the form of train-the-trainer programming for teaching assistants and associates (tas). many instructors rely on tas to lead course or lab sections for large enrollment classes. focusing on this audience to tailor services and programming, whether through hands-on synchronous workshops or the development of asynchronous course materials, may prove invaluable to a growing community of practice in the field and the next generation of instructors. in fact, engaging an audience of teachers who are students simultaneously can be key for the library to target https://doi.org/10.29173/iq1039 14/15 greer, rebecca. & curty, renata g. (2022), investigating teaching practices in quantitative and computational social sciences: a case study. iassist quarterly 46(3), pp. 1-15. doi: https://doi.org/10.29173/iq1039 emerging needs that are not fully represented in the curricula as of yet. this aligns well with the library’s interest in identifying and developing services and programming that support data literacy instruction generally, taking into account the whole data lifecycle. while it is imperative to pair our resources with curricular needs on campus, we are also positioned to advance the work of teachers in their dual roles as research practitioners who contribute to advancing disciplinary practices. data services and data education are logical and more recent outgrowths of the core role libraries have played for generations in educating the community they serve. while positioning the library as central to advancing data instruction in collaboration with departments, we also acknowledge the need to establish partnerships with other campus units and groups that could help us foster statistical literacy and data literacy more broadly at ucsb. examples of potential collaborators include the datalab coordinated by the department of statistics and applied probability, student organizations such as the data analysis and coding (danc) club, and interest groups such as the quantitative methods in the social sciences (qmss) and the ucsb research data community has members from various campus units to discuss data-related services, infrastructure, and instruction. because libraries are often seen as hubs for central services on campus, we believe they may also act as natural facilitators to connect related but still siloed initiatives, while nurturing community building towards shared goals and efforts concerning data instruction. mapping and engaging relevant campus units dedicated to advance data instruction to discuss and develop a more robust framework to support pedagogical approaches and learning assessments at the institutional level, in alignment with the whole data lifecycle continuum, is pivotal for preparing the next generation of professionals capable of exercising their data citizenship more confidently and profoundly. references carlson, j., nelson, m. s., johnston, l. r., & koshoffer, a. (2015). ‘developing data literacy programs: working with faculty, graduate students and undergraduates’, bulletin of the association for information science and technology, 41(6), p14-17. available at https://doi.org/10.1002/bult.2015.1720410608 carmi, e., yates, s. j., lockley, e., & pawluczuk, a. (2020). ‘data citizenship: rethinking data literacy in the age of disinformation, misinformation, and malinformation’, internet policy review, 9(2), p122. available at: http://hdl.handle.net/10419/218938 conte et al. (2012). ‘manifesto of computational social science’, the european physical journal special topics, 214(1), p325-346. available at: https://doi.org/10.1140/epjst/e2012-01697-8 henderson, j., & corry, m. (2020). ‘data literacy training and use for educational professionals’, journal of research in innovative teaching & learning, 14(2), p232-244. available at: https://doi.org/10.1108/jrit-11-2019-0074 oecd (2019). core foundations for 2030: conceptual learning frameworks. available at: https://www.oecd.org/education/2030-project/teaching-and-learning/learning/corefoundations/core_foundations_for_2030_concept_note.pdf (accessed: 12 mar 2022). papamitsiou, z., filippakis, m. e., poulou, m., sampson, d., ifenthaler, d., & giannakos, m. (2021). ‘towards an educational data literacy framework: enhancing the profiles of instructional https://doi.org/10.29173/iq1039 https://doi.org/10.1002/bult.2015.1720410608 http://hdl.handle.net/10419/218938 https://doi.org/10.1140/epjst/e2012-01697-8 https://doi.org/10.1108/jrit-11-2019-0074 https://www.oecd.org/education/2030-project/teaching-and-learning/learning/core-foundations/core_foundations_for_2030_concept_note.pdf https://www.oecd.org/education/2030-project/teaching-and-learning/learning/core-foundations/core_foundations_for_2030_concept_note.pdf 15/15 greer, rebecca. & curty, renata g. (2022), investigating teaching practices in quantitative and computational social sciences: a case study. iassist quarterly 46(3), pp. 1-15. doi: https://doi.org/10.29173/iq1039 designers and e-tutors of online and blended courses with new competences’, smart learning environments, 8(1), p1-26. available at: https://doi.org/10.1186/s40561-021-00163-w ridsdale, c., rothwell, j., smit, m., ali-hassan, h., bliemel, m., irvine, d., ... & wuetherick, b. (2015). strategies and best practices for data literacy education: knowledge synthesis report. halifax, ns: dalhousie university. available at: http://www.mikesmit.com/wp-content/papercitedata/pdf/data_literacy.pdf (accessed: 12 mar 2022). stephenson, e., & caravello, p. s. (2007). ‘incorporating data literacy into undergraduate information literacy programs in the social sciences: a pilot project’, reference services review, 35(4), p525540. available at: https://doi.org/10.1108/00907320710838354 usova, t., & laws, r. (2021).’ teaching a one-credit course on data literacy and data visualisation’, journal of information literacy, 15(1), p84-95. zaidane, a., & koizumi, m. (2019). ‘roles of data librarians and research data services in academic libraries’, information and technology transforming lives: connection, interaction, innovation. proceedings of the xxvii bobcatsss symposium, osijek, croatia, january 2019. available at: http://bobcatsss2019.ffos.hr/docs/bobcatsss_proceedings.pdf (accessed: 30 mar 2022). endnotes 1 rebecca greer is the director of teaching and learning at the university of california, santa barbara library. rrgreer@ucsb.edu. 2 renata curty is the social science research facilitator at the university of california, santa barbara library. rcurty@ucsb.edu. 3 the study was irb approved and was exempt by the ucsb’s office of research in july 2020 (protocol 1-20-0491). 4 curty, renata g.; greer, rebecca; white, torin (2021), teaching undergraduates with quantitative data in the social sciences at university of california santa barbara, dryad, dataset, https://doi.org/10.25349/d9402j https://doi.org/10.29173/iq1039 https://doi.org/10.1186/s40561-021-00163-w http://www.mikesmit.com/wp-content/papercite-data/pdf/data_literacy.pdf http://www.mikesmit.com/wp-content/papercite-data/pdf/data_literacy.pdf https://doi.org/10.1108/00907320710838354 http://bobcatsss2019.ffos.hr/docs/bobcatsss_proceedings.pdf mailto:rrgreer@ucsb.edu mailto:rcurty@ucsb.edu https://doi.org/10.25349/d9402j by 12 iassist quarterly summer 2007 andy boettcher*1 abstract the federal reserve board (board) purchases and creates numerous datasets to support its role in monetary policy, banking regulation, and consumer protection. to better manage these datasets, the board has built a metadata repository called the data and news catalogue (dance), which stores descriptive dataset characteristics. the growing number of datasets and their corresponding security and licensing intricacies motivated a data initiative in which the board’s research community identified enhancements to dance. planned improvements include: the addition of dublin-core standard metadata, the communication of changes in metadata, and the dissemination of metadata on new datasets. the improvements are expected to enhance collaboration between research units of the board, which will in turn enable better research. this paper will chronicle dance’s original role within the organization and its transformation into a knowledge management solution. background on data management the growing interaction between banking, the financial economy, and the real economy has necessitated crossdepartmental projects within the board and the federal reserve system. for example, three different research departments conduct market discipline2 research using similar data. as the number of multi-department research projects increase, data management practices are evolving to facilitate these changes. historically, data management practices were relatively fractured3 across the board as well as the federal reserve system4. research departments that purchased or created datasets tended to silo the data, documentation, expertise, or any combination of the three. if research projects were always intra-departmental, then this approach to data management would suffice. the acquisition of data also presented challenges. to purchase a dataset, a department must confirm that the board does not already have access to a similar dataset. since the board lacked a central metadata facility, attempting to find similarities across undocumented datasets proved particularly frustrating, time-consuming, and inadequate. arguably, more resources were spent determining if similar datasets existed than the cost of purchasing a duplicate dataset. dance history dance was first conceived in 2002 as a tool to encourage using newly purchased datasets to board researchers outside the purchasing group. increasingly, research staff’s forecasting and working papers required purchased datasets, yet knowledge of these datasets remained primarily with the purchasing group, not the larger user group. therefore, dance’s foremost purpose was to remove data silos by serving as a search hub linking locally documented and stored datasets with the larger board and reserve bank community. dance development staff identified several descriptive characteristics that they believed would reduce the silo effect. the metadata fields5 identified the purchased dataset and focused on the following four concepts: description, access rights, contact information, and data location. as altruistic as dance’s initial goals may have been, obtaining underlying metadata on each entry proved to be difficult. for example, some units were hesitant to share metadata due to staff workload concerns. dance staff populated metadata for as many datasets as possible. however, a thorough metadata entry needed input from users with expertise. many of the datasets were purchased by the board’s research and law libraries. thus, efficiencies were gained by having the research and law library staff update and maintain dance entries for library purchased resources. at this point, dance contained the majority of vendor purchased datasets, with the remaining spread amongst multiple departments at the board. since the datasets were not concentrated within one department, the library model could not be used. as a result, dance staff started working with administrative units within each of the research departments at the board. each administrative unit coordinates data purchases for the department, records which datasets are purchased and by whom. through this, metadata for the board’s remaining purchased datasets were populated in dance. each year following, dance staff reconciled entries with each division’s administrative staff. while the annual reconciliation was useful in data and knowledge management at the federal reserve board iassist quarterly summer 2007 13 verifying entries, it was not very timely. to provide more frequent updates, dance staff created a pilot project with one of the administrative units. this project required the data purchasing group to register the dataset with dance, which would then auto populate fields and provide information to the respective administrative department. with this enhancement, dance contained the most current metadata at all times. this process is planned to be incorporated into the new dance structure by implementing similar processes for each division. based on consistent usage statistics, dance appeared to fit its original purpose to broaden board use of vendorpurchased datasets. each user, regardless of department, could view the existence of board-purchased datasets with ease. thus, dance helped foster cross-departmental collaboration and reduced search costs for purchased data. dance revisions in 2006, the research divisions at the board undertook a new data initiative named the research divisions data initiative, or rddi for short. rddi’s purpose was three fold: data cataloging, data storage, and data presentation. before rddi could identify and develop better storage and presentation methods, all datasets needed to have standardized documentation created. dance had basic documentation for purchased datasets, but the research community wanted enhanced metadata for both purchased and board created datasets. as a result, a working group on data documentation was convened to identify new metadata fields for dance. the working group used dublin-core standards as their guide in identifying metadata6 items to better describe a dataset and thus better facilitate search capabilities. the working group started with over 100 items and whittled them down to 38. each metadata item was examined for its usefulness in describing a dataset, its potential as a search criterion, and its maintenance difficulty. for example, the dublin-core item ‘replaces7’ was removed since the working group thought the marginal cost of maintaining that item would be more than the marginal benefit. since dance’s primary purpose is to document purchased datasets, the administrative units at the board requested additional items. the administrative units are required to generate frequent reports and memos on dataset costs and corresponding budget justification. to assist with the generation of these internal publications, an additional ten items were added. once the new metadata items were established, the working group examined several open source software solutions to display them. four options were explored: e-prints, fedora, d-space, and a custom-built solution. each of the open source solutions presented their own set of installation challenges. for example, d-space runs off of java, however, the board provides limited java support. further, the ten additional administrative items were not supported by any of the open source solutions. as a result, the working group concluded that the new dance structure would require a custom, board-built solution. one of the many benefits of a custom-built solution is flexibility. dance was not limited by what metadata items to include or how they were communicated. dance further extends this flexibility with added communication methods, security permissions, and variable-level metadata. these enhancements immediately notify users of new datasets or changes to existing datasets. each dataset also has a unique set of security requirements. these requirements dictate usage rights, how access is granted, and publication rights. many times there are paper or electronic access request forms required for dataset use. a custom-built solution allows dance developers to display the correct request method for each dataset. finally, a custom-built solution enables dance to extend beyond dataset metadata to variable metadata. variablelevel metadata gives users a micro-level description of the dataset’s contents, thus enabling users to determine a dataset’s usefulness for their needs. new dance the new dance is designed not only to accurately describe a dataset, but also around how a user would search/receive metadata related communications, add board specific documentation, and update metadata items. the search screen will provide multiple ways to search and filter for a dataset. first, a persistent search box will be accessible on all pages, allowing users to search for a different dataset without having to return to the main dance page. all the fields will be indexed by the search engine, thus all searches will be free text, unlike library systems which allow users to specify author, title, etc. while the single search box provides a cleaner interface, it provides significant challenges in returning the correct results. dance will incorporate two methods that focus on interaction between the user and the page to meet this challenge. first, the search box will be ajax-enabled8 ajax queries dance for datasets that match the typed search without submitting the search. the user can then modify the search string as needed to pull up the desired results. please see figure 1 for an example of how each additional letter that is added to a search string reduces the returned results second, the user will be able to filter dance by keyword or category. while a similar method is used by the 14 iassist quarterly summer 2007 current dance iteration, the new version will allow for combining filters. the filtering is designed to help the user drill down to the desired dataset. ajax search and filtering utilize direct user interaction to yield the desired results. instead of programmatically trying to determine what the user had in mind, the search methods provide instantaneous feedback so the user may alter search criteria accordingly. in addition to searching from the main page, dance is designed to let a user navigate to different datasets through the result set. each keyword, category, contact name, etc, is linked to all the results for a particular field. for example, a result for ‘industrial production’ includes the keyword ‘semi-conductor.’ a user is able to click on this keyword to see all the datasets that have also been tagged with ‘semi-conductor.’ beyond browsing dataset linkages, users may want to know if there have been any changes to a dataset’s metadata or if there are any recently added datasets. the new dance will communicate changes and additions through two methods: 1) information boxes on the main search page and 2) a subscription service. besides the searching and filtering capabilities on the main page, there will be two information boxes with links to the five most recently changed, as well as the five most recently added, datasets. a subscription e-mail service is also available for a user to register for to receive metadata changes or dataset additions. the new dance will also allow for user-generated content through the addition of wiki9 pages for each dataset. each dataset’s wiki page will be populated initially with vendor-provided documentation and user manuals. anyone from the federal reserve system has the authority to modify a wiki page. with each wiki page, dance hopes to capture and share knowledge across the system. in addition to dataset knowledge, dance plans to include links to computer code and published papers. recently, the board set up a test code repository where anyone can figure 1 share and critique code. dance will link code to the repository designed to manipulate the dataset. to complete the research cycle, dance will include links to published papers that use a particular dataset. currently, each published paper is registered via an online form. dance staff will work with the board’s publications department to include a registration field for what datasets were used in the paper. while the wiki pages allow for anyone to add content, the metadata working group concluded that the core metadata items should only be manipulable by either the data owner or an individual authorized by the data owner. further, the working group requested that certain core metadata items only be displayed to board employees10. to implement the working group’s requests, the new dance will use a two-layer security model; one at the application layer and the other at the database layer. the application security will display the authorized fields based upon the user community. the database security will limit metadata editing capabilities to authorized users. any user across the federal reserve system may request a new entry to dance. the new content entry screen will be the same as what is used for the administration dataset registration process. the user’s request will be submitted to dance staff for review, and added to dance after approval. a new feature to the entry form includes fields for the addition of variable-level metadata for the requested dataset. the new dance will accomplish a new level of understanding and collaboration with respect to board (and potentially reserve bank) datasets. the new user interfaces and content distribution methods are designed to easily communicate available data, how to use the data, and published works using the data. while the new dance reduces many dataset-specific silos, seamless knowledge transfer is dance staff’s long term vision. iassist quarterly summer 2007 15 data and knowledge management vision the dance working group also requested that dance be able to return results from other data search tools, most notably fame.11 the group agreed that the most effective search tool would be independent from the metadata or data location, thus fully eliminating board data silos. in addition to a uniform dataset search, dance staff are working with the board’s research library to create an enterprise wide search tool. in this new vision, data, metadata, documents, papers, programs, and library resources will all be indexed and accessible through a universal interface. all result sets will redirect users to the appropriate individual catalogs for more information. furthermore, a flexible design could include indexing reserve bank libraries and resources. one step toward this integrated vision was taken in spring 2008 with the installation of libx.12 libx is a search-tool that is installed on a browser similar to the quick search google toolbar. libx provides one location to search registered repositories; however, each search is repositoryspecific. for example, a user may want to search ‘stock prices’ in the board research library catalog, dance and the e-journal portal. libx requires the user to search each repository separately. while libx accomplishes the uniform location requirement; a more general enterprise search is still being investigated. there is already some degree of system coordination. for example, several of the reserve bank libraries are searchable through a common gateway. further, there is a system-wide search. however, if you search for the term ‘dance’ most of the results consist of social dance clubs, dance lessons or hosted events with dancing. while there is still a bit of tweaking to be done, this search could be altered to target research related resources. along with simplifying the search process, users must be able to access the datasets. data storage issues are the next phase of rddi. an effort is underway to test storing several managed datasets in a relational database server instead of as individual sas datasets. relationally stored data may be combined with dance to create a data retrieval tool with embedded metadata. the advantage of this setup is that users may build datasets in their preferred format instead of the storage format. the increasing size and number of accessible datasets places further importance on the ability to manipulate data into a research-friendly format. conclusions economic research is increasingly multi-faceted, and, as a result, research projects are extremely data-intensive. therefore, metadata documentation and standards have evolved significantly since dance’s inception. dance started as a communication tool identifying the existence of a dataset with minimal descriptive characteristics. dance’s popularity, along with new data initiatives, led to several metadata enhancement requests to provide deeper searching and administrative capabilities. multiple open-source solutions were explored to handle new metadata items and communication requirements. however, technical requirements, as well as additional business requirements, lead to a custom-built solution for the new dance. the custom solution enabled new features which focused on the ability to foster collaboration between departments. the new features allow users to seamlessly traverse from dataset metadata, to security request forms, to variable-level metadata without changing applications. these efficiencies will enable researchers and regulators system-wide to focus on data analysis and searching for the data. as a result, it is expected that this focus will translate into even more effective policy making benefiting the entire financial and economic systems.. appendix 1 – federal reserve system structure the federal reserve system is comprised of the board of governors and twelve reserve banks. the federal open market committee (fomc) is made up of the board of governors and presidents of the reserve banks. the board of governors, in washington, dc, provides the leadership for the entire system. both the reserve banks and the board conduct monetary and economic research. the combined board and reserve bank research assists the fomc in monetary policy decisions. 16 iassist quarterly summer 2007 appendix 2 – original dance metadata fields metadata field description database the commonly accepted name or title of the database or dataset hyper linked to database specific documentation. vendor the primary vendor name hyper linked to the appropriate website. division(s) the division and co-owning division if applicable. form of access a general description of the network location and access software required. status a boolean indicating whether the board still purchases or maintains the database. description an executive summary of the database as well as any board specific elements. keyword(s) dance staff assigned general descriptive words. data contact the name and phone extension of the primary individual(s) with expertise using the database. license contact the name and phone extension of the individual holding the license agreement. license information one of the board security classifications or a contact person if express consent is required. category a dance staff assigned general data grouping used for search purposes. cost usd cost at the time of purchase. iassist quarterly summer 2007 17 appendix 3 –new dance metadata fields metadata field dublin – core name description title title the commonly accepted name or title of the database or dataset hyper linked to database specific documentation. title short title abbreviated name for drop down search creator creator the primary dataset creator: could be an internal individual or external company publisher publisher the entity that makes the dataset available to the public. may or may not be creator. vendor vendor entity that sold the data to the board data contact data contact primary contact(s) for data questions data requestor data requestor individual(s) who requested the dataset contributing section contributor group(s) with expertise data originationsource original data source; may or may not be the creator or publisher. keyword subject frequently used descriptive terms description description dataset abstract date created date date the dataset was first created at the frs. date range available available the date range the data is available from the vendor. geographical coverage coverage the physical locations the dataset covers. type type abstracted keywords; i.e. micro economics, macro economics. data location physical storage location at the frs. output format format output file format; example: sas dataset input format medium input file format; example: tab-delimited text file identifier identifier auto-generated number uniquely identifying a dataset. data confidentiality rights frs security classifications assigned to the dataset. bibliographic notation bibliographic citation how the dataset should be cited in published works. related resources relation related datasets license agreement license a scanned pdf of the user license. license owner rights holder the individual(s) responsible for the license additional information bulletin board the wiki page for the dataset. update method accrual method method in which the dataset is updated at the frs. update scheduleaccrual periodicity frequency that the dataset is updated at the frs. dataset status accrual policy an active or inactive flag purchasing division division division(s) purchasing the dataset. purchasing section section section(s) purchasing the dataset product url product website url product website vendor url vendor website url vendor website vendor contacts vendor contacts technical and sales contacts for the dataset. cost cost usd dollar cost payment schedule frequency of payment schedule that payments are to be made. contract renewal date contract renewal date the date the contract is to be renewed. purchase justification purchase justification budget justification for purchasing the dataset. purchase order purchase order number linking the dataset to the procurement purchasing system. reason needed need the economic or regulation reason for purchasing the dataset. sole source sole source a flag indicating if the dataset is only available from one vendor. contract length contract length length of time the contract is valid for. record last update modification history log of the time in which a record was modified. record updated by catalog entry maintained by log of individuals who modified a record. 18 iassist quarterly summer 2007 appendix 3 –new dance metadata fields ..(cont) metadata field dublin – core name description title title the commonly accepted name or title of the database or dataset hyper linked to database specific documentation. title short title abbreviated name for drop down search creator creator the primary dataset creator: could be an internal individual or external company publisher publisher the entity that makes the dataset available to the public. may or may not be creator. vendor vendor entity that sold the data to the board data contact data contact primary contact(s) for data questions data requestor data requestor individual(s) who requested the dataset contributing section contributor group(s) with expertise data originationsource original data source; may or may not be the creator or publisher. keyword subject frequently used descriptive terms description description dataset abstract date created date date the dataset was first created at the frs. date range available available the date range the data is available from the vendor. geographical coverage coverage the physical locations the dataset covers. type type abstracted keywords; i.e. micro economics, macro economics. data location physical storage location at the frs. output format format output file format; example: sas dataset input format medium input file format; example: tab-delimited text file identifier identifier auto-generated number uniquely identifying a dataset. data confidentiality rights frs security classifications assigned to the dataset. bibliographic notation bibliographic citation how the dataset should be cited in published works. related resources relation related datasets license agreement license a scanned pdf of the user license. license owner rights holder the individual(s) responsible for the license additional information bulletin board the wiki page for the dataset. update method accrual method method in which the dataset is updated at the frs. update scheduleaccrual periodicity frequency that the dataset is updated at the frs. dataset status accrual policy an active or inactive flag purchasing division division division(s) purchasing the dataset. purchasing section section section(s) purchasing the dataset product url product website url product website vendor url vendor website url vendor website vendor contacts vendor contacts technical and sales contacts for the dataset. cost cost usd dollar cost payment schedule frequency of payment schedule that payments are to be made. contract renewal date contract renewal date the date the contract is to be renewed. purchase justification purchase justification budget justification for purchasing the dataset. purchase order purchase order number linking the dataset to the procurement purchasing system. reason needed need the economic or regulation reason for purchasing the dataset. sole source sole source a flag indicating if the dataset is only available from one vendor. contract length contract length length of time the contract is valid for. record last update modification history log of the time in which a record was modified. record updated by catalog entry maintained by log of individuals who modified a record. iassist quarterly summer 2007 19 footnotes 1. the opinions are of the author and not the federal reserve board. andy boettcher, board of governors of the federal reserve system, washington, d.c. contact: andrew.s.boettcher@frb.gov. this article is based upon a presentation at the iassist 2008 conference at stanford. 2. market discipline is one of the three pillars of the basel ii banking accords. please see http://www.bis. org/publ/bcbsca.htm for details. 3. fractured data management practices have decreased since data cataloging started and are further reduced with recent data initiatives, described later in the paper. 4. please see appendix 1 for brief description of the federal reserve system’s structure. 5. please see appendix 2 – original dance metadata fields for a list and descriptions. 6. please see appendix 3 – new dance metadata fields for a metadata list and descriptions. 7..replaces refines the relationship variable, for example, dataset b was purchased to replace dataset a. 8. ajax stands for asynchronous javascript and xml. see http://www.google.com/webhp?complete=1&hl=en for an example. 9. wikis are software solutions that allow anyone to add or edit content. please see http://en.wikipedia.org/wiki/ main_page 10. see appendix 3 for which user community can view which items. 11. the entire macroeconomic side of research uses fame for data storage. fame provides basic metadata for each stored series. 12. libx is an open-source search tool written by virginia tech. www.libx.org 1/12 mandikiana, brian w., timms-ferrara, lois and maynard, marc (2019) data archiving for dissemination within a gulf nation, iassist quarterly 43(3), pp. 1-12. doi: https://doi.org/10.29173/iq943 data archiving for dissemination within a gulf nation brian w. mandikiana1, lois timms-ferrara2, and marc maynard3 abstract since 2008, qatar university’s social and economic survey research institute (sesri), has been collecting nationally representative survey data on social and economic issues. in 2017, sesri leadership established an archiving unit tasked with data preservation and dissemination both for internal purposes and with the intent of disseminating select data to the public for secondary analysis. this paper reviews the lessons learned from creating a data archive in an emerging economy where both cultural and political sensitivities exist amid diverse groups of stakeholders. challenges have included recruiting trained personnel, developing policies for data selection and workflow objectives, processing restricted and non-restricted datasets and metadata, data security issues, and promoting usage. additionally, there is hope that the presence of the archiving unit adds value for other sesri research staff involved in the design, collection, documentation, and processing of studies. after successfully addressing these challenges over the past year, the archive met its objective to launch a data center at the institute's website (http://sesri.qu.edu.qa) and to make multiple datasets available for public download from it. also, to be discussed are the tools, processes and leveraging of resources that are being implemented as the archiving process continues to evolve. keywords data archive, qatar, middle east, open data, digital preservation workflow, repositories 1. introduction the middle east and north africa (mena) region presents a unique environment for survey research throughout the data lifecycle. research institutes in the region face the challenges of navigating cultural and political sensitivities and the availability of research capabilities. these research institutes also have to figure out how these aspects among other factors affect planning and design, data collection, data processing, data dissemination, and data re-use. in qatar, considerable progress has been made in promoting survey research (gengler, le, and howell, 2018), and has resulted in the collection of several studies. similarly, over the past decade, many research organizations have emerged within the gulf region and the mena region as a whole, generating dozens of datasets. given the insufficient efforts in data sharing initiatives across the gulf region, access to data for replication of published research output to confirm findings is limited. data sharing throughout this region would improve the credibility of statistics generated from the region, improve relevance and reliability of such data, lessen data collection costs incurred due to fielding similar surveys, and build a body of knowledge that policy-makers could use in addressing various issues facing the region. https://doi.org/10.29173/iq943 2/12 mandikiana, brian w., timms-ferrara, lois and maynard, marc (2019) data archiving for dissemination within a gulf nation, iassist quarterly 43(3), pp. 1-12. doi: https://doi.org/10.29173/iq943 government and in particular, national statistics offices throughout the mena region have been the driving force for survey data collection, data processing and sharing of emerging statistics. as more government offices and data collection agencies within the region are adopting open data and fair (findable, accessible, interoperable, and re-usable) data principles, the sharing of data and related metadata is gaining momentum (saxena, 2018). organizations beyond government (private foundations, companies, research centers) are beginning to sponsor data collection, and these institutions require a review of existing data to inform and supplement their efforts on behalf of their stakeholders (elsayed and saleh, 2018). although findings from a recent study by shaon, straube, and chowdhury (2017) highlight deficiencies in capabilities, particularly skills required to execute data sharing activities, efforts to address the lack of data sharing are underway, beginning with the establishment of the first open data archives for survey research in the region. below we give a brief description of the social and economic survey research institute (sesri), an independent survey research organization that recently established a data archive and how it has approached promoting data sharing given the contextual challenges that the organization faces. 1.1 the social and economic survey research institute a brief history the social and economic survey research institute (sesri) is an independent academic research organization at qatar university. since its inception in 2008, with the assistance of the institute for social research (isr) based at the university of michigan, it has developed a robust survey-based infrastructure in order to provide high-quality survey data for planning and research in the social and economic sectors. the data are intended to inform planners and decision makers, as well as the academic research community. since 2008 sesri has conducted more than 50 household surveys on such topics as entrepreneurship, education, food security, marriage delay, social capital, labor, migration, tourism, health, among others. also, the on-going omnibus surveys that cover many timely subjects. with the development of a fully operational and sustainable survey research program at qatar university, sesri has accomplished much in line with its purpose and mission. the mission of the social and economic survey research institute (sesri) is; “to contribute to the development of qatari society by providing high-quality survey data to guide policy formulation, priority setting, and evidence-based planning and research in the social and economic sectors.” the above mission statement highlights the importance of data access. it is from such a background that the need for a microdata data archive emerged. 1.2 sesri data archive: a (very) brief history founded in 2017, the sesri data archive, housed at qatar university is one of the few microdata archives in the persian gulf region providing access to high-quality socio-economic household survey data. the archive was created due to the recognition of the need to maintain and re-use sesri’s own created resources, and the desire to share data to advance evaluation of social and economic policies in qatar, the gulf region, and beyond. furthermore, angel-urdinola, hilger, and ivins (2011) https://doi.org/10.29173/iq943 3/12 mandikiana, brian w., timms-ferrara, lois and maynard, marc (2019) data archiving for dissemination within a gulf nation, iassist quarterly 43(3), pp. 1-12. doi: https://doi.org/10.29173/iq943 highlighted the need to enhance access to microdata in the middle east and north africa region. since its establishment, the archive has supported sesri survey data management, preservation, internal use, and external dissemination. the sesri data archive's mission is: “to advance globally accepted standards of economic and social survey research data for; storage, management, preservation, and dissemination." few studies (angel-urdinola, hilger, and ivins, 2011; saxena, 2016; saxena, 2017; thompson, 2009) report microdata archiving initiatives in the middle east and north africa or the gulf region. in this paper, we describe a case study of the establishment and development of the sesri data archive in the state of qatar and explore the approach and challenges faced during the first year of operation. after briefly reviewing the state of survey data sharing in the gulf region, we preview the challenges confronted and, in some cases, anticipated, followed by initial development of the archive in terms of staffing and working policy development. next, we review the development of the archive data processing workflows. we conclude with some brief comments on the partners' plans for strengthening the infrastructure of the archive and promoting its use and growth. 2. data sharing culture in the gulf region availability of microdata from countries in the persian gulf region has been limited both in terms of comparative analysis across countries, but also within individual countries. several outside efforts to investigate views of societal structures and policy perspectives have taken place over the years. some of the well-established include; the arab barometer (2005-ongoing), the world values survey (1981ongoing), and more recently the pew global attitudes project (2001on-going). survey research efforts to collect and publicly share microdata within individual countries have been less ambitious and met with moderate success (angel-urdinola, hilger, and ivins, 2011; saxena, 2017). although data collection through household survey has been a standard way of supporting governments' decisionmaking processes in the region, data sharing has not been a critical feature in the persian gulf. other institutions involved in social science survey research include the emirates center for strategic studies and research (ecssr) and the bahrain centre for studies and research, bahrain (bcsr). however, when looking at the mena region, notable successes include the economic research forum (erf), focusing on the middle east and north africa (thompson, 2009). in qatar, recent noteworthy data sharing developments have emerged. first, the ministry of information communication and transport in qatar published the open data policy (ictqatar, 2014). this policy stresses the importance of making data available, while at the same time using the linked open data model as a guide to further the degree of openness. it is argued that a data sharing culture will facilitate not only efficient public service delivery but building a less extensively hydrocarbonbased economy that is knowledge-based and shares recent data with the public in an accessible manner (ictqatar, 2014). second, in addition to the open data policy, the data management policy was published to inform issues of data handling in the state of qatar. among other issues highlighted is the need for professionals to support data archival related processes. these policies demonstrate the leadership’s willingness to support a culture of open data sharing in the state of qatar. https://doi.org/10.29173/iq943 4/12 mandikiana, brian w., timms-ferrara, lois and maynard, marc (2019) data archiving for dissemination within a gulf nation, iassist quarterly 43(3), pp. 1-12. doi: https://doi.org/10.29173/iq943 the social and economic survey research institute (sesri) was founded on a principle of data sharing for research and policy planning purposes. furthermore, the research institute can provide a model for how microdata and associated documentation may be shared from within a unique and sometimes challenging environment. sesri is in a unique position in year ten of its research program. it has conducted dozens of high quality, timely and substantively relevant survey research projects that have already seen primary use in public policy planning in qatar, as well as, in peer-reviewed policy-oriented academic journals (al-emadi et al., 2017; diop et al. 2016; gengler and mitchell, 2018). sesri management sees opportunities presented by providing materials for reproducible research and the growth in open-access publishing. the underlying infrastructure for data sharing is still in its infancy. policies to support data use, data management and access have not been fully developed. internal research networks are still nascent and need further development and encouragement. 3. challenges as in many cases, the great promise can be accompanied by significant and diverse challenges as the sesri data archive set down its development path. challenges faced by the archive can be grouped into three areas: technical, organizational, and contractual or legal. technical issues tend to be common to many digital preservation efforts and focus mainly on the status of data and metadata resources and the ability to harness those resources to provide valuable materials for future researchers. organizational and contractual challenges can be unique to a particular context and may require distinctive approaches and innovative solutions to address. the solutions to these challenges can be found by looking at exemplary institutions and proven standards and best practices. 3.1 technical challenges survey metadata is compiled from a variety of sources including published summary results, survey reports, and methodological reports. recovery of complete and updated survey data and documentation files and supporting metadata provided the first primary challenge to the archive. since the surveys were conducted many years ago and are subsequently used for not only primary analysis but also reused in succeeding analyses, identifying complete and authoritative version files required additional effort working with principal investigators (pis), analysts and information technology (it) staff. unfortunately, standard file-naming conventions and version control were not implemented across survey projects and, further, multiplicities4 or multiple versions of the same file were found and needed to be rectified. while files were stored on a shared network drive, external contextual metadata was not captured. the lack of unique and permanent study identifiers also hindered the development of an authoritative survey inventory. for each survey, several relevant files had to be reviewed and evaluated for completeness based on consultation with appropriate research unit staff, technical evaluation of file formats and comparisons of datasets and related documentation (i.e., questionnaires, blaise export, etc.). this tended to be an iterative process due to the loss of context because sufficient metadata was not captured originally at its creation. blaise is the computer-assisted telephonic interviewing (cati)/computer assisted personal interviewing (capi) program utilized at sesri and while blaise exports both arabic and english questionnaires, if available, these files needed to be cross-referenced to typically incomplete englishonly variable labels as stored in the exported stata data file (.dta). https://doi.org/10.29173/iq943 5/12 mandikiana, brian w., timms-ferrara, lois and maynard, marc (2019) data archiving for dissemination within a gulf nation, iassist quarterly 43(3), pp. 1-12. doi: https://doi.org/10.29173/iq943 concerning dissemination efforts, sesri had no external data dissemination channel other than the availability of publications via its main website. 3.2 organizational challenges throughout its research, sesri had developed a generally accepted set of procedures concerning data management throughout the lifecycle of the project, but did not have formal data management policies to guide data use beyond the completion of the project and publication of the final report(s). research and policy units worked in tandem, but data preservation and archival work were not built into the process. lead principal investigators and research assistants worked with the resulting data primarily for analytical purposes and developing scholarly research outputs and did not see the value of maintaining a type of persistent, reusable version of the resulting data files. this situation is not unique; many survey research organizations complete a project and move on immediately to the next one and do not have time to describe and document the just-completed project. the challenge for the archive unit was to re-define the division of labor and re-configure the requirements around the survey files hand-off from the research unit. 3.3 legal and contractual challenges sesri is committed to preserving the confidentiality and privacy of individual survey respondents. as with all its survey projects, protecting respondent confidentiality and privacy is critically important. the archive unit extends this protection beyond the initial collection and use of the data from individual interviews to data re-use and exposure beyond sesri staff. various approaches are employed to enforce this commitment including the creation of a public release version of datasets that suppress appropriate variables or uses other mechanisms to mitigate the risk of violating response confidentiality. additionally, contractual arrangements with funding partners can impose restrictions on the preservation, accessibility, and use of sesri datasets. each survey must undergo a thorough review of underlying funding commitments and business considerations as a critical component when determining if and how the data can be documented, stored and released. 4. recruitment and staff skilled human capital is one of the essential factors in building a data archive. it is recommended that a minimum of four dedicated staff be recruited for a well-functioning data archive. currently, the sesri data archive is under-resourced. the lead archivist has the primary role in managing the day-to-day operations of the data archive. although recruiting more staff was approved, securing local hires has been a challenge for several reasons. first, the concept of microdata archiving is commonly unknown or misunderstood by most people within the region. second, few institutes equip students with a set of skills that are essential for data archiving. finally, academic qualifications such as information technology, information systems, statistics, are not as prominent as engineering studies in the gulf region. on the other hand, recruiting non-local staff to take up the posts is challenging for many reasons. due to changes in policy, there has been more focus on hiring locally. however, only when that fails are international candidates considered. in as much as that provision is available, getting https://doi.org/10.29173/iq943 6/12 mandikiana, brian w., timms-ferrara, lois and maynard, marc (2019) data archiving for dissemination within a gulf nation, iassist quarterly 43(3), pp. 1-12. doi: https://doi.org/10.29173/iq943 multilingual candidates with working knowledge of arabic is yet another challenge. this adds to the length of the recruitment process, and as a result potential good candidates end up getting hired by other companies while the recruitment process is still in progress. for the past one year since the establishment of the sesri data archive, the existing lead data archivist together with external collaborators from other organizations have worked together on many items essential for the normal operation of the data archive. also, financial support, it services support, and human resources services have been provided within sesri and by other departments within the housing institution, that is, qatar university. 5. approach and architecture development of the sesri archive proceeded on two tracks: general policy development and practical data processing workflow implementation. on the policy track, general information gathering and best practice reviews were conducted with a particular focus on generally accepted frameworks to learn and evaluate broad organizational structures and policies. this included a review of data management and preservation policies from a variety of data archives and preservation institutions. data management policies for primary data collection organizations were reviewed and discussed. from these, a draft data management policy was created and discussed with sesri administration and presented in a training setting to sesri staff (research unit, policy unit and it). in summary, the data management policy sets out to achieve the following objectives: to support research collaboration through efficient access to survey data; to warrant that sesri confirms with requirements of funding bodies, especially qatar national research foundation. the national digital stewardship alliance (ndsa) levels of preservation (ashenfelder, 2016) were subsequently reviewed to provide a departure point for further discussions around minimal and aspirational archive goals. the ndsa levels lend themselves to this type of evaluation. in all discussions it was recognized that the archive unit should adhere to and implement methods to meet the minimal requirements (ndsa level one) and develop plans and actions to reach aspirational goals along the continuum on various dimensions, including storage location, data integrity, metadata development, standardized file formats, as well as, data access. a reasonably aggressive schedule to release up to four studies in the first nine months of the archive was set early in the process. these first four studies allowed for discussion around general issues of public release both on the administrative level as well as within the archive team. exploring multiple dimensions of concerns for releasing microdata for the first time required a thoughtful and deliberate process on the one hand, but also a flexible and iterative process on the other. clear communications with stakeholders (including management, unit heads, and pis) were critical. documentation of decision-making processes was also crucial. https://doi.org/10.29173/iq943 7/12 mandikiana, brian w., timms-ferrara, lois and maynard, marc (2019) data archiving for dissemination within a gulf nation, iassist quarterly 43(3), pp. 1-12. doi: https://doi.org/10.29173/iq943 once the studies were selected and agreed upon, an efficient data processing workflow needed to be identified and developed. the team divided this process into three processing phases: 1) pre-ingest evaluation and initial preservation; 2) data processing for preservation; and 3) data processing for public release. it was understood that the public release files would not include confidential variables, contractually restricted variables, as well as variables that are restricted for other reasons. internally, though, the archive unit's charge is to preserve data and documentation from all sesri projects regardless of whether a public release version would be made available. archive staff used the open archival information systems (oais) reference model (consultative committee for space data systems, 2012) for thinking about and planning the processing workflow. while work had already begun on several data files, process tasks were identified and reviewed for inputs, outputs, and dependencies. these were mapped to the oais model for further discussion and modification. 6. data processing workflow data processing, both in terms of data cleaning and metadata development, are critically important to the overall success and sustainability of any data archive. consistent, reliable and reproducible actions are threaded together to provide standard pathways that can be managed efficiently. the ability to successfully clean and preserve one specific data resource must be, itself, preserved through testing and documentation. the short-term effort involved in thinking through required steps, identifying dependencies, and documenting their effectiveness, pays dividends in the long-term as the data collection and the archive unit grows. for the sesri archive workflow, data processing was designed around capabilities and features of stata5, and metadata development efforts rely on colectica6 designer file naming conventions and network file storage structures were established and documented taking into consideration the anticipated data collection scope and versioning. data cleaning and modularized, standardized labeling and recoding schemes were designed and documented by archive unit staff to provide consistent and reliable data handling routines. additionally, guidelines were established concerning code documentation and recording of decision-making rationale for future reference. the current sesri archive unit processing workflow (see table 1) can be thought of as an iterative process with three major phases: 1) pre-ingest review; 2) data processing for preservation; and 3) data processing for public release. each of these phases had several quality control reviews built-in, and each ultimately ends with a newly preserved data resource and appropriate documentation. https://doi.org/10.29173/iq943 8/12 mandikiana, brian w., timms-ferrara, lois and maynard, marc (2019) data archiving for dissemination within a gulf nation, iassist quarterly 43(3), pp. 1-12. doi: https://doi.org/10.29173/iq943 table 1. sesri workflow summary 1 pre-ingest review submission information package (sip) evaluation; review for confidentiality and potential public release ingest study level cataloging; file level technical review preservation files to archive storage; metadata to repository 2 data processing initialize administrative checkpoints; build study metadata; dataset checks and cleaning; build variable metadata; record actions data quality review review dataset and processing log codebook creation generate and review english and arabic codebooks general quality review review codebook and dataset(s) based on processing log preservation files to archive storage; metadata to repository 3 public release screening policy review; contractual review; confidentiality review; variable level screening; final approval prepare public use files remove restricted cases, variables, administrative variables, etc.; prepare usage notes; citation requirements; review terms and conditions preservation files to archive storage; metadata to repository concerning public data release screening, an advisory committee provides oversight and guidance to the archive unit to balance respondent confidentiality concerns and contractual restrictions with the overarching goal to provide open access to sesri research output. the archive unit employs a variety of methods to protect confidential data including de-identifying or anonymize sensitive information and aggregating indicators at appropriate geographic levels. 7. data availability in order for the archive to advance appropriately sesri's goal of sharing data with scholars and thereby contributing to scientific research on important issues, a data portal was created offering public release versions of the metadata and datasets to select surveys, with access provided directly from the website. this portal would become the data centre. 7.1 data center portal for dissemination the sesri website had been well established and designed to introduce the institute to a variety of stakeholders including the qatar university community, policy analysts around the country, media outlets, and the broader research community interested in qatari public opinion. in an organized format, the site is easily navigated with links provided to key analytical reports, sesri sponsored events, training workshops. the data center, however, was a different endeavor for the website and communications unit. as a primary research facility within qatar university, known as the national institute of higher education in qatar, there are inherent dependencies and relationships with stakeholders that must be cultivated across departments and takes time to achieve. one unanticipated mandate was that qatar university web services had to build the data center website and upload the prepared studies https://doi.org/10.29173/iq943 9/12 mandikiana, brian w., timms-ferrara, lois and maynard, marc (2019) data archiving for dissemination within a gulf nation, iassist quarterly 43(3), pp. 1-12. doi: https://doi.org/10.29173/iq943 for dissemination to assure adherence to standards and branding protocols. this involved significant involvement of the sesri communications and archive staff with the university web services team to realize a fully functional data center that in january 2018 became the official public face of the sesri data archive. archive staff is actively adding materials to the data center website with new studies becoming available every two to three months. there are extensive plans for expanding the website to include efficient data discovery tools, including a catalog of publicly released studies once there is a substantial collection uploaded. 7.2 user and access support staffing requirements for providing researcher support are being assessed and is anticipated to grow with the size of the public offerings. a list of frequently asked questions (faqs) has been, and a variety of additional support documentation is being considered as more information is gathered on user needs. plans are in place to secure doi space to assign unique identifiers and manage version support. 7.3 promotion marketing the archive a multistage approach for promoting the use of the archive has been developed, but not yet tested. the plan involves a combination of traditional outreach with public and targeted announcements, press releases, and email communications, along with components of social media. the plan further invites partnerships with organizations of mutual interests and a set of planned paid promotions. 7.4 measuring success given institutional objectives to share data among an intellectual community of research scholars in order to inform discussions that advance social science, the metrics for measuring success include both short term and long term outcomes. in the short run, success will be gauged by data center traffic reports, detailed records of downloaded materials, and the amount of interest in the qatari opinion as measured by an optional form to be added to the mailing list. as the data center becomes more complete, the team plan to build a bibliography of scholarly articles, conference papers, and teaching materials that utilize sesri data and develop a portal for submitting completed works and citation links. 8. discussion the discussion above has shown the processes involved in building a microdata archive. in summary, to increase the success of the data archive establishment, many factors need to be considered. financial investment research institutes and organizations need to plan for the financial investment required for data dissemination. financial investments may range from the cost of data documentation, special handling of restricted data, expert consultation fees to address particular issues, and appropriate data processing of data awaiting dissemination. in situations where the research institution is involved in all parts of the data lifecycle, there are more cost-cutting opportunities. https://doi.org/10.29173/iq943 10/12 mandikiana, brian w., timms-ferrara, lois and maynard, marc (2019) data archiving for dissemination within a gulf nation, iassist quarterly 43(3), pp. 1-12. doi: https://doi.org/10.29173/iq943 workflow documentation data processing workflows must guide data processing staff in a manner that supports consistency and reproducible research. furthermore, similar to data files, it is important to implement good preservation of syntax used for data processing. annual reviews of workflows are equally important. updating these processes can significantly contribute to efficiency. confidentiality a common concern amongst many research institutions is meeting the requirements for maintaining confidentiality. data centers planning to share data will need to invest time in this process. notably, besides removing data such as respondent residential areas, information on the day, month, and year of birth, names, other exploratory statistical techniques and aggregation can be applied to reduce the risk of re-identifying survey participants. data publications many researchers have proposed data publications as a way of promoting data sharing. this strategy could help unlock data enclaves in the gulf region were for many year data sharing has not been part of the norm. hsu et al. (2015) highlights the rapid rise in the data publications as a concept and incentivizing strategy for data sharing. however, for this strategy to work hsu et al. (2015) proposes making data papers part of the mainstream academic output, such as journal articles. training a feature of most statistical related training programmes is that students usually work with clean data. in order to build capacity, students should receive training data management as well as sound exploratory data analysis techniques. similar to shaon, straube and chowdhury (2017), in qatar and other countries in the region, research institutions need to provide training opportunities to students. besides, data centers should encourage researchers to participate in data management related programmes as a way of retooling as well as keeping up to date with the evolving statistical programming software. 9. conclusion at sesri the need for an institutional social sciences data archive has been recognized. qatar's open data policy and data management policy suggest that there is much needed support at the national level to promote data sharing. given the limited microdata archives in the gulf, the oais reference model is critical in promoting a better understanding of archival processes in the region. data archiving is beneficial to the growth of quantitative research in the gulf. equally important, conducting this case study has allowed us to share insights on setting up a data archive, the challenges involved, and lessons learned in the process. ways to improve prospects for data sharing could include data management plans with understandable terms of use for gathered data files. equally important, particularly in the social science field, sharing data anonymizing techniques can contribute towards reducing the risks of data sharing. moreover, when data archive staff offer training on best practices and standards for efficient hand-offs to depositors, it dramatically improves ingest processes. https://doi.org/10.29173/iq943 11/12 mandikiana, brian w., timms-ferrara, lois and maynard, marc (2019) data archiving for dissemination within a gulf nation, iassist quarterly 43(3), pp. 1-12. doi: https://doi.org/10.29173/iq943 the potential for growth in data sharing exists in gulf countries. re-use of nationally representative household survey data needs to be encouraged in gulf states to further data quality, smoother access and metadata documentation that supports analysis. research institutions in the gulf need to build on the sparse data dissemination programmes that have been established thus far and making data related products interoperable. these are foundations upon which they may build data archives or at least deposit survey data with established archives in the gulf region. thus, we need to start implementing archival processes, tap into innovative infrastructures that are emerging to preserve survey data for generations to come. references al-emadi, a., kaplanidou, k., diop, a., sagas, m., le, k.t. and al-ali mustafa, s. (2017) ‘2022 qatar world cup: impact perceptions among qatar residents’, journal of travel research, 56(5), pp. 678694. angel-urdinola, d.f., hilger, a. and ivins, i.b. (2011) ‘enhancing access to micro-data in the middle east and north africa’ arab world brief, 4. washington, dc: world bank. online: https://openknowledge.worldbank.org/handle/10986/9451 ashenfelder, m. (2016) expanding ndsa levels of preservation. available at: https://blogs.loc.gov/thesignal/2016/04/expanding-ndsa-levels-of-preservation/ (accessed: 14 march 2018). consultative committee for space data systems (ccsds) (2012) reference model for an open archival information system (oais). magenta book. issue 2. june 2012. https://public.ccsds.org/pubs/650x0m2.pdf diop, a., le, k.t., johnston, t. and ewers, m. (2016) ‘citizens’ attitudes towards migrant workers in qatar’, migration and development, 6(1), pp. 144-160. doi: 10.1080/21632324.2015.1112558 diop, a. al ansari, m., le, k.t., elmaghraby, e., al bloshi, a., al qassas, h., al khulaifi, b. and mustafa, s. (2017) ‘from the "fareej" to metropolis: qatar social capital survey ii.' sesri, qatar university. available at: http://hdl.handle.net/10576/6386 (accessed: 20 january 2018) elsayed, a. m. and saleh, e. i. (2018) ‘research data management and sharing among researchers in arab universities: an exploratory study.’ ifla journal, 44(4) pp. 281–299. doi: 10.1177/0340035218785196. gengler, j., le, k. t., and howell, d. (2018) ‘survey challenges and strategies in the middle east and arab gulf regions.’ in t. p. johnson, b.-e. pennell, i. a. l. stoop, and b. dorer (eds.), advances in comparative survey methods: multinational, multiregional and multicultural contexts (3mc) (pp. 555-568). new york, ny: john wiley. gengler, j. and mitchell, j.s. (2018) ‘a hard test of individual heterogeneity in response scale usage: evidence from qatar’, international journal of public opinion research, 30(1), pp. 102– 124.available at: https://doi.org/10.1093/ijpor/edw025 (accessed: 20th january 2018) https://doi.org/10.29173/iq943 https://openknowledge.worldbank.org/handle/10986/9451 https://blogs.loc.gov/thesignal/2016/04/expanding-ndsa-levels-of-preservation/ https://public.ccsds.org/pubs/650x0m2.pdf http://hdl.handle.net/10576/6386 https://doi.org/10.1093/ijpor/edw025 12/12 mandikiana, brian w., timms-ferrara, lois and maynard, marc (2019) data archiving for dissemination within a gulf nation, iassist quarterly 43(3), pp. 1-12. doi: https://doi.org/10.29173/iq943 houghton, b. (2016) ‘preservation challenges in the digital age’, d-lib magazine, 22(7/8), pp. 1-6. doi: 10.1045/july2016-houghton hsu, l., martin, r.l., mcelroy, b., litwin-miller, k. and kim, w. (2015) ‘data management, sharing, and reuse in experimental geomorphology: challenges, strategies, and scientific opportunities’, geomorphology, 244, pp. 180-189. ictqatar (2014) open data policy. ministry of information communication and transport, qatar. available at: http://www.ictqatar.qa/en/documents/document/open-data-policy (accessed: 20 january 2018) ictqatar (2015) data management policy. ministry of information communication and transport, qatar. available at: http://www.ictqatar.qa/en/documents/document/data-management-policy (accessed: 20 january 2018)saxena, s. (2018) ‘drivers and barriers towards re-using open government data (ogd): a case study of open data initiative in oman’, foresight, 20(2), pp.206-218. saxena, s. (2017) ‘open public data (opd) and the gulf cooperation council (gcc): challenges and prospects’, contemporary arab affairs, 10(2), pp. 228-240. saxena, s. (2016) ‘integrating open and big data via ‘e-oman’: prospects and issues’,. contemporary arab affairs, 9(4), pp. 607-621. shaon, a., straube, a. and chowdhury, k. r. (2017) ‘setting up a national research data curation service for qatar: challenges and opportunities‘, international journal of digital curation, 12(2), pp. 146–156. doi: https://doi.org/10.2218/ijdc.v12i2.515 thompson, k. a. (2009) ‘data in development: an overview of microdata on developing countries’, iassist quarterly, 33(4), pp. 25-30. end-notes 1 brian mandikiana is the senior research data analyst / lead archivist at the social and economic survey research institute (sesri), qatar university and can be reached by email: bmandikiana@qu.edu.qa. 2 lois tims-ferrara is the chief executive officer at data independence, ellington, connecticut, usa 3 is the director of technology at data independence, ellington, connecticut, usa 4 preservation challenges in the digital age, http://www.dlib.org/dlib/july16/houghton/07houghton.html 5 stata – www.stata.com (viewed 14 march 2018) 6 colectica – www.colectica.com (viewed 14 march 2018) https://doi.org/10.29173/iq943 http://www.ictqatar.qa/en/documents/document/open-data-policy http://www.ictqatar.qa/en/documents/document/data-management-policy https://doi.org/10.2218/ijdc.v12i2.515 mailto:bmandikiana@qu.edu.qa http://www.dlib.org/dlib/july16/houghton/07houghton.html http://www.stata.com/ http://www.colectica.com/ 1/16 saldanha bach, janete; klas, claus-peter (2024) enhancing fair compliance: a controlled vocabulary for mapping social sciences survey variables, iassist quarterly 48(2), pp. 1-15. doi: https://doi.org/10.29173/iq1118 the creative commons-attribution-noncommercial license 4.0 international applies to all works published by iassist quarterly. authors will retain copyright of the work and full publishing rights. enhancing fair compliance: a controlled vocabulary for mapping social sciences survey variables janete saldanha bach1 and claus-peter klas2 abstract the dynamic relationship among survey instruments and study entities like questionnaires, variables, questions, and response formats evolve in social sciences surveys. researchers may need to modify variable attributes such as labels or names, question-wording, or response scales when reusing variables in survey design. therefore, explaining these relations across different waves and studies is necessary to track how variables relate to each other. although standards like data documentation initiative – lifecycle (ddi-lc) and datacite model these relationships, these frameworks fall short of capturing the complexity of variable relationships. the ddi alliance controlled vocabulary for commonality type employs codes—such as 'identical,' 'some,' and 'none'—to outline shifts in entities like variables; however, this approach is insufficient for disambiguating these relationships since they do not differentiate the variable attributes subject to change. we introduce the gesis controlled vocabulary (cv) for variables in social sciences research data to bridge this gap. this cv is designed to enhance semantic interoperability across various organizations and systems. establishing explicit relationships facilitates harmonization across different study waves and enriches data reuse. this enhancement supports advanced search and browse functionalities. the cv, published via the cessda vocabulary manager, seeks to forge a semantically rich, interconnected knowledge graph specifically tailored for social science research. this endeavour aligns with the fair data principles, aiming to foster a more integrated and accessible research landscape. keywords controlled vocabulary, survey variables social sciences, knowledge graphs, longitudinal surveys introduction and motivation in social sciences, research outputs are increasingly characterized by interdependent entities. these entities encompass a wide range, including surveys, questions, response schemas, datasets, variables, and various data types like audio and video files produced during data collection. among these, variables within quantitative social science datasets emerge as a particularly interesting entity. common variables in social sciences surveys include demographic factors such as age, education level, income, marital status, and more. this first approach is motivated in the context of the consortium for the social, behavioural, educational and economic sciences konsortswd3 of the german national research data infrastructure nfdi4. the konsortswd task area 5-measure-1 project5 provides a technical solution to meet the growing demand for data services within the konsortswd’s research data ecosystem, as klas et al. (2022) highlighted. the konsortswd pid registration service aims to assign pids for individual variables in datasets to make data findability and accessibility on the level of inline data objects of studies more efficient. as https://doi.org/10.29173/iq1118 https://www.konsortswd.de/ https://www.nfdi.de/ https://zenodo.org/communities/konsortswd-ta5-m1 https://creativecommons.org/licenses/by-nc/4.0/ 2/16 saldanha bach, janete; klas, claus-peter (2024) enhancing fair compliance: a controlled vocabulary for mapping social sciences survey variables, iassist quarterly 48(2), pp. 1-15. doi: https://doi.org/10.29173/iq1118 a consequence, variables are the most relevant entities for mapping their relations across waves6 and studies. our initial step in the konsortswd project involves identifying and documenting relations between variables. we aim to store these relationships within metadata in the pid registration service. once a variable is documented and assigned a pid, it can be automatically incorporated into relationship maps, such as knowledge graphs (kgs). examples of large-scale graphs include the research graph for connecting research data repositories, as discussed by aryani et al. (2018), the open research knowledge graph (stocker et al., 2018), and the openaire research graph data model (manghi et al., 2019). given that attributes of variables in social sciences, such as labels, names, question-wording, or response scales, are prone to change, it is essential to offer a transparent explanation of their relationships across different waves and studies. this clarity is necessary to track the evolution and interconnections of variables comprehensively. while existing methodologies employ standards like the data documentation initiative – lifecycle (ddi-lc) and datacite to model these relationships, they often do not fully encompass the intricate nature of variable relationships in social sciences. the ddi alliance controlled vocabulary for commonality type7 , for example, adopts codes such as 'identical,' 'some,' and 'none'— to classify changes in entities like variables. however, this system falls short of distinguishing the specific attributes of variables that change. to bridge this gap, we have developed the gesis controlled vocabulary for variables in social sciences research data (outlined in section 3.1). this controlled vocabulary (cv) is designed to augment semantic interoperability across various organizations and systems. it provides a concise textual identifier for each variable relationship and includes detailed descriptions to elucidate the nature of these relationships. this paper introduces this cv, highlighting its capability to represent these connections with machine-actionable features that facilitate the construction of a kg for social sciences. the cv and the kg are tailored for the detailed granularity required in research data, specifically survey variables. motivation scenario assigning persistent identifiers (pids) to the finer attributes of datasets enables individual elements to be referenced and retrieved, complete with the necessary metadata for both machine-actionable and human access. utilizing pids for referencing research data and their detailed entities aligns with the fair 8 principles of data usage, enhancing data reuse and citation, and facilitating applications in kgs. however, the relationships between variables in datasets are notably complex. these relationships encompass various aspects, including but not limited to different versions of variables, derived formats in subsequent waves, variations in labels and naming, and alternative response schemas in questionnaires and surveys. variables may be added or omitted from wave to wave, influenced by the evolving research questions and objectives of the study. additionally, the types of values variables hold—such as numerals, free texts, or controlled vocabularies—contribute to their differentiation. these properties can also change within the same study's lifecycle. for example, a variable's label might be altered from one wave to another while its underlying concept remains consistent. likewise, the values of variables are subject to updates in their cardinalities, categorization, or response schema and scale, often modified to adapt to study evolution requirements or new sociological approaches. in disciplines like social sciences, economics, and behavioural sciences, which explore areas like the social structure of populations, political attitudes, opinions on various societal aspects, and competencies of adults, such variable attributes are susceptible to shifts in the empirical reality of a changing world. studies in these fields must account for this evolution to maintain relevance in https://doi.org/10.29173/iq1118 https://vocabularies.cessda.eu/vocabulary/commonalitytype?lang=en 3/16 saldanha bach, janete; klas, claus-peter (2024) enhancing fair compliance: a controlled vocabulary for mapping social sciences survey variables, iassist quarterly 48(2), pp. 1-15. doi: https://doi.org/10.29173/iq1118 researching society. in this context, updating variables becomes crucial for accurately measuring transformation and societal dynamics. the data documentation initiative (ddi)9 standards, which are prevalent in social science research, employ a set of controlled vocabularies to aid systems in identifying, locating, and accessing data for research. developed and maintained by the ddi10, these metadata elements are integral for research data management. an example is the metadata element basedonobjecttype11, used in scenarios where a new object is created based on an existing one or when the new object represents more than just a version change. yet, there is a need to reference the original object. this feature is particularly crucial for tracking variable relationships across different waves and studies, as it enables detailed mapping of a variable's evolution. the 'basedonobjecttype' element offers a versatile approach to describe the object further. it can encompass multiple aspects: (a) references to any number of objects that serve as a foundation for the new object, (b) a description of how the content from the referenced object was incorporated or altered, and (c) a code for specific typing of the object in line with an external controlled vocabulary. we have applied a created controlled vocabulary (cv) to specify the relationships of variables using the 'basedonobject' element, enhancing the description with a comprehensive set of variables’ attributes. additionally, we have incorporated elements from datacite, precisely the 'relation_type' metadata field and its subfields, as defined by the datacite metadata working group (2021). we have also integrated properties from schema.org12, such as 'isbasedon', 'isbasedonurl', and 'ispartof', to enrich the metadata further and facilitate robust data management and traceability in social science research. recognizing the data documentation initiative (ddi) as a pivotal standard in the social science community for documenting and managing research data, we have adopted the ddi standard as the foundation for our cv codes. the ddi standard encompasses the entire research data lifecycle and provides metadata elements for describing data sets and related objects such as questions, variables, and values in datasets (thomas et al., 2014). in line with this, we propose to expand the descriptions of relationships, starting with basedonobjecttype13 as an initial approach. since our goal is to track variables across different waves and studies, 'basedonobjecttype' emerges as the most fitting relation, especially when creating an object that represents more than just a version change and requires maintaining a reference to the original object. a key feature of 'basedonobjecttype' is its versionable property ('basedonreference_versionable'), allowing any versionable object to be referenced and repeated across multiple base objects. this flexibility is significant because it will enable unlimited repetition and applicability to any digital object. enhancing the definitions and explanations of these explicit relations leads to improved semantic clarity across and between variables, subsequently enhancing data findability and promoting the reuse of research data. standardized terms through controlled vocabularies enable machineactionable functions, further augmenting kgs. our contributions, based on the existing modelling metadata to describe relation types among entities, are as follows: 1. extend the descriptions to elucidate relations for enhanced semantics, facilitating comparability between variable relations across waves (refer to section 3 for details). 2. develop a controlled vocabulary (cv) for variable relations in the social sciences to boost semantic interoperability across organizations and systems (detailed in section 3). 3. establish a comprehensive framework for identifying relational connections of variables, integrating diverse ddi elements such as survey questions, response schema, data papers, https://doi.org/10.29173/iq1118 https://www.ddialliance.org/about/about-the-alliance https://ddialliance.github.io/ddimodel-web/ddi-l-3.3/composite-types/basedonobjecttype/ https://schema.org/product https://ddialliance.github.io/ddimodel-web/ddi-l-3.3/composite-types/basedonobjecttype/ 4/16 saldanha bach, janete; klas, claus-peter (2024) enhancing fair compliance: a controlled vocabulary for mapping social sciences survey variables, iassist quarterly 48(2), pp. 1-15. doi: https://doi.org/10.29173/iq1118 interactive resources (like codes or scripts using the variable), data management plans, or audio/video data (see section 4 for the complete list). this paper is structured in 5 sections: following the introduction, in section 2 we provide the related literature, further explaining the complex relation between variables with fundamental requirements to support kgs and the associated metadata standards. section 3 exemplifies the knowledge graph of variable ties and provides examples of extended descriptions. section 4 discusses the need to explicitly variables' connections with other entities. section 5 concludes and indicates further efforts. related literature surveys are fundamental in social sciences research and serve as a primary method for investigating variables (babbie, 1990). widely recognized as a critical research paradigm, surveys are instrumental in measuring people’s perceptions, intentions, and behaviours (ajzen and fishbein, 2005). variables, the entities that shape social science data, vary among individuals and over time, reflecting a range of values (kaur and mittal, 2021). attitudinal variables, encompassing beliefs, values, opinions, attitudes, and perceptions on specific topics, are central to major european surveys like the international social survey programme – issp (issp research group, 1992), the european social survey ess14, the national educational panel study – neps (roßbach and neps, 2016) and the socioeconomic panel – soep (liebig et al., 2021), among others. these surveys also frequently explore observed behaviours, frequency of actions, and intentions within target groups. furthermore, cross-domain studies often utilize psychological variables (examining aspects such as personality traits, emotional states, motivation, self-esteem, and selfefficacy) (bollen, 2002) and environmental variables (concerning physical or social environments, resource access, and social support) (cox, 2015), interchangeably for social sciences objectives. in datasets, variables are organized as tabular data, structured in columns15 and rows16 to facilitate manipulation and inference, aligning with the study's objectives. this arrangement is typical for variable data that has been collected, archived, and disseminated. survey variables from these studies encompass a broad spectrum of topics, accumulating vast amounts of data from numerous individuals over extended periods. this leads to the generation of thousands of variable units, necessitating scalable data management solutions for large-scale kgs. for example, the soep-core17 is the main component of the german socio-economic panel (soep). this extensive longitudinal study, ongoing since 1984, annually surveys the living conditions and attitudes of over 15,000 households involving about 30,000 individuals. it represents the most comprehensive long-term study of social developments in germany. within this repository are 560 datasets, encompassing 21,280 questions across 309 instruments and 101,574 variables18. relations between variables beyond managing the enormous quantity of variables, many complex relations are also challenging to interpret through textual analysis. relations within variables extend beyond simple variations and include aspects such as different versions, derived formats in new waves, variations in labels and naming, and alternative response schemas through questionnaires and surveys. the variability of a variable goes beyond just the response options provided by individual cases in a survey. here are detailed examples of how variables' relationships manifest within different attributes: a) variables’ name: often, a variable is related to other variables from different studies. for instance, a study on work-life balance may include a variable named ‘work_life_bal’, which correlates with the variable named ‘job_sat’ in a separate study. despite different names, both variables aim to measure the same concept, such as job satisfaction. https://doi.org/10.29173/iq1118 https://ess-search.nsd.no/cdw/conceptvariables https://www.icpsr.umich.edu/web/icpsr/cms/2042 https://www.icpsr.umich.edu/web/icpsr/cms/2042 https://paneldata.org/soep-core 5/16 saldanha bach, janete; klas, claus-peter (2024) enhancing fair compliance: a controlled vocabulary for mapping social sciences survey variables, iassist quarterly 48(2), pp. 1-15. doi: https://doi.org/10.29173/iq1118 b) survey question: variables are frequently used in survey questions across different studies. for example, the question ‘how satisfied are you with your job on a scale of 1 to 5?’ utilizes the variable ‘job_sat’ to gather data. in another wave or study, the same variable ‘job_sat’ might be reused, but the question could be modified to suit a new response scale, like ‘how satisfied are you with your job on a scale of 1 to 7?’ c) scales: how a variable's answers are represented can differ between studies. a variable like ‘job_sat’ might initially use a likert scale with response options from 1 (strongly dissatisfied) to 5 (strongly satisfied). however, in another wave or a different study, the same variable might be measured with an extended likert scale, ranging from 1 (strongly dissatisfied) to 7 (strongly satisfied). this broader range allows for capturing more nuanced responses, for example: likert scale 5 1. strongly dissatisfied 2. dissatisfied 3. neutral 4. satisfied 5. strongly satisfied likert scale 7 1. strongly dissatisfied 2. moderately dissatisfied 3. slightly dissatisfied 4. neutral 5. slightly satisfied 6. moderately satisfied 7. strongly satisfied panel studies often survey the same individuals or groups repeatedly, measuring the same variables across multiple waves to examine changes in opinions over time. however, the dimensions of these measures or other rules may also evolve. variables across series are related not only in terms of their content but also in their quantity, and a one-to-one relationship is not always present. in some cases, multiple variables from a previous series are merged into a single variable in the next, or a single variable is divided into several. researchers may be interested in determining if a specific variable is consistently present across all time points. in cross-sectional studies conducted at a single point in time, variables may also differ depending on the sample or population studied. for instance, a study focusing solely on college students may include more diverse variables or measures than a study encompassing the general population. transparency in documenting variables and any modifications across different waves or samples is critical, regardless of the study design. variables may be modified due to several factors: a) the research question or study goals: variables are selected to answer specific research questions, which may evolve over time, necessitating the measurement of different variables in later study waves; b) the sample or population: variables can vary across studies or waves depending on the population studied. different target groups, like college students, may have distinct variables compared to other demographic groups; c) measurement instruments or methods: variations in survey questions can alter how variables are measured, resulting in differences across studies or waves; d) the societal, political, or economic environment: broader conditions can influence variables. for example, the soep 19 was expanded in 1990 to include east germany post-reunification and in 2016 to incorporate a sample of refugees; e) data availability or quality: new data sources or improvements in data collection can lead to variations in the variables used across studies or waves. variable documentation is essential for providing transparency and provenance information about the lifecycle of variables in studies. however, deriving insights from these sources can be complex and time-consuming. to address this, controlled vocabularies, as developed by the ddi alliance's controlled vocabularies group (cvg), are vital for defining metadata element meanings, improving https://doi.org/10.29173/iq1118 https://doi.org/10.5684/soep.core.v37i 6/16 saldanha bach, janete; klas, claus-peter (2024) enhancing fair compliance: a controlled vocabulary for mapping social sciences survey variables, iassist quarterly 48(2), pp. 1-15. doi: https://doi.org/10.29173/iq1118 consistency, comparability, and efficiency of documentation, and enhancing information retrieval (jaaskelainen, moschner, and wackerow, 2010). the ddi alliance's controlled vocabulary for commonality type, which aims to describe the degree of similarity between items, uses codes like 'identical,' 'some,' and 'none.' for instance, 'identical' indicates that all variable attributes are the same, while 'some' suggests similarity but not complete identity. however, 'some' does not specify which attributes differ, making it insufficient for disambiguating relationships between variables. a third code, 'none,' indicates an absence of comparability where it was expected. we will describe support standards to document better the reuse and adaptation of variables across waves and studies. modelling metadata standards social science research employs various methods and standards to document studies and enable the tracking of variables across different studies or waves. commonly used standards and best practices include codebooks, data dictionaries, metadata schema, and longitudinal tracking. codebooks provide detailed information about variables and data collected in a study. best practices for creating codebooks involve offering explicit descriptions of variables and their measurements, including coding instructions, recording procedures, and maintaining consistent terminology and formatting. the data documentation initiative (ddi) is the standard file format for codebooks, utilizing extensible markup language (xml) for metadata specification. data dictionaries focus more on the structure and format of the data. they are crucial for ensuring consistent data collection and organization. best practices include providing clear definitions for variables, specifying variable names and labels, and adhering to standard data types and codes. data dictionaries often use standard file formats like ddi and statistical data and metadata exchange (sdmx) 20. metadata standards offer comprehensive information about study design, sampling methods, and data collection procedures. the dublin core metadata initiative21 is a standard format for metadata for digital objects in general, while the datacite metadata schema (datacite metadata working group, 2021) is specifically tailored for documenting research data publication and citation. longitudinal tracking is another technical feature that allows researchers to follow individual respondents over time in longitudinal studies. these features help ensure that the same variable is measured for individuals across different study waves, leading to identifiable persons. this practice includes using unique identifiers for individuals, which requires a higher level of data security and privacy to comply with privacy laws such as the general data protection regulation gdpr (european union, 2016) while using standardized protocols for tracking individuals over time. standard file formats, such as the longitudinal data file (ldf) format, are commonly used for longitudinal data. considering the importance of variables and their relationships, it is vital to describe the associated metadata to register these relationships and enable machine-actionable features through persistent identifiers (pids) and controlled vocabulary terms. the metadata schema is designed to cater to the growing needs for interoperability, data mappings, and knowledge graphs. this solution includes a metadata schema for persistent identification and cross-linking of relationships. metadata for variable relations in the konsortswd project, variable relations are identified primarily through the pid registration service. this process requires detailed metadata both at the study or dataset level and, importantly, at the individual variable level. this metadata is crucial for registering a variable and obtaining its persistent identifier (pid) (saldanha bach, klas, and mutschke, 2023). a vital component of this metadata schema is dedicated to capturing the relationships between variables. https://doi.org/10.29173/iq1118 https://sdmx.org/ 7/16 saldanha bach, janete; klas, claus-peter (2024) enhancing fair compliance: a controlled vocabulary for mapping social sciences survey variables, iassist quarterly 48(2), pp. 1-15. doi: https://doi.org/10.29173/iq1118 one of the central metadata fields in this schema is 'related_item', which describes the resource type, in this case, the variable. each 'related_item' must be accompanied by an identifier provided through the 'related_item_identifier' field. this identifier is preferably a pid or a code from a controlled vocabulary. following the identification of the 'related_item', the type of its identifier ('related_item_identifier_type') must be specified. this is important due to the varied syntax used by different pid systems. another critical field in this metadata schema is 'relation_type', designed to define the nature of the relationship between two variables labelled as a and b. this foundational metadata schema, which mirrors the 'relateditem' field from datacite (datacite metadata working group, 2021), consists of the following fields and subfields: ● related_item ○ related_item_identifier ■ related_item_identifier_type ○ relation_type the ddi variable cascade (ddi training group, 2021) categorizes comparability among variables into three layers: represented variable, conceptual variable, and instance variable. however, our current focus within ddi modelling22 is limited to the variables within the context of a dataset23, not extending to these abstraction24 layers. this limitation stems from the service requirement, as our metadata acquisition is solely for variables that are being registered for pids. the extended description within the 'basedonobjecttype' ddi is detailed in the subsequent section. enhancing description quality for knowledge graph relationships a knowledge graph (kg) is an advanced data model that encapsulates knowledge in a graph format, where entities and their interrelations are described in a way that machines can interpret. this model is constructed using a suite of established w3c standards, including the resource description framework (rdf), json for data interchange, the simple knowledge organization system (skos) for organizing knowledge, and the web ontology language (owl) for defining and categorizing web content. these standards are complemented by shared vocabularies and application programming interfaces (apis), which facilitate the integration of data from diverse domains and sources. a key feature of kgs is their consistent use of persistent identifiers (pids). these identifiers play a crucial role in ensuring that entities and their relationships are not only identifiable but also linkable across various data sources and applications. this capability is fundamental for integrating heterogeneous data sources into a cohesive and interconnected knowledge base. it supports semantic search and question answering, allowing users to query and retrieve knowledge using natural language queries. beyond using kg in the high-tech industry, such as google25, microsoft,26 and amazon27, kgs are also widely applied in scientific domains. large projects such as the open academic graph28 used for research for scholarly publications, the linked open data cloud29, interlinked datasets from various domains, and the global biodiversity information facility (gbif)30, a kg of biodiversity data, are some examples. in the social sciences, kgs have become instrumental in various research areas, including understanding societal and political debates, investigating fake news and misinformation (gangopadhyay et al., 2023), and mining knowledge about opinions and interactions from x data, formerly known as twitter (fafalios et al., 2018). the gesis research graph project31 also features a prototype graph that interlinks publications, research data, projects, and people. these initiatives are part of the broader effort to build a kg infrastructure that links social science research data and resources across gesis32. https://doi.org/10.29173/iq1118 https://ddi4.readthedocs.io/en/latest/userguides/variablecascade.html#example https://developers.google.com/knowledge-graph https://www.microsoft.com/en-us/research/group/cognitive-services-research/knowledge-and-language https://aws.amazon.com/neptune/knowledge-graphs-on-aws https://www.microsoft.com/en-us/research/project/open-academic-graph/overview https://www.microsoft.com/en-us/research/project/open-academic-graph/overview https://lod-cloud.net/ https://www.gbif.org/ 8/16 saldanha bach, janete; klas, claus-peter (2024) enhancing fair compliance: a controlled vocabulary for mapping social sciences survey variables, iassist quarterly 48(2), pp. 1-15. doi: https://doi.org/10.29173/iq1118 the core strength of kgs lies in their ability to connect, manage, and elucidate complex relationships, a feature that extends to variables in social science research. kgs are adept at capturing and representing the multidirectional connections between entities. by depicting variables as nodes and their relationships as edges, researchers gain a clearer understanding of how variables are associated and interact, facilitating the extraction of insights and predictions from data. adopting recognized standards and apis in kg construction ensures that the information is interoperable and machine-interpretable. this approach fosters the development of intelligent services and applications capable of automating the analysis and processing of variable data. thus, kgs emerge as vital tools in managing, analyzing, and sharing intricate information in social science research. in this sense, kg design benefits from extended and detailed variable descriptions. accurate descriptions improve data quality, leading to more reliable connections. transparent documentation of variables allows other researchers to understand the scope of the relation. detailed descriptions facilitate the integration of data from multiple sources and studies, enhancing data discoverability and making it easier for researchers to identify and use the data they need. datasets in repositories often consist of various files and sub-collections (wehrle and rechert, 2019; bugaje and chowdhury, 2017), making fine-grained levels, such as individual variables, crucial for research data management. these detailed connections enhance data reuse by enriching the decision-making process for researchers who often select specific variables rather than entire datasets. a primary need for social scientists reusing data is to swiftly comprehend the meanings and values of variables within a dataset (sun and khoo, 2018). however, challenges arise from terminology polysemy, where similar variable concepts may have different names or variables with the same name may represent different concepts. this issue necessitates intellectual effort to understand variables across studies and waves. an extensive research university library's experience with data reuse, including matters of replicability and reproducibility, underscores the importance of providing descriptive information about variables' coverage across datasets and specifying variable data definitions (scoulas, 2020). for example, table 1 illustrates the relationship between variables a and b across different waves, where variable b in wave 2 is basedon the variable a from wave 1, although with a different name. https://doi.org/10.29173/iq1118 9/16 saldanha bach, janete; klas, claus-peter (2024) enhancing fair compliance: a controlled vocabulary for mapping social sciences survey variables, iassist quarterly 48(2), pp. 1-15. doi: https://doi.org/10.29173/iq1118 table 1 variables relations: differences across waves: variable name variables relations variables study program wave variable name variable label isbasedon.hasdifferentvarname b -> a is equal is different is different is equal variable 1 var_1 study#100 wave1 age age in years variable 2 var_2 study#100 wave2 age_group age in years note: isbasedon (ddi-lc) figure 1 depicts the variables' relations across waves regarding the different variable names. variable b is based on variable a because it was generated later in a recent wave. although variable b is based on a, b has a different variable name (age_group) than the original variable a (age). figure 1: representation of the relations in a knowledge graph: different variable name label: 1 = equal; 2 = different. the dotted lines represent relations between entities. figure 2 depicts one example of variables a and b relation where a variable b is basedon a variable a but has a different question wording. in this case, the word ‘daß’ (the german word which means ‘that’) uses the character ß (called eszett ). the eszett letter is used only in german and can be typographically replaced with the double-s digraph ‘ss.’ in a more recent wave, the same word is replaced by the form ‘dass,’ adopting double-s. figure 2 depicts the kg representation of variables' relations across waves regarding question-wording. https://doi.org/10.29173/iq1118 10/16 saldanha bach, janete; klas, claus-peter (2024) enhancing fair compliance: a controlled vocabulary for mapping social sciences survey variables, iassist quarterly 48(2), pp. 1-15. doi: https://doi.org/10.29173/iq1118 figure 2: representation of the relations in a knowledge graph: different question-wording label: 1 = equal; 2 = different. the dotted lines represent relations between entities. figure 3 depicts one example of variables a and b relation where a variable b is basedon a variable a but has a different response schema. likert scale from variables a and b differs from 5 to 7 in each wave, respectively. figure 3 depicts the variables' relations across waves regarding different response schema. figure 3: representation of the relations in a knowledge graph: different response schema label: s = study program | w = wave | var = variable | q = question | rs = response schema | d = dataset the ease of discovering and visualizing dataset variable relations through kgs significantly enhances their comparability across different study waves. this functionality aids in understanding how variables interact within and between diverse datasets of several types. in longitudinal studies, tracking changes in variables across different waves becomes more manageable, facilitating the often costly and time-consuming harmonization process among datasets. kgs, with their search and browse functionalities, also augment data discoverability and findability (wu et al., 2019). they enable (inter)disciplinary data reuse by visually depicting how variables are distributed across multi-wave studies and identifying which variables have been consistently used over time. controlled vocabulary with extended descriptions of relations simplifies finding connections between variables within the same study or across different datasets. we provide concise textual identifications for each relation_type, supplemented by a cv and thorough explanations of these relationships. this https://doi.org/10.29173/iq1118 11/16 saldanha bach, janete; klas, claus-peter (2024) enhancing fair compliance: a controlled vocabulary for mapping social sciences survey variables, iassist quarterly 48(2), pp. 1-15. doi: https://doi.org/10.29173/iq1118 approach extends beyond merely naming and labelling variables. it also facilitates the discovery of relationships inherited within the ddi structure and other potential entities, such as data papers and additional resources (refer to section 4). employing these proposed relationships and the resulting controlled vocabulary leads to the creation of a semantically rich, common framework for social science research. these connections can be effectively represented in a kg across various institutions, in line with the fair (findable, accessible, interoperable, and reusable) principles for variables. this method enhances the understanding of variable relations and promotes the efficient and informed use of research data in the social sciences. controlled vocabulary (cv) for variable relations in the social sciences the cessda (consortium of european social science data archives) vocabulary manager plays a pivotal role in documenting and clarifying the relationships within social science research data. it provides extended descriptions and controlled vocabulary terms that describe links across various waves and studies in conjunction with questions and other related entities. a survey within this framework can encompass multiple waves, and each wave may include multiple surveys. each survey comprises numerous questions that can relate to one or more variables. these variables are defined by a variable name (or a variable id, typically a code), a variable label (which describes the variable), a response schema, and potentially, terms from a controlled vocabulary. our focus is understanding the relationship of a variable from its standpoint, specifically how it can be related to these different attributes. the cv adheres to the ddi property 'basedonreference_versionable.' this property allows for references to any number of objects that form the basis of the variable, a 'basedonrationaldescription' detailing how the content of the referenced object was incorporated or altered, and a 'basedonrationalcode' for specific typing of the 'basedonreference' in accordance with an external controlled vocabulary. this cv33 is published at the cessda cv manager. our initial contribution to this domain includes six relation types, each thoroughly detailed within the cv, listed below: 1. isbasedon.hasdifferentwavevariable 2. isbasedon.hasdifferentsurveyvariable 3. isbasedon.hasdifferentvarnamevariable 4. isbasedon.hasdifferentvarlabelvariable 5. isbasedon.hasdifferentquestionvariable 6. isbasedon.hasdifferentresponseschema table 2 provides examples of changes in the variable’s attributes to summarize the relations. table 2: variables relations extended descriptions * based on the ddi term isbasedon and the controlled vocabulary for variables relations for social sciences research data. label: s (study); w (wave); sy (survey); q (question); vn (variable name); vl (variable label); rs (response schema). 1 = equal; 2 = different related_item proposed relation_type examples vn = variable name vl = variable label q = question rs = response schema sy = survey w = wave s = study variable in waves isbasedon.hasdifferentwave study#100-wave1-variable:v5.-> study#100wave2-variable:v5 = = = = = ≠ = variable in surveys isbasedon.hasdifferentsurvey study#100-wave1-surveya-variable:v5.-> study#100-wave1-surveyb-variable:v5 ≠ = = variable name isbasedon.hasdifferentvarname study#100-wave1-variable:v5-"job_sat".-> study#100-wave2-variable:v7-"job_sat". ≠ ≠ = variable label isbasedon.hasdifferentvarlabel study#100-wave1-variable:v5-"job_sat".-> study#100-wave2-variable:v5-"work_sat" ≠ ≠ = variable question wording isbasedon.hasdifferentquestion study#100-wave1-questionabc-variable:v5. -> study#100-wave2-questionxyz-variable:v7. ≠ ≠ = variable response schema (response values) isbasedon.differentresponsesc hema.istypelikertscale study#100-wave1-qabc-variable:v7likert4points -> study#100-wave2-qabcvariable:v7-likert5points . ≠ ≠ = = equal ≠ different label https://doi.org/10.29173/iq1118 12/16 saldanha bach, janete; klas, claus-peter (2024) enhancing fair compliance: a controlled vocabulary for mapping social sciences survey variables, iassist quarterly 48(2), pp. 1-15. doi: https://doi.org/10.29173/iq1118 the following section addresses relations between variables and other entities within different studies. relations inherited within the ddi framing to explain the multifaceted interactions of variables with different entities, we have pinpointed specific types of connections that form a network of elements. this exploration helps identify which elements correlate most effectively with variables. for instance, a variable is inherently linked to a survey question. we aim to demonstrate how a variable can be connected to various entities beyond its original study context. take, for example, a hypothetical variable named ‘job_sat’. we can visualize its relationships with various entities through the following scenarios: a) research or data papers: academic papers may cite or feature a variable. for instance, a paper exploring job satisfaction might reference ‘job_sat’ as a critical factor in understanding employee well-being; b) landing page: websites can offer detailed metadata about a variable. an example is the webpage ‘www.example.com/job-satisfaction’, which could provide comprehensive metadata and descriptions of ‘job_sat’ c) interactive resources: scripts or codes often utilize variables for data analysis. for example, a python script could use ‘job_sat’ to process survey data, create visual representations, or analyse trends in job satisfaction d) data management plan: such plans might include anticipated use of variables. a workplace wellness study’s plan could specify using ‘job_sat’ for data collection and analysis; e) audio/video data: variables can be incorporated into multimedia formats. for instance, a video presentation on study outcomes might include discussions and visualizations of ‘job_sat’, highlighting its impact on employee happiness. by expanding the kg to encompass these diverse entities, using controlled list values from resource type descriptions, we maintain the kg's interoperability across different domains. linking entities to their associated variables provides a comprehensive overview of their interdependent connections. this approach significantly enhances the data’s findability, accessibility, interoperability, and potential for future reuse. conclusion standard metadata fields are indispensable for effectively registering relationships between variables within studies. these standards are crucial for ensuring interoperability among different systems and enabling automation features. accurately representing possible relation_types and formally documenting them significantly enhances meta-searching and meta-browsing capabilities. this makes it easier to find and access relevant data. the essential requirements are uniquely identifying each variable with a persistent identifier (pid) and clearly defining its relationships using controlled vocabulary terms. such approaches are instrumental in fostering machine-actionable data features, thereby strengthening data's findability and enhancing data's reusability at the variable level. data users can benefit from the ability to correspond variables, exploring their consistency or comparability over time across different waves and studies. with machine-readable and actionable features, complex recommendation systems can be developed. these systems can display relationships between variables and other entries in relationship maps, such as those represented in kgs. while the pid registration service's primary function is not to provide kg visualization, its https://doi.org/10.29173/iq1118 13/16 saldanha bach, janete; klas, claus-peter (2024) enhancing fair compliance: a controlled vocabulary for mapping social sciences survey variables, iassist quarterly 48(2), pp. 1-15. doi: https://doi.org/10.29173/iq1118 inclusion of the 'related_item' field and corresponding subfields in its metadata schema lays the groundwork for documenting variables in a way that enhances kg applications. we propose an extended description for the 'relation_type' description and a controlled vocabulary terminology based on the ddi term 'isbasedon'. this approach enables researchers and other interested parties to quickly locate the most relevant and usable variables for their research needs. for data holders, this method facilitates the maximization of value-added services through the increasing interconnection of research output entities. variables are not only linked to their inherent elements like questions, questionnaires, survey waves, and response scales. still, they can also be input for interactive resources such as scripts or do-files. there is also potential for registering and assigning pids to questions and response schemas from existing surveys for reuse purposes. documenting and defining all these relations accurately, with detailed relation descriptions, will enhance the controlled vocabulary for the social sciences. this, in turn, will foster the reuse of the cessda controlled vocabulary tool among institutions, leveraging these interconnected relationships for broader research and analysis purposes. references ajzen, i. and fishbein, m. (2005), ‘the influence of attitudes on behavior’, in albarracin, d., johnson, b. t. and zanna, m.p. (eds), handbook of attitudes and attitude change, lawrence erlbaum associates, mahwah, nj. aryani, a. et al. (2018) ‘a research graph dataset for connecting research data repositories using rd-switchboard’, scientific data, 5(1), p. 180099. available at: https://doi.org/10.1038/sdata.2018.99. babbie, e.r. (1990). survey research methods, wadsworth publishing, belmont, ca. bollen, k.a. (2002) ‘latent variables in psychology and the social sciences’, annual review of psychology, 53(1), pp. 605–634. available at: https://doi.org/10.1146/annurev.psych.53.100901.135239 bugaje, m. and chowdhury, g. (2017) ‘is data retrieval different from text retrieval? an exploratory study’, in s. choemprayong, f. crestani, and s.j. cunningham (eds) digital libraries: data, information, and knowledge for digital lives. cham: springer international publishing (lecture notes in computer science), pp. 97–103. available at: https://doi.org/10.1007/978-3-319-70232-2_8 cox, m. (2015) ‘a basic guide for empirical environmental social science’, ecology and society, 20(1), p. art63. available at: https://doi.org/10.5751/es-07400-200163 datacite metadata working group (2021) ‘datacite metadata schema documentation for the publication and citation of research data and other research outputs v4.4’, p. 82 pages. available at: https://doi.org/10.14454/3w3z-sa82. ddi training group (2021) ‘variables and the variable cascade’. available at: https://doi.org/10.5281/zenodo.5180568 . https://doi.org/10.29173/iq1118 https://doi.org/10.1038/sdata.2018.99 https://doi.org/10.1146/annurev.psych.53.100901.135239 https://doi.org/10.1007/978-3-319-70232-2_8 https://doi.org/10.5751/es-07400-200163 https://doi.org/10.14454/3w3z-sa82 https://doi.org/10.5281/zenodo.5180568 14/16 saldanha bach, janete; klas, claus-peter (2024) enhancing fair compliance: a controlled vocabulary for mapping social sciences survey variables, iassist quarterly 48(2), pp. 1-15. doi: https://doi.org/10.29173/iq1118 european union. (2016). ‘regulation (eu) 2016/679 of the european parliament and of the council of 27 april 2016 on the protection of natural persons about the processing of personal data and the free movement of such data, and repealing directive 95/46/ec (general data protection regulation)’. official journal of the european union. available at: https://eurlex.europa.eu/eli/reg/2016/679/oj fafalios, p.; iosifidis, v.; ntoutsi, e. and dietze, s. tweetskb: a public and large-scale rdf corpus of annotated tweets. in 15th extended semantic web conference (eswc'18), heraklion, crete, greece, june 3-7, 2018. https://doi.org/10.48550/arxiv.1810.10308 gangopadhyay, s., boland, k., dessí, d., dietze, s., fafalios, p., tchechmedjiev, a., ... & jabeen, h. (2023, may). truth or dare: investigating claims truthfulness with claimskg. in second international workshop on linked data-driven resilience research (d2r2’23) co-located with eswc 2023, may 28th, 2023, hersonissos, greece. available at: https://ceurws.org/vol-3401/paper7.pdf issp research group (1992) ‘international social survey programme: role of government ii issp 1990 international social survey programme: role of government ii issp 1990’. gesis data archive. available at: https://doi.org/10.4232/1.1950 jaaskelainen, t., moschner, m. and wackerow, j. (2010) ‘controlled vocabularies for ddi 3: enhancing machine-actionability’, iassist quarterly, 33(1), p. 34. available at: https://doi.org/10.29173/iq649 kaur, loveleen and mittal, ritu. (2021). ‘variables in social science research’. indian res. j. ext. edu. 21 (2&3), april & july, 2021. url: https://www.researchgate.net/profile/ritu-mittal2/publication/351080413_variables_in_social_science_research/links/6083aa49907dcf66 7bbda5cf/variables-in-social-science-research.pd f klas, c.-p. et al. (2022) konsortswd measure 5.1: pid service for variables report. zenodo. available at: https://doi.org/10.5281/zenodo.6397367. manghi, p. et al. (2019) the openaire research graph data model. zenodo. available at: https://doi.org/10.5281/zenodo.2643199. liebig, s. et al. (2021) ‘socio-economic panel, data from 1984-2019, (soep-core, v36, eu edition) sozio-oekonomisches panel, daten der jahre 1984-2019 (soep-core, v36, eu edition)’. soep socio-economic panel study. available at: https://doi.org/10.5684/soep.core.v36eu. roßbach, h.-g. and neps, national educational panel study, bamberg (germany) (2016) ‘neps starting cohort 6: adults (sc6 6.0.1)neps-startkohorte 6: erwachsene (sc6 6.0.1)’. neps national education panel study. available at: https://doi.org/10.5157/neps:sc6:6.0.1 saldanha bach, j., klas, c.-p. and mutschke, p. (2023) konsortswd measure 5.1: use cases description extended report. zenodo. available at: https://doi.org/10.5281/zenodo.7588944 https://doi.org/10.29173/iq1118 https://eur-lex.europa.eu/eli/reg/2016/679/oj https://eur-lex.europa.eu/eli/reg/2016/679/oj https://doi.org/10.48550/arxiv.1810.10308 https://ceur-ws.org/vol-3401/paper7.pdf https://ceur-ws.org/vol-3401/paper7.pdf https://doi.org/10.4232/1.1950 https://doi.org/10.29173/iq649 https://www.researchgate.net/profile/ritu-mittal-2/publication/351080413_variables_in_social_science_research/links/6083aa49907dcf667bbda5cf/variables-in-social-science-research.pd%20f https://www.researchgate.net/profile/ritu-mittal-2/publication/351080413_variables_in_social_science_research/links/6083aa49907dcf667bbda5cf/variables-in-social-science-research.pd%20f https://www.researchgate.net/profile/ritu-mittal-2/publication/351080413_variables_in_social_science_research/links/6083aa49907dcf667bbda5cf/variables-in-social-science-research.pd%20f https://doi.org/10.5281/zenodo.6397367 https://doi.org/10.5281/zenodo.2643199 https://doi.org/10.5684/soep.core.v36eu https://doi.org/10.5157/neps:sc6:6.0.1 https://doi.org/10.5281/zenodo.7588944 15/16 saldanha bach, janete; klas, claus-peter (2024) enhancing fair compliance: a controlled vocabulary for mapping social sciences survey variables, iassist quarterly 48(2), pp. 1-15. doi: https://doi.org/10.29173/iq1118 saldanha bach, j., klas, c.-p. and mutschke, p. (2023) konsortswd measure 5.1: metadata schema extended report. zenodo. available at: https://doi.org/10.5281/zenodo.7588902 scoulas, j.m. (2020) ‘learning from data reuse: successful and failed experiences in a large public research university library’, iassist quarterly, 44(1–2), pp. 1–15. available at: https://doi.org/10.29173/iq966 stocker, m. et al. (2018) ‘curating scientific information in knowledge infrastructures’, data science journal, 17, p. 21. available at: https://doi.org/10.5334/dsj-2018-021. sun, g. and khoo, c.s.g. (2018) ‘a framework to represent variables and values in social science research data sets to support data curation and reuse’, in f. ribeiro and m.e. cerveira (eds) challenges and opportunities for knowledge organization in the digital age. ergon verlag, pp. 231–239. available at: https://doi.org/10.5771/9783956504211-231 thomas, w., et al. (2014). data documentation initiative: technical specification part i version 3.2. url: https://ddialliance.org/specification/ddilifecycle/3.2/xmlschema/highleveldocumentation/ddi_part_i_technicaldocument.pdf wehrle, d. and rechert, k. (2019) ‘are research datasets fair in the long run?’, international journal of digital curation, 13(1), pp. 294–305. available at: https://doi.org/10.2218/ijdc.v13i1.659 wu, m. et al. (2019) ‘data discovery paradigms: user requirements and recommendations for data repositories’, data science journal, 18, p. 3. available at: https://doi.org/10.5334/dsj-2019-003 endnotes 1 dr. janete saldanha bach is a postdoc researcher at gesis – leibniz institute for the social sciences, unter sachsenhausen 6-8, cologne, germany, and can be reached by email: janete.saldanhabach@gesis.org. https://orcid.org/0000-0001-9011-5837. 2 dr. claus-peter klas is the team leader of data & service engineering at gesis – leibniz institute for the social sciences, unter sachsenhausen 6-8, cologne, germany. 3 konsortswd (consortium for the social, behavioural, educational and economic sciences) is funded by the national research data infrastructure (nfdi) https://www.konsortswd.de/ 4 german national research data infrastructure (nfdi) homepage: https://www.nfdi.de/ 5 project community at zenodo https://zenodo.org/communities/konsortswd-ta5-m1 6 waves are different points in time when data is collected in a research study. waves are typically associated with longitudinal studies, which involve the repeated observation of the same subjects over time. 7 https://vocabularies.cessda.eu/vocabulary/commonalitytype?lang=en 8 fair stands for findable, accessible, interoperable and reusable. it refers to the fair data principles developed by the force 11 community, that recommend data should be shared according to these four concepts. 9 the data documentation initiative (ddi) is an effort to create an international standard for describing data from the social, behavioral, and economic sciences. expressed in xml, the ddi metadata specification now supports the entire research data life cycle. 10 https://www.ddialliance.org/about/about-the-alliance https://doi.org/10.29173/iq1118 https://doi.org/10.5281/zenodo.7588902 https://doi.org/10.29173/iq966 https://doi.org/10.5334/dsj-2018-021 https://doi.org/10.5771/9783956504211-231 https://ddialliance.org/specification/ddi-lifecycle/3.2/xmlschema/highleveldocumentation/ddi_part_i_technicaldocument.pdf https://ddialliance.org/specification/ddi-lifecycle/3.2/xmlschema/highleveldocumentation/ddi_part_i_technicaldocument.pdf https://doi.org/10.2218/ijdc.v13i1.659 https://doi.org/10.5334/dsj-2019-003 mailto:janete.saldanhabach@gesis.org https://orcid.org/0000-0001-9011-5837 16/16 saldanha bach, janete; klas, claus-peter (2024) enhancing fair compliance: a controlled vocabulary for mapping social sciences survey variables, iassist quarterly 48(2), pp. 1-15. doi: https://doi.org/10.29173/iq1118 11 basedonobjecttype available at https://ddialliance.github.io/ddimodel-web/ddi-l3.3/composite-types/basedonobjecttype/ 12 https://schema.org/product 13 https://ddialliance.github.io/ddimodel-web/ddi-l-3.3/composite-types/basedonobjecttype/ 14 https://ess-search.nsd.no/cdw/conceptvariables 15 in a data file, a single vertical column, each being one byte in length. fixed-format data files are traditionally described as being arranged in lines and columns. in a fixed format file, column locations describe the locations of variables. (url: https://www.icpsr.umich.edu/web/icpsr/cms/2042) 16 in general, a ‘line’ in data file terminology refers to a physical unit of data that the computer reads and processes, one at a time. (url: https://www.icpsr.umich.edu/web/icpsr/cms/2042). 17 url: https://www.diw.de/soep 18 url: https://paneldata.org/soep-core/ 19 url: https://doi.org/10.5684/soep.core.v37i 20 https://sdmx.org/ 21 https://www.dublincore.org/ 22 https://ddi4.readthedocs.io/en/latest/userguides/variablecascade.html#example 23 this layer is so-called instance variables in the ddi-cdi model and variables in the ddi-lc model 24 abstraction layers are the represented variable and conceptual variable 25 https://developers.google.com/knowledge-graph 26 https://www.microsoft.com/en-us/research/group/cognitive-services-research/knowledge-andlanguage/ 27 https://aws.amazon.com/neptune/knowledge-graphs-on-aws/ 28 https://www.microsoft.com/en-us/research/project/open-academic-graph/overview/ 29 https://lod-cloud.net/ 30 https://www.gbif.org/ 31 https://researchgraph.org/gesis-research-graph/ 32 https://www.gesis.org/en/research/applied-computer-science/knowledge-graph-infrastructure 33 https://vocabularies.cessda.eu/vocabulary/variables-relations?lang=en https://doi.org/10.29173/iq1118 https://www.diw.de/soep https://doi.org/10.5684/soep.core.v37i https://developers.google.com/knowledge-graph https://www.microsoft.com/en-us/research/group/cognitive-services-research/knowledge-and-language/ https://www.microsoft.com/en-us/research/group/cognitive-services-research/knowledge-and-language/ https://lod-cloud.net/ vol282-3.indd iassist quarterly summer/fall 2004 35 by kristi thompson1 and daniel m. edelstein a reference model for providing statistical consulting services in an academic library setting introduction princeton university library, through its data and statistical services (dss) unit, goes further than many libraries in providing consulting on statistical methods as well as software support for data library users. princeton requires all third and fourth year undergraduates to do independent original research papers and theses in their disciplines of concentration. these requirements create a clientele of students who need to conduct relatively sophisticated statistical analysis, but who may or may not possess the necessary skills required to do so. this paper describes the model we employ at princeton dss to help these patrons, drawing on our experience as data consultants to discuss how it works in practice.2 a call to integrate statistical literacy into data library service data librarians go to great lengths to make data fi les available to their patrons and to help their patrons locate data sources. but helping people fi nd the data fi le that best suits their needs is not enough. providing access is more than making an item available. it means making a resource useful. data fi les have a layer of complexity that can make them more challenging to use than other information sources. while most college students can extract and use the information contained in a book, many do not have the statistical or technical skills required to effectively extract and use the information contained in a data fi le. in other words, a high proportion of our patrons are not statistically literate. giving a data fi le to a patron who does not possess the tools and skills needed to analyze it is about as useful as giving a book to someone who cannot read. much discussion of the problem of statistical illiteracy among students revolves around the need to properly integrate statistical literacy into the academic curriculum. schield, for example, claims that “statistical educators should develop a college-level statistical literacy course for students in majors that do not require a math or statistics course.” (schield 2004)3 this would be a welcome achievement. while we are waiting, the data library has a niche to fi ll in helping make it possible for students to use statistical analysis as a research methodology. and we believe that the library has a powerful model for dealing with the immediate needs of students who need to analyze data, that is, the reference model. an introductory information literacy course, no matter how well taught, is not a substitute for the direct assistance of a qualifi ed and experienced reference librarian. teaching patrons about information is important, but the essential service of a librarian is to help meet the need for information at the point when it occurs. similarly, we believe that while teaching the concepts of statistical literacy is critical, the primary purpose of the library’s consulting facility is to help our patrons use the data resources made available by the library. teaching statistical theory is the role of the faculty. our role is to complement them by helping our patrons overcome the practical knowledge barriers that arise when conducting a data analysis project. the dss consultants provide assistance with the software and the statistical techniques necessary to make use of data fi les. our service also makes it easier for faculty to integrate data analysis into their curricula. in her article “understanding barriers to the use of numeric data in learning and teaching,” rice concluded, “universities should develop it strategies that include data services and support for staff and students, and integration of empirical datasets into learning technologies.” (rice 2001)4 we support this conclusion, but believe that expanded library service and support for statistical theory is as important as it service and integrated learning technology. putting it into practice the service model we have evolved is in many ways similar to traditional library reference service, but we have adapted it to meet the unique challenges of statistical consulting. our primary role is to help students use the data resources made available by the library. much like in a traditional academic library reference transaction, we are trying to help our patrons fi nd an answer to a particular research question. our aim is not to teach statistics in itself, but to provide users with the practical knowledge needed to carry out their research. as a result, we have evolved a very practical, problem-oriented and intuitive approach. teaching statistical theory is the professors’ role. as one student remarked, the dss consultants helped explain things “in a way in which even (his) statistically36 iassist quarterly summer/fall 2004 challenged mind could understand” (data and statistical services client database, 2003).5 the consulting interview each consulting transaction begins with an informal reference interview in which we try to gently extract the information we need from our students before we can start working with them: • what is the research question? • what level of knowledge does the student have? • where are the data and what form are they in? • what are the conventions and standards of the student’s area of study? what is the research question? finding out the actual research question – not just the question the student thinks that he needs help with – is important. sometimes a student will come in and say that she wants to perform a particular type of analysis using some data set. occasionally questioning about what she really wants to fi nd out will reveal that the approach she wants to take is incorrect or not suitable for her data. sometimes a student will have picked up a statistical term from classmates or will vaguely remember something from a lecture that he thinks is what he needs to do. others come in with only a very vague idea of what they want to fi nd out, or with a question that needs to be reformulated into something that can actually be answered with the data available. and occasionally, we have oddities such as the person who wanted to explain the gender of a judge as an outcome of what law school he or she had gone to. as one student noted, we help with “logic” as well as with “seeing the limitations of (her) study” (data and statistical services client database, 2004). what level of knowledge does the student have? the level of knowledge of the student is sometimes readily apparent, particularly in those cases when the answer is “none.” on other occasions we need to ask some gentle probing questions: are you familiar with this type of analysis? have you worked with data fi les before? before we can proceed, we need to know how closely the student’s level of knowledge matches the level at which the analysis needs to be done. it is necessary to establish this at the start so that we do not inadvertently confuse or discourage students by giving them explanations that they do not have the background to understand. where is the data and what form is it in? usually, students come in with the data or knowing how to get it easily. occasionally our questioning will reveal that a student is using the wrong dataset, or needs to merge in additional data. sometimes students have come in wanting to perform a regression on a dataset that was only available as a set of summary statistics. very frequently the data needs to be extracted, reshaped, combined with other data fi les, converted to another format, or recoded before anything else can be done. what are the conventions and standards of the student’s area of study? the discipline that the student is working in will often affect what statistical advice we will give her. for example, we have found that biology students are often required to do nonparametric tests in situations where social scientists would not be. the area of study also can affect which results are reported and what language is used. in addition, the level that the analysis is done at, how rigorous or sophisticated it needs to be, can depend on the student’s discipline, the year she is in, the scale of her project, and the standards of the department. to demonstrate more clearly how the process works in practice, consider the following example of a typical encounter. a student comes in and asks the consultant to show him how to “get means in spss.” rather than immediately providing the answer, the consultant fi rst asks the student about his dataset and research problem, and realizes that in fact he wants to do a t-test for the difference between means of some dependent variable grouped by some subgroup variable, such as gender. and given the type of research project the student is working on and the departmental standards for that type of work, the student needs to control for several other variables, so the consultant explains these issues and helps him to run and interpret multiple regression instead. different approaches for different students our patrons exhibit a wide range of both technical and statistical ability and experience, and we need to take this into account when deciding how to approach each individual problem. we have informally grouped our students into three basic types to help us explain the range of different approaches we need to use. the least sophisticated group consists of students with absolutely no knowledge of data or statistics whatsoever, who have somehow found themselves needing to conduct an analysis. sometimes we encounter students who have collected or come across some data that they want to use, but do not know what to do with it. one example was a philosophy student we worked with last year who had conducted a survey of other students to get their reactions to various moral dilemmas. with students like this we often need to start by getting their data into a useable computer form. then we work with them to fi nd out what they want to know, and teach them in an intuitive way the statistical procedures they need to fi nd that out. these students often have nowhere else to turn because statistics iassist quarterly summer/fall 2004 37 simply is not taught in their department. this group also includes students (and, occasionally, faculty members) who encountered a question in their research, or came across an enticing entry in the library catalog, and came to the library hoping to get a book or nicely formatted table they could read the information from. instead, they were given a data fi le and told that the information was in there somewhere, if they could only decode it. we often start our discussion with these patrons by saying something like “a data fi le is sort of like an excel spreadsheet…” the second group contains students who have some background in the mathematical and theoretical basis of statistics, but have trouble applying it. university statistics courses are often taught in a theoretical and abstract way that leaves students under-equipped to deal with the practicalities of analyzing real-life data critically. some may just need help with the software, while others may understand the mathematics of a regression but have no sense of what variables need to be included, or how to code categorical variables. and many seem to simply have trouble understanding how the theoretical concepts they have memorized can be used to make sense of actual data. an exchange that occurred while helping a student with his third year undergraduate research captures this well. consultant: “your explanatory variable is signifi cant.” student: “great!” pause. “but what does it mean?” the most capable group includes students with a solid level of statistical knowledge who are trying to do something ambitious that requires specialized programming skills or otherwise advanced knowledge. we also occasionally encounter students who need to fi nd a valid statistical test that will let them fi nd out something unusual, or need to learn the correct model to use in some peculiar circumstance. answering this last type of question involves pure library research skills. we keep assorted statistical reference material on hand and have favorite trusted web sites to look at, and we also maintain contacts with graduate students and faculty who can help us with particularly diffi cult questions. with all of these students our work often resembles data counseling as much as data consulting. we work with our patrons to explore the data and their research, and give them the opportunity to discuss what they are doing with someone knowledgeable who is not involved in evaluating them. we ask questions to help them clarify in their own minds what they are doing, and encourage them to explore possibilities. our focus is on helping students to do the best work they are capable of themselves, and simply listening to them is often the most effective way to accomplish this. we want to make sure that students understand what they are doing – not the details of the statistical theory behind it so much as an intuitive understanding of the concept behind the actions they are performing and the results they are getting. for example, if a student needs to do a probit analysis, we will not go into detail about the distribution and how it was derived, but we will give a couple of examples to explain why linear regression breaks down in the case of a binary dependent variable and may do a sketch of the probit curve to help her grasp why it works better. patron management: outreach and intake almost all of our consulting is done on a walk-in basis in our computer lab during open consulting hours. we prefer not to make appointments, as we have found that scheduling an extended block of time to work with a single patron is generally an ineffi cient and unfair way of dividing our time. working together in our computer lab, we can serve eight or more students at a time, answering questions as they arise while encouraging our patrons to work independently as much as they are able. naturally, some students do require larger blocks of time, particularly when they are beginning a project. we hire graduate students as assistants and train them to deal with routine questions so that during these busy periods we can focus on the students who need more involved assistance. close to half of our students come from the department of economics, with politics, sociology, public policy and psychology supplying most of the remainder. however, over the last year or so we have also assisted patrons from history, computer science, geology, philosophy, english, bioethics, engineering, religion and evolutionary biology. our outreach efforts include giving presentations to groups of majors in the social science departments, either by meeting with the students in groups or by attending sessions of departmental workshops, together with either a subject librarian or the data librarian. we also meet with new graduate students and encourage them to send us their students as well as use our services themselves. many of our patrons are also referred to us by the data, economics, and other social science librarians. as word of our service spreads around campus, we also get many customers through word of mouth. our web site has also become an increasingly important promotional tool. working with faculty most of our consultations are with undergraduates doing independent research projects under the guidance of a faculty member or graduate student who advises them. we frequently need to decide what is appropriate to teach or advise students in the area of statistical methodology, given that their advisor is also supposed to offer help in this area, and certainly will be evaluating the choices that the student makes. when it comes to methodology questions, we will give a defi nitive answer if one exists. however, frequently the right answer to a statistical question is a matter of 38 iassist quarterly summer/fall 2004 opinion. when a question of this type arises, we may offer suggestions, but we also discuss possible alternative options, and will strongly encourage students to seek help from a professor. we generally try to get a sense of how much support a student is getting from other sources. this can infl uence how much direct methodological advice we will give. in cases where an advisor is working closely with the student we will defer to or, where necessary, reinforce the professor’s advice. in cases where a student is getting less help, we will do our best to make up the lack. in a few cases we have found ourselves dealing with students whose advisors were giving them advice that was unambiguously wrong. often this is a result of miscommunication, and we are able to fi nd the source of the misunderstanding and resolve the problem. in other cases the problem is an actual lack of knowledge, as in the case of the philosophy student whose professor did not understand the concept of statistical signifi cance. advisors who have found themselves dealing with an area outside their realm of knowledge are often relieved when they learn that the data consulting facility is available to assist their students. if handled with care and tact, incidents such as this can improve our relationships with the various departments with whom that we fi nd ourselves working. conclusion once, we asked a student what the next stage of his research was and why he wanted to perform a particularly complex procedure. he said, “i don’t know what the point of this is. i’ll just fi nish this step and then my advisor will tell me what to do next.” this attitude, and the teaching style which fosters it, is antithetical to our approach. our goal is to make it possible for non-statistically literate students to both conduct and understand analyses using sophisticated statistical techniques. we make it possible for our students to do their work themselves, and encourage learning by doing. we teach students enough to get them started, then have them dive into actual statistical analysis as quickly as possible. most of them start to catch on quickly, and then proceed with their work with increasing confi dence. from that point, we act as a resource they can consult at the inevitable bumps in the road. frequently there are quick questions, and sometimes they need to pull over for more extended consultation. the fl exibility of this approach allows our students to quickly take control of their projects, through acquiring and using the appropriate level of statistical knowledge. often they are encouraged to learn more statistics and move on to more ambitious projects. but, even if not, they leave having accomplished something worthwhile. notes 1 contact: kristi a. thompson, princeton university library. phone: +1 609-258-6053 http://dss.princeton.edu/ email: kristit@princeton.edu. 2 parts of this paper were presented at the iassist conference, 2004, madison, wisconsin by daniel m. edelstein and kristi thompson. 3 schield, milo (2004), “statistical literacy curriculum design.” 2004 iase roundtable, lund sweden. available at: http://www.augsburg.edu/ppages/~schield/milopapers/ 2004schieldiase.pdf 4 rice, robin (2001), “understanding barriers to the use of numeric data in learning and teaching.” iassist quarterly. vol. 25, no. 1. pg. 5-9. 5 data and statistical services client database (20012004) a database used to track usage of the data and statistical services computer lab, in which computer lab users enter details of their lab usage, including optional comments, princeton university. iassist quarterly 2014/2015 25 iassist quarterly use cases related to an ontology of the data documentation initiative by thomas bosch1 and brigitte mathiak2 abstract ontology engineers worked in close collaboration with experts from the statistical domain in order to develop an ontology of a subset of the data documentation initiative. in this paper, we give a brief overview of the ddi ontology’s current status and discuss in detail the most significant use cases associated with the ddi data model’s ontology and therefore various benefits for the statistics community. by means of this ontology, ddi data as well as metadata can be published in the linked open data cloud and as a consequence be combined with an extensive number of datasets from diverse heterogeneous data sources. researchers will have the opportunity to discover both data and metadata related to multiple studies which are interlinked in the web of data. in case a user searches for a specific study and does not know which terms to state, it is necessary to link ddi concepts to external thesaurus concepts. as a result, users’ search tasks are facilitated in a significant manner. semantic web technologies enable the ability to check the consistency of the overall ddi data model and ease the comparison of ddi elements among multiple ddi instances. furthermore, external resources like publications related to specific data can be found and linked, if they are semantically specified. keywords: semantic web, linked data, data documentation initiative, ddi, use cases. introduction statistical domain experts worked closely with linked data community experts to define an ontology of the ddi data model. this work has begun at the workshop “semantic statistics for social, behavioural, and economic sciences: leveraging the ddi model for the linked data web” at schloss dagstuhl leibniz center for informatics, germany, in september 2011 (dagstuhl 2011) and continued at the follow-up workshop in the course of the 3rd annual european ddi users group meeting (eddi11) in gothenburg, sweden (european ddi user conference 2011). a final dagstuhl workshop on semantic statistics took place in october 2012 (dagstuhl 2012). figure 1 depicts the ddi ontology’s conceptual model containing the ddi elements that are seen by diverse experts of the statistical domain as the most important ones to solve problems connected with various use cases the authors of this paper identified. xml schemas, which describe the ddi data model, build the basis of the visualized ddi ontology’s conceptual model. extensions partly borrow from existing vocabularies and partly lead to a new ddi vocabulary. the most important parts of the data model are the three components of the ddi conceptual model “study”, “variable”, and “logicaldataset”. thus, they are highlighted and outgoing relations are displayed in three different colors (bosch et al. 2012). widely adopted and accepted ontologies are heavily reused as they can also address some ddi features. some of the reused vocabularies are: • dublin core ontology (dcmi metadata terms 2012) delineating metadata for citation purposes 26 iassist quarterly 2014/2015 iassist quarterly • simple knowledge organization system (skos) (w3c 2009) describing code lists, category schemes, mappings between them, and concepts like topics • rdf data cube vocabulary, which describes aggregated data like multi-dimensional tables (bosch et al. 2012) we defined a direct and a generic mapping between ddi-xml and ddi-rdf. both ddi-codebook and ddi-lifecycle xml documents can be transformed automatically to an rdf representation, as the syntactic structure is described using xml schemas. bosch et al. (2011) have developed a generic multi-level approach for designing domain ontologies based on xml schemas. xml schemas are converted to owl generated ontologies automatically using xslt transformations which are described in detail by bosch et al. (2012). after the transformation process, all the information located in the underlying xml schemas of a specific domain is also stored in the generated ontologies. owl domain ontologies can be inferred completely automatically out of the generated ontologies using swrl rules. in the following, the authors of this paper describe in detail use cases and benefits associated with an ontology of the ddi data model. we want to answer the question why it is crucially important that an rdf representation of the data documentation initiative has been defined. publish and link ddi data and metadata using an ontology of the ddi data model, ddi data as well as metadata can be published in the linked open data cloud in the form of the standard-based exchange format rdf. the lod cloud comprises approximately 29 billion rdf triples. the number of rdf links of nearly 400 million refers to out-going links that are set from data sources within a topical domain to data sources of other thematic areas (bizer, jentzsch & cyganiak 2012). figure 2 visualizes the current state of the entire web of data with its rdf triples, links between them, and diverse topical sections depicted using different colors. one has to fulfill different conditions before datasets can be published in the lod network. data must be published in accordance with the linked data principles. another precondition is offering rdf data through a sparql endpoint (w3c 2008). ddi instances can be processed by rdf tools without supporting the complex ddi xml schemas’ data structures and can be displayed using mature linked data browsers like tabular (the tabulator 2005), marbles (marbles 2012,) or linksailor (linksailor 2012). after publishing publicly available structured data, ddi data and metadata may be linked with other data sources of multiple topical domains. organizations offering rdf representations of their ddi instances will be additional nodes in this lod network. two major advantages are connected with the publication of ddi data and metadata in the lod cloud and with the relation to other rdf datasets. first, each organization which is part of this continuously growing cloud can search for, find, and operate with the published ddi instances of a specific organization. and figure 1. conceptual model of the ddi ontology (bosch et al. 2012) iassist quarterly 2014/2015 27 iassist quarterly secondly, every node in this lod network can also be processed by individual organizations. in summary, you can reach a broader audience and a broader audience can reach you. linked data search engines like sig.ma (sig.ma semantic information mashup 2012), falcons (falcons 2011), or swse (semantic web search engine 2012) can search for ddi instances which can be found in the directory of all known sources of linked data with an open license (linking open data project) (linkeddata 2012). linked data crawlers such as the publicly available ldspider use rdf links between various data sources to provide extensive search functionalities (isele et al. 1996). even semantic mashups utilize linked rdf data from several data sources. furthermore, the publication of linked data in the lod cloud is the prerequisite of the development of linked data driven web applications. discovery what kinds of problems can’t be solved without an ontology of the ddi data model, what types of problems can be solved in a better way using such an ontology, and what is the associated additional value? requesting multiple, distributed, and merged ddi instances will be possible. the semantic web query language sparql is applied to traverse the rdf graph (w3c 2008). the sparql protocol and query language is similar to sql, the structured query language, within the framework of requesting relational databases. but before executing sparql queries, you have to generate a sparql endpoint (w3c 2008). semantic queries are formulated using simple and intuitive ddi domain concepts without knowledge of complex ddi xml schemas’ structures. in the following program listing, all the questions belonging to a given variable with the variable label ”age” are requested. select ?question where { ?variable rdf:type variable; skos:preflabel ?variablelabel; hasquestion ?question. ?question rdf:type question. filter ( ?variablelabel = ‘age’ ) } the next figure visualizes the rdf representation of this specific sparql query. other examples would be querying all the studies in which variables with a specific variable label exist or to request all the publications belonging to a given topic. bosch et al. (2012) provide a detailed description of the discovery use case summarized in this sub-section. by means of the ddi ontology, researchers can discover both data and metadata belonging to more than one particular study. researchers often wish to know which studies are connected with a specific universe consisting of the three dimensions: time (e.g., 2005), country (e.g., france), and population (e.g., age between 18 and 65). figure 4 depicts the sparql query shown below its visualization. the sparql query’s results are the titles of the studies related to the defined universe. these individual studies are of the type ’study’ and are connected with the mentioned universe via the object property ‘ismeasureof’. this particular study is related to its title using the datatype property ’title’ borrowed from the dublin core namespace. the universe consisting of the three dimensions time, country, and population is defined as from the type ’universe’ figure 2. the linking open data cloud diagram (cyganiak 2011) 28 iassist quarterly 2014/2015 iassist quarterly and is combined with its definition via the datatype property ’definition’. the individual namespaces where the class axioms are specified are shown in the figure in the form of namespace prefixes such as ’ddi’, ’dc’, and ’skos’. select ?studytitle where { ?study rdf:type ddi:study; dc:title ?studytitle; ddi:ismeasureof ?universe. ?universe rdf:type ddi:universe; skos:definition ?universedefinition. filter ( ?universedefinition = “country =’france’ and time = ‘2005’ and population = ‘age: 18-65’” ) } the result of the sparql query is a table including all the titles of the studies which are associated with the given universe. the next step could be to request exactly those studies returned from the first query in which a particular concept (e.g., education) exist. in this case, variables associated with the three-dimensional universe and the returned studies are linked to the ddi element ‘concept’ via the object property ‘hasconcept’. the concept label is realized using the datatype property ‘preflabel’ borrowed from skos. the next figure delineates another frequent research discovery process. researchers want to know which questions -such as ’what is your highest school degree?’ -are linked to specific concepts like ’education’ and a certain universe, as was shown in the three-dimensional universe in our previous example. questions have a connection with their texts using the datatype property ’literaltext’ and are related to concepts via the object property ’hasquestionconcept’. these concepts can have a label which has to be stated in the form of the datatype property ’preflabel’ from the skos namespace. resulting questions are indirectly interlinked to the three-dimensional universe via figure 3. requesting all questions of a given variable figure 4. discovery – study, universe iassist quarterly 2014/2015 29 iassist quarterly relationships from the concepts to variables and from variables to the universes. the sparql query following the figure illustrates in detail the navigation from the queried questions to the concepts and the universes. almost the same sparql query should be performed in order to get each of the variables (e.g., highestschooldegree) which are assigned to particular concepts (e.g., education) and which are linked to a specific universe. select ?question where { ?universe rdf:type ddi:universe; skos:definition ?universedefinition. filter(?universedefinition = “country = ‘france’ and time =’2005’ and population = ‘age:18-65’”) ?variable rdf:type varble; ddi:holdsmeasurementof ?universe; ddi:hasconcept ?concept. ?concept rdf:type skos:concept; skos:preflabel ?conceptlabel. filter(?conceptlabel = “education”) ?question rdf:type question; ddi:hasquestionconcept ?concept; ddi:literaltext ? questiontext. filter(?questiontext = “what is your highest schooldegree?”) } so far, the researcher gets the questions joined with the question text ‘what is your highest school degree?’, the concept ’education’, and the universe with the three dimensions country, time, and population. the same researchers are now interested in the representation as wording and as code of the returned questions. variables are interconnected with their representations which are typed as ‘representation’ as well as ‘skos:conceptscheme’, since the wording (the category) and the code are both represented as instances of the class ‘skos:concept’. two datatype properties are defined for this class: ‘skos:notation’ and ‘skos:preflabel’. the datatype properties ‘skos:notation’ points to the code and ‘skos:preflabel’ to the wording representation. figure 6 shows the class axioms needed to formulate the sparql query below in order to implement the stated discovery sub use case. the ‘where’ clause of the previous sparql query has to be included. select ?question where { <where clause of previous sparql query> ?variable rdf:type variable; ddi:hasrepresentation ?representation. ?representation rdf:type skos:conceptscheme rdf:type ddi:representation. ?codecategory rdf:tyoe skos:concept; skos:inscheme ?representation; skos:preflabel ?category. skos:notation ?code; } to get a first impression of the datasets’ microdata, researchers are interested in descriptive statistics such as standard deviations, absolute or relative frequencies, and minimal, mean, or maximal figure 5. discovery – universe, variable, concept, question, instrument 30 iassist quarterly 2014/2015 iassist quarterly values. variables and values are directly connected with descriptive statistics which are of the type ‘descriptivestatistics’ and may have datatype properties like ‘percentage’ to state relative frequencies, as can be seen in the succeeding figure. if summary statistics (e.g., minimal, maximal, mean values, or standard deviations) have to be stated, instances of the class ’descriptivestatistics’ point to variables using the object property ’hasstatisticsvariable’. if the purpose is to define category statistics like absolute and relative frequencies, descriptive statistics point to skos:concepts representing values as well as categories via the object property ’hasstatisticscategory’. which questions, connected with more than one study and the three-dimensional universe, include particular keywords (e.g., “school”) in the question text? this would be another useful query, if no concepts are defined. to implement this, supplementary filters have to be set in sparql queries like filter regex(?questiontext, “school”, “i”). if access to microdata is limited or to get an overview over the entire microdata, researchers could request the aggregated data (e.g., a two-dimensional table with the dimensions ‘age’ and ‘highest school degree’) for particular studies, variables, universes, and concepts. logical datasets build the link between studies and associated aggregated data which is represented by the rdf data cube vocabulary’s class ‘dataset’. in a similar way, microdata for a specific study, variable, universe, and concept figure 6. discovery – representation figure 7. discovery – descriptive statistics iassist quarterly 2014/2015 31 iassist quarterly may be queried for analysis. the study is interconnected with an instance of the ’datafile’ class across the logical dataset. the classes ’logicaldataset’, ’dataset’, and ’datafile’, as well as their datatype and object properties, are visualized in figure 8. further use cases would be retrieving studies in which specific variables are contained or variables which are included in a specific study. using an rdf representation of the ddi data model, comparable data could be found following the same approach describing the characteristic of a given study. parts of studies could be compared, if the same ‘dataelement’ (study-independent re-usable units of information) is used. many more use cases are conceivable like finding source data related to published aggregates (tables) and finding data related to an organization or person. integration of other ontologies classes, datatype, and object properties of the ddi domain ontology can relate to existing similar classes, object, and datatype properties of other external accepted and widely adopted ontologies. conjunctions of multiple ontologies can be realized using the owl constructs owl:equivalentclass and owl:equivalentproperty (w3c 2004). if, for instance, a concept like ‘question’ is defined in the ddi domain ontology, information about possible answers and respective codes may be provided by other ontologies (see figure 9). the study title could be represented using a datatype property called ”title”. this datatype property could be newly defined in the ddi ontology or reused from already available knowledge representation systems such as dublin core (dublin core metadata initiative 2008) or the semantic web research communities ontology (ontoware.org 2012). uris are used referring to remote resources and reasoners may use additional semantic information defined in other ontologies for deductions (kupfer et al. 2007). as external ontologies may change over time, the referred concepts might not exist anymore. therefore it will be necessary to jump to past versions of respective ontologies (kupfer et al. 2007) using the owl language construct owl:versioninfo. an ontology of the spss data model and respective tools transferring metadata between ddi and spss may be built. now, class axioms (i.e., classes, datatype, and object properties) of the spss ontology and the ddi ontology can be stated as equivalent. as a consequence, there is no need to write transformation code translating rdf instances of the spss ontologies to rdf representations of the ddi figure 8. discovery – logicaldataset, dataset, datafile figure 9. integration of other ontologies 32 iassist quarterly 2014/2015 iassist quarterly ontology, as an spss class ”variable” could be defined as equivalent to a ddi class ”variable”. so far, there is no ontology of the spss data model available, but as this statistics program is commonly used, this task should be executed in the future. the members of the data.gov.uk project’s working group developed the data cube vocabulary which is based on the sdmx (statistical data and metadata exchange 2012) data model (cyganiak, reynolds & tennison 2010). sdmx (statistical data and metadata exchange), a metadata standard describing aggregate data and focusing on quantitative data, is increasingly adopted across the globe (gregory 2011). as both microdata and aggregated data are part of study descriptions, there has to be a link between ddi and sdmx. in the current version of the ddi ontology, two appropriate relations are specified. the object property ”inputvariable” points from data cube datasets to ddi variables and the object property ”hasncube” has the domain class ”logicaldataset” from the ddi namespace and the range class ”dataset” specified in the data cube vocabulary. moreover, toplevel components of the ddi conceptual model could be defined as elements of the iso/iec 11179 metadata standard (iso/iec jtc1 sc32 wg2 2012). in order to realize this, iso/iec 11179 elements may be mapped to an ontology formalizing parts of the iso/iec 11179 metadata standard. expressiveness of ontologies ontologies based on formal logic are more expressive than xml schemas. on that score the ddi data model can be depicted more precisely and additional complexity to describe concepts can be formalized as well. you cannot use xml schemas, for instance, to express that two complex types or classes are disjoint. ontologies can describe data models in greater detail than xml schemas, because they not only describe the syntax but semantics as well. xml schema and owl follow different modeling goals. on the one hand, the xml data model describes the terminology and the syntactic structure of xml documents, a node labeled tree. owl, on the other hand, is based on formal logic and on the subject-predicate-object triples from rdf. owl specifies semantic information about specific domains of interest, describes relations between domain classes, and thus allows the sharing of conceptualizations. more effective and efficient collaborations between individuals and organizations are possible if they agree on a common syntax (specified by xml schemas) and have a common understanding of the domain classes (defined by owl ontologies). xml is intended to structure and exchange documents (document-oriented), but is used to structure and exchange data (data-oriented), a purpose for which it has not been developed. also, xml schema languages like xml schema concentrate on structuring documents instead of structuring data. consistency check of the ddi data model owl reasoning techniques, terminological and assertional owl queries, are executed in order to determine if domain data models are consistent. terminological owl queries can be divided into checks for global consistency, class consistency, class equivalence, class disjointness, subsumption testing, and ontology classification. a class is inconsistent if it is equivalent to owl:nothing, an owl language construct. in general, this indicates a modeling error. are there any objects satisfying the concept definition (stuckenschmidt 2009)? if this question cannot be answered with ‘yes’, the respective concept is not consistent. an ontology is globally consistent if it is devoid of inconsistencies. unsatisfiability is often an indication of errors in concept definitions and for this reason you can test the quality of ontologies using global consistency checks (stuckenschmidt 2009). by means of classification, the ontology’s concept hierarchy can be calculated on the basis of concept definitions (stuckenschmidt 2009). instance checks, class extensions, property checks, and property extensions can be classified to assertional owl queries. instance checks are used to test if a specific individual can be assigned to a particular class (stuckenschmidt 2009). the search for all individuals contained in a given class may be performed in terms of class extensions (stuckenschmidt 2009). role checks and extensions can be defined similarly with regard to pairs of individuals. verifications of class and global consistencies provide means to check the overall consistency of the ddi-l data model and corresponding xml schemas by association of xml schema declaration and definitions with owl domain concepts. if it can be verified that the ddi ontology is consistent, meaning that the ontology does not have any contradictions, it may be derived that the ddi data model is consistent as well. facilitation of ddi elements’ comparability an extension of the actual ddi ontology and therefore its rdf representation will ease the comparability of diverse ddi elements among different ddi instances. in order to realize this, sufficient conditions specifying equality, inequality, and similarity of ddi elements have to be delineated. these conditions may be defined as immutable or recommended by an information system to researchers. thereon, scientists will only choose conditions which are relevant for their individual research questions. figure 10. equivalence of variables iassist quarterly 2014/2015 33 iassist quarterly there is a limited number of ddi-l elements like variables, questions, concepts, codes (values), categories (value labels), and study descriptions which may be compared. variables with different scales can be compared by mapping between these scales or by generating derived variables. one possible application example would be the comparison of the general qualification for university entrance in diverse countries. in figure 10, the two variables ‘grade_usa’ and ‘grade_germany’ are compared. in this example, it is defined that variables are equivalent, if all their values are defined as equivalent. it is also defined that all the value pairs of the two variables (‘a’ and ‘1.0’, for instance) are equivalent. as a consequence, owl reasoners can now derive that these two variables ‘grade_usa’ and ‘grade_germany’ are equivalent, too. both necessary and sufficient conditions may be defined. if an individual is a member of a specific class then it must satisfy the necessary conditions. if, on the other hand, some individual satisfies the sufficient conditions then this individual must be a member of a specific class (horridge 2009). in the example above, the necessary condition could be: if an individual, a particular variable, is a member of the class called ‘gradeusa_equivalent_ gradegermany’, for instance, then it must satisfy the condition that all the values of the variables ‘grade_usa’ and ‘grade_germany’ are equivalent. the sufficient condition, however, could be defined as follows: if some individual satisfies the condition that all the values of the individual ‘grade_germany’ are equivalent to values of the individual ‘grade_usa’ then this individual must be a member of the class ‘gradeusa_equivalent_gradegermany’. in this case, the class ‘gradeusa_equivalent_gradegermany’ is consistent, and this means that this class has at least one assigned individual and therefore these two variables can be seen as equivalent. another application example would be to define necessary conditions determining when two variables with a different number of age classes as values are similar and when they are not. finding and linking external resources like publications related to data publications, which describe ongoing research or its output based on research data, are typically held in bibliographical databases or information systems. adding unique, persistent identifiers established in scholarly publishing to ddi-based metadata for datasets, these datasets become citable in research publications and thereby linkable and discoverable for users. also, the extension of research data with links to relevant publications is possible by adding citations and links. such publications can directly describe study results in general or further information about specific details of a study, e.g., publications of methods or design of the study or about theories behind the study. exposing and connecting additional material related to data described in ddi is already covered in ddi codebook as well as in ddi lifecycle. because related material can vary from e.g., appendices to related sampling methods or instruments to related or outcome publications, the way to represent such information in ddi can vary from elements like ‘relatedmaterials’ or ‘otherstudymaterials’ in ddi codebook to the ‘othermaterial’ element in ddi lifecycle 3.*. in version 3.1 of the ddi metadata standard, the element ‘othermaterial’ is used to reference resources such as publications that are related to the content of the relevant module. this element includes a description, a bibliographic citation (containing 15 dublin core elements like identifier, title, creator, or date), an external reference using a url or a urn, and a reference to the item within the module to which the external resource is related (ddi alliance 2009). thus, all the necessary information characterizing the referenced resources can be stated. figure 11 depicts the xml tree of the ‘othermaterial’ element. various drawbacks are associated with this approach modeling references to external resources. the attribute ‘type’ classifies external resources. to state that the resource is a publication, the value ‘publication’ can be assigned to the attribute. this specific value, however, is not a part of a controlled vocabulary, and the possible values of the attribute ‘type’ are not explicitly defined. for this reason, applications cannot understand and process the type of the external resource. as can be seen in the xml tree of the ‘othermaterial’ element, this section of the overall data model is very complex. the semantics of the ‘othermaterial’ element are not intuitive, so you have to read the documentation to get the semantics. the number of elements for bibliographic citation is limited. in ddi 3.1 you can cite works using only 15 unqualified dublin core elements. with an extension to qualified dublin core you could realize more detailed bibliographic citations (ddi alliance 2009). references are backwards from othermaterial and note in ddi 3.* to the elements using these reusable elements. this seems to be a weighty disadvantage from a modeling perspective. ensuring reusability, it is important to store references to reusable elements in the elements using these reusable elements. using semantic web technologies, you can specify references to external resources semantically. one possible application example would be the definition of semantic references to publications as can be seen in the following figure. the class ‘referencingpublication’ is specified as the class of all the things which can have a reference to a publication via the object property ‘referencespublication’. the class ‘variable’ is a sub-class of ‘referencingpublication’. so it can be derived that every variable can   figure  11.  references  to  external  resources  in  ddi  3.1   figure 11. references to external resources in ddi 3.1 34 iassist quarterly 2014/2015 iassist quarterly also have a reference to a publication. as related publications can vary, possible link predicates can also be ‘backgroundpublication’ for a theoretical background of the study, ‘methodologypublication’ for a methodical background of the study, and ‘resultspublication’ for the representation of main results, e.g., a publication based on study. figure 13. concept relationships in ddi 3.1 figure 12. semantic references to external resources using semantic web technologies iassist quarterly 2014/2015 35 iassist quarterly you are able to describe publications further using classes of different external ontologies such as dc, marc 21, and spar. by this means, information about publications like identifier, title, creator, date, or url can be stated. marc (machine-readable cataloging), for example, is the standard for the representation and communication of bibliographic and related information in machine-readable form (library of congress marc standards 2012). bibliographic citations using 15 unqualified dublin core elements can be expanded to qualified dublin core elements. dublin core represents a very primitive way of bibliographic citing. as a consequence, dublin core has to be connected with other metadata standards. nevertheless, dublin core is a well adopted metadata standard supported by many tools (dublin core metadata initiative 2008). spar (semantic publishing and referencing ontologies) is a suite of complementary ontology modules for creating machine-readable rdf metadata for all aspects of semantic publishing and referencing. spar consists of eight ontologies, encoded in owl 2.0, which can be used either individually or in conjunction. the ontologies are revised, checked, stable, and ready for use (peroni & shotton 2011). as you can see in figure 12, the recommended data model is very simple, intuitive, and generic, implying that this data model can be applied in multiple contexts. to ensure reusability, variables only reference publications and not the other way around. using this data model, both the reference and the referenced resources are defined semantically. this model can be expanded if the classes ‘referencingresource’ and ‘resources’ are specified as super-classes of ‘referencingpublication’ and ‘publication.’ a further example, similar to semantic references to external resources such as publications, would be to define references to notes in a semantic way. concept relationships in ddi-l, users are able to state different types of relationships between concepts. the element ‘variable’ may include the element ‘conceptreference’, a reference to the concept measured by this variable. the element ‘concept’ can contain multiple elements called ‘similarconcept’. the content of this element is the element ‘similarconceptreference’, a reference to another concept that is similar to the one included in the ‘concept’ element description. the ‘similarconcept’ element may incorporate diverse elements called ‘difference’ describing the difference and the type of relationship between the concept referenced in ‘conceptreference’ and the concept referenced by the ’similarconceptreference’ element (ddi alliance 2009). figure 13 demonstrates an excerpt of the ddi 3.1 xml schemas used to describe concept relations. the ‘difference’ element can only contain text and a fraction of html markup components. no controlled vocabulary and no semantics are defined. as a consequence, the content of this element is neither machine-readable nor machine-understandable, and applications cannot know how to handle this kind of content. the ddi data model, defined using semantic web technologies, would be very simple, as you can see in figure 14, and reusable in other contexts. according to this modeling approach, variables can have concepts measured by the appropriate variables. now, you will be able to define all possible types of relations between concepts such as sub-, super-concept, and equivalence relations. as a result, the variety of connections between concepts is both readable and understandable by software components which can process the semantic information from now on in a controlled way. links to external thesauri in the current version of the ddi data model, questions, variables, data elements, descriptive statistics, and other ddi elements can relate to concepts in order to provide information about topics. ddi concepts are organized in so-called concept schemes which are similar to thesauri or classification systems regarding structure and content. when assigning concepts to ddi elements, either already available concept schemes can be reused or new concept scheme can be defined. bosch et al. (2012) describe the thesaurus linkage use case in more detail. the connection between ddi concepts and thesauri as well as knowledge systems’ terms is relevant for two reasons: when ddi entities’ concepts are defined and described, terms included in existing classification systems can be reused and provide search terms recommendation services for users. if researchers are searching for specific entities in studies, they have to state one of the concepts these entities are linked to. therefore, a precise annotation of studies’ content is significant. in many cases, researchers do not know which terms to use in the search process. to solve this problem, information systems can recommend to users suitable search terms from established thesauri or dictionaries like eurovoc (eurovoc 2012), wordnet (princeton university 2012), or lcsh (library of congress – library of congress subject headings 2012), when ddi concepts are mapped to these terms. via such mappings paths from entered search terms to actually used concepts can be detected. the advantage of the reuse of different external thesauri and knowledge organization systems is that these have often been maintained for over decades and consist of well-known and established term corpora figure 14. semantic concept relationships 36 iassist quarterly 2014/2015 iassist quarterly in their specific disciplines. the inclusion of external thesauri not only disseminates the use of such vocabularies, but also promotes the potential reuse of the ddi concepts in other linked data applications. as external thesauri are published in the lod cloud, ddi as linked data can technically be connected with thesauri. concepts in the ddi-rdf data format, linked data thesauri, and other linked data classification systems are typically represented on the web in the skos format. conceptually there are two possibilities to establish a connection between linked data thesauri and ddi-rdf: • ddi concepts can be aligned to skos concepts of other external thesauri using skos properties like skos:exactmatch, skos:relatedmatch. this mapping serves a network of related concepts over different thesauri and classification systems, which can be used to identify equivalent or related concepts. • often ddi metadata doesn’t include concepts because they are not captured. in these cases, after-the-fact relations to external thesauri could be done by means of the semantic web. therefore all questions, variables, data elements, or descriptive statistics in a study would reference directly via the ddi-rdf object properties to concepts from external data sources as their concepts. conclusions several use cases are associated with the development and the usage of an ontology of the data documentation initiative. owl reasoning enables the classification of ddi components such as studies. in a previous step, necessary conditions for the classification of studies have to be defined. ddi 3.1 can be used to depict quantitative data. dealing with qualitative data will be implemented within the scope of the next ddi subversions. one additional goal of the ontology creation is to describe both quantitative and qualitative datasets. examples of qualitative data are pictures, texts and open answers (e.g., ‘others’ as a possible response to the question ‘for what party did you vote?’). metadata of pictures, structure models of texts, and relations from qualitative to quantitative data may be formulated. researchers often do not know which terms to use if they want to search for specific topics. ddi concepts can be annotated as equivalent to concepts defined in thesauri or classification systems. as a consequence, information systems may recommend appropriate search terms in order to build more sophisticated search processes. researchers also want to discover microdata as well as aggregated data using graphical user interfaces on the internet. they can investigate, for example, which variables are connected with a specific question with a particular question text. by means of an rdf representation of the ddi ontology, both ddi data and metadata can be published in the linked open data cloud and be linked to other rdf datasets within the lod cloud. a plethora of tools can be used to process rdf data without knowing the complex ddi xml schemas’ structures. another benefit of an ontology of the ddi would be to define hierarchies and other types of relationships between ddi concepts in a semantic manner. using semantic web technologies, you can specify references to external resources like publications semantically and the comparability of ddi elements is facilitated. other external ontologies can be reused to a large extent, the ddi data model can be defined more precisely, additional more complex classes can be formalized, and owl reasoning techniques can be used to check the consistency of the overall ddi data model. references bizer, c, jentzsch, a, cyganiak, r 2012. state of the lod cloud. available from: < http://lod-cloud.net/state/ [6 may 2015]. bosch, t, mathiak, b 2011. generic multilevel approach designing domain ontologies based on xml schemas. paper presented at the workshop ontologies come of age in the semantic web, bonn, germany. bosch, t, mathiak, b 2012. xslt transformation generating owl ontologies automatically based on xml schemas. paper presented at the 6th international conference for internet technology and secured transactions, abu dhabi. bosch, t, cyganiak, r, wackerow j, zapilko b 2012. leveraging the ddi model for linked statistical data in the social, behavioural, and economic sciences. paper presented at the international conference on dublin core and metadata applications, malaysia. cyganiak, r 2011. the linking open data cloud diagram. available from: http://lod-cloud.net/. [6 may 2015]. cyganiak, r, reynolds, d & tennison, j 2010. the rdf data cube vocabulary. available from: http://publishing-statistical-data. googlecode.com/svn/trunk/specs/src/main/html/cube.html. [14 july 2010]. dagstuhl 2011. semantic statistics for social, behavioural, and economic sciences: leveraging the ddi model for the web. available from: http://www.dagstuhl.de/11372. [6 may 2015]. dagstuhl 2012. semantic statistics for social, behavioural, and economic sciences: leveraging the ddi model for the linked data web. available from: http://www.dagstuhl.de/12422. [6 may 2015]. dublin core metadata initiative 2012. dcmi metadata terms. available from: http://dublincore.org/documents/dcmi-terms/. [6 may 2015]. ddi alliance 2009. ddi 3.1 xml schema documentation. available from: http://www.ddialliance.org/specification/ddi-lifecycle/3.1/ xmlschema/fieldleveldocumentation/. [6 may 2015]. dublin core metadata initiative 2008. expressing dublin core metadata using the resource description framework (rdf). available from: http://dublincore.org/documents/dc-rdf/. [6 may 2015]. european ddi user conference 2011. european ddi user conference. available from: http://www.iza.org/conference_files/eddi2011/ call_for_papers. [6 may 2015]. eurovoc, 2012. available from: <http://eurovoc.europa.eu/>. [6 may 2015]. falcons, 2011. available from: <http://ws.nju.edu.cn/falcons/ objectsearch/index.jsp>. [6 may 2015]. gregory, a 2011. open data and metadata standards: should we be satisfied with “good enough”? technical report, open data foundation. available from: http://odaf.org/papers/open%20 data%20and%20metadata%20standards.pdf [6 may 2015] horridge, m 2009. a practical guide to building owl ontologies using protégé 4 and co-ode tools edition 1.2. university of manchester. isele, r, harth, a, umbrich j & bizer, c 2010. ldspider: an open-source crawling framework for the web of linked data. iswc 2010 posters & demonstrations track: collected abstracts, vol-658. iso/iec jtc1 sc32 wg2 2012, iso/iec 11179, information technology -metadata registries (mdr). available from: <http://metadata-stds. org/11179/>. [6 may 2015]. library of congress 2012. marc standards. available from: http:// www.loc.gov/marc/. [6 may 2015]. library of congress 2012. library of congress subject headings. available from: <http://www.loc.gov/aba/cataloging/subject/>. [6 may 2015]. linked data, 2012. available from: <http://linkeddata.org/>. [6 may 2015]. iassist quarterly 2014/2015 37 iassist quarterly linksailor, 2012. available from: < http://www.w3.org/2001/sw/wiki/ linksailor>. [6 may 2015]. kupfer, a, eckstein, s, störmann b, neumann k, mathiak b 2007. ‘methods for a synchronised evolution of databases and associated ontologies.’ in proceeding of the 2007 conference on databases and information systems iv. marbles, 2012. available from: < http://mes.github.io/marbles/>. [6 may 2015]. ontoware.org 2012, swrc ontology. available from: http://ontoware. org/swrc/. [6 may 2015]. peroni, s & shotton, d 2011. semantic publishing and referencing ontologies (spar) available from: http://sempublishing.sourceforge. net/. [6 may 2015]. princeton university 2012. wordnet – a lexical database for english. available from: <http://wordnet.princeton.edu/wordnet/>. [6 may 2015]. the tabulator, 2005. available from: <http://www.w3.org/2005/ajar/ tab>. [6 may 2015]. semantic web search engine 2012. available from: <http://www.swse. org/index.php> [1 may 2012]. sig.ma semantic information mashup 2012. available from: < https:// www.w3.org/2001/sw/wiki/sig.ma> [6 may 2015]. statistical data and metadata exchange 2012. available from: <http:// sdmx.org/> [6 may 2015]. stuckenschmidt, h (2009). ontologien: konzepte, technologien und anwendungen, springer-verlag, berlin heidelberg. w3c 2004, owl web ontology language overview. available from: http://www.w3.org/tr/2004/rec-owl-features-20040210/. [6 may 2015]. w3c 2008, sparql query language for rdf. available from: http:// www.w3.org/tr/2008/rec-rdf-sparql-query-20080115/. [6 may 2015]. w3c 2009, skos simple knowledge organization system namespace document html variant 18 august 2009 recommendation edition. available from: http://www.w3.org/2009/08/skos-reference/ skos.html. [6 may 2015]. notes 1. thomas bosch | gesis leibniz institute for the social sciences, mannheim, germany |e-mail: thomas.bosch@gesis.org 2. brigitte mathiak | gesis leibniz institute for the social sciences, mannheim, germany |e-mail: brigitte.mathiak@gesis.org 1/9 karcher, sebastian; weber, nicholas (2019) annotation for transparent inquiry: transparent data and analysis for qualitative research, iassist quarterly 43(2), pp. 1-9. doi: https://doi.org/10.29173/iq959 annotation for transparent inquiry: transparent data and analysis for qualitative research sebastian karcher1, nicholas weber2 abstract how can authors using many individual pieces of qualitative data throughout a publication make their research transparent? in this paper we introduce annotation for transparent inquiry (ati), an approach to enhance transparency in qualitative research. ati allows authors to connect specific passages in their publication with an annotation. these annotations provide additional information relevant to the passage and, when possible, include a link to one or more data sources underlying a claim; data sources are housed in a repository. after describing ati’s conceptual and technological implementation, we report on its evaluation through a series of workshops conducted by the qualitative data repository (qdr) and present initial results of the evaluation. the article ends with an outlook on next steps for the project. keywords qualitative data; research transparency; annotation; digital technology the problem: transparency for qualitative research research transparency is rapidly becoming the norm in empirical social science scholarship. journals (e.g. giofrè et al. 2017), funders (although see couture et al. 2018), professional organizations (e.g. apa 2016; apsa 2012), and peers (e.g. freese and king 2018) increasingly expect research to be transparent. in its ethics guidelines, the american political science association (apsa) distinguishes between three components of transparent research (apsa 2012): production transparency, data access, and analytic transparency. for quantitative empirical research, a basic template has become well established on how to provide both access to data and analytic transparency (king 1995). authors enable others to reproduce their findings by sharing data along with the computer code used to produce the analysis presented. a similar template, however, does not exist for qualitative research. this article presents annotation for transparent inquiry (ati), a solution for making work based on qualitative and multi-method data more transparent. ati allows scholars to annotate specific passages in a publication in order to provide extended commentary on arguments made in the text and, when possible, include a link to one or more data sources underlying a claim. linked data sources are stored in a repository. ati presents a unique mix of methodological and technological solutions to the problem of transparency for qualitative research. in the following paper, we begin by describing ati’s origins and its conceptual basis. we then outline the technical implementation and the considerations guiding the choice of technology to implement ati. in a third section we describe our process for evaluating ati based on a set of commissioned pilot studies https://doi.org/10.29173/iq959 1/9 karcher, sebastian; weber, nicholas (2019) annotation for transparent inquiry: transparent data and analysis for qualitative research, iassist quarterly 43(2), pp. 1-9. doi: https://doi.org/10.29173/iq959 and their review by subject-area experts. we present some initial results from this evaluation and conclude by outlining next steps for ati. introducing annotation for transparent inquiry calls for greater transparency in qualitative research are not novel. in a series of articles, political scientist andrew moravcsik (2010, 2014) advanced the idea of “active citation” in which “any empirical citation be hyperlinked to an annotated excerpt from the original source, which appears in a ‘transparency appendix’ at the end of the paper, article, or book chapter” (moravcsik 2014, 50). moravcsik’s idea was taken up by the newly founded qualitative data repository (qdr), which commissioned eight authors to pilot a dedicated interface for displaying and viewing active citations (e.g. crawford 2015; saunders 2015). these pilot projects demonstrated the appeal of using active citation to improve qualitative transparency, but also highlighted limitations of the approach. most importantly perhaps, the active citation viewer developed by qdr required us to obtain copyright permissions from the original publisher of the article in order to create and host a copy of the article separate from the authoritative published version. moreover, active citation was poorly connected to existing standards and trends in scholarly digital publishing. it relied on custom software and relied heavily upon online appendices, at a time when those are falling out of favor, especially for research data (kratz and strasser 2014). by lacking a clear location for shared primary data (e.g., interview transcripts or archival scans), active citation also raised concerns about the digital preservation of such data. based on these concerns, qdr sought to develop a new approach to qualitative transparency that could make use of existing open-source technologies, draw upon emerging web-based standards for scholarly publishing, and create a more robust link between active-citations and their underlying data. in the following sections we describe the conceptual and technical design of this approach, which we have called annotation for transparent inquiry (ati). conceptual implementation using ati, authors can annotate any passage in their work that they wish to add additional information to. every annotation includes at least one (and possibly all) of the following elements (see https://qdr.syr.edu/ati/ati-instructions): ● a source excerpt: typically 100 to 150 words from a textual source (e.g., an excerpt from the transcription for handwritten material, audiovisual material, or material generated through interviews or focus groups); ● a source excerpt translation: if the excerpt is not in english, a translation and indication of its source; ● an analytic note: discussion that illustrates how the data were generated and/or analyzed and how they support the empirical claim or conclusion being annotated in the text; ● data source: a link to the underlying data source that can be shared legally and ethically (currently qdr, but potentially any trusted digital archive); ● full citation: any additional details to describe the source being excerpted or linked and its location. the close link between text and data provided by annotation, both visual and in terms of the data structure, corresponds to the close interaction between text and data in most qualitative research. the https://doi.org/10.29173/iq959 https://qdr.syr.edu/ati/ati-instructions 1/9 karcher, sebastian; weber, nicholas (2019) annotation for transparent inquiry: transparent data and analysis for qualitative research, iassist quarterly 43(2), pp. 1-9. doi: https://doi.org/10.29173/iq959 linked data sources of an ati project are housed in a data repository that also provides the landing page for the data project as a whole as well as additional documentation (see figure 1). beyond standard metadata (keywords, dates, geographic tags, funding, etc.) this landing page also includes a data overview, in which authors explain both general background information about their data (collection strategy, provenance, etc.) and their “logic of annotation.” the logic of annotation provides key information to readers on what types of annotations to expect and where: authors may, for example, focus on annotating controversial claims or claims that are of great importance to their argument. they may also choose to use annotations to provide additional context or “color” to particular passages - which may be of particular interest for researchers working in more interpretivist traditions (e.g., ethnographers who seek to “thicken” a textual description, following geertz 1973). while qdr suggests some options for this, the choice of logic is left to researchers to allow for the wide variety in methodological approaches within qualitative social science. figure 1: data landing page for an ati project (here: smith and holmes-elliott 2018) the initial annotation on the article links to the data landing page, and the data landing page likewise features a prominent link to the article with annotations. this bidirectional linking ensures that readers can easily find and navigate both components of the ati data project. the data landing page is also where a digital object identifier (doi) for the data directs, allowing for separate citations to the data. a set of published pilot studies is available at http://qdr.syr.edu/ati/ati-models. technical implementation one of the major concerns in asking researchers to use new methods is the need to use and learn new tools, which are frequently a source of frustration. qdr therefore opted to not impose any given annotation technology on depositors. instead, authors can generate ati annotations using their own choice from familiar tools including comments or cross-references in word or libreoffice, comments in acrobat, hyperrefs in latex, etc. these formats are then converted into web annotations by qdr.3 annotations are displayed alongside the article using a javascript-based open-source software tool called hypothesis (http://web.hypothes.is). the annotations are accessible in a ‘public group’ (walkerhttps://doi.org/10.29173/iq959 http://qdr.syr.edu/ati/ati-models http://web.hypothes.is/ 1/9 karcher, sebastian; weber, nicholas (2019) annotation for transparent inquiry: transparent data and analysis for qualitative research, iassist quarterly 43(2), pp. 1-9. doi: https://doi.org/10.29173/iq959 peddakotla 2018) that provides a layer visible to any reader worldwide but only editable by qdr personnel. based on articles’ doi, annotations are visible across identical versions of the same article. similarly, annotations are visible on both pdf and html versions of articles (see udell 2017 for details and examples).4 hypothesis provides an ideal technical platform for transparency: the software is open source under a simplified bsd license (hypothesis 2019), the data and annotation model is based on a recently accepted w3c web standard (sanderson, ciccarese, and young 2017), and hypothesis itself is a non-profit whose goals align closely with those of open science advocates. given the open nature of the tools and the annotation data, switching to a different annotation client or supporting multiple clients will remain feasible, reducing dramatically the chance of lock-in. the qualitative data repository has recently developed a tool integrated with the dataverse repository software to ingest annotations via the hypothesis api. this tool allows curators at qdr to enter the identifier for a set of annotations and then import a json file with the annotation content, text anchor, and other relevant information. qdr’s implementation of this tool also transforms the json file into an html page. this functionality enables a viewer of the ati project on qdr to see the full set of annotations that are related to a project without having to access the full article on a publisher's website. this feature also satisfies a use case where a qdr user is searching or browsing the repository and discovers a new ati project. the user can then view all related data files and the annotations (as well as anchored texts) in one single, convenient location. evaluating ati through workshops the qualitative data repository evaluated both the conceptual and technical implementations of ati through two workshops held in 2018. these workshops were sponsored, in part, by grants from the robert wood johnson foundation and the national science foundation. participants in the first workshop, participants were recruited from two populations: a set of authors were invited to use ati to help improve the analytic, data, or production transparency of a recently published journal article. these participants were identified by recommendation from editors of journals, due to their topical expertise in qualitative data analysis, or by identifying unique use cases where ati might improve the access to underlying data. once the authors had been recruited, they were asked to suggest a graduate student or early career researcher to act as a reviewer of their ati project. for the second workshop, authors developed an annotation scheme to improve the transparency of works in-progress instead of applying annotations to previously published articles. to recruit authors of in-progress research qdr created the “ati challenge,” which acted as a call for proposals from earlycareer researchers. applicants to the ati challenge were asked to describe how and why they planned to use annotations to improve the transparency of work that they were preparing for publication in a peer-reviewed journal. a selection committee of qualitative research experts identified the most innovative proposals. similar to the first workshop, authors then suggested a topical expert that would act as a reviewer of their annotations (see table 1 for a breakdown of participants across the two workshops). https://doi.org/10.29173/iq959 1/9 karcher, sebastian; weber, nicholas (2019) annotation for transparent inquiry: transparent data and analysis for qualitative research, iassist quarterly 43(2), pp. 1-9. doi: https://doi.org/10.29173/iq959 evaluation in advance of both workshops, authors were asked to annotate their research articles with the goal of improving the narrative through analytic, data, or production transparency notes. this included depositing relevant data with qdr as described above. reviewers were asked to first read an article and mark passages where they would expect to see annotations or underlying data. reviewers then re-read the article with the authors’ annotations to evaluate how annotations improved the transparency of the argument that was being advanced. reviewers prepared a summary of their experience reading both the original and annotated version of an article, and provided written and oral feedback to the authors at the workshop. data collected from participants qdr asked the authors to keep a written log explaining the procedure they followed for selecting and applying annotations to their articles. authors also provided feedback on their experience as producers of ati through both a pre-workshop survey and a set of in-person focus groups at each workshop. reviewers were similarly asked to answer detailed questions on their experience as consumers of ati, and similarly participated in focus groups at each workshop. below, we present some preliminary observations based on this data. author / annotator reviewer state of manuscript workshop 1 established expert early career published workshop 2 early career established expert in preparation table 1: ati initiative workshops reactions to producing and consuming ati participants at both workshops expressed enthusiasm for access to tools and techniques that support a more transparent qualitative research paradigm.5 authors reported that using ati provided more space to explain and defend their use of a method or a particular data source, and an overall sense of liberation from the limited word-counts that journal publishers often impose. authors also reported that they felt their overall argument was improved by having the ability to justify not just what was included in the narrative, but why certain methods were not used or why some data had not been included or analyzed. reviewers reported an overall positive experience reading and consuming annotated articles. many reviewers observed that the annotations provided answers to initial questions they had posed about why a research design or analytic method was chosen. however, with more insight into these choices the annotations opened up new questions and lines of critique from reviewers. in many ways this should be viewed as a positive outcome -through improving production transparency and easing data access, the reviewers were able to focus more specifically on the arguments that were being advanced by an author’s data analysis. participants also identified a number of challenges facing the broad uptake and use of ati. these challenges can broadly be characterized as uncertainty in selecting what to annotate, incentives for producing ati projects, and the clear identification or signalling of different types of annotations to consumers. https://doi.org/10.29173/iq959 1/9 karcher, sebastian; weber, nicholas (2019) annotation for transparent inquiry: transparent data and analysis for qualitative research, iassist quarterly 43(2), pp. 1-9. doi: https://doi.org/10.29173/iq959 annotation selection authors who applied ati to already completed research articles faced the challenge of having to revisit previous research notes, locate data sources, and defend causal claims that were developed months and in some cases years in the past. in short, the revisiting of completed research articles was instructive as to just how onerous transparent research processes are for both producers and consumers qualitative research. authors that were applying ati to works in progress faced the additional challenge of deciding what constituted an annotation (in addition to the main narrative) and what content remained central to (and indeed a part of) the main narrative. in many instances authors of both completed and in-progress articles reported that they were conflicted as to how important annotations had become to interpreting their argument. likewise, reviewers in both workshops commented on how important an annotation seemed that they questioned why such detail was left out of or not included in the main text. the complexity of deciding which narrative explanations of production or analytic transparency remain central to an argument, and what contextual information can be supplied to support an author’s claims through annotation remains a challenge for future ati implementations. incentives for producers of an ati project, the labor needed to assemble data sources, develop concise and clear explanations of how data were collected, as well as produce analytic notes was similar in time and energy spent to produce the main article that was being annotated. many authors, and their sympathetic reviewers, were concerned with the additional delay that producing an ati project might introduce in already lengthy publication processes. for early career researchers, the potential delay of a publication was a major concern given the novelty of ati. authors participating in the second workshop faced the additional uncertainty of how annotations would be accessed and used as part of the peerreview process for getting their research published in a journal. journal editors and reviewers will not likely be familiar with ati, and this might add additional burden of educating editors and reviewers as to the importance and value of ati as a supplement to a new manuscript submission. workshop participants did not question the overall value of producing and consuming ati projects, but were unclear how these intense efforts would be rewarded. the challenge of rewarding open science practices is not unique to qualitative research (e.g. nosek et al. 2015), but the incentive structure within this domain, like techniques for transparency more generally, is not yet well established. while incentives to reward transparent research efforts, like ati, mature there are a number of practical shortterm steps that can be taken to reduce the burden of creating ati projects. we discuss these future steps in the technical challenges section below. signalling similar to the challenge of selecting which portions of an article to annotate, authors and reviewers also desired a typology of annotations that could act as a signal of the importance or function of an annotation to consumers. workshop participants spent a considerable amount of time in focus groups discussing the merits, design, and potential application of an annotation typology. in the current implementation of ati, a hypothesis-supported annotation provides an in-line text highlight to indicate where an annotation should be anchored (see figure 1 for an example). authors and reviewers at each workshop desired additional signalling mechanisms, such as different colored highlights or tags attached to an annotation in order to alert a reader’s attention to the importance and function of an annotated https://doi.org/10.29173/iq959 1/9 karcher, sebastian; weber, nicholas (2019) annotation for transparent inquiry: transparent data and analysis for qualitative research, iassist quarterly 43(2), pp. 1-9. doi: https://doi.org/10.29173/iq959 passage. reviewers in particular wanted a simple way to sort or sift through multiple annotations quickly to determine how important an author thought the annotation was to an argument, or more simply, to view all annotations that pertained to a particular mode of transparency (e.g. analytic transparency). there are existing vocabularies and data models that provide a scheme for this kind of contextual information about a web-annotation. notably, the w3c open annotation data model provides a vocabulary for typing annotations by “motivation.” (sanderson, ciccarese, and young 2017) ati could implement or extend these “motivation” features so that authors may select from a controlled vocabulary to signal to readers what motivated their application of an annotation. this functionality however remains an open challenge for future ati research. figure 2: in-text signalling and corresponding annotation in o’mahoney (2017) technical challenges the most significant technical challenge for ati (and for annotations of scholarly content to date) is the ability to reliably display annotations. hypothesis provides three mechanisms for displaying annotations: 1. using an extension for the chrome browser 2. using a snippet of javascript code installed by the content owner (in our case: the publisher) on the pages to be annotated 3. by routing readers of annotations through a proxy server that reloads the requested page with annotations loaded and visible (for instance, via.hypothes.is/iassistquarterly.com/ will load the iassist quarterly homepage with the hypothesis annotation bar on the left). the first two options work reliably for scholarly content but are still comparably rare: few users have the hypothesis chrome extension installed and the extension is currently neither available on any other browser nor on any mobile devices. hypothesis annotations are currently enabled only by a small number of publishers (such as ubiquity press and elife as well as osf preprints) and was available for the majority of articles used during the first ati workshop due to a collaboration with cambridge university press. looking ahead, however, adoption by publishers appears to be increasing rapidly. given the limited availability of the first two options, the via.hypothes.is proxy currently plays a major role in making annotations visible. unfortunately, however, it interferes with ip-based authentication, https://doi.org/10.29173/iq959 1/9 karcher, sebastian; weber, nicholas (2019) annotation for transparent inquiry: transparent data and analysis for qualitative research, iassist quarterly 43(2), pp. 1-9. doi: https://doi.org/10.29173/iq959 which researchers use to access the vast majority of articles in the social sciences. once routed through a proxy, a publisher no longer recognizes a request for an article as coming from university x, which has a subscription to the journal, and therefore does not display the full text of the requested article. next steps there are a number of future research and technical developments that will ease the adoption and use of ati by qualitative and mixed-methods researchers. as discussed above, the need to include annotations in a peer-review process may pose challenges in educating and familiarizing journal editors with the role of ati in providing access to an author’s supplementary data. qdr staff will continue to closely examine the data collected during the two ati workshops that encompasses logs and short surveys kept by authors, the feedback received from reviewers, as well as detailed notes from the two workshops that allow for a detailed initial evaluation of ati. using the results of this evaluation we expect to improve documentation and guide future technical innovations. a key area for future work that is emerging relates to bridging the divide between digital annotations and the fact that many researchers prefer reading physical paper (or paper-like) versions of publications. lastly, the labor of assembling an ati project is undeniably a barrier to adoption. to overcome the challenge of transforming annotations that are produced in the typical writing workflow of an author (e.g. comments in a word document), qdr has recently started to design a tool that will allow authors to upload their article (with comments) and data, then transform the comments into annotations that can be displayed using hypothesis, and deposited with qdr for long-term preservation. this tool is in the early stages of development, but we believe it will be an important step in easing the process of assembling these diverse resources into a structured research object the creation and sharing of ati projects. references apa. (2016), “ethical principles of psychologists and code of conduct”, american psychological association, washington d.c., available at https://www.apa.org/ethics/code/ethics-code2017.pdf. apsa. (2012), a guide to professional ethics in political science. 2nd ed., american political science association, washington, d.c., available at http://www.apsanet.org/portals/54/files/publications/apsaethicsguide2012.pdf. couture, j. l., blake, r.e., mcdonald, g., and ward, c.l. (2018), “a funder-imposed data publication requirement seldom inspired data sharing”, edited by jelte m. wicherts, plos one, vol. 13 no. 7, https://doi.org/10.1371/journal.pone.0199789. crawford, t. (2015), “data for: pivotal deterrence and the chain gang: sir edward grey’s ambiguous policy and the july crisis, 1914”, in pivotal deterrence: third-party statecraft and the pursuit of peace, qdr main collection, https://doi.org/10.5064/f6g44n6s. freese, j. and king, m.m. (2018), “institutionalizing transparency”, socius: sociological research for a dynamic world, vol. 4, https://doi.org/10.1177/2378023117739216. geertz, c. (1973), the interpretation of cultures: selected essays, basic books, new york. https://doi.org/10.29173/iq959 https://www.apa.org/ethics/code/ethics-code-2017.pdf https://www.apa.org/ethics/code/ethics-code-2017.pdf http://www.apsanet.org/portals/54/files/publications/apsaethicsguide2012.pdf https://doi.org/10.1371/journal.pone.0199789 https://doi.org/10.5064/f6g44n6s https://doi.org/10.1177/2378023117739216 1/9 karcher, sebastian; weber, nicholas (2019) annotation for transparent inquiry: transparent data and analysis for qualitative research, iassist quarterly 43(2), pp. 1-9. doi: https://doi.org/10.29173/iq959 giofrè, d., cumming, g., fresc, l., boedker, i., and tressoldi, p. (2017), “the influence of journal submission guidelines on authors’ reporting of statistics and use of open research practices.” plos one, vol. 12 no. 4, https://doi.org/10.1371/journal.pone.0175583. hypothesis. (2019), the hypothesis web-based annotation client. javascript, available at https://github.com/hypothesis/client. king, g. (1995), “replication, replication”, ps: political science & politics, vol. 28 no. 3, 444–52. kratz, j. and strasser, c. (2014), “data publication consensus and controversies.” f1000research, vol. 3 no. 94, https://doi.org/10.12688/f1000research.3979.3. moravcsik, a. (2010), “active citation: a precondition for replicable qualitative research”, ps: political science & politics, vol. 43 no. 1, 29–35, https://doi.org/10.1017/s1049096510990781. ———. (2014), “transparency: the revolution in qualitative research”, ps: political science & politics, vol. 47 no. 1, 48–53. nosek, b. a., alter, g., banks, g. c., borsboom, d., bowman, s. d., breckler, s. j., buck, s. et al. (2015), “promoting an open research culture”, science, vol. 348 no. 6242, 1422–25, https://doi.org/10.1126/science.aab2374. o’mahoney, j. (2017), “making the real: rhetorical adduction and the bangladesh liberation war”, international organization, vol. 71 no. 2, 317–48, https://doi.org/10.1017/s0020818317000054. sanderson, r., ciccarese, p., and young, b. (2017), “web annotation data model”, w3c recommendation, world wide web consortium, https://www.w3.org/tr/annotation-model/. saunders, e. n. (2015), “data for: john f. kennedy”, in leaders at war: how presidents shape military interventions, qdr main collection, https://doi.org/10.5064/f68g8hmm. smith, j. and holmes-elliott, s. (2018), “data for: the unstoppable glottal: tracking rapid change in an iconic british variable”, qualitative data repository, https://doi.org/10.5064/f6ope7mj. udell, j. (2017), “federating annotations using digital object identifiers (dois)”, hypothesis blog (blog), june 22, 2017, available at https://web.hypothes.is/blog/dois/. walker-peddakotla, a. (2018), “new collaboration capabilities for annotation: open and restricted groups”, hypothesis blog (blog), october 17, 2018, https://web.hypothes.is/blog/expandingour-groups-capabilities/. endnotes 1 sebastian karcher, qualitative data repository, syracuse university, skarcher@syr.edu 2 nicholas weber, information school, university of washington, seattle, nmweber@uw.edu; authors listed in alphabetical order. the authors would like to thank their colleagues at the qualitative data repositories as well as the authors and reviewers of the two ati initiative workshops. 3 currently the conversion is performed manually, but automated conversion is technically feasible and planned. 4 both of these features depend on the presence and accuracy of metadata in the site header, the citation_doi tag for the former and the citation_pdf_url tag for the latter functionality. 5 in many ways, this was a bias pool of researchers as they had all volunteered to participate based on their belief that not only was ati valuable to their own research activities, but that helping to improve it through the workshop focus groups, might benefit later adopters. https://doi.org/10.29173/iq959 https://doi.org/10.1371/journal.pone.0175583 https://github.com/hypothesis/client https://doi.org/10.12688/f1000research.3979.3 https://doi.org/10.1017/s1049096510990781 https://doi.org/10.1126/science.aab2374 https://doi.org/10.1017/s0020818317000054 https://www.w3.org/tr/annotation-model/ https://doi.org/10.5064/f68g8hmm https://doi.org/10.5064/f6ope7mj https://web.hypothes.is/blog/dois/ https://web.hypothes.is/blog/expanding-our-groups-capabilities/ https://web.hypothes.is/blog/expanding-our-groups-capabilities/ mailto:skarcher@syr.edu mailto:nmweber@uw.edu vol25.4 14 iassist quarterly winter 2001 iassist quarterly winter 2001 15 by robert wozniak * with the acceptance of processable metadata and the exploding growth of todayʼs online data storage capacity, current stateless, largely context-free httpor cgidriven extraction interfaces are quickly proving inadequate for traversing the vast amounts of online social science information. this paper explores ways of taking advantage of the latest technology for the discovery and access to ever-growing amounts of social science data as they are explored for the development of the nhgis project at the minnesota population center at the university of minnesota. before the web, people could only go to experts who understood the data they were interested in. they described what they were after, using what terminology they were capable of, and left it to professionals to translate their request into a language the data extraction system understood. putting an extraction process on the web, while relieving the burden on the professional, has simply shifted the burden of expertise onto the user. without the guidance of a domain expert, users are only able to rely on the informational content displayed on their computer screen. users risk spending their time scouring through a quagmire of documentation (sometimes with little context) and overwhelmed by seemingly inexhaustive and often times irrelevant lists and options. domain experts understand the ontology of their domain and can effectively draw the necessary (even common sense) inferences and deductions from a userʼs request to make a data extraction. it is this intellectual property that is missing in the vast majority of current online data extraction systems. difficult hit-or-miss keyword searches and large selection lists are the norm today. but as data grows in size, comprehension and complexity, this approach becomes a hindrance. it is of paramount importance that organizations and domain experts take advantage of current technology and incorporate as much domain knowledge as possible within their search systems. such advances will accommodate an ever-broadening user base confronted with an ever-growing amount of social science data. tomorrowʼs web-based solutions offer the means of democratizing access to data as well as interactively assisting users in understanding social science data and methodologies. leveraging the development of the ddi, rule-based grammars for middle-tier processing, and xslt-driven interface and documentation generation, the web can be used as a pedagogic device to assist both novice and expert users in compiling meaningful social sciences data in a highly dynamic, personalized and intuitive way. this democratizes access in the best possible way: first, by accommodating both novice and expert level usage; and second, by offering the means by which the novice can expand and improve upon their knowledge of social sciences and quantitative research to become, should they so choose, a domain expert themselves. nhgis is the minnesota population centerʼs first step towards making this next generation of web-enabled technology a reality. the national historical geographic information system (nhgis) overview the nhgis will make accessible the aggregate u.s. census data for all available census years between 1790 and 2000. thereʼs over a terabyte of data with over 300,000 variables for these years, all of which we propose to make accessible online, with real time data views and downloading, to students, policy analysts, journalists and academic researchers. simplifying access to this complex data so that these users will not need specialized training to make use of it is of crucial importance. the united states summary census data are the primary source of statistical information about growth and change of the american population. the great bulk of these data exist in machine-readable form, but they are largely inaccessible. the over a terabyte of data covering the period 1790 through 2000 exist or are in preparation, but they are scattered across dozens of archives and stored in incompatible formats on cd-rom, magnetic tape, or paper. only a small fraction of these data are available on the internet, and even those offer only primitive documentation and extraction tools. moreover, census summary data cannot be effectively exploited without clear definitions of each geographic unit, but high-quality electronic boundary files exist only for the 1990 census year. technological change presents an unprecedented opportunity to make these data readily available for social science emerging from the quagmire: building expert systems technologies for the social sciences 16 iassist quarterly winter 2001 iassist quarterly winter 2001 17 research. bringing the complete census within reach of social scientists will unlock the potential of two centuries of data collection, and will stimulate research in economics, history, sociology, geography and other fields. the project consists of three major components: data and documentation, mapping, and data access. • the data and documentation component gathers all extant machine-readable census summary data; fills holes in the surviving machine-readable data through data entry of paper census tabulations; harmonizes the formats and documentation of all files; and produces standardized electronic documentation according to the recently developed data documentation initiative (ddi) specification. • the mapping component creates consistent historical electronic boundary files for tracts, minor civil divisions, counties and larger geographic units. • the data access component creates a powerful but user-friendly web-based browser and extraction system, based on the new ddi metadata standard. the system provides public access free of charge to both documentation and data, and presents results in the form of tables or maps. this project was in part conceived as an online tool to relieve the burden of data archivists at the machine readable data center of mn (minnesota) from conducting a request for an extraction of the aggregate census data in person or over the phone. as a result the situation lends itself to a traditional expert system development scenario. but it does so without requiring the construction of such a system for the whole of social sciences, nor the whole of that part of the social sciences that lends itself to quantitative research. the restriction of work to well-defined domains within the social sciences as well as the availability of expertise in these domains, make this kind of approach to problems, like those faced with projects like the nhgis, possible. in general, expert systems software performs tasks otherwise performed by a human. in particular, for the nhgis project, the software will function as a component of an online data extraction engine that encapsulates higher level knowledge about the domain of u.s. aggregate census data for the purposes of efficient exact as well as approximate data discovery over large data sets. the nhgis middle-tier is designed to abstract out, make explicit, distribute and leverage expertise of its domain as opposed to automating manual procedures via the traditional development of algorithms. this is one of the key differences between knowledge-based systems like the nhgis compared to current conventional data extraction engines. what behooves one to build such a system? the decline in the cost of data storage during the last five years, as well as the exploding growth and availability of the internet, make it both possible to maintain the entire body of machine readable census data online as well as dramatically slashing the cost of access to that data for the end user. but while the storage and the port of access to this data improve, the mode of access, the underlying data discovery mechanisms employed for this access, must necessarily evolve to improve accessibility to this enormous data store for an ever-widening user base. as the thirst for social science data as well as the storage capacity for this data grow hand in hand, we are faced with the peculiar problem of effectively attaching a drinking straw to a fire hydrant. as a result, brute force and algorithmic methods of pruning search space for discovering data may prove too inefficient or otherwise cumbersome. this situation is more complicated due to the symbolic nature of the metadata as opposed to the ordered or quantifiable nature of data most algorithms apply. but itʼs not simply a question of methods we employ but also a question of how these methods are structured. current methods are procedural in the sense that they are hardwired into the process logic of the middle-tier. as such they donʼt lend themselves as easily to the old ʻplug and play ̓ type scenario where rules can be manipulated at a higher level, untangled with the inner plumbing of a systemʼs process logic, then dropped in and out of the process logic as necessary. they necessitate the work of programmers who translate the higher-level business logic of some expertise into lower level machine code that is then hand woven into the fabric of systems process logic. it is in this sense that the business logic of many systems can be considered “hardwired”. this affects the dynamics of a system, such as its ability to grow, shrink, and adapt. systems of the size and complexity of the nhgis could benefit from a modular design of the middle-tier that accommodates the quick prototyping of the business logic of the system in a non-procedural way. this necessitates the ability to easily add, modify and delete or disable business rules, which, in turn, require tools for the construction of these rules that accommodate usage by domain experts in addition to their systems ̓programmers. for example, we would like the ability to allow our domain experts to say: “actually this kind of data for this geography in 1960 doesnʼt exist, the system need not concern itself with this data at this level, in fact it need not concern itself with this class of data, at this level and all levels beneath it for all years until 1980.” they would then be able to use a tool to prototype a rule that states just that and drop it into the system for further testing. this declarative approach to rule specification are what expert systems technologies allow as a short cut 16 iassist quarterly winter 2001 iassist quarterly winter 2001 17 that can reduce both time and cost. that is, we can ignore the procedural aspects of such a rule, the reinvention of the loop, since an expert system framework takes care of that for us, allowing us to concentrate on the logic of the problem as opposed to the logic of the underlying implementation. in other words, the logic of the middle-tier more closely models the logic of the expert that defines that middle-tier. this approach also encourages tighter development and test cycles, allowing one to develop the systemʼs intellectual infrastructure incrementally in the same manner the knowledge that governs a search is obtained incrementally through experience by domain experts. with over a terabyte of data, 300,000+ variables, real time, online data viewing and downloading and a commitment to accessibility for users of all experience levels, this ability to keep the logic that governs search and presentation of this data in an explicit, higher level form is essential to the middle-tier component not only for its maintainability and modifiability over time but also the testing of its correctness, completeness, and consistency during development. the nhgis knowledge base while rules govern the procedure of search over the data, a knowledge base represents the structure and content of that data and is precisely what a rule base depends on for satisfying some search criteria. the nhgis knowledge base describes what entities exist in the data, it describes what these entities are, their properties and attributes as well as relevant relations that exist among them. in other words, the knowledge base consists of a high level, machine computable specification of what domain the nhgis project ranges over. it deals with defining characteristics of identity and partial identity, it makes explicit a conceptual containment of terms into set theoretic and taxonomic structures, it defines a termʼs attributes or properties, it makes these relationships computable, effectively producing the means by which we can impart semantics to these terms and definitions. in some respects the knowledge base is like an online thesaurus. information of this sort exists in any extraction system at some level but the relationships between entities in these systems are implicit in either the layout of the metadata or hardwired in the procedural code that uses that metadata. the knowledge base, on the other hand, makes these relationships explicit and computable for the extraction process. by making the relationships explicit, the system gets closer to the semantics of the metadata, since it uses the semantic relations to prune search space, build interfaces or morph a search criteria, etc. these semantic relations, for the purposes of the nhgis project, borrow from lexicography, set theory, and philosophy and include: synonymy in general, a definition of synonymy states that for two words, if a property exists such that the substitution of one word for another does not change the meaning of the sentence in which that substitution occurs, then the words can be considered synonymous. for nhgis, we define synonymy as the property that exists between variables where the substitution of one for the other does not change the data for which those variables refer. this property is especially important for questions of comparability of variables across time. hyponymy and hypernymy the relationships of hyponymy and hypernomy classify entities into a hierarchical categorization of classes and instances, they denote in set theoretic terms to what set an object belongs and the attributes it may therefore inherit. for instance, the variable “poverty status” belongs to the set “population characteristic” and may therefore inherit much of its identity from the definition of the term “population characteristic”. in this example, “poverty status” is a hyponym of “population characteristic” and “population characteristic” is a hypernym of the variable “poverty status”. we can use these relationships to broaden the search to include terms of the same class as terms in the input, to assist the dynamic construction of interfaces, and in the translation of input. meronymy this is a part/whole relationship that decomposes an entity into its component parts or the “stuff” from which that entity is made. in the knowledge base we take apart composite entities much like a car mechanic takes apart the engine of a car. while some of these entities are not exhaustively decomposable into a numerable set of atoms, like a car engine can be completely dismantled, many are, and it is the composition of these atomic elements of an entity that often times define that entity itself. for example, the “united states of america” is composed of a numerable set of states, which in turn are composed of a numerable set of counties, etc. until you reach an atomic building block from which the composition of higher level entities are composed. in some cases, like the entity “the united states of america”, this composition comprises its functional definition. decomposition may occur along many lines for some entities in an ontology, but for the purposes of the nhgis, those lines are often times evident if not already defined. for example, the u.s. census bureauʼs hierarchical decomposition of geography for the 1990 summary files describes this kind of relationship (figure 1). 18 iassist quarterly winter 2001 figure 1 we can use this relationship to help determine variable availability for different geographies. for instance, if “income” and “education” variables are not available at the block level the system could use a proximity rule to expand the search to include the next geographic level up (or down) and deduce that these variables exist at that level instead. the system could then offer this result to the user with an explanation as to why it reached that conclusion. other lexicographic relationships we are exploring for use in the nhgis knowledge base include: antonymy similarity polysemy origins of the nhgis knowledge base and the knowledge acquisition bottleneck what keeps the development of an ontological knowledge base feasible, how can it be done and employed with a system as practical as an online data extraction system? how do you not get bogged down acquiring the knowledge for the ontology? i want to close with a few words addressing whatʼs called the “knowledge acquisition bottleneck” and how we propose to handle it for nhgis. while many papers and talks have been given to address this problem, only a couple of points are mentioned here as they pertain to the project. • most if not all of the ontology for our project, as well as much of the rules that govern that ontology, already exist in the census bureauʼs technical documentation for the summary tape files and other forms of metadata. • much of this metadata can be parsed and put into the ontology automatically with scripts and software. the problem then becomes one of devising a clever parsing scheme to handle a document as opposed to mining deeply ingrained, non-systematized intelligence from domain experts. these two facts go a long way to alleviating the knowledge acquisition problem and are something that many knowledge engineers do not have the opportunity to leverage. in fact, given the insurmountable complexities inherent in knowledge acquisition, the absence of this metadata would have been cause enough for us to reconsider. building a rule-based ontological knowledge base for any domain cannot be considered a trivial task but nor can the development of a middle-tier business logic for a project as large and complex as the nhgis. we think, however, that our approach best models the intellectual infrastructure we need to incorporate into the nhgis to successfully mine its data. the addition of this better model in turn solves some of the complexity of the development of the middle-tier since it allows for shorter-term quick prototyping as well as longer-term ease of maintenance and extendibility. in the end, it is our hope that this approach will prove beneficial not only to the nhgis project but to the development of tomorrowʼs webbased solutions for the social sciences in toto. * paper presented at the iassist conference, storrs, ct, june, 2002. robert wozniak, minnesota population center, wozniak@pop.umn.edu. mailto:wozniak@pop.umn.edu 1/3 j.h. smith and j.s. horne (2023) letter to the editors: data privacy and dna data, iassist quarterly 47(3-4), pp. 1-3. doi https://doi.org/10.29173/iq1094 the creative commons-attribution-noncommercial license 4.0 international applies to all works published by iassist quarterly. authors will retain copyright of the work and full publishing rights. letter to the editors: data privacy and dna data j.h. smith1 and j.s. horne2 in reference to: hertzog, l., et al. (2023). data management instruments to protect the personal information of children and adolescents in sub-saharan africa. iassist quarterly, 47(2). https://doi.org/10.29173/iq1044 the manuscript titled "data management instruments to protect the personal information of children and adolescents in sub-saharan africa," focuses on the challenges, complexities and strategies associated with protecting the personal information of vulnerable populations, specifically children and adolescents, within the context of recent data protection regulatory frameworks in subsaharan africa (hertzog, chen-charles, wittesaele, de graaf, titus, kelly, langwenya, baerecke, banougnin, saal & southall, 2023). hertzog et al. (2023) highlight the recent data protection regulations, such as the protection of personal information (popi) act (no. 4 of 2013)3 in south africa, in increasing the need for improved governance and protections in data management and in research and testing involving high-risk and vulnerable groups, such as children and adolescents. it is apparent that the primary objective of hertzog et al. (2023) was to dissect and comprehend what constitutes adequate measures for safeguarding the personal information (any information that may identify a person) of these vulnerable populations and to propose strategies that align effectively with the objectives of data protection regulations. in parallel to the popi act, the access to information act, (no. 2 of 2000)4 must also be considered as it deals with how a person’s personnel information may be requested. importantly, these regulations do not replace the health professions council of south africa (hspca) patient confidentiality and other legal policies (smith, 2023). moreover, hertzog et al. (2023) correctly emphasise the significance of adhering to transnational governance frameworks established by data protection regulations, funders, and institutions. recognising the necessity, we concur that regulatory alignment is required throughout the entire life cycle process, from data acquisition and storage to analysis and dissemination of the results in research and testing frameworks. thus, in order to establish a balance between ethical research and testing practices, and compliance with data protection laws, it is essential to interpret these regulatory frameworks. we welcome the fact that hertzog et al. (2023) shared their experiences and strategies thus seeking to ignite a broader dialogue about enhancing the protection of sensitive personal information for children, and adolescents, in sub-saharan africa. it has the potential to serve as a guide for scientists, community policymakers, and stakeholders engaged in research and testing initiatives with vulnerable populations. it promotes a proactive approach to data protection that complies with regulatory frameworks and advances social and health policies that benefit these populations. we intend to capitalise on the insights presented in the article by focusing on augmenting the data protection measures that are applied to the information contained in the national forensic dna database of south africa (nfdd). the criminal law (forensic procedure) amendment act, (no. 37 of https://doi.org/10.29173/iq1094 https://doi.org/10.29173/iq1044 https://creativecommons.org/licenses/by-nc/4.0/ 2/3 j.h. smith and j.s. horne (2023) letter to the editors: data privacy and dna data, iassist quarterly 47(3-4), pp. 1-3. doi https://doi.org/10.29173/iq1094 2013)5, also known as the "dna act," provided the regulatory framework and established the nfdd. the nfdd contains numerous forensic dna profiles derived from specific categories of persons and crime scene samples organised in distinct indexes, which can be used for comparison searches. these comparison searches are conducted to provide detectives with forensic dna investigative leads and aid in resolving criminal cases (smith, 2022; smith & horne, 2023). notably, the legislators of the dna act's considered the right to privacy guaranteed by section 14 of the south african constitution (1996)6 and alignment to popi. their primary goal was to have an effective nfdd, whilst safeguarding the data of individuals whose buccal samples were collected and whose forensic dna profiles are stored in the nfdd. the dna act establishes explicit restrictions and definitions regarding the permissible purpose for which the nfdd may be used. these purposes include criminal investigations, early exoneration of the innocent during investigations, and identification of missing persons and unidentified bodies. in addition, the dna act is aligned with the regulations of the national health professions act, (no. 56 of 1974)7, which stipulates that individuals must be adequately informed and provide consent, using a signed document acknowledging the reason a buccal sample was collected (smith, 2023). the dna act requirements are consistent with the popi act, which regulates data protection in south africa. the dna act mandates the destruction of buccal samples within 90 days after a forensic dna profile is uploaded to the nfdd to prevent the potential misuse of buccal samples containing an individual’s genetic information for purposes other than those authorised by the dna act, such as research. the dna act also prohibits the storage of any personally identifiable information associated with the forensic dna profile within the nfdd, except metadata such as the unique sample identifier, index code, station and reference numbers, date of profile upload, and a minor indicator, excluding any personal information such as the identity number, name and surnames, and information about a person's predisposition, physical, or medical conditions. the dna act specifies the length of time that each individual's forensic dna profile is stored. for example, the minor indicator is essential for distinguishing forensic dna profiles derived from minors, which cannot be retained on the arrestee index for over 12 months. if a criminal case is dismissed or the defendant is acquitted, these profiles must be promptly deleted. by the popi act's data management requirements, the dna act mandates specific measures to ensure the data integrity and security of the nfdd's information. in addition, it criminalises the misuse or compromise of the data's integrity within the nfdd. the act also established the national forensic oversight and ethical board (nfoeb), which is responsible for overseeing ethical compliance, implementing the act, and preserving data integrity within the nfdd. the nfoeb is also responsible for investigating any complaints regarding dna forensics and the management of the nfdd. the rapid advancement of technology in forensic dna research and testing is unavoidably creating new challenges that threaten the security of personal information. in dna typing, the emergence of multiple parallel sequencing technologies is introducing new markers, including str loci, y-str, xstr, and snps. these developments can potentially integrate dna phenotyping, which would add a layer of complexity (meintjes-van der walt & olaborede, 2023). it is strongly suggested that various stakeholders, such as forensic dna practitioners, public representatives, academia, legal experts, and ethicists convene to discuss the extensive implications these new dna technologies may have for the privacy of individuals. in addition, it is necessary to consider whether and under what conditions these novel technologies should be implemented. to safeguard individuals' privacy and constitutional rights in the face of these technological advancements, it is essential to establish a comprehensive regulatory framework and effective oversight mechanisms. https://doi.org/10.29173/iq1094 3/3 j.h. smith and j.s. horne (2023) letter to the editors: data privacy and dna data, iassist quarterly 47(3-4), pp. 1-3. doi https://doi.org/10.29173/iq1094 references hertzog, l., chen-charles, j., wittesaele, c., de graaf, k., titus, r., kelly, j., langwenya, n., baerecke, l., banougnin, b., saal, w. and southall, j. (2023). data management instruments to protect the personal information of children and adolescents in sub-saharan africa. iassist quarterly, 47(2). https://doi.org/10.29173/iq1044 meintjes-van der walt, l & olaborede, a. (2023). dna phenotyping: a possible aid in criminal investigation. south african journal of criminal justice, 36(1):1-23. https://doi.org/10.47348/sacj/v36/i1a1 smith, j.h. (2022) ‘forensic dna investigation’ in: hr dash, p shrivastava, ja lorente (eds) handbook of dna profiling. new york: springer https://doi.org/10.1007/978-981-16-4318-7_57 smith, j.h. (2023) an exploration of the identification and processing of forensic investigative leads in investigating crime in the south african police service. (2023) dphil (university of south africa). smith, j.h. & horne, j.s. (2023) ‘the value of forensic dna investigative leads in south africa’ 17(4) journal of forensic sciences & criminal investigation, 555969. https://juniperpublishers.com/jfsci/pdf/jfsci.ms.id.555969.pdf endnotes 1 school of criminal justice, university of south africa. email: thejhsmith@gmail.com. cell: 082 728 0819. 2 professor, college of law: school of criminal justice, department of police practice, university of south africa. email: hornejs@unisa.ac.za. cell: 084 582 7829. 3 https://popia.co.za/ 4 https://www.gov.za/documents/promotion-access-information-act 5 https://www.gov.za/documents/criminal-law-forensic-procedures-amendment-act-0 6 https://www.gov.za/documents/constitution-republic-south-africa-1996 7 https://www.gov.za/documents/national-health-act-regulations-taking-buccal-sample-or-withdrawal-bloodliving-persons https://doi.org/10.29173/iq1094 https://doi.org/10.29173/iq1044 https://protect.checkpoint.com/v2/___https:/doi.org/10.47348/sacj/v36/i1a1___.yzjlonvuaxnhbw9iawxlomm6bzoymwu1zmuzyzuyotc4mdi5ndk4owuxowqxnznhndhhnjo2oje0m2q6ndm5zju1mwe0owvmytdhzji1mwy0ywe3otrmndhjnzk1zdu5ythjmwi0ntzlzdm5ztuwytmxzjm0ntnkngm3mzpwolq https://doi.org/10.1007/978-981-16-4318-7_57 https://juniperpublishers.com/jfsci/pdf/jfsci.ms.id.555969.pdf mailto:thejhsmith@gmail.com mailto:hornejs@unisa.ac.za https://popia.co.za/ https://www.gov.za/documents/promotion-access-information-act https://www.gov.za/documents/criminal-law-forensic-procedures-amendment-act-0 https://www.gov.za/documents/constitution-republic-south-africa-1996 https://www.gov.za/documents/national-health-act-regulations-taking-buccal-sample-or-withdrawal-blood-living-persons https://www.gov.za/documents/national-health-act-regulations-taking-buccal-sample-or-withdrawal-blood-living-persons 1/3 rasmussen, karsten boye (2022) editor’s notes: deposit data also qualitative data and support students in obtaining the skills for data-driven research, iassist quarterly 46(2), pp. 1-3. doi https://doi.org/10.29173/iq1047 deposit data including qualitative data and support students in obtaining the skills for data-driven research welcome to the second issue of iassist quarterly for the year 2022 iq vol. 46(2). at last a really real conference took place. i am of course referring to the iassist 2022 conference in june in göteborg, sweden. many iassisters saw each other after a long time. that is typical for yearly conferences, but this was the first since 2019, after delays from covid-19 in 2020 and 2021. great work by the organizers and the participants! also, thanks to the people who participated virtually in this hybrid conference. if you missed some presentations, then hopefully you will be able to find the missing information in issues of the iassist quarterly. before sharing data, the data must be deposited! that is the focus of this issue's first article on repository success. the second article finds that 'others' often only focus on quantitative data when addressing the issue of data deposited and registered at repositories and available from libraries. my presumption is that many of us are 'others' sometimes. the use of the data is the subject of the third article that presents successful design and implementation of workshops and internships that raise the number of social science students capable of data-driven research. quantitative data research has the focus here, but i have confidence that in all areas qualitative data are also being recognized and used including among big data. the first article is 'factors contributing to repository success in recruiting data deposits' by michele hayslett and matthew jansen from the university of north carolina at chapel hill. repositories are often found at universities and similar institutions with a focus on deposit of publications and preprints. however, few repositories enable the preservation of datasets. the authors refer to some studies that outline researchers’ data management needs and how repositories can meet those needs, but few have assessed the success of various approaches, and the literature yields very few assessments of data repositories. this study examines infrastructure for accepting data into repositories and identifies factors influencing recruiting data deposits. in the iassist community the concepts of data sharing and open data are per se positive. researchers often have intentions on becoming depositors, not least because data with a persistent identifier is a plus or demanded for publication. however, the authors present many references for obstacles experienced by researchers who intend to share their data. the study consists of a survey with an undetermined number of probable participants where a small number responded to the survey which naturally limits the conclusions. however, the authors were able to formulate several recommendations. i noticed especially 'offering more advanced curation services' like code checking and offering encryption, or the quality of the offering, and good relationships with researchers that could be interpreted as establishing a relationship early in the research process. jessica hagman, university of illinois at urbana-champaign, and hilary bussell, ohio state university, are the authors of the second article 'going qual in: towards methodologically inclusive data work in academic libraries'. qualitative research is seldom central to the support for research data in academic libraries. this article is a report on data literacy, qualitative research, and academic library infrastructure around qualitative research as experienced by practicing academic librarians. specific support for qualitative research connected to the libraries is often limited to nvivo workshops that also include faculty. participants for the survey were identified using the authors' personal social media accounts, relevant email lists, and targeted outreach combined with snowball sampling. indepth interviews were performed with 13 librarians in the united states. the results show how the interviewees define qualitative data and data literacy. the general understanding was that qualitative research is under-valued because others typically are understanding data as being solely quantitative. https://doi.org/10.29173/iq1047 2/3 rasmussen, karsten boye (2022) editor’s notes: deposit data also qualitative data and support students in obtaining the skills for data-driven research, iassist quarterly 46(2), pp. 1-3. doi https://doi.org/10.29173/iq1047 you can even find this expectation of others among qualitative researchers who believe that academic libraries do not support services for their qualitative data. better information and building of relationships are needed. in the third article the geography moves from north america to the united kingdom. vanessa higgins and jackie carter from the university of manchester present 'developing data literacy: how data services and data fellowships are creating data skilled social researchers'. in the abstract, the authors promise to describe two successful approaches to data literacy training within the social sciences. not an offer we can easily refuse. the first, being delivered by the uk data service, is an extensive training programme of events and web-based materials that focuses on essential foundational data literacy skills. there is a need for quantitative data skills in the uk and this has long been recognized by influential bodies like government and business. the second is a data fellows programme delivered by the university of manchester q-step that has been developed to help undergraduate social science students gain real-world experience by applying their classroom skills in the workplace. the focus is on data analysis and data fellows should become able to critically evaluate and use numerical data. the uk data service training contains events and on-demand web-based training, for instance a session on 'getting started with secondary analysis'. furthermore, it is mostly online, involves no financial cost, and has thousands of attendees; some positive feedback quotes are included in the article. data fellows work in organizations on data-driven research for a two-month paid internship. two case studies of data fellows are presented in the article. both programmes are led by the university of manchester. the paper also discusses next steps in the global development of data literacy skills via the empoderadata project, which is trialling the data fellows programme in latin america. submissions of papers for the iassist quarterly are always very welcome. we welcome input from iassist conferences or other conferences and workshops, from local presentations or papers especially written for the iq. when you are preparing such a presentation, give a thought to turning your one-time presentation into a lasting contribution. doing that after the event also gives you the opportunity of improving your work after feedback. we encourage you to login or create an author profile at https://www.iassistquarterly.com (our open journal system application). we permit authors to have 'deep links' into the iq as well as deposition of the paper in your local repository. chairing a conference session or workshop with the purpose of aggregating and integrating papers for a special issue iq is also much appreciated as the information reaches many more people than the limited number of session participants and will be readily available on the iassist quarterly website at https://www.iassistquarterly.com. authors are very welcome to take a look at the instructions and layout: https://www.iassistquarterly.com/index.php/iassist/about/submissions authors can also contact me directly via e-mail: kbr@sam.sdu.dk. should you be interested in compiling a special issue for the iq as guest editor(s) i will also be delighted to hear from you. karsten boye rasmussen september 2022 https://doi.org/10.29173/iq1047 https://www.iassistquarterly.com/ https://www.iassistquarterly.com/ https://www.iassistquarterly.com/index.php/iassist/about/submissions 3/3 rasmussen, karsten boye (2022) editor’s notes: deposit data also qualitative data and support students in obtaining the skills for data-driven research, iassist quarterly 46(2), pp. 1-3. doi https://doi.org/10.29173/iq1047 erratum: in volume 41 (cumulative 1-4 issue, 2017) of the iassist quarterly, the following appears uncited on page 2 of the article by vlaeminck and podkrajac, “journals in economic sciences: paying lip service to reproducible research?”: “while in other sciences replicability is regarded as a fundamental principle of research and a prerequisite for the publication of results, in economic sciences it is not treated as a top priority. in 2006…” this should instead have been cited as follows: “according to höffler (2017), replicability does not take a high priority in economics. in his opinion, this is in sharp contrast to other sciences where replicability is “regarded as a fundamental principle of research and a prerequisite for the publication of results” (höffler, 2017, p.1). already in 2006…” at the authors’ request we have posted a revised version of the entire article, to include this erratum at the beginning to note the changed version, the change in the text on page 2, and the new citation in the references list. https://doi.org/10.29173/iq1047 iassvol194 23winter 1995 in the united states, depository libraries receive federal publications under the depository library program as described under title 44, chapter 19 of the united states code. depository libraries act as custodians for federal publications in exchange for providing public access to those government publications. while the census bureau started experimenting with the distribution of data on cd rom in 1985, the full scale depository distribution of cd roms started about 1990. there is no single agency within the federal government responsible for coordinating the format of the data distributed. while the government printing office is the agency responsible for distributing the depository cd roms, gpo doesn’t fulfill the role of a publisher. the format of the data files and the software to access those data files, if any, is the decision of the data producer, as is the decision whether or not to provide a given data product to depository libraries. the early momentum behind the depository distribution centered around ms dos. the census bureau began using the dbase format for their files beginning with test disc 2 distributed to depository libraries as an experiment in 1987. most census bureau cd roms are still distributed in dbase format. it was easy enough for depository libraries to provide access to the cd roms either using programs produced by the census bureau when available, or in other instances using dbase. given the nature of the depository library program, depository libraries are reactive rather than proactive. depository libraries receive publications based upon item selection surveys. each item number represents a class of documents or datafiles. depository libraries frequently do not know the file format of, and software access to, the cd rom products before they are selected. this means that as their datafile collection takes shape, they must develop access strategies after the fact. at the university of california, davis, we have made specific decisions regarding the types and levels of public service that we can provide for datafiles. these public service activities include loaning the cd roms; making the cd roms available on public microcomputer stations; making the cd roms available via the network; and providing basic extraction of data subsets for end users. we have limited our service to the provision of files, either as created by the data producer, or custom subsets. while we use a variety of applications to produce these subsets, we do not provide any access to analytical or statistical software, including mapping software. so, for example, while we provide mapping data on a regular basis, we do not produce maps. the original depository cd roms distributions did not include front end software. our options at that point were to work with the cd roms ourselves using dbase or to loan the cd roms. loaning the cds was not seriously considered at this time because very few people had access to cd rom drives. our solution was to work with the datafiles, create subsets, and distribute the data on floppy diskettes. as the data producers began to develop some user-friendly front ends to their data, we began to make these cd roms available on public microcomputers. at present, we have five ms dos microcomputers that can be used by the public to access depository cd roms. four of these cd rom stations are in a public area. we provide access to about sixty cd roms on these four public stations. the primary consideration for mounting a particular cd on these public stations is our subjective evaluation of the front-end software provided by the data producer. we mount only files with appropriate front-end software for unmediated public use on these machines. the fifth cd rom station is a computer in the back of the department that can only be accessed when the department is open. researchers can use this computer to work with cd’s that are not available on the four pc’s in the department’s public reference area, for example, lesser used cd roms such as the census cd roms for other states or incorporating complex software such as the national health interview survey cd roms. this computer is also the only computer with dbase software. where extractions are too complex to be done easily or efficiently with the vendor produced front end, we provide extraction services using dbase software at this station. we did not have a formal policy to loan cd roms until last year when we instituted a three day loan period for cd roms that were not accessed on the public cd rom stations. we are usually able to loan most of our cds due to our arrangement with the law library . as a second depository on the same campus, the law library selected the cd roms and transferred them to our department. two other developments influenced this decision. one was the increase in the number of cd roms that we have received in formats not specific to ms dos environments. these public access to large data sets in a depository library by juri stratford*, government documents department shields library university of california, davis 24 iassist quarterly include flat files, microdata files, and geographic or image files. as we don’t have appropriate software in the department to work with these types of data, we allow users to take them to other sites that do. second was the increase in pcs equipped with cd readers. while it was once rare for end users to have their own cd reader, it is now quite commonplace, and many users would prefer to work with the data on their own systems. the geographic datafiles represented the largest category of files‘ for which we did not offer any computer-based access within the department. we have a large number of departments and labs on campus working with geographic files; and while we don’t have the facilities to provide gis services in our department, we are the largest archive of raw geographic data on campus. this includes the census bureau’s 1990 and 1992 tiger line files, and the u.s. geological survey’s digital line graph series at 1:2,000,000 and 1:100,000. we are also anticipating the receipt of approximately three hundred fifty cd roms representing the u.s. geological survey’s digital orthophotoquads for california, geographic image files, in jpeg format. our network approach was not part of some grand plan, but rather a large number of circumstances coming together at all at once. we decided in spring 1994 to ask for a grant for more computing equipment to support access to geographic data files. the census bureau was starting to distribute the landview software with the tiger line files, and we felt that we could not work with this software on an 80386, our most powerful microcomputer. as 80486s were becoming common, we thought that we would ask for this. our administration increased our request for equipment; it was not coming out of their budget. we ended up with a pentium, a larger monitor, and a color printer. however, when the equipment arrived, it was still not obvious what software we would use or even what public access we would provide to the equipment. before the equipment arrived, the staff support person that i had for computing left to go to another department on campus. we decided that, rather than hire new staff, we would hire a student. when we were able to hire a student who was knowledgeable about both unix and networking, we developed our strategy around this person. so we tried linux on our new equipment. linux is a freely available unix system for intel based pc’s. we decided upon linux for several reasons. first, because of the complete network support offered; and second, because we had a student capable of installing and maintaining the system. we were also intrigued by the possibility of running grass on the system. grass is a free gis system developed by the u.s. army. grass is available for several platforms, including linux, and is supported by other gis projects on campus. while we have not worked with grass at this point, we are cooperating with gis labs on campus that use grass, and are now examining grass as an extraction tool for geographic data. our first objective was to provide anonymous ftp access to the system. we started out with two triple speed cd rom drives, and later replaced them with four double speed cd drives as we decided that the slower cd drives were adequate for throughput on the network. our most heavily used cd’s that could best be accessed in this manner were our 1992 tiger line files for california; these are distributed on three cd’s. with four cd drives, we have dedicated three cd rom drives to tiger leaving one drive free for other data. since implementing anonymous ftp access to the tiger line files, we have also made the files available via the world wide web using an http front end to the anonymous ftp access. our cd rom drives are also available to local users via nfs on an experimental and still restricted basis. we focused on network access to our large depository data sets for several reasons. first, we were not able to provide adequate access to the data within the department, and several appropriate computing facilities were available on campus. second, there were competing demands on campus for the tiger line files that could not entirely be met by loaning the cd roms: several different researchers would request the same files at the same time, and the researchers frequently did not have adequate access to cd rom drives in the gis facilities. third, network access was readily available in the building. fourth, appropriate software, i.e. the linux operating system, was freely available; and finally we were able to hire an experienced student to implement the system. so far, we have been able to distribute the raw tiger files to campus users via anonymous ftp. we have been able to make additional cd roms available on the system as necessary. we have been able to upload large data extractions created on the 80386 system, the fifth public access microcomputer described above, onto the unix system for anonymous ftp access. and finally, we have started to develop an http interface to our anonymous ftp system for access via the www. the networked access to the depository cd roms is still very much an experiment, but we have been satisfied with the success of the project so far. we still need to implement a formal system of communication via email to rotate cd roms through the fourth cd rom drive, and once that we conclude that the system is stable we need to increase our efforts to publicize the system. we do not intend to make a large number of data files available on the network permanently. our objective in developing this system has been to use the network to serve a local community of data users. however, we have no objections to outside use to the extent that outside use of the system does not compete with local user needs. * paper presented at iassist95 may 1995 quebec city, quebec, canada. vol223 10 iassist quarterly ilses: development of tools for an integrated library and (survey) data extraction service. http:// www.gamma.rug.nl/ilses project under the european (ec) telematics for libraries program. partners: progamma (netherlands), za (germany), niwi (netherlands), university of amsterdam (netherlands) and associate partners bsdp (france) and trinity college (ireland). introduction data material collected for empirical research has traditionally been computer stored and electronically distributed by data archives and data libraries. whereas publications from the same research were kept, referenced and given access to by libraries. as content providers data archives could not extend their services with relevant book and journal collections, cross referencing and lending of printed material. libraries could not give access to data related to published research or had the means to expand bibliographic references to also point at data as machine readable outcome of the research process. a situation where data and books are separately referenced without consistent cross linking, have to be searched for in separate catalogues and are given access to by different authorities and with different facilities, has consequences for any one embarking upon new research or in general needing social scientific information. it is not possible to start with general literature searches in libraries and easily trace back publications to the empirical research and collected data that is at the heart of it. neither can data archive catalogues (even when expanded with bibliographies) help with book and article searches starting from particular data collecting efforts. properly linking data and publications would need metadata standards that take such relationships into account and coordinated efforts between authors (proper citation of data sources or writing such metadata directly themselves), the library world (referencing with cross linking in new metadata formats) and the data archives (likewise referencing with cross linking). part of those efforts would also have to be a common catalogue search facility or some form of easy access from one catalogue to information in the other. world wide web techniques for linking electronic resources on the internet but also new metadata initiatives that explicitly hold linking information to related (electronic) resources, have the potential to finally bring data and book together again for searching and retrieval. a recent publication1 is referred to for a more complete treatment, including a few internet related projects that already demonstrate first attempts in this direction. one of these is icpsr’s “publication related archive”2, another the european nesstar project3. in the same publication ilses as integrated library and survey-data extraction service, a system of tools and (internet) facilities, is expanded upon as a current project funded within the library programme of the european commission. ilses addresses the same goal of integrating publication and data. to achieve this, it accommodates both content providers (libraries and data archives) and end-users.4 other approaches and further developments in the uk e(lectronic)lib(raries) program, the open journal project has been working on mechanisms and demonstrators for “citation linking” in the broadest sense. “using citations the links made by authors themselves users can navigate between their current work and a priori work in the archives of the research literature or take a recent paper and move forward, tracking the citations dynamically”5 in particular did the open journal develop internet solutions to create, maintain and give access to “distributed links” between primary and subsequent secondary sources, when electronic information is available that never received embedded links or that simply does not have the internal format to adopt such links.6 these could easily include formats like data material distributed over the web, collections following from digitization or marc type of bibliographic references, which by themselves do not have entries for cross linking. another advantage of the open journal approach could be the fact that information to be can the library and the data archive meet in active support of research in the social sciences? the case of ilses by repke de vries * http://www.gamma.rug.nl/ilses http://www.gamma.rug.nl/ilses winter1998 11 linked in “citation style” fashion, can be arbitrarily distributed over the net and that establishing the links can be a separate, dedicated activity at any point in time. central thesaurus facilities over the internet are another key issue in tying together distributed information. when metadata can be attached to both data and related publications that take indexing terms from a central (domain specific) thesaurus, future searches will bring up related material because of such common terms. the uk data archive has established such an internet based thesaurus facility, that also takes into account web based forms to submit new possible additions to the thesaurus.7 metadata the dublin core metadata definition is both finalizing in details but also has a core set that already finds application in sometimes large scale projects. last years dc 5 conference in helsinki reflect this. both a series of current dc projects were discussed8 and several dc elements received further clarification and more precise definition.9 two projects in particular address the same issue of integrating diverse but related information types one by the australian geodynamics cooperative research centre and one by the uc berkeley digital library catalog.10 it is interesting to see the integration approached in a distributed dc metadata fashion instead of by a central database model. at and following the conference the dc.relation element (among others) received further definition. the now proposed six sub-elements have enormous potential for cross-referencing and thus having build-in links to go from a dc data material description to following dc descriptions of article and book publications especially where these are electronically available as well.11 conclusion the technology, connectivity and development of standards is available enough to start bridging the two infrastructures of access to data and access to publications following analysis of those data or touching upon the same theme. ilses addresses that goal with very concrete tools and solutions. hopefully it can both be useful for end-users and at the same serve as an evaluation of the type of approach chosen, i.e. a central database keeping all the metadata and the linking information. it seems that the library world with its digital library projects and its developing metadata standards and distributed models have an advantage for future solutions and more momentum to realize these. the data archiving world should be quick to take up the challenge and start working with library people on the common cause of giving researchers complete access to all the related information of any type, following from their research activities. from the beginning should this access be complemented by facilities for these same researchers (data collectors, authors, publishers, depositors) to create metadata and linking (citation-) information equally well themselves. references 1. vries, r.e.de: ilses: how library and data archive meet in active support of research in the social sciences. inspel vol. 31 no. 4, 1997. the article is also electronically available in pdf format: <http://www.fhpotsdam.de/~ifla/inspel> 2. icpsr “publication related archive” <http:// www.icpsr.umich.edu/icpsr/other_resources/pra.html> 3. nesstar (networking european social science tools and resources) <http://dawww.essex.ac.uk/projects/ nesstar/> though nesstar’s first goal is integrated access across holdings at different data archives, the choice for z39.50 at least opens their catalog searching to libraries and thus bridges the separate infrastructures 4. ilses <http://www.gamma.rug.nl/ilses> 5. from the iriss’98 conference: hitchcock, steve et al.: webs of research: putting the user in control. <http:// sosig.ac.uk/iriss/papers/paper42.htm> 6. carr, leslie et al.: the distributed link service: a tool for publishers, authors and readers. <http://www.w3.org/ conferences/www4/papers/178/> please note the iriss’98 paper in note 5 has references to further work on these tools by the open journal project 7. the hasset thesaurus project at the uk data archive, jointly with other parties: <http://biron.essex.ac.uk/cgi-bin/ zhasset 8. dc conference october 1997: project presentations <http://linnea.helsinki.fi/meta/projects.html> 9. a very thorough account is given by diann rusch-feja: entwicklungen der dublin core metadaten: bericht ueber den 5. dublin core metadata workshop .... (full title abbreviated) <http://www.mpib-berlin.mpg.de/dok/ dc5ber.htm> 10. project presentations 19 and 25 from the document in note 8 11. see rusch-feja’s article in note 9 ; the relevant section is almost self explanatory and a series of further url’s of mostly english information sources is given. * paper presented at the 1998 iassist/css conference, yale university, new haven, usa, may 19-22, 1998, repke de vries, ilses niwi: repke.de.vries@niwi.knaw.nl http://www.fh-potsdam.de/~ifla/inspel http://www.fh-potsdam.de/~ifla/inspel http://www.icpsr.umich.edu/icpsr/other_resources/pra.html http://www.icpsr.umich.edu/icpsr/other_resources/pra.html http://dawww.essex.ac.uk/projects/nesstar/ http://dawww.essex.ac.uk/projects/nesstar/ http://www.gamma.rug.nl/ilses http://sosig.ac.uk/iriss/papers/paper42.htm http://sosig.ac.uk/iriss/papers/paper42.htm http://www.w3.org/conferences/www4/papers/178/ http://www.w3.org/conferences/www4/papers/178/ http://biron.essex.ac.uk/cgi-bin/zhasset http://biron.essex.ac.uk/cgi-bin/zhasset http://linnea.helsinki.fi/meta/projects.html http://www.mpib-berlin.mpg.de/dok/dc5ber.htm http://www.mpib-berlin.mpg.de/dok/dc5ber.htm mailto:repke.de.vries@niwi.knaw.nl vol30-4.indd 16 iassist quarterly winter 2006 by janet p. stamatel1 introduction although the media regularly reports on crime and violence worldwide, there is not a large body of academic research systematically analyzing cross-national crime patterns and trends or developing rigorous explanations of international variations in crime occurrences. because of quantitative data limitations, the little knowledge that we have on the subject of cross-national crime variation tends to focus primarily on developed countries and often uses dated information (stamatel, 2006). however, there has been a renewed interest in comparative criminology over the past decade due to globalization and technological advancements that have improved the ability to conduct cross-national crime research (howard, et al., 2000). in particular, there has generally been an increase in the amount and the quality of quantitative cross-national crime data available to researchers. this paper reviews the content, data collection methods, geographic and temporal coverage, and accessibility of three main sources of publicly available, quantitative cross-national crime data, with a particular emphasis on recent changes with respect to data availability. these sources are the international police organization (interpol), the united nations crime surveys, and the european sourcebook2. there are also a number of reliability and validity issues to consider when analyzing quantitative cross-national crime data, but these issues are beyond the scope of this paper, and they have been discussed elsewhere in the academic literature (see for example, neapolitan 1997; howard et al., 2000; howard and smith 2003; rubin 2006). interpol interpol data collected by the international police organization (http://www.interpol.int/) are the oldest quantitative cross-national data source. they have been collected annually from interpol member nations since 1950. police representatives are requested to complete a multilingual, one-page tabular form recording aggregate counts of offenses known to the police for the entire country for 14 offenses: murder, sex offenses, rape, serious assault, all kinds of theft, aggravated theft, robbery, breaking and entering, motor vehicle theft, other thefts, fraud, counterfeiting, drug offenses, and the total number of recorded crimes. additionally, the police are asked to identify the percentage of these offenses that were attempts and the percentage of cases solved. lastly, respondents are asked to provide information on the total number of offenders, and the percentage of whom are female, minors, and aliens. the data collection process is voluntary. representatives are given a set of definitions for each of the offense categories and requested to report the numbers for their countries according to these guidelines. it is important to note that interpol does not provide any quality control measures on this data collection process. accordingly, the organization provides a disclaimer in their publications stating that “the information given is in no way intended for use as a basis for comparisons between different countries” and that “the figures must be interpreted with caution” (interpol 1999). however, given the limited number of quantitative cross-national crime data sources, researchers, the media, and others have regularly used interpol data for cross-national comparisons. interpol became concerned about the potential misuse of these data and stopped making them publicly available as of 2000, to the dismay of many comparative researchers3 according to the interpol web site, the crime statistics are currently only available to “authorized police users” (http://www. interpol.int/public/statistics/ics/downloadlist.asp). the first published interpol statistics contained data from 36 nations in 1954, and the number has increased since then. in 1998, 116 countries reported data to interpol, and 93 did so in 1999. however, the composition of countries varies from one year to the next. for example, only 82 countries reported data to interpol in both 1998 and 1999. this represents approximately one-third of all countries in the world. according to neapolitan (1997), 150 countries have reported data at least once to interpol. figure 1 illustrates which countries reported to interpol in 1998 and 1999. the large and populous countries of china, india, and the united states are noticeably absent from the collection during this time, and coverage for the middle east and africa is uneven. interpol data are officially available only in hard copy. until 1992 the reports, called international crime statistics, were published biennially, and then annually from 1993 until 1999 when they stopped releasing the data publicly. an overview of publicly available quantitative cross-national crime data iassist quarterly winter 2006 17 figure 1 countries reporting data to interpol in 1998 and 1999 * *grey shading indicates that the country reported data in 1998. hatchmarks indicate reporting in 1999. white represents no reporting in either year. it is not clear at this time whether interpol will allow researchers to petition for access to their data. some of the interpol data are also available electronically in the correlates of crime: a study of 52 nations, 1960-1984 dataset (icpsr 9258) compiled by richard bennett (1990), who is currently updating the collection4 united nations crime surveys the united nations crime surveys (uncs) (http://www. unodc.org/unodc/en/crime_cicp_surveys.html), officially called the united nations surveys of crime and operations of criminal justice systems, have been collecting quantitative cross-national crime data from united nations member states since 1970. the stated goal of this program is “to collect data on the incidence of reported crime and the operations of criminal justice systems with a view to improving the analysis and dissemination of that information globally. the survey results will provide an overview of trends and inter-relationships between various parts of the criminal justice system to promote informed decision-making in administration, nationally and internationally” (united nations office on drugs and crime, 2007). the uncs sends multi-lingual questionnaires to coordinating officers in united nations member countries. the coordinating officers are typically united nations correspondents who compile the data with assistance from government employees from a variety of relevant departments, such as police and corrections. the questionnaires are designed to record aggregate, national-level figures about crime and criminal justice systems in four areas: police, prosecution, courts, and corrections. the police section includes the number of offenses reported to the police annually for 18 offenses: intentional committed homicide, intentional attempted homicide, intentional homicide committed with a firearm, non-intentional homicide, major assault, total assault, rape, robbery, major theft, total theft, automobile theft, burglary, fraud, embezzlement, drug-related crime, bribery and/or corruption, kidnapping, and total recorded crimes. other data collected include personnel figures and budgets for different components of the criminal justice system, as well as the number of suspects by age and sex who encounter different stages of the criminal justice system. the united nations provides a modest level of quality control on the collected data. the data are considered official statements by national governments about the extent of crime and the operations of criminal justice systems in their countries and, therefore, these data are considered more valid than the interpol data. the uncs collects data in multi-year waves (see table 1). the first five waves were administered every five or six years, then subsequent waves were issued every three years, and most recently, every two years. in response to the demand for more recent cross-national crime data, the uncs has increased the frequency of administration and dissemination. geographic coverage has varied by wave, but on average about 80 countries participated in 18 iassist quarterly winter 2006 any given wave. countries’ participation varies by wave, similar to interpol, thereby making it diffi cult to construct longitudinal data series for a large number of countries. figure 2 illustrates the geographic coverage of the uncs for wave 7, covering 1998 to 2000. while coverage for north america and europe is quite good, it is sparse or inconsistent for other regions of the world. the uncs datasets are available electronically through the website of the united nations offi ce on drugs and crime, although the fi le format available depends upon the wave of data collection (see table 2). the fi rst fi ve waves are also available through icpsr, including a harmonized longitudinal fi le for 1970 to 1994, which can be downloaded or analyzed online (burnham and burnham 1999). european sourcebook the third, and newest, source of quantitative crossnational crime data is the european sourcebook of crime and criminal justice statistics (http://www. europeansourcebook.org/), which was modeled after the sourcebook of criminal justice statistics (http://www. albany.edu/sourcebook/) produced by the united states bureau of justice statistics. this data collection effort began in 1990 in response to the growing demand for accurate and timely crime data, particularly in the context of the growing council of europe, and the concern over the limitations of the other two quantitative cross-national crime data sources. the content of the data collected by the european sourcebook is similar to that of the other sources. in particular, the european sourcebook includes annual information on total number offenses reported to the police in each country for 14 types of crimes: total intentional homicide, completed intentional homicide, assault, rape, total robbery, armed robbery, total theft, theft of motor vehicle, bicycle theft, total burglary, domestic burglary, total drug offenses, total drug traffi cking, and serious drug traffi cking. additionally, homicide victimization data from the world health organization and select measures from the international crime victimization surveys are provided as points of comparison against the offi cial records data reported by the police. like the uncs, the european sourcebook also collects information about the number of offenders by crime type, the percentage of offenders who are female, minors, and aliens, as well as information about caseloads, staffi ng, and dispositions related to prosecution, conviction, and corrections. however, the data about the criminal justice system are considered secondary indicators and are generally collected every fi ve years, rather than annually. although the data for the european sourcebook come from offi cial records like the other two sources, the wave years # of countries 1 1970-1975 64 2 1975-1980 80 3 1980-1985 78 4 1985-1990 100 5 1990-1994 92 6 1995-1997 75 7 1998-2000 92 8 2001-2002 65 9 2003-2004 71 table 1 geographic and temporal coverage in the uncs figure 2 countries reporting data to uncs, 1998-2000 ( grey shading indicates that the country reported data in the 1998-2000 wave) interpol uncs european sourcebook hardcopy 1950-1999 waves1-3 ascii waves 1-2, 4-6 lotus 123 waves 2-3 msexcel wave 6-8 wave 1 spss waves 2-6 pdf waves 6-9 waves 1-3 msword waves1-3 table 2 quantitative cross-national crime data availability iassist quarterly winter 2006 19 european sourcebook differs in the way the data are collected and the high level of quality control imposed on the data collection effort. each country participating in the effort has a national correspondent who is an expert in crime and criminal justice statistics and who is responsible for collecting and checking the data. these experts are typically either ministry of justice employees or academics. in addition to using standard classifications schemes for collecting data across countries, the national correspondents also agree upon certain quality control measures to ensure the accuracy and reliability of the data. documentation for this collection is also more detailed than for interpol and the uncs data sources. for example, the european sourcebook not only provides descriptions of crime definitions, but also detailed explanations of why some countries do not conform to these definitions. the european sourcebook also provides a thorough discussion of the methodological limitations of this data collection. the european sourcebook began collecting data in 1990 and, like the uncs, it is administered in waves. the first wave collected data from 1990 to 1996 and included 36 european countries. the second wave gathered data from 1995 to 2000 from 40 european countries. the third waved covered the years 2000 to 2003 for 37 countries, but it was a limited edition and not all of the tables were updated. in general the coverage for europe is thorough and consistent, although some countries do not report data to this source, such as serbia, montenegro, and bosniaherzegovina. data from the european sourcebook are available electronically from the internet (http://www. europeansourcebook.org/) in a variety of file formats (see table 2). summary the three main sources of quantitative cross-national crime data share the same goal to provide reliable, annual counts of the frequency of occurrence of conventional crimes across countries and, for the uncs and the european sourcebook, measures of the operations of criminal justice systems worldwide. although the crime count data for all of these sources come from police reports, the sources differ in terms of the way the data are collected and the amount of quality control exercised. additionally, since reporting crime data to these international sources is voluntary and depends upon membership to the broader organization, the sources vary greatly in terms of geographic and temporal coverage. it may be tempting to simply combine data from these sources to maximize sample sizes, but researchers have noted considerable inconsistencies across the sources, particularly for certain offenses and time periods (e.g., bennett and lynch, 1990; howard and smith, 2003; gottschalk, et al., 2007). quantitative cross-national crime data collections have improved recently with respect to the frequency of collection, greater electronic availability, and to some extent, improved quality control. nonetheless, there are still numerous methodological concerns regarding these data and researchers should use them carefully and responsibly. references bennett, richard, and james lynch. 1990. “does a difference make a difference? comparing cross-national crime indicators.” criminology 38:153-181. bennett, richard r. 1990. “correlates of crime: a study of 52 nations, 1960-1984 [computer file].” ann arbor, mi: inter-university consortium for political and social research (icpsr 9258) [distributor]. burnham, r.w., and helen burnham. 1999. “united nations world surveys on crime trends and criminal justice systems, 1970-1994: restructured five-wave data [computer file].” ann arbor, mi: inter-university consortium for political and social research (icpsr 2513)[distributor]. european sourcebook of crime and criminal justice statistics. retrieved july 3, 2007 from http://www. europeansourcebook.org/. gottschalk, martin, tony smith, gregory j. howard, and bradley r. stevens. 2006. “explaining differences in comparative criminological research: an empirical exhibition.” international journal of comparative and applied criminal justice 30:209-234. howard, gregory j., graeme newman, and william alex pridemore. 2000. “theory, method, and data in comparative criminology.” pp. 139-211 in criminal justice 2000, volume 4: measurement and analysis of crime and justice. washington, d.c.: u.s. department of justice: national institute of justice. howard, gregory j., and tony smith. 2003. “understanding cross-national variations of crime rates in europe and north america.” pp. 23-70 in crime and criminal justice in europe and north america 1995-1997, edited by kauko aromaa, seppo leppa, sami nevala, and natalia ollus. helsinki, finland: european institute for crime prevention and control (heuni). international police organization (interpol). 1999. international crime statistics. lyons: france: icpointerpol general secretariat. neapolitan, jerome. 1997. cross-national crime: a research review and sourcebook. westport, ct: greenwood press. rubin, marilyn marks. 2006. “assessing the reliability of un and interpol crime statistics: a john jay college analysis,” paper presented at the annual meeting of the 20 iassist quarterly winter 2006 american society of criminology. los angeles, ca. stamatel, janet p. 2006. “incorporating socio-historical context into quantitative cross-national criminology.” international journal of comparative and applied criminal justice 30:177-207. united nations office on drugs and crime, “united nations surveys on crime trends and the operations of criminal justice systems.” retrieved july 3, 2007 from http://www.unodc.org/unodc/en/crime_cicp_surveys.html. endnotes: 1. a previous version of this paper was presented at the iassist meeting in ann arbor, michigan, 2006. janet p. stamatel is an assistant professor in the school of criminal justice and the department of informatics at the university at albany, 135 western ave., albany, ny 12209, jstamatel@albany.edu. 2. the three sources discussed in this paper are all official records data. quantitative cross-national crime data are also collected from victimization surveys (see the international crime victims surveys at http://www.unicri.it/wwd/ analysis/icvs/index.php) or self-report delinquency surveys (international self-report delinquency study). although these are important sources of data, they have limited geographic and temporal coverage and, therefore, they are not used as often for cross-national crime analyses. some researchers also use mortality data from the world health organization (http://www.who.int/whosis/en/) to analyze aggregate homicide victimizations. this is an important source of cross-national homicide data, although it is not as easily accessible as the sources discussed in this paper and it only provides a measure of one criminal offense. 3. dr. rosemary barberet, john jay college of criminal justice, will present a paper at the 2007 meeting of the american society of criminology titled “the contribution of interpol crime data to cross-national criminology” that will examine the impact of the loss of access to these data on comparative crime research. 4. information obtained from personal communication with the author instructions for authors of the iassist quarterly 1/9 karcher, sebastian and sophia lafferty-hess (2019) an epic journey in sharing: the story of a young researcher’s journey to share her data and the information professionals who tried to help, iassist quarterly 43 (1), pp. 1-9. doi: https://doi.org/10.29173/iq942 an epic journey in sharing: the story of a young researcher’s journey to share her data and the information professionals who tried to help sebastian karcher, sophia lafferty-hess1 abstract sharing data can be a journey with various characters, challenges along the way, and uncertain outcomes. these “epic journeys in sharing” teach information professionals about our patrons, our institutions, our community, and ourselves. in this paper, we tell a particularly dramatic data-sharing story, in effect a case study, in the form of a greek drama.2 it is the quest of – a young idealistic researcher collecting fascinating sensitive data and seeking to share it, encountering an institution doing its due diligence, helpful library folks, and an expert repository. our story has moments of joy, such as when our researcher is solely motivated to share because she wants others to be able to reuse her unique data; dramatic plot twists involving irbs; and a poignant ending. it explores major tropes and themes about how researchers’ motivations, data types, and data sensitivity can impact sharing; the importance of having clarity concerning institutional policies and procedures; and the role of professional communities and relationships. just like the chorus in greek drama provides commentary on the action, a chorus of data elders in our drama points out larger lessons that the case study has for research data management and data sharing. where actors in the greek chorus were wearing masks, our chorus carries different items, symbolizing their message, on every entry. keywords data sharing, research data management, qualitative data, collaboration dramatis personae jessica, a young undergraduate researcher sophia and jen, two library data folks at duke university sebastian and dessi, two expert repository folks at qdr (syracuse university) institutional review board staff at duke university the chorus of data elders note: the iq editors have granted permission for this article to be published in a special font to reinforce its presentation as a dramatic play. https://doi.org/10.29173/iq942 2/9 karcher, sebastian and sophia lafferty-hess (2019) an epic journey in sharing: the story of a young researcher’s journey to share her data and the information professionals who tried to help, iassist quarterly 43 (1), pp. 1-9. doi: https://doi.org/10.29173/iq942 prologue through their work, researchers produce a commodity of great value to themselves and others – data. today, funders, journals, and research communities alike are asking researchers to share that data within a repository where it can be well cared for and made available to others (nsf, 2010; plos, 2014; lupia and elman, 2014). sharing data in a repository has many benefits – extending the usefulness of the data past the original research question, preserving the data, enhancing transparency and reproducibility of findings, and increasing opportunities for collaboration. at duke university, a young social science researcher has heard the call. during her undergraduate studies, she has collected interview data and after publishing on her research (van meir, 2017), she is now looking for help allowing others to access these unique, but sensitive, qualitative data. she has turned to the library to assist her on her journey. entry of characters duke university libraries’ staff. duke university libraries provides research data management (rdm) support through their data and visualization services department. rdm staff at duke assist researchers throughout the research data lifecycle from the planning phase to sharing their data in a repository. qualitative data repository staff: the qualitative data repository (qdr) was founded at syracuse university in 2012 to “select, ingest, curate, archive, manage, durably preserve, and provide access to digital data used in qualitative and multi-method social inquiry” (qdr, 2017). qdr is the only domain repository in the us with a sole focus on qualitative social science data. (karcher, kirilova, & weber, 2016, note 4) researcher: a researcher collects or generates academic data that provides the evidence for their research claims. researchers, including students and faculty, are the primary clients of library rdm services as well as repository services. institutional review board: institutional review boards (irbs) support the ethical collection, storage, and dissemination of research involving human participants by reviewing, approving, and amending applications and protocols for research projects. episode 1 but now i will tell the lineage and the names of the heroes, and of the long sea-paths and the deeds they wrought in their wanderings; may the muses be the inspirers of my song! appolonius, the argonautica3 sophia and jen, library data folks, prepare to meet with jessica, an undergraduate researcher. it is a beautiful day in april. as is typical, the consult begins with jessica discussing pertinent details on the type of data she has collected, her motivation to share, and her specific data sharing challenges. let us examine each of these in turn. the data are primarily qualitative interviews with south american sex workers and ngos. these data are not easy to come by and represent a large corpus of around 100 interviews. as far as her motivation, she does not have to comply with any data-sharing mandate but chooses to share. her main https://doi.org/10.29173/iq942 3/9 karcher, sebastian and sophia lafferty-hess (2019) an epic journey in sharing: the story of a young researcher’s journey to share her data and the information professionals who tried to help, iassist quarterly 43 (1), pp. 1-9. doi: https://doi.org/10.29173/iq942 motivation is to give voice to her interviewees, whose stories are rarely listened to. she is particularly concerned with allowing others to benefit from her effort and to reuse the data in new research because as an undergraduate student she is unsure if she will be in a position to publish more on the data going forward. however, she also faces a common challenge she did not initially plan for data sharing when she wrote her irb protocol, and therefore, did not gain consent for archiving and sharing. she has spoken to the irb and they can help. a plan is devised where she will 1) re-consent the participants she has contact information for (around 17), 2) follow a de-identification protocol to remove all direct and indirect identifiers from the interviews, and 3) deposit the data within a protected access-restricted repository. to accomplish this she will need to file an amendment with the irb containing details on where she will ultimately deposit her data including information on the repository security protocols, the deposit agreement, and access procedures. given this background, sophia and jen consider the situation. the plan outlined already addresses some of their common concerns about ethical data sharing. the question remains where can jessica deposit her data that will both comply with the irb requirements and meet her needs? the repository must be able to provide mediated restricted access while also allowing access requests from south american researchers. using their knowledge of domain repositories that support these types of qualitative data, they identify some options. they describe the options, promise to reach out to these repositories, and get back to the jessica with a final recommendation for the best home for her data. back in their offices, they reach out to contacts made through professional networks and communities (such as colleagues met at conferences), and discuss the researcher’s needs. they determine a recommendation – the qualitative data repository. the chorus of data elders enters, each carrying a laptop. interlude 1 what can information professionals learn from this episode? first, there is universal joy at the motivation to share not based in the “stick” of mandates but the “carrot” of contributing to something larger. understanding what affects a researcher’s willingness to share provides a foundation for advocating for data sharing. while strong journal data sharing policies have been found to affect the rate of data sharing (piwowar, 2011; vines et al., 2013), researchers have also self-identified scholarly altruism as a driving factor amongst others (kim and stanton, 2015). in this story, the researcher’s motivation toward reuse suggests that quantifying data reuse (piwowar and vision, 2013) as well as gathering a shared corpus of positive data reuse stories and exemplars could provide useful tools for advocacy and education in future. the undergraduate status of the researcher also presents opportunities to consider how age and experience impacts willingness to share. the relationship between age and data-sharing behavior is a complicated question (tenopir et al., 2011; tenopir et al., 2015), but understanding this relationship opens doors to target our message and services to a new generation of researchers. second, what of the help provided by the library data folks? many academic institutions are increasing, or planning to increase, data management services in the face of growing need (cox and pinfield, 2013; tenopir et al., 2014). local data professionals provide a first line of support for their research community https://doi.org/10.29173/iq942 4/9 karcher, sebastian and sophia lafferty-hess (2019) an epic journey in sharing: the story of a young researcher’s journey to share her data and the information professionals who tried to help, iassist quarterly 43 (1), pp. 1-9. doi: https://doi.org/10.29173/iq942 and draw upon in-depth knowledge of data management resources and best practices. this knowledge base comes from education, hands-on experience, and active engagement with professional communities and networks. (rice & southall, 2016, chap. 1) by harnessing all available resources, including resources external to one’s local institution, data professionals can provide more holistic service delivery. as seen through this story, these resources also include the human relationships established through professional interactions at meetings and conferences. episode 2 meantime from the ship the chiefs had sent aethalides the swift herald, to whose care they entrusted their messages and the wand of hermes, his sire, who had granted him a memory of all things, that never grew dim. appolonius, the argonautica two weeks later, two offices, 1000 miles apart, one in the sun, the other in a snow-covered building. library data expert sophia reaches out to repository expert sebastian at qdr. after determining their recommendation, sophia further discusses with sebastian qdr’s ability to safeguard and provide access to jessica’s data. having ensured that qdr would be able to meet all the requirements posed by jessica and her data, an initial meeting is organized. some weeks later, using skype, jessica and sophia meet with sebastian and dessi (also at qdr) to discuss the data deposit and the required next steps. given the salience of ethical concerns in sharing qualitative data, qdr had frequently advised on irb applications and informed consent language (see kirilova and karcher 2017 for some lessons) and is able to provide advice for the irb amendment and consent language to use when recontacting interviewees. one open question at the outset is the right level of access controls on the data: who will be able to access the data and under what conditions? everyone on the call shares the same two goals: to provide the greatest ease of access to the data possible while also not risking the confidentiality promised to participants. in other words, they want to make the data “as open as possible, as closed as necessary” (h2020 programme, 2016, p. 4). together they assess the risks from a breach in confidentiality (no risk of criminal prosecution, but potentially significant reputational harm) as well as disclosure risks. these turn out to be relatively limited, as interviews contain few indirect identifiers and the researcher is confident in her de-identification protocol. the four agree on light restrictions that require a research plan and an established academic affiliation as conditions for access. these conditions are easy to meet for interested parties and easy for the repository to investigate with no need to contact jessica, thus ensuring access to the data in the long term. our four adventurers also hit on some good luck during the conversation. qualitative data in languages other than english often poses a particular challenge because repository staff is not (or only poorly) able to check de-identification, thus removing an important safety check on the data. as luck would have it, sebastian had conducted extensive research in some of the same areas as jessica and would be able to check transcripts for identification risks in their original language. moreover, qdr had recently published de-identified data from researchers working in a somewhat similar context (dunning and camp, 2015) who had established and published an extensive de-identification protocol, which they could point to as an example for jessica. https://doi.org/10.29173/iq942 5/9 karcher, sebastian and sophia lafferty-hess (2019) an epic journey in sharing: the story of a young researcher’s journey to share her data and the information professionals who tried to help, iassist quarterly 43 (1), pp. 1-9. doi: https://doi.org/10.29173/iq942 the chorus of data elders enters, each wearing a headset. interlude 2 “sharing information,” writes nancy van house (2002) “requires that users and providers trust one another.” this trust is the driver of the story we are telling today. it begins with the duke irb and jessica’s thesis advisor trusting an inexperienced researcher to research a sensitive topic abroad. it continues with jessica trusting her library for advice on sharing data. it continues with sophia and jen trusting qdr and its staff to not just treat the data responsibly, but also to be respectful to jessica and her expectations. this trust goes beyond the institutional trust embodied by such notions as “trustworthy digital repository” (beagrie et al., 2002). it relies on communication and personal interaction. it relies on a faculty advisor and an irb taking undergraduate research seriously and being willing to take the time to guide jessica. it takes the daily outreach of library staff to establish the library as a place that researchers – from undergraduate students to senior faculty – trust for advice on questions of data and data sharing. it takes the personal relationships among data professionals, built through meetings such as iassist, rdap, and idcc, to establish the trust between repositories and libraries that helped initiate the contact between duke and qdr. this trust also relies on different stakeholders playing their role throughout the data lifecycle and interacting constructively. library data professionals are uniquely situated to serve as a point of contact for researchers, build long-term relationships, and provide in-person advice and consultations. irbs safeguard human participants and can play an important role in helping researches navigate the ethics of data sharing. finally, domain repositories can provide specialized advice and data services, especially for complex data and data with complex privacy requirements. combining data and subject-level expertise, they are also in a position to actively collaborate with researchers during appraisal (i.e., assessing whether the data are a good fit and shareable in the repository) and curation to ensure ethical sharing and to help make available high-quality data and metadata. open lines of communication between all these stakeholders creates a “trusted network” that researchers can rely on to help them along their data journey. episode 3 let justice and right, to which we have both agreed, stand firm. appolonius, the argonautica back at duke. a warm north carolina summer. in the weeks following the conversation, jessica receives additional documentation from qdr with details about the proposed handling of sensitive data and suggestions for de-identification procedures as well as informed consent language. armed with this information, she files an amendment with the duke irb. the irb sends some additional inquiries about data handling and ownership, which jessica forwards to qdr, who provides the requested information. https://doi.org/10.29173/iq942 6/9 karcher, sebastian and sophia lafferty-hess (2019) an epic journey in sharing: the story of a young researcher’s journey to share her data and the information professionals who tried to help, iassist quarterly 43 (1), pp. 1-9. doi: https://doi.org/10.29173/iq942 however, in august jessica receives bad news from the irb. since she had now graduated from duke university, and the irb is no longer responsible for oversight of her work, so her irb amendment will no longer be accepted. moreover, by duke policy, undergraduate researchers are not able to take data containing identifiable information with them when leaving the institution (duke campus irb, n.d.). not being able to take the data with interviewees’ identities means that jessica will not be able to re-contact and re-consent her participants. the data cannot be shared. the chorus of data elders enters, each carrying a copy of the common rule. interlude 3 as we watch this epic journey unfold, let us further consider the role of the irb in the data sharing landscape. in this story, the irb performed its due diligence and followed the appropriate procedures to help the researcher amend the protocol for data sharing. recognition among irbs of the relevance of data sharing to their work, awareness of irb’s role in open science, and increasing dialogs between irbs and data professionals will be crucial to allow for the responsible and ethical sharing of human participant data (elman, hoelter, kapiszewski, and kirilova, 2017). this growing awareness has recently been exemplified by the inclusion of explicit data sharing language in the informed consent template proposed by cornell’s irb (karcher, 2017). further initiatives might also include irbs expanding guidelines for qualitative data sharing (jones et al., 2018) or developing joint workshops with data professionals (duke graduate school, 2018). these types of collaborative projects and dialogues can likewise help information professionals have clarity concerning institutional policies and procedures. a final broad theme evoked by this story is the importance of data management and sharing education. education about the “what” and “how” of data sharing builds awareness around good practices, resources, and tools. different pedagogical methods have been used by the rdm community including asynchronous online trainings (edina, 2017), in-person workshops and trainings, data management curricula within the classroom (whitmire, 2015), as well as data stories (dataone, n.d.). however, conceptual education can only go so far, the actual “doing” of sharing data provides essential real world experience and lays the groundwork for future sharing. that is why these “epic journeys in sharing” are important; they are the ultimate learning exercises for both researchers and information professionals. they show us the pitfalls and possibilities for data sharing out in the wild. they build capacity and present opportunities for reflection on what worked and what went wrong. exodus ill-starred one, why art thou so smitten with despair? we know how ye went in quest of the golden fleece; we know each toil of yours, all the mighty deeds ye wrought in your wanderings over land and sea. appolonius, the argonautica given the eventual outcome (the inability to share the data), this story may seem tragic. since shared data on sex workers are rare, the data would have been quite valuable, and significant investment had gone https://doi.org/10.29173/iq942 7/9 karcher, sebastian and sophia lafferty-hess (2019) an epic journey in sharing: the story of a young researcher’s journey to share her data and the information professionals who tried to help, iassist quarterly 43 (1), pp. 1-9. doi: https://doi.org/10.29173/iq942 into making it available between the researcher, duke university libraries, and qdr. yet in spite of the disappointment, this did not feel like a tragedy to those of us involved. in the course of the attempted data deposit, all involved gained important knowledge – about the ethics and logistics of data sharing, about our own approaches working with researchers and irbs, and about the particular opportunities and challenges working with undergraduate researchers. and finally, this is not the end of the story. as we all know, research follows the research data lifecycle: the end of one research product is the beginning of the next. in early 2018, jessica, now in graduate school, contacted dessi for advice on consent language to support data sharing in her next research project. the end (for now) acknowledgements the authors would like to thank their colleagues, jennifer darragh and dessi kirilova, who played a part in this sharing epic journey. particular thanks to jessica van meir, who inspired this story, gave us permission to tell it, and provided comments on a draft. references beagrie, n., doerr, m., hedstrom, m., jones, m., kenney, a., lupovici, c., … woodyard, d. (2002). trusted digital repositories: attributes and responsibilities (rlg-oclc report). mountain view, ca: rlg. retrieved from: https://www.oclc.org/content/dam/research/activities/trustedrep/repositories.pdf cox, a., & pinfield, s. (2014). research data management and libraries: current activities and future priorities. journal of librarianship and information science 46, 299–316. https://doi.org/10.1177/0961000613492542 dataone. (n.d.). data stories. available from: https://www.dataone.org/data-stories dunning, t., & camp, e. (2015). brokers, voters, and clientelism: the puzzle of distributive politics [data set]. https://doi.org/10.5064/f6z60kzb duke graduate school. (2018). rcr forum: developing a good informed consent process. retrieved from: https://gradschool.duke.edu/student-life/events/rcr-forum-developing-good-informed-consentprocess duke campus institutional review board. (n.d.). irb policies: undergraduate students as researchers. retrieved from: https://campusirb.duke.edu/irb-policies/undergraduate-students-researchers edina. (2017). mantra: research data management training. university of edinburgh. retrieved from: https://mantra.edina.ac.uk/ elman, c., hoelter, l., kapiszewski, d., & kirilova, d. (2017, november). irb guidelines and data sharing in the social science: tensions and strategies to address them. presented at the prim&r social, behavioral, and educational research conference, san antonio, tx. https://doi.org/10.6084/m9.figshare.5969104.v1 https://doi.org/10.29173/iq942 https://www.oclc.org/content/dam/research/activities/trustedrep/repositories.pdf https://doi.org/10.1177/0961000613492542 https://doi.org/10.1177/0961000613492542 https://www.dataone.org/data-stories https://doi.org/10.5064/f6z60kzb https://doi.org/10.5064/f6z60kzb https://gradschool.duke.edu/student-life/events/rcr-forum-developing-good-informed-consent-process https://gradschool.duke.edu/student-life/events/rcr-forum-developing-good-informed-consent-process https://campusirb.duke.edu/irb-policies/undergraduate-students-researchers https://mantra.edina.ac.uk/ https://doi.org/10.6084/m9.figshare.5969104.v1 https://doi.org/10.6084/m9.figshare.5969104.v1 8/9 karcher, sebastian and sophia lafferty-hess (2019) an epic journey in sharing: the story of a young researcher’s journey to share her data and the information professionals who tried to help, iassist quarterly 43 (1), pp. 1-9. doi: https://doi.org/10.29173/iq942 h2020 programme. (2016). guidelines on fair data management in horizon 2020. brussels: european commission directorate for general for research & innovation. retrieved from http://ec.europa.eu/research/participants/data/ref/h2020/grants_manual/hi/oa_pilot/h2020-hi-oadata-mgt_en.pdf jones, k., alexander, s.m., bennett, n., bishop, l., budden, a., cox, m., … winslow, d. (2018). qualitative data sharing and re-use for socio-environmental systems research: a synthesis of opportunities, challenges, resources and approaches. digital repository at the university of maryland. https://doi.org/10.13016/m2wh2dg59 karcher, s. (2017, february 9). participant protection, informed consent, and data sharing. retrieved march 21, 2018, from https://qdr.syr.edu/qdr-blog/participant-protection-informed-consent-anddata-sharing karcher, s., kirilova, d., & weber, n. (2016). beyond the matrix: repository services for qualitative data. ifla journal, 42(4), 292–302. https://doi.org/10.1177/0340035216672870 kirilova, d., & karcher, s. (2017). rethinking data sharing and human participant protection in social science research: applications from the qualitative realm. data science journal, 16. https://doi.org/10.5334/dsj-2017-043 lupia, a., & elman, c. (2014). openness in political science: data access and research transparency. ps: political science & politics, 47(1), 19–42. https://doi.org/10.1017/s1049096513001716 national science foundation (nsf). (2010). dissemination and sharing of research results. arlington, va: national science foundation. retrieved from: https://www.nsf.gov/bfa/dias/policy/dmp.jsp plos. (2014). data availability policy. retrieved from: http://journals.plos.org/plosone/s/data-availability piwowar, h.a. (2011). who shares? who doesn’t? factors associated with openly archiving raw research data. plos one 6, e18657. https://doi.org/10.1371/journal.pone.0018657 piwowar, h.a., & vision, t.j. (2013). data reuse and the open data citation advantage. peerj 1:e175 https://doi.org/10.7717/peerj.175 qualitative data repository (qdr). (2017). our mission. retrieved from: https://qdr.syr.edu/ rice, r., & southall, j. (2016). the data librarian’s handbook. london: facet publishing. abgerufen von http://www.facetpublishing.co.uk/title.php?id=300471 tenopir, c., allard, s., douglass, k., aydinoglu, a.u., wu, l., read, e., manoff, m., & frame, m. (2011). data sharing by scientists: practices and perceptions. plos one 6, e21101. https://doi.org/10.1371/journal.pone.0021101 tenopir, c., dalton, e.d., allard, s., frame, m., pjesivac, i., birch, b., pollock, d., & dorsett, k. (2015). changes in data sharing and data reuse practices and perceptions among scientists worldwide. plos one 10, e0134826. https://doi.org/10.1371/journal.pone.0134826 https://doi.org/10.29173/iq942 http://ec.europa.eu/research/participants/data/ref/h2020/grants_manual/hi/oa_pilot/h2020-hi-oa-data-mgt_en.pdf http://ec.europa.eu/research/participants/data/ref/h2020/grants_manual/hi/oa_pilot/h2020-hi-oa-data-mgt_en.pdf http://ec.europa.eu/research/participants/data/ref/h2020/grants_manual/hi/oa_pilot/h2020-hi-oa-data-mgt_en.pdf https://doi.org/10.13016/m2wh2dg59 https://doi.org/10.13016/m2wh2dg59 https://qdr.syr.edu/qdr-blog/participant-protection-informed-consent-and-data-sharing https://qdr.syr.edu/qdr-blog/participant-protection-informed-consent-and-data-sharing https://doi.org/10.1177/0340035216672870 https://doi.org/10.5334/dsj-2017-043 https://doi.org/10.5334/dsj-2017-043 https://doi.org/10.1017/s1049096513001716 https://www.nsf.gov/bfa/dias/policy/dmpfaqs.jsp#1 http://journals.plos.org/plosone/s/data-availability https://doi.org/10.1371/journal.pone.0018657 https://doi.org/10.7717/peerj.175 https://qdr.syr.edu/ http://www.facetpublishing.co.uk/title.php?id=300471 https://doi.org/10.1371/journal.pone.0021101 https://doi.org/10.1371/journal.pone.0021101 https://doi.org/10.1371/journal.pone.0134826 9/9 karcher, sebastian and sophia lafferty-hess (2019) an epic journey in sharing: the story of a young researcher’s journey to share her data and the information professionals who tried to help, iassist quarterly 43 (1), pp. 1-9. doi: https://doi.org/10.29173/iq942 tenopir, c., sandusky, r.j., allard, s., & birch, b. (2014). research data management services in academic research libraries and perceptions of librarians. library & information science research 36, 84– 90. https://doi.org/10.1016/j.lisr.2013.11.003 vines, t.h., andrew, r.l., bock, d.g., franklin, m.t., gilbert, k.j., kane, n.c., moore, j.-s., moyers, b.t., renaut, s., rennison, d.j., veen, t., & yeaman, s. (2013). mandated data archiving greatly improves access to research data. the faseb journal 27, 1304–1308. https://doi.org/10.1096/fj.12218164 van meir, jessica. (2017). sex work and the politics of space: case studies of sex workers in argentina and ecuador. social sciences 6 (2): 42. https://doi.org/10.3390/socsci6020042. whitmire, a.l. (2015). implementing a graduate-level research data management course: approach, outcomes, and lessons learned. journal of librarianship and scholarly communication 3(2), pep1246. https://doi.org/10.7710/2162-3309.1246 1 both authors have contributed equally to this paper and are listed alphabetically. sebastian karcher, associate director, qualitative data repository, skarcher@syr.edu. sophia lafferty-hess, research data management consultant, duke university libraries, sophia.lafferty.hess@duke.edu. 2 the narrative structure of this paper was inspired by the iassist 2018 conference theme of “once upon a data point: sustaining our data storytellers.” 3 all quotes from the argonautica, appolonius’ ancient account of a band of heroes on an improbable quest, are from the english translation by r. c. seaton available on the internet classics archive at http://classics.mit.edu/apollonius/argon.html https://doi.org/10.29173/iq942 https://doi.org/10.1016/j.lisr.2013.11.003 https://doi.org/10.1096/fj.12-218164 https://doi.org/10.1096/fj.12-218164 https://doi.org/10.3390/socsci6020042 https://doi.org/10.7710/2162-3309.1246 https://doi.org/10.7710/2162-3309.1246 mailto:skarcher@syr.edu mailto:sophia.lafferty.hess@duke.edu http://classics.mit.edu/apollonius/argon.html vol282-3.indd 30 iassist quarterly summer/fall 2004 introduction the analysis of numbers, beginning with raw research data and emerging with knowledge, is a vital skill. its role in a solid education is well established, but it takes specialized skills and support systems to provide optimal conditions for data analysis to fl ourish as part of an undergraduate curriculum. as stated in an article elsewhere in this issue of iq, data fi les have a layer of complexity that can make them more challenging to use than other information sources (edelstein & thompson 2005).2 data analysis is now routinely carried on almost exclusively with digital data, using some type of analytical software on a computer. optimally, incorporating such analysis in an undergraduate curriculum involves the coordination of all those resources and the skills necessary to utilize them. in setting up their courses, instructors take the assignment of required readings for granted. the support machinery and the skill sets are all in place: the library stocks the books, the students read the books and are rated on their ability to digest their contents. but what if an instructor wishes to assign students a required data analysis project? since the 1960s, completing an assignment using data has generally meant entering into the world of computers and software. how best to provide the data, computers, software and statistical and technical support that are involved in assignments that involve data analysis: this is a challenge for all colleges and universities. over the years, this challenge has been evolving as the technology itself evolves. this article examines the experience of mcgill university in responding to the changing computing and data support environment. mainframe era instructors who wanted students to do computer-based data analysis when mainframe computers were the only available option encountered many hurdles. mechanisms had to be developed for students to get computer time, codes and passwords and to be initiated into all the intricacies of submitting, correcting and retrieving jobs using the abstruse job-control language of the mainframe. by susan czarnocki 1 and anastassia khouri a library service model for digital data support of course, some students excelled, and produced reams of output. but generally, it was not an experience relished by the average undergraduate. by and large, it would only be those courses which were mandated to teach social science methods and statistics that would include a project involving analysis of primary data. pc era the introduction of the personal computer has made the computer a much more widespread part of our environment. there is an expectation that an undergraduate needs to be able to ‘use’ a computer. but often ‘use’ means interaction at the level of an expensive typewriter and communications device, for example, sending emails or surfi ng the internet. numeracy may not play a role in using a computer. undergraduate courses may be the fi rst time that a student is expected to analyze a problem using numbers and graphs. the personal computer has also dramatically changed access to data. in the ‘mainframe era’, data was being stored at the computing centre because it came on tapes which could only be read at the computing centre. with internet connectivity, every pc, whether at home or on campus, can be a potential channel for acquiring data. the physical archiving of data can be ‘anywhere’ that there is an internet connection. likewise, the analysis tools can now be installed on computers anywhere. so, in terms of access to data and analytical tools, a huge revolution has occurred. supporting the data-access revolution does this revolution, in itself, make the introduction of data analysis into undergraduate classes of 100-200 students a possibility? can data analysis be widely incorporated into the undergraduate curriculum? should it now be possible to promote use of data in any course which involves learning to make and challenge interpretations of numerical information relevant to a particular discipline? while the challenge of availability has receded, there is still the challenge of building up basic statistical literacy and software-related computer skills so that these are not barriers to a student’s progress. theoretically, just as data and software can be available ‘anywhere’, so should support. and in a sense, it is. the internet does mean iassist quarterly summer/fall 2004 31 that on-line software tutorials and guides can be available anywhere, through on-line help and a vast array of helptools. but while these are useful for a portion of the students who have had greater experience with or aptitude for computers and number, many students are not prepared to confront that mode of learning about a topic which is essentially foreign to them. there needs to be local support to help “level the playing fi eld” to some degree for those students who have not had such past experience if dataanalyses projects are required. this implies opportunities for students to obtain training in the tools that are needed and to be able to have access to some assistance in the minutiae of software techniques, producing graphs, etc. which are the major sources of frustration for the uninitiated. there also needs to be suffi cient access to computers and software so that access to them does not become an obstacle. certainly, in terms of effi ciency and reduced frustration levels for the user, these requirements are best met in a unit which can offer one-stop services for all of the elements of dealing with electronic data, from technical to statistical. activities of a data-support unit support for large classes is, of course, just one of many forms of support that a unit can undertake. the creation of easily-usable datasets is more effi ciently done once by a central service, than repeatedly by course instructors in several departments. levels of support that were once of interest mainly to graduate students can now be applied to encouraging undergraduate use of electronic data. if they can receive minimal support in locating and manipulating data-sets they fi nd more interesting, students doing individual research projects can go beyond those heavily-used standard datasets that have to be pre-prepared for instructional purposes. the local data support staff can create usable data sets on more esoteric topics, as requested, allowing the student to focus on analysing the data. in these interactions, the data-support team can offer some words of caution as to the pitfalls if a student has set his/her ambitions unrealistically high, and offer alternative suggestions. institutional home for data support in the world of distributed data and software, where is the best place to provide the assistance for being able to use what is now so readily available? as the data and tools have become much more widely diffused, it is not always evident where the support should be located for teaching how to use these tools. but without such support, it is likely that these data resources will remain underused. the mcgill experience in the 1970s and 1980s the researchers in mcgill social science departments depended on the services of the university computing centre for support for their instructional and research computing services. there was no push to create a specialized data library for those departments, or centrally for the university. during that period, the department of economics operated a statistical consulting service composed of a statistical specialist with assorted graduate students as assistants. it focussed on assistance to professors with large research grants. researchers were generating data from their own research, obtaining it from icpsr, as well as buying it at high prices, often from agencies of the canadian government. they tended to rely on the informal channels within and among departments for fi nding out what data might be available, for example, for graduate student research. some graduate assistant time was used to make a catalogue of data-fi les on tape that were being archived at the computing centre. this catalogue was a simple text fi le, stored on a mainframe account, and although the fi le was ‘public’, almost no one knew of its existence. it was not generally distributed. assistance on accessing data-fi les, or acquiring new ones via icpsr was given to any student seeking such help, but this was done as a favour to other departments, not as a major mandate of the unit. when the statistical consultant in the unit moved on to a position in the university planning offi ce, the social science departments were spurred on to re-evaluate how computer services should be provided for research and teaching. they were able to obtain a budgetary allocation for the creation of the social sciences computing centre (sscc). susan czarnocki, as the fi rst manager of this service, was given a mandate to administer the training required to use the mainframe from the remote terminals in the sscc, and to develop a set of student consultants who could assist others in resolving error messages, getting their printouts, etc. the computing centre had developed a number of procedures to provide support to instructional computing, but very few courses at the under-graduate level in the social sciences attempted to include assignments using data projects. the learning curve for using the mainframe was seen as absorbing too much energy away from the course content. in terms of more general data support services, the sscc manager was also given the role of icpsr offi cial representative (or), and was handed the catalogue of data-fi les and a list of tapes at the computing centre. students seeking to use data sets were told to speak to the sscc manager, but offering of a full-fl edged data-service never became part of the offi cial mandate of the sscc, nor of its expanded successor, the faculty of arts computer laboratory. getting the mcgill library involved during this same time-period, across canada, a wide-range of services were being developed for dealing with data access, refl ecting the variation in budgetary and structural arrangements across universities. there were services at computing centres and services attached to social science departments. those hired to offer these services often 32 iassist quarterly summer/fall 2004 met for the fi rst time as ors for icpsr, or at iassist conferences and began to share their experiences as data support professionals. the announcement of a sharp increase in the costs of acquiring the data fi les for the 1986 canadian census spurred a movement to greater collaboration. the result was a consortium formed through the canadian association of research libraries for purchase of that census data, with the university of toronto data library taking on a central role in the reproduction and delivery of the data to the other members of the consortium. through these negotiations, libraries across canada were becoming engaged with the issues of support for data services at universities. this fi rst consortium agreement formed the precedent for expanding agreements with statistics canada, and the libraries became the channel through which the data was being released. when statistics canada signed the data liberation initiative (dli) in 1996, mcgill opted to retain the library as its channel for participation.3 the head of the libraries asked anastassia khouri, who had been responsible for the computerization of the libraries, to draw up a plan for a pilot-project data service. this lead to the initiation of the mcgill libraries electronic data resources service (edrs) in the fall of 1997. the original remit of the service was to: support teaching and research in all disciplines; develop a core data collection and associated resources to support mcgill research and academic programs that would complement the printed collection available in government documents and the branch libraries; provide access to resources and data via the edrs web site where all electronic data resources services would complement the library’s current and future services. edrs and the faculty of arts computer laboratory (facl) collaborated closely in an effort to support data usage across campus. this was a collaboration involving the two authors of this article. anastassia would often provide students with data fi les, and then send them over to susan at facl so that she could make sure they understood how to use them. edrs was a resource for locating, acquiring, archiving and retrieving data, and at facl, users could obtain some limited one-on-one assistance for utilizing microdata. at that time, students could still choose to use the mcgill mainframe, or personal computers, or both -to complete their analyses. since 1997, edrs has been designing web-based interfaces to promote access to data resources and services. an acquisition budget was allocated and space on a university server has been made available for archiving. we are currently aiming at the development of a data gateway offering access to the thousands of resources that are available on the web.4 during the 1990s, the libraries also saw increasing changes in their mandate as they were moved from reporting to the vice-principal academic to the vice principalinformation systems technologies (vp-ist). when a social scientist was named as vp-ist in april 2000, he was aware of the fragmented situation of data-support, and was interested in fi nding a better approach. creating a unit that was a ‘one-stop shopping’ data service was a somewhat unconventional mandate from the perspective of many librarians. however, the successful experience of developing the edrs within the libraries provided the impetus to develop that facility. it developed along the lines of other data services that were housed in libraries around north america, and in some ways move beyond them to offer a wide range of support for research initiatives and professors interested in much broader inclusion of data utilization in their under-graduate courses. in 2001, susan czarnocki was transferred from the faculty of arts computer lab to join anastassia khouri in the edrs. since that time we have worked to develop a full range of services which aim to assist not only the heavyduty data crunchers, but also to assist any instructor, with courses of any size, to incorporate assignments that utilize data for instructional purposes. social science data social science data is now a signifi cant budgetary item for the libraries. in previous decades, micro-data were often purchased by researchers for their own use, with only infrequent access to such data sets being made available to undergraduates. macro-data tables, such as those contained in statistical yearbooks, were the main source of data for undergraduate research papers. the positioning of these volumes in reserve sections of government documents sections, or reference areas, meant that only the more ambitious students actually incorporated such material in their work. now, the library is full of computers, and all the computers can access primary data sources in the area of international development and fi nance, united nations data, and national archives and statistical offi ces around the world. in canada, statistics canada has developed an interface to canadian statistical and census data called “e-stat” which facilitates the manipulation of macro-data down to the census tract level for a wide range of social and economic characteristics5. this has greatly widened the range of undergraduate courses in which the instructor can pursue the possibilities of requiring students to obtain and analyze numeric data. through edrs, mcgill also participated actively in the design and the creation of sherlock, a shared infrastructure for bilingual access to and manipulation of numeric data. sherlock was developed as a collaborative effort by quebec university libraries.6 current edrs service model: the edrs service model has evolved since 1997 and will continue to adjust to the various changes in the delivery of iassist quarterly summer/fall 2004 33 information services. the edrs is part of a grouping of library services which have digital information as a major component: maps, electronic data and digital government information. historically, the library had maps and government documents as two separate services, operating in separate buildings. the location in separate buildings remains, but the direction of these services has been combined with the edrs to form a unifi ed administrative unit: the government information, maps, and electronic data centre. the aim is to have the staff in each sub-unit have some familiarity with the software and problematics involved in the work of all three sub-units, in order to offer an integrated service to the increasingly frequent research collaborations involving gis and government-generated digital data. the unit provides support for all research and instructional activities involving digital data. the description below focuses on the approach taken with respect to supporting use of data in undergraduate curricula and for graduate students. the edrs facility the edrs occupies a large room near the so-called ‘information commons’ area of the humanities and social science library. the room is equipped with 14 computers for student use. several of the major statistical software packages are available (spss, sas, stata, e-views), along with full documentation. it also provides access to useful tools for electronic data handling such as stattransfer and adobe acrobat. support for group instruction is also provided, with projection facilities, etc. on the premises. the geographic information centre complements those facilities by having 6 gis workstations with gis software (arcview and arcgis in addition to spss). the government information service has an additional 8 workstations. all equipment has “cuttingedge” capabilities to allow easy manipulation, analysis, saving and accessing data and the associated full-text documentation. the libraries also provide small and large electronic classrooms fully equipped with the necessary software. those facilities, when not used for teaching, are available for student use. some copies of imf print resources, such as world bank human development reports, are retained at edrs, so that students can have a quick overview of the types of data that are available from these sources, or what coverage is available for a certain country, etc. before they start to access the data electronically. the print collection of icpsr codebooks from years past is also housed at edrs, along with codebooks and user guides for many of the widely-used statistics canada surveys. various help tools have been developed to help and support teaching and research. edrs is developing a series of ‘how to’ web-pages: as an example ‘how to export a table from a pdf file into an excel spreadsheet’. multidisciplinary research and instruction through combining the resources of those units dealing with digital data, it is able to respond more creatively not only to individual student or faculty research, but to large multi-disciplinary projects involving historical, socioeconomic and gis data. more and more, funding agencies are encouraging the development of research projects that mingle and merge across interdisciplinary boundaries. the combined unit is well-situated to support such projects and is being sought by researchers as a partner in their applications for such grants. whether it is research on the fate of the aboriginal communities of the james bay cree, displaced by hydro-development, or ecological research at the mcgill biological research station in the barbados, there is a need for a blend of digital gis, numeric and offi cial government information. such collaborations make for rich research experiences which can now be extended in some cases down into upper-level undergraduate programs, such as the mcgill school for the environment, and courses on geography. support for large undergraduate classes course web-pages: when requests are received for support for the use of data in a course, the edrs staff will prepare a webpage tailored for the course, highlighting appropriate data-sources. we also offer to make a presentation based on the web-site during class-time, and encourage students to use the edrs equipment and software to complete their assignments7. software training sessions: edrs staff provide a variety of training sessions and workshops on software, such as excel, spss, stata and sas. some sessions are linked to the fulfi lment of a particular assignment, others are offered to any member of the university at the beginning of the academic year, using training spaces in the library. onsite consultation: students seeking help can come to edrs, work on an assignment for which edrs staff has made an in-class presentation, or any other work involving numeric data and get technical assistance if they are having diffi culties. support for graduate students graduate orientation: edrs staff make an effort to schedule sessions for graduate students in appropriate departments at the time of departmental graduate student orientation sessions. these presentations are aimed at demonstrating the wide array of research resources available, and indicating that the edrs is available to help them track down elusive data resources as they begin pursuing their research ideas. class presentations: they also make presentations for graduate level classes, and provide sessions on using whatever software is being used for their projects. 34 iassist quarterly summer/fall 2004 data acquisition: in some cases, graduate student research will stimulate the acquisition of resources not previously acquired. being in close touch with the graduate students allows for a good gage of what resources are lacking. the result of this approach has been to build support for several courses of 150+ students, across several departments, which are able to include assignments using major canadian and international statistical data bases, as well as census data, etc. in the past year, for example, large lecture courses in international development and in urban geography have had assignments requiring utilization of several of these complex electronic data bases. edrs supports this process in several ways: edrs staff work with the professors to package data resources specifi c to the course content; help to assess the feasibility of the data assignment; and receive feedback on how the students were able to complete the assignment. assistance to the students involves edrs staff members: demonstrating usage of these data-bases during a lecture period and helping those who come to the edrs needing extra coaching, to improve their skills in the manipulation of electronic data. in the case of the urban geography class, edrs staff prepared and presented 5 sessions demonstrating how to use excel to fulfi l the assignment. edrs also collaborated with a professor in the history department to make available original data on lebanon that he had collected, as well as train his students to use spss to undertake some simple analyses. conclusion as noted by robin rice in a report about ‘barriers to the use of numeric data in learning and teaching,’ advances in information technology are creating new spaces for learning beyond the traditional classroom, and forms of teaching beyond the traditional lecture (rice 2001).8 the personal computing ’revolution’ has made possible the speedy delivery of numeric and verbal information wherever the electronic network can be made to reach. this is a necessary condition for a great expansion in the ability of students to analyze real data in tacking real-world issues. but this is not a suffi cient condition. for real analysis to result, much infrastructure needs to be in place and many skills need to be acquired. mcgill university is experimenting with providing a unifi ed administrative unit within the library system, for providing the infrastructure and skills-training for making data analysis happen, as part of the regular undergraduate curriculum. the mcgill university data service model has many resemblances to a number of similar services elsewhere, but with a special focus on an integrated service to users, combining data acquisition and retrieval with statistical and software assistance oriented to making students skilled users of data resources. we modify our approach to data dissemination and support as the prevailing campus technology-infrastructure changes. resources are made increasingly available and we are following up every opportunity to augment the list of resources offered to our users. many collections-development strategies are selected and various acquisition methods are applied. our objectives are to augment the connectivity benefi ts of the library network, apply rigorously data and software licenses, to cooperate with our campus community and beyond it, and to coordinate our activities with multiple partners. the data world has been a model of cooperation and partnership which is encouraging the development of a true numeracy revolution. notes 1 contact: susan czarnocki, edrs centre, repath library building, room r-23, 3459 mctavish street, montreal, quebec,canada h3a 1y1. phone: +1 (514) 398-1429 / 398-4702. www: http://www.mcgill.ca/edrs/. email: susan.czarnocki@mcgill.ca 2 edelstein, daniel m. and thompson, kristi (2005), ‘a reference model for providing statistical consulting services in an academic library setting’, paper presented at iassist conference, madison, 2004. 3 data liberation initiative (dli). available at: dli ref http://www.statcan.ca/english/dli/dli.htm 4 http://www.mcgill.ca/edrs/ 5 e-stat. available at: http://www.statcan.ca/english/estat/ licence.htm 6 sherlock. available at: http://sherlock.crepuq.qc.ca/ public/anglais/sherlock.html 7 http://www.library.mcgill.ca/edrs/seminar/health/ nursing05.html) 8 rice, robin (2001), ‘understanding barriers to the user of numeric data in learning and teaching. ‘iassist quarterly. vol. 25, no. 1. pg.5-9. 1/15 hertzog, lucas; chen-charles, jenny; wittesaele, camille; de graaf, kristen; titus, raylene; kelly, jane; langwenya, nontokozo; baerecke, lauren; banougnin, boladé hamed; saal, wylene; southall, john; cluver, lucie; and toska, elona (2023) data management instruments to protect the personal information of children and adolescents in sub-saharan africa, iassist quarterly 47(2), pp. 1-15. doi: https://doi.org/10.29173/iq1044 the creative commons-attribution-noncommercial license 4.0 international applies to all works published by iassist quarterly. authors will retain copyright of the work and full publishing rights. data management instruments to protect the personal information of children and adolescents in sub-saharan africa lucas hertzog1, jenny chen-charles2, camille wittesaele3, kristen de graaf4, raylene titus5, jane kelly6, nontokozo langwenya7, lauren baerecke8, boladé hamed banougnin9, wylene saal10, john southall11, lucie cluver12, and elona toska13 abstract recent data protection regulatory frameworks, such as the protection of personal information act (popi act) in south africa and the general data protection regulation (gdpr) in the european union, impose governance requirements for research involving high-risk and vulnerable groups such as children and adolescents. our paper's objective is to unpack what constitutes adequate safeguards to protect the personal information of vulnerable populations such as children and adolescents. we suggest strategies to adhere meaningfully to the principal aims of data protection regulations. navigating this within established research projects raises questions about how to interpret regulatory frameworks to build on existing mechanisms already used by researchers. therefore, we will explore a series of best practices in safeguarding the personal information of children, adolescents, and young people (0-24 years old), who represent more than half of sub-saharan africa's population. we discuss the actions the research group took to ensure regulations such as gdpr and popia effectively build on existing data protection mechanisms for research projects at all stages, focusing on promoting regulatory alignment throughout the data lifecycle. our goal is to stimulate a broader conversation on improving the protection of sensitive personal information of children, adolescents, and young people in sub-saharan africa. we join this discussion as a research group generating evidence influencing social and health policy and programming for young people in sub-saharan africa. our contribution draws on our work adhering to multiple transnational governance frameworks imposed by national legislation, such as data protection regulations, funders, and academic institutions. keywords data management, children and adolescents, sub-saharan africa, personal information protection, health research introduction researchers curating data collected from children, adolescents, and their caregivers, often face a dilemma. on the one hand, they support the view that special protection and different mechanisms should be in place to mitigate risks associated with the research, even if it slows data collection. on https://doi.org/10.29173/iq1044 https://creativecommons.org/licenses/by-nc/4.0/ 2/15 hertzog, lucas; chen-charles, jenny; wittesaele, camille; de graaf, kristen; titus, raylene; kelly, jane; langwenya, nontokozo; baerecke, lauren; banougnin, boladé hamed; saal, wylene; southall, john; cluver, lucie; and toska, elona (2023) data management instruments to protect the personal information of children and adolescents in sub-saharan africa, iassist quarterly 47(2), pp. 1-15. doi: https://doi.org/10.29173/iq1044 the other hand, they are willing to accelerate the pace of scientific research with this population group because of the enormous benefits of science and technology (caldwell et al., 2004). this dilemma is exemplified in research on infectious diseases which predominantly affect children, such as infantile paralysis, measles, and pneumococci. without collecting data from children, we could not have achieved the development of vaccines which has led to a dramatic reduction in child mortality and improvements in the quality of life, at the same time they relied on clinical trials that are more challenging than those with adult participants (joseph, craig & caldwell, 2015). although this discussion re-gained traction through the contentious debates about the rollout of covid-19 vaccines for children and adolescents (ledford, 2021), it is not new strife. ever since the world experienced the horrors of world war ii and the tuskegee study in the united states, various declarations and codes were constructed to orient ethical research with human participants, from the declaration of helsinki in 1964 to the international ethical guidelines for health-related research involving humans and their continuous revisions (williams, 2008; shrestha & dunn, 2020). with the increased digitalisation of society and the increased risks of making behaviour amenable to datafication (mbembe, 2019), especially in quantitative research, researchers must consider the implications for research with human participants. in collecting data, where we inevitably transform responses into numbers, what are the most efficient and simple mechanisms to safeguard participants' rights? in the context of research involving children and adolescents, how do we accommodate their rights to privacy with other rights, such as the public interest, to build a corpus of knowledge that could benefit this same population? several data protection regulatory frameworks are being elaborated and implemented in this context, and we have seen a drastic increase in these mechanisms across the african continent (daigle, 2021). recent legal frameworks such as the south african protection of personal information act (popi act) (2013) recognise the risks of processing the personal information of minors but also allow exceptions for research-specific purposes14. nevertheless, the scientific community, specifically health researchers, is uncertain about which mechanisms to implement to comply with new regulatory frameworks. it is still unclear how to achieve compliance within their research projects and what concrete changes are needed to research governance structures and processes. stimulated by these recent discussions and the necessary adjustments for conducting research with children and adolescents in sub-saharan africa, this paper outlines data management instruments and practices iteratively developed, tested and refined in a research consortium jointly located in south africa and the united kingdom, with research partnerships in zambia, malawi, nigeria, lesotho, tanzania, and kenya. this paper's objective is to: a) discuss data management instruments aimed at safeguarding the protection of children and adolescents' personal information; b) outline practical mechanisms employed by our team to comply with regulatory frameworks; and c) stimulate a discussion on how to improve the protection of sensitive personal information within research contexts in sub-saharan africa. the vision is not limited to, but primarily focuses on, data collected in https://doi.org/10.29173/iq1044 3/15 hertzog, lucas; chen-charles, jenny; wittesaele, camille; de graaf, kristen; titus, raylene; kelly, jane; langwenya, nontokozo; baerecke, lauren; banougnin, boladé hamed; saal, wylene; southall, john; cluver, lucie; and toska, elona (2023) data management instruments to protect the personal information of children and adolescents in sub-saharan africa, iassist quarterly 47(2), pp. 1-15. doi: https://doi.org/10.29173/iq1044 south africa. it aims to generate scientifically rigorous evidence to influence policy and programmes to support children and adolescents to reach their full potential. the instruments and practices outlined here were constructed to safeguard the personal information of children and adolescents in longitudinal social science studies and randomised trials, including a large cohort of adolescents living with hiv (toska et al., 2016), a cohort of adolescent mothers and their children, and several studies of parenting programs in lowand middle-income countries (lmics) (lachman et al., 2016; cluver et al., 2018). structuring teams and building capacity sustainable data management starts by designing appropriate team structures that vary according to the research project's goals and resources. in the context of transdisciplinary research across different countries, we found that the minimum staff required is one information officer per institution, working in collaboration with research staff and the project's principal investigator. information officers will be able to respond to daily tasks, such as curating collections, securely storing data, digitising paper-based documents, managing staff access credentials and flagging potential data transfer risks, especially during data collection periods. data transfer points generally represent the weakest link in any data security chain. research staff and the principal investigators oversee these tasks and engage teams in capacity-building opportunities to enhance collective comprehension of strategies that safeguard personal information collected from research participants. in reflecting on our field research practices, we identified an opportunity emerging from new data protection regulations such as popi act. they trigger the need for 'capacity-sharing spaces' for researchers and potential research participants. these spaces stimulate researchers to develop better data protection and data management skills through interlinked strategies (figure 1). when possible, for research participants, capacity-sharing spaces about their rights enable them to engage in the making of research and empower them as individuals for future interactions with researchers and research projects that may invite them as participants. it can also be seen as essential to any genuinely informed consent process. https://doi.org/10.29173/iq1044 4/15 hertzog, lucas; chen-charles, jenny; wittesaele, camille; de graaf, kristen; titus, raylene; kelly, jane; langwenya, nontokozo; baerecke, lauren; banougnin, boladé hamed; saal, wylene; southall, john; cluver, lucie; and toska, elona (2023) data management instruments to protect the personal information of children and adolescents in sub-saharan africa, iassist quarterly 47(2), pp. 1-15. doi: https://doi.org/10.29173/iq1044 figure 1. interlinked research governance strategies and data management instruments to safeguard the personal information of children and adolescents in research projects. implementing regulatory frameworks such as gdpr and popi act invites research teams to engage researchers to simultaneously build understanding and move towards alignment and compliance across ongoing research projects. this also creates an opportunity to streamline processes and delivers efficiencies. the first step is establishing internal forums with experts promoting capacitybuilding opportunities, which establishes and maintains agreed working practices and data management instruments. this applies to all levels of research staff but is especially useful for early career researchers and new team members. these sessions should cover various topics, such as special data categories, specific research provisions, and data-sharing platforms. there is also a need to put in place processes that create documents demonstrating data protection compliance for research ethics committees and information regulators. these training sessions have proven to be instrumental in creating an understanding amongst our research team about the implications of new legislation. this, in turn, triggered useful internal audits of existing research governance documents and protocols. secondly, these forums help strengthen collaborations and essential staff selection to establish a data management team within the research projects. in building an understanding of how to unpack regulatory frameworks and adapt existing research governance practices, teams are invited to revise their structure and assess potential adaptations to ensure alignment and compliance with new legislation. these forums are also an opportunity for study leads (or principal investigators) to complete risk assessments to identify the type and format of personal information collected within each study and classify them into different levels (adams et al., 2021). https://doi.org/10.29173/iq1044 5/15 hertzog, lucas; chen-charles, jenny; wittesaele, camille; de graaf, kristen; titus, raylene; kelly, jane; langwenya, nontokozo; baerecke, lauren; banougnin, boladé hamed; saal, wylene; southall, john; cluver, lucie; and toska, elona (2023) data management instruments to protect the personal information of children and adolescents in sub-saharan africa, iassist quarterly 47(2), pp. 1-15. doi: https://doi.org/10.29173/iq1044 finally, the data management team may identify the training needs of data operators (i.e., a person who processes personal information for a responsible party in terms of a contract or mandate without coming under the direct authority of that party) and field workers. forums with experts allow teams to implement data security enhancement processes. for example, in our research context, these forums triggered a change to direct data uploading onto protected servers using end-to-end encryption instead of storing data on password-protected laptops for later transfer, as practised previously. in addition, covid-19 safety requirements and possible remote data collection were a further incentive for researchers to reflect and adapt, particularly in resource-limited settings. insights from the forums underlined that there must be a clear understanding of how to a) safely provide participants with a copy of the informed consent forms (icf); b) maintain and track the process of consent (verbally, text messages, and voice recordings) for each phase of data collection; and c) ensure participants have contact details for the research team, ethics committees and responsible parties. the forums also provided a space for ongoing discussion of how to ensure the confidentiality of sharing icfs, given the high rates of mobile devices being shared in resource-limited settings (he et al., 2020). based on our experience, we would argue that compliance and alignment with data protection regulatory frameworks rely on a sustainable data culture with an adequate team structure that can continuously engage in capacity-building activities. in addition, compliance can only be pursued if the additional administrative tasks engage fieldworkers meaningfully with continuous support in transition moments, such as hiring new staff or conducting data collection, especially with new regulations and national data protection legislation in sub-saharan africa. enhancing data information in ethical applications ethical clearance processes from existing research ethics committees (recs) increasingly request detailed information from researchers about their handling of personal information in the context of the increased digitalisation of society. in some cases, recs will make favourable ethical approvals contingent on the opinion of an information regulator, such as an institution's information officer, who may or may not be involved in recs directly. acquiring clearance for processing personal information from such recognised authorities should be equally important (in terms of timelines, resourcing, and compliance) as receiving ethical approval from existing recs. each must be held in continuous review and monitored simultaneously. in research that includes sensitive data from children and adolescents living under extreme adversities, many in vulnerable contexts, ethical responsibility in research projects is a foremost concern. given the emergence of previously mentioned personal information data protection regulations across sub-saharan africa, this crucial stage in every research project has a new layer of complexity. ethical parameters may be set out in research ethics applications enabling researchers to ensure personal information provided by data subjects is protected following principles of ethical research and protection of personal information. in south africa, for example, section 34 of the popi act prohibits processing minors' personal information unless provisions of section 35 are applicable. to https://doi.org/10.29173/iq1044 6/15 hertzog, lucas; chen-charles, jenny; wittesaele, camille; de graaf, kristen; titus, raylene; kelly, jane; langwenya, nontokozo; baerecke, lauren; banougnin, boladé hamed; saal, wylene; southall, john; cluver, lucie; and toska, elona (2023) data management instruments to protect the personal information of children and adolescents in sub-saharan africa, iassist quarterly 47(2), pp. 1-15. doi: https://doi.org/10.29173/iq1044 meet these special provisions, ethical approvals from recs are crucial to confirm whether the research and processing of personal information are appropriate and for the public interest. relying on more than thirteen years of fieldwork experience in multiple south african provinces (cluver et al., 2015) and experience from other studies with vulnerable populations in similar contexts (bostock, 2002; hensen et al., 2021), it is possible to outline nine considerations which improve ethical parameters in applications in a digital era: 1. submission of ethical applications to suitable recs, including a clear data flow with all the proposed steps of the data lifecycle. 2. negotiation of consent from each individual involved at each stage of data collection. a. consent must be provided by a competent person and, where the data subject is a minor, followed by participant assent; b. consent forms must clearly identify institutions and lead investigators responsible for data management. 3. adhere to the data subject's confidentiality principles, in line with the icf. 4. where possible, de-identify (anonymise) datasets at the earliest opportunity and minimise the risk of re-identification. 5. limit access to personally identifiable information within the research team on a 'need-toknow' basis. 6. put in place appropriate retention of personal information records for historical, statistical, and research purposes, with sufficient safeguards against the records being used for purposes not approved by recs. 7. stress that responsible parties should ensure that personal information is always secure throughout data collection, processing, migration, storage, sharing, archiving, and dissemination. 8. adopt a mechanism for personal information to be withdrawn at request by the data subject and competent person. 9. ensure trained interviewers interview data subjects in private locations to maximise confidentiality. negotiating disclosure through informed consent obtaining informed and voluntary consent from research participants is central to conducting ethical research (strode & slack, 2012), and negotiation for disclosure during interviews must be explored during capacity-building activities for researchers. this negotiation implies, in basic terms, transparently informing data subjects about how their data will be handled and the risks of participating in the research before requesting voluntary consent for participation. the complete understanding of participation will be consolidated through voluntary consent, a mutual agreement where researchers engage with data subjects. in the context of research with minors, consent is also a dialogue process with caregivers regarding their child's rights. this renders informed consent a fundamental instrument for informing data subjects about the risks and benefits of providing personal information. with the rollout of data protection regulations, researchers should also use this process to inform children and their caregivers about their personal information and privacy rights (strode & slack, 2013). https://doi.org/10.29173/iq1044 7/15 hertzog, lucas; chen-charles, jenny; wittesaele, camille; de graaf, kristen; titus, raylene; kelly, jane; langwenya, nontokozo; baerecke, lauren; banougnin, boladé hamed; saal, wylene; southall, john; cluver, lucie; and toska, elona (2023) data management instruments to protect the personal information of children and adolescents in sub-saharan africa, iassist quarterly 47(2), pp. 1-15. doi: https://doi.org/10.29173/iq1044 in sub-saharan africa, where children, adolescents, and caregivers may have low literacy rates (unesco, 2013), fieldworkers and researchers must be cognisant of the unbalanced power relations in place, adopting strategies to mitigate them. the use of accessible languages in information sheets and consent forms is crucial, and reading documents aloud to the data subjects in their chosen language should be the standard practice. this critical moment in data collection is an entry point for fieldworkers to give participants ample opportunities to ask questions and decide about participation. it benefits all parties to consider the impact this will have on the management and analysis of research data as part of the project. if a participant withdraws consent or requests for their data to be removed, for example, this is an agreed right. according to data protection regulatory frameworks such as popia and gdpr, researchers should ensure that data subjects are informed about: a) which data protection regulations govern the handling of personal information in the research, b) the nature of data that will be collected, c) how it will be processed, d) where it will be stored, e) what security measures will be in place to protect the data, e) who will have access to personal information, f) how long their data will be retained, and g) how they may request for their data to be updated or removed. particularly in longitudinal studies, researchers should ensure explicit permission from data subjects is obtained to re-contact them in the future. in one of our projects, hey baby (helping empower youth brought up in adversity with their babies and young children), a longitudinal cohort study of adolescent mothers and their children in the eastern cape, south africa, all data collected were entered into a password protected research electronic data capture (redcap) forms using tablets with log-in passwords. for each form, participant serial numbers, date of birth, and date of administering the questionnaire were used as unique identifiers to ensure that the participants completed the questionnaire anonymously. the study team was trained in how to use redcap, including all the security features. data management plan: building a roadmap for sustainable research the data management plan (dmp) is a core research governance instrument to safeguard data, particularly participants' personal data. it assists researchers in making and enacting decisions about data curation throughout the data lifecycle (i.e., collection, processing, analysing, preserving, sharing, and archiving). carefully planning and agreeing on how data will be managed at the outset, and keeping this in review, minimises data quality and security risks. ensuring data is of good quality and is available appropriately throughout the research process enhances the public benefit of research. in addition, dmps can be used as instruments for articulating the linkages between other data management instruments that may become opaque during large or long-term research projects. it may contain diagrams explaining, for instance, how the data-sharing agreement is linked to personal information impact assessments or plans for capacity-sharing activities to improve researchers' knowledge in certain areas. like other data management instruments, the dmp should be treated as a living document with continuous improvements throughout the research (whyte et al., 2022). significant events in the data lifecycle are opportunities to trigger modifications: a) when substantive changes in data needs arise, https://doi.org/10.29173/iq1044 8/15 hertzog, lucas; chen-charles, jenny; wittesaele, camille; de graaf, kristen; titus, raylene; kelly, jane; langwenya, nontokozo; baerecke, lauren; banougnin, boladé hamed; saal, wylene; southall, john; cluver, lucie; and toska, elona (2023) data management instruments to protect the personal information of children and adolescents in sub-saharan africa, iassist quarterly 47(2), pp. 1-15. doi: https://doi.org/10.29173/iq1044 b) at scheduled time points, and c) at key study stages (e.g., new waves of data collection, data migration, alignment to new national personal information regulatory frameworks). there are four key elements to consider when elaborating a dmp: a) project description, b) data storage, c) data security, and e) data sharing and reuse. each one of these elements triggers different questions that need to be addressed: 1. will you produce original (primary) data and/or use existing (secondary) data? 2. where and how will you acquire your data? 3. what types and data formats will you collect, and how will you describe them? 4. where will you store your data? 5. how will you organise and name your data files? 6. how large are your data? 7. what are the provisions for the secure storage and transfer of sensitive data? 8. is the data safely stored in repositories for long-term preservation? 9. how and where will you share your data during and after the study? 10. what are your plans for long-term data-sharing and preservation? 11. who will be able to access the data, under what conditions and for how long? dmps' requirements will vary from one institution to another. research institutions and funders may require dmps to be a comprehensive piece that brings together relevant information from all instruments used to protect data subjects' personal information. there are challenges in tailoring a dmp for different contexts to this extent. having it as a research governance roadmap may enhance its acceptability across various settings (e.g., institutional regulatory bodies for research, ethics committees, and legal departments) and thereby give it even further perceived value. all anonymised electronic data within the hey baby study were stored on a secure server housed at the university of cape town, south africa, which is located in a secure room with restricted access on a "need-to-access" basis. paper-based data, on the other hand, were stored in secure offices at the centre for social science research at the university of cape town. assessing personal information risks personal information impact assessments (piia) and data protection impact assessments (dpia) are crucial for alignment with regulatory frameworks such as popia and gdpr when processing special category data, such as data from minors with high-risk levels. these assessments support the researcher in considering the necessity and proportionality of processing personal information, including a risk assessment detailing the potential risks to data subjects. maintaining and updating a personal information data workflow, which depicts the flow of all personal information throughout the lifecycle of a project, can be a valuable and practical process for understanding where significant risks to the rights of data subjects may occur. these instruments support researchers in evaluating the level of risk involved in data processing. in research involving children and adolescents and vulnerable populations containing highly-sensitive data such as hiv status or other health information that could lead to stigma, such as hereditary conditions, and behaviour deemed deviant and non-normative, all personal data collected is classified https://doi.org/10.29173/iq1044 9/15 hertzog, lucas; chen-charles, jenny; wittesaele, camille; de graaf, kristen; titus, raylene; kelly, jane; langwenya, nontokozo; baerecke, lauren; banougnin, boladé hamed; saal, wylene; southall, john; cluver, lucie; and toska, elona (2023) data management instruments to protect the personal information of children and adolescents in sub-saharan africa, iassist quarterly 47(2), pp. 1-15. doi: https://doi.org/10.29173/iq1044 as high-risk. in practice, we have therefore found that the piia and dpia should address specific questions: 1. is there a risk of individual identification of data subjects that may lead to loss of privacy and unconsented identification? 2. could personal data being processed lead to stigmatisation, discrimination, bias, trauma, and legal prosecution? 3. is the personal information being processed in a different place than data collection? if that is the case, are the regulatory frameworks equivalent in the third country/place? 4. does the data contain unique identifiers, geographic information, or activity identifiers that may threaten data subjects' confidentiality? 5. what are the mechanisms in place for data de-identification? the fieldwork team within the hey baby study is facing increasing pressure to respond to the rising psychosocial referral needs of the study cohort. however, as a research project, we do not provide psychosocial services directly. still, we refer study participants to access services, including the south african depression and anxiety group, lifeline, etc. the team worked with a local ngo named masithete, which provides psychosocial support in the study community, to revise their existing referral protocol. for example, for an emergency referral such as suicide, the team will refer the participant directly to the hey baby masithethe line, where a counsellor will initiate the first counselling session with the participant. for a non-emergency referral such as emotional abuse, the team will refer the participant to masithethe or lifeline social workers. sharing data and collaborating collaboration & data sharing agreements (cdsa) between collaborating research partners across different institutions clarify terms of (co-)ownership and (joint) responsibility for research data. their success is built on mutual respect, cooperation, trust, and communication. it facilitates transparency and fairness in the treatment of data, ensuring compliance with legal and ethical obligations and that parties take appropriate technical and organisational measures to protect the security and confidentiality of data. in the context of unbalanced power relations between institutions and funders in the global south and north, particularly looking at the sub-saharan african context, the cdsa is a powerful instrument to stimulate change in decolonising scientific endeavours in light of discussions of restitution of archives to their place of origin (sarr & savoy, 2018). this agreement also clarifies roles and responsibilities for processing personal information. in south africa, for example, popia stipulates that south african research institutions may only transfer personal information if the third party is subject to a "law, binding corporate rules or binding agreement" (section 71 (1)(a)), which provide an adequate level of protection for the handling of the personal information. in the context of research projects with children and adolescents in sub-saharan africa, this form of agreement gives a basis for how data can be shared between the collaborating institutions and which legislations govern this. this is something beneficial given the transnational nature of our research. it demonstrates that research data may be transferred to a third party governed by other regulations as https://doi.org/10.29173/iq1044 10/15 hertzog, lucas; chen-charles, jenny; wittesaele, camille; de graaf, kristen; titus, raylene; kelly, jane; langwenya, nontokozo; baerecke, lauren; banougnin, boladé hamed; saal, wylene; southall, john; cluver, lucie; and toska, elona (2023) data management instruments to protect the personal information of children and adolescents in sub-saharan africa, iassist quarterly 47(2), pp. 1-15. doi: https://doi.org/10.29173/iq1044 long as they provide an appropriate level of protection for the personal information of data subjects and comply with the respective laws that govern their research. it is essential to highlight a few crucial elements of this instrument. first, it is a contractual document among all relevant parties. therefore, institutional representatives need to be involved. secondly, it is vital that the agreement clarifies the names of entities and differentiates between research collaboration and individual studies. for example, the contract may include a memorandum of understanding (mou) template that individual studies may use to enter a partnership with ngos for research. additional addendums about processes unique to individuals' studies may be included and governed by the agreement. finally, a data use undertaking (duu) template may be included to ensure that both parties use consistent terms for sharing data with external data users. data lifecycle awareness data protection regulations highlight the importance of implementing security measures to protect personal information throughout the data lifecycle (figure 2). their interpretation and implementation through a series of innovations in research practice greatly impacted our work. this must be adequately resourced and budgeted, and additional support may be required from services within research institutions and technical experts. figure 2. data workflow from collection to open access repository15. during restrictions imposed by the covid-19 pandemic, these innovations in practice continued, and researchers have transitioned from face-to-face to remote data collection, enhancing the digitalisation of data collection processes. this demands proportionate changes in security measures, such as using reliable open-sourced data collection platforms such as redcap and open data kit (odk). both have sufficient technical capabilities and functionalities for data collection processes with end-to-end encryption technologies. research institutions may have preferred software and hardware options and processes for evaluating the level of security of third-party services, devices, and tools, reducing the demand for research teams to resource this expertise internally. both security and information protection compliance should be assessed before use. low-tech techniques (e.g., concealing sensitive information by using unique identifiers or pseudonyms) may also be used to ensure the security of personal data. electronic https://doi.org/10.29173/iq1044 11/15 hertzog, lucas; chen-charles, jenny; wittesaele, camille; de graaf, kristen; titus, raylene; kelly, jane; langwenya, nontokozo; baerecke, lauren; banougnin, boladé hamed; saal, wylene; southall, john; cluver, lucie; and toska, elona (2023) data management instruments to protect the personal information of children and adolescents in sub-saharan africa, iassist quarterly 47(2), pp. 1-15. doi: https://doi.org/10.29173/iq1044 data captured should be submitted to servers daily and encrypted between the data collection device and the data servers. finally, specific protocols may be developed to support compliant data collection, processing, and storage governance, defining procedures and parameters for: a) data retention, b) de-identifying data, and c) managing access credentials to enable responsible parties to establish common and compliant standards. this data management work needs to be well-resourced, as it incorporates additional layers of research governance required by data protection regulations beyond sub-saharan africa. conclusion in this paper, we explored a series of best practices for safeguarding the personal information of children, adolescents, and young people (0-24 years old) (sawyer et al., 2018). this key age group represented nearly half of south africa's population in 2021, and adolescents are the fastest-growing global populational group in the world (unicef, 2019). the focus was to stimulate a broader conversation on how to improve the protection of children's and adolescents' sensitive personal information in sub-saharan africa and inform considerations that need to be addressed by research codes of conduct currently under development (adams et al., 2021). the aim was to discuss possible actions to ensure that new regulatory frameworks effectively build on existing data protection mechanisms for research projects at all research cycle stages. as a research group generating evidence that influences social and health policy and programming for young people in sub-saharan africa, our contribution draws on our work adhering to multiple transnational governance frameworks imposed by national legislation, such as data protection regulations, funders and academic institutions. this has involved the use of the data management instruments outlined. continuous reformulation and adaptations of these strategies are needed, such as including recurrent training in the curricula of our institutions to achieve the constant process of building a sustainable and ethical data culture in our research projects. https://doi.org/10.29173/iq1044 12/15 hertzog, lucas; chen-charles, jenny; wittesaele, camille; de graaf, kristen; titus, raylene; kelly, jane; langwenya, nontokozo; baerecke, lauren; banougnin, boladé hamed; saal, wylene; southall, john; cluver, lucie; and toska, elona (2023) data management instruments to protect the personal information of children and adolescents in sub-saharan africa, iassist quarterly 47(2), pp. 1-15. doi: https://doi.org/10.29173/iq1044 funding the international aids society through the cipher grant (155-hod; 625-tos); claude leon foundation [f08 559/c]; evidence for hiv prevention in southern africa (ehpsa), a uk aid programme managed by mott macdonald; janssen pharmaceutica n.v., part of the janssen pharmaceutical companies of johnson & johnson; the nuffield foundation, but the views expressed are those of the authors and not necessarily the foundation; oak foundation [r46194/aa001] and oak foundation/gcrf "accelerating violence prevention in africa" [ofil-20-057]; the regional inter-agency task team for children affected by aids eastern and southern africa (riatt-esa); the john fell fund [103/757; 161/033]; the philip leverhulme trust [plp-2014-095]; the university of oxford's esrc impact acceleration account [k1311-kea-004]; european research council (erc) under the european union's horizon 2020 research and innovation programme (grant agreement no 771468); research england [0005218]; co-funded by the medical research council (mrc) and the department of health social care (dhsc) through its national institutes of health research (nihr) [mr/r022372/1]; unicef eastern and southern africa office (unicef-esaro); ukri gcrf accelerating achievement for africa's adolescents (accelerate) hub (grant ref: es/s008101/1); the fogarty international center, national institute on mental health, national institutes of health under award number k43tw011434. the content is solely the responsibility of the authors and does not represent the official views of the national institutes of health; ucl's helpage funding; wellspring philanthropic fund. https://doi.org/10.29173/iq1044 13/15 hertzog, lucas; chen-charles, jenny; wittesaele, camille; de graaf, kristen; titus, raylene; kelly, jane; langwenya, nontokozo; baerecke, lauren; banougnin, boladé hamed; saal, wylene; southall, john; cluver, lucie; and toska, elona (2023) data management instruments to protect the personal information of children and adolescents in sub-saharan africa, iassist quarterly 47(2), pp. 1-15. doi: https://doi.org/10.29173/iq1044 references adams, r., adeleke, f., anderson, d., bawa, a., branson, n., christoffels, a., de vries, j., etheredge, h., et al. 2021. popia code of conduct for research. south african journal of science. 117(5/6). doi: https://doi.org/10.17159/sajs.2021/10933. bostock, l. 2002. "god, she's gonna report me": the ethics of child protection in poverty research. children & society. 16(4):273–283. doi: https://doi.org/10.1002/chi.712. caldwell, p.h., murphy, s.b., butow, p.n. & craig, j.c. 2004. clinical trials in children. the lancet. 364(9436):803–811. doi: https://doi.org/10.1016/s0140-6736(04)16942-0. cluver, l., boyes, m., bustamam, a., casale, m., henderson, k., kuo, k. & sello, l. 2015. the cost of action: large scale, longitudinal quantitative research with aids-affected children in south africa. ethical quandaries in social research. 41–56. cluver, l.d., meinck, f., steinert, j.i., shenderovich, y., doubt, j., herrero romero, r., lombard, c.j., redfern, a., et al. 2018. parenting for lifelong health: a pragmatic cluster randomised controlled trial of a non-commercialised parenting programme for adolescents and their families in south africa. bmj global health. 3(1):e000539. doi: https://doi.org/10.1136/bmjgh-2017-000539. daigle, b. 2021. data protection laws in africa: a panafrican survey and noted trends. journal of international commerce and economics. 27. he, e., hertzog, l., toska, e., carty, c. & cluver, l. 2020. mhealth entry points for hiv prevention and care among adolescents and young people in south africa. cssr working paper. (457). doi: https://doi.org/10.13140/rg.2.2.36698.98241. hensen, b., mackworth-young, c.r.s., simwinga, m., abdelmagid, n., banda, j., mavodza, c., doyle, a.m., bonell, c., et al. 2021. remote data collection for public health research in a covid-19 era: ethical implications, challenges and opportunities. health policy and planning. 36(3):360–368. doi: https://doi.org/10.1093/heapol/czaa158. hertzog, l., wittesaele, c., titus, r., chen, j.j., kelly, j., langwenya, n., baerecke, l. & toska, e. 2021. seven essential instruments for popia compliance in research involving children and adolescents in south africa. south african journal of science. 117(9/10). doi: https://doi.org/10.17159/sajs.2021/12290. joseph, p.d., craig, j.c. & caldwell, p.h.y. 2015. clinical trials in children: clinical trials in children. british journal of clinical pharmacology. 79(3):357–369. doi: https://doi.org/10.1111/bcp.12305. lachman, j.m., sherr, l.t., cluver, l., ward, c.l., hutchings, j. & gardner, f. 2016. integrating evidence and context to develop a parenting program for low-income families in south africa. journal of child and family studies. 25(7):2337–2352. doi: https://doi.org/10.1007/s10826-0160389-6. https://doi.org/10.29173/iq1044 https://doi.org/10.17159/sajs.2021/10933 https://doi.org/10.1002/chi.712 https://doi.org/10.1016/s0140-6736(04)16942-0 https://doi.org/10.1136/bmjgh-2017-000539 https://doi.org/10.13140/rg.2.2.36698.98241 https://doi.org/10.1093/heapol/czaa158 https://doi.org/10.17159/sajs.2021/12290 https://doi.org/10.1111/bcp.12305 https://doi.org/10.1007/s10826-016-0389-6 https://doi.org/10.1007/s10826-016-0389-6 14/15 hertzog, lucas; chen-charles, jenny; wittesaele, camille; de graaf, kristen; titus, raylene; kelly, jane; langwenya, nontokozo; baerecke, lauren; banougnin, boladé hamed; saal, wylene; southall, john; cluver, lucie; and toska, elona (2023) data management instruments to protect the personal information of children and adolescents in sub-saharan africa, iassist quarterly 47(2), pp. 1-15. doi: https://doi.org/10.29173/iq1044 ledford, h. 2021. should children get covid vaccines? what the science says. nature. 595(7869):638–639. doi: https://doi.org/10.1038/d41586-021-01898-9. mbembe, a. 2019. bodies as borders. from the european south. 4:5–18. protection of personal information act (popi act). 2013. available: https://popia.co.za/ [2022, march 18]. sarr, f. & savoy, b. 2018. the restitution of african cultural heritage: toward a new relational ethics. ministère de la culture ed. translated by drew burk. paris. sawyer, s.m., azzopardi, p.s., wickremarathne, d. & patton, g.c. 2018. the age of adolescence. the lancet child & adolescent health. 2(3):223–228. doi: https://doi.org/10.1016/s23524642(18)30022-1. shrestha, b. & dunn, l. 2020. the declaration of helsinki on medical research involving human subjects: a review of seventh revision. journal of nepal health research council. 17(4):548–552. doi: https://doi.org/10.33314/jnhrc.v17i4.1042. strode, a. & slack, c. 2012. selected ethical-legal norms in child and adolescent hiv prevention research: consent, confidentiality and mandatory reporting [revised]. durban, south africa: european and developing countries clinical trials partnership (edctp). toska, e., gittings, l., hodes, r., cluver, l.d., govender, k., chademana, k.e. & gutiérrez, v.e. 2016. resourcing resilience: social protection for hiv prevention amongst children and adolescents in eastern and southern africa. african journal of aids research. 15(2):123–140. doi: https://doi.org/10.2989/16085906.2016.1194299. unesco. 2013. adult and youth literacy: national, regional and global trends, 1985–2015. montreal: unesco institute for statistics. unicef. 2019. adolescent demographics. available: https://data.unicef.org/topic/adolescents/demographics/ [2022, february 02]. whyte, a., molloy, l., grootveld, m. & thorley, m. 2022. supporting data management planning. acme-fair issue. (4). doi: https://doi.org/10.5281/zenodo.6346747. williams, j. 2008. the declaration of helsinki and public health. bulletin of the world health organization. 86(8):650–651. doi: https://doi.org/10.2471/blt.08.050955. https://doi.org/10.29173/iq1044 https://doi.org/10.1038/d41586-021-01898-9 https://popia.co.za/ https://doi.org/10.1016/s2352-4642(18)30022-1 https://doi.org/10.1016/s2352-4642(18)30022-1 https://doi.org/10.33314/jnhrc.v17i4.1042 https://doi.org/10.2989/16085906.2016.1194299 https://data.unicef.org/topic/adolescents/demographics/ https://doi.org/10.5281/zenodo.6346747 https://doi.org/10.2471/blt.08.050955 15/15 hertzog, lucas; chen-charles, jenny; wittesaele, camille; de graaf, kristen; titus, raylene; kelly, jane; langwenya, nontokozo; baerecke, lauren; banougnin, boladé hamed; saal, wylene; southall, john; cluver, lucie; and toska, elona (2023) data management instruments to protect the personal information of children and adolescents in sub-saharan africa, iassist quarterly 47(2), pp. 1-15. doi: https://doi.org/10.29173/iq1044 endnotes 1 lucas hertzog is a health sociologist in the who collaborating centre for climate change and health impact assessment at curtin university, australia. this work was conducted within the scope of his previous affiliation with the university of cape town in the centre for social science research, south africa. lucas.hertzog@curtin.edu.au 2 jenny chen-charles is a research coordinator based at the university of oxford in the department for social policy and intervention. jenny.chen@spi.ox.ac.uk. 3 camille wittesaele is a researcher based at the department of infectious disease epidemiology, london school of hygiene & tropical medicine. camille.wittesaele1@lshtm.ac.uk. 4 kristen de graaf is a programme manager based at the university of oxford in the department for social policy and intervention. kristen.degraaf@spi.ox.ac.uk. 5 raylene titus is a junior information officer at the university of cape town in the centre for social science research. raylene.titus@uct.ac.za. 6 jane kelly is a research officer at the university of cape town in the centre for social science research. jane.kelly@uct.ac.za. 7 nontokozo langwenya is an epidemiologist pursuing a dphil in social intervention and policy evaluation at the university of oxford. nontokozo.langwenya@acceleratehub.org. 8 lauren baerecke is a researcher based at the university of cape town in the centre for social science research. lauren.baerecke@uct.ac.za. 9 boladé hamed banougnin is a demographer based at the university of cape town in the centre for social science research. 10 wylene saal is a research officer at the university of cape town in the centre for social science research. wylene.saal@uct.ac.za. 11 john southall is a bodleian data librarian and subject consultant for economics and sociology in the bodleian social science library at the university of oxford. john.southall@bodleian.ox.ac.uk. 12 lucie cluver is a professor of child and family social work at the department of social policy and intervention at the university of oxford and the department of psychiatry and mental health at the university of cape town. lucie.cluver@spi.ox.ac.uk. 13 elona toska is a researcher in adolescent health at the centre for social science research and an associate lecturer at the department of sociology, university of cape town. she is a co-principal investigator of the mzantsi wakho and hey baby studies and leads the uct team of the ukri gcrf accelerating achievement for africa’s adolescents hub. elona.toska@uct.ac.za. 14 popia’s section 35 alongside regulations on prior consent of a competent person for data collection, with specific provisions in section 11. 15 data sharing may also occur, for non-profit use, before data is moved to public repositories, in the case of other researchers approaching our team requesting to access datasets. https://doi.org/10.29173/iq1044 mailto:lucas.hertzog@curtin.edu.au mailto:jenny.chen@spi.ox.ac.uk mailto:camille.wittesaele1@lshtm.ac.uk mailto:kristen.degraaf@spi.ox.ac.uk mailto:raylene.titus@uct.ac.za mailto:jane.kelly@uct.ac.za mailto:nontokozo.langwenya@acceleratehub.org mailto:lauren.baerecke@uct.ac.za mailto:wylene.saal@uct.ac.za mailto:john.southall@bodleian.ox.ac.uk mailto:lucie.cluver@spi.ox.ac.uk mailto:elona.toska@uct.ac.za vol244 4 iassist quarterly winter 2000 toxics release inventory an environmental database by mary j. lee* data background/description this paper describes a publicly available database called the toxics release inventory (tri), which provides a valuable data source of environmental information. the tri data are machinereadable microdata at the industrial facility level and are provided free with unlimited access to the public. the data collection approach for this database is innovative because it uses the information collection provision as the regulatory instrument. following a chemical-release accident in bhopal, india, the u.s. congress passed the emergency planning and community right-to-know act (epcra) in 1986. under these provisions, manufacturing facilities with 10 or more employees in standard industrial classification (sic) codes 20 through 39 are required to publicly disclose their annual toxic release to air, water, and land as well as off-site transfers. the u.s. environmental protection agency (epa) compiled these annual reports into the tri database. because of the mandatory requirement of data provision and its inclusion of a public’s right-to-know provision, this database provides a reliable source of environmental performance information. since the initial data release in 1989, the number of reporting facilities and chemicals has been increased. seven industrial sectors have been added to the original reporting manufacturing industries. these include electric utilities, coal mining, metal mining, chemical wholesalers, petroleum bulk plants and terminals, solvent recovery and hazardous waste treatment, storage, and disposal. new data for 1998 were released in 2000 covering seven industrial sectors. there is a two-year time lag in the release of tri data. for example, the most recent data release in 2000 is for the reporting year 1998. the tri data have become a primary source of environmental performance information for a broad range of user groups. these include social scientists, environmentalists, government officials, investors, consulting firms, journalists, health professionals, etc. the international organizations have also recently joined the group of tri users. data dissemination/search engines data access technological progress has brought major changes in data management and data analysis system. the data storage and data access are much easier due to the faster speed of personal or mainframe computers, larger data storage capacities and the development of the internet. accessibility and media format options for the tri have also changed. the tri data are currently available on floppy diskette, cd-rom, or through the internet. the floppy diskettes contain the most frequently used data elements including each facility’s identification numbers, county, city, state, zip code, sic code, parent company name, chemical name and chemical registry number, total releases to the air, water, land, underground injection and off-site transfers. they also include the longitude and latitude of the facility and federal information processing standards (fips) code. the cd-rom edition is comprised of two cds. disc one has the tri data for 1987-1990. disc two contains data for 1991-1996. the basic features of the cd-rom include: user guide, combining searches using boolean operators, displaying records, exporting records in several formats, creating custom reports and calculating the data using kastat. the tri data on the internet is available for the time span of 1988-1998. tri explorer the tri explorer is a search engine that provides access to the tri data on the internet. the initial version of the tri explorer included onand off-site release data. the latest version added waste transfer and waste management data to the original toxic release data. this search engine allows you to identify facilities and their chemical release. data can be disseminated into release, waste transfer and waste quality reports. data can be grouped according to five criteria: facility, chemical, year or industry type and geographic area at the county, state or national level. waste management reports include recycling, energy recovery, treatment as well as off-site waste transfers. a trends report option is also available for the core chemicals. metadata on the web provides valuable information including detailed data element descriptions. data elements provide both facility identification information and chemical-specific information. iassist quarterly winter 2000 5 other tri sites on the internet on-line searching for the tri data is available using the web sites such as envirofacts: data warehouse and applications and the national library of medicine (nlm) toxnet system. the envirofacts warehouse includes multiple environmental databases that allow you to retrieve environmental information from several epa databases. spatial data are available using the maps on demand applications. toxnet (toxicology data network) is a cluster of databases on toxicology, hazardous chemicals, and other related environmental or public health areas. international development systems international tri-like system there has been a growing global movement for information database on toxic release. international tri-like system is called as pollutant release and transfer registers (prtrs). international organizations started some initiatives to implement the development of prtrs. the prtrs in the world include: canada: national pollutant release inventory (npri), united kingdom: pollutant inventory (pi), mexico: registro de emisiones y transferencia de contaminantes (retc), australia: national pollutant inventory (npi) and czech republic: pollutant release and transfer register (prtr). among three north american prtrs, canada has prtr data starting 1993. mexico collects prtr data from industrial facilities on a voluntary basis. the commission for environmental cooperation (cec), an environmental organization created by the north american free trade association (nafta), compiles the data and publishes an annual report on the north american prtrs. asian countries also participated in this international movement. japan hosted the most recent international conference on prtrs: national and global responsibility in september 1998. indonesia developed similar public disclosure program called program for pollution control, evaluation and rating (proper). indonesia’s national pollution control agency initiated this pollution control program to evaluate the environmental performance of indonesian factories. the philippines followed indonesia’s footsteps. the philippines’ department of environment and natural resources recently started a public disclosure program called ecowatch modeled on indonesia’s proper program. international organization prtr sites the international movement on prtrs stems from the 1992 earth summit, also called the united nations conference on environment and development (unced). several international organizations now have their own home page for the development of prtrs: organization for economic co-operation and development (oecd) prtr homepage, united nations environmental programme (unep) prtr homepage, unitar prtr homepage and world bank prtr homepage. further data usability growing international interest in toxic release information presents the opportunity to combine the various databases and compare each country’s toxic releases and waste management activities. it also provides the possibility of developing an international toxic release database in the future. this would be in addition to the existing data archives and would be beneficial both to academic researchers and government policy makers. wide use of this database will also provide industries with the incentive to improve existing pollution abatement technology. in addition to the unified international database, it is suggested that a comprehensive database may be developed using existing databases from other areas. one such area is public health, where human health risks could be measured using the tri and other health information databases. other health information databases include the hazard information on toxic chemicals, integrated risk information system (iris), and toxfaqs™ by the agency for toxic substances and disease registry (atsdr). another areas are finance and economy, where the financial effects of environmental information can be assessed using the financial databases. they include the center for research in security prices (crsp) database and the standard & poor’s compustat database. conclusion the toxics release inventory database is a valuable resource as a database for researchers in the area of environmental studies, health and business. in addition, it provides policy makers and the public with a reliable source of information on toxic emissions and serves as a regulatory tool for the management of industrial pollutants. in addition, the availability of tri data along with similar efforts in other countries provides an incentive for cooperative efforts in international reporting and analysis of toxic release data. * mary j. lee, 945 flanner hall, laboratory for social research, university of notre dame, notre dame, in 46556 u.s.a. tel (219) 631-4521 e-mail: lee.82@nd.edu mailto:lee.82@nd.edu 1/16 orlowska, daria; fallaw, colleen; feng, yali; garza, livia; hetrick, ashley; imker, heidi & luong, hoa (2021) better data management, one nudge at a time, iassist quarterly 45(2), pp. 1-16. doi: https://doi.org/10.29173/iq1010 better data management, one nudge at a time daria orlowska1, colleen fallaw2, yali feng3, livia garza4, ashley hetrick5, heidi imker6, hoa luong7 abstract how do you help people improve their data management skills? for our team at the university of illinois at urbana-champaign, we decided the answer was "one nudge at a time”. a study conducted by wiley and mischo (2016) found that illinois researchers are aware of data services available but under-utilize them. many researchers do not consider data management as a concern distinct from researching and producing scholarly work products. in 2017, the rds piloted the data nudge – a monthly, opt-in email service to “nudge” illinois researchers toward good data management practices, and towards utilizing data services on campus. the aim of the data nudge was to address the gap between knowing about a service and using it by highlighting best practices and campus resources. the topics covered in the data nudge center around data. some topics are applicable to everyone, such as data back-up, documentation, and file naming conventions. other topics are specific to illinois, like storage options, events, and conferences. after four years, the data nudge has accumulated over 400 subscribers through word-of-mouth, marketing channels on campus and inclusion in subject liaisons' instructional workshops. it receives stable open rates averaging at 52% (compared to 19.44% average industry rate for higher education*) and many compliments from subscribers. we expect the data nudge to continue supplementing workshops and training as an effective means of communication to reach researchers on our campus. in the spirit of re-use, we are in the process of archiving the data nudge topics in a reusable format, readily adaptable by other institutions.  data nudge link: https://go.illinois.edu/past_nudges keywords data nudge, data management, research data service, subject librarians, community engagement, service promotion, research data introduction as an r1 research university, the university of illinois at urbana-champaign boasts over 150 centers, labs, and institutes, in addition to active research programs of more than 2,700 faculty (office of the vice chancellor for research & innovation, 2019). in fy18, the university’s research and development expenditures amounted to $652 million, with $353 million in federal research expenditures (office of the vice chancellor for research & innovation, n.d.). the release of the 2013 office of science and technology policy memorandum calling for “increasing access to the results of federally funded scientific research” bolstered the recommendation to establish specific research data services and infrastructure in illinois’ 2013-2016 strategic plan, under the goal of “foster[ing] scholarship, discovery and innovation.” the strategic plan called for the creation of a “research data service and accompanying research education initiative in the curation, use and dissemination of large amounts https://doi.org/10.29173/iq1010 https://go.illinois.edu/past_nudges 2/16 orlowska, daria; fallaw, colleen; feng, yali; garza, livia; hetrick, ashley; imker, heidi & luong, hoa (2021) better data management, one nudge at a time, iassist quarterly 45(2), pp. 1-16. doi: https://doi.org/10.29173/iq1010 of data” (strategic working group, 2013). since its inception in 2014 with four staff members, the illinois research data service (rds) has aided researchers by offering a digital repository (the illinois data bank), curating deposited datasets to ensure completeness and accessibility, and reviewing data management plans for federal grant applications. around the same time, fearon et al. (2013) surveyed 73 arl member libraries to gain a sense of how research libraries were meeting data service demands driven by an increased emphasis on open data and dmp requirements. the survey indicated that libraries were still in the early stages of development. although data education was not covered as its own category, most education occurred through website resources and consultations (n = 48 for dmps and other research data management) (fearon et al., 2013, pp. 13-14; fearon et al., 2013, p. 42). less than half of respondents reported running workshops for dmps (n=33) (fearon et al., 2013, p. 13) and other research data management education (n=35) (fearon et al., 2013, p. 42). service challenges touched on themes such as campuswide collaboration and engagement and marketing (fearon et al., 2013, pp. 82-90). similarly, a 2014 survey of eighty-one uk libraries by cox and pinefield found that uk academic libraries offered a limited research data services with only twenty-four libraries hosting web resources and twenty-eight offering consultations (cox and pinefield, 2014, p. 18), but many saw data services as a priority, with medium-term goals of consultations and training (cox and pinefield, 2014, p. 22). however, offering these services alone did not completely meet the need to help researchers take advantage of resources to support better data management practices. a study conducted by wiley and mischo (2016) found that illinois researchers were aware of campus data preservation and management services, but rarely used these resources. they further found that researchers did not consider data management as a separate element of scholarly communication workflow, highlighting an area that could benefit from improved outreach. in 2017, the rds piloted the data nudge, a monthly opt-in email service, to prompt and offer support to illinois researchers about data services on campus. the aim of the data nudge was to address the gap between awareness and practice by promoting best practices and campus resources in a bitesized monthly subscription service. while many recent surveys have identified the areas in which libraries are providing data support services, to our knowledge, they have not explicitly gathered information on how data education is being implemented. in a recent article, ohaji, chawner and young (2019) argue that the role of a data librarian encompasses skill building, developing policy, and supporting data infrastructure, with a specific focus on fostering new connections across campus. this suggests that hiring a data librarian as a response to data needs moves past a traditional liaison role, and encourages a new approach to engaging with the research community. a scan of websites from other universities in the data curation network (dcn), a group of data curation collaborators who offer data services similar to the university of illinois, revealed that all universities within the dcn offer a website and consultation services, and most offer library guides and workshops as their primary source of data education (table 1). however, the scan also revealed a plethora of new forms of data education, suggesting that universities are spearheading a more experimental approach to building awareness and competency within their research community. https://doi.org/10.29173/iq1010 3/16 orlowska, daria; fallaw, colleen; feng, yali; garza, livia; hetrick, ashley; imker, heidi & luong, hoa (2021) better data management, one nudge at a time, iassist quarterly 45(2), pp. 1-16. doi: https://doi.org/10.29173/iq1010 universities in the dcn run community workshops on demand, upload online tutorials, and engage members of their research community through social media. in addition to continuing to offer more traditional data education like workshops, they have also ventured into listservs and newsletters. of the five universities offering this service, most use it as a communication tool: washington university offers “regular updates about events and news”, cornell university offers “low traffic mailing list for updates, upcoming events, and news”, john hopkins university offers “the latest data management and sharing information, upcoming data services events and training sessions”, and duke university offers “a biweekly newsletter for upcoming events, exhibits, resources, services, and other library news”. in contrast, the university of illinois has developed a dedicated data listserv offering resources and highlighting best practices: the data nudge. the purpose of this paper is to provide insight into the success of turning a listserv into a tool that not only nudges the research community into action but also connects them with resources across campus. we hope that through describing our processes and publishing our materials openly online, we can offer other university libraries a framework and ready-made content to connect with their own communities. workshop series workshop on demand online tutorial consultation library guide landing page listserv/ newsletter blog posting twitter handle university of illinois x x x x x x university of michigan x x x x x x university of minnesota x x x x x washington university x x x x x x x x cornell university x x x x x x x john hopkins x x x x x x x penn state x x x x x new york university x x x x x x x duke university x x x x x x x x x https://doi.org/10.29173/iq1010 https://www.library.illinois.edu/rds/ https://www.library.illinois.edu/rds/ https://urldefense.proofpoint.com/v2/url?u=https-3a__www.lib.umich.edu_research-2ddata-2dservices&d=dwmfag&c=ociemewdeq_anlsp4ff3gfqsn-e3mlr2t9jcddfozag&r=6thda5v0-66yimmydiwsbtpo4yb5aj6eikq8x9mvzpy&m=h62unftg-7iy-ixgez3oxl3rcux35agx6ssipxctr1u&s=weqag7fnh6g0h62fh5ub9feo-lanjddincgrjsfcu3q&e= https://urldefense.proofpoint.com/v2/url?u=https-3a__www.lib.umich.edu_research-2ddata-2dservices&d=dwmfag&c=ociemewdeq_anlsp4ff3gfqsn-e3mlr2t9jcddfozag&r=6thda5v0-66yimmydiwsbtpo4yb5aj6eikq8x9mvzpy&m=h62unftg-7iy-ixgez3oxl3rcux35agx6ssipxctr1u&s=weqag7fnh6g0h62fh5ub9feo-lanjddincgrjsfcu3q&e= https://www.lib.umn.edu/datamanagement https://www.lib.umn.edu/datamanagement https://library.wustl.edu/services/data/ https://library.wustl.edu/services/data/ https://data.research.cornell.edu/ https://data.research.cornell.edu/ https://dataservices.library.jhu.edu/ https://libraries.psu.edu/research/research-data-services http://library.nyu.edu/departments/data-services/ http://library.nyu.edu/departments/data-services/ https://library.duke.edu/data/ 4/16 orlowska, daria; fallaw, colleen; feng, yali; garza, livia; hetrick, ashley; imker, heidi & luong, hoa (2021) better data management, one nudge at a time, iassist quarterly 45(2), pp. 1-16. doi: https://doi.org/10.29173/iq1010 structure of the data nudge 1. topic selection and development inspiration for data nudge topics comes from a variety of sources, including campus events and consultations with researchers in which it is clear that being aware of information made a positive difference to the research process. increasingly, we receive requests from campus partners supporting research with specific messages they hope to share with the illinois community through the data nudge. while these can be opportunities to share new information about emerging and changing services available to researchers, they can also be challenging as the content must be reworked to match the format, tone, and purpose of the data nudge. however, the seed for a data nudge most often sprouts from personal experiences, when someone on the team learns something in the course of their work and thinks, “i wish i had known that sooner!” for example, while evaluating file types in the illinois data bank and encountering compressed and archive file formats, various team members were surprised by different aspects of the distinctions, characteristics, and options available. thus, the august 2018 data nudge focused on four common compressed file formats widely used in research and offered readers facts about files that do not compress well (figure 1). figure 1. a section of the august 2018 data nudge, “common zipped file formats used in research” regardless of the source of inspiration, the idea for a nudge begins when a proposer pitches a topic to the team. initial acceptance of a topic requires that it can be represented in a concise and visual manner and has dedicated support on campus (only for topics that specific to illinois), demonstrated https://doi.org/10.29173/iq1010 5/16 orlowska, daria; fallaw, colleen; feng, yali; garza, livia; hetrick, ashley; imker, heidi & luong, hoa (2021) better data management, one nudge at a time, iassist quarterly 45(2), pp. 1-16. doi: https://doi.org/10.29173/iq1010 through the inclusion of who to contact with questions. once the pitch convinces the team, accepted topics are scheduled depending on campus events or in an order that allows topics to build on each other. after the brainstorming process, a member of the team commences work on an initial draft of the data nudge, aiming to work ahead 2-4 months before the release date. this work consists of research, collaborating on how to phrase and structure the content, and creating visual elements when appropriate. all draft materials are transferred into a template, where they are checked for size and adjusted using html. any images are uploaded and linked from u of i box, and alt-text is included for accessibility. at least one month in advance of release, the draft data nudge is presented at a weekly rds meeting to receive feedback from other team members. team members consist of partners within and outside of the university of illinois library, including subject liaisons, functional specialists, and it professionals. this mix of diverse perspectives is crucial to crafting messages that are free from jargon and speak across disciplinary or professional boundaries. drafts are displayed on a large monitor and modified live based on group feedback for clarity, brevity and appeal to the target audience. in the case where more research is needed or larger scale rework is required, the group provides suggestions during the meeting, and the team makes adjustments afterwards. if the complexity of the topic still remains high, then it is postponed, and another is chosen that better fits the guiding principles. after several reviews, the content is finalized and scheduled to be sent out on the last tuesday of the month. while we did not start out with a fully-articulated set of guiding principles, a de facto sense of a data nudge style has formed from a consensus of i-know-it-when-i-see-it responses as the team shapes a draft into published form. we are guided substantially by empathy for our busy researchers, and thus we shape content around anticipations of how it will be experienced by our subscribers. we aim to offer pragmatic, specific information that a researcher could use immediately, or describe a practice that could improve the body of research data available in the world. we evaluate drafts in terms of careful accuracy for scan-and-skim and every-word readers, striving to offer a valuable take-away from a glance. we close by highlighting available resources, linking to relevant materials, and providing contact information for those who could answer further questions. the january 2020 data nudge on “illinois box features” is a particularly good example of reworking content based upon empathetic reading. originally, the data nudge content was adapted from text provided by a campus partner. the text provided a high-level overview of useful features, security best practices for data stored in box, and handy training videos to help users get started. because our primary audience is researchers and we want to follow the aforementioned guiding principles, the rds meeting members decided to emphasize and demonstrate features we know to be of high interest to researchers. these includes team folders for collaboration, favorites for quickly locating frequently used folders and files, and security concerns (figure 2). https://doi.org/10.29173/iq1010 6/16 orlowska, daria; fallaw, colleen; feng, yali; garza, livia; hetrick, ashley; imker, heidi & luong, hoa (2021) better data management, one nudge at a time, iassist quarterly 45(2), pp. 1-16. doi: https://doi.org/10.29173/iq1010 figure 2. final version of the january 2020 data nudge, “illinois box features”. 2. data nudge platform during its first year (between 2017-01 and 2018-01), the data nudge was delivered through mailchimp (mailchimp.com), a commercial marketing automation platform and email marketing service. subscribers could view the email in their choice of browser, in html format. beginning in 2018, the university’s technology services implemented dmarc (domain-based message authentication, reporting and conformance) to authenticate emails from third party email services to reduce phishing and increase email security, and mailchimp was among several email services technology services identified as not being configured to work with the new dmarc. to continue supporting our researchers, we switched to the email+ platform in webtools, a resource provided by university public affairs (figure 3). in either case, what was most important was a user-friendly platform for both our subscribers and for us as the content creators and providers. in particular, it was essential that we had access to analytics so we could evaluate usage. https://doi.org/10.29173/iq1010 7/16 orlowska, daria; fallaw, colleen; feng, yali; garza, livia; hetrick, ashley; imker, heidi & luong, hoa (2021) better data management, one nudge at a time, iassist quarterly 45(2), pp. 1-16. doi: https://doi.org/10.29173/iq1010 figure 3. to conform to campus security requirements, content is developed in email+ platform, powered by an illinois service called webtools. 3. marketing channels to bring the data nudge to the campus community, we send announcements to two different ebulletins containing general interest campus information and opportunities for faculty/staff or graduate students at the beginning of each semester. we have created data nudge cards not only to pass out at university events for it professionals, graduate students, and faculty, but to make them available as take-away promotional literature in the library’s digital scholarship center. this card template has been transformed into a digital sign that is displayed regularly at the main library and library branches with heavy rds users such as our campus engineering and life sciences libraries. we also advertise the data nudge on the rds website, with a link to past nudges. another avenue is subject librarian promotion, such as an announcement in a monthly departmental newsletter or through information sheets created for undergraduate instructional sessions. other advertisements include electronic slides displayed on departmental monitors or by posting physical data nudge promotion flyers with a qr code, making it convenient for users to subscribe through their smartphones (figure 4). https://doi.org/10.29173/iq1010 8/16 orlowska, daria; fallaw, colleen; feng, yali; garza, livia; hetrick, ashley; imker, heidi & luong, hoa (2021) better data management, one nudge at a time, iassist quarterly 45(2), pp. 1-16. doi: https://doi.org/10.29173/iq1010 figure 4. data nudge promotional flyer as the data nudge has grown, we have been pleasantly surprised by strong word-of-mouth advertising, with users unexpectedly championing the nudge without prompt at community events. in many cases, internal subscribers moving on from the university of illinois continue to subscribe through a personal email or by resubscribing with a new address. other external subscribers include those who have never been affiliated with the university of illinois, and we can only speculate whether they stumbled on data nudge organically or through shout-outs from external sources such as webinars hosted by current subscribers. 4. usage figures thanks to all of our marketing channels, the number of data nudge subscribers has been steadily increasing. from humble “library friend” opt-ins in 2017, the data nudge has accumulated over 400 subscribers as of december 2020, with 82% university of illinois subscribers and 18% external subscribers, including individuals from government organizations and other academic institutions (figure 5, 6). reports are recorded a week after the release of the emails, with an average of 52% for email open rate and 11.4% for link click rates between january 2017 through december 2020. in comparison, the industry open and click rate (constant contact, 2021) for higher education is around 19.55% and 7.94%, respectively (figure 8). https://doi.org/10.29173/iq1010 9/16 orlowska, daria; fallaw, colleen; feng, yali; garza, livia; hetrick, ashley; imker, heidi & luong, hoa (2021) better data management, one nudge at a time, iassist quarterly 45(2), pp. 1-16. doi: https://doi.org/10.29173/iq1010 another avenue is subject librarian promotion, such as an announcement in a monthly departmental newsletter or through information sheets created for undergraduate instructional sessions. other advertisements include electronic slides displayed on departmental monitors or by posting physical data nudge promotion flyers with a qr code, making it convenient for users to subscribe through their smartphones (figure 4). data nudge, data management, research data service, subject librarians, community engagement, service promotion, research data figure 5. number of subscribers between inception in january 2017 through december 2020. figure 6. breakdown of data nudge subscribers. 0 50 100 150 200 250 300 350 400 450 500 n u m b . s u b sc ri b er s by month data nudge subscribers over time https://doi.org/10.29173/iq1010 10/16 orlowska, daria; fallaw, colleen; feng, yali; garza, livia; hetrick, ashley; imker, heidi & luong, hoa (2021) better data management, one nudge at a time, iassist quarterly 45(2), pp. 1-16. doi: https://doi.org/10.29173/iq1010 figure 7. number of subscribers break down by college. note: “university operations” refers to subscribers who administer or support services, like technology services / research it, ncsa, igb, sibel design center, ovcri, graduate college, etc. “external” refers to subscribers who are located outside of illinois. “unknown” refers to subscribers who we could not locate their department. 0 10 20 30 40 50 60 70 80 90 100 college of medicine institute for sustainability, energy, and environment college of media unknown applied health sciences veterinary medicine fine & applied arts school of social work education college of business school of information sciences prairie research institute university library engineering college of agricultural, consumer and environmental… external university operations college of liberal arts & sciences data nudge subscribers distribution based on college https://doi.org/10.29173/iq1010 11/16 orlowska, daria; fallaw, colleen; feng, yali; garza, livia; hetrick, ashley; imker, heidi & luong, hoa (2021) better data management, one nudge at a time, iassist quarterly 45(2), pp. 1-16. doi: https://doi.org/10.29173/iq1010 figure 8. open and click rate of each data nudge topic compared with the educational industry average rate, according to constant contact (2021). data nudge topic open rate 2017-02 file organization 71.90% 2017-01 backup 65.20% 2018-10 scary data stories from il community 64.09% 2017-06 box tips 63.70% 2019-04 data cleaning 62.59% table 2. top five most popular nudges by open rate. 0.00% 10.00% 20.00% 30.00% 40.00% 50.00% 60.00% 70.00% 80.00% data nudge open & click rate vs. industry open & click rate (%) opens rate(%) clicks rate(%) edu industry ave-open rate(%) edu industry ave-click rate(%) https://doi.org/10.29173/iq1010 12/16 orlowska, daria; fallaw, colleen; feng, yali; garza, livia; hetrick, ashley; imker, heidi & luong, hoa (2021) better data management, one nudge at a time, iassist quarterly 45(2), pp. 1-16. doi: https://doi.org/10.29173/iq1010 5. quotes from users we embed a one-question survey in the footer of each data nudge to learn if subscribers have been nudged into action by that month’s topic. after explaining how they have been nudged, subscribers are provided an optional field for their name and campus address to receive a swag bag. examples of feedback are provided below: "i just tagged my box data!! thanks for the fabulous tip!" “i have gone through your data nudges and i love them. i have been tasked as the lead of a data initiative [...] to improve the current data infrastructure. i would love to hear more about your thought processes.” “[...] i did want to share how much i enjoy the nudge and compliment you on your content. the halloween cartoon yesterday was especially well done, sometimes audiences equate simplistic with simple and here's a great example of how wrong that view is. kudos on your consistently well-crafted messaging.” “i've subscribed to these data nudge newsletters for a while, and i've found them pretty useful. this one [orcid] in particular is super important, especially because the culture of academia makes it so that certain people's work is acknowledged and recognized more than others'. little things like this can go a long way in changing that culture. thanks for sending these out!” “i just wanted to say again how much i enjoy the data nudge. i learn something valuable every time.” adaptability and reusability topics for data nudge data nudge content is under a cc-by license, and all topics are available online: https://go.illinois.edu/past_nudges. we are currently in the process of converting content into a reusable format (html and pdf), readily adaptable by other institutions and archived at ideals, known as illinois institutional repository: https://www.ideals.illinois.edu/handle/2142/109236. some data nudge topics can be easily customized by other institutions, while others are specific to illinois and require modification. nudge topics listed as “generally applicable” may require some changes to linked resources, but otherwise provide universally applicable content. on the other hand, nudge topics listed as “illinois specific” mostly focus on illinois resources and would require significant modification to customize for another institution. “mixed” nudge topics contain nudges that fall into both the “generally applicable” and “illinois specific” categories and require further consideration before adapting (table 2). the following table is a quick guide to determine the amount of institutional customization a data nudge topic will need: topic data nudges generally applicable illinois specific mixed back up data 2017-01, 2017-11, 2019-03 x https://doi.org/10.29173/iq1010 https://go.illinois.edu/past_nudges https://www.ideals.illinois.edu/handle/2142/109236 13/16 orlowska, daria; fallaw, colleen; feng, yali; garza, livia; hetrick, ashley; imker, heidi & luong, hoa (2021) better data management, one nudge at a time, iassist quarterly 45(2), pp. 1-16. doi: https://doi.org/10.29173/iq1010 campus events / resources 2017-09, 2018-04, 2018-11, 2019-09, 2020-08, 2020-09 x cloud storage 2017-06, 2017-07, 2018-03, 2020-01, 2020-03 x documentation 2017-05, 2019-10 x data analysis 2020-04 x data curation 2020-11 x data cleaning 2019-04 x data destruction 2019-11 x data loss 2020-05 x data management plan and practices 2017-03, 2017-10, 2020-02 x data repositories 2017-08, 2018-07, 2019-02 x data sharing 2018-05, 2018-06, 2018-09, 2019-01, 2019-07, 2019-08, 2020-06 x data visualization 2019-05, 2019-06 x file formats 2018-01, 2019-08 x file naming 2017-02, 2018-02 x risk assessment 2017-04, 2020-07 x scary data stories (halloween comic) 2018-10, 2020-10 x table 3. categorization of data nudge topics by reusability. https://doi.org/10.29173/iq1010 14/16 orlowska, daria; fallaw, colleen; feng, yali; garza, livia; hetrick, ashley; imker, heidi & luong, hoa (2021) better data management, one nudge at a time, iassist quarterly 45(2), pp. 1-16. doi: https://doi.org/10.29173/iq1010 discussion/conclusion what started out as a small pilot program intending to periodically nudge researchers about their interest in data has grown into an effective tool for the promotion of effective data management. along the way, we have learned that it takes a long time to craft a short message, it takes a networked team to craft a useful message, and that people really enjoy comics. the biggest lesson learned might be that the time and effort are worth it because our community values and uses our messages. the data nudge has become an outlet when we are bursting with a potential solution that we wish everybody had in their toolboxes. when librarians talk about rds to their units, subscribing to the data nudge has become a concrete step furthering the relationship between our service and crucial collaborators. the archive of past data nudges has become a resource for subscribers to go back and check on information about services or for rds team members to share right-sized guides on topics that come up in consultations. with our curated topics freely available for adaptation, we hope our work can be useful to other universities who want to re-use this content to jump-start a similar effort to engage with their own communities. we hope to someday see the nudge come full circle and adapt topics offered by other institutions in a mutually supportive network. acknowledgements the authors would like to thank their colleagues at illinois for supporting and promoting the data nudge. particular thanks go to elise dunham as the inaugural data nudge lead and for setting a high bar for quality and professionalism. we would also like to thank susan braxton, carissa phillips, dena strong, elizabeth wickes, and qian zhang for their contributions to the data nudge over the years and to our current graduate assistant, lauren phegley, for helping archive data nudges in a reusable format. references center for data and visualization sciences (no date). center for data and visualization science. duke university [online]. available at: https://library.duke.edu/data/ (accessed: 22 february 2020) constant contact. (2020). average industry rates for email as of december 2019. available at: https://knowledgebase.constantcontact.com/articles/knowledgebase/5409-average-industryrates?lang=en_us (accessed: 25 january 2021) cox, a.m. & pinfield, s. (2014). research data management and libraries: current activities and future priorities. journal of librarianship and information science; 46(4), pp.299–316. data services. (no date). data services, john hopkins university [online]. available at: https://dataservices.library.jhu.edu/ (accessed: 22 february 2020) data services. (no date). data services, new york university [online]. available at: https://library.nyu.edu/departments/data-services/ (accessed: 22 february 2020) data services. (no date). data services, washington university in st louis [online]. available at: https://library.wustl.edu/services/data/ (accessed: 22 february 2020) https://doi.org/10.29173/iq1010 https://library.duke.edu/data/ https://knowledgebase.constantcontact.com/articles/knowledgebase/5409-average-industry-rates?lang=en_us https://knowledgebase.constantcontact.com/articles/knowledgebase/5409-average-industry-rates?lang=en_us https://dataservices.library.jhu.edu/ https://library.nyu.edu/departments/data-services/ https://library.wustl.edu/services/data/ 15/16 orlowska, daria; fallaw, colleen; feng, yali; garza, livia; hetrick, ashley; imker, heidi & luong, hoa (2021) better data management, one nudge at a time, iassist quarterly 45(2), pp. 1-16. doi: https://doi.org/10.29173/iq1010 fearon, d. jr., gunia, b., lake, s., pralle, b.e., & sallans, a.l. (2013). spec kit 334: research data management. association of research libraries. available at: https://doi.org/10.29242/spec.334 (accessed: 21 february 2020) holdren, j.p. (2013). increasing access to the results of federally funded scientific research. office of science and technology policy [online]. available at: https://obamawhitehouse.archives.gov/sites/default/files/microsites/ostp/ostp_public_access_me mo_2013.pdf (accessed: 20 february 2020) information technology (2018). email fraud defense. university of illinois at urbana-champaign [online]. available at: https://techservices.illinois.edu/content/email-fraud-defense (accessed: 20 february 2020) mailchimp. (no date). [online]. available at: https://mailchimp.com (accessed: 20 february 2020) office of the vice chancellor for research & innovation (no date). 2019 research report. the university of illinois at urbana-champaign [online]. available at: https://research.illinois.edu/sites/research.illinois.edu/files/upload/research_report_digital.pdf (accessed: 15 march 2020) office of the vice chancellor for research & innovation (no date). by the number. the university of illinois at urbana-champaign [online]. available at: http://research.illinois.edu/researchillinois/numbers (accessed: 20 february 2020) ohaji, i. k., chawner, b. and yoong, p. (2019). the role of a data librarian in academic and research libraries. information research, 24(4), p. n.pag. available at: http://search.ebscohost.com/login.aspx?direct=true&db=lls&an=140844395&site=ehost-live (accessed: 14 february 2020). research data management service group (no date). research data management service group, cornell university [online]. available at: https://data.research.cornell.edu/ (accessed: 22 february 2020) research data service (no date). research data service, the university of illinois at urbanachampaign [online]. available at: https://www.library.illinois.edu/rds/ (accessed: 22 february 2020) research data services (no date). research data services, penn state university [online]. available at: https://libraries.psu.edu/research/research-data-services (accessed: 22 february 2020) research data services (no date). research data services, the university of michigan [online]. available at: https://www.lib.umich.edu/research-data-services (accessed: 22 february 2020) research data services (no date). research data services, the university of minnesota [online]. available at: https://www.lib.umn.edu/datamanagement (accessed: 22 february 2020) https://doi.org/10.29173/iq1010 https://doi.org/10.29242/spec.334 https://obamawhitehouse.archives.gov/sites/default/files/microsites/ostp/ostp_public_access_memo_2013.pdf https://obamawhitehouse.archives.gov/sites/default/files/microsites/ostp/ostp_public_access_memo_2013.pdf https://techservices.illinois.edu/content/email-fraud-defense https://mailchimp.com/ https://research.illinois.edu/sites/research.illinois.edu/files/upload/research_report_digital.pdf http://research.illinois.edu/research-illinois/numbers http://research.illinois.edu/research-illinois/numbers http://search.ebscohost.com/login.aspx?direct=true&db=lls&an=140844395&site=ehost-live https://data.research.cornell.edu/ https://www.library.illinois.edu/rds/ https://libraries.psu.edu/research/research-data-services https://www.lib.umich.edu/research-data-services https://www.lib.umn.edu/datamanagement 16/16 orlowska, daria; fallaw, colleen; feng, yali; garza, livia; hetrick, ashley; imker, heidi & luong, hoa (2021) better data management, one nudge at a time, iassist quarterly 45(2), pp. 1-16. doi: https://doi.org/10.29173/iq1010 strategic plan working group (2013). the illinois strategic plan, the university of illinois at urbanachampaign [online]. available at: https://strategicplan.illinois.edu/2013-2016/goals.html (accessed: 20 february 2020) wiley, c., & mischo, w.h. (2016). data management practices and perspectives of atmospheric scientists and engineering faculty. issues in science & technology librarianship; 85(1) [online]. available at: https://doi.org/10.5062/f43x84nj (accessed: 20 february 2020) endnotes 1 daria orlowska is data librarian at western michigan university, email: daria.orlowska@wmich.edu 2 colleen fallaw is research programmer for the research data service, university of illinois at urbana-champaign, email: mfall3@illinois.edu 3 yali feng is behavioral sciences research and data services librarian at university of illinois at urbana-champaign, email: yalifeng@illinois.edu 4 livia garza is former pre-professional graduate assistant at research data service, university of illinois at urbana-champaign, email: liviag2@illinois.edu 5 ashley hetrick is former assistant director of research data engagement education at research data service, university of illinois at urbana-champaign, email: ahetrick@illinois.edu 6 heidi imker is director of research data service, university of illinois at urbana-champaign, email: imker@illinois.edu 7 hoa luong is associate director of research data service, university of illinois at urbanachampaign, email: hluong2@illinois.edu https://doi.org/10.29173/iq1010 https://strategicplan.illinois.edu/2013-2016/goals.html https://doi.org/10.5062/f43x84nj mailto:daria.orlowska@wmich.edu mailto:mfall3@illinois.edu mailto:yalifeng@illinois.edu mailto:liviag2@illinois.edu mailto:ahetrick@illinois.edu mailto:imker@illinois.edu mailto:hluong2@illinois.edu vol282-3.indd 24 iassist quarterly summer/fall 2004 by ann s. gray1 introduction libraries are full of publications containing statistics and most offer databases with indexes to statistical publications. statistics are important: they are essential for social and economic development; for understanding among peoples; and necessary for any society that seeks to understand itself and respect the rights of its citizens. in order to provide a high level of support and service for these resources, librarians benefi t signifi cantly from having an interest in data and statistical literacy. the goal of the librarian is to direct users of electronic data and statistical resources toward useful information that refl ects the nature of the real world and to help users avoid the possible misuse of data and statistics. thus, as librarians, we need to: • understand statistical publications and electronic resources of national and international statistical organizations which are the primary source of most statistics; • deal with statistical publications and value-added commercial products which may actually hide statistical details from us; • be able to judge the different uses the media make of statistics; • and be able to make informed decisions regarding the use of mapping and other display tools such as charts, graphs, and other presentations of statistics. this paper looks at why data and statistical literacy for librarians are relevant, how data and statistical resources are evaluated, the types of information about data and statistics that one needs to know to provide assistance in the use of statistical resources, and the possibility that training in statistics would be useful, if not necessary, to providing these services2. the importance of statistical literacy data and statistical publications in electronic format were introduced in to u.s. academic libraries in the mid 1980s by u.s. federal statistical agencies. by the early 1990s data held on cd-rom were fairly common in most libraries in the u.s. with this came the the growth of data services as a function of the university library. data in its many forms had a foot in the door at many libraries and libraries were rapidly embracing computer technologies. in the u.s., 1996 saw the beginning of a loosely organized approach to centralized delivery of statistical information from the various federal statistical agencies. lead by the census bureau, an internet service known as fedstats was set up to provide one-stop shopping for accessing federal statistics3. in canada, progress was slowed by statistics canada’s steep price increases in its data products during the 1980s. however, the data liberation initiative (dli) was launched formally in 1996, which contributed signifi cantly to the growth of data support services within canadian libraries.4 with the delivery of statistical tables using the internet, there was a lot of concern over the potential misuse or misunderstanding of statistical data that would appear electronically without explanation. when summary tables and statistical reports were in print, it was expected that the report would include information about the tables that would allow for an evaluation of the information. but when tables became separated from the text and when the data became survey microdata, the question of utility and quality became more complex. not only are data resources being delivered directly to the analytically ‘unsophisticated’ public, they are also being grabbed by the mass media and used in ways that may not always be honest or useful. groups such as the statistical assessment service (stats) and the center for media and public affairs deal with the misuse of data by that sector, and many statistical organizations and government agencies are trying to solve the problems posed by direct access to statistical resources by establishing policies that reduce the chance of misuse.5 there are several reasons why statistical literacy has become important recently. we live in the information age with rapid distribution of news and content, where content is often overlooked in favor of images, and at a time when more and more statistics and data products are being made available to a larger and less data-literate audience. data and statistical literacy for librarians iassist quarterly summer/fall 2004 25 olenski (2003) suggests that we are also living in a time in which new democracies are emerging and recognizing the functions their national statistical agencies ought to play in the life of their countries and the global community6. in 1993 the united nations’ (un) economic commission for europe (ece) adopted a list of fundamental principles of offi cial statistics later endorsed by all as being of “universal signifi cance.”7 noting the importance of offi cial statistical information for development and the necessity that public trust should be based on scientifi c principles and professional ethics, the ece adopted 10 principles. the fundamentals recognize the importance of the quality of the statistics but emphasize that correct interpretation depends on the presentation of information on the sources, methods and procedures of the statistics and that agencies are entitled to comment on erroneous interpretation and misuse of data. agencies should protect the human rights of individuals, laws and regulations that govern the agencies should be public, and standards and concepts should be established to ensure comparability. the principles point to quality, timeliness, costs and respondent burden as key considerations in data collections. statistical literacy is also a concern of the international statistical institute (isi), which has its own literacy project8, and many national statistical organizations, such as statistics canada, the u.s. statistical agencies, and international agencies such as the international monetary fund (imf), have policies and guidelines to support both quality and utility for statistics. librarians already act as intermediaries to statistical resources within their libraries and it seems logical that they could serve the same function for supporting a broader range of data and statistical products. looking at how statistical publications are evaluated, there is a natural progression from this evaluative process to other quality measures that form part of the knowledge base of statistical literacy. evaluation of statistical publications print resources are evaluated based on their quality and objectivity, and on their utility and value. the publisher, author or peer reviewer is often used to determine quality and objectivity. the relevance, timeliness, scope, coverage, and presentation of the information are criteria that are typically used to judge utility and value of the piece within the context of the collection of the library and the purpose for holding the publication. this does not differ substantially from the guidelines set up by statistics canada in its elements of quality and its longer document quality guidelines.9 statistics canada quality guidelines statistics canada stresses the relationship between quality and utility. in order to judge quality and determine utility one needs accuracy indicators. these are possible with a presentation that fosters understanding. the agency accepts responsibility to provide suffi cient information about its statistics so that one can determine if the data or statistic can be used in a specifi c application or for a specifi c purpose. documentation on methodology and on quality are vital to data analysis. because neither the statistical agencies nor the data provider or data supporter can anticipate every future intended use of data, we must rely on the information already supplied about the data. statistics canada supplies information that covers the areas of relevance, accuracy, timeliness, accessibility, interpretability, and coherence. still, statistical literacy is necessary to interprete the information that is provided in order to determine if the statistic is right for a specifi c purpose or application. it is possible to rate quality in terms of accuracy using both expert opinion and measures of data accuracy. accuracy measures typically describe error, such as bias or systematic error and variance, also known as random error. documentation can also be judged based on its completeness and accuracy. “statistics canada’s “highlights, interpretations, statistical test results, and statements of trend, change or signifi cance” are other ways to provide users with a notion of quality.10 interpretability, another cornerstone of statistical literacy, requires information necessary to interpret and utilize statistical information appropriately. this covers the concepts, variables, classifi cations that are used as well as the methodology of data collection and the accuracy indicators and statements relating to the strength and weaknesses of the data. coherence refers to the use of standardized terms, classifi cations, and concepts among the data products, and can be described by information on how classifi cation schemes are related to each other, e.g. sic and naics, standard recodes and, if possible, aggregations.11 quality guidelines u.s. census bureau the quality guidelines of the u.s. census bureau are very similar in nature12. they also include utility and objectivity as goals for the agency, state that they will provide indicators of quality in a timely fashion and indicate that the analytic results of their products should be reproducible following the prescribed methodology (i.e., exact same method of data collection and analysis). that said, in reality it would be impossible for users to actually reproduce the surveys since the underlying microdata would not be available due to confi dentiality issues. other u.s. statistical agencies, following the direction of the offi ce of budget and management, have also tried to create policies and sites that promote good data practices13. thus we can look for statements regarding accuracy and methods in order to interpret the utility of statistical resources, but is that enough? some of the quality 26 iassist quarterly summer/fall 2004 measures that these agencies see as being important are given below. because so much of our current data comes from sample surveys, the quality of the survey method plays an important role in the resulting quality of the data. the following measures are important: • sample size is often regarded as the most likely indicator of quality reported in surveys. all things being equal a larger sample will produce more reliable results with less margin of error; • the sample frame needs to be current and contain information necessary to draw the sample; • response rate remains important even though there are studies that claim it may not. many survey methodologists still regard a high response rate as necessary to reduce nonresponse error: in other words, would those in the sample who did not respond have responded with the same variation as those that did respond?; • sampling error refers to deliberate undercounts or overcounts of groups of persons in the sample; • coverage error looks at the omission of sets of the population from the sampling frame (such as takes place when phone surveys omit non-phone households); • quality of inputs are important for derived or computed measures, such as estimates and indexes; • other types of error, for example, measurement error and processing errors, are less easy to detect. regardless of the format of the publication, this information may or may not be available. in print sources, we often look for other types of quality indicators. two commercial print publications two commercial publications were examined for quality indicators including: sources of data; sources of error; assumptions; adjustments; comparability; defi nitions; and methods. the fi rst, international smoking statistics, is a collection of historical data from thirty economically developed countries14. it includes statements regarding coverage, sources of data, brief notices of possible sources of errors, time period to which the data refer (e.g. midpoint year), and a brief description of the survey scope and methodology, if available. it further covers adjustments made to render the data comparable among the constituent nations, and provides information on the length of cigarettes and thus the amount of tobacco within, and whether hand-rolled or manufactured. however, the publication is associated with the wolfson institute of preventative medicine , an organization that aims to reduce smoking. this could lead one to question its objectivity. the second example, complete economic and demographic data source is both a book and a set of tables available on cd-rom15. the print version contains an entire chapter devoted to how estimates and forecasts are created, as well as defi nitions and sources. however, this chapter and the notes are not present in the cd-rom version of the tables. these two resources provide adequate information on sources, assumptions, comparability, defi nitions and methods provided one knows how and where to look for this type of information and how to evaluate it. data resources on the internet another frequent source of statistics are websites. an example is voter registration data assembled on a site at the university of california, berkeley. if one begins at the webpage for statewide information by assembly district16, there is little indication of how the data came about. data on registration and voting by race is not collected by the state of california nor the u.s. government, but the site has a report and data on the number of registered minorities in the 1996 california elections. the project homepage provides something of an introduction into this large project by a reliable source but there is no link from this page to its home. it appears that voter registration rolls furnished the name, location of residence, political party affi liation and vote history (if the person voted or not in various elections) the classifi cation of ethnicity was based on an analysis of surname. the validity of the latter type of classifi cation should be further investigated by users, for example fi nding out more about the methodology of using residence and surname to classify african-americans, hispanics, and jews . this particular internet resource did not provide much in the way of supporting studies to justify these methods and thus may be of questionable use to the new user. however, an exhaustive search of the site found a citation to a publication about the use of surnames and contextual information to determine race (lauderdale & kestenbaum 2000)17. clearly, a higher level of understanding of statistical methods would be necessary in order to interpret these types of resources effectively and to assist others in the selection of data or data resources for analysis. statistical literacy iddo gal (2002) made some observations on the components and attributes of statistical literacy18. writing in the international statistical review, gal looked at a person’s ability to interpret and critically evaluate statistical information and his or her ability to discuss or communicate reactions to same. to that i add the ability to apply the method of analysis, that is to interpret the information in order to determine the best method of analysis. the continuum of skills runs from the ability to evaluate data and statistics to the ability to communicate iassist quarterly summer/fall 2004 27 meaning and concerns. gal sees his defi nition of statistical literacy as requiring certain knowledge elements. these knowledge elements include literacy skills, statistical knowledge, mathematical knowledge, context knowledge, and critical questions. if we accept these elements as necessary components of statistical literacy, the question becomes one of degree. questions include how much statistical knowledge must one have to be literate? how much mathematical knowledge is needed to be able to communicate about statistics? should explanations of data quality and utility be framed in the scientifi c language of statistics? levels of competence: the standards for success in 1999 the association of american universities created a task force to “defi ne the knowledge and skills students would need in their fi rst year of college.”19 with support of the pew charitable trusts, the standards for success project was launched in 2000 and published, and are available on the internet (conley 2003)20. for the social sciences, the standards include areas dealing with statistical literacy and state that incoming freshmen should know how to interpret data presented in tables and graphs, know the basics of probability theory and the concept of a sample, and know the difference between statistical and substantive signifi cance. successful students are expected to know how to fi nd different sources of information and be able to analyze, evaluate the use them properly. critical thinking is emphasized in evaluation based on quality of materials, credibility of information, presence of bias, and the ability to draw inferences. students who plan on majoring in a social science are expected to have some basic knowledge of a statistical software package, and, depending on the fi eld, understand and know how to use analytic tools, including basic statistics. standards for statistical knowledge is included in the section on mathematics, although knowledge of statistics was not seen as a prerequisite for mathematical course work. rather the knowledge of statistics was included for use in the social sciences, particularly economics, as important for entry level college work. these levels of competence could also work for librarians in the fi eld of social science. understand and use summary data summary data can be presented as counts, calculations, percentages, rates, or ratios in tables, graphs, histograms, polygons, pie charts, or maps. for data in print, statistical literacy would cover some types of evaluation and interpretation as well as the ability to formulate some opinion or concerns regarding the presentation. much statistical information is presented as summary statistics, and exploratory data analysis – what one does prior to the actual tests, models, or other analytic tools – involves various summary statistics. interactive systems now allow us to create our own tables, graphs, charts, and begin the process of data analysis. exploring the data might involve preparing summary statistics about individual variables. the language of statistics refers to the examination of a single measure as univariate analysis. in this process, the distribution provides a sense of variance. the central tendency and dispersion also hint at the shape of the distribution. this might involve fi nding the maximum, minimum, mean, median, mode, and standard deviation. it should be possible to explain these measures using the language of statistics. for example, that the mode is useful where measures are expressed in categories, categorical variables such as 1=yes, 2=don’t know, and 3=no. the category that occurs most frequently is the mode. for continuous variables, such as actual years of age, the mean or median is the statistic of choice. this is not calculus, but it is far from simple because the usefulness of each statistic depends upon the intended purpose as well as how the measure was constructed. but from this basic level, statistical processes can rapidly climb beyond simple mathematical computations. today there are many online systems that allow users to manipulate data, create customized tables, and conduct higher level analysis. the popular online analysis tool, survey documenation and analysis (sda) allows users to determine the following measures: • frequencies or cross tabulations • comparison of means • correlation matrix • comparison or correlations • multiple regression • logit/probit • statistics: eta, r, somer’d, gamma, tau-b, tau-c, chisq(p), chisq (lr) df here we get into advanced descriptive statistics and the statistics of relationships and a higher level of understanding is required. even constructing and understanding a cross-tabulation can quickly become a source of confusion. knowing which method can be used or should be used requires a greater practical knowledge of common use as well as some theoretical understanding. although there are a number of online resources for students and teachers, including textbooks, glossaries and tutorials, the end-user probably wants some targeted, specifi c assistance when confronted with statistical information or data. it is my contention that this is best provided by a knowledgeable person who can engage the user in a structured dialog whereby he or she can determine the best statistic or measure based on “fi tness for use.” this human interaction is being abandoned in favor of 28 iassist quarterly summer/fall 2004 the development of online systems, probably with the belief that such systems lead to higher productivity and decreased cost. this is not to criticize those who promote improvements to online statistical resources, such as the un and the ece particularly when they emphasize the inclusion and organization of statistical metadata through publications and conferences, but rather to recognize that understanding of the need or use precedes the choice of statistic or method21. online help and the digital government project in an effort to help the general public understand the statistical information available from u.s. federal statistical agencies, the national science foundation and other federal statistical agencies have funded a number of projects that look at how statistical information is organized and presented. the govstat project at the school of library and information science of the university of north carolina and the university of maryland is multidimensional but part of the project involves creating some type of visual assistance for understanding statistical concepts.22 this type of online help would be part of a larger architecture of information but it is uncertain as to how it translates into the correct use of statistics. conclusion in addition to teaching students how to make help screens, library schools should teach statistical literacy. librarians need to know to make assessments of quality and utility -how to evaluate data and statistical publications. they should be able to furnish some guidance in interpretation of various types of statistical presentations and be able to point the way on how to interpret the results of analysis. as a group, they may never be statisticians, but they can perform a useful function in communicating meaning and concerns. statistical associations involved in statistical literacy have education programs aimed at high schools. many online resources have been developed to teach statistics, but as hans-joachim mittag noted, there is little “systematic cooperation.” (mittag 2000, p 6)23. there are many individual internet sites devoted to teaching statistics to classroom students (saporta 1999)24. there are very few aimed at data and statistical intermediaries, such as librarians. sociometrics has a data & internet literacy series (sociometrics 2005) that contains one section on understanding data in numbers, words, and pictures25. i have not seen it and cannot provide an evaluation. iassist would benefi t from having some members qualifi ed to teach and advise on the skill set we need to provide credible service in this area. the organization certainly has the ability to develop interactive, online, support to those who would be in the service of others. notes 1 contact: at the time this paper was presented, ann s. gray was the data reference librarian at princeton university library. 2 this paper is based on a talk given at the 2003 iassist conference held in ottawa on may 30, 2003. 3 fedstatss is an internet resource: http://www.fedstats. gov/ 4 dli update. internet resource: http://www.statcan. ca.english/dli/document/update.htm 5 stats: statistical assessment service, george mason university, is an internet resource. http://www.stats.org/ 6 olenski, jozef (2004), “the citizens’ right to information and the duties of a democratic state in modern it environment in the light of the un fundamental principles of offi cial statistics and the isi declaration on statistical ethics,” international statistical review, 71(1): 33-48. 7 united nations (1992), economic commission for europe. “fundamental principles of offi cial statistics in the region of the economic commission for europe”, e1992/32 c(47) available at http://www.unece.org/stats/ archive/docs.fp.e.htm; united nations (1993), economic and social council. statistical commission, “adoption of the agenda and other organizational matters: fundamental principles of offi cial statistics.” e/cn.3/1993/26; united nations (1994), statistics division, “fundamental principles of offi cial statistics” (1994), washington, 1994.2.10. available at: http://unstats.un.org/unsd/methods/ statorg/fp-english.htm. 8 isi (2004), international statistical institute international statistical literacy project (islp). internet resource: http://course1.winona.edu/cblumberg/islphome.htm. 9 statistics canada (1998), “statistics canada quality guidelines”, third edition, october 1998, ottawa; statisitics canada (1998a), “policy on standards”, internet resource: http://www.statcan.ca/english/concepts/policystandards.htm. 10 statistics canada (2000), “policy on informing users of data quality and methodology.” internet resource: http:// www.statcan.ca/english/about/policy-infousers.htm 11 north american industry classifi cation system (naics), was developed in cooperation with the u.s. economic classifi cation policy committee, statistics canada, and mexico’s instituto nacional del estadistica, geografi a e informatica. see http://www.census.gov/epcd/www/naics. iassist quarterly summer/fall 2004 29 html. in the u.s. it replaced the standard industrical classifi cation system (sic). one version of the sic codes and their meanings can be found in standard industrial classifi cation manaual [rev. ed.], washington, d.c. : executive offi ce of the president, offi ce of management and budget; springfi eld, va. : national technical information service, 1987 standard industrial classifi cation (sic) system. see http://www.census.gov/ epcd/www/sic.html. 12 united states census bureau (2004), u.s. census bureau section 515 information quality guidelines. internet resource. http://www.census.gov/qdocs/www/quality_ guidelines.htm 13 federal register (2004), “federal statistical organizations’ guidelines for ensuring and maximizing the quality, objectivity, utility, and integrity of disseminated information.” 67(107): 38467. available at: http://www. whitehouse.gov/omb/fedreg/reproducible.html/ 14 international smoking statisitics: a collection of historical data from 30 economically developed countries (2002), forey, barabara (ed.), london: wolfson institute of preventive medicine, oxford: oxford university press. 15 complete economic and demographic data source (ceeds) (2002), woods & poole economics, inc., washington, d.c. 16 statewide information by assembly districtinternet resource. http://swdb.berkeley.edu/info/statetext/staterpt. html#adstaterpt; statewide databases, institute of governmental studies. university of california, berkeley. internet resource. http://swdb.berkeley.edu/index.html 17 lauderdale, diane s. and bert kestenbaum (2000), “asian american ethnic identifi cation by surname”, population research and policy review (19(3) 283-300, june 2000. 18 gal, iddo (2002), “adults’ statistical literacy: meanings, components, and responsibilities”, international statistical review 70(1):1-25. 19 standards for success organization (1999), association of american universities. internet resource: http://www. s4s,org/02_projectoverview/history.php 20 conley, david t. (2003) understanding university success.: a report from standards for success: a project of the association of american universities and the pew charitable trusts, eugene, center for educational policy research, 2003. available at: http://www.s4s.org/03_ viewproducts/ksus/ mathematics: ksus_math.pdf (page 10) ; social sciences: ksus_social_sci.pdf. 21 united nations (2000), united nations. economic commission for europe. “guidelines for statistical metadata on the internet”, unsc/ece-ces statistical standards and studies, no. 52. 22 govstat (2005), university of north carolina school of library and information science, govstat program. internet resource: http://ils.unc.edu/govstat/ 23 mittag, hans-hochim (2000), “multimedia and multimedia databases for teaching statistics.” invited paper to be presented at the 9th international conference on mathematical education, makuhari/tokyo, july 31 august 6, 2000. available at: http://www.stat.auckland. ac.nz/~iase/publications/10/icme9_07.pdf 24 saporta, gilbert (1999), “teaching statistics with internet: a survey of available resources and the st@tnet project”. 52nd session of the international statistical institute, helsinki, finland. available at: http://www.stat. auckland.ac/nx/~iase/publications/5/sapo0830.pdf 25 sociometrics (2005), “about sociometrics data & internet literacy series” (dil series), available at: http:// www.socio.com/dil/ 1/2 rasmussen, karsten boye (2022) editor’s notes: openness in metadata, dictionaries, and data, iassist quarterly 46(1), pp. 1-2. doi https://doi.org/10.29173/iq1034 openness in metadata, dictionaries, and data welcome to the first issue of iassist quarterly for the year 2022 iq vol. 46(1). iassist quarterly has often published new and further developments in metadata. new submissions in the area are welcome, and the iq expects to continue presenting a flow of interesting articles on this important topic. often these developments in metadata directly mention their relatedness to the data documentation initiative (ddi). many iassist members have a primary role in supporting users of data in research and education. it is clear that the data item '42' needs metadata to be of any use. over a long period of time the developments in improving metadata have been rising to higher levels. among the latest achievements is the work presented in the first paper, on metadata extraction from available documentation using machine learning. the second paper is also drawing on metadata when constructing a dictionary for social terms for searching existing metadata, and for supporting new research's design and development. the last article is on a special type of data geospatial data or geodata and with a special look at those data as open data. the insights of this article bring us positive awareness of central concepts and areas that can be used to achieve greater openness. unfortunately, these days it is not evident in all parts of the world, but i must insist: openness is a good thing! closedness has always failed. it is with openness and free speech, openness in metadata, dictionaries, and data with free research, that we create a better world. the first paper 'engineering a machine learning pipeline for automating metadata extraction from longitudinal survey questionnaires' is written by a group of researchers and developers at several english institutions. suparna de and haeron pereira are in the department of computer science at the university of surrey, harry moss and sanaz jabbari are at the centre for advanced research computing at university college london (ucl), jon johnson and jenny li are at closer at ucl social research institute. the article describes the first results from the project of extracting existing questions in survey questionnaires available as pdfs. the machine learning is a central part of automating the extraction for further use in the metadata model of the data documentation initiativelifecycle (ddi-l). the supervised machine learning is supported by xml marked-up questionnaires that serves as training and validation. the article describes the details of the technical setup and the design of the processes and experiments, including the more complicated dataset and parameter versioning in machine learning. the experiment workflows are also graphically illustrated in the article. the two authors of the second article are ioannis kallas, professor at the university of the aegean, and dimitra kondyli, research director at the national centre for social research (ekke) – institute of social research in greece. their article presents 'a tool to promote research planning and conceptualization: sodanet research infrastructure’s scientific dictionary of social terms'. the research infrastructure presented is based on projects developed by sodanet, the greek infrastructure for social sciences, member of the cessda consortium. the described scientific dictionary of social terms was designed to support conceptualization, design, and management of research searched through sodanet. the dictionary is a computer application that is developed as a collective hypertext product to ensure continuity and validity. it is designed to meet the needs of the greek-speaking scientific community. the dictionary is dynamic and develops gradually as it is supplemented with new terms and definitions through the work of many researchers. the organization of the dictionary is on three levels: terms, definitions, and bibliographic records. the terms are in both english and greek. the article shows the example of the social term 'unemployed person'. as a tool for searching for scientific information, the dictionary also supports the design of new research and has a positive relationship to the ddi – also central in the first article. a follow-up article in the future could be on the project's contributions to the development and design of questions and hypotheses of new research. https://doi.org/10.29173/iq1034 2/2 rasmussen, karsten boye (2022) editor’s notes: openness in metadata, dictionaries, and data, iassist quarterly 46(1), pp. 1-2. doi https://doi.org/10.29173/iq1034 a group of four is behind the third article on 'open geospatial data: a comparison of data cultures in local government'. the authors are karen majewicz (geospatial project manager and metadata coordinator at the john r. borchert map library, university of minnesota), jaime martindale (map & geospatial data librarian at the arthur h. robinson map library, university of wisconsin-madison), and melinda kernik (spatial data analyst and curator at the john r. borchert map library, university of minnesota). the group was supported by yijing zhou (gis & metadata programming intern for the btaa geoportal, university of minnesota). the examples in this article compare the two states of minnesota and wisconsin. however, the issue of open data here open geospatial data concerns us all. the authors explore the gis history, programs, organizations, and legislation of the two states, examining how their different approaches have influenced the availability of open data. the article starts with worthy discussions and definitions or qualifications of 'open geodata' and also the many possible sources for creation of public geodata, as well as its availability and barriers. the case studies of minnesota and wisconsin are thorough in presenting the historical, political and legal developments and the many stakeholders. the differences between the two states are concentrated in the areas of legislation, funding, workflows, and library involvement. the many central concepts mentioned above are well summed up in the title as culture. the paper won the iassist conference 2020/2021 paper competition with the remarks of being well-written and addressing an important topic. submissions of papers for the iassist quarterly are always very welcome. we welcome input from iassist conferences or other conferences and workshops, from local presentations or papers especially written for the iq. when you are preparing such a presentation, give a thought to turning your one-time presentation into a lasting contribution. doing that after the event also gives you the opportunity of improving your work after feedback. we encourage you to login or create an author profile at https://www.iassistquarterly.com (our open journal system application). we permit authors to have 'deep links' into the iq as well as deposition of the paper in your local repository. chairing a conference session or workshop with the purpose of aggregating and integrating papers for a special issue iq is also much appreciated as the information reaches many more people than the limited number of session participants and will be readily available on the iassist quarterly website at https://www.iassistquarterly.com. authors are very welcome to take a look at the instructions and layout: https://www.iassistquarterly.com/index.php/iassist/about/submissions. authors can also contact me directly via e-mail: kbr@sam.sdu.dk. should you be interested in compiling a special issue for the iq as guest editor(s) i will also be delighted to hear from you. karsten boye rasmussen march 2022 https://doi.org/10.29173/iq1034 https://www.iassistquarterly.com/ https://www.iassistquarterly.com/ https://www.iassistquarterly.com/index.php/iassist/about/submissions mailto:kbr@sam.sdu.dk microsoft word 2024-10-retraction-mankone retraction of mankone, a. m. (2023). data protection and right to privacy legislation in kenya. iassist quarterly, 47(3-4). https://doi.org/10.29173/iq1080 for plagiarism (retraction made 10/2024). the editorial staff of the iassist quarterly received a charge of plagiarism for this article in september 2024, from mr. shadrack mutisya, a former subordinate of mr. mankone. upon receiving a draft of the paper from mr. mutisya, editorial staff compared its similarity to the published article using the tool ithenticate and found a 94% level of similarity between the two. mr. mankone has admitted to our editorial staff that he submitted his subordinate’s paper as his own without crediting mr. mutisya. co-authorship being unsatisfactory to mr. mutisya, the iq is retracting the article altogether. notice has been sent to the journal’s indexing services. this being mr. mankone’s only paper in the iq, this is the only paper affected. mr. mankone has been permanently banned from publishing with the iq. michele hayslett and ofira schwartz, editors 1/24 peller, peter (2018) from paper map to geospatial vector layer: demystifying the process, iassist quarterly 42 (3), pp. 1-22. doi: https://doi.org/ 10.29173/iq914 from paper map to geospatial vector layer: demystifying the process peter peller1 abstract with paper map use in decline, one of the strategies that libraries and archives can adopt to make the information contained within them more accessible and usable is to extract features of interest from their scanned raster maps and convert those to geospatial vector data. this process adds valuable unique data to library geospatial collections and enables those previously map-bound features to be used separately in geographic information systems (gis) software for custom mapping and analysis. advances in partially automating most of the process have made this a much more viable option for libraries and archives. although there is no one-size-fits-all automated solution for all maps and map features, this paper provides a complete description of the entire process incorporating examples of the various techniques and software used in selected studies that would be applicable in the library and archive environment. keywords map features, vectorization, raster-to-vector conversion, geospatial vector layers, digitization, feature extraction introduction with the paradigm shift towards digital mapping sources, the use of paper maps has significantly declined over the past fifteen years to the point that much of the valuable information in them is at risk of being forgotten or ignored. many libraries and archives have responded to this challenge by digitizing2 parts of their map collections – making the resulting raster3 images more universally accessible through the internet. some have even gone a step further and have transformed those raster images – through a process called georeferencing – into geospatial raster images that are compatible with geographic information systems (gis) software. the georeferenced maps in the david rumsey map collection (cartography associates, 2017) are an excellent example of how historic geospatial raster images can be used in google earth or google maps as overlays. more recently, a few libraries and archives have taken the next step and have started to experiment with the extraction of specific features of interest from the geospatial raster imagery to create entirely new sources of geospatial vector4 data that can be manipulated in innovative ways that were not previously possible. in order to work with gis software, the extracted features must be point, line or polygon vectors (geospatial vector data) with geographic coordinate information – representing real-world geographic features. these vectors are organized into single-type layers (ex/ all road segments) which are the fundamental units used by gis software.these are normally created from air photos, satellite data, or from ground surveys. historic maps allow us to travel backwards in time to reverse engineer the original data. 2/24 peller, peter (2018) from paper map to geospatial vector layer: demystifying the process, iassist quarterly 42 (3), pp. 1-22. doi: https://doi.org/ 10.29173/iq914 the conversion of raster images to vector data – otherwise known as vectorization – is analogous to ocr technology extracting words from scanned print documents. without ocr it would be impossible to search for specific text in a scanned document or do any kind of textual analysis without manually going through the entire document. the same is true for scanned maps. vectorization extracts pixels delineating physical and cultural features (e.g. contour lines, roads, soil zones) from a geospatial raster image and converts them into vector point, line and polygon layers. the resulting individual geospatial vector layers are searchable and useable for custom mapping and spatial analysis purposes in gis. although it is possible to manually vectorize features from geospatial raster images through heads-up digitizing5, the process is very tedious and labor-intensive, making it a suitable option for only simpler oneoff maps. a great deal of research effort has gone into developing automated solutions, but a fully automated solution that works on all features in all maps still does not exist due to the huge variety and complexity of maps (see figure 1): different features, symbols, colors (including textures), labels, and textual information – all of which can also overlap and intersect each other. nevertheless, progress has been made, and it is now possible to semi-automate the map vectorization process. this development makes it a more viable option for libraries and archives considering map vectorization projects. the goal of this paper is to provide a detailed review of all the steps that comprise the process of converting a map feature on a print map into a geospatial vector layer. this review incorporates examples of methods and procedures that would be most relevant in a library and archives setting. the examples are taken from nine selected studies which include all the published library-based ones (arteaga, 2013; godfrey and eveleth, 2015; marciano et al., 2013; pearson et al., 2013; bracke et al., 2008) and a few others that utilized commercial or open source solutions and were of a more applied nature (brown, 2002; jung, 2009; southall, 2003; whitfield, 2005). for the purposes of this review, two of the selected studies (jung, 2009; southall, 2003) have been split into two separate cases each: one study used two different methods and the other study extracted two different types of features. the eleven cases are summarized in the appendix. 3/24 peller, peter (2018) from paper map to geospatial vector layer: demystifying the process, iassist quarterly 42 (3), pp. 1-22. doi: https://doi.org/ 10.29173/iq914 figure 1. an example of map complexity: overlapping text, intersecting features, different features with the same color, grid line, symbols, textured areas and map damage. (source: map of the dominion of canada, 1929) print map to geospatial vector layer process the success of automated vectorization is dependent not only on the raster-to-vector conversion, but also on everything from the quality of the original print map to all the processes that enhance, isolate and edit the pixels of interest prior to this step, and the subsequent fine-tuning work on the extracted vector layer. the entire process can be broken down into the following steps: 1) scanning 2) georeferencing 3) image enhancement 4) image segmentation 4/24 peller, peter (2018) from paper map to geospatial vector layer: demystifying the process, iassist quarterly 42 (3), pp. 1-22. doi: https://doi.org/ 10.29173/iq914 5) raster editing 6) raster to vector conversion (vectorization) 7) vector editing. 1. scanning usage, storage conditions and time can take their toll on the original paper map: tearing, staining, creases, fading, color loss and color bleeding. although some of these defects can be mitigated in later steps, a poor quality map will nevertheless negatively impact the vectorization process. one option may be to borrow maps that are in better condition from other libraries (bracke et al., 2008). the paper map is scanned using either a digital camera system or a flatbed or roll scanner. scanning resolution is an important factor. on a scanner the terms dots per inch (dpi) or pixels per inch (ppi) are often used interchangeably and basically refer to samples per inch. the studies that reported their scanning resolutions used between 200-600 dpi (bracke et al., 2008; brown, 2002; pearson et al., 2013; southall et al., 2003; whitfield, 2005). the major finding from these studies is that – for vectorization purposes – higher resolutions are not always better (pearson et al., 2013; southall et al., 2003). it really depends on the size of the features of interest and – in the case of lines – their closeness to each other. slightly higher resolutions prevent very close lines from fusing together. on the other hand, higher resolutions tend to capture the paper texture and ink spread which introduces noise and errors into the image. pearson et al’s (2013) research found that a 400 ppi resolution produced 1.37 times more gaps and bridges than the 330 ppi resolution of the same map. a bridge is a link between two lines that shouldn’t be linked and a gap is a break in a line that should be continuous. the other consideration with higher resolutions is the larger file size and its implications for later processing. for most purposes, it seems that a 300 dpi scanner resolution is sufficient and preferable for vectorization purposes. other factors are bit depth, thresholding, and output file type. a 24 bit color is most commonly used with colored maps for vectorization purposes (allord et al., 2014; chiang et al., 2016). bit depth also affects image file size: one study found that 8 bit color was sufficient and generated tiff images that were 26mb in size at a resolution of 200dpi (southall et al., 2003); while another used 32 bit color at a resolution of 600dpi and generated a tiff image that was 730mb in size (bracke et al., 2008). thresholding on a scanner was employed by one united states geological survey (usgs) study to remove greenlines from their mylar maps (whitfield, 2005). as recommended by the usgs (allord et al., 2014), saving output in a lossless compression format, such as tiff, ensures the retention of as much of the original data from the map as possible; this was the output file format used by the majority of the studies. 2. georeferencing in the studies examined, there was a bit of variation as to when georeferencing occurred. it doesn’t really matter, but the one advantage of completing it before the segmentation and raster editing steps is that the georeferenced raster map image can then be overlaid on a base map. this enables one to see how 5/24 peller, peter (2018) from paper map to geospatial vector layer: demystifying the process, iassist quarterly 42 (3), pp. 1-22. doi: https://doi.org/ 10.29173/iq914 well the segmentation worked and also helps to visualize whether the raster editing is being done correctly. georeferencing is the process of assigning the raster map its geographic coordinates and coordinate system; in other words, it associates a map image with its actual location in geographic space. gis software like arcgis is normally used to do the georeferencing (bracke et al., 2008; godfrey and eveleth, 2015; pearson et al. 2013; whitfield, 2005); however, the r2v vectorization software has the georeferencing capability built into it (brown, 2002). basically the process involves using gis software to fit the raster map to a geospatial layer – that does have a coordinate system – using control points. a control point is simply a point on the raster map image that matches a point on the geospatial layer such as a street intersection, boundary, or grid point. when applying control points it is advisable to alternate adding them at opposite sides of the map and distributing them throughout the map. how many control points are created will depend on the map; there will be a diminishing return after a certain number. the main thing is to keep an eye on each point’s residual error and the total error which is calculated by taking the root mean square (rms) of all the residuals. the “residual error is the difference between where the point ended up as opposed to the actual location that was specified” (esri, 2016a). it is advisable to delete and redo any control points that have unacceptable errors. “when georeferencing, the goal is to attain a root mean square (rms) error that is less than or equal to the cell size of the raster file; this cell size represents the accuracy of the data” (esri, n.d.). the bracke et al. study (2008), which used 330 control points, reported a rms of approximately zero; however, it acknowledged that this may have been overkill when dealing with older maps and themes such as soil zones which don’t actually have finite edges. for most purposes, it would probably be sufficient to use a similar number (16) to the whitfield case (2005). in order to save the georeferenced raster image with its new coordinate information it needs to be georectified. this georectification involves transforming the image (scale, skew, rotate, translate, stretch and warp) with a transformation equation. a first order polynomial transformation is sufficient for the vast majority of scanned maps (esri, n.d.). once a raster map in tiff format has been georeferenced and georectified it is usually saved as a geotiff file. georeferencing can be a time-consuming process and a few major projects have actually used crowdsourcing for this. the new york public library (nypl) has adapted a program called mapwarper for this purpose and the british library has used georeferencer for their crowdsourced georeferencing. both of these georeferencing tools are open source. (fleet et al., 2012) 3. image enhancement most studies actually incorporated the image enhancement into the raster editing step; however, a couple of studies did it prior to segmentation which is why it is listed separately here. various techniques can be applied to enhance the raster image created by scanning. if the image does end up skewed, it can often be deskewed with the scanning software. in one case, adobe photoshop and arcgis desktop were used 6/24 peller, peter (2018) from paper map to geospatial vector layer: demystifying the process, iassist quarterly 42 (3), pp. 1-22. doi: https://doi.org/ 10.29173/iq914 to apply blurring, stretching and cubic convolution resampling to “reduce variation and noise by smoothing the inconsistencies with the colors” (godfrey and eveleth, 2015, p 27). photoshop was used to resample the image in one project where the researchers wanted to reduce the resolution of the original scanned image (southall et al., 2003). this latter process is a way to turn an archival quality scanned map into a lower resolution raster image without having to scan it again at a lower resolution. a mask can be used to exclude extraneous parts of the map such as the title, legend, scale, neat lines and annotations outside the actual mapped area; this simplifies the segmentation and reduces some of the raster editing that might otherwise need to be done later (godfrey and eveleth, 2015). 4. image segmentation segmentation is the process used to isolate the feature of interest. it is predominantly based on the color characteristics of the feature. “in particular, focus has been made on line extraction on binary images, and in maps on feature extraction on each colored layer” (lacroix, 2009, p 318). with a pure binary image where black pixels make up the foreground feature of interest and white pixels make up the background or vice versa, there is no need to do any further segmentation. however, with greyscale and color images, there are several ways to accomplish this segmentation. simple thresholding can be done on greyscale images just using the pixel color intensity histogram to identify which pixel values best identify the feature of interest – see figure 2. the situation gets more complicated with color maps. one option is to convert the color maps to 8 bit grayscale and then use the simple thresholding to segment the foreground pixels of interest and background pixels. color image segmentation can also be accomplished with image processing software such as photoshop (southall et al., 2003), gimp (arteaga, 2013) or imagemagick (pearson et al., 2013). the southall project tested a class reduction technique where all the colors were first divided into 100 classes and then manually each one of those was assigned to 1 of the 7 land use categories in the map (southall et al., 2003). r2v, a commercial vectorization tool, has the thresholding function built into it and this was used to isolate geologic contacts and faults in one case (brown, 2002). most of the projects – where color features were involved – used remote sensing techniques in software such as erdas imagine (southall et al., 2003), definiens ecognition/professional (bracke et al., 2008; jung, 2009), and arcgis (godfrey and eveleth, 2015; marciano et al., 2013; southall et al., 2003) to perform the segmentation. 7/24 peller, peter (2018) from paper map to geospatial vector layer: demystifying the process, iassist quarterly 42 (3), pp. 1-22. doi: https://doi.org/ 10.29173/iq914 figure 2. grayscale thresholding histogram with x-axis showing intensity and y-axis showing number of pixels. foreground pixels of interest are approximately between the 60 and 90 intensity levels. remote sensing classification techniques, normally applied to classifying satellite image pixels by their spectral reflectance values, can be adapted for classifying map images. there are two major types of classification: unsupervised and supervised. in unsupervised classification, the number of different classes are specified by the user and the algorithms will cluster similar pixels – based on shared spectral patterns – into the specified number of classes; this method utilizes clustering statistical methods. in supervised classification, the user selects “training samples” of similar pixels from the image and assigns them to a class; based on the training sample, the algorithm groups together pixels which are neighbours and have similar pixel values and separates them from groups of pixels which are dissimilar in value. these training samples can be reused with other maps in the same series provided that the colors in the maps are fairly similar (southall et al., 2003); this benefit can potentially save a lot of time in the classification process. when picking training samples it is important to pick a number of them for each color zone; a single sample is not sufficient to define a good average for a color zone. also, it is recommended that training samples for a color zone should be selected from cluttered areas – which also contain text and unwanted map elements in addition to the homogenously colored area – rather than very clean areas; it is believed that this technique will train the algorithm to ignore some of this noise (southall et al., 2003). ultimately the goal is to minimize the microscopic heterogeneous noise in macroscopically homogeneous zones without actually losing any legitimate microscopic homogeneous zones (bracke et al., 2008). classification can be either pixel-based or object-based; the latter goes beyond just classifying pixels spectrally but also combines that with structural analysis to use shape and spatial characteristics to group 8/24 peller, peter (2018) from paper map to geospatial vector layer: demystifying the process, iassist quarterly 42 (3), pp. 1-22. doi: https://doi.org/ 10.29173/iq914 pixels together into objects. object-based classification can offset some of the problems with variations in the same color as well as the effects of embedded noise (bracke et al., 2008; jung, 2009); however, it does require more expertise and experimentation in order to determine the optimal parameters. depending on the feature colors and segmentation technique used, the output from the image segmentation will be an image with foreground pixels delineating one of the following categories (figure 3): a. lines representing one feature type such as geologic contacts, roads or contours (brown, 2002; jung, 2009; pearson et al., 2013) b. outlines of areas representing one feature type such as geologic formations (whitfield, 2005) c. areas representing one feature type (with one class) such as water bodies or urban areas (jung, 2009) d. areas representing one feature type (with multiple classes) such as building types, soil zones, snow loads or land use (arteaga, 2013; bracke et al., 2008; godfrey and eveleth, 2015; southall et al., 2003) the usual reason for using outlines of areas in category b) above is due to fact that the area pixels are indistinguishable from the background pixels. 9/24 peller, peter (2018) from paper map to geospatial vector layer: demystifying the process, iassist quarterly 42 (3), pp. 1-22. doi: https://doi.org/ 10.29173/iq914 a) b) c) d) figure 3. the four categories of output from the segmentation step with original map on left and segmented features on right: a) single class line feature (rail lines), b) area feature outlines (census divisions), c) single class area feature (lakes), and d) multi-class area feature (first nations treaty areas). 10/24 peller, peter (2018) from paper map to geospatial vector layer: demystifying the process, iassist quarterly 42 (3), pp. 1-22. doi: https://doi.org/ 10.29173/iq914 5. raster editing due to the problems with consistent colors, overlapping text, other intersecting features, and varying line widths, the segmentation process rarely isolates the feature of interest perfectly; therefore, the raster image usually requires some editing before it can be vectorized in an automated way. three of the examined cases, however, did bypass the raster editing step: for arteaga (2013) it was due to the automated process followed and choice of software; for bracke et al. (2008), it was because the segmentation result was exported as a vector geospatial file from the definiens software resulting in automatic vectorization; and, for pearson et al. (2013), it was for the reason that the purpose of their study was to record errors. although, raster editing can be quite laborious, it is usually worth the effort; however, it will depend on the software and processes used. a) b) c) d) figure 4. errors in raster image: a) bridges between two lines that should not be connected; b) gaps and holes in a continuous line; c) text overlapping feature; and, d) intersecting features (dashed power line right of way intersecting solid property lines). 11/24 peller, peter (2018) from paper map to geospatial vector layer: demystifying the process, iassist quarterly 42 (3), pp. 1-22. doi: https://doi.org/ 10.29173/iq914 the usual errors with foreground pixels representing lines and outlines are holes within the pixel group making up a line feature, breaks (gaps) in a continuous line feature, connected lines that should be separate (bridges), and intersecting pixels belonging to other map elements such as text or other features – see figure 4. for foreground pixels representing areas, the errors are leftover pixels from unwanted map elements within the feature areas, holes, and misclassified pixels. the main tools available to repair these issues include morphological operators, gis functions, bulk erase, manual erase, and manual draw/paint. 12/24 peller, peter (2018) from paper map to geospatial vector layer: demystifying the process, iassist quarterly 42 (3), pp. 1-22. doi: https://doi.org/ 10.29173/iq914 a) b) c) d) figure 5. morphological operators: a) pixels representing line with gaps and holes; b) dilate operator applied to line in (a) fixing holes and gaps; c) pixels representing a boundary line that is intersected by a lake feature outline and text with small noise pixel at bottom left; and, d) erode operator applied to line in (c) removing the intersecting text and lake outline as well as noise – a few bits of leftover text remaining that can be easily erased. the use of morphological operators automates the editing of foreground pixels delineating lines, outlines and single class feature areas (jung, 2009). morphological operators are filters – composed of a small array of pixels – that are applied to each pixel in the raster image. if the pixels in the filter match those in the underlying image then it is a “hit” and if not then a “miss”. depending on the operator type, the hit or miss results in a certain action. the two most common operators are erode and dilate which are correspondingly used to remove or add foreground pixels to a raster image – see figure 5. the erode operator can be used to remove text, unwanted intersections, and bridges; the dilate operator can be 13/24 peller, peter (2018) from paper map to geospatial vector layer: demystifying the process, iassist quarterly 42 (3), pp. 1-22. doi: https://doi.org/ 10.29173/iq914 used to fill holes and to close gaps (chiang, 2010; chiang et al., 2005, 2014). care must be taken when applying these operators because fixing one problem can sometimes create another: for example, fixing gaps can create bridges. morphologicial operators can also be applied iteratively to achieve the required clean up (jung, 2009). arcgis desktop includes the arcscan extension which comes with the erode, dilate, opening (erode then dilate), closing (dilate then erode) morphological operators (esri, 2016b). if there are still problematic pixels or missing pixels after applying the morphological operators, the manual erase and manual draw/paint can be used to fix these problems more precisely. with both arcscan and r2v there is also the ability to bulk select connected foreground pixels based on parameters such as area, diagonal length and width; once selected, the foreground pixels can be changed into background pixels and vice versa (able software corporation, 2008; esri, 2016b). arcscan’s magic eraser will bulk erase connected foreground pixels by touching them with the tool or drawing a box around them with it (whitfield, 2005). arcscan also has a gap setting (width & angle) and a hole setting: these will respectively direct the automatic vectorization to leap any matching gaps and to ignore smaller holes. one other very handy element of arcscan is the ability to preview in advance what the vectorization will look like as each edit is made (esri, 2016b). in both arcscan and r2v it is easy to undo changes that don’t result in the desired outcome. with foreground pixels delineating multiple class feature areas, gis functions such as arcgis’ nibble, shrink and expand tools and majority filter (esri, 2016c) are the most automated way to repair them – see figure 6. the southall study (2003) applied the majority filter first to get rid of the bulk of the unwanted noise pixels within the feature areas; this replaces the noise pixels with the value of the majority of their contiguous neighbours. the nibble was then applied to eat up the small leftover noise bits and remove them (southall et al., 2003). another investigation (godfrey and eveleth, 2015) took a slightly different approach since their unwanted map elements had been removed completely from the raster image leaving no data areas. this was due to the iterative process they used of creating a separate raster for the best matching feature class by masking everything else out each time. following each iteration, they ran the shrink and expand to fix inconsistencies with the edges and after merging all the separate raster images together, they applied the expand tool again to fill any remaining no data gaps. although not explicitly stated, it appears that a similar approach for dealing with distortion on the edges was achieved in another case through the application of the boundary clean operation; this smooths boundaries with a combined expand and shrink in one or two passes (marciano et al., 2013). 14/24 peller, peter (2018) from paper map to geospatial vector layer: demystifying the process, iassist quarterly 42 (3), pp. 1-22. doi: https://doi.org/ 10.29173/iq914 a) b) c) figure 6. a) original map with overlapping text and intersecting lines; b) segmented map showing boundary between 2 multi-class areas; c) using majority filter, expand and shrink to remove leftover noise and unwanted map elements. 6. raster to vector conversion (vectorization) the raster to vector conversion depends on the category of output from the segmentation (figure 3). in the case of feature areas, vectorization is based on the colored raster classes created in the segmentation step and can be performed on a single feature type (one or multiple class). for pixels representing lines, the conversion is done on one feature type at a time. automated polygon vectorization is usually achieved through a standard raster to vector conversion process using gis, such as arcgis’ rastertopolygon tool (godfrey and eveleth, 2015; southall et al., 2003) or in nypl’s case the open source gdal’s polygonize tool (arteaga, 2013). the projects that used the definiens software for segmentation/classification were able to directly export the resulting raster as a vector polygon shapefile (bracke et al., 2008; jung, 2009). arcscan can be used to vectorize polygon-like pixel groups that exceed a specified pixel width; this requires a binary image and can only be done on a single feature type at a time. another option is using arcscan to vectorize the boundaries of feature areas as outlines and then later convert them into polygons (whitfield, 2005) in the vector editing stage. lines are inherently more problematic to vectorize in an automated fashion. a line or a boundary on a raster map image consists of a certain thickness of pixels; however, due to quality issues – either on the original map or from the scanning process – this thickness will not be uniform throughout. this makes it more difficult for software to accurately vectorize lines – see figure 7. in addition, corners, junctions and line intersections are harder to interpret – especially low angle intersections. arcscan has three settings – geometrical (preserves angles and straight lines), median (designed for non-rectilinear angles) and none (designed for non-intersecting features) – that can be applied to mitigate some of the issues with 15/24 peller, peter (2018) from paper map to geospatial vector layer: demystifying the process, iassist quarterly 42 (3), pp. 1-22. doi: https://doi.org/ 10.29173/iq914 intersections. the commercial software that was used for line vectorization in the projects examined was either arcscan (whitfield, 2005; jung, 2009) or r2v (brown, 2002). due to the requirement of a binary image, lines of only one color can be vectorized at a time. figure 7. how differing line width can impact vectorization of a right angle intersection. if the majority of errors in the raster map image have been removed or corrected, the automated vectorization should output a reasonable vector line layer. it will never be entirely perfect and some vector editing may be required, but it is a real time-saver (whitfield, 2005). when the raster image simply has too many problems – with noise, missing pixels, and intersecting unwanted pixels – to be edited in a reasonable amount of time, the other option is to use interactive raster tracing to delineate the lines or boundaries. with interactive raster tracing, the user clicks on the line pixels and indicates the direction; the software then automatically traces a line until it encounters a spot – usually an intersection or gap – where it doesn’t know which way to proceed. the user then points it in the right direction and off it goes again to the next ambiguous spot. this process is essentially a semi-automated form of heads-up digitizing. arcscan and r2v include both automated vectorization and interactive raster tracing as well as the option to select areas of the raster image for either automatic or trace vectorization. this allows for a hybrid approach: automatic vectorizing of clean straightforward areas and trace vectorizing of noisier more complex areas. the different areas can then be exported and merged together into one vector line layer. 16/24 peller, peter (2018) from paper map to geospatial vector layer: demystifying the process, iassist quarterly 42 (3), pp. 1-22. doi: https://doi.org/ 10.29173/iq914 7. vector editing vector editing is the final phase of the entire process. there are a number of different operations done in the vector editing step: cleaning up and fixing errors, dealing with leftover noise, filling no data gaps, smoothing lines and polygons, and assigning attribute data to features. a number of strategies were applied to dealing with the leftover noise or no data gaps on multiple-class feature type polygons. southall et al. (2003) ran the arcgis eliminate function to dissolve areas of unwanted map elements (below a minimum threshold) into the polygon with the longest shared boundary. bracke et al. (2008) closed up these gaps with empty filler polygons in arcgis. then, using a spatial join operation twice, they assigned the class from the nearest polygon to each empty filler polygon. the dissolve tool was then applied to aggregate all the smaller filler polygons that intersected or were contained within larger same class feature polygons. as none of the above methods were foolproof in assigning the correct category, some manual recoding was required for a few misclassified polygons. godfrey and eveleth (2013) also dissolved their polygon features to clean things up. smoothing of the vector polygon boundaries was done with the arcgis smooth polygon tool in the godfrey and eveleth project (2013) in contrast to some of the others who had done this as part of the raster editing step. they also had to clip the vector result to the extent of the original mapped area; this was necessitated by the spillover caused by their use of the expand tool to fill in all the no data gaps during the raster editing phase. the nypl project did not do any raster editing prior to vectorization (arteaga, 2013). most of their work was done in the vector editing stage. they used the r software for shape simplification and for polygon exclusion; the latter was based on minimum and maximum thresholds. further polygon exclusion was done by comparing the polygon to the color of the corresponding area on the raster; the white polygons – which corresponded to the background – were removed. although not stated in the arteaga description, it appears that since their initial project was begun, nypl has developed a crowd-sourced tool, building inspector, to assist with the quality control work of checking, and if necessary, modifying building polygons through the adjustment of vertices (new york public library, n.d.). nypl has put together all the script and templates that went into their project into an open source tool called map vectorizer which is available through github (arteaga, 2017). as for the vector editing of the extracted line features (including polygon boundary lines), it is important to compare it to the original map so that any missing or erroneous lines be corrected. these can be fixed with arcgis’ edit tools. in the case of boundary lines these can be converted to polygons using arcgis’ featuretopolygon tool (whitfield, 2005). the vector lines created through vectorization often have too many vertices; this gives the resulting lines a stair-case appearance. arcgis has a smooth line tool that can make the line look better. the r2v software also has a built-in “smooth lines” command to do something similar (brown, 2002). pearson et al. (2013) used the novel approach of smoothing the extracted contour lines by converting the vector back to raster and then re-vectorizing. 17/24 peller, peter (2018) from paper map to geospatial vector layer: demystifying the process, iassist quarterly 42 (3), pp. 1-22. doi: https://doi.org/ 10.29173/iq914 once the vector lines and polygons have been finalized there is still one more task to complete – the addition of attribute data to the features. these can be road names, river names, administrative units, geologic units, etc. in the case of single or multiple class feature polygons, it is a straightforward process to assign each class a proper name by editing the existing class name. where more detail is required or in the case of most linear features, this is accomplished by creating fields in each feature’s attribute table and populating them with their corresponding data. although this is largely a manual process, it can be expedited by the use of lookup tables; one basic field is entered in the attribute table and then it is joined to the lookup table on the key field (whitfield, 2005). this, in essence, automatically transfers all of the corresponding information in the other fields to the feature. discussion the previous sections have highlighted what seems to be a myriad of ways to extract features from a raster map and convert them to a vector geospatial layer, but it is important to keep in mind that the eleven different cases really aren’t that different in their overall scheme; they do the same thing but just differ slightly in which tools are used and when. each study was unique though, and employed some novel techniques worth considering for any map vectorization project. the key point is that a certain amount of experimentation will be required upfront to identify the optimal processes for any vectorization project. the goal of this review was not to judge these methods but rather to report them. in order to state that one method is better than another, a comprehensive comparison of the results from all methods would need to be done on the same map and that was beyond the scope of this paper. even then, some methods may work better for a particular kind of map or the specific feature to be extracted. these kinds of comparative analyses would be potential areas for further research. when evaluating the methods used, the quality of the end result is not the only factor to consider. the time required to complete the whole process is just as important, particularly when vectorizing large numbers of maps. most of the reviewed studies did not mention the specific time involved so it was not possible to compare them by this factor; however, the southall et al. study (2003) did compare the time requirements between its two methods as well as with a full manual approach. the upshot is that achieving greater spatial accuracy will usually require more time; this means balancing the trade-off between accuracy and time taken to best serve the potential map use. the requisite spatial accuracy is determined by the ultimate use of the geospatial vector layer and is impacted by all the steps in the process. although vector layers are scalable – basically to any scale level – they are usually intended for a specific narrow range of scales depending on their purpose. levachkine identified two types of gis: analytical gis and register gis (levachkine, 2004). analytical gis does not require the same high level of accuracy as register gis, because the former entails working with thematic (soils, geology, vegetation, etc.) data which is usually less exact and at smaller scales. on the other hand, exactness is much more of an issue for register gis as it is concerned with topographic (contours), cadastral (properties and buildings), utility or transportation data at larger scales. the extraction of buildings (arteaga, 2013), roads (jung, 2009) and contour lines (pearson et al., 2013) would be categorized 18/24 peller, peter (2018) from paper map to geospatial vector layer: demystifying the process, iassist quarterly 42 (3), pp. 1-22. doi: https://doi.org/ 10.29173/iq914 as register gis; the other projects would all fall into the analytical gis category (bracke et al., 2008; brown, 2002; godfrey and eveleth, 2015; marciano et al., 2013; southall et al., 2003; whitfield, 2005). this review did not comprehensively test the different software; although some experimentation was done to better understand the processes, it was not applied in a systematic way. therefore, the review can’t recommend one software solution over another. the focus, as specified earlier, was on projects that used either commercially available or open source solutions and each program used was identified in its corresponding step. the selected studies were done over a period of time from 2003 to the present and one must keep in mind that the different software have probably evolved over the same period. where doing a specific task with a certain software may not have been possible a decade ago, it may now be possible to do so. through the close examination of these nine studies, it has been demonstrated that there is no easy-touse, one-size-fits-all automated solution – that would work for all maps and all features – for extracting features from a paper map and converting them into vector geospatial layers. the most progress on automation has been achieved in the segmentation and vectorization steps, but there have been some developments in the raster and vector editing steps as well. the type of map and feature (and amount of noise) will largely determine the level of automation that can be exploited, but in all cases some manual intervention and handling are unavoidable. the significant decline in the use of paper maps has prompted many libraries to either put into storage or give away large parts of their collections. efforts to scan and make them more easily and universally accessible have breathed some new life into maps and played an important role in their preservation. the resulting raster map images also provide libraries with a tremendous, largely unrealized opportunity: mining these raster maps for their features of interest and converting those features into geospatial vector layers unleashes all kinds of possibilities for customized maps and spatial analysis using gis, not to mention easier discovery and the augmentation of library geospatial data collections. it is within the means of libraries to accomplish this using currently available software and following the steps and methods outlined above. references able software corporation. (2008). r2v user’s manual: advanced raster to vector conversion software. [online] available at: http://www.ablesw.com/r2v/r2vmanual.pdf [accessed 12 may 2017]. allord, g., fishburn, k. and walter, j. (2014). standard for the u.s. geological survey historical topographic map collection. [online] available at: https://dx.doi.org/10.3133/tm11b03 [accessed 4 may 2017]. 19/24 peller, peter (2018) from paper map to geospatial vector layer: demystifying the process, iassist quarterly 42 (3), pp. 1-22. doi: https://doi.org/ 10.29173/iq914 arteaga, m. (2013). historical map polygon and feature extractor. in: proceedings of the 1st acm sigspatial international workshop on mapinteraction. [online] new york, ny: acm, pp. 66–71. available at: dx.doi.org/10.1145/2534931.2534932 [accessed 18 apr 2017]. arteaga, m. (2017). [software]. map vectorizer. [online] available at: https://github.com/nyplspacetime/map-vectorizer [accessed 21 may 2017]. bracke, m., miller, c. and kim, j. (2008). adding value to digitizing with gis. library hi tech, [online] 26(2), pp. 201-212. available at: http://10.1108/07378830810880315 [accessed 19 apr 2017]. brown, k. (2002). raster to vector conversion of geologic maps: using r2v from able software corporation. [online] available at: https://pubs.usgs.gov/of/2002/of02-370/brown.htm [accessed 5 may 2017]. cartography associates. (2017). david rumsey map collection. [online] available at: http://www.davidrumsey.com/ [accessed 9 may 2017]. chiang, y-y. (2010). harvesting geographic features from heterogeneous raster maps. phd. university of southern california. chiang, y-y., knoblock, c. and chen, c. (2005). automatic extraction of road intersections from raster maps. in proceedings of the 13th annual acm international workshop on geographic information systems. [online] new york, ny: acm, pp. 267–276. available at: dx.doi.org/ 10.1007/s10707-008-00463 [accessed 24 mar 2017]. chiang, y-y., leyk, s., honarvar nazari, n., moghaddam, s. and tan, t. (2016). assessing the impact of graphical quality on automatic text recognition in digital maps. computers & geosciences, [online] 93, pp. 21–35. available at: dx.doi.org/ 10.1016/j.cageo.2016.04.013 [accessed 12 apr 2017]. chiang, y-y., leyk, s. and knoblock, c. (2014). a survey of digital map processing techniques. acm computing surveys, [online] 47(1). available at: https://doi.org/10.1145/2557423 [accessed 11 may 2017]. esri. (2016a). fundamentals of georeferencing a raster dataset. [online] available at: http://desktop.arcgis.com/en/arcmap/10.3/manage-data/raster-and-images/fundamentals-forgeoreferencing-a-raster-dataset.htm#guid-4deee2e1-e031-4cea-9318-8ca707ed31cb [accessed 5 may 2017]. esri. (2016b). getting started with arcscan: arcmap 10.4. [online] available at: http://desktop.arcgis.com/en/arcmap/10.4/extensions/arcscan/what-is-arcscan-.htm [accessed 11 may 2017]. https://pubs.usgs.gov/of/2002/of02-370/brown.htm 20/24 peller, peter (2018) from paper map to geospatial vector layer: demystifying the process, iassist quarterly 42 (3), pp. 1-22. doi: https://doi.org/ 10.29173/iq914 esri. (2016c). generalizing zones with nibble, shrink and expand. [online] available at: http://desktop.arcgis.com/en/arcmap/10.4/tools/spatial-analyst-toolbox/generalizing-zones-withnibble-shrink-and-expand.htm [accessed 12 may 2017]. esri. (2016d). what is raster data?. [online] available at: http://desktop.arcgis.com/en/arcmap/10.3/manage-data/raster-and-images/what-is-raster-data.htm [accessed 25 may 2017]. esri. (n.d.). gis for humanitarian mine action: georeferencing and digitizing web course. [online] available at: https://www.esri.com/training/catalog/57630434851d31e02a43ef72/gis-for-humanitarianmine-action:-georeferencing-and-digitizing/ [accessed 4 may 2017a]. esri. (n.d.). vector. in: gis dictionary, [online] available at: http://support.esri.com/other-resources/gisdictionary/term/vector [accessed 25 may 2017b]. fleet, c., kowal, k. and pridal, p. (2012). georeferencer: crowdsourced georeferencing for map library collections. d-lib magazine, [online] 18(11/12). available at: dx.doi.org/ 10.1045/november2012-fleet [accessed 5 may 2017]. godfrey, b. and eveleth, h. (2015). an adaptable approach for generating vector features from scanned historical thematic maps using image enhancement and remote sensing techniques in a geographic information system. journal of map & geography libraries, [online] 11(1), pp. 18-36. available at: dx.doi.org/10.1080/15420353.2014.1001107 [accessed 24 mar 2017]. jung, w.r. (2009). vector feature extraction using object-oriented image analysis techniques from scanned maps. m.a.western michigan university. lacroix, v. (2009). raster-to-vector conversion: problems and tools towards a solution a map segmentation application. in proceedings of the 7th international conference on advances in pattern recognition, icapr 2009, [online] pp. 318–321. available at: dx.doi.org/ 10.1109/icapr.2009.96 [accessed 24 mar 2017]. levachkine, s. (2004), raster to vector conversion of color cartographic maps. in: llados, j. kwon, y.,eds., graphics recognition. recent advances and perspectives. grec 2003. lecture notes in computer science, [online] 3088. berlin: springer, pp. 50-62. available at dx.doi.org/10.1007/978-3-540-25977-0_5 [accessed 24 mar 2017]. marciano, r., allen, r., hou, c. and lach, p. (2013), “big historical data” feature extraction. journal of map & geography libraries, [online] 9(1/2), pp. 69-80. available at dx.doi.org/10.1080/15420353.2012.732020 [accessed 24 mar 2017]. new york public library. (n.d.). “building inspector”, [online] available at: http://buildinginspector.nypl.org/ [accessed 17 may 2017]. 21/24 peller, peter (2018) from paper map to geospatial vector layer: demystifying the process, iassist quarterly 42 (3), pp. 1-22. doi: https://doi.org/ 10.29173/iq914 olson, j. (2009), which is it? scan or digitize. make up your mind!. journal of map & geography libraries, [online] 5(1), pp. 108-111. available at: dx.doi.org/10.1080/15420350802470913 [accessed 1 may 2017]. pearson, m., mohammed, g., sanchez-silva, r. and carbajales, p. (2013), stanford university libraries study: topographical map vectorization and the impact of bayer moiré defect. journal of map & geography libraries, [online] 9(3), pp. 313–334. available at: dx.doi.org/10.1080/15420353.2013.820677 [accessed 24 mar 2017]. southall, h., brown, n. and burton, n. (2003). digitising the inter-war land use survey of great britain: a pilot project. [online] available at: https://researchportal.port.ac.uk/portal/files/174820/digitising_lusgb_2003_pilot_project_report.pdf [accessed 24 march 2017]. whitfield, t. (2005). capturing and vectorizing black lines from greenline mylars. [online] available at: https://pubs.usgs.gov/of//2005/1428/whitfield/index.html [accessed 4 may 2017]. end-notes 1 peter peller is the director of the spatial and numeric data services unit at libraries and cultural resources, university of calgary and can be reached by email: ppeller@ucalgary.ca. 2 the term “digitize” is a generic term that is sometimes used to individually describe both scanning and vectorization processes which can be confusing (olson, 2009). to avoid confusion, for the rest of the paper the term “scan” will be used for the processing of creating a raster image from a paper map and the term “vectorize” will be used for the process of converting a raster image to a vector file. 3 in its simplest form, a raster consists of a matrix of cells (or pixels) organized into columns and rows where each cell contains a value representing information.” (esri, 2016d) a scanned map is a raster image. 4 vector is “a coordinate-based data model that represents geographic features as points, lines and polygons.” (esri, n.d.) information is associated with each vector feature. 5 heads-up digitizing is a process where lines and boundaries are manually traced with a mouse interactively on the computer screen using gis software. appendix 22/24 peller, peter (2018) from paper map to geospatial vector layer: demystifying the process, iassist quarterly 42 (3), pp. 1-22. doi: https://doi.org/ 10.29173/iq914 lead author arteaga bracke brown godfrey jung 1 jung 2 library project ✓ ✓ ✓ map type thematic thematic thematic thematic topographic topographic number 100s 1 multiple 1 multiple multiple feature(s) buildings soil classes geologic contacts, faults snow load zones contours, roads, rail, streams water, urban, vegetation 1. scanning resolution 600 dpi 300 dpi color depth 32 bit 24 bit output file type tiff tiff jpeg 2. georeferencing software map warper arcgis r2v arcgis definiens definiens 3. image enhancement software arcgis, photoshop operations: (b=blur, bi= convert to binary, c=convolution, cr=crop, m=mask, r=resample, s=sharpen, st=stretch) b, c, m, r, s 4. segmentation software gimp definiens r2v arcgis definiens definiens type: (rs=remote sensing, t=thresholding, o=other) t rs t rs rs rs rs type (ps=pixel supervised, pu=pixel unsupervised, os=object supervised os pu os os rs algorithm (h=heuristic, i=iso cluster, m=maximum likelihood, p=principle components analysis) h i h h output pixel features (l=lines, m=multiclass areas, s=single-class areas) m m l m l s 5. raster editing software r2v arcgis arcscan arcscan operations (c=convert to binary, e=erase, ex=expand, m=majority, ma=mask, mo=morphological operator, n=nibble, p=paint, s=shrink) e ex, ma, s mo mo 6. vectorization software gdal polygonize definiens r2v arcgis arcscan arcscan output (l=lines, p=polygons, po=polygon outlines) p p l p l p 7. vector editing software r, building inspector arcgis r2v arcgis appendix 23/24 peller, peter (2018) from paper map to geospatial vector layer: demystifying the process, iassist quarterly 42 (3), pp. 1-22. doi: https://doi.org/ 10.29173/iq914 operations(a=attribute data, c=clip, d=dissolve, e=eliminate, fp=feature to polygon, m=merge, me=manual edits, p=polygon exclusion, s=smoothing, sp=spatial join, ss=shape simplification) a, me, p, ss d, me, s, sp s c, d, s lead author marciano pearson southall 1 southall 2 whitfield library project ✓ map type thematic topographic thematic thematic thematic number multiple multiple multiple multiple multiple feature(s) neighborhoods contours land use classes land use classes geologic contacts 1. scanning resolution 330-440 ppi 200 dpi 200 dpi 300-400 dpi color depth 8 bit 8 bit output file type tiff tiff tiff 2. georeferencing software arcgis arcgis arcgis arcgis arcgis 3. image enhancement software imagemagick arcgis, photoshop, paintshop pro arcgis, photoshop, paintshop pro arcgis operations: (b=blur, bi= convert to binary, c=convolution, cr=crop, m=mask, r=resample, s=sharpen, st=stretch) c cr ,r, s cr, r, s bi 4. segmentation software arcgis imagemagick paintshop pro, arcgis erdas imagine scanner type: (rs=remote sensing, t=thresholding, o=other) rs t o rs t rs type (ps=pixel supervised, pu=pixel unsupervised, os=object supervised ps ps rs algorithm (h=heuristic, i=iso cluster, m=maximum likelihood, p=principle components analysis) m p output pixel features (l=lines, m=multi-class areas, s=single-class areas) m l m m l 5. raster editing software arcgis / arcscan imagemagick arcgis arcgis arcscan operations (c=convert to binary, e=erase, ex=expand, m=majority, ma=mask, mo=morphological operator, n=nibble, p=paint, s=shrink) c, e, ex, s c ex, m, n, s e, p 6. vectorization software arcscan potrace arcgis arcgis arcscan output (l=lines, p=polygons, po=polygon outlines) p l p p po 7. vector editing appendix 24/24 peller, peter (2018) from paper map to geospatial vector layer: demystifying the process, iassist quarterly 42 (3), pp. 1-22. doi: https://doi.org/ 10.29173/iq914 software arcgis arcgis arcgis arcgis operations(a=attribute data, c=clip, d=dissolve, e=eliminate, fp=feature to polygon, m=merge, me=manual edits, p=polygon exclusion, s=smoothing, sp=spatial join, ss=shape simplification) m e, me e, me a, me, fp vol282-3.indd iassist quarterly summer/fall 2004 55 by aaron k. shrimplin1 and jen-chien yu introduction over the past few years the electronic data center (edc) at miami university libraries has made it a priority to provide support to faculty who want to incorporate numeric data in their courses. to our thinking, too many students are uncomfortable using quantitative and numerical concepts to problem solve. new opportunities now exist for faculty to teach with numeric data and to provide hands-on data analysis in the classroom. typically, data and statistical analysis are taught in methods and statistics courses. however, it is of little benefi t to give students tools without showing them how these tools are useful for solving specifi c problems. the better approach may be to teach them quantitative skills in the presence of a classroom problem. if students are learning about income distribution and poverty, it might be of benefi t for them to be able to develop testable propositions about the working poor, for example, and employ a statistical technique to test their hypotheses. unfortunately, teaching students in this way is often problematic. there are often signifi cant barriers to teaching with numeric data. survey datasets are often very large and complex to use. moreover, most teaching faculty do not have the time and specialized skills necessary to prepare classroom datasets. this paper describes ways in which the edc has tackled these challenges.2 a window of opportunity despite these barriers, we have been able to work with faculty to improve students’ quantitative-reasoning skills. before we highlight a few examples of this collaboration, it is important to call attention to some of the factors that have made this partnership possible. here is a short list of infl uences we think have supported the use of numeric data in the classroom. re-inventing ourselves advances in information technology have given librarians the opportunity to re-invent themselves and grow professionally. as more library services are technologically-enabled, libraries are not only able to reshape and extend existing services, but also to create new services and products. in some cases, library activities are being converted to technological solutions, freeing up resources to focus on new activities. in short, the networked information revolution has afforded, if not necessitated, libraries and their staff the opportunities to step outside their traditional service boundaries and be proactive to the needs of the academic community. it is within this context that we have been able to work alongside faculty in making a direct contribution to student learning. active learning ‘sage on the stage’ is out and ‘guide on the side’ is in. this concept of learning focuses attention on the student’s experience. faculty who adopt this approach to learning and teaching tend to focus on students’ learning outcomes and competencies. this approach, not only changes the way in which faculty teach, but also the way in which students learn. that is, rather than passively accumulating knowledge, students are asked to approach the learning experience differently; they are asked to become more activelyinvolved in the learning process and to demonstrate their ability to use the knowledge that they have acquired. in this environment, librarians who speak with faculty and specifi cally address the learning outcomes important to student success are rewarded with a more active role in the learning process. opportunities then present themselves for the integration of quantitative reasoning skills into the curriculum. innovative tools web-based tools have made it possible for students to query large datasets in ways not possible a few years ago. data analysis and retrieval tools like nesstar, sda, and dataferrett have opened up access in unprecedented ways.3 while making effective use of numeric data in teaching and learning requires specialized skills and a considerable amount of time for preparation, barriers which inhibit the use of data in the classroom and in student projects are being lowered. collaboration recognizing that the conditions were right for a proactive approach to working with faculty to get numeric data into the classroom, we went on a listening tour to better understand the types of data support faculty require in focusing in on student learning outcomes: how sda helped us get data into the classroom 56 iassist quarterly summer/fall 2004 order to teach quantitative skills. as we talked with faculty one-on-one, it became clear that a partnership was in the making. although the vocabulary may have differed, it became obvious that we shared similar points of view: that statistical literacy is important and that students need to be aware of datasets and how to use them. it also became clear that we were two sides of the same coin. faculty had the subject expertise and learning objectives; we had the knowledge of datasets and the tools to explore them. most of the faculty we spoke with wanted similar things. they were looking for ways to: • marry theory and method in an active learning environment; • make tradition courses more analytical; • explore innovative pedagogies, such as creating self-directed learning paths; • improve access to diffi cult-to-use datasets with a user-friendly interface; • add new, exciting features to existing courses. at times it felt as if we were writers making a pitch for a sitcom. time was short and we had to make a lasting impression. our secret weapon was the web-based tool developed at berkeley called survey documentation and analysis, otherwise known as sda. in a few minutes, we could show a professor a customized html codebook and a user-friendly interface for basic data analysis and exploration. we often demonstrated a dataset important to his discipline. professors were impressed with sda for the same reasons it had impressed us. it could be run with a web browser, it made data analysis accessible and available to students, and there were no proprietary, platform-specifi c software that needed to be installed. moreover, students could investigate a basic question using cross-tabulation or another technique without having to learn a high-level language. for those situations where a statistical software program like sas, spss, or stata were appropriate, sda let users download customized subsets with data defi nition fi les for these programs. discovering variables through customized html codebooks was also very useful, particularly for large datasets. examples over the last two years, the edc at miami university libraries has delivered approximately three dozen datasets to the web using sda. the majority of these datasets have been for three departments: gerontology, economics, and political science. gerontology the edc developed a webpage to supplement in-class instruction for the course linking research and practice in gerontology. the course is designed to give graduate students an overview of social research in gerontology. students review the basic assumptions, models, processes, and problems of social research. they also take a more detailed look at applied research, including program evaluation. using our webpage, students learn about existing datasets and gain hands-on experience by performing exploratory analysis and extracting subsets of observations and variables. the subsets are then downloaded for more in-depth analysis using various statistical packages. the following us-based datasets have been delivered to the web with sda and are available for use: • the second longitudinal study of aging,1994-2000: wave 3 survivor file; • the second longitudinal study of aging, 1994-98: wave 3 survivor file; • the second longitudinal study of aging, 1994-98: wave 2 descendent file; • evaluation of long-term care initiatives in ohio; • longitudinal study of aging, 1984-90; • national long-term care survey, 1994; • national nursing home survey 1995, 1997, and 19994. economics the bulk of our work with the economics department has been centered on the current population survey’s (cps) annual demographic files5. to date, we have created datasets for the years 1997 through 2001. access to these public-use fi les has always been diffi cult. perhaps the biggest obstacle to using these data is the sheer size of the surveys. for example, the annual demographic file for 1998 has over 650 variables and over 135,000 individual respondents. a student wanting to use these data to investigate a hypothesis would have to download the raw data fi le, write a data extraction program to obtain the variables needed for the research, and then harmonize the data defi nitions for different years. we used sda to deliver these data fi les to the web to improve access and to support faculty who wanted to incorporate hands-on data analysis into their undergraduate courses. one of the key skills that students should learn in social science classes is the ability to think critically about the ways in which raw data are processed to support analysis. the cps annual demographic files are the most commonly employed dataset for social science analysis of income distribution, labor force behavior, and poverty in the us. in particular, these data fi les contain data for many of the standard government reports on these topics. the problem is that typically students are presented only with analyzed results, despite the fact that many of the key issues of analysis are refl ected as much in what measurements one undertakes as in how those iassist quarterly summer/fall 2004 57 measurements turn out. sda lets students ask their own questions of these data using simple statistical analysis and tabulations. political science a number of different datasets have been created with sda to support political science courses. in most cases, these datasets have served to make introductory courses more analytical while promoting active learning for students. a few sample exercises illustrate how students of politics are using sda to gain a glimpse into the world of social science research: • we discussed presidential approval and the increasing importance of the president to be an effective economic manager. what percent of respondents who favored clinton’s economic performance approved of his general job performance? • we discussed the concept of political socialization. what are the agents of socialization and how has socialization changed over time? • are people less interested in politics today than in the past? • political scientists have noted emerging “gender gaps” over the past several decades. examine the change in party identifi cation among women and men over time. conclusion we have enjoyed stepping outside our traditional service boundaries and working alongside faculty in making a direct contribution to student learning. whether it’s making a gerontology graduate course more quantitative by providing online access to core datasets or making it possible for an economics professor to let his students ask their own questions of a dataset, the edc will continue to make access to numeric data more easy to use by both faculty and students and is committed to supporting faculty who are interested in incorporating data analysis into their courses. notes 1 contact: aaron k. shrimplin, miami university libraries, oxford, ohio 45056. phone: 513.529.6823. internet: http:// www.lib.muohio.edu. email: aaron@lib.muohio.edu 2 this paper was presented at the iassist conference held in ottawa in may 2003 in the session on “advancing research and data literacy: empowering users”. 3 survey documentation & analysis (sda) available at: http://sda.berkeley.edu; nesstar (networked social science tools and resources) available at: http://www.nesstar. org; dataferrett: for thedataweb available at: http:// dataferrett.census.gov/thedataweb/index.html; internet, accessed 22 march 2005. 4 e.g., u.s. dept. of health and human services, national center for health statistics. the second longitudinal study of aging, 1994-2000: wave 3 survivor file [computer fi le]. hyattsville, md: u.s. dept. of health and human services, national center for health statistics [producer & distributor], 2002 scripps gerontology center. evaluation of longterm care initiatives in ohio [computer fi le]. oxford, oh: miami university, scripps gerontology center [producer], 1997. oxford, oh: miami university, scripps gerontology center/electronic data center [distributor], 2003. u.s. dept. of health and human services, national center for health statistics. national health interview survey: longitudinal study of aging, 70 years and over, 19841990. [computer fi le]. hyattsville, md: u.s. dept. of health and human services, national center for health statistics [producer], 1984. ann arbor, mi: inter-university consortium for political and social research [distributor], 1988. manton, kenneth g [principal investigator]. national long-term care survey, 1994 durham, nc: duke university, center for demographic studies [producer & distributor], 1994. e.g., u.s. dept. of health and human services, national center for health statistics. national nursing home survey, 1999 [computer fi le]. hyattsville, md: u.s. dept. of health and human services, national center for health statistics [producer], 1999. ann arbor, mi: inter-university consortium for political and social research [distributor], 2001. 5 ee.g., u.s. dept. of commerce, bureau of the census. current population survey: annual demographic file, 2001 [computer fi le]. washington, dc: u.s. dept. of commerce, bureau of the census [producer], 2001. ann arbor, mi: inter-university consortium for political and social research [distributor], 2001. iassvol201 4 iassist quarterly facilitating access to comparative data by ekkehard mochmann & lorenz gräf1, central archive for empirical social research (za); at the university of cologne mandate of the za as part of a social science infrastructure the central archive for empirical social research (za) at the university of cologne serves as a research, training and resource center for social research. founded in 1960 by the faculty of economics and social sciences of the university of cologne, it soon developed into a data service with a supraregional and international clientele. as a central node in the international data service network it became the starting point for a more comprehensive social science infrastructure, the german social science infrastructure services (gesis e. v.). this association was created as a response to needs formulated by the social science profession in 1986 to provide infrastructural services in all fields of social research with particular emphasis on: • collecting data and making it available for further research. • informing about social science literature and research projects. • development of research methods, teaching instruments and methods consulting for research projects. the core of the za mandate is to facilitate access to already existing data, especially survey data, which can be used for secondary analysis. the holdings cover all fields of empirical social research. beyond survey data there are collections of statistical data, regional data, various types of quantitative historical data and machine readable texts for computer assisted content analysis, as well as party manifestos and other text collections. za provides services in the area of acquisition, processing, documenting and making available data for social research, especially survey data. za offers consulting services for secondary analysis. training in complex analysis methods takes place twice a year in the za spring seminar for empirical social research and the autumn seminar for quantitative historical research. beyond this, za creates ex post statistical time series and supports comparative international studies for the analysis of long term social developments . za holdings of empirical social research data include european time series and comparative studies. the za department zhsf (center for historical social research) develops data bases, in some cases going back to earlier centuries. the za holds nearly 4000 data sets and data collections. even though there is no particular topical restriction, emphasis is on topics such as political attitudes, election studies, education, unemployment, leisure and occupation, media and the environment. among the data sets intensively used are the eurobarometers (a data pool of comparative surveys from european countries taken for more than 15 years), the german general social survey allbus, which is conducted every two years, the international social survey program (issp) for 25 countries from australia, america, europe to japan. similar attention is paid to the monthly politbarometer series provided by the research group elections (fgw: forschungsgruppe wahlen) which is also presented on the second public tv station (zdf) every month and the collection of surveys to the national parliament (bundestag) since 1949. a gesis branch in berlin is now focusing on data and information transfer from and to eastern europe. recently more than 400 data sets from surveys conducted in the former german democratic republic (gdr) since 1975, were included in the za holdings and were processed for secondary analyses. currently emphasis is on supporting initiatives to create infrastructure institutes in eastern europe and to develop a service network for european wide data transfer. the za has access to data held in the social science data archives world wide. international data transfer is coordinated with the council of european social science data archives (cessda) and the international federation of data organizations for the social sciences (ifdo). access to internationally distributed data bases is supported by making use of modern telematic services like wais, www, ftp on the internet and other computer networks. selecting relevant data and solving methodological problems relating to secondary analysis is an essential part of individual consulting. the newsletter za information and the journal historical social research (hsr) inform about new data 5winter 1996 sets, methodological developments, research findings and conferences. a documentation of more than 1000 empirical research projects conducted in germany, austria and switzerland is published annually. organizational priorities in collecting and distributing data from the very beginning the za philosophy was to develop services in close interaction with the scientific community. as a consequence of this philosophy za also supports a small research and training department which focuses on new methodological developments in data collection and analysis. under a guest professor scheme scholars from abroad are supported in their research from planning new surveys to secondary analysis of available data. experts in data management and analysis offer advice from the selection of appropriate data to advanced statistical analysis. already in the 60s erwin k. scheuch, one of the za founding fathers, created a climate for comparative research which was inspired by the standing committee for comparative research of the international social science council, in which he cooperated with stein rokkan and warren miller. this orientation was enforced by the emerging european unification and the globalization of social research. over the years several international research projects have chosen the za as their resource center for creating an integrated data bases. integrating national data sets into internationally comparative data sets includes comprehensive documentation of methodological, technical and historical background of a study and additional interpretation knowledge to facilitate further comparative analysis. currently za serves in this function for the eurobarometers (jointly with icpsr and swedish data services (ssd), the international social survey program (issp), and the major election studies to national parliaments in europe (icore)2 . bringing together researchers working on the data and the data management experience of za provides a unique working environment for creating an integrated fund of knowledge on core topics of european social development. in cooperation with the principal investigators and other european data services za coordinates and creates european data bases, which could otherwise not be made available to the scientific community, relying just on national resources. za strongly supports a policy of labor division between european archives according to topically focused european data collections. under tightening resources this is a must for integrating the european data bases. inspite of intellectual and political efforts there is an ongoing demand for additional european resources to achieve what cannot be covered by the subsidiarity principle: the data service capacities are by and large absorbed by the national demands and there is little leverage to cope with additional international workloads. direct access to the expertise and information banks on social science literature and research projects, as well as to the methodological expertise of its gesis partner institutes complement this infrastructural support for the production and analysis of comparative data bases on europe. the za user survey although the central archive has always made efforts to communicate with its clientele we have found it necessary to get more information about our clientele to face the rapid technological changes which are taking place. in the past, information concerning needs and demands of the clientele were mostly gathered by mail surveys. this procedure involves three major drawbacks. first, only the users of the institution are surveyed. so we would miss the comments of those researchers who did not make use of the data services. second, the findings are often biased because only the most motivated people contribute in these surveys. third, the response rates in mail surveys are low. the installation of a laboratory for telephone surveys at the university of cologne last autumn gave us the opportunity to avoid these drawbacks. in a pilot study to test this telephone facility we could interview social researchers about their research environment and about their impression of the central archive. description of the sampling procedure the target population of this study were all social scientists engaged in empirical research. for this purpose we defined empirical social research as a quantitative approach which is done with the methods of empirical social research, mainly interviewing, observation and content or document analysis (cf. obershall 1972) since there is no list of scientists using this very approach in their research, we could have started our project with a list of institutions known to us as informants for our documentation. but this procedure would have led to some sort of snowball sample resulting in an unpredictable sample structure. furthermore, we wanted to interview even those people who do social research but do not want to appear in our documentation so they do not inform us about their work. eventually, we came to the conclusion that a sample drawn out of the subscribers of the za newsletter would suit our needs best, for they may be assumed to be highly interested in the application of the methods of social research. it was equally important for our purpose that fifty percent of the za users were also subscribers to the newsletter. sampling under the subscribers of the za newsletter gave us the opportunity to get the 6 iassist quarterly »research about sampling 2,765 subscribers of za newsletter 1,375 addresses of social researchers 1,258 telephone numbers 762 interviews conducted (68.4% response rate) 1,114 correct telephone numbers 538 interviews with social researchers scientists who are not familiar with the za services scientists who made use of za service: za users scientists who never ordered a za service random sampling identification of telephone numbers screening (281) 52.9% (225) 42.4% (25) 4.7% the central archive (za) user during interviewing proved to be correct figure 1figure 1 7winter 1996 feedback of those who had already made use of the services of the central archive, and of potential users as well. we planned to collect about 500 interviews. in advance, we estimated that about 25% of the subscribers were not directly involved in empirical research, e.g. librarians or local staff of university computer centers. from over 2,500 subscribers of the za newsletter we drew a sample of 1,375 people. since we only had their addresses we had to find out their phone numbers. using a telephone directory cd-rom and directory inquiries we managed to locate more than 1,100 potential respondents. the survey took place between nov. 28 and dec. 5, 1995. we conducted over 700 interviews. more than 200 respondents were not engaged in empirical social research so we finished with a sample of 538 social researchers. figure 1 shows the details of the sampling procedure. the interviews consisted of three parts. the first part dealt with the institutional affiliation of the researcher. the description of the actual empirical work formed the content of the second part and finally the respondents were asked questions concerning the performance of the central archive. in this paper we will focus on the description of the research community and the za clientele. characteristics of the za clientele empirical social research is done in a variety of disciplines. the readers of the za newsletter are heavily inclined to sociology as shown in figure 2. two in five researchers (38.0%) belong to an institute which is situated in the field of sociology. the relevance of sociology is outstanding. it is mentioned nearly three times as often than is psychology (13.3%), which ranges second. next follows a group of three fields with a proportion of ten percent each: economics, political science and education. these five subjects together form the core of the social sciences. medicine, communication studies and subject area readers of za newsletter za clients 22.7 0.7 3.5 5.0 5.7 5.7 5.7 10.7 14.2 48.2 26.6 3.3 5.7 9.4 13.3 5.9 5.9 9.2 9.4 38.0 in % other statistics market research education psycholog medicine communication studies economics political science sociology 0 10 20 30 4 5 n=458 multiple response figure 2 8 iassist quarterly market research gain nearly 5% each. statistics ranges last in this list of disciplines. there were 27 other subject fields named by the respondents, like geography, history, criminology, social psychology etc. but none gained more than two percent. let us now focus on those interviewees who have formerly received data sets from the central archive. among those people nearly 50% belong to an institute of the research field of sociology. the second most important discipline is political science. psychology which was second among the readership of the za newsletter now follows in the fourth position. this is due to the fact that psychologists do not deal that much with survey data. they prefer experimental data mostly collected from college graduates. they adopt the methods of empirical research but they normally do not need nationwide survey data. political scientists on the other hand very often look for election data or data concerning the nationwide electorate. so it is not surprising that they range second as users of the za data service. ranging third among the users of the za data service are researchers belonging to institutions in the field of economics. they gain a proportion of nearly ten percent. psychology, education, communication science and medicine gain 5% each. obtaining data from the central archive is of less importance for people belonging to market research institutes and to the statistics branch. the former do not care much about surveys carried out by other scientists and the statisticians do not seem to be in particular demand of survey data. as shown in figure 3 two thirds of the users of the central archive work in an academic institutional background. 20.7% of the respondents are employed in publicly financed research institutes. they consist mainly of federal research agencies, like the bundesinstitut für bildungsforschung, or governmentally financed large scale institutes, like the max-planck-institute or the wissenschaftszentrum berlin für sozialforschung. only a small number of them are private nonprofit organizations like the konrad-adenauer-stiftung. 13.9% of the readers of the za newsletter work in private organisations in the commercial sector, mainly within the field of market research. there are only slight differences in the percentage between the researchers who have made use of the za data service and those who have not. while emphasis is on providing services for the academic community, the clientele also includes researchers from public administration and the media. in the za user survey people institutional affiliation readers of za-newsletter za-clients 2.1 10.8 24.3 62.8 5.1 15.7 20.7 60.3 in % none private organizations governmentally financed or nonprofit organizations universities 0 10 2 30 40 50 60 7 figure 3 9winter 1996 who work in the media are underrepresented to some extent. they normally do not subscribe to the za newsletter because they are less interested in methological issues. the survey focused on people who are actually doing social research. but a remarkable part of the za clientele is mainly interested in getting information about the distribution of attitudes in the population. for what purpose do the clients use the data? we asked those persons who had at least once received data from the cologne archive what purpose they followed in examining the data. the question was posed as open ended question and respondents could give multiple responses. nevertheless we got a clear-cut picture. there are two main intentions behind the ordering of data. one third of the population uses the data as a source for secondary analysis under a new research question. another third employs the data as a supplement to own data sets. this completion was mostly sought in time dimension. in most cases this means that researchers who have already got data at the present stage want to make comparisons concerning the same population at some former point in time. the intention to conduct an international or intercultural comparison as a supplement to their own data was mentioned by a smaller fraction. 10.7% of the researchers used the data for teaching purposes in class and 6.4% used the data in order to evaluate indicators used by other researchers. another 10% named other intentions, like information about the distribution of certain opinions in society, compilation of dissertation thesis etc. internet as a research tool in the near future the already heavy use of the internet by the research community will drastically change and hopefully improve the conditions for doing scientific research. computer mediated communication via email will expand the possibility to collaborate and interact with distant colleagues. the flow of information will accelerate and increase when scientists begin to use virtual arenas (multi user dungeons, news groups, mailing lists, virtual conferences, electronic journals etc.) to discuss and distribute new ideas. with more scientists using the internet the demand for quick and easy access to socio-economic data will increase. therefore it is vital for data archives to know how many of their clients have access to the internet and for use of data as a source for secondary analysis 35.0% data presentation in class 10.7% evaluation of indicators 6.4% other 10.0% complementary to own data collection 32.9% both in combination 5.0% n=140 figure 4 10 iassist quarterly how many of them it has already become an ordinary research tool. in november 1995 when we asked german social scientists 56.7% of them had direct access to the internet from their working place and 35.2% made frequent use of it. this low percentage indicates that the internet-revolution has still to gain ground in germany. only one third of the subscribers of the za newsletter have adopted the internet as a research tool. broken down into institutional affiliation we find a big difference between researchers working in the academic context and those who work outside university. already 68.4% of the respondents belonging to university institutes have access to the internet. this figure is nearly twice as high as in non-academic institutes. in private organisations only 32.2% of the employees dispose of a direct connection to the internet. in governmentally financed and nonprofit institutes the adoption rate is higher and amounts to 42.7%. access frequent use ratio n (use/access) university 68.4 44.1 64.5 (320) government / nonprofit 42.7 26.1 61.1 (119) industry 32.2 15.1 46.9 (73) no institutional affiliation 33.3 26.7 80.2 (15) 56.7 35.2 62.1 (527) table 1: access and use of the internet by institutional affiliation access to the internet does not imply that scientists make use of the internet in their daily work. only two thirds of the researchers with access to the internet adopt an internet based service as an ordinary research tool. there is still a lot of hesitation in exploring the usefulness of the internet. the ratio of use to access of the internet is nearly the same in university institutes and in governmentally financed institutes. but in the industrial context only one half of the people with direct access to the internet make frequent use of email, www, ftp or some other internet service. the situation seems even worse if we look at the percentage of scientists in the industrial context using the internet. only 15.1% of them mention the internet as a useful research tool. in germany, at the time of our survey, the internet was still an academic challenge. but we suppose that the low adoption-rates in the industry sector are only due to the fact that the internet is basically an academic invention. with a time lag of a few months we expect that researchers in non-academic institutes will use the internet with the same frequency as their colleagues in university institutes. it is often assumed that one of the major impacts of the use of internet-services will be a gap between young and skilled persons who adopt the new technologies quickly and older people who will be excluded from the new information technologies (cf. negroponte 1995). in our study we do not find support for the thesis of a widening gap between the generations. we do find differences between young and older scientists in the access-rates but there are no differences in adoption-rates. since the internet has not yet arrived in non-academic institutes we confine the analysis of this thesis to institutes in universities. as shown in table 2 nearly three quarters (73.8%) of the young scientists (age < 40) have access to the internet. among the older scientists (age ≥ 40) 65.6% dispose of a direct access. if we focus on those people in universities who dispose of a direct access to the internet the percentage of frequent users among young scientists is nearly the same as among older scientists. 65.9% of the young scientists make frequent use of internet-services compared to 64.0% of older scientists. thus the adoption-rates in the two age-groups are almost identical. if there was an effect of age on adoption we would expect a much higher adoption-rate among the younger scientists. we can conclude from these findings that differences in the percentage of internet-users between age-groups are only the result of differences in access-rates. presumably older people get access to the internet at a later stage of the innovation-process than younger people. but if there is a direct access to the internet the same fraction of researchers will use the internet in the older and in the younger generation. this indicates that the use of internet-services is already a valuable research tool and that it depends mostly on the institutional context in which the scientist works whether he adopts internet-services or does not. but we can expect that in the near future the use of internet-services will be as natural as that of personal computers is now. therefore archives have started to prepare themselves for the coming internet age. 11winter 1996 access frequent use ratio n (use/access) under 40 years 73.8 48.6 65.9 (107) 40 years and older 65.6 42.0 64.0 (212) only university institues table 2: access and use of the internet by age desiderata and recommendations of the za clientele at the end of the interview we asked the respondents if there was anything that the central archive should improve or which services should be introduced. a large fraction of the respondents commented on the information policy. they wanted information to come more frequently and more directly to their working place. another group recommended to give more detailed information. some researchers gave the advice to foster the effort of addressing people outside the core-disciplines of social sciences. the second main topic was the dissemination of data. many of the users wanted to have quick and easy access to the data via ftp and to have more data sets made available on cd-rom. some mentioned the present pricing policy and expressed their wish for reduced charges for data access. as the third main topic some users pointed to topically focused data collections and to a better and easier access to international data. some researchers would be glad if we could offer more surveys from the field of commercial market research and if we could offer more recent data. facilitating access to comparative data using the internet and publishing on cd-rom the central archive has always made the effort to expand its services and to use new technology to disseminate data and to communicate with its clientele as shown in the first chapter of this paper. in response to the answers the researchers gave in the user survey the central archive will strengthen these efforts. we will spread information about new data sets and other relevant news through a mailing list. more detailed and always up-to-date information can be found on our web pages (http:// www.za.uni-koeln.de/) just now and will be developed further in the near future.the question text, codebook information and marginal distributions of the international social survey program (issp) e.g. are searchable in the internet under wais. soon data will be accessible by ftp-transfer. furthermore we will enlarge our collections of data sets available on cdrom. third we are engaged in a multinational project, named ilses which aims at the development of an integrated libraryand dataservice. finally, we have installed a scientific laboratory equipped with all the infrastructure needed for comparative research. also, the european data archives are creating a virtually integrated catalogue of their holdings, accessible via internet. social research labs / large scale facilities as we start aging in the virtual scientific community we learn that the dream of information and data traveling to any place in the world is becoming true, yet it does not provide the ideal research environment for comparative research. researchers may be well informed about major events in their societies that might have had an impact on attitudes and behavior of respondents. the further we progress in time, the more interpretation knowledge must be transferred to the collective memory of researchers in order to provide the context that was decisive in the phase of data collection. this is particularly relevant for information about other societies which are not part of the daily information routine of the researcher. contextual information, cultural background and historic knowledge which may be necessary for sound interpretation of empirical evidence do not automatically travel with the collection of data sets from different societies . bringing together relevant data is still an exercise in systematic selection of comparable variables, data recoding and overcoming transboarder data flow hurdles emanating from data protection and data access regulations. a response to the needs of comparative research may be social science data labs, in which all relevant data and information for a particular research field is at the fingertips.over the past two years za has created a eurolab which provides access to major comparative studies and related background material (e.g. party manifestos, media-reports, event data bases, fact books etc.). the study collections include among others the international social survey program, the eurobarometers and major election studies on national parliaments in europe. the standing committee for the social sciences of the european science foundation had pointed to the need for better integration of the european data base and brought to the attention of the european union that social science data bases are the equivalent to large scale research instruments of the natural and technical sciences. a study panel proposed to acknowledge 12 iassist quarterly the need of social research for large scale facilities where researchers not normally having access, could come to profit from available resources. the institute for social sciences in essex and the zentralarchiv in cologne received recognition as first large scale facilities in europe under the training and mobility program of the eu. this will allow to cover travel and subsistence costs for scholars from eu member and associated states who want to make use of these resources subject to approval of their applications. over the next three years this will allow the za to have scholars and research teams not only making use of the resources, but at the same time enjoying truely comparative research by bringing together their specific knowledge about different countries. thus they can help to validate data and background material. ultimately this will improve the research resources for the scientific community at large, since validated data and background information may be compiled in knowledge basis for general distribution. references: lazarsfeld, paul f., 1962: the sociology of empirical social research; in: american sociological review, volume 27, no. 6 lazarsfeld, paul f., 1972: foreword; in: anthony oberschall (ed.): the establishment of empirical sociology: studies in continuity, discontinuity, and institutionalization; new york marcson, simon, 1972: research settings; in: saad z. nagi and ronald g. corwin (ed.): the social contexts of research; london negroponte, nicholas, 1995: being digital, new york oberschall anthony, 1972: introduction: the sociological study of the history of social research; in: anthony oberschall (ed.): the establishment of empirical sociology: studies in continuity, discontinuity, and institutionalization; new york zuckerman, harriet, 1988: the sociology of science; in: neil j. smelser (ed.): handbook of sociology; newbury park, london, new delhi 1.paper presented at the annual meetings of iassist, may 15, 1996, minneapolis, minnesota. 2. international community for research into elections and representative democrarcy 1/13 magnuson, diana l. (2024). the ipums business process model: instituting a workflow mapping strategy to support archival processes, iassist quarterly 48(4), pp. 1-13. doi: https://doi.org/10.29173/iq1130 the creative commons-attribution-noncommercial license 4.0 international applies to all works published by iassist quarterly. authors will retain copyright of the work and full publishing rights. the ipums business process model: instituting a workflow mapping strategy to support archival processes diana l. magnuson1 abstract the ipums preservation archive is instituting a workflow mapping strategy to further identify ipums process and metadata capture points to expand its holdings in the data archive. drawing on two business process models, the generic statistical business process model (gsbpm) and the generic longitudinal business process model (glbpm), archival staff have created an ipums business process model (ipums bpm). the ipums bpm reflects the use of secondary data sources and the work of harmonization and integration to create a data infrastructure that supports research across time and space. internally, the ipums bpm provides a clear visualization of the ipums workflow from external submission of data, harmonization process, documentation, extraction systems, and archival preservation of metadata. the challenge for archival staff is furthering the understanding and adoption of the ipums bpm within the ipums project groups, and to identify metadata production points that require the intervention of the archive for provenance and preservation purposes. it is part of an on-going effort to clearly define the role of the archive within ipums as an integral part of ipums organization and workflow. this paper identifies the value of instituting this mapping approach to gain a clearer understanding of the role of the archive within project work cycles, points where production and preservaton activities intersect, and opportunities to expand archival holdings. keywords business process model, archive, metadata, preservation introduction ipums at the university of minnesota has created the world’s largest accessible database of census and survey microdata, and geographic summary tables.2 the primary work of ipums is data harmonization—making census and survey data compatible across time and space. ipums integration and documentation makes it easy for researchers to study change, conduct comparative research, merge information across data types, and analyze individuals within family and community contexts. as of this writing, the ipums suite of products contains nine harmonized data collections.3 international data comes from over one hundred national and regional statistical organizations. all data and documentation are freely available to the global public. https://doi.org/10.29173/iq1130 https://creativecommons.org/licenses/by-nc/4.0/ 2/13 magnuson, diana l. (2024). the ipums business process model: instituting a workflow mapping strategy to support archival processes, iassist quarterly 48(4), pp. 1-13. doi: https://doi.org/10.29173/iq1130 the ipums preservation archive has historically functioned as a secondary unit to the main ipums data harmonization product line. as such, our data curation work was often operating in “stealth” mode from the perspective of ipums data product managers (magnuson 2015a). the first ipums product, what is now known as ipums usa, was launched in 1993 (magnuson and ruggles 2022). ipums international followed in 1999, and formal agreements with partner international statistical organizations committed the minnesota population center (now the institute for social research and data innovation) to curate and preserve tens of thousands of ancillary materials used in support of ipums international data harmonization work (ruggles et al. 2003b). in 2001 the minnesota population center hired data curator wendy thomas and under her guidance a nascent manuscript curation workflow began to take shape, culminating in the launch of a public-facing document access system in 2023 (magnuson 2024). 2016 was a watershed year for ipums archival preservation work. the minnesota population center (mpc) reorganized into the institute for social research and data innovation (isrdi), with ipums becoming its own center within the new structure. at this point in time ipums “clarified its mission in terms of data harmonization, access, curation, and preservation” and developed institutional guidelines around data product versioning, preservation, and assigning dois (magnuson and thomas 2023). the role of the archive during this realignment was enhanced but still outside the immediate vision of most internal ipums stakeholders. the effort to achieve coretrustseal (cts) certification, motivated by our institutional realignment and internal commitment to establishing clear versioning rules around our data products and registering dois, began in 2018 and culminated successfully in 2023.4 cts certification required us to pull together policy documentation that was scattered, incomplete, out of date, or simply non-existent for practices already in place. building ipums policy documentation to complete the cts application, especially around our preservation practices, clarified our institutional strengths and identified areas to refine. for the ipums preservation archive, this process documentation provided the blueprint for future action. further, it required ipums to clarify the role of the ipums preservation archive in the overall organization and workflow of ipums. this is an on-going process. the archive is reflected in the ipums workflows of the trusted repositories section of ipums.org but it is not reflected in the organization or staff pages.5 this paper may be of interest for groups that, like the ipums preservation archive, are situated in an organizational setting where preservation work is vital but secondary to the main product. this paper will demonstrate the power of a clear business process model for developing archival goals in an organizational setting in which the archive function is vital but secondary to the main product. expanding ipums preservation archive work four areas of institutional maturation across thirty years of ipums institutional history encouraged the administrative decision to expand the work and holdings of the ipums preservation archive. https://doi.org/10.29173/iq1130 3/13 magnuson, diana l. (2024). the ipums business process model: instituting a workflow mapping strategy to support archival processes, iassist quarterly 48(4), pp. 1-13. doi: https://doi.org/10.29173/iq1130 first, the steady growth of the number, size, and complexity of ipums projects over its thirty-year history necessitated the clarification, expansion, and reconfiguration of ipums archival work. with funding from both the university of minnesota and the national science foundation (nsf), the social history research laboratory at the university of minnesota converted existing public use samples of the u.s. census from 1880 to 1980 (excluding 1890 which was destroyed by fire in 1921) into a single coherent series with extensive documentation. the resulting ipums data series was first disseminated through an anonymous ftp site, and the first dataset was downloaded on november 19, 1993 (magnuson and ruggles 2022). over time, the ipums data integration project expanded to include other u.s. microdata sources: the current population survey (ipums cps), the american community survey (in ipums usa), the national health interview survey and medical expenditure panel survey (ipums health surveys), and survey data on scientists and engineers (ipums higher ed). ipums international (first data release in 2002) harmonizes and disseminates census and survey data from around the world, presently partnering with 103 countries. global health data are harmonized and disseminated as ipums dhs (demographic and health surveys), ipums mics (multiple indicator cluster surveys) and ipums pma (performance monitoring for action). the data collection incorporates historical full-count international census data of the north atlantic population project (naap, first released in 2001; now included in ipums international) and ipums usa full count data (1790-1950). aggregate data are harmonized and disseminated through the national historical geographic information system (ipums nhgis), ipums terra (combining global population, land use, and environmental data; now decommissioned), and the international historical geographic information system (ipums ihgis). another aggregate data collection, the contextual determinants of health (cdoh) provides access to measures of disparities, policies, and counts, at the state and national level, for historical marginalized populations. ipums time use has harmonized data from time diary surveys: american time use survey (atus), american heritage time use study (ahtus), and multinational time use study (mtus). second, increasing expectations of funding organizations regarding ipums’ preservation practices motivated internal effort to align with external standards. ipums is currently supported by a variety of funding agencies and foundations including the national institutes of health (nih, nia), the national science foundation (nsf) and the bill and melinda gates foundation.6 ipums’ commitment to preservation, discoverability, documentation, and dissemination goes beyond our grant promises. we recognize that the unique resources ipums has created and maintains need to be accessible long into the future for new research, scholarly replication, and for unanticipated creative uses of the data. the added impetus of evolving funder expectations around preservation increased the institutional visibility of the archive and its professional concerns. third, the developing expertise of ipums it teams encouraged and supported the decision to expand the ipums preservation archive. across the thirty years of ipums institutional history, a unique partnership between researchers and technologists grew in which systems and tools were collaboratively developed to support ipums data and metadata harmonization, documentation, and dissemination (ruggles et al. 2023, ruggles et al. 2003a, ruggles et al. 2003b, esteve and sobek 2003, fitch and ruggles 2003, block and thomas 2003, hall et al. 1999, ruggles et al. 1996). the ipums it https://doi.org/10.29173/iq1130 4/13 magnuson, diana l. (2024). the ipums business process model: instituting a workflow mapping strategy to support archival processes, iassist quarterly 48(4), pp. 1-13. doi: https://doi.org/10.29173/iq1130 team grew from graduate student support in the 1990s, its first full-time hire in 2000, to (currently) seventeen full-time staff of software developers, ux/ui specialists, data engineers, operations staff, and managers (fabrizio 2023, magnuson 2015b). the ipums it team of designers, developers, and system administrators “collectively build all of the systems required to produce and disseminate the many ipums data products.”7 the longevity of ipums it staff (twice the industry average) is testimony to the vitality of this dynamic and productive partnership between researchers and technologists (fabrizio 2023). lastly, the rigorous cts application process requires applicants to assess ongoing compliance and compliance stretch goals with respect to “the characteristics required to be a trustworthy repository for digital data and metadata.”8 this assessment involved a significant investment of time to develop a sustainable metadata culture within isrdi, including: articulating the archive role within the organization; creating an archival workflow that makes sense in the unique ipums environment;9 producing and leveraging documentation;10 and working to comply with recognized international preservation standards.11 a valuable byproduct of all this work was illuminating the enormous intellectual investment ipums product teams contributed to collecting, harmonizing, organizing, cleaning, documenting, and disseminating ipums’ unique data collections. maturation in these four areas brought the ipums organization to the point where the will and capacity coexisted to expand the archival preservation workflow to include additional metadata produced by ipums data products, as well as key input artifacts, such as source datasets for which ipums is the primary holder. the challenge for archival staff was to leverage this opportunity to communicate to ipums administrators, project managers, and it staff the benefits of adopting the ipums bpm to support the expansion of the ipums preservation archive. developing the ipums business process model (ipums bpm) the ipums business process model (ipums bpm) was fleshed out by data curator wendy thomas as part of the cts application process. thomas was exposed to general business process modeling prior to her hire at the mpc and specifically to the general statistical business process model (gsbpm) and the generic longitudinal business process model (glbpm) through her extensive work with the data document initiative (ddi alliance).12 as noted below, thomas drew on the principles behind the gsbpm and the glbpm to create an ipums business process model (ipums bpm) (thomas 2024, thomas 2018, magnuson 2015a).13 developing an organization model that clearly and accurately reflected the workflow of ipums products and archival processes was a crucial step in developing cts application materials and a key document clarifying “the role of the archive within the organization and provid[ing] the archive with the means to clearly present that role and identify specific touchpoints to the ipums project workflows” (magnuson and thomas 2023). thomas’ exposure to the gsbpm and the glbpm as they were adapted to ddi processes enabled her to see the utility of tailoring the glbpm to an ipums context. while the gsbpm and glbpm are similar, the glbpm sub-levels use the terminology and perspective of social scientists, thus making it a good fit for ipums processes (thomas 2024). https://doi.org/10.29173/iq1130 5/13 magnuson, diana l. (2024). the ipums business process model: instituting a workflow mapping strategy to support archival processes, iassist quarterly 48(4), pp. 1-13. doi: https://doi.org/10.29173/iq1130 the current iteration of our business process model contains nine general process areas that mirror the nine general process areas of the glbpm with sub-levels that are tailored to ipums specific tasks.14 table 1 outlines the ipums bpm and notes tailored glbpm sublevels in bold. table 2 lists the sub-level language changes and additions identified in bold in table 1. as with the gsbpm and the glbpm, the key to the ipums bpm is its flexibility: commonalities between the processes of individual ipums projects can be identified in the model while allowing for differences in the selection and ordering of tasks within each project over time. ipums project managers are not “locked in” to one path through the ipums bpm, thus preserving their project specific workflow and our institutional standards. these nine process “activities” and their respective sub-levels are documented in outline, descriptive, and “map” formats, and described in detail in the ipums archive workflow documentation.15 table 1. ipums business process model (ipums bpm) outline (tailored glbpm sublevels in bold) evaluate/specify needs 1.1 define research needs, coverage and high-level concepts 1.2 evaluate existing data and publications 1.3 establish outputs and needed infrastructure 1.4 identify specific concepts to be harmonized 1.5 plan, create timetable, and identify needed infrastructure 1.6 identify partners 1.7 prepare proposal and obtain funding design/redesign 2.1 identify sources 2.2 design sampling methods 2.3 design capture process 2.4 specify data elements and related metadata 2.5 specify processing/data cleaning methods 2.6 specify evaluation plan 2.7 organize research team 2.8 design infrastructure build/rebuild 3.1 develop data capture processes 3.2 create or enhance infrastructure components 3.3 validate processes and tools 3.4 test production systems 3.5 finalize production systems collect https://doi.org/10.29173/iq1130 6/13 magnuson, diana l. (2024). the ipums business process model: instituting a workflow mapping strategy to support archival processes, iassist quarterly 48(4), pp. 1-13. doi: https://doi.org/10.29173/iq1130 4.1 select sources 4.2 negotiate access and distribution rights 4.3 capture data 4.4 obtain metadata 4.5 create sample process/analyze 5.1 validate data against metadata 5.2 select and restructure data 5.3 clean and anonymize data 5.4 impute missing data 5.5 harmonize selected data 5.6 calculate weights 5.7 calculate aggregates 5.8 validate processed data 5.9 finalize data outputs archive/preserve/curate 6.1 ingest data and metadata 6.2 enhance metadata 6.3 capture process/provenance metadata 6.4 preserve data and metadata 6.5 undertake ongoing curation data/dissemination/discovery 7.1 deploy release infrastructure 7.2 preserve dissemination products 7.3 deploy access control system/policies 7.4 promote dissemination products 7.5 provide data citation support 7.6 enhance data discovery 7.7 manage user support research/publish 8.1 obtain listing of publications based on the data product 8.2 maintain publication database 8.3 manage versioning 8.4 deposit metadata in related systems 8.5 manage disclosure work https://doi.org/10.29173/iq1130 7/13 magnuson, diana l. (2024). the ipums business process model: instituting a workflow mapping strategy to support archival processes, iassist quarterly 48(4), pp. 1-13. doi: https://doi.org/10.29173/iq1130 retrospective evaluation 9.1 establish evaluation criteria 9.2 gather evaluation inputs 9.3 conduct evaluation 9.4 determine future actions table 2. list of specific changes to the glbpm to create the ipums bpm evaluate/specify needs 1.4 “concepts to measure” to “concepts to harmonize” design/redesign 2.3 “collection process” to “capture process” 2.4 added related metadata 2.6 “analysis plan” to “evaluation plan” build/rebuild 3.1 “collection process” to “capture process” collect 4.1 “sample” to “sources” 4.2 “set up collection” to “negotiate access” 4.3 “run collection” to “capture data” 4.4 finalized collection” to “obtain metadata” 4.5 added create sample process/analyze 5.1 “integrate data” to “validate data against metadata” 5.2 “classify and code” to “select and restructure data” 5.3 “explore, validate and clean data” to “clean and anonymize data” 5.5 “construct new variable and units” to “harmonize selected data” 5.8 “anonymize data” to “validate processed data” archive/preserve/curate 6.3 added capture process/provenance metadata https://doi.org/10.29173/iq1130 8/13 magnuson, diana l. (2024). the ipums business process model: instituting a workflow mapping strategy to support archival processes, iassist quarterly 48(4), pp. 1-13. doi: https://doi.org/10.29173/iq1130 data/dissemination/discovery 7.1 “discovery of data and relevant research” to “deploy release infrastructure” 7.2 “access data” to “preserve dissemination products” 7.3 “prepare data” to “deploy access control system/policies” 7.4 “analyze data” to “provide data citation support” 7.5 “prepare research “publications” to provide data citation support” 7.6 “manage disclosure risk” to “enhance data discovery” 7.7 “publish research” to “manage user support” research/publish no change retrospective evaluation no change armed with the ipums bpm and administrative encouragement, archival staff organized an initial meeting with each of the nine ipums project management teams to examine the ipums bpm and begin a conversation about how archival staff was positioning to support ipums project teams more effectively and expand the ipums preservation archive work. after several meetings, it became clear that utilizing the activity “map” visualization of the ipums bpm (figure 1) facilitated an immediate grasp of project specific ipums workflows and archival touchpoints from external submission of data, harmonization process, extraction systems, documentation, and archival preservation of metadata. figure 1. ipums business process model (ipums bpm) as an activity “map” from an archival perspective, the purpose of the meetings was threefold. first, to use the ipums bpm as a locus to identify preservation priorities for each project. second, to expand project managers’ thinking about preserving intellectual contribution relating to process and methodology of developing ipums data collections. historically ipums data collection teams have primarily been (rightly) concerned with preserving the data that is disseminated to users and less intentionally attentive to preserving the pieces of intellectual activity that contributed to the data harmonization process. preserving all these elements is important not only to ipums’ institutional history narrowly, but a https://doi.org/10.29173/iq1130 9/13 magnuson, diana l. (2024). the ipums business process model: instituting a workflow mapping strategy to support archival processes, iassist quarterly 48(4), pp. 1-13. doi: https://doi.org/10.29173/iq1130 significant contribution to social science infrastructure more broadly. lastly, the meetings persuaded ipums project managers of the future utility both internally for staff and externally for data users, of preserving their enormous investment collecting, harmonizing, documenting, and disseminating ipums’ unique data collections. once project teams began talking about the intellectual activity underlying the processes and metadata they were creating, they eagerly identified potential areas for long term preservation (table 3). table 3. preservation priorities of ipums project teams (denoted by “x”) key below* (1) (2) (3) (4) (5) (6) (7) (8) (9) annual reports to funders x x git hub x x grant proposals x x x x x x x x x process documentation x x x x x x x x producer documentation x x x software tools (ipums it) x x syntax x x x training materials (internal) x x x x x x x training materials (external) x x x x x x x x x translation tables x webpages x x x x x x x x x wiki x x x x x x x x x https://doi.org/10.29173/iq1130 10/13 magnuson, diana l. (2024). the ipums business process model: instituting a workflow mapping strategy to support archival processes, iassist quarterly 48(4), pp. 1-13. doi: https://doi.org/10.29173/iq1130 *(1) usa; (2) cps; (3) international; (4) global health; (5) nhgis; (6) ihgis; (7) time use; (8) health surveys; (9) highered; terra (now decommissioned); cdoh; and historical census projects after these productive information gathering meetings, archival staff created five clear action points. first, for each ipums project the archive will identify metadata creation points using the ipums bpm. second, in consultation with project managers and ipums it, archival staff will differentiate between business continuity preservation (short term) and archive preservation (permanent) practices for the content and metadata identified for preservation. next, project metadata capture requests will be prioritized. fourth, archival staff will collaborate with ipums it to create tools for archivist led metadata capture and preservation. finally, a metadata capture schedule will be reviewed with project stakeholders and implemented in collaboration with ipums it. advantages of implementing the ipums bpm as we have previously noted, implementing the ipums bpm has advantages for the projects, the administrative team, the it team, and the ipums preservation archive. first, common vocabulary is used across all organizational entities, streamlining and enhancing archival related communication. second, the ipums it team can better develop tools for use across projects, which in turn advance efficiencies and economies of scale around all aspects of the ipums data acquisition, harmonization, documentation, dissemination, and preservation workflow. at the administrative level, process and tool developments can be identified for use across projects and for future grant development purposes. third, the ipums bpm is flexible enough to provide both institutional continuity and individualized project workflows. lastly, the ipums bpm flags the areas of project metadata production that require the attention of the archive for provenance and preservation purposes (magnuson and thomas 2023). strategic use of the ipums bpm will support and further four ipums preservation archive objectives. first, to provide access to current and previous versions of ipums data. second, to retain internal-use data and metadata for the purpose of provenance and quality assurance. third, to meet archival requirements of external funding agencies, demonstrating compliance with data, metadata, and archival standards, including disciplinary standards of data users and the digital preservation community. finally, to maintain a digital preservation program that is nimbly responsive to the everchanging technological environment. for the ipums preservation archive staff in particular, strategic use of the ipums bpm has continued to clarify the role of archival function within the organization and provide us “with the means to clearly present that role and identify specific touchpoints to ipums project workflows” (magnuson and thomas 2023). further, use of the ipums bpm in conversation with ipums project managers has expanded communication regarding the vital importance of the enormous investment project teams have contributed over the history of the ipums data harmonization production work. preservation of this important intellectual work in developing processes, systems, and tools for data harmonization, documentation, and dissemination is vital for the intelligent use and analysis of the data by contemporary and future researchers. in collaboration with ipums it staff, efforts to create and implement tools for archivist-led preservation processes using the ipums bpm are ongoing. https://doi.org/10.29173/iq1130 11/13 magnuson, diana l. (2024). the ipums business process model: instituting a workflow mapping strategy to support archival processes, iassist quarterly 48(4), pp. 1-13. doi: https://doi.org/10.29173/iq1130 conclusion is your archival function situated in an organizational setting in which your preservation work is vital but secondary to the main product? are you playing preservation “catch-up” with a fast growing “main product” within your organization? the ipums bpm, a workflow mapping strategy employed by the ipums preservation archive, is applicable to other data archive contexts, especially those in which preservation work is secondary to the primary product of the institution. actionable steps to move forward with, develop, and expand archival goals include: • create a business process model (bpm) that reflects the workflow and metadata creation points for your organization. • invest in using the bpm to effectively communicate with organization stakeholders. • align metadata preservation, discoverability, documentation, and accessibility with international standards. • leverage collaborative work with it staff to develop longand short-term goals. identify strengths and weaknesses of your current archival data management plan and how it involvement can help you reach your preservation goals. • maintain a digital preservation program that is responsive to the changing technological environment. • advocate for, and take advantage of, opportunities to solidify the role of the archive within the larger organization; highlight value-added to the organization’s process and products through working with data projects, supporting a better understanding of workflows, preservation, and access needs. • do not waiver from the position that preservation, discoverability, documentation, and accessibility of archival material in all its forms is valuable to the main product of your organization. investing in creating and implementing a business process model will benefit your archival workflow, support your shortand long-term preservation goals, highlight where production and preservation activities intersect, and thus enhance archival support of your organization’s main product. references block, w. and thomas, w. (2003) ”implementing the data documentation initiative at the minnesota population center,” historical methods: a journal of quantitative and interdisciplinary history, volume 36, no. 2. esteve, a. and sobek, m. (2003) ”challenges and methods of international census harmonization,” historical methods: a journal of quantitative and interdisciplinary history, volume 36, no. 2. fabrizio, f. (2023) ”ipums technology: 30 years of innovation, a unique partnership of researchers and technologists,” data-intensive research conference, minneapolis, minnesota. fitch, c.a. and ruggles, s. (2003) ”building the national historical geographic information system,” historical methods: a journal of quantitative and interdisciplinary history, volume 32, no. 1. https://doi.org/10.29173/iq1130 12/13 magnuson, diana l. (2024). the ipums business process model: instituting a workflow mapping strategy to support archival processes, iassist quarterly 48(4), pp. 1-13. doi: https://doi.org/10.29173/iq1130 hall, p.k, fitch, c., canaday, m., ebeltoft-kraske, l., ronnander, c., and thomas. k.m. (1999) ”ipums metadata: documenting 150 years of census microdata,” historical methods: a journal of quantitative and interdisciplinary history, volume 32, no. 3. magnuson, d.l. (2024) ”stewarding our resources: building a sustainable ipums archival document access system,” iassist quarterly, volume 48 number 1. https://iassistquarterly.com/index.php/iassist/article/view/1095/1032 magnuson, d.l. (2015a) wendy thomas interview, university of minnesota, march 24, 2015. magnuson, d.l. (2015b) todd gardner interview, university of minnesota, may 27, 2015. magnuson, d.l. and ruggles, s. (2022) “challenges of large-scale data processing in the 1990s: the ipums experience,” ieee annals of the history of computing, pp. 71-83. https://ieeexplore.ieee.org/abstract/document/9972862 magnuson, d.l. and thomas, w. l. (2023) ”expanding our perspective: building a sustainable metadata culture,” iassist quarterly, volume 42 number 2. https://iassistquarterly.com/index.php/iassist/article/view/1046 ruggles, s., cleveland, l., and sobek, m. (2023) ”harmonizing global census microdata: ipums international,” in irina thomescu-bubrow, christof wolf, kazimierz m. slomeczynski, and j. craig jenkins (eds) survey data harmonization in the social sciences. new york: wiley, pp. 207-226. doi:10.1002/9781119712206 ruggles, s., sobek, m., king, m.l., liebler, c., and fitch, c.a. (2003a) ”ipums redesign,” historical methods: a journal of quantitative and interdisciplinary history, volume 32, no. 1. ruggles, s., king, m.l., levison, d., mccaa, r., and sobek, m. (2003b) ”ipums international,” historical methods: a journal of quantitative and interdisciplinary history, volume 32, no. 2. ruggles, s., sobek, m., and gardner, t. (1996) ”disseminating historical census data on the world wide web,” iassist quarterly, volume 20 number 3. https://iassistquarterly.com/index.php/iassist/issue/view/15 thomas, w. (2024) ”why gsbpm glbpm,” general information, isrdi archive. thomas, w. (2018) ”harmonization business process model,” general information, isrdi archive. endnotes 1 diana l. magnuson is curator and historian at the institute for social research and data innovation, university of minnesota (magn0031@umn.edu). 2 https://www.ipums.org/mission-purpose https://doi.org/10.29173/iq1130 https://iassistquarterly.com/index.php/iassist/article/view/1095/1032 https://ieeexplore.ieee.org/abstract/document/9972862 https://iassistquarterly.com/index.php/iassist/article/view/1046 https://iassistquarterly.com/index.php/iassist/issue/view/15 https://www.ipums.org/mission-purpose 13/13 magnuson, diana l. (2024). the ipums business process model: instituting a workflow mapping strategy to support archival processes, iassist quarterly 48(4), pp. 1-13. doi: https://doi.org/10.29173/iq1130 3 in 2016, as part of an institutional reorganization, all data projects took on the ipums prefix as part of their project name. since not all projects are microdata and some have access conditions that limit their usage, it is inaccurate to describe ipums as a “public use” microdata series. thus, since 2016, ipums is a brand, not an acronym. magnuson, d.l. and thomas, w. l. (2023) ”expanding our perspective: building a sustainable metadata culture,” iassist quarterly, volume 42 number 2. 4 ipums motivation and work to build the cts application are described in magnuson, d.l. and thomas, w. l. (2023) ”expanding our perspective: building a sustainable metadata culture,” iassist quarterly, volume 42 number 2. ipums cts documentation: ”ipums is a trustworthy repository,” https://www.ipums.org/about/more. ipums cts certification: https://dataverse.nl/dataset.xhtml?persistentid=doi:10.34894/cranso 5 https://www.ipums.org/about/more; https://www.ipums.org/; https://www.ipums.org/about/staff 6 https://www.ipums.org/about/funding 7 https://tech.popdata.org/about/about-isrdi-it 8 https://www.coretrustseal.org/why-certification/requirements/ 9 https://www.ipums.org/workflows 10 https://assets.ipums.org/_files/ipums/digital_preservation_framework_may2022.pdf 11 https://dataverse.nl/dataset.xhtml?persistentid=doi:10.34894/cranso 12 https://ddialliance.org/ 13 the general statistical business process model was developed over several years by the joint unece/eurostat/oecd work sessions with v1.0 released in march 2008 (v5.1 was released january 2019) https://statswiki.unece.org/display/gsbpm. drawing on the principles of the gsbpm, the ddi alliance developed the general longitudinal business process model in march 2013 to emphasize the data management needs of longitudal data production. https://ddialliance.org/sites/default/files/genericlongitudinalbusinessprocessmodel.pdf https://assets.ipums.org/_files/ipums/workflows/ipums_archive_workflow_nov2021.pdf 14 https://ddialliance.org/sites/default/files/genericlongitudinalbusinessprocessmodel.pdf, p. 7. 15 https://assets.ipums.org/_files/ipums/workflows/ipums_bpm_outline_nov2021.pdf https://doi.org/10.29173/iq1130 https://www.ipums.org/about/more https://dataverse.nl/dataset.xhtml?persistentid=doi:10.34894/cranso https://www.ipums.org/about/more https://www.ipums.org/ https://www.ipums.org/about/staff https://www.ipums.org/about/funding https://tech.popdata.org/about/about-isrdi-it https://www.coretrustseal.org/why-certification/requirements/ https://www.ipums.org/workflows https://assets.ipums.org/_files/ipums/digital_preservation_framework_may2022.pdf https://dataverse.nl/dataset.xhtml?persistentid=doi:10.34894/cranso https://ddialliance.org/ https://statswiki.unece.org/display/gsbpm https://ddialliance.org/sites/default/files/genericlongitudinalbusinessprocessmodel.pdf https://assets.ipums.org/_files/ipums/workflows/ipums_archive_workflow_nov2021.pdf https://ddialliance.org/sites/default/files/genericlongitudinalbusinessprocessmodel.pdf https://assets.ipums.org/_files/ipums/workflows/ipums_bpm_outline_nov2021.pdf 1/13 sapp nelson, megan & kong, ningning nicole (2020) capturing their “first” dataset: a graduate course to walk phd students through the curation of their dissertation data, iassist quarterly 44(3), pp. 1-13. doi: https://doi.org/10.29173/iq971 capturing their “first” dataset: a graduate course to walk phd students through the curation of their dissertation data megan sapp nelson1; ningning nicole kong2 abstract the data set accompanying theses is a valuable intellectual property asset, both from the viewpoint of the phd student, who can procure employment and build publications and research grants from the work for years to come, and the university, which owns the data and has invested in the work. however, the data set has generally not been captured as a finished product in a similar manner to the published thesis. a course has been developed which walks phd students through the process of identifying an archival data set, selecting a repository or long term storage location, creating metadata and documentation for the data package, and the deposit process. a preand post assessment has been designed to ascertain the level of data literacy the students gain through curating their own dataset. pis for the projects have input into the repositories and metadata standards selected. the university thesis office was consulted as the course was developed, so that accurate procedures and practices are reflected throughout the course. this first of a kind class is open to students of any discipline at a research-1 university. the resulting mixture of data types creates a unique course every time it is offered. keywords data curation, instruction, curriculum, data literacy, dissertation, thesis introduction identifying the appropriate timing within the capture, management, and preservation of a research project’s data for the just-in-time education in data curation has been discussed briefly in the literature but remains unresolved. this lack of resolution is in part due to the necessity of practicing skills in situ with existing data sets as conceptual skills are taught or transferred. phd students, by virtue of the point in their career that they have attained, have produced viable data sets, and have reached a conceptually open space where they understand the necessity of practicing data curation skills. the alignment of timing, existing data sets of intrinsic importance to the individual learner, a self concept within data management (either in a current role as a data manager in a lab or a perception of themselves as a future data steward responsible for the production of data), and extrinsic motivation to capture as many metrics as possible to demonstrate impact of nascent research careers, all align to make metriculating phd candidates prime targets for data curation education. a need for this data curation education was also determined by requests from phd candidates’ academic advisors, particularly when the dataset was not captured but was a valuable output of the research process. this often reflected the fact that the traditional university dissertation submission requirement is not enough to capture the value of the research data. https://doi.org/10.29173/iq971 2/13 sapp nelson, megan & kong, ningning nicole (2020) capturing their “first” dataset: a graduate course to walk phd students through the curation of their dissertation data, iassist quarterly 44(3), pp. 1-13. doi: https://doi.org/10.29173/iq971 dedicated time and appropriate guidance is needed to guide the students through the process of organizing their final datasets, creating data documentation, and sharing their data either within their research lab or with the public. coupled with the recent increasing demands emphasizing the data management skills from the u.s. job market, the curriculum designers developed a data curation course for graduate students to fulfill this need. given this likely audience is also among the busiest and most burdened, providing instruction in data curation of a highly valued data set in an efficient and targeted manner is necessary for knowledge transfer. in the design phase of this course, the curriculum developers have collected information from disciplinary faculty and phd students about their expectations. then their needs were mapped with the data curation process with inputs from many data librarians, to develop the course content and lab activities. while this course is being offered for the first time, the curriculum developers also are closely observing the classroom dynamics and collecting feedback from students in order to improve it for future iterations. brief synopsis of curricular innovation this article reports on an innovative implementation of a graduate level 16 week three credit hour course, designed to assist phd students in the final semester before they deposit their thesis to prepare their data for deposit. the initial pilot enrolled six students from the colleges of agriculture, education, engineering, liberal arts, and polytechnic. the types of data deposited during the pilot included flat files, audio files, 3d image files, gis data, image files, and text files. review of literature research data management has been identified as a key educational need for graduate and postgraduate students. (carlson et al, 2011, doucette and fyfe, 2013 )there have been credit courses developed targeting graduate students that teach data management and data literacy, but frequently with specific audiences in mind. basic introductions to data management and data literacy have been described in the literature, tailored to graduate students in agriculture (carlson and bracke, 2015), engineering (jeffryes and johnston, 2013), science (qin and d’ignazio 2010; frank and pharo, 2016), social sciences (thielen and hess 2017) and health (macy and coates, 2016). many courses focus on data storage and reuse competencies. however, a few courses are now moving into data wrangling using data science tools (pascuzzi and sapp nelson 2018). generally, if data publication and sharing is addressed in these courses, it is the topic of one course session. theses are a special category of institutional and scholarly communication that have required libraries to take special steps to catalog and preserve knowledge for decades, due to limited publication runs. in the past two decades the electronic manifestation of theses and dissertations (etds) have required specific research and technological support to ensure the long term preservation and access of content that was previously available only in print on local institutional shelves.(fineman, 2003) when those dissertations were sitting on the shelves, the data frequently were packaged in the form of tables in text or supplemental data on cd-roms or floppy disks tipped https://doi.org/10.29173/iq971 3/13 sapp nelson, megan & kong, ningning nicole (2020) capturing their “first” dataset: a graduate course to walk phd students through the curation of their dissertation data, iassist quarterly 44(3), pp. 1-13. doi: https://doi.org/10.29173/iq971 into the binding of the dissertation itself, and shelved with the text in the library. (schöpfel et al, 2015a) this made the data available for the purpose of reproducibility, but was not sufficiently flexible to facilitate reuse. with the explosion of digital data, the dissertation is widely available on electronic platforms, but the institution has lost control of the co-linking of the data set with the published dissertation.(collie and witt, 2011) the question of where the responsibility lies for the curation of the dissertation data sets remains unresolved, but generally institutions do not have the staffing to provide hands-on support for full service curation for dissertation data sets. instead, institutions are training phd candidates to curate their own data sets in advance of the deposit of their dissertation (schöpfel et al, 2015b). disciplinary needs for dissertation data curation the initial idea for the development of this course was inspired by a civil engineering faculty member who felt that data was being lost due to phd students and masters level students graduating after publishing theses without the attendant data sets curated in appropriate repositories. library consultations were provided at that time to help his graduating student publish datasets specific to the faculty members’ research lab to a repository, as well as create a web page to guide users to access the data. as part of the library's data curation support, the data usage has been monitored after the publication. below are visit statistics for the dataset one year after the data was published. figure 1: visit statistics for published dataset used to help make use case for data publication course this data curation and publication process (along with the data set altmetrics) helped the faculty to demonstrate his research findings to colleagues and funding agencies, as well as to secure the longevity of the dataset for further research. this process is a good use case showing that libraries can provide data curation education for graduate students at the final stage of their research, providing positive outcomes for themselves and their research advisors. initial proposal following the suggestion from this faculty member, the initial conversation with him identified learning outcomes and a general outline for content that should be included in a course, should it be developed. this initial proposal represents the first tentative thoughts as to what might be appropriate for students who are approaching the problem of archiving their data for long term access and sharing, as prioritized in the civil engineering professor’s original vision. the proposal as written is included here: https://doi.org/10.29173/iq971 4/13 sapp nelson, megan & kong, ningning nicole (2020) capturing their “first” dataset: a graduate course to walk phd students through the curation of their dissertation data, iassist quarterly 44(3), pp. 1-13. doi: https://doi.org/10.29173/iq971 this class is a practicum for final semester graduate students who have data sets to deposit prior to graduation. in this course, the students will receive instruction and mentoring on preparing their thesis data set for deposit and sharing. topics include protecting the intellectual property of the data set, preparing the data set for deposit, and maximizing scholarly impact of the data set. learning objectives students will • critically evaluate their thesis data set for quality in order to enable reproducibility and reuse • compile readme files and documentation in order to enable others understand and use their data sets • create metadata in order to facilitate publication of their data set in a data repository • prepare files with human readable file names in order to share data with thesis supervisors and others. • select data publication venues based upon future impact in order to maximize their scholarly output from their data. • select an appropriate data license in order to preserve their intellectual property. general outline 1. critical evaluation of thesis data sets 2. human readable file names/file structure and file conversion to non-proprietary file formats 3. introduction to documentation as context (readme as index) 4. identification of publication files (subset of thesis data set) 5. file level documentation (all files in thesis data set) 6. file level documentation peer review 7. selecting a data repository 8. metadata (subset of thesis data set) 9. metadata peer review 10. data licensing 11. submission process overview 12. preparation of submission package for repository 13. preparation of submission package for repository 14. peer review of submission package for repository https://doi.org/10.29173/iq971 5/13 sapp nelson, megan & kong, ningning nicole (2020) capturing their “first” dataset: a graduate course to walk phd students through the curation of their dissertation data, iassist quarterly 44(3), pp. 1-13. doi: https://doi.org/10.29173/iq971 15. critical evaluation of thesis data set part two/ course evaluation the proposal did not proceed at that point in time, in part due to a changing landscape in graduate education around data science at the university. a year long process was developed to create an integrated data science initiative, which seeks to instill data science principles within each curriculum at the university at some level, but does not articulate fixed curricular outcomes for those data science principles. this landscape within the university means that individual departments are looking closely at how their existing curricula currently purvey data science, but that there are not many locations that are synthesizing data management and curation from a holistic perspective. given the direction that the university has taken in the development of this program, the libraries has initiated programs to develop curricula to support these underpinning management and curation skills from the undergraduate to graduate levels. in this newly developing landscape, the previously proposed course represented an innovative solution to serve the entire graduate college in educating future data stewards, as well as preserving and protecting the intellectual property produced by the students of the university. given these incentives, the course development process began in summer 2018, and the course was offered for the first time in spring semester 2019. curricular development for the pilot, the proposed course was limited to final semester phd students. this narrowed the scope for the course design: the course designers could focus on the very practical needs of taking a pre-existing data set coming from any discipline on campus that may or may not have had metadata or documentation prepared for an external audience, and bring that data set to publishable state in 16 weeks (1 full semester). additionally, the course developers wished to keep nearly all course activities within the confines of in-class time (one 50 minute lecture and one 120 minute laboratory). therefore, ensuring that students were not simultaneously collecting data, analyzing data, etc. was a priority so that the tasks for the class could be completed within the time scope designed for the class. additionally, since all course components were intended to be concluded within the confines of the class meetings, the format of the course was by definition active learning. any lecture or theoretical content had to be relayed to the students efficiently, clearly, and in the context of practical, hands on, applied tasks that logically come next in the chain of events that are linked together to create a curated data set. the hands on activities also need to dovetail with the thesis deposit process as dictated by the university’s graduate college. at the time of the course development, a new dissertation repository was being developed and procedures were being developed simultaneously with the course. therefore, a series of meetings were held with the thesis repository manager to ensure that the curriculum was accurate and supported the current best practices of the graduate school. https://doi.org/10.29173/iq971 6/13 sapp nelson, megan & kong, ningning nicole (2020) capturing their “first” dataset: a graduate course to walk phd students through the curation of their dissertation data, iassist quarterly 44(3), pp. 1-13. doi: https://doi.org/10.29173/iq971 these constraints meant that the course development of the lecture materials and lab activities started from the initial outline, and then solidified into a “data set designed for re-use” focus that was comprehensive of both reuse within the local (but future) research laboratory, or reuse by an unknown third party. the selling point for students taking the class focused on building their professional portfolios through sharing their data for reuse. ultimately, the course objectives and learning objectives were articulated as: this course walks students through the process of preparing a data set for sharing with both internal and external audiences. students will select authoritative data sets from the data sets that they have prepared in the process of doing their thesis work and/or research projects for sharing and publication, apply metadata to those data sets, create documentation for end users of the data sets, and publish the data sets to internal or external data repositories or storage as appropriate. the learning objectives include: • recognize and evaluate the value of research datasets and the needs for preservation. • prepare dataset packages for sharing and reuse that describes the documentation, workflows, and data enclosed in ways that allow users to determine currency, relevance, authority, accuracy, and purpose of the data set. • understand basic metadata fields, metadata standards, and be able to apply standard metadata to make the data set available to others and, if applicable to comply with disciplinary norms. • recognize disciplinary practices, values and norms related to organizing, sharing in disciplinary data repositories, curating and preserving data. • post data sets with recommended citations, including digital object identifier. • share data in a repository or appropriate storage as agreed upon project primary investigator. in some occasional cases, pre-existing lecture materials created for other courses at our institution were appropriate for use. in one case, the dataone lecture on metadata was considered to be at the correct level of specificity and comprehensiveness that there was no reason to develop a lecture from scratch. however, in most cases, lectures were developed to meet the specific needs in terms of practicality and theory that are unique to this course. the activities are custom created for the course according to the content of each week. overall, each course activity can be considered as one component toward a larger project of curating and sharing the student’s valuable research data during their phd study. the activities are not graded on a weekly basis. in some cases, the assignments accumulate over the course of several weeks prior to submission, due to the work required to compile a completed component of the data submission package. of note, the documentation and metadata each require multiple weeks to create in the class, and therefore those sections of the semester are covered over about one and a half months. https://doi.org/10.29173/iq971 7/13 sapp nelson, megan & kong, ningning nicole (2020) capturing their “first” dataset: a graduate course to walk phd students through the curation of their dissertation data, iassist quarterly 44(3), pp. 1-13. doi: https://doi.org/10.29173/iq971 assignments graded towards the final grade total the course begins with a review of the known information about the students’ research projects and data sets to familiarize the entire class with everyone else’s projects and datasets. students were then introduced to methods of determining the value of data and the research data life cycle so that they can identify the most valuable data records for the purposes of preservation and reuse. after that point, they learn the process of creating meaningful documentation, machine-readable metadata, identify relevant data repositories or data sharing spaces, and prepare the data package for preservation. the concepts of data sharing policies, licensing, embargo, and selecting a data repository are introduced during the semester “just in time” so that students can make appropriate decisions according to the nature and stage of their projects. the course performance is evaluated by seven assignments, two peer review experiences, and attendance. attendance is key to the completion of the final deliverables due to the reliance on inclass time to complete tasks. without actively attending the course, especially the first time when it is offered, it is hard to ensure each student will get enough guidance for their discipline specific dataset. the list of evaluations include: • data information sheet an introduction to the dataset as well as an agreement with the student’s academic supervisor regarding the intent to share the data either internally or externally • data of record an inventory of datasets used or generated for the research • final data of record/ authoritative data an evaluation of the data records using the value of data rubrics • documentation documentation at the project, folder, and data levels, including readme file, data dictionary, and code book the curriculum developers adapted the readme file template in use in the class from https://cornell.app.box.com/v/readmetemplate. additional fields were added to integrate a data dictionary to the template. • metadata standards are selected as appropriate to each project/discipline/repository/dataset and applied as appropriate. • data submission package a completed package of documentation, metadata, data and a preferred citation. • citation a deliberately designed, specified citation for the data set. • peer review of documentation • peer review of submission package • attendance https://doi.org/10.29173/iq971 https://cornell.app.box.com/v/readmetemplate 8/13 sapp nelson, megan & kong, ningning nicole (2020) capturing their “first” dataset: a graduate course to walk phd students through the curation of their dissertation data, iassist quarterly 44(3), pp. 1-13. doi: https://doi.org/10.29173/iq971 peer review processes were built into this course during the data documentation stage and data submission package stage. these are two critical stages for data curation. through a peer review process, students can get helpful feedback from another set of eyes and make sure their data documentation and organization can be understood by others who are not inculcated within their research project. it was not necessary that all the lab activities were evaluated for the course, since there is no right or wrong decision that could be made at multiple points throughout the semester. the curriculum designers were present throughout the semester to provide feedback as decisions were made and to point to best practices and pros and cons of decisions. we only chose to grade the activities where our feedback or suggestions could be helpful in the students’ data curation process and to provide momentum toward the creation of the data submission package. students’ academic advisors’ signatures are required at the beginning and end of the course to make sure the data curation and sharing efforts made through this course align with their research labs’ data management needs and ethics, so that the students can maximize the benefits of the course by handing down best practices to other personnel in their home labs during this process. class schedule (subject to modification) week class topic activities week 1 jan 7 lecture introduction gather individual data information and get pi to fill out data profile section before lab. lab data profile week 2 jan 14 lecture value of data identifying the data of record for preservation lab week 3 jan 21 lecture no class – martin luther king, jr. holiday lab week 4 jan 28 lecture evaluation of data of record initial evaluation of data set lab week 5 feb 4 lecture introduction to documentation 1 st draft of documentation lab week 6 feb 11 lecture peer review of documentation 1 peer review of documentation lab week 7 feb 18 lecture documentation, cont. 2nd draft of documentation lab week 8 feb 25 lecture identifying repository (guest lecture) identifying repository, discussion with data steward lab week 9 mar 4 lecture metadata draft of metadata lab spring break this cell is empty this cell is empty this cell is empty week 10 mar 18 lecture sharing policies and licensing/ embargo (guest lecture) license selection in collaboration with pi/ identification of embargo if any. lab week 11 mar 25 lecture pulling together a data package inventorying data objects to include in data package lab creating a data package week 12 apr 1 lecture creating a data package, part 2 peer review of data package, including pi https://doi.org/10.29173/iq971 9/13 sapp nelson, megan & kong, ningning nicole (2020) capturing their “first” dataset: a graduate course to walk phd students through the curation of their dissertation data, iassist quarterly 44(3), pp. 1-13. doi: https://doi.org/10.29173/iq971 lab week 13 apr 8 lecture submission to purr identify required information to submit; don’t submit yet. lab submission to storage/disciplinary repository week 14 apr 15 lecture attribution/citation/ purl develop a suggested citation; make certain a doi will be assigned for a published data set; get final sign off from data steward/pi lab week 15 apr 22 lecture submission week/ final proofread/ hit submit record your doi in your thesis. lab table 1: course schedule with topics and deliverables data curation professionals who work with research data management within the libraries were enlisted to present the lectures and activities in areas of their specialization that relate to the specific activities the learners are undertaking at a specific point within their curation process. additionally, professional data curators were enlisted to provide feedback on the curriculum, and then to provide specific technical support for data types that require additional layers of access or curation, such as matlab files or gis layers as indexes to multiple data files. assessment the class was developed with a pre-assessment and post-assessment as part of the curricular model. the assessments were based on two pre-existing documents. the first was the data curation profiles, a long-form structured interview modality designed to elicit data curation practices from researchers.(witt, carlson, brandt, and cragin, 2009) the second pre-existing tool that was used as a starting point for the pre and postassessments was a self assessment tool designed for post-docs and early career faculty members to identify research data management skills that they may need to develop in order to be successful data stewards. (carlson, nelson, johnston, and koshoffer, 2015) the resulting assessment contains a brief demographics section that links school, department, years pursuing their phd, and types of funding and support received. a section records number and types of research outputs that exist from the research project independent of the thesis, including conference proceedings and journal articles, and the role that the student participating in the class played in the authorship of these works (solo author, first author, etc.). finally, an extensive section details the research data management process and parameters for the data set that will be curated in the course of the class. this section includes information on grant funding agencies and data sharing requirements, data management plans, licensing agreements for any data that was reused in the process of carrying out the phd research project, preferences for data preservation and sharing, and any requirements that are mandated by local research groups or funding agencies. the pre-assessment and post-assessment questions are similar in topic. the post-assessment measures both the skills the students have put in practice and attitudinal indicators regarding the importance of the practice for the long term preservation and curation of their data set. the intent https://doi.org/10.29173/iq971 10/13 sapp nelson, megan & kong, ningning nicole (2020) capturing their “first” dataset: a graduate course to walk phd students through the curation of their dissertation data, iassist quarterly 44(3), pp. 1-13. doi: https://doi.org/10.29173/iq971 is that the assessments will measure changes in each of the cognitive, affective, and psychomotor domain, when combined with the graded submissions for the course. with the size of the pilot course, nothing will be able to be said about statistical significance or correlation. it is the hope of the curriculum developers that (over the course of multiple semesters) we will be able to determine the impact of the curation of the data sets on altmetrics and and traditional citation metrics (which have not traditionally represented the impact of dissertations well). lessons learned from pilot course each data set brought to the pilot course is unique and represents a different discipline. not only are the data fundamentally different (audio data, microscopy images, gis data, matlab data, open source software outputs, text) but the goals of the students for the data sets are different as well. the articulated goals include creating a reference data set of images, setting up a curation protocol for an ongoing project that will be completed in three years, and sharing for the purpose of developing policy. the students also bring a variety of baseline knowledge of research data management skills and curation principles. some students have been working with high performance computing, while others are using desktop computing software for all aspects of their dissertation project. one has been the designated data manager for a large research lab, others have been working independently on their research throughout their career. all of this is to say that the curriculum designers focus on fundamental principles of data publication was correct because the principles are the only thing the course participants have in common. tools are not held in common. processes are not held in common. if examples are needed for the course, they can be selected from any discipline because connections will have to be drawn for many other people in the class from that example to their work. course preparation for specific aspects of the class such as metadata standards, packaging for software, identifying repositories, and other disciplinary specific topics will have to be customized to each semester’s roster of students and their specific data sets. in that way, the course will never truly be a completed curricula, and will always require a higher workload for the instructor on an ongoing basis. the initial instinct of the course designers were to ensure that the primary investigators signed a document indicating their consent for the phd students to share data. it turned out that this was too limited for large research groups. many primary investigators consider themselves to not be the data stewards/sole proprietors of the data, but to hold it in common with all members of the research laboratory. this in turn means that the pi wants to have conversations with all members of the laboratory about what portions of the data can be shared along with the dissertation. this is a very positive ripple effect from the class, but the one week turnaround time for this deliverable is too short, given the amount of coordination that has to happen for the large labs. https://doi.org/10.29173/iq971 11/13 sapp nelson, megan & kong, ningning nicole (2020) capturing their “first” dataset: a graduate course to walk phd students through the curation of their dissertation data, iassist quarterly 44(3), pp. 1-13. doi: https://doi.org/10.29173/iq971 there was some question whether three credit hours were too many. however, in the final semester prior to deposit, the participants are very busy with final edits on their dissertation, defending their thesis, and a number of other details needed to get the document done. were the class time not provided to carry out the curation activities, it is unlikely that the steps of the process would be completed. therefore, though reducing the number of credit hours to two was an option, it has been rejected. issues remaining to be addressed an unanticipated issue that was identified early in the course was that of the inflexible nature of the software profile of the computers in our computer lab. an image is loaded at the beginning of the semester, before the students are enrolled in the course. no software is permanently installed after that point. however, the instructors don’t actually know what the curation requirements for the course will be, including metadata editors, software packages, etc, until the first week of class. this timeline mismatch is problematic, and leads to the computer lab being significantly less useful than it would otherwise have been. there was not enough time, even with 16 weeks in the class, to do comprehensive image level or audio file level metadata for the largest of the student projects. good quality metadata really requires the time investment of a full time data curator, which is just not available under the current design of the course. unless the students select a data repository that employs a full time data curator that can backfill that level of metadata, the data sets will have collection level metadata, but at some level will still not be machine readable at the individual datum level. relying on the individual researcher to create this level of metadata may well be wholly unrealistic, however. conclusion this small pilot indicates that the format of a three credit hour course is an appropriate venue for data curation and publication education. however, educational research is ongoing regarding the efficacy of the intervention itself. whether the data sets that are produced will be of sufficient quality for reuse; whether repositories will be happy with the level of documentation and metadata produced by the students; whether students will be ready to serve in a directive role as data stewards in their future endeavors; and whether primary investigators will feel that the data has been captured sufficiently are all areas of ongoing research. once these areas have been established, a primary problem that will have to be addressed is that of the scale of the educational intervention. how can institutions of higher education teach all graduating phd and master’s students to capture their own research data in conjunction with the writing of their thesis? it is the question that the curriculum developers started this project with, and it remains an outsized problem that this class does not resolve. if anything, this class points to the complicated nature of providing the customized data curation education that individual students with their own data sets need. further research and publication will be released in the future as these questions are investigated. references carlson, j. and bracke, m. (2015). planting the seeds for data literacy: lessons learned from a student-centered education program. international journal of digital curation, 10(1), https://doi.org/10.2218/ijdc.v10i1.348 https://doi.org/10.29173/iq971 https://doi.org/10.2218/ijdc.v10i1.348 12/13 sapp nelson, megan & kong, ningning nicole (2020) capturing their “first” dataset: a graduate course to walk phd students through the curation of their dissertation data, iassist quarterly 44(3), pp. 1-13. doi: https://doi.org/10.29173/iq971 carlson, j, fosmire, m, miller, c, and sapp nelson, m. (2011). determining data literacy needs: a study of students and research faculty. portal: libraries & the academy, 11(2) p. 629-657. doi: https://doi.org/10.1353/pla.2011.0022 carlson, j., sapp nelson, m., johnston, l. and koshoffer, a. (2015). developing data literacy programs: working with faculty, graduate students and undergraduates. bulletin of the association for information science and technology. http://doi.org/10.1002/bult.2015.1720410608 collie, w. and witt, m. (2011). a practice and value proposal for doctoral dissertation data curation. international journal of digital curation, 6(2), pp.165-175. doucette, l and fyfe, b. (2013). drowning in research data: addressing data management literacy of graduate students. in imagine, innovate, inspire: the proceedings of the acrl 2013 conference (pp. 165-171). fineman, y. (2003). electronic theses and dissertations. portal: libraries and the academy, 3(2), pp.219-227. https://doi.org/10.1353/pla.2003.0032 frank, e. and pharo, n. (2016). academic librarians in data information literacy instruction: a case study in meteorology. http://hdl.handle.net/10642/3470 jeffreys, j. and johnston, l. (2013). an e-learning approach to data information literacy education. in proceedings of the asee annual conference proceedings, 2013. http://hdl.handle.net/11299/156951 macy, k and coates, h. (2016). data information literacy in business and public health: comparative case studies. ifla journal 42(4), 313-327. https://doi.org/10.1177/0340035216673382 qin, j. and d’ignazio, j. (2010). lessons learned from a two-year experience in science data literacy education. in the proceedings of the 31st annual iatul conference. retrieved from https://docs.lib.purdue.edu/iatul2010/conf/day2/5/ research data management service group. (2019). “author_dataset_readmetemplate.txt” retrieved from https://data.research.cornell.edu/content/readme schöpfel, j., primož, j., prost, h., malleret, c., češarek, a., & koler-povh, t. (2015a). dissertations and data: keynote address. in gl17 international conference on grey literature (hal-01285304). amsterdam, netherlands. retrieved from https://hal.univ-lille3.fr/hal-01285304/document schöpfel, j., prost, h., & malleret, c. (2015b). making data in phd dissertations reusable for research. in 8th conference on grey literature and repositories (p. hal-01248979). prague, czech republic. retrieved from https://hal.univ-lille3.fr/hal-01248979/document thielen, j and hess, a. (2017). advancing research data management in the social sciences: implementing instruction for education graduate students into a doctoral curriculum. behavioral & social sciences librarian 36(1), pp 16-30. http://doi.org/10.1080/01639269.2017.1387739 witt, m., carlson, j. , brandt, d. , & cragin, m. (2009). constructing data curation profiles. international journal of data curation, 4(3), pp. 93–103. http://doi.org/10.2218/ijdc.v4i3.117 https://doi.org/10.29173/iq971 https://doi.org/10.1353/pla.2011.0022 http://doi.org/10.1002/bult.2015.1720410608 https://doi.org/doi:10.1353/pla.2003.0032 http://hdl.handle.net/10642/3470 http://hdl.handle.net/11299/156951 https://doi.org/10.1177%2f0340035216673382 https://docs.lib.purdue.edu/iatul2010/conf/day2/5/ https://data.research.cornell.edu/content/readme https://hal.univ-lille3.fr/hal-01285304/document https://hal.univ-lille3.fr/hal-01248979/document http://doi.org/10.1080/01639269.2017.1387739 http://doi.org/10.2218/ijdc.v4i3.117 13/13 sapp nelson, megan & kong, ningning nicole (2020) capturing their “first” dataset: a graduate course to walk phd students through the curation of their dissertation data, iassist quarterly 44(3), pp. 1-13. doi: https://doi.org/10.29173/iq971 endnotes 1 megan sapp nelson is a professor of library science and science and engineering data librarian at purdue university libraries. she can be reached by email: msn@purdue.edu. 2 ningning nicole kong is an associate professor of library science and geographic information specialist at purdue university libraries. https://doi.org/10.29173/iq971 mailto:msn@purdue.edu 1/2 kellam, lynda; emmelhainz, celia (2019) guest editors’ notes: special issue on qualitative research support, iassist quarterly 43(2), pp. 1-2. doi: https://doi.org/10.29173/iq954 guest editors’ notes: special issue on qualitative research support welcome to the second issue of volume 43 of the iassist quarterly (iq 43:2, 2019). four papers are presented in this issue on qualitative research support. this special issue arises from conversations in the qualitative social science and humanities data interest group (qsshdig) at iassist about how best to support qualitative researchers. this group was founded in 2016 to explore the challenges and opportunities facing data professionals in the social sciences and humanities, and has focused on using, reusing, sharing, and archiving of qualitative, textual, and other non-numeric data. in ‘annotation for transparent inquiry (ati),’ sebastian karcher and nic weber present their work on a new approach to transparency in qualitative research by the same name, which they have been exploring at the qualitative data repository at the university of syracuse, new york. as one solution to the problem of ‘showing one’s work’ in qualitative research, ati allows researchers to link final reports back to the underlying qualitative and textual data used to support a claim. using the example of hypothes.is, they discuss the positives and negatives of ati, particularly the amount of time required to annotate a qualitative article effectively and technical limitations in widespread web display. the next article highlights how archived materials can be re-used by qualitative researchers and used to build their arguments. in ‘research driven approaches to archival discovery,’ diana marsh examines what qualitative researchers need from the collections at the national anthropological archives in the united states, in order to improve archival discovery for those not as accustomed to working in the archives. in ‘bringing method to the madness,’ mandy swygart-hobaugh, leader of the research data services team at the georgia state university library, outlines a project created to bridge the gap between training researchers to use qualitative data software and training them in qualitative methods. her answer has been a collaborative workshop with a sociology professor who provides a methodological framework while she applies those principles to a project in nvivo. these successful workshops have helped to encourage researchers to consider qualitative methods while at the same time promoting the use of caqdas software. jonathan cain, liz cooper, sarah demott, and alesia montgomery in their article ‘where qda is hiding?’ draw on a study originally conducted for qsshdig to create a list of qualitative data services in libraries. when they realized that finding these services was quite difficult, they expanded the study to examine the discoverability of library sites supporting qda. this study of 95 academic library websites provides insight into the issues of finding and accessing library websites that support the full range of qualitative research needs. they also outline the key characteristics of websites that provide more accessible access to qualitative data services. we thank our authors for participating in this special issue and providing their insights on qualitative data and research. if you are interested in issues related to qualitative research, then please join the qualitative social sciences and humanities data interest group. starting with iassist 2019 in australia, our interest group has a new leadership team with two of our authors, sebastian karcher and alesia montgomery, taking over as co-conveners. we are certain that they would love to hear your ideas for the group, and we look forward to working with the qualitative data community more in the future. https://doi.org/10.29173/iq954 https://sites.google.com/uncg.edu/iassistqsshdig/ 2/2 kellam, lynda; emmelhainz, celia (2019) guest editors’ notes: special issue on qualitative research support, iassist quarterly 43(2), pp. 1-2. doi: https://doi.org/10.29173/iq954 lynda kellam, cornell institute for social & economic research celia emmelhainz, university of california, berkeley https://doi.org/10.29173/iq954 4 iassist quarterly summer 2007 editor’s notes welcome to the second issue of the iassist quarterly, vol. 31 (2007). in this issue we have three papers from people working at the us federal reserve board. viewed from posterity, it might look as if we at the iq were clairvoyant and in 2007 foresaw the global role for the frb in the financial crisis in the last quarter of 2008. the secret is first of all the fact that volume 31-2 is the second issue of the 2007 volume but is somewhat delayed, and we are writing in november 2008. secondly, the articles from the federal reserve board carry opinions that “are of the authors and not the federal reserve board”. as an author in the iq you are supported in expressing your opinions and not necessarily those of your employer. thirdly, these three articles are not about the financial crisis, but hopefully some of the initiatives that are described in them will help us in the current situation. linda f. powell and andrew boettcher from the board of governors of the federal reserve system (washington, d.c.) are involved in the collection, editing, storage, and dissemination of commercial bank reports of income and condition, and the use of the extensible business reporting language (xbrl-format) for that purpose. their article is called “modernizing financial data collection with xbrl”. xbrl can be thought of as a set of accounting standards coupled with information technology standards that simplifies the exchange of data. what was earlier accomplished through a manual collection is now using xbrl for a call report a regulator-specified report for about 7,700 banks that are required to file a quarterly report, containing over 2,000 variables. the article addresses the challenges: 1) multiple collection and storage sites, 2) difficulties for the industry in implementing changes to the data collection requirements, and 3) improvements to data quality. this involves centralization in the new collection model by submitting the data to the central data repository. the article states that “financial theory suggests that more frequent, reliable, and readable financial statement reports will result in a healthier marketplace”. since the presentation of the article at the iassist 2008 conference in may we have experienced a financial crisis. let us hope for further refinements in this area as the xbrl is being used by government regulators worldwide. at the iassist 2008 conference andrew boettcher presented from the frb a metadata repository called the data and news catalogue (dance). the article “data and knowledge management at the federal reserve board” chronicles the role of dance in the organization and its transformation into a knowledge management solution. when research projects were always intradepartmental the departments tended to silo the data, documentation, and expertise. now, the number of multidepartment research projects is rising and more linking is needed. the dance development staff then focused on the concepts: description, access rights, contact information, and data location. one issue was that all datasets needed to have standardized documentation. each dataset has a unique set of security requirements dictating usage rights, how access is granted (request form), and publication rights. the article also addresses the searching of datasets and the additional feature of allowing user-generated content supported by wiki-pages for the dataset. as the electronic information environment is shifting, the presentation of information from the federal reserve board is becoming far less important. san cannon is chief at the economic information management at the frb; she presented “snippets of data at a glance: using rss to deliver statistics” at the united nations economic commission for europe’s dissemination and communication work session in geneva (switzerland) in may 2008. an early version was also presented at the iassist conference in 2007. instant access to information on a variety of devices meant that few would wait until the frb had information posted on a website; the response, in collaboration with other central banks, was to create rss-cb, a specification for central bank data. this was also a response to how frb content was being “harvested” or accessed by automated processes as well as some “screen scraping” software that was used to pull the latest exchange rate or commercial paper rate from an html table. instead there was developed an alternative format for human readers as well as for machines. the article shows details in examples of coding of the rss with content like the exchange rate for the us dollar and the mexican peso. many international institutions are now producing rsscb feeds, and many are meeting in a central bank online communications group collaborating on a version 1.2 the last article in the iq 31-2 is authored by lynn woolfrey at the university of cape town. her article outlines “the establishment of the african association of statistical data archivests (aasda)”. the introduction explains: “aasda represents practitioners in survey data curation in africa and was established to facilitate co-operation among them with regard to the development and use of best practices in the preservation and sharing of survey microdata in the region. this association was established with the assistance of international organisations promoting optimal management of survey data. these included the international household survey network (ihsn) and the international association for social science information service and technology (iassist).” we are naturally happy that iassist was found of help here and the article shows how ihsn was aware and could take action to improve the survey data production and utilization in developing countries. the focus on establishing a iassist quarterly summer 2007 5 community of practice for sharing african data found realization when the aasda held the inaugural meeting in april 2008. remember to take a look at the website http://iassistdata. org and the iassist blog the iassist communiqué – at http://iassistblog.org. articles for the iassist quarterly are very welcome. articles can be papers from iassist conferences, from other conferences, from local presentations, discussion input, etc. contact the editor via e-mail: kbr@sam.sdu.dk. karsten boye rasmussen, november 2008 1/9 marsh, diana e. (2019) research-driven approaches to improving archival discovery, iassist quarterly 43(2), pp. 1-9. doi: https://doi.org/10.29173/iq955 research-driven approaches to improving archival discovery diana e. marsh1 abstract the national anthropological archives (naa), part of the department of anthropology at the smithsonian’s national museum of natural history, holds some 18,000 cubic feet of materials of relevance to qualitative researchers. these archival collections—manuscripts, fieldnotes, audio recordings, drawings, maps, and still and moving images—are used by not only anthropologists, but increasingly scholars from a range of qualitative research fields. in 2016, the naa received a grant to support a 3-year post-doctoral fellow to conduct research that would lead to the improved discovery and use of archival resources. this article discusses some of the practical ways the fellowship was designed to ask interdisciplinary research questions, and describes how that premise, as well as findings from a pilot study run in the first year, are helping to improve the research experience for our increasingly interdisciplinary users. both the project’s preliminary findings and its overall design may provide valuable insights to qualitative researchers and their institutions. keywords access, anthropology, archives, collections, users introduction the national anthropological archives (naa), part of the department of anthropology at the smithsonian’s national museum of natural history, holds some 18,000 cubic feet of materials of relevance to qualitative researchers. these archival collections—manuscripts, fieldnotes, audio recordings, drawings, maps, and still and moving images—are used by not only anthropologists, but increasingly scholars from fields such as history, art history, and social studies of science. in 2016, the naa received a grant to support a 3-year post-doctoral fellow to join the staff of the naa and the broader scientific staff of the department. the goal of the fellowship was to conduct research that would explore archival access, and ideally lead to the improved discovery and use of archival resources of value to researchers. this article discusses some of the practical ways the fellowship was designed to ask interdisciplinary research questions, and describes how that premise, as well as findings from a pilot study run in the first year, are helping to improve the research experience for our increasingly interdisciplinary users. both the project’s preliminary findings and its overall design may provide valuable insights to qualitative researchers and their institutions. the naa collections the national anthropological archives is the united states’ largest archival repository dedicated to the history of anthropology and the world’s cultures, with over 18,000 cubic feet of historical documents, photographs, audio recordings, and film. it holds one of the world’s largest archival collections of both american indigenous languages and ethnographic film. such collections are crucial to understanding how published research was generated. archival materials include a range of unpublished work by scientists and fieldworkers and their colleagues, such as drafts of manuscripts with annotations and corrections, loose sheets or cards containing https://doi.org/10.29173/iq955 2/9 marsh, diana e. (2019) research-driven approaches to improving archival discovery, iassist quarterly 43(2), pp. 1-9. doi: https://doi.org/10.29173/iq955 raw linguistic documentation, notes containing free-form thoughts and observations, or letters illustrating scholarly relationships and networks. such archival documents provide the full context in which knowledge was obtained, synthesized, and produced. as i have written about elsewhere, it was clear through the range of publications generated from naa’s archival materials that these collections are increasingly being used by a diverse range of qualitative researchers (marsh 2018). raw research materials, such as fieldnotes, personal diaries, correspondence, annotated maps, photographs, sound recording, and video, are now being used not only by anthropologists, but linguists (davis 2010), environmentalists (anderson 2005) and ecological historians (loring & spiess 2007), immigration scholars (schmidt, seguchi, & thompson 2011), apparel scholars (marks 2014), the historians of science (hinsley 1994; rich 2012), musicologists (troutman 2013), and ethnomusicologists (moon 2010), english literature scholars (applegarth 2014), and art historians (naeem 2018). non-academic researchers, such as artists, documentary filmmakers, exhibit designers, journalists, and even children’s book authors are also researching these collections for a range of uses with much wider public exposure. native and indigenous community members are also researching their own histories, languages, and cultures in these collections, especially in service of language revitalization programs.2 yet, little had been done to analyze this apparent shift. furthermore, it was known anecdotally that these collections were not reaching the broadest possible range of researchers because of barriers to their access both in-person and online. research questions this postdoctoral nsf project was therefore driven by three premises: 1) that despite the importance of naa archival collections and their increased digital presence, usage remains below the immense potential that the collections hold; 2) a general institutional desire to see naa collections have more scholarly centrality and citation, as well as overall circulation and secondary use; 3) the hypothesis that collections discovery and access are hindered by current descriptive practices, discoverability tools, and interfaces. preliminary research questions included: 1. how can the naa make its collections more discoverable, accessible, usable to researchers? 2. what difficulties are encountered by anthropological researchers and source communities in seeking information in the archives? 3. what attitudes or understandings about archival research are held by anthropologists and other researchers? 4. how can archivists more effectively involve anthropologists and source communities in the archival processes of collection representation? 5. how can archival descriptive practices better represent elements of the collection to increase discoverability by anthropological researchers? 6. how can “traditional” archives such as the naa better engage with emerging digital data repositories such as the digital archaeological record (tdar), the archive of indigenous languages of latin america (ailla), the open language archives community (olac), https://doi.org/10.29173/iq955 3/9 marsh, diana e. (2019) research-driven approaches to improving archival discovery, iassist quarterly 43(2), pp. 1-9. doi: https://doi.org/10.29173/iq955 and the digital endangered languages and musics archives netork (delaman), and how can we best develop shared understandings of “archives,” and “digital data”? my research sought to establish a better understanding of both archival repository and user needs to improve researcher success in the discovery of archival sources. in addition, a core goal of the fellowship was to fill a gap in professional training between archival studies and anthropology. organizationally, the naa (including the human studies film archives) sits within the collections program of the department of anthropology, within the national museum of natural history at the smithsonian institution. being embedded within an anthropology department (staffed with experts in anthropological research) as well as within an archive (staffed with experts in archival science and practice) was key to bridging this gap. the fellowship was designed to meld hands-on work in the naa with research by joining the naa staff, and work on an ongoing collections assessment to learn the collections. because my background is primarily in anthropology, i spent a good deal of time in the first year of the fellowship reading archival science basics and learning how to help with the naa’s daily work, especially in reference. i also assisted our contract archivist, gabriela sanchez, with the naa’s current comprehensive collections assessment. my task was to assess the “intellectual value” of collections based on documentation quality (based on the types of materials and their uniqueness), researcher interest (based on topical focus and past use), and local importance (based on relationships to other collections or inherent institutional value to the naa). in total, we assessed over 300 collections in the first year. concurrently, i began an environmental scan that included a) informal interviews with naa staff about current users, uses, discovery tools, and access issues to glean naa staff understandings of the grant’s research questions; b) compiling a project bibliography on archival users, access, discoverability and other relevant readings, and c) reviewing previously produced institutional reports and studies relevant to the current project. environmental scan & assessment findings my literature review made clear that this project has novelty due to its disciplinary emphasis on anthropological and indigenous collections and its implementation component.3 in particular, few user studies have the benefit of being undertaken at a repository, rather than by university-based researchers, or of a three-year timeline in which findings can be implemented and reflected upon.4 it is not so much that, as elizabeth yakel (2004, p. 65) noted over a decade ago, “user evaluation has rarely been mentioned as an integral aspect of implementation,” but rather that most studies lack the longitudinal timeline and internal institutional support to directly apply evaluation findings to practical change. findings from year one of our assessment revealed that naa collections have major intellectual access barriers. of 314 assessed collections, 253 (81%) do have at least a collection-level catalog (marc) record (e.g. the beatrice medicine papers). however, only 25% have a more detailed finding aid online. many of these are pdf documents. at the time of writing this piece, due to the recent redesign of the nmnh’s website, many of those pdf finding aids are no longer findable online, and it will take staff time to get them back up. only 15% of the 314 assessed have a fully keyword searchable, ead finding aid (in archivesspace) that comes up in all smithsonian search https://doi.org/10.29173/iq955 http://collections.si.edu/search/detail/edanmdm:siris_arc_272739?q=beatrice+medicine&record=2&hlterm=beatrice%2bmedicine 4/9 marsh, diana e. (2019) research-driven approaches to improving archival discovery, iassist quarterly 43(2), pp. 1-9. doi: https://doi.org/10.29173/iq955 platforms (and therefore has been guaranteed to survive the nmnh website redesign). therefore, on a scale from 1 to 5, where 5 is highest, only 15% of our year one assessed collections are rated a 5 and considered highly accessible (see figure 1). figure 1. percentage of collections accessible by description type of 314 assessed of course, the naa is not alone in this. according to a study of their backlog, the national archives and records administration reports only having 26% of its textual collections processed sufficiently to allow “researchers to easily identify records of interest.” thirty-three percent of its collections records lacked basic elements of intellectual control such as titles or dates (bucciferro 2008). according to a 1998 study, the mean of special collections repositories’ backlog is 33% (panitch 2000). i found that the naa’s collections also had additional major barriers to discovery due to the design of its website. working with two spring break interns from the university of michigan’s ischool, thanhthu nguyen and wendi ding, we also found that the naa website lacks discovery functionality. from google analytics, naa sites have a 50-60% exit rate (percentage of users that leave the site from a page) and a 75% drop off rate (percentage who don’t click through to a next page). in other words, researchers were not aided in finding materials of relevance on our institutional website. for qualitative researchers, this means that both at the naa and elsewhere, many materials of relevance are almost impossible to find, or may not be available for use at all. typically, unprocessed archival collections are not made available to researchers because they have not yet been vetted by archivists for potentially sensitive, personally identifiable, or legally complex materials. many archival repositories have websites and interfaces that are not intuitive to researchers and require institutional knowledge to navigate. moreover, these systems often change, so that researchers may need to learn and re-learn interfaces at institutions with relevant collections over the course of their project. pilot study findings the primary research conducted in 2017-2018 was a pilot study to better understand these naa users and their information-seeking behaviors. the pilot study included: a) exploration, https://doi.org/10.29173/iq955 5/9 marsh, diana e. (2019) research-driven approaches to improving archival discovery, iassist quarterly 43(2), pp. 1-9. doi: https://doi.org/10.29173/iq955 preliminary coding, and analysis of available naa fy2016 users from the naa’s remote reference log, visitor appointment database, and permissions database; b) scheduling and completion of 22 targeted 1-hour interviews and three focus group discussions with user communities identified during the existing user data analysis; and c) transcription of all recorded interviews and focus groups for coding and analysis. an analysis of our fy2016 databases confirmed that the naa has a highly diverse set of users (see table 1). native community-based researchers are now the naa’s second largest user group, and we have almost an equal number of academic (47%) and non-academic (46%) users. in addition, from central fy2017 smithsonian data, the naa serves users from 49 us states and territories and 33 countries around the globe. thus, we are serving a range of qualitative researchers from across different disciplinary and professional backgrounds. table 1. frequency of users in top 7 researcher groups, fy2016 twenty-two participants were recruited from these top user communities. i loosely correlated the number of participants from each designated community to the number of total users from that group in 2016. in total, i interviewed: 6 anthropologists from different subdisciplines, 5 community-based researchers, 4 heritage professionals, 3 historians, 2 filmmakers, and 2 humanities scholars (one an art historian and community member). in total, these researcher interviews include 20 hours of audio .wav recording and three written responses. sixty-four participants were invited to participate in three smithsonian focus group discussions, with a total of 14 smithsonian staff participants and 4 hours of .wav audio recording. all interview and fgd transcripts are transcribed and were coded using tamsanalyzer. key interview findings thus far include that: 1. search tendencies make collections harder to find for community and non-academic users. pathways to naa collections differ by user community. academics and community researchers tended to find out about the naa through word of mouth, either in a fellowship or directly from colleagues. all heritage professionals and filmmakers, and all who identified as photo researchers found out about the naa through online searches. only academics found out about the naa through bibliographic sources. only academic users mentioned using a finding aid to identify relevant collections. many academic users search by specific anthropologists’ (record total frequency of users in top 7 groups fy2016 user group frequency (n=1004) anthropologist 170 community researcher 91 historian 85 heritage professional 82 art historian 41 filmmaker 40 other social scientist 32 https://doi.org/10.29173/iq955 6/9 marsh, diana e. (2019) research-driven approaches to improving archival discovery, iassist quarterly 43(2), pp. 1-9. doi: https://doi.org/10.29173/iq955 creator) names or by collection; all non-academic and community-based users tend to search by cultural group name or subject. 2. researchers lack training in archives. very few researchers receive any training in archival research (the logistics of conducting archival research or how archives are organized), and describe learning “as they go,” even if they attended graduate programs in anthropology or history. 3. current smithsonian search platforms are not intuitive for users, even if they know what they are looking for. multiple entry points at smithsonian and nested nature of naa exacerbates this problem. 4. outdated information and thin description is problematic. users also noted the existence of problematic and incorrect catalogue information, where community members specifically mentioned outdated, problematic, or racist terminology (and collections’ description non-native perspective) as an issue. desire for more depth of collections description was identified by 4 academic participants and was thus a lesser factor than expected. 5. user expectations are shifting. users expect more collections easily accessible digitally, especially through the presence of more digital surrogates online. the availability of digital surrogates and the lack of on-demand or on-site digitization were listed as the top barrier to collections access. 6. gaining on-site access is confusing, difficult, and cost-prohibitive. many interviewees mentioned difficulty in finding out how to contact the naa, our appointment process, our security process, and the prohibitive cost and distance to visit. one community user noted that the security process evokes historical trauma. the confusing nature of the smithsonian’s organizational structure (and what collections are where) adds to the feeling of institutional impenetrability. research during year two of this project will include a survey with a number of professional organizations such as the american anthropological association, the association of tribal archives, libraries, and museums, native american, the national association of tribal preservation officers, and the native american and indigenous studies association. implementation of improvements instigated by this research, we have begun to make a number of practical improvements, however small. as part of the nmnh’s website redesign, ongoing during the first year of my fellowship, we worked with other anthropology department staff on the development of a new appointment form that will allow on-site researchers to self-identify their research interests and disciplinary background or subject expertise. this will allow us to more accurately track our users and their backgrounds, and how their interests are changing through time. we are also partnering on broader initiatives. in may, i assisted with a joint university of maryland-naa workshop aiming to understand the impacts of digitized ethnographic archives. the workshop, run by umd assistant professor ricardo punzalan, brought together for 30 community-based users of the john peabody harrington papers. we are also collaborating with the washington state university center for digital humanities’ mukurtu shared project (a webbased version of the mukurtu cms). this project brought two researchers (mukurtu fellows) to the naa for four months to identify collections of interest to partner native community partners, which will be brought into the mukurtu shared system for community comment and the https://doi.org/10.29173/iq955 https://naturalhistory.si.edu/research/anthropology https://naturalhistory.si.edu/research/anthropology/collections-and-archives-access/anthropology-collections-appointment-request https://naturalhistory.si.edu/research/anthropology/collections-and-archives-access/anthropology-collections-appointment-request http://mukurtu.org/ 7/9 marsh, diana e. (2019) research-driven approaches to improving archival discovery, iassist quarterly 43(2), pp. 1-9. doi: https://doi.org/10.29173/iq955 development of ‘copyright’/access protocols. additionally, ricardo punzalan and i are working together to revitalize the council on the preservation of anthropological records (formerly at http://copar.org/) of which the naa has historically been an active part, which will provide researchers with a resource for connecting to anthropological records and resources across the globe. thus far, we have applied for two grants to process and selectively digitize collections of high potential interest to users. we are working to initiate more cross-smithsonian collaborations to showcase naa collections through initiatives such as smithsonian’s transcription center, which will allow online volunteers to transcribe, and thereby make keyword searchable, our digitized collections. most recently, to address the lack of training in archives for many graduate-educated researchers, gina rappaport, naa photo archivist, and i collaborated with alessandro pezzati at the penn museum archives and guha shankar at the library of congress american folklife center to run an archives 101 workshop at the november 2018 american anthropological association meetings. the workshop leaders were all from repositories with significant anthropological holdings. the workshop aimed to “demystify” archival repositories and jargon, and to provide an orientation to conducting research in archives. we introduced the general principles that govern archival organization and descriptive practices, described the types of records that are found in archival repositories and how they can be used, and helped our participants determine strategies for locating materials of interest to them in archival repositories, especially by searching online catalogs and finding aids. we gave participants time to search for archival materials relevant to their own research interests. we had 25 participants attend and received highly positive feedback. we think this is a model that can be replicated at a wide range of conferences for qualitative researchers. it might also be possible to create online videos, modules, or tutorials to train researchers in archival principles to help a wide range of researchers better intuit how to search for potential collections of relevance. conclusion archival collections are a primary research site for many qualitative researchers, and that interest is growing. yet, it is clear that archival collections like those at the naa are not easily accessible for remote research. as a result, many researchers may not be aware of collections of direct relevance to their projects. it is often through professional networks, fellowships, and in-person research that users may find out about collections without full descriptions online. this project’s interviews illustrated that many qualitative researchers would benefit from training in archival principles and practices. most researchers, without knowing how archives are processed, described, and made accessible by archivists, lack the knowledge of collections’ provenance to intuit where collections of relevance may be held in an archives. other repositories may find it useful to carry out similar projects, interviews, or interdisciplinary fellowships at their institutions. our initial nsf grant proposal aimed to develop new anthropological leadership attuned to the needs and practices of both anthropologists and archivists. applied fellowships of this kind help to bridge disciplinary gaps and to realign archival and disciplinary expectations and can encourage the broader use of archival collections as sites for knowledge production and to demonstrate the continued value of archival data for us all. https://doi.org/10.29173/iq955 http://copar.umd.edu/ http://copar.org/ https://transcription.si.edu/ 8/9 marsh, diana e. (2019) research-driven approaches to improving archival discovery, iassist quarterly 43(2), pp. 1-9. doi: https://doi.org/10.29173/iq955 references anderson, k. (2005), tending the wild: native american knowledge and the management of california's natural resources, university of california press, berkeley, ca. applegarth, r. (2014), rhetoric in american anthropology: gender, genre, and science, university of pittsburgh press, pittsburgh, pa. bucciferro, a. (2008), “attacking the backlog: nara archivists mobilize to make unprocessed records available to the public”, prologue magazine, vol. 40 no. 2, pp. 46-51. davis, j.e. (2010), hand talk: sign language among american indian nations, cambridge university press, cambridge, uk. hinsley, c. (1994), the smithsonian and the american indian: making a moral anthropology in victorian america, smithsonian institution press, washington d.c. loring, s. & spiess, a. (2007), “further documentation supporting the former existence of grizzly bears (ursus arctos) in northern quebec-labrador”, arctic, vol. 60 no. 1, pp. 7-16. marks, d. (2014), “the kuna mola”, dress, vol. 40 no. 1, pp. 17-30. marsh, d.e. (2018), “toward inclusive museum archives: user research at the smithsonian's national anthropological archives”, in yun shun susie chung, anna leshchenko, & bruno brulon soares (eds.), defining the museum of the 21st century: evolving multiculturalism in museums in the united states (pp. 129-142), icom/icofom, paris. moon, k. r. (2010), “the quest for music's origin at the st. louis world's fair: frances densmore and the racialization of music”, american music, vol. 28 no. 2, pp. 191-210, doi:10.5406/americanmusic.28.2.0191. naeem, a., knipe, p., nemerov, a., shaw, g. d., & verplanck, a. (2018), black out: silhouettes then and now, princeton university press, princeton, nj. panitch, j. m. (2000), special collections in arl libraries: results of the 1998 survey sponsored by the arl research collections committee, association of research libraries, washington, dc. rich, j. (2012), missing links: the african and american worlds of r.l. garner, primate collector, university of georgia press, athens, ga. schmidt, r.w., seguchi, n., & thompson, j.l. (2011), “chinese immigrant population history in north america based on craniometric diversity”, anthropological science, vol. 119 no. 1, pp. 919. https://doi.org/10.29173/iq955 9/9 marsh, diana e. (2019) research-driven approaches to improving archival discovery, iassist quarterly 43(2), pp. 1-9. doi: https://doi.org/10.29173/iq955 troutman, j.w. (2013), indian blues: american indians and the politics of music, 1879–1934, university of oklahoma press, norman, ok. yakel, e. (2004), “encoded archival description: are finding aids boundary spanners or barriers for users?”, journal of archival organization, vol. 2 nos. 1/2, pp. 63-77. endnotes 1 diana e. marsh is postdoctoral fellow, national anthropological archives, department of anthropology, national museum of natural history, diana.e.marsh@gmail.com. 2 see, for instance, baldwin, d. (2017), language reconstruction and strengthening community : the role of archival resources. paper presented at the irving k. barber learning centre events; fitzgerald, c.m., & linn, m.s. (2013), “training communities, training graduate students: the 2012 oklahoma breath of life workshop”, language documentation & conservation, vol. 7, 185-206; hinton, l. (2013), “the use of linguistic archives in language revitalization: the native california language restoration workshop”, in leanne hinton & ken hale (eds.), the green book of language revitalization in practice (pp. 419-423), brill, boston, ma; roy, l., bhasin, a., & arriaga, s.k. (2011), tribal libraries, archives, and museums: preserving our language, memory, and lifeways, scarecrow press, lanham, md. 3 there are related studies focused on historians and historical collections, such as: anderson, i.g. (2004), "are you being served? historians and the search for primary sources”, archivaria vol. 58; torou, e., akrivi k., costas v., george l., and constantin h. (2010), “historical research in archives: user methodology and supporting tools”, international journal on digital libraries, vol 11, no. 1, pp. 25-36; toms, e.g. & wendy d. (2002), “‘i spent 1 1/2 hours sifting through one large box...’: diaries as information behavior of the archives user: lessons learned”, journal of the american society for information science and technology, vol. 53 no. 14, pp. 1232-38; for a broad study of native user needs, see association of tribal archives, libraries, museums, and miriam jorgensen. sustaining indigenous culture: the structure, activities, and needs of tribal archives, libraries, and museums. association of tribal archives, libraries, and museums, 2012. 4 for repository-based exceptions, see for instance, altman, b. & nemmers, j. (2001), “the usability of on-line archival resources: the polaris project finding aid”, the american archivist, vol. 64 no. 1, pp. 121-31; conway, p. (1994), partners in research: improving access to the nation's archive, archives & museum informatics, pittsburgh, pa; goggin, j. (1986), “the indirect approach: a study of scholarly users of black and women's organizational records in the library of congress manuscript division”, the midwestern archivist, vol. 11, no. 1, pp. 57-67; on applied implementation approaches based on user studies, see brancolini, k.r. (2000), “selecting research collections for digitization: applying the harvard model”, library trends, vol. 18, no. 4, pp. 783-98; daines, j.g. & nimer, c. l. (2011), “re-imagining archival display: creating user-friendly finding aids”, journal of archival organization, vol. 9 no. 1, pp. 4-31; nimer, c. & daines, j.g. (2008), “what do you mean it doesn't make sense? redesigning finding aids from the user's perspective”, journal of archival organization, vol. 6 no. 4, pp. 21632. https://doi.org/10.29173/iq955 mailto:diana.e.marsh@gmail.com 1/15 scoulas, jung mi; de groote, sandra l.; dempsey, paula r. (2020) learning from data reuse: successful and failed experiences in a large public research university library, iassist quarterly 44(1-2), pp. 1-15. https://doi.org/10.29173/iq966 learning from data reuse: successful and failed experiences in a large public research university library jung mi scoulas1, sandra l. de groote2, paula r. dempsey3 abstract this paper illustrates a large research university library’s experience in reusing data for research collected both within and outside of the library. the purpose of the paper is 1) to demonstrate when, why and how data are reused in a large public research university library, 2) to share tips on what to consider when reusing and reproducing data for research data, including issues of replicability and research ethics, and 3) to share challenges and lessons learned from data reuse and reproducibility experiences. this paper presents five proposed opportunities for data reuse conducted by three researchers at the institution’s library, which resulted in three successful instances of data reuse and two failed data reuses. learning from successful and failed experiences is critical to understand what works and what does not work in order to identify best practices for data reuse. this paper will be helpful for librarians who intend to reuse data for research and publication. keywords data reuse, surveys, internal and external data, data practices, reproducibility 1. introduction every day, academic libraries collect a wealth of data such as gate counts, book circulations, study room reservations, and use of online library resources including website and chat logs, in order to capture users’ activities and behaviors and often to use them for decision making (e.g., staffing). furthermore, these data are often reused by researchers in order to demonstrate the library’s value and impact on students’ academic success (allison 2015; lemaistre, shi & thanki 2018; soria, fransen & nackerud 2013; soria, fransen & nackerud 2017) by examining the relationships between library use, students’ academic performance and retention rate, etc. however, the types of challenges and issues researchers encounter, and how to handle the process of data reuse and ensure reproducibility, are understudied. in addition, when reusing data, it was discovered that many published research projects in the library field did not follow data protection practices (e.g., informed consent, anonymization), which may potentially result in violating students’ privacy (briney 2019). in addition to using internally generated data, library researchers have used data collected by outside entities to demonstrate the library’s impact. examples of large datasets that are widely used in the library field include the association of college and research libraries (acrl) library trends & statistics survey4 containing information on staffing, teaching and collections; integrated postsecondary education data system 5(ipeds) academic libraries survey containing information on library resources, services and expenditures; and the national center for education statistics6 (nces) academic libraries survey. these three datasets can be accessed via acrl metrics,7 an online subscription service. another large dataset is from the association of research libraries (arl) statistics,8 a series of annual surveys containing information on library collections, expenditures, staffing and service activities for arl member libraries. many researchers have used the large https://doi.org/10.29173/iq966 https://acrl.countingopinions.com/ https://acrl.countingopinions.com/ https://nces.ed.gov/ipeds/ https://nces.ed.gov/ipeds/ https://nces.ed.gov/ https://www.acrlmetrics.com/ https://www.arlstatistics.org/home https://www.arlstatistics.org/home 2/15 scoulas, jung mi; de groote, sandra l.; dempsey, paula r. (2020) learning from data reuse: successful and failed experiences in a large public research university library, iassist quarterly 44(1-2), pp. 1-15. https://doi.org/10.29173/iq966 datasets described above to demonstrate the library’s impact. mezick (2007) used data from acrl, arl and ipeds in order to examine the correlations between library expenditure, staffing and student retention. haddow and joseph (2010) also analyzed data from arl, ipeds and nces to investigate the relationships between student retention and library use (workstation use and logging into library resources. stewart (2012) analyzed data from nces and ipeds to compare graduation rate and library expenditure per student from 2004 to 2010. crawford (2015) used data from academic libraries survey and ipeds to measure the relationships between institutional expenditures, library expenditures, library use and students’ graduation and retention rates. however, there are challenges in using large scale data in terms of data reuse and research reproducibility (yan et al. 2019). accessing and analyzing these large data sets requires effort, knowledge, skills and expenses such as online subscriptions to access some datasets (e.g., acrl metrics7 and arl statistics8). in particular, it is critical to understand the research design, codes, and data analysis related to statistical software in order to reuse the data or reproduce the results of studies that were conducted by other researchers. academic librarians are aware of the importance of using evidence-based data for decision making and attempting to determine the library’s impact. scholarship in the field would benefit from more comparative studies that would require sharing data across institutions. however, little is known about questions such as “what types of data can be reused in the library field?” “are there any challenges to consider when reusing data or reproducing research?” “what issues need to be considered when reusing data or reproducing research?” the answers to these questions are critical for librarians to gain awareness of what data is available to them, to learn how to use evidence-based data, to make the results reproducible, and to increase research productivity. the purpose of the paper is to: 1) demonstrate when, why and how data are reused in a large public research university library; 2) share tips on what to consider when reusing and reproducing data for research; and 3) share lessons learned from data reuse and reproducibility experiences from a research perspective. this paper will be useful for librarians who are not familiar with reusing existing data to understand what types of data are available for them in their own institutions, where to begin addressing new problems, and how to transform a failed experience for data reuse and reproducibility to a successful experience. this paper provides practical implications for promoting data reuse and reproducibility practices for librarians. it will be helpful for librarians who intend to reuse data and reproduce it in research for publication. 2. literature review 2.1. data reuse and reproducibility in the data life cycle, there is a sequence of the stages of data life, from data creation to data reuse, with data reuse identified as the last stage of its useful life (briney 2015). in elsevier’s website where it describes research data, “reproducible” and “reusable” data are displayed as the highest stages of its life cycle (elsevier 2019). data reuse refers to using secondary or existing data to examine new problems that were not considered in the original study and generate new findings (yoon 2017; zimmerman 2008). the national science foundation’s (nsf) social behavioral and economic (sbe) division subcommittee on reproducible science defined reproducibility as “the ability of a researcher to duplicate the results of a prior study using the same materials and procedures as were used by the original investigator” (bollen, cacioppo, kaplan, krosnick & olds 2015, p. 3). in the report on how to promote research practices, bollen and colleagues considered it “a minimum necessary condition for a finding to be believable and informative” (p. 4). the benefits of data reuse include validating results and potentially increasing research productivity and effectiveness (yoon & kim 2017). research reproducibility requires accuracy of results, transparency of data collection and analysis (bollen et al. 2015), in order to reproduce the findings. in fact, due to the high percentage of results that cannot be reproduced, published articles in various disciplines from the sciences (e.g., https://doi.org/10.29173/iq966 https://www.acrlmetrics.com/ https://www.acrlmetrics.com/ https://www.arlstatistics.org/home https://www.elsevier.com/about/open-science/research-data https://www.elsevier.com/about/open-science/research-data 3/15 scoulas, jung mi; de groote, sandra l.; dempsey, paula r. (2020) learning from data reuse: successful and failed experiences in a large public research university library, iassist quarterly 44(1-2), pp. 1-15. https://doi.org/10.29173/iq966 neuroscience; gilmore, diaz, wyble & yarkoni 2017) to the social sciences (e.g., psychology; open science collaboration 2015) have entered a “reproducibility crisis” (sayre & riegelman 2018). under some federal (e.g., the national science foundation9) and private funding agencies (e.g., the bill & melinda gates foundation10) data sharing policies, researchers who receive grants are obligated to share their data if it is not sensitive data (e.g., gpas and patrons’ ids) (briney 2015). reproducibility is one of the primary reasons of why data sharing is needed (briney 2015). with the increase in data sharing, data is now frequently accessible and is easier to find for either data reuse or validating results (briney 2015). 2.2. data reuse and reproducibility: behaviors and challenges in a study of survey researchers regarding their perceptions and perspectives of data reuse and reproducibility, they were asked about the main reasons for reusing data (yan, huang & palmer 2019). they found that the top two reasons were to “conduct new analysis” (87%) and to “compare results” (70.5%), and the lowest reason was to “reproduce published articles” (18.5%). yan and colleagues further reported that the top problem related to reproducibility was “not enough detail in the published paper on how study was conducted” (85.2%). increased opportunities for sharing and reuse of research and academic data have raised issues around the ethics of data sharing as well as practical barriers to data reuse and reproducibility. in spite of the increase in data sharing, however, finding datasets remains difficult. briney (2015) shared strategies on how to find data for data reuse: look for published articles, search for a “subject-specific index” in your specialty (p. 164), look for discipline-specific data repositories in your field, and check other resources such as re3data.11 yoon (2016) conducted interviews with 23 researchers who reused social science data and found that barriers to data reuse involved: inaccurate descriptive information about the data, difficulty accessing data, difficulties with data format, software, special analytic programs, problems with samples (i.e., too many missing values), issues with original data analysis, and data cleaning. faniel, kriesberg and yakel (2012) found that novice researchers in the process of matching and merging data from multiple sources had difficulty dealing with different time periods and creating unique identifiers. yoon and kim (2017) further studied what factors influenced researchers’ behaviors in data reuse by conducting a survey of 1,528 participants and found that researchers’ perceived usefulness (e.g., increase in research productivity) was the strongest predictor that affected data reuse intentions. additional factors influencing the reuse of data found in their study included concerns about misinterpreting the data, copyright infringement, availability of internal resources, and availability of data repositories. among the types of data capture and reuse in academic libraries is participation in learning analytics initiatives. perry and colleagues (2018) examined how 54 arl member libraries participated in practices, policies and ethical issues about learning analytics within member libraries. the results related to behavior policies for learning analytics revealed that although 70% of respondents obtained approval from the university’s institutional review board (irb) for their learning analytics project, not all respondents informed students about their learning analytics initiative (perry et al. 2018). libraries’ practices to link data pertaining to students’ identifying information associated with their library use (e.g., database logins from library website, book check outs and login from library computer workstation) and their academic success (e.g., gpa and retention rate) draw attention to ethical issues such as maintaining patron privacy, appropriate data handling and de-identification. as perry, et al show, it is not clear whether or not researchers inform students about reusing data from https://doi.org/10.29173/iq966 https://www.nsf.gov/bfa/dias/policy/dmp.jsp https://www.gatesfoundation.org/how-we-work/general-information/information-sharing-approach https://www.gatesfoundation.org/how-we-work/general-information/information-sharing-approach https://www.re3data.org/ 4/15 scoulas, jung mi; de groote, sandra l.; dempsey, paula r. (2020) learning from data reuse: successful and failed experiences in a large public research university library, iassist quarterly 44(1-2), pp. 1-15. https://doi.org/10.29173/iq966 their library use information such as book check outs and database logins, or whether the researchers obtained irb approval for the research project related to measuring the correlations between students’ library usage and their academic achievement and learning outcomes. for example, in both studies conducted by soria and colleagues (2013; 2017), they used students’ library usage data including book check outs, interlibrary loans, online chats with reference librarians, and database logins from various data sources. however, the authors did not clearly explain the security of the data system, and so it was not clear whether they informed students when reusing the data for their project and how they handled students’ personal information (i.e., identification numbers) in the articles (soria et al. 2013; soria et al. 2017). allison (2015) also examined whether students’ library use (checkouts and off-campus access to library resources) had an impact on students’ gpa. similar to soria and colleagues, allison also reused data collected from the library to address the research question. unlike soria and colleagues, allison explained in her article that “the data were then made anonymous by removing the id number that could be linked back to individual student’ records” (p. 33). however, whether she obtained irb approval or whether she informed the students whose data she used was not addressed. some researchers demonstrated in their data reuse practices that they practiced adequate data protection in their published research projects. a good example of data reuse can be seen in the study by de jager, nassimbeni, daniels and d’angelo (2017). de jager and colleagues studied the correlations between undergraduate students’ library use and their gpa using the data obtained from the institution’s data warehouse in the university of cape town, south africa. in the article, they described each step in obtaining the anonymized data (e.g., library visits and checks out of the library materials) so there was no possibility of identifying students. they also addressed the challenges of obtaining library data from the data warehouse like delays in obtaining the assurance of anonymized data, the process of securing data from various sources, and ensuring the “integrity and completeness of the reported data” (2017, p. 5). the data reuse project by lemaistre, shi and thanki (2018) is another good example. lemaistre and colleagues (2018) at the nevada stage college investigated whether or not students’ use of online library resources was correlated with their gpa using the login data on ezproxy. nevada stage college included a data privacy policy and indicated that students have the option to opt out of their private information being saved by contacting the library via the ezproxy log-in page. in their study, lemaistre and colleagues clearly stated the importance of students’ data privacy to protect students’ confidentiality (2018). while data reuse practices vary by discipline, researchers in the library field tend to reuse data collected directly by their institution and library, focusing on measuring the library’s impact on students’ academic success and learning. however, few studies have examined data reuse practices on other types of data such as survey data. due to data privacy policies, not all institutions and libraries have access to students’ data. when researchers encounter data privacy and policy issues, are there alternatives to measuring the library’s value or address reusing existing data? this paper will demonstrate five cases in data reuse practices and how three researchers in a university library experienced successful and failed data reuse practices by focusing on five points: original data description, purpose of the data reuse, level of accessibility, challenges and lessons learned, and outcomes. 3. case specific examples of library data reuse practices 3.1. case 1: student surveys 3.1.1. background and original data description: beginning in spring 2016, the university library has conducted locally developed biannual surveys for students to: 1) assess current student behavior and satisfaction associated with the use of online resources, library services, and the physical library; 2) examine students’ needs related to library resources and services for improvement or expansion of physical library spaces; and 3) determine if there is any correlation between library use and https://doi.org/10.29173/iq966 5/15 scoulas, jung mi; de groote, sandra l.; dempsey, paula r. (2020) learning from data reuse: successful and failed experiences in a large public research university library, iassist quarterly 44(1-2), pp. 1-15. https://doi.org/10.29173/iq966 students’ academic achievements. both the 2016 and 2018 surveys were approved by the institution’s irb. the university library obtained data about students’ demographic information and their grade point average (gpa) from the office of institutional research. the survey results were used to information decision making about which areas of services and resources are needed for improvement. the data was shared with various stakeholders inside and outside of the university library. also, the results were used to determine which areas of services and resources are needed for improvement. 3.1.2. purpose of data reuse: the proposed need for data reuse was 1) to reproduce the results that were reported by the head of assessment and scholarly communications; 2) to measure the value of the library in terms of students’ success from the 2018 student survey data by examining both quantitative and qualitative data; and 3) to use the 2016 and 2018 student survey data to examine any differences between students’ library website use and satisfaction and students’ use of library spaces and satisfaction. 3.1.3. level of accessibility: the survey data was captured in qualtrics (2018 version). following the completion of the survey it was exported as spss and excel files and stored in the university library box folders, a web-based cloud file sharing management service accessible only to the assessment coordinator advisory committee (ac2). the analyzed data and results are stored in the university library’s data warehouse (also a box folder) where it is available to all university library staff excluding student employees. 3.1.4. challenges and lessons learned: data was saved in both spss and excel files. most of the descriptive information about variables in the data dictionary created by the head of assessment and scholarly communications was clear and accurate. several issues occurred when cleaning and analyzing the original data for reuse. first of all, students’ demographic information was coded differently in the 2016 and 2018 survey data. for example, in the 2016 survey the gender code for female was (1) and the code for male was (0), whereas in the 2018 survey the female code was (1) and the code for male was (2). to eliminate any confusion in the data interpretation, all of the codes in the 2016 data were converted to match those in the 2018 data. second, the survey response scales in the 2016 and 2018 survey data were different. for instance, a 6-point likert scale [e.g., from very difficult (1) to i have not used this (6), including (3) neutral] was used in the 2016 surveys, whereas a 5 point-likert scale [e.g., from i have not used this (0) to very easy (4)] was used in the 2018 survey. these different codes and scales were adjusted by converting the ordinal scales to continuous data based on the previous study (preston & colman 2000). the last challenge was to select which questions overlap in both surveys. based on the issues described here, library faculty will maintain the same response scale and codes used in the 2018 student survey for future student surveys. this will allow us to more accurately capture and compare data for making better decisions and improvements in library services. in order for other researchers to replicate and reproduce the published studies, survey instruments and the procedures of data collection and analysis were included in the publications (scoulas & de groote 2019; scoulas & de groote, under review). 3.1.5. outcomes: re-using data from locally developed student surveys expands the existing literature on academic libraries efforts to demonstrate the library impact on students’ academic success. if user surveys are carefully designed at the beginning, there are a great number of potential benefits of data reuse. using locally developed surveys, academic libraries can make decisions using the evidence-based findings for improvement, demonstrate the library value on students’ academic success and learning outcomes, and examine user’s behaviors and attitudes over time to monitor trends. 3.2 case 2: chat with a librarian 3.2.1. background and original data description: transcripts of chat reference interactions between patrons and the university library reference providers are created in the course of ordinary work https://doi.org/10.29173/iq966 6/15 scoulas, jung mi; de groote, sandra l.; dempsey, paula r. (2020) learning from data reuse: successful and failed experiences in a large public research university library, iassist quarterly 44(1-2), pp. 1-15. https://doi.org/10.29173/iq966 serving patron information needs of the uic community. the transcript data are produced by the libchat platform (springshare12), which is commonly used by academic libraries to provide virtual reference services. other platforms that produce similar data include libraryh3lp13 and questionpoint14 (acquired by springshare from oclc in may 2019). the university library retains the transcripts for 5 years for purposes of follow-up with patrons, training, and quality control. 3.2.2. purpose of data reuse: chat transcripts provide a solid basis for understanding what library users need, how library staff work with patrons, and how reference services might be improved. there is extensive research using chat transcripts. a review of such studies from 1995-2010 found the researchers were most concerned with what level of service was provided, who used the service, what questions were asked, and the ways in which providers responded (matteson, salamon & brewster 2011). most of the studies are single-institution case studies. other than a few studies of consortial data (kwon 2006; meert 2009), it is unusual for studies of chat transcripts to provide cross-institutional comparison (dempsey 2017). a recent study of uic library chat transcripts analyzed the extent to which patrons were referred to subject specialists and how chat providers framed the referral (dempsey 2019). on ongoing study examines in-depth reference interactions to gauge whether librarians from a range of institutions who provide virtual reference believe they should have been referred to a subject specialist. 3.2.3. level of access or reuse data: any university library employee with credentials for the libchat system has access to transcripts for the past 5 years. in order to harvest them for research, however, irb approval is needed. the privacy rights of the patrons and the well-being of the chat providers employed by the library must both be considered as risks in balance with potential benefits. patrons have a right to privacy outlined in a uic library policy (uic library 2017). all identifying information is scrubbed from the data, and therefore researchers and readers of the published reports will have no way to identify individuals represented in the data. however, it is possible, though highly unlikely, that a patron’s topic of research could serve as an identifier. investigators using transcripts must consider carefully the details presented in published quotations from the data. in the case of chat reference providers, they might recognize their own work if quoted in a study. this recognition could cause distress if the analysis identifies shortcomings in the service provided, but no library employee should be judged on the basis of one chat interaction. thus, the analysis must take into account the dignity of people doing challenging work in a fast-paced environment and handle critiques in a way that emphasizes the aggregate and does not harm either of the individuals participating in any one chat interaction. 3.2.4. challenges and lesson learned: privacy. de-identifying data is a significant investment of time, because even though names entered by patrons can be scrubbed automatically in libanswers, patrons and providers frequently use one another’s names and provide e-mails and other identifiers that must be deleted individually. data ownership. who owns these data, and under what circumstances can they be shared? transcripts are created as part of ordinary work flows and most likely represent work for hire by the university library. the effort that goes into de-identifying and cleaning the data in other ways does not confer ownership on the researchers; rather it is done as part of one’s professional responsibilities with the purpose of benefitting patrons and the knowledge base of the profession. these efforts provide more benefits the more widely the data are shared, because multi-institutional data are likely to generate more generalizable findings. as noted above, cross-institutional studies are unusual – sharing data to promote replicating studies across institutions would improve knowledge in the field. if possible, establish and document answers to questions of ownership before embarking on data collection and cleaning and share data in a repository such as the qualitative data repository at syracuse university.15 3.2.5. outcomes: re-using naturally occurring data in the form of chat transcripts has contributed to the scholarly conversation as an empirical basis for establishing best practices for navigating the https://doi.org/10.29173/iq966 https://springshare.com/ https://libraryh3lp.com/ https://www.oclc.org/en/questionpoint.html https://qdr.syr.edu/ 7/15 scoulas, jung mi; de groote, sandra l.; dempsey, paula r. (2020) learning from data reuse: successful and failed experiences in a large public research university library, iassist quarterly 44(1-2), pp. 1-15. https://doi.org/10.29173/iq966 reference interview, teaching information literacy in the virtual context, providing patrons relevant resources, and referring patrons to subject specialists. if used thoughtfully to design training tools, these findings have the potential to improve virtual reference across academic libraries. moving toward wider availability of data for replication in cross-institutional studies would have even more impact. 3.3 case 3: library collections and research productivity 3.3.1. background and original data description: the data for this study came from various sources with different purposes. the scopus16 database is the largest source of abstracts and citations of peer reviewed research literature. this database is typically used to find literature on specific topics but for the case being discussed here, it was used to identify publications by institution and the references used in them. the higher education research and development (herd) survey17 contains information of research and development expenditures at u.s academic institutions. arl statistics8 data contains annually conducted survey data from the arl member libraries about collections, expenditure, staffing and service activities. 3.3.2. purpose of data reuse: demonstrating the value and potential impact of the academic library on research at academic institutions can help support arguments for maintaining or increasing funding to the library. recent studies have not explored funding, collection size, and collection use and their relationship with research output (publications) using existing data, nor explored if new library metrics (database searches, journal article downloads) can predict research output at academic institutions. the purpose of this study was to explore the impact of the research library on faculty productivity by using arl statistics, herd expenditure data, and scopus publication information. 3.3.3. level of access or reuse data: the data from the various sources was merged together into one data set. as noted above, the scopus database, which requires an institutional subscription, was used to identify the number of publications produced at specified institutions, and the number of refences used in them. synthesized data on faculty publications and references is not directly available in scopus and the data needed to be searched for and recorded. for example, to find the number of publications for institution y, a search by institution y was conducted and the results limited by year. the publication data for institution y needed to be displayed in a specific way to capture the number of references included in the publications for a given year. both the number of publications and number of references included in the publications needed to be manually recorded in a spreadsheet. herd data, which indicates the research and development expenditures of an institution, is synthesized annually and shared in a downloaded spreadsheet where each row lists an institution and its data. each institution was searched for in a summary table for a specific year, and the needed data was entered into to a spreadsheet. finally, arl data was also retrieved to provide information about the expenditures and resource use of the libraries included in the study. to retrieve the data, pull-down menus are available for institution and data variable. for each year in the study, the institutions included in the study and variables of interest were selected. these data were then exported in spreadsheet format, that was then merged with the data collected from other sources. a subscription is also required to access arl data. 3.3.4. challenges and lesson learned: because the data for this study was collected from three different resources, it was challenging to ensure the three data sets lined up for each institution. for example, some institutions have separate budgets and administrative lines for their health sciences colleges and libraries. it was not always clear from the collected data sets if it was all locations of an institution, or just certain disciplines or cities where data would be captured. familiarity, particularly with large state institutions that have multiple locations was helpful to understand what academic locations would be included under the title of an institution. close inspection of data coverage was needed to ensure that data sets were representative of the same population. one aspect that we https://doi.org/10.29173/iq966 https://www.scopus.com/search/form.uri?display=basic https://www.nsf.gov/statistics/srvyherd/ https://www.arlstatistics.org/home 8/15 scoulas, jung mi; de groote, sandra l.; dempsey, paula r. (2020) learning from data reuse: successful and failed experiences in a large public research university library, iassist quarterly 44(1-2), pp. 1-15. https://doi.org/10.29173/iq966 wanted to explore with this study was how productive faculty were. we were able to obtain the number of publications for an institution through searches in scopus, and arl provides data on the number of full and part-time faculty. but this data was not sufficient to approximate average faculty productivity at an institution because it was not clear how many publications may have been written by individuals other their faculty, including students, fellows, post docs and staff, nor was it clear that faculty would be defined similarly at all institutions. use of the data was also hindered by the ability to readily find data dictionaries that clearly defined and described the data. knowing who to ask to get a definition of a data variables was important. finally, the data collected by the arl was greatly changed between2014 and 2015. while adding new data measures meant exploring the ability of new measures to assess productivity, because the collection of some data points ceased, it meant limiting the number of years that could be retrospectively be studied, as well as possible changes over time. 3.3.5. outcomes: if data from different sources is collected thoughtfully and carefully, data can be merged and reused to explore new relationships. due to uncertainty with the alignment and comprehensiveness of some of the institutional data between data sources, several institutions were excluded from this study. given that it was not possible to determine the average number of publications per faculty by institution using the re-used data, partial correlation analyses were done holding number of students and faculty constant to explore the impact of libraries. 3.4 case 4: university library’s undergraduate engagement program (uep) data 3.4.1. background and original data description: “finals week relaxation station” is a successful uep program that targets undergraduate students and helps them to manage their stress, which is thought to influence their academic success. since fall 2016, an outreach coordinator has collected data from students who participated in this program by asking them to swipe their id cards when visiting the relaxation station. the data contains the date the program occurred and students’ identification numbers. 3.4.2. purpose of data reuse: it was not clear whether this program is providing beneficial services for the targeted audience. further, there is interest in measuring any correlation of students using the finals week relaxation station on their gpa by comparing groups (i.e., one time use vs. more than one time). the outreach coordinators and the assessment coordinator wanted to use this data not only for their internal use (identifying the users’ characteristics) but also to investigate the impact of the program on students’ academic success. 3.4.3. level of access or reuse data: originally, the data was accessible only to one of the outreach program coordinators. after discussing data reuse with the other outreach coordinator for the purposes described above, the first outreach program coordinator shared the raw data via box folder. the data was saved in excel format and each event per semester was saved in a separate excel file. a total of 6 excel files were created. given that the raw data contained only students’ identification numbers, this information was sent to the office of institutional research (oir) to retrieve detailed information about the students: program, class level, gpa for the beginning of the semester and at the end of the semester. the assessment coordinator combined and organized the data and sent it securely to the oir. the process of merging data into one file and retrieving the data from oir took about a month. the assessment coordinator analyzed the data and shared the descriptive statistics with the outreach coordinators. 3.4.4. challenges and lessons learned: the findings were unexpected and interesting. there was interest in publishing the findings by demonstrating the background of the uep and how program impact was measured. however, a couple of critical issues for data reuse were discovered. first, when collecting data, students were not informed of how their data would be used. second, while this project aimed to establish a sustainable and welcoming culture in the library for undergraduate students’ learning and academic success, this project has not received irb approval for data capture https://doi.org/10.29173/iq966 9/15 scoulas, jung mi; de groote, sandra l.; dempsey, paula r. (2020) learning from data reuse: successful and failed experiences in a large public research university library, iassist quarterly 44(1-2), pp. 1-15. https://doi.org/10.29173/iq966 as a research project. the university library is committed to protecting students’ privacy; we could not overlook the ethical issues that may harm students’ privacy and autonomy. given that the findings from the data indicate that students’ use of the finals week relaxation station increased over time, and many students used it more than one time, this information will be valuable for increasing buy-ins inside and outside of the library in order to demonstrate the impact of the program. to proceed with publication of the results, we need to address the ethics of data reuse. as the historical data cannot be used, the outreach coordinators and assessment librarian will need to explore other methods of data capture if there is further interest in publishing impact studies. 3.4.5. outcomes: while this project may not proceed for publication at this time, the project allowed the outreach coordinators to consider how and when data needs to be collected to better understand the users and outcomes of the program. to this end, the collected data can be utilized to demonstrate not only whether the desired outcomes were met, but also if the uep program has an impact on students’ learning outcomes. 3.5 case 5: faculty survey 3.5.1. background and original data description: since spring 2017, the university library has conducted locally developed biannual surveys for faculty to: 1) assess how faculty members utilize library resources (both online and in print) for their teaching, research or scholarship and 2) examine the university faculty’s level of satisfaction with the library’s programs and services. the university library obtained data about faculty’s demographic information from the office of institutional research (email address, faculty status, the highest fte department, etc.). prior to conducting each survey, all of the documents associated with these proposals were submitted to the irb for approval. the 2017 survey was approved by the irb as a research project, whereas the 2019 survey was determined by the irb as a quality improvement project, stating that the 2019 survey project is considered as having “no intent to produce or contribute to generalizable knowledge,” meaning that “this initiative was deemed not human subjects research and was therefore not received by the institutional review board” based on the objectives that the university library proposed. 3.5.2. purpose of data reuse: the proposed need for data reuse were to: 1) to compare the differences in the university faculty’s library use in 2017 and 2019; and 2) to measure the impact of the faculty’s library use and satisfaction on their research productivity. 3.5.3. level of access or reuse data: similar to the student survey data, the faculty survey data was stored in box folders with access restricted to ac2. after analyzing the data, the summarized data was stored in the secure university library data warehouse where it is available to all university library staff excluding student employees. 3.5.4. challenges and lesson learned: as noted above, when submitting the proposal to the university irb, the researchers indicated that the objectives of the projects were to “identify the library resources and services used by the university faculty for teaching, research or scholarship and examine faculty’s perceived importance and level of satisfaction with library support.” based on the objectives stated in the irb application, the university irb determined “this project as a quality improvement project with no intent to produce or contribute to generalizable knowledge.” in other words, this project is no longer considered as a human research project and was not reviewed by the irb. in our effort to reduce the number of questions asked in the survey, we may have reduced the usefulness of the data collected. it is not clear how this determination may impact the use of this data set with previous and future data sets where the data is considered “generalizable”. additionally, as an afterthought, we realized that as part of the data embedded in the survey, we could have also included information about the number of publications of each faculty member. this would have made the results of the survey more useful to us, and we would also have had data that would have made the results more generalizable. based on the issues that the authors encountered, https://doi.org/10.29173/iq966 10/15 scoulas, jung mi; de groote, sandra l.; dempsey, paula r. (2020) learning from data reuse: successful and failed experiences in a large public research university library, iassist quarterly 44(1-2), pp. 1-15. https://doi.org/10.29173/iq966 we learned the importance of original data collection: whether key variables were included at the beginning. once the data is collected, it is done. we cannot go back to collect the data again. that is, if the original data is not sufficient and does not contain key outcome variables (in this case, publications), this will prevent us from reusing the data to draw a meaningful finding. 3.5.5. outcomes: from the lessons that were learned above, we will be mindful not only about focusing on the immediate questions we want to answer, but also obtaining meaningful data that can be used for future decision-making and trends to demonstrate the value of the library. table 1. summary of case studies conducted by three researchers in a large public research university types of data location of data preservation purpose of data reuse outcomes case 1: student survey survey (numeric and text) university box folder, university library data warehouse 1) to reproduce the results 2) to conduct new analysis 3) to compare the results conference presentations (scoulas and de groote, june 17 2019) (scoulas and de groote, june 18 2019) publication (scoulas and de groote 2019) (scoulas and de groote under review) case 2: chat with a librarian transcripts (text) libanswers, springshare servers to conduct a new analysis publication (dempsey 2019) case 3: library collections and research productivity surveys (numeric) arl statistics, scopus and herd to conduct a new analysis publication in progress case 4: undergraduate engagement program students’ identification information (numeric) university box folder, university library data warehouse to conduct a new analysis unable to publish case 5: faculty survey survey (numeric and text) qualtrics server, university box folder, university to conduct a new analysis and compare the results unable to publish https://doi.org/10.29173/iq966 11/15 scoulas, jung mi; de groote, sandra l.; dempsey, paula r. (2020) learning from data reuse: successful and failed experiences in a large public research university library, iassist quarterly 44(1-2), pp. 1-15. https://doi.org/10.29173/iq966 library data warehouse 4. lessons learned from data reuse experiences as shown in the five case examples described above, our goal for this paper is to share what three researchers in a large public research university library experienced throughout the process of data reuse practices for research: what worked, what did not work, and what to improve for the next research project. below is the summary of the key lessons the authors learned during our data reuse practices. key lessons that can be derived from these five case examples: • instrument. maintain key questions and response scales to compare trends over time. the core questions can be a great asset for taking a longitudinal approach. • ethical issues. be sure to obtain informed consent from users when gathering data, even if you are not sure whether research will be conducted at that time or in the future. • privacy: consider the implications of the research for possible violations of user privacy. this should extend beyond financial or legal ramifications and also take into account the participants' dignity and overall well-being. • the importance of including key variables in the original data collection. carefully design the original research project. if key variables are missing, such as the number of publications per each faculty member in the faculty survey, the data is less likely to be reused for demonstrating the library’s value in the faculty’s research productivity. • documentation of data, data collection and analysis: record every procedure of data analysis and the codes used. in addition, save all of the data analysis output. this information will be critical for data reuse, replication and reproducibility in research. • data ownership. when starting a project with existing data, think ahead about rights and responsibilities surrounding those data to establish whether you can share the data, take them with you to a new employer, etc. (cdl uc318). consider making data available in a repository when possible to promote cross-institutional research. • data coverage and definitions. understand the coverage of the data in order to merge like data sets together. if it is not clear if each data set is using the same source to produce the data or if the definitions of a data variable are not specified, data reuse may not be possible. 5. implications for other librarians while most of the literature highlighted the rigorous research practices of journals (e.g., elsevier), funding agencies (e.g., nsf), and libraries with published articles, few studies discuss and share the individual researchers’ successful and failed experiences in data reuse and reproducibility in the library field. if every researcher is committed to learning and following the full process of data management (e.g., proper data storage, recording all the steps of data analysis, documenting any changes in data analysis, data output, and depositing data in archives) within their organizations, they will share their data confidently. further, other researchers can easily reuse data that is available within their organizations and reproduce the results that were conducted within or outside of their organizations. if a researcher is concerned about the issues of sharing raw data due to confidentiality, at least a summary of statistics (e.g., a matrix of correlations) needs to be presented https://doi.org/10.29173/iq966 https://uc3.cdlib.org/2016/09/08/who-owns-your-data/ 12/15 scoulas, jung mi; de groote, sandra l.; dempsey, paula r. (2020) learning from data reuse: successful and failed experiences in a large public research university library, iassist quarterly 44(1-2), pp. 1-15. https://doi.org/10.29173/iq966 in their publications, which will enable other researchers to reproduce the results using those statistics (bollen et al. 2015). this will be useful for librarians who are interested in becoming involved in research and scholarship activities by reusing data that already exists in their organizations and outside of the library, or by practicing reproducibility. to this end, librarians can have empirical evidence for establishing best practices for navigating various projects that will be beneficial for other librarians across academic libraries. references bollen, k, cacioppo, j, kaplan, r, krosnick, ja & olds, jl 2015, social, behavioral, and economic sciences perspectives on robust and reliable science. report of the subcommittee on replicability in science advisory committee to the national science foundation directorate for social, behavioral, and economic sciences. available from: http://web.stanford.edu/group/bps/cgi-bin/wordpress/wpcontent/uploads/2015/09/nsf-robust-research-workshop-report.pdf. [23 december 2019]. briney, ka 2015, data management for researchers: organize, maintain and share your data for research success. pelagic, exeter, uk. briney, ka 2019, ‘data management practices in academic library learning analytics: a critical review’, journal of librarianship and scholarly communication, vol. 7, no. 1. https://doi.org/10.7710/2162-3309.2268 crawford, ga 2015, ‘the academic library and student retention and graduation: an exploratory study’, portal: libraries and the academy, vol. 15, no. 1, pp. 41–57. https://doi.org/10.1353/pla.2015.0003 de jager, k, nassimbeni, m, daniels, w & d’angelo, a 2017, ‘the use of academic libraries in turbulent times’, performance measurement and metrics, vol. 19, no. 1, pp. 40–52. https://doi.org/10.1108/pmm-09-2017-0037 dempsey, pr 2017, 'resource delivery and teaching in live chat reference: comparing two libraries.’ college & research libraries, vol. 78, no. 7, pp. 898–919. https://doi.org/10.5860/crl.78.7.898 dempsey, pr 2019, ‘chat reference referral strategies: making a connection, or dropping the ball?’, college & research libraries, vol. 80, no. 5, pp. 674-93. https://doi.org/10.5860/crl.80.5.674 elsevier 2019, ‘research data’. available from: https://www.elsevier.com/about/openscience/research-data. [23 december 2019]. faniel, im, kriesberg, a & yakel, e 2012, ‘data reuse and sensemaking among novice social scientists’, proceedings of the american society for information science and technology, vol. 49, no. 1, pp. 1-10. https://doi.org/10.1002/meet.14504901068 gilmore, ro, diaz mt, wyble, ba & yarkoni t 2017, ‘progress toward openness, transparency, and reproducibility in cognitive neuroscience’, physiology & behavior, vol. 176, no. 1, pp. 139–48. https://doi.org/10.1016/j.physbeh.2017.03.040 haddow, g & joseph, j 2010, ‘loans, logins, and lasting the course: academic library use and student retention’, australian academic and research libraries, vol. 41, no. 4, pp. 233–244. https://doi.org/10.1080/00048623.2010.10721478 https://doi.org/10.29173/iq966 http://web.stanford.edu/group/bps/cgi-bin/wordpress/wp-content/uploads/2015/09/nsf-robust-research-workshop-report.pdf http://web.stanford.edu/group/bps/cgi-bin/wordpress/wp-content/uploads/2015/09/nsf-robust-research-workshop-report.pdf https://doi.org/10.7710/2162-3309.2268 https://doi.org/10.1353/pla.2015.0003 https://doi.org/10.1108/pmm-09-2017-0037 https://doi.org/10.5860/crl.78.7.898 https://doi.org/10.5860/crl.80.5.674 https://www.elsevier.com/about/open-science/research-data https://www.elsevier.com/about/open-science/research-data https://doi.org/10.1002/meet.14504901068 https://doi.org/10.1016/j.physbeh.2017.03.040 https://doi.org/10.1080/00048623.2010.10721478 13/15 scoulas, jung mi; de groote, sandra l.; dempsey, paula r. (2020) learning from data reuse: successful and failed experiences in a large public research university library, iassist quarterly 44(1-2), pp. 1-15. https://doi.org/10.29173/iq966 kwon, n 2006, ‘user satisfaction with referrals at a collaborative virtual reference service’, information research: an international electronic journal, vol. 11, no. 2, p. n2. available from: http://www.informationr.net/ir/11-2/paper246.html. [23 december 2019]. lemaistre, t, shi, q & thanki, s 2018, ‘connecting library use to student success’, portal: libraries and the academy, vol. 18, no. 1, pp. 117–40. https://doi.org/10.1353/pla.2018.0006 matteson, ml, salamon, j & brewster, l 2011, ‘a systematic review of research on live chat service’, reference & user services quarterly, vol. 51, no. 2, pp. 172-89. https://doi.org/10.5860/rusq.51n2.172 meert, dl & given, lm 2009, ‘measuring quality in chat reference consortia: a comparative analysis of responses to users’ queries’, college & research libraries. vol. 70, no. 1, pp. 71–84. https://doi.org/10.5860/0700071 mezick, em 2007, ‘return on investment: libraries and student retention’, journal of academic librarianship, vol. 33, no. 5, pp. 561–566. https://doi.org/10.1016/j.acalib.2007.05.002 open science collaboration 2015, ‘estimating the reproducibility of psychological science’, science, vol. 349, no. 6251. https://doi.org/10.1126/science.aac4716 perry, mr, briney, ka, goben, a, asher, a, jones, kml, robertshaw, mb & salo, d 2018, ‘learning analytics’, spec kit 360. washington, dc: association of research libraries. available from: https://publications.arl.org/learning-analytics-spec-kit-360/. [23 december 2019]. preston, cc & colman, am 2000, ‘optimal number of response categories in rating scales: reliability, validity, discriminating power, and respondent preferences’, acta psychologica, vol. 104, no. 1, pp. 1–15. https://doi.org/10.1016/s0001-6918(99)00050-5 sayre, f & riegelman, a 2018, ‘the reproducibility crisis and academic libraries’, college & research libraries, vol. 79, no. 1, pp. 2–9. https://doi.org/10.5860/crl.79.1.2 scoulas, jm & de groote, sl 2019, ‘factors affecting university students’ library visits in person and online using a multiple regression approach’, paper presentation, evidence based library and information practice 10, glasgow, united kingdom, june 17. available from: https://eblip10.org/abstracts/tabid/8487/default.aspx#scoulas. [23 december 2019]. scoulas, jm & de groote, sl 2019, ‘assessing the university library’s impact on students’ academic performance’, poster presentation, evidence based library and information practice 10, glasgow, united kingdom, june 18. available from: https://eblip10.org/abstracts/tabid/8487/default.aspx#scoulas2. [23 december 2019]. scoulas, jm & de groote, sl 2019, ‘the library’s impact on university students’ academic success and learning’, evidence based library and information practice, vol. 14, no. 3, pp. 2–27. https://doi.org/10.18438/eblip29547 stewart, c 2012, ‘an overview of acrl metrics, part ii: using nces and ipeds data’, journal of academic librarianship, vol. 38, no. 6, pp. 342–45. https://doi.org/10.1016/j.acalib.2012.09.018 https://doi.org/10.29173/iq966 http://www.informationr.net/ir/11-2/paper246.html https://doi.org/10.1353/pla.2018.0006 https://doi.org/10.5860/rusq.51n2.172 https://doi.org/10.5860/0700071 https://doi.org/10.1016/j.acalib.2007.05.002 https://doi.org/10.1126/science.aac4716 https://publications.arl.org/learning-analytics-spec-kit-360/ https://doi.org/10.1016/s0001-6918(99)00050-5 https://doi.org/10.5860/crl.79.1.2 https://eblip10.org/abstracts/tabid/8487/default.aspx#scoulas https://eblip10.org/abstracts/tabid/8487/default.aspx#scoulas2 https://doi.org/10.18438/eblip29547 https://doi.org/10.1016/j.acalib.2012.09.018 14/15 scoulas, jung mi; de groote, sandra l.; dempsey, paula r. (2020) learning from data reuse: successful and failed experiences in a large public research university library, iassist quarterly 44(1-2), pp. 1-15. https://doi.org/10.29173/iq966 uic university library 2017, ‘user privacy policy.’ available from: https://library.uic.edu/about/policies#privacy. [23 december 2019]. yan, a, huang, c & palmer, cl 2019, ‘data reuse and reproducibility in earth system science: a survey of current practices, barriers, and expectations’. available from: https://www.essoar.org/doi/pdf/10.1002/essoar.10500464.1. [23 december 2019]. yoon, a 2016, ‘red flags in data: learning from failed data reuse experiences’, proceedings of the association for information science and technology, vol. 53, no. 1, pp. 1–6. https://doi.org/10.1002/pra2.2016.14505301126 yoon, a 2017, ‘data reuse trust development’, journal of the association for information science and technology, vol. 68, no. 4, pp. 946–56. https://doi.org/10.1002/asi.23730 yoon, a & kim, y 2017, ‘social scientists’ data reuse behaviors: exploring the roles of attitudinal beliefs, attitudes, norms, and data repositories’, library and information science research, vol. 39, no. 3, pp. 224–33. https://doi.org/10.1016/j.lisr.2017.07.008 zimmerman, as 2008, ‘new knowledge from old data: the role of standards in the sharing and reuse of ecological data’, science, technology, & human values, vol. 33, no. 5, pp. 631–52. https://doi.org/10.1177%2f0162243907306704 endnotes 1 jung mi scoulas is a clinical assistant professor and assessment coordinator in the richard j. daley library, university of illinois at chicago. correspondence should be addressed to jung mi scoulas and can be reached by email: jscoul2@uic.edu 2 sandra l. de groote is a professor and the head of assessment and scholarly communications in the richard j. daley library, university of illinois at chicago. email: sgroote@uic.edu 3 paula r. dempsey is an assistant professor and the head of research services & resources in the richard j. daley library, university of illinois at chicago. email: dempseyp@uic.edu 4https://acrl.countingopinions.com/ 5https://nces.ed.gov/ipeds/ 6https://nces.ed.gov/surveys/libraries/ 7https://www.acrlmetrics.com 8https://www.arlstatistics.org/home 9https://www.nsf.gov/bfa/dias/policy/dmp.jsp 10https://www.gatesfoundation.org/how-we-work/general-information/information-sharingapproach 11https://www.re3data.org 12https://springshare.com 13https://libraryh3lp.com/ https://doi.org/10.29173/iq966 https://library.uic.edu/about/policies#privacy https://www.essoar.org/doi/pdf/10.1002/essoar.10500464.1 https://doi.org/10.1002/pra2.2016.14505301126 https://doi.org/10.1002/asi.23730 https://doi.org/10.1016/j.lisr.2017.07.008 https://doi.org/10.1177%2f0162243907306704 https://acrl.countingopinions.com/ https://nces.ed.gov/ipeds/ https://nces.ed.gov/surveys/libraries/ https://www.acrlmetrics.com/ https://www.arlstatistics.org/home https://www.nsf.gov/bfa/dias/policy/dmp.jsp https://www.gatesfoundation.org/how-we-work/general-information/information-sharing-approach https://www.gatesfoundation.org/how-we-work/general-information/information-sharing-approach https://www.re3data.org/ https://springshare.com/ https://libraryh3lp.com/ 15/15 scoulas, jung mi; de groote, sandra l.; dempsey, paula r. (2020) learning from data reuse: successful and failed experiences in a large public research university library, iassist quarterly 44(1-2), pp. 1-15. https://doi.org/10.29173/iq966 14https://www.oclc.org/en/questionpoint.html 15https://qdr.syr.edu/ 16https://www.scopus.com/search/form.uri?display=basic 17https://www.nsf.gov/statistics/srvyherd/ 18https://uc3.cdlib.org/2016/09/08/who-owns-your-data https://doi.org/10.29173/iq966 https://www.oclc.org/en/questionpoint.html https://www.scopus.com/search/form.uri?display=basic https://www.nsf.gov/statistics/srvyherd/ https://uc3.cdlib.org/2016/09/08/who-owns-your-data 1/20 castro, joão aguiar; rodrigues, joana; matos, paula mena; sales, célia; ribeiro, cristina (2023) getting in touch with metadata: a ddi subset for fair metadata production in clinical psychology, iassist quarterly 47(1), pp. 1-19. doi: https://doi.org/ 10.29173/iq1008 getting in touch with metadata: a ddi subset for fair metadata production in clinical psychology joão aguiar castro; joana rodrigues; paula mena matos; célia sales; cristina ribeiro1 abstract when addressing metadata with researchers, it is important to use models that include familiar domain concepts. in the social sciences, the data documentation initiative (ddi) is a well-accepted source of such domain concepts. to create data and metadata, that is findable, accessible, interoperable, and reusable (fair), it is necessary to establish a compact set of ddi elements that meet project requirements and are likely to be adopted by researchers inexperienced with metadata creation. over time, we have engaged in interviews and data description sessions with research groups in the social sciences, identifying a manageable ddi subset. together, recent clinical psychology project dealing with risk assessment for hereditary cancer, considered the inclusion of a ddi subset for the production of metadata that are timely and interoperable with data publication initiatives in the same domain. taking the ddi subset identified by the data curators, we present a preliminary assessment of its use as a realistic effort on the part of the researchers, taking into consideration the metadata created in two data description sessions, the effort involved, and the overall metadata quality. a follow-up questionnaire was used to assess the perspectives of the researchers regarding data description. keywords research data management, metadata, fair, data documentation initiative; clinical psychology introduction the fast-paced growth of scientific production and the risk of data being permanently lost (vines, 2014) have prompted funding agencies to define policies to promote data fairness and require data management plans (european commission, 2016a). with the interest in research data management (rdm) globally on the rise (perrier et al., 2017), researchers are increasingly aware of the need to develop adequate skills to organize their data. in this context, metadata production is an essential activity that involves all the stages of well-managed research data to ultimately enable reuse. to promote open science and the adoption of the findability, accessibility, interoperability, and reusability (fair) principles, the european commission expert group on fair data recommended the development of cases to further engage communities and the provision of tools to make metadata production as easy as possible for researchers (european commission, 2018). moreover, the libraries for research data interest group of the research data alliance recognizes that direct training is an effective way to make people aware of the importance of data management best practices (clare et al., 2019). https://doi.org/%2010.29173/iq1008 2/20 castro, joão aguiar; rodrigues, joana; matos, paula mena; sales, célia; ribeiro, cristina (2023) getting in touch with metadata: a ddi subset for fair metadata production in clinical psychology, iassist quarterly 47(1), pp. 1-19. doi: https://doi.org/ 10.29173/iq1008 our experimental work at the university of porto, under the tail project2, focused on the development of an rdm workflow that integrated a set of different tools depending on the requirements of researchers (ribeiro et al., 2018). the tail project regarded researchers as core rdm stakeholders who needed straightforward workflows. in this sense, the availability of tools to support data organization and metadata creation at the beginning of research projects can be a determinant to improve data management practices. thus, and in order to motivate small research groups with no time or funding for data curation, the tail team developed dendro3, an open-source platform designed to help researchers describe their data, fully built on linked open data (rocha da silva, ribeiro and lopes, 2018). dendro includes domain-specific metadata models to address disciplinary data description requirements (castro et al., 2017). the integration of tools in the research workflow designed during the tail project resulted from the assessment of requirements of several researchers from different domains, who were contacted individually over time. this enabled the team to obtain feedback from a diverse panel of researchers. contact with researchers at the university of porto had been previously established with a scoping study (ribeiro and fernandes, 2011) sent through the deans of its 15 schools in 2011. this scoping study can be regarded as a preliminary effort to poll the availability of researchers. during the tail project, some sessions were organized to disseminate the project among researchers. one of these sessions was targeted at a group of researchers affiliated with the faculty of psychology and educational sciences of the university of porto (fpceup). further contacts were made with two researchers working in family psychology and another working in clinical psychology, who had shown motivation to adopt measures leading to better practices and were starting a new project. the work described in this paper results from the collaboration between the tail project and members of together4, a project in the psycho-oncology domain. together ran from july 2018 to june 2021 and was a partnership between the fpceup and the portuguese oncology institute (ipoporto), a state-run institute for healthcare and research. the goal of together was to study the process of psychosocial adaptation of individuals and families enrolled in genetic counselling to assess and manage the increased risk of hereditary cancer syndromes. over a period of two years, individuals enrolled for genetic testing and their families were monitored regarding psychological and relational variables, as well as their needs and preferences for care. a longitudinal design with a mixed qualitative and quantitative methodology included semi-structured interviews, self-report questionnaires, and documentary analysis of clinical records. the project followed a participatory approach, with a collaborative panel of users and professionals involved in methodological decisions and the interpretation and dissemination of results. this project was a first step toward a strategy of family-centered psychological care, indispensable in personalized preventive medicine of hereditary cancer. https://doi.org/%2010.29173/iq1008 3/20 castro, joão aguiar; rodrigues, joana; matos, paula mena; sales, célia; ribeiro, cristina (2023) getting in touch with metadata: a ddi subset for fair metadata production in clinical psychology, iassist quarterly 47(1), pp. 1-19. doi: https://doi.org/ 10.29173/iq1008 table 1: overview of the tail and of the together projects tail the tail project focused on providing researchers with adequate tools to organize, describe, and publish their data. the tail team developed an rdm workflow based on the integration of different tools taking into account the requirements of a panel of researchers built over time. together using a longitudinal mixed methods design, the purpose of the together project was to study the process of psychosocial adjustment of both unaffected individuals undergoing genetic testing and their families. more specifically, it intended to analyze the mechanisms by which psychosocial factors and clinical factors shape psychosocial outcomes of genetic testing. the ultimate goal is to gather knowledge that informed an integrated family-centered care for inherited cancer syndromes. in this paper, we focus on the preliminary steps in the development of a ddi-based ontology to support researchers to describe data. our primary objective is to raise rdm awareness and simplify the description of data to promote sharing of project data. the cooperation between researchers and data curators is the backbone of our approach, so we aim to address the perspective of researchers when they become aware of data management tasks and have to reconcile them with the already demanding research work. this requires providing the support they deem necessary in adopting rdm practices, such as the organization and description of data in a first stage. our interactions with researchers showed that most are inexperienced with metadata and therefore, to achieve our objectives, we opted for a ddi subset that, for the sake of interoperability, takes into account the already ddi-based list of metadata fields recommended by the interuniversity consortium for political and social research (icpsr)5, already ddi-based, since it stands out as the world´s largest social science data repository. adopting the ddi subset in the early stages of research also meant that, if the group decided to make their data available, fair metadata would already be available by the time they proceeded to the deposit stage. next, we provide a brief overview of related works and of the dendro platform used in the description tasks. then, we proceed with a description of the steps in the development of an ontology based on the ddi subset, which is followed by a description of the activities performed in this work, including the interviews and how we prepared the data description sessions. finally, we highlight the organization and metadata practices of interviewed researchers, the metadata created during the data description sessions, taking into account the number of descriptors filled in and the amount of time spent in the activity, and the feedback of participants regarding their experience in metadata production. support for researchers in metadata production https://doi.org/%2010.29173/iq1008 4/20 castro, joão aguiar; rodrigues, joana; matos, paula mena; sales, célia; ribeiro, cristina (2023) getting in touch with metadata: a ddi subset for fair metadata production in clinical psychology, iassist quarterly 47(1), pp. 1-19. doi: https://doi.org/ 10.29173/iq1008 with their domain expertise and proximity to data collection, researchers have the potential to create highly detailed descriptions about their data, which makes them key stakeholders in fair metadata production. even though data curators are experts at making sure that metadata follows a set of rules to ensure data is findable, accessible and interoperable, their limited ability to capture domain-specific knowledge in the metadata can hinder reusability. from interviews carried out with 23 quantitative social scientists who failed data reuse experiences (yoon, 2016), it was found that access and interoperability were chief primary conditions for successful data reuse, whilst understanding data documentation was less of an issue, at least for experienced researchers, though the process was still seen as challenging. the lack of support in reusing data was the most prominent issue in the reported failed data reuse experiences, making it necessary to establish support systems for those willing to reuse data. in another study, 13 social scientists were interviewed to assess which factors influenced the perceptions and experiences of researchers in attempts to reuse data. it was concluded that data documentation was, among other factors, an important enabling factor for data reuse (curty, 2016). an institutional study conducted to evaluate data management skills and including, both graduate students and postdoctoral researchers, concluded that many researchers were frustrated when former colleagues left without providing annotations of the completed work. consistent data description and organization were regarded as challenges given the different workflows, practices, and value concepts of individuals. a practical solution to address this limitation was the provision of a short description to enable group members to understand the research workflow (wiley and kerby, 2018). by comparing the metadata created by researchers and information professionals, white (2014) found that researchers were more focused on the details and produced more granular metadata. however, the same study found that there was a difference between what was created for personal use and what was created in a formal deposit setting, as more descriptive metadata was added in the deposit stage. moreover, a study with researchers from the center for embedded networked sensing also demonstrated that researchers rarely created documentation that was not directly tied to their own personal use, and therefore data sharing with users from outside of their immediate projects was rare (mayernik, 2011). hence, it seems necessary to combine the skills of data curators and researchers to move the production of fair data and metadata forward. there are several initiatives to support metadata creation in a number of disciplines. for instance, the research data alliance, in the context of the metadata standards directory working group6, lists available metadata standards by domain, such as ddi7 for the social and behavioral sciences and the darwin core8 for biodiversity. the directory also includes domain-neutral standards for different functions, namely the common european research information format (cerif)9 for recording research activity and prov10 for data quality and reliability. however, most standards target data description only at the end of the research workflow and their adoption by researchers can be hard (qin and li, 2013). according to qin et al. (2012), it makes more sense to develop specific goal-oriented metadata schemes, as smaller and more specific schemes will likely increase their adoption by researchers. an analysis of several metadata standards corroborated https://doi.org/%2010.29173/iq1008 5/20 castro, joão aguiar; rodrigues, joana; matos, paula mena; sales, célia; ribeiro, cristina (2023) getting in touch with metadata: a ddi subset for fair metadata production in clinical psychology, iassist quarterly 47(1), pp. 1-19. doi: https://doi.org/ 10.29173/iq1008 the idea that these do not follow principles of simplicity and sufficiency, since most of them lacked a minimal set of essential domain elements (willis, greenberg and white, 2012). as a result, disciplinary vocabularies for research data description are mostly underused, so there is a need to implement more effective processes for the adoption of vocabularies by research communities (european commission, 2018). in this context, we opted for a top-down approach to vocabulary development, with the selection of a set of essential domain elements tailored to describe data from the beginning of the research workflow. the goal was for the metadata model to fit the metadata needs of both the together project and of similar projects, and for it to be simple enough to encourage researchers in their first experience describing data. thus, we decided to develop a minimalist ontology having the ddi as the reference. working with a ddi subset enabled an incremental build-up of the ontology according to researchers’ needs, which were identified by engaging them in metadata production activities. this approach was made possible by the existence of a staging platform for data organization and description, which provided researchers with a training environment for metadata production and with the flexibility to combine the most relevant metadata elements for their data. dendro, a staging platform for data description we sought to simplify and promote fair metadata production at the university of porto by embedding the dendro platform in the research workflow. dendro (rocha da silva, 2016) was developed within the tail project and is an open-source, collaborative data organization platform that promotes the description of data from the moment of its creation. dendro follows a file management structure that resembles popular cloud storage environments, with additional collaborative capabilities common in semantic wikis. the effort in creating metadata is reduced by the incremental description of data, since project members can add and fill in new metadata elements at different times to enrich the quality of the metadata records. figure 1 depicts dendro´s data description user interface. the folder and file management panel is on the left, while some of the available vocabularies are accessible on the right-hand side. each vocabulary has its own set of descriptors. https://doi.org/%2010.29173/iq1008 6/20 castro, joão aguiar; rodrigues, joana; matos, paula mena; sales, célia; ribeiro, cristina (2023) getting in touch with metadata: a ddi subset for fair metadata production in clinical psychology, iassist quarterly 47(1), pp. 1-19. doi: https://doi.org/ 10.29173/iq1008 figure 1: dendro data description user interface dendro favors the compliance with the fair guidelines in several aspects. findability is achieved by assigning persistent identifiers and by allowing the creation of rich metadata records. in dendro, researchers can produce a versatile description for each dataset and have the flexibility to combine descriptors from multiple vocabularies. these vocabularies can either be domain-specific or developed with the participation of researchers with whom the tail team collaborates. the latter mostly happens when there are no available standards for very particular applications such as the double cantilever beam and the hydrogen generation vocabularies represented in figure 1. other vocabularies can combine elements from different standards. for example, for the biodiversity evolution studies and biological oceanography we selected descriptors from multiple standards which we combined with new descriptors created based on the specificity of each experiment (castro et al., 2017). finally, some vocabularies can be based on a single standard if it is comprehensive enough to cover the identified data description requirements. a good example of this is our work with the ddi standard, detailed in the next section. the data and metadata are accessible by their identifier, using a standardized protocol that is open, free, and universally implementable. interoperability relies on the use of a formal, accessible, and broadly applicable language for knowledge representation. dendro uses linked open data at the core, which encourages data curators to model ontologies that satisfy the needs of each specific domain while maintaining the interoperability characteristics of the ontology itself. the metadata also meets domain-specific community requirements to make data reusable. besides the full representation of the dublin core terms11, we reuse concepts from disciplinary standards, as mentioned above. moreover, dendro integrates with several data repositories, e.g., ckan instances and eudat's b2share (silva et al., 2018). a package containing the dataset and its metadata can be submitted to the intended repository in the final step of the process. the adoption and combination https://doi.org/%2010.29173/iq1008 7/20 castro, joão aguiar; rodrigues, joana; matos, paula mena; sales, célia; ribeiro, cristina (2023) getting in touch with metadata: a ddi subset for fair metadata production in clinical psychology, iassist quarterly 47(1), pp. 1-19. doi: https://doi.org/ 10.29173/iq1008 of multiple descriptors address the data documentation limitations associated with generalist repositories meant to cover several research communities, as identified by assante et el. (2016). development of the ddi subset ontology in the course of our rdm activities, we established several contacts with experts from a diversity of domains through open data management sessions or word-of-mouth recommendations. since data reuse depends on the knowledge of variables, data collection methodologies, and experimental parameters, engaging with domain experts is a useful approach to identify relevant concepts for the production of domain-specific metadata, as well as a favorable pretext for researchers to improve their rdm awareness. once validated by researchers, those concepts can be formalized as data properties (using the protégé editor, a free and open-source ontology software), with specified rdf:labels and rdf:comments, in a lightweight ontology (castro, rocha da silva and ribeiro, 2014). these labels define the natural language representation of the descriptor and the comment is used for the definition presented in tool interfaces. from the moment we started our contacts with social science researchers from the university of porto, we proceeded to identify a first set of descriptors based on the ddi, namely: data collection methodology; data source; sample size; external aid; kind of data; and universe, as depicted in figure 2 (amorim et al., 2015). further feedback from social scientists focused on the importance of describing the methodology and the need for additional finer descriptors in order to enrich the metadata. figure 2: first set of ddi descriptors implemented in the dendro platform (amorim, 2015) https://doi.org/%2010.29173/iq1008 8/20 castro, joão aguiar; rodrigues, joana; matos, paula mena; sales, célia; ribeiro, cristina (2023) getting in touch with metadata: a ddi subset for fair metadata production in clinical psychology, iassist quarterly 47(1), pp. 1-19. doi: https://doi.org/ 10.29173/iq1008 as we pursued the contacts in the social sciences, there was the need to extend the number of descriptors in the ddi subset included in dendro, so we looked at the metadata recommendations for data deposit in the icpsr repository. we assumed that the metadata fields recommended for the deposit in the icpsr are intuitive for both domain researchers and data curators with less social sciences expertise. moreover, combined with dublin core metadata they offer the guarantee of interoperability and dissemination of data outside the project. the intention is to have data described from the outset and to be ready for deposit. additional descriptors that we added to the ddi subset include analysis unit; dependent / independent dimension; sampling procedure; summary statistics; and time method. we also uploaded the ddi-rdf discovery vocabulary (disco)12 to dendro. approach in order to promote rdm awareness and simplify the description of data by the researchers from the together project, our approach consisted of a set of interactions, using a combination of techniques. the first contact with the researchers from the together project took place before the launch of the project. more specifically, it happened in the general meeting with a group of researchers from the psychology and educational sciences domain, in the context of the dissemination of the objectives of the tail project to the scientific community of the university of porto. the researchers that were preparing the together project expressed their interest in developing knowledge to adopt rdm measures and agreed to collaborate in the proposed activities. figure 3 is a timeline of the contacts with the researchers after the initial general meeting. we scheduled interviews to assess practices and domain-specific metadata requirements in november 2017 and december 2017, and researchers were also involved in data description sessions at the beginning and the end of 2018. between the first interview and the first data description session, we developed the metadata model to make sure that ddi descriptors were available in the dendro platform. figure 3: timeline of the activities and participants in the study. https://doi.org/%2010.29173/iq1008 9/20 castro, joão aguiar; rodrigues, joana; matos, paula mena; sales, célia; ribeiro, cristina (2023) getting in touch with metadata: a ddi subset for fair metadata production in clinical psychology, iassist quarterly 47(1), pp. 1-19. doi: https://doi.org/ 10.29173/iq1008 the members of the together project who participated in this work were its principal investigator (r01) and the co-principal investigator (r02), both senior researchers with a phd. in both cases, the frequency of data usage is low given the leadership, monitoring, and management roles assumed in their current projects. two other postdoctoral researchers, (r03) and (r04), also participated in this work, given their collaboration with r02 in other projects and their frequent activity in the collection and analysis of data. diagnostic interview the method used in the tail project to engage researchers in data management started with a diagnostic in the form of a semi-structured interview based on the data curation profile toolkit, interview sheet (carlson, 2012). this interview sheet is designed to develop the data curation profile of particular projects. the sheet is structured in modules for specific stages in the data life-cycle and takes into account researchers’ practices and perspectives to guide the conversation, namely details about the dataset, data sharing, repository usage, and organization and description of data. the interviews focused on the researchers´experiences in previous projects, as well as their general research interests and background, and were conducted in late 2017. the interview with r01 lasted 92 minutes and the one with r02 took 90 minutes. the interviews were transcribed and coded, in portuguese, using the atlas.ti software. six main code categories, detailed below, were defined to markup relevant statements. other annotations were made freely to highlight important excerpts that did not fit the predefined categories. this task was performed by a single analyst in the context of a doctoral thesis that involved several other interviews. i. demographic information: information such as professional title, data usage frequency, and level of metadata expertise; ii. awareness: statements that show that the interviewee is aware or unaware of a given data management topic, which can be raised during the interview itself; iii. share: statements that indicate interest in data sharing and issues related to data sharing; iv. organization practice: statements that describe tasks performed by the researcher to organize data, both for problem-solving activities and perceived issues; v. annotation practice: statements that encompass activities to document data, from ad-hoc metadata practices to standard metadata usage; vi. reuse perspective: general statements concerning data reuse potential and positive or negative experiences with data reuse. https://doi.org/%2010.29173/iq1008 10/20 castro, joão aguiar; rodrigues, joana; matos, paula mena; sales, célia; ribeiro, cristina (2023) getting in touch with metadata: a ddi subset for fair metadata production in clinical psychology, iassist quarterly 47(1), pp. 1-19. doi: https://doi.org/ 10.29173/iq1008 data description sessions in our workflow, a data description session is an activity designed to introduce and train researchers in metadata creation. when scheduling the sessions with the researchers, they are asked to select a dataset to describe (if possible, one mentioned during the interview) from an ongoing project or a recent publication. the biggest constraints to scheduling these sessions were related to researchers' busy schedules and lack of interest in participating. in the data description sessions, we started by introducing researchers to dendro. we briefly demonstrated features, like creating a new project and filling metadata, adding contributors, making backups and keeping versions. each researcher was then asked to create a folder and upload a dataset. after this step, we explained in detail the choices that could be made in the descriptors panel, recommending ddi and dublin core descriptors according to the characteristics of the datasets generated by the researchers. during the sessions the selection of metadata elements was mostly up to the researcher we stepped in when requested or when we realized that the researcher was not progressing in the task. in case there was an opportunity, we advised them on the meaning of descriptors or the use of metadata. audio from the sessions was recorded with consent and deleted after the transcription of the relevant events and comments, which were then used to complement the analysis of the produced metadata. the audio was also used to mark the moment when researchers started and finished the description, i.e., to ascertain the session duration. with the experience accumulated at the university of porto, we found that these sessions tended to last approximately 30 minutes. in general, the shorter sessions took place in experimental domains, where many metadata fields could be represented with a numerical value. in one case, a researcher from an experimental domain completed 28 fields in 16 minutes. at the other extreme, a researcher from a sociology area completed 21 fields in 75 minutes. here, the metadata had a substantial volume of text and the researcher consulted documents and justified most choices. a few weeks after the sessions, we asked the researchers to fill out a brief online questionnaire to get additional feedback, namely concerning the perceived usefulness of data description for their research purposes, the degree of interest regarding rdm activities, and the most important factors for rdm engagement. we also showed a document to the researchers with the metadata record created in dendro and asked them to judge whether the information was sufficient or if more information might be provided. https://doi.org/%2010.29173/iq1008 11/20 castro, joão aguiar; rodrigues, joana; matos, paula mena; sales, célia; ribeiro, cristina (2023) getting in touch with metadata: a ddi subset for fair metadata production in clinical psychology, iassist quarterly 47(1), pp. 1-19. doi: https://doi.org/ 10.29173/iq1008 results data organization and metadata practices in the interview, r01 addressed a study related to methods of assessing general psychological wellbeing in mental health contexts, with the goal to understanding how individualized measurements can add value when compared with standardized ones. this study generated self-reported data collected over time in a clinical context. metadata was a concept vaguely understood in the form of keywords by r01. to keep track of data, r01 resorted to meaningful file names and the date of recently saved files, although this strategy was recognized as not always effective. for versioning information, a recurring solution was to write a complementary text file that served as an alert when the database was shared or when the information was directly edited. a specific data sharing challenge identified by r01 related to security and confidentiality, particularly when matching codes to combine datasets from questionnaires that were passed on to people with databases from confidential clinical processes. data sharing with external parties is a very delicate and rare issue, requiring tight regulation, particularly if the data is to be reused in a different country. another aspect mentioned by r01 was that the ability to remember important details decreases over time. although data may not be physically lost, the confidence to reuse it may decrease, reinforcing the need to make the information more explicit. in the r02 case, the background for the interview was a study to understand how the dynamics between work and family are linked to the exercise of parenting and the children´s socioemotional development. this researcher had no previous experience with metadata – “that is what we want you to teach us”. an issue for r02 was how to retrieve data in periods of greater workload, something that is overcome by personal strategies and through years of experience. when addressing the ability to remember specific details about the data, r02 compared the direct collaborators to hard discs. one of the things required from collaborators are memos, “which might be metadata”, of things related to the data. in relation to metadata, r02 expressed curiosity towards other methodologies and processes to document data, despite an overall feeling that up to that point the projects the projects had been successful in that respect. nevertheless, the need to share data was present and r02 mentioned the importance of having data accompanied by an “instruction book”. sharing data with outsiders was something r02 was anticipating on a network with partners from europe, the usa, japan, and china, and the expected issues concerned the different cultural contexts in which the databases were created. a specific issue was how to integrate cultural knowledge in the interpretation of data. when asked about the reuse potential of data outside of the original project and domain scope, r02 considered that the question posed an interesting challenge to think about, and was not sure about how data might be reused beyond the psychology and social sciences fields. https://doi.org/%2010.29173/iq1008 12/20 castro, joão aguiar; rodrigues, joana; matos, paula mena; sales, célia; ribeiro, cristina (2023) getting in touch with metadata: a ddi subset for fair metadata production in clinical psychology, iassist quarterly 47(1), pp. 1-19. doi: https://doi.org/ 10.29173/iq1008 both researchers agreed that data documentation is pivotal for data reuse. for r02, data documentation is directly related to the quality of the data itself. a scenario described by r02 was the need to document occurrences that are out of the ordinary— a teenager responding in jest, for instance. this type of behavior from participant behavior is easily spotted by the researcher in charge of data collection and this information is useful to process the data accordingly. similarly, r01 mentioned frequently asking collaborators to annotate data if they felt that the person did not understand the questionnaire. moreover, r01 considers data documentation essential to reuse data in a different context, for instance when clinical data collected by a therapist is aggregated at the service level to evaluate service quality. data description by researchers using dendro the two data description sessions lasted approximately thirty minutes each, excluding the time dedicated to introducing dendro or addressing any questions. session 1 was carried out in january 2018 with two participants, r03 and r04. a dataset with results from descriptive statistics on children's emotion regulation, parents' work-family conflict, and psychological availability was described during this session and published in b2share13. although r02 was the principal investigator of the project that generated the data they opted to delegate the description to the two collaborators that worked with the data regularly, r03 and r04. in this session, r03 and r04 filled in 13 metadata elements in 30 minutes. both researchers had no data description experience but became familiar with the proposed task quite easily. they talked to each other during the session to discuss the meaning of some descriptors. the researchers were very prudent in the selection of descriptors and in the information provided given the perspective of subsequent data publication. they selected metadata elements for the temporal context and methodological information such as sampling procedure, time method, and sample size. in this case, they considered the ddi subset convenient for their data, especially because the concepts are close to the terminology they regularly adopt. session 2 took place in with r01 in november 2018. the metadata recorded in this session pertained to the validation of three tools for assessing the psychosocial impact of genetic testing for cancer risk. in this session, r01 produced a metadata record with 17 descriptors. the metadata values consisted mostly of short texts. this researcher used the description element to contextualize the project and used it again to identify the type of data that made up the uploaded file. despite the incomplete information on the characteristics of the dataset, r01 showed awareness of the need to use the same metadata element at different levels of description. the deviation from from sample design was also filled, showing consistency with what had been said during the interview regarding the need to document the contingencies in the research process. the quality of the metadata may have been hindered by two constraints. first, the data description session https://doi.org/%2010.29173/iq1008 13/20 castro, joão aguiar; rodrigues, joana; matos, paula mena; sales, célia; ribeiro, cristina (2023) getting in touch with metadata: a ddi subset for fair metadata production in clinical psychology, iassist quarterly 47(1), pp. 1-19. doi: https://doi.org/ 10.29173/iq1008 took place amidst a busy schedule; second, as the follow-up questionnaire would reveal, the benefits and the objectives of metadata creation were still not totally clear. table 2 shows the list of descriptors from dublin core and ddi filled in during each session. additionally, it shows the number of times that these descriptors were used by other social scientists in all contacts we established during the tail project, beyond the together project. from the beginning of 2018, we carried out 8 data description sessions with social scientists in dendro, including the two detailed here. in addition to the sessions reported in this work, the other social sciences domains represented in the right column cover studies related to consumption sociology, questionnaires to evaluate fitness trackers, the nutritional status of people with dementia, work psychology, organizational sociology, and fragility assessment. we recognize the interdisciplinary nature of most research projects, but we take into account the typology of the data described in each case and the fact that the ddi is the most suitable vocabulary to represent them. metadata elements related to methodological aspects such as sampling procedure, kind of data, and sample size, are generally the first to be filled in because they are common in the research process and researchers can easily understand these concepts. we also found that, due to their scope, these descriptors are chosen by researchers from domains other than the social sciences who are looking for a less general description of their data. a researcher working with nanoparticle synthesis decided to make a high-level description and after browsing dendro, regarded the concepts represented in the ddi subset as the most appropriate for an immediate registration of metadata. on the other hand, administrative and descriptive metadata elements are not consistently used, which is in line with the results obtained by white (2014), as researchers tend to focus more on scientific details. in dendro, we observed a tendency of researchers to choose elements from their domain ontology and to show little interest in exploring complementary vocabularies that would enrich the quality of the metadata. however, this implies that researchers still need to deepen their general understanding of the benefits of metadata. table 2: descriptors filled in during the data description sessions descriptor use session 1 (r03 + r04) session 2 (r01) total in 8 sessions in social sciences abstract 6 sampling procedure 6 kind of data 6 temporal coverage 5 sample size 5 data collection methodology 5 https://doi.org/%2010.29173/iq1008 14/20 castro, joão aguiar; rodrigues, joana; matos, paula mena; sales, célia; ribeiro, cristina (2023) getting in touch with metadata: a ddi subset for fair metadata production in clinical psychology, iassist quarterly 47(1), pp. 1-19. doi: https://doi.org/ 10.29173/iq1008 creator 4 spatial coverage 4 language 4 methodology 4 universe 4 date created 4 subject 4 audience 3 format 3 analysis unit 2 description 2 relation 2 summary statistics 2 access rights 2 variable 2 time method 1 instrument 1 deviation from sample design 1 coverage 1 overall, the researchers in the two sessions created good quality metadata records, considering that they were able to produce comprehensive and detailed metadata records in their first experience in this kind of activity. the metadata provided by r03 and r04 can easily be enriched with temporal coverage information and the accuracy of the format information can be refined. the record created by r01 lacks subject and temporal metadata. in both cases, it would not take much to improve the metadata records in order to promote search and access to the data. the metadata is rich in terms of information regarding the methodologies and context of data production. only the abstract, sampling procedure, creator, and language co-occurred in the two sessions. this suggests that it is necessary to maintain a subset with an adequate number of descriptors, i.e., large enough to fit the expectations of researchers, and also to have flexible tools that enable researchers to combine suitable descriptors according to the metadata requirements of a given dataset. it also suggests that researchers continue to be involved in training related to the production of metadata to increase their awareness of the benefits that can result from detailed and accurate metadata. https://doi.org/%2010.29173/iq1008 15/20 castro, joão aguiar; rodrigues, joana; matos, paula mena; sales, célia; ribeiro, cristina (2023) getting in touch with metadata: a ddi subset for fair metadata production in clinical psychology, iassist quarterly 47(1), pp. 1-19. doi: https://doi.org/ 10.29173/iq1008 researchers’ perspectives the additional feedback obtained through the online questionnaire showed that the researchers think that the metadata they produced during the session was sufficient. according to r01, there was no need for additional information. both r03 and r04 stated that more information might have been useful but did not provide details. after participating in the session, and despite having created a detailed metadata record, r01 did not identify any particular usefulness in the data description – “i have not yet fully understood the application of the knowledge that resulted from the description of the data. i still considered it as important but it is something abstract”. this researcher thinks that data description is an important subject, but an overly abstract activity. data description was perceived as a very boring activity, slightly difficult and time-consuming. on the other hand, the interest in rdm is very high since it helps to organize, store, and reuse data. as for r03, data description is a somewhat easy and practical task, yet slightly time-consuming. the activity was considered useful to facilitate the dissemination of data to other researchers and to the academic community. moreover, according to this researcher, the reuse of existing databases is also a benefit since it prevents overloading participants with new questionnaires. the degree of interest of r03 regarding data management is moderate, which can be explained by the general feeling that the research group was being successful in data documentation, mentioned by r02 during the interview. the participants agree that they will have more interest if rdm practices bring more visibility to their work via data citation and enable midor long-term data reuse. on top of that, r01 highlights the availability of rdm tools as a top factor for increased motivation. the availability of data description tools was a preference for r03, corroborated by r04. additionally, r01 stated that the developed activities can be improved by practical examples of data description in the clinical psychology domain. after carefully thinking about the collaboration described in this case study, r01 highlighted that if data management activities are not properly integrated into the research workflow, they may be perceived as time-consuming and as an overload, which may prevent researchers from engaging in them. moreover, there is also the need to work on the communication between data curators and researchers. in the words of r01: “on one hand we now have an established communication channel and a work relationship, we are aware of our data management needs, and we wish to be on board in the development of tools that integrate data management in routine research work. however, on the other hand, we found ourselves lost in the transition between the data manager´s and the researchers’ world. for us as researchers, data management terminology is still perceived as abstract and technical. on the other hand, data management activities are perceived as essential but logistic, and the aim is to spend as little time as possible planning and implementing them. essentially, we are aware that there is much to do in order to address rdm routinely in research projects”. https://doi.org/%2010.29173/iq1008 16/20 castro, joão aguiar; rodrigues, joana; matos, paula mena; sales, célia; ribeiro, cristina (2023) getting in touch with metadata: a ddi subset for fair metadata production in clinical psychology, iassist quarterly 47(1), pp. 1-19. doi: https://doi.org/ 10.29173/iq1008 conclusion the work carried out during the tail project focused on the integration of data management tools early in the research workflow. to do so, we established a set of contacts with researchers at the university of porto to assess domain requirements and test such tools. in promising cases, such as the one established with the together project, the collaboration can take place throughout the project as an independent and complementary task. among other aspects, it can make it possible to develop in-depth knowledge that can be applied in recommending best practices for ongoing and future projects with similar characteristics. we envision the collaboration with researchers as an opportunity to establish data management practices tailored for small projects and supported by recommendations from expert communities. documenting successful stories due to the adoption of good rdm practices can have a persuasive effect on others since researchers often ask for practical examples from their domain to understand what is expected from them to comply with the current rdm mandates. this paper described activities to simplify metadata production by researchers, in a context where adequate resources to get them involved in rdm are missing. we implemented a ddi subset in dendro, a staging platform for data description, that was used by researchers in their first contact with metadata creation. good examples can provide valuable insight for the adoption of the ddi subset and lay the foundations for data documentation in projects in the clinical psychology domain. the results showed that participants were able to choose and fill in several descriptors in a reasonable amount of time (half an hour), thus producing a comprehensive metadata record, especially when compared to the metadata that is usually available in generic data repositories. nevertheless, to meet the data deposit metadata a requirement of a disciplinary repository like icpsr further investment is required. the metadata records created were also similar to those we obtained with other social scientists using the ddi representation in the dendro platform. on average, our data description sessions took 35 minutes and resulted in 13 metadata elements. still, data description was considered a timeconsuming activity by the participants and boring by r01. if on the one hand, the communication between r03 and r04 helped them to understand the meaning of some concepts, r01 acknowledged doubts regarding the meaning behind analysis unit and time method. these observations indicate that data description may be perceived more as an additional task than as an activity that saves time and offers other benefits afterward. it should be noted that this feedback is tightly related to the specifics of dendro´s design, and users´ experience of metadata production may vary according to the according to the perceived usefulness and usability of the data description platforms. https://doi.org/%2010.29173/iq1008 17/20 castro, joão aguiar; rodrigues, joana; matos, paula mena; sales, célia; ribeiro, cristina (2023) getting in touch with metadata: a ddi subset for fair metadata production in clinical psychology, iassist quarterly 47(1), pp. 1-19. doi: https://doi.org/ 10.29173/iq1008 our approach to model a domain standard and train researcher in metadata creation was also exploited in other contexts, namely with researchers from the biomedical domains, from a large institute for health sciences and technologies in porto (sampaio et al. 2019). the feedback from these biomedical researchers suggests that the use of a restricted vocabulary favors data description but did not prevent them from identifying the limitations of the model and did not prevent them from arguing about the usefulness of descriptors even more specific to the type of experiments they perform. further developments would benefit from opening these activities to more researchers from the same domains, particularly to assess the quality of the metadata produced. however, reaching new participants is often a laborious task on its own. in addition to their busy schedules, a general belief that current practices are already good enough may prevent some researchers from participating in this type of study. another possibility is to study the attitude of researchers toward data description and the overall quality of metadata over different platforms. to improve the results, it would be useful to have a default list of general descriptors presented to the researchers as they start the description of data, as in most cases they do not know where to begin. it would also avoid browsing the dendro vocabulary list. researchers from experimental domains, for instance, have mentioned interest in having a limited number of high-level elements that can be broken-down into more specific elements according to their initial choices. implementing controlled vocabularies is something that can also improve the experience of researchers. researchers pointed out that it would be helpful to see practical examples of the use of data description and to involve others in collaborative training activities, within research units or through an institutional training plan. the engagement of more and more researchers is likely to encourage others to participate. references assante, m., candela, l., castelli, d. and tani, a. (2016). ‘are scientific data repositories coping with research data publishing?’, data science journal, 15:6, pp.1–24. doi: http://dx.doi.org/10.5334/dsj-2016-006. castro, j.a.; amorim, r., gattelli, r., karimova, r., silva j. r. and ribeiro, c. (2017) ‘involving data creators in an ontology-based design process for metadata models’, developing metadata application profiles, pp. 181-213. doi: 10. 4018/978-1-5225-2221-8.ch008 carlson, j. (2012) ‘demystifying the data interview. developing a foundation for research librarians to talk with researchers about their data’, reference services review. 40(1). doi: 10.110800907321211203603 https://doi.org/%2010.29173/iq1008 http://dx.doi.org/10.5334/dsj-2016-006 18/20 castro, joão aguiar; rodrigues, joana; matos, paula mena; sales, célia; ribeiro, cristina (2023) getting in touch with metadata: a ddi subset for fair metadata production in clinical psychology, iassist quarterly 47(1), pp. 1-19. doi: https://doi.org/ 10.29173/iq1008 clare, c., cruz, m. papadopoulou, e., savage, j., teperek, m., wang, y., witkowska, i. and yeomans, j. (2019) ‘engaging researchers with data management: the cookbook’, open reports series. 8. doi:10.11647/obp.0185 curty, r. g. (2019) ‘factors influencing research data reuse in the social sciences: an exploratory study’, international journal of digital curation. 11(1). doi:10.2218/ijdc.v11i1.401 european commission (2016a) ‘guidelines on fair data management in horizon 2020’, technical report. available at: http://ec.europa.eu/research/participants/data/ref/ h2020/grants_manual/hi/oa_pilot/h2020-hi-oa-data-mgt_en.pdf european commission (2016b) ‘regulation (eu) 2016/679 of the european parliament and of the council of 27 april 2016 on the protection of natural persons with regard to the processing of personal data and on the free movement of such data, and repealing directive 95/46/ec (general data protection regulation)’, official journal of the european union. available at: http://publications.europa.eu/en/publication-detail//publication/3e485e15-11bd-11e6-ba9a-01aa75ed71a1/language-en european commission (2018). turning fair into reality. final report and action plan from the european commission expert group on fair data. doi:10.2777/1524 mayernik, m.s. (2011) ‘metadata realities for cyberinfrastructure: data authors as metadata creators’, proquest dissertations and theses. available at: https://papers.ssrn.com/sol3/papers.cfm?abstract_id=2042653 perrier, l., blondal, e., ayala, p., dearborn, d., kenny, t., lightfoot, d., reka, r., thuna, m., trimble, l. and macdonald, h. (2017) ‘research data management in academic institutions: a scoping review’, plos one, 12(5). doi:https://doi.org/10.1371/journal.pone.0178261 qin, j., ball, a. and greenberg, j., (2012) ‘functional and architectural requirements for metadata: supporting discovery and management of scientific data’, in proceedings of the international conference on dublin core and metadata applications. pp. 62–71. qin, j. and li, k., (2013) ‘how portable are the metadata standards for scientific data? a proposal for a metadata infrastructure’, in proceedings of the international conference on dublin core and metadata applications. pp. 25–34. ribeiro, c. and fernandes, m. (2011) ‘data curation at u. porto: identifying current practices across disciplinary domains, iassist quarterly, 35(4), doi:10.29173/iq893 ribeiro, c., silva, j. r., castro, j. a., amorim, r., lopes, j.c. and david, g. (2018) ‘research data management tools and workflows: experimental work at the university of porto’, iassist quarterly, 42(2), pp.1–16. doi:https://doi.org/10.29173/iq925 rocha da silva, j. (2016) ‘usage-driven application profile generation using ontologies’. phd thesis, faculdade de engenharia da universidade do porto. https://doi.org/%2010.29173/iq1008 https://doi.org/10.1371/journal.pone.0178261 https://doi.org/10.29173/iq925 https://doi.org/10.29173/iq925 19/20 castro, joão aguiar; rodrigues, joana; matos, paula mena; sales, célia; ribeiro, cristina (2023) getting in touch with metadata: a ddi subset for fair metadata production in clinical psychology, iassist quarterly 47(1), pp. 1-19. doi: https://doi.org/ 10.29173/iq1008 rocha da silva, j., ribeiro, c. and lopes, j.c. (2018) ‘ranking dublin core descriptor lists from user interactions: a case study with dublin core terms using the dendro platform’, international journal on digital libraries. doi:https://doi.org/10.1007/s00799-018-0238-x sampaio, m., ferreira, a. l., castro, j. a., ribeiro, c. (2019) ‘training biomedical researchers is metadata with a mibbi-based ontology’, in: garoufallou, e., fallucchi, f., william de luca, e. (eds) metadata and semantic research. mtsr 2019. communications in computer and information science, vol 1057. springer, cham. https://doi.org/10.1007/978-3-030-36599-8_3 vines, t.h., albert, a.y.k., andrew, r.l., débarre, f., bock, d.g., franklin, m.t., gilbert, k.j., moore, j.s., renaut, s., and rennison, d.j. (2014) ‘the availability of research data declines rapidly with article age’, current biology, 24(1), pp.94–97. doi: 10.1016/j.cub.2013.11.014 yoon, a. (2016) ‘red flags in data: learning from failed data reuse experiences’, proceedings of the association for information science and technology, 53(1). doi: 10.1002/pra2.2016.14505301126 white, h.c. (2014) ‘descriptive metadata for scientific data repositories: a comparison of information scientist and scientist organizing behaviors’, journal of library metadata, 14(1), pp.24–51. doi: https://doi.org/10.1080/19386389.2014.891896 wiley, c., and kerby, e. (2018) ‘managing research data: graduate student and postdoctoral researcher perspectives’, issues in science and technology librarianship. doi: 10.5062/f4fn14fj willis, c., greenberg, j. & white, h. (2012) ‘analysis and synthesis of metadata goals for scientific data’, journal of the american society for information science and technology, 63(8). doi: 10.1002/asi.22683 end-notes 1 all authors are affiliated with the university of porto 2 https://www.inesctec.pt/en/projects/tail, supported by european compete grant (poci-01-0145-feder-0167 3 https://github.com/feup-infolab/dendro 4 supported by european compete grant (poci-01-0145-feder-030980) and portuguese national funds fct – fundação para a ciência e a tecnologia, i.p. (ptdc/psi-esp/30980/2017) 5 https://www.icpsr.umich.edu/web/pages/datamanagement/lifecycle/metadata.html 6 https://rd-alliance.org/groups/metadata-standards-directory-working-group.html 7 http://www.ddialliance.org/specification/ https://doi.org/%2010.29173/iq1008 https://doi.org/10.1007/s00799-018-0238-x https://doi.org/10.1007/s00799-018-0238-x https://doi.org/10.1007/978-3-030-36599-8_3 https://doi.org/10.1016/j.cub.2013.11.014 https://doi.org/10.1016/j.cub.2013.11.014 https://doi.org/10.1080/19386389.2014.891896 https://doi.org/10.1080/19386389.2014.891896 https://doi.org/10.1002/asi.22683 20/20 castro, joão aguiar; rodrigues, joana; matos, paula mena; sales, célia; ribeiro, cristina (2023) getting in touch with metadata: a ddi subset for fair metadata production in clinical psychology, iassist quarterly 47(1), pp. 1-19. doi: https://doi.org/ 10.29173/iq1008 8 https://dwc.tdwg.org/ 9 https://www.eurocris.org/cerif/main-features-cerif 10 https://www.w3.org/tr/prov-overview/ 11 http://www.dublincore.org/specifications/dublin-core/dcmi-terms/ 12 http://www.ddialliance.org/specification/rdf/discovery 13 doi: 10.23728/b2share.7b3c66dfa4df4a7f9ba04fbc30cfb8bc https://doi.org/%2010.29173/iq1008 14 iassist quarterly 2015 iassist quarterly iassist quarterlyiassist quarterly abstract academic librarians and data specialists use a variety of approaches to gain insight into how researcher data needs and practices vary by discipline, including surveys, focus groups, and interviews. some published studies included small numbers of business school faculty and graduate students in their samples, but provided little, if any, insight into variations within the business discipline. business researchers employ a variety of research designs and data collection methods and engage in quantitative and qualitative data analysis. the purpose of this paper is to provide deeper insight into primary and secondary data use by business graduate students at one canadian university based on a content analysis of a corpus of 32 master of science in management theses. this paper explores variations in research designs and data collection methods between and within business subfields (e.g., accounting, finance, operations and information systems, marketing, or organization studies) in order to better understand the extent to which these researchers collect and analyze primary data or secondary data sources, including commercial or open data sources. the results of this analysis will inform the work of data specialists and liaison librarians who provide research data management services for business school researchers.. keywords: business, primary data, secondary data, graduate students, research data management introduction a bridge is an apt metaphor for the work of an academic liaison librarian, who acts as a boundary spanner between faculty, students, and the library. much of this boundary spanning activity is driven by traditional liaison responsibilities including reference service, information literacy instruction, and collection development. as canadian academic libraries begin to develop new research data management (rdm) services, liaison librarians have been identified as ‘crucial intermediaries between the library’s services and its researcher community… [who] often have domain-specific expertise and a network of department-specific relationships’ (steeleworthy, 2014, p.7). like many of its canadian peers, brock university library has articulated a desire, through its most recent strategic planning exercise, to explore opportunities to support research data management and curation (brock university library, 2012). as the liaison librarian to the goodman school of business at brock university, i was quite familiar with the challenges of working with complex, and often expensive, commercial sources of numeric business data such as compustat and crsp (hong & lowry, 2007), but less familiar with the data practices of business scholars who generated primary data as part of the research life cycle. in order to bridge the business data divide, i needed to acquire evidence-based bridging the business data divide: insights into primary and secondary data use by business researchers by linda d. lowry1 this study employs content analysis to investigate the research designs and data collection methods iassist quarterly 2015 14 iassist quarterly 2015 15 iassist quarterly insight into business researchers in their dual roles as data producers and data consumers. a key phase in the development of rdm services is the discovery phase, which documents and analyses current researcher data practices that may be shaped by a variety of factors such as discipline, funding source requirements, research team composition, and career stage (whyte, 2014). an independent assessment of the management, business, and finance (mbf) research landscape in canada, commissioned by the social sciences and humanities research council (sshrc), provides some insight into these factors (council of canadian academies [cca], 2009a). the number of business faculty in canada was estimated at just over 2,900 individuals working at 58 different academic institutions (cca, 2009a, p. 14). business is diverse discipline comprised of many subfields, some of which are more research-intensive than others. a bibliometric analysis of canadian mbf research output published between 1997 and 2006 found that the accounting subfield represented 14% of business school faculty but produced only 2% of the research output, while the organizational studies and human resources subfield represented 5% of total business school faculty but produced 11% of the research output (cca, 2009a, p.22). an analysis of research grants administered by sshrc between 2005 and 2008 calculated that just 1.7% of these grants went to mbf research (cca, 2009a, p. 18). the council of canadian academies also examined the level of collaborative activity among mbf researchers and found that: (a) 40% of all papers published between 1996 and 2007 were collaborative; and (b), among the top 25 canadian universities, 45% of collaborative papers had an international co-author. data management plans are not currently required for sshrc-funded research, but researchers who collaborate internationally may find themselves subject to data management and sharing policies required by funding agencies in other countries (corti et al., 2014). this study employs content analysis to investigate the research designs and data collection methods found in one form of academic business research output, the master’s thesis, in order to discover to what extent graduate student business researchers collect primary data, or rely on access to secondary data sources for their analysis, and to explore variations within and between business subfields. in order to distinguish between the terms research strategy (which was not considered in this study), research design, and research method, the following definitions were considered: 1 a research strategy refers to ‘a general orientation to the conduct of social research’ (bryman et al., 2011, p.579). commonly cited strategies are qualitative, quantitative, and mixed methods, while other terms used to describe research strategies include strategies of inquiry, traditions of inquiry, or methodologies (creswell, 2003, p. 13). 2 a research design refers to ‘a framework for the collection and analysis of data’ (bryman et al., 2011, p. 579). examples of research designs described in standard accounting, business, and social science research textbooks include experimental, cross-sectional (survey), fieldwork, case study, and archival (secondary analysis) designs (bryman et al., 2011; neuman, 2003; and smith, 2011). 3 a research method can be defined as ‘simply a technique for collecting data’ (bryman et al., 2011, p. 77) such as selfcompletion questionnaires, structured interviewing, focus groups, structured observation, ethnography and participant observation, content analysis, and secondary analysis. for consistency’s sake, the term ‘data collection method’ will be used in this study when discussing research methods. reliance on primary data collection has implications for the development of research data management services, while reliance on secondary data has implications for data reference support and collection development planning, particularly due to the high cost, proprietary nature, and complex interfaces of many business data sets. this study sets a baseline measurement for data practices at the master’s level of business research, and can be used in future studies to compare current data practices at other career stages, or at other institutions, at the disciplinary or sub disciplinary level of analysis. this paper is structured as follows: section 2 reviews the literature related to methods of discovering researcher data practices; section 3 describes the purpose of the study and the research questions i will be exploring; section 4 describes the study’s procedures including the setting, and methods of data collection; section 5 presents the findings of the content analysis; section 6 discusses the implications of the findings for research data management, reference support, and collection development; section 7 discusses the limitations of the study; and the final section presents suggestions for future research literature review surveys and interviews academic librarians and data specialists have used a variety of approaches to gain insight into current research data management practices such as case studies (e.g., key perspectives, 2010), campus-wide questionnaires (e.g., parham, bodnar & fuchs, 2012), interviews (e.g., carlson, 2012), and focus groups (e.g., mcclure et al., 2014). one study revealed statistically significant differences in research data management practices and attitudes across four research domains (but did not consider discipline-specific distinctions), leading the authors to recommend tailoring data management services using discipline-specific approaches (akers & doty, 2013). another study attempted to examine differences in research data practices by discipline and by methodology but the findings were of limited generalizability due to the low response rate and a survey instrument which confounded research strategies (e.g., qualitative, quantitative, and mixed methods) with research designs (e.g., experimental, survey, field work), and data collection methods (e.g., oral history, textual analysis) (weller & monroe-gulick, 2014). motivated by research funding agency requirements for data management plans, most studies have focused on the data curation behaviors and attitudes (such as data preservation and data sharing) of science researchers (e.g., scaramozzino, ramirez, & mcgaughey, 2012), or if institution-wide, have grouped business scholars with social science or professional schools (e.g., akers & doty, 2013), thus providing little, if any, insight into variations within the business discipline. in the next section, i discuss how deeper insight into a researcher’s choice of data collection methods and patterns of secondary data use within a discipline can be acquired by conducting a content analysis of scholarly research publications such as journal articles, theses, and dissertations. content analysis content analysis, which is a nonreactive or unobtrusive data collection method, enables researchers to overcome some of the weaknesses of survey research, such as low response rates, sampling errors, or unclear question wording (neuman, 2003). several studies provided insight into business research designs 16 iassist quarterly 2015 iassist quarterly at the disciplinary level of analysis. researchers investigating the prevalence of mixed methods research designs in business and management dissertations conducted a systematic content analysis of 186 doctor of business administration theses and confirmed the use of a diverse set research designs and data collection methods (miller & cameron, 2011). in a similar study, mclennan, moyle, and weiler (2013) explored the role of economics in tourism postgraduate research by conducting a content analysis of 118 doctoral dissertations completed in the united states, canada, australia, and new zealand between 2000 and 2010. their examination of the frequency of use of specific research approaches methodologies found that 60% of tourism economics theses used quantitative approaches, 21% used qualitative approaches, and 9% used mixed methods approaches, while their analysis of the data collection methods employed identified a diverse range of techniques including interviews, surveys, case studies, econometric forecasting, observation, and econometric modeling (mclennan, moyle, & weiler, 2013, p. 186). other content analysis studies explored research trends and practices within specific business subfields (e.g., accounting, logistics and supply chain management), thus providing insight into primary and secondary data use at the sub disciplinary level of analysis. an examination of trends in accounting research over 50 year period found that archival research, defined as ‘papers using data from historical market information [such as] stock prices’ (oler, oler, & skousen, 2010, p. 668), has been the dominant research methodology in published accounting papers since the 1980s and comprised more than 60% of all papers published between 2000 and 2007. a review of articles published over a two year span in the journal of business logistics found that 62% of empirically-based studies used primary data from surveys or case studies, while 21% of studies used secondary data methods (rabinovich & cheon, 2011). while logistics and supply chain researchers appear to rely less heavily on secondary sources than do accounting researchers, a broader review of recent research in the logistics and supply chain field identified extensive use of secondary data sources for archival data collection, simulation, content analysis, event studies, and meta-analysis, leading rabinovich and cheon to advocate for the extension of traditional secondary data methods to include logistics research. several studies conducted by librarians also illustrate the value of the content analysis method in uncovering discipline-specific data practices. nicholson and bennett (2009) explored the nature of primary and secondary data use and availability within business ethics research through a content analysis of 48 doctoral dissertations. their analysis revealed that 51% of the dissertations contained only primary data, 12% relied exclusively on secondary data, and 32% collected both primary and secondary data. a review of primary data collection methods identified four main categories: observations, surveys, experiments, and structured interviews, while a review of the secondary data collected identified a range of data types including numeric datasets, corporate annual reports, government filings and regulatory cases (nicholson & bennett, 2009). more recently, williams (2013) analysed the content of 124 journal articles published by 64 faculty members in crop sciences for evidence of data usage and data sharing, in order to identify faculty candidates for data services. an advantage of the bibliographic study (sic) approach was that it revealed a diversity of discipline-specific data practices, but it was time consuming to conduct, because if data sets were used, they were typically not cited in the bibliography, but within the text of the article (williams, 2013, p. 207). in summary, content analysis is an unobtrusive discovery method which can provide insight into the prevalence of various research designs and data collection methods in order to determine patterns of primary and secondary data use within specific disciplines, but few studies have examined variations between or within business subfields. this study attempts to fill that gap by reporting on the findings of an exploratory content analysis study of business master’s theses. purpose the purpose of this study was to investigate the research designs and data collection methods of students in a research-based master of science in management (mscm) program in order to better understand the extent to which these researchers collected and analyzed primary or secondary data. a content analysis of a corpus of 32 master’s theses explored differences between and within subfields of business with respect to research designs and data collection methods. in cases where a thesis used secondary data, attempts were made to identify whether the data sources could be considered open data, or commercial data. this study explored the following research questions: 1 what is the distribution of theses by area of specialization and how does it compare to the distribution of core (supervisory) faculty? 2 what is the overall distribution of theses by research design and by data collection method? what are the patterns of data collection method use within each type of research design? 3 what is the distribution of research designs and data collection methods by area of specialization? 4 what is the overall nature of primary and secondary data collection and use (across all specializations)? 5 what types of secondary sources are used in business research? do these researchers use open data sources, proprietary/ commercial data sources, or both? procedures setting brock university is a large comprehensive university located in canada which offers a wide variety of undergraduate and graduate programs across seven faculties. brock university’s goodman school of business (gsb) is accredited by aacsb international and has undergraduate and graduate degree programs in accounting and business administration, an enrollment of 2500 fte students, and a faculty complement of 95 (brock university institutional analysis & planning, 2014). the gsb launched a research-based master of science in management program during the 2007/2008 academic year with two goals in mind: first, to prepare students to conduct research in industry and government settings, and second, to prepare students for doctoral level studies in business (brock university, 2007). the mscm is a two year program which culminates in a thesis based on independent and original research, and currently offers specializations in accounting, finance, operations and information systems management (o & ism), (formerly known as management science), marketing, and organization studies. the organization studies stream was first offered during the 2010/2011 academic year (brock university, 2010). although brock university does not currently offer a doctoral level degree in business, at one point in time the gsb’s medium to long term plan included the development of a iassist quarterly 2015 17 iassist quarterly research-based doctoral degree in business, perhaps jointly with another university (brock university faculty of business, 2005). method in order to better understand the extent to which business student researchers collect and analyze primary or secondary data, i conducted a systematic content analysis of a corpus of 32 master of science in management theses which were deposited in brock university library’s digital repository2. master’s theses must be published in the digital repository as a graduation requirement, so this sample represented 100% of the mscm degrees awarded since the inception of the program. each thesis was hand coded using a hybrid approach of manifest and latent coding, similar to the approach taken by nicholson and bennett (2009), in order to identify the business subfield, research design, and data collection method employed, and the extent and nature of secondary data use (see appendix a) . the full text of each thesis was reviewed, with particular attention paid to the title page, acknowledgements, abstract, table of contents, methods, and data sections. brock university’s mscm program offers five subfields (referred to as areas of specialization): which are (a) accounting, (b) finance, (c) o & ism, (d) marketing, and (e) organization studies. if the area of specialization was not specifically stated on the title page, a code was assigned based on the topic of the theses, and the home subject area of the student’s thesis advisor (who was often cited in the acknowledgements). each thesis was coded according to the choice of research design and data collection method and the coding form allowed for the possible use of more than one research design and data collection method, as might be the case in a mixed method research strategy. a core list of research designs and data collection methods was compiled after a review of accounting, business, and social science research methods textbooks (see appendices b and c). finally, each thesis was analyzed for evidence of secondary data use. each secondary data source was identified by name and by type (i.e., open or commercial / proprietary). further investigation was required in some cases to determine if a secondary source was a commercial or an open source. findings distribution of theses by area of specialization table 1 presents a comparison of the distribution of theses and core faculty by area of specialization. the largest proportion of theses came from the finance area, followed by the marketing area. the proportions for these two areas were larger than one might expect, based on the distribution of core faculty by area of specialization as currently listed on the program’s website (brock university goodman school of business, 2015). three of the five subject area specializations were under-represented (when compared to the distribution of core faculty) including: accounting, o & ism, and organization studies. the differences in proportions might be a result of several factors such as the relative newness of the mscm program, the growth of the program over time, and variations in student interest in each of the specialized streams. according to the appraisal brief for the mscm program (brock university faculty of business, 2005), at the time the degree program was proposed there were 21 core faculty distributed across four areas of specialization: (a) eight faculty in accounting, (b); five faculty in finance, (c); five faculty in management science, and (d) three faculty in marketing. given the two year length of the program, and the fact that the organization studies specialization was not added until the 2010-2011 academic year, it is not as surprising to have just two theses completed in organization studies. the 2014-2015 graduate calendar notes that the specialized streams may not be offered every year if there is insufficient student interest (brock university, 2014). distribution of theses by research design and data collection method three types of research designs were employed in mscm theses: archival / secondary analysis, survey, and experimental (see figure 1). there were no examples of case study designs, and none of the theses employed more than one research design. the analysis of data collection method use, as shown in figure 2, noted three different types of data gathering methods: archival-empirical/ quantitative, questionnaires, and archival-content analysis. patterns of data collection method use within each type of research design appear in table 2. both examples of theses with experimental designs used questionnaires for data collection, as did all seven of the theses with survey research designs. of the 23 theses which employed the archival / secondary analysis research design, only one engaged in a qualitative content analysis, while the other 22 engaged in the empirical analysis of quantitative data. none of the theses employed more than one type of method for the collection of data. table 1 table 1 distribu.on of theses and core faculty by area of specializa.on area of specializa.on theses (%) core faculty (%) over or underrepresented accoun.ng 5 (16%) 13 (25%) under finance 15 (47%) 8 (15%) over o & ism 3 (9%) 8 (15%) under marke.ng 7 (22%) 7 (13%) over organiza.on studies 2 (6%) 16 (30%) under total 32 (100%) 52 (100%) 1 figure 1 research design use 6% 22% 72% archival survey experimental figure 1 overall patterns of research design use in mscm theses (n=32). 18 iassist quarterly 2015 iassist quarterly distributions of research designs and data collection methods by area of specialization the distribution of research designs and data collection methods by area of specialization are presented in table 3 and table 4. the archival /secondary analysis research design was employed at least once within each area of specialization, but was most heavily used within the finance and accounting specializations. survey designs were employed in four of the five areas of specialization, while experimental designs were used in two of the marketing theses. the marketing area exhibited the widest variety of research designs, while finance used only one type of research design. an analysis of data collection method use by area of specialization revealed widespread use of the archival – empirical/quantitative method, with evidence of use within four of the five areas of specialization. in order to make sense of the patterns of research design and data collection method use by area of specialization, i also examined the course descriptions for each area of specialization in the mscm program. students in all specializations except finance take a two course research methodology sequence which covers topics such as: multivariate statistical techniques, advanced regression analysis, measurement and scaling, survey research and questionnaire design, sampling methods, qualitative research, and structural equation modeling (brock university, 2014). students in the finance specialization take a two course sequence in empirical finance which covers empirical research methods and econometric techniques in investment finance (brock university, 2014). students in the figure 2 data collection method use 3% 28% 69% archival quantitative questionnaire archival content analysis figure 2 overall patterns of data collection method use in mscm theses (n=32). table 2 pa)erns of data collec3on method use by research design research design ques/onnaire archival – content analysis archival – quan/ta/ve total: experimental 2 (100%) 0 (0%) under 2 (100%) survey 7 (100%) 0 (0%) over 7 (100%) archival 0 (0%) 1 (4.3%) under 23 (100%) total 9 (28%) 1 (3.15) 22 (68.7%) 32 (100%) 1 accounting stream also take additional courses which cover accounting theory and research methods in behavioural accounting research and market-based research, while the o & ism specialization includes courses on modeling, data mining, mathematical programming, simulation, and forecasting. looking again at table 3 and table 4, patterns of use begin to emerge, with finance, accounting, and o & ism theses favouring archival designs and quantitative analysis methods, while marketing and organization studies theses used a variety of research designs and data collection methods. insights from a case study of an msc program in finance in the united kingdom confirmed that students were exposed to secondary data and regression analysis as the model to follow in their own research (belghitar & belghitar, 2010, p.578). primary and secondary data collection this study also explored the nature of primary and secondary data collection and use across all areas of specialization, and within each specialization. table 5 presents the patterns of primary and secondary data collection across all areas of specialization. mscm theses showed a greater reliance on secondary data sources, less reliance on primary data collection, and no evidence of combining primary and secondary data collection, when compared to the nicholson and bennett (2009) analysis of business ethics dissertations. primary data collection methods were used in four of the five areas of specialization, and secondary data sources were used in all five areas of specialization (see table 6). the finance area relied exclusively on secondary data sources, as did the majority of table 3 research design use by area of specializa9on, percentage of row totals specializa)on experimental survey case archival total accoun)ng 0 (0%) 1 (20%) 0 (0%) 4 (80%) 5 (100%) finance 0 (0%) 0 (0%) 0 (0%) 15 (100%) 15 (100%) o & ism 0 (0%) 1 (33.3%) 0 (0%) 2 (66.6%) 3 (100%) marke)ng 2 (28.5%) 4 (57.1%) 0 (0%) 1 (14.2%) 7 (100%) organ. studies 0 (0%) 1 (50%) 0 (0%) 1 (50%) 2 (100%) total 2 (6.2%) 7 (21.8%) 0 (0%) 23 (71.8%) 32 (100%) 1 table 4 data collec-on method use by area of specializa-on, percentage of row totals specializa)on ques)onnaire archival – content analysis archival – quan)ta)ve total accoun)ng 1(20%) 0 (0%) 4 (80%) 5 (100%) finance 0 (0%) 0 (0%) 15 (100%) 15 (100%) o & ism 1 (33.3%) 0 (0%)) 2 (66.6%) 3 (100%) marke)ng 6 (85.7%) 0 (0%) 1 (14.2%) 7 (100%) organ. studies 1 (50%) 1 (50%) 0 (0%) 2 (100%) total 9 (28%) 1 (3.1%) 22 (68.7%) 32 (100%) 1 iassist quarterly 2015 19 iassist quarterly accounting and operations and information systems management theses. all the theses which collected primary data relied on some form of a questionnaire for data collection (see table 4). although not a focus of this study, a latent analysis revealed that a variety of methods were used to collect questionnaire data including printed questionnaires and online software packages (e.g., medialab, surveymonkey, and qualtrics). however, in some cases it was not possible to determine if the questionnaires were administered using paper or online instruments. secondary sources types used in business research this study investigated both the nature and types of secondary sources used and revealed that very few theses relied exclusively on open data sources. 28% of theses used commercial data sources exclusively, and 37% of theses used a combination of open and commercial data sources (see table 5). many theses used multiple secondary data sources, including four theses which used five or more secondary data sources (table 5). a detailed listing of these open and commercial secondary data sources, including source type, example names, and frequency of use, is presented in appendix d. the secondary sources included many of the data types identified in the nicholson and bennett (2009) study, but given this study’s canadian location, it was not surprising to find some canadian equivalents, such as sedar (for corporate financial reports and filings) and cansim (for socioeconomic data). some, but not all, of the theses that used secondary sources relied on commercial numeric datasets hosted by the library or business school (e.g., bloomberg professional, cfmrc, compustat, crsp, datastream). other theses used datasets such as the internet retailer’s top 500 guide, which may have been purchased directly the student or provided by the student’s faculty advisor. discussion implications for research data management libraries planning research data services may use a life cycle model to describe the real-world activities of their researchers (carlson, 2014). the research data lifecycle consists of the following data related activities: discovery and planning, data collection, data processing and analysis, publishing and sharing, long term management, and reusing data (corti et al., 2014 p. 17). viewed through the lens of the research data lifecycle, this study’s analysis of business master’s theses identified variations between and within business subfields with respect to research design and data collection method use which have implications for the development of research data management services for business researchers. over 70% of graduate student business researchers could be categorized as data consumers, while less than 30% could be considered data producers. archival / secondary analysis research designs were employed at least once within each business subfield, and constituted the majority of theses in the accounting, finance, and operations and information system management subfields. the discovery and acquisition of secondary data sources are crucial real-world activities for these business data consumers. a minority of theses employed survey or experimental research designs and collected research data via questionnaires. due to the small sample size, clear sub disciplinary patterns could not be found, but the researchers collecting primary data in this study were more likely to be from the marketing subfield. the real-world data activities of these business data producers, such as planning data collection protocols and obtaining informed consent, closely align with social science researchers employing survey or experimental research designs. this segment of researchers could be targeted for future research data management services. unlike our counterparts in the united states and the united kingdom, there is less urgency with respect to developing research data management plans in canada due to a lack of public funding agency data sharing mandates. the tri-agency open access policy on publication, which applies to peer-reviewed journal publications, was announced on february 27, 2015, but the sharing of publication-related research data is only required by canadian institutes of health research funding recipients (tri-agency, 2015). barriers to publishing and sharing business data exist due to the table 5 primary and secondary data collec5on across all areas of specializa5on (n = 32) descrip(on number ( %) theses with primary data collec(on 9 (28%) theses with secondary data collec(on 23 (72%) • collected only open data sources 2 (6%) • collected only commercial data sources 9 (28%) • collected both open and commercial sources 12 (37%) theses with both primary and secondary data collec(on 0 (0%) number of secondary data sources used number ( %) 1 source 3 (13%) 2 sources 6 (26%) 3 sources 3 (13%) 4 sources 4 (17%) 5 sources or more 4 (17%) total 23 (100%) 1 table 6 primary and secondary data collec5on by area of specializa5on, percentage of row totals specializa)on primary secondary total accoun)ng 1 (20%) 4 (80%) 5 (100%) finance 0 (0%) 15 (100%) 15 (100%) o & ism 1 (25%) 3 (75%) 4 (100%) marke)ng 6 (85%) 1(15%) 7 (100%) organ. studies 1 (50%) 1(50%) 2 (100%) total 9 (28%) 23 (72%) 32 (100%) 1 20 iassist quarterly 2015 iassist quarterly proprietary nature of many of the data sets used in the accounting and financial research which rely on econometric methods. while some top economic journals have mandatory data availability policies which require authors to submit the datasets used in their research, most journals do allow exemptions for research based on proprietary or confidential sources such as thomson reuters datastream (vlaeminck, 2012, 2013). implications for reference support and collection development liaison librarians and data specialists often provide data reference support to novice researchers such as undergraduate students (partlo, 2010). according to the ‘seven ages of research’ model, master’s degree students (who are in the first age of the researcher’s lifecycle) are engaged in research for a limited period of time and may not be seeking a career in academia (bent, gannon-leary, & webb, 2007, p. 85). researchers in this early career stage often turn to thesis supervisors for guidance on the research process, academic writing, and information retrieval, to varying degrees of satisfaction (bent, gannon-leary, & webb, 2007, p. 89). librarians serving business schools with research-based master’s degree programs with an accounting or finance emphasis, where students are expected to engage in archival quantitative research, may want to direct their energy toward providing higher levels of support for secondary data discovery and extraction, and then promoting this service to faculty and graduate students (e.g., kellam, 2011). liaison librarians and data specialists who develop data collections to support academic programs in their institutions will need to work closely with disciplinary faculty in accounting, finance, or other subfields that rely heavily on secondary data, to identify key data sources. thesis and dissertation content analysis can provide much needed evidence of student usage, and strengthen the argument that these databases support the curriculum (e.g., when writing the library or data services support portions of program appraisal briefs or reaccreditation reviews). a content analysis of faculty journal publications can uncover hidden data sources (e.g., owned or licensed by faculty members for use by their own research groups, but not listed on the library website or a departmental resources web page), which may be missed by traditional database benchmarking efforts that examine such listings. one must also keep in mind that traditional citation analysis studies which only look at bibliographies (for a review, see hoffman & doucette, 2012), will drastically undercount the use of secondary data sources which are often only cited within the methods sections of dissertations and journal articles. in times of library materials budget cuts, expensive subscription-based financial datasets with high cost per use may be prime targets for cancellation, but are fundamental to the research process for financial scholars. limitations there were three aspects of this study which limited the generalization of the findings to the broader population of academic business researchers. the first limitation was due to the small sample size which, while representative of 100% of the population of brock university’s mscm students, did not reflect the sub disciplinary distribution of core faculty supervisors in the mscm program. the second limitation was due to the focus of the study on only one career stage of researchers. while a graduate student’s choice of research designs and data collection methods may mirror the standard protocols within a discipline, the short duration of master’s degree programs may limit his or her choices (see belghitar & belghitar, 2010, p. 579), and therefore may not be representative of research designs employed within doctoral theses or studies published by faculty in peer-reviewed journal articles. the third limitation was due to the use of a single coder, rather than multiple coders. however, the results were strengthened by the nature of the coding scheme, which relied primarily on manifest coding, and focused on measuring the frequency of occurrence of each variable (neuman, 2003). this exploratory study served as a pilot project to test the thesis coding scheme, which could easily be adapted for broader use at other institutions by using a broader list of business subfields, such as the list of 16 subfields employed by the council of canadian academies in their bibliometric analysis (cca, 2009b, p. 11). conclusion this study used the content analysis method to provide insight into primary and secondary data use by master’s level business students at one institution. the content analysis found variations in the choice of research designs and data collection methods use across the business discipline, as well as variations within business subfields. the results of this exploratory study can serve as a benchmark for future discipline-specific content analysis studies, and the content analysis method can be used to examine variations in data practices within other social science disciplines. the study could be extended to examine business research from multiple academic institutions, multiple career stages (e.g., doctoral students, faculty, and postdoctoral researchers), or a variety of types of research output including doctoral dissertations and peer-reviewed journal articles. the use of multiple coders would allow for the examination of a larger and more representative sample of the broader business research landscape. given the canadian association of research libraries’ interest in developing a collaborative, networked approach to building a research data management infrastructure (for a discussion see canadian association of research libraries, 2013), a collaborative approach to conducting future discipline-specific content analysis studies is recommended. references akers, k.g. & doty, j. (2013). disciplinary differences in faculty research data management practices and perspectives. international journal of digital curation, 8(2), p. 5-26. belghitar, y. & belghitar, g.s. (2010). the role of critical evaluation in finance education: insights from an msc programme. accounting education: an international journal, 19(6), p. 569-586. bent, m., gannon-leary, & webb, j. (2007). information literacy in a researcher’s learning life: the seven ages of research. new review of information networking, 13(2), p. 84-99. brock university (2014). master of science in management program description, brock university graduate calendar, 2014-2015. [online]. available from: http://www.brocku.ca/webca/2014/graduate/mgmt. html brock university (2010). master of science in management program description, brock university graduate calendar, 2010-2011. [online]. available from: http://www.brocku.ca/webcal/2010/graduate/mgmt. html brock university (2007). master of science in management program description, brock university graduate calendar, 2007-2008. [online]. available from: http://brocku.ca/webcal/2007/graduate/mgmt.html brock university faculty of business (2005). appraisal brief master of science (m.sc.) in management. [online]. available from: http://www. brocku.ca/webfm_send/1289 iassist quarterly 2015 21 iassist quarterly brock university goodman school of business (2015). msc in management: core faculty/supervisors. [online]. available from http://www.brocku.ca/business/future/graduate/researchdegrees/ msc/core-faculty brock university institutional analysis & planning (2014). brock facts and supplemental dynamic reports. [online]. available from: http:// brocku.ca/institutional-analysis brock university library (2012). brock university library strategic plan. [online]. available from: https://www.brocku.ca/webfm_send/23579 bryman, a., bell, e., mills, a.j. & yue, a.r. (2011). business research methods, canadian edition. don mills, on: oxford university press. canadian association of research libraries (2013). facilitation, collaboration, and cooperation: a canadian research data management network. [online]. available from: http://www.carlabrc.ca/uploads/scc/canadian_rdmn-dec-2-2013-summary.pdf carlson, j. (2012). demystifying the data interview: developing a foundation for reference librarians to talk with researchers about their data. reference services review, 40(1), p. 7-23. carlson, j. (2014). the use of life cycle models in developing and supporting data services. in ray, j.m (ed.). research data management: practical strategies for information professionals. west lafayette: purdue university press. corti, l., van den eynden, v., bishop, l. & woollard, m. (2014). managing and sharing research data: a guide to good practice. los angeles: sage. council of canadian academies. expert panel on management, business, and finance research (2009a). better research for better business. ottawa: council of canadian academies. [online]. available from: http://www.scienceadvice.ca/en/assessments/completed/ research-business.aspx council of canadian academies. expert panel on management, business, and finance research (2009b). better research for better business. report appendices. [online]. ottawa: council of canadian academies. available from: http://www.scienceadvice.ca/en/ assessments/completed/research-business.aspx creswell, j.w. (2003) research design: qualitative, quantitative, and mixed methods, second edition. thousand oaks, ca: sage. hoffman, k. & doucette, l. (2012). a review of citation analysis methodologies for collection management. college & research libraries, 73(4), p. 321-335. hong, e. & lowry, l. (2007). business data: issues and challenges from the canadian perspective. iassist quarterly, (spring), p. 9-13. kellam, l.m. (2011). numeric data services and sources for the general reference librarian. oxford: chandos publishing. key perspectives (2010). data dimensions: disciplinary differences in research data sharing, reuse, and long term viability. scarp synthesis study. [online]. edinburgh: digital curation centre. available from: http://www.dcc.ac.uk mcclure, m., level, a.v., cranston, c.l., oehlerts, b. & culbertson, m. (2014). data curation: a study of researcher practices and needs. portal: libraries and the academy, 14(2), p. 139-164. mclennan, c.j., moyle, b.d. & weiler, b.v. (2013). the role of economics in tourism postgraduate research: an analysis of doctoral dissertations completed between 2000 2010. journal of applied economics and business research, 3(4), p. 181-191. miller, p.j. & cameron, r. (2011). mixed method research designs: a case study of their adoption in a doctor of business administration program. international journal of multiple research approaches, 5(3), p. 382-402. neuman, w.l. (2003). social research methods: qualitative and quantitative approaches, fifth edition. boston: allyn and bacon. nicholson, s.w. & bennett, t.b. (2009). transparent practices: primary and secondary data in business ethics dissertations. journal of business ethics, 84(3), p. 417-425. oler, d.k, oler, m.j. & skousen, c.j. (2010). characterizing accounting research. accounting horizons, 24(4), p. 635-670. parham, s.w., bodnar, j., & fuchs, s. (2012). supporting tomorrow’s research: assessing faculty data curation needs at georgia tech. college & research libraries news, 73(1), p. 10-13. partlo, k. (2010). the pedagogical data interview. iassist quarterly (winter/spring), p. 6-10. rabinovich, e. & cheon, s. (2011). expanding horizons and deepening understanding via the use of secondary data sources. journal of business logistics, 32(4), p. 303-316. scaramozzino, j.m., ramirez, m.l. & mcgaughey, k.j. (2012). a study of faculty data curation behaviors and attitudes at a teaching-centered university. college & research libraries, 73(4), p. 349-365. smith, m. (2011). research methods in accounting, 2nd edition. los angeles, sage. steeleworthy, m. (2014). research data management and the canadian academic library: an organizational consideration of data management and data stewardship. [online]. partnership: the canadian journal of library and information practice and research, 9(1). available from: https://journal.lib.uoguelph.ca/index.php/perj/ article/view/2990/3278 tri-agency (cirh, nserc, & sshrc). (2015). tri-agency open access policy on publications. [online]. 27th february. available from: http:// www.science.gc.ca/default.asp?lang=en&n=f6765465-1 vlaeminck, s. (2012). data policies of economics journals: research data management in economic journals. [online]. 10th december. available from http://openeconomics.net/resources/ data-policies-of-economics-journals/ vlaeminck, s. (2013). data management in scholarly journals and possible roles for libraries – some insights from edawax. liber quarterly, 23(1), p. 48-79. weller, t. & monroe-gulick, a. (2014). understanding methodological and disciplinary differences in the data practices of academic researcher. library hi tech, 32(3), p. 467-482. whyte, a. (2014). a pathway to sustainable research data services: from scoping to sustainability. in pryor, g., jones, s. & whyte, a. (eds.) delivering research data management services: fundamentals of good practice. london: facet publishing. williams, s.c. (2013). using a bibliographic study to identify faculty candidates for data services. science & technology libraries. 32 (2), p. 202-209. notes 1 linda d. lowry is the business and economics liaison librarian at brock university in st. catharines, ontario, canada. she can be reached by email: llowry@brocku.ca 2 https://dr.library.brocku.ca/ 22 iassist quarterly 2015 iassist quarterly appendix a content analysis coding form variable name codes area of specializa5on (circle one) accoun5ng (1); finance (2); opera5ons and informa5on systems management (3); marke5ng (4); organiza5on studies (5) research design (circle all that apply) experimental (1); survey (2); case study/field research (3); archival / secondary analysis (5); other (specify) (6) data collec5on method (circle all that apply) ques5onnaire (1); structured interview (2); focus group (3); ethnography/observa5on (4); archival-content analysis (5); archival-empirical/quan5ta5ve (6); other (specify) (7) type of secondary data collected open only (specify sources or names of data sets) (1); proprietary only (specify sources or names of data sets) (2); both (3) (specify sources or names of data sets); not applicable (primary data only) (4) 1 iassist quarterly 2015 23 iassist quarterly appendix b descriptions of research designs research design descrip.on examples from textbooks experimental includes classical experimental designs (random assignment, pretest, post-test, experimental group, control group) or quasiexperimental designs (neuman, 2003, p. 247). experimental (bryman et al., 2011; smith, 2011; neuman, 2003). survey “quan.ta.ve social research in which one systema.cally asks many people the same ques.ons, then records and analyzes their answers” (neuman, 2003, p. 546). cross-sec.onal or social survey (bryman et al., 2011); survey (smith, 2011; neuman, 2003). case study / field research case study “entails the detailed and intensive analysis of a single case” (bryman et al., 2011, p. 571); field research is “a type of qualita.ve research in which a researcher directly observes the people being studied in a natural sewng for an extended period” (neuman, 2003, p. 535). case study (bryman et al., 2011); fieldwork (smith, 2011); field research (neuman, 2003). archival / secondary analysis archival: research using secondary sources such as historical documents, texts, journal ar.cles, corporate annual reports, and company disclosures to conduct .me-series, cross-sec.on data analysis, content analysis or cri.cal analysis (smith, 2011); secondary analysis: “research in which one does not gather data oneself, but reexamines data previously gathered by someone else and asks new ques.ons” (neuman, 2003, p.544). archival (smith, 2011); secondary analysis (neuman, 2003). 1 24 iassist quarterly 2015 iassist quarterly appendix c descriptions of data collection methods data collec*on method descrip*on examples from textbooks ques*onnaire “a collec*on of ques*ons administered to respondents” (bryman et al., 2011, p. 579). self-comple*on ques*onnaires (bryman et al., 2011); mail and online surveys (smith 2011); mail and selfadministered ques*onnaires (neuman, 2003). structured interview “a research interview in which all respondents are asked exactly the same ques*ons in the same order with the aid of a formal interview schedule” (bryman et al., 2011, p. 581). structured interviewing (bryman et al., 2011); interviews (smith, 2011); telephone and face-to-face interviews (neuman, 2003) focus group “a form of group interview in which: there are several par*cipants; there is an emphasis in the ques*oning on a par*cular fairly *ghtly defined topic; and the emphasis is upon interac*on with the group and the joint construc*on of meaning” (bryman et al., 2011, p. 575). focus groups (bryman et al., 2011; neuman, 2003). ethnography/ observa*on a composite category which includes ethnography (immersion in a social se]ng) and structured observa*on (observing and recording behaviour) (bryman, et al., p. 574, 581). structured observa*on, ethnography & par*cipant observa*on (bryman et al., 2011); complete par*cipant, complete observer, par*cipant-observer (smith, 2011); nonreac*ve / unobtrusive observa*on; ethnography (neuman, 2003). archival – content analysis “a systema*c analysis of texts (which may be printed or visual) to determine the presence, associa*on, and meaning of images, words, phrases, concepts, and/or themes (bryman et al., p. 375) content analysis (bryman et al., 2011; smith, 2011; neuman, 2003) archival – empirical / quan*ta*ve the empirical analysis of secondary data sources (such as cross-sec*onal or *me-series data). secondary analysis (bryman et al., 2011; neuman, 2003); archival (econometric analysis of crosssec*onal or *me series data) (smith, 2011). 1 iassist quarterly 2015 25 iassist quarterly appendix d open and commercial secondary data sources used in mscm theses classifica(on data type examples number open data sources canadian public company & mutual fund filings socioeconomic data stock exchange websites canadian regulatory filings us government websites bankruptcy cases other academic datasets organiza(onal websites sedar cansim; fred ftse; me; nasdaq iiroc epa ucla-lopucki bankruptcy research database hasbrouck’s liquidity es(mates ontario winery websites 5 3 3 1 1 1 1 1 commercial / proprietary sources library subscrip(ons numeric databases compustat, datastream, cfmrc, bloomberg, crsp, execucomp, fundata, risk metrics, trace, ibes, carbon disclosure project, 8 7 5 5 4 2 2 2 1 1 1 library subscrip(ons bibliographic databases and full text publica(ons lexis-nexus tsx e-review cbca hoover’s, law source, 4 2 1 1 1 other (not libraryhosted) ecommerce data; web search analy(cs, financial trading and ownership data internet retailer top 500 guide spyfu keyword spy econoday, espeed, global hysales, 13f spectrum 3 2 1 1 1 1 1 1 microsoft word 49-2-marchant.docx 1/23 marchant, margaret & belliston, c. jeffrey (2025) data literacy in undergraduate research: a case study from student poster sessions, iassist quarterly 49(2), pp. 1–25. https://doi.org/10.29173/iq1139 the creative commons-attribution-noncommercial license 4.0 international applies to all works published by iassist quarterly. authors will retain copyright of the work and full publishing rights. data literacy in undergraduate research: a case study from student poster competitions margaret marchant1 and c. jeffrey belliston2 abstract at universities, research involving data is often regarded as the domain of graduate students and faculty. however, undergraduate students also work with data within the research process, and it can be a core experience to prepare them for future education and careers. research products from undergraduate students can demonstrate the extent of their data literacy skills and understanding, which are becoming central to success in graduate studies and the world of work. since a leading way for undergraduate students to share research is through posters, this paper examines undergraduate posters at brigham young university (byu) in the context of data literacy skills. the paper defines data literacy and the importance of undergraduate students becoming data literate. this case study shares the byu context for the undergraduate poster competitions and the resulting strengths and gaps in data literacy education followed by suggestions for supporting and encouraging undergraduate research and data literacy development beyond the traditional area of data analysis. keywords data literacy, undergraduates, posters, case study introduction in today’s collaborative scientific enterprise, data ‘have become more valuable as . . . [stand-alone] scholarly product[s] with potential for reuse’ (shorish, 2015, p. 98). as research across disciplines becomes more data-focused, it is essential for researchers, including student researchers, to be data literate. although there is not yet a standard definition for data literacy, there is much consensus among researchers about the main elements of data literacy, including areas of describing, cleaning, analyzing, evaluating, storing, and sharing data, which the current study uses to define data competency categories (carlson et al., 2011; calzada prado and marzal, 2013). the need for undergraduates to obtain data literacy competency is true regardless of whether they pursue graduate study or enter the labor force. in her article describing data information literacy as ‘a critical competency’ for undergraduates, shorish (2015) stated, ‘therefore, as one seeks to create a more informed and productive citizenry, one should seek to expose all college graduates to the skills required to effectively evaluate and use data’ (p. 102). undergraduates studying social sciences and life sciences disciplines at brigham young university (byu) in provo, utah, usa, have been creating posters sharing research results for over a decade. this 2/23 marchant, margaret & belliston, c. jeffrey (2025) data literacy in undergraduate research: a case study from student poster sessions, iassist quarterly 49(2), pp. 1–25. https://doi.org/10.29173/iq1139 trove of products was analyzed using both manual and automated content analysis methods to identify skills and gaps in data literacy competencies among undergraduate researchers. the longitudinal and multi-disciplinary nature of the poster archives allowed for the analysis of changes over time and across fields. specifically, we sought to address 1) how and what kind of data undergraduate students use in their research; 2) what data literacy skills and gaps undergraduates exhibit in their research; 3) heterogeneous effects based on the discipline, analytical methodology, type of data, or experience of the student researchers. the findings guide university faculty, librarians, and others in mentoring undergraduate students in research. students already possess many skills they can harness in future educational and career endeavors. targeted support in weak areas will prepare students for a world ever more saturated with data. literature review data literacy is the ability to find, interpret, analyze, and communicate with and about data. a recent review of the literature on data literacy education demonstrated that there is room for more research (ghodoosi, 2023). our study helps to fill this gap by exposing strengths and weaknesses in students’ data skills, particularly in their ability to communicate about data. undergraduate research research within institutions of higher education has traditionally been viewed as the domain of faculty and graduate students. however, universities are increasingly recognizing the benefit of research experiences for undergraduate students. experiential learning related to research is a high-impact practice because it is particularly effective in helping students develop skills and provides added benefits for students from traditionally underrepresented groups (american association of colleges and universities, 2024). researchers have assessed the positive impact of undergraduate research through multiple modes, including course-integrated, applied research projects (stark et al., 2018; pratoomchat and mahjabeen, 2023), data fellowships (carter, 2021), mentored research (gilmore et al., 2015; ruth et al., 2023), class poster sessions (stegemann and sutton-brady, 2009; kinikin and hench, 2012; altintas et al., 2014; logan, quiñones, and sunderland, 2015; duckworth and halliwell, 2022), or research conference poster sessions (mabrouk, 2009; burress, 2022). students benefit from these experiences by being more engaged and self-motivated in learning (mabrouk, 2009; stegemann and sutton-brady, 2009), reducing learning anxiety (stegemann and sutton-brady, 2009), improving comprehension of concepts (kinikin and hench, 2012; altintas et al., 2014), and gaining transferable skills for graduate school and the workplace (gilmore et al., 2015; carter, 2021; duckworth and halliwell, 2022; pratoomchat and mahjabeen, 2023; ruth et al., 2023). the most significant of these skills include learning to collect, analyze, visualize, communicate, and manage data. the benefits of undergraduate research relating to data literacy skills are most relevant to our study. a study of implementing a real-world research project into undergraduate economics principles courses found that 90% of students agreed that participating in the research improved their skill in collecting, processing, and interpreting data (pratoomchat and mahjabeen, 2023). similarly, ruth et al. (2023) assessed a new undergraduate research program in the social sciences. findings from 3/23 marchant, margaret & belliston, c. jeffrey (2025) data literacy in undergraduate research: a case study from student poster sessions, iassist quarterly 49(2), pp. 1–25. https://doi.org/10.29173/iq1139 surveying undergraduate students and their research mentors pointed to data management as one of the major skills gained, along with collecting and analyzing data. since one of the most common and most accessible forms of sharing undergraduate research experiences is a poster session, this study focuses on that mode. posters have the added benefit of being a product that can be shared and evaluated over time to assess students’ data literacy. research posters and data literacy to date, there have been a few studies on undergraduate research posters that connect to students’ data literacy skills. duckworth and halliwell’s (2022) content analysis of 100 posters from a multidisciplinary virtual poster session found that students are most comfortable with presenting data visually (73%) or through reporting results of data analysis (70%) and are less comfortable evaluating sources, including data sources (42%). this corroborates logan, quiñones, and sunderland whose 2015 report of a longitudinal study of a poster presentation project in a lower-level chemistry course found that data visualization does not come naturally to students, that training on poster and figure design improved the posters, and that having the opportunity to present their findings in this format contributed to students’ feelings of efficacy in their learning. burress’s (2022) evaluation of 58 undergraduate posters in multiple fields used both student selfreports of the use of data practices and proxy evidence from content analysis of the posters. the study focused on visualization, analysis, evaluation, citation, cleaning, and metadata creation. while most students (98%) reported using at least one of these data practices content analysis of the posters did not support their reports. burress suggests this could be due to a lack of understanding of academic terminology and called for consideration of additional proxy evidence from research posters, which our study provides. our research adds to previous studies by evaluating a more comprehensive sample of student work. we compare student research from different disciplines (social sciences and life sciences) over multiple years. a larger sample size allows us to produce more precise estimates of areas of strength and weakness in student data literacy. since approaches for studying text and image data from research posters are inherently more subjective, it is useful to have multiple studies in different contexts and methods to understand what findings are reproducible and generalizable. we also address findings on students’ lack of knowledge of academic terminology by using an automated text search in addition to manual content analysis. brigham young university brigham young university (byu) is a private, faith-based university sponsored by the church of jesus christ of latter-day saints. byu has the dual mission of strengthening students’ faith in the lord jesus christ and providing undergraduates the opportunity to gain a rigorous education in one or more of 198 majors and 113 minors. though having a research 1 carnegie classification (very high research spending and doctorate production), byu is primarily an ‘undergraduate teaching institution.’3 slightly more than 32,000 undergraduates comprise the vast majority of the student body of about 35,000 daytime students. 4/23 marchant, margaret & belliston, c. jeffrey (2025) data literacy in undergraduate research: a case study from student poster sessions, iassist quarterly 49(2), pp. 1–25. https://doi.org/10.29173/iq1139 undergraduate research at byu building on the work of past leadership, current byu president, c. shane reese chose strengthening the student experience as his top strategic goal. in his inaugural remarks he stated, “becoming byu will require enriching the student experience and strengthening our already student-centric approach” (tanner, 2024, p. 306). accordingly, there is a strong emphasis placed on the involvement of undergraduate students in research conducted by byu faculty. evidence of this is shown in the attention given to student mentoring in byu’s rank and status policy. all byu faculty are expected to mentor students and encouraged specifically that, “involving and mentoring students in high-quality scholarship can deepen their learning and expand future opportunities” (byu, 2022b, section 3.3). the procedural documents that guide implementation of the rank and status policy also include student mentorship (where possible) as a criterion for evaluating scholarship (byu, 2022c; byu, 2022d). moving toward concrete and measurable learning outcomes, byu administration encourages its constituent colleges to develop specific outcomes from mentored undergraduate research (e.g., becoming data literate) rather than developing university-wide outcomes. the university provides significant resources to colleges to help achieve the research learning outcomes they set forth (howell, 2024, march). over the 2013–2022 calendar years, byu’s receipt of external research funds averaged $35.7 million usd. of the $39.4 million of external research funds received in 2022, nearly 25% ($9.7 million) were directed to supporting students involved in research. that same year, the university directly supported experiential learning to the tune of nearly $4.1 million, and colleges kicked in an additional $11.2 million for total experiential learning support of $25 million. in the most recent year, experiential learning financial support from all sources increased from just over $25 million usd to $37.3 million spread over 18,207 student experiences (byu, 2022a; howell, 2024, august). in addition to structural support of mentored undergraduate research, leaders at byu frequently provide meaningful rhetorical encouragement. core elements of the university mission include commitment to excellence in research and the development of the full potential of students, which are interwoven in mentored undergraduate research (worthen, 2017). the deep institutional support of undergraduate research, both financially and ideologically, sets the stage for our study of the research works produced by students under faculty mentorship. byu poster sessions individual colleges at byu host opportunities for undergraduates to share their research. poster competitions where the posters are archived enable us to study data literacy skills after the fact. the fhss mentored research conference4 is run every year by byu’s college of family, home, and social sciences (fhss) to support undergraduate mentored research experiences and give students opportunities to practice sharing their research. the goal5 is to prepare students for future careers and/or grad school. the library/life sciences undergraduate poster competition6 is co-hosted by the byu library and the college of life sciences. the goal7 of this poster session is to help students learn and practice communicating research to general audiences (frost, goates, and nelson, 2023). 5/23 marchant, margaret & belliston, c. jeffrey (2025) data literacy in undergraduate research: a case study from student poster sessions, iassist quarterly 49(2), pp. 1–25. https://doi.org/10.29173/iq1139 methodology extending from past research (logan, quiñones, and sunderland, 2015; burress, 2022; duckworth and halliwell, 2022), this study uses student research posters as the data source to study students’ data literacy competencies, gaps in skills, and differences across disciplines and over time. the benefit of using student posters is to examine what students do in practice and how they communicate their work with and understanding of data. this study adds to previous research by utilizing a large sample of student research posters archived over more than 10 years and comprising both social sciences and life sciences. this provides a greater sample to validate trends and differences across groups. sample the data comes from two collections of student posters in brigham young university’s scholarsarchive repository. the fhss mentored research conference archives 240 student posters from social science disciplines published from 2010 to 2023. the library/life sciences undergraduate poster competition has grown each year since it started in 2017. as of the end of 2023, there were 213 archived posters. some of the posters (n = 23) in the collections provided overviews of a topic but did not report on a research study. additionally, a few posters were created by graduate students (n = 9). since these posters were out of the research question scope, they were excluded from the final analysis, making the total sample size 421 posters. the class standing of students submitting posters to the competitions ranged from freshmen to seniors. table 1 reports the proportion of students in each year of school who submitted posters. the online submission form for the life sciences poster competition included information about class standing, which is not a required metadata field in scholarsarchive. thus, the class standing data is more complete for the life science posters. where student class standing information was available, most students creating research posters were juniors and seniors. the majors of students submitting posters ranged from biology to neuroscience within the life sciences and from anthropology to sociology in the social sciences. the most common life science major was biology (n = 62). the most common social science majors were psychology (n = 52) and family life (n = 50). table 1. sample averages for demographic and data literacy characteristics of undergraduate research posters, by poster discipline all posters social sciences life sciences freshmen 0.02 (0.13) 0*** (0) 0.03 (0.18) sophomore 0.05 (0.21) 0.01*** (0.10) 0.09 (0.28) junior 0.17 (0.38) 0.03*** (0.17) 0.31 (0.46) senior 0.32 (0.47) 0.10*** (0.30) 0.53 (0.50) unspecified class standing 0.45 (0.50) 0.86*** (0.35) 0.04 (0.19) 6/23 marchant, margaret & belliston, c. jeffrey (2025) data literacy in undergraduate research: a case study from student poster sessions, iassist quarterly 49(2), pp. 1–25. https://doi.org/10.29173/iq1139 sample 421 210 211 notes: data from undergraduate research posters archived in byu scholarsarchive for the fhss mentored research conference and life sciences competition. t-tests were run to compare the outcomes of the social sciences and life sciences posters. * p < .10, ** p < .05, *** p < .01. data collection six main data literacy competency areas were identified: describing data, cleaning data, analyzing data, evaluating data, sharing data, and storing data. proxy data to measure the presence of each competency was collected using two different content analysis methods. this provides multiple perspectives to view each competency and increases the reliability of the results. first, a manual content analysis of each poster was performed. marchant defined criteria for measuring each data literacy competency category (see appendix a for the full definitions). data was collected by a research assistant reviewing each poster in relation to the defined criteria. before data for all posters was collected, the author and research assistant each coded 20 posters and compared responses. the coding definitions were adjusted as necessary to align the data. after all poster data was collected, a random selection of 30% of the posters was verified by marchant to ensure consistency and accuracy. second, data on term frequencies was collected using the automated adobe acrobat index search feature. marchant selected common terms relating to each phase of the research data process. this was accomplished by 1) close reading of 20 posters (ten each from the social sciences and life sciences poster competitions), with the intent to identify terms used when discussing data and data practices, and 2) research within the sphere of data literacy to identify other terms used to describe working with data (carlson et al., 2011; calzada prado and marzal, 2013). repeated terms were added to the search term list. appendix b lists the terms along with the related data practice category. the proportion of posters using each term8 was identified. this provides a more objective measure of which data practices student researchers participate in. results table 2 reports the descriptive statistics for variables collected manually from the posters. in both poster competitions, there was an emphasis on quantitative data and research methods and primary data collection. of the sections on the poster that relate to data, the most used was a results section (77%), followed by data descriptions (47%). data was visualized in a variety of tabular and graphical ways, with data representation in figures being more widespread. table 2. sample averages for data literacy characteristics of undergraduate research posters, by poster discipline all posters social sciences life sciences methods 0.82 (0.39) 0.80 (0.40) 0.84 (0.37) used quantitative method 0.73 (0.44) 0.66*** (0.48) 0.80 (0.40) 7/23 marchant, margaret & belliston, c. jeffrey (2025) data literacy in undergraduate research: a case study from student poster sessions, iassist quarterly 49(2), pp. 1–25. https://doi.org/10.29173/iq1139 used qualitative method 0.15 (0.35) 0.20*** (0.40) 0.09 (0.29) used mixed methods 0.12 (0.32) 0.14* (0.35) 0.09 (0.29) used primary data 0.71 (0.45) 0.55*** (0.50) 0.88 (0.33) used secondary data 0.30 (0.46) 0.46*** (0.50) 0.15 (0.36) describe data 0.47 (0.50) 0.68*** (0.47) 0.27 (0.44) analyze data 0.77 (0.42) 0.75 (0.43) 0.78 (0.41) clean data 0.20 (0.40) 0.21 (0.41) 0.19 (0.39) evaluate data 0.11 (0.32) 0.17*** (0.38) 0.06 (0.23) include form data citation 0.07 (0.25) 0.08 (0.27) 0.06 (0.23) share data 0.01 (0.10) 0.01 (0.10) 0.01 (0.10) store data 0 (0) 0 (0) 0 (0) number of data tables 0.61 (1.03) 0.91*** (1.19) 0.31 (0.73) number of data visualizations 3.03 (2.30) 2.03*** (1.97) 4.03 (2.16) sample 421 210 211 notes: data from undergraduate research posters archived in byu scholarsarchive for the fhss mentored research conference and life sciences competition. t-tests were run to compare the outcomes of the social sciences and life sciences posters. * p < .10, ** p < .05, *** p < .01. t-tests were also performed comparing posters from the social sciences and life sciences. several significant differences between the two poster competitions highlight differences in logistics and more meaningful, unique characteristics and research patterns. the life science competition incorporates a robust qualtrics survey for collecting metadata on the poster and student creator, which likely helped to provide more consistent coverage of student class standing. the differences in the type of methodology and data used in research between the social sciences and life sciences posters point to differences in research approach between the two domains. while still quantitatively focused, social science research posters showed a greater variety of research and data used, including qualitative and mixed methods and secondary data. students doing research in the social sciences were also more likely to include a data description and discuss the quality of the research data. the more open display of these competencies may be connected to the different types 8/23 marchant, margaret & belliston, c. jeffrey (2025) data literacy in undergraduate research: a case study from student poster sessions, iassist quarterly 49(2), pp. 1–25. https://doi.org/10.29173/iq1139 of data used since students may feel more need to describe and justify their choice to use an outside data source. as we explored the specific sources of secondary data used by social science and life science undergraduate researchers, we identified common sources and types of data used. there were 81 unique data sources used in the social sciences research posters and 35 unique data sources used in the life sciences posters, reflecting the greater percentage of posters using secondary data in the social sciences. as shown in figure 1, students researching in the social sciences more often used data collected by other researchers (either at byu or other institutions). social science posters also used nonprofit or ngo data more frequently. figure 1. type of secondary data used in undergraduate research posters, by poster discipline notes: data from undergraduate research posters archived in byu scholarsarchive for the fhss mentored research conference and life sciences competition. table 3 reports the proportion of articles including each term in the context of using or communicating data. overall, students most used terms related to describing (73%), analyzing (69%), and evaluating (71%) data. a minority of students (19%) used terms related to data management, either sharing or storing data. these results on describing, analyzing, sharing, and storing data align with what we found through the manual coding of posters. the large difference in the evaluating data metrics (11% for manual coding and 71% for automated indexing) indicates the complexity of measuring, teaching, and demonstrating this competency. it also suggests that students may be familiar with terms related to evaluating data (e.g., compare or limitations) but not yet be able to clearly communicate the data quality or how they evaluated it. 23% 21% 21% 19% 4% 4% 8% 4% 10% 24% 8% 8% 0% 46% 0% 10% 20% 30% 40% 50% pe rc en t o f s ec on da ry d at a so ur ce s ( % ) social sciences life sciences 9/23 marchant, margaret & belliston, c. jeffrey (2025) data literacy in undergraduate research: a case study from student poster sessions, iassist quarterly 49(2), pp. 1–25. https://doi.org/10.29173/iq1139 table 3. proportion of articles with data literacy competency terms, by poster discipline term all social sciences life sciences describing 0.73 (0.44) 0.73 (0.44) 0.73 (0.44) average 0.29 (0.45) 0.28 (0.45) 0.29 (0.45) binary 0.03 (0.18) 0.04 (0.20) 0.02 (0.15) categoric* 0.01 (0.08) 0.01* (0.12) 0 (0) continuous 0.03 (0.17) 0.03 (0.17) 0.03 (0.17) discrete 0.01 (0.10) 0.01 (0.10) 0.01 (0.10) frequency 0.13 (0.34) 0.12 (0.32) 0.14 (0.35) longitudinal 0.10 (0.30) 0.18*** (0.38) 0.03 (0.17) mean 0.34 (0.48) 0.40** (0.49) 0.29 (0.45) median 0.03 (0.18) 0.02 (0.15) 0.04 (0.20) nominal 0 (0) 0 (0) 0 (0) ordinal 0.01 (0.08) 0.01* (0.12) 0 (0) population 0.26 (0.44) 0.26 (0.44) 0.26 (0.44) proportion 0.07 (0.25) 0.06 (0.23) 0.08 (0.27) random 0.14 (0.35) 0.14 (0.35) 0.13 (0.34) standard deviation 0.04 (0.19) 0.05 (0.21) 0.02 (0.15) subjects 0.11 (0.31) 0.13* (0.34) 0.08 (0.27) cleaning 0.39 (0.49) 0.31*** (0.47) 0.46 (0.50) calculat* 0.13 (0.34) 0.09*** (0.28) 0.18 (0.39) clean 0.02 (0.15) 0.02 (0.14) 0.03 (0.17) 10/23 marchant, margaret & belliston, c. jeffrey (2025) data literacy in undergraduate research: a case study from student poster sessions, iassist quarterly 49(2), pp. 1–25. https://doi.org/10.29173/iq1139 construct 0.09 (0.29) 0.11* (0.32) 0.07 (0.25) conver* 0.08 (0.28) 0.08 (0.27) 0.09 (0.28) extract 0.06 (0.24) 0.01*** (0.12) 0.10 (0.31) merge 0.07 (0.26) 0.06 (0.24) 0.08 (0.27) normalize 0.05 (0.23) 0.02*** (0.14) 0.09 (0.29) analyzing 0.69 (0.46) 0.72 (0.45) 0.66 (0.47) analy* 0.57 (0.50) 0.60 (0.49) 0.54 (0.50) correlat* 0.27 (0.44) 0.31* (0.46) 0.23 (0.42) p-value 0.05 (0.22) 0.05 (0.22) 0.05 (0.22) regress* 0.19 (0.39) 0.25*** (0.44) 0.12 (0.32) t-test 0.04 (0.20) 0.05 (0.21) 0.03 (0.18) variation 0.09 (0.29) 0.09 (0.28) 0.10 (0.30) evaluating 0.71 (0.46) 0.69 (0.47) 0.73 (0.45) accura* 0.10 (0.30) 0.07* (0.26) 0.12 (0.33) appropriate 0.05 (0.21) 0.08*** (0.27) 0.01 (0.10) authority 0.01 (0.08) 0.01* (0.12) 0 (0) compar* 0.47 (0.50) 0.37*** (0.48) 0.57 (0.50) credib* 0.00 (0.07) 0.00 (0.07) 0.00 (0.07) limi* 0.31 (0.46) 0.34 (0.48) 0.27 (0.45) quality 0.20 (0.40) 0.25*** (0.43) 0.15 (0.35) reliab* 0.05 (0.23) 0.06 (0.24) 0.05 (0.21) sharing 0.19 0.18 0.19 11/23 marchant, margaret & belliston, c. jeffrey (2025) data literacy in undergraduate research: a case study from student poster sessions, iassist quarterly 49(2), pp. 1–25. https://doi.org/10.29173/iq1139 (0.39) (0.38) (0.40) availabl* 0.13 (0.34) 0.13 (0.34) 0.13 (0.33) replicat* 0.06 (0.23) 0.04 (0.20) 0.07 (0.26) repository 0.00 (0.05) 0 (0) 0.00 (0.07) storing 0.19 (0.40) 0.19 (0.39) 0.20 (0.40) archiv* 0.03 (0.16) 0.03 (0.18) 0.02 (0.14) confidential 0 (0) 0 (0) 0 (0) manag* 0.11 (0.31) 0.10 (0.29) 0.12 (0.32) preserv* 0.03 (0.17) 0.01** (0.12) 0.05 (0.21) secur* 0.05 (0.23) 0.07 (0.26) 0.04 (0.19) sample 421 210 211 notes: data from undergraduate research posters archived in byu scholarsarchive for the fhss mentored research conference and life sciences competition. t-tests were run to compare the outcomes between social sciences and life sciences posters. * p < .10, ** p < .05, *** p < .01. when comparing the proportions between social sciences and life sciences posters, the only statistically significant difference in a major data literacy competency category was for data cleaning. students doing life science research posters discussed how they cleaned or processed their data before data analysis more (46% compared to 31%). statistically significant differences between proportions of individual terms suggest that this comes from more discussion of calculating and normalizing variables and extracting data, which all may be more common when collecting primary data. burress’s (2022) findings pointed to differences in data literacy competencies between students who conducted quantitative research or collected their own data. we tested these relationships through two-sample t-tests, comparing posters that used quantitative methods with those that used qualitative or mixed methods (see table 4) and comparing posters that used primary data with those that used secondary data (see table 5). we found that posters that reported on quantitative research were more likely to include method, data description, and results sections on the poster compared to qualitative or mixed method studies. unsurprisingly, posters reporting on quantitative research also included more tables and figures. this suggests that students conducting quantitative research have more practice with skills of describing, analyzing, and visualizing data. on the other hand, posters that reported on research with primary data were less likely to include data descriptions or mention cleaning data. secondary data users more often cited their data than primary data users. 12/23 marchant, margaret & belliston, c. jeffrey (2025) data literacy in undergraduate research: a case study from student poster sessions, iassist quarterly 49(2), pp. 1–25. https://doi.org/10.29173/iq1139 table 4. proportion of articles with data literacy competency practice, by method quantitative qualitative or mixed methods manually coded measures methods 0.84* (0.37) 0.76 (0.43) used primary data 0.72 (0.45) 0.69 (0.46) used secondary data 0.30 (0.46) 0.31 (0.46) describe data 0.50** (0.50) 0.38 (0.49) analyze data 0.82*** (0.38) 0.62 (0.49) clean data 0.21 (0.41) 0.17 (0.37) evaluate data 0.13* (0.34) 0.07 (0.26) include form data citation 0.08 (0.27) 0.04 (0.21) share data 0.01 (0.11) 0 (0) store data 0 (0) 0 (0) number of data tables 0.68** (1.02) 0.42 (1.03) number of data visualizations 3.25*** (2.37) 2.45 (1.96) automated indexing measures describing data 0.78*** (0.42) 0.61 (0.49) cleaning data 0.40 (0.49) 0.35 (0.48) analyzing data 0.72** (0.45) 0.61 (0.49) evaluating data 0.71 (0.45) 0.68 (0.47) sharing data 0.20 (0.40) 0.15 (0.36) storing data 0.18 (0.38) 0.24 (0.43) sample 307 114 13/23 marchant, margaret & belliston, c. jeffrey (2025) data literacy in undergraduate research: a case study from student poster sessions, iassist quarterly 49(2), pp. 1–25. https://doi.org/10.29173/iq1139 notes: data from undergraduate research posters archived in byu scholarsarchive for the fhss mentored research conference and life sciences competition. t-tests were run to compare the outcomes between posters using a quantitative method to those with a qualitative or mixed method. * p < .10, ** p < .05, *** p < .01. table 5. proportion of articles with data literacy competency practice by type of data primary data secondary data manually coded measures methods 0.86*** (0.35) 0.72 (0.45) used quantitative method 0.74 (0.44) 0.71 (0.46) used qualitative method 0.14 (0.35) 0.16 (0.37) used mixed methods 0.12 (0.33) 0.11 (0.31) describe data 0.44** (0.50) 0.55 (0.50) analyze data 0.77 (0.42) 0.77 (0.42) clean data 0.16*** (0.37) 0.29 (0.46) evaluate data 0.11 (0.31) 0.13 (0.34) include form data citation 0.02*** (0.15) 0.17 (0.37) share data 0.01 (0.08) 0.02 (0.13) store data 0 (0) 0 (0) number of data tables 0.52*** (0.99) 0.85 (1.10) number of data visualizations 3.40*** (2.20) 1.93 (1.90) automated indexing measures describing data 0.74 (0.44) 0.70 (0.46) cleaning data 0.40 (0.49) 0.36 (0.48) analyzing data 0.67 (0.47) 0.73 (0.44) evaluating data 0.72 0.65 14/23 marchant, margaret & belliston, c. jeffrey (2025) data literacy in undergraduate research: a case study from student poster sessions, iassist quarterly 49(2), pp. 1–25. https://doi.org/10.29173/iq1139 (0.45) (0.48) sharing data 0.18 (0.38) 0.19 (0.40) storing data 0.18 (0.39) 0.22 (0.41) sample 290 120 notes: data from undergraduate research posters archived in byu scholarsarchive for the fhss mentored research conference and life sciences competition. posters that used both primary and secondary data (n = 11) were excluded. t-tests were run to compare the outcomes between posters using primary data and posters using secondary data. * p < .10, ** p < .05, *** p < .01. we also compared data literacy competencies by student experience, using their class standing as a proxy for experience. as students progress through their coursework and gain experience, we would expect them to grow in data literacy. our findings, reported in table 6, demonstrate that overall, upperclassmen demonstrate similar data literacy competency levels compared to less experienced students. the exceptions are in using mixed methods and describing data, which upperclassmen were more likely to do in their research compared to underclassmen. on the other hand, underclassmen were more likely to share data; all data sharing from students with class standing information came from underclassmen. these results should be interpreted with caution as the sample size for underclassmen is small, limiting the statistical power. table 6. proportion of articles with data literacy competency practice, by class standing upperclassmen underclassmen manually coded measures methods 0.84 (0.37) 0.74 (0.45) used quantitative method 0.77 (0.42) 0.81 (0.40) used qualitative method 0.11 (0.31) 0.07 (0.27) used mixed methods 0.13** (0.33) 0 (0) used primary data 0.85 (0.36) 0.78 (0.42) used secondary data 0.19 (0.39) 0.22 (0.42) describe data 0.31 (0.46) 0.26 (0.45) analyze data 0.79 (0.41) 0.81 (0.40) clean data 0.18 (0.39) 0.22 (0.42) 15/23 marchant, margaret & belliston, c. jeffrey (2025) data literacy in undergraduate research: a case study from student poster sessions, iassist quarterly 49(2), pp. 1–25. https://doi.org/10.29173/iq1139 evaluate data 0.08 (0.27) 0.04 (0.19) include form data citation 0.06 (0.24) 0.04 (0.19) share data 0*** (0) 0.07 (0.27) store data 0 (0) 0 (0) number of data tables 0.37 (0.77) 0.52 (0.98) number of data visualizations 3.73 (2.20) 3.89 (2.62) automated indexing measures describing data 0.75** (0.43) 0.56 (0.51) cleaning data 0.43 (0.50) 0.44 (0.51) analyzing data 0.67 (0.47) 0.59 (0.50) evaluating data 0.72 (0.45) 0.74 (0.45) sharing data 0.19 (0.39) 0.22 (0.42) storing data 0.20 (0.40) 0.11 (0.32) sample 205 27 notes: data from undergraduate research posters archived in byu scholarsarchive for the fhss mentored research conference and life sciences competition. posters by students with unspecified class standing (n = 206) were excluded. t-tests to compare the outcomes between upperclassmen (juniors and seniors) and underclassmen (freshmen and sophomores). * p < .10, ** p < .05, *** p < .01. appendix c provides visualizations that show the changes in data literacy competencies demonstrated in the posters over time. since the year ranges for the social science and life science posters differ and because of unique trends within each discipline, the effects over time were evaluated within each discipline rather than with all the posters together. correlating the manually coded measures with the year led to very weak correlations (less than +/0.30). the graphs in appendix c clearly show the lack of positive or negative trends over time. most of the data literacy competencies we measured fluctuate over time but are within a fairly consistent range. some large jumps exist at the beginning and end of the poster time periods, likely due to smaller sample sizes and greater variance in those years. 16/23 marchant, margaret & belliston, c. jeffrey (2025) data literacy in undergraduate research: a case study from student poster sessions, iassist quarterly 49(2), pp. 1–25. https://doi.org/10.29173/iq1139 a few patterns stand out, contributing to our understanding of students’ data literacy competencies applied to research. looking at the trends in methods used in social science research, we see a decline in quantitative research and an increase in qualitative research. this mirrors trends in the social sciences throughout the second half of the twentieth century (alasuutari 2010). it also demonstrates that students are being exposed to and using a greater variety of data and analysis methods. another trend in the social sciences is a slight increase in data sharing, as measured through data sharing terms identified through automated indexing. this is a positive sign and mirrors trends in social science research, including data management and sharing mandates from major grant funders. in the life sciences posters, there are some trends over time in the number of figures or data visualizations students use. the average number of figures used reached a high of 6.57 in 2019, steadily decreasing to about 3.5 by 2022. this supports anecdotal evidence from the poster conference organizers that student poster design has improved over time and suggests that students may be incorporating simplification, which is a best practice in research poster design (rossi, slattery, and richter 2020; siedlecki, 2017). discussion how and what kind of data do undergraduates use in research? while undergraduate students may still be learning about the research process and data skills, data plays a central role in most undergraduate research. most students use data (89% of posters include a data practice or section from the manual coding; 92% of posters include terms related to data skills from the automated indexing). the majority report their data in some way, commonly in a results section (77%) and/or through tables or figures (96%). in both the social sciences and life sciences, quantitative research and data are the most common, with some growth over time in qualitative and mixed methods research for the social sciences. for individuals supporting social science data and research, this suggests that it is important to be prepared to help students with a wider variety of data types and sources. looking at the most used secondary sources in the social sciences, librarians can support undergraduate researchers by becoming familiar with datasets collected by research groups on their campus and methods for finding data in shared repositories. additionally, our study of secondary data sources suggests that becoming familiar with government sources for data, such as the national longitudinal survey of youth (nlsy) from the bureau of labor statistics or the genetics databases made available by the national center for biotechnical information, can prepare librarians to support researchers from social sciences and life sciences respectively. similar government data sources for research outside the united states include the annual macro-economic database (ameco) from the european commission’s directorate general for economic and financial affairs and the canadian alcohol and drug use monitoring survey (cadums) from health canada. as undergraduate researchers in both the social sciences and life sciences use secondary data, librarians should also be prepared to help students evaluate the quality of data sources. zilinski, sapp nelson, and van epps (2014) provide a great framework for teaching data source evaluation skills. 17/23 marchant, margaret & belliston, c. jeffrey (2025) data literacy in undergraduate research: a case study from student poster sessions, iassist quarterly 49(2), pp. 1–25. https://doi.org/10.29173/iq1139 what data literacy skills are present or missing in undergraduate research? in both the social sciences and life sciences, students demonstrate strength in describing and analyzing data. these skills are central to gaining insights from data and reporting important findings to an audience. these are also skills commonly taught in statistics or research methods courses. students’ strength in these areas suggests that they can conduct research. they can be trusted to be involved and contribute to the process. they will especially benefit when given the opportunity to take ownership of their work, such as by creating and presenting a research poster. the research posters also display gaps in student application of data skills, particularly in the areas of sharing and storing data. one reason for this could be limited space on a research poster. generally, details on data management are not as central to explaining research findings as details on data analysis. additionally, uncluttered research posters are easiest to read (rossi, slattery, and richter, 2020; siedlecki, 2017), and posters are not meant to include everything that might be in a research article, let alone a full data management plan. however, the lack of data sharing and storing mirrors trends in published research articles, where data sharing and management lag behind other data practices despite journal data sharing policies (marchant, 2023). although undergraduate students at non-r1 institutions do not often receive data management training, they can and should learn data management best practices to prepare to succeed in future opportunities (blackwood, 2021). are there heterogeneous effects? additional strengths and gaps in data literacy skills were revealed when we compared posters by discipline, type of analysis, type of data, and experience. students doing research in the social sciences more often included discussion of the quality of the data. one reason behind this difference may be that students in the social sciences also used secondary data more often than life sciences students and may have felt the need to justify their choice of data source. on the other hand, students doing research in the life sciences more often used terms relating to data cleaning practices, suggesting a greater understanding or value of the process of preparing data for analysis. when it comes to the type of method used, students using quantitative methodology more frequently demonstrated data literacy skills in several areas. this was particularly significant (statistically and practically) in the areas of describing and analyzing data. part of this difference may be due to bias toward quantitative methods in our poster coding structure, which is a limitation of the study. future research should explore patterns in how students describe, analyze, and communicate data in qualitative research. burress (2022) found no statistically significant difference in the number of data practices displayed by students using quantitative methods compared to other types of methods. when data from our poster sample is analyzed by looking at the number of data practices in each poster (rather than the proportion of posters displaying a specific practice), we find a statistically significant difference at the 95% level. however, the differences here are not practically significant at a difference of less than half a data practice (manually coded measures: quantitative = 1.68; all others = 1.23; automated indexing measures: quantitative = 2.99; all others = 2.65). for students using primary data, we found that they were less likely to include details about describing, cleaning, and citing their data compared to students who used secondary data. this adds complexity to the findings in burress (2022), where students using primary data displayed more data competencies. our findings suggest unique skills come with using secondary data, including added 18/23 marchant, margaret & belliston, c. jeffrey (2025) data literacy in undergraduate research: a case study from student poster sessions, iassist quarterly 49(2), pp. 1–25. https://doi.org/10.29173/iq1139 needs to describe, clean, combine, and cite data sets. additionally, when we look at the number of data practices displayed in each poster, those using secondary data also included a slightly higher amount (manually coded measures: primary data = 1.48; secondary data = 1.76). the divergence in our results may be due to different proxy evidence for identifying data literacy practices, such as looking for discussion about the actions of data cleaning rather than just the mention of a software name, which students may not reference. while we would expect students to grow in data literacy competency as they progress toward graduation, we found little evidence supporting this idea. posters created by upperclassmen did not differ significantly in the data skills displayed. the only statistically significant differences were in describing data, sharing data, and using mixed methods. of the posters created by upperclassmen, 75% included terms related to data description in the automated indexing measure, compared to 56% of posters from underclassmen students. however, as the manually coded measure for describing data does not show a statistically significant difference, this could point to a limitation in the automated indexing measure. it is only able to identify the presence of terms and not the context in which they are used. another potential limitation in the automated indexing measure is that it can only cover the most used terms for a data practice, but there are many ways a student may describe data. in a similar vein, 7% of underclassmen posters included data sharing according to the manually coded measure, while no upperclassmen posters included data sharing. again, the paired measure, in this case automated indexing, was not statistically significant. upperclassmen also more frequently used mixed methods, with 13% of upperclassmen posters using mixed methods and no underclassmen using mixed methods. this suggests that upperclassmen may be prepared for more nuanced research. overall, our findings suggest that more school experience does not significantly improve data skills. this analysis is limited by the small sample size for underclassmen, which restricts the power and precision of results. it is also important to note that this analysis did not compare the same students over time, and the sample of students doing research at any experience level is biased toward more achievement-oriented students. this may account for our finding that there is little difference by experience level. conclusion through our content analysis of student research posters, we found that students display both strengths and gaps in their data skills. what do these findings mean for students, librarians, and universities? an important implication is that students have the capacity to perform research. they can contribute significantly to research projects and have important skills with data, including describing and analyzing data. giving students opportunities to take ownership of research may help them continue to increase their data competency and confidence. we advocate for more universities and faculty to provide opportunities for undergraduate research and encourage and mentor students throughout the process. it is also clear that there are gaps in student skills. across the board, students displayed little evidence of sharing or managing data. there is conflicting evidence of skills in evaluating and cleaning data. with the additional finding that greater experience in school is not related to greater data literacy skills, there is clear room to improve training in these core data skills. while statistics and methods courses appear to be doing a good job of preparing students to describe and analyze data, more 19/23 marchant, margaret & belliston, c. jeffrey (2025) data literacy in undergraduate research: a case study from student poster sessions, iassist quarterly 49(2), pp. 1–25. https://doi.org/10.29173/iq1139 emphasis can be placed on cleaning data in these courses. condon, exline, and buckley (2023) recommend partnerships between librarians, instructors, and other campus support units to help students overcome the hurdle of learning technical skills. additionally, targeted training on skills relevant to managing and communicating about research data would be beneficial. this is a great place for librarians to provide additional support through workshops, one-on-one training, or online tutorials. one example is the great data visualization and infographic checklist created by kapel and schimdt (2021). incentivizing data practices through updating poster rubrics or awards to specifically include data may also help students recognize the importance of data literacy and more clearly articulate their data practices. our sample of undergraduate research posters is limited to students from one university, so the findings are not fully representative of the broader global undergraduate population. additionally, by conducting only a content analysis, our findings relate solely to the observed practices in the text and images of the research poster. students may exhibit additional data skills that are not articulated in the research poster due to space or relevance constraints. future research can build on this understanding of current data skills and practices in undergraduate research by evaluating student data skills in other contexts. more research is also needed to assess the impact of recommended interventions such as workshops, tutorials, or updated rubrics on improving students’ data literacy. references alasuutari, p. (2010) ‘the rise and relevance of qualitative research’, international journal of social research methodology, 13(2), 139–155, https://doi.org/10.1080/13645570902966056 altintas, n. n. et al. (2014) ‘the use of poster projects as a motivational and learning tool in managerial accounting courses’, journal of education for business, 89(4), 196–201, https://doi.org/10.1080/08832323.2013.840553 american association of colleges and universities (2024) ‘high-impact practices’, https://www.aacu.org/trending-topics/high-impact blackwood, e. (2021) ‘outside the r1: equitable data management at the undergraduate level’, iassist quarterly, 45(2), 1–12, https://doi.org/10.29173/iq1011 burress, t. (2022) ‘data literacy practices of students conducting undergraduate research’, college & research libraries, 83(3), 434–451, https://doi.org/10.5860/crl.83.3.434 brigham young university (2022a) ‘annual report of sponsored research’, https://spo.byu.edu/0000018e-c41e-dbed-afaf-e6ffceaf0002/2022-annual-report-final-pdf 20/23 marchant, margaret & belliston, c. jeffrey (2025) data literacy in undergraduate research: a case study from student poster sessions, iassist quarterly 49(2), pp. 1–25. https://doi.org/10.29173/iq1139 brigham young university (2022b) ‘rank and status policy’, https://policy.byu.edu/view/rank-andstatus-policy brigham young university (2022c) ‘rank and status professional faculty review procedures’, https://policy.byu.edu/view/rank-and-status-professional-faculty-review-procedures brigham young university (2022d) ‘rank and status professorial faculty review procedures’, https://policy.byu.edu/view/rank-and-status-professorial-faculty-review-procedures calzada prado, j., & marzal, m. a. (2013) ‘incorporating data literacy into information literacy programs: core competencies and contents’, libri, 63(2), 123–134, https://doi.org/10.1515/libri-2013–0010 carlson, j., fosmire, m., miller, c. c., nelson, m. s. (2011) determining data information literacy needs: a study of students and research faculty. portal: libraries and the academy, 11(2), 629–657. https://doi.org/10.1353/pla.2011.0022 carter, j. (2021) ‘developing a future pipeline of applied social researchers through experiential learning: the case of a data fellows programme’, statistical journal of the iaos, 37(3), 935– 950, https://doi.org/10.3233/sji-210844 condon, p. b., exline, e., & buckley, l. a. (2023) ‘data literacy in the social sciences: findings from a local study on teaching with quantitative data in undergraduate courses’, evidence based library & information practice, 18(1), 61–75, https://doi.org/10.18438/eblip30138 duckworth, j., & halliwell, c. (2022) ‘evaluation of higher-order skills development in an asynchronous online poster session for final year science undergraduates’, international review of research in open and distributed learning, 23(3), 259–273, https://doi.org/10.19173/irrodl.v23i3.6238 frost, m. e., goates, m. c., & nelson, g. m. (2023) ‘the benefits of hosting a poster competition in an academic library’, college & research libraries, 84(4), 495–512, https://doi.org/10.5860/crl.84.4.495 21/23 marchant, margaret & belliston, c. jeffrey (2025) data literacy in undergraduate research: a case study from student poster sessions, iassist quarterly 49(2), pp. 1–25. https://doi.org/10.29173/iq1139 ghodoosi, b., west, t., li, q., torrisi-steele, g., & dey, s. (2023) ‘a systematic literature review of data literacy education’, journal of business & finance librarianship, 28(2), 112– 127, https://doi.org/10.1080/08963568.2023.2171552 gilmore, j., vieyra, m., timmerman, b., feldon, d., & maher, m. (2015) ‘the relationship between undergraduate research participation and subsequent research performance of early career stem graduate students’, journal of higher education, 86(6), 834–863, https://doi.org/10.1080/00221546.2015.11777386 howell, l. (2024, march 11) personal conversation with c. jeffrey belliston. howell, l. (2024, august 17) personal email communication with c. jeffrey belliston. kapel, s., & schmidt, k. (2021) ‘a student-focused checklist for creating infographics’, reference services review, 49(3), 311–328, https://doi.org/10.1108/rsr-07–2021–0042 kinikin, j., & hench, k. (2012) ‘poster presentations as an assessment tool in a third/college level information literacy course: an effective method of measuring student understanding of library research skills’, journal of information literacy, 6(2), 8696, https://doi.org/10.11645/6.2.1698 logan, j. l., quiñones, r., & sunderland, d. p. (2015) ‘poster presentations: turning a lab of the week into a culminating experience’, journal of chemical education, 92(1), 96–101, https://doi.org/10.1021/ed400695x mabrouk, p. a. (2009) ‘survey study investigating the significance of conference participation to undergraduate research students’, journal of chemical education, 86(11), 1335–1340, https://doi.org/10.1021/ed086p1335 marchant, m. (2023) ‘teaching by example: evidence of data literacy competencies and practices in top economics journal articles’, journal of escience librarianship, 12(3), https://doi.org/10.7191/jeslib.757 22/23 marchant, margaret & belliston, c. jeffrey (2025) data literacy in undergraduate research: a case study from student poster sessions, iassist quarterly 49(2), pp. 1–25. https://doi.org/10.29173/iq1139 pratoomchat, p., & mahjabeen, r. (2023) ‘building research skills through an undergraduate research project on local community’, scholarship and practice of undergraduate research, 7(1), 64–70, https://doi.org/10.18833/spur/7/1/10 rossi, t., slattery, f., & richter, k. (2020) ‘the evolution of the scientific poster: from eye-sore to eye-catcher’, medical writing, 29(1), 36–40. ruth, a., brewis, a., beresford, m., & stojanowski, c. m. (2023) ‘research supervisors and undergraduate students’ perceived gains from undergraduate research experiences in the social sciences’, international journal of inclusive education, 1-18, https://doi.org/10.1080/13603116.2023.2288642 shorish, y. (2015) ‘data information literacy and undergraduates: a critical competency’, college & undergraduate libraries, 22(1), 97–106, https://doi.org/10.1080/10691316.2015.1001246 siedlecki, s. l. (2017) ‘how to create a poster that attracts an audience’, the american journal of nursing, 117(3), 48–54, https://doi.org/10.1097/01.naj.0000513287.29624.7e stark, e., kintz, s., pestorious, c., & teriba, a. (2018) ‘assessment for learning: using programmatic assessment requirements as an opportunity to develop information literacy and data skills in undergraduate students’, assessment & evaluation in higher education, 43(7), 1061– 1068, https://doi.org/10.1080/02602938.2018.1432029 stegemann, n., & sutton-brady, c. (2009) ‘poster sessions in marketing education: an empirical examination’, journal of marketing education, 31(3), 219–229, https://doi.org/10.1177/0273475309344998 tanner, j. s. (2024) ‘envisioning byu: learning and light’, brigham young university: provo, utah. worthen, k. j. (2017, august 28) ‘byu: a unique kind of education’, https://speeches.byu.edu/talks/kevin-j-worthen/byu-unique-kind-education/ 23/23 marchant, margaret & belliston, c. jeffrey (2025) data literacy in undergraduate research: a case study from student poster sessions, iassist quarterly 49(2), pp. 1–25. https://doi.org/10.29173/iq1139 zilinski, l. d., sapp nelson, m., & van epps, a. s. (2014) ‘developing professional skills in stem students: data information literacy’, issues in science and technology librarianship, (77), https://doi.org/10.29173/istl1608 endnotes 1 maggie marchant is the economics, finance, and social science data librarian at brigham young university. she can be reached by email: maggie_marchant@byu.edu. 2 jeff belliston is the associate university librarian for administrative services at brigham young university. he can be reached by email: jeffrey_belliston@byu.edu. 3 the earliest usage of this phrase that can be found on the byu website is in an address by then byu president rex e. lee on 27 august 1990 (https://speeches.byu.edu/talks/rex-e-lee/mt-everestfound-byu-undergraduate-education-can/). in his installation of and charge to the current president on 19 september 2023, elder d. todd chistofferson, chairman of the executive committee of the byu board of trustees stated, ‘since brigham young university is first and foremost an undergraduate teaching institution, i charge you to elevate that core mission’ (https://speeches.byu.edu/talks/d-todd-christofferson/installation-and-charge/). 4 https://scholarsarchive.byu.edu/fhssconference_studentpub/ 5 https://fultonconference.byu.edu/attending-the-conference 6 https://scholarsarchive.byu.edu/library_studentposters/ 7 https://guides.lib.byu.edu/2024postercomp/judging_criteria 8 truncation used to ensure that all forms of a word were included. iasslsl quarterly 3 directions of major archives by bj0m henrichsen' let me start by giving a broad overview of the organizational structure and main services provided by the norwegian social science data services (nsd). based on this description 1 will then close by saying a few words about new services to be developed in the coming years. nsd was formally established in 1971 as an organ of the norwegian research council for science and the humanities (navf). nsd differs from most similar organizations in five ways: it is a federally structured facility with offices at all four universities in norway and in the regional colleges at smaller centers across norway. its headquarters are at the universit}' of bergen; it has built up a wide variety of data resources in all fields of the social sciences: not only data from surveys, but also a large data bank for communes and census tracts, an archive of information about organizations, and a series of files on the 'prepared for opening plenary session. 1assist conference, marina del rey, may 1986 recruitment and careers of various elite groups; it acts as the census bureau's distribution agency to the academic commimity; it has set up a special service responsible for contacts between the research community and the governmental data inspectorate; and it has established a national service for information on current research in the social sciences. in comparison with most other data facilities established in europe and in the u.s. in the last two decades, nsd is probably the one giving the highest priority to book-keeping and "process produced" data. it is deliberately multisectoral and sees its primary task as to link up and to systematize data of different types; in contrast to the typical survey archive, it is not just a repository of separately documented data sets. it is even correct to say that it is only in the last few years that nsd has been active in archiving data from various research projects. among the larger data holdings of the nsd eire: the commune data base 1769 1986 this data base contains statistics on all local administrative units in norway since 1769, and is linked up with a computer cartography facility'. this is the most widely used facility in the nsd and is constantly expanded and improved. a great deal of energy has been invested in developing effective solutions to the problems posed b> changes in boundaries and in the number of units . the base includes detailed documentation of all such changes that have taken place. coordinate matrices for all commune boundaries have been established, and boundary' segments winter 1986 4 iassist quarterly are time coded to allow the production of maps for the imits existing at any particular time period since 1769. as of 1986 the commune data base includes about 29.000 variables for each commune. census tract data base 1950 1980 to allow analyses at a lower level of aggregation, nsd has also organized a system of data for the lowest level of official enumeration: the census tracl this data base includes the censuses of 1950. 1960, 1970 and 1980. census data bank 196019701980 10% of the population are followed through three censuses. the data base includes approximately 483,000 individuals. nordic regional data base the social science research councils of denmark, finland, norway and sweden have funded this data base. data are gathered and organised in systematic time-series for all five nordic countries, including iceland. the regional units of the data are counties: amt for denmark lan for sweden and finland fylker for norway syslu for iceland time series are created for all units from 1850 to 1980 for population census data, and the period 1945 to 1980 for other groups of data. the data base system is composed of four elements: most of them organized in five-year time series: a longitudinal set of data based on population censuses 1850-1970/80 consisting of 100-150 variables organized in ten-year time series; a data set on population movements 1945-1980 consisting of an annual time series for each unit; a set of coordinate matrices for the boundaries of the units. this includes time-specific segments whereever there have been changes in the boundaries of units in the period from 1850 to 1970/80. criminal justice data nsd also has an archive of norwegian criminal justice data from 1860 to 1975. gallup data this collection is based on data from norsk gallup institutt and norsk opinionsinstitutt it contains their monthly surveys from 1964 to the present election studies nsd has taken over the surveys conducted by the norwegian election project data from the following national surveys are available from nsd: 1957. 1965, 1969, 1977, 1981. surveys from the central bureau of statistics some of the most thorough surveys in norway have been carried out by the centra! bureau of statistics from 1967 to the present the data from these surveys are at the disposal of academic users in norway via nsd. a set of data for the post-war period, consisting of 4-500 vanables for each unit. members of parliament winter 1986 iassisl quarterly 5 a data bank has been established containing information about all members of parliament and the govemmenl it covers the period from 1814 and includes information on father's occupation, education, early career, positions in legislative committees, etc. members of official committees this collection includes information on all committees appointed by the various ministries, as well as tjie members of such committees. it covers 1936, 1951, 1966, and every year from 1980 on. voluntary associations the file includes data on the 1300 largest volimtary associations in norway. data are available for the following years: 1964, 1967, 1970. 1976 and 1983. teaching packages nsd has given priority to the establishment of a set of teaching packages for both the universities and the regional colleges. in 1985, we also launched a program to establish working tools for the norwegian high schools. the program has been accepted and is financed by the norwegian ministry of education. our first products under this program are now in use in the norwegian schools. of the ne\.nsd services established in the last five years, i will mention two: secretariat for data protection affairs the norwegian personal data registers act (lov om person-registre m.m.) came into force in 1980. in response to proposals from the social sciences, the research council in 1980 established the secretariat for data protection affairs as a part of nsd. the secretariat was accepted as a broker between the research community (including medicine, the humanities. etc.) and the data inspectorate, and was mandated to provide regular reports to the data inspectorate on all projects funded through the research council for which concession was required in accordance with the provisions of the acl since then, the scctetariat has been given the same mandate for all research carried out at the universities with grants from other sources than the research council. through agreements between the data inspectorate, the research council and the universities, nsd has also been given the responsibility of archiving data, provided there is reason to assume their usefulness in future research. information service for ongoing research in 1984, the research council established an information service for norwegian research, the aim of which is to improve awareness of current research; it is provisionally established for a period of five years. the information service, in addition to general management, consists of one branch responsible for research in the humanities, and one branch responsible for research in the social sciences. the social science branch is located at nsd. the service is active in all fields in which the research council is engaged, i.e. medical science, the humanities, social science and research for social planning. the information is available in a data base for convenient access from users' own terminals. in addition there are printed catalogs for specific research areas. a small country with 4.5 million inhabitants and only four universities must coordinate national activities. for nsd it has meant that not only the research council but also the universities winter 1986 6 iassist quarterly and to a certain extent the regional colleges have chosen to concentrate their means for a social science infrastructure at nsd. today, different parts of the research council cover about 75% of our expenses, while the universities and research projects cover the rest we have today a professional staff of 15 and 4 clerical staff members. including assistants etc., we estimate that about 28 full-year equivalents will be utilized in 1986. given our special relationships to the research council, the universities, the census bureau, and the data inspectorate, we have today a monopoly in the areas in which we are active. we are striving to fulfill our responsibilities, to serve our users in the best possible way, and to make our services easily accessible to the scientific community. until now, our data services have been made operative through local offices at each of the universities, located in the university computer centers. the network among the universities has not been seen as an alternative to direct service with our own staft present at the local university. initiatives have, however, now been taken to make the network function better, and we expect that, within two years, some of our data holdings will be held in bergen only, to be requested via the tmiversity network for local users at other institutions. the services we provide today cover a broad range: from a data base with information on all social science projects and publications based on them, to our own data banks and data from all projects financed by the research council, to projects financed by other institutions, such as the universities and some of the ministries. given that researchers must deposit their data with nsd, they are informed of standards for documentation and data, which means a standardization among widely separated scholars. as the data holdings grow, there is an increased need for researchers to be kept informed of the data holdings and services. we act not only as a distributor of data, but also as a broker of social science information. we think that we have played an important role in giving researchers easy access to information and in introducing new technology. through our work we have prevented duplication of data work, we have made data available free to various users, and we have stimulated cumulative research by making data from earlier projects available to new ones. although our data have been mainly used in the social sciences, users in other fields such as history, medicine etc. are increasingly using our services. our greatest growth potential within the researh community lies in serving these new groups. we are on our way from being a social science service to being a more general service for a broader group of users. in the past, we have concentrated our efforts on providing service to the research community. during the last few years we have also started to serve local and federal agencies. these are now in the same position as the social science commimity five years ago, and they now want access to the services established for researchers. new recruits to governmental agencies often find that they, in their new position, do not have the easy access to data they had had as students. as students they were introduced to our services and now, they are still in need of access. we are discussing ways of serving both researchers and bureaucrats, and beheve that we will agree on a model covering both needs. presently we are also negotiating with the norwegian parliament to make our services available to both members of parliament and their staff. by pooling resources from these different sources, all parties will have access to a much broader range of services. our main efforts in the coming few years will be devoted to planning a shared information system for planners and researchers, and hopefully we can, within a few years, present a system serving a broader community than today.n winter 1986 1/19 weldon, kathleen j. (2020) standards and scoring to increase transparency for archived public opinion data, iassist quarterly 44(3), pp. 1-19. doi: https://doi.org/10.29173/iq974 standards and scoring to increase transparency for archived public opinion data kathleen j. weldon1 abstract faced with increased diversification of methodologies in the polling industry, the roper center for public opinion research center is embarking on a major initiative aimed at increasing methodological transparency across the field of public opinion survey research by increasing minimum disclosure requirements and providing users with transparency scoring for new submissions to the archive. roper center, the world’s largest archive of public opinion survey data, has long enforced disclosure requirements for archival submissions based on transparency standards developed by professional organizations in the polling industry, particularly the american association for public opinion research (aapor). roper center’s new requirements and scoring mechanism expand longstanding policies and procedures to better meet the challenges of today’s research environment. in this paper, roper center’s new standards will be described in the context of the historical development of transparency expectations in the polling community. the paper will also detail the implementation process, providing an account of how standards were translated into actionable ddi-based metadata to drive an automatic scoring system, how new workflows were developed with input from data providers to facilitate maximum disclosure, and how the display of the user interface was designed to ensure the transparency information can be easily viewed and understood. keywords transparency, polling, disclosure, public opinion, standards introduction the u.s. polling community has long demonstrated a commitment to transparency, as encoded in a series of standards adopted by professional organizations in the field since the 1960s. but the rapid proliferation of new methods in polling since the turn of the century have spurred the development of more stringent and complex standards by both the national council on public polls (ncpp) and the american association for public opinion research (aapor). these changes in polling methodologies present unique challenges to the roper center for public opinion research archive. the center has long depended on methodological criteria to evaluate submissions for inclusion in the archive. the policy was described internally as preserving polling that is ‘the best of its time,’ in acknowledgement that methods have changed over time and the most respected polls from the 1930s preserved in the archive used quota methods that would disqualify them for inclusion in the archive if fielded in later years. in 2002, the acquisitions https://doi.org/10.29173/iq974 2/19 weldon, kathleen j. (2020) standards and scoring to increase transparency for archived public opinion data, iassist quarterly 44(3), pp. 1-19. doi: https://doi.org/10.29173/iq974 committee of the roper center wrote a formal acquisitions policy that included the restriction that “[i]nterview data cannot be from exclusively self-selected respondents.” while this wording might be interpreted in several ways, as practice the center followed a policy of not acquiring data from online nonprobability panels. with the rise of online nonprobability panels, interactive voice response (ivr), redirected inbound call sampling (rics), and other new approaches to fielding surveys, the survey research community no longer shares a consensus over what constitutes best practices in the field. this change has complicated roper’s approach to acquisition. in 2018, the board of directors of the roper center approved recommendations for a new policy described in a memo from its acquisitions and transparency committee. the new policy opened roper center’s acquisitions to new methodologies, while creating a stringent set of disclosure requirements for this new collection and developing a system to score transparency across both the longstanding and recently developed collections at the center. this paper will trace the developments in the field of polling that led to this decision and outline the new approaches, including a description of the process of implementation. background: disclosure in polling in the 20th century jane jacobs wrote of professional self-regulation that ‘[a]ll variations have the self-interest of members at their core, usually sincerely construed as advancement of the profession itself.’ (jacobs, 2010, p. 128) in the case of pollsters, self-interest might be closer to self-preservation. unlike any other form of social science research, public opinion polls, which are frequently conducted by media organizations themselves, are released almost immediately following the completion of fieldwork, then discussed at length by media and politicians. the uniquely public role of polling has meant that from the earliest days the profession – a group that includes commercial firms, media organizations, academic research organizations, and nonprofits, with the variation in values and interests that might be expected of such a diverse group – had to invest time and effort in building the trust of politicians, journalists, and the general public. skepticism from these groups ran high, particularly in the late 1940s when polling’s massive failure to predict the winner in the 1948 truman/dewey race nearly destroyed confidence that had been built over the previous two presidential election success. the 1949 publication of lindsay rogers’s the pollsters, a work deeply critical of the role public opinion polling was coming to play in american life, increased the sense that this new industry was not to be trusted. without the support of the media, and by extension the public, the field of polling could not thrive or possibly even survive. george gallup believed full disclosure of methods, sponsorship, and data was essential. describing the commitment of the american institute of public opinion (later the gallup organization) to what would come to be known as transparency, gallup wrote: since the day it was organized the american institute of public opinion has maintained a policy of providing full information https://doi.org/10.29173/iq974 3/19 weldon, kathleen j. (2020) standards and scoring to increase transparency for archived public opinion data, iassist quarterly 44(3), pp. 1-19. doi: https://doi.org/10.29173/iq974 about all of its procedures and operations. a duplicate of every ballot ever collected in its entire history is on record in the files of princeton university for use and study by qualified students. in books, and in countless articles and speeches, we have described our methods, the size of our samples, the limitations of polls in making election forecasts, accuracy, source of revenue-which comes entirely from publications-and our overall philosophy of the place of polls in a democratic society. […] unlike some fields, the polling profession has no trade secrets. we have held that the public has every right to know just how we function. one of the best safeguards which we have imposed upon ourselves is to report in every news release the question or questions asked, the type of cross-section (whole population over 2i, voting population, informed public, etc.), along with the results. (gallup, 1948) gallup’s admirable openness was, as he notes, ‘self-imposed.’ prominent pollsters, gallup included, had discussed the potential value of setting professional standards for reporting from their first meetings together. as described by sidney hollander in a meeting place, the history of the american association for public opinion research (aapor), the idea of establishing a set of reporting standards was raised at the central city conference in 1946, the precursor to aapor’s yearly conference (hollander, 1992). in 1948, at the meeting where the aapor’s constitution was written, a set of disclosure standards was also drafted, though no action was taken to move forward with the adoption. the debate over standards continued without action for twenty years. in 1967, aapor finally made its move. gallup led the charge, concerned that the field was threatened by a proliferation of bad actors using questionable methods and, particularly in the case of the rapidly expanding field of political polling, releasing partial results intended more to influence than to reflect public opinion (gollin, 1992). a set of disclosure standards was adopted by aapor council. another form of pressure had surely influenced this decision. the specter of government regulation that had long hung over the industry had grown more threatening. in 1943, senator gerald nye had proposed a bill that would have required pollsters to disclose sample size and retain records for two years. no action was taken, but the warning bell had been rung. in 1968, as the aapor membership was first learning of the new standards council had committed to the previous year, rep. lucien nedzi of michigan sponsored a bill with real teeth. his legislation set disclosure standards to be enforceable by a fine of $1000 or 90 days in jail or both. his required items for reporting looked similar to the list first suggested in 1948, covering sponsorship and basic methodological details. in an article in public opinion quarterly, rep. nedzi directly addressed the polling community, suggesting the ‘prospect of legislation’ might be as effective as legislation itself, motivating pollsters to self-police. (nedzi, 1971) in 1979, the aapor standards served as the basis for a new set of standards adopted by the national council of public polls (ncpp). over the next few decades, major polls published results with a https://doi.org/10.29173/iq974 4/19 weldon, kathleen j. (2020) standards and scoring to increase transparency for archived public opinion data, iassist quarterly 44(3), pp. 1-19. doi: https://doi.org/10.29173/iq974 methodology statement followed by what became a familiar notation: ‘these statements conform to the principles of disclosure of the national council on public polls.’ peer-to-peer transparency although gallup had boasted of his submission of data punch cards to a repository at princeton (later moved to the roper center), as well as conference presentations and published articles, for decades the primary focus of all debates over disclosure had been public reporting, not data sharing. the audience of concern was the media, and by extension government and the people. in his chapter in a meeting place, albert gollin described the concerns about disclosure that led to the first official standards in the 1960s as ‘a struggle about control over the release of public opinion data to the public as well as about how to educate the press and public concerning the hallmarks of a professionally conducted survey.’ (gollin, 1992) the focus on the media continued into the professional literature on standards. in a 1982 public opinion quarterly (poq) article, miller and hurd noted that the aapor and ncpp polls were primarily intended to provide disclosure guidelines for survey researchers in releasing polls, but also that ‘it is obvious they were also meant to sensitize journalists.’ (miller & hurd, 1982) a number of academic articles over the 1980s and 1990s attempted to measure the success of the ncpp standards by determining what proportion of media reports on polls included the required information. implementation of disclosure in media reporting on polls was also the topic of 1971 and 1980 public opinion quarterly symposiums and a 1979 ncpp/kettering foundation conference. data sharing or requirements intended to explicate methodology at a level of detail required for researcher analysis were not part of the discussion. sharing of methodological information among polling professionals continued just as gallup described, through annual aapor meetings and other conferences, in ad hoc aapor committees, in the pages of poq and other academic journals, and at the roper center archive, which maintained a minimum disclosure requirement for acquisition that closely followed the ncpp and aapor standards. in 2006, everything changed. the national council of public polls created an expanded three-level disclosure standard. (ncpp, n.d.) level one concentrated on the traditional information required with public release of results. the second level focused on information that member organizations had to make available upon written request. the items in this level were far more comprehensive that those at the first level. the third level, which was strongly encouraged, but not required, was the release of datasets. the intended audience for these additional layers of requirements were clearly other members of the polling community. even the most poll-savvy reporters or citizens were not expected to make judgments about weighting methods or disposition codes, much less to wrangle spss files. in 2008, the aapor community found evidence that lack of transparency was preventing the field from identifying the problems that had plagued that year’s primary election polling. the willingness of polling organizations to share detailed methodological information and datasets had helped the industry overcome its failures in the 1948 election. but the report of the ad hoc committee on the 2008 presidential primary polling repeatedly noted the failure of survey organizations to provide https://doi.org/10.29173/iq974 5/19 weldon, kathleen j. (2020) standards and scoring to increase transparency for archived public opinion data, iassist quarterly 44(3), pp. 1-19. doi: https://doi.org/10.29173/iq974 timely and thorough methodological information. (traugott et al., 2009) twenty-one organizations provided at least some information, but three organizations never responded to the committee’s request for data at all. while the majority of responding organizations provided information on weighting and question wording, only seven provided the microdata, which was then deposited at the roper center. just four fulfilled the request for data on the gender and race of interviewers. as a result of these omissions, the committee called for a review of disclosure standards. in 2010 aapor announced the establishment of the transparency initiative (ti), creating a membership program which polling organizations could join by committing to abiding by the new disclosure standards. like the ncpp standards, aapor included a set of additional disclosure items to be made available upon request. in a july 2012 presentation at the rc-33 conference, ti committee chair timothy johnson and paul lavrakas identified the primary problem as ‘inadequate transparency of research methods and statistical methods’ that causes a ‘serious detriment to progress.’ (johnson & lavakras, 2012) the main goal of the initiative was to ‘advance the science and reputation of survey research’, while public education on transparency was secondary. although neither archiving nor sharing of the dataset was required in the standards, by establishing a much greater level of transparency expectation upon request, aapor expanded its focus on disclosure from journalists and the public to peer-to-peer transparency. evolving needs why did ncpp and aapor both increase their requirements so dramatically within in a few short years? the movement toward greater transparency in polling was part of a larger shift toward new expectations of data sharing, replication, and transparency in social science research, as described by herndon and o’reilly in 2016. new requirements were enforced by journals in which polling researchers often publish, like the american journal of political science, which adopted data sharing requirements in 2012, as well as funding agencies that support academic pollsters, like the national science foundation, which incorporated data sharing requirements into large grants in 2011, and nih, which did so even earlier, in 2003. (ajps, n.d.; nsf, n.d.) but polling as a discipline had another driver towards greater transparency. polling returned to the question of transparency standards in the early 2000s when several new and controversial methods, most notably internet panels, began to become mainstream. in a 2005 article, mark blumenthal built a case for increased transparency in polling by tracing recent increases in methodological heterogeneity.(blumenthal, 2005) the first of the internet opt-in panel pollsters, harris interactive, had conducted polls during the 2000 election, but in the next presidential cycle, multiple organizations jumped into the new methods sphere, with online panel pollsters zogby international and british firm yougov, and interactive voice response (ivr, or ‘robocall’) pollsters surveyusa and rasmussen drawing major media attention and enormous internet traffic. blumenthal argued that, not only these new methods, but new dissemination approaches changed the polling landscape during the first decade of the new millennium, as some new polling organizations began to publish their results directly on their own websites, rather than through major media outlets. these organizations were able to receive wide attention for their polls without undergoing the standard vetting process used by most major media organizations. in https://doi.org/10.29173/iq974 6/19 weldon, kathleen j. (2020) standards and scoring to increase transparency for archived public opinion data, iassist quarterly 44(3), pp. 1-19. doi: https://doi.org/10.29173/iq974 response to these developments, blumenthal called upon survey researchers to embrace transparency of methodology, specifically citing rapid rate of change as the primary reason for increased need for disclosure. since 2005, polls based on recently developed methods like ivr and online nonprobability panels have increasingly entered the mainstream, despite ongoing concerns about data quality and accuracy.2 the new york times and the economist both have partnered with yougov, washington post and business insider with surveymonkey, and usa today with ipsos public affairs, all utilizing online non-probability panel methods. these approaches are also expanding the types of populations polled. new organizations have been taking advantage of the lower costs of targeting historically underpolled groups using new methods by developing polling projects focused on these populations, such as latino decisions, asian american decisions, the african american research collaborative, and the american muslim poll. the new aapor disclosure standards included a number of items aimed specifically at new methodologies, including disclosure of use of routers (sites that connect potential respondents with online surveys for which they are eligible) and specific recommendations for the reporting of sampling error estimates in nonprobability polls. when aapor announced its new standards, republican pollster david hill wrote approvingly of the effort in the hill, tying the need for new standards directly to the explosion of new methods: ‘as data collection methods and sampling frames have become more exotic, including robo-calls and online panel surveys, new standards are clearly indicated.’ (hill, 2010) the relationship of new methods and disclosure was also apparent in the report of the ad hoc committee on the 2008 presidential primary polling, which in calling for a review of disclosure standards specifically referenced the new world of ‘more complicated and diverse sampling frames and selection techniques’ and ‘more complicated and diverse statistical adjustments for errors of non-observation.’ (traugott et al, 2009) roper center’s transparency project: new standards the proliferation of new methods polling presented a challenge to the roper center’s traditional approach to collection. the acquisitions policy, adopted in 2002 and most recently reviewed in 2012, specified the use of probability-based methods, while the use of ivr technologies and voter file samples were not specifically prohibited, but in practice had been avoided in collection. the policy also specified required elements of disclosure reflective of the standards of ncpp and aapor before their revisions. the field of polling research had changed, and the roper center had to respond thoughtfully. the need to accurately represent current methods had to be balanced with the center’s reputation as an archive that preserved ‘the best of its time.’ the center also had to weigh increased expectations of disclosure with a commitment to maintain overall transparency in the field by ensuring strict new requirements did not cause current donor organizations to stop sharing data. in june 2018, after several years of deliberation, the acquisitions and transparency committee of the roper center’s board of directors submitted a memo to the full board proposing a bold new transparency project with two major initiatives: a transparency scoring metric to be displayed on all new dataset catalog entries and the establishment of a new collection of surveys conducted using https://doi.org/10.29173/iq974 7/19 weldon, kathleen j. (2020) standards and scoring to increase transparency for archived public opinion data, iassist quarterly 44(3), pp. 1-19. doi: https://doi.org/10.29173/iq974 recently developed methods. the new collection would be open to all methodologies, allowing researchers to analyze these methods and potentially improve upon them. however, in recognition of concerns about possible data quality issues, several conditions would apply. all recently developed methods submissions would need to meet a high bar of transparency. only questions with dataset submissions would be included in the new methods database, in contrast to the longstanding methods collection, which includes topline results without underlying data files. finally, the new collection would be searched and displayed separately from the longstanding methods collection on the roper center website. these safeguards were particularly important for roper center users who are not advanced researchers in the field, a group that includes undergraduates and some media and nonprofit users. the board approved these recommendations. transparency scoring the scoring system groups disclosure elements into ‘core’ and ‘additional.’ core items will be required for all recently developed methods submissions and strongly encouraged for longstanding methods studies. in order to ensure that overall transparency in the field was not reduced by a sudden increase in requirements that might lead longtime data providers to stop sharing data with the center, the committee decided not to change existing requirements for traditional methods studies. over time, the center hopes that the transparency project will provide an incentive for all data providers to adopt disclosure of the core items as recognized best practice. this scoring system was heavily influenced by the aapor and ncpp standards, and overlap across the different standards is significant. figure a, which builds upon work by lois timms-ferrara and marc maynard, provides an overview of the elements included in each of the major proposed and enacted standards from 1968 to today, showing how standards have grown in scope and complexity.(timmsferrara & maynard, 2011) empty cells represent items that are not included in a standard. in recent years, more disclosure has been required or encouraged and more focus has been brought to bear on questions of weighting and sampling, both essential in understanding new methods. nedzi proposal aapor 1967 ncpp (pre2006) ncpp (current) aapor (current) roper summary information survey field organization level 1 immediate core sponsor/funder x x x level 1 immediate core population x x level 1 immediate core dates of interviewing x x level 1 immediate core timing of interviewing in relation to events x topline results x x level 1 core* sample size x x x level 1 immediate core size of any subgroup included in the report x x level 1 immediate core* weighted and unweighted level 2 core* margin of error level 1 immediate https://doi.org/10.29173/iq974 8/19 weldon, kathleen j. (2020) standards and scoring to increase transparency for archived public opinion data, iassist quarterly 44(3), pp. 1-19. doi: https://doi.org/10.29173/iq974 other description of estimated accuracy immediate questionnaire/instrument exact wording of questions/responses x x x level 1 immediate core exact wording of introduction level 2 core interviewer or respondent instructions within 30 days core languages in which survey was offered immediate core complete wording of questions in any foreign languages in which the survey was conducted level 2 any relevant stimuli/visual aids within 30 days core sampling sampling method x level 1 immediate core margin of sampling error x level 1 immediate whether these have been adjusted for design effect due to weighting, clustering, or other factors immediate justification for claims of representativeness core coverage of target population/estimated size of the noncovered population level 2 immediate additional sample design/sampling frame(s) immediate core name of the sample supplier, if sample/frame provided by third party immediate additional proportion of sample provided additional the methods used to recruit the panel or participants, if applicable immediate respondent selection procedure (for example, within household), if any level 2 immediate core description of any quotas or additional sample selection criteria during or post fielding immediate maximum number of attempts to reach respondent level 2 incentives within 30 days other strategies to gain cooperation within 30 days use of breakout routers or chains within 30 days additional details about other types of screening procedures within 30 days interviewing method of interviewing (mode) x x x immediate core response rates x within 30 days core** completion or participation rate (surveys for which a response rate cannot be calculated) core** sample dispositions adequate to compute contact, cooperation and response rates level 2 within 30 days core** https://doi.org/10.29173/iq974 9/19 weldon, kathleen j. (2020) standards and scoring to increase transparency for archived public opinion data, iassist quarterly 44(3), pp. 1-19. doi: https://doi.org/10.29173/iq974 minimum number of completed questions to qualify a completed interview level 2 breakoff rate additional whether interviewers were paid level 2 incentives or compensation provided for participation level 2 within 30 days additional weighting description of weighting procedures (if any) used to generalize data to the full population level 2 immediate weighting benchmark source immediate core variables used to calculate weights immediate core identification of weighting variable in dataset core quality control procedures for managing the membership, participation, and attrition of the panel, if applicable within 30 days methods of interviewer training, supervision, and monitoring, if interviewers were used level 2 within 30 days quality control procedures/data verification within 30 days additional % respondents removed due to quality control checks additional datasets release deidentified raw datasets level 3 core operational and reporting post complete wording, ordering and percentage results of all publicly released survey questions to a publicly available web site for a minimum of two weeks level 3 publicly note their compliance with these principles of disclosure level 3 contact name and information immediate survey organizations reporting results will endeavor to have print and broadcast media include the above items in their news stories and make a report containing these items available to the public upon release x specifications adequate for replication of indices or statistical modeling included in research reports. x within 30 days wordings to describe similar disclosure items vary across different standards. please see individual standards for complete wordings. *this information can be derived from the dataset, a core roper center item. **roper center considers either response rate and aapor definition or disposition codes to calculate the same sufficient to meet core disclolsure. https://doi.org/10.29173/iq974 10/19 weldon, kathleen j. (2020) standards and scoring to increase transparency for archived public opinion data, iassist quarterly 44(3), pp. 1-19. doi: https://doi.org/10.29173/iq974 sources: nedzi, 1969; meyer, 1968; asher, 2001, p. 96; ncpp, n.d.; aapor, 2015. implementation of scoring in order to implement the committee’s recommendations, roper staff had first to map the elements the committee had identified to existing ddi-based metadata in the roper database. in many cases, the mapping was a simple one-to-one connection to existing metadata elements. in some cases, however, a single element from the committee recommendations actually represented several metadata fields, as described in the ddi standard. for example, the list of ‘modes’ as described by the committee included the concepts of both sampling procedure and mode. some committee recommendations expanded the number of metadata fields that the roper center will need to capture. the number of fields related to weighting, for example, has increased from one to four, three required by the new system and an additional notes field. after the full range of necessary disclosure items had been defined by roper staff and verified by the committee, the staff then had to determine the type of metadata field needed. in most cases, the elements of transparency scoring consisted of fields that could be entered in numerical or text formats. but in some cases, the element represented an indicator of whether certain material was included in the archival package: for example, visual aids or complete interviewer instructions. in these cases, the staff determined the most practical approach to be a simple checkbox to indicate presence of this item in the downloadable documentation file. staff also reviewed each element to determine if in any case it might be inapplicable to a particular survey based on methodology or other reasons. for those items, a ‘not applicable’ option would need to be available to avoid surveys being docked points for ‘missing’ inapplicable information in an autoscoring system. finally, staff wrote definitions for each element to ensure that each item was understandable and clear to both data providers and end users. field definition acquisition committee memo item na option field type survey sponsor when applicable, the name of the organization that commissioned the survey. if the same organization funded, designed, and fielded a poll, no sponsor is listed. survey sponsor, including all funding sources yes open text grant funding source funding source for academic or other grantsupported research. survey sponsor, including all funding sources yes open text survey organization the organization that conducted the fieldwork for a survey. field work provider, if outsourced no open text data collection dates the date range during which data was collected from respondents. interview dates no date https://doi.org/10.29173/iq974 11/19 weldon, kathleen j. (2020) standards and scoring to increase transparency for archived public opinion data, iassist quarterly 44(3), pp. 1-19. doi: https://doi.org/10.29173/iq974 universe the population the survey results are intended to represent. also known as "target population." the population of which the results are said to be representative, and the justification for this research claim, and the universe from which the sample was drawn, and the proportion of that universe that had a nonzero chance of participation no open text geographic coverage the geographic area from which data were collected. no list justification for claims of representativen ess a description of the elements of the research design intended to ensure that the survey is representative of the universe it is designed to study. the population of which the results are said to be representative, and the justification for this research claim no open text mode method by which data were collected (such as telephone, in-person, online, etc.) mode: rdd telephone, ivr; listed-sample telephone with live interviewers; listed sample telephone via ivr; other telephone (describe); opt-in online panel; other online (e.g., river samples, mobile apps; hybrid or other (describe)) yes list mode other: description (filtered on previous) method by which data were collected, such as telephone, in-person, online, etc. yes open text sample size the total unweighted number of respondents in the survey. unweighted sample size no numerical sampling procedure: summary the method by which participants in a poll were selected. sampling method: probability, non probability or hybrid and mode: rdd telephone, ivr; listed-sample telephone with live interviewers; listedsample telephone via ivr; other telephone (describe); opt-in online panel; other online (e.g., river samples, mobile apps; hybrid or other (describe)) no open text sampling procedure: respondent selection stage the method by which participants in a poll were selected; specifically, the method by which the individual respondents were chosen. in a multistage sampling process, respondent selection is the final stage of sampling. respondent selection procedure, or absence thereof no open text https://doi.org/10.29173/iq974 12/19 weldon, kathleen j. (2020) standards and scoring to increase transparency for archived public opinion data, iassist quarterly 44(3), pp. 1-19. doi: https://doi.org/10.29173/iq974 sampling frame a list of the items or people forming the universe from which a sample is taken. sample frame and a description of the universe from which the sample was drawn description of all sample weights and sources of weighting targets no open text weight variable name of the variable in the datasets used for weighting the sample. if mutiple weighting schemes were used for different analysis, the variable identified here will be the one used for reporting on the total population, and information on other weights provided in the documentation. yes open text weighting benchmark source data source for benchmarks used to weight the sample yes open text variables used for weighting specific variables used in the calculation of survey weights. response rate calculated to aapor standards, or sample disposition data adequate for the calculation of aapor standard response rates. when aaporstandard response rates cannot be calculated, completion or participation rates shall be provided using another method that is fully disclosed response rate calculated to aapor standards, or sample disposition data adequate for the calculation of aaporstandard response rates. when aapor standard response rates cannot be calculated, completion or participation rates shall be provided using another method that is fully disclosed yes open text https://doi.org/10.29173/iq974 13/19 weldon, kathleen j. (2020) standards and scoring to increase transparency for archived public opinion data, iassist quarterly 44(3), pp. 1-19. doi: https://doi.org/10.29173/iq974 response rate* proportion of contacted respondents who completed the survey. the american association for public opinion research (aapor) provides definitions for six measures of response rates. yes numerical disposition codes* a set of codes or categories used by survey researchers to document the ultimate outcome of contact attempts on individual cases in a survey sample. yes checkbox completion or participation rate the proportion of all cases interviewed of all eligible units ever contacted, used if response rates calculated to aapor standards would be inappropriate for the survey design. yes numerical completion or participation rate details (filter on previous) method for calculation of completion/participatio n rates for surveys for which standard aapor response rates cannot be calculated survey language(s) yes open text survey language(s) languages in which the survey was fielded. survey language(s) no list full question wording with all interviewer instructions, prompts and visual aids a complete survey questionnaire includes all questions, including any screening questions, introductory language, interviewer instructions, and, in the case of some in-person or online polls, visual aids used to illustrate questions. full survey questionnaire with all instructions, prompts, visual aids no checkbox https://doi.org/10.29173/iq974 14/19 weldon, kathleen j. (2020) standards and scoring to increase transparency for archived public opinion data, iassist quarterly 44(3), pp. 1-19. doi: https://doi.org/10.29173/iq974 external sample provider(s) the organization that provided the sampling frame to the field organization, if external sample provider used. sample provider(s), and, if multiple, the share of sample from each provider yes open text proportion of sample provided (filtered on previous) the proportion of the total sample provided by the external sample provider. no numerical use of breakout routers or chains use of online survey routers that screen respondents and direct them to open surveys for which they are qualified or use of chains that direct respondents to additional surveys at the end of completed surveys. use of survey routers or chains yes checkbox breakoff rate the percent of respondents who start the survey but do not finish it. breakoff rate (i.e., the percent of respondents who start the survey but do not finish it) no numerical estimated size of the noncovered population proportion of universe that had a nonzero chance of participation the universe from which the sample was drawn, and the proportion of that universe that had a nonzero chance of participation use of incentives no numerical use of incentives use of incentives provided to survey recipients to reward participation. use of incentives no yes/no https://doi.org/10.29173/iq974 15/19 weldon, kathleen j. (2020) standards and scoring to increase transparency for archived public opinion data, iassist quarterly 44(3), pp. 1-19. doi: https://doi.org/10.29173/iq974 what incentive was provided (filter on previous) specific incentives provided to survey recipients to reward participation. yes open text quality control checks quality control checks performed on the data from the survey. many possible approaches can be taken for quality assurance, such as monitoring online surveys for cases of "speeding" (answering at a rate too fast to allow for adequate comprehension of questions) or "straightlining" (providing identical answers across a range of questions); reinterviewing inperson survey respondents; or random quality control monitoring of telephone interviews. details of quality control checks (e.g., for logic, speeding, straightlining), including how they were performed and results of those checks, including percent of completed interviews excluded or dropped from the analysis details of quality control checks (e.g., for logic, speeding, straightlining), including no checkbox % respondents removed due to checks (filtered on above) percentage of respondents whose cases were removed from the survey before analysis based on quality checks performed. yes numerical at this point, the center data staff shared back with the committee the translation of their work into a plan for a functional scoring system, and after some collaborative revision, the list of elements on which scoring would be based was finalized. the next task was to score one study from the most recent submission from each of thirty-two active data providers. new questions emerged as a result of applying the scoring mechanism to actual studies and were presented by staff to the committee for review. could a survey that was fielded on an omnibus be considered to include ‘all question wordings, including interviewer instructions’ when the sponsor was unable to provide the introductory text and other questions asked on the same instrument? (yes.) is an average response rate for a tracking poll sufficient when a data provider submits a monthly aggregate of daily polls? (yes.) if a multicountry poll offers https://doi.org/10.29173/iq974 16/19 weldon, kathleen j. (2020) standards and scoring to increase transparency for archived public opinion data, iassist quarterly 44(3), pp. 1-19. doi: https://doi.org/10.29173/iq974 different levels of disclosure for different countries, should the highest or lowest level of disclosure be used in scoring? (lowest.) how much information on respondent selection method in an rdd survey is needed to satisfy the requirement? (‘random’ is not enough; method of randomization must be provided.) although refining standards will of course be an ongoing process, this initial effort by center staff in collaboration with the committee ensured that procedures for dealing with the most common issues were in place. display and design with development of the scoring system underway, the center staff and committee also considered the issue of display. to lead users to meaningful engagement with the scoring system, only a button reading ‘transparency details’ will show on search results pages. this button will lead to a page on which a numerical score will be provided, based on the following formula: ((10 points for providing a dataset + 2 points for every other applicable core item + 1 point for every applicable additional item))/(total possible points for all applicable items)) x 10. (results rounded to the nearest .5) studies will also be assigned to one of three descriptive categories. studies with a score >=9 and <=10 ‘greatly exceed requirements;’ scores >=8 and <9 ‘exceed requirements;’ and scores >=6 and <8 ‘meet requirements.’ no study meeting current acquisitions guidelines could score below a 6. these categories were chosen by the committee to frame scoring appropriately: any data provider to the center meets a high standard of transparency, and scoring simply expands upon that baseline. however, the categories are also intended to offer data providers an incentive for offer more information. under the category and numerical score, the elements are provided in a checklist to offer users a quick overview of available documentation. into the future at the time of writing, the center is conducting outreach to current and potential data providers to describe and explain the transparency project; some changes may result from this effort. the center is also inviting feedback from the broader polling research and data archives communities, both now and in the coming years as this project evolves to reflect the rapidly changing polling research environment. development of automated scoring in the ingest system is ongoing. after the completion of that project, scoring will be integrated into the member website. at that point, the display of an overview of available information will be a convenience that should aid research. if the effort is also successful in increasing the information provided by data providers, the opportunity to judge data quality and compare the effects of different methodological approaches should increase. but will the transparency project actually increase transparency? as with so many such questions, it may depend on how success is measured. the results of the aapor ti to date have been hard to quantify. the ti boosts an impressive list of nearly ninety members. however, formal requests for additional information, which are channeled through the aapor transparency committee, have been few and far between. during the tenure of the first chair of the transparency committee, no https://doi.org/10.29173/iq974 17/19 weldon, kathleen j. (2020) standards and scoring to increase transparency for archived public opinion data, iassist quarterly 44(3), pp. 1-19. doi: https://doi.org/10.29173/iq974 request was made, and only three have come through since the second chair took over.(johnson, t. personal correspondence, february 20, 2019; kirzinger, a., personal correspondence, february 11, 2019) informal requests sent directly to survey organizations by researchers, however, are not recorded, and therefore the standards may have had an impact that is currently undocumented. after the 2016 election, aapor once again appointed a committee to review failures in state-level election polling. the committee contacted 59 organizations, an increase over 2008 that likely reflects both the broader geographical scope of the committee’s charge and the proliferation of polling operations. only 35 responded. (kennedy et al, 2018) this low response rate seems to indicate that the ti’s hopes of increasing transparency in polling have not been fulfilled. however, none of the nonresponders was part of the ti. (kennedy et al, n.d.) the polling industry may be moving in two directions at once. as more and more inexpensive online polls are conducted by new organizations with, as mark blumenthal noted back in 2010, little connection to professional associations in the field and no need to rely on the vetting process of major media outlets for dissemination, the overall level of disclosure in the field may decrease. however, those organizations that have embraced the polling evolving commitment to transparency may, under the influence of the ncpp standards, the aapor ti, and roper center’s transparency project, provide far more comprehensive information in both their initial releases and in their archival submissions. the body of well-documented polling datasets preserved for the future should increase substantially. currently non-archiving organizations, in choosing to align themselves with the ‘transparencycommitted’ sector of the polling world, may decide to share their data through the roper center and ensure researcher access to more data now and into the future. these results would represent a major success not only for the roper center’s transparency project, but for the field of public opinion polling as a whole. references aapor, november 2105, code of ethics 2015, viewed 20 december 2019. https://www.aapor.org/standards-ethics/aapor-code-of-ethics.aspx american journal of political science, (n.d.), ajps replication and verification policy, viewed 20 december 2019. https://ajps.org/ajps-verification-policy/ national council on public polls, (n.d.), principles of disclosure, viewed 20 december 2019. http://www.ncpp.org/?q=node/19 asher, h 2001. polling and the public: what every citizen should know (5th edn.), cq press, washington, dc. blumenthal, mm 2005, ‘toward an open-source methodology: what we can learn from the blogosphere’, public opinion quarterly, vol. 69, no. 5, pp. 655–669. https://doi.org/10.1093/poq/nfi059 https://doi.org/10.29173/iq974 https://www.aapor.org/standards-ethics/aapor-code-of-ethics.aspx https://ajps.org/ajps-verification-policy/ http://www.ncpp.org/?q=node/19 https://doi.org/10.1093/poq/nfi059 18/19 weldon, kathleen j. (2020) standards and scoring to increase transparency for archived public opinion data, iassist quarterly 44(3), pp. 1-19. doi: https://doi.org/10.29173/iq974 gallup, g 1948, ‘on the regulation of polling’, public opinion quarterly, vol. 12, no. 4, 733–735. https://doi.org/10.1086/266023 gollin, ae 1992 ‘aapor and the media’, in pb sheatsley & wj mitofsky (eds.) 1992. a meeting place: the history of the american association for public opinion research, american association for public opinion research. herndon, j, & o'reilly, r. (2016). data sharing policies in social sciences academic journals: evolving expectations of data sharing as a form of scholarly communication. databrarianship: the academic data librarian in theory and practice, 22. hill, d 2010, ‘aapor updates poll standards,’ the hill, may 18. hollander, s 1992 ‘survey standards’, in pb sheatsley & wj mitofsky (eds.), 1992. a meeting place: the history of the american association for public opinion research, american association for public opinion research. jacobs, j 2010. dark age ahead. vintage canada, toronto. johnson, t & lakavrakas, p 2012, ‘aapor's transparency initiative,’ paper presented at rc-33 eighth international conference on social science methodology, sydney, australia, viewed 20 december 2019. https://slideplayer.com/slide/6530738/ kennedy, c, blumenthal, m, clement, s, clinton, jd 2018. ‘an evaluation of the 2016 election polls in the united states,’ public opinion quarterly, vol. 82, no. 1, pp. 1-33. https://doi.org/10.1093/poq/nfx047 kennedy, c, blumenthal, m, clement, s, clinton, jd (n.d.), ‘an evaluation of the 2016 election polls in the united states’, american association for public opinion research, viewed 20 december 2019, https://www.aapor.org/education-resources/reports/an-evaluation-of-2016election-polls-in-the-u-s.aspx. ncpp (n.d.), principles of disclosure, viewed 20 december 2019. http://www.ncpp.org/?q=node/19 maynard, m & timms-ferrara, l 2011, ‘methodological disclosure issues and opinion data,’ journal of economic and social measurement, vol. 36, no. 1-2, pp. 19-32. https://doi.org/10.3233/jem-20110340 meyer, p 1968, ’truth in polling,’ columbia journalism review, vol. 7, no. 2, p. 20. miller, mm, & hurd, r 1982, ‘conformity to aapor standards in newspaper reporting of public opinion polls,’ public opinion quarterly, vol. 46, no. 2, pp. 243-249. https://doi.org/10.1086/268716 nedzi, ln 1971, ‘public opinion polls: will legislation help?’ the public opinion quarterly, vol. 35, no. 3, pp. 336341. https://doi.org/10.1086/267917 traugott, m, bolger, g, davis, dw, franklin, c, 2009, ‘an evaluation of the methodology of the 2008 pre-election primary polls,’ american association for public opinion research, viewed 20 december 2019, https://www.aapor.org/education-resources/reports/methodology-2008primary-polls.aspx https://doi.org/10.29173/iq974 https://doi.org/10.1086/266023 https://slideplayer.com/slide/6530738/ https://doi.org/10.1093/poq/nfx047 https://www.aapor.org/education-resources/reports/an-evaluation-of-2016-election-polls-in-the-u-s.aspx https://www.aapor.org/education-resources/reports/an-evaluation-of-2016-election-polls-in-the-u-s.aspx https://www.aapor.org/education-resources/reports/an-evaluation-of-2016-election-polls-in-the-u-s.aspx http://www.ncpp.org/?q=node/19 https://doi.org/10.3233/jem-2011-0340 https://doi.org/10.3233/jem-2011-0340 https://doi.org/10.1086/268716 https://doi.org/10.1086/267917 https://www.aapor.org/education-resources/reports/methodology-2008-primary-polls.aspx https://www.aapor.org/education-resources/reports/methodology-2008-primary-polls.aspx 19/19 weldon, kathleen j. (2020) standards and scoring to increase transparency for archived public opinion data, iassist quarterly 44(3), pp. 1-19. doi: https://doi.org/10.29173/iq974 end-notes 1 kathleen j. weldon is the director of data operations and communications at the roper center for public opinion research, cornell university. she can be reached at kjw93@cornell.edu. 2 for an in-depth exploration of the problems with nonprobability polling, see macinnis, b, krosnick, ja, ho, a, & cho, mj 2018. ‘the accuracy of measurements with probability and nonprobability survey samples: replication and extension’, public opinion quarterly. https://doi.org/10.29173/iq974 mailto:kjw93@cornell.edu 1/2 rasmussen, karsten boye (2019) editor’s notes: sharing qualitative research data, improving data literacy and establishing national data services, iassist quarterly 43(4), pp. 1-2. doi https://doi.org/10.29173/iq972 editor's notes: sharing qualitative research data, improving data literacy and establishing national data services welcome to the fourth issue of volume 43 of the iassist quarterly (iq 43(4) 2019). the first article is authored by jessica mozersky, heidi walsh, meredith parsons, tristan mcintosh, kari baldwin, and james m. dubois – all located at the bioethics research center, washington university school of medicine, st. louis, missouri in the usa. they ask the question 'are we ready to share qualitative research data?', with the subtitle 'knowledge and preparedness among qualitative researchers, irb members, and data repository curators'. the report is obtained through semistructured in-depth interviews with key personnel involved in scientific data sharing: 30 data repository curators, 30 qualitative researchers, and 30 irb staff members in the usa. irb stands for institutional review board, which in other countries might be called research ethics committee or similar. there is an increasing trend towards data sharing and open science, but qualitative data are rarely shared. in health and medicine, qualitative methods are frequently used to explore sensitive topics. the sensitivity of such data necessitates that data are adequately de-identified to protect confidentiality, but this may, in turn, hinder maintaining sufficient contextual detail to enable secondary analyses. de-identifying qualitative data is challenging as sensitive information can be hidden in every corner of data. in contrast, standard methods for de-identification of quantitative data exist at the individual variable level. the article provides insights into the differences in knowledge and preparedness to share qualitative data between the three stakeholder groups. all stakeholder groups lack preparedness for qualitative data sharing. qualitative researchers associate data sharing with quantitative data and are generally unfamiliar with sharing qualitative data, while irb members also have limited experience. among data curators, about half had curated qualitative data, but many only worked with quantitative data. there is a strong need for guidance and standards on qualitative data sharing. the second article is also raising a question: 'how many ways can we teach data literacy?'. we are now in asia with a connection to usa. the author yun dai is working at the library of new york university shanghai, where they have explored many ways to teach data literacy to undergraduate students. these initiatives, described in the article, included workshops and in-class instruction which tempted students by offering up-to-date technology, through online casebooks of topics in the data lifecycle, to event series with appealing names like 'lying with data'. the event series had a marketing mascot a 'lying with data' pinocchio and sessions on being fooled by advertisements and getting the truth out of opinion surveys. data literacy has a resemblance to information literacy and in that perspective, data literacy is defined as 'critical thinking applied to evaluating data sources and formats, and interpreting and communicating findings', while statistical literacy is 'the ability to evaluate statistical information as evidence'. the article presents the approaches and does not conclude on the question concerning 'how many?'. no readers will be surprised by the missing answer, and i am certain readers will enjoy the ideas of the article and the marketing focus. with the last article 'examining barriers for establishing a national data service', the author janez štebe takes us to europe. janez štebe is head of the social science data archives (arhiv družboslovnih podatkov) at the university of ljubljana, slovenia. the consortium of european social science data archives (cessda) is a distributed european social science data infrastructure for access to research data. cessda has many but not all european countries as members. the focus is on the situation in 20 non-cessda member european countries, with emerging and immature data archive services being developed through such projects as the cessda strengthening and widening (saw 2016 and 2017) and cessda widening activities (wa 2018). by identifying and comparing gaps https://doi.org/10.29173/iq972 2/2 rasmussen, karsten boye (2019) editor’s notes: sharing qualitative research data, improving data literacy and establishing national data services, iassist quarterly 43(4), pp. 1-2. doi https://doi.org/10.29173/iq972 and differences, a group of countries at a similar level may consider following similar best practice examples to achieve a more mature and supportive open scientific data ecosystem. like the earlier articles, this article provides good references to earlier literature and description of previous studies in the area. in this project 22 countries were selected all cessda non-members and interviewees among social science researchers and data librarians were contacted with an e-mail template between october 2018 and january 2019. the article brings results and discussion of the national data sharing culture and data infrastructure. yes, there is a lack of money! however, it is the process of gradually establishing a robust data infrastructure that is believed to impact the growth of a data sharing culture and improve the excellence and the efficiency of research in general. submissions of papers for the iassist quarterly are always very welcome. we welcome input from iassist conferences or other conferences and workshops, from local presentations or papers especially written for the iq. when you are preparing such a presentation, give a thought to turning your one-time presentation into a lasting contribution. doing that after the event also gives you the opportunity of improving your work after feedback. we encourage you to login or create an author login to https://www.iassistquarterly.com (our open journal system application). we permit authors 'deep links' into the iq as well as deposition of the paper in your local repository. chairing a conference session with the purpose of aggregating and integrating papers for a special issue iq is also much appreciated as the information reaches many more people than the limited number of session participants and will be readily available on the iassist quarterly website at https://www.iassistquarterly.com. authors are very welcome to take a look at the instructions and layout: https://www.iassistquarterly.com/index.php/iassist/about/submissions authors can also contact me directly via e-mail: kbr@sam.sdu.dk. should you be interested in compiling a special issue for the iq as guest editor(s) i will also be delighted to hear from you. karsten boye rasmussen december 2019 https://doi.org/10.29173/iq972 https://www.iassistquarterly.com/ https://www.iassistquarterly.com/ https://www.iassistquarterly.com/index.php/iassist/about/submissions mailto:kbr@sam.sdu.dk 1/8 katabalwa, anajoyce samuel, bates, jo, and abbott, pamela (2021), potential opportunities and risks of sharing agricultural research data in tanzania, iassist quarterly 45(3-4), pp. 1-8. doi: https://doi.org/10.29173/iq997 potential opportunities and risks of sharing agricultural research data in tanzania anajoyce samuel katabalwa1, jo bates2 and pamela abbott3 abstract purpose: the purpose of this paper was to examine the potential opportunities and risks of sharing agricultural research data in tanzania identified in the existing research literature. design/methodology/approach: the study involved a review of the literature on research data sharing practices. findings: the findings indicate that, research data sharing has significant positive benefits among researchers such as increased high research impact; enhancing international community collaboration among researchers with same interests; improving scientific transparency and accuracy of data (rappert and bezuidenhout, 2016); increasing research output whereby a single dataset can be used to generate more than one article by different authors; and many more. the risks hampering data sharing practices include researchers’ fears that data will be scooped, poached or misused (onyancha, 2016); unreliable electric power; lack of funding to support research data sharing activities; absence of institutional governmental support for data management; perceived lack of benefits for data sharing (leonelli, rappert and bezuidenhout, 2018); and others. however, in tanzania research data sharing is relatively new, thus, there are no governmental agency(ies) mandating or encouraging research data sharing; therefore, there is no research data management; no research open data repositories and no research data sharing policy at any agricultural institution in tanzania. the study recommends that agricultural researchers should be encouraged to share their data, research data policy and data repositories should also be established to support data sharing practices in tanzania. originality and usefulness: from the available literature, this has been the first time that an effort has been made to examine the potential opportunities and risks of sharing agricultural research data in tanzania. the study could be used by agricultural institutions and other institutions to assess the researchers’ needs in supporting research data sharing. also, it can be used by the government and institutions to see the need of establishing open data repositories and open data policies to support research data sharing. keywords research data sharing, open data repositories, research data policy, agricultural institutions, tanzania. introduction the development of information and communication technology (ict) has brought tremendous changes in research. ict supports open science in which data dissemination have triggered the development of open access and open data (bezuidenhout et al., 2017). open access is defined as: ‘free availability on public internet, permitting any user to read, download, copy, distribute, print, search, or link to full-texts of these articles, crawl them for indexing, pass them as data to software, or use them for any other lawful purpose, without financial, legal, or technical barriers other than those inseparable from gaining access to the internet itself’ (budapest open access initiative, 2002). open data refers to ‘data that can be freely used, re-used and redistributed by anyone, subject only, at most, to the requirements to attributed and sharealike’ (open data handbook, 2013). onyancha (2016) defines research data as data from instruments including ‘telescope or raw data from a mass spectrometer to digital maps or full-text documents such as those in the creation of critical editions’ and consists observational data, experimental and computational data. https://doi.org/10.29173/iq997 2/8 katabalwa, anajoyce samuel, bates, jo, and abbott, pamela (2021), potential opportunities and risks of sharing agricultural research data in tanzania, iassist quarterly 45(3-4), pp. 1-8. doi: https://doi.org/10.29173/iq997 various initiatives have been established in supporting open data. initiatives such as elixir, openaire and european open science cloud supports data sharing and re-use (leonelli et al., 2018). moreover, the guardian’s ‘free our data’ campaign forced the uk government to release 8000 open datasets (randalls, 2017). in africa, the african open science platform seeks to foster openness in african science (leonelli et al., 2018). in tanzania, some related initiatives started early in the 2000s. the government’s open data portal was launched in 2014 (twaweza, 2017). the portal hosts data on health, education, water services and tanzania national bureau statistics and efforts are underway to include agriculture and transport datasets (urt, 2017). furthermore, tanzania hosted an africa open data conference where open data sharing was emphasized (urt, 2015). methodology the study involved a review of the literature on research data sharing practices. specifically, a thematic review was conducted using some selected papers on data sharing practices. the researchers consulted databases such as google scholar and google to select the papers used in this study. the keywords used were ‘open data’, ‘open data and africa’, ‘open data and tanzania’, ‘research data sharing’, ‘research data and africa’, and ‘research data and tanzania’. research environments in tanzania tanzania is among low-resource countries which (leonelli et al., 2018) comment that they are challenged with ‘low-resourced environment with intermittent access to a broadband connection and inexpensive outdated instruments’ affecting excellence of research conducted and research goals choices. researchers’ (faculty members) working environments at tanzania’s agricultural institutions are equivalent to other developing countries. researchers work from offices having desktop computers connected with internet. however, not all researchers have computers; forcing others to buy their laptops and software using their own money; many offices have old computers, challenged with unreliable bandwidth and frequent power outage as reported by bezuidenhout et al. (2017) and rappert and bezuidenhout (2016). confirming bezuidenhout's et al. (2017) study concerning institutions’ lacking proxy servers; very few institutions have a proxy server supporting access to information resources outside the ip range of the institution, enabling researchers to continue working from home. agricultural institutions have field-based laboratories and research-based laboratories that are relatively in good conditions. although, institutions are doing laboratories maintenance, buying new equipment whenever funds are available; they are lacking good, modern equipment, chemical reagents and other important tools. therefore, sometimes researchers are forced to send compounds to other laboratories within tanzania or abroad for activity testing. rappert and bezuidenhout (2016) argue that sending compounds abroad entails costs and delays in paper publishing. high teaching and supervision workloads (bezuidenhout et al., 2017; rappert and bezuidenhout, 2016) is challenging researchers at tanzania’s agricultural institutions; limiting time for research activity. researchers have big numbers of undergraduate students to teach (ranging from 100 to 600 students – especially for shared courses) and postgraduate students to supervise, yet same staff might have three to four courses per semester. although research publications are used for promotions, researchers are not given time and support to conduct research and publish. like the researchers in kenya (rappert and bezuidenhout, 2016) some researchers use their own money to pay for data collection and/or for publication fee in some journals to acquire publications to support their promotions. moreover, each academic year researchers are evaluated in terms of research effectiveness using various criteria whereby aside from number of publications and students supervised, funds attracted by a researcher is also applied. very few can engage in projects that attract funds at institutions. thus, many research activities are connected to graduate students as reported by bezuidenhout et al. (2017) and rappert and bezuidenhout (2016). https://doi.org/10.29173/iq997 3/8 katabalwa, anajoyce samuel, bates, jo, and abbott, pamela (2021), potential opportunities and risks of sharing agricultural research data in tanzania, iassist quarterly 45(3-4), pp. 1-8. doi: https://doi.org/10.29173/iq997 researchers acquire research project funds either from costech (a governmental agency) or international funders from abroad. to acquire funds, researchers must write a proposal and win the specific project call. finding time for funds application and establishing collaboration is taxing (rappert and bezuidenhout, 2016). tanzania’s institutions have libraries with both printed and electronic resources on subscription and open access basis. however, due to insufficient funds, only few information resources are through subscriptions. many libraries rely on open access and resources supported by research4life and eifl (electronic information for libraries). contrasting raju's (2018) study regarding humanities and social sciences backgrounds of librarians; in tanzania, there are varied backgrounds including library and information science, social sciences, information studies, icts, and others. nevertheless, librarians’ skills are lacking to cope with changing technology (use of it, artificial intelligence) and ever-changing needs of users. the question is: what are the potential opportunities and risks of sharing agricultural research data in tanzania? in order to address this question, some other sub-questions were asked as follow: does data have any value? are there social and political dynamics influencing research data sharing in tanzania? are there data sharing policies and open data repositories at agricultural institutions in tanzania? is there research data management at agricultural institutions in tanzania? what are the benefits of research data sharing? what are the risks of sharing agricultural research data? potential opportunities for sharing agricultural research data agricultural sector in tanzania agriculture is a mainstream activity and key contributor to overall growth and development in tanzania. agriculture provides about 66.9% of employment, 29% of gdp, 30% of exports and 65% of inputs to the industrial sector (cpf, 2017). as a mainstream activity of many people in sub-sahara africa (ssa), agriculture takes a centre stage in data sharing (onyancha, 2016). data sharing can contribute to the agricultural sector’s improvement. the government predicts that agriculture will lead to growth and transformation of the economy (cpf, 2017). this is an opportunity for agricultural researchers to share data to contribute to economic growth. big data, data as a commodity many organisations depend on data to understand things and make correct decisions. the meteorological organisations, for example, depend heavily on data to explain the changes in climate, why it is changing and what should be done to respond (bates, 2017a). in developed countries like the us, big companies like bayer and monsanto control the food system whereby data collection involves ‘use of sensors, ranging from pieces attached to farm machinery to satellites’ (davidson, 2018). in tanzania, researchers and farmers access data from the ministry of agriculture, and/or through the government’s open data portal. this is an opportunity for researchers and other companies to invest in digital agriculture. data has been regarded as a precious commodity and ‘new oil’ fueling fourth industrial revolution (bates, 2017a). according to bates (2017a), financial markets depend on data and key actors in markets are businesses and government; ‘weather derivatives are types of climate risk product traded’ whereby volumes of weather data required by climate risk industry are treated as a commodity traded by the national meteorological agency. likewise, environment data in the us and uk are traded as market goods rather than public goods (randalls, 2017). in tanzania, the tanzania meteorological agency (tma) is a government agency responsible for meteorological services, climate services and warnings, weather services and advisories information (urt, 2020). not all data are considered public goods (randalls, 2017); some data collected by tma are treated as market goods including data requested by various industries. tma confirms that one of its responsibilities is to ‘collect fees for data, products and services https://doi.org/10.29173/iq997 4/8 katabalwa, anajoyce samuel, bates, jo, and abbott, pamela (2021), potential opportunities and risks of sharing agricultural research data in tanzania, iassist quarterly 45(3-4), pp. 1-8. doi: https://doi.org/10.29173/iq997 rendered’ by the agency (tma, 2020). therefore, there is a commercial value of data and services. bates (2017a) and randalls (2017) questioned whether these practices are good for society, which identifies that there are some possible risks that need to be considered in sharing such research data. social and political dynamics influence on research data sharing social and political forces influence research data sharing. one force comes from journal publishers’ requirement for authors to submit data as they submit papers for publishing, and the other comes from funding agencies and professional societies (bates, 2017b; onyancha, 2016; rappert and bezuidenhout, 2016). in tanzania, researchers are mandated only to submit their data when publishing in journals demanding data submission or to international funders of different projects who require data submission. unlike in the us and south africa where the us national science foundation and national research foundation (nrf) scheme require fund beneficiaries to submit their research products (datasets, software, non-traditional research outputs, final peer-review publications) (onyancha, 2016); in tanzania, there is no governmental agency mandating agricultural researchers for the same. researchers receiving funds through costech must submit reports and publications published out the project but only encouraged to share data, it is not a requirement. thus, tanzania contributes fewer (0.1%) data records in the data citation index (onyancha, 2016). this is an opportunity for tanzania to encourage researchers to share their research data and for the tanzanian government to think of establishing a national mandate to support research data sharing. data sharing policies and open data repositories various regulatory efforts including national funding councils and institutions develop policies to encourage or mandate research data sharing (bates, 2017b). these regulatory frameworks ‘come to shape the collection, use and accessibility of data’ (randalls, 2017). also, governments develop policies supporting data sharing. change of governmental regulations and establishment of the uk government data policy made weather data collected by the meteorological organization, available free for re-use by anyone (bates, 2017a). like in many african institutions (leonelli et al., 2018); in tanzania, there is no data sharing policy (mushi et al., 2020; shao and saxena, 2019), only a draft for open data policy has been developed (urt, 2017). nevertheless, few academic institutions have started the process of establishing the policies which will support research data sharing. despite having institutional repository policies supporting publications sharing in few of tanzania’s academic institutions, there are no institutional open data policies and hence no institutional research open data repositories at various academic institutions including agricultural academic institutions. it would have been important to have open data policies stipulating guidelines regarding data sharing (bezuidenhout et al., 2017). having open data policies will provide an opportunity of establishing research open data repositories at agricultural academic institutions. libraries have a very important role in establishing and implementing research open data repositories in tanzania. mushi et al. (2020) argue that libraries should be involved from the early stage especially in creating a needs assessment of the academic research community, participation in creation of research data management (rdm) policies, creating awareness, training reseachers, and establishing open data repositories. research data management many african institutions have no formal research data management (rdm) (bezuidenhout et al., 2017; leonelli et al., 2018). researchers require data management and curation skills, mentorship skills and knowledge in data engagement activities for enhanced research community sharing norms (bezuidenhout et al., 2017). research data sharing is relatively new in tanzania, few academic institutions are at the initial stage of establishing rdm policies and support, both are important and training for researchers is needed. https://doi.org/10.29173/iq997 5/8 katabalwa, anajoyce samuel, bates, jo, and abbott, pamela (2021), potential opportunities and risks of sharing agricultural research data in tanzania, iassist quarterly 45(3-4), pp. 1-8. doi: https://doi.org/10.29173/iq997 benefits of research data sharing various scholars outlined the benefits regarding research data sharing. these include: increase in high research impact as many prominent scientific researches refer to data than articles because data attract more citations than articles; enhance international collaboration and community among researchers with same interests; improve scientific transparency and accuracy of data hence fostering scientific integrity (onyancha, 2016; rappert and bezuidenhout, 2016); increases research output whereby a single dataset can be used to generate more than one article by different authors; accelerate socioeconomic development in ssa; increase accessibility and availability of research findings (onyancha, 2016); increase researchers’ visibility; fostering identification of research questions and direction; help in attaining intellectual credit and peer recognition through standard citation; and improves data management capacities through greater resourcing (rappert and bezuidenhout, 2016). the perceived benefits of research data sharing among agricultural researchers in tanzania require investigation. risk of sharing agricultural research data scholars have mentioned several risks hampering research data sharing among researchers. these include: researchers’ fears that data will be scooped, poached or misused (bates, 2017b; bezuidenhout et al., 2017; onyancha, 2016; rappert and bezuidenhout, 2016); unreliable power; insecure/absence of ict infrastructure (bezuidenhout et al., 2017; leonelli et al., 2018; rappert and bezuidenhout, 2016); lack of skills on using repositories (onyancha, 2016; rappert and bezuidenhout, 2016; bates, 2017b); lack of fund to support research data sharing activities (leonelli, rappert and bezuidenhout, 2018; onyancha, 2016); absence of institutional governmental support for data management (bezuidenhout et al., 2017; leonelli et al., 2018); inadequate time to deposit data in repositories (bezuidenhout et al., 2017; onyancha, 2016); perceived lack of personal benefits (leonelli et al., 2018; rappert and bezuidenhout, 2016); confidentiality issues (bates, 2017b; rappert and bezuidenhout, 2016); lack of ict support and curation solutions; low bandwidth conditions; lack of agency to actively counter misappropriation of data; lack of confidence in researchers’ data compared to data produced in developed countries; lack of training in data engagement activities (bezuidenhout et al., 2017); technological obsolescence; absence of research data management policies; unequal power relation in the setting of standards for what counts as good science worldwide (leonelli et al., 2018); and fear of losing to those who can publish easily (rappert and bezuidenhout, 2016). the perceived risks of data sharing by agricultural researchers in tanzania should be investigated. conclusion research data sharing is very valuable to various institutions including agricultural institutions in tanzania. research data sharing is being necessitated by different forces. however, in tanzania research data sharing is relatively new, thus, no governmental agency mandating or encouraging research data sharing, therefore, there is no research data management; no research open data repositories and no research data sharing policies at agricultural institutions in tanzania. moreover, the perceived benefits and risks of research data sharing among agricultural researchers in tanzania is not known. this situation calls for further research to fill this gap. table 1: summary of benefits and risks of data sharing benefits of data sharing scholars increase in high research impact as many prominent researches refer to data than articles onyancha, 2016; rappert and bezuidenhout, 2016 enhance international collaboration and community among researchers with same interests onyancha, 2016; rappert and bezuidenhout, 2016 https://doi.org/10.29173/iq997 6/8 katabalwa, anajoyce samuel, bates, jo, and abbott, pamela (2021), potential opportunities and risks of sharing agricultural research data in tanzania, iassist quarterly 45(3-4), pp. 1-8. doi: https://doi.org/10.29173/iq997 improve scientific transparency and accuracy of data hence fostering scientific integrity onyancha, 2016; rappert and bezuidenhout, 2016 increases research output whereby a single dataset can be used to generate more than one article by different authors onyancha, 2016 accelerate socio-economic development in ssa onyancha, 2016 increase accessibility and availability of research findings onyancha, 2016 increase researchers’ visibility; fostering identification of research questions and direction rappert and bezuidenhout, 2016 help in attaining intellectual credit and peer recognition through standard citation rappert and bezuidenhout, 2016 improves data management capacities through greater resourcing rappert and bezuidenhout, 2016 risks of data sharing researchers’ fears that data will be scooped, poached or misused bates, 2017b; bezuidenhout et al., 2017; onyancha, 2016; rappert and bezuidenhout, 2016 unreliable power; insecure/absence of ict infrastructure bezuidenhout et al., 2017; leonelli et al., 2018; rappert and bezuidenhout, 2016 lack of skills on using repositories onyancha, 2016; rappert and bezuidenhout, 2016; bates, 2017b lack of fund to support research data sharing activities leonelli, rappert and bezuidenhout, 2018; onyancha, 2016 absence of institutional governmental support for data management bezuidenhout et al., 2017; leonelli et al., 2018 inadequate time to deposit data in repositories bezuidenhout et al., 2017; onyancha, 2016 perceived lack of personal benefits leonelli et al., 2018; rappert and bezuidenhout, 2016 confidentiality issues bates, 2017b; rappert and bezuidenhout, 2016 lack of ict support and curation solutions bezuidenhout et al., 2017 low bandwidth conditions bezuidenhout et al., 2017 lack of agency to actively counter misappropriation of data bezuidenhout et al., 2017 https://doi.org/10.29173/iq997 7/8 katabalwa, anajoyce samuel, bates, jo, and abbott, pamela (2021), potential opportunities and risks of sharing agricultural research data in tanzania, iassist quarterly 45(3-4), pp. 1-8. doi: https://doi.org/10.29173/iq997 lack of confidence in researchers’ data compared to data produced in developed countries bezuidenhout et al., 2017 lack of training in data engagement activities bezuidenhout et al., 2017 technological obsolescence leonelli et al., 2018 absence of research data management policies leonelli et al., 2018 unequal power relation in the setting of standards for what counts as good science worldwide leonelli et al., 2018 fear of losing to those who can publish easily rappert and bezuidenhout, 2016 references bates, j. (2017a) ‘big data, open data and the climate risk market’, in brevini, b. and murdock, g. (eds) carbon capitalism and communication. palgrave macmillan, cham, pp. 83–93. doi: https://doi.org/10.1007/978-3-319-57876-7_7. bates, j. (2017b) ‘the politics of data friction’, journal of documentation, 74(2), pp. 412–429. doi: https://doi.org/10.1108/jd-05-2017-0080. bezuidenhout, l. m.; leonelli, s.; kelly, a. h; and rappert, b. (2017) ‘beyond the digital divide : towards a situated approach to open data’, science and public policy, 44(4), pp. 464–475. doi: https://doi.org/10.1093/scipol/scw036. budapest open access initiative (2002) budapest open access initiative | read the budapest open access initiative. available at: https://www.budapestopenaccessinitiative.org/read (accessed: 23 april 2020). cpf (2017) country programming profile for united republic tanzania. available at: http://www.fao.org/3/a-bt133e.pdf. (accessed: 23 april 2020) davidson, j. (2018) bayer, monsanto and big data: who will control our food system in the era of digital agriculture and mega-mergers? available at: https://medium.com/@foe_us/bayer-monsantoand-big-data-who-will-control-our-food-system-in-the-era-of-digital-agricultureaae80d991e4d (accessed: 22 april 2020). leonelli, s., rappert, b. and bezuidenhout, l. (2018) ‘introduction : open data and africa’, data science journal, 17(5), pp. 1–3. mushi, g. e., pienaar, h. and deventer, m. van (2020) ‘identifying and implementing relevant research data management services for the library at the university of dodoma , tanzania’, data science journal, 19(1), pp. 1–9. onyancha, o. b. (2016) ‘open research data in sub-saharan africa : a bibliometric study using the data citation index’, publishing research quarterly, 32, pp. 227–246. doi: https://doi.org/10.1007/s12109-016-9463-6. open data handbook (2013) what is open data? available at: http://opendatahandbook.org/guide/en/what-is-open-data/ (accessed: 23 april 2020). raju, r. (2018) ‘from “ life support ” to collaborative partnership: a local/global view of academic libraries in south africa’, college & research libraries news, pp. 30–33. randalls, s. (2017) ‘commercializing environmental data: seeing like a market’, in. routledge. rappert, b. and bezuidenhout, l. (2016) ‘data sharing in low-resourced research environments’, https://doi.org/10.29173/iq997 https://doi.org/10.1007/978-3-319-57876-7_7 https://doi.org/10.1108/jd-05-2017-0080 https://doi.org/10.1093/scipol/scw036 https://www.budapestopenaccessinitiative.org/read http://www.fao.org/3/a-bt133e.pdf https://medium.com/@foe_us/bayer-monsanto-and-big-data-who-will-control-our-food-system-in-the-era-of-digital-agriculture-aae80d991e4d https://medium.com/@foe_us/bayer-monsanto-and-big-data-who-will-control-our-food-system-in-the-era-of-digital-agriculture-aae80d991e4d https://medium.com/@foe_us/bayer-monsanto-and-big-data-who-will-control-our-food-system-in-the-era-of-digital-agriculture-aae80d991e4d https://doi.org/10.1007/s12109-016-9463-6 http://opendatahandbook.org/guide/en/what-is-open-data/ 8/8 katabalwa, anajoyce samuel, bates, jo, and abbott, pamela (2021), potential opportunities and risks of sharing agricultural research data in tanzania, iassist quarterly 45(3-4), pp. 1-8. doi: https://doi.org/10.29173/iq997 prometheus, 34(3-4), pp. 207–224. doi: https://doi.org/10.1080/08109028.2017.1325142. shao, d. d. and saxena, s. (2019) ‘barriers to open government data (ogd) initiative in tanzania : stakeholders ’ perspectives’, growth and change, 50, pp. 470–485. doi: https://doi.org/ 10.1111/grow.12282. tma (2020) majukumu ya mamlaka ya hali ya hewa tanzania. available at: http://www.meteo.go.tz/pages/functions-of-tma (accessed: 21 april 2020). twaweza (2017) icts and national agricultural research systems tanzania. available at: https://www.twaweza.org/go/monitoring-series. (accessed: 23 april 2020) urt (2015) ‘africa open data conference: developing africa throguh open data.’ available at: http://africaopendata.net/wp-content/uploads/2017/06/aodc_report-final_withlogos.pdf (accessed: 18 april 2020). urt (2017) africa open data conference: developing africa through open data. dar es salaam. doi: https://doi.org/10.1016/0144-2449(89)90018-3. urt (2020) tanzania meteorological agency (tma). available at: https://www.kilimo.go.tz/index.php/en/stakeholders/view/tanzania-meteorological-agencytma. (accessed: 24 april 2020) ____________________ endnotes 1 anajoyce samuel is a lecturer at the sokoine university of agriculture. she can be contacted through joykatabalwa@sua.ac.tz 2 dr. jo bates is a senior lecturer at the university of sheffield. she can be contacted through jo.bates@sheffield.ac.uk 3 dr. pamela abbott is a senior lecturer at the university of sheffield. she can be contacted through p.y.abbott@sheffield.ac.uk https://doi.org/10.29173/iq997 https://doi.org/10.1080/08109028.2017.1325142 https://doi.org/%2010.1111/grow.12282 https://doi.org/%2010.1111/grow.12282 http://www.meteo.go.tz/pages/functions-of-tma https://www.twaweza.org/go/monitoring-series http://africaopendata.net/wp-content/uploads/2017/06/aodc_report-final_withlogos.pdf https://doi.org/10.1016/0144-2449(89)90018-3 https://www.kilimo.go.tz/index.php/en/stakeholders/view/tanzania-meteorological-agency-tma https://www.kilimo.go.tz/index.php/en/stakeholders/view/tanzania-meteorological-agency-tma mailto:joykatabalwa@sua.ac.tz mailto:jo.bates@sheffield.ac.uk mailto:p.y.abbott@sheffield.ac.uk 1/10 kim, hyowon, kim, do won & yang, jungwon (2024) understanding motivations and future needs for data deposits at korea social science data archive, iassist quarterly 48(4), pp. 1-10. doi: https://doi.org/10.29173/iq1110 the creative commons-attribution-noncommercial license 4.0 international applies to all works published by iassist quarterly. authors will retain copyright of the work and full publishing rights. understanding motivations and future needs for data deposits at korea social science data archive hyowon kim1, do won kim2, jungwon yang3 abstract korea social science data archive (kossda) has been integral in archiving and disseminating social science data in south korea. as it transitioned into an independent research center under seoul national university, we evaluated current data deposit process of kossda through in-depth interviews with institutional data donors and identified future needs of data depositors. motivations behind data deposit included recognizing data as a public good, addressing resource shortages in data curation, and enhancing data promotion. future needs for kossda included developing qualitative data guidelines, strengthening collaborations with donor institutions and academia as a whole, and automating the deposit process. these results guide the long-term strategy of kossda to remain a leader in social science data archiving, ensuring its alignment with the evolving needs of the academic community. keywords korea social science data archive, kossda, social science data, data deposition introduction the korea social science data archive (kossda) has been a pioneer in archiving and sharing social science data since 2006 in south korea. it has gathered various quantitative and qualitative data that were scattered across individual researchers or research institutions, and made these data publicly available. kossda was designed in accordance with the open archival information system (oais) reference model, which is the international standard model. their data indexing and metadata creation follow the metadata encoding and transmission standards (mets) and data documentation initiative (ddi) 2.5. this initiative has not only facilitated data reuse by the academic community including researchers and students, but also opened data access to the general public. in 2022, kossda transitioned from being a part of the asia center of seoul national university (snu) to becoming an independent research center under the social science college of snu. this significant transition highlights the urgent need to develop a long-term strategy for kossda, positioning it as a leading independent research institution at the forefront of social science data. the importance of establishing long-term strategies is further emphasized by the situations kossda currently faces. notably, there is a lack of national and institutional foundation for the archiving and management of social science data in south korea. even for research funded by national resources, there are no national laws mandating the submission of data management plans. in late 2023, a bill was proposed to establish the 'act on the management and utilization promotion of national research data'. this legislation aimed to promote research data sharing and build an open research https://doi.org/10.29173/iq1110 https://kossda.snu.ac.kr/ 2/10 kim, hyowon, kim, do won & yang, jungwon (2024) understanding motivations and future needs for data deposits at korea social science data archive, iassist quarterly 48(4), pp. 1-10. doi: https://doi.org/10.29173/iq1110 ecosystem, including mandating the submission of data management plans (dmps). however, this bill was discarded due to the expiration of the legislative term4. this legislative setback illustrates the ongoing challenges in establishing systematic management and sharing of research data in south korea, particularly in the social sciences. consequently, research data in social sciences continue to lack systematic management and sharing protocols in the country. in this context of insufficient interest and institutional support by the government, data management and sharing are more dependent on the will and capacity of the research community than on governmental initiatives or regulations. on the other hand, this situation presents an opportunity for kossda to take a leading role, especially at a time when individual researchers and institutions need a focal actor for developing norms about data management, sharing, and reuse. moreover, while institutional depositors are primary data contributors to kossda, they have also begun establishing their own data archive agencies. there has been a growing interest in self-archiving within each institution, with an increased emphasis on the transparency, reproducibility, and usability of research results, which would give these institutional depositors direct control and autonomous management of their data. although kossda has been a part of this open science movement, it is crucial to thoroughly analyze whether its future relationships with these data depositor institutions will lean more towards competition or collaboration. to summarize, as a crucial step towards establishing kossda as a leading independent research institution in the field of social science data archives, this study aims to review the current data deposit process and identify the future needs of depositors. to this end, this study involved conducting oneon-one, in-depth interviews with individual representatives of institutional data donors. these interviews help us understand the motivations behind data donations, evaluate kossda's strategies for data archiving and curation, and comprehend future expectations of data depositors that kossda needs to meet. background a significant difference between the current situation in south korea and that of other developed countries is the role of data sharing requirements from national and funding agencies. countries such as the united states and the united kingdom are actively pursuing policies to promote the sharing and utilization of research outcomes that have received public funding (boehm et al., 2023; kvale & pharo, 2021; smale et al., 2020; williams et al., 2017; whyte & tedds, 2011). in the united states, all federal agencies spending over $100 million annually on r&d are required to submit a public access plan for publicly funded research outputs (choi & lee, 2020). notably, the national institutes of health (nih) has mandated the submission of a data sharing plan for research projects exceeding $500,000 since 2003. similarly, the national science foundation (nsf) has required the inclusion of a data management plan (dmp) for all research projects since january 2011 (choi & lee, 2020). these policies view data sharing as an "economy driver" (smale et al., 2020), ensuring that "full value would be realized from the collected data" (smale et al., 2020). this approach enables follow-up studies and allows for the validation of other research through data reuse. under such institutional support and attention, researchers find that sharing data creates opportunities for collaboration and recognition of their work as data producers (boehm et al., 2023; kvale & pharo, 2021; smale et al., 2020; williams et al., 2017; whyte & tedds, 2011). https://doi.org/10.29173/iq1110 3/10 kim, hyowon, kim, do won & yang, jungwon (2024) understanding motivations and future needs for data deposits at korea social science data archive, iassist quarterly 48(4), pp. 1-10. doi: https://doi.org/10.29173/iq1110 however, in the south korean context, norms for managing and sharing research data are still not well-established. the field of social sciences in korea receives less institutional attention and support compared to scientific and technological fields. yet, even in the scientific and technological fields that receive relatively more national attention regarding research data management and sharing, the situation is far from ideal. according to the ministry of science and the national research council of science & technology (nst), among four major science and technology institutes in south korea, only one has established an online repository for storing r&d research data (cho, 2023). moreover, none of these institutes have linked their repositories to dataon, the national research data archive for science and technology fields (cho, 2023). the south korean context is particularly important given the lack of national interest and institutional support, especially in non-scientific fields. research data management is not prioritized, and data sharing is not mandatory. this situation differs from europe and the united states, where data sharing has been imposed onto the research community by external forces (smale et al., 2020). in this context, the research community in south korea must take the lead. social science archive institutions like kossda have taken the initiative to drive the open science movement within the country. kossda work to convince other social science research institutions to deposit their data, which kossda then quality-controls and makes publicly available. this proactive approach positions kossda as a key player in promoting open science practices in the south korean social science community. however, research institutions are beginning to establish their own archives for the data they produce. this trend of institutions sharing data through their own proprietary archives raises important questions about how to foster future collaboration with these institutions. in light of this evolving landscape, this study aimed to gather opinions and perspectives from data depositors on how data deposits can be sustained and expanded within this changing relationship. method to evaluate data deposit process of kossda and identify any future needs of data depositors, we conducted in-depth individual interviews with representatives from eight institutions from august 25 to september 2, 2022. we specifically targeted representatives responsible for data deposits from each institution. the selection was based on institutions that had deposited datasets with the highest download rates within their respective categories. the institutions spanned various organizations, including: four national/public research institutions, one regional research institution, one private research institution, one university research institution, and one national agency. we prepared interview questions and scripts to gather information on motivations of each institution for data deposits, data sharing and archiving status of each institution, strengths and weaknesses of the current data deposit process with kossda, and expectations about kossda’s contribution related to research data management services. the main questions are presented in table 1. yet, for in-depth exploration of perspectives of data donors, our interviewers were guided to use a semi-structured interview method while they led a conversation with interviewees. each interview lasted between 6090 minutes. https://doi.org/10.29173/iq1110 4/10 kim, hyowon, kim, do won & yang, jungwon (2024) understanding motivations and future needs for data deposits at korea social science data archive, iassist quarterly 48(4), pp. 1-10. doi: https://doi.org/10.29173/iq1110 table 1. key questions on institutional data deposition experience and expectations for kossda 1. could you describe the current state of data management practices within your research institution? 2. what are the key motivations driving your institution to deposit research data? 3. what is your personal perspective on data deposition? 4. since beginning to deposit data, what positive outcomes or benefits has your institution observed? 5. during the data deposition process, what challenges or obstacles does your institution typically encounter? 6. regarding the data deposition process, what improvements or changes would your institution like to see implemented by kossda? we assigned one interviewer who would lead the questions to an interviewee and one note taker who would take detailed notes for each interview session. after each interview, we held a debriefing session with the interviewers. results the interviews revealed underlying motivations and incentives for data deposits at both the institutional and individual levels, which were closely related to the data management situation at each institution (see table 2; the number of participants who mentioned each point is shown in parentheses, along with the corresponding percentage). specifically, these incentives can be categorized as follows: (1) the norm that research data is a public good, (2) the lack of data-curating resources in each institution, and (3) the desire to promote data. [1] the recognition and norm that research data is a public good the ethos that research data is a public good was particularly strong across national or public research institutions. within these institutions, it is well-acknowledged that data produced by government organizations belongs to south korean citizens and that the public has the right to be informed about how tax money has been used. some institutions are developing or have already developed their own platforms for archiving their research data. however, they do not view kossda as a competitor. instead, since research institutions acknowledge that research data is a public good, they aim to foster its preservation, sharing, and reuse through multiple channels. in other words, the data archive platform provided by kossda and that developed and managed by each institution are complementary rather than substitutes. [2] lack of data curating resources in each organization while there is widespread agreement within data donor institutions that data is a public good, the necessary infrastructure and resources to uphold this principle are lacking. this lack of specialized resources for data curation has incentivized these institutions to rely on kossda for data curation. for example, there is a scarcity of researchers who have received professional training in data archiving and curation, and institutional researchers are typically tasked with other responsibilities and administrative workloads, leading to the de-prioritization of data archiving and curation tasks. one of the interviewees had even experienced data loss in their institutions due to insufficient data management resources. therefore, some institutions address their lack of data curation resources by depositing data with kossda, effectively outsourcing this aspect of data management. https://doi.org/10.29173/iq1110 5/10 kim, hyowon, kim, do won & yang, jungwon (2024) understanding motivations and future needs for data deposits at korea social science data archive, iassist quarterly 48(4), pp. 1-10. doi: https://doi.org/10.29173/iq1110 at the same time, a lack of data curation resources may result in fewer data deposits to kossda, rather than encouraging researchers to deposit more. some research institutions lacked the resources to invest in the data review process, which is required for data deposits to kossda. as a result, reducing or facilitating these required steps would reinforce motivations of institutions to deposit data to kossda. table 2. incentive of data curation [3] data promotion research institutions highly value kossda as a social science data archive, and view it as a valuable channel for promoting their data. in particular, they noted an increase in the use of their data after it was deposited in kossda. although a clear causal relationship cannot be established, researchers frequently observe a diversification in the fields of study among authors who publish papers based on their data, when comparing before and after data deposition. for instance, an interviewee mentioned that while their data was primarily used in the field of education in the past, after depositing their data with kossda, they noticed that their data is being used in many other social science domains, including sociology and public administration. the recognition and norm that research data is a public good acknowledgment that government-produced data belongs to citizens (6/8, 75%) kossda seen as complementary, not competitive in data archiving issues; fostering preservation, sharing, and reuse through multiple channels (8/8, 100%) lack of data curating resources in each organization insufficient infrastructure and resources despite agreement on data as a public good (5/8, 62.5%) scarcity of professionally trained researchers in data archiving and curation (6/8, 75%) researchers overwhelmed with their primary research duties; data management and sharing often deprioritized (4/8, 50%) researchers experienced instances of data loss due to insufficient management resources (1/8, 12.5%) some institutions outsource data management to kossda (5/8, 62.5%) lack of resources may also lead to fewer data deposits (5/8, 62.5%) data promotion kossda highly regarded as a social science data archive (4/8, 50%) perceived as a channel to promote data produced by research institutions (5/8, 62.5%) reported increase in data usage after depositing with kossda (3/8, 37.5%) observed diversification in fields of study using data produced by research institutions. for example, data previously used mainly in education is now being utilized in sociology and public administration (2/8, 25%) https://doi.org/10.29173/iq1110 6/10 kim, hyowon, kim, do won & yang, jungwon (2024) understanding motivations and future needs for data deposits at korea social science data archive, iassist quarterly 48(4), pp. 1-10. doi: https://doi.org/10.29173/iq1110 future needs through the interviews, we also identified future needs that both institutions and individuals expect from kossda (see table 3; the number of participants who mentioned each point is shown in parentheses, along with the corresponding percentage). these can be summarized as follows: (1) setting up guidelines for qualitative data, (2) strengthening collaboration including spreading citation standards and enhancing academic collaboration between kossda and donor institutions, and (3) automating the data deposit process. table 3. future needs establishing guidelines for qualitative data develop guidelines for handling and disclosing personal information (5/8, 62.5%) create comprehensive guidelines for collecting, managing, and sharing qualitative data (3/8, 37.5%) implement educational programs for data curators at donor institutions (3/8, 37.5%) strengthening collaboration between kossda and the academic community spreading citation standards: collaborate with journal publishers to establish and spread data citation standards (3/8, 37.5%) promote wider adoption of data citation standards in social science journals and theses (3/8, 37.5%) enhancing academic collaboration: integrate kossda's data into methodology training programs (2/8, 25%) launch joint conferences or journals to promote data reuse (3/8, 37.5%) utilize kossda's seasonal methodology courses for data promotion (2/8, 25%) expand kossda's award program for theses using its data collections (2/8, 25%) automation of the data deposit process implement an automated system for easier and more secure data handling (3/8, 37.5%) provide easy access to usage statistics for donated data (5/8, 62.5%) [1] establishing guidelines for qualitative data the most prominent demand is evident in dealing with qualitative data generated in social science research. at the institutional level, there is a notable lack of common norms or guidelines for managing qualitative data, particularly concerning ethical issues such as the handling and disclosing personal information. this absence of guidelines not only complicates the internal management of qualitative data but also makes depositing this data challenging. the lack of clear standards across different research institutions leads to confusion among researchers. consequently, research institutions and data management representatives often receive inquiries from academic researchers on the use of qualitative data, but they are not equipped with formal training or established protocols to assist them. there is a strong need to create comprehensive guidelines and regulations for collecting, managing, sharing and deposit of qualitative data. https://doi.org/10.29173/iq1110 7/10 kim, hyowon, kim, do won & yang, jungwon (2024) understanding motivations and future needs for data deposits at korea social science data archive, iassist quarterly 48(4), pp. 1-10. doi: https://doi.org/10.29173/iq1110 furthermore, there is a recognized need for an educational program for data curators at donor institutions. [2] strengthening collaboration between kossda and the academic community [2-1] spreading citation standards kossda currently distributes data citation guidelines5 to various users, including data depositors, graduate students, and researchers. however, data donor institutions expressed a desire for these citation norms to be more firmly established and implemented. particularly, for the data citation norms to be spread and established across different social science fields, kossda should collaborate with not only data depositors but also journal publishers. in the south korean context, academic journals lack such guidelines. as of january 25, 2024, prominent journals published by organizations like the korean political science association or the korean sociological association do not include any data citation guidelines in their submission instructions. this disparity highlights the need for wider adoption of data citation standards within academic societies, especially in journals related to social science, as well as in graduate theses and dissertations. as a comparison, in the united states, many academic journals have well-established guidelines for data citation. currently, there are ongoing efforts to develop citation guidelines for various components, including datasets, measurement instruments, and software in south korea. [2-2] enhancing academic collaboration between kossda and donor institutions there is a need for enhancing academic collaboration between kossda and donor institutions. for example, some interviewees suggested using deposited data in coding sessions of kossda's methodology training programs, and collaborating with kossda to initiate joint conferences or journals, aiming to foster the culture of data reuse. all these specific demands are aimed at fostering cooperation to boost data reuse. since 2007, kossda has been at the forefront of promoting social science methodology through its specialized methodology courses conducted across all seasons spring, summer, fall, and winter. each season offers between six to nine classes that cover a wide array of research methodologies, both quantitative and qualitative. particularly, the quantitative research courses are popular, often attracting more than fifty participants per course. with attendees from varied backgrounds, including students, researchers, and professors, data depositors believed that these courses could serve as an excellent platform to promote their data. for instance, instructors can be encouraged to utilize deposited data in their lab sessions during the methodology training programs. they can teach attendees how to access and download data from kossda’s website and analyze it during the lab sessions. hence, the data is not only reused for educational purposes but also has the potential to lead to further research opportunities for junior scholars and students. furthermore, kossda is exploring ways to enhance its contribution to data reuse by establishing joint conferences or journals. currently, kossda awards outstanding graduation theses that use its data collections. however, there have been suggestions that this award program should further expand its scope and impact further. [3] automation of the data deposit process lastly, there is a strong desire for improvements in the data deposit process to enhance its efficiency and security. at the time of the interview, the data deposit procedure was manual and conducted via https://doi.org/10.29173/iq1110 http://www.kpsa.or.kr/contents/bbs/bbs_content.html?bbs_cls_cd=002005002001 http://www.ksa21.or.kr/content/journal/manual.php http://www.ksa21.or.kr/content/journal/manual.php https://kossda.methods.snu.ac.kr/index.php https://kossda.methods.snu.ac.kr/index.php 8/10 kim, hyowon, kim, do won & yang, jungwon (2024) understanding motivations and future needs for data deposits at korea social science data archive, iassist quarterly 48(4), pp. 1-10. doi: https://doi.org/10.29173/iq1110 email. coupled with the lack of specialized resources for data curation within institutions, there is significant interest in automating this system to enable easier and more secure handling. additionally, some institutional representatives wished for easy access to usage statistics for donated data, instead of having to submit separate requests to kossda for this information. recommendations these interviews with representatives of institutions that deposit data in the archive have revealed several key areas where kossda, as a social science data archive institution in south korea, can focus its efforts for long-term development. these areas include: [1] strengthening collaboration with academic journals the interviews reaffirmed the gaps in the korean social science field, particularly regarding data citation practices. moving forward, kossda's long-term development strategy should encompass collaboration not only with research institutions that deposit data to kossda but also with academic journals. this approach differs significantly from more advanced academic circles in europe and the united states, where data citation has already become an established norm. in these advanced academic environments, citing data in publications when using data produced by others has become a common practice. this shift occurred following the introduction of digital object identifiers (dois) for datasets by initiatives like datacite, and the establishment of the fair (findable, accessible, interoperable, reusable) principles as standard in the mid-2010s (silvello, 2018). the discussion in these regions has since evolved beyond mere data citation, now exploring comprehensive ways to acknowledge the contributions of data curation and archiving institutions (buneman et al., 2020). however, in korea, even government agencies and higher institutions such as the national research foundation, which manage publicly funded research, still lack mandatory and specific requirements for research data management and sharing in social science. consequently, the practice of citing open data used in research has not yet been firmly established. to establish such norms within the research community and academia, close collaboration with journals would be most effective. specifically, kossda should collaborate with major social science journals to incorporate data citation guidelines into their submission requirements. this concrete step would help institutionalize data citation practices within the korean academic community. [2] engaging individual researchers while the current interviews focused on institutional representatives, there is a need to study individual researchers' research data management, sharing, and reuse practices to expand research data sharing and reuse. moreover, in an environment where data management plans (dmps) are not standard practice when receiving research funding, kossda should: a) conduct surveys or interviews with individual researchers to understand their data sharing practices and motivations. b) develop strategies to encourage data sharing among individual researchers, considering the complex licensing issues that may arise when research is funded by universities, corporations, or national research foundations. c) provide education and resources to individual researchers on the importance and benefits of data sharing. [3] practical implications for donor-friendly processes simultaneously, kossda needs to improve its deposition process and provide better services to data donors. based on the demand for automation of the data deposit process, the following improvements are suggested: a) implement an automated, user-friendly, and secure system for data deposition, replacing the current manual, email-based process. b) leverage dois to provide regular analyses of citation counts and usage statistics for deposited data. this information should be easily https://doi.org/10.29173/iq1110 9/10 kim, hyowon, kim, do won & yang, jungwon (2024) understanding motivations and future needs for data deposits at korea social science data archive, iassist quarterly 48(4), pp. 1-10. doi: https://doi.org/10.29173/iq1110 accessible to data donors, eliminating the need for separate requests. c) establish a structured timeline for data deposition requests. currently, the timing of data deposition is left to the discretion of the donating institutions. regular reminders and structured timelines can help ensure timely and consistent data deposits. [4] expansion of educational offerings while kossda currently provides education in qualitative and quantitative social science research methodologies, there is an opportunity to expand its educational scope in the areas of data curation and management. this expansion is particularly crucial as researchers often lack awareness and skills in data management. it is imperative to develop educational programs tailored to the diverse needs of researchers at different career stages and roles within research teams. these programs should cater to various groups including undergraduate and graduate students, early-career researchers, established researchers leading projects, and professionals responsible for research data management. such targeted education would address the importance of data management and provide the necessary skills throughout the researcher's lifecycle and within the context of their specific roles in research teams. by broadening its educational offerings, kossda can play a pivotal role in enhancing the data management capabilities of the entire research ecosystem, thereby fostering a culture of effective data curation and sharing within the academic community. conclusion kossda faces the dual challenges of establishing itself as a leader in the data archive field while navigating the evolving landscape of data management and sharing. through interviews with data deposit institution representatives, it has been established that these discussions provide practical evidence for shaping a specific and actionable long-term development agenda. interviews revealed a shared recognition of the importance of data as a public good, the need for improved data curation, and a desire to enhance data visibility and prevent data loss. these insights underscore a clear demand for kossda to establish robust guidelines for qualitative data, strengthen collaborative efforts, especially in creating data citation norms with major journals, and, from a practical standpoint, automate the data deposit process. future needs encompass not only technical and procedural improvements but also a strategic focus on advocacy and education to promote the value of data sharing and management. these future needs collectively position kossda not as a competitor to research institutions in data archiving, but as an integral collaborator and partner within the academic community. this perspective informs the approach to implementing kossda's long-term development plans. it emphasizes the importance of kossda's role beyond its current functions as an archive institution, a platform for social science research methodology education, and a social science research institute. consequently, the research results confirm that the long-term development directions discussed internally at kossda align with those of the academic research community, including researchers and institutions. in response to these future needs, kossda has actively embarked on making improvements. firstly, kossda has initiated research focused on the development of qualitative research guidelines to ensure robust and consistent data curation practices. secondly, in september 2023, kossda introduced the online material deposit system. this upgrade made the entire deposit process digital, including the completion of application forms, pre-review of data through a checklist, file upload, and discussions regarding data usage conditions and disclosure methods. kossda also enhanced the security and stability of data file uploads and integrated new features for viewing online deposit details and data usage statistics. https://doi.org/10.29173/iq1110 https://kossda.snu.ac.kr/deposit/notice 10/10 kim, hyowon, kim, do won & yang, jungwon (2024) understanding motivations and future needs for data deposits at korea social science data archive, iassist quarterly 48(4), pp. 1-10. doi: https://doi.org/10.29173/iq1110 looking forward, kossda must continue to adapt and innovate to meet the evolving needs of its depositors and the broader research community, ensuring that it remains at the forefront of data archiving and management in the field of social sciences. references boehm, r. i., calkins, h., condon, p. b., petters, j., & woodbrook, r. (2023). analysis of u.s. federal funding agency data sharing policies: 2020 highlights and key observations. international journal of digital curation, 17(1). https://doi.org/10.2218/ijdc.v17i1.791 cho, s.-h. (2023). science and technology institutes and government-funded research institutes indifferent to national research project data management. yonhap news agency. https://www.yna.co.kr/view/akr20230927093800017?input=1195m (accessed: august 14, 2024). choi, m.-s., & lee, s. (2020). current status and issues of data management plan in korea. the journal of the korea contents association, 20(6). https://doi.org/10.9728/kca.2020.20.6.220 kvale, l., & pharo, n. (2021). understanding the data management plan as a boundary object through a multi-stakeholder perspective. international journal of digital curation, 16(1), article 1. https://doi.org/10.2218/ijdc.v16i1.746 silvello, g. (2018). theory and practice of data citation. journal of the association for information science and technology, 69(1), 6–20. https://doi.org/10.1002/asi.23917 smale, n. a., unsworth, k., denyer, g., magatova, e., & barr, d. (2020). a review of the history, advocacy and efficacy of data management plans. international journal of digital curation, 15(1), article 1. https://doi.org/10.2218/ijdc.v15i1.525 williams, m., bagwell, j., & nahm zozus, m. (2017). data management plans: the missing perspective. journal of biomedical informatics, 71. https://doi.org/10.1016/j.jbi.2017.05.004 whyte, a., & tedds, j. (2011). making the case for research data management. endnotes 1 phd student in information at university of arizona. she can be reached by email: hyowonkim@ariznoa.edu 2 phd student in information studies at university of maryland. she can be reached by email: dowonkim@umd.edu 3 international government information librarian and associate director of social science and clark library team, university of michigan. she can be reached by email: yangjw@umich.edu 4 http://likms.assembly.go.kr/bill/billdetail.do?billid=prc_a2a3y0z9x0y6g1g5f2d8e4c5d3k5j5 5 please see page 40 for the data citation guidelines. https://kossda.snu.ac.kr/download?file=user_guide.pdf https://doi.org/10.29173/iq1110 https://doi.org/10.2218/ijdc.v17i1.791 https://doi.org/10.2218/ijdc.v17i1.791 https://www.yna.co.kr/view/akr20230927093800017?input=1195m https://www.yna.co.kr/view/akr20230927093800017?input=1195m https://www.yna.co.kr/view/akr20230927093800017?input=1195m https://doi.org/10.9728/kca.2020.20.6.220 https://doi.org/10.2218/ijdc.v16i1.746 https://doi.org/10.2218/ijdc.v16i1.746 https://doi.org/10.1002/asi.23917 https://doi.org/10.2218/ijdc.v15i1.525 https://doi.org/10.2218/ijdc.v15i1.525 https://doi.org/10.1016/j.jbi.2017.05.004 https://doi.org/10.1016/j.jbi.2017.05.004 mailto:hyowonkim@ariznoa.edu mailto:dowonkim@umd.edu mailto:yangjw@umich.edu http://likms.assembly.go.kr/bill/billdetail.do?billid=prc_a2a3y0z9x0y6g1g5f2d8e4c5d3k5j5 https://kossda.snu.ac.kr/download?file=user_guide.pdf newsletter vol. a no. 1 the directory will be useajl to all types of census data users: data archives, libraries, planning organizations, other federal agencies, state and local goverments, and the private sector. by developnent of the directory , the ceta user services division is providing a new information service to the entire user comunity. state and regional data archives (srda) mirray a. straus ihiversity of new hanpshire durham, nh 0382'j (603) 862-1888 a vast anomt of data is available on a state by state basis, or by region. political scientists, demographers, economists, and geographers have used this valu^le resource. other social scieni ists have rarely used this data. part of the reason is that the data is scattered over many sources. social scientists are often not aware of its potential. even those vho realize the potential do not know the full range of variables. they also do not have convenient access. in march 1979 we therefore began to compile a truly ccrprdtensive archive of data on anerican states and regions. this archive (srda) is intended to be the equivalent, for states of the lhited states, of the unan relations area files and other archives of cross-national data. vfell over ?,000 variables have been identified in the preliminary search. other variables are being added continuously. the data files will be available in machine readable form (spss file, card image tape, or cards) or as s printed listing. iwo overlapping archives will make up the srda. the state arc^ve consists of data on each of the 50 states and the district of cblunbia. tv.ere are now about 1,00c variables actually entered in the state archive in the form of an spss system file. the regional archive consists of data on the nine divisions which the u.s. census uses as its main regional classification. all of the variables in the state archive will also be in the regional archive. however, the regional archive will contain additional variables that are not in the state archive. these are variables for v*ilch state by state data could not be obtained. the initial regional archive file will be created sonetime in why use state level data? there are many problans connected with the use of state level data for social science research. anong the most obvious is the ambiguity inherent in a statistic which corisines, for exeinple. new york city and the adirondack rtxntain region. these problems, and also vhat is to be gained fran using state level data, are analyzed in my forthcaning book. state and regional analysis in the newsletter vol. k no. 1 social sciences . among the reasons for using state level data despite these problems are: 1. states are the theoretically appropriate uiit for issues vhlch involve many aspects of goverment, politics, taxes, schools, econcmics, adninistration, and legal issues, as for example in hicks, fv-iedland and jcinson.* however, the uses of sffift data will include many issues for vhich the states are not the natural ixiit. in such cases the decision to use state level, as in many research decisions, involves a trade-off: one accepts the problematic aspects of state data in order to be able to do research v*iich vduld otherwise not be possible or practical. for exanple: 2. historical analysis is possible because sane state level data goes back to colonial times and the nunber of variables available has grown exponentially each generation. 3. causal inferences can aanetimes be more clearly established becaijse the sane data is avail^le for two or more time periods. this permits the use of time series and cross-lagged correlation analysis. ^4. variables can be linked , even thdugh they are located in different surveys and refer to different respondents. this is possible by first converting the individual level data to state level data: for exanple, the percent in each state \to agree that "homosexuality is alveys wrong" fhcm one study, with the percent who oppose the equal rights anakinent from another study. 5. contextual analysis is possible by oonbining variables fran the srda with individual level data fvon specifc surveys. for exanple, kersti yllo has constricted a sexual inequality index for each state and used this to find out if the correlates of a male-ctxninant marriage are the same or different in states viiere women are generally disadvantaged versus those with gi-eater equality between the sexes. types cf dftta 1 . published lists of state by state data . many examples are to be foind in standard almanacs and the statistical abstract of the united states . the largest soiree of pi4)lished data is the u.s. census. 2. aggregated survey data. these surveys were designed for analysis on an individual by individual basis. for the srda. the results are tabulated by state. an exanple is the percent in each state agreeing with sane attitude question. 3. ccnpilations from docunents . many docunents give the state as one itan of information. this can be the basis for conpiling a new variable. for example. who's wx) in america gives the state of birth. it is possible to obtain a measure of the extent to vhich each state has contributed eninent persons by tallying the nimber of eminent anericans bom in each state and dividing that by the population of the state. ^4. indexes fran other variables . ^ ccmbining existing variables it is possible to produce entirely new variables, or to produce an index vhich does a better job of measuring than any one of the variables vhich are conbined to form the index. an exanple would be a measure of the •wioks, alexander, roger friedland and edwin johnson 1978 "qass power and state policy: the case of large business corporations, labor uiions and governmental redistribution in the anerican states." anerican sociological review 13 (3): 302-315 3la/ssist newsletter vol. 4 no. 1 socioeconanic status of each state's population. this could be made up by combining the median income, education, and occupational prestige values from each state. data sources another vay of classifying the data in the srda is according to the accessibility of the source. sections a and e list readily available data sources, each of which contains many variables. the main value of including then in the srda is convenience. researchers using the srda data do not have to pinch, verify, provide variable levels, etc. instead they can acquire a proofed, clean, labeled, ready-to-tln data set. the full value of the s?da, however, will derive fron the inclusion of data from sources which are not readily available and which often will not even be known to researchers. these cone from dozens of books and research reports, reports of goverment agencies, special topic reports by the census, and reports of private organizations such as the institute of life insurance, the audit bureau of circulation, the boy scouts of america, the national wanens political caucus, etc. finally, the least accessible data of all are the state level statistics created by aggregating individual level surveys to provide rates and averages for the states and regions. this is a long and expensive process for which the procedures are now being developed. the initial archive will consist mainly of data given for the 50 states (and the district of colurtiia) in the sources below. a^. milti-topic ccnpendia bacheller, martin a. (ed.) 1979 the hanttond almanac of a ^tlluon facts, records, forecasts. maplewood, n.j.: almanac, inc. book of the states 1978-1979 1978 cotncil of state goverments. lexington, ky. county and city ebta book 1947washington, d.c. bureau of the census. demographic, social and economic profile of states: spring, 1976 1979 current population reports, series p-20, no. 33*^. washington, d.c: bureau of the census. information please 1979 information please almanac, atlas, f, yearbook, 33rd edition. new york: information please rjbl. the world almanac 1979 the world almanac i book of facts 1979. new york: newspaper enterprise association. roswe, arthur e. (ed.) 1978 1978-79 help: the useful almanac. washington, d.c: consuner news. statistical abstract of the uhited states 1878washington d.c: bureau of the census, u.s. department of connsrce. b. specific topics (printed sources ) lalmaiac of anerican rolitics 1978 new york: e. p. cutton lelreau of labor statistics 1/ 1978 handbook of labor statistics. washington, d.c: u.s. department of labor. ! -9newsletter vol. k no. 1 gottfredacxi, michael r., michael j. hindelang. and nicolette parisi (eds.) 1978 sourcebook of q-iminal justice statistics, 1977. washington, d.c.: crimnal justice research center, law ehforcement assistance mministration, u.s. department of justice. grant, w. vance and c. george lind 1979 digest of eaucation statistics 1979. washington, d.c.: u.s. department of health, education, and welfare education division. rosten, leo (ed.) 1975 a guide to the religions of anerica. new york: simon and schuster, vital statistics of the ihited states 1978 washir^ton, d.c.: national center for health statistics, hew. c. aggregated survey data riysical violence in anerican fanilies 1976 survey conducted by response analysis corp. for mrray a. straus, principal investigator. reogram fuelicanob in fr0c3?ess at the untversmr of new hamfjshre straus, mrray a. codebock for the state and regional ifeta ardiives. (anticipated availability for part i, variables 1 to 1,ij99 is march, 1980; for f^rt n. variables 1,500 to 2,999, septnber, 1980). state and regional rteta eljlletin this quarterly pufclication will replace the codebook beginning with variable 3,000. the bllletin will improve on the codebook in three ways: (1) quarterly pialication will make materials availd^le more quickly. (2) in addition to docunentating the source and nature of the data, it will include a printed listing of the statistics for each state and region. (3) t>ie bulletin will include news items dxxit the srdft and occasional ccnmentary and analyses of data included in that issue. the planned plfclication date for the first issue is january 1981. vol21.4bacp fall 1997 15 by ken miller * meaningful relationships introduction the data archive’s thesaurus, hasset (humanities and social science electronic thesaurus), is based upon the unesco thesaurus compiled by jean aitchison (paris; unesco 1977) and has been built up over 18 years so that its coverage reflects the subject matter of the 5,000 datasets held at the data archive. this paper will describe the construction, maintenance and use of the thesaurus as a controlled vocabulary for indexing and a retrieval tool in the data archive’s on-line catalogue biron (bibliographic information retrieval on-line),. it will also outline the data archive’s proposed developments for hasset as an on-line thesaural resource for the social science community in general and as a multilingual free-text retrieval tool within, among others, the nesstar (networked european social science tools and resources) project. thesauri dictionary definitions of thesauri describe them as “a storehouse of information” e.g. a dictionary, or “a list of concepts or words arranged according to sense” e.g. roget’s, or “a list of concepts or words chosen for use in indexing” e.g. the unesco thesaurus. hasset is all this and more and is why we at the data archive consider it just that, a huge asset. its use as a controlled vocabulary means that every dataset whose question or variable covers the same subject material will be indexed by the same concept term. the structured relationships allow the indexer to view candidate terms within a concept hierarchy. the fact that it is machine readable allows instant, easy and consistent maintenance and flexibility in displaying terms in various different ways. it also means that it can be used as a retrieval tool in biron helping the searcher to better define, expand or focus their search hasset there are six basic relationships between the terms held in the thesaurus and these are held in one database table with the simple format of concept term relationship type relationship term. they are 1) use 2) use for uf 3) narrower term nt 4) broader term bt 5) top term tt 6) related term rt. 16 iassist quarterly hence :drinking habits use alcohol use i.e. “drinking habits” is a non-preferred synonym of the preferred term “alcohol use” there will also be the reciprocal entry : alcohol use use for drinking habits alcoholism broader term addiction with the reciprocal entry :addiction narrower term alcoholism i.e. there is a narrower concept “alcoholism” to the subject term “addiction” winter 1997 17 both subject terms “alcoholism” and “addiction” are in two hierarchies, one from the top term “diseases” and one from the top term “social problems”. hence the following entries are found in the database table :alcoholism top term diseases addiction top term diseases alcoholism top term social problems addiction top term social problems n.b. there is no reciprocal entry for a top term relationship, which acts as an aid to understand the scope and meaning of the subject concept under review, and in programming to build up the correct hierarchies. the final relationship is that between two preferred subject terms which are related to each other but are not covered by the nt, bt or tt relationships. the reciprocal entry is also included in the database table. hence:alcohol use related term alcoholism alcoholism related term alcohol use hasset does actually have two other database tables, one which holds a textual clarification of the subject term, known as a scope note (sn), and the second holds a classification code which places the subject term in one fixed hierarchy, so that hasset could be used as a shelving scheme for hard copy documentation, and a marker to show whether the term was taken from the unesco thesaurus or is a data archive new term. 18 iassist quarterly there are, at present, approximately 8,650 terms in hasset; 2,500 of which are non-preferred terms or synonyms. 38,600 relationships exist between these terms and they form 296 hierarchies. approximately half of the terms have been taken from the unesco thesaurus. maintenance & interfaces the tables described above are held in an ingres database and updates to the thesaurus are performed through ‘c’ programs, written at the data archive, which employ embedded sql calls to the underlying tables. the same interface also performs the indexing of the actual datasets with terms from the controlled vocabulary. the program ensures that reciprocal entries are automatically included, terms are correctly positioned in hierarchies with the most appropriate allocation of classification code. the duplication of terms is impossible as is the creation of incorrect relationships between terms. to aid the allocation of terms to the datasets held at the data archive, the indexer has recourse not only to the thesaurus, hierarchical and classification listings described above, but also the scope notes, listings of datasets previously indexed by the term under consideration and a kwic (keyword in context) listing of words from the candidate term. the example below shows the kwic listing for the term “alcohol use”. winter 1997 19 biron & hasset the www interface to both hasset and biron is through dynamically produced html forms from a cgi-bin ‘c’ program with embedded sql calls to the underlying ingres database tables. how then does hasset aid the searcher of the data archive’s on-line catalogue biron. first of all the indexing program ensures that the same subject concept in any dataset held is assigned the same controlled vocabulary term. biron’s first task then is to point the user to the preferred term, if the keyword searched on is not in itself a preferred term; it does this in three ways. firstly it searches the synonyms from the use and uf relationships to see if it can find a match. if it does it will automatically substitute the preferred term and carry out a search immediately. if it cannot match against a non-preferred term then the program produces a kwic listing from the word or words in the search term. finally, if the second option fails, biron will produce another kwic listing, but this time from progressively truncating the search string until a listing is produced, even if it has to be a list of all terms in the thesaurus. hence entering a slight misspelling of “alchol” results in :therefore the searcher is always offered some candidate terms no matter what is entered as the search string. selection is carried out by just clicking on the required term and then the on “search” icon to perform the search. once a search has been carried out the thesaurus is also available to help the searcher redefine their search through the displays described above, by changing to broader or narrower concepts, adding more terms to their search or combining the results from their present search, so that the datasets retrieved also cover the concept of another subject term or terms displayed. 20 iassist quarterly consider the following search:which results in :winter 1997 21 clicking on the thesaurus help icon displays the thesaural entry for “alcohol use”. from which you could select the related term select the button and click on the icon which results in :or you could select more than one related term and click on the extend icon which results in :then from “alcoholism” you could select the top term “social problems” and click on the followed by the to display the full hierarchy 22 iassist quarterly and select up to ten terms from the listing of 269 terms. n.b. n4 indicates a narrower term 4 levels below the selected term. the maximum level for this hierarchy is n6. all 269 terms can be selected for a search by returning to the thesaurus listing and selecting the following options before clicking on the extend icon. which results in:future developments it must be remembered that hasset has been constructed based on the 5,000 datasets held at the data archive as an indexing tool. so therefore the coverage only reflects the subject coverage of these datasets themselves, and because it is a controlled vocabulary it has not been specifically designed as a free text retrieval tool. however, its use within biron has seen an increase in the number of use and uf relationships. so although the study descriptions and dataset documentation have not been trawled for candidate terms, which are then structured into a thesaurus, the data archive and several external winter 1997 23 organisations are experimenting with using hasset as a retrieval tool for free text searching. the cessda (council for european social science data archives) idc (integrated data catalogue) is based on a z39.50wais protocol and uses freewais-sf and sfgate as its search engine and gateway. one of the options when creating an index for a wais database is to have present a synonym file, however since wais indexes every word, apart from stop words such as and, the etc., the synonyms have to be single words themselves. there is also no facility for the narrower / broader type relationships or control over when to apply the synonyms to a search. hence we have selected only single word terms with a use relationship to another single word term and single word top terms that have a rt relationship with other single word top terms. part of the nesstar project will be to investigate how hasset can be employed more fruitfully across the distributed databases of the european data archives and whether a multi-lingual version is a viable option. other organisations have also shown an interest in hasset, namely sosig (social science information gateway), midas (manchester information datasets and associated services), qualidata (qualitative data archival resource centre), the steinmetz archive for the eu-funded ilses project, ibss (international bibliography for the social sciences) and the office for national statistics in the uk. the most advanced of these is sosig who have a test interface on the www which they hope to incorporate into their search facility by june 1997. they have matched terms in the hasset thesaurus against keywords used in their own database records. by keying in a search string and selecting the ‘any related terms’ button and clicking on ‘do look up’ the present test interface will return the number of direct matches and also any term from the hasset relationships that are guaranteed to result in a match in sosig. 24 iassist quarterly the data archive hopes to undertake a project later this year where these participating organisations help convert hasset into a thesaurus resource for the whole of the social science community. control and maintenance of hasset will still remain the responsibility of the data archive, but the other organisations will offer up candidate terms and suggestion position in the hierarchies through a new www interface to hasset. as well the data archive will also review the contents and structure of hasset through analysis of the search logs from biron and a trawl of the study descriptions and recently digitised dataset documentation. it is hoped that this will also make hasset a valuable, universally available, freetext retrieval tool. * paper presented at iassist/ifdo ‘97, odense, denmark, may 6-9,1997. ken miller database programmer, the data archive, university of essex, england. vol264 4 iassist quarterly winter 2002 iassist quarterly winter 2002 5 edwardians online by emma j barker & louise corti* introduction in june 2002, the qualitative data service (qualidata) based at the uk data archive released edwardians online, a pilot, web-based, multimedia resource. the aim of this work is to develop a standard framework for digitizing data collections, and to provide online access to the content of digitized qualitative data collections. edwardians online integrates a wealth of existing primary and secondary materials relating to an oral history study carried out in the 1970s. researchers can perform free-text and thematic database searches of the interview summaries, as well as a sample of the full text transcripts. thematic searches are based on the existing coding schema originally used to classify and analyse the data. linked to this primary material are sound extracts from the audio recordings, images and contemporary photographs. further background material relating to the original research study, such as press reviews and details of publications based on secondary studies of the interview texts are also included. phase ii of the project aims to link other key sources, including maps and census data. 6 iassist quarterly winter 2002 iassist quarterly winter 2002 7 increasing access to qualitative data: digital resources qualidata provides a national service for the acquisition, dissemination and re-use of qualitative data that have been collected using social science research methods. the issue of how to make these data resources accessible to users is a central concern for the service. we are continually seeking ways to meet users ̓requirements. results of earlier work in this area can be seen in qualidataʼs resource discovery hub, where users can search and locate accessible collections of qualitative data across the uk via the online catalogue, qualicat. more recently, in response to user demand, emphasis has been on data development, with a view to providing users with direct access to the content of digitized collections via an online facility. the first steps in this content-oriented direction include an increased focus on depositing digital data in-house with the ukda, and the digitization of ʻclassic collections ̓ for research and teaching resources. edwardians online, based upon a set of oral history interviews, was selected as an appropriate collection for undertaking qualidataʼs first major web-based digitization project. the data collection the interviews were undertaken in the early 1970s as part of professor paul thompsonʼs study of edwardian society. these interviews form the basis for thompsonʼs, the edwardians: the remaking of british society, (1975, 1992). the 444 interviews were originally recorded on audio tapes and later transcribed as typed, paper documents. the original study materials were initially archived, catalogued and disseminated by qualidata. the importance of this collection for secondary use lies in the diversity and broad scope of the interview content and the scale of the collection. the collection has attracted high usage across a variety of research interests and is a valuable teaching resource. users have requested access to both complete interview transcripts and more specific information or extracts from within the documents. due to the length of the interviews, use of the collection can be time-consuming. for example, a typical transcribed interview may be 80 typed pages and an audio recording as long as four hours. phase i of the project a main aim of this project is to produce a prototype methodology which may be developed into a more general application for other examples of social science datasets. research on this collection has focused on the following key areas: • developing a non-proprietary electronic format for preserving the content of qualitative datasets. • developing tools for facilitating the encoding of data in this format • questioning the methods of access and facilities for exploring qualitative data online a standard framework for archiving digital qualitative data resources in pursuit of these issues, a comprehensive application appropriate for interchange that will enable sophisticated on-line searching and information retrieval from encoded texts is required. ideally, the application should meet a number of specific objectives that: • support the encoding of the content of various types of primary data documents produced in qualitative research. • support the encoding of contextual documentation and metadata linked to the primary sources • provide formalised links between the texts and associated audio and video materials, with a view to providing in the long term, integrated, multi media resources. • represent the content of datasets, such as the researcherʼs original analytic schema, annotations and speaker tags. a uniform format for encoding the content of datasets is useful for both data providers and users. it ensures consistency across datasets; supports the development of common publishing and search tools; and facilitates data interchange and comparison between datasets. development of an xml application for qualitative data finding a framework that will enable these functions leads us to consider xml standards and technologies. xml and related tools for creating and processing documents in xml have rapidly been adopted by communities of users for whom semantic tagging for their own application areas is essential. examples where xml tag sets are specially adapted to allow markup of the types of information specific to the user community include the data documentation initiative (ddi) for the social sciences and the text encoding initiative (tei). with increasing recognition of the benefits of xml in creating non-proprietary, cross-platform applications, there has been serious interest in, and calls for, the development of a qualitative data xml markup language from members of the social science research community who are eager to encourage the re-use of social science data. the development of a common framework for marking up the content of qualitative datasets requires support and contributions from various members of the social science community: data creators; qualitative data software 6 iassist quarterly winter 2002 iassist quarterly winter 2002 7 developers; data providers and end users. in particular agreements need to be made on: • types of documents and structures to be marked up • formal definition of a common xml vocabulary and dtd for describing these structures • specification of publishing and analysis tools • test applications with ʻreal ̓datasets edwardians online has aimed to provide the foundations for a broader initiative. research to date considered two options: first, to create a customized application of xml specifically for the purpose of marking up the content of spoken interviews and other types of qualitative material. second, to adapt existing standards, such as the tei and the ddi, thereby opening opportunities for using existing and forthcoming tools for processing the xml texts, in addition to the benefits of using a standard, such as detailed documentation and the expertise and experience of the previous user community. conclusion these ideas will be explored further in phase ii of the project. phase ii will focus on the development of additional search and retrieval functionality, and the encoding of additional features in the interview texts. the presentation of a document type definition (dtd) for a generalized xml application for qualitative datasets is a key milestone in this programme. over the next year we will begin work on adapting and integrating the tei and ddi to produce a prototype dtd for qualitative data. we hope this would become a de facto standard and one that could be used by other data creators and data publishers to encode a broad class of qualitative data. we are keen to encourage testing of the pilot resource. if you would like to complete a simple user evaluation exercise, please go to the web site and look for the evaluate tab. the resource can be found at: http://www.qualidata.e ssex.ac.uk/edwardians. for further details of ongoing and future work contact louise corti at the uk data archive. acknowledgement emma barker, the project officer, left the uk data archive in december. the ukda would like to extend its thanks to emma for the dedication she has shown to phase i of this project. references ddi http://www.icpsr.umich.edu/ddi/org/index.html qualidata, http://www.qualidata.essex.ac.uk/ national social policy and social change archive, http: //www.qualidata.essex.ac.uk/dataresources/nspsca.asp national sound archive, the british library, http://www.bl.uk/collections/sound-archive/ holdings.html#story tei www.tei-c.org/ thompson, p.r. (1975), the edwardians, the remaking of british societ. london: granada * paper presented at the iassist conference, june 2002, in storrs, ct, usa. louise corti is associate director & head, qualitative data service and outreach and training at the uk data archive, university of essex, email: cortl@essex.ac.uk mailto:cortl@essex.ac.uk 1/14 štebe, janez (2019) examining barriers for establishing a national data service, iassist quarterly 43(4), pp. 1-14. doi: https://doi.org/10.29173/iq960 examining barriers for establishing a national data service janez štebe1 abstract a system for monitoring the current situation of data archive services (das) maturity in european countries was developed during the cessda strengthening and widening in (saw 2016 and 2017) and further adapted in cessda widening activities 2018 (wa 2018) projects for continuous monitoring. an assessment of the existing national data sharing culture, the development of the social science sector and its production of high-quality research data, the funders’ research data policy requirements, and the capacity and skills of national grassroots initiatives, provide a framework for understanding the current situation in different countries. methods used in the projects, included desk research of existing documents and a survey, combined with extensive interviews focused on the area of expertise of the informants (individuals from data services, research and decision makers’ representatives from each country). the focus of the paper is the situation in 20 non-member cessda european countries with emerging and immature das initiatives. results show that countries are slowly but persistently removing the key obstacles in establishing a das initiative in their respective countries. the remaining obstacles reside mainly outside the control of the data professional community – namely research funders slowly adopt data sharing policies and incentives for data sharing, including the provision of a sustainable das infrastructure, capable of supporting researchers with publishing and accessing research data. the results show that the lack of expertise and skills of das initiatives, their understanding of tools and services or organizational settings are not such an issue, as more mature das are organising training and mentorship activities. detailed guidance in the das advocacy and planning was prepared in the framework of the above-mentioned pan-european and some past regional projects. the tools and framework of those activities will be referred to in the discussions as a resource that can be used in other countries and continents. keywords strategies for setting up data service, international partnership, data policies, cessda introduction establishing and running a national data archive service (das) can bring many benefits to the scientific community. the consortium of european social science data archives (cessda), as a distributed paneuropean social science data infrastructure, strives for a whole european research area (era) coverage, thus enabling equal opportunities for access to research data, regardless of researchers’ origin. cessda membership is country-based, and a signature from the responsible ministry needs to be obtained. each country nominates a country service provider, which is usually an individual institution/organization that is actively engaged in the national das provision. in order to document and support activities among the cessda partners non-member countries, a system for monitoring individual country situations was developed during the cessda strengthening and widening (saw 2015 – 2017)2 project. the effort was continued within the cessda widening activities (wa 2018)3 project, which established a system of continuous monitoring in order to capture and reflect the most current progress. the monitoring aims to address the problems of less mature https://doi.org/10.29173/iq960 http://cessdasaw.eu/about/ https://www.cessda.eu/about/projects/work-plans/work-plan-2018#wide 2/14 štebe, janez (2019) examining barriers for establishing a national data service, iassist quarterly 43(4), pp. 1-14. doi: https://doi.org/10.29173/iq960 das, and offers to search for solutions for the common problems and support for improvement. the aim is also to increase the visibility of individual country initiatives and organisations that may have spent long periods trying to establish a professional data service for their research community. in this paper, we will examine mainly the results of the most recent wa 2018 project, where one of the tasks was to continue the monitoring of the status of data archive services, led by the adp4. both projects contained a number of other activities, which we will refer to when discussing the results of the monitoring. problem setup the saw and wa 2018 projects developed an unique monitoring approach that examines a range of conditions of establishing and running a das in each of the countries by addressing the wider context of the data-sharing ecosystem. the approach starts with estimating the overall financial position of social sciences in a country, and considers the differences among countries regarding the demand for data sharing services. continuous national studies that produce high-quality data are important in this respect. next, the data sharing culture among the scientific community was estimated, both regarding the readiness to share one’s own data, and the widespread habit of using existing data whenever possible. a broad area of enablers and constraints influence a data sharing culture in the social sciences community. in particular, the research funders add yet another dimension to the ecosystem, by providing the policy framework with requirements for all regarding opening data, by stimulating the data management planning in order to maximize access to high-quality reusable data, and by providing incentives and a support environment to those that prepare and openly share data. finally, a mature ecosystem and an established culture of data sharing enable sustainable operation of a das, which in turn can support its parts to function well. it is the funders and the community of users and data depositors that influence the orientation of the das initiatives in individual countries and who can profit from its efficient functioning. in this respect, it is important that initiatives and small pilot das that arise based on well-developed professional grounds, actively engage with the user's community, and demonstrate and advocate with the decision makers about the justification of their activities. developed das can have a multiplying effect to support further development of social sciences, in particular with providing access to relevant high-quality data that tackles important societal issues. evidence was collected and the situation in each above-mentioned aspects was examined for the range of european countries. each of the aspects had a few pre-set questions, that were adapted during the in-depth interviews with selected national informants, and finally, a country report was written, emphasizing the main, either the positive factors or barriers and weaknesses in the system. the main focus was the das vitality and sustainability, both regarding internal organization and external conditions. stakeholders in individual countries or regions need to determine internal goals while comparing the current gaps in their countries with others and act correspondingly. a group of countries, identified to be at a similar level may consider following similar best practice examples to achieve a more mature and supportive open scientific data ecosystem. the results presented here can motivate in finding https://doi.org/10.29173/iq960 https://www.adp.fdv.uni-lj.si/ 3/14 štebe, janez (2019) examining barriers for establishing a national data service, iassist quarterly 43(4), pp. 1-14. doi: https://doi.org/10.29173/iq960 sustainable arrangements for a particular national situations, matching the interests of both the scientific community and the policy makers. previous studies as a baseline, some past studies were taken into account. the series of cessda widening projects were built upon the continuous experiences that arose from the unesco workshop on social science data archives in eastern europe in 2002 (hausstein and guchteneire, 2002). a group of more than 10 countries’ initiatives to establish the national das were engaged in informal cooperation under the edan the east european data archives network and coordinated by the gesis leibniz-institute for social sciences at that time. the serscida project (2012-14) was the first in a series of projects that provided full-fledged support and activities workflow for establishing a national das from scratch. four well-established cessda european das partnered in the project with the das initiatives from bosnia and hercegovina, croatia and serbia. results of the surveys by the serscida project in bosniaherzegovina, croatia and serbia show that in the absence of data infrastructure and support services, in practice, research data are mostly shared with colleagues and peers within the research group/institution, or not shared at all5, despite the fact that researchers may be very willing to share their research data with the wider scientific or civil community. within the serscida project, working visits to well-established partners were organised6, training modules were delivered for archiving professionals, draft archive policies and business plan documents were developed and prototype webpages were established. in the following years, a series of projects followed with similar aims and approach to serscida, including those mentioned in the introduction, which produced the manuals and guides and knowledge sharing materials that have the potential of wider relevance for the archiving community7. all the projects and initiatives mentioned provided some overview of the conditions and capacities of different organisations, residing in the national contexts. a comprehensive overview of the national open social science data policies was provided for the first time in the ifdo report from 2014 (kvalheim and kvamme, 2014). chuck humphrey’s analysis (humphrey, 2003) of the profiles and organisational settings that collected information of das worldwide, aiming at the proposal of how to establish a national das in canada, started with a similar assumption as we do in the country reports: a comprehensive set of conditions needs to be explored in each national setting, which will in turn help to shape further development steps. that is, there is no development model that fits all. as an inspiration about how to approach describing the complex situation regarding das in the variety of european countries, some of them in the initial stage of considering how to start activities used a metaphor of the data sharing ecosystem: ’it is a complex system involving data collectors, stewards, and users as well as sponsors and stakeholders; emergent and historical transparent technologies; and ever-growing data along with their myriad associated artefacts. the system must be understood in totality in order to optimize the whole and not just the individual components.´ (parsons et al., 2011, p. 557). data sharing culture is a key systemic component that determines the efficiency and sustainability of a data ecosystem. much research has been done in the last decade across disciplines and at an international level on research data sharing culture, data sharing and management practices, on https://doi.org/10.29173/iq960 4/14 štebe, janez (2019) examining barriers for establishing a national data service, iassist quarterly 43(4), pp. 1-14. doi: https://doi.org/10.29173/iq960 barriers and enablers. this published literature provides us with much information on this topic, which is most likely applicable across countries. both detailed qualitative or mixed studies (borgman 2012, eagda, 2014, van den eynden et al., 2014) and comprehensive surveys (sveinsdottir et al., 2013, van den eynden et al., 2016) assessing data sharing practices, barriers and enablers amongst researchers at a local, european or international level some of which focus on specific research disciplines, others look across a range of disciplines – identify numerous perceived or real barriers to data sharing, such as the lack of standards and data infrastructure, the fear of competition, the costs and the absence of rewards to prepare data and documentation, amongst others. among enablers of data sharing generally reported in the literature, we could stress the data sharing expectations of funders, institutions and journals (tsoukala et al., 2016), the established habits in the research community, and the areas where das support is visible, like professional training on data management skills, and enabling data publication for citation. method results and reports from the previous widening activities (serscida, seeds) and in particular, the most recent saw project reports (cessda saw, 2017a,b) were taken into account while assessing the national situation in 2018, with an emphasis on change and progress being made nationally. the mapping contains a review of the elements identified in the introduction of the wider data-sharing ecosystem: the interplay of the structural conditions of social science development, the funders open data policies and strategies, and the data sharing culture and the incentives that increase the data sharing habits of researchers. external stakeholders, in particular funders, can play an important role in improving the national data service sustainability. research funders are the key stakeholders that can help to provide incentives and remove some of the barriers to data sharing. advanced policy recommendations, appropriate funding mechanisms and a strong das can lead to a sustainable datasharing ecosystem. finally, the countries where no formal das exist were analysed regarding the potentials of integration of initial rdm support infrastructure. by identifying proto-activities and open access support activities, we detected actors and institutions that could play a key role in the elaboration of a new national das. the list might be of help to funders and to the cessda main office on a national and an international level when planning further development. a monitoring system has been established consisting of the following steps: step 1: desk research to consult the existing sources of information; step 2: selection of contact(s); step 3: tailoring semistructured interview (country and stakeholder-specific); step 4: contact and carry out the interview, either orally or written. the project group members utilised the guidelines and communication protocol for interviewers that contain a list of suggested interview questions and issues to be addressed, arranged along the content areas (see cessda wa, 2019, appendix 1). monitoring has been based on regular short interviews of cessda partners and other contacts established in non-member countries. the country reports on recent developments summed up information from the previous reports, and the monitoring interview. the project group decided to approach 22 countries among all european research area cessda nonmember countries to monitor in 2018. those that had at least one possible productive contact identified in previous rounds of activities among the relevant stakeholders (policy makers, research https://doi.org/10.29173/iq960 5/14 štebe, janez (2019) examining barriers for establishing a national data service, iassist quarterly 43(4), pp. 1-14. doi: https://doi.org/10.29173/iq960 data expert or similar). a list of countries was determined and distributed between 5 partners for this task in june 2018. the contact info table from the cessda saw project was updated with recent contacts. reporting on the contacts made during the interviewing period was filled in by partners. following the communication protocol, including the initial contact e-mail template, the majority of the requests for interviews were sent between october and november 2018 and realised soon after. the last interviews were conducted in january 2019. map 1: cessda member and partner countries, as of mid-20188 the most recent reports sum up past and existing information gathered on different occasions for each individual country, and add up to what is new from the fresh interviews. results most of the interviewees were social science data researchers or data librarians involved in the organisation of the national data archives service. https://doi.org/10.29173/iq960 6/14 štebe, janez (2019) examining barriers for establishing a national data service, iassist quarterly 43(4), pp. 1-14. doi: https://doi.org/10.29173/iq960 some of the services have a longer time span, having been members of the old cessda (pre-european research infrastructure consortium era). some went through an institutional change, like in ireland and italy, where the seat of the das had moved compared to the old cessda. some have sustained a low level of activity for a longer period of time, such as estonia, poland, slovakia and romania, without being able to make a breakthrough to achieve the status of a national service provider for cessda. the latter is due to not having a ministry ready to sign the agreement to join cessda and to sustainably support the national das. from the recent era, the notable progress of the das initiatives in many countries results from their participation in the various widening projects mentioned before. these projects were important for engaging with stakeholders in countries (funders, researchers, etc.) and for developing the professional competence of the people and organisations involved. the projects helped to create wellelaborated plans for the das in serbia, croatia, and macedonia, which have recently become new cessda member states. development of the social sciences sector in the country the focus of the first part of the interviews was on funding capacities, human resources and infrastructure, international collaboration and national studies as a driver of das demand in the country. most reports that came from economically less developed countries share an impression of the low status of social sciences research, which leads to the generally low supply of high-quality key research data resources and the weak policy and funding support for data service activities. yet the creation of a data archiving service may influence the production of higher quality data in the future, by raising awareness of the importance of data sharing and by pooling the resources around fewer new data-collecting projects due to the wider reuse of existing data. these were among the justifications given in the national reports in favour of establishing the das in less developed countries like albania and kosovo. other perceived benefits include the fact that important national data can be preserved and used for longitudinal studies, such as evidence from ex-ante policy evaluations. secondary data can bring value for teaching and training. those arguments have been part of the reports and national development plans of some other countries, as was the case in croatia and in ireland. rdm policy and support setting one of the key elements of a data sharing ecosystem is national funders’ policies for data documentation and management, facilitating data sharing and ethical and legal frameworks. a clear research data policy in a country prepares a space for an existing and emerging das to function more efficiently. a mature policy is exemplified by interview questions about the requirement to prepare a data management plan (dmp), a recommendation about appropriate place of deposit, selection of data based on quality and reuse potential for long-term curation, and the importance of legal and ethical guidelines to attain clarity on the legal conditions framing the envisaged re-use of research data. the eu commission has been active in setting the open science agenda for its members. the open science policy platform contains references to european open science cloud (eosc), fair data and https://doi.org/10.29173/iq960 http://ec.europa.eu/research/openscience/index.cfm?pg=open-science-cloud http://ec.europa.eu/research/openscience/index.cfm?pg=open-science-cloud http://ec.europa.eu/research/openscience/index.cfm?pg=open-science-cloud http://ec.europa.eu/research/openscience/index.cfm?pg=open-science-cloud http://ec.europa.eu/research/openscience/index.cfm?pg=open-science-cloud http://ec.europa.eu/research/openscience/index.cfm?pg=open-science-cloud http://ec.europa.eu/research/openscience/index.cfm?pg=open-science-cloud 7/14 štebe, janez (2019) examining barriers for establishing a national data service, iassist quarterly 43(4), pp. 1-14. doi: https://doi.org/10.29173/iq960 other initiatives. with the launch of the open research data pilot in horizon 2020 projects, the eosc, the adoption of the digital single markets strategy and the new directive on open data and public sector information, the rdm strategies and implementations are becoming an important factor in the research infrastructure development in countries where no formal das exists. gradual implementation of the eu recommendations can be seen in the eu commission report about current national open science policy activities (dgri, 2018). in the interviews, the topic was addressed from two angles. following from the top-down requirements of the national research data policy aligned with the eu recommendations, the question was about how to support the bottom-up implementation. a notable example arises from croatia, where the croatian initiative for the establishment of the social science data archive referred to a positive acceptance of the rdm training activities offered by them in different regions of the country. they have further speculated that the croatian science foundation may extend the existing requirements for keeping research data from humanities to social sciences as well and that the future cessda service provider (sp) from the country can help to fulfil the requirements by offering a full range of data support services. similarly, in iceland, parallel to the three-year funding to build the das planned in 2019 at the social science research institute, the country is in a process of articulating the first icelandic research infrastructure roadmap and a national policy for open access to data. to conclude, we can observe the discrepancy between the sometimes-isolated policy plans in the countries, as visible in the eu reports about the national settings, and the reality of the low level of data sharing practices. a national das, that follows the cessda overall mission, can fill in that gap, including supporting data citation, and other incentives to the science community. data sharing culture attitudes, perceived barriers and incentives related to data sharing and rdm support and practices was the next topic of the interviews. though data sharing culture is hard to assess objectively just by interviewing one or few informants from a country, the approach was to have informed country experts estimate, for example, the willingness of researchers to share data and the channels they use, and the experiences of researchers trying to obtain data when needed. some of the questions addressed the rewards and career progression that someone could anticipate if they are active in open data sharing practices. the answers reflect a generally low level of systemic data sharing and little or no awareness about best practices in managing and documenting data for reuse. in one of the interviews, an informant from macedonia stated that currently ‘(…)this is done on an individual basis by the involved researchers.’ countries with active initiatives for establishing the das can already show their visibility in their national setting, with researchers starting to think about data sharing from the project’s beginning, and seeking collaboration with the das initiative. as reported in croatia, the collaboration agreement with one of the projects ‘will consist of preparation of the dmp, and final version of data for submission to the das service, with the purpose to further distribute data and promote its usage.’ https://doi.org/10.29173/iq960 https://www.openaire.eu/opendatapilot https://www.openaire.eu/opendatapilot https://www.openaire.eu/opendatapilot https://www.openaire.eu/opendatapilot https://www.openaire.eu/opendatapilot https://www.openaire.eu/opendatapilot https://www.openaire.eu/opendatapilot https://ec.europa.eu/programmes/horizon2020/en https://ec.europa.eu/programmes/horizon2020/en http://ec.europa.eu/priorities/digital-single-market_en http://ec.europa.eu/priorities/digital-single-market_en http://ec.europa.eu/priorities/digital-single-market_en http://ec.europa.eu/priorities/digital-single-market_en http://ec.europa.eu/priorities/digital-single-market_en 8/14 štebe, janez (2019) examining barriers for establishing a national data service, iassist quarterly 43(4), pp. 1-14. doi: https://doi.org/10.29173/iq960 data infrastructure the assessment of data archive proto-activities in countries where no formal cessda membership exists was the main part of the recent cessda widening 2018 reports. the focus of the description of the situation for the countries that do not have a national das was put on exploring the conditions for establishing a data service that could in the future obtain the role of a cessda national service provider. these were labelled as das proto-activities. some of the european countries have a long tradition of research data management (rdm) and data archiving in social sciences, while others are at the very beginning. the overview of the profile and the organisational infrastructure shows that most organisations provide a publically available mission that clearly declares that it carries out the main required functions of a typical data archive. while the ambitions and potential for delivering a fully flagged das service are common to all types of organisations, the analysis also shows a substantial variation in some of the aspects of the maturity of self-assessments among the ‘aspiring’ members. countries should provide long-term funds for the establishment and functioning of the das, in order to be able to fulfil the mission clearly stated in the documents. sustainability of the das or the das initiative is the main problem in most of the countries considered in the reports. the main factor in sustainability is the lack of resources and the related absence of political support. there are groups of countries that follow a similar path and have similar problems. in the first group are countries that are only at the beginning of their activities, and have made no firm decision about how to organise the das. some institutions from those countries participated in some of the past cessda projects and were identified as potential partners, yet they do not show any recent activity on the institutional level, little or no advocacy or involvement in projects, and little or no political support. for some of those, no contacts could be established or there was no productive interview obtained in 2019, even though some of them participated in previous rounds of country reporting activities. the potential for future activities in existing human resources, technological infrastructures and support services (libraries, research institutes, and research information services) were areas addressed in the das proto-activities part of the assessments of the countries with no existing data infrastructure. these are countries like cyprus, where the country report concludes with an observation: ‘currently, there is no institutional and technical infrastructure for data deposit in ssh in cyprus. also, official steps or public initiatives towards establishing a das for the social sciences were not identified.’ the kosovo das initiative representatives participated in the seeds project and continued their participation in the saw, where areas of activities were strategically planned. yet, upon requesting an update, the representatives explained, that currently, they ‘do not have any further comment to the state of the art and that they hope to find the support to boost the archive in kosovo in the future’. a similar situation has been found in montenegro and albania, even though both collaborated in writing the national development plans in previous rounds (cessda saw, 2017b). a subgroup of the countries with no currently active initiative for establishing a das identified are spain and luxemburg, both of which had active social science data archives in the previous decades. the next group consists of countries with a long established but very basic level of das service. the das have been running for some time on minimal financial and human resources. usually, they are represented by a single person with little or no institutional support, seeking to reach the funder and decision-makers to help establish a more robust framework for a functioning das. the situation in https://doi.org/10.29173/iq960 9/14 štebe, janez (2019) examining barriers for establishing a national data service, iassist quarterly 43(4), pp. 1-14. doi: https://doi.org/10.29173/iq960 poland is typical. poland has the polish social data archive, ads9, led by marcin zieliński for more than 20 years, sadly observing that despite a long tradition of making social research in poland, ’most of the data have been already lost because of the lack of financial sources to preserve them and structural possibilities of long time preservation.’ estonia, romania, and latvia have all been long established but unable to reach sufficient funding for a continuously running, mature das and lacking political support for the cessda country membership. for those countries, the expertise has already been acquired in past years, by keeping in contact with the cessda experts’ community and participating in some of the projects. one of the most important ones was the saw project, that in addition to the country’s overall monitoring and planning, as already mentioned, also offered an introductory core trust seal (cts) training that experts from those countries attended (cessda saw, 2017c). voluntary self-assessment regarding the criteria of the cts and organisational maturity self-assessment aimed for the country report was useful, as it actually demonstrated the lack of a sustainability component in its core. as observed, the das in those countries run without regular staff members, mainly on a voluntary basis of the involved individuals or as an in-kind contribution of host institutions. finally, there is a group of countries that are reaching the sufficient level of support from ministries and funders, while also showing a range of activities in improving the das. among them are countries that were involved in all cessda widening projects, like croatia and serbia. both became cessda members early in 2019. those two countries started from scratch with the individuals who were involved in the series of widening projects, acquiring the necessary skills and competences and following the suggested advocacy and planning strategies. it took more than five years since the beginning of the first in a series, the serscida project, to reach success, and this is one of the biggest lessons learned persistence eventually pays off. macedonia reached membership status later this year, and bosnia and hercegovina, and bulgaria, among the balkan countries with no legacy of a das, have good prospects to gain support both nationally and institutionally. organisations in italy, ireland, iceland, and slovakia are active in extending their already established institutional services to nationwide services or consortia and are trying to gain ministry support for cessda membership. belgium is unique within this group, as the country already obtained cessda membership status, parallel to the on-going soda prototype project (social sciences data archive) which aims to set up a data archive in belgium once again, as there was in the past. common to all actively establishing organisations is detailed investigation about the technical and legal issues, in particular about the types of licenses and agreements with users, and the business model of the future. russia is a special case regarding cessda membership, since it is a non-eu country and therefore its inclusion probably demands a specifically developed legal framework. the joint economic and social data archive (jesda) has been running since the year 2000 with ambitions to actively contribute to a data sharing culture in the country by providing an extended training programme. the ukrainian national data bank of sociological data “kyiv archive” as the national service provider for ssh data is in a similar position. the staff from both organisations attend events organised by the cessda lead projects, and actively contribute to the activities, including providing reports about the current situation regarding the das service and its environment. https://doi.org/10.29173/iq960 10/14 štebe, janez (2019) examining barriers for establishing a national data service, iassist quarterly 43(4), pp. 1-14. doi: https://doi.org/10.29173/iq960 discussion visibility proto sps gain their visibility in the national setting by being included in one or more cessda widening project activities. the cessda widening projects, following from the state-of-the-art national reports, provide a framework for ‘proof of concept’ of the the prototype das. national development plans and media packs are being produced in consecutive projects, with the guidance on the general oais structure, actively adapted to specific national settings. partners from institutions that collaborated in the widening projects that have either been working for some time already on the project of establishing a das, or are just starting, through their involvement in the projects catalyse the national discussion and help generate the network among different stakeholders. thus, one important side product of the widening activities, both while preparing national reports and development plans, has been an updated contact list of the people in a country who were consulted in various phases of the project activities. most importantly, when the decision makers and funders’ representatives were involved in the national development plans, they added to the realistic planning and confirmation of the key aspects, including expectations regarding financial and political decisions about joining cessda and fulfilling their membership obligations. the most effective way to gain a momentum of visibility was the opportunity for the new partners to host some of the planned conferences, meetings and workshops of the widening projects in their respective countries with financial support from the cessda projects. this was usually seen as an opportunity for the candidate national cessda sp to demonstrate to the national stakeholders their involvement in the professional community. the first day of a typical widening workshop was composed of an introductory session, where both national decision makers and an institution’s leadership representatives were present, together with similarly profiled invited stakeholders from other countries. the aim was to present and exchange experiences both nationally and in european space10. national development plans (ndp) ndp addressed the series of challenges that a new das has to deal with (cessda saw, 2017b). it starts with a mission and designate community statements. it covers the preservation policy, collection plan and organizational setting, including its board role, the staff composition and financial resources. the tools and services are described, either as a decision already taken or as a topic that needs to be dealt with in the next steps. the last round of the national reports considered the ndp fulfilment. what it showed is that the more detailed and specific to the circumstances of the organisation they are, the more effective they have been in their actual fulfilment. among the countries that provided the ndp drafts during the saw project, kosovo, montenegro and albania didn’t succed in reporting on followup, however, all other countries are showing substantial progress. some of the partners were actively involved in the parallel cts training and provisional selfassessments that were organized as the activity of the dedicated cessda trust projects and continue to provide support for both members and non-members regarding gaining the cts. both ndp and cts evaluation and planning processes follow the approach where guidance contains the range of options, and the organisations themselves realistically decide to choose the processing level and data complexity they feel capable to control, reflecting also on the national traditions of supply and demand for data. thus the policy and organizational framework planning is structured around the decision to deal with anonymized data only, or to cover more complex services of secure access to sensitive personal data as well, including for example qualitative data. most of the partners are engaged with different stakeholders to reach a more stable position, as seen from different projects https://doi.org/10.29173/iq960 11/14 štebe, janez (2019) examining barriers for establishing a national data service, iassist quarterly 43(4), pp. 1-14. doi: https://doi.org/10.29173/iq960 and activities reported on a country level. cessda main office has an active role in those countries, providing support letters or visits to the ministries’ representatives. there is always an opportunity for experts from those institutions to collaborate in the cessda lead projects, and to participate in some of the workshops and training, which helps to sustain the professional contacts and to keep some minimal human resources involved. experts from different countries are seeking support in different areas that are covered in the resource directory, in particular looking for new shared tools for running a das. one of the areas identified during the saw project is the customisation of dataverse to the european setting with a simple docker installation, available for new or aspiring partners. the cessda saw and the wa 2018 projects prepared support packs for partners that address some of the problems identified. the wa 2018 project offered a resource directory (cessda wa, 2018a) that contains references to documents, training, tools and support services from the past cessda projects and from the sp’s, addressing and analysing the needs of the partners that are expected to utilise the resources. the resources can help new partners build a professional basis in data archiving. the original resources from different sps can also help the existing partners upgrade their services, for example in drafting the forms and agreements used in the das in relation to data deposit and access. besides those, the resources that partners demand the most arise from the areas of technical infrastructure and funding and advocating for the das (cessda wa, 2018a), including references to the eu open science initiatives, and more clearly described criteria that the cessda members sp needs to fulfil.11 in the new wa 2019 project the support extends to a mentorship programme, within which interested organisations with an ambition to improve on certain areas may apply to. the mentorship is delivered by the participating cessda sps. experts that are new to the cessda community and participated in the wa 2018 workshops and activities, expressed some concerns that the support information and resources are dispersed. the expert guide on data archiving can fill that gap, customised to reflect the european das12. conclusions parts of the conditions that affect the research data infrastructure concern the financial and institutional statuses of social sciences in each country. in some of the current cessda membership countries, we can find excellence in the development of social science that is supported with a robust and multifaceted research data service. what we often encounter at the other end is a syndrome of underdevelopment, where lack of funds affects every other aspect of the science system, including the infrastructure. where social sciences have a low budget in general, there are usually also poor conditions for a data infrastructure. the impact of gradually establishing a robust data infrastructure in that case can have an even greater impact on building a data sharing culture and improving the excellence and the efficiency of research in general. if data is shared widely, it will have immediate effect through improved quality control and transparency of the research. focus for the countries that do not have a national das was put on exploring the conditions for establishing a data service that could in the future obtain the role of a cessda national sp. these conditions are to a large extent contained in the areas addressed in a wider data sharing ecosystem, such as the financial and structural conditions of the social sciences sectors, the scientific policy https://doi.org/10.29173/iq960 12/14 štebe, janez (2019) examining barriers for establishing a national data service, iassist quarterly 43(4), pp. 1-14. doi: https://doi.org/10.29173/iq960 requirements and norms established in the scientific community. these external stakeholders play their role in influencing the existence of a das and constitute a data sharing cultural environment. there are internal stakeholders that are capable and willing to play a role in the establishment of new future services, and can bring current services to a higher maturity level. for the countries that do not have a running das, it is essential to ground the establishment of its services primarily on internal resources, which means finding the potential for future activities in existing human resources and organisational settings. the key factor in the slow but persistent growth of the cessda membership are enthusiastic and eager individuals that internalize as their mission the formation of a national das. such individuals, at the beginning mainly supported by their home institutions (usually universities or research institutes), communicate and help to articulate the needs of their scientific and academic communities, by demonstrating the advantage of establishing a das. the lobbying activity with funders and decision makers eventually brings a change. the last part is probably the toughest, as evidenced by the many failed missions reported. the ministry staff is prone to fluctuations (frequently in some of the countries with a less stable political situation), and this may cause already obtained agreements fail. the aims of the various widening projects are not fully accomplished, since full european coverage has not yet been obtained. new countries are joining or are about to join, which shows that the cessda widening projects add to catalysing the establishment of data services. the cessda, its national service providers and partners seek to continue their widening activities in some of the current and future projects13. references borgman, c. l. (2012) ’the conundrum of sharing research data’, journal of the american society for information. science and technology 63: 1059-1078. http://dx.doi.org/10.1002/asi.22634; cessda saw. cessda strengthening and widening (2017a) 'deliverable 3.2 country report on development potentials', consortium of european social science data archives, (available at http://cessdasaw.eu/content/uploads/2017/11/d3.2_cessda_saw_v1.3.pdf) cessda saw. cessda strengthening and widening (2017b) ’deliverable 3.4 national development plans for data services in non-cessda member countries in the era', consortium of european social science data archives, (available at http://cessdasaw.eu/content/uploads/2017/11/d3.4_cessda_saw_v1.0.pdf) cessda saw. cessda strengthening and widening (2017c) ’d4.4: report on dsa certification for cessda’, consortium of european social science data archives. http://dx.doi.org/10.18448/16.0069 cessda wa. cessda widening activities 2018 (2018a) ’deliverable 1 – cessda resource directory’, consortium of european social science data archives, (available at https://www.cessda.eu/toolsservices/for-service-providers/resource-directory) cessda wa. cessda widening activities 2018 (2018b) ’deliverable 5 – gap analysis of cessda resources, consortium of european social science data archives, (available at https://www.cessda.eu/about/projects/work-plans/work-plan-2018) https://doi.org/10.29173/iq960 http://dx.doi.org/10.1002/asi.22634 http://cessdasaw.eu/content/uploads/2017/11/d3.2_cessda_saw_v1.3.pdf http://cessdasaw.eu/content/uploads/2017/11/d3.4_cessda_saw_v1.0.pdf http://dx.doi.org/10.18448/16.0069 https://www.cessda.eu/tools-services/for-service-providers/resource-directory https://www.cessda.eu/tools-services/for-service-providers/resource-directory https://www.cessda.eu/about/projects/work-plans/work-plan-2018 13/14 štebe, janez (2019) examining barriers for establishing a national data service, iassist quarterly 43(4), pp. 1-14. doi: https://doi.org/10.29173/iq960 cessda wa. cessda widening activities 2018 (2019) ’deliverable 2: system of monitoring of the state-of-play. consortium of european social science data archives (v1.0), https://doi.org/10.5281/zenodo.3474048. dgri. directorate-general for research and innovation (2018) ‘access to and preservation of scientific information in europe. report on the implementation of commission recommendation c(2012) 4890 final – study’, european commission, https://dx.doi.org/10.2777/642887 eagda. expert advisory group on data access (2014) ‘establishing incentives and changing cultures to support data access’, (available at https://wellcome.figshare.com/articles/incentives_to_support_data_access/5613199) hausstein, b. and p. de guchteneire, eds. (2002) ‘social science data archives in eastern europe : results, potentials and prospects of the archival development’, workshop on social science data archives in eastern europe, berlin. humphrey, c. (2003) 'models of data archiving services. the results of and international survey' iassist 2003-strength in numbers, ottawa, canada, (available at https://iassistdata.org/downloads/2003/g1_humphrey.pdf) kvalheim, v. and t. kvamme (2014) ‘policies for sharing research data in social sciences and humanities. a survey about research funders’ data policies’, (available at http://ifdo.org/wordpress/wpcontent/uploads/2015/07/ifdo_survey_report.pdf ) parsons, m. a. and ø. godøy, e. ledrew, t. f. de bruin, b. danis, s. tomlinson, d. carlson (2011) ‘a conceptual framework for managing very diverse data for complex, interdisciplinary science’, journal of information science 37(6) 555–569, https://doi.org/10.1177/0165551511412705 sveinsdottir, t., wessels, b., smallwood, r., linde, p., kala, v., tsoukala, v., and sondervan, j. (2013, september 30). stakeholder values and ecoystems. zenodo. http://doi.org/10.5281/zenodo.835772 tsoukala, v., angelaki, m., kalaitzi, v., wessels, b., price, l., taylor, m. j., … wadhwa, k. (2016). recode: policy recommendations for open access to research data. zenodo. http://doi.org/10.5281/zenodo.50863 van den eynden, v. and bishop, l. (2014) ’sowing the seed: incentives and motivations for sharing research data, a researcher’s perspective’. a knowledge exchange report, (available at http://www.knowledge-exchange.info/projects/project/research-data/sowing-the-seed) van den eynden, v. and g. knight, a. vlad, b. radler, c.l tenopir, d. leon, f. manista, j. whitworth, l. corti (2016) ‘towards open research: practices, experiences, barriers and opportunities’, wellcome trust, https://dx.doi.org/10.6084/m9.figshare.4055448 https://doi.org/10.29173/iq960 https://doi.org/10.5281/zenodo.3474048 https://publications.europa.eu/en/publication-detail?p_p_id=portal2012documentdetail_war_portal2012portlet&p_p_lifecycle=1&p_p_state=normal&p_p_mode=view&p_p_col_id=maincontentarea&p_p_col_count=3&_portal2012documentdetail_war_portal2012portlet_javax.portlet.action=author&facet.author=rtd&language=en&facet.collection=eupub https://publications.europa.eu/en/publication-detail?p_p_id=portal2012documentdetail_war_portal2012portlet&p_p_lifecycle=1&p_p_state=normal&p_p_mode=view&p_p_col_id=maincontentarea&p_p_col_count=3&_portal2012documentdetail_war_portal2012portlet_javax.portlet.action=author&facet.author=com,ecfin,taskf,oil,giw,oib,repres_nld,repres_lva,jls,erc,markt,mare,regio,rea,bepa,press,bds,elarg,pmo,repres_lit,agri,repres_spa_bcn,spp,echo,eaph,repres_gbr_lon,repres_est,fpi,repres_spa_mad,casstm,cnect,digit,home,ener,repres_hun,ieea,easme,comp,repres_cze,repres_bgr,scr,repres_mlt,repres_prt,repres_cyp,repres_hrv,clima,eahc,repres_swe,repres_svn,del_acc,infso,eaci,ethi,dg18,dg15,dg10,chafea,repres_deu_muc,repres_pol_waw,estat,devco,dgt,epsc,grow,sante,near,fisma,just,com_cab,scad,repres_gbr,repres_pol,taskf_a50_uk,repres_spa,repres_fra,repres_ita,acshhpw,pc_budg,iab,rsb,pc_conj,com_coll,acsh,evhac,pc_mte,repres_deu,repres_svk,justi,repres_deu_bon,scic,repres_fra_par,sj,sg,repres_pol_wro,olaf,repres_deu_ber,ccss,fsu,repres_irl,hr,repres_lux,repres_fin,taxud,commu,sanco,entr,audit,igs,repres_ita_mil,move,budg,repres_rou,rtd,ias,btl,tentea,btb,cmt_empl,dg01b,dg01a,repres_bel,repres_gbr_cdf,env,dg23,dg17,dg07,dg03,dg02,dg01,repres_aut,inea,empl,eac,trade,tren,repres_ita_rom,relex,aidco,repres_grc,eacea,repres_gbr_bel,repres_fra_mrs,repres_gbr_edi,repres_dan,jrc,dev,srss,has,stecf,dpo&language=en&facet.collection=eupub https://dx.doi.org/10.2777/642887 https://wellcome.figshare.com/articles/incentives_to_support_data_access/5613199 https://unesdoc.unesco.org/query?q=conference:%20%22workshop%20on%20social%20science%20data%20archives%20in%20eastern%20europe,%20berlin,%202002%22&sf=sf:* https://unesdoc.unesco.org/query?q=conference:%20%22workshop%20on%20social%20science%20data%20archives%20in%20eastern%20europe,%20berlin,%202002%22&sf=sf:* https://iassistdata.org/downloads/2003/g1_humphrey.pdf http://ifdo.org/wordpress/wpcontent/uploads/2015/07/ifdo_survey_report.pdf https://journals.sagepub.com/action/dosearch?target=default&contribauthorstored=parsons%2c+mark+a https://journals.sagepub.com/action/dosearch?target=default&contribauthorstored=god%c3%b8y%2c+%c3%98ystein https://journals.sagepub.com/action/dosearch?target=default&contribauthorstored=ledrew%2c+ellsworth https://journals.sagepub.com/action/dosearch?target=default&contribauthorstored=de+bruin%2c+taco+f https://journals.sagepub.com/action/dosearch?target=default&contribauthorstored=danis%2c+bruno https://journals.sagepub.com/action/dosearch?target=default&contribauthorstored=tomlinson%2c+scott https://journals.sagepub.com/action/dosearch?target=default&contribauthorstored=carlson%2c+david https://doi.org/10.1177%2f0165551511412705 http://doi.org/10.5281/zenodo.835772 http://doi.org/10.5281/zenodo.50863 http://www.knowledge-exchange.info/projects/project/research-data/sowing-the-seed https://dx.doi.org/10.6084/m9.figshare.4055448 14/14 štebe, janez (2019) examining barriers for establishing a national data service, iassist quarterly 43(4), pp. 1-14. doi: https://doi.org/10.29173/iq960 end-notes 1 janez štebe is head of adp, arhiv družboslovnih podatkov [social science data archives], university of ljubljana, slovenia. he is an associate professor of methodology at the faculty of social science, university of ljubljana: janez.stebe@fdv.uni-lj.si 2 strengthening and widening the european infrastructure for social science data archives project funded by the eu horizon 2020 research and innovation programme under the agreement no.674939. http://cessdasaw.eu/about/ 3 cessda widening activities 2018, project under the call cessda work plan tasks 2018 https://www.cessda.eu/about/projects/work-plans/work-plan-2018#wide 4 arhiv družboslovnih podatkov [social science data archives, university of ljubljana, slovenia]. fors – swiss foundation for research in social sciences, čsda – czech social science data archive, snd – swedish national data service and tarki – foundation as project partners in wa 2018 contributed to the task. 5 see serscida project deliverables: country maping reports for bosnia and herzegovina, croatia and serbia. project web: http://www.serscida.eu/en/; deliverables: http://www.serscida.eu/en/deliverables 6 e.g. see https://www.adp.fdv.uni-lj.si/serscida_wp4_lj2013_working_visit/ 7 including seeds (south-eastern european data services 2015-17; http://seedsproject.ch) project supported by swiss national fonds; saw, wa 2018 and 2019. 8 update on current status of both members and partners with contact information of service providers is available at https://www.cessda.eu/about/consortium/cessda-countries/. 9 polish social data archive, ads: http://www.ads.org.pl 10 see for example strengthening and widening of the european infrastructure of social science data archives wa 2018 workshop in milano, italy: https://www.cessda.eu/widening2018/, and belgrade, serbia: https://www.cessda.eu/belgrade2018/ 11 compare: https://www.cessda.eu/content/download/733/6532/file/wittenberg.pdf. 12 authors opinion is that curating research data, volume two: a handbook of current practice could serve as approximation for that purpuse. 13 see service providers’ tools & services on https://www.cessda.eu/tools-services/for-serviceproviders. https://doi.org/10.29173/iq960 mailto:janez.stebe@fdv.uni-lj.si http://cessdasaw.eu/about/ https://www.cessda.eu/about/projects/work-plans/work-plan-2018#wide http://www.serscida.eu/en/ http://www.serscida.eu/en/deliverables https://www.adp.fdv.uni-lj.si/serscida_wp4_lj2013_working_visit/ http://seedsproject.ch/ https://www.cessda.eu/about/consortium/cessda-countries/ http://www.ads.org.pl/ https://www.cessda.eu/widening2018/ https://www.cessda.eu/belgrade2018/ https://www.cessda.eu/content/download/733/6532/file/wittenberg.pdf https://www.cessda.eu/tools-services/for-service-providers https://www.cessda.eu/tools-services/for-service-providers 1/3 rasmussen, karsten boye (2023) editor’s notes: fair bot. as metadata is data is metadata is data ..., iassist quarterly 47(1), pp. 12. doi https://doi.org/10.29173/iq1086 editor's notes: fair bot. as metadata is data is metadata is data ... welcome to the first issue of iassist quarterly for the year 2023 iq vol. 47(1). the last article in this issue has in the title the fair acronym that stands for findable, accessible, interoperable, and reusable. these are the concepts most often focused on by our articles in the iq and fair has an extra emphasis in this issue. the first article introduces and demonstrates a shared vocabulary for data points where the need arose after confusions about data and metadata. basically, i find that the most valuable virtue of well-structured data – i deliberately use a fuzzy term to save you from long excursions here in the editor's notes – is that other well-structured data can benefit from use of the same software. similarly, well-structured metadata can benefit from the same software. i also see this as the driver for the second article, on time series data and description. sometimes, the software mentioned is the same software in both instances as metadata is treated as data or vice versa. this allows for new levels of data-driven machine actions. these days universities are busy investigating and discussing the latest chatbots. i find many of the approaches restrictive and prefer to support the inclusive ones. likewise, i also expect and look forward to bots having great relevance for the future implementation of fair principles. the first article is on data and metadata by george alter, flavio rizzolo, and kathi schleidt and has the title ‘view points on data points: a shared vocabulary for cross-domain conversations on data and metadata’. the authors have observed that sharing data across scientific domains is often impeded by differences in the language used to describe data and metadata. to avoid confusion, the authors develop a terminology. part of the confusion concerns disagreement about the boundaries between data and metadata; and that what is metadata in one domain can be data in another. the shift between data and metadata is what they name as ‘semantic transposition’. i find that such shifts are a virtue and a strength and as the authors say, there is no fixed boundary between data and metadata, and both can be acted upon by people and machines. the article draws on and refers to many other standards and developments, most cited are the data model of observations and measurements (iso 19156) and tools of the data documentation initiative’s cross domain integration (ddi-cdi). the article is thorough and explanatory with many examples and diagrams for learning, including examples of transformations between the formats: wide, long, and multidimensional. the long format of entityattribute-value has the value domain restricted by the attribute, and in examples time and source are added, which demonstrates how further metadata enter the format. when transposing to the wide format, this is a more familiar data matrix where the same value domain applies to the complete column. the multidimensional format with facets is for most readers the familiar aggregations published by statistical agencies. the authors argue that their domain-independent vocabulary enables the cross-domain conversation. george alter is research professor emeritus in the institute for social research at the university of michigan, flavio rizzolo is senior data science architect for statistics canada. kathi schleidt is a data scientist and the founder of datacove. https://doi.org/ 2/3 rasmussen, karsten boye (2023) editor’s notes: fair bot. as metadata is data is metadata is data ..., iassist quarterly 47(1), pp. 12. doi https://doi.org/10.29173/iq1086 the format discussion in the first article is also the point of the second paper on ‘modernizing data management at the us bureau of labor statistics’. the us bureau of labor statistics (bls) has a focus on time series and daniel w. gillman and clayton waring (both from the bls) view time series data as a combination of three components: a measure element; an element for person, places, and things (ppt); and a time element. in the paper gillman and waring also describe the conceptual model (uml) and the design and features of the system. first, they go back in history to the 1970s and the codd relational model and to the standards developed and refined after 2000. you will not be surprised to find here among the references also the data documentation initiative’s cross domain integration (ddi-cdi). the mission is: ‘to find a simple and intuitive way to store and organize statistical data with the goal of making it easy to find and use the data’. a semantic approach is adopted, i.e. the focus is on the meaning of the data based upon the ‘measures / people-places-things / time’ model. detailed examples show how ppt are categories of dimensions, for instance ‘nurse’ is in the standard occupational classification and 'hospital' in the north american industry classification system. the paper – like the first paper – also refers to multidimensional structures. the modernization described at bls is expected to be released in early 2023. the third paper is by joão aguiar castro, joana rodrigues, paula mena matos, célia sales, and cristina ribeiro where all authors are affiliated with the university of porto. like the earlier articles this also references the data documentation initiative (ddi) with a focus on the concepts behind the fair acronym: findable, accessible, interoperable, and reusable. the title is: ‘getting in touch with metadata: a ddi subset for fair metadata production in clinical psychology’. clinical psychology is not an area frequently occurring in iassist quarterly, but it turns out that the project described started with interviews and data description sessions with research groups in the social sciences for identifying a manageable ddi subset. the project also draws on other projects such as tail, together, and dendro. the tail project concerned the integration metadata tools in the research workflow and assessed the requirements of researchers from different domains. together was a project in the psycho-oncology domain and family-centered care for hereditary cancer. as most researchers showed to be inexperienced with metadata, they concentrated on a ddi subset that meant that fair metadata would be available for deposit. support for researchers is essential as the they have the domain expertise and can create highly detailed descriptions. on the other hand, data curators can ensure that the metadata follow the rules of fair. this was achieved by embedding the dendro platform in the research workflow, where creation of metadata is performed in an incremental description of the data. the article includes screenshots of the user interface showing the choice of vocabularies. the approach and the adoption of a ddi subset produced more comprehensive metadata than is usually available. submissions of papers for the iassist quarterly are always very welcome. we welcome input from iassist conferences or other conferences and workshops, from local presentations or papers especially written for the iq. when you are preparing such a presentation, give a thought to turning your one-time presentation into a lasting contribution. doing that after the event also gives you the opportunity of improving your work after feedback. we encourage you to login or create an author profile at https://www.iassistquarterly.com (our open journal system application). we permit https://doi.org/ 3/3 rasmussen, karsten boye (2023) editor’s notes: fair bot. as metadata is data is metadata is data ..., iassist quarterly 47(1), pp. 12. doi https://doi.org/10.29173/iq1086 authors to have 'deep links' into the iq as well as deposition of the paper in your local repository. chairing a conference session or workshop with the purpose of aggregating and integrating papers for a special issue iq is also much appreciated as the information reaches many more people than the limited number of session participants and will be readily available on the iassist quarterly website at https://www.iassistquarterly.com. authors are very welcome to take a look at the instructions and layout: https://www.iassistquarterly.com/index.php/iassist/about/submissions authors can also contact me directly via e-mail: kbr@sam.sdu.dk. should you be interested in compiling a special issue for the iq as guest editor(s) i will also be delighted to hear from you. karsten boye rasmussen march 2023 https://doi.org/ https://www.iassistquarterly.com/index.php/iassist/about/submissions 1/23 mozersky, jessica; walsh, heidi; parsons, meredith; mcintosh, tristan; baldwin, kari; dubois, james m. (2019) are we ready to share qualitative research data? knowledge and preparedness among qualitative researchers, irb members, and data repository curators, iassist quarterly 43(4), pp. 1-23. doi: https://doi.org/10.29173/iq952 are we ready to share qualitative research data? knowledge and preparedness among qualitative researchers, irb members, and data repository curators jessica mozersky, heidi walsh, meredith parsons, tristan mcintosh, kari baldwin, james m. dubois1 2 abstract data sharing maximizes the value of data, which is time and resource intensive to collect. major funding bodies in the united states (us), like the national institutes of health (nih), require data sharing and researchers frequently share de-identified quantitative data. in contrast, qualitative data are rarely shared in the us but the increasing trend towards data sharing and open science suggest this may be required in future. qualitative methods are often used to explore sensitive health topics raising unique ethical challenges regarding protecting confidentiality while maintaining enough contextual detail for secondary analyses. here, we report findings from semi-structured in-depth interviews with 30 data repository curators, 30 qualitative researchers, and 30 irb staff members to explore their experience and knowledge of qds. our findings indicate that all stakeholder groups lack preparedness for qds. researchers are the least knowledgeable and are often unfamiliar with the concept of sharing qualitative data in a repository. curators are highly supportive of qds, but not all have experienced curating qualitative data sets and indicated they would like guidance and standards specific to qds. irb members lack familiarity with qds although they support it as long as proper legal and regulatory procedures are followed. irb members and data curators are not prepared to advise researchers on legal and regulatory matters, potentially leaving researchers who have the least knowledge with no guidance. ethical and productive qds will require overcoming barriers, creating standards, and changing long held practices among all stakeholder groups. keywords qualitative data sharing, data curators, irb members, research personnel, qualitative research, interviews, attitudes, research ethics background data is a valuable commodity that is costly and labor-intensive to collect (mannheimer et al., 2018). data sharing is one way to maximize the value of data. benefits of data sharing include avoiding duplication of research, enabling secondary analyses, reducing research participant burden, and providing training resources to students (mannheimer et al., 2018, corti, 2000, corti, 2012, dubois et al., 2018, borgman, 2012). data sharing also promotes transparency and enables replication of findings. major united states (us) funders such as the national institutes of health (nih) and national science foundation (nsf), select journals, and private funders like the gates foundation, all have policies requiring some form of data sharing (mannheimer et al., 2018, meyer, 2018). the nih ‘expects and supports the timely release and sharing of final research data from nih-supported studies for use by other researchers’ (national institutes of health, 2003). while the nih recognizes that data sharing may be complicated or limited by ‘institutional policies, local irb rules, and local, state and federal laws and regulations, including the hipaa privacy rule’ they endorse data sharing wherever possible. in some cases, nih mandates data sharing or an explanation to justify why data cannot be shared (meyer, 2018, https://doi.org/10.29173/iq952 2/23 mozersky, jessica; walsh, heidi; parsons, meredith; mcintosh, tristan; baldwin, kari; dubois, james m. (2019) are we ready to share qualitative research data? knowledge and preparedness among qualitative researchers, irb members, and data repository curators, iassist quarterly 43(4), pp. 1-23. doi: https://doi.org/10.29173/iq952 national institutes of health, 2003). even when not required, there is arguably an ethical obligation to share data collected using public funds (dubois et al., 2018, meyer, 2018). notably, these policies do not specify the type of data to be shared, although most were likely written with quantitative data sharing in mind. there is, however, no reason to assume these policies may not apply equally to qualitative data. in fact, the increasing trends towards data sharing and open science suggest that qualitative researchers may find themselves subject to data sharing requirements (meyer, 2018, yoon, 2014, tsai et al., 2016, mannheimer et al., 2018, dubois et al., 2018). differences between qualitative and quantitative data raise challenges–both epistemological and ethical–to sharing qualitative data (tsai et al., 2016, guishard, 2017b, mannheimer et al., 2018, yoon, 2014). ethically, concerns exist regarding informed consent, data ownership, confidentiality, and adequately anonymizing qualitative data while maintaining enough contextual details to enable secondary analyses. qualitative methods are frequently used to explore highly sensitive or stigmatized issues, which increases the harmful effects of potential re-identification of participants (tsai et al., 2016, guishard, 2017a). the health insurance portability and accountability acts (hipaa) privacy rule regulates the collection, handling, sharing, and transfer of protected health information (phi). hipaa provides two methods to de-identify phi: removal of 18 ‘safe harbor’ identifiers3 or expert determination, although other anonymizing tools exist (hipaa 2017, meyer, 2018). in practice, removal of the 18 safe harbor identifiers is commonly used as it is easy to operationalize. however, qualitative data present unique challenges for anonymization that will require more sophisticated tools than ‘rote application of hipaa safe harbor rules’ to ensure adequate anonymization (meyer, 2018). for instance, three demographic variables may not appear to be identifiers when examined independently: ‘female’, ‘hispanic’, and ‘psychiatric nurse’ (and none are considered hipaa safe harbor identifiers). however, if a researcher publishes the hospital where the research occurred, then it may become relatively easy for individuals who work at the hospital to identify this person (dubois et al., 2018). to complicate matters further, qualitative health data may not be considered phi when gathered in a research setting, and therefore may not be subject to hipaa, leaving it in a ‘regulatory twilight zone’ (meyer, 2018). epistemologically, researchers may not view qualitative data that are removed from the original context as being appropriate for analysis by third parties who lack the original researcher’s contextual knowledge and expertise (dubois et al., 2018, mannheimer et al., 2018, tsai et al., 2016). qualitative researchers have expressed concerns that their data are often collected within relationships of trust and that secondary analyses by new researchers are neither feasible nor ethically acceptable (guishard, 2017a). for some, the complexity of qualitative data simply does not lend itself to reuse or sharing (yoon, 2014). the notion of replicating qualitative data may also be problematic because qualitative studies are rarely meant to be representative or generalizable (dubois et al., 2018). it is not suitable to speak of replicating qualitative data in the same manner as replicating quantitative data. however, sharing codebooks, interview guides, and transcripts would allow others to verify the rationale for the claims made in a particular qualitative study (dubois et al., 2018). further, secondary analyses of qualitative data can yield new insights (bishop and kuula-luumi, 2017). these challenges to qualitative data sharing (qds) may be especially pertinent in the us where qds is relatively new (yoon, 2014, corti, 2000, corti, 2012). in europe and australia, policies to promote and https://doi.org/10.29173/iq952 3/23 mozersky, jessica; walsh, heidi; parsons, meredith; mcintosh, tristan; baldwin, kari; dubois, james m. (2019) are we ready to share qualitative research data? knowledge and preparedness among qualitative researchers, irb members, and data repository curators, iassist quarterly 43(4), pp. 1-23. doi: https://doi.org/10.29173/iq952 encourage qds have existed for some time, and researchers are more familiar with the concept (corti and backhouse, 2000, kuula, 2011, yoon, 2014). repositories such as the uk data archive or finnish social science data archive provide national storehouses for social science and humanities data (corti and backhouse, 2000, kuula, 22011, yoon, 2014). data repositories exist in the us, but most are not capable of handling sensitive qualitative data (antes et al., 2017, mannheimer et al., 2018). some repositories may only be able to take data that will be made available open access, prohibiting them from accepting sensitive data (antes et al., 2017). a review of qualitative data repository guidelines found 32 english language social science repositories globally; only 12 repositories had written guidelines for qualitative data, and only 5 were located in the us (antes et al., 2017). the limited available data suggests us social scientists rarely share their data (jeng et al., 2016). a survey of over 1,200 researchers regarding data sharing practices found that those in social sciences and medicine were the least willing to share data compared to disciplines such as atmospheric science and biology (tenopir et al., 2011). this is related to the fact that these data are more likely to be sensitive human subject data. those in the social sciences also reported the lowest level of institutional support for data management, including the necessary processes, tools, funding, and technology (tenopir et al., 2011). the need for our project in light of the relative newness and potential expectation of qds in the us, research is needed to explore diverse stakeholder experiences of and attitudes towards qds. we received a grant from the nih to explore the barriers and facilitators to qds. we conducted 120 semi-structured in-depth telephone interviews with 4 stakeholder groups: 30 qualitative research participants, 30 data repository curators, 30 qualitative researchers, and 30 irb staff members to explore knowledge, barriers, and benefits of qds. while qualitative research can take many forms, we focus on qualitative health data; that is, textual data gathered in an interview or focus group that may include potentially sensitive health information. focusing on transcribed data provides a useful starting point as they are common, useful, and may be easier to anonymize than audio or video files that contain voice prints or facial images (dubois et al., 2018). methods data collection we report findings from qualitative interviews and pre-interview surveys with three stakeholder groups—qualitative researchers, irb staff, and data curators—regarding their preparedness, or lack thereof, for qds. findings from the research participant stakeholder group will be reported in a forthcoming article. we built recruitment lists using publicly available databases or personal contacts supplemented by snowball recruitment. data curators were identified through the open access directory (open access directory, 2016). curators were eligible if they worked at a repository that could accept multidisciplinary or social science data, although not all repositories had experience curating qualitative data https://doi.org/10.29173/iq952 4/23 mozersky, jessica; walsh, heidi; parsons, meredith; mcintosh, tristan; baldwin, kari; dubois, james m. (2019) are we ready to share qualitative research data? knowledge and preparedness among qualitative researchers, irb members, and data repository curators, iassist quarterly 43(4), pp. 1-23. doi: https://doi.org/10.29173/iq952 sets. due to the limited number of qualitative repositories globally, we included curators from international social science repositories in addition to us curators (antes et al., 2017). irb members and staff were identified through the websites of 62 academic medical institutions with nih clinical and translation science awards (ctsa). irb members and staff had to have experience reviewing and approving qualitative research studies. qualitative researchers were identified through personal contacts and publicly available information from organizations that represent minority researchers. we purposively sampled qualitative researchers to ensure one-third of the sample was from racial or ethnic minorities. researchers were eligible if they were the principal investigator or lead of a project involving qualitative data collection. all participants had to speak english. the research was approved by the washington university school of medicine institutional review board (irb) as expedited human subjects research. recruitment we recruited participants via email invitation. after providing informed consent, participants completed a brief survey online administered via qualtrics to obtain demographic details and prior experience before completing a semi-structured telephone interview. interviews explored prior experience and knowledge, perceived barriers and benefits, and preparedness for qds. questions ranged from broad and open-ended such as ‘tell me what you know about qds’ to more focused questions about particular issues we anticipated would arise such as ‘what concerns do you have about cost?’ interview guides for each stakeholder group followed a similar structure, but questions were adapted to be relevant for each stakeholder group. for instance, we asked all three stakeholder groups about their preparedness for qds; researchers were asked about preparedness to deposit qualitative data with a repository, while the question for irb members and data curators was posed as preparedness to advise researchers regarding qds. participants were paid $30 usd following completion of the interview. trained interviewers (parsons, baldwin, walsh) conducted the interviews, which lasted approximately one hour. data analysis interviews were audio recorded and professionally transcribed verbatim before being uploaded to dedoose, a qualitative data analysis software.4 survey data were exported from qualtrics into dedoose and linked to the relevant transcript. the principal investigator, co-investigator, and project manager (dubois, mozersky, walsh) led codebook development for each stakeholder group with input from all interviewers. coding involved a combination of inductive concept coding and deductive structural coding (saldana, 2016). inductive concept coding was used to capture participant views and attitudes about qds that arose spontaneously among stakeholder groups (usually in response to broad, openended questions although not exclusively). structural deductive coding was used to categorize answers to the specific questions regarding the hypothesized benefits and challenges we anticipated based on the limited literature (e.g., concerns about confidentiality or cost). we assigned one coder to each stakeholder group, enabling them to become familiar with the particular data set and codebook. coders kept detailed notes and memos during coding. one author (mozersky) served as the gold standard coder for all stakeholder groups. in the first step, the coder and gold standard coder blind coded a single transcript, discussed discrepancies, and resolved them. we repeated this process until both coders reliably coded without major discrepancies. we held weekly meetings with coders to discuss questions https://doi.org/10.29173/iq952 5/23 mozersky, jessica; walsh, heidi; parsons, meredith; mcintosh, tristan; baldwin, kari; dubois, james m. (2019) are we ready to share qualitative research data? knowledge and preparedness among qualitative researchers, irb members, and data repository curators, iassist quarterly 43(4), pp. 1-23. doi: https://doi.org/10.29173/iq952 or discrepancies. for each stakeholder group, the gold standard coder periodically blind coded interviews to ensure agreement between the gold standard and coders. this blind coding occurred for ~15% of each set of 30 interviews (i.e., 4 – 5 transcripts per set). discrepancies arising during blind coding were discussed with coders and resolved through consensus. results we present our results by stakeholder group, but all groups demonstrate a lack of knowledge and experience with qds. tables i, ii, and iii contain demographic details and prior experience for each group. researchers are the least knowledgeable about qds and were often unfamiliar with the very idea of sharing qualitative data with a repository. curators are highly supportive of qds, but not all had experienced curating qualitative data sets. furthermore, many indicated they would like guidance and standards specific to sharing qualitative data. irb members lacked familiarity with qds plans, as they often had not encountered them. however, in principle, they did not see a problem with qds as long as proper legal and regulatory procedures are followed and that sharing is consistent with the information provided in the informed consent document. these findings are not surprising—qds is relatively new, uncharted, and many have not yet experienced it. researchers table i: researchers (n=30) demographic frequency percenta age (blank) (blank) 20-29 0 0% 30-39 6 20% 40-49 14 47% 50-59 6 20% 60 or older 4 13% sex (blank) (blank) female 26 87% male 4 13% raceb (blank) (blank) asian 1 3% black or african american 7 23% white 20 67% prefer not to answer 1 3% ethnicity (blank) (blank) hispanic or latino 2 7% not hispanic or latino 28 93% region of birth (blank) (blank) united states 28 93% africa 1 3% asia 1 3% education (blank) (blank) doctoral degree 28 93% other degree 2 7% https://doi.org/10.29173/iq952 6/23 mozersky, jessica; walsh, heidi; parsons, meredith; mcintosh, tristan; baldwin, kari; dubois, james m. (2019) are we ready to share qualitative research data? knowledge and preparedness among qualitative researchers, irb members, and data repository curators, iassist quarterly 43(4), pp. 1-23. doi: https://doi.org/10.29173/iq952 degree field (blank) (blank) anthropology 4 13% communications 2 7% psychology 6 20% public health 3 10% social work 3 10% other 8 27% academic rank (blank) (blank) instructor 1 3% assistant professor 10 33% associate professor 8 27% full professor 7 23% other 4 13% years’ experience (blank) (blank) 0-2 1 3% 3-5 3 10% 5-10 9 30% 10 or more 17 57% collecting sensitive datab (blank) (blank) personal health information (phi) 19 63% sensitive non-phi data 17 57% populations studiedb (blank) (blank) healthy individuals 24 80% patients 15 50% children 7 23% pregnant women 3 10% older adults 9 30% individuals with sensitive diagnoses 10 33% economically disadvantaged individuals 21 70% prisoners 4 13% other 13 43% methods usedb (blank) (blank) in-depth interviews 30 100% focus groups 25 83% observations 18 60% community based participatory research (cbpr) 15 50% coding of archival data 9 30% journals written by participants 2 7% other 3 10% has shared data with a repository (blank) (blank) yes 1 3% no 29 97% knows a peer who has shared data with a repository (blank) (blank) yes 4 13% no 26 87% https://doi.org/10.29173/iq952 7/23 mozersky, jessica; walsh, heidi; parsons, meredith; mcintosh, tristan; baldwin, kari; dubois, james m. (2019) are we ready to share qualitative research data? knowledge and preparedness among qualitative researchers, irb members, and data repository curators, iassist quarterly 43(4), pp. 1-23. doi: https://doi.org/10.29173/iq952 anumbers may not add up to 100% due to rounding bparticipants were asked to check all that apply so percentages will not equal 100% lack of knowledge among qualitative researchers, only one individual had experience depositing qualitative data with a repository (table i). four individuals were aware of a peer who had used a qualitative repository. in the early phase of coding, the team created a code for lack of knowledge to capture demonstrations that researchers did not have knowledge or understanding of qds. we created this code because it was evident upon reading transcripts that lack of knowledge permeated researchers’ responses. researchers’ responses were often speculative and did not reflect actual experience with qds. this code captured a wide range of sentiments from having never heard of a repository to being familiar or curious about their existence but not knowing the details of how repositories operate. when asked what they know about sharing qualitative data in a repository, some researchers were entirely unfamiliar with the concept, as this researcher indicated, ‘i'm not aware that there is such a thing…’ [r1]. others associated data sharing with quantitative data: i actually never thought about it before. when i hear ‘data repository,’ i think quantitative data, so i never even considered using it for qualitative [r3]. for some respondents, taking part in the interview for our study was the first time they had encountered the idea of qds. as this researcher notes: we’ve never done it…i don’t know what the process is, or who gets to access it, or how it needs to be deidentified…other than what you described earlier [r23]. other researchers were familiar with the concept and open to the idea of sharing qualitative data but lacked practical knowledge in terms of how, where, or what is actually involved in the process. this researcher stated: my first impulse was, ‘yeah, i could do that,’ and then, you know, i don’t know which data repositories are good. i don’t know what the requirements are. who knows what the irbs going to want and how participants are going to react [r8]. some researchers were uncomfortable with the idea of sharing qualitative data at all, and lack of knowledge played a role in these concerns. as this researcher stated: personally…i don’t feel comfortable sharing my data in a data repository…and i’m not sure what would be in the repository, whether it would just be, you know, interview transcripts, or would it be field notes as well? what would be the kind of data that would be placed in a repository? [r25]. lack of familiarity also leads to a lack of trust. according to this researcher, ‘i've never worked with those people, so i don't even know, can i trust them with the data? that's a big question’ [r22]. researchers frequently expressed a more fundamental concern regarding the purpose or value of qds. that is, they first need to be convinced about the benefits before considering placing data in a https://doi.org/10.29173/iq952 8/23 mozersky, jessica; walsh, heidi; parsons, meredith; mcintosh, tristan; baldwin, kari; dubois, james m. (2019) are we ready to share qualitative research data? knowledge and preparedness among qualitative researchers, irb members, and data repository curators, iassist quarterly 43(4), pp. 1-23. doi: https://doi.org/10.29173/iq952 repository. as this researcher asked ‘why does this make any sense to do? why? why? ’ [r29]. another researcher suggested that qualitative data is shared through journal articles, literature, and intellectual discourse and that a repository was ‘a very narrow notion’ of sharing qualitative data. this attitude is noteworthy because data presented in the literature or at conferences is arguably highly curated and edited, leaving the majority of data sequestered from other researchers or the public (dubois et al., 2018). preparedness for qds we asked researchers how prepared they felt to share qualitative data, and the majority reported they were not prepared for two primary reasons: either 1) they did not see the value or benefit in sharing qualitative data and were resistant, or 2) they did not have enough knowledge regarding how to deposit qualitative data. this researcher stated, ‘i feel prepared, but i feel skeptical’ [r21]. according to this researcher: i would need an incentive to do it, like somebody would have to either pay me or, there would have to be, like, real strong norms in the field that this is what people do…[r11]. another researcher commented, ‘publication is important for me, and i need all this data without sharing…it…’ [r9]. in contrast, other researchers were open to the idea, but they reported not knowing how to go about sharing their data. according to this researcher, ‘i'm interested in doing it…i haven't had experience doing it’ [r6]. another researcher stated, ‘i’ve thought through some of the issues, but i’ve never actually had to do it and think through the logistics.’ [r27] this researcher said, ‘i don’t even know…what my university’s stance is on it, so i would say no. i’m not very prepared at all’ [r13]. while researchers expressed a range of attitudes about qds from supportive to curious to highly suspicious, all researchers lacked knowledge regarding qds, repositories, and the potential to share qualitative data via a repository. irb members table ii: irb (n=30) demographic frequency percenta age (blank) (blank) 20-29 4 13% 30-39 9 30% 40-49 5 17% 50-59 10 33% 60 or older 2 7% sex (blank) (blank) female 23 77% male 7 23% raceb (blank) (blank) https://doi.org/10.29173/iq952 9/23 mozersky, jessica; walsh, heidi; parsons, meredith; mcintosh, tristan; baldwin, kari; dubois, james m. (2019) are we ready to share qualitative research data? knowledge and preparedness among qualitative researchers, irb members, and data repository curators, iassist quarterly 43(4), pp. 1-23. doi: https://doi.org/10.29173/iq952 native hawaiian or other pacific islander 1 3% multiple races 1 3% black or african american 1 3% white 26 87% ethnicity (blank) (blank) hispanic or latino 0 0% not hispanic or latino 30 100% region of birth (blank) (blank) united states 30 100% education (blank) (blank) bachelor's degree 7 23% master's degree 12 40% doctoral degree 10 33% other 1 3% years’ experience (blank) (blank) 0-2 4 13% 3-5 8 27% 5-10 8 27% 10 or more 10 33% reviews sensitive studiesb (blank) (blank) personal health information (phi) 27 90% sensitive non-phi data 28 93% populations reviewedb (blank) (blank) healthy individuals 29 97% patients 22 73% children 25 83% pregnant women 21 70% older adults 23 77% individuals with sensitive diagnoses 24 80% economically disadvantaged individuals 26 87% individuals with cognitive impairment 21 70% prisoners 22 73% other 13 43% has reviewed data sharing plans (blank) (blank) yes 14 47% no 16 53% anumbers may not add up to 100% due to rounding bparticipants were asked to check all that apply so percentages will not equal 100% lack of knowledge among irb members and staff, just over half (16/30) had reviewed a study with a qds plan (table ii). we also created a lack of knowledge code to capture demonstrations of lacking knowledge regarding qds, available repositories, permissions required, data security and access, and references to wanting guidelines, or standards. in contrast to researchers whose lack of knowledge contributed to their https://doi.org/10.29173/iq952 10/23 mozersky, jessica; walsh, heidi; parsons, meredith; mcintosh, tristan; baldwin, kari; dubois, james m. (2019) are we ready to share qualitative research data? knowledge and preparedness among qualitative researchers, irb members, and data repository curators, iassist quarterly 43(4), pp. 1-23. doi: https://doi.org/10.29173/iq952 concerns about qds, irb members supported qds as long as sharing was compliant with regulations and participants had provided informed consent. however, their responses demonstrated a lack of familiarity with the details of qds in practice. similar to researchers, some irb members had never heard of a qualitative repository, reporting they ‘don’t know anything’ or ‘hadn’t thought about it before’ when asked what they knew about depositing qualitative data with a repository. others were slightly familiar with qds but did not know details of how it operated, as this irb member stated: i don’t really know a lot about any repositories of qualitative data at this time, but it is something that’s come on my radar recently that i want to look into more [irb 13]. another irb member stated: i don’t know much about it other than we have study teams that request to do it from time to time but as for what that process entails exactly, i don't know the mechanics of it [irb 11]. irb members and staff were more familiar with quantitative, rather than qualitative, data sharing plans and noted that regulations and guidelines reflected this. this irb member commented, ‘i think the regulations are probably more built with quantitative in mind’ [irb 21]. de-identified data irb members generally assumed that only de-identified qualitative data could be shared similar to how de-identified quantitative data is commonly shared. according to this irb member, ‘i don’t have any problems with it if the data is stripped of identifiers, and participants are aware that it’s going to be shared’ [irb 13]. this individual said ‘we’re gonna make sure that it’s deidentified before we release it to anybody’ [irb 1]. however, the current standards for de-identification of quantitative health data— removal of the 18 hipaa safe harbor direct identifiers—will not suffice to anonymize qualitative data— that is, to ensure that participants could not be deductively identified (meyer, 2018). most irb members were not familiar with anonymizing qualitative data sets. they were also unsure when data might be considered de-identified. this irb member noted the difficulty of determining when qualitative data are considered de-identified, ‘…unless again the data's completely de-identified, and that's the tricky part with the qualitative data’ [irb 17]. another irb member described the competing ideas of what is classified as de-identified data: what does de-identified mean?... my experience is that the word de-identified is tossed around fairly haphazardly and no one really knows what that means. there is a hipaa security rule that goes into de-identified, but most qualitative data isn’t covered under hipaa, so what does that mean in a nonhipaa center? [irb 19]. irb members expressed uncertainty about what constitutes sensitive data or vulnerable populations, as this person stated: https://doi.org/10.29173/iq952 11/23 mozersky, jessica; walsh, heidi; parsons, meredith; mcintosh, tristan; baldwin, kari; dubois, james m. (2019) are we ready to share qualitative research data? knowledge and preparedness among qualitative researchers, irb members, and data repository curators, iassist quarterly 43(4), pp. 1-23. doi: https://doi.org/10.29173/iq952 we don’t have a national, or even an international, set of standards on depositing qualitative data. i think that some of the competing standards would be everything from how do we classify materials as highly vulnerable populations, highly sensitive, all the way to not sensitive, not vulnerable populations [irb 2]. in contrast, a few speculated that identifiable data can be shared, as long as the consent form discloses this, ‘…if participants were told from the beginning that their identifiable data will be shared with other researchers, that i believe it’s okay’ [irb 4]. interestingly, participants did not discuss the notion of ‘a limited dataset;’ such datasets are commonly shared without participant consent even though they include some hipaa identifiers that could, in principle, be used to re-identify participants; secondary users typically must enter into data use agreements that require the protection of data confidentiality and prohibit attempts to re-identify participants (2017). competing regulations: we used to tell people to destroy data we asked participants to think of competing regulations that could be barriers to qds. most described a historical or current tendency of irbs to ask that data be destroyed, or at least identifiers are destroyed, and that this could conflict with the ability to share data. most irb members and staff advised against indicating that data will be destroyed in consent forms but recognized that many people have been given this advice in the past or that different agencies have different standards. they suggested including broad language in consent forms to enable data sharing. this irb member noted: for years and years…researchers would always write into the consent forms that they would destroy certain data, or at least they'd destroy identifiers. but then unfortunately, we discovered that actually runs up against state records retention policies here in [state x] [irb 16]. some irb members described a policy of advising researchers to destroy identifiers after three years, although neither the nsf, nih, nor the common rule require data destruction (meyer, 2018). according to this individual: what we require here is destroy the identifiers and the consent forms after 36 months…and, you know, if they've got a requirement to share it, then obviously there's gonna be some confusion on their end [irb 28]. another irb member described a need to create new norms given the historical tendency of social scientists to destroy data as a way to protect privacy and confidentiality: i think that’s a standard…in social science to try to reassure participants in research that their data isn’t gonna be forever available…there’s gonna have to be some clarification on what the norm is [irb 7]. for some, the interview itself led irb members to consider issues they had not considered before about qds. this irb member wondered: https://doi.org/10.29173/iq952 12/23 mozersky, jessica; walsh, heidi; parsons, meredith; mcintosh, tristan; baldwin, kari; dubois, james m. (2019) are we ready to share qualitative research data? knowledge and preparedness among qualitative researchers, irb members, and data repository curators, iassist quarterly 43(4), pp. 1-23. doi: https://doi.org/10.29173/iq952 we'll often tell people you have to keep these things for three years. but when i think about it, we probably should be saying…de-identify it and keep it forever, if you want [irb 25]. given the historical tendency of irbs to advise data destruction, it is unsurprising that consent forms frequently do not address data sharing. we asked irb members if their irb had a policy on data sharing when consent forms are silent on the topic, and their answers varied widely.5 in some cases, there was no policy, and again our interview was a stimulus for thinking about this issue for the first time. as this irb member commented, ‘it’s a good question. i don’t think we have a set policy’ [irb23]. according to this irb member: i wish we had one [referring to a policy]. you know, other than that hipaa authorization language, which specifies generally who will receive what information, we are not great at explaining that in consent forms [irb 20]. according to this irb member, ‘the last five and half months we've mandated the use of a consent form that does include data-sharing provisions’ [irb 28]. some felt a consent form that was silent on the matter of data sharing provided more flexibility with regard to allowing data sharing compared to a consent form that explicitly says data will not be shared: ‘if it was silent, we…tend to be a little bit more flexible and allowable’ [irb 12]. the assumption is that sharing de-identified data might be allowable when consent forms are silent on the matter, but what constitutes de-identified qualitative data is still not resolved as this irb member stated: …it's a tricky judgment call sometimes if it's long transcripts of information then we might say, no, we want you to go back and get consent for that potentially [irb 17]. in contrast, others reported that if a consent form was silent, then data sharing was not permitted, ‘if it’s silent on data sharing, then it’s not sharing data. they don’t have approval to do so’ [irb 13]. again, this position is contrary to regulations that permit the sharing of ‘limited data sets’ that contain phi without patient consent (hipaa 2017). impact on irb approval status when asked what impact data sharing plans might have on irb approval status, the answers varied widely from completely unsure to no changes, to possible changes including potentially enhancing the application, but no irb member or staff indicated that research would not be approved. some suggested data sharing plans would increase the likelihood of approval because it would clarify and transparently let the irb know about these plans. as this individual stated, ‘we would leap for joy—cuz we don’t normally get something as well thought out as a data-sharing plan’ [irb 19]. another individual stated, ‘i think it would make the approvals much, much quicker’ [irb 20]. others noted that as long as the consent form stated plans to share data, they were likely to approve the data sharing plans. however, given the lack of standards for de-identifying qualitative data, it is unclear what such a review would accomplish. in contrast, others thought qds would raise the level of scrutiny by an irb, potentially making a project ineligible for exempt status, as this person stated: https://doi.org/10.29173/iq952 13/23 mozersky, jessica; walsh, heidi; parsons, meredith; mcintosh, tristan; baldwin, kari; dubois, james m. (2019) are we ready to share qualitative research data? knowledge and preparedness among qualitative researchers, irb members, and data repository curators, iassist quarterly 43(4), pp. 1-23. doi: https://doi.org/10.29173/iq952 …if you're also sharing it, that's an extra layer of concern that i don't think meets the exemption category i know some chairs would regard data-sharing as more than a minor thing on research studies and want to take some of them to the full board…[irb 28]. responsibility and preparedness to advise researchers when asked how prepared they felt to advise researchers, most irb members and staff reported they are not prepared because they lack experience or that data sharing is under the purview of another official or department (e.g., privacy or compliance officers, legal counsel, computing or technology offices) and required a collaborative effort between various officials. irb members and staff described a lack of guidance and standards with regards to qds. according to one individual: there's not really a set of standards in terms of institutions, privacy officers, irbs. i mean, this stuff is all over the place and … people just don’t understand it very well, and there aren't good answers to give sometimes because we don’t have information [irb 20]. this left many irb members feeling unprepared to advise researchers on legal and regulatory matters. when asked who was responsible for providing guidance to researchers on legal and regulatory matters pertaining to qds, most irb members indicated a collaborative effort that required input from both the home institution and the repository. this individual said: i think the institution should have relationships with data repositories. and then, manage that relationship with their researchers, and not just leave it up to a researcher-data repository relationship [irb 08]. irb members expressed concerns regarding variations in policies between the institution, the repository, and state laws. according to this irb member: the more the data repository can provide that information and…even advise on what they know of state laws, other policies that might apply, the better. but…we do our due diligence and make sure that we don't see any other rules that might affect us specifically. so it'd be a team effort [irb 17]. in contrast, some thought the home institutions were responsible for guiding researchers: the home institutions need to be able to guide their researchers about this. i would hope that they would understand what the repositories were doing, and offering, and providing [irb 7]. others thought repositories were primarily responsible: ‘data repository needs to have all of this in order…i would put it more on the…onus of the data repository than the institution’ [irb 25]. most often irb members wondered or speculated about who was responsible, but they were not sure or had not thought about this before. according to this irb member, ‘that’s a very good question. i don't know…i honestly haven’t thought much about that.’ [irb 15]. another irb member stated: https://doi.org/10.29173/iq952 14/23 mozersky, jessica; walsh, heidi; parsons, meredith; mcintosh, tristan; baldwin, kari; dubois, james m. (2019) are we ready to share qualitative research data? knowledge and preparedness among qualitative researchers, irb members, and data repository curators, iassist quarterly 43(4), pp. 1-23. doi: https://doi.org/10.29173/iq952 i would think the repository’s role is just to hold it, and whatever entity is wanting to gather that, i believe the agreement then would be between the home institution and whoever the next institution is. i may be inaccurate about that [irb 24]. according to yet another individual: so, you know theoretically, i think the right answer is the repository if the repository can have a handle on all of the state and local requirements [irb 20]. what is clear from irb members’ answers is that they are not describing current practices but speculating about what they think would happen, demonstrating a lack of experience with qds. curators table iii: curators (n=30) demographic frequency percenta age (blank) (blank) 20-29 1 3% 30-39 12 40% 40-49 9 30% 50-59 6 20% 60 or older 2 7% sex (blank) (blank) female 14 47% male 16 53% raceb (blank) (blank) black or african american 2 7% white 26 86% prefer not to answer 2 7% ethnicity (blank) (blank) hispanic or latino 1 3% not hispanic or latino 27 90% prefer not to answer 2 7% region of birth (blank) (blank) united states 25 84% european union 4 13% oceania 1 3% education (blank) (blank) some college 1 3% bachelor's degree 1 3% master's degree 14 47% doctoral degree 14 47% years’ experience (blank) (blank) 0-2 8 27% https://doi.org/10.29173/iq952 15/23 mozersky, jessica; walsh, heidi; parsons, meredith; mcintosh, tristan; baldwin, kari; dubois, james m. (2019) are we ready to share qualitative research data? knowledge and preparedness among qualitative researchers, irb members, and data repository curators, iassist quarterly 43(4), pp. 1-23. doi: https://doi.org/10.29173/iq952 3-5 8 27% 5-10 4 13% 10 or more 10 33% personally curates data (blank) (blank) yes 14 47% no 16 53% curates sensitive datab (blank) (blank) personal health information (phi) 12 40% sensitive non-phi data 16 53% types of data personally curatedb (blank) (blank) healthy individuals 16 53% patients 11 36% children 6 20% pregnant women 5 17% older adults 10 33% individuals with sensitive diagnoses 6 20% economically disadvantaged individuals 9 30% individuals with cognitive impairment 4 13% prisoners 4 13% other 13 43% types of data deposited in repositoryb (blank) (blank) quantitative 24 80% qualitative 23 77% mixed methods 25 83% other 6 20% types of data curated by repositoryb (blank) (blank) none 2 7% written transcripts 17 57% audio-recordings 9 30% video-recordings 9 30% field notes 12 40% archival data 17 57% journals written by participants 5 17% other 9 30% anumbers may not add up to 100% due to rounding bparticipants were asked to check all that apply so percentages will not equal 100% not all repositories are created equally curators are highly supportive of data sharing, which is not surprising since they work at data repositories. however, not all curators had experience curating qualitative data sets, let alone sensitive data sets. only 14/30 curators had personally curated a qualitative data set (table iii). while all curators worked at repositories that could accept qualitative data, some repositories had never actually received a qualitative data set. others were only capable of accepting data that could be made available open https://doi.org/10.29173/iq952 16/23 mozersky, jessica; walsh, heidi; parsons, meredith; mcintosh, tristan; baldwin, kari; dubois, james m. (2019) are we ready to share qualitative research data? knowledge and preparedness among qualitative researchers, irb members, and data repository curators, iassist quarterly 43(4), pp. 1-23. doi: https://doi.org/10.29173/iq952 access, prohibiting them from accepting sensitive qualitative data. this curator stated, ‘we’re not keeping that data if it’s sensitive right now’ [c7]. storing only non-sensitive data helps repositories avoid the challenges of anonymization and maintaining confidentiality, but it limits the number of repositories that qualitative researchers can use if they have sensitive data. as this curator noted, ‘either we can take it because it doesn’t have sensitive data, or we can’t take it because it does, and, frankly, that’s very unsatisfying’ [c28]. according to this curator: we will point you to curators at [another repository] who can do that for you … so that’s kind of where we were at…super-cautious, not really accepting anything that was a human-subjects base [c26]. some repositories only accept archival materials such as oral histories where the individual agrees to be identified, or supporting materials such as codebooks, interview guides, blank consent documents, summaries of project findings, or published articles that have been annotated for transparent inquiry, but not the human subject data itself (moravcsik, 2013). in contrast, curators at the few repositories who were more experienced with qds more clearly articulated an anonymization process and confidentiality protections through different layers of access to qualitative data. these curators described reviewing qualitative data sets carefully before deposit: we review all data before it’s deposited, so no transcript of an interview goes in the repository without at least one—typically two curators having read through it and checked for both direct and indirect identifiers…and we actually find that we probably, in three out of four cases, that we’ve sent transcripts back to depositors and asked for additional changes typically to indirect identifiers but sometimes even direct identifiers that slip through…that’s all during the preparation process of the initial deposit [c6]. another curator described the process as follows: every piece of information that involves human study subjects is thoroughly scrutinized by a group of people, including members of the institutional review board, the researchers in question, as well as an outside group of individuals within the university...[c10]. curators who have more experience with qds also noted that qualitative data is often only available via restricted access rather than open access in order to protect confidentiality. this curator noted: we have multiple layers of access and permissions before researchers can actually get at much of this data’ [c1]. another curator noted, ‘we go through, with all our data but particularly qualitative data to make sure there’s nothing identifiable from any in those data. many times, we can’t make those data publicly available anyway, so they have to be behind the restricted access barriers [c18]. those curators who were more experienced with accepting and anonymizing qualitative data sets were confident in their ability to protect confidentiality: i’m very confident. we take it very seriously, and we have a number of measures in place from depositing content, which is in a secure environment, to the people who work with the data here https://doi.org/10.29173/iq952 17/23 mozersky, jessica; walsh, heidi; parsons, meredith; mcintosh, tristan; baldwin, kari; dubois, james m. (2019) are we ready to share qualitative research data? knowledge and preparedness among qualitative researchers, irb members, and data repository curators, iassist quarterly 43(4), pp. 1-23. doi: https://doi.org/10.29173/iq952 internally, our curation staff. they’re trained. they also work in a very secure environment—to the application process for reusing the data, to the environment in which people reuse the data. [c20]. given the inability of many repositories to handle sensitive qualitative data, very few curators reported having qualitative-specific guidelines for depositing data. curators working at repositories that could not accept sensitive data did not necessarily perceive important differences between quantitative and qualitative data given that they could only accept open access data. this curator said, ‘we don’t specify anything differently for qualitative versus quantitative. i think the same rules would apply’ [c2]. according to this curator, ‘there are rules just for all data” [c4]. another curator stated: …we have a webpage that has some guidelines for preparing, for de-identifying data…our deposit form for the archive [requires] that data be de-identified. it doesn’t specify qualitative or quantitative since we’ll take any kind of data [c12]. some were unsure whether they had qualitative guidelines at all, as this curator says ‘my sense is that we have internal best practices. i don’t know that we have any published procedures’ [c11]. this curator noted, ‘we have a guide i think that has a section about qualitative studies’ [c20]. in contrast, curators at repositories with experience handling sensitive qualitative data reported having qualitative-specific guidelines to guide depositors. this curator noted: we are dedicated to only qualitative and mixed-method data, so anything that’s purely quantitative is out of scope for us. so all of our guidelines are in fact specific to preparing qualitative data [c6]. according to another curator: the [repository name] has got extensive [qualitative] guidance in detail…every archive has to handle the strategic challenge, basically, of whether or not it’s got the resources to support…a wide range of data types [c17]. even among those repositories that were prepared to handle sensitive qualitative data, some curators reported that they had never used the guidelines because they had never actually received a sensitive qualitative data set. as this curator stated: we have guidelines. they were developed in association with our social science research community. so, we have a nice policy framework for dealing with qualitative data, human subjects, that is sensitive. we’ve never actually needed to use it. no one’s submitted any of that kind of stuff yet [c21]. it’s the wild west: there are no agreed upon best practices while curators were the most familiar and comfortable with data sharing among all the stakeholder groups, curators desired more guidance or standards specific to qds. we created a lack of knowledge code among this stakeholder group to capture any indications of knowledge curators needed regarding legal, regulatory, or logistical issues. one individual commented, ‘we don’t have the tools and the resources to even understand… what is it we’re asking researchers to do in sharing their data’ [c28]. https://doi.org/10.29173/iq952 18/23 mozersky, jessica; walsh, heidi; parsons, meredith; mcintosh, tristan; baldwin, kari; dubois, james m. (2019) are we ready to share qualitative research data? knowledge and preparedness among qualitative researchers, irb members, and data repository curators, iassist quarterly 43(4), pp. 1-23. doi: https://doi.org/10.29173/iq952 curators often spoke of the lack of standards or national guidelines that specified best practices for qds. according to one curator: a lot of faculty and librarians and other curators, too, rely on…word of mouth in their communities. and we need voices of authority to say, ‘this is the situation,’ or…‘this is the best practice [c26]. lack of standards led some to worry that there was no qds oversight. monitoring compliance is difficult as there are no auditing bodies or standards for qds. as this curator commented: there's very little auditing that's happening, so there's a lot of assumptions about how things are happening, but very little auditing of what is actually being done by researchers…the policies are well thought out that exist when they exist, but the problem is that there's not a lot of eyes on the content of the data that's actually being shared to determine if the policies have been followed [c3]. when asked what resources would be helpful to them, many curators spoke of wanting better guidelines to guide qds. this curator described wanting tangible examples of what sharing sensitive qualitative data involves in practice: having some examples that are out there and prominent that we could point to say…this is what we mean when we say, ‘sharing sensitive data.’ this is how it actually took place…what was the workflow? what was the process?…to have some transparent examples to work with…and to help relate what these issues are and how they could actually translate into practice would be really helpful. [c28] curators also noted the challenges caused by variations in institutional, state, and national policies and laws. this curator noted, ‘these things vary by institutions…my university will likely have very different rules regardless of what the professional association or whatever says i should do’ [c25]. according to this curator, ‘on the legal and regulatory side … from country to country and from state to state…things vary’ [c23]. curators desired guidance on dealing with the variation in practices and lack of standards: one of the things that i think is a real challenge is that some of these regulations really vary state to state, even county to county and certainly institution to institution. so i wonder if a centralized service could provide the sort of nuanced level of guidance that would be necessary for answering all of these questions [c30]. responsibility and preparedness to advise researchers when asked how prepared they felt to advise researchers on legal and regulatory matters, the majority of curators were either not prepared or were unsure if they had the knowledge needed. many of the curators, even those who had curated a sensitive qualitative data set, did not feel prepared and mentioned that they would like to be more prepared on legal and regulatory matters pertaining to qualitative data and privacy laws. for instance, this curator who had personally curated a qualitative data set said: https://doi.org/10.29173/iq952 19/23 mozersky, jessica; walsh, heidi; parsons, meredith; mcintosh, tristan; baldwin, kari; dubois, james m. (2019) are we ready to share qualitative research data? knowledge and preparedness among qualitative researchers, irb members, and data repository curators, iassist quarterly 43(4), pp. 1-23. doi: https://doi.org/10.29173/iq952 i don’t feel very prepared…that’s one of my goals for this year is to get trained in the laws surrounding sensitive data in particular, especially with human subjects [c7]. another curator who had not personally curated a qualitative data set said: i feel like i could look up information or consult with other specialists on campus, but i don't feel like i personally have a really good grasp of that right now [c5]. this individual said, ‘i generally say, ‘well, let’s ask irb’…’ [c12]. curators working at repositories experienced in qds more clearly articulated the legal and regulatory responsibilities of repositories. this curator stated, ‘legally, the responsibility remains with the depositors, as per our deposit agreement with them’ [c23]. according to another curator: we're pretty careful when we talk to researchers about our repository to be clear about the function of it and who would have access. but it ultimately is up to the researcher to ensure that they're adequately protecting the subjects of their study. and the repository itself is not—is not protecting your subjects [c30]. according to this curator: we would generally have a policy of not advising people on legal matters. it’s up to them to find their own legal advice if they have those concern. we can only advise them on what our legal obligations are, and what our own limitations would be. but it’s kind of the policy we have not to act as legal advisors ourselves [c13]. as this curator noted: i’m not supposed to provide them with legal advice. i can give them literature. i can suggest that they strive to get legal advice, but i am not qualified to give legal advice, and i would not attempt to do so [c27]. discussion although we used sample sizes that are generally considered adequate for qualitative research studies (n=90 with 30 from each relatively homogenous stakeholder group), we cannot assume that our findings can be generalized to the larger populations of qualitative researchers, irb members, and data curators. nevertheless, as an exploratory study, the project points to important and concerning features of today’s research environment, in which data sharing is typically strongly encouraged and sometimes required. imagine for a moment that tomorrow the nih explicitly states that their data sharing policy applies equally to qualitative data and that qualitative researchers with nih funding are required to share data. would we be ready? https://doi.org/10.29173/iq952 20/23 mozersky, jessica; walsh, heidi; parsons, meredith; mcintosh, tristan; baldwin, kari; dubois, james m. (2019) are we ready to share qualitative research data? knowledge and preparedness among qualitative researchers, irb members, and data repository curators, iassist quarterly 43(4), pp. 1-23. doi: https://doi.org/10.29173/iq952 our results suggest that we are not ready to share qualitative data due to a lack of experience with, and guidance on, qds among all stakeholder groups. researchers are the least knowledgeable about qds. many were not aware that qualitative data repositories existed let alone that some repositories are capable of archiving and providing restrictions on who accesses sensitive qualitative data. while some qualitative researchers expressed willingness to share their data, the vast majority do not know how to share data responsibly or requirements of depositing qualitative data. this is particularly concerning given that many irb members feel ill-prepared to advise researchers on qds, and data curators feel that researchers have the obligation to protect their data and navigate legal and regulatory matters. many researchers also conveyed attitudinal barriers to the endeavor as a whole. future research should also gather evidence of benefits arising from data sharing, such as increased citations or secondary analyses, to increase researcher willingness. the majority of irb members, who many researchers would turn to for guidance, also lacked knowledge of how qualitative repositories operated or the protections they could offer, how to de-identify qualitative data, or who was responsible for ensuring legal and regulatory compliance. they expressed needing guidance for qds and were unprepared to guide researchers about most aspects related to qds. irb members frequently hoped repositories could provide guidance or indicated they would seek advice from other institutional officials like privacy officers or legal counsel. we strongly recommend that irb offices identify the appropriate institutional officials to advise on, and facilitate, qds, so they can provide researchers with timely referrals when the irb does not address qds. irbs have widely varied policies regarding informed consent forms and data sharing permissions. this variation in irb interpretations of federal research regulations is consistent with the findings of earlier studies (abbott and grady, 2011). the most incompatible language, which some consent forms still use, is the commitment to destroy data after a certain amount of time or explicit statements that data will not be shared. as a result, many qualitative data sets that have already been collected may not be shareable. in order to avoid this type of restriction going forward, it is essential that researchers and irb members consider data sharing from the outset and seek consent for data sharing when possible (meyer, 2018, dubois et al., 2018). qualitative researchers and irb members will need template language and guidance on the informed consent process and incorporating language about data sharing and reuse. irb members generally assumed that only de-identified qualitative data can be shared. while removal of the 18 hipaa safe harbor direct identifiers is relatively straightforward, removing indirect identifiers is far more complicated, and there are currently no adequate guidelines or tools for de-identifying qualitative data. standards and tools for de-identifying sensitive qualitative data are urgently needed. at the same time, the majority of repositories are not equipped to handle sensitive qualitative data, lack specific guidelines, and will therefore not be suitable for most qualitative researchers who collect data from human subjects. select experienced repositories have guidelines for sharing sensitive qualitative data and were confident about their repository’s ability to protect confidentiality and provide restricted access. irb members and researchers lacked knowledge of these options because they were unfamiliar with repositories in general. curators at these repositories articulated more clearly that they were not responsible for legal and regulatory compliance which they presumed came from researchers and their institutions or irbs. this leaves us in a challenging situation with both irbs and repositories desiring better guidance and neither feeling adequately prepared to advise researchers on legal and regulatory https://doi.org/10.29173/iq952 21/23 mozersky, jessica; walsh, heidi; parsons, meredith; mcintosh, tristan; baldwin, kari; dubois, james m. (2019) are we ready to share qualitative research data? knowledge and preparedness among qualitative researchers, irb members, and data repository curators, iassist quarterly 43(4), pp. 1-23. doi: https://doi.org/10.29173/iq952 matters. it might be helpful if data curators—even if reluctant to provide legal advice—provided researchers and institutions with examples of how sensitive, human subjects data have been shared (e.g., as limited data sets that require data use agreements). ethical and productive qds will require overcoming barriers, creating standards, and changing long held practices among all stakeholder groups. in the next phase of our project, we will develop tools and resources to support researchers as they begin to enter the uncharted territory of qds. we are developing a software to assist qualitative researchers in anonymizing qualitative data and guidelines on sharing qualitative data for all stakeholders. specifically, our software will provide the most comprehensive electronic method to date to search for, flag, and enable replacement of direct and indirect identifiers such as profession, setting, institution, or race. the software will support researchers but still requires their input and review before a determination can be made about sharing and access controls. we will then recruit 30 qualitative researchers to pilot our anonymization software and actually deposit their anonymized data for archiving with our project partners at the icpsr. our team will also develop a toolkit consisting of the finalized anonymization support software along with guidelines for sharing qualitative data that can be used by all stakeholders. we will disseminate the toolkit to stakeholder groups while evaluating the adoption of qds practices. by providing materials to support qds, our goal is to facilitate qds in an ethical manner and ensure that data are discoverable and usable by others. we do not believe that all qualitative data are appropriate for sharing (e.g., sensitive video footage that cannot be adequately anonymized, sensitive data that participants were told would not be shared, data that would place participants at serious risk of harm). at the same time, there are many advantages to qds, and many datasets could be shared in a responsible manner—with adequate preparation of the relevant stakeholder groups. references 2017. ‘hipaa privacy rule and public health guidance from cdc and the u.s. department of health and human services’*. in: services, d. o. h. a. h. (ed.). abbott, l. & grady, c. (2011) ‘a systematic review of the empirical literature evaluating irbs: what we know and what we still need to learn’, journal of empirical research on human research ethics, 6(1), pp. 3-19. antes, a. l., wwalsh, h., strait, m., hudson-vitale, c. r. & dubois, j. m. (2017) ‘examining data repository guidelines for qualitative data sharing’, journal of empirical research on human research ethics, online ahead of print, pp. 1-13. bishop, l. & kuula-luumi, a. (2017) ‘revisiting qualitative data reuse’, sage open, 7. borgman, c. l. (2012) ‘the conundrum of sharing research data’, journal of the american society for information science and technology, 63, pp. 1059-1078. corti, l. (2000) ‘progress and problems of preserving and providing access to qualitative data for social research —the international picture of an emerging culture [58 paragraphs]’, forum: qualitative social research, 1, art. 2. corti, l. (2012) ‘recent developments in archiving social research’, international journal of social research methodology, 15, pp.281-290. corti, l. & backhouse, g. (2000) ‘esrc qualitative data archival resource centre (qualidata), university of essex, uk’, forum: qualitative social research, 1. https://doi.org/10.29173/iq952 22/23 mozersky, jessica; walsh, heidi; parsons, meredith; mcintosh, tristan; baldwin, kari; dubois, james m. (2019) are we ready to share qualitative research data? knowledge and preparedness among qualitative researchers, irb members, and data repository curators, iassist quarterly 43(4), pp. 1-23. doi: https://doi.org/10.29173/iq952 dubois, j. m., strait, m. & walsh, h. (2018) ‘is it time to share qualitative research data?’, qualitative psychology, 5, pp. 380-393. guishard, m. a. (2017b) now’s not the time! qualitative data repositories on tricky ground: comment on dubois et al. (2017)’, qualitative psychology. hipaa journal. (2017) ‘what is a limited data set under hipaa?’ hipaa journal, available at: https://www.hipaajournal.com/limited-data-set-under-hipaa/ [accessed july 31 2019] jeng, w., he, d. & oh, j. s. (2016) ‘toward a conceptual framework for data sharing practices in social sciences: a profile approach’, proceedings of the association for information science and technology, 53, pp. 1-10. kuula, a. (2011) ‘methodological and ethical dilemmas of archiving qualitative data’, iassist quarterly, 34/35, pp. 12-17. mannheimer, s., pienta, a., kirilova, d., elman, c. & wuitch, a. (2018) ‘qualitative data sharing: data repositories and academic libraries as key partners in addressing challenges’, american behavioral scientist. meyer, m. n. (2018) ‘practical tips for ethical data sharing’, advances in methods and practices in psychological science, 1, pp.131-144. moravcsik, a. e., colin; kapiszewski, d. (2013) ‘a guide to active citation’, qualitative data repository (qdr) center for qualitative and multi method inquiry, syracuse university. national institutes of health. (2003) nih data sharing policy and implementation guidance. available at: http://grants.nih.gov/grants/policy/data_sharing/data_sharing_guidance.htm#funds [accessed: november 16 2018]. open access directory. (2016) data repositories index. available at: http://oad.simmons.edu/oadwiki/main_page [accessed]. saldaña, j. (2016) the coding manual for qualitative researchers. edited by jai seaman. 3 ed. sage publications. tenopir, c., allard, s., douglass, k., aydinoglu, a. u., wu, l., read, e., manoff, m. & frame, m. (2011) ‘data sharing by scientists: practices and perceptions’, plos one, 6, e21101. tsai, a. c., kohrt, b. a., matthews, l. t., betancourt, t. s., lee, j. k., papachristos, a. v., weiser, s. d. & dworkin, s. l. (2016) ‘promises and pitfalls of data sharing in qualitative research’, soc sci med, 169, pp. 191-198. yoon, a. (2014) ‘making a square fit into a circle: researchers’ experiences reusing qualitative data’,proceedings of the american society for information science and technology, 51, pp. 1-4. 1all authors are located at the bioethics research center, washington university school of medicine, st. louis, missouri, usa. correspondence to: james m. dubois, washington university school of medicine, division of general medicine sciences, 4523 clayton avenue, box 8005, st. louis, mo 63110; duboisjm@wustl.edu; tel: 1-314-747-2710. 2 this project was supported by grants from the u.s. national institutes of health, ul1tr002345 and r01hg009351. 3 the safe harbor method requires that the following 18 identifiers are removed: names, geographic divisions smaller than a state, dates, telephone numbers, vehicle identifiers, fax numbers, device/serial numbers, email addresses, web urls, social security number, medical record numbers, ip addresses, biometric identifiers, health plan numbers, full face photographs, account numbers, any other unique identifying number or characteristic, certificate/license numbers. https://doi.org/10.29173/iq952 https://www.hipaajournal.com/limited-data-set-under-hipaa/ http://grants.nih.gov/grants/policy/data_sharing/data_sharing_guidance.htm#funds http://oad.simmons.edu/oadwiki/main_page mailto:duboisjm@wustl.edu 23/23 mozersky, jessica; walsh, heidi; parsons, meredith; mcintosh, tristan; baldwin, kari; dubois, james m. (2019) are we ready to share qualitative research data? knowledge and preparedness among qualitative researchers, irb members, and data repository curators, iassist quarterly 43(4), pp. 1-23. doi: https://doi.org/10.29173/iq952 4 verbatim quotes published in this article have been edited to remove non-language text such as stutters, ums, and ahhs, to facilitate reading. 5 meyer (2018) offers conditions when data sharing may be ethically justified when the consent form is silent on data sharing: the form contains no statement that data will not be shared; data are not sensitive; and restrictions on use/access are in place. https://doi.org/10.29173/iq952 lassist newsletter, vol. 2, no. 1 (winter 1978) resodrce sharing through networks: problems and potentials for the social science community lorraine borman northwestern university abstract t tive gene that two purp to e user be t tool pone syst data such need tent his ter ral man mech ose : nsur com rans nt em, bas cap s, iai paper minal pur£0 y or anism 1) co munit mitte the s nto which e spe abili utili of ti overviews s for acce se user-or the probl s: lirst provide f ntinuing s v; and 2) q to a tu econd mech resource d would pro cific tuto ties, sy zation of mesharing th ssi ien ems eed yst pro or ani ata vid rm ste inf net e utiliz ng and m ted info inheren an inter back to em respo vide dat ial ser sm would bases e browsi g ana d m design ormation works wi ation anipu rmati t in actio the d nsi ve a on ving be t acces ng ca iagno will reso 11 no of co lating on sys such u n moni esigne ness t how th both he int sible pabili stic i not urces t be r mputer ne public d tem. t se could tor which rs of the o the va e system as a lear reduction by the c ties, co nguiry fa be truly will not eached. twor ks ata bas he maio be all would inform rying n is used ning a of a t omputer ntext-s cilitie respons occur. via 1 es th r pre eviat serve ation eeds whic nd re utori info ensit s. ive and nteracrough a raise is ed with a dual system of the h could farence al comrmation ive and without to user the pointroduction utilization of networking facilities to gain access to large scale computer systems and to public data base services has been increasing rapidly in the seventies. he might, in fact, call this period the age of the data base, with hundreds of data bases being generated by various agencies for public consumption. the extent and success of network access to these resources by social scientists may depend, however, on the unified efforts of computer and information scientists, educators, psychologists, linguists, data base producers, and network operators. that nsanot for cal, ingthe icasay rebethe osed redata ers, rier e. our thesis in thi while network access ble for many purpose grown to its fulles many reasons, some but others behaviora oriented. we will problems of hardware tions linkages here that they are rapi duced. the icformat havior and characte potential user commu of both experienced searchers, casual an base and computer may, however, be a to increased and sat s paper is is indispe s, it has t potential technologi 1 and learn not discuss and commun except to dly being ion seeking ristics of nity, comp and novice d frequent system us greater bar isfied usag the sharing of public data bases, accessed and manipulated by a general purpose data management system capable of providing transparent interfaces to specialized statistical, graphical and report generation packages, in a network environment, is the goal. a few problems immediately come to mind. the location of the terminal may be such that the user becomes separated from the computing center environment and from needed user aids. these may be reference materials, directories, and specialized technical documentation, or human contact with the computer operator, programmers, consultants, and subject specialists. automated system support should be available m an interactive mode, or on reguast from the terminal. the network itself can be a useful communication medium to provide a summary of historical information of interest to the user. user supp active syste ronment beco tant. some and off-line phone suppo mented, in all operatio problem are neumann (197 prepared for standards no of support c addressed in providing s needed on a basis, but a ering the ov of preservat storage capa neumann cont interactive rial design, copy docum printing, a of user fee areas, tut feeaback, h information group at th ort in auto ms in the n mes increas capabilit instructio rt have some degree nal networ as, thoug 3, p. 16) the nation tes that "p apabilities an integr upport when highly ind t the same erall syst ion of pro city." maj inues, sho language de integrat entation a nd further dback. t orial desig ave been st systems a e vogelbac mated etwor ingly ies r n an been , on ks. in a al bu roper need ated and ividu time em ec cessi or in uld f sign, ion o nd explo wo of n an udied nd s k co interk enviimporor ond teleimplealmost several emerge. report reau of design s to be manner , where alistic considonomics ng and terest , ocus on tutof hard on-line itation these d user by the ervices mputing 14 lassist newsletter, vol. 2, no. 1 (winter 1978) center, northwestern university. tutorial design is closely related to the problem of computer-aided instruction. a comment often heard (nickerson, 1969, p. 12) and true to a large extent, is that the need of the future is not so much for computer-oriented people as for peopleoriented computers. neumann (1973, p. 17) points out that many authors in the field of man-terminal interaction are still "computer-oriented" people and that there is a need for inter-disciplinary efforts including computer science, linguistics, psychology, and other human-oriented disciplines. the balance of this paper will look at work done at nortnwestern university in the area of system performance and user feedback and will relate tutorials and feedback mechanisms to the problem of network access and resource sharing of public data bases. the first barrier to successful utilization is the difficulty of acquiring the necessary knowledge required to use the system. marcus et al. (1971) working with the intrex system, nave made some general observations regarding user behavior: 1. users often fail to notice even the most explicit instructions. 2. there is not one single method applicable to all users. 3. if there are too many instructional options, they are all ignored; the user prefers to be given instructions only when needed. 4. users do not like to spend time in preparation for system use: they would rather use the system. ind 5. users are constrained bi previous experience ny training. there is barrier against learning anything new. 5. some users fear the machine, either because they think they will appear foolish, or because they fear they may damage the machine (machine fear) . some users do not want to ask for advice (people fear) . 7. users are overawed by the complexity of the system. and assume they need not or cannot understand the system. such attitudes impede learning. the question, though, is how to translate these generalizations into a user support system, and this cannot be accomplished until the designers of systems know how the system is being used by the end user^ which often is quite different rrom that envisioned by the designer. any user-oriented system must be responsive to the continually varying needs of the user community and to the different and differing levels of familiarity with the system. provision must be included in the network protocols and the information system to monitor the user/system interactions to generate feedback to the system designers. study of such usage data will enable, when necessary, intelligent redesign of the user/system interface, and will provide the necessary data for the development of the comprehensive on-line tutorials necessary for network access (borman and dominick in progress). the user cohhuniii we will define the anticipated user community of a nationally available educational network supporting access to social science oriented public data bases as composed of students, faculty, and staff of academic and research institutions. one example of such a network is edunet (edocoh, 1976). within this community will be experienced/inexperienced users, casual/frequent users, programming/nonprogramming users, and others. the diversity of application areas can range from analysis of voting behavior in a recent local election to economic forecasting based on data gathered na' e previoustionwide and spanning 20 years. n the network resouaces we project a network which provides access to not only public data bases but which will also place additional content requirements on the data base before it is accepted as a network resource. these would include: 1. comprehensive description of the data: its general content, its structure, and specific variable in-ana speci formation 15 lassist newsletter, vol. 2, no. 1 (winter 1978) 2. references to both bibliinteraction--are written into the ographic citations and to system based on or ogrammer dicprevious users of the tates. in actuality, we design and data base. implement interactive systems bsed on what we think they should ao, 3. a tutorial, context-orirather than on what the user would ented, developed by the expect them to do. proaucers of the data base and reflecting angiven this situation, and we beticipatory usage. lieve it to be widely true, the need for a feedback. mechanism belt. examples of use: recomes apparent. an interaction search selection, reportmonitor can provide various levels ing, graphical representof information. some have been imation of the data, etc. plemented to report system-oriented data: length of session, central 5. the data. processor and peripheral processor time per search, amount of disk in addition, the network respace used, job cost, etc. these sources would include; 1) an indata can then be used to evaluate formation system equipped with an the impact of the system on the tointeraction monitor and the necestal computing environment, on the sary linkages to process the tutouse of capabilities within the sysrial component of the data base, tem, etc. no monitor has yet been and 2) a simple and "understanding" implemented to capture data on how set of access protocols to allow the system is used by various types the user to enter the network, enof users with the idea that a feadter the inrormation system data back mechanism would be used to base, inquire about general informboth redesign for system efficiency ation such as charging, scheduling, and usability, and to provide the news events, etc. information necessary to produce a context-oriented, data base soecific, learning and reference tutorial. the concept of user tnt:2trictt0h hdttttdhttjtj" the 252sti0n of cai-type tutorials all interactive information systems are originally designed acthe potential utility of cai cording to some pre-established techniques within an information guidelines. these may include opsystem environment is well recogtimization for updating or searchnized and has been described by ing or reporting, or ease of exdominick and borman (1976). the portability from one hardware potential benefits can include: installation to another following the concepts of structured program1. author controlled and ming and modularity. .^ost will user controlled sealso proclaim their responsiveness quences. to the user--that oft cited phrase--user-oriented. 2. dynamic instructional strategies which can adthe system designer begins with just to the experience, preconceptions based on the existperformance, and informamg literature (almost negligible tion-seeking requirements in the area of users interacting of individual users, with numeric data bases and data management systems; larger, but of3. inclusion of various levten not more generally useful, reels of preprogrammed garding use of bibliographic data spelling algorithms and bases; and a few industrially orisynonym recognition which ented studies of the decision-maktend to give users the ing processes) . the aesigner, is illusion of at least a more often than not, not a user of minimum amount of system information systems. given the intelligence, varying pressures upon design time vs. implementation time, design given these potential benefits, usually loses. implementation prowhy have information system tutoriceeds, following the dictates of als failed to live up to expectathe easiest guiaelines— those contions? the problem areas include: nected with hardware, software, and economics. the goal of "user-ori1. tutorials function only ented" is nebulous, so features as stand-alone programs. such as language, diagnostic mesusers, interacting with sages, prompting— everything we the tutorial sequence, taink of as user/system cannot easily and 16 lassist newsletter, vol. 2, no. 1 (winter 1978) 4. immediately transfer control to the information system to try what they just learned, returning back to the tutorial whenever they encounter problems. most tutorials provide only extremely verbose or extremely terse information with minimal user interaction. although many tutorials employ the concepts of teaching by example, it is not clear how effective such training is if the examples are not data base specific. most tutorials rely on user comments concerning user problems, errors, etc. to provide feedback to system designers. these are not sufficient to enable knowledgable evaluation and interface redesign. riqs remote information a an i hens back the form al., ern hens for effi tion icl nte i ve m riq ati 19 uni i ve col cie s. ude: nitia racti tuto echan s sy on qu 76) . versi , au lecti ncy some 1 at on m rial ism stem ery deve ty, toma ng d and temp onit sys is syst lope con ted ata on the t to co or, a tem and exempli the re em (bor d at no tains a on-line both o user data c ordinate comprea feedfied in mote inman, et rthwestcompreaonitor n system interacollected 1. name and department of user 3. a. 5. 6. 7. name and size of aata base time required for entry of query time required to execute the query frequency of use of system commands and capabilities and types of full context of errors rrequency errors sixty five different elements of information are gathered for each 2uery and total session. three evels on monitoring are provided to avoid any invasion of data base privacy or user application privacy. the compleat tutor i tuto proc riqs tori file toma samp reco base othe tion t levi ousl brow sele prov one, brow data sele eral of t the and ei" aval sen a sa t tuto user of lang sear sear the to t and own base base whic lect when raand nost rial to a appr tion cont tion n ad rial esso tuto al; bro tic le f rds, pa r f s. dition in the r has r is mo it al wsmg c retrie iie des tae ssword r o n t e n to the form of been re than so prov apabilit val and cription enforcem specif i d proce mom a fr dev simpl ides ies, disp s and en t o catio ssing tor , a ont-end eloped, y a tupublic and aulay of sample f data ns and f unche ri ate s y lis sing, ctive ides rega se t base ction info he da cont appl es to espe lable ption mple utori rial s in the uaqe, ch ching inter he pr anal data f s, sa h ill ive ever th ics a and how opria and ext-s qstuto ome of ted. tutor ingui capabi rdless hrough s via us r ma tio ta bas ents a icatio whic cially are s of t record al mod for the s data b and strate sea nal li ocesso yze da bases or eac mple s ustrat inquir a user e info re coo error automa te err to al ensiti r has the p it h ial, s iities of ex ava simple ers ca sue e i a no su n area h the rele the n he dat from e repr initi yntax ase s prov id gies rching nkage r. us ta in or an li of t earche e comm y mod enter rmatio rdinat recov tic br or r ec low p ve tut tri robl as f earc bro wh peri ilab lis n s h as desc gges s a data vant ames a el the esen ally and yste es e and mod from ers eit i p he p s ar on u ed ems our hing wsin ereb ence le t or elec th ript ted nd ba an emen data m '3 xamp tip e pr the can her ubli ubii e pr sage is n sy ed ery anch over rese or ia hel stem with seg ing y in ntat 1 in to alprevimodes; , and g mode y any, can public menu t gene size ion of uses, discise may also d dets and base. a full a i n i n g antics query ies o_ s for ovides tutor search their c data c data ovided senvoked p comdiagtutouences to the formaion of f ormafull search by user text entered iftssist newsletter, vol. 2, no. 1 (winter 1978) summary and conclusions the concept of user ii.teraction moiiitoring. multi-mode tutoring, and d feeaback irechaiiism to provide inaividuai user, and individual references data base information, to the tutorial has been presented. in order borman. lorraine; chalice, robert; to interface a diverse user commadillaman, donald; dominick, nity utiiizin>-j a computer network wayne; and kobbe, ruth. riqs witn public data bases via sophisremote information smgry system, ticated processing systems, it ap^evanston, til.: vogelback compears that such a tnree-way interputing center, northwestern unirace system must be aaopted. the versity, 1976) monitor, serving both as a recording device and a feedback mecha, and dominick, wayne. an nism, can ensure system responsiveanalvsis of data captured by^ an ness to the varying neeas of the ih^eraction monitor. national user community; tne tutorial can science foun^aeion grant # di3 provide file browsing capabilities, 75-19u81, in progress. northcontext-sensitive training, selecwestern university, evanston, tive inquiry facilities linked to 111. diagnostics, and internal linkages to the data base manageeent system dominick, wayne and borman, loritself. only by providing such raine. user/system interfacing personalized, and necessary, servin an interactive retrieval, ices will utilization o* public statistical and jraphical analydata bases via time-sharing netsis environment. proceedings of works reach their full potential. the 5th asis mid-year heeting. vandereile university, !iasnville, tn, may, 1976, pp. 34-49. educom. description of edunet. bulletin of the interuniversity communica'eions councirn^educdhf . tt-t17~fart7"t975r7 holmes, d. c. comauters in oil 1967-1987. computer yearbook and ^isectory "{2na 113.) je'd. f. h. srille]" "jdalroif: american data processing) marcus, r. s.. ; beuenfeld, a. r,; and kugel, p. the user interface tor the intrex retrieval system. ihtaractiye bibliograahic searcn (ea. d. 'e. wallferr thonlvale, nj: afips press, 1971) , 159-201 . nickerson, r. s. man-computer interaction: a challenge for human factor research. ergonomics 12, 4 (july, 1969). neumann, a. j. network user information support national bureau of standard tecnnical no^e bwt. (wasnington , d. c: u. j. government printing office, 1973) .. 1/16 swygart-hobaugh, mandy (2019) bringing method to the madness: an example of integrating social science qualitative research methods into nvivo data analysis software training, iassist quarterly 43(2), pp. 1-16. doi: https://doi.org/10.29173/iq956 bringing method to the madness: an example of integrating social science qualitative research methods into nvivo data analysis software training mandy swygart-hobaugh1 abstract it is not uncommon for researchers who wish to delve into qualitative data analysis to be lacking in qualitative methods training. data professionals who support these aspiring qualitative researchers are well positioned to recognize and develop resources, training, and services to address this methods gap. this article describes a specific training session aimed at bridging this gap: a collaboration between a sociology professor and the author that integrates a qualitative methodological framework with specific features of nvivo qualitative data analysis software that complement and facilitate research guided by that framework. this article (1) outlines how this collaboration came to be; (2) describes the roles that the sociology professor and the author play in the collaboration, including specific examples from the training session; and (3) offers a reflection on the experience, including successes and growth possibilities going forward. keywords qualitative research, qualitative data analysis, theorizing, nvivo qualitative data analysis software, data services workshops, data services support and training introduction i want to bring us to that awesome point where you have collected the closetsfull [sic] or tons of data and then you have to do something with them. you face the terrible moment when you want to leave this mortal coil because you are wondering what all of this awful buzzing confusion called “the data” can possibly mean. . . . what is the story? what is in there? —fred davis2 in the summer of 2017, i and two other members of iassist’s qualitative social science and humanities data interest group co-authored a four-post series for iassist’s iblog discussing how data-support professionals can assist researchers who wish to delve into qualitative research but are lacking in methodological training – what we referred to as the ‘qualitative methods gap’ (swygart-hobaugh et al., 2017). in the third post of the series, i highlighted how data-support professionals ‘can draw on their own expertise and also partner with qualitative researchers on campus to offer presentations, workshops, brown bags, etc., aimed at addressing the qualitative methods gap’ (swygart-hobaugh, 2017). i cited a specific example of this type of collaboration: a training session given by a sociology professor and me that bridges a qualitative methodological framework with specific features of nvivo qualitative data analysis software. this article (1) outlines how this collaboration came to be; (2) describes the roles that the sociology professor and i play in the collaboration, including specific examples from the training https://doi.org/10.29173/iq956 https://sites.google.com/uncg.edu/iassistqsshdig/home https://sites.google.com/uncg.edu/iassistqsshdig/home https://iassistdata.org/blog/how-do-i-do-qualitative-research-bridging-gap-between-qualitative-researchers-and-methods-resou https://iassistdata.org/blog/how-do-i-do-qualitative-research-bridging-gap-between-qualitative-researchers-and-methods-res-1 2/16 swygart-hobaugh, mandy (2019) bringing method to the madness: an example of integrating social science qualitative research methods into nvivo data analysis software training, iassist quarterly 43(2), pp. 1-16. doi: https://doi.org/10.29173/iq956 session; and (3) offers a reflection on the experience, including successes and growth possibilities going forward. the birth of a collaboration in august 2014 at the invitation of qsr international, the company that makes nvivo, i created and delivered the webinar, ‘using nvivo 10 for windows for sociological qualitative data analysis’ (swygarthobaugh, 2014). unbeknownst to me, dr. ralph larossa, a georgia state university emeritus professor of sociology, was in attendance. dr. larossa has written a number of articles advancing qualitative methodological inquiry. we had previously conversed about the pros and cons of using computer assisted/aided qualitative data analysis software (caqdas), such as nvivo, atlas.ti., dedoose, etc., at which time he had yet to take the plunge into using caqdas. dr. larossa emailed me to ask whether i concurred that certain aspects of nvivo i had discussed in my webinar lent themselves to advancing a particular methodological framework that he had presented in previous publications and has continued to develop (larossa, 2005, 2012a, 2012b). his ‘levels of theorizing in qualitative analysis’ framework (larossa, 2012b: 649–654), which i will discuss in more detail later in this article, delineates a model for moving qualitative research beyond description into hypothesis and relationship explorations. after reading his articles, i replied that i wholeheartedly concurred that his framework readily linked to the nvivo features i had presented, and that i saw potential for even more nvivo features realizing his methodological framework. i also mentioned that during one-on-one consults and workshops, i sometimes witnessed a paucity of qualitative methods knowledge among aspiring nvivo users, and that i planned to share his articles with campus researchers as a possible framework that might demystify the process of qualitative data analysis. i then proposed that we do a joint presentation: he presenting his methods for qualitative research, and i presenting the mechanics of using nvivo to implement his methods. he replied that he was indeed interested; various meetings and discussions ensued to develop our joint presentation, with the aim of debuting it that following spring semester. interweaving the logics (qualitative method) into the logistics (nvivo mechanics) the following is the title and synopsis of our joint presentation as described in the event posting and promotional communications: the logics and logistics of qualitative research: a framework for exploring concepts, dimensions, and relationships in qualitative data using nvivo research software in this presentation, dr. ralph larossa, professor emeritus of sociology, and dr. mandy swygarthobaugh, librarian associate professor for sociology & data services and team leader for research data services, present both the theoretical-methodological logics and the appliedmethodological logistics of conducting qualitative data analysis (i.e., non-statistical analysis of textual, audio, visual, and/or audiovisual sources). dr. larossa discusses the steps involved in building theoretically-rich qualitative analyses (the logics). dr. swygart-hobaugh outlines the specific features of nvivo qualitative research software that complement and facilitate these analyses (the logistics). our first iteration was presented in march 2015, and we have given the presentation four subsequent times (in october of 2015, 2016, 2017, and 2018), continuously revising toward improvement and to incorporate changes with newer nvivo versions (cycling from version 10 to 11 pro to the current 12 plus https://doi.org/10.29173/iq956 https://www.youtube.com/watch?v=vfqaw61o0rg 3/16 swygart-hobaugh, mandy (2019) bringing method to the madness: an example of integrating social science qualitative research methods into nvivo data analysis software training, iassist quarterly 43(2), pp. 1-16. doi: https://doi.org/10.29173/iq956 version). attendees span the various social sciences plus fields such as business, nursing, and public health. while attendees were predominantly georgia state university affiliates, researchers from other metro atlanta universities, non-profit organizations, and government agencies have attended as well. the presentation is 90-minutes in length, delivered via powerpoint slides, and divided into three parts: (1) a lecture by dr. larossa on his methodological framework, (2) a demonstration by me of relevant nvivo features via screenshots, and (3) question and answer time. attendees receive a handout with our contact information, a list of the readings cited in the presentation, and a link to my nvivo online help guide that includes information about upcoming hands-on nvivo workshops (appendix 1). while dr. larossa draws from several of his own publications and those of others in his lecture, he primarily focuses on explicating his ‘levels of theorizing in qualitative analysis’ framework (larossa, 2012b: 649–654). i have crafted my portion of the presentation to demonstrate specific nvivo features that facilitate the framework that he presents. dr. larossa uses examples from his own research on adviceseeking letter exchanges between parents and the educator and author dr. angelo patri between the mid1920s to the late 1930s (larossa and reitzes, 1993). this dataset consists of approximately 900 letter exchanges in text format and is available from the icpsr data repository (larossa, 2009). i downloaded and imported the letter exchanges into nvivo to use for my demonstration slides in order to make direct connections between examples given in our respective parts of the presentation.3 1. larossa’s framework: levels of theorizing in qualitative analysis this section of the article contains a summary of dr. larossa’s overall framework and then three subsections that correspond to the three ‘levels of theorizing’ he proposes. in each subsection, i briefly summarize his conception of the level and how he presents it in our joint presentation, and then i demonstrate how i tie aspects of the level to specific nvivo features in my portion of the presentation. drawing from his years of experience as a reviewer and as deputy editor for the journal of marriage and family, larossa proposes that qualitative research generally operates along three dimensions: . . . [w]e can visualize a world in which every qualitative manuscript can be plotted by its latitude (where it is with respect to the humanities and sciences), longitude (where it is with respect to the length and number of data excerpts), and altitude (where it is with respect to the level of theorizing). (2012b: 644–645) the ‘altitude’ dimension is where larossa expands on his ‘levels of theorizing’ framework. in his conceptualization, ‘theorizing’ relates to whether a qualitative manuscript should 'be essentially descriptive, with a set of concepts and/or variables presented but with not much said about the possible relationships among those concepts and/or variables,’ or whether it should move to ‘a higher level of abstraction, talk about variable relationships and make, as its centerpiece, the development of hypotheses’ (larossa, 2012b: 648). he then demarcates qualitative analysis theorizing into three levels, including a diagram reminiscent of three stair steps to illustrate the upward movement to higher levels of theorizing (larossa, 2012b: 650, figure 4). 1.1 level 1: concept formation – denoting and connoting as conceptualized by larossa, at level 1 a qualitative researcher ‘denotes’ text segments (in manual coding, by underlining them with a pencil, or highlighting them with a marker, etc.) ‘with the purpose of https://doi.org/10.29173/iq956 4/16 swygart-hobaugh, mandy (2019) bringing method to the madness: an example of integrating social science qualitative research methods into nvivo data analysis software training, iassist quarterly 43(2), pp. 1-16. doi: https://doi.org/10.29173/iq956 singling out text segments that are thought to be important’ and then ‘connotes’ the denoted text segments by ‘linking’ them with an analytical concept to which that segment speaks or resonates (larossa, 2012b: 649). the demarcated text segments are ‘indicators’ in grounded theory methods terminology (larossa, 2005). larossa (2005) argues that when engaging in concept-indicator formation – more broadly termed coding – researchers move along the dialectic path between deduction (their analytical framework and/or research questions inform their gleaning of concept-indicator connections) and induction (they let concept-indicator connections emerge from the data without being fettered by preconceived assumptions/frameworks/research questions). level 1 theorizing (concept-indicator coding) can be understood as the foundation of all qualitative data analysis. in his presentation of level 1, dr. larossa offers examples of denoting and connoting drawn from his own research on the patri letter exchanges. he uses a figure (larossa, 2012b: 650, figure 3) to illustrate how an individual text segment can be denoted (underlined) and then connoted at multiple concepts, be they (1) a priori, from an existing, recognized theory/framework, (2) in vivo, reflecting verbatim the text in the analyzed materials, or (3) juxta vivo, a middle-ground between a priori and in vivo. to demonstrate to researchers the utility of qualitative research software for level 1 theorizing, i first present coding using nvivo nodes as a means for researchers to engage in denoting and connoting (figure 1). figure 1: using nvivo nodes to denote and connote (or code) text. i present nodes as akin to concepts and coded text segments (or nvivo coding references) as akin to indicators as described by dr. larossa. i explain that the basic mechanics of coding in nvivo entails first highlighting the text segment (denoting it as an indicator), then right-clicking it to then either code it at an existing node or a new node (connoting it at an existing or new concept). i also note that in software such as nvivo you can readily code a single text segment/indicator at multiple concepts if it is analytically https://doi.org/10.29173/iq956 5/16 swygart-hobaugh, mandy (2019) bringing method to the madness: an example of integrating social science qualitative research methods into nvivo data analysis software training, iassist quarterly 43(2), pp. 1-16. doi: https://doi.org/10.29173/iq956 warranted – a task that proves more challenging when doing manual coding. for example, in the letter pictured in figure 1 above, the mother in her opening sentence writes, ‘i am in trouble and i know it is my fault’. the researcher can denote that sentence by highlighting it, then connote it by right-clicking the highlighted text segment and check-boxing the existing nodes of ‘faulting oneself’ and ‘being in trouble’. i also describe how you can double-click on a node to then easily access data files and the corresponding text segments coded within data files at that node and other nodes (by using the nvivo coding stripes feature), which allows the researcher to easily move through their data and see similarities/differences of concepts across the different data files that would, again, prove more challenging when coding manually. for example, looking again at the figure 1 screenshot above, the researcher can double-click the ‘faulting oneself’ node to see a listing of the 104 individual text segments/references coded at that node parsed by the 62 letters/files in which those text segments appear, which will allow the researcher to delve into the nuances of how that concept emerges in the letters and how it diverges/converges between and within letters. i also point out the utility of nvivo tallying the number of coding references (text segments/indicators) coded at a node, as a researcher can cite these counts to argue that a node/concept with ‘numerous indicators [is] theoretically saturated’ (larossa, 2005: 846). i next discuss how researchers can use nvivo queries to tease out possible concept-indicator connections. for example, i demonstrate a word frequency query that returns the top 25 most frequently occurring words and their stemmed variations across the patri letter exchanges (figure 2). figure 2: using an nvivo word frequency query to explore potential concept-indicator connections. https://doi.org/10.29173/iq956 6/16 swygart-hobaugh, mandy (2019) bringing method to the madness: an example of integrating social science qualitative research methods into nvivo data analysis software training, iassist quarterly 43(2), pp. 1-16. doi: https://doi.org/10.29173/iq956 i note that it grabbed my attention that boy/boys/boys’ is the second most frequently occurring word, while girl/girls/girls’ does not even make the top 25. i show how i can double-click on boys in the word frequency list to jump to a page that parses out the occurrence of the word across the data files, showing boy/boys/boys’ with the five preceding words and the five succeeding words to give brief context for its use but allowing me to click on a hot-linked data file name to jump to that specific file and see the word used in its full context. i then pose the following analytical questions we might explore: what interpretations might we draw from boy/boys being the second most frequent word grouping, appearing 2586 times, and girl/girls not even making the top 25? is there something about the discussion of boys (and, by association, girls) that alludes to a concept-indicator connection we could code for? does it point to the possibility that the gender of the child plays a critical role in the data? are boys perceived as more trouble (thus necessitating more advice seeking regarding their rearing)? are boys held as more important than girls and thus more worthy of seeking advice about? what story might we tell with this finding if we dig deeper? for my last example of level 1 theorizing i demonstrate how a researcher can use a text search query to explore possible connections between searched words and a concept – i.e., to explore possible conceptindicator connections (figure 3). figure 3: using an nvivo text search query to explore potential concept-indicator connections. for example, i might be curious about how the word ‘help’ and similar words are used in the patri letter exchanges. i can do a text search query across the data files using a boolean logic search for help or aid or assist (and their stemmed variations) to drill down to instances of these words in the letter exchanges. then i can explore whether the word usage indicates a concept i should create and code for in the letters. for example, i might create a concept node called ‘father helps’ and code text segments at this node https://doi.org/10.29173/iq956 7/16 swygart-hobaugh, mandy (2019) bringing method to the madness: an example of integrating social science qualitative research methods into nvivo data analysis software training, iassist quarterly 43(2), pp. 1-16. doi: https://doi.org/10.29173/iq956 when it is implied that mothers hold primary responsibility for rearing children, but fathers can ‘help’ when/if they choose. 1.2 level 2: variable formation – dimensionalizing in his article and in his part of our joint presentation, larossa (2012b) proposes that qualitative researchers move to level 2 theorizing when they bring variables or dimensions of concepts into their consideration and analysis of the data. returning to his own research on the patri letter exchanges, dr. larossa shows how the in vivo concept of ‘faulting oneself’ for the trouble a mother has feeding her baby becomes a variable/dimension when the researcher begins to array possibilities within the larger concept, such as ‘extent of faulting oneself’ or ‘blame taking (yes or no)’ (larossa, 2012b: 653, figure 4). he also mentions how qualitative researchers can incorporate relevant variables into their analysis such as age and marital status (larossa, 2012b). for example, my proposing that the child’s gender might have some bearing on a parent seeking child-rearing advice from patri illustrates that, even while at the first level of theorizing, i was already moving toward level 2 variable/dimension formation by injecting gender-of-child as a variable. in my presentation of using nvivo for level 2 variable/dimension formation, i begin with demonstrating how researchers can create hierarchical node structures (what nvivo calls parent nodes and child nodes) to represent variability/dimensions of a concept (figure 4). figure 4: creating nvivo parent and child node hierarchies to code for variables/dimensions. for example, i created a parent node for the concept of ‘blame taking’ and created child nodes of ‘blaming others’ and ‘blaming self’ as dimensions of that concept. if analytically warranted, i could create deeper node hierarchies to represent dimensions within the child nodes (e.g., under ‘blaming others’ i might add child nodes of ‘blaming spouse’ or ‘blaming child’ or ‘blaming in-laws’, etc.). i discuss the option of https://doi.org/10.29173/iq956 8/16 swygart-hobaugh, mandy (2019) bringing method to the madness: an example of integrating social science qualitative research methods into nvivo data analysis software training, iassist quarterly 43(2), pp. 1-16. doi: https://doi.org/10.29173/iq956 aggregating coding from child nodes, which allows the researcher to code at a child node but then have the coding automatically applied to the parent node (or grandparent node and on up, if applicable). i also highlight nvivo’s flexibility and ease of renaming, rearranging, and merging nodes within and across hierarchies when analytically warranted – a process of ‘fractur[ing]’ or ‘reconstitut[ing]’ concepts that is important in grounded theory methods (larossa, 2005: 846). i next discuss nvivo’s classification sheet feature as a means to define variations in attributes (akin to variables) about the units of analysis, which then can be used for analytical comparisons and/or variablerelationship formation (the next level of theorizing in larossa’s framework). i display a screenshot of an nvivo classification sheet i created to classify attributes about the patri letter exchanges (figure 5). figure 5: using nvivo classification sheets to define attributes/variables for data units. i explain that, similar to an spss statistical file, the rows of an nvivo classification sheet represent the individual data units (e.g., a patri letter exchange data file), and the columns contain attribute/variable data that has been collected and compiled for the corresponding individual data units (e.g., gender of the letter writer, year of the letter, where the letter writer lived within the united states, the gender of the child discussed in the letter, etc.).4 1.3 level 3: variable-relationship formation – typologizing and hypothesizing dr. larossa in his article (2012b) and his part of our joint presentation describes level 3 of his framework as the level at which qualitative researchers explore relationships between the variables and/or dimensions they developed at previous levels. larossa (2012b) argues that this variable-relationship formation takes two forms: typologizing and hypothesizing. when typologizing, ‘two or more variables are cross-listed to create a matrix of cells’ (larossa, 2012b: 653), which allows the researcher to explore the relationship of different variables/dimensions by examining the intersections of their differing https://doi.org/10.29173/iq956 9/16 swygart-hobaugh, mandy (2019) bringing method to the madness: an example of integrating social science qualitative research methods into nvivo data analysis software training, iassist quarterly 43(2), pp. 1-16. doi: https://doi.org/10.29173/iq956 components. larossa posits hypothesizing (proposing correlations and/or cause-and-effect relationships between variables) as another aspect of higher-level theorizing in qualitative research, while recognizing that it is ‘especially tricky in qualitative research’ due to its traditional association with positivist/quantitative research epistemologies that many qualitative researchers approach with skepticism if not outright reject (2012b: 653). he poses that while many qualitative researchers are wary of hypothesis testing due to its adherence to positivist/quantitative tenets and assumptions, hypothesis development – ‘offer[ing] ‘‘plausible suggestions’’ about—not definitive tests of—variable relationships’ – is likely more palatable, and is readily attainable and desirable in order for qualitative researchers to move to level 3 theorizing (larossa, 2012b: 654, see also larossa 2012a). in his part of the presentation, dr. larossa departs from the patri letter exchanges to present the level 3 variable-relationship formation paradigm, using examples from his other research endeavors (larossa and larossa, 1981; larossa and sinha, 2006). i return to the patri letter exchanges to offer several nvivo features that facilitate typologizing and hypothesizing to examine variable-relationship formation: ● using crosstab queries and matrix coding queries to examine the intersections of variable/dimension components (typologizing) or to explore hypothesized relationships between variables/dimensions (hypothesizing) ● using coding queries to examine when text segments have overlapping coded concepts to examine potential relationships between those concepts (e.g., is there a relationship between letter writers’ faulting themselves and feeling ashamed?); creating and coding those text segments as a relationship (e.g., faulting oneself influences feeling ashamed); and creating a chart (under nvivo explore menu) of the relationship broken down by a selected attribute’s values to examine whether the attribute has a moderating influence on this relationship (e.g., is the faulting-oneself-influences-feeling-ashamed relationship more prevalent in women letterwriters over men letter-writers?) ● using nvivo visualizations (under nvivo explore menu) to explore potential relationships: o generating a cluster analysis diagram that clusters selected nodes/concepts together if they code many of the same files (e.g., explore whether files coded at the ‘faulting oneself’ node cluster by coding similarity with files coded at the ‘being ashamed’ node) o generating a comparison diagram of two nodes/concepts that visually displays what files they share and do not share in terms coding (e.g., which files are coded at both ‘faulting oneself’ and ‘being ashamed’ and which are coded at only one or the other of the concepts) for the sake of brevity – and because i believe it to be the most clear-cut and powerful of the aforementioned examples – i will expound upon my demonstration of an nvivo crosstab query as a means for both typologizing and hypothesizing. in my presentation i return to the node/concept of ‘blame taking’ and its dimensions (child nodes) of ‘blaming self’ and ‘blaming others’ and pose this question: is there a relationship between the gender of letter writers and whether they are taking or assigning blame for a child’s issues? i propose that one way to examine the relationship of the gender-of-letter-writer variable (or attribute in nvivo lingo) with the ‘blame taking’ concept variable/dimension is to run a crosstab query (figure 6). https://doi.org/10.29173/iq956 10/16 swygart-hobaugh, mandy (2019) bringing method to the madness: an example of integrating social science qualitative research methods into nvivo data analysis software training, iassist quarterly 43(2), pp. 1-16. doi: https://doi.org/10.29173/iq956 figure 6: using an nvivo crosstab query for variable-relationship formation. i explain that the created crosstab matrix contains the two rows for the different dimensions of the ‘blame taking’ concept and two columns for the gender-of-letter-writer attribute values, and that the number displaying in each cell is the number of associated data files (letter exchanges) that have at least one text segment (indicator) coded at the associated ‘blame taking’ concept dimension (or child node). i stress that, rather than relying on these numbers alone for drawing conclusions (although they do in and of themselves give compelling evidence for a relationship), as qualitative researchers they should explore the coded text segments to delve deeper into, for example, variations in the language/themes used by men versus women when discussing blaming self or others – and accessing the associated text segments is easily achieved by just double-clicking the matrix cell. i then bridge this feature to dr. larossa’s typologizing and hypothesizing by presenting the following: out of the 55 men’s letters with blaming discourse, 44 men’s letters had one or more text segments coded to the ‘blaming others’ node, and only 11 had one or more text segments coded to the ‘blaming self’ node. in contrast, we see the reverse in the women’s letters, with only 14 in the ‘blaming others’ cell and 51 in the ‘blaming self’ cell. given this evidence, we might pose the following variable-relationship hypothesis: among letter writers that engaged in blaming discourse, women were more likely to blame themselves than others, and men were more likely to blame others than themselves. throughout my demonstration, i continuously stress that nvivo does not do the intellectual work for the researcher, nor does any other qualitative data analysis software/tool. i tell them there is no magic button to push to have nvivo produce definitive findings. i stress that they still have to become intimately familiar with the data through reading and re-reading – i.e., as qualitative researchers committed to probing the rich, nuanced meaning within people’s words, they should not just text mine and call that ‘analysis’. i https://doi.org/10.29173/iq956 11/16 swygart-hobaugh, mandy (2019) bringing method to the madness: an example of integrating social science qualitative research methods into nvivo data analysis software training, iassist quarterly 43(2), pp. 1-16. doi: https://doi.org/10.29173/iq956 emphasize that nvivo is primarily a data management and organization tool that facilitates analysis – but the true analysis occurs within researchers’ minds where they draw out the meaning from the data guided by their analytical and methodological frameworks and research questions. reflections and going forward overall, both dr. larossa and i deem this collaboration a success. verbal feedback from attendees is consistently positive, and we often hear from attendees that a colleague who had attended a previous iteration referred them. a comment in response to an evaluation question – ‘what is the most useful thing you learned at the workshop?’ – aptly resonates with the point of the session: i got a great rationale which can link theory to a practical research tool. as there is always room for improvement, below are growth possibilities for future iterations: ● gather more official feedback – while we distribute an online evaluation form to attendees after each presentation, we get very low response rate. one way to improve the response might be to hand out a short printed form instead and have attendees complete it before they leave. ● increase turnout – although our registration is often full at 50 registrants (and with a waitlist), we average around 50% of registrants actually coming. while this low turnout rate is consistent with that of our other research data services workshops, as the leader of the research data services team i am working with my teammates on strategies to improve turnout. ● consider adapting the nvivo portion of the joint session to be a live demonstration, and possibly hands-on – using screenshots to demonstrate nvivo features proves efficient in terms of allowing coverage of more content in a shorter window, permitting visual annotations to facilitate understanding, and avoiding the risk of technological failures. that said, live demonstration and hands-on work would likely be more engaging for attendees and increase learning and retention of the nvivo mechanics presented. however, implementing live demonstration and/or hands-on activities would necessitate significantly more time for the session. currently i encourage attendees to attend my upcoming hands-on nvivo workshops in which i cover many of the features presented. but, perhaps having a multi-part series (or maybe even a webinar series) could be warranted, wherein specific aspects of the levels-of-theorizing framework would be the entire focal point of each individual workshop. weighing the pros/cons of our current delivery method against adding live/hands-on elements is something dr. larossa and i could explore further. after each offering dr. larossa and i say to each other that we both learned more about qualitative methodology and qualitative data analysis software from the experience. in addition, our collaboration has bred collaboration outside of the presentation, often taking the form of “i saw this article and thought of your/our work” scenarios. my work supporting qualitative researchers on the georgia state university campus has unquestionably benefited from this experience. a key benefit of working closely with a widely published and respected methodologist on developing a method-embedded qualitative data analysis software training is that researchers get a clearer sense of how method should be the driving force behind their use of analysis software. my regularly offered workshops are primarily consumed by teaching the mechanics of using nvivo; as such, any integration of methods on my part is often tertiary or surface-level. this joint presentation with dr. larossa allows me to articulate mindfully how method can, should, and must guide https://doi.org/10.29173/iq956 12/16 swygart-hobaugh, mandy (2019) bringing method to the madness: an example of integrating social science qualitative research methods into nvivo data analysis software training, iassist quarterly 43(2), pp. 1-16. doi: https://doi.org/10.29173/iq956 the use of the software’s mechanics. qualitative researchers wedded to manual coding often cite their fear that software such as nvivo, atlas.ti, and dedoose removes the researcher from the analysis process and ‘does all the work’ as their reason for shunning it. this joint presentation makes headway in alleviating these fears, as it demonstrates how qualitative data analysis software is merely a conduit via which the researcher applies their chosen theory and method – it is not the theory or method in and of itself. similarly, aspiring qualitative researchers often learn method divorced from its application. this bridged method-mechanics training session offers a unique space to remove methods training from the vacuum in which it typically is taught. in conclusion, i hope that this example experience of integrating qualitative research methods into qualitative data analysis software training can serve as a model, or at least an inspiration, for those who offer training and support for existing and aspiring qualitative researchers. if anything, it demonstrates that qualitative researchers are eager and grateful for knowledge and skills that help them bring method to the madness that is qualitative data. acknowledgment i would like to thank dr. ralph larossa for his contributions to our joint presentation, his feedback on this article, and his continuing commitment to developing and teaching qualitative methodological inquiry that is rigorous, nuanced, and theoretically-rich yet still graspable. references larossa, r. 2005, “grounded theory methods and qualitative family research”, journal of marriage and family, vol. 67, no. 4, pp. 837–857, https://doi.org/10.1111/j.1741-3737.2005.00179.x. larossa, r. 2009, parenthood in early twentieth-century america project (petcap), 1900-1944, ann arbor, mi: inter-university consortium for political and social research [distributor], available at: https://doi.org/10.3886/icpsr06876.v2. larossa, r. 2012a, “thinking about the nature and scope of qualitative research”, journal of marriage and family, vol. 74, no. 4, pp. 678–687, https://doi.org/10.1111/j.1741-3737.2012.00979.x. larossa, r. 2012b, “writing and reviewing manuscripts in the multidimensional world of qualitative research”, journal of marriage and family, vol. 74, no. 4, pp. 643–659, https://doi.org/10.1111/j.1741-3737.2012.00978.x. larossa, r. and larossa, m.m. 1981, transition to parenthood: how infants change families, sage publications, beverly hills, ca. larossa, r. and reitzes, d.c. 1993, “continuity and change in middle class fatherhood, 1925-1939: the culture-conduct connection”, journal of marriage and family, vol. 55, no. 2, pp. 455–468, https://doi.org/10.2307/352815. larossa, r. and sinha, c.b. 2006, “constructing the transition to parenthood”, sociological inquiry, vol. 76, no. 4, pp. 433–457, https://doi.org/10.1111/j.1475-682x.2006.00165.x. https://doi.org/10.29173/iq956 https://doi.org/10.1111/j.1741-3737.2005.00179.x https://doi.org/10.3886/icpsr06876.v2 https://doi.org/10.1111/j.1741-3737.2012.00979.x https://doi.org/10.1111/j.1741-3737.2012.00978.x https://doi.org/10.2307/352815 https://doi.org/10.1111/j.1475-682x.2006.00165.x 13/16 swygart-hobaugh, mandy (2019) bringing method to the madness: an example of integrating social science qualitative research methods into nvivo data analysis software training, iassist quarterly 43(2), pp. 1-16. doi: https://doi.org/10.29173/iq956 swygart-hobaugh, m. 2014, using nvivo 10 for windows for sociological qualitative data, available at https://www.youtube.com/watch?v=vfqaw61o0rg (accessed 13 january 2019). swygart-hobaugh, m. 2017, “but how do i *do* qualitative research? bridging the gap between qualitative researchers and methods resources--part 3”, in iassist iblog, available from https://iassistdata.org/blog/how-do-i-do-qualitative-research-bridging-gap-between-qualitativeresearchers-and-methods-res-1 (accessed 12 january 2019). swygart-hobaugh, m., conte, j. and cooper, l. 2017, “but how do i *do* qualitative research? bridging the gap between qualitative researchers and methods resources”, in iassist iblog, available from https://iassistdata.org/blog/how-do-i-do-qualitative-research-bridging-gap-betweenqualitative-researchers-and-methods-resou (accessed 12 january 2019). https://doi.org/10.29173/iq956 https://www.youtube.com/watch?v=vfqaw61o0rg https://iassistdata.org/blog/how-do-i-do-qualitative-research-bridging-gap-between-qualitative-researchers-and-methods-res-1 https://iassistdata.org/blog/how-do-i-do-qualitative-research-bridging-gap-between-qualitative-researchers-and-methods-res-1 https://iassistdata.org/blog/how-do-i-do-qualitative-research-bridging-gap-between-qualitative-researchers-and-methods-resou https://iassistdata.org/blog/how-do-i-do-qualitative-research-bridging-gap-between-qualitative-researchers-and-methods-resou 14/16 swygart-hobaugh, mandy (2019) bringing method to the madness: an example of integrating social science qualitative research methods into nvivo data analysis software training, iassist quarterly 43(2), pp. 1-16. doi: https://doi.org/10.29173/iq956 appendix 1 handout given to attendees at joint presentation. https://doi.org/10.29173/iq956 15/16 swygart-hobaugh, mandy (2019) bringing method to the madness: an example of integrating social science qualitative research methods into nvivo data analysis software training, iassist quarterly 43(2), pp. 1-16. doi: https://doi.org/10.29173/iq956 appendix 2 selected readings on qualitative research methods and computer aided/assisted qualitative data analysis software (caqdas). bourdon, s. 2002, “the integration of qualitative data analysis software in research strategies: resistances and possibilities”, forum qualitative sozialforschung / forum: qualitative social research, vol. 3, no. 2, doi: 10.17169/fqs-3.2.850. houghton, c., murphy, k., meehan, b., et al. 2017, “from screening to synthesis: using nvivo to enhance transparency in qualitative evidence synthesis”, journal of clinical nursing, vol. 26, no. 5–6, pp. 873–881, doi: 10.1111/jocn.13443. hutchison, a.j., johnston, l.h. and breckon, j.d. 2010, “using qsr‐nvivo to facilitate the development of a grounded theory project: an account of a worked example”, international journal of social research methodology, vol. 13, no. 4, pp. 283–302, doi: 10.1080/13645570902996301. leech, n.l. and onwuegbuzie, a.j. 2011, “beyond constant comparison qualitative data analysis: using nvivo”, school psychology quarterly, vol. 26, no. 1, pp. 70–84, doi: 10.1037/a0022711. maher, c., hadfield, m., hutchings, m., et al. 2018, “ensuring rigor in qualitative data analysis: a design research approach to coding combining nvivo with traditional material methods”, international journal of qualitative methods, vol. 17, no. 1, doi: 10.1177/1609406918786362. paulus, t., woods, m., atkins, d.p., et al. 2017, “the discourse of qdas: reporting practices of atlas.ti and nvivo users with implications for best practices”, international journal of social research methodology, vol. 20, no. 1, pp. 35–47, doi: 10.1080/13645579.2015.1102454. woods, m., paulus, t., atkins, d.p., et al. 2016, “advancing qualitative research using qualitative data analysis software (qdas)? reviewing potential versus practice in published studies using atlas.ti and nvivo, 1994–2013”, social science computer review, vol. 34, no. 5, pp. 597–617, doi: 10.1177/0894439315596311. zamawe, f.c. 2015, “the implication of using nvivo software in qualitative data analysis: evidencebased reflections”, malawi medical journal, vol. 27, no. 1, pp. 13–15. endnotes 1 mandy swygart-hobaugh is the leader of the research data services team at the georgia state university library, a data services librarian, and a liaison librarian for sociology. she can be reached by email: aswygarthobaugh@gsu.edu. 2 dr. larossa cites this quotation in his part of our joint presentation; it comes from the following: davis f (1974) stories and sociology. urban life and culture 3(3): 310–316. doi: 10.1177/089124167400300305. 3 given i did not have the time to do actual analysis on the 900+ letter exchanges, i artificially manipulated some of the analyses contained in my portion of the presentation to have ideal examples to illustrate nvivo’s features. i inform the presentation attendees of this, and, of course, tell them that fabricating results is not accepted practice. i also tell them that the dataset is available via icpsr if others wish to use the resources for actual analyses, but they should bear in mind that the data will not necessarily produce the results as presented. the word frequency https://doi.org/10.29173/iq956 https://doi.org/10.17169/fqs-3.2.850 https://doi.org/10.1111/jocn.13443 https://doi.org/10.1080/13645570902996301 https://doi.org/10.1037/a0022711 https://doi.org/10.1177/1609406918786362 https://doi.org/10.1080/13645579.2015.1102454 https://doi.org/10.1177/0894439315596311 mailto:aswygarthobaugh@gsu.edu https://doi.org/10.1177/089124167400300305 16/16 swygart-hobaugh, mandy (2019) bringing method to the madness: an example of integrating social science qualitative research methods into nvivo data analysis software training, iassist quarterly 43(2), pp. 1-16. doi: https://doi.org/10.29173/iq956 query and text search query results presented are genuine; the remaining coding and analysis examples have been artificially manipulated. 4 the attributes of gender of letter writer, year (and specific days) a letter exchange took place, and the city where the letter writer lived are clearly indicated in the patri letter exchanges data files. drawing from his experience analyzing the letters (larossa and reitzes, 1993), dr. larossa notes that the vast majority of letter writers refer to the gender and age of the child about whom they were writing. dr. larossa also notes that, while there is little direct reference in the letters to socioeconomic status indicators (i.e., education, occupation, or income of letter writers), he and his co-researcher surmised that those who wrote to patri generally were middle class (based on the quality of the stationery and grammar and spelling in the letters). regarding the attribute of family structure, it is possible that this information could be gleaned from letter content as well. https://doi.org/10.29173/iq956 1/3 dekker, harrison and riegelman, amy (2020) editors’ notes: [title], iassist quarterly 44(1-2), pp. 1-2. doi https://doi.org/10.29173/iq982 advocating for reproducibility as guest editors, we are excited to publish this special double issue of iassist quarterly. the topics of reproducibility, replicability, and transparency have been addressed in past issues of iassist quarterly and at the iassist conference, but this double issue is entirely focused on these issues. in recent years, efforts “to improve the credibility of science by advancing transparency, reproducibility, rigor, and ethics in research” have gained momentum in the social sciences (center for effective global action, 2020). while few question the spirit of the reproducibility and research transparency movement, it faces significant challenges because it goes against the grain of established practice. we believe the data services community is in a unique position to help advance this movement given our data and technical expertise, training and consulting work, international scope, and established role in data management and preservation, and more. as evidence of the movement, several initiatives exist to support research reproducibility infrastructure and data preservation efforts: • center for open science (cos) / open science framework (osf)i • berkeley initiative for transparency in the social sciences (bitss)ii • curating for reproducibility (cure)iii • project tieriv • data curation networkv • uk reproducibility networkvi while many new initiatives have launched in recent years, prior to the now commonly used phrase “reproducibility crisis” and ioannidis publishing the essay, “why most published research findings are false,” we know that the data services community was supporting reproducibility in a variety of ways (e.g., data management, data preservation, metadata standards) in wellestablished consortiums such as inter-university consortium for political and social research (icpsr) (ioannidis, 2005). the articles in this issue comprise several very important aspects of reproducible research: • identification of barriers to reproducibility and solutions to such barriers • evidence synthesis as related to transparent reporting and reproducibility • reflection on how information professionals, researchers, and librarians perceive the reproducibility crisis and how they can partner to help solve it. https://doi.org/10.29173/iq982 https://osf.io/ https://www.bitss.org/ https://www.bitss.org/ http://cure.web.unc.edu/ https://www.projecttier.org/ https://datacurationnetwork.org/ https://ukrn.org/ 2/3 dekker, harrison and riegelman, amy (2020) editors’ notes: [title], iassist quarterly 44(1-2), pp. 1-2. doi https://doi.org/10.29173/iq982 the issue begins with “reproducibility literature analysis” which looks at existing resources and literature to identify barriers to reproducibility and potential solutions. the authors have compiled a comprehensive list of resources with annotations that include definitions of key concepts pertinent to the reproducibility crisis. the next article addresses data reuse from the perspective of a large research university. the authors examine instances of both successful and failed data reuse instances and identify best practices for librarians interested in conducting research involving the common forms of data collected in an academic library. systematic reviews are a research approach that involves the quantitative and/or qualitative synthesis of data collected through a comprehensive literature review. “methods reporting that supports reader confidence for systematic reviews in psychology” looks at the reproducibility of electronic literature searches reported in psychology systematic reviews. a fundamental challenge in reproducing or replicating computational results is the need for researchers to make available the code used in producing these results. but sharing code and having it to run correctly for another user can present significant technical challenges. in “reproducibility, preservation, and access to research with reprozip, reproserver” the authors describe open source software that they are developing to address these challenges. taking a published article and attempting to reproduce the results, is an exercise that is sometimes used in academic courses to highlight the inherent difficulty of the process. the final article in this issue, “reprohacknl 2019: how libraries can promote research reproducibility through community engagement” describes an innovative library-based variation to this exercise. harrison dekker, data librarian, university of rhode island amy riegelman, social sciences librarian, university of minnesota references center for effective global action (2020), about the berkeley initiative for transparency in the social sciences. available at: https://www.bitss.org/about (accessed 23 june 2020). ioannidis, j.p. (2005) ‘why most published research findings are false’, plos medicine, 2(8), p. e124. doi: https://doi.org/10.1371/journal.pmed.0020124 i https://osf.io ii https://www.bitss.org/ iii http://cure.web.unc.edu https://doi.org/10.29173/iq982 https://www.bitss.org/about https://doi.org/10.1371/journal.pmed.0020124 3/3 dekker, harrison and riegelman, amy (2020) editors’ notes: [title], iassist quarterly 44(1-2), pp. 1-2. doi https://doi.org/10.29173/iq982 iv https://www.projecttier.org/ v https://datacurationnetwork.org/ vi https://ukrn.org https://doi.org/10.29173/iq982 1/15 lubaale, yovani, nakabugo, goretti, and nassereka, faridah (2021), learning outcome in literacy and numeracy in uganda: mining uwezo assessment data to demonstrate the importance of maternal education, iassist quarterly 45(3-4), pp. 1-15. doi: https://doi.org/10.29173/iq1001 learning outcome in literacy and numeracy in uganda: mining uwezo assessment data to demonstrate the importance of maternal education yovani lubaale1, goretti nakabugo2, and faridah nassereka3 abstract academic performance in primary education plays a crucial role in obtaining further educational opportunities. despite increased focus on addressing the inequality gaps in access to education, a number of studies have shown that children living in poor families with mothers who have low educational attainments experience less success, both in school and later as adults in the workforce, than children living in more advantaged circumstances. this paper analyses the effect of mothers’ education on the numeracy and literacy learning outcomes among children in uganda. mining data from the 2018 uwezo uganda learning assessment survey, we explore the influence of maternal education on learning outcomes. the findings showed that the proportion of children who demonstrated the ability of competently reading and comprehending a story of primary two level increased with increasing maternal education. whereas only 13.6% of the primary four children whose mothers had never been to school were able to read and comprehend a story (the highest level in literacy assessment), more than four times (50.7%) of the children whose mother had above senior four qualification had similar abilities. a similar trend was seen with performance in numeracy where 31.9% of primary four children whose mothers had no education at all were able to attain the highest numeracy level, compared to 59.1% for children whose mothers’ level of education was beyond senior four. it was further observed that slightly more than one in three (35.6%) of the primary one/two children whose mothers had never been to school were completely non numerate compared to less than one in ten (9.0%) of the children whose mothers had studied beyond senior four who were non-numerate. given the changes in access to schooling and impact on learning yielding from the global covid 19 pandemic, whereas the data mined was collected before this pandemic, there is need for reflection on the home-schooling approach being proposed by government and other stakeholders considering that this is likely to benefit more children whose mothers have higher levels of education than those with less education or never. keywords maternal education, learning outcomes, numeracy, literacy, data mining, uganda introduction sustainable development goal four (sdg 4) is about provision of quality education. this goal ensures that all girls and boys complete free primary and secondary schooling by 2030. it also aims to provide equal access to affordable vocational training, to eliminate gender and wealth disparities, and achieve universal access to a quality higher education (undp, 2000). equal access to education opportunities, the type or quality of schooling, resourcing/financing learners’ achievement in primary education plays a crucial role in obtaining further educational opportunities. childhood education not only affects the achievement and happiness of a person at the individual level, but also shapes the labor force quality and capacity of innovation to determine the potentiality of the development of a nation (zhonglu & zeqi 2018). academic outcomes in primary education plays a crucial role in obtaining further educational opportunities. a number of studies have shown that children living in poor families with mothers who have low educational attainments experience less success, both in school and later as adults in the workforce, than children https://doi.org/10.29173/iq1001 2/15 lubaale, yovani, nakabugo, goretti, and nassereka, faridah (2021), learning outcome in literacy and numeracy in uganda: mining uwezo assessment data to demonstrate the importance of maternal education, iassist quarterly 45(3-4), pp. 1-15. doi: https://doi.org/10.29173/iq1001 living in more advantaged circumstances (hernandez & napierala 2014, gooding 2001). maternal education is a key driver of education attainment of children (birdsall, levine & ibrahim, 2005; browne & barrett, 1991 cited in abuya et al 2018). it has been further revealed that the children of teen mothers who are able to go back to school and complete their education do show improvements compared to children of mothers who are unable to continue going to school (dizon, 2014). this further shows that whatever level a mother goes back to school, her educational achievement has a direct positive impact on the learning outcomes of her children whether born to her before or after attaining higher educational levels (dizon 2014). the effect of parental education has also been found crucial in learning a foreign language. khodadady and farnaz (2012) studied 1352 students’ performance in iran and found that third graders whose fathers and mothers held secondary and higher education certificates performed significantly higher than those having parents with primary education in learning english as a foreign language. against this backdrop, the team thought to mine the uwezo assessment data exploring the effect of maternal education on the learning outcomes of learners. objectives of the paper the objective of this paper was to study learning outcome in literacy and numeracy in uganda; mining uwezo assessment data to demonstrate the importance of maternal education to learning outcome. research question does the performance (literacy numeracy and ethno mathematics) of learners in primary one (p1) / primary two (p2) and learners in primary four (p4) differ by maternal education of the mother. hypotheses to test i. there is no difference in literacy, numeracy and ethno mathematics among p1/p2 learners by maternal education ii. there is no difference in literacy, numeracy and ethno mathematics among p4 learners by maternal education. methodology this section explains the source of data, the analytical framework and data analysis. source of data the source of data for this paper is the 2018 uwezo uganda learning assessment (uwezo 2019), which was the eighth learning assessment (literacy and numeracy) that was successfully undertaken in 324 sampled districts in october 2018. the survey conducted at household level was representative at both the national and district level and focused on assessing children aged 6-16 years. the mined sub-sample consisted of household data from 950 enumeration areas, involving 16,859 households with 45,676 children of whom 29,430 were aged 6-16 years. of these, 8,372 were in primary one and two (p1-p2) and 3,405 in primary four (p4) on which the analysis is based. the reason for selecting those in the primary one and two was because a lot of work has been emphasized at these low levels (republic of uganda, 2020). selecting out the primary four children is because this is the class of transition from the lower primary thematic curriculum to a more subject based curriculum with a focus on use of english as the language of instruction. based on the ugandan education system, age six5 is the official age of entry into primary one and if the learner went through the system without disruption, then they would be in senior four at the age of https://doi.org/10.29173/iq1001 3/15 lubaale, yovani, nakabugo, goretti, and nassereka, faridah (2021), learning outcome in literacy and numeracy in uganda: mining uwezo assessment data to demonstrate the importance of maternal education, iassist quarterly 45(3-4), pp. 1-15. doi: https://doi.org/10.29173/iq1001 sixteen years (republic of uganda, 2008). however, due to other factors (like working mothers, household poverty, distance to the nearest primary school), children’s age of entry in school varies whereby some children begin primary one as early as 5 years and others begin late at the age of 8 years while other repeat classes (rti6 2018). about uwezo assessment the uwezo assessments aim at generating evidence on learning outcomes in basic literacy and numeracy that can be used to influence policy action, reforms and practices towards improving learning outcomes. the 2018 learning assessment adopted a two-stage cluster sampling design with households as the elements and eas as clusters. selection of 30 eas in each of the 32 districts was done using probability proportional to size. selection of households at the second stage was done using simple random sampling following which 20 households were selected from each of the 30 eas per district. the survey involved assessing all children regularly residing in the household whether in or out of school using a simple primary 2 test in literacy and numeracy as well as collecting background information on the children and household. data on selected sdgs was collected including that on the quality of drinking water, health and nutrition. detailed explanation on the methodology can be obtained from uwezo eighth learning assessment report (uwezo 2019). the uwezo assessment tests in literacy and numeracy are benchmarked against the uganda primary 2 curriculum. while developing uwezo tests competencies in order of increasing difficulty are distinguished. for the literacy, tests assessed a child’s competence by determining their ability to; recognize alphabet letters, recognize and read commonly used words, read a short sentence and their ability to read a short story and comprehend it. the numeracy test assesses basic skills in terms of counting and matching, number recognition, operations with whole numbers and every day mathematics in form of word problems. during the assessment, each child is graded according to the highest level achieved. however, for both competencies, non-reader and non-numerate levels are established to categorize children who are unable to recognize letters of the alphabet or count and match and are regarded the lowest gradable level. the mined data categorizes the different learners by their level of performance in both literacy and numeracy. data analysis the analysis compared learning outcome assessment among children in primary one and two combined and those in primary four with maternal education. at the bivariate level the final learning outcome of the child either in literacy or numeracy were cross-tabulated with the maternal education. to show the differences in the learning outcome, more emphasis for literacy was placed on those who attained the story reading and the non-literate while for the numeracy, emphasis was placed on non-numerate and those able to do division. an unadjusted binary logistic regression in which mothers is education was the independent variable while ethno mathematics and the three bonus questions were the dependent variables was run. the results are presented as odds ratios with the corresponding significance levels. the binary logistic (logit) regression model takes the form: ( )3,2,1,,education maternal 1 log setsetsetmaticsethnomathef p p i i =        − ……………1 https://doi.org/10.29173/iq1001 4/15 lubaale, yovani, nakabugo, goretti, and nassereka, faridah (2021), learning outcome in literacy and numeracy in uganda: mining uwezo assessment data to demonstrate the importance of maternal education, iassist quarterly 45(3-4), pp. 1-15. doi: https://doi.org/10.29173/iq1001 where f=β0+ β1ethnomath+ β2set1+ β3set2+ β4set3+µ βithe coefficients µthe error term ρiprobability of the child being vulnerable analytical framework according to zhonglu and zeqi (2018), documented and undocumented daily experience shows that the impact of family socio-economic status on children’s academic achievement is not direct, but rather through the following two paths. first, families with relatively high socio-economic status will strive to secure quality educational opportunities for their children, such as those provided by key schools and markets in the system, which in turn will influence their academic achievements positively. secondly, these key schools are the ones with excellent teachers. in turn, these learners not only have a direct impact on learning outcome but are also affected by the learning attitudes and behaviors through teachers and peers. the result is that better future learning outcomes and further educational opportunities. thirdly, family socio-economic status affects children’s learning behavior and academic performance by affecting parents’ educational expectations towards children and their educational participation. parents’ educational expectations and behavioral support for children are, to a certain extent, also affected by their socio-economic status, resources availability and parents’ ability. this analytical framework is a slight modification from zhonglu and zeqi (2018). based on existing literature, this paper, has mined the uwezo uganda data to try and explain the learning outcomes of children based on their mothers’ level of education (maternal education). whereas the formulation of this analytical framework used data from china, a country with different socio-economic and demographic characteristics to uganda, it is very practical to the ugandan situation. the analytical framework for this paper is displayed in figure 1: the framework shows that the outcome for this analysis is learning competency in numeracy and literacy. there are direct and indirect factors that affect the learning outcomes of children. those that affect the child directly include the child’s learning behavior, parents’ educational participation, and educational opportunities available. on the other hand, family social economic status is a background factor which can go through any of the three to affect the learning outcomes, for example, if the family belongs to the lower social economic status (ses), it may be hard in the era of online learning to afford a computer/tablet or smartphone. lack of the any of these gadgets may impair a brilliant learner from achieving their potential. https://doi.org/10.29173/iq1001 5/15 lubaale, yovani, nakabugo, goretti, and nassereka, faridah (2021), learning outcome in literacy and numeracy in uganda: mining uwezo assessment data to demonstrate the importance of maternal education, iassist quarterly 45(3-4), pp. 1-15. doi: https://doi.org/10.29173/iq1001 source: modification of zhonglu and zeqi (2018) figure 1: analytical framework on the relationship learning outcomes limitation of the uwezo 2018 learning assessment data this analytical framework has not been implemented fully in this paper because not all the variables can be found in the dataset for example child learning behavior since uwezo assessments do not collect data on child learning behavior or differences in educational opportunities. little information is available on parental participation. results of the study in order to understand the effect of education on the child’s learning competences, mothers’ level of education was categorized into six groups. those who had never been to school at all, those who stopped in lower primary (primary one to primary fourp1-p4), those in upper primary (primary five to primary sevenp5-p7), those who never completed secondary, (senior to senior threes1-s3), those who completed ordinary level (senior fours4) and those who were educated beyond senior four and above s4. table 1 shows the distribution of learners in p1-p2 and p4 by their mothers’ educational level. table 1: distribution of p1-p2 and p4 children by maternal education p1-p2 p4 total number % number % number % none 1791 21.4 656 19.3 2447 20.8 p1-p4 2232 26.7 974 28.6 3206 27.2 p5-p7 2991 35.7 1250 36.7 4241 36.0 s1-s3 732 8.7 299 8.8 1031 8.8 s4 422 5.0 158 4.6 580 4.9 above s4 204 2.4 68 2.0 272 2.3 total 8372 100.0 3405 100 11777 100.0 source: 2018 uwezo uganda learning assessment parents educational level attainment participation family social economic status control variables: gender, language used at home, region differences in educational opportunities child learning behaviour child learning outcome (competence in numeracy and literacy https://doi.org/10.29173/iq1001 6/15 lubaale, yovani, nakabugo, goretti, and nassereka, faridah (2021), learning outcome in literacy and numeracy in uganda: mining uwezo assessment data to demonstrate the importance of maternal education, iassist quarterly 45(3-4), pp. 1-15. doi: https://doi.org/10.29173/iq1001 the distribution of children by maternal education shows that most children, had mothers who had acquired upper primary education. the table also indicates that one in every five children (20.8%), had mothers who had never been to school. on the other hand, slightly more than one in four, had mothers who had stopped in lower primary (27.2%). less than 20% of the children had mothers who were educated beyond primary seven (s1-s3 8.8%, s4 4.9%, and above s4 2.3%). these results have a lot of implication on the performance of children. this needs deeper reflection more so in the light of the education sector response to the covid-19 pandemic in which home schooling and learning materials are being distributed and parents are expected to assist their children learn at home (republic of uganda, moes, 2020). whereas some schools have been allowed to open, home schooling is going to take become another way of instruction even after covid-19 comes to the end. the 2018 uwezo assessment of learning outcomes in numeracy and literacy among children aged 6-16 years was benchmarked on primary two standard work. the various reading competences from this assessment are displayed in table 2 and figure 2a and 2b. table 2: literacy outcome for p1-p2 and p4 children by maternal education mothers level of education none p1-p4 p5-p7 s1-s3 s4 above s4 total highest level attained in english literacy primary one and two % % % % % % % number non-reader 52.5 50.3 43.6 27.9 24.1 13.3 44.2 3614 letter 33.5 36.0 38.8 42.1 36.4 29.1 36.8 3013 word 11.2 10.8 13.7 21.3 27.2 32.0 14.2 1163 paragraph 1.9 1.5 2.3 4.2 5.3 13.3 2.6 211 story 0.8 1.4 1.6 4.5 7.0 12.3 2.2 176 primary four non-reader 8.3 7.6 7.5 4.1 3.9 3.0 7.1 237 letter 35.6 33.3 27.7 20.9 20.1 9.0 29.5 984 word 30.3 32.2 31.1 29.4 20.8 19.4 30.4 1014 paragraph 12.2 12.3 13.1 16.9 16.9 17.9 13.3 444 story 13.6 14.6 20.6 28.7 38.3 50.7 19.7 657 source: 2018 uwezo uganda learning assessment https://doi.org/10.29173/iq1001 7/15 lubaale, yovani, nakabugo, goretti, and nassereka, faridah (2021), learning outcome in literacy and numeracy in uganda: mining uwezo assessment data to demonstrate the importance of maternal education, iassist quarterly 45(3-4), pp. 1-15. doi: https://doi.org/10.29173/iq1001 figure 2a: proportion of non-readers by maternal education figure 2b: proportion of children who reached story level by maternal education source: 2018 uwezo uganda learning assessment figure 2a displays the proportion of children aged 6-16 years in p1-p2 combined, and those in p4 who were categorized as non-readers and could hardly read anything from the tasks given. a non-reader in this context implies a child with hardly any mastery of reading. among the primary one and two children, the proportion of those who could not read anything was highest among children whose mothers had never acquired any education (52.5%) and lowest among children whose mothers had above senior four level of education (13.3%). the same trend is observed among the primary four learners, were about 8.3% of the primary four pupils whose mothers had never been to school were classified as non-readers while just 3% of the those whose mothers had been educated above senior four were classified as non-readers. the ability to read a story was the highest level that would be reached by a learner in relation to the literacy assessment. results indicate that less than one percent of children in primary one and two whose mothers had never been to school were able to read a primary two story. however, this proportion grows exponentially as the mothers’ education increases. slightly more than one in ten children in primary one and two were able to read a story whose mothers had above senior four level of education. looking at primary four children overall 7.1% of children were categorized as non-readers. however, the category of children contributing most to this result are children whose mother had never attended school (8.3%), those whose mothers stopped in lower primary (7.6%), and those whose mothers were in upper primary (7.5%). it is surprising to note that only 13.6% of the children in primary 4 whose mothers had never been 52.5 50.3 43.6 27.9 24.1 13.3 44.2 0.0 10.0 20.0 30.0 40.0 50.0 60.0 none p1-p4 p5-p7 s1-s3 s4 above s4 total non-reader p1-p2 0.8 1.4 1.6 4.5 7.0 12.3 2.2 0.0 2.0 4.0 6.0 8.0 10.0 12.0 14.0 none p1-p4 p5-p7 s1-s3 s4 above s4 total story reader p1-p2 8.3 7.6 7.5 4.1 3.9 3.0 7.1 0.0 1.0 2.0 3.0 4.0 5.0 6.0 7.0 8.0 9.0 none p1-p4 p5-p7 s1-s3 s4 above s4 total non-reader p4 13.6 14.6 20.6 28.7 38.3 50.7 19.7 0.0 10.0 20.0 30.0 40.0 50.0 60.0 none p1-p4 p5-p7 s1-s3 s4 above s4 total story reader p4 https://doi.org/10.29173/iq1001 8/15 lubaale, yovani, nakabugo, goretti, and nassereka, faridah (2021), learning outcome in literacy and numeracy in uganda: mining uwezo assessment data to demonstrate the importance of maternal education, iassist quarterly 45(3-4), pp. 1-15. doi: https://doi.org/10.29173/iq1001 to school were able to read a primary two story. on the other hand, one in two of primary four pupils whose mothers had an education level of senior four and above were able to read a story. results regarding numeracy are not different from those of literacy except that if one draws a trend line, the gradient for literacy will be stiffer than of numeracy. among p1-p2 children whose mothers had never been to school, 35.6% did not reach the lowest numeracy level that required counting and matching numbers (0-10). this proportion declines gradually among children whose mothers had lower primary (30.5%), upper primary (26.1%) to the lowest proportion among those whose mothers had been educated beyond senior four (9%). among p4 children whose mothers had never been to school, the trend is not smooth. however, there is a lot to note. none of the p4 children whose mothers had education above senior 4 was at the nonnumerate level. overall, about one in thirty among p4 children can be classified as non-numerate. the proportion of non-numerate children being high among mothers with upper primary education being higher than that of mothers without any level of education or lower primary is a surprise and requires further investigation. this can be by looking at data for others years to see if a similar trends or results will be noticeable. however, though the differences are observed, they are not statistically significant. table 3: numeracy outcome for p1-p2 and p4 children by maternal education maternal education none p1-p4 p5-p7 s1-s3 s4 above s4 total p1-p2 % % % % % % % number highest level attained in numeracy non-numerate 35.6 30.5 26.1 18.6 17.0 9.0 27.8 2264 matching 30.7 33.5 32.5 30.6 29.1 24.4 31.8 2596 number rec 10-99 11.8 13.0 13.9 14.6 15.0 12.9 13.3 1086 addition 7.5 9.0 10.9 13.0 12.9 12.4 10.0 816 subtraction 7.9 7.2 8.9 10.1 11.7 17.4 8.7 710 multiplication 2.2 2.0 2.7 3.5 3.6 10.9 2.7 224 division 4.2 4.9 4.8 9.6 10.7 12.9 5.6 459 p4 non-numerate 4.1 3.1 4.6 2.1 2.6 0.0 3.7 121 matching 8.7 8.9 7.3 6.6 6.6 3.0 7.8 259 number rec 10-99 7.6 6.4 6.6 8.7 1.3 4.5 6.6 219 addition 16.7 16.5 14.6 13.1 9.2 9.1 15.1 498 subtraction 20.5 22.7 20.4 17.3 19.1 19.7 20.8 686 multiplication 10.4 8.6 10.9 14.5 17.8 4.5 10.6 352 division 31.9 33.9 35.6 37.7 43.4 59.1 35.4 1171 source: 2018 uwezo uganda learning assessment https://doi.org/10.29173/iq1001 9/15 lubaale, yovani, nakabugo, goretti, and nassereka, faridah (2021), learning outcome in literacy and numeracy in uganda: mining uwezo assessment data to demonstrate the importance of maternal education, iassist quarterly 45(3-4), pp. 1-15. doi: https://doi.org/10.29173/iq1001 figure 3a: proportion of non-numerate children by maternal education figure 3b: proportion of children who reached division level by maternal education source: 2018 uwezo uganda learning assessment while reading a story and answering the comprehension questions is the highest level in literacy, division is the highest level of numeracy assessed by uwezo. among the primary one and two children, there was a gradual rise in the proportion of children who reached the division level. among children whose mothers had never been to school, about one in twenty-five (4.2%) reached the level of division while among those whose mothers had been educated above senior four, slightly more than one in ten was able to reach this maximum level. this is similar to what was observed in relation to performance in literacy (12.3%) and numeracy (12.9%), that is, slightly more than one in ten of the primary one and two children were able to carry out division and story reading tasks. among the primary four children, a reasonable number reached the division level irrespective of the mothers’ level of education. however, there are still noticeable differences. whereas 31.9% of children whose mothers had never been to school were able to reach the division level, this raises to almost double (59.1%) among children whose mothers had acquired education above senior four. 35.6 30.5 26.1 18.6 17.0 9.0 27.8 0.0 5.0 10.0 15.0 20.0 25.0 30.0 35.0 40.0 none p1-p4 p5-p7 s1-s3 s4 above s4 total non-numerate p1-p2 4.2 4.9 4.8 9.6 10.7 12.9 5.6 0.0 2.0 4.0 6.0 8.0 10.0 12.0 14.0 none p1-p4 p5-p7 s1-s3 s4 above s4 total division p1-p2 4.1 3.1 4.6 2.1 2.6 0.0 3.7 0.0 0.5 1.0 1.5 2.0 2.5 3.0 3.5 4.0 4.5 5.0 non-numerate p4 31.9 33.9 35.6 37.7 43.4 59.1 35.4 0.0 10.0 20.0 30.0 40.0 50.0 60.0 none p1-p4 p5-p7 s1-s3 s4 above s4 total divisionp4 https://doi.org/10.29173/iq1001 10/15 lubaale, yovani, nakabugo, goretti, and nassereka, faridah (2021), learning outcome in literacy and numeracy in uganda: mining uwezo assessment data to demonstrate the importance of maternal education, iassist quarterly 45(3-4), pp. 1-15. doi: https://doi.org/10.29173/iq1001 ethno mathematics and additional questions in addition to literacy and numeracy, there were four additional general knowledge /problem solving questions that are asked of every child being assessed. the ethno mathematics question and a set of three additional general knowledge questions which needed logic and application of day to day experiences in problem solving including identifying and naming shapes were further administered to the children. the bivariate analysis between these four additional questions and maternal education is displayed in table 4. table 4: performance outcome in ethno mathematics and general knowledge problem questions by maternal education none p1-p4 p5-p7 s1-s3 s4 above s4 total primary one and two can do ethno mathematics 24.1 23.5 23.8 32.0 30.4 48.1 25.4 1982 bonus question set one 14.2 13.5 18.3 27.9 32.5 48.5 18.4 1495 set two 7.7 8.5 10.4 13.6 14.9 27.4 10.3 822 set three 6.1 6.1 8.2 9.2 9.7 16.7 7.5 604 total 1791 2232 2991 732 422 204 8372 primary four can do ethno mathematics 62.2 62.8 63.6 68.6 65.6 80.6 64.0 2069 set one 46.8 45.8 49.2 59.7 58.9 77.9 49.7 1644 set two 28.0 29.3 29.7 33.1 40.1 60.0 30.6 1004 set three 18.2 18.2 21.6 27.9 25.3 37.9 21.0 689 total 656 974 1250 299 158 68 3405 source: 2018 uwezo uganda learning assessment in relation to ethno mathematics, the proportion of those who could do the ethno mathematics question did not follow a specific pattern. nonetheless, there is some evidence to suggest that the higher the maternal education the higher the proportion of those who could do the ethno mathematics tasks. the results for the set one and set three question which was concerned with the identification of objects showed a similar pattern with ethno mathematics but at lower proportions. for example, 62.2% and 46.8% compared to 65.6% and 58.9% of the children passed ethno mathematics and set one question among mothers with no education and those above senior four level respectively. performance on the two problem solving questions which mostly assessed logical reasoning reveal that maternal education has an influence on the children’s abilities. this revealed the largest percentage points gap of between 28% and 40% in this category of questions whose mothers had never been to school and children whose mothers educated beyond senior four. binary logistic regression in order to explain further the differences in the performance, we carried out an unadjusted logistic regression with maternal education as the independent variable and ethno mathematics and additional questions as the dependent variable. the results are displayed in table 5. children whose mothers had never been to school were taken as the reference category. the performance in ethno mathematics by p1-p2 children and that of p4 children reveals some interesting results different from those observed as the bivariate level. for the p1-p2 children, considering those whose mothers had never been to school, https://doi.org/10.29173/iq1001 11/15 lubaale, yovani, nakabugo, goretti, and nassereka, faridah (2021), learning outcome in literacy and numeracy in uganda: mining uwezo assessment data to demonstrate the importance of maternal education, iassist quarterly 45(3-4), pp. 1-15. doi: https://doi.org/10.29173/iq1001 significant differences in performance in ethno mathematics are observed after the mothers who completed primary school. the implication is that performance in ethno mathematics does not differ among children born to mothers who have never been to school and those in lower primary and upper primary though still results shows that even the little education received is important. results are statistically significant when between children whose mothers have never been to school and those with above primary education. among the primary four children, a significant difference was only observed among the children whose mothers have higher education that is above senior four (p=0.004). the meaning for this is that in primary four, performance in ethno mathematics does not vary much by the maternal education especially that this was primary two standard question. for the additional general knowledge questions revealed, children in primary one and two showed that they comprehended these questions better when mothers had some level of education than their counterparts whose mothers have never been to school. the odds ratios from the logistic regression show that the higher the maternal education the higher the ratios implying that as the maternal education increases, the better the performance even in the general knowledge /problem solving questions. on the other hand, little difference is observed in the learning outcome between a mother who has never been to school and those with lower primary education among the primary four children. in general, irrespective of the class of the child, little difference has been observed on these set of four questions, no differences were registered between children born to mothers who have never been to school and those who stopped in lower primary. overall major differences in learning outcomes are observed among children when mothers are beyond primary. table 5: binary logistic regression of the general knowledge questions by maternal education ethno set1 set2 set3 odds ratio p>z odds ratio p>z odds ratio p>z odds ratio p>z p1-p2 none 1.00 1.00 1.00 1.00 p1-p4 0.97 0.652 0.94 0.504 1.11 0.379 1.01 0.945 p5-p7 0.98 0.811 1.35 0.000 1.40 0.002 1.38 0.008 s1-s3 1.48 0.000 2.34 0.000 1.89 0.000 1.57 0.007 s4 1.37 0.010 2.91 0.000 2.10 0.000 1.66 0.010 above s4 2.92 0.000 5.69 0.000 4.51 0.000 3.10 0.000 _cons 0.32 0.000 0.17 0.000 0.08 0.000 0.06 0.000 p4 none 1.00 1.00 1.00 1.00 p1-p4 1.03 0.813 0.96 0.686 1.07 0.556 1.00 0.998 p5-p7 1.06 0.568 1.10 0.332 1.09 0.446 1.24 0.087 s1-s3 1.32 0.066 1.68 0.000 1.27 0.114 1.73 0.001 s4 1.16 0.446 1.63 0.008 1.73 0.004 1.52 0.050 above s4 2.52 0.004 4.01 0.000 3.86 0.000 2.74 0.000 _cons 1.65 0.000 0.88 0.112 0.39 0.000 0.22 0.000 source: 2018 uwezo uganda learning assessment https://doi.org/10.29173/iq1001 12/15 lubaale, yovani, nakabugo, goretti, and nassereka, faridah (2021), learning outcome in literacy and numeracy in uganda: mining uwezo assessment data to demonstrate the importance of maternal education, iassist quarterly 45(3-4), pp. 1-15. doi: https://doi.org/10.29173/iq1001 discussion the research question to be answered by the study was “does the performance (literacy numeracy and ethno mathematics) of learners in primary one (p1) / primary two (p2) and learners in primary four (p4) differ by maternal education of the mother”. awan (2013) says that education is the most important factor which plays a leading role in human resource development. it promotes productive and informed populace and creates opportunities for the socially and economically deprived sections of society. numerous studies such as gooding, 2001; rana; nadeem; saima; 2015) have revealed parental education more so maternal education to be a strong predictor of children’s education and behavior outcomes. with 22% of ugandan women compared to 16% of men having no education (uganda bureau of statisticsubos, 2017), it is worth investigating the effect of maternal education on learning outcomes of children. mining the uwezo uganda data is one way in which this effect has been demonstrated in light of using evidenced based results. mining the uwezo uganda data has been able to show that in terms of education outcome, children whose mothers have never been to school and those with lower primary are among most vulnerable. if the country is to get out of illiteracy and innumeracy to achieve sdg 4, there is need to come up with a special program to assist these children. there will be no development if these children are not assisted and as a country, it is the beginning of dualism. similarly, according to hernandez and napierala (2014), policies and programs aimed at increasing educational and economic opportunities usually target the less advantaged especially those with low income. this analysis shows that not only poverty should be targeted but that is just part of the problem hence the need to look at other factors like parental education especially maternal education. educational programs that aim improve education outcome by targeting the poor but not looking at the entire child will yield limited impact or take long to be realized without consideration of developing policies that promote women’s education. to focus simultaneously on both children and mothers will foster long-term learning and economic success for lowincome families. improving the maternal education should be the best way to improve learning outcomes of children. this will therefore bring in another aspect; how do we maintain children in the school especially the girl child in the era of universal primary education and universal secondary education. there is need to brake this cycle if uganda is to develop. whereas the study was conducted in 2018 just before covid19 pandemic, there are a lot of lessons that can be learned in this new normal, especially in uganda which has had long period of school closure. since the beginning of covid19 (march 2020), the government has been encouraging home schooling for learners due to the lock down of country. considering the importance of maternal education and with over 80% of the children in uganda with mothers who have not gone beyond primary seven, this causes a lot of concern. the lower levels of maternal education make it had for the parents especially the mothers to guide the children. conclusion in mining the uwezo 2018 data, we have been able to explore the effect of maternal education on child performance by looking at the literacy and numeracy outcomes of children. we have explored the three areas of assessment done by uwezo that is literacy, numeracy and comprehension. the results from the mined data shows that as maternal education increases the likelihood for the child to reach a higher level of assessment increases. there were two hypotheses to be tested based in this paper. the first one was that there is no difference in literacy, numeracy and ethno mathematics among p1/p2 learners by maternal education. results presented in table 2, 3, and 4 and figures 2a, 3a and 3b showed that performance among the p1/p2 pupils varied by the maternal education. we therefore reject the null hypotheses and accept the alternative that https://doi.org/10.29173/iq1001 13/15 lubaale, yovani, nakabugo, goretti, and nassereka, faridah (2021), learning outcome in literacy and numeracy in uganda: mining uwezo assessment data to demonstrate the importance of maternal education, iassist quarterly 45(3-4), pp. 1-15. doi: https://doi.org/10.29173/iq1001 there is a difference in the p1/p2 pupils in literacy, numeracy and ethno mathematics by the maternal education. the second hypothesis was that there is no difference in literacy, numeracy and ethno mathematics among p4 learners by maternal education. like the first hypothesis, results presented in tables 2, 3, and 4 and figures 2b, 3b, and 4b show that the performance among p4 learner varied by maternal education. recommendation data mining and utilization: the source of data for this paper was the 8th annual learning assessment, conducted in 2018. however, uwezo uganda has more than 8 datasets each having more than 40,000 learners. this is a big resource which researchers have not yet exploited to explain different reasons for declining levels of education in uganda. uwezo uganda should try as much as possible to popularize this data to the general public including researchers, academicians and policy makers since it is easily available free from www.uwezouganda.org. this data can be mined for further analysis and will generate a lot of knowledge for the country. additional analysis: more data mining to look at maternal education controlling for other factors like household socio-economic status, rural–urban residence and should be carried. another study can mine the performance of children living with their biological mothers controlling for maternal education with those not living with their biological mothers. government policy: data was collected on child attendance of early childhood education (ece). this is in line with the government policy on promoting ece. uwezo data can be analyzed further and provide critical insights which can be used to inform policy. similarly, further analysis of this data can be used to report on the sdgs. there is need for further analysis on learning outcome to inform policy and make meaningful decisions in this regard. this study has demonstrated the importance of maternal education, the government through either the ministry of education and sports or ministry of gender labour and socio welfare should resume adult studies that used to take place in the 1970s. however, these adult studies should target the use of technology to help parents, especially mothers. it can cover basics like creating and having an email, how to use online platform and so on. limitation of the data: there were some important variables which could enrich analysis. we are proposing that one or two questions can be added to the assessment tool to enrich this analysis like mother /female caretaker occupation. this is because an educated working mother may not have enough time to be with her children like an educated mother who is not working. https://doi.org/10.29173/iq1001 http://www.uwezouganda.org/ 14/15 lubaale, yovani, nakabugo, goretti, and nassereka, faridah (2021), learning outcome in literacy and numeracy in uganda: mining uwezo assessment data to demonstrate the importance of maternal education, iassist quarterly 45(3-4), pp. 1-15. doi: https://doi.org/10.29173/iq1001 references abuya b.a., mumah j., austrian k, mutisya m., kabiru c., (2018) mothers’ education and girls’ achievement in kibera: the link with self-efficacy. accessed: 19th september, 2021, from https://doi.org/10.1177/2158244018765608 awan, a.g. (2013) “relationship between environment and sustainable development: a theoretical approach to environmental problems” international journal of asian social science, vol 3 (3) 741-761. birdsall n., levine r., & ibrahim a. (2005) towards universal primary education: investments, incentives, and institutions. european journal of education, 2005, wiley online library. accessed: 19th september, 2021, from https://doi.org/10.1111/j.1465-3435.2005.00230.x dizon j (2014) mother's education crucial to academic success of children. accessed: jan 2021, from https://www.techtimes.com/articles/20024/20141113/mothers-education-crucial-to-academicsuccess-of-children.htm gooding, y. (2001) "the relationship between parental educational level and academic success of college freshmen " (2001). retrospective, theses and dissertations. 429. accessed: may 2020, from https://lib.dr.iastate.edu/rtd/429 hernandez d.j. and jeffrey s. napierala, (2014), mother’s education and children’s outcomes: how dual-generation programs offer increased opportunities for america’s families khodadady e and farnaz farrokh alaee. (2012). parents education and high school achievement in english as a foreign language rana m. asad khan; nadeem iqbal; saima tasneem (2015). the influence of parents educational level on secondary school students academic achievements in district rajanpur journal of education and practice www.iiste.org; issn 2222-1735 (paper) issn 2222-288x (online); vol.6, no.16, 2015 republic of uganda, ministry of education (2008) the education (pre-primary, primary and post-primary) act, 2008 acts supplement to the uganda gazette no. 44 volume ci dated 29th august, 2008. printed by uppc, entebbe, by order of the government. accessed: october 29th, 2020, from https://www.parliament.go.ug/documents/1257/acts-2008 republic of uganda, ministry of education and sports (2020) draft national inclusive education policy (unpublished) republic of uganda ministry of education and sports (2020): framework for provision of continued learning during the covid-19 lockdown in uganda rti (2018) uganda early years study milestone 3: final report. accessed: october 29th, 2020, from https://assets.publishing.service.gov.uk uganda bureau of statistics (2017) women and men in uganda, facts and figures. uganda bureau of statistics. undp (2000) sustainable development goals, goal 4, quality education. accessed: january 13th, 2021, from https://www.ug.undp.org/content/uganda/en/home/sustainable-developmentgoals/goal-4-quality-education.html uwezo (2016) are our children learning? uwezo uganda sixth learning assessment report. kampala: twaweza east africa uwezo (2019) are our children learning? uwezo uganda eighth learning assessment report. kampala: twaweza east africa uwezo uganda (2020) adapted strategy (2020-23) approved by uwezo uganda board of directors on 20th march 2020 zhonglu li1 and zeqi qiu2 (2018) how does family background affect children’s educational achievement? evidence from contemporary china. https://doi.org/10.29173/iq1001 https://doi.org/10.1177/2158244018765608 https://doi.org/10.1111/j.1465-3435.2005.00230.x https://www.techtimes.com/articles/20024/20141113/mothers-education-crucial-to-academic-success-of-children.htm https://www.techtimes.com/articles/20024/20141113/mothers-education-crucial-to-academic-success-of-children.htm https://lib.dr.iastate.edu/rtd/429 file:///c:/users/oschwart/stokes/iq/iq45_3_4/www.iiste.org https://www.parliament.go.ug/documents/1257/acts-2008 https://assets.publishing.service.gov.uk/ https://www.ug.undp.org/content/uganda/en/home/sustainable-development-goals/goal-4-quality-education.html https://www.ug.undp.org/content/uganda/en/home/sustainable-development-goals/goal-4-quality-education.html 15/15 lubaale, yovani, nakabugo, goretti, and nassereka, faridah (2021), learning outcome in literacy and numeracy in uganda: mining uwezo assessment data to demonstrate the importance of maternal education, iassist quarterly 45(3-4), pp. 1-15. doi: https://doi.org/10.29173/iq1001 endnotes 1prof. yovani m.lubaale, advisor (meal) and data management, uwezo uganda, email: ylubaale@gmail.com 2 dr.goretti nakabugo, executive director, uwezo uganda, email: gnakabugo@uwezouganda.org 3 ms.faridah nassereka, senior program officer, assessment, action and research, uwezo uganda, email: fnassereka@uwezouganda.org 4 some of the 32 districts have been subdivided since then like arua, jinja among others. 5 section 10 of the 2008 education act sub section 3a “primary education shall be universal and compulsory for pupils aged 6 (six) years and above which shall last seven years”. 6 rti research triangle institute. https://doi.org/10.29173/iq1001 mailto:ylubaale@gmail.com mailto:gnakabugo@uwezouganda.org mailto:fnassereka@uwezouganda.org 1/20 lee, c., gonzalez, s., & payne, k. (2024) research analysis: a world data system and canadian coretrustseal cohort needs assessment, iassist quarterly 48(2), pp. 1-20. doi: https://doi.org/10.29173/iq1084 the creative commons-attribution-noncommercial license 4.0 international applies to all works published by iassist quarterly. authors will retain copyright of the work and full publishing rights. research analysis: a world data system and canadian coretrustseal cohort needs assessment caroline leei, sarah gonzalezii, karen payneiii, meredith p. goinsiv abstract from july 2022 to december 2022, the world data system (wds) international technology (ito) and international program (ipo) offices conducted a review of strategic plans and technical roadmaps of all current wds members and the set of canadian repositories that participated in the digital research alliance of canada's coretrustseal certification support and funding pilotv (digital research alliance of canada, 2022). in this paper, we describe how a new organizational assessment method was designed and utilized to identify the needs and challenges faced by the wds and canadian cts pilot members. our method relied on reviewing public-facing documentation provided by the repositories, with a priority on strategic plans and technical road maps. in total, we reviewed 95 sources of information, including 33 strategic plans and 3 technical roadmaps describing a total of 95 out of the original 147 target organizations. in this paper, we also describe our assessment tool and the overarching challenges and goals we identified through the usage of this tool. finally, we will describe the limitations of our methodology and provide recommendations from the world data system on how best to assist the wds members and the cohort of canadian data repositories based on our findings. keywords world data system, coretrustseal, strategic plans, technical roadmaps, needs assessment, data repositories acknowledgements this work is supported by the digital research alliance of canada and ocean networks canada in partnership with the world data system international program office hosted by the university of tennessee oak ridge innovation institute, supported by a cooperative agreement (de-sc0021915) with the u.s. department of energy office of science (pi suzie allard). 1. introduction the world data system (wds), an interdisciplinary body of the international science council, is a global consortium of data repositories and affiliated organizations. the wds grew out of the world data centers, which were established in 1957 largely in response to the need to store large amounts of data created during the international polar and geophysical years (1932 and 1957, respectively). today, the wds is a consortium of over 120 data distribution centers and related entities in four different membership classes: regular, network, partner, and associate. wds members are charged with https://doi.org/10.29173/iq1084 https://alliancecan.ca/en/coretrustseal-certification-support-cohort-funding https://alliancecan.ca/en/coretrustseal-certification-support-cohort-funding https://creativecommons.org/licenses/by-nc/4.0/ 2/20 lee, c., gonzalez, s., & payne, k. (2024) research analysis: a world data system and canadian coretrustseal cohort needs assessment, iassist quarterly 48(2), pp. 1-20. doi: https://doi.org/10.29173/iq1084 responsible data stewardship and analysis, serving a wide range of research domains in over 29 countries. the overall objectives of wds are defined in its constitutionvi as follows: 1. enable universal and equitable access to quality-assured scientific data, data services, products and information 2. promote long-term data stewardship 3. foster compliance to agreed-upon data standards and conventions 4. provide mechanisms to facilitate and improve data access operationally, activities of the wds are conducted under the leadership of two complementary and coordinated offices: the world data system international program office (wds-ipo) hosted at the university of tennessee oak ridge innovation institute and the world data system international technology office (wds-ito) hosted by ocean networks canada at the university of victoria. this report details how these offices worked together to test out a new method of needs assessment for wds members and a cohort of canadian data repositories. as part of our commitment to understanding the current state of the wds membership, a needs assessment was conducted centered on a systematic review of the existing technical roadmaps and strategic plans of current wds member repositories, as well as the current set of repositories participating in the digital research alliance of canada's coretrustseal certification support and funding pilot (digital research alliance of canada, 2022). the latter set of repositories was included in this review in part because any repository that receives coretrustseal certification is eligible to apply for wds membership, and identification of their needs helps the wds create targeted programs that will encourage new applications. the goal of this assessment was, first and foremost, to identify the needs and challenges faced by wds members and canadian repositories. the outcome of this activity enables the wds to identify ways in which they can be of service to these member repositories, such as providing insight with respect to creating guiding architecture plans for integrating data repositories with other global research services or helping prioritize funding streams to projects that support common infrastructure needs across organizations. additional goals for this assessment include expanding wds member profiles, identifying wds datasets that support polar research, and identifying which members have satisfied coretrustseal requirements. ultimately, our priority was to assess whether a comprehensive needs assessment could be completed from a review of public-facing strategic plans. previous studies, such as those by ashiq et al. and thoegersen & borlund, utilized key search parameters to identify relevant documentation on their research topics of research data management (rdm) practices and researcher attitudes toward data sharing (ashiq et al., 2020; thoegersen & borlund, 2021). while more traditional needs assessments and https://doi.org/10.29173/iq1084 https://alliancecan.ca/en/coretrustseal-certification-support-cohort-funding https://alliancecan.ca/en/coretrustseal-certification-support-cohort-funding 3/20 lee, c., gonzalez, s., & payne, k. (2024) research analysis: a world data system and canadian coretrustseal cohort needs assessment, iassist quarterly 48(2), pp. 1-20. doi: https://doi.org/10.29173/iq1084 repository surveys consisted of a questionnaire that was then sent to a target audience (joo & peters, 2019; khan et al., 2021; payne & urquidi diaz, 2020), we sought to investigate whether a review of strategic plans would allow us to assess the needs and challenges of the wds and canadian cts cohort without the need to send a survey to each member, mitigating the administrative burden for the member organization. the core of strategic planning, as stated by bryson, is defined as “the identification and resolution of strategic issues,” ultimately yielding goals, policies and plans to address an organization's strategic challenges (bryson, 2018; george et al., 2019). therefore, based on this definition, we focused our needs assessment around the review of strategic plans. 2. methods we conducted a multi-staged process to create the instrument that guided our review of the target repositories public facing documentation, beginning with generating an initial list of characteristics we wanted to capture from each document. this stage of instrument development was informed by both the current wds action plan and by reviewing a series of strategic plans from the wds members. the initial list of characteristics covered areas such as mission and vision statements, long-term goals and strategic priorities, guiding principles, references to polar data or resources, funding sources and known challenges and obstacles encountered, among others. the next phase of instrument development was informed by the criteria under development by the global biodata coalition (gbc)vii (global biodata coalition, 2022; durinx et al., 2017). the gbc is an initiative designed to support funders by identifying and prioritizing fundamental data resources, referred to as core biodata resources, that should be maintained as part of the worldwide life science infrastructure. the criteria used to evaluate and assess these core resources was first piloted in europe as a set of qualitative and quantitative indicators and processes for identifying elixir core data resources (global biodata coalition, 2022; durinx et al., 2017). we examined the 2022 version of the global core biodata resources: concept and selection processviii and extracted criteria that could be applied to our review of wds and canadian cts candidate repositories (global biodata coalition, 2022). our initial list of characteristics was augmented with criteria from the gbc. in this phase, we included quantitative questions in the form of yes/no responses, such as “does the organization reviewed refer to a mission statement?” as well as qualitative responses recording more detailed information, such as the entire mission statement for the reviewed organization. together, our initial criteria, augmented with the gbc criteria, constitute the reviewer instrument. the instrument allowed multiple researchers to review strategic plans and roadmaps simultaneously, with all the responses being directed to a standardized spreadsheet. we implemented the instrument using google forms. we tested the instrument to ensure that responses created within the instrument fed correctly into the designated spreadsheet and all results were saved to a collective workspace. the second and primary reason for testing was to ensure that the information requested in the reviewer instrument was robust enough to successfully capture all the necessary data to achieve the defined objectives. we used five strategic plans for testing the instrument. three separate researchers conducted test reviews using the five strategic plans. feedback generated through the testing process showed the need for further https://doi.org/10.29173/iq1084 https://doi.org/10.5281/zenodo.5845116 https://doi.org/10.5281/zenodo.5845116 https://doi.org/10.5281/zenodo.5845116 https://doi.org/10.5281/zenodo.5845116 4/20 lee, c., gonzalez, s., & payne, k. (2024) research analysis: a world data system and canadian coretrustseal cohort needs assessment, iassist quarterly 48(2), pp. 1-20. doi: https://doi.org/10.29173/iq1084 questions to adequately capture sufficient information about the common themes, issues, and goals faced by the wds and canadian cohort repositories. additional questions were added, including questions related to repository partnerships with indigenous groups; references to diversity, equity and inclusion; references to sustainability and long-term funding sources; and any references to the repository's plans for obtaining new partnerships for research and funding. following the culmination of testing, the reviewer instrument was finalized. the entire reviewer instrument can be found in the world data system and canadian coretrustseal cohort need assessment: reviewer instrument and supplemental information (lee et al., 2023). 2.1 identification of canadian cohort and world data system members the members of the wds review team compiled a list of the organizations that went through the digital research alliance of canada's coretrustseal certification support and funding pilot to form the canadian cts cohort membership list (digital research alliance of canada, 2022). we included the canadian cts cohort because these repositories are in the early stages of completing their coretrustseal certification with the aid of our funder, digital research data alliance of canada. once they complete their coretrustseal certification, these repositories will be eligible to apply for wds membership. the world data system membership listix was created by second author sarah gonzalez after an in-depth membership audit completed in may 2022. the membership list includes four classes of members: regular members are data repositories that have achieved coretrustseal certification, network members are networks of certified repositories, and partner and associate member designations are assigned by wds for organizations that support research data. the number of members as of december 2022 are: ● wds regular members (84) ● wds network members (10) ● wds partner members (11) ● wds associate members (20) 2.2 collection of strategic plans/ roadmaps and other documentation for canadian cts cohort members the wds review team made a good-faith effort to find the strategic plans and/or technical roadmaps for the canadian cohort. unfortunately, we could not find their plans on their public-facing websites. per our request, the lead of the canadian cohort pilot program asked that these members send their strategic plans directly to the researchers by september 2022. as a result of this request, we received two strategic plans from the canadian cohort. investigation into additional public-facing documentation for the canadian cts cohort, such as annual reports or data policies, yielded no results on their public-facing websites. therefore, we assessed information on each cohort member from their organizational information on the re3datax website or directly from the cohort member’s website as part of phase two of the review process. figure 1 shows the hierarchy that was used when collecting public-facing documentation and levels utilized for the canadian cts cohort and the wds members. https://doi.org/10.29173/iq1084 https://doi.org/10.5281/zenodo.7738214 https://doi.org/10.5281/zenodo.7738214 https://doi.org/10.5281/zenodo.7738214 https://alliancecan.ca/en/coretrustseal-certification-support-cohort-funding https://alliancecan.ca/en/coretrustseal-certification-support-cohort-funding https://worlddatasystem.org/members/ https://worlddatasystem.org/members/ https://www.re3data.org/ 5/20 lee, c., gonzalez, s., & payne, k. (2024) research analysis: a world data system and canadian coretrustseal cohort needs assessment, iassist quarterly 48(2), pp. 1-20. doi: https://doi.org/10.29173/iq1084 2.3 collection of strategic plans/ roadmaps and other documentation for wds members we repeated the process of searching for strategic plans and/or technical roadmaps for the wds members. priority was placed on the collection of wds members' strategic plans and/or technical roadmaps. however, other documentation, including annual reports or data policies, journal articles or other sources of documentation (fact sheets, statutes), were additionally pulled from public-facing websites where they existed. additionally, coretrustseal applications were pulled from the coretrustseal (cts) websitexi. wds requires that regular members complete cts certification but does not require cts certification for other membership types. the researchers searched for strategic plans, roadmaps, or other documents published on each wds member’s website, but only in the case of the canadian cts cohort were the actual websites themselves consulted as a form of information for assessment. figure 1. hierarchy of documents that were prioritized in this assessment. 2.4 documentation review our review of the collected documents was conducted in two phases: first, a review of available strategic plans and technical roadmaps, followed by a review of additional documents that fell outside of the scope of strategic plans and technical roadmaps. reviews of the collection of strategic plans commenced in october 2022. at the time of this writing (december 2022), each document has been reviewed once. during the testing process, feedback mentioned the presence of multiple questions that could indicate similar items, i.e., long-term goals, strategic priorities, activities, strategies employed, et cetera. the goal of including these items was to capture the terminology the organization used to find common themes. a degree of discretion was required on the reviewer's part to identify where specific information should be https://doi.org/10.29173/iq1084 https://amt.coretrustseal.org/certificates https://amt.coretrustseal.org/certificates 6/20 lee, c., gonzalez, s., & payne, k. (2024) research analysis: a world data system and canadian coretrustseal cohort needs assessment, iassist quarterly 48(2), pp. 1-20. doi: https://doi.org/10.29173/iq1084 recorded within the instrument. in our summary below, overarching themes from these multiple related questions have been synthesized to identify overlapping goals and themes of the organizations. an additional note is that the category of ‘maybe’ was included in our review to indicate where it was difficult to discern the relevancy of the information provided. an example of this is where an organization may reference an indigenous group but not expand on how they are involved with the organization or in instances where the information was not clearly defined, such as in identifying long-term goals for the organization. in these instances, best discretion was used, and where there was doubt, a maybe was indicated instead of a yes. the instrument contains a final question where comments or additional feedback may be added by the reviewer. these comments included additional relevant information on the source of the assessment or any comments that the reviewer thought would be beneficial in the analysis of the data. during the second phase of the review process, additional documents that fell outside the scope of strategic plans or technical roadmaps were assessed. this included documents such as coretrustseal applications, annual reports, and data policies, among others. the decision was made on an individual basis for organizations that had multiple alternative documents, for example, an annual report and a data policy. the decision on which document to use was made by the researcher after determining which document contained the most relevant information based on the goals of this assessment. 2.5 analysis ultimately, all collected documents were reviewed and summarized by a single reviewer, which minimized the amount of variance when reviewer discretion was necessary. the analysis and summarization of reviewer responses followed a straightforward approach, with the primary goal of identifying commonalities and differences between each of the wds member types and the canadian coretrustseal cohort. responses were organized based on quantitative or numerical yes/no type answers and qualitative or long-form responses. long-form answers were summarized where possible using a web-based text summarizing toolxii; additionally, word statisticsxiii were also calculated to aid in identifying keywords or phrases that may help identify common themes. responses were divided by wds membership type. the intent behind this decision was the belief that when identifying goals, challenges, or needs, there would be commonalities within each membership type. in addition to these tools, we also created a more traditional researcher interpretation and summary of the literature where we identified the most relevant criteria and responses based on the goals of this assessment, the results of which can be found below. 3. results in this section, we provide the results of our assessment and discuss the common themes that emerged amongst all wds and canadian cts cohort members. we take full responsibility for any mistakes or unintentional mischaracterizations within this report, and therefore, we do not identify repositories by name. we begin by providing an overview of the organizations reviewed by summarizing the number of organizations, types of documents, domains (with particular attention paid to polar activities due to ongoing investments in this area by the wds), and references to diversity, equity and inclusion, particularly in the context of indigenous communities, for all organizations reviewed in aggregate. in https://doi.org/10.29173/iq1084 https://resoomer.com/en/ https://voyant-tools.org/ 7/20 lee, c., gonzalez, s., & payne, k. (2024) research analysis: a world data system and canadian coretrustseal cohort needs assessment, iassist quarterly 48(2), pp. 1-20. doi: https://doi.org/10.29173/iq1084 subsequent sub-sections, we break out our analysis by canadian cts cohort and wds member type. in the broadest view, across all member types, the two most significant concerns of data repositories were sustainability and creating effective partnerships with other entities in the research data space. moreover, we suggest that more granular details of our findings could be made after additional consultation with and consent from each repository. figure 2. comparison of the total number of target organizations we attempted to review, and the number of organizations actually reviewed. all of the organizations that were included in the review had one document reviewed per organization. of the 125 wds members, 73 documents were reviewed, with one document reviewed per organization for a total of 73 organizations. of the 22 canadian cts cohort members, 22 documents or sources of information were reviewed, with one per organization for a total of 22 organizations. in total, between the wds and canadian cts cohort members, 95 documents were reviewed, for a total of 95 organizations reviewed. the goal of this assessment was to review strategic plans or technical roadmaps; however, due to the lack of available strategic plans or technical roadmaps, the type of documents under review was expanded. in total, only 36 out of the 95 reviewed documents were a strategic plan or technical roadmap. therefore, documents such as coretrustseal applications, annual reports, data policies, and journal articles were included within the review as well and ultimately accounted for a more significant number of the documents reviewed. unfortunately, we found it difficult to find any additional documentation for the canadian cts cohort. the decision was made to include resources such as https://doi.org/10.29173/iq1084 8/20 lee, c., gonzalez, s., & payne, k. (2024) research analysis: a world data system and canadian coretrustseal cohort needs assessment, iassist quarterly 48(2), pp. 1-20. doi: https://doi.org/10.29173/iq1084 re3data6 repository information and the canadian cts cohort member’s website in instances where no other information could be identified. figure 3 shows a breakdown of the document types included in this review for the wds and canadian cts cohort members. figure 3. document types included within the review for all repositories (both wds and canadian cts cohort members). [document types not shown: other report: 1(1.1%), technical reference manual: 1(1.1%), newsletter: 1(1.1%), magazine: 1(1.1%)] figure 4 shows the domains served by the wds and canadian cts cohort members with documents included within this assessment (95 out of 147). however, it should be noted that many organizations are multi-disciplinary and fall into various domains. the most common domains for the wds and canadian cts cohort members that were included in this assessment are biological sciences and earth and environmental sciences. https://doi.org/10.29173/iq1084 https://www.re3data.org/ 9/20 lee, c., gonzalez, s., & payne, k. (2024) research analysis: a world data system and canadian coretrustseal cohort needs assessment, iassist quarterly 48(2), pp. 1-20. doi: https://doi.org/10.29173/iq1084 figure 4. domains of the wds and canadian cts cohort reviewed. [values not shown: ocean sciences: 17 (6.6%), astronomy and space sciences: 16 (6.3%), atmospheric sciences: 7 (2.7%), medical and health sciences: 9 (3.5%), mathematics: 3 (1.2%)] we also summarized the organizational references to polar activities, references to indigenous groups, and references to advancing diversity and inclusion within their organization. 15/95 or 15.8% of organizations reviewed indicated that they had references to polar activities. the polar activities mentioned included the collection of research data from both the arctic and the antarctic, with datasets on cryosphere processes, sea ice, terrestrial snow and ice, glaciers, and climate change effects, among others. references to indigenous groups were found in 10/95 organizations and included activities such as partnering and collaborating with indigenous groups, safeguarding indigenous data, supporting decolonization, land acknowledgements, and supporting the care principles for indigenous data governance (collective benefit, authority to control, responsibility and ethics) and the first nations principles of ocapxiv (ownership, control, access and possession) (carroll et al., 2020; fnigc, 2022). 18/95 organizations referenced advancing diversity, inclusion and equity, including some specific references to fostering a diverse and inclusive work environment and increasing participation amongst underrepresented and minority communities by providing education and reducing barriers. https://doi.org/10.29173/iq1084 https://doi.org/10.5334/dsj-2020-043 https://doi.org/10.5334/dsj-2020-043 https://fnigc.ca/ocap-training/ https://fnigc.ca/ocap-training/ 10/20 lee, c., gonzalez, s., & payne, k. (2024) research analysis: a world data system and canadian coretrustseal cohort needs assessment, iassist quarterly 48(2), pp. 1-20. doi: https://doi.org/10.29173/iq1084 3.1 wds regular members as previously stated, we have chosen to discuss each membership type individually with the belief that common themes amongst goals, challenges and needs would be more apparent within each membership class. the criteria presented for each membership type in tables 1 5 show the criteria that were found to be the most relevant when identifying goals, priorities, needs and challenges for the organizations reviewed. as of december 2022, the wds had a total of 84 regular members. we reviewed 48 documents describing 48 of the wds regular members. table 1. highlighted criteria for wds regular members wds regular do the documents reviewed mention: yes no maybe total mission statement 36 11 1 48 vision statement 11 36 1 48 guiding principles 10 32 6 48 long-term goals 17 23 8 48 strategic priorities 11 32 5 48 planning/outreach for obtaining new partnerships 14 30 4 48 obstacles/ challenges 18 30 0 48 technical requirements 19 29 0 48 note: for the criteria above, the category of ‘maybe’ indicates organizations that may have reference to the specific criterion, but it was difficult to determine the relevancy of information. yes no maybe n/a total having long-term and sustainable funding 25 0 8 15 48 note: for the above criterion, ‘n/a’ indicates organizations that did not mention funding in the document reviewed, whereas ‘maybe’ indicates organizations that referenced funding, but the longevity or sustainability of the funding source could not be confirmed. table 1 displays the criteria that were used to evaluate common themes, goals, challenges and needs and the number of references that were found in the documents reviewed for wds regular members. the most common theme that arose when looking at their mission statements, vision statements, and guiding principles was the desire to advance research, provide high-quality data products, increase collaboration amongst partnerships and within the user community, and ensure the long-term preservation of data. the last goal focuses on acting as a data steward and data custodian and promoting the exchange of data by ensuring data is open and easily accessible. additional commonalities are the desire to protect the environment and address global change and challenges. many of their guiding principles highlighted the need to be innovative, accountable, connected, respectful, and part of a community. unsurprisingly, when assessing their goals and strategic priorities, many of them mirrored the driving factors defined in their mission and vision statements. however, more specific goals and priorities emerged, such as the desire to develop techniques for data collection, research to advance technology, fill data gaps, and provide actionable research that improves understanding and education. https://doi.org/10.29173/iq1084 11/20 lee, c., gonzalez, s., & payne, k. (2024) research analysis: a world data system and canadian coretrustseal cohort needs assessment, iassist quarterly 48(2), pp. 1-20. doi: https://doi.org/10.29173/iq1084 organizations also spoke repeatedly about their desire to be involved in international communication with like-minded organizations strengthening their external relationships and building long-term partnerships. currently, the largest obstacles for wds regular members fall into two categories, underfunding or lack of funding and struggles with infrastructure. many organizations referenced a lack of funding for all their desired activities and projects, insufficient funding for continuous upward growth, lack of funding for staff, and lack of funding to improve infrastructure. another common theme was the desire to move to a cloud-based infrastructure; however, migrating to newer technologies was slow. funding presented itself as the leading challenge amongst wds regular members. 3.2 wds network members as of december 2022, the wds had a total of 10 network members. we reviewed 7 documents describing 7 of the wds network members. approximately half of the wds network members with documents that had been included within this review had a mission, vision, or guiding principles for their organization. the most common themes expressed by the network members reviewed showed a desire to create infrastructure that provides high-quality data, products and services, such as observatory networks and online environments that would support researchers using their data. of the four organizations that have either goals or strategic priorities, or both, the majority fall under the domains of remote sensing and astronomy and space sciences; therefore, the common goals represented are under the themes of improving networks of observatories, telescopes or other observation platforms. other common goals expressed are the desire to strengthen further contacts that make use of the products they provide, facilitate data exchange, and expand their membership base and integration with other scientific communities. community outreach was another priority that was highlighted, especially user feedback about ways they can improve products and develop new products that better serve their users. table 2. highlighted criteria for wds network members wds network do the documents reviewed mention: yes no maybe total mission statement 4 3 0 7 vision statement 3 4 0 7 guiding principles 2 5 0 7 long-term goals 1 4 2 7 strategic priorities 4 3 0 7 planning/outreach for obtaining new partnerships 4 2 1 7 obstacles/ challenges 5 2 0 7 technical requirements 6 1 0 7 note: for the criteria above, the category of ‘maybe’ indicates organizations that may have reference to the specific criterion, but it was difficult to determine the relevancy of information. yes no maybe n/a total having long-term and sustainable funding 0 0 3 4 7 https://doi.org/10.29173/iq1084 12/20 lee, c., gonzalez, s., & payne, k. (2024) research analysis: a world data system and canadian coretrustseal cohort needs assessment, iassist quarterly 48(2), pp. 1-20. doi: https://doi.org/10.29173/iq1084 note: for the above criterion, ‘n/a’ indicates organizations that did not mention funding in the document reviewed, whereas ‘maybe’ indicates organizations that referenced funding, but the longevity or sustainability of the funding source could not be confirmed challenges and technical requirements identified by the wds network members reviewed mainly revolved around their ability to meet the technical needs of their organization. mentioned challenges were made between meeting the data accuracy, resolution, and timeliness requirements of what users’ desire from their data products and what is actually feasible from the organization with current economic and organizational circumstances. additionally, when looking at the distribution of organizations that have stated they have long-term funding and are sustainable in table 2, three organizations had brief mentions of funding, but it was difficult to discern the terms of funding or what the references to sustainability implied, for instance, sustainability for the organization or sustainability for the products they provide. therefore, based on the information collected from the wds network members, technical needs appear to be their most significant challenge moving forward. 3.3 wds associate members as of december 2022, the wds had a total of 20 associate members. we identified and reviewed 9 documents describing 9 of these members. the majority of wds associate members with documents that were included in the review had either a mission or vision statement for their organization. the most common desires for these organizations were to be a supporting body for researchers and academic institutions, promote open data and the formation of openly shared data ecosystems, form international communities devoted to advancing scientific knowledge and research, embrace diversity and inclusivity, and support the next generation of researchers. similar to the wds regular and network members described above, the goals and strategic priorities of the wds associate members echo the desires expressed in the mission and vision statements and guiding principles. more specifically, the common goals are to ensure financial sustainability, create products and services that have a substantial impact on the end user, attract funders and partnerships, and improve outreach and facilitate engagement. table 3. highlighted criteria for wds associate members wds associate do the documents reviewed mention: yes no maybe total mission statement 7 2 0 9 vision statement 4 5 0 9 guiding principles 4 5 0 9 long-term goals 6 2 1 9 strategic priorities 6 3 0 9 planning/outreach for obtaining new partnerships 6 1 2 9 obstacles/ challenges 6 3 0 9 technical requirements 2 7 0 9 note: for the criteria above, the category of ‘maybe’ indicates organizations that may have reference to the specific criterion, but it was difficult to determine the relevancy of information. https://doi.org/10.29173/iq1084 13/20 lee, c., gonzalez, s., & payne, k. (2024) research analysis: a world data system and canadian coretrustseal cohort needs assessment, iassist quarterly 48(2), pp. 1-20. doi: https://doi.org/10.29173/iq1084 yes no maybe n/a total having long-term and sustainable funding 2 0 2 5 9 note: for the above criterion, ‘n/a’ indicates organizations that did not mention funding in the document reviewed, whereas ‘maybe’ indicates organizations that referenced funding, but the longevity or sustainability of the funding source could not be confirmed the largest identified challenge faced by wds associate members is funding, both lack of overall funding and the need for consistent funding. these included themes such as not obtaining funding and not being able to cover the operating costs of their organization. also noted was the issue of project-based funding. it was stated that only securing funding for specific projects diminishes the ability of the organization to perform other core tasks relating to its operation. only 2 out of the 9 organizations with documents that were included within this review identified themselves as having a long-term funding source; the two “maybe” organizations in table 3 reflect references in the reviewed documents that indicated that while they have project-based funding, they would require more resources to ensure that they can support all of their core operations. 3.4 wds partner members as of december 2022, the wds had a total of 11 partner members. we identified and reviewed 9 documents describing 9 of these members. out of the 9 organizations included as part of this review, all 9 members reviewed had some form of mission, vision, or guiding principles for their organization. many of the same themes as the other wds membership types were common amongst the wds partner members. these included the desire to advance user engagement, support international cooperation among organizations within the same or similar domains, build stakeholder relationships, and support the needs of data users. research objectives were also included in many of the wds partner mission and vision statements. these objectives included their desire to collect and validate scientific data and provide tools to disseminate data to progress research. progressing research and enabling the international scientific community were seen as particularly important to advancing their ability to solve key societal challenges. the goals and strategic priorities of the wds partner members are similar to the other wds membership types and echo their core mission and vision statements. one of the most common goals amongst wds partner members is strengthening and promoting the role and impact of the organization. steps for achieving this goal include the creation of activities that promote engagement, such as annual events, meetings with partner organizations, and the publication of white papers and other informative reports, as well as other opportunities that facilitate the promotion of the products and services of the organization. table 4. highlighted criteria for wds partner members wds partner do the documents reviewed mention: yes no maybe total mission statement 7 1 1 9 vision statement 8 1 0 9 https://doi.org/10.29173/iq1084 14/20 lee, c., gonzalez, s., & payne, k. (2024) research analysis: a world data system and canadian coretrustseal cohort needs assessment, iassist quarterly 48(2), pp. 1-20. doi: https://doi.org/10.29173/iq1084 guiding principles 5 3 1 9 long-term goals 6 3 0 9 strategic priorities 7 2 0 9 planning/outreach for obtaining new partnerships 7 1 1 9 obstacles/ challenges 7 2 0 9 technical requirements 6 3 0 9 note: for the criteria above, the category of ‘maybe’ indicates organizations that may have reference to the specific criterion, but it was difficult to determine the relevancy of information. yes no maybe n/a total having long-term and sustainable funding 0 2 2 5 9 note: for the above criterion, ‘n/a’ indicates organizations that did not mention funding in the document reviewed, whereas ‘maybe’ indicates organizations that referenced funding, but the longevity or sustainability of the funding source could not be confirmed of the 9 wds partner members that were included in this assessment, 7 of them identified obstacles or challenges faced by their organization or technical requirements needed by their organization. the most common challenges faced were related to funding, open sharing, and data access. lack of funding is one of the largest challenges faced by not only wds partner members but also among the entire wds membership. the most common challenge described by wds partner members relating to funding was the need for long-term funding to ensure sustainability. members described that short-term funding did not allow any long-term planning to be accomplished by the organization, and they had concerns about the lack of support. as seen in table 4, 4/9 organizations referenced funding, with 2/4 stating that they did not have long-term funding. the other two organizations made reference to funding, but information was not provided about the length of time the funding was guaranteed. data access was another concern; one wds partner member described that they rely on pay-walled data sources, while others had concerns about the lack of services and practices relating to data sharing and reuse. 3.5 canadian coretrustseal cohort members a total of 22 repositories participated as part of the canadian coretrustseal cohort. we reviewed 22 documents or websites describing all 22 of these cohort members. unlike the wds members reviewed above, the review of the cts cohort included websites due to a lack of alternative information. of the 22 canadian cts cohort members, 12 organizations have a mission statement, vision statement or guiding principles or a combination of the three for their organization. similarly to the wds members, the most common themes expressed by the canadian cts cohort members is the desire to be a global leader supporting research, promoting knowledge and data sharing with regards to the fair (findable, accessible, interoperable, reusable) principles, and to build and maintain strong partnerships. preserving, curating, and disseminating data was also a common theme expressed in the documentation of the canadian cts cohort members. guiding principles included common subjects such as accessibility, sustainability, integrity, leadership, interoperability, and sharing. of the organizations that listed their goals or strategic priorities, the most common themes were increasing the reach and role of the organization through advancing partnerships and increasing engagement with stakeholders and partners https://doi.org/10.29173/iq1084 15/20 lee, c., gonzalez, s., & payne, k. (2024) research analysis: a world data system and canadian coretrustseal cohort needs assessment, iassist quarterly 48(2), pp. 1-20. doi: https://doi.org/10.29173/iq1084 that share specific goals, such as advancing scientific research in specific domains. expanding data holdings was another goal mentioned by the canadian cts cohort, as well as ensuring collections are high-quality, reliable, easy to access, and meet the needs of the user community, as well as the goal of improving infrastructure to meet other technological goals. table 5. highlighted criteria for canadian cts cohort members canadian cts cohort do the documents reviewed mention: yes no maybe total mission statement 11 6 5 22 vision statement 4 18 0 22 guiding principles 8 14 0 22 long-term goals 4 18 0 22 strategic priorities 4 17 1 22 planning/outreach for obtaining new partnerships 3 17 2 22 obstacles/ challenges 1 21 0 22 technical requirements 0 22 0 22 note: for the criteria above, the category of ‘maybe’ indicates organizations that may have reference to the specific criterion, but it was difficult to determine the relevancy of information. yes no maybe n/a total having long-term and sustainable funding 3 0 2 18 22 note: for the above criterion, ‘n/a’ indicates organizations that did not mention funding in the document reviewed, whereas ‘maybe’ indicates organizations that referenced funding, but the longevity or sustainability of the funding source could not be confirmed. we were able to identify challenges faced by one organization in the canadian cts cohort. specifically, the challenge that technological change had brought increased cyber-security risks. this may be a challenge experienced by more than one of the cohort; however, we cannot draw any more conclusions with our current data. table 5 shows which canadian cts members have sustainable and long-term funding based on our assessment. four organizations identified that they had a funding source, and one organization possibly had a funding source, but it needed to be clarified about the longevity of that source. 18 members did not have any reference to sustainability or long-term funding in the information sources that were included in this assessment. therefore, we are unable to draw any conclusions on whether funding is a challenge for the canadian cts cohort. 4. discussion in broad strokes, we identified the need for sustainable funding and a desire to create effective partnerships as priorities for all organizations reviewed. in addition, all organizations have technical challenges, and the scope of those needs will need to be clarified with in-depth interviews with organization representatives. the canadian cohort goals and priorities matched similarly with the rest of the wds members. however, it was difficult to identify challenges for the canadian cohort. only one organization stated they had a challenge, and it was related to concerns about cyber-security. https://doi.org/10.29173/iq1084 16/20 lee, c., gonzalez, s., & payne, k. (2024) research analysis: a world data system and canadian coretrustseal cohort needs assessment, iassist quarterly 48(2), pp. 1-20. doi: https://doi.org/10.29173/iq1084 ultimately, our research was limited by our use of publicly available strategic plans for the 147 world data system and canadian coretrustseal cohort. of the 95 documents that were included in this assessment, only 36 of them were a strategic plan or technical roadmap. ultimately, our assessment included a much larger percentage of other types of documents, as seen in figure 3. therefore, because many of these documents were not designed to include criteria such as goals, priorities, or mission statements like strategic plans, we were largely unable to identify challenges or needs for a larger amount of our target audience than originally intended. therefore, the results of this research can only be used to summarize the common themes expressed by those organizations who referenced them and not the whole of the wds or canadian cts cohort membership. finally, our recommendations, therefore, are based on this subset of organizations. based on our findings, our recommendations for future research are that a targeted call for strategic plans/action plans or technical roadmaps is sent to each organization. additionally, we would recommend that an assessment survey instrument be developed based on the review instrument we created. this instrument would be sent to the data manager for each canadian cts cohort member and wds member. each wds member organization completes a report to the wds scientific committee, and this survey could be incorporated as part of this report per wds membership guidelines. the survey would create a 1:1 assessment for all of our identified criteria in this research assessment and remove the ambiguity of the maybe responses outlined above. the information included in the needs assessment survey would be gbc criteria; challenges; goals; diversity initiatives; and adherence to fair, trust, and care, among other criteria. we also recommend that this needs assessment be sent to each new wds member organization as an onboarding measure. for the canadian cts cohort members, we suggest that they take part in brief interviews with our researchers. only 3 of the 22 cohort members have a strategic plan, and this group may be in the process of creating planning documentation as they compile their cts application materials. it should be noted that a strategic plan is not a requirement for coretrustseal certification, but in preparing a strategic plan, the cohort members will be creating an outline of information that will be needed for the cts certification criteria. for example, the cts requirements for certification include the organization’s mission, vision, sustainability plan, and data provenance guidance, all items that may also be included in a strategic plan. the interviews will allow us to assess where the cohort members are in terms of documentation for their repository. finally, we can expand the documentation being assessed to include other certifications for data repositories. the focus for this analysis was coretrustseal, but documentation from applications for certification to niso, department of energy, or other certifying organizations may allow us to more readily gather data aligned with gbc identifiers and other factors for assessment of goals and challenges. again, we suggest a targeted call to all of the organizations to request certification documentation from other certifying bodies for data repositories. https://doi.org/10.29173/iq1084 17/20 lee, c., gonzalez, s., & payne, k. (2024) research analysis: a world data system and canadian coretrustseal cohort needs assessment, iassist quarterly 48(2), pp. 1-20. doi: https://doi.org/10.29173/iq1084 5. conclusion this needs assessment provided insight into the needs of the target repositories without an additional time commitment of those organizations. the assessment method built on the work they have already done and in that respect, it is a valuable activity for the wds. however, this process does lack some detail and further engagement with these organizations will be necessary. the method we developed could be part of the preparatory work for any consortia that is engaging in strategic planning. based on the analysis from this report, world data system will recommend moving forward in assisting canadian repositories and wds members in several ways. these recommendations are based on the world data system action plan 2022 through 2024xv (wds, 2022). the current world data system action plan goals are: 1. provide services and support to existing and new members 2. develop value narratives for wds members 3. provide global leadership and agenda setting 4. enhance access, quality and accessibility of data worldwide in light of these goals, and as a result of this work we can recommend to following actions be taken by the wds: 1. through this analysis, we have identified key challenges for data repositories around transparency, reproducibility, user-focus, sustainability, and technology (trust). in line with our action plan goals, wds could help repositories create a strategic plan/action plan or technical roadmap that clearly outlines how they are achieving their goals and addressing their challenges. these plans or roadmaps will follow the guidelines as outlined in the trust principlesxvi (lin et al., 2020). a clear strategic plan including trust principles will aid in implicitly stating each member’s value proposition. 2. wds should launch a webinar/workshop series on overcoming common challenges as identified through our research. these challenges include identifying and applying for sustainable funding, creating partnerships, open sharing, data access, technical needs, building/maintaining infrastructure, and handling the increase in data volume and format. 3. wds could create a resource area as part of the world data system website to include ways that repositories may more effectively communicate how they are implementing trust, fair and other data standards. this could include information on how each repository may communicate how they are meeting global biodata coalition (gbc) criteria or other discipline-specific data criteria or principles (global biodata coalition, 2022). 4. wds should provide communication through newsletters to the cohort and members regarding funding opportunities to enable data repositories to create sustainability plans and should also promote and engage in new funding models in development, including unesco open science funding models. 5. wds is developing a customer relationship management system for its members. this system should enable wds members to communicate effectively in a membership forum to allow https://doi.org/10.29173/iq1084 https://worlddatasystem.org/about/action-plan/ https://www.nature.com/articles/s41597-020-0486-7 https://doi.org/10.5281/zenodo.5845116 18/20 lee, c., gonzalez, s., & payne, k. (2024) research analysis: a world data system and canadian coretrustseal cohort needs assessment, iassist quarterly 48(2), pp. 1-20. doi: https://doi.org/10.29173/iq1084 members to connect to data experts among the membership, permitting our members to collectively find answers to common challenges. 6. wds could partner with codata and rda to provide working groups and educational opportunities for its members. they are already collaborating on scidatacon and international data week, the largest gathering of data experts. at these events, wds and its partners will provide information on best practices related to finding sustainable funding, creating partnerships, open sharing, data access, technical needs, building/maintaining infrastructure, and handling the increase in data volume and format. 7. the wds has generated a list of the organizations that supported polar research identified in this review and, going forward, should include them in outreach activities about relevant ongoing polar data support programs headed by the wds international technology office (ito). 8. wds should investigate opportunities to support member technical challenges identified in this review or in other scoping and needs assessment activities, for example supporting migrations to cloud infrastructure and advancing cyber-security. it is the recommendation of the research team that wds should commit to an ongoing needs assessment for member data repositories to be compiled on an annual basis. finally, wds will need to implement evaluation and monitoring tools to gauge the efficacy of its efforts in education and resource building to overcome data repository challenges. https://doi.org/10.29173/iq1084 19/20 lee, c., gonzalez, s., & payne, k. (2024) research analysis: a world data system and canadian coretrustseal cohort needs assessment, iassist quarterly 48(2), pp. 1-20. doi: https://doi.org/10.29173/iq1084 references ashiq, m., usmani, m., & naeem, m. (2020). a systematic literature review on research data management practices and services. global knowledge, memory, and communication, 71(8/9), p 647-671. https://doi.org/10.1108/gkmc-07-2020-0103 bryson, j. m. (2018). strategic planning for public and nonprofit organizations: a guide to strengthening and sustaining organizational achievement. john wiley & sons. carroll, s. r., garba, i., figueroa-rodríguez, o. l., holbrook, j., lovett, r., materechera, s., parsons, m., raseroka, k., rodriguez-lonebear, d., rowe, r., sara, r., walker, j. d., anderson, j., & hudson, m. (2020). the care principles for indigenous data governance. data science journal, 19. https://doi.org/10.5334/dsj-2020-043 digital research alliance of canada. (2022). coretrustseal certification support cohort & funding. digital research alliance of canada. https://alliancecan.ca/en/coretrustseal-certification-supportcohort-funding durinx, c., mcentyre, j., appel, r., apweiler, r., barlow, m., blomberg, n., cook, c., gasteiger, e., kim, j.h., lopez, r., redaschi, n., stockinger, h., teixeira, d., & valencia, a. (2017). identifying elixir core data resources. f1000research, 5, 2422. https://doi.org/10.12688/f1000research.9656.2 fnigc. (2022). the first nations principles of ocap®. the first nations information governance centre. https://fnigc.ca/ocap-training/ george, b., walker, r. m., & monster, j. (2019). does strategic planning improve organizational performance? a meta‐analysis. public administration review, 79(6), 810–819. https://doi.org/10.1111/puar.13104 global biodata coalition. (2022). global core biodata resources: concept and selection process. https://doi.org/10.5281/zenodo.5845116 joo, s., & peters, c. (2019). user needs assessment for research data services in a research university. journal of librarianship and information science, 52(3), 633–646. https://doi.org/10.1177/0961000619856073 khan, n., thelwall, m., & kousha, k. (2021). are data repositories fettered? a survey of current practices, challenges and future technologies. online information review, 46(3), 483–502. https://doi.org/10.1108/oir-04-2021-0204 lee, c., gonzalez, s., & payne, k. (2023). world data system and canadian cts cohort needs assessment reviewer instrument and supplemental information. world data system. https://doi.org/10.5281/zenodo.7738214 https://doi.org/10.29173/iq1084 https://doi.org/10.1108/gkmc-07-2020-0103 https://doi.org/10.1108/gkmc-07-2020-0103 https://doi.org/10.5334/dsj-2020-043 https://alliancecan.ca/en/coretrustseal-certification-support-cohort-funding https://alliancecan.ca/en/coretrustseal-certification-support-cohort-funding https://alliancecan.ca/en/coretrustseal-certification-support-cohort-funding https://doi.org/10.12688/f1000research.9656.2 https://doi.org/10.12688/f1000research.9656.2 https://fnigc.ca/ocap-training/ https://doi.org/10.1111/puar.13104 https://doi.org/10.1111/puar.13104 https://doi.org/10.1111/puar.13104 https://doi.org/10.5281/zenodo.5845116 https://doi.org/10.5281/zenodo.5845116 https://doi.org/10.5281/zenodo.5845116 https://doi.org/10.1177/0961000619856073 https://doi.org/10.1177/0961000619856073 https://doi.org/10.1177/0961000619856073 https://doi.org/10.1108/oir-04-2021-0204 https://doi.org/10.5281/zenodo.7738214 20/20 lee, c., gonzalez, s., & payne, k. (2024) research analysis: a world data system and canadian coretrustseal cohort needs assessment, iassist quarterly 48(2), pp. 1-20. doi: https://doi.org/10.29173/iq1084 lin, d., crabtree, j., dillo, i. et al. the trust principles for digital repositories. scientific data, 7, 144 (2020). https://doi.org/10.1038/s41597-020-0486-7 moher, d. (2009). preferred reporting items for systematic reviews and meta-analyses: the prisma statement. annals of internal medicine, 151(4), 264. https://doi.org/10.7326/0003-4819-151-4200908180-00135 payne, k., & urquidi díaz, a. (2020). world data system member survey 2019. zenodo. https://doi.org/10.5281/zenodo.3840406 thoegersen, j. l., & borlund, p. (2021). researcher attitudes toward data sharing in public data repositories: a meta-evaluation of studies on researcher data sharing. journal of documentation, 78(7), 1–17. https://doi.org/10.1108/jd-01-2021-0015 wds. (2022). action plan 2022 2024. world data system. https://worlddatasystem.org/about/actionplan/ i caroline lee, research associate for the world data system international technology office. itora3@oceannetworks.ca ii sarah gonzalez, program manager for the world data system international program office. sgonzal4@utk.edu iii karen payne, director for the world data system international technology office. ito-director@oceannetworks.ca iv meredith p. goins, executive director, world data system international program office mgoins2@utk.edu v digital research alliance of canada's coretrustseal certification support and funding pilot https://alliancecan.ca/en/coretrustseal-certification-support-cohort-funding vi https://worlddatasystem.org/about/constitution/ vii global biodata coalition (gbc). https://globalbiodata.org/ viii global core biodata resources: concept and selection process https://zenodo.org/record/5845116#.zbi1cnbmjd8 ix world data system members. https://worlddatasystem.org/members/ x re3data registry of research data repositories. https://www.re3data.org/ xi coretrustseal certified repositories. https://amt.coretrustseal.org/certificates xii resoomer. https://resoomer.com/en/ xiii voyant tools. https://voyant-tools.org/ xiv the first nations principles of ocap® https://fnigc.ca/ocap-training/ xv world data system action plan 2022-2024 https://worlddatasystem.org/about/action-plan/ xvi the trust principles for digital repositories https://www.nature.com/articles/s41597-020-0486-7 https://doi.org/10.29173/iq1084 https://doi.org/10.1038/s41597-020-0486-7 https://doi.org/10.1038/s41597-020-0486-7 https://doi.org/10.7326/0003-4819-151-4-200908180-00135 https://doi.org/10.7326/0003-4819-151-4-200908180-00135 https://doi.org/10.7326/0003-4819-151-4-200908180-00135 https://doi.org/10.5281/zenodo.3840406 https://doi.org/10.1108/jd-01-2021-0015 https://doi.org/10.1108/jd-01-2021-0015 https://worlddatasystem.org/about/action-plan/ https://worlddatasystem.org/about/action-plan/ https://worlddatasystem.org/about/action-plan/ mailto:ito-ra3@oceannetworks.ca mailto:ito-ra3@oceannetworks.ca mailto:sgonzal4@utk.edu mailto:ito-director@oceannetworks.ca mailto:mgoins2@utk.edu https://alliancecan.ca/en/coretrustseal-certification-support-cohort-funding https://worlddatasystem.org/about/constitution/ https://globalbiodata.org/ https://zenodo.org/record/5845116#.zbi1cnbmjd8 https://worlddatasystem.org/members/ https://www.re3data.org/ https://amt.coretrustseal.org/certificates https://resoomer.com/en/ https://voyant-tools.org/ https://fnigc.ca/ocap-training/ https://worlddatasystem.org/about/action-plan/ https://www.nature.com/articles/s41597-020-0486-7 i)^$sist newsletter vol.1, no. 1 i amendments article 7 amendments to these statutes may be proposed by any member with the support of five signatures. all amendments will be submitted to the general assembly for approval along with the election ballot. amendments approved by a majority of the members voting will be incorporated into the constitution. termination the association may be dissolved by a majority of the members. remaining funds will be transferred to the international social science council. transitional norm the ad hoc steering committee will serve until after the first meeting of the genertl assembly, but in any case not later than december 31, 1978. the ad hoc committee will arrange for a regular election to be held as soon as possible before that date. secretariat reports canadian secretariat report i sharon chappie data clearing house for the social sciences the data clearing house for the social sciences is the canadian secretariat for lassist. activities of lassist are published in the data clearinghouse bulletin . a large campaign for lassist membership was conducted by the data clearinghouse. over five hundred forms were sent out. to date, 53 individuals and institutions have expressed an interest in affiliating with lassist; 9 others wish to be retained only on the mailing list. supplementary membership campaigns are being considered. interest in the established action groups is as follows: data archive registry: 13; data archive development: 13; data acquisition: 14; data documentation: 16; classification: 8; process-produced data: 16. coordinators for each of the groups have been appointed. initial meetings are planned for the coming months. sist newsletter vol.1, no. 1 some of the respondents to the membership campaign letter recommended the creation of other action groups, such as the hardware/software technology action group (data organization and management) and a cross-file indexing action group. other suggestions included a group to interface between information data bases and potential users and a group to examine the relationship between bibliographic data bases, statistical data bases, and lassist. initial response to the lassist mailing indicates that there is considerable enthusiasm for such an organization in canada. it will be up to the canadian members of lassist to decide whether their needs can best be served by merging with the us organization in a north american lassist movement or by working closely with the americans in matters of conmon interest, but continuing to maintain a separate secretariat and a separate voice at the international level. west european secretariat report per nielsen danish data archives in europe, the lassist membership campaign in september, 1975, resulted in an immediate re'ponse from more than 80 interested individuals — and names are continuously cominj in. in the spring of 1976, a number of european social science data archives agreed to provide information dissemination facilities for lassist. consequently, all individuals having indicated an interest in lassist will receive mailings according to the following geographical division, where bracketed numbers indicate the approximate "membership" size in may, 1976: greece, italy, and spain: archivio dati e programmi per le scienze sociali (adpss), milan (7); belgium, france, and french-speaking switzerland: belgian archives for the social sciences (bass), louvain-la-neuve (8); denmark: danish data archives (dda), copenhagen (7); eire, israel, and the united kingdom: ssrc survey archive, essex (21); finland, norway, and sweden: norwegian social science data services (nsd), bergen (13); the netherlands: steinmetzarchief , amsterdam (7); austria, the federal republic of germany, and german-speaking switzerland: zentralarchiv (za), cologne (14). until recently, the za has also disseminated information to half a dozen potential members in eastern europe; these people (and, we hope a lot more to be brought in) will now be served by dr. ostrowski of the polish academy of sciences, warsaw. prior to the edinburgh meetings an overall mailing was carried out, including information on the action groups in europe. within a few months, we hope that the action groups have defined the priority of tasks so that workshops and other substantive activities accomplishing the objectives of lassist can be scheduled for 1977. vol274.indd 4 iassist quarterly winter 2003 editorʼs notes this is the fourth issue of the iassist quarterly vol. 27. with this issue we end the 2003 volume of the iq. the coming issues in vol. 28 (2004) will hopefully be out somewhat earlier next year. as we go to press at end of 2004, this is a good time to wish you all a good 2005. celia russell, keith cole, m.a.s. jones, s.m. pickles, m. riding, k. roy, and m. sensier are presenting an article on the samd-project (“the seamless access to multiple datasets”) based in manchester at mimas (“manchester information and associated services”) in collaboration with the supercomputing center at the same uk-university. the article describes the project to demonstrate the use of grid technologies for data retrieval, manipulation, and analysis in an economic application. the author indicates that grid technologies have superceded the world wide web for the handling of large and complex datasets in the physical sciences, and companies like microsoft, sun, ibm are developing grid technologies for use in business to business communications. the paper is about the first project to successfully use grid technologies in a social science context and looks at the implications of this for future social science quantitative research (e-science). this paper was presented by celia russell at the iassist 2004 conference in the session on “new avenues for data dissemination”. from the session on “changes in the way data archives process data” at the same iassist 2004 conference a paper by anne sofie fink kjeldgaard, søren priisholm, and birgitte grønlund jensen (at the danish data archives) shows “data processing in danish data archives”. at the danish data archives (dda) great effort is taken to preserve the data sets in a way that meets the needs of the secondary researcher. for this reason data processing is a core operation in the dda and great importance is attached to producing reliable and useful documentation of the preserved data files. also from manchester comes an article from angela dale, ccsr, at the university of manchester on “research access to microdata: an attempt to provide a context”. the article mentions that among the “aims and objectives of national statistics” is “to provide researchers, analysts and other customers with a statistical service that assists their work and studies”. however, statistical offices have to tread a careful balance between providing the data needed by all sections of society and maintaining the confidence of the general public. experiences of some countries have shown that the public can lose confidence in the national statistical office and that the process of data collection then will be undermined and may not recover. microdata are now being provided and used, it is thus very important that these data do not reveal any information for too precise identification. furthermore, a future breach of confidentiality should also be prevented – including using other, new, and unknown access and analysis methods. consequently a balance is required. this has to be carried out in such a way that a reasonable amount of detail is released in order not to hinder research and at the same time securing the confidential information i.e. hindering the specific identification. the more detail a dataset contains the more restricted the access to the file has to be. we invite you to visit the iassist website at www. iassistdata.org. there you will find information on previous and coming conferences. among other features of the website is the possibility to access the iassist quarterly as a pdf-file. papers for the iassist quarterly are most welcome. papers can be from iassist conferences, from other conferences, from local presentation, etc. for further information contact the editor via e-mail: kbr@sam.sdu.dk. this issue also presents the “iassist call for papers”. the iassist (and ifdo) conference will be in edinburgh 25th to 27th may 2005. karsten boye rasmussen, november 2004. 1 1/17 bonifacio, flavio; akullo, winny nekesa (2021), differences in data-sharing attitudes and behaviours, extended version to african data curators and data management experts, iassist quarterly 45(3-4), pp. 1-17. doi: https://doi.org/10.29173/iq993 differences in data-sharing attitudes and behaviours, extended version to african data curators and data management experts flavio bonifacio1 and winny nekesa akullo2 abstract this article reports the results of a survey conducted between 16th november and 8th december 2020 among african data curators and data experts about different aspects of data sharing. the sample of respondents has been extracted from participants to the 1st iassist africa regional workshop held on 11th -13th january 2021, kampala, uganda and other data experts and practitioners. first, we recall the main results of a previous article published by iq3 about the same argument in order to introduce the new survey. after that we analyse the new findings comparing them with the previous results, splitting the samples between africans and not africans. the idea of a new survey the idea came from a previous survey conducted at the end of 2017 and published in the article “differences in data-sharing attitudes and behaviours”, vol. 42 no. 3 (2018): iassist quarterly. the questions regarded different aspects of data sharing: tools used in building metadata, problems encountered in order to share the data, the propensity to share the data, and the satisfaction obtained from different working tasks. this article extends the survey to african data practitioners. the occasion has been the 1st iassist africa regional workshop, 11th -13th january 2021, kampala, uganda. winny nekesa helped me to reach the attendees and other interested persons by email, sending them the link to the questionnaire. thirty-one people took part in the survey. above all, this conclusion of the previous survey attracted my attention: “we have to broaden the scope of service promotion, moving from ‘developed countries to ‘developing countries’, where data curation is less practised, to younger people, involving women in greater responsibility and more remunerative roles. how to do that is a matter that goes beyond the scope of this article, but one suggestion is to transform the selfreferential meetings into open symposiums, for example moving from the usual locations (for iassisters usa, canada, north europe) to, perhaps, less easy locations, such as southern europe, africa or asia, and to less easy environments, outside the university, in an open public and private space.” we think that the iassist africa regional workshop that took place in uganda may be seen as a step moving in this direction. 2021 survey, goal and general description the goal of the work already reported in the previous survey “is to analyse attitudes toward data sharing in order to find strategies to expand data sharing and to show the best paths to gain new ‘premium followers’, as we call the best performers in data sharing (below). this pragmatic point of view comes from empirical evidence that is compliant with other more detailed analysis coming from wider surveys. we confirm what other sources say: ‘results show that researchers in different regions have different perceptions about data and different data behaviours’4. although in literature there are many other well documented perspectives from which to examine data sharing, covering a range of aspects5 from the political to the technical, we think that one of the most pressing issues is how to expand best practices in sharing data.” in the actual survey we add the description of perceptions about data and data behaviours among african data experts to the findings of the 2017 survey. the 2017 survey showed the existing differences between “italians” as representative of a community where data sharing is less practised, and other more attentive communities. https://doi.org/10.29173/iq993 2/17 bonifacio, flavio; akullo, winny nekesa (2021) differences in data-sharing attitudes and behaviours, extended version to african data curators and data management experts, iassist quarterly 45(3-4), pp. 1-17. doi: https://doi.org/10.29173/iq993 we start with a short description of background variables. contrasting with the non-africans sample, we note that africans work more often in the public administration, are slightly more male than female, are younger (more than one half less than 40), more frequently they have a master degree and studied more in the scientific field while not africans studied more in the social sciences field. considering the use of metadata standards, an important indicator of data familiarity, africans use them less than not-africans. if we split the not-african between italians and not-italians, the africans place themselves in the middle. except for gender and educational qualifications, all other relations are significant. in the following figures one asterisk (*) means significance level at alpha level of 10%, two asterisks (**) means significance level at alpha level of 5%. three asterisks (***) means significance level at alpha level of 1%. fig.1 – significance (**) https://doi.org/10.29173/iq993 3/17 bonifacio, flavio; akullo, winny nekesa (2021) differences in data-sharing attitudes and behaviours, extended version to african data curators and data management experts, iassist quarterly 45(3-4), pp. 1-17. doi: https://doi.org/10.29173/iq993 fig.2 – not significant fig. 3 – significance (**) fig. 4 – not significant https://doi.org/10.29173/iq993 4/17 bonifacio, flavio; akullo, winny nekesa (2021) differences in data-sharing attitudes and behaviours, extended version to african data curators and data management experts, iassist quarterly 45(3-4), pp. 1-17. doi: https://doi.org/10.29173/iq993 fig. 5 – significance (***) fig. 6 – significance (*) fig. 7 – significance (***) https://doi.org/10.29173/iq993 5/17 bonifacio, flavio; akullo, winny nekesa (2021) differences in data-sharing attitudes and behaviours, extended version to african data curators and data management experts, iassist quarterly 45(3-4), pp. 1-17. doi: https://doi.org/10.29173/iq993 data sharing attitudes and behaviours most of the questions have been extracted from imdsm2017 (see the above quoted 2017 survey report, differences in data-sharing…)6 and from the survey changes in data sharing and data reuse practices and perceptions (cf. differences in data-sharing…) among scientists worldwide and grouped into three main conceptual frameworks: 1. items influencing data sharing propensity (6 items) 2. items influencing work satisfaction (7 items) 3. problems emerging in data sharing (21 items) we report the results for each framework starting from a synthetic overview of the respondent answers frequencies for the african sample, ordered in descending order of agreement (agree somewhat, agree strongly are jointly considered as a single answer for ordering the frequencies). then we analyse in deeper detail only the distributions of the items showing greater, significant differences between africans and not africans. finally, we will show a synthetic index for each framework7, summarizing all the information in a single measure. items influencing data sharing propensity (6 items) in the following figure we report the items used to measure the sharing propensity. the most appreciated item refers to the use of other researcher datasets, while one of the less appreciated ones refers to placing all “my” data into a central repository. fig. 8 – the following statements relate to sharing scientific data. tell us how much you agree with each statement… ** https://doi.org/10.29173/iq993 6/17 bonifacio, flavio; akullo, winny nekesa (2021) differences in data-sharing attitudes and behaviours, extended version to african data curators and data management experts, iassist quarterly 45(3-4), pp. 1-17. doi: https://doi.org/10.29173/iq993 details the only distribution showing a significant difference between africans and not africans (α=5%) relates the opinion to place data into a central repository. looking at the distributions we observe a greater concentration of africans, both among the most disagreeing and the most agreeing, with respect to not africans. interesting is the fact that about one half of africans agree strongly with placing their own data into a central repository. fig. 9 significance (**) sharing propensity index the index summarizing all the items (data sharing propensity index) does not show significant differences. therefore, we can say that from the data sharing propensity point of view, africans and not africans show similar behaviour. fig. 10 – sharing propensity index https://doi.org/10.29173/iq993 7/17 bonifacio, flavio; akullo, winny nekesa (2021) differences in data-sharing attitudes and behaviours, extended version to african data curators and data management experts, iassist quarterly 45(3-4), pp. 1-17. doi: https://doi.org/10.29173/iq993 items influencing work satisfaction (7 items) in this case four items referring to the work satisfaction show a significant relation with the origin of the respondents. one of them (tools for preparing my documentation) is at 1% α level. also significant are processes for storing, searching and collecting data. looking at the order of items the most satisfying is the process for cataloguing-describing data while the process for analysing data is the least satisfying. fig. 11 – the following statements relate to how you collect and use research data. tell us how much you agree with the following ways to complete this sentence: i am satisfied with the… i am satisfied with… fig. xx significance (**) ** ** *** ** https://doi.org/10.29173/iq993 8/17 bonifacio, flavio; akullo, winny nekesa (2021) differences in data-sharing attitudes and behaviours, extended version to african data curators and data management experts, iassist quarterly 45(3-4), pp. 1-17. doi: https://doi.org/10.29173/iq993 details interesting is the fact that for almost all the items africans show stronger agreement. fig. 12 significant (**) fig. 13 significance (***) https://doi.org/10.29173/iq993 9/17 bonifacio, flavio; akullo, winny nekesa (2021) differences in data-sharing attitudes and behaviours, extended version to african data curators and data management experts, iassist quarterly 45(3-4), pp. 1-17. doi: https://doi.org/10.29173/iq993 fig. 14 – significance (**) fig.15 significance (**) https://doi.org/10.29173/iq993 10/17 bonifacio, flavio; akullo, winny nekesa (2021) differences in data-sharing attitudes and behaviours, extended version to african data curators and data management experts, iassist quarterly 45(3-4), pp. 1-17. doi: https://doi.org/10.29173/iq993 work satisfaction index this causes a higher concentration on higher scores of satisfaction index, although not yet significant. fig. 16 work satisfaction index problems emerging in data sharing (21 items) in this context most of the item distribution differences between africans and not africans are significant. there are five items that reach three stars of significance, α=1%: i would lose control of the data, my data change too quickly, my data are old they don’t answer to the questions researchers ask today, i spent a lot of money on this research and it is not economically convenient to share it, it could affect negatively my career. noteworthy is the fact that at least three items refer to personal views which directly contrast the data sharing positive vision. referring to the order of items in the table, on the top we find problems concerning data propriety, the shortage of funds, privacy considerations; on the bottom the aspects purporting personal interests instead of more general views. https://doi.org/10.29173/iq993 11/17 bonifacio, flavio; akullo, winny nekesa (2021) differences in data-sharing attitudes and behaviours, extended version to african data curators and data management experts, iassist quarterly 45(3-4), pp. 1-17. doi: https://doi.org/10.29173/iq993 fig. 17 – how much do you agree with these statements about the reasons that prevent data sharing? i would share the data but… details in all distributions we note that africans seem to have more problems in sharing data than not africans: they fear more to lose control of the data, they think they have old or not affordable data, or think it is not convenient to share their data in terms of economic or personal values (career), as we can see in the following more significant tables. also, in almost all other not reported less significant item tables this is true. ** *** ** ** *** ** *** *** ** ** *** ** * ** ** ** ** https://doi.org/10.29173/iq993 12/17 bonifacio, flavio; akullo, winny nekesa (2021) differences in data-sharing attitudes and behaviours, extended version to african data curators and data management experts, iassist quarterly 45(3-4), pp. 1-17. doi: https://doi.org/10.29173/iq993 fig. 18 – significance (***) fig – 19 significance (***) https://doi.org/10.29173/iq993 13/17 bonifacio, flavio; akullo, winny nekesa (2021) differences in data-sharing attitudes and behaviours, extended version to african data curators and data management experts, iassist quarterly 45(3-4), pp. 1-17. doi: https://doi.org/10.29173/iq993 fig – 20 significance (***) fig – 21 significance (***) https://doi.org/10.29173/iq993 14/17 bonifacio, flavio; akullo, winny nekesa (2021) differences in data-sharing attitudes and behaviours, extended version to african data curators and data management experts, iassist quarterly 45(3-4), pp. 1-17. doi: https://doi.org/10.29173/iq993 fig – 22 significance (***) sharing problems importance index as we did before, we summarize the answer to the items in one single index, the sharing problems importance index. in this case also the synthetic index shows a significative difference (α=5%) between africans and not africans: africans often recognise a high problems importance evaluation. fig. 23 sharing problems importance index significance (**) https://doi.org/10.29173/iq993 15 15/17 bonifacio, flavio; akullo, winny nekesa (2021), differences in data-sharing attitudes and behaviours, extended version to african data curators and data management experts, iassist quarterly 45(3-4), pp. 1-17. doi: https://doi.org/10.29173/iq993 the model coding v3 (not italy or africa) as 1, otherwise coding v3 as 0, coding v4 (africa) as 1, otherwise coding v4 as 0, coding v2 as 1 when the respondent works at university or in the public sector, otherwise coding v2 as 0, and taking data sharing propensity index as dependent variable (v1) we obtain the model schema of the model reported below (direct and indirect effects). it seems to be the most appropriate and not reducible model. all the direct and indirect effects are significant at alpha=5%. the indirect effect of the country of origin (v3, v4) influences the sharing propensity via the work sector, which is positive. the total effect of v3 and v4 (country) over v1 (data sharing propensity) is given by 0.27+(0.25*0.27)+0.27+(0.38*0.27) that equals 0.71. in other words, in this sample africans more often work at university or in the public sector than not africans and not italians that work more at university too (not africans and not italians means northern europa, canada, usa, etc., respondents): they both add to their own higher sharing propensity also the fact that they work at university (or in the public sector). we indeed know that working at a university has its own positive effect on data sharing propensity too, as the model shows. fig. 24 – path model with sharing propensity as dependent variable (v1) and work sector (v2) and country of origin (v3, v4) as independent variables https://doi.org/10.29173/iq993 16 16/17 bonifacio, flavio; akullo, winny nekesa (2021), differences in data-sharing attitudes and behaviours, extended version to african data curators and data management experts, iassist quarterly 45(3-4), pp. 1-17. doi: https://doi.org/10.29173/iq993 conclusions africans experts are more likely to be male, are younger, with more scientific study, working more in university or public administration. they appear to be at the same time more attentive to problems that may arise in the data sharing activity and more satisfied. they seem also to have a higher propensity to share data. in general, the distributions on the attitude items have not such a different shape if compared with the other “not african” experts. a warning comes from the work: africans seem recognise more problems in sharing data, as said above. an effort may be necessary to make the issues of data sharing more attractive. maybe it is only a feeling of mine but i observed more enthusiasm in the african sample. this will help in order to reach the goal. acknowledgements i want to thank first ms. winny nekesa akullo that helped me to collect data from respondents sending them the link to the questionnaire. without her i could not have done anything. second, i want to thank all those that had the patience to answer me. third, i want to thank also the others that did not answer this time but that surely will do so next time, and susan phillips for reviewing my english. references bonifacio, flavio (2017) ‘working across boundaries – public and private domains’, part 3 – a follow-up survey, (available at http://doi.org/10.5281/zenodo.1120237) doorn, peter and tjalsma, heiko (2007) ‘introduction: archiving research data’, springer science+business media b.v. hatcher, larry (1994) ‘sas system factor analysis and structural equation modelling’, sas institute horton, laurence (2016) ‘lse research data management data sharing objections faqs and noughts and crosses game’, (available at http://doi.org/10.5281/zenodo.61978) kim, youngseek and m. stanton, jeffrey (2012) ‘institutional and individual influences on scientists’ data sharing practices’, journal of computational science education, volume 3, issue 1 noble, susan; russel, celia and wiseman, richard (2012) ‘mind the gap: global data sharing’, iassist quarterly, vol. 35, n° 3 qualtrics (2010) ‘the 1936 election – a polling catastrophe’, available at (https://www.qualtrics.com/blog/the-1936-election-a-polling-catastrophe) rasmussen, karsten boye (2014) ‘social science metadata and the foundations of the ddi’, iassist quarterly, vol. 37, n° 1 ribeiro, cristina and matos fernandes, maria eugenia (2012) ‘data curation at u. porto: identifying current practices across disciplinary domains’, iassist quarterly, vol. 35, n° 4 tenopir, carol; d. dalton, elizabeth; allard, suzie; frame, mike; pjesivac, ivanka; birch, ben; pollock, danielle and dorsett, kristina (2015) ‘changes in data sharing and data reuse practices and perceptions among scientists worldwide’, (available at https://doi.org/10.1371/journal.pone.0134826) yang, meng-li (2013) ‘strategies of promoting the use of survey research data archive’, iassist quarterly, vol. 36, n° 1 https://doi.org/10.29173/iq993 http://doi.org/10.5281/zenodo.1120237 http://doi.org/10.5281/zenodo.61978 https://www.qualtrics.com/blog/the-1936-election-a-polling-catastrophe https://doi.org/10.1371/journal.pone.0134826 17/17 bonifacio, flavio; akullo, winny nekesa (2021), differences in data-sharing attitudes and behaviours, extended version to african data curators and data management experts, iassist quarterly 45(3-4), pp. 1-17. doi: https://doi.org/10.29173/iq993 17 endnotes 1flavio bonifacio, metis ricerche srl, via camerana 6, i-10128 torino, italy. contact e-mail: flavio.bonifacio@metis-ricerche.it. 2 winny nekesa akullo, public procurement and disposal of public assets authority, kampala, uganda. contact e-mail: winny.nekesa@yahoo.com 3 bonifacio, flavio, differences in data-sharing attitudes and behaviours, iassist quarterly, vol. 42 no. 3 (2018) 4tenopir, carol; d. dalton, elizabeth; allard, suzie; frame, mike; pjesivac, ivanka; birch, ben, et al. (2015) 5kim, youngseek and m. stanton, jeffrey (2012), doorn, peter and tjalsma, heiko, (2007), noble, susan; russel, celia and wiseman, richard, (2012), ribeiro, cristina ad matos fernandes, maria eugenia, (2012) yang, meng-li, (2013), rasmussen, karsten boye, (2014) 6 we refer to the annex 2: imdsm2017, iassist members data sharing mail, fall 2017. the mails reported suggestions about how to ask questions 7 for the method used building the indexes (pca analysis) and for questions concerning the sample see the quoted bonifacio, flavio, differences in…. while the subsample for italians and (not africans and not italians) have been weighted, the subsample of africans is not weighted. https://doi.org/10.29173/iq993 mailto:flavio.bonifacio@metis-ricerche.it mailto:winny.nekesa@yahoo.com 1/2 schwartz, ofira & hayslett, michele (2024), evaluating new technologies and organizational structures, iassist quarterly 48(4), pp. 1-2. doi: https://doi.org/10.29173/iq1149 the creative commons-attribution-noncommercial license 4.0 international applies to all works published by iassist quarterly. authors will retain copyright of the work and full publishing rights. editors’ notes: evaluating new technologies and organizational structures welcome to the last issue of iassist quarterly for 2024, iq 48(4). we are excited to share news of several developments that we have been working on over the last few months: the iassist qualitative social science and humanities data interest group (qsshdig) is planning an iassist quarterly special issue dedicated to the complexities of sharing qualitative data. for this special issue, we invite submissions of abstract proposals focused on the ethical challenges, methodological concerns, and labor involved in making qualitative data and research materials publicly available. the full cfp and details on how to submit an abstract can be viewed on the iassist quarterly website: https://iassistquarterly.com/index.php/iassist/announcement/view/7 . the deadline for proposing articles is january 31st (full articles won’t be needed until later). we are delighted to welcome minglu wang as a new iq editorial board member (as of october 2024). minglu is the research data management librarian in the open scholarship department at york university libraries, york, ontario, canada. among other qualifications, she brings experience as a member of the editorial board for acrl’s college & research libraries (c&rl) (2019–2025), and she led the project group for that board to investigate a data policy for c&rl. a new feature recently enabled on the ojs platform allows reviewers to link their profile with their orcid id. we mentioned last time that this will enable auto-loading of your articles to your orcid profile, but the other effect is that it provides an opportunity for reviewers to receive credit and be acknowledged for their professional contributions. note that the credit will merely note that you have served as a reviewer for the iq—it will not indicate which article(s) you reviewed. unfortunately, the iq editorial team had to retract a paper from publication this fall due to plagiarism. the paper titled “data protection and right to privacy legislation in kenya” by mankone, a. m. (2023), was published in iq, 47(3-4). the full retraction notice can be found here. this new issue of iq 48(4) presents four excellent papers. the first two evaluate methods to enhance findability of data deposited in data repositories. the subsequent two papers focus on organizational structure and improving organizational workflows. kokila jamwal in ”boosting data findability: the role of ai-enhanced keyword” examines the use of artificial intelligece (ai) to supplement keywords that may be missing or inaccurately defined as a method to improve metadata and boost data findability. the author suggests that using this relatively https://doi.org/10.29173/iq1149 https://iassistquarterly.com/index.php/iassist/announcement/view/7 https://orcid.org/ https://iassistquarterly.com/index.php/iassist/issue/view/155 https://iassistquarterly.com/index.php/iassist/article/view/1080 https://creativecommons.org/licenses/by-nc/4.0/ 2/2 schwartz, ofira & hayslett, michele (2024), evaluating new technologies and organizational structures, iassist quarterly 48(4), pp. 1-2. doi: https://doi.org/10.29173/iq1149 new technology may reduce the time and effort required by data repositories staff for data curation and may enhance data findability and usability. co-authors knut wenzig and xiaoyao han are examining the findability of data deposited in data repositories that are using ddi metadata standards. their paper ”state of ddi cloud” invetigates the availability and the comprehensive element usage of ddi standards across 29 repositories registered on re3data.org. based on their findings they provide recommendations for various stakeholders including the repositories, dataverse developers, re3data.org, and the ddi alliance. the article ”the ipums business process model: instituting a workflow mapping strategy to support archival processes” introduces the ipums workflow from external submission of data, harmonization process, documentation, extraction systems, and archival preservation of metadata. author diana magnuson explains the value of instituting this mapping approach, and demonstrates the power of a clear business process model for developing archival goals in an organizational setting in which the archive function is vital but secondary to the main product. in ”understanding motivations and future needs for data depoists at korea social sciences data archive”, authors hyowon kim, do won kim and jungwon yang evaluate the current data deposit process of the korea social science data archive (kossda). the data archive recently transitioned into an idependent researh center under seoul national univerity. using interviews with stakeholders, they identify future needs and suggest a long-term strategy to ensure that the archive meets the needs of the academic community it supports. wishing you a happy holidays season, and peace, health, and happiness in the new year. ofira schwartz and michele hayslett, december 2024 https://doi.org/10.29173/iq1149 16 iassist quarterly 2014 iassist quarterly linking thesauri elsst as a hub for social science data terms by lorna balkan and lucy bell1 abstract without controlled index terms, data retrieval within a data catalogue becomes at best hit and miss. the uk data archive manages two thesauri: the multilingual elsst thesaurus, and the monolingual hasset thesaurus, from which elsst is derived. over the last year, through funding from the uk’s economic and social research council (esrc), the archive has developed both of these thesauri, plus their management applications, and created linkages between them. extending iso 25964, the archive has developed a way of mapping its social science thesauri to facilitate cross-national data retrieval. it has also created skos formats and is developing a new and innovative application for thesaurus management which combines term visualisation with tree structures. the two thesauri now share a clearly defined, common set of core concepts, but have room for divergence. the new application will allow terms to be promoted to the core set, where they exhibit partial or exact equivalence, or be demoted to ‘non-core’. the application allows authorised language equivalents to be added to elsst terms. bundled suggestions will also be made for changes to the thesaurus terms or structure, linked to the tree structure. this paper describes the new application and the processes used within the archive for thesaurus management. keywords: : thesaurus management, thesaurus mapping, skos, multilingual thesauri introduction traditionally, thesauri have held a key role in data archiving as aids for searching and browsing. more recently, they are finding new applications in the world of linked data and the semantic web. the uk data archive2 has wide expertise in thesaurus development. it has been developing the monolingual social science thesaurus humanities and social science electronic thesaurus (hasset3) for over forty years, and for the last 10 years has been managing it in tandem with the related multilingual european language social science thesaurus (elsst4). recently it has received funding from the uk’s economic and social research council (esrc) for a 5-year project (2012-2017), cessda-elsst5, to update and develop both thesauri, and implement a new management application and processes. this paper describes work to date, challenges faced, and solutions adopted. background the work reported in this paper was motivated by two main concerns. the first was to update the thesauri managed by the uk data archive so that they can be exploited and shared more easily. the second was to find a way of managing them more efficiently. thesauri and other controlled vocabularies also play a vital role in linked data. iassist quarterly 2014 17 iassist quarterly the changing role of thesauri thesauri are a type of controlled vocabulary. controlled vocabularies “mandate the use of predefined, authorized terms that have been preselected by the designer of a vocabulary, in contrast to natural language vocabularies, where there is no restriction on vocabulary” (andritsos and keilty, 2014). specifically, thesauri belong to the class of knowledge organization systems (koss) that are “controlled vocabularies, which are organized and structured via different types of semantic relationships” (golub and tudhope, 2009). traditionally thesauri have been used to support humanmediated access to information. increasingly, they and other controlled vocabularies are being used for machine-to-machine communication and to underpin web services (dextre clarke and zeng, 2012). web services include terminology services, defined as ”a group of abstract services, presenting and applying vocabularies, their member concepts, terms and relationships, describing the meaning of terms and facilitating semantic interoperability. this is done for purposes of searching, browsing, discovery, translation, mapping, semantic reasoning, subject indexing and classification, harvesting, alerting etc.” (tudhope et al., 2006). examples of terminology services include the nerc vocabulary server6, umls7 and the bioportal8. terminology services may be used in isolation or in combination with a wide range of other web services (see golub and tudhope (2009) for a discussion of use cases). thesauri and other controlled vocabularies also play a vital role in linked data. linked data is ”an approach to data integration that employs ontologies, terminologies, uniform resource identifiers (uris) and the resource description framework (rdf9) to connect pieces of data, information and knowledge on the semantic web” (marshall et al., 2012). vocabularies published as linked data can be linked to other vocabularies, allowing databases indexed with one vocabulary to be searched using another (méndez and greenberg, 2012). some researchers, including shiri, also predict that controlled vocabularies will have a role to play in the new big data landscape: ”general purpose and domain-specific controlled vocabularies published as linked open vocabularies can not only be used to organize and represent structured data such as linked data repositories and semantic web applications, they can also be used to index, organize and analyze unstructured textual information that exists in several big data sources.”(shiri, 2014). to enable thesauri and other vocabularies to be fully exploited requires them to be interoperable. recent years have seen the emergence of new standards designed to promote interoperability. amongst the most important is the publication of iso 2596410, the new international guidelines on thesaurus construction. it is in two parts. part 1 covers the construction of monolingual and multilingual thesauri, while part 2 is concerned with interoperability and mapping between different vocabulary types, including thesauri. crucially iso 24964 contains an explicit data model that clearly distinguishes between concepts and the terms used to represent the concepts. dextre clarke and zeng (2012) argues that ”to perform on the semantic web, computer software needs an explicit data model that distinguishes between terms and concepts”. interoperability also requires common encoding schemes. simple knowledge organization system (skos11), which uses rdf, is emerging as the preferred standard in which to encode thesauri and other koss. many prominent thesauri have already been converted to skos and made available as linked data (see for example agrovoc (caracciolo et al., 2013), eurovoc12 and gemet13). key goals of the cessda-elsst project were, therefore, to convert the thesauri from a term-based to a concept-based model, following iso 25964, and to make them available as skos-based linked data. thesauri at the uk data archive thesaurus development has been a key activity of the uk data archive since its inception. hasset is the in-house thesaurus of the uk data service, and is used to index and search the service’s data collection, which, with over 6,000 datasets, is the largest social science data collection in the uk. hasset was originally derived from the unesco thesaurus, but has been developed in-house at the archive for over 40 years. it currently contains 4,743 preferred terms, and is available to external users under license. elsst is a multilingual thesaurus that is used in the cessda data portal. the consortium of european social science data archives (cessda) is a body that promotes the acquisition, archiving and distribution of electronic data for social science teaching and research in europe. elsst began life in 2000 as part of the eu-funded language independent metadata browsing of european resources (limber14) project, with the aim of enhancing cross-border data discovery and utilization. elsst has been developed further over the years through additional funding from the esrc and the university of essex and via other eu-funded projects including the multilingual access to data infrastructures of the european research area (madiera15) and cessda-preparatory phase project (cessdappp16). it is currently available in 9 languages (danish, english, finnish, french, german, greek, norwegian, spanish, swedish) with more in progress, including czech, lithuanian and romanian. it was originally derived from hasset and english continues to be the source language. it currently contains 3,286 english preferred terms. like hasset, it is available to external users under license. although the two thesauri have much in common (almost all elsst terms are also in hasset), historically, they have been managed on different platforms. not only has this led to much duplication of effort, since often the same information has to be entered in two different places, it has led to the two thesauri diverging without this being desired or obvious. one of the main aims of the cessda-elsst project was, therefore, to bring the two thesauri together onto one platform, and possibly merge them. it was expected that this would result in resource efficiencies and improved quality assurance. the design of the new system had to take account of the fact that the two thesauri work to differing time scales. hasset is constantly updated so that indexers within the uk data archive and members of the general public may browse the current version at all times, while elsst is moving towards an annual version release. another difference concerns validation procedures – hasset terms are validated in-house, while additions or changes to elsst source language terms require the approval of its international translators committee, of which the uk data archive is the chair. (translations of source terms are, however, the sole responsibility of translators or their institutions.) 18 iassist quarterly 2014 iassist quarterly the conversion of both thesauri to a concept-based model following iso 25964 was expected not only to improve their interoperability but to bring efficiency gains for their management. as dextre clarke and zeng (2012) points out: “benefits of adopting the [data] model include easier implementation by computers, consistency enforced in thesaurus construction and mapping, greater interoperability between thesauri and with other vocabularies, and enhanced performance at all stages of the thesaurus through development, management, and exchange.” iso 25964-2 also proved helpful to articulating the relationship between hasset and elsst. thesaurus development work development work on the thesauri was broken into a number of sub-tasks, some of which overlapped, and some of which are still ongoing. the main ones are described below. revising and updating terms and structures a major and ongoing part of the project is to review and update the terms and structure of both thesauri. as part of this work, a thorough review of the top terms is being undertaken, with a view to reducing them in number. reducing the number of top terms (currently 298 in hasset, 218 in elsst) will enhance the usefulness of the thesauri as a browsing aid. another focus of this sub-task is to review the terms themselves, removing redundancy where it has occurred, and ensuring that terms are up-to-date. up-to-date terms are of particular importance for automatic indexing. hasset has already been used for automatic indexing (el haj et al., 2013), and further experiments are planned in future. as terms are revised or added, scope notes are added where possible. these serve to define the semantic boundaries of the term and are useful not just to the thesaurus developers (in particular to elsst developers when looking for a translation of an english term), but also to users of the thesaurus. iso 25964 will be followed wherever possible. from term-based to concept-based model as mentioned above, an important part of the cessda-elsst project was to move the thesauri from a term-based to a conceptbased model, following iso 25964. thus preferred terms become labels for concepts, and in a multilingual thesaurus like elsst, different language versions of preferred terms are just alternative labels for the same concept. each concept has its own guid (globally unique identifier). in a term-based thesaurus semantic relationships are established between the terms themselves. in a concept-based thesaurus certain semantic relationships (e.g. hierarchical and associative relationships) are established between the concepts, and others (e.g. equivalence) between terms (dextre clarke and zeng, 2012). the concept-based model has advantages over the term-based model, not just for thesaurus management, but also for indexing. documents are associated with concepts, not terms, thus changes involving preferred and non-preferred terms do not impact on indexing (pastor-sanchez et al., 2009). skos in order to promote interoperability, converting both thesauri to simple knowledge organization system (skos) was a crucial goal of the project. the main objective of skos is to enable the easy publication of koss for the semantic web. like iso 25964, skos is concept-based, and, according to the iso 25964 homepage, care has been taken by iso 25964 developers to maintain compatability with skos (see iso 25964). elsst is being converted to skos as part of the cessda-elsst project – hasset was already converted to skos during the skos-hasset17 project (bell, 2013). the skos version of hasset is available to external users under license, and the skos version of elsst will be available in early 2015.the skos versions of both thesauri are implemented using guids and brightstardb18 for the triple stores, and published via pubby19 , which provides a browseable, meaningful view of the thesaurus (bell, 2012). defining the basics an important prerequisite to designing the new thesaurus management system was to establish all the possible elements and relationships in each thesaurus. the three basic relationships within a thesaurus are as defined in iso 25964: equivalence, which holds between a preferred term (pt) and a non-preferred term or use for (uf); hierarchical, which holds between a broader term (bt) and a narrower term (nt); and associative, which holds between related terms (rts). in the legacy systems, scope notes in both thesauri included information about meaning as well as usage and history of a term. in the new system, this information will be separated into different types of notes, namely scope notes, use notes and history notes respectively. historically, scope notes in elsst have also included ’translation notes’ that describe any difference in meaning between the english source term and its equivalent in another language. in the new model, translation notes will be also recorded in a separate field. these changes were introduced to make it easier for both users and developers to distinguish the different types of information, and additionally, to help developers keep track of the differences between hasset and elsst concepts. axioms and constraints once the basic elements of the thesauri were established, the next step was to define the axioms and constraints that hold among concepts within and between each thesaurus. an example of the former type of constraint is the requirement that a term may be a use for (uf) to only one pt within the thesaurus (or same language version of the thesaurus, in the case of elsst). inter-thesaural axioms and constraints were refined and updated as a result of the thesaurus alignment exercise, described below. thesaurus alignment exercise the aim of the alignment exercise was to see whether the two thesauri could be merged. all the terms and relationships that were in elsst, not hasset, were examined, and the differences between the two thesauri resolved wherever possible. further alignment work will look at the terms and relationships that are in hasset, not elsst, to see whether they should be brought into elsst. results from the first part of the alignment exercise suggested that, instead of forcing the two thesauri to merge, their common set of core concepts should be kept identical wherever possible, but allowed to diverge in clearly defined ways. rather than being seen in terms of merging, their relationship can best be described in terms of a mapping. in this way, both thesauri can retain their integrity and identity. iassist quarterly 2014 19 iassist quarterly mapping elsst to hasset elsst and hasset will have a set of shared or ’core’ concepts. noncore concepts will also be possible in both thesauri. a mapping relationship, based on the equivalence relationship defined in iso 25964-2, will hold between all core concepts. iso 25964-2 defines three types of mapping between thesauri: equivalence, hierarchical, and associative. an equivalence mapping is established when matching concepts are found in two or more different vocabularies potentially with different preferred term labels. equivalence may be ’simple’ (when the two thesauri contain concepts that are identical in scope), or ’compound’ (where a concept represented in one vocabulary with just one preferred term may be represented in another vocabulary by a combination of two or more concepts/terms). simple equivalence mappings may also be either exact or inexact. exact equivalence arises when ”the concepts can be used interchangeably across all the applications that can be envisaged for the mapping” (iso 25964-2, section 11.2). inexact equivalence, by contrast, arises when the concepts are equivalent in some contexts but not others, or where concepts have overlapping scopes or small differences of connotation. the equivalence mapping between elsst and hasset is defined as follows. core concepts will exhibit either ’exact’ or ’close’ equivalence. for ’exact equivalence’ to hold, the concepts must have the same preferred terms (pts), broader terms (bts), scope note and scope note source. ’close equivalents’ will only be required to have the same pts and bts. in both cases, all other associated metadata may differ, including: use for (ufs), narrower terms (nts), related terms (rts), use notes, etc. note that these definitions are stricter than the definition of iso 25964-2 equivalence mapping, since they demand identity at structural and term/linguistic level, as well as semantic level. the iso 25964-2 definition of equivalence, by contrast, is entirely semantic. the elsst/hasset ‘exact equivalence’ corresponds to ‘exact simple equivalence’ in iso 25964. we may expect any difference between the scope notes of ‘ close equivalents’ in elsst and hasset also to be small enough for ‘exact simple equivalence’ to hold in iso terms. that is to say, the difference in meaning between ’exact’ and ’close’ elsst/hasset equivalents will have no significant impact on information retrieval. not all restrictions on the relationship between core concepts in the two thesauri can be captured by axioms and constraints (for example, when scope notes and other metadata may differ), and developers will rely on the reporting functions of the thesaurus management system to keep track of all the differences between the two thesauri. work continues on when shared concepts may differ. thesaurus management system all elsst partners contributed to the requirements gathering of the new thesaurus management system. they were also invited to give feedback on the first prototype, and their comments were used to inform the final version. the system is now complete and has received excellent feedback from the testing by translators. figure 1 visual graph view of the concept nurses 20 iassist quarterly 2014 iassist quarterly the central feature of the new thesaurus management system is that the two thesauri now share the same database. this is essential to managing them in an efficient manner. however, the two thesauri will have separate user interfaces. hasset terms will be linked to the studies at the uk data service that have been indexed with them, while elsst will provide a link to its multilingual equivalents. the user interfaces otherwise have identical features. a novel feature of both the user and management interfaces is the implementation of a visualisation tool for navigation, in addition to the traditional tree structures. as data are becoming ever bigger and more complex, data visualisation is becoming increasingly popular as an alternative and more user-friendly way of viewing data. data visualisation of knowledge organization systems in particular is a lively research topic (see for example katifori et al., 2007 for an overview) and has been the focus of many recent workshops (see for example slavic et al. 2013). the visualisation solution adopted for the new thesaurus management system is an interactive tool which presents concepts to the user in a colourcoded, expandable graph (see figure 1). a well-designed management interface is key to the smooth management of both thesauri. management permissions are controlled via shibboleth20, and authorised users can, as appropriate, suggest, discuss, and implement changes or translations, in the relevant thesaurus. core concepts can be ‘demoted’ to non-core concepts in either elsst or hasset, and conversely, non-core concepts in hasset can be ‘promoted’ to core concepts. an innovative feature is the suggestions area where a number of proposed changes can be bundled together as one suggestion, since a change to one term is often associated with a change to another. the bundled suggestions are also linked to the tree structure. the suggestions area also provides a place where changes to terms can be discussed and agreed with external partners. since it is essential for the uk data service developers to keep track of the differences between hasset and elsst core concepts, and since not all of these can be captured by axioms and constraints, reporting functions play a vital role. to facilitate this, terms will be deprecated, but never deleted. technologies used for implementing the system include ajax scripting for the user interface, to minimise post backs and maximise the user experience, linq to xml, and solr. conclusion and future directions the cessda-elsst project has made good progress to date. development work in this first phase has concentrated on converting the thesauri to a concept-based model, creating a skos version of elsst, aligning the two thesauri, and defining the relationship between them. these in turn have enabled the design and implementation of the new thesaurus management system. archive staff and elsst translators are now keen to use the system in their everyday work. future work will include refining the structure, and updating the content of the two thesauri further, as well as possible enhancements to the management system, based on user feedback. the project is expected to produce significant gains for both end users and developers. uk developers will benefit from a new and improved management application, which will allow them to keep track of the differences between the two thesauri and thus save time and effort. they will also have a more user-friendly platform for managing suggestions and changes to elsst with their international partners. end users will benefit from the updated content and structure of both thesauri, and improved access to them via the new user interfaces. compliance with iso 25964 and conversion to skos will not only enhance elsst’s status as a hub for social science data terms but open it up to the world of the linked data and the semantic web. acknowledgements the authors wish to gratefully acknowledge the work and contributions of all individuals who have contributed to the work reported in this paper, including colleagues at the uk data archive and elsst partners. references andritsos, periklis and keilty, patrick (2014) level-wise exploration of linked and big data guided by controlled vocabularies and folksonomies. advances in classification research online. 24(1). doi:10.7152/acro.v24i1.14670 (available at http://journals.lib. washington.edu/index.php/acro/article/view/14670/12310) bell, darren (2012) from tuples to triples: applying skos to hasset – a technical overview. skos-hasset blog. (available at http://hassetukda.wordpress.com/2012/12/20/ from-tuples-to-triples-applying-skos-to-hasset-a-technical-overview/) bell, lucy (2013) skos-hasset: help for users. (available at http://www. data-archive.ac.uk/media/393116/skos-hasset-help.pdf ) caracciolo, caterina; stellato, armando; morshed, ahsan; johannsen, gudrun; rajbhandari, sachit; jaques, yves; and keizer, johannes (2013) the agrovoc linked dataset. semantic web, 2013 4(3) pp. 341-348. (available at http://eprints.rclis.org/20648/) dextre clarke, stella and zeng, marcia lei (2012) from iso 2788 to iso 25964: the evolution of thesaurus standards towards interoperability and data modeling. information standards quarterly. 24(1). (available at http://www.niso.org/publications/isq/2012/v24no1/ clarke/) el-haj, mahmoud; balkan, lorna; barbalet, suzanne; bell, lucy and shepherdson, john (2013) an experiment in automatic indexing using the hasset thesaurus. the 5th computer science and electronic engineering conference (ceec’13), ieee xplore. 17-18 september 2013, university of essex. doi: 10.1109/ ceec.2013.6659437 (available at http://ieeexplore.ieee.org/stamp/ stamp.jsp?tp=&arnumber=6659437) golub, koraljka and tudhope, douglas (2009) terminology registry scoping study (trss) final report. (available at http://www.jisc. ac.uk/media/documents/programmes/sharedservices/trss-reportfinal.pdf ) katifori, akrivi; halatsis, constantin; lepouras, george; vassilakis, costas and giannopoulou, eugenia (2007) ontology visualization methods – a survey. acm computing surveys, 39(4). (available at http://dl.acm. org/citation.cfm?id=1287621) marshall, m. scott; boyce, richard; deus, helena f.; zhao, jun; willighagen, egon l.; samwald, matthias; pichler, elgar; hajagos, janos; prud’hommeaux, eric and stephens, susie iassist quarterly 2014 21 iassist quarterly (2012) emerging practices for mapping and linking life sciences data using rdf a case series. web semantics: science, services and agents on the world wide web. volume 14. july 2012. pp. 2–13. (available at http://www.researchgate.net/ publication/253234505_emerging_practices_for_mapping_and_ linking_life_sciences_data_using_rdf__a_case_series) méndez, eva and greenberg, jane (2012) linked data for open vocabularies and hive’s global framework, el profesional de la informacion. may/june 2012. 21(3). (available at http://www. elprofesionaldelainformacion.com/contenidos/2012/mayo/03_eng. pdf ) pastor-sanchez, juan-antonio; martínez mendez, francisco javier and rodríguez-muñoz, josé vicente (2009) advantages of thesaurus representation using the simple knowledge organization system (skos) compared with proposed alternatives. information research. 14(4). december 2009. (available at http://files.eric.ed.gov/fulltext/ ej869364.pdf ) shiri, ali (2014) linked data meets big data: a knowledge organization systems perspective, advances in classification research online, 24(1). doi:10.7152/acro.v24i1.14672. (available at http://journals.lib. washington.edu/index.php/acro/article/view/14672/12312) slavic, aida; akdag saha, almila and davies, sylvie (eds.) (2013): classification and visualization: interfaces to knowledge: proceedings of the international udc seminar, 24-25 october 2013, the hague, the netherlands. ergon verlag, würzburg. tudhope, douglas; koch, traugott and heery, rachel (2006) terminology services and technology: jisc state of the art review. (available at http://www.ukoln.ac.uk/terminology/jisc-review2006. html) notes 1. both authors work at the uk data archive, university of essex. lorna balkan is cessda-elsst co-ordination officer. she can be contacted at balka@essex.ac.uk. lucy bell is functional director, data access. she can be contacted at lajbell@essex.ac.uk. this paper was presented in the ‘harmonization, thesauri and indexing‘ session at iassist 2014 2. http://www.data-archive.ac.uk 3. http://hasset.ukdataservice.ac.uk 4. http://elsst.ukdataservice.ac.uk 5. http://ukdataservice.ac.uk/about-us/projects/cessda-elsst/details. aspx 6. http://www.bodc.ac.uk/products/web_services/vocab/ 7. http://www.nlm.nih.gov/research/umls/ 8. http://bioportal.bioontology.org/ 9. http://www.w3.org/rdf/ 10. http://www.niso.org/schemas/iso25964/ 11. http://www.w3.org/2004/02/skos/ 12. http://open-data.europa.eu/en/data/dataset/eurovoc 13. http://www.eionet.europa.eu/gemet/exports/en/rdf/ 14. http://www.data-archive.ac.uk/about/projects/limber 15. http://www.data-archive.ac.uk/about/projects/madiera 16. http://www.data-archive.ac.uk/about/projects/cessda-ppp 17. http://www.data-archive.ac.uk/find/our-projects/skos-hasset 18. http://brightstardb.com/ 19. http://www.w3.org/2001/sw/wiki/pubby 20. https://shibboleth.net/ l^ssist newsletter vol.1, no. 4 discussion paper/ michel melton an overview of display terminals michael melton battelle columbus laboratories columbus, ohio overview display terminals can be generally classified into three categories: (1) dumb, (2) semi-stupid, and (3) intelligent. dumb terminals offer a limited number of functions and most closely resemble teletypes. semi-stupid terminals usually offer a certain amount of features, most likely, data input editing and formatting. some terminals in this class can be tailored to fit a specific application via a limited programming capability. intelligent terminals are supported by software programs. the vendors typically provide an operating system, an assembler or a compiler, i/o utilities, and one or more application programs such as text editing or data entry. programmable terminals provide the user with the highest degree of flexibility, permitting the terminal to be tailored to the user's environment. some programmable terminals are available as turnkey terminals, which means that the program is furnished by the vendor and the terminal is ready to use as soon as it is installed. cost is usually proportional to capability. dumb display terminals are the least expensive at $1000 to $2000. intelligent terminals range upward from $6000. expensive items include increased memory, additional display units, and peripherals such as diskette or disk storage and printers. semi-stupid terminals are priced somewhere inbetween. some brand names in dumb terminals include applied digital data systems, beehive, infoton, and lear siegler. some of the semi-stupid terminals include hewlett-packard 2640, ibm 3270, icc 40+, and the univac uniscope. some intelligent terminals are applied digital data systems 70, the beehive b800, datapoint, four-phase systems incoterm, sycor, and univac uts400. most terminals introduced on the market in the last three years have been microprocessed-controlled (i.e., intelligent). microprocessors are cheap, cut design, development, and production costs, and they lend themselves to a variety of applications that can be implemented by the vendor or the user. computer programs which control the functions of an intelligent terminal are called firmware. the user can control the terminal's functions by either changing a set of parameters or by adding new firmware to the microprocessor. display functions some different terminal display functions are: (1) color few display terminals offer color but some offer up to eight colors (2) reverse video a negative image of the data, data normally displayed in white or dark background is displayed in black or white background (3) programmable brightness level (4) character or field blinking 19 sist newsletter vol.1, no. 4 (5) roll or scroll data is rolled up or down the screen, permits the user to scan a large volume of data (6) paging data is stored on pages (a full screen) user is able to review any selected page. editing functions editing features include: (1) character deletion (2) line insertion (3) line deletion (4) erase (5) character repeat. external i/o devices can add flexibility to the applications possibilities for display terminals. a cassette tape drive or diskette drive can be used to store display formats, data to be transmitted, or user programs. a printer can provide hard copy. selecting a terminal some questions you should ask yourself when selecting a display terminal are: (1) what are the essential parameters for a display terminal that will satisfy your needs? (2) who supplies the terminals with the features you desire? (3) maintenance provisions? (4) talk to users concerning problems encountered when installing it, failures that have occurred, and any incompatibilities. discussion paper/alice robbin the issue of confidential data: the need for formulation of policy by the data archive and library by alice robbin data and program library service university of wisconsin-madison the issue of confidential data i. during the last decade there has been increased concern about the problems of confidentiality involved in the collection and dissemination of individual microdata. concern has revolved around the government's perceived need to collect increasing amounts of information at a microdata level for social policy formulation and evaluation, types of information which potentially compromise iassist quarterlyiassist quarterly abstract ontology engineers and experts from the social, behavioral, and economic sciences developed a data discovery ontology covering a subset of both the ddi codebook and lifecycle models, and implemented a rendering of ddi xml instances to rdf (resource description framework). the main goals associated with the design process of the ddi ontology were to reuse widely adopted and accepted ontologies like dublin core (dc) and simple knowledge organization system (skos) and also to define meaningful relationships to the rdf data cube vocabulary. now, organizations have the possibility to publish their ddi data and metadata in rdf and link it with many other datasets from the linked open data (lod) cloud. as a consequence, a huge number of related ddi instances can be discovered, queried, connected, and harmonized. the combination of ddi metadata (as well as data) from several organizations, based on this rdf discovery (disco) vocabulary, will enable powerful derivations of implicit knowledge out of explicitly stated pieces of information. keywords: semantic web, linked data, ontology design, ddi data documentation initiative: background overview the ddi specification describes social science data, data covering human activity, and other data based on observational methods measuring real-life phenomena. ddi supports the entire research data lifecycle. ddi metadata accompany and enable data conceptualization, collection, processing, distribution, discovery, analysis, repurposing, and archiving. metadata is structured information that describes, explains, locates, or otherwise makes it easier to retrieve, use, or manage data (niso press, 2004). ddi does not invent a new model for statistical data. it formalizes state of the art concepts and common practice in this domain. ddi focuses on both microdata and aggregated data. it has its strength in microdata -data on the characteristics of units of a population, such as individuals or households, collected by, for example, a census or a survey. statistical microdata are not to be confused with microdata in html, an approach to nest semantics within web pages. aggregated data (e.g., multidimensional tables) are likewise covered by ddi. they provide summarized versions of the microdata in the form of statistics like means or frequencies. publicly accessible metadata of good quality are important for finding the right data. this is especially the case if access to microdata is restricted due to potential risk of disclosure of respondent identities. ddi is currently specified in xml schema, organized in multiple modules corresponding to the individual stages of the data lifecycle, and includes over 800 elements (ddi lifecycle). a specific ddi module (using the simple dublin core namespace) allows for the capture and expression of native dublin core elements, used either as references or as descriptions of a particular set of metadata. this is used for citation of the data, parts of the data documentation, and external material in addition to the richer, native ddi. this approach supports applications that understand the dublin core xml, but do not understand ddi. ddi is aligned with other metadata standards as well, with sdmx6 (time-series data) for exchanging aggregate data, iso/iec 11179 (metadata registry) for building data registries such as question, variable, and concept banks (iso/iec, ddi-rdf discovery – a discovery model for microdata by thomas bosch1, olof olsson2, benjamin zapilko3, arofan gregory4, and joachim wackerow5 ddi has its strength in the domain of social, economic, and behavioral data iassist quarterly 2013 17 18 iassist quarterly 2014/2015 iassist quarterly 2004), and iso 19115 (geographic standard) for supporting gis (geographic information system) users (iso 19115-1:2003, 2003). goals ddi supports technological and semantic interoperability in enabling and promoting international and interdisciplinary access to and use of research data. structured metadata with high quality enable secondary analysis without the need to contact the primary researcher who collected the data. comprehensive metadata (potentially along the whole data lifecycle) are crucial for the replication of analysis results in order to enhance research transparency. ddi also enables the reuse of metadata of existing studies (e.g., questions, variables) for designing new studies, an important ability for repeated surveys and for comparison purposes. ddi supports researchers who follow the above mentioned goals. ddi users a large community of data professionals, including data producers (e.g., of large, academic international surveys), data archivists, data managers in national statistical agencies and other official data producing agencies, and international organizations use the ddi metadata standard. the ddi alliance hosts a comprehensive list of projects using the ddi7. academic users include the uk data archive at the university of essex8, the dataverse network at the harvard-mit data center9, and the inter-university consortium for political and social research (icpsr) at the university of michigan10. official data producers in more than 50 countries include the australian bureau of statistics (abs)11 and many national statistical institutes of the accelerated data program for developing countries12. examples of international organizations using ddi are unicef, the multiple indicator cluster surveys (mics)13, the world bank14, and the global fund to fight aids, tuberculosis and malaria15. ddi history and versions the ddi project, which started in 1995, has steadily gained momentum and evolved to meet the needs of the social science research community. in 2003, the ddi alliance was established to develop and promote the ddi specification and associated tools, education, and outreach program. the ddi alliance is a selfsustaining membership organization whose institutional members have a voice in the development of the ddi specification. to ensure continued support and ongoing development of the standard, ddi has been branched into two separate development lines. ddi-codebook (formerly ddi2) is a more light-weight version of the standard, intended primarily to document simple survey data for archival purposes. encompassing all of the ddi-codebook specification and extending it, ddi-lifecycle (formerly ddi3, first version published in 2008) is designed to document and manage data across the entire data lifecycle, from conceptualization to data publication and analysis and beyond. data lifecycle the common understanding is that both statistical data and metadata are part of a data lifecycle (figure 1 displays this lifecycle -it is described in more detail on the ddi alliance website16). multiple institutions are involved in the data lifecycle, which is an interactive process with multiple feedback loops. data documentation is a process, not an end condition where a final status of the data is documented. rather, metadata production should begin early in a project, and metadata should continue to be captured at the source as data come into being. the metadata can then ideally be reused along the data lifecycle. such practice would incorporate documentation as part of the research method (jacobs et al., 2004). a paradigm change would be enabled: on the basis of the metadata, it becomes possible to drive processes and generate items like questionnaires, statistical command files, and web documentation, if metadata creation is started at the design stage of a study (e.g., survey) in a well-defined and structured way. limitations ddi has its strength in the domain of social, economic, and behavioral data. ongoing work focuses on the early phases of survey design and data collection as well as on other data sources like register data. the next major version of ddi will incorporate the results of this work. it will be opened to other data sources and to data of other disciplines. related work with respect to documenting data, there are several relevant metadata standards like sdmx (statistical data and metadata exchange) for the representation and exchange of aggregated data, iso 19115 (iso 19115-1:2003, 2003) for geographic information, and premis17 for preservation purposes. the metadata registry standard iso 11179 (iso/iec, 2004) addresses the modeling figure 1. ddi data lifecycle iassist quarterly 2014/2015 19 iassist quarterly of metadata, e.g., reference models, and registries. however, there are as yet few adequate rdf-based vocabularies for documenting data. ddi-rdf for discovery, or disco, has a clearly defined focus on describing microdata, which has not been covered to this extent by other established vocabularies yet. therefore it fits well alongside other metadata standards on the web and can clearly be distinguished. connection points to classes or properties of other vocabularies ensure equivalent or more detailed possibilities for describing entities or relationships. an rdf expression of the simple dublin core specification exists which could be used for citation purposes (dcmi, 2008). furthermore, the dcmi metadata terms (dcmi, 2010) have been applied when suitable for representing basic information about publishing objects on the web as well as for haspart relationships. for representing concepts that are organized in ways similar to thesauri and classification systems, classes and properties of simple knowledge organization system (skos)18 have been used. some aspects of ddi-rdf are already similarly represented in other metadata vocabularies, e.g., data management and documentation. the vocabulary of interlinked datasets (void)19 represents relationships between multiple datasets, while the provenance vocabulary20 provides the possibility to describe information on ownership and can be used to represent and exchange provenance information generated in different systems and under different contexts. in this context, a study can be seen as a data-producing process and a logical dataset as its output artifact. data catalog vocabulary (dcat)21 is an rdf vocabulary designed to facilitate interoperability between data catalogs published on the web. by using dcat to describe datasets in data catalogs, publishers increase discoverability and enable applications easily to consume metadata from multiple catalogs. an established rdf metadata vocabulary, which seems similar to ddi-rdf at first glance, is the rdf data cube vocabulary (cyganiak et al., 2010). this model maps the sdmx information model to an ontology and is therefore compatible with the cube model that underlies sdmx. it can be used for representing aggregated data (also known as macrodata) such as multidimensional tables. aggregate data are data derived from microdata by statistics on groups or aggregates, such as counts, means, or frequencies. a dataset presented with the data cube vocabulary consists of a set of values organized along a group of dimensions, which is comparable to the representation of data in an online analytical processing system. in the data cube vocabulary associated metadata are added. ddi as linked data statistical domain experts (core members of the ddi alliance technical implementation committee, representatives of national statistical institutes, national data archives) and linked open data community members have chosen the ddi elements that are seen as most important to solve problems associated with diverse identified use cases around data discovery. widely accepted and adopted vocabularies are reused to a large extent. there are features of ddi that can be addressed through other vocabularies, such as: describing metadata for citation purposes using dublin core, describing aggregated data like multidimensional tables using the rdf data cube vocabulary22, and delineating code lists, category schemes, mappings between them, and concepts like topics using skos. this section serves as an overview of the conceptual model for the disco vocabulary. more detailed descriptions of all the properties are given in the specification23 and a conference paper (bosch et al. 2012). class  overview ´ unionª variablequestion instrument questionnaire dcat:dataset logicaldataset skos:concept analysisunit skos:concept universe study studygroup 1..* product 0..* 0..* ingroup 0..1 1..* variable 0..* 0..*universe1 1..* containsvariable 0..* 0..* question 1..*0..* universe 1 0..* analysisunit 0..1 0..* universe 1 0..*question0..* 0..* analysisunit 0..1 0..* universe 1..* figure 2. ddi-rdf discovery vocabulary (disco) 20 iassist quarterly 2014/2015 iassist quarterly overview figure 2 provides a diagram of the conceptual model containing a small subset of the ddi-xml specification24. to understand the ddi discovery vocabulary, there are a few central classes, which can serve as entry points. the first of these is study. a study represents the process by which a dataset was generated or collected. literal properties include information about the funding, organizational affiliation, abstract, title, version, and other such high-level information. in some cases, where data collection is cyclic or ongoing, datasets may be released as a studygroup, where each cycle or ”wave” of the data collection activity produces one or more datasets. this is typical for longitudinal studies, panel studies, and other types of ”series”. in this case, a number of study objects would be collected into a single studygroup. datasets have two representations: a logical representation, which describes the contents of the dataset, and a physical representation, which is a distributed file holding that data. it is possible to format data files in many different ways, even if the logical content is the same. logicaldataset represents the content of the file (it is organized into a set of variables). the logicaldataset is an extension of the dcat:dataset. physical, distributed files are represented by the datafile, which is itself an extension of dcat:distribution. when it comes to understanding the contents of the dataset, this is done using the variable class. variables provide a definition of the column in a rectangular data file, and can associate it with a concept and a question (the question in the questionnaire which was used to collect the data). variables are related to a representation of some form, which may be a set of codes and categories (a “codelist”) or may be one of other normal data types (datetime, numeric, textual, etc.). codes and categories are represented using skos concepts and concept schemes. data are collected about a specific phenomenon, typically involving some target population, and focusing on the analysis of a particular type of subject. these are respectively represented by the classes universe and analysisunit. if, for example, the adult population of finland is being studied, the analysisunit would be individuals or persons. unique identifiers for specific ddi versions are used for easing the linkage between ddi-rdf metadata and the original ddi-xml files. every element can be related to any foaf:document (ddi-xml files) using dcterms:relation. any entity can have version information (owl:versioninfo). however, the most typical cases are the versioning of the metadata (the ddi or the rdf file), the versioning of the study (as a study goes through the lifecycle from conception through data collection), and the versioning of the data files. every logicaldataset may have access rights statements (dcterms:accessrights) and class  logicaldatasets  and  datafiles dcat:dataset logicaldataset -­‐   dcterms:title    :rdf:langstring -­‐   ispublic    :xsd:boolean dcat:distribution dcterms:dataset datafile -­‐   casequantity    :xsd:nonnegativeinteger -­‐   dcterms:description    :rdf:langstring -­‐   owl:versioninfo    :string descriptivestatistics categorystatistics -­‐   cumulativepercentage    :xsd:decimal -­‐   frequency    :xsd:nonnegativeinteger -­‐   percentage    :xsd:decimal -­‐   weightedcumulativepercentage    :xsd:decimal -­‐   weightedfrequency    :xsd:nonnegativeinteger -­‐   weightedpercentage    :xsd:decimal summarystatistics -­‐   invalidcases    :xsd:nonnegativeinteger -­‐   maximum    :xsd:decimal -­‐   mean    :xsd:decimal -­‐   median    :xsd:decimal -­‐   minimum    :xsd:decimal -­‐   mode    :xsd:decimal -­‐   standarddeviation    :xsd:decimal -­‐   validcases    :xsd:nonnegativeinteger -­‐   weightedinvalidcases    :xsd:nonnegativeinteger -­‐   weightedmean    :xsd:decimal -­‐   weightedmedian    :xsd:decimal -­‐   weightedmode    :xsd:decimal -­‐   weightedvalidcases    :xsd:nonnegativeinteger 0..* statisticsdatafile 0..* 0..* datafile 0..* figure 3. logicaldatasets and datafiles iassist quarterly 2014/2015 21 iassist quarterly licensing information (dcterms:license) attached to it. studies, logical datasets, and data files may have spatial (dcterms:spatial), temporal (dcterms:temporal), and topical (dcterms:subject) coverage. studies and studygroups a simple study supports the stages of the full data lifecycle in a modular manner. as noted above, a study represents the process by which a dataset was generated or collected, and a number of study objects can be collected into a single studygroup. studies may have multiple disco:instrument relationships to instruments and may have disco:datafile connections with 0 to n datafiles. studies are associated with 0 to n variables using the object property disco:variable. studies may have multiple logicaldatasets (disco:product). studies or studygroups (the union of study and studygroup) may have an abstract (dcterms:abstract), a title (dcterms:title), a subtitle (disco:subtitle), an alternative title (dcterms:alternative), a purpose (disco:purpose), and information about the date and time the study was made publicly available (dcterms:available). disco:kindofdata describes the kind of data documented in the logical product(s) of a study (e.g., survey data or administrative data). disco:ddifile leads to foaf:documents which are the ddi-xml files containing further descriptions of the study or the studygroup. creators (dcterms:creator), contributors (dcterms:contributor), and publishers (dcterms:publisher) of studies and studygroups are foaf:agents which are either foaf:persons or org:organizations whose members are foaf:persons. studies and studygroups may be funded by (disco:fundedby) foaf:agents. the object property disco:fundedby is defined as sub-property of dcterms:contributor. universe is the total membership or population of a defined class of people, objects, or events. analysisunit is the particular type of subject being analyzed, for example, individuals or persons. studies and groups of studies must have 1 to n universes which are subclasses of skos:concepts. for universes one can state definitions using skos:definition. the union of study and studygroup may have 0 or 1 analysisunit reached by the object property disco:analysisunit. analysisunit is specified as a sub-class of skos:concept. logical datasets, data files, descriptive statistics, and aggregated data as noted, datasets have a logical representation, which describes the contents of the dataset, and a physical representation, which is a distributed file holding that data. it is possible to format data files in many different ways, even if the logical content is the same. logicaldataset represents the content of the file (its organization into a set of variables). the logicaldataset is an extension of dcat:dataset. physical, distributed files containing the microdata datasets are represented by datafile, which are sub-classes of dcterms:datasets and dcat:distribution. an overview of the microdata can be given either by descriptive statistics or aggregated data. descriptivestatistics may be minimal, maximal, mean values, and absolute and relative frequencies. qb:dataset originates from the rdf data cube vocabulary25, an approach to map the sdmx information model to an ontology. a dataset represents aggregated data such as multidimensional tables. summarystatistics pointing to variables and categorystatistics pointing to categories and codes are both descriptive statistics. variables, variable definitions, representations, and concepts when it comes to understanding the contents of the dataset, this is done using the variable class. variables provide a definition of the column in a rectangular data file, and can associate it with a concept, and a question. variable is a characteristic of a unit being observed. a variable might be the answer to a question, have an administrative source, or be derived from other variables. variabledefinitions encompass study-independent, reusable parts of variables like occupation classification. questions, variables, and variabledefinitions may have representations. representation is defined as a sub-class of the union of rdfs:datatype (e.g., numeric or textual values) and skos:conceptscheme, as for example questions may have as their response domain a mixture of a numeric response domain containing numeric values (rdfs:datatype) and a code response domain (skos:conceptscheme) -a set of codes and categories (a ”codelist”). codes and categories are represented using skos concepts and concept schemes. skos defines the term skos:concept, which is a unit of knowledge created by a unique combination of class  variables variable -­‐   dcterms:description    :rdf:langstring +   skos:notation    :rdfs:literal -­‐   skos:preflabel    :rdf:langstring variabledefinition +   dcterms:description    :rdf:langstring -­‐   skos:preflabel    :rdf:langstring skos:concept -­‐   skos:definition    :rdf:langstring -­‐   skos:notation    :rdfs:literal -­‐   skos:preflabel    :rdf:langstring representation 0..* skos:narrower 0..* 0..* skos:broader 0..* 0..* representation 0..* 0..* concept 1 0..* representation 1 0..* basedon 0..1 0..* concept 1 figure 4. variables 22 iassist quarterly 2014/2015 iassist quarterly characteristics. in the context of statistical (meta)data, concepts are abstract summaries, general notions, or knowledge of a whole set of behaviors, attitudes, or characteristics which are seen as having something in common. concepts may be associated with variables and questions. a skos:conceptscheme is a set of metadata describing statistical concepts. skos:concept is reused to a large extent to represent ddi concepts, codes, and categories. data collection the data for the study are collected by an instrument. the purpose of an instrument, e.g., an interview, a questionnaire, or another entity used as a means of data collection, is in the case of a survey to record the flow of a questionnaire, its use of questions, and additional component parts. a questionnaire contains a flow of questions. a question is designed to elicit information on a subject, or sequence of subjects, from a respondent. the next figure visualizes the datatype and object properties of instrument and question. one can describe (dcterms:description) instruments and associate labels (skos:preflabel) to instruments. instruments may have multiple external documentation files of the type foaf:document. questionnaires are special instruments having at least one collection mode (disco:collectionmode) which is a skos:concept. questionnaires must contain at least one question. questions have a question text (disco:questiontext), a label (skos:preflabel), exactly one universe (disco:universe), multiple concepts (disco:concept), and at least one response domain (disco:responsedomain). use cases this section describes the scenarios that the ddi-rdf discovery vocabulary was designed to support. these are not formal uml use cases -instead, they are scenarios for the possible use of the vocabulary, based on an analysis of existing search interfaces and known behaviors for those looking for research data. the process around these discovery scenarios is to posit the thinking of the researcher/user seeking to find data, to identify needed classes and properties in the vocabulary, and then to render the search as it might be implemented. enhancing discovery of data by providing related metadata many archives and government organizations have large amounts of data, sometimes publicly available, but often confidential in nature, requiring applications for access. while the datasets may be available (typically as csv files), the metadata which accompanies them is not necessarily coherent, making the discovery of these datasets difficult. a prospective user has to read related documents to determine if the data are useful for his/her research purposes. the data provider could enhance discovery of data by providing key metadata in a standardized form. this would allow the creation of standard queries to programmatically identify datasets. the ddi-rdf discovery vocabulary would support this approach. link publications to datasets publications, which describe ongoing research or its output based on research data, are typically held in bibliographical databases or information systems. by adding unique, persistent identifiers established in scholarly publishing to ddi-based metadata for datasets, these datasets become citable in research publications and thereby linkable and discoverable for users. and in addition the extension of research data with links to relevant publications is possible by adding citations and links. such publications can directly describe study results in general or further information about specific details of a study, e.g., publications of methods or design of the study or about theories behind the study. exposing and connecting additional material related to data described in ddi is already covered in ddi. in ddi-rdf, every element can be related to any foaf:document using dcterms:relation. researchers may also want to search for publications where specific questions are discussed. discovering studies using free text search in study descriptions the most natural way of searching for data is to formulate the information need by using free text terms and to match them against the most common metadata, like title, description, abstract, or unit of analysis. a researcher might search for relevant studies that have a particular title or keywords assigned to them in order to further explore the datasets. the definition of an analysis unit might help to directly determine which datasets the researcher wants to download afterwards. a typical query could be ‘find all studies with questions about commuting to work’. searching for studies by publishing agency researchers are often aware of the organizations that disseminate the kind of data they want to use. this scenario shows how a researcher might wish to see the studies disseminated by a particular organization, so that the datasets that comprise them can be further explored and accessed. “show me all the studies for the period 2000 to 2010 disseminated by the esds service of the uk data archive” is an example of a typical query. class  data  collection question -­‐   questiontext    :rdf:langstring -­‐   skos:preflabel    :rdf:langstring questionnaire instrument -­‐   dcterms:description    :rdf:langstring -­‐   skos:preflabel    :rdf:langstring foaf:document representation 0..* externaldocumentation 0..* 0..* question 1..* 0..* responsedomain 1..* figure 5. data collection iassist quarterly 2014/2015 23 iassist quarterly searching for datasets by accessibility this scenario describes how to retrieve datasets that fulfill particular access conditions. many research datasets are not freely available, and access conditions may restrict some users from accessing some datasets. it is common to want to search only for those datasets that are either publicly available, or that have specific types of licensing/access conditions. access conditions vary by country and institution. users may be familiar with the specific licenses that apply in their own context. it is expected that the researcher looking for data might wish to see the datasets that meet specific access conditions or license terms. here, a researcher is using a tool that will generate a sparql query that returns the titles of datasets that are, for example, publicly available under the canadian data liberation initiative community policy. optionally it would also be possible to provide links to the rights statement and the license. there is a paper26 describing further possible use cases in detail. researchers can search for studies by producer, contributor, coverage, universe (i.e., study population), and data source (e.g., study questionnaire). social science researchers can search for datasets using variables, related questions, and classifications. furthermore, one can search for reusable questions using related concepts, variables, universe, and coverage, or by text. rdf from codebook and lifecycle we have implemented a direct and a generic mapping between ddi-xml and ddi-rdf. ddi-codebook and ddi-lifecycle xml documents can be transformed automatically into an rdf representation corresponding to the ontology. the direct mappings are realized through xslt stylesheets27 . bosch and mathiak (2011) have developed a generic approach for designing domain ontologies. xml schemas are converted to ontologies automatically using xslt transformations, which are described in detail by bosch and mathiak (2012). after the transformation process, all the information located in the underlying xml schemas of a specific domain is also stored in the generated ontologies. domain ontologies can be inferred automatically out of the generated ontologies in a subsequent step (bosch 2012). in this section, only the direct approach is described in detail. the structure of ddi-codebook differs substantially from ddilifecycle. ddi-c is designed to describe metadata for archival purposes, and the structure is very predictable and focused on describing variables with the option to add annotations for used question texts, etc. ddi-l on the other hand is designed to capture metadata from the early stages in the research process. a lot of the metadata can be described in modules, and references are used between, for example, questions and variables. ddi-l enables capturing and reuse of metadata through referencing. the disco vocabulary is developed with this in mind -the discovery of studies, questions, and variables should be the same regardless of which version of ddi was used to document the study. ddi-l has more elements and is able to describe studies, variables, and questions in greater detail than ddi-c. however, the core metadata for the discovery purpose is available in both ddi-c and ddi-l. the transformation can be automated and standardized for both. that means that regardless of the input -ddi-c or ddi-l -the resulting rdf is the same. this enables an easy and equal search in rdf resulting from ddi-c and ddi-l. also, interoperability between both is increased. creating triples from ddi xml via xslt there is a huge ecosystem of tools exporting ddi-xml. this makes it possible to act on the output in a standardized way via xslt. xslt is implemented in a wide variety of environments and is a good method for making the transformation from ddi-xml to disco. the flexibility of xslt allows us to generate one conversion process for both ddi-c and ddi-l, which can be detected automatically inside the xslt by paths and nodes of the input files. this corresponds to the goal to generate a consistent and equal disco output independently of the ddi input. the goal of making this implementation is to provide a simple way to start publishing ddi as rdf. xslt is also easy to customize and extend so users can take the base and add output to other vocabularies if they have specialized requirements. it can also be adjusted if special requirements to the input are given. keeping the xslt as general as possible, we provide the basis for a broad reusability of the conversion process. the implementation can also be used as a reference to show how elements in ddi-c and ddi-l map to disco. the current version of the xslt can be found at <https://github.com/linked-statistics/ ddi-rdf-tools>. future work on the mapping and ddi-rdf xslt currently, we have created two separate xslt files for the conversion of ddi-c and ddi-l. according to the flexibility of xslt we aim to merge them into one generic conversion xslt that automatically detects which ddi input is given. also, we plan on including parameters into the conversion process in order to select and define particular languages and uri prefixes. since the work on the conceptual model of disco is currently not finished, the finalized mappings of ddi to disco have to be included into the xslt. future work on integrated use of disco and related rdf vocabularies the description of the relationship of aggregated data to the original microdata by disco, rdf data cube, and prov will be further explored. another focus will be how data portals can benefit of the combined use of disco with dcat, and the new rdf vocabulary on physical data description (phdd)28. conclusions in this paper, we introduced the ddi-rdf model, an approach for applying a non-rdf standard to the web of data. we developed an rdfs/owl ontology for a basic subset of ddi to solve the most frequent and important problems associated with diverse use cases (especially for discovery purposes) and to open the ddi model to the linked open data community. there are two implementations of mappings between ddi-xml and ddi-rdf: a direct mapping and a generic one, which can be applied within various contexts. the most important use cases associated with an ontology of the ddi data model are to find and link to publications related with particular data, to map terms to concepts of external thesauri, and to discover data and metadata that are interlinked with more than one study. diverse benefits are connected with the publication of ddi data and metadata in the form of rdf. users of the ddi social science metadata standard can query multiple, distributed, and merged ddi instances using established semantic web technologies. 24 iassist quarterly 2014/2015 iassist quarterly members of the ddi community can publish ddi data as well as metadata in the linked open data cloud. therefore, ddi instances can be processed by rdf tools without supporting and knowing the ddi-xml schemas’ data structures. after publishing public available structured data, ddi data and metadata can be connected with other data sources of multiple topical domains. acknowledgements the work described in this paper was started at the first workshop on “semantic statistics for social, behavioral, and economic sciences: leveraging the ddi model for the linked data web” 29 at schloss dagstuhl leibniz center for informatics, germany in september 2011. this work was continued at three meetings: a follow-up working meeting in the course of the 3rd annual european ddi users group meeting (eddi11)30 in gothenburg, sweden, in december 2011; a second workshop on “semantic statistics for social, behavioral, and economic sciences: leveraging the ddi model for the linked data web” 31 at schloss dagstuhl leibniz center for informatics, germany in october 2012; and a follow-up working meeting at gesis leibniz institute for the social sciences in mannheim, germany, in february 2013. this work has been supported by contributions of the participants of the events mentioned above: archana bidargaddi (nsd norwegian social science data services), thomas bosch (gesis leibniz institute for the social sciences, germany), sarven capadisli (bern university of applied sciences, switzerland), franck cotton (insee institut national de la statistique et des études économiques, france), richard cyganiak (deri, digital enterprise research institute, ireland), daniel gillman (bls bureau of labor statistics, usa), arofan gregory (odaf open data foundation, usa and ddi alliance technical implementation committee), rob grim (tilburg university, netherlands), marcel hebing (soep german socioeconomic panel study), larry hoyle (university of kansas, usa), yves jaques (fao of the un), jannik jensen (dda danish data archive), benedikt kämpgen (karlsruhe institute of technology, germany), stefan kramer (ciser cornell institute for social and economic research, usa), amber leahey (scholars portal project university of toronto, canada), olof olsson (snd swedish national data service), heiko paulheim (university of mannheim, germany), abdul rahim (metadata technologies inc., usa), john shepherdson (uk data archive), dan smith (algenta technologies inc., usa), humphrey southall (department of geography, uk portsmouth university), wendy thomas (mpc minnesota population center, usa and ddi alliance technical implementation committee), johanna vompras (university bielefeld library, germany), joachim wackerow (gesis leibniz institute for the social sciences, germany and ddi alliance technical implementation committee), benjamin zapilko (gesis leibniz institute for the social sciences, germany), matthäus zloch (gesis leibniz institute for the social sciences, germany). references bosch, t., cyganiak, r., wackerow, j., and zapilko, b. 2012. leveraging the ddi model for linked statistical data in the social, behavioural, and economic sciences. international conference on dublin core and metadata applications, 46–55. bosch, t. 2012. reusing xml schemas’ information as a foundation for designing domain ontologies. proceedings of the 11th international semantic web conference, part ii (berlin, heidelberg, 2012), 437–440. bosch, t. and mathiak, b. 2011. generic multilevel approach designing domain ontologies based on xml schemas. proceedings of the iswc 2011 workshop ontologies come of age in the semantic web (ocas) (bonn, germany, 2011), 1–12. bosch, t. and mathiak, b. 2012. xslt transformation generating owl ontologies automatically based on xml schemas. 6th international conference for internet technology and secured transactions (icitst) (abu dhabi, united arab emirates, 2012), 660 –667. notes 1. thomas bosch gesis leibniz institute for the social sciences, mannheim, germany e-mail: thomas.bosch@gesis.org 2. olof olsson snd swedish national data service, gothenburg, sweden e-mail: olof.olsson@snd.gu.se 3. benjamin zapilko gesis leibniz institute for the social sciences, köln, germany e-mail: benjamin.zapilko@gesis.org 4. arofan gregory open data foundation, tucson, usa e-mail: agregory@opendatafoundation.org 5. joachim wackerow gesis leibniz institute for the social sciences, mannheim, germany e-mail: joachim.wackerow@gesis.org 6. http://sdmx.org/ 7. http://www.ddialliance.org/ddi-at-work/projects 8. http://www.dataarchive.ac.uk/ 9. http://thedata.org/ 10. http://www.icpsr.umich.edu 11. http://www.abs.gov.au/ 12. http://www.ihsn.org/adp 13. http://www.childinfo.org/mics3_surveys.html 14. http://data.worldbank.org/ 15. http://www.theglobalfund.org/ 16. http://www.ddialliance.org/what 17. http://www.loc.gov/standards/premis/ 18. http://www.w3.org/2004/02/skos/ 19. http://www.w3.org/tr/void/ 20. http://www.w3.org/tr/prov-o/ 21. http://www.w3.org/tr/vocab-dcat/ 22. http://www.w3.org/tr/vocab-data-cube/ 23. http://rdf-vocabulary.ddialliance.org/discovery 24. http://www.ddialliance.org/specification/ 25. http://www.w3.org/tr/vocab-data-cube/ 26. http://www.ddialliance.org/resources/publications 27. https://github.com/linked-statistics/ddi-rdf-tools 28. http://rdf-vocabulary.ddialliance.org/phdd.html 29. http://www.dagstuhl.de/11372 30. http://www.iza.org/eddi11 31. http://www.dagstuhl.de/12422 6 iassist quarterly 2015 iassist quarterly abstract as data-sharing becomes more prevalent throughout the natural and social sciences, the research community is working to meet the demands of managing and publishing data in ways that facilitate sharing. despite the availability of repositories and research data management plans, fundamental concerns remain about how to best manage and curate data for long-term usability. the value of shared data is very much linked to its usability, and a big question remains: what tools support the preparation and review of research materials for replication, reproducibility, repurposing, and reuse? this paper describes key curation tasks and new data curation software designed specifically for reviewing and enhancing research data. it is being developed by two research groups, the institution for social and policy studies at yale university and innovations for poverty action, in collaboration with colectica. the software includes curation steps designed to improve the research materials and thus to enable users to derive greater value from the data: checking variable-level and study-level metadata, verifying that code can reproduce published results, and ensuring that pii is removed. the tool is based upon the best practices of data archives and fits into repository and research workflows. it is open-source, extensible, and will help ensure that shared data can be used. keywords data curation, curation software, data sharing, social science, randomized controlled trials . introduction over the past 10 years, many scientific communities have embarked on discussions of data-sharing and reproducibility. from biology (vines, 2014) to epidemiology (peng, 2006) to economics (hammermesh, 2007) to political science (king, 1995), researchers are calling for more data sharing. research funders and journals have been encouraging data sharing and adopting data access policies in greater numbers over the past decade. for example, in the uk, all of the research councils have adopted data-sharing policies (see data curation center’s useful summary3 of all of these policies). wellcome trust in the uk has led a joint statement4 of purpose on data-sharing principles, which includes over 15 funders. in the us, the office of science and technology policy memorandum of 2013 5 stipulated that us funders receiving $100m or more in federal research funds adopt data-sharing polices, and the government is working to facilitate code sharing6. major foundations such as the bill and melinda gates foundation7 and the laura and john arnold foundation8 have also adopted data-sharing polices. a number of journals are instituting policies in which they require researchers to share the data and code underlying the published research results (see this list of social science journals with a data sharing policy9 and this journal data policy review 10). there is much variety across policies. funder policies differ in their timeframes, whether data should be made openly available or simply available on request, which materials should be shared, and in many other ways (for an overview, see wykstra, 2013). likewise, journals vary in whether data should be available openly. some journals, for example the american economic review11, require researchers to post the data on the journal website, whereas other journals merely ask researchers to note in the article where they shared the data or that they make it available upon request. while the language and particulars may vary, a constant theme running through these discussions is the desire for scientists to be able to examine each other’s work. can others dig into the analysis and data; can others understand the study in enough detail to try to repeat it? in this paper, we focus on an issue which is crucial for examining others’ work: that of the usability of shared data. by “data” here, we mean not just the datasets new curation software: step-by-step preparation of social science data and code for publication and preservation by limor peer1 and stephanie wykstra2 research will be more credible if others can have full access to all aspects of scholarly work iassist quarterly 2015 7 iassist quarterly themselves but the related materials as well: the analysis code, the metadata, documentation, and instruments. we refer to preparing these materials for public use as data curation. after a description of this project and a discussion of the value of data sharing, we discuss the relation of reproducibility, re-use, and data curation, describe key curation tasks, and present new curation software, developed with colectica12, aimed at helping with review and enhancement of research materials. background: data from randomized controlled trials (rcts) in the social sciences the impetus for the collaboration around data curation between the institution for social and policy studies (isps)13 at yale university and innovations for poverty (ipa)14 is a focus on a particular way of doing social science research: field experiments. both organizations collect data from social science research that measures the impact of interventions – such as voter mobilization campaigns and microfinance programs – via randomized controlled trials in the real world. isps has been involved with close to 100 such studies, mostly in political science, and ipa in about 300 studies, working with researchers in development economics, among other fields. studies linked with isps and ipa have been published in such journals as the american political science review, political analysis, american political research, public opinion quarterly, american economic review, american behavioral scientist, and the quarterly journal of economics. data from these studies are mostly quantitative, often gathered from a combination of administrative records, surveys, and observation, and of potentially high value for researchers, educators, policy makers and students. data are often generated to address a particular research question and linked to a publication that describes the results of a particular experiment. datasets underlying published articles or books span time periods and continents, and vary in scale in terms of the number of observations and variables. since rcts are relatively new to the social sciences, metadata standards are still emerging. the data documentation initiative (ddi15), the primary social science metadata standard, now has a working group 16 charged with updating the standard to capture the unique characteristics of this research method. high quality descriptive metadata is essential to facilitating the interpretation of social science studies. isps has supported a data archive17 since 2010 (peer and green, 2012). the archive includes research output by isps-affiliated researchers, with emphasis on experimental design and methods. research output includes data and code and is typically deposited at the end of the project, coinciding with manuscript publication. research output is organized as a complex object around a study, with multiple files of various sorts related to each study, including data, code, output, and other files. study-level metadata are compiled from information provided by depositors (e.g., researchers) via a deposit agreement form, and from associated materials (e.g., published article). for variable-level metadata, isps uses stat/transfer to produce make available xml files based on ddi version 3.1 for datasets. ipa has also launched a new repository to share data from rcts (both from ipa studies as well as rcts from other groups). the repository is hosted by harvard’s dataverse18. ipa shares isps’ approach to curation but differs in that it also requests that researchers share the full collected datasets, and works to help them prepare the larger datasets, as opposed to only the data underlying the published research results. in addition, ipa is working with research staff on the ground within its country offices, to improve code and data management processes early on in the study workflow and improve later data usability. the value of data-sharing there are two primary sources of value from sharing data: reuse and transparency. first, sharing data permits others to use the data for further purposes. it is currently more common to re-use data from large-scale survey-based studies such as demographic and health surveys19, than to re-use data from experimental studies. however, as data sharing becomes more prevalent there is promise that scientists will conduct additional analyses, such as secondary analysis and meta-analysis and formulate new questions. this is the logic expressed in a 2013 ostp memo20 to all government agencies, which states with respect to government-funded studies that, “the results of that research become the grist for new insights and are assets for progress in areas such as health, energy, the environment, agriculture, and national security.” there is some evidence that, at least in one field, studies that made data available received more citations than similar studies for which data were not made available, as measured by number of citations (piwowar, 2013). the idea driving research transparency is that research will be more credible if others can have full access to all aspects of scholarly work that led to publication. as king (1995) put it, “the only way to understand and evaluate an empirical analysis fully is to know the exact process by which the data were generated and the analysis produced” (p.444). an essential component is the ability to reproduce computations and analyses by using the shared code and data. re-analysis of this kind is often assigned in methods courses in the social sciences, in which it is also often recommended that “replicators” go beyond simple re-analysis to conduct robustness checks and delve into the analytical decisions made in the published research (king, 1995). access to code in addition to the data is increasingly recognized as critical in all computational sciences (i.e., those which rely heavily on analysis of quantitative data) and as contributing to the credibility of the research (stodden et al., 2013). data use and data curation in order to glean full value from shared data, for re-analysis or any future re-use, the data must be usable in the long-term. the usability of data simply means that it can be “independently understandable” by future scientists (peer, green and stephenson, 2014; peer, 2014a; peer and green, 2015). preparing files for long-term use starts with good documentation. it is strongly recommended that a standards-based, structured, open and machine-readable metadata scheme, such as ddi for social sciences data, is used (e.g., starr et al., 2015; u.s. government, 2012 ; w3c, 2015). preparing data files includes, but is not limited to, ensuring that variables are clearly named and labeled. variables created via original data collection should be linked to the source, e.g., survey questions. numeric data with value codes should be labelled clearly. code should be commented to indicate which operations the code carries out (e.g., variable-construction and cleaning, producing tables). in the context of a study, additional documentation may be required. it is recommended that researchers provide readme files documenting the files which are shared, with instructions about running the files and any other information about them (see this useful guide21). sufficient study-level metadata is critical to understanding the study and its context. information such as: time period, geographic area, 8 iassist quarterly 2015 iassist quarterly sampling frame and selection method, sample size, study methodology, and data collection method should be provided. if there is a publication, the data and the published research results should be clearly linked. and, of course, open and persistent access to files is a precondition for long-term usability. most basically, the files should be in a sustainable location, preferably a data repository which offers long-term preservation. files should also be available in non-proprietary formats to increase accessibility. if they are in proprietary formats, there may also be a greater chance they will become unusable over time due to software updates. we refer to the process of reviewing and enhancing research outputs for the purpose of long-term usability as “data curation.” according to the digital curation center22 curation involves “maintaining, preserving and adding value to digital research data throughout its lifecycle.” our goal in undertaking data curation is to ensure that users may have persistent access and be able to correctly interpret and re-use these materials without the need to contact original researchers. we see particular value in curation when the intended re-use of the research materials is to fully evaluate an empirical analysis, that is, to reproduce research results (peer, 2011)23. specifically, we use the “data quality review” framework to focus on specific curation tasks (peer, green and stephenson, 2014). key curation tasks for data from rcts in the social sciences to ensure research transparency, long-term usability, and ongoing persistent access to research outputs generated by a specialized community, certain data curation tasks need to take place. data archives such as the inter-university consortium for political and social research (icpsr 24) and uk data archive (ukda25) have established practices that are tried and tested to ensure “that data are accurate, complete, well documented, and that they are delivered in a way that maximizes their use and reuse” (peer, green and, stephenson 2014, p.16). the isps curation workflow is based on the icpsr pipeline (peer, 2014b), and has been adapted for research output from rcts in the social sciences (peer and green, 2012). the existing isps data archive workflow26 has gone a long way toward satisfying the curation needs of isps, but the partnership with ipa presents an opportunity to build a modular, open-source curation tool that could be adapted by our organizations to changing needs, research methods, dissemination platforms, and preservation solutions. on the basis of data archives’ best practices and the isps curation workflow, we have identified eight key curation tasks that are designed to improve the research materials and thus to enable users to derive greater value from the data by, for example, checking variable-level and study-level metadata and fiqure 1 fiqure 2 iassist quarterly 2015 9 iassist quarterly ensuring that personally-identified information is removed. the curation tasks also include the review of code files -statistical and other programming scripts -for the purpose of checking whether scientific results can be reproduced with the code and data provided. the eight tasks are as follows (see appendix): 1 check for missing labels. 2 review observation count. 3 identify potential data errors. 4 compare questionnaire, codebook, and data. 5 ensure there is no personally-identifiable information (pii) in data file. 6 confirm code executes. 7 confirm code replicates reported results. 8 create preservation and open formats digital curators and research teams who strive to meet the demands of these tasks also need to track and confirm the completion of the tasks as well as to capture all the useful metadata generated throughout the processes. the prime objectives for this project are: to automate as many of the curation tasks as possible, to technically integrate these curation tasks, and to do so using a structured but flexible workflow. these eight curation tasks constitute core requirements for the software we describe here27. new curation software working with colectica, a software development group specializing in data and metadata tools for social sciences, we have developed new software that structures and tracks the curation workflow, helps automate parts of the data pipeline, captures all metadata throughout the process, and pushes out relevant information to pre-determined destinations (i.e., a user, the archive administrators, a web based dissemination system, or preservation systems). the tool was developed to fit into repository and research workflows. the software is primarily open-source, extensible, and can be easily integrated with other systems. it is written in c# and runs on the asp.net mvc framework. it leverages ddi lifecycle (also known as ddi 3.2) and combines several off-the-shelf components with a new, open source web application that integrates the existing components to create a flexible data pipeline. default components include stattransfer28 , colectica repository, and bagit file packaging format, but the software is developed so each of these can be swapped for alternatives. key curation software characteristics: 1 the software supports specific curation tasks for different file types and provides automatically-generated information helpful in performing each task. the file categories relevant to curation are data files, code file, and other. 2 the software facilitates the production of descriptive metadata at the study, file, and variable levels, and maps and stores all metadata in ddi lifecycle format. study-level metadata is used to inform the catalog record. the software identifies file types and, for data files with recognized formats, such as .dta, r, and .sps, produces further metadata for each variable. 3 the software allows editing metadata in the web-based interface and automatically updates fileand variable-level metadata, and the catalog record. the software can produce a new version of a data file with changed metadata (data not changed). all changes to files are tracked using a git platform. 4 the software also allows downloading files for further review or editing for specific curation tasks. the revision management check-out/check-in system recognizes files that have been updated offline and then checked as a new version, and all relevant metadata is updated, with versions of metadata corresponding to new versions of files. all versions are stored. the system records which curation task was completed. for data files, the software also provides the option to create and store revised summary statistics. 5 the software allows viewing other files in a window to compare documentation to metadata in the system. 6 notes can be entered at every step, at every level of metadata. 7 each curation step can be approved or rejected. all activity around these curation tasks is tracked, and each task must be performed before a catalog record is published. this allows curation progress to be seen at a glance, and ensures a reliable record of the curation process. discussion we conclude with a few thoughts about this new curation software and its benefits for the scientific research community. the main advantage of this tool is that it helps codify and automate a series of curation tasks that prepare data and code for re-use. many of these tasks are not new: they have been described in numerous best-practice documents and guides. yet, as far as we know, there are no tools that facilitate curation of research data prior to their ingest into a repository.29 the advantages of this tool include its functionality to create a consistent workflow for key curation tasks that can integrate several software environments, automatic metadata production, presentation of missing variable and value labels, versioning of files to control what has been done, and the ability to add notes to document changes and enhancements to files and to track the entire process. we think this unified curation workflow can help bring about many of the recommended fiqure 3 10 iassist quarterly 2015 iassist quarterly curation practices that the research data management community has been advocating for, at scale. in terms of the application of the software, it is our recommendation that this software be used as close as possible to the research process. in-house trained curation staff at research labs or centers are best positioned to understand the research, the data, and the analyses and can create additional documentation if possible and necessary. in-house staff also typically benefit from access to, and communication with, researchers in case questions come up. researchers may be asked to provide more information if documentation is incomplete or clarifications are needed. information and curation specialists, such as data and subject liaison librarians, in conjunction with statistical experts in the researcher’s institution or professional society are also in a good position to undertake this type of curation. if no curation was done on research outputs intended for re-use or preservation, we urge repositories, journals, and funders, to facilitate or make use of this software before research outputs are preserved or disseminated. this curation tool is flexible enough to allow modification. we have described major steps that we have identified within our own research groups, and we have customized the software to aid with these steps. however, the steps may be modified according to the needs of particular research groups, repositories, or researchers using the software. for example, differences between isps and ipa in research management and infrastructure led to their using the tool differently. generally, we foresee circumstances in which some curation tasks may not be relevant (e.g., stand-alone data may not require regenerating the results by running the code to produce tables) and conversely, instances in which additional curation tasks may be added to fulfill specific repository, lab, or discipline requirements and standards. there are multiple mechanisms for sharing data these days, and most involve some level of curation. this tool can be used by researchers or labs who wish to self-deposit into a general data repository, by established data archives that are looking to automate and integrate disparate curation processes and systems, by journals or funders who wish to review research outputs before they disseminate them along with publications, and by institutional repositories and other archives who plan to preserve these research outputs and ensure they can be persistently accessed and usable. if used by data archives or by general data repositories such as dryad, figshare or dataverse, this tool may be used to support a service model of managing data with a view toward usability and replication. we have argued that data curation is an essential part of the movement towards open data. data curation is a key component of usability, without which long-term re-use is not possible. therefore, we hope that the open source software will be taken up by a number of groups in addition to our own. acknowledgments we thank our partners, jeremy iverson and dan smith at colectica, for collaborating with us on this project. we would also like to acknowledge ann green and niall keleher for their significant contribution to the conception and early development of this entire endeavor. references asendorpf, jens b., conner, mark, de fruyt, filip, de houwer, jan, denissen, jaap j. a., fiedler, klaus, fiedler, susann, funder, david c., kliegl, reinhold, nosek, brian a., perugini, marco, roberts, brent w., schmitt, manfred, van aken, marcel a. g., weber, hannelore and wicherts, jelte m. (2013) recommendations for increasing replicability in psychology. european journal of personality, 27: 108–119. (available at http://onlinelibrary.wiley.com/doi/10.1002/ per.1919/full#per1919 or doi: 10.1002/per.1919 borgman, christine l. (2010) research data: who will share what, with whom, when, and why? ratswd working paper no. 161. (available at ssrn: http://ssrn.com/abstract=1714427 or http://dx.doi. org/10.2139/ssrn.1714427) collberg, christian, proebsting, todd, moraila, gina, shankaran, akash, shi, zuoming, and warren, alex m. (2014) “measuring reproducibility in computer systems research,” department of computer science, university of arizona, technical report. (available at: http:// reproducibility.cs.arizona.edu/tr.pdf ) hammermesh, daniel, (2007) viewpoint: replication in economics. canadian journal of economics, 40, no. 3: 715-733. holdren, john (2013) increasing access to the results of federally funded research, office of science and technology policy. (available at http://www.whitehouse.gov/sites/default/files/microsites/ostp/ ostp_public_access_memo_2013.pdf ) king, gary (1995) replication, replication. ps: political science & politics, 28: 444-452. peer, limor (2011) building an open data repository: lessons and challenges (available at ssrn: http://ssrn.com/abstract=1931048 or http://dx.doi.org/10.2139/ssrn.1931048) peer, limor (2014a) why ’intelligent openness’ is especially important when content is disaggregated. isps lux et data blog (available at http://isps.yale.edu/news/blog/2014/12/why-intelligent-opennessis-especially-important-when-content-is-disaggregated) peer, limor (2014b) mind the gap in data reuse: sharing data is necessary but not sufficient for future reuse. lse impact blog. (available at http://blogs.lse.ac.uk/ impactofsocialsciences/2014/03/28/mind-the-gap-in-data-reuse/ ) peer, limor and green, ann (2012) building an open data repository for a specialized research community: process, challenges and lessons. international journal of data curation, 7(1): 151-162. (available at http://www.ijdc.net/index.php/ijdc/article/view/212) peer, limor and green, ann (2015) research data review is gaining ground. isps lux et data blog (available at http://isps.yale.edu/ news/blog/2015/03/research-data-review-is-gaining-ground) peer, limor, green, ann and stephenson, elizabeth (2014) committing to data quality review. international journal of data curation, 9(1): 263-291. (available at http://www.ijdc.net/index.php/ijdc/article/ view/9.1.263/358) peng, roger, francesa dominici, scott segar (2006) reproducible epidemiologic research. american journal of epidemiology, 163(9): 783-789. piwowar, heather, rs day, db fridsma (2007) sharing detailed research data is associated with increased citation rate. plos one 2(3). starr j, et al. (2015) achieving human and machine accessibility of cited data in scholarly publications. peerj computer science 1:e1 (available at https://dx.doi.org/10.7717/peerj-cs.1 ) stodden, victoria et al. (2013) setting the default to reproducible: reproducibility in computational and experimental mathematics. icerm workshop paper. (available at http://stodden.net/icerm_ report.pdf ). united states government (2012) a primer on machine readability for online documents and data data.gov iassist quarterly 2015 11 iassist quarterly blog (available at https://www.data.gov/developers/blog/ primer-machine-readability-online-documents-and-data) vines, tim et al. (2014) the availability of research data declines rapidly with article age. current biology, 24(1): 94-97. w3c (2015) data on the web best practices. working draft. (available at http://www.w3.org/tr/2015/wd-dwbp-20150625/ ) wykstra, stephanie (2013), data access policies landscape, figshare, (available at http://dx.doi.org/10.6084/m9.figshare.827268 appendix curation tasks detail we describe here the eight curation tasks essential to our work, with an explanation of how each is handled by the software. note that users of the software may choose a combination of one or more of any of the tasks, and may also extend them with apis to other software. these may include the generation of descriptive statistics, a data dictionary and a set of variable frequency distributions. (see figures 2, 3.) 1. check for missing labels. a. rationale: all variables should be properly identified with a name and a description that provides additional information. nominal (or categorical) variables should have numeric or string labels for each category (see more on icpsr guide 30). b. description: for data files. the software displays variables and value labels and alerts of any missing labels. upon ingest, the software analyses data files and extracts variable-level metadata for each column. this metadata is stored in ddi 3.2 format. for any variables without labels, and for categorical data without value labels, the software prompts the curator to enter labels. c. technical: this variable-level metadata extraction is built on the data import functionality found in colectica designer and colectica repository. the software currently supports reading variable-level information for stata, rdata, and csv files. other formats can be supported using stat/transfer to convert the file to a supported format, or by extending the data ingest capabilities of the curation software. 2. review observation count. a. rationale: the goal is to view the number of cases (observations) for every variable and for the dataset as a whole. this is helpful in providing the basis for subsequent curation tasks. b. description: for data files. the software displays # of observations, summary statistics including frequencies. when ingesting data files, the curation software determines the number of observations, calculates summary statistics, and stores this information about the data file with the metadata. c. technical: this task is completed using the curation web application, but may require viewing and reconciling documents in other readers or editors. 3. identify potential data errors. a. rationale: unlikely or impossible values for interval variables and undefined or incorrect values for nominal (categorical) variables make it difficult for future users to interpret the data. also out of range and missing values. this is intended to check the overall integrity of the data. the ukda guide31 provides examples of such anomalies. b. description: for data files. the curation web application provides a variable-level metadata browser. this shows details of each variable in a data file, including summary statistics and value labels for categorical variables. curators are responsible for reviewing each variable to ensure there are no obvious errors in the data. c. technical: this task can be completed using the curation web application, but may also require viewing or editing the data in a statistical software package. 4. compare questionnaire, codebook, and data. a. rationale: this task can be carried out at the same time as other tasks related to data files. performing tasks 1-3 may be informed by other documentation that was deposited along with the data files. comparing summary and descriptive statistics along with data dictionary or codebook helps ensure that question text, labels, response categories and value labels are consistent. b. description: for data files. the software allows viewing other files in a window to compare documentation to fileand variable-level metadata in the system, or importing the information from a questionnaire if it can be read by colectica designer. for each variable, curators are shown summary statistics and label information, and can review other documents, including links to publications based on the data. curators are responsible for flagging instances where the observation count reported in publication does not match the observation count in the data file, where variables are missing or transformed, or where label information is missing or incomplete. curators are responsible for ensuring all information is consistent. c. technical: this task is completed using the curation web application, but may require viewing and reconciling documents in other readers or editors. 5. ensure there is no personally-identifiable information (pii) in data file. a. rationale: human subject data are prevalent in the social sciences. any entity that shares human subject research data has responsibility to protect respondent confidentiality and to minimize the risk of identifying individuals. guidance on confidentiality32 is provided by icpsr33 and the australian national data service (ands)34. b. description: for data files. using the curation web application’s variable-level metadata browser, curators are responsible for reviewing each variable to ensure it does not contain personally identifiable information (pii). if a curator finds variables containing names, social security numbers, or other identifiable information, they are responsible for removing the columns and submitting a new version of the data file. c. technical: this task can be completed using the curation web application, but may also require viewing or editing the data in a statistical software package, as well as also viewing other documentation. in addition to visual inspection of the data and documentation, code may be used to identify variable names that indicate pii (social security numbers, phone numbers, addresses, etc). future development of the software could enable display of a list of predefined types of variables or data and api integration with anonymization software (e.g., anonimatron35, qualanon36 ). follow up may 12 iassist quarterly 2015 iassist quarterly require curator to create new program file that produces a new data file. 6. confirm code executes. a. rationale: by code we are referring to computational workflows used in the research process, including data collection, cleaning, and analysis. in some disciplines, researchers are just warming up to the idea of sharing their data and are not used to providing code. but even in disciplines such as computer science, where code is commonly shared, problems with code builds exist (collberg et al., 2014). the goal here is to test the code with the given data to identify any potential errors in the script itself. b. description: for code files. curators are responsible for ensuring that all source code submitted executes without errors. the curation web application provides links to download the source code and any dependencies, such as data files. a web-based preview with syntax highlighting is also available. c. technical: to complete this step, curators must use the appropriate statistical software (e.g. stata, r, spss, or sas). the curation software tracks the versions of the code files for changes made by the curator. 7. confirm code replicates reported results. a. rationale: after confirming that the script is error free, “an assessment is made about the purpose of the code (e.g., recoding variables, manipulating or testing data, testing hypotheses, analysis), and about whether that goal is accomplished.” (peer et al., 2014). the main goal is to check whether the code, in conjunction with the data provided, produces the results reported. the idea is that, “researcher b… obtains exactly the same results (e.g. statistics and parameter estimates) that were originally reported by researcher a (e.g. the author of that paper) from a’s data when following the same methodology” (asendorpf et al., 2013). confirmation that results can be replicated, and any additional annotation created in the process, helps inform future users exactly how results were generated. note that the focus here is on the regeneration of results and not on the correctness of the methodology, analysis, or interpretation of the results. b. description: for code files. curators are responsible for ensuring that statistical programs produce the results reported in any related publications. the curation web application provides links to download the source code and any dependencies, such as data files. the curator analyzes the output of these programs, reviews any numbers, tables, and charts included in publications, and ensures they match the output. c. technical: to complete this step, curators must use the appropriate statistical software (e.g. stata, r, spss, or sas). the curation software tracks the versions of the code files for changes made by the curator. 8. create preservation and open formats. a. rationale: the goal is to create file formats that easily lend themselves to reuse via technology. files trapped in licensed formats (e.g., xls, .dta) will not be available for use by non-licensed software mechanisms, and may be less usable over time as software is outdated. the ukda guide37 explains preservation formats. b. description: for all files. the curation software automatically converts supported data files to csv format for preservation. this occurs during publication after curation has been performed, reviewed, and approved. c. technical: for file types in proprietary formats that are not specifically supported by the curation software, curators are responsible for creating a file in the appropriate preservation format, and uploading the file to the catalog record. this can be accomplished using conversion software such as stat/transfer, or saving documents to text or pdf formats. for code file, curators may write r scripts that replicate any statistical code provided by the researcher in a licensed statistical program. notes 1 limor peer is associate director for research at yale university’s institution for social and policy studies, http://isps.yale.edu/. contact: limor.peer@yale.edu. 2 stephanie wykstra is research manager of research transparency at innovations for poverty action, http://www.poverty-action.org/ 3 http://www.dcc.ac.uk/resources/policy-and-legal/ overview-funders-data-policies 4 http://www.wellcome.ac.uk/about-us/policy/spotlight-issues/datasharing/public-health-and-epidemiology/wtdv030690.htm 5 https://www2.icsu-wds.org/files/ostp-public-access-memo-2013.pdf 6 https://government.github.com/ 7 http://www.gatesfoundation.org/how-we-work/ general-information/open-access-policy 8 http://www.arnoldfoundation.org/sites/default/files/pdf/ guidelines%20for%20research%20funded%20by%20ljaf%20 11-12-2013%20ma%20-%20july%2016%202015.pdf 9 https://jordproject.wordpress.com/project-data/ social-science-journals-that-have-a-research-data-policy/ 10 https://docs.google.com/spreadsheets/d/1iwe-hnmjugv9jrg22f9t og0bfh3fdnuuad65jvohuc0/edit?pli=1#gid=160911802 11 https://www.aeaweb.org/aer/data.php 12 http://www.colectica.com/ 13 http://isps.yale.edu/ 14 http://www.poverty-action.org/ 15 http://www.ddialliance.org/ 16 http://www.ddialliance.org/alliance/working-groups#governance 17 http://isps.yale.edu/research/data 18 http://thedata.harvard.edu/dvn/dv/socialsciencercts 19 http://dhsprogram.com/ 20 https://www2.icsu-wds.org/files/ostp-public-access-memo-2013. pdf 21 http://data.research.cornell.edu/content/readme 22 http://www.dcc.ac.uk/digital-curation/what-digital-curation 23 to be clear: performing data curation doesn’t ensure that studies may be deeply examined and re-analyzed. if only a subset of the data and code is shared, for example, there may easily be limits to how thoroughly others can examine the original researcher’s method of arriving at results. likewise, re-use may be very limited if only a small subset of originally collected data is shared, as is often the case when researchers share their data in accordance with journal requirements. however, we argue that data curation is a very helpful, if not sufficient, condition for reproducibility and re-use. 24 https://www.icpsr.umich.edu/icpsrweb/content/datamanagement/ lifecycle/ingest/enhance.html 25 http://www.data-archive.ac.uk/media/54770/ukda081-ds-quantitati vedataprocessingprocedures.pdf 26 http://or2013.net/content/repository-data-re-user-hand-curatingreplication/index.html iassist quarterly 2015 13 iassist quarterly 27 other isps and ipa requirements included a workflow management dashboard, version tracking, integrating metadata production with data and code review and cleaning, creating preservation metadata, secure upload, storage and access, persistent identifier assignment, easy transition to public dissemination of content, and preference for open source solutions. 28 separate licenses may be needed for some software. 29 software exists that helps with curation of other digital objects, e.g., bitcurator (http://www.bitcurator.net/bitcurator-access/), ladybird (http://ladybird.library.yale.edu/). 30 http://www.icpsr.umich.edu/icpsrweb/content/deposit/guide/ chapter3quant.html#labels 31 http://www.data-archive.ac.uk/media/54770/ukda081-ds-quantitati vedataprocessingprocedures.pdf 32 https://www.icpsr.umich.edu/icpsrweb/content/datamanagement/ confidentiality/ 33 http://www.icpsr.umich.edu/icpsrweb/content/deposit/guide/ chapter5.html 34 http://ands.org.au/guides/sensitivedata.html 35 http://sourceforge.net/projects/anonimatron/ 36 https://www.icpsr.umich.edu//icpsrweb/dsdr/tools/anonymize.jsp 37 http://www.data-archive.ac.uk/create-manage/format/formats 1/11 magnuson, diana l; thomas, wendy l. (2023) expanding our perspective: building a sustainable metadata culture, iassist quarterly 47(2), pp. 1-11. doi: https://doi.org/10.29173/iq1046 the creative commons-attribution-noncommercial license 4.0 international applies to all works published by iassist quarterly. authors will retain copyright of the work and full publishing rights. expanding our perspective: building a sustainable metadata culture diana l. magnuson1 and wendy l. thomas2 abstract the institute for social research and data innovation (isrdi) at the university of minnesota submitted an application for approval to the core trust seal (cts) in june 2022. in the course of the protracted process of preparing isrdi materials for the application, we learned five lessons that expanded our perspective on the role of the archive within our organization and committed the institute to building a sustainable metadata culture. after reviewing the specialized nature of isrdi as it developed over time, clarifying and documenting the processes that developed as the institute matured and expanded, and applying the standards and guidelines supported by the cts, isdri staff are now better positioned to identify areas of future process development and to address outstanding needs for documenting and preserving the institute’s work. these lessons are applicable to research organizations responsible for preserving a record of their work in the midand long-term. keywords core trust seal, archive, metadata, business process model, preservation introduction in june 2022, the institute for social research and data innovation (isrdi) submitted an application to the core trust seal (cts) for its ipums projects.3 application to the cts for its professional approval culminates years of effort within the institute for ensuring access to our signature collection of harmonized census and survey data from around the world. in the course of this protracted work, isrdi learned five valuable lessons for data archivists organizationally positioned in a larger institutional context: • institutional history situating our institutional history in a larger social science context sharpened our understanding of our unique contribution to social science infrastructure and to data archiving. • building the cts application building our cts application clarified our institutional strengths and illuminated areas to refine. • business process model developing a business process model documented roles and responsibilities of organizational components (project, administration, and archive) and highlighted metadata production and curation points. • leveraging documentation leveraging documentation produced for the cts application will support future funding applications, enumerate data archive responsibilities, identify cross-project technical systems, educate current staff, and facilitate onboarding new employees. https://doi.org/10.29173/iq1046 https://creativecommons.org/licenses/by-nc/4.0/ 2/11 magnuson, diana l; thomas, wendy l. (2023) expanding our perspective: building a sustainable metadata culture, iassist quarterly 47(2), pp. 1-11. doi: https://doi.org/10.29173/iq1046 • preservation preserving our data products and unique intellectual property relating to the processing and methodology that contributed to the development of our data products is an essential contribution to social science infrastructure. these lessons are shaping the way we internally conduct our data archival work, externally relate to our funders and data collaborators, and prepare for future data harmonization projects. we believe our experience can be a guide for other organizations aiming to build a sustainable metadata culture. this paper presents the value of the cts review and submission process in helping a non-traditional archive define its place within a research organization and clarifies the archive’s role in supporting the standing of its parent organization with funders, data providers, and the research community. institutional history while producing an application to submit to the cts we came to appreciate the value of reflecting on our institutional history and situating that history in the larger context of data archiving and social science infrastructure. this intellectual exercise sharpened our understanding of our unique story and contribution to social science. this is particularly important within non-traditional archives where the focus of the organization may be research or providing a specialized data product. the review process forces the archive to clearly explain its role and archival activities within the larger organization. over the last thirty years, ipums has created the world’s largest accessible database of census microdata. the institute for social research and data innovation and its flagship data project, ipums, has its roots in the 1880 historical census project, a eunice kennedy shriver national institute of child health and human development (nichd) funded project to create a 1-in-100 public use microdata sample of the 1880 u.s. census of population. housed in the history department at the university of minnesota, historical demographers and co-principal investigators steven ruggles and russell menard conceived of extending back the series of public-use microdata samples already in existence (1900, 1910, 1940, 1950, 1960, 1970) (magnuson and ruggles, 2022). once the completed 1880 pums was disseminated, researcher feedback was overwhelmingly enthusiastic.4 funding to complete the decennial population series--1850-1870 and 1920-1930 and updates to 1900 and 1910--would come to the university of minnesota between 1992 and 2002 (magnuson and ruggles, 2022).5 by 1991 ten machine readable public use microdata samples covering the decennial censuses of population from 1880 and 1990 were publicly available or under development (for 1850, 1880, 1900, 1910, 1940, 1950, 1960, 1970, 1980, and 1990). the nagging problem facing the research community was the difficulty of using the data as a time series because the various datasets were created at different times, by different investigators, employing different formats, record layouts, coding schemes, and producing different documentation. between 1985 and 1991, steven ruggles “developed a set of fortran programs that recoded selected variables into a common format across the available census samples, created subsets of the samples that were of manageable size, and pooled multiple censuses into a single file” (magnuson and ruggles, 2022). initially ruggles used a “lowest common denominator” approach for variable codes when combining samples, which naturally resulted in significant loss of information. despite these limitations, the program could be customized to meet the requirements of any research question. demand for customized data sets steadily increased in-house at the university of minnesota, as well as coming from a few researchers at other universities. clearly there was a user base for time series microdata if the compatibility issues could be resolved. in 1991, steven ruggles was awarded a national science foundation grant to create a single integrated series that would “maximize comparability and minimize information loss” (ruggles 1992-95; 1991a, 1991b). he proposed to name the finished product the “integrated public use microdata series,” and https://doi.org/10.29173/iq1046 3/11 magnuson, diana l; thomas, wendy l. (2023) expanding our perspective: building a sustainable metadata culture, iassist quarterly 47(2), pp. 1-11. doi: https://doi.org/10.29173/iq1046 thus ipums was born.6 key technical innovations emerging from ipums included the first structured metadata system for data integration and the first interactive web-based system for user-customized data extraction. in 1993, ipums data were disseminated through an anonymous file transfer protocol (ftp) site and two years later the ipums website launched its own data extract system. hypertext variable-level documentation became available in 1997 (magnuson and ruggles, 2022). in 1999, the university of minnesota graduate school issued a call for competitive applications to receive funding for interdisciplinary centers. collaborators from units representing geography, history, public affairs, industrial relations, and health services successfully made their case to the university for establishing an interdisciplinary population center. two smaller population centers merged to become one, and the minnesota population center (mpc) thus emerged from a “strategic positioning process” that sought to prioritize and foster highly collaborative and interdisciplinary activities at the university (magnuson 2015, lawrenz and paller 2006). beginning in 2000, the mpc was a university-wide interdisciplinary cooperative for demographic research at the university of minnesota. the center had three main goals: “to foster connections among population researchers across disciplines, to develop large-scale collaborative research projects, and to provide infrastructure for demographic research” (ruggles 2011). in 2016, the mpc was reorganized in recognition of the diverse development of population research infrastructure at the university of minnesota. the isrdi became the parent organization of four centers: the mpc, ipums, the life course center, and the minnesota research data center.7 ipums separated from the mpc to become a co-equal center within the newly constituted institute. over the course of our thirty-year history, the institutional entities that produce and disseminate ground-breaking ipums data products and technological innovations have formed an integral part of social science infrastructure as we know it today. at present, the ipums suite of products contain nine harmonized data collections.8 data comes from the united nations statistical division, the united states census bureau international division, and over 100 national and regional statistical organizations. building the cts application the time-consuming undertaking of building ipums policy documentation to complete the cts application clarified our institutional strengths and identified areas to refine. the cts process is intended to be on-going, including continuing to support and implement policies and processes, refine activities as needed, and to respond to changes in the environment over time. to do this the archive needs to obtain the initial buy-in of the parent organization as well as maintain on-going support for archival work. recognizing our institutional strengths assure the parent organization that the goal of the archive is to improve and support the organization through its work. 2016 was a watershed year for the mpc as it reorganized into the isrdi. as a co-equal entity within isrdi, ipums clarified its mission in terms of data harmonization, access, curation, and preservation. this included expanding the use of metadata standards such as dublin core, ddi, and iso-19115 in describing our data for archival purposes. at the same time, external funding organizations were increasing requirements for adherence to standard archival practice using the open archival information system (oais) model and digital object identifiers (doi).9 to address these external concerns, we began an internal assessment of our data products, metadata, and archival practices with respect to those standards. for microdata projects, variable definitions, source data for harmonization, data collection forms, collection instructions, and sampling information were relevant. for aggregate data, the table and dimension descriptions, data source, universe, geographic https://doi.org/10.29173/iq1046 4/11 magnuson, diana l; thomas, wendy l. (2023) expanding our perspective: building a sustainable metadata culture, iassist quarterly 47(2), pp. 1-11. doi: https://doi.org/10.29173/iq1046 definitions, and imputation information were important to capture and preserve.10 our internal assessment revealed that we captured an extensive amount of metadata, but we did not capture changes to the metadata over time in a structured way. developing clear guidelines regarding why and how we would be assigning doi’s, requirements for a versioning policy for each project, and capturing preservation copies for the archive that met oais standards, were the initial points of discussion. the needs of each project were reviewed and commonalities were documented. communicating these issues and concerns across projects and administrative units was a challenging but important part of the assessment process. ultimately, this process began to nurture a sustainable metadata culture within our organization. after a roughly three-year internal assessment, the decision to adopt the practice of using digital object identifiers (doi) was made in 2016. datacite (datacite.org) was selected as the university of minnesota was already a member. registering dois with datacite required decision points around the following tasks: determining at what level to assign a doi; capturing data and metadata for specific versions of our data products; providing persistent access to each identified version of our data products; and maintaining and providing access to those versions over time. the discussion of these issues was done in an iterative fashion and involved input from all of the ipums project groups. our goal was to establish clear versioning rules around our data products while allowing each project the flexibility to decide when in their project workflow a version was triggered. once guidelines were established, adhering to these requirements had a number of important internal payoffs. first, the digital object identifiers were persistent and unique. unique, persistent identifiers provided support for our users to be able to accurately reference the data obtained from the ipums system. second, references and related publications became trackable for our internal processes. use of those dois by researchers made it much easier for ipums to track the research based on our products and to use this information in applying for continued funding for ipums projects. third and most obviously, our data and metadata were captured and preserved, making our preservation work more accurate and complete. perhaps the greatest pay off was the recognition by ipums project managers of the need to capture and retain the content of each product version in a format that could be preserved and accessed for the purpose of research replication. finally, these developments motivated our organization to apply for the data seal of approval (now core trust seal).11 while initially ipums focused on the value of the first two points, the preservation and access requirements of obtaining a doi brought the role of the archive into clearer focus within the organization. as we dug into the cts application process, we quickly recognized that our policy documentation was scattered and incomplete, a byproduct of rapid institutional growth from 1991 to 2016. pulling existing materials together, assessing policy documentation that needed to be updated, and crafting new documentation to reflect practices already in place, was time consuming but necessary to document our workflow. business process model developing an organizational model that clearly and accurately reflected the workflow of our data projects and archival processes was a crucial step in developing our cts application materials. the process clarified the role of the archive within the organization and provided us with the means to clearly present that role and identify specific touchpoints to ipums project workflows. it also encouraged ipums project managers to look at the commonalities of their work across ipums projects, rather than just the uniqueness of their specific project. the cts application process requires applicants to describe their archival responsibilities within their organization using an oais model.12 using the oais model, we worked to identify where our archive obtained submissions (both external and internal), what actions we took once we obtained those https://doi.org/10.29173/iq1046 5/11 magnuson, diana l; thomas, wendy l. (2023) expanding our perspective: building a sustainable metadata culture, iassist quarterly 47(2), pp. 1-11. doi: https://doi.org/10.29173/iq1046 submissions, and how we delivered the products to users. these information packages were integrated into the ipums business process model (figure 1) to clarify where the oais information packages originated and how they moved through the process model from external and project sources, to the archive for management and user access. the oais model helped us to establish an expanded workflow model of the collection, harmonization, and publication work done within the various ipums projects, and importantly, align that workflow with the role of the archive. the new workflow model made clear how and where the archive interacted with the projects in terms of submitting data to the archive, packaging that data for persistent access, and delivering archived data to users once a dataset is replaced in the ipums live data access system by a new version. figure 1. ipums implementation of the oais model the oais model provided only a general model, however, and we soon determined that it was not flexible enough for our detailed, project-specific workflows; ipums is not a standard archive and thus the oais model could not reflect the full range of our activities. “the primary activities of ipums focus on acquiring data from an external producer, processing the data and related metadata to integrate it for the purposes of comparative research, providing a means of access to facilitate that research, and then delivering customized packages of data and metadata to the consumer.”13 to identify the commonalities between the processes of individual ipums projects while allowing for differences in the selection and ordering of tasks within each project over time, data curator wendy thomas drew on two business process models, the generic statistical business process model (gsbpm) and the generic longitudinal business process model (glbpm), to serve as templates in the creation of the ipums business process model (ipums bpm). the gsbpm was designed to model a standard framework for statistical organizations that could be used modernize statistical products through the use of harmonized language and shared methodologies.14 the glbpm is a modification of the gsbpm, developed to focus “on the longitudinal survey process as employed in longitudinal data gathering by academic, governmental, and private research organizations.”15 the ipums business process model is a customization of the gsbpm and the glbpm, “reflect[ing] the use of secondary data sources and the work of harmonization and integration to create a data infrastructure that supports research across time and space.16 internally, use of the ipums bpm provides a clear https://doi.org/10.29173/iq1046 6/11 magnuson, diana l; thomas, wendy l. (2023) expanding our perspective: building a sustainable metadata culture, iassist quarterly 47(2), pp. 1-11. doi: https://doi.org/10.29173/iq1046 visualization of our workflow from external submission of data, harmonization process, extraction systems, and archival preservation of metadata.17 the upper levels of the ipums bpm also proved useful in identifying points where metadata is being produced by the projects (shown in green). (figure 2) figure 2. ipums business process model we are instituting a workflow mapping strategy to further identify ipums process and metadata capture points for the data archive. currently our activity map has nine activity areas with subactivities within each area. we have added depth in several of these activity paths to provide detail on specific activities. these activity paths will be expanded as we work with the individual projects to ensure that they each see their set of activities and process paths through the model. our nine data projects have individualized project processes imposed by the “needs and constraints of their data sources and goals.”18 the mapping approach has several advantages for the projects, the administrative team, the it team, and the data archive. first, a common vocabulary is used across projects, administration, and it. because research staff sometimes move between projects, a common vocabulary streamlines those transitions. second, the technical team can readily identify tools that can be developed and used across projects, developing efficiencies and economies of scale around data/metadata management, preservation, and delivery. third, by establishing which products perform similar activities, ipums administration can identify process and tool developments that could benefit all projects. further, the activity mapping approach allows each project to identify their own path through the process activities, thus preserving their individualized workflow, while maintaining our institutional standard. finally, activity mapping identifies the areas of metadata production that require the attention of the archive for provenance and preservation purposes. (figure 3) figure 3. sip, aip and dip activity areas https://doi.org/10.29173/iq1046 7/11 magnuson, diana l; thomas, wendy l. (2023) expanding our perspective: building a sustainable metadata culture, iassist quarterly 47(2), pp. 1-11. doi: https://doi.org/10.29173/iq1046 the combined oais and ipums bpm makes it clear at what point a new version of data is deposited in the archive as a submission information package (sip). the model also accommodates any steps needed to meet the needs of individual projects--for example, handling the difference in creating snapshots for microdata and aggregate data products. significantly, the model clarifies the point at which the content of the sip becomes the custody of the archive and is no longer actively managed by the individual project. the content is then organized by the archive in an archive information package (aip) for the purposes of management and future dissemination as a distribution information package (dip) through a system separate from the ipums live data access system. leveraging documentation the utility of leveraging the documentation collected, refined, and/or produced for the cts application for various institute purposes became evident as we organized our materials. for example, some of the documentation we produced for the cts application will be used to support future funding application efforts. our organization can efficiently demonstrate to potential funders the preservation policy practices constructed and maintained for our data production and metadata capture. it is particularly important that ipums can certify that it follows international standards; our work involves harmonization of official statistical data and contributing organizations need to know that their data are being responsibly handled. further, because we took the time to thoughtfully think through and enumerate our data archive responsibilities, we will use this information and the visualizations we created to educate current staff, to inform stakeholders, and to onboard new employees. the focus of the cts on documenting and providing access to the documentation of their methods and processes makes the role of the archive transparent to the parent organization, the user, funding agencies, and future archive staff. staffing in non-traditional archives is often limited with low turnover. the danger of losing institutional knowledge of the reasoning, methods, and processes of the archive is very real as staff age out of their positions. transparency becomes a means of retaining institutional coherence. lastly, our documentation provides valuable institutional and procedural history, both of which are often overlooked or attempted piecemeal long after processes have changed or been discontinued. some of this documentation is publicly available, and other pieces are posted for internal use.19 the cts certification process provided justification for providing access to documentation on our processes in a more consistent and organized manner. it resulted in the addition of a working paper series focusing on ipums methodologies and development work. it has also highlighted the need to provide a process for routinely capturing the metadata currently provided on web pages and associating the content with the archived data files for past versions. preservation preserving our data products and resultant metadata is an obvious role of the ipums data archive. it became clearer to us as we worked to construct our cts application that an important additional activity of our data archive is preserving the enormous intellectual investment that went into collecting, integrating, organizing, cleaning, documenting, and distributing our unique data products. our project managers have historically been primarily concerned with preserving the end product that is disseminated to users, and less attentive to preserving the pieces of intellectual activity that contributed to the data harmonization process. while the projects all operate within the nine activity areas identified on the activity map, as noted, each project follows its own unique path that can be preserved by the data archive. the major importance of this is in obtaining buy-in from the projects themselves. the purpose of the model in identifying common processes does not require individual process steps to occur in the same order for each iteration of a project or between two or more projects. recognizing the individuality of each project reduces the concern of the project managers https://doi.org/10.29173/iq1046 8/11 magnuson, diana l; thomas, wendy l. (2023) expanding our perspective: building a sustainable metadata culture, iassist quarterly 47(2), pp. 1-11. doi: https://doi.org/10.29173/iq1046 that they are being forced into a model as opposed to a model being developed based on the work that they do. preserving the intellectual property relating to the processing and methodology that contributed to the development of our data products is a key and significant contribution to social science infrastructure. while general archives have a relatively clear definition of what they need to preserve, the nontraditional archive often has a broader mandate. they are responsible for ensuring that the work done in the development of research and/or product development is also preserved. this information provides context for the traditional data and metadata and is a resource for future researchers in terms of methodology, design, and decision making in the creation of data products or areas of research. guide for moving forward it is vital that a non-traditional archive define its place within a research organization and clarify its role in supporting the standing of its parent organization with funders, data providers, and the research community. for the ipums archive, the cts certification process was the opportunity to clearly articulate its role within the organization. as noted, questions from our funders provided the impetus by our administration to consider cts certification, which in turn facilitated our organization considering the work of ipums more broadly, recognizing the value of preservation and future access to the unique products and by-products of ipums projects. the obvious significance of this to the archive was acknowledging its functions as an integral activity within ipums. drawing on earlier work conducted within the archive that explored process models from government statistical agencies and longitudinal survey projects (organizations that produced one or more statistical products on a regular basis and followed similar processes for creating each consecutive iteration), archive staff sought to adapt these process models to the ipums context. the archives’ interest in these processes centered on capturing metadata generated by these processes and ensuring that metadata was being captured and preserved in a consistent format. the cts certification application provided a framework for identifying gaps in our documentation, processes that needed clarification, and justification for ensuring that we had a complete and coherent workflow for ensuring access to our products far into the future. as a first step, the cts submission process required a review of the current ipums processes and encouraged us to clearly document all our activities. we modeled our workflows according to commonly used approaches (oais and gsbpm). the adaptation of these models helped us to present the role of the archive in a way that was understandable to the ipums organization and product groups. specifically, we were able to: • clarify the role of the archive in ipums as a means of preserving the input to, and work of, each project within ipums • designate the touchpoints between the archive and the project teams to ensure that information was systematically passed into the archive as part of the project’s production flow • identify the value-added information provided by the archive to the content deposited to it in terms of content, organization, and adherence to international standards • identify gaps in the archive process that required attention to meet the requirements of cts certification overall, this process has been instrumental in presenting the importance of the role of the archive in ensuring that the intellectual work of the ipums projects is not lost over time and that the lessons of https://doi.org/10.29173/iq1046 9/11 magnuson, diana l; thomas, wendy l. (2023) expanding our perspective: building a sustainable metadata culture, iassist quarterly 47(2), pp. 1-11. doi: https://doi.org/10.29173/iq1046 this important work in data harmonization and integration are not lost to future researchers. data archives often serve a secondary role in research organizations, performing important work that is difficult to articulate and identify within the primary workflow of the research projects themselves. it is a bit like plumbing, its importance is only recognized when it fails. the cts certification process is a valuable means of specifying and justifying the work of the archive in a research organization and brings that work to the attention of the organization in terms that they can understand. for the research organization, the cts certification can be used to reassure data contributors that the organization understands and abides by international standards in the preservation of their data. finally, in preparing for cts certification, the organization goes through the process of updating, completing, and providing access to information on its processes and policies. these documents can be referenced by applications for funding, showing the overall organization and care for the data being created, and the commitment of the organization to both data quality, as well as long term care and access. conclusion the lessons we learned as part of the cts application process are applicable in other data archive contexts as well, especially those in which preservation activities are not viewed as the primary function of the institution. in the ipums context, creation and distribution of harmonized datasets from census and survey data has always been the main focus of the projects, as required by our funders. ipums evolved from a single project (1880 pums) to a suite of products intended to be supported over time with the capacity to add new harmonized data products. our institutional context has led to repositioning from a unit based in a small (history) department to an interdisciplinary research institute within the university of minnesota. healthy institutions change over time, responding to a myriad of contingencies, both internal and external. documenting this change provides a clear history of intent within the organization and offers a possible roadmap for other organizations experiencing similar growth, change, and development. the maturation of the role of the data archive within ipums reflects this dynamic growth and touches on the common issues of describing new functions, specifying the role of the archive in developing a sustainable metadata culture within the organization, and clarifying areas of management as specialization occurs within each contributing project. the multi-year process of preparing the ipums’ application for the cts encouraged us to provide models of our archival practices and clarify the details of our processes. discussions with the project groups continue to clarify the role of the archive within ipums and to pinpoint where the project workflow intersects with the archive, improving overall communication. the disruption of the ongoing pandemic also reinforces the importance of preserving institutional history that is both clear and accessible to support the inevitable transition in personnel that occurs over the life course of an institution. references lawrenz f. and paller, m.s. (2006) “transforming the university: recommendations of the task force on collaborative research, university of minnesota digital conservancy, https://hdl.handle.net/11299/567. magnuson, d.l. (2014) steven ruggles interview, university of minnesota, january 9, 2014. magnuson, d.l. (2015a) wendy thomas interview, university of minnesota, march 24, 2015. magnuson, d.l. (2015b) “curating our social science infrastructure: the mpc/ipums institutional history as a case study,” presented at the social science history association, baltimore, usa, november 12-15, 2015. https://doi.org/10.29173/iq1046 https://hdl.handle.net/11299/567 10/11 magnuson, diana l; thomas, wendy l. (2023) expanding our perspective: building a sustainable metadata culture, iassist quarterly 47(2), pp. 1-11. doi: https://doi.org/10.29173/iq1046 magnuson, d.l. and ruggles, s. (2022) “challenges of large-scale data processing in the 1990s: the ipums experience,” ieee annals of the history of computing, pp. 71-83. ruggles, s. (1991a) “integration of the public use samples of the u.s. census,” 1991 proceedings of the american statistical association, social statistics section. alexandria, va, asa, pp. 265-370. ruggles, s. (summer 1991b) “the u.s. public use census microdata files as a source for the study of long-term social change,” iassist quarterly. https://doi.org/10.29173/iq703. ruggles, s. (1992-1995) “integrated public use microdata series,” ses-9118299, nsf. ruggles, s. (2011) ”minnesota population center: self study report,” university of minnesota, june 7, 2011. van hook, j.l., bleakley, c.h. and hummer, r.a. (2016) ”external review of the minnesota population center,” june 14-16, 2016. endnotes 1 diana l. magnuson is curator and historian at the institute for social research and data innovation, university of minnesota (magn0031@umn.edu). 2 wendy l. thomas is retired curator at the institute for social research and data innovation, university of minnesota (wlt@umn.edu). 3 https://isrdi.umn.edu/. https://www.coretrustseal.org/. https://www.ipums.org/missionpurpose. 4 steven ruggles, interviewed by diana l. magnuson, university of minnesota, january 9, 2014. 5 steven ruggles, interviewed by diana l. magnuson, university of minnesota, january 9, 2014. 6 steven ruggles, interviewed by diana l. magnuson, university of minnesota, january 9, 2014. 7 steven ruggles, ”mpc strategic plan,” may 2, 2016 (email in possession of author). steven ruggles, institute name,” august 19, 2016 (email in possession of author). van hook, j.l., bleakley, c.h. and hummer, r.a. (2016) ”external review of the minnesota population center,” isrdi institutional archive, june 14-16, 2016. in 2016, all data projects took on the ipums prefix as part of their project name. since not all projects are microdata and some have access conditions that limit their usage, it is inaccurate to describe ipums as a “public use” microdata series. thus, since 2016 ipums is a brand, not an acronym. https://www.ipums.org/mission-purpose. 8 https://www.ipums.org/. 9 wendy thomas, interviewed by diana l. magnuson, university of minnesota, march 24, 2015. https://assets.ipums.org/_files/ipums/workflows/ipums_archive_workflow_nov2021.pdf. https://doi.org/10.29173/iq1046 https://doi.org/10.29173/iq703 mailto:magn0031@umn.edu mailto:wlt@umn.edu https://isrdi.umn.edu/ https://www.coretrustseal.org/ https://www.ipums.org/mission-purpose https://www.ipums.org/mission-purpose https://www.ipums.org/mission-purpose https://www.ipums.org/ https://assets.ipums.org/_files/ipums/workflows/ipums_archive_workflow_nov2021.pdf 11/11 magnuson, diana l; thomas, wendy l. (2023) expanding our perspective: building a sustainable metadata culture, iassist quarterly 47(2), pp. 1-11. doi: https://doi.org/10.29173/iq1046 10 https://assets.ipums.org/_files/ipums/workflows/ipums_archive_workflow_nov2021.pdf, p. 5. 11 https://www.coretrustseal.org/about/history/data-seal-of-approval-synopsis-2008-2018/. 12 https://assets.ipums.org/_files/ipums/workflows/ipums_archive_workflow_nov2021.pdf, p. 12. 13 https://assets.ipums.org/_files/ipums/workflows/ipums_archive_workflow_nov2021.pdf, p. 11. 14 https://statswiki.unece.org/display/gsbpm/gsbpm+v5.1. 15 https://ddialliance.org/sites/default/files/genericlongitudinalbusinessprocessmodel.pdf. 16 https://assets.ipums.org/_files/ipums/workflows/ipums_archive_workflow_nov2021.pdf, p. 16. 17 https://www.ipums.org/workflows. 18 https://assets.ipums.org/_files/ipums/workflows/ipums_archive_workflow_nov2021.pdf, p. 17. 19 https://www.ipums.org/about/more. https://doi.org/10.29173/iq1046 https://assets.ipums.org/_files/ipums/workflows/ipums_archive_workflow_nov2021.pdf https://www.coretrustseal.org/about/history/data-seal-of-approval-synopsis-2008-2018/ https://assets.ipums.org/_files/ipums/workflows/ipums_archive_workflow_nov2021.pdf https://assets.ipums.org/_files/ipums/workflows/ipums_archive_workflow_nov2021.pdf https://statswiki.unece.org/display/gsbpm/gsbpm+v5.1 https://ddialliance.org/sites/default/files/genericlongitudinalbusinessprocessmodel.pdf https://assets.ipums.org/_files/ipums/workflows/ipums_archive_workflow_nov2021.pdf https://www.ipums.org/workflows https://assets.ipums.org/_files/ipums/workflows/ipums_archive_workflow_nov2021.pdf https://www.ipums.org/about/more 1/18 costanzo, l.., and cooper, a., (2024) developing institutional research data management strategies in canada: setting the foundation for stronger partnerships and collaborations, iassist quarterly 48(2), pp. 1-18. doi: https://doi.org/10.29173/iq1096 the creative commons-attribution-noncommercial license 4.0 international applies to all works published by iassist quarterly. authors will retain copyright of the work and full publishing rights. developing institutional research data management strategies in canada: setting the foundation for stronger partnerships and collaborations lucia costanzo1 and alexandra cooper2 abstract the government of canada’s tri-agency formally launched the research data management (rdm) policy in march 2021 with the objective of supporting “canadian research excellence by promoting sound data management and data stewardship practices”. a central component of this policy requires postsecondary institutions eligible to administer canadian institutes of health research (cihr), the natural sciences and engineering research council of canada (nserc), or the social sciences and humanities research council of canada (sshrc ) funds to create an institutional rdm strategy by march 2023. a national survey was developed to gauge institutions’ readiness for developing an institutional rdm strategy required by the tri-agency. the survey emphasized increasing participation from diverse institutions to ensure that future support and resources are developed to address the distinct needs of institutions. recommendations from the survey report included increasing tri-agency involvement as institutions developed their institutional rdm strategies, encouraging institutions to collaborate, and the development of forums to provide support for disciplinary societies to have rdm conversations. as a result, three panel discussions covering the active stages (initial, planning, and execution) of developing an institutional rdm strategy were successfully delivered through the digital research alliance rdm (alliance rdm) to a diverse range of institutions. recognizing the needs of smaller institutions including cegeps, colleges, and polytechnics, an additional panel discussion was developed and delivered to this audience. keywords research data management (rdm), institutional strategy, tri-agency, digital research alliance of canada, partnerships, collaborations introduction institutions in canada are being required to develop institutional strategies as part of the evolving research data landscape. this article explores the tri-agency research data management (rdm) policy and the two surveys conducted to assess institutional preparedness before the policy’s release and oneyear after, and how recommendations from the second survey fostered stronger collaboration among stakeholders. copies of both the 2019 survey and 2022 survey can be found in the appendix. https://doi.org/10.29173/iq1096 https://science.gc.ca/site/science/en/interagency-research-funding/policies-and-guidelines/research-data-management/tri-agency-research-data-management-policy-frequently-asked-questions https://creativecommons.org/licenses/by-nc/4.0/ 2/18 costanzo, l.., and cooper, a., (2024) developing institutional research data management strategies in canada: setting the foundation for stronger partnerships and collaborations, iassist quarterly 48(2), pp. 1-18. doi: https://doi.org/10.29173/iq1096 background in order to place research within a canadian context, let's first look at some definitions. canadian research institutions fall into two categories: post-secondary and research institutions. post-secondary institutions include universities (degree granting institutions), colleges (certificate and diploma granting institutions that provide primarily technical, academic and/or vocational programs), and cegeps (similar to colleges but are only in quebec) and research institutions. researchers in canada have access to digital tools and services through the digital research alliance of canada (the alliance), a non-profit organization funded by the government of canada to serve canadian researchers. it integrates, champions and funds the infrastructure and activities required for advanced research computing (arc), research data management (rdm) and research software (rs). a group within the alliance that supports researchers with research data management (rdm) is the network of experts (noe), a national network of professionals supporting rdm across canada. several expert groups are affiliated with the noe and one of these groups is the research intelligence expert group (rieg) whose mandate is to provide evidence to guide the development of rdm, provide support for institutions around strategies, policies, capacity, and resources related to rdm, and identify gaps in the rdm landscape. finally, government funding for research is supplied by the tri-agency, which is composed of the canadian institutes of health research (cihr), the natural sciences and engineering research council of canada (nserc), and the social sciences and humanities research council of canada (sshrc). in march 2021, the tri-agency released the research data management policy that aims to promote sound rdm practices to support research excellence in canada. in the policy, the tri-agency outlined three key requirements: 1. institutional rdm strategy: all research institutions eligible for tri-agency funding were required to develop an institutional rdm strategy by march 2023. the published strategies can be found on the tri-agency rdm policy website. 2. data management plans: researchers must include data management plans in grant proposals. the phased introduction of this requirement began in spring 2022, with further funding opportunities being introduced over the next few years (current funding opportunities requiring dmps). 3. data sharing and access: researchers are expected to deposit all digital research data, metadata, and code supporting research conclusions in journal publications and preprints resulting from agency-supported research. while data sharing is not mandatory, researchers must provide appropriate access to the data, adhering to ethical, cultural, legal, and commercial requirements, as well as the fair principles and disciplinary standards. this requirement will be implemented gradually over the next few years. in anticipation of the policy's release, rieg conducted a survey in 2019 to evaluate institutions' preparedness in developing their institutional rdm strategies. the initial survey results indicated that some institutions had initiated the process of formulating a policy, but most had not due to the lack of information regarding the policy's content or the desire for a supportive community of practice to aid in the writing process. in march 2022 (one year after the policy's release), rieg conducted a follow-up https://doi.org/10.29173/iq1096 https://science.gc.ca/site/science/en/interagency-research-funding/policies-and-guidelines/research-data-management/published-institutional-research-data-management-strategies https://science.gc.ca/site/science/en/interagency-research-funding/policies-and-guidelines/research-data-management https://science.gc.ca/site/science/en/interagency-research-funding/policies-and-guidelines/research-data-management 3/18 costanzo, l.., and cooper, a., (2024) developing institutional research data management strategies in canada: setting the foundation for stronger partnerships and collaborations, iassist quarterly 48(2), pp. 1-18. doi: https://doi.org/10.29173/iq1096 survey to assess the progress of institutions in creating their institutional strategies and identify any additional resources required to complete their strategy. 2019 survey results the 2019 survey was developed prior to the tri-agency policy being released and was intended to measure progress, if any, was being made and what other supports portage and other stakeholders could provide. the survey received eighty-eight completed submissions, with the majority of respondents coming from universities, followed by cegeps, and the remaining from research centres and other institutions. regionally, ontario, quebec and the west institutions provided about eighty-five percent of the responses and the atlantic region about fifteen percent. to start preparing for the policy’s release, most respondents reported they had assessed their institution’s capacity, reviewed support material for strategy development, and formed, or were in the process of forming, a working group to develop their institution’s strategy. few respondents indicated they were in the final stages of planning partly due to the delay in the release of the tri-agency’s final policy (see figure 1). about the delay, one commenter noted « compte tenu du report de l’entrée en vigueur d’une politique des conseils sur la gdr, les travaux institutionnels sur cette question ont ralenti ; lorsque les attentes exactes des conseils seront connues, il sera plus facile de s’y remettre. » [translation: given the delay in implementation of the rdm policy, institutional work on this issue has slowed; when the exact requirements of the tri-council are known, it will be easier to continue the momentum]. figure 1. status of institutional research data management strategies in development by institution. most respondents indicated their institution had formed one or more working groups to start developing their institutional strategy. for institutions that had formed working groups, on average more than four different offices were represented. figure 2 presents the range of offices represented in these groups. https://doi.org/10.29173/iq1096 4/18 costanzo, l.., and cooper, a., (2024) developing institutional research data management strategies in canada: setting the foundation for stronger partnerships and collaborations, iassist quarterly 48(2), pp. 1-18. doi: https://doi.org/10.29173/iq1096 figure 2. the range of offices or departments represented on working groups tasked with developing institutional strategies at the institutions of respondents. portage3 developed a range of tools to support institutions in writing their institutional strategies. most respondents were using the portage institutional strategy template and guidance document, which the majority of institutions rated as being helpful or very helpful. as well, many indicated that they were using existing policies at their own institutions to inform their strategies. suggestions from respondents for additional guidance and support frequently referenced the need for sample strategies and best practices documents (see figure 3). communication tools that institutions and/or departments could adopt in their outreach and promotion were also suggested, for instance, the national science foundation’s “dear colleague” series of letters. figure 4 summarizes interest in the various methods of receiving additional support, with webinars and online documentation being of especially high interest. in-person training and local or regional workshops were rated more favorably than national-level training, indicating a possible gap in more localized support/connection. comments reflected interest in knowing what actions peer institutions in their region/province were taking. https://doi.org/10.29173/iq1096 https://alliancecan.ca/en/services/research-data-management/learning-and-training/training-resources#heading-institutional-strategies-guidance https://www.nsf.gov/news/news_summ.jsp?cntn_id=127138 5/18 costanzo, l.., and cooper, a., (2024) developing institutional research data management strategies in canada: setting the foundation for stronger partnerships and collaborations, iassist quarterly 48(2), pp. 1-18. doi: https://doi.org/10.29173/iq1096 figure 3. the range of resources being used to develop institutional strategies by survey respondents. total affirmative responses for each resource type are presented. figure 4. interest in the method of delivery of training and guidance to support the development of institutional strategies. https://doi.org/10.29173/iq1096 6/18 costanzo, l.., and cooper, a., (2024) developing institutional research data management strategies in canada: setting the foundation for stronger partnerships and collaborations, iassist quarterly 48(2), pp. 1-18. doi: https://doi.org/10.29173/iq1096 2019 survey recommendations this survey found that many institutions were in the beginning of developing their strategy, they were also hesitant to move forward ahead of any final tri-agency policy. it was anticipated that the announcement of the policy would spur the institutions into action. the following recommendations provided suggestions that the tri-agency, portage, and other organizations supporting research data management could take to support institutions. ● collect and share strategies as they become available and make them available in a single online location to improve accessibility. ● develop communities of practice so that stakeholders and peers have opportunities to meet counterparts at regional and national levels to receive additional guidance, compare notes, learn from one another’s approaches, and develop strategies. ● provide more explicit best practices with clear guidance on how to best meet the requirements outlined in the new policy. 2022 survey results the survey received a total of ninety-two responses. invitations to participate in the survey were distributed through listservs primarily targeting rdm professionals. institutions were encouraged to collaborate with relevant stakeholders involved in developing their institutional rdm strategy while completing the survey. the largest number of respondents were from universities followed by colleges/cégeps then other institution types which included research centers and research hospitals. many respondents were from ontario, which has a greater representation across all institutions compared to other regions. this was followed by both the western and quebec regions which had a similar number of responses. the atlantic regions had the lowest responses. similar to the 2019 results, the main offices and services contributing to the survey were the research office and library followed by ethics, it, researchers, and cio. in consultation with survey stakeholders, the following office and services were added to the question about institutional stakeholdersexecutive management, indigenous office/representative/council, legal, privacy office, graduate studies, and records management. to get a better idea on how to support institutions in the development of their institutional rdm strategies, institutions were asked to report on the barriers they face when meeting the agencies rdm policy requirement. in figure 5, the top 3 barriers are: • lack of time and budget – which is a barrier facing many institutions • lack of understanding and awareness of the agencies expectations • and the lack of supporting materials – more resources are needed to be developed https://doi.org/10.29173/iq1096 7/18 costanzo, l.., and cooper, a., (2024) developing institutional research data management strategies in canada: setting the foundation for stronger partnerships and collaborations, iassist quarterly 48(2), pp. 1-18. doi: https://doi.org/10.29173/iq1096 figure 5: barriers to meeting the tri-agency rdm policy as reported by institutions. knowing these barriers illustrates there is a need to strengthen the relationship between the agencies and institutions and for collaborations to develop materials between all partners including the network of experts. however, lack of rdm knowledge and there being no barriers are at the bottom of the list for institutions. to determine which resources were used to develop institutional rdm strategies, institutions were asked to report which they have used. in figure 6, the top resources utilized included: • alliance rdm resources developed through the network of experts existing policies and resources from other institutions; • consulting fellow colleagues at other institutions and maturity assessment models which include mamic, rise; • the development of these resources relies on existing collaboration between institutions, alliance rdm team, agencies, and data professionals. https://doi.org/10.29173/iq1096 https://zenodo.org/records/5745493 https://digitalresearchservices.ed.ac.uk/resources/rise-framework 8/18 costanzo, l.., and cooper, a., (2024) developing institutional research data management strategies in canada: setting the foundation for stronger partnerships and collaborations, iassist quarterly 48(2), pp. 1-18. doi: https://doi.org/10.29173/iq1096 figure 6: resources institutions used to develop their institutional rdm strategies. the survey asked institutions where in the process they were in developing an institutional rdm strategy in early 2022. the coordination of a working group or committee at an institution did not seem to pose much of a challenge, however, for the remaining processes, institutions appeared to have greater difficulties. the most challenging or difficult step in the process was estimating the cost of rdm related activities. addressing disciplinary approaches to rdm were reported as difficult, like the other challenges, but also had the highest response of not applicable (almost one quarter of institutions) (figure 7). https://doi.org/10.29173/iq1096 9/18 costanzo, l.., and cooper, a., (2024) developing institutional research data management strategies in canada: setting the foundation for stronger partnerships and collaborations, iassist quarterly 48(2), pp. 1-18. doi: https://doi.org/10.29173/iq1096 figure 7. challenges to developing an institutional rdm policy as reported by institutions. 2022 survey recommendations in 2022, institutions were actively engaging in the process of creating institutional rdm strategies in order to meet the triagency deadline of march 2023 to have a strategy posted. major stakeholders identified in both the 2019 and 2022 surveys included: • libraries who support researchers with the curation, preservation, storage and reuse of research data; • research offices who have a broad responsibility for administration of sponsored research and related policies and services; • office of cio who is responsible for almost all areas of it; • ethics board who governs the standards of conduct for researchers; • researchers who are the ones conducting the work; and • it/systems who provide the technical infrastructure including storage, security, high performance computing, etc. additional key stakeholders also consulted in the 2022 survey included: • indigenous office/representative/council who represent the rights of indigenous peoples to control data from and about their communities and lands, articulating both individual and collective rights to data access and to privacy; • privacy office who ensure institutional compliance regarding personal information; and • legal office who manage the legal risks arising through the institution's activities and objectives. by thoughtfully expanding the stakeholder groups to include the new stakeholders, the institutions created a more integrated rdm ecosystem needed to support researchers creating and using research data requiring specific support for care, management or sharing. (figure 8) https://doi.org/10.29173/iq1096 10/18 costanzo, l.., and cooper, a., (2024) developing institutional research data management strategies in canada: setting the foundation for stronger partnerships and collaborations, iassist quarterly 48(2), pp. 1-18. doi: https://doi.org/10.29173/iq1096 figure 8: the new integrated rdm ecosystem. enhancing collaboration also requires building upon the rdm community of experts that includes network of experts, alliance rdm, stakeholders, and tri-agency. by increasing the involvement of the tri-agency in developing these resources, not only is the rdm community enriched, but needed clarity is provided on the expectations of what an institutional rdm strategy should be along with a central location to share all published. institutions have and will save staff time through the collaborative creation of additional support and resources, as well, this enhanced collaboration allows for the development of forums and support for disciplinary societies to have rdm conversations, express their needs and concerns, and encourage grassroots rdm initiatives within disciplines and across disciplines within the global open science context. also recommended in the survey report was that additional resources be developed to address the barriers and challenges encountered by institutions when developing institutional strategies. in response, a three-part online panel discussion series was developed to guide institutions along the three active stages of rdm strategy development: 1. initial stage discussing steps used to complete this stage including forming a working group/committee, reviewing available support material, and assessing institutional rdm capacity; 2. planning stage discussing how to envision the future state of rdm and creating either a roadmap or action plan; 3. execution stage discussing how to create a draft strategy document and articulate the institutional path forward through a roadmap or action plan during these sessions, it was determined that a fourth session was needed for colleges, cegeps, polytechnic schools, and other research institutions to discuss their barriers, challenges and shared experiences outside of the university context. response to these sessions was enthusiastic with about 80-100 people attending each one. moving forward, additional projects will be developed to continue the collaboration created as a result of the institutional strategy survey report. as of the summer of 2023, a working group composed of rieg, the alliance-rdm, university of ottawa heart institute, and the tri-agency, has begun https://doi.org/10.29173/iq1096 11/18 costanzo, l.., and cooper, a., (2024) developing institutional research data management strategies in canada: setting the foundation for stronger partnerships and collaborations, iassist quarterly 48(2), pp. 1-18. doi: https://doi.org/10.29173/iq1096 reviewing/summarizing all the submitted institutional rdm strategies with an aim to summarize the content covered and assess gaps and needs across institutional type and size and geographic location in canada. this working group will also include additional members of the rdm community in canada. a second project will begin in late 2023/early 2024 to rerun the rdm capacity survey originally run in 2019. the survey will evaluate the current efforts of canadian research institutions in developing and allocating human, organizational, infrastructure, and fiscal resources for research data management (rdm) on their campuses. information from the results of the rdm institutional strategy review working group will be used to help guide the rewriting of the capacity survey. abstract 1: 2019 survey questionnaire survey introduction the following questionnaire was designed by the portage research intelligence expert group (rieg) to gather information on the successes and challenges of canadian research institutions in developing an institutional strategy for research data management (rdm) in response to the tri-agency’s draft research data management policy. this bilingual questionnaire surveys the progress made by canadian research institutions in developing an institutional strategy for rdm on their campus, and solicits suggestions for additional support that portage network and other stakeholders could provide to assist with these efforts. this survey consists of 5 questions and is expected to take 5 minutes of your time to complete. information gathered by this survey will be used to summarize the state of institutional strategy development across canada, and inform the creation of additional resources to support canadian institutions in developing rdm strategies. a report summarizing our survey results will be shared publicly with the research community soon after the survey closes. individual survey responses will not be shared publicly. about us: the research intelligence expert group (reig) is working to gather evidence to guide the development of best practices in rdm in canada, and inform stakeholder communities about issues in related policy and practices. this questionnaire was developed by members of rieg’s strategic planning working group: ● shahira khair, university of victoria ● mark leggott, research data canada ● will meredith, royal roads university ● tatiana zaraiskaya, university of new brunswick for more information on the objectives and membership of this expert group, please consult the portage website. if you have any comments or questions about this survey, please contact portage@carl-abrc.ca. https://doi.org/10.29173/iq1096 https://doi.org/10.14288/1.0388722 12/18 costanzo, l.., and cooper, a., (2024) developing institutional research data management strategies in canada: setting the foundation for stronger partnerships and collaborations, iassist quarterly 48(2), pp. 1-18. doi: https://doi.org/10.29173/iq1096 survey questions 1. contact info (open text fields) ● name ● institution ● position title ● office/service department (e.g. library, research office, systems, faculty, etc.) ● e-mail ● phone 2. check all that apply to indicate where you are in your process of developing institutional strategy for rdm: (check all that apply) ● have not started yet ● reviewing available support material (e.g. portage institutional strategy template and guidance document) ● formed a working group/committee ● assessing institutional capacity ● defining a desired end state ● creating a roadmap/action plan ● creating a draft strategy document ● draft strategy document currently under review by university administration ● finalized the institutional strategy document. ● posted the institutional strategy document on the website. if selected, please provide a link. (open text field) 3. please provide any additional comments on your progress (open text field). 4. if you have formed a working group (wg) or committee to develop an institutional rdm strategy for your institution, what stakeholders are currently involved in this wg or committee? (check all that apply) ● have not formed a wg or a committee ● library ● research office ● cio ● ethics board ● researchers ● it ● other (please specify) (open text field) 5. if any of options b through h have been selected above, then... please provide contact information for the appropriate member who we could contact for further information about this working group. (open text fields) ● name ● position ● e-mail 6. what resources are you using to develop an institutional strategy? (check all that apply) ● portage strategy template and/or guidance document ● existing policies from other institutions https://doi.org/10.29173/iq1096 https://portagenetwork.ca/wp-content/uploads/2018/03/portage-institutional-strategy-template-v4-en.pdf https://portagenetwork.ca/wp-content/uploads/2018/03/portage-institutional-strategy-guidance-v4-en.pdf https://portagenetwork.ca/wp-content/uploads/2018/03/portage-institutional-strategy-template-v4-en.pdf https://portagenetwork.ca/wp-content/uploads/2018/03/portage-institutional-strategy-guidance-v4-en.pdf 13/18 costanzo, l.., and cooper, a., (2024) developing institutional research data management strategies in canada: setting the foundation for stronger partnerships and collaborations, iassist quarterly 48(2), pp. 1-18. doi: https://doi.org/10.29173/iq1096 ● guidance documents ● consultants ● workshops ● other (please specify) (open text field) 7. if you selected a) in the previous question, please rank your perceived usefulness of the portage tools (rating table: not useful, somewhat useful, very useful). ● portage strategy template ● guidance document 8. please provide any suggestions for improvement (open text field) 9. would you be interested in receiving additional guidance during the process of developing institutional strategy? please rank the following (rating table: not interested, somewhat interested, interested): ● a workshop at your institution ● a workshop at a regional venue ● a workshop at a national venue ● a webinar ● online documentation with tutorials ● other please specify (open text field) 10. would you be interested in acting as a pilot site to develop a strategy, which could be used for exemplar purposes? this could include being part of a small group of institutions working with a core team to develop sample institutional strategy documents, which would then be shared with the broader community. appendix 2: 2022 survey questionnaire tri-agency institutional rdm strategy survey introduction thank you to all institutions who have already completed the survey. for those who haven’t yet had a chance to complete the survey, we have extended the deadline by one week; the survey will now close on tuesday, april 12th, 2022. on march 15th, 2021, the tri-agency research data management policy was launched. in preparation for these changes, the digital research alliance of canada (the alliance) research data management (rdm) research intelligence expert group (rieg) is conducting a brief survey of all canadian research institutions (colleges, universities or other research institutions receiving research grant funds from the tri-agencies) to determine challenges and successes in developing an institutional strategy for rdm. this bilingual questionnaire surveys the progress made by canadian research institutions in developing an institutional strategy for rdm on their campus and solicits suggestions for additional support that the alliance and other stakeholders could provide to assist with these efforts. this survey was previously conducted in june 2019 and a report summarizing the findings was generated by rieg. https://doi.org/10.29173/iq1096 14/18 costanzo, l.., and cooper, a., (2024) developing institutional research data management strategies in canada: setting the foundation for stronger partnerships and collaborations, iassist quarterly 48(2), pp. 1-18. doi: https://doi.org/10.29173/iq1096 the survey consists of 8 questions and is expected to take 15-20 minutes of your time to complete. we encourage this survey to be completed in coordination with institutional stakeholders who may be involved with developing your institutional rdm strategy (one response per institution). the survey will remain open until april 5th, 2022 and a pdf version is available for previewing. information gathered by this survey will be used to summarize the state of institutional strategy development across canada, and inform the creation of additional resources to support canadian institutions in developing rdm strategies. individual survey responses will not be shared. they will be used for program planning and evaluation only. a report summarizing the survey results will be shared publicly with the research community soon after the survey closes. any questions can be directed to alexandra cooper, chair of rieg (coopera@queensu.ca) or lucia costanzo, research, intelligence, and assessment coordinator (lucia.costanzo@engagedri.ca). thank you for taking the time to complete and share this survey. survey questions 1. province choose one ● alberta ● british columbia ● manitoba ● newfoundland and labrador ● new brunswick ● northwest territories ● nova scotia ● nunavut ● ontario ● prince edward island ● quebec ● saskatchewan ● yukon ● other 2. institution type check one ● university ● college/cégep ● research institute ● government ● research hospital ● other 3. institution name https://doi.org/10.29173/iq1096 15/18 costanzo, l.., and cooper, a., (2024) developing institutional research data management strategies in canada: setting the foundation for stronger partnerships and collaborations, iassist quarterly 48(2), pp. 1-18. doi: https://doi.org/10.29173/iq1096 4. office/service(s) contributing to survey response check all that apply ● library ● research office ● office of cio ● ethics board ● researcher ● it/systems ● legal (office) ● other 5. what barriers are you realizing in meeting the tri-agency rdm policy requirements for an institutional strategy? check all that apply ● lack of institutional understanding and awareness of tri-agency expectations. please explain ● lack of rdm knowledge please explain ● lack of resources (time, budget, personnel, etc) please explain ● lack of availability of support materials. please explain ● none ● other. please explain 6. check all that apply to indicate where you are in your process of developing an institutional strategy for rdm: ● not yet started ● formed a working group/committee ● reviewing available support material (e.g. alliance-rdm (portage)) ● assessing institutional rdm capacity ● envisioning the future state of rdm ● creating a roadmap/action plan ● creating a draft strategy document ● articulating the institutional path forward through a roadmap or action plan 7. process part two (check one) ● draft strategy document currently under review ● if you have launched your strategy, please provide a link 8. please provide any additional comments on your progress. 9. if you have formed a working group or committee to develop an institutional rdm strategy for your institution, what stakeholders are currently involved in this working group or committee? check all that apply ● have not formed a working group or a committee [if checked, no other options can be checked]. 10. if you have formed a working group or committee to develop an institutional rdm strategy for your institution, what stakeholders are currently involved in this working group or committee? ● institutional library systems ● research office/institutional research https://doi.org/10.29173/iq1096 16/18 costanzo, l.., and cooper, a., (2024) developing institutional research data management strategies in canada: setting the foundation for stronger partnerships and collaborations, iassist quarterly 48(2), pp. 1-18. doi: https://doi.org/10.29173/iq1096 ● office of the cio ● research ethics board ● researchers ● it services ● graduate studies ● executive management ● records management ● privacy office ● legal services ● indigenous office/representative/council ● other (please specify) 11. what resources are you using to develop an institutional strategy? check all that apply ● alliance rdm (portage) institutional strategies resources list which ones ● existing policies and resources from other institutions list which ones ● maturity assessment models (e.g., mamic, rise) ● consultants ● workshops ● fellow colleagues at other institutions ● other (please specify) 12. please rank your perceived usefulness of the alliance rdm (portage) tools. have not used not useful somewhat useful very useful strategy development template, v3 (released november 2021) strategy template and guidance, v2(prior to november 2021) maturity assessment model in canada (mamic) videos discussion prompts brief guide primer ● please provide any suggestions or other tools you would like to see developed. 13. what challenges or difficulties have you encountered during the process of developing your institutional strategy? please rate the following will not be doing not started not difficult difficult very difficult not applicable https://doi.org/10.29173/iq1096 17/18 costanzo, l.., and cooper, a., (2024) developing institutional research data management strategies in canada: setting the foundation for stronger partnerships and collaborations, iassist quarterly 48(2), pp. 1-18. doi: https://doi.org/10.29173/iq1096 coordination of the working group or committee engaging researchers addressing disciplinary approaches assessing institutional capacity creating a roadmap/action plan defining a desired end state estimating the cost of rdm related activities, services, and infrastructure ● other, please specify 14. if you have answered difficult or very difficult, please explain why. 15. would additional guidance and support be helpful as you continue to develop your institutional strategy? check one ● yes ● no 16. what type of guidance or support would your institution need? please rank the following options. not interested somewhat interested interested a workshop for your institution a workshop at a regional venue a workshop at a national venue a webinar online documentation with tutorials videos ● other, please specify. 17. would you like to provide additional comments? https://doi.org/10.29173/iq1096 18/18 costanzo, l.., and cooper, a., (2024) developing institutional research data management strategies in canada: setting the foundation for stronger partnerships and collaborations, iassist quarterly 48(2), pp. 1-18. doi: https://doi.org/10.29173/iq1096 references costanzo, lucia, & cooper, alexandra. (2023, june 1). setting the foundations for stronger partnerships and collaborations for developing institutional rdm strategies in canada. iassist 2023, philadelphia, pa, usa. zenodo. doi: https://doi.org/10.5281/zenodo.8010869 costanzo, lucia and cooper, alexandra. (2022). “institutional rdm strategy survey report”. zenodo [online]. doi: http://doi.org/10.5281/zenodo.6830003 portage research intelligence expert group. (2019). “portage research intelligence expert group institutional rdm strategy survey summary of results”. zenodo [online]. doi: https://doi.org/10.5281/zenodo.3962831 tri-agency (accessed 2023). tri-agency research data management policy, (available at https://science.gc.ca/site/science/en/interagency-research-funding/policies-and-guidelines/researchdata-management/tri-agency-research-data-management-policy) 1 university of guelph 2 queen’s university 3 the alliance rdm was known as portage at the time of this survey. https://doi.org/10.29173/iq1096 https://doi.org/10.5281/zenodo.8010869 http://doi.org/10.5281/zenodo.6830003 https://doi.org/10.5281/zenodo.3962831 https://science.gc.ca/site/science/en/interagency-research-funding/policies-and-guidelines/research-data-management/tri-agency-research-data-management-policy https://science.gc.ca/site/science/en/interagency-research-funding/policies-and-guidelines/research-data-management/tri-agency-research-data-management-policy 1/14 lockhart, j, xesi, x & chiware, e (2024) working towards securing and building a trusted institutional research data repository through the coretrustseal process: case of cape peninsula university of technology data repository, iassist quarterly 48(3), pp. 1-14. doi: https://doi.org/10.29173/iq1111 the creative commons-attribution-noncommercial license 4.0 international applies to all works published by iassist quarterly. authors will retain copyright of the work and full publishing rights. working towards securing and building a trusted institutional research data repository through the coretrustseal process: case of cape peninsula university of technology data repository janine lockharti, xabiso xesiii and elisha r t chiwareiii abstract in support of the open science movement and as a signatory of the berlin declaration, the cape peninsula university of technology has since 2013 developed various systems, infrastructures and workflows to support open access and good research data management practices at the institution, providing a highly functional environment. institutional policies that include a research data management policy and an open access policy, data deposit guidelines and data deposit platforms are currently in place and utilized by affiliated postgraduate students and researchers from faculties, research units and entities as well as researchers from academic support units in alignments with fair principles. the requirement of postgraduate students to submit their research data with their theses for graduation purposes has increased the advocacy and publishing of datasets. the purpose of this paper is therefore to highlight the initial developmental trajectory of a research data repository and what was achieved to date. this includes the selection of the platform through the ilifu project in the western cape, the implementation and strengthening of the repository review workflows to include a number of stakeholder players to ensure the quality and integrity of the data as well as ethics approval checks, the development of the data management planning tool and a more recent upgrade to include a section for the south africa’s protection of personal information act (nr 4 of 2013) compliancy, advocacy, training and processes that the institution has embarked on to secure the research data platform through proper preservation methodologies and approaches. some challenges are discussed and how these were addressed. the paper also outlines the process of how the institution embarked on applying to have the research data repository certified as trustworthy through the coretrustseal. keywords research data management, research data repositories, academic libraries, cape peninsula university of technology, coretrustseal, trusted digital repository (tdr), data preservation https://doi.org/10.29173/iq1111 https://creativecommons.org/licenses/by-nc/4.0/ 2/14 lockhart, j, xesi, x & chiware, e (2024) working towards securing and building a trusted institutional research data repository through the coretrustseal process: case of cape peninsula university of technology data repository, iassist quarterly 48(3), pp. 1-14. doi: https://doi.org/10.29173/iq1111 introduction research data repositories have been on the rise in the last few decades as various international and national organizations, funding agencies, publishers and research communities demand effective and efficient online access to digitally stored research data. government agencies are mandating funding recipients to make research data publicly available in approved repositories (hutchison et al., 2021). research data repositories and archives are important components of the research infrastructure, providing resources and services to research communities. the growth and development of research data repositories is also seen as a direct response to several international and national declarations on open science that aim to advance discoverability and use of scientific research publications and data in an open and transparent manner. the provision of research data repositories is at several levels, including: institutional, regional, national, international and discipline specific. institutional level research data repositories are often maintained to store internally generated research data whose curation, preservation and dissemination is governed by clear policies and guidelines. at regional, national and international levels data repositories can come in different forms, including discipline specific to generalist types of repositories. research data repositories have had the positive effect on the research processes by allowing researchers from different institutions and disciplines to share research workflows, experimental methods and data. to ensure that there is a global standard on the discoverability and usability of research data, key stakeholders from industry, academic, funding agencies and publishers have designed and endorsed a concise set of principles known as fair (findable, accessible, interoperable, reusable) data principles (wilkinson et al., 2016). south africa stands out in sub-saharan africa with its advances and sustained funding for the development of research infrastructure across institutions and at national level. the level of funding and development is enabling the development and integration of data repositories into existing research infrastructures (chiware and becker, 2018). according to the registry of research data repositories (re3data)iv, south africa has the most registered repositories than any other country in sub-saharan africa with more than 17 institutional and discipline specific platforms. most of the research data repositories are found in academic institutions which are the main knowledge production centres in the country. the growth of south african repositories has also been bolstered by the main government research funding agency, the national research foundation, which in 2015 mandated that all its grant recipients must deposit their research outputs, including research data, into trusted institutional repositories (nrf, 2015). a significant increase in the number of data repositories and data deposits in the past ten years is supported largely by several universities’ research data management (rdm) and open science policies. one of the key issues in the development of research data repositories at the global level is a roadmap to achieving trustworthy digital repositories (tdr) status. johnston (2012) outlines that a trusted digital repository is “a set of metrics that are used to certify that a given repository is an appropriate custodian of a collection of digital assets”. furthermore, johnston (2012) emphasizes that a trustworthy digital repository must be a stable and sustainable platform, following a clear set of policies and procedures for the sound management of digital assets, housed within secure technical environments. faundeen (2017) suggests that in order to secure and gain certification for digital https://doi.org/10.29173/iq1111 3/14 lockhart, j, xesi, x & chiware, e (2024) working towards securing and building a trusted institutional research data repository through the coretrustseal process: case of cape peninsula university of technology data repository, iassist quarterly 48(3), pp. 1-14. doi: https://doi.org/10.29173/iq1111 repositories, it is important to follow the guidelines set by national and international organizations and establish national policies and data governance guidelines. bak (2016), was of the view that “the notion of trust within trustworthy digital repositories standards culture is itself evolving in the positive direction that emphasizes user perceptions of trust rather than seeking to establish objective evidence of trust”. yoon (2014) also emphasized that much attention has been paid to establishment of iso standards towards trustworthy digital repositories, with very little attention on the users who are equally important. it is important for data repositories to maximize research outcomes and facilitate collaboration and sharing, as well as, ensuring the quality of the data and accompanying services (mehnert et al., 2019). johnston (2012) clearly states that to be certified as a trustworthy digital repository, organizations must undergo an audit which will ensure that their repository meets all criteria of certifying body. johnston outlines the need for overall information management processes, access, data security systems, and risk management parameters. in this paper we describe the technical and non-technical process around the historical development of esango, the research data repository at the cape peninsula university of technology (cput) powered by figshare, and the roadmap towards striving to achieve the status of a trustworthy research data repository through the coretrustsealv process. cape peninsula university of technology research data management services the cape peninsula university of technology’s research, technology and innovation (rti) 10-year blueprint (2012) outlined the key role that the university library was to play in supporting research, technology and innovation at the institution, which included: “curation, dissemination and promotion of the traditional outputs of research in terms of articles and theses, and curation of research data and innovation output, including enhanced research data management systems”. this recognition of the library’s role in the provision of research data management services provided the basis on which rdm services were developed at the institution (chiware and mathe, 2015). since 2013, cput has put in place policies and developed systems and workflows to support good rdm practices supporting a strong open access environment at the institution. these guideline and infrastructure were used by students and researchers at all levels (chiware and mathe, 2015). research data environment at the beginning, research data management (rdm) at cput was placed in the library in a division called knowledge, information and technology services (kits). this division was instrumental in creating platforms, systems, and processes for research data management. to advance the adoption of rdm practices, cput libraries established collaboration with several institutional stakeholders to develop policies, build infrastructure, train library staff, and conduct awareness and advocacy campaigns with academic staff and researchers. tripathi et al. (2017) highlights the importance of data management plans (dmp) and the role of libraries in supporting researchers in storing and accessing their data. according to wilkinson et al. (2016), good rdm is important for knowledge discovery and innovation, and for subsequent data and knowledge integration and reuse by the community after the data publication process. to encourage data discoverability and reuse, in 2020 cput’s higher degrees committee (hdc) mandated that as part of the graduation process, master https://doi.org/10.29173/iq1111 4/14 lockhart, j, xesi, x & chiware, e (2024) working towards securing and building a trusted institutional research data repository through the coretrustseal process: case of cape peninsula university of technology data repository, iassist quarterly 48(3), pp. 1-14. doi: https://doi.org/10.29173/iq1111 and doctorate students must share the datasets used in their research on esango, the cput research data repository. in addition to ensuring that essential research data is kept accessible, available for future reference, and verification and strengthens the transparency of the research this requirement also increased advocacy for rdm practices at the institution. cput has taken a stance in applying the four foundational principles of good rdm that is findability, accessibility, interoperability and reusability, known as fair principles, in managing research data (ntja, 2022). the driving force in promoting good rdm stewardship is due to the impact it has to highquality digital publications that facilitate and simplify the ongoing process of discovery, evaluation, and reuse in teaching, learning and research. policy framework understanding the importance of data for scientific advancement and human development, cput’s administration was interested in guaranteeing that data produced by researchers at the institution are of high quality and are openly accessible for reuse by other researchers. cput also recognized that good research data management procedures are essential for a productive and efficient research process. for example, it is important to follow the ethical and legal guidelines when working with sensitive data. good rdm responds to the open science and open data movements, which urge for more transparency and efficiency in research to accelerate the scientific enterprise. in 2019, the south african government published a white paper on science, technology, and innovation designed to strengthen the national system of innovation (nsi) through research output sharing. in response to this white paper and to establish an rdm policy landscape at cput, in 2020, the university formed the policy working group (pwg) comprised of faculty and staff from across the different disciplines. the pwg reviewed the university’s policies in place from 2013. the policies were revised and published as open access (oa) policy (2021a) and research data management (rdm) policy (2021b). the objective of the cput rdm policy is to govern research data management, promote reproducibility and ensure compliance by all university staff and students. the policy aims to establish guidelines and procedures for the management, ownership, sharing, access, storage, preservation, and ethical handling of research data within the university community. it requires all individuals affiliated with the university to adhere to the defined principles and practices of responsible research data management. having a policy in place is only the first step towards good rdm practices. often researchers find these policies to be cumbersome and are therefore reluctant to adopt them. researchers may be more likely to adopt rdm policies when they are required by external funding agencies or publishers. offering training, consultation, and collaboration to university staff and researchers can also help with achieving full implementation of the policies. building trust in the value of repositories to the researchers’ work may also help with their adherence to rdm policies. this trust is often built based on the value of the repositories to researchers’ work, the visibility of reused data and how institutions, funders and publishers respond to the compliance mandates (curty, 2016; swauger and vision, 2015). data repository https://doi.org/10.29173/iq1111 5/14 lockhart, j, xesi, x & chiware, e (2024) working towards securing and building a trusted institutional research data repository through the coretrustseal process: case of cape peninsula university of technology data repository, iassist quarterly 48(3), pp. 1-14. doi: https://doi.org/10.29173/iq1111 ilifuvi is a regional node in the western cape, south africa, known as a tier ii node, in the national infrastructure. it supports research mostly in the astronomy and bioinformatic fields. ilifu is funded partly by the department of science and technology (dst) through their data-intensive research initiative of south africa (dirisa). the regional project involved four universities, cape peninsula university of technology, university of cape town, university of the western cape and stellenbosch university. the research data management and open science component of the project involved the development of policies and guidelines for research data management, sharing, reuse, governance, and quality. as part of this process, participating institutions worked collaboratively to negotiate for a platform that would be suitable for storage of research data and that a country wide licence was negotiated to make it more affordable for all universities in the country, and to also enable archiving of all south african universities’ research data through one platform, see figshare south africavii. each of the four institutions listed above have its own instance and contract with figshare as a proprietary software. the cput research data repository, called esango, went live in early 2018. with the launch of esango, cput worked on strengthening and implementation the repository review workflows and policies. data review workflow according to mayernik et al. (2015) peer review is critical to the scientific communication system. he sees reviewing as both a community responsibility and an opportunity to polish and expand one's understanding of cutting-edge research. adding research data to the publication and peer review queues will put additional strain on the scientific publishing system, however, it will also increase the trustworthiness and value of individual datasets, strengthen findings based on cited datasets, and improve transparency and traceability of data and publications. the data review process at cput involves several stakeholders and is based on coretrustseal (cts) requirements. the review process, see figure 1, includes the digital scholarship librarian, metadata librarian, ethics manager and in cases of postgraduate students, their research supervisor. the process is as follows: https://doi.org/10.29173/iq1111 6/14 lockhart, j, xesi, x & chiware, e (2024) working towards securing and building a trusted institutional research data repository through the coretrustseal process: case of cape peninsula university of technology data repository, iassist quarterly 48(3), pp. 1-14. doi: https://doi.org/10.29173/iq1111 figure 1: esango review process workflow the standard submission form had to be adapted by figshare so that the above workflow could be implemented at cput. it was important to know whether the submission was for graduation purposes, i.e., is it postgraduate students who submits research data related to their thesis (which is submitted through a different process). therefore, the following question has been added to the submission: “is this dataset for graduation purposes?” if yes, then the submitter needs to add the research supervisor’s e-mail address. adding this to the submission form provided the library team the needed information to identify postgraduate student research data submissions and to get the research supervisor’s contact information so that they can be included in the review process. the library team does a basic level curation, but the quality and information in the actual dataset should be reviewed by an expert in the field. preservation a preservation strategy is another important aspect to consider in securing and building trust in a research data repository, and one of the cts requirements. to meet these requirements, cput acquired a cloud-based digital preservation and data management solution. the system called arkivum, a leading digital preservation solution, was selected as a fully managed software as a service (saas), which manages the ingest process, safeguarding of data, preservation, supports over 100 formats, and providing discovery and access. arkivum is a supporter of the digital preservation coalition (dpc) and member of the national data stewardship alliance (ndsa). while figshare is a research data repository, arkivum is a digital archiving and preservation software that archives and preserves data, therefore guaranteeing the longevity of the research data. figshare and arkivum had to develop and adapt their products so that they work seamlessly together. 1. researcher uploads dataset(s) to esango. 2. digital scholarship librarian receives the email notication, does a first level quality check of the metadata and datasets. 3. in cases of postgraduate student research data for graduation purposes, assigns the dataset to the research supervisor. 4. the research supervisor verifies that the student has submitted the correct datasets, check quality and assign back to the digital scholarship librarian. 5. digital scholarship librarian receives the datasets and assign them to the metadata specialist. 6. metadata specialist does a further review on the metadata and assign the datasets to the digital scholarship librarian. 7. digital scholarship librarian assign the datasets to the ethics manager. 8. ethics manager reviews the datasets as per the developed ethics checklist and assign to the digital scholarship librarian. 9. researcher receives an email notifaction that the datasets have been published. 10. datasets gets ingested to arkivum (data preservation platform). https://doi.org/10.29173/iq1111 7/14 lockhart, j, xesi, x & chiware, e (2024) working towards securing and building a trusted institutional research data repository through the coretrustseal process: case of cape peninsula university of technology data repository, iassist quarterly 48(3), pp. 1-14. doi: https://doi.org/10.29173/iq1111 promotion of rdm to promote effective rdm practices, cput libraries offered several training options. first, faculty support teams were trained on the esango data repository as well as on the data management planning (dmp) tool. this was followed by a series of online presentations (during lunchtimes) and workshops for faculty, aimed at highlighting various rdm tools and services available to researchers. one of the key tools discussed was the dmp tool, which is an essential aspect of responsible research data management. rdm training sessions always include a discussion of esango and the dmp tool. the presentations were designed to provide researchers and postgraduate students (who are required to submit a dmp with their research proposal) with practical guidance on how to manage research data effectively and efficiently. during 2023 alone close to 30 training sessions were offered to cput researchers and postgraduate students. the workshops were well received, and the uptake and growth can be seen in the statistics of the dmp tool and the submission of datasets on esango. statistics obtained from the dmp tool suggests that as of the writing of this paper, 1,365 users registered and 1,360 dmp plans were started. through these initiatives, the library has continued to play a vital role in supporting research excellence at cput. securing a trustworthy data repository status general overview trust requirements in research data repositories are growing and as stated by crabtree (2020) “trust in research data repositories is critical as they provide the evidence for past discoveries as well as the input for future discoveries”. the coretrustseal (cts) is an internationally recognized standard for trustworthiness in digital repositories, ensuring that data is being managed in a secure and reliable manner. the cts offers a process for core level certification based on the cts 16+ requirements that reflect certain characteristics of trustworthy data repositories. these cts requirements are a good assessment tool and are helpful in identifying gaps that need to be addressed. l’hours, kleemola and de leeuw (2019) outline the history and background of the cts and the formation of the requirements. as the cts was officially launched in 2017, and it can take several years to get certified, to date there aren’t many publications outlining this certification process in practice. corrado (2019) examines the issue of trust in digital repositories. the author indicates that it is not clear if designated communities are influenced by certificates, however, repositories who meet the requirements may have a better foundation for building trust. the cts requirements are reviewed every few years and adapted as needed. the certification process consists of several submissions and may take a few years to complete. each submission of documentation to cts is reviewed by two referees and the organization applying for certification has to respond. the time from submission to receiving reviewers’ comments is about three months and it takes another three months to resubmit. the process can continue for up to five submissions. so far, the cts has certified over 160 repositories around the world, one of which is on the african continent in south africa. coretrustseal process data sharing is becoming an essential component of scientific research and scholarly publication. this necessitates informed and intentional planning, from early study planning to data and metadata collection, interoperability, deposit in data repositories, and curation (austin et al., 2016). magnuson and thomas (2023) discuss their cts application and list the five valuable lessons they learned: https://doi.org/10.29173/iq1111 8/14 lockhart, j, xesi, x & chiware, e (2024) working towards securing and building a trusted institutional research data repository through the coretrustseal process: case of cape peninsula university of technology data repository, iassist quarterly 48(3), pp. 1-14. doi: https://doi.org/10.29173/iq1111 institutional history, building the cts application, business process model, leveraging documentation and preservation. by applying for the cts certification, cput demonstrated its commitment to promoting responsible rdm. to achieve the certification, cput has implemented a rigorous review process ensuring that its rdm policies and procedures meet the cts requirements. by obtaining the cts certification, cput will be positioned to better serve its researchers and enhance the visibility and impact of their research. the work on the cts application was part of the larger ilifu project undertaken during 2020. the four universities in the region worked on this together, holding regular meetings and supporting each other through the process. however, each university had their own figshare data repository and applied separately for the cts. cput’s goal was to get the research data repository, esango, which is powered by figshare, to meet the core requirement of cts and achieve a secured and trustworthy data repository. the 17 cts trustworthy data repositories requirements (2020-2022) are detailed in table 1. number requirement description r0 repository type, brief description, designated community, level of curation performed, insource/outsource partners. provide context of the repository r1 mission/scope the repository has an explicit mission to provide access to and preserve data in its domain. r2 licenses the repository maintains all applicable licenses covering data access and use and monitors compliance. r3 continuity of access the repository has a continuity plan to ensure ongoing access to and preservation of its holdings. r4 confidentiality/ethics the repository ensures, to the extent possible, that data are created, curated, accessed, and used in compliance with disciplinary and ethical norms. r5 organizational infrastructure the repository has adequate funding and enough qualified staff managed through a clear system of governance to effectively carry out the mission. r6 expert guidance the repository adopts mechanism(s) to secure ongoing expert guidance and feedback (either inhttps://doi.org/10.29173/iq1111 9/14 lockhart, j, xesi, x & chiware, e (2024) working towards securing and building a trusted institutional research data repository through the coretrustseal process: case of cape peninsula university of technology data repository, iassist quarterly 48(3), pp. 1-14. doi: https://doi.org/10.29173/iq1111 house, or external, including scientific guidance, if relevant). r7 data integrity and authenticity the repository guarantees the integrity and authenticity of the data. r8 appraisal the repository accepts data and metadata based on defined criteria to ensure relevance and understandability for data users. r9 documented storage procedures the repository applies documented processes and procedures in managing archival storage of the data. r10 preservation plan the repository assumes responsibility for long-term preservation and manages this function in a planned and documented way. r11 data quality the repository has appropriate expertise to address technical data and metadata quality and ensures that sufficient information is available for end users to make quality-related evaluations. r12 workflows archiving takes place according to defined workflows from ingest to dissemination. r13 data discovery and identification the repository enables users to discover the data and refer to them in a persistent way through proper citation. r14 data reuse the repository enables reuse of the data over time, ensuring that appropriate metadata are available to support the understanding and use of the data. r15 technical infrastructure the repository functions on well-supported operating systems and other core infrastructural software and is using hardware and software technologies appropriate to the services it provides to its designated community. r16 security the technical infrastructure of the repository provides for protection of the facility and its data, products, services, and users. table 1: coretrustseal trustworthy data repositories requirements 2020-2022 going through the submission and revision process streamlined and strengthened the internal workflows and led to better understanding of what it may take to build a trustworthy repository for cput. most of the requirements for cts focussed on the cput’s internal environment regarding https://doi.org/10.29173/iq1111 10/14 lockhart, j, xesi, x & chiware, e (2024) working towards securing and building a trusted institutional research data repository through the coretrustseal process: case of cape peninsula university of technology data repository, iassist quarterly 48(3), pp. 1-14. doi: https://doi.org/10.29173/iq1111 policy, expertise, workflows, and preservation strategies that was in place. however, there were instances that documentation from figshare was required. cput’s application process is currently still in progress. the scoring was done between 0 to 4 and described as seen in table 2. score description 0 not applicable 1 not considered this yet 2 the repository has a theoretical concept 3 the repository is in the implementation phase 4 the guideline has been fully implemented in the repository table 2: coretrustseal scoring categories challenges as much as libraries can implement new services, platforms and tools, the key challenge is often the advocacy and uptake of the services by the researchers and the university community at large. one of the important drivers in university environments is the policy landscape. policies are key, but also take longer to develop, and review periods are often every three years and done consultatively as per the policy on policy development. once a policy is in place, advocacy is a key driver and this involves several different strategies which could include, being on the agenda of key university committee structures, regular communication via university communication channels, developing training by the library, sometimes as part other key departments training, e.g., the directorate research development and centre for postgraduate studies. since it was a new aspect to deal with the cts process, it has been challenging. it was difficult to incorporate some of the discipline-specific repositories’ requirements into cput’s esango research data repository since it is a generalist repository. however, working through the submission and revision process, helped the team improve and update the university’s research data repository as well as the data deposit and preservation processes. while it was a steep learning curve, much was learned about good rdm practices. a review process is essential to ensure the quality of research data. since library staff may not have subject expertise, libraries are not able to provide more than a basic level curation of the datasets deposited. this highlights the importance of including discipline-specific specialists as part of the review process and this may be challenging to do. to overcome this challenge, cput decided to include the postgraduate students’ research supervisors as part of the data review process. there are still hurdles in bringing students’ research data and procedures into the review process and the library https://doi.org/10.29173/iq1111 11/14 lockhart, j, xesi, x & chiware, e (2024) working towards securing and building a trusted institutional research data repository through the coretrustseal process: case of cape peninsula university of technology data repository, iassist quarterly 48(3), pp. 1-14. doi: https://doi.org/10.29173/iq1111 has put in place webinars that equip students with the necessary skills on how to manage their research data throughout the research lifecycle. lessons learned embarking on this process may look daunting. however, establishing a small working group that includes two or three people, makes achieving this goal doable with limited additional workload for each person in the team. a systematic approach worked well for our team. we scheduled one-hour sessions over several days to go through each requirement and to write how we plan to meet each requirement. this approach was followed when addressing the reviewers’ comments and making the necessary changes and/or notes to update workflows and guidelines. additionally, it was helpful to look at other cts approved repository documentation, to get a sense of what is required. if the repository uses proprietary software, it is helpful to ask them for support and to provide documentation and policies for aspect they are responsible for. future trends future trends could include more adoption by university libraries and research entities of certified research data repositories as a standard and thus ensuring improved research data management practices. this will lead to higher quality research data as per the fair principles. this will also lead to the requirement of increased preservation demands within university libraries; therefore, development of additional skills and preservation experts may be needed. an increase in the requirements of research dataset submissions from publishers and funders are expected, especially as governments put in place structures, white papers and statements regarding research funding from public funds. conclusion rdm practices have evolved over the last decade and more publishers and funders require datasets as part of the publishing process and many governments have put in place strategies to ensure a good science, technology, and innovation landscape. to ensure best practices for rdm services and tools, it will be of excellent value to start with a self-assessment to measure a research data repository against the cts requirements. this will assist with identifying gaps within the rdm environment at the university or research entity and lead to enhancements and standardization. a further step would be to submit the application to certify the research data repository for cts approval. references austin, c. c., brown, s., fong, n., humphrey, c., leahey, a. & webster, p. 2016. research data repositories: review of current features, gap analysis, and recommendations for minimum requirements. iassist quarterly, 39, 24-24. doi: https://doi.org/10.29173/iq904 bak, g. 2016. trusted by whom? tdrs, standards culture and the nature of trust. archival science, 16, 373-402. doi https://doi.org/10.1007/s10502-015-9257-1 https://doi.org/10.29173/iq1111 https://doi.org/10.29173/iq904 https://doi.org/10.1007/s10502-015-9257-1 12/14 lockhart, j, xesi, x & chiware, e (2024) working towards securing and building a trusted institutional research data repository through the coretrustseal process: case of cape peninsula university of technology data repository, iassist quarterly 48(3), pp. 1-14. doi: https://doi.org/10.29173/iq1111 cape peninsula university of technology. 2012. research, technology and innovation (rti) 10-year blueprint [online]. available: https://www.cput.ac.za/storage/research/research_directorate/rti_blueprint.pdf [accessed 10 january 2024]. cape peninsula university of technology. 2021a. open access (oa) policy [online]. available: https://www.cput.ac.za/storage/library/pdf/research-support/cput-open-access-policy.pdf [accessed 11 january 2024]. cape peninsula university of technology. 2021b. research data management (rdm) policy [online]. available: https://www.cput.ac.za/storage/library/pdf/research-support/research-datamanagement-policy.pdf [accessed 09 may 2023]. chiware, e. and mathe, z. 2015. academic libraries' role in research data management services: a south african perspective. south african journal of libraries and information science, 81, 1-10. doi: https://doi.org/10.7553/81-2-1563 chiware, e. and becker, d.a. 2018. research data management services in southern africa: a readiness survey of academic and research libraries. african journal of library, archives, and information science, 28(1), 1-16. corrado, e.m. 2019. repositories, trust, and the coretrustseal. technical services quarterly, 36(1), 61-72. doi: https://doi.org/10.1080/07317131.2018.1532055 crabtree, j. d. 2020. evidence for trusted digital repository reviews: an analysis of perspectives. phd thesis. university of north carolina. available at: https://cdr.lib.unc.edu/concern/dissertations/2n49t9931 (accessed: 17 january 2024) curty, r.g. 2016. factors influencing research data reuse in the social sciences: an exploratory study. internal journal of digital curation, 11(1), 96-118. doi: https://doi.org/10.2218/ijdc.v11i1.401 department of science & technology. 2019. white paper on science, technology and innovation: science, technology and innovation enabling inclusive and sustainable south african development in a changing world. [online]. available: https://www.dst.gov.za/images/2019/white_paper_web_copyv1.pdf [accessed 11 january 2024]. faundeen, j. 2017. developing criteria to establish trusted digital repositories. data science journal, 16, 22-22. doi: https://doi.org/10.5334/dsj-2017-022 hutchison, v. b., norkin, t., langseth, m. l., ignizio, d. a., zolly, l. s., mcclees-funinan, r. & liford, a. 2021. leveraging existing technology: developing a trusted digital repository for the us geological survey. international journal of digital curation, 16, 23-23. doi: https://doi.org/10.2218/ijdc.v16i1.741 https://doi.org/10.29173/iq1111 https://doi.org/10.7553/81-2-1563 https://doi.org/10.1080/07317131.2018.1532055 https://cdr.lib.unc.edu/concern/dissertations/2n49t9931 https://doi.org/10.5334/dsj-2017-022 https://doi.org/10.2218/ijdc.v16i1.741 13/14 lockhart, j, xesi, x & chiware, e (2024) working towards securing and building a trusted institutional research data repository through the coretrustseal process: case of cape peninsula university of technology data repository, iassist quarterly 48(3), pp. 1-14. doi: https://doi.org/10.29173/iq1111 johnston, w. 2012. digital preservation initiatives in ontario: trusted digital repositories and research data repositories. partnership: the canadian journal of library and information practice and research, 7(2). doi: https://doi.org/10.21083/partnership.v7i2.2014 l’hours, h., kleemola, m., de leeuw, l. 2019. coretrustseal: from academic collaboration to sustainable services. iassist quarterly, 43(1). doi: https://doi.org/10.29173/iq936 magnuson, d. l., thomas, w. l. 2023. expanding our perspective: building a sustainable metadata culture. iassist quarterly, 47(2). doi: https://doi.org/10.29173/iq1046 mayernik, m. s., callaghan, s., leigh, r., tedds, j. & worley, s. 2015. peer review of datasets: when, why, and how. bulletin of the american meteorological society, 96, 191-201. doi: https://doi.org/10.1175/bams-d-13-00083.1 mehnert, a. j., janke, a., gruwel, m., goscinski, w. j., close, t., taylor, d., narayanan, a., vidalis, g., galloway, g. & treloar, a. 2019. putting the trust into trusted data repositories: a federated solution for the australian national imaging facility. international journal of digital curation. 14(1) doi: https://doi.org/10.2218/ijdc.v14i1.594 national research foundation. 2015. statement on open access to research publications from the national research foundation (nrf) funded research. [online]. available: https://chelsa.ac.za/wpcontent/uploads/2018/12/nrf_open_access_-statement_19jan2015_v6.pdf [accessed 11 jan 2024]. ntja, b. 2022. enhancing the role of the libraries in south african higher education institutions through research data management: a case study of cape peninsula university of technology, faculty of humanities, library and information studies centre (lisc). http://hdl.handle.net/11427/37692 swauger, s., vision, t.j. 2015. what factors influence where researchers deposit their data? a survey of researcher submissions to data repositories. international journal of digital curation, 10(1), 68-81. doi https://doi.org/10.2218/ijdc.v10i1.289 tripathi, m., shukla, a. & sonkar, s. k. 2017. research data management practices in university libraries: a study. desidoc journal of library & information technology, 37(6), 417-424. doi https://doi.org/10.14429/djlit.37.6.11336 wilkinson, m. d., dumontier, m., aalbersberg, i. et al. 2016. the fair guiding principles for scientific data management and stewardship. scientific data, 3(160018) doi https://doi.org/10.1038/sdata.2016.18 yoon, a. 2014. end users’ trust in data repositories: definition and influences on trust development. archival science, 14, 17-34. doi https://doi.org/10.1007/s10502-013-9207-8 https://doi.org/10.29173/iq1111 https://doi.org/10.21083/partnership.v7i2.2014 https://doi.org/10.29173/iq936 https://doi.org/10.29173/iq1046 https://doi.org/10.1175/bams-d-13-00083.1 https://doi.org/10.2218/ijdc.v14i1.594 http://hdl.handle.net/11427/37692 https://doi.org/10.2218/ijdc.v10i1.289 https://doi.org/10.14429/djlit.37.6.11336 https://doi.org/10.1038/sdata.2016.18 https://doi.org/10.1007/s10502-013-9207-8 14/14 lockhart, j, xesi, x & chiware, e (2024) working towards securing and building a trusted institutional research data repository through the coretrustseal process: case of cape peninsula university of technology data repository, iassist quarterly 48(3), pp. 1-14. doi: https://doi.org/10.29173/iq1111 endnotes i janine lockhart is the library manager: scholarly communication and digital scholarship at cput libraries lockhartj@cput.ac.za. ii xabiso xesi is the digital scholarship librarian at cput libraries xesix@cput.ac.za. iii prof elisha r chiware is the director at cput libraries chiwaree@cput.ac.za. iv https://www.re3data.org/ v https://www.coretrustseal.org/ vi https://www.ilifu.ac.za/ vii https://southafrica.figshare.com/ https://doi.org/10.29173/iq1111 mailto:lockhartj@cput.ac.za mailto:xesix@cput.ac.za mailto:chiwaree@cput.ac.za https://www.re3data.org/ https://www.coretrustseal.org/ https://www.ilifu.ac.za/ https://southafrica.figshare.com/ vol274.indd iassist quarterly winter 2003 9 by anne sofie fink kjeldgaard, søren priisholm, birgitte grønlund jensen * preservation of knowledgedata processing in the danish data archives abstract high quality secondary analysis of sample surveys depends on the quality of the primary data sets and their preservation. the researchers performing the secondary analysis must be able to access as much information about the data sets as possible. in the danish data archives (dda) great effort is taken to preserve the data sets in a way that meets the needs of the secondary researcher. for this reason data processing is a core operation in the dda and great importance is attached to producing reliable and useful documentation of the preserved data files. data processing is a core activity in the danish data archives1 (dda). it is the performance of data processing that makes dda unique in comparison with alternative data preservation efforts in the danish research world. despite the importance attached to the activity of data processing, it is an activity that is invisible to outsiders. in this article we set out to discuss the advantages of data processing for the depositor of the data set, for the end-users when performing secondary analysis on the data set and for the data archive. in this light we will then describe our preservation strategy in detail. finally, we will guide readers through the data processing process step by step and point to future development. the advantages of data processing data processing has advantages for the depositors, the end-users and the data archive. these advantages can be traced back to the effort of collecting and integrating all available information about a data set. the physical product of the data processing process in the dda is a data documentation publication (ddp) consisting of a study description and a codebook with frequency tables for all variables in the data set. the study description holds all information about the creation of the data set and facts about its preservation in the dda, as well as restrictions for access to the data set. the codebook consists of all available documentation about the data file and a copy of the original questionnaire. data deposition although it ought to be a straightforward task to add together the information for the study description and the codebook, most often it is a timeconsuming job. it becomes especially difficult if it has been a while since the actual survey was carried out. in order to make it as easy as possible to gather information about the study, the dda urges researchers to deposit their data as early as possible in the research process. after a data set has gone through data processing a data documentation publication (ddp) is created. the depositor gets the message that her study has a ddp, i.e., a study description and a codebook. she can then be sure that her data set is preserved for the future and she will have no need for storing the data elsewhere. data location as soon as the data material has a study description it becomes searchable in the dda̓ s data catalogue on the internet. a future challenge in this respect is to allow users not only to search the study description, but also to search the whole ddp. at the moment, the dda is taking part in the madiera project2, which, among other things, will offer users this opportunity. secondary analysis as regards data analysis, data processing offers essential advantages for end-users. first and foremost, it becomes straightforward to get access to all information about the origin of the data set and its contents because all information is at hand in the ddp. however the ddp is not just a collection of information provided by the depositor. during the data processing process, several additions, standardisations and checks are made. for example, if divergence between the data and the questionnaire is found, the person in charge of the data processing will make a comment about this in the codebook, thereby eliminating the need for future users of this particular data set to spend time finding “wild codes” themselves. data archiving internally, data processing has the advantage that users of the data sets seldom need guidance upon having received a data set. as a consequence, the time spent on providing user services is reduced in spite of increasing use of our data sets over the years. 10 iassist quarterly winter 2003 preservation strategy if the dda is to promise depositors long-term preservation of their data sets, data and documentation must be preserved in a simple format in order to ensure that all content can be read and analysed in the future. obviously this promise must be made without knowledge of the composition of future computers and their statistical software. therefore, any kind of software and hardware dependence has to be removed and the data set has to be documented as completely as possible. the majority of the data sets in the dda̓ s collection are one-off cross-section sample surveys. in these cases there is only one data file and one documentation file to be matched. although in practice many data sets relate to each other, e.g., surveys replicated in time, these are handled as separate data sets for the time being. the dda is part of the metadater project3, which is working on a solution to reflect relationships between studies, by both study naming conventions and methods of storage. in the dda, the data material is said to consist of two parts: a data file and a documentation file. this means that the dda has one job concerned with preservation of the data file and another job concerned with preservation of the documentation file. the data file is a data matrix that constitutes the substance of the quantitative data material. this data matrix contains codes for respondents/cases, questions, categories for answers, and additional codes the researcher might have added to the data set. the documentation file is the information that describes the origin and contents of the data. in other words, documentation is the information users need to make sense of the data matrix. an example is that in the data matrix you can see that respondents in column no. 47 have answered either 1 or 2. in the documentation you can see column 47 contains information about the respondentʼs sex (male = 1 or female = 2) as answers to question no. 30: “are you male or female?” in the questionnaire. all this documentation is gathered together in the codebook. for dda, the data preservation strategy has two strands. both the data matrix and the documentation have to be preserved technically and physically, and the semantic information/documentation in the study description and the codebook has to be preserved in a way that ensures that the data stay ʻunderstandable ̓to future users. technical and physical preservation at a very early stage, the dda and other data archives decided to use the system-independent format osiris iii (symbols 0-9) as the technical preservation format. osiris iii is an extremely simple format that was developed to describe existing data materials useful for sample surveys. as soon as dda receives data material, the data are transferred to a dda server. in most cases it is possible to make a conversion of the data to osiris format at once, thereby creating a long-term preservation file at this stage. data are processed in order of priority according to novelty, demand from users, relation to other studies, etc. from the osiris format, data can easily be converted to up-to-date formats like the current versions of sas, spss, stata, etc. as stated above, it is an advantage to users to get data material that has been processed; otherwise they must cope with the original data file themselves. data are physically stored on a server with a daily backup routine. although technical and physical preservation is just as important to data preservation as semantic preservation, it is most often much more straightforward, depending on the complexity of the material, of course. semantic preservation the semantic preservation is carried out to ensure that all available information about the data set is collected and merged with the data matrix. as mentioned above, the documentation of a data set has two main components: a study description describing information about the depositor, the production of the survey and access conditions, and a codebook with identifying keys to translate the codes in the data matrix into something meaningful. the ddp consists of the study description and the codebook with frequency tables. the study description the study description consists of a large amount of background information that the user must take into account when drawing conclusions on her analysis of the data set. the study description has information about the following subjects: it has always been an objective to prepare a study description upon reception of a data set. however, in practice it often happens that the study description is not prepared until the data material is processed. the unfortunate consequence of this is that there is a delay before the study description appears in dda̓ s search catalogue. dda recently introduced a new programme that was developed in-house for the preparation and preservation of study descriptions. from now on, study descriptions will be placed in a database structure, which allows us greater flexibility than we had with the fixed text format we used before. the new programme offers a template for structuring information inputs. this allows everyone in the archive to take part in the preparation of the study description and to gather the information step by step. the programme has benefits to the end-users too, since iassist quarterly winter 2003 11 table 1: subjects in the study description general information: title and year of study, status of the study, classification of the study in cluster(s), relevant keywords for the study, language employed in the present study description identifications and references: bibliographic reference, local archive where the study is stored, archive where the study was originally stored, depositor (donor), date of deposit, primary investigator (research organization), data collector, research initiator, funding agency analysis conditions: research topic (abstract), kind of data, units of observation, number of units (cases), dimensions of data set, completeness of study stored, time period covered, time dimensions, definition of total universe (universe sampled), sampling procedures, geographical area covered, dates of data collection, method of data collection, type of research instrument, actions to minimize losses, data gathering staff, characteristics of data collection situation noted, weighting re-analysis conditions: current data representation, applicable analysis packages, language(s) of written material, control operations performed by primary investigator, control operations performed by archive, accessibility, access directing authority references to the relevant publications/results/studies: publications/reports from primary investigator variables included: basic characteristics, residence, household characteristics, characteristics of parental family/household, occupation, education information in the database will be become simultaneously searchable on the web. the codebook the dda codebook is the main product of the data processing process with its translation of the data file into human understandable information. it consists of a description of variables containing information about the positions of every single variable in the data file, the exact text from the questionnaire and a number of other essential pieces of information. it also explains any discrepancies between the data and the documentation that were found during processing. an osiris codebook is a text file in punch card format, which means that each line must not exceed 80 characters. it is, however, possible to insert a continuation card (kcard). the first 10 columns are assigned as follows: column 1: card type; column 2-5: variables number; column 6-9: reference number; and column 10: typically blank. various types of cards for different kinds of information are available to the data processor. it should be mentioned that osiris is a preservation format – not a presentation format. the dda has developed a separate presentation format for codebook information. data processing step by step the data processing process is described below and illustrated in the accompanying flowchart. 1. allocation of study number to the data set when the dda and the depositor have made a deposition arrangement, the data material gets allocated a unique study number in dda̓ s journal. this number will tie together all parts of the data material stored in different places on dda servers. 2. reception of the data set (data0.org) the dda has no required format for data sets. the archive has a wide selection of software programmes to convert the original data to an osiris temporary file. this file is preserved until data processing has been completed. a check is performed to ensure that the essential documentation (questionnaire, data files and documentation files) has been transferred to the data archive. the original questionnaire is scanned and stored as a pdf file. 3. control of data (data1.org, createdata.sas) a check is performed to test if the number of respondents and variables are in accordance with the documentation from the researcher. frequency tables for all variables are printed, in order to perform a more thorough control of the material. 4. data recoding and documentation recoding and modification is performed in a sas script. the first step is to insert two standard variables: the dda study number (v1) and a sequence number (v2). the value labels are then modified so they do not exceed 24 characters and have expressive names. variables are recoded for filters and formats, e.g., length and numbers of decimals. values for missing data are recoded into the dda standard format. three categories are used: • not answered (9,99 ... etc.) – the respondent has not answered the question • inappropriate (10,100 ...etc) – the respondent should not answer (due to a filtering variable) • no participation (11, 101 ...etc.) – the question 12 iassist quarterly winter 2003 was not presented for the respondent (this typically occurs when surveys with different variables and/or respondents are merged together) by the use of standard missing codes, the work for the endusers is made easier. this is very important when making comparisons between several data sets. in parallel to the recoding, a codebook is edited in an editor (kedit). the job here is to write the exact text from the questionnaire from top to bottom into the documentation file. for all variables, the codebook contains frequencies in figures as well as in percentages. finally, the variables are renamed to a continuous variable row. the data processor then makes a profound check of the data to ensure that the data definition is in accordance with the original data and the documentation. dda does not recode data if there is a discrepancy, but makes a note in the codebook about this. such a note can save researchers or students using the data set hard work. the questionnaire in pdf format is converted into a text file. 5. data processing of data and codebook by running a script on the data and codebook and by merging the codebook and data file, the data documentation publication (ddp) is produced. beforehand a preface and a study description are inserted in the codebook. the final files are data.osi, which is a simple data matrix, and ddp.las, a codebook file that contains print codes for page shifts and page numbers. 6. preservation and user services the processed material is stored in two parts – a data file and a documentation file. the two files and the sas script file are preserved on the server. 7. proofreading internally and approval externally when the ddp is produced, there will be a proofreading of the produced material by a staff member who has not been involved in the data processing process so far. if the proofreader finds mistakes, the material is given back to the data processor, who corrects it accordingly. afterwards, the proofreader goes through the material a second time to make sure that mistakes are corrected. the material is now send out to the depositor for her approval of the processed data set. 8. user services when a user requests the data material for secondary analysis, the data and the documentation files are merged by use of a program (osi2spss) that creates a spss file. if the user requests a format other than spss, further conversion must be made. when uploading to the internet, the two files are merged by the use of a program for xml and html files. the flowchart below shows the data processing step by step. • all steps marked with italics indicates that the data processing step is done manually. • all filenames are written in bold characters. • all scripts/programs are put in quotation marks. future development in data processing from the flowchart (table 2) it should be obvious that data processing is a very time-consuming activity for the dda. at the moment the average amount of time spent on data processing is about 150 hours per data material. however there is great variation, as data processing of some studies only takes about 50 hours and, obviously, some take much more. in the data processing process it is codebook editing and recoding and renaming of variables that are the great time consumers, whereas running scripts is something that is quickly done. in order to reduce the time spent on data processing, it is planned to develop a new programme for data processing. the objectives for this development will be to minimise the time spent on data processing and to implement a higher degree of standardisation of the process. for example, at the moment, 10 scripts are used in every data processing process. hopefully this can be cut down to a single one. similarly, the number of working files may be cut down from 16 to one. with the introduction of the new programme, we can look forward to a substantial reduction of the time spent on data processing. we will then be able to process more materials than we can today, and thereby be able to provide our users with more high quality materials faster. * authors: researcher anne sofie fink kjeldgaard, data archivist søren priisholm and information science officer birgitte grønlund jensen, the danish data archives. contact: anne sofie fink kjeldgaard, asf@dda.dk. iassist quarterly winter 2003 13 table 2: data processing of a study in the danish data archives1 data documentation data0.org* (original data from the depositor, e.g., spss or sas files)* formal control (check of machine readability, number of variables and respondents are right according to documentation) data1.org* (checks data in *.por or *.sav format) dandata.sas (creates new sas script file using the editor kedit) use dandata.sas to read data1.org into sas = creates frequency tables (sasdata.sd2) print sasdata.sd2 or the frequency tables. at this point it can be estimated how many resources to be put into data processing of the study, e.g., by checking the quality of data documentation. changing labeland variable names and recoding of variables according to dda’s guidelines (this includes re-coding for filters). this is a “trial and error”-process. saving and running the script dandatax.sas*) using sas – over and over again. output file is sasdatax.sd2 running “sasosir” script on sasdatax.sd2 converted files (data.osi and dict.osi) data.osi* = datamatrix dict.osi = dictionary (list of labelnames or tcards) manual check of dict.osi (all variables are recoded and label names are max. 24 characters long.) cdbk.txt (create new codebook file – using an editor such as kedit). the questionnaire in pdf format is converted to a machine readable text file. comparison with dandatax.sas. run “kodebyg” script on cdbk.txt (this script moves the codebook text in accordance with the osiris format) cdbk.osi (converted codebook-file) run “merget” script on dict.osi,cdbk.osi dicb.osi (codebook with t-cards/labels) run “sasmarg” script on dicb.osi, sasdatax. sd2 dicbm.osi* (codebook with correct margins) copy/paste writing of fortale.osi (a preface/short introduction to the study) into dicbm.osi 14 iassist quarterly winter 2003 data documentation .. from last page... from last page.. run “dicbdok” on dicbm.osi dicbm.dda* (codebook with ‘guide to the codebook format’) writing of study-description s1234gb.sd run “ddasdf” script on s1234gb.sd s1234gb.dda (change item-number and code to standard-text.) copy/paste s1234gb.dda into dicbm.dda dicbm.dda run “ddalist” on dicbm.dda dicbm.las (or ddp.las – the print file without any internal codes, but with printing information such as ‘new page’, ‘page number’, etc.) print dicbm.las dda preserves 1. data.osi 2. 3. doku.osi (=dicbm.dda) 4. dandata.sas (consists of all used sas-statements = documentation of the process) 5. data1.org/data0.org iassist quarterly winter 2003 15 footnotes 1 the dda was established in 1973 as a national data bank for quantitative research carried out primarily in the social sciences but also in medical science and history in denmark. in 1993 the dda became an independent unit in the danish state archives. at present the archive has 13 full-time employees. the dda collects, preserves and disseminates machine readable research data. 2 read more about the madiera project here: www. madiera.org. 3 read more about the metadater project here: www. metadater.org. 4 rasmussen, karsten boye, 1996, oparbejdning med sas i dansk data arkiv (data processing using sas in the danish data archives), version march 1996, dda, odense, p. 21. http://www.madiera.org http://www.madiera.org www.metadater.org www.metadater.org boosting data findability: the role of ai-enhanced keywords 1/12 jamwal, kokila (2024) boosting data findability: the role of ai-enhanced keywords, iassist quarterly 48(4), pp. 1-12. doi: https://doi.org/10.29173/iq1127 the creative commons-attribution-noncommercial license 4.0 international applies to all works published by iassist quarterly. authors will retain copyright of the work and full publishing rights. boosting data findability: the role of ai-enhanced keywords kokila jamwal1 abstract in today’s data-driven world, finding relevant data in a vast expanse of information is increasingly important. researchers have been exploring various methods to improve the findability, accessibility, interoperability, and reusability of data, for example, by using controlled vocabularies to enhance data findability. although the use of controlled vocabularies is growing, challenges remain for findability when users provide their own keywords, known as user-defined keywords or do not provide keywords at all. finding data in data archives based on metadata fields with user-defined or missing keywords is challenging, or even impossible. here, we show the use of artificial intelligence (ai) techniques from the subfield of deep learning to automate the assignment of keywords using controlled vocabulary, leading to improved data findability. the main results demonstrate that ai automation performs well on the test set. in addition, we comapre our deep learning model against large language model (llm) on the task of automated topic assignment. automated topic assignments will reduce the time and effort required for data curation, enhancing data findability and usability for data producers and consumers. the application of ai to automate metadata assignment offers practical solutions for improving data findability and reusability, not only in research data archives but across various datadriven domains. overall, this approach highlights the potential of ai in addressing data findability challenges, paving the way for more efficient and effective data discovery and utilization in the era of big data and information abundance. keywords fair, user-defined keywords, text classification, ai, findability, controlled vocabulary introduction research data generation has seen an upward trend. researchers generate research data to validate their findings and hence actively contribute to this upward trend (steiner, 2023). in addition to this, new data collection methods like apis and web scraping have added to the exponential growth in the volume of daily-generated research data. due to this, managing research data has become more challenging. metadata plays a crucial role in managing research data. research data with rich metadata is easier to manage than research data with limited or no metadata. supporting the idea of rich metadata, fair principles have caught the research community's attention. the fair principle focuses on making data findable, accessible, interoperable, and reusable. focusing on the findability aspect of fair principles, we investigate gesis search. gesis search2 is a search platform from gesis (gesis – leibniz institute for the social sciences, based in germany, is an infrastructure institute for the social sciences). the gesis search platform allows researchers to find https://doi.org/10.29173/iq1127 https://creativecommons.org/licenses/by-nc/4.0/ 2/12 jamwal, kokila (2024) boosting data findability: the role of ai-enhanced keywords, iassist quarterly 48(4), pp. 1-12. doi: https://doi.org/10.29173/iq1127 surveys and social science research data. gesis search provides necessary metadata fields along with research data in order to make research data findable for the researchers. in our paper, we focus on the 'topics' metadata field. the 'topics' metadata field generally takes in values from a controlled list of topics, known as controlled vocabulary (cv). controlled vocabularies improve data findability 3,4,5,6,7. however, some studies in gesis search have either user-defined or no keywords, which leads to poor findability, as shown in figure 1. figure 1: limited findability of research data due to missing and user-defined values in 'topics' metadata field in figure 1, the data provider submits a study into the gesis archive. however, the study might have missing or user-defined topics. when data users search for studies based on keywords in gesis archive, the search results may be negative, leading to limited findability. there is a need to overcome this gap so that the findability of the studies can be further improved. for example, in one of the studies from gesis search, data depositors used ‘migration cost’ as a user-defined topic, which is not easily findable. however, replacing it with ‘migration’ from cv would make it more findable. figure 2 shows the result of searching for a user-defined topic versus a cv topic on gesis search. there are no results for the user-defined topic, but the cv topics group similar studies together, making such studies more findable. artificial intelligence is beneficial in labeling data automatically. we specifically use the deep learning subfield of ai, which focuses on learning from complex data to make predictions. leveraging ai techniques, we use an automatic topic classification model to label missing or user-defined keywords in the 'topics' metadata field. these labels take values from a set of cv sources. in particular, the ‘topics’ metadata field can have more than one label, making it a multi-label classification problem. as an input feature, we utilize text from the ‘abstract’ metadata field and encode this text with a sentence transformer. the encoded abstract is then passed through a multi-label classification model, which outputs the labels from the cv sources. we will discuss the definitions and approach in detail in the following sections. our work helps to improve data findability and reusability through ai techniques. in summary, our contributions are as follows: 1) we investigate the presence of missing or user-defined keywords for the ‘topics’ metadata field for gesis research data. https://doi.org/10.29173/iq1127 3/12 jamwal, kokila (2024) boosting data findability: the role of ai-enhanced keywords, iassist quarterly 48(4), pp. 1-12. doi: https://doi.org/10.29173/iq1127 2) we propose a deep learning-based multi-label classification model1 for automatic labeling of the ‘topics’ metadata field with values from cv. we evaluate the performance of our model based on various metrics such as precision (micro), recall (micro), f1 score (micro) and hamming loss, to measure the fraction of incorrectly predicted labels. these metrics will be discussed in detail in ‘experimental setup’ section. 3) we compare our deep learning-based classification model with a large language model (llm), chatgpt 3.5. our evaluation results on multiple subsets of test data demonstrate that our model trained for topic classification performs better than the llm. figure 2: gesis search results based on filter ‘topic’ for ‘migration cost’ vs ‘migration’. ©gesis we structure the rest of our paper as follows: first, we formalize our problem statement. then, we present our proposed multi-label topic classification model in the ‘approach’ section. we then describe our ‘experimental setup’ section, including information on datasets, parameter settings and evaluation metrics. ‘evaluation results’ section assesses our approach using various metrics. in ‘use case’ section, we compare our model with llm before wrapping up in ‘conclusion’ section. definitions & problem formulation we consider our topic assignment problem as a multi-label classification problem. table 1 provides an example of a multi-label dataset. table 1 shows a multi-label dataset for movie reviews. each movie review is an input instance of the dataset, and the corresponding set of labels in column ‘label’ is the output label. each output label can contain more than one value. 1 https://github.com/kokila134/multi-label-classification https://doi.org/10.29173/iq1127 https://github.com/kokila134/multi-label-classification 4/12 jamwal, kokila (2024) boosting data findability: the role of ai-enhanced keywords, iassist quarterly 48(4), pp. 1-12. doi: https://doi.org/10.29173/iq1127 movie review (x) label (y) "the movie is great mix of comedy and romance.", ["comedy", "romance"] "i loved great action sequence in the climax of beautiful love story" ["action", " romance"] table 1: example of multi-label dataset each study in gesis search contains metadata fields to describe the study. figure 3 shows a study and its associated metadata fields in gesis search. in this paper, we are interested in assigning cvs to the ‘topics’ metadata field. for this, we only consider the ‘abstract’ metadata field as an input feature as it generally contains the most textual information about the study. let a dataset 𝐷 = {(𝑥1, 𝑦1), (𝑥2, 𝑦2), . . ., (𝑥𝑛, 𝑦𝑛)}, where 𝑥𝑖 ∈ 𝑋 is an input instance and 𝑦𝑖 ⊆ 𝑌 is the set of labels associated with 𝑥𝑖. figure 3: a study and its metadata fields in gesis search. ©gesis input instance (𝑥𝑖): each input instance is the set of input features passed to a model for learning the patterns. an input instance is an entry in the dataset. output labels (𝑦𝑖): output labels, for each input instance, is a list of labels associated with the provided input instance. input feature: an input feature is an attribute of a specific input instance. in a dataset, especially one in a tablular format, an input feature refers to the value in a specific column for a given input instance (row). the goal is to learn a function 𝑓(𝑥) that maps each input instance 𝑥 to its corresponding set of labels 𝑦. multi-label classification (𝑓(𝑥)): multi-label classification is a text classification problem in which each instance may belong to several predefined categories or classes simultaneously. in our case, these categories are values from cv sources. https://doi.org/10.29173/iq1127 5/12 jamwal, kokila (2024) boosting data findability: the role of ai-enhanced keywords, iassist quarterly 48(4), pp. 1-12. doi: https://doi.org/10.29173/iq1127 figure 4 shows a snapshot from our initial dataset, which contains two columns: abstract and label. each row depicts a study in gesis search. an input instance and a list of output labels correspond to each study. figure 4: snapshot of our initial dataset approach in this section, we discuss our approach in more detail. we first discuss our data preprocessing pipeline in section ‘data preprocessing’. next, we describe our embedding generation technique to embed the input instances. finally, we discuss the multi-label classification model which outputs the class labels from cvs. the overall pipeline for our approach is illustrated in figure 5. figure 5: overall pipeline for our approach data preprocessing gesis search is a platform that allows researchers to find information about social science research data, publications on research data, and open-access publications. the studies in gesis search include metadata about the study or research data. each study contains metadata fields in descriptive, methodological, and bibliographic metadata categories. in descriptive metadata, the studies contain ‘abstract’ and ‘topics’ metadata fields among others. the ‘abstract’ metadata field allows the researchers to summarize the study they are publishing. the ‘topics’ metadata field provides keywords from a set of controlled vocabulary (cv), which provide a better picture of the study. for the ‘topics’ metadata field, the research users generally select more than one value from the cvs. however, some studies either have user-defined values in the ‘topics’ metadata field or have no value. considering that, we filter studies based on the number of cv topics. we select the studies containing more than one cv topics for training purposes. later, we remove the duplicate studies at the end of the data preprocessing pipeline. a controlled vocabulary is a controlled list of values from which the relevant values are selected for a specific metadata field. multiple sources of controlled vocabularies have been used in social science research studies at gesis, such as the stw thesaurus for economics and thesoz thesaurus. in our paper, we select cessda topic classification, dprex, iso 3166-1 country codes, kategorienschema https://doi.org/10.29173/iq1127 6/12 jamwal, kokila (2024) boosting data findability: the role of ai-enhanced keywords, iassist quarterly 48(4), pp. 1-12. doi: https://doi.org/10.29173/iq1127 wahlstudien, stw thesaurus for economics, thesoz thesaurus, and unesco thesaurus as the controlled vocabulary sources8. the values from these sources are combined to create a unified list. the duplicated values are removed and converted into lowercase for mapping to cv topics as shown in figure 6. figure 6: data preprocessing pipeline embedding generation deep learning models do not understand raw textual data, so we need to convert the textual data into numerical form. embeddings are numerical representations of data that captures semantics of the data, making it possible to understand real-world data effectively. in this paper, data in the ‘abstract’ metadata field is in text format. we utilize the textual information from the ‘abstract’ metadata field and capture its semantic meaning through our embedding generation component. to achieve this, we apply a sentence transformer (reimers and gurevych, 2019) to transform text in the ‘abstract’ metadata field into embeddings. since most of the text in the ‘abstract’ metadata field across studies is multilingual (mostly in german and english), we utilize a multilingual sentence transformer called ‘distiluse-base-multilingual-cased-v1’ to generate the embeddings for ‘abstract’. these embeddings are vector representations that encapsulate the semantic essence of original sentences. figure 7 depicts an example of embedding generation. figure 7: sentence transformer based text to embedding generation multi-label classification model the multi-label classification model is the last component of our approach. it takes the embeddings generated for each study as input and passes them through a series of fully connected layers (fc) to https://doi.org/10.29173/iq1127 https://huggingface.co/sentence-transformers/distiluse-base-multilingual-cased-v1#:~:text=sentence-transformers%2fdistiluse-base-multilingual-cased-v1%20this%20is%20a%20sentence-transformers%20model%3a%20it%20maps,used%20for%20tasks%20like%20clustering%20or%20semantic%20search. 7/12 jamwal, kokila (2024) boosting data findability: the role of ai-enhanced keywords, iassist quarterly 48(4), pp. 1-12. doi: https://doi.org/10.29173/iq1127 provide a list of output labels associated with the input instance. figure 8 illustrates the multi-label classification model. the multi-label classification problem differs from other classification problems9. the basic form of classification problem is the binary classification problem, where the output label for an input instance is only one of the two available output labels. other than binary classification problem, there is multiclass classification problem. in a multi-class classification problem, an input instance can have only one associated output label from among more than two output labels. due to mutually exclusive classes in binary and multi-class classification, we deal with them differently than multi-label classification problems. multi-label classification problems are generally dealt with through problem transformation or algorithm adaptation methods (pant et al., 2019). as the name suggests, problem transformation methods transform the multi-label problem into a set of binary classification problems, which are then handled using algorithms for single-class classifiers. the algorithm adaptation methods adapt the algorithms and perform multi-label classification directly instead of figure 8: multi-label classification model simplifying the problem as a binary classification problem. we employ the problem transformation method by dividing the task into a series of binary classification problems. at the end of this component, we get the predicted output labels for each input instance. experimental setup in this section, we discuss our experimental setup including details on the datasets used, parameter settings, and the evaluation metrics used to evaluate the performance of our approach. all experiments are conducted on a laptop with 16 gb of ram, and an intel core i7 processor. the software environment includes python 3.11.9, tensorflow 2.12.0, and other necessary libraries. datasets we extract the dataset from gesis search via elasticsearch10. elasticsearch allows the users to store, search, and analyze huge volumes of data quickly. each row in our dataset represents a research study, and the column ‘abstract’ acts as a feature of the study. the ‘topics’ is the set of labels associated with the study. for our experiments, we extract 64,791 studies from gesis search, which depicts the https://doi.org/10.29173/iq1127 8/12 jamwal, kokila (2024) boosting data findability: the role of ai-enhanced keywords, iassist quarterly 48(4), pp. 1-12. doi: https://doi.org/10.29173/iq1127 number of instances in our dataset. among these, the total number of studies with missing values in the ‘topics’ metadata field is 15,776. considering that, we select 44,417 studies containing more than one cv topics. later, we de-duplicate the studies, resulting in 36474 studies at the end of the data preprocessing step. the number of unique labels from the cv sources equals 1,221. for more clarity on the dataset, we extract the following dataset properties for our multi-label dataset (sorower, 2010). • distinct label set (dl): it is the total number of distinct label combinations in a dataset • proportion of distinct label set (pdl): it is the measure of the number of distinct label sets per input instance • label cardinality (lcard): it is the average number of labels per input instance. lcard is a measure of ‘multi-labelledness’ • label density (lden): it is measured as lcard normalized by the number of labels, which means lden is the ratio of lcard to the number of distinct labels. the values for each of these properties are depicted in table 2. property value distinct label set (dl) 7,237 proportion of distinct label set (pdl) 0.198415 label cardinality (lcard) 5.966387 label density (lden) 0.002699 table 2: properties of our multi-label dataset since distinct label set (dl) and proportion of distinct label set (pdl) depict the diversity and coverage of labels in a dataset, they are essential to determine if the model needs to handle a wide range of labels or a more limited subset. the value for dl is 7,237 and since the number of unique studies is 36,474, value for pdl is 0.198415. label cardinality (lcard) and label density (lden) provide information about the average label load per input instance. high lcard and lden values indicate that dataset instances are likely to have multiple labels, necessitating robust multi-label handling capabilities in the model. our dataset averages 5.966387 labels per instance, and the number of unique labels is 2,210, therefore the value of lden is 0.002699. parameter settings deep learning models use different hyperparameters. the hyperparameters control the learning process of a model. in order to get the best model results, these hyperparameters can be optimized using different techniques. we have employed one such hyperparameter optimization technique called grid search, which exhaustively searches the best parameter over all possible combinations of hyperparameters. during the multi-label classification model step, the output embeddings of size 512 each, created from the embedding generation step, are taken as input. the model consists of a series of fully connected (fc) layers. the fully connected layers consist of neurons. neurons in deep learning models are nodes through which data and computations flow. neurons receive one or more input signals, perform some calculations, and send output signals to the next layer of neurons. based on the https://doi.org/10.29173/iq1127 9/12 jamwal, kokila (2024) boosting data findability: the role of ai-enhanced keywords, iassist quarterly 48(4), pp. 1-12. doi: https://doi.org/10.29173/iq1127 hyperparameter optimization, the number of neurons in fully connected layer 1 (fc1) is 1,024, and in fully connected layer 2 (fc2), it is 256. the size or number of neurons in each layer represents the number of features leaned from the input data by this layer. both fully connected layers use rectified linear unit (relu) as the activation function. activation functions add non-linearity to the model and help the model learn complex relationships in data. for the classification layer, the sigmoid function is used as the activation function, with an output size equal to the number of unique topics. we split the dataset into training and testing sets for the multi-label classification problem, with 80% data in training and 20% in testing sets, respectively. the training set is passed to the multi-label classification model to learn patterns in the data, whereas the test set tests the model's performance. during training, we can set other parameters like batch size, learning rate, and number of epochs. batch size is the number of training samples processed before the model's internal parameters are updated. the learning rate controls the rate or speed at which the model learns. the number of epochs signifies the number of times the entire training set is passed through the model. based on the hyperparameter optimization for our model, the batch size is 128, the learning rate is set to 0.001 for the adam (adaptive moment estimation) optimizer. optimizers are algorithms or methods used to change the attributes of neural network such as weights and learning rate for better learning by the model. the number of epochs is set to 200. we use early stopping as a regularization technique to avoid overfitting. regularization techniques are methods used to improve a model's ability to generalize to new data by adding constraints to the learning process. early stopping is a regularization technique used to prevent overfitting in training deep learning models. overfitting occurs when a model becomes too specialized to the training data, capturing noise or irrelevant patterns, leading to poor generalization of unseen data. early stopping monitors the model's performance on a validation set during training and halts the training process when the model's performance degrades, indicating it has started to overfit. if the model overfits, it will not perform well for the new unseen data. evaluation metrics the performance of our model for multi-label classification is evaluated using following metrics (sorower, 2010): • micro-averaging: it is calculated by aggregating the contributions of all classes to compute the average metric. it considers the proportion of each class in the overall population. o precision: the ratio of correctly predicted instances for a class to the total instances predicted as that class. o recall (sensitivity or true positive rate): the ratio of correctly predicted instances for a class to the total instances that belong to that class. o f1 score: f1 score is the harmonic mean of precision and recall, providing a single metric that balances both. • hamming loss: hamming loss measures the fraction of incorrectly predicted labels (both false positives and false negatives) over the total number of labels. evaluation results in this section, we evaluated our model’s performance through the evaluation metrics defined in section ‘evaluation metrics’. the result of our evaluation is provided in table 3. the value of all the https://doi.org/10.29173/iq1127 https://databasecamp.de/en/ml/relu-en 10/12 jamwal, kokila (2024) boosting data findability: the role of ai-enhanced keywords, iassist quarterly 48(4), pp. 1-12. doi: https://doi.org/10.29173/iq1127 metrics used for evaluation ranges between 0 and 1. the higher the precision, recall, and f1 score value, the better the model's performance. however, a lower value for hamming loss depicts better performance. the precision value is 68%, recall is 56%, f1 score is 61%, and hamming loss is 0.27%. we asked human annotators to validate the performance of our model. we provided the annotators with a model-labeled sample of 20 instances, with ten random instances containing no label and ten random instances containing user-defined labels. the annotators got text in the ‘abstract’ metadata field and the labels predicted by our model. we asked two human annotators to label the predictions per instance into two categories: the number of correct labels and the number of incorrect labels. the annotators indicated the number of labels they consider correct and the number they consider incorrect in the two categories, respectively. in case of disagreement between the two annotators, we asked a third annotator to resolve the disagreement. ultimately, we took the labels generated by the third annotator after resolved annotations. our model predicted 74% correct and 26% incorrect labels based on this sample. this experiment validated the performance of our model; however, we can further improve the performance by including more cv sources and input instances. metrics value precision (micro) 0.6828 recall (micro) 0.5577 f1 score (micro) 0.6140 hamming loss 0.0027 table 3: evaluation results for our model use case: a comparison of our model with llm we compared our model’s performance with large language models (llms). as llm, we used openai’s chatgpt 3.5. we randomly took multiple samples, each containing 30 studies from the test data to carry out this experiment. to compare the performance of our model with llm, we labeled the input instances using our model and using llm with relevant labels from the cv. we compared our model with llm during this experiment; the detailed comparison is in table 4. we provided the chatgpt with a prompt, including the problem statement, one ‘abstract’ text at a time and the set of all possible values from cv. we asked chatgpt to provide five labels for each ‘abstract’ text. a sample prompt and result from chatgpt is provided in the figure 9. we averaged the performance of our model and llm across the samples and carried out the friedman test (bogatinovski et al., 2022) to calculate the significance of our results. table 4 provides the mean and standard deviation of our model and llm across various metrics. according to the results, our model performed significantly better for the multi-label classification task on the test data samples. metrics value (our model) value (llm) precision (micro) 0.85 ± 0.031 0.35 ± 0.02 recall (micro) 0.60 ± 0.10 0.31 ± 0.16 f1 score (micro) 0.69 ± 0.08 0.31 ± 0.08 hamming loss 0.02 ± 0.004 0.0528 ± 0.0012 table 4: comparison of evaluation results for our model vs llm https://doi.org/10.29173/iq1127 11/12 jamwal, kokila (2024) boosting data findability: the role of ai-enhanced keywords, iassist quarterly 48(4), pp. 1-12. doi: https://doi.org/10.29173/iq1127 figure 9: sample chatgpt prompt and results conclusion researchers have focused on making their data fair. data findability and reusability are among the main pillars of fair principles. using controlled vocabulary improves findability. a controlled list helps to restrict the values for a particular metadata field, ensuring further data findability and reusability later. however, missing and user-defined topics hinder data findability and reusability. we have devised an approach to automatically label studies for the ‘topics’ metadata field. the ‘topics’ metadata field labels are taken from a set of cv sources. our topic classification model efficiently classifies the studies into multiple topic labels based on the ‘abstract’ metadata field. our model can help people across the whole lifecycle of research data. for data depositors, it can help make their data findable and reusable. our model can help save the time and effort required for data curation and can easily manage larger volumes of data to help the data curators. our model allows data users to enjoy improved findability with enhanced topic discovery. we compared our model with chat gpt 3.5 for the same task, and our model proved to work better with the random sample from the test data. in the future, we will try to improve our classification model and examine the performance of different open-source llm models for assigning topic labels to a study. acknowledgements i would like to thank libby bishop (gesis) for her suggestions and comments on the manuscript. https://doi.org/10.29173/iq1127 12/12 jamwal, kokila (2024) boosting data findability: the role of ai-enhanced keywords, iassist quarterly 48(4), pp. 1-12. doi: https://doi.org/10.29173/iq1127 references bogatinovski, j., todorovski, l., džeroski, s., & kocev, d. (2022). comprehensive comparative study of multi-label classification methods. expert systems with applications, 203, 117215. pant, p., sai sabitha, a., choudhury, t., & dhingra, p. (2019). multi-label classification trending challenges and approaches. emerging trends in expert applications and security: proceedings of iceteas 2018, 433-444. reimers, n., gurevych, i. (2019). sentence-bert: sentence embeddings using siamese bertnetworks. in: emnlp-ijcnlp 2019. acl. sorower, m. s. (2010). a literature survey on algorithms for multi-label learning. oregon state university, corvallis, 18(1), 25 steiner, g., (2023). the exponential growth of research data. analytical science magazine, vol. 3 may/23 (https://analyticalscience.wiley.com/content/article-do/exponential-growth-research-data) [accessed 27/06/2024] endnotes 1 kokila jamwal, gesis-leibniz institute for the social sciences, unter sachsenhausen 6, 50667 cologne, germany, kokila.jamwal@gesis.org 2 gesis search (https://search.gesis.org/?source=%7b%22query%22%3a%7b%22bool%22%3a%7b%22must%22%3 a%7b%22match_all%22%3a%7b%7d%7d%2c%22filter%22%3a%5b%7b%22term%22%3a%7b%22t ype%22%3a%22all%22%7d%7d%5d%7d%7d%7d&lang=en) [accessed 15/05/2024]. 3 metadata for data management: a tutorial: controlled vocabularies (https://guides.lib.unc.edu/c.php?g=8749&p=44502) [accessed 19/06/2024]. 4 controlled vocabularies (https://www.cms.hu-berlin.de/en/dl-en/datamanen/share/documentation/controlled-vocabularies) [accessed 19/06/2024]. 5 controlled vocabulary for dam (https://digitalassetmanagementnews.org/dam-guruarchive/controlled-vocabulary-for-dam/) [accessed 19/06/2024]. 6 controlled vocabulary policy (https://www.toronto.ca/wp-content/uploads/2022/05/8f0acontrolledvocabularypolicy2021.pdf) [accessed 19/06/2024]. 7 controlled vocabulary (https://www.informedbyte.com/services/controlled-vocabulary) [accessed 21/06/2024]. 8 gesis controlled vocabulary service (https://lod.gesis.org/en/) [accessed 20/05/2024]. 9 difference: binary vs multiclass vs multilabel classification (https://vitalflux.com/difference-binarymulti-class-multi-labelclassification/#:~:text=to%20summarize%2c%20binary%20classification%20is%20a%20supervised% 20machine,predict%20one%20or%20more%20classes%20for%20an%20item.) [accessed 15/04/2024]. 10 elasticsearch (https://www.elastic.co/elasticsearch) [accessed 21/05/2024]. https://doi.org/10.29173/iq1127 https://analyticalscience.wiley.com/content/article-do/exponential-growth-research-data mailto:kokila.jamwal@gesis.org https://search.gesis.org/?source=%7b%22query%22%3a%7b%22bool%22%3a%7b%22must%22%3a%7b%22match_all%22%3a%7b%7d%7d%2c%22filter%22%3a%5b%7b%22term%22%3a%7b%22type%22%3a%22all%22%7d%7d%5d%7d%7d%7d&lang=en https://search.gesis.org/?source=%7b%22query%22%3a%7b%22bool%22%3a%7b%22must%22%3a%7b%22match_all%22%3a%7b%7d%7d%2c%22filter%22%3a%5b%7b%22term%22%3a%7b%22type%22%3a%22all%22%7d%7d%5d%7d%7d%7d&lang=en https://search.gesis.org/?source=%7b%22query%22%3a%7b%22bool%22%3a%7b%22must%22%3a%7b%22match_all%22%3a%7b%7d%7d%2c%22filter%22%3a%5b%7b%22term%22%3a%7b%22type%22%3a%22all%22%7d%7d%5d%7d%7d%7d&lang=en https://guides.lib.unc.edu/c.php?g=8749&p=44502 https://www.cms.hu-berlin.de/en/dl-en/dataman-en/share/documentation/controlled-vocabularies https://www.cms.hu-berlin.de/en/dl-en/dataman-en/share/documentation/controlled-vocabularies https://digitalassetmanagementnews.org/dam-guru-archive/controlled-vocabulary-for-dam/ https://digitalassetmanagementnews.org/dam-guru-archive/controlled-vocabulary-for-dam/ https://www.toronto.ca/wp-content/uploads/2022/05/8f0a-controlledvocabularypolicy2021.pdf https://www.toronto.ca/wp-content/uploads/2022/05/8f0a-controlledvocabularypolicy2021.pdf https://www.informedbyte.com/services/controlled-vocabulary https://lod.gesis.org/en/ https://vitalflux.com/difference-binary-multi-class-multi-label-classification/#:~:text=to%20summarize%2c%20binary%20classification%20is%20a%20supervised%20machine,predict%20one%20or%20more%20classes%20for%20an%20item https://vitalflux.com/difference-binary-multi-class-multi-label-classification/#:~:text=to%20summarize%2c%20binary%20classification%20is%20a%20supervised%20machine,predict%20one%20or%20more%20classes%20for%20an%20item https://vitalflux.com/difference-binary-multi-class-multi-label-classification/#:~:text=to%20summarize%2c%20binary%20classification%20is%20a%20supervised%20machine,predict%20one%20or%20more%20classes%20for%20an%20item https://vitalflux.com/difference-binary-multi-class-multi-label-classification/#:~:text=to%20summarize%2c%20binary%20classification%20is%20a%20supervised%20machine,predict%20one%20or%20more%20classes%20for%20an%20item https://www.elastic.co/elasticsearch newsletter vol.1, no. 4 u.s. department of health, education and welfare. standardized micro-data tape transcripts. (dhew pub. no. 76-1213). washington: gpo 1976. 30 pp. describes data bases for health statistics available from the federal government. prices, technical configurations. available from gpo. u.s. department of labor. bureau of labor statistics. bls data bank files and statistical routines , (draft brochure). describes 30 data files available from the bureau of labor statistics. describes publications. write: commissioner, bls, dept. of labor, washington, dc 20212. u.s. general accounting office. 1976 congressional sourcebook series (opa 76-23) federal information sources and systems; a directory for the congress . washington, dc: gpo 1976. 456 pp. paperback. an indexed reference guide to over 1,000 federal sources and information systems in 63 federal agencies, which contain budgetary, fiscal and program-related data. available from gpo. u.s. national archives and records service. catalog of machine-readable records in the national archives of the united states . washington: gpo, 1977. 37 pp. describes 99 files created by federal agencies and retained in the national archives. prices, available technical configurations. write: national archives (nnr), washington, dc 20408. university of waterloo. leisure studies data bank. april, 1977. 23 pp. lists 24 files created by private and federal canadian agencies. in french and english. write: dept. of recreation, university of waterloo, waterloo, ontario, canada. n2l 3g1 . discussion paper/gary m. grandon using spss mult response to generate filtered marginals gary m. grandon social science data center the university of connecticut, storrs with the release of spss's version 7.0 this year, several additions have been added to its already extensive battery of analysis programs. mult response is one of these programs. it generates frequency counts and bivariate tables for "dummy" coded multiple response questions. these multiple response variables are a real nuisance to the analyst and the program at face value provides a means for their interpretation. considerations for the analysis of multiple response data are not the subject of this paper though, but rather the use of this program for quite another purpose; the generation of "filtered marginals." filtered marginals are typically basic frequency counts and percentages for specified variables for each of a number of subpopulations within a study. spss has in the past provided *select if statements to facilitate the processing of subpopulations. each such statement followed by frequencies procedure statements will generate filtered marginals. the inherent problem with this type of coding 30 sist newsletter vol.1, no. 4 is that each *select if statement requires a complete reading of the data set. for large studies this procedure can be quite inefficient! another alternative to the *select if procedure would be to create a new variable with codes indicating each filter group (subpopulation) . then the very fast crosstabs procedure can be used with the new filter group variable as one variable in the bivariate table and the variable upon which frequencies are to be generated as the other. a large number of frequencies can be generated in this manner with only one read of the data. the problem with this crosstabs procedure is that it requires that each filter group or subpopulation be discrete. that is, any one subject can only be in one filter group. this restriction is removed by using the mult response technique suggest in this paper. the following setup using the spss osiris interface to shorten the presentation illustrates the use of the mult response procedures to perform filtered marginals: 1 run name osiris vars input medium n of cases weight if if mult response statistics read input data finish yankelovich—combined studies-1976 v7,v8,v11,v13,v18,v28,v64,v75 tape unknown v75 ((v7 eq 1 or 2 or 3) and (v8 eq 6) and (vll eq 1) and (v18 eq 1)f1=1 ((v7 eq 5 or 6) and (v8 eq 1 or 2) and (vll eq )3) and {v18 eq 1)f2=1 groups=filters (fl f2(l))/ variables=v13(1,4),v28(1,3),v64{1,3)/ tables=filters by vi 3 to v64/ 1 not only does the technique illustrated above produce the same kind of results as the *select if process but it formats the output in such a way as to allow the user to scan differences across subpopulations within a single tabular display. a single procedure statement like the one above can process 20 filter groups and up to 100 analysis variables. each frequency variable will have its own table. a resource expenditure comparison between the *select if procedure and the mult response procedure was performed on the 1976 yankelovich combined study archived at the social science data center of the university of connecticut. the study has a n of 7,977 subjects. nineteen filter groups were identified and frequencies for 45 analysis variables were obtained using both techniques. table 1 provides relevant resource comparisons for the techniques. the crosstabs technique was not included due to its unique group membership constraint. clearly costs are much lower for the mult response procedure. sist newsletter vol.1, no. 4 '^select if and mult response techniques for filtered marginals. resource *select if mult response number of jobs 2 1 kbyt-sec 21,016 24,957 elap-kbs 180,493 73,752 reader excp** 93 31 printer excp 14,630 3,844 public disk excp 10,476 360 private disk excp 10 5 tape excp 72 36 tape mounts 2 1 cost $31.70 $13.39 elapsed time 13.48 min. 6.23 min. *comparisons were performed on the university of connecticut's research computer center ibm 360/65 ibm 370/155 systems running under os and shared spool hasp. local costs may vary from installation to installation. **each excp represents the "execution of a channel program" and indicates the movement of one block of data. the mult response procedure can further be extended to filtered bivariate tables by including a second by_ statement on the tables card. this technique is not illustrated due to space limitations. readers are referred to their local installations for further documentation of spss version 7.0. discussion paper/riohard c. roistacher the following article describes the work of richard c. roistacher and barbara noble at the center for advanced computation at the university of illinois. they are involved in the development of guidelines for the descriptive materials which accompany a data file. a source documentation style manual by richard roistacher center for advanced computation university of illinois urbana, illinois barbara noble and richard roistacher of the university of illinois' center for advanced computation are currently developing a style manual for the documentation of machine readable data. the manual, which is being developed as part of a project funded by the u.s. department of justice's law enforcement assistance administration, is presently available in draft form. the manual, conforming to the 1/8 adeyeye, sophia vivian.& oladokun, taofeek abiodun (2023) application of emerging technologies for research support in nigerian academic libraries: trends, problems and prospects, iassist quarterly 47(3-4), pp1-8. doi: https://doi.org/10.29173/iq1069 the creative commons-attribution-noncommercial license 4.0 international applies to all works published by iassist quarterly. authors will retain copyright of the work and full publishing rights. application of emerging technologies for research support in nigerian academic libraries: trends, problems and prospects sophia vivian adeyeye1 & taofeek abiodun oladokun2 abstract academic libraries in the modern era are often asked to justify their existence in tertiary institutions by showing how the institution and society at large have benefited from the library services. one main area that librarians often point to is the research productivity of members of academic institutions. however, studies have shown that research productivity among nigerian scholars is low, which means that academic libraries have to be more innovative in supporting researchers in their domains. emerging technologies offer innovative ways of supporting research activities by providing tools and resources that streamline the research process and ensure proper visibility for research outputs of academic library clients. this article, which is based on a review of previous studies, explores various areas where academic libraries in nigeria can apply emerging technologies, the likely challenges, and strategies that can be adopted to ensure sustainable use of emerging technologies in academic libraries. it has been found that emerging technologies can enhance existing library services and create new ones, such as data mining, data management, and scholarly communication, among others. however, although steps are being taken by academic librarians in nigerian tertiary institutions to leverage technology in providing the needed support for researchers, the pace of technology adoption is still slow and the range of technologies being adopted is limited compared to available options. this state of affairs has been attributed to challenges such as lack of infrastructure, librarians’ skills, and a negative attitude towards change. the study recommends a multidimensional approach to the application of emerging technologies in nigerian academic libraries keywords artificial intelligence, emerging technologies, library automation, library services, research support. introduction the use of emerging technologies in academic libraries has become widely accepted, especially in the provision of various research support services. the modern academic library is no longer expected to play the role of information custodian but rather operate as facilitators in the creation, dissemination and use of knowledge. academic librarians are laying more emphasis on research-centered services such as information organization and retrieval, citation management, data management, electronic publication, and other services capable of enhancing the quality of research output in academic institutions (sewell & kingsley, 2017; moruf & dangani, 2020). the ability of academic libraries to achieve their aim of assisting researchers to effectively use available information resources for the https://doi.org/10.29173/iq1069 https://creativecommons.org/licenses/by-nc/4.0/ 2/8 adeyeye, sophia vivian.& oladokun, taofeek abiodun (2023) application of emerging technologies for research support in nigerian academic libraries: trends, problems and prospects, iassist quarterly 47(3-4), pp1-8. doi: https://doi.org/10.29173/iq1069 creation of new knowledge now rests on how well librarians are familiar with and able to utilize emerging technologies to support all aspects of research. the emergence of information technology and the remote access to information resources and other research tools have led to questions about the relevance of library services in the research process (momoh & folorunso, 2019; arumuru, 2020). available information, however, indicates that, despite the touted ‘unlimited’ access to information resources and research tools, the research output of researchers in developing countries such as nigeria is still below expectations (okagbue et al., 2018; orji & anunobi, 2019; oluwasanu et al., 2019). with the importance of research to national development, this indicates a gap that academic librarians can fill by providing the right library and information services capable of not only enhancing the research productivity of researchers in nigeria but also contributing to the visibility and impact of the research output emanating from nigerian tertiary institutions. however, although steps are being taken by academic librarians in nigerian tertiary institutions to leverage technology in providing the needed support for researchers in their domains, the pace of technology adoption is still slow and the range of technologies being adopted is limited compared to available options (bakare, 2023). according to moruf and dangani (2020), academic library users are demanding broader range of services delivered in accurate and efficient manner. so it is important for librarians to do all they can to meet these expectations. the purpose of this article is therefore to highlight various emerging technologies available to academic librarians for research support, areas in which academic librarians can apply emerging technology to support researchers, the issues surrounding the use of emerging technology in academic libraries, and the strategies that can be applied to overcome the challenges. emerging technologies in academic libraries the concept of emerging technology would seem straightforward to describe based on its name as a new technology that is just coming to people’s attention based on its uniqueness or usefulness. however, rotolo, hicks, and martin (2015) are of the opinion that the definition of what constitutes an emerging technology depends on the perspective of each scholar. this is reflected in the submission of saibakumo (2021), who posited that emerging technologies in the context of library and information science are those technologies that are recently being introduced into the librarianship profession based on the recognition that they can help improve service provision. from this perspective, emerging technology in the context of libraries and information science can be described as technological innovations that have recently been discovered to be relevant to service provision in libraries. there are many examples of these technologies that can be used to support the activities of researchers in order to boost research productivity. emerging technologies relevant to research support, according to moruf and dangani (2020), include bibliographic citation management software such as mendeley, etc; instructional system design software such as blackboard, edmodo, etc.; electronic copyright management systems; classroom management software such as moodle, google classroom, canva, etc.; library automation software such as koha, greenstone, d-space; electronic resource management software; and integrated search software, among others. saibakumo (2021) also identified emerging technologies that include qr (quick response) barcode technology, cloud computing, robotics, and artificial intelligence (ai). in https://doi.org/10.29173/iq1069 3/8 adeyeye, sophia vivian.& oladokun, taofeek abiodun (2023) application of emerging technologies for research support in nigerian academic libraries: trends, problems and prospects, iassist quarterly 47(3-4), pp1-8. doi: https://doi.org/10.29173/iq1069 addition, odeyemi (2019) identified ambient intelligence and data mining as new emerging technologies that have been introduced into libraries to improve the efficiency of librarians and satisfy the needs of users by bridging the information gap. these studies have shown that there are numerous emerging technologies relevant to the needs of librarians, especially in supporting researchers. as pointed out by moruf and dangani (2020), technologies are emerging at a rapid pace and potentials adopters such as librarian can decide on which ones to adopt. it is therefore important to match the available technologies with the relevant services that libraries can provide to researchers in the 21st century. application of emerging technology in research support as emerging technologies have been found to enhance research and information management, academic libraries have the option to adopt these technologies to enhance their service delivery, especially in the area of research support. some of these services include information resource management, research data management, copyright advisory, citation management, scholarly communication, data analysis services, research preservation and curation, and digital literacy training, to name but a few (das & banerjee, 2021; keller, 2015; sewell & kingsley, 2017). information resource management the main role of academic libraries is to provide the necessary information resources needed by researchers. with the aid of emerging technologies, libraries can move from being information custodians concerned with storing information to information access providers. this is done through the use of integrated library management systems. an integrated library management system (ilms), according to sheik and olugbenga (2019), is a complex program or database that can be used to automate regular library routines such as cataloging and classification, charging and discharging, and remote access to information. a typical ilms has an opac (online public access catalogue), which acts as a search engine for library users to find the available information. although the ilms has been in existence for a while, the application of emerging technologies, particularly artificial intelligence, can now be used to turn the opac into an expert system (asemi & asemi, 2018). as a result of the integration of ai, modern-day opac has transformed into a real information system that can provide electronic information resources, track user preferences, and make necessary recommendations. the most important factor is that it makes the carefully selected information resources by librarians available to the researchers round the clock and online. data mining (bibliometric): the amount of existing literature has grown to the extent that no individual can be able to synthesize and engage with it. thus, researchers try to find ways in which emerging technologies can speed scientific discovery by incorporating artificial intelligence into research workflows, such as the use of emerging technology to automate the information search. data mining tools are one of the most useful emerging technologies that can be used by information professionals to bring order into the chaotic world of information. the use of tools such as rapid miner, orange, weka, etc. (educba, 2021) offers various advantages to librarians. according to lone and khan (2014), data mining allows libraries to effectively support researchers by providing insight into the composition of database collections. with data mining, librarians can easily identify emerging trends in research and use the information to provide current awareness services and selective dissemination of information to researchers, thereby saving them a lot of preliminary work in their research. https://doi.org/10.29173/iq1069 4/8 adeyeye, sophia vivian.& oladokun, taofeek abiodun (2023) application of emerging technologies for research support in nigerian academic libraries: trends, problems and prospects, iassist quarterly 47(3-4), pp1-8. doi: https://doi.org/10.29173/iq1069 digital reference services: online information resources have become the first choice of researchers when seeking information resources for their research. however, they often face challenges in using these resources (hagiwara et al., 2022; akande & popoola, 2022). librarians have found that with technology, they can extend the reference services that they traditionally render within the library to the digital space. many libraries are now taking advantage of ai-enabled chatbots to render roundthe-clock digital reference services even when librarians are not available (wan, 2022; panda & chakravarty, 2022). the use of chatbots offers researchers the opportunity to get help with using electronic information resources whenever they need help. this can enhance productivity by ensuring that minor issues hindering effective use of information systems are solved as they arise. research data management: one of the practices that has been known to enhance research reproducibility is research data management. emerging tools in research data management include the open science framework, dmp tool, researchworks archive, redcap, perma.cc, orcid, and several others (library guides, 2022). in the digital world, research data, both qualitative and quantitative, has become an essential commodity among researchers. calvert and kennedy (2020) reported that the collection, exchange, and preservation of data have become common practices among researchers. this has made libraries in the developed world take advantage of emerging technologies for data management and dissemination. some libraries have created data repositories, while others are partnering with large data repositories such as zenodo and dryad. the idea is to free researchers from data management tasks so that they can focus more on research. calvert and kennedy (2020) found that librarians, particularly in the united states, are making advances in research data management by leveraging emerging technologies. the application of emerging technology in data management offers several advantages, which include a reduction in the cost of data curation, providing opportunities for researchers to discover and access research data from local and global sources, and adding value to data through expert processing and organization, which means that researchers are provided with a body of related data relevant to a given study. the net effect of reducing the task of researchers to allow them to dedicate more time to research. issues in the use of emerging technologies for research support in nigeria the benefits accrued from the application of emerging technology notwithstanding, available evidence from nigeria shows that only a few academic libraries have made appreciable attempts at incorporating them in library and information services provision. for one, most librarians in nigeria are not clear about what constitutes emerging technology and consequently, their specific application to research. agbetuyi and isah (2021) grouped emerging technologies into opac, mobile-based technology, web 2.0 technology, institutional repositories, and cloud computing technology. but the authors did not go into specifics. the study by adeoye, oladokun, and opalere (2021) was more specific as it examined the readiness for digital reference services among librarians. it was found that the majority of the librarians were ready but there are institutional issues affecting digital reference services in nigerian libraries. in the same vein, rotimi, et al., (2022) found that nigerian libraries are yet to exploit the abundant opportunities offeredby emerging technology rendering innovative library services. the study found that library services in nigerian academic libraries are still dominated by manual operations with limited application of basic technologies. also, examining the readiness of nigerian academic librarians to adopt robotic technologies in rendering library services owolabi et al. (2022) reported that, neither the libraries nor the academic librarians show adequate level of readiness to incorporate the emerging https://doi.org/10.29173/iq1069 5/8 adeyeye, sophia vivian.& oladokun, taofeek abiodun (2023) application of emerging technologies for research support in nigerian academic libraries: trends, problems and prospects, iassist quarterly 47(3-4), pp1-8. doi: https://doi.org/10.29173/iq1069 technology in their services. the same could be said of other emerging technologies that could be adopted to support research activities in tertiary institutions. there are various reasons for this state of affairs . abayomi et al., (2021) examined the level of awareness and perception of academic librarians regarding the adoption of artificial intelligence (ai) in library operations. the finding showed that the librarians were quite aware of the benefits of ai tools but they did not support its adoption for fear of being made redundant and losing their jobs. in addition, misau (2021) reported that while emerging technologies such as expert systems in reference services, technical, indexing, acquisition, pattern recognition and robotics are relevant to academic library operation, there is a need for academic librarians to be properly trained and sensitized on the use of these technologies. however, the factors affecting the adoption and use of emerging technologies for research support are not limited to librarians’ skills or attitude. according to bawack and nkolo (2018), adoption of emerging technologies has not taken root in developing countries due to institutional factors such as lack of a clear cut policy, infrastructural deficit, and the dearth of innovative library managers. the implication of all this is that a multidimensional approach is required to stimulate the widespread adoption of emerging technologies to facilitate effective research support in nigerian academic libraries. the strategy to be adopted to ensure the adoption of emerging technologies must take cognizance of the various interwoven issues affecting the adoption of emerging technologies. strategies to boost the application of emerging technologies for research support in academic libraries. as an innovation that has come to change the existing status quo, the successful adoption and application of emerging technologies in research support, even in developed countries, requires a holistic approach that must be driven by library management. according to calvert and kennedy (2020), academic library managers must secure the support of key decision-makers in their institutions in order to create an official policy that mandates and supports the use of technology. this is important because of all the infrastructure and logistics required for the effective use of technologies. academic librarians are experienced in traditional research and information services, which offers them leverage in the use of technologies for research support. however, effective use of emerging technologies to support researchers and conduct their own research requires some upskilling. it is important that academic librarians are well versed in the technologies they are introducing to their clients. in addition, it is important that academic librarians always stay a pace or two ahead of their clients in terms of their skills and knowledge about emerging technology. the level of skill acquisition required must, however, be coordinated to ensure that the skills acquired match the overall aims of the library and the institution it serves. in essence, academic libraries must take charge of the skill development of their personnel instead of leaving every librarian to fend for themselves. conclusion. the current level of research productivity in nigerian tertiary institutions shows that there is a big role for the academic library to play in boosting research productivity. providing support for researchers in the 21st century has, however, evolved beyond merely stocking a large collection; the main trend today is the provision of access to information tailored as close to the needs of the researcher as https://doi.org/10.29173/iq1069 6/8 adeyeye, sophia vivian.& oladokun, taofeek abiodun (2023) application of emerging technologies for research support in nigerian academic libraries: trends, problems and prospects, iassist quarterly 47(3-4), pp1-8. doi: https://doi.org/10.29173/iq1069 possible. to meet this need, it requires the application of emerging technologies. these are new digital innovations used to create, synthesize, organize, and make information available for easy access. this study has, however, shown that despite the availability of various technologies, most of which are freely available online, the application of emerging technologies is progressing at a slower pace which suggests a need for strategic intervention from all stakeholders. it is expected that the various recommendations made in this study and other relevant studies would lead to an increased pace of technology adoption innigerian academic libraries.. references abayomi, o.k., adenekan, f.n., abayomi, a.o., ajayi, t.a. and aderonke, a.o., 2021. awareness and perception of the artificial intelligence in the management of university libraries in nigeria. journal of interlibrary loan, document delivery & electronic reserve, 29(1-2), pp.13-28. https://doi.org/10.1080/1072303x.2021.1918602 adeoye, a.a., oladokun, t.a. and opalere, e., 2022. readiness for digital reference services: a survey of libraries in ibadan metropolis, nigeria. international information & library review, 53(4), pp.306-314. https://doi.org/10.1080/10572317.2020.1841526 agbetuyi, p. a., & isah, a. (2021). the implementation of emerging technologies for sustainable academic libraries: a comparative analysis between developed and developing countries [conference session]. proceedings of the 2nd international conference on ict for national development and its sustainability, faculty of communication and information sciences, university of ilorin, ilorin, nigeria2021. http://repository.futminna.edu.ng:8080/jspui/bitstream/123456789/6232/1/524-131799-1-10-20210312.pdf akande, f.t. and popoola, s.o., 2022. awareness and use of electronic resources as predictors of scholarly publication output of researchers in national agricultural research institutes in nigeria. mousaion: south african journal of information studies, 40(1), pp.18-pages. https://hdl.handle.net/10520/ejc-mousaion-v40-n1-a9 arumuru, l. 2020. re-positioning the 21st century university libraries in nigeria: the role of librarians and the need for innovative services for sustainable development. library philosophy and practice (e-journal). https://digitalcommons. unl. edu/libphilprac/3936 bakare, o.d., 2023. emerging technologies as a panacea for sustainable provision of library services in nigeria. in global perspectives on sustainable library practices (pp. 1-21). igi global. https://www.igi-global.com/chapter/emerging-technologies-as-a-panacea-forsustainable-provision-of-library-services-in-nigeria/313365 bawack, r. and nkolo, p., 2018. open access movement: reception and acceptance by academic libraries in developing countries. library philosophy and practice, pp.0_1-24. https://core.ac.uk/download/pdf/188141049.pdf https://doi.org/10.29173/iq1069 https://doi.org/10.1080/1072303x.2021.1918602 https://doi.org/10.1080/10572317.2020.1841526 http://repository.futminna.edu.ng:8080/jspui/bitstream/123456789/6232/1/524-13-1799-1-10-20210312.pdf http://repository.futminna.edu.ng:8080/jspui/bitstream/123456789/6232/1/524-13-1799-1-10-20210312.pdf https://hdl.handle.net/10520/ejc-mousaion-v40-n1-a9 https://www.igi-global.com/chapter/emerging-technologies-as-a-panacea-for-sustainable-provision-of-library-services-in-nigeria/313365 https://www.igi-global.com/chapter/emerging-technologies-as-a-panacea-for-sustainable-provision-of-library-services-in-nigeria/313365 https://core.ac.uk/download/pdf/188141049.pdf 7/8 adeyeye, sophia vivian.& oladokun, taofeek abiodun (2023) application of emerging technologies for research support in nigerian academic libraries: trends, problems and prospects, iassist quarterly 47(3-4), pp1-8. doi: https://doi.org/10.29173/iq1069 calvert, scout, and mary l. kennedy. "emerging technologies for research and learning: interviews with experts." 2020. https://doi.org/10.29242/report.emergingtech2020.interviews das, a. and banerjee, s., 2021. optimising research support services through libraries: a review of practices. library philosophy and practice, pp.1-43. https://digitalcommons.unl.edu/libphilprac/5515/. data mining tool. (2021, november 1). educba. https://www.educba.com/data-mining-tool/ hagiwara, y., ishita, e., watanabe, y. and tomiura, y., 2022. identifying scholarly search skills based on resource and document selection behavior among researchers and master’s students in engineering. college & research libraries, 83(4), p.610. https://crl.acrl.org/index.php/crl/article/download/24257/33409 keller, a. 2015. research support in australian university libraries: an outsider view. australian academic and research libraries, 46(2), 73–85. https://doi.org/10.1080/00048623.2015.1009528 library guides: tools for research: research data management. (2022, march 2). library guides at university of washington libraries. https://guides.lib.uw.edu/research/tools/rdm momoh, e.o. and folorunso, a.l., 2019. the evolving roles of libraries and librarians in the 21st century. library philosophy and practice (e-journal). https://www.academia.edu/download/60614755/fulltext20190916-78160-m1rokp.pdf moruf, h.a. and dangani, b.u., 2020. emerging library technology trends in academic environmentan updated review. science world journal, 15(3), pp.13-18. https://www.ajol.info/index.php/swj/article/view/202961 okagbue, h.i., opanuga, a.a., oguntunde, p.e., adamu, p.i., iroham, c.o. and adebayo, a.o., 2018. research output analysis for universities of technology in nigeria. int. j. educ. info. technol, 12, pp.105-109. . https://www.researchgate.net/profile/hilaryokagbue/publication/328364925_ oluwasanu, m.m., atara, n., balogun, w., awolude, o., kotila, o., aniagwu, t., adejumo, p., oyedele, o.o., ogun, m., arinola, g. and babalola, c.p., 2019. causes and remedies for low research productivity among postgraduate scholars and early career researchers on non-communicable diseases in nigeria. bmc research notes, 12(1), pp.1-6. https://link.springer.com/article/10.1186/s13104-019-4458-y. orji, s. and anunobi, c.v., 2019. citation frequency of research output of academic librarians in federal universities in south nigeria using google scholar. journal of library and information sciences, 7(2), pp.1-9. http://jlisnet.com/journals/jlis/vol_7_no_2_december_2019/1.pdf . owolabi, k.a., okorie, n.c., yemi-peters, o.e., oyetola, s.o., bello, t.o. and oladokun, b.d., 2022. readiness of academic librarians towards the use of robotic technologies in nigerian https://doi.org/10.29173/iq1069 https://doi.org/10.29242/report.emergingtech2020.interviews https://digitalcommons.unl.edu/libphilprac/5515/ https://www.educba.com/data-mining-tool/ https://crl.acrl.org/index.php/crl/article/download/24257/33409 https://doi.org/10.1080/00048623.2015.1009528 https://guides.lib.uw.edu/research/tools/rdm https://www.academia.edu/download/60614755/fulltext20190916-78160-m1rokp.pdf https://www.ajol.info/index.php/swj/article/view/202961 https://www.researchgate.net/profile/hilary-okagbue/publication/328364925_ https://www.researchgate.net/profile/hilary-okagbue/publication/328364925_ https://link.springer.com/article/10.1186/s13104-019-4458-y http://jlisnet.com/journals/jlis/vol_7_no_2_december_2019/1.pdf 8/8 adeyeye, sophia vivian.& oladokun, taofeek abiodun (2023) application of emerging technologies for research support in nigerian academic libraries: trends, problems and prospects, iassist quarterly 47(3-4), pp1-8. doi: https://doi.org/10.29173/iq1069 university libraries. library management. vol. 43 no. 3/4, pp. 296305. https://doi.org/10.1108/lm-11-2021-0104 panda, s. and chakravarty, r., 2022. adapting intelligent information services in libraries: a case of smart ai chatbots. library hi tech news. vol. 39 no. 1, pp. 1215. https://doi.org/10.1108/lhtn-11-2021-0081. rotimi adesina egunjobi, d., ogunniyi, s.o. and ajakaye, j.e., 2022. challenges and prospects of reference services in federal university libraries in south-west, nigeria.library philosophy and practice electronic journal 7155 https://digitalcommons.unl.edu/libphilprac/7155/ rotolo, d., hicks, d. and martin, b.r., 2015. what is an emerging technology?. research policy, 44(10), pp.1827-1843. https://doi.org/10.1016/j.respol.2015.06.006 saibakumo, w.t., 2021. awareness and acceptance of emerging technologies for extended information service delivery in academic libraries in nigeria. library philosophy and practice, pp.1-11. https://digitalcommons.unl.edu/libphilprac/5266/ sewell, c. and kingsley, d., 2017. developing the 21st century academic librarian: the research support ambassador programme. new review of academic librarianship, 23(2-3), pp.148-158. https://www.tandfonline.com/doi/pdf/10.1080/13614533.2017.1323766 sewell, c. and kingsley, d., 2017. developing the 21st century academic librarian: the research support ambassador programme. new review of academic librarianship, 23(2-3), pp.148-158. https://www.tandfonline.com/doi/pdf/10.1080/13614533.2017.1323766 sheik, m., & olugbenga, c. o. (2019). a study on emerging technology trends in academic libraries: an overview. icrlit–2019: e-proceedings on reshaping of librarianship, innovations and transformation, 163-168. wan, s., 2022, march. developing an engati-based library chatbot to improve reference services. in innovation and experiential learning in academic libraries: meeting the needs of twenty-first century students . maryland: rowman and littlefield https://doi.org/10.29173/iq1069 https://doi.org/10.1108/lm-11-2021-0104 https://doi.org/10.1108/lhtn-11-2021-0081 https://digitalcommons.unl.edu/libphilprac/7155/ https://doi.org/10.1016/j.respol.2015.06.006 https://digitalcommons.unl.edu/libphilprac/5266/ https://www.tandfonline.com/doi/pdf/10.1080/13614533.2017.1323766 https://www.tandfonline.com/doi/pdf/10.1080/13614533.2017.1323766 network management of machinereadable data file records by mari 1 yn nasat i r dubl in , ohio users of machi nereadab le data files (mrdf) bear the responsibility to share information about data files as well as to share access to data files. anyone needing a particular data file would find it useful to be able to sit down at a terminal and call up a record of that file by author or title or a variety of control numbers, find out where that data file is and at the same time find out how to access the file, wherever it happens to be. this presumes that someone has entered that record into the database in order to be able to retrieve it. the bibliographic networks such as oclc (online computer library center), rl i n (research libraries i nformat ion network) , wln (washington library network), utlas (university of toronto library automated systems) and n0ti5 (northwestern online total information system) provide a mechanism for participants to share information resources while the networks take care of storing, maintaining and managing machinereadable records of them. at present, these records are bibliographic descriptions of books, serials, audiovisual media, maps, manuscripts, music scores, and sound recordings, making each item not only uniquely identifiable but also locatable for loan or purchase by means of institutional holdings symbols. at least one bibliographic network--oclc--wi 1 1 be incorporating records of mrdf into its bibliographic database to facilitate utilization of machine-readable data. the two primary considerations in this important system enhancement are 1) the applicability of existing standards for describing and accessing mrdf records and 2) the feasibility of integrating mrdf records with records of the traditional information sources which comprise a bibliographic network database. standards for describing and accessing mrdf the need for standards to enable libraries, archives, and other information based agencies to cooperate in the description and access of machine-readable data files already being collected by academic and other institutions was demonstrated over a decade ago when the descriptive cataloging committee of ala's cataloging and classification section established the subcommittee on rules for mrdf. over a period of years the subcommittee drafted position papers making recommendations on cataloging mrdf which became the basis of chapter nine ("machine-readable data files") of the second edition of anglo-american cataloging rules (aacr2) .^ aacr2 attempts to integrate cataloging of mrdf with all other types of materials. the subcommittee's recommendations supplemented by recommendations made by the joint steering committee of aacr2 were followed in 1978 by the national conference on cataloging and information services for mrdf, organized to develop standards for bibliographic control of mrdf. two of the recommendations of the conference were j) the design of a marc (machine-readable cataloging) format for mrdf and 2) the definition of products and services to be derived from cataloging mrdf.^ in response to the first conference recommendation, the network development office of the library of congress (lc) produced machine-readable data files: a marc format (mrdf/marc) . -^ the us/marc formats are standards for representing bibliographic and authority information in machinereadable form. an individual marc format is a set of conventions for encoding a particular type of machine-readable record. the structure of a us/marc record is determined by the ameri can nat ional standard for information interchange on magnetic tape (ansi z39-2 1979) and by documentation format for bibliographic information interchange on magnetic tape (iso 2709 1973) • the content of certain data elements in a marc record is speci f i ed by the us/marc formats, while the content of traditional catalog elements is defined by other guidelines such as aacr2. in response to the second recommendation, the inter-university consortium for political and social research (icpsr) at the university of michigan has created an automated cataloging system for mrdf from which such products as the guide to icpsr's resources, facsimilies of catalog cards, bibliographic citations, indexes, and authority lists of authors and titles will be derived.^ the system is still in the development stage and is not yet available to the public. other agencies have also derived products and services from cataloging mrdf. at the federal level, standards were developed for bibliographic citations and abstracts in the production of directories of mrdf. two examples of applications of these standards are the directory of data files issued by the bureau of the census and american national standard for computer program abstracts (ansi x3.88 1981). the existence of stanards thus demonstrated, the question of applicability must be addressed. in the intervening years since the adoption of aacr2, microcomputer and video game software have emerged which were not considered either in aacr2 or in mrdf/marc. it seems more appropriate to extend the definition of mrdf to include microcomputer software than to violate input standards of the bibliographic networks by forcing these materials indiscriminately (and un ret r ievably ) into the network formats for books or audiovisual media which were not constructed to handle them. therefore, since input standards are rule-driven, cataloqers and policy makers have been meeting to propose solutions to the problem of rules for physical description of microcomputer software. a special cataloging task force has been formed with the charge to develop a statement regarding mi crocofiiputer software description, to be published in lc's cataloging service bulletin and to constitute national policy until aacr2 can be revised. this policy will be reflected in modifications to mrdf/marc. for assistance in the description of mrdf, catalogers and data librarians will have at their disposal sue dodd's new cataloging machine-readable data files: an interpretive manual^ as well as a manual of microcomputer software cataloging examples which is expected to be published by the minnesota aacr2 training series. integrating a hrdf fopjiat into a bibliographic network the feasibility of integrating mrdf records into the database of a bibliographic network, which in netowrk terminology is called implementing a mrdf format, is dependent upon 1) the existence of data elements in a bibliographic record which are appropriate to mrdf users, 2) network systems and services already in place which meet the needs of mrdf users, 3) products which can be generated from these systems and services, and h) a viable mechanism for extending treatment of traditional information sources to new formats. data elements data elements present in a network bibliographic record needed by mrdf users are as follows: a data file description showing existence and source of data; a detailed abstract which includes the genesis and history of the file so as to link modified files; a keyword structure; links between data files and the software created to manage or operate them, including the presence and bibliographic citation of accompanying documentation, software compatibility, and linkage with other files or programs; and the applicability of the data to solving specific problems or analytic needs . dodd has grouped these data elements for mrdf into six categories, or levels. level 1, bibliographic identity, includes such bibliographic elements as author, title, edition, distributor, notes, etc. level 2, data abstract, is a descriptive summary, abstract, or subject analysis of the contents of the file. level 3i classification, provides the classification codes, indexing, or descriptors. level k, technical in.'ormation for access, spells out the physical characteristics needed to access the mrdf such as recording density or software compatibility. level 5, analysis or use, gives a citation of documentation and related reports as well as such data collection information as how/when the data were collected, the unit of analysis, and sampling procedures. finally, level 6, archiving or maintaining, indicates records on processing, storage, use, and modifications to the mrdf.^ while this may sound like more information than could be incorporated into a single record, a bibliographic record currently provided by oclc, and presumably by the other networks, can accommodate over aooo characters, considerably in excess of the approximately 2500 required, according to dodd, for an expanded bibliographic mrdf record. systems and services the data elements just discussed, along with numerous others, form the basis of the systems and services already offered in different combinations by the bibliographic networks for managing, with appropriate variations, records of books, serials, audiovisual media, maps, manuscripts, music scores and sound recordings. prominent among these systems and services are online union catalogs; shared cataloging; serials control; subject retrieval; patron access; interlibrary loan (ill); authority control, services for converting existing bibliographic and location data into machine-readable form and adding them to the online union catalog; and acquisitions, ordering, and fund accounting which currently interface with a directory of libraries, publishers and vendors, soon to interface directly with the vendors themselves, and easily extendable to producers and distributors of mrdf. this powerful combination of services will provide mrdf users with a systematic way of identifying and locating mrdf for in-house control as well as for acquiring and accessing mrdf from external sources . additional services not to be overlooked are the training programs, workshops, and materials provided both by the instructional coordinators of the bibliographic networks themselves and by the independent contracting network offices through which most institutions participate. it is interesting to note that the independent networks that contract with oclc to provide services to their member libraries are already getting requests for the mrdf format from a variety of sources. government and research institutions are of course spurred on by the i98o census. data archives are underutilized and eager to add their holdings to the oclc database. "silicon valley" type organizations have a wide range of mrdf to control. academic and public libraries, schools, and media centers, stimulated by the microcomputer craze, are asking to use oclc for cataloging, ill (especially for documentation), acquisitions, and collection development in the lively area of microcomputer software, games, and instructional materials as well as mrdf. in fact, dialog (available through oclc's affiliated online services) is adding the international software directory and the microcomputer index to its long list of databases. products a mrdf record unambiguously tagged in an online union catalog can serve as an organic record from which catalog cards, accessions lists, and institution-specific archive tapes can be generated just as they are for other types of materials. the potential exists for such additional products as union lists, bibliographic citations, catalog entries, data abstr-'cts, card or tape distribution of mrdf records, and subject access either printed ur on microform. '^ implementation of a mrdf format although all of the bibliographic networks have the capability of implementing new formats in response to and support of lc policy and precedent, and despite mary magrega's paper describing the areas of activity which would be involved at utlas in implementing a mrdf format"* , oclc is the only bibliographic network which has actually made a commitment to implement a mrdf format and is already involved in the analysis iphase. analysis is the first of five stages of an oclc cataloging maintenance project, the developmental process decided upon as the most viable mechanism for expediting the mrdf implementation for two reasons. one is staff and resource allocation. the dther is that mrdf/marc will be distributed by lc as an update to its marc formats for bibliographic data (mfbd) on which many cataloging maintenance changes are oased.lo the product of the analysis phase will be functional specifications written oy a library systems analyst serving as product manager of a team of highly skilled catalogers, programmers, technical writers, quality control and testing staff. the functional specifications will provide the basis for the four remaining stages of the cataloging maintenance process: development of software, preparation '3f user documentation, creation of scenarios for testing the changes, and finally i" nstal lat ion of online system and related offline production changes. one of the benefits of installing the mrdf format as a cataloging maintenance project is that subsequent modifications distributed by lc through mrbd updates will be easily accomodated. conclusion a major goal of integrating mrdf into traditional library and information services by incorporating mrdf records into a bibliographic network is to communicate the availability of machinereadable materials and facilitate their utilization. it has been shown that existing standards for the description and access of mrdf, with minor modifications already in progress, are adequate for the bibliographic control of a wide range of machine-readable materials currently being collected. it has also been shown that data element descriptions, bibliographic network systems, services, and products, and a network mechanism for implementing new formats all are firmly established and easily applicable to mrdf ensuring the feasibility of the management of mrdf records by the bibliographic networks. what remains to be demonstrated is the mrdf users will avail themselves of this developing capability by contributing and accessing mrdf records by means of the various systems, services and products discussed. the vast numbers of inquiries already received suggest that they will. it is anticipated that this will be a cooperative venture in which the time and money spent by every user will be compensated for by benefits far outweighing the costs. references 1. american library association. anglo-american cataloguing rules, 2d ed. chicago: ala. 1978. 2. sue a. dodd. "toward integration of catalog records on social science machinereadable data files into existing bibliographic utilities: a commentary." library trends 30, n.3 (winter i982): 3^2. 3. sue a. dodd. working manual for cataloging machine-readable data files (chapel hill: social science data library, university of north carolina, 1976, mi meo) . ^4. "the us/marc formats: underlying principles." library of congress information bui let i n (in press). 5. carolyn geda of icpsr, private communication. 6. sue a. dodd. cataloging machine readable data files: an interpretive manual . 7dodd, "toward integration," pp. 350-351. 8. marilyn nasatir. "machine-readable data files and networks." informat ion technology and libraries (in press). 9. mary k. magrega. "a b ib 1 i ogranh i c utility's implementation of the marc format for machine-readable data files." (toronto: university of toronto library i automated systems, i98i, mimeo). 10. united states. library of congress. automated systems office. marc formats for bib] iograph i c data . washington: library of congress, i98o. icdbhss/83 international conference on data bases in the humanities and social sciences i the icdbhss/83 will be held in rutgers, the state university, on the douglas' college campus, at new brunswick, new jersey, from june 10 to 12. it is intended that the conference be broad in scope and that there be a true exchange of ideas and information", not only between the social sciences and the humanities, but also between the various fields in each area. speakers from twenty-four countries will present papers concerning data bases, their structure, accessibility, and uses in the fields of archaeology, art, artificial intelligence, criminology, education, history, law, lexicology, linguistics, literature, music, philosophy, sociology, and religion, to mention only the broader areas. most of the fields just mentioned extend over to other fields, e.g., arthistory, soc io-h i story , mus i c1 i brary , 1 i terary1 i ngui st i cs . the language of the conference will be english. for additional information contact: dr. i\obert allen bishop house, room 306a rutgers university new brunswick, n.j. o8903 ,, (201) 932-7335 or (201) 932-7505 index of public opinion poll questions compiled dennis gilbert at the university of louisville is compiling the american public opinion index, which will provide a tooical listing of all questions asked in american public opinion polls, some information about where to get the results of the survey, and some information about the survey in which the questions were asked. the first volume will include questions asked in 1951. the second volume will cover i982. by going beyond abstracts of polls to question level indexing, gilbert hooes to provide researchers with the kind of detail they need to directly access the particular topic in which they are interested. like any other index, apol will not give researchers substantive information, but instead will tell where to find it. consequently, )ol1ing agencies will retain control over their findings, especially those not yet )ubl i shed. for further information, write to dennis r.ilbert, 501 hiqhwood dr., louisville, rvy ^0206 or call (502) 893-252710 1/8 mushi, gilbert exaud (2021), research data management and services: resources for different data practitioners, iassist quarterly 45(3-4), pp. 1-8. doi: https://doi.org/10.29173/iq995 research data management and services: resources for different data practitioners gilbert exaud mushi1 abstract the emergence of data-driven research and demands for research data management (rdm) has created interest in global academic institutions and research organisations. some of the libraries, especially in developed countries, have started offering rdm services to their communities. although lagging, some academic libraries in developing countries are planning or implementing the service. however, the level of rdm awareness is deficient among researchers, librarians and other data practitioners. this paper aims to present available open resources for different data practitioners, particularly researchers and librarians. it includes training resources for researchers and librarians, data management plan (dmp) tool for researchers, a data repository available for researchers to freely archive and shares their research data to the local and international communities. a case study with a survey was conducted at the university of dodoma to identify relevant rdm services so that librarians could assist researchers in making their data accessible to the local and international community. the study findings revealed a low level of rdm awareness among researchers and librarians. over 50% of the respondent indicated their perceived knowledge as poor in the following rdm knowledge areas; dmp, data repository, long term digital preservation, funders rdm mandates, metadata standards describing data and general awareness of rdm. therefore, this paper presents available open resources for different data practitioners to improve rdm knowledge and boost the confidence of academic and research libraries in establishing the service. keywords research data management, open data, data management plan, data training, data repositories 1. introduction sustainable development and economies in the fourth industrial revolution are purely characterised driven by data-intensive research and innovation. the massive data generated through research, socialeconomic activities, and advanced technologies trigger the need to capture, integrate, and interpret new competitive knowledge. governments and federal granting agencies, especially in developed countries such as the national science foundation (nsf), national institute of health (nih) in the us, australian research data commons and e-science core programme in the uk, are championing this movement by mandating rdm practices in research organisations and academic institutions (chiware and mathe, 2016; tang and hu, 2019). all research organisations using public funds must manage their research data so that valuable data can be stored, shared and accessed for the long term. researchers must submit data management plans (dmps) alongside their proposals as a mandate from their organisations and research funding agencies. many research organisations and academic institutions in the us, uk, australia and other developed countries have somewhat experience offering data management services to their research community and internationally. national research foundation (nrf) in south africa mandates researchers and research institutions to effectively manage data during research and share valuable data to national and https://doi.org/10.29173/iq995 2/8 mushi, gilbert exaud (2021), research data management and services: resources for different data practitioners, iassist quarterly 45(3-4), pp. 1-8. doi: https://doi.org/10.29173/iq995 international communities (mushi, 2017b). having realised the potential benefits of data management and sharing, some of the research institutions established the rdm initiative without the mandates of the federal funding agencies or governments. data literacy plays a significant role in the growth of research and innovation, sustainable economic development and social well-being of society. thus, attracting local and international collaboration projects to solve global challenges such as food security and climate change. 2. statement problem rdm has been a subject of interest for over a decade, with several academic and research libraries providing and planning for the project. the tremendous growth of academic libraries offering rdm services has been reported, especially in the developed countries (cox et al., 2017). although lagging, research and educational institutions in developing nations plan to be part of this changing research environment. a case study conducted at the university of dodoma, tanzania, on identifying relevant rdm services for the library depicts a low awareness of rdm by researchers and the academic community. however, most researchers and postgraduate students were willing to practice rdm and share their research data with the local and international community. therefore, this paper presents available open resources for data practitioners, particularly researchers and librarians. these include training materials and online courses, dmps tools, and open data repositories for researchers to archive their valuable data freely. 3. literature review the new data-intensive and technology-based research environment jointly require different stakeholders' responsibilities, unlike a traditional research environment whereby in most cases, a researcher or group of researchers solely conduct research and manage data through its life cycle. it involves researchers, librarians, data analysts, ict personnel and more – it depends on the nature of research and organisation culture. researchers and librarians need to acquire skills (data literacy) and knowledge to navigate this new paradigm effectively. researchers need to present a data blueprint showing how data will be collected, stored, preserved and shared to meet research funders requirements (tenopir, birch and allard, 2012; tang and hu, 2019). librarians and researchers work together with librarians’ roles and responsibilities revolving around the research data life cycle. there are different librarian roles in each stage of the data life cycle which begins with planning. planning requires researchers to show a data roadmap such as how data will be collected, secured, stored, shared and reused. it involves the use of so-called data management plans (dmps). data collection follows after a project that requires librarians to ensure data are collected based on respective dmps using correct delimiters and consistent codes in data collection. furthermore, librarians assure the quality of the collected data and describe data by assigning proper metadata to enable data finding, understanding and re-usability. they are also responsible for data preservation and deposit to a trusted data repository while ensuring the use of the friendly format and support data discovery and reuse (tenopir, birch and allard, 2012; dataone, 2020). librarians have for a long time served researchers, especially with literature and information literacy skills. thus, making it a possible and favourite group to support researchers with rdm services. however, the most common setback for the librarians in this new role is the lack of rdm skills and knowledge (chiware and mathe, 2016; yoon and schultz, 2017). a study conducted in tanzania also observed low awareness https://doi.org/10.29173/iq995 3/8 mushi, gilbert exaud (2021), research data management and services: resources for different data practitioners, iassist quarterly 45(3-4), pp. 1-8. doi: https://doi.org/10.29173/iq995 of the rdm practices and associated benefits (mushi, 2017). mushi recommended advocacy of the services and skills needed to researchers and librarians through seminars, workshops and training. 4. methodology a case study used a purposive sampling technique to interview two (2) university managers (director of research and publications and director of library services). the study aimed to gather information on the managers’ perception and future of rdm services in academic libraries. using the snowball sampling technique, a researcher collected data from six (6) researchers and six (6) postgraduate students who could have experience with international funders’ mandates when applying for research grants. data were collected from 12 respondents from the university of dodoma using google forms online survey. the majority of respondents had never had encountered any rdm practice mandate because of the use of public funds from the government channelled to the university. some other researchers were selffinancing their research projects (mushi, 2017a). 5. results and discussion 5.1 rdm awareness and skills the study observed the low awareness of rdm practices among researchers, with more than 88% of researchers storing and handling research data in their devices. they also rated their rdm skills and knowledge as poor. they suggested training interventions in different rdm skills areas, including a guide on using dmps, a guideline for data appraisal, trusted data repositories, and other rdm practices presented in figure 1 below. figure 1. training asked for by the researchers (source: mushi, 2017b) 5.2 rdm implementation model building from jones et al. (2013), the study developed an rdm implementation model for implementing rdm services in academic libraries. this model has four implementation phases whereby phase one is on ‘strategy, policy, procedures and infrastructure’. the phrase stands as a master plan for sustainable and effective rdm services at the institution. it involves developing a policy that specifies the responsibilities of different stakeholders in the organisation, developing rdm guidelines and necessary infrastructure. phase two is about ‘awareness creation, skill development and repository content’. phase three insists on ‘management of active data’ while phase four is about ‘data selection and preservation. librarians, policymakers, and researchers can read about this four-phase rdm implementation model in the article https://doi.org/10.29173/iq995 4/8 mushi, gilbert exaud (2021), research data management and services: resources for different data practitioners, iassist quarterly 45(3-4), pp. 1-8. doi: https://doi.org/10.29173/iq995 entitled ‘identifying relevant research data management services for the library at the university of dodoma, tanzania’ (mushi, pienaar and deventer, 2020). 5.3 identified rdm resources for different data practitioners the study objective also aimed to identify different rdm resources available that can be used by various data practitioners when managing data through its life cycle. it includes training resources, online tools and forums that researchers can use, librarians, ict personnel supporting rdm, senior managers, research funding agencies and other research stakeholders as described below. 5.3.1 rdm training materials and courses several institutions and organisations offer training materials to improve data literacy among data practitioners, especially researchers and librarians. the materials address the common challenge of lack of skills and knowledge to manage the new role of this exciting subject. macdonald and rice (2013) argued that librarians need to acquire new skills to build capacity for supporting rdm services at the institutions. these include the following: mantra free online course. it provides online materials and training on rdm skills targeting graduate, postgraduate students, early researchers and professionals from different research disciplines. it offers training on various modules covering data services through its life cycle, including dmp, organising data, files format and transformation, documentation and metadata citation, storage and security, data preservation, rights and access, sharing protection and licensing. edinburgh university of the uk maintains this online training programme. training materials are licensed under the creative commons attribute, which means they can be used and shared while constantly acknowledging the source. access: https://mantra.edina.ac.uk/ dcc training and reference materials. free online training offered by the uk experienced trainers indepth subject experts in digital curation and research data management. it has two training approaches; one is an inclusive training event for all, and the second is customised private training requested by the individual or an organisation. dcc can tailor the latter approach to meet the requirements of the requesting individual or organisation. it targets different data practitioners and research stakeholders from institutional managers, researchers, students, professionals, librarians and other research support staff. digital curation center (dcc) of the uk maintains the training resources. access: http://www.dcc.ac.uk/training university of minnesota libraries consultation and workshop. apart from providing data management consulting services to its community, it offers learning materials in various formats such as templates, slides, and videos, which improves skills in general data management, developing a dmp, data sharing, access and ownership, and organising and storing data. resources are helpful to all research staff dealing with data and supportive services. access: https://www.lib.umn.edu/datamanagement/workshops supporting data management infrastructure for humanities (sudamih). sudamih is an rdm support infrastructure for the humanities helping institutions to manage research data effectively addressing technical requirements. it facilitates the creation of online databases to collect data in humanities and support different data services such as text, image, and geo-data. although materials are meant to be used in uk based institutions, they are also valuable for other institutions wishing to implement rdm services in humanities. the project is maintained by the oxford university of the uk and funded by the joint information system committee (jisc). https://doi.org/10.29173/iq995 https://mantra.edina.ac.uk/ http://www.dcc.ac.uk/training https://www.lib.umn.edu/datamanagement/workshops 5/8 mushi, gilbert exaud (2021), research data management and services: resources for different data practitioners, iassist quarterly 45(3-4), pp. 1-8. doi: https://doi.org/10.29173/iq995 access: http://sudamih.oucs.ox.ac.uk/index.xml rdmrose. the training project was for information professionals from uk leading ischool targeting lis educators and provide continued professional development (cpd) learning materials in rdm. although this project was closed, training materials remained available and valuable for information professionals planning rdm services. access: https://www.slideshare.net/rdmrose essentials 4 data support. provide support and training courses on data storing, managing, archiving and sharing. the site is maintained by research data netherlands (rdnl) and offers various courses suitable for librarians, archivists, researchers, managers and related research stakeholders in rdm services. access: https://researchdata.nl/ digital humanities data curation (dhdc). provides different resources, including articles and rdm guides which are helpful to researchers and librarians in humanities. resources provided in this platform are published under the creative commons no-commercial attribution license to encourage broad access and reuse of globally educational resources. access: https://guide.dhcuration.org/about/ coursera – research data management and sharing. provides opportunities for data practitioners such as librarians and researchers to join their platform and access training courses and materials freely. financial aids are available for some individuals who could meet the criterion provided for the course participants. the university north caroline jointly offers the courses at chapel hill and the university of edinburgh. access: https://www.coursera.org/learn/data-management new england collaborative data management curriculum (necdmc). it is an educational and rdm training resource that teaches rdm best practices explicitly to students and researchers in health sciences and engineering. although the teaching curriculum is designed based on nsf data plans recommendations, the materials are still valid and valuable to other institutions around the globe as it addresses many common challenges of rdm services. materials are published under creative commons attribution-non-commercial share-alike, giving freedom to reuse, share, and modify while acknowledging the source. the platform is maintained by lamar soutter library of massachusetts medical school in partnership with other libraries in the new england region. access: https://library.umassmed.edu/resources/necdmc/index cessdafree online expert tour guide to data management. offers training materials that aim at imparting skills in managing data in all stages of the research data life cycle. materials are organised in different formats, such as texts and videos. furthermore, the platform offers a guide to researchers on managing data and making their data fair (findable, accessible, interoperable and reusable). it is designed and maintained by european experts and is mainly on social sciences research datasets. access: https://www.cessda.eu/training/training-resources/library/data-management-expert-guide uk data archive. experienced platform offering data management resources, including training and publications of all matters related to research data. it provides opportunities for librarians and researchers to gain and improve their skills in rdm services. the long experience of training in rdm led to the https://doi.org/10.29173/iq995 http://sudamih.oucs.ox.ac.uk/index.xml https://www.slideshare.net/rdmrose https://researchdata.nl/ https://guide.dhcuration.org/about/ https://www.coursera.org/learn/data-management https://library.umassmed.edu/resources/necdmc/index https://www.cessda.eu/training/training-resources/library/data-management-expert-guide 6/8 mushi, gilbert exaud (2021), research data management and services: resources for different data practitioners, iassist quarterly 45(3-4), pp. 1-8. doi: https://doi.org/10.29173/iq995 handbook published by sage publication ltd. in 2014 which is entitled “managing and sharing research data: a guide to good practices”. access: https://www.ukdataservice.ac.uk/manage-data/training.aspx dataone training resources. it is a freely available education resource for librarians and researchers on rdm. a comprehensive package covering all significant aspects of data management, general under the creative commons zero license (cc0), gives users enormous freedom to use, share, and edit even without acknowledging the sources. resources that are accessible include sample assignments used for seminars or workshops, slides and handouts are covering essential rdm services, dmp, policy issues, metadata and data protection. the us national science foundation funds the platform (nsf). access: https://www.dataone.org/education-modules mit libraries rdm workshops. mit libraries offer many different workshops to build capacity in managing research data. the topics cover general rdm overview, dmp, data storage and sharing, file organisation and version control. the target group of these workshops includes postgraduate and graduate students, academic staff, researchers, librarians, and other teams working with data to improve their skills to manage their new roles. access: https://libraries.mit.edu/data-management/services/workshops/ 5.3.2 dmp tools the data management plan is an essential tool required by most research funders nowadays. it provides a roadmap of data through its life cycle while meeting standards set by the institution of funding organisation. librarians should have the skills to assist researchers in creating an impressive dmp for their research projects. most institutions already offering rdm services have developed their own dmp tool reflecting their requirements and that of their federal government agencies. however, researchers and institutions willing to practice data management can use dmp tools open for use by the research community globally. these tools are not tailored to meet a specific institution of funder requirements but address common issues of dmp that are accepted by many funders and institutions. some of these dmp tools have options for librarians and researchers to develop dmp based on particular funder requirements. moreover, other dmp tools such as dmponline maintained by digital curation center allow institutions to customise the template to provide tailored support for their users (chargers apply for institutions in this service). most dmp tools are free to use by individual researchers and librarians. below is the summary of dmp tools librarians and researchers can use to develop a data management plan (see table 1 below). table 1. data management plan tools s/n dmp tool organisation accessed at: 1 data management planning tool (dmptool) university of california digital curation centre of the california library https://dmptool.org/ 2 dmponline digital curation centre http://www.dcc.ac.uk/dmponline 3 perdue data curation profile toolkit perdue university libraries http://datacurationprofiles.org/ https://doi.org/10.29173/iq995 https://www.ukdataservice.ac.uk/manage-data/training.aspx https://www.dataone.org/education-modules https://libraries.mit.edu/data-management/services/workshops/ https://dmptool.org/ http://www.dcc.ac.uk/dmponline http://datacurationprofiles.org/ 7/8 mushi, gilbert exaud (2021), research data management and services: resources for different data practitioners, iassist quarterly 45(3-4), pp. 1-8. doi: https://doi.org/10.29173/iq995 5.3.3 data repositories there are also various data repositories offering free data archiving services for researchers. it provides opportunities for librarians and researchers to deposit their research data in a manner that can be fair (findable, accessible, interoperable, and reusable). librarians need to be aware of these resources to assist researchers in preparing data for archiving, especially assigning good metadata for easy discoverability of data. there are many discipline-specific public data repositories. however, this article focuses on generalist data repositories that researchers and institutions can archive research data (see table 2 below). table 2. generalist repositories s/n data repository name fees/cost accessed at: 1 zonedo free (donations are encouraged) http://help.zenodo.org/ 2 mendeley data free up to 10 gb datasets https://data.mendeley.com/ 3 dryad digital repository $120 usd for first 20 gb, and $50 usd for each additional 10 gb http://datadryad.org/ 4 figshare 100 gb free per scientific data manuscript. http://figshare.com/ 5 harvard dataverse free http://dataverse.harvard.edu/ 6 open science framework free http://osf.io/ source: (spinger nature, 2020) 6. conclusion all rdm resources discussed in this article provide opportunities for the libraries, researchers and other data champions to improve their data literacy and begin practising and offering rdm services in a limited resource environment of academic and research libraries. the resources address many data management challenges, particularly the required skills of researchers, librarians and research support staff. developing rdm services in the institution requires many resources more than that of an institutional repository. using dmp tools and data repositories services offered by other institutions encourages researchers and institutions, especially from developing countries, to be part of the changing research environment which is more data-intensive and collaborative. the discussed resources are essential to increase data literacy which will eventually stimulate data management practices for the sustainable development of societies. https://doi.org/10.29173/iq995 http://help.zenodo.org/ https://data.mendeley.com/ http://datadryad.org/ http://figshare.com/ http://dataverse.harvard.edu/ http://osf.io/ 8/8 mushi, gilbert exaud (2021), research data management and services: resources for different data practitioners, iassist quarterly 45(3-4), pp. 1-8. doi: https://doi.org/10.29173/iq995 7. reference chiware, e. r. t. and mathe, z. (2016) ‘academic libraries’ role in research data management services: a south african perspective’, south african journal of libraries and information science, 81(2), pp. 1–10. doi: https://doi.org/10.7553/81-2-1563. cox, a. m. et al. (2017) ‘developments in research data management in academic libraries: towards an understanding of research data service maturity.’, journal of the association for information science and technology, 68(9), pp. 2182-2200. doi: https://doi.org/10.1002/asi.23781. dataone (2020) best practices, dataone: data observation network for earth. available at: https://www.dataone.org/best-practices (accessed: 23 april 2020). jones, s., pryor, g. and whyte, a. (2013) ‘how to develop research data management services a guide for heis’, digital curation centre, (march), pp. 1–22. available at: http://www.dcc.ac.uk/resources/how-guides. macdonald, s. and rice, r. (2013) ‘diy’ research data management training kit for librarians, digcurv: lifelong learning programme. available at: http://datalib.edina.ac.uk/mantra/libtraining.html. mushi, g. e. (2017a) ‘identifying and implementing relevant research data management services for the library at the university of dodoma, tanzania [data set].’ available at: https://zenodo.org/record/3553837. mushi, g. e. (2017b) identifying relevant research data management services for the library at university of dodoma, tanzania. the university of pretoria. mushi, g. e., pienaar, h. and deventer, m. van (2020) ‘identifying and implementing relevant research data management services for the library at the university of dodoma, tanzania’, data science journal, 19(1), pp. 1–9. doi: https://doi.org/10.5334/dsj-2020-001. spinger nature (2020) recommended repositories. available at: https://www.nature.com/sdata/policies/repositories#general (accessed: 29 april 2020). tang, r. and hu, z. (2019) ‘providing research data management (rdm) services in libraries: preparedness, roles, challenges, and training for rdm practice’, data and information management, 3(2), pp. 84–101. doi: https://doi.org/10.2478/dim-2019-0009. tenopir, c., birch, b. and allard, s. (2012) ‘academic libraries and data services in academic libraries, college & research libraries, 46(3), pp. 61–75. available at: http://www.ala.org/acrl/sites/ala.org.acrl/files/content/publications/whitepapers/tenopir_birc h_allard.pdf. yoon, a. and schultz, t. (2017) ‘research data management services in academic libraries in the us : a content analysis of libraries ’ websites’, college & research libraries, 78(7), pp. 1–19. available at: https://crl.acrl.org/index.php/crl/rt/printerfriendly/16788/18346. endnotes 1 gilbert exaud mushi is an assistant lecturer and a librarian at the sokoine university of agriculture in tanzania. he works in the directorate of sokoine national agricultural library and can be reached at gilbert.mushi@sua.ac.tz. https://doi.org/10.29173/iq995 https://doi.org/10.7553/81-2-1563 https://doi.org/10.1002/asi.23781 https://www.dataone.org/best-practices http://www.dcc.ac.uk/resources/how-guides http://datalib.edina.ac.uk/mantra/libtraining.html https://zenodo.org/record/3553837 https://doi.org/10.5334/dsj-2020-001 https://www.nature.com/sdata/policies/repositories#general https://doi.org/10.2478/dim-2019-0009 http://www.ala.org/acrl/sites/ala.org.acrl/files/content/publications/whitepapers/tenopir_birch_allard.pdf http://www.ala.org/acrl/sites/ala.org.acrl/files/content/publications/whitepapers/tenopir_birch_allard.pdf https://crl.acrl.org/index.php/crl/rt/printerfriendly/16788/18346 mailto:gilbert.mushi@sua.ac.tz 14 iassist quarterly 2016 / vol 40 no 4 iassist quarterly abstract in recent years, the storage of qualitative data has been a challenge to data archives using repositories that are based on relational databases, as large files cannot really be represented well in these structures. most of the time, two or more structures have to be in place e.g. a fileserver that includes versioning for large files and a relational database for the tabular information. these structures necessitate the handling of multiple systems at the same time. with the arrival of hadoop and other big data technologies, qualitative data and quantitative data can now be stored as mixed mode data in the same structures. this paper will discuss our findings in developing an early prototype version of mmrepo at the university of applied sciences eastern switzerland htw chur. our prototype of mmrepo is a combination of the invenio portal solution from cern with a hadoop 2.0 cluster using the ddi 3.3 beta metadata scheme for data documentation. keywords mixed mode, qualitative, quantitative, big data, repository introduction storing different kinds of data from different domains and enhancing them with metadata has been the main workflow of research data centers, data archives or scientific repositories for many years. nevertheless, the usage of different file formats, metadata standards or it infrastructures is challenging for data managers and it managers in those facilities. mixed mode data – meaning data derived from qualitative research (e.g. open interview formats, ethnographic studies using video and audio) and quantitative research (e.g. questionnaires, cognitive tests) – which arise from the trend of combining different research designs (punch, 2009), poses a particularly significant challenge. data from qualitative research differ in size and structure very much from data derived from a quantitative design. this can be illustrated by the following example: an observation of a classroom full of students doing a computer-based test by filming high definition videos from multiple angles and recording different audio tracks. this is actually a mixed mode design as it combines a qualitative ethnographic study with a quantitative design (the computer-based test). the qualitative design will have as a result several gigabytes of video and audio files leading to processes like transcription for scientific processing. in contrast, the computerbased tests result in datasets that are often stored in formats for statistical packages (e.g. spss, stata, sas, r) and that contain variables, including variable and value labels. both types of data will be documented with different metadata standards with hardly any overlap between them due to the difference in domains. storage of qualitative data and quantitative data in repositories from an it perspective, the interesting question is this: can different research data types be stored within the same technical infrastructure (e.g. file servers, relational databases, data warehouses) to enable search functionalities for users within the same frontend across all different data types? the next chapter therefore looks at the current processes of storing these kind of data in selected repositories, gives examples and explanations on why different systems exist in the same organization, and explains the objectives and design considerations of a big data driven approach. mmrepo storing qualitative and quantitative data into one big data repository by ingo barkow1 catharina wasner2 fabian odoni3 vol 40 no 4 / iassist quarterly 2016 15 iassist quarterly the current state of the art in storing mixed mode data as mixed mode data vary considerably in size, documentation and type, storing quantitative data and qualitative data in one structure is a challenge. in most data archives, these are the most common ways to handle mixed mode data. 1.) storing both types in a relational database if a relational database is used as the data storage mechanism, the quantitative data can be ingested in its tabular format (e.g. by importing an excel table or spss file into a database table). the associated metadata could be stored in database tables as well using table joins or referential integrity to connect metadata and data thus allowing for variable shopping baskets or personal extracts (see amin et al., 2011). this means by using these features of some repository systems (e.g. questasy from centerdata) the user does not have to download a scientific use file (suf) and clean it from unnecessary variables but can choose in the portal which variables should be exported from the system. in many of the organizations connected to iassist or the ddi community (e.g. gesis, dipf, iab, centerdata) storing of tabular data in relational databases is therefore still a preferred way to document and store quantitative research. however, while a relational database is advantageous for strong quantitative data, it does not work well for qualitative data. typically, qualitative data would be stored in a relational database as a binary large object (blob) or an object of similar type directly in the table, which increases table size dramatically. the database would then be linked to the content of a file server. sometimes, hybrid technologies like file streams would be used (e.g. sql server 2014; see mistry and misner, 2014). all these technologies do not combine very well. relational database systems are not very good at handling blobs. there are limitations in size per single cell (usually 2gb – see sql server blob varbinary(max) datatype; mistry and misner, 2014) and the inflation of size normally leads to performance issues as database servers are optimized for handling small atomic data like short strings or numbers. the external linkage between relational database and file server also does not work very well as the two systems are separated. if users perform changes on the file server, the database server uses the information where the files have been stored, usually leading to dead links. file streams as a hybrid technology try to avoid this problem by letting the database server handle the file server automatically. this technology is only available in enterprise database servers like sql server or oracle. unfortunately, file streams have some technical disadvantages like e.g. the data outside of the table will not be part of the backup environment of the database server anymore. furthermore, the outsourced data on the external file server is not a part of a database transaction (meaning the database server will not be able to roll back a transaction in case of a technical problem or break-off from the user side). in summary, filestreams fix the problem of broken links between file servers and database servers, but do not include the features relational databases are known and used for. 2.) storing all data as files on a file server the other option would be to handle quantitative data and qualitative data in their file-based state on a file server. metadata would be provided one of three ways: by an external relational database, be attached as attributes of files, or by adding additional files that contain the metadata information. while this is a good way to handle qualitative data, the advantages of processing structured tabular information of quantitative data within a relational database are lost. in particular, quantitative data is simply processed in its file form so users can download it, while advanced features like variable shopping basket, personal extracts, and search functionalities for variables or basic tabulation from a portal solution would be heavily limited. advantages of storing mixed mode research data in the same technical infrastructure when talking about the advantages of storing mixed mode research data in one infrastructure first the question has to be answered why it can be problematic to have different storage structures for qualitative and quantitative data. from an it perspective two different kinds of repositories also mean handling two times a complete set of hardware and software infrastructure. this means the it administration, it support and software development have a much higher effort, as this can be completely separate systems (different servers, different operating systems, different repository software, different frontend). if the user is supposed to be provided with one portal solution to search through two or even more repositories (basing on different kinds of data) a meta search solution has to be provided. this means the user types in a search request and the meta search solution divides this into different search requests running on the different back ends, collecting the results and collating them into one common result. consolidating a multitude of different systems is a huge effort and a simpler one stop solution in the backend would lead to huge advantages as only one system has to be catered from hardware and software side. examples from data centers to get a clearer picture of how data repositories handle different kinds of data, two organizations from germany were selected as examples – the german institute for international educational research (dipf)4 and the leibniz institute for social sciences (gesis)5. both institutions archive and distribute research data and publications and run one or more research data centers accredited by the german data council (ratswd)6 dipf is currently running three different repositories for different purposes (see bambey et. al. 2013 and bambey et. al. 2012): qualitative data (e.g. school observations in video) are stored in the medienarchiv7 questionnaires and answer schemes are stored in the database for quality of schools (daqs)8 data documentation regarding the framework program for educational research in germany (over 300 projects) is stored in the metadata database of the verbund forschungsdaten bildung9 those three separate repository systems are derived from former projects and are run by three different organizations – dipf, gesis and iqb.10 from a technological point of view they are completely different and optimized for their respective content: qualitative data, quantitative data, and metadata on quantitative and qualitative data. 16 iassist quarterly 2016 / vol 40 no 4 iassist quarterly a similar approach can be seen at gesis where the following structures can be found: quantitative data, questionnaires and study documentation are stored in the data catalogue (dbk)11 variable-level information is stored in the zacat portal12 metadata on microdata are stored in the missy system13 full-text social science documents are stored in the social science open access repository (ssoar)14 historical studies and time series are stored in histat15 time series data from social indicators are stored in the online information system simon16 the diversity of gesis systems is due to independent developments in different departments but as well to the requirements of different target groups and diverse digital resources. an integrated search function across all sources is under development.17 the diversity of repository systems in dipf and gesis developed organically with the organizations over the years. this pattern can be seen in many similarly-sized institutions around the world. the development of integrated repository systems was constrained by the previously mentioned limitations of relational database and file server data storage systems. the separation of storage based on data types is therefore valid and has to be seen as a product of the times in which they were developed. many of the repository systems have been running for several years or in some cases even decades. vision and objectives of a big data driven repository nevertheless, the question remains whether all research data can be unified in one system by using more modern approaches, not least because the it administration of multiple existing systems alone takes up many resources. one candidate for this is big data technology. possible benefits of using big data technology to store different research data types include • the use of cluster based file systems. big data file systems like hdfs2 automatically split files and workload across multiple servers including redundant copies. this means it has inbuilt fail save and processing capabilities already from the software side (no need to use expensive hardware). • access to robust semantic search systems like solr and elasticsearch. as big data was originally developed for handling large unstructured data within search engines it comes with sophisticated search capabilities which go beyond a simple string matching, but also features understanding of underlying concepts in e.g. a full-text search to improve the results for the users. • applicability of text mining or natural language processing. as large quantities of unstructured data can be stored and processed in parallel big data offers the possibility to analyze these data with advanced syntactical or semantic methods. furthermore, big data technologies can manage unstructured data like qualitative data. having all data types in one system would be less costly and resource intensive because no meta search platform as described before has to be set up and only one system has to be developed and maintained. another advantage from user perspective is new methods of data analytics and data science can be used e.g. for analyzing data across different data types with the possibilities like text mining or natural language processing offered within the big data solution. these can be combined with classical statistical analysis from the sphere of social sciences and offer a scientific value add. design considerations of using big data as a unified repository big data solutions like hadoop18 were originally developed as search engine companies like yahoo or google were not able to store the masses of data needed to offer their services. while relational databases or data warehouses rely heavily on clear data structures, the design paradigms of big data technology are different. instead of having structured data on expensive cluster hardware, big data technology was designed to allow parallel processing of unstructured data on inexpensive hardware. this basis was extended over the years from a more file systems based approach like hadoop 1.0 to a multilayer platform as can be seen from figure 1. the addition of services is especially interesting in big data developments like hadoop 2.0. one additional service in hadoop 2.0 is hbase, a non-relational (nosql) and column-oriented database modeled after google’s bigtable19. it runs on figure 1 – the development from hadoop 1.0 to hadoop 2.0 (hortonworks 2013) vol 40 no 4 / iassist quarterly 2016 17 iassist quarterly top of hdfs2, the native file system of hadoop designed to store data among multiple computers, a cluster, by breaking up the data into blocks and distributing them throughout the cluster20. another additional service is hive which enables the use of sql-like queries on the cluster. hadoop 2.0 could be the technical basis of a unified repository, whereby tabular content like metadata or qualitative data can be stored within hive and large quantitative datafiles within hdfs2. this is the design idea which led to the mmrepo prototype project at the university of applied sciences eastern switzerland htw chur. mmrepo as a prototype of a unified repository as described in the chapter before, mmrepo was started as a prototype project to experiment with metadata, qualitative data and quantitative data within a single big data-based repository infrastructure. at the time of writing, mmrepo is a small project meant to test the performance and feasibility of the big data approach and must be considered a work in progress. the project started on january 1st, 2016 and its first phase ended on september 30th, 2016. the biggest obstacle in this respect was invenio 3.0 final is as of today not released yet (march 2017). this means a large part of the final frontend testing had to be postponed to a later phase of the project. the current plan of invenio specifies the release of the final version of 3.0 for summer 2017. the frontend testing will therefore be performed at a later date in the follow-up project called lifecyclelab and not be part of this paper. structural design of mmrepo the following test scenario has been set up to test the feasibility of the project. as the system’s backend, a hadoop 2.0 cluster is used with hbase and hdfs2 as services. this backend will be combined with an invenio21 3.0 beta frontend in the second project phase as the project is too small to develop portal functionalities by itself. the advantage of invenio in this context is that it offers a modular framework for repositories where the data storage can be exchanged with something else (in our case a big data cluster). furthermore, invenio offers advanced features like semantic search or versioning, which benefit especially qualitative data. the structural design can be seen in the following figure. 2 the quantitative data and qualitative data used in the project are test data from previous studies at htw chur plus sample data from dipf and the research data center (fdz) of the german federal employment agency (ba) at the institute for employment research (iab)22 . therefore, a variety of test data has been imported into hbase 2.0 and hdfs2. as the metadata schema for the quantitative data, data documentation initiative (ddi)23 lifecycle 3.3 beta is used. for documenting the qualitative data, several internal groups at htw are in favor of adopting the metadata encoding and transmission standard (mets)24 . however, the final decision has not been made yet. hardware layout of the mmrepo prototype testing as mmrepo uses hadoop 2.0 as its backend, for proper testing, it is necessary to have a cluster environment. only then, the advantages of parallel processing like the mapreduce algorithm can be exploited. unfortunately, the project is much too small to employ even the least expensive servers. to simulate a multitude of nodes, the decision was made to set up the system on small third generation raspberry pi microcomputers. raspberry pi 3 offers a quadcore 1.2 ghz arm processor with 1gb of ram plus sdhc card support and is an excellent less-costly solution for prototyping. a huge advantage for the hadoop solution is additionally the support of lan and wireless lan meaning two separated networks. for the current test setup the cluster network (lan) and the user access (wlan) can be separated as mixing the network traffic figure 2 – software structure of mmrepo 18 iassist quarterly 2016 / vol 40 no 4 iassist quarterly from users and between server nodes might influence the results by slowing down reaction times. from the hardware layout of big data solutions it does not make sense to test the capabilities by setting up e.g. two large powerful servers with inbuilt redudant hardware (e.g. multicore, multiple redudandant hard disks, fail safe networking) as the hadoop software layout rather demands a high number of cheap servers with inexpensive hardware which can operate in parallel. the redudancy and the parallel processing is provided by the big data software itself. we also decided against employing virtualization (e.g. vmware vsphere, citrix xenserver). if we tested the setup e.g. by setting up eight virtual servers in reality all virtual machines might end up on the same physical machine or might be shifted within several physical servers with different hardware layouts during the test (e.g. vmware – vmotion between nodes) thus influencing our results. a prototype setup using the cheapest possible hardware is therefore closer to the specifications hadoop was originally intended for. as we want to test primarily the search capabilities using an array of multiple servers we chose this setup for the prototype. in later phases of this project this prototype will be replaced by an array of inexpensive servers. results from the prototype in late autumn 2016 we started to test the backend of the prototype after setting up the hardware with a limited number of raspberry pi nodes (one master, seven slaves) after it was clear invenio 3.0 will not be released before the end of this project phase. we therefore decided on testing qualitative data and quantitative data on the backend of the cluster. installing hadoop on the inexpensive hardware was uneventful as there are in the meantime multiple how-tos to be found (e.g. from ibm25). some settings had to be modified manually in the environment variables, as raspberry pi is an unfamiliar hardware for hadoop, but the how-tos provided enough support. to test the setup and set the correct block size for hadoop the following datasets were implemented: us census data (2013_acssf_all_states_all_tables) o size: 469 mb o format: tables (*.csv) o usage in hadoop: tables stored in hbase (not relationally connected) wikipedia – us version (11/2016) o size: 49 gb o format: sql dump o usage in hadoop: relational database in hbase wikipedia – us version (06/2008) o size: 7.2 gb o format: static html dump o usage in hadoop: files in hbase the selection of example datasets shows a mix of qualitative and quantitative data, but also the first problem in the prototypical setup. the amount of data used does not qualify as big data in a classical sense (4 vs – volume, velocity, veracity, value – see e.g. meyer, 2013) as the volume is only several gigabytes while in big data repositories we rather aim at hundreds of terabytes or exabytes in the near future. nevertheless, as the cluster is far from powerful from a systems’ performance perspective the amount of data should be sufficient. to see if the hadoop cluster functions properly we started by implementing simple word count operations. the raspberry pi based cluster worked according to specifications, but ran very slowly. part of the performance could be optimized by changing the block sizes. also, the hadoop software is not optimized to the raspian operating system running underneath it, so the full performance of cpu and chipset is not used. as a first result it can be said the raspberry setup is interesting to explore the possibilities of hadoop on a real cluster especially for teaching purposes in an university environment, but it is not sufficient for performance testing or even productive setup. from a conceptual perspective the results were much more promising. the ideas of storing files into the cluster-based filesystem hdfs and the tables into the nosql database hbase worked fine. nosql database in this respect means “not only sql” a non-relational database layout which offers more flexibility in database layout while losing some transaction capabilities. the mysql database dump from wikipedia was imported into hbase by using sqoop. to see if the overall setup works we created search requests for hdfs2 and hbase using solr which is embedded into hadoop. as hbase is built on top of hdfs2 as a column-oriented non-relational database figure 3 – raspberry pi 3 (taken from www.raspberrypi.org) vol 40 no 4 / iassist quarterly 2016 19 iassist quarterly system it worked well with solr so essentially we were able to use one hardware platform, one software platform and one search engine for the purpose of searching through qualitative and quantitative data as a first step. nevertheless, there is a limitation. although solr searches through files and tables it does not mean the results are meaningful per se. in the end we get text extracts from documents and rows of tables as result sets which can be considered an intermediate step from a presentation point of view. a real frontend search solution needs adaptations in the presentation layer to have a user-friendly version of the results. currently the whole setup from the backend to the portal is very crude and can be considered a proof of concept for further funding, but not an out-of-the-box usable solution. nevertheless, the basic functionality is already available and therefore our project can continue. we are currently certain to be able to exchange our currently separated repositories into one big data solution, although the real development work on much better hardware has to start first. references amin, a., barkow, i., kramer, s., schiller, d. & williams, j. (2011). representing and utilizing ddi in relational databases. minnesota : ddi working paper series [doi:http://dx.doi.org/10.3886/ddiothertopics02]. bambey, doris; rittberger, marc (2013): das forschungsdatenzentrum (fdz) bildung des dipf: qualitative daten der empirischen bildungsforschung im kontext. standards und disziplinspezifische lösungen. in: huschka, denis; knoblauch, hubert; oellers, claudia; solga, heike (2013) (hrsg.): forschungsinfrastrukturen für die qualitative sozialforschung. berlin: scivero-verlag, s. 63-71. url: http://ratswd.de/dl/ downloads/forschungsinfrastrukturen_qualitative_ sozialforschung.pdf (20.03.2015). bambey, doris; reinhold, anke; rittberger, marc (2012): pädagogik und erziehungswissenschaft. in: neuroth, heike; strathmann, stefan; oßwald, achim; scheffel, regine; klump, jens; ludwig, jens (hrsg.): in: langzeitarchivierung von forschungsdaten. eine bestandsaufnahme. boizenburg: hülsbusch, s. 111-135. hortonworks (2013). apache hadoop patterns of use. andreas meier (2013). relationale und postrelationale datenbanken – leitfaden für die praxis, springer-verlag. mistry, ross and stacia misner (2014). introducing microsoft sql server 2014. punch, k. f. (2009). introduction to research methods in education. london: sage. notes 1. ingo barkow is an associate professor for data management at the university of applied sciences eastern switzerland htw chur and can be reached by email: ingo.barkow@htwchur.ch 2. catharina wasner is a research associate at the university of applied sciences eastern switzerland htw chur 3. fabian odoni is a research associate at the university of applied sciences eastern switzerland htw chur 4. http://www.dipf.de 5. http://www.gesis.org 6. http://www.ratswd.de 7. http://www.fachportal-paedagogik.de/forschungsdaten_bildung/medien.php?la=de 8. http://daqs.fachportal-paedagogik.de 9. http://www.forschungsdaten-bildung.de 10. https://www.iqb.hu-berlin.de 11. https://dbk.gesis.org 12. http://zacat.gesis.org 13. http://www.gesis.org/missy 14. http://www.ssoar.info 15. http://www.gesis.org/histat 16. http://gesis-simon.de 17. http://www.gesis.org/en/research/applied-computer-and-information-science/information-retrieval/ 18. http://hadoop.apache.org/ 19. http://www-01.ibm.com/software/data/infosphere/hadoop/hbase 20. http://www-01.ibm.com/software/data/infosphere/hadoop/hdfs 21 http://invenio-software.org 22 http://fdz.iab.de 23. http://www.ddialliance.org 24. http://www.loc.gov/standards/mets 25. alan verdugo building a hadoop cluster with raspberry pi https://developer.ibm.com/recipes/tutorials/ building-a-hadoop-cluster-with-raspberry-pi/#r_overview 1/11 akinyoola, oladoyin grace (2023) knowledge and perception of librarians towards cloud-based technology in academic libraries in southwest, nigeria, iassist quarterly 47(3-4), pp. 1-11. doi: https://doi.org/10.29173/iq1083 the creative commons-attribution-noncommercial license 4.0 international applies to all works published by iassist quarterly. authors will retain copyright of the work and full publishing rights. knowledge and perception of librarians towards cloud-based technology in academic libraries in southwest, nigeria oladoyin grace akinyoola abstract this paper investigated knowledge and perception of librarians towards cloud-based technology in academic libraries in southwest, nigeria. the population comprised all professional and nonprofessional librarians in academic libraries in southwest, nigeria. one hundred and thirty-two (132) librarians from southwest, nigeria were selected using simple random sampling technique. four research questions were answered. a structured questionnaire titled “qkplctlsn” was used for data collections. the reliability coefficient of the instrument yielded r = 0.76 and r =0.82. the findings of this research revealed that, librarians in academic libraries in southwest, nigeria have knowledge about cloud-based technology and they have been using one or more applications in academic libraries but their perception towards the use of cloud-based technology was negative. in view of this, the following recommendations were made; librarians should embark on staff development programs that would enable them keep pace with the latest technology, library management should encourage librarians to attend seminars, conferences and workshops that would enhance their technological skills, there should be adequate funding from the government and provision of stable power. keywords academic libraries, librarians, cloud-based technology. introduction the 3rd and 4th industrial revolutions ushered in change and growing of information at a tremendous speed due to information explosion brought by information technology (it). it brought about the emergence of the web-based technologies which is also known as cloud-based or cloud computing globalization of networks and internet. with the advent of it, more libraries become automated as the basic need towards advancement and many libraries now have virtual library units. further development in libraries is characterized by the emergence of e-books, e-journals, internet usage, web tools application, consortium, and so on. among the emerging area within the field of information technology is cloud based technology. cloud based technology is the movement from desktop application to web applications. the latest technology trend in the field of librarianship is the use of cloud-based technology for achieving library functions and services like managing collections and moving from desktop to webbased services https://doi.org/10.29173/iq1083 2/11 akinyoola, oladoyin grace (2023) knowledge and perception of librarians towards cloud-based technology in academic libraries in southwest, nigeria, iassist quarterly 47(3-4), pp. 1-11. doi: https://doi.org/10.29173/iq1083 to mention just few. it is, therefore, necessary for the professionals to be aware of it and be able to apply it in the library. cloud computing is a webbased technology which facilitates the sharing of resources and services over the internet rather than having these resources and services on desktops and local servers. kaushik and kumar (2013) described cloud-based technology as the combination of servers, networks, connections, applications and resources. kamba (2017) described it as a utility package that delivers computing as a service rather than as a product, shared resources and software applications. data access/retrieval and information storage are provided to networked computers without the user knowing the location or architecture of the computing infrastructure. cloud computing provides computing service over the internet which has completely changed the way one can use the power of computers irrespective of geographical location. cloud computing provides a shared pool of resources, including data storage space, networks, computer processing power, and specialized corporate and user applications. it helps peoples to access their e-mail, social networking site or photo services from anywhere in the world, at any time, at minimal or no cost (swapna and biradar, 2017). united states national institute of standards and technology (2012) defines cloud computing as a model for enabling convenient, on-demand network access to a shared pool of configurable computing resources like networks, servers, storage, applications, and services that can be rapidly provisioned and released with minimal management effort or service provider interaction. seena and sudhier (2013) defined cloud computing as a technology that enables the migration of desktop application to web-based applications such as communication tools like, gmail, google calendar, google talk and google+ and productivity tools like google documents such as text files, spreadsheets, and presentations. cloud-based technology services in the library: pros and cons academic libraries are now using cloud-based technology as data/information storage device in order to preserve their data and information so as to have access to them anywhere and anytime. academic libraries are libraries attached to post-secondary schools. they exist in universities and colleges of education, polytechnics and nursing schools. some notable cloud-based technology used in the academic libraries are online computer library center’s worldshare management services (wms) which allows libraries to manage entire collection management life cycle in a cloud-based application, google apps, which allows migration from desktop to web-accessible applications and storage of information and information resources, oss labs uses amazon’s elastic cloud computing platform in offering koha integrated library software and dspace institutional repository hosting and software maintenance subscription services for libraries, duraspace is a collaboration of the dspace digital library software and fedora commons, it is an open source repository application that allows capturing, storing, indexing, preservation, and distribution of digital materials including text, video, audio and data. dropbox is designed as an invisible application. it automatically backs up and syncrotinizes files across all devices and keeps files in the cloud so it can be accessed from computers anywhere in the world. they are not limited to these alone. library services have been experiencing transformation with the use of cloudbased technology application in many areas like, software application, data access, storage and retrieval, building digital https://doi.org/10.29173/iq1083 3/11 akinyoola, oladoyin grace (2023) knowledge and perception of librarians towards cloud-based technology in academic libraries in southwest, nigeria, iassist quarterly 47(3-4), pp. 1-11. doi: https://doi.org/10.29173/iq1083 library/repositories, file storage, library automatic and housekeeping activities, library websites hosting, building community power through social networking, searching library data and searching scholarly content (achugbue, (2018) and swapna and biradar, (2017). a number of authors including muhammad and abddullah (2021), annie (2004) lazarus and euchanu (2021) and achugbue (2018) claim that many librarians in nigeria are aware of cloud computing and that it could be accessed anywhere. cloudbased technology according to sahu (2015) has the capacity to increase and decrease hardware or software resources consumption. libraries only paid for what they used, storage is been controlled by the service provider hence, it can be adjusted according to the needs of the library and extend it to various services, resources could be accessed anywhere and test and evaluation of resources is possible, sharing of resources of one or more libraries is possible with online public access catalogue (opac). cloud-based technology also keeps software up-to-date that is, whenever the server is updated, any library using the service will be updated automatically. moreover, musungwini (2016) reported that googledocs can be used by academics in various ways; it has ability to work on files anywhere, anytime and provide quick feedback from people simultaneously. abidi and abidi (2012) maintained that using cloud-based technology minimizes library expenses as capital expenditure done on infrastructure will mainly be converted into operational expenditure. saving the time of the users is paramount in the library, hussaini, vishstha, garba and jimah (2017) affirmed that using cloud-based technology saves the time of the users and service providers. alhami and khaparde (2014) and mcmanas (2016) asserted that librarians have been using one or more cloudbased technology to access and store information for future usage regardless of locations. challenges/constraints with the usage of cloud-based technology challenges of cloud-based technology include the following: many libraries and individuals are scared of using cloudbased technology because of privacy and security issues. since libraries do not have ownership of the servers that housed them definitely information may be prone to theft and loss yuvaraj, (2015), hussaini, et.al (2017), agandi & gull (2013) and sahu, 2013). pal (2013) asserted that if data managed through cloudbased technology are lost, a library also lost physical or local back up. another challenge is unstable internet connection or inadequate internet subscriptions, it is not possible for academic libraries to implement cloud –based technology without adequate internet connection and subscription and all these require huge costs (nasir & bashis (2013), agandi & gull (2013), (swapna & biradar (2017), and hussaini, et.al (2017). cost is another issue in implementing cloud-based technology. hussaini et.al (2017) corroborated this, that budget constraint in some libraries is an impediment. funds are needed to purchase equipment and other infrastructural facilities for implementation but may not be available to some libraries. technical know-how on the part of librarians is also a major challenge; some librarians in africa still prefer traditional ways of doing things. some have phobia for using technology. this, to a greater extent, is an impediment to the use of cloud – based technology in academic libraries. statement of the problem as academic libraries continue to acquire resources almost every day in order to provide services for their users, it has been observed that traditional ways of providing services coupled with the use of desktop services are inadequate in this forth industrial revolution. more so, these types of practices have been https://doi.org/10.29173/iq1083 4/11 akinyoola, oladoyin grace (2023) knowledge and perception of librarians towards cloud-based technology in academic libraries in southwest, nigeria, iassist quarterly 47(3-4), pp. 1-11. doi: https://doi.org/10.29173/iq1083 affecting librarians because it is time consuming and duplication of efforts in the library. it is pertinent for librarians to keep pace with cloud – based technology in solving aforementioned challenges. cloud-based technology is seen as evolving paradigm with a lot of benefits that can assist librarians and non professional librarians to provide timeless services to users at any time. it offers great opportunities to librarians to keep and maintain records of data and also rendering efficient services. it is against this backdrop, the researcher investigates knowledge and perception of librarians towards cloud-based technology in academic libraries in southwest, nigeria. purpose of the study the main purpose of the study is to investigate knowledge and perception of librarians towards cloud – based technology in academic libraries in southwest, nigeria. research questions the following research questions were answered in this study. 1. what is the level of knowledge of librarians towards cloud-based technology in academic libraries in southwest, nigeria? 2. what is the perception of librarians towards cloud-based technology in academic libraries in southwest, nigeria? 3. what is the level of usage of cloud-based – technology by librarians in academic libraries in southwest, nigeria? 4. what are the challenges facing librarians towards the use of cloud – based technology in academic libraries in southwest, nigeria? methodology the descriptive survey method was adopted. the population of the study covered all librarians in public academic libraries in southwest, nigeria. it cut across both professional and para-professional librarians. simple random sampling technique was used to select one hundred and thirty-two (132) respondents. research instrument used for data collection was questionnaire. instrument titled “qkplctlsn” was used for data collection. the data collected were analyzed using simple percentage and frequency counts to indicate the number of respondents who strongly agree/agree to positive statements and disagree/strongly disagree to negative statements. the reliability co-efficient of the knowledge of cloud-based technology was pilot-tested among computer lecturers in emmanuel alayande college of education, oyo who were not part of the real respondents and it yielded r = 0.76 while the reliability of the perception yielded r = 0.82 using test-retest reliability. the raw data collected were later analysed using simple percentage. results this deals with the results of the analysis which will be discussed below. https://doi.org/10.29173/iq1083 5/11 akinyoola, oladoyin grace (2023) knowledge and perception of librarians towards cloud-based technology in academic libraries in southwest, nigeria, iassist quarterly 47(3-4), pp. 1-11. doi: https://doi.org/10.29173/iq1083 research question 1: what is the level of knowledge of librarians towards cloud-based technology in academic libraries in southwest, nigeria? information collected from respondents on librarians’ knowledge was subjected to descriptive analysis of simple percentage in order to answer research question1. table 1: this shows the results of the level of the knowledge of the librarians towards cloud-based technology s/n sa a d sd n % n % n % n % 1. i am aware of cloud-based technology 121 91.7 11 8.3 00 00 00 00 2. information can be stored using cloud based technology 128 97 04 3 00 00 00 00 3. i have seen many libraries using cloud based technology 126 95.5 06 4.5 00 00 00 00 4. cloud-based technology is easy to use and manage 101 76.5 31 23.5 00 00 00 00 5. information stored with cloud-based technology can be accessed anywhere and anytime. 130 98.5 02 1.5 00 00 00 00 table 1 above shows responses of librarians towards their knowledge of cloud-based technology. as could be seen in item in table 1, all the librarians have knowledge of cloud-based technology. this was indicated by 91.7% and 8.3% respectively. 97% and 30% of the librarians agreed that cloud-based technology can be used to store information. moreso, 95.5% and 4.5% believed strongly that many libraries in nigeria are now using cloud-based technology. the librarians also believed that the implementation of cloud-based technology is easy to use and manage. this was indicated by their responses of 76.5% and 23.5%. going through the table in item 5, 98.5% and 1.5% of librarians believed information stored with cloud-based technology can be accessed anywhere and anytime. the results revealed that the librarians have knowledge of cloud-based technology. research question 2: what is the perception of librarians towards cloud-based technology in academic libraries in southwest, nigeria? table 2: shows the results of the perception of the librarians towards cloud-based technology https://doi.org/10.29173/iq1083 6/11 akinyoola, oladoyin grace (2023) knowledge and perception of librarians towards cloud-based technology in academic libraries in southwest, nigeria, iassist quarterly 47(3-4), pp. 1-11. doi: https://doi.org/10.29173/iq1083 s/n sa a d sd n % n % n % n % 1. high-technical knowledge is required for cloud-based technology in libraries 101 76.5 28 21.2 03 2.3 00 00 2. i feel cloud-based technology can be used to build community power 93 70.5 39 29.5 00 00 00 00 3. i feel most academic libraries can afford the cost of using cloud-based technology. 42 31.8 24 18.2 66 50 00 00 4. cloud-based technology is easy to maintain and use in academic libraries 02 1.5 18 13.6 26 19.7 86 65.2 5. i feel cloud-based technology will enhance storage and preservation of information resources 89 67.4 29 22 14 10.6 00 00 6. cloud-based technology can be trusted 28 21.2 16 12.1 64 48.5 24 18.2 table 2 above shows the perception of librarians towards cloud-based technology. in item 1 above, the librarians believed that cloud-based technology requires high technical knowledge in academic libraries. this has shown that 76.5% and 21.2% of them agreed to this statement while 2.3% did not agree. also, the librarians are of the view that cloud-based technology can be used to build community power because 70.5% and 29.5% agreed to this statement. average librarians believed most academic libraries can afford the cost of using cloud-based technology because 31.8% and 18.2% agreed while 50% disagreed. in item 4, 1.5% and 13.6% librarians believed that cloud-based technology maintenance is not difficult while 19.7% and 65.2% disagreed with this statement. more so, 67.4% and 22% agreed that cloud-based technology will enhance storage and preservation of information resources but 10.6% did not support this idea. whether cloud-based technology can be trusted, 21.2% and 12.1% believed cloud-based technology can be trusted while 48.5% and 18.2% disagreed with the statement. in view of the above statements and the results, it was concluded that the perception of librarians toward cloud-based technology in academic libraries is negative, this can be attributed to librarians’ belief that cloud-based technology requires high technical technology and cannot be trusted. https://doi.org/10.29173/iq1083 7/11 akinyoola, oladoyin grace (2023) knowledge and perception of librarians towards cloud-based technology in academic libraries in southwest, nigeria, iassist quarterly 47(3-4), pp. 1-11. doi: https://doi.org/10.29173/iq1083 table 3: shows the level of the usage of cloud-based technology among librarians s/n sa a d sd n % n % n % n % 1. i use google apps to migrate from desktop to web-based technology in the library 61 46.2 22 16.7 31 23.5 18 13.6 2. i use koha software for storage and preservation of materials in the library 40 30.3 44 33.3 21 15.9 27 20.5 3. i use drop box to store data/information in the library 12 9.1 34 25.8 26 19.7 60 45.5 4. i use dspace software in the library 06 4.5 14 10.6 62 47.1 50 38.8 5. polaris library automation system is not new to me as storage device 02 1.5 02 1.5 78 59.1 50 37.9 6. i use cloud-based technology such as google docs like gmail, google drive, text files, spreadsheet to store information in the library 88 66.7 44 13.2 00 00 00 00 table 3 above shows the level of usage of cloud-based technology by librarians in academic libraries. in item 1, many librarians have been using google applications to migrate from desktop to cloud-based technology in academic libraries. this has shown that 46.2% and 16.7% agreed with this statement while 23.5% and 13.6% disagreed with this statement. also, the librarians were of the view that they use koha software for storage and preservation of materials in academic libraries because 30.3% and 33.3% of them agreed to the statement, while 15.9% and 20.5% disagreed. 9% and 25.8% agreed that they have been using dropbox to store data/information while 19.7% and 45.5% disagreed with this statement. few libraries use dspace with responses of 4.5% and 10.6% while 47% and 38% of them disagreed with the statement. in item 5 only 15% and 1.5% believed polaris library automation system (plas) is not new to them as storage device in academic libraries while 59.1% and 37.9% disagreed with the statement. many librarians agreed that they have been using cloud-based technology such as google documents, like gmail, google drive, text files, spreadsheet, text files to store information in the library. the above result showed that many librarians have been using one or more cloud-based technologies in academic libraries in nigeria. https://doi.org/10.29173/iq1083 8/11 akinyoola, oladoyin grace (2023) knowledge and perception of librarians towards cloud-based technology in academic libraries in southwest, nigeria, iassist quarterly 47(3-4), pp. 1-11. doi: https://doi.org/10.29173/iq1083 table 4: shows the challenges of the cloud-based technology among librarians s/n sa a d sd n % n % n % n % 1. maintaining cloud-base technology requires huge cost 78 59.1 48 36.4 06 45 00 00 2. lack of policy as regards cloud-based technology 45 34.1 81 61.4 06 4.5 00 00 3. technical know-how and inadequate in house expert. 56 42.4 74 56.1 02 1.5 00 00 4. poor maintenance culture 110 83.3 22 16.7 00 00 00 00 5. erratic power supply 123 93.2 09 6.8 00 00 00 00 6. unwillingness on the part of librarians to embrace change. 99 75 33 25 00 00 00 00 7. phobia for technology 89 67.4 19 14.4 24 18.2 00 00 8. poor ict infrastructural facilities. 77 58.3 45 34.1 10 7.6 00 00 table 4 above revealed perceived challenges associated with the use of cloud-based technology by librarians in nigeria. 59.1% and 36.4% of the respondents agreed that the maintenance of cloud-based technology is high while 4.5% disagreed with the statement. 34.1% and 61.4% agreed that lack of policy as regards cloud-based technology is one of the challenges while 4.5% disagreed with the statement. 42.4% and 56.1% indicated that technical know-how and inadequate in-house experts affect the use of cloud-based technology. 83.3% and 16.7% of the respondents believed poor maintenance culture is an issue with cloud-based technology. erratic power supply was considered as major challenge with 93.2% and 6.8% of the respondents agreeing with this statement. 75% and 25% believed that unwillingness to embrace change affects cloud-based technology. 67.4% and 14.4% of the respondents are of the view that phobia for technology is one of the challenges associated with cloud-based technology while 18.2% of the respondents disagreed with the statement. 58.3% and 34.1% believed that poor information and communication technology infrastructural facilities affect the use of cloud-based technology. this results and findings, therefore, indicate that there are a lot of challenges militating against the use of cloud-based technology in academic libraries in nigeria. discussion research question one revealed that the level of librarians’ knowledge towards cloud-basedtechnology https://doi.org/10.29173/iq1083 9/11 akinyoola, oladoyin grace (2023) knowledge and perception of librarians towards cloud-based technology in academic libraries in southwest, nigeria, iassist quarterly 47(3-4), pp. 1-11. doi: https://doi.org/10.29173/iq1083 in academic libraries in southwest, nigeria is high. this finding is in agreement with muhammad and abdullah (2021) who did a study on awareness and adoption of cloud computing and reported that many librarians in nigeria are aware of cloud computing. this finding also corroborates annie (2014) study who that many librarians have a positive attitude towards cloud-based technology. the finding is in agreement with lazaru and eucharia (2021) who found out that the librarians are highly aware of cloud-based technology. however, the finding is in contrary to finding of achugbue (2018) who revealed that the level of librarians’ awareness of cloud-based technology in universities libraries in south-south, nigeria was below average. result from research question 2 revealed that librarians’ perception of cloud-based technology is negative. this finding corroborates the finding of achugbue (2018) who reported that the perception of librarians towards the adoption of cloud-based technologies in university in nigeria is poor as a result of poor understanding of the concept. this finding is in agreement with lazarus and eucharia (2021) who reported that higher percentage of librarians in nigeria do not have positive perception towards cloud based technology, as a result of low technical knowledge of the different cloud-based applications. however, this finding is in contrary opinion to muhammad and abdullahi (2021) who reported that librarians’ perception towards cloud computing is high irrespective of computer literacy and age. findings also revealed that librarians have been using one or more cloud-based technologies and found it easy to use and it enhances library services. this finding corroborates the views of alhami and khaparde (2014), achugbue (2018) and mcmarus (2016) that revealed librarians are using one or more of cloud based technologies to access information and to store information for later retrieval anytime and anywhere. also, this finding is in consonance with the study of lazarus and eucharia (2021) who found out that librarians have been using these cloud-based technologies to access information, enrich their social knowledge, easy to use and enrich their skills in solving academic challenges. swapna and birader (2017) study revealed that librarians have not been using some cloud-based technologies like drop-box, polaris software and others which this finding also corroborates. the study also revealed that there are a lot of challenges towards the use of cloud-based technologies by librarians in academic libraries in nigeria which include, poor maintenance culture, erratic power supply, high cost of purchasing equipment, lack of policy as regards cloud-based technology technical know-how, phobia for technology, unwillingness to embrace change and the like. this finding corroborates finding of achugbue (2018) and dosumu (2015) that the cost of inserting and managing cloud-based technology in nigeria is high and other challenges revealed by this study. the finding is in line with the finding of nazir and rashis (2013) who found out that there are lots of impediments which have made cloud-based technology vulnerable to security threats. conclusion cloud-based technology has been seen as paradigm shift from traditional ways of storing information to online storage and preservation of data /information. librarians need to integrate cloud-based technology into library operations system and ensure its maximization in providing resources for users’ satisfaction. provision of quality services to users at all times regardless of locations can only be achieved if the https://doi.org/10.29173/iq1083 10/11 akinyoola, oladoyin grace (2023) knowledge and perception of librarians towards cloud-based technology in academic libraries in southwest, nigeria, iassist quarterly 47(3-4), pp. 1-11. doi: https://doi.org/10.29173/iq1083 librarians can keep abreast of cloud-based technology in this forth industrial revolution. recommendations the following recommendations were emerged; i. it was recommended that staff development programs should be embarked upon by librarians in order to enable them keep abreast of cloud-based technology and other new developments. ii. library management should encourage librarians to keep pace with technologies that can help library operations and be ready to sponsor librarians to seminars, workshops and conferences that can help develop information technological skills. iii. adequate funding to be taken serious by government of nigeria. this will enable librarians to equip libraries with equipment and other facilities needed for the growth of new technology. iv. in order to implement cloud-based technology there should be adequate power supply. nigeria government should ensure uninterrupted power supply and in the absence of this, library management should make provision for alternative power supply. v. policy on the usage of cloud-computing should be developed. references abidi, f. & abidi, h. j. (2012). cloud libraries: a novel application for cloud computing. international journal of cloud computing & services science, 1(3), 79-83. achugbue i. e. (2018). librarians’ awareness and perception towards the adoption of cloud-based technologies in public university libraries in south-south. retrieved from https://hdl.handle.net/29.500.12493/153. agandi, a. b. & gull, k. c. (2013). security issues with possible solutions in cloud computing: a survey. international journal of advanced research in computer engineering and technology , 2( 2), 652-661. alhamdi, f. a & khaparde, v (2010). collaboration in the computing among students of library and information science department of basasaheh ambedkar marathwada university, aurangabad. international journal of advance library and information science, 2(12), 82-92. retrieved from http://dx.doi.org/10.15640/jlis.v2n2a1 annie, l. (2014). using cloud computing in higher education: a strategy to improve agility in current financial crisis in communication of ibima.procedia engineering, 13, 44-51. hussaini, s., vishistha, r., garba, a. & jimah, h. (2017). cloud computing in nigeria university library system: an overview.. new delhi: indian federation of united nations associations. kamba, i. (2017). librarians awareness and perception towards the adoption of cloud computing in nigeria. retrieved from http://idr.kab.ac.ug. kaushik, a. & kumar, a. (2013). application of cloud computing in libraries. international journal of information dissemination and technology. 3(4): 270-273. lazarus c. n. & eucharia, k. (2021). awareness and use of cloud computing: its implications in selected academic libraries in imo state, nigeria. journal of information and knowledge management, https://doi.org/10.29173/iq1083 https://hdl.handle.net/29.500.12493/153 http://dx.doi.org/10.15640/jlis.v2n2a1 http://idr.kab.ac.ug/ 11/11 akinyoola, oladoyin grace (2023) knowledge and perception of librarians towards cloud-based technology in academic libraries in southwest, nigeria, iassist quarterly 47(3-4), pp. 1-11. doi: https://doi.org/10.29173/iq1083 12(1), 62-75, retrieved from doi: https://dx.doi.org/10.4314/iijikm.v12i1 . mark, j. n. (2016). cloud computing for education: a new dawn. international journal of information management, 51, 31-37. retrieved from researchgate.net. mcmanas, b. (2016). the implications of web 2.0 for academic libraries. electron journal of academic and special librarianship, 3(3), 1-10. retrieved from https://www.ajol.info>article. muhammad, f & abdullah, s. (2021). awareness and adoption of cloud computing in digital and university libraries for effective service delivery. world journal of innovative research, 11(1), 87-95. musungwini, s. (2016). analysis of the use of cloud computing among universities lecturers: a case study in zimbabwe. international journal of education development using information and communication technology , 12(1), 53-70. retrieved from researchgate.net/public national institute of standard technology. (2012). the nist definition of cloud computing. retrieved from http://csrc.nist.gov/publications/nistpubs/800-145/sp800-145.pdf nazir, i. & rashis, s. (2013). security threats with associated mitigation technique in cloud computing. international journal of apple information system, 5(7), 16-19. retrieved from researchgate.net pal, s. k. (2013). cloud computing and library services: challenges and issues. retrieved from https://www.academia.edu/14138994/cloud_computing_and_library_services_challenge _and_issues sahu, r. (2015). cloud computing: an innovative tool for library services. retrieved from http://eprints.rclis.org/29058/1/r%20sahu.pdf seena, a. & sudhier, i. (2013). microreviews: types of cloud computing. library philosophy and practice. retrieved from http://www.webpages.uidaho.edu/~mbolin/safdar-mahmood-qutab.htm swapna, g. & biradar, b. s. (2017). application of cloud computing technology in libraries international journal of library and information studies, 7(1), 5261. yuvraj, m. (2013). cloud computing application in indian central library: a study of librarians use. library philosophy and practice. retrieved from https://digitalcommons.unl.edu/libphilprac/992 . https://doi.org/10.29173/iq1083 https://dx.doi.org/10.4314/iijikm.v12i1 http://www.ajol.info/ http://csrc.nist.gov/publications/nistpubs/800-145/sp800-145.pdf https://www.academia.edu/14138994/cloud_computing_and_library_services_challenge_and_issues https://www.academia.edu/14138994/cloud_computing_and_library_services_challenge_and_issues http://eprints.rclis.org/29058/1/r%20sahu.pdf http://www.webpages.uidaho.edu/~mbolin/safdar-mahmood-qutab.htm https://digitalcommons.unl.edu/libphilprac/992 editor's notes 42 2 1/2 rasmussen, karsten boye (2018) editor’s notes: metadata is key the most important data after data, iassist quarterly 42 (2), pp. 1-2. doi https://doi.org/10.29173/iq933 editor's notes metadata is key the most important data after data welcome to the second issue of volume 42 of the iassist quarterly (iq 42:2, 2018). the iassist quarterly has had several papers on many different aspects of the data documentation initiative for a long time better known by its acronym ddi, without any further explanation. ddi is a brand. the iassist quarterly has also included special issues of collections of papers concerning ddi. among staff at data archives and data libraries, as well as the users of these facilities, i think we can agree that it is the data that comes first. however, fundamental to all uses of data is the documentation describing the data, without which the data are useless. therefore, it comes as no surprise that the iassist quarterly is devoted partly to the presentation of papers related to documentation. the question of documentation or data resembles the question of the chicken or the egg. don't mistake the keys for your car. the metadata and the data belong together and should not be separated. ddi now is a standard, but as with other standards it continues to evolve. the argument about why standards are good comes to mind: 'the nice thing about standards is that you have so many to choose from!'. ddi is the de facto standard for most social science data at data archives and university data libraries. the first paper demonstrates a way to tackle the heterogeneous character of the usage of the ddi. the approach is able to support collaborative questionnaire development as well as export in several formats including the metadata as ddi. the second paper shows how an institutionalized and more general metadata standard in this case the belgian encoded archival description (ead) is supported by a developed crosswalk from ddi to ead. both these ddi papers originated from the european ddi conferences (eddi). based on an agreement between eddi and iq, eddi18 mentions the possibility to publish articles in the iq. see the entries on “full paper” and “communications (short paper)” and the author guidelines of the call for papers at http://www.eddiconferences.eu/ocs/index.php/eddi/eddi18/schedconf/cfp. iq 42:2 is not a ddi special issue, and the third paper presents an open-source research data management platform called dendro and a laboratory notebook called labtablet without mentioning ddi. however, the paper certainly does mention metadata it is the key to all data. the winner of the paper competition of the iassist 2017 conference is presented in this issue. 'flexible ddi storage' is authored by oliver hopt, claus-peter klas, alexander mühlbauer, all affiliated with gesis the leibniz-institute for the social sciences in germany. the authors argue that the current usage of ddi is heterogeneous and that this results in complex database models for each developed application. the paper shows a new binding of ddi to applications that works independently of most version changes and interpretative differences, thus avoiding continuous reimplementation. the work is based upon their developed ddi-flatdb approach, which they showed at the european ddi conferences in 2015 and 2016, and which is also described in the paper. furthermore, a web-based questionnaire editor and application supports large ddi structures and collaborative questionnaire development as well as production of structured metadata for http://www.eddi-conferences.eu/ocs/index.php/eddi/eddi18/schedconf/cfp http://www.eddi-conferences.eu/ocs/index.php/eddi/eddi18/schedconf/cfp 2/2 rasmussen, karsten boye (2018) editor’s notes: metadata is key the most important data after data, iassist quarterly 42 (2), pp. 1-2. doi https://doi.org/10.29173/iq933 survey institutes and data archives. the paper describes the questionnaire workflow from the start to the export of questionnaire, ddi xml, and spss. the development is continuing and it will be published as open source. the second paper is also focused on ddi, now in relation to a new data archive. 'elaborating a crosswalk between data documentation initiative (ddi) and encoded archival description (ead) for an emerging data archive service provider' is by benjamin peuch who is a researcher at the state archives of belgium. it is expected that the future belgian data archive will be part of the state archives, and because ddi is the most widespread metadata standard in the social sciences, the state archives have developed a ddi-to-ead crosswalk in order to re-use their ead infrastructure. the paper shows the conceptual differences between ddi and ead both xml based and how these can be reconciled or avoided for the purpose of a data archive for the social sciences. the author also foresees a fruitful collaboration between traditional archivists and social scientists. the third paper is by a group of scholars connected to the informatics engineering department of university of porto and the inesc tec in portugal. cristina ribeiro, joão rocha da silva, joão aguiar castro, ricardo carvalho amorim, joão correia lopes, and gabriel david are the authors of 'research data management tools and workflows: experimental work at the university of porto'. the authors start with the statement that 'research datasets include all kinds of objects, from web pages to sensor data, and originate in every domain'. the task is to make these data visible, described, preserved, and searchable. the focus is on data preparation, dataset organization and metadata creation. some groups were proposed a developed open-source research data management platform called dendro and a laboratory notebook called labtablet, while other groups that demanded a domain-specific approach had special developed models and applications. all development and metadata modelling have in sight the metadata dissemination. submissions of papers for the iassist quarterly are always very welcome. we welcome input from iassist conferences or other conferences and workshops, from local presentations or papers especially written for the iq. when you are preparing such a presentation, give a thought to turning your one-time presentation into a lasting contribution. doing that after the event also gives you the opportunity of improving your work after feedback. we encourage you to login or create an author login to https://www.iassistquarterly.com (our open journal system application). we permit authors 'deep links' into the iq as well as deposition of the paper in your local repository. chairing a conference session with the purpose of aggregating and integrating papers for a special issue iq is also much appreciated as the information reaches many more people than the limited number of session participants and will be readily available on the iassist quarterly website at https://www.iassistquarterly.com. authors are very welcome to take a look at the instructions and layout: https://www.iassistquarterly.com/index.php/iassist/about/submissions authors can also contact me directly via e-mail: kbr@sam.sdu.dk. should you be interested in compiling a special issue for the iq as guest editor(s) i will also be delighted to hear from you. karsten boye rasmussen june, 2018 https://www.iassistquarterly.com/index.php/iassist/about/submissions 6 iassist quarterly 2016 iassist quarterly iassist quarterlyiassist quarterly towards common metadata using gsim and ddi 3.2 by mogens grosen nielsen and flemming dannevang1 common metadata are used together with local metadata that are specific to a statistical domain. abstract since the 1990’s, statistics denmark has provided users with rich metadata, including classifications, quality declarations, and variables and concept systems. in 2011, we initiated projects aimed at a common integrated metadata system in order to improve quality and facilitate dissemination of statistics. at the beginning of 2015, statistics denmark launched a ddi-based system handling concepts and quality information for 237 statistics. at the ongoing project statistics denmark are including variables, concepts and classifications in one common metadata system. gsim and ddi 3.2 are used as standards for building models at the conceptual, logical and implementation levels; colectica is the software of choice. we demonstrate in this paper that the road towards common metadata in a statistical context requires a) improvement and precision in the terminology we use when we talk about metadata, b) better understanding of the role of metadata in relation to our users, and c) greater recognition of the role of metadata in production of statistics. the paper presents various perspectives on metadata in our context, including metadata in a helicopter perspective, metadata in business processes, and metadata implemented in colectica using gsim and ddi 3.2. keywords frame of reference, metadata, gsim, ddi, colectica introduction since 2011 statistics denmark has been working on building a common metadata system based on ddi and other standards among which quality standards from eurostat play a key role. in january 2015, the standardised description of quality was in place for 237 statistics. the ongoing work focuses on variables, concepts and classifications. as in all other metadata systems the important part is to work towards reuse of common metadata. by common metadata we mean metadata about concepts, classification, and variables that are stored only in one place and reused across various domains and in all relevant business processes, including dissemination processes. civil status, firm, and income concepts are examples of common metadata in our system. common metadata are used together with local metadata that are specific to a statistical domain. metadata introduces a high degree of complexity and requires a great deal of effort from various disciplines and from the whole organisation. the claim in this paper is that the iassist quarterly 2016 7 iassist quarterly path towards common metadata in a statistical context requires a) improvement and precision in the terminology that is used when we talk about metadata b) better understanding of the role of metadata in relation to our users, and c) greater recognition of the role of metadata in the production of statistics. in section 2, we first give a short history of the development of metadata at statistics denmark and then provide a description of the challenges with regard to opinions on metadata. we also touch upon the difficulties related to the introduction of the generic statistical information model (gsim) and the generic statistical business process model (gsbpm). in section 3, we offer a helicopter perspective on metadata in order to give an idea of how to approach metadata, focusing on metadata as frames of reference. section 4 introduces metadata and business processes. we first introduce a general model. hereafter we present a model on how to integrate metadata in business processes. the last part of the section focuses on processes related to user needs. section 5 gives a longer introduction to models and terminology used in gsim. it can be seen as our approach to integrate not only gsim, but also ddi 3.2 in a colectica [4] implementation. metadata at statistics denmark: present situation and challenges for several years, statistics denmark operated with separate classifications, quality declarations, and variables and concept systems. some of these were founded on international standards, but they lived their own lives without integration. in 2010, we decided to improve the integration of these systems, including going from separate metadata to creation and definition of common, reusable metadata. for example, concepts are partly defined in both variable and concept sub-systems. they must be common in order to get consistent definitions to be reused in various areas. on the dissemination side, we wanted to use common metadata together with other metadata to give users easier access to our products (metadata used to direct users to the right products) as well as better information about what we offer. in 2010, we discussed whether to keep the existing systems and build an integration component on top of those or build a new integrated metadata system. at a metis meeting in geneva in 2011, it became clear to us that ddi could help us with the integration, reuse and dissemination of metadata. note that the integration was the main driver in relation to dissemination. we wanted to give the users easy ways to navigate from, for example, a variable in a dataset to common metadata on classifications and code lists relevant for this variable. we launched a pilot study using colectica, which is based on ddi. the results from the study were promising. in 2012, statistics denmark got an eu grant focusing on improved horizontal and vertical integration of metadata. we used this opportunity to introduce both ddi as a metadata standard and colectica, which is a ddi-based metadata tool. in addition to ddi, we introduced the sdmx reference metadata on quality. this was basically a list of quality concepts and corresponding values. this implementation was and is compliant with the single integrated metadata structure (sims). the sims standard is developed and used by eurostat and member states. it covers metadata for quality reporting and as such it is only a small brick in the bigger sdmx framework used by eurostat. the ddi 3.1 standard had been improved and it was possible using colectica to “plug in” sims into our metadata system. in january 2015, the standardised description of quality was in place for 237 statistics at statistics denmark. they are now an integrated part of our dissemination system. in spring 2015, a new strategy on quality and metadata was approved by the top management at statistics denmark. the vision focuses on fulfilment of user needs, implementation of quality, and efficiency. regarding metadata, the strategy stresses the use of standards, end-to-end production, reuse and active metadata. we have since introduced new projects as part of the implementation. in the project on quality mentioned above and the present ongoing projects on concepts, variables and classifications, we are experiencing difficulties in people’s perception of metadata and especially the role of metadata in the production processes. it has been difficult to allocate resources and to change processes (including some stovepipe production) towards using common standards. in addition, there is a widespread opinion that documentation is something you attach to your statistical products after they have been published. there is little attention paid to how the statistics and documentation about the statistics are produced and should be used. user needs are often introduced too late in the production process or without the efforts necessary [1], [2]. our observations also show that there is little awareness among our subject matter experts of the benefits of using common models, both in terms of precise definitions of information objects like population, statistical unit, etc., but also in how to apply common work procedures introduced in gsbpm. we have tried to introduce gsim, although the general opinion is that gsim seems to be too complicated. in addition, there is a widespread opinion that gsim seems mainly to fit into organisations having a lot of resources for automation. in order to create common metadata we must handle these challenges. in section 4, we focus on how to integrate use and production of metadata into production processes (4.1) and how to improve the dissemination of metadata (4.2). section 5 focuses on the related challenges achieving common metadata terminology. all three aspects of metadata need to be addressed in order to get metadata to support not only automation but also improvements in the information about our statistics that we disseminate to our users. . the helicopter perspective: metadata as compatible frames of reference the term “statistical system” is being used in many different ways. from a helicopter perspective, it is helpful to use the following distinction: a statistical system as a system of statistics or a statistical system as a system for producing statistics [5]. in short, a system of statistics is about the statistics we produce and a system for producing statistics is how we produce the statistics. 8 iassist quarterly 2016 iassist quarterly figure 3.1 reality, information and data. working towards a situation with common metadata, we need some basics on reality, data, and information. the diagram below shows the interplay [5]:: in short, the mental process can be expressed by the following equation introduced by børje langefors: “i = i(d, s, t)” )” where i is the information (or knowledge) produced from the data d and the preknowledge s, by the interpretation process i, during the time t.[ ...] in the general case, s in the equation is the result of the total life experience of the individual. it is obvious from this that not every individual will receive the intended information from even simple data.” [6 ] but what about common and reusable metadata which we are aiming at? “sharing of data (over time and space) is a proxy process for sharing of information. sharing of information is fundamentally impossible. we can only do our best to improve the chances that different persons sharing the same data will interpret them in the same or at least similar ways. how can we do that?” we use “compatible frames of reference: a person’s interpretation of data depends on the person’s frame of reference, which consists of concepts and information in the person’s mind. if two persons have the same or at least compatible, frames of reference, it seems likely that they will interpret the same data in similar ways”; [5] the challenge is to ensure compatible frames of reference. this is done by creating metadata: data about data. metadata can be shared and communicated between users. “communication of metadata is subject to the same fundamental difficulties as communication of the basic data that they describe, but even so, adequate metadata will reduce the range of possible interpretations of the data that they describe, and thus improve the chances of different persons making similar interpretations of the same data.” [5] metadata has an essential role in facilitating compatible frames of reference. gsim and other models can help in creating these frames of reference by introducing common terminology. sundgren notes that communication about metadata shares the same difficulties as communication in general. these difficulties can be viewed at all levels: from global to simple communication between two people. much research in this field introduces terminology on social systems and organisational learning [7], [8]. it is beyond the scope of this paper, but may be the most important aspect on the road towards a successful understanding and implementation of metadata. based on the considerations above, we can distinguish between two levels of terminology: general terminology on metadata, and domain specific terminology on metadata. the first level is terminology with regard to metadata constructs themselves (e.g., what is a “variable” compared with a “represented variable” or “concept”). the second level (closer to most end users) is instances of metadata constructs (e.g., the specific variable “income of business” and its definition, etc.). both are associated with considerations of clarity and consistency. for example, “income of business” may vary as depending on what is included and excluded in the calculation of “income”. this is a key consideration. often there exist multiple “standard” definitions e.g., definitions associated with accounting standards vs. definitions associated with the international system of national accounts (sna). sometimes the challenge may be to explain how one definition relates to another, and/or to explain data quality considerations when data collected using one frame of reference (e.g. accounting / business reporting) are used to produce estimates for data using a different frame of reference (sna). the table below shows relations between general metadata terminology and domain specific metadata with respect to frames of reference of producers and users of statistics. in section 5, we will focus on terminology related to metadata. we remedy this by introducing gsim at several levels in order to have metadata terminology as a frame of reference. note that the introduction in section 5 is directed to metadata experts inside nsis. whether we succeed depends on the ways in which we communicate. by going from a very abstract level to the logical and physical levels we believe that it will be easier to establish the needed terminology. metadata and business processes how should we handle the role of metadata in the production of statistics and in relation to end users? we must establish processes that provide the right knowledge for the user. but more precisely, how should we build processes and interact with users? in order to move iassist quarterly 2016 9 iassist quarterly table 1 frames of reference of producers frames of reference of various users general terminology for sta6s6cal metadata 1. complex metadata terminology for producers inside nsi’s (examples: instance variable vs represented variable; classificaaons vs code-lists) 2. simplified metadata terminology used both by internal and external users. (examples: classificaaon, variable, concept, populaaon used for disseminaaon e.g. search-tools) domain specific sta6s6cal metadata 3. domain specific metadata (examples: detailed descripaon of the definiaon of income of person) 4. domain specific metadata differenaated and communicated with respect to frames of reference of various users (examples: short and detailed descripaon of income directed to various user segments) 1 table 3.1 different kinds of frames of reference related to general metadata terminology and domain specific statistical metadata. in this direction we must have a general picture of statistical organisations. if we do not want silos depicted in a traditional organisational diagram, what should we have instead? the high level model of the process-centric organisation described below is a starting point. tthe construction of the model is inspired by the idea of value chains introduced by porter in the 1980’s [9] and ideas on business process management [10] but also more recent ideas and practices related to gsbpm and gsim. this way of thinking implies that you organise all processes from start to finish in such a way that each process adds value to users. this is reflected in distinctions between core processes, figure 4.1 business process perspective with environment elements. 10 iassist quarterly 2016 iassist quarterly high level management processes and support processes. quality, metadata, it, etc., are seen as supporting processes. the focus must be on core processes delivering value to users. management and support processes must be designed to assist the core processes. the production processes, including processes on metadata, must be designed to fulfil goals for the organisation – including goals on cost effectiveness. the model gives an overall framework for the production of statistics. when combining the model with sources of inspiration above, it will be possible to build a complete framework for how to manage metadata, how to work with surveys, how to use standards on processes, metadata, etc., and how to build it solutions. the ideas on end-to-end production have existed for many years in the international statistical community, but very few nsis have managed to walk successfully down that road. examples of more large-scale implementations can be found in canada, australia and sweden. they have all put a great deal of resources into building advanced cross-department solutions. the challenge for smaller nsis with fewer resources is how to benefit from these ideas using standards and standard solutions taking small manageable steps in the right direction. integration of metadata in business processes we must make sure that metadata is integrated in our processes in order to fulfil user needs, vision and goals outlined for statistics denmark. the diagram below shows an overall diagram of the business process architecture including how we expect to integrate metadata into gsbpm. the gsbpm phases are marked with a red rectangle. the boxes in the bottom show the production and use of metadata in the gsbpm phases. the overall idea is to start with users in the needs phase outlined in the gsbpm. the next phase starts with the design of the outputs to be disseminated in the dissemination phase. these inputs will drive what will need to be collected, derived, etc. during the statistical production. in this way what users need in the end (“documentation”) is driving the definition of metadata to be used during statistical collection, processing and analysis. with the structural design for “documentation” having been established early in the process, the assembling should become much more straightforward. another important point behind the diagram is the idea of metadata driven production [13] [14]. the australian bureau of statistics define metadata driven production as ‘configurable, rule-based and modular ways of producing statistics’ [13]. the sub processes for is phase 4-8 are typically the starting point when designing and implementing components for the metadata driven production. subject matter staff and support staff define metadata on variables, concepts business rules etc. in phase 2. these metadata are used in phase 3 in the construction of components. these component are ”plugged into” production process systems in phase 4 to 8.. figure 4.2 diagram showing gsbpm business processes, workflow and the use of metadata iassist quarterly 2016 11 iassist quarterly metadata and user needs according to the model presented in section 3 it is important to handle frames of reference. how do we do that? in order to establish frames of reference we must have dialogues with users in workshops where we invite users, trying to achieve a double purpose. the first purpose is to learn as much as possible about how users wish to use statistics and the problems they seem to encounter. the second purpose is to make them understand terminology and ways to use metadata to improve their search and use of information, and allow them to better understand the way in which we handle metadata. an example of the improved dissemination of metadata is the implementation of summary and detailed levels in the dissemination of quality information at www.dst.dk. in this way, we try to target both the general public and users who need detailed information (e.g. researchers). in general, we must establish processes that consider various users’ input. this will primarily happen in gsbpm phase one (needs) and in gsbpm phase seven (disseminate). this aspect is discussed in several papers [1] [2] [3] including a model on how to find out about user needs for metadata. metadata and the implementation of a gsim-compliant ddi model in colectica introduction to mapping the claim in the paper is that we need a) improvement and precision in the terminology we use when we talk about metadata, b) better understanding of the role of metadata in relation to our users, and c) better understanding of the role of metadata in production of statistics. the last two parts have been discussed in the previous sections. besides these, there is a need for improvement and precision of terminology we use when creating and using metadata about statistics. this is where gsim comes in. the gsim model is complementary to gspbm and together they aim to improve processes, communication and automation as depicted in figure 4.2. another important aspect of gsim is, that it works as a reference model of information objects that can aligned with the ddi and sdmx standards. this gives us the benefit of the work on these standards as well as access to standard software. the aim of this section is to introduce gsim and related standards in order to move towards a common terminology. the introduction is mainly complex metadata terminology for staff with knowledge of and experience in modelling statistical metadata. as discussed in section 3, this must be supplemented by a more intuitive and narrow terminology that must be used in relation to end users. end user terminology must only include intuitively known terms respecting the end user frame of reference like statistical unit, property of a unit, population, unit-type, variables at the dataset, etc. as written above, the content and form of metadata depend on the common understanding of metadata terminology reached in discussions with end users, e.g., presentations in metadata portals. it must also be noted that we are mainly focusing on a part of gsim telling us about systems of statistics. this means that we mainly use elements like concepts, variables, classifications, populations, and unit type to name the most important terms. we are aware that gsim has the ambition of having metadata about systems for producing statistics. in a broader scope this must be taken into account, e.g., how definitions of variables for output are derived from definitions of input variables. the information objects in gsim are described by definitions, attributes and relationships; however, the model is not fully developed in terms of attributes. areas of the model such as classifications have enhanced sets of attributes whereas the model in other areas just includes the attribute’s name and description. when aligning ddi with gsim, we aim to use only ddi elements and associations which are present in gsim. on the attribute level, this is not possible due to the simplicity of gsim. we expect that the two models will become closer during the next couples of years, and by applying this strategy we will more easily be able to migrate towards newer versions of the standards, both at a logical and at a physical level. the development of information-systems typically involves three levels of abstraction. these contain requirements of modelling, use case modelling, class definition, etc. modelling traditionally includes a three-tiered approach [11]: • conceptual level – the basic entities of a proposed system and relationships between them • logical level – specifies entities with a full set of attributes and their relationships without implementation details • physical level – defines the physical structure for a technology specific format the models at each of the three levels of abstraction correspond to model driven architecture (mda) concepts. in our implementation we have an additional top level, the gsim conceptual model, which is mapped into a gsim-compliant ddi conceptual model. level scope of model and standards used conceptual 1 selected elements from gsim concept and structure area: variable, concept, dataset etc. conceptual 2 selected terms from ddi 3.2 complying with gsim terms logical selected elements from ddi 3.2 used for implementation physical logical model extended with colectica implementation details figure 5.1 modelling at various levels 12 iassist quarterly 2016 iassist quarterly the process of implementing gsim-compliant ddi in colectica involves the following steps: 1) mapping a. high-level mapping from gsim to ddi: the purpose of the high-level mapping is to ensure compliance between gsim and ddi in terms of association and cardinality. this mapping can be quite complex, as seen in classification. see paper about the copenhagen mapping [12]. an easier example is mapping variables from gsim. see figure 5.3 below. b. low-level mapping, including adjustments from gsim to ddi: here the models are mapped on the attribute level. ddi provides some generic mechanisms for implementing user attributes. 2) creating the logical ddi 3.2 model. once the mapping is in order, it is then possible to create the logical ddi model. workflowrelated logic is added. figure 5.2 gsim variables figure 5.3 dd1 variables aligned with gsim iassist quarterly 2016 13 iassist quarterly 3) creating the physical model: the gsim-compliant ddi model implemented in colectica is a common effort between colectica and statistics denmark so that the revised ddi model in (2) is followed. the metadata also have to be organised physically in a way that facilitates reuse. mapping fig 5.2 below shows selected gsim objects and their relations. figure 5.3 shows that the gsim variables aligned with ddi 3.2 using the ddi terminology. one should note that both unit type and population are mapped into universe. this is a fault in the ddi model as they are clearly not the same. in the gsim model, the instance variable is associated with the represented variable which again is associated with the variable. in ddi there is an additional association between instance variable and variable. hence, we must prohibit the use of the additional association. in the following we will deal with the mapping from gsim to ddi. in order to become familiar with the terminology we will use two simple cases. please note that for communication purposes the cases deal with physical representations rather than treating gsim at a logical level. this masks how elements in data structures are reused but gives the reader a handle to a concrete example from the real world. brno name adress noofemp economicactivity 17150413 statististics denmark sejrøgade 9 2100 københavn ø denmark 450 22.12.01 30500435 buena noche pizzaria nørrebrogade 55 2200 københavn n denmark 1 57.01.12 figure 5.4 br_frozen_2014.xlsx figure 5.5 conceptual model showing gism concepts and example data from the business register 14 iassist quarterly 2016 iassist quarterly case 1: unit-dataset from the business register. this is a case where we focus on microdata, which is called unit data in gsim. a so-called “frozen” version of the business register is disseminated in various forms once a year. one example would be an excel spreadsheet for the year 2014, “sheet1” in the spreadsheet file; “br_frozen_2014.xlsx”. in cell 2,4 a value of 450 is found, which is the number of employees at statistics denmark. the column header name is “noofemp” . in addition, the dataset has columns with id, name, address and economic activity. now we want metadata about the number 450. the diagram below shows selected information objects from gsim. the gsim concepts are shown with a black font and the object value is in red. for example, 450 is marked with red and put into the box called datum in the figure. note also that the gsim concepts are marked with bold in the text below. so how are we to interpret the number 450? the gsim model comes in handy here: in gsim the value 450 is an attribute of the concept datum. the datum lives within a placeholder: the data point which lives in a data set structured by a data structure. for a unit data set a data point will measure one unique instance variable, “actcomp_noofemployees”. so now we have an intuitive understanding of what we are measuring, namely a number in a one-dimensional dataset that measures the instance variable “actcomp_noofemployees”. but we need more information and gsim elegantly explains it all while making it all reusable. to understand an instance variable it must be associated with a population and a representation, the represented variable. in this case the population is “all active companies within the period” and the represented variable takes its meaning from “comp_noofemployees”. to understand the representation of the datum value is quite easy. this is not coded but simply a value domain of positive integers. so are we there? not yet. we need an explanation of the content of the variable “comp_noofemployees”. every represented variable takes meaning from a conceptual variable, which measures a concept on a unit type. so now we are back to basics. when the statistics were designed years ago, someone expressed a wish to measure the number of employees on all active companies (the population) every year. they elaborated that a company should be specified as a legal entity with a business register id (the unit type). furthermore, the concept “number of employees” should be specified as people working full-time all year. we now have a conceptual interpretation for the datum, but it would be nice to know which company we are dealing with (company identification). let us therefore proceed to the structural part of gsim. a dataset in gsim is structured by a data structure, which contains data structure components. in this case the dataset is our excel spreadsheet, which is structured by “br_frozen_layout”. the layout contains a logical record, which is a reference to a data record independent of its physical location. the layout also points to three types of data structure components: an identifier component is the unique identifier for the unit, here “br_id” with the value “17150413”. the thing we are measuring is stored in the measure component, “comp_noofemployees”. with metadata about content, population, unit-type, etc., we now have information about value 450 in cell 2.4! case 2: a dimensional dataset from population data about the population in denmark are disseminated in various forms. one version would be a dimensional dataset (a cube) for the year 2014, “sheet1” in the spreadsheet file “sc_2014.xlsx”. in cell 2.2 a value of 14500 is found. the column header name is “gender”, the first column is labelled civil status. the sheet name (not shown here) tells us the measure is livingpersonsincph2014_noof as in case 1 the selected objects are shown in the following object diagram where the gsim concepts are with the colour black and the object value with the colour red. (see figure 5.7) now we want information about the number “14500”. the understanding of a dimensional dataset is different from a unit dataset based on the data structure. as the business register example with unit data, the data point is interpreted by its instance, representation and fig 5.6 sc_2014.xlsx iassist quarterly 2016 15 iassist quarterly figure 5.7 conceptual model showing gism concepts and example data about population variable. for dimensional data, the measure component is connected to as many identifier components as the number of dimensions, each of which has its own representation and concept. this is shown in the figure below, which shows the conceptual part of gsim for the two dimensions civil status and gender, (see figure 5.8). so aside from understanding livingpersonsincph2014, we have to explain gender f and civil status unmarried. the instance variable “genderpersonsincph2014” takes its meaning from the represented variable “persongenderrepresentation” measured by a “gendercodelist”. the variable takes its meaning from the “persongendervariable”. dealing with civil status is equivalent. implementation of ddi and colectica in the start of 2015 statistics denmark completed the implementation of quality declarations in colectica. in spring 2015, together with colectica, we completed the modelling of classifications, which we are prototyping now. feeling confident, we finally took the bold move to modelling and prototyping the conceptual part of gsim variables as seen in our two cases for three statistics, and all of this is available in our coming portal. the structural part of linking it up with our statistics bank will be a new exciting project coming years. conclusion in the beginning of the paper, we claimed that the road towards common metadata in a statistical context requires a) improvement and precision in the terminology that is used when we talk about metadata b) better understanding of the role of metadata in relation to our users, and c) greater recognition of the role of metadata in the production of statistics. based on this, we went through the history of metadata at statistics denmark and presented the challenges we have had with the existing situation of the metadata work. it was shown that widespread myths about metadata make it difficult for both producers and user of metadata to benefit from common metadata. for example, a widespread opinion regarding metadata is that documentation is something you attach to your statistical products after they have been published. there is only a little reflection on how the statistics and documentation about the statistics are produced and should be used. user needs are often introduced too late in the production process or without the efforts necessary. in addition, there is little awareness about using common models, both in terms of precise definitions of information objects like population, statistical unit, etc., but also in how to apply common work procedures introduced in gsbpm. in order to pave the road towards common metadata we gave a helicopter perspective on how to approach metadata focusing on the importance of metadata understood as a frame of reference. 16 iassist quarterly 2016 iassist quarterly figure 5.8 gender and civil status variables our claims regarding better understanding of the role of metadata in relation to our users and better understanding of the role of metadata in the production of statistics were dealt with in section 4. we stressed the importance of business process modelling and how metadata should be integrated into the business processes. this included a short note about metadata-driven production. based on this, we talked about the importance of metadata in relation to users and how these aspects of metadata should be closely connected to what is going on in the business processes when talking about needs and dissemination. section 5 dealt with the last part of our claim, namely the need for improvement and precision in the terminology used when we talk about metadata. as with all areas of “modern life” the use of terminology plays an important role as well. we gave a longer introduction to models and terminology used in gsim mainly for information model experts. we introduced the gsim model as a complementary model to gspbm. together they aim to improve processes and communication, but also to support it development and automation. the introduction in section 5 in this paper contains information models for people with a frame of reference that requires knowledge of terminology and experience in modelling metadata. this is only one kind of use that can be distinguished from a more intuitive and narrow terminology that must be used in relation to end users. end-user metadata terminology must only include intuitively known terms respecting the end user frame of reference like statistical unit, property of a unit, population, unit-type, variables at the dataset, etc. building and implementing common metadata requires that many disciplines must be in play. in the paper, we argued in favor of a need for improved understanding of the role of metadata in relation to users, metadata in relation to production processes and improvement and precision and communication of common terminology. to succeed in creating common metadata requires that these three aspects are well elaborated and communicated by users and producers of metadata. references [1] thygesen, lars (2013) and nielsen, mogens grosen; how to fulfil user needs – from industrial production of statistics to production of knowledge. statistical journal of the iaos, volume 29, number 4 / 2013. iaos press [2] nielsen, mogens grosen and thygesen, lars (2011) how do end users of statistics want metadata? paper at metis workshop on statistical metadata, 5-7 october 2011 [3] nielsen, mogens grosen and thygesen, lars (2014) implementation of eurostat quality declarations at statistics denmark with cost-effective use of standards. paper presented at european conference on quality in official statistics, vieanna 2-5 june 2014. iassist quarterly 2016 17 iassist quarterly [4] colectica a tool for statistical metadata, developed by colectica. link: www.colectica.com [5] sundgren, b. (2004b). statistical systems – some fundamentals. statistics sweden [6] langefors b. (1995). essays on infology summing up and planning for the future. lund: studentlitteratur [7] espejo, raul (2000): self-construction of desirable social systems in kybernetes, vol. 29 no. 7/8, mcb university press [8] bednar, peter m (2000) a contextual integration of individual and organizational learning perspectives as part of is analysis, school of computing and management sciences, sheffield hallam university department of informatics, lund university, vol 3, no. 3 [9] porter, michael (1985). competitive advantage: creating and sustaining superior performance, the free press [10] harmon, paul (2007). business process change – a guide for business process managers and bpm and six sigma professionals. massachusetts, usa. [11] sparx systems (2011). from conceptual model to dbms, link to website: http://community.sparxsystems.com/white-papers/669-data-mode ling-from-conceptual-model-to-dbms [12] iverson, jeremy; mogens grosen nielsen & dan smith (2014). the copenhagen mapping implementing the gsim statistical classifications model with ddi lifecycle. link: http://cdn.colectica.com/thecopenhagenmapping-draft.pdf [13] aurito rivera, abs; simon wall, abs; michael glasson, abs: “metadata driven business process in the australien bureau of statistics”. work session on statistical metadata (geneva, switzerland 6-8 may 2013) [14] pedro revilla, josé luis maldonado, francisco hernández and josé manuel bercebal national statistical institute, spain. “implementing a corporate-wide metadata driven production process at ine spain”. work session on statistical metadata (geneva, switzerland 6-8 may 2013) notes 1 mogens grosen nielsen, chief adviser, mgn@dst ..dk and flemming dannevang, senior adviser, fda@dst ..dk are employed at statistics denmark. mogens grosen nielsen is head of metadata. flemming dannevang is working with the metadata task and other tasks. 1/26 alter, et al. (2020) provenance metadata for statistical data: an introduction to structured data transformation language (sdtl), iassist quarterly 44(4), pp. 1-26. doi: https://doi.org/10.29173/iq983 provenance metadata for statistical data: an introduction to structured data transformation language (sdtl) george alter1, darrell donakowski1, jack gager2, pascal heus2, carson hunter2, sanda ionescu1, jeremy iverson3, h v jagadish1, carl lagoze1, jared lyle1, alexander mueller1, sigbjorn revheim4, matthew a. richardson1, ornulf risnes4, karunakara seelam1, dan smith3, tom smith5, jie song1, yashas jaydeep vaidya1, ole voldsater4 abstract structured data transformation language (sdtl) provides structured, machine actionable representations of data transformation commands found in statistical analysis software. the continuous capture of metadata for statistical data project (c2metadata) created sdtl as part of an automated system that captures provenance metadata from data transformation scripts and adds variable derivations to standard metadata files. sdtl also has potential for auditing scripts and for translating scripts between languages. sdtl is expressed in a set of json schemas, which are machine actionable and easily serialized to other formats. statistical software languages have a number of special features that have been carried into sdtl. we explain how sdtl handles differences among statistical languages and complex operations, such as merging files and reshaping data tables from “wide” to “long”. keywords metadata, provenance, statistical data acknowledgment acknowledgement: the continuous capture of metadata for statistical data project is funded by national science foundation grant aci-1640575. https://doi.org/10.29173/iq983 2/26 alter, et al. (2020) provenance metadata for statistical data: an introduction to structured data transformation language (sdtl), iassist quarterly 44(4), pp. 1-26. doi: https://doi.org/10.29173/iq983 introduction structured data transformation language (sdtl) is a language for representing data transformation commands found in statistical analysis and data management software. sdtl describes changes to a dataset at both the fileand variable-level. since sdtl is structured and machine actionable, it can be queried to produce histories of each variable in a dataset and to answer questions like: • which original variables were used to construct this derived variable? • which commands were used in the construction of this derived variable? • which derived variables were affected by this original variable? sdtl can be translated into natural language, so that researchers do not need to understand the specific software used to process the data. sdtl can also be incorporated into versions of the prov model that have been extended to describe provenance at the variableand command-level, like provone (cuevasvicenttín et al., 2015). sdtl extends the capabilities of tools used by data repositories and data producers to document and describe data. data repositories specializing in the social sciences maintain documentation in the data documentation initiative (ddi) metadata standard (vardigan, 2008). online data catalogs draw upon content stored in ddi xml, and codebooks are translated from xml into pdf or other formats. ddi includes features for recording data provenance, but there was no systematic way to describe data provenance before sdtl. variable histories expressed in sdtl can be formatted, searched, and queried. data users who download ddi xml from repositories can use automated tools to create new codebooks reflecting their changes to the data. by building sdtl into their workflows, data producers can use sdtl to create documentation and to audit the command scripts that manage their data. in the future, sdtl may also be used to translate command scripts from one statistical analysis package to another. sdtl was developed to work with five leading statistical packages: spss, stata, sas, r, and python (ibm corp., 2019; python software foundation, 2019; r core team, 2013; sas institute, 2015; statacorp., 2020). the continuous capture of metadata for statistical data (c2metadata) project, which created sdtl, set out to automate the creation of variable-level provenance metadata by translating scripts used by statistical analysis software into a format compatible with metadata standards like the data documentation initiative (ddi) (vardigan, 2008) and ecological metadata language (eml) (e.h. fegraus, 2005). (see alter et al., 2020.) our goal was to create a history for each variable showing its derivation from earlier variables and all of the ways that it has been modified. sdtl serves as an intermediate language that represents other languages in a more convenient format. since sdtl is expressed in a structured format (e.g., json) with tags and delimiters, its syntax is obvious and unambiguous, and sdtl is easily read by computer programs without elaborate parsing algorithms. sdtl may be used to translate between statistical languages, but it is designed for documentation and description and not as an operational language. the ddi alliance, which maintains international standards for metadata, has adopted sdtl as one of its suite of products (ddi alliance, 2020). ddi metadata is widely used by data repositories serving the social sciences for data discovery tools, catalogs, and codebooks. sdtl provenance descriptions can be https://doi.org/10.29173/iq983 3/26 alter, et al. (2020) provenance metadata for statistical data: an introduction to structured data transformation language (sdtl), iassist quarterly 44(4), pp. 1-26. doi: https://doi.org/10.29173/iq983 inserted into existing data derivation fields in the ddi metadata standards. the ddi alliance will assure that the sdtl is maintained and expanded in an orderly way. we provide here an introduction to sdtl focusing on important features of the source languages that it can represent and ways that sdtl handles differences among them. a user guide and detailed descriptions of sdtl commands are available at c2metadata project (2020b). statistical packages as data processing platforms sdtl inherits a number of assumptions about how data are transformed from the languages used in statistical analysis packages, and it is helpful to understand how those programs work. statistical packages differ in important ways from two other tools often used for managing data, spreadsheets and relational databases, such as sql. all three tools encourage users to think of data as rectangular matrices, “tables,” but they each have different capabilities and limitations. 1. rows and columns in statistical packages, each row is an individual/entity or an observation of an individual/entity, and each column is a variable describing an attribute of an individual/entity. as in a relational database, columns are named and are referenced by their variable names. statistical packages typically do not allow users to put more than one type of information in each column as a spreadsheet does. 2. metadata statistical packages attach more metadata to each variable than either a spreadsheet or a relational database, even if it is less metadata than most researchers need. users control the data type (numeric, string, date, etc.) and display format of every variable. columns can have both variable and value labels that appear on output. a variable label is a brief description of its content. value labels are text descriptions of the categories in variables. statistical packages encourage the use of integer codes for categorical information, but they will display the corresponding value label in output tables. for example, a variable may be coded as 1 for ages 0 to 15, 2 for ages 15 to 65, and 3 for ages 65 and above, but tables can show these categories with labels “children,” “working ages,” and “older ages.” 3. variable lists and variable ranges one of the most common features of statistics software packages is the use of “variable lists” and “variable ranges” to simplify the application of data transformation commands to multiple variables. for example, common value labels (e.g. 1=”yes”, 2=”no”, 3=”na”) may be applied to hundreds of variables with a single command. variable ranges refer to a group of adjacent columns by identifying the first and last variable in the range. for example, “var01 to var04” in spss or “var01-var04” in stata or sas will apply a command to var01, var02, var03, and var04, assuming that the columns appear in that order in the data. a variable list may include both individual variables and variable ranges, such as “var01 var05 var11-var32 var51-var72”. https://doi.org/10.29173/iq983 4/26 alter, et al. (2020) provenance metadata for statistical data: an introduction to structured data transformation language (sdtl), iassist quarterly 44(4), pp. 1-26. doi: https://doi.org/10.29173/iq983 4. order of rows and columns matters commands in statistical packages can take advantage of the order of columns and rows. the meaning of a variable range, such as “age to income” depends upon the order of the columns. statistical packages also process rows in sequential order, which is often used in data processing scripts. commands that merge files or aggregate within groups may only operate on data sorted in advance. statistical packages can also use values from earlier or later rows in computing variables. for example, if the data consist of annual observations of a country or region, the spss lag function can be used to access the value in the previous year. in stata, a command can test whether the current row applies to the same person or place as the preceding row by using syntax like “if districtid == districtid[_n-1]”. this is not possible in sql relational databases, which do not permit operations that depend on the sequential ordering of rows, but it is possible in spreadsheets, which allow both absolute and relative cell references. sdtl includes a command (sortcases) to change the order of rows, but it does not currently support a command to change the order of columns. when we tested commands that sort columns in spss and stata, we discovered that they apply different sort sequences to variable names. stata is case sensitive and sorts variable names in ascii order. spss is not case sensitive for variable names, but it sorts names beginning with lowercase before uppercase of the same letter. for example, spss sort order: aa8 aa9 aa7 aa6 id xx3 xx4 xx2 xx1 xx5 stata sort order: aa6 aa7 xx1 xx5 xx2 aa9 aa8 id xx4 xx3 5. missing values the value of a variable may not be available for all rows, and all statistical packages have features for handling these “missing values.” statistical calculations may exclude cases with missing values on any variable, or they may adjust for missing values in some way. some statistical packages also allow users to identify more than one type of missing value. in survey research some questions do not apply to all respondents (e.g. “how many years have you been married?”), and respondents may respond “don’t know” or simply refuse to answer. researchers need to distinguish between “does not apply”, “don’t know”, and “no response”. there are also important differences among statistical packages in the ways that missing values in logical expressions are processed. spss and r use three-valued logic in which a logical expression may be true, false or missing. sas and stata use two-valued logic (true or false) by processing missing values as either negative or positive infinity. thus, if the value of varx is missing, the logical expression “varx > 0” will be false in sas but true in stata. 6. dataframes and files when a statistical package is in operation, data may exist only in computer memory or in temporary storage space. a data transformation script may create any number of temporary instances of the data and save only a few of them for later use. the c2metadata project adopted the convention of https://doi.org/10.29173/iq983 5/26 alter, et al. (2020) provenance metadata for statistical data: an introduction to structured data transformation language (sdtl), iassist quarterly 44(4), pp. 1-26. doi: https://doi.org/10.29173/iq983 using “dataframe” for working versions of data stored in memory or temporary storage to distinguish them from data in files that will persist after the transformation script is completed. elements of sdtl the elements of sdtl, called “types,” are divided into groups as shown in figure 1. commands are found in commandbase, which is divided into two parts: transformbase for commands that change data or metadata and informbase for types that generate messages or comments. expressionbase consists of elements used to construct numeric, text, or logical expressions within commands. types that describe variables are in variablereferencebase, which is a sub-category of expressionbase. the last group in figure 1, “types for complex properties,” is used when a property of a type has more than one subproperty. the “bases” shown in figure 1 are hierarchical, and types inherit properties from higher levels. for example, the messagetext property is available to all types in commandbase.6 tables 1-5 list the sdtl types under the headings shown in figure 1. figure 1. sdtl types hierarchy transformbase the types belonging to transformbase are commands that change data or provide information about the data to a user. all commands in transformbase also inherit properties from commandbase: properties inherited from commandbase: command the name of a command sourceinformation information about the source of the transform command. message adds a message that can be displayed with the command. properties inherited from transformbase: producesdataframe identifies the dataframe which this transform produces. consumesdataframe identifies the dataframe which this transform acts upon. https://doi.org/10.29173/iq983 6/26 alter, et al. (2020) provenance metadata for statistical data: an introduction to structured data transformation language (sdtl), iassist quarterly 44(4), pp. 1-26. doi: https://doi.org/10.29173/iq983 in table 1 the commands in transformbase are arranged into six functional sub-groups. there are only four sdtl commands that add or modify variables without changing the structure of a dataframe (group a), and compute and recode are by far the most frequently used in data transformation scripts. compute assigns the value of an expression to a variable. recode converts a continuous variable into categories. commands in group b operate only on metadata (names, labels, data type, display properties). load and save in group c read and write data from files into dataframes. the commands in group d modify the structure of a dataframe by changing the number of rows or columns. commands that control the execution of a script (group e) are discussed below. table 1 transformbase: sdtl types that change a dataframe a. commands that create variables or change the values of a variable aggregate an aggregation summarizes data using aggregation functions applied to data that may be grouped by one or more variables. the resulting summary data is added to each row of the existing dataset. the sdtl collapse command is used when the summary data is used to create a new dataframe with one row per group.. compute assigns the value of an expression to a variable. recode describes recoding values in one or more variables according to a specified mapping. the recode command can either describe a recoding of one or more individual variables, or a range of variables. when one or more individual variables are described, a new variable name can be specified. in this case, the original variable is left alone, and a new variable is created with the recoded values. setmissingvalues defines values that are treated as missing values for a list of variables. b. commands that change the metadata associated with a variable or dataframe rename rename changes the name of a variable or list of variables. setdatasetproperty changes a property of a dataframe. https://doi.org/10.29173/iq983 7/26 alter, et al. (2020) provenance metadata for statistical data: an introduction to structured data transformation language (sdtl), iassist quarterly 44(4), pp. 1-26. doi: https://doi.org/10.29173/iq983 setdatatype sets the data type of a variable or list of variables. setdisplayformat sets the display or output format for a variable or list of variables. setvaluelabels describes the assignment of labels to categorical values. setvariablelabel describes the assignment of a label to a variable. c. commands that read or write files load load data from a file. save writes a dataset to a file. d. commands that change the structure of a dataframe appenddatasets combines datasets by concatenation for datasets with the same or overlapping variables. collapse a collapse command summarizes data using aggregation functions applied to data that may be grouped by one or more variables. the resulting summary data is represented in a new dataset. see aggregate for adding summary variables without changing the number of rows. dropcases rows that match the selection condition are deleted in the dataset. other rows are retained. dropvariables deletes variables from the dataset. keepcases rows that match the selection condition are retained in the dataset. other rows are deleted. keepvariables variables to be retained in the dataset. variables not on the list are deleted. mergedatasets merges datasets holding overlapping cases but different variables. the merge may be controlled by keys or grouping variables. newdataframe creates a new empty dataframe. numbers of rows or columns may be specified. all values are assumed to be missing. reshapelong creates a new dataset with multiple rows per case by assigning a set of variables in the original dataset to a single variable in the new dataset. https://doi.org/10.29173/iq983 8/26 alter, et al. (2020) provenance metadata for statistical data: an introduction to structured data transformation language (sdtl), iassist quarterly 44(4), pp. 1-26. doi: https://doi.org/10.29173/iq983 reshapewide reshapewide is not supported in the current version of sdtl, because it depends on values in the data. however, it may be useful when values of the index variable are available in the metadata file or the data can be processed. sortcases sorts rows in the dataframe in a specified order. e. commands that control the flow of operations in a script doif a set of commands that are performed when a logical expression is true. may also include elsecommands to be performed if the logical expression is false. the commands in doif are performed once, and it expects a logical condition that applies to the entire dataframe. use ifrows for commands that are performed on each row depending upon values on those rows. execute this command causes the system to execute preceding commands before continuing to process the command script. ifrows a set of commands that are performed on each row in the dataframe when a logical expression is true for that row. may also include elsecommands to be performed if the logical expression is false. use doif for a logical condition that applies to the entire dataframe and commands that are performed once. loopoverlist a loop creates multiple versions of a set of commands by iterating over a list of variables, numbers, or strings. loopwhile loopwhile iterates over a set of commands under the control of one or more logical expressions. since the logical conditions typically depend upon values in the data, commands executed in a loopwhile cannot be anticipated and expanded in sdtl. informbase table 2 shows informational commands that do not describe changes to the data. although sdtl does not include commands that analyze data, these commands can be transcribed verbatim in an sdtl script with the analysis command. unsupported is used for commands that our parser cannot translate into sdtl. invalid is used when the parser recognizes a command in the source language but its syntax does not conform to expectations. notransformop was created for commands in the source language that do not play a role in sdtl. for example, r and python install libraries that may change the operation of https://doi.org/10.29173/iq983 9/26 alter, et al. (2020) provenance metadata for statistical data: an introduction to structured data transformation language (sdtl), iassist quarterly 44(4), pp. 1-26. doi: https://doi.org/10.29173/iq983 commands. even though the parser has translated these commands into sdtl, the library may be relevant information for some data users. table 2. informbase: commands that provide information analysis describes an analysis command. an analysis command does not result in any data transformation. comment describes a source code comment. invalid describes an invalid command. a command is invalid if it uses incorrect syntax, or is otherwise not allowed by the executing system. message inserts message text in the sdtl file. notransformop notransformop is used for a command in the original script that provides important information but does not have a function in sdtl. for example, “library()” in r loads a package of r functions. since the parser detects the library, the sdtl will reflect the library that is used, and commands derived from the library will be translated in the sdtl script. however, it is useful to know which library is active for auditing the r script, even if it does not perform any data transformations. unsupported describes an unsupported command. an unsupported command is valid syntax, but not supported by the parsing application. expressionbase the sdtl types in table 3 (expressionbase) are used in expressions, which may be numeric, text, datetime, or logical. the most powerful of these types is functioncallexpression, which is a reference to the function library discussed below. variablereferencebase (table 4) is a subcategory of expressionbase used to describe the variables used in an expression. https://doi.org/10.29173/iq983 10/26 alter, et al. (2020) provenance metadata for statistical data: an introduction to structured data transformation language (sdtl), iassist quarterly 44(4), pp. 1-26. doi: https://doi.org/10.29173/iq983 table 3. expressionbase: sdtl types used in expressions booleanconstantexpression booleanconstantexpression takes values of true and false. datetimeconstant describes a date or date-time combination using an iso 8601 compliant string. functioncallexpression an expression evaluated by reference to the function library. groupedexpression a group of expressions to be evaluated before expressions outside of the group. used to control the order of operations in a formula. iteratorsymbolexpression the name of an iterator symbol used as an index in describing the actions of a loop. missingvalueconstantexpression a missing value constant. some languages allow multiple missing value constants. numberrangeexpression defines a range of numeric values. numericconstantexpression a numeric constant. numericmaximumvalueexpression represents the largest numeric value supported by a system. numericminimumvalueexpression represents the smallest numeric value supported by a system. stringconstantexpression a text string. stringrangeexpression defines a range of string values. timedurationconstant describes a duration of time using an iso 8601 compliant string. unhandledvaluesexpression represents any values not previously handled (for example, in a set of recode rules). valuelistexpression wraps a list of other expressions. variablereferencebase sdtl types used to describe variables. see table 4. https://doi.org/10.29173/iq983 11/26 alter, et al. (2020) provenance metadata for statistical data: an introduction to structured data transformation language (sdtl), iassist quarterly 44(4), pp. 1-26. doi: https://doi.org/10.29173/iq983 table 4. variablereferencebase: sdtl types used to describe variables in expressions allnumericvariablesexpression an expression that represents all numeric variables in the dataset, similar to `_all` in spss or stata. alltextvariablesexpression an expression that represents all text variables in the dataset, similar to `_all` in spss or stata. allvariablesexpression an expression that represents all variables in the dataset, similar to _all in spss or stata. compositevariablenameexpression a composite variable name is used to describe a variable name that is computed. variablelistexpression a list of variables, which may include variable names (variablesymbolexpression) and variable ranges (variablerangeexpression). variablerangeexpression a list of variables in adjacent columns defined by the variable names of first and last columns. variablesymbolexpression a reference to a variable. table 5 includes types that were created to represent complex properties of other commands. for example, every type in commandbase uses sourceinformation to show the original language of each command and its location in the command script. appenddatasets and mergedatasets, which operate on more than one file use types appendfiledescription and mergefiledescription to capture a number of properties associated with each file. https://doi.org/10.29173/iq983 12/26 alter, et al. (2020) provenance metadata for statistical data: an introduction to structured data transformation language (sdtl), iassist quarterly 44(4), pp. 1-26. doi: https://doi.org/10.29173/iq983 table 5. types for complex properties in sdtl commands appendfiledescription describes files used in an appenddatasets command. dataframedescription describes a dataframe in the consumesdataframe or producesdataframe types. provides the name of the data frame and a list of variables (columns). dataframedescription can also define dimensions in dataframes that have hierarchical indexes, data cubes, or multi-indexes. functionargument describes the arguments in a function as specified in the sdtl function library. iteratordescription describes an iteration process consisting of an iteratorsymbolexpression and a list of values it takes. mergefiledescription describes files used in a mergedatasets command. recoderule describes how values will be recoded. recodevariable describes a variable that will have its values recoded. renamepair variable names before and after a variable is renamed. reshapeitemdescription describes a new variable created by reshaping a dataset from wide to long. sortcriterion describes a criterion by which cases are sorted, including the variable name and whether to sort ascending or descending. sourceinformation sourceinformation defines information about the original source of a data transform. valuelabel associates a label with a value in a categorical variable. conditional execution by row or by file/dataframe statistical languages have some commands that are executed sequentially on every row and other commands that apply to an entire file or dataframe. the compute command illustrated above is an example of the first type. compute creates or modifies a variable that will appear on every row in the data. new variables are usually computed from other variables on the same row, but we describe calculations that aggregate across rows in our discussion of the “function library” below. in contrast, https://doi.org/10.29173/iq983 13/26 alter, et al. (2020) provenance metadata for statistical data: an introduction to structured data transformation language (sdtl), iassist quarterly 44(4), pp. 1-26. doi: https://doi.org/10.29173/iq983 commands that load or save files or modify metadata, such as data type and display format, do not change the number or contents of rows in the dataframe. the difference between row-level and file/dataframe-level commands becomes very important when the action is conditional on the value of a variable or other parameter. consider these commands in the stata language: replace vary=3 if varx>5 /*** version 1 *****/ if varx>5 replace vary=3 /*** version 2 ****/ although they appear to be the same, they have very different outcomes. the condition in version 1, “if varx>5”, is a qualifier within a stata command (“replace”) that is executed sequentially on each row in the dataframe. some rows will be set to 3 and others will not be changed, depending upon the value of “varx” on each row. in version 2 the “replace” command is nested in an “if” command, which is a program flow command designed for use in stat scripts (“do-files”). the “if” command is not evaluated separately for each row; it is evaluated only once using the value of “varx” on the first row in the dataframe. consequently, if “varx>5” is true for row one, “vary” is set to 4 for all rows, and if “varx>5” is false for row one, no rows are changed regardless of the value of “varx” on other rows. table 6 illustrates the results of these commands where only row 1 satisfies the condition “varx>5”. table 6. examples of conditional execution by row and dataframe in stata initial values version 1 (sdtl ifrows): replace vary=3 if varx>5 version 2 (sdtl doif): if varx>5 replace vary=3 row varx vary varx vary varx vary 1 9 11 9 3 9 3 2 4 11 4 11 4 3 3 1 11 1 11 1 3 sdtl includes two ways of applying conditions to commands. the sdtl command ifrows is used for conditions that should be evaluated sequentially on every row. doif in sdtl is used for flow control in scripts where the condition is evaluated once before executing a command or group of commands. both ifrows and doif can be applied to a group of commands, and both include an elsecommands property for commands to be performed if the condition is false. function library although the number of data transformation commands in statistical packages is small, the power of these commands is magnified by “functions,” which are available in every language. functions are https://doi.org/10.29173/iq983 14/26 alter, et al. (2020) provenance metadata for statistical data: an introduction to structured data transformation language (sdtl), iassist quarterly 44(4), pp. 1-26. doi: https://doi.org/10.29173/iq983 available in most computer languages as a convenient way to invoke common operations. in the same way that a mathematical equation may use “sine(x)” to refer to the corresponding trigonometric function, “sine(varx)” may be used in a statistical package to insert the sine of variable “varx” in a computation or comparison. there are thousands of functions in statistical packages, and programming c2metadata parsers and updaters to reproduce all of them would have been an enormous job. fortunately, our goal is to describe data transformations not to perform them. we devised a simple way to add an unlimited number of functions to sdtl with minimal impact on the code required to translate a script into sdtl. this was accomplished by creating a function library, which serves as a crosswalk between sdtl and the various statistical packages. the function library is a file that can be maintained in a spreadsheet and accessed as a json file. functions in computer languages normally have two parts: a function name followed by parameters enclosed in parentheses. the function invokes program code that replaces the function with a value computed from the parameters. the computed value of a function may be a number, text, or logical (boolean) constant. for example, sine(varx) will return the sine of an angle equal to the value of varx, and gt(varx, vary) will return true if varx is greater than vary and false otherwise. each parameter is used in a specific way by the computer code that evaluates the function. parameters may be specified in two ways. sometimes, parameters are given in a defined order separated by a delimiter, usually a comma, which is included even if the parameter is omitted. parameters may also be identified by name. for example, in stata “std(varx), mean(10) std(3)” will standardize the values of varx so that the transformed values have mean=10 and standard deviation=3. in this case the first parameter (varx) is given by position, but the other two parameters (“mean” and “std”) are specified by name. in sdtl parameters may be specified by position or by name. we currently use exp1, exp2, exp3, … as parameter names in sdtl, but these names are arbitrary and meaningful mnemonics could be used. some functions operate on a list items of the same type, which makes them appear to have an indeterminate number of parameters. for example, mean(varx, vary, varz) would compute the mean of three variables. to avoid parameter lists of indefinite length, the sdtl function library uses the variablelistexpression and valuelistexpression types in sdtl. a variablelistexpression packages a list of variables into a single sdtl type that is treated as one parameter in an sdtl function. a variablelistexpression has a single property (variables) defined as a json array that can consist of any combination of individual variables (variablesymbolexpression) or variable ranges (variablerangeexpression). the function library maps the names and parameters of sdtl functions to functions in other languages. every function is described with the sdtl name of the function and the order and names of its parameters. the sdtl function is also mapped to the same function in spss, stata, sas, r, and python. this table compares functions that compute a random number from a uniform distribution in sdtl and five other languages: https://doi.org/10.29173/iq983 15/26 alter, et al. (2020) provenance metadata for statistical data: an introduction to structured data transformation language (sdtl), iassist quarterly 44(4), pp. 1-26. doi: https://doi.org/10.29173/iq983 sdtl random_variable_uniform(exp1, exp2) spss rv.uniform(exp1, exp2) stata runiform(exp1 exp2) sas ranuni(seed) r runif(n, min=exp1, max=exp2) python numpy.random.uniform(low= exp1, high= exp2) sdtl and most of these languages specify the minimum (exp1) and maximum ( exp2) of the range of the random number. in sas the range is always 0 to 1, which is the default range in other languages. the function library entry for sas specifies that 0 and 1 are passed to sdtl as values for parameters exp1 and exp2. computer programs often use mathematical formulas to approximate random numbers, and the sas version of this function allows users to specify a “seed” for its random number generator. since the seed is specific to the implementation in sas, it is not included in sdtl. in r the “runif” function creates a vector of random numbers of length “n”. we assume that the random number will be either a single number used in an expression or a vector added to the dataframe as a new variable, which makes this parameter unnecessary in sdtl. the function library partitions functions into four groups corresponding to different sdtl commands: function library group sdtl command meaning horizontal compute calculates a value from variables on the same row. rows are processed sequentially. vertical aggregate calculates a new variable by aggregating across rows in a group. every row in the group has the same value. collapse collapse calculates a new variable by aggregating across rows in a group in a new dataframe with one row per group. logical doif ifrows keepcases dropcases functions used in logical conditions. https://doi.org/10.29173/iq983 16/26 alter, et al. (2020) provenance metadata for statistical data: an introduction to structured data transformation language (sdtl), iassist quarterly 44(4), pp. 1-26. doi: https://doi.org/10.29173/iq983 horizontal functions operate sequentially by row using variables appearing on each row. vertical and collapse functions operate on groups of rows by aggregating values within columns (see figure 2). vertical functions used in an sdtl aggregate command add new variables (columns) to a dataframe by applying the result of a computation to every row in a group. the collapse command does the same computation, but it reduces the number of rows by creating one row per group. for example, suppose that groups are defined by variable “yearsofeducation,” and we compute mean(annualincome). the aggregate command will add mean annualincome to every row, and the collapse command will create one row for every value of yearsofeducation including both yearsofeducation and mean annualincome. thus, vertical functions do not change the number of rows in the dataframe, and collapse functions create a new dataframe with fewer rows. figure 2. illustrations of aggregate and collapse all functions have unique names in sdtl, but other languages sometimes use the same function name in different contexts with different outcomes. a good illustration is a function for computing means, which has three different meanings in both spss and stata. https://doi.org/10.29173/iq983 17/26 alter, et al. (2020) provenance metadata for statistical data: an introduction to structured data transformation language (sdtl), iassist quarterly 44(4), pp. 1-26. doi: https://doi.org/10.29173/iq983 compute aggregate collapse sdtl mean(exp1) agg_mean(exp1) col_mean(exp1) spss mean(exp1) context: compute mean(exp1) context: aggregate with mode=addvariables mean(exp1) context: aggregate stata rowmean(exp1) context: generate, replace mean(exp1) context: egen “(mean)” statistic option context: collapse flow control, loops, and macros statistical software packages include extensive programming capabilities. stata and sas have powerful macro features, and r and python are very capable programming languages. there are two ways of handling these programming features in sdtl. first, whenever possible the parser will expand macros and other programming code into simpler commands. for example, if a loop applies a compute command to four variables, it can be converted into four compute commands. this may make the sdtl long and verbose, but it simplifies the work of finding which commands affect every variable. second, sdtl does include types for describing loops (loopoverlist, loopwhile), which are the most common kind of flow control, and iteratorsymbolexpression was created to describe an index used in a loop. sdtl does not have arrays, which may be used in loops, but it does have functions that operate like arrays. the variablearraydereference and valuearraydereference functions allow an sdtl script to use an expression to select an entry in a list. the first parameter of each function points to the position of an entry in a variable or value list given as the second parameter. the operation of these functions can be illustrated by this simplified example, in which “[age, sex, education, income]” is a list of variable names: variablearraydereference( 3, [age, sex, education, income]) the value of this function would be “education”, which is the third item in the list. since the contents of the list is not stored anywhere, the full list must be repeated every time that the function is used. however, the index parameter could be an iteratorsymbolexpression in a loop. appending, merging, and updating datasets the five languages covered by the c2metadata project offer a wide variety of ways of combining datasets. appenddatasets is used to concatenate rows (cases) from datasets that have the same columns (variables) (figure 3). mergedatasets combines columns from datasets that have the same https://doi.org/10.29173/iq983 18/26 alter, et al. (2020) provenance metadata for statistical data: an introduction to structured data transformation language (sdtl), iassist quarterly 44(4), pp. 1-26. doi: https://doi.org/10.29173/iq983 rows. these operations are complicated by features that resolve conflicts, such as merging files with overlapping column names or unmatched rows. datasets are usually merged by joining rows with the same keys, but some statistical packages will merge rows sequentially when keys are not specified. mergedatasets can also be used to update a dataset by replacing its current values with values from a different dataset. both appenddatasets and mergedatasets use subtypes (appendfiledescription, mergefiledescription) to describe actions that apply to specific input datasets. for example, the merge commands in spss and sas allow users to rename variables, select variables, and select cases at the time of the merge without changing the input dataset. figure 3. appending datasets r (dplyr) and python (pandas) use “joins” like those in sql to merge dataframes. rows in the output dataset are created by comparing one or more key variables specified in a “by” parameter. joins in r and python are implicitly cartesian joins that create every possible combination of rows with the same keys. for example, caseid=2 is repeated in ds_a and caseid=1 is repeated in ds_b. the cartesian join of ds_a and ds_b by caseid is ds_c (figure 4), in which there are two rows for both caseid=1 and caseid=2. note that caseid=3 and =4 are not included in ds_c, because they do not exist in both input datasets. ds_c is the result of an “inner” join, and the unmatched rows can be included by specifying “outer”, “left”, or “right” joins. following the model of sql, r and python are agnostic about the order in which the data are sorted, and all joins are cartesian. https://doi.org/10.29173/iq983 19/26 alter, et al. (2020) provenance metadata for statistical data: an introduction to structured data transformation language (sdtl), iassist quarterly 44(4), pp. 1-26. doi: https://doi.org/10.29173/iq983 figure 4. cartesian inner join in spss, sas, and stata merging is often a sequential process on files that are sorted before they are merged. even when the merge involves matching on key variables, spss, sas, and stata require the input files to be sorted before they can be merged, and the user must determine whether keys are unique (one-to-one) or repeated (one-to-many or many-to-many). a sequential merge of ds_a and ds_b without keys produces ds_d (figure 5), which is very different from the result of a cartesian join (figure 4). figure 5. sequential merge https://doi.org/10.29173/iq983 20/26 alter, et al. (2020) provenance metadata for statistical data: an introduction to structured data transformation language (sdtl), iassist quarterly 44(4), pp. 1-26. doi: https://doi.org/10.29173/iq983 sdtl uses three properties (mergetype, newrow, and update) to represent all of these possibilities. these properties are found in mergefiledescription, which means that they are specified for each input dataset. the mergetype property describes how rows from the input datasets are combined in the output data. most merge types (e.g. onetoone, onetomany) involve matching rows on key variables, which are specified with mergebyvariables (in mergedatasets) and mergebynames (in mergefiledescription). sequential merges assume that the input files are already sorted. the newrow property determines when the rows contributed by an input file generate a row in the output file. when newrow is true, all rows in this dataset are included in the output dataset, regardless of whether they were matched to another input dataset on the mergebyvariables. when newrow is false, only rows that have been matched are included. an inner join is represented in sdtl by setting newrow to false on all input datasets, and newrow is true for all input datasets to describe an outer join. left and right-joins are created by using true and false on different inputs. mergetype sequential match rows from each input dataframe in the order in sequential order. onetoone create one row for each value of the mergebyvariables. if a combination of the mergebyvariables is repeated, only one row is matched. rows with repeated combinations of the mergebyvariables may or may not be included in the output file depending on the newrow property. onetomany create a row in the output dataframe by matching rows in this dataframe to every row in other dataframes with the same value of mergebyvariables. note that onetomany implies that one of the other input datarames is set to manytoone. manytoone create a row in the output dataframe by matching all rows in this dataframe to the one row in the other dataframe with the same value of mergebyvariables. cartesian create a new row in the output dataframe for every possible combination of rows having the same value of mergebyvariables. this is equivalent to a many to many merge. unmatched create a new row for every row that cannot be matched on the mergebyvariables sasmatchmerge sas uses a merging approach that combines matching keys and sequential merges within groups. https://doi.org/10.29173/iq983 21/26 alter, et al. (2020) provenance metadata for statistical data: an introduction to structured data transformation language (sdtl), iassist quarterly 44(4), pp. 1-26. doi: https://doi.org/10.29173/iq983 newrow true always include rows from this dataframe, even if the mergefilevariables do not match a row in any other dataframe. false only include rows from this dataframe, if the mergefilevariables match a row in another dataframe. there is even more diversity in the responses of different languages when the datasets to be merged contain a variable (column) with the same name. r and python follow sql by including both variables with modified names, which can be handled by using the renamevariables property in the mergefiledescription. however, spss, stata, and sas will include only one variable in the output data, and they may use the omitted variable to update values in the included variable. the update property of mergefiledescription is used to specify how values from the omitted version of the variable will be handled. if update is set to ignore, a variable that is also found in the master dataset will have no effect on the output dataset. if update is set to fillnew, values from the repeated variable will only appear on new rows not found in the master dataset. updatemissing replaces only missing values in the master dataset, and replace changes all values on matched rows in the master dataset. update master this dataframe is the master dataframe. ignore if a column with the same name exists in the master dataframe, ignore the values in this dataframe. fillnew if a column with the same name exists in the master dataframe, use the values from this dataframe only in new rows created from this dataframe. updatemissing if a column with the same name exists in the master dataframe, use values from this dataframe when the value in the master dataframe is missing. replace if a column with the same name exists in the master dataframe, use values from this dataframe. reshapelong, reshapewide, and compositevariablenameexpression all of the statistical packages covered by the c2metadata project have commands to reshape files between “wide” and “long” formats. figures 6 and 7 illustrate the difference between wide and long format for data describing a mother and her children. in the wide format (figure 6) there is one row for each mother, and each child is described by two variables, age and sex. data for each child are identified by including birth order in the variable name, e.g. age1, age2, etc. the wide format requires a https://doi.org/10.29173/iq983 22/26 alter, et al. (2020) provenance metadata for statistical data: an introduction to structured data transformation language (sdtl), iassist quarterly 44(4), pp. 1-26. doi: https://doi.org/10.29173/iq983 column for every variable for each child, and we must allow enough variables to describe the largest family in the data. if one woman had 20 children, the dataset in figure 6 would have 40 columns: age1, sex1, …, age20, sex20. consequently, datasets in wide format usually have many empty cells. in long format, figure 7, there is a row for each child and information about mothers is repeated on the rows for each of their children. the long format includes an additional variable, birthorder that uniquely identifies children within each family. since the information in each format is identical, the choice between wide and long depends upon the types of analysis to be performed and convenience. figure 6. wide format figure 7. long format the information in figures 5 and 6 can also be stored in separate datasets for mothers and children by using the motherid variable as a key for linking children to their mothers. in a relational database the two-table approach would be used to remove repetition and “normalize” the database. however, unlike sql, most statistical analysis software cannot compute results on data contained in more than one table. data from the mothers table and the children table would need to be merged before any analysis is performed. sdtl includes features for operating on wide and long format data. the compositevariablenameexpression is used to describe repeated variable names in wide-format data, such as age1, age2, etc. composite names consist of a “stub” (e.g. “age”, “sex”) and an index value. composite names are described in a reshapeitemdescription, which is a complex property used in the reshapewide and reshapelong sdtl commands. https://doi.org/10.29173/iq983 23/26 alter, et al. (2020) provenance metadata for statistical data: an introduction to structured data transformation language (sdtl), iassist quarterly 44(4), pp. 1-26. doi: https://doi.org/10.29173/iq983 the c2metadata project has implemented the reshapelong but not the reshapewide command. when we convert data from wide to long, we know how many rows to create, because each row corresponds to a set of identical variables described in the metadata file. but we cannot reshape data from long to wide format without knowing the maximum number of columns to create, which is not included in the metadata file of a long format dataset. since the scope of the c2metadata project has been limited to metadata-only operations, reshapewide is not currently supported. pseudocode library and translator the pseudocode library is a simple and extensible way to create natural language versions of sdtl scripts. every type in sdtl consists of a set of properties. each of these properties can be resolved into text -a variable name, a number, or a string. the pseudocode library is a set of templates for sdtl types with text to include before and/or after each property in a command. templates look like this: “starting text {property1} more text {property2} even more text” the translation involves inserting text created from each property into the corresponding space in the template, where property names surrounded by curly brackets. for example, the pseudocode templates for the rename command and the renamepair type are: rename rename variables: {renames} renamepair \n\t from {oldvariable} to {newvariable}; in this case, renamepair is a complex type used to fill the renames property of the rename command. if we rename vara to varalpha, the renamepair becomes. \n\t from vara to varalpha; and the rename command becomes rename variables: \n\t from vara to varalpha; if we evaluate \n as a new line and \t as a tab, we get rename variables: from vara to varalpha; pseudocode templates for functions are included in the function library. limitations and future developments the ddi alliance has created an sdtl working group to manage sdtl as one of its suite of standards. modifications and additions to the sdtl standard will follow an orderly process with opportunities for community review and a published calendar for new versions. this framework assures that sdtl will evolve in response to new developments in source languages and new applications of the language. https://doi.org/10.29173/iq983 24/26 alter, et al. (2020) provenance metadata for statistical data: an introduction to structured data transformation language (sdtl), iassist quarterly 44(4), pp. 1-26. doi: https://doi.org/10.29173/iq983 the function library provides a simple way to expand the reach of sdtl without changing the language itself. the function library can be updated to include new functions in sdtl or to map additional functions in source languages to existing sdtl functions. in most cases, additions to the function library do not require changes to the program code in applications that translate source languages into sdtl. version 1.0 of sdtl is being released with two limitations that are due to the limited scope of the c2metadata project. first, the c2metadata project adopted a metadata-only approach. we assume that the pre-transformation data are well described in a standard metadata schema, such as ddi or eml, and we do not access the data at any time. for this reason, sdtl can describe reshaping data from long to wide, but c2metadata parsers do not support that command. when data are changed from long to wide, the number of columns in the new dataframe depends upon the values of the index variables in the original dataframe. the only way to know the range of these index variables is to inspect the actual data, and this requires integration of sdtl into statistical analysis software. we hope that this integration will happen in the future, especially for the open source packages r and python. second, sdtl does not yet describe variables created by statistical analysis commands. sdtl was created to describe data and not tables, graphs or other analytical results. since statistical analysis packages have many more analysis commands than data transformation commands, representing analysis commands was not on the agenda of the c2metadata project. however, we acknowledge that analysis commands can also produce data. for example, estimated regression models are often used to construct predicted values and residuals. in view of the number and diversity of analytical commands, sdtl may be linked to an external ontology of statistical tests, such as the stato ontology (isa commons, 2020). sdtl was designed to document data transformations, and it is not intended to be an executable language. sdtl provides enough information for a human to understand changes to a file or a variable, but this may not be sufficient for a computer to perform these operations. in addition, a command script may be translated into sdtl in more than one way. sdtl, like other complex languages, often provides several methods for accomplishing a specific result. for example, the functions performed by the sdtl recode command can also be achieved by ifrows and compute statements or by the cut() function found in some languages. discussion sdtl provides a new level of transparency for data processed and managed by statistical analysis packages. sdtl was created to simplify the automated creation of provenance metadata at the variable level. the c2metadata project is providing open-source code for translating spss, sas, stata, r, and python into sdtl, as well as code for translating sdtl into natural language (c2metadata project, 2020a). software applications that create data catalogs, codebooks, and tools to reconstruct data provenance can read sdtl rather than interpreting each of the different statistical languages. for data producers, these tools simplify the process of describing the steps in preparing raw data for publication. data repositories will receive more detailed machine-actionable metadata to improve the documentation their collections. researchers will be able to understand how variables were created regardless of the software used in their production. https://doi.org/10.29173/iq983 25/26 alter, et al. (2020) provenance metadata for statistical data: an introduction to structured data transformation language (sdtl), iassist quarterly 44(4), pp. 1-26. doi: https://doi.org/10.29173/iq983 references alter, g., donakowski, d., gager, j., heus, p., hunter, c., ionescu, s., . . . voldsater, o. (2020). automating the capture of data transformation metadata from statistical analysis software. icpsr. university of michigan. ann arbor mi. retrieved from http://hdl.handle.net/2027.42/156014 c2metadata project. (2020a). gitlab repository: c2metadata. retrieved from https://gitlab.com/c2metadata c2metadata project. (2020b). structured data transformation language. retrieved from http://c2metadata.gitlab.io/sdtl-docs/ cuevas-vicenttín, víctor, b ludäscher, p missier, k belhajjame, f chirigati, y wei, and b leinfelder. 2016. "provone: a prov extension data model for scientific workflow provenance." retrieved from http://jenkins-1.dataone.org/jenkins/view/documentation%20projects/job/provonedocumentation-trunk/ws/provenance/provone/v1/provone.html ddi alliance. (2020). structured data transformation language. retrieved from https://ddialliance.org/products/sdtl/1.0 e.h. fegraus, s. a., m.b. jones, m. schildhauer. (2005). maximizing the value of ecological data with structured metadata: an introduction to ecological metadata language (eml) and principles for metadata creation. bulletin of the ecological society of america, 86, 158–168. ibm corp. (2019). ibm spss statistics for windows, version 26.0. armonk, ny: ibm corporation. isa commons. (2020). stato: an ontology of statistical methods. retrieved from http://statoontology.org/ python software foundation. (2019). python language reference, version 3.8. beaverton, or. retrieved from https://www.python.org/ r core team. (2013). r: a language and environment for statistical computing. vienna, austria: r foundation for statistical computing. retrieved from http://www.r-project.org/ sas institute. (2015). sas®9.4 product documentation. cary, nc: sas institute inc. retrieved from http://support.sas.com/documentation/94/index.html statacorp. (2020). stata statistical software: release 16.1. college station, tx: statacorp lp. vardigan, m. (2008). beyond the codebook: documenting survey research on the web. paper presented at the international conference on survey methods in multinational, multiregional, and multicultural contexts (3mc), berlin, germany. https://doi.org/10.29173/iq983 http://hdl.handle.net/2027.42/156014 https://gitlab.com/c2metadata http://c2metadata.gitlab.io/sdtl-docs/ http://jenkins-1.dataone.org/jenkins/view/documentation%20projects/job/provone-documentation-trunk/ws/provenance/provone/v1/provone.html http://jenkins-1.dataone.org/jenkins/view/documentation%20projects/job/provone-documentation-trunk/ws/provenance/provone/v1/provone.html https://ddialliance.org/products/sdtl/1.0 http://stato-ontology.org/ http://stato-ontology.org/ https://www.python.org/ http://www.r-project.org/ http://support.sas.com/documentation/94/index.html 26/26 alter, et al. (2020) provenance metadata for statistical data: an introduction to structured data transformation language (sdtl), iassist quarterly 44(4), pp. 1-26. doi: https://doi.org/10.29173/iq983 endnotes 1 university of michigan 2 metadata technologies north america 3 algenta technologies 4 norwegian centre for research data 5 norc 6 we show sdtl types in italic font beginning with uppercase, like compute. properties within types are in italic font beginning with a lowercase letter, like sourceinformation. https://doi.org/10.29173/iq983 ibi^sist newsletter vol.1, no. 3 lassist secretariat report: united states judith rowe princeton university the report of the united states secretariat covers the current status of membership enrollment as well as proposed membership recruitment activity. specifically, it addresses the goal of recruiting for action group activity all of the professional staff of each data archive and data library. it includes a general summary of both secretariat and action group activity, as well as a report on the first north american action group conference in cocoa beach. other topics include an initial report on the plans for next year's conference and some activities and programs planned by other organizations which would be of interest to lassist members. among the latter are the annual conferences of special libraries association (sla), association of public data users (apdu) and the american association for information science (asis), as well as the inter-university consortium for political and social research (icpsr) workshop for data librarians. data archive registry-survey of past effort, suggestions for the future lisa lasko canadian consortium for social research the mandate of the lassist action group for data archive registry is described. the major portion of the paper is devoted to a brief review, description, and evaluation of some of the more important existing directories. these are: social science data archives in the united states , published by the council of social science data archives in 1967; a directory of information resources in the united states: social sciences , 2nd edition, published by the library of congress in 1973; the second edition of the encyclopedia of information systems and services , edited by anthony kruzas and published in 1974; and the directory of data bases in the social and behavioral sciences , edited by vivian sessions and published in 1974. recent developments in data archive registries, such as the unesco sponsored directory of data services and the directory of data centres , to be published by the data clearing house for the social science in canada are also discussed. essential elements and user requirements of social science data archive directories are dealt with in reference to these directories. finally, a summary of the problems facing the data archive registry group, both abroad and in canada, is given, along with several proposed options include a discussion of the viability of a data archive registry group in canada; unique contributions to be made in this area, one of which might be a compilation of a list of subject headings for social science data archives; and, the production of a directory of data archive personnel. 20. microsoft word iq49_1-editors-notes-final 1/2 schwartz, ofira & hayslett, michele (2025), reflecting on past practices and research to innovate, iassist quarterly 49(1), pp. 1-2. doi: https://doi.org/10.29173/iq1155 the creative commons-attribution-noncommercial license 4.0 international applies to all works published by iassist quarterly. authors will retain copyright of the work and full publishing rights. editors’ notes: reflecting on past practice and research to innovate welcome to the first issue of iassist quarterly for 2025, iq 49(1). we are excited to welcome two new members to our editorial team, mary carter, the finance and operations research librarian at princeton university, and jessica (jess) hagman, the social sciences research librarian and an assistant professor at the university of illinois urbana-champaign. mary and jess have graciously volunteered to serve as our new managing editors and will share the responsibility. they have already taken an active role in the production of this issue. the editorial team together with the editorial board continue to develop policies for authors, reviewers, and the editorial team. we hope to share these policies with the iassist community in the near future. the current issue, iq 49(1), presents three excellent papers. all three review services offered by data librarians or issues important to them, and identify opportunities to incorporate innovative approaches to enhance those services for users and researchers. author madison golden shares the adaptive approach she uses as a research data librarian. the article ”adaptive data governance for research data management” builds on the author’s experience working in data governance at a corporation as well as her more recent experience as a research data librarian in an academic institution. the author introduces four styles of data governance that provide a framework for librarians and data governance specialists alike to prioritize competing needs and guide researchers through the data lifecycle. this approach offers increased flexibility in data management practices, continuous improvement of services and resources, efficiency, and empowerment of researchers and related stakeholders. in the article ”literature review on the competencies of data literacy for middle-grade learners” author semi yeom reviews the literature related to data literacy guidelines and practices for students in k-12 school settings, focussing on middle-grade learners. the author identifies eight main competencies that are important for data-literate adolescents, and highlights the two competencies that were pointed out by researchers as essential skills for academic success and critical engagement in our increasingly datadriven world. the article ”support for computer-assisted qualitative data analysis software in arl libraries” authored by paul pival surveys support for computer-aided qualitative data analysis software (caqdas) among members of the association of research libraries (arl). by visiting institutional websites and libguides, the author tries to understand the level of qualitative data analysis expertise and support provided to 2/2 schwartz, ofira & hayslett, michele (2025), reflecting on past practices and research to innovate, iassist quarterly 49(1), pp. 1-2. doi: https://doi.org/10.29173/iq1155 researchers at these academic institutions. peer institutions that do not offer such services are encouraged to explore this possibility to better support their researchers. we hope you enjoy the reading! we are looking forward to seeing many of you in june, at iassist 50th anniversary conference, the ”best iassist ever” in bristol, uk. ofira schwartz and michele hayslett, march 2025 4 iassist quarterly 2016 / vol 40 no 3 iassist quarterly editor’s notes being international and proud of it! iassist is proud of being international. these days some us of find it important to emphasize how international collaboration has improved and made our lives more efficient. in the small but around-the-globe-reaching world of iassist, many national data archives have come into existence as well as continuing their development, through friendly international support and spreading of knowledge and good practices among iassisters. so let us cherish the 'international' in iassist. we are proud of the lead 'i' for 'international' in the iassist acronym and have no intention of changing that to 'n' for 'national'. it is also my impression that data archives all over the world simply don't have the facilities for storing 'alternative facts' as they are shy of all kinds of documentation. welcome to the third issue of volume 40 of the iassist quarterly (iq 40:3, 2016). four papers with authors from three continents are presented in this issue. the paper 'demonstrating repository trustworthiness through the data seal of approval' is a summary of a panel session at the iassist 2015 conference in minneapolis with panel members stuart macdonald, ingrid dillo, sophia lafferty-hess, lynn woolfrey, and mary vardigan. the paper has an introduction from dans in the netherlands where the data seal of approval (dsa) originated. cases from the us and south africa are presented and the future of the dsa including possible harmonization with other systems is discussed. dsa certifications are basically consumer guidance, clearly assisting all the involved parties. depositors and funding bodies will be assured that data are reliably stored, researchers can reliably access the data repositories, and repositories are supported in their work of archiving and distribution of data. the second article brings us to the actual use of data. from the uk data service, rebecca parsons and scott summers in 'the role of case studies in effective data sharing, reuse and impact' take us into positive narratives around secondary data. the background is that although the publishing of data is now recognised by funders, the authors find that ‘showcasing’ brings motivation for data sharing and reuse as well as improving the quality of data and documentation. the impact of case studies is all-sided and research, depositing data, and the brand recognition of the uk data service are among the areas investigated. the future is likely to include new case studies developed for use in teaching in schools, with easy linking to datasets, as well as for researchers being assisted to build their own portfolios. the appendix presents case studies on research and impact. in the third article, we are situated in data creation. muhammad f. bhuiyan and paula lackie from carleton college in minnesota write on 'mitigating survey fraud and human error: lessons learned from a low budget village census in bangladesh'. as the 'fraud' term implies, they are looking into the problem of data creators being too creative, but more importantly they are investigating the essential area of data quality. the authors explain how selected technological assets like the use of geographic information systems (gis) and audio-capturing smart pens improved data quality. the use of these tools is exemplified through many scenarios described in the paper. furthermore, a procedure of daily monitoring and fast transcription lead to quick surveyor re-training and dismissal of others, thus minimising data errors. for those interested in false data and its detection, the introduction in particular has valuable references to literature. in the last paper the difficult task of handling images is addressed in 'image management as a data service' by berenica vejvoda, k. jane burpee, and paula lackie. vejvoda and burpee work at mcgill university in montreal. you have already met lackie from carleton college in relation to the third paper above. the 'images' in the article are digital images, and the authors suggest that the knowledge of digital data services across the 'research data lifecycle' also benefits the management of digital images. digital images are numerical data, and the article compares the data, metadata, and paradata of a survey respondent to the information on a digital image. considerations from normal data concerning system formats and storage space also apply to management of images. in the last section the paper introduces copyright issues that are complicated, to say the least. just as reuse of normal data can have ethical angles, it is even more apparent that images can have complicated issues of privacy and confidentiality. papers for the iassist quarterly are always very welcome. we welcome input from iassist conferences or other conferences and workshops, from local presentations or papers especially written for the iq. when you are preparing a presentation, give a thought to turning your one-time presentation into a lasting contribution. we permit authors 'deep links' into the iq as well as deposition of the paper in your local repository. chairing a conference session with the purpose of aggregating and integrating papers for a special issue iq is also much appreciated as the information reaches many more people than the session participants, and will be readily available on the iassist website at http://www.iassistd ata.org. authors are very welcome to take a look at the instructions and layout: http://iassistdata.org/iq/instructions-authors authors can also contact me via e-mail: kbr@sam.sdu.dk. should you be interested in compiling a special issue for the iq as guest editor(s) i will also be delighted to hear from you. karsten boye rasmussen january 2017 editor http://www.iassistdata.org http://www.iassistdata.org http://iassistdata.org/iq/instructions-authors http://www.sam.sdu.dk vol21.3 36 iassist quarterly this is a reprint of a paper that first appeared in the iassist quarterly vol. 21:2 without tables. we apologize to the authors for this error of omission and reproduce the paper here in its entirety. introduction researchers who work with large sequential datasets are often limited in the kinds of analytic strategies they can use because of the sheer size of the data. automated techniques for analyzing sequences were developed in the 1960s by scientists studying dna, rna, and proteins. in a classic volume on sequence analysis, sankoff and kruskal (1983) demonstrated its potential application for subjects as diverse as bird songs and macromolecules. in other work, andrew abbott developed “optimal matching” for sequence analysis in the field of sociology. in this paper, we describe a technique for analyzing sequences using regular expression matching (rem). this technique allows researchers to examine patterns in longitudinal data by condensing sequences of events into smaller, more tractable units. we also briefly discuss the development of a database structure that facilitates this kind of analysis. although all sequence analyses compare linear arrangements of symbols, whether in human behavior or dna, they differ in their assumptions about what makes two sequences similar or different. sequences in their original form often contain too much detail for useful comparison, since the possible permutations of occurrences can be limitless. therefore, in all cases researchers must create the rules that define sequence similarity for their analyses. methods for determining sequence similarity are often referred to as sequence-matching algorithms. these algorithms are mathematical, and compare sequences without reference to the semantic or theoretical structures that created them. when using such methods, researchers who wish to place their analyses in an appropriate context categorizing event sequences using regular expressions by lisa sanfilippo & john van voorhis* must carefully define what events represent the phenomena of interest. the technique described in this paper was developed as an alternative to existing algorithms and allows researchers to identify sub-patterns of events within sequences at the start of their analysis, based on theoretical or practical considerations. because this technique operates on a single sequence at a time, it is faster than processes that require comparing many sequences to one another. the project to illustrate rem, we will describe how we used it to analyze the sequence of events that led to a child’s placement into foster care in three states: illinois, michigan, and missouri. we were looking for systematic demographic and geographic differences among children that correlated with the events they experienced in the child welfare system. rem was developed to describe and compare the pathways the children took through this system. the data were derived from the administrative data systems of the illinois department of children and family services, the michigan family independence agency, and the missouri department of social services. preliminary data processing we received two data extracts from each state: one covering investigations of child abuse and neglect in the child protection system (cps), and the other, services such as foster care to children in the child welfare system (cws).1 we began by creating a project database for each state with the same essential structure. each state’s database contained tables for cps and cws data and one table for demographic information on the children. next, we created an event table that contained all of the administrative events for all of the children in each system. we then transformed each child’s events into a sequence variable or “history.” finally, we used regular expression matching to formulate “careers” by reducing the history sequences. at each step in the process, we preserved enough information from the previous step to retain flexibility in the subsequent steps. as the categories became broader at each step, the comparability of the data across states increased. fall 1997 37 creating the event table in this analysis we focused on four key administrative events: (1) indicated investigation, an investigation in which credible evidence of abuse/neglect was found, (2) unfounded investigation, an investigation in which no credible evidence of abuse/neglect was found, (3) case opening, when a case was opened for child welfare services, and (4) placement, when a child was placed in a foster home or institution. we created one record for every event a child experienced in either the cps or the cws. we then coded every record with a number denoting a particular event type (see “event codes” in table 1). these records contained the child’s id, an event date, and an event type code (see table 2). creating the history sequences we transformed each child’s event records into a single sequence of codes, since as separate records the table structure was not appropriate for sequence analysis. to make the programming and its interpretation easier, we used only single-character codes in the history sequence. although each history code represented a single event, a given code value could represent more than one type of event (see history codes” in table 1). we first reviewed a frequency distribution of the history sequences to identify the most common sequences and to see the repetition of patterns within and among sequences. this review also revealed data entry errors that we could correct or eliminate, such as children receiving services before their birth or children being born multiple times. although we had anticipated that the variation in the patterns between sequences would make them unsuitable for analyses in their present form, we had not foreseen the amount of variation in the length of the sequences. for example, examining the distribution of event sequences revealed that many children experienced only one event, while others experienced up to fifty. this wide variation in length made it difficult to make meaningful comparisons among cases and suggested that we needed a method that would not rely solely on whole-sequence comparison. therefore, we focused our attention on identifying the subpatterns which we had observed in the sequences. creating the career sequences one goal of our research was to elucidate the connection between cps investigations and a child’s subsequent placement in foster care. we had three initial questions: (1) what sequences of investigations never resulted in a child welfare case opening and placement? (2) what 38 iassist quarterly sequences of investigations resulted in the child’s first placement? and (3) what sequences of events resulted in the child entering the system without an investigation? because of our extensive work with the illinois data and our contact with all three states regarding current and past practices and policies, we had some knowledge of what the most common patterns of events might be. the following examples illustrate how this prior knowledge provided us with clues about what patterns to focus our attention on: • we understood that the number of investigations a child experienced was not a critical factor in the caseworker’s decision to place the child in foster care. we knew that children with histories composed solely of unfounded investigations were almost never provided with services, despite repeated contact with the department. therefore, we believed that the number of indicated investigations would predict placement better than the raw number of investigations. • we knew that, in one state, caseworkers were reluctant to remove children from their homes after only one indicated investigation unless they were in imminent danger. thus, we expected that a child with one indicated investigation would be less likely to be placed into foster care than a child who had two or more indicated investigations. • in all three states, we knew it was possible for children to experience a case opening and placement without an investigation of abuse or neglect, but we had no information on the frequency of such occurrences. • our prior analyses of the foster care data indicated that once in foster care, a child could move between placements numerous times before being returned home. although the placements could be of different types, the child was still living away from his or her parents. as a result, we chose to treat a series of placements without a return home as one career event. regular expression matching it became apparent in looking at the sub-patterns that they could be represented by regular expressions, a notation used widely in the computer science field for specifying and matching sequences.2 (see appendix.) we created a file listing the regular expression patterns we had decided to analyze along with a “career” code for each pattern which is shown in table 4. we grouped the patterns in passes because we knew that certain patterns occurred only at the very beginning of the history and we needed to control the generation of the matching program. the first pass was used to remove any events that occurred before a child was born. since we were especially interested in the first series of investigations, we created a pass that only matched to initial investigation subsequences. the last pass, which was applied repeatedly until the history was exhausted, contained all of the subpatterns we were investigating. from this pattern file we generated a series of programs to transform the data. we used the awk programming language for both our program generator and the matching programs themselves. an awk program is composed of a series of pattern and action pairs. it automatically reads through data files one line at a time, and each line is matched against the patterns in the order they are listed in the program. when a line contains data that matches one of the patterns, the action associated with that pattern is executed. the patterns may contain regular expressions, while the actions are written in a language similar to the c programming language. in our project, the program generator read the pattern file containing the sub-patterns of interest to us and generated a series of programs that used those regular expression patterns to process the history data. each program in the series corresponded to a particular pass in the pattern file. if a pattern matched to the beginning of a history sequence, the matching characters were removed and the career code for that pattern was appended to the career sequence. the child’s id, history, and career were then passed to the next program for the next pass. the final program passed the data back to itself until the history sequence was empty or until a fixed number of passes had been run. if the history sequence was completely matched, a lower case ‘x’ was appended to the career to indicate completion. an upper case ‘x’ was appended if more history remained after the maximum pass limit had been reached. analyzing the career sequences since our analysis was limited to examining the subpatterns that led to a child’s first placement, we did not analyze children’s entire careers. instead, we only analyzed the first four career events after a child’s birth. because the rem approach simply recoded the original history sequences, it preserved the unit of analysis, thus allowing us to attach explanatory variables such as year of first entry into the system, sex, race, and region3. once this information was stored in one file, we aggregated the data by creating a crosstabulation which contained frequencies for every combination of the career sequences and the explanatory variables. these files were relatively small fall 1997 39 (fewer than 1,000 records) allowing us to import them into a spreadsheet program for final analysis and presentation. conclusion the rem technique described in this paper departs from more common pattern matching methods in that it incorporates theory and practice into the actual matching process. using this technique, researchers can test their assumptions about the structure of a sequence. it is an iterative technique that allows the analyst to explore patterns in the data and to compare them across populations simply and quickly. because the process of developing the career file is split into several steps (i.e., creating the event table, creating the history sequences, and pattern matching), it provides many opportunities to check the data and to ensure that the processes are transforming the data correctly. rem allows the researcher to take a very large dataset and to represent it in a much smaller form, while maintaining the critical details of event order and sequence. for example, in our illinois database we began with an event file of over 5 million records. transforming this file into history sequences, career sequences, and finally into a crosstabulation, decreased the size of the file by a factor of 5,000, making it significantly easier to work with. the rem technique, as written in awk, can save the researcher hours of processing time, in large part due to: 1) the way awk reads data files (i.e., it automatically reads a file one record at a time) and 2) the minimal programming it requires. performing the same analyses using a statistical software package would have required much more extensive programming and perhaps more important, would have restricted the kinds of questions we could have asked in exploring the original data. 40 iassist quarterly future directions clearly, rem has a much wider application than what we have illustrated with our project. our analysis did not utilize rem to its fullest potential. for example, instead of analyzing just the initial sequence of sub-patterns, rem could be used to analyze full careers. we could run a similar process against the career sequences to further shrink the number of categories. finally, we did not explore the sub-patterns in as much detail as we could have. for example, we included specific placement event types in our event table and history sequences but did not treat them as separate types. in the future, we can easily compare differences in children’s histories following specific types of substitute care placements (e.g., home of a relative, private foster home, group home, etc.) based on this project’s current database. appendix regular expressions in general, a character in an awk regular expression matches itself. some characters with special meanings in our pattern file are listed below along with some examples of their use. see the references for more details. references abbott, andrew. 1995. sequence analysis. annual review of sociology, 21:93-113. abbott, andrew & alexandra hrycak. 1990. “measuring resemblance in sequence data: an optimal matching analysis of musicians’ careers.” american journal of sociology, 96(1): 144-185. abbott, andrew & john forrest. 1986. “optimal matching methods for historical sequences.” journal of interdisciplinary history, 16(3): 471-494. aho, alfred v., brian w. kernighan, & peter j. weinberger. 1988. the awk programming language. reading: addison-wesley. aho, alfred v., jeffrey d. ullman. 1979. principles of compiler design. reading: addison-wesley. forrest, john & andrew abbott. 1990. “the optimal matching method for anthropological data: an introduction and reliability analysis.” journal of quantitative anthropology 2:151-170. friedl, jeffrey e. f. 1997. mastering regular expressions. sebastopol: o‘reilly & associates, inc. sankoff, david & joseph b. kruskal eds. 1983. time warps, string edits, and macromolecules: the theory and practice of sequence comparison. reading, ma: addison-wesley. notes 1 . child protection systems: in illinois, the child abuse and neglect tracking system; in michigan, the protective services management information system; and in missouri, the child abuse and neglect data system. child welfare services systems: in illinois, the child and youth centered information system; in michigan, the children’s services management information system; and in missouri, the alternative care tracking system. 2 . regular expressions (res) can recognize patterns which are left linear. patterns, such as a balanced sequence of parentheses, cannot be recognized by res regular expression examples: special characters used in regular expressions: fall 1997 41 because such patterns require “going-backwards” or maintaining information outside of the re. for further information see aho, kernighan, and weinberger 1988 in the references. 3 . in all three states we differentiated the major urban area from the balance of the state. * paper preseneted at the iassist/ifdo 1997 annual conference odense, denmark may 7, 1997 indexing machine -readable data files for a social science data archives by jacqueline mcgee rand corporation "it is still true that the best retrieval system is the expert human mind. " this paper was presented at the lassist annual conference, may 19-22, 1983, in philadelphia, pennsylvania. introduction in the recent past much has been written and discussed about the problems of cataloging and bibliographic control of social science data. many of these problems may have been resolved with the implementation of the angle-american cataloging rules ii, chapter 9 (aacrii) and the marc format for bibliographic control (2). however, there are a number of reasons these solutions may not yet be universally implemented. for instance, the aacrii and the marc format may be very familiar to library staff, but all archives are not staffed by librarians. many archives are suffering from a shortage of staff and financial resources. federal agencies produce a major portion of the data archived and used for secondary analysis and these agencies are also financially depressed. researchers and programmers who use these machine-readable data files (mrdf) are not as aware of the problems related to the acquisition or storage of data and their interests do not necessarily correspond to the interests of the data archivist. technological changes occur so frequently procedures may become obsolete by the time implementation occurs. and finally, so many new commercial firms are installing social science numeric data bases online and the interests of these firms do not lie in the same directions as those of the data archivist. to assist the novice who may be overwhelmed by some of these problems, it is the hope of the author this paper will provide some examples of simple record keeping. rowe and byrum previously described a user-oriented system for the documentation and control of mrdf (3). this system was comprised of four parts. first, a standard catalog entry, second, a data abstract or description form, third, documentation codebooks and lastly, the records of physical and logical characteristics of the data set. it is not the purpose of this paper to offer an alternative system for the documentation and control of mrdf but to provide a practical example of implementing such a system. this example will provide an illustration for the person who has just received responsibility for the safekeeping of a collection of mrdf or to establish an archive and isn't sure where to begin. the first item in the system by rowe and byrum, the standard catalog entry, was described before the anglo-american catalog rules ii, chapter 9 were implemented. data librarians located in a traditional library are already familiar with the rules for cataloging, but may not be familiar with the aacrii, chapter 9. the anglo-american cataloging rules ii, chapter 9 describes the standard rules for cataloging mrdf. it is not within the scope of this paper to argue the pros and cons of the acceptability of the aacrii. there can be no doubt that a uniform standard defining mrdf is necessary in order to alleviate present confusing practices and the proliferation of titles for one data file. certainly implementation of the aacrii and the agreement of the marc format were giant strides in the cataloging and bibliographic control of mrdf. standard catalog entry rowe and byrum state "standard catalog entries, constitute the primary records by which computerreadable data files should.be controlled and accessed." it is with this one area of their discussion that i disagree slightly. the standard catalog entry requires extensive staff time and financial resources and need only be considered as necessary under certain conditions; if the required resources are available and may be allocated to such an endeavor; if the data archive or data bank is situated in a library or a library is available and willing to participate; if the data holdings are original data from the institution responsible for the establishment of the archive. it is hoped non-originating archived data will be catalogued by the originating institution. however, the federal government is responsible for a major portion of the data files held in many archives and current fiscal restraints on most federal agencies probably will not permit such a project in the near future. there is, however, an ongoing cataloging project at michigan's inter-university consortium for political and social research (icpsb) which may resolve the problem of cataloging federal data (4). icpsr is certainly one of the largest, if not the largest, of the data archives in the united states. when this project to catalog their holdings is complete, it may be possible to consider a union catalog. data abstract or data description form it is the data abstract or data description form described by rowe and byrum which should be given priority in the development of an archive record system. the data abstract or data description form in a standard format is an absolute necessity and should be the core of the documentation for the archiving of mrdf. aldrich has proposed a similar abstract form for the documentation of federal mrdf (5). it was from a description by aldrich the following example was derived. changes made in the form were for the benefit of the user and do not reflect a disagreement with her proposed standards. since this form includes an abstract summarizing the data set or file being archived, this document shall be referred to here as a data base profile. with some slight variations, the items contained in the form generally should contain the items shown on the next page. if it is not possible or feasible for the archive or library to catalog the holdings of the archive according to aacrii at least the information supplied in the abstract or data base profile form will conform to the standards for describing mrdf. if at some future time cataloging is possible, the information for the catalog abstract form for the docwnentation of federal medf file #: an identifying number for the individual archive file name: a title file source: producer, distributor, processor principle investigator: primary researcher type of file: survey data, microdata, administrative records, process records, geographic records, software universe: total universe the records describe sample size: number of observations or records sample unit: household, person or unit being measured restrictions: none, or any restrictions placed on the distribution abstract: a summary description of the data set or file. each abstract held in the archive should contain the same information. care should be taken not to omit any portion of the required information and the information should appear as closely as possible in the same paragraph. this assures an easier search for the individual . looking for specific information as to size of the sample, purpose of the study, key variables, etc. references: a descriptive listing of the hard-copy documentation available for use with the data file, i.e., codebooks, survey instruments, dictionaries, etc. related printed reports: known reports where the data is described or where the data file has been used tape specifications: the physical characteristics of the data file, the tape numbers, the logical record length, the blocksize, data set names and density entry will be readily available. copies of the data base profile may be stored in computer format as well as in hard-copy. if the data is stored in computer format it would be possible to devise a simple online search capability. if the data librarian wishes to be bibliographically correct, study the aacrii and include in the profile the pertinent information from the aacrii as well as the information required or deemed necessary for the institution housing the data library (7). many of the elements of the data base profile may be utilized to produce a catalog. for instance, each abstract when extracted from the profile provides summaries of the archive holdings. the rand data facility catalog uses the abstracts in such a way. each abstract then written, therefore, should include the following (6): • data base identification number • source • name • date the information was collected • subject • geographic level of the data (lowest) • population or sampling unit • number of observations or number of logical records • key variables indicies may then be derived from the information given on a data base prof i te . source and name index the source and name index is derived from the file name and the file source as given in the data base profile. often these items are sorted as separate indices; an author index and a title index. an index should lead a user to the information he is seeking with as little effort as possible, and so we have combined these indices. on the data base profile and in data file records we use the most correct name for a data file. the correct name may be drived by using the aacrii rules. since we also wish our individual indices to assist the user in his search we also include in our local archive index those aliases or acronyms when they are commonly used, if the index is to prove useful. in order not to have a great many "see ." included in the index where aliases or acronyms or common usage names are listed, the correct identifying number of a particular data file is used as the pointer. keyword index this index is certainly one of the most difficult to construct. a thesaurus would be helpful; however, the keyword index discussed here was developed from individual data files. as mentioned earlier, one of the manditory sections of the data base profile is a list of key variables. at the time of archiving a new data file, the abstract is written and the key variables extracted. these variables or subject categories are added to the end of the keyword index. a copy of the keyword index is kept online. new keywords are added at the end of the old index and a short sas program sorts the keyword index by words or by identifying file numbers whenever necessary. geographic levels and major subject variables in the keyword index using census bureau designations, the lowest geographic level of the data is assigned as a second element of the keyword index. by using the lowest geographic level it is possible when searching for data to weed out those data files not useful to the researcher. a third element for the keyword index is a prescribed list of major subject variables. for each data file the keyword index will include at least one, but not more than three major subject categories. it is then possible to produce tables listing data files by geographic areas and major subject categories. librarians may want to use the library of congress subject headings (lcsh) list. the data base profiles may be produced as printed copies to be given to researchers interested in using a particular data file and can be used as documentation for bibliographic citations in research papers for the creating of a catalog of holdings. the data base profiles stored in a partitioned data set at rand are soon to be converted to a total online system using the ibm info/sys csd data retrieval system. the oz info/sys has the capability of handling multiple data bases and has been utilized at rand recently to create an information system for a library of software models as used in one department (8). with the data stored in oz it will be possible to search the system for a particular data base by title, by keword, or keyword combinations. requesting data bases by keywords produces a "hit list" of all data bases that contain the requested keywords. the hit list can then be accessed in either of two forms: one is an abreviated form that lists only a few specifics about the data bases on the list (including the data base number and title); the other is a full description of the data base. it is possible to get print copies of an oz screen, of the individual data base descriptions and of the hit list of keywords. references (1) michael phillipson, "the classification, storage and retrieval of survey data, reader in machine-readable social data,: h. white, ed. , information handling services, englewood, colorado. (2) see sue dodd, "cataloging machine-readable data files: an interpretive manual," american library association, december, 1982. (3) john d. byrum, jr. and judith rowe, "an integrated, user-oriented system for the documentation and control of machine-readable data files," library resources & technical services 16(3), summer, 1972. (4) carolyn geda, "marc formal applied to machine-readable data files: a pilot project of icpsr," paper prepared for delivery at the annual meeting of the society of american archivists, september 1-4, 1981, berkeley, california. (5) barbara aldrich, "proposed standards for bibliographic entry and abstracts for federal machine-readable data files," redraft 2, october, 1978. (6) don trees, "the rand computation center: rand's data facility: a guide to resources and services, " bcc-1555/18, santa monica, april, 1981. (7) anglo american cataloging rules ii, chapter 9, machine-readable data files, p201-216. (8) as described by william fowler, the rand corporation, santa monica, california, 1983. 10 4 iassist quarterly 2015 iassist quarterly editor’s notes standardization supporting diversity of use welcome to the fourth issue of volume 39 of the iassist quarterly (iq 39:4, 2015). this issue of the iassist quarterly brings four papers from the iassist 2015 conference in minneapolis. all four have their focus on research data. we take off with the development of curation software, and continue with the special problems of handling video streams because of the current heterogeneous policies requiring much effort from researchers. next follows recommendations for good practices for research data repositories and for university programs in including data management. these three papers overall discuss issues of standardization and how that can benefit researchers and other users of data. the fourth paper demonstrates that standardization in building data collections and software does not necessarily lead to standardized use, by describing the use of data in geographic information systems (gis) in very diverse subject areas. the first paper is ‘new curation software: step-by-step preparation of social science data and code for publication and preservation’ by limor peer and stephanie wykstra. the conference presentation was in the session ‘curation and research data repositories’. limor peer is associate director for research at yale university’s institution for social and policy studies (isps) and stephanie wykstra is research manager of research transparency at innovations for poverty action (ipa). the curation software for reviewing and enhancing research data and code is being developed by research groups at these two institutions in collaboration with colectica. social research carried out at isps and ipa includes field experiments. the paper begins with a discussion of the value of data sharing, and leads on to a description of key curation tasks and the support offered by new curation software. there is an appendix with a detailed description of the curation tasks and also many references and useful links. purdue university has a research project using big data and visual analytics based on over 60,000 publicly accessible video feeds (e.g. weather or traffic cameras). the large quantity of data raises questions about its management. this is described in the paper ‘comparing policies for open data from publicly accessible international sources’, that was presented in the conference session ‘data professionals’. the authors line c. pouchard, megan sapp nelson, and yunghsiang lu are assistant or associate professors at purdue. such data sources currently have heterogeneous policies for data use. the paper compares different policies and describes the implications for open access. the discussion uses examples of the law relating to data privacy in the us and eu such as the uk data protection policy based on the european union data protection directive. fifteen policies were analyzed and are presented in graphics. to illustrate the research burden, the paper describes policies having different use restrictions, all of which the researcher has to abide by when collecting from the various video feeds. the authors propose a template of standards for the components of the policies, as implementation of standards could support scientific use and reuse of video data. the third paper ‘research data repositories: review of current features, gap analysis, and recommendations for minimum requirements’ is by a group of canadian authors: claire c. austin, susan brown, nancy fong, chuck humphrey, amber leahey and peter webster. the work was presented in the conference session ‘data repository models and infrastructure’ by amber leahey. the group surveyed 32 canadian and international online data platforms and compared features and services. alongside this work research data canada developed and published guidelines for deposit and preservation of research data. as with the other papers in this volume, this paper carries extensive literature references as well as excellent documentation of the work through the text and appendices. in the conclusion, the authors recommend that best practices for data management should be incorporated widely in university studies to support research by building capacity and skills. students directly experiencing use of research data in their study programs is addressed in the fourth paper ‘teaching users to work with research data: case studies in architecture, history and social work’ by aaron addison and jennifer moore, who both work at washington university in st. louis. this work was presented at the ‘training data users ii’ session at the 2015 conference. the teaching of research data was based on using gis in three contexts: primary collection, digital data reuse and mined textual data. the examples from different disciplines addressed subjects ranging from climate change through reconstruction of history to data for villages in india. there is a strong argument that visualization and location with gis supports the understanding of data. further support of gis is found in the long list of learning outcomes experienced. the paper delivers a thorough description of the examples used in the problem-based teaching in the various disciplines articles for the iassist quarterly are always very welcome. they can be papers from iassist conferences or other conferences and workshops, from local presentations or papers especially written for the iq. when you are preparing a presentation, give a thought to turning your one-time presentation into a lasting contribution to continuing development. as an author you are permitted ‘deep links’ where you link directly to your paper published in the iq. chairing a conference session with the purpose of aggregating and integrating papers for a special issue iq is also much appreciated as the information reaches many more people than the session participants, and will be readily available on the iassist website at http://www. iassistdata.org. iassist quarterly 2015 5 iassist quarterlyiassist quarterly authors are very welcome to take a look at the instructions and layout: http://iassistdata.org/iq/instructions-authors authors can also contact me via e-mail: kbr@sam.sdu.dk. should you be interested in compiling a special issue for the iq as guest editor(s) i will also be delighted to hear from you. karsten boye rasmussen january 2016 editor iassisl quarterly 41 using spires: the icpsr experience the icpsr guide database contains information about the data collections in the icpsr archive. the database contains the data holdings section of the guide to resources and services , which is published annually. the quarterly additions and updates to the holdings that are announced in the icpsr bulletin and icpsr "hotline" are incorporated into the database after each announcement the icpsr variables database includes complete text for all questions used in selected surveys in the icpsr holdings. included at the time of this writing are euro-barometers 3-21, american national election studies, 1948-1984, cbs news/new york times polls, abc news/washington post polls and selected aging and health-related studies. in addition to question texts, variable codes and frequencies, where available, are included. by janet k. vavra' inter-university consortium for political and social research the university of michigan the inter-university consortium for political and social research (icpsr) has prepared a number of databases in the spires (stanford public information retrieval system) database management system. these databases are icpsr guide. icpsr variables, icpsr rollcalls and smis. below is a description of each of these databases. there is also a brief discussion of how each of the databases was developed. 'paper presented at the iniernauonal associaiion for social science information service and technolqgv (lasslst) conference held in marina del rey, california, may 21-24, 1986 the icpsr rollcalls database contains information on roll call votes taken in the last twenty years of the united states congress. in addition to the roll call items, yeas and nays are recorded for each roll call. the smis database has entries from the survey methodology information system which was developed by the bureau of the census. included are citations of journal publications. bureau of the census documents, publications, research reports, and conference papers that deal with methodological aspects of the design and conduct of surveys. there are citations in the database through 1982. the databases are part of several on-site services available to users through the consortium data network (cdnet). while giving users immediate access to information about the holdings, the databases also allow easier management of information about these holdings. cunently the archive has over 20,000 machine-readable files in approximately 1,400 study collections [these holdings represent over 6 million variables]. this number continues to fall 1986 42 iassist quarterly grow as nearly 150 titles are added to the holdings annually. without automation it is difficult to identify the collections in the holdings and to get any understanding of their contents. the icpsr guide, icpsr variables and icpsr roll calls databases were developed to help facilitate access to more complete and reliable information about the holdings. further, they seek to provide this access in a timely and economical manner. an appendix to this paper lists the indexes and elements found in each of these databases. all the databases can be accessed on-line through cdnet or can be installed locally either in spires or converted into other database management systems. database preparation all the input information used in each of the four icpsr databases was already machine-readable so there was no need to automate large quantities of raw information. however, as one would expect, none of the machine-readable files were in a form that could be input directly into the respective spires databases. the input files for each database had to be reformatted to make them acceptable to spires. the icpsr guide to resources and services has been a machine-readable text file since about 1973. this file is updated and used each year to prepare the annual hardcopy document ii was this file that was used as input for the icpsr guide database. the file was prepared for spires input with a series of editing procedures using the edit capability on the icpsr prime minicomputer. the edit procedures removed characters from the text that are unacceptable to spires, such as semi-colons (spires uses semi-colons as terminators), and replaced them with acceptable characters. using the print control characters that existed in the original file as a guide, the edit procedures inserted the appropriate spires element names and terminated the fields with semi-colons. these procedures handled most of the conversion into spires format the spires programs that enter the information into the database identified any remaining problems. these were handled on-hne as the entries were added to the database. information in the icpsr variables and icpsr rollcalls databases comes from osiris machine-readable codebooks. before these files are acceptable to spires, they must go through several steps. the first step is a simple set of editing procedures that search for unacceptable characters and replace them with acceptable ones. the second step creates several files which handle such conditions as references to question numbers, for which full question text is supplied eariier in the codebook, and references to notes found at the back of the codebook. this step results in three files: the question text file, a note file and the codebook file now without the notes at the end. the third and final step performs several operations: 1. it merges the appropriate question text into the variable description where references to the question text are made, and when it is possible to locate the question number in the codebook, 2. it inserts appropriate note information where references to given notes appear and 3. it generates the appropriate element names and delimiters to make the file acceptable to spires. final enor checking is done b\ the spires input programs as the reformatted codebook files are input into the database. fall j 986 iassist quarterly 43 the original information for the smis database was supplied on several magnetic tapes by the bureau of the census. the information on each tape was written onto a line file and editing procedures were set up to make the information compatible with spires. each file was scanned for characters unacceptable to spires. these were then changed to acceptable ones as they were encountered. in order to facilitate the conversion, icpsr retained the element names assigned by the bureau of the census and also followed the database structure for smis that had been developed by the bureau of the census. the edited files were used as input into the spires database. cunenlly the icpsr guide database contains approximately 1,400 entries, the icpsr variables database has over 47.000 entries, the icpsr rollcalls database has approximately 18,000 entries and the smis database contains over 7,000 entries. all of the databases remain dynamic as more records are added and as some existing ones are updated. database features in order to facilitate the searching of the databases, spires was instructed to create "indexes" for each database. the indexes consist of individual words that are found in a particular element or set of elements. the indexes are generated by spires as entries are added to a database. it should be noted that this type of index consuuction is not like a manually-generated classification scheme. since the indexes are "word" indexes, search terms (or values) cannot be longer than one word. two or more search commands, each using a word, are used when searching for a phrase. the databases accept truncated value searches, allowing the user to obtain a valid search result without complete information for a given item. a user may not know an author's full name, but by inserting a "#" sign at the end of the search command line, the user can instruct spires to save all records that begin with the value listed in the search command. for example, if the user is not certain whether an author's name is williams or williamson, by asking spires to retrieve all authors with the name williams#, the user will get all authors whose last name begins with williams including williamson. the databases accept search commands in either upper or lower case or in mixed mode. this frees the user from having to remember the appropriate mode for the database or rurming the risk of invalid search results simply because commands were given in the wrong case. selected special characters in the body of the text are ignored as a user searches the databases. this allows the user to locate all entries, for example, with california in the title, irrespective of whether california is listed alone or in brackets or parentheses. each database has custom formats that display the results of searches in a more readable form than the routine spires default format output can further be processed through these formats so that it is more readable by sorting a string of values, by forcing output from a given field to upper or lower case, by providing more meaningful headings for fields, or by rearranging the order in which fields are output so that the information fits together in a more logical manner. users who are not happy with the indexes that have been created can ignore index searching and switch to sequential searching. since in sequential searches all informauon in each entry is examined in sequence to determine whether or not it meets the search criteria, this method of searching is more time consuming and expensive. frequently users switch to sequential fall j 986 44 iassisl quarterly searching as part of an index search rather than in place of the index search. being able to use both modes of searching gives the user more flexibility and also difterent ways to test the results of searches. conclusion the icpsr experience with the spires database management system has been favorable. the capacity of the system is basically unlimited so that databases can continue to expand without much fear that the maximum number of entries will be exceeded. in fact, it is more likely that the capacities of the storage devices will sooner be exhausted. the system has many capabilities and yet can be used quickly and effectively by inexperienced users. finally, it has not been an especially difticult or expensive task to prepare existing machine-readable information into input files that are acceptable to spires.n fall j 986 iassist quarterly 45 appendix icpsr spires databases below is a list of the four icpsr spires databases that are available through cdnet. a brief description of each database, a list of the indexes which can be used to search each one and a list of the elements found in each is included. icpsr guide database archival holdings section of the guide to resources and services issued and collections announced in bulletins that have been issued since the publication of the guide. database is most up-to-date list of icpsr data holdings. goal records: collection simple index: ad, added simple index: up, updated simple index: t, titlew, titleword, tw, tword simple index: s, subjectword, subword, su, sword simple index: p, piauth. piname, piw, piword, pw key : studyno, icpsrno, idno. mrdf.num, studno, 037a element: date-added, da, dad, dadd , dateadd element: date-updated, dateup, du, dup, dupd element: investigator, invest, pi, prin. invest, 199a element: title, ti , titl, 245a element: summary, desc, descrip, sum, 520b element: subject. term, s.term, sterm, sub, sub. term, subj, subj.term, subject, 653a element: arch. filter, a. filter, afilter, arch, archfilter, archive, filter structure: element : element : element : element : element : element : element : classif icpsr. classifl, classifl, icpsr. classif2, classif2, icpsr. class1f3, 972a 972b classif3, 972c element : element : element : element : element : icpsr. class if4, classifa, 972d icpsr. class1f5, classif5, 972e extent. collect, e.col, e. collect, ecol, ecollect, ext, extent, extent. file, extentcollect, files, 300a series. name, s.name, ser, ser.name, series, seriesname, sname , 440a series. info, s.info, ser. info, seriesinfo, sinfo, 500a restrictions, limitations, limits, rest, restrict, 506a data. type, d.type, datatype, dtype, 516a time. period, chron. cover, t. period, timeperiod, tper, tperiod, 523a date. of. collect, collect. date, d.coll, d. collect, datadates, datesofcollect, dcoll, dcollect, 523b fall 1986 46 iassist quarterly element: funding. agtncy , f. agency, fagency , fund, funding, fundingagency, sponsor, 536a element: grant. number, g. number, gno, gnumber, grant, grant. num, grantno, grantnumber element: data. source, d. source, datasource, dsource, source. data, 537a element: data. format, d.form, d. format, dform, dformat, form, format, 538a element: collect. note, c.n, c.note, cnote. coll. note, collectnote, collnote element: sampling, sam, samp, 567a element: universe, uni, univ, 567b element: related. pubs , primary. pub, primary. pubs, r.pubs, rel.pub, rel.pubs, rpubs, 581a element: classno, class, classnum, cno, cnum, icpsr. class, 962a structure: part element: partno, pno, 245p element: part. name, p. name, partname. pname, 565p element: file. struct, f. struct, filestruct, fstruc, fstruct, struc, struct, 538s element: case. count, c. count, casecount, cases, ccount, 565a element: variable. count, v. count, var, variablecount, variables, vars , vcount, 565b element: lrecl element: records. per. case, r.case, r.per.c, rcase, reg. per. case, records. case, recordspercase, recpercase, recs. per. case , recspercase, rpc, rperc, 565c fall j 986 iassist quarterly 47 appendix icpsr variables database descriptions and codes of variables found in the american national election studies 1948-1984, euro-barometers 3-21, quality of american life, 1971 and 1978, media polls, and healthand aging-related surveys.... database under development as studies continue to be added. goal records: variable simple index: icpsrnum, id, studynum simple index: ad, added simple index: up, updated simple index: chronological, coverage, period. time, timeperiod simple index: s, sub, subject, v, var, variable key : id-num element: date-added, da, dadd, dateadd element: date-updated. dateup, du, dup element: studynum, sn, snum element: studydsnum, dataset, ds element: varmum, v*, vnum element: varname, vn, vname element: time-period, time, years element: vardesc, desc, description, q. question element: varcodes, code, codes element: varfreqs, freq, freqs element: study-name fall 1986 48 iassist quarterly appendix icpsr rollcalls database variables found in the machine-readable collection of the united states congressional roll call voting records [icpsr 0004]. contains most recent congressional sessions. database under development as congresses continue to be added. goal records: roll. call simple index: ds, icpsrds, id simple index: ad, added simple index: up, updated simple index: chronological, coverage, period, time, timeperiod simple index: s, sub, subject simple index: date. voted, v, voted key : id-num element: date-added, da, dadd, datkadd element: date-updated, dateup, du , dup element: studynum, sn, snum element: studydsnum, dataset, ds element: dsdates, dsd element: varnum, v» , vnum element: varnahe, vn, vname element: rcsource, rcs, sou element: rcdate, rcd, rcdat element: rcnum, rcn element: rcproposer, rcp, rcprop element: rcvotes, rcv, votes element: vardesc, desc, description, m, motion element: rcnotenum element: note-text, nn, note fall 1986 iassist quarterly 49 appendix smis database contains entries from the smis (survey methodology information system) that was originally developed by the bureau of the census. included are journal publications. bureau of the census publications and documents, research reports, etc. that deal with methodological aspects of the design and conduct of surveys. goal records: title simple index: ad, added simple index: up, updated simple index: t, titleword, tl, tw simple index: s, sub, subjectword, subword simple index: a, auth, author simple index: c, con, confer, conference key element element element element element element element element element element element element element element element element element element element element element element element e 1 eme n t element element element dt no date-added date-updated t-title, t dt-doctype , kw-keyword, kw d-date, d cj-litcode, cj nr-no.of.refs, nr v-volume, v i-issue, 1 pi-pages, pi fg-financ. source, fg fc-agen.cntrl.no, fc fp-project.no, fp ptpart. number, pt ln-language, ln se-series.name, se df-document.form, df av-availability, av dl-distrib. limit, dl rs-agency.rpt.no, rs rm-monitor.no, rm ro-distri. rpt.no, ro ab-abstract, ab cn-conference, cn pu-publisher, pu vr structure: au element: na-author, na element: ra-editor, ra element: on-authaffil , on fall 1986 50 iassist quarterly appendix structure: parent element: mt-soukc .title, mt element: md-sourc.date, md element: mkw-sourc .keywrd, mkw element: mnr-sourc.no.ref, mnr element: mv-sourc. volume, mv element: mi-sourc. issue , mi element: mp i-sourc. pages , mpi element: mfg-sourc.suppor, mfc element: mse-sourc. series, mse element: mdf-sourc.docfrm, mdf element: mdl-distrib .lim, mdl element: mav-sourc. avail, mav element: mrs-sourc.agenno, mrs element: mrm-sourc.mon, mrm element: mro-sourc.distri, mro element: mab-sourc.abstr, mab element: mcn-sourc. confer, hon element: mpu-sourc.pub, mpu element: hpt-sourc.partno, mpt ' element: mln-sourc.lang, mln element: mfc-sourc.ctrlno, mfc structure: mau element: mna-sourc. author, mna element: mra-sourc. editor, hra element: mon-sourc.affili, mon fall 1986 1/11 chenoweth, megan & john kubale (2025). using common data elements to foster interoperability of research on health disparities, iassist quarterly 49(2), pp. 1-11. doi: https://doi.org/10.29173/iq1112 the creative commons-attribution-noncommercial license 4.0 international applies to all works published by iassist quarterly. authors will retain copyright of the work and full publishing rights. using common data elements to foster interoperability of research on health disparities megan chenoweth1 and john kubale2 abstract common data elements (cdes) are standardized questions, variables, or measures with specific sets of responses that are used across multiple studies. they are organized around a particular research topic or question, validated, and defined via a consensus building process. their use fosters comparability of results and findings across studies. cdes are more common in national institutes of health (nih)-funded clinical and biomedical research than in social, behavioral, and economic (sbe) research. yet the community-driven, consensus-building approach to defining cdes makes them well suited to measuring complex social phenomena. the social, behavioral, and economic covid coordinating center at icpsr (sbe ccc) is leading the effort to establish cdes for sbe research into the effects of the covid-19 pandemic. we are collaborating with fifteen nih-funded research teams who are examining pandemic-related health disparities related to race, ethnicity, sex, geography, income, and other factors. in this article, we discuss ways in which cdes support research into health disparities and describe our process for identifying, validating, and building consensus on cdes related to covid public health policies. keywords common data elements, health disparities, research interoperability, social science research introduction common data elements (cdes) are standardized questions, variables, or measures with specific sets of responses that are used across multiple studies to ensure consistent data collection (national library of medicine, no date). cdes originated in clinical health research, and they remain more widely used in some of the clinical sciences than in the social sciences (sheehan et al., 2016). initiatives such as the phenx toolkit3, radx-up4, and the all of us research program5 have begun to expand cdes from the clinical sciences into the social sciences. in particular, the covid-19 pandemic has offered an opportunity to rapidly develop cdes for understanding the experiences of covid-19 (krzyzanowski et al., 2021; carrillo et al., 2022; gleason, tamburro and signore, 2023). to date, however, cdes are still less common within social, behavioral, and economic (sbe) research than clinical or biomedical research. this is particularly true when it comes to studying the effects of covid-19 beyond diagnosis and symptoms, for example, the economic or educational impacts of mitigation policies implemented by federal, state, and local governments in 2020. to support interoperability and consistency in sbe research into covid-19, the national institutes of health (nih) established the social, behavioral, and economic covid coordinating center (sbe ccc). administered by icpsr, a social science research data repository at the university of michigan https://doi.org/10.29173/iq1112 https://www.phenxtoolkit.org/ https://radx-up.org/ https://allofus.nih.gov/ 2/11 chenoweth, megan & john kubale (2025). using common data elements to foster interoperability of research on health disparities, iassist quarterly 49(2), pp. 1-11. doi: https://doi.org/10.29173/iq1112 institute for social research, sbe ccc was established in 2021 as a coordinating center to foster collaboration across a consortium of 15 nih-funded research teams (the sbe covid consortium) examining different social, behavioral, and economic outcomes related to the covid-19 pandemic. a primary aim of sbe ccc is promote the use of cdes in social science research on covid, including the creation and establishment of new cdes for sbe research. in this article, we will provide an introduction to the purpose and structure of cdes, discuss their potential applications to sbe research, and describe sbe ccc’s process for establishing new cdes for use in measuring covid-19 mitigation policies. about cdes the benefits of cdes for standardizing data collection and improving reproducibility of studies are widely recognized, especially within the clinical sciences (rubinstein and mcinnes, 2015; lapinlampi et al., 2017; kush et al., 2020). in the absence of standardization efforts, it is common and likely for key concepts in scientific research to be measured differently across studies; examples can be found in a variety of topics such as human anatomy (ioannou et al., 2011), social isolation in aging (evans et al., 2019), and the covid-19 era labor market (maas, 2022). cdes make it possible to compare results across studies, resulting in research that is more interoperable. in some cases, when using cdes it becomes possible to combine results across small studies, making these studies collectively more meaningful and impactful. also, they save time, money, and effort during data collection because they offer ready-made questions and responses instead of requiring researchers to create their own. cdes also have their limitations. for example, uptake is limited, and gaps in subject matter coverage exist. also, because cdes have arisen from different nih-funded centers and initiatives, they can be siloed. different cdes from different sources may exist to measure the same concept, without accompanying guidance on how to select the most appropriate cde for a given research question (kush et al., 2020). despite these limitations, many researchers and funders recognize the value that cdes bring to research studies. several nih centers encourage and sometimes require the use of cdes in funded research (meeuws et al., 2020; wandner et al., 2022). in 2015, the national library of medicine (nlm) created a searchable repository6 of cdes to bring together cdes established by multiple different centers and initiatives across nih. figure 1 illustrates a typical common data element from the nlm repository. a typical cde might be a single question from a survey questionnaire or a component of a clinical protocol. each cde contains a definition, standardized wording for the question, a data type, and for value lists, a set of permissible responses. concepts represented by the question itself and its permissible values are typically linked to concept identifiers found in controlled vocabularies like the national cancer institute thesaurus7. some cdes are part of forms (also known as bundles) or sets of data elements that are only valid when asked as a set. if a cde is part of a form, it is flagged that way in the cde repository, with a link to the other questions in the same form. https://doi.org/10.29173/iq1112 https://cde.nlm.nih.gov/ https://ncithesaurus.nci.nih.gov/ncitbrowser/ https://ncithesaurus.nci.nih.gov/ncitbrowser/ 3/11 chenoweth, megan & john kubale (2025). using common data elements to foster interoperability of research on health disparities, iassist quarterly 49(2), pp. 1-11. doi: https://doi.org/10.29173/iq1112 figure 1. screenshot from the nlm cde repository showing a common data element for the concept "employment status." the screenshot shows the standard question text, definition, possible values, and links to controlled vocabulary. the path from data element to nlm-endorsed cde typically follows three steps. first, cdes originate not as abstract concepts, but as variables in research studies. one common origin for cdes is in existing instruments or scales which have already been validated and are widely used within a given discipline, such as the kessler screening scale for psychological distress (k6) (kessler et al., 2010). for measures of new and emerging concepts where validated measures do not already exist, as is the case for covid-19, cdes can be derived from other sources. one example is the set of cdes for the study of covid-19 taken from the radx-up initiative8 in 2021. the next step in defining cdes for a given research area is to convene a working group of individuals with expertise in that area. working group members are subject matter experts and leaders in their fields. their role is to establish consensus on which data elements are the best candidates to become cdes for a given domain of study. this consensus building phase is typically followed by a process for gathering feedback on proposed cdes. redeker et al. (2015) provides a useful case study for the process of creating cdes for the research domain of symptom science, a field of interest within nursing research. the authors describe how a group of experts from the national institute for nursing research worked as a group to select a list of symptoms (e.g. pain, fatigue), identify validated measures for these symptoms (e.g. the promis pain and promis fatigue scales), and https://doi.org/10.29173/iq1112 https://cde.nlm.nih.gov/cde/search?selectedorg=radx-up 4/11 chenoweth, megan & john kubale (2025). using common data elements to foster interoperability of research on health disparities, iassist quarterly 49(2), pp. 1-11. doi: https://doi.org/10.29173/iq1112 establish consensus on which measures were best suited to become cdes (e.g. based on appropriateness for study aims, cost, and participant burden). the final step to becoming a cde is review by nih’s cde governance committee. typically, this is initiated either by an nih-sponsored project or within nih. but in some instances, if there is a mission-critical gap and available funds, nih-funded investigators could potentially work with their program officers to submit cdes to the nih cde governance committee for review and inclusion in the nlm cde repository as nih-endorsed cdes. cdes address the four key principles of open science and of good data management – findability, accessibility, interoperability, and reusability (fair)9 – at multiple stages of the research data lifecycle. during the study design phase, a researcher can search the cde repository to find commonly used variables to incorporate into their study. using measures that are common across studies improves the interoperability of their work and enhances the reusability of their study by making their data suitable for analysis in comparison with other studies or for use in meta-analyses. during the data sharing and preservation stage of a study, novel measures created or identified in study design can be submitted to the nlm cde repository. this supports findability and accessibility by making new measures available for researchers while again supporting interoperability and reusability by promoting inclusion of those measures into new studies. potential for cdes in studying health disparities cdes originated in clinical research and are more commonly found in some (though not all) clinical disciplines. they were not created for the purpose of measuring complex social phenomena. figure 2 illustrates this history by showing the distribution of cdes by their source – the center or initiative within nih that was responsible for adding them to the nlm cde repository. of the centers and initiatives represented in this figure, many major contributors of cdes – national institute of neurological disorders and stroke (ninds), national cancer institute (nci), and national heart, lung, and blood institute (nhlbi) – are in the clinical sciences. only the eunice kennedy shriver national institute of child health and human development (nichd) has funded social and behavioral research that has helped establish cdes. https://doi.org/10.29173/iq1112 5/11 chenoweth, megan & john kubale (2025). using common data elements to foster interoperability of research on health disparities, iassist quarterly 49(2), pp. 1-11. doi: https://doi.org/10.29173/iq1112 figure 2. bar graph illustrating the relative contributions of cdes by source. sources correspond to nih centers or initiatives. (ninds = national institute of neurological disorders and stroke. loinc = logical observation identifiers, names, and codes. nhlbi = national heart, lung, and blood institute. nci = national cancer institute. promis = patient-reported outcomes measurement information system. nichd = eunice kennedy shriver national institute of child health and human development) this disparity in cde uptake is understandable to a certain extent given the different nature of the concepts being studied in the clinical vs. social sciences. while biomedical concepts such as, for example, blood pressure and body weight pose their own challenges for measurement and interpretation, their definition is more straightforward than many concepts of interest in the social sciences, such as anxiety or wealth. yet despite this challenge, social science researchers have recognized the benefits of standardized measures and worked to establish shared measures. examples include the diagnostic and statistical manual of mental disorders (dsm) in psychology and the north american industry classification system (naics) in economics. similarly, cdes bring value to the study of health disparities and the complex concepts that must be measured in social research into health disparities in several ways. first, cdes support interoperability in research, even when applied to hard-to-measure concepts. to consider a few examples, studies of covid-related health disparities may entail defining and measuring sociodemographic characteristics (race, gender identity, essential worker status), mental or behavioral health outcomes (wellbeing, stress), and policy interventions (mask mandates, economic assistance). each of these concepts represents one or more variables that needs to be defined and quantified before analysis is possible. studies that measure concepts in incompatible ways may reach different and even conflicting conclusions simply because of those differences in measurement. to give an example, two studies of the impact of essential worker status on mental health during covid-19 might reach different conclusions if one uses self-reported essential worker status while another uses job-based industry codes. even if these two ways of defining essential worker status contain inherent limitations, use of the same cde or shared measure can reduce the 0 2000 4000 6000 8000 10000 12000 14000 16000 18000 20000 ninds loinc nhlbi promis / neuro-qol nci nlm nichd other n u m b er o f c d es cde source (as specified in nlm repository) number of cdes by source https://doi.org/10.29173/iq1112 6/11 chenoweth, megan & john kubale (2025). using common data elements to foster interoperability of research on health disparities, iassist quarterly 49(2), pp. 1-11. doi: https://doi.org/10.29173/iq1112 likelihood that differences found by these two studies are due to differences in measurement of essential worker status. second, cdes support specificity of research measures for a given research topic. several of the examples provided above represent complex identities, such as race and gender identity, that require a great deal of nuance in definition and measurement. measuring these topics in the same way across all research domains and contexts may not be appropriate. because cdes are domainspecific, they can help researchers identify the best measures for their research domain. finally, cdes provide a mechanism for establishing and sharing consensus. definitions of social phenomena shift in response to changes in consensus, understanding, and best practices. cdes arise from consensus within a given research domain, offering a process for capturing current consensus on how to measure these concepts and disseminating it across studies. furthermore, while efforts to date have focused primarily on creating new cdes, it is possible in the long term to follow these same processes to revise cdes in response to evolving research. for example, the national institute of neurological disorders and stroke (ninds) maintains a form10 that can be used to recommend either new cdes or substantial revisions to existing cdes. cdes for social, behavioral, and economic covid research one of the first projects launched by sbe ccc was to define a set of cdes for capturing data about covid-19 mitigation policies undertaken by state and local governments during the covid-19 pandemic. these policies include mask mandates, school closures, business closures, and food and rental assistance programs. defining cdes for policy measures was a critical first effort because health policy in general, and covid mitigation policies in particular, have been shown to affect health and to do so inequitably depending on race, income, employment, health insurance status, and other factors (koppaka, 2011; thompson, 1993; nana-sinkam et al., 2021; park, 2021). covid mitigation policies were also not equally distributed geographically during the pandemic.11 nor were they static, as policies changed over time. finally, the tracking of these policies was done through many data sources, each measuring these policies in different ways. all of these factors undercut the ability to harmonize measures and reach consensus regarding the effectiveness of specific policy interventions. the process to define cdes for covid mitigation policies was as follows: 1. working with the 15 research teams that comprise the sbe covid consortium, sbe ccc conducted an inventory of policy measures used in their studies, focusing on: a. types of policies being measured, b. level of geography (usually state or county), c. data source(s) used. 2. we searched published literature for other articles measuring similar concepts, identified additional covid mitigation policy measures and data sources, and combined them with the inventory results to create a list of covid mitigation policy measures and the data sources they were derived from. 3. we documented common concepts found across measures and data sources. for example, one study contained a variable “days of exposure to gym closure mandate” measured in number of days. another study contained a similar variable, “gym closure in effect as of https://doi.org/10.29173/iq1112 https://ninds.nih.gov/ninds-cde-project-request-form 7/11 chenoweth, megan & john kubale (2025). using common data elements to foster interoperability of research on health disparities, iassist quarterly 49(2), pp. 1-11. doi: https://doi.org/10.29173/iq1112 [date]” with values ranging from 0 (no restrictions) to 3 (closed). both measures imply that the underlying data pertains to gym closures, has a start an end date, is a mandate rather than a recommendation, and is at the state level. 4. we held a meeting with sbe covid consortium members to review findings and obtain feedback on the completeness of our list of common concepts. based on this discussion, we added some details to the list of concepts identified in the literature review; for example, a policy might apply to only a portion of the population (e.g. unvaccinated individuals) rather than everyone. 5. we drafted a list of cdes based on the common concepts identified in steps 3 and 4. each cde included a variable name and label, question text, a predefined format, and possible values. for example, the cde for “policy type” has question text that reads “what type of covid-19 mitigation policy was enacted?” and possible values “mask policy,” “social distancing policy,” and “business closure policy,” among others. 6. draft cdes were shared again at a meeting of sbe covid consortium members and nih stakeholders for final validation. 7. we worked with our nih program officer to submit the cdes to nlm’s repository. 8. nlm’s cde governance committee approved the submission, and sbe ccc’s covid mitigation policy cdes were published in the nlm cde repository in july 2024. table 1 lists cde names, their brief definition, and an example of the format or some possible values for each. detailed information about each cde, including question text and a complete list of possible values, can be found in the nlm cde repository12. and on sbe ccc’s website.13 common data element definition example values start date the date on which a covid-19 mitigation policy was enacted or went into effect. 4/1/2020 end date the date on which a covid-19 mitigation policy was repealed, superseded, or invalidated. 12/31/2020 geographic level the type of governing body or jurisdiction (in the u.s.) that is implementing the policy, e.g. state government, county government, or school district). state government county government school district coverage area the full name of the jurisdiction (e.g. the name of the state, county, or school district) to which the mitigation policy applies. california alameda [county] oakland unified [school district] policy type the type of covid-19 mitigation policy that may be enacted by governments and municipalities. mask policy business closure work from home policy target population the population and/or group(s) for whom a covid-19 mitigation policy is intended to affect. all individuals employees of businesses unvaccinated people https://doi.org/10.29173/iq1112 https://cde.nlm.nih.gov/formview?tinyid=euqjdxif8k https://cde.nlm.nih.gov/formview?tinyid=euqjdxif8k https://www.icpsr.umich.edu/files/sbeccc/sbe-ccc-common-data-elements-policy.pdf 8/11 chenoweth, megan & john kubale (2025). using common data elements to foster interoperability of research on health disparities, iassist quarterly 49(2), pp. 1-11. doi: https://doi.org/10.29173/iq1112 setting the location or environment in which the covid-19 mitigation policy applies. all retail businesses restaurants schools regulation type the type of regulation a reflected in a covid-19 mitigation policy (i.e., a mandate, guidance, recommendation, or limitation/restriction on future policies). requirement recommendation restriction table 1. covid mitigation policy cdes created by sbe ccc. unlike many efforts to define cdes, which employ existing questions or previously validated measures, sbe ccc’s approach entailed identifying common concepts found in datasets used to collect data on covid mitigation policies. we adopted this approach for two reasons. one practical reason is that covid mitigation policies were a novel phenomenon and there were no previously validated measures or consensus on measurement best practices. another reason was that this approach allowed us to identify commonalities across data sources and reveals some underlying harmonies across variables that differ in question text and possible values. sbe ccc found this approach useful and worth applying to future efforts to define new cdes for social, behavioral, and economic disciplines. next steps and future research throughout the process of defining and publishing our covid mitigation policy cdes, sbe ccc has conducted outreach to researchers through social media, mailing lists, and presentations at conferences such as the society for epidemiological research annual meeting. one opportunity for future research is to review data sources identified in the inventory of policy measures and compare how completely they capture the concepts outlined in our policy cdes. sbe ccc’s cdes provide a framework for highlighting commonalities, differences, and harmonization challenges across data sources. a side-by-side comparison of the policy datasets identified during our inventory will help future researchers foresee harmonization issues and select policy datasets that are best suited to their research questions. additionally, sbe ccc plans to review the existing cdes in the nlm repository that relate to covid, with the aim of identifying overlap and as a next step in establishing consensus on preferred measures. finally, sbe ccc has launched a covid measures archive, a searchable database of variables used in social, behavioral, and economic studies of the covid pandemic14. measures come from the 15 sbe covid consortium members, as well as other major studies that incorporated questions about covid. measures in the covid measures archive can be used to identify commonalities and differences across studies, and to identify more opportunities for novel measures to be promoted as cdes in nlm’s repository of cdes. funding acknowledgement research reported in this publication was supported by the national institute on aging and the office of behavioral and social science research of the national institutes of health under award number u24ag076462. references carrillo, g. a., cohen-wolkowiez, m., d’agostino, e. m., marsolo, k., wruck, l. m., johnson, l., topping, j., richmond, a., corbie, g., & kibbe, w. a. (2022). standardizing, harmonizing, and protecting data collection to broaden the impact of covid-19 research: the rapid https://doi.org/10.29173/iq1112 https://www.icpsr.umich.edu/web/sbeccc/search/variables 9/11 chenoweth, megan & john kubale (2025). using common data elements to foster interoperability of research on health disparities, iassist quarterly 49(2), pp. 1-11. doi: https://doi.org/10.29173/iq1112 acceleration of diagnostics-underserved populations (radx-up) initiative. journal of the american medical informatics association, 29(9), 1480–1488. https://doi.org/10.1093/jamia/ocac097 evans, i. e. m., martyr, a., collins, r., brayne, c., & clare, l. (2019). social isolation and cognitive function in later life: a systematic review and meta-analysis. journal of alzheimer’s disease, 70(s1), s119–s144. https://doi.org/10.3233/jad-180501 gleason, j. l., tamburro, r., & signore, c. (2023). promoting data harmonization of covid-19 research in pregnant and pediatric populations. jama, 330(6), 497. https://doi.org/10.1001/jama.2023.10835 ioannou, c., sarris, i., salomon, l. j., & papageorghiou, a. t. (2011). a review of fetal volumetry: the need for standardization and definitions in measurement methodology. ultrasound in obstetrics & gynecology, 38(6), 613–619. https://doi.org/10.1002/uog.9074 kessler, r. c., green, j. g., gruber, m. j., sampson, n. a., bromet, e., cuitan, m., furukawa, t. a., gureje, o., hinkov, h., hu, c., lara, c., lee, s., mneimneh, z., myer, l., oakley‐browne, m., posada‐villa, j., sagar, r., viana, m. c., & zaslavsky, a. m. (2010). screening for serious mental illness in the general population with the k6 screening scale: results from the who world mental health (wmh) survey initiative. international journal of methods in psychiatric research, 19(s1), 4–22. https://doi.org/10.1002/mpr.310 koppaka, r. (2011). ten great public health achievements—united states, 2001—2010. morbidity and mortality weekly report (mmwr), 60(19). https://www.cdc.gov/mmwr/preview/mmwrhtml/mm6019a5.htm krzyzanowski, m. c., terry, i., williams, d., west, p., gridley, l. n., & hamilton, c. m. (2021). the phenx toolkit: establishing standard measures for covid‐19 research. current protocols, 1(4), e111. https://doi.org/10.1002/cpz1.111 kush, r. d., warzel, d., kush, m. a., sherman, a., navarro, e. a., fitzmartin, r., pétavy, f., galvez, j., becnel, l. b., zhou, f. l., harmon, n., jauregui, b., jackson, t., & hudson, l. (2020). fair data sharing: the roles of common data elements and harmonization. journal of biomedical informatics, 107, 103421. https://doi.org/10.1016/j.jbi.2020.103421 lapinlampi, n., melin, e., aronica, e., bankstahl, j. p., becker, a., bernard, c., gorter, j. a., gröhn, o., lipsanen, a., lukasiuk, k., löscher, w., paananen, j., ravizza, t., roncon, p., simonato, m., vezzani, a., kokaia, m., & pitkänen, a. (2017). common data elements and data management: remedy to cure underpowered preclinical studies. epilepsy research, 129, 87– 90. https://doi.org/10.1016/j.eplepsyres.2016.11.010 maas, s. (2022). pandemic school closures and parents’ labor supply. nber digest, 4. https://www.nber.org/digest/202204/pandemic-school-closures-and-parents-labor-supply meeuws, s., yue, j. k., huijben, j. a., nair, n., lingsma, h. f., bell, m. j., manley, g. t., & maas, a. i. r. (2020). common data elements: critical assessment of harmonization between current https://doi.org/10.29173/iq1112 https://doi.org/10.1093/jamia/ocac097 https://doi.org/10.3233/jad-180501 https://doi.org/10.1001/jama.2023.10835 https://doi.org/10.1002/uog.9074 https://doi.org/10.1002/mpr.310 https://www.cdc.gov/mmwr/preview/mmwrhtml/mm6019a5.htm https://doi.org/10.1002/cpz1.111 https://doi.org/10.1016/j.jbi.2020.103421 https://doi.org/10.1016/j.eplepsyres.2016.11.010 https://www.nber.org/digest/202204/pandemic-school-closures-and-parents-labor-supply 10/11 chenoweth, megan & john kubale (2025). using common data elements to foster interoperability of research on health disparities, iassist quarterly 49(2), pp. 1-11. doi: https://doi.org/10.29173/iq1112 multi-center traumatic brain injury studies. journal of neurotrauma, 37(11), 1283–1290. https://doi.org/10.1089/neu.2019.6867 nana-sinkam, p., kraschnewski, j., sacco, r., chavez, j., fouad, m., gal, t., auyoung, m., namoos, a., winn, r., sheppard, v., corbie-smith, g., & behar-zusman, v. (2021). health disparities and equity in the era of covid-19. journal of clinical and translational science, 5(1), e99. https://doi.org/10.1017/cts.2021.23 national library of medicine. (n.d.). common data elements: standardizing data collection. https://www.nlm.nih.gov/oet/ed/cde/tutorial/index.html park, j. (2021). who is hardest hit by a pandemic? racial disparities in covid-19 hardship in the u.s. international journal of urban sciences, 25(2), 149–177. https://doi.org/10.1080/12265934.2021.1877566 redeker, n. s., anderson, r., bakken, s., corwin, e., docherty, s., dorsey, s. g., heitkemper, m., mccloskey, d. j., moore, s., pullen, c., rapkin, b., schiffman, r., waldrop‐valverde, d., & grady, p. (2015). advancing symptom science through use of common data elements. journal of nursing scholarship, 47(5), 379–388. https://doi.org/10.1111/jnu.12155 rubinstein, y. r., & mcinnes, p. (2015). nih/ncats/grdr® common data elements: a leading force for standardized data collection. contemporary clinical trials, 42, 78–80. https://doi.org/10.1016/j.cct.2015.03.003 sheehan, j., hirschfeld, s., foster, e., ghitza, u., goetz, k., karpinski, j., lang, l., moser, r. p., odenkirchen, j., reeves, d., rubinstein, y., werner, e., & huerta, m. (2016). improving the value of clinical research through the use of common data elements. clinical trials, 13(6), 671–676. https://doi.org/10.1177/1740774516653238 thompson, l. v. (1993). the social context of health-related behavi. american sociological association. wandner, l. d., domenichiello, a. f., beierlein, j., pogorzala, l., aquino, g., siddons, a., porter, l., & atkinson, j. (2022). nih’s helping to end addiction long-termsm initiative (nih heal initiative) clinical pain management common data element program. the journal of pain, 23(3), 370–378. https://doi.org/10.1016/j.jpain.2021.08.005 endnotes 1 megan chenoweth is a data project manager at icpsr at the university of michigan. she can be reached at mmchenow@umich.edu. 2 john kubale is a research assistant professor at icpsr at the university of michigan. 3 https://www.phenxtoolkit.org/ 4 https://radx-up.org/ 5 https://allofus.nih.gov/ https://doi.org/10.29173/iq1112 https://doi.org/10.1089/neu.2019.6867 https://doi.org/10.1017/cts.2021.23 https://www.nlm.nih.gov/oet/ed/cde/tutorial/index.html https://doi.org/10.1080/12265934.2021.1877566 https://doi.org/10.1111/jnu.12155 https://doi.org/10.1016/j.cct.2015.03.003 https://doi.org/10.1177/1740774516653238 https://doi.org/10.1016/j.jpain.2021.08.005 mailto:mmchenow@umich.edu https://www.phenxtoolkit.org/ https://radx-up.org/ https://allofus.nih.gov/ 11/11 chenoweth, megan & john kubale (2025). using common data elements to foster interoperability of research on health disparities, iassist quarterly 49(2), pp. 1-11. doi: https://doi.org/10.29173/iq1112 6 https://cde.nlm.nih.gov/ 7 https://ncithesaurus.nci.nih.gov/ncitbrowser/ 8 https://cde.nlm.nih.gov/cde/search?selectedorg=radx-up 9 for a summary of fair data principles, see https://force11.org/info/the-fair-data-principles/ 10 https://ninds.nih.gov/ninds-cde-project-request-form 11 as an example, see the map of mask mandate policies by state as of november 9, 2020: https://abcnews.go.com/health/states-mask-mandates-map/story?id=74168504 12 https://cde.nlm.nih.gov/formview?tinyid=euqjdxif8k 13 https://www.icpsr.umich.edu/files/sbeccc/sbe-ccc-common-data-elements-policy.pdf 14 https://www.icpsr.umich.edu/web/sbeccc/search/variables https://doi.org/10.29173/iq1112 https://cde.nlm.nih.gov/ https://ncithesaurus.nci.nih.gov/ncitbrowser/ https://cde.nlm.nih.gov/cde/search?selectedorg=radx-up https://force11.org/info/the-fair-data-principles/ https://ninds.nih.gov/ninds-cde-project-request-form https://abcnews.go.com/health/states-mask-mandates-map/story?id=74168504 https://cde.nlm.nih.gov/formview?tinyid=euqjdxif8k https://www.icpsr.umich.edu/files/sbeccc/sbe-ccc-common-data-elements-policy.pdf https://www.icpsr.umich.edu/web/sbeccc/search/variables using common data elements to foster interoperability of research on health disparities abstract keywords introduction about cdes potential for cdes in studying health disparities cdes for social, behavioral, and economic covid research next steps and future research funding acknowledgement references 1/17 l'hours, hervé; kleemola, mari; de leeuw, lisa (2019) coretrustseal: from academic collaboration to sustainable services, iassist quarterly 43 (1), pp. 1-17. doi: https://doi.org/10.29173/iq936 coretrustseal: from academic collaboration to sustainable services hervé l'hours, mari kleemola, lisa de leeuw 1 abstract national and international digital repositories must design and deliver sustainable services supporting a range of scientific and data management activities while reducing costs and avoiding duplication of effort. the coretrustseal, launched in 2017, defines requirements and offers core level certification for trustworthy digital repositories (tdr) holding data for long-term preservation. this paper traces the journey of the coretrustseal through the data seal of approval (dsa), icsu world data system (wds), research data alliance (rda) working groups and community engagement, toward becoming a sustainable service supporting global data infrastructure. we outline the design and delivery of the service, current activities, the benefits of certification to a range of communities, and future plans and challenges. as well as providing a historical narrative and current and future perspectives, the coretrustseal experience offers lessons for those involved in developing standards and best practices or seeking to develop cooperative and community-driven efforts bridging data curation activities across academic disciplines, governmental and private sectors. keywords trust, trustworthy digital repositories, tdr, certification, archives, preservation introduction the coretrustseal2 is a not-for-profit foundation that authors and maintains the 16 coretrustseal trustworthy digital repository (tdr) requirements, and the audit procedures and process necessary to attain coretrustseal tdr certification. coretrustseal is governed by the cts board that is drawn primarily from the assembly of reviewers, which in turn consists of volunteer reviewers designated by cts certified repositories. the coretrustseal provides a benchmark for those seeking assurance either for their repository or for the data they produce, own or use, to ensure the data will be actively preserved as digital assets for the long term. this paper presents the context and process the coretrustseal has taken toward providing a sustainable service. the board acknowledges that though much has been accomplished, further work remains. this paper seeks to engage in the spirit of openness and community that has created coretrustseal by sharing its experience so others may benefit from our experience in developing common standards, products and services for our data communities. beyond the formalisation of its requirements, processes and governance bodies, the coretrustseal remains a community-created, community-driven entity with all the complexity, collaboration, promise and compromise that entails. https://doi.org/10.29173/iq936 2/17 l'hours, hervé; kleemola, mari; de leeuw, lisa (2019) coretrustseal: from academic collaboration to sustainable services, iassist quarterly 43 (1), pp. 1-17. doi: https://doi.org/10.29173/iq936 standards and preservation tdr certification and the oais model the word preservation may not always be clearly defined, being sometimes confused with digitisation, or mistaken for good storage practice (harrower and cassidy, 2017). for the coretrustseal and related tdr standards, the underlying concepts which define an organisation capable of delivering long-term digital preservation are derived from their common reference: the open archival information system (oais) model3. the oais model makes it clear that a repository shall: ● obtain sufficient control of the information provided to the level needed to ensure long term preservation. ● determine, either by itself or in conjunction with other parties, which communities should become the designated community and, therefore, should be able to understand the information provided, thereby defining its knowledge base. ● ensure that the information being preserved is independently understandable to the designated community. in particular, the designated community should be able to understand the information without needing special resources such as the assistance of the experts who produced the information. (oais, 2012). the consultative committee for space data systems (ccsds) originally developed the oais reference model (ccsds 650.0-m-2 with a parallel standard process as iso147214) as part of its suite of standards for space data systems, but it became the de facto standard for a wider range of disciplines. at the time of its original publication in the late 1990s, oais documentation noted the need for some form of compliance certification. trustworthy repositories audit & certification (trac): criteria and checklist of the research library group (rlg)5 was developed as part of a wide consortium brought together by the national archives and records administration (nara) and rlg (giaretta 2011, 463). cssds then took this forward as the audit and certification of trustworthy digital repositories standard (iso16363)6. in germany, the network of expertise in long-term storage of digital resources awards the nestor seal against their criteria for trustworthy digital archives (kriterienkatalog vertrauenswürdige digitale langzeitarchive)7 (din31644). together the tdr standards trac/iso16363, nestor seal, coretrustseal (previously data seal of approval) all derive their key concepts from the oais’ responsibilities to a defined designated community. designated community it is not sufficient to ensure that data are stored in (and made available from) environments which ensure bit-level integrity checks, multi-copy/multi-site redundancy and low-risk disaster recovery methods. though these functions are vital nodes in our (research) data networks, trustworthy digital repositories are defined by ensuring the availability of data which is also understandable and usable by their designated community for the long term, ensuring the full value of data assets is assured over time. https://doi.org/10.29173/iq936 3/17 l'hours, hervé; kleemola, mari; de leeuw, lisa (2019) coretrustseal: from academic collaboration to sustainable services, iassist quarterly 43 (1), pp. 1-17. doi: https://doi.org/10.29173/iq936 the designated community is defined, in part, as an “identified group of potential consumers who should be able to understand a particular set of information” (oais, 2012). these requirements make it clear that even if repositories are open to the public, their designated community must almost certainly be more clearly and narrowly defined. the composition and needs of the designated community will change over time. new scientific discoveries, or changes to the common software tools they use, may necessitate a change to data provision through emulation or file format migration. any effective data steward must respond to changes in the knowledge base or technical requirements of their users. the trustworthy digital repository is designed to respond to these changes for the long term, ensuring data, metadata and documentation remain fit for purpose through each round of change. to obtain the necessary expertise to fulfil this mission, applicants for tdr are likely to be disciplinary repositories or other organizations focused on a defined collection, with a specific topical area, theme, or type of data. sharing expertise and effort there is one essential aspect of digital preservation that the oais model does not address directly: preservation of digital assets requires a lot of resources. this is captured by hedstrom (1998, 190) who defines digital preservation as ”the planning, resource allocation, and application of preservation methods and technologies necessary to ensure that digital information of continuing value remains accessible and usable”. giaretta (2011, 8) puts it more bluntly: “the really foolproof solution for digital preservation: money… enough of it, and for an indefinite period.” in a world with limited resources, one way to reduce costs is to share the expertise and the effort of preservation. since coretrustseal certifications are public, they provide a growing knowledge base of repository practice. certification also helps to identify the repositories one can trust to have expertise in long term preservation of digital assets and with whom one might wish to collaborate and thus share efforts and costs. the data seal of approval (dsa) sixteen guidelines for social science and humanities data when the netherlands’ data archiving and network services (dans) was established by the royal netherlands academy of arts and sciences (knaw) and the netherlands organisation for scientific research (nwo), they assigned it the task of developing a seal of approval for data, to ensure that archived data can still be found, understood and used in the future. in 2008 the first edition of the data seal of approval, written by laurents sesink, rené van horik and henk harmsen, was presented in the international conference on preservation of digital objects (harmsen, 2008). the criteria for the data seal of approval were aligned with national and international guidelines for digital data archiving such as nestor, trac, and digital repository audit method based on risk assessment (drambora)8 published by the digital curation centre (dcc) and digitalpreservationeurope (dpe). foundations of modern language resource archives of the max planck institute9 and stewardship of digital research data: a framework of principles and guidelines published by the research information network10 were also taken into account. in distilling the https://doi.org/10.29173/iq936 4/17 l'hours, hervé; kleemola, mari; de leeuw, lisa (2019) coretrustseal: from academic collaboration to sustainable services, iassist quarterly 43 (1), pp. 1-17. doi: https://doi.org/10.29173/iq936 minimum set of requirements from these sources the data seal of approval sought to ensure high quality and reliable management of data for the future without requiring the implementation of new standards, regulations or heavy investment.. the result was the sixteen dsa guidelines for the application and verification of quality aspects regarding the creation, storage and (re-)use of digital research data in the social sciences and humanities. these served as the basis for granting a ‘data seal of approval’ by the data seal of approval board. three stakeholder groups the guidelines were based on input from three stakeholder groups: data producers (quality of the research data content, formats, documentation), data repositories (storage quality, organisation of processes, technical infrastructure and assurance of availability), and data consumers (quality of data use in terms of access regulations, codes of conduct and licences). together the dsa guidelines were intended to (dsa 2010): ● “give researchers the assurance that their research results will be stored in a reliable manner and can be reused ● provide research sponsors with the guarantee that research results will remain available for reuse ● enable researchers a reliable means to assess the repository where research data are held. ● allow data repositories to archive and distribute research data efficiently” fundamental to the guidelines was that sustainable archiving entails research data are reliable, accessible on the internet in a usable format, can be referred to, and that relevant legislation with regard to personal information and intellectual property of the data is taken into account (dsa, 2010). international board the response from the research data community made it clear that the dsa guidelines were of relevance far beyond the netherlands and beyond the social sciences and humanities. in january of 2009, dans convened a workshop to transition the dsa governance to an international board for its further development. the initial international board consisted of dans (henk harmsen, laurents sesink, lisa de leeuw), icpsr (mary vardigan), uk data archive (matthew woollard), cines (olivier rouchon), mpi nijmegen (paul trilsbeek), nestor (natascha schumann) and the polar research centre (hans pfeiffenberger). this original dsa board was primarily european-based and with a tendency towards social science data though the community of seal recipients as a whole remained more diverse and international. under this board, the second revision of the data seal of approval was undertaken, leveraging lessons learned from the initial tranche of successful (and unsuccessful) applicants. in 2015, after a number of changes to those voluntarily serving on the board, the dsa seal holders were given the opportunity to vote for their representatives drawn from the growing dsa community. representatives of dans (ingrid dillo), strasbourg astronomical center (francoise genova), university college dublin (john howard), finnish social science data archive (mari kleemola), uk data archive https://doi.org/10.29173/iq936 5/17 l'hours, hervé; kleemola, mari; de leeuw, lisa (2019) coretrustseal: from academic collaboration to sustainable services, iassist quarterly 43 (1), pp. 1-17. doi: https://doi.org/10.29173/iq936 (hervé l’hours), cines (marion massol), gesis-leibniz institute for the social sciences (natascha schumann) and mpi nijmegen (paul trilsbeek) took their elected position in january 2016. beyond academic research data the data seal of approval envisaged that data demanded by the social sciences and humanities were of potentially broad provenance noting their guidelines were “of interest to researchers and institutions that create digital research files, to organizations that archive research files, and to users of research data” (dsa, 2010). interest from a wider academic sphere made it clear that the certification of digital archives is not only important for scientific archives of primary research data, but also for cultural heritage institutions such as public libraries, museums and archives. the big data revolution encompasses both the capacity for collection and for analysis of data at previously impossible scales. the range of data and the research projects themselves are increasingly heterogeneous. public-private research partnerships are increasingly common, and data originally conceived for other purposes, including governmental, administrative and other data including social media data, are increasingly used within research. these changes make a narrow definition of research data increasingly difficult to justify as data are created, curated, stored and used by a wider range of actors across the data/research data lifecycle. trust in these actors is increasingly critical to our trust in data generally and in scientific data specifically. with these changes in mind, the dsa board undertook a revision of the dsa guidelines in 2010. wording was clarified to encompass the full range of data of interest to the widest possible group of data users and guidance was extended to address the need for more specific examples. this exercise was repeated in 2013 resulting in the second version of the dsa guidelines (dsa 2014-2017). wider adoption of tdr principles the added value of the dsa process was first recognized by the approximate 60 individual repositories seeking or achieving dsa at the time. other adopters included the european research infrastructure consortia (eric) as defined by the european strategy forum on research infrastructures (esfri)11 within the european research area (era) and innovation union. they encompass: major scientific equipment, resources such as collections, archives or scientific data, e-infrastructures such as data and computing systems, and communication networks. the erics represent a broad range of potential applicants covering a wide range of subject, disciplinary and scope initiatives. not all of them are directly comparable from technical or workflow perspectives, but all seek a trust relationship within and between their components (l’hours et al, 2018). in this context infrastructures such as cessda12, clarin13 and dariah14 used dsa, and are using the coretrustseal requirements. clarin has made certification mandatory for a large part of its centres. all cessda service providers must seek certification, and the dsa/coretrustseal requirements are mapped directly to their common statutes where possible. dariah is using the requirements in their assessment of national contributions to the infrastructure. the emergence of the fair data principles (findable, accessible, interoperable, reusable, see wilkinson et al. (2016)) alongside tdr standards as important for collaboration in the evolving european open science cloud (eosc) means that the role of coretrustseal continues to evolve and advance. https://doi.org/10.29173/iq936 6/17 l'hours, hervé; kleemola, mari; de leeuw, lisa (2019) coretrustseal: from academic collaboration to sustainable services, iassist quarterly 43 (1), pp. 1-17. doi: https://doi.org/10.29173/iq936 a stepped framework of rigour and trustworthiness to achieve greater harmonization between the different trust initiatives and their audit criteria, a memorandum of understanding (mou) was signed in 2010 establishing a european framework for audit and certification15 between the data seal of approval, audit and certification of trustworthy digital repositories standard (iso16363) and nestor seal criteria for trustworthy digital archives (din31644). the mou acknowledged a hierarchy of basic (now referred to as core), extended and formal certification in order of complexity and audit rigour (schumann 2012): 1. basic certification is granted by obtaining the dsa. 2. extended certification requires completing the dsa and an externally reviewed self-audit based either on iso 16363 or din 31644. 3. formal certification requires completing the dsa and a full external certification based either on iso 16363 or din 31644. the memorandum clarified the position of the dsa, and by extension, the coretrustseal, within the wider context of tdr standards. the mission of the coretrustseal remains to offer a basic/core certification, which is low barrier to enter, community driven, and integrated into a stepped framework of improvements for those organisations seeking a more rigorous path. standards integration through the research data alliance when a research data alliance (rda) interest group on the certification of digital repositories16 was created, there was recognition of the value of core certification but also a concern that a proliferation of such efforts may be unhelpful. both the data seal of approval and the membership criteria of the international council for science’s world data system (icsu-wds) were identified as core efforts with aligned goals. both held a multidisciplinary remit, though for historical reasons their primary applicants differed with icsu-wds coming from the earth and space sciences, so the partnership also offered an opportunity to serve a wider community. in 2013, the repository audit and certification working group17 was proposed, with a vision of realizing efficiencies, simplifying assessment options, stimulating more certifications, and increasing impact on the community (rickards et al, 2016b). the central focus was a dsa-wds partnership with representatives of both communities involved. in keeping with the transparency principles of the rda, the interest and working group efforts were open to the full rda membership for both participation and communication. the working group undertook an analysis and comparison of the two sets of procedures and criteria with a view to creating a single set supporting the goals of both sources. the process and governance structures for core certification were aligned. the coverage, content and wording of the requirements were reviewed and revised and the resultant common requirements and procedures were run through a test bed process involving current wds members and dsa recipients. these results were published, and their findings integrated into a second revision of the working group outputs. process and requirements, including an introduction on the benefits of certification, background information, guidance text, and a glossary were made publicly available on the rda website for comment during the lifetime of the working group. https://doi.org/10.29173/iq936 7/17 l'hours, hervé; kleemola, mari; de leeuw, lisa (2019) coretrustseal: from academic collaboration to sustainable services, iassist quarterly 43 (1), pp. 1-17. doi: https://doi.org/10.29173/iq936 with an agreed set of procedural and criteria references, the wds and dsa began negotiations to merge the data seal of approval and the wds membership process into a single independent entity. an alignment plan was developed to define the initial governance entities and relationships of both parties. a joint position on the future of core certification was agreed upon, including short term alignment with cooperative, parallel activities, and longer-term alignment towards a single entity. at this point the wds board and dsa board18 began applying the common requirements to new applications and renewals. figure 1: dsa timeline. coretrustseal current activities community-based non-profit organization key factors in the alignment plan were the formation of an interim board19, the development of a common branding plan, board statutes and business plan, as well as the creation of additional guidance to support reviewer training including an online tool to support the application and administration processes. on the 11th of september 2017 the new coretrustseal organisation was announced as a “communitybased non-profit organization promoting sustainable and trustworthy data infrastructures, [] governed by a standards and certification board consisting of members drawn from the assembly of reviewers (by election) and the wider repositories stakeholders (appointed)”. from 2018, the coretrustseal is a legal foundation entity under dutch law governed by a standards and certification board composed of 12 elected members representing the assembly of reviewers. the first board elections were held in july 2018 and the board 2018-2021 consists of: https://doi.org/10.29173/iq936 8/17 l'hours, hervé; kleemola, mari; de leeuw, lisa (2019) coretrustseal: from academic collaboration to sustainable services, iassist quarterly 43 (1), pp. 1-17. doi: https://doi.org/10.29173/iq936 directors ● chair—jonas recker (gesis-leibniz institute for the social sciences, germany) ● vice-chair—hervé l’hours (uk data archive, united kingdom) ● secretary—mari kleemola (finnish social science data archive, finland) ● treasurer—ingrid dillo (data archiving and networked services, the netherlands) members ● jonathan crabtree (odum institute data archive, usa) ● robert r. downs (ciesin-sedac, university of columbia, usa) ● john faundeen (usgs eros centre, usa) ● wim hugo (south african environmental observation network, south africa) ● reyna jenkins (ocean networks canada) ● dawei lin (immport repository, dait-niaid-nih, usa) ● mustapha mokrane (world data system, france) ● paul trilsbeek (max planck institute for psycholinguistics, the netherlands) ex officio ● rorie edmunds (world data system) ● ilona von stein (data archiving and networked systems) the range of existing wds, dsa and dsa-wds certifications have been integrated into the application and certification process of the coretrustseal as part of the certification renewal process. within the ever growing coretrustseal community new volunteers are being sought on an ongoing basis for the pool of reviewers. an introduction to the coretrustseal, the 16 requirements, extended guidance and a supporting glossary are all available20. coretrustseal application process in brief organisations with data expertise for a defined collection may seek certification against the 16 coretrustseal requirements. the first step is to create an account in the coretrustseal application management tool. the actual application process begins when the applicant submits their selfassessment, i.e. their statements for each requirement via the tool. each self-assessment statement against each requirement needs to be supported by evidence. the coretrustseal board then assigns two independent peer reviewers taken from the community of coretrustseal holders. by undertaking this responsibility peer reviewers become eligible for election to the coretrustseal board. members of the coretrustseal community are asked to volunteer to join the reviewer pool. the comments and feedback from the two peer reviewers are assessed by the board and a coretrustseal is either granted for a period of three years, or the application is returned to the applicant for further work. the self-assessments and reviewers’ final comments are published online once the coretrustseal is awarded. each applicant pays a fee of 1000 euros to cover the cost of the operation, maintenance and development of the certification service. certification fee one of the drivers behind the effort to create a single core certification was an acknowledgement of the human and financial investment required to deliver such services. https://doi.org/10.29173/iq936 9/17 l'hours, hervé; kleemola, mari; de leeuw, lisa (2019) coretrustseal: from academic collaboration to sustainable services, iassist quarterly 43 (1), pp. 1-17. doi: https://doi.org/10.29173/iq936 the provision of a fee for coretrustseal certification is a key element of the business model selected to ensure sustainability of the requirements, procedure and the service. the board understood that any cost implication would present an issue for some members of the community. in order to ensure longevity and confidence in the criteria, however, an approach was required which went beyond a reliance on periodic project funding and in-kind contributions from participating organisations. the board has been careful to communicate21 that the fee is for administrative purposes and is in line with the not-for-profit foundation status of coretrustseal. since the certification is valid for three years, the cost averages approximately 300 euros per year of certification. the fee-based model ensures that the cost of operation can be met for the maintenance of the standard, supporting procedures, associated tools, ongoing training of reviewers and engagement with the community. the cost of core certification activities remains entirely met by the volunteer efforts of those joining the pool of reviewers including those elected and appointed to the board. such activities include the dual peer review of self-assessment statements and associated supporting evidence, followed by review, feedback and decisions from the board. overview of the coretrustseal requirements coretrustseal take a ‘whole organisation’ perspective to reviewing data repositories. it starts off by asking for contextualizing background information and then focuses on the organisational infrastructure (mission, licences, continuity of access, sustainability, confidentiality/ethics, skills and guidance), digital object management (integrity, authenticity, appraisal, storage, preservation, quality, workflows, discovery, identifiers, re-use) and technology (technical infrastructure and security). the 16 coretrustseal requirements22: 1. the repository has an explicit mission to provide access to, and preserve data, in its domain. 2. the repository maintains all applicable licenses covering data access and use and monitors compliance. 3. the repository has a continuity plan to ensure ongoing access to and preservation of its holdings. 4. the repository ensures, to the extent possible, that data are created, curated, accessed, and used in compliance with disciplinary and ethical norms. 5. the repository has adequate funding and sufficient numbers of qualified staff managed through a clear system of governance to effectively carry out the mission. 6. the repository adopts mechanism(s) to secure ongoing expert guidance and feedback (either inhouse, or external, including scientific guidance, if relevant). 7. the repository guarantees the integrity and authenticity of the data. 8. the repository accepts data and metadata based on defined criteria to ensure relevance and understandability for data users. 9. the repository applies documented processes and procedures in managing archival storage of the data. 10. the repository assumes responsibility for long-term preservation and manages this function in a planned and documented way. https://doi.org/10.29173/iq936 10/17 l'hours, hervé; kleemola, mari; de leeuw, lisa (2019) coretrustseal: from academic collaboration to sustainable services, iassist quarterly 43 (1), pp. 1-17. doi: https://doi.org/10.29173/iq936 11. the repository has appropriate expertise to address technical data and metadata quality and ensures that sufficient information is available for end users to make quality-related evaluations. 12. archiving takes place according to defined workflows from ingest to dissemination. 13. the repository enables users to discover the data and refer to them in a persistent way through proper citation. 14. the repository enables reuse of the data over time, ensuring that appropriate metadata are available to support the understanding and use of the data. 15. the repository functions on well-supported operating systems and other core infrastructural software and is using hardware and software technologies appropriate to the services it provides to its designated community. 16. the technical infrastructure of the repository provides for protection of the facility and its data, products, services, and users. guidance for self-assessment and review as the coretrustseal has evolved there has been an increased need to provide more elaborate training and guidance for new reviewers to ensure consistency across the applications. in 2017, supported by a small project financed by the research data alliance (rda), extended guidance was created by the coretrustseal board. the extended guidance gives a more detailed perspective on what reviewers should expect as evidence against each of the requirements. as the reviewer pool grows and more repositories are certified, the reviewers’ experiences and observations are captured by the coretrustseal board to ensure the extended guidance remains in line with the latest developments. as a publicly available document, the extended guidance is also a useful tool for applicants23 benefits of certification over time coretrustseal and other certification efforts have sought to demonstrate the benefits of certification to a variety of stakeholders. for tdr these lie not only in the certification itself but in the process by which different repository actors communicate, share and document their knowledge as they prepare and manage the relevant evidence. in a review of their own data seal of approval process, the finnish social science data archive24 noted: the use of models and metrics to assess our procedures and policies have raised our awareness about the challenges of digital preservation, revealed existing and possible problems and weaknesses as well as strengths, steered and initiated minor and major changes in our operations, and resulted in improved documentation. as a consequence, many of our processes are now better and more efficient or, they will be better – some of the bigger changes will take time to implement. we are also able to better manage risks, provide more trustworthy services for the research community, and demonstrate fsd’s trustworthiness to our stakeholders. (kleemola 2015.) https://doi.org/10.29173/iq936 11/17 l'hours, hervé; kleemola, mari; de leeuw, lisa (2019) coretrustseal: from academic collaboration to sustainable services, iassist quarterly 43 (1), pp. 1-17. doi: https://doi.org/10.29173/iq936 during the years data seal of approval and world data system have been conducting their certifications, user experiences have been collected via case studies, presentations at conferences and/ or engagement with the community. furthermore the dutch network digital heritage (nde) conducted a survey on the benefits of the data seal of approval (waterman and sierman, 2016). the following conclusions/benefits can be drawn from their experiences and remain valid for the coretrustseal. ● performing a self-assessment does not take much time; on average, two to four days. it mainly depends on the level of existing documentation and its disclosure. ● although most documentation is intended to be publicly accessible, an exception can be made for documentation containing privacy-sensitive and confidential information, such as a long-term vision. ● the certification process is very useful as an evaluation of internal procedures, which can be reviewed and updated where necessary. the current state of affairs, which can also serve for future accreditation, is made visible. additionally, the procedures and documentation are evaluated, tested and approved by an external professional and the coretrustseal is very helpful in determining strengths and weaknesses. ● the coretrustseal reaffirms the necessity and usefulness of succession/long-term planning and helps to get these issues higher on the agenda of management. ● the coretrustseal contributes to a reliable image. it can be used to improve reputation, but also as a benchmark for comparison. it clarifies what constitutes a digital repository and its business, and it creates transparency for the community in the area of sustainability. ● the coretrustseal increases the confidence of users: it shows that standards are being used, just like the ones being used by traditional museums or repositories. ● the coretrustseal helps to build a community: 'we' all work according to the same standards. ● the coretrustseal emphasizes the need to conform towards the oais standards. ● interaction with the peer reviewer is perceived as significant. ● the requirements are sufficiently generic to be applied to scientific data as well as publications. ● because of its general approach the coretrustseal is perceived as a less 'threatening', detailed and time-consuming procedure than more comprehensive standards, such as iso or trac. the focus is on increasing awareness and transparency; coretrustseal takes a community's and peer reviewers point of view rather than a top-down approach. ● the coretrustseal is a solid foundation for applying for din 31644 certification. ● by renewing the coretrustseal, the data repository will show its progress. (waterman and sierman, 2016) a recent study by donaldson et al (2017) validates these claims. their findings demonstrate that dsa certification has allowed the repositories to: ● build stakeholder confidence of their stakeholders in them, ● improve their documentation, ● gain assurance that they are following best practice, ● demonstrate their transparency, https://doi.org/10.29173/iq936 12/17 l'hours, hervé; kleemola, mari; de leeuw, lisa (2019) coretrustseal: from academic collaboration to sustainable services, iassist quarterly 43 (1), pp. 1-17. doi: https://doi.org/10.29173/iq936 ● improve their processes, ● raise awareness about the importance of digital preservation, ● spend less time on audit and certification as compared to audit and certification through other programs, ● improve communication among staff members, and ● join a community of repositories who have demonstrated their commitment to digital preservation and following best practice.” donaldson et al (2017) future plans certification as a service for complex partnerships though neither the oais model nor the extant tdr standards preclude the idea of a tdr being a complex partnership of organisations, or a mixture of in-house and third party resource, neither do they address it in detail. the coretrustseal acknowledges the rapid expansion of data management partnerships by asking for context about the organisation structure and acknowledges the increased provision of third party services by asking about outsourcing. both of these questions speak to the scope of the repository in terms of control of and responsibility for data and acknowledge the possibility of shared responsibility for the coretrustseal requirements. in an ideal world, each outsource partnership would be to a similarly certified trustworthy entity, but not all partners will seek such certification and in some cases appropriate certification may not yet exist. complex partnerships provide a challenge for applying the coretrustseal. while outsourcing may provide cost savings and access to systems, services and expertise at scale not otherwise available to the applicant, it also introduces a more complex range of relationships and dependencies which can increase bureaucracy and risk. the coretrustseal board is actively investigating how best to define the acceptable scope of outsourcing including which requirements it might apply to and the level of control (and supporting evidence) applicants must provide to support assurances of trustworthiness. the provision of clear service level agreements is one key element of trust in complex partnership and outsourcing models but the coretrustseal must keep up with the evolving nature and technical realities of modern repositories. widening the evidence and certification community the heritage of tdr standards leads to an inevitable focus on oais-defined repositories undertaking active data preservation by data/disciplinary experts for designated community having a defined knowledge base. not all data assets are held in such repositories, however, and the coretrustseal is increasingly receiving queries about how more general purpose repositories or institutional repositories with a broad disciplinary remit might be better supported. many data assets may be stored and managed in such environments for some part of their lifecycle, thus forming a portion of the data provenance critical to ensuring an unbroken chain of trust from data creation/collection to use. a general purpose institutional repository with appropriate disciplinary expertise to define and support preservation for a designated community could apply for the coretrustseal but would need https://doi.org/10.29173/iq936 13/17 l'hours, hervé; kleemola, mari; de leeuw, lisa (2019) coretrustseal: from academic collaboration to sustainable services, iassist quarterly 43 (1), pp. 1-17. doi: https://doi.org/10.29173/iq936 to provide evidence related to a particular part of the data collection. the notion of providing clearer collection profiles to support better certification is already on the coretrustseal radar as repositories undertake a range of curation levels (from storage of untouched deposited data to active participation in data quality improvement) and a range of responsibilities for digital objects (from harvesting metadata for resource discovery to active training and management of access to sensitive data through secure remote environments or safe rooms). the challenge is to set clear, common criteria for defining the data collections while retaining the ‘core’, low barrier to entry mission of coretrustseal. from a full data lifecycle perspective, there is a strong relationship between different repository types and the increased tendency for complex partnerships and outsourcing. both present challenges for our trust in data across a range of data stewards over time. repository host institutions, providers of metadata entry systems, data storage, and discovery and access systems may all contribute (through paid and unpaid relationships) to the overall infrastructure of people, processes and technologies necessary to ensure valuable digital assets are maintained. one potential solution is for coretrustseal to engage with a wider variety of actors, including product and service vendors, to identify how they could support their clients with standardised evidence to support one or more aspects of the coretrustseal requirements. this would lower the barrier to entry of certification; support standardised, transparent evidence of practice; and provide an additional ‘ready for tdr’ incentive to potential partners of participating product and service providers. conclusions the coretrustseal has grown from two complementary approaches to a single set of guidelines ensuring that data repositories can be trusted as stewards for the long term. it has grown and adapted to changing circumstances and continues to do so. the notion of a ‘trustworthy digital repository’ stems from the need to move beyond de facto trust in partner organisations to act as responsible stewards of data, towards de jure assertions of their trustworthiness. standardisation, audit and certification are partly the kind of natural progression towards professionalization experienced by all mature service models. certification also provides clear labelling of trustworthy ‘nodes’ in the (research) data lifecycle where outsourcing, third party relationships, and complex partnerships make the overall technical and human infrastructure of repository services more opaque. despite a membership and history which is predominantly academic, the coretrustseal aspires to a generalised assurance of data preservation across disciplinary and specialist boundaries to ensure that digital data remain accessible to, and understandable by, those interested in seeking to use them for analysis, policy or profit. the notion of trust is critical across the data lifecycle. the ‘repository’ or ‘archive’ model has historically defined itself as a distinct part of the lifecycle, but increasingly some repository standards and best practice have been adopted into more general research data management guidance. https://doi.org/10.29173/iq936 14/17 l'hours, hervé; kleemola, mari; de leeuw, lisa (2019) coretrustseal: from academic collaboration to sustainable services, iassist quarterly 43 (1), pp. 1-17. doi: https://doi.org/10.29173/iq936 organisations which consider themselves as repositories are increasingly ‘full lifecycle’ actors as they are engaged with data producers pre-deposit and with researchers during the data use phase. the deluge of data is a defining change to society and the coretrustseal acknowledges these changes with a broad remit for certification of trustworthy digital repositories. coretrustseal and the coretrustseal community are growing and thriving. the dsa-wds collaboration and aligning of the two certification procedures has proved successful and the coretrustseal has become an independent certification organisation that supports a variety of repositories. today (2018), over 130 seals have been awarded and more are in process. certification standards like the coretrustseal are also playing their part in the european open science cloud (eosc)25 and collaborative research infrastructure developments. in addition, coretrustseal certification is instrumental in helping data repositories adhere to the fair principles26 and coretrustseal is currently working through the eu ict standardisation27 process which will permit it to be referenced for procurement purposes. the formalisation of the coretrustseal requirements, processes and governing bodies must sit alongside a flexible and responsive community-driven approach to change, if it is to continue to adapt to the rapidly evolving needs of the research data community and their collaborative partners. sources [all links accessed 18 september 2018.] donaldson, dillo, downs and ramdeen (2017). the perceived value of acquiring data seals of approval. http://dx.doi.org/10.2218/ijdc.v12i1.481 dsa (2010). quality guidelines for digital research data (“dsa booklet”). https://assessment.datasealofapproval.org/sitemedia/files/dsa_booklets/dsa-booklet_2010.pdf giaretta, david (2011). advanced digital preservation. heidelberg: springer-verlag berlin. harmsen, henk (2008). data seal of approval assessment and review of the quality of operations for research data repositories. international conference on preservation of digital objects (ipres 2008), 29-30 september 2008, london. http://www.bl.uk/ipres2008/presentations_day2/34_harmsen.pdf harrower, natalie & cassidy, kathryn (2017). why storage is not preservation: a conversation, surrounded by conservation. dri blog. http://dri.ie/why-storage-not-preservation-conversationsurrounded-conservation hedstrom, margaret (1998). digital preservation: a time bomb for digital libraries. computers and the humanities 31(3): 189-202. doi 10.1023/a:1000676723815 https://doi.org/10.29173/iq936 http://dx.doi.org/10.2218/ijdc.v12i1.481 https://assessment.datasealofapproval.org/sitemedia/files/dsa_booklets/dsa-booklet_2010.pdf http://www.bl.uk/ipres2008/presentations_day2/34_harmsen.pdf http://dri.ie/why-storage-not-preservation-conversation-surrounded-conservation http://dri.ie/why-storage-not-preservation-conversation-surrounded-conservation 15/17 l'hours, hervé; kleemola, mari; de leeuw, lisa (2019) coretrustseal: from academic collaboration to sustainable services, iassist quarterly 43 (1), pp. 1-17. doi: https://doi.org/10.29173/iq936 kleemola, mari (2015). improving the quality of digital preservation using metrics. iassist quarterly 2015. www.iassistdata.org/sites/default/files/iqvol_39_2_kleemola.pdf oais (2012). reference model for an open archival information system (oais). the consultative committee for space data systems (ccsds). https://public.ccsds.org/pubs/650x0m2.pdf rickards et al. (2016a) developments in the certification of data centres, services and repositories through an rda/wds/dsa partnership http://www.vliz.be/imisdocs/publications/296599.pdf#page=15 rickards, lesley; mary vardigan; ingrid dillo; françoise genova; hervé l'hours; jean-bernard minster; rorie edmunds; mustapha mokrane (2016b). dsa–wds partnership: streamlining the landscape of data repository certification. scidatacon 2016, denver, co., 2016, 11–13 september 2016 (session auditing of trustworthy data repositories). https://doi.org/10.5281/zenodo.252417 schumann, natascha (2012). tried and trusted. experiences with certification processes at the gesis data archive. iassist quarterly fall winter 2012, 23-27. http://www.iassistdata.org/sites/default/files/iqvol36_34_schumann.pdf waterman, kees, & sierman, barbara. (2016). survey on dsa-certified digital repositories. report on the findings in a survey of all dsa-certified digital repositories on investments in and benefits of acquiring the data seal of approval (dsa). https://doi.org/10.5281/zenodo.1188256 wilkinson, mark et al. (2016). the fair guiding principles for scientific data management and stewardship. scientific data 3, article number 160018. https://doi.org/10.1038/sdata.2016.18 1 hervé l'hours, uk data archive, uk data service, university of essex, united kingdom; mari kleemola, finnish social science data archive, university of tampere, finland; lisa de leeuw, data archiving and networked services, netherlands. 2 coretrustseal website: https://www.coretrustseal.org/ 3 reference model for an open archival information system. https://public.ccsds.org/pubs/650x0m2.pdf 4 iso 14721:2012 (ccsds 650.0-p-1.1). https://www.iso.org/standard/57284.html 5 trustworthy repositories audit & certification: criteria and checklist. http://www.dcc.ac.uk/resources/repository-audit-and-assessment/trustworthy-repositories 6 iso 16363:2012 (ccsds 652.0-r-1). https://www.iso.org/standard/56510.html https://doi.org/10.29173/iq936 http://www.iassistdata.org/sites/default/files/iqvol_39_2_kleemola.pdf https://public.ccsds.org/pubs/650x0m2.pdf http://www.vliz.be/imisdocs/publications/296599.pdf#page=15 https://doi.org/10.5281/zenodo.252417 http://www.iassistdata.org/sites/default/files/iqvol36_34_schumann.pdf https://doi.org/10.5281/zenodo.1188256 https://doi.org/10.1038/sdata.2016.18 https://www.coretrustseal.org/ https://public.ccsds.org/pubs/650x0m2.pdf https://www.iso.org/standard/57284.html http://www.dcc.ac.uk/resources/repository-audit-and-assessment/trustworthy-repositories https://www.iso.org/standard/56510.html 16/17 l'hours, hervé; kleemola, mari; de leeuw, lisa (2019) coretrustseal: from academic collaboration to sustainable services, iassist quarterly 43 (1), pp. 1-17. doi: https://doi.org/10.29173/iq936 7 zertifizierung, kriterienkatalog vertrauenswürdige digitale langzeitarchive. http://dx.doi.org/10.18452/1523 8 drambora. http://www.dcc.ac.uk/resources/repository-audit-and-assessment/drambora 9 wittenburg, p., broeder, d., klein, w., levinson, s. c., & romary, l. (2006). foundations of modern language resource archives. in proceedings of the 5th international conference on language resources and evaluation (lrec 2006) (pp. 625-628). http://www.mpi.nl/publications/escidoc-58934 10 stewardship of digital research data: a framework of principles and guidelines (2008). http://www.rin.ac.uk/system/files/attachments/stewardship-data-guidelines.pdf 11 https://ec.europa.eu/research/infrastructures/index_en.cfm?pg=eric-landscape 12 consortium of european social science data archives cessda. https://www.cessda.eu/ 13 european research infrastructure for language resources and technology clarin. https://www.clarin.eu/ 14 european research infrastructure consortium for the arts and humanities dariah. https://www.dariah.eu/ 15 mou to support working on standards for trusted digital repositories. http://www.trusteddigitalrepository.eu/memorandum%20of%20understanding.html 16 rda/wds certification of digital repositories ig. https://www.rdalliance.org/groups/rdawds-certification-digital-repositories-ig.html 17 repository audit and certification dsa–wds partnership wg. https://rdalliance.org/groups/repository-audit-and-certification-dsa%e2%80%93wds-partnershipwg.html 18 the members of these boards were: ciesin-sedac (alex de sherbinin), cines (marion massol), dans (ingrid dillo, lisa de leeuw), finnish social science data archive (mari kleemola), gesis-leibniz institute for the social sciences (natascha schumann), institute of remote sensing and digital earth, china (guoqing li), international service of geomagnetic indices (aude chambodut), mpi (paul trilsbeek), south african environmental observation network (wim hugo), strasbourg astronomical data center (françoise genova), ukda (hervé l'hours), university college dublin (john howard), wdc geomagnetism, university of kyoto (toshihiko iyemori), wds-ipo (mustapha mokrane, rorie edmunds), world glacier monitoring service, university of zurich (isabelle gartner-roer). 19 the interim coretrustseal board consisted of the following members: ciesin-sedac, university of columbia (robert r. downs), dans (ingrid dillo and ilona von stein), finnish data archive (mari kleemola), gesis-leibniz institute for the social sciences (jonas recker), mpi (paul trilsbeek), ocean networks canada (reyna jenkyns), south african environmental observation network (wim hugo), ukda (hervé l'hours), university college dublin (john howard), usgs eros centre (john faundeen), wdc geomagnetism, university of kyoto (toshihiko iyemori), wds-ipo (mustapha mokrane, rorie edmunds). https://doi.org/10.29173/iq936 http://dx.doi.org/10.18452/1523 http://www.dcc.ac.uk/resources/repository-audit-and-assessment/drambora http://www.mpi.nl/publications/escidoc-58934 http://www.rin.ac.uk/system/files/attachments/stewardship-data-guidelines.pdf https://www.cessda.eu/ https://www.clarin.eu/ https://www.dariah.eu/ http://www.trusteddigitalrepository.eu/memorandum%20of%20understanding.html https://www.rd-alliance.org/groups/rdawds-certification-digital-repositories-ig.html https://www.rd-alliance.org/groups/rdawds-certification-digital-repositories-ig.html https://rd-alliance.org/groups/repository-audit-and-certification-dsa%e2%80%93wds-partnership-wg.html https://rd-alliance.org/groups/repository-audit-and-certification-dsa%e2%80%93wds-partnership-wg.html https://rd-alliance.org/groups/repository-audit-and-certification-dsa%e2%80%93wds-partnership-wg.html 17/17 l'hours, hervé; kleemola, mari; de leeuw, lisa (2019) coretrustseal: from academic collaboration to sustainable services, iassist quarterly 43 (1), pp. 1-17. doi: https://doi.org/10.29173/iq936 20 coretrustseal data repositories requirements. https://www.coretrustseal.org/whycertification/requirements/ 21 coretrustseal administrative fee. https://www.coretrustseal.org/apply/administrativefee/ 22 https://www.coretrustseal.org/why-certification/requirements/ 23 coretrustseal extended guidance. https://www.coretrustseal.org/wpcontent/uploads/2017/01/20171026-cts-extended-guidance-v1.0.pdf 24 fsd: http://www.fsd.uta.fi/en/ 25 the european open science cloud. https://eoscpilot.eu/eosc 26 force11. the fair principles. https://www.force11.org/group/fairgroup/fairprinciples 27 european comisssion. ict standardisation. https://ec.europa.eu/growth/industry/policy/ict-standardisation_en https://doi.org/10.29173/iq936 https://www.coretrustseal.org/why-certification/requirements/ https://www.coretrustseal.org/why-certification/requirements/ https://www.coretrustseal.org/apply/administrative-fee/ https://www.coretrustseal.org/apply/administrative-fee/ https://www.coretrustseal.org/why-certification/requirements/ https://www.coretrustseal.org/wp-content/uploads/2017/01/20171026-cts-extended-guidance-v1.0.pdf https://www.coretrustseal.org/wp-content/uploads/2017/01/20171026-cts-extended-guidance-v1.0.pdf http://www.fsd.uta.fi/en/ https://eoscpilot.eu/eosc https://www.force11.org/group/fairgroup/fairprinciples https://ec.europa.eu/growth/industry/policy/ict-standardisation_en vol223 12 iassist quarterly midas is a jisc designated national data centre for the uk higher education community providing on-line access and support for a range of large and complex datasets, such as censuses, surveys, time series databanks, bibliographic and full text databases. in this context, midas is part of the developing jisc funded national distributed electronic resource which is seeking to promote and extend access to electronic information and services to the entire uk higher education community. the expectations of users have changed considerably and we have had to rethink how we deliver data and information to the researcher’s desktop. it is not enough to promote awareness of the data resources and their potential applications in teaching and research. we also have to convince the users that their time is being used efficiently, that they can easily identify the data that they want, extract it and where appropriate put it into a suitable format for secondary analysis. for us this means creating appropriate interfaces for the data simple enough for a once-off selection and versatile enough for more sophisticated use. this paper addresses the influence of the web and the expectations of its users on the services provided by midas. it shall describe some the new interfaces to data and information which will be of particular interest in both research and teaching. midas overview midas (http://www.mimas.ac.uk/) is a jisc (http:// www.jisc.ac.uk ) designated national data centre for the uk higher education community providing on-line access and support for a range of data and information resources. along with the other jisc funded data centres (bids and edina), data services and projects, midas is part of the developing distributed national electronic resource (dner) which is seeking to promote and extend access to electronic information and services to the entire uk higher education community (http://www.jisc.ac.uk/cei/ dner_colpol.html). midas is based at manchester computing at the university of manchester. manchester computing also hosts other projects, such as copac (http://copac.ac.uk/copac/) and the superjournal project (http://www.superjournal.ac.uk/sj/). the data and information resources available through midas include on-line access to the uk census of population; government and other continuous surveys; national and international time series databanks; digital map data; satellite images; chemical information systems; bibliographic data resources and electronic journals. a key component of the service is the provision of a range of specialist support services, such as documentation, training and research support, relating to the data and information resources available on midas. many of the services are run in collaboration with other organisations, such as the data archive at the university of essex. midas also provides access to a range of software packages and the large scale computing resources (i.e. memory, disk space and cpu time) required by users wishing to undertake complex data analysis. midas also provides facilities and support to projects wishing to exploit the internet to provide wider network access to data and/or information resources for teaching and research purposes. two key data sharing and gateway services that have developed considerably over the past couple of years include netec (http://netec.mcc.ac.uk), which is a collection of projects which aim to improve the usefulness of electronic networks in economics, and genuki (http://www.genuki.org.uk), which provides information relating to the study of genealogy and family history. the service continues to grow both in terms of the range of services offered and numbers of registered and active users. the history of midas and the web historically, manchester computing has always made efforts to ensure that service specific information is available on-line. prior to the installation of a unix platform for the midas service in 1993, the on-line help systems were developed using proprietary tools and access was restricted to users logged onto the system. with the advent of a unix based service, the on-line information system moved to gopher. this client-server approach to providing access to information about the various datasets, software and other services available via midas proved the role of the web in the provision of national data and information services: the midas experience by julia chruszcz, keith cole and anne mccombe * http://www.mimas.ac.uk/ http://www.jisc.ac.uk http://www.jisc.ac.uk http://www.jisc.ac.uk/cei/dner_colpol.html http://www.jisc.ac.uk/cei/dner_colpol.html http://copac.ac.uk/copac/ http://www.superjournal.ac.uk/sj/ http://netec.mcc.ac.uk http://www.genuki.org.uk winter1998 13 extremely flexible. it could be used by users actually logged onto midas as well as users with local access to a gopher client. until relatively recently, the gopher server was used as the primary method of providing access to information. with the advent of the web, the initial function of the midas home page was to simply provide an alternative interface to the information held in the gopher. however, the midas web site (http://www.mimas.ac.uk/) soon developed into a much more sophisticated information resource which was more suited to meeting the diverse information needs of data set users. service specific web pages were established; documents, reports and newsletters were made available in an html format; experimental web interfaces to datasets, and information were also developed together with the provision of links to related web based resources. compared to the graphical user interface offered by the web, the hierarchical text based gopher system soon started to look antiquated. as use of the web and access to web browsers became more widespread, it seemed appropriate to transfer the midas information server from gopher to web. increasingly, the service specific midas web pages are starting to become information gateways in their own right. a good example are the web pages relating to the statistical packages on midas (http://www.mimas.ac.uk/ stats/). the decision to transfer the information server from gopher to web was not greeted with universal acclaim. initial user feedback indicated that whilst the majority of midas users had easy access to web browsers there was still a sizeable minority of users with only telnet access and/or without access to a web browser who would be inconvenienced if the gopher service were withdrawn. consequently, it was necessary to install lynx – the non-graphical web browser on the midas unix server to provide an alternative method of accessing information held as web pages on midas. responding to changing user expectations there are currently a variety of different interfaces to the various data and information resources available on midas. in part, this reflects the specialist software requirements of many of the datasets. for example, the majority of the government and other continuous surveys are supplied to midas in sir database format. indeed, some surveys, such as the british household panel study (bhps), are actually supplied in multiple formats (e.g. sir, spss, sas and stata). although some of the services are entirely web based (e.g. jstor, copac and ons databank) or use specific client-server software packages (eg beilstein crossfire). the majority still require the user to physically log onto a remote unix server – either via telnet or x-windows and run one or more application packages. there are a number of problems with this mode of working. the user is required to have a basic competence with the unix operating system, a reasonable prior knowledge of the structure of the data as well as expertise in one or more specialist data extraction and analysis packages. whilst this might not pose a problem for computer literate and network aware researchers, it is does represent a significant barrier to less experienced users, such as undergraduate students, across a broad range of disciplines who might have extremely limited data requirements. the expectations of users have changed dramatically over the last couple of years and we have had to rethink how we deliver data and information to the researcher’s desktop. it is not enough to promote awareness of the data resources and their potential applications in teaching and research. we also have to convince the users that their time is being used efficiently, that they can easily identify the data that they want, extract it and put it into a suitable format for secondary analysis and that they can obtain informed advice on types of use. for us this means creating appropriate interfaces for the data – simple enough for a once-off selection and versatile enough for more sophisticated use. over the last couple of years midas has endeavoured to respond to changing user expectations by trying to develop more web based interfaces to many of the data and information resources. a number of the web based interfaces to midas services have been developed in house using cgi programming techniques. for example, a web based interface to the ons time series databank for the uk has been developed (http://www.mimas.ac.uk/ons/) in order to facilitate greater use of the data in teaching. this simple interface permits searching by keyword, browsing of series and saving of extracted series in different formats. it is important to note that many of these web interfaces represent additional interfaces rather than replacements for existing interfaces requiring a unix login. the aim is to provide more appropriate interfaces to new categories of users rather force existing users into new modes of working. whilst it is relatively easy to develop simple search and retrieval interfaces to data and information resources – providing it does not require too much restructuring of the underlying datasets it is much harder to develop more sophisticated interfaces. midas is currently exploring the use of java tools to develop a web enabled version of the cartographic data visualiser (cdv) which uses scientific visualisation techniques for exploratory spatial data analysis (http://www.mimas.ac.uk/janus/cdv/). however, building complex web based interactive data extraction and visualisation systems is currently resource intensive and might not be worthwhile given the advent of commercially supported products. in addition, it is clear from the feedback from users that the ease of use of an http://www.mimas.ac.uk/ http://www.mimas.ac.uk/stats/ http://www.mimas.ac.uk/stats/ http://www.mimas.ac.uk/ons/ http://www.mimas.ac.uk/janus/cdv/ 14 iassist quarterly interface is often more important than its functionality. one of the major brakes on developing more web based interfaces is the current absence of web enabled versions of many of the packages currently used on midas for providing access to data. this is changing gradually but perhaps too slowly for most users. for example, the sas/ intrnet software product (http://www.sas.com/software/ components/intrnet.html) can be used to develop webenabled sas applications running on a central server. similarly, esri map objects internet map server (http:// www.esri.com) provides a set of tools that can be used for serving maps over the web. both of these products will be evaluated for use on midas. is accessibility the only constraint? the web has a major role to play in improving accessibility to electronic data and information resources. whilst accessibility is a major problem that requires addressing, it is not the only constraint on more widespread and effective use of data and information resources. for certain datasets, increased availability and improved interfaces have not necessarily resulted in the expected growth in numbers of users that might have been expected. one factor that plays an important role is the general lack of awareness amongst the wider academic community about the availability, content, scope, structure of key data resources combined with a lack of understanding about their potential application in both teaching and research. although some users may have a good awareness about individual datasets this may not be matched by an equivalent understanding of complementary data resources. in addition, the level and type of use of certain data sets can also be adversely affected by a lack of appropriate secondary analysis skills. for example, an absence of spatial data handling skills frequently acts as a major constraint on the of use of key spatial data resources, such as 1991 census digitised boundaries. the combined problems of accessibility, awareness and usability are being addressed by the kinds (http:// www.mimas.ac.uk/kinds/) project with respect to the spatial data resources available on midas (http:// www.mimas.ac.uk/maps/). an integrated set of web based tools have been developed which enable relatively inexperienced users to search, browse and work directly with large and complex spatial data sets, such as the bartholomew’s digital map data for great britain. these tools include a spatial search engine; data set browsing; a help system which provides access to a knowledge base of geographical terms and concepts together with mapping/ data download facilities. feedback from users indicates that the complex registration process for many of the copyright/commercial data sets held on midas does act as a major deterrent to use – particularly for teaching purposes. although registration is frequently seen as part of the cost of obtaining free or low cost access to commercially valuable datasets there is a need to negotiate unrestricted academic access to more sample datasets to facilitate the development of teaching and learning materials. if the web is to be used to build virtual learning environments it is important that it is populated with meaningful sample datasets. developing sustainable data services use of data and information resources in academic teaching and research will inevitably generate a range of derived materials. these derived materials could include derived data sets; program code/algorithms; methodological notes; data quality reports or teaching materials. it is regrettable that large amounts of this value added material are frequently discarded and not recycled back to other users. whilst funding bodies such as jisc are concerned about the long term preservation of electronic materials that constitute the dner the issue of building dynamic and sustainable databases of derived material has not received much attention to date. midas has worked with a number of projects and users wishing to make derived datasets and/or software tools more widely available – particularly where this adds value to the services hosted by midas. this has happened particularly with respect to the 1991 census area and interaction statistics held on midas. an example is the 1981 and 1991 census population surfaces and associated access software (http://census.ac.uk/cdu/surpop/). the kinds project is also currently exploring developing databases of derived material relating to the spatial data resources available on midas. on the basis of experience to date, it is clear that the web offers great potential for building collaborative virtual user communities around key data and information resources. however, there are a considerable number of technical, legal, organisational, quality assurance and cultural issues that require resolution before such dynamic web enabled databases of derived materials can be developed. exploiting emerging technologies keeping in contact with both existing and potential users of midas services and alerting them to new developments is a major problem. not all users subscribe to the relevant email distribution lists at mailbase (http:// www.mailbase.ac.uk), there are problems with crossposting and managing these lists can also be problematic as email addresses constantly change. similarly, not all users visit the ‘latest news’ section of the midas web site on a regular basis. one possible solution is the use of push channel technology as an automated method of keeping users notified about new developments. a variety of different products are now available for delivering customised news feeds to users. however, use of push channel technology is still relatively underdeveloped in the uk academic community despite its potential. http://www.sas.com/software/components/intrnet.html http://www.sas.com/software/components/intrnet.html http://www.esri.com http://www.esri.com http://www.kinds.ac.uk/kinds/ http://www.kinds.ac.uk/kinds/ http://www.mimas.ac.uk/maps/ http://www.mimas.ac.uk/maps/ http://census.ac.uk/cdu/surpop/ http://www.mailbase.ac.uk http://www.mailbase.ac.uk winter1998 15 another emerging technology is the use of plugin and helper applications which embed additional functionality within web browsers and also enable users to access files in more specialised formats. for example, the adobe acrobat reader plugin enables users to view, navigate and print documents, such as electronic journal articles, held in pdf format. from the midas perspective, one of the key benefits of plugin technology is the extent to which it can significantly reduce interface development times. however, less technically competent users do express concerns about the concept of having to locate, download and install plugins – although in the future these tasks may be performed automatically by the browser. midas has also been experimenting with use of other types of plugin. as part of an esrc funded project, midas has been developing an entirely web based interface to 1991 census area statistics. a major component of this interface is the map based front end to the 1991 census area statistics database. by downloading and installing the autodesk mapguide viewer (http:// www.mapguide.com/) the user is able to embed limited desktop mapping functionality within the web browser. this enables the user to interact dynamically with a multilayered map which incorporates digital map data from the 1:1,000,000 bartholomew europe dataset and digitised 1991 census output area boundaries. by panning and zooming the user is able to identify and select areas for which 1991 census area statistics are to be extracted. accessibility for all? it has become increasingly apparent that whilst new digital telecommunications technologies can significantly widen access to data and information resources on a world wide basis they can also serve to reduce accessibility for other groups. it is in that context that many organisations such as the world wide web consortium (w3c) are looking at ways in which access and usability can be improved for people with disabilities (http://www.w3.org/wai/) or with access to text only browsers. it is clear from the ongoing debate on accessibility and usability issues as part of the design of html 4.0 (http://www.w3.org/wai/references/ html4-access) that it will be increasingly important for data and information service providers, such as midas, to look at the way in which documents are structured and whether there is an over reliance on graphics for information presentation (e.g. frames), navigation aids and/ or resource discovery (e.g. clickable image maps). graphics and multimedia have an important role to play but this should not be at the expense of textual content and/or network efficiency. however, in an era of rapid technological development there will always be a tension between the desire to exploit leading edge technology in interface development and the requirements of low technology users and/or those with special needs. conclusion the web has already made a tremendous impact on the way in which data centres, such as midas, deliver data and information to the desktop. in order to respond to changing user expectations it will be strategically important for midas to continue to develop more web based interfaces to its services – although these might not represent the only access method. however, for certain datasets this may not be achievable in the short term due to the absence of web enabled versions of all the key data access/analysis packages used by the service. therefore, for certain data resources it may be necessary for midas to develop its own web based interfaces – which may require a restructuring of the underlying data formats. irrespective of the access method, the web does offer considerable scope for promoting more widespread and effective use of the unrivalled data and information resources that have been made available uk academic community. in this context, it is vitally important that the jisc funded data centres and data services continue to develop and enhance the knowledge base relating to particular data and information resources and provide access to more didactic materials. this will facilitate more effective resource discovery as well as contributing to the development of virtual learning environments and hopefully start to narrow the gap between actual and potential use of the dner in teaching and research. acknowledgements keith cole and anne mccombe both of whom work for midas have made significant contributions to this paper. keith cole as well as being a manager of the midas services is also director of the esrc/jisc 1991 census dissemination unit and the assistant director (technical) of the cathie marsh centre for census and survey research (ccsr) at the university of manchester. keith has presented a similar paper specifically geared to social scientists at iriss ’98 in the uk. references midas (http://www.mimas.ac.uk/) joint information systems committee (http:// www.jisc.ac.uk ) distributed national electronic resource (dner) (http:// www.jisc.ac.uk/cei/dner_colpol.html). copac (http://irwell.mimas.ac.uk/sj/) superjournal project (http://www.superjournal.ac.uk/sj/). netec (http://netec.mcc.ac.uk) genuki (http://www.genuki.org.uk) http://www.mapguide.com/ http://www.mapguide.com/ http://www.w3.org/wai/ http://www.w3.org/wai/references/html4-access http://www.w3.org/wai/references/html4-access http://www.mimas.ac.uk/ http://www.jisc.ac.uk http://www.jisc.ac.uk http://www.jisc.ac.uk/cei/dner_colpol.html http://www.jisc.ac.uk/cei/dner_colpol.html http://irwell.mimas.ac.uk/sj/ http://www.superjournal.ac.uk/sj/ http://netec.mcc.ac.uk http://www.genuki.org.uk 16 iassist quarterly office for national statistics databank (http:// www.mimas.ac.uk/ons/) sas/intrnet software product (http://www.sas.com/ software/components/intrnet.html) esri map objects internet map server (http:// www.esri.com) kinds (http://www.kinds.ac.uk/kinds/) mimas: the spatial side (http://www.kinds.ac.uk/kinds) surpop v 2.0: introduction (http://census.ac.uk/cdu/surpop) mailbase (http://www.mailbase.ac.uk) autodesk mapguide (http://www.mapguide.com) web accessibility initiative (wai) (http://www.w3.org/ wai) wai resource : html 4.0 accessibility improvements (http://www.w3.org/wai/references/html4-access.html) * paper presented at the 1998 iassist/css conference, yale university, new haven, usa, may 19-22, 1998, julia chruszcz, keith cole and anne mccombe, university of manchester, manchester computing, oxford road, manchester m13 9pl. http://www.mimas.ac.uk/ons/ http://www.mimas.ac.uk/ons/ http://www.sas.com/software/components/intrnet.html http://www.sas.com/software/components/intrnet.html http://www.esri.com http://www.esri.com http://www.kinds.ac.uk/kinds/ http://www.kinds.ac.uk/kinds/ http://census.ac.uk/cdu/surpop/ http://www.mailbase.ac.uk http://www.mapguide.com/ http://www.w3.org/wai/ http://www.w3.org/wai/ 1/2 hayslett, michele (2022) editor’s notes: the work continues, iassist quarterly 46(4), pp. 1-2. doi: https://doi.org/10.29173/iq1076 the work continues welcome to the final issue of the iassist quarterly for the year 2022 – iq volume 46(4), our eagerlyawaited special issue on systemic racism in data practices. this issue represents more than you might think: the culmination of more than two years of the intellectual hard work of writing, of course, but that in itself is not unusual for any journal issue. however. the global pandemic exploded just after the conception of this special issue and hit all of us hard, wreaking not only physical destruction of lives but also unleashing social upheaval, job insecurity, housing insecurity, and major mental health challenges. social injustice erupted during the pandemic, shocking and enraging many of us with its violence and disregard for human dignity. i was privileged to witness the genesis of this issue, and i helped recruit our guest editors, trevor watkins and jonathan cain. i salute their perseverance, patience and courage, and that of the article authors, in bringing this content to fruition. many involved in this issue faced multiple personal challenges, from the loss of family members to repeated moves, job changes, and more in the process of trying to get this work done. some were unable to surmount the many obstacles and were forced to withdraw their proposals. so i do not think it is hyperbole to say this is the hardest issue we have ever produced. trevor and jonathan, thank you again for spearheading this important work. some good things have come from the societal call for racial justice for iassist, including this issue of the iq. iassist has initiated several new ventures to advocate for diversity and equity, both within our organization and among researchers generally: we restructured our membership fees to allow half price for people joining from lower income countries. iassist also sponsored diversity scholarships for members to attend the american library association conference and the icpsr summer program in quantitative methods in 2022. a new anti-racism resources interest group which focuses on compiling anti-racism resources has been working for more than two years and recently collaborated with the professional development committee to present a webinar on varying national approaches to collecting (or not collecting) data about race and ethnicity (see this page for the webinar recording as well as the essays members have written). the group welcomes contributions of essays for additional countries and suggestions of other webinar topics. looking ahead, the 2023 conference theme is diversity in research: social justice from data, sure to result in some fascinating presentations (and future iq papers!). and here at the iq, we’re already contemplating a second special issue in this area around the role of social justice in data services. we invite volunteers who would like to serve as guest editors to contact us. and so the work continues. the iq editorial team is happy to welcome a new volunteer, phillip ndhlovu, as our managing editor with this issue. phillip is the deputy librarian at the gwanda state university library in filabusi, zimbabwe. we thank him profusely—his role is key to producing every issue and his participation enables ofira and me to focus on learning the editor’s role. we welcome suggestions for new features or columns, and encourage you to reach out if you are interested in becoming involved. from all of us on the iq editorial team, we wish you a much better year in 2023. and meanwhile, enjoy the hard work herein of your colleagues. read on for trevor and jonathan’s guest editors’ notes describing the enclosed articles. for the iq editorial team, https://doi.org/10.29173/iq1076 https://iassistdata.org/community/antiracism-resources/ https://iassistdata.org/community/antiracismresources-ig/essays/ 2/2 hayslett, michele (2022) editor’s notes: the work continues, iassist quarterly 46(4), pp. 1-2. doi: https://doi.org/10.29173/iq1076 michele hayslett – december 2022 karsten boye rasmussen ofira schwartz-soicher https://doi.org/10.29173/iq1076 6 iassist quarterly winter/spring 2010 by kristin partlo1 the pedagogical data reference interview abstract this essay reflects on the reference interview on several levels. if we accept that academic reference in general has a pedagogical role, then it is necessary to adjust the standard model of the reference interview to reflect that value. within the specific context of academic data reference, undergraduates as a group require more instruction during the reference interview because they are less prepared than graduate students and faculty to ask for what they need. a strict service model does not meet their needs. a successful model appropriately balances the tension between instruction and service. this balance will vary from one institution to another based on different user groups and institutional goals, with implications for resource allocation. data librarians on the one hand and general and subject reference librarians on the other bring distinct sets of knowledge and experience to bear on the challenge of "assessing the user's need," which can be a rich point of collaboration and referral between them. keywords: reference interview, bibliographic instruction, undergraduates a data reference pedagogy as a social science data librarian at carleton college, a small undergraduate campus with a distributed data support model, i consider myself to be situated somewhere between a general reference librarian and a data specialist, with a foot firmly planted in each professional culture. i have found the reference interview2 to be a rich site for examining and bridging the practices, values and expertise of the two specializations, especially as we have reflected on an appropriate level of research data support for our liberal arts campus. because my work is at a confluence of many different levels of the organization, namely librarians, staff, students and faculty, i have come to realize that my reflections may be extrapolated to larger campuses where more individuals and organizations are involved in supplying data support across the campus. in a 2002 article in the journal portal: libraries and the academy, james elmborg called for a vocabulary and theoretical underpinnings for discussing reference work as a teaching activity. how does one teach well at the desk? since what sets academic librarians apart from the rest of the profession is a recognized role in "participating in the teaching roles of their institutions," it is important to conceive of the reference interview not just as providing a service but also as engaging in teaching. to achieve this end, elmborg proposes that we draw from the theories of cognitive and social constructivism and the language of writing instructors to reposition our practice so that along with meeting the users' needs it is also our goal to help create self-sufficient learners. he argues, “it is the role of the teacher to identify where the student is in his or her development … and then provide guidance and collaboration in ways the student can internalize” (p. 462). further, he expands upon the context of the reference interview, framing it not as an isolated interaction at a service counter, but rather as part of the socialization of new researchers into the community of scholarship: . …if we accept the central notion that knowledge and meaning get negotiated in social contexts among members of a discourse community, then our responsibility within that community is to participate in discourse, to engage our students with meaningful talk about their research, to help them develop a language of inquiry that will allow them to articulate to themselves how to proceed with present and future research challenges (p. 461). however, in the field in general, and in the materials from which reference librarians are taught, the reference interview is based on an assumption of identifying a user's existing need and providing them with relevant information – or perhaps a search strategy – suitable to that need. it does not reflect a goal of teaching, only of providing a service that meets an information need. i find that this dynamic is further complicated when one considers the data reference interview. first, consider the way the reference interview is taught in library schools. the popular reference textbook by bopp & smith outlines the following five elements of the reference interview: 7 iassist quarterly winter/spring 2010 • open the interview expressing openness and approachability establish that you want to help • negotiate the question learn the context of the patron: how much detail at what level is needed use of open and closed questions and active listening • search for information • communicate the information to the user • close the interview express willingness to provide further help refer if necessary3 bopp & smith's articulation of a service ethic and emphasis on understanding the user's question are central to the way reference librarians conceive of their work. however, this model leaves out such important pedagogical elements as encouraging the student to participate in the process, explaining the judgements and decisions made to determine relevant information, determining the learning stage of the student, attempting to create a dynamic, student-centered conversation (elmborg p. 460), and fostering exploration and independence as a researcher in the student. one paragraph in particular highlights the disconnect between a broader reference model as articulated by bopp & smith and the pedagogical model espoused by elmborg. it should be obvious that information given to users must be at an appropriate intellectual level and free of jargon from either the field of librarianship or from any other field with which the user is unfamiliar. users will not always say that the material presented to them is unclear or too difficult for them to master, so librarians must assess the user’s abilities during the search process to avoid giving the user the right information in the wrong package. (bopp & smith, p. 58) this useful advice on meeting the patrons where they are and applying empathetic attention to their context and intellectual level still stops short. the goal here is to meet users where they are, but does not challenge them to become self-sufficient learners and searchers. elmborg pushes us to consider instead that the reference librarian should intentionally help students to understand and apply the language of research in their field, rather than protect them from it. i propose that helping students understand disciplinary terminology is especially called for in the case of working with undergraduates and data. it is not the librarian's job to match the information to students' level of understanding. rather, it is the student's job, with the help of a librarian and their professors, to strive to raise their level of understanding to the terminology and concepts used with data and their documentation. if we extend this critique to the data reference interview, a similar question arises. in the absense of a standard data reference textbook, i have constructed a composite of the advice i have encountered through fora such as iassist presentations and listserv discussions, the iq, individual data librarians and the icpsr class, " providing social science data services: strategies for design and operation," taught by jim jacobs, chuck humphrey, and diane geraci. establish what the patron needs: • statistics or data? • what is the subject or topic? • what is the unit of analysis? • geographic constraints or units? • time constraints (a range of years; monthly, quarterly or annually)? • do they need cross-sectional or longitudinal data? time series? • opinion or demographic data? financial or administrative data? these questions that must be answered prior to finding data are highly helpful in establishing goals and a structure for a data reference interview, but similar to the bopp & smith model, this advice is focused on service. i still question: what would a data reference model be that blended pedagogy and service, imbuing the data reference interview with the goal to teach? undergraduates as users of data one step toward addressing my question is to consider undergraduates as a distinct user group at a particular stage in their learning. we are careful to adapt teaching techniques to suit distinct learning stages in the classroom, so why not the same care at the reference desk? frequently, working with undergraduates can feel alien, a bad fit for the data reference interview. my own experience of applying the model by asking the questions referred to above has been to uncover more questions. students with fluid, emerging research questions can not specify such things as geographic or time constraints. the constraints introduced by the availability of data will have a stronger impact on their research question than vice versa. students new to quantitative research often do not fully understand the concept of unit of analysis, or the difference betwen cross-sectional and longitudinal data or, worse, the 8 iassist quarterly winter/spring 2010 difference between data and statistics. i am sure many of us who have worked with undergraduates have experienced significant breakdowns of communication that catch one off-guard. reference interactions are rife with problems of semantics and definitions (e.g., what exactly is meant by "raw data"), of process and of expectations (fig. 1). it is easy for these experiences to lead to underestimate undergraduates and to let them become caricatures in our minds. we might conclude that undergraduates are too lazy or impatient to pay sufficient attention to detail, that they lack the motivation, that they rarely if ever actually need real data, or that they're overconfident and always leave their work for the last minute. it is necessary, though, if we are going to take seriously the task of helping undergraduates learn, to reconceive of some of these patterns and use them as a way to understand where they are developmentally as researchers. these negative qualities are sometimes strategies that students have developed because they have been successful. working to deadline has provided motivation and adrenaline to boost creativity. abundant confidence gets them through the constant barrage of new ideas they get on a daily basis in classes. everyday web searching rewards clicking links quickly instead of reading the page first. new researchers need to be given convincing cues that these strategies are not inherently wrong, but will no longer work in the context of long-term research projects. since undergraduates are emerging as researchers they do not have experience to guide them in areas more experienced researchers take for granted. they have not yet mastered the process of literature review and data search strategizing. they are still learning how to form researchable questions. the process of working with data is vague. they lack fluency in the language of quantitative research. in fact, when they come to us, they may have never encountered data "in the wild" before, having always been provided data with their assignments, and having no idea of the amount of decision-making, cleaning and arranging that goes into preparing data once it is found and accessed. all of these considerations aside, undergraduates often simply are not doing the same thing as advanced researchers. they have different, but equally legitimate, motivations for looking for data. they may be required to find data on a topic not their own. they are likely to be working on short-term projects in which they are investing limited time and attention. they may be looking for data not as evidence per se but rather in order to demonstrate newly-learned statistics skills (i.e., their criteria are structural rather than topical such as a dataset with at least two continuous variables and a categorical variable). also, format can unevenly influence data selection. students familiar only with one statistical package may only be willing to look for data files formatted for that package. figure 1 9 iassist quarterly winter/spring 2010 combining discovery & instruction i do not have a formal proposal for building a data reference pedagogy, but i can share some of the strategies used at carleton college. these strategies are inspired by the idea of a teaching reference interview and follow the general principle of combining discovery with teaching whenever possible. recognizing that most novices will not typically have a strategy for searching for data beyond google, i try to model good search behavior, emphasizing process and figure 2 continuously narrating my decisions. i use visualizations whenever possible to help speed understanding of complex ideas. i help students take notes by making very concrete suggestions and taking notes with them during the consultation, and suggest approaches for dealing with uncertainty. on a broader, programmatic level, the carleton reference librarians have tried to embed data instruction into information literacy instruction whenever possible. the quantitative reasoning initiative on our campus emphasizes figure 3 10 iassist quarterly winter/spring 2010 teaching qr or numeracy across the curriculum. this cross-curricular emphasis has the added advantage of creating opportunities for all librarians to provide instruction on finding quantiative information. all the librarians integrate quantitative sources into their online research guides so that data are presented as just one type of information among many. students repeatedly receive the message that employing data as evidence in making arguments is critical to argument for all scholars, not just "quant geeks." for our web based finding aids, we have intentionally prioritized integration into course guides over creation of standalone data-specific materials, which could run the risk of becoming a data silo. perhaps most important, the other social science librarian, danya leebaw, and i have developed a data reference worksheet (see appendix) that prompts students and the librarian assiting them through a brainstorming process. the worksheet provides a place to jot down the suggested resources into not just a list of places to look, but within a structure that suggests a method. filling out the form together demonstrates that the librarian doesn't just come up with ideas out of thin air, but rather out of a particular thought process necessitated by the information landscape of data production and publication. below are two examples of ways that i regularly prompt students to actually take notes because otherwise they often tend to click and click and go round in circles and get frustrated. data reference outside the data center although the majority of readers of this article will not share my context of working in a small liberal arts college, i believe there are important reasons for all data specialists to think about how our reference model serves undergraduates. we know that more and more undergraduates are using data. they're being introduced to it in their classes, they're using it in independent research and internships. but are the data specialists prepared for the teaching intensive work of helping undergraduates? further, are undergraduates more likely to take their questions to the general reference desk or to a subject specialist who has knowledge of their assignment and subect area? is there support for data across the campus, outside the data center? what does the data reference interview look like at the general reference desk and how does it differ from and complement the way it works in the data center? for data specialists who are examining data support across the campus and are thinking about ways to train other librarians, it would help to frame that experience as a collaboration. to truly support undergraduates, it is necessary to combine the expertise of the data specialists with the expertise of subject and general reference librarians, namely their familiarity with the practice of teaching undergraduates and their knowledge of what is being taught in assignments. data and non-data librarians can bring their complementary areas of expertise to bear on the puzzles of developing a pedagogical data reference model and helping undergraduates find and access data. in conclusion, i have tried to show that the model of the data reference interview needs to reflect the tension of providing both service and instruction. it is not sufficient for the model to include only the elements of helping a patron to determine their need and then either getting them to the data or pointing them in the right direction. rather, there is an essential third dimension introduced by the commitment to reference as a site of teaching. especially in the case of undergraduates, the data reference interview should take a pedagogical approach, helping new researchers develop the capacity to do it on their own, creating self-sufficient learners and nascent social scientists. acknowledgements this article would not have been possible without the insightful discussion and generous sharing of ideas by colleagues, danya leebaw and heather tompkins. i would also like to thank paula lackie and carolyn sanford who established the groundwork for and continue to provide support for data at carleton and who provided helpful comments on early drafts of this paper. references bopp, richard e. and linda c. smith. 2001. reference and information services: an introduction. 3rd edition. englewood, colorado: libraries unlimited, inc. elmborg, james k. 2002. teaching at the desk: toward a reference pedagogy. portal: libraries and the academy 2 (3): 455-464. green, samuel s. 1876. personal relations between librarians and readers. library journal 1: 74-81. notes 1 kristin partlo, reference and instruction librarian for social science and data, gould library, carleton college. one north college street, northfield, mn 55057. email: kpartlo@carleton.edu. 2 reference librarians have been concerned with defining professionally the practices and goals of reference service since samuel green’s article in the first issue of library journal in 1876. specifically, the notion of the reference interview, or the preoccupation with the interaction and communication between the librarian, the patron, and his or her need for information, emerged and crystallized in the 134 years since. 46 iassisl quarterly providing local data services by r. de vries ' steinmetz archives/swidoc. amsterdam, the netherlands the steimnetz archive, as the dutch national data archive for the social sciences, is a somewhat special case in the context of this workshop. we function both as a data clearinghouse and as a data library per se. as data library, the archive operates in an area that, in other countries, would be seen as a "regional" (as opposed to "national"); the following are some reasons why the steiiraietz can serve as an example of a data library providing local data services. dependency on someone else's computer center. i.e., decisions on installation of software packages, policy on how to handle mass storage problems and the safekeeping of magnetic tapes, participation in networks, etc., are all beyond our direct control; we can argue, but have no real influence on such decisions, let alone the means of implementing them on our own. dependency on the willingness of research 'paper prepared for presentation at the ifdo/iassist conference. workshop on data services, amsterdam, mav 20 23, 1985. organizations and their funding bodies to deposit data (survey or otherwise) in the archive's holdings. again, we have no financial means with which to buy large datasets, nor the manpower, in the case of published statistical data for example, to generate new data from published sources. a small staft (5). a strong emphasis on documentation and reference service, aided by a reference database containing study descriptions of every stored dataset "documentation" here refers to the original questionnaire, research report, print-outs of frequencies, etc given the national role of the archive a less than desirable situation, but in the context of "local data servcies" quite reasonable: we caimot give access to our holdings through a network, nor is the reference database available online. data exchange is via magnetic tape, and available datasets are brought to the attention of potential users via regular newsletters and a published catalogue. a service that is in my opinion typical of local data services, data exchange on floppy discs, is possible but we have hardly any experience with it as yet if one sees a data library as a "local" service backed by a central data clearinghouse or central acquisition and processing centre, then indeed one expects an organisation with a small staft, using external computers and software, concentrating on reference as well as actual dissemination in as friendly a manner as possible, generating subsets, documentation and other special requests. also, seen geographically, the data librar>' should be within one day's travelling distance for its users, for consultation purposes. given this defmiiion, the summer 1986 iassist quarterly 47 steinmetz archive does serve as an example of a data library. are there any other local data services in holland for the social sciences? for survey data: no. for statistical information at the level of cities or regions: yes. there are several specialised databases ovmed by government organizations for planning and policy making, and by university departments, for research and training. these are, on the whole, "local" in the sense of being accessible only to their own people, not to outside researchers, either academic or otherwise. in the second part of this account, i will briefly outiine our approach in the following areas: data acquisition, data dissemination, storage and maintenance of data, dociraientation and reference services. data acquisition is accomplished by routinely checking registers of ongoing research, social science periodicals, and reports of finished research. this is facilitated by the steinmetz archive's participation in the social science information and documentation centre. there is no human network of researchers or fund raisers in the field who could report to the archive interesting projects or data. (nor would i would expect the kind of data library that provides local data service to rely on such a network for data acquisition.) dissemination. users in amsterdam, where the steinmetz archive is located, have direct access to the data through the local university-owned computer centre (sara). from a local termimil, a user can get a copy of a datasei by simply starting a job, that has only one variable: the steinmetz number given to the particular dataset central logging of these jobs and who has started them, is automatically reported to the archive, thus enabling a monthly overview of this type of usage. other users receive the data, and often an spss setup, on tape. this arrangement makes, of course, no provision for users without at least access to a minicomputer with a tape drive and statistical package, such as sf*ss. for example, we are imable by these means to provide service to schools. data transfer to such users should be through floppy discs, a service that we have not really started yet as an archive, with an obligation to disseminate 10 to 15 year old datasets, it requires that we have strong data storage and maintenance systems. this we achieve with a system of multiple tape backups and a tape refreshing scheme to guarantee that no tape is physically more than three or four years old. all tapes are stored in the computer centre. one might expect a data library per se to be more relaxed in these matters; whatever gets lost can be replaced upon request from a central data organisation, but this is not so in our case. developments in "laser disc" technology could ease local mass storage problems, and at the same time ensure long term reliability. how is the user introduced to these masses of carefully preserved data? through documentation and reference services. the steinmetz archive, as mentioned previously, helps users find data suitable to their needs through a catalogue, which is easily, and regularly, produced from a reference database, various indices on microfiche and paper, and through the original documentation produced by the principal investigator. introductions to the principles of empirical research and the data available for secondary analysis by means of making "teaching packages" available and giving lectures at schools and colleges, are other means to assist users. the steinmetz does not give lectures but does offer a teaching package together with the relevant data; the package was developed by an outside institution. data libraries, which should need to put less effort into such activities as acquisitions, processing and maintenance, might profitably put more efibrt into actively getting users acquainted with computer-assisted analysis, data sources, etc. n summer 1986 vol274.indd iassist quarterly winter 2003 5 by by celia russell, keith cole, m. a.s. jones, s.m. pickles, m. riding, k. roy, m. sensier* grid technologies for social science: the samd project abstract the seamless access to multiple datasets (samd) project is designed to demonstrate the benefits of grid (e-science) technologies for dataset manipulation and analyses in a social science context. grid technologies run over existing internet infrastructures and offer a faster alternative to the world wide web for the transfer and analysis of large datasets. under the samd project, a web-delivered social science dataset was made available for large-scale data analysis through a grid architecture. using an exemplar problem drawn from the uk social science community, the project demonstrates how the integration of a single sign-on environment, grid technologies and access to high performance computational resources can significantly speed up computationally intensive queries and streamline data gathering and analysis. the approach can be generalised to virtually any kind of problem involving data retrieval and analysis. the paper also discusses how this could allow social scientists to significantly scale up their quantitative research inquiries. keywords: e-science; e-social science; grid technologies; time series data; multivariate analysis; grid security infrastructure; digital certificates; authentication; authorisation; single sign-on. background e-science grids run on existing internet hardware but offer a more powerful infrastructure than current web technologies. they have been described as a massive extension or even the next generation of existing web services. just as the web is designed to view html documents worldwide, grid technologies are designed to provide seamless access to large scale datasets and to share large scale computing resources, including specialised facilities and visualisation technologies. a number of factors make quantitative social science research well suited to grid based research strategies. human behaviour takes place in a dense social and economic context, yet limitations in computing power and problems of data access inevitably result in models that oversimplify environmental conditions. furthermore, many social scientists now wish to develop more complex research questions by combining datasets from more than one source, perhaps over different geographies or time periods. the datasets themselves are growing rapidly in size; the uk 2001 census datasets are estimated at over 20 gigabytes, twice the size of the previous release. moreover, analysis of data on themes such as trade, economic governance, human development or industrialization requires research strategies that can accommodate a high degree of interdependency between what may be hundreds of disparate variables. the seamless access to multiple datasets (samd) project was funded by the ukʼs economic and social research council (esrc) and the department of trade and industry. the manchester information and associated services (mimas) and the supercomputing, visualization and escience centre at manchester computing in the university of manchester were responsible for implementation. the project was designed to demonstrate the benefits of grid technologies for dataset manipulation and analyses in a social science context. firstly, the project showed how an existing social science dataset can be made available for large-scale data analysis via the grid. secondly, the project demonstrated how the development of parallelised software tools can significantly speed up computationally intensive analyses. using an exemplar problem drawn from the uk social science community, the project showed how the integration of access to both data and high performance computational resources within a single sign-on environment enables the automation of complex workflows and can facilitate the scaling up of social science research applications. samd also demonstrated the successful incorporation of emerging grid technologies into an existing social science data service. the problem samd was based around a genuine econometric research question drawn from the academic community and is based on a typical social science databank, the national statistics time series data. mimas hosts a web interface to this databank, which is widely used in teaching and research. the research question examines the asymmetries in the response of uk gross domestic product to interest rate changes. sensier, 6 iassist quarterly winter 2003 osborn and öcal1 found that interest rate effects on gross domestic product are larger when lagged growth has been high and when there has been a substantial increase in the interest rate. they did not fi nd the opposite effect for low growth phases and decreases in the interest rate. the authors used a non-linear regression, namely the smooth transition regression model, to model this asymmetry. to help fi nd the starting parameters, a computationally intensive 5-dimensional array search program was used. the program typically took several hours to run on a standard mainframe. before the existence of samd, the required data were collected over the web, returned to the userʼs pc and then sent manually to the mainframe computer for analysis. altogether, the collection, collation and analysis of the data took around a working day. moreover, the user had to access a variety of resources to acquire and analyze the data, all of which required authentication via multiple usernames and passwords. the samd solution the samd project used globus 2.0 (globus is the open source software used to create grids) and gridftp to create a specifi c grid architecture incorporating the national statistics time series data. a number of performance improvements, including parallelisation, were made to the analysis code to represent a high performance computing facility. the demonstrator application transfers multiple time series from mimas (a national data centre based at manchester computing) to a high performance computing (hpc) engine for computationally intensive analysis, and then returns the results to the user. the entire operation requires only a single sign-on (grid-proxy-init) on the userʼs workstation. in step (1), the application on the userʼs workstation searches for and requests one or more time series via https with gsi authentication. in (2-3), cgi programs running on the web server verify that the user has permission to access the requested dataset. in (4-7), the requested data are extracted from the mimas data repository and copied to a short-lived fi le, and an xml “ticket” is returned to the application. based on information supplied in the ticket, the application uses gridftp to initiate a third-party fi le transfer from mimas to the hpc engine. steps (1-7) are repeated for each group of time series. in (9-10), globus mechanisms are used to launch an analysis run on the hpc engine and to retrieve the results. a housekeeping task cleans up temporary fi les. figure 1 samd architecture iassist quarterly winter 2003 7 a graphical user interface allows the user to search the databank, extract the data and transfer the resulting time series directly from the databank to an high performance computing facility via grid ftp. the results are then returned to the user. the process is further streamlined by a security model which gives the user access to all the resources to which they have permission through a single sign on procedure. as a result of these advances, the data collection and analysis, which had previously taken around a working day, is reduced to a matter of minutes. the project team also developed a command line shell script that runs the essential steps of data retrieval, transfer and analysis. the scripts enable the query to be scheduled to rerun automatically (whenever the datasets are updated, for example). the shell scripts also demonstrate that the principles of the project can be reapplied to other applications without the need to develop a specialised graphical interface. future research applications by substantially decreasing the time required to locate, transfer data and perform a complex analysis, samd showed how a grid approach could allow the social scientist to significantly scale up their research. the demonstration problem used in the samd project will now be extended to include the effects of international influences (such as us monetary policy decisions) on business cycles in germany, france italy and the uk. in addition to the national statistics time series data, this new model will incorporate data from the oecdʼs main economic indicators and the imfʼs international financial statistics. the use of grid technologies allows the models to be rerun automatically whenever the databases are updated. a grid strategy also enables users to include a much larger number of data points in a computationally intensive analysis. for instance, a current cluster analysis of the samples of anonymised records (a dataset derived from the 1991 and 2001 uk census) uses just 1% of the available data. using a grid approach allows the entire dataset to be used. more generally, using the grid permits researchers to attempt more computationally intensive analyses (including non-statistical modelling methods), or drop assumptions to create a model that more closely reflects the complexities of the real world. the samd architecture also encourages the cross–analysis of multiple datasets, an issue of increasing importance as researchers develop more complex multi-level models, or investigate the interdependencies between economies, societies and the environment. although samd was built around a single databank and computing facility, the architecture can be simply extended to include multiple databanks and computing resources. a single sign-on procedure gives the user access to any resource to which they have an entitlement wherever it is physically located. as a result, databanks hosted by a range of data providers can be accessed and searched seamlessly as an apparently single resource. project outcomes samd was based on a particular substantive problem, but its approach and methods can be generalised to virtually any kind of social science research involving data retrieval and analysis. one aim of the project was to reduce the technical barriers that have slowed the diffusion of grid technologies into social science applications. this was achieved by developing features such as the desktop graphical user interface and transparent access system. these generic elements can be modified and redeployed in future social science grid applications the samd web site (http://www.sve.man.ac.uk/research/ atoz/samd) contains an overview of the project and a page from which various resources developed in the project can be downloaded. these resources include the patches to mod_ssl, shell scripts, presentations and handouts. samd was the first of the esrc e-social science pilot projects to be successfully completed. as such, it provides a concrete example of the benefits e-science could offer in a social science context. consequent to the project completion in 2002, the esrc developed a broader e-social science strategy with the establishment of a national centre for e-social science at manchester. this £7.5 million programme will stimulate the acceptance of grid technologies in social science applications. the five elements of the esrc programme are presented in appendix 1 appendix 1 the esrc e-social science strategy as part of the broad cross council e-science programme, esrc has been developing an e-science strategy to stimulate adoption and use by social scientists of new and emerging grid-enabled computing and data infrastructure, both in quantitative and qualitative research. the esrc escience programme currently consists of the following five integrated components: 1) the national centre for e-social science (ncess) ncess is the key component of the esrc e-science strategy. the ncess will have a distributed structure, consisting of a co-ordinating hub based at the university of manchester in collaboration with the uk data archive at the university of essex, and a set of research-based nodes distributed across the uk. there is an overall budget of £4.5 million for the nodes, which will be commissioned in 2004 and begin work in april 2005. see (http://www.ncess. ac.uk/). http://www.sve.man.ac.uk/research/atoz/samd http://www.sve.man.ac.uk/research/atoz/samd http://www.ncess.ac.uk/ http://www.ncess.ac.uk/ 8 iassist quarterly winter 2003 2) pilot demonstrator projects a programme of 11 small-scale pilot projects began in autumn 2003. these projects are exploring the potential application of grid technologies within the social sciences.). 3) scoping studies the esrc commissioned four scoping studies aimed at identifying the key issues that the ncess should address through its research programme. these scoping studies are available via the esrc website at http://www.esrc.ac.uk/ esrccontent/researchfunding/esciencecentre.asp 4) a training and awareness programme co-funded by the joint information systems committee (jisc) and the esrc, this programme will highlight the potential of e-science within the social science community and develop materials and courses to equip researchers to exploit grid technologies. see (http://www.jisc.ac.uk/index. cfm?name=circular_2_03) 4) establishing a network of access grid nodes for uk social science esrc is establishing a network of access grid nodes (agns) for the uk social science community. this network will facilitate active participation in remote group-to-group interactions for large-scale technical collaborations, e.g., distributed meetings, collaborative teamwork sessions, seminars and training. * paper presented at the iassist conference, madison, may 2004, by celia russell (mimas, manchester computing, university of manchester, kilburn building, oxford road, manchester, m13 9pl united kingdom). contact: celia.russell@man.ac.uk (web-site: url: http:// www.esds.ac.uk/international).) footnotes 1 sensier, m., osborn d.r. and öcal n. (2002) ʻasymmetric interest rate effects for the uk real economyʼ, with centre for growth and business cycle research discussion paper series, university of manchester, no. 10. forthcoming in oxford bulletin of economics and statistics, september 2002. http://www.esrc.ac.uk/esrccontent/researchfunding/esciencecentre.asp http://www.esrc.ac.uk/esrccontent/researchfunding/esciencecentre.asp http://www.ji http://www.ji vol274.indd 20 iassist quarterly winter 2003 iassist call for papers this is fi rst call for papers for the annual conference of the international association for social science information services and technology (iassist), being held in may 2005 in collaboration with the international federation of data organisations (ifdo). proposals for papers, sessions and poster/ demonstrations should be submitted by 10th january 2005. the iassist/ifdo conference is being hosted in edinburgh, scotland uk over the period wednesday 25th to friday 27th may, 2005; this is preceded by workshops earlier in the week and followed by a highland weekend, for a less formal exchange of views. details may be accessed via the iassist website at http://www.iassistdata.org, or directly at http://datalib.ed.ac. uk/iassist/ .iassist is an international organization of professionals working in and with information technology and data services to support research and teaching in the social sciences. typical workplaces include data archives/libraries, statistical agencies, research centres, libraries, academic departments, government departments, and non-profi t organizations, see the iassist website, above, for further information. ifdo was established in 1977 in response to advanced research needs of the international social science community. the purpose of ifdo is to stimulate and coordinate worldwide data services and thus enhance social science research. for further information please visit http://www.ifdo. org/. for over thirty years, iassist conferences have brought together data professionals, data producers, and data analysts from around the world. the annual conference, which itself moves about the world, and was last in europe in 2001, is the forum for presentation of papers covering both new and persistent issues relating to access to data, documentation of data, and digital preservation, with special emphasis on the social sciences. the social sciences have a long history of data sharing activity which may make the conference of interest to colleagues in disciplines where open access practices to data have been on the policy agenda, with clear overlaps with digital curation, data publishing and e-science/ cyberinfrastructure initiatives. the iassist quarterly (iq), available online from the iassist website and in print, is another important means of communication for the data community, and a suitable place for publication of papers presented at the conference. of special note is the iassist publication award, intended to promote the associationʼs five year strategic plan, involving a cash prize for the winning paper. for further details see: http://www.iassistdata.org/publications/pubaward.html . the theme for the 2005 conference, evidence and enlightenment, highlights the need for empirical data in a society that wishes to know itself, and of the role that the iassist membership have in ensuring that researchers have continuing access to the data necessary for furthering scholarship and understanding. it also hints at intellectual activity during the latter half of the 18th century in europe, and in scotland in particular, which “would generate the basic attitudes and habits of mind that characterise the modern age” arthur herman, the age of enlightenment, 2002. iassist quarterly winter 2003 21 iassist / ifdo 2005 edinburg as iassist enters its fourth decade, the organization and the data community must confront a range of socio-economic and organisational challenges as well as technological opportunities. we seek submissions of papers, poster/demonstration sessions, and panel sessions on topics that address these issues, especially those that bear on: a.. data access b.. data documentation c.. data preservation d.. data use and current research activity. additional topics might also include data, information and statistical literacy, gis and spatial data, and such as data publishing, annotation, provenance and authenticity in digital curation. for other key topics see previous iassist conferences at http://www. iassistdata.org/conferences/index.html . procedure the deadline for paper, session, and poster/ demonstration proposals is 10th january 2005. the conference program committee will send notifi cation of the acceptance of proposals on or before 10th february 2005. please send submissions, including proposed title and an abstract (recommended length 150 words) to: iassist05@ed.ac.uk. proposals for complete sessions, conventionally of a panel or of three/four papers within a 90 minute session, should contain information on the focus of the session, the organizer or moderator, and possible session participants. the session organizer or moderator will be responsible for the arranging and securing session participants. make plans to come to edinburgh for the iassist/ifdo conference in the week commencing 22nd may 2005 to help us make the 31st iassist conference ʻthe best conference ever! ̓further information on travel and accommodation is available at links from the iassist website . online registration is scheduled to open on 17 january 2005. 4 iassist quarterly 2016 / vol 40 no 4 iassist quarterly editor’s notes when things get digital and huge. doing the things right and doing the right things. welcome to the fourth issue of volume 40 of the iassist quarterly (iq 40:4, 2016). there is a lot of management involved in the data management carried out at data archives and with data collections. the phrase 'doing the things right and doing the right things' belongs to fathers of modern management and is used to distinguish management vs. leadership, efficiency vs. effectiveness, and tactics vs. strategy. the winning authors of the 2016 lassist paper competition used the article 'more product, less process: revamping traditional archival processing' (mark a. greene and dennis meissner, 2005) as their starting point for investigating the 'more product, less process' (mplp) approach for digital data. the winning paper 'more data, less process? the applicability of mplp to research data' is written by sophia lafferty-hess and thu-mai christian. the authors work at the odum institute for research in social science, university of north carolina at chapel hill as research data manager and assistant director of archives. the paper was presented in the session 'data management archiving/curation platforms' at the iassist 2016 conference in bergen. in their paper lafferty-hess and christian set out to apply the principles and concepts formulated in mplp to the archiving of digital research data. they discuss data quality, usability, preservation and access, leading to the question: what is the ‘golden minimum’ for archiving digital data? in terms of data archiving, spending too much effort on doing the things right may bring the trade-off problem that the resources are not sufficient to do all the things. users in the digital world retrieve and consume lots of information by themselves, but digital data comes in forms that are seldom directly consumable without additional processing. when the authors also relate the phrase 'golden minimum' to the phrase 'good enough', management is again brought into the discussion. in my view, the short formulation of herbert simon's ‘satisficing’ concept in his theory of bounded rationality is 'good enough is best'. lafferty-hess and christian are aware that shifting responsibility for certain data curation tasks from the data archive to the data producer and to the data user can present problems. their best advice and hope for the future is that additional 'future research will help us build better understanding of the connection between user needs and data curation processes'. the following paper 'mmrepo storing qualitative and quantitative data into one big data repository' is authored by ingo barkow, catharina wasner and fabian odoni, working at university of applied sciences eastern switzerland htw chur where barkow is associate professor and wasner and odoni are research associates. they describe a prototype of their mmrepo project that addresses the problem of storing qualitative large binary objects with regular quantitative data in order to achieve the advantage of storing mixed mode data in the same infrastructure, whereby only one system needs to be provided and maintained. linking to the first paper they are looking into the efficiency problem of doing the things right. when you are efficient you can do more things right. the project is trying to achieve this by combining cern’s invenio portal with a hadoop 2.0 cluster and ddi 3.3. the prototype was successful and the project continues. the paper was presented at the iassist 2016 conference in the session ‘technical data infrastructure frameworks’. aidan condron works with the big data network support team at the uk data service. at the iassist 2016 conference he presented ‘data science: the future of social science?’ at the session ‘big data, big science', and has submitted this presentation as the paper 'servicing new and novel forms of data: opportunities for social science'. these ‘new and novel’ forms are, for example, social media data that present potential resources for researchers but also pose challenges for access provision and analysis. the paper introduces data service as a platform (dsaap), which is a project to establish technological infrastructure support. as with the mmrepo project, the dsaap project will include both familiar and new and novel forms of data. the novel forms of data are often huge, and 'hadoop' solutions are also at play here using a data lake built through use of open source software. the article also gives several demonstrations through graphs of energy consumption based on 3.7 billion datapoints. after the presentation and the paper, the big data network support team will standardise and generalise the procedures developed from their dsaap project. submissions of papers for the iassist quarterly are always very welcome. we welcome input from iassist conferences or other conferences and workshops, from local presentations or papers especially written for the iq. when you are preparing a presentation, give a thought to turning your one-time presentation into a lasting contribution. we permit authors 'deep links' into the iq as well as deposition of the paper in your local repository. chairing a conference session with the purpose of aggregating and integrating papers for a special issue iq is also much appreciated as the information reaches many more people than the session participants, and will be readily available on the iassist website at http://www.iassistdata.org. authors are very welcome to take a look at the instructions and layout: http://iassistdata.org/iq/instructions-authors authors can also contact me via e-mail: kbr@sam.sdu.dk. should you be interested in compiling a special issue for the iq as guest editor(s) i will also be delighted to hear from you. karsten boye rasmussen july 2017 editor electronic reference systems in the year 2000: the symposium on advanced information processing & analysis, march 24 26, 1992 by lee a. gladwin' , centerfor electronic records (nnx) national archives and records administration, washington, dc 20408 "i'm drowning in open source information!" is the cry of intelligence analysts, a cry not unfamiliar to many other researchers. how does one navigate through a turbulent paper sea? confronted with a wealth of new open source material (e.g. newspapers, television, technical journals, wire service bulletins, etc.), analysts, who should be interpreting incoming information, currently spend most of their time either sifting and reading documents or writing reports based upon them. contributing to their problems are those of inefficient technological transfer, shrinking resources (funds and p)eople), overlapping r&d efforts, and the lack of system integration. unlike most researchers, however, intelligence community has an organization, the advanced information processing & analysis steering group (aipasg), through which to present its needs to potential contractors and enough funding to attract bidders. aipasg held its annual meeting march 24 26, 1992 in reston, virginia. as with other researchers, analysts need assistance in scanning the data, selecting and organizing documents relevant to their problem, and printing the results. what they do not need is to spend time battling a recalcitrant retrieval system, reading system manuals, attending software workshops, or talking with customer service about software problems. through panel presentations, contractors were acquainted with technical developments in the five areas of need identified by aipasg: intuitive user interfaces document processing, organization and management transparent access to multiple data bases and sources collaborative communications automated data understanding the first need is for a powerful but easily used interface which does not require exhaustive efforts to accomplish simple procedures. any system should be developed in close consultation with the users, not in a vacuum. document processing addresses the need to organize, sort and link documents relevant to an intelligence problem before routing them to the analyst studying those problems. central to the solution of the problem of coping with massive amounts of material is to shift the focus from document retrieval to managing discrete pieces of information, using visual indicators to point to other possibly relevant sources. transparent access to multiple databases addresses the need to know what data is out there and how to retrieve it collaborative communications concerns the institutional barriers to sharing information; i.e., security, organizational territoriality, etc. symposium topics addressed these five needs and such key technologies as continuous speech recognition, natural language and graphical user interfaces, on-line tutors with user modelling capabilities, optical character recognition, automated data extraction, and multiple database correlation. some of the technologies relating to this future system are described in the concluding portion of this report. natural language/text processing papers presented in this session addressed the needs to translate text from foreign languages and then extract information and present it to the user in a meaningful manner. programs were described for translating news items written in spanish and japanese, parsing the text syntactically and semantically, and displaying information in an on-screen template. adrian kleiboemer, mitre, called for creation of reusable environments which could be ported easily to any number of applications without having to build a new natural language frontend (nlf) for every new application as is currently being done. natural language frontends are currently available for searching large databases without forcing the user to learn sql. natural language inc. developed a frontend for oracle which takes short phrases and even senlassist quarterly tences, parses them into sql statements, and runs them against the database. they may be used in conjunction with graphical user interfaces (guis). optical characters recognition and neural networks the cia currently is engaged in a five six year research program aimed at translating source documents, in varying conditions and all languages, into machinereadable format. there are two forms of ocr enhancement digital and repair. digital enhancement clarifies the image using bit mapping and a gray scale. since it simply prints what is there, letters may be broken or run together. digital enhancement cannot recognize letters or words. repair enhancement techniques seek to reconstruct letters and words from faded, damaged or crooked images (i.e., paper orientation). an ibm program was described which uses neural networks to segment pages into regions (e.g., return address, recipient address, stamp, logo, signature), identifies and classifies these segments, and deduces whether the document is a letter, form, article, etc. neural networks (computer simulated biological neurons) are used in pattern recognition tasks to identify printed or written text images. applications are currendy underway at the us post office to read handwritten addresses. document processing, organization & management papers presented in this session dealt with the "derivation and use of statistical procedures for retrieval and data extraction". richard m. tong presented a paper entitled "automatic document retrieval using cart [classification and regression tree]" which classifies and retrieves documents containing at least one of 15 specific sub-concepts of "civil unrest" e.g., "labor union". their initial results show better reuneval results for conceptbased search than by using standard key word searches in a test involving one-thousand news items. relevant to this finding was a remark made by paul thompson, an attendee, who observed that boolean or probabilistic ranking approaches to document retrieval are insufficient to assure a document's relevance. simply adding terms to an sql query only serves to increase errors. data bases & information retrieval large scale information reuieval (lsir) addresses the need to retrieve information from "data sources" that are complex and distributed globally or through different departments of an organization. potential technologies for dealing with an "indexless encyclopedia" are objectoriented databases, hypermedia, natural language processing, parallel processing, expert systems, and information visualization. in addition to meeting the five needs, it was suggested that intelligent user interfaces be developed which could model user expertise, and, in the event of error, infer what the researcher was attempting to do and provide greater assistance in information extraction; i.e., an electronic reference guide sensitive to the researcher's facial expressions and body language. while this electronic reference guide is not yet fully integrated, many of the components are under development and were discussed at the symposium. some day an electronic reference guide will help researchers navigate their ways through text, graphics, art, musical recordings, still and motion pictures in search of information relevant to the problem at hand. undiscussed were questions about the implications of this technology for researchers, research methodology and the reference room in the year 2000. perhaps now would be a good time to begin thinking about these implications. ' article based on notes taken at the symposium on advanced information processing & analysis, held in reston, virginia, march 24 26, 1992. summer 1991 1/17 oluwaseun obasola & rukayat atinuke usman (2024) digitising old yoruba newspapers at kenneth dike library, iassist quarterly 48(2), pp. 1-17. doi: https://doi.org/10.29173/iq1041 the creative commons-attribution-noncommercial license 4.0 international applies to all works published by iassist quarterly. authors will retain copyright of the work and full publishing rights. digitising old yoruba newspapers at kenneth dike library oluwaseun obasola1 and rukayat atinuke usman2 abstract the kenneth dike library and the nigeria national archives are especially rich in ancient collections, particularly those unique to southwestern nigeria, home to many people of the yoruba extraction. these facilities house print and non-print materials such as personal notes and written collections of prominent persons, old manuscripts, ancient and modern maps, journals, and old yoruba newspapers. many of these print materials, especially the newspapers, are deteriorating. in a bid to prolong shelflife, access to these old materials is limited. as newspapers serve as gateways to the past, this restricted access can impact the research experience of users. the paper begins by presenting the project framework, which was designed before the project began. it goes on to detail the nuances involved in the several stages of the digitisation process and considers the aftermath of digitising the papers in terms of ownership, storage, backup, and access. this project revealed two things: first, though digitisation solves the problem of access and preservation, it is still necessary to preserve the original materials to prevent loss due to technical issues. second, funding, and international partnership work hand in hand with digitisation, as it is a capital-intensive activity. last, the paper contributes to the ongoing debates on the cultural, and socio-political discourses entwined with the technical processes of digitisation. the highlighted project was sponsored by the european research council (erc) in collaboration with local partners. the website, https://yorubaprints.wordpress.com/yoruba-erc-project/ raises awareness for the project. keywords digitisation, yoruba, old newspapers, kenneth dike library introduction libraries worldwide have been increasingly focused on digitising their collections. digitisation solves the problem of preserving older materials (liu, 2012), and allows libraries to make their collections available twenty-four hours a day, seven days a week (hampson, 2001). this helps libraries to meet a fundamental demand of the twenty-first century. the evolution from industrialisation to post-industrialisation has directed the world’s interest to a knowledge-driven economy. thus, there is a corresponding increase in the need to supply and access knowledge, both new and old. as for newly created written knowledge, writers and publishers concentrate on publishing e-books or articles instead of printing them in hard copies. climate change has made it even more imperative to seek creative ways of publishing without paper. driving the paperless vision conserves trees, and contributes to saving the planet. these are increasingly known as "born digital” resources. https://doi.org/10.29173/iq1041 https://yorubaprints.wordpress.com/yoruba-erc-project/ https://creativecommons.org/licenses/by-nc/4.0/ 2/17 oluwaseun obasola & rukayat atinuke usman (2024) digitising old yoruba newspapers at kenneth dike library, iassist quarterly 48(2), pp. 1-17. doi: https://doi.org/10.29173/iq1041 the constant creation of knowledge, however, does not obliterate the relevance of old print materials. their relevance is especially conspicuous in the humanities, where “old” knowledge is the foundation on which newer research can stand, providing anchors for future works. the all-important nature of relics and old material necessitates their preservation. it is for this reason that libraries, as repositories of knowledge, turned to digitisation, engaging technology to digitise cultural heritage. some of these old materials include ancient manuscripts, pamphlets, personal notes and letters, newspapers, and so on. however, what does the term “digitisation” mean? though this article is primarily focused on documenting the technical and ethical processes and implications of the kenneth dike library (kdl) experience in digitising old yoruba newspapers, it is necessary to define how we use “digitisation” in this context. in most articles engaging digitisation, words like – photography, preservation, conservation, accessibility, online presence, digital humanities, and heritage – frequently appear, and are sometimes used interchangeably. king (2005) points out that digitisation can be traced back to the twentieth century when photography debuted for research purposes. this implies that digitisation is not merely taking pictures of “endangered” materials for storage. corrado and sandy (2017) added to the list of what digitisation is not, in the first chapter of their book. succinctly, digitisation is not merely about having backups, as backups are more about continuity and continuity can be guaranteed via various methods. digitisation is not merely about increasing access, as access can be ensured by creating digital or even physical surrogates. neither is digitisation an afterthought resulting in a rush for preservation departments to preserve their holdings. rather, it is a carefully conceived process mediated by an intersection of experts in various departments of technical and policy teams. though grycz (2006) acknowledges that the term “digitisation” suffices in the realm of digital imaging because this kind of imaging involves capturing the pictures of images, he also asserts that in the context of the activities of archives and libraries, digitisation is a “process”. it is a process or series of activities that involves taking images of materials, recording this data, refining these images as needed, storing them in databases, and multiplying or reproducing them to guarantee access and longevity. digitisation can thus be placed at the intersection of these terms and activities. the kenneth dike library (kdl), named after the first indigenous principal and former vice-chancellor of the university of ibadan, professor kenneth dike, was founded in 1948. up until 1970, the kenneth dike library was the major national depository for nigerian publications. from june 1970, it has shared the privilege of a central depository library with the national library of nigeria which became the senior partner and the centre for the compilation of the national bibliography. the library contains approximately 700,000 volumes of books and receives over 6,000 separate journals and other serials. the library collection also consists of the africana collection, arabic books, and manuscript, government publications, maps, publications ordinance, theses and staff publications. through macarthur foundation’s funding (2002-2007), kenneth dike library received great boosts in all aspects of its functions and services to the immediate community and beyond. the kenneth dike library has a digitisation chamber provided by the macarthur foundation as a part of the systems unit in 2010. the kdl has more than a decade of experience in digitisation. it digitises both old and new publications. the library has large deposits of national and regional collections because of its former role as a national library before evolving into an academic library. https://doi.org/10.29173/iq1041 3/17 oluwaseun obasola & rukayat atinuke usman (2024) digitising old yoruba newspapers at kenneth dike library, iassist quarterly 48(2), pp. 1-17. doi: https://doi.org/10.29173/iq1041 these numerous collections, with different risk exposure levels, create a strain on the digitisation chamber, which has limited human and capital resources. many of these collections, which are stored on different floors and segments of the library, urgently need to be digitised so they are not lost forever. this article focuses on old yoruba newspapers held in kenneth dike library, the main library at the university of ibadan, and at the nigeria national archives, ibadan. beyond advocating the urgent need to digitise delicate materials, this article contributes to the advancement of the nascent field of digitisation in africa. the yoruba-prints project revealed that digitisation projects are largely isolated from each other, and also rarely engages in knowledgesharing, consequently creating a paucity of relevant and contextual knowledge that can guide or provide illumination to other digitisation projects. adversely, this has produced a vicious cycle of trial and error in digitisation projects, thereby slowing down the digital humanities drive in nigeria. this article provides detailed and hands-on approach to digitising delicate print materials. further, it contributes to the global discourse in digital humanities by interrogating the philosophical debates underlying the politics of memory, heritagisation, and digitisation such as questions on what is digitised, how, why, and for whom. we discuss this in the section titled ‘the socio-cultural and political questions in digitisation’. finally, the article contributes to the future of digitisation in nigeria, so forthcoming projects can learn from the successes and challenges of this project, thus furthering the frontiers of digitisation, specifically for projects dealing with fragile print materials and generally for any project involving the use of photographs. the need to digitise and early efforts libraries with collections of old newspapers are increasingly faced with the dilemma of meeting user needs by providing access to these fast-disappearing materials, and also preserving the old materials for future use (king, 2005). these obligations are opposed to each other, as providing unrestrained access to already low-quality or fragile papers, would lead to the rapid deterioration of the materials, which would then shorten their lifespan, so they would not be able to reach future users. mieczkowska and pryor (2002), while giving a brief history of newspaper production in britain, explained the reason for printing newspapers on (lower quality) wood-pulp paper. the shortage of linen and old cotton in the 1850s the canvas formerly used to write/print newspapers, forced production to shift to using wood pulp for paper. however, wood-pulp paper had a crucial limitation. it contained a chemical property – lignin, which turned the paper’s colour from white to yellow upon interaction with light. this yellowing resulted in reducing the paper shelf life. though the production of acid-free paper through laser technology has removed the problem of yellowing papers, papers used for producing newspapers are still very brittle and susceptible, with the earliest papers facing the most critical threat of irredeemable destruction (corrado & sandy, 2017). the period in which britain started printing on paper was not too distant from when the printing press found its way to nigeria via british missionaries. thus, the type of paper used in printing newspapers in nigeria was not much different from what was obtainable in britain. in ancient islamic centres of learning, direct hand copying of manuscripts was the method scholars adopted to produce multiple copies, improve access, and preserve knowledge. the digital turn was heralded when more recent scholars, particularly in the global north (europe and america), used microfilms to preserve early newspapers. king (2014) asserts that microfilms were quite viable in attempting to preserve the content of newspapers because if the process of creating quality microfilms (i.e., using the photographed negatives of the newspapers) is properly followed, reeled https://doi.org/10.29173/iq1041 4/17 oluwaseun obasola & rukayat atinuke usman (2024) digitising old yoruba newspapers at kenneth dike library, iassist quarterly 48(2), pp. 1-17. doi: https://doi.org/10.29173/iq1041 out, and stored afterwards, newspapers can then be accessed without damaging the microfilms. they also have other advantages, such as the ability to last more than two hundred years, minimal technical know-how, and reduced storage space consumption. however, microfilm as a medium has challenges too, such as the complicated procedure of loading the reels into reading machines, it is hard on the eyes to view for prolonged periods, and it is impossible to search texts except by manually reading every single page and also requires electricity. in recent times in nigeria, the need to preserve historical records and heritage materials has moved the staff of galleries, libraries, archives, and museums (glam) to explore several options in partnerships and methods of digitising forms and volumes of heritage materials. however, the sparse availability of literature documenting these efforts has left these digitisation efforts nearly unseen (pickover, 2009). thus, this article was conceived in response to the recognition of the importance of documenting and publishing the nascent digitisation projects in nigeria, and also to hauswedell et al’s (2020) call for countries and institutions (especially in the global south) to document the processes and dynamics of digitisation within their contexts. further, this article contributes to the ongoing debates on the ethical, moral, cultural, and socio-political discourses entwined with the technical processes of digitisation. the socio-cultural and political questions in digitisation many authors and scholars in digitisation have called for a reflection on deeper level philosophical and political questions on what is digitised, how, why, and for whom they are selected to be digitised on the one hand, and where and how one can access the digitised materials on the other hand (bianchi, 2006; pickover, 2009; ugah, 2009; hauswedell et al., 2020; zaagsma, 2022). this is important because the underlying discourse proves a relationship between archives (information in general), knowledge, social memory, and power (zaagsma, 2022), which also introduces digitisation into the politics of heritagisation. this article engages these discourses by contextualising heritage and archives digitisation in two contexts. first, the nigerian-yoruba or african methods of historicising, and second, the global north-south power dynamics. finally, similar to zaagsma (2022), we achieve this by transposing the what, how, and why questions of digitisation in the traditional archiving processes to the digital domain; processes such as selection, collection development, cataloguing and classification, circulation and access3. in a bid to address these critical questions and advance the debates, we emphasise that this is not an attempt to justify any form of politicisation in digitisation efforts, nor to provide solutions to the complex issues, but to explain our reasons for and limitations in making these choices. the following sub-headings address these questions. selection – the criteria for selecting the materials to be digitised entails a re-selection out of preselected and already catalogued old newspapers in the kdl and national archives. therefore, determining what is picked and what is left out, according to hauswedell et al. (2020), is inherently politicised. why were old yoruba newspapers selected for digitisation? why not select the oldest newspapers from each nigerian region, given the multiplicity of tribes in the country? where are the earliest copies of published newspapers and why are they not part of this digitisation efforts? with over fifty thousand volumes of 19th and 20th century newspapers, kdl, like many other organisations, does not have the required technical, staff, and financial strength to digitise everything. selecting old yoruba newspapers as the focus of this project was based first, on a hierarchy of needs and the requirement of a project funded by the european research council. early newspapers are https://doi.org/10.29173/iq1041 5/17 oluwaseun obasola & rukayat atinuke usman (2024) digitising old yoruba newspapers at kenneth dike library, iassist quarterly 48(2), pp. 1-17. doi: https://doi.org/10.29173/iq1041 among the most vulnerable print materials held in libraries/archives; thus, they are prioritised for preservation above other forms of print. yoruba newspapers top this list because they represent the earliest history of local/vernacular newspaper production in nigeria and africa (omu, 1967). additionally, kdl and the national archives in ibadan are particularly vested with yoruba newspapers, under the collection allocation strategy in nigeria. subjecting newspapers already in a precarious state to transport from one region of the country to another would further endanger them. neither would it be cost-effective to transfer already over-burdened experts from ibadan to other regions to digitise other regional newspapers. in a world where what is available online is increasingly becoming “the only accessible” data, we are aware of how the availability of one set of data builds memory politics, particularly for coming generations (britz & lor, 2004). however, we hope that digitisation of other critical materials in nigerian history will follow very quickly, and even more importantly, that those digitisation processes are documented and published. this would build a framework for a holistic analysis of digitisation efforts in glam institutions in nigeria. classification and metadata – metamorphosing catalogued print materials into digitised materials necessitates a re-classification of the newly digitised materials to ensure that they are findable and accessible online (pickover, 2009). how can the newly digital materials be organised and described without reproducing the existing power asymmetry in knowledge politics, while maintaining a practical approach to findability? for searchability, we adapted the dublin core metadata schema for the digitised newspaper. although the metadata has its own limitations (translational issues), it was quite easy to adapt since a previous project in the library had already successfully used the schema. we adapted nine elements out of the fifteen elements of the dublin core metadata schema: title, date of publication, publisher, keywords, description/abstract, naming conventions, type of document, and volume and issue number. translational challenges – the schemas of available metadata are not quite cultural-diversitysensitive (ma, 2020). despite the knotty questions on translational challenges, and the issues we considered on the ethics of using summary or “umbrella” english words4 as keywords; we planned to adopt the dublin core metadata schema to define the keywords in yoruba and english because the intent of the digitisation project is to widen accessibility to all library users, some of whom may not be speakers of the yoruba language. we figured out the ‘best’ english words that capture the meanings of some yoruba words/phrases that google translate could not adequately capture. again, since the metadata is purely descriptive, the keyword search can serve as a general navigation route for non-yoruba speakers/readers, while an advanced search would find specific yoruba words. though this does not solve the problem, it offers an interim approach to coping with the wider problem of african knowledge production in an academic clime dominated by non-african research methodologies. history of newspapers in nigeria (yoruba land) the modern printed newspaper has a nearly 400-year existence (ugah, 2009). through this period, newspapers have been used for recording and storing information, happenings of people, events, and places. apart from merely recording texts, they also contain photographs showing the everyday life of the people they serve. the ability of newspapers to provide investigative journalism, and educate readers, is partly because they have been owned by people with strong ideological drives. newspapers have also served as agents of social change and revolution (curras, 1987). local newspapers are a https://doi.org/10.29173/iq1041 6/17 oluwaseun obasola & rukayat atinuke usman (2024) digitising old yoruba newspapers at kenneth dike library, iassist quarterly 48(2), pp. 1-17. doi: https://doi.org/10.29173/iq1041 primary source of history and community data. newspapers occupy a pivotal position in transmitting information on several levels (onwubiko, 2005). they transmit information from the past to the present and provide regular and up-to-date information (ugah, 2009). they can show the gaps and strides of a society, and they can play entertainment roles. mostly, newspapers are a rich source of information about politics, ideas and thoughts, commercial, economic, legal, and just about everything about the era they were published. their relevance to issues attracts a wide body of researchers. newspaper production in nigeria began in the late nineteenth century in south-western nigeria, among the yoruba. thus, the earliest history of newspaper publication in nigeria parallels the history of yoruba newspaper production. the term 'yoruba' describes a language of the niger-congo family, and the people who speak it. they predominantly occupy the current southwestern part of nigeria, a part of benin republic and togo (ogundiji, 2003; barber, 2014). within nigeria, the yoruba are famous for ancient art works such as the ife bronze figures, and a history replete with pockets of external influences; from arabic traders and scholars, and later on, western/western-formed missionaries. the arabic influence led to the foremost attempt to put yoruba in written text. thus, the earliest form of written yoruba was formed by adapting an arabic script called ajami before the 1850s (olumuyiwa, 2013). however, ajami neither reigned for long nor spread widely. extensive yoruba texts began with the early stages of colonisation, in the mid-nineteenth century, with missionaries of the church missionary society playing pivotal roles such as writing the tonal language in latin text (ogundiji 2002; olumuyiwa, 2013). the periods between the entrance of the printing press, the development of latinised yoruba and inception of yoruba newspapers coincided roughly around the mid-nineteenth century. the entrance of the newspaper in nigeria began with the establishment of iwe irohin fun awon egba ati yoruba, (popularly known as iwe irohin), a yoruba newspaper based in egbaland, abeokuta, southwestern nigeria (omu, 1967). it was founded by a member of the anglican laity, rev. henry townsend in 1859 and was published every fortnight (omu, 1967). its content was limited to church activities and social events. it neither reported crimes nor entertainment. the advent of christian missionary activities, across south-west nigeria, abeokuta, in particular, steered the content and language towards people represented in these communities. this created a number of limitations. first, it was limited to the yoruba elite, as formal education was not yet commonplace in the 1850s. second, its concentration on the people of egba, and then the yoruba race, generally, meant the side-lining of other regions. however, these problems were justifiable because nigeria as we now know it, was not established. its regions were autonomous. as the church was primarily established for religious purposes, it held no obligation to print judicial, crime, or entertainment features in its paper. despite this, it featured strong opinions on anti-slave trade and promoted civil rights and education (maringues, 2001). omoloso and abdulrauf-salau (2014) assert that the newspaper, though not the first in africa, is known to be the first “vernacular” newspaper on the continent. it was a widespread notion in colonial nigeria and indeed, africa, that local or indigenous languages were vernacular, inferring that they were inferior to the english language and other colonial languages. reverend townsend, initially did not appear bothered about the second-fiddle reference given to the newspaper he founded because “his motive of establishing a local language newspaper was in keeping with some of the main functions of the press, which are to inform and educate” (omoloso and abdulrauf-salau, 2014) and to “beget the habit of seeking information by reading” (alabi, 2003). in 1860, just one year after publication started, an english https://doi.org/10.29173/iq1041 7/17 oluwaseun obasola & rukayat atinuke usman (2024) digitising old yoruba newspapers at kenneth dike library, iassist quarterly 48(2), pp. 1-17. doi: https://doi.org/10.29173/iq1041 translation was attached to the newspaper, making it a weekly published bilingual paper. sadly, less than a decade later, the publication came to an abrupt end. this unfortunate development, as opined by salawu (2004) paved the way for many other newspapers indigenous, bilingual, and english. for example, iwe iroyin eko which started in 1888, published monthly by andrew thomas was the immediate successor of iwe irohin (omoloso & abdulrauf-salau, 2014). newspapers written solely in the english language were already being published before iwe irohin folded. in 1863, the anglo african, edited by robert campbell, was the second published newspaper. while the nineteenth century heralded the production of newspapers in yoruba land and nigeria as a whole, the twentieth century represented the glory days of yoruba news production. in 1922, adeoye deniga founded eko akete, a newspaper that lasted for more than a decade and was far more diversified in content compared to earlier publications. this project engaged part of the earliest printed eko akete from 1922 into the 1930s. eleti ofe followed in 1923, iwe iroyin osose and eko igbein in 1925, and akede ekoin in 1927. these daily and/or weekly papers were not intended to serve as a permanent storage of information (stoker, 1999). thus, many are now in a deplorable condition and are being eaten up by worms and insects. librarians and researchers have urgently searched out means to preserve them. pre-digitisation framework at kenneth dike library according to holley (2003), strategic planning before undertaking any form of digitisation project is important because digitisation affects some critical aspects of the library, such as current infrastructure ict equipment and software, organisation structure, staffing, and service delivery. before the project commenced on the 7th of september 2021, a framework to determine the number and variety of collections held by the library was drafted, with projected work procedures and timelines designed. the framework was agreed upon by the library, sponsors, and other partners for the project. procurement of new equipment and employment of ad-hoc staff were also gradually put in place from about three months before the digitisation process began. 1. number and variety of newspaper collections the kenneth dike library and the nigeria national archives, ibadan, are both historically significant as holders of old and rare collections in the country. the university of ibadan is the premier university in nigeria and among the earliest in west africa with the capability and capacity to hold many significant collections. the city of ibadan is also relevant as a publishing hub, and historically, as a metropolis of the yoruba, in proximity to other big towns famous for publishing early yoruba newspapers, such as lagos and abeokuta. the library has more than fifty thousand volumes of old newspapers, and about five thousand are old yoruba newspapers. out of this number, those most vulnerable to extinction were first selected for digitisation. an opportunity cost model was used to maximise the limited available resources as digitisation is labour and capital-intensive. samples were picked from the 1920s (the year of the yoruba newspaper publication boom) into the 1970s when indigenous newspaper publication was already on a precarious and steep decline. samples include: ⚫ eleti ofe1950 to 1965 ⚫ gbohun-gbohun1971 ⚫ imole owuro1974 to 1978 ⚫ irawo obokun1955 https://doi.org/10.29173/iq1041 8/17 oluwaseun obasola & rukayat atinuke usman (2024) digitising old yoruba newspapers at kenneth dike library, iassist quarterly 48(2), pp. 1-17. doi: https://doi.org/10.29173/iq1041 ⚫ irohin yoruba – 1951 to 1973 ⚫ irohin imole1952 to 1966 ⚫ irohin onigbagbo1963/64 these volumes were packed in labeled boxes in the reference section of the kenneth dike library. some boxes did not contain the complete publications for the marked year because some publishers were inconsistent, and other volumes were simply unavailable. during this stage of determining the numbers, type, and state of the collections available, it did not occur to the researchers to check for duplicates. it was later discovered upon the commencement of the project that some issues were repeated thrice. collections from the archives, on the other hand, did not have any cases of duplication. they were bound; backdated and had different publications from the collection at kenneth dike library. samples drawn from there include: • eleti ofe1924 to 1927 • the yoruba news – 1924 to 1928 • eko akete – 1922 to 1925 • akede eko – 1931 to 1937 in total, this stage took about 2 weeks. 2. projected work equipment, procedure, and timeline the next stage addressed the technical and systems-related concerns. this included the type of software and equipment that would be compatible with the project and what would produce the best outcome. since the digitisation chamber of kdl had been engaging in digitising recent print materials before this project, the choice of software intuitively had been resolved. however, the decision on the exact type of equipment to be used was determined by the nature of the collections. scanning using a sheet-feeder scanner, for example, was impracticable because of the size and brittle nature of the old newspapers. also, loosening the bound collections would subject the materials to deterioration. while the use of a flatbed scanner may have been helpful, the massive size of some of the old newspapers made this impossible. gbohun-gbohun 1971 for instance, measures 53cm x 35.5cm when closed, which was impossible to capture even using a large flatbed scanner which measures 45.5cm x 30cm without cutting out some parts of the large newspapers. the content of the metadata that would be generated to describe the collection was also considered at this stage. the dublin core metadata schema used by the university of ibadan for digital objects was adapted for the newspaper digitisation project. 3. approval processes the approval obtained for the project was in a two-part policy brief – internal and external. the internal aspects took keen consideration of the organisational aspect of the project. the collections were to be retrieved from the reference section of the kenneth dike library, meaning that an internal memo needed to be written to coordinate activities between the reference and digitisation sections. similarly, coordination was required to get approval to work with the collection from the archives. as there was no digitisation policy in place at the university of ibadan, the project team had to work with https://doi.org/10.29173/iq1041 9/17 oluwaseun obasola & rukayat atinuke usman (2024) digitising old yoruba newspapers at kenneth dike library, iassist quarterly 48(2), pp. 1-17. doi: https://doi.org/10.29173/iq1041 the digitisation manual/guidelines developed by the system unit team at the kenneth dike library. the kdl digitisation manual was followed throughout the duration of the digitisation project. 4. procedure for digitisation after taking a long-range view of the process, we developed an action plan to achieve the digitisation of the newspapers. it is important to note that the framework drafted was constantly updated, as the plan and the reality of executing the plan turned out to be different. for example, the framework considered using the software, macromedia fireworks for editing the captured images of the newspapers. however, after a test run, it was discovered that the newspaper titles edited using this software were not accessible after converting them to tiff. simply put, they could not be viewed. this necessitated reviewing the software and switching to using lightroom for editing the images. this does not mean that the macromedia fireworks software is not suitable for digitisation. other digitisation works such as digitisation of book materials and other texts, carried out in the digitisation chamber of the kenneth dike library are edited and cleaned using this software. the processes below sequentially describe how the old yoruba collections were digitised, while simultaneously noting the lessons learned on the way. i. set up – this project began by setting up the purchased hardware components. except for the copy stands, cameras, laptops, and uninterrupted power supply (ups), other equipment used in the project was locally sourced. the capturing/placement boards and laptops were bought in the netherlands and shipped to nigeria. they include: uninterrupted power supply (ups) standing led lights copy stands a server laptops cameras solar power installations router the project did not start with the procurement of all the listed equipment. for example, the server, solar power installation, and laptop arrived several weeks after the commencement of the project. the project commenced with the initial procurement of the following materials: two canon 4000d digital cameras, two placement boards, and two laptops. later, a router for internet connections, solar installations for constant electricity, and a server were procured. a pack of nose masks and a bottle of hand sanitizer were procured for the staff to observe covid-19 protocols. it was important to protect the equipment from the dust on the old volumes of papers. a space within the digitisation chamber was cleared to make room for the new equipment. the capturing boards were laid on a regular worktable. each board had a stand attached to the top of its mid-section, where a digital camera was fixed. the camera was then connected to the laptop/server through a wire connection. before the cameras were fixed, the settings such as the light scale, size, etc. were set on the same parameters for both cameras as a quality control measure, to ensure similar output. factors like different lighting in the areas where the two capturing boards were placed, weather conditions such as rainy seasons, when it got dark due to clouds, and most importantly, the different hues of yellow, brown, and grey of different newspapers meant adjusting the camera settings until we got a desired pictorial result. also, acquiring devices from outside the country, i.e., the laptops, was not a promising idea, as just https://doi.org/10.29173/iq1041 10/17 oluwaseun obasola & rukayat atinuke usman (2024) digitising old yoruba newspapers at kenneth dike library, iassist quarterly 48(2), pp. 1-17. doi: https://doi.org/10.29173/iq1041 two weeks into the project, they malfunctioned. the laptops tended to overheat while in use, perhaps because of the climate differences between tropical nigeria and the cold. figure 1: schematics showing the framework adopted for the digitisation of yoruba newspapers at kdl https://doi.org/10.29173/iq1041 11/17 oluwaseun obasola & rukayat atinuke usman (2024) digitising old yoruba newspapers at kenneth dike library, iassist quarterly 48(2), pp. 1-17. doi: https://doi.org/10.29173/iq1041 figure 2: a camera, copy stand, two light stands and a computer system (server) with ups used for the digitisation of yoruba newspapers at the kenneth dike library, university of ibadan (2021). software several kinds of software were bought and installed on the computers used for the execution of the project. they include: ⚫ lightroom used for editing and converting images captured with the cameras into cleaner versions. ⚫ adobe acrobat 7 converts images in tiff format (tagged image file format) as saved in lightroom, to pdf (portable document format). ⚫ abbyyfinereader files are converted into searchable files. it changes the newspaper form from images into text (i.e., ocr optical character recognition). ⚫ microsoft office applications word and excel were employed. excel was used for the collection of the metadata gathered from the newspapers while word was used to type documents. ⚫ anti-virus “bitdefender” was bought to protect the computers and files from viruses. ii. acquisition and dusting the newspapers or print materials were first acquired from the national archives and/or the reference section of the university of ibadan library. they were then cleaned https://doi.org/10.29173/iq1041 12/17 oluwaseun obasola & rukayat atinuke usman (2024) digitising old yoruba newspapers at kenneth dike library, iassist quarterly 48(2), pp. 1-17. doi: https://doi.org/10.29173/iq1041 with soft dry towels to remove dust, carefully spread out to stretch wrinkles from squeezed papers, and placed on the capturing boards. iii. sorting and capturing the newspapers were stored in boxes in the reference section. each box was examined to check for duplicate copies and to confirm the year(s) of publication. the sorting of duplicates is another way in which the planning stage differs from that of the execution. while compiling boxes for this project, the labels on the boxes were considered to be accurate and properly sorted. but on commencement of the project and more thorough examination, it was discovered that some newspaper issues had many duplicates, sometimes, even packed into different boxes. this slowed down workflow for a while, as sorting duplicates was not designed in the initial framework drafted before the project began. after the duplicates were removed, the remaining issues were sequentially arranged so each staff could clearly mark the issues they discovered. this made later collation of all the captured images easier. one after the other, the newspapers were placed on the capturing boards. the camera was then set at a height that adequately covered the face of the print material and a picture was captured using the canon 4000 d camera. the image was saved in a document format known as raw or cr.2. to avoid a mix-up among the newspaper issues, separate folders were created for new newspaper brands and their different issues. iv. renaming – cameras automatically title the file of the images taken during the capturing process. however, these titles do not reflect the standardised naming convention designed for each volume. each captured image/file had to be renamed to reflect the naming convention. the naming convention involved using the essential details of a newspaper in a shortened format. for example: “eo_jun_17-23_1961_xiii_694” represents – eleti ofe (newspaper title), june 17 to 23 (date), 1961 (year), xii (volume) and 694 (issue). codes like this were designed for each type and volume of newspapers digitised. v. editing – it involved cropping, cleaning, and sometimes rotating the pages of the captured newspapers until they appeared in the standard form (portrait and straight view). lightroom software by adobe was utilised for this stage. each page was cropped, cleaned, lightened, and adjusted as required, and then compiled by issues in different folders. a crucial factor that needed harmonising during editing was the setting of pre-sets. a pre-set is the adjustment of all elements, such as texture, shadow, brightness, and clarity. these options in the lightroom software were saved and titled according to each newspaper brand. a pre-set is usually designed on the server and exported to each laptop using the router, so they are accessible on multiple laptops for use by other staff. this is to control quality and ensure a uniform appearance for all the files. however, this approach was plagued with challenges. the yellow shades of the newspapers, even if they were of the same title, were different. some appeared yellowish-brown, while some tilted towards grey and orange. in fact, various hues were observed across different papers, sometimes contained within a particular volume. this meant that the pre-set had to be manually adjusted, issue by issue, and page by page, to produce a good result. the advantage of this manual process is that after a while, one becomes familiar with the elements of the software and the newspaper. knowing what to adjust, and to what degree, becomes relatively easy. it must be noted that it was easier to edit on the server than on other laptops used for the project. this is because the server hosts all of the original captured newspapers. the other laptops needed to be connected to the server via the router, or directly copied using a hard drive to the memory of the laptops, and then had to be re-transferred to the server after editing. hard drives were used for copying instead of flash drives because the latter are highly susceptible to corruption. after editing, lightroom saved the files in tiff. vi. converting to searchable format the tiff format in which lightroom saves the files cannot be viewed by some image or document viewers nor can they be seen when readers search texts. to enable these features, the pictures were first converted into pdf using adobe acrobat 7.0 and then converted from an image to searchable texts using optical character recognition (ocr), also known as searchable pdf, with abbyyfinereader. each of these formats was carefully collated into different https://doi.org/10.29173/iq1041 13/17 oluwaseun obasola & rukayat atinuke usman (2024) digitising old yoruba newspapers at kenneth dike library, iassist quarterly 48(2), pp. 1-17. doi: https://doi.org/10.29173/iq1041 folders and labeled appropriately. early versions of the files, i.e., cr.2/raw and tiff were not discarded. they were kept as a check in case of any dysfunction that might arise while transforming file formats and as an additional backup. after conversion to ocr, the team vetted the searchability of the words. searches using the english keyboard did not yield particularly distinct search results. we attribute this to the latinised rendition of the yoruba alphabet by the software. vii. collation of metadata metadata is an important part of information storage, sorting, retrieval, and use. the thought of having to retrieve a piece that has no title, or any form of description is mentally challenging. this illustrates the necessity of generating metadata for collections. according to corrado and sandy (2017), metadata is structured information about a resource. debates exist among digital librarians most especially, on the classification of metadata, but corrado and sandy (2017) stick with four types: descriptive, technical, administrative, and structural metadata. this project was concerned with descriptive metadata, the basic information that makes it easy to retrieve a document. upon completing the digitisation of a box or bounded copy, the information/metadata from the newspapers was extracted and organised on excel spreadsheets. the details include the title of the newspaper, date of publication, publisher, keywords, description/abstract, naming conventions generated for each paper, type of document, and volume/issue number. each excel spreadsheet had columns created for each piece of information. separate excel files were created for each newspaper brand, while separate sheets within a file were opened for different newspaper years (year of publication), to make for a neat and clear organisation. to save time, metadata could be created by a staff thereby creating specialization in the workflow. in the alternative, each staff engaged in the digitisation process could enter metadata into a pre-designed template and convert the same to pdf. the latter method was used in this exercise. viii. uploading digital objects into the institutional repository (ir) this involves the compilation of the final format of each file (i.e., ocr) and collecting their metadata to be inputted into the online institutional repository of the kenneth dike library, university of ibadan. the online ir of the university is hosted by dspace and administered by the staff of the systems unit of the library. it is helpful to have a stable internet connection, to facilitate prompt uploading of files, for access at any time, from anywhere in the world. challenges the digitisation of old yoruba newspapers at the kenneth dike library had many challenges. the challenges are highlighted as follows: 1. image recognition deficit – the pages of newspapers are usually dotted with several images and photographs which may add meaning to a written piece or stand as its own body of knowledge. this project was technically limited in having fully detailed images like the ocr done on texts. 2. ocr limitation – significant variation in font sizes; elaborate or tiny font, and blurred texts, often make it difficult for the ocr to conduct a credible search of the affected words. this could lead to a false search. the settings did not affect the reading of the characters when converted to ocr. however, when we tried searching for keywords in yoruba with ascent marks using the search box, it was a bit difficult to get accurate results. searching using the installed yoruba keyboard helped. 3. different screen types of the computer systems used a total of five computer systems were used, though not more than three were active at a time. none of these systems were the same. they had different specifications such as screen resolution, capacity, dpi/ppi (dot per inch/pixel per inch), and so on. also, the laptops were business editions; they were not graphics-dedicated nor did they have nvidia (a feature that is usually on graphics computers). this caused occasional software problems. for example, pages could suddenly re-arrange themselves in a mixed form of https://doi.org/10.29173/iq1041 14/17 oluwaseun obasola & rukayat atinuke usman (2024) digitising old yoruba newspapers at kenneth dike library, iassist quarterly 48(2), pp. 1-17. doi: https://doi.org/10.29173/iq1041 portrait and landscape view, the software could hang up for a while or even suddenly close or stop responding. 4. irregular columns – the text layout in columns of some of the newspapers is sometimes bent, bending into the page borders, even outside of printed areas. this makes cropping and page size harmonisation a little difficult. thus, we sometimes had no choice but to leave some pages and texts bent. 5. loss of text many reasons account for this: segments eaten up by bookworms, brittle and lost parts of the papers, creased papers (this also sometimes affects ocr), and the spine curvature of bound volumes may make text difficult to be captured clearly. 6. indexing indexing of these old papers is laborious for several reasons. many of the papers lack consistency in frequency of publication. for example, within the archival materials, one can find eleti ofe, a supposed weekly publication as written on the paper, dated 17th to 23rd of february 1924, and the next issue is dated 19th of february 1924. the publishers sometimes repeated issue numbers when this occurred. the publisher and editor could also change abruptly. these subjects are relevant to enable a keyword search and systematic arrangement of these materials. this makes metadata collation very tasking and important. publisher inconsistency thus limited our ability to provide an authoritative digital data navigation keyword guide for users. these problems also made the use of indexing software impracticable. 7. format migration issues this is mostly caused by software problems. the page orientation and colour consistency of some newspapers could be distorted while changing them from one format to the other. the page arrangement also gets distorted if the pages exceed nine (9) pages while editing on the dell xps laptop. they are also varied while copying from one system to the other, especially if the different formats are archived on several different systems/computer units. data could be lost in the format migration process. conclusion digitisation is only the first step in the efforts being made by scholars towards digital preservation. corrado and sandy (2017) corroborate this as they stated that digital materials are unlike “physical collections of books, manuscripts, or artifacts that can be neglected for years without significant loss or additional expense,” because digital materials can be affected by bit rot, software malware and many other types of “digital accidents and disasters”. the first step towards preserving the digitised items is to store them on an open-source digital library software such as the one already employed by the kenneth dike library, dspace. it continues with digital asset management as technology is constantly evolving. file formats, for example, may become obsolete at any time. with an already set up stand-by digitisation chamber/unit at the kenneth dike library, the digitised materials are unlikely to become obsolete without warning. staff are trained and re-trained in the face of technological development and trends. these staff, in conjunction with the systems unit, keep a watchful eye on these trends. repository maintenance/administration, access provision, access control, and user support are provided by the systems unit team of the university library, thus, reasonably securing the future of the digitised materials. intellectual property rights data protection regarding this project is not typically concerned with the local distribution of these newspapers, either in physical or digital form. it is concerned rather, with the protection of a national and cultural heritage. in local settings, the old yoruba newspapers were published and widely circulated particularly in the western parts of nigeria; their content was meant to be seen by its audience, not kept away for covert use, thus, there is little to no copyright breach with the digitisation and local circulation of the papers. however, these papers have evolved from https://doi.org/10.29173/iq1041 15/17 oluwaseun obasola & rukayat atinuke usman (2024) digitising old yoruba newspapers at kenneth dike library, iassist quarterly 48(2), pp. 1-17. doi: https://doi.org/10.29173/iq1041 being mere articles carrying details of the everyday significant happenings of the past to becoming subjects of national and cultural heritage. thus, opening access to them beyond the community they were created to serve, might be crossing the line of copyright and intellectual property rights. in putting this reality into perspective, the digitised collections are not only kept online but in the opensource institutional repository of the kenneth dike library. the library has the purview of granting access and access control. we are aware of the limitations this could create, such as debarring access to users outside of kdl, thus limiting the overarching purpose of accessibility. or the issue of excluding this heritage from the pool of digitised global resources or heritages. this is an emerging and rapidly evolving sector, especially in africa. hopefully, as newer and better methods emerge, holistic solutions to the complex issues in digitisation can be adequately addressed. copy-right law there was no need to seek permission from authors of articles in the newspapers or the publishers; as the library is protected by the fair use act under the copyright law (second schedule), as long as the materials digitised are used exclusively for teaching, learning, and research. further, most of the newspapers digitised are in the public domain as the date of first publication for some was over 50 years. acknowledgement the project was funded by the european research council starting grant 2020 – “yoruba print”. references alabi, s. (2003). the development of indigenous language publications. issues in nigerian media history: 19002000ad (akinfeleye r. and okoye i., eds.). lagos, nigeria: malthouse press limited. anthropology and education. philadelphia: university of pen. barber, k. (2004). bibliographical survey of sources for early yoruba language and literature studies, 1820-1970. research in african literatures, 35(1), 203-204. https://doi.org/10.1353/ral.2004.0004. bianchi, c. (2006). making online monuments more accessible through interface design. digital heritage—.(pp. 445–466). routledge britz, j., & lor, p. (2004). a moral reflection on the digitization of africa’s documentary heritage. ifla journal, 30(3), 216–223. https://doi.org/10.1177/034003520403000304 corrado, e. ., & sandy, h. (eds.). (2017). digital preservation for libraries, archives, and museums (2nd ed.). rowman and littlefield. curras, e. (1987). information as a fifth vital element and its influence on the culture of the people. journal of information science, 13(3), 27–36. grycz, c. j. (2006). digitising rare books and manuscripts. in digital heritage (pp. 33–68). routledge hampson, a. (2001). practical experiences of digitisation in the builder hybrid library project. program: electronic library and information systems, vol. 35 no. 3, pp. 263-275 https://doi.org/10.1108/eum0000000006950 https://doi.org/10.29173/iq1041 https://doi.org/10.1353/ral.2004.0004 https://doi.org/10.1177/034003520403000304 https://doi.org/10.1108/eum0000000006950 16/17 oluwaseun obasola & rukayat atinuke usman (2024) digitising old yoruba newspapers at kenneth dike library, iassist quarterly 48(2), pp. 1-17. doi: https://doi.org/10.29173/iq1041 hauswedell, t., nyhan, j., beals, m. h., terras, m., & bell, e. (2020). of global reach yet of situated contexts: an examination of the implicit and explicit selection criteria that shape digital archives of historical newspapers. archival science, 20(2), 139–165. https://doi.org/10.1007/s10502-020-09332-1 holley, r. (2003). developing a digitisation framework for your organisation (feel the fear and do it anyway). lianza, october, 1–8. https://doi.org/10.1108/02640470410570820 king, e. (2005). digitisation of newspapers at the british library. the serials librarian. 49, 1–2, 165– 181. https://doi.org/10.1300/j123v49n01_07 liu, a. (2012). the state of the digital humanities: a report and a critique. arts and humanities in higher education. 11, 1–2, 8–41. https://doi.org/10.1177/1474022211427364 ma, r. (2020). translational challenges in cross-cultural digitization ethics: the case of chinese marriage documents, 1909–1997. libri, 70(4), 269–277. https://doi.org/10.1515/libri-2020-0088 maringues, m. (2001). the nigerian press: current state, travails and prospects. in k. amuwo, d. c. bach, & y. lebeau (eds.), nigeria during the abacha years (1993-1998) (1–). ifra-nigeria. https://doi.org/10.4000/books.ifra.640 mieczkowska, s., & pryor, k. (2002). digitised newspapers at norfolk and norwich millennium library. collection building, 21(4), 155–160. https://doi.org/10.1108/01604950210447395 ogunbiyi, i. a. (2003). the search for a yoruba orthography since the 1840s: obstacles to the choice of the arabic script. sudanic africa, 14, 77-102. http://www.jstor.org/stable/25653397 olúmúyìwá, t. (2013). yoruba writing: standards and trends. journal of arts and humanities, 2(1), 40-51. https://doi.org/10.18533/journal.v2i1.50 omoloso, a. i., & abdulrauf-salau, a. (2014, may 4). indigenous language newspapers in nigeria from 1914–2013: a review. in amalgamation national conference of the department of political science & department of history and international studies, ibrahim badamasi babagida university, niger state, nigeria. omu, f. i. a. (1967). the “iwe irohin”, 1859-1867. journal of the historical society of nigeria. 4, 1, 35–44. https://www.jstor.org/stable/41971199 onwubiko, p. (2005). using newspapers to satisfy the information needs of readers at abia state university library, uturu. african journal of education and information management. 7, 2, 66– 88. pickover, m. (2009, july 1). contestations, ownership, access, and ideology: policy development challenges for the digitization of african heritage and liberation archives. 1–10. first international conference on african digital libraries and archives (icadla-1), addis ababa, ethiopia. https://wiredspace.wits.ac.za/server/api/core/bitstreams/f76d3475-e18a-4a48-9ab80fbb605a49d0/content salawu, a. (2004). the yoruba and their language newspapers: origin, nature, problems and prospects. studies of tribes and tribals, 2(2), 97-104. 2, 2, 97–104. https://doi.org/10.1080/0972639x.2004.11886508 stoker, d. (1999). should newspaper preservation be a lottery? journal of librarianship and https://doi.org/10.29173/iq1041 https://doi.org/10.1108/02640470410570820 https://doi.org/10.1177/1474022211427364 https://doi.org/10.1515/libri-2020-0088 https://doi.org/10.1108/01604950210447395 http://www.jstor.org/stable/25653397 https://doi.org/10.18533/journal.v2i1.50 https://www.jstor.org/stable/41971199 https://wiredspace.wits.ac.za/server/api/core/bitstreams/f76d3475-e18a-4a48-9ab8-0fbb605a49d0/content https://wiredspace.wits.ac.za/server/api/core/bitstreams/f76d3475-e18a-4a48-9ab8-0fbb605a49d0/content 17/17 oluwaseun obasola & rukayat atinuke usman (2024) digitising old yoruba newspapers at kenneth dike library, iassist quarterly 48(2), pp. 1-17. doi: https://doi.org/10.29173/iq1041 information science. 31(3). https://doi.org/10.1177/096100069903100301 ugah, a. (2009). strategies for preservation and increased access to newspapers in nigerian university libraries. library philosophy and practice(e-journal), 270. https://digitalcommons.unl.edu/libphilprac/270 zaagsma, g. (2022). digital history and the politics of digitization. digital scholarship in the humanities 38(2), 1–22. https://doi.org/10.1093/llc/fqac050 endnotes 1 oluwaseun obasola is the digitisation librarian at the kenneth dike library, university of ibadan, nigeria. she can be reached by email: seunobash@gmail.com 2 rukayat atinuke usman is a postgraduate student at the institute of african studies, university of ibadan, nigeria. 3 the problem of access was implicated in many sub-headings but fully addressed towards the end of the article. 4 for example, in the abstract of each newspaper volume, words like “yoruba culture” were used to describe a range of events already tagged in the yoruba language like the egungun festivals, naming ceremonies, coronations, excerpts on indigenous hairstyles, and many others. https://doi.org/10.29173/iq1041 https://doi.org/10.1177/096100069903100301 https://digitalcommons.unl.edu/libphilprac/270 https://doi.org/10.1093/llc/fqac050 bid bringing integration to data by karsten boye rasmussen ' dansk data, denmark. abstract the purpose of data archives is not only to store data materials for posterity, but also and equally important to fulfil the needs of the present users. the growing demand for studies along with a number of other factors force the data archives to change their strategy for retrieval and dissemination of studies in order to serve their customers in the best and cheapest way possible. this calls for effectiveness and new thinking in the retrieval and dissemination procedures. this paper outlines the state-of-the-art at the danish data archives (dda) by describing the format of machine-readable records and the present utilization of machine-readable documentation at the archive. a projection of the growing service rate shows the pressing need for a more effective system for direct servicing of the users, the idea being to integrate the study description, the variable documentation and the data and as the distributing medium use cd-roms, as the production costs are reasonable. the remaining problem is obtaining or financing the development costs of a suitable retrieval system. these will most probably be rather high, and as probably a lot of archives have the same need for a retrieval system, it could be a good idea for the archives to cover the development costs between them. background: machine-readable records to ensure the right perspective 1 shall shortly outline the records kept at the ddal the data are social science quantitative investigations (typically surveys). after the processing at the archive of the raw data and the (typically) paper documentation (questionnaire, coding instructions, reports etc.) each survey consists of three files: documentation of the study (study description) rectangular (flat) character data file documentation of the variables the files of the finished survey study description i shall not go into details about the study description. discussions concerning item numbers, subitem numbers and especially the introduction of new item numbers have been the subject for discussion at many earlier lassist conferences. the "standard study description scheme" has gradually been improved^ but it is still close to the original scheme agreed upon more than ten years ago*. the scheme is used at several archives'. at the dda all study descriptions are maintained in both a danish and an english version. the actual format and the many items makes the standard study description scheme a rather complex instrument. the design is specifically intended to be a machine-readable record at the study level. the study description format contains coded as well as text information. it is of importance that there is no limitation to the amount of text in the text fields. summer 1991 45 title title of study class ready-made for analysis or just deposited access there may be restrictions year the year the survey was carried out cases the number of cases in the survey nation which country or countries donor the depositor of the data institute place of employment of the depositor sponsor finance (social science research council) collector who carried out the field work abstract description in free text keywords controlled vocabulary areas of the study description the data file data stored at the dda are kept as rectangular character data files (not as system files). one record per case. interlinked datasets are either divided into separate files (like tables in a relational database) or patted into one (greatly redundant) file. the data for each observation is placed as a record in the data file. and each record holds information of the variables stored in fields (fixed columns). the data per se will not be of any use without the documentation. 020912030602031003900300566020ii061103902010 02081210101010101020020051108011021103103100 02063203030102010330062034408011051103302000 0207120306020303033005005230 10 1 104 1 103 10300 1 02071201050202010130020051308011011103306000 01044204050204040440022021208011011103101000 01071201050203010190010054308011051103103001 02051299101010040490023031405011061103203001 02072301010102010190110063208011011103302000 02043204050203030430030012208011041103202010 02042312050203121290011022108011021103203000 01034199101010999910030011105011061103201000 02051204020103010430010012408011051102303000 02061204020103020430010053408011011103303000 01073203010102030390990054408011011103103000 character data file documentation of the variables the character data file needs documentation in order to be of any use in its own right without the proper documentation the user is not able to distinguish one flat file from another. in order to exploit the information in the data file we need variable documentation. the variable documentation can be thought of as a relational table where each field describes certain characteristics of the variable. this idea is already in production in several databases. in ibm's sql database* several system tables keep track of the information stored in the database. an sql database may thus be viewed as a complete data archive, where each study is a separate table. the available tables are then stored in the systable-table and the documentation of the fields of each table (data) is stored in the syscolumns-table. 46 {assist quarterly the project of making a complete description of the fields in the variable documentation lies outside the scope of this paper, so the ust below shows only a selection of possible fields. the list shows that the documentation provided by the most widely used social science data packages (sas and spss'') does not have the necessary facilities for a complete documentation. both packages support only the first part of the field ust. name the variable name or number labe short text description place extract the data from these columns missing definition of missing data format output format / categories question the complete questionnaire text study identification of the study qid identification in original questionnaire filter reference to filtering codingspecial coding instructions processing unofficial comments check checking performed checklog generated notes on the checking fields for a variable definition the most important field for the user is the question field consisting of free text description exactly as it appeared in the questionnaire. this field provides opportunity for performing full text retrieval. neither sas nor spss are capable of storing and utilizing the complete questionnaire text. the only field for giving a text description of the variable is the variable label. the label in both packages is in practice* limited to 40 characters resulting in definitions like: label income = 'prs mon ave inc main-jb -tax dk 894-904' ; the variable label, 40 characters 40 characters is not muchi the variable income actually covers "personal average monthly payment on main job, after taxation, in danish kroner, april 1988-april 1989". this is certainly not a lot of text for a questionnaire text or a created variable, but even then this cryptic label presented above is a common result when the text is restricted to 40 characters. this is one of the reasons for the dda to use the data-archaic format provided by osiris'. the format was developed by the icpsr. the virtues of the osiris codebook format is that there are no limitations to the amount of description for each variable. there can be several lines describing the variables and the categories. there can even be references to notes. the format uses a sort of tag in the first column of each "card" yes, it is that old to identify which portion of the variable documentation is described: summer 1991 47 t label, columns, missing q question x explanation j unorficial comments f frequency c,b category text g reference to note s,e introduction m note text the osiris codebook types using the format will result in a variable description looking like the following. this is certainly not as straightforward as the syntax used by sas or spss. but normally this lay-out is not written by human beings. a machine-readable format of this complexity and rigidity is best written by machines. in the actual production of machine-readable documentation at the dda a pre-processor software is used. 48 assist quarteriy t0093 eec vote today 018400020 0100000090000009 q00930093 if there was a referendum on joining the eec today would k0093 you vote yes or no? in the danish election study, 1975 "don't know" is included in "4. don't want to answer". variable basis pre 1971: post 1971: des 1973: 157 des 1975:90 des 1977: 211 des 1979: 166 des 1981:259 1971-1 1971-2 1973 1975 1977 1979 1981 x0093 k0093 j0093 k0093 k0093 k0093 k0093 k0093 k0093 k0093 x0093 k0093 k0093 k0093 k0093 k0093 k0093 k0093 k0093 k0093 k0093 k0093 c0093 f0093 c0093 f0093 c0093 f0093 c0093 f0093 c0093 f0093 c0093 f0093 c0093 f0093 100 100 45 37 30 41 41 44 38 50 41 40 1 3 3 2 2 13 2 1 1 9 15 13 16 1 ) 9 1 2 1 wgtn= 1302 1302 533 1600 1602 3192 1500 2558 weighted: 255801. would vote yes unweighted: 280702. would vote no unweighted: 2807 weighted: 15003. would return a blank ballot paper unweighted: 150 weighted: 28104. don't want to answer unweighted: 281 weighted: 68805. don't know unweighted: 688 weighted: 20809. not ascertained unweighted: 208 weighted: 260411. the variable is not included unweighted: 2604 weighted: 3231 3542 189 297 931 237 2604 example of osiris codebook' the osiris software itself has not been used for many years at the dda, only the format lives on. but around this format a great many procedures and programs have been developed for checking the data against the documentation, for presenting the documentation, and for further utilizing the documentation. utilization of machine-readable documentation the main reason for the production of machine-readable documentation is the reuse of information. safely stored documentation can be used in a variety of ways. this chapter describes the current activities at the dda. the present status implies summer 1991 49 some inexpediencies that will be clarified in details in the next chapter. printout conversion to other formats for analysis retrieval dissemination the utilization of machine-readable documentation most of these evident virtues of machine-readable documentation goes for the study level (the study description) as well as the variable level (the codebook). printoui in compiling a complete documentation a printout of the study description in a readable format (for human beings) is placed as an introduction to the codebook. printouts of several studies form a catalogue. as the study descriptions are rather voluminous, a catalogue of study descriptions normally includes only the most important items, but the entries are heavily indexed (by persons, subjects, and keywords). however, it is expensive to produce printed catalogues, and it is difficult to keep up with the immediate obsolescence of the catalogue. to make a printout at the variable level is rather straightforward. the problem is to make the documentation as accessible as possible to the user. most of the inconvenience of the rigid osiris format is overcome by getting rid of the redundant information in the codebook'^ 50 {assist quarterly dda-0658 danish election studies, continuity file 1971-1981 var.93 eec vote today start pos. 184, missing data: = 9 or >= 9 if there was a referendum on joining the eec today would you vote yes or no? in the danish election study, 1975 "don't know" is included in "4. don't want to answer". 1971-1 1971-2 1973 1975 1977 1979 1981 1. . 45 37 30 41 41 2. 44 38 50 41 40 3. 1 3 3 2 2 4. 13 2 1 1 5. 9 15 13 16 9. 1 9 1 2 1 11. 100 100 wgtn=1302 1302 533 1600 1602 3192 1500 unweight md marg % % 2558 28 39 01. would vote yes 2807 30 43 02. would vote no 150 2 2 03. would return a blank ballot paper 281 3 4 04. don't want to answer 688 7 11 05. don't know 208 2 . 09. not ascertained 2604 28 . 11. the variable is not included 9296 100 99 weighted respondents: 11031 conversion to other packages (sas and spss) very few users if any make their analyses on delivered datasets using osiris. instead most users in denmark use spss or sas. to the user the osiris codebook produces only a printed codebook. the abihty of software programs to read system files from other packages has appeared to be quite unstable. with every new release and every new operating system platform this conversion often proves to be the tricky part. summer 1991 however, with the osiris codebook being a machine-readable variable documentation it is possible to convert the documentation file without affecting the data file. the rigid format of the codebook makes this a relatively straightforward process. presently the dda uses a micro computer program called osi-spc" for doing the conversion. the idea is to leave the data file unaltered. the same data file can be described by both osiris, sas and spss. the user then receives the osiris documentation and the data file, and with the help of the conversion program, the user is able to convert the documentation to the preferred package. and the conversion will produce a sas-program or spss-controlcards, which the user would otherwise have had to write himself. title 'danish elections, continuity file 1971-1981' . data list nle='l:\u0658\mini.dat' nxed table/ var3 8 var93 184-185 . variable labels var3 'survey id var93 'eec vote today value labels var3 1 'danish pre-election study, 1971' 2 'danish post-election study, 1971' 3 'danish election study, 1973' 4 'danish election study, 1975' 5 'danish election study, 1977' 6 'danish election study, 1979 (no. 20)' 7 'danish election study, 1979 (no. 21)' 8 'danish election study, 1981' / var93 1 'would vote yes' 2 'would vote no' 3 'would return a blank ballot paper' 4 'don t want to answer' 5 'don t know' 9 'not ascertained' 11 'the variable is not included' . save file=' l:\u0658\spssmini' . spss/pc+ generated program from osi-spc 52 lassist quarterly title 'prsetup two variables from dda-0658' ; title2 'danish elections, continuity file 1971-1981' ; libname library 'l:\u0658' ; proc format library=library ; value var3l 1 = 'danish pre-election study, 1971' 2 = 'danish post-election study, 1971' 3 = 'danish election study, 1973' 4 = 'danish election study, 1975' 5 = 'danish election study, 1977' 6 = 'danish election study, 1979 (no. 20)' 7 = 'danish election study, 1979 (no. 21)' 8 = 'danish election study, 1981' ; value v93l 1 = 'would vote yes' 2 = 'would vote no' 3 'would return a blank ballot paper' 4 = 'don"t want to answer' 5 = 'don"t know' 9 = 'not ascertained' 11 = 'the variable is not included' ; run; libname u0658 'l:\u0658' ; nlename datain 'l:\u0658\mini.dat' ; data u0658.minidat ; infile datain lrecl=185 ; attrib var3 label = 'survey id length = 3 format = var3l. attrib var93 label = 'eec vote today length = 3 format = var93l. input var3 8 var93 184-185 ; run; sas setup generated by osi-spc the development of osi-spc is pan of the archive policy. the osiris codebook format is used because of its unlimited capabilities for storing text description. with the help of conversion program(s) the needs of the users of the future can be fulfilled as well. to support new analysis packages we need only to make adjustments to the software (contrary to developing a new conversion program) to stay in pace with the development of new software packages. documentation retrieval i reckon that all archives make use of some kind of retrieval system (possibly many systems) in order to search and identify datasets of interest to the users. at the dda we have made several systems at the dataset level as well as at the variable level. the system used for searching study description excerpts (ddaguide) has the most extensive documentation and it is the summer 1991 53 only system available to outside users. ddaguide is running on a mainframe host''', and uses a dialect of the common command language (ccl, promoted by the eec). retrieval at the study level is the first step in searching for a suitable dataset: find year>1981 cases>1000 word=tv. index=election year>1981 study carried out after 1981 cases>1000 with more than 1000 cases word=tv. words beginning with "tv" index=election "election" found as a keyword ddaguide retrieval in study descriptions this request can be regarded as a construction of 4 sets, which are all combined (with logical and) in a resulting set (possibly empty). normal boolean logic (and, or, andnot) can be applied for building new sets. the result is a vector of study numbers, and then the codebooks for these studies must be searched to secure that they contain variables that come close to the user's need. dissemination of studies at present studies are distributed to the user by means of many different media: mainframe taf)es, diskettes, and electronic network as preferred by the user. compared with the situation only a few years ago a lot of the analysis is now taking place on micro computers, and consequently many users prefer the data to be delivered on diskettes. the pressing need for more effective servicing the growing demand for delivery of studies forces the dda to make the retrieval and selection as well as the practical dissemination more effective. in this chapter i shall take a look at the present technical obstacles that make the service less effective. both the retrieval and the dissemination are carried out at the dda and require a lot of manpower at the archive. about 25 pet. of the work of the archive is devoted to servicing'^ and because of the growing rale of data deliveries this number is expected to increase. because no extra funding is available the success of data deliveries will drain the potential for processing new studies to becoming available in fully documented form. it is thus a necessity to make the retrieval and dissemination procedures more effective. more effective retrieval tools the present retrieval tools suffer from the following drawbacks: 54 lassist quarteriy dispersed retrieval facilities ddaguide searches only study descriptions updating is irregular and cumbersome retrieval only available on one specific mainframe no thesaurus (actual words searched) line mode user interface drawbacks of presently used retrieval tools having several poorly integrated and sometimes individual retrieval facilities makes it more than a one-man-job to select the appropriate studies. using ddaguide will narrow the number of studies to be investigated further, but these studies will have to be scrutinized at the variable level. this means that the codebooks must be searched, but at present no system is available for searching the codebooks, therefore the most obvious choice is to ask the people at the archive responsible for the selected studies. secondly, some studies (series of studies like gallup polls etc.) have some extra machine-readable indexing of variables, but these indexes are not integrated into the codebooks and are made by a third person. this shows poor integration and a massive use of manpower, which eventually introduces the risk of performing erroneous retrieval. the solution is fu^t to integrate the index with the codebooks. the retrieval of codebooks must then be integrated with the retrieval of study descriptions. it ought to be possible to select a set of studies at the study level. furthermore the variables of these studies should be searched within the same system. superficially there is not much difference whether the unit of retrieval is a study or a variable. but this proves to be wrong, as this example illustrates'*: setl: housing set2: city 1 and 2: variables with both "housing" and "city" var: found within the same single variable study: found 2 variables within same study combination (and) of retrieved variables a retrieval system for codebooks has to have a kind of extra parameter to distinguish whether the combinations of variables take place at the variable level (both subjects found within the same variable) or at the study level (two variables covering both subjects found within the same study). updating text databases presents major problems. adding new studies or variables is no problem, the obstacle exists in replacing a text entry. this is caused by the most used algorithm where the actual text is subdivided and scattered throughout the data base. the simplest solution is to build the data base from scratch every time replacement or updating is required. this demands computing power, which in turn is becoming still less expensive. earlier retrieval data bases were typically placed on large mainframe hosts. the performance of recent microcomputers makes this kind of machine a perfect choice for the integrated retrieval system. summer 1991 unfortunately an integrated system for retrieval at both the study and the variable level is not available at the dda at present some of the design mostly in the form of wishful thinking is presented in this paper, but the actual development requires funding. we are not so stubborn as to insist on making this development. presently we have not come across a system fulfilling our demands, but we try to experiment with systems and we are looking forward to the perfect system being presented. also a lot of archives must be investigating into the same problems, so maybe a more integrated and common effort is required to reach the goal. bringing the data to the user if we assume that the user has decided on which studies he wants, the distribution of studies to the user can take place in a lot of different ways. presently the dda places the studies on the requested medium (diskettes, mainframe tapes or using electronic mail). this demands active work from the employees at the dda. but in addition to sending data to the user there exist several other distribution methods that require a more active role of the user. mainframe in network (e-mail) bulletin board service diskettes, magnetic tapes, cd-rom distribution media mainframe in network if the users are all connected to the same machine the solution is simply to place the archive files on this computer. this is the solution of campus-archives, but it is seldom efficient for archives with national coverage. however, as mainframes are being connected with networks this solution has turned out to be a big success. viewed from the icpsr, access through network (cdnet) in 1989 "accounts for almost three quarters of the icpsr data orders"'"'. the success is totally depending upon the users' access to, knowledge of and familiarity with the networking procedures. although the dda is the national data archive for denmark, in a lot of respects it would be a great mistake to compare the dda and the icpsr. in this context the big difference is that the icpsr is serving trained staff at member institutions. this staff is then handing out the data to the users, which on their side are accustomed to using the campus computer. at the dda we are serving the user directly. the users are using a lot of different mainframes and are seldom aware of nor interested in the networking techniques. bulletin board service all the users can be expected to be using some kind of micro computer. it would seem rational to make the data archive holdings available on a remote basis for pc-users. the most common form to be utilized is then running a bulletin board service (bbs) with download facilities. however, most of the users can not h>e expected to possess the necessary technical facilities (especially modem connections). as a side effect, the investigations into bbs have drawn our attention towards data compression techniques. in a recent byte article comparing data compression software'* the virtue of data compression is seen as a saver of space on the hard di.sk, but for long the greatest potential for data compression has been the dissemination of data via bbs by modem (normal modem without built-in data compression facilities). however, the data compression technique is fully apphcable to diskettes. this example shows what is saved by using the pkzip" data compression: 56 lassist quarteriy 1.696.531 bytes original data tile 231.494 bytes zlp-nie 259.868 bytes exe-nie data compression (pkzip family version) due to the limited variation of the bytes in the data file the compressed data file gains 86 percent of the space required by the original file. this means that a file that is too big for a single one of the largest diskette formats available (ibm 1 .44 megabytes) will now fit an old floppy disk (360 kilobytes). the zip-file is converted back to the original format by an unpacking program. to make sure that the user will be able to unpack the zip-file even if the user does not possess the unzip program the zip-file can be extended to a self-unpacking program (an exe-file produced by the zip2exe program). this will cost around 30.(x)0 bytes. furthermore there exists a family version (a program that is capable of running both under dos as well as os/2), so the produced exe-file can be unpacked in any of the two environments. by using the most common baud rate (24(x) baud) it would take approximately 20 minutes to download the exe-file. diskettes with data compression presently data compression combined with the mailing of diskettes seems to be the best alternative for the dissemination of a minor collection of studies to users: using difterent mainframes neither confident with the use of electronic networking nor with the use of bbs using micro computers (dos or os/2) user profile for using data compression and diskettes the data archive on a disk (cd-rom^») the real potential of a data archive lies in the user's opportunity to perform comparative analyses on more than a single study. this means that most secondary analyses will need several studies, and that a single diskette will prove insufficient a radical solution would be to put the complete archive on a disk which could be distributed. this has become relevant with the marketing of optical disks. also worm-disks exist in a lot of different and incompatible formats. but with the growing number of cd-rom players the cd-rom medium seem to be the standardized and perfect solution for bringing out the archive to personal computers. summer 1991 57 the cd-rom is available for both os/2 and dos" cd-rom holds about 640 megabytes of memory" the number of cd-rom applications is rising the number of cd-rom players is rising the media is read-only cd-rom figures normally read-only media are thought of as some kind of second class media, but read-only is the perfect attribute for the stable archive data. in this chapter i shall present some storage considerations as well as some different viewpoints on the implementation of a cd-rom solution. 600 megabytes is still a limit user benefits and demands depositors' reactions archive staff reactions cost of production implementation of cd-rom is 600 megabytes enough? lets take a look at the size of data stored at the dda. in the following table the unit is files archived at the dda. but each file is in addition weighted with the number of bytes that it occupies. totally we are talking about 5398 files occupying 2953 megabytes or 3 gigabytes. the main archive is placed on mainframe tapes. each file belongs lo one of the categories: finished, process (being processed, i.e. an intermediary file that is stored for extended security) or original (the file received). likewise each file also belongs to one of the package-categones (osiris, spss or sas) and finally the file contains either data or documentation (dict/dicb/misc)". of these files only a small fraction of the 3 gigabytes are of interest to the cd-rom project namely studies finished and stored in an osiris version: at present 586 data files (341 megabytes) together with the documentation files (491 dicbs (dicuonary-codebooks) totalling 94 megabytes and 546 dicts totalling 6 megabytes^). at present these approximately 5(x) finished studies will occupy around 450 megabytes. a cd-rom will hold from 540 to 640 megabytes. {assist quarteriy siaius f1 ucess ri ii!>ueu kji giiiai all 1 n sum n sum n sum n sum osi cdbk 21 4mb 20 4mb 41 9 mb data 480 412 mb 586 341mb 1412 1417 mb 2478 2171 mb dice 21 12 mb 491 94 mb 125 41mb 637 148 mb dict 302 3 mb 546 6 mb 227 6 mb 1075 16mb div 448 158 mb 40 11mb 80 48 mb 568 218 mb all 1272 590 mb 1663 455 mb 1864 1519 mb 4799 2565 mb sas div 2 mb 9 mb 1 mb 12 mb all 2 mb 9 mb 1 mb 12 mb spss ctrl 30 2mb 90 9 mb 104 9 mb 224 21mb data 13 35 mb 83 133 mb 146 184 mb 242 353 mb div 7 mb 5 mb 109 11mb 121 12 mb all 50 38 mb 178 142 mb 359 206 mb 587 387 mb all cdbk 21 4mb 20 4mb 41 5 mb ctrl 30 2mb 90 9 mb 104 9 mb 224 21mb data 493 447 mb 669 475 mb 1558 1602 mb 2720 2525 mb dice 21 12 mb 491 94 mb 125 41mb 637 148 mb dict 302 3 mb 546 6 mb 227 6 mb 1075 16mb div 457 158 mb 54 12 mb 190 60 mb 701 231mb all 1324 629 mb 1850 598 mb 2224 1725 mb 5398 2953 mb at least 100 megabytes of the space on a cd-rom would be left free. and with compression techniques the required space would be less than 100 megabytes. so a cd-rom disk could hold approximately 2000 surveys and still have plenty of free space. bringing retrievable documentation to the user for each study the documentation should be available to the user as well. but the user and the archive staff need a retrieval system integrating the documentation at the study level as well as at the variable level as mentioned above. the retrieval demand is the reason for setting aside some free space on the cd-rom for the distribution of retrieval software, indexed files, and other assisting materials. summer 1991 59 integration of study and variable level documentation subsetting of variables conversion to other packages logging of search criteria inclusion of search criteria human interface user needs for integrated retrieval even though a cd-rom caii hold a lot of information there is no guarantee that the user has similar abundance of space available on his hard disk. often the user will need only selected variables from a study. in addition to this type of selection, the user should have the option of transferring the background variables. this calls for the possibility to mark off the background variables in the documentation. apart from hardware limitations there may be some software limitations" too. a lot of patience is required when building datasets with a great number of variables on a pc^. retrievals may become quite complicated, so it is a necessity that the user is given the possibility of tracking down the searches and combinations. a log will provide documentation of the actual reuieval for documentation. similarly a session should be able to start off where a previous session was left. the line-mode ccl search syntax is developed for dumb terminals. but the increased use of micro computers has accustomed users to more intuitive search systems. many search systems now employ pull-down menus as well as windowing systems". the graphic user interface standardized in ibm's cua^ uses radio buttons, push buttons, check boxes, list boxes etc. as well. at present i have no knowledge of the cua standard used in an actual implementation of retrieval software, but probably it has already been marketed. as most users are conservative about learning new computer languages for analysis it should be possible to convert the documentation to the user's preferred package as exemplified earlier in this paper. an integration of a (new) analysis package could be counter productive for the user and would definitely involve some further costs. no answer to questions like: "when was the proportion of social democrats among men higher than among women?" analysis can be carried out in other processes (os/2) with the analysis package preferred by the user analysis package separated from retrieval depositors reaction the studies stored at the dda are not all directly available to the user. although the data are placed at the archive, the depositor still has the formal rights to the material, and the depositor has the possibility of assigning access categories to the dataset. the access categories are assigned to the data only, the documentation is always available without special permits. lassist quarterly no access restrictions whatsoever no access restrictions to scientific use no access restrictions, but consultation with access directing authority is strongly advised no publication without written permission from access directing authority no use of data without written permission from access directing authority available only after special arrangement with access directing authority; generally not yet available access restrictions for dda studies one solution to this problem would be to incorporate only studies without access restrictions (the first three categories) in the bid-projecu as this would be a dramatic solution it is not advisable. another solution would be to work for a transfer of all studies to the category of free access. but this work has already been carried out in so far that most studies after a while are transferred to a less restrictive category. but even though the data are freely available, the depositor (the access directing authority) is presently being informed of whom the datasets are disseminated to. all in all this implies some kind of password, so that it will be possible for the user only to access files that he has been positively assigned to. the password ought to be distinct for every combination of user and dataset. but user information can not be placed on the cd-rom, and the password would then have to be distinct only for the dataset as dataset protection passwords are not a standard feature of the operation systems for micro computers, the passwords would be implemented in the form of a key for the encryption of the data file. a user would then be able to pass-on the key to other users. this shows that the security is not perfect, but then it does not have to be perfect. the same thing is hapf)ening now, when a few users are actually distributing the data they have received from the dda. this is in contradiction with the agreement between the archive and the user and therefore illegal. but unless the analysis system is an integrated part of the retrieval system the spreading of data can not be prevented. passwords for combination of user and study encryption of study delay in data processing serious delay in analysis obstacles of access categories but every restriction introduces drawbacks for the user. first of all the data would have to be both decompressed and decrypted unless a program would be able to do this in a single pass. secondly and most seriously the user would have to contact the archive to receive permission. this would delay the actual analysis by several days for the studies placed in the most restricted categories that requires the interaction of the depositor. summer 1991 as proven above it is impossible to totally prevent the "pirating" of data. and in my opinion it should be legalized and encouraged. the data and documentation of social science research which in denmark is most often funded by the state should be regarded as public property. any use of machine-readable material should imply the same consideration as the use of otlier sources of material. this means that all sources should be quoted, and at the same time the original investigators given credit. "piracy" encouraged sources quoted original investigators given credit the future of free information archive staff reaction the bid system is intended for the user. but the system should be used at the archive as well. at the archive it would be possible to maintain a more updated version in order to give access to the newest studies. it must be foreseen that some staff work would be transferred into giving advice on the use of the bid cd-rom system. but a major part of the present service work at the archive would disappear and leave more resources for bringing up studies to the highest documentation level. cost ofproduction the production costs of cd-roms have gone down very fast over the last year. recent announcements mentions prices as low as 30.000 dkr^' for the total production of 300 cd-roms. the cost would be expected to be even less in the us. but recently a figure of use 10.000 (70.000 dkr) for the complete process has been mentioned'". data preparation retrieval software pre-mastering mastering pressing of disks licence to software cost elements of cd-rom production the difference in pricing is caused by the exclusion or inclusion of some steps in the process. the low costs are obtained if you simply have some files and want them available on a cd-rom. the receiving company will then do the mastering and pressing of the disks. so this is comparable to the actual printing costs. 62 (assist quarteriy if all data and documentation are available the expensive process will be to develop or apply retrieval software. ready-made retrieval tools for cd-rom production are available. for companies using computers the software can be acquired, indexes constructed, and the system tested locally. the technical requirements are a medium-sized pc, lots of disk space (maybe optical) and some sort of medium for copying the laige quantities of data lo the mastering company (tape). do not forget that software is under the law of copyright. this means that although you can locally set up a nice system using the software, you are not allowed to put the retrieval software on the cd-rom. a licence or royalty fee is required in order to distribute the retrieval software to the users. conclusion a lot of different companies and software packages are available to chose among when deciding on the retrieval software. but the "plastic"-cd-software 1 have seen have all been limited to typical library systems. they could only handle the retrieval from one level of information (eg. bookcards) and not the hierarchy of studies and variables retrieval on two levels (study and variable) start of other external processes: conversion to data analysis packages decrypting the data file decompressing the data file subsetting the data file storage of search profiles the need for open data archive software is the price of integration this need for retrieval software that can be further developed is arisen from the idea to integrate the study description, the variable documentation and the data. the conclusion of this paper is that the total integration of these parts makes it even more pressing to make data, documentation and the retrieval software freely available without bureaucratic hindrance. this paper has been an advertisement for a retrieval instrument open for further development and not the introduction of the product from a concluded development. but the need must be very similar at many data archives, and there should be a potential for covering the development costs between the archives. footnotes: ' presented at the "lassist 90" conference held may 30 june 2 at the radisson hotel, poughkeepsie, new york, u.sa. i the dda documentation standards were earlier presented in my paper "data on data", in "suigr89", proceedings of the sas european users group international conference, sas gmbh. the bid project has in an earlier version been described in "beskrivelsens integration med data" in dda-nyl 48 (in danish). \ "standard study description scheme". latest update 1988, available from the dda (doc00364)". *. "study description guide and scheme" by per nielsen, copenhagen, dda, 1975. summer 1991 ^ the use of the standard study description and similar vehicles among european archives described in: "from localization to cataloguing of data sets". (workbook of the first cessda expert seminar). ed. astrid bogh lauritzen, 1987, danish data archives, odense, denmark. *. "ibm os/2 ee 1.1. database manager programming guide and reference", 1988, ibm. (appendix c). the database manager is part of the extended edition for os/2. \ "sas language guide for personal computers. version 6 edition". 1985, sas institute inc., gary, nc, usa. "spss/ pc+ version 2". 1988, spss inc., il, usa. the mainframe versions of sas and spss do not differ significantly with respect to documentation facilities. ^ "spss-x user's guide" 3rd edition, spss inc., chicago. spss-x supports up to 120 characters, spss-pc+ (now version 3) places the limit at 60 characters, but most procedures print out only the first 40 characters. '. "osiris iii. volume 1, system and program description" 1973, university of michigan, usa. (appendix d describes the osiris datasei). '°. this codebook is untypical for studies at the dda. first of all it is translated into english. the codebooks at the dda are mostly in danish. secondly the study actually consist of 7 studies tjiat have been merged, this accounts for the tabulations showing responsepercentages at different points of time. the study is now available through icpsr (dda-0658 "danish election studies, continuity file 1971-1981", lcpsr-8946). ". the latest versions of the catalogue of holdings at the dda are: "danish data guide 1986" and "danish data guide update 1988". they are both in english. '^. the reason for the atypical layout is explained in note 9. ". osi-spc runs under dos and os/2 and is available on application to the dda. the program is capable of subsetting variables from an osiris documentation. '*. the ddaguide retrieval database runs under ibm vm/cms. '^ "dda annual report 1988", in dda-nyt49 (in danish). ". i shall not mention the standard retrieval problems: how to ensure that the employed retrieval terms actually cover the subject searched for; and how at the same time to obtain both a high level of precision and a high level of recall. these aspects are covered in "information retrieval experiment", karen s. jones (ed.), 1981. ". "icpsr annual report 1988-1989". the ordering of datasets is free but the use of the cdnet database search is a charged-for service. ". "saving space" by steven j. vaughan-nichols, byte march 1990. ". pkzip (pkware inc.), ver. 1.01 dos and os/2 family mode. the latest and much faster version is 1.10 for dos only. ^. the cd-rom "bible" is: "cd rom. the new papyrus", microsoft press, 1986. in this collection of early papers leonard laub's "what is cd rom?" is recommended as an introduction to the media. ^'. ibm has just introduced scsi-interface in their ps/2 family and also marketed a cd-rom player supporting this format ". the cd-rom maximum capacity is from 540 megabytes to 640 megabytes depending on the software used for processing and the software (and hardware) used for reading. ". the study descnptions are left out of this calculation, as their size is relatively unimportant lassist quarteriy ^. a few of the finished studies do not include a codebook, only the dictionary (variable location and label etc.) ". the spss pc+ data list can not read more than 200 variables nor read ascii files with a record length of more than 1024 bytes. ("spss/pc+ version 2". 1988, spss inc., il, usa. page c-38). in practice pc sas (dos) has limitations to the number of variables too. ". both sas and spss are just now being released in os/2 versions that do not have the memory problems of the pcdos version. thus software limitations concerning the number of variables will disappear as the operating systems demand more hardware. ". the ready-to-use software guides from nonon and microsoft also have the potential for setting up new user defined retrieval systems. wordcruncher from electronic text corporation is especially designed to search great quantities of text information. ^. "saa common user access advanced interface design guide", ibm, 1989 (sc26-4582-0). ". "compact data nyt", april 1990 (newsletter in danish). '". "do-it-yourself cd-roms" by wayne rash jr., byte may 1990 summer 1991 65 iassist quarterly fall 2009 37 the 37th international association for social science information services and technology (iassist) annual conference will be hosted by simon fraser university and university of british columbia and will be held in vancouver, canada, may 31 june 3, 2011. the theme of this year's conference is data science professionals: a global community of sharing. social science benefits from professional practices that enable sharing of data, information, and knowledge with a global community. this theme is intended to stimulate discussions about ways in which sharing data, information, and knowledge can contribute to research and to professional practices that enable scientific progress. submissions are encouraged that offer improvements for creating, documenting, submitting, describing, disseminating, and preserving scientific research data. we seek submissions on the theme outlined above, and encourage conference participants to propose papers and sessions that would be of interest to themselves and other attendees. below is a sample of possible topics that may be considered: * open data and the development of knowledge communities data sharing, access and management in the future identifying and reducing barriers to data sharingissues of confidentiality in sharing sharing professional data science skills, knowledge, & techniques within &across discipline citation of research data and persistent identifiers metadata facilitating data sharing emerging research infrastructures and data sharing new data partnerships in knowledge communities sharing resources and data through social networks identifying user needs and customizing data services to meet the needs the evolving data librarian profession data science practices that support global use and understanding of research data open (linked) data and digital repositories preservation for sharing, recovering data for contemporary use proposals on other topics related to the conference theme will be considered too. papers will be selected from a wide range of subjects to ensure a broad balance of topics. the program committee welcomes proposals for: individual presentations (typically 15-20 minutes) sessions, which could take a variety of formats (e.g. a set of three or four presentations, a discussion panel, a discussion with the audience, etc.) posters/demonstrations for the poster session workshops (pre-conference workshops that blend lecture and hands-on instruction). [note: a separate call for workshops is forthcoming.] iassist 2011 call for papers 38 iassist quarterly fall 2009 iassist 2011 again this year we would be interested in receiving submissions for presentations in formats successfully introduced in last year's conference, in particula pecha kucha (a presentation of 20 slides shown for 20 seconds each, with a heavy emphasis on visual content). round table discussions (as these are likely to have limited spaces, an explanation of how the discussion will be shared with the wider group should form part of the proposal). session formats are not limited to these ideas and session organizers are welcome to suggest other formats. proposals for complete sessions should list the organizer or moderator and possible participants; the session organizer will be responsible for securing both session participants and a chair. all submissions should include the proposed title and an abstract no longer than 200 words. longer abstracts will be returned to be shortened before being considered. please note that all presenters are required to register and pay the registration fee for the conference; registration for individual days will be available. a web form for submission of proposals will be available on the conference web site on october 18, 2010. deadline for submission: november 29, 2010. notification of acceptance: february 3, 2011. please note that the conference program committee may not be able to accept all proposals. conference papers will be considered for articles in the iassist quarterly. this applies to individual papers as well as selections of papers from sessions that could form special issues of the iassist quarterly. for more information about the conference, including travel and accommodation, see the conference web site at: http://www.rdl.sfu.ca/iassist/ online conference registration is scheduled to open in early february, 2011. make plans to come to vancouver for iassist 2011 : 31 may 3 june 2011! . questions may be sent to the program planning co-chairs, bob downs, ernie boyko and tuomas j. alateräi at iassist2011@gmail.com 1/13 kallas, ioannis and kondyli, dimitra (2022). a tool to promote research planning and conceptualization: sodanet research infrastructure’s scientific dictionary of social terms, iassist quarterly 46(1), pp. 1-13. doi: https://doi.org/10.29173/iq1021 a tool to promote research planning and conceptualization: sodanet research infrastructure’s scientific dictionary of social terms ioannis kallas1 & dimitra kondyli2 abstract this article examines the contribution of sodanet research infrastructure’s scientific dictionary of social terms to empirical social research3. the article records the dictionary functional specifications in regarding to terms, definitions and bibliographic records and analyzes the management issues in user access in relation to the basic functions (search, import, modification and deletion of digital content). in addition, the functions of the dictionary as a research planning tool are analyzed (providing opportunities to search for scientific information necessary to design a new research), conceptualization (providing access to the different meanings of a term through the different definitions given) and scientific documentation. finally, the function of the dictionary as an element of a research infrastructure is evaluated. keywords scientific dictionary, research infrastructure, documentation, research design tool, conceptualization. 1. introduction the research infrastructure of sodanet, a member of cessda, has developed a series of functions aimed at serving the local and international research community based on oia4, as the majority of cessda data archives, meaning interoperability standards to allow access to digital resources for elearning, open data and other services. among them, the scientific dictionary of social terms was designed to serve the conceptualizing, designing and managing of research that can be searched through the sodanet portal5. as a scientific dictionary of social terms, it is based on the scientific discourse of the social sciences. therefore, according to foucault’s discourse theory (1987), the dictionary should meet the following criteria: 1. refer to terms and concepts that are constructed and used in the context of particular social science practices. 2. its terms and concepts should be constructed and used in scientific decisions, such as research cases, laws or scientific regularities and scientific theories. 3. its development should be the work of scientists who have the respective competence to ensure the required validity of its content. the purpose of the dictionary is the terminological and conceptual support of social research and especially empirical research, which organizes the production and analysis of empirical evidence based on the coding of evidence of experience, with the help of theory. a significant difficulty encountered in empirical social research is the connection of sociological theory with the empirical basis that ensures its empirical control (schnell et al, 2014, p. 13). this difficulty can be addressed based on the formulation of appropriate research questions and research hypotheses, in tandem with the precise conceptualization of the mentioned objects and phenomena to which the analysis is directed, in order to enable the appropriate codification of the collected information. a scientific dictionary of terms can substantially facilitate this process, especially if it provides access not only to the appropriate terms and their definitions, but also to the bibliography that theoretically substantiates them, as well as to the empirical research that uses them. https://doi.org/10.29173/iq1021 2/13 kallas, ioannis and kondyli, dimitra (2022). a tool to promote research planning and conceptualization: sodanet research infrastructure’s scientific dictionary of social terms, iassist quarterly 46(1), pp. 1-13. doi: https://doi.org/10.29173/iq1021 since, according to althusser (1978), modern scientific research is organized as theoretical production, its development is firmly based on access to both the raw material (empirical data and evidence), as well as the means of production: mainly theoretical tools, such as concepts, theories, and methods, that are already available. access to both existing raw material and means of production is enhanced drastically by the development of documentation infrastructure. scientific research, as a collective endeavor, is supported by written communication, which is organized with the help of a global system of scientific publications, but also a system organizing access to and management of scientific texts. with the development of the internet, both access to information and communication have changed dramatically, as scientific texts change from print to digital and their production and management is mechanically supported with the help of information infrastructure. this development helps make scientific activity even more collective. the proposed dictionary was designed to function as an it application, and in fact as a sub-element of a research infrastructure, the sodanet6 infrastructure, which supports the management of empirical social research and its data. as an infrastructure element, the dictionary has the following features: 1. it is developed as a computer application and therefore its operation is mechanically supported 2. it is developed as a collective product through strict procedures of organizing collective work, which ensure both its continuity and its validity. 3. it develops as a hypertext. this means that: a. it is based on a combination of digital text and data that serve as nodes for its management. b. the digital content that is produced is dynamic and evolves with the help provided by the correlations between the terms, which are constantly created by the authors of the terms enriching and clarifying their meaning. c. it is open to additions and new correlations. this dictionary, in addition to being a tool for conceptual documentation, aspires to become a tool for conceptualizing empirical research that supports the strict definition and description of objects and phenomena mentioned in social research, as well as a research planning tool. 2. the functional specifications of the scientific dictionary of social terms the dictionary is dynamic and develops gradually over time through the work of many different researchers. it is constantly supplemented with new terms and definitions introduced by different researchers, certified as entry writers. this way of development allows not only the continuous enrichment but also the evolution of the terms, that is, the expansion of their meaning through the formulation of new definitions. thus, the meanings of a term increase and change over time. this is achieved thanks to the special design of the dictionary based on two elements: first, the multilevel organization of its content, and second, the strict organization and management of the access of those who contribute to its development, thus ensuring its scientific validity. the information provided by the dictionary is organized on three levels: at the level of terms, at the level of definitions and at the level of bibliographic records. the three levels form a hierarchy, so one can insert, modify or delete the information at one level while the upper levels remain unchanged. at the terms level, the dictionary provides for the introduction of scientific terms related to social research. the following information is provided for each term: the term, some comments on the term, the type of term, and the relevant terms. more specifically: https://doi.org/10.29173/iq1021 3/13 kallas, ioannis and kondyli, dimitra (2022). a tool to promote research planning and conceptualization: sodanet research infrastructure’s scientific dictionary of social terms, iassist quarterly 46(1), pp. 1-13. doi: https://doi.org/10.29173/iq1021 1. the scientific term is provided, which must be mentioned in both english and greek. the pair of greek and english terms, which must be unique, is considered a distinct term. this is because sometimes the same term in english corresponds to more terms in greek or vice versa. it is permissible to formulate new terms that have not yet been recognized in the scientific community, as long as they appear at least once in the bibliography or in some research. 2. the dictionary provides for the introduction of term-level comments by the various entry authors. comments are required if there are alternative terms, either in greek or in english, that could be used instead of the suggested term. for example, the term “segmentation” is found in the greek bibliography both as “κατάτμηση” and as “τμηματοποίηση”. the purpose of the dictionary is to standardize the terminology by choosing one of the equivalent terms as the basic term. the other alternative terms, although not used, are mentioned in the comments of the term. 3. the dictionary distinguishes terms into two main categories: a. terms relating to the different objects or phenomena to which the observation of social empirical research is directed. the terms in this case are distinguished into phenomena, objects, characteristics, relations between objects. b. terms concerning methodology and epistemology. these terms do not describe phenomena or objects of reality to which observation is directed, but methods, processes or concepts that support recording or analysis and therefore refer to epistemology or methodology (terms such as sampling, conceptual analysis, structure, system, etc.). the categorization of terms can be more detailed and can be based on one of the available thematic thesauruses (e.g. elsst7). 4. for the same term it is allowed to formulate many different definitions, produced in the context of different theories and theoretical approaches and introduced by different entry authors. 5. for each term the introduction and management of relevant terms is supported, while allowing navigation from term to term. the relevant terms are explicitly inserted either by the entry author proposing the term for the first time, or by the content manager. the following information is provided at definition level: the text of the definition, some comments on the definition, the entry author, and references in the bibliography are provided for each definition. the following also apply: 1. the scientific term must be constructed in the context of scientific decisions. to ensure this condition, each definition must be substantiated by at least one reference to the bibliography. it is not allowed to insert it in the dictionary without this documentation. of course, it is possible to have more references for the same definition. 2. each definition is recorded in the dictionary by a certified scholar who is recognized as the author of the entry and who is not necessarily the author of the definition. the author of the entry refers to the author of the definition directly or indirectly by referring to at least one bibliographic reference. a definition is inserted with quotation marks only if it is quoted. in this case the reference to the bibliography must definitely include the page of the mentioned text. each definition corresponds to a single scientific term, to which it assigns a specific identity. the formulation of a definition is made separately in greek and english. these formulations are recorded as distinct definitions. but because one is a translation of the other, they have the same definition code (and the codes are given automatically by the system). 3. each definition may be supplemented by additional clarifications or information in the form of a comment if this is deemed useful by the entry author. the comment may include a more https://doi.org/10.29173/iq1021 4/13 kallas, ioannis and kondyli, dimitra (2022). a tool to promote research planning and conceptualization: sodanet research infrastructure’s scientific dictionary of social terms, iassist quarterly 46(1), pp. 1-13. doi: https://doi.org/10.29173/iq1021 detailed description, explanations or even clarifications if this is deemed necessary. additional references to the bibliography are given in the comments. more than one comment for the same definition is allowed even by different entry authors. 4. each definition is characterized as either a nominal or a functional definition. a definition is characterized as nominal when it gives meaning to a term, i.e., without specifying how this definition can be confirmed empirically (kallas, 2015, p. 190). a definition is characterized as functional when it does not simply give meaning to a term, but when it corresponds precisely to a process of observation, measurement or processing, i.e., to a way of empirically confirming this meaning (kallas, 2015, p. 193; babbie, 2011, p.781). for the documentation of the nominal definitions, we refer to theory, while for the documentation of the functional ones, we refer to empirical research. as an example of a functional definition, an unemployed person is defined by eurostat, according to the guidelines of the international labour organization, as: someone aged 15 to 74 (in italy, spain, the united kingdom, iceland, norway: 16 to 74 years); without work during the reference week; available to start work within the next two weeks (or has already found a job to start within the next three months); actively having sought employment at some time during the last four weeks. visual representations of the aforementioned functional definition concerning unemployed person, available via sodanet’ dictionary follow below (figures 1 and 2). figure 1: english definition https://doi.org/10.29173/iq1021 https://www.sodanet.gr/metadata/socialterms/search?social-term-inputtext=unemployed+person&social-term-inputstate=definition#terms-search-results 5/13 kallas, ioannis and kondyli, dimitra (2022). a tool to promote research planning and conceptualization: sodanet research infrastructure’s scientific dictionary of social terms, iassist quarterly 46(1), pp. 1-13. doi: https://doi.org/10.29173/iq1021 figure 2: definition translated in greek 5. at the level of bibliographic records, the following information is provided: since the comments and definitions, in order to be scientific, refer to the bibliography, each reference is documented with the corresponding bibliographic record that meets the following specifications: 1. bibliographic records refer both to scientific publications, i.e., books and articles that have been published and circulated in printed or digital form, and to texts documenting empirical research, such as research reports, working papers or some internal texts of repositories (in the specific case of the sodanet infrastructure), which are gray bibliography available in digital format with a unique identifier (doi) and stored in repositories. 2. the entry of bibliographic records follows the apa standard8. 3. bibliographic records can be continuously enriched even without adding new references. however, the enrichment is achieved mainly either through the introduction and modification of terms and definitions by the authors of the entries, or by the addition of new comments to a definition, even by other entry authors. 3. managing dictionary access to ensure its validity, a scientific dictionary is developed by scientists. to the extent that this dictionary is meant to be the work of many authors and is constantly open to input from new authors including beginning researchers as well, it must be able to effectively manage their access to it. access management is not just about controlling who has rights, but also what kind of rights they have to the basic functions of searching, inserting, modifying and deleting digital content. dictionary access is divided into the following levels: 1. the level of the ordinary user 2. the level of the entry authors 3. the level of the content managers of the repositories maintained by the operators of the sodanet network9 4. the level of the sodanet infrastructure administrator access to the dictionary as a simple user is free to anyone interested. they use only the dictionary application, having only the right to search. they cannot modify or delete the contents of the https://doi.org/10.29173/iq1021 6/13 kallas, ioannis and kondyli, dimitra (2022). a tool to promote research planning and conceptualization: sodanet research infrastructure’s scientific dictionary of social terms, iassist quarterly 46(1), pp. 1-13. doi: https://doi.org/10.29173/iq1021 dictionary or insert new content. ordinary users are allowed to search terms from the dictionary in both greek and english. the search for terms can be done in the following ways: 1. by term. in this case, the exact term in greek or english must be entered for search 2. by keyword. in this case the term must be entered in greek or english as a keyword, i.e., without its wording being accurate and complete. one part of the word is enough. in this case all the terms that meet the criterion show up 3. by definition. in this case the term must be entered as a word in greek or english. in this case the search for the word is not done in the terms but in the definitions, and as a result, all the terms whose definitions contain the keyword show up. such a search makes it easier to find terms related to what we are looking for, even if they have not been explicitly identified as relevant terms by the author of the corresponding entry. many different scholars are allowed to introduce content into the dictionary once they have been certified as entry authors, thus ensuring the scientific validity of the dictionary. therefore, the author of an entry is responsible for any new term or definition that is introduced, which suggests either a new term and at least one definition, or a definition of an existing term. entry authors are not necessarily the creators of definitions. thus, entry authors are responsible for the wording of a definition, without, however, necessarily having created the definitions they introduce. the introduction of definitions includes at least: a) the introduction of the term if it does not already exist, or the selection of an existing one, b) the introduction of at least one definition for the specific term, c) the possibility of inserting a comment on the specific definition if the author of the entry deems it useful, d) the introduction of bibliographic records to which the definition or comment refers, if they have not already been entered; and e) the introduction of at least one reference to bibliographic records. entry authors are responsible for the definitions they suggest. therefore, the modification of the definitions and the data related to them (reports, comments, etc.) is done only by the authors of the definitions. this is ensured by the finalization of the definition by the content administrator who certified the entry author. once the definition is finalized, modification is impossible. if the entry author wants to make a modification, they must ask the content administrator to remove the finalization. the digital content managed by the sodanet infrastructure is distributed across many distinct repositories. specific research bodies that make up the sodanet network are responsible for the maintenance of the contents of each repository. each body appoints a content manager of its own repository, who is responsible for the digital content of that repository. the content manager may be assisted by a scientific committee. the content manager has all the responsibilities of the entry author and in addition the following responsibilities: certification of new entry authors, as well as revocation of certification; finalization of a definition, as well as revocation of the finalization when the entry author wants to make a modification; deletion of a definition after informing the entry author, as well as modifying the elements of a term. only the administrator of research infrastructure has the right to delete terms from the dictionary. the overall management of the infrastructure is the responsibility of the infrastructure manager, who has the technical ability to intervene in both the applications and the content of the infrastructure. however, their intervention in the contents of the repositories is institutionally prohibited without the permission of the competent body. regarding the management of the dictionary, the infrastructure https://doi.org/10.29173/iq1021 7/13 kallas, ioannis and kondyli, dimitra (2022). a tool to promote research planning and conceptualization: sodanet research infrastructure’s scientific dictionary of social terms, iassist quarterly 46(1), pp. 1-13. doi: https://doi.org/10.29173/iq1021 manager has the responsibilities of certifying the content managers at the request of the competent body and deleting terms for all the repositories. deleting a term follows a standard interdependent procedure that consists of deleting references for each term definition, term comments, definitions, and finally the term itself. 4. the dictionary as a tool of scientific documentation this dictionary is a tool for scientific documentation for two reasons: first, because it supports the documentation of new scientific terms, and second, because it provides access to information on scientific terms and their definitions. more specifically: 1. it provides access to social science terms as well as corresponding definitions. the content of the definitions is documented and supplemented with comments and references in the bibliography, which are constantly enriched. the documentation should even include gray bibliography and mainly references to empirical research reports where required. 2. access to search its digital content is open to all without restrictions. however, the ability to insert, modify, or delete digital content is controlled to ensure validity. 3. the function of the dictionary is mechanically supported, thus ensuring increased potential for searching and navigating its digital content, as well as the possibility of maintaining it for a long time. for example, it is possible to search for relevant terms that are not directly stated as such, but that are indirectly related to others as long as they are referred to, in their definitions. a thematic classification and search of terms is also possible with the help of a thesaurus (e.g. elsst) (see above). 4. it is set up collectively under the supervision of a network of universities and research organizations, the sodanet network. this choice was based on the view that, because science is a collective endeavor, the development of a dictionary of scientific terms and definitions cannot be the work of individual scientists, but of the scientific community collectively. thus, the operation of this dictionary supports the potential participation of the entire scientific community in the production of terms and definitions. the development of the digital content of the dictionary is implemented gradually, through the action of many independent researchers and research teams, which are constantly producing new terms and mainly new definitions in their current scientific activity. this process is supervised by the sodanet infrastructure repositories. the sodanet network oversees the development and operation of the dictionary by one or more network operators wishing to undertake this task. each body appoints a content manager, who is responsible for overseeing content. the content administrator then certifies those scholars who want to undertake the introduction of terms and definitions as entry authors. 5. the dictionary, as a result of both its collective structure and its mechanical management, is constantly evolving, expanding and improving its digital content through the constant addition of new terms, definitions, comments and references, and through the contribution of many different scholars. this dictionary was designed to meet the needs of the greek-speaking scientific community. however, in order to meet this goal with scientific competence, it must provide its documentation in both greek and english. this is necessary since science, as an internationalized practice, is based on the ability of all scientists to communicate regardless of nationality. in order to achieve this, it is necessary to have a language that will function in practice as an international scientific language, as a lingua franca. the english language is now recognized as such. the terms are mandatory in both greek and english because the dictionary aims to contribute to the development of a bilingual terminology. their definitions and comments are mandatory in greek. in many cases the definitions are found in the english bibliography. in that case they are introduced in english and translated into greek. however, https://doi.org/10.29173/iq1021 8/13 kallas, ioannis and kondyli, dimitra (2022). a tool to promote research planning and conceptualization: sodanet research infrastructure’s scientific dictionary of social terms, iassist quarterly 46(1), pp. 1-13. doi: https://doi.org/10.29173/iq1021 definitions are not introduced in other languages, which means that definitions originating in other languages must be translated into english or greek before being entered. bibliographic references refer to a specific item (book, publication, gray bibliography) and are formulated in the reference language of the item. it ensures the validity of its content, ensuring that the introduction and modification of its digital content is done only by those who have the required scientific competence and taking into account certain restrictions. the introduction of definitions must be accompanied by the corresponding bibliographic reference. the introduction of new terms and definitions as well as their modification is allowed only to scientists approved by the bodies responsible for the development of the dictionary. the modification of the terms, definitions and references in the bibliography after their finalization is communicated and approved by the content managers of the sodanet network. 5. the dictionary as a tool for conceptualizing empirical research empirical research is basically composed of three categories of research processes, which concern conceptualization/codification, data production and analysis. conceptualization is defined as the mental process through which vague and inaccurate concepts take on a more specific and rigid nature (babbie, 2011, p.778). in empirical research, the concept refers mainly to the definition but also to the strict theoretical description of the mentioned objects and phenomena. conceptualization is necessary for the codification of the indications of reality. conceptualization is based on the search for appropriate theories and hypotheses and the selection of appropriate scientific terms, usually after a systematic review of the literature. conceptualization is based on both theory and conceptual analysis. in any case, however, it is based on nominal or functional definitions of the terms it uses. access to well-documented scientific terms and definitions is important to support conceptualization and codification in both quantitative (kallas, 2015, p.196-202; schnell et al, 2014, p. 412) and qualitative research (braun and clarke, 2012, p. 57, willig, 2015, p. 161-162, tsiolis, 2014, p. 107). it is therefore obvious that a dictionary like the one proposed can be a tool of conceptualization. the conceptualization is based on scientific concepts and not everyday concepts, the construction of which is done either with the help of theory or with the help of the conceptual analysis of the available empirical evidence. each term of the dictionary is a focal point of at least one theoretical approach associated with the corresponding nominal definition. given that in the social sciences there is no single scientific example, a term is usually not a focal point for a single theoretical approach, but functions as a meteoric signifier, as it takes on different meanings through different definitions given in the context of different theoretical approaches. easy access to the different meanings of a term through the different definitions given is important for conceptualization. 1. the purpose of the concept is to describe what is said with the help of scientific terms and decisions, that is, with concepts and decisions that are inscribed in the conceptual framework of the social sciences. the search for appropriate concepts and decisions is therefore crucial. in the dictionary, each term is theoretically documented both with the required bibliographic references and with the necessary comments. therefore, the conceptualization is theoretically substantiated and supported with the help of the dictionary through access to the appropriate bibliography, but also the appropriate clarifications and information provided by the commentary. 2. concepts often get their exact meaning from their relationships with concepts. the existence of correlations between the terms therefore provides important help for conceptualization. not only correlations based on experience provided to us by conceptual analysis, but also correlations based on theory that reinforce analogical thinking. the dictionary allows one term to be related to others and thus makes it possible to navigate from term to term. the correlation of the terms can be direct, that is, introduced by the authors of the entries and the https://doi.org/10.29173/iq1021 9/13 kallas, ioannis and kondyli, dimitra (2022). a tool to promote research planning and conceptualization: sodanet research infrastructure’s scientific dictionary of social terms, iassist quarterly 46(1), pp. 1-13. doi: https://doi.org/10.29173/iq1021 content managers. an indirect correlation arises from the search of all the terms that include the term we are looking for in their definitions. 3. conceptualization aims to produce not only concepts that fit into the theoretical framework of the social sciences, but also concepts suitable to support the codification of empirical research data. support for this process requires access not only to nominal definitions but mainly to functional ones. in addition to nominal definitions, which are linked to a theoretical framework through references to theory, the dictionary also provides functional definitions, which are linked to the specific uses of a term in the context of specific empirical research with a corresponding reference to them. by linking to specific empirical research, the dictionary supports codification as well as conceptualization. 6. the dictionary as a research design tool the dictionary can also be used as a research design tool, mainly because it is a tool for searching for scientific information necessary for the design of new research. the scientific terms to which this dictionary gives us access are the key points in the construction of scientific discourse. they can therefore be used as keys for searching and navigating the digital content of both this dictionary and other empirical research documentation infrastructures. using one or more terms, the researcher can: search and compile a bibliography, locate relevant empirical research and information on units of analysis and observation as well as methods. 1. the dictionary-supported search and navigation can be used in conjunction with the thematic classification thesauruses commonly used by empirical research documentation infrastructures such as the elsst. to the extent that the dictionary allows the thematic classification of its terms, with the help of such thesauruses, one can use the dictionary to identify the available scientific terms for each thematic category of a thesaurus as well as the thematic categories related to a specific term. thus, one can expand one's search criteria, starting with a term and identifying the relevant thematic categories, and from there identifying new related terms. 2. the dictionary provides information not only on terms and definitions, but also on the theoretical and research context in which these terms are used. this is achieved both through the commentary on terms and definitions and through references in the bibliography that include not only theoretical texts, but also research reports. the purpose of the comments is to include a definition in a specific theoretical framework and to link it to specific theories through access to the appropriate bibliography. this bibliography is not limited to the publication containing the stated definition, but is supplemented by other references that provide comments or clarifications related to the term or definition, allowing a fuller understanding of the term. 3. in addition to supporting search and navigation in the available digital stock of social science knowledge, the dictionary also provides information that is directly useful for the design of empirical research, both theoretically and methodologically: theoretically, by contributing directly to the theoretical overview of the research field by providing terms, definitions and bibliographic references and supporting the conceptualization, and methodologically by providing information on methods and research procedures, provided that the terms of the dictionary are not limited to the definition of objects of observation and phenomena but also extend to the definition of methods, research procedures or even epistemological approaches that are necessary for the construction of the methodology. to facilitate this process, the dictionary distinguishes the terms into theoretical and methodological. one can thus immediately seek information on terms relating to methods, such as their definitions, some critical references for their definition and description, as well as references to empirical research that contributed to the development of these methods. https://doi.org/10.29173/iq1021 10/13 kallas, ioannis and kondyli, dimitra (2022). a tool to promote research planning and conceptualization: sodanet research infrastructure’s scientific dictionary of social terms, iassist quarterly 46(1), pp. 1-13. doi: https://doi.org/10.29173/iq1021 7. the data model of the scientific dictionary of social terms the scientific dictionary of social terms was designed as part of a broader documentation infrastructure of empirical research. the dictionary becomes part of a broader documentation infrastructure mainly because it is designed to be shared with tools provided by the various social research documentation infrastructures. this is achieved to the extent that the data model on which its development is based is part of a more general data model used to develop empirical research documentation infrastructures. this more general model is described in figure 3 below. yellow indicates the entities that are directly related to the dictionary data model and have already been implemented in the corresponding application. the entities that are regularly implemented in research infrastructures for documentation of empirical research are generally shown in green, usually according to the ddi (2012), without, however, the dictionary being linked to them yet. the two models, yellow and green, allow correlations between their entities (such as those shown in red) without the need for the two models to be implemented in a common computer application. thus, the connection of the definitions with the units of analysis and observation, as well as the methods of analysis and recording, stems from the fact that the latter share with the former a term and a definition. the function of the dictionary as an element of a wider documentation infrastructure is based on the idea that scientific terms, and not just some words of everyday language, should be the key words for navigating the scientific discourse and searching for scientific information. the dictionary therefore provides access to the scientific terms of the social sciences, effectively providing access to the appropriate keywords to search for the digital content corresponding to a scientific discourse. even if the dictionary is not fully integrated into the computer system of a research infrastructure, it can still be used in parallel with it, providing the researcher with the appropriate terms to then use as keywords in any other application. https://doi.org/10.29173/iq1021 11/13 kallas, ioannis and kondyli, dimitra (2022). a tool to promote research planning and conceptualization: sodanet research infrastructure’s scientific dictionary of social terms, iassist quarterly 46(1), pp. 1-13. doi: https://doi.org/10.29173/iq1021 figure 3: entity-relationship diagram showing the interconnection of concepts, terms, research and coding books. 8. conclusion the scientific dictionary of social terms operates as an independent application within the sodanet research infrastructure, with the ultimate goal of being linked to its research. it supports the terminological and conceptual support of social research and especially empirical research, which organizes the production and analysis of empirical evidence based on the coding of evidence of experience, with the help of theory. the dictionary is dynamic and develops gradually over time through the work of many different researchers and through processes provided within a research infrastructure. it is bilingual, can be thematically linked to the elsst, and it can also contribute to the design of new research together with other available digital resources, to design the research questions and hypotheses of a new research. https://doi.org/10.29173/iq1021 12/13 kallas, ioannis and kondyli, dimitra (2022). a tool to promote research planning and conceptualization: sodanet research infrastructure’s scientific dictionary of social terms, iassist quarterly 46(1), pp. 1-13. doi: https://doi.org/10.29173/iq1021 references althusser, l. (1978) for marx. athens: grammata press (in greek). babbie, e. (2011) introduction to social research. vogiatzis, i. (translat.) zafeiropoulos, k. (eds). athens: kritiki (in greek). braun v. and clarke, v. (2012) ‘thematic analysis’ in cooper h. (eds.) apa handbook of research methods in psychology. washington: american psychological association, pp. 51-77. cessda elsst thesaurus (2021) elsst – european language social science thesaurus. available at: https://elsst.cessda.eu/ (accessed: 20th february 2021). ddi (2012). ddi-codebook 2.5. available at: https://ddialliance.org/specification/ddi-codebook/2.5/ (accessed: 24th february 2021). foucault m. (1992) l’archéologie du savoir. 2nd ed. paris: gallimard. kallas, j. (2015) theory, methodology and research infrastructures in social sciences. athens: kritiki press (in greek). schnell, r., hill, p. and esser e. (2014) methods of empirical social research. nagopoulos, n. (translat.) nagopoulos, n., giosos, j. and sakellario, a. (eds). athens: propompos (in greek). tsiolis, g. (2014) methods and analysis. techniques in qualitative social research. athens: kritiki press (in greek). willig c. (2015) qualitative survey methods in psychology. introduction. avgita, e. (translat.) tseliou, e. (eds). athens: gutenberg (in greek). endnotes 1ioannis kallas is professor at the university of the aegean and can be reached via email: i.kallas@soc.aegean.gr 2 dimitra kondyli is research director at the national centre for social research (ekke) – institute of social research in greece and can be reached via email: dkondyli@ekke.gr or dkondyli@gmail.com 3 this article is implemented in the framework of the ‘sodanet in action’ sub-project 2: university of the aegean, which is funded by the operational programme competitiveness entrepreneurship and innovation. the operation is co-financed by the european regional development fund (erdf). in addition to the authors, dr. d. paraskevopoulos, member of the research team, contributed to the editing of this article. 4 open archives initiative protocol for metadata harvesting enables harvesting metadata from a data repository and must support metadata in dublin core. it seeks to develop and promote interoperability standards that aim to facilitate the efficient dissemination of content and has its roots in the open access and institutional repository movements. more https://www.cessda.eu/training/training-resources/library/tutorial-access-anddissemination/data-discovery and https://www.openarchives.org/pmh/ https://doi.org/10.29173/iq1021 https://elsst.cessda.eu/ https://ddialliance.org/specification/ddi-codebook/2.5/ mailto:i.kallas@soc.aegean.gr mailto:dkondyli@ekke.gr mailto:dkondyli@gmail.com https://www.cessda.eu/training/training-resources/library/tutorial-access-and-dissemination/data-discovery https://www.cessda.eu/training/training-resources/library/tutorial-access-and-dissemination/data-discovery https://www.openarchives.org/pmh/ 13/13 kallas, ioannis and kondyli, dimitra (2022). a tool to promote research planning and conceptualization: sodanet research infrastructure’s scientific dictionary of social terms, iassist quarterly 46(1), pp. 1-13. doi: https://doi.org/10.29173/iq1021 5 it was developed in the framework of the project for the development of research infrastructures for the social sciences entitled ‘epae aegean: application development and data processing and documentation’ and continued in the project ‘sodanet in action’. 6 sodanet is the greek research infrastructure for social sciences, member of the consortium for european social science data archives (cessda eric). cessda has been on the european strategic forum for research infrastructures (esfri) roadmaps since 2006, became an esfri landmark in 2016, and as of 2017 it has been assigned european legal status as an eric: european research infrastructure consortium. 7 τhe european language social science thesaurus (elsst) is a broad-based, multilingual thesaurus for the social sciences, owned and published by the consortium of european social science data archives (cessda) and its national service providers. it is currently available in 14 languages and it includes 3,000 concepts covering the core social science disciplines: politics, sociology, economics, education, law, crime, demography, health, employment, information and communication technology and, increasingly, environmental science (cessda elsst thesaurus, 2021). the elsst is also a controlled vocabulary for the social sciences and comprises a structure that consists of terms that relate hierarchically (broader/narrower) or non-hierarchically (related and synonymous). thus, the increase in bilingual terms is planned to be linked to individual hierarchies of the european dictionary, adding value to existing terms and acting as an incentive to enrich sodanet's scientific dictionary terms. 8 apa (american psychological association) style is one of the most commonly used to cite sources within the social sciences. 9 the network of organizations which constitute the greek research infrastructure are one research centre , the national centre for social research and six university departments, namely university of aegean-faculty of sociology, university of athens-department of political science and public administration, panteion university department of political science and history, democritus university of thrace department of social policy, university of crete department of sociology, university of peloponnese: department of social & educational policy. see more at https://www.sodanet.gr/ https://doi.org/10.29173/iq1021 https://www.sodanet.gr/ instructions for authors of the iassist quarterly 1/2 schwartz, ofira & hayslett, michele (2024), building infrastructure and networks – rewards and challenges, iassist quarterly 48(3), pp. 1-2. doi: https://doi.org/10.29173/iq1136 the creative commons-attribution-noncommercial license 4.0 international applies to all works published by iassist quarterly. authors will retain copyright of the work and full publishing rights. editors’ notes: building infrastructure and networks – rewards and challenges welcome to the third issue of iassist quarterly for 2024, iq 48(3). as we are moving towards an open research environment, institutions are building infructractures that will enable sharing data and other research resources with a wider audience. the authors of the three papers in this issue offer our readers the benefit of their experience by sharing what they have learned through the process of establishing new infrustractures and networks. the article ”future models and architecture of data repositories in african universities,” describes the existing landscape of data repositories in african universities. chigwada and chiware use a review of existing literature to identify requirements for establishing an institutional data depository, and also identify successes and challenges. based on their research they offer a roadmap for universities in africa that are interested in establishing a data repository. the second article titled ”working towards securing and building a trusted institutional research data repository through the coretrustseal process: case of cape peninsula university of technology data repository” seems like a natural extention of the previous one. the three authors, lockhart, xesi and chiware (a co-authors of the previous paper), describe the process of establishing a research data repository at cape peninsula university of technology. they provide details about the journey, starting with developing open access (oa) and research data management (rdm) policies, identifying tools and developing the infrustructure needed for a an institutional data repository and data preservation, and developing a training program for faculty, students and staff. additionally the authors comment on their experience, challenges and lessons learned from the application for coretrustseal (cts) certification for their newly created repository, esango. in their article, ”building human networks to drive forward innovations in international data access: introducing the international secure data facility professionals network (isdfpn),” authors wiltshire, lichtwardt and bishop describe the motivation and process of establishing the international secure data facility professional network (isdfpn), a forum that brings together international colleagues to share expertise and experience, and to collaborate in developing trusted research environments (tres). we hope you enjoy reading. ofira schwartz and michele hayslett, september 2024 https://doi.org/10.29173/iq1136 https://ukdataservice.ac.uk/about/research-and-development/international-secure-data-facility-professionals-network-isdfpn/ https://ukdataservice.ac.uk/about/research-and-development/international-secure-data-facility-professionals-network-isdfpn/ https://creativecommons.org/licenses/by-nc/4.0/ 2/2 schwartz, ofira & hayslett, michele (2024), building infrastructure and networks – rewards and challenges, iassist quarterly 48(3), pp. 1-2. doi: https://doi.org/10.29173/iq1136 submissions of papers for the iassist quarterly are always very welcome. authors may take a look at the instructions and layout: https://iassistquarterly.com/index.php/iassist/about/submissions. we are available via e-mail for questions, or proposals for special issues: editor.iassistquarterly@gmail.com. https://doi.org/10.29173/iq1136 https://iassistquarterly.com/index.php/iassist/about/submissions mailto:editor.iassistquarterly@gmail.com ^ sist newsletter vol.1, no. 1 i united states secretariat report judith rowe computing center i, princeton university the united states secretariat has received positive letters of intent from 170 people, more than half of whom are already involved in action group activity. by popular demand, we are organizing a north american working conference in florida this february. the accomodations will be pleasant and inexpensive and the meeting will provide an opportunity for the ag's to work intensively on their projects. full details will be sent out shortly. i would like to take this opportunity to thank each of our ag coordinators for the enthusiastic efforts they have already expended and to encourage all of you not only to join lassist, but to participate in its activitites. action group reports data archive registry canadalisa lasko, canadian (.onsortium for social research, institute for behavioral research, york university, 4700 keele street, downsview, ontario m3j 1p3 europejoseph bonmariage, belgian archives for the social sciences, university of louvain, sh-2, 1348 louvain-la-neuve, belgium united statesdavid nasatir, behavioral sciences graduate program, california state college, domingus hills, california 90747 mandate a directory containing names, addresses, types of holdings, and dissemination policies of existing data archives and libraries throughout the world will be compiled. a supplementary directory listing archival personnel and other individuals with relevant expertise would be developed. in addition, a central location would be designated to maintain all published lists of archival holdings and study descriptions. these documents would be produced as reference documents for users of archives or institutions interested in consulting with individuals in the field of archiving. activities and plans with respect to the first part of the mandate, the action group will focus initially on developing the directory of data archives and data libraries. sist newsletter vol.1, no. 1 certain definitional problems are now being addressed by members of the us action group: what do we mean by a social science data archive? how does it differ from a social science data library? joseph bonmariage has suggested that there are different levels of production and dissemination. the issue of what constitutes "social science" has been raised. david nasitir suggested that "everything is {or can be) grist for the social scientific mill." he has proposed that by social science we mean "used by social scientists" rather than "produced by social scientists." this broadens the definition to include, for example, original or secondary distributors of process-produced data, a major social science resource. the directory would be developed in machine-readable form in order to facilitate updating and would either be published annually or available on a subscription basis. it will include entries for both archives and libraries and would include an appendix listing known organizations for which no further information is available. a questionnaire is being prepared for distribution by the first of the year and a preliminary version of the directory should be ready in april, 1977. in the meantime, a "primitive biliography" of existing data archive "registries" is being compiled which might appear in the directory in the form of an annotated bibliography. data acquisition canadapierre lacasse, centre d-^ recherches en am^nagement regional, universite de sherbrooke, sherbrojke, quebec europemarcia taylor, social science research council survey archive, university of essex, wivenhoe park, p.o. box 23, colchester, essex, england c04 3s0 united statesdonald harrison, national archives (nnr), washington, d.c. 20408 mandate this action group addresses the problems of data acquisition for archives with particular emphasis on the necessary relationship between the collectors and creators of data and the archives. recommended procedures for the acquisition of data would be developed with the intent of assisting researchers at critical points during the data collection process to ensure and promote the transfer of high quality data to the public domain for further academic investigation. the group will also survey different acquisition policies already in use, with attention to both legal and procedural problems connected with the acquisition and de-acquisition of data . [editor's note: the original mandate included the statement, "a survey of confidentiality laws, implications, and problems existing in machine-readable data would be conducted. a report summarizing the survey and resolving the problems when possible would be produced." this activity has been transferred to the action group on process-produced data. the under-lined section of the mandate of the data acquisition action group reflects an enlarging of the scope of activities to be undertaken.] 1 0. microsoft word 49-1-yeom-final.docx 1/21 yeom, semi (2025) literature review on the competencies of data literacy for middle-grade learners, iassist quarterly 49(1), pp. 121. doi: https://doi.org/10.29173/iq1123 the creative commons-attribution-noncommercial license 4.0 international applies to all works published by iassist quarterly. authors will retain copyright of the work and full publishing rights. literature review on the competencies of data literacy for middle-grade learners semi yeom1 abstract in today’s data-driven world, it is crucial for students to be data literate; able to view, understand, and reason with data in multimodal forms representing real-world phenomena. despite its importance, data literacy is rarely integrated into k-12 curricula, and its definition remains unclear for this age group. this paper reviews existing literature to define the competencies relevant to adolescent learners and highlights those crucial for middle-grade students. a literature review of theoretical and empirical discussions on data literacy concepts, instructional practices, and assessments revealed eight key competencies. among these, two were identified as most critical for middle-grade students: interpreting data representations and evaluating claims based on data representations. this paper aims to serve as a conceptual and practical guide to enhance data literacy in educational settings, providing a foundation for educators and researchers to collaboratively support middle-grade learners. keywords data literacy, k-12 school settings, literature review, middle-grade learners introduction in today's technology-driven world, data—numerical, textual, algorithmic, and visual—shape how we perceive and interact with our surroundings. algorithms guide online experiences, such as news articles featuring statistical insights or personalized advertisements based on browsing habits. big data, including numbers, texts, images, times, and locations, plays a significant role in individual and organizational decision-making (fontichiaro & oehrli, 2016; gould, 2017). data literacy involves critically analyzing, interpreting, and using data in various contexts. it encompasses understanding the social and cultural dimensions of data, which are deeply rooted in power structures and individual identities (twidale et al., 2013). researchers define data literacy as the ability to interact with and contextualize data, recognizing that interpretations are shaped by personal and cultural perspectives (english & watson, 2018; gunter, 2007; van’t hooft et al., 2012;). students today frequently encounter multimodal representations of real-world phenomena, and data literacy may provide them with the necessary tools to engage with these forms of information thoughtfully. with data literacy, students can study real-life problems and develop evidence-based solutions (erwin jr., 2015; yates et al., 2021;). this skill enables them to recognize the factors and contexts underlying 2/21 yeom, semi (2025) literature review on the competencies of data literacy for middle-grade learners, iassist quarterly 49(1), pp. 121. doi: https://doi.org/10.29173/iq1123 presented data and to apply data purposefully in diverse situations. students learn to identify messages and intentions behind data representations and evaluate their validity before drawing conclusions (philip et al., 2016). these critical skills empower them to make informed decisions and participate in public discussions about issues affecting their lives (gordon et al., 2016). in a world filled with streams of data, not all of which are neutral or truthful, data literacy can help students discern accurate information and evaluate sources critically (forzani, 2018; turton & martin, 2020;). beyond personal benefits, it fosters their ability to engage in civic activities and public discourse effectively (ercegovac, 2015). despite its importance, data literacy remains underrepresented in k-12 education. few instructional programs or initiatives integrate data literacy into curricula due to challenges such as limited classroom technology, insufficient teacher training, and rigid subject boundaries (gunter, 2007; van’t hooft et al., 2012). in many schools, opportunities to explore numerical and statistical data are scarce and often confined to science or mathematics classes (deahl, 2014). moreover, the emphasis on meeting core requirements and preparing for high stakes testing leaves little room for incorporating data literacy practices into lesson plans (ridsdale et al., 2015). another barrier to teaching data literacy is the lack of a clear, consistent definition for k-12 students. while researchers largely agree on its core elements, subtle differences in focus exist. some emphasize posing questions about data, using tools for representation, extracting relevant information, and evaluating inferences (english & watson, 2018; van’t hooft et al., 2012). others define it as accessing, analyzing, generating, and assessing data-based representations and inferences (love, 2004; vahey et al., 2012). organizations like the oceans of data institute highlight skills such as collecting, synthesizing, and visualizing data, while others underscore ethical considerations, such as understanding privacy issues and the intent behind data collection (gould, 2017). critical examination of data sources, their limitations, and their reliability is another recurring theme (lipton & wellman, 2012; yates et al., 2021). collectively, these competencies enable individuals to make data-informed decisions in everyday life (sorapure, 2019). although these competencies are foundational, a comprehensive framework defining data literacy and its proficiency levels remains elusive. scholars argue that the field must clarify critical competencies to support effective instruction for k-12 students (bowler & shaw, 2024). the challenge is compounded by overlaps with other literacies, such as statistical or informational literacy, and the multidisciplinary perspectives brought by researchers, educators, and data professionals (shields, 2005; twidale et al., 2013). middle-grade students, in particular, would benefit from targeted data literacy education. while attention to data literacy has grown in higher education as a core learning outcome (cubarrubia, 2019), studies focusing on middle school learners remain limited. early exposure during this developmental stage can lay a strong foundation for advanced data literacy skills in later years. recent efforts have emphasized the importance of introducing these concepts during secondary education (wolff et al., 2016). however, the lack of a clear understanding of essential competencies for adolescent learners presents a significant obstacle. before developing robust instructional models, it is crucial to identify the specific competencies middle-grade students need to attain data literacy. educators require a clearer understanding of these competencies to design effective curricula and integrate data literacy practices into school settings. 3/21 yeom, semi (2025) literature review on the competencies of data literacy for middle-grade learners, iassist quarterly 49(1), pp. 121. doi: https://doi.org/10.29173/iq1123 to address this need, this paper investigates existing literature to identify and categorize data literacy competencies relevant to adolescent learners. by synthesizing empirical and theoretical studies and cross-referencing findings with nationwide middle-grade learning standards, this study aims to clarify the essential competencies for middle-grade students and provide a foundation for data literacy education in k-12 settings. my research questions for this paper are as follows: 1. what does it mean to be data literate for adolescent learners? 2. what data literacy competencies would students in middle-grade years benefit from? methods i conducted a literature review to explore theoretical and empirical discussions on the concepts, instructional practices, and assessments of data literacy. this review focuses on identifying practices relevant to data literacy for school-aged learners. i searched education, library science, anthropology, and behavioral science databases to reflect the interdisciplinary nature of the field. using key terms aligned with my research questions, i included related terms with similar meanings. the review encompasses peer-reviewed journal articles published in english, including studies from international contexts, such as australia (callingham et al., 2017; english & watson, 2018;). a flowchart detailing the search and selection process is provided in figure 1, adapted from surrain and luk’s (2017) procedure. inclusion criteria of the studies the initial search (tier 1) yielded 58 articles, which i screened based on inclusion criteria. i included studies addressing data literacy or its complementary fields, such as information and statistical literacy (calzada prado & marzal, 2013; shields, 2005;). i focused on articles targeting adolescent learners or k-12 students and excluded those with overlapping datasets unless they raised distinct research questions. this process narrowed the selection to 23 articles. a subsequent review of these articles' literature sections (tier 2) identified seven additional studies meeting the same criteria. coding the studies to address my first research question, i analyzed the terms and definitions of data literacy discussed in theoretical frameworks, teaching practices, and assessments. i aggregated verb-noun phrases describing knowledge, skills, and abilities (ksas), categorized them by overarching themes, and consolidated a list of competencies. for example, discussions on learners formulating data-related questions were categorized as pose questions. to guide future research, i also coded characteristics of empirical studies, such as participants’ demographics, settings, and theoretical approaches. 4/21 yeom, semi (2025) literature review on the competencies of data literacy for middle-grade learners, iassist quarterly 49(1), pp. 121. doi: https://doi.org/10.29173/iq1123 figure 1: search process and selection criteria for my second research question, i cross-referenced these competencies with middle school learning standards, including common core state standards (ccss), next generation science standards (ngss), and middle school learning standards (msls) for subjects like english language arts (ela), science, social studies, and mathematics. each competency was broken into specific skills and aligned with relevant subject standards. for example, the competency pose questions was matched with standards on formulating research questions. i created a table linking competencies with specific standards codes and descriptions to ensure alignment and noted instances where a standard corresponded to multiple competencies. 5/21 yeom, semi (2025) literature review on the competencies of data literacy for middle-grade learners, iassist quarterly 49(1), pp. 121. doi: https://doi.org/10.29173/iq1123 characterization of the studies table 1 summarizes the participants, contexts, theories, methods, and subjects in the 22 empirical studies analyzed. participant numbers ranged from five to 7,000, with diverse racial, ethnic, and socioeconomic backgrounds. most studies included both native and non-native english speakers, though some omitted demographic details. in the u.s., more studies were conducted in urban than rural settings. studies were conducted between 2001 and 2022, with 10 based in the u.s. and 12 internationally. settings included regular classrooms (n=11) and after-school programs (n=6). table 1: analysis of characterization of the focal studies participants (n ranged 5 to 7,000) language english-as-l1 english-as-lx both unidentified race/ethnicity included european american african american latino/a asian native american unidentified no. of studies n=2 n=10 n=3 n=7 n=4 n=4 n=4 n=8 n=1 n=7 social class middle lower-middle lower unidentified no. of studies n=4 n=2 n=2 n=14 contexts u.s. rural urban unidentified state in the u.s. mid-west north-east south-west west unidentified n=10 n=1 n=4 n=5 n=2 n=2 n=2 n=3 n=1 non-u.s. suburban urban unidentified name of country turkey, canada, france, australia, germany, columbia, ecuador, indonesia, china, uk, korea in-school settings regular classroom after-school unidentified n=12 n=1 n=1 n=10 n=11 n=6 n=5 major theories statistical literacy information literacy data literacy n=10 n=9 n=12 6/21 yeom, semi (2025) literature review on the competencies of data literacy for middle-grade learners, iassist quarterly 49(1), pp. 121. doi: https://doi.org/10.29173/iq1123 new literacy studies/multiple literacies media literacy n=2 n=2 methodology qualitative case study testing content analysis quantitative survey testing mixed-methods n=10 n=3 n=2 n=2 n=10 n=4 n=6 n=2 interview participatory action research interpretive microanalysis phenomenography n=1 n=1 n=1 n=1 school subjects math science/stem social studies* english not specified *including history, political science n=7 n=5 n=4 n=2 n=3 other subject involved digital media literacy note. stem: science, technology, engineering, and mathematics the studies often referenced overlapping concepts of data, statistical, and information literacy, with some adopting broader frameworks like new literacy studies. methodologies varied; ten studies used quantitative methods, ten studies used qualitative methods, and two studies employed mixed methods. six quantitative and two qualitative studies used assessments as a method. mathematics, science, and social studies were the primary subjects, although three studies did not link practices to specific disciplines. findings the results of my literature review are organized according to the two research questions: 1) what are the data literacy competencies for adolescent learners; and 2) what are the benefits of developing data literacy competencies during adolescence. what does it mean to be data literate for adolescent learners? in response to the first research question, i synthesized ksas that are emphasized in research on adolescents’ data literacy (see table 2) into eight competencies: pose questions, access/collect, transform, manage/handle, analyze, interpret, evaluate, answer questions, and present/communicate. appendix a illustrates the relevant articles that mentioned each competency and corresponding learning standards. 7/21 yeom, semi (2025) literature review on the competencies of data literacy for middle-grade learners, iassist quarterly 49(1), pp. 121. doi: https://doi.org/10.29173/iq1123 table 2: data literacy by competency pose questions  investigate authentic problems  formulate/articulate data-based questions  identify the audience and context for the problem  define your goals and how they can be achieved with data access/collect  know how to select research and statistical methods and tools that align to purposes  demonstrate an understanding of validation  document methods and tools  understand a wide variety of tools for accessing data  investigate the source of information  explore the data that are available transform  clean, transform, manipulate & synthesize data  understand a wide variety of tools for converting and manipulating data  synthesize information from multiple sources/map data across heterogeneous sources  understand how representations in computers can vary and why data must sometimes be altered before analysis manage/handle  understand issues of data privacy, confidentiality, ownership, and handling process  understand the importance of the provenance of data and how data are stored  understand ethical use of personal data  understand ways of protecting and documenting data to maintain reproducibility  understand issues of data quality and credibility markers  assess risk and bias involved in conducting the study analyze  develop an analysis plan and conduct exploratory analyses  know how to use analytic software  understand some aspects of predictive modeling interpret  understand the data representations  interpret information from data  develop data-based inferences and explanations  explore patterns in data with a skeptical but open mind 8/21 yeom, semi (2025) literature review on the competencies of data literacy for middle-grade learners, iassist quarterly 49(1), pp. 121. doi: https://doi.org/10.29173/iq1123  produce explanations, comparisons and predictions based on the variability in the data  compare the results with other findings evaluate  understand statistical arguments  evaluate data-based claims, inferences and explanations  use appropriate data, tools, and representations to support critical thinking  use data as part of evidence-based analytic thinking answer questions  answer data-based questions  make statistically sound decisions  discuss and document findings  form a judgment and conclusions and draw arguments from data  discuss limitations of the study present/communicate  create and construct basic descriptive representations and visualizations of data to answer questions about real-life processes  translate and present information into different data representations  communicate solutions and recommendations  understand how to do citation pose questions. this competency entails identifying problems to solve based on data, with an emphasis on audience and context. it requires defining clear goals and determining how data can be used to achieve those objectives. this competency aligns with wolff et al.’s (2016) ppdac cycle— problem, plan, data, analysis, and conclusion—where posing questions serves as a critical initial step in framing data-driven inquiries in k-12 contexts, as demonstrated in various studies (ercegovac, 2015; van’t hooft et al., 2012;). kim et al. (2016) also highlighted the ability to “formulate questions to search for information” (p. 445) when they asked students to self-assess their information literacy. access/collect. gathering information and data requires selecting appropriate research methods, validating sources, and documenting tools used for data access (cunningham et al., 2018). this includes evaluating sources, understanding data relevance, and exploring available data. ridsdale et al. (2015) highlight “collection” as a core competency for data literacy, emphasizing the ability to search for, assess, and contextualize data. aillerie et al. (2016) explore how teenagers practice this skill on social networking sites. programs like city digits, where public-school students analyze state lottery data to build evidence-based arguments, demonstrate how such competencies can be applied to real-world contexts (deahl, 2014). transform. this competency includes cleaning and synthesizing information from diverse sources, critical for handling raw data. it requires a deep understanding of multiple data conversion and manipulation tools, the ability to synthesize information from multiple sources, and the capability to 9/21 yeom, semi (2025) literature review on the competencies of data literacy for middle-grade learners, iassist quarterly 49(1), pp. 121. doi: https://doi.org/10.29173/iq1123 map data across different platforms. additionally, one needs a grasp of varied computer representations and rationales for altering data before analysis. data literacy projects, like those at the university of maine, encourage students to visualize and interpret real data sets, reinforcing the importance of transformation skills (wolff et al., 2016). cohen et al. (2017) also emphasized that transforming data requires an understanding of technology systems, an essential part of preparing students for stem fields. manage/handle. this competency involves understanding data privacy, quality, and credibility markers. examples of credibility markers include data relevance to users’ interests and authorship. understanding authorship requires knowledge about authorial intention, which involves selecting data or visualizations based on intended meanings, contexts, and sociocultural codes (hullman & diakopoulos, 2011). understanding credibility markers, authorship, and data relevance, as discussed in the literature, equips students to handle data ethically and critically evaluate sources and interpretations (utomo, 2021). analyze. this competency focuses on hands-on data analysis, creating an analysis plan, and using tools proficiently. additionally, it includes aspects of predictive modeling, allowing students to extract meaningful insights from data. this competency is reinforced by curricula such as thinking with data, which engages students in problem-based exercises to compare learning gains in data literacy between those exposed to the curriculum and those who are not (van’t hooft et al., 2012). zalles (2005) also highlighted students’ application of analytical skills in real-world data in the epa phoenix performance task about air quality (zalles, 2005). interpret. this competency is for understanding processed data and visualizations, exploring and extracting information from mapped data, graphs, pie charts, and emerging forms of visualizations (gordon et al., 2016). it requires skills to extract insights from visual representations. body of research emphasizing the necessity of interpreting variability and terminology used in data, a skill practiced in both k-12 and postsecondary contexts (ben-zvi & arcavi, 2001; mandinach & gummer, 2012; wolff et al., 2016; yolcu, 2014;; ). for instance, ben-zvi and arcavi (2001) emphasized the importance of understanding variable values in data representations such as tables and graphs. evaluate. this competency encompasses assessing the trustworthiness of data claims, which is essential for informed judgments and data-driven insights (gunter, 2007). it includes detecting loaded language or omitted facts based on the author’s purpose or context. the competency allows critical investigation into arguments derived from data and identifying missing information to strengthen those arguments (büscher, 2022). it supports critical inquiry into the evidence behind claims, connecting data to the meanings conveyed, and questioning the legitimacy of those connections (cunnimgham et al., 2018). seroff (2017) described exercises where students evaluate limitations in data collection methods about traffic fatalities, reinforcing the importance of questioning data validity and generalizability. answer questions. this competency focuses on addressing problems or questions identified at the beginning of data-related tasks (utomo, 2021). drawing conclusions based on data requires, understanding the limitations of data from collection methods or unwarranted assumptions. it necessitates forming judgments and making arguments grounded in data, while also discussing limitations inherent in the study. project-based learning further facilitates this competency by addressing authentic societal problems through the examination of data and the use of digital tools 10/21 yeom, semi (2025) literature review on the competencies of data literacy for middle-grade learners, iassist quarterly 49(1), pp. 1-21. doi: https://doi.org/10.29173/iq1123 (erwin jr., 2015; wu et al., 2020). van’t hooft et al. (2012)’s thinking with data project assessments also include performance tasks that gauge students’ ability to provide well-supported answers and suggest what additional data would be needed to answer questions. present/communicate. this competency involves effectively communicating data-based understandings and findings to others (wu et al., 2020). it includes translating information into various data representations and following proper citation practices. creating and employing different graph types is essential for middle school students (hunter-thomson, 2019). for instance, city digits and the globe integrated investigation assessments emphasize skills in representing data through various formats, such as graphs and maps, to communicate findings effectively (wolff et al., 2016; zalles, 2005). what are the benefits of developing data literacy competencies during adolescence? developing data literacy competencies during adolescence, particularly interpret and evaluate, can help students gain essential skills for academic success, make informed decisions, and engage critically with the increasingly data-driven world. these competencies enable students to analyze, interpret, and evaluate data representations across various contexts, fostering critical thinking and problemsolving abilities. strengthening data literacy at this stage prepares students for future educational and professional settings where data-driven reasoning is crucial. interpret is one of the most emphasized competencies in existing literature, with the largest number of corresponding learning standards for middle-grade students. this competency is a fundamental competency that helps students comprehend and analyze different forms of data representation, including graphs, tables, and multimedia formats. thirteen articles, including those by callingham et al. (2017) and büscher (2022), identified interpret as a core competency of data literacy. yolcu (2014), in defining statistical literacy within a three-tiered framework, highlighted the ability to interpret statistical information and messages as the second tier. projects like city digits demonstrate the practical application of these skills by engaging students in interpreting data to analyze social issues, such as studying lottery data, thereby strengthening their interpretive abilities in meaningful contexts (deahl, 2014). findings also revealed that many middle-grade learning standards across various disciplines emphasized interpret as a key competency. for instance, the ccss highlight the ability to determine central ideas, information, or conclusions in literacy for history, social studies, science, and technical texts. similarly, ngss underscore students’ ability to interpret graphical displays of data to identify linear and nonlinear relationships. in mathematics, the msls focus on the ability to discuss and understand the correspondence between data sets and their graphical representations. evaluate is also among the competencies most frequently addressed in the studies. this competency allows students to critically analyze claims based on data and assess the validity of conclusions drawn from data-driven arguments. evaluate is prominently featured in research, with 21 studies identifying it as essential for data literacy. for instance, womack (2015) emphasizes the skills required to critically evaluate information and integrate it into students’ knowledge and value systems. chin et al. (2016) measured evaluation in their choicelets assessment by testing students' ability to critique sources of visually presented information. similarly, van’t hooft et al. (2012) used problem-based exercises requiring students to evaluate air quality data and make evidence-based recommendations. the importance of evaluate is further underscored in today’s data-driven society, where critical skills are 11/21 yeom, semi (2025) literature review on the competencies of data literacy for middle-grade learners, iassist quarterly 49(1), pp. 1-21. doi: https://doi.org/10.29173/iq1123 needed to assess the credibility of messages in media and web-based information (bussert-webb et al., 2017; cuervo sánchez et al., 2021; leu et al., 2013 ). numerous learning standards emphasize the importance of evaluate across disciplines, highlighting its multidisciplinary relevance. the ngss stress the ksas needed to assess the merit and validity of ideas and methods, while the ccss call for the ability to evaluate content presented in diverse formats and media. additionally, the ccss underscore skills in distinguishing between facts and reasoned judgments derived from research findings, particularly in the contexts of history, social studies, and science and technical texts. these two competencies are interrelated, with many researchers classifying them under one category. for example, interpret and evaluate belong to the “conclusion” step in the ppdac cycle defined by wolff et al. (2016). this step encompasses interpreting data to understand patterns and evaluating the validity of explanations based on that data. additionally, both competencies align with the “data evaluation” category identified by ridsdale et al. (2015), which encompasses the ksas needed to assess graphical representations of data. leu et al. (2013) also emphasize that evaluate should be assessed alongside a related competency, further supporting the focus on these closely linked but distinct competencies. discussion and implications this paper contributes significantly to the conceptualization of data literacy by synthesizing different arrays of definitions and competencies. data literacy encompasses a range of ksas, making it challenging to define and tailor for students across grade levels. to address this, i conducted a literature review of data literacy and related constructs, such as information and statistical literacy, to aggregate and categorize competencies. among these, two key competencies—interpreting data representations and evaluating claims based on data—emerged as most emphasized in the learning standards for middle-grade students. by cross-referencing competencies with standards like ccss and ngss, i highlighted their relevance to core subjects such as math, science, english, and social studies. the findings have significant implications in developing a practical framework of data literacy education applicable to middle school settings. they can guide educators in initiating data literacy instruction by providing clear target competencies aligned with learning standards. this shared understanding can help implement data literacy practices across subjects, offering insights for designing interdisciplinary curricula. curricula focusing on the development of interpret and evaluate skills can be incorporated into syllabi aimed at fostering other competencies. they can also serve as a foundation for helping students develop data literacy skills throughout high school and higher education. the study highlights the need for future research to address gaps in current data literacy education. few studies have been conducted in rural settings or included minoritized populations, such as native american students. additionally, most studies employing instructional interventions lacked assessment tools to measure data literacy, limiting their ability to evaluate the impact on students’ competencies. only two studies in this review utilized mixed-method approaches, which are essential for capturing both quantitative outcomes and qualitative insights into learners’ experiences. this finding underscores the need for more research focusing on underrepresented populations and incorporating validated assessments and diverse methodologies. the paper emphasizes the 12/21 yeom, semi (2025) literature review on the competencies of data literacy for middle-grade learners, iassist quarterly 49(1), pp. 1-21. doi: https://doi.org/10.29173/iq1123 importance of integrating instruction and assessment practices to enhance data literacy among diverse middle-grade students. references aillerie, k., & mcnicol, s. (2016). information literacy and social networking sites: challenges and stakes regarding teenagers’ uses. essachess–journal for communication studies, 9(2-18), 89–100. ben-zvi, d., & arcavi, a. (2001). junior high school students’ construction of global views of data and data representations. educational studies in mathematics, 45, 35–65. https://doi.org/10.1023/a:1013809201228 bowler, l., & shaw, c. (2024). trends in data literacy, 2018-2023: a review of the literature. information research an international electronic journal, 29(2), 198-205. https://doi.org/10.47989/ir292822 büscher, c. (2022). design principles for developing statistical literacy in middle schools. statistics education research journal, 21(1), 1–16. https://doi.org/10.52041/serj.v21i1.80 bussert-webb, k., & henry, l. (2017). promising digital practices for nondominant learners. international journal of educational technology, 4(2), 43–55. callingham, r & watson, j.m. (2017) the development of statistical literacy at school. statistics education research journal, 16(1), 181–201. calzada prado, j. c., & marzal, m. á. (2013). incorporating data literacy into information literacy programs: core competencies and contents. libri, 63(2), 123–134. https://doi.org/10.1515/libri-2013-0010 chin, d. b., blair, k. p., & schwartz, d. l. (2016). got game? a choice-based learning assessment of data literacy and visualization skills. technology, knowledge and learning, 21(2), 195–210. https://doi.org/10.1007/s10758-016-9279-7 cohen, j. d., renken, m., & calandra, b. (2017). urban middle school students, twenty-first century skills, and stem-ict careers: selected findings from a front-end analysis. techtrends, 61, 380-385. https://doi.org/10.1007/s11528-017-0170-8 cubarrubia, a. p. (2019, october 18). we all need to be data people. the chronicle of higher education, 66(7). https://www.chronicle.com/article/we-all-need-to-be-data-people/ cuervo sánchez, s. l., foronda rojo, a., rodriguez martinez, a., & medrano samaniego, c. (2021)media and information literacy: a measurement instrument for adolescents. educational review, 73(4), 487-502. https://doi.org/10.1080/00131911.2019.1646708 cunningham, v., & williams, d. (2018). the seven voices of information literacy (il). journal of information literacy, 12(2). http://dx.doi.org/10.11645/12.2.2332 13/21 yeom, semi (2025) literature review on the competencies of data literacy for middle-grade learners, iassist quarterly 49(1), pp. 1-21. doi: https://doi.org/10.29173/iq1123 deahl, e. s. (2014). better the data you know: developing youth data literacy in schools and informal learning environments. [master’s thesis, the massachusetts institute of technology]. https://doi.org/10.2139/ssrn.2445621 english, l. d., & watson, j. (2018). modelling with authentic data in sixth grade. zdm mathematics education, 50(1–2), 103–115. https://doi.org/10.1007/s11858-017-0896-y ercegovac, z. (2015). data-driven society begins with data-savvy youth. bulletin of the association for information science & technology, 42(1), 42–48. https://doi.org/10.1002/bul2.2015.1720420111 erwin, r. w., jr. (2015). data literacy: real-world learning through problem-solving with data sets. american secondary education, 43(2), 18–26. https://www.jstor.org/stable/43694208 fontichiaro, k., & oehrli, j. a. (2016). why data literacy matters. knowledge quest, 44(5), 21–27. https://eric.ed.gov/?id=ej1099487 forzani, e. (2018). how well can students evaluate online science information? contributions of prior knowledge, gender, socioeconomic status, and offline reading ability. reading research quarterly, 53(4), 385–390. https://doi.org/10.1002/rrq.218 gebre, e. (2022). conceptions and perspectives of data literacy in secondary education. british journal of educational technology, 53(5), 1080–1095. https://doi.org/10.1111/bjet.13246 godaert, e., aesaert, k., voogt, j., & van braak, j. (2022). assessment of students’ digital competences in primary school: a systematic review. education and information technologies, 27, 9953–10011. https://doi.org/10.1007/s10639-022-11020-9 gordon, e., elwood, s., & mitchell, k. (2016). critical spatial learning: participatory mapping, spatial histories, and youth civic engagement. children’s geographies, 14, 1–15. https://doi.org/10.1080/14733285.2015.1136736 gould, r. (2017). data literacy is statistical literacy. statistics education research journal, 16(1), 22– 25. https://doi.org/10.52041/serj.v16i1.209 guler, m., gursoy, k., & guven, b. (2016). critical views of 8th grade students toward statistical data in newspaper articles: analysis in light of statistical literacy. cogent education, 3(1), http://dx.doi.org/10.1080/2331186x.2016.1268773 gunter, g. a. (2007). building student data literacy: an essential critical-thinking skill for the 21st century. multimedia & internet @ schools, 14(3), 24–28. hullman, j., & diakopoulos, n. (2011). visualization rhetoric: framing effects in narrative visualization. ieee transactions on visualization and computer graphics, 17(12), 2231–2240. https://doi.org/10.1109/tvcg.2011.255 hunter-thomson, k. (2019). data literacy 101: how do we set up graphs in science? science scope, 42(2), 78–82. 14/21 yeom, semi (2025) literature review on the competencies of data literacy for middle-grade learners, iassist quarterly 49(1), pp. 1-21. doi: https://doi.org/10.29173/iq1123 kim, e. m., & yang, s. (2016). internet literacy and digital natives’ civic engagement: internet skill literacy or internet information literacy? journal of youth studies, 19(4), 438–456. https://doi.org/10.1080/13676261.2015.1083961 kostelnick, c. (2004). melting-pot ideology, modernist aesthetics, and the emergence of graphical conventions: the statistical atlases of the united states, 1874–1925. in c. a. hill & m. helmers (eds.), defining visual rhetorics (pp. 215–242). routledge. leu, d. j., forzani, e., burlingame, c., kulikowich, j., sedransk, n., coiro, j., & kennedy, c. (2013). the new literacies of online research and comprehension: assessing and preparing students for the 21st century with common core state standards. in s. neuman (ed.), quality reading instruction in the age of common core standards (pp. 219–236). international reading association. lipton, l., & wellman, b. (2012). got data? now what? creating and leading cultures of inquiry. solution tree press. love, n. (2004). taking data to new depths. journal of staff development, 25(4), 22–26. https://all4ed.org/wp-content/uploads/2013/09/lovejsdarticle1.pdf mandinach, e. b., & gummer, e. s. (2012). navigating the landscape of data literacy: it is complex. wested. https://eric.ed.gov/?id=ed582807 manduca, c. a., & mogk, d. w. (2002). using data in undergraduate science classrooms: final report on an interdisciplinary workshop held at carleton college. https://d32ogoqmya1dw8.cloudfront.net/files/usingdata/usingdata.pdf oceans of data institute (odi) (2015). building global interest in data literacy: a dialogue. educational development center. https://oceansofdata.org/our-work/building-globalinterest-data-literacy-dialogue-workshop-report philip, t. m., olivares-pasillas, m. c., & rocha, j. (2016). becoming racially literate about data and data-literate about race: data visualizations in the classroom as a site of racial-ideological micro-contestations. cognition and instruction, 34(4), 361–388. https://doi.org/10.1080/07370008.2016.1210418 ridsdale, c., rothwell, j., smit, m., ali-hassan, h., bliemel, m., irvine, d., kelley, d., matwin, s., & wuetherick, b. (2015). strategies and best practices for data literacy education: knowledge synthesis report. dalhousie university. https://dalspaceb.library.dal.ca/server/api/core/bitstreams/7c0db868-79ac-4c3c-9ad0f2a93cf2ccc6/content shields, m. (2005). information literacy, statistical literacy, data literacy. iassist quarterly, 28(2-3), 6–11. https://doi.org/10.29173/iq790 seroff, j. (2017). using data in the research process. in k. fontichiaro, j. a. oehrli, & a. lennex (eds.), creating data literate students. michigan publishing. 15/21 yeom, semi (2025) literature review on the competencies of data literacy for middle-grade learners, iassist quarterly 49(1), pp. 1-21. doi: https://doi.org/10.29173/iq1123 sorapure, m. (2019). text, image, data, interaction: understanding information visualization. computers and composition, 54, 1–16. https://doi.org/10.1016/j.compcom.2019.102519 statewide longitudinal data systems grant program. (2015). slds data use standards: knowledge, skills, and professional behaviors for effective data use, version 2. u.s. department of education national center for education statistics. https://eric.ed.gov/?id=ed595040 surrain, s., & luk, g. (2019). describing bilinguals: a systematic review of labels and descriptions used in the literature between 2005–2015. bilingualism: language and cognition, 22(2), 401-415. https://doi.org/10.1017/s1366728917000682 turton, w., & martin, a. (2020, january 5). how deepfakes make disinformation more real than ever. bloomberg. https://www.bloomberg.com/news/articles/2024-02-09/fighting-deepfakeswhats-being-done-biden-robocalls-to-taylor-swift-ai-images twidale, m. b., blake, c., & gant, j. (2013). towards a data literate citizenry. in proceedings of iconference 2013 (pp. 247–257). https://hdl.handle.net/2142/38385 vahey, p., rafanan, k., patton, c., swan, k., hooft, m., kratcoski, a., & stanford, t. (2012). a crossdisciplinary approach to teaching data literacy and proportionality. educational studies in mathematics, 81(2), 179–205. https://doi.org/10.1007/s10649-012-9392-z vahey, p., yarnall, l., patton, c., zalles, d., & swan, k. (2006, april). mathematizing middle school: results from a cross-disciplinary study of data literacy. presented at the american educational research association annual conference, san francisco, ca. van’t hooft, m., swan, k., cook, d., stanford, t., vahey, p., kratcoski, a., rafanan, k., & yarnall, l. (2012). a cross-curricular approach to the development of data literacy in the middle grades: the thinking with data project. middle grades research journal, 7(3), 19–33. viegas, f. b., & wattenberg, m. (2006). communication-minded visualization: a call to action. ibm systems journal, 45(4), 801–812. https://doi.org/10.1147/sj.454.0801 wested. (n.d.). developing assessments of data literacy. wested. https://www.wested.org/project/data-literacy-assessment-development/ wolff, a., gooch, d., cavero montaner, j. j, rashid, u., kortuem, g., (2016). creating an understanding of data literacy for a data-driven society. the journal of community informatics, 12(3), 9–26. https://doi.org/10.15353/joci.v12i3.3275 womack, r. (2015). data visualization and information literacy. iassist quarterly, 38(1), 12–17. https://doi.org/10.29173/iq619 wu, d., yu, l., yang, h. h., zhu, s., & tsai, c. c. (2020). parents’ profiles concerning ict proficiency and their relation to adolescents’ information literacy: a latent profile analysis approach. british journal of educational technology, 51(6), 2268-2285. https://doi.org/10.1111/bjet.12899 16/21 yeom, semi (2025) literature review on the competencies of data literacy for middle-grade learners, iassist quarterly 49(1), pp. 1-21. doi: https://doi.org/10.29173/iq1123 yates, s., carmi, e., lockley, e., wessels, b. & pawluczuk, a. (2021). me and my big data: understanding citizens data literacies. nuffield foundation. https://openaccess.city.ac.uk/id/eprint/26756/ yolcu, a. (2014) middle school students’ statistical literacy: role of grade level and gender. statistics education research journal, 13(2), 118–131. https://doi.org/10.52041/serj.v13i2.285 zalles, d. (2005). designs for assessing foundational data literacy. https://serc.carleton.edu/files/nagtworkshops/assess/zallesessay3.pdf zoellick, b., schauffler, m., flubacher, m., weatherbee, r., & webber, h. (2016, april). data literacy: assessing student understanding of variability in data. paper presented at the annual meeting of the national association for research in science teaching. baltimore, md. endnotes 1 semi yeom, university of maryland, college park, usa, can be reached at syeom@terpmail.umd.edu 17/21 yeom, semi (2025) literature review on the competencies of data literacy for middle-grade learners, iassist quarterly 49(1), pp. 1-21. doi: https://doi.org/10.29173/iq1123 appendix a: competencies of data literacy as a function of relevant learning standards and literature competency relevant standards specific examples relevant literature pose questions ● ccss math g6-12 ● msls mathematics/social studies/science/ela standards ● ngss g6-8 ● make sense of problems and persevere in solving them. ● formulate/address questions/pose and define problems and design/execute studies. ● develop the ability to refine & refocus broad & illdefined questions. ● use conjectures to formulate new questions & studies to answer them. vahey et al. (2012); van’t hooft et al. (2012); ben-zvi & arcavi (2001); erwin jr. (2015); chin et al. (2016); wolff et al. (2016); guler et al. (2016); kim et al. (2016); utomo (2021) access/collect ● ccss writing g6-12 ● msls mathematics/social studies/ela ● ngss g6-8 ● ccss math g7 ● conduct short as well as more sustained research projects based on focused questions, demonstrating understanding of the subject under investigation. ● gather relevant information from multiple print and digital sources. ● collect/gather data about a characteristic shared by two populations or different characteristics within one population. ● use random sampling to draw inferences about a population. erwin jr. (2015); van’t hooft et al. (2012); shields (2005); deahl (2014); calzada prado et al. (2013); ridsdale et al. (2015); ercegovac (2012); wolff et al. (2016); aillerie et al (2016); cuervo sánchez et al. (2021); cunningham et al. (2018); kim et al. (2016); bussert-webb et al. (2017) transform ● msls ela ● ccss reading/speaking & listening/writing g6-12 ● ccss literacy in science and ● synthesize data from a variety of sources. ● integrate content presented in diverse formats and media, including visually and quantitatively, as well as in words. ● integrate quantitative or technical information expressed in words in a text with a version of that information expressed visually. calzada prado et al. (2013); vahey et al. (2012); van’t hooft et al. (2012); english & watson (2018); gould (2017); erwin jr. (2015); chin et al. (2016); shields (2005); gordon et al. (2016); wolff et al. (2016); cuervo 18/21 yeom, semi (2025) literature review on the competencies of data literacy for middle-grade learners, iassist quarterly 49(1), pp. 1-21. doi: https://doi.org/10.29173/iq1123 technical texts g6-8 sánchez et al. (2021); cohen et al. (2017); wu et al. (2020) manage/handl e ● ccss writing g6-12 ● know how to avoid plagiarism. ● assess the credibility and accuracy of each source. calzada prado et al. (2013); gould (2017); erwin jr. (2015); deahl (2014); gordon et al. (2016); ercegovac (2012); utomo (2021) analyze ● ngss g6-8 ● ccss reading/literacy g6-12 ● ccss math g6-12 ● ngss g6-8 ● ccss literacy in history/social studies g6-8 ● msls math ● investigate chance processes and develop, use, and evaluate probability models. ● develop, use, and revise models to describe, test, and predict more abstract phenomena and design systems. ● analyze human behavior in relation to its physical and cultural environments. ● extend quantitative analysis to investigations, distinguishing between correlation and causation, and basic statistical techniques of data and error analysis. ● investigate patterns of association in bivariate data. ● identify patterns in large data sets. ● analyze the relationship between a primary and secondary source on the same topic. ⃰ ● specify relationships between variables and clarify arguments and models. ⃰ ● use observations about differences between two or more samples to make conjectures about the populations ⃰ ● develop understanding of statistical variability ⃰ calzada prado et al. (2013); gould (2017); ben-zvi & arcavi (2001); english & watson (2018); erwin jr. (2015); chin et al. (2016); shields (2005); zalles (2005); zoellick et al. (2016); ridsdale et al. (2015); guler et al. (2016); yolcu (2014); buscher (2022); cuervo sánchez et al. (2021); bussert-webb et al. (2017) 19/21 yeom, semi (2025) literature review on the competencies of data literacy for middle-grade learners, iassist quarterly 49(1), pp. 1-21. doi: https://doi.org/10.29173/iq1123 interpret ● ccss literacy in history/social studies/science and technical texts g6-8 ● msls math/social studies/science ● ccss math g6-12 ● ngss g6-8 ● ccss writing g6-12 ● determine the central ideas or information or conclusions. ● determine the meaning of keywords, symbols, domains-specific words and phrases in a specific context. ● discuss & understand the correspondence between data sets & their graphical representations. ● construct sound historical interpretations. ● compare and contrast the information gained from experiments, simulations, video, or multimedia sources with that gained from reading a text on the same topic. ● interpret graphical displays of data and/or large data sets to identify linear and nonlinear relationships. ● analyze the relationship between a primary and secondary source on the same topic. ⃰ ● use observations about differences between two or more samples to make conjectures about the populations. ⃰ ● use evidence to generate explanations. ⃰ ● use appropriate tools strategically such as diagrams, two-way tables, graphs, flowcharts and formulas to identify important quantities in a practical situation and map their relationships. ⃰ deahl (2014); vahey et al. (2012); van’t hooft et al. (2012); calzada prado et al. (2013); twidale et al. (2013); zalles (2005); ridsdale et al. (2015); wolff et al. (2016); guler et al. (2016); callingham et al. (2017); yolcu (2014); buscher (2022); cuervo sánchez et al. (2021) evaluate ● ccss reading/ speaking & listening g6-12 ● ngss g6-8 ● ccss literacy in ● evaluate content presented in diverse formats and media, including visually and quantitatively, as well as in words and orally. ● evaluate the merit and validity of ideas and methods. calzada et al. (2013); gould (2017); english & watson (2018); erwin jr. (2015); shields (2005); deahl (2014); vahey et al. (2012); van’t hooft et al. (2012); calzada prado et al. (2013); 20/21 yeom, semi (2025) literature review on the competencies of data literacy for middle-grade learners, iassist quarterly 49(1), pp. 1-21. doi: https://doi.org/10.29173/iq1123 history/social studies/science and technical texts g6-8 ● msls social studies/ science ● analyze the author’s purpose in providing an explanation, describing a procedure, or discussing an experiment in a text. ● distinguish among fact, opinion, and reasoned judgment based on research findings in a text. ● question and identify gaps in data. ● propose alternative explanations and critique explanations and procedures. ● delineate the argument and specific claims in a text, including the validity of the reasoning as well as the relevance and sufficiency of the evidence. twidale et al. (2013); womack (2015); zalles (2005); ercegovac (2012); wolff et al. (2016); aillerie et al (2016); callingham et al. (2017); yolcu (2014); cuervo sánchez et al. (2021); cohen et al. (2017); kim et al. (2016); bussert-webb et al. (2017) answer questions ● msls science ● ngss g6-8 ● ccss writing g6-12 ● use evidence to generate explanations. ⃰ ● use mathematical concepts to support explanations and arguments. ⃰ ● draw evidence from literary or informational texts to support analysis, reflection, and research. ⃰ ● construct explanations and design solutions supported by multiple sources of evidence consistent with scientific ideas, principles, and theories. ● construct a convincing argument that supports or refutes claims for either explanations or solutions about the natural and designed world(s). vahey et al. (2012); van’t hooft et al. (2012); calzada prado et al. (2013); gunter (2007); twidale et al. (2013); womack (2015); zalles (2005); guler et al. (2016); callingham et al. (2017); cohen et al. (2017); wu et al. (2020); utomo (2021) present/com municate ● msls math/ela ● ccss speaking & listening g6-12 ● ngss g6-8 ● ccss literacy in history/social ● prepare for and participate effectively in a range of conversations and collaborations with diverse partners, building on others’ ideas and expressing their own clearly and persuasively. ● use mathematical, computational, and/or algorithmic representations of vahey et al. (2012); gould (2017); english & watson (2018); erwin jr. (2015); van’t hooft et al. (2012); chin et al. (2016); shields (2005); deahl (2014); gunter (2007); twidale et al. (2013); womack (2015); zalles 21/21 yeom, semi (2025) literature review on the competencies of data literacy for middle-grade learners, iassist quarterly 49(1), pp. 1-21. doi: https://doi.org/10.29173/iq1123 studies/science and technical texts g610 ● ccss math g6-12 phenomena/provide evidence or design solutions to describe and/or support claims and/or explanations. ● present information, findings, and supporting evidence such that listeners can follow the line of reasoning, and the organization, development, and style are appropriate to task, purpose, and audience. ● create appropriate graphical representations of data. ● use spoken, written, & visual language and communicate their discoveries in ways that suit their purpose & audience. ● translate quantitative or technical information expressed in words into a visual form. ● provide an accurate summary of the source distinct from prior knowledge or opinions. ● cite specific textual evidence to support analysis. (2005); ercegovac (2012); guler et al. (2016); callingham et al. (2017); cohen et al. (2017); utomo (2021) note. * = standards relevant to more than one competency. a few competencies have overlapping standards with each other because, in the standards, different competencies are clustered in one sentence. 41 iassist quarterly winter/spring 2010 the 37th international association for social science information services and technology (iassist) annual conference will be hosted by simon fraser university and university of british columbia and will be held in vancouver, canada, may 31 june 3, 2011. the theme of this year's conference is data science professionals: a global community of sharing. social science benefits from professional practices that enable sharing of data, information, and knowledge with a global community. this theme is intended to stimulate discussions about ways in which sharing data, information, and knowledge can contribute to research and to professional practices that enable scientific progress. submissions are encouraged that offer improvements for creating, documenting, submitting, describing, disseminating, and preserving scientific research data. we seek submissions on the theme outlined above, and encourage conference participants to propose papers and sessions that would be of interest to themselves and other attendees. below is a sample of possible topics that may be considered: * open data and the development of knowledge communities data sharing, access and management in the future identifying and reducing barriers to data sharingissues of confidentiality in sharing sharing professional data science skills, knowledge, & techniques within &across discipline citation of research data and persistent identifiers metadata facilitating data sharing emerging research infrastructures and data sharing new data partnerships in knowledge communities sharing resources and data through social networks identifying user needs and customizing data services to meet the needs the evolving data librarian profession data science practices that support global use and understanding of research data open (linked) data and digital repositories preservation for sharing, recovering data for contemporary use proposals on other topics related to the conference theme will be considered too. papers will be selected from a wide range of subjects to ensure a broad balance of topics. the program committee welcomes proposals for: individual presentations (typically 15-20 minutes) sessions, which could take a variety of formats (e.g. a set of three or four presentations, a discussion panel, a discussion with the audience, etc.) posters/demonstrations for the poster session workshops (pre-conference workshops that blend lecture and hands-on instruction). [note: a separate call for workshops is forthcoming.] iassist 2011 call for papers 42 iassist quarterly winter/spring 2010 iassist 2011 again this year we would be interested in receiving submissions for presentations in formats successfully introduced in last year's conference, in particula pecha kucha (a presentation of 20 slides shown for 20 seconds each, with a heavy emphasis on visual content). round table discussions (as these are likely to have limited spaces, an explanation of how the discussion will be shared with the wider group should form part of the proposal). session formats are not limited to these ideas and session organizers are welcome to suggest other formats. proposals for complete sessions should list the organizer or moderator and possible participants; the session organizer will be responsible for securing both session participants and a chair. all submissions should include the proposed title and an abstract no longer than 200 words. longer abstracts will be returned to be shortened before being considered. please note that all presenters are required to register and pay the registration fee for the conference; registration for individual days will be available. a web form for submission of proposals will be available on the conference web site on october 18, 2010. deadline for submission: november 29, 2010. notification of acceptance: february 3, 2011. please note that the conference program committee may not be able to accept all proposals. conference papers will be considered for articles in the iassist quarterly. this applies to individual papers as well as selections of papers from sessions that could form special issues of the iassist quarterly. for more information about the conference, including travel and accommodation, see the conference web site at: http://www.rdl.sfu.ca/iassist/ online conference registration is scheduled to open in early february, 2011. make plans to come to vancouver for iassist 2011 : 31 may 3 june 2011! . questions may be sent to the program planning co-chairs, bob downs, ernie boyko and tuomas j. alateräi at iassist2011@gmail.com vol31_1.indd 14 iassist quarterly spring 2007 by lynn woolfrey* the role of survey data archives survey data archives perform the dual functions of facilitating data sharing and assisting in safeguarding the quality of the data shared. dedicated archives for the sharing of social survey data evolved as a result of the development of quantitative research within the social sciences, involving the collection of social statistics through surveys. quantitative social research and the establishment of survey data archives to house and disseminate the output of this research were in turn aided by the invention of the computer and its increasingly sophisticated technologies. these two roles, safeguarding data quality and enabling data sharing, including providing training in survey research and data analysis, are integral to the work of survey data archives. worldwide, survey data archives extend the work of national statistical agencies through these functions. a network of survey data archives in asia, australasia, western and eastern europe and north america, facilitates research cooperation and data sharing on a regional and global scale. this is not the case in africa. however, survey data archives may be the most appropriate facilities for ensuring the survival and usage of survey data in africa, and an african survey data archive network could optimize technological developments for the provision of access to african survey data. this article looks at the possibility of establishing such a survey data archive network in africa, advantages that would accrue from this, and obstacles to such a development. the value of data sharing for policy formulation findings in data archiving literature confirm the importance of social survey data as a policy tool for social and economic development1. the data from these studies provides vital information on the state of nations. government agencies now rely on social statistics generated by surveys conducted by state and private institutions as a basis for policy formulation, particularly related to economic growth and poverty alleviation. survey data archives, which began as storage facilities for survey datasets, have gradually taken on a more active role in reprocessing and disseminating these data in the interests of research and policy-making. the value of the survey data archive’s role as curator of quantitative national data in africa cannot be overestimated, as much african survey research is lost to posterity through lack of data management and preservation structures for data products2. economic and social development in african countries is linked to their ability to promote applied research. governments need to be aware of the utility of information and their inability to act effectively without it. these countries may be left behind in the knowledge economy if they are unable to tap into global knowledge using modern technology, generate indigenous knowledge and disseminate this knowledge to ensure its practical use3. the establishment of national survey data archives in african countries can be seen as part of the infrastructure enabling these countries to utilize knowledge that is produced locally and through international survey research. survey data archives in africa could act as intermediaries for data sharing from private and official survey research, and ensure the survival and quality of african survey data for re-use4. failure to support secondary usage of survey data in africa could mean this resource will be underused, and opportunities for national growth and for improving the quality of the survey research process through reexamination of data will be lost5. data archives and data quality in africa as the value of survey data is realised, demands are concurrently being made to ensure the high quality of survey products on which development policies are based. government agencies, academic institutions and the private sector have come to expect optimal usage of survey data, particularly in relation to the relevance, consistency and comparability of survey datasets. here again, survey data archives have a role to play. for example, the datafirst archive in south africa has begun to provide a valueadded service with regard to national statistical data, connecting the national statistical agency with data users, such as academics and government policy-makers. this archive provides feedback to statistics south africa on the relevance and usability of its datasets, and notifies data users on problems with and changes to national datasets6. data sharing organisations in africa in africa, national statistical agencies have limited financial a survey data archive network in africa possibilities and practicalities iassist quarterly spring 2007 15 and staff resources for survey administration and analysis to feed into national policy. consequently they do not actively share the microdata from their surveys. however, with the growing recognition that statistical data is an important national resource for scientific investigation and sound national decision-making, international organisations such as the world bank, the united nations and the international monetary fund, have initiated programs to stimulate production of african national statistics and their usage to formulate national policies. these have included the world bank’s african household survey capability programme in 1973 and the addis ababa plan of action for statistical development in africa (aapa) in the 1990s7. the world bank funds the international household survey network, launched in 2004 to promote survey research in developing countries8. the un initiated african symposium on statistical development (assd), was convened in 2006 as a platform for the exchange of information on statistics and technical assistance with the collection and dissemination of african census data. at this symposium, delegates pledged to promote “knowledge management in statistics” on the continent9. in 2007, africa’s national statistics agencies, in collaboration with the un economic commission for africa and the african development bank, recommended a draft charter on statistics for the continent, acknowledging the vital role quantitative data plays in development on the continent10. support for national statistical agencies on the continent is vital. however, establishment of a network of survey data archives in africa would be a complementary development, ensuring independent quality-monitoring of national datasets, and national and regional data sharing. in africa, three survey data archives have been established for these purposes. two of these are in south africa – the south african data archive (sada), and datafirst, an archive at the university of cape town. sada is a government repository, which was established in 1993 with the assistance of consultants from the danish data archive11, and is now part of the national research foundation. datafirst is a grant-receiving initiative with the dual roles of facilitating data sharing for research, and training students in quantitative analysis of survey data. a third nascent survey data archive has been established within the ethiopia central statistical agency12. a development towards establishing data archives in other african countries, and possibly facilitating linkages among them, occurred at the 2007 conference of the international association for social science information service and technology (iassist). delegates from africa and those interested in data management in africa met and decided to create an electronic mailing list for sharing ideas. members of this group included staff from datafirst (south africa) and representatives from statistical agencies in cameroon, ethiopia, the gambia, ghana, mozambique, niger and uganda (the listserve address is iassist-africa@lists.carleton.edu ). group activities thus far involve exchanging ideas on data management in africa, and discussions regarding collaboration on staff training. the demand for data in africa developments towards the establishment of survey data archives in africa reflect an increasing regional demand for high-quality statistics on which to base social and economic policy. however, it is difficult to quantify this demand to support the value of establishing data management organisations in the region. examination of statistics maintained by survey data archives in some countries can indicate the demand for secondary data13. these include statistics of requests for data, derived from online data usage forms, and records of dataset preparation and distribution. additionally, one can judge the extent of usage of shared data by assessing the number of documents produced from the holdings of national survey data archives. many survey data archives have web-based bibliographies of publications that are based on their survey data. from these we can ascertain at least the quantity of works produced with shared data. however, frequency counts cannot measure the quality of the academic output from secondary data use14. systematic statistics on secondary use of survey data in africa have not been compiled. as the south african data archive (sada) does not publish statistics of dataset distribution, judging demand for datasets in their collection is hindered. however, one significant case in south africa demonstrates the value academics and policymakers find in re-using south african survey data. the south african living standards measurement survey (lsms) was conducted in south africa by the world bank and the university of cape town. this was the first household survey conducted in south africa, with the final dataset released in 1995. the data have been heavily used since its release, both for academic and teaching purposes. many of the reports emanating from secondary analysis of the data from this survey have been used in policy-making15. by january 2007, the south african lsms data had spawned 78 monographs and 39 journal articles, by international and local academics, published in south africa and abroad. forty-six unpublished papers were produced using the data. these publications represent work commissioned by government bodies and international funding agencies, readers for academic courses, and publications for local and international journals and conferences. in addition, the dataset is used pedagogically; for example as part of an annual course on survey analysis at the university of cape town16. mailto:iassist-africa@lists.carleton.edu 16 iassist quarterly spring 2007 obstacles to the establishment of a survey data archive network in africa the lack of country infrastructures for establishing facilities for data archiving and regional data sharing is a significant obstacle to the development of a survey data archive network in africa 17. a further obstacle is the lack of an educated and skilled workforce to appreciate the advantages of global and indigenous information and to utilise new technologies to access this information. thus, in africa there has been poor archiving of survey microdata, resulting in the loss of valuable data in many countries. the current situation in africa is that the results of african social surveys and censuses are often archived outside the continent. these data are archived and disseminated either by foreign research organizations conducting survey research in africa or by international organizations funding development projects in the region. historical obstacles obstacles to effective survey data management in africa mirror those experienced by the newly established survey data archives in eastern europe. these problems include the lack of a data-sharing culture, making regional cooperation in survey research difficult18. historical animosity between african governments and academia has hampered the use of applied research for policy development. early on african academics saw the value of research for economic development. in 1981, the southern african development research association (sadra) emphasised that “...research in the region [should] provide a necessary base for the policy choices governments must make to promote development19” however, the attitude of many government officials remains uncooperative; even basic statistical data collected by governments, such as census information, has not been easily available to researchers20. african academics have also advocated regional cooperation in africa for research and dissemination of results, given the scarce resources in the region21. historically, however, many african governments and research bodies have maintained relationships primarily with their colonizing powers, rather than attempting to forge regional ties22. data sharing is mutually beneficial for government statisticians and the academic community, both regionally and nationally. the public must perceive academic research based on national survey data as an extension and enhancement of official statistical information. this will increase their trust in the official data and in the national agencies responsible for its compilation. academic research and policy formation will benefit through the process official data-collection23. networked survey data archives could facilitate this to aid national and regional development in africa. lack of resources here also, the situation in africa reflects the problems of data sharing in eastern europe. funding for the establishment and upkeep of survey data archives is meagre in both africa and eastern europe, and there is a paucity of skilled staff to manage data archives. in contrast to eastern europe, in africa the situation is further exacerbated by comparatively low levels of education among the general population, with the resultant lack of a critical mass of researchers to promote data sharing24. this deficiency also means african researchers are unable take advantage of the new information and communications technology for development25. this includes a paucity of skilled personnel to manage any technology infrastructure. government policy-makers on the continent are also unskilled in using survey and census results26. technical and logistical obstacles difficulties also arise in the identification and location of suitable data for secondary analysis, as information about survey data in africa may not be readily available to researchers. technological infrastructure is needed to facilitate the archiving and sharing of research data in african countries. internet access is limited, and internet users are few; africans comprise just three percent of internet users worldwide27. this is partially due to a low level of internet research skills among african researchers. however this is exacerbated by the cost of accessing information in these countries, which is often higher than in the rest of the world. government regulatory mechanisms for telecommunications often make access to internet resources expensive and inefficient. computer equipment is usually more expensive in africa than in countries of the developed world28. to compound this problem african countries often have bandwidth problems29. furthermore, african countries need a reliable electricity source upon which internet infrastructures depend, and this is not always available. even where internet access is available, information on african surveys is scarce. a search of websites of african statistical agencies revealed that none of these national agencies provide a comprehensive list of the surveys and censuses they conduct. requests for a list of surveys are met by referrals to online catalogues of publications. it would seem that these organizations are still in the pre-technology mode concerning data products. that is, planning for budgets and training tends to focus on sound methods of data collection and analysis, but does not support data sharing. in these statistical agencies, the end product of national censuses and sample surveys is envisioned (and presented to the public) as a series of reports on the findings. the production of microdata files from survey research to enable local and international researchers to conduct secondary analysis is not an integral part of the process. in south africa, the national statistical iassist quarterly spring 2007 17 agency does not publish a list of their surveys on their website. however, online analysis of some of their surveys is now possible through their nesstar server. the survey data archives in south africa provide information on south african survey datasets, and both sada and datafirst post their dataset holdings on their websites30. data producers in africa are not yet harnessing new technologies in the interest of data sharing for development. however, these organisations may benefit from entering the information technology revolution at a late stage. decreasing computer and communication costs allow “technology leapfrogging” in these countries, because they do not have to bear the costs of technology development and experimentation. rather, these countries can begin to participate in a computer and telecommunication market with versatile and less expensive products for data sharing31. any survey data archives established in the region can also benefit from developments in data management worldwide. language issues language is another barrier to identifying and sharing african datasets. any regional coordination in data sharing in africa must account for language difficulties. unlike the countries of latin america, which have spanish as a common language (except brazil and some parts of the west indies), africa has three languages for the purpose of research – english, french and portuguese. arabic is also spoken in some african countries. ideally, this and the major regional languages should be considered in providing access to african data catalogues. a possible solution to these logistical problems could be the creation of an africa-wide data portal such as the one developed by the multilingual access to data infrastructures of the european research area (madiera) project. the aim of this european commission-funded project is to develop the european social science research infrastructure for data, and these efforts launched the madiera portal in 2006. this portal is a multilingual front-end access to the data holdings of thirteen social science survey data archives in europe, which the researcher may use in any of nine languages32. the creation of an african data portal would necessitate an audit of existing african datasets to provide information on these for the portal, and links to the relevant data suppliers. this could be one of the first tasks of a survey data archive network established on the continent. data ownership issues in africa other barriers to sharing survey data through an african survey data archive network include the dissuading motivations inherent in academia. this is not unique to african research, but is particularly relevant here because much of the survey work conducted in africa is undertaken or at least co-produced by foreign organisations. principal investigators from these organisations wish to have exclusive access to data to maximise their academic advantage. they consider themselves “owners” of the final data product. because much of this survey data is in the private domain, no institutional or peer pressure exists to expand data availability. foreign researchers conducting surveys in africa have little interest in devoting resources to preparing datasets for the benefit of academic competitors. thus, they maintain exclusive rights to data collected in africa, at least for extensive embargo periods. this often results in the survey microdata from these studies being unavailable to researchers in these african countries. for example, the afrobarometer project has used their superior funding and technical resources to harvest data from african countries, but does not place this data in the public domain for at least two years. african researchers not involved in the project are unable to utilise this data for research and policy development during the embargo period33. in africa, networks of survey data archives could counter the dissuading factors inherent in the academic reward system. in south africa, sada and datafirst have assisted in bringing the issue of data sharing to the attention of funding agencies, and have received funding support because of their perceived nature as valuable resources for academic advancement, as repositories of empirical information for development, and for their role in promoting scarce quantitative skills in developing countries34. hopefully, african funding agencies will in the future require, through research grant stipulations, that grantees prepare their data for secondary use and release it in a timely manner. this positive trend is occurring in some countries outside of africa, facilitated by their respective survey data archives35. confidentiality issues protection of confidential data is one of the fundamental principals espoused by government statistical agencies, internationally and in africa, as a necessary condition to ensure the trust of respondents, and therefore the accuracy and reliability of data collected36. other organisations conducting surveys are also obligated to adhere to these principles. however, african researchers are increasingly requesting access to unit records from official surveys, and sharing records could compromise respondent confidentiality. these records offer great value to researchers needing empirical information. an important recent finding regarding statistical data is that a truly accurate picture of the national economy of a country cannot be obtained from analysis of only aggregate data37. thus, access to unit records becomes important for research that impacts government policy. other research needs call for a compromise between data access and data protection, for example, when information that links responses to individuals is needed for secondary analysis, as with 18 iassist quarterly spring 2007 epidemiological or panel studies38. restrictions on the use of these linkages can hamper research. as formal structures for data sharing, survey data archives are able to support methods of protecting the confidentiality of respondent information. archive staff can accomplish this by assisting in making the datasets anonymous. they can also facilitate wider usage of restricted access datasets by providing secure facilities for use by researchers, as in some parts of the world39. thus they can play an important role in mitigating the confidentiality concerns of principal investigators and government statistical agencies in africa. africa-wide comparative research is possible using linked datasets from different countries, such as the demographic and health data collected for the measuredhs project. however, ensuring data confidentiality is compounded in comparative and cross-national research. issues of regional concern in african research could be effectively handled by a continental network of survey data archives. these networked archives could manage and monitor adherence to ethical principals in complex cross-national research, while standardizing data and access procedures40. building on existing infrastructure those interested in african survey data archives must utilise existing structures, including current african networks, to provide the necessary resources for establishing a continental survey data archive network. data exchange structures in africa include government agencies and policy-making and regulatory bodies, intergovernmental scientific organizations, publicly-funded research institutions such as universities, publicly-funded data management institutions, such as data centres and libraries, and national and international ngos41. the wealthy institutions of higher education that support survey data archives in north america are scarce on the african continent. however, african countries may be able to obtain funding from international grant agencies to support the development of survey data archives. this assistance would be similar to the structures that support nascent survey data archives in eastern europe42. international funding agencies have been willing to support data sharing for improved research and governance in africa, for example the world bank’s sponsoring of the development and use of their microdata management toolkit for processing and distributing census microdata on the continent43. international funding has also directly enabled the establishment of the datafirst survey data archive in south africa44. however, support from international funding agencies may not facilitate a coordinated effort towards data-sharing on the continent. the development of a continental network for sharing african social survey data could rather be facilitated by building on the existing government statistical agency structures, and promoting collaborative efforts among them. this could mean physically situating african survey data archives within the government statistical offices, or as departments at african universities. in western europe and north america, most official statistical agencies play a supporting role for data archive networks. generally, they are one of the many suppliers of survey data to survey data archives. in africa, and in some countries in eastern europe, government statistics departments have a less clearly defined role for regional data sharing. they are historically the main producers and suppliers of survey information in countries with poorly developed research structures. for example, in lithuania the main holder of empirical data is the lithuanian department of statistics45. it may therefore be practical to establish survey data archives as semi-autonomous units within africa statistical agencies. such developments are already underway in the region. one example is the ethiopian central statistical agency, where a survey data archive has been established and is producing national micro-datasets for secondary analysis, using the world bank’s microdata management toolkit46. establishing these archives in the national statistical agencies would allow these organizations to utilize their existing data management infrastructures, and could improve official data management practices when data users interact directly with national data suppliers and provide regular feedback. however, situating african survey data archives in national statistical agencies would circumvent independent quality control of the data produced. quality monitoring of datasets is a vital role played by survey data archives. thus any archive administered and funded by a national statistical agency would have its objectivity compromised. these survey data archives would need to collect all survey datasets available nationally, both publicly-funded and commercial, to properly fulfil the role of national survey data archives. statistics offices worldwide have traditionally only archived and disseminated their own survey datasets. they may be resistant to house datasets from independent surveys, or unprepared to promote secondary usage because they view them as inferior to their own survey products. conversely, other data producers may be reluctant to deposit their datasets with a government organization, doubting their objectivity and ability to assign their data equal weight to potential data users. another problem is that situating survey data archives in a government department may lead to the perception among researchers that data produced has a government bias. this is particularly valid in an african context, with a history of intolerance between government and academia. this may restrict usage of datasets from archives established within national statistical agencies. establishing survey data archives within african iassist quarterly spring 2007 19 universities could counter fears of bias in the data produced47. another advantage of housing these facilities at institutions of higher learning is that african academics, the main users of survey data, could easily access survey microdata. this will also ensure that there is dialogue between these data users and government agencies collecting and supplying the survey data, concerning the requirements for secondary analysis. finding indigenous solutions wherever they are situated within african countries, survey data archives can play a vital role in regional data sharing for development. quantitative data from surveys can be used for “problem-solving, policy-oriented research” in africa, or applied research for solving the development problems in african countries48. re-use of these data can spur creation of indigenous knowledge for growth. best practices from survey data archives worldwide can be utilised to develop an effective data sharing network in africa, but one should avoid simply imitating unapplicable practices and trends in creating this research and policy resource. existing data support projects funded by international organisations should be examined critically to ensure they suit local requirements. technological innovations can only provide advantages to african countries where skills are developed and technical resources made available. a third necessity in ensuring innovations are optimized is establishing institutions that connect data producers to data users. the development of a network of african survey data archives on the continent could fulfil this condition, and act as both a catalyst for the production of high-quality statistical data, and a facility for sharing these data for practical use. * lynn woolfrey. lynn.woolfrey@uct.ac.za. lynn woolfrey is data manager at the datafirst survey data archive, university of cape town. references 1 mwase, n. 1986. social science research in eastern and southern africa. international social science journal 38(1):145 2 personal communication with professor robert mccaa, ipums international, 2006 3 marshalling technology for development: proceedings of a symposium. 1995. washington: national academies press:19, 61-62 4 [sada] south african data archive website, 2007. [online]. available: http://www.nrf.ac.za/sada 5 musgrave, s. 2003. the metadata dynamic: ensuring data has a long and healthy life. association for survey computing. [online]. available: http://www.essex.ac.uk/ hhs/staff/musgrave/ascmetadata(revised)pdf 6email communication between datafirst staff and the staff of statistics south africa 7 kiregvera, b. 2001. novel statistical capacity building initiatives: addis ababa plan of action and paris21. international statistical institute. [online] available: http:// isi.cbs.nl:1-2 8 [ihsn] the world bank international household survey network website. [online]. available: http://www.internati onalsurveynetwork.org 9 african symposium on statistical development (assd) website, 2007 10 african charter paves way for professional data crunching across the continent, article in the business report, 21 june, 2007 11 lesaoana, m.a. 1997. data archiving in africa: the south african experience. iassist quarterly volume 21 number 1:4-7 12 mudesir seid, y. 2004. implementing a national data archive in ethiopia: challenges and experiences. paper presented at the 2006 iassist conference. may 2006, ann arbor, michigan. [online]. available: http://www. iassist.org. 13 clubb in feinberg, e., martin m..e. and l. m.l. straf. eds. 1985. sharing research data. washington: national academies press.49-50 14 boruch in feinberg, e., martin m..e. and l. m.l. straf. eds. 1985. sharing research data. washington: national academies press:114 15 wilson, f. and horner, d. 1995. lessons from the project for statistics on living standards and development: the south african story. (unpublished paper). 16 datafirst statistics, 2007 17 marshalling technology for development: 31 18 hausstein, b. and de guchteneire, p. 2002. social science data archives in eastern europe: results, potentials and prospects of the archival development. bergisen gladbach: e. ferger verlag:59 19roma declaration on research and development in southern africa, 1981 in mwase:139 (repeat) http://www.nrf.ac.za/sada http://www.essex.ac.uk/hhs/staff/musgrave/ascmetadata(revised)pdf http://www.essex.ac.uk/hhs/staff/musgrave/ascmetadata(revised)pdf http://www.internationalsurveynetwork.org http://www.internationalsurveynetwork.org http://www.iassist.org http://www.iassist.org 20 iassist quarterly spring 2007 20 mwase:139 21 mwase:145 22 regionaliza�on of social sciences in la�n america, asia and africa. 1973. interna�onal social science journal 25(4):557:559 23 cook in statistical confidentiality and access to microdata. proceedings of the seminar session of the 2003 conference of european statisticians. 2003. new york: united nations. [online]. available: http://www.eurostat.eu 24 hausstein and de guchteneire:59-60 25 bits of power: issues in global access to scientific data. 1997. washington: national academies press:41 26 assd website, 2007 27 internet world stats website, 2007 [online]. available: http://www.internetworldstats.com 28 adam, l. electronic networking for the research community in ethiopia in bridge builders: african experiences with information and communication technology. 1996. washington: national academies press 29 bits of power:41-42 30 sada website, 2007[online]. available: http://www.nrf. ac.za/sada and datafirst website, 2007. [online]. available: http://www.datafirst.uct.ac.za 31 bits of power: 27 32 [madiera] multilingual access to data infrastructures of the european research area website, 2007. [online]. available http://www.madiera.net 33 afrobarometer website, 2007. [online]. available: http:// www.afrobarometer.org 34 personal communication with matthew welch, director of datafirst, 2007 35 guide to social science data preparation and archiving: best practice throughout the data life cycle. 2005. 3rd edition. ann arbor: inter-university consortium for political and social research 36 cook in statistical confidentiality and access to microdata. proceedings of the seminar session of the 2003 conference of european statisticians. 2003. new york: united nations. [online]. available: http://www.eurostat. eu:1 37 lane in statistical confidentiality and access to microdata. proceedings of the seminar session of the 2003 conference of european statisticians. 2003. new york: united nations. [online]. available: http://www.eurostat. eu:12 38 feinberg, e., martin m..e. and m.l. straf. eds. 1985. sharing research data. washington: ational academies press:20 39 dunne, s. and austin, e.w. 1998. protecting confidentiality in archival data resources. icpsr bulletin 19(1):1-6 40 taylor, m. 1994. ethical considerations in european cross-national research. international social science journal 46(4):523-532. 41 bits of power:41-42 42 hausstein and de guchteneire:58 43ihsn website 44 personal communication with matthew welch, director of datafirst, 2007 45 hausstein and de guchteneire 46 mudesir seid, 1-5 47 lesaoana, m.a. 1997. data archiving in africa: the south african experience. iassist quarterly volume 21 number 1:6 48 mwase:139 http://www.eurostat.eu http://www.internetworldstats.com http://www.nrf.ac.za/sada http://www.nrf.ac.za/sada http://www.datafirst.uct.ac.za http://www.madiera.net http://www.afrobarometer.org http://www.afrobarometer.org http://www.eurostat.eu http://www.eurostat.eu http://www.eurostat.eu http://www.eurostat.eu iassist quarterly winter 2011 5 iassist quarterly editor’s notes depositing data and placing it on the market this issue (volume 35-4, 2011) of the iassist quarterly (iq) is the last of the 2011 volume. many iassist members are now getting ready and looking forward to this year’s conference. probably it will turn out to be “the best ever!” and with interesting papers for the coming issues of the iq. this issue focuses on aspects of the depositing of data, the guidelines, regulations and formalities involved, and also on the connections between the deposit and the re-use of data. the paper entitled “examination of data deposit practices in repositories with the oais model: social science context” is written by ayoung yoon and helen tibbo from the school of information and library science at university of north carolina at chapel hill. the paper examines the requirements for depositing data in selected data repositories by analyzing the forms and guidelines for such deposits. the open archival information system (oais) an iso standard is used as a framework for this examination. the authors emphasize and reference others in arguing how “research data need to be available for use beyond the purposes for which they were initially collected, to make the results of studies using publicly funded data available to the public, to enable others to ask new questions of extant data and advance solutions for complex human problems, to advance the state of science, to reproduce research, and to expand the instruments and products of research to new communities”. the authors use a method of content analysis in examining the requirements that exist within depositors’ guidelines and deposit forms. the analysis is based upon 14 documents from 16 social science data repositories. the analysis is not looking into the actual content but registering the “required”, “optional” and “not mentioned” requirements. it turned out that the documents varied significantly, including such surprises as not all repositories asked for the title or a description of the data study. the second paper is authored by cristina ribeiro and maria eugénia matos fernandes from the university of porto (universidade do porto). as the title outlines “ data curation at u.porto: identifying current practices across disciplinary domains” we are now turning from comparing depositing at different repositories to the differences in data curation between different disciplines. the study has involved researchers collecting their views on data curation and data. the article includes a presentation of the university of porto and the paper draws information from a local information system called sigarra (information system for the aggregated management of resources and academic records). this system supports authors in making their intellectual output centrally available as they sign contracts with publishers, so they maintain the right to self-archive their work in institutional open repositories that include a data repository prototype. the interviews with the researchers found that the design of a data repository should be determined by researchers’ needs. from curation and depositing of data, the authors laurence horton and alexia katsanidou from gesis (gesis-leibniz institute for the social sciences in cologne) take a further step in their paper “purposing your survey: archives as a market regulator, or how can archives connect supply and demand?”. the authors start with the statement that researchers who are data creators and researchers who are data re-users have different needs and that archives mediate between them. the paper outlines the gesis plan to create a research data management and archive training centre for the european research area. in their paper the authors give examples of how the re-use of data now has strong political support. the european commission has committed itself to an open data policy and this is accompanied by statements like “taxpayers have already paid for this information, the least we can do is give it back to those who want to use it in new ways…” and “your data is worth more if you give it away”. the arguments presented for data preservation and sharing are the technological and financial benefits. there are, however, continued obstacles that prevent data sharing viewed from the supply side of the social sciences. there are restrictions through law and ethics, and also a lack of incentives to share data. as a publisher the iassist quarterly supports the need to have institutional open repositories as mentioned in the second paper. we also support “deep links” where you link directly to your paper published in the iq. articles for the iq are always very welcome. they can be papers from iassist conferences or other conferences and workshops, from local presentations or papers especially written for the iq. if you don’t have anything to offer right now, then please prepare yourself for a future iassist conference and start planning for participation in a session there. chairing a conference session with the purpose of aggregating and integrating papers for a special issue iq is much appreciated as the information in the form of an iq issue reaches many more people than the session participants and will be readily available on the iassist website at http://www.iassistdata.org. authors are very welcome to take a look at the instructions and layout: http://iassistdata.org/iq/instructions-authors authors can also contact me via e-mail: kbr@sam.sdu.dk. should you be interested in compiling a special issue for the iq as guest editor(s) i will also be delighted to hear from you. karsten boye rasmussen may 2012 editor microsoft word 49-4-marmier.docx 1/19 marmier, a., stepanovic, s., & mettler, t. (2025). in the data steward’s shoes: an autoethnographic exploration of everyday challenges, iassist quarterly 49(4), pp. 1-19. doi: https://doi.org/10.29173/iq1172 the creative commons-attribution-noncommercial license 4.0 international applies to all works published by iassist quarterly. authors will retain copyright of the work and full publishing rights. in the data steward’s shoes: an autoethnographic exploration of everyday challenges auriane marmier1, stefan stepanovic2 and tobias mettler3 abstract the role of data stewards (dss) in academic institutions has become increasingly complex as research data management (rdm) policies evolve under the pressures of open science, data protection regulations, and funding mandates. this paper examines the challenges dss face through an autoethnographic approach, analysing four cases that highlight tensions between global compliance requirements and researchers' practical needs. the findings illustrate that dss operate in a “buffer zone,” mediating between top-down global imperatives, such as the principles of findability, accessibility, interoperability, and reusability (fair), national data-sharing policies, and legal constraints, as well as bottom-up pressures from researchers prioritising knowledge production, academic freedom, and project-specific requirements. rather than offering generalisations or fixed solutions, this paper provides a practice-based perspective that seeks to open the debate on the current positioning of dss within academic institutions. by highlighting recurring frictions and underexplored issues, it identifies key areas for reflection and improvement, such as integrating dss into institutional decision-making and promoting more flexible, context-sensitive rdm frameworks. this study contributes to a growing conversation on how data stewardship can evolve to better support both regulatory compliance and research innovation. keywords data stewardship, research data management, autoethnography, lived experiences introduction over the past decades, the academic landscape has undergone significant transformations. the proliferation of the internet and rapid technological advancements have led to an unprecedented explosion in global internet traffic and data production, fundamentally changing research processes and priorities. as highlighted by authors (emmott & rison, 2008; hey & trefethen, 2003; peter lyman & hal r. varian, 2003), these developments have reshaped how knowledge is produced, stored, and disseminated. the ability to generate, process, and share large volumes of data has become a defining 1 fors swiss centre of expertise in the social sciences, auriane.marmier@fors.unil.ch 2 university of lausanne, stefan.stepanovic@unil.ch 3 university of lausanne, tobias.mettler@unil.ch 2/19 marmier, a., stepanovic, s., & mettler, t. (2025). in the data steward’s shoes: an autoethnographic exploration of everyday challenges, iassist quarterly 49(4), pp. 1-19. doi: https://doi.org/10.29173/iq1172 characteristic of contemporary research, necessitating new methodologies and tools to address the challenges of scale, complexity, and interdisciplinary collaboration. the ethical and legal dimensions of research have also evolved significantly. historical ethical failures, such as the tuskegee syphilis study, the milgram experiment, and the stanford prison study, have highlighted the need for strict ethical standards and robust regulatory frameworks to safeguard research participants and their data (smale et al., 2020). these lessons have notably led to the establishment of regulations such as the general data protection regulation (gdpr) in europe and the new federal data protection act (nfdpa) in switzerland (the context of this study) which have profoundly influenced researchers’ practices (european union, 2016; the swiss confederation, 2023). such regulations impose rigorous requirements for researchers regarding data protection, ensuring privacy and security while accommodating the growing demand for data sharing and collaboration. simultaneously, the open science movement has gained attention, advocating for the democratisation of research and the transparency of scientific findings. notably, as lord & macdonald (2003) emphasise, one of the main missions is to make research outputs more findable, accessible, interoperable, and reusable (fair). these principles have emerged as a cornerstone of the open data movement, fostering a culture of openness while addressing challenges in data management (wilkinson et al., 2016). however, balancing data protection with the principles of openness presents significant challenges for both researchers and institutions. to address these challenges, universities and higher education institutions have developed comprehensive research data management programs that provide structured approaches for handling data throughout its lifecycle. the growing demand for effective data management has, in turn, led to the establishment of the role of dss. although terminology varies across the literature—with dss sometimes referred to as data custodians or research librarians, among others (rousi et al., 2024; unece, 2022; verhulst, 2025)—there is general agreement that they are professionals who advise researchers on good data management practices and help ensure the quality of data assets (mons, 2018; peng et al., 2016; whyte et al., 2023). they often act as an intermediary between those who generate data and those who use it (rosenbaum, 2010). they are domain experts tasked with overseeing data collection, storage, sharing, and preservation. their role is critical in ensuring alignment with institutional policies and regulatory frameworks while supporting researchers in adopting best practices for data management (peng et al., 2018; perrier et al., 2017; plotkin, 2020). the establishment of data stewardship roles has sparked interest in understanding the profession’s characteristics (mons, 2018; wendelborn et al., 2023) and initiatives to develop their competencies (data stewardship, curricula and career paths eosc association, 2023; oladipo et al., 2022; wildgaard et al., 2020). the specific structures, processes, tasks, and requirements of dss across various domains have been closely examined (arend et al., 2022; peng et al., 2016; whyte et al., 2023; york et al., 2018) and captured in diverse models describing the role of dss (edmunds et al., 2016; peng et al., 2015). considerable focus has been placed on how academic libraries structure their service models (cox et al., 2019; hackett & kim, 2024; pinfield et al., 2014; tenopir et al., 2012) and on how data stewardship works in higher education (rousi et al., 2024). 3/19 marmier, a., stepanovic, s., & mettler, t. (2025). in the data steward’s shoes: an autoethnographic exploration of everyday challenges, iassist quarterly 49(4), pp. 1-19. doi: https://doi.org/10.29173/iq1172 however, while these studies provide valuable insights into institutional frameworks, they often lack an empirical grounding in the lived experiences of dss. this gap limits our understanding of how global imperatives translate into concrete practices at the local level and how these dynamics shape the ds role. given that dss operate in a complex environment, interacting with diverse stakeholders, including researchers, administrators, and it professionals (perrier et al., 2017), a closer examination of their daily experiences is essential for improving both practice and policy. to do so, the paper adopts an autoethnographic approach grounded in the daily practice of a ds embedded within a researchintensive university. through four case studies drawn from lived experience, this paper illustrates recurring challenges that many dss will recognise: (1) navigating conflicting data storage policies and external partner contracts; (2) preserving research data after the unexpected departure of a principal investigator; (3) managing legal risks related to insecure data collection tools; and (4) addressing the unintended consequences of requiring data management plans for infrastructure access. the aim of this paper is not to present a set of universal best practices, but to offer a grounded, practice-based analysis of the tensions encountered by a ds in their work. because projects and situations are highly specific, ready-made solutions are rarely applicable. instead, the analysis aims to highlight the tensions experienced by the author and consistently echoed in discussions with other dss at both institutional and national levels. in doing so, the intention is therefore to open a practical dialogue on how to evolve data stewardship roles and frameworks to better support both compliance and research realities. the paper begins by describing the autoethnographic approach employed in the study. we then present results from the ds/co-author’s experience, highlighting practical challenges and institutional pressures. to reflect the personal and contextual nature of the autoethnographic approach, this section, which describes lived experiences and case studies, is written in the first person (i’) by the coauthor responsible for the data collection. this stylistic choice emphasises the subjective, reflective nature of the study. the final part of this study offers a comprehensive discussion of the challenges encountered. unlike the section on the ds’s experience, the rest of the paper uses the collective ‘we’ to emphasise the collaborative nature of the research and the joint interpretation of the findings. methodology this study adopts a practice-oriented autoethnographic approach based on the professional experience of one of the co-authors, who has worked as a data steward in a research-intensive swiss university since the formal introduction of the role in 2023. autoethnography offers a structured way to reflect on and analyse one’s own lived experience to shed light on broader institutional, cultural, and operational challenges (hammersley & atkinson, 2019; poulos, 2021). it is particularly well-suited to the ds role, which sits at the intersection of policy, infrastructure, and research practice. the paper presents four case studies in which the ds was directly involved as an active support actor. their role included advising researchers, facilitating dialogue with institutional services, proposing alternatives, and advocating for exceptions. while not in a position of formal authority, the ds acted as a mediator and problem-solver, navigating institutional constraints while responding to researchers' needs. these case studies were selected because they exemplify recurring tensions that dss commonly face across institutions and disciplines: navigating legal constraints, ensuring data preservation, complying with technical infrastructure requirements, and managing divergent 4/19 marmier, a., stepanovic, s., & mettler, t. (2025). in the data steward’s shoes: an autoethnographic exploration of everyday challenges, iassist quarterly 49(4), pp. 1-19. doi: https://doi.org/10.29173/iq1172 interpretations of data protection rules. rather than isolated or extraordinary events, they look representative of common dilemmas that arise in everyday data stewardship practice. the cases draw on multiple data sources collected over the course of a year, including a reflective journal (maintained in the form of a logbook during project support), participant observation in institutional meetings and working groups, informal discussions with colleagues, meeting notes, and exchanges with other dss from the university network and the swiss research data support network (srdsn). these peer interactions were crucial in confirming that the described tensions were not unique to the author's personal situation, but were widely recognised by other dss facing comparable challenges. in addition, the ds also analysed internal documents, such as institutional policies, communication records, and guidance materials, to contextualise each situation within the broader data governance environment. this combination of self-reflection, observation, and document analysis ensures that the findings are grounded in concrete, everyday professional practice. the data steward experience my thoughts about the role of ds began long before this study. after only six months as a ds at a swiss university, i began to question the real value of my work. nine months earlier, i had almost completed my dissertation, which focused on the challenges open government data providers face in meeting open access requirements and maximising their data’s value. with this background, i applied for a ds position in the faculty of law, criminal justice and public administration. at this university, the ds function follows a decentralised model: each faculty (i.e. social and political sciences, theology and sciences of religions, business and economics, geosciences and environment, and biology and medicine) hosts its own ds position. dss are hired as staff members rather than faculty, typically holding a phd, and are attached directly to their faculty rather than to a central service. the job description and interview made it clear that my role as a ds will be varied. and so, it was! my main duties included supporting researchers in research data management (rdm) to meet the requirements of different funders and align with open science ideals. this involved ensuring compliance with university and local authority rules. my role also involved advising researchers on best practices within these legal and institutional limits, as well as teaching rdm (e.g., one-on-one sessions, webinars, workshops, and lectures). when my expertise wasn’t enough, i was responsible for finding someone with more knowledge in the relevant field. for instance, if researchers needed specific advice on the ethical aspects of a project involving sex workers, or if they needed help anonymising data from phones confiscated from criminals, my job was to connect them with the right experts. another important aspect of my role was collaborating with other dss from different faculties. we met biweekly to exchange ideas, compare approaches, and discuss institutional challenges. this enabled us to refine our practices, learn from each other’s expertise, and strengthen our collective capacity to support researchers. beyond the university, i was also an active member of the srdsn, which organised biannual meetings and other events where dss from across switzerland shared their experiences and practices. finally, i was also responsible for coordinating with various university central services involved in rdm, including legal services, the it service, and the computing division. the aim was to match researchers’ 5/19 marmier, a., stepanovic, s., & mettler, t. (2025). in the data steward’s shoes: an autoethnographic exploration of everyday challenges, iassist quarterly 49(4), pp. 1-19. doi: https://doi.org/10.29173/iq1172 needs with the university’s services, while ensuring compliance with relevant legislation. i was constantly navigating between two worlds (1) the bureaucratic one, including other dss and central services (legal, it, computing), and (2) the researcher one, including phd students, post-docs, and professors. this job seemed perfect for me. although my dissertation focused on government data, i had a good understanding of the challenges researchers face in managing data, as well as the institutional problems that may arise. my dissertation showed that the biggest challenges in open data often come from managing the data itself, ensuring data protection, meeting legal requirements, addressing storage security, and handling anonymisation. i was excited by the prospect of tackling these issues and discovering research projects in fields different from my own, while also helping researchers achieve their goals. however, six months into the job, i began to question the true meaning of this work. i encountered these challenges while helping researchers manage their data. half of my time was spent discussing their problems, understanding their needs and working through the solutions proposed by the university. the other half was spent trying to adapt solutions proposed by international, national, regional or institutional bodies to researchers’ practices. most of my activities were devoted to discussing with the university’s central services involved in research data management, why the solutions proposed in many cases didn’t fit most projects. these discussions made me realise that most of the actual open research data and rdm recommendations came from a bureaucratic perspective, following a "one size fits all" approach, whereas researchers’ projects depend on the reality on the ground and are therefore unique. at that moment, i realised that i was navigating between two separate environments that would be difficult to connect. when institutional policies clash with field realities drawing on frequent exchanges with other dss, many of the problems encountered in rdm are related to the sensitivity of the data used. working in a faculty that includes researchers in law, public administration, and criminal justice, i dealt with sensitive data issues daily and was drawn into a compelling case study that illustrates the challenges of balancing data storage requirements with the terms and conditions of secondary data reuse. the data to be stored was secondary data, including messages from platforms such as whatsapp, facebook messenger, and twitter, extracted from criminals’ mobile phones. given the sensitivity of the research and the origin of the data, the researchers sought my advice on how to store the data in a way that satisfies the data owners (i.e. the police) and complies with legal and university regulations. according to the applicable national legislation and the guide to technical and organisational data protection measures (federal data protection and information commissioner, 2024), the data controller (i.e. the university) is required to take measures to ensure the security of files and personal data, in particular against loss, destruction and unlawful processing. to comply with these regulations, the university encourages researchers to use the storage infrastructure provided by the computing division. the solution was therefore a cloud storage system for sensitive data. this cloud-based solution ensured compliance with swiss legislation, specifically regarding the location of data centres in europe and the encryption of data, with the university retaining ownership of the encryption keys. 6/19 marmier, a., stepanovic, s., & mettler, t. (2025). in the data steward’s shoes: an autoethnographic exploration of everyday challenges, iassist quarterly 49(4), pp. 1-19. doi: https://doi.org/10.29173/iq1172 from the perspective of open science and funders, recommendations for storing sensitive data only specify compliance with security requirements and stress the accessibility of metadata. however, from the researcher’s point of view, the solutions offered by the regulations and infrastructure created problems. by using the university’s solutions, the researchers were in breach of their contract with the police, which stipulated that the data had to be stored with a high level of security and that cloud storage was not secure enough. in addition, the collaboration between the researchers and the police was not new, and much of the information exchanged was facilitated by networks and old acquaintances. for the researchers, breaking the contract with the police meant not only breaching terms but also risking credibility and trust with the police, with serious implications for future collaboration. moreover, using a cloud solution increased the risk of a data breach, with consequences not only for the criminals but also for the researchers. after several meetings, discussions, and reflections on the risks, the main challenge for the researchers was not compliance with university rules, but rather preventing data leakage. based on this, i proposed a "least bad alternative": an offline computer in a secure office within a secure building, accessible only by authorised personnel with key access. disaster or security breaches were still possible, but considered less likely than an internet-based breach. as this solution required the support of the university’s it and computing division, i brought it to their attention. however, they neither validated nor rejected the proposal. instead, they maintained their position that researchers should use university storage infrastructure. they also noted that in the event of a breach, the researchers would be solely responsible. what i noticed was that there was no room for alternative solutions, no options outside strict compliance. for the university, cloud storage was the only option, even if it meant breaking a contract and taking on more risks. as a result, there was no agreement between the researchers and the university’s research support services. both parties were bound by different constraints, but from the researchers’ point of view, only one was binding—the contract with the police. in the end, the researchers ignored the university’s requirements. attracted by the offline computer solution, they found in-house resources to implement it along with security measures such as encryption and key access. how to preserve years of research when the pi suddenly passed away? another important issue discussed with other dss is the preservation of data when researchers retire or leave a university. this case retraces the latter, involving a professor who had amassed 30 years of data but passed away suddenly. throughout his tenure, he devoted his entire career to gathering data on public administration fields with a special emphasis on the municipal level. one of his signature research projects, launched in 1988, was a global monitoring survey conducted every five to six years. three decades of research material on a single topic can represent a valuable resource. however, for this material to be reusable, it must be accessible. therefore, shortly after the pi passed away and about a year before completing his phd, the doctoral student associated with this survey expressed concern about the future of the material. after learning about open science and storage solutions, he contacted me. his request was simple: with less than a year to finish his thesis, he could no longer preserve the data or continue the survey waves. several institutions were potentially interested in 7/19 marmier, a., stepanovic, s., & mettler, t. (2025). in the data steward’s shoes: an autoethnographic exploration of everyday challenges, iassist quarterly 49(4), pp. 1-19. doi: https://doi.org/10.29173/iq1172 continuing the work, so he sought a solution to store the research materials until future successors emerged. in such contexts, dss from my university usually provide two solutions. the first is to follow open science and funders’ recommendations by publishing the data and research material in repositories without embargoes or other restrictions. in this way, anyone could reuse the data and continue the study. the second is institutional: a long-term storage (lts) solution, i.e. archiving research material on magnetic tapes. the process is simple, but costly, making it suitable for data that does not need to be retrieved regularly, only occasionally. while the first solution may seem attractive, it is not suitable for all data. repositories are not always designed to host 30 years of material, and few can accept personal or sensitive data. after some discussion, i learned that this project contained sensitive data. anonymisation or other depersonalisation techniques could have been organised, but for 30 years of material, this would have been extremely expensive and time-consuming. given the needs of the phd student, i proposed the second solution: institutional lts. the second solution was simpler. according to the university’s policy, the principal investigator (pi) had to contact the it department and request storage space. less than 24 hours later, the space should be open. however, in this situation, there was no pi, and the phd student was not recognised as such. we contacted the relevant service to explain the situation, but after three meetings, it became clear that there was no way around the rule: only a pi could request storage, and without one, it was impossible to store research data on the institution’s servers. i had no idea what to do with this massive amount of data, nor whom to turn to. from the first meetings, i knew it would be challenging, so i contacted everyone who might be able to help, but none proved useful. after a final coffee with the phd student, he told me we had done all we could, and that if no one stepped forward before he completed his phd to take over the survey, it would no longer be his responsibility. i was as discouraged as he was. less than a year later, i heard that a new pi had been appointed and agreed to "take care of this data". american survey software here, i share my experience with researchers who sought my expertise in constructing a questionnaire to collect personal and sensitive data. their study investigated the accessibility of swiss administrative procedures for war refugees. to do this, they designed a questionnaire that included questions about the accessibility of procedures for entering switzerland. the questionnaire was complemented by socio-demographic questions providing insight into respondents’ profiles. despite the researchers’ efforts to limit or eliminate questions relating to sensitive and personal data, certain variables acted as strong indirect identifiers and could lead to re-identification. according to the researchers, if this were to happen, participants could be perceived as betrayers in their home countries, with serious consequences. to minimise the risk, we implemented a secure data processing strategy that covers data collection, storage, and analysis. the issue was the survey software used to collect the data. it was hosted in the united states, a country not on the list of jurisdictions providing adequate data protection under swiss law. according to article 16 of the new federal act on data protection (nfadp) on data transfer abroad, personal data 8/19 marmier, a., stepanovic, s., & mettler, t. (2025). in the data steward’s shoes: an autoethnographic exploration of everyday challenges, iassist quarterly 49(4), pp. 1-19. doi: https://doi.org/10.29173/iq1172 may be transferred if the federal council recognises that the third country ensures adequate protection. if the federal data protection and information commissioner (fdpic) deems that the third country in question does not fulfill the requirements, then the data may still be transmitted, as under current law, provided that adequate data protection is ensured in another manner, such as an international treaty, data protection clauses, which must be submitted to the fdpic in advance, or binding corporate rules (the swiss confederation, 2023). for example, researchers must inform participants of safeguards such as contractual clauses, exceptions by the data controller, or data encryption. however, like many researchers, few were aware of the processing rules (i.e. collection, analysis, storage, exchange, destruction, etc.) for personal and sensitive data. when they contacted me, the questionnaire was already ready for distribution. all questions and adjustments had been finalised. furthermore, it was not feasible to implement clauses regarding encryption, unique key ownership, or the location of data storage, given the company's economic size and the ownership of the chosen solution. at this stage, following discussions with my ds colleagues, the only option was to change the software and use a more secure solution that could provide guarantees. this meant learning new survey software and re-implementing the questions. yet, given the intensity of research competition, the "publish or perish" rules, and the difficulty of changing habits, i realised that adopting more secure but less user-friendly software would be a challenge. having made my recommendations, i no longer had any reason to monitor this project, and the researchers had no reason to keep me informed. as a result, i never received confirmation that they had changed the software to mitigate the risks associated with it and comply with article 16 of the nfadp. requiring a data management plan (dmp) to access the university’s data storage system i cannot speak for other universities, but at mine, a significant portion of dss’ time is devoted to dmps. based on my exchanges with other dss, it appears that at least 50% of their job involve dmp-related activities (reviewing, workshops, corrections, concept explanations, etc.). while i understand their utility, this isn’t always the case for researchers in the social sciences. the concept of data management originated in the 1960s in the sciences to manage research data collection and analysis for aeronautical and engineering projects. over the next two decades, the use of dmps expanded across engineering and scientific disciplines. until the 2000s, they were primarily employed for projects of high technical complexity (smale et al., 2020). however, the 2010s marked a significant shift: funding agencies began increasingly requiring dmps in proposals and evaluations (smale et al., 2020). at the same time, data protection laws, such as the gdpr and the nfadp, became more stringent, and new regulations emerged. so, writing a dmp has become more than just creating a document; it is an opportunity to develop a comprehensive strategy for managing data and ensuring compliance with legal, ethical and security requirements set by funders. yet, in the world of social science research, the dmp is often seen as just another funder's requirement, a roadmap to guide projects. consequently, many researchers view it as another bureaucratic task with little intrinsic value. researchers at the university, and in my faculty in particular, are no exception. yet, they must write a dmp to meet funder requirements and secure 9/19 marmier, a., stepanovic, s., & mettler, t. (2025). in the data steward’s shoes: an autoethnographic exploration of everyday challenges, iassist quarterly 49(4), pp. 1-19. doi: https://doi.org/10.29173/iq1172 grants, but also to follow the ethics committee’s guidelines, request storage space, and (if applicable) plan for publishing data. for storage access, the process is straightforward. the pi must submit a request, providing basic project information (i.e. name, id, estimated storage size, start and end dates, data type—normal, personal, or sensitive), and confirm the existence of a dmp. if a dmp exists, access to storage is granted within 24 hours. if not, access is still provided; however, the dmp must be uploaded later, and the faculty ds is responsible for ensuring this is done. despite the simplicity of this process, together with dss from other faculties, we often encountered two major challenges. firstly, only pis can request storage space, meaning phd students or postdocs cannot do it themselves. this is problematic because pis often delegate administrative tasks to their students. but what happens if the pi is on parental leave, long-term sick leave, sabbatical, holiday, or, as i have experienced, passes away? i faced two such cases (and other dss reported similar situations), and in both, phd students were left without access to secure storage. secondly, the dmp requirement can be a barrier. when already required for grants, the process is straightforward. but if no dmp exists, expecting a researcher to develop one just to access storage is ambitious. creating a dmp is time-consuming, and many researchers prefer free commercial drives (e.g. google drive, dropbox) over institutional requirements. as a result, the responsibility for uploading a dmp falls to the ds, caught between convincing researchers to comply and negotiating with it staff. after a year of trying to reconcile it needs with researcher realities, nothing has changed except my lack of motivation. my suggestion to remove the dmp requirement and instead add targeted questions to the storage request form, which would have provided it personnel with the necessary information to comply with security rules, was not even discussed. discussion the case studies above demonstrate the multifaceted role of the ds, extending beyond rdm advice. in many cases, the ds appears to navigate tensions between researchers and institutional policies, balancing the demands of data access with compliance requirements. the structural rigidity of institutions further complicates this role, as contingency planning is often lacking, leaving dss exposed to unforeseen challenges. furthermore, the ds often operates in a grey zone, able to recommend best practices but lacking enforcement power. rigid, one-size-fits-all rdm policies often fail to account for disciplinary diversity, limiting the ds’s ability to support research needs effectively. these dynamics were evident in both the author’s cases and in biweekly exchanges with other dss across faculties, as well as in discussions within the srdsn. similar challenges also appear in the broader literature, which highlights the lack of consensus on the ds role and its often-intermediary position (rousi et al., 2024; unece, 2022; verhulst, 2025). taken together, these cases show the ds as an intermediary, navigating a buffer zone between the imperatives of global compliance and the practical realities faced by researchers. the ds operates in a street-level role (lipsky, 1980), caught between the abstract, formal logic of institutional mandates and the day-to-day complexities of research practice. although they are expected to implement policies that align with fair and care principles, the gdpr and open science imperatives, dss often lack the formal authority to adapt these mandates to local needs. instead, they 10/19 marmier, a., stepanovic, s., & mettler, t. (2025). in the data steward’s shoes: an autoethnographic exploration of everyday challenges, iassist quarterly 49(4), pp. 1-19. doi: https://doi.org/10.29173/iq1172 exercise discretion within constraints, seeking pragmatic workarounds and advocating for more context-sensitive approaches but without power to reshape the systems they uphold. grounded in lived experience and collective exchanges with other dss, figure 1 illustrates the systemic pressures they face. these pressures stem both from above, through institutional and regulatory imperatives, and from below, through the diverse, practical demands of researchers. framing the model in this way ensures that the dynamics described resonate beyond a single perspective. the following discussion examines how these forces shape the ds’s position and activities within this buffer zone (figure 1). figure 1. data steward challenge: global imperatives vs researchers’ reality top-down pressures from academic bureaucracy from an academic perspective, the study shows that three major imperatives shape the pressures faced by dss in their work: the rise of the open science movement, the evolving landscape of funding and journal policies (academic actors), and the implementation of data protection regulations. the first factor that seems to influence universities in their implementation of rdm programs is the growing global emphasis on the open science movement, particularly its mandates regarding open data. initially designed to promote transparency, shareability, and reproducibility in research (fecher et al., 2014), open science has become a key institutional ranking variable that shapes university policies and funding decisions (borgman, 2018; peng, 2018; wałek, 2019). many universities have integrated open science principles not only to enhance research integrity but also as a strategic response to funders and rankings (allen & mehler, 2019; araujo et al., 2024; nosek et al., 2015). 11/19 marmier, a., stepanovic, s., & mettler, t. (2025). in the data steward’s shoes: an autoethnographic exploration of everyday challenges, iassist quarterly 49(4), pp. 1-19. doi: https://doi.org/10.29173/iq1172 research institutions are progressively assessed based on their adherence to open science standards, such as the implementation of fair principles (wilkinson et al., 2016). in addition, the increasing demand for data sharing and open access publishing has led to incentives for compliant researchers, such as a data reuse award or data works! prize (data science at nih, 2024; fors, 2024). yet, despite such measures, the push for openness often collides with researchers’ priorities, i.e. time constraints, sensitive data, or inadequate institutional resources. this institutionalisation of open science thus places dss in a complex position, requiring them to act as both facilitators and enforcers of data sharing policies. while supporting researchers comply with requirements and policies, they must also navigate disciplinary differences, ethical concerns, and hesitations about premature sharing or misuse. as noted earlier, another factor that appears to influence ds activities is the evolving landscape of funding and journal policies regarding open data. major research funders, such as the usa’s national institutes of health (nih), have implemented data-sharing mandates as funding prerequisites. for instance, the nih’s data management and sharing (dms) policy, effective january 25, 2023, requires institutions to share data to accelerate biomedical discovery (national institutes of health, 2023). similarly, academic journals are increasingly adopting data-sharing policies (piwowar et al., 2018). publishers like taylor & francis and sage publications now expect authors to share supporting data. taylor & francis, for example, has a basic data-sharing policy that encourages authors to make their data open, provided it does not violate ethical or privacy concerns (taylor & francis, 2025). sage publications similarly supports and encourages research data being shared, discoverable, citable, and recognised as an intellectual product of value (sage editorial policies, 2025). however, as seen in this study and others (tenopir et al., 2020), despite funders' requirements and openness incentives, researchers often struggle with the practicalities of implementing open science mandates, especially with sensitive data, disciplinary variation, or proprietary constraints. this places dss under significant pressure, as they are expected to support researchers on highly specific issues without being involved in shaping funders’ policies. as a result, they find themselves in this buffer zone—listening to researchers’ needs while aligning them with the requirements of funders. consequently, dss must continuously adapt to shifting policies while acting as both facilitators and negotiators, advocating for researchers’ needs within institutional frameworks while ensuring compliance with external mandates. particularly visible throughout this autoethnography, the implementation of data protection regulation has fundamentally altered the landscape of research data governance. regulations such as the gdpr and the swiss nfadp impose stringent compliance requirements on universities, extending their obligations not only to institutional but also to researchers and dss. additionally, technical recommendations, such as technical and organisational measures (toms) (federal data protection and information commissioner, 2024), define steps for data security, risk mitigation, and privacy. universities must now establish data protection policies that align with gdpr, nfadp, and other jurisdiction-specific regulations (golla, 2017). one of the major challenges of data protection regulations is the tension between legal compliance and research efficiency. while laws emphasise privacy and security, researchers often prioritise data accessibility and usability. this creates a compliance paradox: security measures often clash with research needs. this was particularly evident in one case, where researchers’ contract with the police made full compliance with institutional and federal recommendations difficult without compromising relations with data providers. as a result, 12/19 marmier, a., stepanovic, s., & mettler, t. (2025). in the data steward’s shoes: an autoethnographic exploration of everyday challenges, iassist quarterly 49(4), pp. 1-19. doi: https://doi.org/10.29173/iq1172 this situation places dss at the centre of the tension between data protection regulations and researchers’ practices, with no straightforward solution to satisfy both sides. beyond acting as advisors, dss often serve as mediators, balancing legal constraints with the needs of researchers. this requires them to interpret evolving regulations, translate legal jargon into practical guidance, and sometimes even advocate for researchers when institutional policies become overly restrictive. however, dss often lack formal authority in policy-making, leaving them in a reactive role rather than a proactive one. bottom-up pressures from researchers’ reality from the researchers’ perspective, the study shows that factors such as prioritising the advancement of knowledge, maintaining academic freedom, and meeting project-specific needs seem to shape the pressures dss face daily. at the heart of academic research is the pursuit of knowledge, driving researchers to generate, analyse, and share data that contributes to scientific progress. as borgman (2017) notes, we show that this commitment often leads researchers to prioritise methodologies and outputs that maximise knowledge production, sometimes at the expense of institutional requirements for compliance and data governance. while research institutions implement policies to enhance data integrity, security, and accessibility, researchers may view them as restrictive, particularly when they introduce administrative burdens that slow down scientific workflows (allen & mehler, 2019; levin et al., 2016). this compliance–efficiency tension places dss in a delicate balance. on the one hand, they must comply with institutional policies and external regulations to ensure that research data are managed following legal and ethical standards. on the other hand, they must meet researchers’ needs for flexibility, speed and efficiency—especially in fast-moving fields where delays in data processing or publication can have significant consequences. dss are thus caught between institutional imperatives of control and standardisation and researchers’ expectations of autonomy and minimal bureaucracy. in practice, this means dss often act as negotiators rather than enforcers, translating compliance requirements into workable researcher-friendly workflows. academic freedom is a foundational principle of research institutions, granting researchers the autonomy to explore novel ideas, select methodologies, and communicate findings without undue restrictions. however, contemporary research governance, particularly in data management, adds oversight often perceived as encroaching on this freedom (borgman, 2018). as this paper shows, one key area of tension arises in the choice of tools and infrastructure. institutional policies often mandate the use of specific data storage solutions, collaboration platforms, or security protocols to ensure compliance with regulations such as the gdpr or national data protection laws (corti et al., 2014; université de lausanne direction, 2019). while these policies aim to enhance security and standardisation, they seem to limit researchers’ ability to use external software or cloud services that better align with their workflows. for instance, researchers conducting international collaborations may find institutional policies prohibitive if they restrict the use of widely used global platforms like google drive or dropbox in favour of local institutional storage (perrier et al., 2017). furthermore, some researchers express concerns that stringent data governance policies could introduce a culture of surveillance, where institutions control what data can be collected, stored, or shared (borgman, 2018). this perceived oversight can discourage innovative or controversial research, stifling academic freedom. balancing compliance with academic autonomy is a delicate task that requires careful 13/19 marmier, a., stepanovic, s., & mettler, t. (2025). in the data steward’s shoes: an autoethnographic exploration of everyday challenges, iassist quarterly 49(4), pp. 1-19. doi: https://doi.org/10.29173/iq1172 consideration. in this context, dss navigate the intersection of regulation and academic freedom, advocating for flexible policies that accommodate diverse research needs while ensuring legal and ethical standards. with more institutional flexibility, dss could develop tailored solutions that maintain both data governance integrity and academic freedom. as shown in the four cases, research projects vary widely in scope, methodology and data requirements, highlighting the inadequacy of a one-size-fits-all approach to rdm. disciplinary diversity necessitates tailored strategies addressing field-specific challenges. in disciplines like genomics or climate science, standardised sharing norms facilitate consistent practices. conversely, ethnography or qualitative social sciences handle highly contextual, sensitive data requiring customised rdm strategies (mannheimer et al., 2018; zenk-möltgen et al., 2018). furthermore, the nature of qualitative data, which often contains personal narratives and contextual details, raises ethical concerns regarding privacy and consent when sharing. a study analysing variations across research domains arts and humanities, social sciences, medical sciences, and basic sciences also revealed significant distinctions in data management actions and attitudes (akers & doty, 2013). these differences highlight the necessity for rdm services that are sensitive to the specific needs of each discipline. additionally, the rise of big data sources, such as social media and blogs, presents new challenges. ethical questions around using publicly available data without explicit consent remain unresolved, complicating sharing and reuse (mannheimer et al., 2018). given this complexity, dss must navigate a landscape where standardised rdm protocols may not suffice. they face pressures to develop flexible, discipline-specific strategies that respect ethics and meet project needs. this necessitates a nuanced understanding of diverse methodologies and a commitment to ethical data management practices. implications and limits implications for research data management and data stewardship this study emphasises the evolving and multifaceted role of dss in academic institutions. dss are positioned at the intersection of institutional mandates and researchers’ practical needs. they are expected to ensure compliance with open science policies, funding requirements and data protection regulations, while also supporting researchers in their pursuit of knowledge. this dual responsibility places dss in a 'buffer zone', where they must balance sometimes conflicting priorities, often without the formal authority to influence the rules they are tasked with implementing. rather than offering a catalogue of issues, this paper provides an overview of recurring tensions. through an autoethnographic lens, it brings forward a perspective rarely voiced but increasingly relevant in practice. while the findings are based on an autoethnography, the issues raised have broader relevance and invite reflection within other settings. the goal is not to generalise, but to initiate debate on the role and function of dss in research environments, and to identify areas for improvement. this paper presents a practice-based analysis of structural frictions in data stewardship and calls for rethinking how the ds role is integrated, supported, and empowered, enabling institutions to transition from compliance to enabling research environments. one such area is the need to integrate dss more effectively into institutional decision-making processes. currently, dss are often tasked with implementing policies without having been involved 14/19 marmier, a., stepanovic, s., & mettler, t. (2025). in the data steward’s shoes: an autoethnographic exploration of everyday challenges, iassist quarterly 49(4), pp. 1-19. doi: https://doi.org/10.29173/iq1172 in their development, which limits their ability to adapt them to disciplinary, methodological, or ethical realities. giving dss a formal voice in data governance structures would increase the relevance of policies and foster institutional coherence. another key implication is the need for more flexible, researcher-centred data management frameworks. rigid, one-size-fits-all policies often fail to reflect the diversity of research practices, particularly in the social sciences. allowing dss to propose tailored solutions and encouraging institutions to accommodate exceptions, where justified, would improve the relevance and uptake of rdm policies. finally, as funders’ and journals’ requirements continue to evolve, institutions must strengthen internal support structures. this includes ongoing training, technical resources, clearer internal guidance, and better coordination. limitations and future research directions while this study provides valuable insights into the lived experiences of dss, it also has certain limitations. first, the autoethnographic approach is inherently subjective, relying on the experiences of a single ds. although this method enables a nuanced exploration of challenges in practice, it does not offer a broader, institution-wide perspective. future studies could complement these findings with comparative analyses or surveys of multiple dss. second, the case studies highlight tensions between institutional compliance and research practice, but they primarily focus on researchers from the faculty of law, criminal justice and public administration. expanding this analysis to life sciences, humanities, or interdisciplinary research environments could provide a more holistic view of how dss adapt to different disciplinary challenges. finally, the study acknowledges that dss often operate without formal authority in institutional decision-making. future research could explore strategies for empowering dss, including potential governance models where dss play a more active role in policy development, researcher advocacy, and institutional data strategy planning. examining how different universities structure their ds roles could provide actionable recommendations for institutions seeking to improve their rdm frameworks. references akers, k. g., & doty, j. (2013). disciplinary differences in faculty research data management practices and perspectives. international journal of digital curation, 8(2), 5–26. https://doi.org/10.2218/ijdc.v8i2.263 allen, c., & mehler, d. m. a. (2019). open science challenges, benefits and tips in early career and beyond. plos biology, 17(5), 1–14. https://doi.org/10.1371/journal.pbio.3000246 araujo, p., bornatici, c., ochsner, m., & heers, m. (2024). assessing and enabling open research data practices in swiss higher education institutions: a comprehensive landscape analysis. https://doi.org/10.5281/zenodo.12755475 15/19 marmier, a., stepanovic, s., & mettler, t. (2025). in the data steward’s shoes: an autoethnographic exploration of everyday challenges, iassist quarterly 49(4), pp. 1-19. doi: https://doi.org/10.29173/iq1172 arend, d., psaroudakis, d., memon, j. a., rey-mazón, e., schüler, d., szymanski, j. j., scholz, u., junker, a., & lange, m. (2022). from data to knowledge big data needs stewardship, a plant phenomics perspective. the plant journal : for cell and molecular biology, 111(2), 335–347. https://doi.org/10.1111/tpj.15804 borgman, c. l. (2017). big data, little data, no data: scholarship in the networked world. (the mit press, ed.). the mit press. borgman, c. l. (2018). open data, grey data, and stewardship: universities at the privacy frontier. berkeley technology law journal, 33(2), 365. https://doi.org/10.15779/z38b56d489 corti, l., eynden, v. van den, bishop, l., & woollard, m. (2014). managing and sharing research data : a guide to good practice (sage publishing). sage publishing book webpage. https://doi.org/https://doi.org/10.25607/obp-1540 cox, a. m., pinfield, s., & rutter, s. (2019). the intelligent library: thought leaders’ views on the likely impact of artificial intelligence on academic libraries. library hi tech, 37(3), 418–435. https://doi.org/10.1108/lht-08-2018-0105 data science at nih. (2024). the 2024 dataworks prize. https://datascience.nih.gov/news/2024dataworks-prize data stewardship, curricula and career paths eosc association. (2023). https://eosc.eu/advisorygroups/data-stewardship-curricula-and-career-paths/ edmunds, r., l’hours, h., rickards, l., trilsbeek, p., vardigan, m., & mokrane, m. (2016). core trustworthy data repositories requirements. https://doi.org/10.5281/zenodo.168411 emmott, s., & rison, s. (2008). towards 2020 science. science in parliament, 65(4), 31–33. european union. (2016). general data protection regulation (gdpr). official journal of the european union. https://gdpr-info.eu/ fecher, b., friesike, s., fecher, b., & friesike, s. (2014). open science: one term, five schools of thought. opening science, 17–47. https://doi.org/10.1007/978-3-319-00026-8_2 federal data protection and information commissioner. (2024). guide to technical and organisational data protection measures (tom). fors. (2024). fors data re-use award. https://forscenter.ch/fors-data-reuse-award-2024/ golla, s. j. (2017). is data protection law growing teeth? the current lack of sanctions in data protection law and administrative fines under the gdpr. journal of intellectual property, information technology and e-commerce law, 8(1), 70–78. https://www.jipitec.eu/jipitec/article/view/192 hackett, c., & kim, j. (2024). planning, implementing and evaluating research data services in academic libraries: a model approach. journal of documentation, 80(1), 27–38. https://doi.org/10.1108/jd-01-2023-0007 16/19 marmier, a., stepanovic, s., & mettler, t. (2025). in the data steward’s shoes: an autoethnographic exploration of everyday challenges, iassist quarterly 49(4), pp. 1-19. doi: https://doi.org/10.29173/iq1172 hammersley, m., & atkinson, p. (2019). ethnography: principles in practice. in ethnography: principles in practice: fourth edition. taylor and francis. https://doi.org/10.4324/9781315146027/ethnography-martyn-hammersley-paulatkinson/accessibility-information hey, t., & trefethen, a. (2003). the data deluge: an e-science perspective. grid computing: making the global infrastructure a reality, 809–824. levin, n., leonelli, s., weckowska, d., castle, d., & dupré, j. (2016). how do scientists define openness? exploring the relationship between open science policies and research practice. bulletin of science, technology & society, 36(2), 128–141. https://doi.org/10.1177/0270467616668760 lord, p., & macdonald, a. (2003). data curation for e-science in the uk: an audit to establish requirements for future curation and provision. mannheimer, s., pienta, a., kirilova, d., elman, c., & wutich, a. (2018). qualitative data sharing: data repositories and academic libraries as key partners in addressing challenges. american behavioral scientist, 63(5), 643–664. https://doi.org/10.1177/0002764218784991 mons, b. (2018). data stewardship for open science: implementing fair principles (1 st edition). chapman and hall/crc. https://doi.org/10.1201/9781315380711 national institutes of health. (2023). data management & sharing policy overview | data sharing. https://sharing.nih.gov/data-management-and-sharing-policy/about-data-management-andsharing-policies/data-management-and-sharing-policy-overview#after nosek, b. a., alter, g., banks, g. c., borsboom, d., bowman, s. d., breckler, s. j., buck, s., chambers, c. d., chin, g., christensen, g., contestabile, m., dafoe, a., eich, e., freese, j., glennerster, r., goroff, d., green, d. p., hesse, b., humphreys, m., … yarkoni, t. (2015). promoting an open research culture. science, 348(6242), 1422–1425. https://doi.org/10.1126/science.aab2374 oladipo, f., folorunso, s., ogundepo, e., osigwe, o., & akindele, a. (2022). curriculum development for fair data stewardship. data intelligence, 4(4), 991–1012. https://doi.org/10.1162/dint_a_00183 peng, g. (2018). the state of assessing data stewardship maturity – an overview. data science journal, 17, 7–7. https://doi.org/10.5334/dsj-2018-007 peng, g., privette, j. l., kearns, e. j., ritchey, n. a., & ansari, s. (2015). a unified framework for measuring stewardship practices applied to digital environmental datasets. data science journal, 13, 231–253. https://doi.org/10.2481/dsj.14-049 peng, g., privette, j. l., tilmes, c., bristol, s., maycock, t., bates, j. j., hausman, s., brown, o., & kearns, e. j. (2018). a conceptual enterprise framework for managing scientific data stewardship. data science journal, 17, 15. https://doi.org/10.5334/dsj-2018-015 17/19 marmier, a., stepanovic, s., & mettler, t. (2025). in the data steward’s shoes: an autoethnographic exploration of everyday challenges, iassist quarterly 49(4), pp. 1-19. doi: https://doi.org/10.29173/iq1172 peng, g., ritchey, n. a., casey, k. s., kearns, e. j., privette, j. l., saunders, d., jones, p., maycock, t., & ansari, s. (2016). scientific stewardship in the open data and big data era roles and responsibilities of stewards and other major product stakeholders. d-lib magazine, 22(5–6), 1–1. https://doi.org/10.1045/may2016-peng perrier, l., blondal, e., ayala, a. p., dearborn, d., kenny, t., lightfoot, d., reka, r., thuna, m., trimble, l., & macdonald, h. (2017). research data management in academic institutions: a scoping review. plos one, 12(5). https://doi.org/10.1371/journal.pone.0178261 peter lyman, & hal r. varian. (2003). how much information. in journal of electronic publishing (vol. 6, issue 2). https://doi.org/10.3998/3336451.0006.204 pinfield, s., cox, a. m., & smith, j. (2014). research data management and libraries: relationships, activities, drivers and influences. plos one, 9(12), e114734. https://doi.org/10.1371/journal.pone.0114734 piwowar, h., priem, j., larivière, v., alperin, j. p., matthias, l., norlander, b., farley, a., west, j., & haustein, s. (2018). the state of oa: a large-scale analysis of the prevalence and impact of open access articles. peerj, 2018(2), e4375. https://doi.org/10.7717/peerj.4375/supp-1 plotkin, d. (2020). data stewardship: an actionable guide to effective data management and data governance (academic press). https://doi.org/10.1016/c2019-0-03988-x poulos, c. n. (2021). essentials of autoethnography. american psychological association. https://doi.org/10.1037/0000222-000 rosenbaum, s. (2010). data governance and stewardship: designing data stewardship entities and advancing data access. health services research, 45(5p2), 1442–1455. https://doi.org/10.1111/j.1475-6773.2010.01140.x rousi, a. m., boehm, r. i., & wang, y. (2024). data stewardship: case studies from north american, dutch and finnish universities. journal of documentation, 80(7), 306–324. https://doi.org/10.1108/jd-12-2023-0264 sage editorial policies. (2025). research data sharing policies. sage publications inc. https://us.sagepub.com/en-us/nam/research-data-sharing-policies smale, n., vietnam, r., csiro, k. u., denyer, g., magatova, e., & barr, d. (2020). a review of the history, advocacy and efficacy of data management plans. international journal of digital curation, 15(1), 1–29. https://doi.org/10.2218/ijdc.v15i1.525 taylor & francis. (2025). understanding our data sharing policies author services. https://authorservices.taylorandfrancis.com/data-sharing-policies/ tenopir, c., birch, b., & allard, s. (2012). academic libraries and research data services: current practices and plans for the future. association of college and research libraries. http://hdl.handle.net/11213/17190 18/19 marmier, a., stepanovic, s., & mettler, t. (2025). in the data steward’s shoes: an autoethnographic exploration of everyday challenges, iassist quarterly 49(4), pp. 1-19. doi: https://doi.org/10.29173/iq1172 tenopir, c., rice, n. m., allard, s., baird, l., borycz, j., christian, l., grant, b., olendorf, r., & sandusky, r. j. (2020). data sharing, management, use, and reuse: practices and perceptions of scientists worldwide. plos one, 15(3), e0229003. https://doi.org/10.1371/journal.pone.0229003 the swiss confederation. (2023). federal act on data protection. fedlex. https://www.fedlex.admin.ch/eli/cc/2022/491/en unece. (2022). data stewardship (version 1, 2 june 2022) – consultation paper. https://unece.org/sites/default/files/2022-06/data_stewardship_ver_1_020622%20%20for%20consultation.pdf université de lausanne direction. (2019). directive de la direction 4.5 sur le traitement et la gestion des données de recherche. https://www.unil.ch/unil/fr/home/menuinst/universite/cadre-legalet-reglementaire/directives-internes-de-lunil.html#recherche verhulst, s. g. (2025). data stewardship decoded: mapping its diverse manifestations and emerging relevance at a time of ai. https://arxiv.org/abs/2502.10399v1 wałek, a. (2019). data librarian and data steward – new tasks and responsibilities of academic libraries in the context of open research data implementation in poland. przegląd biblioteczny, 87(4), 497–512. https://doi.org/10.36702/pb.634 wendelborn, c., anger, m., & schickhardt, c. (2023). what is data stewardship? towards a comprehensive understanding. journal of biomedical informatics, 140, 104337. https://doi.org/10.1016/j.jbi.2023.104337 whyte, a., green, d., avanço, k., di giorgio, s., gingold, a., horton, l., koteska, b., kyprianou, k., prnjat, o., rauste, p., schirru, l., sowinski, c., torres ramos, g., van leersum, n., sharma, c., méndez, e., & lazzeri, e. (2023). catalogue of open science career profiles minimum viable skillsets. https://doi.org/10.5281/zenodo.8101903 wildgaard, l., vlachos, e., nondal, l., larsen, a. v., & svendsen, m. (2020). national coordination of data steward education in denmark: final report to the national forum for research data management (dm forum). https://doi.org/10.5281/zenodo.3609516 wilkinson, m. d., dumontier, m., aalbersberg, ij. j., appleton, g., axton, m., baak, a., blomberg, n., boiten, j. w., da silva santos, l. b., bourne, p. e., bouwman, j., brookes, a. j., clark, t., crosas, m., dillo, i., dumon, o., edmunds, s., evelo, c. t., finkers, r., … mons, b. (2016). the fair guiding principles for scientific data management and stewardship. scientific data 2016, 3(1), 1–9. https://doi.org/10.1038/sdata.2016.18 york, j., gutmann, m., & berman, f. (2018). what do we know about the stewardship gap. data science journal, 17, 1–19. https://doi.org/10.5334/dsj-2018-019 19/19 marmier, a., stepanovic, s., & mettler, t. (2025). in the data steward’s shoes: an autoethnographic exploration of everyday challenges, iassist quarterly 49(4), pp. 1-19. doi: https://doi.org/10.29173/iq1172 zenk-möltgen, w., akdeniz, e., katsanidou, a., naßhoven, v., & balaban, e. (2018). factors influencing the data sharing behavior of researchers in sociology and political science. journal of documentation, 74(5), 1053–1073. https://doi.org/10.1108/jd-09-2017-0126/full/pdf history and statistical analysis: a case study by patricia e. prestwich ' department of history university ofalberta in my historical work, i am an occasional, or perhaps even accidental, user of computerized statistical analysis. my level of competence in this area can be best conveyed by the fact that i first met my colleague. chuck humphrey, of the university of alberta, when i walked into his office and told him that i had copied the records of 14,000 french mental patients. 1 then asked whether he thought that i could analyze them using index cards. i was fortunate to find an expert who understood what i was trying to do with my data and who could make the computer work for me. historians have long been reluctant to engage in extensive statistical analysis, which they often dismiss as "number-crunching." in part, their reluctance stems from a genuine fear of obliterating the particular and the personal—aspects that, for many of us, are an essential part of history. obviously, however, this reluctance also stems from ignorance or fear of the technology and methodology. the field in which i am now working—the social history of medicine or, specifically, the social history of madness—illustrates how slowly historians can turn to computerized statistical data. this is a relatively new field and until seven or eight years ago, most people in the field concentrated on the analysis of historical documents, particularly the writings of doctors, using many of the theories about power inspired by foucault and sociologists. a primary interest has been the nineteenth century psychiatrist hospital, or asylum.or "madhouse". it has been seen as the symbol of social control, of the ways in which the bourgeoisie in general and psychiatrists ( or "mad doctors") in particular deflected any challenges to their ]x)wer by labelling it as deviation. but, as a number of historians in different countries began to point out, much was being theorized about the asylum without any detailed evidence of how it functioned or whom it supposedly controlled. for the past few years, a small number of studies have emerged which look at asylum records and try to understand the complex functioning of this institution. these studies , although few in number, have already begun to challenge many of the predominant theories about the asylum and about nineteenth century psychiatric medicine. most of these studies contain some statistical analysis, although even historians of the asylum are still cautious in this respect. my own research is the study of a parisian asylum, sainteanne, from its opening in 1867 as the first of the new model asylums, until the end of the first world war. sainte-anne may not be a typical asylum — although no one is sure now what a typical nineteenth century asylum was. like most public asylums in the nineteenth century, it was for the poor, in this case the working class and petty bourgeois of paris. sainte-anne was, however, the only parisian asylum that was not in the suburbs, but the city itself—an important factor in considering the relations between families, the asylum and the psychiatrists. it was also the teaching hospital for the faculty of medicine of the university of paris and its doctors were among the most eminent in france. the nineteenth century asylum generated masses of printed statistics— in fact the main occupation of nineteenth century medical institutions seems to have been the compilation of statistics. this is not only a reflection of their institutional character but of the fact that by the end of the nineteenth century doctors seemed to be more interested in the diagnosis, or rather the classification, of mental illness than in its treaunent. data on mental patients became an important means of both refining and justifying their classifications. but, of course, much of the published statistical material is not useful today becau.se we ask different questions. to give some specific examples, the asylum recorded and printed extensive statistics on the occupations, marital status, age, sex, and diagnoses of their patients but always in separate charts, so that it is difficult to make any correlations. ( for example, we know how many single women were interned, and how many employees, but not how many single women employees.) they recorded the length of stay of those admitted for the first lime ( probably with a view to giving a rosier picture of cure rates) but not of those who had been readmitted, although readmissions constituted a significant proportion of their patients. in the printed statistics, there is no correlation between diagnosis and length of stay, or between length of stay and result of treatment (i.e. death, transfer or release). so, for example, it is impossible to tell from the printed records whether a male depressive lassist quarleriy would stay as long as a male alcoholic or a female depressive and what chances each had of release. thus, while the printed material is sometimes useful for verification, it was essential for me to compile my data from the original records. these records are the registres de la loi, the legal register that must be retained permanently for every patient admitted to psychiatric hospital in france. they are highly confidential documents and even today are not computerized because the french have very strict legislation about privacy of 'information. (today, a clerk enters the details by hand; in the nineteenth century it was often mental patients who did this work.) the registers give the basic demographic data on each patient— age, occupation, marital status— as well as date of entry, date of exit, legal status, and result of treatment ( ie death, transfer or discharge.) there are also three diagnoses for each patient: an admitting diagnosis, a diagnosis after 24 hours and a diagnosis after 2 weeks. usually, the diagnoses are by different doctors. the records often contain incidental information on the circumstances under which the patient was interned ( e.g. as a result of a suicide attempt, family violence, or strange behaviour) and sometimes some observations of the patient's behaviour while interned. i have selected the registers only for sainte-anne itself. the hospital also had an admissions bureau which saw almost every patient that was interned in the paris region. the patients came through this bureau and were sent on to the various parisian asylums. appro.ximately 3000 patients per year passed through the bureau of admissions and their records are intact, including many of their medical files. to collect the data, would, however, be an immense project that could only be undertaken by team effort. my data come from the patients that were transferred from the admissions bureau to sainteanne itself. the asylum was built for 500 patients, but by the 1890s usually held about 1000 patients. i have transcribed the registers for every second year from 18671927, for a total of 14,000 patient records. this sample is considerably larger than in comparable historical studies of public asylums, which usually select only certain years. i collected such a large sample in part to deflect criticism that my sample would be unrepresentative, but also because i felt that with a larger sample i could begin to ask certain questions about internment patterns that could not be asked with a smaller sample. even now, i have certain problems; for example, 1 have only 238 cases of senility for the period 1873-1913 and so for some of the detailed analysis, my sample is extremely small. of course, even on the basis of selecting every second year, my sample is not complete, because certain registers could not be found. the registers are stored, in a very disorganized fashion, in a basement room, lit by a 40 watt bulb and covered in dust and rat poison ( the basements of sainte-anne connect with the catacombs of paris.) with the help of a hospital worker, or occasionally, a patient, 1 had to haul these large registers up from the basement. 1 simply did not find all the years that i wanted. or, as often happened, since one year would be spread over several registers, i would find only part of a year. the registers were also difficult to read, because, apart from the dust and yellowing paper, the ink had faded and the handwriting was not always decipherable. although these registers offer some very difficult problems of interpretation, they are an important source for the type of social history that 1 am trying to write. my goal is to write a book on the asylum as a social institution, i.e. as part of a specific historical community. i want to understand what roles this medical institution played in the lives of families, patients, nurses, and doctors.i want to understand what power these different groups had and how they interacted. the statistical data is merely the beginning of my analysis. the data, in some cases, will give me specific answers, but in most cases, it will direct me to the nonstatistical literature. for example, analysis of the statistical data is helpful simply to clear away some of the myths about the nineteenth century asylum and to establish who got interned, for what diagnosis and for how long. social historians of medicine, who have read only the qualitative material, have postulated that the asylum was the dumping ground for the " inconvenient" in society, those who simply did not fit into the developing industrial society. patients in public asylums were certainly not middle-class, but as the analysis of occupations at sainteanne shows, neither were they the dregs of society. there were very few labelled as "vagabonds"(1.5%) and in fact, most gave their occupations as skilled workers ( carpenters, seamstresses, etc.) or as employees.(43% and 16% respectively, but the figure is probably higher if one counts part of the 17% who were women listed as " no occupation and who are usually the wives of skilled workers or employees.) the proportion of unskilled workers, such as day labourers or domestic servants, in my data was only 14%. (again if wives are counted, it might be higher.) it is also clear that, once inside the asylum doors, patients were not necessary doomed to perpetual confinement. after about 1 860, there was a great deal of political and public hostility toward asylums, which were labelled as "modem bastilles", where people languished in unjust internment. although doctors certainly had extensive legal powers, an analysis of the length of stay of patients over 40 or 50 years paints a more complicated picture. at sainte-anne, in the period up to the first world war, about 45% of all patients were released, 30% died and summer 1991 25% were transferred. the length of stay for those who were released is shorter than one would expect. release: 25% 50% 75% au: 40 days 94 220 cut to : 800 days : 36 78 162 of course, these statistics can only be interpreted by relating them to the diagnoses. for example, the 30% death rate, which was higher for men than for women, is directly related to the high number of male patients interned for general paralysis, the third and fatal stage of syphilis . (general paralysis made up 22% of male internments. eighty-seven % of gp cases were men and the death rate at the asylum itself was about 75%. the analysis of the data is useful simply to give some idea of how patients were diagnosed and, although my analysis of this aspect is not finished, there seem to be fairly discrete diagnosis, with not too much overlap. the most common diagnoses were general paralysis , alcoholism, depression, persecution and old age in various forms. it is revealing to compare what doctors faced in the asylums— quite often what they would label " banal" or "uninteresting" problems— and what they discussed in their medical literature, which was usually the unusual, exotic or, as they said the " interesting". the question of what interested doctors can be approached in another way through the data, for i have records not only from the asylum itself, but from the teaching clinic at sainte-anne. by comparing the patterns of diagnosis of the asylum and the clinic, i hope to make some deductions about what interested doctors and how comprehensive an education medical students received. aside from giving certain basic information about who was interned and why, the data can also begin the process of answering some of the questions about the role of families in the whole process of internment. one of the important aspects of the data is that admissions were divided into two types. the first was placement officicl (po)—a legal internment which involved police action. usually the person was taken to the local police station and then to the pouce dispensary, where a police doctor made the final decision as to whether the person would be sent to the bureau of admissions at sainte-anne. but by the 1880s, there was a second type of admission, the placement volontaire (pv), which allowed families and even friends to intern someone without going through the police, although this involved paying the internment expenses in most cases. the pv admissions will give some insights into family behaviour, that is, what behaviour was considered so unacceptable or intolerable as to lead to internment and, conversely, under what conditions would families request the release of patients. this is not to imply, of course, that family decisions were not involved in placement legal. it is clear from the records that a number of families, presumably the poorer ones, would simply call in the police to deal with an intolerable family situation, such as an alcoholic father or a senile elderly relative. but the pv admissions give much clearer evidence of the family's role because they usually indicate who interned the patient ( a mother, spouse, friend, etc. ) . also, because a patient interned "voluntarily" could be released at the insistence of a family member, even if the doctor objected, these files give some insights into the complex relationship between doctors and families. one good example of family power comes from an examination of data on patients who were transferred. transfer of patients from sainte-anne to more distant asylums became increasingly necessary as the asylum became overcrowded in the latter part of the nineteenth century. transfers were strongly resisted, both by patients and families, because it usually meant transfer to poorer care and at a distance that made family intervention impossible. my analysis of length of stay shows that pv patients stayed considerably longer ( i.e., in terms of years) than po patients before they were transferred and that, significantly, this pattern was true for both men and women. 1 would argue that here is a clear indication of effective family infiuence. a third aspect that emerges from the analysis of the data is the gendered nature of the asylum. although feminist historians have speculated a great deal about the tendency to label women as mad if they did not conform to societal norms, there has been relatively little analysis of the asylum from the point of view gender. again, statistical analysis is useful to clear away some myths. women, for example, were not interned more frequently than men, nor did they have a lower release rate. but, they did stay longer and consequently, they had a higher rate of transfer. these differences are clearly related to different patterns of diagnosis. women and men, on the whole, were diagnosed differently. the clearest example is between alcoholism and depression. nearly 30% of the men, but only 10 percent of the women were diagnosed as alcoholic, whereas approximately 30% of the women were diagnosed as depressive, and only 10 % of the men. men and women therefore had different experiences in the asylum. why women were labelled as depressive and men as alcoholic is a question that cannot be answered by the statistical data, of course, but can only be explored by examining more traditional written sources. this is my first foray into this type of analysis and i clearly have much still to learn. (although i now admit lassist quarteriy the superiority of the computer over index cards!) i wish that i had had some idea of the possibilities of computer analysis before i began to collect the data, but that was impossible. i obtained access to these records purely by chance; i recognized the their richness in terms of social history, but i simply had to trust that i would eventually find the right people and the right techniques to help me use the data. whether i will ever use this type of analysis again will depend on the research project. my real problem now is to integrate this statistical analysis into a broader, more traditional narrative and to convey this analysis effectively to my audience of historians, who for the most part still skip the statistical sections in any book. ' paper presented to lasslst conference, may 17, 1991, edmonton alberta. summer 1991 1/12 mwalubanda, joseph mathew (2021), the development of institutional repositories in east africa countries: a comparative analysis of tanzania, kenya, and uganda, iassist quarterly 45(3-4), pp. 1-12. doi: https://doi.org/10.29173/iq1012 the development of institutional repositories in east africa countries: a comparative analysis of tanzania, kenya, and uganda joseph mathew mwalubanda1 abstract this paper aims at examining the growth of ir in the east african region (tanzania, kenya, and uganda) from 2010-2020. this study adopted a content analysis methodology. data for this study was extracted from opendoar (directory of open access repository), roar (registry of open access repository) and repository websites to identify the language used, subject covered, software used and types of content that are found in east african repositories. the findings of this study reveal that east african region has a total number of 66 repositories, which are registered in opendoar. kenya is a leading country in the region by having 42 repositories, followed by tanzania with 14 repositories and uganda with 10 repositories. the findings show that there is an increase in number of repositories in the region from 4 in 2010 to 66 in 2020. however, the growth is low compared to other parts of the world particularly, europe, asia, and america. the study shows the need for librarians, researchers, stakeholders, and east african governments to come together to address the challenges that hinder the growth of repositories in the region. likewise, mandate policy formulation, training, financial support, oa awareness and technical support are needed in order to overcome those challenges. keywords institutional repository, open access, content growth, institutional repository software, items types, institutional repository language, subject covered in repository, east african region. introduction with the development of computer and information technology, people changed the way of sharing and exchanging information. this development has, as a result, enhanced the communication media in the scholarly world, and led to the rise and use of institutional repository as a mean of communication by different institutions. ukwoma and mole (2017, p.117) define ‘institutional repository (ir) as an online platform for preservation and disseminating the intellectual output of an institution’. in other words, a combination of different sets of services in providing information regarding theses and dissertations, e-print, and technical report, among others, is another way of describing institutional repository (ratanya, 2017). as noted by ratanya (2017, p. 276), ‘in the age of electronic publishing and digital content, academic institutions are increasingly realizing the importance of irs as a vital infrastructure for scholarly communication’. indeed, this is in line with the united nations sustainable development goals (unsdgs) regarding promotion of quality education and supporting innovation in the society. the task of supporting the institutions in sharing, disseminating, and preserving information that they produce remains the main role of institutional repositories. furthermore, irs and open access (oa) movement has provided room for different authors and institutions to communicate their findings to the society. through the link and oa tools and services (doaj, doar, roar, sherpa-romeo and sparc) provided to the resources they have, irs and oa provide the easy way of sharing information between institutions and the communities they serve. as noted by okpala (2017, p.5) ‘most universities are now aiming to provide oa to their local contents vis institutional repositories’ and this shows the link between irs and oa movements to most of the academic institutions. in this regard, most of the https://doi.org/10.29173/iq1012 2/12 mwalubanda, joseph mathew (2021), the development of institutional repositories in east africa countries: a comparative analysis of tanzania, kenya, and uganda, iassist quarterly 45(3-4), pp. 1-12. doi: https://doi.org/10.29173/iq1012 institutions are striving to achieve unsdgs by enabling their users to obtain information necessary in their pursuit of education and research. according to okoroma (2018, p.289) universities and other academic institutions all over the world are using ir as a mean of bridging the gap between authors, researchers and other information users. in addition, these serve as the preservation tool of knowledge produced by specific institutions. indeed ir has become a good way of sharing information and it assists universities in accomplishing their goals and objectives of serving the community. in this regard, in order for institution to achieve unsdgs such as quality education, innovation, and infrastructure, use of ir becomes inevitable. accordingly, as tapfuma and hoskins (2019, p.1) observe, ‘oa journals and irs are alternative channels for disseminating and communicating research findings’. however, in order to facilitate the use of irs by organizations and universities, policies are inevitable. these policies will in turn enable these organizations determine the type of information to be hosted by their respective irs. statement of the problem the establishment and implementation of irs has gained momentum globally in recent years. the developed countries have a greater number of institutional repositories compared to the developing countries. this may be observed in different registries of repositories such as the directory of open access repository (doar) and registry of open access repository (roar). data from those repositories show that most of the african countries are lagging behind in the establishment and implementation of irs. dlamini & snyman (2017) note that africa as a continent is struggling in the implementation of irs both in terms of establishing and use. different factors such as lack of funds, poor infrastructure, lack of government support and lack of expertise have been attributed to this phenomenon. this study analyzes and compares the characteristics of irs in east africa. the study uses openly available web resources to determine the software used in different irs; type of materials that are hosted in the irs; type of subject that are covered in the irs and the growth of the irs in the east african region. objectives 1. to determine the software used in different irs. 2. to determine the type of materials that are hosted in the irs. 3. to determine the type of subjects that are covered in the irs. 4. to identify the most used languages in east african repositories. 5. to determine the growth of irs in east africa region. methodology data for this study was collected from the directory of open access repository (opendoar), registry of open access repository (roar) and repository websites. opendoar and roar have been used as the source of data in different studies in the field of repositories and oa (ezema & onyancha, 2017; verma & shulka, 2014; and gul, bashir and ganaie, 2019). given the importance and usability of opendoar in this study, it is very important to provide some design details and characteristics of its records that are found in the opendoar database. this will allow easy understanding of the information found in the database especially in the growth trend, subject, software, and language used. opendoar which is maintained by the university of nottingham was developed in collaboration with lund university under the umbrella of sherpa services. https://doi.org/10.29173/iq1012 3/12 mwalubanda, joseph mathew (2021), the development of institutional repositories in east africa countries: a comparative analysis of tanzania, kenya, and uganda, iassist quarterly 45(3-4), pp. 1-12. doi: https://doi.org/10.29173/iq1012 opendoar consists of different records in its database such as: • description: this place describes the repository and kind of service that is offered • software: shows the type of archived software used by the specific repository when known • organization: it shows the origin of the parent organization and it provides official website • size: shows the number of records that are hosted in the repository • subjects: show the broad description of the subject by following the library of congress classification scheme • content types: show the types of materials that are hosted in the repository based on the controlled vocabulary like articles, conferences, etc. • languages: show language used in the content types hosted in the repository • policies: provide a brief description of the types of policies of the specific repository (like metadata policy, preservation policy, data re-use policy, etc.) • remarks: it is a place where additional information is provided about the repository (information like partnership may be found here) • repository url: is the place where users will find the url address of the specific repository • opendoar id: this is the unique number of the registered repository in opendoar database therefore, this kind of information is the basis for this study because it provides different characteristics of the repositories. roar is maintained by the university of southampton, uk and is the part of eprints.org. roar data is limited to the number of repositories, software used in archiving materials, types of repositories and the number of records. the core data for this study come from opendoar, and i supplement it with data from roar and the repositories websites. the data that was collected from opendoar database is up-to-date and reliable for this study as it has been updated monthly. opendoar and roar have been used to identify different repositories that are found in the east african region (tanzania, kenya, and uganda). as mentioned above, opendoar provides information about the country of originality for every repository. this has assisted the researcher in collecting and scrutinizing data for this study. moreover, both opendoar and roar provide information about registered repositories in the world. information, like software used, type of ir, number of ir by country, content types, and location of ir are found in those two directories. data were extracted from opendoar, roar and repositories websites and coded in excel spreadsheet. the study also explores the literature and existing data to know in detail about the growing trend of ir. this exploratory study was completed by using different published literature like books, journal article, and report. review of literature growth of irs in the world the history of the irs can be traced from 1990s where many institutions around the world started to implement the use of ir in their libraries. at the global level, there have been notable improvements regarding the number of existing ir from 1991 (wyk & mostert, 2010) to 4,580 (ogungbeni, obiamalu, & obiora, 2019). out of 4,580 repositories, europe has 1,798, north america 1,064, asia 917, south america 447, africa 173, australia 86, oceania 3, and unknown locations 92. this means that europe is the leading continent in the world with a large number of repositories. this increase can be attributed to the open access movement in the world where different institutions are trying to assist their societies to have free access to information more easily (wyk & mostert, 2010). as raju, adam & powell (2015, p. 142) observe, ‘these repositories quickly evolved into a platform for libraries to publish and showcase institutions entire range of scholarly output including articles, theses, https://doi.org/10.29173/iq1012 https://www.eprints.org/uk/ 4/12 mwalubanda, joseph mathew (2021), the development of institutional repositories in east africa countries: a comparative analysis of tanzania, kenya, and uganda, iassist quarterly 45(3-4), pp. 1-12. doi: https://doi.org/10.29173/iq1012 dissertations, and journals’. the promotion of open access movement increases this number of irs in the world regardless of the challenges that developing countries are facing in implementing it. growth of irs in africa despite the fact that, africa is far behind in the establishment and use of irs as stated in different international registries of repositories such as roar, and the opendoar, recent trends indicate substantial increase in irs in the region. for example, a study conducted by ogungbeni, obiamalu, & obiora (2019) reveals an increasing number of irs in africa from 7 in 2006 to 62 in 2012 and 173 in 2018. this development might have been motivated by the existence of open access movement and free open source software like dspace, e-print and others. despite this promising trend in the continent, the growth rate of irs in the east african region (tanzania, kenya, and uganda) is still very low. kenya is the leading country in east africa with 42 irs, followed by tanzania with 14 and uganda with 10 irs (opendoar, 2020). kakai, masoke, & okelloobura (2018) point to challenges like lack of awareness, poor infrastructure, culture, technology skills, lack of support from institutions and government, and lack of funds as factors that make most countries in east africa to be behind compared to other places like europe and america. content-type of materials in irs empirical evidence shows that most of irs content types are journal articles (abrizah, noorhidawati & kiram 2010; ezema & onyacha, 2017; shijitha & majeed, 2018; dhanavandan & tamizhchelvan 2014). apart from journal articles, some of the irs particularly in developing countries were found to host dissertations and theses (ejikeme & ezema 2019; verma & shulka 2014). this can be due to the primary functions of irs which is to support parent organization activities. subject covered by irs the choice of which subjects are to be included in irs depends on many factors including the nature of the institution and the audience intended. most of the irs are owned by academic institutions and therefore they support the work of parent organizations in preserving and disseminating research and other publications which are produced in those institutions. studies suggest that most of irs are of multidisciplinary in nature (ezema & onyacha, 2017; shajitha & majeed, 2018). in this case, the nature of the content includes a number of subjects. software used by irs in making sure that content management is effective, different irs have the option of selecting a proper software to use. that software can be open source, built-in or proprietary in nature but they serve the same purpose of managing the content of irs. evidence from different studies identified dspace, gnu eprints and opus as the common software used in conent management (shajitha & majeed 2018; mezbah-ul-islam 2014; ogungbeni, obiamalu, & obiora 2019; ezema & onyacha, 2017; tapfuma & hoskins 2019; sharma, meichieo, & saha 2008; oguche 2018; and sahu & parabhoi 2019). the existence of different content management software, requires each institution intending to implement and manage irs to conduct intensive research to determine type of software to be used for their repositories. literature suggests that different conditions and criteria may be considered when selecting appropriate software for use. this includes the ‘needs of the user, functionality of the software, technical specifications, repository and system administration, content management, dissemination, archiving and system maintenance’ (smith, 2015, p.9). https://doi.org/10.29173/iq1012 5/12 mwalubanda, joseph mathew (2021), the development of institutional repositories in east africa countries: a comparative analysis of tanzania, kenya, and uganda, iassist quarterly 45(3-4), pp. 1-12. doi: https://doi.org/10.29173/iq1012 language one of the important elements in irs is the language to be used target audience. yet, the language to be used in irs varies according to the nature of intended audience and the language they use. however, the most dominant languages used in a good number of irs include english, followed by french and german (ibrahim & beigh 201; ezema & onyancha, 2017). other languages are used depending on the area where ir is located and the functions that it serves to the society. in this case, apart from english, french and german as common languages, hindi, sanskrit, arabic, dyuthi, malayalam, kannada, bangla, marathi, bengali, tamil, gujarati, sanskrit, sinhalese and pashto were found to be used though at a minimal rate (gul, bashir & ganaie 2020; shajitha & majeed 2018). policy it is the task of any irs to provide guideline on how information will be deposited and presented. further such guidelines should also specify individuals who have rights to deposit in the irs. apart from preservation and presentation, these guidelines should also address specific areas like, meta data, data re-use, content policy and submission policy. as elaborated by nunda & elia (2019, p.3), ‘a clear institutional repository policy and managerial issues are crucial in the operability and sustainability of institutional repositories as they guide on the type of content deposited, preserved, withdrawn and the day-to-day interoperability of institutional repositories’. the importance of ir policy cannot be underestimated as it acts as a proven guidelines which assist members of the community to have easy access and use of irs without violating key issues like copy rights. despite the importance and relevance of ir policies, recent studies by gul, bashir & ganaie (2020); sahu & parabhoi (2019) conducted in asia and german, switzerland and austria respectively, revealed that most of the irs do not have policies for metadata re-use, data re-use policy, content policy, submission policy, and preservation policy, while small number of irs have one among the mentioned policies. although there are no clearly stipulated reasons for the absence of ir policies in the mentioned countries, the reasons might be associated with what was presented by xia et al 2012 and lynch 2003, who showed the little contribution of the said ir guiding policies. to them ir policies have been the reason for the failure of irs in many institutions as some of the policies are trying to force people to use irs. in order to ensure effectiveness and usefulness of the irs policy, there is a need for considering different stakeholders when institutions are deciding to impose policy in the implementation of irs. this will enable the institutions to come up with relevant and reliable policy which serves the institutions, authors and community at large that are likely to use irs as the platform for sharing scholarly communication. one of the important elements to be stated in the policy should be on the role of different actors working on the implementation of irs. likewise, there has to be collaboration and coordination between different sections and fields so that institutions will achieve the planned goals for the establishment of irs (smith, 2015). discussion & findings number of ir in east african countries the findings reveal that kenya is the leading country in east africa and has 42 (61%) repositories, followed by tanzania 14 (21%), and uganda which has only 10 (15%). this correlates with the good economy (gdp per capita) that kenya has compared to tanzania and uganda. furthermore, oa policy formulation, and collaboration with different organizations such as kenya library and information services consortium and electronic information for libraries with libraries in kenya assisted much in the establishment of irs in most universities in kenya (chilimo, 2015). the total number of irs in all three countries which have been registered in opendoar up to 18 may 2020 was 66. https://doi.org/10.29173/iq1012 6/12 mwalubanda, joseph mathew (2021), the development of institutional repositories in east africa countries: a comparative analysis of tanzania, kenya, and uganda, iassist quarterly 45(3-4), pp. 1-12. doi: https://doi.org/10.29173/iq1012 table 1: number of irs in east african countries sn countries number of ir percentage % 1. tanzania 14 21% 2. kenya 42 61% 3. uganda 10 15% total 66 100% source: directory of open access repository (18-may-2020) types of ir the findings reveal that a total of 63 (95%) irs were established by universities, research-based institutions and higher learning institutions, while government irs were 2 (3%), and disciplinary irs was 1 (2%). these findings seem to correlate with what was noted by ezema & onyancha (2017, p.110): that ‘the major challenge of african governments is lack of interest in education and research’. this is demonstrated by the number of government repositories presented in opendoar. the majority of irs are institutionally based because of the research activities and the need for promoting scholarly communications in society. this is confirmed to the previous studies of (verma & shulka, 2014; ibrahim & beigh, 2019; ezema & onyancha, 2017; sahu & parabhoi, 2019 and singh, 2014) that irs leads as the type of repository in most places. table 2: types of irs in east african countries s/n type of ir number percentage (%) 1. disciplinary 1 2% 2. government 2 3% 3. institutional 63 95% total 66 100% source: directory of open access repository (18-may-2020) software used in irs in implementing and developing an ir, institutions may opt to use either open source or proprietary software. there are several different types of software available for use into irs. the study established that most of the institutions opt to use dspace 60 (91%), followed by eprints 3 (4.5%), drupal 1 (1.5%), and greenstone 1 (1.5%). however, 1 (1.5%) repository did not specify the type of software used. the findings show that dspace is the leading software for most of the repositories. the use of dspace can be attributed to the fact that it is free open-source software, it is cheap to install and to maintain, and it has user-friendly features. the same findings have been revealed by ezema & onyancha (2017), ejikeme & ezema (2019), singh (2014), verma & shukla (2014), and oguche (2018). however, this is different from other studies such as those done by ibrahim & beigh (2019) and sahu & parabhoi (2019) where the preferred software was opus and eprints. therefore, the use dspace software by most of east african irs may have been influenced by the fact that it is open access software and it easy to use and install. https://doi.org/10.29173/iq1012 7/12 mwalubanda, joseph mathew (2021), the development of institutional repositories in east africa countries: a comparative analysis of tanzania, kenya, and uganda, iassist quarterly 45(3-4), pp. 1-12. doi: https://doi.org/10.29173/iq1012 table 3: software used in east african countries irs s/n software number percentage (%) 1. dspace 60 91% 2. eprints 3 4.5% 3. drupal 1 1.5% 4. greenstone 1 1.5% 5. unknown 1 1.5% total 66 100% source: directory of open access repository (18-may-2020) language used in east african region (tanzania, kenya and uganda) the commonly used official language are kiswahili and english. the findings reveal that english is the dominant language in east africa as 66 (100%) repositories host materials in english. however, the study also found that 3 repositories had materials in kiswahili while 2 repositories had materials in french. publications in local languages were few with all language such as swahili which is dominant language used in the east african community. the publication of electronic resources in english in all the repositories can be attributed to the fact that english used as a medium of instruction in education systems and as the official language in the region. furthermore, the need to disseminate information globally has also influenced the use of english by most of these irs. as ezema and onyacha (2017, p.107), observe ‘in scholarly communication, language of research publication is critical to international scientific publication’. for this reason, it is important to use the language that is understood by many people in order to facilitate easy access of research works. this is consistent with previous studies which had also confirmed that english is the dominant language used in most repositories (ezema & onyancha, 2017; sahu & parabhoi, 2014; ibrahim& beigh, 2019; verma & sulka, 2014 and shajitha & majeed, 2018). table 4: language used in irs s/n languages number of ir 1. swahili 3 2. english 66 3. french 2 source: directory of open access repository (18-may-2020) policies table 5 indicates the policies of irs on these three countries. the policy documents have included metadata policy, data policy, preservation policy, content policy, and submission policy. the findings show that 42 (63.6%) repositories do not have written policies and 24 (36.4%) had their policies. the majority of irs did not define policies for their repositories. sahu & parabhoi (2019) reveal the same findings about a comparative study of german, switzerland, and austria on the oa repository. thus, https://doi.org/10.29173/iq1012 8/12 mwalubanda, joseph mathew (2021), the development of institutional repositories in east africa countries: a comparative analysis of tanzania, kenya, and uganda, iassist quarterly 45(3-4), pp. 1-12. doi: https://doi.org/10.29173/iq1012 most of the repositories experience difficulties for lack of guidance in those areas. this can be the challenge of many irs in the world because having policy documents needs involvement of many stakeholders and support from parent organizations in order to contribute to the sustainable maintenance and management of the ir. table 5: policies in irs s/n policies documents no. of repositories percentage 1. yes 24 36.4% 2. no 42 63.6% total 66 100% source: directory of open access repository (18-may-2020) content types the study established that institutional repositories in the east african countries had the following type of content, namely, journal articles 57 (86%), theses and dissertations 54 (82%), conference and workshop papers 43 (65%), reports and working papers 37 (56%), books, chapters and sections 28 (42%), learning objects 19 (29%), bibliographic references 16 (24%), other special item types 19 (29%) and datasets 1 (2%). journal articles are the leading content type on most repositories in the east african countries. most of these repositories are hosted by research institution and universities as they publish journal articles to support activities of their parent institutions. previous studies also confirm journal articles as the dominant content type in most of the repositories (ibrahim & beigh, 2019; ezema & onyancha, 2017; shajitha & majeed, 2018; verma & shulka, 2014; singh, 2014; and ejikeme and ezema, 2019). table 6: content types of east african countries irs s/n content type no. of repositories percentage 1. journal article 57 86% 2. conferences and workshop papers 43 65% 3. theses and dissertation 54 82% 4. reports and working papers 37 56% 5. books, chapters and sections 28 42% 6. learning objects 19 29% 7. other special items 19 29% 8. bibliographic references 16 24% 9. datasets 1 2% source: directory of open access repository (18-may-2020) https://doi.org/10.29173/iq1012 9/12 mwalubanda, joseph mathew (2021), the development of institutional repositories in east africa countries: a comparative analysis of tanzania, kenya, and uganda, iassist quarterly 45(3-4), pp. 1-12. doi: https://doi.org/10.29173/iq1012 subjects covered in ir subject coverage in opendoar has been arranged using broad subject description of the library of congress classification scheme. thus, other repositories with wide range of subject have been registered as multidisciplinary repositories and others have been registered based on the subjects that they cover. the findings show that majority of subjects in irs were multidisciplinary in nature and followed by other subjects as shown in figure (6) below. this can be attributed to the fact that most of the repositories are institutions by nature and that is the way they register their institutions in opendoar. these findings are consistent with previous studies on this subject (ibrahim & beigh, 2019; ezema & onyancha, 2017; singh, 2014; verma & shulka, 2014; shajitha & majeed, 2018, and abrizah, noorhidawati and kiran, 2017). again, the multidisciplinary nature of these irs may also be linked to the fact that these repositories are irs by nature and thus they are designed to support universities as they offer services to their communities in different areas of specialization. figure 1: subject coverage in east african countries repositories source: directory of open access repository (18-may-2020) growth of irs in east african countries the findings for the growth of irs in east african countries from 2010 to 2020 show that in 2013, 2015, and 2019, there was a rise of many irs in the region. the first ir in east africa was established in 2007 by makerere university in uganda followed by other universities from kenya and tanzania in 2008 and 2009. in 2010 most of the universities in kenya, uganda and tanzania had started the process of establishing the irs in their institutions. however, for tanzania and uganda, the growth of irs is low compared to kenya, and there is no registration of new ir in the opendoar and roar for 2011, 2014, 2016, 2017 and 2020. in kenya, the growth of ir registration and establishment is high compared to tanzania and uganda because every year from 2010 to 2020 there was new registration in opendoar and roar. the total number of repositories in these three countries was 66 by 18-may-2020. these data were collected from the opendoar and roar. this represents a positive development when compared with 2017 data in which only 35 repositories existed in the region. 0 5 10 15 20 25 30 35 40 45 m u lt id is ci p lin ar y sc ie ce g e n er al a gr ic u lt u re , f o o d … b io lo gy a n d … c h e m is tr y an d … ea rt h a n d … ec o lo gy a n d … m at h em at ic s an d … p h si cs a n d … h ea lt h a n d … te ch n o lo gy g en e ra l a rc h it ec tu re c iv il en gi n e er in g c o m p u te rs a n d it a rt s an d … g eo gr ap h y an d … h is to ry a n d … la n gu ag e an d … p h ilo so p h y an d … so ci al s ci en ce … b u si n es a n d … ed u ca ti o n la w a n d p o lit ic s li b ra ry a n d … m an ag em en t an d … subject in east african countries institutional repositories https://doi.org/10.29173/iq1012 10/12 mwalubanda, joseph mathew (2021), the development of institutional repositories in east africa countries: a comparative analysis of tanzania, kenya, and uganda, iassist quarterly 45(3-4), pp. 1-12. doi: https://doi.org/10.29173/iq1012 figure 2: the growth trend of repositories in east african countries source: open directory of open access repository (18-may-2020) conclusion the study will assist stakeholders in the east african region (tanzania, kenya and uganda) to see their initiatives on supporting the development of ir in region. however, more support from stakeholders is needed in providing training to librarians, good it infrastructure, collaboration between library organizations and stakeholders. government support is also needed in making sure the sector is having good development because good ir environment is critical for these countries to achieve unsdgs especially innovation, research and providing quality education. again, more work still needs to be done in order to understand the driving forces behind the growth of repositories, and the reasons that influence the development of such repositories. due to methodological approach this study could only speculate the reason why most of the materials are not in local languages like kiswahili and how the content of repositories grows from time to time in the region. therefore, this kind of study was based on growth numbers of irs in the region. future studies should therefore investigate the content growth and general development of repositories in east african region. there is need to identify the content of repositories which is growing and the underlying motivation instead of just looking at the bare numbers of material and repositories. furthermore, the use of local languages in materials hosted in the repositories in the east african region should be examined to identify the various languages in use. however, future studies should adopt a methodology, such as the use of interviews and questionnaires in order to give the researcher in-depth information about the study. 0 2 4 6 8 10 12 2010 2011 2012 2013 2014 2015 2016 2017 2018 2019 2020 growth of institutional repositories in east african countries tanzania kenya uganga https://doi.org/10.29173/iq1012 11/12 mwalubanda, joseph mathew (2021), the development of institutional repositories in east africa countries: a comparative analysis of tanzania, kenya, and uganda, iassist quarterly 45(3-4), pp. 1-12. doi: https://doi.org/10.29173/iq1012 references abrizah, a, noorhidawati, a, & kiran, k 2010, ‘global visibility of asian universities open access institutional repositories’, malaysian journal of library & information science, vol.15, no.3, pp.53-73. chilimo, w 2015, ‘green open access in kenya: a review of the content, policies, and usage of institutional repositories’, https://www.researchgate.net/publication/327187267. dhanavandan, s, & tamizhchelvan, m 2014, ‘institutional repositories in south asian countries: a study on trends and development’, brazilian journal of information science: research trends, vol.8, no.1/2. dlamini, nn, & snyman, m 2017, ‘institutional repositories in africa: obstacles and challenges’, library review, vol.66, no.6-7, pp.535-548, doi: https://doi.org/10.1108/lr-03-2017-0021. ejikeme, an, & ezema, ij 2019,’ the potentials of open access initiatives and the development of institutional repositories in nigeria: implications for scholarly communication’, publishing research quarterly vol.35, no.1, pp.1-16. ezema, ij, & onyancha, ob 2017,’ open access publishing in africa: advancing research outputs to global visibility’, african journal of library, archives & information science, vol.27, no.2,pp. 97115, https://www.ajol.info/index.php/ajlais/article/view/164661. gul, s, bashir, s, & shabir, ag 2019,’evaluation of institutional repositories of south asia’, online information review, vol.43 no.1, pp.192-212. doi: https://doi.org/10.1108/oir-03-2019-0087 ibrahim, s, & beigh, in 2019,’contribution of uk open access repositories to opendoar’, library philosophy and practice, pp.1-10, https://digitalcommons.unl.edu/libphilprac/2592/ kakai, m, musoke, mgn, & okello-obura, c 2018, ‘open access institutional repositories in universities in east africa’, information and learning science, vol.119, no.11, pp.667-681, doi: https://doi.org/10.1108/ils-07-2018-0066 lynch, ca 2003, ‘institutional repositories: essential infrastructure for scholarship in the digital age’, portal: libraries and academy, vol.3, no.2, pp327-336, doi. https://doi.org/10.1353/pla.2003.0039. nunda, im, & elia, ef 2019, ‘institutional repositories adoption and use in selected tanzania higher learning institutions’. oguche, d 2018, ‘the state of institutional repositories and scholarly communication in nigeria’, global knowledge, memory and communication, vol.67, no.1/2, pp.19-33. doi: https://doi.org/10.1108/gkmc-04-2017-0033. okoroma, fn 2018, ‘awareness, knowledge, and attitude of lectures towards institutional repositories in university libraries in nigeria’, digital library perspectives, vol.34, no.4, pp.288307. doi: https://doi.org/10.1108/dlp-04-2018-0011. ogungbeni, ji, obiamalu, ar, & obiora, ku 2019, ‘open digital repositories: prospects of african countries within the global information space’, library philosophy and practice (e-journal). https://digitalcommons.unl.edu/libphilprac/2444. okpala, h.h 2017 ‘access tools and services to open access: doar, roar, sherpa-romeo, sparc and doaj’, informatics studies, vol.4, no.3, pp. 05-20.opendoar, 2020, ‘the directory of open access repositories’, https://v2.sherpa.ac.uk/opendoar/search.html. pinfield, s, salter, j, bath, pa., hubbard, b, millington, p, anders, jh, et al. 2014, ‘open access repositories worldwide, 2005-2012: past growth, current characteristics, and future possibilities’, journal of association for information science and technology, vol.65, no.12, pp.2404-2421. doi: https://doi.org/10.1002/asi.23131. rahman, mm, & mezbah-ul-islam, m 2014, ‘issues and strategy of institutional repositories (ir) in bangladesh: a paradigm shift’, the electronic library, vol.32, no.1, pp.47-61. doi: https://doi.org/10.1108/el-02-2012-0020. https://doi.org/10.29173/iq1012 https://www.researchgate.net/publication/327187267 https://doi.org/10.1108/lr-03-2017-0021 https://www.ajol.info/index.php/ajlais/article/view/164661 https://doi.org/10.1108/oir-03-2019-0087 https://digitalcommons.unl.edu/libphilprac/2592/ https://doi.org/10.1108/ils-07-2018-0066 https://doi.org/10.1353/pla.2003.0039 https://doi.org/10.1108/gkmc-04-2017-0033 https://doi.org/10.1108/dlp-04-2018-0011 https://digitalcommons.unl.edu/libphilprac/2444 https://v2.sherpa.ac.uk/opendoar/search.html https://doi.org/10.1002/asi.23131 https://doi.org/10.1108/el-02-2012-0020 12/12 mwalubanda, joseph mathew (2021), the development of institutional repositories in east africa countries: a comparative analysis of tanzania, kenya, and uganda, iassist quarterly 45(3-4), pp. 1-12. doi: https://doi.org/10.29173/iq1012 raju, r, adam, a, & powell, c 2015, ‘promoting open scholarship in africa: benefits and best library practices’, library trends, vol.64, no.1, pp.136-160. https://www.semanticscholar.org/paper/promoting-open-scholarship-in-africa%3a-benefitsand-raju-adam/99af69dbae19a5b69bc654148e58d9ad0d549ec4 registry of open access repository, http://roar.eprints.org. registry of open access repository mandates and policies, https://roarmap.eprints.org. rutanya, fc 2017, ‘institutional repository: access and use by academic staff at egerton university kenya’,library management journal, vol.38, no.4-5, pp.276-284. sahu, rr, & parabhoi, l 2019, ‘open access repository: a comparative study of germany, switzerland and austria’, library philosophy and practice, pp.1-9, https://digitalcommons.unl.edu/libphilprac/2511/ shajitha, c, & kc, am 2018), ‘content growth of institutional repositories in south india: a status report’, global knowledge, memory and communication, vol.67, no.8, pp.547-565. doi: https://doi.org/10.1108/gkmc-02-2018-0018 sharma, aj, meichieo, k, & saha, nc 2008, ‘institutional repositories and skills requirements, a new horizon to preserve the intellectual output: an indian perspective’, planner, pp.336-353. smith, ina 2015, ‘open access infrastructure: open access for library schools’, unesco publications. tapfuma, mm, & hoskins, rg 2019, ‘usage of institutional repositories in zimbabwe’s public universities’, south africa journal of information management, vol.21, no.1, pp.1-9. doi: https://doi.org/10.4102/sajim.v21i1.1039. ukwoma, sc, & mole, ajc 2017, ‘utilization of institutional repositories for searching information sources, selfarchiving and preservation of research publication in selected nigeria universities’, afr. j. arch & info. sc, vol.27, no. 2, pp.117-130. verma, nk, & shulka, a 2014, ‘evaluating growth and development of open access repositories: a case study of opendoar’, international conference on knowledge organization in academic libraries, pp.59-67, jaipur. wyk, b, & mostert, j 2010, ‘towards enhanced access to africa’s research and local content: a case study of the institutional depository project, university of zululand, south africa’, afr.j. lib & sc, vol.21, no.2, pp.139-151. xia, j, gilchrist, sb, smith, nxp, kingery, ja, radecki, jr, wilhelm, ml et al. 2012, ‘a review of open access self-archiving mandate policies’, portal: libraries and the academy, vol.12, no.1, pp.85102. doi: http://dx.doi.org/10.1353/pla.2012.0000 endnotes 1 joseph mathew mwalubanda is a librarian and researcher at the department of library, tanzania institute of accountancy and can be reached by joseph.mwalubanda@tia.ac.tz or mwalubandajoseph88@gmail.com https://doi.org/10.29173/iq1012 https://www.semanticscholar.org/paper/promoting-open-scholarship-in-africa%3a-benefits-and-raju-adam/99af69dbae19a5b69bc654148e58d9ad0d549ec4 https://www.semanticscholar.org/paper/promoting-open-scholarship-in-africa%3a-benefits-and-raju-adam/99af69dbae19a5b69bc654148e58d9ad0d549ec4 http://roar.eprints.org/ https://roarmap.eprints.org/ https://digitalcommons.unl.edu/libphilprac/2511/ https://doi.org/10.1108/gkmc-02-2018-0018 https://doi.org/10.4102/sajim.v21i1.1039 http://dx.doi.org/10.1353/pla.2012.0000 mailto:joseph.mwalubanda@tia.ac.tz mailto:mwalubandajoseph88@gmail.com vol223 4 iassist quarterly data liberation, bridges to cross by richard boily * abstract in canada, the use of statistical data (micro data files and major databases) for teaching and research is an important phenomena that does not seem to be losing strength in the near future. this situation is a major consequence of the data liberation initiative (dli), established in 1996 as a partnership among statistics canada, other federal departments and canada’s academic community. the idea of providing affordable access to canadian information results from a co-operative effort among the humanities and social science federation of canada (hssfc), the canadian association of research libraries (carl), the canadian association of public data users (capdu) and the canadian association of small university libraries (casul). less than two years after its inception more than 50 universities have joined the consortium, which is a clear indication of a true willingness to make data more available. this illustrates the fact that the high cost of buying data was an obstacle to its availability, especially in small universities where the absence of a minimum number of students results in a higher cost/benefit ratio related to data acquisition. however, there are still many obstacles to free numerical data use. if some canadian universities have a long history in data services (carleton university’s data centre celebrated its 30th anniversary in 1996), such a tradition does not exist everywhere, especially in small universities. to maximise use of data files, increased education at the reference staff level and at the consumer level, including professors, must occur. data usage requires a good knowledge of data extraction and associated analytical instruments. how can these tools be made accessible to customers who are not able to manipulate data files, but who have a definite need for the information? how can we satisfy different needs for different types of users? how can data be included in the academic curriculum? how can data librarians play their educational role and how can this role be balanced with professorial responsibilities? fortunately, interesting answers are unfolding. introduction numerical data collected from various surveys conducted throughout the nation by organisations such as statistics canada constitute an information source that is both important and extremely powerful in understanding social phenomena. access to these data is necessary for teaching and academic research. in fact, this access is essential to intellectual freedom and democracy. in canada, the conditions of accessibility to numerical data have undergone major changes within the past two past years, due to the implementation of the data liberation initiative (dli) by statistics canada. the data liberation initiative is a management framework that modifies the conditions of data access. these changes coincide, and certainly not by accident, with the arrival of new technologies: the development of both powerful personal computers and their increased data storage capacity and of user-friendly software (excel and spss), in addition to the advent of the internet. such events bear directly on the theme of this conference, notably global access and local support. with new parameters defined by dli, canadian data are potentially more accessible than ever to canadians. even if dli is successful, the fact remains that widespread numerical data use in canadian universities is uncommon and many obstacles exist that would allow the situation to change. the objective of this presentation is to examine problems of accessibility to numerical data within the context of dli. this presentation is composed of three parts. first, a brief history of the origins of dli and its role will be given. this will also entail a review of the objectives. then, the problems of developing numerical data use within the context of dli will be addressed. finally, it will be shown that there are elements of dli that represent an opportunity to improve the democratisation (accessibility) of data in canada. fall 1998 5 1.dli origins traditionally, statistics canada publishes the statistical information that it has collected in the form of aggregated data tables. these documents are largely distributed to libraries via the government publication deposit program. however, numerical data files that have been excluded from the deposit program and until recently, were available only at a very high price. such a situation has been strongly discredited by the research community, notably by professor paul bernard, professor of sociology at the university of montreal and a member of the national statistics council. in 1991, professor bernard asserted that, “the genuine exercise of democracy increasingly requires that citizens get access to complex information and have the skills required to understand it”. in 1993, following the opinions voiced by professor bernard and others, several individuals representing the social sciences and humanities research council of canada, the association of universities and colleges of canada, the canadian association of research libraries and the canadian association of public data users united under the auspices of the social science federation of canada. their specific objective was to develop a strategy to render canadian survey data more accessible to the research and teaching community. the work of this group led to a proposal that rapidly passed through the various levels of the federal government and thus, was accepted by statistics canada. the dli received official recognition from the treasury board of canada in february 1996. it was subsequently included as part of the canadian government’s science and technology strategy in march of the same year. what is the dli? the dli is a five-year project among universities, statistics canada and several federal departments. under the agreement, participating universities pay a known and affordable yearly fee ($12,000 for carl members or $3,000 for casul members), that gives them access to all standard data products provided by statistics canada. ftp on the internet is the primary method of accessing these files. if files exist only on cd-rom each participant is entitled to a copy. however, if they exist in both forms, users may choose one or both types of files. participating libraries must make acquired data available to their users while, at the same time, insuring that they do not use it for commercial purposes. objectives of the dli. as stated on the dli web site, itself, “timely access to data is essential if researchers are to focus on canadian problems and students are to learn to analyse canadian information. without affordable data for research and training, canada risks producing innumerate graduates and basing its policy decisions on incomplete information. independent analyses enhance public debate and policy making on questions relevant to all canadians. the federal government invests large amounts of public money in data collection. the dli can ensure a valuable return on this investment by distributing data to the university community, which will encourage analysis and put more information in the public domain”. before the inception of dli, we experienced the embarrassing situation where canadian researchers, who needed to develop methodological expertise or study a particular social phenomena, had to work with american data because statistics canada data were either too expensive or not available at geographically specific levels. unfortunately, this situation still exists. 2.numerical data use in the context of dli, or, how do things happen now? although dli has been in existence barely two years, it is showing positive results that are measurable and it seems to be fulfilling expectations. there is an excellent participation among canadian universities in dli that exceeds even the most optimistic expectations. before dli’s inception, barely 15 to 20 universities offered any form of numerical data service. this obviously does not take into account individual professors and researchers who ordered files from statistics canada. it is even probable that the acquisition of these data has been made through the library. it does not take account either of the various numerical data files, notably the 1986 and 1991 canadian census data distributed on cd-rom, that were made available by several libraries well before the inception of dli. for the past several years, we have offered training sessions on cd-rom census data search at our library. however, the fact remains that these transactions were not integrated within a real data service. today, more than 60 institutions participate in the dli consortium, which is nearly 80% (61/78) of canadian table 1. university participation to dli universities offering data participating universities services before dli in the data liberation initiative (dli) 1996 1997 1998 between 15 and 20 50 59 61 6 iassist quarterly universities. fifty of those affiliated have been with dli since its first year. however, it does not follow that all 61 participating institutions offer a numerical data service. this number simply indicates that there is, at minimum, a dli representative in each of these universities and that this person is involved, to one degree or another, in datarelated activities that may lead to the implementation of a data service. we have discussed institutional participation, but what happens to the data use level? in march 1998, more than 1,000 cd-roms had been delivered to participating universities. the number and growth of ftp file transfers rose from 10,000 in 1996 to 25,000 in 1997. obviously, these transfers were not carried out exclusively on data files, but also on command files and text files (e.g., code books, readme files, etc.). also, transfers do not focus only on micro data files. aggregated census data accounts for a high proportion of transactions. looking at the previous data about dli, one can advance some observations: 1. perhaps some data would not have been ordered due to cost, even by libraries that already have data use experience; 2. one can suppose that the global cost would have been far more important if all these data had been acquired individually; 3. it is probable that several files have been downloaded or that cd-rom products have been ordered by libraries that, until now, had little used numerical data. therefore, it seems that dli contributes largely to data distribution and that the program replies to a real need. the objective to make numerical data accessible at a reasonable price is, therefore, partially attained. on the other hand, it is important to remember that dli was not conceived merely to reduce the cost to those already using data. the real objective is to increase data use by the whole community, to see a real expertise developed in data use, and to enable a greater number of studies that focus on canadian society to be conducted. this is the real meaning of accessibility and the democratisation of data. context of numerical data use even if statistics canada survey data are potentially available (data can be downloaded at any given time by whoever needs it) and the price no longer constitutes an obstacle to use, both the intrinsic complexity of the data and the means of exploiting it remain major constraints for its use. the majority of users are, in fact, incapable of manipulating raw data. their need for a data service before the arrival of massive data sets occurs is more important than ever. however, the costs associated with setting-up and maintaining such a service are considerable and largely exceed the cost of the data. these costs constitute a major obstacle to the democratisation of data. variety of resource persons the participation of new institutions in dli and, the consequent arrival of new representatives are promising events for the development of new data services. on the other hand, all those who have accepted to be responsible for data files do not have, necessarily, the same level of competence. not all have the same interest or desire to develop the expertise required by this new function. the dli is a young program that has experienced accelerated development (50 members the first year), probably due to a copy-cat effect. as we have previously stated, it is necessary to remember that statistical information distributed by statistics canada is traditionally published first as working documents and then, these documents are often integrated into governmental publications. for librarians who specialise in governmental publication reference, the responsibility of numerical data management constitutes a sizeable challenge. in many cases, and i have to recognise that it was true with me, new dli representatives came to the job previously unaware of what was entailed with numerical data files. numerical data exploitation and use, especially micro data, supposes a knowledge of computers and statistics that is often deficient in both the users (students at all levels and good number of professors) and the library personnel. in terms of helping the clientele, it seems apparent that consultation services must be collaborative efforts involving both the computer service personnel and the professors and researchers. among them are specialists in computer science and statistics. but, if one believes that users in search of numerical information are better served in a library (and do not forget that these are dli participating libraries), it appears problematic to me to cd-roms delivered to files downloaded from the ftp paricipating univesities to dli site of statistics canada 1035 1995 1996 1997 887 10173 24384 table 2. numerical data use in the context of dli. fall 1998 7 think that data information specialists must always cater to other professionals. collaboration is essential, but total dependence should be limited. training needs are an important, yet considerable cost of rendering data more accessible. training to satisfy their training needs, dli representatives can count on a continuous training program put in place by the dli external advisory committee. in addition to this national program, local organisations (such as the council of prairie and pacific university libraries’ (copul) consortium of library electronic data services (accoleds), or the working group on data of the conference of rectors and principals of quebec universities) also organise training activities. since the inception of dli, several training activities have already taken place. at the initiative of the advisory board, 4 dli initiation workshops (in fact, the same workshop repeated 4 times) have taken place in four canadian cities during 1997. these workshops brought together 120 individuals who until then knew almost nothing about numerical data. in evaluating the workshop, participants stated their preference for the organisation of additional workshops on more specific themes. as a result, new workshops will take place this spring on the use of spss for data processing and numerical analysis. in fact, one of these workshops was held two weeks ago in montreal. conducting training sessions raises both challenges and opportunities for librarians interested in promoting numerical data use. it is a challenge because it is necessary to develop a certain level of competence before being able to teach. but challenge aside, the possibility of teaching users represents a privileged opportunity to assume our role as information specialists. instead of waiting to be asked for information, we can create a demand for information. how can users ask to use numerical data if they do not know it exists or if they can not use it? far from me to suggest that librarians replace other professionals or professors. however, i believe that a solid knowledge of data and of the tools of exploitation from which arises a need for training is necessary. first, it allows us to exercise the educational aspect of our job and, second, it enables us to become knowledgeable spokespersons alongside professors and other professionals. it is only with a solid knowledge of data that we can become counsellors for users, orienting them towards the best data sources or advising them on the best manner of exploitation. the competence of data librarians is as necessary to establishing bonds with professors and researchers, as is identifying the main research areas for which data are required. there certainly is not unanimity among colleagues on the way in which these new responsibilities will be handled nor, consequently, on the usefulness of advanced training. some will never be able (nor want) to develop statistical expertise. the problem is complex and there is certainly no correct reply, but it will have to be considered and a great deal of progress will have to occur. no matter what others decide to do, some libraries, such as the carleton university data centre, have already specialised in numerical data and offer complete data services that include computer and statistical assistance. such levels of service are not widely found. even large universities do not always offer complete data services. all the new libraries that now play a role in data specialisation have not and will not develop such an expertise. but is there a middle road and where is it situated? if data accessibility is essential to democracy, the inequality of services offered constitutes an important obstacle to exercising our rights. equality of access within the dli framework, the choice has been made to insure numerical data development in libraries. considering information needs in general, researchers in small or regional universities are no longer penalised with regard to information accessibility. with the development of new technologies, such as the internet and the emergence of periodicals and other electronic publications, the availability of large bibliographical data bases on either cd-rom or the internet (uncover) and finally, with the development of increasingly specialised inter-library loan services (ariel), the disparities have lessened between the large and small universities, between urban universities and those in the regions. this is true even if the level of document availability remains variable at the local level. the situation with regard to numerical data availability is entirely different. it is obvious that students at universities that offer well structured data services are in a favoured position over their colleagues who do not have such access. the other libraries simply do not offer the students the same level of support. in fact, one can assume that regional disparities were even greater before the inception of dli and that the program will attenuate these differences. in this regard, the question of training for librarians and their perceived role in offering these new services is an important consideration. regional disparities another aspect related to the disparity of services offered has to do with the regionalisation of the data. regional universities are often located in areas that are characterised by a low population density. research work related to regional problems require data over geographical areas that are not comparable to data that defines large metropolitan 8 iassist quarterly regions. this is the problem experienced at the university of quebec at rimouski, which has masters and ph.d. programs in regional development. the users from these groups need not just numerical data, but numerical data at a specific geographical level. unfortunately, the data available to these researchers are often over too wide a geographical level. even if more specific data exists, they are not available for their use. this is the problem, for example, with the large survey of consumer finances and with the national population health survey. if the objective of dli is to increase data access in all universities no matter where they are located, the question of geographical data specificity is important, even though we recognise that solutions will not come directly from dli. indeed, the mandate of dli is to give access to standard data products provided by statistics canada. the notion of standard products is also linked to the question of data confidentiality. as a citizen, one can only rejoice in observing that statistics canada respects norms of strictest confidentiality and that these principles should never be challenged. from the academic viewpoint, the problem is not less important. on the other hand, it is probable that the successes of dli will increase demand for this level of data. the question of confidentiality is also important from the standpoint of the installation of data services. it is not sufficient to simply initiate users to the data and to analytical instruments, such as sas or spss, if one cannot also provide data that satisfies their research needs. there is a risk of losing hard-gained credibility if users do not have access to data that they know exist, especially after they invest considerable energy in learning complex instruments. conclusion previous statistics have shown that dli has had real success; a success that exceeds the hopes of many. but beyond all the figures relating to file transactions and considering the current context of data use described here, one of major successes of dli has been to enlarge data access to a greater number of individuals interested in using numerical data in canadian libraries. these people share considerable amounts of information (via the two list servers that help in administration of the program). several of them have had the opportunity to meet during workshops. the community is, therefore, in the process of widening and there exists a tangible willingness among colleagues who are more experienced to share their expertise. finally, i would like to mention that universities in quebec are full participants in this process. all participate in dli. members of the sub-workgroup on numerical data files are presently working on the development of an identification and data extraction system that will facilitate and stimulate data use. *paper presented at the iassist conference, may 21, 1998, at yale university, new haven, connecticut. richard boily, université du québec à rimouski, and member of the conference of rectors and principals of quebec universities (crepuq) working group on data 4 iassist quarterly fall & winter 2007 editor’s notes welcome to the iq volume 31 double issue 3&4. this marks the end of the 2007 iassist quarterly. this also marks the end of printing of the iq. with this issue the iq will only be available on the web-site, available for reading on-line and also for downloading and printing. the effort of bringing earlier issues of the iq as scanned version to the web-site is continuing. this double issue is the work of the authors and their articles are introduced below. we are presenting an integrated double issue of high quality. we should also give a special thanks to the editors of the issue. gretchen gano is the writing guest editor of this iq as you can see below. gretchen gano is the assistant curator librarian for public administration & government information and coordinator, data service studio at new york university libraries. gretchen gano collaborated on this issue from the start with former iassist president ann green. together with the authors a great issue has been made. if some of you would be interested in compiling issues for the iq as guest editor(s) please contact me. if you don't have anything to offer right now, then please prepare yourselves for the coming iassist 2009 if you are acting as chair for a session there. that is an obvious opportunity to bring quality sessions to more people than the session participants and also making the memory more sticky. take a look at the website http://iassistdata.org and the iassist blog the iassist communiqué – at http:// iassistblog.org. articles for the iassist quarterly are very welcome. articles can be papers from iassist conferences, from other conferences, from local presentations, discussion input, etc. contact the editor via e-mail: kbr@sam.sdu.dk. karsten boye rasmussen december 2008 introduction to the iassist quarterly special issue on data services and institutional repositories welcome to this special issue of the iassist quarterly. this time we turn our attention to the broader environment for digital data collections. contributors provide examples and analysis about how data makes its way between and among institutional collections in a complex data ecosystem. data does not exist, nor does it move anywhere on its own, of course. each author in this issue identifies challenges faced by those who participate in and support the research enterprise all along the way. the complexity of this ecosystem seems to increase as more institutions wish to share collections with one another and to migrate research data across digital repository systems. data services and repository staff recognize the importance of designing systems to increase faculty participation and to move the data management partnership further “upstream” in the lifecycle to when the plan for designing questions and collecting data is first hatched. system designers find themselves adjusting existing technologies to optimize environments to care for research data as it is created, used, shared, and saved over the long term. iassist quarterly 2015 39 iassist quarterly abstract a tailored approach is ideal for teaching users to work with research data, which often varies significantly by domain and project depending on methodology, available data sources and intended outcomes. in this paper and presentation, three distinct contexts will be put forth, each using geographic information systems (gis) and focused problem-based learning (pbl) approaches to teach research data use: primary collection, digital data reuse and mined textual data. in each illustration, researchers are not only working to implement a functional methodology, but also to engage students in practices that equip them with theory, tools and skills to advance their own research trajectory. further, these examples are from researchers in distinctly different disciplines: an architect working on climate change in the st. louis region, three historians reconstructing history with data from texts and a professor of social work collecting data for villages in india. the data and gis services (dgs) team at washington university in st. louis (wustl) has partnered with each project presented to support analyses, visualization, management, preservation and sharing of research data. methods, challenges and opportunities are discussed.. keywords research data, problem-based learning, gis, higher learning introduction what is research data? the definition of research data is nebulous because what data means to each knowledge domain differs. the federal government defines research data as ‘the recorded factual material commonly accepted in the scientific community as necessary to validate research findings.’ (white house, 2015). however, expanding the definition to include scholarly communities is more appropriate for an academic environment. it is widely agreed that data can be collected in a number of ways including by observation, data mining, modeling and from referential sources. while the form and function of data varies between disciplines, as research grows more collaborative and cross-disciplinary, interest in new and combined approaches to data and technology grow in tandem. quantitative skills are increasingly sought in many professional fields (uttl, 2013). while the call for ‘data science’ skills is resounding, learning outcomes are often difficult to define due to the noted variability (wlodarczyk and hacker, 2014). barriers to technical learning self-efficacy the american psychological association defines self-efficacy as an individual’s belief in his/her ability to produce specific performance attainments, which affect that individual’s efforts and likelihood of success (american psychological association, 2015). students who have not worked with data in the past have expressed insecurity in learning to use data and technical skills. some students state ‘i’m not so good with computers’ and whether that assessment is true or not, it speaks to their self-efficacy. a pedagogical approach that focuses on students progressively mastering increasingly difficult tasks through a sequence of steps impacts efficacy. social learning is also impactful as it allows students to observe peers struggling with materials and mastery of tools, as are they, which gives them a better sense of self-efficacy (bandura, 1982). technical skills working with research data requires a number of skills that many students have not learned in previous educational experiences. these teaching users to work with research data: case studies in architecture, history and social work by aaron addison1 and jennifer moore2 40 iassist quarterly 2015 iassist quarterly skills can include basic data finding and cleaning, data and file management, analyses and visualization. skill levels vary, but even some very bright students have trouble with mundane tasks (e.g., downloading, moving and unzipping a folder), which are not intuitive. connection to real world scenarios skills taught in a vacuum are usually not easy for students to retain. students receiving one-off instruction on data usage without application to a problem in their domain may have difficulty developing skills to work meaningfully with research data. blended approaches to pedagogy learning data science through gis geographic information systems (gis) are a combination of technologies, both software and hardware that allow users to describe, analyze, manage and visualize space using a myriad of data, including but not limited to spatial data. but this should not minimize the importance of space in gis; in fact, reed (2014, p. 280) calls location the ‘great data integrator’. location offers students who are new to data a familiar anchor for understanding. throughout life, everyone uses data to make decisions and maps are commonly used to help audiences understand various data problems. many individuals are unwittingly consuming data visualizations and analyses daily, but learning to make maps and see the data at the foundation allows students a better understanding (drennon, 2005). the foundation of problem solving with gis is spatial thinking. spatial thinking is the process by which students use the concept of space and the tools of representation to answer questions. the first function of spatial thinking is to define space and describe the objects or movement within it. secondly, it functions to analyze the structure of objects in the space. these functions, facilitate making inferences, predictions and building arguments. by learning to apply spatial thinking to problems, students develop an understanding of the nature of data, how and where to find or create it, how to assess it and attribute it, how to build arguments and look critically at the arguments made in spatial representations (national academies press, 2005). keenan and fontaine (2012) describe combining methods of inquiry-based pedagogy, where students must be able to pose questions, find and assess reliable data and understand the interconnectedness within the data (e.g., objects in their environment) with student-centered instruction. problem-based learning and gis the approach of problem-based learning (pbl) directs students to work through complex, real life problems and attempt to solve them, often collaboratively. as barrows (1986) explains, there is not one single pbl methodology, but the common thread in all of them is using problems as the basis of instruction. often, pbl requires students to self direct, digest, reflect and identify knowledge and tools they may need. instructors are tasked with creating problems with specific, yet implicit, learning outcomes that students can reach through the problem solving activity. instructors may help students develop necessary skills to solve the problem, but don’t include a step-by-step guide. this approach requires instructors to be ready to let students fall off-course and find their way back. research suggests that pbl may facilitate a better conceptual understanding of their discipline as well as the development of soft skills (allen et. al., 2012). groups of students working collaboratively on a problem may increase their understanding and belief in their abilities to work through table research data learning outcomes using gis data information literacy sourcing, understanding and using research data data cleaning correcting or removing records in a dataset; making the dataset operable data and file management organizing data in a structure that makes it easy to access and share combining datasets adding or joining relevant data to an extant, primary dataset querying datasets creating expressions to draw out particular data from a dataset analysis skills showing relationships, regression, clustering, etc. through analytical processes data modeling developing a structure that reflects how data and datasets will relate to others data collection gathering and inputting data data creation inputting or drawing new datasets in a workspace deriving data creating or modifying a dataset from a selection of a larger dataset visualization displaying selected data in a way that communicates an intended message assessment of arguments critically evaluating the products of others using data based on acquired knowledge and skills data attribution identifying and expressing data sources 1 it. inherently, pbl allows for students to take problems in any direction; it’s noted that this can make a classroom chaotic (white, 1996). combining spatial thinking and pbl allows students from a number of disciplines to work with and understand data and analysis required to solve problems. drennon (2005, p.397) found that by integrating pbl into a spatial project students became competent with data and gis analyses ‘almost by accident’. because gis is a complex tool on its own, integrating it into pbl requires scaffolding to give students footing in basic gis skills and tasks before introducing the problem. scaffolding, first articulated by wood, bruner and ross (1976), introduces students to tasks just out of their capability and assists them in completing the tasks; as tasks are sequentially mastered, assistance for that task fades. howarth (2011) suggests that gis learning requires introduction to core, sequential skills working through guided problems and solutions first. following the introduction to basic tools, the instructor may hand students a problem, which can be solved by building on those skills and with minimal guidance. this approach to research data learning through problem-based, scaffolded gis instruction can bridge the gap between novice students and research data. gis learning encompasses many of the needed outcomes and through the problem based approach iassist quarterly 2015 41 iassist quarterly students connect to tangible problem solving and therefore better retain the technical skills. problems and praxis – case studies history and mined textual data close reading combined with gis building on constructivist and active learning theory, calandra (2005) described a model of learning using digital history resources and methods to help students develop their own understanding of history. to that end, technology employed in the classroom must be authentic, flexible, scaffolded and foster creative, independent thinking. three case studies are presented in this history section, all of which are based on a problem that students were introduced to and investigated in a history course; one uses historical legislation, another uses a holocaust memoir and another is based on research about a territory in dispute. history problem 1 turning historical legislation into digital data the first case study is based on a project developed by a faculty member to visualize the establishment and growth of the early u.s. federal government. the project has huge amounts of data to be digitized, attributed and mapped. in the first phase of the project a graduate student built a working model from mined legislative documents and existing legislative data. this student had been involved in a summer internship in the humanities digital workshop (hdw), which supports long-term projects in the digital humanities on campus. the student had some basic experience with digital data, but not gis. with assistance from the data and gis services (dgs) team, which operates out of the libraries, this student built a proof of concept mapping project. working from that proof of concept, a team formed between the faculty, the hdw and dgs to develop a framework that would drive the project vision forward and also to introduce a pedagogical experience for undergraduates. the students would be introduced to research data, gis skills and methods through the lens of history. meeting these students in a familiar domain of study and working on content they understand creates an opportunity for conceptual understanding of how this data fit into the historical equation. upwards of thirty hours were dedicated to scaling the project goals into a classroom learning experience, including designing a flexible, data model that could fit into the scope of the class project, data preparation, looking for synergy with an existing database and working through example legislations to understand the parameters of the problem. the data/gis portion of this class took place over six sessions. basic skills were introduced in session one, and sessions two through four were designed to equip students with the specific skills needed to work through a piece of legislation in sessions five and six. the importance of sequence and scaffolding became apparent during the classroom experience. in some cases students moved into tasks before they were able to tackle the technology, but the students did have a very clear idea of the conceptual purpose of the exercise and what the data could do. by the final session, students became more comfortable manipulating the data. while using gis to input and visualize this data made sense to the students, they also became aware of the hiccups that come along with that. students were particularly engaged by the organization of the data model. they had questions regarding how to delineate between data that can be represented through a spatial feature and data as an attribute of a feature (e.g., when is a jurisdiction a boundary and when is it a level in the bureaucracy). learning outcomes included: data creation, data and file management, data editing, data literacy and data modeling. history problem 2 making an argument using modern digital data and historical maps the second case study was built on a historical problem that has bled into modern-day disputes over the senkaku/diaoyu islands off the coast of east asia. currently taiwan, japan and china have claims to these uninhabited islands. the professor teaching this class wanted to introduce the dispute to freshmen and ask them to make arguments for different sides of the dispute. part of that argument had to be expressed in a map they created that included both modern data and historical maps. this course took place over eight sessions after intense study of texts on the conflict. session one through three focused on data and gis skills, which were presented sequentially and based on the functions of spatial thinking, to describe, analyze and make inferences in space. while moving through these skills students were reminded what functions the skills relate to. sessions one through three focused purely on description, placing data, georeferencing, data editing and creation. sessions four and five focused on understanding relationships through spatial and attribute data queries, basic analyses (e.g., buffering features), data finding and metadata. in session six, skills were reviewed. session seven and eight served as guided practice so students could use the data and skills acquired to work in teams on their problem, to build an argument for a particular side of the dispute. file management was a challenge to these students. concepts of file naming, hierarchy and versioning were introduced in session one, but were difficult in practice. working with zipped folders also presented challenges. while instructors focused on scaffolding the data representation and analyses methods, more attention to basic practices of file management and versioning was needed and will be emphasized in future sessions. learning outcomes included: data creation, data literacy, data and file management, data editing, data analysis, digitization, and making an argument using data. history problem 3 using a memoir and digital data to recreate a journey & human experience in this case study students began working from a text written by a holocaust survivor chronicling her experience moving from her home, between camps and finally home again. students were tasked with visualizing this experience using something more than dots on a map; the goal was to make a map that depicted something students deemed impactful about the journey, for example emotions (e.g. hope or fear), languages spoken, separation from families. etc. students used methods like applying number values that described levels of emotion felt at each location based on their assessment of the text and then applied meaningful symbology to express it. instruction took place over three sessions followed by three open guided-practice sessions. in session one, students learned to organize their data into features with attributes that they placed in a table; between session one and two they populated that table with whatever they deemed impactful. in session two they combined their data with established geodata and described the area. in session three students learned techniques for visualizing a map that tells a story. in this case, with only three short sessions to learn skills, students were left alone to tackle their problem. some students made use of group open editing sessions designed for guided practice, but many needed individual appointments as well. in this case greater attention to the pace of scaffolding was needed, but in the process students did become more familiar with data and produced a unique map based on the attributes they selected 42 iassist quarterly 2015 iassist quarterly to visualize. learning outcomes included data literacy, data and file management, combining datasets, data creation, editing and visualizations. social work – teaching through primary collection in the field the field of social work has continued its move towards evidencebased research over the past several years. this shift has required students to develop proficiencies in primary data collection and data analysis. although these skills are important, there continues to be debate about what research skills social work students need. for students in the uk, requirements encompass capabilities to interpret information through collection and analysis and an ability to assess materials and data (macintyer and paul 2011). data in the social work domain originates in numerous forms and much of it is applicable for gis analysis. felke (2014) makes an argument for the importance of gis to the social work student’s dossier, and he embedded within his own undergraduate course data literacy, data creation, aggregating data, data management and visualization outcomes. for some students such skills led to employment, and for others the skills were utilized for further study. a social work case study introduced students to primary data collection fieldwork for a large-scale project in india to map villages with gps devices and using colloquial understanding of the village and locals’ sense of place. the field experience required students to contemplate how real world phenomenon will be represented in gis and what will be valuable in their research. once back in the classroom, all data points were mapped using google earth and gis files generated for use in arcgis. faculty worked closely with dgs on design and delivery of workshops to develop methodology for teaching students and locals to research, collect, process manage and archive data, including de-identification of data. a key aspect of this project was for students to learn the value of place in the context of analysis and decision-making. gis was not the focus of their research, but rather a data visualization and analysis tool to support evidence based research. learning outcomes include data literacy, data and file management, data collection and cleaning. architecture digital data reuse aggregation for collaborative workshops monsur and islam (2014, p.49) argue that architects and landscape architects are in a position to make better and more informed designs based on available digital data. in this case, data is not limited to features in the landscape, but also includes sociocultural, socio-economic, behavioral or demographic information and the contextual relationships across a regional area. however, while using data, as a part of architectural decision-making is not necessarily common practice, monsur argues that data and gis methods should be an integral part of the planning process both to design efficiently and to decide whether it makes sense to design at all. in the sam fox school of design and visual arts at wustl numerous faculty have embraced spatial data and gis approaches in their pedagogy. while several instructors have invited the dgs team in for talks and one time instruction on using gis methods and data, some have dedicated large portions of instruction to data analysis using gis. one example where this approach was successful involved two faculty members collaborating to deliver a multidisciplinary workshop focused on flooding and climate change in the st. louis region. the workshop also investigated the regional relationship to the changing environment, both culturally and economically. students utilized sourced and extant landscape feature data as well community data they collected. source data used was primarily publicly available hydrologic, levee, boundary, agricultural and soil data. faculty created base maps and used georeferenced photographs to record data from members of the community and other stakeholders. students participated in various scenarios throughout the workshop and used the data to work through problems, including proposing future architectural development and addressing agricultural, ecological and navigational challenges. the workshop outcomes were models of potential future consequences based on different data-driven decisions. workshops like this one not only incorporate the idea that data can be applied to structural planning, but they illustrate that useful data comes in a variety of forms, from river levels to microdata from local residents. it introduced the concept that data can be sourced or organically generated and that it requires management. lastly, it presented the idea that factoring new data variables could change the parameters and outcomes of student design. this problem-based experience provided students tools and skills required to explore potential directions of architectural and landscape development, which created a rich learning experience for everyone involved. learning outcomes include data creation, sourcing data, data and file management, data editing, data literacy, data analysis, digitization, and making an argument using data. conclusion as shown through these case studies, research data learning outcomes were achieved through problem-based gis learning. this evidence demonstrates how this approach can be an effective method to introduce students to skills and tools needed to work with research data. however, each case presented is an example of the first cycle of instruction. at least four of the five cases will be repeated, offering an opportunity for enhancement. instructors have a better sense of the students’ starting point and, following further assessment, can tailor materials in collaboration with faculty to better match learning outcomes. determining the lasting effects of teaching the use of research data through gis problem based learning will require more local assessment, but the techniques outlined here are working toward a higher rate of long-term retention. gis can be applied to all sorts of problems; therefore it is a very suitable medium to deliver this material in various domains. drennon (2005) asserted that gis in research data management and modeling furthers scientific enquiry. with the growing call for skills around research data use, visualization, analysis, management and sharing from every corner of professional and academic life, the types of learning experience described can launch students forward to meet that call in various disciplines. acknowledgements we would like to acknowledge the faculty members involved, dr. derek hoeferlin, dr. peter kastor, dr. tabea linhard, dr. lori watt, dr. gautam yadama as well as doug knox and bill winston. we would also like to thank our colleagues who reviewed our manuscript, ruth lewis, cynthia hudson-vitale and bill winston. bibliography allen, d. e., r.s. donham and s. a. bernhardt. (2012) problem-based learning. new directions for teaching and learning, 2011 (128) p. 21–29. available from: http://onlinelibrary.wiley.com/doi/10.1002/ tl.465/abstract iassist quarterly 2015 43 iassist quarterly american psychological association. (2015) self-efficacy teaching tip sheet. available from: http://www.apa.org/pi/aids/resources/ education/self-efficacy.aspx bandura, a. (1982) self-efficacy mechanism in human agency. american psychologist 37 (2): 122–47. available from: http://dx.doi. org/10.1037/0003-066x.37.2.122 barrows, h. s. (1986) a taxonomy of problem-based learning methods. medical education 20 (6): 481–86. available from: http://onlinelibrary. wiley.com/doi/10.1111/j.1365-2923.1986.tb01386.x/abstract drennon, c. (2005) teaching geographic information systems in a problem-based learning environment. journal of geography in higher education 29 (3): 385–402. available from: www.tandfonline. com/doi/pdf/10.1080/03098260500290934 felke, t.p. (2014) building capacity for the use of geographic information systems (gis) in social work planning, practice, and research. journal of technology in human services 32 (1-2): 81–92. available from: http://dx.doi.org/10.1080/15228835.2013.860365 howarth, j. t. and d.sinton. (2011) sequencing spatial concepts in problem-based gis instruction. procedia social and behavioral sciences, international conference: spatial thinking and geographic information sciences 21 (2011): 253–59. available from: http://www. sciencedirect.com/science/article/pii/s1877042811013656 keenan, k. and d. fontaine (2012) listening to our students: understanding how they learn research methods in geography. journal of geography 111 (6): 224–35. available from: http://dx.doi. org/10.1080/00221341.2011.653651 monsur, m. and z. islam. (2014) gis for architects: exploring the potentials of incorporating gis in architecture curriculum. arcc conference repository. available from: http://www.arcc-journal.org/ index.php/repository/article/view/250 macintyre, g. and s. paul. (2013) teaching research in social work: capacity and challenge. british journal of social work 43 (4): 685–702. available from: http://bjsw.oxfordjournals.org/content/ early/2012/03/22/bjsw.bcs010 national academies press, ed. (2005) learning to think spatially. washington: national academies press. available from: http://www. worldcat.org/oclc/63674143 rickles, p. and c. ellul, (2014) a preliminary investigation into the challenges of learning gis in interdisciplinary research. journal of geography in higher education 0 (0): 1–11. available from: http:// dx.doi.org/10.1080/03098265.2014.956297 reed, carl. (2014) ogc standards and geospatial big data. in: karimi, h. (ed.) big data: techniques and technologies in geoinformatics. boca rotan: crc press. available from: www.crcnetbase.com/doi/ abs/10.1201/b16524-15 uttl, bob, carmela a. white, and alain morin. (2013) the numbers tell it all: students don’t like numbers! plos one 8 (12). available from: http://journals.plos.org/plosone/article?id=10.1371/journal. pone.0083443 white, h. b. (1996) dan tries problem-based learning: a case study. in: richlin, l. (ed.), to improve the academy vol. 15 (pp. 75 91). stillwater, ok: new forums press and the professional and organizational network in higher education. the white house. (2015) federal register notice re omb circular a-110. available from: https://www.whitehouse.gov/omb/ fedreg_a110-finalnotice wlodarczyk, t. w. and t.j. hacker. (2014) problem-based learning approach to a course in data intensive systems. 2014 ieee 6th international conference on cloud computing singapore 2014. ieee. available from: http://ieeexplore.ieee.org/xpl/articledetails.jsp?r eload=true&arnumber=7037788 wood, d., j.s. bruner and g. ross. (1976) the role of tutoring in problem solving. journal of child psychology and psychiatry 17 (2): 89–100. available from: http://onlinelibrary.wiley.com/ doi/10.1111/j.1469-7610.1976.tb00381.x/abstract notes 1. director of data and gis services at washington university in st. louis, 1 brookings dr., campus box 1169, st. louis, mo 63130; aaddison@wustl.edu 2. gis & data projects manager & anthropology librarian at washington university in st. louis, 1 brookings dr., campus box 1061, st. louis, mo 63130; j.moore@wustl.edu data abstracts by john b. kolp laboratory for political research un i ve rs i ty of 1 owa union members and leaders: an international perspective the data files described below contain information on the opinions, attitudes, and background characteristics of labor/trade union members and leaders in five countries. the complete addresses for the data archives holding these materials are included at the end of this section. the abstracts have been constructed from documents supplied by these archives. union representation elections and the role of the national labor relations board jeanne herman brett, julius g. getman, and stephen b. goldberg population: sample of workers cases: 1239 variables: 16^4 archive: icpsr and ssda (illinois) this study is based on extensive interviews conducted with over 1,000 workers who participated in union representation elections in the united states. the original research investigation examined the influence of the national labor relations board (nlrb) on these elections. workers were questioned about their past experiences with unions, their feelings towards unions, and their observations of pressures exerted by companies, unions, or the nlrb before and after the representation election. data concerning the background and other characteristics of each of the workers were also collected. united auto workers: work group influence and political participation (detroit area study. 1961) warren miller and donald stokes time period: winter i96i population: daw workers cases: a19 archive: icpsr ' as part of the i96o-i96i detroit area study, uaw workers were interviewed in 'the winter of i96i. respondents were asked how long they had worked on their job, kwhat their job duties were, and whether they were satisfied with their job. another set of questions covered their length of union membership, their union activity, their conceptions of what the role of their union should be and their satisfaction with the job their union was doing. political questions covered the good and bad points of political parties, the kennedy-nixon debates, and political issues facing the nation, party identification, past and present vote in state and national elections, and political participation. the social structure of the work group was probed and the respondent was questioned about the importance of politics in work group relationships. demographic variables included class, age, organizational membership, religion, education, occupation, income, and race. new york state teachers survey, 1975 louis harris and associates, inc. time period: february, 1976 population: new york state united teachers union members archive: ssdl/harris date center (unc) the survey consists of eight samples of new york union teachers: new york city upstate. new york city chairpersons, upstate local presidents, suny, uup, and cuny psc. survey of teacher's union members investigates attitudes toward higher education in the state and the job the organization is doing. questions include rating of ntsut in representing its members' interests, satisfaction with leadership, publications, and membership benefits. other areas covered include tenure, united farm workers, abortion, detente. middle east, national health plan, tuition, and collective bargaining. survey of christian national trade-union, 1967 j. ramond (vrije unlversit) time period: 196? population: members, board members, and potential members of christian national trade-union (netherlands), born between 1907 and 19^2. cases: 582 archive: steinmetz a study of members, board members, and potential members of the union about the meaning of such a trade-union for them, the goals of such a trade-union, functioning, etc. '? dutch steel workers survey, 1973 p. van den eeden time period: 1973 population: workers at dutch national steel hooqovens cases: 200 variables: 55 archive: steinmetz working situation at dutch national steel hoogovens , attitudes to work-situation and policies, process of innovation and consequences for job security, work contracts, reorganizations, shift system, wage system, workload, independence, scope of control, problems at work and solutions, protest action, work experience, participation in strikes, negotiations, role of trade-unions, type of actions against the strike by management, role of judge in breaking strikes, role of worker councils, having a say in work the department contracts; attitude to worker-management relations, to management policies, motivation to be trade-union member. substudy 3 concerned function-groups and how one is alloted a place therein, which points of view are most important, independence, information, experience, opinions about own functiongroup, social contacts, type of judgements/ratings, influence, perception of opinions of boss and colleagues, attitude of the trade-unions toward reorganizations and management policy, and solving conflicts through actions and strikes. i er^.ployee opinion survey, 197^ time period: 197^^ population: employees of s tork-apparatenbouw (netherlands) cases: hs variables: 173 archive: steinmetz ' opinions of workers/employees on industrial relations; details on works council, copartnership, participation, working conditions, job satisfaction, trade unions; strikes, workers taking over plant, who must play a role in achieving industrial democracy; most opinions asked in relation to and sometimes after exposure to tv program about subjects as above mentioned. ii union survey on education and work david stern (yale) population: 16^4 accountants, 2u college office assistants, a27 social service supervisors, and 90 nurse's aides in new york city area. cases: 895 archive: ssda (yale) data were prepared for a study entitled "education, pay and job satisfaction," which was funded by the national institute of education. the same questionnaire was administered to accountants, college office assistants, and social service 10 supervisors. a similar but different questionnaire was used with nurse's aides. the codebook for the pooled file shows the relationship between the tagged variables for each file. attitudes of industrial clerical workers to white collar unionisation monica p. shaw, p. bowen , and v. elsy time period: december 1 972-decembe r 1973 population: clerical employees in six firms cases: 575 archive: ssrc survey archive (essex) the purpose of the study was to investigate the attitudes and reference groups of clerical employees in different employment situations. the firms were selected to give as wide a range as possible of different employment and trade union situations. attitudinal and behavioural questions included type of firm, job, department and employment history. satisfaction with pay and work conditions, assessment of job satisfaction and fairness of oay , comparison of pay with other workers. experience of regrading (opinion and assessment), prospect of change in clerical work, nature and effect of changes. assessment of present and desired relative influence of different types of employees in the firm. trade union membership (past and present), reasons for leaving/changing trade unions, whether office held in union, ideal type of union. opinion on trade union affiliation with labour party, ideal characteristics of a trade union for clerks, reasons for joining trade union, frequency of attendance of union meetings at work/outside work (reasons), trade union literature read. respondent's perception of union function (present and desired), satisfaction with union representation, whether other/no union membership preferred, willingness to participate in official industrial action. r's assessment of union's recruitment methods and growth, whether respondent felt involved in union affairs, industrial action appropriate for clerks, opinion on pressure for membership. experience of problems at work (type, outcome, personnel involved). knowledge and opinion of equal pay act, source of information, attitude to women at work/various social classes. background variables are age, sex, marital status, number of children, school leavi ng age, educational qualifications, political support, income, subjective social class. employee survey of working conditions k. w. redder time period: 1973 population: sample of members of a number of danish trade unions cases: 6931 variables; 272 archive: danish data archive the aim of the survey is to map the experience of members of trade/labour unions regarding working conditions. the questionnaire asked for former and present job; with regard to the latter, information was sought as to the length of time that r 11 had held this job, what inconveniences were connected with it, whether protective measues were orescribed and used, and about contentment and stress on the job. the last part of the questionnaire probed for possible illness or ailments and asked r whether the causes for these were to be found in r's working conditions or e 1 sewhere. workplace industrial relations: trade union officers office of population censuses and surveys time period: april-june 1973 population: sample of trade union officers cases: 127 archive: ssrc survey archive (essex) the purpose of the study was to monitor the effects, at workplace level, of changes in legislation in the industrial relations field; also the changes which may have taken place in the structure and procedure of employing organisations and trade unions. questions focused on responsibilities at workplace including: no. of members/stewards for whom responsible, extent and type of contact, no. and purpose of meetings, degree of efficacy felt, type of issues discussed, satisfaction with officers from other unions. relations with management including: amount of contact, type of issues raised, areas of disagreement, opinion of management. also, existence of and satisfaction with written agreements, attitude to strikes and other forms of pressure, assessment of stewards' work, amount of consultation and agreement/disagreement over proposals for change. assessment of climate of industrial relations. background variables include official title, which union, length of service, total no. of members for whom responsible, age, sex. the role of full-time trade union officers in northern ireland n. robertson time period: march 1-october 19, 1973 population: all full-time trade-union officers in northern ireland cases: 81 archive: ssrc survey archive (essex) the purpose of the study was to collect data to produce a profile of full-time trade union officers in northern ireland, revealing their backgrounds, the nature of their work, the problems they encounter and their attitude towards certain contemporary issues in industrial relations. to assess hov; the role of the full-time officer might be deficient and how it might be more effective, and to explore certain problems in the operation of trade unions peculiar to northern ireland. attitudinal and behavioral questions include nature and method of v/ork. opinions on inter alia : union aims, officers' salaries, induction and training, facilities provided, place of shop stewards, optimal forms of branch organization, forms of collective bargaining, union communications, union staffing, industrial disputes, causes and remedies, relationship with employers and managers, impact of government policy. background 12 variables include age, place of birth, school education, further education and training, original occupation, type of union, method of appointment, size of const! tuency, terms and conditions of employment, work load, outside committments. union leaders in chile henry a. landsberger time preiod: 1962 cases: 231 archive: icpsr variables: 175 questions in the study explored the development of awareness, interest, and involvement in the union as well as objectives for the union and self as a union leader, and r's participation in other organizations. respondents were asked about relations between the firms and the unions and between the unions and federations. also included were items on union tactics, level of interest, and involvement of other union members and officials. the study sought the respondents' attitudes toward the chilean labor movement, opinions as to what steps the country should take to continue social and economic progress and of the roles workers and industries should take to further national economic progress. several items probed perception of the personality of most chilean workers. personal data were also gathered including the effect the leadership role has had on respondents' personal lives and cynicism about other people. questions were asked about past family involvement in the union, respondents' career plans, and expectations for self and children. standard demographic information included age, marital status, education, parent's financial status, education, and regional background. archive addresses danish data archive odense university niels bohrs al le 2 5 inter-university consortium for political and social research p.o. box ]2h8 ann arbor, michigan ^islos social science data archive social science library yale un i vers i ty box 1958 yale station new haven, connecticut 06520 social science data library/ louis harris data center university of north carolina manning hal 1 026a chapel hill, north carolina 2751^* srl data archive survey research laboratory 1005 w. nevada street university of illinois urbana, i 1 1 inois 6i8oi ssrc survey archive university of essex v/ivenhoe park, colchester essex, england steinmetzarchief herengracht ^10-^12 1017 bx amsterdam netherl ands 13 18 iassist quarterly 2013 iassist quarterly abstract this is written in appreciation of the pioneering contribution made by sue dodd to what we would now call metadata standards for research data files. it describes two occasions when i had good cause to cite her work, the first when writing in 1984/5 about data libraries and how these might develop in the uk. the context is the early years of edinburgh university data library and the visit by sue dodd to present at a seminar and workshop in london and edinburgh. the second occasion for citation was almost 30 years later, when writing about digital preservation of scholarly statement. that gives opportunity to place her work in the context of the new forms of scholarly publication in which research data form an increasing part, with new need to ensure appropriate citation for webbased resources. keywords: cataloguing, metadata, seriality, web, registries, history introduction i have this sense of having met sue dodd for the first time on three separate occasions: through her writing in the iassist quarterly (iq); when we spoke on the telephone; and finally when we met in person at the start of her visit to the uk in 1985. i recall those moments with a smile. her writings, voice and warm sense of person have continued in my thoughts, her mix of charm, insight, dogged determination and encouragement. we all have access to her writing and those ideas and insights live on in our practice. we surely all have mixed thoughts when we realise that some variant of the following abstract could have been written yesterday: in the last two decades … agencies … have invested heavily in the collection of … data, contributing to the proliferation of … data. however, … the ability to produce data [has] progressed much more rapidly than our capacity to organize, classify, and reference its availability. … the purpose of this article is twofold: (1) to outline some of the information components associated with … data files, and (2) to provide guidelines, examples, and a uniform vocabulary for the creation of a bibliographic reference. (dodd, 1979) there is little doubt at the prescience of the advice that “information stored in a computer-readable form will soon become a legitimate library resource available to those patrons who need it” (op cit). however, even with the arrival of the web and the passage of time, research data is only now top of the agenda for libraries, and seemingly with a supply-side perspective, rather than having focus on the demand-side for the data needed for secondary analysis. i first cited sue’s work in 1985; i found the need to do so again when writing an article for serials review almost 30 years later. the interest in making those two citations serve as temporal bookends for the two parts of this appreciation, labelled parts a & b: part a has its focus on the first article, “towards the development of data libraries in the uk” (burnhill, 1985). not surprisingly, when i began writing about data libraries i gave emphasis to the importance of cataloguing data – the term metadata then had other meaning – and i would cite the work of sue dodd. part b has its focus on the other, “tales from the keepers registry: serial issues about archiving & the web” (burnhill, 2013), issued almost 30 years later when writing about digital preservation of scholarly statement. a legacy of inspiration and an enduring smile by peter burnhill1 iassist quarterly 2013 19 iassist quarterly i want to use this as opportunity to say something about the early years of edinburgh university data library which has now been operating for some 30 years. i also wish to say something of the new forms of scholarly publication in which data form an increasing part. perhaps what is persistent is the concern to ensure that researchers, students and their teachers can have access, both ease and continuity of access, to the resources that they need for their scholarship. my first encounter with sue dodd the very first time i met sue was through her writing. it was 1984 and i had just been appointed to develop the data library at the university of edinburgh. i had landed a very good job at a young age to lead a small team of two and a half full time equivalent staff, to take charge of the data library and advised that i would need to win external funding for its development. i was reading the iq collection that my predecessors had been collecting in order that i might understand the varied institutional settings in which data libraries were set. i began at the beginning, with volume 1 issue no. 1 of what was then called the iassist newsletter (november 1976).2 what stood out was the importance of standards for cataloguing datasets and the key role being played by sue who was listed as the us chairperson of the classification action group. the report of activity stated that the action group in the us gave emphasis “on the library cataloguing of machine-readable data files in public multi-media catalogues,” and noted: sue dodd has used the rules recommended by the american library association’s subcommittee on the cataloguing of machine-readable data files to prepare a draft version of a working manual for cataloguing machine-readable data files which will be tested by members of the us action group. the other actions noted were a committee to investigate a national union catalogue of catalogued mrdf, use of marc and a critical review of controlled vocabularies – the latter to interact with the european members of the classification action group led by the data archives in europe which had their focus on study descriptions. my background in my new role as ‘principal consultant (data)’ was that of a statistician and social scientist but i would go on to work with a number of forward thinking individuals in internationally well-regarded computing service organizations in edinburgh. the largest of these computing organizations was edinburgh regional computing centre (ercc) which operated the network and the mainframes for universities of glasgow, strathclyde and many a research institute across scotland as well as the large research and teaching base of the university of edinburgh. the university’s computer science department and the ercc had pioneered the development of multi-access computing, supporting a system known as emas that allowed its users to make use of commands in the english language (not ibm jcl) and to program within this operating system, including use of a form of hypertext in a system called view. this enabled us to escape much of the tyranny of magnetic tapes being experienced elsewhere. file transfer and remote log-on to computers hosted in national and regional computing centres were becoming routine for the initiated, as was email (and i still retain access to folders of email from that time). in the uk, sercnet was being re-launched as janet as the internet backbone for uk research computing. responsibility for application software was with another group at edinburgh, the program library unit (plu). this had been set up in 1969 with a national (and international) role for ‘knowledge based software facilities (or data)’ also converting and distributing ibm mainframe source code software to run under the operating systems used for the british manufactured icl hardware. the founding director of plu, marjorie barritt was clearly the farsighted-genius, with commitment to ‘data handling software’. just prior to my joining, plu had merged with the ercc database group to form a software house called the centre application software technology (cast). cast was a relatively shortlived organisation merging into ercc in 1989 to become the computing service, but for those five years cast provided the data library with a loving nursery. there had already been positive activity to establish the operation of a university data library by trevor jones, a lecturer in sociology, and by audrey stacey who was the computing expert (jones and stacey, 1984) with policy support from deputy librarian peter freshwater. researchers had petitioned for centrallymanaged university wide provision of access to large-scale datasets, typically the decennial population censuses for scotland, the annual agricultural censuses for england & wales and for scotland, the general household surveys and a range of digitized boundaries being used in what were still path-breaking ways to do computerized mapping. trevor left to work for caci in the emerging and lucrative geo-demographic industry, creating the vacancy that i had applied to fill.3 part a (1985). towards the development of data libraries in the uk a visit by geoffrey hamilton from the british library to peter freshwater, the deputy librarian at the university, led to an invitation to present a paper by a member of the uk committee of librarians and statisticians. this was a joint standing consultation body of the library association and royal statistical society that was responsible for publishing a series on statistical sources, such as ‘a union list of statistical serials in british libraries’ 4. geoffrey hamilton was leading an initiative on indexing the statistical tables published in government documents and he was intrigued at the discovery of activity to catalogue the datasets behind those tables. i set about re-reading those early issues of the iq in order to research the topic. the resultant paper, entitled “towards the development of data libraries in the uk” (burnhill, 1985), was duly presented to the committee. the opening page begins with a quote from sue dodd when offering a definition of ‘data’ to complement a media-based definition of ‘library’: data has been described as “a general term used to denote any or all facts, numbers, letters and symbols which refer to or describe an object, idea, condition, situation or other factor” (s. dodd 1982). clearly this is quite wide and describes much that anyone would want to analyze. the word library is derived from the latin word liber, originally the rind between the wood and the bark, the medium on which the information was recorded before the invention of paper. at one time the reader of a book had to know how to treat that particular medium, but after a 20 iassist quarterly 2013 iassist quarterly while all that was needed were literacy and the right to use a library. access software and analysis software now free the researcher from having to worry too much about the physical characteristics of machine-readable data held in a data library. re-reading that now, i would take issue with what was said, by sue and by myself. however, perhaps that planted the seed for the view i took later to separate ‘data’ from the ‘digital’, regarding the former as only being so if it (they?) could be regarded as having evidential value for some enquiry, and the latter prompting the question ‘what is different about the digital?’ with focus on the malleability of the medium. i made another reference to the work of sue dodd on page 9 in the section on ‘documentation’ and then again when discussing the value of the abstract, before placing her words centre stage when discussing cataloguing of machine-readable data files. this was an opportunity to combine my new found ‘cataloguing’ knowledge with some of the practices i had learnt from my time working as a survey statistician and researcher with the scottish education data archive. the stated purpose for my report to the uk committee of librarians and statisticians was to highlight the existence of the data behind those statistical tables in government publications, and of the value of what i termed ‘an online metadatabase’. i also wanted to think aloud and see what was wanted of a ‘data library’ from the different perspectives of a data analyst and of a data producer. in this paper i look at data libraries from each of two directions: from the point of view of those who want to use the data, and from the point of view of those who generate the data; that is, from the point of view of data analysts and data producers. the paper also includes a rough historical sketch of the development of data libraries in the academic (mostly social scientific) sector; a discussion of the importance of bibliographic control and the provision of an on-line meta-database. (‘data about data’), and highlights the trend towards access to the data that produce statistical tables. although not formally published that article is now, belatedly, in the university’s institutional repository – scanned from a printed copy – and reportedly still being downloaded every month (burnhill, 1985). in what now looks like a ‘use case workflow’, i wrote: when using a data library the data analyst may be motivated either by the need to provide information for managers and decision makers, or by the wish to contribute towards some longer term research enterprise. either way, the data analyst asks something like the following series of questions: 1 would the problem in hand benefit from empirical evidence? 2 are there data available which could shed light on this problem? 3 where is the database located? 4 how may i negotiate access? • permissions; mode of access; payment or funding implications 5 what is the provenance, status and quality of the data? • questionnaire; target population; sampling scheme; non-response 6 can i obtain codebooks and allied documentation? 7 how may i re-cast my problems so that these data can contribute? 8 what software is available for data retrieval, manipulation, analysis and presentation? 9 could i use this software myself? 10 how may i obtain hard copy of the results from the analysis? 11 what would be the cost in time and money? regrettably, i look back on that paper as something of a ‘failed manifesto’ as the development of data libraries in the uk was much delayed – even now they exist in very few universities. however, the paper was influential at the time as evidence in the joint enquiry by the esrc (uk) and nsf (us), alongside a contribution from alice robbin, a past iassist president (1979-82) and then director of the data and program library service, university of wisconsin-madison. the esrc leadership was provided by howard newby, previously a director of the data archive at essex who would go on to be chairman and chief executive of the economic and social research council (esrc), and then ceo of the higher education funding council for england (hefce). my second encounter with sue dodd the second time i first met sue dodd was when i spoke to her in person on the ‘phone. i had come to the conclusion that there was insufficient knowledge in the uk ‘anglo’ part of aacr2 about the new chapter 9. i decided that i should try to persuade sue to come to visit the uk and that the best way to achieve that was to reach out to her by tracking down her number at chapel hill, north carolina, which i then dialled. the voice at the end was slightly taken aback, as transatlantic calls were far from usual, for either of us. i established that she was interested in participating in the two seminars i then proposed, one to be held at the university in edinburgh and one in london under the auspices of rss/la committee of librarians and statisticians. sue was not at the iassist conference in amsterdam, may 1985, the first i attended. however, i did meet a number of the other names i had come across in those issues of the iq. i also began to see some differences and divisions in the european approach being taken, with the practice of the national data archives in europe, and that adopted in the us/canada in which there were many university-based data libraries. my third encounter with sue dodd the third time i first met sue was the delight of meeting her in person when she did indeed accept our invitation to travel to the uk. i recall that she noticed the jet lag but was determined to be positive and helpful. we took the opportunity to enjoy a travelling exhibition of the terracotta warriors that was visiting edinburgh. i learnt later of her graduate studies about china.5 advertisements for the two meetings had been distributed over the summer of 1985, including this one: seminar on bibliographic control of statistical data files as the number of machine-readable statistical data files increases it is becoming ever more difficult for data users to find out about all the data which may be relevant to their work. the need for a comprehensive register, or national bibliography, of data files is becoming apparent. how could this be prepared? could it be compatible with bibliographies and library catalogues of printed material? how might it relate to output from the european access project with which the esrc data iassist quarterly 2013 21 iassist quarterly archive is involved? what is the role of data libraries in making data accessible to the user community?” in order to provide an opportunity for discussion of these and related questions, the committee of librarians and statisticians is organising a seminar at the city university, london on monday 23 september 1985. the principal speaker will be sue dodd, a data librarian at the university of north carolina, whose pioneering work in developing standards for cataloguing machine readable data files has earned her an international reputation. other speakers include marcia taylor and bridget winstanley (esrc data archive), peter burnhill (university of edinburgh data library services) and geoffrey hamilton (british library). . . . while she is in the united kingdom, sue dodd will also lead a workshop on “computer-based catalogues for describing computer files and their documentation” on friday 20 september 1985 at the university of edinburgh, 18 buccleuch place, edinburgh. the title for the edinburgh workshop centred on what i still think is still moot, namely whether ‘data file and documentation’ necessarily and collectively constitute a multi-part object – indeed, whether there is any simple object where data files are concerned. unlike many data libraries in north america all data files at edinburgh were online and spinning on disc, not stored physically on tapes held in labelled tape racks. moreover there was an online ‘catalogue’ of what was held in the data library. just before sue visited, alison bayley had joined the data library as a part-time programmer.6 alison was developing the online information service ‘datalib’ enabling users to navigate a form of hypertext in ‘eview’ (called simply view in emas) to find information on services, facilities, filenames, access restrictions, etc. this had many descriptive fields of our own making. i recall that sue’s visit prompted an attempt to create a catalogue record for the small area statistics from the 1971 population census for scotland in the university library’s (opac) catalogue. this led to interesting discussion with peter berwick, the library’s head cataloguer, when it was suggested that we change the title in order to improve the way in which the item would be filed. there was nothing of a title found in the ‘item in hand’: what had been received had no header file with a descriptive title. we were introduced to ‘toward integration of catalog records on social science machine-readable data files into existing bibliographic utilities: a commentary’ (dodd, 1982a). the seminar at city university was interesting, attracting a wide variety from the library world as well as the data archive at essex. i recall that sarah tyacke was there, then deputy map librarian at the british library. the next year she became director of special collections in the library and subsequently keeper of public records and chief executive of the national archives where she oversaw the development of new strategies for dealing with the preservation of born-digital records. through her visit contact was made with ray templeton of the library association who had been working on standards for cataloguing the recent phenomena of software for microcomputers, (templeton and witten, 1984). ray and i were later to share the task of editing a guide that resulted from the esrc computer files cataloguing group (burnhill and templeton, 1989) which drew much from sue’s cataloging machine-readable data files: an interpretive manual (dodd, 1982b). the knowledge derived from sue’s work had practical application as the data library participated in the esrc regional research laboratory (rrl) initiative as part of rrl scotland (burnhill, carruthers and messer, 1988; burnhill and ewington, 1992). the ‘rrl initiative’ provided an opportunity to engage with the developing field of geographic information systems. particularly significant was a symposium sponsored by the uk association for geographic information on ‘metadata in the geosciences’ in 1990. this brought together several disciplines having interest in ‘metadata’ and its relation to ‘cataloguing information’, especially as this might relate to spatially-referenced data.7 metadata was characterised within the database community as the data dictionary that gave formal definition for the objects in the database. there was the beginning of understanding that additional metadata were required to support resource discovery. the term ‘actionable metadata’ was used to go beyond that needed for data discovery to include information that could be read and acted upon by software, not only metadata to identify relevant data for a user but also to retrieve the relevant data from a (remote) database and produce a predefined product such as a map or table (burnhill, 1991; medyckyj-scott et al 1995). there was attempt to juxtapose these new metadata requirements with the cataloguing fields from aacr2 chapter 9 that had their focus on ‘identification and availability’, ‘subject and content’, ‘characteristics of the media’ and ‘access and management’ (burnhill, 1991). returning from the 1995 iassist conference, hosted in québec, canada, i learnt that the university of edinburgh had decided to respond to a national (uk) call for a third national datacentre (at that time there were bids, at bath, and midas, at manchester) and wished to put forward the data library as the basis of that bid. three weeks later the bid went in. two months later we learnt that the university was successful, and we were given five months to be up and running and delivering online services. we launched edina, the poetic name for edinburgh, on 25 january 1996, on burns night, starting with biosis previews, a bibliographic database. that event might signal the date when my energies finally shifted away from the sharp focus on the social science data file. the prior contact with database experts and working with geospatial and mapping data had already prompted the beginnings of that shift. i have come to remember the strap-line for the 1990 iassist conference as “words, numbers, pictures, sounds: all will be digital and accessed from afar.” in fact, although i recall proposing the strap-line in the programme committee, it was actually “numbers, pictures, words, and sounds: priorities for the 1990’s.” during the early 1990s, my management responsibilities broadened, to be required to deliver computing support given to the library: staff in the data library began to learn more about text, and to carry out project work that led us to launch salser8, ‘probably the first webbased national union catalogue of serials’. edina continues today with a very broad range of services, <http://edina.ac.uk>, and with the mission to develop and deliver online services as part of the ‘jisc family’9 in order to enhance research and education in the uk, and beyond. the best way to appreciate the present spread of activity is to download the ‘community report’; perhaps the best way to appreciate the 22 iassist quarterly 2013 iassist quarterlyiassist quarterly variety of activity over the years is to dip into the online archive of past issues of ‘‘edina newsline’. the data library continues to flourish and have purpose: it has its data catalogue10 as well as a set of services geared at benefiting researchers, students and their teachers at the university of edinburgh.11 my colleagues in the data library, which together with edina form part of information services at the university, also contribute nationally and internationally. examples include significant contribution to the university’s focus on research data management (rice et al, 2013) and mantra12 , an online course designed for researchers or others planning to manage digital data as part of the research process. that includes a module on metadata and documentation in which three broad categories of metadata are described as part of training for future researchers: • descriptive common fields such as title, author, abstract, • keywords • administrative preservation, rights management & technical metadata • structural how components of a set of associated data relate to one another, such as a schema describing relations between tables in a database. active participation in iassist continues, including recent secondment of stuart macdonald to cornell university and the temporary addition at edinburgh of laine ruus, one of the famous names i read about in those early editions of the iq alongside sue dodd, and whom i also cited in that first article, “towards the development of data libraries in the uk” (burnhill, 1985): what is needed is a union catalogue of all known disseminators of mrdf, and some efficient means to access information on what new data files are being created. the movement by icpsr and the roper centre towards on-line remote access to their inventories is a major step towards information retrieval.” (l. g. m. ruus, 1980). part b (2013). tales from the keepers registry: serial issues about archiving & the web fast forward some thirty years and i look back to when there was again need to cite the work of sue dodd. during those thirty years the digital medium was no longer confined to those machine-readable data files that sue had focused upon: the digital medium had become the norm for scholarly statement, as with much in everyday life. the privileged access to the internet had given way to mass engagement with the web as an arena of interaction. invitation to contribute an article for serials review had prompted me to renew my acquaintance with the writings of sue dodd. i was writing about the arrangements being made in order that we might know what e-journals were being kept safe and what remained at risk. i had been asked to report on progress being made to ensure continuity of access to scholarly literature given the shift from print to digital format for all types of continuing resources, particularly journals, and the need to archive not just serials but also ongoing ‘integrating resources’ such as databases and web sites. my principal reason for citing dodd (1982a, 1982b) was to place her work within the history of aacr2, in part also to alert today’s librarians to the work of social science data librarians now that research data from all disciplines was being listed high on their agenda. i would like to use this occasion to alert social science data librarians to some ideas being taken forward now that scholarly content is issued as online resources, either issued in parts or changing over time. the article (burnhill, 1985) contains three stories which centre on the keepers registry which monitors the extent of e-journal archiving. the first tale: the keepers registry the first story in the “tales from the keepers registry” describes the problem of e-journal preservation, as noted in a number of reports over the past 10 to 15 years and the emergence of organizations willing to act as ‘digital shelves’. it also described the role of keepers registry as a global monitor on who is looking after what (how and with what terms of access). the registry has enabled the generation of statistics that indicate the extent of archiving for e-journals is cause for concern. today researchers in the social sciences – as in all disciplines ranging from physics to philosophy rejoice in the good news that scholarly statement is made available in ways that can be accessed any-time, any-place, and increasingly by any person and for any purpose. that advance had been greatly assisted by the emergence of the web, the principle arena for interaction across the internet. authors can make their content available very readily, via publishers or directly (with or without explicit licence). consumers of that content can shorten the time and effort required to discover, locate, request and access what they require (according to the licence). that is true for the produce of scholarship and for the resources that scholarship requires. the bad news is that so much of this scholarly content is not in the custody of research libraries. academic and research libraries continue to play a part but their role as intermediaries has been challenged, not least in their role as stewards of scholarly content that exists in digital form. libraries depend upon e-connections; they do not have their own e-collections. the shift to journal content that is digital, online and held remotely has challenged the essential responsibility that libraries have in figure 1 keepers registry iassist quarterly 2013 23 iassist quarterlyiassist quarterly ensuring continuity of access to scholarly content for their patrons. following reports from studies and projects around the world, a small number of organizations stepped forward to act as long-term archives for e-journal content. those reports noted the potential value of a resource that could address ‘who was looking after what, how, and what are the terms of access?’ the study commissioned by the jisc in 2007 (sparks, look, muir and bide, 2008) confirmed the feasibility and the perceived need for an e-journal preservation registry, indicating that such a registry could be built around the serials union catalogue (suncat), the national union catalogue in the uk developed at edina (burnhill, halliday, rozenfeld & kidd, 2004). the keepers registry has now emerged as a global online facility13, designed and built by edina at the university of edinburgh in collaboration with the issn international centre in paris. the basics of the design are illustrated below, taken from burnhill et al (2009), and show how the identifier for serials, the international standard serial number (issn) and the issn register is at the heart of the facility, against which the leading archiving organizations report on which serial titles (having issn) each is looking after: reporting metadata on how, to what extent and with what terms of access. the real heroes in this first tale are those digital preservation agencies, the ten archiving organizations that are contributing to the registry. as shown in the graphic above, the two main web-scale organizations of clockss and portico were in from the start, as was the global lockss alliance. the library of congress and hathitrust are among those to have joined since alongside the archeological data service (uk), having a discipline-based archiving responsibility. considering the complexities of research data, it might be supposed that the preservation of e-journal content was easy and that the problem was solved. unfortunately, that does not seem to be the case, as revealed by analyzing the archiving metadata that is aggregated in the keepers registry, as reported on the blog for the keepers registry.14 currently only about 22,000 e-serial titles of the 113,092 issn assigned to ‘online serials in issn register are reported as being ‘kept safe’ by the archiving organisations reporting into the registry and there are many ‘missing volumes and issues’. in 2013, the simple coverage statistic is 19%, an increase from 2011 when it was 17% (being 16,558 / 97,563), and what is interesting is that the numerator and denominator are both increasing, as archiving organisations ingest more titles and as issn is assigned to an increasing number of ‘points of issue’, about which more later. even if one narrows the focus to those serials that are considered important to libraries, the lists provided by cornell, columbia and duke universities, about 75% of e-serials (having issn) should be regarded as ‘at risk’ the second tale: metadata matters the second story in the “tales from the keepers registry” is about the variety of metadata issues that had to be addressed during the peprs project, including a number that remain unresolved (burnhill et al, 2009). typically serials are ‘well-published’ with rich metadata made available to archiving organizations by publishers. however, there are challenges relating to identifiers; variants in publisher information (naming and identification, and reference to issuing bodies) and variability about ‘holdings’ information relating to issues, volumes, and other buckets of digital stuff. the role of the issn has been key, the international standard identifier for a stream of content. the serial provides an entity which is ‘economic’ from an information management point of view, with discrete objects (typically as articles) made available in parts (typically as issues and volumes). nevertheless, the article (file) remains the ‘object of desire’, being accorded its own identifier, the digital object identifier (doi). attention is also given to the search for the universal holdings format in order to enumerate the extent of issued content, and thereby to check what is held and what may be missing. the importance of another identifier is becoming plain. until recently there was no universally accepted identification scheme for publishers, with name variants seen as part of the more general quest for authority files for personal and corporate names. a variety of name expressions is perhaps always to be expected, not just because of language differences. however, there is now a prospective solution with the emergence of the international standard name identifier (isni), an iso (international organization for standardization) standard (iso 27729), whose scope is identification for public identities15. the purpose of isni is to assist disambiguation of the public identities involved throughout the creation, production, management, and content distribution chain. that includes both organizations and persons (whether living or dead): there is a special allocation of isni numbers made available for assignment as orcid.16 the first two tales were presented in serials review in ways that were intended to engage serial librarians. the applicability of all of this for data librarians may not be self-evident but i would like to argue that there is much to be gained by considering how those matters might extend beyond such ‘well-published’ material as journals, especially to those social science data files that are generated from periodic enquiry and process. the emphasis is on identification rather than a full ‘bibliographic record’, and on the simplicity in a ‘data registry’ of knowing ‘who is doing what’. the central idea for a registry such as an e-journal preservation registry, the model for which might be generalized and adapted for other purposes, is part of a four-point proposition: 1. assign an identifier at the ‘point of issue’ for a stream of digital content 2. ensure that (digital) content is archived routinely, and that arrangement is made to have others/peers do that for you too 3. tell someone what you are doing and what you hold (and how) 4. publish the terms of access for the archived content (now and when triggered as orphaned). the third tale: where data and journal content collide the third story in the “tales from the keepers registry” was also written with the serials librarian in mind. the intention was to look beyond the conventional journal to the new research objects that have now become to be recognized and to the implications of the dynamics of the web. there is focus on the implications for citation, for notions of fixity, and for broader matters of digital preservation. i wished to highlight for the serials librarian some of the consequences for scholarly statement now that the web was becoming a principal arena for scholarly communication. not merely a dominant means to access, the web also enables rich aggregations of linked content into what have been termed ‘research objects’ having two classes: archived objects and publication objects that “are intended as a record of activity, and 24 iassist quarterly 2013 iassist quarterly should thus be immutable” and citable (bechhofer, de roure, gamble, goble and buchan, 2010). this can be seen to have built upon an attempt “to distill some core characteristics of a future scholarly communication system” (van de sompel, payette, erickson, lagoze & warner, 2004) with both registration (and ultimately preservation) of a scholarly asset being central to its success within a workflow or pathway through various service hubs. what are data librarians to make of these new scholarly objects that are growing in significance as part of the new information infrastructure for scholarship enabled by the web? thirty years ago it was important for so very many reasons to highlight the special case of social science data files as resources for scholarship and to contrast these with the apparent simplicity and fixity of what appeared as scholarly statement, as articles in journals and books on shelves. in the interim, scholarly statement has become digital and therefore malleable, with the characterization made above, it is now also extended to include data as intrinsic to that statement. in that telling of the third tale i wanted to point out to serial librarians that the shift to a broader view of scholarly works in digital format should not necessarily be regarded as completely new and alien, noting that sue dodd had made important observation thirty years ago in the pre-web era of the internet. and we are reminded that she wrote that “there is no doubt that machine-readable data will play an even greater role in research and development programs of the future. more and more data needed for government and private research will appear in computerized form.” (dodd, 1982a, p352); “in the near future, libraries will have no choice but to become more involved with computerized files and programs.” (op cit, p355). she was of course writing in the context of the publication in 1978 of aacr2 chapter 9 on ‘machine-readable data files,’ renamed ‘computer files’ in the revision published in 1988. on the other hand, this third tale could be interpreted and re-stated as a story that reflects upon the value of the concept of ‘seriality’ for data librarians and archivists. i have become convinced that this is a key concept for the structure of metadata for much that is issued on the web and indeed for much of what we were and still are interested in for ‘secondary data analysis of machine-readable data files’. complete revision of chapter 9 saw it become ‘electronic resources’ in the 2001 amendments that were confirmed in aacr2 2002, which also saw chapter 12 on ‘serials’ renamed ‘continuing resources,’ driven by a wish to harmonize across aacr2 and other serials bodies, including issn. the motive was common belief in the usefulness of the concept of seriality for what was, following widespread adoption of the web, being recognized as important points of issuance of content. the term ‘integrating resources’ was used to signify what was updated over time (differing from serials that are issued in separate discrete parts). the manifesto noted above and described by bechhofer et al (2010) is also reminiscent of work by hunter and choudhury (2006) and hunter (2006) that focus, respectively, upon the preservation of composite digital objects using semantic web services and the use of scientific publication packages (spps) for linking the raw data, their associated contextual and metadata on provenance, as part of publishing and dissemination of scientific results and selective preservation of scientific data. there is determined focus upon a “unit of scholarly communication” that is not “journals and their contained articles.” this evokes what are referred to as compound units, “aggregations of distinct information units that, when combined, form a logical whole” and can be represented in a manner (oai-ore17) that enables them to be accessed and processed by machines and agents (van de sompel & lagoze, 2007). seriality of issuance as such is not utilized in the argument put forward by bechhofer et al (2010). however, now that the web is recognized as an important point of issuance of scholarly content, both of scholarly product and of resource for scholarship, there is need for identification and ‘minimally-sufficient’ description of that stream, recognizing that some content is issued in separate discrete parts, and some changes (or is retrospectively updated/ modified) over time. what is particularly interesting about the article on research objects cited above was how it was made available; it was issued as a reviewed conference paper in nature precedings. at first sight, nature precedings resembles a journal, but it is not. launched in 2007 and closed in 2012, it acted as an open access preprint repository for the life science community. it was an integrating resource and as such assigned an issn, 1756–0357. the issn assignment policy now is being extended to online repositories as first point of issue for an increasing number of scholarly works. it may yet extend to repositories, such as figshare18, that exist to make research data and other forms of research output publically available. one wonders whether that issn assignment policy should and could extend to social science data archives. this third tale mentioned a project being carried out jointly by the research library at los alamos national laboratory and edina and the language technology group at the university of edinburgh.19 that investigation into what is termed ‘reference rot’ is now underway (sanderson, van de sompel, burnhill and grover, 2013). reference rot describes when content referenced at the end of the link has evolved, has changed dramatically, or has disappeared completely; it is more than ‘link rot’. an engaging overview is given in a talk by van de sompel (2011) about the use of the memento tool to access prior versions of web resources available from web archives and content management systems by using their original uri and a constructed ‘date-time stamp’ for the desired version, a bit like ‘time travel for the web’. preliminary work examining the survival of web-based content cited in articles in two scholarly repositories noted that 28% of the resources referenced by the articles in an institutional repository had been lost, and 45% (66,096) of the urls (in arxiv) that were found to still exist had not been archived (sanderson, phillips and van de sompel, 2011). it may be fitting to end this appreciation on the topic of citation. the contrast with the fixity associated with earlier printed format for scholarly statement is obvious. that contrast with the past is less obvious for the dataset, despite the suggestion made by dodd (1982a) to “conceptualize a singular mrdf to be an ‘inert file’ … that conceptually becomes the ‘item in hand’ to be described”. that was clearly said with the librarian of the early 1980s in mind. however, today’s data librarians and data archivists might be reassured to note, that dodd (1982a) also drew attention to the “dynamic data base [as] one that is characterized by its fluid and constantly changing nature. it may be represented by economic time series, or bibliographic data bases, and may be corrected, revised retrospectively, updated, merged, partitioned, and blocked iassist quarterly 2013 25 iassist quarterly into subfiles without changing its bibliographic identity.” although this latter observation predates the arrival of the web it should underscore our recognition that the web is dynamic. what may have existed, as indicated by citation, at the moment of reference can and does change. once more we must pay renewed attention on how to cite the (web-based) data resources that are issued beyond the traditional journal literature. references bechhofer, sean, d. de roure, m. gamble, c. goble, & i. buchan. (2010, july 6). research objects: towards exchange and reuse of digital knowledge. nature precedings. doi:10.1038/npre.2010.4626.1 <http://precedings.nature.com/documents/4626/version/1> burnhill, peter. (1985) towards the development of data libraries in the uk, edinburgh: centre for application software and technology 1985. available from: <https://www.era.lib.ed.ac.uk/ handle/1842/2510> burnhill, peter, ann carruthers, and anne messer. (1988) bibliographic control of research data. edinburgh: regional research laboratory for scotland. burnhill, peter and templeton, ray. (eds.) (1989) cataloguing computer files in the uk: a practical guide to standards. (joint report to the computer files cataloguing group, economic and social research council.) colchester: esrc data archive, university of essex. burnhill, peter. (1991) metadata and cataloguing standards: one eye on the spatial, in metadata in the geosciences. ian newman, david medyckyj-scott, clive ruggles and david walker (eds). loughborough: group d. burnhill, peter, and ewington, heather. (1992) catalogue of digitized boundary files for england and wales held by edinburgh university data library. working paper 33. edinburgh: regional research laboratory for scotland. burnhill, peter, leah halliday, slavek rozenfeld, & tony kidd. (2004) suncat: a modern serials union catalogue for the uk. serials 17 (1), 61-67. available from: <http://uksg.metapress.com/content/ c6y0ltjtxlhrgfr9/> burnhill, peter, francoise pelle, pierre godefroy, fred guy, morag macgregor, christine rees and adam rusbridge. (2009) piloting an e-journals preservation registry service (peprs). serials 22 (1), 53-59. available from: <http://uksg.metapress.com/link. asp?id=350487p5670h0v61> burnhill, peter. (2013) tales from the keepers registry: serial issues about archiving & the web. serials review 39 (1), 3–20. available from: <http://www.sciencedirect.com/science/article/pii/ s0098791313000178> and <https://www.era.lib.ed.ac.uk/ handle/1842/6682> dodd, sue a. (1979) bibliographic references for numeric social science data files: suggested guidelines. journal of the american society for information science 77–82. dodd, sue a. (1982a) toward integration of catalog records on social science machine-readable data files into existing bibliographic utilities: a commentary. library trends 30 (3). <http://hdl.handle. net/2142/7215> dodd, sue a. (1982b) cataloging machine-readable data files: an interpretive manual. chicago: american library association. hunter, j. (2006). scientific publication packages – a selective approach to the communication and archival of scientific output. international journal of digital curation 1 (1) 33-52. doi:10.2218/ijdc. v1i1.4 hunter, j. & choudhury, s. (2006). panic – an integrated approach to the preservation of composite digital objects using semantic web services. international journal on digital libraries: special issue on complex digital objects. 6 (2), 174-183. <http://www.springerlink. com/content/p480171432571j25/?mud=mp> jones, trevor and stacey, audrey. (1984) data library service: introductory guide. edinburgh: centre for applications software and technology, edinburgh university. medyckyj-scott, d. j., c. monckton, p. burnhill. (1995) progress towards standards for spatial metadata, in rowley, j. (ed.), standards supplement to the agi ’94 conference, london: association for geographic information. rice, robin, cuna ekmekcioglu, jeff haywood, sarah jones, stuart lewis, stuart macdonald, and tony weir. (2013) implementing the research data management policy: university of edinburgh roadmap. the international journal of digital curation. 8 (2) 194-204 <http://dx.doi. org/10.2218/ijdc.v8i2.283> ruus, laine g.m., (1980), user services in a data library. iassist newsletter 4 (2), 29-33 sanderson, robert, herbert van de sompel, peter burnhill and claire grover. (2013) hiberlink: towards time travel for the scholarly web. proceedings of the 1st international workshop on digital preservation of research methods and artefacts (dprma ‘13). acm. <http://www. deepdyve.com/lp/association-for-computing-machinery/ hiberlink-towards-time-travel-for-the-scholarly-web-7bjrwfvcpa> sparks, sue, hugh look, mark bide, & adrienne muir. (2010) a registry of archived electronic journals. journal of librarianship and information science, 42(2), pp. 1-11. templeton, ray and witten, anita. (1984) study of cataloguing computer software: applying aacr2 to microcomputer programs. london: british library; distributed by publications section, british library lending division. van de sompel, herbert, s. payette, j. erickson, c. lagoze, and s. warner (2004). rethinking scholarly communication: building the system that scholars deserve, d-lib magazine. 10 (9). <http://www.dlib.org/ dlib/september04/vandesompel/09vandesompel.html> van de sompel, h., lagoze, c. (2007, august). interoperability for the discovery, use, and re-use of units of scholarly communication. ctwatch quarterly. 3 (3). <http://www.ctwatch.org/quarterly/ articles/2007/08/interoperability-for-the-discovery-use-and-re-useof-units-of-scholarly-communication/> 26 iassist quarterly 2013 iassist quarterly van de sompel, h., r. sanderson, m.l. nelson, l. balakireva, s. ainsworth, and h. shankar. (2009, november 6). memento: time travel for the web. arxiv preprint. <http://arxiv.org/abs/0911.1112> van de sompel, h. (2011, december 2). time travel for the scholarly web. talk given at stm innovations seminar 2011 ( <http://www. stm-assoc.org/events/stm-innovations-seminar-2011/>). <http:// river-valley.tv/time-travel-for-the-scholarly-web/> ) notes 1. peter burnhill is director, edina & data library, information services, university of edinburgh, causewayside house, 160 causewayside, edinburgh eh9 1pr. <p.burnhill@ed.ac.uk> 2. the first issue of what was to become the iq was based upon reports of an iassist meeting held as part of the international political science association in august 1976. it had been hosted in edinburgh which was also to host the iassist conference on two later occasions, in 1993 and 2005. 3. prior to that i had been working for almost five years as a survey statistician and researcher at the centre for educational sociology in the university’s social science faculty, funded by the scottish education department. with colleagues i was designing and conducting surveys of school leavers and helping with a collection of survey datasets known as the scottish education data archive doing a lot of what we now call data curation. 4. a union list of statistical serials in british libraries, committee of librarians and statisticians. london, library association, 1972. 5. typical of her generosity sue insisted in buying me a figurine of one of those chinese warriors that still has a pride of place on the mantelpiece at home. 6. twenty years later alison and i would do a joint presentation at iassist 2003, entitled ‘getting to know the score: using the first 20 years to plan the next’, found at <http://datalib.library.ualberta.ca/ conferences/2003/presentations/> 7. that was my first encounter with david medyckyj-scott who eventually joined edina and data library in 1995/6 in order to lead the development of digimap and of metadata for geo-spatial systems more generally. 8. salser is the union catalogue of serials holdings for scottish universities, the municipal research libraries of edinburgh and glasgow, numerous smaller scottish research libraries and the national library of scotland. it was launched in 1994 and is available at <http://edina.ac.uk/salser/description.html> 9. jisc (formerly the joint information systems committee, and still commonly referred to as jisc) is owned by the representative bodies of uk universities, colleges and skills organizations, <http:// www.jisc.ac.uk/> 10. <http://datalib.edina.ac.uk/> 11. <http://www.ed.ac.uk/schools-departments/information-services/ services/research-support/data-library> 12. <http://datalib.edina.ac.uk/mantra/> 13. <http://thekeepers.org>, 14. <http://thekeepers.blogs.edina.ac.uk/> 15. <http://www.isni. org/> 16. orcid (open researcher and contributor id) is an alphanumeric code to uniquely identify scientific and other academic authors. <http://orcid.org/about/>; <http://www.isni.org/content/ isni-other-identifiers> 17. open archives initiative object reuse and exchange (oai-ore) defines standards for the description and exchange of aggregations of web resources. sometimes called compound digital objects, these may combine distributed resources with multiple media types including text, images, data, and video. <http://www.openarchives. org/ore/> 18. <http://figshare.com/> 19. this project as funded by the andrew mellon foundation was called ‘time travel for the scholarly web’ (tt4sw); the hiberlink project website is at <http://hiberlink.org> iassist quarterly 2013 27 iassist quarterly keywords from vol 37 n0. 1-4. courtesy tagxedo.com i^$> st newsletter vol.4 no. 1 apprmches tc djformation aboltt cdeus bureau machine-reaeaele mta files ( 1 ) ferbara aldrich ifeta user services division eureau of the census u.s. dept. of camerce during the last several years, the canrlnity of census eureau data file tisers has expanded greatly and the nurber of files available has increased dranatically. this phenanenon will continue v*ien files are released fran the 1977 econanic censuses and the census of goverments and the 19s0 census of fbpulation and housing. as the nuriber of data files distributed by the census eureau increases, there is need for a systematic method of providing information about existing holdings of data files and updating that information as new files are released. vhlle current methods of amouicing new file releases reach a very large segnent of the user camrity, the need exists for a single reference pifclication containing information about the oanplete collection of files available fran the eureau. this is the need the eureau proposes to meet with a new publication now available. li this paper, the current means of data information dissemination will be discussed, followsd by a detailed explanation of directory of data files a new census bu-eau puslication prepared by the ebta user services division. currently, archivists and researchers obtain initied information about eureau files fran several sources. data files released during the year are described briefly in f^rt ii of the annual bureau of the census catalog . more current releases are annoinced in data user news (published monthly by the data user services division) and in non-eureau publications such as american dempgraphics and review of public data use . another less formal means of "spreading the word" dxxit new data files is the data archivist/researcher network. particularly active at major universities, this informal information network operates outside the constraints of printing and mailing delays. although each of these mechanisms for providing information about new files serves a particular audience, none are standardized regarding the level of data information provided or the format in viiich the information is presented. nor is there a method of maintaining this information in an easily consulted, curulative form. ibwever, this information need will be met by the directory of data files (2), prepared within the etta user sa-vices division of the census bureau. this directory is a emulative reference guide (l)prepared for the 1979 international association for social science information service and tednology (lassist) annual conference, may 7-10, 1979, ottawa. ontario. updated march. 1980. (2)the directory of data files was prepared by molly abranowitz and barbara aldrich uider the direction of larry carbaugh, chief. custcmer services e>-anch, teta user services division. -3newsletter vol. 4 no. 1 for all publicly avail^le machine-readable data files. it serves a two-fold purpose by indicating that a particular file is available and providing brief but concise information regarding its iniverse, sifcject-matter , geographic level, and size. the directory is a subscription itan available through the u.s. goverment printing office. subscribers to the basic publication will autonatically receive quarterly updates ccntaining information about new files. the directory is in a loose-leaf format so all update pages can be entered in the appropriate place. this publication does not replace any of the current systems of data information dissemination discussed earlier nor is it intended to replace the technical dccunentation as the source of detailed information about individual files. an abstract is written for each file in the directory . each abstract provides enou^ information so that the user cm determine if the file is of interest or if the more detailed technical dccunentation for the file is needed. the directory reflects the efforts of the eureau to provide authority titles for files and a standardized franework for information about than. this effort conplements the on-going develofment of standards for the eureau 's technical docunentation . the directtyy begins with a general introduction to census data and a review of file ordering procedures. the next ch^ter is divided into program areas: agriculture, eccncmics, general, geographic reference, goverment, population and housing, ar«j software. a general overview is provided for each. following this, individual censuses and surveys are outlined, followed by abstracts of the available files produced in these efforts. for exarple, the pqxilation and housing section contains a general outline of the elireau's ro''e in collection of population and housing data. this is followed by an overview of the census of fbpulation and tousing and ^stracts of the files that are products of that toisus. these abstracts are a major portion of the directory . information provided in the abstract includes the bibliographic entry for the file, type of file, descriptions of universe, subject-matter, and geographic coverage, file size, related reference materials, printed reports end machine-readable data, and file availability. the btbl .tografhic entry for the file ocntains the authority title, the statement of responsibility, place of production, producer and distributor, and the date of production. examfle: current fbpulation survey, cdtober 1968-1978 / conducted by the bureau of the census for the bureau of labor statistics. washington : the bureau, 1978. (all subsequent exanples are based on this file.) the biblicgraphic entry is followed by a statanent of type cf file and the unit cf cbservancn. within the frajiework of census bureau data, the options for type of file are microdata, simary statistics, geographic reference data and software. type cf file: microdata, uiit of observation is individuals within housing mits. the next element in the abstract is the universe descriftion. if the data are sarple data, a brief description of the saipling technique is also provided. ia)$)$)|$t newsletter vol. 4 no. 1 universe descriftion: the iniverse consists of persons 3 years old and over viio are the civilian non-institutional population of the u.s. living in housing units. a probability satiple is used in selecting housing units. a subject-matter descriftion follows. this is not intended to be a conplete list of the individual variables on the files, but rather to provide an overview of the type of data available on the file. subject-matter description: the basic infonnation collected monthly provides data on labor force activity the week prior to the survey. conprdtensive data are available on the erplojment status, occupation and industry of persons ^u years and over. characteristics such as age, sex, race, marital status, household relationship, educational backgrotrd and spanish origin are shown for each person tj and over in the hous^xjld survey. in addition, supplanental surveys are sonetijnes taken at the sane time the basic data are gathered. the current bopulation survey, october 1968-78, files contain slpplanentary data on school enrollment. items include current school enrollment, type of school in which anrolled and grade attending for all persons 3 and older. other school enrollment data varies betwsen supplanents. the current ft^jjlation slirvey, october 197^^ contains post-secondary school infonnation as well as the school enrollment data. the gbografhic coverage elanents provide the information ocnceming the levels of geographic identification used in the file. c3e0grafhic coverace: in the years 1968-72, geographic identification includes regions, divisions, 19 states (in the remainder several states are grouped together for identification) and 19 eisa's. in the years 1973-75 geographic identification includes regions, divisions, 13 states and s'^ s^ea's. begiming in 1977. geographic identification includes regions, divisions, all states and 44 smsa's. the file size elanent reports the nurber of files and logical record coirt of the file(s). in the case of a series like the cps, vhen the size varies ft-cm year to year, the approximate range is given. file she: 1 file per year, logical record couit ranging fran approximately 130,000 to 1'46,0o0. the elanent related reference materials is designed to provide the user with a guide to more information about the files. it references the technical docunentation as well as other materials regarding data-collection or sanpling. it also provides information on hew these materials can be obtained. reference materials: u.s. bureau of the census. the current population survey : design and methodology. (technical paper ijou.s. bureau of the census) superintendent of ebcunents no: c3,212:i«. goverment printing office, washington, d.c. 201402. price $3.25. newsletter vol. 4 no. 1 this canprehensive document gives detailed information on the cps program including sanple design and rotation, survey operations, and preparation and accuracy of estimates as well as sanpling errors. thirteen appendices are provided liiich contain very useful information to a cps user. "qjrrent population survey, october (year) tednical etocunentation" . available fy-on the clistaner services e^cinch. this is a guide to the nbchine-readd)le data file. it has general information about the data, specific content information and a codebook. a separate abstract element related printed reports cites pifclicatlons which contain actual data produced fran the files being abstracted. related printed reports: u.s. bureau of labor statistics. employment and earnings , noveufcer (year ). the anployment information in section a of this docunent is derived fvcm the current population survey of the previous month. e&ta f>-cm the current fbpulation survey, cbtober file is usually pitollshed in the cuxent population reports p-20 series. the publications for specific years can be fotrd in the bureau of the census catalog . exanples of sane of these reports are listed below: ctollege plans of high school seniors: cbtober 197'), p-20 no. 28^4. cbllege plans of high school seniors: october 1975, p-20 no. 299. major field of study of college students: october 197h, p-20 no. 289. school ehrollment — social and economic characteristics of students: october 1975, p-20 no. 303. hjrsery sdxol and kindergarten enrollment of children and labor force status of their mothers: october 1967 to october 1976. p-20 no. 318. school ehrollment — social and eboncmic characteristics of students: october 1976. p-20 no. 319. the last elsnent of the abstract is file avatlabtltty. this section indicates how the file is packaged and sold (by census division. state, all files as a single trit) and identifies individual file order nunbers for each file. appendix c of the directory is a chart arranged by file order ntmber vhlch indicates the nuitier of reels required providing several density and blocking options. appendix a provides a glossary of census geographic concepts and appendix b is a bibliography of census publications relating to madilne-readable data files. this pitlication, issued october 1979. is sold by the u.s. govertment printing office. the cost (vhich includes the quarterly updates) is $11. orders may be placed with sifcscriber services section (publications), bureau of the census, washington, d.c. 20233. -6 1/2 rasmussen, karsten boye (2020) editor’s notes: countries closing down reproducibility keeping science open, iassist quarterly 44(1-2), pp. 1-2. doi https://doi.org/10.29173/iq981 countries closing down reproducibility keeping science open karsten boye rasmussen welcome to volume 44 of the iassist quarterly. here in 2020 we start with a double issue on reproducibility (iq 44(1-2)). the start of 2020 was in the sign of corona. though we are now only in the middle of the year, we can say with confidence that 2020 will be known for the closing down of nearly all public life. from our very own world this included the move of the iassist 2020 conference to 2021. the closing down of societies took different forms and this will and should be long debated and investigated, because many civil rights in open society were put on instant standby by governments, with various precautionary measures. fortunately, many countries are now in the processes of opening up. hopefully, we are now more careful, keeping socially distant, executing better sanitation, etc. we are also eagerly expectant of science breakthroughs: the vaccine, the better treatment, the cure. but corona science extends beyond health and biology. social science in particular has an obligation to make us better prepared to take necessary measures and to uphold democracy. social science will always have the problem of reliability that you cannot step into the same river twice: survey data collected at one time will not in a subsequent data collection bring the same results, even with the same panel of respondents. reproducibility has many more forms than exact data collection, though, and is foundational for open science and an open society. science needs to be transparent in order to be challenged and improved. fellow scientists as well as laymen should have the possibility of performing analyses to find whether results can be reproduced. i am therefore very happy to send my thanks to harrison dekker and amy riegelman for taking the initiative to create this special issue of the iassist quarterly on reproducibility. harrison dekker is a data librarian at university of rhode island and amy riegelman a librarian in social sciences at the university of minnesota. together, amy and harrison reviewed the papers submitted for their special issue and wrote the introduction in the following pages. in addition to expressing my great appreciation to them, i also want to thank all the authors who submitted papers for this issue. thanks! let's keep science open again! submissions of papers for the iassist quarterly are always very welcome. we welcome input from iassist conferences or other conferences and workshops, from local presentations or papers especially written for the iq. when you are preparing such a presentation, give a thought to turning your one-time presentation into a lasting contribution. doing that after the event also gives you the opportunity of improving your work after feedback. we encourage you to login or create an author login to https://www.iassistquarterly.com (our open journal system application). we permit authors 'deep links' into the iq as well as deposition of the paper in your local repository. chairing a conference session with the purpose of aggregating and integrating papers for a special issue iq is also much appreciated as the information reaches many more people than the limited number of session participants and will be readily available on the iassist quarterly website at https://doi.org/10.29173/iq981 https://www.iassistquarterly.com/ 2/2 rasmussen, karsten boye (2020) editor’s notes: countries closing down reproducibility keeping science open, iassist quarterly 44(1-2), pp. 1-2. doi https://doi.org/10.29173/iq981 https://www.iassistquarterly.com. authors are very welcome to take a look at the instructions and layout: https://www.iassistquarterly.com/index.php/iassist/about/submissions authors can also contact me directly via e-mail: kbr@sam.sdu.dk. should you be interested in compiling a special issue for the iq as guest editor(s) i will also be delighted to hear from you. karsten boye rasmussen june 2020 https://doi.org/10.29173/iq981 https://www.iassistquarterly.com/ https://www.iassistquarterly.com/index.php/iassist/about/submissions mailto:kbr@sam.sdu.dk 1/2 rasmussen, karsten boye (2020) editor’s notes: knowing what to do and how to do it: high transparency and careful curation of data and metadata, iassist quarterly 44(3), pp. 1-2. doi https://doi.org/10.29173/iq984 editor's notes: knowing what to do and how to do it: high transparency and careful curation of data and metadata welcome to the third issue of volume 44 of the iassist quarterly (iq 44(3) 2020). transparency is a prerequisite for valid analysis of data. full disclosure of all aspects of the creation process is necessary for the evaluation of a data collection. the roper center has collaborated, assembled and developed standards, and performed scoring of datasets to facilitate the evaluation of data. it is easy to say that all aspects of data collection are important, but with more knowledge about the process of data curation phd students become aware of how it is important for their research. the cessda metadata office is a tool supporting the realization of high transparency and successful research data management. this is an extremely short overview of the three submissions found in this issue, below a little more information, and finally enjoy the articles. the first paper in this issue is first for a reason. 'standards and scoring to increase transparency for archived public opinion data' by kathleen j. weldon won first prize among the papers submitted for the iassist 2019 conference. you do remember iassist conferences, i hope. annual conferences have been regular as clockwork, but due to covid-19 the 2020 conference in göteborg, sweden, is postponed to april 7–9, 2021. see more at https://www.iassist2021.org/. kathleen j. weldon is the director of data operations and communications at the roper center for public opinion research, cornell university – the world’s largest archive of public opinion survey data. new standards for archiving data at roper center are described in the paper in the context of their historical development. what is now termed ‘transparency’ was an early commitment by george gallup to 'full disclosure of methods, sponsorship, and data'. after some years, standards were adopted by professional organizations. at the roper center, submissions have long been evaluated in order to preserve ‘the best of its time’ polling. however, with new technologies new survey methods were used without a consensus of best practices. this led to the development of standards by the national council on public polls (ncpp) and the american association for public opinion research (aapor). the new standards are the basis for roper center developing a system for scoring transparency. it is the author's hope and expectation that the future will bring an increase in the preservation of well-documented polling datasets. many phd students produce data collections or compile and use archival data for their thesis. at purdue university, a school of information studies course was developed for phd students covering topics such as identifying an archival data set, creating metadata and documentation, selecting an appropriate long-term storage location, and planning for the deposit process. the paper 'capturing their “first” dataset: a graduate course to walk phd students through the curation of their dissertation data' describes the background, the development, and the content and structure of the course, and concludes with assessments and the insights gained for the library. naturally some practical issues evolved through the pilot course; specifically, the broad spectrum of disciplines in phd projects demands much customization of the course curricula. the authors megan sapp nelson https://doi.org/10.29173/iq984 https://www.iassist2021.org/ 2/2 rasmussen, karsten boye (2020) editor’s notes: knowing what to do and how to do it: high transparency and careful curation of data and metadata, iassist quarterly 44(3), pp. 1-2. doi https://doi.org/10.29173/iq984 and ningning nicole kong are respectively professor and associate professor of library science at purdue university libraries. the course was established by request from phd advisors, who found that datasets were valuable outputs not being captured in the thesis deposit process. this course also meets the demand for data management skills in the job market. gesis leibniz institute for the social sciences in germany and the uk data archive are leading the cessda metadata office project. the metadata office covers several areas such as metadata models, schemas, profiles, controlled vocabularies, and thesauri. the paper 'the matter of meta in research data management: introducing the cessda metadata office project' reports on the project. cessda stands for consortium of european social science data archives, and with the many european languages it follows that language issues are the focus. the project builds upon ddi (data documentation initiative) and also includes collaboration with the ddi alliance on translations of ddi vocabularies through the use of iso language tags in metadata and controlled vocabularies; this is exemplified in some screen shots. the paper also presents the information covered in the cessda metadata model and some examples of the characteristics of elements of the model. the authors are all participants in the project; andré förster and kerrin borschewski as consecutive heads of the project at gesis, sharon bolton as head of the project at uk data service, and taina jääskeläinen at the finnish social science data archive. submissions of papers for the iassist quarterly are always very welcome. we welcome input from iassist conferences or other conferences and workshops, from local presentations or papers especially written for the iq. when you are preparing such a presentation, give a thought to turning your one-time presentation into a lasting contribution. doing that after the event also gives you the opportunity of improving your work after feedback. we encourage you to login or create an author login to https://www.iassistquarterly.com (our open journal system application). we permit authors 'deep links' into the iq as well as deposition of the paper in your local repository. chairing a conference session with the purpose of aggregating and integrating papers for a special issue iq is also much appreciated as the information reaches many more people than the limited number of session participants and will be readily available on the iassist quarterly website at https://www.iassistquarterly.com. authors are very welcome to take a look at the instructions and layout: https://www.iassistquarterly.com/index.php/iassist/about/submissions authors can also contact me directly via e-mail: kbr@sam.sdu.dk. should you be interested in compiling a special issue for the iq as guest editor(s) i will also be delighted to hear from you. karsten boye rasmussen september 2020 https://doi.org/10.29173/iq984 https://www.iassistquarterly.com/ https://www.iassistquarterly.com/ https://www.iassistquarterly.com/index.php/iassist/about/submissions mailto:kbr@sam.sdu.dk vol183&4 13fall/winter 1994 introduction jim jacob’s work on levels of service and levels of access provides an excellent starting point for exploring options for cooperative support of access to numeric files2. jim divides types and levels of service into four basic categories. in the first, general data services, he delineates the full range of services that an institution might offer in connection with machine-readable data. services are laid out in a hierarchical manner. the three remaining lists, library data services, reference data services, and computing services, outline services that fall into each of these more specialized categories. while he does not directly address cooperative support, he outlines quite clearly the possible levels of support and service. he also makes the equally important point that it is not necessary — and probably not desirable — for a library to attempt to provide “full service” for machinereadable information on its own. that being the case, building partnerships to provide enhanced levels of service makes a great deal of sense for many libraries. however before setting out to forge these partnerships, a library must take stock and determine exactly what levels of support it can provide in-house, where to draw the line, and what sort of partnerships it might logically seek. in the data archives community, there is no one model for the provision of data services—no right way to do it. each data library or archive seems to have its own unique structure, procedures, and services. therefore, traditional libraries entering into the data services arena will do well to review their goals and objectives and formulate a mission statement for data services. such a statement requires deliberate managerial decisions on what levels of staffing and funding are available to commit to data services. it must also determine what levels of service are desirable given its mission and supportable given the resources at hand. from there, formal collection development and public service policy statements are appropriate and useful tools for communicating these decisions to the library’s clientele. once the library delineates the role it can and will play, it can seek partnerships that will strengthen and complement its services. doing the latter without the former may result in difficulties when the objectives of cooperation are unclear and the division of responsibilities and authority between cooperating organizations is ambiguous. just as there is no one right way to deliver data services, the right way for any given library will be governed to a great extent by its larger institutional context. a library department seeking to define its role in providing data services must understand its place within the library and the role of other actual or potential data service providers within that library. if the library services (or is considering servicing) datasets in other reference units, it should consider the pros and cons of establishing additional decentralized data services. there may be economies of scale in consolidating services where subject and/or technical expertise is strongest. there may also be an established philosophy within the library that dictates one approach over another. the library must also understand and take into account its role in relationship to other organizations within the institution, as well as externally. for example, if a campus has other strong units with a history of established data services, the library will need to be aware of those units and their services when deciding its role. if the campus administration has funded other units to provide some types of data services (for example, gis systems), the library may want to develop arrangements for housing and/or servicing any geographic files it acquires in those units rather than duplicating services available elsewhere. once it has completed it deliberations and come to some decisions on the levels of service it will provide, a library may want to review its selection of machine readable items. for example, the library may decide not to support files without their own extraction software. in that light, (without agreements to support them elsewhere on campus), the library would want to ensure that it had not selected items like the current population survey or the american housing survey through the depository library program. conversely, if a library is not selecting or acquiring files that another unit is willing to support, perhaps they should be acquired. collaborative alliances may take one or more forms. they may be purely informational and informal: they options for cooperative support of access to numeric files by jean slemmons stratford1 institute of governmental affairs university of california, davis 14 iassist quarterly may be for the sharing of expertise, information, or solutions to problems; they may also be more formal and involve a division of labor or resources in support specific files or classes for files. anything other than the most informal of collaborations will benefit from a written agreement. such a document can clarify many aspects of the arrangement. it should include information on what the aims of cooperation are, how the collaboration will work, what the division of labor will be, what each party’s level of commitmentis in terms of resources, services, whether commitments are ongoing or for a set period of time, etc. with these points in mind, here is an overview of some of the many possible sources for strategic alliances to enhance support of data collections. the campus icpsr official representative icpsr, the inter-university consortium for political and social research, is a consortium of nearly 400 institutions worldwide. one of the primary functions of the consortium is to support a central repository and dissemination service for machinereadable social science data. for many members, the primary benefit of membership in icpsr is access to the consortium’s vast data collections. many of the data series held and distributed by icpsr will be familiar to librarians in their printed forms. the icpsr membership and data distribution is handled on each member campus by an “official representative.” currently, there is a trend toward housing the icpsr membership and data collection within a support unit such as the member institution’s library or computer center. however, historically icpsr ors have come from other areas as well. ors include among their ranks not only librarians and programmer/analysts but teaching faculty in a variety of social science disciplines and academic staff from research institutes and programs. there is a substantial body of expertise in the organization in use of social science machine-readable data. in an institution where the icpsr membership is handled outside the library, this would be an excellent first place to look for strategic alliances. however, the range of options in servicing library datafiles may be limited. two options come most readily to mind: 1) informal collaboration and sharing of expertise, 2) expanding access to data sources by including the icpsr collection in the library’s opac regardless of physical ownership and location of the collection. the latter has been done successfully at several institutions and a variety of approaches have been used. computing facilities as jim notes, users of machine-readable information must have access to computing services. jim divides these services into four basic categories: data storage services, copying and subsetting services, data retrieval services, and data analysis services. provision of even the most basic data storage services will require some access to appropriate hardware and software. given the dramatically short life of computer products, computing services is one area where cost may quickly outstrip a library’s resources. therefore, the library will benefit from a clear understanding of what levels of service it can support in-house and what other institutional resources are available to provide computing services. a library may acquire datafiles on a number of storage media from floppy diskette and cd-rom to various and sundry tape formats. the range and type of media acquired will determine whether the computing facilities needed to support even basic data storage services are minimal or more extensive. if a library limits its acquisitions to diskettes and cd-roms, the equipment required to verify and backup datasets will be manageable. however, equipment and software must still be available and kept up-to date to perform these simple procedures. for the other levels of services (copying, subsetting, retrieval and analysis), the equipment requirements escalate rapidly. additional hardware and a broader range of software are required for these latter services. as the hardware and software requirements increase, the human resources that must be devoted to servicing the files are also dramatically increased. clearly, this is an area where collaboration may be in order. on a typical campus, there are several places to look for partnerships. centralized computing facilities are a likely possibility and may have resources to commit to supporting access to the library’s datafiles. most such facilities are better placed than any library can hope to be for the simple reason that they have budgetary resources committed to maintaining and upgrading a volume of hardware and software. more formal collaboration with centralized facilities might include something as simple as providing end users with access to lab equipment and ensuring that the lab provides support for appropriate software packages for use with the library’s datafiles. greater collaboration might encompass shared access to and support of equipment, delegated support for data storage services (for example, the library acquires and catalogs the datafiles which are housed and retrieved in the lab), or the provision of copying and subsetting services to end users by referral. other computer labs may also be maintained by computing intensive departments, colleges or institutes. for example, the research emphases in many geography departments may make it feasible for them to maintain their own gis labs. c)n my own campus, there is a 15fall/winter 1994 college-supported computing facility for social scientists. in these specialized labs, direct access to equipment by outside users may be more problematic. however, libraries will still benefit from informal ties with their personnel, as these staffs frequently have expertise with appropriate hardware and software. in some cases, even these “closed shops” may be willing to provide some level of pubic access to depository datasets where access to the files is important to their own teaching or research mission. for example, at uc berkeley’s lawrence berkeley lab, they have mounted many library datafiles received through depository distribution on their cdrom network, allowing some pubic access to the campus community, because they considered the files important to their own research. clearly, if a campus computing facility is currently providing the type of in-depth support that jim characterizes as data retrieval and analysis service to a library’s primary clientele, the library would be wise to establish an arrangement to make referrals to that service rather than try to develop such capabilities in-house. such services are so costly and labor intensive that most libraries would make better use of their resources in other areas. computing support groups another important source for informal collaboration and communications in support of computing services are computer users groups. many areas, and even some campuses, have grassroots “user groups” where computer users can share information, expertise, and mentor less sophisticated users. these groups may be organized around computing platforms (ibm, mac, etc.), software packages, or specific tasks (network administration), etc. they may meet for informal discussion, organize training sessions, or sponsor local (or even national) experts as speakers. some may have online mailing lists. in addition, there are news groups and list servers on the intemet that deal with technical issues of interest to data users and providers. again, these may be dedicated to a specific type of hardware or software, aspect of computing support, or substantive data issue. these groups can be invaluable in troubleshooting specific problems. data libraries libraries should be aware of all data libraries that exist on their campus or in their local area. these may be found within computing centers, academic departments or schools, and research units. on-campus likely places to support such libraries include centralized computing centers, teaching department such economics, political science, psychology, and geography, college-level computing facilities in the social sciences and health related fields, and research units concerned with quantitative or survey research. there may be multiple narrow subject-oriented collections in various locations on campus. off-campus data libraries may be found in other academic institutions, city or regional planning agencies, business libraries, or research organizations. as with other computing services, they may be publicly accessible or be “closed shops” with a specific clientele. these libraries may be formally staffed and structured or run by staff or students with other primary responsibilities. if the library is interested in providing what jim characterizes as “the lowest possible level of service,” passive referral services, staff will need to be aware of the existence, holdings, and accessibility of these collections. informal collaboration and communication will also strengthen library services as staff draw on the (sometimes substantial) discipline specific expertise in these facilities. one other possible form of cooperation is for the library to include the data library’s holding in the campus opac. subject experts another important source of informal collaboration are subject experts in datadependent disciplines. library personnel will benefit immeasurably from contact with these data users. in most institutions they will tend to be members of the faculty engaged in quantitative teaching or research in disciplines such as statistics, economics, political science, sociology, psychology, management, organizational studies, public health, civil engineering, agricultural economics, education, history, anthropology, etc. others may be in these same disciplines in post doctorate or research appointments. many will or should be users of the library data collections. while most collaboration will be informal, this group will be the constituency best qualified to assist users with areas such as advanced datafile recommendation and datafile use advisory services. when users have advanced questions as to the content of a particular datafile and its suitability for a specific research application, or seek advice on specific research methodologies, statistical techniques or software, referrals to other more expert users in their department or subject discipline may be the only means of providing assistance. while most experts would be unwilling to enter into a formal agreement to provide public consulting on such matters, many would consider it professional courtesy to provide minimal assistance to a colleague. statistics labs libraries in institutions with statistics labs (or the equivalent) may wish to develop cooperative relationships with these facilities. the discipline of statistics influences the method of inquiry in almost every discipline from agriculture and engineering to social and medical sciences. campuses sometimes provide centralized laboratories in support of teaching 16 iassist quarterly and research involving statistical methods. these facilities may incorporate computing equipment and range of both specialized and general purpose statistical software, as well as consulting on statistical methods and research design. any library considering the provision of data analysis or advisory services will want to investigate the existence of such facilities on its campus. again, collaboration may be informal and the statistics facility may only serve as a referral point for more complex methodological questions. local contacts options may also exist for inter-institutional cooperation. many data producers have their own distribution networks. local members of those networks can be of great assistance and may have access to datafiles outside the library’s holdings. libraries will benefit from knowledge of and contact with any such contacts in their local area or region. relevant networks include the state census data center network, the business and industry data center network (both part of u.s. bureau of the census), the bea’s regional economic measurement users group, and data centers receiving files on deposit from the national center for health statistics data tape program. another obvious option for inter-institutional cooperation is other local libraries with machine-readable collections. cooperative support of service for datafiles may take many forms, including sharing of expertise and coordinating referrals between institutions. for example, a public library with limited data holdings is likely to benefit immensely by communication and collaboration with a larger academic institution nearby that has more extensive resources for its data services. conversely, the large academic depository will benefit from close ties to a local public collection where the general public may be referred for basic assistance. more creative arrangements might include coordinated collection development and selection of datafiles within various subject disciplines. conclusion this list is not exhaustive. it is meant to be suggestive of the types of relationships a library might seek to develop and some logical places to look for partnerships. a library’s options for collaboration will be varied, and one library’s options will differ from anothers given their differences in institutional setting, mission, and resources. it should be clear, however, that no library is likely to be in a position to “do it all.” financial and personnel resources will be a primary limiting factor. even if these resources were limitless (especially unlikely in the current economic climate), there will be certain roles that are inappropriate within the traditional library model. as jim suggests, more complex data analysis services are may fall into this category of service. most librarians would agree that it is not their role to evaluate the reliability of print sources, or to interpret research results or statistical tables for end user. for most libraries it will then follow that even with appropriate technical or subject background some activities are rightly outside the scope of the library’s public service mission. these activities may include advising on research methodology, analytical procedures, sample design, statistical techniques, as well as software selection, and result interpretation. unless a library has access to a comprehensive data analysis service, these activities should be avoided and specifically excluded from its public service policy. 1. paper presented at iassist in san francsisco, may 1994. 2. jim jacobs, data services and collections (hand-out prepared for the iassist/godort workshop, public service for numeric datafiles: issues for depository," held february, 1994 at ucla). vol272.indd 4 iassist quarterly summer 2003 editorʼs notes welcome to the second issue of the iassist quarterly vol. 27. again we present three articles in an iq issue. the articles are elaborations from a conference, from a local presentation, and a librarianʼs thoughts on information products management. at the iassist ottawa conference in may 2003 the session on “the roots of historical censuses: the archivistʼs perspective” held several perspectives. cara downey from the library and archives of canada presents in her paper on “the census of canada” an “archival perspective”. according to cara downey it all started in 1666 when intendant jean talon enumerated the 3.215 inhabitants of new france. it is said that “the census has been a wonderful source of information about canadians”. the paper includes a chronological description of the development of the canadian census, and the paper also discusses the aspects of archival value. newpaper clippings form the foundation of much research. juan linz is sterling professor emeritus of political and social science at yale. the newspaper clippings of juan linz form the foundation of the archive of the spanish transition to democracy for the period 1975-1983. martha peach has in spain presented the archive, but here in the iq the collection and the process of digitalization are described in english. martha peach is director of the library, centre for advanced studies in the social sciences, juan march institute, madrid, spain. the title of the presentation is subtitled “a project of analysis, digitalization and event history database”. the iassist quarterly mostly uses presentations and longer statements from the iassist conferences and iassist members. the iassist community is not that common, and we can use and welcome some input from outside. so it is with pleasure i announce the article from the librarian rajashekhar d. kumbar (jansons school of business, karumathampatti in india). from rajashekhar d. kumbar we have received an article with the title “information products management in the internet age”. the article takes its base in the electronic market place of today, with digital goods. the abstract states that “information or knowledge goods are a peculiar kind of commodity. management of these information products requires librarians to deal with information not just as a set of objects or artifacts such as data or files, but also as a process that extends from information identification (sensing), collection and organization through its processing, maintenance and use.”. this is a promising point of departure for a discussion and development within iassist. remember to visit the iassist website on www. iassistdata.org. papers for the iassist quarterly are most welcome. papers can be from iassist conferences, from other conferences, from local presentation, etc. so please contact the editor (kbr@sam.sdu.dk). karsten boye rasmussen, march 2004. http://www.iassistdata.org/ http://www.iassistdata.org/ mailto:kbr@sam.sdu.dk 1/2 schwartz, ofira & hayslett, michele (2025) from fair principles to data competencies: evolving library support for data-driven scholarship, iassist quarterly 49(3), pp. 1-2. doi: https://doi.org/10.29173/iq1178 the creative commons-attribution-noncommercial license 4.0 international applies to all works published by iassist quarterly. authors will retain copyright of the work and full publishing rights. editors’ notes: from fair principles to data competencies: evolving library support for data-driven scholarship dear iassisters, welcome to iassist quarterly vol. 49 no. 3. we trust that you had a pleasant and restorative summer, and that the academic year is off to a good start. we are grateful to the thoughtful reader or author who recently recommended iq for inclusion in scopus. we’ve received notice of this recommendation directly from scopus. at this time, however, we have chosen to delay the review process. a successful application requires a comprehensive set of editorial and ethical policies, and we are still in the process of developing and strengthening these to meet scopus’ standards. the editors and editorial board of iq are actively working to ensure our policies not only align with scopus’ criteria but also provide a clear, consistent, and equitable experience for authors, reviewers, and readers. we will continue to share policy updates on the journal’s website. we value your engagement and welcome your input and feedback. as the research landscape continues to evolve, librarians and other data support professionals are increasingly called upon to develop tools, competencies, and services that foster transparency, usability, and accessibility in data-driven scholarship. the four articles in this issue highlight a shared commitment to empowering diverse research communities through innovative approaches to data stewardship and support. all of them underscore the importance of adaptable infrastructure, inclusive training, and responsive guidance. together, they reflect the growing role of libraries in advancing equitable and effective research data practices across disciplines. this issue opens with the winning submission of the 2024 iassist conference paper competition, titled “how are we fair-ing? creating a fair self-assessment checklist for data repositories,” authored by lauren phegley and lynda kellam. the paper outlines an initiative undertaken by the penn libraries research data team to develop a self-assessment tool based on the fair principles—findable, accessible, interoperable, and reusable. this tool is designed to help data repository teams evaluate whether their repository’s policies, infrastructure, and documentation support the deposit of fair-compliant data. the authors also discuss the intended applications of the tool and share key insights gained throughout its development. meryl brodsky and hannah chapman tripp, conduct a retrospective review, exploring the data-related competencies required by liaison and subject librarians to effectively support academic researchers. in their article “data competencies for liaison librarians: a scoping review,” the authors map data-related https://doi.org/10.29173/iq1178 https://creativecommons.org/licenses/by-nc/4.0/ 2/2 schwartz, ofira & hayslett, michele (2025) from fair principles to data competencies: evolving library support for data-driven scholarship, iassist quarterly 49(3), pp. 1-2. doi: https://doi.org/10.29173/iq1178 competencies over a ten-year period (2012-2022) with particular attention given to the skill sets that liaisons, or non-data librarians, may need to develop or hone. the study identifies key competencies that are critical for librarians supporting researchers with data-related inquiries. it underscores the importance of structured training and offers guidance on prioritizing specific skill areas. in their article “adventures in data visualization support assessment: when the gap you were trying to identify turns out to be a chasm,” authors meg miller, grace o’hanlon and hafizat sanni-anibire present findings from two studies conducted by a group of librarians at the university of manitoba. these studies explore emerging patterns in research mobilization methods and assess the structural support needed to enhance libraries’ ability to effectively serve their users. authors alison sizer, andreas mastrosavvas, and oliver duke-williams present an overview of the (u.k.) office for national statistics’ longitudinal study (ons-ls) and the unique opportunities it offers for longitudinal research on the population of england and wales. their article, “the ons longitudinal study – opportunities for longitudinal research on the england and wales population,” details the scope and content of the ons-ls and introduces the centre for longitudinal study information and user support (celsius), highlighting its role in assisting both current and prospective users. the article also includes a comparative analysis of the ons-ls alongside other census-based longitudinal studies within an international context. enjoy the reading! michele hayslett and ofira schwartz, september 2025 https://doi.org/10.29173/iq1178 editors’ notes: from fair principles to data competencies: evolving library support for data-driven scholarship instructions for authors of the iassist quarterly 1/15 phan l. et al. (2021) a model for data ethics instruction for non-experts, iassist quarterly 46(4), pp. 1-15 doi: https://doi.org/10.29173/iq1028 a model for data ethics instruction for non-experts leigh phan1, stephanie labou2, erin foster3, ibraheem ali4 abstract the dramatic increase in use of technological and algorithmic-based solutions for research, economic, and policy decisions has led to a number of high-profile ethical and privacy violations in the last decade. current disparities in academic curriculum for data and computational science result in significant gaps regarding ethics training in the next generation of data-intensive researchers. libraries are often called to fill the curricular gaps in data science training for non-data science disciplines, including within the university of california (uc) system. we found that in addition to incomplete computational training, ethics training is almost completely absent in the standard course curricula. in this report, we highlight the experiences of library data services providers in attempting to meet the need for additional training, by designing and running two workshops: ethical considerations in data (2021) and its sequel data ethics & justice (2022). we discuss our interdisciplinary workshop approach and our efforts to highlight resources that can be used by non-experts to engage productively with these topics. finally, we report a set of recommendations for librarians and data science instructors to more easily incorporate data ethics concepts into curricular instruction. keywords data ethics, data literacy, data science, data privacy, algorithmic bias, community engagement introduction the last decade has seen an unprecedented increase in the availability and prevalence of technological solutions used to simplify complex decision-making processes. importantly, use of these technologies is so pervasive that it influences nearly every aspect of modern life. these solutions span from algorithms that regulate content and marketing that a typical person sees in their social media feed (orlowski, 2020), to resumé reading software for job applicants (borsellino, 2018), to recidivism predictors for former criminals (angwin et al., 2016; o’neil, 2016), to insurance eligibility algorithms that approve funding for vulnerable populations (obermeyer et al., 2019). certain models have yielded substantial benefits; for example, pandemic prediction models for informing public policy decisions, among others (alzahrani et al., 2020). however, other models are filled with explicit or implicit biases that exacerbate racially biased decision-making, amplify stereotypes or misinformation, and make it extremely difficult to identify those accountable for failing to properly train these technologies to identify bias (o’neil, 2016; noble, 2018 benjamin, 2019). while the need for thoughtful exploration of algorithms within an ethical framework is increasingly recognized (holmes et al., 2021), academic curriculum has lagged in incorporating these aspects as requirements. this issue is similar to the increased need for data literacy training, more generally, across academic domains. disciplines relating to science, technology, engineering, and math (stem) tend to incorporate some data literacy instruction into their curriculum; however, disciplines in the social sciences and humanities generally do not have the resources to do so (dennis et al., 2021). https://doi.org/10.29173/iq1028 2/15 phan l. et al. (2021) a model for data ethics instruction for non-experts, iassist quarterly 46(4), pp. 1-15 doi: https://doi.org/10.29173/iq1028 as a result of this curricular gap, academic libraries are increasingly called upon to deliver data literacy education through community instructional approaches like the carpentries (carpentries, 2021), or consultations catered to individuals and small groups (ucla dsc, 2021; acrl report, 2021). in addition to this disparity that exists in data science and digital literacy training, ethical uses in data or data ethics are generally not required for degree completion. while data scientists can technically come from any discipline, research data ethics practices are almost certainly not covered at an undergraduate level. even curricular requirements for computer science majors at uc campuses, where data science is most commonly taught, reveals there is limited (ucla curriculum, 2021; uc berkeley curriculum, 2021), optional (uc irvine curriculum, 2021) or no formal required training (ucsd curriculum, 2021) in data ethics as part of their curricular instruction. historically uc libraries’ instruction whether through the carpentries or other models was limited to their respective campus communities. this siloed approach to instruction changed when the 2020 covid-19 pandemic induced a complete disruption of in-person instruction. in response to the move to virtual instruction, uc campuses pivoted rapidly to remote learning environments. despite the challenging circumstances, remote learning presented new opportunities for collaboration across the uc campuses and allowed for a broader reach of previously campus-specific events. a compelling example of this is the uc libraries co-hosting uc love data week in february of 2021, which was comprised of a week of free events for uc affiliates on a wide range of topics including: data publishing, data repositories, data cleaning, and text mining to name a few (universities of california, 2021). as love data week organizers from uc berkeley, ucla, and uc san diego, we saw this collaborative uc love data week as an opportunity to present to a wider audience across the uc system on the subject of data ethics. as data information professionals, we are familiar with some aspects of data ethics, but none of us have received formal training or education on this topic. however, we felt that the importance of the topic, in addition to our professional experience with ethical issues in data, created a unique opportunity to begin a dialogue among uc affiliates and lay the groundwork for a cross-campus community interested in continuing conversations and learning around this subject. for uc love data week 2022, we broadened the scope of the initial workshop to feature student speakers from different disciplines. importantly, taking into consideration the highly diverse body of learners across the uc system, we sought to design a workshop that allowed non-experts from a wide variety of disciplines to engage with the topic of data ethics. workshop approach in this section, we will go into detail about how we structured our 2021 workshop before moving into how we incorporated lessons learned from the first iteration to inform our 2022 workshop. we grounded our initial workshop as part of uc love data week 2021 with the overarching question: who is impacted by your research? due to the time allotted for this workshop to cover a topic as broad as data ethics (e.g., only 1 hour in 2021), this guiding question allowed us to help attendees consider ethical considerations in their own research, beyond the confines of the workshop. leading with this guiding question also allowed us to limit the scope and introduce critical topics as applied to research data analysis in various disciplines (phan et al., 2021). to inform discussions, we provided examples of ethical issues related to data that are well-described in the academic literature. we focused on https://doi.org/10.29173/iq1028 3/15 phan l. et al. (2021) a model for data ethics instruction for non-experts, iassist quarterly 46(4), pp. 1-15 doi: https://doi.org/10.29173/iq1028 transferable concepts to help attendees build context for questions of data ethics that might arise in their respective disciplines (figure 1). figure 1: background and interests of attendees for 2021 workshop. a. learners were asked to indicate the topic of most interest to them in the workshop. b. learners for the 2021 workshop were asked to indicate their primary professional background and disciplinary areas. for the 2021 workshop, we created a pre-workshop survey to gather information about our workshop attendees' backgrounds and interests as they relate to topics in data ethics in order to best customize workshop content. the workshop attracted participants from a wide variety of disciplines, including 40.5% from social sciences and humanities, 21.6% from health or physical sciences disciplines, and 37.8% self-described as other or interdisciplinary, across many levels of experience. while preparing for the initial workshop, we focused on three aspects of data ethics: data privacy, algorithmic bias, and engaging communities in research. from the pre-workshop survey, 16.2% were interested in the topic of data privacy, 37.8% were interested in algorithmic bias, and 45.9% were interested in community engagement and social justice (figure 1) so we framed our workshop content around these three topics. our workshop began with lecture content that described real world examples in which these three ethical topics were at play. following the lecture portion of the workshop, we https://doi.org/10.29173/iq1028 4/15 phan l. et al. (2021) a model for data ethics instruction for non-experts, iassist quarterly 46(4), pp. 1-15 doi: https://doi.org/10.29173/iq1028 hosted breakout discussions in which we asked attendees to bring up ethics issues that they observe in their fields of study, related to one of the three data ethics themes discussed in more detail below. data privacy: striking a balance between privacy and transparency there have been a number of high profile data breaches in the last several years (larson, 2017; barrett, 2019; alder, 2020; mihalcik, 2020). for this reason, we expected most attendees to be familiar with the general concept of data privacy, and that breaches of privacy have consequences. therefore, we chose to focus more narrowly on examples that highlighted disparities in the consequences of data breaches on individuals, based on political context, economic and social status. for example, individuals experiencing poverty can be subjected to high levels of police surveillance, but have few tools to protect themselves from breaches of that information (green and gilman, 2018). vulnerable populations that work with researchers, like people who identify as transgender or sex workers, may be at extreme risk in the event that their identities are revealed in the event of a data breach (sandy et al., 2019; sinha, 2017). while topics such as personally identifiable information (pii) are often covered in human-subjects research trainings such as those created by institutional review boards (irb), we wanted to go beyond the technicalities described in those trainings and emphasize the importance of data privacy beyond merely a mandatory training and study design approval before research can begin. importantly, we highlighted the study published by rocher and colleagues in 2019 which showed that it is possible to re-identify the vast majority of the american population using only 15 demographic attributes (rocher et al., 2019), and showcased other examples where supposedly anonymous data was reidentified. with that context in mind, and following discussion of disparities in consequences of data disclosure, we provided a brief overview of methodologies for anonymizing data, removing personally identifiable information (pii), and aggregating data as a means of protecting the identities of individuals associated with studies. this workshop took place in february 2021 and we could safely assume every participant had heard of the decadal census taking place in 2020. the us census serves as an example of a heavily relied upon data set used for informing public policy and determining state and federal representation. beyond being a dataset that affects literally every us resident in a practical sense, it is also used for a wide array of interdisciplinary research and as such is applicable to many researchers. this made the 2020 census an ideal case study to discuss how geographic scales can impact data reidentification; trade-offs between microdata and aggregated data, in terms of variables made public; and how differential privacy, specifically statistical methods for introducing noise into a dataset, can be used to preserve individual participant privacy (see abowd, 2018). importantly, we did not offer a single one-size-fits-all conclusion for this portion of the workshop, but rather, encouraged attendees to think critically about their data, weighing the importance of data privacy vs transparency. as academic spaces grapple with a reproducibility crisis (baker, 2016), partially due to incomplete data transparency, researchers must also be mindful that data shared is not done so at the expense of participant privacy. to connect this idea to workshop attendees’ own work, we concluded this portion with the following questions: • what are the minimum variables needed for meaningful analysis? • could they avoid collecting unnecessary identifiable data? https://doi.org/10.29173/iq1028 5/15 phan l. et al. (2021) a model for data ethics instruction for non-experts, iassist quarterly 46(4), pp. 1-15 doi: https://doi.org/10.29173/iq1028 • what are the best practices or accepted de-identification methods in their research domain? • what would be some privacy-preserving statistical methodologies that are most suitable or needed for their work? research models and algorithmic bias to build on the concept of how data-driven research impacts vulnerable populations beyond aspects of privacy, we next covered predictive research models themselves, providing examples of how bias in algorithms can lead to bias in results. we began with a conceptual example that the audience would be familiar with, by explaining that researchers often use predictive models to understand complex phenomena. in this case, we used two examples of such models: infection prediction rates of covid19 and the trajectory of a hurricane. importantly, we pointed out to learners that predictive models rely on specific assumptions in order for their predictions to be fulfilled correctly; similarly, algorithms follow the same principles. if the assumptions are biased, or fail to acknowledge bias in the data collection, the outcome of a predictive model will be biased. we unpacked three examples where assumptions in algorithms lead to racial or unjust bias. these examples included a variety of commonly-applied technologies, including algorithms used for prediction of recidivism, health insurance financial allocation, and facial recognition. in our first example, we introduced the compas ('correctional offender management profiling for alternative sanctions') recidivism algorithm, a software used to assess a defendant’s risk of committing another crime within two years. as part of the compas system, once defendants are booked in jail, they are required to answer a questionnaire; the algorithm incorporates the questionnaire results to predict the defendant’s likelihood to reoffend. in 2016, angwin and colleagues investigated compas and found that it resulted in high false positives, predicting black defendants to be at a higher risk of recidivism than they actually were two years later (angwin et al., 2016). we emphasized that rather than a single flawed algorithm, compas is one of numerous examples of algorithms which exacerbate systemic racial bias in the justice system by using data informed by racially biased policing to train algorithms that predict recidivism or criminality and failing to protect individuals from false-positives. (o’neil, 2016). our second example focused on an algorithm commonly used to simplify funding approvals for insurance claims. in 2019, obermeyer and colleagues uncovered that an algorithm designed to provide additional support for people with significant healthcare needs was racially biased. the algorithm designers used money spent on healthcare as a proxy for healthcare needs, assuming that the more money spent, the sicker the person. however, the designers failed to recognize that disparities in access to healthcare results in less healthcare spending for black patients on average compared to white patients (obermeyer et al., 2019). as a consequence, the algorithm assumed black patients were healthier on average than white patients, even though the black patients suffered from more chronic health conditions than the white patients at any given health score created by the algorithm. by using the number of chronic conditions, rather than money spent, to determine a health score, obermeyer and colleagues were able to more than double the insurance approval rates for black patients from 17.7% to 46.5% (obermeyer et al., 2019). for this example, we highlighted to our learners that the choice of convenient, seemingly effective proxies, for difficult to access information can be an important source of algorithmic bias in contexts such as this. https://doi.org/10.29173/iq1028 6/15 phan l. et al. (2021) a model for data ethics instruction for non-experts, iassist quarterly 46(4), pp. 1-15 doi: https://doi.org/10.29173/iq1028 in our final example, we highlighted facial recognition software, specifically use cases of facial recognition in law enforcement. the gender shades project spearheaded by buolamwini and colleagues assessed the accuracy of multiple commercial facial recognition software. while many of the facial recognition software had high accuracy, all companies performed better on male than female subjects, and generally on lighter-skinned subjects overall. the authors discovered that most of the facial recognition algorithms were trained on caucasian samples, resulting in poor accuracy, underrepresentation and misrepresentation by gender and race (buolamwini and gebru, 2018). in addition to gender shades, we introduced fairface, a project led by karkkainen and colleagues, who developed a more balanced racial composition compared to several other datasets used commercially for facial recognition (karkkainen and joo, 2021). in contrast to fairface, we presented examples of amazon’s rekognition software at one point used by the u.s. immigration and customs enforcement agency which mistakenly matched u.s. members of congress with a mugshot database. juxtaposing examples of improvements to facial recognition accuracy along with cases that demonstrate the negative impact of inaccurate models allowed workshop participants to critically examine the realworld consequences of algorithms on communities impacted by these models. community engagement and research impact a critical component for holistically considering appropriate methods for data anonymization and creating ethical and equitable algorithms is engagement with the communities impacted by the work. therefore, our last topic focused on providing our learners with examples of how research projects across disciplines have constructively engaged with the communities affected by their research. we highlighted three ways this engagement has taken place: (1) reporting results to study participants, (2) utilizing community expertise, and (3) enabling broader social change through policy development. as an example for reporting results to study participants, we highlighted the for healthy kids! project. this project, published by thompson and colleagues in 2017, was an examination of pesticide exposure amongst primarily immigrant farmworkers and non-farmworkers in a region of the state of washington, united states. the authors used a community based participatory research approach to engage with the community consistently throughout the data collection process, such as using town halls and community boards to inform the community about the dangers of chronic pesticide exposure. following the study, the research team reached out to the individual study participants to provide their results. while the authors were unable to reach all study participants, their efforts highlight the importance of using existing community infrastructure and direct communication to make a greater impact to study participants in a vulnerable community (thompson et al., 2017). in our second scenario, as we had discussed the dangers of racial bias associated with predictive policing software earlier, we wanted to note a constructive example of community engagement with a historically criminalized population. in the case of the safe lab led by dr. desmond patton out of columbia university, this research group studies social media communication from gang-involved and affiliated youth and how that might result in 'off-line' instances of community violence (frey et al., 2019). in the case of this particular work, the safe lab actively works with community domain experts (often former gang members) to ensure that the social media communications are not misinterpreted by the researchers and accurately analyzed (patton et al., 2019). additionally, to address community concerns about the use of social media data as a means of surveillance, particularly by police, the safe https://doi.org/10.29173/iq1028 7/15 phan l. et al. (2021) a model for data ethics instruction for non-experts, iassist quarterly 46(4), pp. 1-15 doi: https://doi.org/10.29173/iq1028 lab has established a code of ethics that helped the researchers explore different modalities of obtaining consent and did not allow data to be shared with groups that engage in punitive and criminalizing actions. in our final example, we introduced learners to the los angeles-based million dollar hoods project, led by dr. kelly lytle hernández, which is a community-based research initiative focusing on the human and fiscal costs of mass incarceration (lee, lytle hernández and tso, 2018). this initiative engages deeply with communities in and around los angeles county to 1) gain access to relevant data on incarceration in these regions and 2) produce outputs such as data driven reports and dashboards that show the disproportionate effect of incarceration. additionally, this research initiative emphasizes building skills of community members to allow for them to actively engage with research focused on their needs. an example of this work is the big data for justice summer institute, held at ucla, which provides training on working with data and relevant tools for both university students and community members (big data for justice summer institute, 2022). most compellingly, the million dollar hoods project has informed changes to local and state policy around issues of mass incarceration by providing data-driven evidence of disparity and discrimination. a cornerstone of this work is 'centering' the voice of the communities most affected by these issues to inform the research questions and contribute to the research process. together, these scenarios demonstrate that direct community engagement improves the distribution of research results, generates new pathways to build expertise amongst communities, and can help improve public policy on issues of direct relevance to the respective communities. incorporating data ethics into data literacy instruction at times it may seem that ethical issues in data collection and modeling are so pervasive that engaging with the topic feels daunting, especially for those without any formal training on these issues. fostering intentional engagement within individual sub-disciplines can lower the barrier and amplify engagement with these critical issues. in doing so, we as librarians can take ownership over how data ethics impacts us as researchers, resource providers, and data users directly, and provide a higher level of support and guidance on best practices for patrons. while we might not feel like experts as instructors on these topics, by asking critical questions (as indicated above) we were able to feel empowered to find resources that impact our own work (such as census data, or facial recognition datasets) while presenting examples that would inspire our audience to consider ethics in their own disciplines. furthermore, making connections between themes in data ethics and examples that are relevant to learners’ backgrounds can create entry points to integrating ethics in existing library data literacy curricula. understanding and learning from your audience’s background in preparation for the initial workshop in 2021, we conducted a pre-workshop survey which allowed us to gauge our audience’s backgrounds and their interest in the themes we planned to present, and to better prepare for facilitating breakout discussions. being attentive to learners’ backgrounds and customizing examples to match learner interest (as expressed in the pre-workshop survey) creates an environment more conducive to learners making connections between themes within data ethics and their own field(s) of study. presenting a variety of examples which resonate with the audience not https://doi.org/10.29173/iq1028 8/15 phan l. et al. (2021) a model for data ethics instruction for non-experts, iassist quarterly 46(4), pp. 1-15 doi: https://doi.org/10.29173/iq1028 only lowers the barrier to entry for learners to engage in these topics, but also provides a framework for them to identify ethical considerations within their own research beyond the workshop setting. with the knowledge that uc love data week 2021 was well-attended among uc graduate students and staff, we planned the 2022 data ethics & justice workshop with this audience in mind. building on the lessons learned and feedback from the 2021 workshop, we designed the 2022 iteration to highlight the work of three graduate researchers from various disciplines whose work intersects with data ethics. these presentations were used to demonstrate real-world ethics consideration in both research and industry and also provided a framework for breakout room discussions, themed to the issues brought up by each speaker. our first speaker for the 2022 data ethics & justice workshop was a phd candidate in gender studies researching racism and disinformation through social media data analysis. to begin the workshop, they discussed their experience gathering twitter data, which led them to question the ethics of mass social media data collection. they recommended protecting the privacy of social media users, and offered ethical frameworks for those conducting research using social media data. the second speaker was a master's candidate in statistics whose work experience included medical billing and information security at a global healthcare software company. they emphasized the importance of cybersecurity on personal and organizational levels and provided an example of a real-world healthcare system whose database was compromised. this speaker concluded by giving recommendations for protecting our own personal privacy and data as well as those in our own organizations. our third speaker was a phd candidate in computer science who focuses on human-computer interaction and investigates ways to improve privacy communication. they presented on limitations of consent for ethical data collection, providing examples from their own research in which individuals may either feel pressured to provide their personal information, unknowingly give more of their own personal information, or how social media sites can make inferences about an individual based on the information their friends provide. following the presentations, each speaker paired with a workshop organizer to facilitate discussion in a breakout room, focused on the data ethics aspect discussed in the respective speaker’s presentation. we aimed to lower the barrier for engagement by encouraging attendees to join discussion rooms whether or not they felt ready to contribute to the discussion, highlighting the value of learning through listening. in addition, we created a notes document accessible for all participants to share discussion points. following 25 minutes in breakout rooms, all participants reconvened for a group recap of breakout room main points and concluding thoughts before the workshop ended. https://doi.org/10.29173/iq1028 9/15 phan l. et al. (2021) a model for data ethics instruction for non-experts, iassist quarterly 46(4), pp. 1-15 doi: https://doi.org/10.29173/iq1028 figure 2: background of workshop attendees for 2022 data ethics and justice workshop. a. learners were asked to indicate their academic status from a drop-down list. b. learners were asked to indicate their primary professional background and disciplinary areas. in designing both our initial and 2022 workshops, we focused on a common theme of understanding the audience’s background. for example, since a large portion of 2021 attendees were graduate researchers and staff, we recognized the benefit for the audience to learn directly from peer researchers. this approach provided scholars the opportunity to share their work while creating a forum for discussion around data ethics across disciplines.the demographics of registrations for our 2022 workshop suggest that there is high interest from scholars, staff, and community members outside the university. overall, there is evidence of increasing interest on the part of staff and interdisciplinary researchers. https://doi.org/10.29173/iq1028 10/15 phan l. et al. (2021) a model for data ethics instruction for non-experts, iassist quarterly 46(4), pp. 1-15 doi: https://doi.org/10.29173/iq1028 figure 3: registrants’ disciplinary areas for 2022. learners were asked to indicate their professional background and disciplinary areas. this information was collected solely for the 2022 workshop. work with disciplinary instructors to identify curricular gaps considering the growing interest, and limited required training, in data ethics and justice across disciplines and professional fields, as library professionals, we can incorporate data ethics training into data literacy curricula through partnerships with disciplinary instructors. moving beyond one-time workshops, strategic partnerships between data science instructors, instructors that specialize in a particular discipline, and librarians can be a first step to identifying an approach to data ethics training that leverages all instructors’ respective expertise. working with disciplinary instructors can help set the scope of the audience’s background, and instructors can incorporate data ethics examples in standard data literacy curricula. for example, as data science and data literacy instructors are often called on to teach research data management, instructors can incorporate topics of data privacy and examples of privacy violations (and how to avoid them) into existing training modules. domain-specific examples and experience from instructors, complemented by data literacy fundamentals from librarians, can also serve to emphasize for learners the idea that researchers can and should directly engage with the community they are studying. learners get a better sense of the appropriate level of privacy required for the data being collected in their field of study, as well as any disciplinary standards and best practices for data de-identification. furthermore, when developing models or software, it can be critical to put faces to the data, so to speak, by involving communities who will be impacted by the models themselves. these communities can also inform selection of proxies for difficult to measure parameters, as demonstrated in the examples of algorithmic bias. https://doi.org/10.29173/iq1028 11/15 phan l. et al. (2021) a model for data ethics instruction for non-experts, iassist quarterly 46(4), pp. 1-15 doi: https://doi.org/10.29173/iq1028 these concepts apply to a wide range of disciplines and fields, and allow researchers to make more informed decisions throughout the research lifecycle. collaborate with offices of research cross-institutional collaborations may present opportunities to further expand the reach of data ethics curricula and data literacy overall. such collaborations can develop by promoting data ethics-related instructional workshops and events with cross-campus colleagues and units. within a single campus, forming partnerships and lines of communication with offices of research could help identify gaps unseen by individual disciplinary experts. furthermore, working with offices of research and other campus-wide, domain-agnostic entities could help in identifying existing platforms such as seminar series and journal clubs both within and across disciplines which could be a useful method of bringing these workshops directly to researcher communities. while we did not directly engage offices of research in our 2021 or 2022 workshops, we identified them as potential collaborators for future workshops of this nature. while the library is a domain-agnostic campus entity suited to providing training in ethical issues in data, offices of research are also well-versed in this area and can provide additional real-world and campus-specific examples relevant to learners. conclusion in response to increased demand for data science professionals and cross-disciplinary data science training (dennis et al., 2021), community-developed and collaborative models of providing computational training have emerged (e.g., the carpentries), but standardized training for data ethics remains scarce. the workshops described here serve as starting points for data services providers and instructors to expand their knowledge and develop critical competencies to incorporate ethics discussions in their own work with patrons as well as during formal instruction and training settings. through teaching this workshop, we discovered that there is substantial interest across the uc system in these topics, and an increasing interest among staff and particularly among fellow librarians. we developed an approach to teaching these topics with limited formal training which includes noting our non-expert status for learners and found that we could build on this work as we continue to engage with our research communities. the time for workshops such as this is now. there are a number of major societal consequences when research is conducted without deep consideration of the communities involved in the research process. or, as we described in our workshop: who is impacted by your research? as librarians and data services practitioners play an increasingly important role in data literacy instruction, data analysis guidance, and data management best practices, intentional engagement with these topics can begin to address some of the downstream consequences of the gaps in this training. we at the universities of california libraries, expect to build on the lessons learned from teaching these workshops and work within and outside our institutions to further identify and fill curricular gaps in data ethics training. https://doi.org/10.29173/iq1028 12/15 phan l. et al. (2021) a model for data ethics instruction for non-experts, iassist quarterly 46(4), pp. 1-15 doi: https://doi.org/10.29173/iq1028 references abowd, j. m. (2018). the u.s. census bureau adopts differential privacy. proceedings of the 24th acm sigkdd international conference on knowledge discovery & data mining, 2867. available at: https://doi.org/10.1145/3219819.3226070 acrl report. (2021). arl/carl joint task force on research data services releases final report. association of research libraries. available at: https://www.arl.org/news/arl-carl-jointtask-force-on-research-data-services-releases-final-report/ alder, s. (2020, august 17). healthcare data leaks on github: credentials, corporate data and the phi of 150,000+ patients exposed. hipaa journal. available at: https://www.hipaajournal.com/healthcare-data-leaks-on-github-credentials-corporatedata-and-the-phi-of-150000-patients-exposed/ alzahrani, saleh i., ibrahim a. aljamaan, and ebrahim a. al-fakih. “forecasting the spread of the covid-19 pandemic in saudi arabia using arima prediction model under current public health interventions.” journal of infection and public health 13, no. 7 (july 2020): 914– 19. available at: https://doi.org/10.1016/j.jiph.2020.06.001 angwin, j., larson, j., mattu, s., & kirchner, l. (2016, may 23). machine bias. propublica. available at: https://www.propublica.org/article/machine-bias-risk-assessments-in-criminalsentencing baker, m. (2016). 1,500 scientists lift the lid on reproducibility. nature, 533(7604), 452–454. available at: https://doi.org/10.1038/533452a barrett, d. (2019, july 29). capital one says data breach affected 100 million credit card applications. washington post. available at: https://www.washingtonpost.com/nationalsecurity/capital-one-data-breach-compromises-tens-of-millions-of-credit-cardapplications-fbi-says/2019/07/29/72114cc2-b243-11e9-8f6c-7828e68cb15f_story.html benjamin, r. (2019). race after technology: abolitionist tools for the new jim code. polity. big data for justice summer institute. (n.d.). ucla bunche center. available at: https://bunchecenter.ucla.edu/programs-events/thurgood-marshall-lecture-2/ (retrieved march 10, 2022) borsellino, r. (2018). get your resume past the robots and into human hands. the muse. available at: https://www.themuse.com/advice/beat-the-robots-how-to-get-your-resume-pastthe-system-into-human-hands buolamwini, j., & gebru, t. (2018). gender shades: intersectional accuracy disparities in commercial gender classification. conference on fairness, accountability and transparency, 77–91. available at: https://proceedings.mlr.press/v81/buolamwini18a.html carpentries. (2021). the carpentries. the carpentries. available at: https://carpentries.org/index.html https://doi.org/10.29173/iq1028 https://doi.org/10.1145/3219819.3226070 https://www.arl.org/news/arl-carl-joint-task-force-on-research-data-services-releases-final-report/ https://www.arl.org/news/arl-carl-joint-task-force-on-research-data-services-releases-final-report/ https://www.hipaajournal.com/healthcare-data-leaks-on-github-credentials-corporate-data-and-the-phi-of-150000-patients-exposed/ https://www.hipaajournal.com/healthcare-data-leaks-on-github-credentials-corporate-data-and-the-phi-of-150000-patients-exposed/ https://doi.org/10.1016/j.jiph.2020.06.001 https://www.propublica.org/article/machine-bias-risk-assessments-in-criminal-sentencing https://www.propublica.org/article/machine-bias-risk-assessments-in-criminal-sentencing https://doi.org/10.1038/533452a https://www.washingtonpost.com/national-security/capital-one-data-breach-compromises-tens-of-millions-of-credit-card-applications-fbi-says/2019/07/29/72114cc2-b243-11e9-8f6c-7828e68cb15f_story.html https://www.washingtonpost.com/national-security/capital-one-data-breach-compromises-tens-of-millions-of-credit-card-applications-fbi-says/2019/07/29/72114cc2-b243-11e9-8f6c-7828e68cb15f_story.html https://www.washingtonpost.com/national-security/capital-one-data-breach-compromises-tens-of-millions-of-credit-card-applications-fbi-says/2019/07/29/72114cc2-b243-11e9-8f6c-7828e68cb15f_story.html https://bunchecenter.ucla.edu/programs-events/thurgood-marshall-lecture-2/ https://www.themuse.com/advice/beat-the-robots-how-to-get-your-resume-past-the-system-into-human-hands https://www.themuse.com/advice/beat-the-robots-how-to-get-your-resume-past-the-system-into-human-hands https://proceedings.mlr.press/v81/buolamwini18a.html https://carpentries.org/index.html 13/15 phan l. et al. (2021) a model for data ethics instruction for non-experts, iassist quarterly 46(4), pp. 1-15 doi: https://doi.org/10.29173/iq1028 curriculum, ucla. (2021). 2020-2021 computer science curriculum. computer science curriculum, ucla. available at: https://www.seasoasa.ucla.edu/curric-20-21/83-compsci-cur20.html curriculum, ucsd. (2021). b.s. computer science | computer science. ucsd computer science major requirements. available at: https://cse.ucsd.edu/undergraduate/bs-computer-science curriculum, ucb. (2021). requirements: upper division | computing, data science, and society. available at: https://data.berkeley.edu/degrees/data-science-ba/upper-division curriculum, uci. (2021). computer science, b.s. < university of california irvine. available at: http://catalogue.uci.edu/donaldbrenschoolofinformationandcomputersciences/depart mentofcomputerscience/computerscience_bs/#requirementstext frey, w. r., patton, d. u., gaskell, m. b., & mcgregor, k. a. (2020). artificial intelligence and inclusion: formerly gang-involved youth as domain experts for analyzing unstructured twitter data. social science computer review, 38(1), 42–56. available at: https://doi.org/10.1177/0894439318788314 green, r., & gilman, m. (2018). the surveillance gap: the harms of extreme privacy and data marginalization. n.y.u. review of law & social change. available at: https://socialchangenyu.com/review/the-surveillance-gap-the-harms-of-extremeprivacy-and-data-marginalization/ dennis et al. 2021. from a data archive to data science: supporting current research. in herndon, j. (eds.). data science in the library: tools and strategies for supporting data-driven research and instruction. (1st ed., pp. 99-109). facet publishing. holmes, w., porayska-pomsta, k., holstein, k., sutherland, e., baker, t., shum, s. b., santos, o. c., rodrigo, m. t., cukurova, m., bittencourt, i. i., & koedinger, k. r. (2021). ethics of ai in education: towards a community-wide framework. international journal of artificial intelligence in education. available at: https://doi.org/10.1007/s40593-021-00239-1 karkkainen, k., & joo, j. (2021). fairface: face attribute dataset for balanced race, gender, and age for bias measurement and mitigation. 1548–1558. available at: https://openaccess.thecvf.com/content/wacv2021/html/karkkainen_fairface_face_a ttribute_dataset_for_balanced_race_gender_and_age_wacv_2021_paper.html larson, s. (2017, july 12). verizon customer data leaked through an online security hole. cnnmoney. available at: https://money.cnn.com/2017/07/12/technology/verizon-data-leakedonline/index.html lee, e., lytle hernández, k., & tso, m. (2018). policing transitional-aged youth in culver city. the million dollar hoods project. available at: https://www.culvercity.org/files/assets/public/documents/city-manager/public-safetyreview/policing-transitional-aged-youth-in-culver-city-for-website.pdf https://doi.org/10.29173/iq1028 https://www.seasoasa.ucla.edu/curric-20-21/83-compsci-cur20.html https://cse.ucsd.edu/undergraduate/bs-computer-science https://data.berkeley.edu/degrees/data-science-ba/upper-division http://catalogue.uci.edu/donaldbrenschoolofinformationandcomputersciences/departmentofcomputerscience/computerscience_bs/#requirementstext http://catalogue.uci.edu/donaldbrenschoolofinformationandcomputersciences/departmentofcomputerscience/computerscience_bs/#requirementstext https://doi.org/10.1177/0894439318788314 https://socialchangenyu.com/review/the-surveillance-gap-the-harms-of-extreme-privacy-and-data-marginalization/ https://socialchangenyu.com/review/the-surveillance-gap-the-harms-of-extreme-privacy-and-data-marginalization/ https://doi.org/10.1007/s40593-021-00239-1 https://openaccess.thecvf.com/content/wacv2021/html/karkkainen_fairface_face_attribute_dataset_for_balanced_race_gender_and_age_wacv_2021_paper.html https://openaccess.thecvf.com/content/wacv2021/html/karkkainen_fairface_face_attribute_dataset_for_balanced_race_gender_and_age_wacv_2021_paper.html https://money.cnn.com/2017/07/12/technology/verizon-data-leaked-online/index.html https://money.cnn.com/2017/07/12/technology/verizon-data-leaked-online/index.html https://www.culvercity.org/files/assets/public/documents/city-manager/public-safety-review/policing-transitional-aged-youth-in-culver-city-for-website.pdf https://www.culvercity.org/files/assets/public/documents/city-manager/public-safety-review/policing-transitional-aged-youth-in-culver-city-for-website.pdf 14/15 phan l. et al. (2021) a model for data ethics instruction for non-experts, iassist quarterly 46(4), pp. 1-15 doi: https://doi.org/10.29173/iq1028 mihalcik, c. (2020, march 31). marriott discloses new data breach impacting 5.2 million guests [2020]. cnet. available at: https://www.cnet.com/tech/services-and-software/marriottdiscloses-new-data-breach-impacting-5-point-2-million-guests/ noble, s. u. (2018). algorithms of oppression: how search engines reinforce racism. new york university press. obermeyer, z., powers, b., vogeli, c., & mullainathan, s. (2019). dissecting racial bias in an algorithm used to manage the health of populations. science, 366(6464), 447–453. available at: https://doi.org/10.1126/science.aax2342 o’neil, c. (2016). weapons of math destruction: how big data increases inequality and threatens democracy (first edition). crown. orlowski, j. (2020). the social dilemma [documentary]. netflix. available at: https://www.netflix.com/title/81254224 patton, d. u., pyrooz, d., decker, s., frey, w. r., & leonard, p. (2019). when twitter fingers turn to trigger fingers: a qualitative study of social media-related gang violence. international journal of bullying prevention, 1(3), 205–217. available at: https://doi.org/10.1007/s42380-019-00014-w phan, l., labou, s., ali, i., & foster, e. (2021, august 27). ethical considerations in data. available at: https://doi.org/10.5281/zenodo.5297228 rocher, l., hendrickx, j. m., & de montjoye, y.-a. (2019). estimating the success of re-identifications in incomplete datasets using generative models. nature communications, 10(1), 3069. available at: https://doi.org/10.1038/s41467-019-10933-3 sandy, j. e., herman, j., keisling, m., mottet, l., & anafi, m. (2019). 2015 u.s. transgender survey (usts) (icpsr 37229). resource center for minority data. available at: https://doi.org/10.3886/icpsr37229.v1 sinha, s. (2017). ethical and safety issues in doing sex work research: reflections from a field-based ethnographic study in kolkata, india. qualitative health research, 27(6), 893–908. available at: https://doi.org/10.1177/1049732316669338 thompson, b., carosso, e., griffith, w., workman, t., hohl, s., & faustman, e. (2017). disseminating pesticide exposure results to farmworker and nonfarmworker families in an agricultural community: a community-based participatory research approach. journal of occupational & environmental medicine, 59(10), 982–987. available at: https://doi.org/10.1097/jom.0000000000001107 ucla dsc. (2021). about | ucla data science center. data science center, ucla. available at: https://www.library.ucla.edu/about-0 https://doi.org/10.29173/iq1028 https://www.cnet.com/tech/services-and-software/marriott-discloses-new-data-breach-impacting-5-point-2-million-guests/ https://www.cnet.com/tech/services-and-software/marriott-discloses-new-data-breach-impacting-5-point-2-million-guests/ https://doi.org/10.1126/science.aax2342 https://www.netflix.com/title/81254224 https://doi.org/10.1007/s42380-019-00014-w https://doi.org/10.5281/zenodo.5297228 https://doi.org/10.1038/s41467-019-10933-3 https://doi.org/10.3886/icpsr37229.v1 https://doi.org/10.1177/1049732316669338 https://doi.org/10.1097/jom.0000000000001107 https://www.library.ucla.edu/about-0 15/15 phan l. et al. (2021) a model for data ethics instruction for non-experts, iassist quarterly 46(4), pp. 1-15 doi: https://doi.org/10.29173/iq1028 universities of california. (2021). uc love data week. uc love data week. available at: https://uclove-data-week.github.io/uc-love-data-week.github.io/ _________________ endnotes 1 leigh phan, data scientist, university of california los angeles, ucla library data science center, email: leighphan@ucla.edu. 2 stephanie labou, data science librarian, university of california san diego, uc san diego library, email: slabou@ucsd.edu. 3 erin foster, service lead research data management program, university of california berkeley, uc berkeley library & research it, email: edfoster@berkeley.edu. 4 ibraheem ali, sciences data librarian, university of california los angeles, louise m. darling biomedical library, ucla library data science center, email: ibraheemali@ucla.edu. https://doi.org/10.29173/iq1028 https://uc-love-data-week.github.io/uc-love-data-week.github.io/ https://uc-love-data-week.github.io/uc-love-data-week.github.io/ mailto:leighphan@ucla.edu mailto:slabou@ucsd.edu mailto:edfoster@berkeley.edu mailto:ibraheemali@ucla.edu mi^sist newsletter vol. 1, no. 1 lasslst constitution i. name article 1 the name of the association shall be: international association for social science information service and technology. the association will hereafter be referred to by the acronym: lassist: i headquarters article 2 the headquarters of lassist will be located with the designated treasurer. objectives article 3 the objectives of lassist are: 3.1 to encourage and support the establishment at local and national levels of information centers for data base reference, maintenance, and dissemination. 3.2 to foster international dissemination and exchange of information on significant developments in information centers for statistical and textual 11 machine-readable data bases. 3.3 to coordinate on an international level programs, projects, and general procedural efforts which provide an international forum for the discussion of problems relating to information centers. ,3.4 to promote the development of professional standards and encourage the ii establishment of training courses for data center personnel. activities article 4 to accomplish the objectives of lassist the following activities are envisioned: 4.1 action groups organized to find solutions to specific problems and/or to develop and compile relevant materials for specific projects. 4.2 workshops , seminars or training sessions in any area consistent with the lassist objectives stated above. ^\sist newsletter vol.1, no. 1 4.3 newsletter to be published and circulated regularly to all lassist members. 4.4 any other activities which advance the association's objectives. membership article 5 5.1 anyone interested in supporting the objectives of lassist may apply for individual voting membership. 5.2 other categories of membership may be instituted by the steering committee. 5.3 dues will be established by a majority vote of the steering committee. governance article 6 the association shall consist of a general assembly composed of all voting members and an executive body to be known as the steering committee. the general assembly will establish the general policies of the association and elect members of the steering committee. the steering committee will implement policies, develop activities and future directions for the association and elect officers. 6.1 the general assembly will be organized by regions which will be constituted by the steering committee and lassist members within each region will elect a regional secretary who will serve as the administrative officer for that region. 6.2 the general assembly will meet at least once every three years. 6.3 the steering committee shall be composed of ten members elected by the general assembly and the secretaries of all approved regions. 6.4 a nominating committee of three members will be appointed by the steering committee to prepare a slate for the election of the ten at large steering committee members. any member who receives the support of five other members may have his name placed on the ballot. 6.5 elections will be held by mail and the ten candidates receiving the largest number of votes will be elected. 6.6 the steering cormittee will elect from among the members designated by the plenary general assembly a chairperson, two vice-chairpersons and a treasurer and will appoint a newsletter editor. i)^$sist newsletter vol.1, no. 1 i amendments article 7 amendments to these statutes may be proposed by any member with the support of five signatures. all amendments will be submitted to the general assembly for approval along with the election ballot. amendments approved by a majority of the members voting will be incorporated into the constitution. termination the association may be dissolved by a majority of the members. remaining funds will be transferred to the international social science council. transitional norm the ad hoc steering committee will serve until after the first meeting of the genertl assembly, but in any case not later than december 31, 1978. the ad hoc committee will arrange for a regular election to be held as soon as possible before that date. secretariat reports canadian secretariat report i sharon chappie data clearing house for the social sciences the data clearing house for the social sciences is the canadian secretariat for lassist. activities of lassist are published in the data clearinghouse bulletin . a large campaign for lassist membership was conducted by the data clearinghouse. over five hundred forms were sent out. to date, 53 individuals and institutions have expressed an interest in affiliating with lassist; 9 others wish to be retained only on the mailing list. supplementary membership campaigns are being considered. interest in the established action groups is as follows: data archive registry: 13; data archive development: 13; data acquisition: 14; data documentation: 16; classification: 8; process-produced data: 16. coordinators for each of the groups have been appointed. initial meetings are planned for the coming months. vol31_1.indd iassist quarterly spring 2007 by by eun-ha hong and linda lowry* business data: issues and challenges from the canadian perspective introduction this paper explores the issues and challenges that we have faced as canadian academic business librarians when working with business data. as this is an exploratory study, we hope only to start a discussion among data librarians about some key challenges facing the academic community related to supporting the teaching and research use of business data. our paper begins with a brief discussion of general data trends, followed by a detailed exploration of business data trends and trends in canadian business education. we discuss challenges and issues related to working with business data from both the collections and reference service perspectives, including the pros and cons of providing business data services and support within the library environment. we conclude by suggesting some measures that both academic business librarians and data librarians can take to address some of these challenges. general data trends halliwell observed that over the last twenty-five years, canada’s quantitative researchers have benefited from a substantial increase in data supply, but there has been an even larger increase in the demand for data1 while he is referring primarily to the supply of and demand for government-produced data sets (from statistics canada or other government agencies), his observations could also apply to the demand for data for academic research, particularly in the social sciences. davis and vickery recently identified several current trends that indicate the growing importance of data sets (over journal articles) as a unit of information currency, including the transformation of data sets into valuable economic commodities, an increase in legislative initiatives to protect data sets, the sheer growth and manipulability of data sets, evolving publication expectations placed on authors, and public/private partnerships developing around data sets2 business data trends as canadian academic business librarians, we have also observed an increase in the demand for business data. before we explore business data trends in more detail, it would be prudent to define what we mean by business data. business researchers either gather their own data (via surveys or other means) or rely on secondary sources of data. these secondary data sources may be divided into two types: social science data and commercial data. an example of a social science data set frequently used by business researchers is statistics canada’s cansim database. researchers seeking international data also rely on data from international governmental and nongovernmental organizations such as the international monetary fund or the organization for economic cooperation and development. most of these data producers are not-forprofit organizations who offer their data sets to universities at little or no cost. on the other hand, commercial data sources are used to obtain accounting and financial data such as company financials and stock market data. an overview of the core numeric business databases used in canada can be found in table 1. these commercial data producers include data aggregators such as thomson financial’s datastream or standard & poor’s compustat, the university of chicago’s center for research in stock prices (crsp) and financial industry sources such as the toronto stock exchange. many of these commercial data producers are profit-seeking organizations, so while they may offer an academic discount, their primary target market is not the university sector and their data is often very costly to acquire. database name company financials equities fixed income derivatives / futures macro-economics cfmrc (tsx) canada crsp us stocks us compustat north america us canada datastream advance global global global global global note: company financials includes financial statements (balance sheets, income statements) & financial ratios. equities includes stock prices (open, close, return, volume, beta).fixed income includes bonds and treasury bills. derivatives / futures also includes commodities, options & warrants. macroeconomics includes items such as interest rates, exchange, gdp table 1 core numeric business databases, data types, and geographic coverage 10 iassist quarterly spring 2007 trends in canadian business education in this section, we examine the current situation in canadian postsecondary business education and identify some key trends witnessed at our own institutions. an increasing number of research-oriented graduate programs in business are being offered at canadian universities which are driving the growth in demand for business data. these programs include thesis-based master’s degrees in business economics, business administration, and management, and future plans for phd programs. increasingly, canadian business schools are seeking accreditation from the association to advance collegiate schools of business international (aacsb), partly as a means of staying competitive in the national and international marketplace for both students and faculty, and in order to increase the emphasis on research3. there is an increasing focus on the research productivity of business school faculty. this is partly in response to the results of a study conducted by erkut in 2002 to measure the output and input of canadian business school research which found, among others, that the paper output of canadian schools is relatively low and declining, there are significant differences among canadian business schools in research output and impact, and a few ‘stars’ produce most of the impact4. perhaps in response, some canadian business schools have implemented research performance management incentive reward systems to improve the quantity and/or quality of publications5. canadian business schools are also engaged in bidding wars for newly minted accounting and finance phds who do empirical research, as there is intense competition from the private sector in these fields6. in addition, faculty and students are increasingly demanding remote access to business data, as well as ‘one-stopshopping’ research portals such as wharton research data services, where scholars can access data sets from multiple vendors within a uniform interface.7.. what is the current context for postsecondary business education across ontario’s twenty universities? most offer an undergraduate degree in business as well as an mba program. however, fewer than half of these have doctoral degree programs in business administration, which has implications for business faculty and their access to research support and resources. a 1996 study by tompkins, hermanson and hermanson on expectations and resources associated with new finance faculty positions revealed important differences in expectations and resources between three types of institutions: aacsb-accredited doctoral schools, aacsb-accredited non-doctoral schools, and non-aacsb schools8. they found substantial differences in data resource availability, whereby new faculty at doctoral institutions could expect plentiful research resources and excellent database access, compared to new faculty at non-doctoral institutions who could expect ‘reasonable’ research resources, including database access, and new faculty in non-accredited, non-doctoral business schools who could expect moderate to poor research resources and poor database access.9 as noted earlier, data demand in the academic disciplines of finance and accounting can often only be met by commercial data sources with high costs that are often difficult to obtain given the funding situations of the researcher’s parent institution. a recent study explored the effects of database choice on international accounting research.10 in accounting and finance research it is the previous literature, specifically benchmark papers, that set the standard for database choice. in the united states, data comes primarily from compustat and crsp, but the choice is less clear-cut in international accounting research due to factors such as a lack of tradition and a wider choice of databases.11 due to budgetary constraints, universities will only invest in databases where there is an assurance of usage and therefore the type of research conducted by phd students and faculty is driven by database availability. business data issues from the collections perspective librarians involved in developing data collections encounter five main business models: the institutional membership model (e.g. icpsr); the serials continuation model (i.e. by subscription); the one-time payment model; the ad-hoc arrangement; and the lack of a business model (where no process exists to deal with the needs of academic libraries).12 while business data producers typically follow the serials continuation model, many business school faculty mistakenly believe that the one-time payment model is in place, thus resulting in mismatch between funding sources (one-time research grants) and funding needs (ongoing subscriptions). it has been our experience that some faculty who were conditioned to use particular databases for research as doctoral students are now demanding that their institutions acquire these, even if their own institution does not offer a doctoral program in their discipline. in other cases, new hires are being promised new business databases by business school deans who may not realize the financial implications of such promises. many libraries lack a formal collection development policy for numeric data which often limits the level of support for numeric data files.13 ideally, business data collection development decisions are made proactively and collaboratively such that faculty and librarians work together to investigate and select data sets for purchase in order to meet current and/or future teaching and research needs. however, our experience tells that this is not always the case, as many data sets are purchased unilaterally by business faculty without any library involvement. these unilateral decisions may be partly due to a lack of understanding among business faculty of library collection development practices, or may arise in cases where the library is unwilling or unable to commit funds to purchase a particular database. so, how prevalent are numeric business databases in iassist quarterly spring 2007 11 ontario academic libraries? we compared the holdings of four core numeric business databases (cfmrc, compustat, crsp, datastream) across 17 ontario universities by examining library web sites (see table 2). the most frequently held database was cfmrc, followed by compustat. only five universities held all four of these products. interestingly, we found universities with doctorallevel business programs that do not hold any of these four products. while it is hard to know why they do not, it could be related to the nature of their programs (perhaps lacking a focus in accounting or finance), or these data products may be available in the business school, but not listed on the library’s web site. the latter reason, although rare, further complicates the support that librarians are able to provide for business data as the librarians they may not be able to access these data sets from the library. business data issues from the reference service perspective in this section, we continue our exploration of business data by examining issues from the reference service perspective. a recent study by bennett and nicholson that investigated the interaction between business library services and research data services in academic institutions described numerous models for the location and provision of both types of services.14 a snapshot of business librarianship at ontario universities revealed that 13 of 20 universities employed one or more business librarians. of these, five universities (all with doctoral programs in business) had a separate business library. the lack of a standard service model for academic business librarianship leads to inconsistent levels of service across the academic business community. for example, there is no consistency in the hiring practices of librarians, particularly business librarians. unlike faculty, where the student-faculty ratio informs the need for additional faculty positions, increases in student-librarian ratios do not often result in additional librarian positions. this means that in some universities there is no dedicated business librarian, even when the size of the business student population warrants the need. this is a concern, because unlike the humanities or social science disciplines where subject librarians can easily apply their expertise to other disciplines when necessary, business subjects require expert knowledge even more so when librarians are dealing with business data. variations in the presence or absence of business and data librarians have influenced each university’s ability to provide support for business data. our own informal investigation of service models to support the use of business data for teaching and research in ontario found that in some academic institutions it is the business librarian who provides research support, trains new users and troubleshoots problems, while in other academic institutions it is the data librarian, if there is one. oftentimes, there is cooperation between business and data librarians where they exist within the same institutions. however, in some institutions, business database support is not provided by the library; instead, someone with business database expertise, for example a faculty member or researcher, trains new users and acts as a champion for business data. just as university libraries vary greatly in terms of size, staffing levels and budgets, the data needs of their faculty and student clientele also vary. business faculty and doctoral students have similar needs which can be quite complex in contrast with the fairly straightforward needs of mba or undergraduate business students. pros and cons of business data services one might ask, should libraries be providing these types of services? or is it the responsibility of the business school? for example, in the library literature, we have found two completely contrasting views on the compustat database: (1) “[compustat’s] format is more appropriate as an instructional aid and/or laboratory application than as a library reference product”15, (2) “compustat is an indisputably valuable research and education tool”.16 the arguments against libraries providing business data services stem from the basic fact that numeric databases are not a medium that libraries are used to taking care of (compared to bibliographic databases). most librarians (business or otherwise) do not have the background or capability to understand and support numeric business data users. in addition, these databases are costly to subscribe to, difficult to use and time consuming to support, as users can have complex questions that take a long time to resolve. even if a library offers a data service, the data librarian is likely to concentrate on social science data and may not be familiar with business data. finally, it may not be possible to offer good service in particular library environments, as it can be especially challenging for librarians such as ourselves who are ‘solo’ business librarians in a general or central academic library, and who have few, if any, colleagues to provide backup reference support in our absence. that being said, there are good reasons for advocating for libraries to provide these services. according to rebecca smith: “libraries no longer have a monopoly on the provision cfmrc: 12 universities subscribe 5 universities have 4 of these compustat: 9 universities subscribe 3 universities have 3 of these crsp: 8 universities subscribe 4 universities have 3 of these datastream: 8 universities subscribe 5 universities have none of these note: based on a review of the databases holdings of 17 ontario universities with business or management programs as listed on library web sites. table 2 numeric business database holdings in ontario universities 12 iassist quarterly spring 2007 and distribution of information. as we face competition for information stewardship, it is the quality of service we provide that will determine our fate... users can already go around the library for some of their information needs... if users continue to perceive academic librarians as knowledgeable and service orientated the library will remain central to a university’s mission.”17 the business librarian, and the library in general, risks being marginalized by the business school and business data users. business schools that bypass the library are usually the ones that have paid for the data themselves; usually when there is cost-sharing between the library and the business school, the library becomes more involved in providing services. cynthia lenox identified the following success factors for supporting compustat in academic libraries: the degree of involvement of the teaching faculty and incorporation into the curriculum; continuing product improvement and ease of use; and the amount of computer support and training offered.18 academic libraries wishing to provide reference support for business data users face two hurdles: low levels of data and statistical literacy amongst reference librarians coupled with a lack of business subject expertise. data librarians know that the acquisition of statistical knowledge is important in order to support general data users, and have advocated for statistical literacy among librarians.19 data and statistical literacy, and financial literacy in particular, have also been identified as key skills to be acquired by business information professionals, not only within the business information community but from the business school.20 while some information professionals expressed doubts about the legitimacy of their role in promoting and using numeric databases, allen foster advocated not only data retrieval skills, but a broad understanding of the nature of the data retrieved and how to manage that data.21 with these skills, business information professionals can broaden their role to include establishing policies, training users, providing reference services and liaising with vendors.22 conclusion and some suggestions at the very least, academic business librarians are urged to become statistically literate, particularly if their library is not equipped to support numeric business databases through traditional data library services. data librarians are urged to become financially literate, so that they can be capable of working with financial data. that way, if and when the library decides to offer such a service, they will be ready, as numeric business data comes with a steep learning curve. canadian academic business librarians and data librarians must both recognize that the demand for business data services is only going to grow as canadian business schools realize that “becoming a distributor of knowledge generated elsewhere is not an option for canadian business schools; we must strive for research leadership in business just as our fellow researchers do in science and medicine”.23 such changes can be successfully met if business and data librarians embrace this change because, according to marydee ojala, “transformational librarians look at the changes in the information world and take advantage of those changes to enhance their roles”.24 in addition, we advocate for a strong partnership between academic business librarians and data librarians to support business data use. bobray bordelon, speaking on the topic of collaboration between social science librarians and data librarians, suggested that since no one librarian can know enough to meet all the users’ needs, subject and data librarians should form partnerships.25 we believe these new academic business librarian-data librarian partnerships could take place within the same institution, or across organizational boundaries. just as the data liberation initiative supports the use of statistics canada data across canadian universities, a similar ‘community of practice’ could be created to support the use of core numeric business databases. perhaps iassist might be the natural home or sponsor for such a community of practice?26 finally, we advocate for better relationships between commercial business data vendors and academic users, including better academic price discounts, better products, better documentation, provision of usage data, and more training. this has already been advocated in the iassist community with respect to international economic data.27 a critical mass of academic business data users will be needed to convince commercial data vendors to change their business models and their products to better suit the academic market. * contact: eun-ha hong is the business and economics librarian at wilfred laurier university (ehong@wlu. ca); linda lowry is the business and economics librarian at brock university (llowry@brocku. ca). this paper is based on a talk presented at the iassist conference in montreal in may 2007. footnotes 1 cliff halliwell, “desperately seeking data,” horizons 8, no.1 (2005): 31-37. 2 hilary m. davis and john n. vickery, “datasets, a shift in the currency of scholarly communication: implications for library collections and acquisitions,” serials review 33, no.1 (2007): 26-32. 3 margaret mckee, albert j. mills and terrance weatherbee, “institutional field of dreams: exploring the aacsb and the new legitimacy of canadian business schools,” canadian journal of administrative sciences 22, no.4 (2005): 288-300. 4 erhan erkut, “measuring canadian business school iassist quarterly spring 2007 13 research output and impact,” canadian journal of administrative sciences 19, no.2 (2002): 97-123. 5 linda m. manning and jacques barrette, “research performance management in academe,” canadian journal of administrative sciences 22, no.4 (2005): 273-287. 6 gordon pitts, “business academics write their own ticket,” the globe and mail, march 28, 2007: e1. 7 davis and vickery, “datasets, a shift,”:30. 8 james g. tompkins, heather m. hermanson and dana r. hermanson, “expectations and resources associated with new finance faculty positions,” financial practice and education, spring/summer 1996: 54-64. 9 ibid.,:63. 10 juan manual garcia lara, beatriz garcia osma, and belen gill de albernoz noguer, “effects of database choice on international accounting research,” abacus, 42, no.3/4 (2006): 426-454. 11 ibid., 427. 12 davis and vickery, “datasets, a shift,”: 29. 13 william h. walters, “building and maintaining a numeric data collection,” journal of documentation, 55, no. 3 (1999): 271-287. 14 terrence b. bennett and shawn w. nicholson, “interactions between the academic business library and research data services,” portal: libraries and the academy, 4, no.1 (2004): 105-122. 15 alexia strout-dapaz and dennis odom, “intercepting departmental fumbles and running with the ball” (paper presented at the acrl 9th national conference, detroit, mi, april 8-11, 1999), http://www.ala.org/ala/acrl/ acrlevents/strout99.pdf (accessed april 21, 2007). 16 cynthia lenox, “success factors for supporting compustat in academic business libraries,” business & finance bulletin, 109 (1998): 41-49. 17 rebecca smith, “product management: a new skill for reference librarians?” reference & user services quarterly, 39, no.3 (1999): 121-127. 18 lenox, “success factors,”: 46-57. 19 ann gray, “data and statistical literacy for librarians,” iassist quarterly, summer/fall 2004: 24-29. 20 rita marcella, “view from a business school: an interview with professor rita marcella,” business information review, 24, no.1 (2007): 30-35. 21 allan foster, “online numeric databases: four years later,” business information review, 5, no. 3 (1989): 3-12. 22 ibid.,7. 23 erkut, “measuring canadian business school research output and impact”, 119. 24 marydee ojala, “journeys and transformations,” online, september/october 2006: 5. 25 bobray bordelon, “experience-education-interest: a collaborative approach to data reference and interpretation”, presentation at the digital library federation workshop on social science data archives, january 1999, http://www. diglib.org/collections/ssda/ssdaresults.htm (accessed april 29, 2007). 26 for overview of communities of practice, see: etienne wenger, “communities of practice: a brief introduction,” http://www.ewenger.com/theory/ (accessed february 5, 2008). 27 bobray bordelon, “cross-national & intergovernmental data: paying for one-stop shopping,” iassist quarterly, fall 2005: 5-7. 12 iassist quarterly 2014 iassist quarterly iassist quarterlyiassist quarterly abstract while institutions, methodology and geography all present barriers for communication and development of infrastructure, sometimes the greatest barriers may be in reaching not across the world but across the hallway. engaging in the work of unified infrastructure requires finding language that bridges modes of inquiry and meaning, so that all participants see their place in the whole. this work of finding shared language involves translation at many levels. data librarians know that not everyone means the same thing by ‘data’ and increasingly they seek language that spans the practices of social science, sciences, humanities, and performing arts. this paper aims to highlight some of the ways in which data professionals are already adept at translation. drawing on examples from work as a subject librarian and data professional at an undergraduate institution, i will elaborate on ways in which translation permeates the daily work of data librarians, from helping new researchers learn the language and methods of a field, to supporting faculty as they expand their teaching and research across disciplines. additionally, librarians’ role as semi-outsiders within the institution situates them well to help drive conversations spanning disciplinary modes of thinking, in which faculty may also find themselves as semi-outsiders. keywords: data profession, language, translation, data theory, critique introduction in the first of this two-part series, justin joque (page 7: from data to the creation of meaning part 1: unit of analysis as epistemological problem) discussed the ways in which the problem of data harmonization is not just technical but also political, ideological, and infrastructural. in this second part, i would like to dwell further on the ways in which the expertise of data librarians is not just technical but also cultural in the sense that much of their work is about communication, specifically translation. though my approach is further removed from the texts and language of philosophy, it is my hope that i can use and build on the problem that joque articulated. namely, i propose that employing the metaphor of translation to describe the work of data librarians highlights a less obvious aspect of the expertise that they bring to the work of aligning infrastructure and data. the work of the data librarian can be seen as situated at the point where the efficiency of data meets the human work of interpretation, decision-making and communication. as enthusiasts for the potential benefits of making data reusable, librarians are deeply familiar with the ways in which consistent methods and standards open the doors for datasets to become valuable beyond their initially intended use. but in working with patrons who wrestle with other priorities, librarians also know that not everyone from data to the creation of meaning part ii: data librarian as translator by kristin partlo1 ...the work of data librarians involves bridging systems of meaning and acting as translators. iassist quarterly 2014 13 iassist quarterly is willing to make changes to their workflow to employ those methods and standards, or follow good data lifecycle management practices. through their relations with scholars and students across disciplines and levels of expertise with different goals and values, the work of data librarians involves bridging systems of meaning and acting as translators. in this paper i will expand on this idea of translation and its implications through the lens of my work as a data librarian and subject liaison at a small liberal arts college. though my job is idiosyncratic, that quality is shared by many data librarian positions, so it is my hope that there are threads here that will resonate with others in the field. translation and data while the idea of translation may for many readers invoke the google translate tool, anyone who has used it knows its limitations. it is handy for getting the gist of a text in an unfamiliar language, but it cannot fully capture the meaning and nuance of the original text, nor is it reliable enough for much beyond casual use. likewise, anyone who has attempted to travel with a phrasebook or to translate with a dictionary runs immediately into similar problems. though on the simplest level, translation might seem to be a sign for sign replacement, not unlike assigning value labels in a dataset (male is 1, female is 2), it is actually a more complicated process of re-describing from one system of meaning, value, culture and experience to another. the catch is that some or most of the meaning needs to remain intact after the transformation. a recent review in the london review of books aptly demonstrated this complexity while discussing a new translation of finnegan’s wake into chinese: “there’s plenty of finnegans wake that i’d be stumped to put into mandarin. browsing at random: ‘the fall (bababadalgharaghtaka mminarronnkonnbronntonnerronntuonn-thunntrovarrhounaw nskawntoohoohoordenenthurnuk!) of a once wallstrait oldparr is retaled early in bed and later on life down through all christian minstrelsy.’ i’m not sure this is convertible into any language, even an indo-european one, but dai’s translation has been a hit in china, as the western media reported widely at the time of publication.” (yun, 2014) even if it were possible to render this example word for word in another language, there are other things going on in this text that would be lost. the successful translator must be deeply familiar with not just the spoken and written forms of the original and the target languages, but also the culture and history, even, like in the case above, the sounds of the language when spoken and the associations they invoke in a listener or reader. the expertise of a translator comes from experience in both worlds of meaning, of the original and target languages, and the work of a translator involves slogging through decision after decision, interpretation, and awareness that the translation will never capture all of the original. rather than the mechanized replacement of one word for another, or google’s more sophisticated statistical analysis of previous translations2, rich translation requires the work of a human. this work of interpretation and decision-making is messy and, even when it is not error-prone, any translation is imperfect and involves a loss of meaning from the original. yet translation is necessary because, despite the loss, something is also gained, some new meaning or understanding made possible by shepherding an idea or concept from one context to another. because of the inevitable loss, translation requires the arbitration of gain and loss of meaning which, again, requires deep familiarity with both the origin and target contexts of meaning. it is possible in some cases that the loss of meaning is greater than the gain, leading to the conclusion that translation is not possible or desirable. for example, when a data librarian helps a patron dig through documentation to become familiar with a dataset in order to reuse it for their research, they are judging whether it is possible to translate that dataset into the context of the new work. sometimes the decision can be that the data are not a good fit because the loss would be too great to justify the translation. just as textual translation involves fluency of multiple languages and their cultures, data librarians must be familiar with the disciplinary contexts in which data are created and used, the languages, practices, ontologies and classification debates that inform them. working without these fluencies can lead to mistranslations that can set work back or cause librarians to lose credibility with faculty researchers, instructors or students. as appealing as it might sound, librarians know that datasets are not like so many apps in the data archive app store ready to be plugged into any research project. datasets have their own ecosystems of sense and values and rules, and they require documentation in machine and human readable form to allow for informed decision-making about their careful reuse. the ability of researchers to make such decisions depends in part on the work of data librarians, who collect, assure the quality of, and help users interpret that documentation. the more familiar data librarians are with the types of research projects that produce and use sharable data, the better job they will do. just as translation always involves some loss of meaning, data librarians’ efforts to build systems and services are informed by a need to balance gain and loss, measuring efficiencies against decreased ease of use or meaningfulness in particular contexts. translating day to day translation, with its technical and cultural aspects, shares this dual quality with both data and the expertise of data librarians. anyone familiar with the history of the u.s. census knows that data themselves are cultural and political artifacts even as they are created for an analytical purpose. likewise, data librarians, valued for a certain technical expertise, also have a cultural expertise that is present in and built out of their day to day work. data librarians translate between datasets and users, students and their professors’ assignments, metadata and repositories, researchers across disciplines, and librarians and other professionals. in nearly every aspect of their work collecting, describing, teaching, providing reference assistance, building systems and informing campus data management policy librarians work among and between cultures of data use that are distinct with their own languages and worlds of meaning that overlap in some ways and not in others. the daily work of translation is well illustrated by looking at the work of service-oriented roles. providing data reference services involves listening to patrons’ questions and translating what they say into statements of need or inquiry that can either be addressed directly or through referral. the reference interview process involves empathizing with the patron and understanding as much as one can about the context of the question not just trying to take it at face value. this need is then matched in particular to an understanding of the collection and how it is organized as well as more generally to the landscape of scholarly communication and the search tools available. furthermore, in my case working with undergraduates, the question must also be interpreted in light of what i know about the professor’s goals for the assignment. librarians take questions stated in the language of a novice and provide a bridge to the works organized according to the systems of disciplines and experts. the patron might wish that everything 14 iassist quarterly 2014 iassist quarterly were organized according to the logic of their own research topic, or that search engines could be sophisticated enough to anticipate their needs, but given that impossibility, it is clear that a human must be present and ready to help those transactions take place between the language of the patron looking for data and the language of the collection, repository, disciplinary literature, or dataset documentation. libraries as institutions attempt to place the works of all the academic disciplines in one collection, yet in their own languages. librarians who tend these collections and translate their value to scholars and students exist in a place between and among the disciplines. a universalizing conception of the library, like those discussed by joque (see page 7 joque, j. (2014) from data to the creation of meaning part i: unit of analysis as epistemological problem. iassist quarterly [online] 38(2). available from: http:// iassistdata.org/iq/issue/38/2. [accessed: 4 march 2015] ), places the library outside and above the disciplines, organizing them within an overarching ontology. focusing on the work of librarians as translators shifts the focus of the work from crafting the universal system to something more liminal, running through the spaces in between the disciplines. situated in this way, data librarians must always be translating, building technical infrastructure while also building, participating in, and constituting cultural infrastructure.3 by cultural infrastructure i mean the social norms, practices, and expectations in which our systems function and make sense as well as the cast of characters who enact them. viewing the work in this way has implications for how data librarians organize and prioritize their time, form partnerships, develop expertise, and explain the nature of their work to their bosses. easily seen as a disadvantage, existing in a space of imperfect translation also opens up the potential to help frame issues in new ways. librarians, rarely as fluent in any one disciplinary language as the teaching faculty, are at a disadvantage when speaking to faculty in their own disciplinary languages. however, when those same faculty step outside their own home context, for example when doing interdisciplinary work, it becomes easier for them to rely on others and for librarians to offer help. when one knows that one is learning, the expectations are changed and it allows space for imperfect articulation. a barrier of authority and fluency is removed. librarians, who are accustomed to finding themselves in this liminal space, can take advantage of and recognize this inversion as an opportunity to make themselves understood. data librarians can empathize with the uneasy feeling of communicating in a language that is not their first and are positioned well to anticipate how and where they can help. sometimes imperfect translations serve to draw people out of their native language into unfamiliar territory making it easier, when all goes well, to see commonality. for example, as part of a gap assessment, the research data services and support group on my campus wrote a document (2012) articulating the points where students working with quantitative information in any class might run into trouble and seek assistance. it was a simple idea, but was complicated by conflicting uses of terms like ‘analyze,’ ‘collect,’ and ‘data’ in different disciplines. in the end, it was written in imperfect general language, not aligned with any one of the disciplines, but was meaningful enough to trigger wide engagement on an issue that had not gained traction in the past. by finding language specific enough for people to see their experience in it, but general enough to draw people out of their own disciplinary perspectives, it allowed a different kind of conversation take place across disciplines. two other examples illustrate further this idea of accessing an in-between space of meaning. first, being at a teaching college, it is sometimes more fruitful to raise an issue with faculty as a pedagogical question first. the language of teaching and learning is one in which faculty expect to see multiple disciplines reflected in close proximity. i have had better success engaging faculty about how to teach students to manage their data than i have had talking with faculty about their own research or teaching data. by shifting the conversation outside of my primary expertise and outside of the faculty’s research area into a shared second language, we are able to find common ground. in a similar move, in order to engage the topic of data across the curriculum with librarians who do not work with traditionally quantitative fields, i have shared an article by boyd and crawford (2011), “six provocations for big data,” which draws out some of the broader questions about how big data in research are fundamentally changing the ways researchers ask questions across disciplines. boyd and crawford’s language is broad, for example, they pose that big data “reframes key questions about the constitution of knowledge, the processes of research, how we should engage with information, and the nature and the categorization of reality” (boyd & crawford. 2011, p. 3). by broadening the topic and effectively raising the stakes from the relatively narrow concept of big data to issues of epistemological change, we were able together to see the impact of these ideas on all of our areas of expertise. implications of considering data librarians as translators iit is widely recognized that metadata has the greatest chance of being meaningful if it is written in the language of the creator of the data. the creator not only has the most intimate knowledge of the data, but also speaks the language of the discipline or scholarly community in which the project emerged. when researchers speak to each other within their own field, they draw on the literature, they know which terms are contentious and which are clear. furthermore, they are familiar with and can appeal to shared values. for example, stephanie hampton and her co-authors (2013) make what is basically an ecological argument in favor of data stewardship and reuse. through discussion of multiple examples of research that made use of existing data to solve stubborn problems of measurement, she demonstrates how researchers operate not alone but in a system and within an environment of existing data. by framing her argument in this way, she appeals to the professional commitments of ecology, such as reuse and tending to systems, to make a case for sharing data. it would have been difficult for a non-ecologist to be persuasive in this way without this degree of disciplinary fluency. if part of the work of a data librarian is to translate the appeal for good data practices into the disciplinary languages of faculty, then it follows that part of the job is to develop these fluencies. one might attempt to do so directly in areas with affinities to one’s own expertise, or one might turn to subject librarians to get closer, by proxy, to the ideal of speaking fluently across all of the disciplines. both of these options take time, which is easier to justify when the translator role is a visible part of the work. finally, making visible the data librarian’s interpretive work as a translator highlights the data librarian’s teaching role. unlike iassist quarterly 2014 15 iassist quarterly technical solutions that can be set into action and observed from a distance, bringing about cultural change involves educating researchers and teachers, emerging scholars, and other professionals about the value of managing data according to established (and emerging) good practices. it is not enough to present these practices as they have emerged in the social sciences, in their native language and context. instead, data librarians (with subject liaison partners) do the work of making these practices appear relevant by translating them into language that is meaningful in other contexts and workflows and that speaks to the relevant intellectual motivations and values. working in this way is slow going, decentralized, and requires room for failure, miscommunication and mistranslation. a current focus in my own work is bringing knowledge of social sciences data management to the digital humanities initiative on campus, with which i am peripherally involved. in my own liaison areas, where i have confidence that concepts of data management work their way to greater and lesser degrees into methodology instruction, my approach to education and outreach is to complement what the professors are already teaching or aspire to teach their students. in the digital humanities, the humanities librarians and i are working on finding language and metaphors to help scholars see their existing practices with materials, digital or otherwise, as amenable to data management. for example, we have used summarized versions of the data curation profiles interview instruments (carlson 2010) and used them as a discussion exercise in several settings with other librarians and with the undergraduate digital humanities interns to introduce the concepts of data management and reshape them into a meaningful framework for re-applying the model and thinking about what counts as data in the digital humanities. nearly always, the term data gets replaced with something like research materials, but that replacement is not sufficient to make the disciplinary leap. without translation of these concepts in a very concrete manner to questions and considerations familiar to individuals in the humanities, too many people see data management as something that does not apply to the kind of work they do even as their work becomes increasingly digital. infrastructure is something most people don’t see or think about until it breaks down. through their work, data librarians make visible the challenges of aligning infrastructure, both technical and cultural. the work of data stewardship is not a back room problem, but one that is tied up in cultures of research, teaching, and processes of scholarly communications. data librarians engage in the cultural work of translation in many ways, and that skill is a part of their expertise. such expertise is needed in developing data management policy at the institutional level and in the broader culture increasingly interested in big data. professionals with the detailed knowledge of data structures and practices can help translate the value of integrating best practices to those who teach, those who collect data in the field, those who fund the research and the institutions that support it, and those who are learning to become tomorrow’s researchers. acknowledgements i am deeply indebted to my colleagues at gould library, carleton college, who unfailingly and generously offer creative and insightful dialogue and critical editing, especially heather tompkins and iris jastram. i am grateful to my employer and my library for supporting and valuing the work of contributing to the published scholarship in our field and to iassist for providing a forum for it. finally, this paper would not have been written without the prompt from (and the ensuing lively discussion with) justin joque (see page 7 joque, j. (2014) from data to the creation of meaning part i: unit of analysis as epistemological problem. iassist quarterly [online] 38(2). available from: http://iassistdata.org/iq/ issue/38/2. [accessed: 4 march 2015] )to collaborate on a project drawing on our shared background in continental philosophy. references boyd, d. and crawford, k. (2011) six provocations for big data. a decade in internet time: symposium on the dynamics of the internet and society. 21 september 2011. oxford: the oxford internet institute. [online] available from http://papers.ssrn.com/ sol3/papers.cfm?abstract_id=1926431. [accessed: 3 june 2014] carlson, j. (2010) data curation profiles toolkit [online] available from: http://datacurationprofiles.org/. [accessed: 2 june 2014] costa-jussà, m.r. & farrús, m., 2014. statistical machine translation enhancements through linguistic levels: a survey. acm computing surveys, [online] 46(3), pp.42:1–42:28. available from: http://dl.acm. org/citation.cfm?doid=2518130. [accessed: 25 july 2014] hampton, s. e. et al. (2013) big data and the future of ecology. frontiers in ecology and the environment. [online] 11 (3). p. 156-162. available from: http://www.esajournals.org/doi/abs/10.1890/120103. [accessed 3 june 2014] joque, j. (forthcoming) from data to the creation of meaning part i: unit of analysis as epistemological problem. iassist quarterly [online] ? (?). available from: ?. [accessed ?] research data services and support group (2012) 10 points in working with quantitative data when students need or seek support. [online] northfield: carleton college. available from https://apps.carleton.edu/campus/library/assets/10_points_data_ support_13.10.09.pdf [accessed 30 july 2014] yun, s. (2014) short cuts. london review of books. [online] 36 (7). p. 24. available from: http://www.lrb.co.uk/v36/n07/sheng-yun/short-cuts. [accessed 3 june 2014] notes 1. kristin partlo is reference & instruction librarian for social sciences and data at carleton college in minnesota, usa. she can by reached email: kpartlo@carleton.edu. this paper was presented at the 2014 iassist conference in toronto, ontario, canada on 4 june, session 3j, along with its companion paper by justin joque, “from data to the creation of meaning part i: unit of analysis as epistemological problem.” 2. readers interested in learning about statistical machine translation methods used by google and other translation software will find a literature review and useful insights into how linguistics and computer science concepts are used together in this multidisciplinary field in a recent survey by marta costa-jussà and mireia farrús (2014). 3. my colleague, heather tompkins, frequently uses this expression in an analogy about the current state of digital humanities being like finding yourself in a car on a rope bridge. the technical tools may be there to make certain projects possible, but the cultural infrastructure is not yet developed sufficiently to plan well and prevent disasters. 6 iassist quarterly 2016 iassist quarterly abstract this paper reports findings about inter-organizational influence and collaboration relationships among social science data archives over time, focusing on activities of institutions affiliated with the journal international association of social science information services and technology quarterly (iassist quarterly). we examine how archives interacted from 1976-2014 by tracing relationships described in articles published in iassist quarterly. introduction this paper reports findings about inter-organizational influence and collaboration relationships among social science data archives over time, focusing on activities of institutions affiliated with the journal international association of social science information services and technology quarterly (iassist quarterly). we examine how archives interacted from 1976-2014 by tracing relationships described in articles published in iassist quarterly. . keywords social science data archives, history, iassist quarterly, social network analysis, collaboration introduction research disciplines increasingly rely on data archives and repositories to share results, advance their work, and support large-scale collaboration. while there have been numerous studies that examine the technologies and practices of how research fields and researchers develop and use data archiving and archives, there has been less attention paid to how data archives themselves as information institutions have adapted over time to evolving research trends, institutional changes, and funding models. social science data archives (ssda) are exemplars of long-lived information infrastructures (broadly defined as the computing and technological resources and their supporting institutions that are designed to advance scientific inquiry) that have successfully adapted to such changes (heim, 1980; o’neill adams, 2006). understanding inter-organizational relationships among ssda, funders and partner institutions over time is essential to understanding how ssda have evolved to serve their user communities. in this paper, we examine how such relationships have evolved, focusing on activities of institutions that appear in articles published in iassist quarterly (iq) from 1976-2014. our larger goal is to better understand how ssda have cooperated and competed to achieve their goals and draw out lessons learned that can be applied to the development and maintenance of contemporary cyberinfrastructures for research in the social sciences. we hope that lessons learned by ssda may be useful to similar infrastructure projects in other fields. this paper explores the following research questions: 1 which institutions are most influential as depicted in iq articles? 2 which institutions collaborate the most as depicted in iq articles? 3 to what degree are international relationships represented in iq articles? 4 to what extent are highly collaborative institutions collaborating with each other? methods social network analysis (sna) is a data collection and analysis approach useful for examining patterns in connections between people or social institutions. the results of sna analysis are often depicted as a web of ‘nodes’ (people or institutions) and relationships or ‘links’ which represent the links between nodes. as hansen, schneiderman and smith (2011) describe, sna analysis may trace: a. the number of unique links connected to a node. nodes that have more links connected them to other nodes may be more important or influential. social science data archives: a historical social network analysis by kristin r. eschenfelder, morgaine gilchrist scott, kalpana shankar, greg downey1 iassist quarterly 2016 7 iassist quarterly b. changes in the patterns of connections between nodes. c. variation in the types of links or between nodes. to answer our research questions, we use sna to analyze links between organizations or institutions (the nodes or our social network) as represented in iassist quarterly (iq) articles from 1976 through the end of 2014. we obtained back-issues of iq from the iassist website, from archive.org or from the university of wisconsin-madison library. we did not include papers from the annual iassist conference proceedings because we only had access to the full text of conference presentations after 2000 (http://www.iassistdata.org/conferences). we first identified all articles by issue according to the volume number and date listed on the bottom of the article. for each paper we identified the following nodes in each paper: • institutional home of author or co-author of iq article, • institutions mentioned in the paper as influencers or collaborators (more on this below) by iq authors • funders of projects described by iq authors. we identified the following major types of relationships, or links, as explicitly described in the papers: 1 influencing relationship: when one node mentioned another node as being influential, including co-authorship when the authors were at different institutions. social network analysis typically refers to nodes that have more attached links as having a higher ‘degree of influence’ because a link between institutions represent opportunities for influence or evidence of influence. we examined the iassist networks to see which institutions were the most connected to other institutions. arguably these linkages represent the potential for one institution to influence the practice of the other institution through sharing of knowledge or resources. 2 collaborating institution relationship: when one node described a collaboration with another node. (see below) 3 funding relationship: when a node provided funding (explicitly stated in the article). funding relationships are shown in green on network graphs. 4 international relationship: we specifically noted data provider or collaborating institution relationships that crossed national boundaries. because the consortium of european social science data archives (cessda) is a pan-national organization, any relationship with cessda was marked as international. we tracked influencing, collaborating, funding and international relationships among nodes in iq using the tool nodexl, an open source template for conducting sna with data from microsoft excel (nodexl 2015). we also tracked change over time of all of the four types of relationships. in order to show change over time, it is a common practice in longitudinal social network analysis papers to analyze data in multiyear sections rather than year by year. data within a year often do not provide sufficient nodes and links to show a network of relationships. at the same time, analyzing the data as one large set (1976-2014) cannot show change. we determined that five year sections were sufficient to show both network relationships and change over time. operationalizing influence and collaboration we coded the iq articles in two ways: first broadly for influence, and then more narrowly for collaboration. collaboration relationships are a narrower subset of influence relationships. influence: first, we coded broadly for relationships indicating influence among data archives. for the purpose of this project influence between nodes was included by one node’s reference to another node as: • an inspiration or model, • part of a larger grouping of affiliated organizations, • part of a collaboration, • a provider of data, • a provider of resources (staff, software), or • a co-author on an iq paper from a different institution as we describe below, and as we depicted in our (eschenfelder et al., 2015) iassist 2015 conference paper, the influence analysis resulted in a large loose network of relationships in the iassist community. we argue that influence patterns show which data archives other archives talk about, or which data archives are most influential in the larger community. collaboration: in this narrower analysis, we examined a subset of influence relationships describing collaboration between data archives. to do so, we first developed a definition of collaboration and coding rules by conducting comparative analysis of three random samples of approximately ten iq articles. from the analysis, the team identified the following types of collaboration relationships: • new collection/service: institution y gets data from institution x when institution y creates a new data collection or service. or x allows y to make x’s data accessible through a portal. • new entity/project: institutions x and y create a new entity ‘project z’ related to data. • cross national surveys: institution x describes participating in a cross national survey. or a cross national survey project describes getting participation from nations x, y and z. • software/processing: institution x gets software from institution y and creates a new data product or service. or, institution x uses institution y’s processing power to manipulate or manage data to provide a service. 8 iassist quarterly 2016 iassist quarterly • collaborate to collect data: institutions y and y collaborate to obtain data from individuals (pis’ projects or from research subjects themselves). • research about practice: institution y collects data from other data archives in order to compare practices and reports results. • mergers: reports on acquisition of another data archive or its collection. taking over a previous archive. • learning materials: institutions x, y and z collaborate to create learning materials. in order for the description of an influence relationship in iq to count as a collaboration, the depiction also needed to meet significance criteria. to count as significant, a description of a collaboration had to: have a header, have its own paragraph, be visually separated from other text, be mentioned multiple times within the article, or have at least 4 lines of text devoted to it. iq articles included many descriptions of influence relationships that we did not count as collaboration including: x and y co-authoring the iq paper; providing curated links to data that lives at other institutions; compiling, indexing and coding data created through surveys run by others; simple descriptions of data deposit; descriptions of institution x getting data from institution y and x writing a paper; mere descriptions of data available at institution x (even if the data is from other institutions); collaboration of subunits within the same larger organization; aspirational or planned collaborative activities; x describing how they contract out services to y; and descriptions of teaching activities. coding data for influence and collaboration the research team first coded the relevant articles to identify influence relationships. a team member read through each article included in iq, identified the names of nodes, and entered information about co-author, influence and funder relationships between nodes into an excel template. any given node typically had several links (influence relationships) and each link was listed in a separate row. to code articles for the collaboration code, three members of the research team co-read and co-coded five separate samples of iq influence articles from across different time periods in order to ensure we had agreement on the application of the coding rules. one team member then coded the remainder of the iq articles and entered information about collaborations into a new excel template. each collaboration relationship was listed in a separate row. institutions and relationships: 1976-2014 figure 1 and table 1 summarize data on the types of relationships reported in iq articles from 1976-2014 including influence, collaborator, funder and international relationships. the number of nodes, or ssda appearing in relationships grew over time from 12 in the 19761980 time period to 111 in the 2011-2014 period. the number of all types of relationships also grew over time, but not steadily. in particular, the number of funder relationships has been very up and down over the years. moreover, the number of collaborator relationships reported in iq grew steadily, but then dropped in the 2006-2010 reporting period. most nodes (or ssda) in the data have a very small number of relationships. a small number of ssda have a higher number of relationships. in this paper we focus on those ssda with the higher number of relationships. figure 1: types of relationships over time iassist quarterly 2016 9 iassist quarterly time period nodes # influence links # collaborator links # funder links overall 1976-2014 653 285 196 145 1976-1980 12 13 5 0 1981-1985 37 27 10 6 1986-1990 81 69 26 12 1991-1995 76 54 21 17 1996-2000 108 77 37 25 2001-2005 130 119 53 47 2006-2010 98 81 13 15 2011-2014 111 130 31 23 table 1: relationships over time which ssda has had the most influence? we examined the data for which ssda had the most influence. influence links included when one node referred to another node in an iq article as an influence or inspiration, as part of a larger grouping of affiliated organizations, as part of a collaboration, as a provider of data, or a provider of resources (staff, software). one problem is that many influence relationships in our data stem from large projects with many partners. we call these instances ‘projects.’ these projects are depicted in only one or two iq articles, but include a high number of relationships. overall 1976-2014: ssda with the highest number of influence links to other ssda # of influence links # articles in which the institution is coded influence ratio: articles/ links links/articles inter-university consortium for political and social research (icpsr) 30 22 0.73 1.36 zentralarchiv fur empirische sozialforschung (za) 16 9 0.56 1.78 ukda 32 17 0.53 1.88 university of edinburgh 13 6 0.46 2.17 international social survey program (issp) 21 7 0.33 3 us federal reserve 14 4 0.29 3.5 university of minnesota 15 4 0.27 3.75 pennsylvania state university 19 2 0.11 9.5 the pew forum 19 2 0.11 9.5 table 2: 1976-2014 nodes with most influence links we created a measure called ‘influence ratio’ to depict those institutions with the most influence. this measure helped us identify institutions with many influence relationships over time, in addition to those with many influence relationships. the influence ratio captures this measure of an institution’s influence across different articles. the ratio divides the number of total articles in which institution x appears by the total number of influence links between institution x and others. table 2 reports the institutions with the 10 iassist quarterly 2016 iassist quarterly highest influence ratio from 1976-2004. large data archives appear at the top with icpsr leading with a ratio of .73, followed by the za and ukda with ratios of .56 and .53 respectively. which ssda has collaborated the most? we examined the data to see which ssda has the most collaborative relationships. collaborative relationships were a narrower subset of influence relationships that involved a specific set of relationship types including: creating new collections or services, new projects, cross national surveys, sharing of software or processing, new data collections, comparisons of practices or mergers. because we were interested in reporting on institutions with many collaborative relationships over time, we focus our analysis on those institutions depicted in more than one iq article. table 3 lists the institutions with the most collaborative relationships in terms of the same article/links ratio used above. again, large data archives appear at the top of the list with icpsr having a collaboration ratio of 1 and ukda having a ratio of .85. the us census bureau had the next highest ratio with .43. overall 1976-2014: institution with the highest number of links to other nodes (c) # of collaboration links (coding 2) number of iq articles in which institution is mentioned collaboration ratio (articles/ links) icpsr 6 6 1 ukda 13 11 .85 us census bureau 7 3 .43 international social survey program (issp) 7 2 .29 integrated library and surveydata extraction service (ilses) 11 2 .19 east asian business and development (eabad) archive 7 2 .29 table 3: institutions with the most collaborating relationships across multiple iq articles 1976-2014 who has funded whom? we examined funding relationships and nodes to determine which funders had the most relationships from 1976-2014. table 4 below shows the funding nodes with the most relationships in the 1976-2014 period. jisc was the top funding node with 18 reported relationships, followed by nsf and escr with 12 relationships each. funder number of funding links 1976-2014 number of iq articles in which funder is mentioned joint information systems committee (jisc) 18 14 national science foundation 12 7 economic and social research council (esrc) 12 6 higher education support project (hesp) 3 2 national institute on aging 3 2 library of california 6 1 national institute of child health and human development 4 1 california department of finance 3 1 iassist quarterly 2016 11 iassist quarterly california digital library 3 1 juan march institute, spain 3 1 us bureau of the census 3 1 table 4: top funding nodes table 4 shows several instances of projects with many funding links that are only reported in one article (e.g., library of california). these instances represent iq articles that describe projects in which a funder funded a project with multiple partner nodes. in the next section, we continue by providing more detailed data on those nodes with the most influence and collaboration relationships during each of eight five year time periods of our study. we also report on international relationships during each period. period 1: 1976-1980 the early time period of 1976-1980 saw the lowest number of nodes (n=12) and the lowest number of influence and collaboration links (n=13, n=5). at this stage in iq’s history, most articles tended to simply describe activities at an author’s institution, and few described cooperative activities. influence: roper had the highest number of influence relationships (6), followed by university of iowa (4) and then yale, williams college and the university of connecticut (all of which were affiliated with roper center – 3 each). collaboration: in the narrower collaboration measure, only university of iowa and the roper center showed more than one collaboration link during this period. 1976-1980: institutions with the highest number of relationships influence collaboration most relationships roper (6) (2 each) roper center; 2nd most relationships university of iowa (4) 3rd most relationships yale university; williams college; university of connecticut (3) table 5: 1976-1980 number and type of relationships international: iq articles did not depict any international collaborations during the 1976-1989 period. period 2: 1981-1985 the second time period saw a growth in nodes (n=37) and influence and collaboration relationships (n=27, n=10). this period also saw a rise in international relationships. influence: in this period, icpsr and the us census bureau had the highest number of influence links (6) followed by rand corporation (4). collaborations: iq articles still reported very few collaboration relationships. four institutions reported two collaborations each: center for human resource research, indian council of social science research (icssr), the international federation of data organizations for the social sciences (ifdo), norwegian social science data services, zentralarchiv fur empirische sozialforschung (za). 1981-1985: institutions with the highest number of relationships influence collaboration top most relationships icpsr; us census bureau (6 each) (2 each) icssr, ifdo, nsd, za 2nd most relationships rand (4) 3rd most relationships (all with 3) (norc); nsd, nara; us doe; bonneville table 6: 1981-1985 number and type of relationships international collaborations: this period saw three international collaborations. both involved partners in europe or partnerships with international organizations such as ifdo. 12 iassist quarterly 2016 iassist quarterly year collaboration relationship 1981 norwegian social science data service and european consortium on political research (ecpr) 1985 two projects between zentralarchiv fur empirische sozialforschung and the international federation of data organizations for the social sciences (ifdo) table 7: 1981-1985 international collaborations period 3: 1986-1990 the third period showed growing interactivity among ssda. nodes increased to 81. influence relationships increased to 69 and collaboration relationships grew to 26. influence: the international social survey programme (issp) had the highest number of influence links (13) followed by australian national university (6). collaboration: the east asian business and development (eabad) archive had the most collaboration links (5) stemming from a large multi-institution project. the us census had 4 reported collaborations in this period. 1986-1990: institutions with the highest number of relationships influence collaboration top most relationships international social survey program (issp) (13) east asian business and development (eabad) (5) 2nd most relationships australian national university (6) us census bureau (4) 3rd most relationships (5 each)national opinion research center (norc); east asian business and development research archive (eabad) (3 each) ukda; tarki (hungarian social science information center) 4th most relationships (4 each) icpsr; university of amsterdam; university of alberta; office of population census and surveys uk; hunter college cuny; university of mannheim table 8: 1986-1990 number and type of relationships international: this period’s iq articles described an international collaboration between celade in chile and the icrc in canada in 1989. in 1990 the east asian business and development (eabad) archive at uc davis reported on a collaborative project involving numerous partners in asian nations. the ecpr, an international scholarly political science association located in essex uk reported collaborations with two european universities. year international collaboration relationship 1989 united nations latin american demographic center(celade) in chile and international development research center of canada (icrc) iassist quarterly 2016 13 iassist quarterly 1990 east asian business and development (eabad) archive at uc davis and a series of partners including university of hong kong, national university of singapore, tunghai university in taiwan and the china credit information service ecpr (uk) and both university of mannheim and university of amsterdam table 9: 1986-1990 international collaborations period 4: 1991-1995 the fourth time period saw a decline in reported activity among ssda. the number of nodes fell from 81 to 76, the number of influence links fell from 69 to 54, and the number of collaborations fell from 26 to 21. influence: university of manchester and the manchester computing center had the highest number of links in this period (10). collaboration: university of california davis had the most collaboration links in iq (5). 1991-1995: institutions with the highest number of relationships influence collaboration top most relationships university of manchester and manchester computing center (10) uc davis (5) 2nd most relationships university of wisconsin-madison (8); (2 each) bringham university suny; east asian business and development (eabad); icpsr; lehman college; roads; us census bureau 3rd most relationships university of illinois urbana campaign (7) 4th most relationships (5 each) us census bureau; university of california davis 5th most relationships university of missouri st louis (4) table 10: 1991-1995 number and type of relationships international: this period’s iq articles described only one international collaboration between the university of ulster and the united nations university (a think tank and post graduate education institution associated with the united nations and located in japan). year international collaboration relationship 1995 university of ulster and the united nations university (japan) table 11: 1991-1995 international collaborations period 5: 1996-2000 activity grew again during the 1996 to 2000 time period. nodes grew to 108, influence links grew to 77 and collaborations grew to 37. influence: in this period, the ukda had 11 influence links and the za had 7 links. collaborations: the integrated library and survey-data extraction service (ilses) appeared as the largest collaborator during this period with eleven links. the data liberation initiative followed with seven links. 14 iassist quarterly 2016 iassist quarterly 1996-2000: institutions with the highest number of relationships influence collaboration top most relationships ukda (11) ilses (11) integrated library and survey-data extraction service (ilses) (11) 2nd most relationships zentralarchiv fur emprische sozialforschung (za) (7) data liberation initiative (7) 3rd most relationships (all with 6) university of minnesota; conference of rectors and principals of quebec universities issp (5) 4th most relationships (5 each) university of quebec; sociometrics corporation royal statistical society (4) table 12: 1996-2000 number and type of relationships international : international collaborative relationships described in this period include three descriptions of integrated library and survey-data extraction service projects with multiple european partners in 1997, 1998 and 2000 year international collaboration relationship 1997 integrated library and survey-data extraction service (ilses) and partners including progamma, groningen, swidoc in amsterdam, university of amsterdam, zentralarchiv fur empirische sozialforschung (za), trinity college and cidsp in grenoble 1998 integrated library and survey-data extraction service (ilses) and partners including progamma, groningen, zentralarchiv fur empirische sozialforschung (za), university of amsterdam, bsdp, trinity college 2000 international social survey program (issp) and partners including national opingion research center (norc), zentrum fur umffragen, methoden und analysen (zuma), zentralarchiv fur empirische sozialforschung (za), national center for social research, research school of social sciences table 13: 1996-2000 international collaborations period 6: 2001-2005 the sixth period saw growth in reported activity as nodes rose to 130, influence relationships rose to 119, and collaborations rose to 53. influence: penn state and pew had the highest number of influence links with 18 each. collaboration: the association of religious data archives had 20 reported collaborations in this period. 2001-2005: institutions with the highest number of relationships influence collaboration top most relationships (18 each) pennsylvania state university; the pew forum; also association of religion data archives (arda) (20) association of religion data archives (arda) (20) iassist quarterly 2016 15 iassist quarterly 2nd most relationships ukda (11) (7 each) sociological data archive (sda) the institute of sociology in prague; ukda 3rd most relationships (10 each) deakin university; university of ballarat collection of census data and resources (uk) (5) 4th most relationships czech academy of the sciences (9) 5th most relationships (7 each) university of manchester; princeton university table 14: 2001-2005 number and type of relationships international: iq articles in this period described more examples of international collaborations than previous periods. they described two cross-atlantic collaborations involving archives in the us and european partners. year international collaboration relationship 2001 sociological data archive (sda) at the institute of sociology in prague and nine other european collaborators. institute for quality of life research bucharest (iqlr), ukda and za. 2002 collection of census data and resources (uk) and the frauhofer institut autonome intellligente systeme (germany) 2003 finish social science data archive (finland) and icpsr (usa) 2005 association of religion data archives (usa) and 20 other partners, mostly based in us but some european table 15: 2001-2005 international collaborations period 7: 2006-2010 the seventh time period was another period of decline. articles only mentioned 98 nodes, and 81 influence links. and we saw a large decline in collaboration links described from 53 in the earlier period to merely 13 in this period. influence: the finnish social science data archive and the university of ljubjana showed the most influence links with 10 and 9 links respectively. collaboration: university of edinburgh, university of oxford, university of south hampton and london school of economics all showed three collaborations each. 2006-2010: institutions with the highest number of relationships influence collaboration top most relationships finnish social science data archive (10) (3 each) university of edinburgh; university of oxford; university of south hampton; london school of economics 2nd most relationships university of ljubjana (9) 3rd most relationships university of minnesota (7) 4th most relationships (6 each) university of windsor; university of albany; university of cape town 16 iassist quarterly 2016 iassist quarterly 5th most relationships (5 each) ukda; us federal reserve; mit table 16: 2006-2010 number and type of relationships international: the number of international collaborations described in iq also dropped in 2006-2010. we found only one: between the african association of statistical data archivists and the international household survey network in 2007 year international collaboration relationship 2007 african association of statistical data archivists (aasda) and the international household survey network (ihsn) table 17: 2006-2010 international collaborations period 8: 2011-2014 growth increased in the final analysis period with 111 nodes, 130 influence relationships and 31 collaborations. influence: the lithuanian center for social research and vilnius university scored highest in terms of number of influence links with 24 each. collaboration: the node with the highest number of collaborators was the nestor memorandum of understanding (13). this stems from one large project with many european collaborators. 2010-2014: institutions with the highest number of relationships influence collaboration top most relationships (24 each) lithuanian center for social research; vilnius university nestor memorandum of understanding (13) 2nd most relationships (11 each) grottingen state university; cologne university of applied sciences (3 each) american national election study (anes); lithuanian humanities and social science data archive (lida); voices project/equalan; wienmer institute for social science data documentation and methods (wisdom) 3rd most relationships ukda (9) 4th most relationships (8 each) us federal reserve; wienmer institute for social science data documentation and methods (wisdom) 5th most relationships council of european social science data archives (cessda) (7) table 18: 2011-2014 number and type of relationships international: this period’s iq articles described three major international collaborations. two were largely european, but one stretched between taiwan and norway. iassist quarterly 2016 17 iassist quarterly year international collaboration relationship 2011 voices project/equalan and 3 european partners, lithuanian humanities and social science data archive (lida) and 2 european partners 2012 center for survey research, taiwan (srda) and the norwegian social science data services nestor memorandum of understanding among 13 european partners table 19: 2011-2014 international collaborations summary discussion in this section we answer the paper’s research questions by examining our data over time. research question 1: who is the most influential ssda as depicted in iq? the most influential ssda (according to our measure of influence) are icpsr, the za, ukda, and the university of edinburgh. the issp project was also very influential but not at the level of the other four. in other words, within the iq environment, these nodes were most often mentioned, across different articles, as having an influence via collaboration, data sharing, or just serving as an example for others. overall 1976-2014: ssda with the highest number of influence links to other ssda influence ratio: articles/links inter-university consortium for political and social research (icpsr) 0.73 zentralarchiv fur empirische sozialforschung (za) 0.56 ukda 0.53 university of edinburgh 0.46 international social survey program (issp) 0.33 us federal reserve 0.29 university of minnesota 0.27 pennsylvania state university 0.11 the pew forum 0.11 table 20: nodes with the most influence relationships research question 2: who collaborates the most as depicted in iq? the ssda with the most collaborative relationships according to our measure is icpsr, with ukda close behind it. the us census bureau also appears prominently. this means that as depicted in iq articles, icpsr and ukda had the most relationships that met our criteria for collaboration described earlier. overall 1976-2014: institution with the highest number of links to other nodes (c) collaboration ratio (articles/links) icpsr 1 18 iassist quarterly 2016 iassist quarterly ukda .85 us census bureau .43 international social survey program (issp) .29 integrated library and survey-data extraction service (ilses) .19 east asian business and development (eabad) archive .29 table 21: most collaborative ssda research question 3: to what degree are international collaborations represented in iq articles? as shown in table 22, the number of international collaborative links varies widely across the time periods of the study. international collaborative relationships have grown over time but seen two periods of decline (1991-1995 and 2006-2010). time period # all collaboration links # international collaboration links ratio intl/all 1976-1980 5 0 0 1981-1985 10 3 .3 1986-1990 26 8 .31 1991-1995 21 4 .19 1996-2000 37 16 .43 2001-2005 53 30 .57 2006-2010 13 1 .07 2011-2014 31 18 .58 table 22: collaborations and international collaborations over time to examine whether the proportion of international collaborative relationships to overall collaborative relationships has grown in iq articles, we plotted a ratio of international collaborations over all collaborations by time period. collaboration downturns: the results shown in figure 2 suggest that the proportion of international collaborations is growing over time, with periods of downturn (e.g., 1991-1995, 2006-2010). this pattern of growth and dips in growth parallel the overall pattern of collaboration activity shown in figure 1 and table 22 above. other kinds of analysis might yield more insight into whether these collaboration downturns were due to financial pressures or other factors; for example as economies go into recession, or governments cut funding on education fewer resources may be available to support travel to collaborate with other ssda. alternatively, ssda staff may be more under resourced and have less time to write up articles describing collaborations to iq. research question 4: to what extent are highly linked institutions interlinked with each other? we took the institutions with the most relationships from table 21 and looked to see to what degree they had collaboration relationships with other ‘most collaborative’ nodes versus collaborative relationships with other nodes. we found that the ‘most collaborative’ nodes did not collaborate with each other; rather, they collaborated with other nodes with lower collaboration scores. this suggests that highly connected nodes are mostly serving as hubs, linking outwards toward less connected nodes. conclusions the data generated by this analysis provide one limited view of the relationships among ssda and related institutions. there was partial overlap between most influential and most collaborative ssda. our analysis show that the most influential ssda include icpsr, za, ukda and university of edinburgh. the most collaborative ssda were icpsr and the ukda followed by the us census bureau. our analysis also found that the top ranked collaborative ssda did not collaborate with each other but with other institutions. top iassist quarterly 2016 19 iassist quarterly figure 2: collaborations and international collaborations over time funders including jisc, snf and esrc. depiction of influence relationships grew most steadily in iq articles. collaboration and funding relationships grew over time, but showed periods decline in 1991-1995 and 2006-2010. because the data set stems from articles published in iq, the data cannot represent the many collaborations that were never described in iq. for example, renewing a contract with a major partner may be very important, but because it is not a new project, it may not merit a write-up in iq. because iq is the publication of record of the iassist organization, one can argue that the most important and innovative relationships would likely be included because iassist members would seek to share information about their projects with other members. secondly, while the data set of relationships is incomplete, it is still representative of the most innovative and noteworthy relationships in the study period. future research could enrich understanding of influence and collaboration relationships between ssda by adding data from iassist annual meeting conference proceedings. at this time, only data from 2000 and later is available via the iassist website. it is much more difficult to draw conclusions about lack of collaboration or competition from this data sets or reasons why relationships might have ended. to learn more about these issues, we plan to rely more on historical documentation from case studies of individual ssda. ssda are complex institutions with a long history (heim, 1980; o’neill adams, 2006). this iq analysis provides one view of how ssda have developed and maintained relationships with their funders, and each other over time to innovate, serve their user bases, and grow their products and services. this paper has described and summarized inter organizational relationships among ssda and funders from 1976-2014 as depicted in articles published in iassist quarterly. as one component of a larger study about the history of social science data archives (ssda) and the field of social science data archiving, the data provides an important broad view of relationships among data archives in the shaping of a research discipline. we would suggest that this study, even though it is partial in scope, represents the kinds of institutional analyses that could yield insight for other kinds of data repositories and institutions in how they are developing and leveraging their own institutional networks and the implications such networks have for growth and maintenance. references eschenfelder, k.r.; gilchrist scott, m.; shankar, k.; leclere, e.; lin, r.; downey, g. 2015. social science data archives: a historical social network analysis’ iassist conference 2015, minneapolis mn. hansen, derek l.; shneiderman, b.; smith, marc a., 2011. analyzing social media networks with nodexl insights from a connected world. burlington, mass: morgan kaufmann. heim (mccook), k., 1980. social science data archives: a user study. unpublished dissertation, university of wisconsin-madison. nodexl, 2015. nodexl: network overview, discovery and exploration for excel [online] available at: < http://nodexl.codeplex.com/> [accessed may 2015]. o’neill adams, m., 2006. the origins and early years of iassist. iassist quarterly, vol 30 fall issue, pp 5-14. acknowledgements this project has been supported by the alfred p. sloan foundation, the university of wisconsin alumni research foundation, university college dublin seed funding scheme, and the american society for information science and technology history fund. notes 1. corresponding author is professor kristin r. eschenfelder, school of library and information studies, university of wisconsin-madison, eschenfelder@wisc.edu 2. as explained by iq editors, the year dates of publication for an article are not always precise because publishing is sometimes behind schedule. analysis by five year periods may therefore be a better depiction of trends than analysis by specific year. 3. where there were more than two authors, all authors’ institutions were listed as collaborating with one-another. when one author was associated with multiple institutions, each institution was listed as having all the relationships referenced in the article. however, they were not listed as collaborating with each other. oais model 6 iassist quarterly winter 2011 iassist quarterlyiassist quarterly abstract given the significance of the role of data in research and the value of data for long-term use, researchers have been discussing the need for archiving and curating research data for future studies. to make data reusable, managing data in a reliable way and making them understandable to users is significant. this paper examines the current requirements for depositing data in selected data repositories by analyzing the forms and guidelines for such deposits. the open archival information system (oais) is used as a framework for examining current requirements. examining current data deposit requirements provides an opportunity to validate current data collection and management practices and provides insights into ways to improve such practices. . keywords: social science data repository, data deposit, depositor requirements, ingest, oais model. introduction the definition of “data” varies by discipline, and data can come in various formats and types. the national research council (1999) defines data as “facts, numbers, letters, and symbols that describe an object, idea, condition, situation, or other factors” (p. 15). the national science board (2005) uses the term “data” to refer to “any information…including text, numbers, images, video or movies, audio, software, algorithms, equations, animations, models, simulations, etc.” (p. 13). the national science foundation classifies data into four types: (1) observational data (e.g., weather measurements and attitude surveys); (2) computational data (e.g., results from computer models and simulations); (3) experimental data (e.g., results from laboratory studies); and (4) records (e.g., from government, business, and public and private life) (borgman, 2010, p. 19). given the importance of the role of data in research and the value of data for long-term use, researchers have been discussing the need for archiving and curating research data for future studies. curating data (1) enables reuse of data for new research and new science; (2) enables retention of unique data that are impossible to recreate; (3) makes more data available for research projects; (4) enhances the ability to validate research results; (5) promotes the use of data in teaching; and (6) should be done for the public good. that data should be shared is almost universally agreed upon (faniel and zimmerman, 2011). research data need to be available for use beyond the purposes for which they were initially collected, to make the results of studies using publicly funded data available to the public, to enable others to ask new questions of extant data and advance solutions for complex human problems, to advance the state of science, to reproduce research, and to expand the instruments and products of research to new communities (borgman, 2010; hey and trefethen, 2003; hey, tansley and tolle, 2009). despite the potential benefits of data reuse, controversies surround data sharing practices. some argue over the ethics of sharing data and the methodological reasons examination of data deposit practices in repositories with the oais model social science context by ayoung yoon1 and helen tibbo2 the oais reference model became an iso standard in 2003 (iso 14721:2003) iassist quarterly winter 2011 7 iassist quarterly for not allowing it (carlson and anderson, 2007, p. 636). others raise questions about how data collected or constructed by one researcher can be trusted or even understood by another, as data reuse generates a disconnection of the data from the people they represent, as well as from the researchers who collect them. thus, to fill the gap generated by this disconnection and to make data reuse a common practice in scholarly communities, an explicit context for the production and establishment of appropriate systems for quality checks and assessments is essential (carlson and anderson, 2007, pp. 643-644). tthis paper aims to understand the current requirements for depositing data in data repositories by analyzing the forms and guidelines for such deposits. the moment of deposit in repositories is key for trustworthy data management and long-term preservation. what is deposited in repositories is referred to as the submission information package (sip) in the reference model of an open archival information system (oais), which is the first step in a data management cycle within the repository setting. examining current data deposit requirements provides an opportunity to validate current data collection and management practices and provides insights into ways to improve such practices. data deposits and the role of sip in the oais reference model for data curation the keys to data curation are documenting, referencing, and indexing data with long-term value, enabling others to find and use them easily, accurately, and appropriately (national academy of science, 2009, p. 7). because data without any a≠ccompanying necessary information concerning how and within what context they were created can be useless, all data should be well documented, associated with related materials, and linked to publications or other subsequent materials. annotation is also significant in data curation to document changes that occur over time, allowing data to retain their long-term value (lord and macdonald, 2003, p. 45). for these actions to occur for curation purposes, data must be placed in a repository (lord and macdonald, 2003). thus, an administrative framework must be developed that can provide mechanisms or channels for data deposit. the oais reference model, which became an iso standard in 2003 (iso 14721:2003), provides procedures and requirements for data when they are deposited in repositories and is useful for managing any type of digital object in a “trusted” way. the oais reference model provides a framework that outlines archival concepts for long-term preservation and access, as well as relevant presentation information on digital objects (ccsds, 2002). in the oais reference model, data from a producer3 or creator packaged for deposit are referred to as a submission information package (sip). within oais, sips are transformed into one or more archival information packages (aip) for preservation. aips are comprised of content information4 and the associated preservation description information (pdi).5 later, information from one or more aips becomes part of a dissemination information package (dip), which is the information package sent to the consumer in response to a request to the oais, enabling consumers to find and order the content information they are interested in (see figure 1, ccsda, 2002). each information package (sip, aip, and dip) has its own role and significance in oais for long-term preservation and access. the implementation of the aip can vary depending on the archives, but all required information contained in the aip is essential for long-term preservation and access and to ensure that archival holdings remain valid. considering the exact information content of the sip and dip and their relationship to the corresponding aip, all relationships and procedures depend on agreements between archives, information producers, and consumers (ccsds, 2002, p. 4-33). however, performing all necessary transformations of information is difficult without attaining proper sips, since sips provide a complete set of content information and associated pdis to form an aip, thus defining the fundamental significance of sips. thus, the interaction between a producer (or a depositor) and repositories is particularly critical during the process of acquiring information for a sip. ross and mchugh (2006) discuss the significance of the depositors’ role in this process as well as the interaction between depositors (producers) and repositories. they insist that “depositors will be able to verify whether they are adequately informed when processes are completed and consulted about changes to repository procedures and services.” according to them, the significance of a producer’s role is determined by “the nature of the repository and its relationship with depositor” (ross and mchugh, 2006). in the oais reference model, the first interaction between oais and producer occurs when the oais preserves the data products created by the producers. the producer first establishes a submission agreement with the oais, which identifies the sips to be submitted and sometimes reflects a mandatory requirement to provide information to the oais, in contrast to sometimes voluntary offerings of information. according to the oais model, even if there is no formal submission agreement, such as in the case of websites, a virtual submission agreement can exist to specify file formats or other subject matter that the site will accept (ccsds, 2002, p. 2-9). this process of transferring information between a producer and a repository is well defined by the producer-archive interface methodology abstract standard (paimas: iso 20652). paimas describes four main phases of the interaction: preliminary, formal definition, transfer, and validation phase (ccsds, 2004). in the preliminary phase, all necessary preliminary information for data archiving is examined, for figure 1. oais functional entities (ccsds, 2002, p. 4-1) 8 iassist quarterly winter 2011 iassist quarterly instance, definition, volume of data, intellectual property, associated cost, and capability needs for ingest process. then, a producer and a repository set the preliminary agreement. this phase should be undertaken as early as possible, even before data creation. based on this phase, an entire process is detailed in the formal definition phase and results in the creation of a data dictionary, data model, and submission agreement. the transfer phase occurs when actual data transfer from the producer to the repository takes place, based on the previously planned agreement. when this sip is received, the validation phase is followed, which can be automatic for some systematic parts such as file sizes or more in-depth for issues such as completeness of submission based on the plan (ccdsd, 2004, pp. 2-3 2-4). the audit and certification of trustworthy digital repositories (2011) also makes several recommendations regarding deposit and ingest processes to develop a trusted digital repository. similarly to what is noted in paimas, the repository should clearly specify the information that needs to be associated with specific content information at the time of its deposit, and should communicate clearly what producers need to provide. although the repository is responsible for ensuring that it can extract information from sips and for verifying each sip for completeness and correctness, it is recommended that the repository provide the producers or depositors with appropriate responses at agreed-upon points during the ingest process. this continuous interaction is important to ensure that the producer can verify that there are no inadvertent lapses in communications, which might result in loss of sips (ccsds, 2011, p. 4-2, 4-6). in the oais reference model, a typical sip consists of the data inventory forms and actual data, or the content information. the inventory forms include (1) pdi (e.g., treatments, parameters measured, research subjects and ids, date/period of collection, collection location, analysis phase, and comments) and (2) descriptive information (e.g., title, description, keywords, principal investigator’s and co-principal investigator’s names). content information is the original target of preservation in oais, and it refers to content data objects as well as representation information. it usually consists of physical samples, spreadsheets, final science reports, published articles, procedural documents, crew logs, photographs, videotapes, analog tapes, digital or printed images, and other types of digital data files (ccsds, 2002, p. a-13). in the oais model, content information allows the data to be fully interpreted into meanings that can be understood by a designated community. if multiple data submission sessions exist, all representation information for each file should be provided, such as how frequently data submission sessions (e.g., one per month for two years) will occur and whether any access restrictions to the data exist (ccsds, 2002, p. 2-9). because it is well known that compliance with oais would be aligned with a concept of “trusted digital repository,” efforts have been made to build a system or process in compliance with oais. however, since the oais model intends to deal with digital objects in a general sense, archival communities or repositories need to translate oais concepts and terminology into their specific context. of course, the elements of sips can differ depending on the nature of the sips. for instance, in a social science context, a typical example of a data object is a numeric survey data file and the associated technical information (codebook) that makes up the representation information used to understand and interpret codes in the data file. representation information should not only include the information used to understand the numeric data (e.g., a codebook), but should also include information to enable the understanding of interpretive information. thus, documentation on original instruments and explanations of methodology are needed to allow users to understand the question flow and determine how questions relate to variables in the resulting data file (vardigan and whiteman, 2007). while efforts are made to understand data archiving processes in a certain repository and map them into the oais model to conform with the archival responsibilities of a trusted oais repository (vardigan and whiteman, 2007), examining this process is worthwhile in larger contexts such as social science data repositories. to respond to the growing need for the archiving and preservation of research data, examining the current status of data management practices, particularly in the sip context, is critical to building more trusted repositories. methods as previously noted, this study examines the requirements for depositors when they submitted data to repositories. to analyze the current practices among data repositories, a content analysis methodology was used, and a protocol was developed to examine the criteria or requirements that exist within depositors’ guidelines or deposit forms. for this study, data depositors’ guidelines or deposit forms were collected from social science data repositories in the united states. data can be deposited either in the institutional repository (ir) or in disciplineor domain-specific data repositories, but this study limits its scope to domain-specific data repositories that contain social science data. while both ir and domain-specific repositories aim to preserve research materials and provide access to them, they are significantly different. irs focus more on publication-related materials from multiple subject areas within a single organization, whereas domainspecific repositories manage collections grouped by type, subject, or discipline-oriented research needs (green and gutmann, 2007, pp. 39-40). in addition, because the diverse nature and types of data from different domains can affect data management requirements, this study only focuses on social science data. social science data repositories, which are not a part of irs, were initially identified from the three lists provided by mcgraw-hill ryerson, data on the net, and international federation of data organization for the social science6 the three lists provide names of 47, 85, and 32, respectively, (including redundant names across the lists) social science data repositories in the world. data repositories outside the u.s. were first excluded from these lists, which left 46 repositories in the u.s. among those 46 repositories, government organizations that only deal with census data and do not receive data from researchers were excluded. after eliminating them, publicly available depositors’ guidelines or deposit forms were collected from the data repositories’ websites, but few social science data repositories have publicly available deposit guidelines or forms. it was also unclear whether some repositories accept data from individual researchers or only from government or research institutions. among repositories that mentioned data deposits, some did not provide information about the manner in which researchers could deposit data. if the organizations mentioned that they receive data from researchers but do not provide information regarding deposits, the repositories were asked if they had written guidelines or forms for depositors. when they were asked about depositing guidelines or forms, a few had forms but only provided them when asked; others either did not yet have a procedure or were in the process of developing one. three repositories share one deposit guideline through the partnership, thus they are counted as one repository in this study. throughout this process, 14 documents from 16 repositories were collected in october 2011. iassist quarterly winter 2011 9 iassist quarterly to conduct the content analysis, an initial protocol was developed based on the sip elements of the oais model. the initial protocol included requirements regarding (1) descriptive information (project or study level), (2) actual content (data) and related information, and (3) information on files. each category contained detailed elements. however, since the collected guidelines or forms contained different elements or requirements, the protocol was modified throughout the coding process. the resulting elements that are seen in tables 1-3 reflect the oais sip data categories and contain the specific items found in the depositor guidelines. findings all 16 repositories are university affiliated, having partnerships with either university libraries or departments. however, as already noted in the methods section, they are not part of university irs, but rather social science domain-specific repositories. learning about repositories’ characteristics from the information publicly available on their websites was difficult because what and how much information was shared on the web differed greatly among repositories. for example, not all 15 repositories explicitly displayed information about collection size. numbers of staff in the repositories were generally between five to nine for repositories which provided that information, but in cases in which social science archives are run as parts of university libraries, it was hard to determine the exact numbers of staff who work for the repositories. except for one repository, all provide online search systems or online catalogs. study level descriptive information requirements project or study level information includes information about research projects that produce data submitted to repositories. descriptive information about the projects creating the data is significant as it provides provenance for the data. the terms used to refer to this information vary, but the concepts are similar. the 14 collected deposit forms varied significantly. while some asked for all detailed information about a study and provided specified requirements, others had only generic requirements and asked for metadata. in this case, the data depositors determined the metadata that should be provided. surprisingly, not all repositories asked for the title and description of the study. a study’s title is fundamental; by not always asking for title information, repositories may be assuming that titles would come with submissions or it would not be necessary for all cases since they require to submit titles of data, as can be seen in table 2. while a description of the study would enhance the understanding of the data and provide more context, only four repositories require this information, and two repositories state that it is optional. some elements of a study are not necessarily the same as the information on the data being deposited, such as the subject (or area of investigation) and the time period of study. the area of investigation refers to the topical subject area on which the research was conducted, and the time period of study refers to the entire study duration, which is different from the data collection time period. only one repository required subject terms or keywords. three different categories of personnel information may be required: information on the principle investigator (pi) or co-principle investigator (co-pi), information on the data producer (if different from the pi), and information on the depositor (donor or contact person). each of these categories usually requires a home address, telephone number, e-mail address, and fax number. some repositories asked for information on the affiliated institution. one repository specifies all three and asks for information in case they are different, but usually repositories do not differentiate among pis for investigators, data producers, and donors or depositors of the data. the definition of donor is sometimes not well defined and could refer to either the person who deposited or who owns the data. repositories that include a depositor agreement form with the deposit form do not ask for duplicate depositor information. interestingly, one repository requires donors to indicate that they are willing to help potential users with any problems that they would have. five repositories require affiliated agency and funder information. three repositories require a grant number with the name of the grant agency, if the research was supported by a grant. content (data) and related information requirements repositories list the actual content required to be submitted with the data, as well as information associated with the data. in general, more requirements are found on deposit forms regarding actual data and related information. these requirements include descriptive information about the data, the actual data being submitted, some contextual information usually referred to as “supporting materials” or “document description,” and provenance information, which tracks changes to the data from the moment of creation. eight repositories ask for descriptive titles of and types of data, and seven repositories require data collection dates. three repositories require either one or more than three subject terms to describe the content of the data, and note that they use the submitted subject terms as subject categories in their data catalog. among the actual files that need to be submitted to repositories, all repositories naturally require the data file. submitting a codebook and instrument is either required or encouraged by more than half of the repositories examined in this study. however, there are variations regarding the requirements for creating a codebook. while some repositories simply suggest, “submit a codebook,” three repositories table 1. requirements for descriptive information found in deposit forms (n=14 unique deposit forms) table 1. requirements for descriptive information found in deposit forms (n=14 unique deposit forms) table 1. requirements for descriptive information found in deposit forms (n=14 unique deposit forms) table 1. requirements for descriptive information found in deposit forms (n=14 unique deposit forms) requirements frequencyfrequencyfrequencyrequirements required optional not mentioned title of study description of study subject/area of investigation time period of study principal investigator (co-principal investigator) data producer (of creator), if different subject term agency/funder identifier copyright check donor/contact person/depositor study metadata in general (not specified) 4 4 4 2 6 3 1 5 1 3 4 2 2 3 1 10 8 10 11 8 11 13 9 13 8 10 10 iassist quarterly winter 2011 iassist quarterly provide detailed guidelines on what the codebook should include and how researchers should prepare it. one of the repositories requires that researchers “list all variables, variable descriptions, and information to understand variables” (r05). another repository emphasizes the significance of a well-prepared codebook, since “it is critical to interpret data and output files” (r10), and asks for the “location of variables in data, name and value, exact question wordings with exact meanings, value labels, missing data codes, etc.” (r10). three repositories ask for a data dictionary that describes indexed or other constructed variables. types and scales of variables and technical information about variables, which refer to information such as rows/columns of variables, variable length, numbers of variables, and weighted variables, are sometimes required in a codebook. other repositories do not state that this information should be included in a codebook, but ask that it be provided as separate documentation. one repository specifically requires information regarding the relationship between variables or tables in a data set. one repository asks for a methodological abstract in a codebook, and half of the repositories examined in this study (seven) require separate methodology documentation. the content of the methodology section also varies depending on the repository; some just require a description of the methods, and some ask about the mode of data collection (e.g., face-to-face, telephone survey, random digit dialing, computerassisted telephone interview, mail, web survey), time span covered by the data, and dates the data were collected. seven repositories also ask for information on sampling, which includes coverage, sampling techniques, response rate, or procedures. although tracking changes to data is critical, only three repositories require documentation on data edit/ cleaning procedures, or information on how the data were changed from creation to the moment of deposit. in addition, four repositories require de-identification, although this process should be required for any data containing personal information, such as names, addresses, telephone numbers, and social security numbers. de-identification is a commonly required practice in social science research, but only four repositories ask for de-identified data or check to see whether de-identification was properly done. seven repositories require or encourage submitting final reports or publications if such documents result from the submitted data. three ask for proper citations for the reports or publications along with the actual reports or publications. five of these six repositories ask for final products and require or encourage providing information on analysis performed on data. an oais recommendation calls for checking for access restrictions on the data when they are deposited. half (seven) of the repositories require providing use of restriction information. table 2. requirements for content (data) and related information found in deposit forms (n=14 unique deposit forms) table 2. requirements for content (data) and related information found in deposit forms (n=14 unique deposit forms) table 2. requirements for content (data) and related information found in deposit forms (n=14 unique deposit forms) table 2. requirements for content (data) and related information found in deposit forms (n=14 unique deposit forms) requirements frequencyfrequencyfrequency required optional not mentioned description about content included title of data data collection date types of data subject terms for data final report/publication generated by data data file codebook instrument data dictionary data collection methodology types and scales of variables technical information about variables sampling data edit/cleaning procedure relationship between documents/tables/variables analysis performed on data de-identification use of restriction check 4 8 7 5 3 4 14 7 6 2 7 4 5 7 3 2 4 4 7 3 1 1 1 2 1 10 5 7 9 11 7 0 6 7 11 7 10 8 7 11 12 9 10 7 file requirements given that the repositories studied are social science data repositories, most either have a requirement for data file formats, particularly regarding statistical data, or state the “preferred” file format for submission. one repository has “no required format” (r09). three mention that the format should be “open standard” (r01), “user-friendly format” (r04), or “in ease of use” (r10). preferred formats or accepted file types were usually ascii, spss, sas, stata, excel, and arcgis. however, only four repositories require information on the version of the software. one repository specifies the versions of the software that it accepts (for instance, spss version 7.x to 16.x (r11)). r10 states that it strongly prefers ascii to maximize the use across different software packages because “files created with older versions may limit readability and usability in the future.” three repositories require spreadsheets with cvs but in tabor comma-delimited format, and one (r13) states that the file “should be easily converted to open or non-proprietary formats meeting iso standards.” only one repository requires submitting information about the platform environment, which affects the software being used. half of the repositories (seven) examined in this study have a required format for text document files (both text as data and text as documentation about data). other repositories do not specify the media to be submitted (paper versus digital format) and assume that all files are digital; one repository requires both paper and digital format, whereas another states that it does not accept paper. the last repository states that it will take paper if that is the researchers’ only option for submission. txt and pdf are the most common file formats preferred by the repositories, but most repositories accept other formats, including word files (doc), ascii, rtf, xml, and odt (opendocument text). only three repositories mention image/audio/ video file formats, possibly because those formats are not as common as data or text files in the social science repositories. two of the repositories prefer tiff, jpeg (one in particular mentions jpeg2000), iassist quarterly winter 2011 11 iassist quarterly and gif files, but the other accepts a greater variety of formats such as png, bmp, pcd, and pcd. in general, except for the file formats, not much information is required and not many requirements exist regarding files. although file compression is known to possibly affect bits of information (heydegger, 2008; panzer-steindel, 2007; wright, miller and addis, 2009), only one repository has requirements about file compression, stating that files can be compressed using 7-zip and winzip (r13). three repositories specify delivery methods and media formats for depositors to use, and two repositories have a system that allows depositors to directly upload all necessary files, although they also receive files from depositors. cds are common across repositories (r03 specifies “ibm compatible cds”), and other delivery methods include ftp and e-mail attachments. among the three repositories that ask for “data edit/ cleaning procedures,” only one requires data file version and update frequency information. the repository does not ask for all different versions of a data file, but does ask for the version of the submitted data file and how frequently it is updated, if it is updated. regarding data file naming, while one repository requires a list of data file names, two ask that depositors follow a specific schema. one repository recommends using a consistent and descriptive file naming scheme, enabling files to be easily identifiable for reference purposes as well as to facilitate operation of the database system. the other provides a way to describe file names, which should consist of author(s), short name of data, years, and other information. discussion and conclusion since this study examined only domain-specific, non-ir social science data repositories, the findings may not be generalized across all social science data repositories in the united states. for instance, characteristics of small-scale data repositories that are affiliated with university departments or collections that are a part of an ir might be qualitatively different, and thus might employ different practices in accepting data from individuals. the findings of this study, however, reflect current deposit requirements and practices for universityaffiliated, social science data repositories. overall, the requirements for data deposit, both regarding the content that should be submitted and the information that should be provided to repositories, vary from repository to repository. requirements range from minimal wherein a repository just asks the user to submit data; to more elaborate guidelines for researchers regarding how to prepare data for deposit, with detailed requirements about file naming, file format, and all necessary information that should be accompany the data. as already discussed, the oais model describes the sips as consisting of inventory forms, which are comprised of pdi and descriptive information, and content information, which contains content data objects as well as representation information. the oais states that the pdi must include information “describing the past and present states of the content information, ensuring it is uniquely identifiable, and ensuring it has not been unknowingly altered” (ccsds, 2002, p. 4–27)., because the pdi ensures that information stored is described sufficiently so it can be accurately retrieved for future users, having a requirement for it is significant for deposits. for content information, the four categories of pdi (reference information, context information, table 3. requirements for files and related information found in deposit forms (n=14 unique deposit forms) table 3. requirements for files and related information found in deposit forms (n=14 unique deposit forms) table 3. requirements for files and related information found in deposit forms (n=14 unique deposit forms) requirements frequencyfrequency required not-mentioned data file format document file format image file format audio file format video file format file compression data file size data file naming software name software version platform data file version data file update frequency numbers of file delivery (media) format 11 (1*) 7 3 3 3 1 4 (2**) 3 4 4 1 1 1 1 3(2***) 2 7 11 11 11 13 8 11 10 10 12 12 12 12 9 *one repository mentions that it has no required format*one repository mentions that it has no required format*one repository mentions that it has no required format **two repositories mention that there is no restriction on file size. **two repositories mention that there is no restriction on file size. **two repositories mention that there is no restriction on file size. ***two repositories ask depositors to deposit directly to their system. ***two repositories ask depositors to deposit directly to their system. ***two repositories ask depositors to deposit directly to their system. provenance information, and fixity information) are critical to the integrity of the information as well as being a good practice for preservation, according to preserving digital information: report of the task force on archiving of digital information (1996). the collected elements from the deposit forms in this study include elements for creating pdi, which must all be presented in the aip later. provenance information documents the history of the content information, its origins, and chain of custody (task force on archiving of digital information, 1996, p. 16). among the deposit forms examined in this study, some descriptive information about data, data processing information (e.g., cleaning or editing history), data file versions, and updating information is part of provenance information. context information about the relationships of the content information to its environment (ccsds, 2002, p. 4–28) would include the technical context of information, linkages among information, and social environment factors (task force on archiving of digital information, 1996, p. 19). among the deposit forms examined in this study, some requirements for files (e.g., formats, software information, platforms, etc), the relationship between documents/tables, and the use of restriction checks would satisfy the efforts to document the context information of data. reference information would include study-level descriptive information as well as some descriptive information of the data (e.g., data title, data collection date, data producer, etc.) so repositories can create bibliographic metadata as well as proper citations. fixity information exists to check if the content information has been altered in an undocumented manner (ccsds, 2002, p. 4-28). while it is relatively easy for a creator of digital objects to alter or retract previously released information (task force on archiving of digital information, 1996, p. 14), checking the number of files, measuring byte counts, recording these counts, or recording length can be one way to ensure fixity once content is within a repository. not much fixity information is required of depositors, but some elements are discussed—for instance, the numbers of files and data file size. 12 iassist quarterly winter 2011 iassist quarterly the requirements for content information also varied across repositories, but in general, there are more requirements for ci than the other types of required information. these more extensive requirements concerning representation information may be necessary, however, since it is critical to understanding not only what variables in a data file mean, but also the actual sequence of bits that makes up the file types, which makes it possible to render the file in the future (vardigan and whiteman, 2007, p. 77). complete content information will allow the full interpretation of data, as the oais model suggests. although the components identified from the deposit forms collected in this study include minimum elements for inventory forms and content information, questions persist regarding how many repositories will adopt these elements and require them for deposit, and how much these requirements reflect compliance with the oais model. as already discussed, since the number of forms collected in this study is small, it is hard to make generalizations from the findings. however, the findings suggest implications for developing good practices for data deposit by examining current practices and mapping them into the oais model. by employing good practices when data come to repositories, repositories enhance users’ trust, as “trust in data was intended to strengthen as good practices and standards are established” (carlson and anderson, 2007, p. 645). future studies as this study solely relies on the collected documents, it may provide a limited view of sips and the data deposit process. for instance, to examine the full process of communication between depositors and repositories, it is necessary to know how repositories follow up on submitted data. both the oais model (ccsds, 2002) and the audit and certification of trustworthy digital repositories (ccsds, 2011) state that it is a repository’s responsibility to verify each sip for completeness and correctness so all information can be extracted for aip and dip. in this study, seven repositories mention proof-edit or verification processes, while others do not mention any such things at all, although it is still possible they are doing so internally. among those seven repositories, two state that they “do not edit or proof read the contents of deposited files” (r01) or “provide comments about the quality” (r09). the other four mention that they will verify the accuracy of final files, and depositors can be contacted to reformat or reorganize the data so the repository can meet its archival needs and goals. one repository says all submitted materials and accompanying metadata are subject to the approval of the repository, and metadata can be revised to enhance access. thus, examining internal archival processes in data repositories is essential to fully understand current data deposit practices. for instance, close examination of metadata after data is processed in repositories and comparison with metadata when it is deposited would give an insight about what information is added. interviewing data managers or archivists would be necessary in order to fully understand how decisions about what additional information is needed are made and how missing information is acquired. references committee for a study on promoting access to scientific and technical data for the public interest, national research council. (1999). a question of balance: private rights and the public interest in scientific and technical databases. washington, dc: national academy press. borgman, c. l. (2010). research data: who will share what, with whom, when, and why? presented at the china-north america library conference, beijing. available at http://works.bepress.com/ borgman/238 carlson, s., & anderson, b. (2007). what are data? the many kinds of data and their implications for data re-use. journal of computermediated communication, 12(2). available at http://jcmc.indiana. edu/vol12/issue2/carlson.html consultative committee for space data system (ccsds). (2002). reference model for an open archival information system (oais). washington, dc, usa: the consultative committee for space data systems. consultative committee for space data system (ccsds). (2004). producer-archive interface methodology abstract standard. washington, dc, usa: the consultative committee for space data systems. consultative committee for space data system (ccsds). (2011). audit and certification of trustworthy digital repositories. washington, dc, usa: the consultative committee for space data systems. faniel, i. m., & zimmerman, a. (2011). beyond the data deluge: a research agenda for large-scale data sharing and reuse. international journal of digital curation, 6(1). available at http://ijdc. net/index.php/ijdc/article/view/163 green, a.g. and m. gutmann. (2007). building partnerships among social science researchers, institutionbased repositories and domain specific data archives. oclc systems & services: international digital library perspectives 23: 35-53. heydegger, v. (2008). analyzing the impact of file formats on data integrity. proceedings of archiving 2008, bern, switzerland, june 24-27. hey, t., trefethen, a. (2003). the data deluge: an e-science perspective. in f. berman, g.c. fox, & t. hey, (eds.), grid computing: making the global infrastructure a reality. new york: wiley. hey, t., & trefethen, a. (2008). e-science, cyberinfrastructure, and scholarly communication. in g.m. olson, a. zimmerman, & n. bos, (eds.), scientific collaboration on the internet. cambridge, ma: mit press. interuniversity consortium for political and social research (icpsr). (december 2009). principles and good practice for preserving data. ihsn working paper no 003. national academy of science. (2009). ensuring the integrity, accessibility, and stewardship of research data in the digital age. washington, dc: nas. available at http://www.nap.edu/catalog. php?record_id=12615 national science board. (2005). long-lived digital data collections. available at http://www.nsf.gov/pubs/2005/nsb0540/ lord, p. & macdonald, a. (2003). data curation for e-science in the uk: an audit to establish requirements for future curation and provision. twickenham, england: 17-55 panzer-steindel, b. (2007). data integrity. april 8, 2007. available at http://indico.cern.ch/getfile.py/access?contribid=3&sessionid=0&resid =1&materialid=paper&confid=13797 ross, s., & mchugh, a. (2006). the role of evidence in establishing trust in repositories. d-lib magazine, 12. doi:10.1045 july2006-ross task force on archiving of digital information. (1996). preserving digital information. report of the task force on archiving of digital information. the commission on preservation and access. available at http://www.eric.ed.gov/ericwebportal/ contentdelivery/servlet/ericservlet?accno=ed395602 vardigan, m., & whiteman, c. (2007). icpsr meets oais: applying the oais reference model to the social science archive context. archival science, 7, 73-87. doi:10.1007/s10502-006-9037-z wright, r., miller, a., & addis, m. (2009). the significance of storage in the “cost of risk” of digital preservation. international journal of digital curation, 4(3). available at http://www.ijdc.net/index.php/ ijdc/article/view/138 iassist quarterly winter 2011 13 iassist quarterly notes 1. a doctoral student at the university of north carolina at chapel hill, school of information and library science. 216 lenoir drive cb #3360 100 manning hall, chapel hill, nc 27599-3360, usa. ayyoon@ email.unc.edu 2. an alumni distinguished professor at the university of north carolina at chapel hill, school of information and library science. 216 lenoir drive cb #3360 100 manning hall, chapel hill, nc 275993360, usa. tibbo@email.unc.edu 3. according to the oais definition, a producer is the role played by those persons, or client systems, that provide the information to be preserved (ccsda, 2002, p. 2-2). the interuniversity consortium for political and social research (icpsr) (2009) supports a producer’s role in data preservation, as it “generates or is responsible for data to be preserved and provides the data to the archive or unit responsible for preservation” (p. 7). 4. the oais model defines content information as “the set of information that is the original target of preservation. it is an information object comprised of its content data object and its representation information. an example of content information could be a single table of numbers representing, and understandable as, temperatures, but excluding the documentation that would explain its history and origin, how it relates to other observations, etc.” (ccsda, 2002, p. 1-8). 5. the oais defines pdi as “the information which is necessary for adequate preservation of the content information and which can be categorized as provenance, reference, fixity, and context information” (ccsds, 2002, p. 2-11). 6. a list provided by mcgraw-hill ryerson: http://www.socsciresearch.com/r6.html; a list provided by data on the net: http://3stages.org/c/es2.cgi?search=dataarchive&file=/data/data. html&print=notitle&header=/header/archive.header; a list provided by the international federation of data organizations for the social science: http://www.ifdo.org/network/index.html 1st newsletter vol. a no.2 public policy areas; and (3) it can relate sets of data files to those of similar focus that have been described previously. if those functions are enough to justify its conhnued publication, then the newsletter can go on for some time in the future however, it is important to raise the possibility that the print medium is now an out-dated mode of communication in the field of data reference. on-line systems at individual archives and projects whereby archives exchange data descriptor tapes certainly hold the promise of a much improved data reference system. networking and other developments, such as improved cataloguing of machine-readable data files, also offer interesting possibilities. it is hard to imagine that one system will dominate the field of data reference in the future, and therefore one important function that organizations like lassist can play is to recommend ways in which new and old reference systems can be effectively integrated in the future. on-line reference tools for the hard sciences gordon h. wood canada institute for scientific and technical information national research council of canada ottawa, ontario i ir toduction it is one thing to know or suspect that collections of machine-readable numerical data pertinent to one's discipline or problem may exist, it is something else to find that data, gain access, and use it profitably in ones research. the purpose of this paper is to review briefly the methods presently used to generate and access numeric data bases relevant to the so-called "hard saences " speaal emphasis will be given to the areas of on-line data retirieval and manipulation--areas where the state of the art in the "hard sciences" is generally conceded to be ahead of that in the "soft sciences." the reader wishing an inventory and description of the many scientific/technical data bases that are available worldwide is referred to the references at the end of the paper. for the sake of clarity, it is useful to define a few terms as they will be used in this paper. a. (scientific/technical) numeric data base an ordered collection of numbers whose values: 1. correspond to various properties, parameters or attributes of elements, substances or systems. 2. are critically evaluated by experts prior to their being included in the data base. b. numeric data base system a numeric data base system consists of one or more machine-readable scientific/technical numeric data bases as defined above plus: 1. programs for searching, retrieving and organizing the data according to user selected criteria and, usually, 2 programs to manipulate the data. (in general the latter property is what sets a numeric data system apart from a simple handbook or compendium. for example, a search routine may retneve data giving the co-ordinate positions of the atoms in a given crystal a simple command permits the user to calculate the various interatomic distances and the angles between the bonds joining the atoms. another command generates a two dimensional dravking of the crystal projected along any desired axis or plane.) newsletter vol. 4 no.2 c. numeric data base syslem a numenc data base network consists of one or more interconnected data base systems to which access is gained from a variety of remote locations by appropriate communication links. ii numeric data base creation experience has shown that an essential factor in the long term acceptance and success of a numenc data base system is the existence of an associated data evaluation center--a place where the data relevant to a particular data base arc processed, both initially and in a continuing sense, for inclusion in the file. ideally such a center should be located in an active research laboratory environment and be assured of long-term, stable finanaal and human resources. in practice, data evaluation centers are often co-ordinated and partly funded by national bodies established for that purpose such as the national standard reference data syslem in the u.s.a. and the science research council in the u.k. a framework for international co-operation in provided in part by the committee of data for science and technology (codata) and sub-groups of ma|or international scientific organizations such as the international unions of pure and applied chemistry and of pure and applied physics the functions of a data center may be summarized as follows: a. search the world literature, both published and non-published b. retrieve and index papers and reports within the area of interest. c. extract the numencal data. d. check and evaluate the data with respect to accuracy, overall quality of the work, consistency with previously published values, etc. e. cast the data into the desired format and merge into the file it is perhaps instructive to briefly consider the rationale behind some of the attributes and functions desired for a data center long-term support with respect to funding and personnel is important because new data are continuously being generated, technical advances tend to make old data obsolete (eg too imprecise, too inaccurate, or too narrow in scope) and key personnel, being subject to the vagaries of the human condition, never last forever. critical evaluation filters out inaccurate or poorly documented data. the distilled product of evaluation is a compact, reliable data base which is tractable to handle and provides a real benefit to the researcher who perhaps has neither the resources, time, nor inclination to search the vast open literature himself. clearly a data base system is no better than the data on which it is founded if researchers lack confidence in the quality and reliability of the data, they are not likely to use a data base system no matter how sophisticated or elegant the accompanying software may be. having a data center located in an active research environment helps to assure that the workers responsible for compiling and evaluating the data stay at the fore-front of their field with respect to theory and experiment. the credibility of the data base is not likely to exceed the scientific credibility of those who produce it. ill dissemination of data bases and data base systems the information contained in numenc data bases is made available to the "hard science" community in three major modes or "packages" which roughly parallel the definitions given earlier. naturally, variations and permutations of these modes arc also possible. a subscription to a network for a fee, which may consist of some combination of annual suhscnption, connect time and characters transmitted charges, a user connects appropriately to a node of a network and is given access to all of the data base systems for which he has paid most applications do not require a very "smart" terminal and the user needs to master only a few simple commands to operate on-line. because of the small capital investment and the wide variety of information potentially available, this mode of dissemination is most likely to suit the worker who has a considerable range of research interests which tend to fluctuate both in intensity and focus. usually a network will, of course, support batch and quasi "batch on-line" tasks as well. two such networks already functioning are the chemical information system in the united states, which has eleven data bases presently available with eight more under test 11], and the direct information newsletter vol. 4 no.2 access network for europe. a similar network on a smaller scale, which exists now in embryonic form, is being assembled in canada by the national research council inibally, five data bases are planned: three are crystallographic, covering the areas of metals, organics and inorganics; the fourth is thermochemical; the fifth, already functional, is a program for comparing an unknown infra red spectrum against a collection of some 100,000 or more spectra of known compounds. b. lease or purchase of a data base system in this mode, a user obtains a tape copy of the complete data base and its relevant software for mounting on a nearby, typically institutional, computing system. assuming the software is adequate for his needs and compatible with the local computer, the user need not have extensive computer expertise nor a "smart" terminal. researchers benefiting from this means of dissemination would be those who tend to work almost exclusively in one discipline and for whom the storage and usage costs of a local computer would be more favourable than online charges. update tapes would normally be supplied periodically by the data base supplier and the user would agree to distribute the tapi s no further than his own institution. c. lease or purchase of all or part of a data base with the proliferation of small yet powerful computers into many laboratories, the option of obtaining a tape copy of an entire data base, or a selected sub-set of it, for private use is becoming more popular. in this distribution mode, the user either requests an actual tape or copies what he needs on-line and proceeds locally form there using his own customized search, retrieval and manipulation routines. such a user, of course, needs not only to support a mini-computer and its ancillary equipment but needs considerable programming expertise as well. again, update tapes would be made available from the supplier and restrictions concerning extra institutional use would usually apply an example of a user for whom this option might be of value would be a lecturer in thermochemistry who leases the data covering a few hundred compounds of interest and generates an appropriate suite of programs. the pedagogical value is clear. his students are able to tackle practical problems, rather than artificially contrived ones, without tedious calculations causing them to lose sight of the chemical principles involved. rv data bases of numeric data bases the media employed in cataloguing and marketing numeric data bases for the hard sciences are basically the same as those covering numeric data bases in general first, there are compendia and directories [2,3,4] which cover a broad range of data bases but provide limited detailed information about any one in particular. second, there are a number of publications [5,6,71, devoted to on-line non-bibliographical data base methods, technology and management, which periodically review the current world-wide inventory of numeric data bases. third, there are diffuse or indirect sources which, while less systematic than those just mentioned, often provide the most useful information. in this category are found scientific conferences, journals of the various learned societies, and the very effective "grapevine" formed by scientists and information specialists. v conclusion the future of on-line numeric data bases in the hard sciences appears very promising indeed. with computer technology improving so rapidly because of pressures from many sectors, it is likely the greatest impediment to the growth of numeric data bases as reference tools will be only the lack of imagination and interaction in the supplier/user component of the activity. references see, e.g.. heller s.r. & milne g.w.a american uboratory 12, 33 (1980). computer-readable data bases, a directory and data source book. martha e. williams, et al, eds. american society for information science, (1979). directory of online data bases. ruth n. landau, et al. eds. cuadra associates, santa monica (issued semi-annually, quarterly updates.) eusidic database guide. alex tomberg, ed. learned information, new york (1978). on-line review, learned information, york. (published cjuarterly.) online and database (2 journals). pemberton, ed. online, inc., weston, (published quarterly.) codata buuetin. codata secretariat, (published irregularly.) new j.k. ct. paris 1/2 rasmussen, karsten boye (2021) editor’s notes: data management for students, researchers, and data science projects, iassist quarterly 45(2), pp. 1-2. doi: https://doi.org/10.29173/iq1018 data management for students, researchers, and data science projects welcome to the second issue of iassist quarterly 2021 (iq vol. 45(2) 2021). data management is the focal point of the articles in this issue of the iq. aspects of data management have often been the central topic of many earlier iq articles. metaphorically simply said: data is what iassist members breathe data management is how we breathe it. the first two articles concern raising the attention and knowledge of data management among students and researchers. in both articles it turned out that more than the target group could benefit from the efforts. when instructing students, faculty at the university gained as well because the data management workshop was well integrated with the teaching. and when nudging researchers and graduate students in illinois with data management hints, the data nudge also extended to other people. the positive reception also from persons not directly targeted by the efforts can be viewed as a sign that there is in general a great need for data management skills. some of the skills are highly specialized as exemplified in the third article on databook but the general skills are necessary at all levels of the data society. in the first article elizabeth blackwood from the channel islands campus of california state university describes how undergrad students at a university without 'very high research activity' benefit from knowledge of data management in the article 'outside the r1: equitable data management at the undergraduate level'. the project addressed the problem that students and faculty at these educational institutions often lack data management instruction. the article describes in much detail how a workshop was prepared and planned. you will find meticulous descriptions with headings and subpoints of the flow and structure of the data management workshop, with headings like 'what data', 'why management', 'data management plans' and more. because of its success, the workshop and student assignments have been integrated as permanent parts of the undergraduate course. it was furthermore demonstrated when reality kicked in with covid-19 that the workshop was transferred to remote delivery without difficulty. the term 'equitable' in the title is because knowledge of data management is part of the essential preparedness for access to both graduate school and ultimately to the digital society with interesting career paths also for students that were 'outside the r1'. the second article is titled 'better data management, one nudge at a time'. a good question and the answer in the same sentence! the data management nudging was implemented and described by a team at the university of illinois at urbana-champaign consisting of daria orlowska, colleen fallaw, yali feng, livia garza, ashley hetrick, heidi imker, and hoa luong. the team had noticed that although researchers at illinois are aware of data services, the services are under-utilized. by releasing a monthly data nudge email the researchers became more alert to best practices and also to the resources on campus. the data nudge has high ratings and many email compliments, and now has more than 500 subscribers. the article presents the background of data services at some universities, and details on the distribution and growth in the number of subscribers as well as of some of the topics used in the data nudge. at the top of the most popular topics is 'file organization'. all past nudges can be found on the web. i am a great fan of using the iso date, so i like the data nudge of 2021-05-25 (not 5/25, 2021). the leftmost carries most importance. i am not https://doi.org/10.29173/iq1018 2/2 rasmussen, karsten boye (2021) editor’s notes: data management for students, researchers, and data science projects, iassist quarterly 45(2), pp. 1-2. doi: https://doi.org/10.29173/iq1018 talking politics here, just what we expect from numbers. good data management respects expectations. among the tables in the article, you will learn that 'data sharing' is a topic often promoted in the data nudge – important data management topics need to be addressed regularly. the third article is 'databook: a standardised framework for dynamic documentation of algorithm design during data science projects' – an extensive description of a system to secure documentation and reproducibility, which are fundamental concepts in data management. anna nesvijevskaia proposes a framework called databook for data science projects, based on points of critique of other platforms. the author finds that traditional data management tools are unable to absorb the final algorithmic model while data science platforms tend to structure only technical aspects, which excludes some project stakeholders. as data science is often developed in the context of business, projects involve people with many different skills, and the division between data skills and business skills when data scientist meets decision-maker can be difficult to bridge. the article is extensive and includes many references and databook builds upon earlier attempts, including the most successful – the cross-industry standard process for data mining (crisp_dm . the framework developed has been tested in projects at the french company quinten during 2017 to 2020. further details of the databook are shown in an appendix. anna nesvijevskaia is researcher at the laboratory dicen ile de france and a partner at the private data science company quinten. enjoy the reading! submissions of papers for the iassist quarterly are always very welcome. we welcome input from iassist conferences or other conferences and workshops, from local presentations or papers especially written for the iq. when you are preparing such a presentation, give a thought to turning your one-time presentation into a lasting contribution. doing that after the event also gives you the opportunity of improving your work after feedback. we encourage you to login or create an author profile at https://www.iassistquarterly.com (our open journal system application). we permit authors to have 'deep links' into the iq as well as deposition of the paper in your local repository. chairing a conference session or workshop with the purpose of aggregating and integrating papers for a special issue iq is also much appreciated as the information reaches many more people than the limited number of session participants and will be readily available on the iassist quarterly website at https://www.iassistquarterly.com. authors are very welcome to take a look at the instructions and layout: https://www.iassistquarterly.com/index.php/iassist/about/submissions authors can also contact me directly via e-mail: kbr@sam.sdu.dk. should you be interested in compiling a special issue for the iq as guest editor(s) i will also be delighted to hear from you. karsten boye rasmussen september 2021 https://doi.org/10.29173/iq1018 https://www.iassistquarterly.com/index.php/iassist/about/submissions mailto:kbr@sam.sdu.dk iassist quarterly / [serial] international association for social science information ser\ ice and technology lassist quarterly volume? no. 2 spring 1983 iassist 1983 annual conference digitized by the internet archive in 2010 with funding from university of north carolina at chapel hill http://www.archive.org/details/iassistquarterly72inte editorial information the lassist newsletter represents an international cooperative effort on the part of individuals managing, operating, or using machine-readable data archives, data libraries, and data services. the newsletter reports on activities related to the production, acquisition, preservation, processing, distribution, and use of machine-readaole data carried out by its members and others in the international social science community. your contributions and suggestions for topics of interest are welcomed. the views set forth by authors of articles contained in this publication are not necessarily those of lassist. information for authors the newsletter is published four times yearly. articles and other information should be typewritten and double-spaced. each page of the manuscript should be numbered. the first page should contain the article title, author's name, affiliation, address to which correspondence may be sent, and telephone number. footnotes and bibliographic citations should be consistent in style, preferably following a standard authority such as the university of chicago press manual of style or kate l. turabian's manual for writers . if the contribution is an announcement of a conference, training session, or the like, the text should include a mailing address and a telephone number for the director of the event or for the organization sponsoring the event. book notices and reviews should not exceed two doublespaced pages. manuscripts should be sent in duplicate to the editor: elizabeth stephenson institute for social science research university of california 405 hilgard avenue los angeles, california 90024 u.s.a. (213) 825-0716 or (213) 825-0711 book reviews should be submitted in duplicate to the book review editor: kathleen m. heim graduate school of library science university of illinois at urbana-champaign 329 main library urbana il 61801 u.s.a. (217) 333-1000-info or (217) 333-2306-office key title: newsletter international association for social science information service and technology issn united states: 0145-238x copyright © 1982 by lassist. all rights reserved. i a s s i s t newsletter volume 7 number 2 spring 1983 in this issue: page calendar 1 lassist 1983 annual conference 5 book revi ews 10 cataloging machine-readable data files: an interpretative manual 10 numeric databases 13 announcements 15 calendar may, 1983 census bureau training course "microdata from the i98o census" contact: dorothy chin user training branch data user services division bureau of the census washington, d.c. 20233 (301) 763-1510 atlanta, georgia may 10 dallas, texas may 2k san francisco, california may 26 may, 19-22, 1983 (assist annual conference contact: sue dodd , program chair institute for research in social science room 25, manning hall, 026a university of north carolina chapel hill, north carolina 2751^ (919) 966-33^46 philadelphia, pennsylvania june 6-8, i983 6th international conference on computers and the human i t ies contact: sarah k. burton department of english p.o. box 5308 north carolina state university raleigh, north carolina 2765o raleigh, north carolina june 6-8, i983 6th annual international acm sigir conference "research and development in information retrieval" contact : michael mcgi 1 1 national science foundation 1800 g. street, nw washington, d.c. 20550 (202) 357-955'* washi ngton , d.c june 10-12, 1983 new brunswick, new jersey international conference in data bases in the humanities and social sciences contact: professor robert f. allen room hll alexander library rutgers university new brunswick, new jersey o8903 u.s.a. july 15-22, 1983 . oxford, united kingdom annual conference of the international society for political psychology contact: professor betty glad department of political science university of illinois at urbane champa i gn 361 lincoln hal 1 702 south wright street urbana, illinois 61801-3696 u.s.a. j^'y 25-29, 1983 detroit, michigan loth annual conference on computer graphics and interactive techniques contact: siggraph '83 conference office 1 1 east wacker drive chicago, i i 1 inois 6o6oi (312) 6i»4-66lo august 11-14, 18-21, 1983 toronto, canada international time series meetings contact: o.d. anderson 9 ingham grove lenton gardens nottingham ng7 2lq england august 15-18, 1983 annual meeting: american statistical association august 11-lk, 1983 1983 public health conference on records and statistics contact: gail f. fisher, ph.d. room 2-28 center building 3700 east-west highway hyattsville, maryland 20782 washington, d.c, august 28 september 2, i983 international peace research association 10th general conference contact: ipra secretariat faculty of law university of tokyo bunkyoku, tokyo 1 13 japan gyor , hungary september 5-9, 1983 international economic association 7th world congress contact: i ea secretary general professor l. fauvel 23 rue campagne premiere 75014 paris france madrid, spain september 11-15, 1983 i5th annual conference of the society for information management san diego, california september 19-23, 1983 ifip world computer congress contact: afips 1815 n. lynn street arlington, virginia 22209 (703) 558-3600 paris, france september 25-30, i983 international society for criminology 9th international congress "relationship between criminology and publ i c pol i cy" contact : interconvention c/o mrs. i . hoi p.o. box 80 a-1107 vienna austria ens te 1 ner vienna, austria october 16-23, 1983 training course in accessibility and dissemination of non-bibliographic data in science and technology contact: dr. s. schwarz, director royal institute of technology library s-100 44 stockholm sweden stockholm, sweden december 12-15, 1983 chi '83 conference "human factors in computing systems" contact: raoul n. smith gte laboratories, inc. 40 sylvan road waltham, massachusetts 02254 (617) 466-4044 (617) 890-8460 boston, massachusetts get twe latest information and improve the quality of your data service . . . attend the lassist annual conference may 19-22, 1983 warwick hotel downtown philadelphia 3 workshops willyoubenefit??? % have you ever thought about % would you like the latest establishing a program or data information on choosing the right center at your organization but microcomputer and targeting its didn't know how to go about it? applications? % has your experience with complex -^ has choosing the right microcomputer data files such as the fames data software for economic and time series and the panel study of income dynamics data been a problem? left you frustrated? if you answered "yes" to any of the above questions, then these workshops are for you'.) , _, u __ m-9 r iiii ii iiiiiiiiiiiiiii 1 — i ir i i i i i _ "™™ t^tj workshop #1: planning a social science data center this workshop will consist of a panel of representatives from data centers in a variety of institutional settinos such as a computer center, traditional library, and a research organization. the panelists will discuss issues to be considered in establishing a program or data center with reference to their particular experiences and settings. topics covered will include planning, space requirements, budgets, staff and services, institutional support, organizational structure, sources of data, and user services. workshop #2: microcomputers and social science databases this workshop will provide an overview of one organization's experiences in choosing a microcomputer and establishing its applications. representatives of commercial developers of software for microcomputers will discuss particular software and its applications. the emphasis in software applications will be on time series analysis of economic data. workshop #3: complex social science data files topics covered in this workshop will include the use of complex data files on smaller computers--specifically using the panel study of income dynamics and the national longitudinal surveys of labor market experience (parnes data) on an hp-2000 data management system using large social science data files--with particular reference to the income survey development program test panel, and a survey of research uses to which the national longitudinal surveys of labor market experience have been put and areas which remain unexplored. note: minimum attendance levels have been established for each workshop. a workshop may be cancelled by april 27 if its minimum registration has not been met. see the registration form for more details. \^ preliminary program outline lassist annual conference may 19-22, 1983 warwick hotel philadelphia, penn. data services panel: planning and resource management systems papers on . . . how demographics can be used to guage the market potential for products and services use of data products in human service planning economic information system and its use for planning in state and local government data resources and sharing among state agencies panel: indexing for access to social science data papers on . . . indexing machine-readable data files for a social science data archive do we need a referral center for social science information? public opinion questions as a resource database impindix: an on-line index for social science data files tutorial in abstracting planned for may 20th a special session entitled "abstracting in perspective" will be offered and cover new developments, problems and prospects for abstracting material for machinereadable data files. chair: inez sperr, migration information and abstract service (mia) and president, national federation of abstracting and information services (nfais) participants: martha cornog, special projects coordinator, national federation of abstracting and information services (nfais) catherine minecci , editorial department, biosciences information service (biosis) sara strachan, associate editor, population index sist for more information on proposed panels and papers, contact the appropriate track leader: hardware and software bill gammell, roper center, university of connecticut, storrs, conn. 06268 data files pat doyle, mathematica policy research, 600 maryland ave., s.w., washington, d.c. 20007 data services peter allison, elmer holmes bobst library, social science center, 70 washington square south, new york, n.y. 10012 6 preliminary progpj\m outline hardware and software panel : data storage papers on . . . database management systems and archival acquisitions program for locatina and mounting files from a tape archive on-line databases vs. tape-stored databases panel: microprocessing papers on . . . old law new technology/new technology new law evaluating hardware computer implementation: the utility of defining utility data files panel: cross-national multi-purpose surveys papers on . . . women in development: a project of the international demographic data center the world fertility program and conditions under which different countries release their data panel: census and census-related files papers on . . . fifty years of public use data, 1940-1980 surveys conducted between the censuses: current population survey and the annual housing survey panel: philadelphia social history project papers on . . . building an individual -level , small-area data archive, 1850-1980 the study of ethnic populations in u.s. cities, 1880-1980 the public use sample of the 1910 u.s. census: (1) an overview, (2) sample design lassist conference will also feature a blue-ribbon panel on the development of "an information network for locatihg afid describing machine-readable data files' iassist annual conference may 19-22, 1983 warwick hotel philadelphia, penn. sist lassist 1982 registration form workshops conference thursday, may 19 friday, may 20 sunday, may 22, 1983 deadlines hotel registration april 22 limited number of rooms available. contact the warwick hotel by april 22, 1983. ^ay 5 (use this form) see ^lay 5 (use this form) other side workshop registration(") conference registration a $10 fee will be charged for conference registration after may 5, 1983. circle rate which applies and complete registration information iassist member non-member $70 conference and workshops (may 19-22): $90--" includes refreshment breaks, saturday evening banquet and speaker, plus conference and workshop. $40 workshops only (may 19): $50 includes workshop, refreshment breaks and workshop materials. $60 conference only (may 20-22): $70 includes refreshment breaks, saturday evening reception, plus conference sessions. $35 conference (any one day) : $40 includes refreshment breaks and one day's conference sessions. (excludes saturday reception). "workshops are by preregis t rat ion only. see other side. '^•this fee includes an lassist memberbhip for 1983. registration information name (last) (first) (initial) institution mailing address (street address) (city) (state or region) (zip) (country) telephone (area code) (number) amount enclosed (in u.s. $) make check payable to lassist 1983. workshops (participants must register prior to may 5th to attend) i i wish to register for the following workshop. (check one): microcomputers & social science data bases [ planning the social science data center accessing complex data files (e.g. parnes , psid, etc.) ^''experienced user only i registration for conference and workshop registration, mail this form and check to: iassist 83 c/o gert lewis rutgers university ccis hill center po box 879 piscataway, nj 08859 note: you must make your own hotel reservations. if you wish to stay in the warwick, you must contact the hotel before april 22, 1983. book reviews sue a. dodd , cataloging machine-readable data files: an interpretive manual . chicago, i 1 1 inois : american library association, december ]h, 1982 xx, 2a8 pages illustrated and indexed. lc 82-11597 isbn 0-8389-9365-7$35.00 this long-awaited contribution to cataloging practices moves mrdf to bibliographic legitimacy at last. designed to be used with chapter 9 of the second edition of the anglo-american cataloging rules , the manual deals with descriptive cataloging and includes the general application of the international standard bibliographic description. judith s. rowe notes in the forward that bibliographic control of mrdf is the cornerstone of the development of other products which will provide still greater access to mrdf, and, as a consequence, the information community might well view this publication as analogous to the first moonlanding. data professionals have not been able to get around the fact that without adequate widely practiced procedures for identification and integration into the larger bibliographic record, mrdf remain elusive and mysterious. no matter that researchers need them, policies are developed through their analysis, or that they often provide the basis for governmental deliberation, as long as this resource eludes the bibliographic net it does not, to most librarians or information users, exist. although intrepid data professionals have developed networks and individual contacts which have somewhat mediated the lack of bibliographic access to mrdf, these efforts have not been able to rectify the gaping hole in the bibliographic fabric. the rather convoluted genesis of dodd's masterpiece is in itself important background to the issues at hand and it is described both in rowe ' s forward and in dodd's introduction and acknowledgements. additional information on the professional activity and deliberations which corroborate the manual is described in dodd's commentary, "toward integration of catalog records on social science machinereadable data files into existing bibliographic utilities," ( library trends 30 [winter 1982]: 335-36].). the manual consists of ten chapters. the first, "what are machine-readable data fields?" begins with a delightful scenario in which a prominent sociologist uses a library's data base on mrdf to identify five files on attitudes toward and knowledge of the law, accesses them at a distant university through the ednet data network system, locates and runs statistical tests on them, and has the results delivered in order to write a paper. the first step toward making this scenario a reality is the cataloging of data files and computer programs--a step already taken at yale, princeton and the university of british columbia--a step that is a possibility for all libraries with the aid of the new manual . the manual provides examples for three categories of computerized information: numeric files, text files, and computer programs. dodd uses the aacr2 definition of mrdf: "any information encoded by methods that require the use of a machine (typically, but not always, a computer) for translation." wei 1 -formul ated sections summarize what makes computers work, computer storage, auxiliary storage devices, processing, input and output devices, and physical characteristics of mrdf. the second chapter, "data files versus documentation," describes documentation, its functions and its importance to cataloging. the major components of 10 documentation (title page, preface or introduction, processing summary, item or variable level content, and appendices) are described in detail along with wellselected illustrative examples of each. characteristics of mrdf in terms of cataloging rules are discussed in the third chapter, "why are mrdf hard to catalog?" problems analyzed include lack of bibliographic control, production (rather than publication), difficulty in determining editions, release date versus operational date, and the fluid nature of mrdf and the physical description area. ' specific rules for cataloging mrdf are outlined in chapter ^4 which is organized to replicate the numbering system for aacr2 sections 9.0 9.10. sections begin with a summary quote from specific rules in aacr2 and are followed by an interpretive discussion with examples. examples include the prescribed international . standard bibliographic description and a "guide to isbd punctuation and spacing and mrdf" is provided in appendix a of the manual . appendix d provides cataloging examples in a card format complete with main entries, headings, and access points. ' chapter 5, "concept of main entry and choice of access points," offers guidelines and discussion on how rules prescribed for print and nonprint materials can be applied to mrdf. one of the desired outcomes of cataloging mrdf and integrating information on computerized materials into a multimedia collection within a library setting is to provide the library user with enough information to make a choice among many formats of stored information. dodd uses aacr2 chapter 21, "choice of access points" and offers interpretations for the rules. ' "step-by-step cataloging examples" are presented in chapter 6 for the three broad types of mrdf--numeri c, text, and computer programs, along with primary sources of documentation. the selection of examples are fascinating and demonstrate the richness and variety of mrdf. they include the american national election study, 197^ ; symap; visitrend + visiplot; maxlink; a machine transcription of the work of thucydides; and webster's dictionary. the 38 figures accompanying these examples provide lucid representation of the cataloging process. ' serially issued mrdf such as censuses, reported proceedings, statistics reported on a regular basis, or periodical materials compiled in a machine-readable format require interpretation as derived from chapter 9 of aacr2 and modified by chapter 12 where appropriate. dodd follows the procedure used in earlier chapters quoting from the rules and providing interpretations for mrdf. bibliographic and source data bases with their constantly changing contents are treated in chapter 8, "time series dynamic data bases." dodd recommends that these be handled along the lines of a loose-leaf catalog with an open production date and incomplete size of file entry. using citibase as an example, the manual gives lucid directions for this type of mrdf. in chapter 9, "guidelines for bibliographic conventions," dodd notes, "the value of cataloging is ultimately proved not by how well each mrdf is uniquely defined, but by how efficiently the user is directed to the needed resource." the chapter is directed at data producers and distributors who have the responsibility 11 for providing descriptive information on available mrdf and for seeing that such information reaches its intended audience. sections on title and title construction, title page equivalent for mrdf, edition, edition responsibility statement, producer, generator, distributor, series, internal and external title page, bibliographic citation, and treatment of mrdf for abstracting along with specific and general instructions for writing mrdf abstracts. the final chapter, "multilevel recordkeeping for mrdf," describes four levels of recordkeeping and provides examples of forms currently in use by libraries providing their patrons with information and access to mrdf. catalog record, data abstracts, documentation, and physical characteristics are all treated in detail. the manual has four appendices: a) "guide to punctuation according to isbd(g) for mrdf;" b) "checkl ist for cataloging mrdf;" c) "worksheet for cataloging mrdf;" and d) "card format and cataloging examples for mrdf." these are followed by a thorough "glossary" derived from major sources such as aacr2 , american national dictionary for information processing , and international organization for standardization vocabulary of data processing , and a well organized index. cataloging machine-readable data files: an interpretative manual is a required purchase for data professionals, librarians, data users, data producers, and data distributors. it is a complete guide to integration of mrdf into all levels of library records. in addition, this manua 1 is unique among manuals because it is interesting to read and provides the best comprehensive introduction to the topic i have seen. it is also recommended for students beginning library and information studies--wi thout it they begin their careers in the dark ages. a final note: throughout the manual , dodd weaves the opinions and decisions of the ala resource and technical services division's cataloging and classification section descriptive cataloging subcommittee on rules for machine-readable data files. her continued emphasis on consensus decisions not only enriches the manual with a quality of col legi al i ty , but demonstrates the enormity of the work required to develop this final product. it has been a long wait but the manual lives up to the expectations of the entire information community. rush out and get yoursl 12 numeric databases special double issue of drexel library quarterly 18 (summer/fall 1982) edited by charles r. claydon and dagobert soergel. available from school of library and information science, drexel university, philadelphia, pa. 19104. $11.00. issn 0012-6160. this special issue of drexel library quarterly greatly enriches literature on numeric databases available to librarians and information professionals. it begins with the premise that users of information services who need numeric data are growing more and more dissatisfied with references to documents and expect to receive the original data in a format suitable to the problem at hand without having to go to another source. the editors, charles r. claydon of battel le's columbus laboratories and dagobert soergel of the university of maryland, intend this issue to enable librarians to transfer their knowledge of the organization and use of bibliographic databases to the domain of numeric databases. john b. fried and gabor j. kovacs provide an overview of numeric databases in the 80s with sections on videotext, software and hardware, telecommunications, and office automation. mary c. berger and judith wanger discuss reasons why libraries and information centers have not accepted numeric databases as easily as those which are bibliographic. in their article, "retrieval, analysis, and display of numeric data," they note that the provision of bibliographic services may have absorbed so much of the capacity of library and information centers to use computer-based technology that there is rel uctance to take on new obligations. costs, limited interest, and targeting of the end user toward specific applications that are job related are all additional reasons which contribute to lack of access to numeric databases in library settings. berger and wanger describe the types of databases associated with online numeric databases, coverage of subjects and types of applications, a review of basic functions, and characteristics of the system interfaces. illustrative figures accompany their discussion. organized collections of physical or chemical properties of materials expressed in numeric form are described by sherman p. fivozinsky of the office of standard reference data, nbs. he provides a short history of u.s. programs providing phys i cal /chemi cal databases focusing on the nbs standard reference data program and summarizes foreign and international efforts such as codata, lupac, and iaea. fivozinsky sees librarians as the future identifiers and providers of these services with scientists and engineers the direct users. the nih/epa chemical information system (cis) physical and chemical databases are described by stephen r. heller of the epa. major categories discussed include spectral databases and toxi cologi cal and environmental databases. access to cis is explained and future plans outlined. r. gubiotti, h. pestel and g. kovacs discuss issues concerning the design and use of numeric databases for the technical community, database management considerations, the products and services of information analysis centers, and information flow and indexing philosophies at battelle information analysis center. major social science machine-readable databases are assessed by jacqueline m. mcgee and donald p. trees who enumerate the basic criteria which determine database utility, and for each of eight subject areas they list ten databases which 13 meet these criteria and which are considered of major importance. an appendix lists available databases by subject area. the list is of value to those building a representative collection of social science databases. general principles of the structure and construction of numeric databases are organized into two sections in dagobert soergel's paper. section 1 introduces the concept of a data point and discusses how data points are related to each other and can be stored in a nonredundant way; section 2 on database construction discusses the collection of numeric data and their integration into the database st ructure. in "metasystems for integrated access to numeric data files" david m. liston, jr. and james l. dolby discuss the concept of a metadata system operated as a parallel counterpart to a numeric data system to enable analysts, decisionmakers, problem solvers, and system managers to learn enough about numeric data to enhance the likelihood that it will be used validly and appropriately. data element linkage information, a particular type of metadata, and a methodology for representing them are focused on in robert r. v. wiedrekehr's paper, "methodology for representing data element tracings and transformations in a numeric data system." an overview of the potential for utilizing database management system (dbms) philosophy within numeric database environments is described by wayne d. dominick and peggy c. weathers who discuss the major features, functions and characteristics of generalized dbms; highlight the applicability of these capabilities to numeric database environment needs and numeric database user needs; survey major current applications of dbms technology; and identify a number of research-oriented and applications-oriented future needs within this information processing environment. w. bruce ewbank describes objective measures that can be used to estimate the probable value to a user of a given database as well as the systems used to access the database. he suggests standards designed to guide a potential database user through a systematic evaluation process which should help ensure that the database is indeed appropriate for a particular application. the "use of numeric databases in reference and information services" is explored by edward p. bartkus with attention to interfacing differences among the types of databases. he delineates options to libraries which include minimal scope, clearinghouse function, coordination function, and expert specialist services. each option is discussed in detail. together the articles in numeric databases comprise a wide-ranging exploration of technical and service issues in the provision of access to numeric data. the predominance of authors from the private sector underscores that fact that public provision of these services through universities or libraries may be declining as an option. attempts to alert librarians in traditional settings to their responsibility in this area and the development of tools to integrate dataf i les (such as dodd's manual reviewed elsewhere in this issue or the winter, i982 issue of library trends , data libraries for the social sciences ) may permit a linkage of the two styles of service. the more highly developed attention to the issue on the part of private sector entities such as battelle, caudra associates, inc., king research, \k or e. i. dupont de nemours & co. indicates that if libraries default on the responsibility the private sector will be there to take up the challenge. this special issue of drexel library quarterly is highly recommended. association for computing f^,achinery: special interest group on computer and human action (acm/sigchi) sigchi, formerly sigsoc focuses on user behavior: how people communicate and interact with computer systems. the scope of the group is in the study of the human-computer interaction process, and includes research in and development efforts leading to the design and evaluation of user interfaces. topics of interest also include cognitive functions involved in the interactive process; interaction between hardware, software, the task, and the user; and, promoting an understanding of the relationship between studies of user psychology and the technology of computing and systems design. sigchi serves as a forum for the exchange of ideas among computer scientists, psychologists, social scientists, systems designers, and end-users. for information, contact lorraine borman, vogelback computing center, northwestern university, evanston, il 60201. announcements belgian archives for the social sciences (bass) has the pleasure to announce the release of the inventory of available archives. this inventory, available both in french and english, includes the full list of data archived in the bass. since they were set up in 1969, the belgian archives for social sciences, better known as bass, have built up in their data bank a collection of more than 200 data sets relating to social sciences. for information, the bass have become the official depositary of euroharometeres opinion surveys which are regularly carried out by the european community. data collected in the framework of the national programme of research in social sciences have been confided to the care of the bass as well. through association with other archives in europe and throughout the world, the bass contribute to the scientists' international cooperation by furthering information exchange and data diffusion. this new inventory follows the 1978 inventory of archives available in the bass. it includes a full list of recorded data until november 1st, 1981. any 15 archived item is accompanied by the following information: title of survey, name of the authors, time of data-collecting, surveyed population, number of unities, number of variables, accessibility degree as well as a short description of the themes taci<led in the survey. to make the inventory easier to consult, an index of authors' names, a geographical index as well as a thematic filing of the archives have been developed. full information concerning data access and data deposit modes has been included as we 1 1 . further information relating to this inventory or the bass activities can be obtained by applying to: jean-claude deheneffe, bass, place montesquieu, 1 b. 18. b13^8 louvai n1 a-neuve , belgium. 16 international association for social science infor^tion service and tecfmology i a s s i s t \ssociation internationale pour les services et techniques d' information en sciences sociales international headquartere social science computing laboratory the university of western ontario london nba 5c2 canada membership and subscription fees (effective january 1, 1981) the international association for social science information services and technology (lassist) is a professional association of individuals who are engaged in the acquistion, processing, maintenance, and distribution of machine readable text and/or numeric social science data. the membership includes information system specialists, data base librarians or administrators, archivists, researchers, programmers, and managers. their range of interests encompass hardcopy as well as machine readable data. paid-up members enjoy voting rights and receive the lassist newsletter and benefit of reduced fees for attendance at regional and international conferences sponsored by lassist. membership fees are: regular membership: $20 per calendar year student membership: $10 per calendar year subscriptions to the newsletter are available. institutional subscriptions do not convey voting rights or other membership benefits, other than receiving the newsletter . institutional subscription: $35 per calendar year (which includes one volume of the newsletter) the treasurer international headquarters, lassist social science computing laboratory the university of western ontario london, canada n6a 5c2 17 administrative committee president: sue gavrel , machine readable data archives, public archives of canada, 395 wellington street, ottawa, ontario kia 0n4 , canada regional secretaries asia : naresh nijhawan, indian council of social science research, data archive, canada : open east europe : krzysztof zagcrski , instytut folosofii i socjologii, polskiej academii nauk, nowy swiat 72, palac staszica, 00-330 warszawa, poland west europe : menk schrik. steinmetz archives. herengracht 410-412, 1017 bx amsterdam, the netherlands united states : judith s. rowe, computer center, 87 prospect street, princeton university, princeton, new jersey 08544, u.s.a. members-at-large john de vries, department of sociology, carleton university, ottawa, ontario, kis 5b6, canada sue a. dodd, social science data library, manning hall, university of north carolina, chapel hill, north carolina 27514, u.s.a. carolyn geda, inter-university consortium for political and social research, p.o. box 1248, university of michigan, ann arbor, michigan 48106, u.s.a. jackie mcgee, the rand corporation, 1700 main street, santa monica, california 90406, u.s.a. nancy carmichael mcmanus, social science research council, 1^38 corcoran street, nw, washington, d.c. 20009, u.s.a. ekkehard mochmann, zentral archi v fur empirische sozial forschung. university of cologne, bachemerstrasse 40, d-5000 cologne 41, federal republic of germany david nasatir, california state university, dominguez hills, california 90747, u.s. a per nielsen, danish data archives, niels bohrs al 1 e 25, dk-5230 odense m, denmark laine g. m. ruus, swedish social science data service, storgatan 13, s 411 24 goteborg, sweden don trees, the rand corporation, 1700 main street, santa monica, california 90406, u.s.a. ex officio treasurer : ed hanis, social science computing laboratory, university of western ontario, london, ontario n6a 5c2, canada editor: elizabeth stephenson, institute for social science research, university of colifornia, 405 hilgard avenue, los angeles, ca 90024, u.s.a. international association for social science information service and technology judith s. rowe, u.s.a. secretariat princeton university computer center u^s postage 87 prospect street paid princeton, new jersey p^^:,?js'tor: 08544 u.s.a. l'association internationale pour les services ft techniques dtnformation en sciences sociale 19.3 28 iassist quarterly holding hands or opening gateways? the uk was, and is, very fortunate in that an early commitment was made to providing a nationwide integrated academic computer network janet. this made it easier for successful national initiatives to be devised and implemented. however janet was based on x25 protocols, not the internet protocols known as “ip”. this hindered international integration it became clear that arguments over the relative merits of protocols were irrelevant because the internet was going to be the de facto standard for international networking. so the uk started a rapid transition to driving on the same side of the road, in networking terms, as the rest of the world. fortunately the existence of janet means that, for academic institutions, this transition will be completed in a relatively short time. in june 1992, some time before this transition was decided upon, but when it was already looking inevitable to many of us, i was appointed by the uk economic and social research council with a brief to support uk social scientists in the use of computer networked information. this was, and is, a very broad brief; and since there was not a precedent for a job of this kind i was to some extent improvising, making the rules up as i went along. the infrastructure for networked communication had been built by the technicians. so, initially, it was the technicians who tended to use it. it wasn’t until the infrastructure was there, and the user base expanded significantly away from the original designers and builders, that the way people were going to use it “in real life” began to emerge. it no longer seems surprising to us that that a medium designed for the rapid exchange of large data files for “serious” work is now most popularly used for exchanging short messages. it may well be that an infrastructure intended for file transfer of software and complex datasets will be used mainly for magazine publishing and the promotion and delivery of new service industries. as it has been with infrastructure so with information. whilst the providers and users were the same small band it was not clear what the problems or the potential were in this area. moreover, the provision of networked information has largely been the realm of the technical specialist, not the information specialist. there are clues here to understanding the subsequent rather uneven development in this area. it has become easier to provide information, and much development has gone into the interface for users, with tools such as mosaic and netscape, but the information processing side has lagged behind. i entered this arena at a time when there was clearly sprouting enthusiasm amongst social scientists who had previously regarded computers with fear and thought of email as just another way of increasing their workload. this enthusiasm was fed by global consciousness-raising in many fora and by my own small efforts. but often this enthusiasm did not progress from my visit it did not translate into “real work”. it was fine to go step by step through the maze with hand-holding documentation and a supportive guide but not so easy to navigate to unknown territory after a few weeks had elapsed. even for the brave there were more obstacles the origins of the infrastructure mentioned above meant that relevant social science information was scattered and sparse, often seeming to occur incidentally. my workshops and demonstrations were offering a glimpse of the possibilities rather than handing out a tool which could immediately increase the efficiency and productivity of researchers. it was all very well to look at all this fancy stuff, and the feedback from my sessions was always very good, but people still regarded this “internet stuff” as a plaything. as information provision slowly became easier, some academics were of course delighted to rediscover the joys of the second hand bookshop. spending an hour browsing the networks might, or might not, uncover the odd jewel amongst the dust and chaos. once found, the jewel could be copied or printed out and squirrelled away with the other printouts and photocopies. but only the most organised users made a note of their path as they went, so after a few days directing a colleague to rediscover the jewel might be impossible. even now with the facilities of browser programs one person’s hot list is another person’s cold shower. i’m not too interested in the schedule for evening classes in a college in the mid-west of the usa but it might be very useful to the right user. making some personal details about oneself available over the net might perform an important function to personalise and humanise discourse in a collaborative project where the participants have never met, for example. but i do not want to keep tripping over these details when i’m looking for something else. so system administrators, responding to complaints, hit upon the extraordinary idea of organising information according to subject headings. it took a little while for everyone to realise that librarians have been doing something similar for years, and by that time a host of idiosyncratic infant subject classification schemes were sprouting. in setting up the social science information gateway, we resolved to attempt at least to share the underlying classification system with sosig a move towards subject-based services. by nicky ferguson1 esrc visiting fellow in networked information social science information gateway project, university of bristol. england. 29fall 1995 other uk national service providers. the return of the cataloguer with the advent of client or browser software giving users a graphical interface to networked information, the possibilities for junk or vanity publishing seemed to expand dramatically, the idea of making pictures, text and sounds available across the world -do-it-yourself multi-media publishing was irresistible. combine this with the relative ease of creating html hypertext mark-up language, the building block of the world wide web, and you have an explosive combination. while the development of publishing was bounding ahead, the users of information were not so well provided for. browsing was more exciting instead of showing users meteorological data in tables, i could now bring up on their screens satellite photographs in glowing colour. but even with subject categories and fancy graphics, all we have really done is to give the second hand bookshop a facelift. you may know which shelf to look on, if you’re lucky some of the books may have glossy covers, but the essential problem of locating relevant and useful texts remains. one way we have tried to deal with this at sosig is by providing a searchable catalogue of information about each of the over 500 resource centres at which we point. this is quite different from the various so-called robots or automated search mechanisms which rely on highly resource intensive scouring of the networks and fairly crude automated examination of the resources themselves. we rely on human intervention to describe, classify and organise social science resource centres. in this way we also introduce an element of quality control. for each resource centre which appears anywhere on the subject menus, a form or template has been filled out this contains a description and keywords as well as appropriate technical information such as the url (network address) and the udc (classification) number assigned to that resource. the user can then search through this information using an on-screen form. a dynamic list of hits will then be returned listing appropriate resource centres, describing them and pointing directly to them. various options are provided and others (including boolean search options) will be added in the near future. roads to the future the ideal for such services is that they should be distributed so that centres of expertise are responsible for relevant subject areas. to answer the obvious question that this raises about our own activities, it is probably neither feasible nor desirable in the long term for us to attempt to take responsibility for describing and organising all the social science resource centres in the world, it is surely better for centres of excellence within the different social sciences to take responsibility for their own areas and for us to coordinate these efforts, but as a medium term solution the current sosig is certainly preferable to a totally centralised model. aiming for a distributed model, however, creates its own problems. it demands the ability to search across different servers which in turn implies that the resource descriptions will be in (preferably an internationally accepted) standard form. when the catalogue databases become large, as they undoubtedly will, manipulating the descriptions and templates will also become a problem if we rely on the relatively unsophisticated tools we use at present. in addition this system of describing and searching for networked resources should not be idiosyncratic it should be adaptable and aim for future integration with other resources such as opacs and citation indices. for these reasons, in collaboration with ukoln, the uk office for library and information networking at the university of bath, and loughborough university of technology, we have recently been funded to develop a system for allowing linked and geographically distributed resource discovery services to be set up. roads -resource organisation and discovery in subjectbased serviceswill allow users to search across different subject-based servers and will develop searching mechanisms based on emerging internet standards such as whois++. it will also investigate integration with other standards such as z39.50 and marc (in its various incarnations). as well as expanding the knowledge base and the capabilities of services such as sosig, roads will provide a packaged solution for information providers who wish to set up a subject-based service. we also hope to encourage centralised national service providers to focus their effort on the (initially many) subject areas not covered by these distributed services, so that a good coverage can be achieved in a relatively short time. thus roads will help to achieve the goal of a scaleable system for resource discovery, cataloguing, description, organisation and quality control. we have no illusions that roads will be a so-called killer application for networked information there will not be such an application, rather a number of different approaches will emerge and possibly merge. moreover sometimes a user will not find the obscure object of desire; or perhaps wishes to comprehensively survey networked resources on a topic without necessarily having regard to quality or currency; or to search across different languages and character sets. for these reasons, the roads partners intend to collaborate with european partners, not only to develop roads further, but also to develop complementary systems, including a comprehensive automated indexing system for european world wide web servers. thus, if the ergonomic nut crackers fail to break open the shell and reveal the kernel, we will provide the back-up of a well-designed hammer. this european collaborative proposal, codenamed desire, has recently been shortlisted for funding by the relevant european funding agency. 30 iassist quarterly we hope that all three of these initiatives will promote the design and building of subject-based information gateways (sbig’s), the implementation of which will result in a distributed resource discovery service based on rich descriptions and a quality controlled approach organised around subject centres of excellence. these efforts will be complemented by a comprehensive approach to european www index design, the implementation of which will result in a european discovery service based on automated indexing and an automated harvesting technology. references and further information 1 paper presented at iassist95 may 1995 quebec city, quebec, canada. bubl the bulletin board for libraries <url: http://www.bubl.bath.ac.uk/bubl/> figit the follett implementation group on information technology <url: http://lamin.bath.ac.uk/figit/figit-2-94.html> guardian (1995). online march 30, 1995 niss national information services and systems <url: http://www.niss.ac.uk/> roads resource organisation and discovery on subjectbased services <url: http://ukoln.bath.ac.uk/ukoln/roads/ roads.html> sosig social science information gateway <url: http://www.sosig.ac.uk/> ukoln the uk office for library and information networking <url: http://www.ukoln.bath.ac.uk/ukoln/> further reading may be found at: <url: http://ukoln.bath.ac.uk/ukoln/roads/ related.htm> vol21.4bacp 4 iassist quarterly a web-based archive of psychological experiments: challenges for client server interactions by ruediger oehlmann * abstract: the virtual psychology laboratory (vplab) is an ongoing project which will provide psychology educators, students, and researchers with a tool that supports the analysis, modification, and re-execution of previously archived experiments. the archive will include experimental materials, designs, procedures, and results which will be submitted by active researchers and educators. a user will be able to access vp-lab via the world wide web using one of the widely used web browsers. the paper identifies several requirements of providing psychological experiments in the world wide web. in addition, it discusses an approach to web programming which satisfies all the requirements. introduction with an increasing body of psychological research, archives became interested in providing their services to the psychological community. first attempts in this endeavor involved archiving abstracts of psychological research papers and developing a suitable thesaurus (american psychological association 1992, 1994). an approach which goes far beyond archiving abstracts is currently taken by the virtual psychology laboratory (vplab) project, a collaboration between the school of psychology at the university of cardiff, the data archive at the university of essex, and the computers in teaching initiative, centre for psychology at the university of york. it is the objective of the vp-lab project to provide facilities for archiving complete psychological experiments which will be submitted by active researchers and educators. the user will be able to access vp-lab via the world wide web using one of the widely available web browsers. the implementation of such a comprehensive service requires an understanding of the elements of psychological research. given that psychology is such a diverse discipline, it is not surprising that the field relies on a wide range of methods. most psychological research methods have the objective of answering empirical questions about behavior or experience by controlled observation. a basic instrument of the empirical approach is the experimental method (davis 1995, coolican 1994). we should note that in psychology as well as in other social sciences other methods are used as well. for example during interviews and surveys, typically the researcher has less control over the environment in which the investigation takes place. the control over the experiment can be viewed as the characteristic difference between the experimental method and other research methods. the advantage of the experimental method in psychology is that it is potentially capable of providing statistical quantifiable evidence for relationships between cause and effect. thus the control of variables is an essential aspect of any experiment. there are two categories of variables: independent and dependent variables. variables which the experimenter manipulates are referred to as independent variables, whilst variables which describe the experimental result are referred to as dependent variables. the statistical quantifiable evidence for relationships between cause and effect is reflected by these two types of variables. typically, such relationships are formulated as hypotheses which form part of a larger theory. by making observations, the psychologist can make a statistical judgement as to whether or not a prediction is correct; that is predictions can be tested against the evidence . the experimental hypothesis proposes that the change in the behavior as it is measured in the dependent variable(s) is actually caused by changes in conditions of the experimental setting. these changing conditions are described in terms of the independent variables. events which change the experimental setting in an objective way and which are capable of evoking a response from participants of the experiment are referred to as stimuli. usually, the evidence cannot be obtained from the total population to which the hypothesis might apply. therefore a sample of the relevant group is taken. assumptions need to be made concerning the representativeness of this sample. the selected members of the sample or participants are put into a controlled situation; i.e. a situation where all relevant independent variables are controlled by the experimenter and where the dependent variables can be measured. however, controlling the independent variables does not suppress the variability in people’s behavior. this variability will influence the measurements taken from the winter 1997 5 dependent variable. therefore the experimenter will be faced with a whole range of different scores by participants. the question to be decided is whether the differences in the scores are the result of manipulating the independent variables, or are the result of chance fluctuations in people’s performance as stated by the null hypothesis. in other words, are the score changes significant in support of the experimental hypothesis? the experimenter decides this question by performing an analysis of data which involves suitable statistical tests. a considerable amount of psychological research follows the general schema of the experimental method as it is outlined above. currently, researchers obtain information about previous experiments mainly from journal publications and conference proceedings. this method has two disadvantages: 1. identification and evaluation of relevant previous work can be very time consuming. 2. even journal publications often may not contain the degree of detail which is required to repeat or to modify a previous experiment. the aim of the vp-lab project to overcome these limitations is addressed by making all the details available in the world wide web which allows a comparatively fast access to archived information. in addition, this approach enables the archivist to provide more detailed information about an experiment than is possible in a journal paper. in this paper, we will argue that: 1. the web-based provision of detailed information about psychological experiments requires a) interactive content, b) secure access, c) interface-database connectivity, and d) platform independence. 2. given the current state of the art, the programming language java satisfies these requirements. in the remainder of this paper we will address these claims by inspecting the various components of a psychological experiment in detail (section 2). this inspection will indicate the requirements for a web-based provider of psychological experiments. in section 3, we will describe an approach to web programming which is based on the programming language java. finally, in section 4, we will discuss the merits of this approach from the perspective of providing executable psychological experiments on the world wide web. psychological experiments in this section, we will describe the various components which constitute a psychological experiment. from this description, we will then derive technical requirements which have to be satisfied if we wish to provide psychological experiments in the world wide web. typically, descriptions of psychological experiments identify the experimental components: materials, design, procedure, and results. before we discuss these components, we consider an experiment which will provide the basis for the subsequent sections. an example experiment the example we choose is an experiment1 which has been described by klein (1994). the experiment was based on previous work (stroop, 1935), in which the reaction of participants who had to identify the colors of ink they saw was measured. in the experimental condition, the participants were presented with words which represented colors such as the words red, blue, and yellow. however, the words were printed in a color which differed from the word’s meaning. for example, the word red was printed in the color blue. in the control condition, participants just had to identify color spots. the reaction time in the experimental condition was significantly longer than in the control condition. this was explained with an interference of the color identification by the processing of the word meaning. klein studied such inference by varying the meaning of the stimulus words. he compared conditions in which the word was either a nonsense syllable (e.g., bjb), a word that implied a color (e.g., grass), or the name for a color (e.g., red). then he measured the time his participants required to identify the color under each of these conditions. materials in the introduction, we have emphasized that changes of the experimental setting described in independent variables play a central role in any experiment. sometimes this setting involves a complete well designed environment. for example, a developmental psychologist might place a child in a play room with particular toys. often, changes are achieved by presenting participants with pre-defined text, image, or sound samples. this type of materials can be presented by using a computer with multi-media capabilities. for example, the word stimuli we described in the previous section can be displayed on a computer screen. in addition to these types of information, materials include computer programs to generate and present suitable multimedia files. the programs may be stand-alone programs written in a multi-purpose programming language, or they may have the form of script files which can be read and executed by a commercial experiment generator. 6 iassist quarterly the materials characterized in this section are different from data which can typically found in social science archives. first, stimuli which are contained in the materials are suitable to control the independent variables of an experiment. we have emphasized above that this aspect does not play such a central role in studies which are based on surveys rather than experiments. the need to control independent variables will have consequences for indexing stimuli. for example, a researcher who wishes to use a given stimulus in another experiment has to ensure that the independent variables in this experiment are controlled appropriately. therefore, the variables to be controlled have to be considered during retrieval of the stimuli. the second important difference to data usually found in social science archives is the need to archive computer programs. moreover, these programs have to be indexed in a suitable way. the support provided by current psychology thesauri such as apa’s thesaurus (american psychological association 1994) is very limited. moreover, the task of archiving programs raises the problem of maintaining programs over long periods of time. design for every experiment, it is crucial to group participants and to present stimuli in a way which avoids bias towards a particular experimental result. these general principles of the experimental setup in terms of conditions, participants, and variables used are described as design (kirk 1995, leon & austin 1996). for example the layout of klein’s experiment mentioned above can be characterized as a three levels of one-factor, between-subjects design. it is based on one independent variable, the word presented to the participant. the three levels are given by the experimenters choice between a nonsense word, an implied color word, and a color name. this design is referred to as between subjects design, because subjects or participants were randomly assigned to each condition without regard to the participants in other conditions, and each participant serves in only one condition. in contrast, a within-subject design would be a design in which the same subjects are repeatedly measured in different conditions, or each subject in one condition is matched with each subject in another condition. in addition to topical retrieval goals, such design characteristics could be used by vp-lab to retrieve an experiment. however, similar designs are used by a number of different experiments. so design characteristics alone will not uniquely describe a particular experiment, additional characteristics are required. procedure whilst the experimental design describes the general arrangements made for an experiment, the procedure provides the detailed steps to be followed in performing the experiment. therefore the procedure often follows a temporal sequence. they begin with summarizing the instructions given to the subjects and proceed through the tasks performed by the subjects in the order in which they were performed. results during the experiment all the responses of the subjects are carefully recorded. often this type of information may be obtained automatically using computers. these raw data are then analyzed to determine the significance of the result. the statistical method used for data analysis is determined by the chosen experimental design type. for example, the method used by klein in his experiment of the stroop effect is the one way, between-subjects analysis of variance. this method should be employed in a between-subject design with two or more levels of one factor. the method of analysis of variance is a procedure which enables the experimenter to determine whether significant differences exist in an experiment involving two or more sample means (greene & oliveira 1995). such statistical methods can be used to index an experiment. a student may later retrieve the experimental data as example data for the given analysis method. requirements: i have provided information about the typical components of a psychological experiment because this provides important indications of the requirements a web-based experiment provider has to address. interactivity interacting with stimuli may not be just a two step process of presenting stimuli and recording the response; it may be a sequence of interactive steps. the presentation of a stimulus may even depend on a previous response. furthermore, a participant might be required to interact with a specified part of a stimulus such as a particular area in an image. therefore the presentation has to include executable content. this is a feature which allows different responses and supports different reactions of the stimulus depending on the response. platform independence typically, an experiment is developed by using a particular computer system, usually one known or accessible to the researchers. there is no need for them to consider portability to other systems. however, a web-based experiment will be used by a large number of users with different computers and operating systems. interface-database connectivity the various experimental components have to be stored in a database. rather than just addressing the issues of interactivity and platform independence, a web-based provider of psychological experiments has to address the question of how to connect an interactive and platform independent browser with the database. a researcher who interacts with vp-lab may be merely interested in a particular type of stimuli rather than a complete experiment. therefore the connection has to be achieved in winter 1997 7 a flexible way which supports the interaction between user interface and single experimental components. security in all web applications, security issues ned to be taken very seriously because networked computer systems are by definition more vulnerable to attacks. in addition to these general considerations, providers of psychological experiments have to ensure the confidentiality of participants. often psychological data obtained from a single participant reveal highly personal information. this problem appears only to a smaller degree in journal publications because they contain the results of the data analysis rather than the raw data. in addition to the protection of the rights of participants, we have to consider the protection of the rights of the researcher who deposits the experiment. sometimes in the course of research, highly sophisticated software has been developed over a long period. the researcher may wish to distribute the experiment in terms of stimuli, design, procedures and results, whilst restricting the distribution of the software for generating the stimuli. providing psychological experiments via the web the requirement that we have identified in the previous section can be addressed by programs written in the programming language java. interactivity java is a programming language developed in 1995 by sun microsystems (gosling, joy, & steele 1997). it is used to create executable content which can be distributed through networks. the java concept distinguishes between two types of code; stand-alone programs referred to as application and pieces of code that are linked to a web page and sent as executable content through the internet. this type of program is referred to as applet. in order to view java content on the web, a user’s browser must be java enabled; i.e., the browser has to be integrated with a java interpreter. an increasing number of web browsers have this feature. an early example of this type of browser is netscape navigator. the content downloaded by a web browser can include a variety of multimedia documents. if the browser receives a user request, it downloads content that describes a web page (figure 1). the web page can contain a particular hypertext tag called applet. when downloading a web page containing an applet tag, the java-enabled browser knows that a special kind of java program called an applet is associated with that web page. the browser than downloads another file of information, as named in an attribute of the applet tag, that describes the execution of that applet. this file of information is written in what are called bytecodes. the java-enabled browser interprets these bytecodes and runs them as an executable program on the user’s computer. the resulting execution then drives the animation, interaction, or further communication which is again displayed by the browser. the overall pattern for the use of content is selection of content by the user, downloading, executing, and displaying of content by the browser. this process in itself already provides considerable support for interactions between a user and psychological experiments described as executable content. however, java has two other supportive features: it is object-oriented and multithreaded. the term objectoriented means that components of an experiment can be represented as separate objects (classes in the java terminology) which store the functionality of the components in the form of methods. this concept supports a highly modular approach because changes in one experimental component could be made independently of other components. moreover, experimental components would be the building blocks of an experiment on the implementational level rather than just on the conceptual level. the term multithreaded refers to a pseudo-parallel approach. typically, computers have only a single processor; therefore parallel execution of several programs in the strict sense is not possible. however, java programs can direct the processor in those parts of the program to be executed next. these parts can be very small and a fast switch from one part to the other gives the impression that these parts are executed in parallel. therefore multithreading also supports rapid interactions between a user and a pre-stored experiment. platform independence we pointed out that java applets are transported over the internet in a particular form which we referred to as bytecode. every computer that can execute java code has a machine specific interpreter that translates bytecode into machine code. this is the reason why java programs are machine independent; i.e. the same program can be executed on a unix system, a pc, or a macintosh computer. interface-database connectivity the java language includes several tools that extend the language to different tasks. one of these tools supports connections between java code and relational database systems. this tool is referred to as java database connectivity (jdbc) and defines every aspect of making data-aware java applications and applets (jepson 1996, patel & moss 1996). in using this tool, the developer does not need to be concerned about the database-specific syntax when connecting to and querying different databases. another advantage of this approach is that changes to the applet code can be minimized when the database system is changed. the jdbc concept has recently been extended to a java based three-tier architecture. typical client server interactions are based on two types of computer systems: a central server that maintains the database and a number of clients that maintain the user interface to the database. in a three-tier architecture this concept is enhanced by an 8 iassist quarterly additional middleware server that connects with the database server on one side and with a number of clients on the other (symantec 1996). this approach has several advantages. for example, the client systems are easy to manage because no application or database software is required to be installed on the client side. any system with a java enabled web browser will work as a client with no additional software. the system is easy to program because all code is executed on the client system. the developer of application programs accesses neither the middleware server nor the database server. security it is the basic principle of the applet approach that java code is downloaded to and executed by the client computer. this involves a considerable potential security risk. therefore the java language includes the most sophisticated security system of any programming language (breedlove et al. 1996, morrison et al. 1996). the complexity of this system is clearly beyond the scope of this paper. however we will discuss some of the basic ideas. java programs can be viewed as a set of classes which in turn can be built-in or user-defined. we focus on the security check for user-defined classes. each class is checked by three different sub-systems: the java verifier, the class loader, and the security manager. the java verifier is used to check the bytecode to make sure that the safety features of the java language are followed. after incoming code has been checked by the verifier, the protections in the java class loader are invoked. an important function of the class loader is to avoid overriding of filesystem source classes. it also makes it impossible for a file system source class to access a network source class by accident. finally, the security manager is used to provide a flexible access control mechanism. any time a figure1: java’s view of client-server interaction winter 1997 9 non-built-in class accesses a system resource such as a file, it must first ask permission from the security manager. ordinary classes are prevented from bypassing the security manager because none of the system resources are available to them. conclusions in the introduction to this paper, we have described the objective of the vp-lab project as to provide an archive of psychological experiments which can be performed over the world wide web. in context of this task we have argued that 1. the web-based provision of detailed information about psychological experiments requires a) interactive content, b) secure access, c) interface-database connectivity, and d) platform independence. 2. given the current state of the art, the programming language java satisfies these requirements. we have considered typical components of psychological experiments to illustrate the need for interactive content, platform independence, database connectivity, and secure access. finally, we have addressed these requirements from the perspective of the general purpose programming language java. we will now evaluate this perspective in terms of the identified requirements. this evaluation will be biased because our viewpoint is how useful the language is for supporting the provision of psychological experiments rather than java’s general usefulness for other types of applications. nevertheless, several of our requirements can be regarded as general requirements for numerous web applications. interactivity java code is downloaded rather than executed on a web server. therefore performing a psychological experiment will not depend on the load imposed by other users on the web server. it just depends on the capabilities of the client computer. downloading java applets can be used to modify the interface of the web browser. this has two advantages. first, vp-lab can integrate a standard web browser with a state-of-the-art interface for browsing its archive and supporting the psychologist in various experimentation tasks. second, in performing a pre-stored experiment, the interface can be changed automatically to the type of screen used in the original experiment. platform independence java programs are executed by the client machine; therefore a very important issue is whether they can be executed by machines of different types under different operating systems. otherwise, the usability of these programs would be very limited. java applets are already available in the internet and their platform independence can be tested. the same applet can be downloaded and executed by a pc, a macintosh, and a unix system. in addition, an increasing number of web browsers have become java-enabled. interface-database connectivity the jdbc-based three-tier approach reduces the programming effort for the actual link between the application program and database because a considerable proportion of this link is already provided by the middleware server and needs only to be adapted to the particular application. this adaptation is performed on the client side rather than on the side of the middleware server or database server. in addition, the middleware server is embedded in a high-level development environment that reduces the programming effort further. security java controls have by definition no capabilities to read or write to the local file system or to make any operating system calls. furthermore, java uses a runtime verification of its code. using bytecode has two advantages from the security perspective. bytecode can be checked for security violations and it allows accessing of methods and variables by name rather than by number. this makes it easier to determine what is being used and to protect from misuse. all these considerations led to the decision to implement vp-lab in java. we should note however that vp-lab is a short-term project that requires a decision about facilities which are immediately available. in a few years time the situation may change. however, we expect that future web programming languages will incorporate java concepts such as machine independence and multi-layer security mechanisms. however, even more important than these technical aspects will be the fact that psychologists do not need to be aware that they are interacting with java applets or with the world wide web. all they will see is an archive of experimental materials, designs, procedures and results. references american psychological association (1992). psycinfo psychological abstracts information services users reference manual. washington, dc: author american psychological association (1994). thesaurus of psychological index terms. 7th ed. washington, dc: author. 10 iassist quarterly breedlove, b. et al. (1996). web programming unleashed. indianapolis, in: sams publishing. coolican, h. (1994). research methods and statistics in psychology. 2nd ed. london: hodder and stoughton. davis, a. (1995). the experimental method in psychology. in g. m. breakwell, s. hammond, and c. fife-schaw, research methods in psychology. london: sage publications. gosling, j., joy, b., & steele, g. (1997). the java language specification. reading, ma: addison-wesley. greene, j. & d’oliveira, m. (1995). learning to use statistical tests in psychology. milton keynes, uk: open university press. heiman, g. (1995). research methods in psychology. boston, ma: houghton mifflin company. jepson, b. (1986). java database programming. new york: john wiley. kirk, r. (1995). experimental design. procedures for the behavioral sciences. 3rd ed. pacific grove, ca: brooks/ cole klein, g. s. (1964). semantic power measured through the interference of words with color-naming. american journal of psychology, 77, 576-588. leong, f. & austin, j. (1996). the psychology research handbook. thousand oaks, ca: sage publications. morrison, m. (1996). java unleashed. indianapolis, in: sams publishing. patel, p. & moss, k. (1996). java database programming with jdbc. scottsdale, az: coriolis group books. stroop, j. r. (1935). studies of interference in serial verbal reactions. journal of experimental psychology, 18, 643662. symantec (1996). evaluating network database architecture. white paper. author. 1 this description follows closely a review given by heiman (1995). * paper presented at iassist/ifdo ‘97, odense, denmark, may 6-9,1997. ruediger oehlmann,university of essex, the data archive, psychology unit colchester co4 3sq, uk oehlmann@essex.ac.uk mailto:oehlmann@essex.ac.uk vol18172 4 iassist quarterly introduction twenty-five years ago, i wrote a proposal to the national science foundation for a workshop on the management of a data and program library to promote the establishment of local university social science data archives. although organized in only a few weeks, the workshop attracted 100 people from 50 different institutions in the u.s. and canada. we even had one person from sweden, indicating the deep roots of swedish aarchival data activity. a few institutions already had data archives, and within a very short time, all did. in preparation for my remarks here today, i went back to the proceedings of the workshop we published shortly afterward2. in the “preface” to the proceedings, i outlined what i thought were the benefits of establishing a data archive. for faculty, the benefits appeared to be greater productivity, the potential for investigating new kinds of problems, and the ability for scholars with limited funding and resources to access high quality data. graduate students could be given a greater opportunity to gain research experience, and they shared with faculty the possibilities of examining new kinds of problems and high quality data. however, i reserved my greatest expectation of benefits for undergraduates. because of the expense and time required for empirical social research, undergraduates in the sixties rarely experienced a genuine introduction to the social sciences as research disciplines. instead they read brief and, often, highly simplified synopses of research that gave little indication of how quantitative social science is done. the ready availability of data and program libraries would, i thought, make possible realistic introductions to the social sciences. looking back now more than two decades, i think it fair to say that data archives have made a significant difference to faculty and graduate students and have not had much impact on undergraduates. for faculty and graduate students, data archives have vastly increased the amount of comparative and over-time research published. better data also fostered greater sophistication in social scientific theories and analytical techniques. twentyfive years ago, economists aside, most social scientists were content with theories based upon cross-sectional data, cross-tabulations, ordinary least squares, and path analytic techniques which required no more than ordinary least squares. we now find publications and dissertations dealing with dynamic theories based upon event history analysis, partial adjustment models, dynamic lisrel type models, and more3. most of this work was made possible by the collections of data archives. however, the impact of data archives on undergraduate education has been modest. i will acknowledge that data archives significantly altered patterns of instruction in undergraduate methods and statistics courses, but in most colleges and universities, these are “required” courses segregated from the substantive foci of their disciplines. in all too many cases, the links between these courses and substantive courses are left to the imagination of the students. the vast majority of undergraduate courses today are taught in a manner little different than decades ago. faculty lecture about research that students find summarized in their texts, and the investigatory process that produced the results about which students read is about as much an enigma today as it was then. as a result, the skills taught in the methods and statistics quickly grow stale. ironically, the achievements of data archives in enriching faculty and graduate student work often has led to an institutional success which works against serving undergraduate education. twenty years ago, data libraries had little or nothing to do with conventional libraries. they were creations of faculty members seeking to provide research and instructional resources for themselves, their colleagues, and their students, and they were housed in faculty offices or a few rooms down the hall. by and large, librarians did not understand computers and data and often were uninterested. much has changed. at contemporary meetings of inter-university consortium for political and social research official representatives, there seem to be as many librarians as faculty—certainly they form a substantial fraction of those attending the meetings. libraries are moving to accept data archives as important parts of their collections—a change that represents institutional commitments to data libraries as scholarly resources. what, then, is the irony? it is that even as libraries take on this new function, university and collegial support for libraries is either static or declining. data services since the 1960s: where are we going? by david elesh1 center for public policy, social science data library, temple university, 5spring/summer 1994 from 1975-76 to 1985-86, current fund expenditures for libraries in institutions of higher education fell from 3.1 percent of total expenditures to 2.6 percent4; and this occured during a period in which library costs for books and periodicals increased roughly 60 percentent faster than overall inflation. quite simply, the cost pressures universities have faced for more than a decade show no signs of improving soon, despite all the professed concern about the state of american education. this means that libraries are unlikely to have the kinds of staff resources required to help undergraduates use data resources. in fact, there is the real possibility that all users will suffer. but the picture is not completely bleak. much is changing in social scientific instruction, and an increasing number of undergraduates are gaining experience in doing quantitative social science. the introduction of microcomputers is slowly—very slowly—transforming undergraduate education. texts increasingly come complete with analytical software and databases, and independent instructional packages and databases such as showcase have been adopted widely. both types of innovations make it possible to introduce real, quantitative social scientific work to both lower and upper division undergraduates and usually find enthusiastic acceptance. but neither involves or leads to use of data archives. why? first, it is important to recognize that data archives were conceived in an era of mainframe computing and were meant to be analyzed by mainframe statistical packages. the fundamental meaning of this statement is that the knowledge base required to use these resources is simply much larger than for a pc. the architecture, organization, and funding of mainframe computing are designed for the researcher, not the instructor, and certainly not, excepting computer science students, the student. mainframe operating systems are far more sophisticated than those available on pcs. mainframe statistical packages, while powerful, are typically intimidating in their complexity, and even those of us who routinely have required undergraduates to learn these packages sufficiently to get through our methods and statistics courses know that we must sacrifice some content to allow time for teaching basic computing skills. the learning curve for mainframe computing is a great deal steeper than for pcs. in the past, we could defend the loss of statistical or methodological subject matter in the belief that knowledge of spss or sas and the like formed part of the research skills we were trying to impart. those of use who used the computer in our undergraduate instruction learned to create program and/or system files that shortcut many of the procedures we expected graduate students to learn. we used archived files because there were few alternatives, and setting a file up for a class was little different than setting one up for our own research use. but the skill level and time it requires to do these things are significant, and many social scientists simply did not and still do not have them. nor would they or their students find much help in the organization of computing. because the machine or machines were located centrally, it was and is almost universally true that consultants were as well. one had to go to the computing center to use computers or seek assistance in using them. at the same time, the available consultants were and are typically programmers unfamiliar with statistical software and social science data. to a very large extent, users must learn the consultants’ language in order to obtain help; they do not learn users’ language. a few consultants might learn spss, sas, bmdp, or the other statistical packages, but the demand for their services always exceeds the supply because social science users of the computer are greatly outnumbered by users in computer science, engineering, and the physical sciences, and the latter have the influence that attends greater external funding; thus central computing budgets favor the latter over the former. local data archives often tried to fill the void by offering assistance in use of the computer as well as of the data. staff became expert in the use of tapes and the manipulation of large and complex files; in some institutions, they provide and have provided the primary consultative assistance in these areas. against this background, it is not surprising that analyses of archived data did not spread widely in undergraduate instruction. while i think there is little doubt that the introduction of microcomputers can transform undergraduate instruction, i have some doubt that data libraries will be significant actors in the transformation. micros eventually will succeed in changing social scientific instruction because they significantly lower the slope of the learning curve for computing. students find pc operating systems easier to learn, and unlike the mainframe world, there are statistical packages specifically designed for instructional use which require far less faculty and student time to learn. however, generally these packages incorporate data which has been tailored for them, and the tools necessary to include other data sets are omitted. one can even find software that allows the student to place a disk in the lowliest pc, turn it on, and find him or herself in a menu driven analytical package capable of multivariate crosstabulations on a substantial number of variables 6 iassist quarterly with adequate samples and with virtually instantaneous response. microcomputers also introduced a new market structure for computing and data. in the mainframe environments, computing is provided and funded centrally as a university or college function, and instructional costs are, at least partially, borne by tuition. data files are also provided centrally and usually cost users nothing. however, with the introduction of microcomputers, the cost of the hardware, software, and data are increasingly being borne by the user as direct charges. where universities or colleges supply microcomputing laboratories, the number of these institutions that have introduced “laboratory” fees to cover these costs grows with each passing year. and, as noted, software and data increasingly come either from text publishers or other third party vendors. clearly, publishers are seeking to make it significantly easier for students and faculty to analyze data. clearly, too, if my history of the past quarter century or so is correct, greater ease-of-use is necessary if data analysis is to spread to substantive subjects. although i have no hard evidence, i suspect that the effort to produce greater ease of use is producing instructional software that is increasingly valuable for research purposes—the development of analytical graphical displays is one example— which is an interesting reversal of direction for the traditional flow of technology. it is possible for data libraries to participate in this transformation, but they will have to change their traditional modes of operation in several ways. first, they will have to work with faculty members to identify analytical software that is easy-to-use and capable of analyzing and presenting data in a way that the faculty member finds useful. second, they will have to create files for that software. typically, this will mean creating files on a mainframe, exporting them in ascii, downloading them to a micro, and modifying them for the analytical program. the program may be a statistical package, a spreadsheet, a graphics program, a database program, or something else. the choices are larger in the micro world, and faculty demands are and can be expected to be diverse. third, data libraries should attempt to develop expertise in exemplars of a number of software types— e.g., statistical packages, databases, spreadsheets, graphics—because it will be necessary if they are to be able to provide support for the files they create and because faculty are likely to ask for recommendations. fourth, as networks expand and take on some of the functions of mainframes, data libraries will have to learn how to create and maintain data servers for users at all levels of sophistication. fifth, data archives should look to the creation of display-formatted tables resident as files on disk as reference works for their most heavily utilized files. while some tables on many subjects will be available on cd-roms from a number of vendors, it should be possible to create tables from archival holdings that serve the needs of particular programs at a cost significantly lower than would be required to manufacture a cdrom; software exists to compress such files and expand them as they are called by programs. all of these possibilities for data archives require new investments—albeit at a relatively modest level—at a time when funding for new ventures is difficult. given the cost pressures higher education now faces and will likely to face during the next decade, it is more likely that funds for these initiatives will come from a more efficient utilization of existing resources than from new ones. one is supposed to close discussions of the future on an optimistic note, and i will try to do so. the transformation of computing offers substantial opportunities for using the data in our archives more broadly. we can move beyond our traditional support of faculty and graduate student research to make more of an impact on undergraduate instruction. but it will take initiative and a careful marshalling of resources. otherwise, the past is, at best, all too likely to be prologue. 1.paper presented at iassist 1990 in poughkeepsie, new york. 2.workshop on the management of a data and program library, proceedings eds magaret oneil adams, david elesh, and alice robbins. madison, wi, 1990. 3. i do not wish to re-open old, and typically, fruitless, debates about the relationship between theory and empirical research. i simply wish to note that neither theory nor research was much concerned with dynamic relationships in the 1960s. 4. u.s. office of education, digest of educational statistics, washington, d.c., government printing office, 1990, p. 301. 19.3 4 iassist quarterly a variety of forces have converged to promote research based on secondary analysis of existing social science data sets. a growing number of federal agencies have officially encouraged data sharing by requesting, often requiring, that their grantees place data sets collected with public funds in the public domain. at the same time, declines in university and federal research budgets have put expensive primary data collection out of the reach of many social scientists. advances in microcomputer technology have allowed powerful data analyses to be performed quickly and economically. data archive centers dedicated to the preparation of data sets for public use have also begun to emerge. several challenges await the data archivist working to provide users with clean, useable data for secondary analysis. the user must be provided with data of high quality, documented in clear and comprehensive fashion. while paper documentation retains its value, paperless (electronic) documentation is becoming increasingly important, in light of burgeoning use of the internet and mass storage media such as the cd-rom. hand-in-hand with the growth of the national movement toward secondary data analysis comes the need for powerful yet user-friendly ways to search through the massive amounts of available data and documentation to retrieve studies or variables of interest. to reduce hard-disk storage burdens and statistical analysis time, such search and retrieval of variables would ideally be linked with data extract capabilities, so that analysis files containing only variables or cases of interest to the user can be created on demand. this article documents the latest advancements of one data archive center, sociometrics corporation, in meeting these challenges. the sociometrics data library currently houses five topically-focused data archives: the data archive on adolescent pregnancy and pregnancy prevention, the american family data archive, the data archive of social research on aging, the maternal drug abuse data archive, and the aids/std data archive. together, these five data archives include over 200 data sets, which have been chosen for technical quality, scientific merit, substantive utility, relevance to social policy, demand for secondary data analysis, and disciplinary balance by a panel of experts in each archive’s substantive field. each data set in each topically-focused archive is made publicly available with a printed and bound user’s guide and a standard set of machine-readable files—raw data, spss and sas program statements that fully document the variables and values in the data file, an spss dictionary, and spss frequencies—with the explicit goal of providing the user with clear documentation and ready-to-use data files (see card and mckean, 1993 for a discussion of standard file preparation). this paper will focus on the recent development of three software features accompanying the data sets which simplify secondary analysis for social scientists. while the examples used will be drawn from the aids/std data archive, the software described is generic to the data archives comprising the sociometrics data library. software to facilitate the selection and analysis of variables search and retrieval software. the first feature, search and retrieval software, allows the user to examine the contents of an entire topically-focused data archive and retrieve variables using a variety of search strategies: (1) searches by full-text keyword, including variable names, variable labels (question descriptors), and value labels (response descriptors); (2) searches by substantive topic and type codes that have been assigned to each variable during the archiving process; and (3) searches by study name, author, or assigned data set number. standard boolean operators such as “and,” “or,” and “not” can be used to conduct any search. electronic instrument-variable link. the second feature, an electronic link between study variables and graphic images of the data collection questionnaire, allows the analyst to select variables and view the instrument pages associated with the selected variables. alternatively, the user may browse through the entire collection of instruments page-by-page, or search the instrument database by substantive keywords. the instrument-variable link allows analysts to examine questionnaire skip patterns and item context on-screen, a process which enhances the variable selection process and reduces the need for paper copies of instruments. data extracting. the third feature, data extract software, allows the user to produce, with a few keystrokes, spss or sas program statements for any subset of selected variables and then create an active or system file for statistical analysis. this the development of software to facilitate use of archived data sets by josefina j. card and elizabeth a. mckean1 sociometrics corporation 5fall 1995 capability permits analysis of even the largest of data sets to be conducted on most microcomputers. it also saves significant preparation time in writing and re-writing spss and sas program statements to define variables used in a given analysis. in the following section, the capabilities of these search & retrieval, instrument link, and data extract software programs will be illustrated with a sequence of searches from the aids/std data and instrument archive, which contains over 14,000 variables from 11 major investigations of the incidence and prevalence of specific sexual behaviors, contraceptive use and std preventive behavior, aids/std knowledge, and attitudes regarding contraception and std prophylaxis. search and retrieval of variables from the aids/std data and instrument archive in the first example, a search by topic across the 11 studies comprising the aids/std archive illustrates how the search and retrieval software can be used to find all variables indexed under the substantive topic of hiv/aids. after initiating the program and pressing the f3 select menu to request a search by topic, the f4 search menu is used to perform the search using the wildcard query topic? (figure 1). figure 1 wildcard searches such as the topic search requested in figure 1 produce a scrollable menu of retrieved items. in this example, the 62 substantive topics available in the aids/std archive are displayed in the menu. the third of the six screens comprising this menu is shown in figure 2. figure 2 shows that 630 variables in the 11 studies comprising the aids/std archive have been indexed under the topic of hiv/aids. descriptions of each of these “hit” variables can be viewed by highlighting the hiv/aids topic line in the scrollable menu and pressing enter. variables records for all 630 “hit” variables will be retrieved, and the first variable record in the set will be displayed on the screen. figure 3 shows the first variable record from the topic = hiv/aids search set. as figure 3 shows, the variable text record provides the variable name and label (line 1), study name (line 3), author or investigator names (line 4), topic and type codes assigned to the variable (lines 5 and 6), and value labels (line 7 ff.). the text record also contains the following on-screen instructions (line 2) for viewing the instrument page containing the original item from which the variable was derived: “press alt + i to see image of actual questionnaire.” in figure 4, the instrument page for hvb01021: a:7a been to hiv testing site is shown in the upper left corner of the screen. the lower right portion of the figure 4 screen contains instructions for zoom enlargement of the instrument page image, which can be 6 iassist quarterly figure 2 magnified to one of three sizes and printed out. in figure 5 the instrument page for variable hvb01021 is magnified at zoom level 2 and has been cropped to fit the page. to allow the user to quickly locate the original item on the graphic page image, all variable labels include the questionnaire item number. in this example, variable hvb01021: a:7a been to hiv testing site, is from item 7.a. on the “a” questionnaire from data set 01, the california survey of aids knowledge, attitudes and behavior: 1987. figure 5 shows that item 7.a. is figure 3 7fall 1995 figure 4 part of a compound question assessing aids-related behaviors. because one of the primary goals of a search and retrieval session is to select a subset of variables for analysis, the next step in our example will show how a search set of demographic variables including age, race, sex, and community of residence can be combined with the set of 630 hiv/aids variables already retrieved, in order to produce a prototypic set of analysis variables. under the f4 search by topic function, the user can select and retrieve multiple topics at one time with the f9 (group ÷ ) key. our second topic search retrieves all variables in the aids/std archive that assess age, race or ethnicity, gender or gender role, neighborhood or community, and region or state — a total of 717 variables indexed under these five different topics. figure 6 shows the fifth of a set of six topic-menu screens, showing how such selection was done for two of the five topics (race/ethnicity and region/state). the 717 “hit” variables from this second search covering five demographic topics are combined with 630 hit variables from the first hiv/aids-topic search using the boolean operator or (figure 7). a new set of 1,347 variables results (figure 8, “set 3”). to save the selection of demographic and hiv/aids variables comprising set 3 in a file that can be used by sociometrics’ data extract software, the user simply selects the option “transport a set” from the f5 sets menu and provides a filename for the transported set (figure 8). transforming a variable search set into a statistical analysis package command file the data extract software uses the transported set file to create extract command files in user’s choice of the spss or sas statistical analysis package. figure 9 shows the on-screen summary produced after the sample transport file has been read. each study from the aids/std data and instrument archive contains some of the demographic and hiv/aids variables from the search. 8 iassist quarterly figure 5 extract command files may be produced for each data set in turn. the program requires between 30 seconds and 3 minutes to produce an ascii command file that will create an spss/pc+ , spssx, or sas active file with variable names, variable labels, and value labels. figure 10 shows a sample extract spss-pc+ command file for data set 01, the california survey of aids knowledge, attitudes, and behavior: 1987. to save space, only a sample of variables from this file is presented; lines edited out of the spss-pc+ command file in figure 10 are noted with a series of dots (....). the resulting extract command file can be used to figure 6 9fall 1995 figure 7 create an spss/pc+ active file, or with minor editing, a system file on which analyses of the 100 age, race, gender/gender role, neighborhood/community, region/state, and hiv/aids variables in the california survey of aids knowledge, attitudes, and behavior: 1987 can be performed with ease. next step the last decade has witnessed paradigmatic changes in the way social science data sets are stored, delivered to users, and analyzed. many of these changes have been brought about by rapid technological developments in microcomputer hardware and software, optical storage devices, and the growth of the internet and on-line digital libraries. this paper has briefly described one state-of-the-art data archive consisting of high quality data in a field of current interest. the aids/std data and instrument archive is representative of all data archives produced at sociometrics, which combine high quality, figure 8 10 iassist quarterly paperless documentation with powerful, yet easy-to-use software for variable search and retrieval, viewing of the original questionnaire page containing the item, and data extraction. as the need for data sets for secondary analysis grows, social scientists will demand more sophisticated methods of data storage and retrieval. data archivists will continue to face the challenge of providing users with increasingly well-prepared, paperless, and machine-searchable data collections. in anticipation of user needs, sociometrics is developing several new enhancements for future data archives. first, programming is being developed for electronic links between variables and descriptive statistics as well as technical notes such as skip logic, scale variables, case weights, and other study-specific information associated with each variable in an archive. because of the overwhelming amount of information that typically accompanies a data set, it is important that accurate and complete documentation at the level of the individual variable also be accompanied by print-on-demand capabilities. providing print-on-demand capabilities for user-selected portions of documentation (e.g., user guide sections, questionnaire pages) in a variety of character formats including ascii, microsoft word, wordperfect and postscript, constitutes a second forthcoming technological enhancement. finally, sociometrics is preparing figure 9 to launch socionet, an on-line data library service, available 24 hours a day to internet users. it is hoped that these developments will move the field of data sharing and secondary data analysis ever closer to the vision of the successful enterprise-of-the-future described by stanley m. davis in his book future perfect: the ability to meet users’ needs—in this case their data information needs—any time, any place, any where (davis, 1987). references card, j. j. and mckean, e. a. (1993). harnessing advances in technology: new opportunities for research and teaching in social gerontology. gerontology & geriatrics education, 14,2, 63-76. davis, stanley m. (1987). future perfect. reading, massachusetts: addison-wesley. 1. we wish to thank our colleagues, drs. kathryn muller, bill farrell, and eric lang, for their comments on an earlier draft of this paper. 11fall 1995 lassist newsletter, vcl. 2, nc. 2 (spring 1978) ecitorifll comment if you wish to make suggestions for review materiai., please notify kathleen, or if you wish to write an over vie w reviews, also let her know. for d variety of reasons no reviews apthe spring issue, 1s78, of the geared in volume 2, number 1, nor iassist lewsletter contains three ao any reviews appear m the presaretcies arawn from the papers ent issue. we should be back on given at the meeting in itasca in track by number 3, however.. february. the three are related in that they each are directed to some of the problems and opportunities inherent in the archival enteroub mistake prise. beginning with a discussion of the need for cataloging standalice hobbin informed us that a ards, sue dodd (university of north strange type of error occurred in carolina) sets forth some of the the first paragraph of her article criteria implicit in the need for a (vol. 2, no. if. it makes no sense means of retrieving and using mato say, "computer tecnnology has chine readable datasets in a meanmade it possible for the tradiingful manner. viewing the arcaiticnal library to service the user val problem from an cr ganiza tional more quickly and (many enthusiasts perspective, william gammell (uniof on-line data bases would add) versity of connecticut) describes sore compli cated than when the refthe goals and objectives of the new erence librarian relied on manual organizational structure of the methods for searching and retrievroper center, inc. one of the priing information." the correct exmary points empnasized by gammell pression should be "more fully." is the continued desire of the one typo was, apparently, not sufhoper center staff to improve docuficient, for in the final paramentation standards sc that the argraph, instead of ". . . networking chive will have expanded usefulcreates the potential for independness. many of the problems facing ence for computing centers . .," the roper center (and many of the it should read. ". . . networking opportunities as well), are characcreates the potential for independteristic of large, general purpose ence ^om computing centers." archives, sucn as foper or the icpsr. cn a more restricted level. we might note that proofreading yet facing many of the same probis a difficult and arduous task, lems, are special purpose arcnives and while we make every erfort to such as that maintained by ncrc at provide clean copy, it is probably the university of chicago. in his inevitable that some errors occur, paper, patrick bova (nofc) seeks to while we will continue to try to aelineate some of the problems and improve our track record, i senopportunities available to special ously doubt that this, or any other purpose archives, and to describe publication will achieve 100% accuthe services the noec library can racy. please bear with us, howmake available to the survey reever, search community. sources of papers book review editor in issues no. 1 and no. 2 o: during the current year ms. volume 2, the papers published m kathleen heim will continue to the new sle tter were extracted from serve as book review editor for the those given at tae itasca confernewsietter. kathleen served ably ence in february. because papers auring~i:"ee editorship of alice robgiven at meetings are or ten not bin and has consented to continue polished products, several of tae m that capacity. readers cf the completed articles have required newsletter should ce aware that as. occassional (and sometimes extenhermts~'aa3ress has changed frcm the sive) copv editing. ordinariiiy we university of wisconsin to: would have returnea copy edited papers to the author (s) for review, ms. kathleen m. fieim but time exigencies for the first graduate school cf library two issues have precluded that posscience sibility. in doing the copy edituniversity of illinois at ing we have attempted to stay— as urbana-champaign much as possible— with the author's 329 mam library original meaning and phraseology. urbana, il 6i8oi i hope we have succeeded. to the present time, however, we r.ave not i2 :^.". .^tsl nfc wsi^ tt-j l , vcl. i (spling 197d) atis . can oi ter fro ir o the the wou lt;t co thti t m 1 ffl e eu ir id ud ai.y drticies spt t; a zor tht ijt^^sielt eirdininj two iesukji tcl m ttie current ntiuue to uublidh ; fapcts iro:r, itasca, he upl-sdia meeting, t. if any authors ither the north ac ropean meetings cart pacers ana subirit invite them to do sc o t jar, .ect in d, pa : f-i :i ca dl.y for the we ions afpers pers n or e vse we forthcoiiing deadlines i caii your attention to tne ract that tae deadlines for tae summer and fail issues of tae newsletter are rapidly approaching. if you have any materials roc the newsletter piease send them as early as possible. the deadlines for nos . 3 and 4 are as follows; number j.: number u : august 15, 1976 novemter 15, 1978 if there are any substantial delays m publication, the timing of the newsletter will be off, and, pereaps, flie usefulness of the publication will be less. ;wsletter rchhat during the p iication and di 2, number 1 o have received a concerning the the journal, format seem to though few hav with enthusiasm ber of lines pe 3y printing eig the lines are c a bit difficult on tne other ha ran about thi lines to the in forty-two pages to tne inch, of publication have one of two terial with th spread out a bi with the lines of cost conside cation, in its not average mor pages per issue comment from r eriou since the pubstribution of volume f the newsletter we number of comments forraat and style of comments on general be positive, ale been overwhelmed concerning the numr vertical inch (8) . ht lines to the inch ompressed and may be , at times, to read, nd. number 1 , which rty pages at eight ch, would have been long at six lines because of the cost of the newsletter we choices: less mae vertical spacing t, or more material compressed. because rations, the publipresfcnt formshould e than about thirty i will appreciate eadets on this matt. w. n. 3j instructions for authors of the iassist quarterly 1/39 alter, george; rizzolo,flavio; and schleidt, kathi (2023) view points on data points, iassist quarterly 47(1), pp. 1-39. doi: https://doi.org/10.29173/iq1051 view points on data points: a shared vocabulary for cross-domain conversations on data and metadata george alter1, flavio rizzolo2, kathi schleidt3 abstract sharing data across scientific domains is often impeded by differences in the language used to describe data and metadata. we argue that disagreements over the boundary between data and metadata are a common source of confusion. information appearing as data in one domain may be considered metadata in another domain, a process that we call “semantic transposition.” to promote greater understanding, we develop new terminology for describing how data and metadata are structured, and we show how it can be applied to a variety of widely used data formats. our approach builds upon previous work, such as the observations and measurements (iso 19156) data model. we rely on tools from the data documentation initiative’s cross domain integration (ddi-cdi) to illustrate how the same data can be represented in different ways, and how information considered data in one format can become metadata in another format. keywords metadata, data sharing, data interoperability acknowledgments this paper grew out of the codata working group on semantic integration, and we are grateful to simon cox (chair) and other members of the group for their encouragement. rob atkinson, alejandra gonzalez-beltran, larry hoyle, hylke van der schaaf, chris schubert, and john wieczorek provided helpful comments and guidance on earlier drafts of this paper. this paper would not have been possible without the path-breaking work of the ddi alliance cross domain integration working group. in memoriam to herbert schentz, we carry on the good fight towards semantic data interoperability. https://doi.org/ 2/39 alter, george; rizzolo,flavio; and schleidt, kathi (2023) view points on data points, iassist quarterly 47(1), pp. 1-39. doi: https://doi.org/10.29173/iq1051 problem statement although the value of sharing data across scientific domains is rapidly increasing, conversations about data are often very difficult. there are many types of scientific data, and each discipline has developed its own standards, procedures, and language about data. combining data from multiple sources becomes a very frustrating process when common terms, like ‘observation’ and ‘attribute’, are used differently across scientific domains. differences in use of the term “metadata” are especially problematic. we will show that there is no fixed boundary between “data” and “metadata,” and that information viewed as data in one discipline may be metadata in another. to achieve the fair goal of interoperability (wilkinson, et al., 2016), we must overcome not only different ways of structuring data but also different ways of conceptualizing and describing data. this paper defines terms describing fundamental aspects of metadata that can be applied consistently across all disciplines. our work builds on and extends the ddi alliance’s cross domain integration (ddi-cdi) model, which describes how data can be arranged in different ways (ddi alliance, 2020a). ddi-cdi is an important departure from ddi’s original focus on describing data arrayed as ‘variables’ (columns) and ‘cases’ (rows) (vardigan, heus, & thomas, 2008). ddi-cdi provides a bridge to data structures used in scientific domains that organize data around other concepts, such as ‘observations’ and ‘features.’ however, ddi-cdi has little to say about the metadata accompanying each data structure. consequently, information appearing as data in one data structure may disappear in the transition to a different data structure. we extend ddi-cdi by applying its constructs to metadata as well as data. metadata is data too, and we show how ddi-cdi can be applied to data structures containing metadata. in our extended version of ddi-cdi, information is not lost when data are moved from one data structure to another, because we map transitions to and from data and metadata. we illustrate this approach by showing how data organized in “long” format can be translated into “wide” and “multidimensional” formats. in particular, we sketch the path from an individual data-point expressed in accordance with the ogc observations & measurements data model (iso 19156; cox, 2011) into the tabular representations described by the ddi codebook format and the multidimensional (n-cube) format described by the sdmx (statistical data and metadata exchange) standard (statistical data and metadata exchange (sdmx), 2013, 2021). we believe that confusion about data restructuring is often due to what we call “semantic transposition.” we define semantic transposition as the relocation of the representation of a characteristic from the data structure to the data content, that retains isomorphic coherence between representations. this occurs because information about the meaning of a measured value may be either internal or external to a data set. for example, if the measured value is 32, we need to know whether the characteristic being measured is age, temperature, or something else. some scientific domains are accustomed to including this information in the same data array as the measured values, but others provide a separate “metadata” file that attaches meanings to areas in the data array. semantic transposition occurs when the data are restructured in a way that moves information into or out of https://doi.org/ 3/39 alter, george; rizzolo,flavio; and schleidt, kathi (2023) view points on data points, iassist quarterly 47(1), pp. 1-39. doi: https://doi.org/10.29173/iq1051 the data set, shifting information about the characteristic measured by the value from the data array to the descriptive frame. in other words, the boundary between ‘data’ and ‘metadata’ is flexible. semantic transposition has been discussed in the computer science literature, where it is described as data-metadata translation (hernández, papotti, & tan, 2008; papotti & torlone, 2009). a common example is the transformation of stock ticker data, which is an illustration of the "pivot-unpivot" problem. a stock ticker reports data with three columns: time, company identifier, value. for the purposes of analysis, these data are often transformed (pivoted) to a matrix with one row per observation time and values arranged in separate columns for each company. this transformation involves transposing the company identifier from data to metadata (column name). computer scientists see this as a problem of mapping across database schemas, but in this case the target schema depends upon the content of the data, which cannot be specified in advance (hernández et al., 2008). several ways of formalizing and automating data-metadata translation have been proposed (beine, hames, weber, & cleve, 2014; britell, delcambre, & atzeni, 2016; wyss & robertson, 2005; xue, shen, nie, kou, & yu, 2013). the pivot-unpivot transformation is common in statistical analysis, where it is called "long to wide," and we link the choice between long versus wide data formats to data cultures found in different scientific domains. the analysis presented here is particularly important for data stewards serving the social sciences. the standard format of data in the social sciences is the ddi-cdi “wide” data structure, which is a rectangular matrix of columns (“variables”) and rows (“observations”). unlike other data structures, “wide” format does not assign roles to different columns. in particular, there is no way to indicate that one column describes an attribute of another column, such as an indicator of data quality or an estimation method. social science data repositories have responded to this problem by developing a robust and detailed metadata standard, ddi, which provides a number of ways to annotate important aspects of a “wide” data array. consequently, metadata plays a much broader role in the social sciences than in other scientific domains, and extending ddi-cdi to trace semantic transposition is important for the interoperability of social science data with data from other domains. overview we proceed in steps intended to bridge practices and understandings in multiple scientific domains. first, we begin by defining basic concepts. this step is essential, because different domains often use the same words to mean different things. although we borrow freely from various sources, we offer our own definitions to provide a consistent and comprehensive terminology. second, we describe a theory of observation based on the iso 19156 observations and measurements (o&m) standard. since o&m was primarily developed from experience in the ecological and earth sciences, we use an example from the social sciences to emphasize its universality. third, we examine four common data structures described by ddi-cdi. we begin with a “simple observation” consistent with o&m and show how it can be represented in ddi-cdi constructs. then, we trace the movement of aspects of an observation as it is transformed from long to wide to multidimensional data structures. we represent each data structure in a tabular format to show how information flows from one data structure to another, even though the data may not be tabular in https://doi.org/ 4/39 alter, george; rizzolo,flavio; and schleidt, kathi (2023) view points on data points, iassist quarterly 47(1), pp. 1-39. doi: https://doi.org/10.29173/iq1051 practice. this discussion illustrates semantic transposition as defined above, and we use the terms and concepts of ddi-cdi to introduce variable definition data structures and dimension definition data structures, which complement and explain each data format. we also provide an appendix with machine-actionable descriptions of the steps involved in transforming data from long to wide to multidimensional using structured data transformation language (sdtl). sdtl is an independent language for describing data transformation commands in standard metadata formats, like data documentation initiative (ddi) and ecological metadata language (eml) (san gil, vanderbilt, & harrington, 2011). our conclusion reflects on the difficulty of combining data from different scientific domains and the importance of semantic transposition as a source of misunderstanding. basic concepts before we dive into the details of the cdi methodology and apply it to data concepts, we introduce a set of basic terms. we are working across a wide range of scientific domains ranging from environmental to social sciences. each domain has a unique entrenched terminology that is often at odds with usage in other domains. for example, the “characteristic” under observation may be referred to as the “property” or “variable” in other communities. this section defines the terms used in this document together with the meanings we apply to them in the hope that this clarification helps to avoid subsequent misunderstandings. instance value we use instance value to refer to the smallest atomic unit of data. an instance value may be a number, text, boolean, or any other data type. an instance value may result from a measurement process, or it may describe the measurement process itself. “child” and “-10” can be instance values measuring age and temperature, but the text strings “age” and “temperature” may also be instance values. characteristic value we use characteristic value to refer to an instance value that contains a measured value describing a single characteristic of a specific entity. the entity may be a person, place, thing, total for a region or a year, etc. a characteristic value does not explain its own meaning. the value “-10” may refer to a temperature, which could be measured on the celsius or the fahrenheit scale, or it may be the difference between a test score and the mean of all scores. “hazel” may be a person’s name or an eye color. characteristic values are only meaningful if they are accompanied by additional descriptive information, such as the characteristic being described and methodological details of the data acquisition process. characteristic characteristics explain the meaning of a characteristic value. “name” and “eye color” are both characteristics that could result in an instance value of “hazel.” our understanding of a characteristic value must be informed by knowledge of the characteristic that it describes. some commonly used terms for this concept are attribute, parameter, variable, observed property (or just property), measurand, analyte. individual domains become even more specific, with geology field observations utilizing terms such as strike and dip, lithology, alteration state, etc. https://doi.org/ 5/39 alter, george; rizzolo,flavio; and schleidt, kathi (2023) view points on data points, iassist quarterly 47(1), pp. 1-39. doi: https://doi.org/10.29173/iq1051 entity an entity is a physical, digital, conceptual, or other kind of thing with fixed aspects, that has separate and distinct existence, real or abstract (provenance working group, 2013; iso/tc 211 terminology maintenance group, 2020). this broad definition is consistent with emerging usage in communities promoting data documentation and exchange. we prefer these definitions to domain-specific terms like “unit of observation,” “statistical unit,” and “feature of interest.” an entity may also encapsulate a temporal aspect, such as a moment or period of time. thus, a person who held three jobs during a calendar year may be modeled as three entities (pertaining to their employment), each of which existed for a part of the year. similarly, a household is an entity composed of persons (individual entities) who live in a defined space or share a common budget. thus, entities may be defined as composites of physical, temporal, and conceptual units. when data are transformed or re-organized, we often change the entity that they describe. this is most apparent when data are aggregated, as in one of the examples below. if we count the number of women enumerated in a census, the characteristic “number of women” refers to an entity defined by the geographic coverage and date of the census. time may also be used to create more disaggregated entities. suppose that the heights of a group of school children were measured several times. these data can be arranged in a wide format showing multiple heights for each child, or they can be in a long format where the entity is an observation for one child on a specific date. (see below for formal definitions of wide and long formats.) value domain a value domain is the set of allowed values for a characteristic, qualifier, or identifier. the value domain for temperature in kelvin, for example, would be all real numbers greater than or equal to 0. name has a value domain that includes “hazel,” “fred,” and “wilma.” the value domain of eye color includes “hazel,” “brown,” “blue,” and “green.” qualifier we use the term qualifier to refer to additional information about a characteristic value. data are often annotated with information about the measurement procedure, the instrument that was used, date and time of measurement, geographic location, verification or validation procedures, etc. qualifiers are attributes of attributes. information of this kind is often essential for comparing measurements from different studies or for deciding how much confidence to place in a specific characteristic value. identifier identifiers associate an instance value with an entity or a type of characteristic or a type of qualifier. an id number for a subject or feature is the most common kind of identifier, but identifiers may also pertain to the type of a characteristic or qualifier. we may characterize identifiers by their functions as ● entity identifiers https://doi.org/ 6/39 alter, george; rizzolo,flavio; and schleidt, kathi (2023) view points on data points, iassist quarterly 47(1), pp. 1-39. doi: https://doi.org/10.29173/iq1051 ● characteristic identifiers ● qualifier identifiers key a key is an identifier or set of identifiers that uniquely reference a characteristic value. thus, a person has only one place of birth, and any combination of instance values that uniquely associates an individual with a place of birth can be used as a key. a key also references any qualifiers and identifiers associated with its characteristic value. keys are dataset specific, and the number of identifiers required to compose a key can change if the structure of the data is modified. for example, if a study collects heights of school children, an identifier for each child can be used as a key. if the study re-measures the same children at a later date, a key for the combined data requires both the child’s id and the date of measurement. thus, a child’s id is not a unique key when children are measured more than once. note that the key points to a unique characteristic value (height) and not to a person (child). this example also shows that a qualifier can be used as an identifier within a key. date of measurement is a qualifier of height, but date of measurement also serves as an identifier when it is part of a key. in other words, when height is measured more than once, date of measurement plays two roles. it is both a qualifier that affects the interpretation of height and an identifier that distinguishes among multiple measurements of the same child. a key can be associated with more than one characteristic value (age, place of birth, date of birth, mother’s name), but only one characteristic value for each characteristic can be associated with a specific key. procedure a procedure is the underlying methodology utilized to ascertain the value of a characteristic value. knowledge of this methodology is often necessary to understand the applicability of the available data to a specific use case. datapoint a datapoint is a container for an instance value, which may be a characteristic value, a characteristic, a qualifier, or an identifier. one can think about a datapoint as a cell in a matrix, such as a spreadsheet. some cells in the spreadsheet are the characteristic values that we plan to study, but other cells contain explanatory information about what characteristic was measured when, where, and for whom. full simple observation we use the term full simple observation to refer to a characteristic value with all of its associated identifiers and qualifiers. this is the most atomic level of usable data, because it brings together a measured value with information about what was measured and how measurement was performed. the components of a full simple observation may appear in the same place, as in a row of a spreadsheet, or in linked locations, such as tables in a relational database. https://doi.org/ 7/39 alter, george; rizzolo,flavio; and schleidt, kathi (2023) view points on data points, iassist quarterly 47(1), pp. 1-39. doi: https://doi.org/10.29173/iq1051 data set a data set is a collection of datapoints that have been organized in a known way. the structure of the data set tells us which datapoints are characteristic values, characteristics, qualifiers, identifiers, etc. thus, the format of a data set sets our expectations about each of its datapoints. data structure a data structure describes the roles played by various datapoints in a data set. the data structure indicates which datapoints are characteristic values, characteristics, qualifiers, and identifiers. a “logical” data structure describes the roles played by the datapoints in a data set. a “physical” data structure shows how datapoints are formatted into rows, columns, and files. a logical data structure may be instantiated in more than one type of physical data structure. some physical data structures do not include all of the information required to make a data set usable. for example, data formatted as comma separated values (csv) without semantically meaningful column headers is useless without accompanying documentation of the characteristic and role played by each column of datapoints. if this documentation is machine actionable, it can be described as a data structure in its own right. metadata the preceding discussion avoided using the word “metadata.” one could say that instance values are data while characteristics, qualifiers, and identifiers are metadata, but we do not consider statements of that kind helpful. as we mentioned above and will illustrate below, characteristics can be datapoints within a data structure or they can be supplied elsewhere. if “temperature” and “name” are included in a data structure, they can be processed like any other datapoint. for example, one can count the number of characteristics in a dataset. in other words, the difference between “data” and “metadata” depends upon how a datapoint is used and not on how it is provided. we argue here that the assignment of content to “data” versus “metadata” is arbitrary. most scientific domains are accustomed to data structures that specify which concepts are embedded in the data content and which concepts are part of the descriptive frame. these decisions are often motivated by technical considerations about types of data and modes of analysis, as well as user perspectives on the usage of the data, but the same data can be represented in alternative data structures. reshaping data into a different data structure is an act of semantic transposition that determines which content will be provided in the data content and which will go into the metadata (descriptive frame). for example, the characteristic measured in a datapoint may be provided in the data or associated with a column name that points to a description of the characteristic in the metadata. cross domain integration from ddi alliance (ddi-cdi) the data continuum different use cases entail the use of different structures for data representation. in some cases, precise metainformation detailing the data acquisition process is essential to understanding the applicability of the provided data to the task at hand. in other cases, simplified structures can be of great benefit, reducing the resources required for data provision, transport and use. a modern data provision landscape should encompass both aspects, while ensuring a degree of continuity between these alternative viewpoints on the same data source. https://doi.org/ 8/39 alter, george; rizzolo,flavio; and schleidt, kathi (2023) view points on data points, iassist quarterly 47(1), pp. 1-39. doi: https://doi.org/10.29173/iq1051 while each of the various data structures utilized is consistent in itself, issues arise when data is transformed from one format to the other, as often required to support a wide array of use cases. it becomes difficult to maintain semantic coherence across structures for those cases when it proves necessary to drill down into the details of the data. in order to expose data in different structures, with differing depths of information contained, it would be advantageous to be able to provide links between parallel concepts, allowing a user to traverse between these different structures. ddi-cdi offers a way to encapsulate the underlying essence of the data being provided. in the following sections, we will describe the relevant concepts from this emerging data alignment model. we focus on two ways that ddi-cdi helps us to characterize data structures. first, a data structure can be defined by its keys. as we discussed above, a key is a set of one or more identifiers that point to a characteristic value. data structures differ in the number and types of identifiers, i.e., keys, required to uniquely identify a characteristic value. second, ddi-cdi describes the roles that instance values play in different data structures. the role played by an instance value may differ across data structures. for example, an instance value that is part of a key in one data structure may not be part of a key in another data structure. ddi-cdi refers to roles as “components”, which will be described below. we differ from ddi-cdi in several ways. ● we use the same concepts to describe the data content as the descriptive frame, i.e., metadata. we emphasize that metadata is also data. transferring data from one data structure to another often involves moving information (instance values) from the data section (data content) to a metadata section (descriptive frame), i.e., semantic transposition. ● ddi-cdi does not specify relationships between qualifiers and characteristic values. we consider this relationship essential for understanding differences among data structures. ● although we use ddi-cdi concepts to describe logical data structures, we also provide examples showing simplified physical data structures. we hope that these examples will help readers to see beyond an abstract discussion of concepts to practical applications. ● we introduce a data structure called “nested name-value pairs,” which is not included in the ddi-cdi. “nested name-value pairs” is similar to the “key-value” data structure included in ddi-cdi, but name-value pairs may be nested, which is not possible with “key-value” pairs. we consider “nested name-value pairs” a reasonable extension of ddi-cdi, and we hope that it will be added to the ddi-cdi specification. theory of observation what is an observation? most information we have about our surroundings can be seen to be the outcome of observations or measurements upon our world. the essential characteristics of an observation have been elaborated within the standard iso 19156 observations and measurements (o&m) (iso 19156), leading to a richly structured model as follows. the essence of an observation is a relation assigning a value (the range of the observation relation) to an entity (the domain of the observation relation). in addition, various additional pieces of observational metainformation are linked to this relation via the observation object, including: https://doi.org/ 9/39 alter, george; rizzolo,flavio; and schleidt, kathi (2023) view points on data points, iassist quarterly 47(1), pp. 1-39. doi: https://doi.org/10.29173/iq1051 ● the property or characteristic of the entity for which a value is being provided, e.g. color, temperature. in a simplified structure, this property would be the name of the relation between the entity and the value; ● the procedure used in obtaining the value for the entity. this can be essential for interpreting the value provided, as different methodologies can deliver vastly different results in dependence on external factors; ● temporal information pertaining to the observation, specifically the phenomenon time and the result time; ● spatial information on where the entity being observed was located at the time of observation. in a similar vein, information such as the measurement device utilized, the person performing the measurement or the facility in which this act took place is often provided, as well as references to other observations providing essential contextual information are foreseen within this model, but omitted here for brevity. in the figure 1 below, we show the conceptual model underlying the update of o&m, to be released as iso 19156:2022. the core of the observation consists of 2 associations to the left: domain and range; domain associates the observation with the feature-of-interest, the entity upon which the observation provides a value for a characteristic, while the range associates the observation with the actual value for this characteristic pertaining to the feature-of-interest. the observableproperty provides the characteristic under investigation, while the observingprocedure describes the measurement methodology. in addition, information on the observer, e.g., a sensor or human providing the value of the characteristic, the host, e.g., the facility the sensor mounted at, as well as deployment information linking an observer to a host can be provided. figure 1. class diagram for observation and measurement model this leads to a precise but complex representation of all aspects of the observational process deemed relevant to later interpretation and use of the data. while access to these details may be essential to understanding the applicability of the data, when it comes to further processing steps, this excess baggage proves cumbersome; simpler formats are required. https://doi.org/ 10/39 alter, george; rizzolo,flavio; and schleidt, kathi (2023) view points on data points, iassist quarterly 47(1), pp. 1-39. doi: https://doi.org/10.29173/iq1051 an example relational observation in the following example, we will concern ourselves with the gender of an individual. in the simplest representation, we could expect an entity of type person to have an attribute or operation gender, providing this characteristic as a string value, ideally referencing a standardized vocabulary. in figure 2 below, we have modeled this example as a uml interface person with the operation gender() of type gendervalue (a data type providing a string value representing the gender of the individual); an instance simple1001 has been derived from this interface, and gender provided as a reference to a uri representing the value “female”. figure 2. simple instance diagram suppose that we want to add information about how gender was determined. the concept of gender can be reified from an attribute to a class, which can have more than one attribute. (see olivé (2007, chapter 6) for a definition of reification.) in figure 3, we show the genderdetermination class with two attributes observingprocedure and gender value. reification of the attribute “gender” to the class genderdetermination allows us to show that a specific measurement procedure, “external_observation,” applies only to the determination of gender and the resulting value “female.” https://doi.org/ 11/39 alter, george; rizzolo,flavio; and schleidt, kathi (2023) view points on data points, iassist quarterly 47(1), pp. 1-39. doi: https://doi.org/10.29173/iq1051 figure 3. reified instance diagram rather than creating a dedicated class for every attribute, o&m further abstracts the reified class to the concept of observation in which the characteristic being represented is provided via the observableproperty association and class. in figure 4 below, we show the full o&m representation of our gender example, where explicit interfaces are provided for the relevant observational metainformation concepts. note the semantic transposition of the denotation of the measurement characteristic “gender” from the name of an attribute on person to the content of the name of the observableproperty; we will repeatedly observe this semantic transposition or flip-flop between data content and data structure for the provision of the characteristic under investigation as we further analyze the various structures commonly used for the representation of observational data. https://doi.org/ 12/39 alter, george; rizzolo,flavio; and schleidt, kathi (2023) view points on data points, iassist quarterly 47(1), pp. 1-39. doi: https://doi.org/10.29173/iq1051 figure 4. instance diagram with identifiers and qualifiers in order to bridge the gap between a relational data store and existing external vocabularies, most interfaces foresee a link attribute by which a reference to the corresponding concept can be stored. alternatively, the link can also be utilized to reference any external source providing additional information on this entity. we are aware that gender determination is a sensitive topic, but the changing understanding of gender illustrates our point. when sex was considered an immutable biological characteristic, all measurement procedures were expected to yield the same result. as we recognize the right of individuals to determine gender for themselves, the procedure used to ascertain gender becomes more important. we do not expect everyone to self-identify with the biological sex ascribed to them at birth or to be limited to the binary choice between female and male. measurement procedures have important consequences. the true power of such a richly structured representation becomes clear when we add additional observations on the gender of this individual over time. while the observation shown above describes the gender determination made at birth, additional determinations could be made over the individual’s lifetime, following various methodologies. continuing this example, we now add a gender self-determination observation in figure 5. as this observation is performed by the subject, there is no need to provide information on the observer and host; the following diagram illustrates this observation. https://doi.org/ 13/39 alter, george; rizzolo,flavio; and schleidt, kathi (2023) view points on data points, iassist quarterly 47(1), pp. 1-39. doi: https://doi.org/10.29173/iq1051 figure 5. instance diagram with identifiers and qualifiers for first alternative procedure a final gender determination is performed after the death of the individual in an attempt to clarify the contradictory gender markers available (figure 6). for simplicity we have defined the observer as the coroner responsible for this step, while in a real-world system, this object should probably represent the dna analysis equipment. this example illustrates the issues encountered when such information is oversimplified. if we only consider the information pertaining to gender being exposed via the simple interface, there are two options for provision of the data, neither proving satisfactory: ● the gender attribute changes over time, thus returning different values for the same individual for three time-instances as follows: ○ 19320303t14:15:00: female ○ 19690713t19:45:00: male ○ 20051203t08:30:00: female ● an undetermined value for gender of the individual the type of representation required depends on the actual use case. when looking for suitable data, the complex richly structured representation provided by the o&m standard is often essential to allow a domain expert to determine if the data is fit for purpose, but once the data has been vetted and deemed appropriate, simpler representations allow for more efficient data transfer and portrayal. https://doi.org/ 14/39 alter, george; rizzolo,flavio; and schleidt, kathi (2023) view points on data points, iassist quarterly 47(1), pp. 1-39. doi: https://doi.org/10.29173/iq1051 figure 6. instance diagram with identifiers and qualifiers for second alternative procedure common data structures conventions and data used in this document data takes many forms in its transformation from primary microdata to highly aggregated data stores. during this process, depending on the type of structure utilized, individual concepts can shift from being part of the content section to the descriptive frame. in this section, we describe the different structures involved at the different levels of this process before going into the details of tracing the individual concepts throughout this process. in order to illustrate this process, we use a simple example dataset, providing data pertaining to the following characteristics on two individuals: ● name ● gender ● born ● died ● refarea https://doi.org/ 15/39 alter, george; rizzolo,flavio; and schleidt, kathi (2023) view points on data points, iassist quarterly 47(1), pp. 1-39. doi: https://doi.org/10.29173/iq1051 we use ddi-cdi concepts to represent the logical data structure of a relational observation in figure 7. for ease of representation within this paper, we have utilized the following graphical representation of the ddi-cdi concepts: ● data points, represented as blue boxes, with the instance value provided in the darker blue oval contained ● variables, represented as pink boxes, with the value domain provided in the darker red oval contained ● roles of elements within a structure, also referred to as data structure components, are represented as green ovals. ○ dark green ovals are used for elements that play the same role in every structure. ○ light green ovals are used for elements playing a role that is specific to a particular format. ● keys, represented by golden key symbols, indicating which concepts must be combined to form the key for a specific observation additional concepts we have defined for this paper ● golden rounded boxes are used to indicate characteristics and characteristic values. these are terms defined in this paper showing continuities across data structures that can be lost in the terminology differences between scientific domains. ● pink and red rounded boxes introduced in figure 11 are used to represent features of data arrays, such as column headers and variable names, that are part of the descriptive frame. named arrow representations: ● “identifies” (in blue): indicates the data points that are uniquely identified by the given key. ● “identifies” (in green): indicates the variable a given instance value references in a variable descriptor component. ● “has" (in green): indicates the data structure component that is part of a given data structure. ● “is defined by”: indicates the variable that gives meaning to a given data structure component. ● “has value from”: indicates the value domain from which an instance value is taken. ● “name”: indicates the name of a data structure component. ● “refers to”: indicates the reference value a descriptor points to. unnamed arrow representations: ● yellow lines show aggregation, indicating the instance value that is part of a given key. https://doi.org/ 16/39 alter, george; rizzolo,flavio; and schleidt, kathi (2023) view points on data points, iassist quarterly 47(1), pp. 1-39. doi: https://doi.org/10.29173/iq1051 ● dashed associations (in red): terminology mapping that relates the notions of characteristics and characteristic values to variables and instance values, respectively. ● aggregation (in red): indicates a data structure component name is part of a header. relational representation of observations we start with a richly structured relational representation of observations, following the o&m model previously described. in figure 7, we show the determination of the characteristic gender utilizing the procedure dna analysis to determine that the individual with identifier 1001 has the value female. concepts such as observer and host have been omitted for brevity. this example includes only one characteristic value, “female”, which is associated with the characteristic “gender”. in figure 7, the instance value “female” of variable “gender” is a measure component. “gender” appears as an instance value that belongs to a value domain together with other characteristics, such as age, height, hair color, place of birth, etc. ddi-cdi calls this role a variable descriptor component. in other words, a variable descriptor identifies the characteristic measured by a variable. in this example, the variable descriptor assigns a name to the variable, but it could also provide a uri referencing a variable description in an ontology. the variables on the right side of figure 7 are qualifiers, which are called attribute components in ddi-cdi. procedure and time tell us important things about the characteristic value (“female”): how it was determined and when it was measured. figure 7. logical data structure of a simple observation (one procedure and one time per measure) the key for this observation includes two instance values “1001”, which is an identifier for the person, and “gender” which is the characteristic. the variable descriptor is part of the key, because we may have additional observations on other characteristics. the instance value of time is not part of the key in figure 7, because we are representing an observation that only occurred once. if there are multiple observations on the same person at different times using different methodologies, as shown https://doi.org/ 17/39 alter, george; rizzolo,flavio; and schleidt, kathi (2023) view points on data points, iassist quarterly 47(1), pp. 1-39. doi: https://doi.org/10.29173/iq1051 in our initial observation description above, time and procedure are included in the key. in figure 8, the yellow aggregation lines between the key symbol and the instance values for time and procedure indicate the composition of the key. figure 8. logical data structure of a simple observation (key supporting multiple procedures and times per characteristic) tracing data through alternative physical data structures to illustrate these concepts, we show how the same data are represented in three different data structures: long, wide, and multidimensional. ddi-cdi characterizes the data content in each of these formats, but we will show that these models are incomplete without also explicitly exposing the descriptive frame (metadata). in each step from long to wide to multidimensional, information moves from the data content to the descriptive frame. in other words, transposing data to a different data structure results in semantic transposition as well. recognizing that metadata are also data, we can use the tools provided by ddi-cdi to create variable description data structures for the descriptive frames associated with each data structure. the direction of this transposition can go both ways depending on the restructuring; in some cases, information moves back from the descriptive frame to the data content. thus, our models retain descriptive information that appears to disappear (or reappear) if we focus only on the data content. the examples below are presented in tabular form, because it is simple to display in print and easy to understand. tabular formats (e.g., csv, spreadsheets) are widely used due to their flexibility and simple ingestion by a wide range of analysis tools. however, the same logical data structures could be implemented in normalized relational databases, json, rdf, or other physical formats. https://doi.org/ 18/39 alter, george; rizzolo,flavio; and schleidt, kathi (2023) view points on data points, iassist quarterly 47(1), pp. 1-39. doi: https://doi.org/10.29173/iq1051 long format long format, also referred to as narrow or stacked data, or in its most primal form as entity–attribute– value model (eav) data, is most closely related to relational observations as described in the section above. for each characteristic value, there is one row in the table. in the purest eav format, the data consists simply of triples with the following structure: ● entity: the person, object, or thing that is the target of the value provided ● attribute: the characteristic (also called variable or property) that is being described by the value ● value: the characteristic value assigned to the entity table a provides an example of data encoded in eav format, in which the entity is personid and the attribute is called property. personid property value 1001 name abigail 1001 gender female 1001 born 03.03.1932 1001 died 01.12.2009 1001 refarea newport 1011 name benjamin table a: example long format: entity–attribute–value model https://doi.org/ 19/39 alter, george; rizzolo,flavio; and schleidt, kathi (2023) view points on data points, iassist quarterly 47(1), pp. 1-39. doi: https://doi.org/10.29173/iq1051 figure 9. logical data structure of long format: entity-attribute-value figure 9 shows the logical data structure of a slice of the data in table a. figure 9 includes one measure component (gender) and one identifier component (person id). the measure component must be linked to two keys: person id and property, which is a variable descriptor component showing the characteristic measured in the characteristic value. the value “female” is linked to two value domains. on one hand, “female” is drawn from the value domain of gender. on the other hand, “female” is drawn from the value domain of the eav column “value”, which is the union of all value domains of attributes in the data including gender. long format can also be extended to provide additional information qualifying each characteristic value. a simple example is the provision of source information for each datapoint, which was included as a qualifier in figure 6, as shown in table b. https://doi.org/ 20/39 alter, george; rizzolo,flavio; and schleidt, kathi (2023) view points on data points, iassist quarterly 47(1), pp. 1-39. doi: https://doi.org/10.29173/iq1051 personid property value source 1001 name abigail birth register 1001 gender female dna analysis 1001 born 03.03.1932 birth register 1001 died 01.12.2009 kin report 1001 refarea newport drivers license 1011 name benjamin birth register table b: example long format – source https://doi.org/ 21/39 alter, george; rizzolo,flavio; and schleidt, kathi (2023) view points on data points, iassist quarterly 47(1), pp. 1-39. doi: https://doi.org/10.29173/iq1051 figure 10. long format extended to include a qualifier: verification figure 10 extends figure 9 in two ways. we show two characteristics (name and born) for two observations (1011 and 1061), and we have added an attribute component (source). the attribute component in figure 10 is linked to a measure component by sharing the same two keys: “1011, name”. this means that its instance value (birth register) applies to the name variable of person 1011. we include only one attribute component to avoid further complications in a busy diagram, but we could add attribute components with the procedures used to ascertain each of the other three instance values: “1011, born”; “1061, name”; “1061, born”. long format can be further extended to include the full breadth of information in an observation in a lossless manner. table c adds columns for two more attribute components depicted in figure 6: time and source. note that these additional columns are both qualifiers that modify the characteristic value (i.e., “female”), and they are linked to the characteristic value by a two-part key, “1001, gender”. personid property value time source 1001 gender female 5.12.2009 dna analysis4 table c: example long format full simple observation wide format wide (or unstacked) data is structured with a separate column for each characteristic, such that each row contains all of the datapoints pertaining to one observed entity. table d exactly corresponds to table a, but the datapoints are arranged horizontally rather than vertically. consequently, characteristic values in table d are linked to only one key, personid. personid gender name refarea born died 1001 female abigail newport 03.03.1932 01.12.2009 1011 male benjamin cardiff 01.08.1929 02.06.2006 table d: example wide format simple representation table d poses a problem that was implicit in our discussion of long format, but critical for understanding wide format. what is the special status of the first line in table d? we can make this question more conspicuous by re-writing the same data as table e. column headings like “var01” have no intrinsic meaning, and we could just as easily write the same data matrix without any headings at all. clearly, if table e is not accompanied by additional information, it is unusable. one might infer the meaning of var01 and var02, but it is impossible to interpret the other columns. https://doi.org/ 22/39 alter, george; rizzolo,flavio; and schleidt, kathi (2023) view points on data points, iassist quarterly 47(1), pp. 1-39. doi: https://doi.org/10.29173/iq1051 var00 var01 var02 var03 var04 var05 1001 female abigail newport 03.03.1932 01.12.2009 1011 male benjamin cardiff 01.08.1929 02.06.2006 table e: example wide format arbitrary column headings table e is only meaningful if it is accompanied by table f. table f is sometimes described as a codebook or a variable inventory, and it is often distributed as a text file or a spreadsheet. for our purposes, table f is a simplified version of a metadata file. scientific domains that distribute data in formats like table e have developed more elaborate standards for providing metadata in machine actionable formats, like xml, json-ld, and rdf. data repositories serving the social sciences rely on metadata in one of the data documentation initiative (ddi) standards, and repositories serving the ecological sciences use ecological metadata language (eml) among other standards. our point is that table e requires table f to supply information contained in the “property” column of table a. variable name variable label variable description var00 personid person identification number var01 gender gender var02 name first name var03 refarea location of principal residence var04 born date of birth var05 died date of death table f. variable descriptions in figure 11 we provide logical data structures for both tables e and f. the left side of figure 11 shows the data content. as we noted above, wide format includes only one key (personid); when we provide other characteristic values (name, refarea, born, died), they are all linked to the same key. the right side of figure 11 is a variable description data structure, that attaches meanings to the arbitrary variable names in table f. notice that the variable description data structure on the right side of figure 11 has essentially the same structure as the wide data structure on the left side of the diagram. both structures use a key to reference a characteristic of an entity. in the wide data structure on the left, the entities are persons who have measured characteristics, such as gender and age. entities in the variable description data structure on the right are variables, which have labels, descriptions, and other attributes. https://doi.org/ 23/39 alter, george; rizzolo,flavio; and schleidt, kathi (2023) view points on data points, iassist quarterly 47(1), pp. 1-39. doi: https://doi.org/10.29173/iq1051 in figure 11, variable descriptions are linked to data in the wide data structure through their variable names (var00, var01), which appear in the header of table e. headers may or may not be stored with the data array. for example, column names may be provided in the first row of a csv file or in a separate document, such as a codebook or ddi xml file. we show the header elements of a data structure in pink and red to indicate its ambivalent position. although we did not show a header in our discussion of long format, we return to this issue below. figure 11. wide data structure with variable description data structure when we consider the wide data structure on the left side of figure 11 by itself, “var01” is a measure component. however, when figure 11 is considered as a whole, “var01” is a bridge to information on the right side of figure 11. in the variable description data structure “var01” is a propertyid, which is used to identify the characteristics of a variable. “var01” is a key referencing the property “gender”, which is the meaning of “var01” in the wide data structure. in ddi-cdi terminology, property id is a variable descriptor component and property is an attribute component. these components are present in long format shown in figure 7, but they are not part of a wide data structure. this means that figure 11 taken as a whole has all of the components found in figure 7 for long format above. thus, figure 11 provides another way of answering the question posed above: what is the status of the first row in table d? the first row in table d is a set of variable descriptions. unlike table a, these descriptions are not included in the data array itself. rather, they belong to a separate data/metadata array that must be provided to make the data in table d meaningful. the importance of breaking down the distinction between data and metadata is clear if we consider going from table e to table a. even though the measured values are the same in both tables, the column headings in table e are arbitrary and do not identify the contents of each column (i.e., the characteristics), which is done in the descriptive frame (table f). in contrast, characteristics are given in the data content in table a. the only way to fill the “property” column in table a is to refer to a separate “metadata” table such as table f. thus, transposing data from tables e and f to table a https://doi.org/ 24/39 alter, george; rizzolo,flavio; and schleidt, kathi (2023) view points on data points, iassist quarterly 47(1), pp. 1-39. doi: https://doi.org/10.29173/iq1051 requires moving instance values from the descriptive frame to the data. to automate data integration across disciplines, the “i” in fair, metadata must be provided in standard machine-actionable formats. adding qualifiers to wide format unlike long format, wide format does not provide an easy way to associate qualifiers with characteristic values. since long format uses two keys, the entity identifier and the characteristic, qualifiers are unambiguously linked to the characteristic values that they describe. in wide format, qualifiers must be linked to characteristic values through their variable names. this can be accomplished in a variable description data structure (metadata) or by including a descriptor for the characteristic in the variable name, as we show in table g. in such cases, care must be taken to assure that the metaformat defined for these qualifiers (e.g., the characteristic name plus “_source” in table g) is explained to data users. personid name name _source gender gender_ source refarea refarea_ source born born _source died died _source 1001 abigail birth register female dna analysis newport drivers license 03.03.19 32 birth register 01.12.20 05 kin report 1011 benjamin birth register male selfreport cardiff drivers license 01.08.19 29 selfreport 02.06.20 06 death register table g: example wide format enriched representation full complex observation multidimensional formats multidimensional data structures, also known as data cubes and multi-indexes, are often used to organize and view large data sets. the axes in a multidimensional data structure are properties (characteristics) of groups of observations, and a specific observation is identified by specifying the intersection of a set of dimensions. gross national product, for example, may be indexed by nation and year. multidimensional data structures are often used by official statistical agencies for measures obtained by aggregating over persons, households, businesses, or other units of observation. for example, the average number of persons per household may be indexed by region, urban/rural residence, and categories of household income. however, multidimensional format is also used for non-aggregated data, such as environmental values that can be located in space and time. for example, sea water temperature may be indexed by longitude, latitude and date. some air quality components offered by the copernicus atmosphere monitoring service add a vertical component in addition to longitude, latitude and date, providing values calculated from multiband satellite imagery. in this paper, we focus only on those multidimensional data structures providing aggregated data. we extend our example by showing how the individual-level data in table g can be converted from wide to multidimensional format through aggregation. table h provides an example in which persons in a wide data structure, like table e, are counted by age and gender. the construction of table h from table e involves several data transformation steps before aggregation. since dimensions must consist of mutually exclusive categories, properties with continuous value ranges must be transformed into related properties with discrete values. we have recoded age into two categories, young and old. we provide standardized syntax for describing data transformations based on structured data https://doi.org/ 25/39 alter, george; rizzolo,flavio; and schleidt, kathi (2023) view points on data points, iassist quarterly 47(1), pp. 1-39. doi: https://doi.org/10.29173/iq1051 transformation language in the appendix. in contrast to the transformation from long to wide formats above, aggregation inherently results in the loss of information, and extraction of the primary data is no longer possible. age young old gender male 3 7 female 6 4 table h: example multidimensional format: number of persons by age and gender when data from table a or e are converted to table h, we perform a semantic transposition that is independent of the process of aggregation. consider the characteristic value “6” in the southwest corner of table h. the characteristic measured as “6” is the number of young, female persons in the set of observations covered by table h. we know this from the title of table h, “...number of persons by age and gender,” not from anything in the table itself. in the long and wide formats “female” was a characteristic value, but here it is a location on the dimension gender. notice that table h has labels for rows as well as columns and that each dimension has two levels of labels, a dimension name (age) and a category name (young). we will refer to “female” and “young” as facets to distinguish them from characteristic values. since labels are part of the descriptive framework of a data structure, facets are not data. “number of young, female persons” is a characteristic, which is composed of a measure (number of persons) and two facets (young and female). as we saw above, labels can be arbitrary tokens that point to descriptions, which are provided in table i. we call table i a dimension description data structure to distinguish it from the variable description data structure that accompanies wide format. table i is linked to table h by a compound key consisting of both the dimension and facet columns. in multidimensional format, “female” has transitioned from data to description. as a row label, it is part of the compound characteristic “number of young, female persons.” https://doi.org/ 26/39 alter, george; rizzolo,flavio; and schleidt, kathi (2023) view points on data points, iassist quarterly 47(1), pp. 1-39. doi: https://doi.org/10.29173/iq1051 dimension facet description gender male identified as “male” by source female identified as “female” by source age young younger than age 15 old age 15 or older table i. dimension description data structure for multidimensional data figure 12 follows the characteristic “gender” and the characteristic value “female” from long to wide to multidimensional format in smaller steps to illustrate the semantic transpositions taking place. in long format (figure 12.a) “gender” and “female” are both instance values in the data, and they are clearly identified as a characteristic and characteristic value by their location in the “property” and “value” columns respectively. when we move to wide format (figure 12.b), “gender” is semantically transposed to become a column label, the meaning of which is explained in the variable description data structure, while “female” remains a characteristic value in the data array. figures 12.c and 12.d move from wide to multidimensional format in two steps. the first step (figure 12.c) converts the characteristic gender into two characteristics female and male. unlike gender, which has a value domain with text values (“female”, “male”), the characteristic values of the female and male characteristics are either true or false (shown as 1 and 0), but this does not reduce the information content in the table. when we perform a similar transformation by converting birth and death dates to true/false values for young and old, we are losing information by converting exact dates and ages into broader categories. the data structure shown in figure 12.c will be unfamiliar to most readers, because it is rarely explicit. most software designed to operate on wide format data can go from figure 12.b to figure 12.d in one step, as we show in the appendix. we include figure 12.c, because it shows the semantic transposition of “female” from a characteristic value to a characteristic without aggregation. in the next step (figure 12.d), we aggregate by counting the number of persons in each of the four possible combinations of gender and age. at this stage, the identities of individuals are subsumed under a new characteristic (count), which is the number of persons with each combination of gender and age group. since values of gender and age group (figure 12.d) uniquely identify values of count, they have become identifiers that form composite keys. https://doi.org/ 27/39 alter, george; rizzolo,flavio; and schleidt, kathi (2023) view points on data points, iassist quarterly 47(1), pp. 1-39. doi: https://doi.org/10.29173/iq1051 in terms of information content, figure 12.d and figure 12.e are identical. at this stage the difference between wide and multidimensional is in the capabilities of the software in which they are implemented. software designed for n-cubes and multi-indexes treat identifiers (e.g., gender and age) as dimensions that facilitate the selection of individual characteristic values or subsets of characteristic values. in ddi-cdi gender and age (figure 13) are designated dimension components to reflect this additional functionality. for example, figure 12.e can be sliced by gender to extract a subset of males by age group. although our example has only two dimensions, gender and age group, we could have added more dimensions, like reference area, to produce a 3, 4, or higher dimensional table. higher dimensional tables have practical applications in data retrieval, but we use only two dimensions to simplify our presentation. figure 12. the trajectory of characteristic “gender” and characteristic value “female” from long to wide to multidimensional format figure 12.c: wide format after recoding figure 12.e: multidimensional format figure 12.b: wide format figure 12.d: wide format after aggregation figure 12.a: long format https://doi.org/ 28/39 alter, george; rizzolo,flavio; and schleidt, kathi (2023) view points on data points, iassist quarterly 47(1), pp. 1-39. doi: https://doi.org/10.29173/iq1051 the differences between figure 12.d and figure 12.e are clearer when we view them from the perspective of the data user. a user viewing figure 12.d through a spreadsheet or statistical analysis package will see “female” and “male” (as well as “young” and “old”) as characteristic values under the “gender” (or “age”) characteristic. when software enables the multidimensional aspect of figure 12.e, the user sees “female” and “male” as facets on the “gender” dimension within the descriptive frame of the dataset, where they can be combined into keys to identify subsets of data, like “young females”. the transition from wide format to multidimensional format is also depicted in figure 13 using ddicdi components to identify the roles of variables in each format. as we saw in figure 12, one variable, gender, moves unchanged to multidimensional format, and two new variables, age and number of persons, are derived from variables that appear in wide format. age is derived from dates of birth and death as described above. gender and age, which would have been measure components in wide format are dimension components in multidimensional format. figure 13 describes number of persons as a count of values of person id, as one would in an sql aggregation command. however, counts occur at the intersection of dimensions as in an sql group by clause. (see appendix for sdtl notation.) figure 13 ddi-cdi representation of multidimensional format table j provides a subset of the variable description data structure for the multidimensional data in h. we describe here two variables, “count” and “gender”, each of which has three properties (name, description, and valuedomain). we also show that the values within a valuedomain may have names and descriptions. the information represented in table j is more complex than previous tables, and we present it as “nested name-value pairs.” we use “nested name-value pairs” to refer to nonrectangular data structures such as xml and json. metadata in standards like ddi, sdmx, and ecological markup language (eml) are often shared in xml or json. ddi-cdi does not describe nonhttps://doi.org/ 29/39 alter, george; rizzolo,flavio; and schleidt, kathi (2023) view points on data points, iassist quarterly 47(1), pp. 1-39. doi: https://doi.org/10.29173/iq1051 rectangular data arrays like table j, but we use concepts from ddi-cdi to represent table j in figure 14 below. a json representation of such a structure is provided in appendix 3 of this document. variable name “count” descriptivetext “number of persons” valuedomain (set of non-negative integers) variable name “gender” descriptivetext “gender as reported in source document” valuedomain code notation “male” descriptivetext “identified as ‘male’” code notation “female” descriptivetext “identified as ‘female’” table j. variable description data structure for multidimensional format the simplest name-value pairs consist of a property and a string, such as name: “count”. however, a group of name-value pairs can be nested inside a value. to understand this data structure, it helps to read table j from right to left. the first three lines of the table present three simple name-value pairs: name: “count”, descriptivetext: “number of persons”, valuedomain: (set of non-negative integers). taken together, these three pairs are the value for the first variable, the property named in the leftmost column. the variable named “gender” is described with three levels of nesting. reading from right to left and bottom to top, the notation and descriptivetext properties are simple name-value pairs, which are nested in a code. codes are nested inside a valuedomain, and valuedomain is nested inside a variable. the nesting of codes inside a valuedomain makes “gender” more complex than “count”, but both variables have the same three properties: name, descriptivetext, and valuedomain. in figure 14 we add logical data structures for the descriptive information required to interpret a multidimensional data structure. the panel in the center shows the cube data structure, which is the data array illustrated in figure 12e and the outcome of the procedures shown in figure 13. the data consists of three variables. gender and age are dimension components, and count is a measure component. these variables are linked to variable descriptions through the headers which accompany the data array. in other words, the meaning of the values measured in a data cube must be defined in a different data object, such as a metadata file in sdmx or ddi format. https://doi.org/ 30/39 alter, george; rizzolo,flavio; and schleidt, kathi (2023) view points on data points, iassist quarterly 47(1), pp. 1-39. doi: https://doi.org/10.29173/iq1051 the panels on the left and right of figure 14 represent the variable description data structure found in table j. notice that the pink boxes at the top are the names in the name-value schema used in table j. the panel on the left of figure 14 describes count, the measure component in the cube data structure. recall that count is the name of the variable created by aggregating over rows in groups defined by values of gender and age. count appears in the header of the data array produced by the aggregation step (figure 12d), although it is not shown in our illustration of a data cube (figure 12e). to add descriptive information about the measure in this data cube (center panel), we link count in the header of the cube data structure to count as the name of a variable in the variable description data structure (left panel). count has two other properties, descriptive text and value domain, which are linked to name: “count” by descending from the same variable. figure 14. multidimensional data structure with variable description data structure the panel on the right illustrates the description of a facet (“male”) within a dimension (“gender”). note that the row and column headers for table 11e have two levels, which are both found on the right side of figure 14. the outer headers (“gender” and “age:) are the names of variables serving as dimensions. the inner headers (“male”/ “female” and “young”/ “old”) are facets of the data cube, which are codes within the valuedomains of their respective variables. the nesting of codes within a valuedomain, which we showed in table j is also present in figure 14. properties for the code named “male” are linked to the valuedomain of variable “gender”, and any number of codes may be part of a valuedomain. converting these data to multidimensional format is an additional step in the semantic transposition of datapoints from the data array to the descriptive frame. we showed above that characteristics (e.g., “gender” and “name”), which were data in long format, become metadata in wide format. in this section, we showed characteristic values (e.g., “female”, “young”) moving from the data array to the descriptive frame as facets in multidimensional format. the functions that facets perform are not identifiable from the data, but from the descriptive frame (metadata) associated with the software. users recognize that gender rendered as a dimension is a different view of the same underlying data as gender presented as a measure. https://doi.org/ 31/39 alter, george; rizzolo,flavio; and schleidt, kathi (2023) view points on data points, iassist quarterly 47(1), pp. 1-39. doi: https://doi.org/10.29173/iq1051 long format revisited we now apply to long format two insights from our discussion of wide format. first, we show that long format also has an implied variable description data structure. the column headers in long format can be arbitrary text that is explained in an accompanying document or dataset. second, instance values in a long format data structure may point to explanations in the variable description data structure. moreover, the variable description data structure may include global resources by using uris as instance values. this means that the variable description data structure may not be a single physical data file. it could be an array of resources, including published documents, distributed web services, and a history of common practices and traditions. figure 15. long data structure with variable description data structure there are two things to note in figure 15, which shows a long data structure with a related variable description data structure similar to the one that we showed in figure 11 for wide format. first, among the pink ovals showing the column headers on the left side of the figure, we have used “var10” as a column header. “var10” is then identified as “source” in the variable description data structure on the right side of the figure. as we saw with wide format in figure 10, the column headers in a data table may not provide information about the meaning of values in that column. in figure 12, the meaning of “var10” is explained in an accompanying variable description data structure, which is not limited to the technical requirements of a column header. “var10” could be associated with other attributes, like a definition, citation, instrument model, etc. we also see that the characteristic (variable descriptor component) measured in figure 15 is given as “http://.../gender” with a value of “http://.../gender/female” (see the second and third blue boxes from the left). these uris are resolved to “gender” and “female” in the variable description data structure on the right side of the diagram. we use this to show not only that a characteristic in a long data structure can be resolved in the variable description data structure, but also that the variable description data structure could be a web service rather than a data set. when the variable “gender” is a link to an online controlled vocabulary, it can resolve not only to a definition of the variable but also to a value domain. in addition, the landing page for the controlled vocabulary can contain a link from the variable to the concept measured by the variable. https://doi.org/ 32/39 alter, george; rizzolo,flavio; and schleidt, kathi (2023) view points on data points, iassist quarterly 47(1), pp. 1-39. doi: https://doi.org/10.29173/iq1051 connecting the dots as the need to share data across scientific domains increases, differences in languages and practices for constructing and describing data become more important. this paper offers concepts and terminology that can bridge these differences. we bring together two traditions, the observation and measurement model (iso 19156), which offers a rich relational framework for contextualizing the data creation process, and the elaborate descriptive structures utilized in the ddi and sdmx standards. in particular, we build upon the new ddi-cross domain integration model. we use ddicdi to characterize three common data formats long, wide, and multidimensional. however, describing the format of the data is insufficient without simultaneously describing how information about the data is provided, and the o&m and ddi/sdmx traditions diverge radically in that respect. we call the process of changing the representation of both data content and data description “semantic transposition,” and we show that the ddi-cdi model can be applied to data description as well as data content. in other words, we show that the boundary between data and metadata is flexible and permeable. in our terminology, the key difference between long format data and wide or multidimensional format data is the handling of the characteristic associated with a characteristic value. a characteristic value is the outcome of an observation, i.e., a number or descriptive text, that is associated with the characteristic, i.e., a property or attribute. in long format, which we have used to represent the o&m approach, the characteristic is included in the data array where the characteristic value is found. ddi and sdmx were created to annotate wide and multidimensional formats, where the characteristic is considered “documentation” to be provided separately from the “data.” semantic transposition occurs when re-formatting data implies moving the characteristic from the data content to the data description frame or vice versa. we bridge this gap by showing that the wide and multidimensional formats imply the existence of a parallel variable description data structure linking each characteristic value to a characteristic. the variable description data structure is itself a data array that can be described by the ddi-cdi model. the central contribution of the ddi and sdmx standards is to convert the variable description data structure from the realm of paper into a machine-actionable object. thus, we can trace both the characteristic and the characteristic value as the data are transformed from long to wide to multidimensional. we are not arguing that the ddi-cdi standard needs to develop a separate specification for metadata. rather, ddicdi should be applying the same concepts to data structures containing data and metadata. the terminology provided here is designed to simplify that process by avoiding terms that are used in ambiguous and inconsistent ways, like “attribute”. from our point of view, characteristic values (data) and characteristics (metadata) are both datapoints that can be acted upon by machines as well as people. when we translate data across disciplines, we must recognize that semantic transposition is often reorganizing structures but maintaining the content of the entire dataset -data and metadata. https://doi.org/ 33/39 alter, george; rizzolo,flavio; and schleidt, kathi (2023) view points on data points, iassist quarterly 47(1), pp. 1-39. doi: https://doi.org/10.29173/iq1051 data stewards in the social sciences should be aware that reliance on wide format data is a disadvantage in a world moving toward fair. as we showed, wide format does not have a way of associating qualifiers with characteristic values. since there are no inherent relationships among columns in a wide data structure, nothing links a measure of data quality with the variable that it describes. the only way to make this connection is in the metadata. even if the relation between these variables is well described in the metadata, parsing xml metadata is not in the skill set of social science researchers. in effect, this requires human intervention and prevents fully automated analysis, which is one of the goals of fair. recognition of the fluid boundary between data and metadata is essential for achieving the interoperability promised by the fair principles. standards, like o&m, ddi, and sdmx create different ways of encapsulating information and place different boundaries between “data” and “metadata.” as we have shown, interoperability often requires semantic transposition, moving information from “data” to “metadata” or the reverse. thus, interoperability requires mappings between structures that have different understandings of what information belongs in “data” and “metadata.” ddi-cdi has created a language for describing these mappings, but it should be extended to support semantic transposition. in practice, the creation of these mappings is further inhibited by different data cultures based on incompatible vocabularies. our domain-independent vocabulary is intended to enable this much needed cross-domain conversation. references beine, m., hames, n., weber, j. h., & cleve, a. (2014) ‘bidirectional transformations in database evolution: a case study at scale’, paper presented at the edbt/icdt workshops. britell, s., delcambre, l. m., & atzeni, p. (2016) ‘facilitating data-metadata transformation by domain specialists in a web-based information system using simple correspondences’, paper presented at the international conference on conceptual modeling. cox, s. j. d. (2011) ‘iso 19156:2011 geographic information – observations and measurements’, retrieved from http://doi.org/10.13140/2.1.1142.3042 ddi alliance. (2020a) ‘ddi-cross domain integration: detailed model’, retrieved from https://ddialliance.org/specification/ddi-cdi ddi alliance. (2020b, december 1, 2020) ‘structured data transformation language’, retrieved from https://ddialliance.org/products/sdtl/1.0 hernández, m. a., papotti, p., & tan, w.-c. (2008) ‘data exchange with data-metadata translations’, proc. vldb endow., 1(1), 260-273. doi:10.14778/1453856.1453888 iso 19156:2022 (2022) ‘observations, measurements and samples’, forthcoming. iso/tc 211 terminology maintenance group. (2020) ‘iso/tc 211 multi-lingual glossary of terms: entity’, retrieved from https://isotc211.geolexica.org/concepts/1948/ https://doi.org/ https://ddialliance.org/specification/ddi-cdi https://ddialliance.org/products/sdtl/1.0 https://isotc211.geolexica.org/concepts/1948/ 34/39 alter, george; rizzolo,flavio; and schleidt, kathi (2023) view points on data points, iassist quarterly 47(1), pp. 1-39. doi: https://doi.org/10.29173/iq1051 olivé, antoni. (2007) ‘conceptual modeling of information systems’, springer berlin / heidelberg. papotti, p., & torlone, r. (2009) ‘schema exchange: generic mappings for transforming data and metadata’, data & knowledge engineering, 68(7), 665-682. doi:https://doi.org/10.1016/j.datak.2009.02.005 provenance working group. (2013) ‘the prov namespace’. retrieved from http://www.w3.org/ns/prov#entity san gil, i., vanderbilt, k., & harrington, s. a. (2011) ‘examples of ecological data synthesis driven by rich metadata, and practical guidelines to use the ecological metadata language specification to this end’, international journal of metadata, semantics and ontologies, 6(1), 46-55. statistical data and metadata exchange (sdmx). (2013, november 11, 2021) ‘iso 17369:2013 statistical data and metadata exchange (sdmx)’. retrieved from https://www.iso.org/standard/52500.html statistical data and metadata exchange (sdmx). (2021) ‘sdmx’. retrieved from https://sdmx.org/ vardigan, m., heus, p., & thomas, w. (2008) ‘data documentation initiative: toward a standard for the social sciences’, international journal of digital curation, 3(1). wilkinson, m. d., dumontier, m., aalbersberg, i. j., appleton, g., axton, m., baak, a., blomberg, n., boiten, j.-w., da silva santos, l. b., bourne, p. e. (2016) ‘the fair guiding principles for scientific data management and stewardship’, scientific data, 3. wyss, c. m., & robertson, e. l. (2005) ‘a formal characterization of pivot/unpivot’, paper presented at the proceedings of the 14th acm international conference on information and knowledge management, bremen, germany. https://doi.org/10.1145/1099554.1099709 xue, j., shen, d., nie, t., kou, y., & yu, g. (2013, 10-15 nov. 2013) ‘inferring and propagating pivot dependencies in schema transformation between data and metadata’, paper presented at the 2013 10th web information system and application conference. https://doi.org/ https://doi.org/10.1016/j.datak.2009.02.005 https://sdmx.org/ 35/39 alter, george; rizzolo,flavio; and schleidt, kathi (2023) view points on data points, iassist quarterly 47(1), pp. 1-39. doi: https://doi.org/10.29173/iq1051 appendix 1: semantic transposition semantic transposition: (aka flip-flop) in refactoring from a conceptual model, the lateral transposition of the representation of a characteristic from the data structure to the data content, that retains isomorphic coherence between representations semantic transposition is the concept, semantic transposition refactoring (str) the process ● simple: { "@context": { "wind-speed": "http://...wind-speed", "foi": "http: //...ogc/o&m/foi", "geom": "http://...ogc/geom" }, "foi": 1, "geom": "xxx", "wind-speed": 5 } ● complex: { "@context": { "foi": "http: //ogc.../o&m/foi", "observedproperty": "http://...ogc/obsprop", "result": "http://...ogc/result", "geom": "http://...ogc/geom" }, "foi": 1, "geom": "xxx", "observedproperty": "http://...wind-speed", "result": 5 } https://doi.org/ http://...wind-speed/ http://...ogc/geom http://...ogc/obsprop http://...ogc/result http://...ogc/geom http://...wind-speed/ 36/39 alter, george; rizzolo,flavio; and schleidt, kathi (2023) view points on data points, iassist quarterly 47(1), pp. 1-39. doi: https://doi.org/10.29173/iq1051 appendix 2: data transformations and aggregations as described above in the section on multidimensional formats, to obtain counts within categories, such as age groups, properties that provide a continuous value range must be transformed or recoded to a related property represented by a discrete set of values. once all relevant properties have been transformed to ones suitable for grouping, the aggregation can be performed. we use a simplified version of structured data transformation language (sdtl; ddi alliance, 2020b) to describe data transformations. sdtl provides machine-actionable descriptions (i.e., provenance metadata) of variable-level processes like this. (see https://ddialliance.org/products/sdtl/1.0) transformation in some cases, the properties being provided for individuals circumscribe the actual property of interest. in our primary demographic data, we have dates of birth and death, but we are actually interested in the age attained by the individual at the time of death. age at death is implicit in the data, and we must perform a calculation on the values provided for birth and death dates to obtain a new property for each individual. sdtl for computing age at death from birth and death dates: command: compute variable: age expression: function: division argumentname: exp1 argumentvalue: function: subtraction argumentname: exp1 argumentvalue: died argumentname: exp2 argumentvalue: born argumentname: exp2 argumentvalue: timedurationconstant: timedurationvalue: “p365.25d” classification, recoding and transformation under “recoding” we understand the process of assigning discrete values, usually from a classification system, to serve as a representative of a property that has been provided by a continuous value range. a simple example of this concept pertains to the age of an individual calculated above from the dates of birth and death. in order to provide a discrete age axis within our multidimensional representation of the data, this property must be transformed to a related property with a discrete value range. in our example, the final age_group property is defined with the following values: ● child: age <= 14 years https://doi.org/ 37/39 alter, george; rizzolo,flavio; and schleidt, kathi (2023) view points on data points, iassist quarterly 47(1), pp. 1-39. doi: https://doi.org/10.29173/iq1051 ● adult: 15 <= age <= 64 ● old: age >= 65 recoding is simply a matter of determining which of the defined groups the value provided for the individual belongs to, this value is then assigned to the individual via the new age_group property. sdtl for recoding into age groups: command: recode recoded variables: source: age target: age_group rules: recoderule: fromvalue: 0 to: 14 label: child recoderule: fromvalue: 15 to: 64 label: adult recoderule: fromvalue: 65 to: numericmaximumvalueexpression label: old aggregation sdtl description of the aggregation process: command: collapse groupbyvariables: gender, age_group aggregatevariables: compute count = col_count(name) producesdataframe: dataframedescription: dataframename: cube20 rowdimensions: gender, age_group https://doi.org/ 38/39 alter, george; rizzolo,flavio; and schleidt, kathi (2023) view points on data points, iassist quarterly 47(1), pp. 1-39. doi: https://doi.org/10.29173/iq1051 appendix 3: json representation variable description data structure for multidimensional format the variable description data structure for multidimensional format could be represented in json as shown below. { "variable": { "name": "count", "descriptivetext": "number of persons", "valuedomain": ["(set of non-negative integers)"] } }, { "variable": { "name": "gender", "descriptivetext": "gender as reported in source document", "valuedomain": [{ "code": { "notation": "male", "descriptivetext": "identified as ‘male’" } }, { "code": { "notation": "female", "descriptivetext": "identified as ‘female’" } }] } } https://doi.org/ 39/39 alter, george; rizzolo,flavio; and schleidt, kathi (2023) view points on data points, iassist quarterly 47(1), pp. 1-39. doi: https://doi.org/10.29173/iq1051 endnotes 1 george alter is research professor emeritus in the institute for social research at the university of michigan. he can be reached by email: altergc@umich.edu. 2 flavio rizzolo is senior data science architect for statistics canada. he can be reached by email: flavio.rizzolo@statcan.gc.ca. 3 kathi schleidt is a data scientist with a specialty in environmental informatics. she is the founder of datacove e.u. she can be reached by email: kathi@datacove.eu. 4 ideally, this should be a uri providing further information such as possible values the methodology can provide. https://doi.org/ mailto:altergc@umich.edu mailto:flavio.rizzolo@statcan.gc.ca mailto:kathi@datacove.eu vol21.1 13spring/summer 1994 the new unit was established on november the 1st 1996. years of work and planning by the danish data archives finally led to its formation half a year ago. sadly the initiator, the director of the danish data archives, per nielsen, died less than two months after it’s formation. the purpose of the unit is to strengthen the registration and storage of medical research data in denmark. the new unit will provide a professional storage function to the medical research community and promote access to collections of medical research data for secondary analysis. the unit is established as a 5-year collaboration project between the danish national research foundation and the danish data archives. in an effort to strenghten the research development capabilities in denmark, the danish national research foundation was established in 1991. since its formation the foundation has worked to improve the conditions for research development in denmark in the entire scientific landscape. this is achieved by giving large and concentrated grants to unique danish research at the international level. several new projects have thus been started, the establishment of the danish unit for registration and storage of medical research data being one of the most recent. denmark has a unique tradition for keeping and maintaining population based registers, which places the new unit in an almost ideal environment for developing a professional archive for medical research data. the unit is physically located at the danish data archives, as a fully integrated part of the institution. the aquisition, processing and subsequent storing of medical research data closely follows the principles already developed at the danish data archives. a simple copying of the existing routines is not enough though, as the, at times, different nature and sheer volume of medical research data dictates the development of new and refinement of existing routines. since its establishment in 1973, the danish data archives have collected data and documenta tion from social science and historic studies, and until now the institution has concentrated it’s efforts on these areas of research. consequently medical research studies constitutes only a small fraction, about 6% of the total contents, primarily in the form of social medicine and occupational health studies. about 8000 articles are published every year by danish medical researchers. although not all of these articles represent a seperate research study, even a conservative estimate of 2 articles per study gives a yearly volume of 4000 studies. add to this a considerable backlog of old studies and you begin to realize the magnitude of the task. this necessitates the development of new, timesaving procedures for registration and storage of data and documentation. an important consideration in the development work is to ensure that the high quality of registra tion and storage of data and documentation already attained at the danish data archives is preserved. the staff of the unit at present consists of 3 persons: the project leader, medical doctor and ph.d. kirsten kyvik, who also is a member of the management group for the danish twin register, senior researcher, medical doctor and specialist in community medicine and occupational health, peter heine jorgensen and informatics assistant, birgit wich. it is the first time that medical staff has been employed at the danish data archives and indeed at the danish state archives. the reason of course being inherent in the nature of the project. the logic being that a medical staff is best suited to deal with medical research data and documentation. furthermore it is expected that this will counteract any reluctancy or mistrust on the part of the donor of the medical research data. the employment of medical doctors in the staff also solves the confidentiality issue often associated with handling medical research data. another task presents itself to the unit regarding confidentiality. the data may contain sensitive, identifiable information about the persons in the study base. the danish data protection agency under the danish justice department controls and regulates the legal aspects of data storage. the agency has granted the unit permisson a presentation of the eras project: the danish unit for registration and storage of medical research data by peter heine jorgensen* 14 iassist quarterly to store sensitive data from medical research. furthermore permission to release the original data and documentation to the donors is given. thus the primary researchers are able to continue their study on the same study base, even years after completion of the first study. naturally other researchers will only have access to the data in an anonymous form. the unit is at present engaged in an effort to persuade the data protection agency to accept that storage at the unit can be regarded as being equal to deletion. it is hoped that this can be stipulated in the standard agreement made between the agency and the medical researchers. the target of the unit is medical research in denmark and existing as well as coming danish medical researchers. a broad acceptance and support from the medical research community is instrumental in achieving a succesful result, ie. the creation of professional archives for medical research data. consequently the initial action will focus on information and dialog. agreement regarding submission of research data to the new unit will be investigated in accordance with the danish medical research community, the scope of the danish unit for registration and storage of medical research data is to assist researchers and research institutions in making data and documentation from studies readily available. the new unit will perform it’s own research and development with the purpose of making the new archive as well functioning as possible, thus facilitating secondary analysis. the possibility of new crosslinks between the existing social science and historic data and the new medical research data represents a unique opportunity to perform secondary analysis bridging several major research fields. the units field of activitivity is at present confined to danish medical research, but the findings could have an international bearing as the principles and procedures of the fully established unit could be implemented in other countries. a more concerted implementation could be possible in the framework of the european union. * paper presented at iassist/ifdo ‘1997 conference, may 6th may 9th odense, denmark. peter heine jorgensen, m.d., specialist in community medicine and occupational health. 20 iassist quarterly 2016 / vol 40 no 4 iassist quarterly abstract for social and economic researchers, many useful but previously unavailable sources of data have become at least potentially accessible in recent years. these ‘new and novel’ forms of data (nnfd), such as social media data or smart meter data, represent potentially invaluable resources for researchers, but pose challenges for access provision and analysis. this short article introduces data service as a platform (dsaap), a project currently underway at the uk data service to establish a technological infrastructure supporting data archivists and social and economic researchers in managing and analysing both familiar and new and novel forms of data. it presents an overview of nnfd in social science contexts, introduces the dsaap system, and sketches short examples of dsaap capabilities in analysing nnfd, drawn from an associated ukds project smarter household energy data: infrastructure for policy and planning, before concluding with some reflections on the potential value added to social scientific research by data service as a platform. keywords big data, data science, social science, data archiving, hadoop introduction the uk data service2 (ukds) big data network support3` (bdns) team is currently engaged in a major project to develop data service as a platform (dsaap), a technological infrastructure supporting data archivists and social and economic researchers in managing and analysing both familiar and new and novel forms of data (nnfd). nnfd encompasses very large datasets, sometimes referred to as ‘big’ data, but not all, or all aspects of nnfd are necessarily ‘big’ in this way. this short paper presents an overview of nnfd in social science contexts, introduces the dsaap system, and sketches short examples of dsaap capabilities in analysing nnfd, drawn from an associated ukds project smarter household energy data: infrastructure for policy and planning,4 before concluding with some reflections on the potential value added to social scientific research by data science techniques and data service as a platform. an accompanying video demonstration can be viewed here.5 new and novel forms of data although the term ‘big data’ has been popularised in recent years, nnfd are not just about the size of the files, but about a broader view of what, how, when, and where data can be collected, stored, linked and analysed to further research6 (oecd, 2013). many new servicing new and novel forms of data: opportunities for social science by aidan condron1 data sources on social and economic activity have emerged such as social media, smart meter and other household consumption data, internet usage, sensor and footfall readings, and many others. while much, if not all, of this this data is not collected specifically for the purposes of social and economic research, there many exciting possibilities for reuse by researchers, presenting a potentially invaluable resource. even if some of these data have been collected for some time now, they represent new and novel forms of data in a social science context. some of these datasets are very large, but not all, or all aspects are necessarily ‘big’ in this way. in any case, the most advantageous approach to ‘big data’ is often to downscale or reduce to manageable sizes, whether by aggregation or mining for the more valuable elements (siems & wolf, 2007). smart meter data, for example, contains a huge number of readings, but these are more useful in context of associated geodemographic data on much smaller numbers of households from which the readings are drawn. analysis and findings are enriched by linking to other data sources such as meteorological data, which does tend to be voluminous, and housing stock and energy performance certificates, which are more compact. in making a case for analysis of nnfd as a progressive factor in social science, and in thinking big about data, it’s not just a matter of scale, but of innovation in assessing what data sources can be harnessed to answer research questions, and how linking and triangulation can enhance analyses and findings. a big data approach doesn’t always mean using massive files and processing power! data service as a platform dsaap our concept of dsaap is a ‘data lake’ system capable of storing data of any kind, which will cater for both traditional data and nnfd drawn from current and future ukds holdings. the data lake is a secure, format-agnostic repository providing a powerful, scalable, suite of tools for data processing. it is built on hadoop,7 a software framework that facilitates high-speed processing of datasets of any size by distributing data across networks (referred to as clusters) of linked computers (referred to as nodes). kdnuggets, a leading data science website, offers the following working definition of a data lake.8 a data lake is a storage repository that holds a [potentially] vast amount of raw data in its native format, including structured, semi-structured, and unstructured data. the data structure and requirements are not [necessarily] defined until the data is needed’. this differs from data warehouses, where data is strictly formatted and structured to meet specific, pre-defined reporting functions. vol 40 no 4 / iassist quarterly 2016 21 iassist quarterly developing dsaap is a staged, medium-term enterprise, involving installing cloud-based and on-premises hadoop infrastructure, establishing ingest pipelines to populate the data lake (which will recombine, store and tag datasets in a resource descriptive framework (rdf) triplet9 and key-value pair format) and providing access channels and user endpoints for researchers to work with data. data stored within it are assigned randomly machine generated globally unique identifiers10 (guid)s at the lowest possible level of granularity, down to the field, record, or even data point level. by drawing on w3c standardised vocabulary services,11 this tagging facilitates dynamic reassembly, interlinking, and querying of data according to user requirements. access channels and endpoints are designed to cater for user requirements, and are determined by engagement with the research community. the approach adopted in developing dsaap is pro-active, moving from making data available through downloads bundles from catalogues to scoping and providing solutions for secure data access, management, linking, and analysis within the dsaap environment. approved researchers should be able to log into dsaap work areas with access to the data they will work with, and as functionality develops, dsaap will facilitate self-service for researchers and other users, unifying data tools and querying, linking and analysing data across a complex environment. the strategic impetus driving the project is to create a twenty-first century data solution for the ongoing data access and curation communities. in developing the dsaap, bdns has adopted the philosophy of the open data platform initiative (odpi),12 focusing on developing a standardised data working environment facilitated entirely through open-source software. in referring to an ‘open data platform’, this does not mean that the platform support is limited to open data. openness refers to the data lake’s open source software build, and to the ongoing cross-community technological development of which the dsaap itself is a part and an exemplar, and in which dsaap developers and users are members. support will be provided for users across a wide range of technical expertise, catering for novices who needs to navigate and query data using point and click or drag and drop interfaces, researchers interested in applying traditional techniques such as linear regression or analysis of variance (anova), analysts interested in employing machine learning algorithms such as random forest, nearest neighbours, or text mining techniques, up to technologists or developers who want to develop their own bespoke data tools. in keeping with the ukds public service ethos and with the spirit of open science, it is hoped that researchers using the dsaap will participate in knowledge sharing and functionality development by contributing to shared repositories of software code and analytical techniques. while capable of scaling as necessary to accommodate nnfd, all dsaap-based data storage and access provision will be governed by established research data management (rdm) principles which underpin all ukds archiving and curation work.13 ukds data is protected by enterprise-grade security and governance, with access governed by the ukds three-tier classification of open, safeguarded, and controlled data.14 if sensitive data is ingested to dsaap, research will be regulated within the ‘five safes’ 15 framework of safe people, safe projects, safe settings, safe outputs, and safe data. as working with very large datasets or linking multiple datasets can pose risks to anonymity and privacy, rigorous machine actionable statistical disclosure control (sdc) checks should be applied when data is ingested, to determine appropriate levels of access security, and to all potentially disclosive outputs. sample use case: exploratory data analysis with household energy data dsaap capabilities in generating value from nnfd is illustrated by work on smarter household energy data: infrastructure for policy and planning16 (shed), an associated project in partnership with the university of cape town’ datafirst17 and university college london’s energy institute,18 which ‘focuses on data infrastructure and brings together data professionals, energy researchers and policymakers in sa and the uk’. dsaap and shed intersect, as the shed infrastructure component will be provided by dsaap, while shed will provide dsaap use cases, from ingest to endpoint, of data valuable in household energy consumption researchers. shed project work has involved scoping the household energy field through reading and undertaking engagement with the research community, canvassing requirements through roundtable meetings19 and collaboration with individual researchers working on household energy. while dsaap is developing discipline-agnostic, generic systems and tools of value to a wide user base, this association with specific research and datasets facilitates pilot project initiation, data processing test cases, and analytical proofs of concept. more general, generic data analysis work has also been carried out by ukds staff on a large dataset collected during the energy demand research project (edrp), a series of experimental trials involving smart meters on household energy consumption during 2008-2010 which was deposited by the department of energy and climate change (decc) with the ukds for curation in late 2014.20 the edrp data presented an initial technical challenge for assessment and curation, as it included files over 12 gb in size containing hundreds of millions of records, far too big to be opened with familiar desktop software, and indeed too big to be fully loaded into the memory of standard pcs, regardless of the software used. a dsaap prototype facilitated loading, opening, and employing exploratory data analysis (eda) techniques to explore the data.21 pioneered by john tukey in the late 1970s (tukey, 1977) and now accepted as an important component in data science, eda involves initial, non-hypothesis driven, investigation of data, developed by generating summary statistics, plotting variable distributions and time series, and transforming data, and is often used as a crucial first step towards understanding data, particularly large, unfamiliar datasets (marsh, c & elliott 2008, o’neil and schutt, 2013: 34-40). exploring the edrp datasets, which previously seemed impenetrable, is relatively easy with dsaap. geodemographic information on the households included in the study consists of fifteen variables including a household anonymous identifier, data on the types of fuel available to the household (electricity or electricity and gas), energy consumption pricing structure (whether fixed rate or time of use tariff(tout)), geographic region, and socioeconomic status as defined by the acorn classification system. once data is loaded, variables 22 iassist quarterly 2016 / vol 40 no 4 iassist quarterly can be viewed as standard data tables, familiar to anyone who has worked with spreadsheets, spss, or any type of database software, as seen in figure 1 below. apache zeppelin23 is an analytical tool integrated into dsaap which supports multiple language backends, such as python, scala, and structured query language (sql), and provides powerful visualisation tools. the user interface is accessed through standard web browsers such as google chrome or mozilla firefox, meaning that users with dsaap log in credentials need not install any additional software to work with data on the system. with non-controlled data, this can be done from any internet-connected location. the data table view is accompanied by a series of interactive graphic views, where variables can be selected and manipulated with clicks or drag and drop, with zeppelin dynamically generating visualisations such as bar, line, or scatter plots. these types of functionality speed and ease eda and other forms of analysis, particularly when working with nnfd. as a first eda step with the edrp data, univariate, bivariate and multivariate distributions of these variables were rapidly and dynamically visualised using ‘out of the box’ features. figure 2 below shows a bivariate bar chart produced with drag and drop commands, graphing the study population of households by geographical region and types of fuel used. while other software packages can certainly produce bar charts, dsaap facilitates linking to and working with data on a much greater scale. while approximately 14,000 households are included in this study, data was collected at half-hourly intervals over a thirty month period, generating 413,000,000 records on electricity usage alone (a figure which was itself unknown before loading and counting in dsaap). apache hive,24 a data warehousing interface included with dsaap is configured for aggregation and analysis of large data sets like this. figure 3 below, a visualisation generated by hive, graphs mean household electricity usage over an average twenty-four hour period in december 2009, demonstrates dsaap capabilities in drawing meaning and value from the data. ‘dual’ households with both gas and electricity installed are represented by the blue line, while ‘eleconly’ households with households relying solely on electricity for energy, including heating are represented by the orange line. as might be expected, and can be seen clearly from the graph, dual usage households consume noticeably less electricity on a december day than electricity only households. figure 1 data table viewed in zeppelin figure 2 zeppelin cross tabulation comparing electricity only and dual use households vol 40 no 4 / iassist quarterly 2016 23 iassist quarterly this graph demonstrates some of the power of the dsaap system. the graph represents a series of aggregations drawn from two linked data tables and hundreds of thousands of data points in a simple and easily readable output. figure 4, below is a cross section of a 3 x 12 grid, extending this analysis across the twelve months for the years 2008-2010, an output drawing on over 3.7 billion data points. figure 3 comparing mean energy consumption, december 2009 figure 4 energy curves, june-september, 2008-2010 24 iassist quarterly 2016 / vol 40 no 4 iassist quarterly while household energy curves provide striking visualisations, which are recognised as useful analytical devices in the field (palmer et al. 2014), it should be stressed again that every visualisation is based on a data table, which is produced by querying the underlying data. the graphs above are generated by selecting, subsetting, pivoting and plotting specific variables from edrp, all standard hive features, and display dsaap’s power to recombine and represent data at different levels of analysis, from high-level national and annual aggregations down to the fine granularity of single households and hourly intervals. the energy curves are based on linking two data tables, one with data on household energy consumption, and some with data on the households themselves. both tables were included with the edrp dataset, and the common anonymised household identifier facilitated easy linkage. linking to other data sources is also facilitated. figure 5 below, a line over bar graph, plots mean electricity consumption against mean daily temperature in the east midlands region of england, displaying a strong negative correlation, with household energy consumption rising as temperature falls. this graph is based on spatial and temporal aggregations of energy consumption (mean monthly household energy consumption across geographical region) linked to open data from an external source, the meteorological office. despite being derived from a large number representing a modest level of complexity, it presents a clear, readily understandable visualisation of the information. any of the outputs produced including derived data tables, descriptive statistics, and visualisations can be easily stored on dsaap or, security permitting, downloaded for use or reproduction elsewhere. conclusion: new and novel forms of data and opportunities for social science the examples above illustrate some basic dsaap capabilities. visually subsetting by categorical data, aggregating and pivoting large sets, displaying correlation, and linking to internal and external data sources are shown, but assuming availability of data, almost any imaginable analysis, visualisation, or derived data product required by social or economic researchers could be generated. the usefulness of these data products is not determined by the technology, but by the research agenda and design. to return to the household energy consumption example, hourly data at the household level are required for answering questions about daily consumption patterns, but monthly aggregations at district level are more useful for exploring questions on seasonal or regional variations in energy consumption. the useful level of detail or granularity or is useful is determined by analysis performed and research questions asked, as are the appropriate tools to use. the examples above have shown hive’s utility in managing and analysing large numbers of datapoints, and some zeppelin capabilities in dynamic plotting and graphing. interoperability between these and other dsaap features provides for a powerful analytical platform. the key driver of this work is not just to understand these particular datasets, but to develop transferable understanding, expertise, and tooling, intended for wider use in working within the new data environment, from ingest to access and analysis. this is a research community oriented and involved project, scoping interest and requirements within the social science community, and working actively with researchers to develop the most useful systems. community engagement is a two-way process. collaboration with energy researchers has developed use cases based on actual research, while demonstrations of dsaap features whether on video or conducted live have generated interest and ideas on how the systems can be used and how nnfd can be leveraged. moving forward, the big data network support team will standardise and generalise procedures developed from dsaap work. for example, work on assessing, aggregating and visualising time series data developed with reference to the shed project can be adapted and scaled to cover time series data from other areas. this institutional learning will be implemented through dsaap and also disseminated through training, seminars, lectures, conference papers, and knowledge exchange programmes and partnership, all of which are already underway. the overall aim is to establish a general-purpose data services system for social and economic research, supporting both established and emerging analytical techniques, and both traditional and new and novel forms of data. references ‘new data for understanding the human condition: international perspectives’ oecd global science forum report on data and research infrastructure for the social sciences, brussels, february 2013. figure 5 line over bar monthly mean temperature and energy consumption vol 40 no 4 / iassist quarterly 2016 25 iassist quarterly marsh, c & elliott, j, ‘exploring data: an introduction to data analysis for social scientists’, polity, cambridge 2008. o’neil , cathy & rachel schutt, ‘doing data science: straight talk from the frontline’, o’reilly media, inc., new york 2013. palmer, et al., ‘further analysis of the household electricity survey energy use at home: models, labels and unusual appliances’, cambridge architectural research limited, cambridge 2014. siems, k & wolf, d, ‘burning the hay to find the needle – data mining strategies in natural product dereplication’ chimia international journal for chemistry, volume 61, number 6, june 2007, pp. 339-345(7) tukey, john w, ‘exploratory data analysis’, pearson, london 1977. notes 1. dr aidan condron | senior officer, collections development and producer relations | big data network support | uk data service | university of essex | wivenhoe park | colchester co4 3sq | t +44 (0) 1206 874254 | e acondron@essex.ac.uk 2. https://www.ukdataservice.ac.uk/ 3. https://bigdata.ukdataservice.ac.uk/ 4. https://www.ukdataservice.ac.uk/about-us/our-rd/smarter-household-energy-data 5. https://www.youtube.com/watch?v=0hbcayuwwdy 6. https://www.oecd.org/sti/sci-tech/new-data-for-understanding-the-human-condition.pdf 7. http://hadoop.apache.org/ 8. http://www.kdnuggets.com/2015/09/data-lake-vs-data-warehouse-key-differences.html 9. https://www.w3.org/tr/rdf11-concepts/ 10. https://betterexplained.com/articles/the-quick-guide-to-guids/ 11. https://www.w3.org/2013/04/vocabs/ 12. https://www.odpi.org/ 13. http://www.data-archive.ac.uk/curate 14. https://www.ukdataservice.ac.uk/get-data/data-access-policy 15. http://blog.ukdataservice.ac.uk/access-to-sensitive-data-for-research-the-5-safes/ 16. https://www.ukdataservice.ac.uk/about-us/our-rd/smarter-household-energy-data 17. https://www.datafirst.uct.ac.za/ 18. http://www.bartlett.ucl.ac.uk/energy/ 19. https://ukdataservicesmartenergydata.wordpress.com/ 20. https://discover.ukdataservice.ac.uk/doi?sn=7591#1 21. https://www.youtube.com/watch?v=0hbcayuwwdy 22. http://acorn.caci.co.uk/ 23. http://zeppelin.apache.org 24. https://hive.apache.org/ 25. http://www.carltd.com/sites/carwebsite/files/report%203_models,%20labels%20and%20unusual%20appliances.pdf vol282-3.indd iassist quarterly summer/fall 2004 39 by louise corti1 introductionintroduction through a number of strategic investments by various funding organisations, the uk academic community has access to a unique and expansive range of digital data resources. the economic and social data service (esds), supported by the economic and social research council (esrc) and the joint information systems committee (jisc) is a national data service. esds provides access and support for an extensive range of key economic and social data, both quantitative and qualitative, spanning many disciplines and themes.2 it comprises a number of specialist data services that promote and encourage data usage in teaching and research. while individual datasets are used extensively in academic research, they are signifi cantly under-used in learning and teaching programmes within higher education (he), at both undergraduate and postgraduate levels, and are rarely used in further education (fe). as a service provider of the esds, the uk data archive (ukda) is in a strong position to offer its data resources to the learning and teaching communities for developing materials that might be more appealing to teachers than raw data. 3 this activity requires advice and input from instructors in the classroom on how to develop the pedagogic aspects of learning resources: which content to extract; how to contextualise and apply raw data; where to position such resources in the learning process; and on the usability and functionality of the digital resources created. this paper describes the ukda survey data in teaching project (sdit), funded under the jisc exchange for learning (x4l) programme.4 the project’s goals were to increase the use of real data sources in the classroom, and in a more ambitious sense, to help improve the data literacy of those studying social sciences, from school students age 16-19 to postgraduates. the project created a set of free teaching and learning data and statistics-oriented resources based on the study of crime in society, and were based on learning strategies that encourage the teaching of research methods within a substantive context. this paper addresses both the positive experiences and challenges that arose from running the project.5 background: under-use of data in the classroom there is a widespread concern in the uk and further afi eld, including in the us, that levels of data literacy are at an all-time low. previous investigating into levels of data literacy reveal that the uk is lacking signifi cantly in a stock of individuals with good quantitative data analysis skills. the jisc 5/99 programme task force on the use of numeric data in the learning and teaching project, centred around how quantitative data are currently used in the classroom, offered valuable insights and evidence into the benefi ts and barriers surrounding the adoption of data handling in social science teaching practise (rice et al. 2001)6. the enquiry looked into the use of numeric datasets in learning and teaching within uk higher education, and into the barriers faced by postgraduate and undergraduate teachers who wish to introduce students to the use of empirical datasets in the classroom. the few he lecturers using data in their teaching used them either to add an empirical dimension to the subject, to teach statistics or data analysis methods, or to teach numeracy or critical thinking skills. certainly, the main focus is on using data to teach research methods, with very few using ‘live’ social data to illustrate substance. for teachers, the barriers to using data cited in the report were largely to do with lack of awareness of data sources, lack of access to suitable data both physically and conceptually, and lack of time available to prepare data and build its use into courses. the ukda has years of experience working with lecturers in the he sector, but mostly in responsive or passive mode. that is, a teacher requests a survey on health of older people for her course on gerontology, and the ukda has previously done little to service the request other than advise on the most suitable choice of data or teaching datasets already prepared (by other teachers and then redeposited). however, user surveys suggest that even those who do access data do not go on to use them in their teaching for the reasons mentioned above. the esrc has also recognised the problem regarding the defi cit of quantitative research methods skills of social science survey data in teaching project (sdit): enhancing critical thinking and data literacy 40 iassist quarterly summer/fall 2004 postgraduates, and has recently introduced strategies to try to redress the situation. an example is the introduction of mandatory quantitative methods training courses for phd students (esrc 2001)7. turning to mathematical education of post-14 school age children in the uk, the recent report, making mathematics count, suggested that current curricula fail to meet the needs of learners and satisfy the requirements and expectations of employers and higher education institutions (smith 2004)8. the department for education and skills (dfes), with guiding support from bodies such as the royal statistical society, is thus seeking to improve ways to give young people confi dence with numbers and in data handling across the curriculm, not only within the confi nes of mathematics education per se. an example of a participatory project for school children is the censusatschool project (2003)9. last year, the tomlinson report (2004), which looked at uk secondary education, proposed major reforms for a new diploma-based system designed to offer more specialized work-related learning that would also better prepare students for higher education10. key to the recommendations was the requirement of all students to complete core communication and numeracy skills elements. evidence from the us suggests the same worrying situation (steen 2004), and various forums have been established, such as the national numeracy network (nnn) by the mathematical association of america (maa), to help support schools and colleges that are exploring ways to infuse quantitative literacy into their curricula11. the department of education has already set up a "fund for the improvement of post-secondary education" (fipse) which, in 2004, awarded carleton college in minnesota an award to support of a three-year project on education in quantitative reasoning12 and, other us initiatives have also become more prominent to help foster statistical literacy at the graduate level, for example, the statlit website13. it is now widely accepted that the benefi ts of working with real life data sources are signifi cant. these include understanding how statistics are created, how published tables and graphs (e.g. opinion polls in newspapers) are interpreted and how to begin to manipulate and analyse data. not only does practical knowledge about survey methods and secondary analysis teach students how research is actually conducted, it informs critical assessment of arguments based on the interpretation of survey data. data are never isolated from theory, and it is never the case that data ‘speak for themselves’. introducing such concepts early on in post-16 education is one way to address the quantitative skills concerns expressed above, and students gain a tangible and marketable skill that they can use in future employment. from a pedagogical point of view, these inquiries provided the driving force behind the sdit project. there are still relatively few instances of good examples of repackaging the rich stock of national data resources amongst the social science teaching communities. the jisc collection of historical and contemporary census data and related materials (chcc) project is one of the fi rst funded activities to look at creating e-learning resources for statistics in context based on census data (chcc 2003)14. the sdit project adds to these initial building blocks. aims of the sdit project the sdit project was funded under the exchange for learning (x4l) programme, which has been motivated by the drive to make the most of the considerable investment that has taken place in a range of jisc resources for teaching and learning. the programme is exploring a range of strategies, methods, tools and metadata standards that will enable the repurposing of e-learning materials. pedagogical outcomes are at the heart of the programme, with a focus on learning activities and outcomes, as is the challenge to elucidate strategies that will encourage sustainability and widespread adoption of e-learning materials. in attempting to meet the programme’s key challenges, the sdit project aimed to consider how repurposing existing data resources housed at the ukda for teaching and learning might increase their use. simplifying and re-packaging complex data, for example, the larger national government survey datasets held by the ukda, is one way of opening up their accessibility to the classroom15. the initial aims of the project were: fi rst to develop, pilot and evaluate a set of survey data-base resources for use in the teaching of social science courses at both he and fe levels; and second, to document the experiences, process and outcomes of the project itself. however, a longer-term aim of the sdit project was to fi nd ways of helping teachers and learners work towards improving the data literacy of gce (general certifi cate of education) ‘a’ level and university students to: • enable a better understanding of the use of social science data as applied to real life problems; • enhance skills in manipulating numerical data textbooks, newspapers, reports and databases; • conceptualise the characteristics of quantitative data so that they can be used to support substantive arguments; iassist quarterly summer/fall 2004 41 • increase the ability and confi dence of students in producing and communicating data; • become critical consumers of these data. the key objectives of the project were to: • produce stand-alone and integrated teaching datasets from more complex socio-economic datasets; • develop an integrated web-based learning and teaching resource that links together resource discovery tools, and data exploration and extraction tools (nesstar) with teaching materials that help address substantive issues for social science teaching; • gain evaluation and feedback from piloting this repurposed content in the area of social science teaching in the he and fe sectors; • provide a model for improving the productivity of teachers by reducing the resources (time and burden) required to incorporate data related resources into learning and teaching courses; • improve access to key primary data sources and related resources for the learning and teaching communities; • promote increased and more effective use of a national data services for problem-based learning in the classroom, at all educational levels. and more specifi cally, the project’s deliverables included: • the creation of new and easily accessible dataset-based on the 2000 british crime survey conducted by the home offi ce16; • accompanying engaging and uncomplicated substantive learning materials and a user guide to resource discovery; • an intuitive and fl exible means of delivering these materials, via nesstar17 and via the ukda download service; • tutor’s guides with scenarios of usage; • high quality and standardised descriptions of the resources, linked to other key web resources in this area (such as sosig resources18, the virtual training suite (vts) internet tutorials19, the offi ce of national statistics20); • an evaluation and awareness raising strategy, together with reports from evaluations with project advisors, tutors and students; • a fully documented ‘warts and all’ report on the processes used to repurpose and pilot the learning materials, and ideas on how, seeing a prototype, the model could be applied, or repurposed, in a wider sense to other subject areas based on existing ukda data resources. thus the focus of the content of the sdit resources was to integrate the mechanics of data analysis with theoretical material. the developments and outputs were conceived to be intrinsically linked and aimed to ensure that students gain a better understanding and knowledge of the nature, context, extraction, manipulation, visualisation, statistical analysis and interpretation of key social science data sources. these skills must be grounded in substantive and intellectual reasoning in relation to the taught subject or curriculum. methodology for methodology’s sake is never a useful way to introduce quantitative reasoning to the social scientist. rather, furnishing the beginner with the applied skills to be able to appreciate the potential of survey data to answer research questions is a more appealing approach. repurposing in the case of this project meant re-packaging complex data with educational narratives and exercises into discrete ‘chunks’ which could then be tried and tested by tutors and incorporated into their teaching in a fl exible way. one further key outcome initially envisaged was to ensure that the resources encouraged an easier route to accessing data, metadata, learning materials, and data exploration/visualisation tools, and thereby encourage student-centred learning and the development of appropriate skills for undertaking project-based work. as such, the resources might be as applicable to distance as to class-based learning. the small scale project ran over eighteen months with a team of four part-time staff. 42 iassist quarterly summer/fall 2004 the teaching and learning resources created x4l sdit uses the study of crime in society to show how existing data sources can be utilised to answer questions about crime. crime is a popular topic taught across the curriculum, and is relevant to a range of social science disciplines, such as sociology, politics, management and general studies, psychology, media and citizenship studies, as well as to public services diplomas. the topic investigating crime was designed to dovetail with a variety of fe and he social science syllabi in which social research methods are often taught. thus the resources developed are appropriate for the uk advanced or ‘a’ level syllabi (taken at age 18) but are also highly applicable for undergraduate and postgraduate learning. the outputs created were a variety of free teaching and learning resources relating to social science and statistics, based on learning strategies that encourage the teaching of research methods within a substantive context. the following resources were developed: four learning modules on the use of crime data; two appendices on sampling and statistical inference; a glossary of statistical terms, and two resource discovery guides, one on the use of the nesstar online data exploration system, and another on how to fi nd data and documentation (resource discovery) in the ukda. the latter two general user guides to exploring and accessing the data collections housed at the ukda have wider more general appeal and are less syllabus oriented. the modules covered the following areas: • module 1: tracking crime: police recorded crime figures, trends and reasons for change; • module 2: theories about crime: public perceptions of crime rates; • module 2 appendix: crime and political parties, aimed at politics students; • module 3: gathering evidence: how to investigate crime statistics; • module 4: examining evidence: how to interrogate crime statistics; • module 4 appendix: reliability of results; • module 5: resource discovery searching for evidence: sources of crime data; • module 6: guide to using nesstar. iassist quarterly summer/fall 2004 43 starting from simple graph reading skills, the most "complicated" data analysis task reached in these modules is actually only a three-way cross-tabulation and resulting measures of association. the concepts of reliability, sampling, and statistical confi dence discussed along the way are perhaps the most complex. appendix a sets out the data literacy and data analysis concepts covered and skills to be learned in the modules. figure 1 sets out the skills to be learned as one moves sequentially through the six modules. figure 1. flow diagram of stepwise learning skills used in sdit project modules were designed to be used as part of standard classroom teaching or as additional/self-paced learning activities and were created in a number of formats to suit different pedagogical needs: • online, interactive self-paced modules hosted (long term) at the ukda web site • printable and reproducible hard copies: bound paper workbook with accompanying cd-rom microsoft word fi les adobe pdf fi les microsoft powerpoint presentations which can be used to provide slides or handouts the project also created freely available new teaching datasets and access to data exploration software, for which a guide to accessing them for the data handling exercises was also provided: 44 iassist quarterly summer/fall 2004 • a restriction-free teaching version of the british crime survey dataset available in multiple formats (spss, stata, nsdstat and tab delimited (suitable for ms excel) and available from three systems21: via the freely available online browsing system, nesstar: the data can be explored directly online using very simple point-and-click procedures or the dataset can be downloaded from the site and imported into other software, such as excel or spss; via the ukda online web download/ ordering system; via the nsdstat data analysis software available from the x4l project website or sdit cd-rom22 via the x4l project website or sdit cd-rom. for promotional purposes, the resources listed above comprising the modules, guides, data and data analysis software, were printed and bound as an sdit resources pack, with an accompanying cd-rom. a tutor guide to accompany the resources was also been prepared which gives model answers and suggestions for classroom exercises, ways of using data resources in teaching, providing an exemplar/model of how such resources could be applied to other topics e.g. health, race etc, model answers to the quizzes, spss syntax for data analysis exercises, and a mapping of the resources to key skills levels 3 and 4, that is, appropriate to ‘a’ level23. finally, case studies of using the sdit resources, based on feedback and usage of the resources in the classroom by tutors and students were included. experiences of repurposing data resources for learning data literacy from the start, the project considered it critical that the pedagogical concerns drive the content of the resources which should then be ‘translated’ to the teams building and implementing the resources. the modules for he and fe were authored and piloted primarily by lecturers who are responsible for teaching quantitative skills in social science (in political science iassist quarterly summer/fall 2004 45 and sociology). the fe teacher was formally bought out from his teaching for 25 days over the life of the project and the he tutor, who helped construct the initial grant application, was donating his time voluntarily. the extraction of teaching datasets, writing nesstar user guides, designing and building the web interface (and cd-rom) for the resources, undertaking evaluation activities, and creating promotional materials, was delegated to staff at the ukda. a close working relationship was built between the ukda project staff and the two tutors, with regular brainstorming sessions and progress meetings. the experience of working in partnership with tutors also helped at the evaluation stage when rich feedback could be obtained from their trialing the resources in their own teaching. moreover, the pedagogical aspects of the resource creation and implementation could be documented (by them) in the tutor guide and in fi nal reports. lessons learned the fi rst and most important lesson learned from the project was that drafting resources that aim to be relevant, appealing, fl exible, light-weight, ‘discrete’ rather than courseware, and suitable for web-based delivery, takes up exponential resources: time for coordinating, authoring, evaluating, rewriting, mounting on web and promoting. the process of deciding upon a suitable topic and authoring the material to be generic enough to answer key learning concepts, yet specifi c enough to meet the needs of syllabi, actually proved to the most complex of the tasks. thus, while the aim of promoting ‘customizable’ repurposing is valuable, the production of new content was far from straightforward, requiring considerable subject knowledge and creative skill, as well as technical ability, to supply freely downloadable software and comply with current web standards. the second lesson learned was that while we had initially envisaged that reaching out to fe teachers across the uk would be a matter of procedure, it was not. teachers are hidden away, embedded within the burrows of their own colleges, and rarely have time to venture out into the world outside the conventional classroom. there are some six hundred fe colleges within the uk and there are no e-mailing lists to which they systematically subscribe, even within their own disciplines. finding ready and willing evaluators for the resources was a major challenge! lack of foresight in these two matters (drafting of the resources and evaluation), meant that an extension of 2.5 months to the original 16 months project duration, was requested (and granted) from the funders. we would certainly recommend that other repurposing projects allow more time to author materials and allow for iterative editing between rounds of testing. it is also impossible to rely simply on the goodwill of tutors, and involvement of tutors needs to take into consideration an extended period of buy-out of their time. that said, even then the repurposing will probably require a single core project staff member to pull together the content / testing / mounting for delivery / user guide preparation / promotion and user support. in the case of the sdit project, mid-term, we had to buy in a new project offi cer, with prev. the original evaluation plan was conceived as a small and informal programme of work, focusing on feedback from the two teachers involved in the project testing the resource themselves, and on a small group of their students. the local fe tutor with whom we worked was unable to test out the web-based resource according to the planned timetable. the 16-19 syllabus timetables are extremely rigid, and to conduct real-time class-based evaluation that coincides with appropriate points in the curriculum is no simple task. in retrospect, we would have benefi ted from allowing more lead time for securing evaluators and testers to cope with the infl exibility. earlier reaching out to fe teachers and information and learning technology staff and detection of ‘champions’ would have been better done earlier, but the whole 16-19 learning territory area was unfamiliar to the ukda. nonetheless, the local fe tutor’s experiences provided us with a useful comparison of three ways of teaching a single topic: with and without the new learning pathways and resources that had been constructed in this project (written up as case study in the tutor guide). finally, we did not anticipate the need for extended checking/proofi ng of all materials, after each editing iteration, and the time taken to undertake promotional and outreach work, which mushroomed as news of the project spread. feedback from teachers, students and library support staff the major pedagogic challanges for e-learning uncovered in this project concern the building of suitable content, interest value, comprehensiveness, complexity, sequencing, relevance and fi t to syllabus, key skills, and positioning in the learning process. the he and fe staff and students, and information professionals who helped advise on and evaluate the project, were asked to consider these issues during the phases of evaluation. on the whole, evaluators considered the learning materials and associated guides created by the project to be impressive and highly useful. moreover, they felt that the project had delivered a neat and fl exible model that both answered the data literacy challenge in hand and demonstrated the concept of repurposing. 46 iassist quarterly summer/fall 2004 pre-prepared materials can save teachers considerable time and effort, and also offer ideas of how to utilize data sources in their own teaching. while the sdit resources span both 16-19 and undergraduate/postgraduate levels, the levels of learning are demarcated: the more advanced fe students can investigate the latter modules, while many undergraduates (or postgraduates) may benefi t from seeing these as a basic revision session. however, in the evaluation phase of this project, some fe level students did undertake all six modules without a problem, and likewise, some postgraduates who took on the whole six modules found the experience highly complementary to their existing knowledge base. case studies from the teacher and student evaluations have been written up and are published in the sdit tutor guide25. these were found to be highly instructive, providing exemplars of how to the resources can be used on the shop fl oor and may entice the more curious or more cautious teachers to think positively about incorporating such approaches. a fi nal round of intensive publicity has resulted in recent feedback suggesting that the resources are indeed offering teachers greater awareness of the potential and relative ease of utilising raw data held by ukda, and that they wanted the printed and cd-rom sdit resources off-the-shelf (not necessarily the web versions!) . we hope they might be inspired to use the repurposing methodology to create their own similar materials in the area of data literacy instruction, although it is questionable as to whether they would have adequate support to ‘indulge’ in the areas of technical competence and web standards-compliance, as we did for the sdit web resources. a new project recently started at essex is taking the sdit model and rewriting the substantive content to cover investigating health issues. while this is being done by lecturers and an ex-data archivist under a small-scale dedicated teaching and learning grant, there is continued liaison with the sdit team at the ukda. thus, adaptation of e-learning resource by teachers is seen to be most fruitfully supported by the holders of basic resources, for example, data service providers such as esds. high quality learning materials may be most usefully developed by teachers working alongside esds and e-learning implementers to build subject-specifi c, or concept-specifi c sustainable dataoriented resources. we believe that those that can map the resources to a common syllabus will be requested most. the mapping of modules and resources to syllabi and key skills (levels 3 and 4) in the sdit resources was attractive to fe tutors. in this way it would be easier to quickly identify appropriate materials from learning banks or repositories for use in their own teaching. finally, technical issues to do with hosting and sharing web resources also arose as important matters for he and fe institutions. the incompatibility of different virtual leaning environments (vles), such as blackboard can cause huge problems for importing web resources that have been described and sequenced using a content package.26 this suggests that resources should be kept simple, probably avoiding the use of complicated style sheets. sustainability of e-learning resources one of the features of the x4l programme under which this sdit project was supported was the building of models to support longer-term e-learning resources. metadata about the sdit web and non-web based resources (using the uk common metadata framework for learning objects) was submitted to the test-bed learning and teaching repository, now set up as a service known as the jorum online repository for learning and teaching materials 27. although an explicit exit strategy was not part of the project plan, nor a requirement of the x4l programme, it was always considered by the sdit team that it would be desirable to have an appropriate continuation strategy that ensured the ongoing maintenance (hosting and functionality) and promotion of the resources, and fi ndings pertinent to the remit of the ukda’s longer term strategic objectives. it is harder to consider the resources that might be required to continually update the content of data to refl ect currency of data and trends, but it is clear that in ten years time statistical results based on 2000 crime data will appear rather out-of-date and, quite probably, unappealing to the student. indeed, the issue of keeping content current is a challenge that many e-learning resources, and learning repositories, need to consider. the sdit project produced exemplar, rather than one-off static resources and a methodology/template for repurposing social science data to help teach key concepts in introductory data literacy and statistics. our own view is that it is probably best for teachers to engage in the revision/updating process so that the learning materials meet the specifi c needs of their own topics taught. where possible, data services like esds can help contribute by providing the appropriate data, but could certainly not engage in the extent of involvement that was required to create these sdit resources. not engage in the extent of involvement that was required to create these sdit resources. not equally, web-based resources must keep pace with the ever-changing web standards. and, on metadata matters, the work iassist quarterly summer/fall 2004 47 involved in having to remap resources to new evolving metadata scheme or new thesauri, could also pose serious barriers to maintenance. it is evident, from this small scale project, a toe-dipping exercise, that embedding strategies are required to get e-learning resources for social science used on a widespread, taken-for-granted basis. there is a need to reach out to teachers on the ground, via discipline-specifi c workshops and road-show events, in addition to information and learning technology staff, within educational institutions. new different teaching and learning styles are often slow to fi lter into mainstream educational practice and can often be better implemented from a top-down directive rather than a more passive bottom up approach via a handful of pioneering tutors. this work requires promotion and encouragement from organisations and policy makers in fi elds of education and ilt, and also national examining boards, to help utilise jisc resources and programme outputs like the x4l, and to help join up the many disparate e-learning initiatives. joined-up initiatives? sadly, many of the recent e-learning initiatives, at least in the uk, that have arisen over the past fi ve years appear far from being joined-up and very few have given rise to sustainable subject-oriented products. different funding organisations have been competing to pioneer e-learning across the learning spectrum. this parallel universe results in replication and noncollaboration, and often the untimely death of the expensive e-resources that have been created, simply because sustainability has not been built in to the programme plan. even in 2004, it was evident from undertaking this project that competitive streams still exist. we found out about two large-scale projects working broadly in the area of statistical literacy who knew nothing about the jisc activities until we contacted them. and, the attempts of our project tried to engender mutual recognition between two such programmes was met with disinterest. a second example is with repositories of learning materials that are cropping up. many are replicating or ignoring work already done by data archives, for example, in the case of establishing vocabularies for social science data resources, designing teaching materials that support particular social science data, and attempting to cover user support for the materials. quality assurance procedures for the submission of specifi c materials (sometimes covering analytical or technical matters) into these repositories is also worryingly absent: in other words the potential for garbage in–garbage out. mutual experiences can be shared by joint events and thematic publications. fortunately, the association for learning technology (alt) does have a lively cross-disciplinary community that has well-attended lively events and a regular news outlets28. future support for data literacy e-resources it is critical that funding bodies, like the jisc, help create further opportunities through networking and funding to create further novel and engaging teaching materials in the area of data literacy that is currently still lacking in a choice of eresources. such future e-learning programmes should be oriented towards methodologies and strategies for embedding learning objects into teaching and learning practise. this sdit project’s start coincided with the initiation of the esds, for which proactive and dedicated support for users of four particular types of data was strategically built in to its remit. the data types are: complex household surveys; longitudinal panel and cohort studies; multi-country macro databanks; and qualitative data. this is over and above the core esds services of data archiving, dissemination, and preliminary user support. through its four help desks, esds is able to offer reactive tailored and sometimes, analytic support to users, and run a full programme of outreach events that deal with: resource discovery; data confrontation; getting started with data analysis software packages; conceptually more diffi cult issues like weighting and linking data; and best practice in data creation and management29. all of these areas touch on data literacy. at the time of writing this paper, a second application to the jisc was submitted under phase two of the exchange for learning programme, looking at creating sustainable data literacy video tutorials with two other institutions. the follow-on programme provides an opportunity to build on what x4l has done on repurposing and extend this to re-use and to explore issues associated with buy-in and embedding at an individual and institutional level. while the grant was turned down, the sdit project staff continues to seek to foster new collaboration with educational communities and pursue funding opportunities for e-learning initiatives based on ukda resources, where this is considered to be benefi cial to the remit of the ukda. equally, it might be worth pursuing a business plan in order to further continued 48 iassist quarterly summer/fall 2004 productive collaboration of ukda with teachers. repurposing requires devoted and skilled staff with complimentary talents to author the materials, pilot and convert resources to web-based media. writing honed to appropriate pedagogic levels, and adequate technical skills are instrumental to the success of the content and learning objects created. this model would thus need to support dedicated core staff to make links with teachers expressing an interest in partnership. at a minimum, personnel devoted staffi ng would require: • one senior offi cer trained in social science methodology and data analysis to provide: o reactive and proactive to support to users; o help prepare data; o collaboratively draft associated resources; o run promotional road shows; • and one part-time web and metadata offi cer to: o create web-based standards-compliant materials hosted at persistent websites like ukda; o prepare metadata for resources to go into the national learning repositories. a budget should also not skimp on costs to cover promotional outputs and activities. indeed, data service providers would be better suited to having this type of staff available in ‘permanent’ positions in-house, as relying on short-fi re funding means that expertise is quickly lost as fi xed-term contract project staff move on (the skilled senior ukda staff working on the sdit project have left due to contracts expiring). conclusion from the ukda’s point of view, this project has enabled us to promote ukda’s existing portfolio of data and online data access tools, such as nesstar, to the teaching and learning communities, which will, hopefully foster new interests in utilizing survey data. if there is one key area from which the ukda benefi ted, it was the chance to engage with the world of further education, to help appeciate better its infrastructure, it and pedagogical concerns surrounding core social science teaching and data literacy. exploring ideas and gaining feedback about repurposing data collections for the fe community should have a pay-off for economic and social data service and for its funders if data resources that are ‘enhanced’ for educational use become popular, or even mainstream within this sector. any further uptake of social science data and associated tools, by the teaching and learning communities would be a signifi cant achievement. appendix a: overview of learning modules iassist quarterly summer/fall 2004 49 module 1 skills covered four social reseach methods./introductory statistics modules: module 1. tracking crime: police recorded crime figures, trends and reasons for c hange this module looks at the trend in recorded crime. it charts the trend in crime for each of the last three political administrations, and concludes with an exercise linking policy decisions with possible explanations for changes in crime levels. • find out what has been happening to crime rates • find out how crime is measured • examine the effectiveness of different governments on crime • try to fi gure out what you would do ü line graph reading skills ü interpretation of trends ü internet usage ü group discussion skills ü problem analysis and evaluation skills module 2. theories about crime: public perceptions of crime rates the module considers an alternative method of measuring crime to the previous module, looking at the british crime survey, and compares the two measures of crime levels. it then shifts emphasis to look at perceptions of crime trends, and examines different theories as to why the public perception of crime levels may not match the actual risk of victimisation. • there are different ways to record crime • the offi cial report says that although crime is really falling, the public think it is increasing • actually when we look at time graphs the position is complex • a usual explanation is that the media create unnecessary worry • there are other factors involved though, such as social class there is also an appendix to module 2 for government students which looks at uk party policy on crime. ü comprehension of basic measurement guidelines ü trend comparison ü more complex graphical analysis (stacked bar charts, time indices, paired bar charts) ü understanding of theoretical concepts and evaluation of evidence ü understanding of simple statistical concepts 50 iassist quarterly summer/fall 2004 module 3. gathering evidence: how to investigate crime statistics this module is concerned with the concepts of operationalisation and validity, and with basic descriptive statistics. it shows how to use the nesstar site to fi nd out information about the british crime survey, to constructively criticise the validity of data used in reports, and to use the simple computer program to generate descriptive statistics, frequency tables and graphs. • learn about devising measures for concepts • learn how to describe large sets of numbers using one or two numbers • learn how to make and interpret straightforward tables and graphs • learn how to cite sources properly • learn how to use an undemanding data analysis program ü understanding of concepts of operationalisation and validity ü understanding of content and usage of metadata ü use of internet to explore metadata ü understanding of basic descriptive statistics and frequency tables ü use of computer program to generate descriptive statistics, graphs and univariate tables module 4. examining evidence: how to interrogate crime statistics this is a skills-based module concerned with explaining the analysis of associations between two variables. • learn how to alter data • learn how to examine associations between two variables • learn how to present and analyse data in tables there is a separate appendix to module four which looks at statistical signifi cance, and shows how to use nsdstat to investigate this. ü understanding of concepts of association and independence ü use of computer program for recoding data ü use of computer program for construction of 2-way table ü analysis of 2-way tables glossary also there is a short glossary of statistical terms, which are referred to in the four modules. the electronic versions of the materials (e.g. in word or the web) can link directly to this glossary. in addition, there are two general guides to fi nding and investigating data and documentation on the uk data archive site: iassist quarterly summer/fall 2004 51 module 5. searching for evidence: sources of crime data this module shows how to search the uk data archive website to fi nd out which studies have been conducted on any given topic • get a tour of some of the resources available on the web, including • uk data archive (ukda) • social science information gateway (sosig) • learn how to search for other data • get a list of links to other useful sites ü resource discovery on the web ü finding surveys at the uk data archive (ukda) ü exploring the social science information gateway (sosig) module 6. browsing and analysing evidence: a guide to using nesstar this guide shows how to use the on-line interactive nesstar website to obtain information about studies, such as data collection details, related publications and even the questionnaire itself. the guide also shows how to use the site to establish which variables are in a dataset, and to produce tables and graphs from the data. • get an introduction to the online data system nesstar • find out how to access and browse data using nesstar • become familiar with the british crime survey dataset • learn how to produce tables and graphs online ü accessing and browsing data using online resources ü familiarity with the british crime survey dataset ü producing tables and graphs 52 iassist quarterly summer/fall 2004 appendix b: acronym check alt association for learning technology chcc collection of historical and contemporary census data and related materials dfes department for education and skills esds economic and social data service esrc economic and social research council fe further education gce general certifi cate of education he higher education jisc joint information systems committee nna national numeracy network maa mathematical association of america sdit survey data in teaching project ukda uk data archive vts virtual training suite x4l exchange for learning programme notes 1 contact: louise corti, uk data archive, colchester, essex co4 3sq, uk. phone: +44 1206 872145. survey data in teaching project (sdit) internet http://x4l.data-archive.ac.uk/. email: corti@essex.ac.uk 2 economic and social data service (esds) web site: http://www.esds.ac.uk 3uk data archive (ukda) website: www.data-archive.ac.uk 4 jisc exchange for learning programme (x4l) available at: http://www.jisc.ac.uk/index.cfm?name=programme_x4l 5 part of this paper was presented at the iassist conference held in ottaowa in may 2003 in the session on “advancing research and data literacy: empowering users”. 6 rice, robin et al. (2001), report on use of data in teaching and learning, edinburgh university. available at: http:// datalib.ed.ac.uk/projects/datateach.html 7 economic and social research council (esrc) (2001), postgraduate training guidelines. available at: http://www.esrc. ac.uk/esrccontent/postgradfunding/postgraduate_training_guidelines_2001.asp 8 smith, a. (2004), ‘making mathematics count’. (smith 2004). available at: http://www.mathsinquiry.org.uk/ 9 censusatschool (2003), nottingham trent university, created by the royal statistical society (rss) centre for statistical education. available at: http://www.censusatschool.ntu.ac.uk/ 10 tomlinson report (2004), the final report of the working group on 14-19 reform, department for education and skills. available at: http://www.14-19reform.gov.uk/ 11 steen, lynn (2004), quantitative literacy: why numeracy matters for schools and colleges, maaonline: http://www. maa.org/ql/index.html iassist quarterly summer/fall 2004 53 12 information about the carleton college quantitative literacy grant entitled “quantitative inquiry reasoning and knowledge to strengthen the educational foundations of citizenship” can be found at: http://webapps.acs.carleton.edu/news/ ?content=content&module=&id=72301 13 see http://www.statlit.org/ 14 collection of historical and contemporary census data and related materials (chcc) (2003), universities of manchester, leeds, glasgow and essex, sept 2003. available at: www.chcc.ac.uk 15 the esds holds a rrange of key uk survey series deposited by the offi ce of national statistics and other government departments, including the labour force survey, the general household survey, the health survey for england and the expenditure and food survey. see http://www.esds.ac.uk/government/surveys/ 16 home offi ce. research, development and statistics directorate and national centre for social research, british crime survey, 2000 [computer fi le]. 2nd edition. colchester, essex: uk data archive [distributor], december 2003. sn: 4463. 17 nesstar: the ukda’s online survey data browsing and download system. uk data archive catalogue available at: http://nesstar.esds.ac.uk/webview/index.jsp 18 social science information gateway (sosig) aims to provide a trusted source of selected, high quality internet information for researchers and practitioners in the social sciences, business and law. it is part of the uk resource discovery network (rdn). available at: http://www.sosig.ac.uk 19the rdn virtual training suite (vts) is a set of free online tutorials designed to help students, lecturers and researchers improve their internet information literacy and it skills. available at: http://www.vts.rdn.ac.uk/ 20 national statistics online (ns) provides information about britain’s economy, population and society at national and local level, through summaries and detailed data releases. available at: http://www.statistics.gov.uk/ 21 higher education funding councils. joint information systems committee. exchange for learning programme. survey data in teaching project, british crime survey, 2000 : x4l sdit teaching dataset [computer fi le]. home offi ce. research, development and statistics directorate, national centre for social research, [original data producer(s)]. colchester, essex: uk data archive [distributor], may 2004. sn: 4918. available at: http://www.data-archive.ac.uk/fi ndingdata/sndescription. asp?sn=4918&key=bcs&catg=xmlall 22 nsdstat is a demonstration version of very simple and user-friendly data analysis software, which is utilised in the last two of the teaching modules were created. this is a very easy windows-based program that enables students to examine the british crime survey data for themselves, and produce their own tables and graphs. this program was developed for use in schools and colleges by the norwegian social science data archive (nsd), and is the analytical engine behind the nesstar website. a program automatically installs the sdit teaching version of the data. 23 the key skills document maps the six modules to the six uk qca key skills levels 3 and 4 requirements for 16-19 years olds in education. the mapping of modules and resources to syllabi and key skills (levels 3 and 4) requirements for 16-18 year olds undertaken in the project was attractive to fe tutors (see tutor guide for mapping), suggesting that all elearning resources should be mapped to the syllabus as precisely as possible. the investigation of crime from a sociological perspective, and the application of statistics and data handling covered in the resources address: application of number, information technology and problem solving, while the group exercises cover communication. the web resources, which can be used as self-paced exercises, may further contribute to improving own learning and performance. the mapping can be found at: http://x4l.data-archive.ac.uk/learning/keyskills.asp. see also http://www.qca.org.uk/14-19 for further information about the development of the national key skills. 24 dr. jon mulberg was the main project offi cer for the last 12 months of the project. he is the author of figuring figures: an introduction to data analysis (2001), harlow: prentice hall. 25 sdit case studies available at: http://x4l.data-archive.ac.uk/learning/tutorsguide.pdf 54 iassist quarterly summer/fall 2004 26 more information about vles can be found at: http://www.jisc.ac.uk/index.cfm?name=mle_related_vle. for content packaging (e.g. ims) see: http://www.cetis.ac.uk/members/x4l/articles/relevantspecs. 27 the ukcmf (uk common metadata framework) is a subset of the learning object metadata (lom) produced by ieee. see http://www.cetis.ac.uk/members/x4l/faq/metadata/20030213162013. the jorum online repository for learning and teaching materials is available at: http://www.jorum.ac.uk. 28 the association for learning technology (alt) is a professional and scholarly association which seeks to bring together all those with an interest in the use of learning technology. it has over 200 organizations and over 500 individuals in membership. available at: http://www.alt.ac.uk/ 29 esds programme of events available at: http://www.esds.ac.uk/news/esdsforthevents.asp controlled access as a means of balancing respondent privacy and analytical utility in survey research by phi 1 1 ip a. windel 1 bonneville power administration u.s. department of energy this paper discusses issues related to the privacy and confidentiality of survey data and describes procedures for protecting the privacy and confidentiality of survey respondents. of primary concern are the impacts of these procedures on the research value of the data. a method for striking a balance between these two competing interests is presented together with a case study in its implementation at the bonneville power administration. the challenge to individual privacy sensitivity to issues involving individual privacy lies at the root of our nation's political and legal systems. the advent of very large, extremely fast electronic data processing machines presents a challenge of unprecedented magnitude, because these machines have made it possible for governments to assemble and readily access vast quantities of information concerning individuals. lest the likelihood of such occurances be too readily dismissed, it should be recalled that during world war ii, the department of war and the department of state inquired about access to individually identifiable records for the u.s. bureau of the census in an effort to identify americans of japanese ancestory living on the west coast. (1) more recently, in the state of new jersey, law enforcement officials requested individually identifiable information concerning participants in the new jersey negative income tax experiment . (2) these are but two of many incidents that could be cited as evidence of the need for limitations on and close scrutiny of governmental uses of information pertaining to individuals. the u.s. government is not the only beneficiary and potential abuser of data pertaining to individuals. advertising and door-to-door sales companies, debt collection agencies, and even electric utility companies regularly use government collected data which derive from individuals, could benefit greatly from access to government data which permit identification of individual sources, and therefore represent a potential threat to individual privacy. the federal government has sought to limit the increasing threat to individual privacy through appropriate legislation, including the privacy act of 197'* (pub. l. 93-579)in response to this and other legislation, federal agencies have regularly attempted to avoid invasions of individual privacy from data collections through data reporting techniques that make it impossible to identify individual respondents. for example, the u.s. bureau of the census does not report data for individual respondents. furthermore, where the characteristics reported for geographic clusters would enable identification of individuals, the data for that cluster are suppressed. and, indeed, the u.s. bureau of the census has been exemplary in maintaining the highest ethical standards in the conduct of surveys and in protecting the privacy of survey respondents. the energy information administration, u.s. department of energy (eia/doe) relies on a different technique to protect the privacy of respondents in their energy consumption sample surveys. among other objectives, these annual surveys are designed to produce a data base for use by analysts in accounting for variations and changes in energy consumption among individual residential units. (2) as a result, data for individual respondents (residential units) are available on computer tapes. in an effort to protect the privacy of the respondents, the eia/doe suppresses certain information (for example, geographic location other than census region and climate zone). in addition to the interview responses, the data from the eia/doe surveys also include actual billing histories for the primary fuels used by the unit (electricity, natural gas, and fuel oil), including delivery or billing period dates and amounts of fuels consumed. since the fuel suppliers could use this information to identify their own customers, the eia/doe "masks" this data by systematically altering the dates and fuel amounts. thus, the billing dates are randomly altered by a factor of ± 3 days and the fuel amounts are randomly altered by a factor of ± ^ percent. the impacts of privacy protection on research utility when properly implemented, both data suppression and data masking techniques are effective methods of protecting the privacy of individuals. however, the application of these techniques to a data set can also have significant impacts on the utility of the data for analytical purposes. for example, decennial census data obviously cannot be used to analyze the relationships between different characteristics of individuals or households. the smallest unit of analysis is the block (in smsa areas), and even here, the amount of data reported is limited and suppression often has significant impacts on the results. in the case of the ela/doe's energy consumption surveys, the suppression of geographic identifiers prohibits state level analyses, among other desirable topics of investigation. in addition, however, the masking techniques used by the eia/doe (i.e., alteration of the billing history information) may have even more serious effects on the analytical utility of the data. it is likely that these techniques have little or no effect on the estimates of total values. however, the effects on subsample totals and on the results of analytical explanations (for example, multiple regression) are unknown and this author is not aware of any published attempts to estimate the possible effects. as a result, the analytical utility of the data may be severely 1 imi ted. the point of this discussion is that there is a strong inverse correlation between the traditional procedures for protecting individual privacy and the utility of the data for analytical purposes. as is too frequently the case, in order to protect ourselves against unethical usages of data, we have restricted and in some cases prohibited legitimate and profitable usages of the data. the task, and the point of this paper, is to design procedures which strike a balance between these two competing interests. the privacy act of 197^ the privacy act of is^^ (pub. l. 93-579) is most frequently cited as justification for the suppression and masking of data. sponsoring governmental agencies either simply assume that the act prohibits the publication of data in which individual respondents might be identified or, if they understand the act, fear that the procedures required by the act in order to provide access to identifiable records will adversely impact the resulting data. the first instance is a clear and simple misinterpretation of the act, for it does not prohibit the publication of individually identifiable data. (^) what the act requires is "informed consent." respondents must be informed of the authority for and the purposes of the collection, what uses will be made of the data and by whom, and the effects on the respondent, if any, for not participating, prior to giving their voluntary consent to participate in the data col lect ion . (5) since the act has been in place for nearly ten years, it is likely that few experienced sponsoring agencies continue to suffer under this misinterpretation of the act. a more frequent reason for engaging in data suppression and masking is likely the sponsoring agency's hesitancy to inform respondents that certain users may be able to identify them. the agency fears, first, that response rates will be adversely impacted; second, that in an effort to avoid refusals, interviewers may avoid clearly informing the respondents; and third, to avoid the second problem, the sponsoring agency must require the respondent to read and sign a consent form, which, in turn, will have further adverse effects on the response rates. in certain instances, the sponsoring agency's interest in complete confidentiality for its respondents is entirely justified. for example, in surveys of individuals who engage in illegal activities, or of individuals who have been the victims of personal crimes such as rape or family abuse, or of individuals whose behavior might be considered unethical or reprehensible by others (for example, extramarital sexual relationships), complete respondent confidentiality is required in order to obtain accurate information, to achieve respectable response rates, and, in some cases, to protect the life of the respondent. in data collections involving less sensitive topics, however, such guarantees of complete confidentiality are neither necessary (in terms of insuring high response rates and a high degree of response validity), nor advantageous (in terms of producing data of maximum analytical utility). concerning the effect on response rates, a major study by the census bureau and the national academy of sciences in 1976 included an experimental design in which respondents were assigned to one of five treatments in which the nature of the statement concerning confidentiality was systematically varied. (6) although the study concluded that there was a statistically significant difference in the refusal rates between respondents presented with offers of complete confidentiality, on the one hand, and those presented with statements that answers might be publicly available, the difference was only i percentage point (1.8% versus 2.8%). (7) in addition, a higher percentage of refusals occured prior to the reading of the confidentiality promise (2.9%)as turner concludes. obviously, there are other factors besides confidentiality conditions that made someone refuse a census bureau survey, and our evidence does not support the notion that confidentiality concern is the principal mot i vator . (8) as we emphasized previously, the effect of confidentiality conditions on response rates is likely to vary depending on the sensitivity of the survey subject matter. unfortunately, there are no well-defined studies which precisely document these effects. the census bureau commissioned another experiment in an effort to determine whether varying levels of confidentiality have any impacts on the validity of responses. (9) conducted by response analysis corporation in november, 1976, the study involved a matched sample of 500 households in taylor, michigan. the results of this small experiment suggest that the varying conditions of confidentiality have no significant impacts on the validity of responses to nonsensitive questions, including income. (10) in general, then, there does not seem to be any advantage in offering guarantees of complete confidentiality to respondents in nonsensitive surveys. the validity of the data is no greater than it would be if the confidentiality guarantees were less stringent. however, the application of procedures to insure complete confidentiality can severely diminish the analytical value of the data. thus, in many circumstances, unconditional guarantees of complete confidentiality are distinctly disadvantageous. controlled access in the two cases discussed previously, the procedures for protecting individual privacy were implemented unconditionally. that is, everyone except the sponsoring agency was provided with the same suppressed or masked copy of the data. as a result, even users who can guarantee restricted access and who use the data for statistical purposes only, are prohibited access beyond the single published level. the result can only be a restriction of unknown extent on our ability to understand social processes. this seems not only wasteful, but tragic, especially in view of the crises which all societies currently are facing. in place of the procedures described thus far, we are proposing here a technique which we shall call "controlled access." strictly speaking, this technique does not "replace" suppression and masking. rather, it differentiates between users on the basis of some well-defined criteria, and grants them varying levels of access to the data. thus, certain users may be provided unrestricted access to the data, others may be granted partial access, while still others are permitted access only to completely confidential versions of the data. clearly, the user screening criteria, the procedures for applying the criteria, and the procedures for enforcing the contingent user restrictions are key elements in a controlled access environment. among the user screening criteria, it is likely that the sponsoring agency will want to include the nature of the user's analytical objectives, the ability of the user to control access to the data, the user's potential for invading the privacy of respondents, and the potential harm that would result to the respondent from such invasions. for example, a sponsoring agency might restrict full access to users interested only in statisical analyses, who present little potential threat to the respondents, and who agree in writing not to contact any of the respondents. further, the sponsoring agency might provide such access only on the agency's own premises. amoung the procedures required by the privacy act prior to the establishment of a new system of records is the the designation of a records system manager. depending on the anticipated demand for the data, the frequency with which the sponsoring agency collects data, the sensitivity of the data, and the potential harmful effects resulting from abuse, the agency may wish to leave all judgments in the hands of the system manager or, on the other hand, may wish to establish an elaborate mechanism of review committees and appeal processes. with regard to enforcement procedures, users who are provided access to other than fully confidential versions of the data should be required to sign contractual agreements. the agreements should clearly specify the data to be provided to the users as well as the applicable restrictions on the distribution and permissible uses of the data. an example the bonneville power administration (bpa) is a power marketing agency within the u.s. department of energy. to assist in resource acquisition and transmission contruction planning and decision-making, bpa develops forecasts of energy demand. the forecasts are produced by relatively sophi s i tcated , data intensive computer simulation models. to support these models, bpa conducted an "energy consumption" survey in 1979personal interviews were conducted with approximately 4,000 residents of the pacific northwest region (washington, oregon, idaho, and montana), and fuel billing histories were obtained for those respondents who signed waiver forms. in designing the survey, no provisions were made either for maintaining the list of respondent names and addresses or for providing the raw data to the participating electric utilities and natural gas companies. as a result, the electric utilities and natural gas companies were unable to obtain copies of the raw data which were of analytical utility to themselves. since the utilities had voluntarily invested some of their own resources in the survey (the utilities selected samples of their own customers and provided the fuel billing histories for their customers), utility analysts and managers were less than pleased with the result. at about the same time, bpa joined the eia/doe in an experimental survey of commercial buildings in the pacific northwest. the original purpose of the survey was to test the feasibility of using utility billing records as a sampling frame as compared with more traditional areal sampling techniques. in return for using three pacific northwest areas as the test sites, bpa contributed sufficient funds to insure the completion of the fieldwork and processing of the data. unfortunately, the terms for delivery of thedatawere not entirely clarified prior to the initiation of the survey so that bpa and the participating electric utilities were seriously disappointed when they discovered that the final data were fully suppressed and masked using the usual eia/doe procedures. as a result of these experiences, when bpa began preparations for a second pacific northwest residential energy survey, several electric utilities clearly and firmly articulated their desires for guaranteed access to data of maximum analytical utility to themselves. indeed, unless bpa would satisfy these desires, several utilities made it clear that they would refuse to participate in the survey. thus, the machine-readable copy of the data to be made available to the participating electric utilities should contain a code identifying the serving utility, together with the respondent zip code, all interview responses and complete, unaltered billing histories for each survey respondent. given this information, an interested electric utility could identify its own customers by matching the billing history information from the survey data with their own master records. with the exception of certain potentially idiosyncratic cases, however, the respondents would not be identifiable to any other agency or organization. thus, from the standpoint of potential invasions of respondent privacy, the user audience can be easily and clearly divided into two groups: the electric utilities and natural gas companies, on the one hand; and all other users, on the other hand. (11) conveniently, several other criteria divide the user audience into the same two groups. for example, apart from bpa and the respondents themselves, only the electric utilities and natural gas companies have invested resources in the survey--the electric utilities assisted in the selection of the customer samples and the electric utilities and natural gas companies will be asked to provide billing histories for their customers. second, the billing data which the electric utilities and natural gas companies collect and maintain for all their customers is itself proprietary. thus, the electric utilities and natural gas companies are experienced at, and have procedures in place for restricting access to certain data sets. third, the electric utilities and natural gas companies have legitimate analytical interests in the data--to support their own planning and decision processes. on the other hand, there is a potential for the electric and natural gas companies to abuse the privacy of the survey respondents based on the data from the survey. for example, the survey inquires about the presence of various conservation measures in the dwelling unit. the utilities could use this information to target conservation promotion campaigns. in an effort to prevent such abuses, bpa has developed an agreement which each requesting utility is required to sign prior to receipt of the data. by signing the agreement, the utility agrees to restrict access to the data to its own employees whose official duties require access; to refrain from contacting the respondents as a result of their participation in the survey; and to refrain from discriminating against the respondents. the agreement was reviewed by bpa's general counsel and by analysts and attorneys of several local utilities prior to final implementation. all other interested analysts will have access to a version of the data in which elements by which the user could identify individual respondents will be suppressed or otherwise masi<ed. that is, respondent zip codes will be removed and elements such as respondent race, household income, and dwelling unit size will be examined to determine whether they enable the identification of individual respondents. if so, certain categories of these variables will be collapsed or, if necessary, the elements wi 1 1 be removed from the data set. since the electric utilities and natural gas companies are the only users capable of identifying individual respondents through the billing history data, it will not be necessary to alter this data in order to protect the privacy of survey respondents. as required by the privacy act, the respondents will be fully informed of the authority for and objectives of the survey, who will have access to the data, and the purposes for which the data will be used and that each respondent's serving electric utility and, where applicable, serving natural gas company may be able to identify them. this information will be presented verbally and in writing at the outset of the interview. in addition, near the end of the interview, each respondent will be asked to sign a form authorizing the respondent's serving electric utility and, where applicable, natural gas company, to release the respondent's billing history to the fieldwork contractor and, ultimately, to bpa. at this time, the respondents are once again informed, both verbally and in writing, that the information may be provided to their electric utility or natural gas company and that these companies may be able to identify them. thus, the signed authorization form serves the dual purpose of authorizing release of the respondent's billing history information and documenting the respondent's informed consent to participate in the survey. the fieldwork for the second pacific northwest residential energy survey was scheduled to begin may 23, 1983thus, what effects, if any, the proposed confidentiality statements will have on response rates is yet to be determined. needless to say, the bpa staff will monitor the response rates closely, and there are plans to conduct an analysis of the response rates as soon as possible following the completion of the fieldwork. with regard to the bpa-utility agreements, generic copies have been distributed to all the participating electric utilities. to date, ten of the 57 participating utilities have expressed an interest in obtaining the data and a willingness to sign the agreement .( 1 2) several other utilities reviewed a previous draft of the agreement and, after submitting comments, expressed a willingness to sign. summary and conclusions the right to individual privacy is one of the tenets of the american political and legal system. the federal government has sought to protect the individual right to privacy through legislation like the privacy act of is^*. in response to this and other legislation, federal agencies which collect and publish data have sought to avoid privacy invasion through data suppression and masking techniques. the unconditional use of such techniques has the unfortunate consequence of impairing further analysis of the data. this paper has offered an alternative procedure which seeks to balance the interest in protecting the privacy of the individual respondents with the interest in maximizing the analytical utility of the data. the procedure involves the provision of differential access to the data based on some well defined criteria. thus, users with legitimate interests in conducting statistical analyses, who present little threat to the privacy of the respondents, who contractually agree not to violate the privacy of the respondents, and who either can demonstrate an ability to limit access or agree to use the data on the sponsoring agency's premises, may be granted acccess to the complete set of data. users who do not satisfy all of these criteria may be granted access only to partially or fully suppressed and/or masked versions of the data. for purposes of illustration, the controlled access system developed by the bpa for its second pnwres was presented. this survey is just now going into the field, so what effects, if any, the confidentiality statements have on the response rates is not yet determined. to date, none of the utilities expressing interest in obtaining the data have expressed any hesitation to signing the agreement. it will likely be several years before we know whether any of the utilities have violated the agreements, or whether there are any other problems with enforcing the terms of the agreements. notes and references see a.g, turner, "what subjects of research believe about confidentiality," in j.e. sieber (ed.). the ethics of social research: surveys and experiments new york: spr i nger-verl ag , 1982, pg. 152. see d.t. campbell and j.s. cecil, "a proposed system of regulation for the protection of participants in low-risk areas of applied social research," in j.e. sieber (ed.), the ethics of social research: fieldwork, regulation and pub) icat ion . new york: spr i nger-ver lag , 1982, pg. 110. (3) the eia/doe also conducts surveys of commercial energy consumption. for purposes of illustration, we focus here only on the residential surveys. the techniques used to protect respondent confidentiality in the commercial surveys are basically similar to those used in the residential survey data bases . (4) see 5 u.s.c. sec. 552a and american statistical association, "report of ad hoc committee on privacy and confidentiality," the american statistician , vol. 31, no. 2, pgs. 59-78. (5) see 5 u.s.c. sec. 552a, c. (6) reported in a.g. turner, op. cit., pgs. 15'*ff(7) op. cit., pg. 159. (8) ibid. (9) op. cit. , pgs. i60ff . (10) op. cit. , pgs. 160-161. (11) since bpa markets and transmits electric power only, our attention thus far has focused solely on the electric utilities. however, the fuel supplier survey portions of the end-use surveys include natural gas billing history data and the natural gas companies frequently are interested in obtaining copies of the final data. (12) it should be noted that many of the utilities do not have the machinery or personnel to conduct statistical analyses, and are therefore not interested in obtaining a copy of the final data. 14 iassist quarterly winter 2011 iassist quarterly abstract the university of porto is the largest portuguese university with more than 60 research centers that generate a significant part of portuguese scientific production. u.porto is currently concerned with the curation of and the access to the scientific data generated by its researchers. researchers are motivated to keep their data assets alive as integral part of their published results, and the scientific impact derived from open datasets is also becoming apparent. we have followed the recommendations from well-known actions in research dataset auditing to lead a short study on available data at u.porto. the study has involved researchers from a diversity of disciplines, collecting their views on data curation and sample data. as a result we have identified some generic use cases to inform the development of a data repository prototype. our contacts with the researchers have revealed a great diversity of situations, from groups where data curation was already integrated in the research practice to others who were struggling to incorporate it into their workflows. our experiment was focused on data auditing and use case identification, but we also concluded that in many groups there is a strong concern with the premature exposure of the data. the sample datasets provided by the researchers are being transformed into preservationfriendly archives to be part of a data repository. we will extend the repository infrastructure with data search facilities and expect feedback from the researchers to help define the research data management services at u.porto. keywords: : data curation, management of research data, data repositories introduction the university of porto (u.porto )2 is currently concerned with the curation of and the access to scientific data generated by its researchers. a steady growth in research activity in all domains, international cooperation initiatives and access to data that is either generated by local projects or available via joint projects has generated many ad-hoc data archives. research cycles of projects and scholarships are very short-term from the data assets point of view: data generated in one project may, if there is no continuation project, be abandoned and lost in less than five years. the researchers’ perspective on the longevity of such data is, in general, quite optimistic and the lack of national mandates for data curation favors the continuation of this state of affairs. in this work we have followed the recommendations of pioneering actions in scientific dataset auditing to lead a short study on available datasets at u.porto. an analysis of current initiatives in this area has shown that close contact with researchers is essential for getting a clear view on their needs (ribeiro, et al. 2010). our study involved researchers from a diversity of disciplines, collecting their views on data curation and sample data (rocha da silva, ribeiro and correia lopes 2011). as a result, we have identified some generic use cases to inform the development of a data repository prototype. our contact with researchers revealed a great diversity of situations. there are areas where some form of data curation is already embedded in current practice, mainly due to the requirements of publication venues or the need to share data in international initiatives. some researchers are motivated and aware of both the value of their data and the existing threats on it, but are still struggling to incorporate data curation into their workflows. others, faced with the possibility of having their data curated in a repository, were extremely cautious with respect to privacy issues. data curation at u.porto: identifying current practices across disciplinary domains by cristina ribeiro, maria eugénia matos fernandes 1 u.porto iassist quarterly winter 2011 15 iassist quarterly our study focused on data auditing and use case identification, but we also asked researchers for samples of their data. the datasets provided by the researchers are being transformed into preservation-friendly archives to be part of a data repository. we are extending an existing repository infrastructure with data search facilities and expect feedback from the researchers to help define the data services for the u.porto data repository (rocha da silva, ribeiro and correia lopes 2011). in this paper we provide a short overview of research at u.porto, an outline of the goals for the data curation project at u.porto and describe its preliminary results. we conclude with some reflections on the project results and the perspectives for the management of research data at u.porto. u.porto: a research university u.porto is the largest portuguese university. it comprises 14 schools, a business school, 30 libraries, 12 museums and about 70 r&d units, 31 of which have been regularly classified at the top ranks by a panel of international experts as part of the portuguese research units evaluation. its population consists of about 30,000 students, more than 2,366 teachers and researchers (76% phd) and 1,689 technical and administrative staff. u. porto offers a large range of courses covering all levels of higher education and all the major areas of knowledge. there are over 670 training programs, including undergraduate, masters, integrated master, doctoral, continuing education and specialization courses. the number of foreign students under mobility programs represents more than 8% of the total number of students. u.porto aims at becoming a national and international reference by the high level of its students and the production and dissemination of knowledge. it can be said that the target of being among the top 100 highereducation european institutions for its 100th anniversary in 2011 has been reached. the physical dispersion is a characteristic of u. porto, as the buildings of the university—schools, rd&i institutes, student residences, sports and cultural facilities—are located in three separate areas of the city of porto. moreover, there are research institutes and centers spread throughout the city and some of them even beyond its geographical boundaries. the shortcomings of this geographical dispersion have been practically overcome by the sigarra information system (information system for the aggregated management of resources and academic records). sigarra originated in the engineering school in 1996 and its success led to the transversal implementation in the university from 2003 on. currently, all the u.porto schools, as well as the rectorate, the social services and some of the research centers use sigarra. this integrated system was conceived to facilitate the production, flow, storage and access to the information managed by the institution— contents of pedagogical, scientific, technical and administrative nature—and to promote internal cooperation and the cooperation with external academic, scientific and business communities. the sigarra system interacts with other applications and systems within the university, such as the libraries, the e-learning services, the student administration and the financial management systems and also with u.porto institutional repository, built on a dspace platform. figure 1 illustrates the integration between the information system and the open repository. the creation of the u.porto institutional repository complemented the information management strategy by the end of 2007. the interface connecting the information system and the open repository guarantees that the intellectual production of the academic and scientific community is transferred automatically from the publications module of sigarra to the open repository. the authors just have to register and self-deposit the full text of their publications on their institutional pages, defining that they are public. the same interface also assures the connection between other applications used within the university to register and catalogue the library collections—such as aleph—and the open repository, thus enforcing consistency of data across different applications and systems. from the moment it was created, the number of publications of the open repository of u.porto has grown steadily. at the beginning of 2008, the repository had almost 1,800 full-text and open access publications. three years later, the number of records has evolved to more than 18,000. one of the missions of u.porto is the creation of cultural, artistic and scientific knowledge within the academic community, composed by teachers, researchers and students. this concern has increased in the recent past due to a great variety of factors. beyond some aspects already mentioned above—such as the functionalities of the publications module of the information system and the benefits of the interconnection between sigarra and the open repository—, one cannot ignore the emphasis that has been placed on the recommendations made to the authors to make their intellectual outputs available, stressing the fact that these works are created in the context of their teaching and research activities. it is also important to highlight the suggestions made to the authors to have them consider, whenever possible, the “sparc author addendum”, when they sign contracts with publishers, so they maintain the right to self-archive their work in institutional open repositories, as well as the advice given to researchers to use the university-recommended format to register their affiliation. figure 1 the sigarra information system and the institutional repository full text open access sigarra open repository thematic repository scientific data repository institutional repository •ingest •storage •preservation •access •dissemination open repository 16 iassist quarterly winter 2011 iassist quarterly considering the last decade, 1/5 of the portuguese scientific production was generated at u.porto. current figures show that u.porto is responsible for more than 21% of the portuguese scientific articles indexed in the isi web of science. goals of the data curation project u.porto is currently concerned with the curation of and the access to scientific data generated by its researchers. there is a growing awareness of the fragility of personal digital archives and researchers feel that they need to keep their data assets alive as the research workflow becomes more sophisticated. the possibilities of scientific impact derived from open datasets are also becoming evident. as a result of an identification task, we present a preliminary study on datasets that are being used in current research at u.porto. the emphasis has been on diversity, picking examples from life sciences, engineering, social sciences and arts. the identification also provides insight on current models for data curation, both formal and informal, and on the sensitivity of researchers with respect to open access to their data (scientific data curation at u.porto, 2011). the study has been complemented by the development of a data repository prototype. the purpose of the development is twofold: to provide services which address some of the requirements identified with researchers in a working tool and to establish the basis for a second round of interaction with the researchers, this time using the repository platform to illustrate the use cases in data curation and to test them with their end-users. we were quite aware, from the start, of the many challenges of the project, but also of its strengths. the study conducted in the context of the national repository project (ribeiro, et al. 2010) located similar initiatives (rice 2009, martinez-uribe 2009) and existing recommendations that have helped to establish the main lines of the data audit experiment (university of glasgow, dcc 2009). the commitment of the rectorate and digital university services of u.porto to the development of the institutional repository has provided a solid ground for supporting an experimental data repository. on the other hand, and in spite of the absence of mandates for data curation plans in national projects, we were able to find many researchers concerned with the management of their data and committed to sharing them within their research groups and in the context of international projects in which they are involved. data auditing and dataset collection ithe data audit at u.porto has followed the recommendations issued by similar initiatives, namely the methodology proposed in the data asset framework (university of glasgow, dcc 2009). considering that this is the first approach to data curation at the university level, we have decided to give preference to the diversity of domains. the choice of research groups to include in the study has followed a mixed strategy, selecting some groups due to personal contacts by the team members and others resulting from a call issued by the university rectorate and digital university services to the directors of schools and research institutes. the first contact with the researchers led to an appointment of interviews at their laboratories, based on a script that allowed for many open questions. in cases where the researchers were willing to provide sample datasets, a follow-up interview was scheduled to discuss data formats, the definition of data and their terms of use. we adopted the recommendations of the data asset framework (university of glasgow, dcc 2009) to prepare an “interview guide” (u.porto 2011) and a “comprehensive questionnaire” (u.porto 2011) that were used to collect the researchers profiles, some general information on their datasets, preservation actions and expected use cases for the scenario of a university-level data repository. there was no imposition on researchers to provide data, but most (8 out of 13) volunteered to provide sample datasets, knowing that the data would be used to design and prototype the system and that there was no agenda for a repository service, so they could not expect any immediate benefits from the collaboration. table 1 lists the nature of the collected datasets and the access conditions established by the researchers. interviews were a rich source of information for their needs, where we can highlight data preservation and data exchange with research partners, either internally at u.porto or externally in international projects and partnerships. the collected datasets provide a first view on the research data at u.porto, with data obtained from science, engineering and social sciences resulting from either automatic acquisition or direct collection by the researchers and access conditions ranging from open data to data useable in research but whose origin must be kept anonymous due to pending contracts. most of the datasets under consideration were originally created as a result of research projects, but there were also data collected by external institutions with which u.porto holds service contracts and data collected by national institutes, such as the census data created by the national statistics institute. the interviews with the researchers confirmed our initial assumption that the design of a solution for a data repository should be determined by researchers needs, rather than by any abstract data management convenience (borgman 2011). the interviews showed more concern with functionalities such as data browsing and querying than with strict data preservation or management. for the 8 sample datasets provided by the researchers we created basic dublin core descriptions to ease their deposit into the upcoming data repository. future directions for data management at u.porto the data audit at u.porto has exceeded our expectations with respect to the commitment of researchers with data curation. in some areas table 1. domains and access conditions for datatable 1. domains and access conditions for datatable 1. domains and access conditions for data domain dataset access astronomy gravimetry free chemical engineering pollutant analysis contract pending mechanical engineering material fracture embargoed civil engineering high-speed railways embargoed educational science interviews embargoed psychology interaction records embargoed economy population embargoed ecology plant distribution embargoed iassist quarterly winter 2011 17 iassist quarterly with established practices of deposit in international repositories, the data curation problem can be considered solved, but this is not the case in most domains. the resources required for this small-scale experiment are indicative of the effort required for setting up a data curation service at an institution with the size of u.porto. the use cases identified in this study are being used to define the requirements for the u.porto data repository. the data samples are the basis for the design of data models where the tradeoff between generality and usefulness must be considered to make the curation process practicable. an experimental repository is being developed to test the requirements. as soon as we have a platform where some datasets are deposited and can be queried, researchers can explore it, detect the shortcomings of the proposed approach in their own domain and engage in future developments. this work has raised even more issues than initially expected and many questions remain unanswered. we have observed that in several areas researchers are willing to participate in data curation, even in a scenario where they cannot expect any immediate benefits. this proves that we will be able to stimulate their cooperation in the following steps, but there must be some perceived gain for the researchers in order for this commitment to be sustained. a scenario where people are motivated to participate and get no practical results may ultimately compromise this and future initiatives. the technological support for a research data repository is another open issue. the maturity of software for institutional repositories shows that we do not have to start from scratch and that basic functionality can be taken for granted. but, on the other hand, the use cases for research data are much less clear and less uniform than those for an institutional repository. another issue worth reflection and experimentation is the nature of data curation services. there are currently no data curation services in portugal so there is no experience with respect to their integration in a research institution. libraries are experienced with many of the issues in curation, but not equipped with the highly technological expertise it requires. computing centers have complementary expertise, but their mission is centered in very different services. maybe the most critical aspect for the success of a data curation project is compliance with researchers needs. institutional repositories have flourished due to the adoption of repository technology, originally created to satisfy very specific needs, by the more traditional library community. there are currently no well-established generic platforms for research data management but many custom-designed systems already exist. experience and successful developments will show whether generic platforms can cater to the needs of researchers in different domains or if they have to be more specialized by discipline. references “u.porto comprehensive questionnaire.” updata. 2011. http://sciencedata.up.pt/doc/ (accessed november 2011). “interview guide (in portuguese).” updata. 2011. http://sciencedata. up.pt/doc/ (accessed november 2011). scientific data curation at u.porto. edited by joão rocha silva. 2011. http://sciencedata.up.pt/updata/ (accessed november 2011). university of glasgow, dcc. the data asset framework implementation guide. october 2009. http://www.data-audit.eu/ (accessed november 2011). borgman, christine l. “the conundrum of sharing research data.” journal of the american society for information science and technology, 2011: 1-40. hey, tony, stewart tansley, and kristin tolle, . the fourth paradigm: data-intensive scientific discovery. microsoft, 2009. martinez-uribe, luis. using the data audit framework: an oxford case study. university of oxford, http://ie-repository.jisc.ac.uk/300/, 2009. oecd. oecd principles and guidelines for access to research data from public funding. http://www.oecd.org/dataoecd/9/61/38500813.pdf, 2007. ribeiro, cristina, eloy rodrigues, maria eugénia matos fernandes, and ricardo saraiva. repositórios de dados científicos: estado da arte (in portuguese). project report, porto: rcaap, 2010. rice, robin. disc-uk datashare project: final report. project report, edinburgh: university of edinburgh, 2009. rocha da silva, joão, cristina ribeiro, and joão correia lopes. “ updata a data curation experiment at u.porto using dspace.” proceedings of 8th international conference on preservation of digital objects, ipres 2011. ipres, 2011. notes 1. cristina ribeiro, dei-faculdade de engenharia da universidade do porto/inesc tec, rua dr. roberto frias, s/n, porto, portugal, mcr@ fe.up.pt. maria eugénia matos fernandes, reitoria da universidade do porto, universidade digital praça gomes teixeira, porto, portugal, efernand@reit.up.pt. 2.u.porto: homepage. http://www.up.pt/ iassist quarterly fall & winter 2007 by libby bishop* moving data into and out of an institutional repository: off the map and into the territory abstract given the recent proliferation of institutional repositories, a key strategic question is how multiple institutions— repositories, archives, universities and others—can best work together to manage and preserve research data. in 2007, green and gutmann proposed how partnerships among social science researchers, institutional repositories and domain repositories should best work. this paper uses the timescapes archive—a new collection of qualitative longitudinal data— to examine the challenges of working across institutions in order to move data into and out of institutional repositories. the timescapes archive both tests and extends their framework by focusing on the specific case of qualitative longitudinal research and by highlighting researchers' roles across all phases of data preservation and sharing. topics of metadata, ethical data sharing, and preservation are discussed in detail. what emerged from the work to date is the extremely complex nature of the coordination required among the agents; getting the timing right is both critical and difficult. coordination among three agents is likely to be challenging under any circumstances and becomes more so when the trajectories of different life cycles, for research projects and for data sharing, are considered. timescapes exposed some structural tensions that, although they can not be removed or eliminated, can be effectively managed. introduction institutional digital repositories are growing at a rapid rate driven by factors such as institutions promoting their intellectual capital, researchers’ seeking greater control over output dissemination, research councils and other funders requiring data to be offered for deposit, and the growing support for open access initiatives (esrc data policy, 2000). a key strategic question is how multiple institutions—repositories, universities and others—can best work together to manage and preserve research data. a number of uk reports have begun to recognise the urgency of this issue (gibbs, 2007; lyon, 2007). this paper uses the timescapes archive to examine the challenges of working across institutions in order to move data into and out of institutional repositories. this paper will first describe key features of a framework green and gutmann (2007) propose for how partnerships among social science researchers, institutional repositories and domain repositories should best work. next, the timescapes project and archive will be described with a focus on its distinctive features. in a number of areas, the timescapes archive both tests and extends their framework. each topic of metadata, ethical data sharing, and preservation will be discussed in detail. lessons learned and next steps make up the conclusion. green and gutmann framework the green and gutmann framework lays out roles and relations among researchers, institutional repositories and domain repositories in relation to the life cycle for social science research. they argue that institutional repositories can play a key role by mediating between researchers (as depositors and users) and domain repositories. as definitions are not fixed in this area, i am following their typology. the distinguishing features of institutional repositories are that they: cover diverse disciplines, have tended to focus on research outputs rather than data, are committed to simplifying deposit procedures (at times employing less elaborate metadata and documentation), and believe in sustainability but often lack resources for long-term preservation. domain repositories, by contrast, often have a disciplinary theme (e.g., social science), are focused primarily on data, not outputs, and are committed to preservation while also assuring usability of their collections (green and gutmann, 2007; 4). green and gutmann present a comprehensive model describing roles for all three agents across all phases of research. this paper will highlight the parts of their model most relevant for the timescapes archive. the institutional repository is especially active early in the research life cycle. early on, it can raise issues of data sharing and archiving to researchers, well before most projects will have considered such matters. it is also well positioned to mediate between researchers and the domain repository. the domain repository can provide information and support in several ways areas: advice and forms regarding the gaining of consent for data sharing, technical advice on collecting metadata, and assisting the transfer or sharing of data and documentation to a domain repository for 14 iassist quarterly fall & winter 2007 preservation. the green and gutmann model is based on “cooperation and specialisation” among the agents. “the next step in the evolution of digital repository strategies should be an explicit development of partnerships between researchers, institutional repositories, and domain-specific repositories” (green and gutmann, 2007, 16). the timescapes project and archive the timescapes project is a £4.5 million, five year esrcfunded study designed to shed light on the dynamics of personal relationships over the life course and the identities that flow from those relationships. timescapes entails a consortium of five universities conducting seven empirical projects that investigate the life course. over 400 participants will contribute data and the archive will hold at half a terabyte of data. a key objective of this initiative is to establish a working archive of data as a resource for sharing among researchers, other authorised users, and for future historical purposes. the archive is developing a partnership between institutional and domain repositories by implementing a structure of “disaggregated preservation” (knight and hedges, 2007). an institutional repository at leeds will receive incoming content and support data preparation, metadata collection and enhancement, and data sharing. this repository will extend the existing midess system at leeds, using digitool software, and is designed to accommodate multi-media file formats. this system is now called the leeds university digital objects (ludos) . the repository will send appropriately prepared (e.g., compliant with oai-pmh and ddi standards) to the uk data archive (ukda) for preservation. dissemination versions of files (whether produced at leeds or ukda) will be available from both locations. thus the repository will have primary responsibility for ingest and dissemination with preservation “disaggregated” to the ukda. a key element in this design choice was the fact that several jisc projects had successfully used similar designs including sherpa dp2 (knight, 2005; knight and hedges, 2007), the preserv project (hitchcock, et al., 2007) and the pledge project (mit libraries, 2008). knowledge of these projects and consultation with their project staff were critical in providing reassurance that the disaggregated preservation model was reasonable for timescapes. among other reasons for adopting disaggregated preservation was the obvious one of not recreating preservation services if existing ones fit for purpose are available (knight and hedges, 2007). ludos is shifting out of its pilot phase and will establish this preservation service, along with associated digitisation, collection managements and preservation policies. in sum, the timescapes archive will provide extended functionality by integrating with existing technical and administration infrastructures at the university of leeds library, university of leeds information and systems support, and at the ukda. distinctive features of the timescapes archive although the differences between qualitative and quantitative data are often overdrawn, in regards to archiving, there are some distinctions that affect each stage of data processing from ingest through to preservation. when considering timescapes, it is useful to consider its distinctive products and processes. its product is an archive of qualitative longitudinal data. several features of this data pose particular challenges. the timescapes collection will be predominantly qualitative data rather than numeric data. much of it will be in traditional forms, such as interviews, but the collection with include other file formats such as images, audio and video. the data are also longitudinal and will be dynamically incrementing over the five years of the project. the most important implications are the capture of appropriate metadata for objects in the collection, and the personal and sensitive nature of the data that will require special treatment to comply with requirements of the data protection act and to meet ethical, as well as legal, confidentiality commitments to research participants. with respect to its process, the timescapes project is also distinctive in its ambition to simultaneously conduct and synchronise primary research, preservation and data reuse. by doing so, the project promotes a central role for researchers not only in the primary research project, but in the data archiving process as well. attention to the repository and researcher interaction has been highlighted by the jisc digital repository review's objective of better integrating repository and researcher work practices (jisc, 2005, p. 8). not only the data, but the methodology and structure of the programme are qualitative. the definition of qualitative is highly contested, but typically, research begins with questions or aims, not usually formal hypotheses. the data gathering process is often emergent and flexibly adapted during the course of the research (mason, 2002; silverman, 1985). although the green and gutmann model can encompass both, there is a feature of qualitative research that matters, and that is its iterative nature or what berg calls “spiralling” (2004). for example, in a typical survey, a sample would be established early and not change unless drop outs required replacements. in qualitative research, after an initial phase of data collection, new themes could emerge, calling for new participants, questions, or data types. the general point is that the timing and phasing the relationships among institutional and domain repositories and researchers are complex and made more so by the non-linear nature of the qualitative research process. the specific forms of this complexity will be elaborated in discussions below on metadata, ethical use of data, and preservation. iassist quarterly fall & winter 2007 15 metadata – challenges for qualitative data while all digital materials require good metadata, data (as distinct from research outputs) pose challenges for adequate metadata in part because of the more complex file formats involved (heery and powell, 2007). because it is data, it needs more extensive metadata and contextual material to render it “independently understandable” (to meet oais standards) than textual research outputs. unlike research outputs with relatively standardised formats, qualitative research data and documentation are highly diverse. firstly, they are diverse in technical file formats (txt, doc, tif, jpeg, wav, etc.). even within a format, a text document can be, to name just a few: interview, focus group, diary, field note, analytical note, and memo. it is also generally accepted that qualitative data needs extensive contextual information to enable effective reuse (fielding, 2004). much of this may fall into familiar metadata categories such as “interview” for type of data, but ideally context should also include information about the project background and even the social and institutional conditions in the wider environment that might have shaped project design (bishop, 2006). not only is there a great deal of metadata to capture, but the knowledge of that metadata is widely distributed. each agent—institutional repository, domain repository and researcher—brings specialised knowledge to metadata production. detailed descriptive metadata intended to support resource discovery is usually known best by the producer of the data. “the metadata required to access, understand, and manipulate scientific datasets will continue to be largely the preserve of domain-experts” (heery and powell, 2007). but there is evidence of problems of getting depositors to provide adequate metadata. a disproportionate share of processing staff time is devoted to collecting and correcting adequate metadata (beagrie, et al., 2008). similar problems have been encountered in repositories of learning objects (ryan and walmsley, 2003). in contrast, the domain repository expertise lies in domain knowledge and technical expertise in resource discovery and preservation: “…the domain-specific repository has specialized knowledge of data management approaches to data in a specific scientific field, for example, domain-specific metadata standards (the ddi in the case of the social sciences), as well as the ability to expose the research products to the field in a way that will have the greatest impact” (green and gutmann, 2007; 16). green and gutmann (2007) suggest that institutional repositories broker key relationships and thus enable the creation of more and higher quality metadata. defining a metadata schema for timescapes the starting point for defining a metadata schema for the timescapes archive was a commitment to openness and using (or building upon) existing standards. one requirement was conformity with the open archives initiative protocol for metadata harvesting (oai-pmh). in building a schema for timescapes, we used ukda standards as a starting point. the study description and catalogue record created at the ukda for each dataset follow the international standard for social science data, the data documentation initiative (ddi). the study metadata is also mapped to the dublin core standard, and is compliant with the open archives initiative (oai) and z39.50 for metadata harvesting and sharing. the ukda is generally compliant with the oais reference model with some “additions and alterations” based on the specific type of material processed (ukda preservation policy, 2008). timescapes is attempting to both meet existing requirements and to anticipate changes (e.g., expanded use of xml, mets and multi-media data) at the ukda. the timescapes metadata schema has been produced by the project's technical officer in consultation with the leeds library staff members. the metadata standards used were chosen because they are being actively used and they are supported by organisations that are responsible for creating standards for digital archives and preservation systems. image metadata is captured using the “niso metadata for images in xml” (mix) standard developed by the library of congress . preservation metadata is captured using the “preservation metadata: implementation strategies” standard developed by the library of congress . as with the mix standard, digitool can automatically generate premis-compliant metadata recorded during the deposit and ingest process. descriptive metadata is used to support several methods for searching the timescapes archive. the first is the use of logical collections which are a feature of digitool that allow the setting up of “slices” (e.g., women) through the metadata in the form of predefined searches on metadata fields. we have also created a set of baseline metadata elements that will be used to support the creation of logical collections and other searching. this will be mapped to the metadata object description schema (mods) standard developed by the library of congress. we have also created a specification for detailed metadata providing very detailed information about the subject (personal characteristics, employment, education, living arrangements and so on). an xml schema is in development that will be used to capture this metadata. we are using microsoft infopath software to create a userfriendly form and interface to capture this metadata. metadata foregrounds researchers' roles in archiving process a central focus of work has been diverse efforts to engage researchers in collecting and providing metadata. activities 16 iassist quarterly fall & winter 2007 included providing researchers with transcription and spreadsheet templates and involving researchers in defining the descriptive metadata and the initial filters to appear on a resource discovery page. what emerged from the experience with metadata in the timescapes project is the extremely complex nature of the coordination required among the agents. the first point of timing coordination was between the institutional repository and the domain repository. we began the metadata schema intending to follow and, as much as possible, fully replicate standards in use at ukda. this was largely successful, but illustrative differences appeared. first, the ukda had not settled on metadata specifications for audio and video. and though it is in the process of doing that work, the scheduled completion date is later than what timescapes required. the situation is similar regarding item-level metadata. this is metadata that applies to a unit smaller than the full dataset or collection, such as a single interview for a qualitative collection. currently, the ukda provides some item-level metadata in the form of a spreadsheet that is part of the documentation for a study. this metadata does not enable the user to search for key categories (e.g., gender) across all collections. (qualidata online does have this capability, but currently holds only a small share of total holdings.) work is proceeding in these areas, but it is not expected to be completed in time to meet timescapes need for itemlevel metadata for resource discovery and, even more urgently, for access control. finally, there is the role of the researchers and the research life cycle to add to the mix of institutional repository and domain repository data life cycles. because researchers are using a flexible, emergent model for data collection, they can not be certain about the scope of data to be collected. this posed a challenge when the metadata specification called for some points to be nailed down earlier than was comfortable for some researchers. for example, in choosing filters for resource discovery, researchers complained that they had not yet started collecting data, how could they be expected to know how to search it or to define key subject categories? ultimately, these requirements did not cause lasting problems. although initial choices had to be made to develop the search, few of these choices are permanent. it is clear that a key point of communication involves making clear to researchers that decisions must be made, but also making clear which choices create “lock-in” and which are more flexible. generally, this has been achieved by good communication among the research archivist, technical officer and researchers. ethical data sharing the collection, use, publication and dissemination of data is subject to an extensive array of guidelines for its legal and ethical use, ranging from requirements of the data protection act, to review ethics committees’ requirements, to guidelines issued by various professional bodies such as the british sociological association and the medical research council. although timescapes data pose particular ethical challenges, even relative to qualitative data generally, the project is explicitly focused on developing strategies to enable sharing and preservation of even this most challenging data. domain repositories and the ukda in particular, have long history and great expertise in finding ways to ethically share and reuse data. there are three inter-related strategies available to make data shareable: gaining informed consent for sharing and archiving, altering data to protect identities (e.g., anonymisation) and controlling access. timescapes is following the ukda model of integrating all three of these strategies and extending ukda procedures in selected areas to address the particular requirements of qualitative longitudinal data. this paper looks in detail the area of informed consent with a brief note on anonymisation. informed consent is an ethical requirement for most research. it must be considered and implemented throughout the research life cycle, from the inception of planning to publication and including making provision for data sharing. researchers and archives share strong commitments to ethical use of data, from point of collection through to preservation. but as we have seen with metadata, agents occupy different locations in research and data life cycles which influence the particulars of how these commitments are perceived and acted upon. if research data to be archived at the ukda contains personal or sensitive data about informants, explicit consent is needed, ideally in writing, for such data to be processed by ukda. if a person can not be identified—if the data are fully anonymised—then the dpa no longer applies. the challenge for much qualitative data is that it can be difficult to assure absolute anonymisation, and thus the safest stance from a legal perspective is to have written consent in place. for researchers, the legal framework is far from clear (thomas and walport, 2007). the vast majority of researchers are deeply committed to ethical use of data and see these standards as more binding than any legal formalities. however, it is not surprising that researchers emphasise their areas of experience: recruiting, contact with participants, and publication. in most cases, their attention to data sharing is a lower priority. as more funders recommend data sharing, this issue is growing in prominence, however, most funder mandates remain voluntary recommendations with infrequent formal enforcement. iassist quarterly fall & winter 2007 17 the process of producing a model consent form acceptable to all the projects took a great deal of time, including drafting the form, holding consultations, incorporating feedback and keeping the wording of the timescapes form aligned with changes and updates in ukda policies. this process of engaging researchers has yielded a standardised consent form that covers areas of consent for participation, research outputs and data sharing and archiving. the outcome has been positive, though not ideal: most but not all the teams have agreed to use the form, although some will use recorded verbal consent or defer seeking consent for archiving. a similar issue arose concerning the second strategy for data sharing: anonymisation. a set of guidelines were drafted with instructions about what content to anonymise and formats for doing so. the system had to be easy to teach to transcribers (some projects works with pools of transcribers and have little or no control over their quality), and to make it possible to easily convert files to xml. as with the metadata case, informed consent showed that understanding agents' locations in their respective life cycles is essential for successful collaboration. the ukda's duty is to meet legal requirements regarding data sharing, and it recommends written informed consent. researchers want to minimising burdens on participants and thus they prefer fewer formal procedures and more discretion. as in the metadata example, the role of researchers as participants in design processes featured in dealing with consent. preservation it is probably the area of preservation where the timescapes archive is most clearly attempting to follow the strategy of “cooperation and specialisation” advocated by green and gutmann. the logic behind this strategy is to maximise effective use of resources, and in theory, it is a compelling argument. at the experiential level in timescapes, we are convinced is it the right approach, however, it has proved important not to underestimate the challenges of coordination. the platform for the timescapes archive, ludos, is a development involving the university library and information systems services. the time invested has been well spent, enabling the creation of a partnership for the parallel development of the platform and archive. nonetheless, there have been hurdles to overcome. timescapes is committed to compliance open source, non-proprietary tools where available. however, the midess project at leeds was in place and already running digitool, a proprietary package. choosing a new software platform such as fedora would have forced timescapes outside of the existing library-managed repository system at leeds. the library stood to benefit from a large project committed to depositing in its repository and timescapes needed an institutional base at leeds. we deemed that a vital key to long-term sustainability for the timescapes archive would be its embeddedness in the wider library and is infrastructures at leeds, and we agreed to accept the limitations of the proprietary software application. regarding the ukda relationship, the agreement to collaborate with the ukda for preservation was obvious; as an esrc funded project, timescapes is obliged to offer its data for deposit. irrespective of any mandate, the ukda goals of preservation, balancing authenticity with usability are shared by timescapes, making the ukda an ideal long-term home for the collection. timescapes is currently actively engaging with the ukda, taking the lead in some areas and following ukda policy in others. for now, we are following a “pull” model, that is, letting the service provider offer guidance on preservation, then attempting to implement the necessary procedures as possible within the timescapes project. in addition to the metadata cooperation already discussed, we have recently sent data already deposited at ukda (from a “feeder” project for timescapes) to leeds. the data and metadata are being used to test the draft metadata schema and infopath ingest and deposit procedures. the objective is that timescapes data will be processed to a very high level of quality, standard compliant, and more “depositready” than the typical dataset received at ukda. researcher-centred archiving green and gutmann's (2007) central insight is their understanding of critical roles for all three agents, researchers and institutional and domain repositories in preservation and situation those agents in their own life cycles. they focus in particular on the ways institutional and domain repositories can collaborate. timescapes extends this model by highlighting researchers’ roles across all phases of data preservation and sharing, including selecting which data are to be preserved, enriching that data with contextual and other metadata, and identifying and promoting opportunities for reworking the archived data. in doing so, timescapes is (implicitly if not always explicitly) following some principles of cooperative, or participatory design, at least to the extent that users’ (i.e., researchers') needs are met and that high standards of usability are achieved. typically, participatory design projects have a dimension of user-empowerment. this is a consideration in timescapes in the following way. repositories and archives may be used in a managerialist fashion to reduce researcher autonomy. this can come about through greater centralised control of research and teaching resources by requiring sharing on conditions not discussed or negotiated with researchers. one objective of timescapes is to promote alternatives to this top-down model. this is not to say that 18 iassist quarterly fall & winter 2007 researchers’ interests should dominate the world of data sharing, but they should be equal partners—along with repository and archive experts, university administrators, and others, including some segments of the public with interests in data sharing—in the development of such systems. in addition to engaging users in metadata collection, interface design, and consent and anonymisation guidelines, timescapes is also seeking researcher involvement very early in promoting the archive for reuse. one of the sychronicities of timescapes is that fact that we are planning and designing for secondary analysis while primary data are being collected. this is unconventional and uncomfortable for researchers. their response is (though usually voiced more diplomatically): “why hassle me about reuse now; i have not even recruited my sample for the primary research yet?” it is hoped that this early planning will help to assure an active community of researchers committed to reusing data as soon as materials become available. several factors about timescapes may enable this vision to be achieved. the programme will produce the largest dataset of its kind with a substantive focus on the life course and a methodological focus in qualitative longitudinal methodology. these thematic foci allow timescapes to target potential re-users from an existing community of researchers, many of whom already have a history of sharing and collaboration. our strategies for building a community of users include encouraging affiliated projects (where affiliates will be required to deposit and share their new data and to re-use timescapes data), securing secondary analysis studentships, and providing mobile and in-house training workshops and a help desk. we also aim to showcase data sharing and re-use among the seven timescapes projects. data sharing sessions have already taken place among projects with common interests (e.g. parenthood, childhood, older lives). active promotion of the archive as a specialist data resource has begun, starting with an extensive consultation with potential end users. however, timescapes is already acting as a magnet for researchers interested in developing affiliated projects. the key to this interest lies in the thematic focus of the archive, enabling researchers who share broad substantive interests (e.g., family, relationships, life course) to also share data management protocols, methodologies, research data, and outputs. conclusion green and gutmann have been excellent tour guides and their map has proved itself an accurate one as timescapes has navigated the difficult terrain of archiving qualitative longitudinal data. they clearly grasp the need for all three agents, researchers, institutional repositories and domain repositories, to engage as equal partners in producing and sharing data. particularly useful insights are gained by situating agents in the contexts of their local practices: research and data life cycles. the timescapes archive has deepened the usefulness of this framework by providing details of how cooperation and specialisation can work in a live project. and timescapes has extended their model by using the case of qualitative longitudinal data to demonstrate the necessity for finelytuned timing if coordination is to work. the metadata example showed how this coordination can happen. in its role as domain expert, the ukda has demonstrated the value of relying on international social science standards such as ddi. timescapes is advancing development in the areas of audio and video, specific to its needs. the institutional repository at leeds, by working closely with researchers, has greatly expanded and regularised the metadata that will be collected for timescapes. this can only enhance resource discovery for the timescapes data and the potential for comparative, mixed methods research with other, large-sample quantitative, data. informed consent demonstrated the power of specialisation with the ukda focused on compliance with complex legal requirements and researchers on their ethical responsibilities. finally, the model of disaggregated preservation adopted clearly showed that institutions can specialise, yet still work together to use valuable data sharing resources in the most efficient way possible. even the best maps and guides can not remove every obstacle from a journey. coordination among three agents is likely to be challenging under any circumstances, and becomes more so when the trajectories of different life cycles, for research projects and for data sharing, are considered. timescapes exposed some structural differences that, although they can be managed, can not be removed or eliminated. repositories, both institutional and domain, tend toward needing fixity and formality in areas such as standard and guidelines. timescapes methodology tends toward less fixity and formalisation, especially in early phases of work. so researchers, by and large, are pushing to keep things loose while repositories need to nail things down. the tension is, of course, compounded in timescapes because of its commitment to engage researchers throughout the data sharing process. what have we learned so far? first, that the tensions of coordination are not “solvable”; no amount of planning or anticipation will remove them. the tensions are inherent in the different roles and perspectives of the various agents. if all agents need to participate (and all evidence suggests the benefits are worthwhile), then effort has to be put toward managing the tensions constructively. this management takes resources, and that much at least can be planned for. cross-institutional teams need to be used with regular, substantive meetings, not mere semi-annual formal iassist quarterly fall & winter 2007 19 sessions. in the face of day to day frustrations, much remains positive. if we are to create archives that researchers will not merely use, but actively support and fight for, then these archives have to be built, from the beginning, with researcher input. equally important, if those archives are to be sustainable and obtain long-term funding, we have to embed archives in on-going institutions and demonstrate efficient use of all-too-scarce (and given prospective economic conditions, likely to become more scarce) resources for this valuable endeavour. *contact: libby bishop, research liaison office-ukda at university of essex and research archivist at university of leeds. e-mail: e.l.bishop@leeds.ac.uk references beagrie, n., chruszcz, j., and lavoie, b. “keeping research data safe.” jisc report, may 2008. http://www.jisc.ac.uk/ publications/publications/keepingresearchdatasafe.aspx berg, b. qualitative research methods. boston: pearson education, 2004. bishop, l. (forthcoming 2008) 'archiving for the future: the archivist as researcher' in: m. brugidou, et al. (eds.) secondary analysis in qualitative research: challenges for human and social sciences. paris: lavoisier. bishop, l. (2006) “a proposal for archiving context for secondary analysis”. methodological innovations online 1(2). http://sirius.soc.plymouth.ac.uk/~andyp/viewarticle. php?id=26. esrc data policy, april 2000 http://www.esrcsocietytoday.ac.uk/esrcinfocentre/ images/datapolicy2000_tcm6-12051.pdf). fielding, n. (2004). 'getting the most from archived qualitative data: epistemological, practical and professional obstacles', international journal of social research methodology, 7(1), pp. 97-104 green, a. and gutmann, m. (2007) “building partnerships among social science researchers, institution-based repositories and domain specific data archives”. oclc systems & services: international digital library perspectives 23(1): 35-53. http://deepblue.lib.umich.edu/ handle/2027.42/41214 [open access version] gibbs, h. (2007) disc-uk datashare: state-of-the-art review. disc-uk, august 2007. http://www.disc-uk.org/ docs/state-of-the-art-review.pdf heery, r. and powell, a. (2006). digital repositories roadmap: looking forward. bath: ukoln/eduserv. http://www.ukoln.ac.uk/repositories/publications/ roadmap-200604/ hitchcock, s., brody, t., hay, j.m.n., and carr, l. (2007) “digital preservation service provider models for institutional repositories: towards distributed services,” d-lib magazine 13(5/6), january 2007. http://www.dlib. org/dlib/may07/hitchcock/05hitchcock.html joint information systems committee (jisc), digital repositories programme. (2005) digital repositories review. http://www.jisc.ac.uk/media/documents/programmes/ digitalrepositories/digitalrepositoriesreview2005.pdf knight, g. (2005) “sherpa-dp oais report: an oais compliant model for disaggregated services”. arts and humanities data service report. http://ahds.ac.uk/about/ projects/sherpa-dp/sherpa-dp-oais-report.pdf knight, g. and m. hedges. (2007) “modelling oais compliance for disaggregated preservation services.” the international journal of digital curation 2(1) june 2007. http://www.ijdc.net/ijdc/article/viewarticle/25/0 lavoie, b. (2004) the open archival information system reference model: introductory guide. digital preservation coalition (dpc) and oclc. http://www.dpconline.org/ docs/lavoie_oais.pdf lyon l. (2007) dealing with data: roles, responsibilities and relationships, consultancy report. bath: ukoln. http://www.jisc.ac.uk/media/documents/programmes/ digitalrepositories/dealing_with_data_report-final.pdf mason, j. qualitative researching. london: sage, 2002. thomas, r. and m. walport. data sharing review. ministry of justice. http://www.justice.gov.uk/reviews/ datasharing-intro.htm mit libraries. pledge. [accessed 16 sept 2008]. available from http://pledge.mit.edu ryan, b. & walmsley, s. (2003) implementing metadata collection: a project’s problems and solutions. learning technology, vol. 5, no. 1, jan. 2003. http://lttf.ieee.org/ learn_tech/issues/january2003/index.html#3 sherpadp2. project overview. [accessed 25 nov 2008]. http://www.sherpadp.org.uk/sherpadp2.html silverman, d. qualitative methodology and sociology. aldershot: gower, 1985. 20 iassist quarterly fall & winter 2007 uk data archive. preservation policy, 2008. http:// www.data-archive.ac.uk/news/publications/ ukdapreservationpolicy0308.pdf footnotes 1 http://ludos.leeds.ac.uk/ludos/ 2 http://www.loc.gov/standards/mix/ 3 http://www.loc.gov/standards/premis/ 4 http://www.loc.gov/standards/mods/ lassist newsletter, vol. 2, no. 1 (winter 1978) the network-based scientific cohhunity economic climate and social structure richard c. roistacher center for advanced computation university of illinois, urbana, il 61801 abstract the effort is made in this paper to identify those elements of the scientific community requiring computer services; evaluate the economies of using computing network facilities to satisfy those needs; and to provide some suggestions as to the substantive advantages of network use. economic climate several tors make way to o search.. t market has of talent universiti university trial sett tutions a ph.d.s wit research, smaller in facilities ganizing search. s cient res institutio research f snaring is such shari by allowin research f basis. h costs wil expensive. social networki rganize he decli resulte from th es into , govern mgs. re now h firsthowe ve stitutio and a for the ince the ources t n with a acilitie necessa ng has g outsid acilitie owever, 1 make and econo n g an at scien ti ning acad d in a di e major a wide va ment, an many smal staffed b class tra r, many ns lack tradition conduct re are no o provid complet s, some ry. in t been acco ers to u s on a increasin travel ev mic trac ic emic sper rese riet d m 1 in y y inin of rese of of t su e e e se for he p mpii se m visi g en f active rejob sion arch y of dusstioung g in the arch orreffivery t of m of ast , shed a 3or ting erg v more at of en of c have and dropp years conti the e compu trave cilit acces small tions the ergy om pu been comp ing f a nue xten tati 1 an ies. s t and sa has tati dec utat by t ren inde t t on d t it o re ser me time increas on and lining, ion cos one hal d which finitely hat comra can su he dupli is poss search vice ori tnat ed, t commu commu ts ha f eve is exp unicat fastitu cation ible t faciii ented the cost he costs nication nication ve been ry five ected to hus, to ion and te for of fao expand ties to institucommunication networks and data bases are natural monopolies; they function best when a single source serves the largest possible clientele. groups of scientific workers are probably best suited to an environment of pure competition in which there are no barriers to the setting up of new groups. the necessary communication network is already rapid and c been satis numbe task commu ers i them, them they in p iy. § ommerci estabii factory r of cl in est nit ies s to i f inan with t need. lace a everai al da shed a servi ients. ablish of sci nf orm ce th he sp nd is major ta arch nd are ces to thus, ing net entif ic people, em , an ecific expanding university ives have providing a growing the major work-based researchorganize d provide resources ing stil new the eval and and such tech ronm resu soci exte cien netw supp v e n t h incr ea 1 an scien dissera uation govern instit thing n o 1 o g y ental it in al sci nt tha t way orking ort. ougn singl evertific inati find ment ut ion s as asse impa an i ence t net to co wil reso y sc mcr kno on ings ope al prog ssme ct ncre res work nduc 1 re urces arce , easing wledge of res into r a t i on reguir ram ev nts. stat em asin ea ing is t such ;ing trch. are th nee an earc comm s. emen alua and ents dema t an res fin growere is d for d for h and ercial legal ts ror tiop. s, enviwill nd for o the eff iearch. ancial users of scieniiiic netwoms in this discussion, users of scientific networks will be classified by their relation to the network, rather tnan by job title or institutional affiliation. scientists a enti whos disc edge in u cies dust sibl of a prof but tion n ob fie e p over m nive and ry. e th mate essi with s as viou rese rima y of os t rsit lab net e s ur s onal out sci s cl arch r y 1 new scie ies, orat work uppo cien tr ins enti ientele net wor nterest scien ntists gove ories, ing als rt and tists , aining tit ut 10 sts. tor k is is tif ic are e rnraen and o mak inte peop in nal a a scipeople in the knowlmployed t agenin ines posgration le with science ffilia19 lassist newsletter, vol. 2, no. 1 (winter 1978) scientist-professional social structures in ~ ~ :^et5dlk£liied~5r0ups an increasing number of doctoral level graduates are going directly the physical and computational into social service and clinical resources required to support a positions. while many people in scientific network are organized in such positions are not active proda formal fashion, since they are ucers of research, they are often subject to a host of financial and avid consumers and critics of scilegal constraints. however, the entific research. scientistprosocial structure of network-based fessionals have many interests scientific groups is still not which are not shared by pure and clear. we have had little experiapplied scientists, and would probence with forming and operating ably form their own network, based such groups, and there is no regroups.. however, it would be easy guirement for any strict division to maintain communication between into types. however, it is possigroups of scientists and scientistble to guess at some types of professionals when dictated by comgroups which may evolve, mon interests. practitioners task groups task groups are network-based practioners, as used here, are research projects. several people people whose interest is primarily would use the network to coordinatem the use, rather than in the distheir joint work on a scientific covery or analysis ot specific problem. although task groups knowledge. a network of profeswould consist of people at differsionals would obviously wish to ent locations and perhaps different make use of scientific research reinstitutions, their research would suits. occasionally, groups of probably be funded and operated as professionals might wish to obtain a joint project, consulting help trom scientists. contacts between network-based groups of scientists and groups of professionals woula probably tend consortia to be mediated by institutional facilities such as librarians and groups of workers might form transfer agents, more tban by perconsortia for the purpose of consonal acquaintance. tracting for data. the members of the consortium would pay for the data collection and for the establishment of a common data base, but librarians and transfer agents would pursue the independent analysis of their own data. these are people whose primary interest is in providing services to network clients. reference librarians would be able to act much invisible colleges as they do now, but would have better access to potential clients and the most obvious form of organibetter knowledge of current issues. zation, and the one which is presthe term "transfer agent" applies ently most common, is the "invisito someone wno would be part organble college." groups of people izationax development specialist with common interests use the netand part extension agent. the work to exchange messages, manutransrer agent would act as referscripts, screeds, broadsides, and ral service, social director, and gossip. people pay their own exspreader of information. transfer penses and pursue their own interagents will become increasingly ests, using whatever public facilinecessary as the demand for informties they need. the invisible ation and the complexity of ir:tormcollege forms a ground for recruitation sources both increase. ing members into more formally organized groups such as task groups it can be expected that librariand consortia, ans and transfer agents, having issues of their own, will form their own network-based groups. network journals by combining the document processing, information retrieval, and communications facilities of a computer network with a system of edi20 lassist newsletter, vol, 2, no. 1 (winter 1978) tors and referees, it is possible archival analysts to produce a network-based journal. several schemes have been proposed a second clientele for a refor the establishment and operation search network is social scientists of such a journal. one of the most who analyze machine-readable archiattractive schemes would allow the val data. such clients would find journal to "publish" all submisaccess to a network highly rewardsions, but with the addition of a ing for both research and communireferee's score. a network journal cation activities. since the netwould allow formal and impersonal work provides data archive and contact between network clients on computational facilities which are the basis of evaluated research equivalent to those at a major unifindings, rather than personal acversity, archival analysts will quaintance. find tne network's facilities equal or superior to those of their own institutions. archival analysts shoula find affiliation with a netbrain trust work both professionally and socially rewarding. the ability to send messages to large groups of people makes possible the use of an invisible college as a source of otherwise unobtamalocal firoducers o_f data ble information. someone with a question may begin by asking cola third category of scientists leagues or the reference librarian. consists of those who can meet if all else fails, it is possible their own instrumentation needs to broadcast a message asking for locally. instrumentation as used help with the problem. not only here includes all those facilities may someone out there have an anand activities required to produce swer, but other people with the data either in machine-readable same problem may now share the anform or in form suitable for key swer. one of the chief tasks of entry. scientists in this category transfer agents may be to bid out are able to raise their own rats, and pass on questions and answers. use their own instruments, or administer and code their own questionnaires. such scientists can use the network solely as a commuosers of a scientific network nication medium, to analyze their own data, or to collaborate and membership in network-based reshare locally produced data with search communities, while far betothers. ter than isolation and inactivity, will provide neither for all needs of scientists nor for the needs of all scientists. the most obvious data contractors shortcoming of scientific networks is that there is no way that they can provide direct access to labo__ ratories and instrumentation. scitists who can contract for data, entists can be classifiea into five work of this type requires the excategories with respect to their act specification of data collecdependence on instrumentation. tion procedures, but does not necessarily require that the investigator actually operate the eguipment. many physical scientheoret icians tists essentially contract for data, buying time on such facilisome scientists, either mathematies as telescopes and nuclar reacticians or theoreticians, require tors. survey researchers often buy access only to libraries and colinterview time and questions on naleagues. participation in a nettional surveys, work-based community provides them with a peer group and an audience. most of today's data contracting a network-based clientele of theofacilities are national resources, reticians would be relatively easy so huge and costly as to be beyond to support. in fact, social supthe means of most universities, port is perhaps the only unique examples of such facilities are the service a network would provide to mt. palomar telescooe, the stanford the unaffiliated theoretician. linear accelerator^ and the national election survey. it might be feasible to establish data contracting facilities equivalent to those at a major university. such facilities would provide "contract a fourth category of potential network clients consists of scienlassist newsletter, vol. 2, no. 1 (winter 1978) data on a scale and at a price inte rna l funding affordable by the individual or by small groups of researchers. conaccess to networks will be diftract laboratories staffed by reficult for those who operate ensearch assistants, technicians, and tirely within institutional budga supervisor would perform experiets, without control of their own ments according to protocols profunds. most universities greet the vided by remote researchers. while very mention of purchasing external it is not clear which, if any, computing services with fear and sorts of research could be pursued loathing. deans and department in such a fashion, the feasibility heads do not see why "real" money of such contract facilities must be should be used to buy computing investigated. when "free" computing is available locally. there are no easy answers ft second way in which scientists to such questions, and the individmight successfully contract for ual scientist or faculty member is data is through the formation of in a poor position to change insticonsortia. a network-affiliated tutional constraints, scientist could propose that a group of researchers jointly support the costs of contracting for data, which they could then analyze extern al funding jointly or as individuals. one strategy suitable to external funding is to have the sponsor write a special condition to the hands-on experimentalists grant requiring the principal investigator to nave discretion over a fifth type of scientist rewhere computing money may be spent. ?uires personal access to equipment another alternative is to have the ound only in the research laborafunding agency execute a separate tories of a major university. it is contract with the network facility, difficult to see how such sciensetting up an account for the retists could use network facilities searcher. the researcher's univerfor anything other than scientific sity thus never exports any computand social communication. ing' funds because it never imported any computing funds. financi ng terminal costs computer conferencing it has a policy of denying faculty members access to professional the on line conference not only none of these things will happen allows for consulting over a netunless there is money available to work and for scientific communicapay for them. people's major cotion of a new and powerful kind, nern is with the acquisition of but also provides a point of access terminals and with paying for netto networking. by definition, a work access. the price of termilocal computing facility cannot nals has declined precipt iously provide access to a conference takover the last five years. a termiing place on a remote computer, nal which cost about $3000 five thus, the would-be conference user years ago can now be bought for cannot be told that the remote site less than $1000. it is unlikely has no unique facility. also it that the decline in the prices of should be difficult for a univerterminais will continue to be as sity administration to state that steep; however, the price of excellent printing terminals is fast approaching that of ordinary office meetings, electric typewriters, and will probably soon be below the cost of there are no easy answers or present typewriters. thus, the ofeasy predictions as to when netrice typewriter of the near future works will grow large enough and will be nothing more than a compowerful enough to attract a major puter terminal lacking communicarraction of scientists as users, tions equipment. turning the ofhowever, the example of telenet, fice typewriter into a terminal growing from five to 85 cities in will require nothing more than the two years, gives some cause for opinstallation of a oneor two-huntimism. dred-dollar communications circuit board. 22 instructions for authors of the iassist quarterly 1/15 wenzig, knut & han, xiaoyao (2024) state of the ddi cloud, iassist quarterly 48(4), pp. 1-15. doi: https://doi.org/10.29173/iq1116 the creative commons-attribution-noncommercial license 4.0 international applies to all works published by iassist quarterly. authors will retain copyright of the work and full publishing rights. state of the ddi cloud knut wenzig1 and xiaoyao han2 abstract as the ddi community continues to grow, an increasing number of repositories are providing their metadata in various ddi formats. however, the current landscape of ddi metadata standards usage is not well understood. understanding this landscape is crucial as it helps identifying usage patterns, improve interoperability, and guide future developments. to address this research gap, we investigated the availability and comprehensive element usage of ddi standards across 29 repositories registered on the platform re3data.org, using the oai-pmh api. by analyzing approximately, a quarter of a million metadata records in ddi-codebook format, we summarized statistics on the usage of popular ddi elements and their distribution across repositories. our findings may have implications for the deployment of ddi metadata and the further development of these standards. they also inform researchers and data stewards about how ddi-codebook is utilized by the community. overall, this investigation underscores the value of openly available metadata in supporting research and achieving the goals of the fair data movement. keywords ddi, metadata harvesting, oai-pmh, ddi-codebook, data catalogues introduction over the years the diagram for the lod cloud (mccrae 2024) shows the success story of linked open data: when it started in 2007 the lod cloud reported the existence of 12 interconnected datasets (resolvable and accessible rdf data with at least 1000 triples). in 2023, 16 years later, 1,314 datasets build the lod cloud showing that the idea of publishing interconnected data is attracting an increasing number of users. driven by the goals of standardization, linking and re-use, the standards of the ddi alliance follow a similar mission in the world of research data: from ddi-codebook (ddi alliance 2014) to ddi-lifecycle (ddi alliance 2020), the ddi standards are constantly becoming more extensive and complex. we are interested in which parts of the standards are used to inform users and development. therefore, we collected and analyzed ddi metadata in the wild. as a result, we created an overview of the ddi cloud: sources where ddi metadata can be found, the structure of published metadata itself, as re-use is intended (there can and should be links between the various entities) so the term “cloud” is appropriate here, too. https://doi.org/10.29173/iq1116 https://creativecommons.org/licenses/by-nc/4.0/ 2/15 wenzig, knut & han, xiaoyao (2024) state of the ddi cloud, iassist quarterly 48(4), pp. 1-15. doi: https://doi.org/10.29173/iq1116 version 2.5 of the standard ddi-codebook is the latest version of ddi-codebook and has been last published with some backward compatible bug corrections in 2014. version 2.6 is currently under development and about to be released soon. a ddi-codebook compliant xml file (see figure 1) can cover 5 main areas, which are located on level 2 within the element <codebook> on level 1: 1. the optional and repeatable description of the xml file itself is located within the element <docdscr>, where for example bibliographic information for the file can be stored in an element <citation>. 2. the mandatory and repeatable element <stdydscr> describes the study and holds information about the data collection (or compilation) and general information including title, abstracts or keywords. 3. the element <filedscr> is used to describe the files that comprise the collection documented by the ddi-codebook xml. it is optional and can be repeated for multiple files. 4. the optional and repeatable element <datadscr> holds information about variables and their categories. 5. other materials that are related to the collection/study can be described using the element <othermat>. <codebook[1]> <docdscr[0-n]> ... bibliographic information describing the ddi document itself ... </docdscr> <stdydscr[1-n]> ... information about the data collection, study, or compilation ... <citation[1-n]> <titlestmt[1]> <titl[1]>text of the title</titl> </titlestmt> </citation> </stdydscr> <filedscr[0-n]> ... information about the data file(s) ... </filedscr> <datadscr[0-n]> ... description of variables ... </datadscr> <othermat[0-n]> ... other materials that are related to the study ... </othermat> </codebook> figure 1: ddi-codebook metadata as pseudo-xml-code with some remarks on the payload information. the cardinality in the square brackets is not part of the xml code: [1] stands for mandatory and non-repeatable element, [1-n] means mandatory and repeatable, [0-1] means optional and non-repeatable, [0-n] is optional and repeatable. https://doi.org/10.29173/iq1116 3/15 wenzig, knut & han, xiaoyao (2024) state of the ddi cloud, iassist quarterly 48(4), pp. 1-15. doi: https://doi.org/10.29173/iq1116 as figure 1 shows, at least the title of the study in the element <titl> is mandatory. it is the only mandatory element within the standard. all together ddi-codebook 2.5 specifies 252 different elements: 243 global and 9 local ones. in social sciences tabular data are widely used. the most prominent but proprietary data formats in social sciences (like spss, stata, sas) already contain per default different metadata like variable labels and value labels, often combined with multilingual features. providing those metadata would be straightforward, relatively inexpensive and undoubtedly in line with each aspect of the four fair principles (wilkinson, dumontier, aalbersberg, et al. 2016), which deal in the end with the availability of metadata: (1) if the metadata, which typically form the basis for data catalogues, are richer and contain information on variables, findability would increase; (2) the availability of fine-grained metadata would be possible, even if the data are no longer available; (3) more metadata would be published using a formal, accessible, shared, and broadly applicable language for knowledge representation, which would increase interoperability; (4) data are described with more relevant attributes, which would contribute to reusability. this kind of metadata can be stored below element <datadscr>. while ddi-codebook covers the needs of social science data archives, ddi-lifecycle for example can describe questionnaires and introduces better options for metadata re-use. methodology to obtain a comprehensive overview of the usage of ddi metadata elements, we decided to use re3data.org (https://www.re3data.org) as our primary platform for data collection. as a global registry of research data repositories, re3data.org records a growing number of data providers covering a wide array of academic disciplines. these data providers offer research data and its metadata (ideally exposing metadata via interfaces) and/or service providers (e.g., a portals) that harvest the metadata of research data from data providers as a basis for building value-added services. (see table 1 for numbers and a comparison to 2017.) the registry went live in autumn 2012 and has been funded by the german research foundation (dfg). since the end of 2015 re3data.org is managed under the auspices of datacite. the family of ddi standards is ranked third of all reported metadata standards, after dublin core, and the data cite metadata schema. as of 2024, re3data.org lists 280 registered repositories that report using ddi metadata standards. only ddi standards are suited for tabular data since they enable storing fine granular metadata on variable level like variable and value labels or descriptive statistics. https://doi.org/10.29173/iq1116 https://www.re3data.org/ 4/15 wenzig, knut & han, xiaoyao (2024) state of the ddi cloud, iassist quarterly 48(4), pp. 1-15. doi: https://doi.org/10.29173/iq1116 inclusion criteria to make the metadata readily available, re3data.org allows for reporting various api options. swords, rest, and oai-pmh are the most common apis under these repositories. we compared these apis for our research purposes and found that sword and rest are not as feasible or convenient as oai-pmh: sword (cottage labs 2021) is in the first place a deposit protocol, but one could also retrieve metadata from an object. none of the repositories (besides one, which does not use ddi) that offer sword access provide a valid url with details for the sword access, therefore, we could not check whether ddi metadata could be retrieved this way. the single repository, which provides actual access via this protocol, requires users to have credentials – which makes sense as sword is used to automate deposit processes. the usage of a rest api (fielding 2000) would need customized access for each repository. the responses to these queries would structurally need to be harmonized, as they typically will not result in a standardized output. nevertheless, we checked manually the rest apis of the few repositories, that reported using ddi and some software other than dataverse (which also allows metadata access via oai-pmh, see below). in the documentation of those apis, we did not find any indication that ddi metadata will be available via rest. for these reasons, the two most frequently mentioned protocols are not suitable for harvesting and analyzing ddi codebook metadata. oai-pmh (open archives initiative 2015) stands out as a straightforward and efficient protocol for harvesting metadata from repositories. its standardized approach proves particularly advantageous for large-scale metadata harvesting from multiple repositories, aligning perfectly with our research objectives. consequently, we chose oai-pmh as the method to acquire the required metadata. out of the 280 registered repositories that report using ddi as a metadata, only 85 report using ddi as a metadata standard and support oai-pmh. the repositories do not report whether they make ddi metadata available via oai-pmh, however, they state independently whether they use ddi as a metadata standard and/or whether they offer an oai-pmh access to their metadata. table 1 gives an table 1: overview of the number of repositories listed in re3data.org by type, metadata standard used and provided api for 2017 and 2024. source: own research and wenzig (2017) 2017 2024 research data repositories on re3data.org 1,989 3,179 among data providers 1,794 2,916 among service providers 765 1,075 repositories that report use of dublin core 175 561 repositories that report use of data cite 78 386 repositories that report use of ddi 116 280 among report sword 20 116 among report rest-api 28 115 among report ddi 14 85 https://doi.org/10.29173/iq1116 5/15 wenzig, knut & han, xiaoyao (2024) state of the ddi cloud, iassist quarterly 48(4), pp. 1-15. doi: https://doi.org/10.29173/iq1116 overview of repositories on re3data.org by type, top-3-metadata standards used and api for the years 2017 and 2024. before embarking on metadata harvesting through the api, we conducted essential preparatory work. initially, we manually scrutinized the validity of the api addresses provided for each repository listed on re3data.org. the inclusion criteria for repositories in our research are as follows: 1. the repository is registered on re3data.org. 2. the repository provides ddi as one of the metadata standards. 3. the metadata is available through the oai-pmh api. 4. there is a valid link pointing to the api endpoint. finally, we identified 29 repositories that satisfied these inclusion criteria, forming the basis for further processing. we reported issues such as incorrect links on websites and inaccuracies in ddi metadata standards categorization to re3data.org. they either rectified the inaccuracies or provided updated information. table 2 shows how the number of 85 repositories reporting the usage of ddi and offering an oai-pmh api breaks down to the 29 repositories we analyzed. data collection once the valid repositories were identified, we employed a python script to conduct metadata harvesting. the script interacted with the api interfaces of each repository, enabling us to traverse the xml metadata tree. to harvest metadata from the oai-pmh api, two essential pieces of information were required: the api address and the metadata prefix. after manually validating links and collecting prefixes, the information underwent further processing through a python script for metadata collection. the oai-pmh protocol does not require that every item should be available in all formats supported by the repository. since we received information that different items might actually be accessible when using different prefixes to retrieve records, we decided to include all records we could access using the various ddi-style prefixes. this includes the prefixes with different styles and languages, like ddi25, ddi-c, ddi25-en, oai_ddi 25-nl. table 3 shows a complete list of those metadata prefixes we found and identified as related to ddi standards. because only two repositories publish metadata in ddi-lifecycle, we excluded ddi-lifecycle from our study. (the prefixes in brackets, which can be found in table 3 indicate, that there are metadata in ddi-lifecycle.) table 2: status of the repositories, that report on re3data.org to use ddi as a metadata standard and oai-pmh as an api status of the repository number of repositories url to oai-pmh in re3data.org is obviously wrong or a duplicate 36 ddi metadata unavailable (ddi metadata is not available via oai-pmh) 9 the answer of oai-pmh query contains permanent errors 3 merge to upper level of repository 8 oai-pmh with ddi codebook 2.5 up and running (one of them offers also ddi lifecycle 3.2) 29 https://doi.org/10.29173/iq1116 6/15 wenzig, knut & han, xiaoyao (2024) state of the ddi cloud, iassist quarterly 48(4), pp. 1-15. doi: https://doi.org/10.29173/iq1116 three python libraries were utilized for this task: requests, xml.etree, and pandas. the requests library, a web-scraping tool, facilitated http requests and handled responses, allowing us to query online apis and retrieve xml metadata information. xml.etree was employed to parse xml elements and prepare the data for extraction. pandas provided convenient data structures and functions for manipulating and analyzing structured data. once the xml file was parsed through xml.etree, the metadata elements were passed to pandas for the results, which enabled us to store the records as a csv file. this process resulted in data from 29 different repositories, encompassing 259,606 ddi-codebook entries. the data were collected in july/august 2024. during the research we observed significant variability over time, as new repositories were added, entries were updated, repositories were temporarily not available or changed their publication strategy and did not provide oai-pmh access any more. new repositories can be suggested via https://www.re3data.org/suggest. https://doi.org/10.29173/iq1116 7/15 wenzig, knut & han, xiaoyao (2024) state of the ddi cloud, iassist quarterly 48(4), pp. 1-15. doi: https://doi.org/10.29173/iq1116 result: how ddi-codebook 2.5 is used? from the 252 different ddi-codebook 2.5 elements 119 elements were used in our sample. ddicodebook also imports other schemas like dublin core, xhtml, or xml. but this feature is rarely used, and only one element from dublin core (<coverage>) is found3 in 64 records we analyzed. within the ddi-codebook payload, every ddi-codebook 2.5 compliant xml file has the element <codebook> on level 1 as shown in the introduction. this element is mandatory and non-repeatable. on level 2 there are the elements <docdscr>, <stdydscr>, <filedscr>, <datadscr>, and <othermat>: <codebook> is used in every record once. nearly all records use <docdscr>, only some use it twice. <studydscr> is used once within all records. <filedscr> appears in 20.9% of the records. over 90% of table 3: number of records, ddi related prefixes and oai-pmh-endpoint (linked to identification) of repositories found on re3data.org, which use ddi as a metadata standard and provide an oai-omh-endpoint to this metadata repository records ddi related prefixes oai-pmh-endpoint (link to identification) harvard dataverse 92,267 oai_ddi https://dataverse.harvard.edu/oai open forest data 79,762 oai_ddi https://dataverse.openforestdata.pl/oai cessda data catalogue 30,289 oai_ddi25 https://datacatalogue.cessda.eu/oai-pmh/v0/oai gesis data archive 20,689 oai_ddi25, oai_ddi25-de, oai_ddi25-en, (oai_ddi32) https://dbkapps.gesis.org/dbkoai/ easy 9,804 oai_ddi25_en, oai_ddi25_nl https://easy.dans.knaw.nl/oai dataversenl 7,446 oai_ddi https://dataverse.nl/oai finnish social science data archive 3,934 oai_ddi25, ddi_c https://services.fsd.tuni.fi/v0/oai swedish national data service 2,943 ddi25, (ddi33) https://snd.se/oai-pmh texas data repository 2,139 oai_ddi https://dataverse.tdl.org/oai darus 1,578 oai_ddi https://darus.uni-stuttgart.de/oai austrian social science data archive 1,540 oai_ddi https://data.aussda.at/oai repository for open data 1,458 oai_ddi https://repod.icm.edu.pl/oai swissubase 979 oai_ddi25 https://www.swissubase.ch/oai-pmh/v1/oai cora. repositori de dades de recerca 902 oai_ddi https://dataverse.csuc.cat/oai social science japan data archive 867 oai_ddi25 https://ssjda.iss.u-tokyo.ac.jp/direct/oai2/ heidata 586 oai_ddi https://heidata.uni-heidelberg.de/oai icrisat 446 oai_ddi https://dataverse.icrisat.org/oai social data repository (rds) 436 oai_ddi https://rds.icm.edu.pl/oai macromolecular xtallography raw data repository 433 oai_ddi https://mxrdr.icm.edu.pl/oai ucla 314 oai_ddi https://dataverse.ucla.edu/oai libradata 243 oai_ddi https://dataverse.lib.virginia.edu/oai trolling 171 oai_ddi https://dataverse.no/oai repositório de dados de pesquisa unifesp 76 oai_ddi https://repositoriodedados.unifesp.br/oai asu library research data repository 74 oai_ddi https://dataverse.asu.edu/oai debreceni egyetem adattár 59 oai_ddi https://adattar.unideb.hu/oai university of warsaw research data repository 59 oai_ddi https://danebadawcze.uw.edu.pl/oai unb libraries dataverse research data repository 56 oai_ddi https://dataverse.lib.unb.ca/oai osnadata 34 oai_ddi https://osnadata.ub.uni-osnabrueck.de/oai repositorio de datos académicos universidad nacional de rosario 22 oai_ddi https://dataverse.unr.edu.ar/oai total 259,606 https://doi.org/10.29173/iq1116 https://dataverse.harvard.edu/oai?verb=identify https://dataverse.openforestdata.pl/oai?verb=identify https://datacatalogue.cessda.eu/oai-pmh/v0/oai?verb=identify https://dbkapps.gesis.org/dbkoai/?verb=identify https://easy.dans.knaw.nl/oai?verb=identify https://dataverse.nl/oai?verb=identify https://services.fsd.tuni.fi/v0/oai?verb=identify https://snd.se/oai-pmh?verb=identify https://dataverse.tdl.org/oai?verb=identify https://darus.uni-stuttgart.de/oai?verb=identify https://data.aussda.at/oai?verb=identify https://repod.icm.edu.pl/oai?verb=identify https://www.swissubase.ch/oai-pmh/v1/oai?verb=identify https://dataverse.csuc.cat/oai?verb=identify https://ssjda.iss.u-tokyo.ac.jp/direct/oai2/?verb=identify https://heidata.uni-heidelberg.de/oai?verb=identify https://dataverse.icrisat.org/oai?verb=identify https://rds.icm.edu.pl/oai?verb=identify https://mxrdr.icm.edu.pl/oai?verb=identify https://dataverse.ucla.edu/oai?verb=identify https://dataverse.lib.virginia.edu/oai?verb=identify https://dataverse.no/oai?verb=identify https://repositoriodedados.unifesp.br/oai?verb=identify https://dataverse.asu.edu/oai?verb=identify https://adattar.unideb.hu/oai?verb=identify https://danebadawcze.uw.edu.pl/oai?verb=identify https://dataverse.lib.unb.ca/oai?verb=identify https://osnadata.ub.uni-osnabrueck.de/oai?verb=identify https://dataverse.unr.edu.ar/oai?verb=identify 8/15 wenzig, knut & han, xiaoyao (2024) state of the ddi cloud, iassist quarterly 48(4), pp. 1-15. doi: https://doi.org/10.29173/iq1116 these records contain it only once, while the remainder include it up to 8 times. <datadscr> is only used in 0.7% of all records. finally, 47.9% of the records use <othermat> (an element which can contain itself and therefore can exist in multiple locations), 50% of these records use it once (median equals 1), 90% use it up to 16 times (90th percentile equals 16), the highest usage per record is 64,491 (see table 4). these results and information on the usage of all level 3 elements in ddi-codebook 2.5 can also be found in table 4: like <othermat> also the elements <citation> and <notes> can be used on different locations within the tree structure of a ddi-codebook xml. as we did not account for the location in our analysis, we reported it for the first possible occurrence and then referenced this line. within the element <filedscr>, information like file name, file type, fingerprint, dimensions, or a file description can be stored. only 5 repositories provide information at this level. therefore, only 20.9% of the analyzed records use this element (and the elements within). only one repository uses this element two or more times, for example, to document a collection of multiple files. within <datadscr>, only the element <var> is used (if one disregards <notes>). it is used only in 0.7% of the records. if used, the mean use is 116.3 times, the median use is 119, the 90th percentile of usage numbers is 337.3, the maximal usage is 3,188. only one repository deploys this feature of ddicodebook and delivers information not only on dataset but on variable level. the rare use of <filedscr> and <datadscr> is surprising: information on files should be easily available for the repositories and it would make sense to inform the data users in advance about the file dimensions they can expect. if the data files are available as stata, spss, or sas files, the metadata for the area within <datadscr> would be very inexpensive to extract, because most of the work has been done during the data curation process. a complete list of all found elements and their usage statistics can be found in the appendix: table 5 and the complete dataset is also available online (wenzig and han, 2024). the dataset includes one repository that used the element <conops>, which is not part of ddi-codebook. the same dataverse driven repository also used in one record <conops>, which is part of the standard. https://doi.org/10.29173/iq1116 9/15 wenzig, knut & han, xiaoyao (2024) state of the ddi cloud, iassist quarterly 48(4), pp. 1-15. doi: https://doi.org/10.29173/iq1116 recommendations 1. repositories: the repositories should enrich the published metadata with information from the datasets by supplementing the metadata with information that is typically already encoded in stata, spss, or sas files, e.g., variable labels and value labels. figure 2 shows a possible re-use of that information in a ddi-codebook file. 2. dataverse developers: none of the records containing the element <filedscr> have been published by a repository that reports using dataverse software. we recommend that dataverse should expose this preexisting information about files via oai-pmh. 3. re3data.org: when we tried to access the different apis using the given link found on re3data.org, we had to learn that the quality of data is often poor. instead of an endpoint or a website with detailed information about the repository’s api, there is too often only a link of the generic documentation of the software’s apis and no server specific information. re3data.org should consider ensuring that only table 4: usage statistics (usage in records, mean/median/90th percentile/max of use per record, number of repositories with element found) for all elements in ddi-codebook up to level 3 element line level 1 level 2 level 3 usage in records mean median 90th percentile max used in repos 0 codebook 100.0% 1.0 1 1.0 1 29 1 docdscr 99.7% 1.0 1 1.0 2 28 1.1 citation 100.0% 3.1 2 3.0 421 29 1.2 guide 0.0% 1.3 docstatus 0.0% 1.4 docsrc 0.0% 1.5 controlledvocabused 0.0% 1.6 notes 78.4% 14.4 2 9.0 64,491 24 2 stdydscr 100.0% 1.0 1 1.0 1 29 2.1 citation see line 1.1 2.2 studyauthorization 0.0% 2.3 stdyinfo 100.0% 1.0 1 1.0 2 29 2.4 studydevelopment 0.0% 2.5 method 95.8% 1.0 1 1.0 2 27 2.6 dataaccs 99.9% 1.0 1 1.0 2 29 2.7 othrstdymat 91.7% 1.0 1 1.0 2 26 2.8 notes see line 1.6 3 filedscr 20.9% 1.0 1 1.0 8 5 3.1 filetxt 11.6% 1.2 1 2.0 192 5 3.2 locmap 0.0% 3.3 notes see line 1.6 4 datadscr 0.7% 1.0 1 1.0 1 1 4.1 vargrp 0.0% 4.2 ncubegrp 0.0% 4.3 var 0.7% 163.3 119 337.3 3,188 1 4.4 ncube 0.0% 4.5 notes see line 1.6 5 othermat 47.9% 22.1 1 16.0 64,491 22 5.1 labl 48.6% 39.6 1 20.0 64,491 23 5.2 txt 47.4% 22.2 1 16.0 64,491 22 5.3 notes see line 1.6 5.4 table 0.0% 5.5 citation see line 1.1 5.6 othermat see line 5 https://doi.org/10.29173/iq1116 10/15 wenzig, knut & han, xiaoyao (2024) state of the ddi cloud, iassist quarterly 48(4), pp. 1-15. doi: https://doi.org/10.29173/iq1116 endpoints of registered oai-pmh providers (https://www.openarchives.org/register/browsesites) would be specified. 4. ddi alliance: obviously, there are use-cases for multiple metadata prefixes related to a single standard when providing access to the metadata of the repository. the specification of the oai-pmh protocol does not allow to qualify the multiple usage of a single standard, apart from encoding the use-cases in the name of the metadata prefix. the specification document also states: “communities should adopt guidelines for sharing metadataprefixes, metadata schema and xml namespace uris of metadata formats.” (open archives initiative 2015, section 3.4) the ddi community should consider providing guidance on which metadata prefixes should be used and what should be done, if one standard is used by more than one prefix. we recommend that the ddi alliance starts a discussion about the use and structure of metadata prefixes. 5. repositories: while trying to access the oai-pmh server, we encountered several issues with data providers. first, repositories should try to improve server stability. occasionally, a repository can respond very slowly or even disconnect, while at other times it works fine. second, the metadata quality is not consistent, and it may not always be pre-checked by the repositories. as a results data collection may fail due to small errors in the xml. limitations • in ddi-codebook some elements are allowed on multiple locations. for example, the element <notes> can be used within all five elements on level 2 and 16 other locations. the element <othermat> even can contain itself, theoretically infinite number of times. while the location <codebook> ... <filedscr> <filetxt> <filename>filename</filename> </filetxt> </filedscr> <datadscr> <var name="variablename" files="filename"> <labl>variablelabel</labl> <catgry> <catvalu>value</catvalu> <labl>valuelabel</labl> </catgry> ... </var> ... </datadscr> ... </codebook> figure 2: additional codebook information, that can easily be extracted from stata, spss oder sas files. https://doi.org/10.29173/iq1116 https://www.openarchives.org/register/browsesites 11/15 wenzig, knut & han, xiaoyao (2024) state of the ddi cloud, iassist quarterly 48(4), pp. 1-15. doi: https://doi.org/10.29173/iq1116 of usage may be of interest, in this first approach we only counted the usage numbers of the elements independently of their location. • while all elements in ddi-codebook support a basic set of attributes (id, xml:lang, source, elementversion, elementversiondate, ddilifecycleurn, ddicodebookurnno), the element <var> supports more than 30 attributes. the analysis of attribute usage has been out of the scope of this analysis but may be of interest in future research. • we did not perform any content analysis or quality checks. data professionals who harvest ddi metadata to provide it aggregated in catalogues, often report inconsistent usage of the different elements. however, the usage of elements we miss in nearly all records (e.g., those that describe variables in datasets) should be more straightforward. summary we analyzed 259,606 metadata records in ddi-codebook format from 29 different repositories. while all records contained information at the study level, and almost half of the records described other material, only 5 repositories (20.9% of the records) provide information on files and only one repository (0.7% of the records) provided detailed information at the variables level. while we did not collect information about where in the schema elements are used (if allowed on multiple locations) and did not analyze the usage of attributes, insights on the use of the standard might be valuable for developers and users of the standard. there is a lack of availability of fine-grained metadata, however: “the sine qua non to greater automation of cross-domain data combination and analysis and fine-grained and responsive access control is sufficiently detailed, standardized, and interoperable metadata. there are no short cuts: data and metadata are hard.” (hodson/gregory 2023, p. 12) the high correlation between the usage of dataverse as a repository software and providing ddi metadata draws attention to dataverse, because the efforts to include options to edit variable metadata in dataverse (lubitch 2023) will be relevant for the community. https://doi.org/10.29173/iq1116 12/15 wenzig, knut & han, xiaoyao (2024) state of the ddi cloud, iassist quarterly 48(4), pp. 1-15. doi: https://doi.org/10.29173/iq1116 acknowledgements we gratefully acknowledge the anonymous reviewers and tom hartl for their invaluable feedback and insightful comments, which significantly enhanced the quality of this manuscript. any remaining errors are solely our responsibility. references all links checked on august 10, 2024. cottage labs (2021). sword 3.0 specification. https://swordapp.github.io/swordv3/swordv3.html ddi alliance (2014). ddi-codebook 2.5. https://ddialliance.org/specification/ddi-codebook/2.5/ ddi alliance (2020). ddi lifecycle 3.3. https://ddialliance.org/specification/ddi-lifecycle/3.3/ fielding, r.t. (2000). architectural styles and the design of network-based software architectures – dissertation. https://www.ics.uci.edu/~fielding/pubs/dissertation/fielding_dissertation_2up.pdf hodson, s., & gregory, a. (2023). worldfair project (d1.3) first policy brief (version 1). https://doi.org/10.5281/zenodo.7853170 lubitch, v. (2023). enhancing ddi support in the open source dataverse repository software [presentation]. https://doi.org/10.5281/zenodo.10636123 mccrae j.p. (2024). lod cloud diagram. https://lod-cloud.net/ open archives initiative (2015). the open archives initiative protocol for metadata harvesting, protocol version 2.0 of 2002-06-14. https://www.openarchives.org/oai/openarchivesprotocol.html wenzig, k. (2017). next time try recycling what reusable metadata (should) look like. https://doi.org/10.5281/zenodo.1084106 wenzig, k. and han, x. (2024). state of the ddi cloud additional datasets (information on 259606 ddi-codebook records). https://doi.org/10.5281/zenodo.13255674. wilkinson, m., dumontier, m., aalbersberg, i. et al. (2016). the fair guiding principles for scientific data management and stewardship. scientific data 3, 160018. https://doi.org/10.1038/sdata.2016.18 https://doi.org/10.29173/iq1116 https://swordapp.github.io/swordv3/swordv3.html https://ddialliance.org/specification/ddi-codebook/2.5/ https://doi.org/10.5281/zenodo.7853170 https://doi.org/10.5281/zenodo.10636123 https://lod-cloud.net/ https://www.openarchives.org/oai/openarchivesprotocol.html https://doi.org/10.5281/zenodo.1084106 https://doi.org/10.5281/zenodo.13255674 https://doi.org/10.1038/sdata.2016.18 13/15 wenzig, knut & han, xiaoyao (2024) state of the ddi cloud, iassist quarterly 48(4), pp. 1-15. doi: https://doi.org/10.29173/iq1116 appendix i table 5: list of all found ddi-codebook elements. also available as csv file in wenzig and han (2024). element usage in records mean median 90th percentile max used in repos abstract 100.0% 1.2 1 2.0 13 29 accsplac 0.4% 1.8 2 2.0 2 6 actmin 0.0% 1.0 1 1.0 1 5 alttitl 0.7% 1.2 1 2.0 2 17 anlyinfo 73.6% 1.0 1 1.0 2 23 anlyunit 16.9% 1.6 1 3.0 24 16 authenty 100.0% 2.1 1 4.0 338 29 avlstatus 6.0% 1.1 1 1.0 2 7 biblcit 73.6% 1.2 1 2.0 379 23 caseqnty 5.1% 1.0 1 1.0 1 1 catgry 0.6% 886.0 641 1739.0 13567 1 citation 100.0% 3.1 2 3.0 421 29 citreq 4.4% 1.5 2 2.0 2 15 cleanops 0.0% 1.0 1 1.0 1 9 codebook 100.0% 1.0 1 1.0 1 29 colldate 56.8% 2.1 2 2.0 162 25 collectortraining 0.0% 1.0 1 1.0 1 7 collmode 21.6% 1.7 1 2.0 74 17 collsitu 0.0% 1.0 1 1.0 1 9 collsize 0.1% 1.0 1 1.0 1 6 complete 0.0% 1.0 1 1.0 1 3 concept 14.7% 6.0 5 10.0 86 5 conditions 8.5% 1.0 1 1.0 2 11 confdec 0.1% 1.0 1 1.0 1 10 conops 0.0% 1.0 1 1.0 1 4 contact 73.2% 1.0 1 1.0 12 23 copyright 12.2% 1.7 1 4.0 5 4 dataaccs 99.9% 1.0 1 1.0 2 29 dataappr 0.0% 1.0 1 1.0 1 1 datacoll 95.8% 1.0 1 1.0 2 27 datacollector 5.5% 1.0 1 1.0 1 11 datadscr 0.7% 1.0 1 1.0 1 1 datakind 28.5% 1.2 1 2.0 35 25 datasrc 1.2% 1.5 1 1.0 67 16 depdate 69.1% 1.0 1 1.0 1 22 depositr 64.5% 1.0 1 1.0 2 23 deposreq 3.6% 1.6 2 2.0 2 7 deviat 0.0% 1.0 1 1.0 1 5 dimensns 5.1% 1.0 1 1.0 1 1 disclaimer 0.9% 1.0 1 1.0 1 6 distdate 92.5% 1.9 2 2.0 419 28 distrbtr 96.2% 2.3 2 3.0 9 28 diststmt 100.0% 2.1 2 2.0 418 29 docdscr 99.7% 1.0 1 1.0 2 28 eastbl 41.3% 1.0 1 1.0 18 12 estsmperr 0.0% 1.0 1 1.0 1 3 extlink 12.1% 1.2 1 1.0 55 24 filedscr 20.9% 1.0 1 1.0 8 5 filename 4.2% 1.4 1 2.0 192 3 filetxt 11.6% 1.2 1 2.0 192 5 filetype 5.1% 1.0 1 1.0 1 1 frequenc 0.1% 1.0 1 1.0 1 10 fundag 4.4% 1.6 1 2.0 23 3 geobndbox 41.2% 1.0 1 1.0 19 13 geogcover 55.4% 2.1 2 2.0 638 22 geogunit 13.0% 1.0 1 1.0 28 15 grantno 6.3% 1.3 1 2.0 19 21 holdings 59.8% 2.4 1 4.0 421 22 idno 100.0% 2.5 2 4.0 423 29 keyword 84.4% 7.5 3 13.0 721 27 labl 48.6% 39.6 1 20.0 64491 23 method 95.8% 1.0 1 1.0 2 27 nation 53.3% 2.3 1 2.0 390 24 https://doi.org/10.29173/iq1116 14/15 wenzig, knut & han, xiaoyao (2024) state of the ddi cloud, iassist quarterly 48(4), pp. 1-15. doi: https://doi.org/10.29173/iq1116 element usage in records mean median 90th percentile max used in repos northbl 41.3% 1.0 1 1.0 19 13 notes 78.4% 14.4 2 9.0 64491 24 origarch 0.3% 1.0 1 1.0 1 5 othermat 47.9% 22.1 1 16.0 64491 22 othid 10.0% 2.7 2 4.0 130 16 othrefs 5.3% 1.2 1 2.0 14 12 othrstdymat 91.7% 1.0 1 1.0 2 26 partitl 12.0% 1.3 1 2.0 3 6 proddate 41.1% 1.1 1 1.0 2 22 prodplac 48.5% 1.0 1 1.0 1 20 prodstmt 97.0% 1.2 1 2.0 2 28 producer 49.9% 1.1 1 2.0 14 24 qstn 0.7% 163.3 119 337.3 3188 1 qstnlit 0.7% 220.3 133 468.0 6376 1 relmat 7.1% 6.4 4 14.0 201 18 relpubl 20.4% 2.8 1 3.0 417 26 relstdy 0.7% 2.1 1 4.0 23 18 resinstru 1.0% 1.0 1 1.0 1 9 resprate 0.4% 1.9 2 2.0 2 9 restrctn 21.3% 1.7 2 2.0 6 15 rspstmt 100.0% 1.1 1 1.0 2 29 samplesize 0.2% 1.0 1 1.0 1 7 samplesizeformula 0.0% 1.0 1 1.0 1 2 sampproc 19.9% 1.7 1 2.0 14 17 serinfo 1.7% 1.9 2 2.0 4 16 sername 3.3% 1.4 1 2.0 4 17 serstmt 3.7% 1.4 1 2.0 4 19 setavail 77.6% 1.0 1 1.0 2 23 software 1.2% 1.4 1 2.0 12 14 sources 73.1% 1.0 1 1.0 1 22 southbl 41.2% 1.0 1 1.0 18 12 specperm 0.5% 1.0 1 1.0 1 7 srcchar 0.0% 1.0 1 1.0 1 8 srcdocu 0.1% 1.0 1 1.0 1 9 srcorig 0.6% 1.0 1 1.0 1 12 stdydscr 100.0% 1.0 1 1.0 1 29 stdyinfo 100.0% 1.0 1 1.0 2 29 subject 100.0% 1.0 1 1.0 2 29 subtitl 0.9% 1.0 1 1.0 2 18 sumdscr 100.0% 1.0 1 1.0 2 29 targetsamplesize 0.2% 1.0 1 1.0 1 7 timemeth 16.3% 1.5 1 2.0 10 14 timeprd 34.8% 2.0 2 2.0 9 21 titl 100.0% 2.9 2 3.0 421 29 titlstmt 100.0% 3.0 2 3.0 421 29 topcclas 30.1% 4.8 3 10.0 128 24 txt 47.4% 22.2 1 16.0 64491 22 universe 15.4% 1.3 1 2.0 6 16 usestmt 98.7% 1.0 1 1.0 1 24 var 0.7% 163.3 119 337.3 3188 1 varqnty 5.1% 1.0 1 1.0 1 1 verresp 5.1% 1.0 1 1.0 1 1 version 91.8% 1.1 1 1.0 4 27 verstmt 95.8% 1.0 1 1.0 2 27 weight 5.3% 1.0 1 1.0 3 6 westbl 41.2% 1.0 1 1.0 19 13 https://doi.org/10.29173/iq1116 15/15 wenzig, knut & han, xiaoyao (2024) state of the ddi cloud, iassist quarterly 48(4), pp. 1-15. doi: https://doi.org/10.29173/iq1116 endnotes 1 diw berlin/soep, kwenzig@diw.de 2 diw berlin/soep, xhan@diw.de 3 example: https://easy.dans.knaw.nl/oai/?verb=getrecord&identifier=oai:easy.dans.knaw.nl:easydataset:115768&metadataprefix=oai_ddi25_en https://doi.org/10.29173/iq1116 mailto:kwenzig@diw.de mailto:xhan@diw.de https://easy.dans.knaw.nl/oai/?verb=getrecord&identifier=oai:easy.dans.knaw.nl:easy-dataset:115768&metadataprefix=oai_ddi25_en https://easy.dans.knaw.nl/oai/?verb=getrecord&identifier=oai:easy.dans.knaw.nl:easy-dataset:115768&metadataprefix=oai_ddi25_en iaiiijist newsletter, vol. c", no. i (bummer lyyb) buuk ht;view3 social science data archives: 6. "social science arcnives applications and potential and confidentiality." special issue ot american kicnard l. hol'ferbert . behavioral scientist. volume ly (marcn-april ly/'b). tdited this special issue of the ameriby richard 1. hofferbert and can behaviroal scientist is devoted jerome m. clubb. information to the developments, problems, and about availaoility of issue implications lor research and may be obtained from: uavlin instruction resulting from the publications, inc., ijb^l growth and diversification of data alondra boulevard, santa fe archives since the late 19ios. the springs, california yubyu. archives examined are multiple serinquiries from the u.k., vice organizations devoted basiturope , the middle tast, and cally to acquisition of data from africa should be sent to sage diverse sources, organization and publications ltd., st. documentation of data for use by george's house, 44 hatton garpersons other than those responsiden , london ecindkh. ble for original data collection, and dissemination of these data in ttie contents include: machine-readable form to users not physically proximate to the archive 1. "machine-readable data itself, rather than local, single production by the federal university-based services. government: access to and utility for social michael w. traugott and jerome kesearch." michael w. m. clubb consider some of the major traugott and jerome m. categories of federally produced clubb. data resources, means of access, difficulties confronted in tnier t! . "i'lne less ubvious t-uncuse, and developments that may tions of archiving survey eventually provide more effective research data." warren access to the resources of the fedi. . miller. eral government by social scientists. in his analysis of the hisi. "i'he historian and social torian's relation to social science science uata archives in data, allan g. bogue complements the united states." traugott and clubb by outlining the allan g. bogue. development of academic archives as well as efforts by public agencies. 4. "uata services in western europe: heflections on stein po;<kan reviews the situavariations in the condition in the seventeen political tions of academic instisystems of western europe keeping tution-building ." stein in mind the need to begin with an kokkan. elementary analysis of the "institutional landscape" of eacn coun"2. "instructional appiicatry. he perceives the data service tions of uata archive as a response to the intellectual resources." betty a. challenge of quantitative methods nesvold. and statistical techniques in the social sciences and to the technoiaiisist newsletter, vol. 2, no. 3 (summer 1978) logical challenge of newly research. because of their cendeveloped computer systems. hokkan traility in providing machine-readpresents a schema that graphically able files of social science data, depicts the developments of data hofferbert calls upon the major archives as well as a portrayal of archives to assume a leadership the norwegian situation and elucirole in implementing procedures to dates the strategies necessary to prevent problems involving the consystematize the storage of informafidentiality ol' data, tion . 'ihis issue of the american behawniie the three essays treated vioral scientist should stand as a above delineate the state of the major critical and evaluative art of data archives and the statement on the role of social mechanisms for access to them, the science data archives in the midremaining three pieces in the abs seventies. it is unique in its special issue present topics that contribution to the data archive run accross the boundaries of time literature because it is authored and nation. warren k. miller's mainly by users of data archives discussion of the "less obvious rather than by archivists themfunctions" of archiving survey data selves. it, in tandem with the considers tne effect of the data drexel library quarterly special archive on the sociology of the issue on data archives (reviewed in social sciences. he views the lassist newsletter 1 ; 4: :iy-iti ; , pro"invisible college" of scholars as vides a dual perspective on tne being replaced by a collectivity of role ol data archives in the inforscholars who are joined together by mation system of the social scienvirtue of shared access to archived tist. highly recommended as an bodies of data central to their important set of statements by common intellectual endeavors. active social scientists. betty a. nesvold's argument that social science instruction should luedke, james a., jr.; kovacs, include training in research methgabor j.; and fried, john b. ods and experiences with machine"numeric data bases and sysreadaole data much in the same way terns." in annual heview ot as beginning chemistry students are information science and lechtrained in a laboratory is based on nology, volume ^2, pp. the supposition that modes of liy-lttl. t.dited by martha t. . learning should be matched with williams. white plains, new modes of discovery. she surveys york: knowledge industry tne few available packages for tne publications, inc. (for the teaching of statistical methods and american society for informaoffers examples of custom construetion science), 1977. issn: tion of data based instructional oobb-i^^ul). isnb: materials. u-9 t4^ib-1 1 j . cuutn: ahlsbc. lc catalog card numthe final essay by kichard 1. ber : bfa-^b09d . price: *35.00. hofferoert confronts the problein of confidentiality and data archives. this first annual keview of hofferbert states that the poteninformation science and technology tial noncooperation that could (apist) to include a chapter result from public breaches of condevoted to numeric data bases and fidentiality could destroy social systems ranges across data bases science credibility and cripple at) iaii:ii:il' newsletter, vol. c", no. i caummer lyya; lor science, technology, the social sciences, economics, and business. ihe review addresses issues relevant to the development and use ol computerized scientific, technical, social, economic, and linanacial numeric data bases and systems empnasizing on-going rather than historical activities. stuttgart: klett-cotta, lyyy. isbn: 3-1i^-9no^u-b. price; ha chara numer sioii ; bases bases eting n yin s raphy to as types data ably ble cusse data bases tusta sltt archi graph tpp. jor c ter ic a an and , an and sess for ther data s th and by t, h -l-l-. ves at topics su istics of data base survey o d systems; system de d use ; th initial isin while this numeric d ertain se the social e is a se bases wh e nature o then cite discipl in tgls, klst: unsite, are mentio the end -i jbj . rveyed inc ijumer ic system di f numeric uumer ic velopment , e huture; s; and bib survey in ata bases o ctions focu sc iences . ction on av ich first f social sc s available e: educa uls; demogr uualabs. ned in one of this se lude : data; scusuata data markacroliogtends f all s on notailadisience data tionaphyuata par action the lasslst national associat the sigs library document not hig science locate t ot numer viders o cal data informat inter are use ion o oc of asso s hou altho hligh data hem 1 ic da f soc in ion p nation note r grou f pub acm , ciatio ndtabl ugh t t the arch n the ta and ial per spe rot ess al , d i ps lie and n ' s e t his rol 1 ves bro pla cien ctiv iona eff s such data the go uouo sur v e ol ader ces ce s e f is. orts of well as as the users , american vernment ht; (p. ey does social it does context the protatistior other paul mull of critical topic of wliicn emerge cial session ings of th association has discusse in an early newsletter 17-^1 ; and d as "all data lected lor tific resear surveys) , b ducts or tra ines of priv tions or per this importa er has essays process d from during germ in bie d proce issue (. volume efined that a statist ch le.g ut are ces of ate or sons ." nt volu ed de -pr the t an ief ssof 1 , the re/ ica in th pub th me ited vote od uc gua he 1 soc eld. prod the n gen were 1 o ce stea e da lie e co mcl a volume d to the ed data ntum spe97b meetiological muller uced data lasslst umber ^ : eric term not colr sciennsuses or d by-proily routorganizantents of ude : 1. "die basis d interak theor ie iknglis changin sociolo between icisum , scheuch trates ical re ing the the "c researc broad view of process mcreas search within wechselnd er soziol tion und h trans. g data gy: int theory a ") by which on the me asons fo data ba lassicial h and theoretic the deve es leadi ed use o el icte social re e d ogiez.w1s timpir base eract nd tm hrwin con thodo r exp se be " su gives al o lopme ng to f non d searc aten zur chen ie ." "the of ions pirk. cenlogandyond rvey a verntal an -redata h mullt", paul j., ed . die analyse prozeb-produzier ter daten tenlgish trans.: the analysis of process-produced dataj. "die buchfuhrung der verwaltungen als sozialwissenschaftliche datenbasis." ^knglish trans.: "administrative bookkeeping as a social science data base,") by wolfgang bick and paul j. muller. la^iiiist newsletter, vol. ^, no. i (summer 197b) presented in a snortened form at the yuantum/iishauonterence, "wuantitication and methods in social science research: possibilities and problems with the use of historical and process-produced data," university of cologne, august lu-li', 1977the essay focuses upon the need to study the representational nature of administrative bookkeeping, ana to work out the kinds of approaches that seem most promising with the use of these kinas of data. :5. "vjrenzen und mogiichkeiten der verwendung von strafakten als urundlage kr iminologischer forschung." (english trans.: "possibilities and problems with the use of punishment pecords as a data base for criminology,") by ul wiebke steffen. examines the usefulness of thse data not for the analysis of clients' behavior, but for propositions about the data generating organizations . 4. "verknupfung und generierung von mikrodaten." cenblish trans.: "linkage and generating of microdata,") dy klaus kortmann and hans-jurgen krupp. reports on the record linkage procedures employed in linking up the various west german censuses . 5. "prozeti-produzier te uaten in der rechtssozologie . " (englisli trans.: process-produced data within the sociology of the law,") by volkmar gessner, barbara rhode, gerhard strate and klaus a. ziergert describes the project design for a multi-level and multi-file analysis of the insolvency situation of various business and concentrates on the dillerent "images" insolvency has in various record-keeping systems . b. "datenverarbeitung als guellenkritik.'" by fcrdmann weyrauch reports on the use of sampling techniques for the analysis ol' medieval tax lists. 7. "mobilitat und soziale der wurttembergischen fabrikarbeiterschaft im 19. jahrhunder t ," by peter borscheid and heilwig schomerus discusses the analyses of various 19th century records, especially the potentialities of records kept at the end or beginning of a marriage (teilungen or inventuren) to construct quantitative life histories of earning, life style and consumption, oriented to the expectation of income (a la m. hriedraan's "permanent income hypotheses.") this volume is a substantial contribution to an expanding comprehension of what constitutes relevant material for research. data archivists should obtain this item and pass it on to their clientele. highly recommended. »/ laijsiot newsletter, vol. 2, wo. j tsummer l^/ttj white, howard u. ed . deader in "archive-library helations," a ilachine-headadle ijocial data. critical issue for information proenglewood , colorado: informafessionals, is the third section tion handling services, 1977. which considers a spectrum of laddress: information hanorganizational structures to encomdling services, library and pass the overlapping missions of t.ducation division, box 1154, tnese two institutions. intellecb.nglewooa , lolorado dulluk tually archives are a coherent comliorary ol congress catalog ponent in the social science inforcard number: i ( -id'i i2 . isbn: mation model, but practically u-91 uyy^'-yo-^' . price: ^ly.uu neither archives nor libraries have been able to integrate their serine treatment of social data as vices for maximum benefit to the a topic for the headers in librariuser, anship and information science series is formal recognition by the bibliographic control is exainformation community that the mined in section four, "indexing field of data archiving is a major and cataloging social science component ot the social science uata." various strategies of docuinformation system. howard d. mentation as well as actual pro»jhite, who edited the drexel gress made in the cataloging of library quarterly issue on machine-readable records are machine-headable social science treated by leading planners, libradata tsee lasslst newsletter: rians, and archivists. 1;'^:i7-jtt) and who did his doctoral thesis on library/data archive the final section, "the world of relations (see lasslst newsletter the data specialist," examines the 1;3:'^9-il), tias compiled an antholmanagement ot data archives as the ogy ot major formative articles on frontier ot librar ianship. 'i'he the history, rationale, and strucspecial skills and wide range of ture of data archives. technological and subject expertise needed to organize and run archives the header is divided into five on a daily basis are articulated sections. the first, "numerical and discussed. data in the social sciences," introduces key problems of social the header is a basic sourcebook science data in machine-readable for the practicing and beginning torm. articles locus on the potendata archivist. the former will tial of archives for research and find codifaction of practice and archive under-util ization due to some of the philosophical bases on lack of users' knowledge about winch the field is built. the lattheir holdings. the second secter will obtain a preliminary surtion, "major data suppliers," provey of the range of skills and vides descriptions of major files resources needed for effective data such as the u.s. census data of archiving, machine-readable records in the u.s. national archives; and cenwhite has pulled together writters such as the inter-university ings by prominent data archivists consortium for political and social and social scientists which provide research, the hoper public opinion a multifaceted approach to this center, the national upinion emerging tield. working archivists research center, and the project will want the book as a reference talent data bank. and intellectual rationale tor ad laii^iar newsletter, vol. 2, no. i (summer 19/b) their daily work; students of lidrary and information science will turn to the feader as the tirst collection ol widely scattered materials to be made conveniently available; and social scientists will find the header a coherent introduction to the expanding array of resourcesand services to facilitate extended analysis. the reader is highly recommended to all three groups. tt9 iassist quarterly 63 ensuring racial representation on jury panels: an empirical and simulation analysis hiroshi fukurai' university of california, riverside edgar w. butler' university of california, riverside jo-ellan huebner-dimitrius^ california state university, long beach 'paper prepared for presentation at the international association for social science information service and technology (iassist) conference held in marina del rey, california, may 23, 1986. california law specifically states that persons listed for service in the court: shall be fairly representative of the population in the area served by the cotirt and shall be selected upon a random basis (section 9. 203). introduction geographic groupings which overlap with racial and economic groupings constitute recognizable classes.^ rural residents, for example, might be underrepresented due to excuses based on distance to the courthouse; often selection officials acquiesce in the reluctance of rural residents to serve. excuses based on willingness to travel great distances have so reduced the jur}' pool that remedial action is required even without proof of geographic cohesiveness.^ the notion of vicinage or geographical locality requirement of jury selection has been traced at least as far back as to charlemagne (charles the great) in 768 a.d. he instituted several reforms, one of which was the establishment of "inquisito". one of the requirements of the inquisito was that 13 to 66 witnesses be chosen from the neighborhood where they woiild have knowledge of the matter in dispute (moore ^see thiel v. southern pacific co., 328 u.s. 217 (1946), state v. holstrom, 43 wis. 465, 168 n.w. 2d 574 (1969). and state v. cage, 337 so. 2d 1123 (la. 1976). the federal act requires that selection procedures "ensure that eacn coimty', parish or similar political subdivision within tne district or division is substantially proportionally represented in the master jury wheel for that judicial district division, or combination of divisions" (u.s. 1968, section 1863 (b) (3)). ^ see united states v. fernandez, 480 f. 2d 726, 732-33 (2d cir. 1973). fall/winter j 987 64 iassist quarterly 1973). the requirement of a jury from the vicinage, found in the magna carta as well as the sixth amendment, is also based on the notion that jurors should be selected from local residents. the meaning of this requirement is often unclear, as when a case from one division in a federal district court is tried in another division, or when grand jurors are selected from only one division.' ciurently, federal law determines the nature of prospective jurors by specifying two key concepts in jury venire or panel selection procedures: (1) "a random" selection of jurors, and (2) the inclusion of special geographic districts wherein a panicular court convenes, i.e., vicinage requirements (u.s. 1968, section 1861). ln california, as in most states, the law similarly requires: (1) a random selection of jurors and (2) selection from "judicial districts of the respective counties" (ca. 1981, section 197, 206). recent federal and california supreme court decisions are such that any substantial violation of these basic requirements of jur>' selection in representativeness is a prima facie case of discrimination.' subsequently, an increasing number of challenges concerning the underrepresentation of "cognizable groups," e.g., minorities, have been brought claiming violation of the sixth amendment, a representative jur\' selected from a fair cross section of the communiri'.' ' see u.s. 1968, section 1861 and house report at 1801. ^ ' see lis, 90th congress senate report no. 891 19"e7rij^90th congress house report no. tu76 1968;^t¥etaie law joufnart970: kaifvt wj2: de cam iw^, chevigny 1975; alker ' hosticka, and michel! 1976; kairys, kadane, and lehoczky 1977; alker and barnard 1978' heyns 1979; butler 1980a, 1980b, and 1981; butler and fukurai 1984; fukurai and buder 1985. ' for example in california, see people v white 43 cal. 3d 740 1954; people v. king 49 cal. rptr. 562 1966; people v. sirhan 7 c^. 3d 258 1978; people v. wheeler 148 cal. rptr. 890 1978; people v. estrada 155 cal. one of the major problems in jiuy challenges is to establish a prima facie case for the imderrepresentation of minorities. part of the problem is due to the ambiguous relationship between random selection and vicinage requirement (area or district). past supreme coim cases have dealt with the systematic imderrepresnetation of cognizable groups, e.g., blacks and hispanics; however, the coim has not addressed the extent to which the area served by the court relates to the random selection of potential jurors. vicinage or geographic representativeness has rather been dealt with, along with the random selection of jurors, without geographic representativeness being clearly demarcated. in duren v. missouri, for example, the u.s. supreme court held that a three-prong test must be applied to establish a prima facie case of discrimination: (1) the group alleged to be excluded is a 'distinctive' group in the community, (2) the representation of this group in venires from which jurors are selected is not fair and reasonable in relation to the number of such persons in the communit>', and (3) this underrepresentation is due to systematic exclusion of the group in the jun'-selection process (duren v. missouri 439 u.s. 357 364 1978). however, the cotm did not spell out a clear cut relationship between juror representativeness and the vicinage requirement, e.g., what is the "commtmity"? "(cont'd) rptr. 731 1979; people v. graham 160 cal. rptr. 10 1979; people v. harris 36 cal. 3d 36. 201 cal. fpn. ih 679 r 2d 433 1984. in federal supreme court see alexander v. louisiana 405 u.s. 625 1972; peters v. iciff 407 u.s. 493 1972; tavlor v. louisiana 419 u.s. 522 1975; duren v. missouri 439 u.s. 357 1979; city of mobile, ala v. bolden 466 u.s. 55 1980. fall/winler 1987 iassisl quarterly 65 vicinage requirement the vicinage or geographical requirement of jur>' trials is an essential element of the sixth amendment as it pertains to the jury selection process. for illustrative purposes, we will use los angeles count}' and its twent>' mile radius rule. however, the process itself is generalizable to all areas of los angeles count)' and any other bounded area such as a county, judicial district, etc. the first provision for a jury trial in a vicinage can be found in article ei of the constitution. article iii. section 2 notes: the trial of all crimes, except in cases of impeachment, shall be by jury; and such trial shall be held in the state where the said crimes shall have been committed; but when not committed within any state, the trial shall be at such place or places as the congress ma\by law have directed. early in the 1970s, the los angeles count}' board of supervisors developed a policy that no juror had to travel more than twent}' miles from his/her house to the courthouse. the count}' board adopted this rule because of convenience for prospective jurors and economic reasons for the county. subsequently, in california, the legislature defined the judicial district in los angeles count}as being within a twent}' mile radius from each courthouse. the california code of civil procedure states that: each court shall adopt rules supplementar}' to such rules as may be adopted by the judicial council, governing the selection of persons to be listed as available for service as trial jurors. the persons so listed shall be fairiy representative of the population in the area served bv the court, and shall be selected upon a random basis. such rules shall govern the duties of the court and its attaches in the production and use of the juror lists. in counties with more than one coun location, the rules shall reasonably minimize the the distance traveled by jurors. in addition, in the count}' of los angeles no juror shall be required to serve at a distance greater than 20 miles from his or her residence (ca. 1981, section 7. 203). despite the explicit rule of random selection of potential jurors from the judicial district defined within 20 mile radius, recent jur}' venire challenge cases have argued the following two points: (1) there is a significant underrepresentation of prospective minority jurors and (2) there is an overrepresentation of particular neighborhoods with high concentrations of anglos (hevns 1979; butler 1980a, 1980b, 1981; butler and fukurai 1984; huebner-dimitrius 1984; fukurai and butler 1985; fukurai and butler 1986). these studies have shown that census tracts with a high anglo concentration are consistently overrepresented and consequently jur}' venires have consisted of a large number of potential anglo jtirors, and an underrepresentation of minorities. this is apparent for all superior court districts in los angeles coimt}' except the central district (heyns 1979). it is theoretically possible to have race/ethnic representation on juries, yet not have a fair cross section of the communit}' or areas served by the court from which jurors are being drawn to serve on juries (heyns 1979; huebner-dimitrius 1984; fukurai 1985; fukirrai and butler 1985). generally however, racial and geographic representativeness are highh" correlated; therefore, it is possible to ensure the cross-section representation of minorities by controlling the random selection of geographic areas. currently, the ovenepresentation of particular neighborhoods contributes to a substantially greater chance of anglos serving on fall/winter 1987 66 iassist quarterly juries, while a random selection of neighborhoods would ensure the fair representation of minorities as prospective jurors within a districl in this paper, we present an analytic strategy that will overcome racially disproportionate jury venires. rather than first focusing on the selection of individual potential jurors, random selection of neighborhoods is examined, i.e., census tracts from which prospective jurors are being drawn to serve on juries. our analysis demonstrates the extent to which neighborhood representativeness could rectify the disproportionate underrepresentation of minorities currentiy the case in most jur}' venires. the main thrust of this paper, thus, is threefold: (1) to propose a geographc sampling strategy to overcome imderrepresentativeness of minorities, (2) to illustrate our strategy using simulation techniques, and (3) to show the extent to which geographical randomness can help ensure that racially proportionate jury venires are obtained. by simulating the los angeles cotmt\' selection process, a comparison between the actual jury composition and the simiilated jur,composition is examined to show the extent to which the proposed geographic samphng strategy is superior to the current selection procedures employed in los angeles cotmty and elsewhere. data two data sets were linked to serve as the foundation for the simulation of the jury selection process: (1) 1980 u.s. census bureau data and (2) jury impanelment hsts for a retrial of the hams case (36 cal. 3d 36 201 cal. rptr. 782 679 p. 2d 433 1984).' eight jiu7 impanelment lists were obtained to delineate neighborhoods (census tracts) from which jurors were being drawn to the long beach superior court and to determine whether or not the panels represented a fair cross section. the impanelment period imder investigation, while not ideal, was lengthy enough to determine whether or not jury venires were representative of the commimity population. these eight panels were typical of panel data available for other time periods, including the first harris trial 1979. empirical analysis figure 1 depicts racial composition of the long beach judicial district using a variety of definitions of the area served by the court map a illustrates the areas served by the superior court as presented in the harris retrial. (ed.note. figures and maps have collected together at end of article) this variation in the definitions of the area served by the court shows that there is a potential for either conscious or imconscious manipulation of minority representation on jury panels. thus the particular "area served by the court" becomes important in jury challenges. that is, if it is to be determined whether or not jurors represent a fair cross section of the community, the area served by the court must be clearly delineated or a valid comparison cannot be made. six different areas served by the court emerged during the harris retrial. the first was los angeles coimry as a whole. in the first hams empirical analyses of people v. harris (36 cal. 5d 36, 201 cal. rptr. 782 679 ?. 2d 433 1984) were performed at university of '(cont'd) california. riverside. in people v. harris, the motion of respondent for leave to proceed in forma pauperis was granted; however, the writ of certiorari by the prosecution to the federal suprerne court was denied on oct 29, 1984. fall/winter 1987 iassist quarterly 67 trial tlie prosecution argued that los angeles county-wide data were the proper comparison. the harris opinion rendered by the california supreme court concluded that the parties, however, presented evidence and argued this case on the assumption that all jtmes in los angeles coimty must be representative of the entire county. the principal question before us is whether evidence based on total countywide population figures, rather than jtiry-eligible population, is adequate to make out a prima facie case; for the reasons explained in this pinion, we conclude that it is. the state has not attempted to rebut this prima facie showing by arguing that the long beach juries need only represent those persons living within 20 miles of the courthouse, and has not attempted to show that such juries were truly representative of that limited area (harris 36 cal. 3d 36. 201 cal. rptr. 782 679 p. 2d 433 1984). a second area served by the court is within a 20-mile region, as delineated by california state law; that is, any juror may be excused from being sent to a particular courthouse that is further than 20 miles from his/her residence. in 1978, the 20-mile region delineated by the jur>services division in los angeles county was for the most part a 20-mile straight line from the courthouse. however, in 1983 the area served by the court was reduced to a 15-mile direct line, presumably on the basis that the driving distance would be 20 miles. thus, a third definition of the area serviced by the court was considered. a fourth area served by the court was empirically delineated. this area was determined by delineating those census tracts from which jurors were summoned for eight panels. a fifth area sen-ed by the court could not be determined geographically but is obviouslv different from the others. this fifth area is a subset of the fourth area which was geographically determined. for each juror summoned, knowledge of their census tracts and address was made available, thus the area served by the court could be empirically determined. all of the potential jurors who showed up at the long beach cotirthottse came from these impanelments and thus were a subset of the impanelment lists. however, between the impanelment or summons stage and the junvenire stage, there was between a 40-50 percent dropout thus, they are similar but not the same. unfortunately we were unable to delineate areas of residence at this stage. finally, during the course of the harris retrial, the prosecution argued that a sixth area was more important than these other five. this specific area was known as the long beach superior court district, as defined by the los angeles board of supervisors. this area is used by the legal system in allocating trials. thus, if a person commits a crime in this bounded area and it becomes necessary to have a trial, it typically, but not invariably, will be assigned to tjie long beach courthouse. however, this particiilar area is not coterminous with any of the other five areas served by the court obviously, all the areas are within los angeles count)', but otherwise they have nothing in common. table 1 shows the racial composition of the eligible hispanic population and impanelment hsts using the 15-mile radius definition of the area served by the court thus, while 20.9 percent of potential jurors at a 15-mile level were hispanic, only 10.1 percent of jurors impaneled and summoned to the long beach superior court were hispanic. underrepresentation of hispanic jurors was inevitable because of the under-selection of hispanically dominant census tracts. table 1 thus shows that more than one-half of the potential hispanic jurors were underrepresented on the impanelment list z scores and fall/winter 1987 iassist quarterly chi-square values show that hispanic jurors are statistically underrepresented on the impanelment list; thus, the hispanic composition on the impanelment lists is significantly different from the racial composition of the long beach judicial district, as defined by the 15-mile radius. the underrepresentation of both hispanic and black jurors on the eight panels under investigation was consistent table 2 shows the racial composition of minority jurors on both impanelment lists and census tracts from which potential jurors are simimoned. census tracts with a high concentration of anglos are overrepresented whereas minority dominated census tracts are underrepresented. table 3 shows the representation of census tracts on eight impanelments. potential jurors from one census tract were represented thirty-six times, while fiftv-one census tracts were represented less than five limes. table 3 also indicates that one-half the potential jitrors came from twenty-three census tracts (5.2%) of the total of 439 tracts in the long beach judicial district defined by a 15-niile radius. a high concentration of anglos. in any case, and for whatever reason, anglo dominant census tracts are clearly ovenepresented on the jury impanelment lists. one dubious explanation is that anglos are more qualified than minority groups for jury duty.* however, the proportion of qualified juiors is the same for the impaneled census tracts and the long beach superior court judicial district as a whole. table 5 shows the proportion of qualified jurors in the impanelment list and the long beach judicial district while 19.9 percent of jurors in impaneled census tracts are qualified jurors, 19.1 percent of those in the long beach judicial district are equally qualified. thus, the percentage of qualified jurors has no bearing on the underrepresentation of minority jurors. further, a random method of selecting jurors has not been exercised, i.e., impaneled census tracts are clustered in particular regions characterized by an anglo population. many minority dominant census tracts are not included in the impanelment list, even though the proportion of qualified jurors is the same in both impaneled and non-impaneled census tracts. table 4 indicates the average representation of census tracts on eight panels. tlie table shows the extent to which tract representation is related to the racial composition of the census tract for example, census tracts which were selected less than the average number of times had three and ten percent higher black and hispanic populations respectively. one-half the overrepresented census tracts had four and fourteen percent less black and hispanic population, respectively. maps 1 to 8 illustrate the census tracts from which actual potential jurors were summoned. these maps show that the census tracts were concentrated in particular regions, i.e., the lower portions which border on orange county. from previous tables, it should be obvious by now that these census tracts are characterized by ' research indicates the ovenepresentation of anglo jurors is necessary since criminality is inherent in some minorit^ groups (hepburn 1978; cullen and link 1980; turk 19§1; kramer 1982). thus, minority groups "take[s] a permissive view of crime within its dorder. as a result, the black community is vulnerable to its own criminal element as well as to the criminal element of the white communitv" (the yale law journal 1970, p.534). further^ researchers suggest that count\' clerks responsible for selecting names from master files purposely exercise systematic selection rather than random selection in creating raciallv disproportionate jury pools (alker and barnard 1978; levine and schweber-koven 1976). because so many different persons use individual discretion to decide who should be excused and who should serve, the possibility of individual prejudice influencing excuses and exemptions is great (van dyke 1977, p.391). fall/winter 1987 iassist quarterly 69 geographic random selection one means by which to rectify the disproportionate representation of census tracts is to implement the random selection of census tracts within a judicial district, but geographically defined. such random selection should provide a hst of census tracts equally distributed within the limited, spatially bounded context, i.e., 15-mile radius, or whatever. since our analysis shows that qualification of particular racial populations does not have a bearing on the selection of anglo-dominant census tracts, the random selection of tracts provides a foundation for equally selecting various racial/ethnic groups within them thus resulting in a fair cross-section of the population, vis-a-vis minorities. a simulated random selection of census tracts was carried out in the following manner. each census tract within a 20-niile radius of long beach judicial district was given a unique number. a series of random numbers were generated for the selected number of census tracts for each of eight panels. those eight individual simulations were conducted to conespond to the actual eight impanelments as previously empirically analyzed and used in the harris retrial. maps 9 through 16 illustrate the simulated mapping of census tracts randomly selected within the long beach judicial district each map shows that selected tracts are evenly distributed in space. the nimiber of potential jtirors also shows that using this process, minority groups would have an eqtial chance of selection for jiiry service. table 6 shows the racial composition of selected census tracts for each of the simulated panels. within the 2(>-mile radius 26.7 percent of potential jtirors were hispanic and 14.8 percent for black. a z scores statistical test for difi'erences in racial composition between census tracts derived by random selection and the 20-mile radius district was then carried oul not one of the scores was significant, suggesting that each of the randomly selected samples of tracts had a racial composition similar to that of the 20-mile radius judicial district this, of course, is in stark contrast to the actual impanelments analyzed in the first section of this paper. map 17 shows the mapping of all census tracts in eight panels using the simtilation method. the map shows that random selection of census tracts provides a virttial equally distributed hst of tracts from which potential jurors would have been summoned. such random selection also provides an imbiased racial representation. critique the results of our simulation clearly show that the process we have suggested is far superior to the current process in ensuring a fair cross section of jurors. one question, of course, is whether or not this process is allowable under current federal and state stames. our response is that not only is it allowable, but the results of the simulation imply that our process shotild be mandated by law. another argument that possibly could be made against the proposed process is that qualification varies by district however, our evaluation of the actual juror qualification rate for los angeles county compared with the long beach district (20-mile) showed that the qtialification rate was virttially identical. even if there had been some variation, such variation could be fined into the system. .another possible objection to the randomized geographical process is that it would increase fail/winter 1987 70 iassisl quarterly the overall mileage driven by jurors. this is true. any system that results in a fair cross section will result in more aggregate miles driven because of the very fact that the jurors would be from all areas of the district rather than concentrated in certain areas . this is a necessary' part of a system that results in a fair cross section of the community — jurors must come from all parts of the community . the proposed system does away with the idea of selecting only jurors from areas closest to the court, and in fact, requires just the opposite. that is, jtirors are drawn from all areas of the district however, the district could still fall within the state law mandated 20-mile region for los angeles county. in los angeles coimty, a particular problem that must also be dealt with is the overiapping of judicial district boundaries. this problem is amenable to statistical sampling methods. however, even if some ovenepresentation should occur, it would be substantially less than is now occurring using non-random selection of areas. finally, the analysis presented here represents only pan of a year and thus might be considered static. a dynamic jury selection process involves selecting jtirors periodically. however, jurors also are qualified only periodically. thus a dynamic system of juryqualification could use the same technique described in the simulation section. that is, the jury qualification process could also be accomplished by the random selection of census tracts, and the mailing out of questionnaires periodically throughout the year. conclusions our analysis conclusively shows that currently there is a systematic and biased selection method employed in the long beach judicial district and elsewhere in los angeles county. the racial composition of actual impaneled census tracts indicates that (1) selected census tracts are clustered in regions with high concentrations of anglos and (2) hispanic and black potential jurors are systematically weeded out in the selection process because of biased impanelment lists. one possible reason for such systematic selection of anglo dominant census tracts might be that anglo jurors in particular census nacts are more qualified than their minority counterparts. however, the proportion of qualified jurors from anglo dominant census tracts was the same as that of the judicial district as a whole. we suggested an alternative sampling strategy of random selection of census tracts which provides a representative list of tracts from which potential jurors could be summoned. our simulation analysis showed that selected census tracts could provide a list of potential jurors that would be unbiased, i.e., racially representative. that is, the impanelment lists would have a racial composition similar to the judicial district our method of randomly selecting census tracts is clearly superior to the selection method currently employed in los angeles county, because the potential jurors coming from the selected tracts are evenly distributed and have an equal chance of being selected. the random selection of census tracts, thus, is congruent with requirements established by both the federal jury selection and service act in 1968 and the california code of civil procedure in 1981.' ' federal jury selection and service ka was passed in 1968 to guarantee that "all litigants in federal courts entitled to trial by jury shall have the right to grand and petit juries selected fall/winter 1987 iassist quarterly 71 cullen, francis t. & bruce g. link. 1980. crime as an occupation. criminology 18:399-410. bibliography alker, hayward r., jr. & joseph j. barnard. 1978. procedural and social biases in the jun' selection process. the justice system journal 3:220-241. alker, haywaid r., jr., carl hosticka mitchell. 1976. jury selection as a biased social process. the law and societ\' review 9:9-41. butler, edgar w. 1980a. torrance superior court panels and population analysis. university of california. riverside. butler, edgar w. 1980b. van nuys superior court panels and population analysis: may 7, 1979 yhrough september 24, 1979. universityof california, riverside. butler, edgar w. 1981. the 1980 los angeles count>' jiu^' selection study: compton superior cotirt university of california, riverside. butler, edgar w. and hiroshi fukurai. 1984. an evaltiation of jury panel selection procedures: the north valley superior court, los angeles county. universityof california, riverside. cherigny, paul g. 1975. the attica case: a successful jiu^challenge in northern ciu-. criminal law bulletin 11:157-172. de cani, john s. 1974. statistical evidence in jury discrimination cases. the journal of criminal law and criminology 65: 234-238. fukurai, hiroshi. 1985. institutionalized racial inequaliry: a theoretical and empirical examination of the junselection process . unpublished dissertation. university of california, riverside. fukurai, hiroshi & edgar w. butler. 1986. assimilation and internal colonialism models of juryselection, [forthcoming] fukurai, hiroshi & edgar w. butler. 1986. the juryselection process: institutionalized inequalit>' . [forthcoming] hepburn, john r. 1978. race and the decision to arrest: an analysis of warrants issued. jotimal of research in crime and delinquency 15:54-73. heyns, barbara. 1979. 1979 jiir\' analysis. (superior court, county of los angeles, no. a-344097, joseph piazza, defendant) huebner-dimitrius, jo-ellan. 1984. the representative jun': fact or fallacy? unpublished dissertation. claremont graduate school . kairys, david. 1972. jtiror selection: the law, a mathematical method of analysis, and a case smdy. american criminal law review 12:771-806. '(cont'd) at random from a fair cross section of the community in the district or division wherein the court convenes" (u.s. 1968, section 1861). kairys, david, joseph b. kadane, & john p. lehoczky. 1977. jury representativeness: a mandate for multiplesotu-ce lists. california law review 65:776-827. fall/winter 1987 iassist quarterly kramer, ronald c. 1982. from 'habitual offenders' to career criminals'. law and htrnian behavior 6:273-293. levine, adeline gordon & claudine schweber-koren. 1976. jury selection in erie county': changing a sexist system. law and society reviewll:43-55. moore, lloyd e. 1973. the iury: tool of kings, palladium of liberu' . cincinnati: the w.h. anderson company. turk, austin t. 198l the meaning of criminality in south africa. international ioumal of sociology and law 9:123-135. u.s. 90th congress senate report 1967. no. 891. u.s. 90th congress house report 1968. no 1076. van dyke. jon m. 1977. jury selection procedure . massachusetts: ballinger publishing company. anon. 1970. the case for black juries. yale law ioumal 79:531-550. people v. graham . 160 cal. rptr. 10 (1979) people v. harris. 36 cal. 3d 36, 201 cal. fprt 782 679 r 2d 433 (1984) people v. fcing . 49 cal. rpti. 562 (1966) people v. sirhan. 7 cal. 3d 258 (1978) people v. wheeler, 148 cal. rptr. 890 (1978) people v. white. 43 cal. 3d 740 (1954) peters v. kiff 407 u.s. 493 (1972) state v. holstrom. 43 wis. 465, 168 n.w. 2d 574 (1969) taylor v. louisiana. 419 u.s. 522 (1975) thiel v. southern pacific co. . 328 u.s. 217 (1946) united states v. fernandez . 480 f. 2d 726, 732-33 (2d cir. 1973). cases cited alexander v. louisiana . 405 u.s. 625 (1972) citv of mobile. ala v. bolden . 466 u.s. 55 (1980) duren v. missouri 439 u.s. 357 (1979) people v. esrrarlfl 155 cal. rptr. 731 (1979) fall/winter 1987 iassist quarterly 73 map a the area served by the uwu lieacii sui'kiuoli cdtlkt tn the haiihis kf.tkiai., 19u3 lege! id i area \mm juror fall/winter 1987 74 lassist quarterly figure 1 la county 20-mile radius x m 15-mile radius m summons area panels, may 1985 la county 20-hile radius (j 15-mile radius b sunnons area panels, may 1985 long beach juror panels amd 1980 u.s. cmsus data fdr dirterem" areas servm by tie cairf: black and spanish populations 23.0 26.7 20.9 15.9 5.c b 5 10 15 20 25 30 (percent) 11.0 1^ .8 16.-1 6.4 6.2 15 (percent) fall/winter 1987 iassist quarterly _ 75 table 1 long 3each sisrric: iriibie -udius '-iscs cis?arl:y dlspari:-/ scare talue 10.:; -10.5 -51. 7 -7.5" •*/i-</o5 through 6/l2,'8s significant at <<<0.000l level. fall/winter 1987 76 iassist quarterly table 2 .-.iscacic lis: '— :— i; 9.c:: 10.;:. -i. 9 7.0* -2.5 ^ 5-1-=; ic.o 16.1 -1.7 7.7 -2.3 2 :-s-:: 13.0 12 . 5 -1.9 3. 3 -2.7 ' 5-l5-:5 11.0 11.3 -:.. 7.5 -2. 3 5 5-:--=5 12.0 li.2 -:.2 5.2 -2.0 6 11.0 13.7 -2.i 4.7 -3. 2 7 6-3-;i 7.2 13 . 2 -3.s 9.2 -2.2 3 6-12-55 8.0 12. 7 -2.7 5.2 -2. 6 toi.u. ic.i 15.9 * -7.5 * -7.6 percencages are calcuiacej on che basis of all included census tracts fall/winter 1987 iasstst quarterly 77 table 3 lokg 3£ac;-i jupixios. colat district arrii. 24, 195 3 rhs.ocgh jlt.i'e 12, 198 5 no. of tees on 3 panels no. of cinsus tsji.cts f5l£quehcy 10 8 11 9 13 3 9 3 9.34 7.47 10.28 3.41 12.15 7.47 8.41 7.47 1.36 2.30 2.30 5.60 6.54 0.93 0.93 1.36 0.93 0.93 0.93 0.93 0.93 0.93 median 6 meaa 7.5 fall/winter 1987 •yg _ iassist quarterly table 4 * 430 census craccs lacluded '.os; 3«ic.-.* 3 t:aas " 'lza ' tiz=s ; ti-ss oli:rlcc or mors ar less 3c hori ar lias 16. is 4.3; ".i; 4.6j 3.3j; 20.9 9.2 19.5 9.2 22.9 fall/winter 1987 iassist quarterly 79 map 1 long beach: april 24, 1985 legend i area 8xiij3 juror fall/winter 1987 iassisl quarterly map 2 long beach: may 1, 1985 legerid: area ki321 juror fall/winter 1987 iassist quarterly 81 map 3 long beach: may 8. 1985 legemd: area i i eebffih juror map 3 fall/winter 1987 82 iassist quarterly map 4 long beach: may 15. 1985 mmj^j'/^^. "v' legemd: area l exju juror fall/winter i9s7 iassist quarterly 83 map 5 long beach: may 22, 1985 legehds area sues juror fall/winter 1987 map 6 long beach: may 29. 1985 iassist quarterly legerid: area czz] ecehj juror fall/winter 1987 iassist quarterly 85 map 7 long beach: june 5. 1985 .ay^~ legemd: area bus juror iall/iv inter 1987 86 iassist quarterly map 8 long beach: june 12, 1985 legemd: area juror fall/winter j 987 iassist quarterly _ gy table 5 long 3s:ach superior col'rx long beach judicial dlscrict lapaneljieai: llscs ~ no. percent no. percenc local jurocs 243,274 looz 62,753 louj; qualified jurors 46,436 19.11 12,459 19. 9z 1. total number of census tracts are 439. 2. total number of census tracts are 107. * source: los angeles jury supervisor ray arce and his computer consultants, 1985 fall/winter 1987 iassist quarterly map 9 long beach superior court district legemd: area ksia jurors map 9 fall/winter 1987 iassist quarterly map 10 long beach superior court district legend: area ^htsi jurors fall/winter 1987 90 iassist quarterly map 11 long beach superior court district legend: area i i ^m jurors fall/winter j 987 iassist quarterly 91 map 12 long beach superior court district legend: area l giss3 jurors map 12 fall/winter 1987 92 iassist quarterly map 13 long beach superior court district map 13 fall/ winter 1987 iassist quarterly 93 map 14 long beach superior court district legend: area e212 jurors map 14 fall/winter 1987 94 iassisl quarterly map 15 long beach superior court district legend: area c enilj jurors map 15 fall/winter 1987 iassist quarterly 95 map 16 long beach superior court district legend: area i i sills jurors map 16 fall/winter 1987 96 iassist quarterly table 6 long iz.-.ch: s pa-tels 5v i.^'i ipail 2'.. 19s5 to suy.z 1935 census trace z scores census trace 1 4-24-35 26.7: 0.0 2 5-1-85 23.9 -0.5 3 5-8-85 29.8 0.7 4 5-15-85 24.6 -0.5 i 5-22-s5 24.7 -0.5 6 5-29-35 30.2 0.8 7 5-5-85 25.8 -0.2 a 6-12-35 25.9 -0.2 17. i: 0.6 12.5 -0.6 15.9 0.3 13.9 -0.3 11.3 -1.0 8.7 -1.7 13.3 -0.4 12.0 -0.7 13.7 fall/winter 1987 iassist quarterly 97 map 17 long beach district: 8 panels legemo: area i i sum jurors fall/winter 1987 iassist quarterly appendix a long beach superior court district legehd: area i i outside district mw>m 10 mile radius fck-ffim 5 mile radiuse^ 15 mile radius appendix a fall/winter 1987 instructions for authors of the iassist quarterly 1/11 johnson, nastasha, sapp nelson, megan, and yngve, katherine (2022) deficit, asset, or whole person? institutional data practices that impact belongingness, iassist quarterly 46(4), pp. 1-11. doi: https://doi.org/10.29173/iq1031 deficit, asset, or whole person? institutional data practices that impact belongingness nastasha johnson1 , megan sapp nelson2, and katherine yngve3 abstract given the capitalist model of higher education that has developed since the 1980s, the data collected by institutions of higher education on students is based on micro-targeting to understand and retain students as consumers, and to retain that customer base (i.e. to prevent attrition/dropouts). institutional data has long been collected but the authors will question how, why, and for whom the data is collected in the current higher education model. the authors will then turn to the current higher education focus on equity, diversity, inclusion, and particularly on the concept of belongingness in higher education. the authors question the collective and local purposes of institutional data collection and the fallout of the current practices and will argue that using existing institutional data to facilitate student belongingness is impossible with current practices. we will propose a new framework of asset-minded institutional data practices that centers the student as a whole person and recenters data collection away from the concept of students as commodities. we propose a new framework based on data feminism that intends to elevate qualitative data and all persons/experiences along the bell-shaped curve, not just the middle two quadrants. keywords data management, critical methodologies, asset mindedness, institutional data introduction beginning in earnest in the 1980s, but accelerating in the 2000s, higher education has been transformed under a market-oriented, “academic capitalist” model. academic capitalism as defined by rhoades & slaughter “[blurs] the boundaries between the for-profit and not-for-profit sectors, and a basic change in academy practices changes that prioritize revenue generation, rather than the unfettered expansion of knowledge, in policy negotiation and in strategic and academic decision making” (rhoades & slaughter, 2004). in the case of the united states, higher education shifted to accommodate the political philosophies and demands of state and federal legislatures that were decreasing support while simultaneously requiring further alignment of the curriculum with the future work economy (schulze-cleven & olson, 2017). the rise of neoliberalism in higher education at the same time introduced a commitment to the concept of higher education as a marketplace, not only of ideas but of students (olssen & peters, 2005). this alignment with a market-oriented model resulted in a students-as-customers ethos as higher education institutions were required to rely on tuition for day-to-day operational funding, and therefore, the development of robust data gathering on the customer pool in order to induce the customers to stay with the company, i.e. increase student retention as a part of the return on investment for the student, the state and the federal government (oblinger, 2012). this student retention effort stems from the mindset that the lack of persistence by a student is the fault of the student and not the institution itself. that is, the neoliberal mindset suggests there is an implicit deficit in every student who cannot persist and earn a degree. bok eloquently suggests that “it's like making them [students] do a play without a script” in his germane work (bok, 2010). the data https://doi.org/10.29173/iq1031 2/11 johnson, nastasha, sapp nelson, megan, and yngve, katherine (2022) deficit, asset, or whole person? institutional data practices that impact belongingness, iassist quarterly 46(4), pp. 1-11. doi: https://doi.org/10.29173/iq1031 collected is being sought in order to identify those at the highest risk of leaving the higher education institution and to prevent those students from being lost (and their tuition dollars along with them) from the matriculation process. in the process, all students who enter the university system have data collected that similarly assume the potential for loss from the system, because all students are the market for the university “product”, i.e., educational services. higher education shrouds the capitalistic nature of the industry that it is and is able to retain an aura of service for the public good (dorn, 2017). therefore, the data practices and the data collected are also shrouded. “much of it is useless to graduating high school students trying to understand whether they are likely to succeed in a specific program or major at a specific institution. [the ipeds metrics] also [do] not effectively track student outcome data with regards to workforce specific indicators, such as employment metrics and post-college earnings. finally many of these ipeds metrics are not disaggregated by key student characteristics, such as race/ethnicity, gender, income status, or veteran status because reporting more disaggregated metrics requires a greater level of effort by institutions reporting the data.” (krishnamoorthi and kaissi, 2020) “institutional data, the administration, and transactional data collected by the university, are already collected, and maintained at the individual level in application, registration, degree databases, and matched to the survey results” (university of california at berkeley, 2022). that is, institutional data comes in many forms and from many sources, including self-reported and publicly available data. it may come from applications for admission or scholarship, fafsa data, micro-targeting individuals and specific groups using social media, or even transcripts for high school. because of the variety of sources and the ways that they were harnessed, the possibilities of biased, inaccurate, or incomplete data are boundless. some institutions have made attempts to define and control institutional data and how it is used. but those efforts do not capture what could equitably be done to ensure appropriate use. for example, at indiana university, institutional data is defined as: • “it is subject to a legal obligation requiring the university to responsibly manage the data. • it is substantive and relevant to the planning, managing, operating, documenting, staffing, or auditing of one or more major administrative functions, or multiple organizational units, of the university. • it is included in an official university report. • it is clinical data or research data that meets the definition of "university work" under the university's [intellectual property policy]. • it is used to derive any data element that meets the above criteria.” (indiana university, 2022) statement of the problem academic capitalism uses student data as a foundational resource in decision-making, resource allocation, and predictive analytics. but because of the intent of the institution to use the student data for capitalistic purposes, institutional data collection methods fall short of providing institutions with appropriate and adequate data for other purposes, such as those to create more diverse and equitable educational experiences. academic capitalism leans into the notion that colleges and universities are a part of a global knowledge market, and students are commodities within and of that market. in this paper, we propose that academic capitalism’s micro-targeting falls short of creating individual footprint digital footprints to create narratives that foster belonging and thereby success. we propose https://doi.org/10.29173/iq1031 3/11 johnson, nastasha, sapp nelson, megan, and yngve, katherine (2022) deficit, asset, or whole person? institutional data practices that impact belongingness, iassist quarterly 46(4), pp. 1-11. doi: https://doi.org/10.29173/iq1031 an alternative view of institutional data collection methods that is asset-focused that will more clearly align with an academic institution’s articulated goals to facilitate belongingness, diversity, equity, and inclusion. literature review belongingness in the diversity, equity, and inclusion goals for higher education institutions, one of the data facets that those institutions may seek to assess is the concept of belongingness. belongingness is a core human need as defined by maslow (1981), centered on the connection of the self to one’s surroundings, community, possessions, or nearby objects or people. hagerty et al. (1992) have defined a ‘sense of belonging’ as the experience of personal involvement in a system or environment so that persons feel themselves to be an integral part of that system or environment. though both of these definitions are germane to the psychological understanding of belonging, for the purposes of the paper, strayhorn’s definition is especially appropriate as “the sense of belonging refers to students’ perceived social support on campus, a feeling or sensation of connectedness, and the experience of mattering and feeling cared about, accepted, respected, valued by, and important to the campus community or others on campus such as faculty, staff, and peers (strayhorn, 2018). baumeister and leary (1995) identify belongingness as a core motivator for much of what human beings do. hausmann, schofield, and woods (2007) found that belongingness in higher education was positively associated with peer group interactions, interactions with faculty, peer support, and parental support, along with institutional commitment and intention to persist, while glass and westmont (2014) found that leadership programs, cultural events, and community service enhanced belongingness. interestingly, academic integration and student background were not associated with sense of belongingness (hausmann et al., 2007) . although feelings of group membership, social connectedness and belongingness have been well defined and heavily researched in primary and secondary school contexts since at least the 1990s, university-level research into and data collection about student belongingness has received far less attention (ingram, 2012; slaten et al, 2017). much of the information we have about belongingness, particularly as it relates to college students of racialized identities, exists either as part of a very context-specific program evaluation study (hausmann et al., 2007) or as data disaggregation performed on results from national-level studies (stebleton et al., 2010). although definitely correlated and often conflated, belongingness, retention, and success are not at all the same thing. qualitative studies of the minority experience in u.s. universities often strongly suggest that, for students of marginalized identities, “success” means finding a way to graduate as fast as possible from an institution at which one has never truly felt welcomed or felt that they belonged (davis et al., 2004). retention, in the educational context, describes the enrollment and successful completion of the required courses, from year to year, for a student to graduate from a specified program and is often conflated with the term persistence which is the continuation to the path to graduation (nces, 2022). failure of retention often implies a deficit on the part of a student, rather than a systemic failure of the higher education institution. however, a student’s sense of belongingness impacts their success, retention, and persistence (haussman 2007). measuring belongingness has been relatively well received with assessments such as the scale of ethnocultural empathy (wang, 2003) and the social connectedness scale (lee, 1995), however, those measures do https://doi.org/10.29173/iq1031 4/11 johnson, nastasha, sapp nelson, megan, and yngve, katherine (2022) deficit, asset, or whole person? institutional data practices that impact belongingness, iassist quarterly 46(4), pp. 1-11. doi: https://doi.org/10.29173/iq1031 not account for the data collected otherwise, and the deficit framing of the interpretation of those data. belongingness is fundamentally different in that it is student-focused, asset-focused, and descriptive of the life experience of the student (and hence, nearly impossible to describe quantitatively in easy metrics, especially when it rebuffs the purpose, paradigm, and sources of the data collected.) organizations within higher education institutions that gather data regarding belongingness are scattered across the institution. both data gathering and data transparency vary widely depending upon the area of the institution that is gathering the data. in the u.s., institutional data researchers began to coalesce as a profession in the late 1960s, and are now incentivized by federal and state regulatory structures, which in turn are tied to compliance and funding mechanisms that focus almost exclusively on educational outputs (e.g. retention, graduation and post-graduation employment) (volkwein, 2008). within this system, student belongingness often matters at the institutional level only inasmuch as it contributes to those bottom-line considerations (chirikov, 2013). as a consequence, the metric of student belongingness frequently tends to exist outside the power hierarchy, often falling under the purview of units such as academic assessment or co-curricular assessment, whose professional support structures began to coalesce in the 1990’s (delaney, 2009) and in the 2010s (bresciani, 2006), respectively. many u.s. universities periodically administer a large-scale survey instrument to students to gain actionable information about the student experience. typically this survey is developed by a nationally recognized entity, for a group of fee-paying consortium members and is administered locally out of an individual consortium member’s central institutional research or academic assessment office. generally, because of a combination of cost and complexity, such a survey is administered to an institution’s student body every two or four years, rather than annually. some universities also regularly collect belongingness data via either survey or qualitative methods as part of the continuous quality improvement cycle for student support services, such as new student orientation, suicide prevention efforts or mentorship programs. depending on the institutional context, these smaller program evaluation efforts inconsistently feature as part of an institutional level inclusion report or other regular cycles of student outcomes transparency. one of the earliest institutional consortium instruments was the college student experience survey (known as the cseq or csxq), which was operated out of indiana university from 1979-2014. the national survey of student engagement (nsse) is another staple in the assessment of national trends and localized aggregate data about students’ collegiate experiences, including involvement in high-impact practices (hips), overall engagement (kuh, 2003). critical methodologies according to finocchiaro (1979), critical methodology is the “study of the nature of criticism and of proper methods of criticism.” the author goes on to say “…critical methodology is the distinction between practice and theory, i.e. between what one does and one’s reflections on what one does” (p. 364). critical methodology is a discourse that critiques data collection and analysis methods in order to understand the implications of those data handling decisions on the resultant data sets and analysis. by understanding the data collection and analysis process through this critical lens, we can better articulate the range of ways that the data’s story can be understood. does current institutional data reflect the stories of the students enrolled in our institution, and thereby tell us what they need from https://doi.org/10.29173/iq1031 5/11 johnson, nastasha, sapp nelson, megan, and yngve, katherine (2022) deficit, asset, or whole person? institutional data practices that impact belongingness, iassist quarterly 46(4), pp. 1-11. doi: https://doi.org/10.29173/iq1031 us to belong? and secondarily, does it tell us what we need as an institution to ensure their belonging? critical methodologies call for the examination of current methods of collecting institutional data, the theories undergirding those practices, and they propose different pieces of evidence be collected based in the narratives of the students. schrag (1980) called for critical reflection of the scientific method that “deconstructs the layers of methodological and metaphysical conceptualization that surround man’s inquiries about himself and his world so as to reopen the text of everyday life and make visible its language, thought, and praxis (p. 126).” it is the everyday lives of our students and prospective students where belonging is defined and created. we currently do not have the institutional data to presume what that could mean on our campuses. schrag describes “methodological naiveté” as the mistaken assumption that “more universality” empirical methods are the only valuable option for measuring student belongingness or student experience, pretentiously (schrag, 1980). we have simply lost the origin of the questions that we are asking, both of the complex and nuanced social dynamics of people collectively and of individuals specifically (schrag, 1980). the tension between the academic capitalist tendency to measure belongingness as an individual focused trait, and the psychology-based definition of belongingness has implications for both the student experience on campus and the efforts to measure belongingness as a metric. belongingness in the writings of bell hooks is the human collective moving towards community (hooks, 2009). in her book, belonging: the culture of place, she reflects on the duplicity of space and nature, being both black and white, rich and poor, invited and excluded (hooks, 2009). her backdrop is appalachia but her realization of the intricacies of going to a new place cannot be ignored. college students migrate to new homes when they move onto college campuses. but their understanding of their new surroundings is defined by what they perceive as safe and comfortable, that is, what feels like home. institutional data at best may minimize their home to zip code, high school attended, and family income. however, those variables do not define belonging in their previous home or the new campus home. it is storytelling and personal narratives like those demonstrated by bell hooks in belonging that can uncover the nuances of the lived experiences of the students with whom we are invited to our campus communities, but also who have been trusted to nurture and guide towards success (denzin et al., 2008). the constructs of nurture and success are subjective, nuanced, and can shift along the student developmental cycle from high school graduate to college graduate. the variables and metrics are not static, which requires a paradigm shift from neoliberal to student-centered. institutional data methodologies can gather snapshots over an academic lifecycle at a university and can center the origin of the data (the person) rather than the data itself (abes, et.al, 2019). weberpillwax suggests that research data belongs to the source community (weber-pillwax, 2004). we suggest that institutional data perhaps should also belong and represent the interests of those entering the institution. data feminism, coined by d’ignazio and klein (2020), criticizes the collection and use of data science and data ethics at the intersection of feminist thought and ideology, specifically challenging power and elevating the pursuit of justice for all. data feminism is defined by seven guiding principles: • examine power • challenge unequal power structures and working toward justice • elevate emotion and embodiment with multiple forms of knowledge • rethink binaries and hierarchies https://doi.org/10.29173/iq1031 6/11 johnson, nastasha, sapp nelson, megan, and yngve, katherine (2022) deficit, asset, or whole person? institutional data practices that impact belongingness, iassist quarterly 46(4), pp. 1-11. doi: https://doi.org/10.29173/iq1031 • embrace pluralism • consider context • make labor visible (d’ignazio and klein 2020). in the context of institutional data methodologies, data feminism challenges who has the power to decide what data is gathered and for what purposes, regardless of the source of the data, which is inherently biased. current practices elevate one capitalistic voice, centered in western and patriarchal thought, that by nature decenters marginalized experiences and authentic voices (d’ignazio and klein 2020). the numbers are not enough and do not give the complete story of lived experiences, and cannot be minimalized to a single quantitative measure. belonging requires that a person or a set of people be valued and matter to an institution, which includes the nuances of their experiences and shared pursuit of subjective growth and success. data feminism would tell us that to measure belonging as a data point requires that the data not be viewed in a vacuum and pursues the understanding of context and social norms and expectations (d’ ignazio and klein 2020). current capitalistic methods of collecting institutional data rely on the market and micro-targeting to tell the story, which is in direct contrast to the centering of a variety of forms of knowledge and ways of knowing, which will encapsulate the entirety of the embodiment of community. in order to reframe data collection away from demographic data and micro-targeting/marketing focused institutional data collection, additional strategies for data collection are needed. those data collection strategies can be developed using multiple existing theories that center the whole student, rather than a shorthand that captures quantitative, descriptive statistics about the person in a deficitdriven frame. campus ecological theory & qualitative network theory campus ecology is the study of interactions and interdependence between students and their campus environment, which is dynamic in fostering learning, engagement, and belonging (banning & bryner, 2001). though the campus ecology movement began in the 1970s, it is not as robustly studied contemporarily. as an interdisciplinary field, literature can be found in student affairs, counseling psychology, environmental psychology, and developmental psychology (banning & bryner, 2001). developmental psychology and constructs of belonging are interwoven and cannot be easily disentangled, but instead respected and normalized within the traditional capitalistic measures of institutional data methodology, institutional and community belonging experienced by students (cabrera et. al, 2016). campus ecology provides a tool that will allow us to examine the communities that arise and provide a sense of belonging for students within the higher education, campus-focused environment (renn 2003). using the findings of a campus ecology-focused analysis, we can gather qualitative focused data assets that reflect a more accurate understanding of how students engage with the campus environment. the use of this theory gives a starting point that allows campuses to gather asset-based data that reflect both the student experience and the impact of systemic factors. as noted in johnson, sapp nelson, and yngve 2022, the theories listed above are useful starting points, but to actually collect usable data that describes belongingness that is asset-based and meaningful, the theories must be tied together and focused on the lived experience of the students. qualitative network theory centers the students and their experiences from within the community, rather than outside of the community, as is done with institutional data methodologies (ahrens 2018). this theory https://doi.org/10.29173/iq1031 7/11 johnson, nastasha, sapp nelson, megan, and yngve, katherine (2022) deficit, asset, or whole person? institutional data practices that impact belongingness, iassist quarterly 46(4), pp. 1-11. doi: https://doi.org/10.29173/iq1031 centers the student as the source of the data describing the sense of belongingness in the campus environment, rather than the campus or higher-education environment itself. asset-mindedness vs deficit-mindedness asset-mindedness is a way of thinking that focuses on strengths, whereas deficit-mindedness is a way of thinking that focuses on deficits or shortcomings (pendakur, 2020). as it pertains to institutional data, the data collected about students and their families are often seen through the lens of what resources or gaps need to be filled, rather than what attributes and strengths the student, their family, and their perspectives bring to the institution. for example, does the family or the student need additional resources to fill in gaps (e.g. lower household income, high school courses on the transcript, highest degree obtained by a parent) could be deficit-framed, whereas does the family have lived experiences and perspectives that are valuable to the campus community (e.g. ability to make new friends, exposure to different lifestyles, ability to connect with others) is asset-framed towards the contributions that the family can bring. table 1 summarizes the characteristics of these mindsets. shifting institutional data methods toward the theories described above and framing them to center on individuals and their experiences enables the institution to frame a variety of students as assets and contributing members to the community rather than potential deficits that will be lost due to attrition. switching the framework of data collection fosters flexibility in the perspectives that decision makers of the institution have by providing a wider range of insights into the assets and strengths that the students bring to the organization. leaning into non-capitalistic models create campus and collegiate experiences where all can experience success and not just those for whom existing systems have been built. table 1. comparison between assets and deficits mindedness asset-based mindedness deficit-based mindedness • strengths driven • opportunity focused • internally focused • what is present that we can build upon • may lead to new, unexpected response • needs driven • problem focused • externally focused • what is missing that i must go find • may lead to downward spiral of burnout source: adapted from: https://www.memphis.edu/ess/module4/page3.php conclusion linda tuhiwai smith, an indigenous education scholar at the university of waikato who has penned germane texts on indigenous methodologies and critiques of western research methodologies posed eight questions to researchers as they embark on new projects: what research do we want done, whom is it for, what difference will it make, who will carry it out, https://doi.org/10.29173/iq1031 https://www.memphis.edu/ess/module4/page3.php 8/11 johnson, nastasha, sapp nelson, megan, and yngve, katherine (2022) deficit, asset, or whole person? institutional data practices that impact belongingness, iassist quarterly 46(4), pp. 1-11. doi: https://doi.org/10.29173/iq1031 how do we want the research done, how will we know is worthwhile, who will own the research, and who will benefit (smith, 2017) with these questions in mind, we must ask ourselves and our institutions: what research do we want to be done, how do we disentangle admissions data, belonging research, student success data, and the like and how do we be more specific about what we are doing and why? we must also realize that mixed methods are ok and perhaps even necessary to tell an adequate story about our students and their lived experiences. we must ask ourselves for whom we are researching and for whom are building our higher education institutions are we researching for ourselves strictly, the accrediting agencies, for our competitors or peers, or for industry e.g. (to demonstrate how “employable” our students are)? we must also ask ourselves who will do the data collection and if they are the most representative, unbiased, and/or suitable researchers for the type of projects that we are embarking on. is there room for representation? should we trust a third-party vendor to capture the data in a manner that matters for our students and our community? and is our work student-centered, assetbased or does it benefit the institution and create gaps for our students to address on their own? these are questions that can be addressed with deliberative decision making about data collection at the institutional level. the constructs presented in this paper present new ways of thinking about institutional data practices and methodologies. references abes, e. s., jones, s. r., & stewart, d. l. (eds.). (2019). rethinking college student development theory using critical frameworks. stylus publishing, llc. banning, j. h., & bryner, c. e. (2001). a framework for organizing the scholarship of campus ecology. colorado state university journal of student affairs, 10, 9-20. baumeister, r. f., & leary, m. r. (1995). the need to belong: desire for interpersonal attachments as a fundamental human motivation. psychological bulletin, 117(3), 497-529. https://doi.org/10.1037/0033-2909.117.3.497 bok, j. (2010). the capacity to aspire to higher education:‘it's like making them do a play without a script’. critical studies in education, 51(2), 163-178. bresciani, m. j. (2006). outcomes-based academic and co-curricular program review: a compilation of institutional good practices. stylus publishing, llc. cabrera, n. l., watson, j. s., & franklin, j. d. (2016). racial arrested development: a critical whiteness analysis of the campus ecology. journal of college student development, 57(2), 119-134. chirikov, i. (2013). research universities as knowledge networks: the role of institutional research. studies in higher education, 38(3), 456-469. d'ignazio, c., & klein, l. f. (2020). data feminism. the mit press. https://doi.org/10.29173/iq1031 https://doi.org/10.1037/0033-2909.117.3.497 9/11 johnson, nastasha, sapp nelson, megan, and yngve, katherine (2022) deficit, asset, or whole person? institutional data practices that impact belongingness, iassist quarterly 46(4), pp. 1-11. doi: https://doi.org/10.29173/iq1031 davis, m., dias-bowie, y., greenberg, k., klukken, g., pollio, h. r., thomas, s. p., & thompson, c. l. (2004). “a fly in the buttermilk”: descriptions of university life by successful black undergraduate students at a predominately white southeastern university. the journal of higher education, 75(4), 420-445. delaney, a. m. (2009). institutional researchers' expanding roles: policy, planning, program evaluation, assessment, and new research methodologies. new directions for institutional research, 2009(143), 29-41. denzin, n. k., lincoln, y. s., & smith, l. t. (2008). handbook of critical and indigenous methodologies. sage. dorn, c. (2017). for the common good: a new history of the higher education in america. cornell university press. glass, c. r., & westmont, c. m. (2014). comparative effects of belongingness on the academic success and cross-cultural interactions of domestic and international students. international journal of intercultural relations, 38, 106-119. https://doi.org/10.1016/j.ijintrel.2013.04.004 hagerty, b. m. k., lynch-sauer, j., patusky, k. l., bouwsema, m., & collier, p. (1992). sense of belonging: a vital mental health concept. archives of psychiatric nursing, 6(3), 172-177. https://doi.org/10.1016/0883-9417(92)90028-h hausmann, l. r. m., schofield, j. w., & woods, r. l. (2007). sense of belonging as a predictor of intentions to persist among african american and white first-year college students. research in higher education, 48(7), 803-839. https://doi.org/10.1007/s11162-007-9052-9 hooks, b. (2009). belonging: a culture of place. routledge. indiana university. (2022). iu data management. https://datamanagement.iu.edu/dataclassifications/public-data.html johnson, n., sapp nelson, m., and yngve, k. (2022). the impact of “academic capitalism” on “belongingness”: institutional data’s impact on diversity, equity and inclusion. research data and preservation 2022 summit. https://rdapassociation.org/page-18210 krishnamoorthi, r. & kaissi, b. (2020). the college transparency act: strengthening transparency, equity, and student success in american higher education. harvard journal on legislation, 57(1), 2-24 kuh, g. d. (2003). what we're learning about student engagement from nsse: benchmarks for effective educational practices. change: the magazine of higher learning, 35(2), 24-32. lee, r. m., & robbins, s. b. (1995). measuring belongingness: the social connectedness and the social assurance scales. journal of counseling psychology, 42(2), 232-241. maslow, a. h. (1981). motivation and personality. prabhat prakashan. national center for education statistics. (2022). undergraduate retention and graduation rates. https://doi.org/10.29173/iq1031 https://doi.org/10.1016/j.ijintrel.2013.04.004 https://doi.org/10.1016/0883-9417(92)90028-h https://doi.org/10.1007/s11162-007-9052-9 https://datamanagement.iu.edu/data-classifications/public-data.html https://datamanagement.iu.edu/data-classifications/public-data.html https://rdapassociation.org/page-18210 10/11 johnson, nastasha, sapp nelson, megan, and yngve, katherine (2022) deficit, asset, or whole person? institutional data practices that impact belongingness, iassist quarterly 46(4), pp. 1-11. doi: https://doi.org/10.29173/iq1031 condition of education. u.s. department of education, institute of education sciences. retrieved october 5, 2022, from https://nces.ed.gov/programs/coe/indicator/ctr. oblinger, d. (2012). game changers: education and information technologies. educause. https://www.educause.edu/ir/library/pdf/pub7203.pdf olssen, m., & peters, m. a. (2005, 2005/01/01). neoliberalism, higher education and the knowledge economy: from the free market to knowledge capitalism. journal of education policy, 20(3), 313-345. https://doi.org/10.1080/02680930500108718 pendakur, v. (2020). designing for racial equity in student affairs: embedding equity frames into your student success programs. change: the magazine of higher learning, 52(2), 84-88. renn, k. a., & arnold, k. d. (2003). reconceptualizing research on college student peer culture. the journal of higher education, 74(3), 261-291. rhoades, g., & slaughter, s. a. (2004). academic capitalism and the new economy: markets, state, and higher education. jhu press. schrag, c. o. (1980). radical reflection and the origin of the human sciences. purdue university press. schulze-cleven, t., & olson, j. r. (2017). worlds of higher education transformed: toward varieties of academic capitalism. higher education, 73(6), 813-831. smith, l. t. (2017). towards developing indigenous methodologies: kaupapa māori research. critical conversations in kaupapa māori. wellington: huia publishers. stebleton, m. j., huesman jr, r. l., & kuzhabekova, a. (2010). do i belong here? exploring immigrant college student responses on the seru survey sense of belonging/satisfaction factor. strayhorn, t. l. (2018). college students’ sense of belonging: a key to educational success for all students. routledge. university of california at berkeley. (2022). institutional data. office of planning and analysis. retrieved 3/20/2022 from https://opa.berkeley.edu/institutional-data volkwein, j. f. (2008). the foundations and evolution of institutional research. new directions for higher education, 141, 5-20. wang, y. w., davidson, m. m., yakushko, o. f., savoy, h. b., tan, j. a., & bleier, j. k. (2003). the scale of ethnocultural empathy: development, validation, and reliability. journal of counseling psychology, 50(2), 221. weber-pillwax, c. (2004). indigenous researchers and indigenous research methods: cultural influences or cultural determinants of research methods. pimatisiwin: a journal of aboriginal & indigenous community health, 2(1). https://doi.org/10.29173/iq1031 https://nces.ed.gov/programs/coe/indicator/ctr/undergrad-retention-graduation https://www.educause.edu/ir/library/pdf/pub7203.pdf https://doi.org/10.1080/02680930500108718 https://opa.berkeley.edu/institutional-data 11/11 johnson, nastasha, sapp nelson, megan, and yngve, katherine (2022) deficit, asset, or whole person? institutional data practices that impact belongingness, iassist quarterly 46(4), pp. 1-11. doi: https://doi.org/10.29173/iq1031 endnotes 1 nejohnson@purdue.edu 2 mrsapp@purdue.edu 3 kyngve@purdue.edu https://doi.org/10.29173/iq1031 mailto:nejohnson@purdue.edu mailto:mrsapp@purdue.edu mailto:kyngve@purdue.edu 4 iassist quarterly 2008 editor’s notes welcome to the iassist quarterly (iq) volume 32 2008. we have collected the 1, 2, 3, and 4 issues into a single issue for 2008 in order to catch up with our schedule. in the text below mostly by cutting and pasting i am giving a short appetizer for the articles in this issue of the iq. this type of editorial is among the few areas where plagiarism is actually welcomed. nikos askitas is the head of the international data service center of the institute for the study of labor in germany (iza). at the 2008 iassist conference he presented what is here an article on the “data documentation and remote computing at the international data service center". the data documentation of the idsc that started with translation of german metadata into english has developed into a detailed, in depth, searchable and standardized information service, especially helpful for comparative research. the datasets are in the areas: employment and wages, education and training, and demographics and migration. the documentation is available in html, pdf, and as ddi-files. this documentation production is explained in the first part of the article. in the second part of the article the idsc experience with “remote computing” is described. germany uses the concept of “factual anonymization” and the production of “scientific use files”. however, such files are not allowed for export. instead idsc supplies interfaces to scientists with both local and remote support for which idsc has developed special software (josua). the article “a documentation model for comparative research based on harmonization strategies” by john kallas from university of the aegean at mytilene on lesvos, greece and apostolos linardis from the national centre for social research in athens, is proposing a documentation model for both longitudinal and crosscultural studies. different harmonization strategies are examined and three documentation models are proposed. the authors have chosen the term “cross-cultural” rather than “cross-national” as cultural discrepancies may exist within the same nation. the article underlines the importance of the data element, the concept, the universe, and the classification as they are study components where even small changes may affect the overall comparability. this is leading to looking at the stages for the different types of harmonization strategies: ex ante input, ex ante output, and ex post. most documentation processes at data archives are ex post harmonization. the authors are aware that the proposed study documentation procedure is laborious for the researchers; however, the positive side is the benefits in searching and locating the data. at the 2009 iassist conference in the session “protecting privacy while preserving access: restricted use data and disclosure considerations”, sharon bolton and matthew woollard gave a presentation that is now an article titled “strengthening data security: an holistic approach” and they are advocating exactly that. the authors both work at uk data archive (ukda) as data services manager and head of digital preservation and systems. the holistic approach to data security includes “the education of data creators in the reduction of disclosure risk, the integration of robust and appropriate data processing, handling and management procedures, the value of emerging technological solutions, the training of data users in data security, and the importance of management control, as well as the need to be informed by emerging government security and digital preservation standards”. the background is a massive governmental data loss that hit front pages and has resulted in reports and laws with criminal penalties for the disclosure of confidential information. these lessons as well as the laws are relevant for the archival society. the ukda had an audit of its “in-house data handling” which resulted in existing good practice being identified and additional methods developed. these were collated into a comprehensive set of data security procedures with effect for both ukda staff and the users. at the same iassist conference in the session “sharing data: high rewards, formidable barriers” carina carlhed and iris alfredsson from respectively mälardalen university, sweden and the swedish national data service (snd) presented a report from an investigation carried out earlier in 2009. the report has been turned into an article for the iq with the title: “swedish national data service’s strategy for sharing and mediating data. practices of open access to and reuse of research data the state of the art in sweden 2009”. the report is based upon a joint project between snd and four university libraries that carried out a national survey of existing databases and database research, as well as attitudes towards data sharing among researchers. this was carried out by email questionnaires sent to professors and doctoral students. in general the results show that doctoral students expressed great uncertainty about questions of amounts of reusable digital data, while professors emphasize lack of resources for researchers to document and make their data accessible for others. the groups consider the most effective interventions for enhancing accessibility to digital data to be that research grants should include funds for preparing the data for sharing and archiving, and that archiving data for use by the scientific community is acknowledged to be of scientific merit. a similar study was carried out in finland and compared to this swedish study. we hope to present the finnish study in a later issue of the iq. the swedish research council founded in 2006 a database infrastructure committee (disc) to promote iassist quarterly winter summer 2008 5 the development of an effective infrastructure for sharing research data. a product of this initiative has been the formation of the swedish national data service (snd) that also is described in the article. the article further describes the procedures of the surveys and there might be followers for doing similar user investigations among other data organizations. the survey contains questions as to the knowledge of plans such as the roadmap “the swedish research council´s guide to infrastructure” (2007) and the “oecd guidelines on open access to research data from public funding” (2007). answers to these questions exhibited a low level of knowledge, as did questions about making own data available. read more in the article, and also about reasons given for not reusing digital data, and the seven suggested obstacles to sharing digital data. should you be interested in compiling issues for the iq as guest editor(s) please contact me. if you don’t have anything to offer right now, then please prepare yourselves for the coming iassist 2010 if you are acting as chair for a session there. that is an obvious opportunity to bring quality sessions to more people than the session participants and also making the memory more sticky. take a look at the website http://iassistdata.org and the iassist blog the iassist communiqué – at http:// iassistblog.org. articles for the iassist quarterly are very welcome. articles can be papers from iassist conferences, from other conferences, from local presentations, discussion input, etc. contact the editor via e-mail: kbr@sam.sdu.dk. karsten boye rasmussen october 2009 18 iassist quarterly 2010 / 2011 iassist quarterly qualitative research in ireland: archiving strategies and development by dr. jane gray and dr aileen o’carroll1 abstract the irish qualitative data archive (iqda) was established in 2008 with initial funding for three years under the fourth cycle of the irish government’s programme for research in third level institutions (prtli4). iqda aims to become the central access point for irish qualitative social science data, including interviews, pictures and other non-numerical material. we have established protocols to ensure that newly generated qualitative data are documented and stored in ways that facilitate sharing and re-use through online access. currently we are developing our digital infrastructure, piloting a number of initial collections and building a catalogue of irish qualitative research. the catalogue has already begun the process of mapping potentially available data, and on the basis of that survey we provide an overview of the kinds of qualitative data that could be archived in ireland. this paper also reports on: the funding situation for social science research and policies in relation to archiving, iqda’s role in national social science research, our progress to date and the potential obstacles to and benefits of qualitative archiving in ireland. information about the archive can be found at www.iqda.ie keywords: qualitative data, longitudinal research, archiving, secondary analysis, ireland introduction the irish qualitative data archive was established in 2008, in response to growing concerns that the potential of social science data being gathered by various institutions and agencies was being lost, as much of the data collected did not have a life beyond the specific projects for which it was obtained, thereby limiting potential use and re-use. many state agencies are increasingly aware of how the lack of archiving policies is a barrier to knowledge production. in june 2008, the higher education authority (hea, the statutory body with responsibility for higher education in ireland) introduced a policy on open access to published research. they argued that: the intellectual effectiveness and progress of the widespread research community may be continually enhanced where the community has access and recourse to as wide a range of shared knowledge and findings as possible. this is particularly the case in the realm of publicly funded research where there is a need to ensure the advancement of scientific research and innovation in the interests of society and the economy, without unnecessary duplication of research effort (higher education authority, 2008). as a consequence, researchers obtaining part or all hea funding are required to lodge publications resulting from their results in an open-access repository as soon as possible2 . the policy further states, “data in general should as far as is feasible be made openly accessible, in keeping with best practice for reproducibility of scientific results (ibid).” however there are barriers to implementing this policy in practice. forfás is the national policy advisory board for enterprise, trade, science, technology and innovation in the republic of ireland. data archives and repositories were identified as an area needing attention in the arts, humanities and social sciences in a report jointly produced by the hea and fórfas, research infrastructure in ireland, building for tomorrow (2007). the report argued that, “the absence of data storage and archive facilities and their ability to be updated with fresh datasets is a serious and continuing impediment to high quality social science research in ireland.”(royal irish academy,n.d.a) when the report was published, the only archive for social science data was the irish social science data archive, based at the university college dublin, which focused exclusively on medium and large quantitative data sets. as will be seen below, in terms of qualitative data, a number of small data archives existed but these tended to be linked to a particular project rather than oriented towards the creation of a general collection. the sense that inadequate investment in data archives was a weakness in the humanities and social science research (hss) infrastructure was further emphasised by the royal irish academy in advancing humanities and social sciences research in ireland (2007). they argued that “a key resource issue for the hss is the availability of, and access to, data sets, iassist quarterly 2010 / 2011 19 iassist quarterly including the capacity to generate new data, as well as the importance of ensuring widespread availability of, and access to, previously gathered data.” (royal irish academy, 2007: 15-17) in addition to facilitating secondary analysis of existing data archives, they maintained that archives (along with libraries and museums have a key role in “ensuring that ireland’s cultural heritage is recorded and maintained for posterity”(ibid) it is within this context of this increasing need to build on the wealth of previous research that the irish qualitative data archive (iqda) was established in 2008. it forms part of the irish social science platform (www.issplatform.ie), an organisation which brings together irish academics from 19 disciplines in 8 institutions in ireland. the iqda is initially funded by the higher education authority under the fourth cycle of the programme for research in third level institutions (prtli4). 1. qualitative research in ireland according to conway (2006), social research in ireland has drawn predominantly on quantitative methods, with qualitative methods such as ethnography, participant observation and archival research receiving much less attention, though he notes that qualitative studies are evident in some sub-fields such as the sociology of religion. in 1988 in a statement on the social sciences arising out of a conference hosted by the royal irish academy, damian hannan criticised the weakness of qualitative research in ireland (cited in o’dowd, 1988). since this time there has been an increasing use of qualitative methods. most undergraduate and post-graduate social science courses contain both quantitative and qualitative research components and in many universities postgraduate theses are based on qualitative methods. the iqda is currently mapping the extent of qualitative research in ireland. we have created an online catalogue of qualitative research. this is an ongoing project but by march 2011 the catalogue contained information on over 480 projects using qualitative research methods. the constituency using qualitative methods has also expanded beyond social science departments in universities. the iqda has presented information seminars on archiving to audiences varying from postgraduate social scientists, to nursing students, to oral historians. many irish government research bodies incorporate a qualitative element to their research processes. for example, the women and crisis pregnancy study, conducted by mahon et al (1998) and commissioned from the department of health and childcare, was based on in-depth interviews and the crisis pregnancy agency continues to commission mixed methods research. the project ‘growing up in ireland,’ (described in more detail below) also combines quantitative and qualitative methodologies in its research design. in compiling entries to the iqda catalogue, we found that qualitative research had been produced by various commissioners of research, such as the national children’s office, the employment research centre in trinity college, combat poverty agency, the crisis pregnancy agency, focus ireland, the national consultative committee on racism and inter culturalism (nccri), the equality authority, the national disability authority and the immigrant council. the methods used include ethnography and community studies and the use of in-depth interviews and focus groups. recently there has been increasing interest in biographical, life history and oral history approaches and the use of more varied types of research methods such as the use of photographs (quinlan, 2008) or texts written by young people (o’connor et al, 2002). additionally in recent years, there has been an increase in longitudinal research, including qualitative longitudinal projects, and those using mixed methods. we have identified a number of longitudinal projects (see appendix 1) with an integral qualitative component based in ireland (some have been completed, others are on-going). major government and eu-funded quantitative longitudinal studies are archived at the irish social science data archive . we are in the process of archiving some of the qualitative longitudinal projects mentioned above. 2. archiving and the research culture all new qualitative social science data generated within the irish social science platform www.issplatform.ie) will be deposited electronically and made available online through the iqda. this represents a major step forward in the development of a culture of data sharing and archiving within the social science research community. the iqda is engaged in the ongoing development of protocols (see below) that will frame the parameters and standards for archiving qualitative social science data within the irish research community. however, more work needs to be done in order for data sharing and archiving for reuse to become an integral part of the research culture in ireland. in particular the cost requirements of preparing data for archiving need to be built in to research funding, and funding agencies need to change their approach away from simply facilitating the production of data towards supporting its use and re-use. additionally projects based on secondary use of archival data need to eligible for funding (even where no ‘new data’ is being produced); this would require a shift in funder priorities in favour of data analysis (with less emphasis being placed on the generation of new data). 3. current archiving strategies the remit of the iqda is to archive irish qualitative data. this remit is interpreted broadly to include • research by irish researchers on ireland and irish issues, • research on ireland by visiting researchers, • research by irish and non-irish researchers on the irish diaspora • research by irish researchers living outside of ireland, • research on northern ireland (as appropriate, in co-operation with the northern ireland qualitative archive (niqa)) the iqda is committed to archiving new qualitative social science data generated within the irish social science platform. we will also archive selected ‘legacy’ research projects. we currently have nine projects either deposited or in various stages of preparation for deposit addressing a range of issues 1. life histories, 20th century ireland 2. integration of new migrants to ireland 3. career trajectories of returning irish migrants 4. social life in new irish suburbs 5. growing up in ireland: the national longitudinal study on children 6. protestants and irishness in ireland 7. irish women’s work experiences during the second world war 8. raccer: re-use and archiving of complex community-based evaluation research 9. nirsa photographic archive of the people and places of ireland 9. ned cassidy photographic archive of the irish built environment since 1970 a variety of methodologies are represented, including life story interviews, life history calendars, social network schedules, qualitative interviews, observation, shadowing and visual methods, focus group interviews, key-informant interviews, children’s essays, drawings and other visual methods (for further information see the appendix). the iqda is primarily a digital repository as it lacks the resources to store other types of data. therefore it accepts interview transcripts, field notes, and other research documentation, audiovisual files and photographs in digital format only. we provide advice on the preferred 20 iassist quarterly 2010 / 2011 iassist quarterly formats for archiving and on the technological options available to researchers. 4 infrastructure for the management and re-use of data a key goal of the irish qualitative data archive is to provide a national infrastructure for the management and re-use of qualitative longitudinal data. there are a number of components to this goal. firstly it involves implementing digital storage solutions that enable depositing of and access to data. this includes the setting of metadata standards to facilitate data retrieval. secondly protocols that meet the ethical responsibilities of qualitative researchers must be in place. thirdly, it will be necessary to create networks of researchers around the archive to facilitate and encourage the re-use of the data. these issues will be discussed in the following section. 4.1 technical database structure one of the first goals of the iqda was to design a digital infrastructure that would integrate two previously existing photographic archives with a new catalogue of qualitative research in ireland as well as a database containing non-visual data. the existing photo archives were housed in two different databases, built specifically for each set of data. most qualitative archives appear to adopt the approach of building such bespoke solutions. generally the difficulties in this approach are that such solutions become increasingly difficult to maintain over time and that there is a certain amount of redundancy in effort, as internationally, each archive reinvents its own ‘wheel’. iqda was designed to use fedora commons, a general purpose open source digital repository, in preference o building bespoke databases. fedora commons has the advantage of being an open source project that is widely used and as such there is a large community of support to draw on if future problems arise. in addition, there is considerable documentation on its use. it is a general purpose solution which allows almost all types of digital objects to be stored; additionally it can preserve relationships with other objects. it is being run on a virtual gnu/linux server. the front end of the database is powered by fez, an open source interface to fedora. in a piloting phase we tested two other systems: isadora (a drupal plug-in; drupal is an open source content management system) and elated z but decided that fez provided the best solution. fez was initially designed by developers at the university of queensland, and has since been transferred to the sourceforge (a webbased source code repository) so that it now is developed by a number of different organisations. one note of caution must be observed however: fez is designed with library projects in mind and as such the default is often to openness and accessibility – something that is not necessarily always desired with more sensitive qualitative social science material. therefore it was necessary to modify the fez operation such that it defaulted to the highest levels of security (see below). indeed one factor in the decision to implement fez was its ability to allow us to set access controls over each individual piece of data. finally we use drupal as the content management system which powers the iqda web presence. this was chosen because, again, it is an open source software, supported by a community of users, with many tutorials on its use available online. additionally its modular system allows the site to be upgraded with little difficulty and the capability of the site to be expanded over time (for example, we have added a modification which allows us to display selected photographs from our archive online). it is hoped that drupal’s powerful capabilities will enable us to expand the site such that it can facilitate the development of networks of irish researchers. a goal of the archive is to become a central access point for information about other archives in ireland (such as the migrant lives and women’s oral history projects described above) and elsewhere. the drupal end of the site collates information about these archives, as well as information on preparing data for archiving, upcoming events and issues of note to qualitative researchers. in the future we plan to develop a component on using qualitative data in teaching. 4.2 metadata metadata, often described as ‘data about data’, is vital for enabling data retrieval (either through browsing or searching collections). metadata standards can facilitate robust data management and assist in the uploading of data. further, shared metadata standards allow coordination with other data sets and harvesting of the data (to ensure, for example, that search engines like google are able to identify and return searches appropriate to the archive). a key challenge of the iqda was to introduce shared standards across the pre-existing databases that we inherited and the newly created ones. dublin core is an international standard for metadata, designing the minimum numbers of elements necessary to allow objects in a networked environment to be discovered. it is a very minimal standard. the iqda added to this standard, drawing on the metadata applied in other national qualitative archives (particularly esds qualidata a specialist service of the esds led by the uk data archive, and the henry murray archive based at harvard university, boston in the u.s.). metadata relating to the content of the material archived is in the first instance added by the researcher depositing the data. they are encouraged to use hasset (humanities and social science electronic) thesaurus when specifying key words. the data documentation initiative (ddi) is a “metadata specification, an emerging international standard for the content, presentation, transport, and preservation of documentation about datasets in the social and behavioural sciences” (www.cessda.org). it is our intention to provide contextual documentation that is “marked up” according to the data documentation initiative (ddi) and we are tracking developments in the generation of qualitative ddi standards. in addition to being searchable by key word, we are currently drawing on expertise available to us within the national institute for regional and spatial analysis to develop a map based interface which will encourage search by geographical place. for confidentiality reasons, most interview data will be searchable only by broad geographical area (see below), however for photographs it will be at the level of the townland or village. 4.3 access and confidentiality a key concern of the iqda is to meet the ethical commitments that researchers make with those who participate in their research. as with quantitative data there is always some degree of risk that confidentiality could be breached, and so archives have an obligation to ensure that interviewees are protected through anonymisation, withdrawal of sensitive data etc. the remit of the irish qualitative data archive includes the establishment of procedures and protocols, in line with international best practice appropriate to qualitative data. to that end we have developed an ethical use framework drawing on best practice developed at esds qualidata, timescapes, a multidisciplinary longitudinal qualitative project based at the university of leeds and the henry murray archive in the us. our best practice handbook is now available (http://www.iqda.ie/sites/default/files/iqda_best_ practice_handbook.pdf ) there are four interconnecting components to this framework. the first is informed consent to archive data obtained at the time of the fieldwork. the iqda has prepared pro-forma letters and forms that have been used in previous irish and uk studies. the second is the use of a rigorous anonymisation protocol. such a protocol has been developed by the iqda based on the experience of irish and uk research projects. as opitiz and witzel (2005) outline, it is difficult to develop a general solution to the problem of anonymisation because such data iassist quarterly 2010 / 2011 21 iassist quarterly are extremely heterogeneous in terms of the themes and areas of life covered. the protocols that we have developed alert researchers to issues they need to be concerned with and suggest possible solutions. the third component is a rights management framework which includes depositor and end-user licenses and legal agreements, in which the user undertakes not to breach confidentiality by using identifiable information in published work or to try to contact research subjects and agrees to ethical use and re-use of the data. the third and final component is a system of options for access and user restrictions; for example, access to very sensitive data may be closed for a period of time. in addition, along with all social science researchers in ireland, we subscribe to the ethical standards imposed by professional organizations such as the sociological association of ireland (www.ucd.ie/sai). part of our role involves alerting researchers to these standards, thus promoting high quality research procedures. our work here is also informed by the respect project (www.respectproject.org), which was funded by the european commission’s information society technologies (ist) programme, to draw up professional and ethical guidelines for the conduct of socio-economic research. 5. challenges for the future in order to fulfil its mandate to become the central access point for all qualitative social science data generated within the irish research community, the iqda will require additional and sustainable funding and must face a number of key challenges in the future. key to the success of the archive will be to encourage use and re-use of data though a process of training, publicity and the development of networks of researchers. in addition continued liaison with the social science research community and other stakeholders will ensure that the archive will develop in a way that matches the needs of researchers. we especially aim to work with post-graduate students, to foster a culture of data archiving and re-use, and ensure that archiving is built in at an early stage of future projects. this work must be matched by liaison with agencies responsible for the development and implementation of social science policy, in order to promote the use of qualitative social science research findings and funding for the processes of preparing data for archiving, which will ensure long term value for money in state funded projects. in addition there is a considerable volume of legacy research in ireland, currently sitting in offices and under desks. we have been contacted by researchers asking about the possibility of adding these to the archive. one challenge will be to address problems such as the absence of informed consent and outdated digital formats when archiving of such legacy data. we will continue work with other digital archiving projects in ireland, and with the recently announced national audiovisual repository (navr), a multi-institutional project funded under prtli5, in which iqda is a funded partner (http://www.ria.ie/research/navr.aspx) to develop and implement a common set of methods, policies and standards and continue ongoing development of policies and guidelines that comply with national and international law. we aim to broaden the development of our it infrastructure and data management tools especially to facilitate deposit, search and retrieval of data. currently we are developing sample ‘soundscapes’ drawn from data within the archive in order to raise the profile of the resource. we are additionaly part of equalan, the newly formed european network of qualitative researchers and archivists. equalan is committed to promoting and implementing a strategy for preserving and organising qualitative and qualitative longitudinal data resources, and we intend to work within this network to develop shared approaches to archiving and to promoting re-use of our data. we will continue to promote best standards and to educate the research and higher education community about the desirability of qualitative social science data archiving. as we have shown above there is an emerging interest in qualitative longitudinal research. archiving such research requires intense, ongoing management and interaction with researchers, but potential rewards are high. however, the principal barrier anticipated is the securing of sustainable funding into the future. lack of such security hinders long term management and planning. appendix 1: longitudinal projects with qualitative dimensions 1. life histories and social change in 20th century ireland in this project qualitative interviews, live history calendars and retrospective social network schedules were collected from a large sample of irish people from three key birth cohorts: 1929-1934; 1949-1954 and 1969-1974. the research design links retrospective qualitative longitudinal data is linked to a panel study. the interviewees were selected from the nationally representative sample of people interviewed from 1994-2001 for the irish part of the european community household panel. the data will be archived in the irish qualitative data archive. http://sociology.nuim.ie/lifehistory.shtml .2. growing up in ireland – qualitative module of the national longitudinal study on children in this government funded study prospective qualitative data is linked to a panel study. the study follows two cohorts of children, aged nine months and nine years. the study was officially launched in 2007 and will be completed in 2013. there is an embedded qualitative module, the data from which will be archived in the irish qualitative data archive. data from the first wave are now available. http://www.growingup.ie/ 3. towards a dynamic approach to research on migration and integration this is a prospective longitudinal qualitative study of migrants to ireland will track sixty migrants from two migrant cohorts over a two year period through interviews, observation, shadowing and visual methods. the project will run until december 2010. the data will be archived in the irish qualitative data archive. http://geography.nuim.ie/staff/gilmartinmary 4. the process of youth homelessness: a qualitative longitudinal cohort study this prospective qll study uses a life history method to investigate the experience of youth homelessness based on young people’s accounts of becoming and being homeless. criteria for inclusion include: a) being between 12-22 years and; b) being homeless or in insecure accommodation. forty young people were recruited with the co-operation of statutory and voluntary agencies with responsibility for providing services and interventions to young people who are without a home. the recruitment strategy aimed to include relevant 22 iassist quarterly 2010 / 2011 iassist quarterly diversity across key variables including age, gender and geographical location. phase i: baseline interviews were conducted with 40 homeless young people, based on the life history model, which prioritises young people’s accounts and experiences, both past and present. phase ii: follow-up life history interviews are ongoing at present. http://www.tcd.ie/childrensresearchcentre/index.php?id=121&prid=16 5. migrant careers and aspirations this is a prospective qualitative longitude panel study of polish migrants to ireland. their sample included 22 men and women aged between 22 and 38 years who almost all arrived in ireland post 2004, following enlargement of the european union. interviews were conducted every four months over a period of two years, including interviews with migrants who have returned to poland. 6. irish centre for migration studies life narratives collection this is a retrospective qualitative longitudinal study, data from which has been collated in an online digital archive. this archive contains four collections of life narratives centring on different aspects of migration in the irish experience. http://migration.ucc.ie/oralarchive/testing/index.html 7. women’s oral history project this retrospective qualitative longitudinal study documents the working lives of munster women during the period 1936-1960, through the collection of oral histories. the project is a study of the stories of women who engaged in paid work in the period 1936-1960. http://www.ucc.ie/wisp/ohp/index.html 8. leaving school in ireland this prospective longitudinal study tracks the school and post-school experiences of a cohort of young people who took part in the postprimary longitudinal study (ppls) in 2002. in 2010 in-depth interviews will detail the influences on young people’s post-school choices and pathways. http://www.esri.ie/research/research_areas/education/ leaving_school_ireland/ references conway, b. (2006). “foreigners, faith and fatherland: the historical origins. development and present status of irish sociology”. sociological origins. 5(1) data documentation initiative. [online]. available at http://www.cessda. org/sharing/managing/3/. [accessed 24th june 2009] higher education authority. (2008). [online]. available at http://www. hea.ie/files/files/file/open%20access%20pdf_.pdf . [accessed 24th june 2009]. irish social science data archive. [online]. available at http://www.ucd. ie/issda/ [accessed 24th june 2009] irish social sciences platform. [online]. available at http://www.issplatform.ie/index.html [accessed 24th june 2009] mahon, e., conlon, c. & dillon, l. (1998). women and crisis pregnancy. dublin: the stationery office o’dowd, l (1988). the state of social science research in ireland. dublin: royal irish academy o’connor p, lane c and haynes, a. (2002). ‘young people’s ideas about time and space’. irish journal of sociology. 11 (1). pp43-61 opitz, d. & witzel, a. (2005). ‘the concept and architecture of the bremen life course archive’. forum: qualitative social research. 6(2). [online]. available at http://www.qualitative-research.net/index.php/ fqs/article/view/460/0. [accessed 24th june 2009] quinlan, c. (2008). home and belonging: a study of women in prison in ireland in belongings edited by m. corcoran and p.share. ireland: institute of public administration. pp129-138 respect project. [online]. available at: http://www.respectproject.org/ main/index.php. [accessed 24th june 2009] royal irish academy. (2007). advancing humanities and social sciences research in ireland. [online]. available at: http://www. ria.ie/policy/pdfs/website.pdf?id=26&cat=policy+reports+ p 17. [accessed 24th june 2009] royal irish academya. (n.d). [online]. available at http://www.ria.ie/ policy/humanities-socialsciences.html [accessed 24th june 2009] royal irish academyb (n.d)[ online]. available at http://www.ria.ie/ourwork/research/navr-announcement.aspx [accessed 24th june 2009] sociological association of ireland. [online]. available at http://www. ucd.ie/sai/sai_ethics.htm. [accessed 24th june 2009) notes 1. contact details: dr. jane gray, department of sociology and national institute for regional and spatial analysis (nirsa), nui maynooth, county kildare, ireland. jane.gray@nuim.ie http://sociology.nuim.ie/janepersonalpage.shtml http://nuim.academia.edu/janegray dr. aileen o’carroll, national institute for regional and spatial analysis, nui maynooth, county kildare, ireland aileen.ocarroll@nuim.ie 2. all researchers must lodge their publications resulting in whole or in part from hea-funded research in an open access repository as soon as is practical after publication, and to be made openly accessible within 6 calendar months at the latest, subject to copyright agreement. http://www.hea.ie/files/files/file/open%20access%20 pdf_.pdf page 2 3. for further information see, http://www.ucd.ie/issda/ 4. for further information see, http://www.esds.ac.uk/qualidata/about/ introduction.asp 5. for further information see, http://www.murray.harvard.edu/ 6. for more information see, http://www.data microsoft word 42-4-phegley.docx 1/9 phegley, lauren & lynda kellam (2025). how are we fair-ing? creating a fair self-assessment checklist for data repositories, iassist quarterly 49(3), pp. 1-9. doi: https://doi.org/10.29173/iq1152 the creative commons-attribution-noncommercial license 4.0 international applies to all works published by iassist quarterly. authors will retain copyright of the work and full publishing rights. how are we fair-ing? creating a fair self-assessment checklist for data repositories lauren phegley1 and lynda kellam2 abstract in 2023, a team from a local grant-funded medical data repository requested guidance from penn libraries on evaluating the extent to which their repository was fair-enabling. they wanted to know if their repository had the policies, infrastructure, and documentation to allow data that is deposited to be fair – findable, accessible, interoperable, and reusable. after a consultation with the repository team, our research data experts discovered that many of the current self-assessments of the fair guidelines were for data creators rather than data repository managers. in addition, we wanted a selfassessment tool similar to the process and guidance created by coretrustseal but focusing explicitly on the fair principles. in answer to their request, the penn libraries research data engineer conducted a literature review and coalesced current guidance and assessment tools on the principles. after this review of the existing documentation, a small team developed a fair principles selfassessment tool for repository teams. in addition to several iterations of the tool, we also met with the repository team for feedback on making the tool more understandable. our conversation provided insights into the challenges of explaining the fair principles to those without information science or data backgrounds. the discussion and creation of this self-assessment tool helped develop a more transparent and trustworthy repository. this paper will discuss our process for developing the assessment, the goals for utilizing the tool, and the lessons learned. reporting our findings as they currently stand will prompt the research data management field to ruminate on the adoption of fair principles for data repositories. we also intend to encourage conversation on the usability of the fair principles for professionals without an information science or data background. keywords fair principles, data repository, data sharing, assessment introduction the first publication on the fair principles was in the 2016 article “the fair guiding principles for scientific data management and stewardship” by wilkinson et al. the intent of the fair principles is to create discipline agnostic guidance to enhance the reproducibility of data. the four foundational principles – findability, accessibility, interoperability, and reusability – are a minimal set of guiding practices that improve the experience for the data creator and downstream users. the fair principles were created as guidance for any individual or system that encounters the data throughout its 2/9 phegley, lauren & lynda kellam (2025). how are we fair-ing? creating a fair self-assessment checklist for data repositories, iassist quarterly 49(3), pp. 1-9. doi: https://doi.org/10.29173/iq1152 lifecycle. each of the words that make up the fair acronym are considered high level principles and have specific indicators (labeled f1, f2, etc.) to use in evaluating if a dataset meets the principle. findability is focused on the ability for machines and humans to identify the data through high quality metadata and indexing. the accessibility principle expects a free, open, and standardized way of accessing the data through their identifier. the interoperable principle expects the entities data and metadata to follow a shared, defined vocabulary that assists in having broadly understood meanings. the reusable principle focuses on how well defined the licenses, provenance, and metadata are for future users. to assist users of the assessment, we linked each of the fair principle indicators to a go fair resource page that elaborates and explains its purpose and how it is used. go fair is a stakeholder driven initiative dedicated to helping implement the fair guiding principles, and one of the most thorough resources for learning more about the fair principles (n.d.). in february of 2023, a grant-funded medical data repository team reached out to the penn libraries’ research data & digital scholarship unit seeking a way to assess their implementation of the fair principles. a small group consisting of the research data engineer, the director of research data & digital scholarship, the director of the holman biotech commons, and the bioinformatics librarian at the penn libraries met with the repository team and an individual from the perelman school of medicine’s computing services to discuss their goals. upon meeting, the data repository team explained they were invited to a symposium by their grant funder to discuss how their repository aligns with the fair principles. the issue was translating the high-level, conceptual fair principles into specific criteria that they could use to evaluate their repository. our team decided that we would search for an existing fair principles assessment that met their needs. our search for an assessment instrument for fair-enabling data repositories that fit the team's needs uncovered a multitude of resources that did not quite fit. the data repository team needed a tool that did not expect previous information science training or require a detailed understanding of the fair principles. since the publication of the fair principles (wilkinson et al., 2016), there has been a plethora of research into these ideals and a deep commitment to improving the data sharing landscape. despite this, many resources are jargon-filled, created by the information science profession for the information science profession, or focused on assessing data objects rather than the repository. based on our findings, we decided to build upon the hard work of many scholars of the fair principles to create an assessment instrument for data repositories to evaluate how fairenabling they are (phegley et al., 2024). our goal was to create an accessible self-evaluation tool that could be used by repository staff to assess their current level of fair adherence, allowing them to reflect on how to take realistic, actionable steps toward improvement. we were motivated to create this new resource instead of using a pre-existing assessment for multiple reasons. first, we uncovered more questions than answers during our investigation into finding an appropriate assessment. the fair principles research sphere is bursting with scholarly products. yet it seemed that many did not delineate between the data creator’s and the repository team’s responsibility for enacting the fair principles. we wanted a resource that focused on evaluating the repositories’ fair-enabling practices rather than the choices made by the data depositor. second, we believe in the importance of the fair principles and working towards increased adoption amongst data repository teams. we want to support data repository managers’ capacity to conduct their own evaluations and improve their understanding of the fair principles. creating and sharing a new low 3/9 phegley, lauren & lynda kellam (2025). how are we fair-ing? creating a fair self-assessment checklist for data repositories, iassist quarterly 49(3), pp. 1-9. doi: https://doi.org/10.29173/iq1152 barrier to entry tool that acts as a stepping stone towards a better data sharing environment allowed us to enact our values. development of the assessment our intention was to build a fair assessment instrument that, as david et al. (2020, p. 3) states, “must be realistic and pragmatic – what should be measured, and how to explicitly find the information needed.” there is a rich offering of resources and assessments from scholars of the fair principles, and we conducted a literature review to find materials that support creating this instrument. the research data alliance (rda) fair data maturity working group created two resources that we relied on heavily for the specific questions and interpretation of the fair principles into the data repository context: fair data maturity model specification and guidelines (2020) and results of an analysis of existing fair assessment tools (2019). the rda resources, in addition to other materials, gave us a firm foundation of criteria that encapsulate the fair principles (behnke et al., 2020; hahnel and valen, 2020; l’hours, 2022; murphy et al., 2021). our assessment instrument is structured as a manual self-assessment checklist divided into four core fair sections (findable, accessible, interoperable, and reusable). each category includes detailed fair principle indicators as developed by wilkinson et al. (2016). we linked the indicators to the associated go fair description page to provide additional context (n.d.). the indicators follow the wilkinson et al. labeling that is commonly used for referencing the fair principles sub-sections. for example, indicator f1 is the first indicator of findable. each of the indicator sections has associated criteria describing a fair practice (“dataset has an assigned identifier”) (figure 1, label a). most of the indicators have multiple criteria that are labeled according to the fair principles indicator they respond to. for example, criteria f1-01 states, “dataset has an assigned identifier,” and criteria f1-02 is “related works (connected literature, data, authors, project, and code) are able to be connected via persistent identifiers”. researchers then tick the checkbox(s) that best describes their repository’s current adoption level of that fair practice (figure 1, label b). 4/9 phegley, lauren & lynda kellam (2025). how are we fair-ing? creating a fair self-assessment checklist for data repositories, iassist quarterly 49(3), pp. 1-9. doi: https://doi.org/10.29173/iq1152 figure 1. criteria f1-01 from the implementation checklist figure 1: the first indicator for a findable (f1) repository has criteria (f1-01) that allow evaluators to identify the repositories current adoption level. the evaluator is then able to explain their level of implementation and provide a self assessment. we intentionally built this tool as a manual self-assessment that allows for nuanced evaluation and thoughtful self-reflection. inspired by the coretrustseal self-evaluation process, we wanted repository staff to use the levels of implementation to reflect on current repository practices (coretrustseal standards and certification board, 2022). as such, there is a space at the bottom of each fair criteria for a narrative self-reflection to explain why they have implemented certain processes and how they want to improve (figure 1, label c). to know how to improve, the data repository team needs to know where they stand. reflection encourages the team to engage with the fair process in a deeper way than would be possible with an automated tool. undoubtedly, this translates into the checklist taking more time than an automated tool. the benefit of this process is that the repository team can learn the fair principles overall rather than relying on an automated tool to tell them what might be wrong without context. while automated tools tools are useful, they do not assist in building a team’s understanding of the fair principles in relation to their own repositories. additionally, we created an evaluation that is not prescriptive in the fair-enabling criteria. in other words, the checklist can be applied to repositories without having to worry about an all or none approach to fair implementation. not all repositories can achieve the same implementation levels due to policy requirements, budget constraints, repository infrastructure limitations, and disciplinary practices. all the criteria contribute towards creating a fair-enabling repository, but certain criteria are more crucial than others. this is why each criterion has an associated priority level of essential, important, or useful (figure 1, label d). more than one adoption level checkbox can be checked at a time to create a flexible yet reliable evaluation, as more than one scenario might be true in a repository. the order of the adoption level checkboxes tends to increase in complexity and comprehensiveness. this allows repository teams to 5/9 phegley, lauren & lynda kellam (2025). how are we fair-ing? creating a fair self-assessment checklist for data repositories, iassist quarterly 49(3), pp. 1-9. doi: https://doi.org/10.29173/iq1152 see the next logical step for increased fair adoption for that criterion. they can provide a detailed explanation in the self-assessment section about current implementation, reasoning for why increased implementation may not be viable, and discuss their plans, if they want to improve. the level of implementation is assessed by repository personnel based on their own understanding of their repository. due to the differences in repositories, there is currently no standard for the differences between implementation levels. as described in the assessment section, feedback from the data repository team showed us that it is difficult for individuals who are not experts in the fair principles to evaluate their own level of implementation. developing a standard for the levels of implementation of the fair principles is a future opportunity for development. this approach allows data repositories with various constraints to participate in a fair evaluation. for example, a specialist repository, such as a gut microbiome data repository, may only encourage standard compliance for criteria r1.3-02, “repository requires or encourages data to comply with an applicable community standard.” gut microbiome data can take many forms (genomic data, proteomic data, etc.), where each data type may or may not have its own community standard. because of this, it may be more realistic to encourage compliance with an applicable standard, rather than requiring data to comply with a standard. this is an example of a situation where an automated fair evaluation would miss the important context that a manual self-evaluation adds. assessment feedback upon presenting the first iteration of the assessment to the data repository team, we met with them to get feedback on their experience using the tool. the data repository team amounted to about four people, all of them with extensive biomedical experience but less information science experience. overall, they found our tool to be useful in figuring out where they were and where they wanted to go, but they had trouble interpreting the technical elements. they found that the assessment required a level of technicality they could not understand and requested additional descriptive information to help them. they suggested that it might be a more meaningful activity if an outside expert could act as a second reviewer. this is the existing model of the coretrustseal process and while we agree this is a beneficial process, we want the assessment to be accomplished without expecting a secondary reviewer. not all repository teams have a data librarian or fair principles expert who can assist with this process. they also had questions about how to evaluate their implementation level, as they themselves were not experts in the fair principles. we were inspired by the coretrustseal practice of the repository team evaluating their own level of implementation (coretrustseal standards and certification board, 2022). this suggestion led us to create two versions of our checklist, one where the reviewer can evaluate the level of implementation for each criterion and one where there is no level of implementation section. we also had the opportunity for an expert in the fair principles to review the tool and provide feedback before we released the updated version on the university of pennsylvania’s institutional repository, scholarlycommons. the feedback from the team and the fair expert helped us understand what we could improve in the assessment and the information that researchers need to understand the fair principles overall.3 6/9 phegley, lauren & lynda kellam (2025). how are we fair-ing? creating a fair self-assessment checklist for data repositories, iassist quarterly 49(3), pp. 1-9. doi: https://doi.org/10.29173/iq1152 addition to the fair landscape our assessment builds on the existing fair principles research by shifting the focus of what is being evaluated. while the originating fair principles article by wilkinson et al. (2016) states, “the principles define characteristics that contemporary data resources, tools, vocabularies, and infrastructures should exhibit to assist discovery and reuse by third-parties”, many tools currently are oriented towards data objects. only a few resources are oriented toward supporting fair-enabling data repository infrastructures. we anticipate that as the fair principles continue to gain attention and have increased uptake, especially outside of the core library and data communities, the principles will continue to be adopted as a standard for other adjacent assets. a second aspect of this assessment is that it differentiates between the data creator and the repository team regarding who is enacting the fair principles. our tool is made to evaluate the fair-enabling adoption levels of a repository, not to evaluate any individual dataset. for criteria that would traditionally be the responsibility of the data creator, such as data format or documentation, our assessment allows repository teams to indicate that they encourage or require a certain action. assessment tools should not assume that the creators of the datasets will update the datasets to comply with fair principles or that the individuals running the repository have the time or tools to do it themselves. we cannot fully assess a repository and its containing datasets as one unit until we realize sharing fair data has to be a conversation between the data creator and the repository manager, as neither can be fair-enabling alone. from the beginning of this endeavor, we have approached the fair principles as a goal to achieve rather than a marker of success or failure. we do not want repository managers to assume they fail if they are not fully fair-enabling. we want the tool to enable honest evaluation of the fair status of the repository and provide a plan with concrete steps towards further adoption of the fair principles. our goal was to take a step towards making the fair principles in this assessment usable for those without an information science or data background. approaching assessment from a place of shame or admonishment for not fully implementing the fair principles negates the tool's usefulness by removing the possibility for learning and growth. lessons learned the first lesson we learned was that the current fair principles information ecosystem expects a certain amount of background knowledge prior to introducing the principles. few materials attempt to explain topics like “protocols” in an accessible way. most materials seem to be written by information professionals for information professionals, which defeats the purpose of trying to increase adoption among all data stakeholders. implementing tools to create or support faircompliant data will only be useful if the individuals implementing them understand the entire fairification process (david et al., 2020). we need to create information and tools that support researchers beginning their fair journey without expecting knowledge on topics such as machine actionability or metadata. if we believe in the importance of fair data, we need to create scaffolded learning material and tools. this lesson was informed by our experience reviewing the fair principles literature and the repository teams’ feedback on our assessment. our experience demonstrated that there is a conceptual knowledge gap between the concept of data and the application of the fair principles. the repository 7/9 phegley, lauren & lynda kellam (2025). how are we fair-ing? creating a fair self-assessment checklist for data repositories, iassist quarterly 49(3), pp. 1-9. doi: https://doi.org/10.29173/iq1152 team had a deep expertise with data, specifically bioinformatics data, but that knowledge did not directly translate into knowledge of the fair principles. even the research data engineer on the penn libraries team who was familiar with the fair principles had to spend many hours untangling jargonfilled sentences, defining terms, and attempting to weave the information into shareable knowledge. the repository team let us know that while using the assessment, there were certain questions that they could not answer due to confusion over the criteria’s meaning. we attempted to provide more clarification without bogging down the tool by linking to each of the principles developed by the go fair foundation (n.d.). the team did not find that the go fair explanations clarified the criteria and had to skip sections of our assessment due to this. they also wanted to see examples of repositories that had fully implemented the fair principles to emulate. we agree this would be an incredible learning tool, but we are unaware of a repository that fulfills these criteria in a way that can be demonstrated to individuals without requiring administrative-level access and an in-depth walkthrough by a repository staff member. this feedback demonstrates that there is a knowledge gap between the intentionally high concept fair principles and the specific actions required to implement those principles. the second lesson we learned was that fair researchers and educators need to separate what is attributed to a fair-enabling data repository and the dataset aligning to the fair principles. the ability to share a fair dataset is a collaboration between a repository’s infrastructure and its datasets. the checklist was developed for repository managers to assess the repository’s infrastructure and policies and evaluate how much they allow full fair adherence. for example, when building checklist criteria for fair principle r1: “(meta)data are richly described with a plurality of accurate and relevant attributes,” the criteria asks if the repository requires or encourages documentation to accompany datasets (phegley et al., 2024). documentation is incredibly important for fair data, but we did not want the repository to be deemed less fair-enabling because the depositor chose not to create rich documentation. instead, the repository shows it is fair-enabling by requiring or encouraging rich documentation with all datasets. meaningful fair data sharing only happens when the repository is built to support fair data sharing, and the data is also created to meet the fair principles. future directions throughout the process of creating the fair data repository checklist, we realized the importance of scaffolded fair principles instruction. as research data management educators, we need to create accessible information on the fair principles that professionals without an information science or data background can use. unfortunately, directing someone to the fair principles does not address what the principles mean in context or how to implement them. we lack jargon free guidance with example case studies and visuals that scaffold towards a comprehensive understanding of the fair principles. we hope this article and the associated checklist will prompt conversation in the research data management field on evaluating the fair principles in a data repository. the fair assessment checklist is an early iteration of what we hope to improve based on feedback from data repository teams, fair experts, and the research data management community. we are especially eager to hear from individuals who have used the checklist to evaluate their data repository. as our community works 8/9 phegley, lauren & lynda kellam (2025). how are we fair-ing? creating a fair self-assessment checklist for data repositories, iassist quarterly 49(3), pp. 1-9. doi: https://doi.org/10.29173/iq1152 towards supporting fair-enabling data repositories, we look forward to this assessment being used as a stepping stone towards a more robust understanding of the fair principles. references bahim, c., dekkers, m. and wyns, b. (2019). “results of an analysis of existing fair assessment tools”. https://doi.org/10.15497/rda00035. behnke, c., bonino, l., coen, g., le franc, y., parland-von essen, j., riungu-kalliosaari, l., and staiger, c. (2020). “d2.3 set of fair data repositories features”. fairsfair. https://doi.org/10.5281/zenodo.3631527. coretrustseal standards and certification board. (2022). “coretrustseal requirements 2023-2025”. https://doi.org/10.5281/zenodo.7051011. david, r. et al. (2020). “fairness literacy: the achilles’ heel of applying fair principles”, data science journal, 19, p. 32. https://doi.org/10.5334/dsj-2020-032. go fair foundation. interpreting fair. https://www.gofair.foundation/interpretation (accessed: 2024-04-10). hahnel, m. and valen, d. (2020). “how to (easily) extend the fairness of existing repositories”, data intelligence, 2(1–2), pp. 192–198. https://doi.org/10.1162/dint_a_00041. l'hours, h. (2022). “report on a maturity model towards fair data in fair repositories (d4.6)”. https://doi.org/10.5281/zenodo.6699520. murphy, f., bar-sinai, m. and martone, m.e. (2021). “a tool for assessing alignment of biomedical data repositories with open, fair, citation and trustworthy principles”. plos one. edited by f. naudet, 16(7), p. e0253538. https://doi.org/10.1371/journal.pone.0253538. phegley, l., kellam., l., de la cruz gutierrez, m., and rajpal, n. (2024). “fair assessment checklist for data repositories”. https://doi.org/10.48659/dtqc-2a45. research data alliance fair data maturity model working group. (2020). “fair data maturity model: specification and guidelines”. https://doi.org/10.15497/rda00050. wilkinson, m.d. et al. (2016). “the fair guiding principles for scientific data management and stewardship”. scientific data, 3(1), p. 160018. https://doi.org/10.1038/sdata.2016.18. endnotes 1 lauren phegley is the research data engineer at the university of pennsylvania. her email is lphegley@upenn.edu and orcid is 0000-0001-7897-1841. 2 lynda kellam is the director of research data and digital scholarship at the university of pennsylvania. her orcid is 0000-0002-3263-859x. 9/9 phegley, lauren & lynda kellam (2025). how are we fair-ing? creating a fair self-assessment checklist for data repositories, iassist quarterly 49(3), pp. 1-9. doi: https://doi.org/10.29173/iq1152 3 the authors are incredibly thankful to laurence horton, research data specialist at the digital curation centre, for his enthusiastic feedback on the assessment. iassist quarterlyiassist quarterly abstract with support from the national science foundation, two long-running social science studies – the american national election study and the general social survey – partnered with the inter-university consortium for political and social research (icpsr) and norc at the university of chicago to improve their metadata and build demonstration tools to illustrate the value of structured, machine-actionable metadata. the partnership also involved evaluating the studies’ data collection workflows to determine where in the data life cycle metadata could be captured at source to avoid metadata loss and costly procedures to recreate the metadata later. this article reports on the experience and knowledge gained over the course of the project and also includes recommendations for others undertaking similar work. keywords: metadata, documentation, data documentation initiative (ddi), data life cycle, data dissemination, tools, workflows background the metadata portal project, a collaboration among the general social survey at norc at the university of chicago, the american national election study at the university of michigan, and the inter-university consortium for political and social research, with technical support provided by metadata technology north america, was funded by the national science foundation (collaborative research: metadata portal for the social sciences, ses-1229957) under the metadata for long-standing large-scale social science surveys (meta-sss) project to meet the following objectives: • to develop rich, structured metadata compliant with the data documentation initiative (ddi) standard for two premier time series studies in the social sciences — the gss and the anes • to showcase tools that can be built upon the foundation of rich metadata • to analyze and improve the projects’ workflows • to capture more metadata at the source the two-year project resulted in enhanced study and variable-level ddi markup for both data series as well as a portal linking to prototypes of several useful tools, including a robust search, a variable bank and shopping cart to generate subsets, a crossstudy concordance and concept tagging tool, and a tool that displays routing paths through a survey. agreement was also reached to transition both data series to new workflows that enable the export of documentation in ddi format from computer-assisted interviewing systems. about the studies and their distribution there are 58 separate studies comprising the anes: the traditional biennial time series studies (with preand post-election surveys in years of presidential elections) creating rich, structured metadata: lessons learned in the metadata portal project by mary vardigan1, darrell donakowski2, pascal heus3, sanda ionescu4, and julia rotondo5 both groups maintain a “master” or “canonical” version of the questionnaire... 16 iassist quarterly 2014 iassist quarterly going back to 1948; pilot studies; panel studies; and special studies of different types. topics cover voting behavior and the elections, together with questions on public opinion and attitudes of the electorate. icpsr and anes are co-distributors of most of the anes studies. there is one cumulative file for the gss that spans the years 1972 to the most recent wave (currently 2012). gss content encompasses a standard core of demographic, behavioral, and attitudinal questions, plus topics of special interest. the gss is distributed by norc, the roper center, and icpsr. project activities 1. file inventory the first major phase of the project involved inventorying the relevant files held by all of the partners to ensure a shared understanding of the files in scope for markup and distribution. some interesting findings were noted during the inventory: • icpsr distributed the data in more formats than anes and had some existing ddi files • for a few years of the series, icpsr had grouped multiple data files into single studies, while anes had kept them separate • there were a few files that only the anes was distributing • anes distributed in general more documentation than icpsr did – e.g., they had additional documents on methodology • codebook content was basically the same across the distributing organizations, but the formats sometimes differed. both icpsr and anes distributed pdf files, but other .txt files were sometimes available. • it was noted that while the cumulative file had been the standard gss product since 1977, icpsr was still disseminating single-year gss files for 1972-1977. it was recommended that icpsr rethink its practice of distributing these single-year files as they had all been subsumed into the cumulative file and may have been revised. at the least, icpsr was advised to add a note to the metadata records for these files to indicate that the data may have changed and that the cumulative file was the authoritative data source for those years. • it was noted that the gss cumulative file had long data records (5000 variables), which raised the issue of whether the file should be reshaped or subsetted to facilitate analysis these findings and the differences noted across distributors led to a need to define which versions of the files to consider the “authoritative” versions going forward. 2. file comparison and defining canonical versions to fully understand differences across the holdings of the distributing partners and to manage project content, a central source for the metadata was required; to that end, the project partners established an irods (integrated rule-oriented data system6) based file repository. this provided a flexible central system for the file comparison and conversion work, with the capabilities to add metadata and rules to files, sort, search, get notifications of newly deposited files, etc. organization and logic of the irods repository were critical, so metadata technology worked with the partners to structure the repository optimally; they also suggested file naming conventions for data and documentation that included the name of the series, type of file (e.g., pilot), and year. to complement irods functionality, a central spreadsheet for all project studies was set up with identifiers and the capability to enter study-level metadata. based on a model description provided, icpsr’s ddi xml metadata records were imported as a batch into the spreadsheet, and study-level metadata from the anes and gss websites were added as well. metadata from these sources were later integrated. using metadata technology tools -such as sledgehammer and caelum7 -as well as shell scripts, files were parsed, analyzed, and compared, with the following findings noted for anes: • file sizes differed across the anes and icpsr holdings, but this was to be expected -in general the number of variables and frequencies agreed across the two organizations • study titles differed across anes and icpsr -it was decided to use the anes titles but to retain icpsr titles tagged as alternative titles • differences, albeit small, were discovered in variable names, labels, category labels, etc. there were many variations in sas code, most likely having to do with the way icpsr and anes produced the sas scripts. in some instances the anes version of the ascii file for study was not fixed but delimited. • for one study, icpsr was distributing an older version of the data than anes • the anes 2012 time series had just been released, and it was decided to include it in the time series covered by the project based on the file comparisons, the project settled on using the anes version of the ascii file and corresponding sas syntax file as the master data/metadata to build the core ddi, updating the variable-level metadata from other sources for substantive differences, especially for the value labels. a decision was made to investigate only major/substantive differences and not typos, minor differences in labels, or file locations and widths, which would not be relevant since the project was designed only for metadata. 3. the ddi markup process the goal of the project was to generate a complete library of ddi markup for all of the anes and gss with study-level metadata and variable-level metadata including basic variable descriptions, categories with values and labels, and frequencies, with additional variable-level information to be added when possible. to convert files to ddi format, metadata technology used its parsing and extraction tools to produce the core ddi markup in an automated way. extensive effort also went into developing custom parsers for extracting metadata from legacy text files available on the anes website, and combining various metadata sources into a final ddi xml document for each study. which ddi specification to use before the markup process could begin, a decision had to be made about which version of ddi to use, even though the tools available were agnostic as to the ddi version. the ddi alliance distributes two main product lines. ddi codebook (ddi-c) is designed to include all of the elements of a typical social science codebook needed to facilitate effective data analysis. ddi lifecycle (ddi-l) has a broader focus: to document and manage data across the entire life cycle, from conceptualization to data publication and analysis and beyond. in the end the project settled on ddi codebook version 2.5 for three main reasons. first, icpsr had been using ddi-c for many years and already had existing study descriptions and some variable-level information in ddic. second, most of what the project aimed to accomplish with the ddi metadata library could be done using the simpler of the standards. a final rationale was that version 2.5, which had recently iassist quarterly 2014 17 iassist quarterly been released, was seen as a bridge to ddi-l if a conversion to the more complex standard were required later. other projects making similar decisions might consider these three factors as they determine which ddi product line will best suit their needs. automating markup as noted, metadata technology used parsing tools, custom development, and scripts to produce the markup in an automated way as much as possible. this was not always straightforward, however. anes osiris-like codebooks had a lot of rich detail at the variable level that the project wanted to capture, including interviewer instructions, lead-ins to questions, forward and backward question flow, etc., but automating the conversion of these different variable components was difficult as the formatting was not always uniform. pdf format was another barrier to markup as each file had to be converted to editable text to extract the needed metadata. identifying patterns in the text so that “families” of study documentation could be parsed together in a more efficient way was an effective solution for some of the heterogeneity encountered. in the end most of the markup was done programmatically with manual markup performed when necessary. in terms of content to include at the variable level, the project ultimately used the following elements: variable name variable id variable label – short variable label – long variable group literal question summary statistics category label category value category frequencies notes (substantive notes relevant for data analysis) sha1 hash for question text and value labels a major innovation was that mtna computed a hash for variable classifications based on a string composed of all codes and categories. this hash would enable reuse of identical classifications and also facilitate conversion from ddi codebook to ddi lifecycle. 4. building tools database and search at the end of the processes described above, new ddi metadata were produced for 58 anes surveys (79,521 variables) and the cumulative gss 1972-2012 dataset (5,558 variables). html reports were also produced for each study. these metadata were loaded in a basex8 database for querying and retrieval and indexed with apache solr9 to facilitate full and faceted searches. both systems are available over a public rest api. a web-based application leveraging these two services was built as a proof of concept. the tool allows users to search and select variables and collect them in a shopping basket, which can in turn be used for generating data subsetting scripts for spss/sas/stata and producing customized codebooks in html/pdf. visualization of question routing the project included an exploratory effort to capture/document question flow through an older survey. the process for marking up this information involved selecting a study to use as a test case, tagging the documentation in ddi, and running an icpsr-created tool called rug (reverse universe generator) to capture the flow. the markup to highlight “system missing” had to be done manually in conjunction with the codebook as the available syntax did not carry this information. input to the tool was an ascii data file and ddi 2.5 variable descriptions. the ddi variable-level metadata had system missing values flagged at the category level. the ddi “missing type” attribute on the category element was used. the assumption was that for each relevant variable a unique code was assigned to system missing values, and no other types of missings (dk, na, etc.) on the same variable carried the same code. the system missing flagging was done manually to create a working input for the tool. (this could be automated if the same code, and preferably the same label, were consistently assigned to system missing values across a dataset.) the weight variables were also flagged using the attribute “weight” on the variable element. this markup was also done manually, but was not so onerous, as there were a limited number of weight variables in a typical dataset. the rug tool identified the system missing code on each given variable and then regressed that variable on all of the other variables to find perfect matches between the system missing code and other codes on the searched variables. it only looked for single-variable dependencies and nested variables. it did not check for complex universe logic based on multiple variables. rug generated variable-level universe information (ddi-c universe element, a child of the variable element) pointing to the source of the dependency as well as the relevant categories involved. it assumed that “correlation equals causation” and automatically created universe metadata when a correlation was found between the system missing values of a particular variable and one or more categories of another variable. in some instances, the independent variables were not found. a closer analysis of the data and original documentation (used in creating the ddi) showed that on quite a few variables the system missing category included cases that were not truly system missing, but represented responses like ‘no comment’, ‘no second mention’, ‘no pro or con’, etc., from respondents who were actually asked the question. this was an important finding that directly impacted the expected tool output: for a satisfactory output, the input data and documentation need to contain accurate system missing information, assigning unique codes for system missing values and documenting them in an unambiguous way. the overall assessment of the rug tool effort was that it was an interesting experiment but that rug would be of limited use as a production tool, given the constraints related to the data and metadata input. legacy datasets would not be good candidates for the tool, as they would require case-by-case evaluation as well as metadata editing and perhaps even data reprocessing, adding to the time spent. the rug experiment highlighted the importance of high quality, complete, and accurate data documentation and clean datasets. the major lesson learned was that question routing markup should ideally be part of the documentation deposited with archives to avoid this costly retrofit work. concept comparisons the project also investigated how anes and gss operationalize some important concepts by building an integrated crosswalk of concepts. the concept comparison was generated by first using the anes cumulative file and then looking at the gss for comparable concepts. since the anes cumulative file groups variables in just a few broad topics, the work was expanded to 18 iassist quarterly 2014 iassist quarterly include the anes core utility, which offers more granularity although it only covers recent time series going back to the 1990s. related to this work on concepts, a prototype concept tagging tool was built. this permits the individual user to tag variables by concept and then build a crosswalk to compare variables over time or across different studies. it is also possible to create public lists so that an organization can apply its own authoritative tagging to its content and make it publicly available. 5. re-envisioning the process: markup at the source the process to mark up legacy documentation chronicled above was laborious and time-intensive, involving a large team of people, specialized tools, and manual work. much of the work performed, e.g., adding question text to variables, was in essence restoring information to its original state as found in the cai interview environment. a logical solution to avoiding this scenario in the future is to capture the metadata at the source – from the original cai instrument -and export to xml when the data are exported from the interview software. this would eliminate costly work on legacy materials as metadata would be harvested once at the source. to explore this idea further, the grant included funding to hold a workshop for all three partners to explore changes in the workflow of data collection and dissemination and how the projects might transition to capturing metadata at the source. the focus of the meeting, held april 24-25, 2014, in ann arbor, michigan, was for anes and gss to share their processes and compare them, applying what had been learned about creating metadata over the life of the project in order to capture more and better metadata – ideally, with less work. anes and gss staff made presentations about their surveys describing workflow processes, opportunities for capturing metadata and paradata, and challenges they faced in creating a harmonized process for the anes and gss surveys. the anes discussion focused on several of the new tools created to enhance researcher and staff experience with the survey. specifically, the questionnaire development tool sparked discussion of the workflows involved in creating the time series questionnaires of the anes. this tool permits collaboration on the questionnaire: users of the tool can draw from a pool of questions previously used and construct new questions. they can specify question provenance, randomization, timings, etc. the group also discussed the two main survey databases used by anes staff – the questionnaire database and the variables database – to facilitate the creation of the questionnaire and the ultimate codebook. the questionnaire database feeds into the variable database, as do observations and notes from the field. anes staff want to include as much information as possible in these databases in order to generate a subset for a well-constructed codebook, but often the metadata information they need from the data collectors is not automatically provided – anes staff must push for its release. it was noted that mappings between ddi and these two internal anes databases would be worthwhile. though the type of paradata desired by anes staff may be consistent – timing of questions, timing of mailings, etc. – the systems and formats used to document these events are not consistent across vendors, nor is the functionality always immediately usable. for example, data collectors can provide timestamp information on every key stroke but that doesn’t necessarily help researchers understand how long each question took to administer and answer. the anes process during data collection involves very close monitoring by the anes staff, and the staff works on the documentation continuously. the gss discussion focused first on a high-level overview of the gss workflow process over the three-year cycle, describing areas of overlap in the pre-production cycle of the next round and the post-production cycle of the current round. discussion then moved to norc’s early attempts to map the gss to the gsbpm (generic statistical business process model), a reference model developed by national statistical organizations from around the world to standardize the production of official statistics. this mapping work involved interviewing norc gss staff, doing an environmental scan of gss dissemination sites to examine the types of metadata available to researchers, and creating a visual flow of the gss work process. during the discussion on creating post-production documentation, the group found similarities between the anes and the gss, though the processes were different. both groups maintain a “master” or “canonical” version of the questionnaire, which is then used to check data collected from the field. actions are taken when deviations are found – either by adding notes in the databases for the codebooks in anes or by doing post-production checks against the capi in gss to see if the discrepancy was due to a mechanical failure that can be resolved easily or not. metadata capture – and specifically where in the process it is captured – was a theme that arose frequently during the workshop. for both the gss and anes it appeared that most metadata were entered into codebooks and databases manually rather than through any automated process. both surveys recorded metadata during data collection, but this information was often ad hoc observations and did not always follow specified pre-planned processes. anes discussed a previous round where they had created rigid procedures for call note documentation for the data collectors. they found that with a strong process in place, they got very useable data – but this was the result of a lesson learned from a previous request for call note documentation that led to unusable information (not consistent, incomplete, etc.) being delivered.     design   collect   data     disseminate   metadata   figure 1: current workflow iassist quarterly 2014 19 iassist quarterly three draft high-level harmonized workflows were created during the workshop to better understand the existing processes and to re-envision new processes. the standard case with metadata as a final input but not really incorporated into the system until the dissemination phase appears in figure 1: then an ideal workflow with a data and metadata done in sync was brainstormed: the group discussed potential delays in the design phase, benefits from having data and metadata being generated simultaneously though remaining conceptually separate, and how useful the workflow shown in figure 2 would be to the end-user – for example, if anes specified the type of metadata they wanted from their data collector, it might not be formatted in a useful way for the researcher. a third version of the workflow was offered that brought paradata and auxiliary data into the process: 6. capturing data transformations under the auspices of the metadata portal project, a second meeting was convened to discuss the creation of a tool that would update ddi xml metadata when the associated data file changed. such a tool would be necessary if the above future workflow were implemented. projects often make changes to the data after they are exported from cai software, and there is a need to keep the metadata synchronized as data are transformed. also important is recording the provenance and history of changes to the data in the documentation so that users are informed of data transformations over time.     design   collect   disseminate   metadata   data     figure 2: future workflow     design   collect   metadata   data     paradata   disseminate   auxiliary   data   figure 3: future workflow with enhancements the meeting was comprised of technologists and developers, most of whom had already created tools to work with ddi xml. a consensus emerged at the meeting that the group should develop a standardized data transformation language (sdtl) that would be neutral in terms of statistical software packages. this would enable the creation of a tool to capture data changes and publish them in the synchronized documentation. the tool could be invoked at any point in the data life cycle when a codebook was needed. good progress was made on the sdtl during the meeting and plans to pursue funding for development of the tool were communicated and discussed. conclusion and next steps project participants learned a lot during the course of the project about the challenges of converting legacy documentation to machine-actionable form. this is laborintensive work that should only be done once. the goal for the future is to capture machine-actionable metadata from the source and to have this marked-up documentation deposited in archives. we also hope to encourage others to leverage the xml documentation produced to create new tools beyond those created for the metadata portal. over the project period, the partners identified several next steps, described below, to carry on the work begun by the collaborative research project. exporting ddi xml from cati-capi programs as shown in the graphics above, the ideal workflow involves exporting ddi xml along with the data from the cai system and then maintaining them in parallel through to dissemination. one way to accomplish this goal is to require export from the cai systems when commissioning a survey. anes goes through a bidding process to identify the data collection firm for each wave of the survey, so they could stipulate in the request for bids that the data be delivered along with ddi xml documentation. gss uses an internal system at norc for data collection but could also work on an xml export. both partners expressed an interest in transitioning to this kind of process, starting with the next data collection cycle. interestingly, after the project ended a related effort -the “survey metadata: barriers and opportunities” meeting held june 26, 2014, in london -resulted in a published ddi profile for questionnaire documentation as well as a collaborative statement calling upon the survey design, production, and archiving communities to take leadership in facilitating survey metadata exchange through adoption of shared metadata standards for questionnaire and data description. the profile and statement for endorsement are available at http://www.ddialliance.org/survey-metadatareusability-and-exchange. this kind of best practice supports and 20 iassist quarterly 2014 iassist quarterly validates the vision for the future coming out of the metadata portal project. documentation data transformations work on a standard data transformation language (sdtl) was started but is not yet complete. there is a commitment from the participants in the initial meeting on this topic to continue to develop the sdtl as it is an essential foundation for tools to capture provenance and data transformations across statistical packages. a meeting to continue the work will likely take place in 2015 exporting ddi xml from anes databases a mapping for the anes codebook database has been completed, and the next step would be to use the mapping to export ddi xml from the database. this could also be done with the anes questionnaire development tool. if the questionnaire were marked up in ddi, this could serve as input to the cai process. capturing paradata the workshop involving the project partners had a strong focus on paradata – which paradata items the anes and gss capture, how they use paradata internally, and which types of paradata they make available or would like to make available. it was decided that this is a fruitful area for further exploration and collaboration. resolving versioning issues one of the key lessons learned on the metadata portal project was the importance of establishing canonical versions of the data for the two data series. anes and icpsr are co-distributors of the anes data, while norc, the roper center, and icpsr all distribute the gss. this situation with multiple distributors means that the data can easily get out of sync. as we learned during the workshop held to analyze the workflows and business processes of the anes and the gss, the gss workflow is particularly vulnerable to versions being out of sync. norc sends its final file to roper, which makes some changes to the data (related to missing values) to integrate it with their ipoll database. icpsr then gets the data from roper. norc may make changes and issue errata, but these changes are not pushed out to the other distributors. while we are not aware that this situation has resulted in divergent analytic results because of different files being used, it is possible that this is occurring and we simply do not know about it. this is a wider problem that affects any dataset with multiple distribution points. for example, icpsr and its counterparts in europe co-distribute several datasets, giving rise to the potential for different versions in circulation. with open access to data becoming the norm around the world, this problem is likely to escalate. the participants on the metadata portal project considered a range of solutions to address this issue. a simple fix is for the gss to push out notifications when there is an updated gss file available. anes integrates version numbers into its data files as variables, which is another simple solution. also needed are tighter versioning controls and rules for these series and persistent identifiers to the data to uniquely identify them. this should be coupled with online access to all versions of the data for replication purposes. to complement these fixes, creating a checksum registry to hold the authoritative version of the data could be useful. other distributors could compare their versions with the authoritative version to ensure that they have identical files. one possibility is to explore the use of the universal numeric fingerprint10 as the checksum. this is a solution that might have wider adoption. ultimately, the best solution is for the co-distributors to send users to a central source for data downloads, but this will take some time and cultural and technological changes to implement for anes and gss. resolving differences between portal metadata and anes/icpsr metadata while, as noted above, a central authoritative version of all files is the end goal, during the transition to that state studyand variablelevel metadata available on the metadata portal will differ from what anes and icpsr currently provide on their websites. we will need to address this issue with the goal of minimizing the number of different versions and confusion for users. it is likely that anes and icpsr will need to update their collections to incorporate metadata enhancements that resulted from the project. references american national election studies. http://www.electionstudies.org/. accessed october 1, 2014. general social survey. http://www3.norc.org/gss+website/. accessed october 1, 2014. generic statistical business process model. http://www1.unece.org/ stat/platform/display/metis/the+generic+statistical+business+proc ess+model. accessed october 1, 2014 inter-university consortium for political and social research. http:// www.icpsr.umich.edu/. accessed october 1, 2014. meta-sss site. http://metasss.mtna.us/. accessed october 1, 2014. metadata technology north america. http://www.mtna.us/. accessed october 1, 2014. roper center for public opinion research. http://www.ropercenter. uconn.edu/. accessed october 1, 2014. notes 1. mary vardigan is an assistant director at icpsr. email: vardigan@ umich.edu 2. darrell donakowski is director of studies for the anes. email: dwdonako@umich.edu 3. pascal heus is vice president of mtna. email: pascal.heus@ metadatatechnology.com 4. sanda ionescu is a documentation specialist at icpsr. email: sandai@umich.edu 5. julia rotondo is a senior research analyst at norc. email: rotondojulia@norc.org 6. http://www.irods.org 7. http://www.openmetadata.org/site/?page_id=362 8. http://www.basex.org 9. http://lucene.apache.org/solr/ 10. http://thedata.org/book/universal-numerical-fingerprint 42 iassist quarterly 2010 / 2011 iassist quarterly abstract in germany as in many other countries there is an abundance of experience with secondary analysis of quantitative data. in particular, the gesis ‘data archive and data analysis’ in cologne has for more than 50 years supported and promoted this tradition of social science and multidisciplinary research by providing the opportunity to use a wide range of social science data for secondary analysis. a similar picture cannot be drawn for the area of qualitative research in germany. in spite of the growing relevance of qualitative methods since the 1970s, there is no widespread culture of data sharing in qualitative research nor can one find an institution providing a user-oriented data service for qualitative material on a nationwide scale. in view of this situation the archive for life course research (allf) at the university of bremen addresses itself to the task of improving the unsatisfactory methodological and data-related conditions through planning for national archival development. as a first step, a nationwide feasibility study on archiving and secondary use of qualitative interview data has been conducted. drawing on the results of the feasibility study, this contribution reports on the culture of sharing and archiving qualitative research data in germany, the support for such a service infrastructure, already existing archiving infrastructure, and last but not least, the development planning for the next two years. due to the allf’s holdings and the particular value of this kind of data, this overview of the german situation includes a special attention to longitudinal data. keywords: data sharing, archiving, qualitative data, longitudinal data, germany 1. introduction in germany as in many other countries there is an abundance of experience with reor secondary analysis of quantitative data, which includes cross-cultural or longitudinal analysis of large comparative datasets (e.g. eurobarometer, european and world values surveys, european social survey, allbus). in particular, the gesis data archive for the social sciences (formerly the central archive for empirical social research,) in cologne has for more than 50 years supported and promoted this tradition of social science and multidisciplinary research by sharing and archiving qualitative and qualitative longitudinal research data in germany by irena medjedović & andreas witzel1 iassist quarterly 2010 / 2011 43 iassist quarterly providing the opportunity to use a wide range of social science data for secondary analysis. a similar picture cannot be drawn for the area of qualitative research in germany. there is no widespread culture of secondary use of the existing unique and rich data of qualitative research – especially for transcripts of qualitative interviews – nor can one find an institution providing a user-oriented data service for qualitative material on a nationwide scale. moreover there is no systematic scientific research on the possibilities and limitations of the reuse and revisiting of existing qualitative information. these shortcomings are surprising in view of the growing relevance of qualitative social science methods since the 1970s. this is reflected in the increasing amount of qualitative data material being collected and the rapid spread of computing in the sciences, as well as advancements in the development of qualitative data analysis software. these developments, along with appropriate data services, facilitate secondary use in teaching and research. in view of this situation the archive for life course research (allf) at the university of bremen addresses itself to the task of improving the unsatisfactory methodological and data-related conditions through planning for national archival development. as a first step towards establishing a qualitative data-sharing culture and infrastructure, in a collaborative research project with the gesis data archive for the social sciences, allf conducted a nationwide feasibility study on archiving and secondary use of qualitative interview data. 2. culture of data sharing and archiving – results from the feasibility study the german research foundation (dfg) financed a cooperation project for allf and the gesis data archive to explore the feasibility of a service infrastructure for qualitative research and to examine the desirability of such infrastructure among the scientific community. the feasibility study, carried out in 2003-5, aimed to explore whether and to what extent social science researchers can be considered as potential data depositors, on the one hand, and future reusers of qualitative data for research and academic teachings, on the other. for this purpose, it combined a nationwide quantitative (n=430) and a qualitative (n=36) survey of qualitative researchers, using the results to inform the criteria and concepts for archiving qualitative data. the importance of establishing an archive became immediately apparent from the fact that much research data is in danger of becoming lost. the feasibility study sought to identify the whereabouts of research data from about 1,100 german projects with a total of 80,000 qualitative interviews. the results were re-assuring at first glance: data had been lost from only 13% of all reported projects. but taking into consideration that 60% of the reviewed projects had just finished in 2003-2004 or were still ongoing and that the period under review comprised only the last ten years, the amount of unrecoverable data is already substantial. given the situation described above, it seemed surprising that data from roughly one quarter of the projects was described as already archived. however, further inquiries through expert interviews carried out as part of the feasibility study revealed that material described as archived had simply been stored in a room in their institution, which does not fulfil the basic standards of a professional archive. that often means that only original audio tapes or partly transcribed interview texts exist, the material is often not anonymised, it is kept with inadequate physical security, and that there is no public access to data, or accompanying documentation or cataloguing. besides the feared loss of important empirical data, the lack of an archive hinders the development of a culture of secondary analysis in qualitative social research. it is not surprising that reuse, especially of other researchers’ data, happens rather infrequently. for instance, results of the feasibility study show that more than one third of the respondents have experience reusing qualitative data, but that in the majority of cases (56 %) this is reusing their own data. a further 20% of secondary use involved data from other sources, e.g. collected by colleagues. most respondents argued there was no reason or special cause to carry out secondary analysis, indicating that there is little actual experience with the reuse of qualitative data and thus, there is very likely to be a misconception of the advantages of secondary analysis. the findings concerning the under-utilisation of secondary analysis and lack of insight into its potential as a research method were re-inforced through the expert interviews where a widespread tentativeness or lack of knowledge of the method or the preconditions for its use was expressed. this revealed the need for clarification on the value and purpose of reuse in the context of archival work. a second group of 18% of the respondents referred to an existing demand for data for secondary analyses and the lack of an archive. this group had not had experiences with reusing qualitative data, either because adequate material was not available or accessible, or because they did not know where to find it. they had at least implicitly considered such an approach, but failed due to the absence of archive or lack of information about reusable data. despite existing uncertainty, lack of knowledge and scepticism concerning the opportunities and advantages of reusing qualitative data material, 80% of the respondents were in favour of the idea of building up an infrastructure for archiving their research as a source of qualitative data in germany. part of the feasibility study was also to take stock of qualitative material in germany. analysis showed a large number of projects based on qualitative interviews, with 60% of the project leaders willing in principle to pass on their data to others for reor secondary analysis. moreover, 65 % of the respondents could imagine conducting secondary analysis in the future. “just taking the number of project managers interviewed in the feasibility study who signalled a willingness to give their data to an archive, this already adds up to more than 400 data sets which in principle could be archived and thus potentially could be made available for secondary use to the scientific community. over 60 % of these datasets derive thematically from sociology, political science and educational research and, according to the primary investigators, they are to a high extent usable for further research projects (90 %), and for teaching and dissertations (in each case 75 %).” (opitz & mauer, 2005, chapter 12, translated from the german) 3. culture of data sharing and archiving – the (new) scientific debate compared to the situation some years ago, there is a new emerging debate in the german scientific community about data sharing and secondary analysis of qualitative research data. more methodological work and advice on secondary analysis has been sought (lüders 2005), archiving and re-analysis as means for ‘verification’ is being discussed (reichertz 2007a, b; flick, 2007; eberle, 2007), and a handbook on 44 iassist quarterly 2010 / 2011 iassist quarterly qualitative methodology deals with secondary analysis for the first time (medjedović, 2010). this development is mainly a result of (1) the feasibility study which increased awareness, and (2) of our own contributions by publications in scientific journals and books , presentations at national and international symposia, and last but not least an annual workshop on secondary analysis of qualitative data at the ‘berlin meeting on qualitative methods’ , the main annual event on qualitative research methods in the german-speaking area. 4. qualitative longitudinal data as the feasibility study has shown, there is a well established qualitative research culture in germany. however, tracking individuals over time via longitudinal research designs is not a widespread practice among qualitative researchers, although qualitative longitudinal (ql) research is not new (witzel, 2010). the collaborative research centre 186 (sfb 186)7 “status passages and risks in the life course” (e.g. heinz 2001) at the university of bremen (1988-2001) was a landmark in the history of german research. the sfb 186 was a research programme with longitudinal projects on different transitions and status passages in the life course. it is remarkable that most of the sfb 186 projects carried out mixed method research, combining quantitative and qualitative methods during the research process. in a time frame of more than 12 years some projects interviewed their respondents up to five times. in the course of the sfb 186, allf was founded upon the recommendation of the german research foundation (dfg) to secure and make available the extensive qualitative and predominantly longitudinal data material to prospective users. as of now, allf holds approximately 700 qualitative interview transcripts (digitised, anonymised, and documented) from the sfb 186. data from a further six longitudinal studies are not processed yet and remain in paper and audio format. a rough overview of ql studies in germany shows a relatively large number of existing longitudinal and even panel studies in the social sciences using qualitative interviews, often in a mixed methods approach. taking our studies in the archive together with those from a retrieval in the databank of the gesis (http://193.175.239.23/owsbin/owa/r.gen_form), we found 36 studies (n>10) in the last ten years. some of these studies have a rather long duration. for instance, the hamburg biographical and life-course-panel (friebel et al. 2000) started in 1980 with the first wave (n=252) and finished with the seventeenth wave (n=138) in 2006. 5. existing qualitative archiving infrastructure currently in germany there are only a few decentralised archives for qualitative materials which partly concentrate on specific topics like psychotherapy, biographies in transition, political culture, social movements, documentations and party manifestos, memoires of war and post-war time, or natural and environment-protection history, as well as archives for different types of qualitative material like oral-history data, letters, photographs, diaries, biographies, essays, correspondences as well as audio and video recordings. though in most german archives it is possible to search electronic catalogues, up to now most of the data itself has not been accessible in a digital or machinereadable format. furthermore, the archives lack basic standards of data management and preservation, and therefore are not visible to prospective users from the research community. some of the archives are affiliated with university departments, and many are non-profit associations. up to now, there are no policy or procedure materials between these qualitative archives that could be shared. concerning long-term preservation and long-term availability of digital resources in general, nestor – the german network of expertise in digital long-term preservation – is concerned with these issues and provides guides and workshops for libraries, archives, museums and other institutions and individuals involved in long-term preservation and archiving of digital resources (see, http://www.langzeitarchivierung.de/). 6. development planning based on the results of the feasibility study, allf intends to establish a central national service organisation for archiving and disseminating qualitative data (qualiservice). though centralised, this service infrastructure will also utilise the benefits of specialised resources and archiving, whether by integrating and supporting already existing archives or by thematically and methodically centred data acquisition for our own data holdings. this includes also special attention to longitudinal data which are of particular value for reuse. as cooperation with experienced archives such as the gesis-data archive is indispensable, a conjoint application for building up qualiservice within the gesis-leibniz institute for the social sciences – the german institution for social science infrastructure services – was submitted in october 2008. unfortunately, due to the new restructuring of gesis and therefore new foci of activity, this application was not successful at that time. development priorities thus, our priority for the upcoming years is to realise the development of basic infrastructure for qualiservice and corresponding data management standards locally at the university of bremen. 1. therefore, we have applied for project funding at the dfg8 together with the escience-insitute 9 (at the university of bremen), the library of the university of bremen (staatsund universitätsbibliothek bremen)10 and the gesis data archive for the social sciences. 2. furthermore, we are making contributions to the scientific debate on archiving and secondary analysis of qualitative data by presentations at national and international symposia as well as publications in scientific journals and books. to demonstrate the potential of reusing data as well as to meet unsolved questions and objections surrounding secondary analysis, a research project started to conduct an exemplary secondary analysis that combines several relevant qualitative studies in a thematic field of family/partnership/gender. 3. last but not least, we have initiated a network for supporting existing specialist qualitative archives in germany (initiativgruppe qualitativer archive – langzeitarchivierung und erschließung qualitativer dokumente und daten)11. the idea results first of all from the enlargement of the data holdings through further qualitative data (e.g. group discussions, texts, documents) and through improvements in the accessibility of data from other german-speaking archives through a central mediation function between those archives and researchers searching for appropriate data. barriers 1. in germany the discussion about the need for a policy for archiving qualitative data is a rather recent phenomenon and includes contributions of the bund-länder-kommission (2006), the iassist quarterly 2010 / 2011 45 iassist quarterly bundesministerium für bildung und forschung (bmbf 2009) and the german research foundation (dfg 2006; kluttig 2008). the realisation first of all depends on overcoming the lack of cooperation between academic experts, libraries, archives, and information professionals. similar to the deutsche initiative für netzwerkinformation (arbeitsgruppe “elektronisches publizieren” dini 2009), a report of the alliance of german science organisation (2008, s. 6) suggests that cooperating academic and information specialists should develop technical standards and define the division of labour related to the process through pilot projects. this should then facilitate the establishment of reliable and accessible archives for primary research data as well as the creation of international, interdisciplinary, and inter-operable access interfaces. the german council of science and humanities (wissenschaftsrat 2011) which provides advice to the german federal government and the state (länder) governments on the structure and development of higher education and research argued recently, that the qualitative social sciences and the humanities should have a comparable position of the development of science-infrastructure like the quantitative social and economic sciences. 2. unlike the esrc in great britain, the german research foundation (dfg), the predominant funding organisation in the social sciences in germany, does not require researchers to offer copies of their data to an archive within three months after the funding period has expired. there is no national policy that mandates archiving and sharing of research data. although the german research foundation (deutsche forschungsgemeinschaft, dfg), since 1998, recommends data “shall be securely stored for ten years in a durable form in the institution of their origin” (dfg 1998, recommendation 7), the responsibility for the data rests with the individual researcher.12 often the data is stored in offices or at home, where as a rule it is not accessible for others and where the long-term storage is uncertain. 3. the splitting of the methods section of the german sociological association (dgs) into a qualitative and a quantitative branch illustrates the fierce competition between qualitative and quantitative research. also the current situation of allf is indicative: unlike in the uk, the usa, or finland, allf is not part of a national data archive, collecting and offering both qualitative and quantitative data. in view of the fact that there are many mixed methods studies, allf tries to overcome this gap and facilitate access for users by providing a simplified and improved reference system. this is why allf from the very beginning has tried to establish cooperation with the gesis-leibniz institute for the social sciences. assistance of existing organisations a good deal of work in building up qualiservice will require establishing standards for the management of data (and production of metadata) throughout the data archiving life cycle. for this purpose we hope to rely on already existing expertise, e.g. provided by cessda and iassist. also we appreciate the cooperation with esds qualidata13 timescapes archive14 and the institute for qualitative research, freie universität berlin/ina 15. references alliance of german science organisation (2008). priority initiative ‘digital information’.[online]. available athttp://www.dfg.de/ forschungsfoerderung/wissenschaftliche_infrastruktur/lis/download/allianz_initiative_digital_information_en.pdf. [accessed 22nd june 2009 ] arbeitsgruppe „elektronisches publizieren“ dini (2009). positionspapier forschungsdaten. dini-schriften 10. [online]. available at: http:// edoc.hu-berlin.de/series/dini-schriften/2009-10/pdf/10.pdf. [accessed 15th march 2010 ] berlin declaration on open access to knowledge in the sciences and humanities (2003). [online].available at: f: http://oa.mpg. de/openaccess-berlin/berlindeclaration.html. [accessed 22nd june 2009] bund-länder kommission (2006). ‚neuausrichtung der öffentlich geförderten informationseinrichtungen, abschlussbericht‘. in: materialien zur bildungsplanung und forschungsförderung, heft 138. [online]. available at: http://www.blk-bonn.de/papers/heft138. pdf: [accessed 22nd june 2009] dfg (1998). recommendations of the commission on professional self regulation in science. proposals for safeguarding good scientific practice. [online]. available at: http://www.dfg.de/aktuelles_presse/ reden_stellungnahmen/download/self_regulation_98.pdf.[accssed 22nd june 2009] dfg (2006). dfg-schwerpunktinitiative „digitale information“. stichwort: primärdaten. [online]. available at: http://www.dfg.de/ forschungsfoerderung/wissenschaftliche_infrastruktur/lis/digitale_ information/primaerdaten/index.html.[accessed 22nd june 2009] eberle, t. (2007). ‘die crux mit der überprüfbarkeit sozialempirischer forschung. forschungspragmatik vs. elaborierte methodologische gütestandards’. erwägen – wissen – ethik, 18(2), 217-220. flick, u. (2007). ‘diversifizierung, güte und kultur qualitativer sozialforschung’. erwägen – wissen – ethik, 18(2), 222-224. friebel, h, epskamp, h knobloch, b,montag, s, and toth, s (2000). bildungsbeteiligung: chancen und risiken. eine längsschnittstudie über bildungsund weiterbildungskarrieren in der „moderne“. opladen: leske + budrich. german council of science and humanities (wissenschaftsrat). (2011). informationsinfrastrukturen für die wissenschaft – eine öffentliche aufgabe, pressemitteilung vom 31.01.2011: http://www.wissenschaftsrat.de/index.php?id=345&l= heinz, w. (ed.) .(2001). statuspassagen und lebenslauf (4 bände). weinheim: juventa. kluttig, t. (2008). bericht über das dfg-rundgespräch „forschungsprimärdaten“ am 17.01.2008. [online]. available at: http:// www.dfg.de/forschungsfoerderung/wissenschaftliche_infrastruktur/ lis/download/forschungsprimaerdaten_0108.pdf.[ accessed 22nd june 2009] lüders, c. (2005). ‘herausforderungen qualitativer forschung’. in: uwe flick, ernst von kardorff & ines steinke (eds.). qualitative forschung. ein handbuch. reinbek: rowohlt, 632-642. medjedović, i. (2007). ‘ sekundäranalyse qualitativer interviewdaten problemkreise und offene fragen einer neuen forschungsstrategie. journal für psychologie, 15(3).[online]. available at: http://www. journal-fuer-psychologie.de/jfp-3-2007-6.html. [accessed 15th march 2010] 46 iassist quarterly 2010 / 2011 iassist quarterly medjedović, i. (2010). ‘sekundäranalyse’. in mey,g.and mruck, k. (eds.). handbuch qualitative forschung in der psychologie. wiesbaden: vs verlag für sozialwissenschaften. 304-319. medjedovic, i. (2011). ‘secondary analysis of qualitative interview data: objections and experiences. results of a german feasibility study’. forum qualitative sozialforschung / forum: qualitative social research, 12(3), art. 10. available at: http://nbn-resolving.de/ urn:nbn:de:0114-fqs1103104. [accessed 04th november 2011] medjedovic, i. and witzel, a. (2007) ‘secondary analysis of interviews: using codes and theoretical concepts from the primary study’. forum qualitative sozialforschung / forum: qualitative social research, 6(1). available at: http://www.qualitative-research.net/fqstexte/1-05/05-1-46-e.htm. [accessed 15th march 2010] medjedović, i. and witzel, a. unter mitarbeit von mauer, reiner, mochmann, ekkehard & watteler, oliver (2010). wiederverwendung qualitativer daten. archivierung und sekundärnutzung qualitativer interviewtranskripte. wiesbaden: vs verlag für sozialwissenschaften. opitz, d. and mauer, r. (2005). ‘erfahrungen mit der sekundärnutzung von qualitativem datenmaterial – erste ergebnisse einer schriftlichen befragung im rahmen der machbarkeitsstudie zur archivierung und sekundärnutzung qualitativer interviewdaten’. forum qualitative sozialforschung / forum: qualitative social research. 6(1 ). [online]. available at: http://www.qualitative-research.net/fqs-texte/1-05/051-43-d.htm. [accessed 15th march 2010]. : opitz, d. and witzel, a. (2005). ‘the concept and architecture of the bremen life course archive’. forum qualitative sozialforschung / forum: qualitative social research. 6(2). [online] available at: , http://www.qualitative-research.net/fqs-texte/2-05/05-2-37-e.htm. [accessed 15th march 2010]. reichertz, j. (2007a). ‘qualitative sozialforschung – ansprüche, prämissen, probleme’. erwägen – wissen – ethik. 18(2). pp195-208. reichertz, j. (2007b). ‘qualitative forschung auch jenseits des interpretativen paradigmas? vermutungen’. erwägen – wissen – ethik. 18(2). pp276-293. witzel, a. (2008). ‘basic considerations about an archive concept for qualitative interview-data’. in: maslo, i, kiegelmann, mand huber, g. (eds.). qualitative psychology in the changing academic context. qualitative psychology nexus vi. tübingen: zentrum für qualitative psychologie e.v., 71-76. [online]. available at:http://www.qualitativepsychologie.de/files/nexus_6.pdf. [accessed 15th march 2010] witzel, a and mauer, r.. (forthcoming) ‘basic considerations about an archive concept for qualitative interview-data’. in: dargentas, m, brugidou,m le roux, d and salomon, a. (eds.). l’analyse secondaire en recherche qualitative : une nouvelle pratique en sciences humaines et sociales. paris: lavoisier. witzel, a. (2010). ‘längsschnittdesign’. in: günter, m. and mruck, k. (eds.). handbuch qualitative forschung in der psychologie. wiesbaden: vs verlag für sozialwissenschaften.pp 290-303. notes 1. contact: irena medjedović, imedjedovic@uni-bremen.de, institute labour and economy (iaw), university of bremen, germany, http:// www.iaw.uni-bremen.de; andreas witzel, awitzel@bigsss.uni-bremen. de, archive for life course research (allf), university of bremen, germany, http://www.lebenslaufarchiv.uni-bremen.de/. 2. http://www.gesis.org/en/institute/gesis-scientific-sections/ data-archive-for-the-social-sciences/ 3. project “archivierung und sekundärnutzung qualitativer daten – eine machbarkeitsstudie”, 2003-2005. team: prof. karl f. schumann, dr. andreas witzel, irena medjedović, diane opitz, britta stiefel (bremen); prof. wolfgang jagodzinski, dr. ekkehard mochmann, reiner mauer (cologne). for more details see: medjedović 2007, 2011; medjedović & witzel 2010; opitz & mauer 2005; witzel & mauer 2011. 4. reusing other researchers’ data seems to be less common than the quantitative results indicate, as not all of those cases of stated reuse in the questionnaire turned out to be such in the face-to-face interview. this points out an unfamiliarity with secondary analysis and its definition. 5. for details on the publications of the allf members see the publication list at: www.lebenslaufarchiv.uni-bremen.de. 6.. see: http://www.qualitative-forschung.de/methodentreffen/ 7. collaborative research centres are long-term university research centres in which scientists and academics pursue ambitious joint interdisciplinary research undertakings. 8. within the dfg funding programme ‘scientific library services and information systems’ (lis) 9. see: http://www.escience.uni-bremen.de 10. http://www.suub.uni-bremen.de 11. members: allf; interviewarchiv „jugend im 20. jahrhundert“, posopa e.v., neu zittau bei berlin (http://www.posopa.de/ html/02_unsere_einrichtungen/023_interviewarchiv.htm); archiv deutsches gedächtnis, institut für geschichte und biographie der fernuniversität hagen, lüdenscheid (http://www.fernuni-hagen. de/geschichteundbiographie/deutschesgedaechtnis/); archiv kindheit-jugend-biographie (akjb), siegener zentrum für kindheits-, jugendund biographieforschung (size), universität siegen (http:// www.uni-siegen.de/fb2/size/ueber_size/archiv_kindheit_jugend_ biographie.html?lang=de). 12. there is an open access movement which discusses applying the open access principles also to data (see: berlin declaration on open access to knowledge in the sciences and humanities 2003; alliance of german science organisations, priority initiative ‘digital information’ 2008). 13. http://www.data-archive.ac.uk/ 14. http://www.timescapes.leeds.ac.uk/the-archive/ 15. http://www.qualitative-forschung.de/institut/ piredeu iassist quarterly summer 2010 19 iassist quarterly a user-driven and flexible procedure for data linking by cees van der eijk and eliyahu v. sapir1 abstract over the past decades, social scientists enjoyed a rapid increase in the availability of various types of data. many of these pertain to different aspects of the same multifaceted phenomenon. in the field of electoral studies, for example, there are data about citizens, political elites, party manifestos, media and the context within which elections take place. researching such multifaceted phenomena requires this diverse information to be analysed jointly, rather than separately. that, in turn, requires the linking of separate datasets (also known as data fusion, or conflation), which is largely a relational database (rdb) management problem. in spite of its large and rewarding potential for empirical research, data-linking is practiced relatively little in the social sciences. this is largely caused by lack of relevant training amongst social scientists to cope with the methodological and technical difficulties involved in data linking. this article presents an approach for facilitating data linking without requiring additional training of researchers. it defines necessary rdb operations as a structured series of user-choices, to be included in a user interface that generates and implements the rdb operations once all choices have been made. the advantage of this approach over the public dissemination of integrated datasets constructed by ‘experts’ is that it does not assume that ‘one-size-fits-all’; it is flexible and tailored to the needs of end-users. this approach can be applied in a wide variety of contexts. an implementation is under development for the piredeu project in the field of comparative electoral research. . keywords: data linking, data integration, relational database, piredeu, electoral research. introduction: an embarrassment of riches compared to only a few decades ago, comparative social researchers now enjoy the availability of a wealth of datasets. if they are interested in behaviours, attitudes and orientations of citizens, they will find that in addition to sundry ad hoc surveys, many regularly conducted national election studies, general social surveys, household, labour-market, crime, and other surveys are available in an increasing number of countries. moreover, they will find an ever growing number of explicitly comparative surveys that are repeatedly conducted in multiple countries, thus enabling comparisons across national contexts as well as over time. these include, amongst many others2 , the comparative study of electoral systems (cses, covering up to 38 countries), the world value studies (wvs, up to 87 countries) and european value studies (evs, up to 45 countries), the international studies of political psychology (ispp, up to 45 countries), the european social surveys (ess, up to 31 countries), and the european election studies (ees, up to 27 countries). this wealth of empirical material is not restricted to any particular social discipline, nor is it only made up of mass surveys. for reasons that will become apparent later in this paper, we are particularly interested in studies relating to elections, parties and public opinion, but our observations of that domain can easily be generalised to other large areas of social scientific inquiry. when reviewing our own field of interest, we find, increasingly, that in addition to the many available mass surveys, political and social elites of various kinds are surveyed as well, in single countries or comparatively across a number of countries. beyond the domain of surveys, we find data derived from party manifestos covering almost all parties that ever competed in democratic general elections after world war ii. these have more recently been complemented by the euro-manifesto program that codes the contents of the manifestos of all parties that have ever competed in direct elections of the european parliament (ep). content analyses are also increasingly used to generate systematic data about media communications, and include projects the promise that these data hold is much greater if they can be linked to each other 20 iassist quarterly summer 2010 iassist quarterly that yield data comparable over time and across countries. other extensive data sets have become available for yet other organisations and institutions, such as social movements and pressure groups. at the level of states, there is also an abundance of data pertaining to sundry economic indicators, political and social indicators, formal institutional arrangements, government performance, and so forth. each of these data collections by itself provides rich possibilities for empirical research, and indeed we see increasing numbers of publications making use of this potential. yet the promise that these data hold is much greater if they can be linked to each other. in recent years, many of the principal investigators of these large data collection efforts have referred to this larger potential as part of the justification for investing in these costly enterprises. indeed, many of the most interesting questions in social research do not pertain only to citizens, or only to elites, or only to media, and so forth. rather, they have to do with the interactions between various kinds of actors, organisations and institutions, which are affected by the characteristics of different contexts, or with the social, political and economic consequences of these contextualised interactions. as a case in point, questions about the quality and functioning of representative democracy pertain simultaneously to citizens, political parties, political elites, and mass media, among other things. seemingly simple concepts such as accountability and representation relate to the interaction between citizens and (political) elites, as well as various types of processes that involve the media (for example, the effects on citizens and elites of agenda-setting, framing, priming, and spin and hype). to the extent that the functioning of representative democracy is affected by economic developments, all of these interactions and relations have to be contextualised in economic terms. from a dynamic perspective this leads to questions as to whether ‘the economy’ is an autonomous factor affecting the behaviours and interactions of the various actors in democratic processes, or whether it is endogenous and the consequence of these behaviours and interactions. important questions that require empirical information from a variety of different actors, groups, organisations, institutions and contexts exist in all social sciences. they may focus on, for example, social integration, crime, traffic and mobility, or the efficiency of markets, but they all have in common that they cannot be adequately addressed using information pertaining to only one of the interacting actors and institutions. our ability to address important multi-facetted questions has not only been increased by the availability of abundant relevant data, but also by advances in multivariate analytical methods, software and affordable computing power. complex models that until recently were beyond the computing infrastructure available to most researchers can currently be estimated on standard personal computers using generally available software. of particular importance for empirical social research are the advances in ordinaland nominal-level multivariate analysis, latent structure modelling, structural equation modelling, dynamic modelling and multi-level methods. in view of the wealth of relevant data and the availability of tools to analyse them, one might expect current social science literature to abound with publications that join together information from multiple data sources in order to more effectively address the important and broad-ranging questions referred to previously. yet, such publications are fewer than one would expect,3 which creates a somewhat puzzling and embarrassing situation. diagnosis in principle, many separate datasets can be linked in ways that would allow important research questions to be addressed in more powerful ways than analysing each of these resources separately. if these separate datasets relate (by way of their units or their variables) to the same real-world objects, such as countries, political parties, media outlets, and the like, they can be seen as component parts of relational databases (rdbs). the methodology of rdbs is well developed and provides a multitude of ways to generate joint information from different components that provides a richer base for analyses than the sum of the separate parts. moreover, rdb software is widely available. why, then, do we see so very few efforts of integrating or linking different datasets? a number of factors contribute to this state of affairs, including the following (without claiming to be exhaustive): a) lack of harmonization. linking of separate datasets in a rdb requires the same objects being identified in the same way in each of them. frequently this is not the case. as a case in point, identification codes of political parties differ more often than not between successive editions of a series of national election studies, each of which pertains to a different election. even data infrastructures that pride themselves on their over-time comparability of coding often fall short in this respect.4 this problem is not limited to the identification of parties, but also to countries, regions, media outlets, and so on.5 as a consequence, any linking has to be preceded by a complex and costly data harmonization stage. without dedicated resources it is impossible for most researchers to undertake such projects, and funding agencies see little glory in providing grants to produce such ‘continuity’ datasets. the few that do exist are not extended when new studies are released and become outdated, thus losing their relevance. analysts that do aspire to the simultaneous use of data from different sources thus have to make their own tailored ‘solution’, which is often too narrowly focused to suit the needs of others. it is therefore not surprising that, when faced with such obstacles and with publication requirements from their own universities, many opt for the short-term option to analyse a single dataset and forego the potential riches that could be gained in the long term from data-linking. b) limitations of ‘statpacks’ and other statistical software. much of the statistical software used in the social sciences is bundled in so-called statpacks such as spss, sas, stata and r. these packages contain a wealth of statistical procedures, as well as extensive procedures for data management, such as recoding and creation of new variables. yet, they are fundamentally geared towards the analysis of ‘flat’ rectangular data matrices, and their capabilities for managing rdb information range from non-existent to extremely limited and cumbersome. as a consequence, after separate databases have been harmonised and linked in a rdb structure, dedicated rdb management tools must be used to generate the (rectangular) data matrices that lend themselves to analysis with the analytical software social researchers have been trained to use. this not only requires additional work, but requires working with software that is often unfamiliar to many researchers in the social sciences. c) lacunae in social science research training. research training in the empirical social sciences traditionally focuses on questions of general research design, data collection methods, and multivariate statistical data analysis with statpacks and similar software. many researchers, therefore, are well-versed in advanced multivariate modelling procedures, yet woefully untrained in recognising the potential benefits of rdb in managing data productively. when confronted with multiple datasets, it seems that many otherwise excellent researchers see only a collection of separate datasets, each of which can be analysed in sophisticated ways, but individually. what they often do not see are the ways in which separate datasets iassist quarterly summer 2010 21 iassist quarterly can be linked in a rdb, which, in turn, can be used to generate new rectangular data matrices that provide better platforms for addressing the substantive research questions they wish to pursue. none of these three obstacles to linking and merging data from different sources is insurmountable, yet they are not easily overcome as they are rooted in entrenched traditions of training and acquired routines. unleashing the full potential of linkable data thus requires more than just pointing out the benefits to be gained. it also requires infrastructure, in the form of software that offsets the lack of rdb familiarity amongst social scientists and that can be used without extensive further training. this paper describes an attempt to work around the obstacles that often prevent analysts from linking data from various sources. although this attempt is, in principle, not limited to any particular kind of research problem, or to any particular collection of datasets, we nevertheless present it in the context of its development, the piredeu program. moreover, as our work is still ongoing, our presentation in this paper relates to ‘work in progress’, with some parts having been developed already, and others still under development (see also van der eijk and sapir 2010). piredeu piredeu is a program of research funded by the european union (eu) under the seventh framework programme from 2008 to 2011. this three-year design study assesses the feasibility of upgrading the existing european election studies to a research infrastructure for studies into citizenship, political participation, and electoral democracy in the european union. the scientific and technical feasibility of this infrastructure is elaborated by means of a pilot study conducted in the context of the 2009 elections to the european parliament.6 in contrast to many other comparative research programs about elections that only investigate voters, or only parties, piredeu considers its subject matter, electoral democracy, as a complex set of interactions between voters, parties, candidates, media and relevant institutions (such as election rules). it therefore collects empirical information concerning different units and it does so for each one in the most appropriate manner. the various data components of piredeu are: • a voter study that conducts surveys of representative samples of approximately 1,000 respondents each from the electorates of the 27 member states of the eu.7 apart from translation and reference to country-specific institutions, the questionnaires were the same for all countries. the questionnaires covered three main themes: 1) electoral behaviour and party preferences; 2) political attitudes and orientations; and 3) background characteristics and media usage. • a candidate study that conducts surveys of the candidates listed on the ballots of the european parliament elections of 2009 in each of the 27 member states of the eu.8 with the same provisos as mentioned for the voter survey, the questionnaires were identical for the candidates from each country and from each party. • a manifesto study that consists of coding the content of the election manifestos of the political parties vying for votes in the 2009 european parliament elections.9 the coding units are sentences or quasi-sentences, each of which is classified in one of many content categories relating to a wide range of policy domains: external relations, freedom and democracy, political system, economy, welfare and quality of life, fabric of society and social groups, and european integration (cf. wüst and volkens 2003). after aggregation to the level of parties, this provides data on the relative emphasis that parties place on the substantive topics reflected in the coding categories. • a media study that consists of coding the content of media items during the three weeks leading up to the european parliament election.10 all news items were coded for the most important tv news programs and the most important newspapers in each of the 27 eu member states. each news item was coded on a large range of characteristics, including topics, actors displayed, use of frames, and physical characteristics (length, placement, embellishments, etc.). • a contextual information study that brings together information at the level of the 27 member states of the eu, regarding the results of the european parliament election and other recent elections, electoral procedures, voting rules and other relevant institutions, the incumbent government, and economic conditions.11 in line with piredeu’s perspective on electoral democracy, these various datasets are seen as providing complementary information about a complex and multifaceted reality. and although it is perfectly feasible to analyse each of the datasets in isolation, one of the main tasks of the program is to link or integrate them in user-friendly ways, thereby promoting a fuller utilization of the joint potential of the data, and, ultimately, better research on electoral democracy in europe. in order to achieve this objective, great care was given to assuring that the variables by which the data from the different components can be linked – the ‘keys’ in rdb terminology – were coded in exactly the same way in each component. this relates in particular to the identification of countries, political parties and media outlets. all components are structured by country and all relate to political parties. the voter study includes questions about party choice, party preference and perception of parties; the candidate study explores the relationship between candidates and the party for which they are listed on the ballot, and other questions about their own and other parties; the figure 1. the piredeu relational data base 22 iassist quarterly summer 2010 iassist quarterly manifesto study explores the relationship between each manifesto and its author (i.e., party); the media study examines coded news items in terms of parties being mentioned and evaluated; and, lastly, the contextual data include the identification of the party/parties that form the incumbent government, and the results of various elections in each of the countries. the identity of media outlets is crucial in the media study to identify the outlet from which each coded item was taken, and in the voter study to identify media outlets that respondents use for their information. these keys define the primary relationships in the piredeu rdb, as illustrated in figure 1. the solution that we chose for allowing data from the various components to be linked in a rectangular data matrix that can be analysed with the statistical software mostly used in social science research will be described below, but to clarify the task at hand we first present an example of a substantive research question that requires information from all components to be merged into a single analysable file. a substantive example of the need for linking data from piredeu components suppose we are interested in the orientations and behaviour of candidates, and, more particularly, in how salient various issues are for candidates. this information is specific for the candidates, and available from the candidate survey. were we to analyse issue salience for candidates only from the data in the candidate survey, we would model the variance in the dependent variable (salience of issues for candidates) in a multi-level model with candidates nested within parties, which, in turn, are nested within countries. as independent variables at the candidate level we can use all other variables collected in the candidate survey, including various political orientations and background characteristics, and the identity of the party and the country of each candidate. in the absence of further information about the traits of these parties and countries, their impact would be modelled in a random effects specification. such a random effects model would be unsatisfactory because it would only tell us that some of the variance in issue salience at the level of individual candidates can be attributed to parties and to countries, but we would be in the dark as to the form of this relationship. that is, which kinds of parties and which kinds of countries have a positive or negative impact on issue salience? the model could be made more informative by adding information about parties and about countries, thus allowing a mixed model specification in which the explicated characteristics of parties and countries are modelled as fixed effects and the remaining variance at that level as random effects. this additional information can be derived from the other data components of piredeu.12 in the absence of any infrastructure or specific tools, such information would have to be added manually by the analyst. from the manifesto study one could, for example, derive how much emphasis parties place on each of the issues in their manifestos. this information would be added to the candidate dataset via a tedious procedure involving a large number of conditional statements. with some 200 political parties this procedure would be prone to error, and would take several hours to accomplish even for an experienced data analyst. in a similar way, one can add country information to the candidate file (which would be somewhat less onerous as there are only 27 countries). additional information can be added that originates from the voter study or from the media study. the work involved becomes increasingly more complex if the theories that we want to test involve a wider set of relationships between candidates, parties, voters, media and contexts. consider the following elaboration, which, although used here for its illustrative value, would substantively be neither unrealistic nor excessively complex. the dependent variable – the salience of various issues for candidates – can be seen as a function of: • the difference between the candidates’ personal views on issues and those of his/her party as expressed in its manifesto. this expectation reflects partly the tendency to reduce cognitive dissonance and partly the political expediency of downplaying differences in views between oneself and one’s party. to test this, it would be necessary to add information from the manifesto data to the candidate dataset. • the salience of the various issues for various media outlets, moderated by the extent to which potential voters for the candidate’s party are exposed to those media outlets. this expectation could be based on the notion that, politically, candidates cannot afford to ignore issues that are played up in the media, particularly if their own potential voters are exposed to the contents of those media. to test this, one has first to arrive at a measure of issue salience for media outlets, which has to be derived by some form of aggregation from the news items that have been coded. this outlet-issue-salience then has to be added to the candidate survey using country as key, as candidates are not directly linked to media outlets. subsequently, the voter survey has to be used to distinguish potential voters for each of the parties, and then to determine for each of these groups the extent to which they are exposed to media outlets. this information has to be added to the candidate dataset using media outlet and party as keys. • the salience of various issues for voters, moderated by their propensity to vote for the party of the candidate in question. this would be based on the expectation of candidates being responsive to voters in general, and in particular to voters who are likely to vote for their party. to implement this, one would have to use the voter study to distinguish voters according to their propensity to vote for each of the parties, and then to assess how salient the various issues are to each of these groups. this aggregated information would then have to be added to the candidate survey using party and country as keys. • the effects of the factors listed in the previous bullet points are potentially moderated by country-contextual factors, such as the (temporal) location of the ep election in the domestic electoral cycle (for theoretical foundations of this expectation see van der eijk and franklin, 1996).this would require retrieving the relevant information from the contextual information dataset, and then adding said information to the candidate dataset using country as key after performing all these data operations, the candidate dataset, extended with information from the datasets pertaining to manifestos, voters, media and contexts, can be used to perform the desired multilevel mixed effects regression with cross-level interactions. again, without specific tools or infrastructure to help accomplish these tasks, the required data management would easily take days of errorprone and tedious work. moreover, all investments in that work would only be relevant for this particular question, and similar, but substantively different operations, would have to be performed for other substantive questions. managing the linking problem in our view, any attempt to facilitate productive linking of data from different sources (or in our particular case, from different components of the piredeu program) has to recognise the following considerations and constraints: • in terms of the outcome of the linking process – a data matrix that can be analysed by the kind of statistical software used by social scientists – there is no single or one-size-fits-all solution. what has iassist quarterly summer 2010 23 iassist quarterly figure 2. flowchart of user decisions involved in merging voter study data into manifesto study data — [mas & ◄ vs] to be linked, and how exactly, is different for various substantive research questions. any attempt to impose a single solution would be futile, as researchers will not use it if it does not fit their aims and theoretical and conceptual perspectives.13 as a consequence, facilitating linking has to take the form of providing tools to accomplish the task, and not providing a linked dataset. this holds equally for linking data pertaining to different observation units, as for linking data pertaining to the same kind of observation units (e.g., repeated studies of the same kind).14 • the outcome of the linking process must be a datafile that lends itself to analysis with the statpacks and other statistical software that is ubiquitous in social science research. as such statistical software is generally not capable of handling rdb structures, the linking process must result in a flat, rectangular data matrix. • many excellent empirical social science data analysts have not been trained in rdb management. therefore, tools to facilitate data linking should not assume such familiarity and must provide, in a structured manner, the kinds of options available. as a consequence, relevant tools must disaggregate the linking process into successive tasks of limited complexity and clarify the options available at each for the user. in accordance with these considerations, we chose to develop tools for linking data across the various piredeu data components in the form of a structured user interface that guides the user through a set of choices. the entire sequence of choices generates a syntax that specifies the required rdb operations and, when implemented, yields as an outcome the desired dataset with linked data. actually, in view of the software considerations referred to above, it produces a tailored dataset into which information from other datasets is merged. we will illustrate this approach by focusing on only two of the piredeu data components, the voter study (vs) and the manifesto study (mas) (see figure 1). linking other combinations of data components operates along the same lines and will not be elaborated in this paper.15 linking and merging the voter study (vs) and manifesto study (mas) the units in the vs are individual respondents. the primary key in the vs is the respondent-id. the units in the mas are political parties, and the primary key in the mas is the party-id. the relationship between these two datasets is defined by a number of foreign keys in the vs that relate to survey questions about parties (each of these foreign keys is coded in the same way as the primary key in the mas). these questions are about different matters, including actual party choice made in particular elections, attractiveness of parties as options to choose in a particular election, generalised affect, and perceptions of parties, among other things. what they all have in common is that their possible responses are defined in terms of parties. these foreign keys can be used for merging information from both datasets. in view of the considerations discussed in the previous section, this merging has to result in a ‘flat’ data matrix that can be analysed by statistical software packages. this merging can therefore take two different forms, resulting in two different linked datasets: (1) a vs dataset (units are individual respondents) with information from the mas merged into it, or (2) a mas dataset (units are political parties) into which information from the vs is merged. owing to the difference in the character of their units, these two forms cater to different research questions, and thus to different groups of researchers. moreover, because the keyed link between the vs and the mas is of a one-to-many type, the actual merging process is somewhat different when going from vs to mas than the other way around. merging mas data into the vs. the mas offers information about the political parties that respondents mention in their answers to survey questions. as this information is not respondent-specific, it is identical for all respondents who refer to the same party in response to a question. thus, for this information, one can regard the respondents as being nested, so to speak, in the parties they mention in their responses. merging is in this case a simple operation, as each mention of a party by a respondent relates to only a single case in the mas data. when new variables are added to the vs, they contain the desired information from the mas, linked by the correspondence between the chosen foreign key in the vs and the primary key in the mas. merging vs data into the mas. this kind of merging provides information about the composition of the groups of respondents who relate to the various parties in terms of choice, affect, particular perceptions, and so forth. this linked information may include anything available in the vs, such as respondents’ views on political issues, their social characteristics, media usage, political behaviour, and so on. merging in this case is somewhat more complex than in the previous case, as the relationship between each party and respondents generally involves multiple respondents. the variables to be added to the mas file, therefore, have to contain summarizing information about the relevant group of cases in the vs. the analyst has to decide which of various possibilities is most desirable. obviously, this is partly dependent on the measurement level of the relevant information in the vs. means and variances would be relevant for interval level 24 iassist quarterly summer 2010 iassist quarterly variables,16 but that level of information is rarely available in survey data. for ordinal level data the ordinal equivalents of these summarizing parameters are available. at nominal level, summarisation is limited to proportions in all (or only in some) of the categories. the user interface for linking thus contains a set of choices for the analyst that specify the desired mode (or modes) of summarising vs data to be merged into the mas. flow of end-user decisions. in order for the user linking interface to generate the desired tailored dataset, the analyst will be guided through a series of structured choices which are reflected in the flowchart in figure 2. as a very first step, the user has to decide on the units that are to populate the desired merged file. in the example presented here, where we consider only the vs and the mas, the question is whether we want to obtain a file of voters with merged data from party manifestos, or, alternatively, a file of parties and their manifesto data, into which voter information is merged. figure 2 has been specified for the situation where the integrated file has parties as units; obviously a very similar, yet in detail, somewhat different flowchart would specify the decisions to be taken for the choice of individual respondents as units in the integrated file. once the choice of units has been made, one should decide on the number of countries one would like to include in the integrated file. this number could range from 1 (i.e., a single country database as the final product) to 27 (i.e., all member states included). next, the user should determine which parties will be included in the integrated file. this number is in the range of 1 (a single party dataset) to all parties participating in the ep elections’ (over 200). the next step is to define which respondents to include in the information to be merged (all respondents can be used to provide the information to be merged, or a selection that can be specified in terms of group criteria and weights). then, the user should determine whether or not the variables to be merged should be recoded and, if so, how. once these steps are completed, the user needs to define the aggregation parameters he/she would like to use in aggregating the vs data into the mas data. in the user interface, the possible choices will be presented in the form of drop down lists or of user-specifiable values. linking and merging data between more than two sources in the previous section we described the linking between two sources of different data, one with survey respondents as units, the other with party manifestos as units. as illustrated in figure 1, however, the piredeu data consist of five different components. between each pair of these, the linking and merging of information proceeds along the logic described in the previous section, possibly with minor modifications necessitated by unique characteristics of each of these five components. when merging data from more than two sources we can distinguish two possibilities, one of which presents its own challenges. we use the notation [primary (or recipient) & secondary (or donor)] to denote the merging of information from the secondary into the primary dataset. the easiest way to merge data between more than two sources consists of parallel merging, which is using the same source as primary dataset vis-a-vis all other ones as successive secondary datasets. in the case of the piredeu data, this could, for example, involve the mas as primary data, into which information is merged, in successive rounds, from other sources such as the vs, the cs and the ms. in other words: [mas & vs] + [mas & cs] + [mas & ms]. the merging operation for each of these successive rounds is similar, and proceeds along the lines described in the previous section. in which order the successive rounds of merging are executed is immaterial for the final result. this example will result in a mas with a great number of additional variables which contain information from the other data components. a more complex way consists of sequential merging, where, for example, in a first round, vs data are merged into the mas, while in the second, mas information (including the data that have their origin in the vs) are merged into the: cs:[cs & [mas & vs]]. this implies that the identity of the primary dataset changes between successive rounds of merging. in these situations, the order of the successive rounds of merging is essential, because, as discussed earlier: [mas & vs] ≠ [vs & mas]. the design of the user interface for multiple parallel merging is not intrinsically more complex than it is for single merging operations. however, in the case of the interface for multiple sequential merging, it is more difficult. the additional difficulty is not technical, but didactic in nature: the user interface is intended to allow users who are not used to rdb management to merge data from different components of a rdb by guiding them through a series of structured questions, the answers to which generate the syntax of the required rdb operations in the background. the challenge will thus be to implement possibilities for sequential merging, while keeping the interface simple and comprehensible concluding remarks we believe that our approach to producing ‘integrated’ datasets has a great advantage over other approaches. it does not invest in the production of a specific end-product, but rather in the tools to be used that will allow such a product to be achieved. the implication is that, once the tools are produced, different ‘integrated’ datasets can easily be produced from the same original empirical material. or, stated differently, it does not produce a straightjacket to which researchers have to adapt themselves, but rather it allows the production of datasets that are tailored to the specific needs of individual researchers. this approach to linking and merging data is in principle also applicable for other data than those collected in the context of the piredeu program. as a case in point, many studies that are conducted repeatedly – such as national election studies – strive to unleash the potential for longitudinal comparison by making available longitudinally integrated datafiles, often referred to as so-called continuity studies. but the construction of such datasets is never straightforward, as they invariably require many decisions being made to cope with unavoidable differences in operationalisations, coding schemes, and the like. irrespective of what decisions are made, they can never be optimal for all the different research projects that would require such over-time comparable data. in these contexts too, it may be advantageous not to invest in the production of datafiles that will not be well suited for at least some researchers, but rather in a flexible interface that allows analysts to tailor the longitudinal data integration to the specific needs of their research. references blumler, jay g. 1983. communicating to voters: television in the first european parliament elections. london: sage. iassist quarterly summer 2010 25 iassist quarterly braun, daniela, slava mikhaylov and hermann schmitt. 2010. ees (2009) manifesto study documentation advance release. mannheim: mzes. [http://www.piredeu.eu/]. van der brug, wouter, and cees van der eijk, eds. 2007. european elections and domestic politics: lessons from the past and scenarios for the future. notre dame, ind.: university of notre dame press. czesnik, mikolaj, michal kotnarowski and radoslaw markowski. 2010. ees (2009) contextual dataset codebook, advance release. warsaw: swps. [http://www.piredeu.eu/]. connolly, william e. 1999. the terms of political discourse. princeton: princeton university press. ees. 2009a. european parliament election study 2009, candidate study, advance release. july/2010. [http://www.piredeu.eu/]. ees. 2009b. european parliament election study 2009, contextual data, advance release. 16/05/2010. [http://www.piredeu.eu/]. ees. 2009c. european parliament election study 2009, manifesto study, advance release. 22/07/2010. [http://www.piredeu.eu/]. ees. 2009d. european parliament election study 2009, media study data, advance release. 31/03/2010. [http://www.piredeu.eu/]. ees. 2009e. european parliament election study 2009, voter study, advance release. 7/4/2010. [http://www.piredeu.eu/]. van der eijk, cees, and mark n. franklin, eds. 1996. choosing europe? the european electorate and national politics in the face of union. ann arbor: university of michigan press. van der eijk, cees, and eliyahu v. sapir. 2010. linking electoral data about citizens, political parties, mass media and countries. a general approach with a user application for the 2009 european election studies. paper for piredeu user community conference ‘auditing electoral democracy in the european union’, brussels, 18-20 november 2010 [available at http://bit.ly/emvwj0]. van egmond, marcel h., eliyahu v. sapir, wouter van der brug, sara b. hobolt and mark n. franklin. 2010. ees 2009 voter study advance release notes. amsterdam: university of amsterdam. [http://www. piredeu.eu/]. giebler, heiko, elmar haus and bernhard weßels 2010. 2009 european election candidate study – codebook, advance release. berlin: wzb. [http://www.piredeu.eu/]. reif, karlheinz, and hermann schmitt. 1980. nine second-order national elections. a conceptual framework for the analysis of european election results. european journal for political research vol. 8, pp 3–44. schmitt, hermann, and jacques thomassen, eds. 1999. political representation and legitimacy in the european union. oxford: oxford university press. schuck, andreas, georgios xezonakis, susan banducci, and claes h de vreese. 2010, ees (2009) media study data advance release documentation. exeter: university of exeter. [http://www.piredeu. eu/]. thomassen, jacques, ed. 2009. the legitimacy of the european union after enlargement. oxford: oxford university press. wüst, andreas m., and andrea volkens. 2003. euromanifesto coding instructions. mzes working paper 64, mannheim: mzes [http://www. mzes.uni-mannheim.de/publications/wp/wp-64.pdf ]. notes 1. paper presented at 36th annual conference of iassist in the panel ‘virtual research environments: tools for presenting and storing data’, cornell university, ithaca (ny), june 1-4, 2010. please address all correspondence to: methods and data institute university of nottingham university park, law & social science building nottingham ng7 2rd, uk cees.vandereijk@nottingham.ac.uk eliyahu.sapir@nottingham.ac.uk 2. the website of the cses lists a large number of comparative data collection projects in the general field of elections, parties and public opinion, with their respective url’s; see: http://www.umich. edu/~cses/about.htm. 3. quite common, however, are publications in which separate and unconnected analyses from single datasets are ‘linked’ in a narrative fashion. some of these are excellent and generate important insights. yet, as will become obvious below, this nevertheless falls far short from linking the diverse data before the analysis stage and then analysing the merged data. 4. as a case in point, coding of parties is not fully comparable across successive editions of the european social survey (ess), mainly by not anticipating that national party systems change over time. in 2008, code 11 for the variable asking about the party the respondent voted for in the last general elections in the netherlands (prtvnl) pertains to the pvv (freedom party), while in 2002 the same code pertains to ‘other party’. such incomparabilities are particularly large in countries with instable party systems. 5. we do not want to suggest that no useful efforts at harmonization are undertaken at all. some of the most productive ones pertain to harmonization efforts aimed at making the coding of, e.g., educational attainments comparable across countries (the unesco initiated isced codes). 6. detailed information about the piredeu program, including questionnaires and coding schemes, can be obtained from its website: http://www.piredeu.eu/. as the focus of this paper is on data linking, we refrain here from presenting substantive information about the specific character of european parliamentary elections, and we refer to the relevant literature, e.g., reif and schmitt (1980), van der eijk and franklin (1996), schmitt and thomassen (1999), van der brug and van der eijk (2007), and thomassen (2009). 7. ees (2009e); van egmond et al. (2010). 8. ees (2009a); giebler, haus & weßels (2010). 9. ees (2009c); braun, mikhaylov & schmitt (2010. 10. ees (2009d; schucket al. (2010). 11. ees (2009b); czesnik, kotnarowski & markowski (2010. 12. such additional information can, of course, also be derived from external sources, such as the world bank, oecd, eurostat, and so on. for our example however we focus only on various piredeu data components. 13. this is similar to the futility of occasional proposals to ‘standardise’ the observation of essentially contested concepts (cf. connolly 1999), or to standardise questionnaire items in survey research. 14. this implies that many of the attempts to provide ‘continuity’ files for, e.g., national election studies, are suboptimal at best. in the process of producing such files many operational decisions have to be made which are not of a technical and innocuous nature, but which have conceptual and theoretical implications. if analysts do not subscribe to these implications, the resulting continuity file will be less desirable, and they have to either repeat the same work on their own terms or, more frequently, abandon the project that required such linking. 15. for full elaboration of the linking solution, see van der eijk and sapir (2010). 16. of course, many other summarising measures could also be relevant for interval level variables, such as x-tiles and x-tile ranges, skew and kurtosis, proportions above/below specified cut-off values, etc. lassist newsletter, vci. 2, nc. i (spring 197d) the nohc liebaey as infgr.iaticn center aijc data archive patrick 3ova national cpinion research center introduction tn to pr igin searc its i munit probl with tnis libra itsel many organ that which publi e ob j ovide ot t h cen nf crm y it ems the paper ry , w t, s ways, izati dart ' inf c. ective a desc he nat ter (n ation b serves encount data ar descr e will ince th the en. t of th ormatio or rip ion orc ase an ere cfli ibe als e pub this t ion al op ) ll , the d so ve. s t he o ref libra lie a ncrc agenc passe pa or t inio brar use me o th er t ry i rm libr y \s t the scribed the se classif as a c book closely search researc inf orma all ma n03c, n search ety of chive f generat it is service versity people researc about a univers data fa of the houses estmgl text. general form of rect re and gro cial sc ncrc as nse y lib onven and alig inter h pro tion nner cec s in ge publi or n ed wi def i data of c acqui h as ata 1 ity cilit sec data y eno the u soci the suit wth lence libra a spe that rarie tiona ref e ned t est grams of f i of tuaie neral cs. orc d tn n nitei arc h 1 c a g re a wel ocate maint y (wi lal from ugh i ni ver al su i] orc of t of no ry c cldl word s. 1 1 renc t of t ce w info s, , t it i ata orc i° hive ata 1 a d el ains thin sci icps n th sit y r vey lib he rc a an best librae is us it is ibrary w e coll he cur re he intr it is th hich su rmation and surv c a wide £ the da and for particip ct the for th althoug from no £ inior sewhere . a se the di ences) b and, e presen copies the p rary is recent h nd of t per is ne orn rey and t comf the cially thougn € nobc o norc £, m of the ary is hro jgh the te dey, in ea to active ith a ection nt ream ural e norc pplies about ey revarita ardata at ion local e unih many bc for nation the parate vision wnich intert conor the resent a dilatory he sohistory cf norc the history of noec includes changes in the size of nosc and in the number, complexity and subject matter cf studies undertaken. nohc was founded in 194 1 as a place where the then new technique of opinion sampling could be appxied in the neutral setting of a university. from 1941 to about 1960 the center nad a modest budget and a small staff in the cnica,.j headquarters and in a new york cit; office. norc could be characterized in that period as a researj.i center housed in an ola mansion or. campus; the staff was small enouga to "meet every day in the dinir.: room of that house ror coffee di.l cake. the during of naf of fo the u. tween these occupa shirle t udes series which by the tratio versit no tha lona reig s . 194 late tion y st towa of some ce n st ytable t per 1 sur n af f depa 5 an l\' ar st r d me healt are nter udies st lod vev3 dlf s rtme d 19 the rest udy ntal h-r e stil for (ch uuies were on t cond nt of 57 (m 1947 iqe st of pop iilne latea 1 be in healt as) a conduct ci the serie,; he conduct acted for state beore about nortn-hatt udy, the ular attiss , and a surveys of g repeated h administ the unia li serv icago nt a inten dex, fairs riod is in on te £, p e typ e tod dent ete. eful ow at out 1 e inc arch nds o brary ex e tiie p and new nd usefu ance of by subj studie from 19 dex cont xt with lus seme e and si ay, alth that it the i in acce roper c 25, and iuded , tool in f anaiys isted 1 rot ess io york. 1 activ a quest ect, of s conduc 45 to a ains th percent mf orma ze. it ough we is any tern ind ssing th enter) , since t tne fil itself, is. n tnose nal staf one im ity was ion or the for ted in pril, 1 e rull q age mar tion on is stil are not longer ex is g ese stu which nu he raargi e is a for cer days f in porthe item eign tne 957. uesginsa n1 in concomuite aies mber nals retain bes librar als u dook a relate the s have a tant h done i rather survey then tained these used f have n taem u user i the f i are pr that t laes y ma sed nd 1 d f tudi con ealt n th ext res prod , i fil rom ot p to nter les. obao ime the i n t a i in 3 ourna o the es be cise h-rel e 19 ensiv ults ucing nclud es s time m ade date est. of ly a perio gues aed tudi 1 co su ing col 1 ated 50s) e ri fro th ing till tc t an due we cou uni d. tion file es llec b jec qon ecti sur ies m ma em fore exi ime , atte to in rse, que inde s of and a tion t mat e (we on or vey r in ad of p ny a were ign s st a alth mpt t cost tend sin resou x, the mater ismall closely ter or still imporesear ch dition, oil and gencies mainour ces. nd are ough we o keep and low to kee'j ce they rce for n t w j x >_' 1 1 o r 2 (3pril;.j 1978) r. a x 1 t t .-. t 1 n were norc mac h does li i o u r n a result ether tact m i ri i n m tact t f uncti even in present norc, tion ac the 01 publ tae per academe occur re t he lib plexity creased well be the li and div sionai new sta tne lib tion ; request creasm .po. as t,.dt ti. ^tj idt: dl le ior and s j ert bee e i s -s u 1 t r. e i a as weli cpjinion r r 3 m s t zatiga th li the tiv 196 ic lod an yon tra €rs st f f rar as br it nd w as d th ear^ aria e vej. y wa n e t a s o r v t y r austi at no ^ re s :» 1 o n as p u news, w u^les oiis. d r a r y w n a t e v e r t ncfc t; st ua y 19 to i. ijegan or pub quite i 1 ur d r e a a _: 1 i z r. e ts dlt s tnat is re that txlsll .11 c.i c by nok publi as pr publi was o y air s, wh wor ki lie m low. os cna repr ese change at 'io h 1 c n ha y — t.-. studie q the t d h e a 1 1 ry cox iried. aif gre mambers y f or h public and qu ref err nged ncrc's ion, j her aspe ever ai : nsequenc , umber an n jed ntd t: d ot; pc s< opic r. st iect th w a cam isto co lit esti ed t n progre s t roil i udies, s ions ex e nosc p r. d c h a n g e to dep lical in act i n c r ens wer c the li id i., a my jn st air ; riles time lea se'o gallup ng t.ie cvared c and c coacbabiy c con1 1 e n a ector . en the rg for for raastyie ust as cts or things e s for d comss i.nerated c taat panaed rof esed end on formaeased, e mbrary. the conte.^eoraey recced or nofc the brief preceding dis has largely to ao with the i research programs at k05c. ncec founded its survey r service (srs) to i:oriiiai.ly research services, chief! collection (but includm phases of survey research), social science community. t ter had previously contrac conduct studies zoi axt sponsors (a good example stouffer communism study in conjunction with the gallup zation) , but srs was active moted as sucn. the major teristic or srs tnat sho noted is that in most cases role was tecnnica^. — we nai xittle to ao with the planning of these studies a tie to do with anaj.y2ing d reljortmg the results. tni acleristic has implications library ir. its mrormation a aichive activities, as we see . coincident with the founding oi s;is were tiie start of the great societv poverty prograns which inc^uata luiit-m evaluation appropriations. sks was available to c ussion n ternal in 196j eseardh provide y data g ail to the he cented to r amural is t .1 e 1554 in o e g a n 1 ly procnar acula be norc s usually i nitia 1 n d iitata and s c .-; a r for the no data scall carry out some of the massive data collection programs needed and became quite busy at it. the pub iences c one co iences re atten e public iped alo ans-acti ere wer cm in hi a u a t e creased the co ces -bc libra re resp rger aua lie nange uld s wante ti ve appa ng b on an e oth gher stude sophi mpute that r y a onsiv ience image or d m the ay that d the p to its r rentiy y such a ps_ych er curr educatio nt popui sticatio r i h t h had eft nd force e t o a the social 19b0s, too the social ublic to be esults, and began to be, magazines as oio^y today. en"es -"the n and in the ation, the n in the use e soft selects on the d it to be larger and one a grea the da to 196 i n g da person pe rson ail, s as muc we see to be haps. at nor dr amat chivin state roper ter in symbol activi which deemph releas ve ys w ticipa devel t impa ta arc peop ta bu o^m hiv le manne impier h qata now. more than c the ic ac g was depar public 1959. ic ot evolve asize ing t h as sig te oppo a ts or w str man fir t r th tme ok our on d d arc e s nil e r our delve nnm opos uaie d c uld ivin ovid ded f el chiv e ar the uec emph nad g an it io s, m ondi not g p e tn to t t w e (w chiv sta ision asize to dj arcn n. any q tion. get roper e run ry to ere u e hop es wh dies) ent in t on the n e movem were cer n a mor sea to a times there si demand hat ther a i g h t f o y studie s t and p elated e deposi nt stud inion r his act policy ar chivi uring th hiving, tate cep leant be evised p to try norc as with t i ve is our dat uite old the funds to ly, and ds itsel send t sable t ed to fi ich woul ue 1 9 chc l ent. tainl e pe n ag were , mply for e was r w a r d d are robab to da ting ies esear is mo (if ng a e 196 the artme cause olicy 60s had ibr ar y: prior y sharrson to ency to after was not them as tended , pertoday. ly most ta arof the in tne ch cenre than not our t norc os: to act of nt surit an, at le a data ha fact an ex pan a are and man or ganiza do the could f, so we hose stu o a log nd appro d liouse ast, arthat si ve norc y m tion arnot deaies ical primany aithougn data archiving nas been aeemphasized, we have not stopped archiving machine readable data. the old studies mentioned above — although not aix are worth equal attention -still require much work beiore we would inflict them on any arcnive. we thereiore provide sample survey data for secondary use, but we do not coileet or keep data frcm other sources, ex^5 lassist newslettdr, vci. 2, nc. 2 (spring 1978) cfcpt those which may have been acquired tor a specitic project.. just as our data archive activity is closely idtntiriej with norc production, so is our information activity. and just as data arc the result of a long line of activity, so too do our lof ormat ion services cover all ascects of these events, not only to florc staff but also to the interested public. the mously and mor only in studies conduct sponds ation o get so spond . duction to scho tereste guestio cial s within search, met hods procedu mg) a matter. our erated— acce of r e t r are wor cation subject start, tion t cases, request user as require we have are sup study d the which r teresti con trac of data 2-3 ye contact to seek serves least o the ori with s working passes sets ar usually will ge and all make th ers, i cases i best, since tape (e sur act i e pe th but ing to m n mo aeon we to cl g d cien the q (e res nd vey ve t ople e re ais them est rc p e on als ncrc roup n vi e re ces, pur uest (eg: witn move hese ar suit o in re-ju ro je th and s th si ti eiv b view ions que in par ment day e in s cf th the es ts cts e 3 rovi £ur at ng ut o] hav stio terv ticu is s, an terest thes e me t h libra for i or tr taff t de an vey re might nohc. ever t mainly surve e to d n wor iewer lar s f robl inf or ieval king of a or depe he u the ed, best d, we a co plied irect em w mati at are on s stu met ndin ser resu ana we sup pyon or p ith u on is prese prim olut 1 d y on hod g on neea its o here can. ply t in ly wi ermis sing nor the usu nt our m itive, ens. t a part is jus what in £. in f a stu we info if da hem, as any case th spon si on. enourd more ed not e many cds ot ry ren i o r m ies to c reintrosear en be inthe he sostay y rec with ding) , tramub ject c genai one ethods but we be loicular t the formamost dy are rm the ta are suming , data £or or data fo eside a ng iim ts call after ars) ,with permis useful f whic ginal cmeone in the mere an e "wil ha ppen t tired cws us e data ncludin t turns origina we will xcept t no bo. f o a our the sion pur h is anax who sam d mo led" sis of bian aval gar out 1 c no mayb xtram rc a ai r pub et t feel origi to poses to ysis is e are re o to that givin ket p lable chi ve that opy o t ha etc ural re i thou lie ime ing na 1 use lear and pr a. f th norc the g pe ermi to s. we f t ve u cop studies n an ingh most release (usually is that sponsor the data not the n about tc talk es umably as time ese data what sponsor emission ssion to all ccmin some have tne he data, sed the y a few tiiras) and since many original sponsors lack the institutional setting needed to care for tapes and documentation for any length of time. these aata sets are treatol in the same manner as norc's own data.. the i:ipact of user de.'iand we of how of the sfaouiu fiae na sc lent set so concer of doc at nor ersity which dence mentin change analys docurae turn user libra af f ec ture o i f ic e me pr ned wi umen ti c: ho or compet over t g of s s in is ha ntat io now to aem and ry (an t tha f s urv nte r pr ior iti th tw ng or w t o c in for e with he ar urveys da ta ve cr n prob af f e d pe t wor ey r ise es . o has a r c h i ope w matio and chivi , an pro eated lems. onsi cts rhap k) esea may we ic vi ng ith n tak ng a d ho cess ad deration the work s how it and how rch as a help us will be problems sur ve vs the divrecuests e precend docuw recent ing and ditionai in serv quit r r cm help soph have quir of te pass the to g on da anot list ten pear day prag sa ti tent ne o the e a e di hig wit isti ser emen n fi ive care et a her of just s t to mati sf y for f the norc l wide va verse i h schoo h a deb cated ious a ts. a nd ours way , t f ai wo data 3 use is thing things does n hat in day rou c rout those ser vie f undamen ibrary riety o nfor mati 1 stude ate topi profess nd compl s a res elves re o immedi rk that et in sh quite o to do to do , ot get a the hur tine, e only who are tai prob is that f users on need nts who c, to q ionals ex data ult we q acting, ate dema is regu ape for ften si in tne and most one. it ly-buriy we take and try most in j.ems we with s — need uite wno reuite in a nds. ired secmply long ofap-. ot the to sisim iqtl of the archivist now, everyone deserves attention, one would presume the high schooler as much as the potential user of one of our surveys (which may lie "unarchi ved" , in unusable shape) . wny should we pay special attention to the archive function when limited resources are stretched in the first place? we suggest that the scientific nature or survey research demands that special attention be paid data for secondary use. tne overriding philosophy of the norc library in its role as conservator of norc surveys is that the entire history of a survey should be preserved so 46 i a s 6 1 .s t k t w s * ci 1 1 <j r , v c 1 . i, i (.op 11 1. j 19 7d) dt .t . a 1 1 1 ;ate tory oi survuy.s " ' :ien :iai ;h s ;pea to :e wl te veys tfed, an .1 1; ap n ce 0x1 is . an 1 lie a t;. 1-= i.d lo rfcij uii cfc! is tc i ^y in that aagests the table. ( jver te time is a mat ion i£ needed for an v i r. if sat lect ! xpe cou > .lab ;£ib . t. one ist.ie ricse reie , rocumentatiot; and its uses idea ii e n t a t i data, b ixcat lo lished least , concius this id cai at stages uce wr our co whicn r of sarv ing in field k s p e c 1 f i publica cation en ol y t he n m findi to i ions eal 1 nor or sj 1 tt en deboo ecogn t y do forma crk a cat lo tions and c then a s way) oraer ntexl draw s no c , s rveys doca ks co izes camen tion nd t ns, and odmg £udy wou to cr , igen n f r t wh ince do men t ntai the tati on rain bibl tne ins ne rinai (not o id p e r m 1 test th at tn tly jad cm tue ciiy imp tae v in fact ation. n infor broader en by 1 the s in g, au lograpfii usual da tr uction doc jniy of t repe p jde ver y ce the data . ractiari js pr odtaus nation n a t a r e ncl udampie , est ion es of ta ioalthough the view of documentation outlined above does not solve the daily crunch, it does give a basis for deciding wnat documentation should include. it is tco bad that there is really very little pressure for good documentation even from the aost sophist icatea users. the rare exception is tne graduate student or professor working in the area of survey methodology. in t largely documen vey . traditi which m ment mi one kne at abou widcspr y s i s " s transit data pr terns, " for mat i data a older s be kep jobs to now a p p udl res ant do w o r k . ra o r e a o bv tne di v 1 .; ja le wo depe ;at 10 i n 1 1 1 :ns ide ssing . ab ; t r.e :ad u )f t wa i on f ;ce = s :here ) n d b ;alys 'stem in cthe :ar s :arcn ;£ mo cnt ;umen ;cr k l res rk of naent n pro cece c; pr it po inio oat o same se oi re p rom o ing has out d is o s re g crae r pe to de er or st o canno ta tio organ tare ti arc on quce ntly oce s ssi b rma t tner 1 1 sta dcka pera to c been a ta dire r t rson tha re f hi t, h nth izat hiving, the guai d by eac , ther sing at le to s ion fro st udie me. wl tistical aes an tot cont f e ra ting a loss f tccessi ur veys. d t r a t r c comma £. tne t tne 1 n search a £ or he c w 6 v e r , an is re ion. t regjires ue are ity of h s u r t were nobc upplem' what £ aone th the anald tne rolled s yscf inr. g and the ecords r 1 c a t e trena divid£slstr own expect qui red he infewer to n a litt esea . a ed v as e u nohc main docu ant may proauce bout his le need, rca, to s a conariacles used by n usable. library tain the ment the der ived be reorig inal records ana tends not intelligible informatio work, because there is at tnat point in the r communicate with others sequence tne construct anu final fern of data researchers tend to b for these reasons the emphasizes the need to original file and to derivation of import variables so that they constructed using th ja ta.. cider and more recant surveys pose very different sets of problems for archiving at nohc. an example of each will point out some of these. an example of an older study is the hay, 196'4 occupational prestige study which is the basis ror the hoag e-s iegel-hossi prestige score (used in the 3eneral social survey). the most remarkable and revealing tning is that this study was not archived until february, 1976, que to study director reluctance to allow secondary use. in tne not so distant future, studies such as these will be virtually irretrievable, for the simple reason that much of the data are saved only on aunched cards (which are warped now) with multiple punches, and the collective memories of the people involved with the studies tend to be more and more vague. almost all norc data prior to mid-1960s can be expected to contain some iaultiple punches which require special processing that increases cost (and tediousness, too) of archivingthis problem also makes it dixficult to create fresh copies of data. time, in general, is a crucial factor in data sotrage, since even tape copies can be expected to become unreadable after a while. there is also the problem of reconstructing documentation for these older studies. in many cases we must rely on memories or oersonal fixes to recreate some crucial piece of information. finally, with the immanent loss of the last rew pieces of unit record equipment at norc, we can expect that the solution for the problems of archiving older data sets will become even more difficult. the continuous national survey (cns) was archived in december. 1975 and is an example, although extreme, of some problems with more recent studies (tais study was conducted rrom april, 197 3 to may, 1974). although tne data are documented and usable, they are only dvdilaole in the form or spss system files or character coded data derived from the spss files. the original data were not processed by 47 lassist newsletter, vci. 2, nc. 2 (spring 1978) the tr rather, rield f ent irel ware. the or tireiy of the usable mentdti way to spss sy back to questio but agg piex do system ally mo associa aditi they ormat y wit which igina lest; chara becau on. veri stem the nnair revat cumen files re st tea w onal me were pu and pr h custom was ne 1 data an in cter cod se of a there i f y the files ( seven or es) . ing prob tation r as oppo raight f ith orig ans nched ocess desi ver u have terme ed d lack s vi guali short eigh a ies lem i eguir sed t orwar inal at in ed gnea sed bee diat ata of rt ua ty o of t th s s s th ed f o th d ma data noac. freealmost sof tagain . n ene form is unaocully no f the going cusand erious e comor the e usute rial conciasicns archival problems will ccntine at norc -of that there is no doubt,. if there was serious concern with archiving as a part of the methodology in the social sc ences, ana perhaps more of a reco nition that a scientific enterpri requires good documentation (a the means to get it) , we wouicl more optimistic. as it is, we e pect to see more archival probie with recent studies and increasi difficulty with aata from old studies. s such that with of n fair sure must cont publ ad ju more job nopc perh tion brar that a glum as a s the car orc info ly well, s for cope. inue to ics as st ou r time w of study studies aps a fa 's data ies. note peci e o rmat co ser v for be a we c do can ster arc should , let ai libr f and c ion and nsiaer 1 ice wit the f u s respo an be, orities the al cumen t a fin a t rate, hives not end us point ary mand ommunica data , w ng the p a whicn ture we nsive to but we ana s 1 impor tion so heir way into tae nd data on out ated tion e do reswa will our may pena tant that , at nali48 by 25 iassist quarterly winter/spring 2010 data in development: an overview of microdata on developing countries abstract finding quality microdata on developing countries can seem problematic as their national infrastructures may not support large-scale surveys. in fact a variety of organizations are collecting and distributing data, though the types of data and reasons for collection often differ from those in the most developed countries. much of the data is collected by groups involved in, or interested in researching, the field of international development. this paper provides an introduction to the different groups involved with collecting data on developing countries and to the data they collect, including population health and welfare surveys, program assessments, finance, and opinion data. keywords: developing countries; development; international; foreign aid, survey data as all data librarians know, a good datum can be hard to find. when the data being sought deal with countries generally referred to as developing, less developed, or low and middle income, finding good data can seem like even more of a challenge. the less developed countries form an extremely heterogeneous group, and from a data perspective, they are united primarily by what they lack. in developed countries, the librarian can usually expect national governments and related institutions such as central banks to collect demographic and business data, and national archives or major research institutions to archive surveys. in less developed countries, national governments and institutions may not collect much detailed data beyond the census and other material strictly required for internal use, or may not have developed the infrastructure to process and distribute the data they do have for public consumption. and yet, "there has been a spectacular increase in the availability and quality of data from developing countries in recent years." (bureau for research in economic analysis of development). data is being collected and distributed, but in many cases the groups collecting the data, and the purpose behind the collection, are specific to the field of international development. this paper looks at data on developing countries and the field of international development which concerns itself with them. for the sake of keeping this overview to a manageable length and also of highlighting the sources will be of by kristi thompson1 use to the most people, the discussion will be limited to microdata available from sources that do not require a paid subscription. development is a process, not a permanent state. developing countries are not united only by their lack of the economic advantages that distinguished the most developed countries. they are united by the fact that they are subject to development. "developing countries is an international practice. the essence of this practice is the mobilization and allocation of resources, and the design of institutions, to transform national economies and societies, in an orderly way, from a state and status of being less developed to one of being more developed." (gore) this practice of developing countries has developed itself into a highly diverse international field, with its own sets of standards and practices and key players large and small. people and institutions active in the practice of development collect data, both to assist with their own operations and for fundraising purposes, to "demonstrate that they can perform effectively and are accountable for their actions." (degomme and guha-sapir) an initial challenge is simply to define the countries under consideration. "there is no established convention for the designation of 'developed' and 'developing' countries or areas in the united nations system." however, "in common practice, japan in asia, canada and the united states in northern america, australia and new zealand in oceania, and europe are considered 'developed' regions or areas,"(united nations statistics division) leaving most of the planet still under development. a more fine-grained categorization is the united nations’ human development index ranking, which divides member states into very high, high, medium and low human development. the world bank uses gross national income to divide countries into high income, upper middle, lower middle and low income. the un measure takes into account a broader range of factors, including life expectancy and education as well as income. for purposes of this paper, and where relevant, i will use the groupings from the un 2009 human development report2 to distinguish level of development. the different groups active in collecting data on developing 26 iassist quarterly winter/spring 2010 countries include national governments, intergovernmental organizations (igo's), non-governmental organizations (ngo's) and other – this last being a catchall category containing such groups as academics and for-profit private sector firms. the bulk of the data considered in this paper comes from intergovernmental and nongovernmental organizations. in the interest of brevity, data from individual governments is considered here only insofar as it appears in other compilations. intergovernmental organizations include such familiar large and well-funded organizations as the world bank, the united nations and the world health organization. these are valuable sources both because they conduct surveys and collect data, and because they serve as compilers and standardizers of country-level macroeconomic data. intergovernmental organizations tend to conduct large-scale, nationally representative surveys that are standardized across a number of countries. non-governmental organizations are more heterogeneous, and include larger and better funded organizations such as demographic and health surveys and the international food policy research group along with a myriad of smaller and less known organizations. the data they collect is similarly heterogeneous, but nongovernmental organizations are more likely to provide subnational surveys targeted towards a particular context or issue. “ngo surveys aim at assessing a local situation for needs and programming… (while) un surveys tend to be large scale snap shots of a situation that serves as a point of reference”.(degomme and guha-sapir) surveys conducted by academics, or collaborating groups of academics, range from the large-scale, standardized, cross-national world values survey to the small, very specific local assessments available from mit’s poverty action lab. survey catalogs there are two databases that compile surveys on developing countries, the international household survey network catalog, and the bureau for research and economic analysis of development (bread) / mcarthur survey database. the international household survey network catalog, which is maintained by the world bank data group, contains 4147 surveys at the time of writing. surveys include major population welfare surveys, censuses, firm-level economic surveys, and others, and the catalog can be browsed by country or survey series. the other database, which is available from the bread website, is smaller, claiming to hold only about 500 surveys, and lists as its focus surveys on poverty and health. it can be searched by survey location and module. both of these databases provide information and links to access survey microdata where available. a third database, the complex emergency database (ce-dat), has information about specialized surveys carried out by non-governmental organizations in smaller populations such as refugee camps, often under crisis conditions. at the time of writing it contained 2713 surveys. while microdata is not available from the database, the citations will assist in finding the appropriate contact person to inquire about availability, and the database includes summary statistics for the indicators and has an interface for mapping and charting them. major population health and welfare surveys there are three major, long-running international data collection programs that focus on surveying developing countries: the demographic and health surveys (dhs), the world bank's living standards measurement study (lms) surveys, and unicef's multiple indicator cluster surveys (mics). the world health organization’s world health surveys cover both developing and developed countries and are also worth a look. each has advantages and drawbacks in terms of geographic coverage and topical focus. i have limited this comparison to these four series because each provides relatively uniform data collection across a number of countries, enabling cross-national comparisons. i have excluded survey programs such as the food policy research institute household surveys, which are not uniform enough to consider as a group, and the world fertility surveys, which were last conducted in the 1980’s and are now quite dated. the data from all these surveys is largely available to researchers as microdata, though individual surveys may be unavailable and registration or application may be required. the demographic and health surveys (dhs) program is the largest of the major population welfare survey series. the dhs program is primarily funded by the u.s. agency for international development, and as such has a broadly defined focus on aid, development and policy improvement. they have conducted over 240 surveys in about 90 high, medium and low human development index countries in africa, asia, latin america, the middle east and europe. the earliest surveys date to the mid-1980’s, and multiple waves have been done in many countries. the primary surveys focus on fertility, family planning, maternal and child health, gender issues, and nutrition, and the surveys also cover household and respondent characteristics including education and school attendance, employment, and income and family wealth, making them useful for a range of analyses. dhs also conducts special modules on key topics such as aids and malaria, and scaled-down surveys called the key indicators surveys that are used to assess smaller sub-national populations that may be targeted by special initiatives. the focus is on women and children and men are often excluded. the living standards measurements study surveys are conducted by the world bank, and perhaps not surprisingly, it has excellent coverage of consumption and income. around 90 surveys have been completed in over 40 countries between 1985 and the present. most of the surveys cover countries with high or medium development, with only one (malawi) classified as low, 27 iassist quarterly winter/spring 2010 and another, iraq, that is currently not classified. while the focus is on economic measures ranging from income and employment to debt and purchasing behavior, there are a number of health, demographic and social measures. some of the surveys are available directly from the world bank web site, others are distributed by the government of the country in which they were conducted. the world health surveys were done in 2002, and followup studies such as the who study on global ageing and adult health are being conducted. microdata has been released for all the countries in the original surveys and is available upon signing a data use agreement. there are detailed questions on health, mortality, health insurance and use of health services along with some basic demographic and wealth variables. employment is included, but not income. a different questionnaire is used for countries in the highest income group. the multiple indicator cluster surveys, done by unicef, are currently on their fourth wave, and have been done every five years since 1995. the mics are particularly notable for their coverage of the poorest countries; in the 2005 wave, 11 of the 55 countries they surveyed were in the low development group. they also surveyed some subnational populations. however, the surveys are limited in their usefulness for general analysis due to the selection of variables available. the mics were developed specifically to track progress on the world fit for children plan of action and the millenium development goals, and cover indicators relating to child health and welfare, women’s reproductive health, and some basic household variables, but do not include the usual economic and demographic variables such as income and employment. comparing the major population health and welfare survey series data for assessing, modeling and targeting development while much of the data collected on developing countries is used to provide evidence to assess developmental progress, data focusing on the evaluation of specific approaches is rarer. the abdul latif jameel poverty action lab was formed at mit as a network of professors around the world who use randomized evaluations to answer questions on poverty alleviation. “what makes j-pal's work innovative is that such randomized studies haven't typically been used in evaluating poverty-alleviation programs, or even in the wider field of economics.” (standish) they offer training on randomized evaluation and maintain a database that currently holds information about 234 randomized assessments, with microdata available for 13 of them, on topics ranging from textbook provision to microfinance. the international food policy research institute (ifpri) is another source for random evaluations and has released data for studies such as the comparing food versus cash for education program, along with their more standard household welfare datasets. the international food policy research institute also collects data to construct social accounting matrices, which are currently available for about 30 countries. social accounting matrices are used as the basis for various economic models that can be used to estimate the effects that different policy or aid approaches will have before dhs lsms whs mic time period focus demogr aphics labour and income variables countries covered 1985 present health, particularly reproductive good some 90+ low, medium and high development 1985 present consumptio n and income good good 40+ mostly high and medium development 2002 health and health systems good some, no income 70 ranging from low to very high development every 5 years from 1995 child health and welfare, reproductive health limited, mostly househol d head limited 65+ low and medium development 28 iassist quarterly winter/spring 2010 actually carrying them out. finance the world bank claims to provide “the world's most omprehensive company-level data in emerging markets and developing economies” (world bank group) and i would not attempt to dispute that claim. they have conducted surveys in over 125 countries, including survey projects such as the world business environment surveys, as well as more specialised projects such as the recent threeround financial crisis surveys and the management, organisation and innovation survey work. firm-level microdata is available to researchers at no cost from these surveys, in contrast to the international monetary fund, which also conducts surveys but only releases macrodata, and some of that only to paying customers. the world bank’s enterprise survey portal also allows users to construct charts online. a couple of regional sources for financial microdata are the economic research forum, focusing on the middle east and north africa, and oxford university’s centre for the study of african economies, which focuses on sub-saharan africa. the economic research forum has released data on micro and small enterprises in egypt, lebanon, morocco and turkey, as well as an egypt labour market panel study. the methodology documents for the micro and small enterprises data provides interesting insights into some of the difficulties in collecting and analysing data on developing economies; for example, in turkey the researchers were unable to weight the rural sample because no more authoritative data on rural enterprises existed.3 the centre for the study of african economies has comparative cross-national firm-level datasets primarily of manufacturing firms, as well as a couple of household panel surveys, and some more specialized datasets from working papers. opinions, attitudes, and values the data collections discussed so far have been factual, consisting of theoretically objective measures of demographic and economic variables that can be used to implement and assess development programs. opinion and values surveys may not appear to have the same practical, development-oriented application that most of the data discussed so far have. development is done, in the end, for people, to improve their lives as they actually experience them, not merely to increase gni or some other measure. opinion data can “delve deeper than, for example, official poverty data, by reporting on people’s experiences in obtaining basic human needs and their own perceptions of whether or not they feel poor.”(corporacion latinobarometro) in addition, development happens in a political as well as an economic context. it is ideally done in cooperation with and with the support of governments and the people in the area being developed. two of the largest and best-known cross-national opinion surveys are the world values surveys and the global barometers. the world values surveys are an international collaboration among academics who are attempting to survey the “basic values and beliefs of the publics of more than 80 societies.” (world values survey) five waves have been completed, conducted between 1981 and 2008, with a sixth currently in progress. "since each national group funded its own survey, its first wave was largely limited to relatively developed societies." (world values survey) by the second wave, the researchers behind the project had decided that "it was important to include societies across the entire range of development, from low income societies to rich societies," (world values survey) and additional researchers and sources of funding were found. the global barometers are done more modularly, with separate afrobarometer, arab barometer, asian barometer, east asian barometer, and latinobarometro surveys, and a eurasia barometer under development. (another one, america’s barometer, covers some of the small latin american nations and is available for subscription through the latin american publib opinion project (lapop).) while the global barometers were inspired by eurobarometer, there is no formal association.4 the barometers are designed to be a comparative survey of attitudes and values toward politics, power, reform, democracy and citizens' political actions. they are repeated at approximately three year intervals, and include a core module of questions asked across regions, plus additional region-specific questions. “whereas the wvs addresses deep-seated, semi-permanent cultural values, the gb is concerned with tracking emerging political and economic attitudes, which are often subject to rapid change.”(corporacion latinobarometro) the pew global attitudes project is conducted by the pew research centre, a u.s. based non-governmental organization. it is a series of opinion surveys, each covering anywhere between five and 50 countries, that have been conducted between 2001 and the present. key areas of interest include “attitudes toward the u.s. and american foreign policy, globalization, terrorism, and democracy.”(pew research center) while there is a decided focus on american foreign policy (for example, the data formed a basis for the book america against the world: how we are different and why we are disliked 5), the data also includes questions of broader interest such as global reactions to issues in the news. the multinational surveys mentioned above are generally too large and slow to implement to track responses to local issues under rapidly changing circumstances. local polls, whether done by local media organizations, academics, or others are more likely to track these highly situational opinions, but the decentralised nature of local polling as well as language barriers make them particularly difficult to 29 iassist quarterly winter/spring 2010 access. worldpublicopinion.org, which is managed by the program on international policy attitudes at the university of maryland, compiles and analyzes some of these opinion polls, and they also conduct their own local polls. their own studies are performed by a network of research centres in 25 countries, and many of their datasets are available to download. conclusion this overview has only touched on some of the sources of data available on the developing world. the choice to focus on microdata means that macroeconomic and many finance sources have been neglected; limiting myself to freely available sources means that resources in archives such as icpsr or subscription services such as polling the nations were left out; most smaller projects covering only one or a few countries are included only insofar as they appear in one of the survey databanks. still, it should be clear by now that a great variety of data on developing countries and development is available, whether it is collected by practitioners of development trying to improve their outcomes, researchers studying developing countries and various types of comparative social science, or other interested parties from both within and outside the developing world. while much of this data is gathered to serve relatively narrow purposes, collectively it can provide the researcher with a window, however imperfect, into the minds and lives of a large and often unheard portion of the world's population appendix: list of data sources in the order they were mentioned. survey catalogs bread/mcarthur/ccpr survey database: http://ipl.econ. duke.edu:8080/survey/ international household survey network’s central survey catalog: http://www.internationalsurveynetwork.org/home/ index.php?q=activities/catalog/surveys complex emergency database (ce-dat): http://www. cedat.be/ major population health and welfare surveys demographic and health surveys: http://www.measuredhs. com/ living standards measurement study:http://www. worldbank.org/lsms/ world health surveys: http://www.who.int/healthinfo/ survey/en/index.html multiple indicator cluster surveys: http://www.childinfo. org/mics.html data for assessing, modeling and targeting development mit’s abdul latif jameel poverty action lab :http://www. povertyactionlab.org/evaluations?filters=type:evaluation international food policy research institute (ifpri) surveys: http://www.ifpri.org/datasets finance world bank enterprise surveys:http://www. enterprisesurveys.org/ economic research forum data:http://www.erf.org.eg/ cms.php?id=datasets oxford university’s centre for the study of african economies:http://www.csae.ox.ac.uk/ opinions, attitudes, and values world values surveys: http://www.worldvaluessurvey. org/ global barometer:http://www.globalbarometer.net/ worldpublicopinion.org: http://www.worldpublicopinion. org/ pew global attitudes project: http://pewglobal.org/ references b r e a d bureau for research in economic analysis of development. "data from developing countries." available online: http://ipl.econ.duke.edu/dthomas/dev_data/index. html, accessed 18 oct 2010. corporacion latinobarometro. "global barometer – strategy." available online: http://www.globalbarometer. net/strategy.htm. accessed 18 oct 2010. degomme, olivier and debarati guha-sapir. "mortality and nutrition surveys by non-governmental organisations. perspectives from the ce-dat database" emerging themes in epidemiology 2007 vol. 4, no. 11. available online: http://www.ete-online.com/content/4/1/11. accessed 18 oct 2010. gore, charles. "the rise and fall of the washington consensus as a paradigm for developing countries." world development vol. 28, no. 5, pp. 789-804, 2000 pew research center. "about the project." available online: http://pewglobal.org/about/. accessed 18 oct 2010. standish, sarah. "researching better ways to end poverty" (blog post). 10 dec. 2009. available online: http://www.globalenvision. org/2009/12/03/rigor-science-now-economics-too 30 iassist quarterly winter/spring 2010 united nations statistics division. "standard country and area codes classifications." (see footnote c.) available online: http://unstats.un.org/unsd/methods/m49/m49regin. htm the world bank group. "enterprise surveys." available online: http://www.enterprisesurveys.org/ world values survey. "building a worldwide network of social scientists." available online: http://www. worldvaluessurvey.org/wvs/articles/folder_published/ article_base_51. accessed 18 oct 2010. endnotes 1. contact: kristi thompson, data librarian, leddy library, university of windsor, windsor, ontario, n9b 3p4. phone:1 519 253 3000 x3858 email: kathomps@uwindsor.ca this paper is based on the presentation data in development given by the author, iassist 2010. 2 united nations development programme. "human development report 2009." available online: http://hdr. undp.org/en/media/hdr_2009_en_complete.pdf 3: ozar, semsa. "micro and small enterprises in turkey: uneasy development". available from the data download area of the economic research forum web site; contact http://www.erf.org.eg/cms.php?id=mse_database for access. 4 see http://www.globalbarometer.net/background.htm 5 see http://pewglobal.org/americaagainsttheworld/ vol263 4 iassist quarterly fall 2002 iassist quarterly fall 2002 5 editor’s notes welcome to the iassist quarterly vol. 26 issue 3. three articles are presented in this issue. at the amsterdam conference in 2001 in the session with the theme “thematic archives” the paper titled “social science information service in poland. an attempt to present the state of art” was presented. the authors of the paper are teresa wildhardt from the cracow pedagogical university main library and anna sokolowska-gogut from cracow university of economics, main library. the paper addresses the questions: 1) can we observe any direct relation between transformation and growing demand for social science information service? if so 2) what kind of information is searched for most frequently? and 3) are social science information services ready to supply users with necessary information? the paper tries to answer these questions by analyzing a questionnaire that was answered by 30 scientific libraries in poland and contains a brief characterization of information sources, staff and categories of users. special regard has been given to polish central statistical office (cso). the staff members are well educated in the relevant subject areas and have technical knowledge on how to obtain information from a variety of sources and are servicing a range of users from individual students and researchers to various kinds of institutions. the paper also describes the many information services from the cso. the paper concludes among other things that the information need has increased in poland since 1990 and that the main obstacle is the scarcity of funding. the following paper is by james reid, geoservices delivery team at edina, edinburgh university data library in scotland. the paper entitled “geoxwalk – a gazetteer server and service for uk academia” was presented at the 2002 joint digital libraries conference in july 2002 in portland, oregon. the geoxwalk project was conceived as a development project to build a shared service that would service the jisc ie (the joint information systems committee – information environment) by providing a mechanism for geographic searching of information resources. the paper cites the ukʼs national geospatial data framework (ngdf) in their estimation that as much as eighty per cent of the information collected in the uk today is geo-referenced in some form. geography is frequently used as a search parameter, and there is an increasing demand from users, data services, archives, libraries, and museums for more powerful geographic searching. however, there are serious obstacles to meeting this demand. clearly, no single system of spatial units and coding will suit all purposes, as people conceptualize geographic space in different ways and different servers deploy differing geographic naming schemes. ideally, users should not be forced to have to explicitly convert from one ʻworld ̓view to another. the described product are able to perform these translations (or ʻcross-walks ̓– hence the name geoxwalk). the technical issues are described in the paper but are also mentioning some outstanding issues that will require further research. the last paper concerns the uk census. this paper is related to the paper “let us bring you to your census: recent developments in uk census data provision” by lucy bell in iassist quarterly 26-2. the title of the paper is “learning and teaching with the uk census” and is authored by dr mark brown, deputy director ccsr (centre for census and survey research), university of manchester, and cressida chappell, head of the history data service, uk data archive, university of essex, and dr jackie carter, chcc project manager, mimas, manchester computing, also at university of manchester. the paper states that the uk academic community has access to an electronic collection of historical and contemporary census data and resources (chcc). individual datasets have been used extensively in research but they have been widely under-used in learning and teaching. to address this issue, the joint information systems committee (jisc), under its learning and teaching programme, has funded a project to develop learning and teaching materials. the units are designed to encourage a “pick and mix” approach offering for the teachers. the units can be used for both classroom-based and online learning, and can thus be used both by teachers and directly by students for independent learning. several units are described in the paper. the project incorporates user-friendly web-interfaces and draws upon other projects such as the ddi and nesstar. several web-sites are mentioned in the paper and further information can be found at http://www.chcc.ac.uk. access iassist at the web on www.iassistdata.org. papers for the iassist quarterly are most welcome. please contact the editor (kbr@sam.sdu.dk) about submissions. karsten boye rasmussen, february 2003 http://www.chcc.ac.uk http://www.iassistdata.org/ mailto:kbr@sam.sdu.dk vol183&4 17fall/winter 1994 this paper will discuss why and how to use gopher servers on the internet to provide access to locally developed data. this includes formatting the data, and establishing the links to the data on the server. the responsibilities involved with providing information on the internet will also be discussed. reasons to mount data on the internet the internet is a vast source of information and chaos, why would anyone want to add to it? some possible reasons for mounting data on the internet include: -the internet allows remote access to the data. the data and its users are not limited to physical locations. this means that people from other institutions, as well as your own, can get to your data. -the internet allows 24 hours a day, 7 days a week access to your data (barring the usual network outages or maintenance down time for the data server). -by providing the data on the internet, it is by definition in electronic format. this allows for further manipulation or massaging of the data. it allows users to take advantage of the computer’s abilities, such as searching and sorting. -the internet allows for quick and easy publishing or updating of your data. reasons not to mount data on the internet most of the reasons for not mounting data on the internet are based on privacy and legal issues. -the data is copyrighted and can not be re-distributed. -the data is of sensitive nature and should not be accessible to just anyone, i.e. the world. -the data is already out on the internet, in a number of different places. -the internet location for your data is unreliable or not maintained by anyone. why use a gopher server currently, gopher clients are widely distributed across the internet community and are available for most types of computer hardware and operating systems. most gopher client and gopher server software does not require high level computers on which to run, unlike other internet tools such as mosaic. gopher clients, as a rule, are easy to use. they provide a common interface to many different types of resources. the gopher protocol provides the capability to perform searches on databases and files. currently, this is mostly primitive string or character-by-character searches. gopher servers have the ability to link or point to other gopher servers. this linking ability makes it relatively easy to create subject oriented gopher servers. gopher servers as a point of access by julie a. fore1 assistant automation librarian indiana university ruth lilly medical library 18 iassist quarterly reserves collection by author abbas, abul k. cellular and molecular immunology. abdellah, faye g. patient-centered approaches to nursing. new directions in patient-centered nursing; guidelines for systems of service, education, and research. ackermann, uwe. essentials of human physiology. [...] fig. 1 — “reserve collection by author” printout from pro-cite. indiana university ruth lilly medical library gopher server pilot project the indiana university ruth lilly medical library has been maintaining a database of its permanent reserve collection holdings using a bibliographic database management system called pro-cite, made by personal bibliographic systems (pbs). the library uses pro-cite to keep this database because the software allows the library to provide a number of different printouts for the library patrons to use. the reserve collection is shelved (mainly) in title order. pro-cite allows the library to generate printouts of the collection in shelflist (title) order, as well as lists sorted by author or subject heading. (fig. 1) the library patrons make great use of these printouts, as do the library circulation staff. the permanent reserve collection database is small at about 225 records. this made it a perfect pilot for testing how well the library’s various pro-cite databases would make the transition from in-house use only to internet accessible information. the pilot project started with an analysis of the data and data fields already in the pro-cite database. (fig 2) rec# 780 auth abbas, abul k.//lichtman, andrew h.//pober, jordan s. titl cellular and molecular immunology plpu philadelphia publ saunders date 1991 extn xi, 417 p isbn 0721630324 call qw 568 a122c 1991 desc cellular immunity/immunity—molecular aspects/immunity, cellular/ lymphocytes -immunology fig. 2 — example of a bibliographic record in pro-cite 19fall/winter 1994 after the evaluation of the electronic data, a decision was made as to which data elements would be most valuable to a person accessing the database over the internet. i decided to use basic bibliographic citation fields, i.e. author, title, place of publication, publisher and date; as well as the subject heading information. the data from these six fields were then exported from pro-cite using pro-cite’s import/export utilities. pro-cite created a standard comma ( “,” ) delimited file. (fig 3.) the gopher server software being used by the ruth lilly medical library, ka9q nos, requires database files to be in dbase iii or dbase iv format. while most current database management programs such as paradox by borland and r:base by microrim can save data in dbase iii format, we chose to use the dbase iii program for the next part of the pilot program. a database structure was created in dbase iii using the six fields exported from the pro-cite database. to keep things simple, the pro-cite field labels were used as the field labels in the dbase database. while pro-cite, for the most part, does not use fixed field lengths, dbase requires fixed field lengths. we made educated guesstimates on what the dbase field lengths should be. figure 4 shows the final structure for the reserves dbase database. the dbase iii import function was used to convert the pro-cite produced comma-delimited file into a dbase iii database. a paper report of the new database was then created to verify two things. first, that the data was correctly transmitted from pro-cite to dbase iii. second, to verify that the data, itself, was correct and complete. the data had indeed transferred correctly, but it was found that some of the records in the original database contained incomplete information. once the dbase database had been cleaned up, and an ascii text file bibliography was generated from it using r&r report writer by concentric data systems, the data was ready to be transferred to the actual microcomputer running the gopher server software. in the case of the ruth lilly medical library gopher server, this meant taking down the gopher “abbas, abul k.//lichtman, andrew h.//pober, jordan s.”,”cellular and molecular immunology”,”philadelphia”,”saunders”,”1991",”cellular immunity/immunity—molecular aspects/immunity,cellular/lymphocytes—immunology” “abdellah, faye g.”,”patient-centered approaches to nursing”,”new york”,”macmillan”,”<1960>”,”nurse-patient relations/education, nursing” fig. 3 — sample of pro-cite export file in comma delimited format. structure for database: c:reserves.dbf number of data records: 225 date of last update: 4/21/94 field field name type width dec 1 auth character 130 2 titl character 200 3 plpu character 50 4 publ character 50 5 date character 8 6 desc character 254 ** total ** 693 fig. 4 — final dbase iii file structure 20 iassist quarterly server, that is, exit out of the server program. then, using the dos copy command to move the files from the library’s novell file server to the gopher server’s dos-based microcomputer. once the actual files were residing on the gopher server’s hard drive, a suitable access point in the gopher’s menu structure had to be found. finally, the gopher server’s menu configuration files had to be edited to include the pointers to the files. the most logical place to include the reserve collection information was in the menu with all the other files specific to the ruth lilly medical library, i.e. the files containing the library’s hours, policies, journal holdings, etc... (fig. 5) fig. 5 — “library indexes, catalogs and information” menu from ruth lilly medical library gopher in the ka9q gopher server software, the menus are designed using directories and subdirectories on the server’s hard drive and “ginfo” files (which possibly stands for “gopher information” or “gopher index file”). for every menu on the server (seen by a gopher client) there is a corresponding directory on the hard drive of the gopher server and in that directory a ginfo file. the ginfo file contains five elements: 1) the text shown on the menu to a gopher client, 2) a code for the type of resource that is being pointed to (text file, database, directory, macintosh binhexed file, uuencoded file, gif file, etc...), 3) the name and path of that resource (for example /server/reserve.db/reserves.dbf), 4) the internet address of the gopher server that provides the resource (for example gopher.medlib.iupui.edu) and finally, 5) the port for that gopher server, usually port 70. an example of a ginfo file is seen in figure 6. the ginfo file is were the telnet or ftp links to other internet sites are described. 1medical library reserve list 1c:/server/reserve.db gopher.medlib.iupui.edu 70 1medical library information 1c:/server gopher.medlib.iupui.edu 70 1medical library journal holdings 1c:/pub gopher.medlib.iupui.edu 70 1indiana university libraries catalog 1c:/pop/catalog gopher.medlib.iupui.edu 70 1other libraries’ catalogs 1/libraries yaleinfo.yale.edu 7000 1carl journal title index 1c:/library/carl gopher.medlib.iupui.edu 70 1first search indexes (password required) 1c:/library/first 134.68.85.17 70 . fig. 6 ginfo file for the “library indexes, ....” menu of the ruth lilly medica l library gopher server. 21fall/winter 1994 figure 7 illustrates what might be found in a directory on a gopher server, note the presence of the ginfo file. volume in drive m is bvol directory of m:\gopher\library ginfo 523 05-13-94 8:48p carl <dir> 05-23-94 10:35a first <dir> 05-23-94 10:35a 3 file(s) 523 bytes 191,299,584 bytes free fig. 7 — listing of the files in the sub-directory containing the ginfo file from figure 6. the second line of the ginfo file shown in figure 6 is the pointer to the library’s reserve collection menu. as one moves through the gopher’s menu structure, the medical library reserve list menu eventually appears. (fig. 8) fig. 8 — “medical library reserve list” menu of the ruth lilly medical library gopher server. the “medical library reserve list” menu allows a gopher client to browse or page through a text file bibliography of the reserve collection, ftp (file transfer protocol) the bibliography back to the user, or perform a character search on the database either using the title field, the author field, or the descriptor field. the ginfo file for this menu determines which function is performed on which file. there are only three files in the directory for this menu. the ginfo file, the actual database file called reserves.dbf, and the text file bibliography called reserves.txt. (fig. 9) 22 iassist quarterly volume in drive m is bvol directory of m:\gopher\server\reserve.db reserves dbf 156,160 04-21-94 3:46p reserves txt 30,989 04-25-94 4:17p ginfo 530 05-21-94 2:12p 4 file(s) 188,210 bytes 191,299,584 bytes free fig. 9 — directory listing of the reserve.db subdirectory figure 10 illustrates the ginfo file for the medical library reserve list menu. 0rlml reserves collection list 0c:/server/reserve.db/reserves.txt gopher.medlib.iupui.edu 70 5get reserves collection list (text file) 5c:/server/reserve.db/reserves.txt gopher.medlib.iupui.edu 70 7search the reserves collection by title qc:/server/reserve.db/reserves.dbf~titl gopher.medlib.iupui.edu 70 7search the reserves collection by author qc:/server/reserve.db/reserves.dbf~auth gopher.medlib.iupui.edu 70 7search the reserves collection by keyword qc:/server/reserve.db/reserves.dbf~desc gopher.medlib.iupui.edu 70 . fig. 10 — ginfo for “medical library reserve list” menu figure 11 shows what the text file bibliography looks like when viewed by a gopher client. the bibliography file was created so that the users could have access to a formatted file that they could browse through. it was decided that the dbase database format was not very easy to browse (fig. 13), nor was it in a file format that most people could use once they had it back at their own computer. the double slash marks (//) in the author field are left over formatting codes from pro-cite. these codes will be removed the next time the database needs significant updating. 23fall/winter 1994 fig. 11 — browsing the “rlml reserves collection list” option the ka9q gopher server search capabilities are currently string or character based searches. this means that when “search the reserve collection by author” is selected off the “medical library reserve list” menu, a dialog box will appear asking the user to enter the words to be searched for in the author field of the database. in the example illustrated by figures 12 and 13, the user asked the gopher to search for all occurrences of the word “sid” in the author field. as the results of the search show (fig. 13), the gopher does not care where the “word” “sid” appears in the author field. it found the letters, or characters, “s-i-d” in the word “president” and in “sidney”. it is expected that gopher-based searching will improve in the future. if not, then some other internet tool will take gopher’s place. fig. 12 — gopher search dialog box 24 iassist quarterly fig. 13 — results of searching for “sid” in the author field of the reserve collection database internet responsibility it is not enough to just mount a database on the internet. it is necesary to take responsibility for it and for the gopher server on which it resides. there are a number of points to keep in mind when setting up and maintaining servers on the internet. -keep your server up and running. no one can use your data if your server or your network is down. -when (not if) you take your server down for routine maintenance, i.e. on a routine schedule, post this information on your server. 25fall/winter 1994 -if there are limitations to your server or your data, post this on your server and on any public announcements you send out about your server. some examples of limitations might include, access only during non-business hours like 5 p.m 6 a.m est, a limited number of simultaneous users, the fact that passwords are required for access to certain files or services, or that only users from a certain place (campus, university, etc...) are permitted access to a resource. -keep your data current and accurate. if this is not possible, indicate on the server that the data is old/out-of-date or not necessarily accurate. -if you move your resource to a new internet site or remove it from the internet, announce this. place a notice stating the new location of the resource in the old location of the resource. post announcements to appropriate listservs and newsgroups. -if you are keeping copies (mirrors) of your resource at more that one location, keep them current and announce their locations as well. technical information about the indiana university ruth lilly medical library gopher s erver url: gopher://gopher.medlib.iupui.edu port 70 the iu rlml gopher server is currently running on a gateway 2000 386-25 mhz processor with 4 mb of ram. the computer has a 300 mb hard disk, of which approximately 50 mb is being used. the computer is running dos version 5.0 and is attached to a 4 mbps token ring lan. the server is backed up weekly to a novell netware 3.11 file server. the gopher server operating system is ka9q nos, a dos-based network operating system. ka9q supports gopher; pop2, pop3, and smtp mail server protocols; ftp, anonymous ftp, telnet and finger; cso name server functions; ntp (time) server functions; and www server functions. the indiana university ruth lilly medical library currently is not supporting the mail server functions but is experimenting with the other capabilities of the ka9q software. in the future, the indiana university ruth lilly medical library gopher server will be switched to a 10 mbps ethernet lan. it may be switched to a unix based computer, and it may be given additional worldwideweb (www or w3) functionality and resources. places to find more information newsgroups for gophers and other information servers: comp.infosystems.gopher comp.infosystems.www comp.infosystems.wais frequently asked question (faq): gopher faq can be retrieved via anonymous ftp from the following site: rtfm.mit.edu:/pub/usenet/news.answers/gopher-faq or via gopher from: 129.130.10.5 port=70, path=0/frequently asked questions (faq)/gopher-faq ka9q nos (network operating system) dos-based gopher server software. ka9q mailing list: send an email to ashok ashok@biochemistry.cwru.edu and ask to be added to the mailing list. this address is an individual, so be nice. ka9q manual: the user manual is available via gopher from the following site: cases.pubaf.washington.edu, port 70, in 1c:\manual university of minnesota — the top gopher: gopher://gopher.tc.umn.edu port 70 questions or comments for the gopher development team, send e-mail to: gopher@boombox.micro.umn.edu news about new gopher servers and software, subcribe to the gopher-news mailing list: gopher-news-request@boombox.micro.umn.edu the most recent releases of gopher software is available via anonymous ftp from: boombox.micro.umn.edu in the /pub/gopher directory. 1. paper presented at iassist 94 in san francisco, may 1994 vol252 24 iassist quarterly summer 2001 here we present the activities of the social science data archive (adp) in slovenia. there are still few such archives, which form the basic infrastructure of national research work, to be found in central and eastern europe. in the process of setting up, a data archive may draw on the experience and support of similar institutions abroad, but the receptivity of the home environment is essential, as discussed in the introductory section. the kinds of past and present research that set a special stamp to the adp’s collection are described in the following in order to spur interest in using the material it is available. introduction the social science data archive (adp) was established on 8th july 1997 by the senate of the social sciences faculty (fdv) of the university of ljubljana. as of autumn 1999 the adp is located in the faculty’s new building, on the second floor, above the above the joze goricar central social sciences library. the work of the adp is supervised by the adp council. an advisor in the ministry for science and technology (mzt), which is responsible for the information infrastructure, is authorised to follow the adp’s activities. as a specialised scientific information centre in the field of the social sciences, the archive is financed from the budget under a contract with the ministry. in the case of special development, research or educational projects the archive seeks support offered in public notices. the archive also derives part of its income from services rendered and remuneration for the use of its database (see arhiv dru~boslvonih podatkov 2000). besides general restrictions related to ethical rules for the use of data, the archive may impose special restrictions on access to particular units of material for different types of users. materials are usually available free of charge and without restriction for educational purposes, while each unit of material is expressly labelled to indicate whether or not the author’s permission must be obtained for use for public or profit-making purposes. users are obliged to cite the author and the archive when publishing any of the material. the adp has been a member of the council of european social science data archives (cessda -http:// www.nsd.uib.no/cessda/) since 1999. cessda provides support in the establishment of new archives, particularly in the form of advice and job-training. the adp has adopted the method of work common in archives with a long tradition such as the german za zentralarchiv für empirische sozialforschung and the uk da united kingdom data archive with modifications to suit conditions in a small country. co-operation with cessda also means that slovenian data stored in the adp is at the disposal of researchers and other users around the world. similarly, the adp mediates access to material in other countries for its local users. the main obstacle to greater international utilisation of slovenian research material is language. however, englishspeakers are advised to consult the adp study descriptions, which are also written in english and contains short summaries, descriptors and other information about the research, to choose materials. the next step is to contact the adp staff who will advise them or arrange for a translation of the required material. when the study is a slovenian part of some comparative international research in most cases an equivalent to it can be found in the english originals without much difficulty. the adp’s basic function is to store and protect data from damage so that it will be available for secondary analyses in research or for teaching purposes. for the foreign as well as the local user the first question, of course, is how to get data on the data, namely on the study materials the adp houses that are pertinent to his or her current research purposes. this can be done by searching a new catalogue on the internet, nesstar networked social science tools and resources (http://www.nesstar.org/). in view of its small size, the adp has opted for the greatest possible standardisation of its procedures to assure a comparable form of documentation and description of the data stored. it was amongst the first to adopt the metadata standard developed in the framework of the data documentation initiative (http://www.icpsr.umich.edu/ddi/ codebook.html). together with the document type definition (dtd) this has allowed the machine-readable social science codebook to be filled in xml extensible markup language (stebe, omerzu 1999). standardisation of procedures enables the efforts of larger archives to the social science data archive in slovenia by janez stebe & irena vipavc 1 iassist quarterly summer 2001 25 construct tools and user-tailored services like nesstar to be pooled. nesstar which was being developed with eu sponsorship and through the co-operation of several partner social sciences data archives, in effect amounts to a virtual data library. the adp has been able to join this project thanks to its admission to cessda and because it has adopted the new ddi dtd data description standard completely. the catalogue allows to search in several archives simultaneously at the level of research description, methodology, and variables of the database, on-line exploratory statistical analyses of the selected database and the ordering data by electronic means for deeper analysis are possible. the data holdings the importance of an archive is judged by the value of the material it houses. the adp rests in particular on the rich and long tradition of the slovenian school of empirical sociology which took shape through the work of the institute for sociology and philosophy (isf) established in 1959. the isf launched several classical-type studies, which blazed the way both thematically, and personally for particular lines of research which continue to this day, although under different institutional conditions. the main research topics of the isf were as follows: local communities and spatial sociology, quality of life and social stratification, customs, lifestyles, and especially the influence and use of the mass media, attitudes and values, family sociology, the study of fertility, youth, industrial sociology and so on. thus, surveys conducted in the sixties and seventies represent a unique source of data on various social phenomena that may serve as starting-points for comparisons over time (stebe 1999). researchers like stane saksida, katja boh, zdravko mlinar and niko tos were trained in this period and went on very actively into further empirical research which steadily built up long thematic series. amongst the most outstanding research is the mass communication media survey (mks; vreg et al. 1962) which was one of the first strictly-designed empirical surveys on mass media and more broadly, on lifestyle, personal use of time, and quality of life in slovenia. thematic links to this research may be traced along the line of social stratification to the survey in the seventies on social stratification and mobility in yugoslav societyssm (saksida et al. 1974) which covered slovenia and macedonia. in the eighties there was an extensive comparative project encompassing the federal republics and provinces in the former-yugoslavia. the class structure of the former yugoslav society was studied in the eighties with the survey "class composition of contemporary yugoslav societies (kb)" (jambrek, tos et al. 1987). at the end of the eighties and in the nineties this topic was echoed in various sections of related surveys on the level of living survey (lol) (svetlik et al. 1994), the international social justice project, 1991 (isjp 1991) and surveys on the topic of inequality in the context of the international social survey programme (issp 1992,1999) in which slovenia has been participating fully through the slovenian public opinion survey (sjm) since 1990. another extensive project, related in type to the foregoing in that data on habits and changes in social structure predominates, is the migration project dating back to the seventies (tos 1974, 1976a, 1976b). this data is stored in the adp. the adp is endeavouring to gradually include all major past studies into its collection. this undertaking is obstructed in many cases because data and documentation has been lost, even in cases when the adp has managed to reach the main researcher. this is unfortunately the case with the mks and ssm survey cited above. data was best preserved when international comparative research was involved. thus, thanks to foreign archives already in operation in the sixties, slovenian data for the following surveys has been preserved: tbs the time budget survey (szalai et al. 1966), an exceptionally well-crafted, extensive and well conducted survey on the way time is used; iw2000 – images of the world in the year 2000 (boh 1967) which yielded unique data on the youth of the time through thematically-based questions on future expectations and values regarding peace; democracy and local governance international studies of values in politics a frequently cited survey which also dealt with methodological problems in comparative research of most different social and political frameworks of local democracy such as in the former yugoslavia (and slovenia) and india on one side and the usa on the other. (mlinar et al. 1966 1991); in the same way, political participation (barbie boh et al. 1971) was a major project at the time when comparative research was flourishing. the participation of slovenian researchers in them shows that they were able to communicate with researchers abroad and capable of achieving the quality standards required in conducting research and presenting data. the importance of continuity for the safeguarding of data is demonstrated by two examples of research that were designed longitudinally and therefore ensured the preservation of the raw data for their own purposes. the first is the slovenian public opinion survey (sjm) (tos 1968 1999), which is the best-known and most extensive empirical research in slovenia. in type it is comparable to the general social surveys abroad and to a great extent it has served in the omnibus design as an infrastructure for fieldwork on individual thematic chapters of deepened conceptually based research. in addition to this, the survey contains topics related to current affairs at different cross-sections and so it reflects politologically relevant attitudes and opinions at particular times. some of these are more interesting for comparisons over time while others are related to the institutional context of the previous system that collapsed after 1989. another major survey that has tos 26 iassist quarterly summer 2001 been preserved in the adp is the “level of living survey in slovenia (lol)”. the most recent replication of it was in 1994 (svetlik 1994). conceptually it is modeled on the scandinavian surveys on subjective assessments of quality or satisfaction with living in various respects such as residential conditions, employment, and leisure time, in comparison with indicators of objective living conditions. both series are now fully accessible through the adp. the archive’s own contribution is a cumulative slovenian public opinion 1990-1998 database (tos 1999) which has been created by combining series of more than 100 identical variables from the sjm surveys from the period between 1990 and 1998. thematically it is encompassed in the politbarometer section. other domestic serial surveys which continuously generate new data are the research on the internet (ris) (vehovar 1999 for the latest of the series) and the youth (ule 1985-1999). the criteria of relevance of a study for secondary analysis include: the data refers to the general population, the sample is random and sufficiently large, the topics are unique and refer to important issues both substantively and in terms of applicability, and most notably, comparability across time and space. in line with the policy of giving priority to research that according to these criteria is the most interesting for re-use, the focus is placed on acquiring new data from extensions of serial surveys and from international comparative surveys that include slovenia. thus the adp promptly obtains data from the central and eastern eurobarometer survey (ceeb) (cunningham 1997, latest in the series), the international social survey programme (issp 1999), the international crime victimization survey (icvs) (dijk, mayhew 1997). it also applies for other well-known surveys within international projects like the fertility behaviour of slovenians (kozuh-novak et al.1998), the aufbruch new departures†‘97: international research on religion and attitudes toward the church (tos et al. 1997), the comparative study of electoral systems (cses 1996) and the european (world) values survey (tos 1999). to increase the diversity of the subject fields the adp is negotiating with the criminology institute of the law faculty, the pedagogic institute, the andragogic centre, the centre for local communities, and the defense studies centre, which are also major producers of empirical social surveys in slovenia, to deposit their data in the archive. the adp is also linking up with the sicris (slovenian current research information system which makes available information on current research projects in slovenia; see: http://sicris.izum.si/). the authors of research projects financed from public research funds are obliged to make empirical data from the research available to other interested users. the adp will collect and disseminate it. the very best gauge for the adp, that data is worth collecting, is of course the expressed interest in it shown by users looking for data from particular surveys. in acquiring new material it is important to widen the circle of donors outside the academic domain especially to market research institutions and government agencies which have many interesting databases suitable for re-use. an agreement has been reached with the republican office of information on the use of data from surveys on the eu and the monthly politbarometer series of surveys (pb). often social scientists have difficulties obtaining access to government statistical data and one of adp’s priorities is to define with the office of statistics the accessibility of its data in a form suitable for use in the social sciences. since 1990 independent commercial market research institutes have been established in slovenia, which produce different kinds of surveys for particular clients which are also interesting for re-use. references: janez stebe 2000. adp social science data archive in slovenia. in: za informationen 47, zentralarchiv für empirische sozialforschung an der universität zu köln. arhiv druzboslovnih podatkov = social science data archive. 2000. cenik = cost. ljubljana: dokumenti adp. http://rcul.uni-lj.si/~fd_adp/dokumenti/cenikadp.htm. barbie, a., boh, k., verba, s., nie, n. h. , and kim, j. 1971. political participation and equality in seven nations, 1966-1971 [computer file]. ljubljana: producer institut za sociologijo in filozofijo. = university of ljubljana (slovenia). institute of sociology and philosophy. ann arbor, mi: producer and distributor inter-university consortium for political and social research, 2000. icpsr studyno = 7768 boh, k., galtung, j., ornauer, h. 1967. podoba sveta v letu 2000. images of the world in the year 2000. [computer file]. ljubljana: producer institut za sociologijo in filozofijo. = university of ljubljana (slovenia). institute of sociology and philosophy. colchester, essex: distribution the data archive. sn: 67024. ljubljana: distribution adp social science data archive, 1998. adp idno: iw200067. comparative study of electoral systems (cses) planning committee. 1996. mednarodna raziskava volilnih sistemov = the comparative study of electoral systems. [data file for slovenia]. ljubljana : production faculty of social science, cjmmk, junij 1996 : production, distribution adp social science data archive, 2000. adp idno: csessi96. cunningham, g. 1997. central and eastern euro-barometer 8, 1997. bruselj: komisija evropske skupnosti. european commission producer. köln : production , distribution za zentralarchiv für empirische sozialforschung, 1999; ljubljana : distribution adp social science data archive, 1999. adp idno: ceeb97. iassist quarterly summer 2001 27 van dijk, j.j.m, and mayhew, p. international victimization survey, 1988, 1992 and 1997 [computer file]. the hague, netherlands: dutch ministry of justice, 1997. ljubljana : distribution adp social science data archive, 1999. adp idno: icvs97. international social justice project (isjp). international social justice project, 1991 [computer file]. ann arbor, mi: duane f. alwin, david m. klingel, and merylin dielman, university of michigan, institute for social research, program in socio-environmental studies; ljubljanna: rus, veljko in antoncic, vojko, faculty of social science, center za proucevanje druzbene blaginje = social welfare research centre [producers], 1993. international social survey programme (issp): international social survey programme 1992: inequality ii [computer data file]. köln : production , distribution za zentralarchiv für empirische sozialforschung, 1993; ljubljana : distribution adp social science data archive, 1998. adp idno: issp92. international social survey programme (issp): international social survey programme 1999: inequality iii [computer data file]. köln : production , distribution zentralarchiv für empirische sozialforschung, 2000; ljubljana : distribution adp social science data archive, 2000. adp idno: issp99. kozuh-novak, m., obersnel-kveder, d.,cerni_istenic m.sircelj, v. in vehovar, v. 1998. rodnostno vedenje slovencev. ljubljana: znanstvenoraziskovalni center sazu. mlinar, z., jacob, p., teune, h. in jerovaek, j., makarovic, j. 1966 1991. vrednote lokalnih voditeljev mednarodna raziskava o politicni participaciji.= democracy and local governance international studies of values in politics. production international social science council. ljubljana : distribution adp social science data archive, 1998. adp idno: dlg66-91. stebe, j. 1999. “izkoriscanje zapuscine slovenske empiricne sociologije za danaanje namene v okviru sekundarne analize.” druzboslovne razprave. xv, no. 3031, october. stebe, j., omerzu, m. 1999. “izkusnje z uporabo xml-ja pri opisovanju druzboslovnih podatkov.” in: cene bavec, matjaz gams. informacijska druzba is’99: zbornik mednarodne multi-konference, october, pp. 59-62. jambrek, p., tos, n. in skupina. 1987. razredna bit sodobne jugoslovanske druzbe, 1987 = class composition of the contemporary yugoslav societies, 1987 [computer file]. ljubljana: production, fspn, cjmmk : distribution adp social science data archive, 1999. adp idno: kb87. svetlik, i. et al.: kvaliteta zivljenja v sloveniji 1994 : retrospektivna studija 1974-1994 = level of living survey in slovenia 1994. retrospective study 1974-1994. ljubljana: production faculty of social science, center za druzbeno blaginjo = social welfare research centre, 1994. ljubljana: distributionb adp social science data archive, 1998. adp idno: lol94. saksida, s., caserman, a. and petrovi, k. 1974. “social stratification and mobility in yugoslav society.” v: toronto 1974. some yugoslav papers presented to the 8th world congress of i. s. a.. ljubljana: university of ljubljana. szalai, a., boh, k. and saksida, s. 1966. time budget study. [computer data file]. budapest: hungarian academy of sciences. ljubljana: producer institut za sociologijo in filozofijo. = university of ljubljana (slovenia). institute of sociology and philosophy. köln : distribution za zentralarchiv für empirische sozialforschung, 1993; ljubljana : distribution adp social science data archive, 1998. adp idno: tbs66. tos, n. et al.: serija slovensko javno mnenje 1968 – 1999 = slovene public opinion survey series. [separated data files]. ljubljana: production university of ljubljana, cjmmk : distribution adp social science data archive, 2000. adp idno: sjm plus year and number. tos, n. et al. 1974. socioloski vidiki migracij slovenskih delavcev v zr nemcijo. zdomci. = migrations of slovene workers in the federal republic germany. emigrants. ljubljana: production university of ljubljana. cjmmk : distribution adp social science data archive, 1999. adp idno: migzdo74. tos, n. et al. 1976. socioloaki vidiki migracij slovenskih delavcev v zr nemcijo. povratniki. migrations of slovene workers in the federal republic germany. remigrants. ljubljana: production university of ljubljana, cjmmk : distribution social science data archive, 1999. adp idno: migpov76. tos, n. et al. 1976. socioloaki vidiki migracij slovenskih delavcev v zr nemcijo. pari. = migrations of slovene workers in the federal republic germany. pairs. ljubljana: production university of ljubljana, cjmmk : distribution adp social science data archive, 1999. adp idno: migpar76. tos, n. et al.: slovensko javno mnenje = slovene public opinion survey, 1990-1998. 1998. repeated questions from politbarometer series. [data file]. ljubljana: production university of ljubljana, cjmmk : production, distribution adp social science data archive, 1999. 28 iassist quarterly summer 2001 adp idno: sjmpb_98. tos, n., de moor, r., inglehart, r. 1999. eurpean/world values survey. [computer data file]. ljubljana: production university of ljubljana, cjmmk : distribution adp social science data archive, 2000. adp idno: evs_99. tos, n., zulehner, p. m., tomka, m. 1997. aufbruch new departures, 97: international research on religion and attitudes toward church. [computer data file]. ljubljana: production university of ljubljana., cjmmk : distribution adp social science data archive, 1999. adp idno: abr97. ule, m. in skupina. 1985-1999. mladina: serija raziskav. [separated data files]. ljubljana: production faculty of social science, center za socialno psihologijo = social psychology research centre : distribution adp social science data archive, 2000. adp idno: mla and year. vehovar, v. 1999. raba interneta v sloveniji. research on internet in slovenia. ljubljana: center za metodologijo in informatiko pri fakulteti za dru~bene vede, univerza v ljubljani. ljubljana: distribution social science data archive, 1999 [distribution]. adp idno: ris99. vreg, f., barbie, a., jezernik, m. in tos, n. 1962. mks anketa o masovnih komunikacijskih sredstvih. = survey on mass communications in slovenia 1962. ljubljana: producer inatitut za sociologijo in filozofijo. = university of ljubljana (slovenia). institute of sociology and philosophy. 1. janez stebe, university of ljubjana, adp socia science data archive, kardeljeva pl. 5, ljublijan si-1000, slovenia. janez.stebe@guest.arnes.si 1/20 hennesy, cody; kubas, alicia; mcburney, jenny (2023) taking count: a computational analysis of data resources on academic libguides. iassist quarterly 47(2), pp. 1-20. doi: https://doi.org/10.29173/iq1040 the creative commons-attribution-noncommercial license 4.0 international applies to all works published by iassist quarterly. authors will retain copyright of the work and full publishing rights. taking count: a computational analysis of data resources on academic libguides in the u.s. cody hennesy1, alicia kubas2, and jenny mcburney3 abstract the libguides platform is a ubiquitous tool in academic libraries and is commonly used by librarians to compile and share lists of recommended social science numerical data resources with users. this study leverages the machine-accessible nature of the libguides platform to collect links to data and statistical resources from over 10,000 libguide pages at 123 r1 research institutions in the united states. after substantial data cleaning and normalization, an analysis of the most common resources on those guides provides a unique window into the data repositories, libraries, archives, statistical data platforms, and other machine-readable data sources that are most popular on academic library guides. results show that freely available resources from u.s. government agencies are among the most common to be included on data and statistical resources guides across institutions. resources requiring paid licenses or memberships for full access, such as statistical insight (proquest), social explorer, and icpsr are linked to most frequently overall, regardless of the percentage of institutions that include them. findings also suggest that libraries are more likely to share traditional licensed statistical resources (e.g., cambridge’s historical statistics of the united states) and collections of simple charts and graphs (e.g., statista) than more robust and complex microdata resources (e.g., ipums). keywords data reference, libguides, data librarian, web scraping introduction over the past several decades, as user expectations for data support have grown, academic libraries have increasingly served as hubs for data and statistical resource discovery and assistance. at the same time, the tools that librarians use to share online resources with their users have shifted significantly, with the libguides platform becoming a ubiquitous resource discovery platform in academic libraries. library guides focused on numerical data and statistics resources have proliferated at a rapid pace, alongside more traditional disciplinary library guides, and currently represent a significant platform for researchers to navigate the sometimes-vexing world of data resource discovery. indeed, finding and utilizing datasets in the social sciences can prove challenging for researchers who are not yet familiar with how data are organized, with the varieties of data types such as “published statistics, microdata, macrodata, survey data, longitudinal data, geospatial data,” or with where datasets on particular topics can be found (rice and southall, p. 52, 2016). along similar lines, data resources can be difficult for even information professionals to understand, as bauder points out: “trying to find data https://doi.org/10.29173/iq1040 https://about.proquest.com/en/products-services/statistical-insight/ https://www.socialexplorer.com/ https://www.socialexplorer.com/ https://www.icpsr.umich.edu/web/pages/ https://hsus.cambridge.org/hsusweb/hsusentryservlet https://www.statista.com/ https://creativecommons.org/licenses/by-nc/4.0/ 2/20 hennesy, cody; kubas, alicia; mcburney, jenny (2023) taking count: a computational analysis of data resources on academic libguides. iassist quarterly 47(2), pp. 1-20. doi: https://doi.org/10.29173/iq1040 can be a frustrating experience” for librarians who are more often familiar with bibliographic discovery tools (p. 11, 2014). the directories of data resources shared on libguides, then, are not only intended as guides for researchers who are new to data topics but also help librarians themselves keep track of the large number of agencies, commercial vendors, nonprofits, and repositories that provide access to data and statistical resources. while several books and many articles have been published in library and information science outlets to help library workers acquaint themselves with key data resources, the libguides platform itself has become a compelling primary source for understanding the real-life recommendations of data librarians. this study leverages the widespread presence of online library data guides to compile lists of the most common freely available and licensed data and statistics resources shared by academic libraries. to gather resources from library guides, the scope of institutions included was first limited to a subset of universities and colleges in the united states designated as r1 institutions, described as ‘doctoral universities very high research activity’ by the 2018 carnegie classification of institutions of higher education (american council of education, 2022). libraries at r1 institutions—where doctoral students and faculty often require access to social science data—consistently devote resources to data research support, and frequently use libguides to organize access to data resources. due to the popularity of springshare’s libguides platform in u.s. academic libraries it was possible to automate the collection of data resources from the relevant guides. in 2021, for example, 91% of 799 academic libraries surveyed, including 95% of 132 doctoral (r1) institutions, used libguides (neuhaus, et al.). this ubiquity makes the platform an excellent primary source for macro-analyses of the kinds of resources that librarians share for specific disciplines or topical research areas. the shared html structure across the platform provides consistent access points for systematic downloads across institutional boundaries, creating opportunities for quick and thorough data collection. ultimately, resources from 10,448 guide pages related to data and statistics were collected from 123 of 131 r1 institutions (93.89%). metadata related to 186,952 non-unique links were initially collected from these guides and, after substantial cleaning and normalization, 64,131 unique resources were analyzed to compile lists of the most common data resources across several categories. the final compilation of top resources leveraged both urls and normalized link names to identify the most common data resources at this specific subset of academic libraries in the u.s. this is the first study of its kind to examine and compile data resources from real-world library guides as a path to explore the resources that data librarians find most essential to share with their users. literature review as the available resources for data and statistics have grown and evolved over the last few decades, so have the strategies that librarians use for discovering these resources and highlighting them for users and other librarians in the field. the rise of the internet and the ability to provide online data access in more accessible files and formats vastly increased the variety and availability of data and statistical resources (kellam and peter, 2011). in the 1980s and 1990s, the focus shifted from solely relying on print data and statistical resources to incorporating more online websites and resources. this ushered in more library trade publications where librarians focused on website lists and recommended sites for statistical resources, such as link-up and online, both magazines from information today (berinstein, 1998; o’leary, 2000). https://doi.org/10.29173/iq1040 3/20 hennesy, cody; kubas, alicia; mcburney, jenny (2023) taking count: a computational analysis of data resources on academic libguides. iassist quarterly 47(2), pp. 1-20. doi: https://doi.org/10.29173/iq1040 in the current landscape, there are many potential avenues for finding data and a proliferation of sources and formats to search or consult along the way. data can encompass both quantitative and qualitative data as well as nonnumeric sources like textual data, audio files, and other corpora (johnson, 2019). as for more traditional numerical data and statistical information, sources span government surveys and databases, international agencies, non-profits, private organizations, researchers and academic journal articles, data repositories, and more (johnson, 2019; kellam & peter, 2011; bauder, 2014; geraci et al., 2012). in particular, government data from local to international levels is becoming more commonly shared via open data portals (huck, 2020). librarians have compiled data guides and composed specific resource reviews to help others navigate the complex world of data reference. multiple book-length guides exist on the topic: kellam and peter compiled a comprehensive list of basic sources in 2011, covering paid, free, and hybrid resources, while bauder published a reference guide to freely available, online data sources in 2014. more recently, johnson published a practical guide for data librarians in 2019, which includes a lengthy appendix of freely available data sources. in addition to these more exhaustive lists of data sources, resource reviews for specific resources have been published in choice, the charleston advisor, and other venues throughout the 2000s and continue today (o’leary, 2000; carroll, 2001; stark and laguardia, 2003; feldmann, 2011; geck, 2013; jakub, 2014; verma, 2014; werner, 2014; geck, 2015; rodriguez, 2016; geck, 2020). data resources are also sometimes published as part of larger lists of recommended reference resources (etkin and coutts, 2015). in addition, recommendations for data sources in particular disciplines or subject areas have been published in a variety of venues on business, education, public health, and art topics (e.g., boslaugh, 2007; smith, 2008; brass education committee, 2012; mcnulty, 2013). iassist quarterly published a double issue in 2009/2010 focused on discipline-specific data with articles providing overviews and lists of sources on various topics, including the american community survey, microdata on developing countries, and sources for international labor data, among other topics (bordelon). librarians often rely on library research guides to help organize the complex landscape of data and statistics resources across different disciplines, and surface both subscription and free online resources for researchers’ use. the most widely used platform for library guides is springshare’s libguides, which is used by over 5,700 institutions across 105 countries (springshare, 2022). as early as 2011, kellam and peter noted that libguides are a key resource for many university libraries, and that “the libguides ‘community search’ function can be a quick way to find a librarian-created research guide to locating statistics on a particular geography or topic, such as health, business, or finance” (p. 100). since libguides are so widely used, the existing library literature includes many best practices for designing and organizing research guides and addresses how librarians have attempted to optimize libguides for sharing data resources specifically (hoffman, 2015; wheatley et al., 2020). guides, in general, have been used for data collection management purposes—identifying existing collection gaps, for example—and tracking statistics on frequently used resources (rice and southall, 2016). because many data resources are born digital and available online, highlighting resources on a platform like libguides is essential for access and use by researchers (ibid.). however, library research guides also present significant challenges, especially with regards to keeping resources up to date and the proliferation of broken links. librarians describing their methods for guide maintenance note that this work can be quite tedious and is often a low priority when more pressing work arises (ornat et al., 2021). however, because libguides are considered an important https://doi.org/10.29173/iq1040 4/20 hennesy, cody; kubas, alicia; mcburney, jenny (2023) taking count: a computational analysis of data resources on academic libguides. iassist quarterly 47(2), pp. 1-20. doi: https://doi.org/10.29173/iq1040 portal for researcher discovery, many librarians have made efforts to address these entropic forces, holding events such as libguides parties to encourage better guide maintenance (ibid.). varying institutional contexts, combined with the data needs of specific institutional users, leads to a wide variety of data service models. some data reference services include dedicated data librarians with different areas of data expertise; others have a single designated data librarian working in conjunction with other subject librarians; while others have no official data librarians and rely on their subject librarians to gain data expertise in their subject areas (bordelon, 2009/2010; geraci et al., 2012; rice and southall, 2016; foster et al., 2019). additionally, reference service points may receive data-related questions regardless of who is currently staffing the desk. research guides for data and statistics can be essential for library staff who are helping researchers answer data-related questions. data and statistics libguides are frequently used as data directories by both library patrons and library staff. library and information science scholars have frequently looked to libguides as primary sources for identifying resources that are key to particular fields, often using manual counts of resources from dozens of disciplinary guides to embark on content analyses. of these studies—looking at guides in theater (furay, 2018), theology (van dyk, 2015), electrical engineering (osorio, 2014), geology (dougherty, 2013), physics (mccormick, 2020) and nursing (stankus and parker, 2012)—anywhere from 37 to 100 guides have been examined at a time, often from a sample taken from a larger collection of relevant guides. along similar lines, common resources have been compiled from topical guides devoted to 3d printing (horton, 2017) and those aimed at physician assistants (johnson and johnson, 2017). many of these studies break resources down by format, reporting on the most common databases, books, ebook collections, journals, and websites that were listed on the subject guides. of those looking at databases, research studies have compiled and analyzed lists ranging from 72 to 143 unique databases. while this study does not discriminate by resource format, by leveraging computational methods it significantly expands upon the range of both the number of guides (10,448 guide pages) and number of unique resources (64,131) under study. it is also the first to examine the presence of data and statistical resources on academic libguides. methodology data and metadata from data and statistical resource links that are shared on libguides were processed in three broad steps: data collection, cleaning, and analysis. data collection and analysis were performed using python in a jupyterlab computing environment, while portions of the cleaning and normalization processes were performed using both python and openrefine. while data scientists joke that 90% of data research projects consist of data cleaning, with only 10% of the work involving analysis, this project, unfortunately, reflected an even more lopsided balance. a significant amount of the labor for this study involved the collection, cleaning, and normalization of inconsistently structured data. the analysis itself consisted of a comparatively straightforward compilation of sums, averages, and percentages to show the most common data and statistics resources on academic library guides. data collection data were ultimately collected from 123 carnegie r1 institutions using a list of libguides urls previously compiled by hennesy and adams in a study of organizational practices for libguides at academic libraries (2021). while 124 of 131 r1 institutions used the libguides platform, one of those https://doi.org/10.29173/iq1040 5/20 hennesy, cody; kubas, alicia; mcburney, jenny (2023) taking count: a computational analysis of data resources on academic libguides. iassist quarterly 47(2), pp. 1-20. doi: https://doi.org/10.29173/iq1040 institutions—columbia university—was dropped from the study since the libguides search interface on the site did not function such that guides related to data and statistics resources could be accurately identified. the final dataset under analysis reflected data and statistics guides from 123 institutions4. data and metadata were scraped from libguides using the requests, beautiful soup, and selenium python packages, following a determination that the research constituted a fair use (reitz, 2015; richardson, 2015; selenium developers, 2021). the data was collected in two general steps on november 18 and december 8, 2021. first, urls for all guide pages that related to data or statistics were extracted. no distinction was made between various guide types (e.g., course, subject, or topic guides), and urls were compiled at the level of individual pages, also often referred to as ‘tabs.’ for an economics guide with a tab called data resources, for example, the url for the data resources page/tab was collected. second, the resource links from every guide page collected in the previous step were compiled. throughout both web scraping steps, several ‘pre-cleaning’ processes were also undertaken to increase the relevance of the compiled data. in the first phase of data collection, the terms data and statistic5 were systematically submitted as keyword searches of the libguides platform at each institution, using selenium to page through the full set of search results. text string matches for the terms data or statistic were compared against the title of each search result, which represented a single page/tab from a guide. when a match was found, the url and title for the guide page (and parent guide) was collected and stored in a pandas dataframe (reback et al., 2022). an initial set of 22,514 guide pages was refined in several ways. first, the results were deduplicated to eliminate 7,029 identical guide pages that were retrieved by searches for both terms. second, the authors scanned the initial results to identify keywords in guide and page titles to identify and remove guides that were unlikely to provide directories of data or statistical resources. the authors identified text strings from the initial search results that indicated guides focused on topics such as software tools, data visualization, data management, data sharing, data publishing, data analysis, text mining, citations, reproducibility, workshops, tutorials, and more. matching on those text strings eliminated an additional 4,233 guide pages. this likely eliminated some relevant resources from the final dataset, but it was an essential step to improve the overall signal to noise ratio. finally, 118 guide pages including the term mathematics in the title were removed (unless they also included the term data), due to the common appearance of subject guides dedicated to mathematics and statistics, which primarily focused on bibliographic resources. the second phase of data collection involved scraping content links from the 11,134 guide pages compiled in the previous step. links from the page header and footer, guide navigational links, and links in librarian profile boxes were skipped, while the collection focused solely on links shared in the primary content boxes of each guide. each link was ultimately represented as a row in a pandas dataframe with columns for the resource url (e.g., https://libguides.asu.edu/ipoll) and name (e.g., ‘roper center for public opinion research (ipoll)’), along with the url from which the resource was collected, and libguides’ site id for the institution. https://doi.org/10.29173/iq1040 https://libguides.asu.edu/ipoll 6/20 hennesy, cody; kubas, alicia; mcburney, jenny (2023) taking count: a computational analysis of data resources on academic libguides. iassist quarterly 47(2), pp. 1-20. doi: https://doi.org/10.29173/iq1040 data cleaning data cleaning steps were taken to both reduce the number of irrelevant resource links in the final dataset and to normalize the names used to reflect those resources, so that they could be counted accurately. initial steps to drop irrelevant rows from the dataset were conducted in python, followed by a series of filtering, faceting, and clustering efforts in openrefine, with the aim of renaming resources consistently across the dataset. beginning with a set of 227,639 resource links, 4,146 resources were dropped because either the url or name of the resource was empty (common for image links). another 1,119 resources were removed when a resource url matched a guide url that was already present in the dataset to exclude links to other libguides from the analysis. the resource names were then converted to lowercase and stripped of extra whitespace to enable names to be clustered regardless of small differences in text strings. resources were exported to a google sheet where the authors manually explored the data, seeking to identify keywords that could be used to filter out common irrelevant resources. regular expression matches for specific text strings were used to remove 24,110 links with terms related to specific software tools, and resources related to data visualization, data management, data sharing, data publishing, data analysis, text mining, qualitative data analysis, citation tools, workshops, and tutorials. regular expressions were also constructed to identify substrings from resource urls that identified irrelevant resources, excluding 10,594 more links. url substrings were identified, for example, to remove links that did not point to other websites (e.g., urls that started with mailto:), root domains that were irrelevant to the project (e.g., libapps, youtube, and creativecommons.org), and frequently occurring links from specific institutions (e.g., links to library services). an additional 536 resources were removed that had urls that did not begin with http or had malformed url schemes, and 182 resource names that had no alpha-numeric characters (usually punctuation marks) were removed. by dropping resources from the dataset as outlined above, the number of guide pages ultimately reflected in the final analysis was also reduced from 11,134 to 10,448. at this stage, python was used for an initial round of resource name normalization, before exporting the data to openrefine. url schemes and common proxy prefixes were split from resource urls to create a list of simplified urls. the resource names for the 1,000 most common simplified urls in the dataset were re-assigned using a python dictionary which replaced the name as extracted from the libguide with a controlled term for the resource. a resource with the string dataplanet.sagepub.com in the url, for example, was assigned the name data planet (sage), and resources pointing to cdc.gov/nchs were renamed nchs (u.s. national center for health statistics). next, the dataset was exported to a csv file and imported into openrefine for further normalization of the resource names. similar resource names were identified and merged using the ‘cluster & edit’ feature, with the default ‘key collision’ method and ‘fingerprint’ keying function selected. each cluster was examined and similar resource titles such as global health (ipums) and ipums global health were merged under a single name. while this unsupervised method had a similar outcome as the normalization of names using regular expressions in python, it had the added benefit of identifying similar titles that the authors had not specifically looked for. next, text facets in openrefine were created of the most common url domains in the dataset. resources for the 150 most common domains were selected one at a time, faceted by their simplified urls, and then renamed en masse. by creating a facet for resources with a domain of www.census.gov, for example, resources such as www.census.gov/acs/www and https://doi.org/10.29173/iq1040 https://dataplanet.sagepub.com/ https://www.cdc.gov/nchs/ https://adminliveunc-my.sharepoint.com/personal/mhayslet_ad_unc_edu/documents/iassist/pubs%20cmte/iq/content/47-2/www.census.gov https://adminliveunc-my.sharepoint.com/personal/mhayslet_ad_unc_edu/documents/iassist/pubs%20cmte/iq/content/47-2/www.census.gov/acs/www 7/20 hennesy, cody; kubas, alicia; mcburney, jenny (2023) taking count: a computational analysis of data resources on academic libguides. iassist quarterly 47(2), pp. 1-20. doi: https://doi.org/10.29173/iq1040 www.census.gov/programs-surveys/acs could be simultaneously renamed u.s. census: american community survey. throughout this process a list of common proxied resources was compiled, including resources such as social explorer and passport (euromonitor international). resource names were filtered using a variety of terms common to each specific resource or platform, and then renamed using the preferred form. filters for the terms euromonitor, passport, and gmid, for example, were each used to find possible matches for euromonitor international’s passport platform, all of which were then renamed passport (euromonitor international). finally, the ‘cluster & edit’ feature was repeated at the end of the normalization process to correct for resources that had been imprecisely renamed during cleaning. throughout the normalization process in openrefine far too many judgments about specific resources were made to document in full here. each decision was guided by the intention of “extract[ing] the useful signal from the noise,” with regards to how librarians named links to data and statistics resources on their libguides (au, 2020). ultimately, 16.81% of the unique resource names in the dataset were normalized, reducing the number of unique resource names from 77,094 (following the initial data cleaning in python) to 64,131. data analysis a csv of the normalized data was exported from openrefine and merged with the original dataframe in pandas, representing 186,952 resources from 10,448 unique guide pages6. common pandas functions were used to count, group, and sort values in the dataset in a variety of combinations. descriptive statistics were generated to understand the scope of the data, looking at the number of unique resource names before and after normalization, the number of guides and resources included in the final dataset, and the average, maximum, and minimum number of resources included from each institution. tables were generated by sorting the data by the most common resource names, domains, top-level domains, and simplified urls. the percentage of institutions that included each of the 500 most common resources was calculated and sorted to show the resources that were most commonly found across institutions (see table 1). to provide a better picture of the freely available websites reflected in the dataset, a separate count by root domain was generated (see table 2). root domains exclude both url paths and subdomains, allowing for the compilation of resources such as www.census.gov, data.census.gov/cedsci, and factfinder.census.gov under the shared root domain of census.gov. an unfiltered list of the most common root domains, however, also includes many urls for library catalogs (exlibrisgroup.com), single sign-on servers (openathens.net), proxies, and link resolvers, often appearing under the root domain for an academic institution (e.g., harvard.edu, upenn.edu). because this kind of root domain obscures the resource they point to, they were excluded from table 2, essentially dropping licensed resources from this view of the data. for example, harvard.edu had a total of 4,480 results, but was excluded because the two most common harvard full domains were nrs.harvard.edu (1,120 results) and id.lib.harvard.edu (1,069 results), neither of which help us identify links to data or statistics resources. additionally, while harvard dataverse is an important resource, the dataverse.harvard.edu domain had only 478 results, while other top harvard domains pointed primarily to library guides (382 results) and the library catalog, hollis (123 results). on the other hand, umich.edu remained on the top root domain list since the most common university of michigan domain in the dataset was for icpsr.umich.edu (2,222 results), a relevant data resource. finally, to better understand the role of paid resources as sources of data and statistics on r1 libguides, a spreadsheet of the 500 most common resources in the dataset was annotated manually https://doi.org/10.29173/iq1040 https://adminliveunc-my.sharepoint.com/personal/mhayslet_ad_unc_edu/documents/iassist/pubs%20cmte/iq/content/47-2/www.census.gov/programs-surveys/acs https://www.socialexplorer.com/ https://adminliveunc-my.sharepoint.com/personal/mhayslet_ad_unc_edu/documents/iassist/pubs%20cmte/iq/content/47-2/www.census.gov https://exlibrisgroup.com/ https://www.openathens.net/ https://www.harvard.edu/ https://id.lib.harvard.edu/ https://dataverse.harvard.edu/ https://umich.edu/ https://www.icpsr.umich.edu/web/pages/ 8/20 hennesy, cody; kubas, alicia; mcburney, jenny (2023) taking count: a computational analysis of data resources on academic libguides. iassist quarterly 47(2), pp. 1-20. doi: https://doi.org/10.29173/iq1040 to note whether a resource was paid, free, or hybrid. the hybrid tag was applied to resources that included any content that was not freely available, even when most of the content was free and public (e.g., icpsr). the free tag was used for both open resources, as well as those that required the creation of a free account to download data. the twenty most common paid and hybrid resources were then extracted (see table 3). results at least one guide page with data or statistic in the title of a tab or the guide was found for all 123 institutions in the study. a mean of 84.94 guide pages were included from each institution, with a maximum of 336 pages from a single institution, a minimum of one from another, and a standard deviation of 64.15. statistics reflecting the number of unique resources across institutions, unfortunately, are not accurate reflections of the number of actual data and statistical resources due to a significant amount of noise that remained present in the resource lists following cleaning and normalization. for example, many links that appeared in small quantities, such as non-data resources (e.g., duo authentication), local services or spaces (e.g., borchert map library), and vaguely named resources (e.g., history or publications) were not removed during cleaning and so are included in the summary statistics for resources overall. we can still make founded observations from the top results, however, since specific ‘noisy’ resources are too uncommon to show up in the most popular 500 results. it's harder to make claims based on the entire dataset due to this long tail of noisy results, however. bearing that in mind, there was a mean of 791.96 unique resources found per institution, with a maximum of 3,284 and a minimum of 21 unique resources found at specific institutions. the standard deviation of unique resources per institution was 638.44, suggesting a normal range of 153.52 to 1,430.40 data and statistical resources per institution. table 1 lists the twenty data and statistics resources that were most common on libguides across r1 institutions, while an online supplementary appendix7 provides a sortable list of the 200 most common resources by overall count. the two resources that appeared most frequently across institutions were icpsr and data.gov, both of which were included on guides from 94.31% of institutions. icpsr was also the resource with the most overall links (1,834) across the entire dataset. of the top twenty data and statistics resources that were most common across institutions, fifteen were links to free and public resources, ten of which were to u.s. government websites (all following references to “top” or “common” resources in this paragraph refer to their frequency across institutions). the most common u.s. government resource, data.gov, compiles open data from a variety of government agencies. the u.s. census bureau was the source of three of the top ten resources, while most of the other top u.s. government sites represented agencies dedicated specifically to statistics and analysis: the national center of health statistics, national center for education statistics, bureau of labor statistics, bureau of justice statistics, and bureau of economic analysis. freely available resources among the twenty most common that were not from u.s. government sites were from the united nations (undata), the european commission (eurostat), the world health organization (the who global health observatory), the world bank (world bank data), and the international monetary fund (imf data). five of the top twenty resources were for platforms on which at least some data requires a paid subscription or membership: icpsr8, roper center ipoll, social explorer, statistical abstracts of the united states (proquest), and historical statistics of the united states (cambridge). when sorting the same list by the overall number of links each resource had across guides, however, seven of the top ten resources were for platforms from which at least some data required paid access: https://doi.org/10.29173/iq1040 https://www.icpsr.umich.edu/web/pages/ https://chennesy.github.io/lg_data/ https://www.icpsr.umich.edu/web/pages/ https://data.gov/ https://www.icpsr.umich.edu/web/pages/ https://data.gov/ https://data.un.org/ https://ec.europa.eu/eurostat https://www.who.int/data/gho https://www.who.int/data/gho https://data.worldbank.org/ https://www.imf.org/en/data https://www.icpsr.umich.edu/web/pages/ https://ropercenter.cornell.edu/ipoll/ https://www.socialexplorer.com/ https://about.proquest.com/en/products-services/statabstract/ https://about.proquest.com/en/products-services/statabstract/ https://hsus.cambridge.org/hsusweb/hsusentryservlet 9/20 hennesy, cody; kubas, alicia; mcburney, jenny (2023) taking count: a computational analysis of data resources on academic libguides. iassist quarterly 47(2), pp. 1-20. doi: https://doi.org/10.29173/iq1040 icpsr (1,834 links), proquest’s statistical insight (1,532 links), social explorer (1,430 links), sage’s data planet (1,318 links), statista (1,104 links), roper center ipoll (1,090), and proquest’s statistical abstracts of the united states (1,058 links). while statistical insight was the second most common resource by link count (1,532), it was present on guides from only 61.79% of the institutions. notably, all links to the eighth most common data resource across institutions, u.s. census bureau’s american factfinder, were dead at the time of data collection, since the site was decommissioned in march 2020 (united states census bureau, 2021). table 1: most common data and statistics resources, sorted by percentage of institutions including them resource name % of institutions overall count example url9 1 icpsr10 94.31% 1834 www.icpsr.umich.edu 2 data.gov 94.31% 954 www.data.gov 3 u.s. census bureau 93.50% 956 www.census.gov 4 u.s. census: data 90.24% 1145 data.census.gov/cedsci 5 undata (united nations) 88.62% 1054 data.un.org 6 nchs (u.s. national center for health statistics) 86.99% 856 www.cdc.gov/nchs 7 nces (u.s. national center for education statistics) 86.99% 546 nces.ed.gov 8 u.s. census: american factfinder 83.74% 720 factfinder.census.gov/faces/na v/jsf/pages/index.xhtml 9 bls (u.s. bureau of labor statistics) 82.93% 535 www.bls.gov 10 roper center: ipoll 80.49% 1090 proxy2.library.illinois.edu/login ?url=ropercenter.cornell.edu 11 eurostat (european commission) 80.49% 501 ec.europa.eu/eurostat 12 social explorer 78.86% 1430 www.socialexplorer.com 13 bjs (u.s. bureau of justice statistics) 78.86% 437 www.bjs.gov 14 statistical abstracts of the united states (proquest) 78.05% 1058 www.libraries.rutgers.edu/inde xes/statabus 15 who: global health observatory (world health organization) 78.05% 491 www.who.int/gho/en 16 world bank: data 76.42% 545 data.worldbank.org 17 historical statistics of the united states (cambridge) 74.80% 684 hsus.cambridge.org/hsusweb 18 imf data (international monetary fund) 74.80% 445 data.imf.org 19 bea (u.s. bureau of economic analysis) 74.80% 422 www.bea.gov 20 cdc: data and statistics 73.17% 361 www.cdc.gov/datastatistics https://doi.org/10.29173/iq1040 https://www.icpsr.umich.edu/web/pages/ https://about.proquest.com/en/products-services/statistical-insight/ https://www.socialexplorer.com/ https://dataplanet.sagepub.com/ https://dataplanet.sagepub.com/ https://www.statista.com/ https://ropercenter.cornell.edu/ipoll/ https://about.proquest.com/en/products-services/statabstract/ https://about.proquest.com/en/products-services/statabstract/ https://about.proquest.com/en/products-services/statistical-insight/ 10/20 hennesy, cody; kubas, alicia; mcburney, jenny (2023) taking count: a computational analysis of data resources on academic libguides. iassist quarterly 47(2), pp. 1-20. doi: https://doi.org/10.29173/iq1040 the analysis above highlights specific webpages and services from major platforms and agency websites but does not provide as clear of a view of the overall popularity of freely available organization websites. table 2 shows the most common url root domains for resources in the dataset by total count, providing a broader perspective of the sources of data and statistical resources in the dataset, while excluding resources that require authentication. it is still the case that u.s. government sites predominate, representing six of the top ten root domains by count. several websites that did not appear in the list of twenty most common resources across institutions (table 1), however, appear here: usda.gov (1,853 results) and nih.gov (1,570 results). this makes sense as the most common full domains for each (nass.usda.gov and ncbi.nlm.nih.gov) reflect only 28.60% and 30.38%, respectively, of the domains associated with the root. in fact, the usda.gov root domain is associated with 69 different full domains, with a mean count of 26.86 per domain. this contrasts with the root domain of bls.gov, which is only associated with four full domains, with a mean count of 516.5 per domain. in other words, usda and nih are agencies with a wide variety of sub-agency pages and services that librarians refer to on their guides. while no single page from these agencies is highly represented, their offerings are important sources of data and statistics when considered holistically. by compiling these counts by root domain, we gain a better sense of the overall importance of parent organizations such as the u.s. census bureau, centers for disease control and prevention, department of education, as well as the world bank, united nations, and the world health organization as sources of data and statistics. table 2: most common root domains (excluding proxies) root domain count top full domain for root count of full domain (% of root domain) 1 census.gov 9,670 census.gov 7,656 (79.17%) 2 cdc.gov 6,072 cdc.gov 5,382 (88.64%) 3 umich.edu 3,691 icpsr.umich.edu 2,222 (60.20%) 4 ed.gov 3,122 nces.ed.gov 2,546 (81.55%) 5 worldbank.org 2,705 data.worldbank.org 1,120 (41.40%) 6 un.org 2,636 unstats.un.org 1,076 (40.82%) 7 bls.gov 2,066 bls.gov 1,928 (93.32%) 8 usda.gov 1,853 nass.usda.gov 530 (28.60%) 9 nih.gov 1,570 ncbi.nlm.nih.gov 477 (30.38%) 10 who.int 1,411 who.int 1,202 (85.19%) table 3 presents a subset of the twenty most common non-free resources, sorted by the number of times they were found across all guides. non-free resources here include databases that are only https://doi.org/10.29173/iq1040 https://www.usda.gov/ https://www.nih.gov/ https://www.nass.usda.gov/ https://ncbi.nlm.nih.gov/ https://www.usda.gov/ 11/20 hennesy, cody; kubas, alicia; mcburney, jenny (2023) taking count: a computational analysis of data resources on academic libguides. iassist quarterly 47(2), pp. 1-20. doi: https://doi.org/10.29173/iq1040 accessible to institutions with subscriptions, as well as ‘hybrid’ platforms that require payment for some subset of data access, though some hybrid platforms provide significant amounts of freely available data. overall, 20.28% of the 500 most common resources (by overall count) were either hybrid or fully paid resources, while 79.72% of the resources were entirely free (see table 4). the paid and hybrid resources represent a diverse group of platforms in terms of publishers or vendors: while two proquest databases (statistical insight and statistical abstract for the united states) appear in the top twenty across institutions, no other single data provider shows up more than once on the list of top paid/hybrid resources. the most common paid or hybrid resources by overall count, as noted previously, are icpsr, statistical insight, social explorer, data planet (sage), and statista. a few databases that primarily collect literature resources (e.g., oecd ilibrary and ebsco’s business source complete) also appear here. paid resources, taken overall, have a significantly higher number of links across guides than free resources do. the mean number of links for a paid resource was 174.43, compared to 104.18 for a free one, and 318.00 for a hybrid resource (see table 4). table 3: most common non-free (paid and hybrid) resources, sorted by overall count resource name % of institutions overall count 1 icpsr 94.31% 1834 2 statistical insight (proquest) 61.79% 1532 3 social explorer 78.86% 1430 4 data planet (sage) 50.41% 1318 5 statista 64.23% 1104 6 roper center: ipoll 80.49% 1090 7 statistical abstract of the united states (proquest) 78.85% 1058 8 oecd ilibrary 69.11% 851 9 historical statistics of the united states (cambridge) 74.80% 684 10 simplyanalytics 44.72% 672 11 policymap 43.09% 645 12 wharton research data services (wrds) 43.90% 470 13 polling the nations 45.53% 313 14 passport (euromonitor international) 34.96% 249 15 china data online 33.33% 209 16 business source complete (ebsco) 34.15% 191 17 data citation index (web of science/clarivate) 32.52% 177 18 economist intelligence unit (eiu) 21.14% 175 19 ibisworld 43.09% 168 20 europa world plus 26.83 161 https://doi.org/10.29173/iq1040 https://about.proquest.com/en/products-services/statistical-insight/ https://about.proquest.com/en/products-services/statabstract/ https://www.icpsr.umich.edu/web/pages/ https://about.proquest.com/en/products-services/statistical-insight/ https://www.socialexplorer.com/ https://dataplanet.sagepub.com/ https://www.statista.com/ https://www.oecd-ilibrary.org/ https://www.ebsco.com/products/research-databases/business-source-complete https://www.ebsco.com/products/research-databases/business-source-complete 12/20 hennesy, cody; kubas, alicia; mcburney, jenny (2023) taking count: a computational analysis of data resources on academic libguides. iassist quarterly 47(2), pp. 1-20. doi: https://doi.org/10.29173/iq1040 table 4: paid, hybrid, and free resource types, of the top 500 resources resource type % of top 500 resources n resources in top 500 number of links to resources overall mean number of links per resource overall paid 17.47% 87 15,175 174.43 hybrid 2.81% 14 4,452 318.0 paid & hybrid combined 20.28% 101 19,627 194.33 free 79.72% 397 41,361 104.18 discussion comparing the inclusion of freely available data and statistical resources to those that require some level of payment for user access highlights that librarians share links to paid and hybrid resources more often than they share free resources on their guides. across all 123 r1 institutions in this analysis, librarians included links to paid resources (with a mean of 174.33 links per resource) and hybrid resources (a mean of 318.0011) on their guides far more often than they included free ones (a mean of 104.18). this phenomenon is especially evident when looking at popular paid resources such as proquest’s statistical insight, which was the second most common by count, with 1,532 links, despite only being available at 61.79% of institutions. it makes sense for institutions with monetary investments in data and statistical resources to make concerted efforts to inform their users of their availability. but the fact that free resources are less commonly shared suggests that they may be underrepresented on resource guides. in this case, a data librarian’s incentives for sharing a particular resource (e.g., improving the library’s return on investment) is less than ideally aligned with users’ needs for finding data and statistics. another contributing factor is likely the fact that libguides ’ a-z database lists, which make it easy to include items across guides at a particular institution, usually focus on licensed resources instead of freely available websites. it is often logistically easier for a guide editor to include and maintain links to licensed resources from the a-z list on their guides, regardless of their research value. electronic resource librarians often manage these “database assets” for their libguides instance. the central management allows a change to a database url or name, for example, to immediately propagate across all guides at the institution. in practice that level of management rarely extends to the kinds of publicly available websites that are most common on data and statistics guides, and instead focuses on subscription resources. for this reason, public data resources are inconsistently named and described across libguides, making it difficult to fully account for the data and statistics resources that are common across academic libraries. one potential area for improvement would be for electronic resource librarians to add the major free data and statistics resources identified by this study to their local a-z lists. a potentially broader opportunity for libraries on the libguides platform would be to take advantage of the ability for librarians to share guides and content links, not just institutionally, but across the entire community of other libguides institutions. this platform attribute, in theory, could support a https://doi.org/10.29173/iq1040 https://about.proquest.com/en/products-services/statistical-insight/ 13/20 hennesy, cody; kubas, alicia; mcburney, jenny (2023) taking count: a computational analysis of data resources on academic libguides. iassist quarterly 47(2), pp. 1-20. doi: https://doi.org/10.29173/iq1040 network of well-defined and managed links to popular resources that librarians could quickly plug into their own guides and then forget about, with ongoing maintenance managed centrally. in practice, due to technical and usability issues related to finding and maintaining resources across libguides instances, however, this would likely require significant development by springshare to be implemented successfully. a lack of organization of links to free sites both within and across institutions contributes to the widespread presence of broken links to obsolete data and statistics resources. in december of 2021, for example, 83.74% of institutions were still sharing links to the u.s. census bureau’s american factfinder site, a site which was decommissioned in march 2020 (united states census bureau, 2021). american factfinder was the eighth most common data resource shared across the institutions analyzed, and 720 links to the resource were still present on institutional guides 20 months after the link ceased to point to a functioning website. the continued inclusion of american factfinder on data libguides after its relatively high-profile decommissioning suggests that maintenance and link rot are as problematic on data and statistics guides as previous research has found for the broader libguides context (ornat et al., 2021). several data and statistical resources were less popular than the authors would have expected, especially compared to other popular resources that seem to be of relatively little research value. free and high-value microdata from ipums, who provide well-maintained granular survey and census data from around the world, were not completely absent from the top resource lists, but links to their resources were not common. there were only 334 links across all guides to the most popular ipums resource collected, the national historical geographic information system (nhgis), which was found at 62.60% of institutions. the main ipums website (ipums.org) was only available from 50.41% of the institutions, with 249 links overall. this is striking compared to the popularity of proquest statistical insight, which was linked to 1,532 times, even though it primarily functions as an index to statistical resources that aren’t available directly from the proquest interface. it’s difficult to understand the enduring popularity of this resource, given the difficulty most users would have finding the sources that are cited without the help of a librarian. this seems to reflect a broader trend in which complex numerical datasets that require more sophisticated data tools or methods to analyze are included on guides less frequently than collections of more traditional tabular statistics (e.g., cambridge’s historical statistics of the united states) or simple charts and graphs (e.g., statista). after icpsr and statistical insight, the three most common paid/hybrid resources—social explorer, data planet (sage), and statista—provide platforms that present data in more accessible formats via dynamic data visualizations, maps, and/or tables. while in some cases the underlying data is exportable to csv or json, these platforms provide outputs such as charts and graphs that pre-digest the data in ways that make data analysis outside of the platform unnecessary. the relative popularity of social explorer may also stem in part from the fact that it only recently (and quietly) transitioned from robust long-term free access to a more limited-term public access via trial accounts. free data resources that are popular for data science applications, such as kaggle and the uci machine learning repository, are other notable absences on academic library guides. the google-owned kaggle provides a repository of user-generated datasets that are frequently used in online tutorials in data analytics and statistics and are often used in data science competitions for developing accurate machine learning models. despite this, only 47 links to kaggle were found on these guides, appearing at only 23.57% of the institutions studied. while kaggle is a free commercial platform, its relative https://doi.org/10.29173/iq1040 https://www.ipums.org/ https://www.nhgis.org/ https://www.ipums.org/ https://about.proquest.com/en/products-services/statistical-insight/ https://about.proquest.com/en/products-services/statistical-insight/ https://hsus.cambridge.org/hsusweb/hsusentryservlet https://www.statista.com/ https://www.icpsr.umich.edu/web/pages/ https://about.proquest.com/en/products-services/statistical-insight/ https://www.socialexplorer.com/ https://dataplanet.sagepub.com/ https://www.statista.com/ https://www.socialexplorer.com/ https://www.kaggle.com/ https://archive.ics.uci.edu/ https://archive.ics.uci.edu/ https://www.kaggle.com/ 14/20 hennesy, cody; kubas, alicia; mcburney, jenny (2023) taking count: a computational analysis of data resources on academic libguides. iassist quarterly 47(2), pp. 1-20. doi: https://doi.org/10.29173/iq1040 absence here might be understood to reflect the relative paucity of guides focused on computer science and data science resources, or as part of a larger pattern in which “unruly” sites are rarely included on data guides. as huck mentions, “the ‘wild west’ of data discovery encompasses author websites and sharing platforms such as github and kaggle. it is considerably harder to find useful data through these websites, because your best search tool is a regular web search” (2020). a similar lack of visibility extends to most official data repositories: while icpsr is a significant exception, the next most common data repository was harvard dataverse, found at 51.22% of institutions, with 234 direct links to the repository (and an additional 244 links to datasets within the repository). this trend likely reflects the fractured landscape in which user-created datasets are stored in any number of local institutional repositories, few of which are large enough to show up across a wide range of guides. rather than link to specific repositories, in fact, many guides include re3data.org, a directory of data repositories, which was linked to 339 times at 66.67% of institutions. while it’s likely that each institution includes links to their preferred local data repository, these fall under the radar when compiled across institutional boundaries. finally, it’s worth noting that the most common data and statistics resources identified across libguides in 2021 are largely accounted for in previously published directories of data resources. all of the resources coded as freely available in table 1, and all of the resources except for ncbi.nlm.nih.gov in table 2, are included in bauder’s 2014 reference guide as either major or minor sources. while bauder, whose book-length guide does not include any paid resources, often identifies resources at a greater level of specificity (pointing to a specific survey from a government agency, for example), a surprisingly high percentage of the most popular free web resources found on data libguides in 2021 were already present in this 2014 collection. while it seems that many of the demands and services related to data librarianship have evolved significantly over the previous decade, the general landscape of free data and statistical sources appears to be fairly stable. limitations and future directions despite the common shared infrastructure across libguides from different institutions, naming practices for links to data and statistical resources across guides are extremely inconsistent. the fact that the urls for licensed resources are generally obscured behind a number of proxies and single sign-on services makes it difficult to accurately account for the presence of paywalled resources using their urls alone, while the prominence of textual links that use terms that are irrelevant to the resource at hand (e.g., “here,” “2017,” “%”) complicates the task of counting the presence of specific resources by name. the data cleaning processes implemented here adjust for those challenges by leveraging both strings from urls and from named links to cluster and merge the same resources together. but given the messiness of the data and the size of the dataset, the long tail of resources shared on these guides was often unor undercounted. essentially, any resource that did not turn up on initial lists of the most common thousand or so urls and resource names was likely to not be normalized, and therefore fell through the cracks in the overall analysis. further study would be required to highlight data and statistical resources that are less commonly shared but have high potential value for academic library audiences. while iterating through the time-consuming series of data cleaning steps detailed above, it occurred to the authors how little consistency there was in how libraries name and describe resources on libguides. given the rich history of consortial library collaborative projects to create shared cataloging https://doi.org/10.29173/iq1040 https://github.com/ https://www.kaggle.com/ https://www.icpsr.umich.edu/web/pages/ https://dataverse.harvard.edu/ https://www.re3data.org/ https://ncbi.nlm.nih.gov/ 15/20 hennesy, cody; kubas, alicia; mcburney, jenny (2023) taking count: a computational analysis of data resources on academic libguides. iassist quarterly 47(2), pp. 1-20. doi: https://doi.org/10.29173/iq1040 frameworks and controlled vocabularies it’s striking that each library on the libguides platform must create their own database assets. for example: of the 1,834 resource-links for icpsr found in this study, there were 131 different link names used, pointing to 220 different urls, some of which were broken. the proliferation of broken links on guides especially signals the potential upside of a collaborative cross-institutional system for creating and sharing database assets, though the hodgepodge of institutional and paid proxy platforms complicates this prospect. it’s a surprise, however, given the prevalence of springshare’s libguides a-z database tool, that springshare does not offer a central service for managing database content across guides. there is potential for an organized effort on the part of a library collective, consortium, or professional organization to partner with springshare to pilot some level of cooperative maintenance of database assets. the time-gains across libraries could be substantial, as would improvements in quality control regarding ongoing maintenance. finally, limiting the initial scope of library guides to be collected to r1 institutions provides an incomplete view of broader trends across academic libraries. one compelling direction for future research in this area would be to look more holistically across guides from different types of academic libraries and consider differences in the kinds of resources that are shared by institution size, type, location, and so forth. for example, do smaller institutions rely more on free and hybrid data resources than academic libraries at r1 institutions? are smaller institutions more likely to take advantage of the community aspect of libguides and repurpose guides from institutions with more resources to devote to compiling data resources? the authors unsystematically noted geographical differences in data resources shared at different institutions (e.g., state data repositories are almost always only available on guides provided by institutions within the state), but a more purposeful look at key differences in resource sharing across libguides could make for a fruitful examination. conclusion there were 10,448 different published libguides pages related to data or statistics topics at 123 r1 institutions in 2021. librarians and other data experts clearly find libguides to be an essential tool for organizing and sharing these kinds of resources. while most common free resources shared on libguides were duplicative of resources compiled for earlier published bibliographic data guides, librarians shared links to licensed data resources more frequently on libguides than they shared free ones. the continued relevance of data and statistics sources noted in earlier studies suggests that resources from many major government agencies and intergovernmental organizations are, as a whole, relatively stable and could be managed more intentionally on libguides. a lack of consistency in the titles and links to data resources both within and across institutions, along with a preponderance of dead links to outdated sites, likewise suggests that academic libraries would have much to gain from some degree of centralized management of their data resources. local inclusion of the most common free data and statistical resources identified in this study on libguides’ a-z database lists, for example, would help ensure that those resources are more visible to users and are better maintained over time. along similar lines, a current over emphasis of licensed data resources on libguides undersells the value of free resources from government sources and points users instead to platforms geared towards a more traditional bibliographic market (e.g., proquest statistical insight). u.s. academic libraries could also better promote data formats such as ipums microdata, which was not found on https://doi.org/10.29173/iq1040 https://www.icpsr.umich.edu/web/pages/ https://about.proquest.com/en/products-services/statistical-insight/ https://www.ipums.org/ 16/20 hennesy, cody; kubas, alicia; mcburney, jenny (2023) taking count: a computational analysis of data resources on academic libguides. iassist quarterly 47(2), pp. 1-20. doi: https://doi.org/10.29173/iq1040 guides from almost half of the institutions studied. it’s likely that there are still data knowledge gaps in libraries at many institutions such that ipums and other valuable resources fall through the cracks while formats that are easier to understand are emphasized instead, leaving the needs of researchers with higher-level data skills unmet. one relatively simple stopgap measure, which could be implemented locally, would be for “accidental data librarians” to draw more freely from guides at institutions that have more robust data knowledge and support. the somewhat shocking preponderance of dead links to american factfinder on these guides—when news of the retirement of the site was widely shared among government documents and data librarian communities— similarly points to the benefit of under-resourced data librarians drawing from guides maintained at other institutions, where librarians may be better equipped to monitor the field of data and statistics resources. overall, libraries at r1 institutions showed remarkable consistency in recommending a similar core list of data resources across their guides. data librarians at many institutions, however, would also benefit from reviewing the list compiled in table 1 (or this more in-depth online table) to ensure that key free resources are available to their users. while local needs differ, it’s hard to imagine that users at the seven institutions (5.69%) that do not link to data.gov on any of their guides would not benefit from that resource, or that users at the 33 institutions (26.83%) without links to the cdc data & statistics platform would not benefit from those tools. while it’s not feasible for an individual librarian to keep up with an ever-expanding catalog of data and statistical resources, the compilation of common data and statistics resources collected here provides a snapshot of data resources that a network of library data peers deems fit to share and promote on their own guides. author statements hennesy: conceptualization, methodology, data collection, cleaning, and analysis; writing: all sections of original draft except for literature review; co-editing: full paper. kubas and mcburney: review of methodology, analysis, and data outputs; data cleaning assistance; writing: literature review; co-editing: full paper. references american council of education (2022) carnegie classification of institutions of higher education: basic classification description. available at: https://carnegieclassifications.acenet.edu/classification_descriptions/basic.php (accessed: 14 june 2022) au, r. (2020) ‘data cleaning is analysis, not grunt work’, counting stuff. available at: https://counting.substack.com/p/data-cleaning-is-analysis-not-grunt (accessed: 19 january 2022). bauder, j. (2014) the reference guide to data sources. chicago: american library association. bordelon, b. (ed.) (2009/2010) ‘the subject content and how researchers use the data’, iassist quarterly, 33(4)/34(1). doi: https://doi.org/10.29173/iq883 https://doi.org/10.29173/iq1040 https://www.ipums.org/ https://chennesy.github.io/lg_data/ https://data.gov/ https://www.cdc.gov/datastatistics/index.html https://carnegieclassifications.acenet.edu/classification_descriptions/basic.php https://counting.substack.com/p/data-cleaning-is-analysis-not-grunt https://counting.substack.com/p/data-cleaning-is-analysis-not-grunt https://counting.substack.com/p/data-cleaning-is-analysis-not-grunt https://doi.org/10.29173/iq883 17/20 hennesy, cody; kubas, alicia; mcburney, jenny (2023) taking count: a computational analysis of data resources on academic libguides. iassist quarterly 47(2), pp. 1-20. doi: https://doi.org/10.29173/iq1040 boslaugh, s. (2007) secondary data sources for public health: a practical guide. cambridge: cambridge university press. dougherty, k. (2013a) ‘the direction of geography libguides’, journal of map and geography libraries, 9(3), pp. 259–75. doi: https://doi.org/10.1080/15420353.2013.779355 eclevia, m.r., fredeluces, j.c.l.t, maestro, r.s., and eclevia jr., c.l. (2019) ‘what makes a data librarian?: an analysis of job descriptions and specifications for data librarian’, qualitative & quantitative methods in libraries, 8(3), pp. 273–290. available at: http://qqmljournal.net/index.php/qqml/article/view/541. (accessed: 14 june 2022) foster, a.k., rinehart, a.k., and springs, g.r. (2019) ‘piloting the purchase of research data sets as collections: navigating the unknowns’, portal: libraries & the academy, 19(2), pp. 315–328. doi: https://doi.org/10.1353/pla.2019.0018 furay, j. (2018) ‘performance review: online research guides for theater students’, reference services review, 46(1), pp. 91–109. doi: https://doi.org/10.1108/rsr-09-2017-0037 garrison, b. and exner, n. (2018) ‘data seeking behavior of economics undergraduate students: an exploratory study’, reference & user services quarterly, 58(2), pp. 103–113. doi: http://dx.doi.org/10.5860/rusq.58.2.6930 geraci, d., humphrey, c., and jacobs, j. (2012) data basics: an introductory text. available at: https://3stages.org/class/2012/pdf/data_basics_2012.pdf. (accessed: 14 june 2022) hennesy, c. and adams, a.l. (2021) ‘measuring actual practices: a computational analysis of libguides in academic libraries’, journal of web librarianship 15(4), pp. 219–242. doi: https://doi.org/10.1080/19322909.2021.1964014 hoffman, s. (2015) ‘data reference and instruction in journalism and the social sciences’, dttp (documents to the people): a quarterly journal of government information practice & perspective, 43(2), pp. 14–17. available at: https://journals.ala.org/index.php/dttp/issue/viewissue/603/360. (accessed 14 june 2022) horton, j.j. (2017) ‘an analysis of academic library 3d printing libguides’, internet reference services quarterly, 22(2/3), pp. 123–131. doi: http://dx.doi.org/10.1080/10875301.2017.1375059 huck, j. (2020) ‘identifying, accessing and evaluating data’, information outlook: the magazine of the special libraries association, 24(1), pp. 4-6. available at: https://scholarworks.sjsu.edu/cgi/viewcontent.cgi?article=1000&context=sla_io_2020 (accessed: 25 march 2022). jackson, r. and stacy-bates, k.k. (2016) ‘the enduring landscape of online subject research guides’, reference & user services quarterly, 55(3), pp. 219-225. doi: https://doi.org/10.5860/rusq.55n3.219 johnson, c.v. and johnson, s.y. (2017) ‘an analysis of physician assistant libguides: a tool for collection development’, medical reference services quarterly, 36(4), pp. 323–333. doi: https://doi.org/10.1080/02763869.2017.1369241 johnson, e.o. (2019) working as a data librarian: a practical guide. denver: libraries unlimited. https://doi.org/10.29173/iq1040 https://doi.org/10.1080/15420353.2013.779355 http://qqml-journal.net/index.php/qqml/article/view/541 http://qqml-journal.net/index.php/qqml/article/view/541 https://doi.org/10.1353/pla.2019.0018 https://doi.org/10.1108/rsr-09-2017-0037 http://dx.doi.org/10.5860/rusq.58.2.6930 https://3stages.org/class/2012/pdf/data_basics_2012.pdf https://3stages.org/class/2012/pdf/data_basics_2012.pdf https://3stages.org/class/2012/pdf/data_basics_2012.pdf https://doi.org/10.1080/19322909.2021.1964014 https://journals.ala.org/index.php/dttp/issue/viewissue/603/360 http://dx.doi.org/10.1080/10875301.2017.1375059 https://go-gale-com.ezp3.lib.umn.edu/ps/retrieve.do?tabid=magazines&resultlisttype=result_list&searchresultstype=multitab&hitcount=1&searchtype=advancedsearchform¤tposition=1&docid=gale%7ca684967313&doctype=article&sort=relevance&contentsegment=zxbk-mod1&prodid=suic&pagenum=1&contentset=gale%7ca684967313&searchid=r1&usergroupname=umn_wilson&inps=true https://go-gale-com.ezp3.lib.umn.edu/ps/retrieve.do?tabid=magazines&resultlisttype=result_list&searchresultstype=multitab&hitcount=1&searchtype=advancedsearchform¤tposition=1&docid=gale%7ca684967313&doctype=article&sort=relevance&contentsegment=zxbk-mod1&prodid=suic&pagenum=1&contentset=gale%7ca684967313&searchid=r1&usergroupname=umn_wilson&inps=true https://scholarworks.sjsu.edu/cgi/viewcontent.cgi?article=1000&context=sla_io_2020 https://doi.org/10.5860/rusq.55n3.219 https://doi.org/10.1080/02763869.2017.1369241 18/20 hennesy, cody; kubas, alicia; mcburney, jenny (2023) taking count: a computational analysis of data resources on academic libguides. iassist quarterly 47(2), pp. 1-20. doi: https://doi.org/10.29173/iq1040 joo, s. and schmidt, g.m. (2021) ‘research data services from the perspective of academic librarians’, digital library perspectives, 37(3), pp. 242–256. doi: https://doi.org/10.1108/dlp-10-2020-0106 kalinowski, a. and hines, t. (2020) ‘eight things to know about business research data’, journal of business & finance librarianship, 25(3/4), pp. 105–122. doi: https://doi.org/10.1080/08963568.2020.1847548 kellam, l.m. and peter, k. (2011) numeric data services and sources for the general reference librarian. cambridge: chandos. kellam, l.m. and thompson, k. (eds) (2016) databrarianship: the academic data librarian in theory and practice. chicago: association of college and research libraries. mccormick, a. (2020) ‘collection development for librarians in a hurry: a survey of the physics resources of the libraries of the association of american universities’, issues in science and technology librarianship 96. doi: https://doi.org/10.29173/istl68 mcnulty, t. (2013) art market research: a guide to methods and sources, 2nd edn. jefferson, north carolina: mcfarland. nelson, m.r.s. (2020) ‘adding data literacy skills to your toolkit’, information outlook: the magazine of the special libraries association, 24(1), pp. 10-11. available at: https://scholarworks.sjsu.edu/cgi/viewcontent.cgi?article=1000&context=sla_io_2020 (accessed: 25 march 2022). neuhaus, c., cox, a., gruber, a.m., kelly, j., koh, h., bowling, c. and bunz, g. (2021) ‘ubiquitous libguides: variations in presence, production, application, and convention’, journal of web librarianship, 15(3), pp. 107–127. doi: https://doi.org/10.1080/19322909.2021.1946457 ohaji, i.k., chawner, b. and yoong, p. (2019) ‘the role of a data librarian in academic and research libraries’, information research, 24(4). available at: http://informationr.net/ir/244/paper844.html. (accessed: 23 march 2022) ornat, n., auten, b., manceaux, r. and tingelstad, c. (2021) ‘ain’t no party like a libguides party: ’cause a libguides party is mandatory’, college & research libraries news, 82(1), pp. 14–17. doi: https://doi.org/10.5860/crln.82.1.14 osorio, n. (2014) ‘content analysis of engineering libguides’, in 2014 asee annual conference & exposition proceedings. indianapolis: asee, pp. 24.318.1-24.318.23. doi: https://doi.org/10.18260/1-2--20209 reback, j. et al. (2022) pandas (1.4.0rc0). computer software. https://zenodo.org/record/5824773. (accessed: 20 january 2022). reitz, k. (2015). requests (2.7.0). computer software. available at: https://pypi.org/project/requests/2.7.0/ (accessed: 14 june 2022) rice, r. and southall, j. (2016) the data librarian’s handbook. chicago: american library association. richardson, l. (2015). beautiful soup (4.8.1). computer software. available at: https://beautifulsoup-4.readthedocs.io/en/latest/ (accessed: 14 june 2022) https://doi.org/10.29173/iq1040 https://doi.org/10.1108/dlp-10-2020-0106 https://doi.org/10.1080/08963568.2020.1847548 https://doi.org/10.29173/istl68 https://go-gale-com.ezp3.lib.umn.edu/ps/retrieve.do?tabid=magazines&resultlisttype=result_list&searchresultstype=multitab&hitcount=1&searchtype=advancedsearchform¤tposition=1&docid=gale%7ca684967313&doctype=article&sort=relevance&contentsegment=zxbk-mod1&prodid=suic&pagenum=1&contentset=gale%7ca684967313&searchid=r1&usergroupname=umn_wilson&inps=true https://go-gale-com.ezp3.lib.umn.edu/ps/retrieve.do?tabid=magazines&resultlisttype=result_list&searchresultstype=multitab&hitcount=1&searchtype=advancedsearchform¤tposition=1&docid=gale%7ca684967313&doctype=article&sort=relevance&contentsegment=zxbk-mod1&prodid=suic&pagenum=1&contentset=gale%7ca684967313&searchid=r1&usergroupname=umn_wilson&inps=true https://scholarworks.sjsu.edu/cgi/viewcontent.cgi?article=1000&context=sla_io_2020 https://doi.org/10.1080/19322909.2021.1946457 http://login.ezproxy.lib.umn.edu/login?url=https://search.ebscohost.com/login.aspx?direct=true&authtype=ip,uid&db=lls&an=140844395&site=ehost-live http://informationr.net/ir/24-4/paper844.html http://informationr.net/ir/24-4/paper844.html https://doi.org/10.5860/crln.82.1.14 https://doi.org/10.18260/1-2--20209 https://zenodo.org/record/5824773 https://pypi.org/project/requests/2.7.0/ https://beautiful-soup-4.readthedocs.io/en/latest/ https://beautiful-soup-4.readthedocs.io/en/latest/ 19/20 hennesy, cody; kubas, alicia; mcburney, jenny (2023) taking count: a computational analysis of data resources on academic libguides. iassist quarterly 47(2), pp. 1-20. doi: https://doi.org/10.29173/iq1040 selenium (4.2.0). computer software. available at: https://pypi.org/project/selenium/. (accessed: 14 june 2022) semeler, a.r., pinto, a.l. and rozados, h.b.f. (2019) ‘data science in data librarianship: core competencies of a data librarian’, journal of librarianship and information science, 51(3), pp. 771–780. doi: https://doi.org/10.1177/0961000617742465 smith, e. (2008) using secondary data in educational and social research. new york: open university press. springshare (2022). libguides community. available at: https://community.libguides.com/ (accessed 14 june 2022) stankus, t. and parker, m.a. (2012) ‘the anatomy of nursing libguides’, science & technology libraries, 31(2), pp. 242–255. doi: https://doi.org/10.1080/0194262x.2012.678222 united states census bureau (2021). transition from aff. available at: https://www.census.gov/data/what-is-data-census-gov/guidance-for-data-users/transitionfrom-aff.html (accessed: 15 january 2022) van dyk, g. (2015) ‘finding religion: an analysis of theology libguides’, theological librarianship, 8(2), pp. 37–45. doi: https://doi.org/10.31046/tl.v8i2.384 wheatley, a., chandler, m. and mckinnon, d. (2020) ‘collaborating with faculty on data awareness: a case study’, journal of business & finance librarianship, 25(3/4), pp. 281–290. doi: https://doi.org/10.1080/08963568.2020.1847553 endnotes 1 cody hennesy (chennesy@umn.edu) is the journalism & digital media librarian at the university of minnesota, twin cities. 2 alicia kubas (akubas@gpo.gov) is a librarian at the us government publishing office. 3 jenny mcburney (jmcburne@umn.edu) is a social sciences librarian at the university of minnesota, twin cities. 4 a list of the institutions included are available, along with the full replication data for this study, in the data repository for the university of minnesota: https://conservancy.umn.edu/handle/11299/228216. 5 these forms of the terms were chosen, in part, because the libguides search interface returned guides with matches such as statistics and statistical for the keyword statistic, and terms such as datasets for the keyword data. 6 python code used for data analysis and the generation of tables is available at https://github.com/chennesy/lg_data/ 7 https://chennesy.github.io/lg_data/ 8 while most datasets on icpsr are freely available, some data from specific datasets (e.g., the american national election study, and the u.s. transgender survey) are only available to users at institutions with paid memberships. for this reason, the authors tagged platforms such as icpsr and statista as ‘hybrid’and tended to group them together with ‘paid’ resources instead of grouping them with freely available sites such as census.gov. 9 the example url column in table 1 refers to the most common url associated with a particular resource in the dataset. any number of urls could be associated with a particular resource. most https://doi.org/10.29173/iq1040 https://pypi.org/project/selenium/ https://doi.org/10.1177/0961000617742465 https://community.libguides.com/ https://doi.org/10.1080/0194262x.2012.678222 https://www.census.gov/data/what-is-data-census-gov/guidance-for-data-users/transition-from-aff.html https://www.census.gov/data/what-is-data-census-gov/guidance-for-data-users/transition-from-aff.html https://doi.org/10.31046/tl.v8i2.384 https://doi.org/10.1080/08963568.2020.1847553 https://github.com/chennesy/lg_data/ https://chennesy.github.io/lg_data/ https://www.icpsr.umich.edu/web/pages/ https://www.icpsr.umich.edu/web/pages/ https://www.statista.com/ 20/20 hennesy, cody; kubas, alicia; mcburney, jenny (2023) taking count: a computational analysis of data resources on academic libguides. iassist quarterly 47(2), pp. 1-20. doi: https://doi.org/10.29173/iq1040 urls for roper center ipoll, for example, are proxied at an institutional level so that the most common url is specific to the institution that linked to ipoll the most often. 10 normalized (uncapitalized) resource names are retained in the tables to enable clearer references to resources in the original dataset. 11 the mean number of links per hybrid resource skews high due to the high number of links to the top three resources (icpsr, statista, and oecd ilibrary) and the relatively few hybrid resources overall. the median number of links for hybrid resources was 67.50, compared to 66.00 for paid, and 53.00 for free. https://doi.org/10.29173/iq1040 https://ropercenter.cornell.edu/ipoll/ https://www.icpsr.umich.edu/web/pages/ https://www.statista.com/ vol25.4 4 iassist quarterly winter 2001 iassist quarterly winter 2001 5 introduction: the growth of open source software open source software (oss) has grown tremendously in scope and popularity over the last several years, and is now in widespread use. oss has a long history of supporting technology infrastructure – the fundamental tasks of managing host names and addresses, sending information across the internet, delivering web pages, and relaying electronic mail are all primarily based on oss, and have been for many years. some of the most well-known technology businesses, like amazon and yahoo, are based on oss, and other technology companies like ibm have made heavy investments (oʼreilly 1999; sandred 2001, chap. 11; lerner and tirole 2002, secs. 1 and 3). last year, the oss operating system linux was used on one third of all servers (making it the second most popular server operating system)), and its use is expected to continue to grow rapidly (broersma 2002). oss is not limited to basic infrastructure. there are tens of thousands of oss projects providing everything from games to statistical packages to digital photography editing. directories of oss projects, such as gnu (<http://www.gnu.org/>) and sourceforge (<http:// www.sourceforge.net/>) now list over 50,000 projects, and the numbers continue to grow. the growth of oss has gained the attention of research librarians (frumkin 2002) and created new opportunities for libraries. we might well ask, what distinguishes oss from commercial software? what are the advantages and disadvantages of oss software? out of the thousands of packages available, which are most useful in a library environment? in this essay, i first discuss the primary features of oss, and where these features particularly benefit libraries. i also provide capsule summaries of oss projects and resources that may be of particular interest to the library community. what is open source software? open source softwareʼs distinguishing feature is the broad rights it awards the consumer. usage of software, as intellectual property, is restricted by copyright law and patent law.2 both commercial software and oss award the consumer certain rights to use it. most commercial software licenses give the consumer only limited use-rights – such as the right for a single user to run the software on a single system for a limited period. in contrast, oss provides broad rights to use, modify, and distribute the software. although the exact details of the rights will vary by license, all oss provides a number of broad rights (see perens 1999 and the open source initiative definition at: <http://www.opensource.org/docs/ definition.php>). these rights fall into three broad categories: 1. rights to use without discrimination. unlike commercial software (and even some ʻacademically licensed ̓software), oss may be used for any purpose, by anyone, at any time. for example, the same oss used to run an academic website can also be used to run an e-commerce business. there are no annual license fees, restrictions on the numbers of users or systems, restrictions for non-commercial use, restriction to a particular country, expiration dates, or other artificial limits on use. 2. full rights to create derived works. oss not only permits one to use the software, but permits one to create new software from it. a. source code availability. the source code for the software is made on the same terms as the binaries used to run it. b. free modification and redistribution. consumers not only have the right to examine the source, but to freely modify and redistribute modified (or unmodified) copies. c. integrity of authorship. oss may require that previous authors be acknowledged, and that modifications be clearly labeled, and separately packaged and named from the original software when redistributed. this maintains the integrity of software. open source software for libraries: from greenstone to the virtual data center and beyond by micah altman* http://www.opensource.org/docs/definition.php http://www.opensource.org/docs/definition.php 6 iassist quarterly winter 2001 iassist quarterly winter 2001 7 3. no traps. modified copies of oss must be redistributable under the same license as the original. the license cannot be restricted to a single product, and it must not restrict the distribution of other independently licensed software. (for example, the license must not insist that only oss software appears on a distribution cd-rom.) all major open source licenses honor these rights, including the gnu general public license (gpl), apple public source license, w3c license, mozilla license, and bsd license. although oss offers broad rights, users of both oss and commercial software user may still be under some restrictions: § government restrictions (such as the u.s. export controls on encryption software) may prevent the distribution of software, although the license allows it. § software (both commercial and non-commercial) that runs afoul of commercial patents may be restricted by the patent holder -regardless of the license between the consumer and the author. in addition, some open source licenses (most famously the gpl) prohibit the merging of open source and commercial softwarethis limits the ability of the consumer to intermingle open source and commercial source code in the same piece of software, but does not otherwise limit the combined distribution and use of commercial and open source software. still, in practice, even under the most restrictive of current oss licenses, oss products can still be sold (as long as the source is also made available for free), and one is free to sell documentation, support, installation and other services for oss. advantages and disadvantages of oss general advantages and disadvantages many librarians are now considering oss because of its low purchase costs. unlike commercial software, there are no initial purchase fees, licensing fees, or upgrade fees. furthermore, oss is generally not tied to proprietary hardware, so the hardware costs associated with oss tend to be lower. other direct costs for oss are often lower than those for comparable commercial software. since the original supplier of the software has no monopoly on the information, the market for support and maintenance of oss is more competitive than that for commercial software with comparable user bases. thus support and maintenance costs are often lower (kenwood 2001). many advocates3 of oss development argue that it leads to faster software development and more reliable software. raymond (1999, 39) argues that successful oss projects update their software quickly and frequently, paying close attention to users ̓bug reports and responding rapidly. moreover, bugs are fixed quickly because of the exposure oss provides; that is, “given enough eyeballs, all bugs are shallow” (p. 41). the open source initiative (<http: //www.opensource.org/>) states this particularly clearly: “when programmers can read, redistribute, and modify the source code for a piece of software, the software evolves. people improve it, people adapt it, people fix bugs. and this can happen at a speed that, if one is used to the slow pace of conventional software development, seems astonishing.” a recent mitre study offers some support for these practical claims, finding that oss often has advantages in reliability, frequency of bug fixes, extensibility and support (kenwood 2001). oss comes with a number of risks as well. first, when considering any software for long-term use, one must consider longevity, and oss is no exception. oss projects may fragment into incompatible versions or stagnate (particularly after losing a lead developer). one should consider the size of the user base and the number and activity level of the developers before incorporating software into long term plans. it is important to note, however, that commercial software suffers similar risks, and that the relative longevity of commercial versus oss software is an open question (lerner and tirole 2002). oss offers the opportunity for users to continue to contribute to a project even after the original developers leave. in contrast, development of commercial software (especially specialized software) is frequently abandoned completely when a business goes out of business or is acquired, or even when the business creates a new product. second, oss is often not as user-friendly as commercial counterparts. many open source projects (with notable exceptions, such as gnome, greenstone, and the virtual data center) are not cognizant of usability (hovater 2002), and the incentives for producing oss, while emphasizing utility, may de-emphasize user-interface development (lerner and tirole 2002). library-specific advantages of oss in addition to these general advantages, there are a number of reasons that libraries, in particular, may prefer to use oss over commercial software: preservation, privacy and auditing, community resources, and open standards. as librarians, we are sometimes the stewards of unique collections. the preservation of digital objects is currently intimately tied to software that presents those objects. complete preservation of complex digital objects, especially, is likely to require preservation of the software needed to use those objects (granger 2000). since commercial software is usually distributed only as a binary that will run only on a single hardware platform (and often only under a single version of a particular operating system), 6 iassist quarterly winter 2001 iassist quarterly winter 2001 7 commercial software is very difficult to preserve over the long run without developing hardware emulation (and possibly operating system ʻemulation ̓as well). oss, in contrast, can often be recompiled, or at least ported, to new hardware and operating systems. librarians have a strong tradition of defending the privacy of users. increasing numbers of commercial packages, by both major and minor vendors, quietly collect information on the systems and habits of people who use them. whether as part of ʻadware, ̓ ʻspyware, ̓ʻlive ̓updates, or digital rights management, these applications send usage information back to the vendor (millman 2001). furthermore, similar features are increasingly added by vendors in order to ʻtransparently ̓(and in many cases, silently) install new software on the clientʼs system, or to disable software that is the subject of payment or intellectual property disputes. although this behavior can be limited with placement of network firewalls, it can be challenging to audit commercial software for this type of behavior. in contrast, oss is easily audited. moreover, since the source code is available, it is relatively easy for a community of users to check for and disable ʻspyware ̓and other remote reporting and installation functions. scholarly standards and exchange both the library and the academic community have a history of sharing information and of using open standards. for example, both citations and cataloging, two of the mainstays of librarianship and academics, are based upon standards that are fundamentally open. the use of oss ensures that all software standards, both explicit and implicit, will continue to be open to inspection, and allows others to build upon previous work done in the community. as lessig (1999) has made clear, software infrastructure has important implications for the types of community values that software can support and encourage. open academies and open libraries demand open source infrastructures. figure 1: example results screen from greenstone digital library an overview of stand-alone solutions for libraries there are several packages that aim to offer stand-alone catalogs or complete digital library solutions: greenstone, koha, rib, sitesearch (which is, unfortunately, not really open source) and the virtual data center. in this section, i draw thumbnail sketches of each. in the next section, i discuss other software tools and resources that are likely to be useful to librarians who wish to construct their own library applications or digital libraries. greenstone greenstone is a package for creating, managing and distributing collections of documents. it runs on linux, unix and windows platforms. collections created through greenstone can be used on-line or distributed on cd-rom. among the features it provides are multi-lingual interface support, full-text and fielded searching, browsable indexes, customized formatting, metadata extraction and a z39.50 client. greenstone was developed as part of the new zealand digital library project, run by the department of computer 8 iassist quarterly winter 2001 iassist quarterly winter 2001 9 figure 3: example navigation, results and management screens from rib figure 2: example splash screen from koha catalogue science, university of waikato, new zealand. the software is in its first major production release, and is actively maintained and updated (through sourceforge). it is available from < http://www .greenstone.org >. koha the koha system is a full catalogue, opac, management, and acquisitions package. it does not, however, support document distribution and indexing. it runs on linux and is accessed primarily through a webbased interface. among the features it provides are simple and fielded searches, reading lists, acquisitions management (including budgets and pricing information), circulation management, and patron management. koha was made in new zealand by the horowhenua library trust and katipo communications ltd. the software is in its first major production release, and is actively maintained and updated (through sourceforge). it is available from <http: //www.koha.org/>. rib repository in a box (rib) is a software package for creating web-browsable metadata collections. it runs on linux, unix and windows. collections created through rib are accessed through the web, and can be configured to interoperate with other similar remote rib collections. among the features rib provides are searching, browsable indexes, repository federation, and a user-friendly java-based management gui. http://www.library.org.nz/ http://www.library.org.nz/ http://www.katipo.co.nz http://www.katipo.co.nz 8 iassist quarterly winter 2001 iassist quarterly winter 2001 9 rib was created by the rib development team at the university of tennessee under direction by the national hpcc software exchange (nhse). the software is mature (in its second full production release) and seems to be actively maintained, although updates are infrequent. it is used in a number of nasa, doe, dod and nsf research centers. it is available from <http://www.nhse.org/rib/>. sitesearch sitesearch is an enhanced opac and distributed catalog search system. it runs on both unix and windows but is not documented to run on linux. among the features it provides are cross-catalog searching over world wide web and z39.50 sources, z39.50 client and server, interoperability hooks for interlibrary loan and document delivery services, relatively advanced search history and result set handling, and marc support. sitesearch was developed by oclc. the software is mature, and has been actively maintained, although the maintenance model is being changed. it is available from <http://www.sitesearch.oclc.org/>. sitesearch offers many of the benefits of oss within an academic library environment, but, unfortunately, sitesearch is not fully open source. the license under which it is distributed by oclc prohibits commercial use, and may limit the availability of third-party support. this license is incompatible with many popular oss licenses, which would hinder integration of sitesearch with other oss solutions. virtual data center the virtual data center (vdc) software is a comprehensive, open-source digital library system. the vdc software provides a complete system for the management and dissemination of federated collections of quantitative data. it runs on linux. collections created through vdc are accessed through the web, and can be distributed across multiple servers, or virtually include selected parts of other data archives. among the features it provides are simple and fielded searching (at all levels of granularity – collection, study and data), data and documentation delivery, data extraction (variable and row selection), data format conversion (data documentation initiative, sas, spss, stata, splus, csv), on-line data analysis (descriptive statistics, exploratory data analysis graphs, crosstabs), archival format and filesystem-independent storage, open archives initiative service, z39.50 service, persistent naming, distributed operation, distributed virtual collections, metadata harvesting, federated authentication and authorization, and on-line gui management tools.(see figure 5) vdc was developed by the harvard-mit data center and harvard university library as part of the digital libraries initiative, sponsored by the national science foundation and other agencies. the software is in beta release, and is actively maintained and updated (through sourceforge). it is available from < http://thedata.org >. other resources for oss in libraries the preceding projects provide stand-alone oss catalogs or complete digital libraries. in addition to these, there are a host of open source resources that are useful to libraries, including website toolkits, indexing engines, databases, and clients, servers and software libraries for specialized technologies such as z39.50, usmarc, and ariel. oss meta-sites. these sites provide directories of opensource projects. § the “open source software for libraries” (< http://www.oss4lib.org/ >) uniquely specializes in library applications, and is particularly useful. § sourceforge (< http://sourceforge.net/ >), freshmeat (< http://freshmeat.net >), and the free software foundation (fsf; < http://gnu.org/ >) provide huge catalogs of open source projects. fsf is the oldest oss site, and hosts thousands of projects. sourceforge hosts nearly 50,000 oss projects. § “the impoverished social scientistʼs guide to free statistical software and resources” (< http: //data.fas.harvard.edu/micah_altman/socsci.shtm l>) is a collection of pointers to oss tools for data analysis and manipulation collected by the author. § the “freegis site” (< http://www.freegis.org/ >) is a collection of pointers to free gis tools and toolkits. § a large collection of searching, harvesting and indexing tools is cataloged at < http://www.searchtools.com/ tools/tools-opensource.html >. specific tools and toolkits. a number of tools and toolkits may be of specific interest to libraries that are building their own tools, or who wish to supplement current tools. full-text searching of material in web-based catalogs can be provided by using indexers such as swish-e or ht://dig. more structured catalogs can be built using xml and xml databases such as dbxml, xindice, or cheshire, or using sql databases such as postgressql or mysql. information portals can easily be built with toolkits such as slash and jetspeed.4 finding aids, pathfinders, and other dynamic, organized web applications can be built with open source application servers such as zope and gist. conclusions for the last several years, open source software has dominated the infrastructure of internet and web services. oss continues to grow in this and other areas, and there are now http://data.fas.harvard.edu/micah_altman/socsci.shtml http://data.fas.harvard.edu/micah_altman/socsci.shtml http://data.fas.harvard.edu/micah_altman/socsci.shtml 10 iassist quarterly winter 2001 iassist quarterly winter 2001 11 f ig ur e 5: e xa m pl e n av ig at io n an d r es ul ts s cr ee n fr om v d c 10 iassist quarterly winter 2001 iassist quarterly winter 2001 11 over 50,000 open source software applications available for instant download. among these are a number of high-quality packages that provide stand-alone digital library and opac functionality, as well as a host of other applications and toolkits that would be of great use in the development or enhancement of library services. the most popular open source software projects produce software that is quite often more stable, secure, auditable, and extensible than commercial alternatives. using oss also makes the preservation of digital objects easier and less risky. moreover, using oss guarantees that the standards and protocols used in the library will always be open to examination, and helps the library community to build upon previous successes. bibliography [1] broersma, matthew, 2002, “will linux survive the dot-com crash?,” zdnet (uk), january 2, 2002, < http://www.zdnet.com/filters/printerfriendly/ 0,6061,2835454-92,00.html>. [2] frumkin, jeremy, ed., 2002, “special issue: open source software,” information technology and libraries 21(1) <http://www.lita.org/ital/ital2101.html>. [3] granger, stewart, 2000, “emulation as a digital preservation strategy,” d-lib magazine 6(10). < http: //www.dlib.org/dlib/october00/granger/10granger.html> [4] hovater, j., d. kiskis, m. krot, i. holland, and m. altman, 2002, “usability testing of the virtual data center,” in proceedings of the joint conference on digital libraries ʼ02, acm press: new york. [5] kenwood, carolyn a., 2001, “a business case study of open source software,” report # m p 0 1 b 0 0 0 0 0 4 8, mitre corporation: bedford, ma. < http: //www.mitre.org/support/papers/tech_papers_01/kenwood_ software/> [6] lerner, josh, and jean tirole, 2002, “some simple economics of open source,” journal of industrial economics, 50 (2002) forthcoming [7] lessig, lawrence, 1999, code and other laws of cyberspace, basic books: new york. [8] millman, howard, 2001, “how to keep vendors from quietly violating your privacy,” new york times, january 18, late edition final, section g, page 9, column 1. [9] o’reilly, tim, “hardware, software and infoware,” in open sources, edited by chris dibona, sam ockman and mark stone, o’reilly and sons: sebastapol, caq [10] perens, bruce, 1999, “the open source definition”, in open sources, edited by chris dibona, sam ockman and mark stone, o’reilly and sons: sebastapol, ca. [11] raymond, eric, 1999, cathedral and the bazaar, o’reilly and sons: sebastapol, ca. [12] sandred, jan, 2001, managing open source projects, john wiley & sons: new york. [13] stallman, richard, “the gnu operating system and the free software movement,” in open sources, edited by chris dibona, sam ockman and mark stone, o’reilly and sons: sebastapol, ca. footnotes 1 this material is based upon work supported by the national science foundation under grant no. 9874747. 2 to a much lesser extent, some pieces of software may additionally be governed by trademark law. 3 some advocates, most notably richard stallman of the free software foundation, argue for oss on ethical grounds. for stallman, free software is a “stark moral choice” (stallman 1999, 55), and restrictions on the distribution of information in general, and software in particular are harmful to society. (see <http://www.gnu.org/ philosophy/why-free.html>.) 4 the scout toolkit, which was released in beta as this article was going to press, also looks promising because of its attention to metadata. (<http://scout.cs.wisc.edu/research/ spt/>) * paper presented at the iassist conference, june 2002, in storrs, ct, usa. micah altman, harvard university, micah_altman@harvard.edu. http://www.zdnet.com/filters/printerfriendly/0,6061,2835454-92,00.html http://www.zdnet.com/filters/printerfriendly/0,6061,2835454-92,00.html http://www.lita.org/ital/ital2101.html 1/12 blackwood, elizabeth (2021) outside the r1: equitable data management at the undergraduate level, iassist quarterly 45(2), pp. 1-12. doi: https://doi.org/10.29173/iq1011 outside the r1: equitable data management at the undergraduate level elizabeth blackwood1 abstract universities within the california state university system are given the mandate to teach the students of the state, as is the case with many regional, public universities. this mandate places teaching first; however, research and scholarship are still required activities for faculty seeking to achieve retention, tenure, and promotion, as well as important skills for students to practice. data management instruction for both faculty and undergraduates is often omitted at these institutions, which fall outside of the r1 designation, the carnegie classification for “very high research activity.” this happens for a variety of reasons, including personnel and resource limitations. such limitations disproportionately burden students from underrepresented populations, who are more heavily represented at these institutions. these students have pathways to graduate school and the digital economy, like their counterparts at r1s; thus, they are also in need of research data management skills. this paper describes and provides a scalable, low-resource model for data management instruction from the university library and integrated into a department’s capstone or final project curriculum. in the case study, students and their instructors participated in workshops and submitted data management plans as a requirement of their final project. the case study will analyze the results of the project and focus on the broader implications of integrating research data management into undergraduate curriculum at public, regional universities. by working with faculty to integrate data management practices into their curricula, librarians reach both students and faculty members with best practices for research data management. this work also contributes to a more equitable and sustainable research landscape. keywords research data management, academic libraries, undergraduates, equity, hispanic-serving institutions introduction the past several years have brought about a much needed diversity, equity, and inclusion (dei) reckoning in many libraries (gibson et al, 2020, pp 75). 2020 was no different. many institutions face a call to analyze, audit, and improve their dei initiatives or consciously create them if they do not yet exist. as academic libraries attempt to improve equity in all service areas, research data management (rdm) should not be overlooked. as with all areas of society, this service area is not free of structural inequality and equity concerns. this paper will document a course-integrated data management workshop for undergraduate students, but the primary focus will be to highlight equity challenges in data management instruction for universities that are not research-focused and do not meet the carnegie classification of “research 1” or “research 2”, as well as methods to combat those challenges. rdm has had a home in academic libraries for more than a decade, making a regular and significant appearance in the literature in 2008 (delserone, 2008, pp. 203-206; henty, 2008, pp. 2-3). by 2014, it became commonplace to associate this service with academic libraries in the united states. this date was not arbitrary as the u.s. federal government stated its intentions for accessing publicly funded https://doi.org/10.29173/iq1011 2/12 blackwood, elizabeth (2021) outside the r1: equitable data management at the undergraduate level, iassist quarterly 45(2), pp. 1-12. doi: https://doi.org/10.29173/iq1011 data just a year prior (holdren, 2013, p. 3). as bethany latham astutely highlighted in her column for the journal of academic librarianship, money, primarily that associated with federal grants, was the primary driver for rdm’s quick rise to fame (latham, 2017, p. 1). over the past decade, scholars have contributed to significant literature on developing successful rdm programs, providing the best practices for these services to users, and standardizing these best practices. this literature has provided the building blocks for smaller libraries to implement successful programs. however, the literature has been nearly silent on the equity challenges that face rdm stakeholders at institutions that fall outside of the carnegie classification of ‘very high research activity (r1)’ or ‘high research activity (r2).’ 2 according to this classification, there were one-hundred-and-thirty-one r1 institutions and onehundred-and-thirty-five r2 institutions in the united states in 2018 (carnegie classification of higher education, 2017,).3 nearly four-thousand-three-hundred postsecondary institutions granted degrees in the 2018 academic year, most of which fall outside of these research classifications (moody, 2019). r1 and r2 institutions represent only six percent of all degree granting institutions. thus, how do libraries serve the rdm needs of the faculty, students, and researchers at the other ninety-four percent of universities, especially those that serve marginalized communities? this is not to claim that all r1 and r2 institutions provide adequate rdm services to their faculty and students. much research states that this not is the case (radecki and spring, 2020, pp. 4). this is also not to claim that r1 and r2 universities do not serve diverse populations. in 2020, ten r1 universities were designated as hispanic serving institutions (hsis), like the campus described in this paper (excelencia in education, 2020, pp. 1-15). however, there is no research documenting rdm at institutions that fall outside of these classifications. additionally, it is well documented that regional institutions serve more diverse populations than their r1 and r2 counterparts (u.s. news & world report, 2019). this paper documents a case study of the provision of rdm services at an institution that falls outside of the r1 and r2 classifications and outside the current landscape of study for rdm. below, the author will describe strategies for 1) creating partnerships with faculty members who face structural challenges in providing rdm instruction for students, 2) recognizing and adjusting for the technical deficits facing students, and 3) implementing practical solutions. these strategies can be implemented to serve all universities, including r1 and r2. background of the case study the subject of this paper is best understood with a clear grasp of the system and institution that it took place within. the california state university (csu) system is the largest public university system in the united states. the csu has an enrollment of nearly five-hundred thousand students across twenty-three campuses. the system is known for teaching practices, which are often prioritized over research activities in contrast to the heavily funded university of california system. none of the twenty-three campuses fall into the r1 or r2 classification, with most serving regional populations. this study took place at csu channel islands (csuci), the newest university within the csu system, established in 2002. csuci is located in camarillo, ca, between santa barbara and los angeles, and serves the tri-county area (ventura, santa barbara, and los angeles counties). as of fall 2020, csuci had a total enrollment of just over seven-thousand students. more than sixty percent are first generation college students or students who are the first in their families to attend a four-year academic institution. fifty-five percent of the total enrollment are from historically underrepresented https://doi.org/10.29173/iq1011 3/12 blackwood, elizabeth (2021) outside the r1: equitable data management at the undergraduate level, iassist quarterly 45(2), pp. 1-12. doi: https://doi.org/10.29173/iq1011 groups and csuci is recognized as an hispanic serving institution (hsi) and receives federal funding to support that designation. despite its mandate for teaching, research occurs across the csu system. at csuci, this research is almost universally supported by undergraduate students as there is little focus or funding for graduate programs and graduate research assistance. utilizing undergraduate students as research assistants creates significant demands on faculty and results in general needs for undergraduate rdm training and instruction, both to support faculty research and as a means to prepare students for careers in the digital economy or graduate school. despite the clear need and plethora of literature on the topic, there is no guidebook or best practice for doing this work with a community that is majority marginalized and faces severely limited resources. this work requires adjustments, sensitivity, and a substantial time investment (for both librarians and discipline faculty) in order to make an effective impact that meets undergraduate students on their level. brief review of the literature the following review of the literature focused on two specific areas: a survey of the rdm work that academic libraries are performing with undergraduates and a survey of where the literature stands on the equitable provision of this work. the research in this review was limited specifically to american institutions due to this paper's focus on the carnegie classification as a tool for measuring resources. additionally, the author chose to focus on american institutions due to the specific diversity, equity, and inclusion (dei) conversations that are currently on-going in institutions of higher education. current work with undergraduates as mentioned in the previous section, faculty at csuci demonstrated the need for undergraduate education around rdm. this experience was validated across much of the literature; though such work has been practiced in a variety of ways. some practitioners found it challenging to reach undergraduates through course-integrated instruction and sought extra-curricular approaches to rdm education (carlson et al, 2015, pp. 16-17; clement et al, 2017, pp. 8-9; cook et al, 2020, p. 3). this approach often required funding for food to attract students to events and workshops or for other programming. others found pathways for course integration through both discipline faculty and library-specific courses (ball and medeiros, 2012, p. 182; mooney et al, 2014, pp. 376-379; reisner, vaughan, and shorish, 2014, p. 1943; zhang and gall, 2017, p. 2). course integration models for rdm instruction rely heavily on relationships with discipline faculty, but prove to be more equitable for those practitioners working on “shoestring” budgets, allowing practitioners to “take advantage of existing resources and tools,” and capitalize on need (henderson et al., 2014, p. 17). the workshop curriculum described later in this paper relied heavily on the work of other practitioners, especially those who have developed frameworks for teaching rdm (sapp nelson, 2017, pp. 4-7; piorun, 2012, pp. 48-49). specifically, megan sapp nelson’s work to develop an rdm competency matrix was extremely helpful in creating a tangible assignment attached to the workshop curriculum. her use of matrix domains for teaching “personal information management,” “team data management,” and “research enterprise management,” translated well to department faculty and capstone students. while outside the context of academic libraries, richard ball and norm medeiros’ work on instruction related to the values and skills of documentation and readmes also offered a foundation for the graded assignment (ball and medeiros, 2012, p. 182). additionally, most authors https://doi.org/10.29173/iq1011 4/12 blackwood, elizabeth (2021) outside the r1: equitable data management at the undergraduate level, iassist quarterly 45(2), pp. 1-12. doi: https://doi.org/10.29173/iq1011 relied in some form on the dataone modules and curriculum for teaching and tools, demonstrating this resource’s sustained value to the rdm community (shorrish, 2015, p. 10; cook et al, 2020, p. 3; mooney et al, 2014, p. 375; sapp nelson, 2017, p. 7; kafel, creamer, and martin, 2014, p. 61; reisner, vaughan, and shorish, 2014, p. 1945). the later described workshop activities were developed specifically for this capstone course, but relied heavily on examples from the literature. the data curation activity was modeled after the undergraduate rdm curriculum developed for the chemistry department at james madison university, with it’s instructions and description left “intentionally minimal” (reisner, vaughan, and shorish, 2014, p. 1944). additionally, the data management plan activity and corresponding assignment relied on the dmptool, a resource designed for institutions with limited resources (henderson et al, 2014, pp. 14-16). equitable provision of rdm while performing this review of the literature, the author also reviewed the types of institutions that dominate the conversation around rdm in academic libraries. of the articles surveyed and mentioned in this paper that specifically refer to undergraduate work, the research came from sixteen r1 and r2 institutions and eight private liberal arts colleges (most of which were supported by consortia that included r1s or r2s). the only outlier in this survey was the work of yasmeen shorish and her colleagues at james madison university, a public, masters-comprehensive university that falls outside of the carnegie designations. her work spoke specifically to the challenges that many academic libraries face in providing rdm services at smaller institutions (shorrish, 2012, pp. 265-270). however, even this outlier does not adequately represent the situation of many regional public institutions (like those in the csu system), which have limited master’s students and serve primarily underrepresented populations. additionally, in a dedicated search of all the literature reviewed for this article, not one author discussed rdm in terms of equity or in the provision of this work for students from marginalized backgrounds. research data management course integration while faculty across csuci’s campus voiced calls for rdm training through a variety of localities (especially administrative offices supporting grants), it was up to the librarians to provide outreach around rdm to ensure that relevant faculty were aware of the service. this outreach primarily consisted of unsolicited emails directed to faculty members and requests for time on the agendas of departmental faculty meetings. it was a fifteen-minute presentation at such a faculty meeting that inspired the described partnership with the department of environmental science and resource management (esrm) at csuci. what began as a departmental rdm training to refresh faculty members with best practices ultimately became a partnership between the library and the department’s required capstone project curriculum. the esrm capstone project is a two-semester, research-based course where students work in groups to answer a scientific research question through data collection and analysis. librarians were asked to partner with discipline faculty in the development of an assignment that would support rdm instruction. in the assignment, students were asked to submit a graded data management plan during the first semester of the course following a librarian-led rdm workshop. this plan would become a core component of the final submission in the second semester of the course. the integration of the graded assignment greatly increased student engagement and was a core component of the workshop. https://doi.org/10.29173/iq1011 5/12 blackwood, elizabeth (2021) outside the r1: equitable data management at the undergraduate level, iassist quarterly 45(2), pp. 1-12. doi: https://doi.org/10.29173/iq1011 workshop description a librarian delivered the workshop to three separate sections of the capstone course, reaching a total of 71 students and three faculty members. the workshop filled a fast-paced hour, but could have utilized more time. the workshop took place in a computer lab classroom to ensure that every student had access to a computer and would be able to fully participate in the activities. students participated in the workshop as pre-existing groups, already assigned for their capstone projects. prior to the workshop, the instructing librarian hosted the documentation and data on a stable library webpage to ensure that students could access the material both during and after the class session. the outline and description of the workshop are summarized below: introduction to research data management for undergrads 1. what data should we manage? a. to begin the workshop, students were asked to define “data” and compile a list of the types of data that they would be collecting in their capstone projects. these data types were collected and written on a dry erase board in the classroom. b. once the list was compiled on the board, students were asked to add any software or hardware needed to collect or process their data. 2. why should we manage this data? a. the instructing librarian then addressed the importance of research data management, focusing specifically on 1) reproducibility, 2) transparency, and 3) integrity. b. students were then presented with several examples of recent data management and preservation debacles in the news and in research (hern, 2020; vine et al, 2014). 3. data curation activity a. after this introduction students were asked to perform a data curation activity to demonstrate the importance of these principles. students were asked to download a zipped folder that contained all of the files from a completed research study that originated in an open access repository. the folder contained more than 1800 unorganized files that did not follow any particular file name structure. the folder also contained multiple readme files. students were given “intentionally minimal” instructions and asked to explore the files and try to answer the following questions in groups: i. how many files did this study produce? ii. what is the topic of this study? iii. what type of data did the researchers collect? iv. how would you open these files? v. what information would need to replicate this study? b. the students responded very well to this activity, vocalizing the file types that they recognized and expressing their frustration at the complexity of the files. the instructing librarian walked around the room, provided assistance when asked, and often suggested students search for a readme within the files. c. after about ten minutes of exploration, most students were able to identify the rough subject of the study, a few data types, and several software tools that might be helpful https://doi.org/10.29173/iq1011 6/12 blackwood, elizabeth (2021) outside the r1: equitable data management at the undergraduate level, iassist quarterly 45(2), pp. 1-12. doi: https://doi.org/10.29173/iq1011 for opening files. the instructing librarian then revealed the title and documentation for the unorganized folder by sharing the link to the open access repository record. d. students were then asked to discuss their thoughts on the activity based on the earlier content. 4. data management plans a. as a response to the somewhat frustrating data curation activity, the instructing librarian then introduced data management plans (dmps) by 1) defining dmps, 2) demonstrating their purpose, especially in the context of the previous activity and previously mentioned examples, and 3) describing their components. 5. data management plan activity a. building on the content, students were asked to set up accounts using the dmptool (dmptool.org). the instructing librarian offered assistance as students signed up. b. once the students had access to their accounts, each group was assigned a section of the plan to familiarize themselves with and then report out to the rest of the class. the sections included: data collection, documentation and metadata, ethics and legal compliance, storage and backup, selection and preservation, data sharing, and responsibilities and resources. each section provided guiding questions and resources for students to explore within the dmptool. c. students were given fifteen minutes to discuss this within their groups. during the sharing portion of the activity, the instructing librarian asked guiding questions that related to each section and provided instruction on the best practice for each area. d. this section of the workshop provided a majority of the research data management instruction, but was led by the students. e. the discipline instructor also introduced the graded data management plan assignment. 6. file name and digital stewardship best practices a. to demonstrate a necessary component of their dmp assignment, the instructing librarian also gave a brief overview of file naming best practices. this included examples of 1) creating unique file names, 2) creating consistent file naming structures, and 3) avoiding special characters and spaces within file names. 7. careers in data management a. lastly, the instructing librarian described several career paths for students who found themselves interested in the topic of data management and provided a list of job titles and education requirements for interested students to investigate. equity evaluation it was necessary to dedicate significant time and sensitivity to equity considerations in preparation of the workshop. having trained in research data management at only r1 institutions, the workshop designer needed to critically evaluate any assumptions made in the workshop, especially regarding students’ available resources and personal technology. as demonstrated through the aforementioned analysis, the literature and best practices for rdm often make many assumptions with regards to https://doi.org/10.29173/iq1011 7/12 blackwood, elizabeth (2021) outside the r1: equitable data management at the undergraduate level, iassist quarterly 45(2), pp. 1-12. doi: https://doi.org/10.29173/iq1011 these aspects, as the majority of this work is done at well-funded research universities. adjustments were needed to ensure that the workshop was equitable for both csuci students and faculty with limited resources. the financial burden of rdm was the primary equity issue facing students in the course. it touched many aspects of the workshop, but was most evident through students' lack of access to dedicated computers. across the three sections of the course, several students did not possess their own computers. these students relied on some combination of cell phones, tablets, loaner laptops from the library (which have significant download and permissions restriction), desktop computers located in the library reading room, desktop machines available within specific campus department labs, or shared computers between family members at home. this created a variety of challenges at several points throughout the workshop. these challenges are outlined below: 1. storage workshop designers and discipline faculty knew that it was out of the question to suggest that students spend additional funds on materials for rdm in their capstone course, specifically because the course was a graduation requirement, not an elective. in addition, the campus did not provide dedicated server space for student projects in the environmental science and resource management department. these factors, along with the lack of personal computers, forced most students to access and manage their data almost universally in low-cost or free cloud storage. nearly all of the student groups stored their data in campus-provided google drive accounts, either because they, themselves did not have consistent access to a computer or an external storage device or because even one member of the project group did not have consistent access to a computer or external storage. this made cloud storage the primary (and typically only) form of storage. most student groups had no feasible option for traditional backups. this was not a problem reserved only for the students. faculty members in the department faced similar storage issues. apart from business-level dropbox accounts provided by the university, faculty members had no department or campus servers, making it difficult for them to store their own research, let alone provide assistance or impart best practice to their students. workshop designers needed to respectfully mitigate these issues, while still providing instruction on the best practice. rather than briefly explaining the need for multiple backups, this workshop required more time, explicit instruction, and discussion about what a good backup procedure might look like in their situation. the instructing librarian specifically explained the importance of multiple backups and asked students to discuss how they would manage two to three backups within their group, either in different platforms of cloud backups, external storage, or on personal computers. although it was not ideal and it was an additional burden on the student groups, most groups were able to negotiate a workflow for data storage and backups. the students were also able to document the workflow in real time using the dmptool, immediately demonstrating the value of a data management plan. asking students to solve this problem, led them to conversations about how the students would communicate with one another and schedule backups and syncs with any non-networked storage options. https://doi.org/10.29173/iq1011 8/12 blackwood, elizabeth (2021) outside the r1: equitable data management at the undergraduate level, iassist quarterly 45(2), pp. 1-12. doi: https://doi.org/10.29173/iq1011 2. file naming file renaming presented another major disruption. students who created or collected data that uploaded directly to google drive or dropbox (often from an external tool like a drone or water probe) faced significant issues with file naming or renaming due to the proprietary nature and usability of the cloud interfaces. there is no easy or granular way to rename files in bulk in these systems. even though several plugins and extensions exist for batch file renaming within google drive, they are less reliable than their client-based counterparts. file naming best practices are typically a major component of basic data management instruction and thus the instructing librarian dedicated more workshop time to this topic so that students could plan for work arounds in their data management plans. to combat this challenge, the instructing librarian asked students to think through their data collection workflows and document how digital data files would be created. students using external tools recognized that they would need to create file naming schemas prior to the data collection process and set up their tools to produce files with names to match their schemas. some students found that they needed to locate and consult with user manuals for the tools that they would be using. this was a time-consuming component of the workshop. the instructing librarian offered additional options for later consultation for student groups who were not able to finish this within the time allotted or those who were not yet prepared for their data collection process. 3. time rdm practitioners are keenly aware that this work is time consuming at its best and painstaking at its worst. to combat this, researchers at well-funded institutions often rely on automation and syncing tools to make tasks less burdensome. researchers with access to these tools spend more time analyzing data and less time moving and processing it. this was not a luxury afforded to students in this course. it was important for workshop designers to be realistic about time costs with students as many have a variety of other demands that require their time. csuci is often referred to as a “commuter school,” with many students coming from the far reaching edges of the county for their courses, while also balancing child care, elder care, and jobs. only a small number of students live on campus with regular and consistent access to labs for processing and analyzing data. additionally, the campus library is not open on a twenty-four-hour schedule, as is the case at many larger institutions, limiting the hours for students to access library computers and software. it was necessary for the instructing librarian to make these time costs abundantly clear to students early in the process, so students could plan realistically and accordingly. for example, students were instructed to prepare for the time costs involved in moving large amounts of data, both from lab computers to cloud storage and back; to prepare for the amount of time they would need to reserve for machine and equipment access, if using loaned equipment; and to prepare for the amount of communication needed across the group when data syncing could not be assumed across all group members. assessment although formal data was collected during the first iteration of the workshop through pre and post survey, and interviews with both faculty and students, the author was unable to utilize the data for this case study due to time constraints and workload issues related to the campus institutional review board during the covid-19 pandemic. this leaves much opportunity to revisit the study and continue https://doi.org/10.29173/iq1011 9/12 blackwood, elizabeth (2021) outside the r1: equitable data management at the undergraduate level, iassist quarterly 45(2), pp. 1-12. doi: https://doi.org/10.29173/iq1011 formalized, quantitative assessment on future iterations of the workshop in the post-pandemic environment. without data collection, assessment of the workshop came from on-going conversations and collaboration with the discipline faculty. it was clear early in the discussion with discipline faculty that the workshop’s biggest weakness was the hour-long time constraint. this did not provide enough time for students to gain depth of knowledge on any aspect. discipline faculty agreed to test the next iteration of the workshop in a ninety-minute class period. apart from the time constraints, faculty course instructors were immediately pleased with the result of the workshop and saw improvement of the capstone assignments over the course of the year. they also requested that the instructing librarian visit the course again during the second semester to check-in with students on their data management plans, answer lingering questions prior to submission, and provide on-going mentorship. based on the first year’s success, the discipline faculty elected to continue the partnership and make both the workshop and assignment a permanent part of the course. in addition, the workshop proved to be easily scalable for remote instruction. during the spring semester of 2020 and the entirety of the 2020-2021 school year, all twenty-three campuses of the csu system remained closed due to the covid-19 pandemic. the workshop ran smoothly in this environment and also worked well as recorded asynchronous instruction. it should be noted that the instructing librarian saw more requests for separate group consultations for rdm during the pandemic, most likely caused by students feeling less inclined to speak during online instruction. on a broader level, the library also saw an increase in faculty requests for assistance and consultation in rdm (data management plans required for grants, specifically) from not only the environmental science and resource management department, but across the university. this was an unexpected, but welcome surprise. it spoke to the success of the course integration and highlighted potential areas for the library to build more relationships with discipline faculty and reach more students. conclusion for library workers, it should always be a priority to provide equitable access to library services. for research data management there is much work to be done. equitable service does not have to contend with best practice; however, it must contend with time allotted to provide this service. this time constraint exists in both the planning and execution of services. this case study provides a variety of valuable lessons for providing rdm services outside of the r1. the first and most important lesson was to establish a strong and communicative partnership with the department and faculty members requesting rdm services. they provide access to students and hold the most direct knowledge of their students’ personal equity challenges. this case study was only possible because discipline faculty welcomed and collaborated with the library. the relationship began as a cold email to a department chair and blossomed into a repeated course integration. the inclusion of a data management plan as a graded assignment was an important piece of the capstone project that pushed the workshop beyond a typical information literacy session. it also integrated rdm (and the library) into the rest of the students’ research and coursework. most importantly, it provided the students with valuable experience and skill building that could serve them well beyond their undergraduate careers. https://doi.org/10.29173/iq1011 10/12 blackwood, elizabeth (2021) outside the r1: equitable data management at the undergraduate level, iassist quarterly 45(2), pp. 1-12. doi: https://doi.org/10.29173/iq1011 the second lesson was that this level of course integration demanded that the librarian instructor dedicate significant thought to the equity challenges that come with providing rdm services. the author had to think about each step of the teaching workflow and ask herself if it was possible to accomplish without a dedicated personal machine or any financial strain. this process was challenging and time-consuming, but allowed for the space to consider the areas that practitioners can improve equity in this service area. while designing the workshop, the author found it important to remember why we provide this service to an undergraduate population. if rdm practitioners hope to diversify academia (and specifically academic libraries), they must expose more students at diverse institutions to this work. references ball, r., medeiros, n. (2012) teaching integrity in empirical research: a protocol for documenting data management and analysis. journal of economic education. 43(2), 182-189. doi: https://doi.org/10.1080/00220485.2012.659647 carlson, j., nelson, m.s., johnston, l.r., koshoffer, a. (2015) developing data literacy programs: working with faculty, graduate students and undergraduates. bulletin of the association for information science and technology; 41: 14-17. doi: https://doi.org/10.1002/bult.2015.1720410608 carnegie classifications of institutions of higher education. (2017) standard listings. https://carnegieclassifications.iu.edu/lookup/standard.php#standard_basic2005_list clement, r; blau, a; abbaspour, p.; gandour-rood, e. (2017) team-based data management instruction at small liberal arts colleges. ifla journals; 43(1). pp 105-118. doi: https://doi.org/10.1177%2f0340035216678239 cook c, magle t, shimon h, adamus t. (2020) dinner and data management: engaging undergraduates in research data management topics outside of the curriculum. journal of escience librarianship; 9(1): e1176. https://doi.org/10.7191/jeslib.2020.1176 delserone, l. (2008) at the watershed: preparing for research data management and stewardship at the university of minnesota libraries. library trends; 57(2). fall 2008. doi: https://doi.org/10.1353/lib.0.0032. retrieved: https://digitalcommons.unl.edu/libraryscience/245/ excelencia in education. hispanic serving institutions (hsis) 2019-20. https://www.edexcelencia.org/hispanic-serving-institutions-hsis-2019-2020 gibson, a.n., chancellor, r.l., cooke, n.a., dahlen, s.p., patin, b. and shorish, y.l. (2021). struggling to breathe: covid-19, protest and the lis response. equality, diversity and inclusion; vol. 40 no. 1, pp. 74-82. henderson, m; raboin, r.; shorish, y.; van tuyl, s. (2014) research data management on a shoestring budget. bulletin of the association for information science and technology; vol. 40(6). retrieved: http://scholarscompass.vcu.edu/libraries_pubs/23 https://doi.org/10.29173/iq1011 https://doi.org/10.1080/00220485.2012.659647 https://doi.org/10.1002/bult.2015.1720410608 https://carnegieclassifications.iu.edu/lookup/standard.php#standard_basic2005_list https://doi.org/10.1177%2f0340035216678239 https://doi.org/10.7191/jeslib.2020.1176 https://doi.org/10.1353/lib.0.0032 https://digitalcommons.unl.edu/libraryscience/245/ https://www.edexcelencia.org/hispanic-serving-institutions-hsis-2019-2020 http://scholarscompass.vcu.edu/libraries_pubs/23 11/12 blackwood, elizabeth (2021) outside the r1: equitable data management at the undergraduate level, iassist quarterly 45(2), pp. 1-12. doi: https://doi.org/10.29173/iq1011 henty, m. (2008) dreaming of data: the library’s role in supporting e-research and data management. australian library and information association biennial conference, alice springs. http://hdl.handle.net/1885/47617 hern, alex. (2020) covid: how excel may have caused loss of 16,000 test results in england. the guardian; oct. 6. https://www.theguardian.com/politics/2020/oct/05/how-excel-may-have-causedloss-of-16000-covid-tests-in-england holdren, j. (2013) increasing access to the results of federally funded scientific research. executive office of the president, office of science and technology policy. https://obamawhitehouse.archives.gov/sites/default/files/microsites/ostp/ostp_public_access_me mo_2013.pdf kafel d, creamer at, martin er. (2014) building the new england collaborative data management curriculum. journal of escience librarianship; 3(1): e1066. https://doi.org/10.7191/jeslib.2014.1066 latham, bethany. (2017) research data management: defining roles, prioritizing services, and enumerating challenges. the journal of academic librarianship; 43(3), pp 263-265. https://doi.org/10.1016/j.acalib.2017.04.004 moody, josh. (2019) a guide to the changing number of u.s. universities. us news & world report. https://www.usnews.com/education/best-colleges/articles/2019-02-15/how-many-universities-arein-the-us-and-why-that-number-is-changing mooney, h; collie, w. aaron; nicholson, shawn; sosulski, marya r. (2014) collaborative approaches to undergraduate research training: information literacy and data management. advances in social work; 15(2). https://doi.org/10.18060/15089 piorun m, kafel d, leger-hornby t, najafi s, martin e, colombo p, lapelle n. (2012) teaching research data management: an undergraduate/graduate curriculum. journal of escience librarianship; 1(1): e1003. https://doi.org/10.7191/jeslib.2012.1003 radecki, j; springer, r. (2020) research data services in us higher education: a web-based report. ithaka s+r. https://doi.org/10.18665/sr.314397 reisner, b., vaughan, k.t.l., shorish, y. (2014) making data management accessible in the undergraduate chemistry curriculum. journal of chemical education; 91 (11), pp 1943-1946. https://doi.org/10.1021/ed500099h sapp nelson, m. (2017) a pilot competency matrix for data management skills: a step toward the development of systematic data information literacy programs. journal of escience librarianship; 6(1): e1096. https://doi.org/10.7191/jeslib.2017.1096. shorish, yasmeen. (2012) data curation is for everyone! the case for master’s and baccalaureate institutional engagement with data curation. journal of web librarianship; vol. 6 (4). pp. 263-273, https://doi.org/10.1080/19322909.2012.729394 https://doi.org/10.29173/iq1011 http://hdl.handle.net/1885/47617 https://www.theguardian.com/politics/2020/oct/05/how-excel-may-have-caused-loss-of-16000-covid-tests-in-england https://www.theguardian.com/politics/2020/oct/05/how-excel-may-have-caused-loss-of-16000-covid-tests-in-england https://obamawhitehouse.archives.gov/sites/default/files/microsites/ostp/ostp_public_access_memo_2013.pdf https://obamawhitehouse.archives.gov/sites/default/files/microsites/ostp/ostp_public_access_memo_2013.pdf https://doi.org/10.7191/jeslib.2014.1066 https://doi-org.proxy.library.ucsb.edu:9443/10.1016/j.acalib.2017.04.004 https://www.usnews.com/education/best-colleges/articles/2019-02-15/how-many-universities-are-in-the-us-and-why-that-number-is-changing https://www.usnews.com/education/best-colleges/articles/2019-02-15/how-many-universities-are-in-the-us-and-why-that-number-is-changing https://doi.org/10.18060/15089 https://doi.org/10.7191/jeslib.2012.1003 https://doi.org/10.18665/sr.314397 https://doi.org/10.1021/ed500099h https://doi.org/10.7191/jeslib.2017.1096 https://doi.org/10.1080/19322909.2012.729394 12/12 blackwood, elizabeth (2021) outside the r1: equitable data management at the undergraduate level, iassist quarterly 45(2), pp. 1-12. doi: https://doi.org/10.29173/iq1011 shorish, yasmeen. (2015) data information literacy and undergraduates: a critical competency. college and undergraduate libraries; vol. 22 (1), pp. 97-106, doi: https://doi.org/10.1080/10691316.2015.1001246 u.s. news & world report. campus ethic diversity. (2019) https://www.usnews.com/bestcolleges/rankings/national-universities/campus-ethnic-diversity vines, t., albert, a., andrew, r., débarre, f., bock, d., franklin, m., gilbert, k., moore, j.s., renaut, s., rennison, d. (2014) the availability of research data declines rapidly with article age. current biology; 24(1) pp 94-97, https://doi.org/10.1016/j.cub.2013.11.014 zhang, q; gall, daniel. (2017) developing good data management habits early in academic life: data literacy education for undergraduate students. poster presented at the research data access and preservation summit, april 19, 2017. seattle, wa. retrieved from: https://ir.uiowa.edu/lib_pubs/209 endnotes 1 elizabeth blackwood is digital curation and scholarship librarian at california state university channel islands and can be reached by email at elizabeth.blackwood@csuci.edu. 2 carnegie classifications and definitions are available at https://carnegieclassifications.iu.edu/definitions.php 3 carnegie classifications include any institution that conferred at least one degree during the 20162017 academic year and were reported to the national center for education statistics ipeds. this data is intended to be a snapshot of a given time frame and is not updated yearly. https://doi.org/10.29173/iq1011 https://doi.org/10.1080/10691316.2015.1001246 https://www.usnews.com/best-colleges/rankings/national-universities/campus-ethnic-diversity https://www.usnews.com/best-colleges/rankings/national-universities/campus-ethnic-diversity https://doi.org/10.1016/j.cub.2013.11.014 https://ir.uiowa.edu/lib_pubs/209 mailto:elizabeth.blackwood@csuci.edu https://carnegieclassifications.iu.edu/definitions.php 1/7 tunmibi, sunday & olatokun, wole (2023) security and preservation of election data in nigeria in the fourth industrial revolution 47(34), pp. 1-7. doi https://doi.org/10.29173/iq1054 the creative commons-attribution-noncommercial license 4.0 international applies to all works published by iassist quarterly. authors will retain copyright of the work and full publishing rights. security and preservation of election data in nigeria in the fourth industrial revolution sunday tunmibi & wole olatokun1 abstract a fraud-free and credible election is a necessary ingredient to the growth of democracy. election malpractices and violence, from 1959 till date, have offered major challenges to the nigerian political system. to achieve a sustainable democracy in nigeria, it is important to build public trust by ensuring the security and preservation of electoral data. the world has gradually moved into the fourth industrial revolution (4ir), an era where artificial intelligence, big data, internet of things, robotics, blockchain, cloud computing and 3-d printing technologies dictate the pace of activities in all walks of life. this paper suggests specific 4ir technologies solutions to electoral data security and preservation challenges. it also suggests that the nigerian government announce policies to serve as catalysts for the independent national electoral commission and stakeholders to harness these developments to ensure that electoral processes benefit from these technologies. keywords election, fourth industrial revolution, data security, data preservation introduction it is generally believed that a free and fair election is crucial to the sustenance of democracy. however, the potential for election malpractices overshadow electoral process in nigeria. electoral malpractices or fraud are committed with an intention to influence an election in favour of a candidate(s) by means such as illegal voting, bribery, cheating and undue influence, intimidation and other acts of coercion exerted on voters, falsification of results, fraudulent announcement of a defeated candidate as winner with or without altering the recorded results (ogbeidi, 2010). for the independent national electoral commission (inec) to conduct credible elections, it needs to adopt relevant technologies to secure and preserve electoral data. the world is gradually moving into the fourth industrial revolution (4ir). the 4ir represents a developing environment in which disruptive digital technologies such as the internet of things (iot), cyber-physical systems (cps), block chain, artificial intelligence (ai), cloud computing, big data analytic, robotics, self-monitoring analysis and reporting technology (smart) technologies, and 3-d printing technologies are changing our approach to work and lifestyle (xu, david and kim, 2018). the 4ir evolved from the previous three industrial revolutions. the first industrial revolution, which began in the 18th century, enabled mechanized water and steampowered production, rather than purely human and animal power. the second industrial revolution, https://doi.org/ https://creativecommons.org/licenses/by-nc/4.0/ 2/7 tunmibi, sunday & olatokun, wole (2023) security and preservation of election data in nigeria in the fourth industrial revolution 47(34), pp. 1-7. doi https://doi.org/10.29173/iq1054 between the late 19th century and 20th century, introduced gas, oil and electric power, as well as more advanced communication technology (telephone and telegraph), for mass production of goods and automation of manufacturing process. the third industrial revolution, which began in the middle of the 20th century, was characterized with the advent of electronics, telecommunications, information technology and programming for automation of production process. the fourth industrial revolution of the present era is a digital revolution, where activities are carried out and controlled digitally (david, nwulu, aigbavboa and adepoju, 2022). there is no empirical evidence on the extent of applicability of the 4ir technologies to securing and preserving electoral data in nigeria. however, there have been various attempts by scholars to share insights into electoral malpractices in nigeria (agbaje and adejumobi, 2006; animashaun, 2010; oromareghake, 2013; ojukwu, mazi and maduekwe, 2019). agbaje and adejumobi (2006) noted that electorates have neither voice nor power, and their mandate is freely stolen by the political barons in nigeria. according to animashaun (2010), elections in nigeria have been marred by malpractices and the conduct of credible elections has remained an albatross. oromareghake (2013) reported that postcolonial elections in nigeria have been impaired by acrimony and rigging by desperate political office contenders. similarly, ojukwu, mazi and maduekwe (2019) noted that the 2019 general election was marred by vote manipulations and has credibility deficit. scholars have also conducted studies on the possibility of the application of election forensics to detect irregular patterns and fraud in nigeria electoral data (beber and scacco, 2008; tunmibi and olatokun 2020; tunmibi and olatokun 2021). brief history of elections in nigeria election fraud, from 1959 till date, has challenged the nigerian political system. according to edoh (2004), incidents of violence, and stuffing of ballot boxes as well as obstructions and intimidation of opponents were reported during the nigeria’s 1959 parliamentary elections. awopeju (2011) noted that elections for the western house of assembly in 1965 ended in violence as a result of widespread rigging. due to widespread election rigging and violence, the first military coup took place in nigeria on january 15, 1966. according to oromareghake (2013), election rigging was also reported during the elections organised by the military in 1979. the observed rigging during the election brought about the phrase “stolen presidency”, which has since become part of nigeria’s political vocabulary. the faulty election ushered in a civilian administration, governed by the national party of nigeria. election rigging was also reported in the 1983 elections. animashaun (2010) noted that there was mayhem in the two southwest states of oyo and ondo as a result of the massive manipulation of votes in favor of the ruling national party of nigeria. violence erupted due to the perceived manipulation of the governorship polls in these two states. in addition to the heavy human and material losses suffered by political opponents, the headquarters of the electoral body, federal electoral commission (fedeco), in oyo and ondo states were set on fire. on the 31st of december 1983, the military intervened once more and took over the government. it was not until may 1999 that democracy was restored. according to osinakachukwu and jawan (2011), the polity had been so damaged that people no longer showed interest in politics due to the prolonged reign of military dictatorship. the lackadaisical attitude shown towards the 1999 elections by nigerians gave the military junta a free hand to manipulate the elections to give power to their prefered candidate, obasanjo (osinakachukwu and https://doi.org/ 3/7 tunmibi, sunday & olatokun, wole (2023) security and preservation of election data in nigeria in the fourth industrial revolution 47(34), pp. 1-7. doi https://doi.org/10.29173/iq1054 jawan, 2011). the 2003 elections also failed to meet basic international standards. agbaje and adejumobi (2006) noted that the 1999 and 2003 elections, like virtually all the other preceding elections in nigeria’s post-colonial history, were classic cases of electoral fraud. the 2007 elections were described as the worst in nigeria’s history ranking among the worst conducted anywhere in the world in recent times (onebamhoi, 2011). the elections were characterized by late arrival of electoral materials in the various polling units, inadequate polling materials, voters’ registration problems, no secrecy of the ballot, ballot paper problems, snatching of ballot boxes and destruction of ballot materials, violence, use of security agencies to intimidate voters and rig elections, no voting in some polling centres, use of government officials to commit electoral fraud, and omissions of some parties’ logo and candidate names on the ballot paper to disenfranchise their opponents’ supporters (kia, 2013). although, the 2011 presidential election was commended by observers as one of the most successful in nigeria’s political history, cases of stuffing of ballot boxes, under age voting and outright falsification of election results were also reported in some states (yusuf and zaheruddin, 2015). similar to the previous elections, the 2015 presidential elections were impaired by vote buying, bribery, violation of electoral rules and other irregularities, while the phenomenon of money politics reached its zenith in nigerian politics (sule, 2019). the european union election observation mission (2015) also reported that the 2015 general elections were marred by malpractices, despite being largely peaceful. likewise, the european union election observation mission (2019) reported that nigeria’s 2019 general elections were marked by severe operational and transparency shortcomings, electoral security problems, and low turnout. in addition, journalists were subjected to harassment and scrutiny of the electoral process was largely compromised with some independent observers obstructed in their work by security agencies. ojukwu, mazi and maduekwe (2019) reported that the 2019 general elections were, to a large extent, fraudulent as they were marred by the inflation of election result figures, multiple voting, falsification of election results, delay in commencement of voting, and result manipulation. all of which resulted in the subversion of the will of people. technologies deployed by inec in the fourth republic in contrast to the collapsed first (1960-1966) and second (1979-1983) republics, and the aborted third republic (1993), democracy has been sustained in the fourth republic (1999 till date). the independent national electoral commission (inec) was established by the 1999 constitution of nigeria and has been able to conduct successive elections to sustain the nation’s democracy. the first and historic nigeria general election in the fourth republic was conducted in 1999. voters registration was done manually with pen on a form provided by inec. there was no database of voters nor was any information communication technology (ict) introduced to reduce multiple registrations (the carter center, 1999). hence, the electoral process of 1999 allowed all forms of malpractice. in 2003, inec introduced the optical magnetic recognition (omr) forms to be used simultaneously with the manual approach of 1999 election. inec also added a database for scanned records from the omr forms which were processed to produce the voters’ register. there is a unique number on each omr form for every registered voter. this number, together with register voter’s thumb print and other necessary details, were used in printing the temporary voters card (tvc). inec attempted to address the issue of multiple voting by introducing the afis (automated finger prints identification system) to clean the register (ayeni and esan, 2018). https://doi.org/ 4/7 tunmibi, sunday & olatokun, wole (2023) security and preservation of election data in nigeria in the fourth industrial revolution 47(34), pp. 1-7. doi https://doi.org/10.29173/iq1054 in 2007, inec introduced the direct data capture machines (ddcm) for voter registration. the motive for the procurement of ddcm was to eliminate multiple registrations and multiple voting. unlike the omr technology, ddcm allows the capturing of a voters’ photograph, which created a more robust database for the electronic voters’ register. in the 2011 general election, inec procured more ddcms and applied more effective afis in an attempt to eliminate multiple registration. inec also introduced the use of electronic mail for transmitting results from local governments and states to the headquarters in abuja. although not fraud-free, the overall approach seemed to produce a more credible general election than the previous elections in the fourth republic (nwagwu, 2016). more sophisticated icts were procured by inec for the 2015 elections. for the afis, inec ensured that two finger prints were captured during the voter registration exercise. inec introduced the permanent voter cards (pvcs) to replace the tvc. the commission also introduced the smart card reader (scr) which allows the accreditation of voters, authentication for identifying a voter’s face, verifying the validity of the pvc, and authenticating fingerprints. with this process, no voter can be accredited twice because the voter identification number (vin) is stored in the scr after the accreditation of the pvc. although reports revealed that inec developed the e-collation website before the 2015 elections, the implementation of e-collation did not commence until 2016 (ayeni and esan, 2018). inec maintained the continuous voters registration for the 2019 general elections, procured more scr and improved on the afis. in recent elections, inec adopted full implementation of electronic transmission of results, in real time, from polling units to collation centers. the need to align electoral process with the 4ir different studies have revealed challenges to nigeria’s electoral process, including data security and voter authentication issues (agbaje & adejumobi, 2006; onebamhoi, 2011; chuks & arnesh, 2019). also, observations have shown many deficiencies in the deployment of some of the current technologies adopted by inec. for example, there are cases of rejected fingerprints for the biometric recognition process during the 2015 and 2019 elections, which could be due to the poor quality of the fingerprint scanners. there are also cases of inadequate biometric verification attributed to poor picture quality. hence, there is a need to develop a smart technology infrastructure to aid in the delivery of credible electoral process. inec, government and electoral stakeholders should enhance the electoral system by promoting the alignment of the electoral process with relevant 4ir technologies. this will aid the preservation and security of nigeria electoral data. there are three 4ir technologies that could be of great benefits to inec. these technologies are the internet of things, cloud computing and artificial intelligence. with the adoption of smart biometric technologies (the internet of things), inec should be able to properly identify voters, which will increase elctorates’ confidence, renew interest in the electoral process and increase voter turnout. likewise, cloud computing could be integrated into electoral system technology for cloud storage of electoral data. with a cloud storage system, inec is guaranteed a continuous stream of electoral data from different units into servers managed by a hosting platform. hence, by implementing cloud computing, inec could solve a major challenge of ballot tampering, as all data will be secured in cloud storage. artificial intelligence, using machine learning techniques, could also be adopted to detect patterns of electoral manipulations in some polling stations and the extent of the manipulations across polling units within a ward or state. https://doi.org/ 5/7 tunmibi, sunday & olatokun, wole (2023) security and preservation of election data in nigeria in the fourth industrial revolution 47(34), pp. 1-7. doi https://doi.org/10.29173/iq1054 hacking and cybercrime are potential concerns in the modern era. hackers can hack into insecure 4ir electoral technologies to manipulate results and undermine voters’ confidence in the election process. stakeholders must realize that both foreign and local actors could be interested in tampering with election results in favor of an acceptable candidate. therefore, it is important for the nigerian government to introduce policies that support the adoption of 4ir electoral technologies to secure the electoral data. these policies include, but are not limited to, support of a sophisticated artificial intelligence based electoral data security center; smart biometric technologies for voters’ verification, authentication and accuracy; cyber security solutions using the internet of things, and artificial intelligence for machine learning. regular internal and external vulnerability assessment and penetration test for the whole platform; and training and necessary skill acquisition programs for operators must be done on an on-going basis. conclusion this study focused on the security and preservation of election data in nigeria in the fourth industrial revolution. the authors have shown, citing different studies, that all elections conducted in the fourth republic were marred with electoral malpractices. although inec has invested in different technologies in order to improve the credibility of the elections, much more can still be done. despite the introduction of biometric technology such as smart card reader (scr), inec still struggles with voters’ authentication concerns. hence, inec and electoral stakeholders need to rethink and consider strategies to align the electoral process with the 4ir implementations. it is our opinion that alignment with the implementation of 4ir would not only help to secure and preserve nigeria electoral data, but also help inec to meet the international standard for the provision of viable, successful and generally acceptable electoral process. we therefore recommend that inec should align the electoral process with 4ir technology implementations such as the internet of things, cloud computing and artificial intelligence. the nigerian government should put relevant policies in place to harness the developments in 4ir as well as support inec and other stakeholders to acquire appropriate 4ir electoral technologies so as to secure and preserve nigeria electoral data. in addition, we recommend a continuous update of the electronic voters’ register (evr) to clean the database of illegitimate voters. references agbaje, a. and adejumobi, s. 2006. do votes count? the travails of electoral politics in nigeria. africa development 31.3:25-44. animashaun, k. 2010. regime character, electoral crisis and prospects of electoral reform in nigeria. journal of nigeria studies 1.1: 1-33. awopeju, a. 2011. election rigging and the problems of electoral act in nigeria. afro asian journal of social sciences 2: 1-17. ayeni, t. and esan, a. 2018. the impact of ict in the conduct of elections in nigeria. american journal of computer science and information technology 6:1. doi: 10.21767/23493917.100014. https://doi.org/ 6/7 tunmibi, sunday & olatokun, wole (2023) security and preservation of election data in nigeria in the fourth industrial revolution 47(34), pp. 1-7. doi https://doi.org/10.29173/iq1054 beber, b. and scacco, a. 2008. what the numbers say: a digit-based test for election fraud using new data from nigeria. paper prepared at the annual meeting of the american political science association, boston, ma. cantu, f. and saiegh, s. m. 2011. fraudulent democracy? an analysis of argentina’s infamous decade using supervised machine learning. political analysis, 19.4:409–433. https://doi.org/10.1093/ pan/mpr033 chuks, m. and arnesh, t. implications of industry 4.0 in nigeria electoral system. proceedings of the international conference on industrial engineering and operations management, toronto, canada, october 23-25, 2019. edoh, h. 2004. corruption: political parties and the electoral process in nigeria. in m. jibo and a.t. simbine. eds. contemporary issues in nigerian politics, ibadan: jodad publication. european union election observation mission. april 13, 2015. second preliminary statement on 2015 general elections of the federal republic of nigeria. http://eueom.eu/files/pressreleases/english/130415-nigeria-ps2_en.pdf. european union election observation mission. 2019. nigeria 2019 final report: general elections 23 february, 9 and 23 march 2019. https://www.ecoi.net/en/file/local/2020744/nigeria_2019_eu_eom_final_reportweb.pdf. kia, b. 2013. electoral corruption and democratic sustainability in nigeria. iosr journal of humanities and social science 17.5: 42-48. nwagwu, e. j. 2016. information and communication technology and administration of 2015 general elections in nigeria. mediterranean journal of social sciences 7.4: 303-316. ogbeidi, m. m. 2010. a culture of failed elections: revisiting democratic elections in nigeria, 19592003. haol 21: 43-56. ojukwu, u.g., mazi, m. and maduekwe, v.c. 2019. elections and democratic consolidation: a study of 2019 general elections in nigeria. direct research journal of social science and educational studies 6.4: 53-64. onebamhoi, o. n. 2011. curbing electoral violence in nigeria: the imperative of political education. international multidisciplinary journal, ethiopia 5.5: 99-110. oromareghake, p. b. 2013. electoral institutions/processes and democratic transition in nigeria under the fourth republic. international review of social sciences and humanities 6.1: 19-34. osinakachukwu, n. p. and jawan, j. a. 2011. the electoral process and democratic consolidation in nigeria. journal of politics and law 4.2: 128-138. sule, b. 2019. the 2019 presidential election in nigeria: an analysis of the voting pattern, issues and impact. malaysian journal of society and space 15.2: 129-140. https://doi.org/ https://doi.org/10.1093/%20pan/mpr033 http://eueom.eu/files/pressreleases/english/130415-nigeria-ps2_en.pdf https://www.ecoi.net/en/file/local/2020744/nigeria_2019_eu_eom_final_report-web.pdf https://www.ecoi.net/en/file/local/2020744/nigeria_2019_eu_eom_final_report-web.pdf 7/7 tunmibi, sunday & olatokun, wole (2023) security and preservation of election data in nigeria in the fourth industrial revolution 47(34), pp. 1-7. doi https://doi.org/10.29173/iq1054 the carter center national democratic institute for international affairs. summer 1999. special report series-observing the 1998-99 nigeria elections. https://www.cartercenter.org/documents/1152.pdf. tunmibi, s. and olatokun, w. 2020. application of digits based test to analyse presidential election data in nigeria, commonwealth & comparative politics doi: 10.1080/14662043.2020.1834743. tunmibi, s. and olatokun, w. 2021. monte carlo simulation of vote counts from nigeria presidential elections. cogent social sciences 7.1: 1914397. doi: 10.1080/23311886.2021.1914397. xu, m., david, j. and kim, s. h. 2018. the fourth industrial revolution: opportunities and challenges. international journal of financial research 9.2: 90-95. yusuf, i. and zaheruddin, o. 2015. challenges of electoral processes in nigeria’s quest for democratic governance in the fourth republic. research on humanities and social sciences 5.22: 110. zhang, m., alvarez, r. m. and levin, i. 2019. election forensics: using machine learning and synthetic data for possible election anomaly detection. plos one 14.10: e0223950. https://doi.org/10.1371/journal.pone.0223950 endnotes 1 sunday tunmibi holds a ph.d. in information science obtained from university of ibadan and he is presently a lecturer at lead city university, ibadan, nigeria. wole olatokun is a professor of information science at university of ibadan. he is also honorary professor at the university of kwazulu-natal and the university of johannesburg, south africa. the lead author can be reached by email: sundaytunmibi@gmail.com. https://doi.org/ https://www.cartercenter.org/documents/1152.pdf https://doi.org/10.1371/journal.pone.0223950 10 iassist quarterly fall 2011 iassist quarterly abstract remote data access, defined as the ability of a researcher to access and evaluate restricted micro data via a secure internet connection from his home desktop computer at any time, has not been implemented by a german research data centre (rdc) so far. privacy regulations and especially the problem of access control are reasons why german rdcs are not able to offer restricted data via remote data access to the research community. by initiating the “rdc-in-rdc” approach, the research data centre (fdz) of the german federal employment agency (ba) at the institute for employment research (iab) in nuremberg, germany, aims to bring data access in germany closer to the ideal perception of remote access. the basic idea is to allow remote data access from designated institutions with comparable standards at locations other than nuremberg. in a first step, access to ba and iab data will be granted from four sites in germany and one site in the us. moreover, the rdcin-rdc approach represents a change of paradigms in two respects. first, data access will be decentralised and the fdz literally brings its data closer to the researchers. second, data of the fdz will be accessible from abroad so the dissemination of micro data will be no longer restricted to national borders. the rdc-in-rdc approach may therefore be regarded as a first step towards remote access in germany and may also represent a blue print for an intensified international data sharing. keywords: micro data, remote data access, international data access introduction fostered by the rapid developments in technologies and methodologies, statistical institutions and authorities have experienced a growing demand for high-quality micro data by both the scientific community and policymakers over the past years. despite the fact that the dis-semination of micro data for scientific purposes is part of their legal mandate, the preservation of the confidentiality in the data (i.e. to prevent the disclosure of single entities) stands above all when outsiders are granted access to micro data by the statistical authorities. in order to ensure privacy for individuals and to serve the needs of the scientific community, statistical authorities usually apply a combination of different access strategies (see lane et al. 2008). these strategies may include for example the approval of projects by (statistical) authorities and/or scientific boards, the training of researchers, the anonymisation of data or the establishment of ‘safe’ settings for on-site use. widely-used examples of these strategies are public use files for off-site use or research data centres (rdcs) and secure data enclaves in order to allow on-site analyses of confidential micro data. the most efficient and for researchers most convenient type of off-site use is remote data access defined here as means by which an approved researcher may access restricted micro data for her approved project via a secure internet connection (see grim et al. 2009 and hundepool et al. 2009). she is able to do all preparations of the data and analyses off-site but the restricted micro data never leave the safe setting of the statistical authority or an rdc. after the program codes of the researcher have been processed with the data, the outputs are screened and sent back. depending on the prevailing national data protection acts, remote access systems may even allow researchers to actually see the underlying data. statistical authorities in many countries have undertaken efforts to make micro data accessible for the scientific community and to establish remote data access the researchdata-centre in research-datacentre approach a first step towards decentralised international data sharing by stefan bender, jörg heining1 rdc-inrdc iassist quarterly fall 2011 11 iassist quarterly systems over the past years. in contrast to other countries in europe, north america or oceania which have already succeeded in the implemetation of remote access systems, however, germany still lags behind this development. legal concerns still prevent the implementation of such access ways to confidential data. besides the varying stages of development of remote access sys-tems, the diversity of national data protection legislations leads to considerable differences in the scope of performance and services provided by the particular data access sytems. the systems may be limited in terms of statistical analysis tools or only provide data access to a limited number of, or only parts of, data products. moreover, remote access to micro data is usually restricted to national borders. as pointed out by ahmad et al. 2009/2010, the limited enforceability of contractual terms and penalties abroad, virtually restricts data access to resident researchers due to high transaction costs for non-residents. with the research-data-centre in research-data-centre (rdc-in-rdc) approach the research data centre (fdz) of the german federal employment agency (ba) at the institute for employment research (iab) in nuremberg, germany tries to overcome the existing legal barriers and to bring micro data access in germany closer to the ideal perception of remote access. the basic idea of this approach is to allow data access from designated national and international institutions with comparable standards as the fdz site in nuremberg. by using a secure internet connection, researchers can link to a server and access the whole scope of micro data available for on-site use in nuremberg. in a first step, fdz data may be accessed from four rdc sites of the statistical offices of the länder2 in germany. moreover, a fifth site at the michigan center for the demography of aging (micda) enclave3 at university of michigan’s institute for social research (isr) in ann arbor, michigan, usa represents the international component of the rdc-in-rdc approach. the sole aim of the project is not merely the facilitation of access to fdz data in germany or the us. it is intended to gather experiences and by doing so, to build expertise among statis-tical authorities, researchers and data protection officials in decentralised ways of data access. the rdc-in-rdc approach may not be equivalent to remote data access, but it may serve as stepping stone, especially for countries where legal concerns are still hindering the establishment of decentralised access ways to restricted micro data. moreover, due to its international aspect and the insights gained from it, this project may be beneficiary for statistical authorities all over the world. it may represent a blue print for data sharing beyond na-tional borders. the paper is organized as follows: section 2 provides a short overview of the international state of developments with regard to remote access, as well as a brief description of both the german situation and the fdz. the technical implementation of the rdc-in-rdc approach is sketched in section 3. finally, section 4 concludes. applications of remote data access international developments the german “research data centre movement” is quite a recent development (see kvi 2000 or bender et al. 2009). other countries, often with less stringent data protection legislation, have a longer tradition of operating rdcs and have already implemented remote data ac-cess systems or are currently working on it. moreover, some countries have also taken first steps towards international data sharing. several examples of remote data access systems are briefly described in what follows. one of the oldest remote data access systems for micro data is the lissy system of the lux-embourg income study (lis)4. the project began in 1983 and was extended to include the luxembourg employment study (les). the main aim of lissy has always been to make mi-cro data of a large number of countries available for comparative social research. lissy is a fully automated system running 24 hours a day, seven days a week. the users submit their statistical requests under the form of specific statistical package programs (spss, sas, stata) via the internet mailing system or a secure graphical user interface. lissy will automatically process jobs and return the outputs to the e-mail address given during the registration process. although not operated by a national statistical authority, the ipumsinternational project (integrated public use microdata series)5 is another old example of a remote data access system. impus is a collaboration of the minnesota population center, national statistical offices, and international data archives. it was set up in 1999 in order to obtain frequency counts from diverse censuses that are in compliance with data protection regulations. data are made available through a data extraction system in which users select the variables and samples they desire. they download the data and analyze them on their local computer. the cornell restricted access data center (cradc) was established 1999. as part of the cornell institute for social and economic research (ciser) in ithaca, ny, the cradc provides secure access to confidential research data. researchers of the cornell university can acquire, house, and use restricted data in cradc’s secure computing environment. after signing a cradc data user agreement, researchers can access confidential data hosted at cradc by using either a windows terminal services client, a terminal services client software or a remote desktop client6. besides the impus-international project and the cardc another example of a non-governmental institution providing access to confidential micro data via remote data access systems in the us is the national opinion research center (norc)7. norc is private entity and located at the university of chicago. the norc data enclave runs a remote data access system which mainly provides access to firm data of several governmental and non-governmental data producers and collectors, including the annie e. casey foundation or the u.s. department of agriculture among others. several governmental authorities in the us have successfully established online systems providing researchers with frequency counts and tabulations, too. in this context, the data analysis system (das) of the national center for education statistics (nces) stands out. besides tabulations the researcher may also calculate simple covariance analyses online8. the implementation of an own remote access system is currently also considered by the us census bureau. in collaboration with external experts from academia the so called microdata analysis system (mas) is planned, offering access for limited statistical analyses on full census micro data sets (see foster et al. 2009/2010). statistics denmark first disseminated micro datasets to researchers in 1986 under an “in-house researcher arrangement”. in 2001, remote data 12 iassist quarterly fall 2011 iassist quarterly access was introduced and 55 access points had already been set up by the end of 20039. statistics netherlands, too, has a long tradition of making micro data available to researchers (since the early 1990s). after the demand by researchers for on-site access had reached a very high level, remote data access was introduced in 2006 (onsite@home). researchers can access dutch micro data by means of some special software which is installed on a regular desktop computer, located in a separate and lockable room at the researcher’s institution. by 2009, this special software has been installed on 45 terminals, one even located in italy. (see hoeve 2009/2010). the australian bureau of statistics also operates a remote data access system (radl), which was set up in april 2003 (see tam et al. 2009/2010). the radl system works in three steps. researchers submit their programs via a secure website, where they are first checked for illegal commands. if this check finds no such commands the program is run and the out-come is automatically checked. there is an additional audit process in which output is manually inspected to ensure that the analysis using the micro data does not violate any legal regulations10. since 2009, the australian bureau of statistics and statistics new zealand are providing mutual access to anonymised micro data by using the radl system. australian data may be accessed in new zealand and vice versa (see upfold et al. 2009/2010 and tam et al. 2009/2010). statistics sweden has had a remote data access system (mona) since 2005. this system provides researchers with the possibility of remote access from any computer with internet access11. the mona system is based on communication between a terminal server and a terminal client. by using a secure internet connection users access a terminal server where they can start applications remotely. for more extensive processing a batch environment is available. in the united kingdom, two remote data access applications are currently operated (see ritchie 2009/2010). the virtual microdata laboratory (vml) by the office of national statistics (ons) allows ons and governmental staff access to micro data through their desktop computers. researchers from other institutions use designated thin terminals at government offices instead. to overcome this disadvantage for non-ons and non-governmental staff, the secure data service (sds) hosted by the uk data archive has become fully operational in 2010. the sds enables safe and secure remote access for approved researchers to the data of the british household panel survey. the sds operates using thin-client and citrix technologies, whereby data are available only via a controlled network (see wright 2009). statistics canada introduced the so called “real time remote access“ (rtra) in 2010. rtra is partially based on the radl model developed by the australian bureau of statistics. researchers will submit their requests through a secure portal to a protected server located on the secure statistics canada network. after a check for forbidden commands, the syntax will be processed with the data. disclosure control will be automated as well as notifications to the submitting researcher (see goldmann 2009/2010). also in 2010, the french remote access centre casd (centre d’accès sécurisé distant aux donnés) became operative. designed and developed by the national institute of statistics and economic studies (insee), casd provides access to household data in france. the casd system is exceptional since it is a hardware-based solution using the so-called sd-box (patent pending). after being installed in the researcher’s institution, the sd-box provides a secure biometric access between the researcher and a secure server hosting confidential data. about thirty research projects in france and one project in the united kingdom have already used casd (see gadouche 2011). the implementation and operation of remote access systems is not limited to national authorities or organisations. comparable developments are also taking place on the transnational level. to access micro data sets of the european union (eu)/eurostat researchers still have to visit the safe centre of eurostat in luxembourg. in order to facilitate data access the essnet-project “decentralised access to eu microdata sets” 12 was established. the idea is to develop decentralised access by which a researcher from a certain eu member state can use european datasets in his member state. the concept of research data centres which has been already realized in some european countries as well as (the concept of ) the safe centre of eurostat could be examples for a decentralised access to european micro data sets. the essnet-project has showed first results of allowing access to european micro data in safe centres (on site). it included the methodology, guidelines and requirements which are essential to implement an access to european micro data in safe centres in the member states. a follow-up project is planned. for an overview on all these activities on the european level see bujnowska et al. 2009/2010. the data without boundaries (dwb) project is another transnational initiative which started in may 2011 and is funded by the 7th framework programme for research and technological development (fp7) of the european commission. the project brings together data archives, national statistical institutions and universities. the objective of dwb is to develop an integrated model where the best solutions for micro data access are available irrespective of national boundaries but are flexible enough to fit national arrangements. hence, dwb aims to achieve standardization and harmonization of micro data access methods as a concerted effort on a european scale. situation in germany access to restricted micro data stemming both from administrative processes and surveys was rather limited in germany until ten years ago. in 2001, the commission to improve the informational infrastructure by cooperation of the scientific community and official statistics (kommission zur verbesserung der informationellen infrastruktur zwischen wissenschaft und statistik, kvi) finally recommended the foundation of research data centres for public producers of micro data in germany (see kvi 2001). the establishment of rdcs at the statistical offices, the german pension insurance fund and the federal employment agency (bundesagentur für arbeit, ba) resulted in standardised access methods for restricted data collected by the federal statistical office, its regional offices and by the labour and social security administration (see bender et al. 2009). german rdcs currently provide two methods of access to restricted or weakly anonymous micro data: either by controlled remote execution or on-site use at the premises of the rdc. controlled remote execution is a limited mode of remote data access. it means that external researchers send evaluation programs to the rdc, where rdc employees conduct the evaluations, check the results for compliance with data protection regulations and send the tested results to the researcher. in contrast to remote data access, remote execution is general not automated. hence, it may be regarded as a sequence of single tasks conducted by rdc employees rather than as an integrated iassist quarterly fall 2011 13 iassist quarterly and automated system as operated by statistics canada or the australian bureau of statistics. remote execution is inefficient in two respects. first, as the researchers have no direct contact with the data, they sometimes program “blindly”. as a consequence, programs have to run several times until the desired evaluation is obtained. second, the level of support required from the rdc staff is high. research visits for on-site use at special separate workplaces for guest researchers at the rdc avoids these problems as the researcher has direct access to the data. however the researchers have to travel to the rdc, a growing number of them even from abroad. this often entails high travel and accommodation expenses. although not strictly forbidden by law, the implementation of true remote data access systems which provide access to weakly anonymised micro data in germany was hindered by the concerns of data protection officers so far. their main concerns focus on the additional information available from the internet and on the question of access control. since remote access systems allow data access from outside a safe environment like an rdc, the usage of additional information cannot be controlled. nowadays, a vast amount of (additional) information is easily accessible via the internet for everyone. it is almost impossible to prevent information from the internet from being used to disclose a single entity in the data. as a consequence, only absolutely anonymised data may be accessed by remote access systems in germany (see schaar 2009). moreover, when processing confidential data outside an rdc via a remote data access system, the problem of how to ensure access control arises. it is questionable whether in the environment of the researcher’s home office or workplace only the approved researchers will have access to the confidential data. since the mere visual inspection of the data represents a transmission according to german law (§ 67, subpar-agraph 6, number 3, second sentence of the social code x), an individual, for example a family member may get unauthorized access to confidential micro data by just glancing on the computer screen at the home office of the approved researcher. the research data centre (fdz) of the german federal employment agency (ba) at the institute for employment research (iab) when the fdz was founded in december 2003, there had been no systematic access to social data up until that point. following a positive evaluation by the german council for social and economic data in april 2006, the fdz was permanently established as an independent research data centre of the ba at the iab. an evaluation by the german council of science and humanities in 2007 confirmed that the fdz was an internationally unique institution: “the research data centre (focusing on methods and data access) is an internationally visible, indispensable service institution, unique in europe and a prime example to other institutions, possessing large datasets of scientific importance.” (wissenschaftsrat (german council of science and humanities) 2007, p.55) the fdz prepares individual datasets developed in the sphere of social security and in employment research and makes them available for research purposes – primarily for external researchers. with its website (http://fdz.iab.de), the documentation and working tools available online, and its workshops and users’ conferences, the fdz makes it easier for external researchers to work with the datasets. the micro datasets available at the fdz include the iab establishment panel, the sample of integrated labour market biographies (siab), the ba employment panel (bap), the establishment history panel (bhp), the linked-employer-employee data from the iab (liab) and the panel study “labour market and social security’” (pass) among others. the fdz serves not only the national but also the international market. one important step towards internationalisation in 2007 was releasing web pages in english and having almost all of the data documentation translated (see bender et al. 2009). in addition to this, members of the fdz have given numerous talks on the fdz, the projects of the fdz and the available data at international conferences and foreign universities. as a consequence, the number of users from abroad constantly increased over the past years. several of these international data users also participated in one or more of the four users’ conferences orga-nized by the fdz. besides these activities the fdz is also involved in several international projects focusing on the creation of new data products or the further development of data access ways. for the project blue-enterprise and trade statistics13 (blue-ets) the fdz cooperates with the university of southampton and the italian institute of statistics (istat) in the development of better test data for complex linked-employeremployee-data. these new test data not only reproduce the structure and content of the original and confidential data, they also share the same statistical properties. due to the higher resemblance of this new kind of test data with the original data, researchers can prepare their program codes for remote execution in a much more efficient way. the fdz is also (co-)organizer of the “workshop on data access” (wda). representatives of several national and international rdcs meet at wda to discuss new developments and to exchange practical experiences. three workshops have been held already. other international projects of the fdz include the data without boundaries (dwb) project as described above and, of course, the rdcin-rdc approach. implementation of rdc-in-rdc the central idea of rdc-in-rdc approach is to enable data access from other rdcs or institutions (called “guest-rdc” in the following) which share comparable security standards as the rdc (called “data-rdc” in the following) where the data are actually stored, but which are located at different sites. in doing so it does not matter whether the guest-rdcs are located in germany or abroad. the data are accessed in a similar way to the on-site use at the rdcs. the only difference is that the guest researcher’s room is not at the local (data-)rdc (for instance in nuremberg) but at another (guest-)rdc. in the pilot project the fdz is the data-rdc. the guest-rdcs can be institutions which fulfill the security requirements of the fdz. these include all german rdcs14 as well as comparable institutions in other countries. in order to improve access to the data of the ba and the iab in germany, the rdc of the german statistical offices of the länder and the fdz are working together on this project. the statistical offices of the länder in the federal states of berlin/brandenburg, bremen, northrhine westphalia and saxony are participating in the project as pilot locations. data access for researchers abroad is to be improved by means of cooperation between fdz and the micda data enclave at university of michigan’s institute for social research (isr) (see figure 1). in the following sections various aspects regarding the “rdc-in-rdc” mode of data access are explained in more detail, in particular the division of tasks between the data-rdcs and the guest-rdcs and issues concerning the technical implementation. 14 iassist quarterly fall 2011 iassist quarterly applying for data access the work of the german rdcs is influenced by different legal framework conditions (social code and federal statistics act). for instance, by legal definition the data available at the fdz are so called social data. the dissemination of social data is regulated by the social code (sozialgesetzbuch sgb). on the other hand, data from the statistical offices are not defined as social data and are made available on the basis of the federal statistics act (bundesstatistikgesetz bstatg). because of this difference in the legal definitions, access procedures differ, too. when applying for social data, for example, researchers have to out-line to what extent the project is related to the social security system in germany. this is definitely not necessary when applying for data of the statistical offices since they are by definition no social data. it will therefore be very difficult to standardise the respective access mechanisms or to transfer these different regulations between the rdcs. as a consequence it is necessary for users to continue to submit their applications for data access to the data-rdc. at the fdz all external researchers continue to submit an application for data access in accordance with the social code . after the application has been approved by the german federal ministry for labour and social affairs (bundesministerium für arbeit und soziales bmas), the fdz concludes a data access agreement with the institution of the user and the user itself. in this agreement the user undertakes to comply with the data protection regulations recorded in the agreement and to bear the consequences stipulated by german law if the agreement is breached. access to the authorised data in order to enable access to the requested data from a guest-rdc it is necessary to develop a new technical concept: the requested data can be accessed from dedicated workstations at the guestrdcs (see also figure 2). for this, the same security criteria must be fulfilled at the guest-rdc as apply at the data-rdc. data access should occur via a secure data line. from the dedicated workstation at the guest-rdc, the researcher logs onto a server of the data-rdc using the remote desktop function and a password. a researcher working at the guest-rdc has the same access rights as the researchers conducting analyses at guest workstations in the data-rdc. in this case, the researcher obtains access to certain servers and to certain directories within the local guest network on these servers. he or she is thus not given the opportunity to intrude into the home network of the data-rdc. similar to a research visit at the data-rdc, the researcher at the guest-rdc can only look at results on the computer screen. it must be guaranteed that only authorised users work with the data at the guest-rdc. at the guest-rdcs located in germany, the supervision is to be performed by employees of the guest-rdc whose data security expertise is regarded as equivalent to that of the staff at the data-rdc. the task of the staff at a guest-rdc is essentially data access control, i.e. they ensure that only the individuals named by the data-rdc gain access to the data. for data protection reasons the employees of the guest-rdc themselves do not gain access to the data, nor may they access the guest researchers’ directories. hence, a researcher may only print results or transmit them electronically after approval by a member of staff of the data-rdc. a different solution for access control has to be found for the guest-rdcs which are located outside germany. according to the requirements made by the data protection experts at the ba/iab, the us site has to be supervised by trained on-site fdz employees in order figure 2: scheme of the rdc-in-rdc approach. while the conclusion of the use agreement and the output control take place at the data-rdc, physical access control is necessary at the guestrdc abroad iassist quarterly fall 2011 15 iassist quarterly to ensure access control and to guarantee the compliance of german data protection regulations in the us. since the hours of work in an us site differ considerably from the local office hours in nuremberg, trained employees of the fdz with administrator rights are required in order to maintain regular operation, too. technical implementation for data access control in the context of the rdc-in-rdc approach a thin client solution using citrix software will be used for data access control and data access restrictions (see figure 3). other methods, for example biometric authentication, webcam monitoring or hardware authentication are possible and used internationally (on this issue see also grim et al. 2009 or rowland 2003), but bear either technical disadvantages or, in case of webcam monitoring, are not compatible with german law. within this technical solution, normal pcs are turned into so-called thin clients15 . thin clients may be regarded as conventional computers which are only able to perform a limited scope of tasks. thus, for instance, all possibilities for external copying (onto usb, cd-rom, dvd), for internet access (including wireless access) or for printing are disabled, the existing software is restricted and access to data is also limited. by running special software (for example citrix), the thin client guarantees that a user only uses the approved drives and directories and also prevents him/her from being able to install additional programs. this is a component of the common solution for remote data access in other countries. researchers connect from the data-rdc via a citirx access gateway to a terminal server which stores the data. access to this terminal server is provided by an encrypted ssl connection (see figure 3). output control and transmission output control continues to be performed at the data-rdcs for several reasons. first, the employees of the guest-rdc have no legal right to access the data. in addition, different legal framework conditions apply for the different rdcs, which influences the monitoring of output for statistical confidentiality. the regulations for monitoring statistical confidentiality can therefore not easily be standardised. therefore, control remains at the data-rdc and the data-rdc transmits the monitored output to the respective researcher. conclusion remote access is regarded as an efficient and convenient method of data access which has already has been implemented in several countries such as the united states, sweden or the netherlands. due to legal restrictions, germany still lags behind this development. by initiating the rdc-in-rdc approach the fdz aims to bring data access in germany closer to the ideal perception of remote access. by allowing data access from designated institutions with comparable standards but locations other than nuremberg. in a first step, access to ba and iab data will be granted from four sites in germany and one site in the us. moreover, the rdc-in-rdc approach represents also a change of paradigms in two respects. first, before the implementation of the rdc-in-rdc approach, researchers had to come or connect to a rdc in order to access sensitive micro data. now, “access is distributed rather than data” (ritchie 2009/2010, p. 113). by establishing a decentralised way of data access the fdz literally brings data access to the researchers. second, data access will also be possible for non-resident researchers. thus, the dissemination of micro data will be no longer restricted to national borders. the successful implementation of the rdc-in-rdc approach may not only serve as a stepping stone for statistical authorities in germany on their way to remote data access. it may also serve as a role model for other countries with a comparable state of development in terms of micro data access. even countries with well established remote data access procedures may benefit from the experiences gained by the rdc-in-rdc approach. because of its international dimension, it may represent a blue print for shifting data access beyond national borders leading to intensified international data sharing. references ahmad, n., de backer, k., yoon, y. (2009/2010) an oecd perspective on microdata access: trends, opportunities and challenges, in: statistical journal of the iaos: journal of the international association for official statistics, 26, 3-4, 57 63 anderson, o. (2003) from onsite to remote data access – the revolution of the danish system for access to micro data, united nations statistical commission and economic commission for europe conference of european statisticians working paper no. 29. bender, s., himmelreicher, r., zühlke, s., zwick, m. (2009) improvement of access to data set from official statistics in: building on progress – expanding the research infrastructure for the social, economic, and behavioral sciences, budrich unipress ltd., opladen & farmington hills, mi, pp. 215 – 230 bound, j. (2008) michigan center on the demography of aging proposal to the national institute of aging (p30 ag012846-16) borchsenius, l. (2005) new developments in the danish system for access to micro data, invited paper to the joint unece/eurostat work session on statistical data confidentiality (geneva, 9-11 november 2005) bujnowska, a., museux, j. (2009/2010) release of european union microdata, ess projects on remote access, in: statistical journal of the iaos: journal of the international association for official statistics, 26, 3-4, 89 – 94 foster, l., jarmin, r., riggs, l. (2009/2010) resolving the tension between access and confidentiality: past experience and future plans at the u.s. census bureau, in: statistical journal of the iaos: journal of the international association for official statistics, 26, 3-4, 119 – 128 figure 3: technical implementation of the rdc-in-rdc approach. 16 iassist quarterly fall 2011 iassist quarterly gadouche, k. (2011) technological aspects concerned in widening access to confidential data in france, paper presented at the new techniques and technologies conference (ntts 2011), brussels, 23 february 2011, www.ntts2011.eu goldmann, g. (2009/2010) from a seed to a forest: microdata access at statistics canada, in: statistical journal of the iaos: journal of the international association for official statistics, 26, 3-4, 75 – 87 grim, r., heus, p., mulcahy, t., ryssevik (2009) secure remote access system for an upgrated cessda ri, metadata technology, cessda ppp, http://www.cessda.org/project/doc/cessda_ri_sra_final.pdf hoeve, f. (2009/2010) microdata access in the netherlands, in: statistical journal of the iaos: journal of the international association for official statistics, 26, 3-4, 95 100 hundepool, a., domingo-ferrer, j., franconi, l., giessing, s., lenz, r., longhurst, j., schulte norholdt, e., seri, g., de wolf, p (2009) handbook on statistical disclosure control – version 1.1. a [eurostat] centre of excellence for statistical disclosure control, http://neon. vb.cbs.nl/cenex/cenex-sdc_handboo k.pdf (2009), accessed 15 october 2009 kommission zur verbesserung der informellen infrastruktur zwischen wissenschaft und statistik (kvi) (2000) wege zu einer besseren informellen infrastruktur, nomos, baden-baden. lane j., heus p., mulcahy t. (2008) data access in a cyber world: making use of cyber-infrastructure, in: transactions on data privacy, 1, 216 ritchie, f. (2009/2010) uk release practices for official microdata, in: statistical journal of the iaos: journal of the international association for official statistics, 26, 3-4, 109 – 117 rowland, s. (2003) an examination of monitored, remote microdata access systems, presented at the nas workshop on access to research data: ‘assessing risks and opportunities‘, www7.nationalacademies.org/cnstat/rowland_paper.pdf. schaar, p. (2010) data protection and statistics – a dynamic and tension-filled relationship in: building on progress – expanding the research infrastructure for the social, economic, and behavioral sciences, budrich unipress ltd., opladen & farmington hills, mi, pp. 629 642 söderberg, l.-j. (2005): mona – microdata on-line access at statistics sweden, united nations statistical commission and economic commission for europe conference of european statisticians working paper no.3. tam, s., farley-larmour, k.,gare, m. (2009/2010) supporting research and protecting confi-dentiality. abs microdata: current strategies and future directions, in: statistical journal of the iaos: journal of the international association for official statistics, 26, 3-4, 65-74 thygesen, l., anderson, o., schnoor, o. (2003) the danish system for access to microdata; from on-site to remote access, paper presented at swedish workshop on microdata, stockholm, www7.nationalacademies.org/cnstat/rowland_paper.pdf upfold, j., ng, p. (2009/2010) new zealand’s approach to the provision of access to micro-data, in: statistical journal of the iaos: journal of the international association for official statistics, 26, 3-4, 95 101 wissenschaftsrat (german council of science and humanities) (2007) stellungnahme zum institut für arbeitsmarktund berufsforschung (iab), nürnberg, drs. 8175-07 wright, melanie (2009) esrc secure data service: a new vision for secure data access, talk given by at new services for social science research: the administrative data liaison service and the secure data service, royal statistical society, london, 14 december 2009, http://securedata.ukda.ac.uk/news/publications.asp notes 1. stefan bender, institute for employment research (iab), regensburger strasse 104, 90478 nuernberg, germany, tel.: +49-911-179-3082, fax: +49-911-179-1728, stefan.bender@iab.de; jörg heining (corresponding author), institute for employment research (iab), regensburger strasse 104, 90478 nuernberg, germany, tel.: +49-911-179-1752, fax: +49-911-179-1728, joerg. heining@iab.de. we would like to thank all of the involved project partners for collaboration and support, especially ramona voshage, statistical office of berlin-brandenburg, sylvia zühlke, statistical office of north rhine-westphalia and maggie levenstein, institute for social research, university of michigan. for many helpful comments and discussions we would also like to thank peter jacobebbinghaus, institute for employment research (iab), karen scott-leuteritz as well as the participants of statistics canada’s 2010 international methodology symposium and the new techniques and technologies in statis-tics (ntts 2011) conference by eurostat. the project underlying this report was funded by the federal ministry for education and research (bundesministerium für bildung und forschung) (grant number 01uw1002). the authors are responsible for the content of this publication. 2. a detailed description of these institutions is given in bender et al. 2009. 3. bound 2008 provides a description of the micda enclave. 4. http://www.lisproject.org/ 5. https://international.ipums.org/international/ 6. http://ciser.cornell.edu/cradc/what_is_cradc.shtml 7. http://www.norc.org/dataenclave 8. http://nces.ed.gov/dasol/ 9. http://www.dst.dk/homedk/tilsalg/forskningsservice.aspx (in danish); see also anderson (2003), borchsenius (2005) or thygesen et al. (2003). 10. http://www.abs.gov.au/websitedbs/d3310114.nsf/home/ curf:+remote+access+data+laboratory+(radl) 11. http://www.scb.se/pages/list____257147.aspx (in swedish); see also söderberg (2005). 12. http://www.safe-centre.eu/ 13. http://www.blue-ets.istat.it 14. an overview of the german rdcs is given on the website of the german data forum 15. (http://www.ratswd.de/eng/dat/fdz.html). http://en.wikipedia.org/wiki/thin_client mag the impact of future social and technological trends on the dissemination of census bureau information by donald l. day ' school ofinformation studies syracuse university abstract this study examines social and technological trends that may impact the dissemination of u.s. census information via the depository library program in the year 2000 and beyond. the study looks beyond currently emerging systems to examine a limited list of future issues in technology, regulation, funding, access, and user demand. it examines information dissemination in the broad, societal context, rather than concentrating narrowly upon the means of delivery. its main objectives are to pinpoint key issues, to stimulate an appreciation of the inextricable nature of information in postindustrial society, and to recommend policies and directions for further research. introduction nature of the topic this study examines social and technological trends that may impact the dissemination of u.s. census information via the depository library program in the year 2000 and beyond. importance of the topic establishing the social and technological context within which census information might be disseminated in the future would facilitate rational policy making regarding that dissemination. it also might save time and money, in that decisions about systems with long lead times could be made so as to implement the systems in a timely and efficient manner. scope and objectives the original impetus for the study was a concern about the sheer volume of paper products absorbed by the depository library i*rogram (as high as 25 percent of all documents printed from the 1980 census), and an interest in how new technologies may make it possible to reduce costs while increasing the availabihty and variety of data (in particular, census data). the study looks beyond currently emerging systems to examine a limited list of future issues in technology, regulation, funding, access, and user demand. it examines information dissemination in the broad, societal context, rather than concentrating narrowly upon the means of delivery. its main objectives are to pinpoint key issues, to stimulate an appreciation of the inextricable nature of information in postindustrial society, and to recommend pohcies and directions for further research. point of view this study was fielded under the presumption that government will be required to continue providing public access to federal information as part of its commiunent to maintaining the informed citizenry that is central to participatory democracy. report outline following a brief review of study methodology, this report presents an overview of future trends (the societal context), a discussion of study findings, and pohcy recommendations for key issues requiring public debate. study method design the current study was conceived as a qualitative, exploratory and descriptive effort to identify issues of concern. ehie interviewing was to be interspersed with a review of literature in the future studies field in an iterative process that would develop perspective and deepen focus and selectivity in data collection. the interviews were in-depth but informal, as recommended by marshall & rossman (1989). they were conducted in subjects' normal work environment (a "natural setting") to ensure ease of discussion and immediate access to reference materials. procedure subjects were selected because of their expert knowledge of federal information dissemination policies and of developing technology. interviews were conducted in two rounds, separated by approximately four weeks. the first set concentrated on preliminary data gathering and focused upon the depository library program. the second set was guided by an outline of concerns (interrogative research questions) developed from the literature and from the earlier interviews (see below). sessions were recorded to ensure the completeness of notes and to help recognize nuances that might have been overlooked at the time of initial data collection. only one subject lassist quarterly was distracted by the presence of a taping machine; the others were largely oblivious to its presence. the scheduling of multiple interview events separated by a review of the literature and by conceptual outlining worked very well in focusing the research and in identifying issues that were not apparent at the outset. an improvement in procedure that might be useful in other qualitative research on this or other topics would be submission of interim reports to key sources for critique. this technique might help sources to feel involved in the research effort, lend to focus their comments during interviews, and encourage them to act as agents of the researcher in obtaining material useful to the study. key research questions the key research questions drafted from an analysis of the literature and during the interviewing process were as follows. 1. what will be the leading edge information technologies in the first decade of the next century? 2. when and to what degree will depository library materials (especially census data) be distributed via cd-rom or other machinereadable media? 3. what technological developments will affect patrons' remote electronic access to depository libraries? 4. what software and data structures will be required for electronically disseminated census data, to facilitate rapid and effective searches and retrieval? 5. what will be the sponsorship and impact of standardization efforts to facilitate network access to federal government information? 6. to what extent will anti-trust concerns inhibit development of data integration protocols and telecommunications software necessary for widespread network access to federal government information? 7. how will the distribution of government information be controlled, under whose auspices and with what objectives? 8. how will data integrity be maintained without impeding widespread electronic dissemination of information? 9. which sponsors of information production, dissemination and use will support high technology access, under what conditions and with what goals? 10. what are the prospects that congress will choose to privatize depository library distribution? what impact would that have upon the quality, quantity, availability and cost of census bureau information? 1 1 . what will be the minimum skill levels required of users and depository librarians in accessing electronically disseminated information? 12. to what degree might user fees and other costs of accessing electronically disseminated information disenfranchise individuals? 13. what impact will changes in work force composition and employment arrangements have upon the types of census information sought by users? overview of future trends the economy although estimates vary widely, some projections forecast a period of modest economic prosperity for the united states in the next two decades, including a strong rise in "knowledge industries." demographers predict an enlargement of the middle class, with fewer very poor or very wealthy. cultural homogenization is anticipated, despite an increase in non-english speakers and the swelling ranks of citizens over 65, with each group having its own unique perspective and needs (cetron, 1988). information consumption may be influenced significantly by an increase in middle class affluence. the incomes of middle-aged citizens will rise about three-fourths. onethird of middle-aged households will have annual incomes of $50,000 or more, in constant dollars (new american, 1986). more sophisticated, better-educated consumers with work experience will have disposable income for travel, leisure and luxury. spending will continue to shift toward service industries (cetron, 1988). education and training four percent of the labor force may be in job retraining programs in the coming decade (cetron, 1988). there may be more rigorous educational standards at all levels and a greater concern for human rights and personal freedom (caddy, 1987). work force composition summer 1991 by the year 2000, manufacturing will employ only nine percent of the labor force, with services tajcing 88 percent, partly because productivity in automated industries may increase fivefold. seventy percent of u.s. homes may have computers in 2000, facilitating the potential for widespread remote access to federal government information (cetron, 1988). changes in the work force may be a key influence during the next decade. the mandatory retirement age may be 70 by the year 2000. union members will comprise less than 10 percent of the labor force (versus 29 percent in 1975 and 18 percent in 1985). the work force will be dynamic, with people changing careers an average of every 10 years. a shortage of low-wage workers will force businesses to automate and to seek foreign workers. the ranks of the self-employed will grow at a faster rate than salaried workers and more mid-career professionals will become entrepreneurs (cetron, 1988). some expect a crisis of consumer confidence in the u.s. technological infrastructure as widespread hacker activity plagues computer networks (1988 ten-year forecast, 1988). also anticipated are a decrease in the divorce rate, an increase in marriages and family formation, and a heightened role for reugion. do-il-yourself activities will be popular, because a 32-hour work week will create more leisure time and due to the high cost of services. protracted adolescence may be more common, although there will be far fewer young people than at present. a decrease in the size of federal government will be accompanied by growth in state and local governments (cetron, 1988). knowledge industries the united states is becoming a postindustrial society. in such a system, telecommunications and computers are vital to the exchange of information and knowledge (bell, 1978). multimedia information networks permeate everyday life. new information and processing devices increase productivity, despite initial retraining losses. knowledge industries grow in importance. as the central role of information accelerates, major policy issues will include privacy, the part government plays in information dissemination, intellectual property rights and functional literacy. neural networks (combinations of electronic and photonic circuits used in optical computers) may be one of the key new technologies. the science of integrating diverse telecommunications and computing systems may become an important driver of technological innovation (bezold and olson, 1986). information technologies despite the increasing use of computers in a wide variety of applications, there also may be an increase in the need for paper. tenner (1988) reported that from 1959 to 1986 u.s. consumption of writing and printing paper increased 320 percent while real gnp rose only 280 percent. he believed that electronic information supplemented rather than replaced paper, and noted that using paper is more efficient, legible, and secure than working with monitor displays. tenner also predicted that increases in the number of office workers will cause a corresponding growth in the use of photocopiers and facsimile machines (which use paper). massive increases in storage technology will take place, with commercial system capacities in the hundreds of megabytes. optical disks (some erasable) will emerge as the medium of choice in applications such as census data reuieval that require high storage capacity, fast access, removability, non-contact recording, and long life. engineers predict that current disk capacity will increase tenfold and that the 600 megabyte disk that now sells for s200 is likely to cost only s25-50 (freese, 1988). knowledge in postindustrial society the economics ofinformation social organization will be shap)ed by intellectual technology in postindustrial america. since information and knowledge are not depleted in the sense that goods are in an industrial economy, knowledge will be considered a social product. its cost, price and value will be assessed in a way vastly different from that for industrial goods, in accordance with what is known as the "knowledge theory of value" (beu, 1978). knowledge, even when it is sold, remains with the producer. it is a "collective good" — once it has been created, it is available to all. there is little incentive for any single person or enterprise to pay for the production of knowledge unless a proprietary advantage (such as a patent or copyright registration) can be obtained. thus, government policy in regard to intellectual property and contractor marketing of publicly funded products will be key in the management of future information dissemination technology. bell (1978) believes that a reduction in incentives for individuals or companies to produce knowledge will cause the responsibility for and costs of satisfying information needs to fall to government whether information dissemination is "privatized" and in what manner may affect the availability of that information significantly. the first infrastructure placed in service by industrializing economies is transf)ortaion. fully industrial systems concentrate on energy utilities. societies entering the lassist quarleriy postindustrial phase need advanced telecommunications. therefore, the major technological problem for america in the next decade will be emplacement of an appropriate digital information network, carried over fiber optic cable (bell, 1978). social impacts the u.s. as postindustrial state is likely to have a vastly different social structure than at present. bell (1978) believes that this new order may be characterized by: 1. centrality of theoretical knowledge as the basis of innovation. of society, but detrimental to others. 2. will have a positive impact primarily in the middle-class suburbs, with a negative impact in central cities. 3. will not be properly understood and regulated until considerable damage has been done in major urban development. 4. will reduce the economic viability of the central city by accelerating delocalization of business and commerce. 2. creation of new intellectual techniques to engineer solutions to economic (and even social) problems. 3. the spread of a (technical and professional) knowledge class. 4. the change from goods to human services. 5. a change in the character of work (people must learn to live with one another, since interaction among groups will be key). 6. the employment of women in expanded human services. 7. science as the societal standard bearer. 8. political units comprised of either vertical organizations of individuals into scientific, technological, administrative, and cultural centers, or of institutions arrayed as economic, government, university, or social complexes. 9. meritocracy (an emphasis on education and skill). 10. scarcities of information and of time. 11. the economics of information. bell (1978) also believes that information by its nature is collective, not private. in postindustrial america, the optimal social investment in knowledge may require that we follow a cooperative strategy to increase, spread and use knowledge. a more pessimistic view of the impact of advanced telecommunications technology is taken by eldredge (1978). hebelieves that new technology will: 1 . will be highly beneficial to some segments 5. will affect the service sector most, because its processes involve paper transactions that are particularly sensitive to technological substitution. the depository library program it is within this context of an increasingly central role for the development and dissemination of information that we consider future management of the depository library program. the federal government has a long history of providing increasing amounts of information to the public as part of its responsibility to maintain an informed citizenry. in particular, substantial volumes of information ranging from census data to contract studies of government activities to congressional hearings have been disseminated in a network of some 1400 libraries as part of the depository library program. under the program, government-printed material is distributed to a limited number of regional depository libraries. additional "select" libraries also archive some subset of these materials for their patrons.^ recently, census data have been distributed to a few depository libraries on cd-rom in an effort to assess the medium's potential, as well as to gauge user reaction.' bureau of the census material also is available to institutions and to the public at state data centers operated specifically for that piupose. discussion of key research questions in terms of findings elite interviewing and the literature review for this study identified five major areas that may affect future dissemination of census data as part of the depository library program. these areas were selected based upon the emphasis they were given by interviewees and upon the author's experience in information systems design. 1 . technology summer 1991 2. regulation 3. funding 4. access 5. user demand technology leading edge technologies what will be the leading edge information technologies in the first decade of tfie next century? the media used in dissemination of federal information in the year 2000 may be a mixture of optimized currentday technologies (mcgee, 1990). data input, the bane of current full-text efforts, may no longer be a problem. research during the past few years has made it possible for many agencies (e.g., the air force) to use improved optical scanning devices to make large amounts of printed text machine-readable (mcgee, 1990). once standards are established for sharing scanned input, an enormous amount of digital data will be available, although coordination and retrieval problems will need to be addressed. information dissemination also may include specially engraved static memory chips for advanced personal computers. response time typical of such devices would be significantly shorter than with mechanical access systems— an important consideration for complex, natural language queries of large databases. inexpensive, high-capacity chips will be available. (in a recent demonstration, ibm technicians wrote patterns at singleatom resolution using a scanning tunnelling microscope.'' thus far, vast sums of money have been spent on advanced technologies without any real understanding of how people might interrelate with these devices (weiner & brown, 1989). current improvements center on the abihly to gather, store, and catalog information. much information dissemination in the year 2000 will be via fiber optic cable, which due to its enormous capacity will compete effectively with relatively limited capacity direct-broadcast satellite transmission in many real time access applications. much of the internal transmission capacity of most telephone companies already has been converted to fiber, and a number of large businesses have access to fiber networks. all subscribers are certain to be fiber-connected in 25 years, with many equipped during the 1990s (weinstein & shumate, 1989). there will be an explosion of communications options as the fiber optic infrastructure is installed, making possible single transmission, multiple-service options such as online catalog ordering and public opinion polling (due to the signal capacity of fiber optic cable). today, the integrated services digital network (isdn) makes possible voice, data and image transmission over the same phone lines. isdn is likely to be deployed widely by the end of century, making possible high fideuty audio and five-second per page facsimile. early in the next century, isdn will give way to the broadband integrated services digital network (bisdn)— an intelligent networks supporting high quality, simultaneous and ondemand video services, home telemetry, and high speed data and image communication among faxes, workstations and computers. the telecommunications network for the deaf (tnd) will convert speech into text and vice versa to assist handicapped subscribers. these and other protocols may be tested soon in the nren high-speed network already in use at many libraries. cd-rom and on-line technology early in the next century will assist the user in sorting through the thicket of available information by means of "information grazing": an intelligent, user-selectable fdtering system capable of passing only items of interest in order to manage information overload. machine-readable media when and to what degree will depository hbrary materials (especially census data) be distributed via cd-rom or other machine-readable media? it is abeady within the means of depository libraries to provide public access to census data via cd-rom. recently, the bureau of the census has been engaged in a test distribution of files on cd-rom to depository libraries throughout the country (j. stratford, personal communication, february 21, 1990). however, the test discs have not been received well by all depositories. data on one disc were formatted in a different pattern for each file, making it necessary for users to apply different techniques to access each datasel these variations in format have decreased the usefulness of the disc (k. chiang, personal communication, april 12, 1990). also, a virus was distributed on the 1988 county and city data book cd-rom, though prompt action by the depository library program appears to have averted serious problems for member libraries (harm, 1990). despite such initial problems, cd-rom seems ideally suited to dissemination of census information which by lassist quarterly its nature contains large amounts of data. in the mid1980s, grolier placed its 21volume, nine-million word academic american encyclopedia on a single compact disc. eighty percent of the disc's capacity remained unused. in addition to enormous space savings, use of cd-rom made possible high-speed keyword searches (cornish, 1985). in the future, cds will not necessarily be distributed to all institutions in the depository library program. much in the way that only regional libraries receive full dissemination now, some expect discs bearing federal information to be distributed only to regional depository libraries. select libraries could request discs on loan as required by their patrons. (mcgee, 1990). remote electronic access what technological developments will affect patrons' remote electronic access to depository libraries? cd-rom will by no means be the only electronic access to federal information in the years ahead. combinations of new and existing technologies will facihtate the widespread availability of census data early in the 21st century. integration of facsimile devices with home entertainment centers will allow users to receive hardcopy on demand of a wide range of federal government information, if online access to depository databases is allowed. installation of fiber-optic cable to homes and businesses will make possible a dynamic, interactive census process in which users not only access more current database information but also can register data about themselves more easily and more frequently than is possible with current, paper-based census techniques. the year 2000 may be the last "traditional" census, as citizens enter the data collection, processing, analysis and dissemination loop more actively. such dynamic census systems would raise issues of quality control, user cooperation, and prevention of abuse by commercial marketing interests. extension of on-line access to the full range of depository data also may call into question the need for regionally distributed depository archives. most depository libraries already are involved in networks. there are 20 regional networks, plus some cdrom disuibution. some federal information specialists feel that depository library information should be on-hne, if only to reduce the expense of storing information that is used infrequently. now, it is enormously expensive and wasteful to print and distribute materials that few if any patrons use: "much of government information is a record going nowhere." (mcgee, 1990). regardless of storage capacity, it still will be necessary to keep truly unused information from clogging the system. it has been suggested that ubrarians help define the characteristics of a filter to be used in deciding what depository information should be part of on-line and cdrom databases. information use could be evaluated at regional libraries, or such institutions could delegate responsibility to subject specialists. however, many librarians resist the filtering of information, preferring to make decisions on a case-by-case basis rather than allow data to be withheld at centralized distribution points (mcgee, 1990). software and data structures what software and data structures will be required for electronically disseminated census data, to facilitate rapid and effective searches and retrieval? the complexity and volume of federal information in all formats will require sophisticated indexing and retrieval software, but existing government databases are seriously lacking in such tools (mcgee, 1990). software used currently to access the test cd-roms distributed by the bureau of the census also may be inadequate ("until there is adequate access software, the cd is not much of an improvement on a stack of microfiche."— k. chiang, personal communication, april 12, 1990). efficient software will need to be developed to manage the volume of machine-readable information available in the future. when the mass of retrievable information is more than the brain can process effectively, the result is not faster decision making, but instead a delay in or even abdication of decision making (weiner & brown, 1989). clarke (1985) noted that without adequate indexing, many users would be completely overwhelmed by a virtually limitless selection of information resources, and chose to select nothing. he also pointed out that with the latest techniques, it would be possible to put the whole of human knowledge into a shoe box. the problem, of course, is to get it out again; anything misfded would be irretrievably lost. within the next decade, artificial intelligence will be applied to such problems (diebold, 1985). some data in future federal databases may need to be in a format suitable for manipulation by spreadsheet and statistical programs such as lotus 1-2-3 or sas (mcgee, 1990). the widespread availability of relatively inexpensive bitmapped displays, faster processors, and a demand for three-dimensional graphics and photographic-quality images may force some network databases to be stored in summer 1991 tokenized or vector formats to be regenerated at user workstations. (graphics regeneration has been available for some time on more expensive minicomputers such as the dec microvax ii, using the ansi/iso-standard graphical kernel system.) such formats allow specialized processors in users' machines to plot images from line endpoint data rather simply displaying full-screen files downloaded from host computers. this capability would accelerate display speeds significantly, partly due to new data compression techniques and the capacity of fully digital, high speed networks to download enormous amounts of data. regulation the history of innovation is replete with examples of new technology that was not implemented to its potential because of social or political factors that frustrated its use. future data dissemination technology could suffer a similar fate if issues of standardization, anti-trust, control of distribution and data integrity are not addressed adequately. standardization what will be the sponsorship and impact of standardization efforts to facilitate network access to federal government information? the architecture of future data access systems and the standards coordination required for integration of diverse computing hardware are key concerns that will need to be addressed as technology makes large volumes of federal data available in machine-readable format. as clarke (1985) observed, another problem is to decide whether we mass produce the shoe boxes [databases!, so that every family has one, or whether we have a central shoe box linked to the home with wideband communications. information specialists at the library of congress believe that a single, coordinated, national database with remote access is not likely. interface standards will make possible an extensive, distributed architecture populated with vastly different hardware in a highly interconnected system. (substantial interconnection ah-eady exists among ntis, doe and medline.) they feel that the centralized database concept is obsolete. in this view, future online access to federal information would more likely be via highly distributed selective centers, linked to each other using standard protocols (bortnick & relyea, 1990). estabushment of standards for such on-line networks is a key concern, since integration of computers from a multitude of vendors operating under vastly different operating systems may be involved. equipment already in place at user sites could be linked to provide reasonably economical service, if sufficient standards were developed by government and applied as part of the access system. standards may not even be imposed by the federal government, but instead by international entities, owing to the substantial interconnection even now between federal government data stores and overseas sources. european networks have tended to be further advanced than those in the u.s., therefore are more likely to drive any movement toward standardization (bortnick & relyea, 1990). if an institution could not link to the system using existing hardware because its plant was antiquated, funding could be requested from corporations, the department of commerce and/or the national science foundation to bring the site to minimum levels for participation in the system. otherwise, it would be presumed that institutions could tap into the on-line system with existing equipment and software, if they met standards (bortnick & relyea, 1990). at present, standardization even within the federal government is difficult. many agencies have material in electronic form, but there is no coordination of data formats or machine compatibility. most electronically stored data are not available outside their host agency. although the office of management and budget (0mb) or general services administration (gsa) may have the authority to force coordination of data formats throughout the federal government, it remains to be seen whether that authority will be applied, and with what result (powell, 1990). anti-trust concerns to what extent will anti-trust concerns inhibit development of data integration protocols and telecommunications software necessary for widespread network access to federal government information? traditionally, federal anti-trust law has prevented firms in competition with each other from cooperating in ways that may be necessary for development of standardized protocols, equipment and software for the fully integrated information systems of the future. however, anti-trust law has changed in recent years to allow cooperative research and development among firms, largely in response to competitive market pressures from overseas (where such cooperation is common) (bortnick & relyea, 1990). there are proposals in congress to extend this liberalization to include product development and even production. the availability of seamless networks for user access to on-line federal information would be affected lassist ouarterty significantly by greater cooperation among service suppliers, if they were not constrained by anti-trust regulation. federal anti-trust enforcement has been dormant in recent years, in part because there have been few substantive changes in business activity from old, established patterns. however, the new information technology constitutes just such a substantive change, implying the need for changes that would facilitate cooperative development of technologies— especially protocols and access software— that would facilitate interconnection (bortnick & relyea, 1990) control of distribution how will the distribution of government information be controlled, under whose auspices and with what objecuves? telecommunications policy centers around who provides services, under what conditions, and at what prices. a jurisdictional struggle is under way now among the government printing office (gpo), omb, national technical information service (ntis) and others over control of federal information dissemination. the gpo is funded by the legislative branch; clashes take place between the executive branch and congress over information policy. each side has its allies in congress. (both the house and senate have committees with oversight authority for each agency involved in the policy debate.) paperwork reduction is the concern of the senate committee on government affairs^ statistical policy is set by the bureau of the census, and the depository library program is managed by the joint committee on printing. commerce and science committees also are involved, because of technology issues. it is difficult to make poucy with such fragmentation (powell, 1990), and just as difficult to implement it because of the tug of war among the actors in their attempts to influence appropriations. house bill 3849 (introduced in january) is one manisfestation of the struggle. this legislation attempts to prevent public monies from being used by the executive branch lo generate and distribute information products and services without involvement of gpo. it broadens the legal definition of "documents" to include information products "in any tangible format, medium or substate", and auempts to interpose the superintendent of documents between the depository libraries and any government agency that might issue information products (goverrunent, 1990). this oversight problem must be resolved before the depository system can upgrade, regardless of the benefits of technology. some observers feel that the depository library program will need to be removed from gpo auspices before its technological potential can be fully realized, in part because the gpo work force is hesitant to diversify from traditional media (bortnick & relyea, 1990). some parts of the library community feel threatened by the new attempts to centralize and standardize the new technology. in 1985, omb's office of information and regulatory affairs issued bulletin a130, which dealt with dissemination of federal information in electronic format and was part of the effort to reduce the volume of government printing (powell, 1990). some librarians felt the policy would undercut the depository library system (powell, 1990). initially, the cost of advanced technology may buttress such resistance by those who favor hardcopy dissemination (mcgee, 1990). questions of access to federal information networks also need to be addressed before practical implementation of rapidly developing technology. some within the federal community feel that capabilities such as full-text retrieval should be available lo all citizens, not merely to depository hbraries (mcgee, 1990). some concern has been expressed within government regarding regulation specifically of the on-line dissemination of federal information due to provisions of the export administration act (pl%-72). they feel that the competitiveness of american industry might be impaired if information were available to foreign competition. also, dod is concerned about the "mosaic theory" potential of widespread, machine-readable federal information (i.e., what new knowledge can be gained from machine correlation of public information). therefore, policy will need to be made regarding whether online information should be available only to u.s. citizens. restriction may not be feasible regardless of poucy because of the difficulties of controlling information transfer in an environment of total interchangeability among data formats. release in one format would be tantamount to release in all others (bortnick & relyea, 1990). further frustrating auempts to limit overseas access to federal information is the fact that a great deal of information in existing federal databases (e.g., ntis or doe) is from foreign sources. it is unlikely that overseas concerns would be willing to continue providing substantial amounts of information to a u.s. database if they were not allowed in turn to access that database. in any case, controls cost money which might not be forthcoming, depending upon the perceived importance summer 1991 of such restrictions compared to the perceived need for convenient and widespread public access to federal government information. data integrity how will data integrity be maintained without impeding widespread electronic dissemination of information? data in a freely accessed, on-line system could be copied, modified then redistributed, without the knowledge of the issuing agency. original data may even be subject to accidental or malicious corruption. it might be difficult to protect important data whose accuracy would be presumed to be the responsibility of the bureau of the census (bortnick & relyea, 1990). with the advent of new technology, however, a different model of responsibility might be in place. in the past, a publisher may have been expected in certain circumstances to notify the trade media (e.g.. publishers weekly) about serious errors discovered after a book was distributed. easily modified works such as those marketed in three-ring notebooks or with spiral bindings could have been distorted in the field, totally unknown to the original publisher. in the analogous future situation with machine-readable material, the aspect of the technology that makes data easy to corrupt also makes unauthorized changes relatively easy to correct, if intrusions can be detected. funding sponsors which sponsors of information production, dissemination and use will support high technology access, under what conditions and with what goals? funding is a natural barrier to the widespread dissemination of information. who should be responsible for capital equipment and telecommunications costs incurred if patrons or depository librarians are to have on-line access to census data? in all probability, future funding for information dissemination would follow existing models. these include (1) a prototype database access system sponsored by the library of congress, (2) the regional supercomputer program, and (3) the federal information center program. in the library of congress model, 12 libraries, corporalions and agencies across the u.s. use their own equipment and staff, and pay their own telecommunications charges. the library maintains the database and provides access to its computers (bortnick & relyea, 1990). the internet communications network operates on a similar model. universities across the world provide their own equipment and software; the u.s. government maintains the relatively small interconnection "backbone" of the system. in the regional supercomputer program model, access to information is restricted to institutions that provide fmancial and support for the system. contracts between government and the universities hosting the centers include specific provisions for use and funding. centers typically solicit financial support from two or three distinct sources (bortnick & relyea, 1990). under this model, depository libraries might maintain regional cdrom or online database centers which loan discs or provide logon accounts to institutions that pay to support the system. whether the subscribing institutions in turn charge their paffons for access is a major poucy option. some at the library of congress feel that the depository library program should be subsumed by the gsa's federal information center (fic) program, making it the central point from which to access many types of integrated media. these centers are not funded solely by the legislative branch, but instead are supported by many sources under a "plurality of funding" concept that is more acceptable politically (bortnick & relyea, 1990). under the fic model, a private firm (biospherics, incorporated of beltsville, maryland) manages regional offices which field telephoned inquiries regarding a wide range of government services, programs and regulations. by terms of the contract, biospherics is is required to provide access to the hearing and speech impaired (federal, 1990). under this model, depository ubrary materials presumably would be accessed on-line using regional fic database systems. depository libraries are not very high on the poutical agenda in congress, since their funding comes directly from the legislative branch budget. many legislators believe that government should provide the information, but not the means for its distribution. in the future, the federal subsidy for information dissemination might be limited to the current cost of paper distribution (bormick & relyea, 1990). telecommunication charges for the dissemination of federal information may be regulated in the future by policies such as the fcc open network architecture (ona) initiative. ona's intent is to prevent monopoly control by local telephone carriers of connections to national networks. the initiative requires modular pricing of discriminible services, to promote competition among suppliers. 12 lassist quarterly if on-line access to depository library information is to be provided in the home, then charges for terminal electronics and for the in-ground installation of fiber-optic cable must be comparable to the total cost of a typical telephone line installation today: si,000 to si,500 per home. network operation also must be economically feasible. that it may be is demonstrated by research at bellcore which has resulted in experimental models of broadband, digital networks (weinstein & shumate, 1989). during the more than a century that the depository library program has provided public information to participating institutions, the government has funded the printing and dissemination of materials. changes in perception of government's proper role— and in its abiuty to support initially expensive dissemination technology — make it doubtful that future funding will come from washington alone. funding expanded network access in particular is a serious concern. some federal information specialists believe that future on-line systems will need to be feebased. federal libraries already pay for on-line access, in a manner similar to that used to charge commercial customers for access to commercial databases such as dialog. despite these concerns, some analysts believe that basic information utilities of the future will be economical to the point of being taken for granted. their concern is that such a cheapening of information will risk potential devaluation of product quality and intellectual creativity (diebold, 1985). it should be noted, however, that low cost does not necessarily translate into easy access. regardless of cost, there still will be a need to filter the vast amounts of information that will be available in order to retrieve only what is needed. privatization what are the prospects that congress will choose to privatize depository library distribution? what impact would that have upon the quality, quantity, availability and cost of census bureau information? there are some who feel that 0mb 's circular a130 (and the corresponding requirement for a-76 studies)' is too favorable to the private sector, because it urges agencies to contract for information services (powell, 1990). this sentiment underlies hr 3849, the government printing office improvement act of 1990, which if passed would require that depository libraries apply through the superintendent of documents for any government information product, and that such application include the specific cost sharing arrangements proposed among users, the depository libraries, the issuing agency and federal appropriations (government, 1990). joint government-industry initiatives may be inevitable, not only because of federal funding constraints, but also because private industry holds advanced search and indexing software necessary for the management of comprehensive, interrelated data stores. possession of such tools by firms such as dialog and nexus place them in an excellent position to bid contracts for on-line or cd-rom access to federal information (bormick & relyea, 1990). joint initiatives are prompted in part by the concern of private industry regarding "unfair" competition with the government if federal information is placed on-line at subsidized rates. the declining federal budget also may prompt partnerships with private industry in the future. federal outlays will be a declining proportion of gnp for the rest of the century, with the growth rate of the budget steadily declining (new american, 1986). in recent years, nearly 30 percent of federal spending has gone to pay old-age benefits to the 1 1 percent of the population currently over 65. far more will be receiving benefits in the future, as the overall u.s. population ages. in 1986, interest on the national debt (the fastest growing portion of the budget) accounted for 1 8 percent of all federal spending (longman, 1988). these and other pressures on the federal treasury may force future information dissemination to be self-sufficient, by means of the involvement of commercial partners. access skill requirements what will be the minimum skill levels required of users and depository librarians in accessing electronically disseminated information? projections of median education and skill levels in the future point to a wide divergence in patrons' basic understanding of new technology and the information that it will provide. a dichotomy is developing between a small, undereducated and underskilled younger generation and a larger, educated and jobexperienced retirement cohort (longman, 1988). longman also reported that today's younger generation is not only comparatively small (due to low birth rates in the last two decades), but also that an alarming proportion of youth lack the basic skills that employers require. only 30 percent of today's 17year-olds are classified as "adept" readers (i.e., competent enough to go on to college or to cope with business environments). in the early 1980s, a 12-nation study found that u.s. average comprehensive scores on seven school subjects always were in the lower third. summer 1991 13 since 1973, the poverty rate among americans under 18 has increased more than 50 percent. each year, hundreds of thousands of young people reach working age without the basic knowledge they need to learn even the simple skills necessary for success in an entry-level job. the implications for their use of depository libraries and access to census data are serious: either data access must be simplified to the extent that the unskilled can retrieve needed information or depository librarians will experience a significantly increased workload acting as intermediaries for such patrons. otherwise, the unskilled will become disenfranchised (see below). nearly all economists agree that the industrialized nations are moving toward information-driven, rather than energy-driven, economies in the next century. the intellectual skills of the labor force in such systems will become increasingly important to maintaining a comparative advantage in international trade. skills deficiencies among younger patrons may create a frusuated, information underprivileged class with comparatively little real political power. fortunately, placement of cd-rom devices in schools and the development of improved interfaces may make possible gradations of technological u.ser-friendliness in the next 10-20 years, if either funding or equipment and software donations to schools can be arranged. literacy either in the traditional sense or in a technological sense will decrease in importance. for example, the library of congress reading room reconstruction currently under way will include installation of touch screen terminals. it is hoped that such devices will help non-technically sophisticated patrons access the collections (bonnick & ralyea, 1990). there is some question, however, whether library staff will be skilled enough with the new technology. part of the startup costs associated with cd-rom or even online access as an integral part of the depository library program may be the training of staff so they can assist others in using the systems. disenfranchisement 1990). they ask 1 . is widespread dissemination of public information essential to the preservation of an acceptable level of shared values among citizens? 2. would acccess charges imperil that acceptable level of shared values and create an information underclass? future technology may provide the flexibility to manage funding problems, making possible the continued widespread availability of information. expensive, on-line access is not essential. more economical cd-rom media will be cheaper, and enhanced video systems costing no more than a tv does today will provide cheap, multimedia access (bortnick & relyea, 1990). to the extent that direct broadcast satellite reception and public access cable channels are available, those technologies may serve to democratize information dissemination. (direct broadcast is in use over india today. as minimum antenna diameter shrinks to less than a meter, home satellite reception may become much more widespread than it is today. also, within the decade there may be enough cable capacity that almost any group that wants its own channel can have it for special broadcasts (new american, 1986).) however, if market forces are allowed free reign, history suggests that large businesses and institutions will have disproportionate access to the new technology (and to the information it provides), because they are most able to afford capital and staffing costs and will have the most influence in information systems development. if funding continues to be a problem, younger individuals may be disadvantaged in comparison to older users, companies and universities. user demand what impact will changes in work force composition and employment arrangements have upon the types of census information sought by users? to what degree might user fees and other costs of accessing electronically disseminated information disenfranchise individuals? since the founding of the depository library system, it has been presumed that participatory democracy dictated widespread dissemination of public information. however, the increasing cost of the distribution effort is causing this premise to be reexamined. some information specialists question the societal costs and benefits of universal information suffrage (bortnick & relyea, the aging population the number of americans over age 75 will grow almost 35 percent by the year 2000. the u.s. population 65 or over in 2035 will be 22 percent, vs. 12 percent in 1985 (haub & von cube, 1987). currently, one of three americans is between 27 and 42 (the baby boom generation). the oldest boomers are just 19 years away from reaching the current average age of retirement (longman, 1988). one consequence of changes in age distribution will be an increase in two-generation geriatric famihes diuing the 1990s— adult children in their lassist quarleriy 60s and 70s caring for parents in their 90s (outlook, 1989). the population of retirees will be characterized by a higher average level of education, a friendliness toward business and free enterprise, entrepreneurism, a skepticism toward big government, an inclination to regionalism, and a preference for decentralized regulation (new american, 1986). between 1970 and 1980, life expectancy at 65 increased more than nine percent (life expectancy at birth increased only three percent, thus the fastest growing segment of the population is the age group over 80). both the absolute number and the proportion of the population under 25 are declining. by 1995, people 16 to 24 will be only 16 percent of the population (they were 25 percent in 1980) (longman, 1988). serious attempts to slow the human aging processes and prolong life expectancy may begin within the next decade (outlook, 1989). by the year 2(xx), life expectancy will be 72.9 years for males and 80.5 years for females (new american, 1986). also, the physical design and environment of america (information services included) will start to change in the 1990s to accommodate a middle-aged and older population. (for example, traffic lights will change more slowly, allowing more time for less-agile people to get across intersections). the graying of america may have a significant impact upon the types of information sought by individuals (and institutions operating on their behaloseniors are likely to have a special interest in census data because of its link to medical and social security benefits, its bearing upon invesunent and savings decisions, and its potential as ammununition in lobbying congress. they will have more time to focus upon and use available data, and will constitute a trained, educated clientele able to make drastically increased demands upon the system. changes in the work force by the year 2(xx), 95 percent of all jobs will be in service industries that require workers who are familiar with computers and other information processing technologies. the u.s. may move toward a dual economy, where professionals, scientists, engineers, technicians and other skilled employees are on one end of the spectrum and a large number of blue-collar workers, clerks, and service workers are on the other (lamm, 1985). telecommuting and other fiexible-place, flexible-lime work schedules will become increasingly common as employers acknowledge modem realities such as single parent households and the stress of urban commuting (outlook, 1989). changes in information need, therefore in the use of depository library material, are certain to follow such changes in lifestyle. hobbies, travel, and intellectual pursuits may become higher priority for some, while others seek inforamtion in self-help, quality of life areas. retirement may become a thing of the past, as seniors remain in jobs at all levels in the workplace. many people will return to work after a sabbatical, act as consultants or become "senior apprentices" to learn new skills for a second or even third career (outlook, 1989). multilingual services work force and patron demographics also will drive the development of multiungual information services in the next century's depository program. the population growth of 12 percent anticipated for the next 15 years will result almost entirely from high level immigration and from the higher-than-replacemenl birthrate of new immigrants (mostly hispanic). in the mid-1990s, mexico's proportion of young job-seekers is due to about double, while the kinds of entry-level jobs they seek at home will shrink drastically. the result will be enormous pressures at the border which will not be entirely unwelcome, as u.s. employers struggle to deal with a shortage of american-bom youngsters in the labor force. the user community of the future will include a steadily rising percentage of racial minorities, many using english as a second language. by the year 2000, 1 1 f)ercent of the population will be hispanic and four percent asian; by 2020, hispanics will be 15 percent (new american, 1986). patrons whose first language is not english are likely to need information in their native language, regarding cijturally specific subjects at variance with those sought currently by typical depository library patrons. immigration, employment and family data may be more important to this group than to the population at large. the growth in non-english patrons also may add impetus to the development of non-culturally specific interfaces. even with new storage technology, simultaneous storage of text in asian and hispanic languages as well as english could be prohibitively expensive, as demonstrated by the canadian experience with french and english. however, automatic translation systems being developed as part of the next generation of computers may help solve the multilingual problem. by the year 20()0, computers with automatic language translation and voice-synthesis capabilities may enable people to speak in one language that listeners will hear translated into another language (outlook, 1989). key issues requiring public debate a number of key issues settled out of the literature summer 1991 15 review and interviews conducted for this study. these are issues that either policymakers, the bureau of the census or the depository hbraries need to address before developing technology forces less than optimal solutions. policymakers should future information dissemination be oriented toward individual users or toward businesses and institutions? political and fiscal reality dictates that dissemination decisions weigh business and institutional concerns over those of individuals. these may include predominant formats, time of availability for on-line services, content of information disseminated, and fee structures. however, an effort should be made to ensure that institutions served by depository dissemination are reasonably responsive to individuals' requests, even if fees are charged for services rendered. should joint ventures with private industry be pursued as a means of funding future dissemination in the face of a shrinking federal budget? fiscal reality may force government into partnerships with public database agencies, and to lake advantage of advanced indexing and retrieval software and to defray data generation costs. however, in order to protect the public's right to access government information, some regulation of user fees charged by industry partners may be necessary. what policies should be adopted regarding intellectual property rights in data analysis, access software development and copyright protection? partnerships with industry will require special copyright provisions to protect the rights of government's partners, despite the fact that the products withheld would be generated in part with public funds. contract provisions should state explicitly that key material is not "work for hire", allowing contractors rather than the government to retain copyright. patents should be filed jointly by contractors and the government. these protections will need to extend specifically to indexing and retrieval software and network protocols developed in support of national on-line access systems. government will no longer be able to claim its right to contractor source code and algorithms as a condition of working with private industry. how and where should advanced indexing and retrieval software be procured for access to machine-readable data? if government enters into a partnership with private industry for database management, the vital access software will come from existing high quality products held by commercial firms. otherwise, a major contracting effort will need to be staged to have suitable software developed for use with electronically disseminated federal information. since technology will be evolving rapidly, the decision to develop software under government auspices would be a long term commitment of funds and manpower. what role should be played in coordination of federal information dissemination policy to eliminate fragmentation of jurisdiction over media, content and formats? it is clear that a single, centralized agency is needed to coordinate electronic dissemination of federal information. a quasi-govemment entity including representatives from major actors such as the gpo, congressional agencies and depository libraries could be organized under an institute for information policy & research (as described in hr744, 1985). the tendency to date has been for each agency to select its own standards and media without regard for activities elsewhere in government further, conflicting funding and advocacy situations within congress and the executive branch have frustrated standards development. the bureau of the census should mount a concerted effort to help resolve the jurisdictional confusion that currendy exists. bureau of the census how should demands for multilingual presentation beaddressed? there are two politically acceptable options. either data should be disseminated through the depository library program in engush and spanish, or it should be formatted in a markup language compatible with automatic translation from english to a handful of user native languages. it may be necessary to disseminate in two languages until the technology is commonly available to perform the automatic translation. to what extent is the census bureau liable for ensuring the integrity of data disseminated in machine-readable formats? the census bureau's liability for the integrity of electronically disseminated data is a question for serious legal consideration in a relatively new area of the law. a comprehensive study of legal issues related to data integrity and other aspects of the new technology needs to be conducted. these include tort liability for damages due to the use of inaccurate or corrupted data and product 16 assist ouarteriy liability in case dissemination includes destructive computer viruses or defective media which physically damage depository library equipment. would the census bureau be accountable for invasion of privacy or threats to defense or industry confidentiality that might result from the ability to manipulate data in machine-readable format (the "mosaic" issue)? this issue may either be a moot fxjint, overtaken by extensive networic integration worldwide, or a major impediment to widespread dissemination. the census bureau needs to examine not only its legal uability regarding privacy issues, but also its political defenses against corporate and dod initiatives that are certain to be fielded <end indent both>as the technology makes machme manipulation of public data possible. to what extent should the census bureau be involved in estabhshment of network protocol and human interface standards both within government and within industry? the history of technological innovation has proven that those who are not actively involved in standards setting activities incur both real financial costs and intangible political losses when standards drafted by others are implemented. despite considerable expenditures of manpower and other resources, the bureau of the census must attempt to guide the establishment of protocol and interface standards. how should responsibility and costs be divided for creation and maintenance of on-line access networks? when on-line access becomes feasible, the bureau of the census should implement the supercomputer centers model to create and support the system. businesses and institutions should provide hardware and staff support, while government maintains network interconnects and polices protocol standards. should the census bureau abandon the depository library program in favor of alternative means of data dissemination, or be a driving force in effecting a restructuring of the program in keeping with new information needs and dissemination technology? the existing depository library program may need substantial revision (e.g., removal of select libraries, assessment of user fees, and collection specialization in terms of media and content). if required, this might be accompushed via a cooperative arrangement with federal agencies and other depository libraries. however, given that near future electronic dissemination is likely to be via cd-rom rather than on-une database access, the distribution channels and procedures of the current program may be useful into the next century. if direct access by patrons using integrated networks or through data centers becomes predominant. census bureau participation in the depository library program may be abandoned. depository libraries what types of training should be provided for depository library staff to better enable them to deal with the chajlenges of new technology? member libraries need to train staff in two key aspects of new technology application: (1) how to operate and maintain local hardware and software themselves and (2) how to best instruct and guide patrons in use of the technology. neither task will be easy, since there may be significant differences among librarians in their degree of technological sophistication (and motivation). preparation of a training cadre for the depository library program should be undertaken immediately, funded jointly by the bureau of the census, congress and the gpo. these master instructors in turn should brief library training officers who could provide ongoing familiarization at each program site. to what extent should collection acquisition, operating and other funds be diverted to the purchase of hardware and software to support the use of electronically disseminated information? it may be difficult to convince librarians to spend limited funds on equipment to support user access rather than on collection acquisition. however, in the long term funds spent on electronic equipment will result in more extensive and productive access to existing collections. because of the density of electronically disseminated data, money invested in cd-rom and associated printers and services will quickly expand a library's actual collection even as funds devoted to traditional acquisition decline. expenditures for hardware, software and services related to electronic media should be a significant line item in each faciuty's budget further research should the census bureau decide upon the medium and content of information disseminated based upon extent and type of use research? given the resistance of the library community and the seeming lack of space concerns for projected media, it would seem unwise to attempt limiting the kinds of summer 1991 17 information disseminated through the depository library program. however, it would be prudent to undertake a systematic and ongoing study of usage patterns in case budget or technological constraints make filtering necessary in the future. patron preferences for specific media in accessing certain types of information should be evaluated. conclusion this study was fielded under the presumption that government will be required to continue providing public access to federal information as part of its commitment to maintaining the informed citizenry that is central to participatory democracy. the nature of that access, however, is entwined in a host of social, economic and technology issues that must be addressed promptly if the pace of change is not to overwhelm policymakers as well as information intermediaries and users. information will be central to the knowledge economy of postindustrial america. however, the population will be split between a relatively affluent, educated retirement community and a smaller unskilled, undereducated younger group less able to deal with sophisticated information access. the generation and dissemination of knowledge will be dependent upon the degree of protection for intellectual property in an environment that features easy unauthorized copying of proprietary materials. the technology predominant at depository libraries will be cd-rom, with on-line database services a distant second, used mainly for current updates of timely information. before emerging technology can approach its potential, problems of efficient indexing and retrieval software, hardware compatibility and protocol standards must be resolved. perhaps key in this effort will be a relaxation of anti-trust regulation to facilitate cooperative research, development and manufacture by major players in the telecommuncations and computing industries. ultimately, the twin issues of funding and regulation underlie all concerns regarding future information technology. given the declining resources of federal government, privatization and user fees seem inevitable. privatization brings with it the prospect of reduced access to public information, and user fees the near certainty of disenfranchisement for a new underclass: the information disadvantaged. these problems are not inu-actable. however, significant changes in the way we fund, generate, conux)l and disseminate public information will be forced by technological change. if indeed those who do not learn from the past are condemned to repeat it, those who do not anticipate the future may be destined to live it amidst laments of what might have been. references bell, d. (1978). the postindustrial economy. in j. fowles, handbook of futures research. westport, conn.: greenwood press, 507-514. bezold, c. & olson, r. (1986). the information millenium: alternative futures. washington: information industry assn. bortnick, j. & relyea, h. (1990, march 30). interview at the madison building, library of congress, washington, d.c. caddy, d. (1987). exploring america's future. college station, tex.: texas a&m univ. press. cetron, m. (1988). into the 21st century. the futurist, 22:4, 29-40. clarke, a. (1978). communications in the future. in j. fowles, handbook of futures research. westport, conn.: greenwood press, 637-652. cornish, e. (1985). the hbrary of the future. the futurist, 19:6,2,39. diebold, j. (1985). new challenges for the information age. the futurist, 19:3,68. eldredge, h. (1978). urban futures. in j. fowles, handbook of futures research. westport, conn.: greenwood press, 617-636. federal information center program (1990). [a background paper.] (available from the gsa information resources management service, washington, dc 20405). freese, r. (1988). optical disks become erasable. ieee spectrum, 26:2,41-45. government printing office improvement act of 1990. [hr 3849]. (available from the superintendent of documents, washington, dc). harm is averted by quick response to computer virus (1990). adminisu-ative notes, ii (april 13), 1. newsletter of the federal depository library program. haub, c. & von cube, a. (1987). the united states 18 assist quarterly population data sheet (6th ed.)washington, d.c.: population reference bureau. in marien, m. (ed.), future survey annual (item nr. 8369). washington, d.c.: world future society. hemon, p. & mcclure, c. (1987). federal information policies in the 1980's: conflicts and issues. norwood, nj.: ablex. lamm, r. (1985). megatraumas: america at the year 2000. boston: houghton mifflin. longman, p. (1988). the challenge of an aging society. the futurist, 23:5, 33-37. marshall, c. & rossman, g. (1989). designing qualitative research.newbury park, calif.: sage. mcgee, milton (1990, march 30). interview at the adams building, library of congress, washington, d.c. the new american boom. (1986). kiplinger washington letter. washington, d.c: the kiplinger washington editors, inc. 1988 ten-year forecast. (1988). menlo park, calif.: institute for the future. outlook '90 and beyond. (1989). the futurist, 23:6, 53-60. powell, elizabeth (1990, february 16). interview at the han senate office building, washington, d.c. tenner, e. the revenge of paper. the new york times, march 5, 1988, 27. weiner, e. & brown, a. (1989) human factors: the gap between humans and machines. the futurist, 23:3, 9-11. weinstein, s. & shumate, p. (1989). beyond the telephone: new ways to communicate. the futurist, 23:6, 812. additional sources the following sources were identified in the literature search for this study, but are not referenced in the final report cornish, e. (ed.) (1982). communications tomorrow. bethesda, md.: world future society. didsbury, h. (ed.) (1982). communications and the future. bethesda, md.: world future society. dowlin, k. (1984). the electronic library. new york: neal-schuman. gorman, m. (ed.) (1984). crossroads. (proceedings of the first national conference of the library and information technology assn., sept 17-21, 1983, baltimore, md.). chicago: american library assn. naisbiu, j. (1982). megatrends. new york: warner books. pasqualini, b. (ed.) (1987). dollars and sense: implications of the new online technology for managing the library. chicago: american library assn. ' presented at the lassist 90 conference held in poughkeepsie, n.y. may 30 june 2, 1990. the author may be contacted at 4-290 center for science and technology, syracuse university, syracuse, ny 13244-4100, or at d01dayxx@suvm.acs.syr.edu. 2 hemon, p., mcclure, c, & g. purcell (1985). gpo's depository library program. norwood, nj.: ablex. 'some information intermediaries believe it is important that federal information be made available via the latest technologies (j. stratford, personal communication, february 21, 1990). however, problems with inconsistent file formatting and a lack of satisfactory retrieval software have made tests of experimental census distribution on cd-rom less than a complete success (k. chiang, personal communication, april 12, 1990). problems of information policy making within a web of overlapping agency jurisdictions also have frustrated modernization efforts. "hudson, r. (1990, april 5). ibm researchers "write" with atoms on a metal surface. the wall street journal, p. b4. ' reauthorization of the paperwork reduction act (1989). hearings before the subcommittee on government information and regulation of the committee on governmental affairs united states senate (senate hearing 101-166; document 19-630). washington, d.c: u.s. government printing office. *omb circular a-76 mandated that agencies determine whether it would be more beneficial to continue to perform functions with government employees or to contract them out to the private sector (government, 1990). ferrarotti, f. (1986). five scenarios for the year 2000. new york: greenwood press. summer 1991 19 vol264 4 iassist quarterly winter 2002 iassist quarterly winter 2002 5 editor’s notes again a warm welcome to the iassist quarterly iq vol. 26 issue 4. three articles are presented in this issue. back at the iassist 2002 conference in storrs, ct in the session “the new frontier for archives” kevin schürer from the uk data archive in essex presented the project on “edwardians online”. the paper by emma j. barker and louise corti is presented now in the iq. the qualitative data service at the archive has released a pilot, web-based, multimedia resource. we are far from the rows and columns of numbers normally associated with data archives and the article also gives you a presentation of the “qualidata” unit. the data collection consist of interviews that lasted up to four hours and transcribed to 80 typed pages. it is surprising that the qualitative methods are combined with great numbers. you will find that there in the 1970ʼies were carried out 444 qualitative interviews of this kind! you can learn more about the project in the article, and obviously you can have a look at the web-site. in the next article – also from a data archive anne sophie fink gives us the story about “a danish research portal on the internet”. the article is about the web presence of the danish data archive structured around data production, data archiving and data usage. the archive is viewed not only as a collector and disseminator of data but as an enabler of flows of communication (and data) among the actors. basically the broader perspective should permit users to suggest and add content to the web site. the paper was presented at the storrs iassist conference and was in line with the conference theme of “accelerating access”. anne sophie fink presented the article in the session called “research portals”. from the same 2002-conference and from the same session jessica eustace from the university of manchester presents the paper “reaching your end-user with mimas”. the full session title was called “research portals, or the truth is out there” which calls for agents – intelligent ones! as the portal should be a facility to provide one-stop-shop for users ̓needs in e-journals, bibliographic databases and other research support facilities. the paper highlights how the cross-disciplinary services at mimas can support the creation of a research paper. the paper describes services like the archives hub and copac that brings access to data and merged catalogues, as well as the citation indexes in the web of science and the access to journals through the jstor facility. in the article jessica eustace use the stages of the research paper (introduction, literature review, etc.) to illustrate how researchers make use of the services. an obvious follow-up of these articles that describes facilities on the web for support of researchers will be an evaluation of these sites and portals. i will be looking forward to receiving articles for the iq describing methodologies for the evaluation as well as empirical results of evaluation by stakeholders. access iassist at the web on www.iassistdata.org. papers for the iassist quarterly are most welcome. please contact the editor (kbr@sam.sdu.dk) about submissions. karsten boye rasmussen, august 2003 http://www.iassistdata.org/ mailto:kbr@sam.sdu.dk iassist quarterly fall winter 2012 13 iassist quarterly abstract digital curation3 is currently not very well covered by university curricula in the german speaking countries. nevertheless there is a strong demand for well-educated staff in this field. as part of the project “nestor”, a transnational partnership of academic institutions in germany, switzerland, and austria, a comprehensive qualification program based on e-learning tutorials, schools, seminars, and publications has been established to meet this demand. keywords: training, education, digital preservation, digital curation, nestor. the nestor network nestor, the network of expertise in long-term storage of digital resources in germany is a cooperative initiative of libraries, archives, museums and other parties interested or involved in digital preservation. from 2003 until 2009 nestor was funded by the german ministry of education and research (bmbf). after 2009 it was transformed from a project into a sustainable membership organization, funded by the partners involved. at the moment (autumn 2013), 14 partner institutions are formal members of the nestor network. however, the formal members are only the core of the network. most of the nestor activities are organized and conducted in the open nestor working groups (wg). currently about 60 institutions are involved in several working groups focusing on topics such as media, cooperative long-term preservation, standardization, policy, cost, etc. these groups are constituted and terminated based on need. only the wg qualification4 is based on a more formal cooperation and therefore organized in a different way than the other wgs (see nestor mou group homepage). the nestor wg qualification the nestor wg qualification is acting since 2005 and organized around a memorandum of understanding (mou) first signed in 2007 and renewed in 2009. the mou provides an effective level of commitment but is at the same time less formal than a regular contract. this form was chosen because a formal contract bringing together so many different institutions might have involved too much bureaucracy. the mou has been signed by 12 partner institutions from the educational sector and one coordinating nestor partner (see table 1). the group is involved in education for libraries, archives, museums, and it in the german speaking countries (germany, austria, switzerland). (see table 1 page 15) the wg’s objective is to stimulate and promote qualification in the field of digital preservation. the involved institutions cooperate in the development of course-materials, mutually accept credit points of courses regarding digital preservation and seek to establish a cooperative and distributed master’s degree program in digital curation as a long-term objective. early on, in 2006, the wg’s agenda was shaped by the results of a survey on the coverage of digital preservation topics in bachelorand master’s programs provided and planned by lis departments in germany, austria, and switzerland (see oßwald and scheffel, 2008). the lecturers from each institution stated that none was able to cover the topic in its entirety or to keep pace with the variety of the current developments in the whole field. moreover, subjectrelated needs (e.g. of museums) resulted in a focus on special topics of digital curation in each institution. this initial situation was a good starting point for a cooperation between the different institutions and lecturers. project activities the nestor wg qualification – hereafter called mou group – offers five major lines of activities in digital curation-related qualification to meet existing needs: • nestor seminars • nestor schools • nestor publications • development of e-learning tutorials • development of a cooperative curriculum digital curation training the nestor activities by stefan strathmann1 and achim oßwald2 sub goettingen fh koeln 14 iassist quarterly fall winter 2012 iassist quarterly nestor seminars and schools the nestor seminars on special topics (or for special audiences) take place occasionally. however, they constitute a strongly limited channel for the distribution of knowledge because of the demands they pose on the lecturers’ time and the topics covered, which often address a smaller audience with a higher degree of specialization. the mou group tried to expand the coverage and to relieve the lecturers by producing a video dvd with two recorded introductory seminars for self-study. as this approach was not very succsessful, the group decided that producing some essential publications was a more effective way of distributing knowledge (see below). the series of nestor schools offered by the mou group is a success story. since 2007 seven nestor schools were conducted. for the duration of three to six days attendees and reknowned lecturers come together for an intense exchange of ideas in the privacy of a remote location. the courses are a combination of lectures and practical exercises, augmented with several social activities. the participants, a heterogeneous audience comprised of professionals from libraries, archives, museums, companies, and administration as well as students, are presented a perfect opportunity to build and enlarge their professional networks, to make contacts with colleagues, and to discuss ideas and conceptions. upon completing the course, attendees receive a certificate equivalent to 2 ects credits (european credit transfer and accumulation system). the school table  1:  ins&tu&ons  par&cipa&ng  in  the  nestor  mou  group partner homepage university  of  applied  sciences  –  hochschule   für  technik  und  wirtschac  berlin hep://www-­‐en.htw-­‐berlin.de university  of  applied  sciences  chur,   switzerland,  informa&on  science hep://www.htwchur.ch/en.html cologne  university  of  applied  sciences,   faculty  of  informa&on  and  communica&on   sciences,  ins&tute  of  informa&on  science hep://www.oi.p-­‐koeln.de/en-­‐ index.htm university  of  applied  sciences  –  hochschule   darmstadt hep://www.h-­‐da.de gesis  –  leibniz-­‐ins&tute  for  the  social   hep://www.gesis.org/en georg-­‐august-­‐universität  göxngen,   göxngen  state  and  university  library   (coordina&ng  nestor  partner) hep://www.sub.uni-­‐ goexngen.de/en humboldt-­‐universität  zu  berlin  -­‐  faculty  of   arts  i  -­‐  ins&tute  of  library  and  informa&on   science hep://www.hu-­‐berlin.de/ leipzig  university  of  applied  sciences,   department  of  media  and  communica&on hep://www.htwk-­‐leipzig.de/en/ archives  school  marburg  -­‐  university  of   applied  studies  for  archival  science hep://www.archivschule.de/ university  of  applied  sciences  potsdam,   department  of  informa&on  science hep://www.p-­‐potsdam.de/ stuegart  state  academy  of  art  and  design hep://www.abk-­‐stuegart.de stuegart  media  university hep://www.hdm-­‐stuegart.de/ vienna  university  of  technology,  austria,   ins&tute  of  socware  technology  and   interac&ve  systems,  digital  preserva&on   group hep://www.tuwien.ac.at/en/ tuwien_home/ iassist quarterly fall winter 2012 15 iassist quarterly events have been conducted in cooperation with several digital preservation projects (delos, dpe, digcurv) and with some support from industrial partners (pdf/a competence center, emc, sun microsystems). throughout, the events received very high evaluation marks from the participants. nestor publications and tutorials because of a lack of course materials and to reach a broader audience compared to the nestor seminars, the nestor mou group initiated the production of the nestor handbook (neuroth, et al., 2010). this german-language “encyclopedia” bundles the recent state of knowledge on digital long-term preservation and its various components. it collects contributions from over 50 authors on more than 630 pages and has been revised and expanded several times. beside the open access online edition (version 2.3), a printed book (version 2.0) is available. the handbook tries to cover the whole field of digital preservation and contains chapters on topics like: state-of-the art of legal aspects, preservation policies, the oais (open archival information system) reference model, trusted digital repositories, formats, relevant standards important in digital curation (e.g. premis), strategies of digital preservation, etc. the handbook is very well accepted as a standard work on digital preservation in the german speaking community and it is used by students as well as by practitioners in 2012 a comprehensive baseline study on research data management and curation was published (online and print) in german language (neuroth, et al., 2012). it is the analysis of a structured survey regarding the curation of research data in germany and holds reports from eleven academic disciplines (medicine, astronomy, humanities, and social sciences among others.). a condensed english-language version, entitled “digital curation of research data. experiences of a baseline study in germany” has just been published online and in print (neuroth, et al., 2013). since 2007 digital preservation e-learning modules – based on the e-learning platform moodle – have been developed within student projects. created by students for students, these tutorials are mutually exchanged between the nestor mou partners and used in their university courses. several tutorials, dealing with topics like ‘introduction to digital preservation’; ’the oais model’; ’formats’; ’digital preservation of gis-data’ etc. have been maintained and updated regularly. adhering to the concept “e-learning tutorials by students for students” adopted by the nestor mou group means that the modules can only be developed in combination with university courses. this approach has some implications: the development of e-learning tutorials proceeds very slowly and must be integrated in the regular curriculum. the initial input given by the teachers is extremely high. at the same time the acceptance of using these tutorials is high as well. despite this acceptance it has become apparent, however, that there are issues to secure the quality of the tutorials regarding the content – which needs frequent updating – as well as the didactics, which should respond to various preconditions of students of different levels. therefore the tutorials are expected to become obsolete very soon. at the moment, the group discusses other ways of organizing the teaching material to keep the time and effort needed reasonable. the new concept is not only a big challenge but also a great chance to cooperate and collaborate! towards a shared digital preservation curriculum as already mentioned none of the partners active in the nestor working group is able to set up a preservation-centered curriculum on its own. instead, the partners have agreed to realize this on a cooperative basis. accordingly, as a long-term goal the mou group intends to set up a cooperative curriculum. this plan received support by experts and the administration of several universities in the german speaking countries. in the memorandum of understanding the thirteen partner institutions declared their intention to cooperate in developing building blocks for a digital preservation curriculum and adjusting modules focusing on selected topics. one first step in the direction of a cooperative curriculum is the agreement that credits earned in digital preservation-related courses can be transferred between the universities involved and are therefore accepted in the local curricula. the curriculum development activities led to the engagement of the mou group in the context of the digital curator vocational education europe project (digcurv; http://www.digcur-education. org/). the mou group was a founding partner of the project and contributed in various ways to the development of the digcurv curriculum framework (see digcurv, 2013). the cooperation between nestor and the educational institutions in the mou group results in several synergies. the nestor mou group has stimulated reflections and awareness of digital curation even in those universities where the topic has been part of courses or programs for years. as a consequence results of the nestor network have influenced the content and the quality of lessons and courses provided. not by intention but as a secondary effect the nestor mou group has initiated intraand cross-sectoral cooperation where competition seemed to dominate. on the national level the nestor network in germany has gained insights and created links to universities and qualification bodies. this has improved the awareness and understanding of issues related to qualification for preservation and curation activities. initiatives, promotional programs (e.g. by the german research foundation) and activities now address qualification issues regularly. summing up since 2007 the mou group – comprising of 13 higher education institutions in the german speaking countries germany, austria, and switzerland – has initiated and contributed to the awareness of the importance of qualification issues in the field of digital preservation and curation in the countries involved. members of the mou group and their universities have provided curriculumbased courses, a variety of qualification events like, for example, the nestor school events, and publications such as the nestor handbook, a state of the art publication on issues, methods and best practice solutions in the field of digital preservation and curation. they are still involved in improving the curriculum-based qualification as well as further collaborative education activities in the field. as a side effect of the mou group´s activities, qualification in the field of digital preservation and curation have become a regular topic of related programs and activities in and beyond the nestor context. references digital curator vocational education europe (digcurv), 2010. homepage – digcur. [online] available at: <http://www.digcureducation.org/> [accessed 8 january 2014] 16 iassist quarterly fall winter 2012 iassist quarterly digcurv, 2013. resources/curriculum framework. [online] available at: <http://www.digcur-education.org/eng/resources> [accessed 8 january 2014] neuroth, h., oßwald, a., scheffel, r., strathmann, s., and huth, k. eds., 2010. nestor-handbuch: eine kleine enzyklopädie der digitalen langzeitarchivierung. version 2.3. [online] available at: <http://www. nestor.sub.uni-goettingen.de/handbuch/index.php> [accessed 8 january 2014] neuroth, h., strathmann, s., oßwald, a., scheffel, r., klump, j., and ludwig, j. eds., 2012. langzeitarchivierung von forschungsdaten: eine bestandsaufnahme. [online] available at: <http://www.nestor. sub.uni-goettingen.de/bestandsaufnahme/index.php> [accessed 8 january 2014] neuroth, h., strathmann, s., oßwald, a. and ludwig, j. eds. 2013. digital curation of research data: experiences of a baseline study in germany. [online] available at: <http://nestor.sub.uni-goettingen.de/ bestandsaufnahme/index.php?lang=en> [accessed 8 january 2014] nestor mou group, n.d. homepage. [online] available at: <http://nestor. sub.uni-goettingen.de/education/index.php?lang=en> [accessed 8 january 2014] oßwald, a. and scheffel, r., 2008. lernen und weitergeben ausund weiterbildungsangebote zur langzeitarchivierung. in: neuroth, h., oßwald, a., scheffel, r., strathmann, s., and huth, k. eds., 2008. nestor handbuch eine kleine enzyklopädie der digitalen langzeitarchivierung: version 1.2. [online] available at: <http:// nestor.sub.uni-goettingen.de/handbuch/artikel/nestor_handbuch_ artikel_141.pdf> [accessed 8 january 2014] notes 1. stefan strathmann is head of the competence center on digital preservation and research data at the research and development department of the göttingen state and university library. he is one of the representatives of the nestor mou group on qualification in digital preservation. he can be reached by email: strathmann@sub. uni-goettingen.de. 2. achim oßwald is professor at the institute of information science, cologne university of applied sciences. his focus of research and teaching is it applications in the lis field including digital curation issues. he is one of the representatives of the nestor mou group on qualification in digital preservation. he can be reached by email: achim.osswald@fh-koeln.de 3. in the context of this article we will use the terms ‘digital preservation’ and ‘digital curation’ synonymously. 50 — iassist quarterly database directories by jim jacobs prepared for the 1986 iassist conference, marina del ray, calif., may 22-24, 1986. introduction the following is a highly selective list of directories of [american] 1 machine-readable data files. emphasis is on those directories which are most current, most complete, or are unique in some useful way. this list is current as of may 1986. directories of online databases computer-readable databases: a directory and data sourcebook . edited by martha e. williams. chicago: american library association, 1985. 2 volumes. describes over 2800 publicly available databases, most of which are available online. gives rather detailed descriptions of each database. a subject index lists databases in 550 categories. this directory can also be searched, full-text, on dialog (file 230). volume one covers databases in science, technology, and medicine; volume two covers business, law, the humaniues, and social sciences; multi-disciplinary databases are listed in both volumes. data base directory. 1984-85 . white plains, ny: knowledge industry publications, inc.. 1984, in cooperation with the american society for information science. identifies and describes machine-readable database, both bibliographic and non-bibliographic, which are available for public access online in north america. lists fewer databases (about 1700 versus more than 2700) than directory of online databases or computer readable databases . well indexed by subject, producer and vendor. also available for searching full-text on brs (database label: k.ipd). directory of online databases . quarterly, cumulative. santa monica, ca: cuadra associates, inc. lists and describes machine-readable databases available online to the public. the spring 1985 issue lists 2760 databases, only slightly fewer than computer readable databases , which includes a few which are not available online. published quarterly and available online on westlaw. good subject and other indexes. 'editor's note summer 1987 iassist quarterly — 51 directory of periodicals online: indexed. abstracted, and full text washington, d.c.: federal document retrieval, 1985-86. 3 volumes. this directory is useful if you have the name of a particular periodical and you want to know if it is indexed or abstracted online, or available for full text searching online. will cover 25,000 periodicals when all 3 volumes are published. specialized directories apdu membership directory . princeton, nj: association of public data users. annual. lists and provides profiles of members in this organization of data users, producers, and distributors. all members are organizations. catalog of machine-readable records in the national archives of the united states. washington, dc: national archives and records administration, 1977. describes holdings of machine-readable data in the national archives. arranged by the same record groups as the national archives guide . these files are not available online, but most can be purchased on tape. a new edition is due in 1986. data acquisitions . storrs, ct: roper public opinion research center. 1983(irregular). the roper center is an archive of sample survey data from over seventy countries. data acquisitions lists and describes new surveys in the archive. currently there are over 900 studies available through the center in machine readable form. a newsletter. data set news roper , announces new acquisitions. a directory of computerized data files . national technical information service. washington, d.c.: government printing office, (annual). lists and describes over 1000 federal databases which are for sale from ntis on computer tape. this catalog does not indicate online availability although some files may be available through vendors. indexed by subject and agency. directory of databases in the social and behavioral sciences . vivian s. sessions, ed. new york. ny: sciences associates. 1974. although dated, this directory is unique and is still valuable for identifying organizations that collect data files, and the types of files they collect many entries are for local data centers collecting locally produced data. some examples: "historical data on the social welfare policies in europe", (1850-1965), "polish immigration in the u.s." (1776), "urban transportation study, amarillo texas", (1940). few, if any, of the data files listed here are available online. summer 1987 sex on the racks: issues of data collection and access by daniel c. tsang ' machine-readable data files librarian main library university of california, irvine california in the midst of the aids crisis, researchers seeking accurate empirical data about the prevalence of high-risk sexual behaviors, or the population of homosexuals in the united states, are increasingly frustrated at the lack of reliable and accurate data. instead, researchers are reaching back some forty years, relying on kinsey-era non-generalisable data to estimate the number of homosexuals expected to come down with aids (fay et al, 1989,243). this absence of reliable new data (except for a 1970 national study) since alfred kinsey's landmark studies of male and female human sexuality in the 1940s and 1950s is due to many factors, including long-standing taboos over certain sexual practices as well as political opposition. most recently, this led to congress scuttling, last year, a planned national survey of sexual habits of americans, after conservative congressmen, such as california's william e. dannemeyer (who said the survey was "more apropos for the pages of a pornographic magazine"), mounted a successful campaign to oppose federal funding of a national opinion research center (norc) national sex survey (associated press, 1989; booth, 1989a, 1989b; dannemeyer, 1989; hayden, 1989; "kinsey ii," 1989; peterson, 1989; specter, 1989, 1990). on the other hand, the u.s. census bureau, in its 1990 decennial census, has been able to gather information on domestic partnerships, so that for the first time in u.s. history, lesbian and gay couples who live with each other are being counted in a national census (vobejda, 1990). that this unprecedented exercise in data collection will not be 100% successful is apparent from the concerns some gay activists have voiced about whether they trust the government to maintain confidentiality of the data; the memory of what happened in world war ii, with the release of census data to the military regarding the distribution of japanese americans, and their subsequent internment, remains too real to many americans. despite the national setback regarding a sex survey, the centers for disease control is attempting to gather data at the municipal level; a number of cities have been slated for projects assessing behaviors that are considered high risk for aids, although not without having to overcome local opposition from residents who fear a repeat of the tuskegee experiment, when the public health service deliberately did not treat hundreds of black men for syphilis (boffey, 1987; boodman, 1988a and 1988b). ironically while congress has focused its attention on stopping the national sex survey, it has overlooked the invasion of law enforcement agents into an arena traditionally the domain of social scientists, and allowed the proliferation of sex surveys aimed, not at promoting public health, but at entrapping those suspected of being interested in pornography. hundreds of americans have been sent to prison, in part because they filled out a bogus sex survey commissioned, surreptitiously, by u.s. customs or the u.s. postal inspection service. our sex police have, in fact, mastered desktop publishing, and sent questionnaires to thousands of americans, asking about intimate details of their sex lives, including whether they are interested in sex with children or with animals (see appendix a for an example). respondents — who may well have been indulging in taboo fantasies — are subsequently sold child pornography published by these law enforcement agencies— and after their homes or businesses are raided, sent to jail for 10 years or more for receiving pornography. the questionnaires they have filled out prevent them from claiming entrapment— because their answers on the questionnaires indicate they are predisposed to the crime of receiving pornography (bull, 1987; johansen, 1988; lee, 1987; stanley, 1988, 1989; tsang 1987). it is thus not surprising that, while law enforcement agents can conduct these surveys without any oversight by congress or by human subjects review boards, independent researchers — not connected to the criminal justice establishment— are finding out that certain sexual practices— such as childhood sexuality— are taboo and cannot be researched without law enforcement involvement. in fact, an increasing number of researchers have themselves been arrested, their research confiscated, their careers destroyed, all because they picked a subject too taboo to research (sonenschein, 1987). furthermore, as sexologist john money has suggested, "the only way a researcher can get government funding is to be against sex" (money, 1990). sex data collection, therefore, is a highly politicized iassist quarterly endeavor, especially in the post-meese commission era. it raises important ethical issues— not only about the ethics of breaching confidentiality (as in data on sexual partners)— but also the ethics of hiding the true (law enforcement) purpose of a sex survey (tsang, 1989). politics aside, all sexual behavior studies are faced with two major technical difficulties, viz., bias in the selection of subjects and the reliability of the responses (forman and chilvers, 1989, 1 140). kinsey and his colleagues did not conduct a national sample, but instead relied on participants on selected campuses and in specific organizations. with homosexual acts still criminalized in half the united states and discrimination against persons with aids rampant, respondents may be unlikely to admit to illegal activities or tell the truth (wolpert, 1989). further the specter of big brother asking these questions and the inability to convince everyone of the confidentiality of one's responses makes the reliability of the data particularly questionable. missing data will be a common result (reinisch, 1988). an undercount is also the likely result for the data librarian or the data user, this means that one should approach these data with more than the usual skepticism. it may also mean that researchers are less willing to part with these data. the universe of sharable existing data on human sexual behavior is rather small. first, there is very little tradition among sex researchers (like other researchers) of citing the availability of their data in their research. although the national science foundation has now mandated that grantees deposit or make available their data, this has not caught on in the sex research community. furthermore, much of what passes for empirical sex research is either of the case study variety, or college sophomores' reports of their sex habits. this type of data may be of less interest to other researchers than something more systematically gathered. with the increase in good aids data, however, more interest in sharing data can be expected. if the data is publicly supported, one can well argue that they should not be made inaccessible. privately supported data will be harder to get, of course; it took almost two decades before a kinsey institute study of sexual behavior, conducted by norc in 1970, was finally released to other researchers, in part because of a dispute over which researcher would be listed as the primary author (booth, 1988; reinisch, 1988; klassen, 1988). in the united states, excluding funding agencies such as the federal government, there are two main sources that collect and distribute datasets that have material relating to this topic. the inter-university consortium for political and social research (p.o. box 1248, ann arbor mi 48106; (313) 763-5010) is the major source for much social science empirical data, and thus it would surprise no one that in fact, in many of the datasets archived at the consortium, there are data on sexual attitudes and behaviors. the icpsr collection is now accessible in a number of ways, by looking up a keyword in the annual icpsr guide to resources subject index (on cdnet or printed from tape), or by searching rlin, the research libraries group's cataloging database. the 1989/90 subject index lists 38 studies under the keyword "sexual," two studies under the keyword "sexually," two more under the keyword "homosexuality," and three under the keyword "homosexuals." in addition, 11 studies are listed under "aids." rlin's mdf subfile (for machine-readable data files) is perhaps a better source since each icpsr (and nonicpsr study) in the database is fully analyzed by subject, and one can productively search under the subjects aids, homosexuality, homosexuals, lesbians, sexual attitudes, rape, or sex offenders. the bulk of the sex data files listed in rlin are bibliographic files from the westlaw database concerning civil rights laws; however, dozens of icpsr datasets also show up, most of those concerning sexual attitudes, and not behavior. among recent datasets distributed by icpsr on sexual behavior are the national lesbian health care survey, 1984-1985 (icpsr study 8991) and dangerous sex offenders (icpsr study 8985); norc's annual general social survey, distributed by icpsr, also now includes questions on sexual behavior, and the 1988 general social survey was used in the recent analysis of the 1970 kinsey institute data (fay et al, 1989). icpsr is also the site of the midwest aids biobehavioral research center, an nih-funded project to collect survey questionnaires used in aids research. thus far, some 5,000 questions have been collected, from over 40 surveys, in the hope of providing a database of questions so that some standardization and comparison studies will take place (michael traugott, personal communication, 25 may 1990). a second major source of data is the data archive on adolescent pregnancy and pregnancy prevention (from sociometrics corporation, 170 state st., #260, los altos ca 94022-2812; (415) 949-3832), now available in part on cd-rom. its 1990 catalog of products listed only six studies under the keyword "sexual" and two other studies under keyword "sexuality." however, a subject search of its database, under the topic "sexuality," produced an 88-page printout, listing descriptions and variables for 28 studies. among them: the 1983 cuyahoga county, ohio, familial communication and adolescent sexual behavior project, with 940 variables, including preteens and teens answering questions about homosexuality, masturbation, and oral sex (study alfall/winter 1990 49 a2). another study focuses on sexual behavior of minority teens and preteens (project redirection, study 91-94). in addition, a new national archive on child abuse and neglect at cornell university operated by its family life development center (e200 mvr hall, ithaca ny 14853-4401; (607) 255-7794), but physically located within the facilities of the cornell institute for social and economic research, is a source for a growing number of data files on child sexual abuse. the archive is funded by a grant form the national center on child abuse and neglect. a list of existing or proposed aidsor hiv-related data files appears as appendix b (four pages) in the u.s. office of science and technology's 1988 report, a national effort to model aids epidemiology . the principal investigators are identified, as are the cities or institutions where the data are based. among the datasets named are aids surveillance (centers for disease control), norc's general social survey, and an aids behavioral research clearing house at temple university. with aids already having claimed the lives of more people in the united states than the number of americans killed in the vietnam war, a sense of urgency pervades recent calls for the establishment of a national aids or hiv database. the previously cited report from the office of science and technology policy, which advises the u.s. president, specifically called for the creation of a directory of relevant aids databases, support for access to significant local databases, and enhanced public access to national databanks. researchers are unwilling to release data they themselves are still analyzing, and public-use aids case data is not released by city but rather only in terms of six larger geographical regions in the u.s. (see also layneetal, 1988,511). it also called for the creation and adoption of standards and guidelines for data collection, documentation and release. in an appendix on long-term prospects for data management, the report argued against the concept that sharing of information implies the centralization of data. it elaborated as follows: — data from distinct sources are often not suitable for pooling because they were not collected under similar circumstances. — data from different studies require different protection measures.— quality control is best exercised by the people who have the original responsibility and authority over the data contents and collection.— longitudinal studies require dynamic updating; remote compilations are unlikely to remain consistent.— flexibility to incorporate new data elements, as they appear to be useful, is difficult to achieve in central repositories dealing with many sources.— technology and tradeoffs of systems versus personnel costs are moving computing toward distributed paradigms. the report envisions a decentralized network-based hypertext-formatted information sharing system, best illustrated by this example of an end user sitting in front of a computer at the information mode, one may locate the title of a publication. clicking on the title can produce the abstract, kept at that node. a click on the abstract can cause the text of the paper to be fetched, and the section tides of the paper will be displayed. clicking a section name will obtain the corresponding section of the paper. clicking a graph can produce the underlying values. a numerical result can be clicked on to show the algorithm or program used to obtain the result, from the workstation where the computation was performed. clicking on another marker corresponding to the data will obtain the data, subject to privacy constraints, for display on the screen. clicking a reference cited can continue this browsing process. text so structured may also be annotated for further private of public use (p. 63). although the report did not call for the establishment of a national aids data center, contrary to a story in the chronicle of higher education (turner, 1989), researchers involved in the report did separately call for such a center (turner, 1989; layne et al, 1988). as envisioned, such a center would house a national hiv database, "in its most complete form, a storehouse of raw data on hiv infection and the aids epidemic" (layne et al, 1988, 512; see also hirons et al, 1989). proponents also propose that the database would furnish a standard agreement governing procedures on sharing raw data. they note that the creation of a national hiv database would require an "extraordinary level of commitment on the part of the research community. individual researchers and institutions will have to share and protect large quantities of confidential data on the intimate behaviour of individuals. they will also have to share data that could otherwise be hoarded to build their own careers. but such a database is needed— and it is needed soon" (layneetal, 1988,512). supporters of the idea of a national center argue that the "current lack of a national aids data base center to collect, analyze and distribute the available data is a severe block to our understanding [of aids transmission]." as researchers who use mathematical models to understand the aids epidemic, they believe establishing a center "will encourage closer collaborations between modellers and data collectors" (hyman and stanley, 1987, viii.3). 50 iassist quarterly given the taboo concerning sex research, the united states, not surprisingly, appears to lag behind other countries in the area of data collection and access. the world health organization has been at the forefront of collecting global data on sexual practices, as part of a multinational study of aids (booth, 1989). in addition to geneva, where who is based, some of the data will be archived at essex, england, at the esrc [economic and social research council] data archive, where the coordinator of who's homosexual response studies is situated (apm coxon, personal communication, 27 may, 1990). essex has also been the site of the computerized aids register, which with funding from the medical research council, lists ongoing research on aids. in summary, improved data collection and access to sex data will only occur if sufficient funding, political support, public trust and researcher commitment all materialize. otherwise, sex on the racks will more likely be found in some sex club, and not in a data archive. references associated press, 1989. "sullivan orders changes in sex-survey questionnaire," the orange county register (8 april), a 10. boffey, philip m., 1987. "u.s. to test for aids in 30 cities; household sampling put off," the new york times (3 december), 10. boodman, sandra g., 1988a. "aids study to involve 800 d.c. households: city officials say they were not consulted on federal project," the washington post (28 july), a 1, a 18. boodman, sandra g., 1988b. "federal aids study in d.c. postponed: city officials say household survey would have been unfair," the washington post (29 july),al,a12. booth, william, 1988. "the long, lost survey on sex," science 239:4844 (4 march), 1084-1085. booth, william, 1989a. "asking america about its sex life," science 243:4889 (20 january), 304. booth, william, 1989b. "u.s. probe meets resistance," science 244:4903 (28 april), 419. booth, william, 1989c. "who seeks global data on sexual practices," science 244:4903 (28 april), 418^19. bull, chris, 1987. "feds nab 150 in pom sting," gav community news (8-14 november), 1, 12. dannemeyer, william e., 1989. "proposed 'sex survey'," science. 244:4912 (30 june), 1530. (letter.) fay, robert e. et al, 1989. "prevalence and patterns of samegender sexual contact among men," science 243:4889 (20 january), 338-348. forman, david, and clair chilvers, 1989. "sexual behaviour of young and middle aged men in england and wales," british medical journal 298 (29 april), 1137-1142. hayden, tom, 1989. '"magic bullets' and deadly taboos: what we don't want to know of sexual behavior may kill us," los angeles times (7 june), n, 13. hirons, g. et al, 1989. "an interactive relational database for hiv and the immune system," in r. a. morrisset, ed., ve conference internationale sur le sida: le defi scientifiaue et social: v international conference on aids: the scientific and social challen ge: montreal. quebec. canada. june 4-9. 1989 (ottawa, ont.: international development research centre), 652. (abstract.) hyman, james m. and e. ann stanley, 1987. "using mathematical models to understand the aids epidemic." paper presented at the los alamos center for nonlinear studies conference on nonlinearity in biology and medicine, may 18-22. revised version in mathematical biosciences 90:1-2 (july/august 1988), 415-473. johansen, bruce e., 1988. 'the meese police on porn patrol," the progressive (june), 20-21. "kinsey n," the orange county register (21 march 1989), b6. (editorial.) klassen, albert d., 1988. "'lost' sex survey," science 240:4851 (22 april), 375-376. (letter.) layne, scott p. et al, 1988. "the need for national hiv databases," nature 333 (9 june), 511-512. lee, kevin, 1987. "sex research or law enforcement?" nambla bulletin . " 8 (april/may), 3-4. money, john, 1990. "sex: the good, the bad and the kinky," plavbov (july), 46-49. peterson, larry, 1989. "dannemeyer fights proposed sex study," the orange county register (18 march), b1.b7. reinisch, june machover, 1988. "kinsey sex surveys," science 240:4854 (13 may), 867. (letter.) sonenschein, david, 1987. "on having one's research seized," the journal of sex research 23:3 (august),408414. specter, michael, 1989. "funds for sex survey blocked by house panel: aids researchers say data is essential," the washington post (26 july), a3. specter, michael, 1990. "what's america doing in fall/winter 1990 51 bed? we need a national sex survey to fight aids effectively, so why is congress ducking it?" jm washington post (25 february), bl. stanley, lawrence a., 1988. "the child-pornography myth," plavbov (september), 41-44. stanley, lawrence a., 1989. "the child porn myth," cardozo arts & entertainment law journal 7:2, 295358. tsang, daniel c, 1987. "moral panic in north america: implications for sex professionals." paper presented at the annual conference of the society for the scientific study of sex, western region, 27-29 march, beverly hills, california. tsang, daniel c, 1989. "ethical dilemmas in sex research and therapy." paper presented at the annual conference of the society for the scientific study of sex, western region, marina del rey, california, 22-25 march. turner, judith axler, 1989. "creation of $6-million national center to collect and analyze data on spread of aids is urged," the chronicle of higher education (25 january), a4. u.s. office of science and technology policy, 1988. a national effort to model aids epidemiology . washington, d.c.: office of science and technology policy. vobejda, barbara, 1990. '"unmarried partner' category to provide first census data on gay," the washington esskll march). a6-a7. wolpert, stuart, 1989. "shaking down the aids data: people don't always tell the truth about their sexual habits." ucla magazine (spring 1 ). 13-14. appendix a sample questionnaire from a u.s. postal inspection service sting operation in the 1980s. 1 paper presented at the 16th annual conference of the international association for social science information service and technology (iassist), poughkeepsie, new york, may 30-june 2, 1990. 52 iassist quarterly appendix a sample questionnaire from a u.s. postal inspection service sting operation in the 1980s. 3 s ? ] ! •5 c 2 j ? *£ ill ! ! £ e i^zi i 3v i. tj-d ll ii c i 1 1 -os 5 a * 6 e "5 i if -s hi fffilllilw a. u. 2 u u. i § e * £ i i o 5 i i 1 i i s < <§ > u> * i * i * s _ — c 3 u it!!ll||lil5 f 8 2zut: bi2 if, § ° r ; 5 f ^ b k.i-5. pi if; si1 111 iii: hi ! i -?iil ih ..lit! "8.3 if 8 8 p -s £ m=a8 1 i*ii 5 t -o -5 o i • & e ° "d * 2 1 l x i•nfeti!" if btllg lllllli c b 5 c 2 5 e £"-£"£ = £ £ ill ei ii op 2 8: 6£<e| |sss; &«gg! s o d £ ; t z £io u ||| £ = *" o h> z i i™ ». < w iii < 3 e o °^ 2 o £& sksss .i£ x iog : x e -j 2 z o 2 m i 8 3b fall/winter 1990 53 ^ ri^jbrary^mformation^cieiice^ current research is an international quarterly journal offering a unique current awareness service on research and development work in library and information science, archives, documentation and the information aspects of other fields the journal provides information about a wide range of projects, from expert systems to local user surveys. fla and doctoral theses, post-doctoral and research-staff work are included each entry provides a complete overview of the project, the personnel involved, duration, funding, references, a brief description and a contact name. full name and subject indexes are included other features include a list of student theses and dissertations and a list of funding bodies. each quarter, an area of research is highlighted in a short article current research is available on magnetic tape, as well as hard copy, and can be searched online on file 61 (sf = cr) of dialog subscription: uk £86.00 overseas (excluding n. america) £103.00 n.america us$195.00 write for a free specimen copy to sales department library association publishing 7 ridgmount street london wc1e 7ae tel: 01 636 7543x360 iassist quarterly library & information science abstracts international scope and unrivalled coverage lisa provides english-language abstracts of material in over thirty languages. its serial coverage is unrivalled; 550 titles from 60 countries are regularly included and new titles are frequently added rapidly expanding service which keeps pace with developments lisa is now available monthly to provide a faster-breaking service which keeps the user informed of the rapid changes in this field • extensive range of non-serial works including british library research department reports, conference monographs and development proceedings and • wide subject span from special collections and union catalogues to word processing and videotex, publishing and reprography • full name and subject indexes provided in each issue abstracts are chain-indexed to facilitate highly specific subject searches • available in magnetic tape, conventional hard-copy format, online (dialog file 61) and now on cd-rom twelve monthly issues and annual index subscription: uk £157.00 overseas (excluding n. america) £188.00 n.america us$357.00 write for a free specimen copy to sales department library association publishing 7 ridgmount street london wc1e 7ae tel: 01 636 7543 x 360 fall/winter 1990 55 vol23/2 18 iassist quarterly abstract we are currently seeing a new culture emerging in the social sciences, of a new form of secondary analysis that of primary qualitative data. it has come about largely as a result of the moves by british social science funding organisations towards formalising archiving policies of data created in the course of research they fund. funders want added value from research and believe in sustaining a solid research base for the future, in the form of the preservation of empirical findings. now, this includes qualitative data in addition to quantitative. however, not only is this is a new methodological approach for traditional qualitative researchers it is also challenging the way qualitative researchers view ownership of ‘their’ raw data. new ideas about sharing and providing access to qualitative data are emerging and in the uk, this is being championed by the qualidata resource centre at the university of essex. this paper seeks to address a number of issues. from an archival point of view, how do qualitative data differ from quantitative data? second, what might the implications be for the acquisition, preservation, dissemination and re-use of qualitative data archives for data archives? thirdly, i want to discuss the kinds of procedures required to document and provide access to qualitative data. inherent in this are the special problems relating to confidentiality of some qualitative materials, and i will suggest ways of overcoming these. finally, i want to raise a number of questions relating to how the traditional data archives might want to consider acquiring, storing and disseminating qualitative data. is it in their interest to acquire them? what kind of infrastructure needs to be in place to accomplish this? background to archiving qualitative data in the uk the esrc qualitative data archival resource centre (qualidata) is supported by the economic and social research council (esrc) and is located in the department of sociology at the university of essex. the centre was established in 1994 in order to redress the balance in the bias towards archiving quantitative data from british social science research. it currently has funding up until the end of september 2000. our relationship to the uk data archive is one of a younger sibling. the data archive was set up in 1967 by the economic and social research council (esrc) in order to retain the most significant machine-readable data from the research, which it funds. in order to achieve this, esrc instigated a ‘datasets policy’ whereby all machinereadable data generated from esrc awards should be offered for archiving. there was, however, a significant loophole in this policy. although the advances of word processing now mean that most research of any kind is machine-readable, until recently most machine-readable data was statistical, based on surveys. qualitative research was paper-based. thus the data archive received only a proportion of the raw research data funded by the esrc. as paul thompson, director of qualidata, stated in his 1991 pilot report to the esrc, 'there was no intellectual reason for this'. qualitative and quantitative research are equally based on comparison. classic re-studies include not only rowntree's three surveys of poverty in york, and llewellyn smiths' repeat of booth's poverty survey in london, but also the two successive multi-method community studies of banbury, or, to take an anthropological instance, the controversial restudy and reinterpretation by oscar lewis of redfield's tepotzlan in mexico1. it is not therefore clear why the early social science research council (ssrc) did not feel the need to provide for the archiving of non-machine readable research data. perhaps it was simply felt that there were enough existing archives to ensure that significant material was saved. but in practice, that was certainly not the case. some qualitative material was archived, but usually in special temporary deposits. thus the interviews on which professor george brown’s notable studies of the social origins of depression, are based, were for many years held at his medical research council unit, of which the longterm future remained until very recently uncertain. similarly, the material from paul thompson’s national study of ‘family life and work experience before 1918’, a unique and unrepeatable set of 444 interviews with men text, sound and videotape: the future of qualitative data in the global network by louise corti* summer 1999 19 and women born before 1918, were kept on a short-term basis in a special room at the sociology department at essex, and consequently became the basis of a series of books and articles by visiting scholars, but had no secure future. more generally, little attempt of any kind was made to archive research material. when a small pilot study commissioned by the esrc was carried out in 1991, it was revealed that 90% of qualitative research data was either already lost, or at risk, in researchers’ homes or offices. even with the 10% ‘archived’, it turned out that many of the so-called archives had none of the basic requirements of an archive, such as physical security, public access, reasonable catalogues, or with recorded material, listening facilities. it was estimated that to create a resource on the scale of that at risk would cost at least £20 million. for the older material, moreover, the risk was acute, and the need for action especially urgent. qualidata’s mission qualidata was set up by the esrc with a dual mission. the first was a rescue operation aiming to seek out the most significant material created by research from past years. the second was to work with the esrc and the data archive to ensure that for current and future projects the unnecessary waste of the past does not continue. qualidata is not an archive itself: it is both a clearinghouse and an action unit. its role is to locate and evaluate research data, catalogue it, organise its transfer to suitable archives, and publicise its existence to researchers and encourage re-use of the collections. we maintain a catalogue, qualicat, located on the world wide web, which provides information both about qualitative datasets archived by the centre and those identified by the centre as having already been archived. the catalogue structure follows that of cessda very closely, with some new and modified fields to suit the characteristics of qualitative data. the centre consults with the esrc and other funding bodies on, the now explicit, qualitative aspects of the datasets policy and provides advice to researchers on the implications of archiving for research, both through organised workshops and through individual consultations. the centre also aims to provide a general stimulus to the practice and standards of qualitative research, especially in documenting social science research in britain, as well as encouraging a more active interface between qualitative and quantitative research. how do we define qualitative data? qualidata is concerned with research data arising from the range of social science disciplines, including sociology, social policy, anthropology, social and economic history, political science, social and human geography and social psychology. we define qualitative data as data collected using a qualitative methodology, which contrasts markedly to the traditional quantitative approach. qualitative research is defined by openness and inclusiveness, aiming to capture participants’ lived experiences of the world and the meanings they attach to these experiences from their own perspectives. moreover, a qualitative perspective encompasses a diversity of methods and tools rather than a single one. our definition of qualitative includes in-depth or unstructured interviews, field and observation notes, unstructured diaries, personal documents, photographs and so on, in typed, hand-written, images, audio and video format and either as a digital or non-digital representation. where do we put the data? one of qualidata’s ongoing objectives is the selection of public repositories suitable and willing to receive research material. given that a high proportion of archives used by earlier researchers had proved to be inadequate, a proper evaluation of each potentially suitable archive is essential. a programme of visits to key national archives took place during the first six months of the project, and one of our on-going activities is to liase with new repositories which have special collecting priorities. meeting with traditional archivists raises a number of interesting points about how these professionals view the acquisition and cataloguing of qualitative data collections, and about their relationships with traditional librarians. although we did have a professional archivist on the team at the beginning, essential for gaining credibility with traditional archivists, we are now, primarily a team of social scientists who have adopted a cross-fertilised approach of data archiving and traditional archiving. repositories willing to accept qualitative deposits from qualidata include: • the data archive, university of essex • renowned university archival repositories across britain • british library of political and economic science, london school of economics • the modern records centre, university of warwick • national social policy and social change archive, university of essex • british library (sound archive and manuscripts) • specialist institute libraries • institute of criminology, university of cambridge • contemporary medical archives centre, wellcome institute, london • british universities film and video council, london • national museum archives • imperial war museum, london • labour history archive, manchester • science museum, london 20 iassist quarterly each repository specialises in a number of fields of research. some had not acquired qualitative research data before, but were very keen to begin. furthermore, some have now acquired valuable collections of qualitative social science data and wish to keep acquiring data from us in their particular areas of interest. evaluating qualitative data for archiving qualidata is a small unit: two fulltime and two part-time senior staff; and four part-time processing officers. masses of data are out there, and the suitability of data for archiving is assessed according to a set of criteria developed by qualidata. potential depositors are first invited to submit a sample of data, such as a transcript, to qualidata, together with some documentation about the project. this includes the following requirements for datasets: • of a sufficiently qualitative nature • in good physical condition, e.g. good quality recordings, abbreviations explained etc. • can be made freely accessible for academic use • perceived as having potential for secondary analysis • be able to fit in with existing collections • sufficient documentation to enable informed reuse • copyright, confidentiality and informed consent situation is satisfactory • resources needed to make material available do not outweigh potential for re-use (if the requirements of archiving are taken into consideration from the outset of a project, it is possible to keep extra work to a minimum. for example, esrc applicants are now encouraged to include in their schedule and budget the necessary resources by which to prepare data for archiving) • a suitable repository can be found (although if the materials are considered very high priority then qualidata will house them temporarily). processing the data the centre undertakes processing work necessary both to ensure that data archived conform to legal and ethical guidelines, for example to abide by commitments of confidentiality given to research participants, and to achieve the greatest practicable accessibility and usability for the data. any acquiring organisation will know that some collections of data arrive in a very disorganised state whereas others will be immaculately filed, indexed and labelled. the amount of time and resources required to document material from a previous qualitative study very much depends on how old the material is and much there is. qualidata does accept hand-written material, such as field notes, but where totally illegible, may need to be retyped. this is an expensive process and is only done in the most exceptional circumstances e.g. where the material is felt to be particularly valuable. we also encounter problems with audio-recordings without summaries or transcripts, as transcripts are almost always requested by researchers. in extreme cases, summaries may be carried out by qualidata. digitisation is also sometimes undertaken to give greater accessibility of datasets. preservation of confidentiality and informed consent in qualitative data since the archiving of qualitative data is fairly recent in terms of the history of social science, i would like to outline some of the procedures we have set up for safeguarding the anonymity of informants. the research community has long recognised the importance of respecting the rights of research participants. these rights take two principal forms: the right to have their identity protected (if so desired); and the right to make an informed decision about the uses made of the data that they provide. personal information should be kept confidential, whether or not a pledge of confidentiality has been given to research participants, and should be stored in a secure manner according to the provisions of the uk data protection act (1998). various professional and commercial organisations within the field of social science research have their own ethical guidelines and rules of conduct. whilst some offer more detail with regards to issues like interviewing in difficult circumstances and preservation of anonymity, all present issues regarding the kind of ethical judgements researchers must make when embarking on a research project. the principal for preserving privacy, as articulated for example, in the british sociological association (bsa) statement, is that of the anonymisation of data. however, only one set of guidelines discusses issues relating to the sharing of research data. qualidata has undertaken considerable consultation within the research community, as well as liasing with potential depositors of data, concerning the issues of confidentiality and informed consent. these have undoubtedly been the most frequent causes of concern in the archiving of data. qualidata has a deep concern both for the rights of participants and the professional integrity and peace-ofmind of researchers, and therefore both the issues of confidentiality and informed consent must be addressed in the context of archiving qualitative material. however, in many ways, adhering to guarantees of anonymity is always problematic. the very nature of qualitative data lends itself to descriptions of the interviewees, their lives and their surroundings, and in doing so, presents a dilemma to the researcher in how much to reveal. is it really possible to completely disguise a workplace or a village or the central characters in the drama? i believe that future re-users of a qualitative dataset are presented with similar, if not the summer 1999 21 same, issues as the first authors, concerning respecting the rights of participants. we have produced information sheets relating to the issues of confidentiality and informed consent and confidentiality, consent and copyright in the interviewing of children, both available from the centre upon request or via its www site. these information sheets describe the current legal and ethical situation and suggest solutions by which to respect the rights of participants. of course, qualidata recognises that some datasets cannot be ethically archived, particularly those that address sensitive issues. the options used by qualidata for preserving confidentiality, where appropriate are: • anonymisation of material is just one option available for helping make qualitative data accessible as a future research resource. it can include the removal of identifiers; the use of pseudonyms; and the use of other techniques for disguising the link between individual identifiers and data. it is, of course, important to arrive at an appropriate level of anonymisation to ensure that the data is not distorted to a degree, which devalues their potential for reuse. • a period of closure. where appropriate, a specified period of closure can be applied, although some archives are naturally resistant to accepting material that cannot be used for a long period of time. the saving grace for extremely sensitive materials is that time is of the essence. in 50 years time any tensions should have dissipated, and the information will become history. • restricted access (operated by the archive). access to the data can be restricted to bona fide researchers for genuine research purposes. • restricted access (operated by the depositor). it is possible to make it a condition of deposit whereby all potential secondary researchers must liase with the depositor to discuss their intentions for secondary analysis. the depositor may choose to only give access when satisfied that the data will be used in an appropriate manner in each case. traditional archivists are well used to this approach. • user undertaking not to disseminate any identifying information. most archives operate user undertakings not to breach confidentiality by using identifiable information in published work. this condition is, of course, more effective if used in conjunction with restricting access to bona fide researchers. such a written undertaking does have contractual force in law. furthermore, the good reputation of a secondary user depends upon abiding by these undertakings. • re-contacting participants. it is possible for investigators to go back to research participants to obtain consent for deposit in a public archive, this being something with which qualidata can sometimes assist. this is very time consuming but usually productive. • gaining informed consent in writing for material to be placed in an archive (at the time of fieldwork, but usually after an interview). qualidata has a sample informed consent form, which is also available upon request. this also allows for transfer of copyright depositors have absolute control in setting the terms and conditions for access. an agreement is then set up between the deposit and recipient repository to implement these terms and conditions. secondary users given access to the data must be made aware of such terms and conditions, and should abide by them. in this respect, as data archivists, we place much emphasis on the responsibility of the secondary user. why are qualitative researchers sceptical about sharing and re-using qualitative data? i would like to digress for a moment or two and consider why qualitative researchers show such scepticism towards archiving. this is simply because there has not been an established culture in social science for re-using someone else’s qualitative data. oral historians do use other sources, but this is because they are primarily social historians. to establish why sociologists have not used colleagues’ data, we must first recognise that qualitative researchers are a different breed from the ranks of the quantitative brigade. some, but not all, see the concept of secondary analysis as purely about number crunching, and others feel very threatened by the idea of sharing or making data accountable. there are a number of reasons for this doubt and worry. 1. it is far more interesting to do your own fieldwork, even if it is extremely costly and possibly may be replicating previous studies of similar populations (at the expense of the taxpayer!) 2. generally, qualitative social ‘scientists’ are just not used to making their findings accountable. they are worried about others seeing their data, and possibly picking holes in them. some argue that certain approaches used in qualitative research, for example, grounded theory (glaser and strauss 19672 ) which opposes the scientific paradigm of testing hypotheses, do not lend themselves to verification. 3. many researchers we have spoken to feel very strongly that, through fieldwork, they have established a special bond with their interviewees. many also have promised informed consent at the time of interview 22 iassist quarterly which precludes the use of the participants’ contributions for anything other than their own eyes or, at least, the current piece of research. 4. some researchers are concerned that their material cannot be used sensibly without the accumulated background knowledge which they have acquired during its collection. this is particularly so with longitudinal studies of a group where the researcher feels that a special rapport has been developed without which the material may be meaningless. thus the essential contextual experience of ‘being there’ cannot be shared. we believe there is a solution to each of the negative points raised above: 1. to gain a more informed approach and to stop the proliferation of repetitive work, new studies should make more attempts to delve into earlier related research and try to include some comparative element. in order to be able to accomplish this, a firm bedding of archives across the uk needs to be cultivated on a regular basis and nurtured thereafter. 2. if we are to accept the label ‘scientist’, then we should adopt the scientific model of accountability, reliability and validity. the quality of social research is highly variable, and in the uk there are no quality control standards for qualitative studies (except for market research 3 ). we believe it is bad practice for raw data not to be available for future scholars and, furthermore, detrimental to the progress of history. as far as i am aware, it is unheard of for a social science journal to cite access to the original source of data, as is necessary in most natural scientific journals. 3. interestingly enough, the complete protection of anonymity that researchers sometimes offer their participants is untenable a first publication which a journalist then seizes upon may undermine this promise with a misguided stroke of a pen. in essence it is impossible to promise total anonymity. in contrast, we have found that when recontacting participants to gain permission for archiving, the majority seem to be in favour, even though this wasn’t mentioned at the time of fieldwork. our experiences tell us that, providing their contribution is not abused, for example, their identifying characteristics are not cited (if they choose them not be), they are happy for serious scholars of the future to look at the raw materials. most people do believe that research is for the public good, and that their contribution will be used in some way to create a better informed society, and even go some way towards implementing policy changes. contractual archival policies mean that investigators must now either rethink negotiations about informed consent and be prepared to discuss with their participants, at some stage, access to data beyond their own team. 4. the ‘being there counts’ argument is understandable but also an easy opt out of being prepared to share data. indeed, there are instances where research data are, in a sense ‘re-used’, by the investigator themselves. for example, some principal investigators who write the final articles resulting from a project have employed research staff or a field force to collect the data. similarly, for those working in research teams, sharing one’s own experiences of the research is essential. both rely on the fieldworkers and co-workers documenting detailed notes about the project and communicating them to each other. of course, audio and videotape recordings enhance the capacity to re-use data without having actually been there. for archives, documentation of the research process provides some degree of the context, and whilst it cannot compete with being there, field notes, letters and memos documenting the research can serve to help aid the original fieldwork experience. what about the format of data? we deal with all formats. much qualitative data nowadays is digital in the sense that the text is word-processed or hand-written material is scanned, or audio-visual material is in digitally recorded form. qualidata has developed standards for the documentation of qualitative digital data in liaison with the uk data archive. generally materials are reduced to their simplest form, ascii, tiff4, but the data archive also accept rich text format (rtf) and adobe’s portable document format (pdf). we put digital data alongside paper-based materials in repositories or, where possible, offer it to the data archive at essex. data from mixed methods studies are usually offered first to the data archive, for example, so those indepth interview transcripts sit alongside the statistical dataset. the data archive are experienced in handling, storing and disseminating textual data, and presently have the advantage over some traditional repositories in being able to keep up with changing media and storage technologies. however, for acquisition by the data archive, textual data must be, as far as possible, anonymous. preservation of confidentiality is addressed below. far more qualitative researchers are now using digital data. the last three years have seen a huge growth in the use of computer-assisted qualitative data analysis software (caqdas) packages in qualitative research. caqdas software, such as nudist and atlas-ti, is rapidly becoming the accepted tool for handling the description and interpretation of qualitative data. for qualidata issues about preservation of data from these packages is something we summer 1999 23 have had to address with some urgency. these are proprietary software packages and in the past it has not been possible to import and export data from one package to another. qualidata has developed guidelines on what to keep for archival purposes i.e. reducing the data to its simplest form ascii text or rtf. as expected, in the past year we have seen software developers taking steps to encourage sharing between packages, for example adding export and import facilities to their programmes, and even beginning to build xml export features. digitisationwhere are we going and what are we keeping? the data archive in the uk archives primarily numerical, textual data: documentation for datasets is now stored in image format mostly in the form of pdf files; and more recently they have also begun to acquire image based datasets. qualidata is currently working on a large-scale digitisation project. this is the preservation of professor george brown’s life’s collection of research data. the major focus of the work is on the role of psychosocial factors in the onset, course and chronicity of, and recovery from, clinical depression (a major public health problem).4 the distinctive feature of george brown’s approach has always been the ability to combine both qualitative and quantitative aspects of the same data. the publications resulting from brown’s team reflect this duality in combining a host of statistical tables with a wealth of case history material. thus the surveys above are all coded and the statistical data for each project will be archived with the data archive here at essex. qualidata is image scanning the paper schedules, many of which contain a great deal of annotation in hand written form. the original tiff4 files and a final pdf file for each case (patient) will still archived. pdf has been chosen by the data archive as a suitable archival format, as have many other institutions in britain. however, we can never be sure whether this format may become extinct, and at the very least we would hope if it did, that conversion to the new formats would be an option. perhaps we can allow ourselves to relax just a little, as we move into a climate of technological sharing and interoperability. but when do we throw away the paper? a number of options come to mind, in no particular order: • when our physical storage space is full up • when we are confident we have a permanent representation • when the paper starts degrading in my own experience, thinking back to the forty four filing cabinets worth of george brown’s data, i am terrified of getting rid of any of them! they are going to be available in electronic book form, and as safe as they could be in a prestigious data archive, but what if….? to avoid this panic and to appease our sense of sentimentality, our current strategy is to keep samples of original data, so for each project we will select about ten cases and these will be placed with a suitable academic (paper based) repository. if scholars still want to set eyes upon the original documents, they can! can traditional archives cope with digital nonnumerical data? well, in short, some can and some can’t! some of our host repositories have the facilities to provide copies of, say transcripts on disk, whereas others just can’t provide that service. this is usually simply a case of under resourcing. it is not uncommon in the traditional british archive world to see one, or at best two, archivists responsible for sorting, cataloguing, housing, and providing access to archives. this leaves little time for digitisation programmes and resources may not stretch to obtaining high-powered computers for storage. reviews of electronic documents in personal papers and organised records held by archival repositories in britain highlight problems of staffing, software, hardware, expertise and dissemination. the other side of the picture, and of course, an ironic one, is the increasing lack of physical storage space for paperbased archives. many archives are full up with paper documentation, and those with inadequate storage facilities are using hot or damp basements for storage. microfilming and digitising saves on storage space, but does not necessarily represent a cheaper option: filming and scanning are expensive operations and the maintenance of electronic records in the long-term involves periodic transfers of data to new media and software. technological changes and the ever-reducing cost of computer storage will undoubtedly mean that digitisation becomes a more attractive option over time, not least because it allows the records themselves to be disseminated electronically. with the dawning of the age of the digital library, and closer relationships being forged by academic libraries and archives with it departments, and new centrally funded programmes, i don’t imagine archivists will turn away machine-readable versions of transcripts for much longer. problem areas for archiving qualitative data video recording and other image (such as photos), and to a lesser extent audio data, all present added difficulties for archiving and it is preferable that participants play a key role in the decision to archive. audio-tape recordings tape recordings of interviews are almost always used in qualitative studies. these may be individual interviews, focus groups, observation and naturally occurring conversation. for some projects, full transcription is 24 iassist quarterly essential, for others summaries may suffice. methods of transcription also vary: sociologists generally want to capture the words, whereas linguistics are more concerned with recording other contextual features of the interview, such as pauses, laughter, tears etc. in terms of re-use potential of data, the ideal is to retain the original tape recordings. there is really no substitute for listening to people’s own words; a transcription is a subjective interpretation of the real-life conversation. in reality, it is often not possible to archive audiotapes where the material is ‘sensitive’, without either restricted access, a period of closure and/or retrospective permission from participants. anonymising tape recordings in the same way as for the transcripts is vastly time-consuming and prohibitively costly. blanking out of identifying information on analogue media is also rather pointless as it distorts the data. perhaps digital audio data may be less problematic. new software is now available for which researchers can edit, anonymise, label and copy their own data with far more ease. again, this is still labour intensive and in the uk there is still no concensus about what the best audio format is for archival purposes. current popular options are minidisc, r-dat and cd-r, but there is still no consensus on the relative longevity of these media. the even more problematic case of video-data everything discussed with reference to audio data is worse for video data, with the added complexity of faces. we have not yet been able to archive much interview video data, as researchers have been very anxious about the possibility of identifying participants. there is no way around seeking permission to archive video data, and we are advising that permission is sought either before or after interview, depending on the sensitivity of the research and context of the interview setting. however, it is still evident, at least in the uk, that only a few branches of social science have taken on board the use of video methods: social anthropologists; socio-linguists and discourse analysts and educationalists. so, should the traditional data archives acquire and store audio, video and multi-media data? as technology moves forward many data archives across the world will have to begin considering the storage of digitised and indexed data from audio, video and multimedia data. we will see great improvements in storage options and indexing facilities for audio and video data. dvd is an exciting but volatile format and surely will replace audio and video cd. since all windows operating systems will be supporting it, it looks likely to dominate the market. whilst it is still very expensive, inevitably costs will drop. but i would like to pose the question: is it in the interest of the traditional social science data archives to take this route? for example, accepting and storing digitised audio and video of qualitative data creates serious issues regarding confidentiality and access, and also indexing. whilst most data archives do not accept photos, audio or video-tapes, there are other specialist archives in britain set up to receive and deal with these formats of data (although not with a social science remit). these have established standards and have dedicated working groups e.g. the digital archiving working group run by the bl, pro and jisc, and research libraries group. we are seeing guidelines emerging for the preservation on each and every kind of media. new types of data clearly require specialist staff for evaluation, processing and documentation. the reason that the uk data archive is able to acquire textual and image qualitative material is that qualidata acts as the front-line, engaging in evaluation, processing and documentation of these data. thus the staff time and expertise to deal with qualitative data are not required of the data archive’s own personnel, who are busy enough with their own specialist roles. with this infrastructure in place, the data archive can provide access to a greater range of social science data. an alternative model might be for the social science data archives’ to act in the role of brokers, where storage and access of social science data in say, audio and video formats, can be negotiated through data archives established systems, but not necessarily either processed or stored there. there are now smaller embryonic "qualidatas" growing across europe. however they are typically run by academics based in sociology departments, and usually have no links with their own country’s data archive community. i am helping to build a network of these centres and hope that the data archive community will begin to take on board the contemporary and historical significance of qualitative data. to do this we all need to communicate and debate the issues i have addressed in this paper. 1. paul thompson, ‘report to the esrc on ‘the archiving of qualitative interviews: a pilot survey’, november 1991. 2. glaser, b.g. and strauss, a.l. (1967), ‘the discovery of grounded theory: strategies for qualitative research’, chicago: aldine. 3. bs 7911 is the trademark for the standard adopted by the market research society in 1988 for ‘specification for organizations conducting market research’. this came about partly as a result of the hugely varying quality of qualitative studies in this arena. summer 1999 25 4. the archive will include twelve collections, based on distinct projects dating from 1969 to the present. the earliest and probably best known study to many social scientists and clinicians is the camberwell study, conducted from 1969-75 and providing the basis for the eminent book, ‘social origins of depression’, by brown and harris. the team pioneered the life events and difficulties schedule (leds), a survey instrument used to record stressful experiences and significant life events. * paper presented at:international association for social science information service & technology, building bridges, breaking barriers: the future of data in the global network, toronto, may, 1999. louise corti, deputy director and manager, qualidata, department of sociology, university of essex, colchester co4 3sq, uk. e-mail: cortl@essex.ac.uk tel: +44 1206 873058 url: www/essex.ac.uk/qualidata/ mailto:cortl@essex.ac.uk http://www/essex.ac.uk/qualidata/ 19.3 21fall 1995 “but it’s really an information ocean, not a highway. if you think of it as an ocean, then you have to consider the kind of tools that are used, who builds the boats, who designs them, and whether youære surfing or diving. if you have a message in the bottle, how do you get the bottle to the people who need it?” peter gabriel, new york times, july 13, 1994 the internet provides a global infrastructure for data and information publishing that has the potential to revolutionize how data and information are accessed and used. while a variety of methods exist for taking advantage of the capabilities of the internet for this purpose, two problems with data and information access have been exacerbated by an explosion of tools and resource servers. first, different tools and approaches each have individual advantages, leading to the use of different methods at different locations. second, resources are distributed across many locations and may often be difficult to locate. ciesin’s information system approach is designed to solve both of these problems by providing a method for locating and integrating heterogeneous, distributed data and information resource servers. the implement of technologies such as those utilized by ciesin could potentially transform the organization of how data and information are provided and provide users and providers alike with new services and capabilities. the technical core of the internet is a set of protocols and standards for computer to computer communication. initially, internet technology supported three primary high level services: access to remote systems (telnet), exchange of files (ftp), and electronic mail (smtp). over time more sophisticated standards and protocols have evolved that offer other services as well. this technology allows the creation of client-servers systems capable of greatly enhancing the dissemination of data and information resources. data and information providers have a variety of types of resource servers to select from, each with its own functionality (see table 1). server functionality world wide web (www) browsing and retrieval of formatted text, images, sounds, movies hypermedia documentbase combined in hyperlinked documents distributed across servers gopher documentbase browsing and retrieval of text and/or graphics files in hierarchical lists distributed across servers wais index full text searching and retrieval of textual documents and/or graphical files database querying of and access to data structured by fields and records application data processing and analysis of data sets utilizing the capabilities of applications such as arc/info and sas ftp archive retrieval of text and/or binary files from a hierarchical list newsgroups/mailing list multi-party, free-form, asynchronous textual discussion table 1. internet server functionality locating and accessing data and information on the internet: methods and organizational impacts by christopher davis1 ciesin 22 iassist quarterly from the user perspective, each of these servers can be accessed at least by a client application that matches the server. some clients, though, provide access to multiple server types (see table 2). www gopher wais database &2 ftp archive newsgroup3 application www x x x x 4 x x browser gopher x x x client wais x client database x5 front end ftp client x news reader x table 2. internet client functionality because of the wide range of functionality offered by www browsers and servers, these systems are more widely used than any other approach for both access and distribution. while www browsers offer access to a range of servers, the www architecture severely limits design flexibility in information systems. current www standards such as html (hypertext mark-up language) and http (hypertext transfer protocol) limit user interface design to what can be accomplished using a basic form interface. this precludes the use of features such as menu bars and dialog boxes and other dynamic windows. this problem is compounded by the fact that http is a stateless protocol. the www browser requests a document; the document is provided by the server, and the connection closes. the server has no way of knowing or tracking what document the user requested once the request is filled, thus global settings and variables are difficult to maintain from one document to another. also, the www browser is a display tool only. the browser has no capabilities for data processing, so all data processing must be performed at the server. this is potentially problematic in two instances. first, the server might become overloaded processing multiple jobs, which could be more easily handled by the client. second, when an image such as a chart is created, the data sent to the www browser is a graphic file which will be a much larger file than the actual data used to generate the image. if the client uses data from the server to create an image, the amount of network traffic would be significantly reduced, leading to improved performance. a final problem is that while www browsers support some types of servers without modifications to the server, database and wais servers require development work on the server to allow access. multiple options are available and documented for wais servers, and several examples exist of providing access to oracle and other relational databases, but access to database and custom servers may involve significant server modification and development work. in some instances, it may not be feasible to develop an interface from www because the server requires a stated protocol. unfortunately, www browsers alone are not a universal solution for internet server access. however, the advantages of utilizing www browsers and servers in information system design should not be minimized. from the development perspective, www browsers exist for all major operating systems, and the development of clients can be anticipated to continue as new operating systems emerge. this removes a major cost of information system development and support. from the user perspective, www browsers provide access to a range of services and are thus more likely to be installed and used on a regular basis than custom clients for a particular information system. this creates a major incentive for resource providers to utilize www servers since they are immediately available to the existing and massive install base of www browser users. 23fall 1995 the development of applications such as www have facilitated the distribution of data and information resources over the internet. while this has lead to an explosion of available resources, a specific resource may be often difficult to locate. a frequent quote overheard on the internet is: "everything you need to know is on the internet. you just can't find it."6 by design, the internet is a cooperative, unmanaged venture. this is one reason why the internet has been so successful, but it is also the source of the grand challenge of locating resources. several efforts have been undertaken to chart the internet or at least provide mechanisms for facilitating the location of data and information resources. one listing of these efforts includes eighteen different searchable catalogs of internet resources.7 some efforts focus on cataloging specific types of resources. for example, the council of european social science data archives (cessda) is developing an interface to the distributed collections of european and other social science data archives. at present, this involves a global map of resources with pointers to individual archive resources.8 the u.s. federal government has created the government information locator service (gils)9 to facilitate the cataloging government resources using a standard system. the g-7 countries have agreed to prototype the gils standards for non-u.s. resources as well, but the gils project is still very much in the early phases of development. another effort is being undertaken by the consortium for international earth science information network (ciesin). ciesin has developed the ciesin gateway as a system for searching distributed metadata collections and providing access to heterogeneous resource servers.10 while these projects are still in development or early phases of implementation, the existing results show that the technology exists for solving the problem of locating resources on the internet. one of the approaches for locating data and information resources that has been implemented is the ciesin gateway. the ciesin gateway was initially developed to meet the requirements of ciesinæs mission as the sedac (socioeconomic data and applications center) in nasa's earth observing system data and information system (eosdis). one of sedac's charges is to serve as a two-way gateway between social science and physical science researchers studying global change. the ciesin gateway fulfills this function by providing a single interface that allows searching of multiple data archives. the ciesin gateway includes a single interface that allows searching of the eosdis ims, nasa global change master directory, and related directories of data, and also allows searching of key social science and other related metadata collections. organizationally, ciesin accomplishes this task through its information cooperative program. the information cooperative provides an institutional umbrella for linking together data centers worldwide. the ciesin gateway provides the technical implementation for the information cooperative by allowing searching of the directories of each of those data centers. f ig ur e 1. c ie s in g at ew ay c on ce pt ua l s ch em at ic 24 iassist quarterly the philosophy of the information cooperative is to encourage each partner data center to maintain their own metadata directories and data archives. this requires that the ciesin gateway be a distributed system, capable of searching multiple systems in parallel. also, since each data center is likely to have different existing information systems, the ciesin gateway must support access to a variety of commercial databases and other internet servers. in addition to searching and displaying high level metadata, the ciesin gateway also provides the capability of accessing on-line resources identified by the metadata. in some instances this is done within the ciesin gateway client, but in others it is accomplished by spawning an external application such as a www browser. a conceptual view of the capabilities provided by the ciesin gateway are described in figure 1. thus, the ciesin gateway combines the capabilities of a search system for locating resources with capabilities for accessing those resources regardless of server type. based on the lessons learned from the development of the original ciesin gateway system, ciesin, in collaboration with brooklyn’s polytechnic university, is working on a new project, raven. raven subsumes the ciesin gateway functionality within a larger framework of access services. raven presents the user with a list of services, such as ciesin gateway, www browser, and interfaces to custom applications and databases (see figure 2). f ig ur e 2. f irs t s cr ee n fr om r av en p ro to ty p e f ig ur e 3. s ea rc h s cr ee n fr om r av en p ro to ty pe 25fall 1995 this approach provides more functionality within one client, and it also enhances the interoperability between separate services. the raven design and development is based on object oriented design techniques that allow new services to be easily built using common components from other services. at the most basic level this includes the networking component of the services, but it also includes capabilities for viewing tables and entering queries (see figures 3-4). raven also uses cross platform development tools to facilitate the development of versions for multiple operating systems. the goal of the raven project is to create a client that provides a single interface to a vast array of resources and systems and provides information system developers with a set of tools for easily creating a user interface appropriate for a particular system. technology such as www and the ciesin gateway provide the technical capabilities that allow organizations to take advantage of the infrastructure of the internet to distribute data and information resources. these capabilities can have several impacts on the organizations that utilize these technologies. first, the internet infrastructure allows for the development of new products and services that allow the creation of digital libraries. digital libraries include network accessible collections of data and information resources. because these resources are available over the internet, the physical location of the user and the library itself becomes irrelevant. also, digital, on-line storage of resources allows enhanced tools for searching, accessing, and analyzing resources. thus, data and information resource providers can utilize internet technologies to expand their user base and the services they provide. a more subtle but significant organizational impact of these technologies affects the very organization of data centers. an example of this is the structure of ciesin's information cooperative. to the user, the information cooperative appears as a single archive, but in fact, it is a collection of archives linked through the internet and the ciesin gateway. this organizational approach is a form of adhocracy, a term first coined by the futurist alvin toffler in his book future shock11. an adhocracy is based on small, specialized organizations that use information networks to coordinate their activities, and thus act like a larger organization. this approach can be more efficient and responsive than traditional hierarchical structures. this change is not limited to data and information resource providers. similar changes are being experienced in a variety of industries, both public and private. as the global economy becomes increasingly information based and dependent, this pattern will increase in frequency. as with other types of organizations, two primary organizational roles emerge: providers and brokers12. providers develop and disseminate resources and provide support on the use and understanding of those resources. brokers work with providers to develop catalogs of resources across providers and to f ig ur e 4. s ea rc h r es ul ts s cr ee n fr om r av en p ro to ty pe 26 iassist quarterly develop information systems to offer a common access system for users to the resources of multiple providers (see figure 5). figure 5. relationship between providers, brokers, and users ciesin, through its information cooperative program, represents one early effort at this approach. the cessda effort is another example. as the technical infrastructure for these types of organizational relationships becomes more widely implemented, other similar efforts will also emerge. the internet and related technologies have already had a significant impact on the practices of distributing and accessing data and information resources. consider this closing thought from christopher locke (1994), "unlike any previous medium, the net's speed and reach seem to enable reaction to events that have not yet taken place. but this is an illusion. we are not seeing into the future, but more deeply into the present."13 this capability will force data and information resource providers to continue to take advantage of new methods both of technology and organization to enhance and improve speedy and effective access and use of data and information. 1 paper presented at iassist95 may 1995 quebec city, quebec, canada. 2 the methods of user access for database and application servers are functionally equivalent. 3 mailing lists are accessed through electronic mail and are functionally separate from any of the other client/server access / distribution methods. 4 the database interface is limited by the forms capability of the html standard and the existence of server side scripts for translating form input into a format understood by the database server and converting the results from the database and converting it to html. 5 a database front-end might be a separate client that runs on a users machine, or it might be a service that is accessed via telnet over the internet. generally, the front end is a custom interface to an off-the-shelf or custom server, usable only with a particular information system. 6 from david lubar's "it's not a bug, it's a feature": computer wit and wisdom, addison-wesley publishing, reading, ma, 1995 who cites the source as "anonymous, but common knowledge to anyone who's been there." 7 http://cuiwww.unige.ch/meta-index.html 27fall 1995 8 http://www.uib.no/nsd/diverse/untenland.htm 9 http://www.usgs.gov/gils 10 http://www.ciesin.org/gateway/gw-home.html 11 random house, new york, 1970. 12 an excellent summary discussion on this is "electronic markets and electronic hierarchies," by thomas w. malone, joanne yates, and robert i. benjamin in computer-supported cooperative work: a book of readings, edited by irene grief, morgan kaufman publishers, san mateo, ca, 1988. additional references are available from the author of this paper. 13 from david lubar's "it's not a bug, it's a feature": computer wit and wisdom, addison-wesley publishing, reading, ma, 1995. 1/2 rasmussen, karsten boye (2022) editor’s notes: we talk data. we do data., iassist quarterly 46(3), pp. 1-2. doi: https://doi.org/10.29173/iq1065 we talk data. we do data. welcome to the third issue of iassist quarterly for the year 2022 iq vol. 46(3). in denmark we sometimes retrieve an old quote from a member of the danish parliament: 'if those are the facts, then i deny the facts'. we have laughed at that for more than a hundred years, but now fact denial is apparently the new normal in many places. and we are not amused. data can become dangerous as facts can be fabricated. therefore, a critical approach to data is fundamental to producing reliable information: facts. the articles in this issue are about teaching students good data behavior, and how researchers with great care and attention can carry out the task of fact production. the first article is about improvement in teaching data: 'investigating teaching practices in quantitative and computational social sciences: a case study' by rebecca greer and renata g. curty. the authors are both at the university of california, santa barbara library, where rebecca greer is director of teaching & learning and renata curty is social science research facilitator. they are investigating data education and present some of the findings from a local report part of a national project into how instructors adapt curricula and pedagogy to advance undergraduates computational and statistical knowledge in the social sciences. the core goal of the instructors concerns 'data thinking' the critical understanding and evaluation of data. many students have a preconceived fear of mathematics that influences other areas. personally, i feel that data thinking is essential to live and participation in society, and i believe that it should be achievable even with a background of math fear. however, for social science students i also expect they have acquired some level of 'data doing'. i agree with the authors that the necessary support for data is more often found in the areas of science, technology, engineering and mathematics than it is in social sciences. however, many iassist members successfully work to relate data to social science students. and the implicit relationship via data to stem areas will furthermore often improve job success for social science students. the local study interviewed instructors and the article presents among other things the learning goals and the explicit skills contained in these goals. the study uses many quotations from the interviewees, including quotes on sharing among the instructors. this leads to how the instructors can be further supported and how the library can support them, including a partnership between the library's research data services and teaching & learning. with the second article we continue at a university. now the focus shifts from teaching to research the other main area of university work, and more specifically the data in research. the article 'research data integrity: a cornerstone of rigorous and reproducible research' is by patricia b. condon, julie f. simpson and maria e. emanuel. all three are in positions at the university of new hampshire, durham, usa. the article starts with the foundation of the four rs of research: rigor, reproducibility, replication, and reuse. the interest in data integrity came from a question at a graduate seminar on the difference between data integrity and data quality. when exploring the data quality component, they found that research data integrity is closely associated with data management as well as with data security. the aims of the article are several, but the first is to establish practical explanations of research data integrity and its components. training and documentation are fundamental and form the surroundings in the proposed research data integrity model that also graphically presents the overlapping areas between the components: data quality, data management, and data security. i find this focus on the sharing between components a structurally clear approach, and with good outcome too. when juggling concepts that often are regarded as being more or less identical, it is clearly positive to make these relationships and distinctions. this positive structural approach is continued as the authors relate research data integrity to the research data lifecycle to produce an implementation schema. the last section is relating research data integrity to the four rs. https://doi.org/10.29173/iq1065 2/2 rasmussen, karsten boye (2022) editor’s notes: we talk data. we do data., iassist quarterly 46(3), pp. 1-2. doi: https://doi.org/10.29173/iq1065 submissions of papers for the iassist quarterly are always very welcome. we welcome input from iassist conferences or other conferences and workshops, from local presentations or papers especially written for the iq. when you are preparing such a presentation, give a thought to turning your one-time presentation into a lasting contribution. doing that after the event also gives you the opportunity of improving your work after feedback. we encourage you to login or create an author profile at https://www.iassistquarterly.com (our open journal system application). we permit authors to have 'deep links' into the iq as well as deposition of the paper in your local repository. chairing a conference session or workshop with the purpose of aggregating and integrating papers for a special issue iq is also much appreciated as the information reaches many more people than the limited number of session participants and will be readily available on the iassist quarterly website at https://www.iassistquarterly.com. authors are very welcome to take a look at the instructions and layout: https://www.iassistquarterly.com/index.php/iassist/about/submissions authors can also contact me directly via e-mail: kbr@sam.sdu.dk. should you be interested in compiling a special issue for the iq as guest editor(s) i will also be delighted to hear from you. karsten boye rasmussen november 2022 https://doi.org/10.29173/iq1065 https://www.iassistquarterly.com/ https://www.iassistquarterly.com/ https://www.iassistquarterly.com/index.php/iassist/about/submissions mailto:kbr@sam.sdu.dk 6 iassist quarterly fall & winter 2007 katherine mcneill* interoperability between institutional and data repositories: a pilot project at mit abstract academic libraries are working in new areas to support the publishing activities of their institution’s faculty members, including helping them to manage and archive research data that they produce. many institutions, such as the massachusetts institute of technology, have multiple locations in which faculty can deposit their data. yet this distributed arrangement presents challenges for searching, unifying collections, and archiving. in order to foster some interoperability between these multiple data repositories, the mit libraries developed a prototype system to bring studies between two such systems, dspace and the institute for quantitative social science dataverse network, by enabling the harvesting and replication of metadata and content across the two systems. this paper will discuss the motivation for this project, details and challenges of the system, and future goals for enhancing interoperability among the two systems. literature review many academic library systems, such as the one at the massachusetts institute of technology (mit), have been developing more services in recent years to support the publishing activities of their faculty. developing institutional repositories (irs) for housing and disseminating the digital research materials produced by an institution is a main area of work. academic librarians play a key role in promoting and facilitating the use of irs (bailey 2005). these activities provide new opportunities for librarians to become partners in publishing with their faculty, which can enrich their relationships and increase the library’s relevance (buehler and boateng 2005; bell, foster, and gibbons 2005). however, many irs are experiencing low rates of faculty contribution (mcdowell 2007). in order to enhance participation, many librarians are working to evaluate the utility of their ir from their faculty’s perspective. some institutions have undertaken projects to study faculty work practices in order to design the repository system which best meets faculty needs. one such project discovered that faculty members must be able to personalize their presence in the ir in order for it to provide them with significant value (foster and gibbons 2005). a recent study indicates that datasets comprise only a very small percentage of items in irs (mcdowell 2007). in this context, many librarians assist faculty members in publishing their datasets, whether it is in their ir, a domainspecific data repository, or another location. for example, purdue university library has established the distributed data curation center (d2c2) to support the curation and archiving of faculty-produced data.1 success in this work requires an understanding of the needs of individual faculty members in order to recommend to them an appropriate system for managing and archiving their data (witt and carlson 2007). moreover, a viable data archiving system must be of tangible benefit to the depositor, not just the secondary data user. one study argues that a requirement for citation of datasets by secondary users would be the best incentive for faculty to prepare their data appropriately for deposit in a data archive (niu 2006). for several years, members of the social science data community have been promoting the need for standards for citing data. some have developed specific standards recommendations designed to interoperate with data repository systems (altman and king 2007). all these studies shed light on how to design data repositories in alignment with the needs of faculty and researchers. a range of different kinds of digital repositories exists: “individual, discipline-based, institutional, consortial, and national” (peters 2002). given this landscape, there often are multiple locations where an individual faculty member can publish and archive data, each of which may have its own approach to and policies regarding archiving and management. these variations in service may make one repository more appealing to a faculty member, and thus implicates the choices she must make (and thus the availability of her research data). how might two kinds of repositories, irs and domain-specific data repositories, come together? green and gutmann envision a collaborative system whereby the ir facilitates communication and exchange of data between the researcher and the domain repository (green and gutmann 2007). the ability for different repositories to exchange metadata and content would provide an important service to enable faculty data to be housed and discovered in more than one system. iassist quarterly fall & winter 2007 7 working examples of systems that can exchange metadata and content among data and/or institutional repositories exist in the field. one such project, bibapp, developed by the university of wisconsin-madison and university of illinois urbana-champaign, has designed software to index, search, and harvest from the web (for one’s local ir) publications of university faculty.2 other repositories3 have the capability to harvest metadata and content from other repositories compliant with the open archives initiative protocol for metadata harvesting (oai-pmh).4 other funded projects aim to develop models for transporting metadata and digital objects among repositories of different structures (fcla digital archives). the harvard-mit data center (hmdc), a member of the institute for quantitative social science at harvard university, has significant experience in developing systems that enable the exchange of metadata and content among data repositories. hmdc has developed and deployed two successive systems of open-source data repository software, the virtual data center and dataverse network software systems (altman et al. 2001 and king 2007). hmdc has utilized these systems to harvest metadata and content from partner data archives, such as icpsr and the roper center for public opinion research. continuing these projects, hmdc now operates the shared catalog for the data preservation alliance for the social sciences (data-pass) (altman et al. 2009). supporting mit faculty as data producers how do the mit libraries support their faculty members so that they can archive and publish their research data? within the mit libraries social science data services program,5 the data services librarian encourages faculty to archive and disseminate data that they produce and helps them to do so. as discussed earlier, the ability to support faculty in their publishing efforts provides new opportunities to be of utility to faculty and share with them the libraries’ expertise in this area (buehler and boateng 2005). mit faculty members in the social sciences have three main options for where they can deposit data that they have produced. dspace, mit’s institutional repository is a dublin-core-based ir system utilizing software developed jointly by the mit libraries and hewlett packard laboratories (smith 2002).6 dspace is committed to preserving not only data sets but also any mit-produced material (e.g., working papers, images, etc.). the harvardmit data center (hmdc) provides mit with its own customized data repository at the institute for quantitative social science (iqss) dataverse network.7 this data documentation initiative (ddi)8 -compliant system is based on the dataverse network software (dvn) developed at harvard (king 2007). mit can load into the iqss dvn any data that mit licenses, purchases or produces. lastly, mit faculty also can deposit their data in the inter-university consortium for political and social research (icpsr) data archive.9 each system has its own advantages and challenges. dspace is more likely to have an established workflow for loading items in a faculty member’s department, yet lacks specific features for working with data. iqss dvn has tools for online data manipulation and enables a high-level of control by the faculty member. moreover, faculty members can create personalized home pages to highlight their data. this feature has been a selling point for some faculty, supporting research documenting the need for personalization (foster and gibbons 2005). icpsr is a formal, full-service archive with staff that can guide depositors in preparing their data for archiving and distribution and can perform additional services such as a confidentiality review and documentation enhancement. different mit faculty members in departments such as economics, history, and political science, have deposited items in each of these repositories, depending upon their individual needs and preferences. in working with these systems and facilitating faculty deposit, the mit data services librarian has been involved in many tasks associated with managing irs (bailey 2005). the data services librarian begins a consulting arrangement with a one-on-one meeting with the faculty member to first to understand her data management and archiving needs (witt and carlson 2007). the data services librarian highlights the benefits of archiving personal research data; discusses with the faculty member which if any of the aforementioned repositories will suit her data; answer questions; and coaches her to a decision as to if, and where, to archive her data. to support this activity, the data services librarian worked with two other mit librarians to develop a web site for faculty on data management and publishing.10 experience has shown however, that in-person meetings (rather than the simple existence of informative web pages) are necessary to give faculty the information and incentive needed to start such a project. while there are benefits to having multiple options for archiving faculty-produced data, this situation also can lead to certain challenges. end users have to search multiple systems and there is no unified collection for a given faculty member. most importantly, it is a challenge to help faculty decide where to put their data. mit has an interest in--and thereby a desire to promote--all three aforementioned systems, yet one cannot expect faculty to deposit in more than one. each system has its strengths, yet it would be beneficial if faculty-produced data could be housed and discovered in all of them. for some time, the mit libraries have been considering how enabling some level of interoperability among these systems could help ease these problems. 8 iassist quarterly fall & winter 2007 pledge project in 2005, the mit libraries had an opportunity to work on a project to develop limited interoperability, in the form of metadata and content exchange, between two of these three systems for data deposit: dspace and iqss dvn. the pledge (policy enforcement in data grid environment) project, a partnership with the san diego supercomputer center and university of north carolina, chapel hill, explored the use of data grid technology (i.e., distributed data storage infrastructure) for replication of content across systems for preservation purposes.11 as part of this grant, mit worked on a specific project to exchange metadata and content between dspace and iqss dvn. in order to address the challenges posed by housing data in multiple locations, mit took on this project with the goal to develop a mechanism to archive, preserve, and provide access in dspace to mit-authored studies in dvn. such a system would allow dspace (with its mission to preserve mit-produced material) to archive the studies and enable users to find the studies from within dspace, while still being able to access the studies separately via dvn, which features unique services tailored to manipulating data. as a result, this system demonstrates how dspace can archive mit content while at the same time allow mit-produced research to be discovered and housed in other specialized systems. in line with the goals of pledge, this project enhances the preservation of these data files through the replication of content. it is important to understand that the interoperability achieved in this project is limited to the exchange of metadata and content, and does not extend to other possible services such as integrated access or shared interfaces. within the past couple of years, mit faculty members have begun to load data that they produce into hmdc. the first mit group to do so was the abdul latif jameel poverty action lab (j-pal), a research lab associated with the department of economics.12 by putting their data in dvn, j-pal could take advantage of all of its data-specific features. but its data was neither housed nor discoverable in dspace, the system for preserving mit-produced research. this set of data was the test case for the pledge project, which was designed to bring the mit-authored studies housed in dvn, such as those from j-pal, into dspace. note: overall, mitauthored studies comprise only a subset of studies that mit stores in hmdc, which also houses materials licensed or purchased (but not produced) by mit. dspace and dvn were selected for the project because they are both home-grown and have mit involvement. the main mit staff member working on the project was the dspace system manager, who previously was a developer of software at hmdc (and thus had a key intersection of skills for the project); the data services librarian provided information and advice on the project. how the system works the dspace system manager designed an agent (working with manual involvement by both the system manager and the data services librarian) to harvest and replicate metadata and content across dvn and dspace. in order to accomplish these tasks, the agent converted metadata and information packages between different formats used by the two systems. therefore, the agent was designed to transform ddi metadata (used by dvn) into mets (metadata encoding and transmission standard), a submission package standard that allows for the exchange of both content and metadata.13 the mets package includes specific items from the ddi record, including mods (metadata object description schema) descriptive metadata14 and premis (preservation metadata) technical metadata.15 this new set of metadata then needs to be converted into a form that can be understood and processed by dspace. dspace ingests digital objects in submission information packages (sips), based on the reference model for an open archival information system (oais) (consultative committee for space data systems 2002). next, the agent produces a stand-alone, self-describing zip package that then is ingested into dspace, creating an item (and associated catalog record) in the repository. the system works in four main steps (see figure 1). iassist quarterly fall & winter 2007 9 1. the system manager sends the url for a particular dvn study to the agent. 2. the agent then, utilizing a given url for a study, harvests a ddi record and all the appropriate study content (i.e., data and related files such as a codebook) from dvn, via oai-pmh, the protocol for metadata harvesting.16 3. the agent packages the content into a sip zip file containing: • mets file (including mods descriptive metadata and premis technical metadata) • ddi file • content file(s) (data file and others associated with a given study, e.g., codebook or other documentation) 4. the agent then sends the sip to the dspace ddi ingest packager. this tool processes the package to produce a dspace item described in a dublin core metadata record. it takes all other items in the package (ddi, data and other files) and attaches them as files associated with the dspace item. figures 2 and 3 are screenshots of a sample data file that has been brought through this system. figure 2 shows the catalog record for one of the data files that j-pal had submitted to dvn. the study record in dvn has some ddi-specific metadata fields: e.g., geographic coverage, unit of analysis, etc. in that system, online subsetting and statistics feature are also available. figure 3 shows the item in dspace after it was processed by the agent. the metadata record is simpler because it is described in the dublin core format. the ddi metadata from dvn now is available as an additional file associated with the item in dspace. this particular study is included in the datasets collection within the j-pal community in dspace. in the future, j-pal could load into dspace other content types that they produce, such as papers, images, etc. (see figures 2 & 3) discussion despite its benefits, the system is not without its challenges. one is the inefficiency of the workflow for selecting and processing studies. since only the mit-authored data files in dvn are being harvested for dspace, a librarian needs to make a manual selection decision to start the process. for now the process begins with the data services librarian notifying the dspace system manager which studies are mit-authored, and then the latter executes the agent. this is a functional, but not particularly efficient or scalable, system. however, with further effort, one might be able to design a more automated system. this system could take the shape of a form in dspace in which one enters the url for a study at dvn, initiating a set of automated processes that execute the workflow fully without further manual intervention. alternatively, the dvn software is designed so that an administrator can create a custom oai set based on metadata criteria such as author or dvn collection. an oai client could be developed that would monitor this set and feed new or updated studies into the dspace ingest system, creating a certain level of synchronization of dvn to dspace. the interaction between these two systems also has implications for licensing. normal dspace loading workflow includes some licensing screens. the mit submission process includes a click through agreement makes mit distribution policy and contributor responsibilities explicit; the author grants the mit libraries the right to distribute his/her content. in the prototype system at mit, an agent loads the studies in the back end without the author viewing those screens (in this test case, agreement to the terms was obtained through an informal email exchange with the author). moreover, dvn has specific pop-up windows whereby secondary data users must agree to text included in the terms of use element in the ddi metadata before the system allows them to download the data file. in dspace, these terms of use are simply described within the ddi file attached to the record of the item. the libraries are considering the implications of these issues. updating is one of the challenges of replication of content. given the fact that data and related files have copies at multiple locations, currently no process exists to update all copies of files simultaneously. for example, if a faculty member corrects and error in a dvn submission, updating a file, there is no automated method of tracking and providing notification of this change to trigger an update of the file in dspace, accordingly. currently, libraries’ staff would need to somehow learn about this event and manually update the file in dspace. investigation may yield a more automated alternative in the future, such as the synchronization system described earlier. in addition, it should be noted that the agent was designed to work with the systems as they were configured at a particular point in time. as each system gains new features, moves to a different platform or infrastructure, or is revised to be compliant with a new version of its metadata standard, someone will need to re-program the agent to keep it up-todate. future work the next step in this project, after the success of the prototype, is to formalize the service. currently the dspace system manager is the only one who knows how to run the agent. the mit libraries therefore now are working to document the system and integrate the service into the normal workflows of dspace staff members in charge of the local repository. in addition, the libraries hope in the future to design more automation around the use of 10 iassist quarterly fall & winter 2007 figure 2 figure 3 iassist quarterly fall & winter 2007 11 the system. however, given the number of mit social scientists known to produce data, the libraries expect to receive a relatively low volume of studies to be processed by this system each year, so manual execution certainly is feasible in the short-term. thinking farther into the future, a major improvement to the system would be to enable data exchange in the other direction, i.e., bringing studies housed in dspace into iqss dvn. many faculty members choose to house their data in dspace given existing workflows in their departments for loading other materials. therefore, allowing access in dvn to studies from dspace would further enhance services and might enable the files originally only housed in dspace to utilize the data-specific features in dvn. however, this is a more complex challenge because of the different metadata specifications used by the systems: dublin core in dspace and ddi in dvn. taking records based on simpler metadata (dublin core) into a system with more complex metadata (ddi) poses a difficult challenge. in one scenario, dvn could simply harvest the dublin core metadata records and map them to a limited set of ddi elements. in this model, studies from dspace could be found via dvn searches; dvn then could direct the user to dspace to download the study. however, more robust data-specific searches would require this metadata to be extended to further elements of the ddi, requiring additional cataloging. moreover, to import data files from dspace into dvn, the format of those data files would determine whether or not they would be compatible with the online analysis features of dvn. more complete interoperability, such as integrated access or shared interfaces, could be explored in the future as well. in addition, it would be a worthwhile effort to explore if this service could be expanded to exchange metadata and content with other data repositories, such as icpsr. this would further extend discovery and utility of mit facultyproduced data and work towards a system of partnerships between local and domain-specific repositories (green and gutmann 2007). conclusion despite these challenges, pledge project success can be of benefit to other institutions. while this system was built to address a particular local need, it can inform other projects to share and replicate content across repositories. the prototype demonstrates the use of packaging standards and strategies for delivery of data to exchange metadata and content between two systems. in addition, this project has lessons for other sets of information. the agent designed is not dvn-specific but can work with any ddi-based system and thus could be used by others. moreover, the project illustrates the ability to harvest both metadata and content across systems based on different metadata conventions, i.e., was not limited to those in ddi format. many researchers now are working with diverse groups of content files and corresponding metadata culled from the web, utilizing the open archives initiative object reuse and exchange (oai-ore) standard (open archives initiative a). this system demonstrates how to bring these complex digital objects into a repository. in conclusion, this project devised a prototype system that accomplishes several goals. it enhances the discovery and preservation of mit-created data files in these two systems used and maintained locally. in addition, it demonstrates how to package and import ddi and related data into a greater variety of systems, holding the promise for more interoperability among systems in the future. acknowledgments i would like to thank mark diggory and sean thomas for their work on developing and sustaining this system, as well as for their input into the content for this article and its associated conference presentations. michele kimpton contributed helpful knowledge regarding interoperability projects in other dspace installations. i also would like to thank micah altman, gretchen gano, and ann green, who read the draft paper and provided helpful comments. * katherine mcneill, data services and economics librarian, massachusetts institute of technology. e-mail: mcneillh@mit.edu references altman, micah, margaret adams, jonathan crabtree, darrell donakowski, marc maynard, amy pienta, and copeland young. 2009. digital preservation through archival collaboration: the data preservation alliance for the social sciences. the american archivist (forthcoming, 72 (1) (spring/summer 2009)). altman, micah, leonid andreev, mark diggory, gary king, akio sone, sidney verba, daniel l. kiskis, and michael krot. 2001. a digital library for the dissemination and replication of quantitative social science research: the virtual data center. social science computer review 19 (4): 458-70. altman, micah, and gary king. 2007. a proposed standard for the scholarly citation of quantitative data. d-lib magazine 13 (3/4) (march/april 2007), http://www.dlib. org/dlib/march07/altman/03altman.html. bailey, jr, charles w. 2005. the role of reference librarians in institutional repositories. reference services review 33 (3): 259-67. bell, suzanne, nancy fried foster, and susan gibbons. 2005. reference librarians and the success of institutional repositories. reference services review 33 (3): 283-90. buehler, marianne a., and adwoa boateng. 2005. the 12 iassist quarterly fall & winter 2007 evolving impact of institutional repositories on reference librarians. reference services review 33 (3): 291-300. consultative committee for space data systems. 2002. open archival information system (oais). [cited september 16 2008]. available from http://public.ccsds.org/ publications/archive/650x0b1.pdf. fcla digital archives. archives for: september 2008. [cited september 19 2008]. available from http://blogs.fcla. edu/index.php/digitalarchive/2008/09. foster, nancy fried, and susan gibbons. 2005. understanding faculty to improve content recruitment for institutional repositories. d-lib magazine 11 (1) (january 2005), http://www.dlib.org/dlib/january05/foster/01foster. html. green, ann g., and myron p. gutmann. 2007. building partnerships among social science researchers, institutionbased repositories and domain specific data archives. oclc systems & services 23 (1): 35-53. king, gary. 2007. an introduction to the dataverse network as an infrastructure for data sharing. sociological methods & research 36 (2): 173-99. mcdowell, cat s. 2007. evaluating institutional repository deployment in american academe since early 2005. d-lib magazine 13 (9/10) (september/october 2007), http:// www.dlib.org/dlib/september07/mcdowell/09mcdowell. html. niu, jinfang. 2006. reward and punishment mechanism for research data sharing. iassist quarterly 30 (4) (winter 2006): 11-5. open archives initiative. a. oai-ore. [cited september 16 2008]. available from http://www.openarchives.org/ore. peters, thomas a. 2002. digital repositories: individual, discipline-based, institutional, consortial, or national? the journal of academic librarianship, 28 (6): 414-7. smith, mackenzie. 2002. dspace: an institutional repository from the mit libraries and hewlett packard laboratories. in proceedings of the sixth european conference on research and advanced technology for digital libraries (ecdl'02), lncs 2458, rome, italy, ed. agosti, m. and thanos, c., 543-549. berlin: springer. witt, michael, and jake r. carlson. 2007. conducting a data interview (conference poster). washington dc, usa. [cited september 16 2008]. available from http://www.dcc. ac.uk/events/dcc-2007/posters/data_interview.pdf. footnotes 1. http://d2c2.lib.purdue.edu. 2. http://code.google.com/p/bibapp. 3. http://www.policyarchive.org. 4. http://www.openarchives.org/oai/openarchivesprotocol. html. 5. http://libraries.mit.edu/guides/subjects/data. 6. http://dspace.mit.edu. 7. http://dvn.iq.harvard.edu/dvn/dv/mit. 8. http://www.ddialliance.org. 9. http://www.icpsr.umich.edu. 10. http://libraries.mit.edu/guides/subjects/datamanagement. 11. http://pledge.mit.edu. 12. http://www.povertyactionlab.org. 13. http://www.loc.gov/standards/mets. 14. http://www.loc.gov/standards/mods. 15. http://www.loc.gov/standards/premis. 16. http://www.openarchives.org/oai/ openarchivesprotocol.html. vol272.indd r vol25.4 12 iassist quarterly winter 2001 iassist quarterly winter 2001 13 by tom smith & michael forstrom* preface on the morning of september 11th america had the tragic misfortune of being the target of the largest terrorist attack in history. four hijacked airlines were crashed into the twin towers of the world trade center, the pentagon, and a field in rural pennsylvania. in the span of a few minutes more than 3,000 people were killed. upon learning about these terrorist attacks, i recalled the study that the national opinion research center (norc) had carried out in 1963 in the days following president kennedyʼs assassination. for the first time in the intervening four decades, i felt that the kennedy assassination study (kas) should be replicated. i consulted with my norc colleague, ken rasinski, and we decided to launch the national tragedy study (nts). in the course of less than two days, we secured norcʼs commitment to conduct the nts; raised funds from the national science foundation, the robert wood johnson foundation, and the russell sage foundation; and settled on the sample design and questionnaire content. by thursday september 13th we were conducting interviews. the nts drew heavily from kas, allowing comparisons of these two great national tragedies, and from norcʼs general social surveys (gsss), providing for more proximate preand post-september 11th comparisons. hunting for the kennedy assassination study (kas) data to select items from kas to be included in nts, we turned to norcʼs librarian, patty cloud and norcʼs archivist, michael forstrom. they immediately pulled the extensive hardcopy files on srs-350, as kas was formally known.1 these files had been prepared by norcʼs long-time librarian (1961-1997), patrick bova. they consisted of 1) codebooks that contained copies of the original questionnaire, coding/data processing instructions and deck/column location information, marginals, and certain select data tabulations; 2) drafts and final versions of papers and articles based on kas; 3) data processing memos, especially about creating work decks for cross tabulating data stored on different decks; 4) correspondence about the study after collection; and 5) miscellaneous related material. with this complete documentation in hand the nts content was quickly finalized. in praise of data archives: finding and recovering the 1963 kennedy assassination study with the nts launched, we now turned to finding the machine-readable, case-level records for kas. cloud and forstrom searched the access master list of norc tapes and found no reference to srs-350, kennedy, or any other term that connected a tape with kas. tom w. smith, co-pi of the gss and nts, recalled that the gss had in the mid-1970s collected copies of various norc and nonnorc studies as part of the social change project (scp). he searched the gss archives and found a tape and related printouts from 1975 with data from kas. with assistance from the fay booker, data librarian from social science research computing at the university of chicago, the data were recovered from this old tape. unfortunately, as was typical when scp acquired data from earlier norc studies, it only contained a sub-set, consisting only of most demographics, the 10 affect-balance scale items, and a scattering of miscellaneous items.2 none of the assassinationspecific items had been extracted and saved. smith also checked with the major data archives where norc typically archived data the roper center, university of connecticut, and the interuniversity consortium for political and social research, university of michigan, and confirmed that norc had not sent them the data. he also followed up a lead in an early norc bibliography (allswang and bova, 1964) that the international data library at the survey research center, university of california berkeley had the kas codebook and questionnaire. they too did not have the data. so the search turned to finding the punched cards, the original machine-readable medium for the data. cloud and forstrom consulted records of both onand off-site records storage. there are about 4,000 cubic feet of records stored at norcʼs headquarters and upwards of 20,000 cubic feet at a remote site, oʼhare records and retention center.3 based both on the storage records and frequent past access to the on-site stored material, there was no indication that any punched cards were stored at norcʼs headquarters.4 cloud and forstrom also did not find any punched cards indicated by the access database listing of 3,669 boxes in the off-site records storage. based on this cloud therefore 12 iassist quarterly winter 2001 iassist quarterly winter 2001 13 reported on september 14th that “the likelihood of our finding cards is virtually nil.” smith then contacted bova who indicated the next day that the cards should still be in storage. so forstrom next consulted hard-copy inventories of the material of bulk storage.5 these inventories from 1997-1999 identified 37 skids of boxes (about 15-20 boxes per skid) as well as bays of filing cabinets. the inventories included no references to kas, but the contents of some skids and cabinets were either cursory or unidentified, so one couldnʼt rule out the punched cards being there. on september 17th bova then supplemented his earlier comments that the cards should be in remote storage with a listing of storage records, including barcode numbers for the kas material from norcʼs previous off-site storage facility, before the holdings were transferred to the oʼhare site. this documentation encouraged us to search on, but unfortunately these earlier control numbers did not match anything in the access database. by september 24th it was decided that if the punched cards were in off-site storage, they had to be in the unidentified, bulk-storage material. from september 26th to october 11th, cloud and forstrom spent 4-5 days at the oʼhare facility going through 37 skids of boxes and the collection of around 50 filing cabinets. no kas material was found. at this point, cloud and forstrom wondered if the kas punched cards could be in boxes in records storage, but not listed among the contents of boxes, if they had been discarded, or if they had been among a set of records lost during the a move to the oʼhare facility. finding the cards under any of these situations seemed unlikely, so smith pursued another line of inquiry. from the hard-copy kas files the names of two individuals who had been sent data in the early 1970s were obtained and from a norc bibliography (bova and worley, 1991) the names of another two non-norc researchers who had published an article using the data in the 1960s were obtained. the two authors were tracked down and indicated that one had had the data, but abandoned them many years earlier. the two data orderers were being traced when this line was then abandoned after a breakthrough. on october 20th, forstrom received from accounting a question about a late fee attached to an oʼhare records storage invoice. this inquiry led forstrom to examine the invoices from oʼhare more closely. they listed 8,348 boxes in general, records storage, but the access database documented only 3,669 boxes. thus, there were 4,679 undocumented boxes in norcʼs general, records-storage holdings at oʼhare. forstrom contacted norcʼs former records manager, connie schumacher, and she indicated that there was a drawer in the norc libraryʼs old card catalog that listed at least some of the oʼhare holdings and that bova had a printed index to punched cards. cloud also found a 1987 memo that referred to holdings of punched cards for about “60 studies in 700 boxes.” these records allowed the old, record-control numbers to be matched to barcodes and locations at oʼhare. oʼhare was contacted on october 24-25th and asked to pull 11 boxes designated as kas punched cards.6 on october 26th forstrom picked up ten boxes from oʼhare, but found that they had pulled one incorrect box. the final box was then retrieved by forstrom on october 29. after six weeks of effort we had the kas data. dealing with multiple-punched cards we now had to convert the cards to a current, machinereadable medium. there were three challenges. first, we had to know how to interpret the data. second, we had to read the cards. third, we had to convert the multiplepunched data7 to single-punched data. the first task was made possible by the detailed codebook that bova had prepared in 1964. we knew what data were in each column and that there were 3 decks per case. next, we had to find the right cards. smith examined the 11 boxes and determined that they consisted of a single copy of cards in deck 2, multiple copies of cards in decks 1 and 3, and several boxes of work decks used in the early analysis. thus, we quickly knew what to read and how to interpret the data. the second and third task depended on finding someone who could read cards and translate or spread multiplepunched data. norc, like most organizations, had given up using cards about 25 years ago and no longer had a card reader. we searched for an organization that could read cards and convert multiple-punches. after considering several possibilities, we settled on november 9th on the national data conversion institute (ndci) in new york city, a firm that norc was already using for old, seven and nine-track tape conversions. since there was only a unique copy of deck 2, we were very concerned about shipping the cards to ndci. we arranged for isabel guzman, a norc administrative assistant, to fly two boxes of cards to new york as carry-ons on november 16th. on november 12th american flight 587 crashed taking off from laguardia, but guzman still was willing to fly to new york and delivered to ndci on november 16th. ndci originally thought the conversion would take only a week or two, but complications soon developed. first, although the 38 year-old cards were in excellent condition, they had problems reading them because of the age of their card reader. they discovered that they “had to refurb our punched card equipment, it had been sitting around so long it got a little rusty.” it took them until the end of december to send a test file of deck three data. the file sent was corrupted, but this fortunately had nothing to do with the original data on the cards. an uncorrupted deck 3 test file 14 iassist quarterly winter 2001 iassist quarterly winter 2001 15 was sent on january 9th. this file revealed that they did not fully understand multiple-punched data and that such data needed to be spread. smith was familiar with spreading multiple-punched data from the 1970s and was able explain to ndci what needed to be done. over the next three weeks, ndci sent various versions of the three decks separately which smith checked against the original documentation and marginals. various data interpretation issues were ironed out and a problem with new variable names being too long was resolved. on january 28th, a merged spss file with all three decks combined into one record per case was received that smith verified as matching the original file. at this point the kas data had been fully recovered, but the spss file was very crude and not user friendly: 1) variables had names that referred to their original column locations rather than meaningful mnemonics, 2) there were no value or variable descriptors, missing values, or other data definition information, 3) some variables had long alpha values where short numerics were needed, and 4) other reformatting was need. after close examination of the data and the original documentation, a final spss system file with detailed labels was finished and the data set was archived at the roper center. lessons norc is to be praised for its thorough documentation of the 1963 kas and for preserving that documentation and the data for 38 years. but norc is not a data archive and it merely stored the information. as a result, the study itself did not appear on any readily accessible listing of norc surveys. it was only because smith had worked with the data on the scp in the mid-1970s that kas came to light after the september 11th terrorist attacks. even within norc many stored records are not well indexed and it took persistent efforts, the assistance of two ex-employees, and a bit of serendipity to unearth the data. moreover, once recovered we had data on a medium that was so antiquated that it took four months of extensive efforts to convert it to a modern, user-friendly format. the lessons from the kas experience are simple, but important. survey data must be sent to survey archives like the roper center and icpsr where the documentation and data will be preserved, backed-up, periodically updated as technologies change, indexed, and made routinely and easily accessible to researchers. failure to archive studies is poor science and a disservice to other contemporary researchers and those in the future. references allswang, john m. and bova, patrick, norc social research, 1941-1964: an inventory of studies and publications in social research. chicago: norc, 1964. bova, patrick and worley, michael preston, norc bibliography of publications, 1941-1991: a fifty year cumulation. chicago: norc, 1991. * paper presented at the iassist conference, storrs, ct, june, 2002. tom w. smith, e-mail: smith@norcmail.uch icago.edu, michael forstrom, national opinion research center, university of chicago. footnotes 1the survey research service (srs) was a division of norc that mostly conducted extramural research during the period 1963-1966. 2scp was mainly interested in identifying studies that included items that had been adopted by the gss to study trends from these baseline studies to the gss. 3the oʼhare storage is equivalent to the contents of about 17 three-bedroom houses. 4on-site storage consists mostly of more recent material. 5norc has two types of storage: records storage contains bar-coded boxes with some indication of the contents and exact location of each box. bulk storage consists of boxes and filing cabinets that are not bar-coded, whose content is less precisely known, and which are stored by skid (for boxes) or by bay (for filing cabinets), and not by individual box or filing cabinet. the access database covered the former, but not the latter. records storage was itself divided into general, records storage and two project specific holdings. only the general holdings were relevant in this case. 6punched cards are stored in corrugated cardboard boxes approximately 3.5 inches tall by 8 inches wide by 21.5 inches deep. they were stored as separate, stand-alone records and not placed in larger storage boxes. 7norc, like many other organizations, multiple-punched data to reduce the number of columns that data took up in order to save space which was very important when data appeared on punched cards and computing storage and analysis were very limited. for example, the study number (350) was entered in the 80th column of each deck. that is, the 3, 5, and 0 punches were all punched in column 80. this had to be converted into three, three-column fields with 350 for all cases. another example was a code-allthat-apply question in which codes 0-9, x, and y in column 28 in deck 2 each stood for a different reason for oswald killing kennedy and people could mention as many reasons as they thought applied. data from this one column had to be spread into 12 columns, which represented each of the 12 reasons as separate variables. lassist newsletter, vol. 2, no. 1 (winter 1978) the impact of computer netw0bkinf5 on the social science data library alice robbin university of hiscor.sin madison abstract this paper explores the factors which have constrained the social science data library's participation in the use of computer networks as a vehicle for accessing information. it also suggests why changes in the situation can be expected and further suggests some of the ways that computer network resource sharing will affect the social science data library structure and services. data and computations, sharing resources, and providing information services" (educom, 1973). at the introduction same time, there continue to be a number of factors which constrain traditional libraries are expethe development of network resource rienced in meeting information sharing and information servicing, needs and in coordinating activithese factors include: ties among users. computer technology has made it possible for the 1. technical considerations, traditional library to service the which revolve around user more quickly and (many enthuprocessor configuration, siasts of on-line data bases would software and communicaadd) more complicated than when the tions (davis, 1972); reference librarian relied on manual methods for searching and re2. financial considerations, trieving information. computer which involve the nature, technology has made it possible for size and distribution of traditional libraries to avoid some monetary support; costly duplication of human and technical resources. 3. organizational and political considerations. while the traditional library is which include the struca relative novice to computer techturing of the network, nology the social science data provision and nature of service organization (data library services, monitoring of or data archive) has been linked to performance, source and the computer and modern technologidistribution oi authority cal developments by the very nature and responsibility (daof the medium of its collection, vis. 1972), degree of machine readable data. the social control at the network science data library, while tied to level and local level, advanced technology, has continued and integration of the to use traditional means for locatlocal efrort into naing, transferring and accessing mational network efforts chine readable data files fmrdfl (educom, 1973); and for communicating its needs ana coordinating activities related to u. legal considerations, mrdf. moreover, the social science which involve federal and data library has neither benefited state legislative refrom the set of experiences of the strictions, which at the traditional libraries in resource federal level prohibit sharing within a networking envimonopolies and restraint ronment nor utilized network comof trade (clavton act of puters to share resources and ex1914) and protect the use pertise and to cooperate for more of communications servefficient allocation of resources. ices as a public utility (federal communications act of 1934, and neumann, 1973) , and which at the constraints im the use of state level are designed cdh7ijtet^n~t¥glktng to protect the outflow of b7~7he lociir state dollars for the iklm^e d5t^ libraii buying of non-state services; and. for several years now, as a growing number of articles, mono5. user considerations, graphs and books attests, computer which include knowledge networking has become an "important of user neeas and mode for remotely gaining access to lassist newsletter, vol. 2, no. 1 (winter 1973) characteristics, ease of these small centers remain system access, use and invisible to policy makers and thus operation, a variety of when decisions are made about comservices to assist in efputer use and about activities ficient and productive which will involve interaction with operations, education and centers outside the home institutraining, and documentations, these centers are never intion. formed. even with these constraints, the one of the results of the lack use of networking by libraries, reof institutionalization is the persearchers and students, particuception of these services as nonlarly in the natural sciences, has legitimate and its staff as unproincreased during the last several fessional. the staff members are years. this has not been the case not viewed as professionals by the for most professional social scienuser community which employs their tists, who have had little or no services, nor do the staff members experience in network use, or for perceive themselves as professional the social science data library, data specialists, although many whose clientele are social sciendata center personnel are indeed tists. experts in data processing and handling. although a staff member may what accounts for the low level in fact be performing the work of a of use of computer networks and why reference librarian or information has there been no network resource specailist who searches and resharing by the social science data trieves selected information upon library? the reasons are strucrequest, as indeed most staff at tural (the result of political relocal data centers do, that staff alities and historical accidents) , member usually does not recognize economic, sociological, and experthe role he perf orms--t ha t is, canientiai. structural factors innot assign a name to the function elude the lack of data services, of being performed. he usually lacks institutionalization of the inrorma methodology for the tasks he is ation service at the local level, performing. professional training and of professionalization. provides tools and products (resources) , an explanation for the efforts at coordination and reactivities performed by certain insource sharing have been made at dividuals, and a metholodogy for national and international levels, task performance. but. in most but only minimal efforts have been cases, the data center staff member made to encourage the development has not been trained as a profesof infrastructures at the local sional. level. major archives have maintained their dominance and have^ why should the lack of profescontrary to public expressions or sionalization affect the use of support, done little to encourage computer networking and resource the development of the local data sharing? reference work implies center. but, networking depends on knowledge of and understanding of the creation of local "nodes" and the nature and potential of availawithout the local effort, national ble information resources. if peonetworking will not be successful. pie are unaware of resources, they cannot utilize them. the informamost organizations which provide tion specialist today is made aware data services are structurally of on-line data resources at the weak, existing as appendages to one introductory course level and reunit of a larger parent organizaceives training in informational tion, rather tnan as an independent (bibliographic) data base creation unit within the parent organizaand manipulation. the housekeeping tion. hith only tenuous funding and maintenance functions performed support, the staff must dedicate by libraries are facilitated by a its efforts to maintaining services networking environment. in other with a continually eroding funding words, our data services personnel base. the staff therefore has few are uninformed of the potential use or no incentives to develop or emof networking facilities because ploy networks to communicate with they have not been trained as proinformation services outside its fessional information specialists, local environment. in generalj these data services are small ana economic reasons also explain operate on the periphery of computwhy the social science data library ing activities of their parent orhas made little use or computer ias5ist newsletter, vol. 2, no. 1 (winter 1978) structu sof twar ineffic sources i n e x p e n the use tor) h revise ter pro assista an ext data, or crea anaiyti ceptabl sis pro society support patabil res, e co lent si ve ran ave docu gra nee enae proc ting c ca e as cess ha was ity. so uld us staf com d st had ment sup de d p essi so pabi pect s h te: and that be wr e of f ti modit af f ( litt ation port lays, roces n g t h f twar litie s of in ot ad en dupl dela special itten, comput me has y , so t and adm le ince / prov and bet as a r s of em and e with s have the d a t her wor ough m ication ys. purpose and for ing rebeen an hat both inistrantive to ide better user esult of locating locating certain been aca analyqs, tne oney to , incomth commi techn purpo servi bly b velop tangi proce mg i the user. which whose short been gener port of p solve at a lecte forma vant diffi -they it is suppl the t for e prese commu cient commu mr from to th a f e varia probl prima use. reali of tr his h are p the d guter le t unive compu side physi ganiz e society has made tment (investment) ology for direct ses, but not for cing purposes. tn ecaase it is far e measurement tools bie byproduct than ss, which informa s. products can b funding agent an but, informati rarely offer a " "product" is rare time after the supplied, have be ate sufficient fu its intended goal, articular data particular proble particular moment d from a very larg tion, most of whic to the problem, cult to justify i must be taken on difficult to demo ying information otal social cost xample. althoug nts a useful te nicating informa ly, the resources nications are not df co seve e pop w to bles. ems h ry c the stica ansmi ome s rohib ata' s fund o tha rsity ting the caliy ation llections ra rai hundred ulation of th literally t although t ave not prove onstraints in social scient lly have the ttmg the in ite. transm itive. utili home site r s which may t researcher prohibits dollar expen campus. wh transfered f to another. an in appl inr is i asie to to tion e of d p on s prod ly v serv en u nds th requ ms, in t e se h is it ntan fait nstr will of r h ne chni tion to aval nge obse e 'j. hous ecnn n t ne ist form issi zing equi be u bee the ditu en rom cos enormous computer icat ions or mation s probar to dejudge a judge a ser vicfered to otentiai er vices, uct" and isibie a ice has nable to to supe supply ired to required ime^ set or inirreleis verv gibles-h alone. ate that reduce esearch, tworking que for ef f iprovide lable. in size rvations s. , from ands of ological o be the tworking does not ssiblity ation to on costs data at res comna vaiiaause his use of res outata are one orts which are calc for ment and data tion data pear tne dist if made prod ment t face to for whic sele user the acce remo or ob ther quir acco er s . sues data more remo incu ulate proce ation over h is b in v prod s to cost ribut this , ho ucer ' rred (sta ssing eadf! ilied estme ucer, have of a ed in physi w do s ca are f f an the rcell th for nt i and no pr data this cal we p iital fair d c r eq ing us, the ncur th oble fil fas tran rote izat ly ompu uest and the cap red e se m ju e wh hion sf er ct ion easy to ter time , docupostage, buyer of italizaby the her apst if ying en it is but, is not the data investheor of writ the h a cted data ssin te u lems e ar e re unti t an ev te a et ic it, e an co d1 us f ll uch wo g th ser ne e ac prog ng s here bout d ho iden cces ally it ac mput ts c es that uld e d wou ed coun ramm yste are who m t wh s ar , at 1 appear countin er bil osts o by the the 1 not be ata fil id. a to be ting on ing of m of m philo should uch, w en net e invol s qui g al ling f ac statu ocal bill e, var re es wh the ost sophi . p^yuich worki ved. on the te easy gorithm system cessing s of a user of ed for but the iety of solved, i c h recurrent computcal isfor the become ng and there are a variety of ical reasons which expla social sciences data lib community has not made u computer networking and sharing. in general, the user community is not p use electronic means to formation and communicate other. computer networ requires a moderate f with computers. como while becoming more comm social sciences, is usual to one or at most two cou mester in each discipline is largely for a class p an exercise in data hand information or data m although the profession scientist may from time frustrated in his inabil cate mrdf sources, in relies on his colleague "invisible college" for i on sources of data and themselves. this actio ates the lack of support library, since it perpe fact that the best data a chived (but if one knows people, one gets access rormation) and reside i hands, and presumes tha brary staff can provide sistance in information s the social scientist. sociologin why the rary user se of the resource library 's repared to access inwith each king still amiliarity uter use, on in the ly limited rses a se, and use roject and ling, not anagement. al social to time be ity to logeneral he s and the nf ormation the data n perpetufor a data tuates the re not art h e right to the inn private t data lilittle aseeking for 5 lassist newsletter, vol. 2, no. 1 (winter 1978) re resou sense the makes whom on k which the who a ratri educo bert tor o tium hinise mariz with have about thems lytic worki needs cial conce bert tists ing, socio commu expla netwo and t what cial its u source sharing rces are known, , deemphasizes invisible colie3 individuals les they know and nowii-g informat are machine bas assistance of re specialists eval. soni.e yea m conference, r (19 73, p. 14 9), f the inter-univ for political r if a political s ed social scie regard to r^lrdf; three needs: ( data, (2) acce elvjs, and (3) capaoilities. ng meet these s ? the discussio scientists was c p t of networking concluaed that were not ready thus, it has be logical aspects nity which prov 'or th which nation rking use: lack he lack of under networking is an scientist could se. implies and the imp e, sin s depenu more dep ion res ed and r i n te r m e dm infor rs ago, ichard h former ersity c esearch, cientist ntists ' he sal 1) mfor s s to t h a vaiiabl kow coul ocial s n group onfusea , a n d h social for ne en perha of the ide the lowlev of expe standing d why t beneiit that n this act of ce it ent on endent ources eguire iaries mation at an of f erdireconsorand , sumneeds d they mation e data e anad n e t cience of soby the or ferscient w o r k ps the user best el of rience about he sofrom new methods of organizing and changes in orientation for of comp in netw must be izing i creatin product product must b standar order t and to ity. a nomic, tial fa will co use, some op ture s network velopme tion t require ment . velopme social ices, ( for des informa profess age the 1 1 o n t h transfe data uter ork inf nfor g an s. 3 de e cr as i o r enha itho see ctor ntin rece timi ocia ing nts hat or thu nt o scie crib tion iona inf at r of 11 net reso rast mati d u to scri eate or etri nee ugh iolo s de ue t nt d sm i 1 sc acti stem info gani s, w f a nee deve ing is a orma cost the brarie works urce 3 r u c t u r on ser tilizi sha bing t d. their eve th the p the st gical scribe o con eveiop s war ience vities from rmatio zation e are n i n f r data lopmen and c 3) re re nee tion, s of infor s to and t ha ring es f o vices ng inf re re he se r there produc e mf roduct ructur and e d in s strain me nts ranted data t the n a bo and seeinq astruc librar t of s ontroi cognit essary pr oauc mation make use o engage , there r organand for ormation sources , esources must be tion 1 n ormation 's utilal, ecoxperienection 1 network suggest for fuiibrary hese derecogniut mrdf manage(1) aeture or y servtandards iing the ion that to m a n recognition and warrant managing access to inrormation, and (5) understanding of various technological advances which the social scientist can utilize to access information and transfer and retrieve data more efficiently and effectively. a that esse tent tion grow data need ophy nati ble, orde tial in vo comi pubi and cial chan uced acti spec age rapi mote file tran sis anal guir tent the his thes ble. are the will chin sear need suff retr and ess. cceptance among researchers secondary analysis of data is ntial to realizing the full poial of expensive data collechas been increasing. the ing cost of creatinq complex riles to meet a variety of s is resulting in a new philosthat these data represent a onal resource, publicly availato be widely' disseminated in r to realize their full potaninter-disciplinary research, iving a variety of data, is bena of increasing importance for ic oolicy planning and making for" sustained analysis of so, political and economic ge. complex data files prodby large-scale data gathering vities nave induced a need for ialists who can organize, raanand document the information, d reductions in the cost of rely accessing tnese complex s is making it unnecessary to sfer to a local site for analypurposed. more sophisticated ytic techniques are being reed in order to realize tae poiai of these complex files, and social scientist is readjusting perspective on requiring that e techniques be locally availathe researcner and analyst finding it necessary to support creation of structures which organize a collection of mae readable data, assist the recher in locating data for his s, and provide the analyst with icient documentation to make ieval of statistics possible to eliminate error in tne procfail trie val recogni data ri the man informa recogni tity o very di possibl evant evaluat ation c process etv mus gies fo mforma society tion of ures ha v tion les ner tion tion f in f f ic e, t bits e th olle its t q r co tion ' s r its th need of tha form ult, o lo of e ju cted elf, evel llec whi ecor elf. to o t h e ther t th atio ind cate inf alit and an op r tmg ch 1 dkee nr or dere col be r co e i e en n 1 eed and orma v of * th d th at io an 3 a ping matio d a lecti organ ilect s a ormou s mak somti seie tion the e col at th nal d ret bypro ana n regrowing ons of ized in ions of growing s auaning it mes imct reland to inf or mlection e socistrater ieving duct of evalua6 lassist newsletter, vol. 2, no. 1 (winter 1978) this has led to a reevaluation experienced data handler who has of the importance of libraries and never paid much attention to the information services and the growquality of documentation and for ing need for individuals who have whom the invisible college has opexpertise in information selection erated effectively to mare it posand retrieval, organization, mansible to obtain tne data he needed, agement , dissemination, and docuthis support comes at an approprimentation. the reference function ate time: during the last several is becoming recognized as a crityears, data information specialists icai activity for the supply of inhave been working on guidelines for formation to the user comraunitv. documenting mrdf, ranging from varthe individual who serves as a retious types of bibliographic deerence librarian for medf will in scriptions, products such as catathe near future be viewed as a prolog records, indexes and fessional. classification schemes and standardized study descriptions, to file support for the establishment of and variable level descriptions, data libraries is the result of software, such as the interchange recognition of the cost of social file, developed by roistacher and research: data libraries represent noble (1976), will obviate the nesavings in scarce resources with cessity of rewriting codebooks fortheir potential for collecting m matted for different statistical one location studies which are of software packages. thus, there is value to a variety of inaividuals; support in the user community for providing a systematic description the data professional's concern of these studies to facilitate loabout documentation and an apparent eating and utilizing data rewillingness to accept ti;e recommensources; reducing duplication of dations of these specialists so purchases; providing centralized that better descriptions of rirdf expertise in rile creation, procwill be available, essing, and description; and, providing a basis for a data services these developments encourage an infra structure to facilitate acatmosohere in which computer netcess to data at a reduced cost workihg can be accepted as a viable through memberships in consortia and desirable means to (1) access and through exchanges, enhance cominformation processing services munications about data resources, outside the local environment, (2) and lay tne foundation for the deshare resources which will provide velopment of data information orodeconomy of scale or operation to a ucts to benefit the user community. number of participants of the social research process, and (3) part of the failure in informashare intellectual resources and tion transfer and retrieval is due cooperate in joint programs (roisto the lack of standards for docutacher, n.d.). networks will promentation. in general, mrdf docuvide a mecnanism for more effective mentation has been poor and bibliocommunication, cooperation and cographic control unaeveloped. sue ordination of information through dodd, of the social science data services such as a social science library at the university of north data librarv. carolina, has commented extensively on this through the the classification action group of lassist (1977 ab, c) as have others involved in the impact of networking the development of documentation ok the socisi standards (nielsen, 1977; mochmann, srieuce citij library 1977; and robbin, 1977). the sociai science data librarian has ofthe social science data library ten been unable to locate files beis a soecial purpose library or incause no title or producer formation service which has been statement was provided. analysts created to respond most directly do not know how to acknowledge tne and immediately to its special use of secondary sources of data in clientele, a diverse community of their publications. study descripusers composed of social science tion providing brief histories of a researchers and students, policy data rile have never contained adeplanners and analysts. it is loguate descriptions of the data. cated in a variety of settings, within government, commercial orin the past year, however, we ganizations, foundations, and acahave seen growing support for docudemic institutions. the data limentation standards. it is signifbrary may be part of a computer leant tnat support for documentacenter, a larger library (faciition standards comes not only from ity) , a (general) information centhe socia.. scientist who does not ter, a data archive (in the eurohave easy or regular access to data pean sense) , a social science or is not iinkea to a major iniormdepartment, a research institute, ation network, but also from the lassist newsletter, vol. 2, no. 1 (winter 1978) market may be ganizat which (in mos tunding stances may be ing arr of spec may not brary • s provisi (which expert! the closely activit vide (e tents a ements library technic clien te els of and tec priat e of the techn ic the inf compreh materia collect ble els materia other , analysi the da agement quired tion. search organi independent the orga library is ases) the lib urce, except en part of it ported by ex ements for t products (w e a byproduct gular activit of specia be the resu re an ion the t c so wh sup ang ial b re on may se) coll as po ies an ff icie nd ret within staf al t le, w sudsta hnical inform libra al as ormati ensive 1 desc ion an ewhere is fro the p s, an ta to and for th zati ser niza imbe rary in t s ac tern he hich of ies) it on , or vice ortion in dded is ' s major hose mtivities al fundcreation may or the lior the services of staff ectio ssibl d is nt) a rieva the f 's raini ho ha ntive skil ation ry ' s sista on. know ribed d fa m one repar d the the analy e dat reflec e the die organized ccess to i 1 of selec collection substanti n g pr o v i ve differi , methodo is, with on the c collecti nee in ut the staf ledye of r in its re tentialiy e transfer location ation of d relation c o m p u te r a sis softwa a in the t s as ntele* s to prots conted elthe ve and des a ng levlogical approontents on and ilizing f has a esource ference a vailaof the to anata for ship of nd manre recol lee the mber dium issio blish on 1 d/or rvice tiona y be velop llect ta m ta a d rec nat io es. ta 1 ecial r ized spons on o cessi d di on re data lib of activi of its n." a da er and pr n machi may pr s much li 1 library responsic ment, da ion, da anagement rchiving ords mana n, and d in gener ibrary wh purpose by three ibie for n behalf oning mat sseminati guest. rary ties coll ta 1 oduc ne ovid ke t t le f ta ta (or geme ata icn libr fun loc of eria ng t engag relate ection ibrary er of r eadabl e inr hat of he data or stud pre para process data a acces nt, dat ref eren howev respon ary is ctions: ating its cl is it a hese m es 1 d to and may info fe or ma the lib y de tion ing naly sion a di ce s er , ds a cha i info lent cgui ater n a the its be a rmaf or m tion trarary sigh and and sis , ing) sseervthe s a ract is rmaele , res, ials the three basic functions of reference, accessioning and dissemination require access to different types of inrormation. the reference function requires access to information about the existence, availability and logical structure of m requ abou iir soft file tena stru rela ture the acce need rele tion rdf. ires t t ical data ware sto nee, ctur tion to diss ss t s f vant of the a acce he his struct to man and rage, qual e and ship of the eminati o infor or effi data f inforraa in is inv and ma lar ba as an tion a organi cializ for d used t ucts . brary ' ers potent it ela tion upon r inform w h ic h collec variet scribi resent about data cessf u manage tions which and ca format produc stat f retrie othe olve nage sis. inte nd t ze i ed p ata ge t s co appr lall ssif to f eque atio it tion y of ng t ace dat reso 1 i ment net per n su ion, e in tra ving r wor q in ment th rmedi he us nr orm resen resou nerat o de llect opria y use ies t acili st. on then th info he co ess t a res urces nf orm req work form pply o forma ined the ccession ss to tory of ure, rel agement compute processi ity of document the lo physical on funct mation cient r rom a la tion. ds, the inf orma processe e dat a ary betw er, to ation t t or pot rees wh e stati velop th ion, the te info ful dat he gat he tate it the lib acquire integr at e librar rmation llection ools for ources t hems el ation ga uires a among o a simil each ot rganizat tion pr in elas informat ing function information the data, ationship of and analysis r hardware, ng and mainthe logical ation, and gical strucstrueture. ion requires about users' etrieval of rger collecdata library tion seeking s on a regulibrary acts een inrormaretrieve and o meet speential needs ich will be stical proda data listaff gathrmation on a resourses. red informas retrieval rary records d materials es into the y produces a products de, which repinformation and for the ves. sucthering and communicarganizations ar function her with inions which oducts, and sifying and ion. w pute mg br ar coll mana enga orga pote tion othe serv zati hat r ne on t v 's ecti geme ges, niza ntia ship r 1 ices on. then i tworki he soc st r on , in nt ind t ions? 1 for of nf orma w i t h i s th ng lal uctu form oces rela n af f the tion n th ef nd re scie re, ation ses i tions etwor ect in data and pa feet o source nee da ser seeki n wai hip to king h g tne libra comnu rent' o f comsharta livices, ng and ch it other as the relary to tation rganinetworking appears ro indicate a trend toward 'centralization of services. the experiences of the 1950 's have shown that centralization of services was not cost effective and has not provided more effective and better user services; however, in the last two years, we have seen a move to centralize organizations which perform similar ias3ist newsletter, vol. 2, no. 1 (winter 1978) services (e.g., the carter administration's efforts to reduce "inefficiencies" in government by reorganizing its bureaucracy). the central computing center appears to be making an effort to exert control, both directly and indirectly, over other organizations which perform functions related to computing. after several years of reassessing its role in the parent organization, the computing center appears to be moving toward'efforts to centralize computing services and computation activities within the darent organization. it apeears that the computing center may e successful in its efforts because of its size (and budget) and the mythology of expertise which the computing center perpetuates. in an era of continuous inflation, there are increasing pressures on the administrators of the parent organization to maintain existing facilities which service the largest number of users and require the largest budgets. the social science data library, specializing in assistance to social scientist who have traditionally not been big users of computing services, is not able to generate large-scale support. thus, i would predict that as networking becomes an integral dart of the activities of a computing facility or large information service, data libraries whicr. do not have strong. independent constituencies within the parent organization, will de subtly and not so subtly pressurea by the parent organization either to be absorbed by a computing center or to have its functions provided by another organization which is responsible for information services. it will be hard to counter the trend toward centralization of services, especially because the computing center provides experience and know how which are strong arguments for extending services to a wider market, one which is covered by a social science facility like the data library. but, the trend should be resisted because it will mean that special needs of a specialized user community will most likely go unmet . in the po comput ties p format frees in a data 1 which and s networ to a c tion prin tent ing rovi ion the free ibra pro ervi king ompu whic ciple, n ial for centers ding com service data 11 market ry can 1 vides th ces at , howeve ting cen h has etwor inde and o puter s. brary situ ook e be the 1 r, po ter , opera king pena ther -rel ne to atio for st owes ses an o ted creates ence for f aciliatea intworking operate n: the a seller products t rate, a threat rganizalike a monopoly, because it affects its utilization and revenue. thus, the computing center will make every effort to control the outflow or computing dollars elsewhere. ratner than entering into a free market environment and upgrading the quality of its services, the computing center retrenches and begins to exert pressure on the parent organization's policy makers to stem the flow of money elsewhere. while it may be impossible to prevent the user commuaity from taking its money elsewhere, the computing center may convince adminstrators to make it difficult or at least inconvenient through a variety of bureaucratic measures to buy services outside. i would predict that this is a short term response and in the long run it will be desirable to develop competitive capabilities which will favor the user by helping to reduce prices for comparable services and to improve service quality and service availability ^neumann, 1973, p. 23). thus, i think that in the long run networking will free the data librarv from overwhelming dependence on its local computing center and at the same time provide the data library with better services from its local computing facility. protectionism has never worked to the long term benefit of the protected— as economic history has shown in the last century . n rela user work face staf time ices tion suit comm syst rath for 1973 meth been leve ing guir tica have prov lar thes carr fund mean use for libr more suit toma pabi expe etw tio ed ba ati uni em er fu a , ods n is is ing ' i ide wit e led ing tn or use ary in / =i ted lit rti orking will also affect t nship between staff a most data libraries ha on a one-to-one, face-t sis with their users. t has provided extensive a >nsuming user support ser training, tutorial inform documentation and human co on, because its us ty has preferred to use t on an "as needed basi than to spend time prepari tare system use tneuman p. 3) and because differe of user assistance ha ecessary to meet differe of expertise. when networ introduced as a means of a new information and stati roducts, the library wi o extend its resources services for users unfami h other operating system services will probably out without addition support. networking wi e development of and great automated interactive too r support services. da staff will probably beco volved in computers as a r nd in the development of a and interactive support c les (and ironically, t se required for the ne nd ve ohe nd vaner he s" ng n, nt ve nt kcs11 to 1s. be al 11 er is ta me 9 lassist newsletter, vol. 2, no. 1 {winter 1978) development will propel the data duplicatinq a copy of a data file library toward the computing center each time there has been a request which has more experience in sysfrom outside its local environment, tems development) . (this is in contrast to a book loan, where one copy of a book cirnetworking will make the user culates and there is no need to d ucommunity even more aware of the plicate a copy of the book each necessity for good documentation time a request is made for its for data, software and operating use.) other reasons which explain systems. while documentation has why data have not been distributed been primarily hard copy, networkon an interlibrary loan basis ining will probably increase the elude the extensive capitalization trend toward automating instrucinvestment that a data library tion, updating through the termimakes in developing its collection, nal. and development of systems didifficulties in the physical transrectories and catalogs of services fer of data and repeated use of the and products. medium on which the data are stored (typically magnetic tape) ; need to probably the greatest opportuprepare a data file in a physical nity that networking presents for structure compatible with the host the data library is in the area of environment to avoid additional retrieval of information about the processing; and, far more human existence and availability of mrdf, time required to prepare data for as documentation becomes automated, an external environment. furthernetworking will make it possible to more, it is not yet economically search for information contained in feasible to transmit data remotely "on-line data bases" of directories from one site to another because of data holdings, contents of data transmission speeds are too slow files, and codebooks and other docfor the quantity of data typically umentation for data files located analyzed by the social scientist, at institutions far away from the and because transmission costs are local data library. still too high. data libraries will have an opobtaining a copy of a data file portunity to participate directly has been the only way that a datam cooperative ventures to create library increases its collection data bases containing information (and thereby justifies the number on the contents of data files which of personnel required to maintain are located at local data centers. the collection, since quantity is there will of course be non-trivial always a more tangible measure of administrative and financial probservice than quality) „ it has allems to be resolved. but, networkways been assumed that when a data ing presents an unprecedented opfile is needed that it must be acportunity to create resources of quired, even when, as in most utility to a wider user community cases, what the user community does tnan now served by small local data is prepare a statistical overview centers and to enhance the quantity of the population in the data file and quality of information now and prepare some inexpensive, preavailable. liminary statistical results. (this is called, "getting a feel as networking becomes a more acror the data" and most researchers cepted activity in the generation begin their projects in just this of statistical products, we can exway.) in many cases, the data pect to see an increasing amount of which are acquired by the data listatistical analysis done remotely, brary on behalf of a user and are although it is doubtful that there reviewed in this manner, are rewill ever be more remote than local jected as not meeting the user's access and analysis of data. reneeds; and the researcher never mote access to data files will afcompletes a detailed analysis of feet the data library: its collecthe file, tion, how much time it allocates to the accessioning process, and how for every data file which is acit (and the user community) pay(s) quired, scarce resources of time for data which are paysically not and money must be allocated to inin the data library's collection. tegrate the file into the collection. this accessioning process is at present, a data library's expensive because the data must be collection grows by acquisition of checked to verify that what is dea copy of data archived and mainscribed as its physical structure tained by other data libraries and actually is and that the descriprepositories. it has never really tive materials (documentation) acbeen feasible for the data library companying the data allow the user to engage in an interlibrary loan to understand the logical structure type of data exchange because the of the data and to carry out stalibrary has had to bear the cost of tistical analysis of the file, and 10 lassist newsletter, vol. 2, no. 1 (winter 1978) because duplicate copies must be made to protect the data for future use. thu plicat costs tion data f and th ting d anothe discou sharin tional cooper tensiv brary a data at uios s, because e copies associate investment rom one 1 e infeasib ata remote r, data 1 raged from g methods librarie ative netw e investme in acquiri file wnic t a few ti or d wi an ocat ilit ibra uti used s th orks nt b h w mes. the need a data th capit d witn ion to an y of tra rom one s ries lav lizing re by more rough re there y the da and maint ill be ac to duf ile , alizamoving other, nsmitite to e been source tradigional is exta liaining cessed in wouxd limin taken acces cause sis w acces ferre attac for a distr could acces user is ac were cover data distr costs charg avail cessre one data anoth cal p inevi unf am data feren ducin vol ve reque inevi indiv elimi view acces the resou tion acces a ne be ary a plac sioni exte as in s of d to hed t ny co ibuto be p s of or f o cesse to p ed th and c ibuto to es i able tworki acquir nalysi e or w ng cou n a e d a tenaea the da the us o acce nsulta r woul aid ei the f i r each d. ay a e cos onsuit r coui all it cross for n ng envi e q only s oft hen acq id be 1 nd det cos ta cou er . a ssing tion t j provi ther fo le by a time t if the one tin t of a ing ass d ca ic s file all etwork ronm af he uisi usti aile t fo id b fee the hat de. r fi n in he d rem e f cces ista ulat s an data rem ent data ter predata had tion and fied bed analyr remote e transwould be data and the data a fee r s t time dividual ata fiie ote user ee wnich sing the nee, the e access d apply files ote acway o trans er : robie tably iliar to b t com g sta d in sttably iduai natin which sioni data rces whicr. sed i access v f reduc f e r fro by elimi ms of da occur b with pr e access gut ing if and proce reducing occur d copies g the bu takes ng proce library ' for acce may be n the f la ing m on nat i ta t eca u oduc ed w envi comp s s i n tl dl ue of rden plac ss ; s a ssio very ture netwo the e ii ng th ransf se pe i ng c ithin r onme uter g th e la to pr a da of e du and, lloca ning intr rking is cost of brary to e physier which ople are opies of a difnt ; retime iue data g s which ocessing ta file; data rering the reducing tion or informaequently the data library has justified its information gathering, acguisiton of new materials, ana retrieval and disseaiination of selectea information on the basis of the information's potential use for a variety of individuals in its user community, a small part o utiiizea on a working makes the cost of a from the infor individual wh cialized infor a reassessment resources for ment tasks wh performs on a lewer resourc records ma nag sioning basica lection, reso cated to ga retrieving s upon request, tomated' user which create which could p wide variety o erse user comra tinuing to services orie sc ientist . ithough o f its co regular b it possi cquiring mation se o require mation an of the a inf or mat ich tne regular es are a ement (w lly is) urces can thering electea and to de support intor mati otentiall f individ unities , provide nted to nly a very llection is asis. netble to shift information rvice to the d the sped to justify llocation of ion manaqedata library basis. if llocated to hich accbsof the colbe realloinf ormation, inf or mat ion veloping aucapaailities on products y benefit a uals in divwhile conspecialized tne social s and rela form comp tech will coram muni idlv sear mess will time shou i fore tion tor not "cos over prod libr in t two inai and bill coram data stjon is, a ne serv an d than der " ices seem tion plis on t get is , more ucce man tion atio uter nolo as unic ty's ch f age re id t ssf u agem ship n net gica sist at in n ee it or i in t ach requ here 1 in ent s t anion work 1 th 3 ^^ as e sho nf or ne n eve est fore f ormati require o coram g or mg is deveiod e data" s and i f f icien uld be mation etwork rvone a and re be sho on gathe s a set unicate ganizati an excel ment w library ts user tiy and easier by placi system w t the sponse rtened. ring or inons. lent hich in com"?o hich same time think e a rea ship bet and loca j ust 1 ting out head and uction ary (and ne socia or ganiza vidual u have qui ties to unities. librar d more 1 i n f o r u t ed-only ices may per at ion in ddva for an on t that m al plann hed if r he basis evened some org than o that netw ssessaent o ween tne da 1 service d n the tec " data and direct co and process t tl u s to t h 1 relation tions whicn sers quite te diifere their dif it may ies will b ike indivi ion will be basis. thi be paid f is perfor nee with a ticipated f he other ha ore effect ing can on esource sha that "eve out in the anizations thers, but or kin f th ta di ata 1 h n i q u pas ing e use ships res diff nt re feren well eg in duals acqu s mea or on med, "blan ut ure nd, i ive a ly be ring ry thi end" will ove g will e relastribuibrar v, es for sing on or data to the r) , but of the pond to e r e n 1 1 v sponsit user be that to re; that ired on n s that ly when rather ket orser vt would nd raa c c o m is done ng will ; that benefit r time. 1 1 lassist newsletter, vol. no. 1 (winter 1978) all organizations will benefit about the same amount. information sharing would argue for its cost being borne by the largest number of potential users possible. certainly, computer networlcing offer this possibility. summary this paper presents a cursory view or the factors which have nstrained the use of computer tworking and resource sharing by e social science data library and s user community. these factors e structural (the result of potical realities and nistorical cidents), economic, sociological d experiential. these factors 11 continue to constrain network e. but recent developments sugst that we will see greater use networking by the data library d social scientists. there is owing acceptance of secondary ta, growing costs of creating mpiex data files, need for spealists who can organize, manage d document these data, rapid rections in the cost of remotely cessing these complex files, recnitiod that collections of data les need to be organized and well cumented, reevaluation of the imrtance of libraries and informaon services, and support for the tablishment of data libraries, ese developments would seem to courage an atmosphere in which tworking would be accepted and en as a viable and aesirable ans to access information procsing services outside the local environment, share intellectual resources, share resources to provide economy of operation, and provide a nechansim for more effective communications, cooperation and coordination of information. he should not expect to see networking affect radical changes in the structure, services, staffing, collection and relationships with other data centers. 3ut, we can expect that networking will produce a push toward centralization of computer-related services within the parent organization in which the social science data library is located, in the future a freer market in which to buy computer services, more user support services which are computer-based, greater demands for library staff expertise about computing by the user community, better documentation of information materials, interactive information products whicn describe the contents of hrdf, remote access for preliminary statistical analysis of data, the shifting of resource support from the library to the user, fees for services rendered rather than for services anticipatea (as memberships in consortia are) , and a reevaluation of the relationship of data distributor-archive to the local data library. the conclusion is that in general computer network and resource sharing should provide better services, although the social science data library may be integrated into a larger information services organization. 12 iassi3t newsletter, vol. 2, no. 1 (winter 1978) references aiities of ks tor princeton , eraniversity il, inc.. , dodd, sua. report of the joint united states-canadian action groups on classification. lassist newsletter. 1,2 (1977), 5-tt7 report of the joint canadian-uniteds states action groups on classification. iassist newsletter. 1,3 (1977), 5-107 the emerging priority in bringing bibliographic control to social science .machine readable data files (mrdf) . iassist newsletter. 1,4 (1977), tt^th: educom. planning for national networking. proceedings of the educo;! sprinq conference, 1913. tprrnce^on7~nj:~ e'dwcuij tee interuniversity coaiaunications council, inc. , ' 1973) . hofferbert, richard. networks and disciplines. coemou themes and concensus; report ang ciscussion ot worlcsitofs. tprince^on, uj: 'e^ts'con, the interur. iversity communications council, inc. , 1973) . mochmann, ekkehard. information access at the data item level. sigsoc bulletin. 6 (2,3). neumann, a. j. review of network management problems and issues. nb3 technical note 795. t^asltingron , dc: u. h. department of commerce, october 1973) , 22-23. network user information support. u., s. department of commerce, december 1973. nielsen, per. information access at the data file level: documentation prerequisites on the file-level data base inauiry process. 3igs0c bulletin. '6 (2,3). robbin, alice. managing inforniation access through documentation of the data base. sigscc bulletin. 6 (2,3). roistacher, richard c. the data intercnange file: a first report. cac document no. 207. urbana, il: the university of illinois, 21 june 1976. and noble, barabara b. computer network support of social research communities. unpublished paper. urbana, il; university or illinois, n,d. 13 research data management tools and workflows: a report from the front 1/17 ribeiro, cristina; da silva, joão rocha; castro, joão aguiar; amorim, ricardo carvalho; lopes, joão correia and david, gabriel (2018) research data management tools and workflows: experimental work at the university of porto, iassist quarterly 42 (2), pp. 1-17. doi: https://doi.org/10.29173/iq925 research data management tools and workflows: experimental work at the university of porto cristina ribeiro, joão rocha da silva, joão aguiar castro, ricardo carvalho amorim, joão correia lopes, gabriel david1 abstract research datasets include all kinds of objects, from web pages to sensor data, and originate in every domain. concerns with data generated in large projects and well-funded research areas are centered on their exploration and analysis. for data in the long tail, the main issues are still how to get data visible, satisfactorily described, preserved, and searchable. our work aims to promote data publication in research institutions, considering that researchers are the core stakeholders and need straightforward workflows, and that multi-disciplinary tools can be designed and adapted to specific areas with a reasonable effort. for small groups with interesting datasets but not much time or funding for data curation, we have to focus on engaging researchers in the process of preparing data for publication, while providing them with measurable outputs. in larger groups, solutions have to be customized to satisfy the requirements of more specific research contexts. we describe our experience at the university of porto in two lines of enquiry. for the work with longtail groups we propose general-purpose tools for data description and the interface to multidisciplinary data repositories. for areas with larger projects and more specific requirements, namely wind infrastructure, sensor data from concrete structures and marine data, we define specialized workflows. in both cases, we present a preliminary evaluation of results and an estimate of the kind of effort required to keep the proposed infrastructures running. the tools available to researchers can be decisive for their commitment. we focus on data preparation, namely on dataset organization and metadata creation. for groups in the long tail, we propose dendro, an open-source research data management platform, and explore automatic metadata creation with labtablet, an electronic laboratory notebook. for groups demanding a domain-specific approach, our analysis has resulted in the development of models and applications to organize the data and support some of their use cases. overall, we have adopted ontologies for metadata modeling, keeping in sight metadata dissemination as linked open data. keywords research workflows, research data management, data publication, metadata, e-science introduction research data are created and used in diverse contexts. they may be generated specifically for a research project, such as sensor data captured in some experiment or interview data from a survey. they include data that are systematically captured for some purpose and then also used in research, such as meteorological data or the security logs of a computational facility. they may be ordinary data, such as a set of web pages, collected ad hoc to assess performance of a search engine. this https://doi.org/10.29173/iq925 2/17 ribeiro, cristina; da silva, joão rocha; castro, joão aguiar; amorim, ricardo carvalho; lopes, joão correia and david, gabriel (2018) research data management tools and workflows: experimental work at the university of porto, iassist quarterly 42 (2), pp. 1-17. doi: https://doi.org/10.29173/iq925 diversity makes research data management (rdm) a rather elusive task, for which researchers have neither well-established processes nor a clear intuition for its usefulness. data generated in large projects within well-funded research areas are already being curated in disciplinary infrastructures. the ncbi (ncbi resource coordinators, 2013) in the life sciences and icpsr (doty et al., 2015) in social sciences are good examples of stable infrastructures supporting large communities and used by researchers to get base data and to contribute new research outcomes. for data on the so-called long tail of science, there are still no general-purpose solutions and even when researchers recognize the value of data management, they typically have no support for making data visible, satisfactorily described, deposited, and searchable. efforts in this area are currently at the project stage, as illustrated by two large eu-funded initiatives, openaire (manghi et al., 2010) and eudat (lecarpentier et al., 2013). one important aspect to take into account is the stakeholders in rdm and the conditions that will foster their collaboration (ribeiro et al., 2015). figure 1 shows stakeholders and tools together with the steps in the research workflow. we concentrate on the roles of researchers, research managers and curators, in their collaboration with developers to build an effective support for dataset description and publication. figure 1 rdm stakeholders, workflow stages and supporting software gather process describe publish researchers research managers curators data providers institutionsdevelopersfunders s ta k e h o ld e rs w o rk fl o w s o ft w a re https://doi.org/10.29173/iq925 3/17 ribeiro, cristina; da silva, joão rocha; castro, joão aguiar; amorim, ricardo carvalho; lopes, joão correia and david, gabriel (2018) research data management tools and workflows: experimental work at the university of porto, iassist quarterly 42 (2), pp. 1-17. doi: https://doi.org/10.29173/iq925 the concern with rdm in the long tail is relatively recent, and several entities provide support for multi-domain datasets (dcc, 2017; ands, 2017; dans, 2017; dataone, 2017; dash, 2017). however, in some areas where large datasets are at the core of research there is a longer record of initiatives around research data, such as large databases identifying individual contributions, with a strong connection with publications; ncbi and icpsr were mentioned before as examples in the life sciences and social sciences, respectively. in other areas, as international groups grow, data sharing becomes more important, and research assessment requires the link to source material, new international infrastructures are being set up. this is illustrated by the projects promoted by esfri, which is supporting a network of long-term research infrastructures of pan-european interest (european strategy forum on research infrastructures, 2016). the work in rdm at the university of porto started with a scoping study, with the collaboration of 8 research groups, using existing recommendations and covering aspects such as the awareness with respect to data curation, the pressing needs regarding existing datasets, current solutions for data storage, the perceived value of legacy data, and the required support for rdm actions (ribeiro and fernandes, 2011). based on this study, we started to design a workflow for research groups, and the tools to improve the effectiveness of the publication and dissemination of research. the tools available to researchers are decisive for their commitment. for groups in the long tail, we focus on data preparation, namely on dataset organization and metadata creation. the workflow proposed for the long tail includes dendro, an open-source research data management platform, and labtablet, an electronic laboratory notebook for automating metadata creation. for groups demanding a domain-specific approach, our analysis has resulted in the development of models and applications to organize the data and support their primary use cases. overall, we have adopted ontologies for metadata modeling, keeping in sight metadata dissemination as linked open data. research groups and rdm requirements there is an overall recognition of the need for data access and reuse. low data reuse can be related to the difficulties faced by researchers in creating rich contextual metadata (faniel and yakel, 2011). this may be related to the fact that researchers tend to produce data documentation that is more focused on their own personal needs (mayernik, 2011). with respect to data documentation, it is possible to follow a traditional approach, describing data according to the well-established models used in publications, and depositing them in existing repositories, as yet another kind of publication. this is the line followed in initiatives such as the european project openaire+, promoting the zenodo repository for gathering and interlinking papers, datasets, software, and the projects where they originate. this approach makes data description a lighter task, but the lack of domain-specific data description can also make datasets harder to find, to understand and to reuse. researchers interested in a dataset, even if they locate it, will probably need to contact the authors to make sense of the data. as an alternative, we propose to empower researchers with expressive metadata models to describe their data. this requires tools, workflows and specialized support. our approach is to define a distributed multi-domain metadata model as part of the curator workflow, assessing rdm requirements at the domain level. for each domain, this includes the identification of domain https://doi.org/10.29173/iq925 4/17 ribeiro, cristina; da silva, joão rocha; castro, joão aguiar; amorim, ricardo carvalho; lopes, joão correia and david, gabriel (2018) research data management tools and workflows: experimental work at the university of porto, iassist quarterly 42 (2), pp. 1-17. doi: https://doi.org/10.29173/iq925 concepts, their representation as descriptors, and their incorporation in the dendro interface, used by researchers to describe their data with comprehensive contextual information (rocha da silva et al., 2016). we have worked with a growing set of research groups from different domains, whom we asked for a moderate commitment to rdm activities. these contacts started with a scoping study (ribeiro and fernandes, 2011) sent through the deans of the 15 schools in the university of porto. people selected by the deans have answered and we conducted 13 interviews with researchers in a broad diversity of areas, following the data asset framework recommendations (daf, 2017). we asked people to identify the nature and goals of their datasets, introduced them to rdm concepts, selected some relevant datasets for further analysis and evaluated rdm tools in the context of the regular activities of the groups. 8 out of these 13 groups provided sample datasets. 3 groups from this subset provided feedback on the prototype data repository we have built based on dspace (rocha da silva et al., 2012). the contacts with these 3 groups continued to the phase where metadata requirements were assessed. in the meantime, through personal contact and project collaborations, other groups got notice of our work and a panel of 11 groups was available by the time the first version of dendro was tested. currently we articulate the more in-depth work with groups, where new metadata models are considered, with a regular collaboration with researchers who need support in the use of standard metadata models, such as dublin core, to make datasets part of their research outputs. considering the areas where we have committed to the design of metadata models, they range from groups that produce experimental data, such as double cantilever beam experiments, to more computational-intensive scenarios, like vehicle performance simulations (castro et al., 2014; castro et al., 2015), or the combinatorial optimization problems in operations research (toledo et al., 2013) and predictive ecology studies in the biodiversity domain (rocha da silva et al., 2014). table 1 shows the groups that collaborated in data management contacts and their domains. table 1. groups involved in domain-specific metadata identification domain experiment affiliation fracture mechanics double cantilever beam engineering school, department of mechanical engineering analytic chemistry pollutant analysis engineering school, department of chemical engineering vehicle simulation bus performance in urban route engineering school, department of informatics engineering hydrogen production hydrogen production via chemical hydrides engineering school, department of chemical engineering biological oceanography interaction between marine ciimar (marine environmental research) and federal university rio grande (furg), brazil https://doi.org/10.29173/iq925 5/17 ribeiro, cristina; da silva, joão rocha; castro, joão aguiar; amorim, ricardo carvalho; lopes, joão correia and david, gabriel (2018) research data management tools and workflows: experimental work at the university of porto, iassist quarterly 42 (2), pp. 1-17. doi: https://doi.org/10.29173/iq925 and estuarine organisms biodiversity species dynamics and distribution, biological communities and ecosystems cibio (research center in biodiversity and genetic resources) social sciences social groups behaviors and beliefs psychology school computational fluid dynamics wind studies engineering school, department of mechanical engineering cutting and packing engineering school, department of industrial management to engage with these groups we applied different techniques, namely interviews, content analysis and data description experiments. non-scripted meetings also took place to get as much feedback as possible from the researchers. semi-structured interviews are usually the technique applied in the first interaction with the researchers, providing knowledge about the domain, their rdm practices, and the overall research data lifecycle used by the researcher or group. the interview script was adapted from the data curation profile toolkit and translated into portuguese. in the first case studies the interview was conducted with little forethought from the subject perspective (fracture mechanics, analytical chemistry). in the remaining cases the interview form was sent by e-mail beforehand, so the researchers had time to read and better prepare their answers, ensuring richer information. content analysis on researchers’ publications serves as a complement to the interview, since publications are good sources of information regarding data collection and analysis. a methodology section traditionally articulates pieces of information that may contextualize data, like experimental configurations, environmental features, and considering the many input variables; this was particularly interesting in the vehicle simulation domain (castro et al., 2015). there were exceptions, as with the hydrogen production group, where content analysis was performed before the first meeting to gain some domain knowledge. in the computational fluid dynamics case it was the main technique used for concept identification, since we did not have the researchers available for interview at the time. here, domain experts were consulted to validate the selected concepts in the first meeting. all the research groups have participated in a data description experiment using dendro, where they were invited to create a folder and upload a file, or more, and pick descriptors from the many metadata models previously loaded. at least two researchers from each group were involved, one https://doi.org/10.29173/iq925 6/17 ribeiro, cristina; da silva, joão rocha; castro, joão aguiar; amorim, ricardo carvalho; lopes, joão correia and david, gabriel (2018) research data management tools and workflows: experimental work at the university of porto, iassist quarterly 42 (2), pp. 1-17. doi: https://doi.org/10.29173/iq925 using a version of dendro with all the available descriptors to choose from. the other used a version with a recommendation system where the interaction with the first set of researchers was used to select a better ranking for descriptors (rocha da silva, 2016). as to the non-scripted meetings, some were upstream, while others were downstream, depending on researcher’s schedule or workplace proximity. we schedule preliminary meetings to present the objective of the following interactions, and to gather preliminary insight on rdm requirements and researchers’ perspectives. with research groups where a more in-depth focus was not possible we run opportunity meetings after the interview took place, or simultaneously. these downstream meetings were useful to discuss and detail the main points from the interview or to specify the information needed to describe the datasets researchers were working on at the time. overall, the level of the engagement with the researchers was not uniform, and depended on factors such as their metadata expertise, the relevance of using a different rdm tool depending on the research group, and the time that researchers have to dedicate to rdm activities. concerning metadata, while most groups were not familiar with the concept, the biodiversity research group was already working on metadata guidelines from the inspire directive in the context of a running project (pôças et al., 2014). for this case we were able to map the inspire concepts and develop a suitable domain metadata model (biome) (rocha da silva et. al, 2014). in other groups, not so familiar with metadata concepts, we took descriptors from metadata standards, such as the data documentation initiative (vardigan et al., 2008) for the social science domains, and the ecological metadata language (fegraus and andelman, 2005) and darwin core vocabularies (wieczorek et al., 2012) for the biological oceanography domain. for others, no ready-to-use vocabularies were identified, and a more in depth collaboration with the domain experts was necessary. concerning the use of tools (dendro and labtablet, detailed in the next section), while all groups carried out descriptions tasks using dendro, in some cases it was pertinent to also experiment with labtablet, depending on the context of the data. a researcher collecting field data or running an experiment in a controlled environment is more likely to use labtablet than one performing computational studies. therefore, researchers from the social sciences and biodiversity domains used labtablet. the time dedicated to work with each group is a dimension where we have no complete control. the collaboration timeline is difficult to predict and our experience has shown that the same amount of interaction or level of engagement can be accomplished over very different periods of time in different groups. for instance, in one of our cases we had three meetings over a period of three months, while the same number of meetings was performed over two weeks in another case, to achieve similar results. other researchers were contacted but our interactions were sporadic, or of a more informal nature. nevertheless, these groups are good complements to the previous ones and help us assemble a more representative set of research domains. data organization and metadata creation with dendro + labtablet the uptake of rdm depends on the existence of clear processes for researchers with respect to data collection, organization, description, deposit and publication. interviews with researchers on their https://doi.org/10.29173/iq925 7/17 ribeiro, cristina; da silva, joão rocha; castro, joão aguiar; amorim, ricardo carvalho; lopes, joão correia and david, gabriel (2018) research data management tools and workflows: experimental work at the university of porto, iassist quarterly 42 (2), pp. 1-17. doi: https://doi.org/10.29173/iq925 rdm practices showed a large gap between the processes used in paper preparation and the routines required to organize data and make them publication-ready. it was clear that appropriate tools might ease the task for researchers in small groups; this led us to design, implement and test tools to support data preparation in the long tail. these tools can have a very broad field of application—data collection, data cleaning, processing, organization, description, deposit, search, are some examples. as rdm is evolving as an integral part of research, we expect that repository platforms, either disciplinary or generic, will become mainstream. however, one should not assume that datasets can be managed as publications already are, particularly when it concerns metadata. research data are often purely numeric—unlike a paper, for example—and emerge from very specific and advanced studies, so the tools for data collection and data processing are likely to be domain dependent. any tools that attempt to capture the production context of a research dataset with these characteristics must be able to capture domain-specific metadata as well, adopting established metadata standards in the corresponding community whenever possible. also, repository managers may be experts at metadata and data management, but they are not the domain experts and thus do not know what domain-specific information is needed to retrieve and reuse the data. to respond to these two issues, we have approached data organization and data description with the goal of simplifying the interface between researchers and repository managers. any tools for preparing and describing datasets should therefore establish a consistent rdm discourse with researchers, while giving them the openness to interface with several repository platforms, so that they can share their data everywhere they want without having to fill in metadata multiple times. dendro is an open-source, ontology-based rdm platform currently in development at inesc tec and the university of porto, whose dependencies are also entirely open-source. it targets researchers as its main users, and helps them deposit and share data both within their research group and with external elements. dendro uses the ‘dropbox’ metaphor for data upload and adds sophisticated data description capabilities. the concepts in dendro include ‘project’—a project is a shared folder where every member can deposit files of any kind, ‘folder’—project members create folders and subfolders to organize resources, and ‘descriptor’—metadata are associated to resources using both domainspecific and generic descriptors. dendro is integrated in the research workflow to cover the steps between dataset creation and publication, and can export data and metadata, or just the metadata, to major repository platforms when the research group is ready to publish them (rocha da silva et. al, 2018). mobile devices have evolved to include advanced capabilities that render them suitable for an array of research-related activities. numerous cases of researchers using their own devices to improve the research workflow prove that this is an important trend. besides an increasing storage capacity, mobile devices are often permanently connected to the internet and have several built-in sensors that can provide contextual data on the researcher's environment effortlessly. this has been the motivation in the development of labtablet, a notebook application with an emphasis on rdm. https://doi.org/10.29173/iq925 8/17 ribeiro, cristina; da silva, joão rocha; castro, joão aguiar; amorim, ricardo carvalho; lopes, joão correia and david, gabriel (2018) research data management tools and workflows: experimental work at the university of porto, iassist quarterly 42 (2), pp. 1-17. doi: https://doi.org/10.29173/iq925 figure 2: production of metadata records using labtablet labtablet is an electronic laboratory notebook, i.e. an application that takes advantage of sensors onboard the mobile device to help researchers describe their data. in some cases the description with the mobile device replaces a process that used paper-based notebooks, making sure metadata are not lost but instead recorded, associated to data and deposited. figure 2 shows the labtablet interface, in a situation where it is used to collect data in a field experiment. a simple tool like this, integrated into frequently used devices, can contribute to get more metadata associated to datasets. moreover, with the mobile device, metadata creation becomes a seamless part of data collection, distributing the effort over the duration of the process and avoiding time-consuming task at the end of a project. the first step in data organization and description workflow using dendro and labtablet is illustrated in figure 3, with a project in the biodiversity domain. after the metadata are synchronized with dendro, we can see the structure of the project folders and files, the descriptors used for this domain (generic dublin core plus a domain-specific subset of the inspire descriptors) and the communication between dendro and labtablet. the generic workflow is as follows: metadata models are passed from https://doi.org/10.29173/iq925 9/17 ribeiro, cristina; da silva, joão rocha; castro, joão aguiar; amorim, ricardo carvalho; lopes, joão correia and david, gabriel (2018) research data management tools and workflows: experimental work at the university of porto, iassist quarterly 42 (2), pp. 1-17. doi: https://doi.org/10.29173/iq925 dendro to labtablet, metadata collection takes place in labtablet, and descriptor values are added to dendro when the two platforms synchronize. figure 3: reviewing and improving metadata records in dendro the availability of domain-dependent descriptors is one of the principles in dendro. as the number of domains grows, so does the number of available ontologies (representing metadata models), making it harder for researchers to locate the more appropriate ones. to cope with this information overload, dendro also offers descriptor recommendation as an advanced feature, assisting researchers (who are not expected to be research data management experts) in the discovery and selection of descriptors that match their description needs. manual descriptor selection is still possible, as the platform allows the user to restrict the ontologies from which the user will add descriptors from, controlling crossdomain descriptor usage. the effectiveness of the dendro platform has been tested at different times. three experiments were run in different conditions, and the results are already available for two of them. a fourth experiment is in the final stage of data collection. the first experiment, dendrodc1, used a preliminary version of dendro, after it had been tested with some researchers from our panel. the subjects, a set of students from an information science course, were asked to fill in dublin core descriptors for datasets published by several organizations. this experiment had two goals: to test dendro in realistic load conditions and to observe the use of a generic metadata model in the annotation of datasets with descriptor recommendation. the system proved robust in lab conditions and showed some improvement in description when recommendation was turned on (rocha da silva et al., 2018). the second experiment, dendropanel, was a more realistic one, and involved the 11 groups in table 1 and a dendro instance configured with the ontologies for the domains in case there were specific ones. https://doi.org/10.29173/iq925 10/17 ribeiro, cristina; da silva, joão rocha; castro, joão aguiar; amorim, ricardo carvalho; lopes, joão correia and david, gabriel (2018) research data management tools and workflows: experimental work at the university of porto, iassist quarterly 42 (2), pp. 1-17. doi: https://doi.org/10.29173/iq925 for each group (each domain) we had two participants, each describing a dataset from the corresponding domain: one used dendro with and the other without recommendation. results were in favor of the use of descriptor ranking, providing a recommendation based on usage data (rocha da silva, 2016). the third experiment, dendrodc2, replicated dendrodc1 with a new set of subjects and the results are not analyzed yet. a fourth experiment, currently collecting data, uses a combination of dublin core and the biodiversity ontology in the description of datasets in this domain, to gain insight on the importance of domain-dependent descriptors. several tests were also conducted using labtablet and researchers in engineering and biodiversity domains. some comprehensive experiments with labtablet are yet to be planned. in the meantime, the application has also been used as the basis for two data-related applications: one for displaying sea conditions for nautic sports (amorim et al., 2016) and the other to support the collection of specimen data in the seabiodata project, mentioned in the sequel (seabiodata, 2017). disciplinary solutions: dedicated data and metadata models as we explore the requirements of research groups in a large research institution, focusing on small groups where data curation needs to be established as an agile process, we were faced with cases that do not fit into the long tail. the solutions for these areas required specific projects supporting the analysis, design and implementation of solutions that satisfy more specific requirements. we now look at three of these projects that ran as separate initiatives, involving teams from the corresponding disciplines, but still have rdm at the core. they provide a view on the different functions covered by a domain-specific system, and extend the focus from organization and description to functions such as data collection, processing and archiving. sensor data from concrete structures the research method in the structural health monitoring, a subfield of civil engineering, is well established. to monitor a bridge, a tower, or other large structure, a research project is set up, the data collection system is designed and deployed on the structure, and data streams with periods of the order of milliseconds start to be produced during months or years. each sample in a data stream contains values from the various sensors in a specific data acquisition system (temperature, acceleration, etc.). each stream is segmented in files for transfer and storage purposes. these raw data files are cleaned and the result is subject to one or more processing routines. both the cleaned data files and the result files are visualized, and conclusions on the health of the structure are incorporated into reports and research papers. after the publications are issued, the data are sometimes discarded. the growing awareness of the value of data as evidence supporting the published conclusions and as source for further studies motivated the research team to launch a digital archive project tailored to their needs. our team has elicited the requirements with the strong involvement of the disciplinary experts and has designed an architecture for the digital archive with three components: a systematic directory structure and naming rules organize all the data files in the file system; a database collects all the metadata; and an application adds management and visualization functionalities. the metadata are organized in five packages: https://doi.org/10.29173/iq925 11/17 ribeiro, cristina; da silva, joão rocha; castro, joão aguiar; amorim, ricardo carvalho; lopes, joão correia and david, gabriel (2018) research data management tools and workflows: experimental work at the university of porto, iassist quarterly 42 (2), pp. 1-17. doi: https://doi.org/10.29173/iq925 1) contextual metadata on the structure characteristics and project details, dates, and funding body; 2) system information on the design, implementation and parameterization of the data acquisition system and its sensors and components; 3) data file structure, measured variables and units, sampling frequency, timestamp, and file location; 4) stakeholders like the structure owner and designer, the research team, and external researchers; 5) documents of any kind, from the project contract and the data acquisition system design to any pictures, as well as published reports and papers. the project resulted in a web application, its functionalities including the management of the five packages of metadata, the automatic ingestion of new data files, a graphic visualization tool to browse the data on intervals, and an asynchronous facility to export data files using oai-pmh. the system developed (da costa et al., 2014) is generic enough for monitoring systems in other fields besides structural health monitoring. wind data from lidars winds@up is a prototype of the e-infrastructure developed in the scope of the windscanner.eu project and includes a repository for experimental data sets, consisting of georeferenced time series data generated by lidar sensors used in open-air field tests (field campaigns). the platform manages the data and corresponding metadata, provides processing capabilities for in-situ data processing and is deployed in the windscanner hub (gomes et al., 2014). the windscanner device is a wind (short and long-range) lidar with scan-head and control software that, in coordinated operation of three units, can be used to measure 3d wind vector field with high accuracy. this is the core technology of the windscanner infrastructure to be used by the corresponding esfri project. windscanners are deployed at existing or planned test facilities, covering different climate conditions and terrains. standard procedures may be applied to collected data and the resulting time-series are stored for further use. data from the time series have gps timestamps used to align values for signals obtained from different devices. besides raw and processed data, the winds@up platform provides storage for metadata and for other resources—research objects—created by researchers in the course of their work. research objects include datasets, archives, photos, charts and scientific papers. metadata for the research objects are organized in three categories: 1) descriptive—title, author, abstract and keywords, which help discovery through searching and browsing; 2) administrative—preservation, rights management, and technical aspects such as format or experimental setup; and 3) structural—how different components of a set of associated data objects relate to one another (e.g. datasets, procedures or results). https://doi.org/10.29173/iq925 12/17 ribeiro, cristina; da silva, joão rocha; castro, joão aguiar; amorim, ricardo carvalho; lopes, joão correia and david, gabriel (2018) research data management tools and workflows: experimental work at the university of porto, iassist quarterly 42 (2), pp. 1-17. doi: https://doi.org/10.29173/iq925 the experiments' raw data are transferred from the devices or data-loggers to the campaign site where a quality assurance process validates the received data. data are also transformed to standard formats, such as the network common data form (netcdf)2. the quality assurance process may require, depending on the campaign, the processing of raw data to produce "clean data". these data are similar to the raw data—time series—but some portions may be purged due to detected errors, and some pre-processing operations may be performed (e.g. 2 or 10 minutes averages). in addition, further documentation should be added, e.g. detailing the process used to clean data. the data and associated metadata, packed in self-describing, machine-independent data formats in netcdf files, are uploaded to the platform repository. at this point, these research objects are described using the same metadata model (captured as an ontology) used for other research objects in the domain. raw data, clean data and processed data are used together with related research objects by a web portal with search facilities suitable for researchers, enabling research collaboration and the reproducibility of results. winds@up is designed as an e-infrastructure, providing dedicated in-house storage for a community. this is one of the differences from the long-tail cases, where external repositories are required to store data. the concepts of campaign, experiment, observed phenomena, and time-series data are common to many domains, and so are the generic metadata elements used to capture them. winds@up is currently a solid prototype, which underwent three development cycles and contributed to the preparation phase of an esfri research infrastructure. further development is expected in the context of the next phase of the esfri national and european wind research infrastructure (windscanner, 2017). the most recent version of the prototype (windsp) has been used recently to plan the experiments of the european project newa (newa, 2017) and has followed the execution of the field campaign in perdigão (witze, 2017), a double-hill field experiment in portugal. seamounts physical and biological data we have been approached by a large team on marine research in order to build a digital archive for their data. marine research is typically organized in campaigns aboard a research ship. a campaign is composed by a set of stations in predefined locations. in a station samples are taken from the soil and the water column using appropriate devices and procedures. the samples may be subject to an onsite study that may be later complemented with physical, chemical and biological analysis, using predefined procedures. the collected data are very diverse in nature, ranging from variables measured in specified units to the identification of the presence or density of biological species, to pictures, video and sound, to actual captured specimens. the georeferencing and the detailed log of the actual sample collection process and involved researchers are very important. at the same time, data streams are received from instruments installed in buoys, from the positioning of ships, and from satellite data out of the campaign scenario. this marine research field, as a subfield of environment research, is heavily controlled by european and portuguese regulations like inspire (european commission joint research centre, 2013) and those issued by snimar (snimar, 2016). the inspire european directive imposes a rather strict api on environmental repositories in order to improve interoperability. this api embodies an abstract data model centered on observations with values for properties, associated procedures and geographic https://doi.org/10.29173/iq925 13/17 ribeiro, cristina; da silva, joão rocha; castro, joão aguiar; amorim, ricardo carvalho; lopes, joão correia and david, gabriel (2018) research data management tools and workflows: experimental work at the university of porto, iassist quarterly 42 (2), pp. 1-17. doi: https://doi.org/10.29173/iq925 references. there are implementations for this abstract data model and we have chosen two: the 52north sensor observation service (sos) implementation and the geoserver web feature service (wfs), web map service (wms), and web coverage service (wcs) implementation. the paradigm for this case is different from the two previous ones, where the data resided in data files with specific formats stored in the file system. in this case, each observation is stored in a database, along with the metadata, thus making it easier to search the data using ad hoc queries. in this project the goal was to develop a client (called seabiodata) for those two servers that maps well with the concepts, data sheets and methods of the marine research team and is able to ingest the data already available in digital form and to present them in visually effective ways. requirements in disciplinary infrastructures in disciplinary platforms, such as the windscanner.eu and seabiodata, besides the immediate requirements we are dealing with, it is likely that, as more research groups participate and use the platforms, more specific functionalities will be requested. this can be regarded as a natural evolution of the infrastructures and even as a sign of their successful adoption. one effect that may result from the definition of disciplinary infrastructures stems from their specialization: datasets are managed in specific structures that, even if they are open, may become hard to discover and search. two current lines counter this effect. one is the existence of research data aggregators, such as re3data, supported by datacite (brase et al., 2015). the other is the trend of linked open data (lod) as applied to metadata. if metadata can be harvested by lod services, the possibilities for data to be available to different kinds of applications increase. conclusions as we look at research data management, we cannot help recognize it is a daunting task. in a way, this is common to all archival endeavor: how can we estimate the valuable assets, when not all can be curated and preserved? with respect to traditional archives, web archiving initiatives have already dealt with the problem of making selection and description a more lightweight, semi-automated task. a similar problem confronts research data archives, where probably some collections will be way better curated than others. at the current point in rdm efforts, we have to struggle to achieve a balance between the effort required for data description, which will render datasets part of the research trail, as publishable and re-usable artifacts, and the immediate rewards that researchers get from their investment in data curation. we argue strongly in favor of tools that can alleviate the researchers’ tasks, embed data curation in the overall research activities, and provide trust on the resulting research outputs. infrastructures and tools appear at a faster pace than researchers can handle, and there is a large gap between their functionalities and the requirements as perceived by researchers. the second main line of action is therefore the analysis of requirements for rdm, the matching of requirements with existing tools, and the identification of the missing links. in the work reported here, there is a focus on solutions for the long tail, but also some cases concerning rdm in areas that require custom-designed solutions. it is important to assume from the start that not all research groups need the same kind of solutions, while looking for common issues. we take the examples of a civil engineering group collecting stream data from sensors, a wind research https://doi.org/10.29173/iq925 14/17 ribeiro, cristina; da silva, joão rocha; castro, joão aguiar; amorim, ricardo carvalho; lopes, joão correia and david, gabriel (2018) research data management tools and workflows: experimental work at the university of porto, iassist quarterly 42 (2), pp. 1-17. doi: https://doi.org/10.29173/iq925 community designing the infrastructure to collect and process large datasets, and a marine and atmosphere institute dealing with the diversity of data used in forecast products, marine research and specimen collection. for each of these groups, we have designed and implemented custom systems dealing with some parts of the research workflows. the field work with diverse researchers and the controlled experiments performed to evaluate the tools have both confirmed the perceived need for a more solid support to the full rdm workflow. in interviews, experiments and informal conversation, researchers have been curious about the concepts they are not familiar with—metadata, ontologies, repositories, preservation— and provided many clues on the ways to address their requirements. moreover, rdm is becoming a strong concern as national and international policies and funding bodies move towards open science. both kinds of motivation are essential to take rdm forward. without genuine interest from researchers, rdm actions may become just another administrative burden with no real commitment to data organization and description. but the institutional requirements are also crucial to provide short-time rewards and bring rdm to the attention of researchers in all areas. with the data organization and description tools reaching maturity, our concern is now the complete research workflow, from data collection to data preservation. in the continuation of the work with the research groups we are setting up partnerships to collaborate on the definition of rdm strategies starting with data management plans and continuing to their execution and evaluation (active dmp, 2017). another important aspect is the identification of researchers and research outputs, and the connection to systems such as orcid for researcher identification and doi for outputs. this brings interoperability with institutional systems, and therefore visibility at institutional level and in international repositories, and enables the management of links between different research results. metadata models, their evolution in communities and the promotion of metadata harvesting and aggregation in data repositories is also a long-time effort where early engagement will contribute to motivate and reward researchers. references active dmp. 2017. research data allianceactive data management plans interest group [online]. available: https://www.rd-alliance.org/groups/active-data-management-plans.html [accessed july 2017]. amorim, r. c., rocha, a., oliveira, m. & ribeiro, c. 2016. efficient delivery of forecasts to a nautical sports mobile application with semantic data services. proceedings of the ninth international c* conference on computer science & software engineering. porto, portugal: acm. ands. 2017. andsaustralian national data service [online]. available: http://ands.org.au/ [accessed july 2017]. brase, j., sens, i. & lautenschlager, m. 2015. the tenth anniversary of assigning doi names to scientific data and a five year history of datacite. d-lib magazine, 21. castro, j. a., perrotta, d., amorim, r. c., rocha da silva, j. & ribeiro, c. 2015. ontologies for research data description: a design process applied to vehicle simulation. metadata and semantics research 9th research conference, mtsr 2015. https://doi.org/10.29173/iq925 http://ands.org.au/ 15/17 ribeiro, cristina; da silva, joão rocha; castro, joão aguiar; amorim, ricardo carvalho; lopes, joão correia and david, gabriel (2018) research data management tools and workflows: experimental work at the university of porto, iassist quarterly 42 (2), pp. 1-17. doi: https://doi.org/10.29173/iq925 castro, j. a., ribeiro, c. & rocha da silva, j. 2014. creating lightweight ontologies for dataset description. practical applications in a cross-domain research data management workflow. ieee/acm joint conference on digital libraries (jcdl), 2014. da costa, f. p., cunha, a. & david, g. 2014. vibest shm: an information system and data repository for structural health monitoring. 9th international conference on structural dynamics (eurodyn). daf. 2017. data asset framework [online]. available: http://www.data-audit.eu/ [accessed july 2017]. dans. 2017. data archiving and networked services [online]. available: http://www.dans.knaw.nl/en [accessed july 2017]. dash. 2017. dashdata sharing made easy [online]. available: https://dash.cdlib.org/ [accessed july 2017]. dataone. 2017. dataone [online]. available: https://www.dataone.org/ [accessed july 2017]. dcc. 2017. dccdigital curation centre [online]. available: http://www.dcc.ac.uk/ [accessed july 2017]. doty, j., herndon, j., lyle, j. & stephenson, l. 2015. learning to curate. bulletin of the american society for information science and technology, 40. european commission joint research centre 2013. inspire data specification on biogeographical regions – technical guidelines 10.12.2013. european strategy forum on research infrastructures 2016. strategy report on research infrastructures. faniel, i. m. & yakel, e. 2011. significant properties as contextual metadata. journal of library metadata, 11. fegraus, e. h. & andelman, s. 2005. maximizing the value of ecological data with structured metadata: an introduction to ecological metadata language (eml) and principles for metadata creation. bulletin of the ecological society of america, 86, 158-168. gomes, f., lopes, j. c., palma, j. l. & ribeiro, l. f. 2014. winds@up: the e-science platform for windscanner.eu. journal of physics: conference series, 524. lecarpentier, d., michelini, a. & wittenburg, p. the building of the eudat cross-disciplinary data infrastructure. egu general assembly conference abstracts, 2013. egu2013-7202. manghi, p., manola, n., horstmann, w. & peters, d. 2010. an infrastructure for managing ec funded research output the openaire project. the grey journal (tgj): an international journal on grey literature, 6. mayernik, m. s. 2011. metadata realities for cyberinfrastructure: data authors as metadata creators. proquest dissertations and theses, 338. ncbi resource coordinators 2013. database resources of the national center for biotechnology information. nucleic acids research, 41, d8-d20. newa. 2017. newa, new european wind atlas [online]. available: http://www.neweuropeanwindatlas.eu/ [accessed july 2017]. pôças, i., gonçalves, j., marcos, b., alonso, j., castro, p. & honrado, j. p. 2014. evaluating the fitness for use of spatial data sets to promote quality in ecological assessment and monitoring. international journal of geographical information science, 28, 2356-2371. ribeiro, c. & fernandes, m. e. m. 2011. data curation at u.porto: identifying current practices across disciplinary domains. iassist quarterly, 35, 14-17. ribeiro, c., rocha da silva, j., castro, j. a., amorim, r. c. & fortuna, p. motivators and deterrents for data description and publication: preliminary results (short paper). on the move to meaningful internet systems: otm 2015 workshops, 2015. 512-516. rocha da silva, j., ribeiro, c. & correia lopes, j. 2018. ranking dublin core descriptor lists from user interactions: a case study with dublin core terms using the dendro platform. https://doi.org/10.29173/iq925 http://www.data-audit.eu/ http://www.dans.knaw.nl/en https://dash.cdlib.org/ https://www.dataone.org/ http://www.dcc.ac.uk/ http://www.neweuropeanwindatlas.eu/ 16/17 ribeiro, cristina; da silva, joão rocha; castro, joão aguiar; amorim, ricardo carvalho; lopes, joão correia and david, gabriel (2018) research data management tools and workflows: experimental work at the university of porto, iassist quarterly 42 (2), pp. 1-17. doi: https://doi.org/10.29173/iq925 international journal on digital libraries, april 2018, pp 1-20, https://doi.org/10.1007/s00799-018-0238-x rocha da silva, j. 2016. usage-driven application profile generation using ontologies. ph.d., universidade do porto. rocha da silva, j., castro, j. a., ribeiro, c., honrado, j., lomba, a. & gonçalves, j. 2014. beyond inspire: an ontology for biodiversity metadata records. on the move to meaningful internet systems: otm 2014 workshops. rocha da silva, j., ribeiro, c. & correia lopes, j. 2012. managing multidisciplinary research data: extending dspace to enable long-term preservation of tabular datasets. ipres international conference on digital preservation. rocha da silva, j., ribeiro, c. & lopes, j. c. 2016. usage-driven dublin core descriptor selection. research and advanced technology for digital libraries: 20th international conference on theory and practice of digital libraries, tpdl 2016. springer international publishing. seabiodata. 2017. seabiodataportuguese seamounts biodiversity data management [online]. available: http://eeagrants.org/project-portal/project/pt02-0017 [accessed july 2017]. snimar. 2016. the snimar metadata profile [online]. available: http://www.snimar.pt/ar/ficheiros/perfilsnimar.pdf [accessed july 2017]. toledo, f. m. b., carravilla, m. a., ribeiro, c., oliveira, j. f. & gomes, a. m. 2013. the dottedboard model: a new mip model for nesting irregular shapes. international journal of production economics, 145, 478-487. vardigan, m., heus, p. & thomas, w. 2008. data documentation initiative: toward a standard tor the social sciences. the international journal of digital curation, 3, 107-113. wieczorek, j., bloom, s., guralnick, r., blum, s., doring, m., giovanni, r., robertson, t. & vieglais, d. 2012. darwin core: an evolving community-developed biodiversity data standard. plos one, 7. windscanner. 2017. windscanner.eua new european distributed research infrastructure [online]. available: http://www.windscanner.eu/ [accessed july 2017]. witze, a. 2017. world's largest wind-mapping project spins up in portugal. nature, 542, 282-283. 1 cristina ribeiro (corresponding author, mcr@fe.up.pt) and gabriel david are associate professors with the informatics engineering department, faculty of engineering, university of porto. joão correia lopes is an assistant professor with the same department. joão rocha da silva, joão aguiar castro and ricardo carvalho amorim are senior researchers at inesc tec. all authors are affiliated with inesc tec. 2 http://www.unidata.ucar.edu/software/netcdf/ https://doi.org/10.29173/iq925 https://doi.org/10.1007/s00799-018-0238-x http://eeagrants.org/project-portal/project/pt02-0017 http://www.snimar.pt/ar/ficheiros/perfilsnimar.pdf http://www.windscanner.eu/ http://www.unidata.ucar.edu/software/netcdf/ microsoft word 49-4-bauder.docx 1/47 bauder, julia & cave, libby (2025). conceptions of data literacy in the statistics education literature, iassist quarterly 49(4), pp. 147. doi: https://doi.org/10.29173/iq1156 the creative commons-attribution-noncommercial license 4.0 international applies to all works published by iassist quarterly. authors will retain copyright of the work and full publishing rights. conceptions of data literacy in the statistics education literature julia bauder1 and libby cave2 abstract data literacy is an increasingly important skill in our data-driven world, and librarians and other information professionals can play a key role in creating a data literate population due to data literacy’s close association with information literacy. however, the definition of data literacy and the attention paid to certain competencies varies greatly between fields: what librarians and statisticians mean by “data literacy” is not the same thing. a scoping review of data literacy articles within the field of statistics education reveals the landscape of data literacy education in statistics, giving librarians and other information professionals a map for coordinating their data literacy work with disciplinary faculty. the areas of data discovery, evaluating and ensuring the quality of data and its sources, and reproducibility are closely examined. these areas are defined and valued inconsistently amongst information professionals and statisticians, but their close associations to traditional library services create an ideal opportunity for libraries and data archives to contribute to data literacy education. keywords data literacy, statistics education, reproducibility, data discovery introduction statistics educators often serve as the primary providers of data literacy education, but there is a disconnect between what statistics educators value in data literacy instruction and what librarians and other information professionals see as foundational data literacy competencies. by reviewing the data literacy literature from journals that publish articles in statistics education, we have found which areas of data literacy are less prioritized in statistics education. these gaps around developing students’ ability to find, critically evaluate, document, and preserve data align closely with the values and duties of librarians, data archivists, and research data management specialists, providing libraries with an opportunity to make substantial contributions to data literacy education. literature review data literacy is still an evolving field, and the exact definition of data literacy has not yet been settled. for example, pinto et al. (2023) found in their systematic review of the literature that discussed both data literacy and information literacy that 45.59% of the 68 included articles provided their own unique definition of data literacy rather than citing a pre-existing definition. 2/47 bauder, julia & cave, libby (2025). conceptions of data literacy in the statistics education literature, iassist quarterly 49(4), pp. 147. doi: https://doi.org/10.29173/iq1156 despite (or perhaps because of) this diversity of definitions, there have been several attempts to establish a consensus definition of data literacy. these efforts have typically involved comparing the competencies that are mentioned in competing definitions of data literacy in search of common themes or areas of overlap. for example, bonikowska, sanmartin and frenette (2019) compared five different data literacy frameworks that were intended for use with broad, general populations of students or working adults. they found twenty-seven different competencies that were mentioned in just those five frameworks. even worse, only five of the competencies appeared in all five frameworks: data discovery, data manipulation, evaluating and ensuring the quality of data and sources, basic data analysis, and data interpretation. extending this line of work, downes (2023) examined twenty different publications that provided lists of competencies that a person needed to achieve to be considered “data literate.” the origins of these twenty publications varied. several were produced by government agencies, such as the australian bureau of statistics and statistics canada; others were written by academic researchers. in those twenty works, downes identified forty different data literacy competencies that were mentioned at least once, and not a single one of those forty competencies appeared in all twenty of the works that he examined. he did, however, find that these competencies tended to cluster into five different data literacy models: the data stewardship model, the analysis and decision-making model, the information literacy model, the science and research data literacy model, and the social engagement model (downes, 2023, p. 109). downes’ (2023) information literacy model for data literacy is an obvious bridge to the library’s domain, and librarians and other information professionals can find strong links from their skillset to the other models. like information literacy, data literacy definitions are conceptualized as both a “specific skill set and a knowledge base, which empowers individuals to transform data into information and into actionable knowledge” (koltay, 2017, p.17). while definitions of data literacy fluctuate in library and information science literature, the most cited data literacy competencies are “access, interpret, critically evaluate, manage, handle, and ethically use data” (pinto molina et al., 2023, p.15). in addition to data literacy’s links to information literacy, librarians and data archivists are also well-situated to assist with data literacy services as they are highly connected to existing library workflows, such as information discovery, dissemination, publication, and subject-specific services (macmillan, 2014). libraries are also experienced in “fostering cross-departmental, crosscampus, etc. communication and collaboration,” which is needed for effective research data management and data literacy education as data needs become ever more interdisciplinary (koltay, 2017, p.8). data discovery bonikowska, sanmartin and frenette (2019) found that two core information literacy skills as they relate to data—data discovery and evaluating and ensuring the quality of data and sources—were two of the five competencies that appeared in all five of the frameworks they investigated. data discovery is the process of finding relevant data to meet a research need (gregory et al., 2018). students competent in data discovery can access data from a range of sources rather than using data they collect themselves or that is directly given to them (ridsdale et al., 2015). however, the term is used inconsistently. librarians use the term in reference to the information-seeking aspect of finding existing data sources, while statistics and math educators often use the term “discovery” to 3/47 bauder, julia & cave, libby (2025). conceptions of data literacy in the statistics education literature, iassist quarterly 49(4), pp. 147. doi: https://doi.org/10.29173/iq1156 refer to finding patterns, actionable insight or other areas of interest in the data at hand (wilson et al., 2021; curley & peterson, 2022; hassad, 2020; roth & temple, 2014). it is also used in the context of the constructivist approach of student-centered discovery within data education (dangol & dasgupta, 2023). the ability to find relevant data is critical to the work of researchers. much like a literature review, data discovery can give researchers an idea of what others in their field have found and highlight gaps in the information landscape. finding data to reuse is more efficient than replicating the data collection process, which isn’t feasible for many researchers. data discovery is also essential for information evaluation, as researchers should be able to trace claims back to the original data, preform their own analysis to confirm claims, and locate and interrogate the accompanying documentation for biases. despite the importance of data discovery to the quantitative research process, previous scholarship has shown that many researchers struggle with this competency. according to sun et al. (2024), researchers often turn to data support specialists for data discovery help for both exploratory searches for new data and for known-item searches. their issues with discovering data lie in a “lack of data search skill, lack of data literacy, and lack of access to data” (sun et al., 2024, p.8). most researchers find datasets through their interpersonal connections or through the data’s citation in text-based sources such as articles (mathiak et al., 2023; million et al., 2024). if that fails, they turn to open web searches (sun et al., 2024). in contrast, data librarians and other data support specialists are adept at data discovery and find that data discovery services make up the bulk of their support interactions. data support specialists are more likely to approach data discovery differently than literature discovery and are more adept at using a variety of sources like search engines, domain repositories, and governmental sources. however, few librarians have the specialized training needed to be confident in their data discovery skills as there is a clear difference in the way data is cataloged, stored, searched for, accessed, and used compared to traditional library materials (huck, 2020; million et al., 2024). evaluation and ensuring quality of data evaluating and ensuring the quality of data is a critical skill for a data literate population and is another common competency across data literacy models and definitions (bonikowska et al., 2019). this skill involves critically considering the trustworthiness of data and its sources, identifying errors in data, evaluating if captured data represents the original information correctly, and determining the quality of data and assessments. data literate people know that when evaluating data, they are evaluating “1) trustworthiness of the measurement 2) trustworthiness of the data processing and 3) trustworthiness of the data integration and visualization” (koedel et al., 2022, p.1). like information evaluation, there are several different data quality models that serve as benchmarks for characteristics of high-quality data. the iso/iec 25012 data quality model, for example, lists and defines 15 characteristics such as accuracy, credibility, currentness, compliance, and traceability (iso/iec 25012, 2008). some data repositories, such as kaggle.com, created data quality assessment scores, but their basis for score calculations are often unclear or evaluate aspects of the dataset that have little bearing on the actual quality of the data. for example, having a cover photo for the dataset increases its quality rating on kaggle.com (chicco et al., 2025). furthermore, these data quality models 4/47 bauder, julia & cave, libby (2025). conceptions of data literacy in the statistics education literature, iassist quarterly 49(4), pp. 147. doi: https://doi.org/10.29173/iq1156 are often dependent on which field of research they were created in, which means that there may be elements that are irrelevant when applied from, for example, soil composition data to pharmacology data. despite these differences, most models agree that high-quality data represents the real-world accurately, and has attributes of “accuracy, correctness, currency, completeness and relevance” (bertino, 2015, p.19). evaluating and ensuring quality of data looks different depending on the role of user. both data creators and data consumers need to actively evaluate and ensure the quality of the data at hand. data creators have several crucial responsibilities: they must ensure the accuracy of their measurements, critically examine their collection process for potential biases, properly address missing data and outliers, reflexively evaluate their own potential biases, and provide comprehensive, "thick description” or context for their data (korstjens & moser, 2018, p. 122). these steps aid in the reproducibility and appropriate reuse of their data. when evaluating data, data consumers must consider the intrinsic data quality and the contextual data quality of the externally produced data (mahanti, 2019). intrinsic data quality considers the elements of the data itself, as discussed above, such as completeness, accuracy and consistency. contextual data quality refers to the users’ own context, such as their research question and purpose for using the data. much of the library literature discusses the quality of data in terms of research data management and working with data creators to better preserve, describe, store, and share their data (giarlo, 2013). evaluating the quality of data and its sources for externally produced data is widely discussed in the library literature, but typically in vague terms. often, library literature will discuss evaluating databased claims or conclusions in media and other sources using information literacy frameworks, such as brungard and smith (2021). while these skills can certainly be applied to data and data sources, they do not necessarily address how to identify the previously discussed data quality models' attributes of high-quality data. other library literature stressed the importance of evaluation of data sources but does not directly state what qualities need to be evaluated (carlson et al., 2015). others state some characteristics of high-quality data as defined by the data quality models, such as arellano douglas et al. (2021), when they provide an evaluation learning objective where “students will critically examine data for accuracy, reliability, bias and context” (p.43). the vagueness about what characterizes high-quality data, and which attributes should be the focus when evaluating data extends outside of the library. sapp nelson (2014) highlighted this in their study in which a faculty member “focused on using critical thinking to evaluate the contents of an externally produced data set for quality”, but “did not describe the actual metrics by which an individual evaluates data quality” (p.232). this provides an opportunity for librarians to collaborate with faculty to better understand and teach data evaluation skills. methodology to create our corpus, we searched 11 journals that were identified by the consortium for the advancement of undergraduate statistics education (cause) as publishing research related to statistics education: technology innovations in statistics education (tise), journal for research in mathematics education (jrme), educational studies in mathematics (esm), mathematical thinking and learning (mtl), international journal of mathematical education in science and technology, international statistical review (isr), the american statistician (tas), mathematics teacher (mt), teaching statistics, journal of statistics and data science education, journal of statistics education, 5/47 bauder, julia & cave, libby (2025). conceptions of data literacy in the statistics education literature, iassist quarterly 49(4), pp. 147. doi: https://doi.org/10.29173/iq1156 and statistics education research journal (cause, n.d.). we decided to limit the search to these journals so as to address the landscape of data literacy specifically in statistics education, and not other disciplines, which may cover different competencies of data literacy. furthermore, articles about data literacy education are not consistently described, making searching a broader corpus difficult and inaccurate. by targeting this curated set of journals that publish articles in statistics education, we also intentionally focus on educational settings that range from primary schools to graduate students. while limiting a scoping review to specific individual journals is not a common practice, it has been done in several fields (maggio et al., 2021; logan et al., 2024; medeiros et al., 2024). we further limited our search to articles published within the last 12 years to evaluate only current literature on the topic. due to the large volume of articles that mention data literacy in passing, we searched for the term in the title, abstract or keywords of the articles. additionally, we experimented with search terms that would include articles that addressed the concept of data literacy education. over 40 different search terms were tested in 12 different search queries in the scopus database, as it indexed all of the identified journals. the results of each search were reviewed, comparing the total number of results and the relevancy of the articles’ titles and abstracts to the research goal and to the other search queries’ results. this review was done by exporting the results to a spreadsheet, manually scanning the results’ titles and abstracts for relevance to our research question, marking the results, and comparing them to the other searches’ results. the final search string, which is reproduced below was chosen for its adequately scoped results that were highly relevant to the research goal. issn ( 1570-1824 or 2693-9169 or 10691898 or 1933-4214 or 0021-8251 or 1573-0816 or 1532-7833 or 1464-5211 or 0306-7734 or 1537-2731 or 1467-9639 ) and pubyear > 2012 and pubyear < 2024 and title-abs-key ( student* or literacy or education or class or classroom or curriculum ) and ( title ( "data literacy" or "statistical literacy" or data ) or key ( "data literacy" or "statistical literacy" or data ) another manual review of the articles was undertaken to remove irrelevant results. articles that focused on purely pedagogical practices, such as articles detailing the effectiveness of flipped classroom course designs, or articles that focused primarily on discussing mathematical concepts, were removed from the corpus as they did not discuss the data literacy competencies that educators wanted students to learn. articles that used simulated data rather than real-world data were also removed from the corpus. simulated data does not provide students with the opportunity to work with several key data literacy competencies, such as evaluating the quality of data and its sources and data collection, and it removes the connection of data from the important real-world contexts. a small number of articles were removed as they either were not research-based articles (e.g., editorials) or their connection to data literacy education was minimal. this left us with a corpus of 260 articles. criteria inclusion exclusion topic data literacy education learning outcomes pedagogical strategy focus; mathematical concepts 6/47 bauder, julia & cave, libby (2025). conceptions of data literacy in the statistics education literature, iassist quarterly 49(4), pp. 147. doi: https://doi.org/10.29173/iq1156 without a data literacy education component; data context use of real-world data use of simulated or synthetic data publication date 2012-2024 articles published outside of the data range the corpus of articles was loaded into the nvivo qualitative data analysis software for analysis. each of the articles was coded to indicate whether the full-text contained any mention of each of the 27 data literacy competencies mentioned in bonikowska, sanmartin and frenette (2019). (see appendix a for a list of those 27 competencies.) if an article contained any indication, even in passing, that the author or authors of the article believed that achieving a given competency was a worthwhile outcome of statistics education, that article was coded as mentioning that competency. results and discussion on our first pass at coding the articles, we identified several codes that could be applied to a nearmajority or a majority of the articles, including basic data analysis, data visualization, data interpretation, and data tools. this is expected, as these competencies are traditionally the focus of data or statistics education. however, interdisciplinary competences were also highly represented. critical thinking was a common theme, which aligns with a shift in statistics education that encourages connecting students to real-world data and its implications (ferris & cheng, 2018; koga, 2022; ben-zvi & garfield, 2004). communication skills were also highly valued, with slightly less than half of the articles mentioning the importance of students’ abilities to present their findings verbally. given the ubiquity of these codes and, in many cases, their lack of a clear connection to the aspects of data literacy that are most relevant to the work of libraries and data archives, we chose not to pursue them further, and instead to focus on our analyses on other data literacy competencies. data discovery data discovery was only mentioned in 30 of the 260 articles (11.53%). of those articles, 12 (4.62% of the overall corpus) mentioned data discovery in passing (i.e., mentioning that students were expected to find an outside data source for an assignment). another 14 articles (5.38%) mentioned data discovery outlined specified data sets or data sources as recommendations for educators to use with their classes. however, 9 of the 35 mentioned sources are no longer widely available for use or are significantly out of date, and another 10 belong to us government agencies that are currently facing mass information suppression. this highlights the issue of providing sources without accompanying skills as sources are prone to disruption, while skills can be more widely applied to various statistical inquiries and can better stand the test of time. only 4 articles (1.53%) in the corpus discussed relevant skills that students would need to facilitate data discovery outside of classroom-provided materials. one of these articles (fergusson & wild, 2021) discussed apis. while articles about apis were typically listed as “data collection” or “data tools” 7/47 bauder, julia & cave, libby (2025). conceptions of data literacy in the statistics education literature, iassist quarterly 49(4), pp. 147. doi: https://doi.org/10.29173/iq1156 in the coding, this article prompted students to find their own data sources using apis and highlighted the importance of combining data from different sources. another article (caballer-tarazona & collserrano, 2020) featured a learning goal that students “become familiar with an official data base and realize that even if data are available, key skills are required to manage the data and extract and understand the available information” (p.309), which is a critical aspect of data discovery. one article (çetinkaya-rundel et al., 2022) provided guidance for student’s data discovery. they stated, “an approach where students are given only guidance, but not a list to pick a dataset from, gives them full control over their project” (p.6), highlighting the importance of data discovery for undertaking the data inquiry process. their first guideline prompts students to think about their questions and what sort of variables and units would be needed in their desired dataset. they go further to state that their question may not have a readily available dataset, so they may need to revisit the question they are asking until they find “a happy medium” (çetinkaya-rundel et al., 2022, p.6). the authors also recommend the services of librarians as they “are helpful in locating data to answer specific questions as well as helping students restate their questions to better match the data available” (çetinkaya-rundel et al., 2022, p.6). this absence of focus on data discovery skills is not surprising when considered in the context of existing literature on data discovery. as discussed above, researchers often turn to data support specialists for data discovery assistance (sun et al. 2024). researchers rely on interpersonal connections or literature searches for their data needs, and this is replicated in the classroom. students and researchers would benefit greatly from librarians’ expertise in strategically searching for data. librarians are well-situated, if not always well-trained, to help with data discovery efforts. data discovery is information-seeking, and as such it requires parallel skills to traditional information discovery, such as source evaluation, search queries, and knowledge of appropriate databases. evaluation and ensuring quality of data in the corpus, 51 articles (19.62%) discussed elements of evaluating and ensuring quality of data. it is important to note that we coded articles that discussed evaluating conclusions, analyses, or claims to the code “evaluating decisions and conclusions based on data,” leaving only articles that discussed the quality of data and datasets. additionally, articles that discussed quality of data in terms of research data management competencies were coded to other competency codes such as data duration and reuse. fourteen of these articles (5.38%) mentioned this competency in passing, vaguely referring to the importance of evaluating data but without any specific discussion of what that entails from either the position of a data creator or a data collector. fifteen articles (5.76%) discussed ensuring the quality of one’s own data, typically in reference to data collection methods (frölich & schellhammer, 2022; zhu et al., 2013), measurement (casleton et al., 2014), variability (roth & temple, 2014), sampling, and evaluating models to see if they accurately represented real-world phenomenon (fleischer et al., 2022). very few of the articles discussed data quality in terms of the data quality models. bilgin et al. (2022) stands out, as they discuss the use of a quality manual and documenting the project in terms of cross industry standard process for data mining (crisp-dm). zhu et al. (2013) did not directly relate back to a specific data quality model, but listed in-depth, systematic quality control measures that would match with previously discussed data quality attributes such as critical examination of their 8/47 bauder, julia & cave, libby (2025). conceptions of data literacy in the statistics education literature, iassist quarterly 49(4), pp. 147. doi: https://doi.org/10.29173/iq1156 collection process for potential biases, properly addressing missing data and outliers, precision in measurements, and detail reporting of context. the remaining 26 articles (10%) discussed evaluating the quality of externally produced data. most of these articles focused on asking critical questions of the data such as who created the data, how they gathered it, and its original purpose, such as delport (2023) with their use of “worry questions,” and lee et al. (2022) with their discussion of the issue of bias, both in what is represented in the data and what is not represented in the data. this focus on bias relates to a broader competency of critical thinking, which was far better represented in the data literacy literature, with 138 (53.01%) articles covering the topic. while evaluating and ensuring quality of data and sources certainly requires a level of critical thinking, it is more specific in the goals of the critical questioning than critical thinking alone does. it is not enough to ask the questions like, “who made this?” and “what sort of biases could be present?.” to truly evaluate the quality of data and its sources, one must be able to find the answers to those questions and, potentially, know how to compensate for the weaknesses in data or its sources. however, directions for asking the questions are rarely followed with instructions on how to find the answers to such questions. few of the articles coded as “evaluating and ensuring quality of data and sources,” discussed how students could find information to answer the questions they were asking of the data source outside of an accompanying data dictionary. while a data dictionary would be important for this task, much of the available data does not come with a data dictionary or codebook. lee et al (2022) suggests that it may “be necessary to reach out external stakeholders or experts to find additional information about the data context” (p.14) to answer questions about the quality of data. other articles suggested comparing findings with other sources, which is an essential skill in both data literacy and information literacy. however, there is a gap in data discovery skills, as previously discussed. without strong data discovery skills, it would be difficult to find an alternative data source that matched the important features (i.e., research focus, method, measurements, categories, etc.) sufficiently to make a meaningful comparison. besides these suggestions, none of the articles prompted students to do their own research on the source of the data, instead relying on educators or the data providers themselves to provide all the necessary context for the data. overall, the articles in the corpus stressed the importance of ensuring and evaluating the quality of found data but did not discuss how to teach practical skills for doing so. çetinkaya-rundel et al. (2022) provide the most practical directions for evaluating found datasets in their guidelines for studentselected datasets. they set basic standards in terms of number of variables and observations that need to be present in a dataset to account for confounding variables. they also require that students use data that includes a “comprehensive” data dictionary, stating, “without these, it is impossible for students to evaluate the reliability, validity, and ethical considerations of the data for their projects” (çetinkaya-rundel et al., 2022, p. 6). the article warns students to be selective about using data from aggregators, citing concerns over varying states of documentation for and the use of sample data analyses in the datasets. still, this article, like the others in the corpus, did not discuss the characteristics of high-quality found data that are mentioned in data quality standards, such as consistent, unambiguous, and current (mahanti, 2019; iso/iec, 2008). some mention these characteristics in passing (e.g., jones, 2020), but there is not a clear discussion of how students can 9/47 bauder, julia & cave, libby (2025). conceptions of data literacy in the statistics education literature, iassist quarterly 49(4), pp. 147. doi: https://doi.org/10.29173/iq1156 spot the presence or lack thereof of these characteristics. while librarians can apply information evaluation skills to assist in this area, some consensus would be needed on what metrics and attributes qualify data as high-quality. reproducibility: a disconnect between librarians and statisticians reproducibility and reproducible research workflows were common themes in these articles. in the included articles, reproducibility typically means computational reproducibility: “a reproducible analysis is one that can be rerun (potentially years later, or by a different person) with the same data to produce exactly the same result” (mcnamara, 2019, p. 381). while computational reproducibility is important, it is not the only form of reproducibility. some other fields strongly emphasize other forms of reproducibility, even to the point of defining “reproducibility” differently. for example, plesser (2018) compiled definitions of reproducibility and the related concepts of “repeatability” and “replicability” from a number of scientific fields, including geophysics, chemistry and computer science. as plesser noted, the idea of reproducibility inherent in “computational reproducibility” “is at odds with the terminology long established in experimental sciences” (plesser 2018, 1), which uses “repeatability” to describe the condition where the same procedure run on the same equipment under the same conditions produces the same results. in those fields, “reproducibility” means the ability of a different research group, using different equipment, to produce the same results. to make the matter even more confusing, some fields use an additional term, “replicability,” which, depending on the field, can mean something closer to “computational reproducibility” or something closer to the experimental sciences’ definition of “reproducibility” (plesser 2018). existing models for data literacy are more closely aligned with the experimental sciences’ definition of “reproducibility” than with the concept of computational reproducibility. no competencies that directly map to the concept of computational reproducibility appear in either of the compilations of data literacy competencies that were discussed in the literature review of this article (downes 2023 and bonikowska et al. 2019). instead, the data literacy models that were compiled in these two reviews emphasize competencies such as data curation, data preservation, and data sharing, which are necessary to allow for research results to be reproduced by other research groups. yet very few of the articles included in this analysis mentioned the data literacy competencies such as these that would allow researchers to make sure that other researchers outside of their own circles, or that the researchers themselves in the long-range future, could access and use the data necessary to reproduce their research. for example, data preservation (ensuring that the data is preserved in its original state for the duration of the student’s or researcher’s period of analysis) and data duration and reuse (ensuring that data is preserved in its original state indefinitely, and in such a way that it can be made available to and reused by other researchers) are rarely discussed explicitly using those terms, although both aspects are inherent in the experimental sciences’ definition of reproducible research. for the purposes of this project, we assumed that, unless the context clearly indicated otherwise, any mention of “reproducibility” in these articles included data preservation as one of the implicit goals. this assumption is due to the typical structure of the projects discussed in the articles, where students were expected to turn in their original data files, along with any code that was used to manipulate or analyze the data, allowing the instructor to reproduce the students’ entire data processing and 10/47 bauder, julia & cave, libby (2025). conceptions of data literacy in the statistics education literature, iassist quarterly 49(4), pp. 1-47. doi: https://doi.org/10.29173/iq1156 analysis procedure. with this generous definition of data preservation, 66 articles (25.38%) were coded as mentioning data preservation. however, these articles rarely explicitly call out “not altering the original data file” as one of the benefits of a reproducible analysis. also, some of the articles discuss reproducible research workflows—being able to run the exact same analysis steps on a different dataset, rather than on the same dataset. where it was clear that reproducible workflows rather than reproducible analyses were being discussed, the article was not coded as “data preservation.” given all of these caveats, it is difficult to say precisely how frequently students are explicitly being taught about the importance of preserving a copy of their original data file. conversely, we did not assume that “reproducibility” referred to data duration and reuse and data sharing unless those competencies were explicitly mentioned. this assumption contributed to a much lower number of articles being coded to the “data duration and reuse” and “data sharing” competencies: 16 (6.15%) and 10 (3.85%), respectively. the concept of metadata creation also appeared infrequently, in 11 articles (4.23%), although typically in the guise of “codebooks” and “data dictionaries”: forms of metadata that are very useful for accurately interpreting a given dataset, but that are less helpful for dataset discovery. data citation was mentioned in only three articles (1.15%), and always in passing. a lack of data citation skills is likely to hamper reproducibility, given that many datasets that may be used in research have contractual or ethical restrictions that prevent them from being freely redistributed by the researchers who use them. if students do not learn how to cite data properly, people in the future who wish to replicate their analysis may be unable to identify and access the specific dataset that was used. similarly, data sharing was mentioned in only 10 articles (3.85%), again usually without any sort of detail about issues to be considered or specific tasks to be completed in order to share data ethically and effectively. a good example of a typical discussion of data sharing in this corpus can be found in donoghue, voytek, and ellis’s article “teaching creative and practical data science at scale,” (2021), which focuses on the skills that students need to learn to be effective data science professionals (see pp. s33-s34). the authors list many specific practices and skills that students must master to be able to carry out reproducible research, but most of these skills are related to the code that implements the analysis rather than the data that is being analyzed. the code must be “well-organized, documented, and tested;” it must be “understandable by other analysts.” the data, however, only needs to be “stor[ed] . . . in a consistent manner.” all of the other skills and practices necessary for ethical and effective data sharing go unmentioned. conclusion although this study, which only examined published journal articles, demonstrated which data literacy competencies are well-covered in the statistics education literature and which are not, it does not and cannot explain why the missing competencies are not covered. future research in this area should draw on a wider range of sources, in particular conversations with statistics educators about their views of the missing data literacy competencies. do statistics educators not think these competencies are important? do they not feel equipped to teach them themselves? do they believe that they are 11/47 bauder, julia & cave, libby (2025). conceptions of data literacy in the statistics education literature, iassist quarterly 49(4), pp. 1-47. doi: https://doi.org/10.29173/iq1156 being covered in classes outside of statistics? without answers to questions such as these, it is not clear how librarians and other information professionals can best position ourselves to help. data librarians and research data management specialists are well-equipped to teach the data literacy skills that are apparently not being covered in statistics education, as these skills are of deep interest to these professions, but becoming empowered to teach these skills to students in statistics classes will require collaboration with statistics educators and a richer understanding of statistics educators’ perspectives on data literacy. references list ben-zvi, d., & garfield, j. (eds.). (2004). the challenge of developing statistical literacy, reasoning and thinking. springer netherlands. https://doi.org/10.1007/1-4020-2278-6 bertino, e. (2015). data trustworthiness—approaches and research challenges. in data privacy management, autonomous spontaneous security, and security assurance (vol. 8872, pp. 17–25). springer international publishing ag. https://doi.org/10.1007/978-3-319-17016-9_2 bilgin, a. a. b., powell, a., & richards, d. (2022). work integrated learning in data science and a proposed assessment framework. statistics education research journal, 21(2), 1-. https://doi.org/10.52041/serj.v21i2.26 bonikowska, a., sanmartin, c., & frenette, m. (2019, august 14). data literacy: what is it and how to measure it in the public service. https://www150.statcan.gc.ca/n1/en/pub/11-633x/11-633-x2019003-eng.pdf brungard, a., & smith, l. (2021). a data discovery project: seeking truth in a post-truth world. in j. bauder (eds.), data literacy in academic libraries. american library association. carlson, j., fosmire, m., miller, c. c., & nelson, m. s. (2011). determining data information literacy needs: a study of students and research faculty. portal: libraries and the academy, 11(2), 629-657. http://dx.doi.org/10.1353/pla.2011.0022 caballer-tarazona, m., & coll-serrano, v. (2020). the raising factor, that great unknown. a guided activity for undergraduate students. journal of statistics education, 28(3), 304–315. https://doi.org/10.1080/10691898.2020.1832006 casleton, e., beyler, a., genschel, u., & wilson, a. (2014). a pilot study teaching metrology in an introductory statistics course. journal of statistics education, 22(3), 1. https://doi.org/10.1080/10691898.2014.11889710 çetinkaya-rundel, m., dogucu, m., & rummerfield, w. (2022). the 5ws and 1h of term projects in the introductory data science classroom. statistics education research journal, 21(2), 1– 19. https://doi.org/10.52041/serj.v21i2.37 12/47 bauder, julia & cave, libby (2025). conceptions of data literacy in the statistics education literature, iassist quarterly 49(4), pp. 1-47. doi: https://doi.org/10.29173/iq1156 chicco, d., fabris, a., & jurman, g. (2025). the venus score for the assessment of the quality and trustworthiness of biomedical datasets. biodata mining, 18(1), 1–31. https://doi.org/10.1186/s13040-024-00412-x consortium for the advancement of undergraduate statistics education (cause) (n.d.). journals publishing research in statistics education. https://www.causeweb.org/cause/research/journals curley, b., & peterson, a. (2022). a fresh shot at statistics in the classroom: three perspectives using world cup soccer player data. journal of statistics and data science education, 30(1), 86–98. https://doi.org/10.1080/26939169.2021.2008283 dangol, a., & dasgupta, s. (2023). constructionist approaches to critical data literacy: a review. proceedings of the 22nd annual acm interaction design and children conference, 112– 123. https://doi.org/10.1145/3585088.3589367 delport, d. h. (2023). the development of statistical literacy among students: analyzing messages in media articles with gal’s worry questions. teaching statistics, 45(2), 61–68. https://doi.org/10.1111/test.12308 donoghue, t., voytek, b., & ellis, s. e. (2021). teaching creative and practical data science at scale. journal of statistics and data science education, 29(s1), s27–s39. https://doi.org/10.1080/10691898.2020.1860725 downes, s. (2023). three frameworks for data literacy. in d. g. sampson, d. ifenthaler, d., & p. isaías (eds.), proceedings of the 20th international conference on cognition and exploratory learning in the digital age (107-115). iadis press. https://files.eric.ed.gov/fulltext/ed636095.pdf ferris, m., & cheng, s. (2018). using twitter to energize the introductory statistics class. technology innovations in statistics education, 11(1). https://doi.org/10.5070/t5111032036 fleischer, y., biehler, r., & schulte, c. (2022). teaching and learning data-driven machine learning with educationally designed jupyter notebooks. statistics education research journal, 21(2), 1–25. https://doi.org/10.52041/serj.v21i2.61 frölich, n., & schellhammer, k. s. (2022). questionnaire design and sampling procedures for business and economics students: a research-oriented, hands-on course. international journal of mathematical education in science and technology, 0(0), 1–19. https://doi.org/10.1080/0020739x.2022.2056722 giarlo, m. j. (2013). academic libraries as data quality hubs. journal of librarianship and scholarly communication, 1(3). https://doi.org/10.7710/2162-3309.1059 gregory, k., khalsa, s. j., michener, w. k., psomopoulos, f. e., de waard, a., & wu, m. (2018). eleven quick tips for finding research data. plos computational biology, 14(4). https://doi.org/10.1371/journal.pcbi.1006038 13/47 bauder, julia & cave, libby (2025). conceptions of data literacy in the statistics education literature, iassist quarterly 49(4), pp. 1-47. doi: https://doi.org/10.29173/iq1156 hassad, r. a. (2020). a foundation for inductive reasoning in harnessing the potential of big data. statistics education research journal, 19(1), 238–258. https://doi.org/10.52041/serj.v19i1.133 huck, j. (2020). identifying, accessing and evaluating data: finding and accessing data can be problematic, but many of the skills used in traditional reference can be applied to data discovery. information outlook, 24(1), 4-6. https://scholarworks.sjsu.edu/sla_io_2020/1 iso/iec. (2008). software engineering—software product quality requirements and evaluation (square)—data quality model (25012:2008). https://www.iso.org/standard/35736.html jones, j. d. (2022). using school mathematics to develop students’ data literacy skills. mathematics teacher: learning and teaching pk-12, 115(8), 576–581. https://doi.org/10.5951/mtlt.2021.0239 koedel, u., schuetze, c., fischer, p., bussmann, i., sauer, p. k., nixdorf, e., kalbacher, t., wichert, v., rechid, d., bouwer, l. m., & dietrich, p. (2022). challenges in the evaluation of observational data trustworthiness from a data producers viewpoint (fair+). frontiers in environmental science, 9. https://doi.org/10.3389/fenvs.2021.772666 koga, s. (2022). characteristics of statistical literacy skills from the perspective of critical thinking. teaching statistics, 44(2), 59–67. https://doi.org/10.1111/test.12302 korstjens, i., & moser, a. (2018). series: practical guidance to qualitative research. part 4: trustworthiness and publishing. the european journal of general practice, 24(1), 120– 124. https://doi.org/10.1080/13814788.2017.1375092 lee, h. s., mojica, g. f., thrasher, e. p., & baumgartner, p. (2022). investigating data like a data scientist: key practices and processes. statistics education research journal, 21(2), 1– 23. https://doi.org/10.52041/serj.v21i2.41 logan, j., webb, j., singh, n. k., tanner, n., barrett, k., wall, m., walsh, b., & ayala, a. p. (2024). scoping review search practices in the social sciences: a scoping review. research synthesis methods, 15(6), 950–963. https://doi.org/10.1002/jrsm.1742 mahanti, r. (2019). data quality: dimensions, measurement, strategy, management, and governance. quality press. http://ebookcentral.proquest.com/lib/grinnellebooks/detail.action?docid=6262212 maggio, l. a., larsen, k., thomas, a., costello, j. a., & artino jr., a. r. (2021). scoping reviews in medical education: a scoping review. medical education, 55(6), 689–700. https://doi.org/10.1111/medu.14431 mathiak, b., juty, n., bardi, a., colomb, j., & kraker, p. (2023). what are researchers’ needs in data discovery? analysis and ranking of a large-scale collection of crowdsourced use cases. data science journal, 22(1). https://doi.org/10.5334/dsj-2023-003 14/47 bauder, julia & cave, libby (2025). conceptions of data literacy in the statistics education literature, iassist quarterly 49(4), pp. 1-47. doi: https://doi.org/10.29173/iq1156 mcnamara, a. (2019). key attributes of a modern statistical computing tool. the american statistician, 73(4), 375–384. https://doi.org/10.1080/00031305.2018.1482784 medeiros, p., shetty, j., lamaj, l., cunningham, j., wanigaratne, s., guttmann, a., & cohen, e. (2024). reported community engagement in health equity research published in highimpact medical journals: a scoping review. https://doi.org/10.1136/bmjopen-2024084952 million, a. j., york, j., lafia, s., & hemphill, l. (2024). data, not documents: moving beyond theories of information-seeking behavior to advance data discovery. journal of the association for information science and technology. https://doi.org/10.1002/asi.24962 plesser, h.e. (2018). reproducibility vs. replicability: a brief history of a confused terminology. frontiers in neuroinformatics 11 (76). https://doi.org/10.3389/fninf.2017.00076 prado, j.c., & marzal, m.a. (2013). incorporating data literacy into information literacy programs: core competencies and contents. libri 63 (2): 123–134. https://doi.org/10.1515/libri2013-0010 roth, w.-m., & temple, s. (2014). on understanding variability in data: a study of graph interpretation in an advanced experimental biology laboratory. educational studies in mathematics, 86(3), 359–376. https://doi.org/10.1007/s10649-014-9535-5 schield, m. (2004). information literacy, statistical literacy and data literacy. iassist quarterly summer/fall: 6–11. https://doi.org/10.29173/iq790 sun, g., friedrich, t., gregory, k., & mathiak, b. (2024). supporting data discovery: comparing perspectives of support specialists and researchers. data science journal, 23(1). https://doi.org/10.5334/dsj-2024-048 towse, j., davies, r., ball, e., james, r., gooding, b., & ivory, m. (2022). lustre: an online data management and student project resource. journal of statistics and data science education, 30(3), 266–273. https://doi.org/10.1080/26939169.2022.2118645 wilkerson, m. h., lanouette, k., & shareff, r. l. (2022). exploring variability during data preparation: a way to connect data, chance, and context when working with complex public datasets. mathematical thinking and learning, 24(4), 312–330. https://doi.org/10.1080/10986065.2021.1922838 wilson, m., ross, a., & casey, s. (2021). a classroom-ready activity on educational disparities in the united states. teaching statistics, 43(s1), s93–s97. https://doi.org/10.1111/test.12252 zhu, y., hernandez, l. m., mueller, p., dong, y., & forman, m. r. (2013). data acquisition and preprocessing in studies on humans: what is not taught in statistics classes? the american statistician, 67(4), 235–241. https://doi.org/10.1080/00031305.2013.842498 15/47 bauder, julia & cave, libby (2025). conceptions of data literacy in the statistics education literature, iassist quarterly 49(4), pp. 1-47. doi: https://doi.org/10.29173/iq1156 appendix a: data literacy definitions the codes used for this review were collected from bonikowska et al., 2019, which compared competencies of data literacy found in data to the people (2018) 1 ; grillenberger & romeike (2018) 2 ; ridsdale et al. (2015)3; sternkopf & mueller (2018)4; and wolff et al. (2016)5. below are the definitions of each competency that we synthesized from the five articles. competencies definitions basic data analysis (select appropriate tool/algorithms/analysis methods for data, knowledge and use of basic summary/descriptive statistics) basic data analysis involves developing and executing plans to examine data using appropriate tools, algorithms and analysis methods including descriptive statistics, hypothesis testing, linear regression, etc. 1,2,3,4,5 critical thinking (aware of high-level issues associated with data, thinks critically when working with data) critical thinking involves being aware of high-level issues and challenges associated with data while applying thoughtful consideration when working with it. 2,3 data culture (psychological barriers, attitudes, etc towards data) data culture refers to the recognition of data's importance and the fostering of an environment that promotes critical use of data for learning, research, and decision-making. it involves overcoming psychological barriers related to data, understanding its potential as an enabler for progress, and securing support from management for data initiatives and resources. 3,4 data collection (gathering data, structure gathered data, critically evaluate the collection process) data collection encompasses the process of gathering information in various formats and complexities to support specific needs. it involves selecting appropriate methods, implementing algorithms, and considering ethical issues and privacy impacts. 1,2,3,5 data conversion (from format to format) data conversion is the ability to transform data from one format or file type to another, requiring knowledge of different data types and conversion methods. 1,3,4,5 data ethics (security, privacy issues) data ethics involves understanding and addressing the moral implications of collecting, analyzing, and using data. it requires considering privacy concerns, potential biases, and the societal impact of data-driven decisions. advanced practitioners can develop ethical frameworks, guide others in ethical data practices, and advocate for responsible data use within organizations. 2,3,4,5 16/47 bauder, julia & cave, libby (2025). conceptions of data literacy in the statistics education literature, iassist quarterly 49(4), pp. 1-47. doi: https://doi.org/10.29173/iq1156 data discovery (ability to find and access data, connect data from different sources, identify useful data) data discovery is the ability to find, access, and identify relevant data from various sources. it progresses from using basic search engines to understanding and selecting from a wide range of data sources, including specialized data portals. advanced skills include assisting others in locating data and formulating assessment criteria for selecting the most relevant data sources for specific informational needs. 1,2,3,4,5 data driven decision making (prioritizes information garnered from data, converts data into actionable information weighs the merit and impacts of possible solutions/decisions, implements decisions/solutions) data-driven decision making (dddm) is the process of using data to inform and guide strategic choices. it involves analyzing relevant data, converting it into actionable insights, and weighing potential outcomes to make informed decisions. those skilled in dddm can also communicate and defend their data-based decisions. 1,3,5 data duration and reuse (structure data in suitable way for storage and other's re-use, curation requirements) data duration and re-use refers to the process of structuring and storing data in a way that facilitates long-term preservation and future utilization by others. this competency involves assessing curation requirements, implementing appropriate storage methods, and ensuring data accessibility while considering ethical and security concerns. 1,2,3 data interpretation (understanding data, read and understand charts & tables, find key points and relationships in data) data interpretation is the ability to understand and extract meaning from data outputs such as analyses and visualizations. it involves identifying key points of interest, recognizing relationships within data, and critically assessing the implications of data outputs. 1,2,3,4,5 data management and organization (store and organize data appropriately for the analysis) data management covers the practices of organizing and storing data for the length of the analysis process. 1,2,3 data manipulation (data cleaning, knowledge that most data is not clean, combine data, decide when it is appropriate to combine data manipulation involves transforming, cleaning, and restructuring data to make it suitable for analysis. at a basic level, data users know that most data are not clean and that cleaning the data is necessary for analysis. skills range from basic sorting and filtering to advanced techniques like appropriately removing outliers and anomalies and deciding when it is appropriate to combine data. 1,2,3,4,5 17/47 bauder, julia & cave, libby (2025). conceptions of data literacy in the statistics education literature, iassist quarterly 49(4), pp. 1-47. doi: https://doi.org/10.29173/iq1156 data, remove outliers and anomalies) data preservation (decide which data to keep or delete, identify appropriate ways to store data) data preservation encompasses determining which data to retain, who should have access, how to ensure the integrity of the data, and how to ethically handle data deletion. effective data preservation ensures data validity, addresses ethical considerations, and maintains data accessibility over time. 2,3 data sharing (decide whom to share data with from a legal or ethical perspective, prepare data for sharing) data sharing involves the practice of making data available to others, both within and outside an organization. it requires understanding of data formats, sharing platforms, and relevant legal and ethical considerations. 2,3 data tools (knowledge of and ability to use tools to collect, store, clean, organize, or analyze data, the ability to choose appropriate tool for the task) data tools are software applications and techniques used for gathering, structuring, and analyzing data. they encompass a range of functionalities, from selecting suitable sensors for data collection and implementing algorithms to download data from web apis, to applying various analysis techniques and visualization methods. those literate in data tools have the ability to choose the appropriate tool for the task at hand. 2,3,5 data visualization (create meaningful graphs and charts, choose appropriate visualizations for the data or analysis) data visualization is the skill of creating meaningful graphical representation of data to facilitate understanding and insight generation. those literate in data visualization not only make charts and graphs but also know the appropriate type of visualization for the data at hand. 1,2,3,4 data citation (knowledge of widely accepted data citation methods creates correct citations for secondary data sets) data citation is the knowledge and application of widely-accepted methods for crediting secondary data sets. it involves creating correct citations for data sources, ensuring proper attribution and enabling others to locate and verify the data used in analyses or research. 3 develop hypotheses (ask questions that can be tested, predict the outcomes of data analysis) the ability to develop hypotheses includes being able to ask questions that can be tested and make informed predictions on the outcome of the data analysis. 5 evaluating and ensuring quality of data and sources (critically consider the evaluating and ensuring quality of data and its sources involves critically assessing the trustworthiness, accuracy, and reliability of data and its origins. this process ranges from identifying errors or problems in datasets to critically evaluating data collection methods, sources, and 18/47 bauder, julia & cave, libby (2025). conceptions of data literacy in the statistics education literature, iassist quarterly 49(4), pp. 1-47. doi: https://doi.org/10.29173/iq1156 trustworthiness of data, identify errors in data, evaluate if captured data represents original information correctly, quality assessment and data checking) potential biases. it also encompasses the ability to verify data quality through multiple layers of checking and connect data from different sources, ultimately ensuring that the data used is representative and suitable for analysis and decision-making purposes. 1,2,3,4,5 evaluating decisions/conclusions based on data (collect follow-up data, compare with other findings) evaluating decisions/conclusions based on data involves assessing the effectiveness of actions or solutions by analyzing follow-up data and comparing results with other findings. this process includes collecting relevant data from various sources, conducting thorough analysis, and using the insights gained to either validate original conclusions or implement new decisions. the ultimate goal is to ensure that decisions are continually refined and improved based on empirical evidence, fostering a culture of data-driven decision-making and continuous improvement within organizations. 1,3 identifying problems using data (knowledge of which questions can be answered by the data) data literate people should be able to identify and describe problems in practical situations using a range of data sources. questions should be formulated precisely and target-orientated to find meaningful answers. 1,3,4,5 metadata creation and use (apply metadata to datasets such as descriptors) metadata creation and use refers to understanding of what metadata associated with data sources are, why such descriptors are important, and the ability to create and assign appropriate metadata descriptors to original data sources. 1,3 presenting data verbally (data storytelling, describing key findings, communicate with others about findings) presenting data verbally involves clearly and coherently describing key insights, datasets, and visualizations in a way that aligns with the audience's needs and familiarity with the subject. this skill progresses from explaining simple data points to effectively using narratives, visualizations, and storytelling to communicate complex data in broader contexts. 1,3,4 undertake data inquiry process undertaking a data inquiry process involves a systematic approach to exploring and analyzing data to answer specific questions or solve problems. it is often done through the ppdac model: problem, plan, data, analysis and conclusion (wolff et al., 2016).5 work with large data sets working with large data sets refers to the ability to effectively handle, process, and analyze substantial volumes of data. this can include datasets that are voluptuous and complex like big data, datasets that 19/47 bauder, julia & cave, libby (2025). conceptions of data literacy in the statistics education literature, iassist quarterly 49(4), pp. 1-47. doi: https://doi.org/10.29173/iq1156 require multiple tools to handle effectively, or that combine diverse data types. 5 knowledge and understanding of data, its uses and applications data literacy involves understanding the nature of data, its various forms, and how it is produced. a data-literate individual should be aware of data's role and influence in society across diverse settings, as well as the ethical considerations associated with its use. 20/47 bauder, julia & cave, libby (2025). conceptions of data literacy in the statistics education literature, iassist quarterly 49(4), pp. 1-47. doi: https://doi.org/10.29173/iq1156 appendix b: article corpus abel, t., & poling, l. (2015). hold my calls: an activity for introducing the statistical process. teaching statistics, 37(3), 96–103. https://doi.org/10.1111/test.12082 ainley, j., gould, r., & pratt, d. (2015). learning to reason from samples: commentary from the perspectives of task design and the emergence of “big data.” educational studies in mathematics, 88(3), 405–412. https://doi.org/10.1007/s10649-015-9592-4 ainley, j., & pratt, d. (2017). computational modelling and children’s expressions of signal and noise. statistics education research journal, 16(2), 15–36. https://doi.org/10.52041/serj.v16i2.183 amdat, w. c. (2021). the chicago hardship index: an introduction to urban inequity. journal of statistics and data science education, 29(3), 328–336. https://doi.org/10.1080/26939169.2021.1994489 aridor, k., & ben-zvi, d. (2017). the co-emergence of aggregate and modelling reasoning. statistics education research journal, 16(2), 38–63. https://doi.org/10.52041/serj.v16i2.184 arnold, p. (2017). statistical literacy in public debate—examples from the uk 2015 general election. statistics education research journal, 16(1), 217–227. https://doi.org/10.52041/serj.v16i1.225 baglin, j., & da costa, c. (2013). comparing training approaches for technological skill development in introductory statistics courses. technology innovations in statistics education, 7(1). https://doi.org/10.5070/t571014007 baldi, b., & utts, j. (2015). what your future doctor should know about statistics: must-include topics for introductory undergraduate biostatistics. the american statistician, 69(3), 231– 240. https://doi.org/10.1080/00031305.2015.1048903 bargagliotti, a., arnold, p., & franklin, c. (2021). gaise ii: bringing data into classrooms. mathematics teacher: learning and teaching pk-12, 114(6), 424–435. https://doi.org/10.5951/mtlt.2020.0343 bargagliotti, a., binder, w., blakesley, l., eusufzai, z., fitzpatrick, b., ford, m., huchting, k., larson, s., miric, n., rovetti, r., seal, k., & zachariah, t. (2020). undergraduate learning outcomes for achieving data acumen. journal of statistics education, 28(2), 197–211. https://doi.org/10.1080/10691898.2020.1776653 bargagliotti, a. e., & anderson, c. r. (2017). using learning trajectories for teacher learning to structure professional development. mathematical thinking and learning, 19(4), 237–259. https://doi.org/10.1080/10986065.2017.1365222 baumer, b. (2015). a data science course for undergraduates: thinking with data. the american statistician, 69(4), 334–342. https://doi.org/10.1080/00031305.2015.1081105 21/47 bauder, julia & cave, libby (2025). conceptions of data literacy in the statistics education literature, iassist quarterly 49(4), pp. 1-47. doi: https://doi.org/10.29173/iq1156 baumer, b., cetinkaya-rundel, m., bray, a., loi, l., & horton, n. j. (2014). r markdown: integrating a reproducible analysis tool into introductory statistics. technology innovations in statistics education, 8(1). https://doi.org/10.5070/t581020118 baumer, b. s. (2018). lessons from between the white lines for isolated data scientists. the american statistician, 72(1), 66–71. https://doi.org/10.1080/00031305.2017.1375985 baumer, b. s., garcia, r. l., kim, a. y., kinnaird, k. m., & ott, m. q. (2022). integrating data science ethics into an undergraduate major: a case study. journal of statistics and data science education, 30(1), 15–28. https://doi.org/10.1080/26939169.2022.2038041 beckman, m. d., çetinkaya-rundel, m., horton, n. j., rundel, c. w., sullivan, a. j., & tackett, m. (2021). implementing version control with git and github as a learning objective in statistics and data science courses. journal of statistics and data science education, 29(s1), s132–s144. https://doi.org/10.1080/10691898.2020.1848485 benakli, n., kostadinov, b., satyanarayana, a., & singh, s. (2017). introducing computational thinking through hands-on projects using r with applications to calculus, probability and data analysis. international journal of mathematical education in science and technology, 48(3), 393–427. https://doi.org/10.1080/0020739x.2016.1254296 berg, a., & hawila, n. (2021). some teaching resources using r with illustrative examples exploring covid-19 data. teaching statistics, 43(s1), s98–s109. https://doi.org/10.1111/test.12258 biehler, r., & fleischer, y. (2021). introducing students to machine learning with decision trees using codap and jupyter notebooks. teaching statistics, 43, s133–s142. https://doi.org/10.1111/test.12279 biehler, r., frischemeier, d., & podworny, s. (2017). elementary preservice teachers´ reasoning about modeling a “family factory” with tinkerplots – a pilot study. statistics education research journal, 16(2), 244–286. https://doi.org/10.52041/serj.v16i2.192 bilgin, a. a. b., date-huxtable, e., coady, c., geiger, v., cavanagh, m., mulligan, j., & petocz, p. (2017). opening real science: evaluation of an online module on statistical literacy for preservice primary teachers. statistics education research journal, 16(1), 120–138. https://doi.org/10.52041/serj.v16i1.220 bilgin, a. a. b., powell, a., & richards, d. (2022). work integrated learning in data science and a proposed assessment framework. statistics education research journal, 21(2), 1-. https://doi.org/10.52041/serj.v21i2.26 boehm, f. j., & hanlon, b. m. (2021). what is happening on twitter? a framework for student research projects with tweets. journal of statistics and data science education, 29(s1), s95–s102. https://doi.org/10.1080/10691898.2020.1848486 22/47 bauder, julia & cave, libby (2025). conceptions of data literacy in the statistics education literature, iassist quarterly 49(4), pp. 1-47. doi: https://doi.org/10.29173/iq1156 boenig-liptsin, m., tanweer, a., & edmundson, a. (2022). data science ethos lifecycle: interplay of ethical thinking and data science practice. journal of statistics and data science education, 30(3), 228–240. https://doi.org/10.1080/26939169.2022.2089411 bolch, c. a., & crippen, k. j. (2022). data scientists’ epistemic thinking for creating and interpreting visualizations. statistics education research journal, 21(2), 1–25. https://doi.org/10.52041/serj.v21i2.21 bradley, s. (2015). handwriting and gender: a multi-use data set. journal of statistics education, 23(1), 1. https://doi.org/10.1080/10691898.2015.11889721 brearley, a. m., bigelow, c., poisson, l. m., grambow, s. c., & nowacki, a. s. (2018). the tshs resources portal: a source of real and relevant data for teaching statistics in the health sciences. technology innovations in statistics education, 11(1). https://doi.org/10.5070/t5111034506 broatch, j. e., dietrich, s., & goelman, d. (2019). introducing data science techniques by connecting database concepts and dplyr. journal of statistics education, 27(3), 147–153. https://doi.org/10.1080/10691898.2019.1647768 brown, m. (2017). making students part of the dataset: a model for statistical enquiry in social issues. teaching statistics, 39(3), 79–83. https://doi.org/10.1111/test.12131 budgett, s., pfannkuch, m., regan, m., & wild, c. j. (2013). dynamic visualizations and the randomization test. technology innovations in statistics education, 7(2). https://doi.org/10.5070/t572013889 budgett, s., & rose, d. (2017). developing statistical literacy in the final school year. statistics education research journal, 16(1), 139–162. https://doi.org/10.52041/serj.v16i1.221 burckhardt, p., nugent, r., & genovese, c. r. (2021). teaching statistical concepts and modern data analysis with a computing-integrated learning environment. journal of statistics and data science education, 29(s1), s61–s73. https://doi.org/10.1080/10691898.2020.1854637 burr, w., chevalier, f., collins, c., gibbs, a. l., ng, r., & wild, c. j. (2021). computational skills by stealth in introductory data science teaching. teaching statistics, 43(s1), s34–s51. https://doi.org/10.1111/test.12277 büscher, c. (2022). design principles for developing statistical literacy in middle schools. statistics education research journal, 21(1), 1–16. https://doi.org/10.52041/serj.v21i1.80 caballer-tarazona, m., & coll-serrano, v. (2020). the raising factor, that great unknown. a guided activity for undergraduate students. journal of statistics education, 28(3), 304–315. https://doi.org/10.1080/10691898.2020.1832006 23/47 bauder, julia & cave, libby (2025). conceptions of data literacy in the statistics education literature, iassist quarterly 49(4), pp. 1-47. doi: https://doi.org/10.29173/iq1156 callingham, r. (2011). assessing statistical understanding in middle schools: emerging issues in a technology-rich environment. technology innovations in statistics education, 5(1). https://doi.org/10.5070/t551000044 callingham, r., & watson, j. m. (2017). the development of statistical literacy at school. statistics education research journal, 16(1), 181–201. https://doi.org/10.52041/serj.v16i1.223 cameron, c., iosua, e., parry, m., richards, r., & jaye, c. (2017). more than just numbers: challenges for professional statisticians. statistics education research journal, 16(2), 362– 375. https://doi.org/10.52041/serj.v16i2.196 carter, j., brown, m., & simpson, k. (2017). from the classroom to the workplace: how social science students are learning to do data analysis for real. statistics education research journal, 16(1), 80–101. carver, r. h. (2011). introductory statistics unconstrained by computability: a new cobb salad. technology innovations in statistics education, 5(1). https://doi.org/10.5070/t551000043 casas-rosal, j. c., caridad y ocerín, j. m., núñez-tabales, j. m., & león-mantero, c. (2019). teaching statistics through the real estate data analyzer software. teaching statistics, 41(2), 58–64. https://doi.org/10.1111/test.12183 casement, c. j., & mcsweeney, l. a. (2022). normalityassessment: an interactive classroom tool for testing normality visually. technology innovations in statistics education, 14(1). https://doi.org/10.5070/t514156556 casey, s. a., albert, j., & ross, a. (2018). developing knowledge for teaching graphing of bivariate categorical data. journal of statistics education, 26(3), 197–213. https://doi.org/10.1080/10691898.2018.1540915 casleton, e., beyler, a., genschel, u., & wilson, a. (2014). a pilot study teaching metrology in an introductory statistics course. journal of statistics education, 22(3), 1. https://doi.org/10.1080/10691898.2014.11889710 çetinkaya-rundel, m., dogucu, m., & rummerfield, w. (2022). the 5ws and 1h of term projects in the introductory data science classroom. statistics education research journal, 21(2), 1–19. https://doi.org/10.52041/serj.v21i2.37 çetinkaya-rundel, m., hardin, j., baumer, b. s., mcnamara, a., horton, n. j., & rundel, c. (2022). an educator’s perspective of the tidyverse. technology innovations in statistics education, 14(1). https://doi.org/10.5070/t514154352 çetinkaya-rundel, m., & rundel, c. (2018). infrastructure and tools for teaching computing throughout the statistical curriculum. the american statistician, 72(1), 58–65. https://doi.org/10.1080/00031305.2017.1397549 24/47 bauder, julia & cave, libby (2025). conceptions of data literacy in the statistics education literature, iassist quarterly 49(4), pp. 1-47. doi: https://doi.org/10.29173/iq1156 chamandy, n., muralidharan, o., & wager, s. (2015). teaching statistics at google-scale. the american statistician, 69(4), 283–291. https://doi.org/10.1080/00031305.2015.1089790 chance, b., ben-zvi, d., garfield, j., & medina, e. (2007). the role of technology in improving student learning of statistics. technology innovations in statistics education, 1(1). https://doi.org/10.5070/t511000026 chance, b., & reynolds, s. (2019). predicting the kentucky derby winner! sort of. journal of statistics education, 27(2), 120–127. https://doi.org/10.1080/10691898.2019.1623137 cobb, g. w. (2007). the introductory statistics course: a ptolemaic curriculum? technology innovations in statistics education, 1(1). https://doi.org/10.5070/t511000028 conti, k. c., & de carvalho, d. l. (2014). statistical literacy: developing a youth and adult education statistical project. statistics education research journal, 13(2), 164–176. https://doi.org/10.52041/serj.v13i2.288 curley, b., & peterson, a. (2022). a fresh shot at statistics in the classroom: three perspectives using world cup soccer player data. journal of statistics and data science education, 30(1), 86–98. https://doi.org/10.1080/26939169.2021.2008283 davies, n., & sheldon, n. (2021). teaching statistics and data science in england’s schools. teaching statistics, 43(s1), s52–s70. https://doi.org/10.1111/test.12276 de oliveira souza, l. d., lopes, c. e., & fitzallen, n. (2020). creative insubordination in statistics teaching: possibilities to go beyond statistical literacy. statistics education research journal, 19(1), 73–91. https://doi.org/10.52041/serj.v19i1.120 de souza oliveira, f. j., & de faria reis, d. a. (2021). the nepso and opinion educative survey in latin america: discussions on statistical literacy in the perspective of this approach. statistics education research journal, 20(2), 1–23. https://doi.org/10.52041/serj.v20i2.316 de veaux, r., hoerl, r., snee, r., & velleman, p. (2022). toward holistic data science education. statistics education research journal, 21(2), 1–12. https://doi.org/10.52041/serj.v21i2.40 delport, d. h. (2021). teaching first-year statistics students with covid-19 real-world data: graphs. teaching statistics, 43(1), 36–43. https://doi.org/10.1111/test.12245 delport, d. h. (2023). the development of statistical literacy among students: analyzing messages in media articles with gal’s worry questions. teaching statistics, 45(2), 61–68. https://doi.org/10.1111/test.12308 depaolo, c. a., robinson, d. f., & jacobs, a. (2016). café data 2.0: new data from a new and improved café. journal of statistics education, 24(2), 85–103. https://doi.org/10.1080/10691898.2016.1196064 25/47 bauder, julia & cave, libby (2025). conceptions of data literacy in the statistics education literature, iassist quarterly 49(4), pp. 1-47. doi: https://doi.org/10.29173/iq1156 dogucu, m., & çetinkaya-rundel, m. (2021). web scraping in the statistics and data science curriculum: challenges and opportunities. journal of statistics and data science education, 29(s1), s112–s122. https://doi.org/10.1080/10691898.2020.1787116 donoghue, t., voytek, b., & ellis, s. e. (2021). teaching creative and practical data science at scale. journal of statistics and data science education, 29(s1), s27–s39. https://doi.org/10.1080/10691898.2020.1860725 dunn, p. k. (2013). comparing the lifetimes of two brands of batteries. journal of statistics education, 21(1), 11. https://doi.org/10.1080/10691898.2013.11889666 dunn, p. k., carey, m. d., farrar, m. b., richardson, a. m., & mcdonald, c. (2017). introductory statistics textbooks and the gaise recommendations. the american statistician, 71(4), 326– 335. https://doi.org/10.1080/00031305.2016.1251972 dunn, p. k., donnison, s., cole, r., & bulmer, m. (2017). using a virtual population to authentically teach epidemiology and biostatistics. international journal of mathematical education in science and technology, 48(2), 185–201. https://doi.org/10.1080/0020739x.2016.1228015 dunn, p. k., richardson, a., prodromou, t., & axelsen, t. (2020). statistics poster competitions: an opportunity to connect academics and teachers. statistics education research journal, 19(1). https://doi.org/10.52041/serj.v19i1.121 engel, j. (2017). statistical literacy for active citizenship: a call for data science education. statistics education research journal, 16(1), 44–49. https://doi.org/10.52041/serj.v16i1.213 engledowl, c. (2019). heat maps: a case for inclusion in secondary statistics instruction. teaching statistics, 41(2), 42–46. https://doi.org/10.1111/test.12177 engledowl, c., & weiland, t. (2021). data (mis)representation and covid-19: leveraging misleading data visualizations for developing statistical literacy across grades 6–16. journal of statistics and data science education, 29(2), 160–164. https://doi.org/10.1080/26939169.2021.1915215 erickson, t. (2013). designing games for understanding in a data analysis environment. technology innovations in statistics education, 7(2). https://doi.org/10.5070/t572013897 erickson, t., & chen, e. (2021). introducing data science with data moves and codap. teaching statistics, 43, s124–s132. https://doi.org/10.1111/test.12240 erickson, t., wilkerson, m., finzer, w., & reichsman, f. (2019). data moves. technology innovations in statistics education, 12(1). https://doi.org/10.5070/t5121038001 estrella, s., vergara, a., & gonzález, o. (2021). developing data sense: making inferences from variability in tsunamis at primary school. statistics education research journal, 20(2), 1–14. https://doi.org/10.52041/serj.v20i2.413 26/47 bauder, julia & cave, libby (2025). conceptions of data literacy in the statistics education literature, iassist quarterly 49(4), pp. 1-47. doi: https://doi.org/10.29173/iq1156 evans, c. (2022). regression, transformations, and mixed-effects with marine bryozoans. journal of statistics and data science education, 30(2), 198–206. https://doi.org/10.1080/26939169.2022.2074923 everson, m. g., & garfield, j. (2008). an innovative approach to teaching online statistics courses. technology innovations in statistics education, 2(1). https://doi.org/10.5070/t521000031 fellers, p. s., & kuiper, s. (2020). introducing undergraduates to concepts of survey data analysis. journal of statistics education, 28(1), 18–24. https://doi.org/10.1080/10691898.2020.1720552 fergusson, a., & pfannkuch, m. (2022). introducing high school statistics teachers to predictive modelling and apis using code-driven tools. statistics education research journal, 21(2), 1-. https://doi.org/10.52041/serj.v21i2.49 fergusson, a., & wild, c. j. (2021). on traversing the data landscape: introducing apis to datascience students. teaching statistics, 43(s1), s71–s83. https://doi.org/10.1111/test.12266 ferris, m., & cheng, s. (2018). using twitter to energize the introductory statistics class. technology innovations in statistics education, 11(1). https://doi.org/10.5070/t5111032036 fiksel, j., jager, l. r., hardin, j. s., & taub, m. a. (2019). using github classroom to teach statistics. journal of statistics education, 27(2), 110–119. https://doi.org/10.1080/10691898.2019.1617089 finzer, w. (2013). the data science education dilemma. technology innovations in statistics education, 7(2). https://doi.org/10.5070/t572013891 finzer, w., erickson, t., swenson, k., & litwin, m. (2007). on getting more and better data into the classroom. technology innovations in statistics education, 1(1). https://doi.org/10.5070/t511000025 fleischer, y., biehler, r., & schulte, c. (2022). teaching and learning data-driven machine learning with educationally designed jupyter notebooks. statistics education research journal, 21(2), 1–25. https://doi.org/10.52041/serj.v21i2.61 forbes, s. (2014a). the coming of age of statistics education in new zealand, and its influence internationally. journal of statistics education, 22(2). https://doi.org/10.1080/10691898.2014.11889699 forbes, s. (2014b). using action research to develop a course in statistical inference for workplacebased adults. journal of statistics education, 22(3). https://doi.org/10.1080/10691898.2014.11889711 forbes, s., chapman, j., harraway, j., stirling, d., & wild, c. (2014). use of data visualisation in the teaching of statistics: a new zealand perspective. statistics education research journal, 13(2), 187–201. https://doi.org/10.52041/serj.v13i2.290 27/47 bauder, julia & cave, libby (2025). conceptions of data literacy in the statistics education literature, iassist quarterly 49(4), pp. 1-47. doi: https://doi.org/10.29173/iq1156 forbes, s. d. (2012). data visualisation: a motivational and teaching tool in official statistics. technology innovations in statistics education, 6(1). https://doi.org/10.5070/t561012851 françois, k., monteiro, c., & allo, p. (2020). big-data literacy as a new vocation for statistical literacy. statistics education research journal, 19(1), 194–205. https://doi.org/10.52041/serj.v19i1.130 freeman, p. e. (2021). facilitating authentic practice for early undergraduate statistics students. the american statistician, 75(4), 433–444. https://doi.org/10.1080/00031305.2020.1844293 frischemeier, d. (2020). building statisticians at an early age – statistical projects exploring meaningful data in primary school. statistics education research journal, 19(1), 39–56. https://doi.org/10.52041/serj.v19i1.118 frischemeier, d., biehler, r., podworny, s., & budde, l. (2021). a first introduction to data science education in secondary schools: teaching and learning about data exploration with codap using survey data. teaching statistics, 43(s1), s182–s189. https://doi.org/10.1111/test.12283 froelich, a. g., & nettleton, d. (2013). does my baby really look like me? using tests for resemblance between parent and child to teach topics in categorical data analysis. journal of statistics education, 21(2), 3. https://doi.org/10.1080/10691898.2013.11889674 froelich, a. g., & stephenson, w. r. (2013). does eye color depend on gender? it might depend on who or how you ask. journal of statistics education, 21(2), 9. https://doi.org/10.1080/10691898.2013.11889680 frölich, n., & schellhammer, k. s. (2022). questionnaire design and sampling procedures for business and economics students: a research-oriented, hands-on course. international journal of mathematical education in science and technology, 0(0), 1–19. https://doi.org/10.1080/0020739x.2022.2056722 fry, k., & makar, k. (2021). how could we teach data science in primary school? teaching statistics, 43(s1), s173–s181. https://doi.org/10.1111/test.12259 gal, i., & ograjenšek, i. (2016). rejoinder: more on enhancing statistics education with qualitative ideas. international statistical review, 84(2), 202–209. https://doi.org/10.1111/insr.12157 garcia-mila, m., marti, e., gilabert, s., & castells, m. (2014). fifth through eighth grade students’ difficulties in constructing bar graphs: data organization, data aggregation, and integration of a second variable. mathematical thinking and learning, 16(3), 201–233. https://doi.org/10.1080/10986065.2014.921132 gardner, k. (2013). a data generating review that bops, twists and pulls at misconceptions. teaching statistics, 35(1), 8–13. https://doi.org/10.1111/j.1467-9639.2012.00522.x 28/47 bauder, julia & cave, libby (2025). conceptions of data literacy in the statistics education literature, iassist quarterly 49(4), pp. 1-47. doi: https://doi.org/10.29173/iq1156 gehrke, m., kistler, t., lübke, k., markgraf, n., krol, b., & sauer, s. (2021). statistics education from a data-centric perspective. teaching statistics, 43(s1), s201–s215. https://doi.org/10.1111/test.12264 gerds, t. a. (2016). the kaplan–meier theatre. teaching statistics, 38(2), 45–49. https://doi.org/10.1111/test.12095 gibbs, a. l., & goossens, e. t. (2013). the evidence for efficacy of hpv vaccines: investigations in categorical data analysis. journal of statistics education, 21(3), 7. https://doi.org/10.1080/10691898.2013.11889688 gil, e., & gibbs, a. l. (2017). promoting modeling and covariational reasoning among secondary school students in the context of big data. statistics education research journal, 16(2), 163-. gómez-blancarte, a. l., chávez, r. r., & chávez aguilar, r. d. (2021). a survey of the teaching of statistical literacy, reasoning and thinking: teachers’ classroom practice in mexican high school education. statistics education research journal, 20(2), 1–18. https://doi.org/10.52041/serj.v20i2.397 gomez-torres, e. (2021). developing “recognition of need for data” in secondary school teachers. statistics education research journal, 20(2), 1-. https://doi.org/10.52041/serj.v20i2.310 gonzalez, o. (2021). teachers’ conceptions and professional knowledge of variability from their interpretation of histograms: the case of venezuelan in-service secondary mathematics teachers. statistics education research journal, 20(2), 1-. https://doi.org/10.52041/serj.v20i2.412 gooding, c. l., lyford, a., & giaimo, g. n. (2022). writing goals in u.s. undergraduate data science course outlines: a textual analysis. teaching statistics, 44(3), 110–118. https://doi.org/10.1111/test.12314 gould, r. (2017). data literacy is statistical literacy. statistics education research journal, 16(1), 22– 25. https://doi.org/10.52041/serj.v16i1.209 gould, r. (2021). toward data-scientific thinking. teaching statistics, 43, s11–s22. https://doi.org/10.1111/test.12267 gould, r., bargagliotti, a., & johnson, t. (2017). an analysis of secondary teachers’ reasoning with participatory sensing data. statistics education research journal, 16(2), 305–334. https://doi.org/10.52041/serj.v16i2.194 grant, r. (2017). statistical literacy in the data science workplace. statistics education research journal, 16(1), 17–21. https://doi.org/10.52041/serj.v16i1.207 green, j. l., smith, w. m., kerby, a. t., blankenship, e. e., schmid, k. k., & carlson, m. a. (2018). introductory statistics: preparing in-service middle-level mathematics teachers for 29/47 bauder, julia & cave, libby (2025). conceptions of data literacy in the statistics education literature, iassist quarterly 49(4), pp. 1-47. doi: https://doi.org/10.29173/iq1156 classroom research. statistics education research journal, 17(2), 216–238. https://doi.org/10.52041/serj.v17i2.167 grimshaw, s. d. (2015). a framework for infusing authentic data experiences within statistics courses. the american statistician, 69(4), 307–314. https://doi.org/10.1080/00031305.2015.1081106 groth, r. e. (2019). applying design-based research findings to improve the common core state standards for data and statistics in grades 4–6. journal of statistics education, 27(1), 29–36. https://doi.org/10.1080/10691898.2019.1565935 groth, r. e., & bergner, j. a. (2013). mapping the structure of knowledge for teaching nominal categorical data analysis. educational studies in mathematics, 83(2), 247–265. https://doi.org/10.1007/s10649-012-9452-4 hahs-vaughn, d. l., acquaye, h., griffith, m. d., jo, h., matthews, k., & acharya, p. (2017). statistical literacy as a function of online versus hybrid course delivery format for an introductory graduate statistics course. journal of statistics education, 25(3), 112–121. https://doi.org/10.1080/10691898.2017.1370363 haldar, l. c., wong, n., heller, j. i., & konold, c. (2018). students making sense of multi-level data. technology innovations in statistics education, 11(1). https://doi.org/10.5070/t5111031358 hardin, j. (2018). dynamic data in the statistics classroom. technology innovations in statistics education, 11(1). https://doi.org/10.5070/t5111031079 hardin, j., hoerl, r., horton, n. j., nolan, d., baumer, b., hall-holt, o., murrell, p., peng, r., roback, p., temple lang, d., & ward, m. d. (2015). data science in statistics curricula: preparing students to “think with data.” the american statistician, 69(4), 343–353. https://doi.org/10.1080/00031305.2015.1077729 hardin, j. s., sarkis, g., & urc, p. c. (2015). network analysis with the enron email corpus. journal of statistics education, 23(2), 2. https://doi.org/10.1080/10691898.2015.11889734 hassad, r. a. (2013). faculty attitude towards technology-assisted instruction for introductory statistics in the context of educational reform. technology innovations in statistics education, 7(2). https://doi.org/10.5070/t572013892 hassad, r. a. (2020). a foundation for inductive reasoning in harnessing the potential of big data. statistics education research journal, 19(1), 238–258. https://doi.org/10.52041/serj.v19i1.133 helenius, r., d’amelio, a., campos, p., & macfeely, s. (2020). islp country coordinators as ambassadors of statistical literacy and innovations. statistics education research journal, 19(1), 120–136. https://doi.org/10.52041/serj.v19i1.125 30/47 bauder, julia & cave, libby (2025). conceptions of data literacy in the statistics education literature, iassist quarterly 49(4), pp. 1-47. doi: https://doi.org/10.29173/iq1156 hicks, s. c., & irizarry, r. a. (2018). a guide to teaching data science. the american statistician, 72(4), 382–391. https://doi.org/10.1080/00031305.2017.1356747 hobden, s. (2014). when statistical literacy really matters: understanding published information about the hiv/aids epidemic in south africa. statistics education research journal, 13(2), 72–82. https://doi.org/10.52041/serj.v13i2.281 horton, n. j. (2015). challenges and opportunities for statistics and statistical education: looking back, looking forward. the american statistician, 69(2), 138–145. https://doi.org/10.1080/00031305.2015.1032435 horton, n. j., alexander, r., parker, m.,piekut, a., & rundel, c. (2022). the growing importance of reproducibility and responsible workflow in the data science and statistics curriculum. journal of statistics and data science education, 30(3), 207–208. https://doi.org/10.1080/26939169.2022.2141001 hourigan, m., & leavy, a. (2016). what do the stats tell us? engaging elementary children in probabilistic reasoning based on data analysis. teaching statistics, 38(1), 8–15. https://doi.org/10.1111/test.12084 hourigan, m., & leavy, a. m. (2020). using integrated stemas a stimulus to develop elementary students’ statistical literacy. teaching statistics, 42(3), 77–86. https://doi.org/10.1111/test.12229 hourigan, m., & leavy, a. m. (2021). interrogating a measurement conjecture to introduce the concept of statistical association in upper elementary education. teaching statistics, 43(2), 62–71. https://doi.org/10.1111/test.12249 hsu, j. l., jones, a., lin, j.-h., & chen, y.-r. (2022). data visualization in introductory business statistics to strengthen students’ practical skills. teaching statistics, 44(1), 21–28. https://doi.org/10.1111/test.12291 hudiburgh, l. m., & garbinsky, d. (2020). data visualization: bringing data to life in an introductory statistics course. journal of statistics education, 28(3), 262–279. https://doi.org/10.1080/10691898.2020.1796399 isoda, m., chitmun, s., & gonzalez, o. (2018). japanese and thai senior high school mathematics teachers’ knowledge of variability. statistics education research journal, 17(2), 196–215. https://doi.org/10.52041/serj.v17i2.166 johnson, r. w. (2017). discovering patterns in interarrival data. teaching statistics, 39(2), 42–46. https://doi.org/10.1111/test.12123 jones, d. l., & scariano, s. m. (2014). measuring the variability of data from other values in the set. teaching statistics, 36(3), 93–96. https://doi.org/10.1111/test.12056 31/47 bauder, julia & cave, libby (2025). conceptions of data literacy in the statistics education literature, iassist quarterly 49(4), pp. 1-47. doi: https://doi.org/10.29173/iq1156 jones, j. d. (2022). using school mathematics to develop students’ data literacy skills. mathematics teacher: learning and teaching pk-12, 115(8), 576–581. https://doi.org/10.5951/mtlt.2021.0239 jones, j. s., & goldring, j. e. (2017). telling stories, landing planes and getting them moving—a holistic approach to developing students’ statistical literacy. statistics education research journal, 16(1), 102–119. https://doi.org/10.52041/serj.v16i1.219 jones, r. c. (2020). data analysis and critical thinking skills training for teachers – the welsh baccalaureate. statistics education research journal, 19(1), 92–105. https://doi.org/10.52041/serj.v19i1.123 kaplan, d. (2007). computing and introductory statistics. technology innovations in statistics education, 1(1). https://doi.org/10.5070/t511000030 kaplan, d. (2018). teaching stats for data science. the american statistician, 72(1), 89–96. https://doi.org/10.1080/00031305.2017.1398107 katrina piatek-jimenez, tibor marcinek, christine m. phelps, & ana dias. (2012). helping students become quantitatively literate. the mathematics teacher, 105(9), 692–696. jstor. https://doi.org/10.5951/mathteacher.105.9.0692 khachatryan, d., & karst, n. (2017). v for voice: strategies for bolstering communication skills in statistics. journal of statistics education, 25(2), 68–78. https://doi.org/10.1080/10691898.2017.1305261 kim, a. y., ismay, c., & chunn, j. (2018). the fivethirtyeight r package: “tame data” principles for introductory statistics and data science courses. technology innovations in statistics education, 11(1). https://doi.org/10.5070/t5111035892 koga, s. (2022). characteristics of statistical literacy skills from the perspective of critical thinking. teaching statistics, 44(2), 59–67. https://doi.org/10.1111/test.12302 konold, c., higgins, t., russell, s. j., & khalil, k. (2015). data seen through different lenses. educational studies in mathematics, 88(3), 305–325. https://doi.org/10.1007/s10649-0139529-8 konold, c., & kazak, s. (2008). reconnecting data and chance. technology innovations in statistics education, 2(1). https://doi.org/10.5070/t521000032 koparan, t. (2019). examination of the dynamic software-supported learning environment in data analysis. international journal of mathematical education in science and technology, 50(2), 277–291. https://doi.org/10.1080/0020739x.2018.1494861 koparan, t., & güven, b. (2015). the effect of project-based learning on students’ statistical literacy levels for data representation. international journal of mathematical education in science and technology, 46(5), 658–686. https://doi.org/10.1080/0020739x.2014.995242 32/47 bauder, julia & cave, libby (2025). conceptions of data literacy in the statistics education literature, iassist quarterly 49(4), pp. 1-47. doi: https://doi.org/10.29173/iq1156 kozak, m., & wnuk, a. (2014). including the tukey mean-difference (bland–altman) plot in a statistics course. teaching statistics, 36(3), 83–87. https://doi.org/10.1111/test.12032 kross, s., peng, r. d., caffo, b. s., gooding, i., & leek, j. t. (2020). the democratization of data science education. the american statistician, 74(1), 1–7. https://doi.org/10.1080/00031305.2019.1668849 kulp, c. w., & sprechini, g. d. (2016). teaching the assessment of normality using large easilygenerated real data sets. teaching statistics, 38(2), 56–62. https://doi.org/10.1111/test.12097 lasser, j., manik, d., silbersdorff, a., säfken, b., & kneib, t. (2021). introductory data science across disciplines, using python, case studies, and industry consulting projects. teaching statistics, 43, s190–s200. https://doi.org/10.1111/test.12243 l’boy, d., & nazim khan, r. (2023). a rasch-model-based hierarchical framework for statistical literacy and learning. international journal of mathematical education in science and technology, 54(9), 1874–1887. https://doi.org/10.1080/0020739x.2023.2261453 le, d. (2013). bringing data to life into an introductory statistics course with gapminder. teaching statistics, 35(3), 114–122. https://doi.org/10.1111/test.12015 leavy, a., & hourigan, m. (2016). crime scenes and mystery players! using driving questions to support the development of statistical literacy. teaching statistics, 38(1), 29–35. https://doi.org/10.1111/test.12088 leavy, a., & hourigan, m. (2018). the role of perceptual similarity, context, and situation when selecting attributes: considerations made by 5–6-year-olds in data modeling environments. educational studies in mathematics, 97(2), 163–183. https://doi.org/10.1007/s10649-0179791-2 lee, h. s., mojica, g. f., thrasher, e. p., & baumgartner, p. (2022). investigating data like a data scientist: key practices and processes. statistics education research journal, 21(2), 1–23. https://doi.org/10.52041/serj.v21i2.41 legacy, c., zieffler, a., fry, e. b., & le, l. (2022). computes: development of an instrument to measure introductory statistics instructors’ emphasis on computational practices. statistics education research journal, 21(1), 7–7. https://doi.org/10.52041/serj.v21i1.63 lem, s., onghena, p., verschaffel, l., & van dooren, w. (2013). external representations for data distributions: in search of cognitive fit. statistics education research journal, 12(1), 4–19. https://doi.org/10.52041/serj.v12i1.319 lesser, l. m., pearl, d. k., weber, j. j., dousa, d. m., carey, r. p., & haddad, s. a. (2019). developing interactive educational songs for introductory statistics. journal of statistics education, 27(3). https://www.proquest.com/docview/2351040760/abstract/92048e7684e14a41pq/1 33/47 bauder, julia & cave, libby (2025). conceptions of data literacy in the statistics education literature, iassist quarterly 49(4), pp. 1-47. doi: https://doi.org/10.29173/iq1156 li, m., mickel, a., & taylor, s. (2018). “should this loan be approved or denied?”: a large dataset with class assignment guidelines. journal of statistics education, 26(1), 55–66. https://doi.org/10.1080/10691898.2018.1434342 lindsay reiten & susanne strachota. (2016). promoting statistical literacy through tuva. the mathematics teacher, 110(3), 228–231. jstor. https://doi.org/10.5951/mathteacher.110.3.0228 loux, t., & gibson, a. k. (2019). using flint, michigan, lead data in introductory statistics. teaching statistics, 41(3), 85–88. https://doi.org/10.1111/test.12187 loy, a., kuiper, s., & chihara, l. (2019). supporting data science in the statistics curriculum. journal of statistics education, 27(1), 2–11. https://doi.org/10.1080/10691898.2018.1564638 lumbantobing, r., & mcfall, t. (2022). an applied statistics teaching lesson that uses nba playoff data to illustrate uncertainty in sporting contests. teaching statistics, 44(3), 104–109. https://doi.org/10.1111/test.12313 macfeely, s., campos, p., & helenius, r. (2017). key success factors for statistical literacy poster competitions. statistics education research journal, 16(1), 202–216. https://doi.org/10.52041/serj.v16i1.224 mackay, j. (2016). discussion: acquiring statistical literacy and thinking. international statistical review, 84(2), 189–194. https://doi.org/10.1111/insr.12154 mackay, j. (2022). data discovery challenge using the covid-19 data portal from new zealand. journal of statistics and data science education, 30(2), 187–190. https://doi.org/10.1080/26939169.2022.2058656 makar, k. (2014). young children’s explorations of average through informal inferential reasoning. educational studies in mathematics, 86(1), 61–78. https://doi.org/10.1007/s10649-0139526-y marla a. sole. (2016). engaging students in survey design and data collection. the mathematics teacher, 109(5), 334–340. jstor. https://doi.org/10.5951/mathteacher.109.5.0334 marron, m. m., & wahed, a. s. (2016). teaching missing data methodology to undergraduates using a group-based project within a six-week summer program. journal of statistics education, 24(1), 8–15. https://doi.org/10.1080/10691898.2016.1158018 maurer, k., & lock, d. (2016). comparison of learning outcomes for simulation-based and traditional inference curricula in a designed educational experiment. technology innovations in statistics education, 9(1). https://doi.org/10.5070/t591026161 mcdaniel, s. n., & green, l. (2012). independent interactive inquiry-based learning modules using audio-visual instruction in statistics. technology innovations in statistics education, 6(1). https://doi.org/10.5070/t561012656 34/47 bauder, julia & cave, libby (2025). conceptions of data literacy in the statistics education literature, iassist quarterly 49(4), pp. 1-47. doi: https://doi.org/10.29173/iq1156 mcgee, m. (2019). deep dive into visual representation and interrater agreement using data from a high-school diving competition. journal of statistics education, 27(3), 275–287. https://doi.org/10.1080/10691898.2019.1632759 mike, k., & hazzan, o. (2022). machine learning for non-major data science students: a white box approach. statistics education research journal, 21(2), 1-. https://doi.org/10.52041/serj.v21i2.45 molnar, a. (2013). discussion: what do instructors of statistics need to know about technology, and how can they best be taught? technology innovations in statistics education, 7(2). https://doi.org/10.5070/t572013896 myint, l., hadavand, a., jager, l., & leek, j. (2020). comparison of beginning r students’ perceptions of peer-made plots created in two plotting systems: a randomized experiment. journal of statistics education, 28(1), 98–108. https://doi.org/10.1080/10691898.2019.1695554 nazzaro, v., rose, j., & dierker, l. (2020). a comparison of future course enrollment among students completing one of four different introductory statistics courses. statistics education research journal, 19(3), 6–17. neumann, d. l., hood, m., & neumann, m. m. (2013). using real-life data when teaching statistics: student perceptions of this strategy in an introductory statistics course. statistics education research journal, 12(2), 59–70. https://doi.org/10.52041/serj.v12i2.304 nicholson, j., ridgway, j., & mccusker, s. (2013). getting real statistics into all curriculum subject areas: can technology make this a reality? technology innovations in statistics education, 7(2). https://doi.org/10.5070/t572013906 nilsson, p. (2013). challenges in seeing data as useful evidence in making predictions on the probability of a real-world phenomenon. statistics education research journal, 12(2), 71– 83. https://doi.org/10.52041/serj.v12i2.305 nolan, d., & perrett, j. (2016). teaching and learning data visualization: ideas and assignments. the american statistician, 70(3), 260–269. https://doi.org/10.1080/00031305.2015.1123651 nolan, d., & temple lang, d. (2015). explorations in statistics research: an approach to expose undergraduates to authentic data analysis. the american statistician, 69(4), 292–299. https://doi.org/10.1080/00031305.2015.1073624 noll, j., & kirin, d. (2016). student approaches to constructing statistical models using tinkerplots tm. technology innovations in statistics education, 9(1). https://doi.org/10.5070/t591023693 noll, j., & tackett, m. (2023). insights from datafest point to new opportunities for undergraduate statistics courses: team collaborations, designing research questions, and data ethics. teaching statistics, 45(s1), s5–s21. https://doi.org/10.1111/test.12345 35/47 bauder, julia & cave, libby (2025). conceptions of data literacy in the statistics education literature, iassist quarterly 49(4), pp. 1-47. doi: https://doi.org/10.29173/iq1156 nowacki, a. s. (2013). data sharing and the development of the cleveland clinic statistical education dataset repository. journal of statistics education, 21(1), 4. https://doi.org/10.1080/10691898.2013.11889660 nowacki, a. s. (2015). teaching statistics from the operating table: minimally invasive and maximally educational. journal of statistics education, 23(1), 6. https://doi.org/10.1080/10691898.2015.11889726 odom, a. l., & bell, c. v. (2017). developing pk-12 preservice teachers’ skills for understanding data-driven instruction through inquiry learning. journal of statistics education, 25(1), 29– 37. https://doi.org/10.1080/10691898.2017.1288557 oslington, g., mulligan, j., & van bergen, p. (2020). third-graders’ predictive reasoning strategies. educational studies in mathematics, 104(1), 5–24. https://doi.org/10.1007/s10649-02009949-0 ostblom, j., & timbers, t. (2022). opinionated practices for teaching reproducibility: motivation, guided instruction and practice. journal of statistics and data science education, 30(3), 241– 250. https://doi.org/10.1080/26939169.2022.2074922 peng, r. d., chen, a., bridgeford, e., leek, j. t., & hicks, s. c. (2021). diagnosing data analytic problems in the classroom. journal of statistics and data science education, 29(3), 267–276. https://doi.org/10.1080/26939169.2021.1971586 peterson, a. d., & ziegler, l. (2021). building a multiple linear regression model with lego brick data. journal of statistics and data science education, 29(3), 297–303. https://doi.org/10.1080/26939169.2021.1946450 phelps, a. l., & szabat, k. a. (2017). the current landscape of teaching analytics to business students at institutions of higher education: who is teaching what? the american statistician, 71(2), 155–161. https://doi.org/10.1080/00031305.2016.1277160 piatek-jimenez, k., marcinek, t., phelps, c. m., & dias, a. (2012). helping students become quantitatively literate. the mathematics teacher, 105(9), 692–696. https://doi.org/10.5951/mathteacher.105.9.0692 podworny, s., husing, s., & schulte, c. (2022). a place for a data science introduction in school: between statistics and programming. statistics education research journal, 21(2), 1-. https://doi.org/10.52041/serj.v21i2.46 poling, l., & weiland, t. (2020). using an interactive platform to recognize the intersection of social and spatial inequalities. teaching statistics, 42(3), 108–116. https://doi.org/10.1111/test.12234 prodromou, t., & dunne, t. (2017). statistical literacy in data revolution era: building blocks and instructional dilemmas. statistics education research journal, 16(1), 38–43. https://doi.org/10.52041/serj.v16i1.212 36/47 bauder, julia & cave, libby (2025). conceptions of data literacy in the statistics education literature, iassist quarterly 49(4), pp. 1-47. doi: https://doi.org/10.29173/iq1156 queiroz, t., monteiro, c., carvalho, l., & françois, k. (2017). interpretation of statistical data: the importance of affective expressions. statistics education research journal, 16(1), 163–180. https://doi.org/10.52041/serj.v16i1.222 rao, v. n. v., legacy, c., zieffler, a., & delmas, r. (2023). designing a sequence of activities to build reasoning about data and visualization. teaching statistics, 45(s1), s80–s92. https://doi.org/10.1111/test.12341 reinhart, a., & genovese, c. r. (2021). expanding the scope of statistical computing: training statisticians to be software engineers. journal of statistics and data science education, 29(s1), s7–s15. https://doi.org/10.1080/10691898.2020.1845109 reiten, l., & strachota, s. (2016). promoting statistical literacy through tuva. the mathematics teacher, 110(3), 228–231. https://doi.org/10.5951/mathteacher.110.3.0228 reston, e. (2013). an outcome-based framework for technology integration in higher education statistics curricula for non-majors. technology innovations in statistics education, 7(2). https://doi.org/10.5070/t572013894 richardson, a. m., & dunn, p. k. (2021). simple interventions to assist students to engage with the language of data science and statistics. teaching statistics, 43(s1), s148–s156. https://doi.org/10.1111/test.12247 ridgway, j. (2016). implications of the data revolution for statistics education. international statistical review, 84(3), 528–549. https://doi.org/10.1111/insr.12110 ridgway, j. (2021). covid and data science: understanding r0 could change your life. teaching statistics, 43(s1), s84–s92. https://doi.org/10.1111/test.12273 ridgway, j., nicholson, j., & mccusker, s. (2013). “open data” and the semantic web require a rethink on statistics teaching. technology innovations in statistics education, 7(2). https://doi.org/10.5070/t572013907 rivera, r., marazzi, m., & torres-saavedra, p. a. (2019). incorporating open data into introductory courses in statistics. journal of statistics education, 27(3), 198–207. https://doi.org/10.1080/10691898.2019.1669506 roepke, t. l., & gallagher, d. k. (2015). using literacy strategies to teach precalculus and calculus. the mathematics teacher, 108(9), 672–678. https://doi.org/10.5951/mathteacher.108.9.0672 rossman, a. j., laurent, r. st., & tabor, j. (2015). advanced placement statistics: expanding the scope of statistics education. the american statistician, 69(2), 121–126. https://doi.org/10.1080/00031305.2015.1033985 37/47 bauder, julia & cave, libby (2025). conceptions of data literacy in the statistics education literature, iassist quarterly 49(4), pp. 1-47. doi: https://doi.org/10.29173/iq1156 roth, w.-m., & temple, s. (2014). on understanding variability in data: a study of graph interpretation in an advanced experimental biology laboratory. educational studies in mathematics, 86(3), 359–376. https://doi.org/10.1007/s10649-014-9535-5 rubel, l. h., nicol, c., & chronaki, a. (2021). a critical mathematics perspective on reading data visualizations: reimagining through reformatting, reframing, and renarrating. educational studies in mathematics, 108(1–2), 249–268. https://doi.org/10.1007/s10649-021-10087-4 rubin, a. (2021). what to consider when we consider data. teaching statistics, 43, s23–s33. https://doi.org/10.1111/test.12275 sabbag, a., garfield, j., & zieffler, a. (2018). assessing statistical literacy and statistical reasoning: the reali instrument. statistics education research journal, 17(2), 141–160. https://doi.org/10.52041/serj.v17i2.163 salcedo, a. (2014). statistics test questions: content and trends. statistics education research journal, 13(2), 202–217. https://doi.org/10.52041/serj.v13i2.291 schield, m. (2017). gaise 2016 promotes statistical literacy. statistics education research journal, 16(1), 50–54. https://doi.org/10.52041/serj.v16i1.214 schwab-mccoy, a., baker, c. m., & gasper, r. e. (2021). data science in 2020: computing, curricula, and challenges for the next 10 years. journal of statistics and data science education, 29(s1), s40–s50. https://doi.org/10.1080/10691898.2020.1851159 shaltayev, d. s., hodges, h., & hasbrouck, r. b. (2010). visa: reducing technological impact on student learning in an introductory statistics course. technology innovations in statistics education, 4(1). https://doi.org/10.5070/t541000041 sole, m. a. (2016a). engaging students in survey design and data collection. the mathematics teacher, 109(5), 334–340. https://doi.org/10.5951/mathteacher.109.5.0334 sole, m. a. (2016b). statistical literacy: data tell a story. the mathematics teacher, 110(1), 26–32. https://doi.org/10.5951/mathteacher.110.1.0026 sole, m. a., & weinberg, s. l. (2017). what’s brewing? a statistics education discovery project. journal of statistics education, 25(3), 137–144. https://doi.org/10.1080/10691898.2017.1395302 soledad fernández, m., pomilio, c., cueto, g., filloy, j., gonzalez-arzac, a., lois-milevicich, j., & pérez, a. (2020). improving skills to teach statistics in secondary school through activitybased workshops. statistics education research journal, 19(1), 106–122. https://doi.org/10.52041/serj.v19i1.124 sønvisen, s. a. (2023). motivation for learning statistics: an example from fishery and aquaculture science. teaching statistics, 45(2), 85–99. https://doi.org/10.1111/test.12334 38/47 bauder, julia & cave, libby (2025). conceptions of data literacy in the statistics education literature, iassist quarterly 49(4), pp. 1-47. doi: https://doi.org/10.29173/iq1156 stander, j., & dalla valle, l. (2017). on enthusing students about big data and social media visualization and analysis using r, rstudio, and rmarkdown. journal of statistics education, 25(2), 60–67. https://doi.org/10.1080/10691898.2017.1322474 stern, d. (2013). developing statistics education in kenya through technological innovations at all academic levels. technology innovations in statistics education, 7(2). https://doi.org/10.5070/t572013905 stern, d., stern, r., parsons, d., musyoka, j., torgbor, f., & mbasu, z. (2020). envisioning change in the statistics-education climate. statistics education research journal, 19(1), 206–225. https://doi.org/10.52041/serj.v19i1.131 stoudt, s. (2022). collaborative writing workflows in the data-driven classroom: a conversation starter. journal of statistics and data science education, 30(3), 282–288. https://doi.org/10.1080/26939169.2022.2082602 stoudt, s., scotina, a. d., & luebke, k. (2022). supporting statistics and data science education with learnr. technology innovations in statistics education, 14(1). https://doi.org/10.5070/t514156264 strachota, s., & reiten, l. (2017). fostering productive statistical skepticism. the mathematics teacher, 111(3), 222–224. https://doi.org/10.5951/mathteacher.111.3.0222 stratton, c., green, j. l., & hoegh, a. (2021). not just normal: exploring power with shiny apps. technology innovations in statistics education, 13(1). https://doi.org/10.5070/t513146468 strayer, j. f., & edwards, m. t. (2015). smarter cookies. the mathematics teacher, 108(8), 608–615. https://doi.org/10.5951/mathteacher.108.8.0608 stump, s. l., bryan, j. a., & mcconnell, t. j. (2016). making stem connections. the mathematics teacher, 109(8), 576–583. https://doi.org/10.5951/mathteacher.109.8.0576 sutherland, s., & ridgway, j. (2017). interactive visualisations and statistical literacy. statistics education research journal, 16(1), 26–30. https://doi.org/10.52041/serj.v16i1.210 tan, k. s., elkin, e. b., & satagopan, j. m. (2022). a model for an undergraduate research experience program in quantitative sciences. journal of statistics and data science education, 30(1), 65–74. https://doi.org/10.1080/26939169.2021.2016036 tena l. roepke & debra k. gallagher. (2015). using literacy strategies to teach precalculus and calculus. the mathematics teacher, 108(9), 672–678. jstor. https://doi.org/10.5951/mathteacher.108.9.0672 theobold, a., & hancock, s. (2019). how environmental science graduate students acquire statistical computing skills. statistics education research journal, 18(2), 68–85. https://doi.org/10.52041/serj.v18i2.141 39/47 bauder, julia & cave, libby (2025). conceptions of data literacy in the statistics education literature, iassist quarterly 49(4), pp. 1-47. doi: https://doi.org/10.29173/iq1156 theobold, a. s., hancock, s. a., & mannheimer, s. (2021). designing data science workshops for data-intensive environmental science research. journal of statistics and data science education, 29(s1), s83–s94. https://doi.org/10.1080/10691898.2020.1854636 thompson, j., & irgens, g. a. (2022). data detectives: a data science program for middle grade learners. journal of statistics and data science education, 30(1), 29–38. https://doi.org/10.1080/26939169.2022.2034489 towse, j., davies, r., ball, e., james, r., gooding, b., & ivory, m. (2022). lustre: an online data management and student project resource. journal of statistics and data science education, 30(3), 266–273. https://doi.org/10.1080/26939169.2022.2118645 trafimow, d. (2016). the attenuation of correlation coefficients: a statistical literacy issue. teaching statistics, 38(1), 25–28. https://doi.org/10.1111/test.12087 tunstall, s. l. (2018). investigating college students’ reasoning with messages of risk and causation. journal of statistics education, 26(2), 76–86. https://doi.org/10.1080/10691898.2018.1456989 ubilla, f. m., & gorgorió, n. (2021). from a source of real data to a brief news report: introducing first-year preservice teachers to the basic cycle of learning from data. teaching statistics, 43(s1), s110–s123. https://doi.org/10.1111/test.12246 utts, j. (2021). enhancing data science ethics through statistical education and practice. international statistical review, 89(1), 1–17. https://doi.org/10.1111/insr.12446 vance, e. a. (2021). using team-based learning to teach data science. journal of statistics and data science education, 29(3), 277–296. https://doi.org/10.1080/26939169.2021.1971587 vance, e. a., glimp, d. r., pieplow, n. d., garrity, j. m., & melbourne, b. a. (2022). integrating the humanities into data science education: reimagining the introductory data science course. statistics education research journal, 21(2), 1–18. https://doi.org/10.52041/serj.v21i2.42 vance, e. a., & smith, h. s. (2019). the asccr frame for learning essential collaboration skills. journal of statistics education, 27(3), 265–274. https://doi.org/10.1080/10691898.2019.1687370 wagaman, j. c. (2017). introductory statistics in the garden. teaching statistics, 39(2), 52–56. https://doi.org/10.1111/test.12125 wang, x., reich, n. g., & horton, n. j. (2019). enriching students’ conceptual understanding of confidence intervals: an interactive trivia-based classroom activity. the american statistician, 73(1), 50–55. https://doi.org/10.1080/00031305.2017.1305294 wang, x., rush, c., & horton, n. j. (2017). data visualization on day one: bringing big ideas into intro stats early and often. technology innovations in statistics education, 10(1). https://doi.org/10.48550/arxiv.1705.08544 40/47 bauder, julia & cave, libby (2025). conceptions of data literacy in the statistics education literature, iassist quarterly 49(4), pp. 1-47. doi: https://doi.org/10.29173/iq1156 watson, j., & donne, j. (2009). tinkerplots as a research tool to explore student understanding. technology innovations in statistics education, 3(1). https://doi.org/10.5070/t531000034 watson, j., & english, l. (2017). reaction time in grade 5: data collection within the practice of statistics. statistics education research journal, 16(1), 262–293. https://doi.org/10.52041/serj.v16i1.231 watson, j., fitzallen, n., english, l., & wright, s. (2020). introducing statistical variation in year 3 in a stem context: manufacturing licorice. international journal of mathematical education in science and technology, 51(3), 354–387. https://doi.org/10.1080/0020739x.2018.1562117 watson, j., fitzallen, n., wright, s., & kelly, b. (2022). characterizing student experience of variation within a stem context: improving catapults. statistics education research journal, 21(1), 9– 9. https://doi.org/10.52041/serj.v21i1.7 weiland, t. (2017). problematizing statistical literacy: an intersection of critical and statistical literacies. educational studies in mathematics, 96(1), 33–47. https://doi.org/10.1007/s10649-017-9764-5 weiland, t, & sundrani, a. (2022). opportunities for k-8 students to learn statistics created by states’ standards in the united states. journal of statistics and data science education, 30(2), 165–178. https://doi.org/10.1080/26939169.2022.2075814 wild, c. j. (2017). statistical literacy as the earth moves. statistics education research journal, 16(1), 31–37. https://doi.org/10.52041/serj.v16i1.211 wilkerson, m. h., lanouette, k., & shareff, r. l. (2022). exploring variability during data preparation: a way to connect data, chance, and context when working with complex public datasets. mathematical thinking and learning, 24(4), 312–330. https://doi.org/10.1080/10986065.2021.1922838 wilson, m., ross, a., & casey, s. (2021). a classroom-ready activity on educational disparities in the united states. teaching statistics, 43(s1), s93–s97. https://doi.org/10.1111/test.12252 witt, g. (2013). using data from climate science to teach introductory statistics. journal of statistics education, 21(1), 12. https://doi.org/10.1080/10691898.2013.11889667 yan, d., & davis, g. e. (2019). a first course in data science. journal of statistics education, 27(2), 99–109. https://doi.org/10.1080/10691898.2019.1623136 yolcu, a. (2014). middle school students’ statistical literacy: role of grade level and gender. statistics education research journal, 13(2), 118–131. https://doi.org/10.52041/serj.v13i2.285 zakari, i. s. (2020). promoting statistics in the era of data science and data-driven innovations. statistics education research journal, 19(1), 226–237. https://doi.org/10.52041/serj.v19i1.132 41/47 bauder, julia & cave, libby (2025). conceptions of data literacy in the statistics education literature, iassist quarterly 49(4), pp. 1-47. doi: https://doi.org/10.29173/iq1156 zapata-cardona, l. (2023). the possibilities of exploring nontraditional datasets with young children. teaching statistics, 45(s1), s22–s29. https://doi.org/10.1111/test.12349 zheng, q., & lu, y. (2016). do you catch undersized fish? let’s go fishing to learn some important concepts in multiple testing. teaching statistics, 38(3), 91–97. https://doi.org/10.1111/test.12107 zhu, y., hernandez, l. m., mueller, p., dong, y., & forman, m. r. (2013). data acquisition and preprocessing in studies on humans: what is not taught in statistics classes? the american statistician, 67(4), 235–241. https://doi.org/10.1080/00031305.2013.842498 ziegler, l., & garfield, j. (2013). exploring students’ intuitive ideas of randomness using an ipod shuffle activity. teaching statistics, 35(1), 2–7. https://doi.org/10.1111/j.14679639.2012.00531.x ziegler, l., & garfield, j. (2018). developing a statistical literacy assessment for the modern introductory statistics course. statistics education research journal, 17(2), 161–178. https://doi.org/10.52041/serj.v17i2.164 appendix c: removed articles aberson, c. (2021). building interactive tutorials for teaching psychological statistics online with learnr. technology innovations in statistics education, 13(1). https://doi.org/10.5070/t513153822 alexis stevens & john stevens. (2016). using mathematics to elect the u.s. president. the mathematics teacher, 110(3), 192–198. jstor. https://doi.org/10.5951/mathteacher.110.3.0192 austin, p. c. (2017). a tutorial on multilevel survival analysis: methods, models and applications. international statistical review, 85(2), 185–203. https://doi.org/10.1111/insr.12214 ayalon, m., watson, a., & lerman, s. (2015). functions represented as linear sequential data: relationships between presentation and student responses. educational studies in mathematics, 90(3), 321–339. https://doi.org/10.1007/s10649-015-9628-9 baffour, b., chandra, h., & martinez, a. (2019). localised estimates of dynamics of multidimensional disadvantage: an application of the small area estimation technique using australian survey and census data. international statistical review, 87(1), 1– 23. https://doi.org/10.1111/insr.12270 beemer, j., spoon, k., fan, j., stronach, j., frazee, j. p., bohonak, a. j., & levine, r. a. (2018). assessing instructional modalities: individualized treatment effects for personalized learning. journal of statistics education, 26(1), 31–39. https://doi.org/10.1080/10691898.2018.1426400 42/47 bauder, julia & cave, libby (2025). conceptions of data literacy in the statistics education literature, iassist quarterly 49(4), pp. 1-47. doi: https://doi.org/10.29173/iq1156 bulmer, m., & haladyn, j. k. (2011). life on an island: a simulated population to support student projects in statistics. technology innovations in statistics education, 5(1). https://doi.org/10.5070/t551000187 carey, m. d., & dunn, p. k. (2018). facilitating language-focused cooperative learning in introductory statistics classrooms: a case study. statistics education research journal, 17(2), 30–50. https://doi.org/10.52041/serj.v17i2.157 çetinkaya-rundel, m., & rundel, c. (2018). infrastructure and tools for teaching computing throughout the statistical curriculum. the american statistician, 72(1), 58–65. https://doi.org/10.1080/00031305.2017.1397549 chad leith, elena rose, & tony king. (2016). teaching mathematics and language to english learners. the mathematics teacher, 109(9), 670–678. jstor. https://doi.org/10.5951/mathteacher.109.9.0670 cronin, a., intepe, g., shearman, d., & sneyd, a. (2019). analysis using natural language processing of feedback data from two mathematics support centres. international journal of mathematical education in science and technology, 50(7), 1087–1103. https://doi.org/10.1080/0020739x.2019.1656831 de oliveira, v. (2020). models for geostatistical binary data: properties and connections. the american statistician, 74(1), 72–79. https://doi.org/10.1080/00031305.2018.1444674 doi, j., potter, g., wong, j., alcaraz, i., & chi, p. (2016). web application teaching tools for statistics using r and shiny. technology innovations in statistics education, 9(1). https://doi.org/10.5070/t591027492 dunn, p. k. (2022). the impact of using artificial data in undergraduate statistics students’ projects due to covid-19 lockdowns. international journal of mathematical education in science and technology, ahead-of-print(ahead-of-print), 1–10. https://doi.org/10.1080/0020739x.2022.2056095 dunn, p. k., marshman, m., mcdougall, r., & wiegand, a. (2015). teachers and textbooks: on statistical definitions in senior secondary mathematics. journal of statistics education, 23(3). https://doi.org/10.1080/10691898.2015.11889744 emily p. thrasher & ayanna d. perry. (2015). high-leverage apps for the mathematics classroom: wolframalpha. the mathematics teacher, 109(1), 66–70. jstor. https://doi.org/10.5951/mathteacher.109.1.0066 english, l. d., & watson, j. m. (2016). development of probabilistic understanding in fourth grade. journal for research in mathematics education, 47(1), 28–62. https://doi.org/10.5951/jresematheduc.47.1.0028 43/47 bauder, julia & cave, libby (2025). conceptions of data literacy in the statistics education literature, iassist quarterly 49(4), pp. 1-47. doi: https://doi.org/10.29173/iq1156 fergusson, a., & pfannkuch, m. (2022). introducing teachers who use gui-driven tools for the randomization test to code-driven tools. mathematical thinking and learning, 24(4), 336–356. https://doi.org/10.1080/10986065.2021.1922856 gabrosek, j., & o’kelly, l. (2018). r-e-s-p-e-c-t: the role of race, gender, and radio consultants on radio airplay in 1960s chicago, il and grand rapids, mi. journal of statistics education, 26(3), 223–233. https://doi.org/10.1080/10691898.2018.1506953 gal, i., & geiger, v. (2022). welcome to the era of vague news: a study of the demands of statistical and mathematical products in the covid-19 pandemic media. educational studies in mathematics, 111(1), 5–28. https://doi.org/10.1007/s10649-022-10151-7 gundlach, e., richards, k. a. r., nelson, d., & levesque-bristol, c. (2015). a comparison of student attitudes, statistical reasoning, performance, and perceptions for web-augmented traditional, fully online, and flipped sections of a statistical literacy class. journal of statistics education, 23(1), 3. https://doi.org/10.1080/10691898.2015.11889723 harraway, j. a. (2012). learning statistics using motivational videos, real data and free software. technology innovations in statistics education, 6(1). https://doi.org/10.5070/t561000186 heinzman, e. (2022). “i love math only if it’s coding”: a case study of student experiences in an introduction to data science course. statistics education research journal, 21(2), 1-. https://doi.org/10.52041/serj.v21i2.43 hermans, l., molenberghs, g., aerts, m., kenward, m. g., & verbeke, g. (2018). a tutorial on the practical use and implication of complete sufficient statistics. international statistical review, 86(3), 403–414. https://doi.org/10.1111/insr.12261 horton, n. j., & hardin, j. s. (2021). integrating computing in the statistics and data science curriculum: creative structures, novel skills and habits, and ways to teach computational thinking. journal of statistics and data science education, 29(s1), s1–s3. https://doi.org/10.1080/10691898.2020.1870416 humenberger, h. (2020). how does the change of a single data point affect the variance, and why? teaching statistics, 42(3), 87–90. https://doi.org/10.1111/test.12230 jankvist, u. t., & niss, m. (2020). upper secondary school students’ difficulties with mathematical modelling. international journal of mathematical education in science and technology, 51(4), 467–496. https://doi.org/10.1080/0020739x.2019.1587530 jeremy strayer & amber matuszewski. (2016). statistical literacy: simulations with dolphins. the mathematics teacher, 109(8), 606–611. jstor. https://doi.org/10.5951/mathteacher.109.8.0606 johnson, r. w., kliche, d. v., & smith, p. l. (2015). modeling raindrop size. journal of statistics education, 23(1), 5. https://doi.org/10.1080/10691898.2015.11889725 44/47 bauder, julia & cave, libby (2025). conceptions of data literacy in the statistics education literature, iassist quarterly 49(4), pp. 1-47. doi: https://doi.org/10.29173/iq1156 kathleen h. offenholley. (2013). bundled-up babies and dangerous ice cream: correlation puzzlers. the mathematics teacher, 106(6), 418–422. jstor. https://doi.org/10.5951/mathteacher.106.6.0418 kuiper, s., & sturdivant, r. x. (2015). using online game-based simulations to strengthen students’ understanding of practical statistical issues in real-world data analysis. the american statistician, 69(4), 354–361. https://doi.org/10.1080/00031305.2015.1075421 laurie h. rubel, michael driskill, & lawrence m. lesser. (2012). decennial redistricting: rich mathematics in context. the mathematics teacher, 106(3), 206–211. jstor. https://doi.org/10.5951/mathteacher.106.3.0206 lesser, l. m., & santos, m. (2023). a survey on how college students in a statistical literacy course apply statistics terms to people. journal of statistics and data science education, 0(0), 1–15. https://doi.org/10.1080/26939169.2023.2193307 lindsay m. keazer & rahul s. menon. (2016). reasoning and sense making begins with the teacher. the mathematics teacher, 109(5), 342–349. jstor. https://doi.org/10.5951/mathteacher.109.5.0342 lu, y., & henning, k. s. s. (2013). are statisticians cold-blooded bosses? a new perspective on the ‘old’ concept of statistical population. teaching statistics, 35(1), 66–71. https://doi.org/10.1111/j.1467-9639.2012.00524.x lübke, k., gehrke, m., horst, j., & szepannek, g. (2020). why we should teach causal inference: examples in linear regression with simulated data. journal of statistics education, 28(2), 133–139. https://doi.org/10.1080/10691898.2020.1752859 macgillivray, h. (2021). statistics and data science must speak together. teaching statistics, 43(s1), s5–s10. https://doi.org/10.1111/test.12281 manage, a. b. w., & scariano, s. m. (2013). an introductory application of principal components to cricket data. journal of statistics education, 21(3), 8. https://doi.org/10.1080/10691898.2013.11889689 marla a. sole. (2017). financial education: increase your purchasing power. the mathematics teacher, 111(1), 60–64. jstor. https://doi.org/10.5951/mathteacher.111.1.0060 marwick, b., boettiger, c., & mullen, l. (2018). packaging data analytical work reproducibly using r (and friends). the american statistician, 72(1), 80–88. https://doi.org/10.1080/00031305.2017.1375986 mccune, d., & tunstall, s. l. (2019). calculated democracy—e xplorations in gerrymandering. teaching statistics, 41(2), 47–53. https://doi.org/10.1111/test.12181 45/47 bauder, julia & cave, libby (2025). conceptions of data literacy in the statistics education literature, iassist quarterly 49(4), pp. 1-47. doi: https://doi.org/10.29173/iq1156 mcdaniel, s. n., & green, l. b. (2012). using applets and video instruction to foster students’ understanding of sampling variability. technology innovations in statistics education, 6(1). https://doi.org/10.5070/t561000177 mcgowan, h. m., & gunderson, b. k. (2010). a randomized experiment exploring how certain features of clicker use effect undergraduate students’ engagement and learning in statistics. technology innovations in statistics education, 4(1). https://doi.org/10.5070/t541000042 mclaughlin, j. e., & kang, i. (2017). a flipped classroom model for a biostatistics short course. statistics education research journal, 16(2), 441–453. https://doi.org/10.52041/serj.v16i2.200 michael j. caulfield. (2012). what if? how apportionment methods choose our presidents. the mathematics teacher, 106(3), 178–183. jstor. https://doi.org/10.5951/mathteacher.106.3.0178 mocko, m. (2013). selecting technology to promote learning in an online introductory statistics course. technology innovations in statistics education, 7(2). https://doi.org/10.5070/t572013893 murawska, j. m., & nabb, k. a. (2015). corvettes, curve fitting, and calculus. the mathematics teacher, 109(2), 128–135. https://doi.org/10.5951/mathteacher.109.2.0128 newfeld, d. (2016). a first assignment to create student buy-in in an introductory business statistics course. teaching statistics, 38(3), 87–90. https://doi.org/10.1111/test.12106 nirmala naresh, suzanne r. harper, jane m. keiser, & norm krumpe. (2014). probability explorations in a multicultural context. the mathematics teacher, 108(3), 184–192. jstor. https://doi.org/10.5951/mathteacher.108.3.0184 north, d., gal, i., & zewotir, t. (2014). building capacity for developing statistical literacy in a developing country: lessons learned from an intervention. statistics education research journal, 13(2), 15–27. https://doi.org/10.52041/serj.v13i2.276 planas, n. (2014). one speaker, two languages: learning opportunities in the mathematics classroom. educational studies in mathematics, 87(1), 51–66. https://doi.org/10.1007/s10649-014-9553-3 porcu, e., alegria, a., & furrer, r. (2018). modeling temporally evolving and spatially globally dependent data. international statistical review, 86(2), 344–377. https://doi.org/10.1111/insr.12266 rethlefsen, m. l., norton, h. f., meyer, s. l., macwilkinson, k. a., smith ii, p. l., & ye, h. (2022). interdisciplinary approaches and strategies from research reproducibility 2020: educating for reproducibility. journal of statistics and data science education, 30(3), 219–227. https://doi.org/10.1080/26939169.2022.2104767 46/47 bauder, julia & cave, libby (2025). conceptions of data literacy in the statistics education literature, iassist quarterly 49(4), pp. 1-47. doi: https://doi.org/10.29173/iq1156 rubel, l. h., & nicol, c. (2020). the power of place: spatializing critical mathematics education. mathematical thinking and learning, 22(3), 173–194. https://doi.org/10.1080/10986065.2020.1709938 rubin, a. (2007). much has changed; little has changed: revisiting the role of technology in statistics education 1992-2007. technology innovations in statistics education, 1(1). https://doi.org/10.5070/t511000027 schindler, m., & lilienthal, a. j. (2019). domain-specific interpretation of eye tracking data: towards a refined use of the eye-mind hypothesis for the field of geometry. educational studies in mathematics, 101(1), 123–139. https://doi.org/10.1007/s10649-0199878-z sun, d. l., & alfredo, j. (2019). a modern look at freedman’s box model. technology innovations in statistics education, 12(1). https://doi.org/10.5070/t5121044395 utts, j. (2015). the many facets of statistics education: 175 years of common themes. the american statistician, 69(2), 100–107. https://doi.org/10.1080/00031305.2015.1033981 victor mateas. (2013). connecting algebra to economics. the mathematics teacher, 107(4), 298– 304. jstor. https://doi.org/10.5951/mathteacher.107.4.0298 wagaman, a. (2016). meeting student needs for multivariate data analysis: a case study in teaching an undergraduate multivariate data analysis course. the american statistician, 70(4), 405–412. https://doi.org/10.1080/00031305.2016.1201005 weiland, t. (2019). the contextualized situations constructed for the use of statistics by school mathematics textbooks. statistics education research journal, 18(2), 18–38. https://doi.org/10.52041/serj.v18i2.138 zhou, j., zhang, z., li, z., & zhang, j. (2015). coarsened propensity scores and hybrid estimators for missing data and causal inference. international statistical review, 83(3), 449–471. https://doi.org/10.1111/insr.12082 ziemer, k. s., pires, b., lancaster, v., keller, s., orr, m., & shipp, s. (2018). a new lens on high school dropout: use of correspondence analysis and the statewide longitudinal data system. the american statistician, 72(2), 191–198. https://doi.org/10.1080/00031305.2017.1322002 47/47 bauder, julia & cave, libby (2025). conceptions of data literacy in the statistics education literature, iassist quarterly 49(4), pp. 1-47. doi: https://doi.org/10.29173/iq1156 endnotes 1 julia bauder is the social studies and data services librarian and director of the data analysis and social inquiry lab (dasil) at grinnell college. she can be reached by email: bauderj@grinnell.edu. 2 libby cave is a term librarian at grinnell college. she can be reached by email: caveelizabeth@grinnell.edu. 2_li_kong_pesja 2 1/15 li, yue; kong, nicole and pejša, stanislav (2017) designing the cyberinfrastructure for spatial data curation, visualization, and sharing, iassist quarterly 41 (1-4), pp. 1-15. doi: https://doi.org/10.29173/iq11 designing the cyberinfrastructure for spatial data curation, visualization, and sharing yue li1, nicole kong2, stanislav pejša3 abstract widely used across disciplines such as natural resources, social sciences, public health, humanities, and economics, spatial data is an important component in many studies and has promoted interdisciplinary research development. though an institutional data repository provides a great solution for data curation, preservation, and sharing, it usually lacks the spatial visualization capability, which limits the use of spatial data to professionals. to increase the impact of research-generated spatial data and truly turn them into digital maps for a broader user base, we have designed and developed the workflow and cyberinfrastructure to extend the current capability of our institutional data repository by visualizing the spatial data on the web. in this project, we added a gis server to the original institutional data repository cyberinfrastructure, which enables web map services. then, through a web mapping api, we visualized the spatial data as an interactive web map and embedded in the data repository web page. from the user’s perspective, researchers can still identify, cite and reuse the dataset by downloading the data and metadata and the doi offered by the data repository. general information users can also browse the web maps to find location-based information. in addition, these data was ingested into the spatial data portal to increase the discoverability for spatial information users. initial usage statistics suggest that this cyberinfrastructure has greatly improved the spatial data usage and extended the institutional data repository to facilitate spatial data sharing. keywords spatial information, data repository, visualization, cyberinfrastructure, gis introduction although spatial information is widely used in many disciplines, it is usually saved in very specific formats that are not familiar to most of the researchers and general information users. to view the spatial data files, users need to have some background knowledge and skills about gis software, which might not even be free. this software requirement hindered the wide adoption of spatial information. to overcome this barrier, the association of research libraries (arl) has set out the gis literacy project to educate and equip librarians with the gis skills necessary to provide access to spatial data in all formats since 1992 (association of research libraries 1999). as a result, the valueadded spatial reference services have benefited a range of users from gis experts in special libraries to casual users in public libraries (gluck & yu 1999; weimer & reehling 2006). in recent years, the emerging technology in web gis has shown promise for general information users to access spatial data via web maps (kong et al. 2014; batty et al. 2010). web maps provides an easy and direct way for any web users to browse spatial information without any software license restriction or learning curve. many spatial data providers, such as us census bureau and the u.s. geological survey (usgs), 2/15 li, yue; kong, nicole and pejša, stanislav (2017) designing the cyberinfrastructure for spatial data curation, visualization, and sharing, iassist quarterly 41 (1-4), pp. 1-15. doi: https://doi.org/10.29173/iq11 started to take advantage of the web mapping technology and provide online spatial data portals for users to view and download information. in academic settings, research-generated spatial data is experiencing an exponential growth due to the availability of new sensor technologies and the broader adoption of spatial thinking skills across disciplines (gregory et al. 2015; matei et al. 2007; kong 2015). although institutional data repositories offer a great platform for data management, curation and sharing, they usually lack the web-based visualization capabilities for spatial data sharing. adding the spatial data visualization function to institutional data repository can benefit any information users by providing an online interactive map, so that users can view the research generated spatial data in their browser in the same way as they can in any other online map tools, such as google maps. for gis professionals, the visualization function can also help them to judge if the data is suitable for their research before downloading. in this article, we introduce the project that we have developed to streamline the spatial data curation, visualization and sharing by connecting our institutional research data repository with the library’s gis server set and spatial data portal. based on our campus users’ needs, we have prototyped the theoretical workflow to curate, visualize and share the research-generated spatial data. with the prototyped workflow, we designed, optimized, and adjusted our existing cyberinfrastructures to fulfill the requirements. in our design, the research data is curated using the institutional repository, visualized using the library’s gis server, and then the generated web map is embedded in the data publication webpage via web mapping api (application programming interface). to promote the data usage, we also share the data via our spatial data portal in addition to the institutional data by ingesting the metadata and link the published dataset from the portal. additionally, we developed strategies to track the data usage statistics for both web map view and data download. we would demonstrate our first successful test case of spatial data publication, using this improved cyberinfrastructure and workflow. the initial statistics show that our cyberinfrastructure design has greatly improved spatial data usage and extended the institutional data repository to facilitate spatial data sharing. background spatial data and gis technology development spatial data is data that is associated with a place on the earth’s surface, implicitly or explicitly (iso/tc 211 2009). it includes not only digital maps in various themes, but also any tabular data with location information such as census datasets, housing, marketing, and social media data. it provides a spatial dimension to simplify complex and dense information, and allows for the review of patterns, relationships and trends which may not be obvious in statistical and non-spatial datasets (knowles & hillier 2008; ridge et al. 2012; galbraith & coonin 2001; goodchild 2000). spatial data serves as an important component in various disciplines as well as promotes interdisciplinary research development (kong et al. 2016). in science, technology, engineering, and mathematics (stem) disciplines such as environmental science, agriculture and civil engineering, spatial data has been widely used since the 1980’s with the emergence of the geographic information systems (gis) software (diamond & wright 1988; goodchild et al. 1993; corwin & wagenet 1995). in the humanities and social sciences, it is only in recent years with the “spatial turn” that researchers began to realize the power of spatial data and started to learn and use this type of data for their research (knowles 3/15 li, yue; kong, nicole and pejša, stanislav (2017) designing the cyberinfrastructure for spatial data curation, visualization, and sharing, iassist quarterly 41 (1-4), pp. 1-15. doi: https://doi.org/10.29173/iq11 2000; warf & arias 2008; bodenhamer et al. 2010). as spatial data have such a broad user population across disciplines and technology levels, it is essential to design effective and efficient infrastructure and systems to curate, preserve, and share research-generated spatial data so that different user groups can benefit from it. due to its nature of recording the relationships between geographic features and their attribute information, spatial data is often saved in a more complex format comparing to other regular numeric or text files. the typical spatial data formats include shapefile, coverage, geodatabase or raster dataset, which require specific gis software, such as arcgis developed by the environmental scientific research institute (esri). although some open source gis software exist, they were designed for professional gis users and developers who already developed the understanding about geoinformatics. with the development of web technology in the early 21st century, a service-oriented architecture and web-based system was introduced to gis technology, extending the task oriented desktop-based system. unlike desktop gis, which is intended for professional users with training and experience in gis, web gis is intended for a broader audience, including people without any knowledge about gis (kidd 2010; batty et al. 2010). it is designed to be simple, intuitive, and convenient to use. visualizing data with web gis technology is a streamlined process which can be completed within a few clicks. google and microsoft launched online global basemap imagery, enabling people to access free geospatial information products from home and mobile devices. this exposure also lead to consumers’ increasing familiarity with information in map form. in 2004, esri released arcgis server, a software that makes geographic information available to anyone with an internet connection. government agencies, universities, and organizations began to build gis infrastructure with web technology to share their data, maps and services online. institutional repository over the last decade there has been a push in the research communities across the disciplines to adjust and modify the foundation of scholarly communication, which often manifests itself as calls for a more open and transparent science. reproducibility, replicability, and verifiability are increasingly being seen as a fundamental attributes of good science and responsible research (casadevall & fang 2010; doorn et al. 2013). access to the data that underlie the scholarly journal publications is an unavoidable prerequisite of transparent science. some research areas have access to dedicated data repositories around which research communities gather and develop data archival practices, data publishing standards, and tools dedicated to data visualization and analysis. these repositories, however, are relatively rare. many researchers work in disciplines that are lacking such an infrastructure and are facing difficult choices when project funders require a mandatory data management plan (dmp) that contains a proviso for access and preservation of data. this provides an opportunity for establishing institutional data repositories, such as the purdue university research repository (purr, https://purr.purdue.edu/). purr allows for data deposit for all purdue faculty, staff, students and their collaborators. purdue researchers can use purr, to share and manage their research data with their collaborators in a secured environment, to publish and disseminate datasets with dois, to track and measure the impact of sharing their data, and to archive their data for long-term preservation and reuse. the purr system has been developed based on hubzero, an open source software platform for building powerful web sites that support scientific discovery, learning, and collaboration (mclennan & kennell 2010). purr is a combination of a virtual research environment (vre), content management system 4/15 li, yue; kong, nicole and pejša, stanislav (2017) designing the cyberinfrastructure for spatial data curation, visualization, and sharing, iassist quarterly 41 (1-4), pp. 1-15. doi: https://doi.org/10.29173/iq11 (built on joomla), and a data publication platform. purr provides secure storage for research data and a safe space for online collaboration. the results of such collaboration can be published as datasets and shared with the research community at large. researchers can publish through purr not only their primary data in a variety of types and formats, such as observational and experimental data, software code, audio and video files, or gis files, but also the necessary documentation, such as the research procedures, code books, technical drawings. the process of data publication and archiving is straightforward in purr (figure 1). it is a team effort between the researchers, the purr administrators, and the library subject liaisons. it starts with the researchers, who select and upload the data files and provide key metadata, including the dataset authors, description, and the license under which dataset can be used and shared. then the dataset is submitted for review by the subject liaisons or the purr administrator. the liaisons assess the technical aspects of the dataset, primarily whether all the necessary domain-specific scholarly attributes are present. once the dataset is approved, it is published and the purr repository assumes custody of the dataset. all published datasets are provided with unique digital object identifiers (doi) so that they can be easily identified, cited and reused. after a 30-day grace period, an archival information package (aip) is created and the archival package is deposited in a dark archive provided through the mataarchive consortium for a long-term preservation. the aip contains not only the descriptive metadata, but also technical and preservation metadata that are essential for a successful preservation and long-term access to the archived data. figure 1. purr data publication and data archiving workflow the usage statistics of the published dataset in purr, i.e. the number of views and downloads, are tracked and visible to the users of purr. figure 2 shows an example of a published dataset in purr. purr allows for version tracking. the individual versions are linked and all versions of a dataset can be accessed in the ‘versions’ section of the published dataset webpage. researchers are also able 5/15 li, yue; kong, nicole and pejša, stanislav (2017) designing the cyberinfrastructure for spatial data curation, visualization, and sharing, iassist quarterly 41 (1-4), pp. 1-15. doi: https://doi.org/10.29173/iq11 to post questions to the authors of the datasets or leave comments in the 'reviews' section of the purr publication. figure 2. example of a published dataset on purr. it allows users to read the metadata information and supporting document about the dataset, download the data, cite the data, as well as view the versions and usage statistics, leave review comments, and ask questions. purr was initially designed to provide generic data publication and archival functions for data in all disciplines. different from many other discipline-specific data repositories which provide data visualization and analysis functions, purr lacked customizations for special datasets, including the spatial data discussed in this article. purr’s development team was working with hubzero team over the last few years to add more specialized data functions into purr, such as visualizing tabular datasets and images. however, spatial data visualization was not available in purr because it requires additional hardware and software support, which was relatively more challenging compared to some other common file formats. gis cyberinfrastructure review and our existing status as geospatial information has increased exponentially in recent years, many academic libraries have started to build their own gis cyberinfrastructure to support the user communities (lage 2007; kollen et al. 2013). these efforts include both setting up their own gis servers to support webbased research data publications, and developing or deploying spatial data portals to facilitate spatial data discovery (stowell bracke et al. 2008; carlson et al. 2011; florance et al. 2015). the gis servers can be used to create and manage web map services and web mapping applications. in an academic library setting, this server environment facilitates mapping of research-generated spatial data online. 6/15 li, yue; kong, nicole and pejša, stanislav (2017) designing the cyberinfrastructure for spatial data curation, visualization, and sharing, iassist quarterly 41 (1-4), pp. 1-15. doi: https://doi.org/10.29173/iq11 the spatial data portals are web applications that facilitate discovery, preview and retrieval of spatial data. at purdue university libraries, the gis team has made efforts in both directions in setting up its gis cyberinfrastructure. this in-house infrastructure provides powerful, reliable, and customizable services for storing, organizing and sharing gis assets (figure3). figure 3. purdue university libraries’ gis cyberinfrastructure includes gis server set for geodatabase, geoporcessing, web mapping, and file storage. this server set is also connected to purdue geodata portal for spatial data discovery. the gis server cluster, which includes an arcgis server, an enterprise geodatabase and two geoprocessing servers, is essential to creating and managing map services, applications and data. the arcgis server is used for publishing web map services and hosting web applications, which helps research data gain a broader impact by sharing the interactive maps online. it also acts as a file server to accommodate file download requests for offline maps and other spatial data sharing purposes. using arcgis and microsoft sql technology, the geodatabase supports multi-user editing, security management and version controls in addition to other functions. the geodatabase and arcgis server are used to manage data from purdue library map collections and purdue research generated spatial datasets. in addition, two state-of-art high-performance compute nodes are dedicated to running computationally intensive geoprocessing jobs, which significantly decreases the processing time for users who have high computing demands. spatial data portals usually provide a web map interface which allows users to search for spatial data by both map location and keyword. connecting the curated spatial data from an institutional repository to the spatial data portal can help information users to develop a better understanding about the background of the dataset as well as correctly cite the data source. it can also increase the discoverability of the research dataset from the institutional data repository. to develop our institutional spatial data portal, we deployed the open geoportal project framework. in addition to connecting the metadata contributed from collaborative universities in the open geoportal project, we also connected our open geoportal instance to metadata shared by our state agencies and our library’s own spatial data, including the map collection and research data. although this geoportal server is a separate server from the gis server cluster, they are connected for through the metadata when the geoportal server needs map preview and download functions from purdue spatial data. this infrastructure is upgraded as needed and has been kept up-to-date since created over four years ago. it has supported various spatial data research projects across disciplines. 7/15 li, yue; kong, nicole and pejša, stanislav (2017) designing the cyberinfrastructure for spatial data curation, visualization, and sharing, iassist quarterly 41 (1-4), pp. 1-15. doi: https://doi.org/10.29173/iq11 figure 4. geodata portal at purdue supports spatial data discovery, preview and download function for data contributed by collaborative universities in open geoportal project, state agencies, and purdue curated spatial data. additional spatial data service needs though the purr and gis cyberinfrastructure provided excellent tools for researchers to curate their research data and share online maps, we still have received additional service requests to meet researchers’ expectations. one particular service area was to create customized spatial data sharing pages, so that the particular data user group could easily preview, download, cite and read more background information about the spatial dataset. although purr provided a convenient way for researchers to curate and share their data, it did not offer a map preview function for potential users to view the map without downloading it. thus, requiring users to download and view gis data with professional software, it limited the data sharing to professional gis users. the geodata portal provided data preview and download functions, but it did not have a specific web link for a particular dataset. the researchers had to describe how to find their dataset in the portal to their potential data users. the gis server cluster offered the opportunity possibility to meet the researchers’ expectations, but we had to make an extra effort to create a customized webpage for the particular dataset. one example of this kind of request is from a researcher whose project was about floodplain mapping using soil information (sangwan & merwade 2015). the research lab had developed an innovative approach to generating floodplain maps for the state. although most of the project is conducted using gis, their findings (i.e. floodplain map) can benefit a larger group of the population beyond gis professionals, including homebuyers, property insurance providers, emergency management agencies, urban planners, etc. it is the research lab expects to widely share their research findings as online interactive maps so that users can zoom in to their area of interest to explore more detailed information about the possibility of floods. in addition, sharing the data as a set of gis files can help gis professionals reuse their findings in other related studies. to accommodate this researcher’s needs, we first investigated the existing servers and created a customized webpage to help disseminate the dataset (figure 5). the webpage was 8/15 li, yue; kong, nicole and pejša, stanislav (2017) designing the cyberinfrastructure for spatial data curation, visualization, and sharing, iassist quarterly 41 (1-4), pp. 1-15. doi: https://doi.org/10.29173/iq11 developed using the arcgis server web map service and javascript api. this page has greatly helped the researcher to communicate his research findings with his potential user groups and funding agencies. the usage statistics have shown that the webpage has received almost 5,000 visits since its launch to the time that we developed this new project to enhance its functionality in 2016 (figure 6). although the initial dataset shared via our webpage is just for indiana, it attracts visitors from many other states in u.s. and countries all over the world. figure 5. customized webpage to disseminate the flood map information as both online interactive maps and downloadable gis files. figure 6. usage statistics of the customized flood map data sharing webpage. with this initial success, we looked for solutions to expand this data sharing solution to our institutional data repository (purr) so that the dataset could be easily cited. another advantage of connecting purr with this data sharing solution is that purr offers data downloading option. this will free up gis server space for file hosting. in addition, purr provides an archival solution for the dataset, 9/15 li, yue; kong, nicole and pejša, stanislav (2017) designing the cyberinfrastructure for spatial data curation, visualization, and sharing, iassist quarterly 41 (1-4), pp. 1-15. doi: https://doi.org/10.29173/iq11 so that the information can be safely preserved for a long period without the hassle of computer software/hardware upgrade. results and conclusion project design theoretical model with an initial assessment of our existing cyberinfrastructure for purr and the gis servers, we have drawn the conclusion that the most feasible solution to connect these two systems is to extend purr with spatial visualization capability supported from the gis servers. two factors helped us to make the decision. first, purr has already been used by many researchers at our institution. by the time this article was written, there were more than 3,500 registered researchers on purr. second, purr provides a great platform to manage users, and a user-friendly interface to curate and publish datasets. in our theoretical model, spatial data users can start the data publication process by uploading their datasets in purr, and publish it with a doi. once they enable the spatial data visualization functionality, map services can be created based on their datasets using our gis server set. the map service makes it possible to visualize the spatial data on the web using the javascript api. then, the visualization will be integrated into purr’s data publication webpage to offer a visuallyappealing and interactive preview. the spatial visualization enables users to explore the features on various scales, and zoom into an area of interest. the map visualization allows for user interactions beyond a textual and static description, and provides users with a convenient way to access more indepth information embedded within the dataset. furthermore, to increase discoverability beyond the institutional data repository, we expect to ingest the spatial data into our geodata portal. thus, general spatial information users can find the data and preview the map and its metadata from the portal. if they are interested in downloading the data, users will be redirected to the dataset webpage in purr. by combining the data sharing and preservation platform of the institutional data repository with a spatial visualization from the gis cyberinfrastructure, this efficient and comprehensive model solves the challenges in spatial data management, increases the discoverability and meets the needs of researchers and collaborators (figure 7). figure 7. theoretical model of connecting institutional data repository (purr) with the gis cyberinfrastructure. 10/15 li, yue; kong, nicole and pejša, stanislav (2017) designing the cyberinfrastructure for spatial data curation, visualization, and sharing, iassist quarterly 41 (1-4), pp. 1-15. doi: https://doi.org/10.29173/iq11 workflow with the theoretical model, we further list the detailed spatial data curation workflows shown in figure 8. first, the spatial dataset is submitted and published in purr and given a doi. to initiate this step, a project is created and shared with all collaborators by the researcher. during this process, project description, funding source, and other information need to be collected. after the storage space is allocated, the dataset can be uploaded and submitted for review by the data management specialist and subject librarian using the purr publication wizard. upon approval, this dataset is published and a doi is assigned, so that the dataset is openly accessible for use and discovery. the dataset is now preserved and archived in purr, and can be cited. while the dataset is curated in purr, it is also copied into our geodatabase and published as a map service using the gis server. a web map application is created with the map service overlain on a basemap. users can pan and zoom in or out to their area of interest. an overview map is included for users to have a broader picture, and users can choose a preferred basemap using the basemap switch. after that, the web map application is embedded in the project description in purr to make the data previewable and interactive, along with a brief introduction on how to use the map and download the data. with this comprehensive insight into the dataset, users can download the data from the repository and cite it with doi. finally, to increase usability spatial data is ingested into our geodata portal. metadata is created using the federal geographic data committee (fgdc) content standard for digital geospatial metadata (csdgm), including the link to the preview layer from the gis server, and the link to the purr page for the dataset. upon users’ download request, they will be directed to the dataset’s webpage in purr. figure 8. workflow for spatial data publication. deployment in the deployment process, we tried to utilize the existing functions in both purr and the gis server set and minimize our development efforts. we took the floodplain map for united states as the first sample dataset for this project. for the first step in the workflow, purr already provides excellent functionality for data submission and publication. therefore, there was no further development requirement in this step. however, since many spatial datasets are large, researchers usually organize them into smaller files by region, which created challenges in purr to organize these individual files into one data collection. for example, when we worked with the floodplain map data for the united states, the researcher divided the map into 50 smaller state-level datasets, because the smaller files make the data uploading and downloading process much easier. in this case, we recommended they 11/15 li, yue; kong, nicole and pejša, stanislav (2017) designing the cyberinfrastructure for spatial data curation, visualization, and sharing, iassist quarterly 41 (1-4), pp. 1-15. doi: https://doi.org/10.29173/iq11 publish each small file separately with an independent doi. then we used purr’s data series function to list all these datasets as one data series (figure 9). to connect the dataset from purr with our gis servers, we have to manually check the quality of the dataset to verify if it fits the map service publication requirement, then initiate the map service publication process. to connect both systems, we developed a javascript-based web application template to visualize the published map service. in this template, the map service information has to be updated based on the new dataset that is published. then, this web map application (or the javascript code) was updated in purr so that the interactive map could be embedded into the dataset’s webpage in purr. to be compatible with the purr system, the map service and web map application were both secured with https connections on our gis server. at the final step, we generated fgdc standard metadata based on the generic metadata information obtained from purr, and updated the data preview link according to the map services created on our gis server. this standard metadata was ingested into our geodata portal’s indexed searching database. the metadata ingestion is an existing function for our geodata portal. in order to make our geodata portal recognize the dataset contributed from purr, we modified the data download link in our portal to redirect the data users to purr’s data publication webpage. with minimum development efforts, including a javascript web mapping application template and the modification or geodata portal download link, we were able to connect purr with the gis cyberinfrastructure and extended our institutional repository by adding the spatial data visualization functionality. 12/15 li, yue; kong, nicole and pejša, stanislav (2017) designing the cyberinfrastructure for spatial data curation, visualization, and sharing, iassist quarterly 41 (1-4), pp. 1-15. doi: https://doi.org/10.29173/iq11 figure 9. the spatial data series publication webpage with an interactive map. initial assessment the datasets have been published in purr for six months. based on the visit statistics during this limited period, there were 5,683 total visits on the published dataset’s webpages and 842 total downloads. the spatial visualization integration has been tested for about one month at the time this article was written. the online interactive map has received more than 500 visits even though it has been available for a very short time period. the statistics suggest that the datasets and the spatial visualization have received great attention by information users. there are no statistics at this time to track users’ activities on this webpage, such as the interactions they had with the online map before they decide to either download the data or leave the webpage. we believe that the spatial visualization can help users to better understand the information contained in the dataset before they decide to either download the data or be satisfied with the information provided on the map. we will develop more logging scripts to capture user activities on the webpage to help us further understand users’ needs. discussion 13/15 li, yue; kong, nicole and pejša, stanislav (2017) designing the cyberinfrastructure for spatial data curation, visualization, and sharing, iassist quarterly 41 (1-4), pp. 1-15. doi: https://doi.org/10.29173/iq11 by connecting the institutional data repository to the gis server, we were able to expand the spatial data sharing capability to serve a broader information user group. the spatial data visualization feature can help any web users to view spatial information related to their interested location without gis software or skills. thus, it helps to communicate the research findings to a broader audience. from a technical perspective, our efforts have suggested that such an expansion is feasible with limited development needs. many academic libraries already have institutional data repositories and gis server environments, or are in the process of creating that cyberinfrastructure. our project can serve as an example to demonstrate the benefit of connecting these two sets of servers to create a better option for sharing spatial data. there are still three manual steps in our workflow to share the spatial data with an interactive map visualization. these three manual steps include creating map services on the gis server, updating the purr data sharing webpage with the web map application, and updating the fgdc metadata with purr publication information. we expect to automate or semi-automate these steps in the future if more resources are available. once this happens, researchers will experience a more streamlined process to upload and share their spatial data using purr interface. a good data sharing practice not only needs an excellent cyberinfrastructure as the platform to support the data sharing mechanisms, but also requires researchers to fully recognize the value of data sharing, and to build up good practices around the data sharing process. in our practice, we have learned that there are several key points that the researchers and their graduate students need to pay special attention before publishing and sharing their spatial data. first, they need to develop good data management practices in order to prepare the data for publication. this includes critical evaluation of their datasets, and creation of human readable documentation in order to enable others understand and use their datasets, etc. in our project, we have worked closely with the researcher’s team to learn about the dataset and created appropriate metadata for the data repository and spatial data portal. second, they need to understand the data sharing license options, data ownership, and data publication options in order to share the data as they desire or as required by their funding agencies. although we have prepared a good environment in this project to facilitate spatial data sharing, more education programs need to be developed in order to improve the data sharing practices. references association of research libraries, 1999. the arl geographic information systems literacy project, washington, dc. available at: http://www.istl.org/06-fall/refereed.html [accessed may 1, 2015]. batty, m. et al., 2010. map mashups, web 2.0 and the gis revolution. annals of gis, 16(1), pp.1–13. available at: http://www.tandfonline.com/doi/abs/10.1080/19475681003700831 [accessed september 3, 2014]. bodenhamer, d.j., corrigan, j. & harris, t.m. eds., 2010. spatial humanities: gis and the future of humanities scholarship, indiana university press. available at: http://site.ebrary.com/lib/purdue/reader.action?docid=10767195&ppg=20 [accessed may 4, 2015]. carlson, j. et al., 2011. determining data information literacy needs: a study of students and research faculty. portal: libraries and the academy, 11(2), pp.629–657. available at: http://muse.jhu.edu/content/crossref/journals/portal_libraries_and_the_academy/v011/11.2. 14/15 li, yue; kong, nicole and pejša, stanislav (2017) designing the cyberinfrastructure for spatial data curation, visualization, and sharing, iassist quarterly 41 (1-4), pp. 1-15. doi: https://doi.org/10.29173/iq11 carlson.html [accessed november 6, 2014]. casadevall, a. & fang, f.c., 2010. reproducible science. infection and immunity, 78(12), pp.4972–5. available at: http://www.ncbi.nlm.nih.gov/pubmed/20876290 [accessed march 22, 2017]. corwin, d.l. & wagenet, r.j., 1995. applications of gis to the modeling of nonpoint source pollutants in the vadose zone: a conference overview. journal of environmental quality, 25(3), pp.403–411. diamond, j.t. & wright, j.r., 1988. design of an integrated spatial information system for multiobjective land-use planning. environment and planning b, 15(2), pp.205–214. doorn, p., dillo, i. & van horik, r., 2013. lies, damned lies and research data: can data sharing prevent data fraud? international journal of digital curation, 8(1), pp.229–243. available at: http://www.ijdc.net/index.php/ijdc/article/view/8.1.229%5cnhttp://www.ijdc.net/index.php/ij dc/article/view/8.1.229/308. florance, p. et al., 2015. the open geoportal federation. journal of map & geography libraries, 11(3), pp.376–394. galbraith, j. & coonin, b., 2001. gis in business: building a core collection for business geographics. reference & user services quarterly, 41(1), pp.9–17. gluck, m. & yu, l., 1999. geographic information systems: background, frameworks and uses in library. in e. a. chapman, ed. advances in librarianship. emerald group publishing limited, pp. 1–38. goodchild, m. et al., 1993. environmental modeling with gis, null, ed., goodchild, m.f., 2000. gis and transportation: status and challenges. geoinformatica, 4(2), pp.127– 139. gregory, i. et al., 2015. geoparsing, gis, and textual analysis: current developments in spatial humanities research. international journal of humanities and arts computing, 9(1), pp.1–14. available at: http://www.euppublishing.com/doi/abs/10.3366/ijhac.2015.0135?journalcode=ijhac [accessed march 19, 2015]. iso/tc 211, 2009. report from stage 0 project 19150 geographic information ontology, available at: http://metadata-standards.org/document-library/documents-by-number/wg2-n1251n1300/wg2_n1292_tc211n2705_report_19150.pdf. kidd, j.c., 2010. web-based mapping: an evaluation of free mapping applications and web gis for library reference services. university of north carolina at chapel hill. available at: https://cdr.lib.unc.edu/indexablecontent?id=uuid:f64751da-cb95-47aa-8d4f3ec5338420af&ds=data_file. knowles, a.k., 2000. special issue: historical gis: the spatial turn in social science history. social science history, 24(3), pp.451–470. knowles, a.k. & hillier, a., 2008. placing history: how maps, spatial data, and gis are changing historical scholarship, esri, inc. available at: https://books.google.com/books?hl=en&lr=&id=vn1v7rzhsqec&pgis=1 [accessed april 12, 2016]. kollen, c. et al., 2013. spatial data catalogs.pdf. journal of map & geography libraries, 9(3), pp.276– 295. kong, n. et al., 2016. an interdisciplinary approach for a water sustainability study. papers in applied geography, 2(2), pp.189–200. available at: http://www.tandfonline.com/doi/full/10.1080/23754931.2015.1116106. kong, n., 2015. exploring best management practices for geospatial data in academic libraries. journal of map & geography libraries, 11(september), pp.207–225. kong, n., zhang, t. & stonebraker, i., 2014. common metrics for web-based mapping applications in academic libraries. online information review, 38(7), pp.918–935. available at: http://www.emeraldinsight.com/doi/abs/10.1108/oir-06-2014-0140 [accessed january 18, 2015]. 15/15 li, yue; kong, nicole and pejša, stanislav (2017) designing the cyberinfrastructure for spatial data curation, visualization, and sharing, iassist quarterly 41 (1-4), pp. 1-15. doi: https://doi.org/10.29173/iq11 lage, k., 2007. cataloging digital geospatial data. journal of map & geography libraries, 3(1), pp.39–55. matei, s.a. et al., 2007. visible past: learning and discovering in real and virtual space and time. first monday, 12(5). available at: http://ojphi.org/ojs/index.php/fm/article/view/1836/1720 [accessed march 19, 2015]. mclennan, m. & kennell, r., 2010. hubzero: a platform for dissemination and collaboration in computational science and engineering. computing in science & engineering, 12(2), pp.48–53. ridge, o., stefanidis, a. & ber, d.o.e., 2012. spatiotemporal data mining in the era of big spatial data : , pp.1–10. sangwan, n. & merwade, v., 2015. a faster and economical approach to floodplain mapping by using soil information. journal of american water resources association, 51(5), pp.1286–1304. stowell bracke, m., miller, c.c. & kim, j., 2008. adding value to digitizing with gis. library hi tech, 26(2), pp.201–212. available at: http://www.emeraldinsight.com/doi/abs/10.1108/07378830810880315 [accessed december 28, 2014]. warf, b. & arias, s., 2008. the spatial turn: interdisciplinary perspectives b. warf & s. arias, eds., new york, ny: routledge. available at: http://reader.eblib.com.ezproxy.lib.purdue.edu/(s(atxzqvrkm1hx0jd2yxokucki))/reader.aspx?p =356413&o=58&u=jwvgaimz8bm%3d&t=1430763856&h=1d90639e400c6ab96d982f7df72 81ef697c95df2&s=36065559&ut=130&pg=1&r=img&c=-1&pat=n&cms=-1&sd=2# [accessed may 4, 2015]. weimer, h. & reehling, p., 2006. a new model of geographic information librarianship : description, curriculum and program proposal. journal of education for library and information science, 47(4), pp.291–302. end-notes 1 yue li is a gis analyst at purdue university libraries. 2 nicole kong is an assistant professor and gis specialist at purdue university libraries. she can be reached by email: kongn@purdue.edu 3 stanislav pejša is a data curator at purdue university libraries. 1/25 silverstein, p, bottesini, j, karcher, s, & elman, c (2025) introducing the journal editors discussion interface, iassist quarterly 49(2), pp. 1-25. doi: https://doi.org/10.29173/iq1146 the creative commons-attribution-noncommercial license 4.0 international applies to all works published by iassist quarterly. authors will retain copyright of the work and full publishing rights. introducing the journal editors discussion interface priya silverstein1, julia g. bottesini2,3, sebastian karcher4, colin elman5 abstract journal editors play an important role in advancing open science in their respective fields. however, their role is temporary and (usually) part time, and therefore many do not have enough time to dedicate towards changing policies, practices, and procedures at their journals. the journal editors discussion interface (jedi, https://dpjedi.org) is an online community for journal editors in the social sciences that was launched in 2021, consisting of a listserv and resource page. jedi aims to increase uptake of open science at social science journals by providing journal editors with a space to learn and discuss. in this paper, we explore jedi’s progress in its first two years, presenting data on membership, posts, and from a members survey. we see a reasonable mix of people participating in listserv conversations and there are no detectable differences among groups in the number of replies received by thread-starters. the community survey suggests jedi members find conversations and resources on jedi generally informative and useful and see jedi primarily as a community to get honest opinions from others on editorial practices. however, jedi membership is not as heterogeneous as would be ideal for the purpose of the group, especially when considering geographic diversity. keywords open science, scholarly publishing, community, social science introduction changing the academic ecosystem is difficult and requires all stakeholders to do their part and work together (rhys evans et al., 2022). within this ecosystem, journals have considerable capacity to incentivize and shape scholarly behavior. their influence offers a broad opportunity for disciplines to take a more intentional approach to open science. open science is “an umbrella term reflecting the idea that scientific knowledge of all kinds, where appropriate, should be openly accessible, transparent, rigorous, reproducible, replicable, accumulative, and inclusive, all which are considered fundamental features of the scientific endeavour” (parsons et al., 2022). research communities in different disciplines have begun to develop stronger norms for openness (silverstein et al., 2024), including psychology (button et al. 2013; nosek et al., 2015; levenstein & lyle, 2018; nosek et al. 2022), economics (christensen and miguel 2018; miguel et al., 2014), education (makel & plucker, 2014; cook et al., 2018; gehlbach & robinson, 2018; mcbee et al., 2018; fleming et al., 2021), political science (lupia & elman, 2014), public health (harris et al., 2018; peng & https://doi.org/10.29173/iq1146 https://dpjedi.org/ https://creativecommons.org/licenses/by-nc/4.0/ 2/25 silverstein, p, bottesini, j, karcher, s, & elman, c (2025) introducing the journal editors discussion interface, iassist quarterly 49(2), pp. 1-25. doi: https://doi.org/10.29173/iq1146 hicks, 2021), science and technology studies (maienschein et al., 2019), and sociology (freese, 2007; freese & king, 2018)6. scholars are socialized into their respective research community, inculcated with the belief that they are building on and contributing to a group endeavor (merton, 1949, 309-316). research is a communal, cumulative endeavor in which scholars learn from each other (bird, 2014; national academies of science, 2019, p24). the validity of a set of findings depends on researchers correctly applying methods for collecting and analyzing data. when scholars “show their work,” other members of their research communities can evaluate their descriptive and causal inferences. hence scholars have a responsibility to make their work evaluable, including making their research transparent and sharing the data and materials that underlie their findings (elman et al., 2018; elman & lupia, 2016). notwithstanding its potential to deliver a range of benefits, the shift towards open science remains a work-in-progress. open science advocates have argued that changing research culture requires following multiple strategies, including: developing initial infrastructure to make open science possible; creating more sophisticated and fine-grained tools to make open science easier; increasing the visibility of open science practices to show that they are customary and expected; generating incentives to reward open science, including linking it to publications, funding and hiring; and institutional stakeholders requiring the researchers they serve to use open science practices (nosek, 2019; mellor, 2021). different institutional stakeholders can influence how research is conducted by pulling on one or more of these levers, including funders, disciplinary associations, data repositories, universities, and publishers and journals. indeed, several of these stakeholders have already encouraged open science practices. for example, us government entities such as the national science foundation (2011, 2019, 2023), the uk’s research excellence framework (used to allocate public funding for universities’ research, www.ref.ac.uk), and the european parliament (2019) and commission (2023) encourage research transparency and data sharing. many academic associations’ ethics guidelines (e.g., the american anthropology association (2012), the american sociological association (2018)), as well as those of the national academies of science (2019) now encourage scholars to adopt a range of open science practices. governments, foundations, and public universities in latin america have fostered a vibrant culture of open access, with between 50% and 90% of periodical articles published in the region appearing open access through platforms such as scielo and redalyc, typically as “diamond” open access, i.e., without any charges to authors or readers (alperin, 2015, pp. 10-17). among the different institutional stakeholders, journals are some of the most influential institutions in the academic ecosystem (silverstein et al., 2024). this influence has not always been beneficial. for example, it is now widely acknowledged that journals’ preference for publishing positive findings helped to incentivize questionable research practices (christensen et al., 2019). but it is this capacity to incentivize and shape scholarly behavior which offers a broad opportunity to promote transparency and reproducibility. editors are particularly influential actors in the academic ecosystem because they are a major vehicle for organizing and disseminating academic communications, promoting knowledge, and producing https://doi.org/10.29173/iq1146 http://www.ref.ac.uk/ 3/25 silverstein, p, bottesini, j, karcher, s, & elman, c (2025) introducing the journal editors discussion interface, iassist quarterly 49(2), pp. 1-25. doi: https://doi.org/10.29173/iq1146 success signals for individual researchers (elman et al., 2018). editors make judgments about the content of their publication and shape their field’s substantive trajectory. because they also implement processes for reviewing, revising, and publishing research, they decide how their publication will address fundamental questions raised by the shift towards open science. as part of this engagement, editors have multiple opportunities to facilitate and instantiate open science, including: directly mandating relevant practices via author guidelines; providing information to authors in faqs and other guidance documents; offering the means to reward openness via badges; and operating as exemplars for other journals as well as the research communities they serve more generally (silverstein et al., 2024). if editors have so many opportunities to facilitate and instantiate open science, does this mean they are leading the way? top factor is a metric that reports the steps that a journal is taking to implement open science practices (https://topfactor.org). out of the 3,200 indexed journals, 2,412 (75%) receive a score of 0 out of 30 (with 30 being the highest score for implementing open science practices), and only 36 (1%) receive a score of 15 or over (top advisory board, n.d.). so, despite their potential for influence, many journal editors have not fully embraced open science practices. naaman et al. (2022) surveyed journal editors on their capability, opportunity and motivation to implement top guidelines, and found that the majority of editors did not see implementing top as a high priority compared with their other editorial responsibilities. they also identified several barriers to implementing open science policies, practices, and procedures, including a lack of time and resources. this can be especially difficult for editors when considering open science policies that have fieldor methodologyspecific nuances. for example, data sharing considerations in the social sciences can be more complex due to the use of human subjects data and resulting concerns ensuring identifiable and/or sensitive information about participants is not made public. data management personnel are very familiar with these concerns, and there are several considerations and solutions that can be implemented through a variety of data repositories. however, journal editors may simply not have the time to search for relevant resources, or the connections to have helpful conversations with experts on this topic. editors with easy access to information and expertise are better positioned to make well-founded decisions at their journals that will encourage authors to adopt open science practices. while editors have multiple potential sources of information and expertise, two may be especially useful. first, editors benefit when they can communicate easily with each other to consider the difficulty, merits, and implications of change. second, editors may benefit from the expertise of other institutional stakeholders. in the case of data sharing and research transparency, editors stand to gain considerably from interactions with domain repositories and other members of the data management community with experience and expertise in practices related to research openness (crosas et al., 2018, 18f.). for instance, data professionals are developing a range of mechanisms to facilitate the responsible sharing of sensitive research data, of which editors should be aware, and be equipped to suggest to authors (harvard university privacy tools project, 2020; kamath & ullman, 2020; levenstein et al., 2018; wood et al., 2020). despite the clear merit of these types of exchanges, it can be difficult for social science journal editors to participate in such dialogues. there are limited opportunities for them to share information with https://doi.org/10.29173/iq1146 https://topfactor.org/ 4/25 silverstein, p, bottesini, j, karcher, s, & elman, c (2025) introducing the journal editors discussion interface, iassist quarterly 49(2), pp. 1-25. doi: https://doi.org/10.29173/iq1146 each other and other stakeholders in the academic ecosystem, pool their collective wisdom, or make their hard-won expertise available to their successors and other new editors. this multifaceted communication gap slows the development of new editorial knowledge and prevents the emergence of what could be a powerful sense of community both among editors, and across institutional actors who share the goal of incentivizing robust social science, including data sharing and research openness. the journal editors discussion interface (jedi) (https://dpjedi.org) was created to help fill this gap. the current paper in this manuscript, we provide an overview of jedi’s history, its aims, and activities, and then describe data on three different aspects of jedi, for the period ranging from its official launch7 in march 2021 to march 2023. we present data concerning membership, activity on the mailing list, and from a community survey conducted a year after launching. based on these analyses, we then consider jedi’s successes and areas for improvement. we hope these details are instructive for others seeking to create structures to promote institutional change towards a more open science. history the qualitative data repository at syracuse university, together with other repositories from the data preservation alliance for the social sciences (data-pass), previously received an early-concept grants for exploratory research (eager) grant from the us national science foundation (nsf, grant #2032661) to establish a proof-of-concept for jedi, an online community of social science journal editors and data professionals that focuses on open science. data-pass – a voluntary partnership of organizations created to archive, catalog and preserve data used for social science research – has a strong, multi-year record working with journal editors to encourage open science practices. prior to the launch of jedi in 2021, data-pass’s engagement with journal editors stretched back to 2016, with a series of annual workshops focused on open science themes. while these workshops were helpful in sharing information about open science with the attending editors, they suffered from some limitations. first, the workshops were only occasional events, while the challenges that editors face are ongoing, constantly changing, and need real-time responses. second, the workshops were mostly structured to follow an ‘outside-in’ model of communication, with the bulk of the time taken up with presentations by repository personnel to journal editors. while the editors appreciated the expertise being provided by the repositories, during the relatively brief discussion periods, journal editors were very eager to share their knowledge and experience with each other, and to be especially open to learning from the shared experiences of their peers. jedi addresses both of these shortcomings, providing an opportunity for continuous editor-to-editor engagement. aims jedi aims to facilitate convergence and consensus on key principles of open science, encourage adoption of a common language and set of norms, and contribute to the generation of innovative solutions and a fund of collective knowledge. given the many demands on editors’ time – and given that most editors face similar processual challenges – there is great value to their interacting with each other about these key issues, and pooling their collective wisdom, sharing lessons, examples, https://doi.org/10.29173/iq1146 https://dpjedi.org/ 5/25 silverstein, p, bottesini, j, karcher, s, & elman, c (2025) introducing the journal editors discussion interface, iassist quarterly 49(2), pp. 1-25. doi: https://doi.org/10.29173/iq1146 insights, and solutions. the benefits can be further multiplied if experts on relevant topics are included in the conversation. jedi seeks to generate that interaction and those benefits through building an online community of social science journal editors and “scholarly knowledge builders''. scholarly knowledge builders are experts in topics relevant to journal editing. when founding jedi, the representatives from the data repositories that form data-pass were the original scholarly knowledge builders, but one aim was to expand this to other open science experts and metascientists studying peer review. bringing together editors and scholarly knowledge builders encourages and facilitates continuous communication and learning. jedi combines features of an online forum and a traditional email listserv. jedi predominantly (but not exclusively) focuses on the aspects of the editorial process that concern data and code, and their management, citation, and accessibility; other aspects of research transparency; and reproducibility, replication, and verification. however, members are free to post about anything concerning editorial functions. this can include “day to day” concerns (e.g., how to secure reviewers), big picture questions (e.g., what does the future of publishing look like), and anything in-between. even when open science isn’t explicitly the topic of conversation, links can often be (and often are) made. for example, when discussing how to secure reviewers, someone may suggest having an open call for reviewers on a journal landing page (which advances transparency and diversity), or how reviewers may be more likely to accept requests if they are being asked specifically to review something relating to their expertise (e.g., the data and code reproducibility package). by offering journal editors the opportunity to draw on the expertise of other editors and scholarly knowledge builders, jedi aims to augment the readiness of the social science publishing community to generate and adopt best practices for open science. jedi offers the editors and editorial staff of social science journals a shared forum in which to ask and answer questions, pool information and expertise, and build a fund of collective knowledge. at the beginning of 2021, a dedicated community manager (ps) was hired to build and manage jedi. the online community takes the form of a google group, which has the benefits of a traditional listserv, but discussion threads are also archived and searchable online. activities jedi launched in march 2021. we have invited hundreds of social science journal editors (predominantly current, outgoing, and incoming editors from top journals across the social sciences [anthropology, criminology, economics, education, environmental science, geography, political science, psychology, and sociology]) and scholarly knowledge builders to join jedi. to help sustain momentum, we have invited some members to help “catalyze” conversation by posting on the listserv on specific dates when things are a bit quiet. we have aggregated conversations in a biweekly newsletter, summarizing threads, inviting input, and highlighting resources. we have curated resources from those recommended by members on the jedi listserv to share on our website (https://dpjedi.org/resources), including those on best practices in open science. this resources page brings together materials for editors in a way that is accessible and easily digestible. as it is crowdsourced through conversations on jedi, it ensures we are covering topics of interest to this group. https://doi.org/10.29173/iq1146 https://dpjedi.org/resources.html 6/25 silverstein, p, bottesini, j, karcher, s, & elman, c (2025) introducing the journal editors discussion interface, iassist quarterly 49(2), pp. 1-25. doi: https://doi.org/10.29173/iq1146 in addition to the google group and resources page, jedi seeks to continue the earlier data-pass tradition of organizing workshops, and has so far hosted two workshops on issues around open science in journal editing (2022: https://dpjedi.org/events/may-the-force-be-with-you; 2023: https://dpjedi.org/events/jedi-2023-annual-meeting). in addition to this, in 2022, ps led a hackathon to work on “a guide for social science journal editors on easing into open science” which has now been published (silverstein et al., 2024). jedi policies, procedures, and practices are decided on by an invited steering committee (meeting twice annually) consisting of the community manager, the associate director, editors of seven social science journals (currently covering anthropology, criminology, economics, education, political science, psychology, and sociology), and representatives from the repositories currently taking the lead on organizing jedi (databrary, harvard dataverse, inter-university consortium for political and social research [icpsr], the roper center for public policy, and qualitative data repository [qdr]). method the data used in the following analyses were obtained in three different ways, outlined below. a detailed description of all variables used in these analyses, and how they were obtained, can be found at https://osf.io/q7s5m. an anonymized version of the datasets and code to reproduce all analyses and figures is available at https://osf.io/sh5ry/. membership data the community manager collects and curates data about all new jedi members using both the information each member entered into the jedi signup form and publicly available information. these data are used to maintain a list of jedi members and include (but are not limited to) each member’s name and contact information, the date when they joined jedi, the institution or organisation they are associated with, journals they currently edit or have edited in the past, and their main field or discipline. based on this information, we can obtain the variables used in this analysis, including member join date, role, members’ location, members’ field or discipline, members’ presumed gender, and member type. a detailed description of how these variables were obtained can be found at s1.1. google group data we obtained this data through google takeout, which provides an mbox file. we also used the list of members in the google group, which can be downloaded directly from the webpage. we used a python script to extract the mbox data, and r to clean the data into a dataframe containing information on each post.8 variables obtained using this method for these analyses include linking variables (email addresses), post-level variables (subject, post text, posting date and weekday, and the post’s order in the thread), thread-level variables (thread subject, a thread id, the total number of posts in the thread), and several timing variables related to when the posts were made (e.g., how long threads remained active for). a detailed description of these variables can be found at s1.2. community survey in march 2022, we designed and launched a community survey in order to solicit impressions from the jedi community. we asked members about their current editorial position or role (e.g., editor-inchief, data professional), how long they had been an editor at any journal, and how many different https://doi.org/10.29173/iq1146 https://dpjedi.org/events/may-the-force-be-with-you/ https://dpjedi.org/events/jedi-2023-annual-meeting https://osf.io/q7s5m https://osf.io/sh5ry/ 7/25 silverstein, p, bottesini, j, karcher, s, & elman, c (2025) introducing the journal editors discussion interface, iassist quarterly 49(2), pp. 1-25. doi: https://doi.org/10.29173/iq1146 roles they had held. we then asked about the informativeness and relevance of the listserv and resources page, as well as the helpfulness of listserv replies. finally, we asked our members where they currently went to get help and learn about editorial practices, how important different community features were to them, the interestingness of different topics that have been explored, and which challenges our members would like help solving. in total, we received 126 responses — a reasonable response rate of 28.6–32.9 percent, based on our estimated membership of 382–440 members at the time — with 125 responses being included in the analyses.9 this estimate was obtained by including only members who joined before 31 august 2022, when data collection for this community survey ended. we explain the uncertainty in our membership numbers in the next section. the project was approved and classified as exempt by the syracuse university institutional review board (irb #: 22-047). a copy of the complete survey, with all questions, can be found at https://osf.io/8zhmr. data exclusions in any community, including jedi, members come and go — and we would like this to be reflected in our data. although we can now identify members who have left, unfortunately, we cannot know exactly when this happened. here, we dealt with this issue in two ways. first, where we describe jedi’s current membership as well as show membership numbers and characteristics over time, we only included data on those members that joined during the target period (before or during march 2023) and were still members as of july 2023 (n = 417). prior to this date, we did not have a procedure in place to track the ebb and flow of members, so it is possible that some members who have left the group between march and july of 2023 (and would otherwise be included in this dataset) are not present. however, we feel that this subset of our data is the best reflection of what our membership looked like during the target period. the second way we dealt with the issue of members who have left the group applies to any result or visualization that relates to posts on the listserv. because posts made by members who have left are still included in this dataset, we chose to keep those members in our data for those analyses. in those cases, the number of unique members included in the analyses is 479. results jedi membership jedi membership has grown at a steady pace (+13.5 members per month on average) since its official launch10, both through active recruitment and organically, increasing from 79 members in march 2021 to 417 members in march 2023. based on the first word of each member’s name — genderize.io11 determined 49.4% of jedi members to be men (n = 206) and 44.6% to be women (n = 186), indicating good gender diversity (figure 1). jedi is also diverse in terms of its members’ principal fields or disciplines, which span a large number of social sciences as well as other sciences, the humanities, and research-supporting fields like library sciences, data & research infrastructure, and publishing. the largest disciplinary groups stems from psychology with 29.0% of jedi members (n = 121), followed by political science with 15.8% (n = 66), economics (10.3%; n = 43), anthropology (8.4%; n= 35) and sociology (7.0%; n = 29; see figure 1 for a full breakdown of jedi member’s fields). there is room for improvement in terms of geographic location, however; jedi members are highly concentrated in the https://doi.org/10.29173/iq1146 https://osf.io/8zhmr 8/25 silverstein, p, bottesini, j, karcher, s, & elman, c (2025) introducing the journal editors discussion interface, iassist quarterly 49(2), pp. 1-25. doi: https://doi.org/10.29173/iq1146 americas (65.0% (n = 271); 60.7% (n = 253) in the us alone), and europe (25.9% (n = 108)). only 9.1% of members are not in europe or in the americas. in line with jedi’s mission, most members of jedi are editors (86.6%, n = 361) — that is, they have held or currently hold an editing-related position at one or more journals — while 13.4% (n = 56) are noneditors, primarily in the field of psychology (23.2%; n = 13) and data & research infrastructure (21.4%; n = 12). https://doi.org/10.29173/iq1146 9/25 silverstein, p, bottesini, j, karcher, s, & elman, c (2025) introducing the journal editors discussion interface, iassist quarterly 49(2), pp. 1-25. doi: https://doi.org/10.29173/iq1146 figure 1. jedi membership at a glance https://doi.org/10.29173/iq1146 10/25 silverstein, p, bottesini, j, karcher, s, & elman, c (2025) introducing the journal editors discussion interface, iassist quarterly 49(2), pp. 1-25. doi: https://doi.org/10.29173/iq1146 listserv posts all data and numbers in this section include everyone who was a member at any point during the target period (n = 479). after excluding newsletters and the initial “how to use jedi” post, from march 2021 until march 2023 (inclusive, 25 months), there were a total of 674 posts on the jedi listserv within 181 threads, giving an average of 7.2 new threads and 27 posts each month. the median thread length was 3 [iqr = 4], or one thread starter and two replies (figure 2). figure 2.histogram showing the length of jedi threads. under 30% of threads get no replies, while most threads get at least few replies. thread “lifespan” — that is, the time between the first post, or thread starter, and the last post on the same thread — is quite varied. threads that received at least one reply, had a median lifespan of 5.85 days, with lifespan ranging from 19 minutes to 50 weeks (iqr: 2.75 weeks). posting behavior over time posting frequency varies from month to month, with clear drops during the northern hemisphere summer and holiday months. overall cumulative posting was similar across both years (figure 3), although the end of the second year saw a drop in posting frequency, possibly related to the transition period between community managers. https://doi.org/10.29173/iq1146 11/25 silverstein, p, bottesini, j, karcher, s, & elman, c (2025) introducing the journal editors discussion interface, iassist quarterly 49(2), pp. 1-25. doi: https://doi.org/10.29173/iq1146 figure 3. posts per month over jedi’s first and second years (left axis), and cumulative posts (blue line, right axis) across each year. march of 2023 (28 posts) is not pictured to allow for a more direct comparison. discussion topics jedi discussions cover a variety of topics. the most popular threads (those with at least ten replies) during this time period were: how to arrange book reviews (16 posts), should all papers be published? (14 posts), positionality statements, maximum word limits in papers, and independently managed journals (12 posts each), triple masked reviewing, open data checking, and how to preserve sensitive data (11 posts each). a full list of topics on jedi in the target period is available at https://osf.io/q6ah5. member engagement and posting behavior jedi as a community aims to welcome both members who actively participate in the discussion and those who prefer to observe from the sidelines. this is reflected in our membership’s posting behavior: 68.3% of jedi members (n = 327) have never posted to the listserv. out of the 31.7% who have, 10.9% have both started a thread and replied to an active thread (n = 52), while 16.9% have only replied to other members’ threads, and a small number of users have only ever started a thread but never replied to any posts on jedi (4%; n = 19). posting behavior by member type as a relatively young community, jedi still relies on its team members and “catalyst” users to maintain engagement. figure 4 shows the number of unique active users — those who have posted at least once in that month (first panel, med = 18) — and the total number of posts (second panel; med = 25) https://doi.org/10.29173/iq1146 https://osf.io/q6ah5 12/25 silverstein, p, bottesini, j, karcher, s, & elman, c (2025) introducing the journal editors discussion interface, iassist quarterly 49(2), pp. 1-25. doi: https://doi.org/10.29173/iq1146 over the target period. the bottom panel shows the true proportion of each member type in jedi, making it evident that both jedi team members and catalysts are overrepresented among active users and post authors. despite this, regular members are clearly still participating in the conversation, and are doing so totally unprompted. it is worth noting that, although they are often prompted to post, jedi team members and catalysts also write unprompted posts. the proportion of each type of user stays relatively stable over time, potentially indicating that soliciting posts is still important for generating engagement at this stage. https://doi.org/10.29173/iq1146 13/25 silverstein, p, bottesini, j, karcher, s, & elman, c (2025) introducing the journal editors discussion interface, iassist quarterly 49(2), pp. 1-25. doi: https://doi.org/10.29173/iq1146 figure 4. the top panel shows the number of unique active users each month of the target period; the middle panel shows the total number of posts each month; and the bottom panel shows the proportion of each member type in the dataset. n = 479. poster representation on threads are some groups more likely to start or contribute to threads than others? we look at this by comparing the “population” distribution of different groups on jedi — the group of all 479 people who were jedi members at some point during the target period — with the distribution of those same https://doi.org/10.29173/iq1146 14/25 silverstein, p, bottesini, j, karcher, s, & elman, c (2025) introducing the journal editors discussion interface, iassist quarterly 49(2), pp. 1-25. doi: https://doi.org/10.29173/iq1146 groups in users who started threads or who made any post in the mailing list during the target period. those proportions are shown in figure 5. figure 5. proportion of different groups (from the top: whether a member is a catalyst, whether they are an editor, the gender associated with their first name, the country in which their main institution or organization is located, and their primary field or discipline. for each plot, the top bar represents the proportion among all jedi members during the target period, the middle bar represents users who started a thread, and the bottom bar represents users who made a post on jedi, whether or not it was a thread starter. respective ns = 479, 182, 675. catalyst members are clearly overrepresented both in those who start threads and respond to threads. this is not surprising; the aim of the catalyst program is to stimulate conversation in the community, and it is clearly working well. the fact that many people who are not catalysts are also participating is encouraging. together with figure 4, this suggests that, although catalysts (and jedi team members) are still initiating many of the conversations on jedi, they are not the only ones doing so, and at least part of the conversation is being driven by other members. further, the fact that a catalyst member started or responded to a thread does not automatically mean the thread was prompted; many people who are invited to become catalysts were already more active within jedi, and so have a tendency to start more threads and participate in conversation more than regular members. the second panel of figure 5 shows that, although non-editor members are in the minority, they are overrepresented among thread-starters and, to a lesser extent, posters. this probably reflects jedi’s history, as many jedi members who are not editors are members of data repositories, and work in https://doi.org/10.29173/iq1146 15/25 silverstein, p, bottesini, j, karcher, s, & elman, c (2025) introducing the journal editors discussion interface, iassist quarterly 49(2), pp. 1-25. doi: https://doi.org/10.29173/iq1146 data & research infrastructure or metascience. we see this as one of jedi’s main advantages — not only do editors get to hear from other editors in similar fields, but they also have access to others with valuable knowledge about editorial practices. in terms of gender, men are slightly overrepresented in starting and responding to threads, but not by much. this seems to be a general trend in online discussions (e.g., jarvis et al., 2022) and not so pronounced at jedi as to demand attention. geographic region groups seem relatively evenly represented as well, although it is clear that most of the traffic on jedi is generated by its us-based member majority. while this is not a problem per se, a more diverse mix of geographic regions would be an asset for the jedi community. it is slightly worrying that countries outside north america, the uk, and western europe seem to be underrepresented in thread starters and posters, which may suggest that not only are jedi members concentrated in the “global north”, but that members in other regions of the world are not participating in jedi exchanges as much. we further discuss steps to promote this in the discussion. finally, different disciplines seem to be appropriately represented in thread starters and posters. data & research infrastructure folks are overrepresented in thread-starters, which probably reflects datapass’s close connection to jedi and data repositories’ continued stewardship of jedi — representatives of five data repositories in the social sciences sit on jedi’s steering committee alongside the seven social science representatives, and regularly post on jedi. psychology is also overrepresented; we speculate on why that may be in the discussion. overall, figure 5 shows a mix among different groups, suggesting jedi is succeeding in facilitating the transfer of knowledge among its members. associations with number of replies on a post starting a post on jedi is often a way to seek out information from one’s peers, so we were interested in exploring whether there are any factors related to the thread starter’s characteristics that are associated with the number of replies received by a given post. as these are exploratory, post hoc analyses, the results of any statistical tests should be taken with a grain of salt. visual inspection of the means and bootstrapped confidence intervals displayed in figure 6 reveals very little difference in the average number of replies among subgroups, which is confirmed by one-way anovas (all ps > .05). that is, given our current data, we cannot detect any differences in the number of replies received by a given poster based on characteristics of the thread starter. https://doi.org/10.29173/iq1146 16/25 silverstein, p, bottesini, j, karcher, s, & elman, c (2025) introducing the journal editors discussion interface, iassist quarterly 49(2), pp. 1-25. doi: https://doi.org/10.29173/iq1146 figure 6. thread length plotted according to the characteristics of the thread starter. average thread length (gray dots) with bootstrapped confidence intervals show no relevant differences among groups. community survey respondents to this survey tended to be experienced editors, and more active than the average jedi member. in terms of their current roles, editors-in-chief, co-editors, and associate editors made up over 86% of respondents. participants also reported experience in multiple editorial roles — over 70% had held 2 or more roles — and many years of editorial experience — almost half of participants had held editorial roles for 5 years or more (a detailed table with the responses can be found at table s1). additionally, in the community survey sample, 57.5% of the respondents reported never having posted to jedi (vs. 68.3% of all jedi users, the number obtained directly from the data), suggesting that this sample of jedi members may be biased towards more active members. responses to the community survey also indicate that this sample of jedi members holds very positive views of jedi and its usefulness. jedi’s informativeness, relevance to their field, and helpfulness was rated highly by participants: across 5 questions, the median response was four on a five-point likerttype scale where 1 = not [informative/relevant/helpful] and 5 = very [informative/relevant/helpful]. https://doi.org/10.29173/iq1146 17/25 silverstein, p, bottesini, j, karcher, s, & elman, c (2025) introducing the journal editors discussion interface, iassist quarterly 49(2), pp. 1-25. doi: https://doi.org/10.29173/iq1146 when indicating where they go for help with editorial questions from a list of nine presented options, the jedi listserv ranked second in number of mentions, with 62, behind only “emailing colleagues” (85 mentions). among the spontaneous, write-in answers, there were around 10 mentions of asking others with relevant knowledge (e.g., “ask other journal editors”, “editorial advisors”, “talk to the publisher”, “prior editors at my journal”). additionally, being able to get honest opinions from people about editorial practices was rated highest in importance (med = 5 on a 5-point scale) among six aspects of jedi. when taken together, these answers suggest jedi members prefer to obtain information on editorial practices directly from exchanges with knowledgeable others. detailed statistics and the distribution of responses to all questions can be found at s2.6. in terms of conversation topics on jedi, “the future of publishing” emerged as the most popular of topics presented to respondents. other topics like finding reviewers, publication decisions, data and code sharing, open peer review, and fraud and fabrication, were also quite popular. interestingly, several participants indicated that the topics “preregistration and registered reports” and “data and code sharing” did not apply to them. other topics mentioned in an open-ended question included open access, diversity, equity and inclusion (dei) issues, preprints, and the relationship between journals and publishers. unsurprisingly, similar themes emerged when participants were asked about specific challenges at their journals they would like to overcome: challenges related to increasing the diversity of reviewers, authors, and editorial board members; implementing data sharing policies and other transparency policies at their journals; and the implementation of open access publishing and its ramifications. detailed statistics on respondents’ interest in different topics can be found at table s3. discussion an online community of editors and representatives of other institutional stakeholders in the research cycle has now been established through jedi. the discussion group offers members the opportunity for continuous dialogue. overall, jedi is meeting its objectives of promoting conversations among journal editors in the social sciences, with a particular focus on open science. in its first two years, jedi membership has grown steadily, and the frequency of posts has stayed stable across both years. topics discussed reflect the breadth of responsibilities of editors. they range from technical or workflow issues – how to find reviewers? how to facilitate data sharing? – to more fundamental questions such as the value of positionality statements or the feasibility of different publication models for journals. this wide range of topics also underlines the case for a community-based approach: no individual point of contact could provide knowledgeable answers (let alone multiple perspectives) on these topics. the majority of jedi members are editors and editorial staff of social science journals. to support editors with questions concerning data and research transparency, personnel from the data services and open science communities – for example, representatives from digital repositories that safely store, publish, and preserve digital social science data – also form part of jedi (scholarly knowledge builders). the current balance between editors and scholarly knowledge builders makes it clear that the group is for editors, with scholarly knowledge builders being a minority of membership. https://doi.org/10.29173/iq1146 18/25 silverstein, p, bottesini, j, karcher, s, & elman, c (2025) introducing the journal editors discussion interface, iassist quarterly 49(2), pp. 1-25. doi: https://doi.org/10.29173/iq1146 our data suggests the conversation is still primarily initiated by jedi team members (those on the jedi steering committee or jedi staff) and jedi catalysts. however, because catalysts are also more likely to have been more active all along (even before agreeing to become catalysts), it is hard to say whether we would see a different pattern without the catalyst program. in spite of their roles, both jedi team members and catalyst frequently make unprompted and spontaneous posts, which suggests jedi is self-sustaining. there is also a sizable and steady proportion of regular members who participate in jedi's discussions. although the majority of jedi members have never posted to the listserv, we do not see this as particularly problematic. data on attrition points to inactive jedi members finding the group worthwhile to be a part of: each email includes a link to unsubscribe from the list, but only around 10% of members have ever left the group. another 3% each choose not to receive emails (but still be able to access the listserv content), or to receive a regular digest of all emails, respectively. overall, the listserv seems to be working well. we see a reasonable mix of people participating in listserv conversations and there are no detectable differences among groups in the number of replies received by thread-starters. the community survey suggests jedi members find conversations and resources on jedi generally informative and useful and see jedi primarily as a community to get honest opinions from others on editorial practices, as intended. it is reassuring that in only two years since launching, jedi was considered by those who filled out the survey to be almost the top place to go for help with editorial questions, behind only emailing colleagues. jedi membership is not as heterogeneous as would be ideal for the purpose of the group. there is low geographic diversity, with over half of members residing in the united states. we continue to seek to expand the geographical diversity of membership. in 2022, we began an outreach program to encourage membership from editors based in the global south and have since sent out invitations to those journal editors, resulting in a slightly increased proportion of new members from the global south. we plan to increase and extend our outreach program further by systematically creating and maintaining a database of journal editors in other regions and inviting them to participate. unequal representation among different disciplines is less stark. although the plurality of jedi members are from psychology, they make up less than a third of membership. the high percentage of jedi members from psychology could be because psychology has been one of the social science disciplines leading the way for some aspects of open science. it could also be because both jedi community managers have a psychology background, naturally impacting on their networks when recruiting new members. finally, the relative size of disciplines may also play a role. in the us, e.g., there are almost three times as many faculty members in psychology as in political science, the second largest discipline both in jedi and in us social science faculty (us bureau of labor statistics, 2023). as part of jedi’s outreach program for 2024, we plan to preferentially invite editors in social sciences that make up the smaller groups in jedi as well as expand our recruiting efforts to adjacent disciplines (e.g., behavioral sciences, law). we will continue to monitor and make efforts to raise jedi’s profile globally and across the other social sciences. https://doi.org/10.29173/iq1146 19/25 silverstein, p, bottesini, j, karcher, s, & elman, c (2025) introducing the journal editors discussion interface, iassist quarterly 49(2), pp. 1-25. doi: https://doi.org/10.29173/iq1146 strengths and limitations our use of different methodologies to understand how our community works gives us a more threedimensional picture of our strengths and where we still have to make improvements to better serve our community. in addition, our open survey materials and code can be adapted by others for use with their own communities. however, there are several limitations to the data we have collected and the conclusions we can draw. firstly, we are limited by the many variables that we didn’t investigate, with regards to all three types of data (membership, posting, community survey). for membership, we are limited by the information we gather upon sign-up, and what can be easily coded based upon publicly available information. this means that we do not, for example, have data on members’ actual (non-inferred) gender, career stage, or the journal impact factor of the journal they are editing for. for posting, we have not coded the content of posts, and instead only use automated strategies for gathering data from google takeout and so cannot know whether anything about the content of the posts themselves is driving the number of replies. in addition, the number of replies is our only variable indicating the “success” of a particular thread, but the number of replies could indicate many things. it could be that there’s just a very clear answer to some questions, and so there isn’t a need for more discussion. however, as jedi’s aim is to facilitate discussion, we still believe that number of replies is a somewhat useful metric. the data from the community survey are limited in several ways. firstly, although we obtained a high response rate, around 30 percent, it is likely that the sample who completed the survey is biased. for example, a higher percentage of people filling out the community survey had previously posted on the listserv compared to the percentage for overall membership. we also had some evidence of inattentive responding (e.g., missing and/or internally incompatible responses), suggesting we take the results of the community survey with a pinch of salt. as the purpose of our community survey was to get quick feedback from members on a variety of aspects related to jedi, we weren’t able to get to much depth on any of the individual topics. for example, we don’t have any information on why members would rather email colleagues about editorial questions than post on jedi. this point also relates to a bigger open question, which is that we do not have any data on why many members never post on the listserv. the answer to this would be very important for jedi strategy, as it is important to know whether members are satisfied “lurking”, or whether there is something that would make them more likely to post. challenges and future directions despite jedi’s success, there are still a number of challenges that we will need to consider moving forward. the issue of geographical diversity is an especially important one that needs to be resolved in order for jedi to meet its goals, as ensuring science is equitable and inclusive is an important facet of open science12. to this end, we will continue to adapt our global south outreach program with the aim of increasing the diversity of jedi membership. we will also make further efforts to increase field diversity of our membership and will make active efforts to invite more scholarly knowledge builders with relevant expertise to join jedi in order to help increase the helpfulness of replies. another issue is the time-limited nature of editorial positions, in that they are usually only for a few years, after which someone else will take over at the journal. although many editors edit at several https://doi.org/10.29173/iq1146 20/25 silverstein, p, bottesini, j, karcher, s, & elman, c (2025) introducing the journal editors discussion interface, iassist quarterly 49(2), pp. 1-25. doi: https://doi.org/10.29173/iq1146 journals over their career, there are many who end their positions and then are no longer editors. when we have reached out to members who have left the group about why they have left, this is the most common reason (that they are no longer an editor). we make clear in our correspondence that we welcome previous editors too, in order to preserve this institutional knowledge, but this is a big ask from someone who will no longer necessarily be benefiting themselves from the discussions and resources. a related issue is that of identity – although we have not collected data on this (it would be very interesting to do so), it is likely that most of our editor members have many other identities that they hold closer than that of “editor” (e.g. scientist, a member of their individual discipline [e.g. sociologist], academic, researcher). and even if editors do identify as editors during their term, many may no longer identify with this once their term is up. lastly, for many of the issues discussed on the listserv – particularly those related to open science – there is no consensus on best practices. this is further complicated by field and methodological considerations among such a varied group. our resource collection is growing rapidly, but the current static website is not fit for this purpose. sections are created organically based on conversations, but this is not systematic, so there is unequal coverage for different topics. one of our solutions to this has been to create “a guide for social science journal editors on easing into open science” (silverstein et al., 2024) that aims to come to some consensus through input from editors and skbs across the social sciences. in addition to this, we have secured further funding to develop resources where they are currently missing. conclusion jedi has been successful since launching in march 2021. however, there are still many improvements that can be made. we stress the importance of actively working to enhance diversity and inclusivity in any organization looking to make science more open. communities of practice (wenger 1999) have played an important role in fomenting change towards a more open science. jedi, as a community, did not emerge organically but was purposefully created: given the costs of community building and the already significant workload of most journal editors, efforts that bear the costs of building and sustaining communities can play an important role in helping them emerge and endure. the growth and sustained activity in the jedi listserv demonstrate the relevance of building structures that facilitate open communication and consultation among stakeholder groups within the scientific ecosystem. professional associations (such as the european association for science editors [ease] – https://ease.org.uk/) can play a similar role, but most editors in the social sciences fulfill their role part time and for a limited period of time and are thus unlikely to pay to join a dedicated organization. loosely organized communities such as jedi can help to fill this need. as open science practices become more established, they are also becoming more complex and their implementation more nuanced. community-based groups such as jedi that allow for communication and consultation among stakeholders are an essential component of building the capacity for working in this complex environment. acknowledgements we thank diana kapiszewski for portions of text from a grant proposal that informed some of this manuscript, and all jedi members for contributing to the jedi community. https://doi.org/10.29173/iq1146 https://ease.org.uk/ 21/25 silverstein, p, bottesini, j, karcher, s, & elman, c (2025) introducing the journal editors discussion interface, iassist quarterly 49(2), pp. 1-25. doi: https://doi.org/10.29173/iq1146 ethical statement the data collected for this paper are not research data. the project was approved and classified as exempt by the syracuse university institutional review board (irb #: 22-047). funding statement this paper is based upon work supported by the national science foundation under grants no. 2032661 and 2332061. data accessibility the anonymized data and a data dictionary can be found on the osf at https://osf.io/sh5ry/. competing interests ps and jgb are paid consultants on the grants that fund the journal editors discussion interface, ce was the original pi on both of these grants, sk is the current pi on both grants. the authors have no other competing interests to disclose. author contributions priya silverstein: conceptualization, methodology, investigation, writing original draft, writing review & editing. julia g. bottesini: methodology, formal analysis, data curation, visualization, investigation, writing original draft, writing review & editing. sebastian karcher: writing review & editing, supervision, funding acquisition. colin elman: conceptualization, writing review & editing, supervision, funding acquisition. references alperin, j. p. (2015). the public impact of latin america's approach to open access. stanford university. american anthropological association. (2012). 2012 ethics statement. http://ethics.americananthro.org/category/statement/ american sociological association (asa). (2018). asa code of ethics. http://www.asanet.org/codeethics bird, s. j. (2014). socially responsible science is more than “good science”. journal of microbiology & biology education, 15(2), 169-172. https://doi.org/10.1128/jmbe.v15i2.870 button, k. s., ioannidis, j. p., mokrysz, c., nosek, b. a., flint, j., robinson, e. s., & munafò, m. r. (2013). power failure: why small sample size undermines the reliability of neuroscience. nature reviews neuroscience, 14(5), 365-376. https://www.nature.com/articles/nrn3475 christensen, g., & miguel, e. (2018). transparency, reproducibility, and the credibility of economics research. journal of economic literature, 56(3), 920-980. https://doi.org/10.1257/jel.20171350 https://doi.org/10.29173/iq1146 https://osf.io/sh5ry/ http://ethics.americananthro.org/category/statement/ http://ethics.americananthro.org/category/statement/ http://ethics.americananthro.org/category/statement/ http://www.asanet.org/code-ethics http://www.asanet.org/code-ethics http://www.asanet.org/code-ethics https://doi.org/10.1128/jmbe.v15i2.870 https://www.nature.com/articles/nrn3475 https://doi.org/10.1257/jel.20171350 22/25 silverstein, p, bottesini, j, karcher, s, & elman, c (2025) introducing the journal editors discussion interface, iassist quarterly 49(2), pp. 1-25. doi: https://doi.org/10.29173/iq1146 christensen, g., freese, j., & miguel, e. (2019). transparent and reproducible social science research: how to do open science. university of california press. cook, b. g., lloyd, j. w., mellor, d., nosek, b. a., & therrien, w. j. (2018). promoting open science to increase the trustworthiness of evidence in special education. exceptional children, 85(1), 104-118. https://doi.org/10.1177/0014402918793138 crosas, m., gautier, j., karcher, s., kirilova, d., otalora, g., & schwartz, a. (2018, march 30). data policies of highly-ranked social science journals. https://doi.org/10.31235/osf.io/9h7ay elman, c., & lupia, a. (2016). da-rt: aspirations and anxieties. comparative politics newsletter, 26(1), 44–52. https://qdr.syr.edu/drupal_data/public/elmanlupia_dart_2016.pdf elman, c., kapiszewski, d., & lupia, a. (2018). transparent social inquiry: implications for political science. annual review of political science, 21(1), 29-47. https://doi.org/10.1146/annurevpolisci-091515-025429 european commission. (2023, february 10). open science. https://research-andinnovation.ec.europa.eu/strategy/strategy-2020-2024/our-digital-future/open-science_en european parliament. (2019, april 17). european research priorities for 2021-2027 agreed with member states [press release]. https://www.europarl.europa.eu/news/en/pressroom/20190311ipr31038/european-research-priorities-for-2021-2027-agreed-withmember-states evans, t. r., pownall, m., collins, e., henderson, e. l., pickering, j. s., o’mahony, a., ... & dumbalska, t. (2022). a network of change: united action on research integrity. bmc research notes, 15(1), 141. farran, e. k., silverstein, p., ameen, a. a., misheva, i., & gilmore, c. (2020, december 15). open research: examples of good practice, and resources across disciplines. https://doi.org/10.31219/osf.io/3r8hb fleming, j. i., wilson, s. e., hart, s. a., therrien, w. j., & cook, b. g. (2021). open accessibility in education research: enhancing the credibility, equity, impact, and efficiency of research. educational psychologist, 56(2), 110-121. https://doi.org/10.1080/00461520.2021.1897593 freese, j. (2007). replication standards for quantitative social science: why not sociology?. sociological methods & research, 36(2), 153-172. https://doi.org/10.1177/0049124107306659 freese, j., & king, m. m. (2018). institutionalizing transparency. socius, 4. https://doi.org/10.1177/2378023117739216 gehlbach, h., & robinson, c. d. (2018). mitigating illusory results through preregistration in education. journal of research on educational effectiveness, 11(2), 296-315. https://doi.org/10.1080/19345747.2017.1387950 https://doi.org/10.29173/iq1146 https://doi.org/10.1177/0014402918793138 https://doi.org/10.31235/osf.io/9h7ay https://qdr.syr.edu/drupal_data/public/elmanlupia_dart_2016.pdf https://doi.org/10.1146/annurev-polisci-091515-025429 https://doi.org/10.1146/annurev-polisci-091515-025429 https://research-and-innovation.ec.europa.eu/strategy/strategy-2020-2024/our-digital-future/open-science_en https://research-and-innovation.ec.europa.eu/strategy/strategy-2020-2024/our-digital-future/open-science_en https://www.europarl.europa.eu/news/en/press-room/20190311ipr31038/european-research-priorities-for-2021-2027-agreed-with-member-states https://www.europarl.europa.eu/news/en/press-room/20190311ipr31038/european-research-priorities-for-2021-2027-agreed-with-member-states https://www.europarl.europa.eu/news/en/press-room/20190311ipr31038/european-research-priorities-for-2021-2027-agreed-with-member-states https://www.europarl.europa.eu/news/en/press-room/20190311ipr31038/european-research-priorities-for-2021-2027-agreed-with-member-states https://doi.org/10.31219/osf.io/3r8hb https://doi.org/10.1080/00461520.2021.1897593 https://doi.org/10.1177/0049124107306659 https://doi.org/10.1177/2378023117739216 https://doi.org/10.1080/19345747.2017.1387950 23/25 silverstein, p, bottesini, j, karcher, s, & elman, c (2025) introducing the journal editors discussion interface, iassist quarterly 49(2), pp. 1-25. doi: https://doi.org/10.29173/iq1146 harris, j. k., johnson, k. j., carothers, b. j., combs, t. b., luke, d. a., & wang, x. (2018). use of reproducible research practices in public health: a survey of public health analysts. plos one, 13(9), e0202447. https://doi.org/10.1371/journal.pone.0202447 harvard university privacy tools project. (2020). [project homepage]. https://privacytools.seas.harvard.edu/home jarvis, s. n., ebersole, c. r., nguyen, c. q., zhu, m., & kray, l. j. (2022). stepping up to the mic: gender gaps in participation in live question-and-answer sessions at academic conferences. psychological science, 33(11), 1882-1893. https://doi.org/10.1177/09567976221094036 kamath, g., & ullman, j. (2020). a primer on private statistics. arxiv preprint arxiv:2005.00010. https://doi.org/10.48550/arxiv.2005.00010 levenstein, m. c., & lyle, j. a. (2018). data: sharing is caring. advances in methods and practices in psychological science, 1(1), 95-103. levenstein, m. c., tyler, a. r., & davidson bleckman, j. (2018). the researcher passport: improving data access and confidentiality protection. https://doi.org/10.18235/0002027 lupia, a., & elman, c. (2014). openness in political science: data access and research transparency: introduction. ps: political science & politics, 47(1), 19-42. https://doi.org/10.1017/s1049096513001716 maienschein, j., parker, j. n., laubichler, m., & hackett, e. j. (2019). data management and data sharing in science and technology studies. science, technology, & human values, 44(1), 143160. https://doi.org/10.1177/0162243918798906 makel, m. c., & plucker, j. a. (2014). facts are more important than novelty: replication in the education sciences. educational researcher, 43(6), 304-316. https://doi.org/10.3102/0013189x14545513 mcbee, m. t., makel, m. c., peters, s. j., & matthews, m. s. (2018). a call for open science in giftedness research. gifted child quarterly, 62(4), 374-388. https://doi.org/10.1177/0016986218784178 mellor, d. (2021). improving norms in research culture to incentivize transparency and rigor. educational psychologist, 56(2), 122-131. https://doi.org/10.1080/00461520.2021.1902329 merton, r. k. (1949). social theory and social structure: toward the codification of theory and research. the free press. miguel, e., camerer, c., casey, k., cohen, j., esterling, k. m., gerber, a., ... & van der laan, m. (2014). promoting transparency in social science research. science, 343(6166), 30-31. https://doi.org/10.1126/science.1245317 national academies of sciences, policy, global affairs, board on research data, information, division on engineering, ... & replicability in science. (2019). reproducibility and replicability in science. national academies press. https://doi.org/10.17226/25303 https://doi.org/10.29173/iq1146 https://doi.org/10.1371/journal.pone.0202447 https://privacytools.seas.harvard.edu/home https://privacytools.seas.harvard.edu/home https://privacytools.seas.harvard.edu/home https://doi.org/10.1177/09567976221094036 https://doi.org/10.48550/arxiv.2005.00010 https://doi.org/10.18235/0002027 https://doi.org/10.1017/s1049096513001716 https://doi.org/10.1177/0162243918798906 https://doi.org/10.3102/0013189x14545513 https://doi.org/10.1177/0016986218784178 https://doi.org/10.1080/00461520.2021.1902329 https://doi.org/10.1126/science.1245317 https://doi.org/10.17226/25303 24/25 silverstein, p, bottesini, j, karcher, s, & elman, c (2025) introducing the journal editors discussion interface, iassist quarterly 49(2), pp. 1-25. doi: https://doi.org/10.29173/iq1146 national science foundation. (2011). dissemination and sharing of research results. http://www.nsf.gov/bfa/dias/policy/dmp.jsp national science foundation. (2019). dear colleague letter: effective practices for data. https://www.nsf.gov/pubs/2019/nsf19069/nsf19069.jsp national science foundation. (2023). nsf public access plan 2.0 (nsf publication no. 23–104). https://www.nsf.gov/pubs/2023/nsf23104/nsf23104.pdf nosek, b. a. (2019, june 11). strategy for culture change. center for open science blog. https://www.cos.io/blog/strategy-for-culture-change nosek, b. a., alter, g., banks, g. c., borsboom, d., bowman, s. d., breckler, s. j., ... & yarkoni, t. (2015). promoting an open research culture. science, 348(6242), 1422-1425. https://doi.org/10.1126/science.aab2374 nosek, b. a., hardwicke, t. e., moshontz, h., allard, a., corker, k. s., dreber, a., ... & vazire, s. (2022). replicability, robustness, and reproducibility in psychological science. annual review of psychology, 73(1), 719-748. https://www.annualreviews.org/content/journals/10.1146/annurev-psych-020821-114157 naaman, k., grant, s., kianersi, s., supplee, l., henschel, b., & mayo-wilson, e. (2023). exploring enablers and barriers to implementing the transparency and openness promotion guidelines: a theory-based survey of journal editors. royal society open science, 10(2), 221093. https://doi.org/10.1098/rsos.221093 parsons, s., azevedo, f., elsherif, m. m., guay, s., shahim, o. n., govaart, g. h., ... & aczel, b. (2022). a community-sourced glossary of open scholarship terms. nature human behaviour, 6(3), 312-318. https://www.nature.com/articles/s41562-021-01269-4 peng, r. d., & hicks, s. c. (2021). reproducible research: a retrospective. annual review of public health, 42(1), 79-93. https://doi.org/10.1146/annurev-publhealth-012420-105110 silverstein, p., elman, c., montoya, a., mcgillivray, b., pennington, c. r., harrison, c. h., ... & syed, m. (2024). a guide for social science journal editors on easing into open science. research integrity and peer review, 9(1), 2. https://osf.io/preprints/osf/5dar8 top advisory board. (n.d.). top factor. top factor. https://www.topfactor.org/ accessed on march 1, 2025. us bureau of labor statistics. (2024). occupational employment and wage statistics. https://www.bls.gov/oes/ wenger, e. (1999). communities of practice: learning, meaning, and identity. cambridge university press. wood, a., altman, m., nissim, k., & vadhan, s. (2020). designing access with differential privacy. in s. cole, i. dhaliwal, a. sautmann, and l. vilhuber (eds.), handbook on using administrative data for research and evidence-based policy. https://admindatahandbook.mit.edu/book/v1.0/diffpriv.html https://doi.org/10.29173/iq1146 http://www.nsf.gov/bfa/dias/policy/dmp.jsp http://www.nsf.gov/bfa/dias/policy/dmp.jsp http://www.nsf.gov/bfa/dias/policy/dmp.jsp https://www.nsf.gov/pubs/2019/nsf19069/nsf19069.jsp https://www.nsf.gov/pubs/2019/nsf19069/nsf19069.jsp https://www.nsf.gov/pubs/2019/nsf19069/nsf19069.jsp https://www.nsf.gov/pubs/2023/nsf23104/nsf23104.pdf https://www.nsf.gov/pubs/2023/nsf23104/nsf23104.pdf https://www.nsf.gov/pubs/2023/nsf23104/nsf23104.pdf https://www.cos.io/blog/strategy-for-culture-change https://doi.org/10.1126/science.aab2374 https://www.annualreviews.org/content/journals/10.1146/annurev-psych-020821-114157 https://doi.org/10.1098/rsos.221093 https://www.nature.com/articles/s41562-021-01269-4 https://doi.org/10.1146/annurev-publhealth-012420-105110 https://osf.io/preprints/osf/5dar8 https://www.topfactor.org/ https://www.bls.gov/oes/ https://www.bls.gov/oes/ https://www.bls.gov/oes/ https://admindatahandbook.mit.edu/book/v1.0/diffpriv.html 25/25 silverstein, p, bottesini, j, karcher, s, & elman, c (2025) introducing the journal editors discussion interface, iassist quarterly 49(2), pp. 1-25. doi: https://doi.org/10.29173/iq1146 endnotes 1 department of psychology, ashland university; institute for globally disrupted open research and education 2 independent scholar 3 priya silverstein and julia bottesini are joint first authors. 4 maxwell school of citizenship and public affairs, syracuse university 5 maxwell school of citizenship and public affairs, syracuse university 6 for an overview of open research resources and case studies across disciplines, see the uk reproducibility network’s open research across disciplines (https://www.ukrn.org/disciplines/, adapted and extended from farran et al., 2020) 7 jedi was already a group with a few dozen members (but no posts) before its official launch in 2021, so membership at the outset was not zero. 8 while the mbox data contains identifiable information and cannot be shared, the extraction scripts are included with the data for this article. 9 one of the participants completed the survey but did not check the box consenting to participate, and their response was therefore excluded from all analyses. 10 prior to its launch in march 2021, the jedi google group had several dozen members due to initial sign-ups from data-pass workshops from 2016 to 2019. 11 we used genderize.io to guess each member’s gender (see s1.1 for full details on this process). we must note that this is very far from an ideal way of determining members’ genders, and that it is falsely dichotomous. however, we didn’t collect information about gender upon sign-up to the group, but we believed it still important to have some sense of whether or not the group was dominated by male members and/or interactions due to gender asymmetries in editing (liu et al., 2023). 12 it is important to note that we lacked comparison data regarding the actual percentages of editors in the social sciences based in different countries. https://doi.org/10.29173/iq1146 https://www.ukrn.org/disciplines/ https://doi.org/10.1038/s41562-022-01498-1 vol241 iassist quarterly winter 1999 15 authenticity as a requirement of preserving digital data and records by shelby sanett and eun park 1 abstract assuring continued authenticity is an essential but intransigent preservation consideration for digital data and records. several key issues need to be addressed: which intellectual and technical elements of data and records are essential for assuring authenticity; how should these be maintained and represented over time; and how are authentic data and records used in various systems of practice? the authors will address these questions in light of case studies and interviews being conducted with government agencies, academic institutions, and various organizations in america, canada, europe and asia by the interpares (international research on permanent authentic records in electronic systems) project. this article will also discuss initial project findings as they relate to the specific characteristics of authenticity in the preservation of digital data and records. i. introduction why is it important to know that preserved digital data and records2 are authentic? how do we define authenticity? how do we know that received digital data and records are authentic? how are we assured that the digital data and records are as authentic when we retrieve them as they were when they were first stored and preserved? these questions are large in scope. our presentation explores the notion of the significance of authenticity in the management of records and data and reports upon the work of the interpares project currently underway. the records generated by society, whether in the course of government, business or private activity, need to be maintained and preserved as a mechanism for accountability; as evidence of individual and corporate rights; and as a form of long-term memory. in the paper world, documentary forms and procedures have developed over time to ensure that records are capable of serving as evidence of activity to be so, the records must be both reliable and authentic. reliability can be defined as the trustworthiness of the content of the record, which is ascertainable through an examination of the completeness of the record and of the procedures exercising control over its creation3 . charles m. dollar states that “archival science defines authentic records as being what they purport to be — reliable records that over time have not been altered, changed or otherwise corrupted.”4 authenticity guarantees that the record is not changed or manipulated after it has been created or received or migrated over the whole continuum of records creation, maintenance and preservation5 . in the context of records as legal evidence, authenticity is an absolute concept in that it either exists or does not. there is no relative degree of authenticity, while there may be for reliability. the status of being authentic, however, can change at any moment as a result of residual effects of an action or migration that has been performed on the record over time. this is the case for digital data as well. by contrast, authentication is the process of guaranteeing the authenticity of a record6 . if authenticity is the status of being authentic, then authentication is the action or set of activities that demonstrate that something is authentic. when it is created, a record has two indispensable components: its content and the medium to which that content is affixed. with traditional paper records, the content of a record could not be separated from its medium. in the case of an electronic record, however, its content can be separated from the original medium and transferred to another medium or even to multiple other media. even maintaining the same type of medium, an electronic record can be migrated to another hardware and software environment, thus effectively breaking the bond between content and medium. due to the physical separation of the content from the media, as well as the various ways in which the integrity of the record’s content can be compromised during the migration processes, the authenticity of the record is vulnerable. to address and overcome this vulnerability, increasing emphasis is being placed in many communities on the development and implementation of authentication processes to ensure and demonstrate the authenticity of the record. authentication processes have always included both methodological and procedural techniques for assuring authenticity, although with traditional records, these techniques have tended to be more implicit than explicit, for example, through demonstration of an unbroken chain of custody for a record and through archival description7 . there has been 16 iassist quarterly spring 2000 increasing concern, therefore, about understanding (i.e., identifying and defining) the quality and processes associated with authenticity and authentication of information objects within the digital environment. ii. authenticity and the interpares project in recent years, along with the rapid growth of electronic communications and information systems, digital records and data have presented new challenges and opportunities to a variety of communities of records. for example, the legal community is concerned that digital records are legally reliable as evidence. although attorneys on either side of a case may interpret the record differently, the records themselves must somehow be demonstrated as being as authentic when we retrieve them, as they were when they were first stored and preserved. in healthcare, it is critical that digitally stored x-rays retrieved perhaps, in connection with a court case, or to evaluate a treatment decision, are identical in resolution and color when we retrieve it, as when they were stored and preserved. in computer network communication systems, it is important to establish the security of a transmission, a message, a station, or an originator, by ensuring that the sender transmits a message only to an intended receiver and that the message has not been altered in route. to the archival community, the significance of the description, identification and preservation of digital materials is increased as a result of the evidence-based approach to the management of records. these needs and concerns raise several research questions concerning the establishment of the authenticity of digital records and data: which intellectual and technical elements of data and records are essential for ensuring authenticity in different communities of practice? can these the requirements for ensuring the authenticity of data and records be applicable across jurisdictional and technological boundaries? how should authentic data and records be maintained over time? by identifying the requirements for ensuring authenticity, the interpares project also hopes to answer these questions: what is a record and what is data? the overall focus of the project is the long-term preservation of vital organizational records and critical research data created or maintained in electronic systems and which must be preserved permanently for administrative, legal or cultural reasons. the interpares project is a collaborative effort among fourteen countries to develop strategies, policies and standards of authenticity and preservation of electronic records within archives. research is divided into four interrelated investigative domains: (1) the conceptual requirements for preserving authentic electronic records; (2) appraisal criteria and methodology for authentic electronic records; (3) methodologies for preserving authentic electronic records; and (4) development of policies, strategies and standards to ensure preservation of the authenticity of those records. the goal of the first research domain, which is concerned with authenticity, is to identify the elements of electronic records which are necessary to maintain the authenticity of those records over time through an analysis of the elements of physical and intellectual form which may affect the authenticity and nature of an electronic record. task forces in each domain are using methodologies including diplomatic analysis, structured interviews, literature reviews, systems analysis and design, and activity and entity modeling. the four task forces each focus on authenticity, appraisal, preservation and policy development. iii. preservation and the interpares project the importance of determining and analyzing the preservation function, institutional needs and long-term expectations of use and accessibility of electronic records underlies the research questions driving the interpares preservation task force. the first goal of the preservation task force is to identify and develop the procedures and resources required for the implementation of the conceptual requirements and criteria identified in the first two research domains. broadly put, responses to the research questions will incorporate an examination of the present state of longterm preservation either in use, or in development; articulate an understanding of procedural and technical methods of authentication for preserved electronic records; provide data about the principles and criteria for media and storage management required for preservation of authentic electronic records; and lastly, enable the development of a statement of responsibilities for the long-term preservation of authentic electronic records. the second goal of the preservation task force is to model the preservation function and implementation, which will be based on information gathered from responses to the research questions. the institutional investigators working at the various national archival institutions will test models. through an iterative process, results will be brought back to the international team and will be used to further refine the models, which will then be re-tested. this process is expected to reveal certain basic principles upon which strategies, policies and standards for the preservation of authentic electronic records can be drafted. iv. method and findings to date the project uses case studies to analyze requirements for authenticity based on an analysis of features of records and their genesis, using a research methodology, which is derived from diplomatics. diplomatics is an analytical iassist quarterly spring 2000 17 method developed in europe in the seventeenth and eighteenth centuries to determine the authenticity and reliability of historical documents. in the process of its introduction into most european countries, diplomatics grew into a very sophisticated system of ideas and methods about the nature of records, their creation and their relationships with the actions and persons connected to them and with their organizational, social, and legal context8 . the concepts and principles of contemporary diplomatics have been applied in ongoing electronic recordkeeping research9 including interpares and have proven effective in identifying technical and procedural requirements for ensuring the reliability and authenticity of electronic records10. the interpares project has developed a typology of the conceptual requirements for authenticity for different types of electronic records: a case study interview protocol (csip) and the template element data gathering instrument (tedgi). these are the tools used to gather the empirical data, and perform diplomatic analysis of electronic records and systems to create the electronic record typology. the csip is the primary instrument to gather the empirical data. the csip will then provide the data that researchers will need to populate the template for analysis elements for each case study. these protocol instruments have been devised by the authenticity task force to ensure that interviews carried out by the interpares case studies are conducted under comparable conditions at each institution. currently, case studies are being conducted with a variety of institutions in canada, the united states, europe (italy, united kingdom, ireland, sweden, france, and the netherlands), australia, and asia (china and hong kong) as well as a global industry group that includes censa (the collaborative electronic notebook systems association). additional information may also come from supporting documentation provided by the interviewee, additional comments made by the interviewee, external documentation from or about the case study system or organization or other identifiable sources. to date, twelve case studies for round 1 and nine case studies for round 2 have been completed or are underway. the analysis of case studies focuses on the specific characteristics and function of ensuring authenticity in the preservation of digital data and records. among the case studies, there are the multiple case studies that have the similar function and purposes with different situational contexts. for example, there are six registration systems being conducted in six different institutions in five different countries. there are five student records systems in five universities in three countries. the case studies with the same function are examined to identify whether the requirements for ensuring authenticity are applicable across juridical, technological, functional, and cultural contexts. iv. implications for further research the results of the interpares project will be used as the basis for developing further research on electronic recordkeeping systems. a methodological typology derived from a variety of case studies in real-life settings will be the basis for developing further data collection instruments and refining data analysis methods, which can then be applicable across electronic record-keeping systems. an in-depth analysis of different communities of practice would yield more insight into the ways that authentic data and records can be understood, used and managed and how common requirements of ensuring authenticity can be shared across jurisdictional and technological boundaries. as a result of the interpares project findings, it is hoped that standards establishing authenticity of electronic records will be developed that will be applicable across many communities of practice now and in the future. findings from each investigative domain are expected by december 2001. however, given the depth of the problem domains and the ongoing iterative process of designing, testing and analyzing the various requirements and methodologies, it is anticipated that research will continue beyond this date. stay tuned; the results should be quite exciting. acknowledgments the authors gratefully acknowledge dr. anne j. gillilandswetland and dr. michéle v. cloonan for their comments and encouragement while preparing this article. they also acknowledge the funding support of interpares by the united states national historical publications and records commission, the social sciences and humanities research council of canada, the national archives and records administration of the united states, and the italian national research council. references 1. the authors are ph.d. students in the department of information studies at the university of california, los angeles and participants in the interpares project. 2 records are recorded information in any form, including data in computer systems created or received and maintained by an organization or person in the transaction of business or the conduct of affairs and kept as evidence of such activity. digital records are created or received and maintained in digital form by individuals or agencies in the course of conducting business. 3 duranti, luciana. “reliability and authenticity: the concepts and their implications.” archivaria, 39 (spring 1995): 5-10. 4 dollar, charles m. authentic electronic records: strategies for long-term access. cohasset associates, inc. chicago, illinois: 1999, p. 54. 18 iassist quarterly spring 2000 5 duranti, 1995. 6 bearman, david and jennifer trant. “authenticity of digital resources: towards a statement of requirements in the research process.” d-lib magazine (june 1998). available from http://www.dlib.org/dlib/june98/ 06bearman.html, august 1, 2000. 7 gilliland-swetland, anne j. enduring paradigms, new opportunities: the value of the archival perspective in the digital environment (washington, d.c.: council on library and information resources, 2000). 8 duranti, luciana and eastwood, terry. “protecting electronic evidence: a progress report on a research study and its methodology”, archivi & computer, 3 (1995): 213-250. 9 these are: the university of british columbia’s the preservation of the integrity of electronic records project. available http://www.slais.ubc.ca/users/duranti, august 1, 2000; the interpares project international research on preservation of authentic records in electronic systems. available http://www.interpares.org, august 1, 2000; and the u.s. department of defense’s records management task force project. available http://jitc-emh.army.mil/ recmgt, august 1, 2000. 10 duranti, luciana, macneil, heather, and underwood, william. “protecting electronic evidence: a second progress report on a research study and its methodology.” archivi & computer. 6 (1996): 37-70; gilliland-swetland, anne j. and eppard, philip b. “preserving the authenticity of contingent digital objects: the interpares project. d-lib magazine (july/august 2000). available from http://www.dlib.org/dlib/july00/ eppard/07eppard.html, august 3, 2000. * paper presented at iassist 2000 (chicago, 7-10 june 2000). microsoft word 5_gluskerbookreview.docx 1/3 glusker, ann (2017) book review: the data librarian’s handbook, iassist quarterly 41 (1-4), pp. 1-3. doi: https://doi.org/10.29173/iq14 book review: the data librarian’s handbook robin rice and john southall, eds. london: facet publishing. 192 pp. £54.95. isbn 978-1783300471 ann glusker in the acknowledgements for this volume, editors rice and southall note that they were worried, in the course of writing, that “the field of data librarianship was changing faster than we could even fix our knowledge onto the page.” (p.ix). they also note, in their preface, the growing number of books written recently about research data support, as librarians, particularly in academia, are increasingly called to provide this service. with the speed of change and the number of competing volumes available in the area of data librarianship, is it worth obtaining and delving into this one? the answer is a resounding yes. there are still very few books which take a hands-on approach to data librarianship (another excellent example being margaret henderson’s data management: a practical guide for librarians), so this volume makes an important contribution to the field. the chapters included are as follows: data librarianship: responding to research innovation; what is different about data?; supporting data literacy; building a data collection; research data management service and policy: working across your institution; data management plans as a calling card (including 8 vignettes from varied institutions and disciplines); essentials of data repositories; dealing with sensitive data; data sharing in the disciplines; and supporting open scholarship and open science. an extensive list of references and an index are also included. the audience for this book is intended to be both practicing data librarians/research data support professionals, and teachers and students in library and information schools. as such, it’s conceived as a primer, and able to be used for instruction. it includes very convenient “key take-away points” at the end of each chapter, as well as “reflective questions” that can be used with students. the tone of the writing is straightforward, engaging and accessible, and there are practical suggestions throughout the text, including some very useful sets of bullet-pointed lists. an advantage is that the authors have experience in and address both the us and european/british contexts, which gives this work a more international scope than some others. another helpful feature is that, rather than having a list of suggested urls at the end of chapters, they are embedded as they occur in the text, and so as the reader engages with the material there are many side explorations available which enhance understanding (or illuminate a point—try the authors’ suggestion of searching twitter for “lost usb”). in fact, they suggest that readers use the community resource “open research glossary” as a companion to the book (http://www.righttoresearch.org/resources/openresearchglossary/). while the authors cover a wide range of topics, in a well-delineated framework, they also engage directly with the complexity of each topic, demonstrating first-hand insight into the research data life cycle and its challenges. an example is this one, on the challenges of assessing denominators for evaluation: 2/3 glusker, ann (2017) book review: the data librarian’s handbook, iassist quarterly 41 (1-4), pp. 1-3. doi: https://doi.org/10.29173/iq14 “if one-third of a new petabyte store is filled within a couple of months is that a good or bad rate of uptake? if three-quarters of departments have created collections in an open data repository after the first year, is that good or bad? is it more important to get all the departments using it or to get more collections from existing departments? does 30 downloads in a month mean a dataset is popular? if three principal investigators have sought advice from the library about a dmp for a research proposal in a month, is that a good rate of uptake? how many others have written plans without consulting the library or it service; does it matter? how well are referrals working; are there gaps in the referral network? how many researchers or research groups are following and updating their plans after they receive funding?” (p.82) the authors don’t, of course, answer all of these questions, but they give resources and ideas for how to approach them (for example, discussing the use of benchmarking and standards and identifying some relevant web sites), and treat the many other topics in the book with similar thoughtfulness and thoroughness. in fact, in the course of dealing with these aspects of data librarianship, the authors provide some very interesting and even unusual insights, based on their combined 30+ years as research data support professionals. an example is the ways in which the work of data archives and academic research user services are converging: “…as data archives have sought to emulate the reader services role of academic libraries, the latter have also begun to emulate the role of user support” (p.14). another is the identification of drivers of support services as being influenced by “top-down drivers” (funders) as opposed to “bottom-up drivers” (institutions); taking into account this dichotomy can help determine strategy (pp. 69-73). the most interesting to me was the section on understanding how researchers view their research (p. 122). the idea that researchers have an emotional bond to their work, and may not be psychologically ready to share their data, to which they may have personal and deep attachments after the intense work required to create it, was enlightening to me, and is rarely discussed. this (often unacknowledged) attachment can be a crucial barrier to the relationship with the data librarian/professional, and as such we need to understand it, and how to work with and honor it. one topic i would have liked to see discussed in this volume is strategies for continuing professional development, such those outlined by goben and raszewski in their (mainly us-focused) chapter “data 101: learning and keeping current in data management skills”. this is in large part because i’d be so interested in hearing rice and southall’s ideas, while i admit that the purview of their book is more about the direct practice of providing research data support than data librarianship as a professional (and ever-changing) path. nevertheless, the authors make an important contribution to those of us on that path. since training programs in data science and data librarianship are still in their early stages, most of us in this field have had to pick up training on the job (and for many, without any background in creating or using research data). it can often feel like learning a language by ear; you can get quite good at communicating, but without understanding the grammar and structure of the language, there’s only so far that you can advance. this volume provides the equivalent structure, acting as a textbook for the language of data librarianship, and filling in details which enhance our fluency. in sum, this book provides a valuable reference source both for the beginner and the more experienced practitioner, giving background and suggestions for practice that may be new to them. i highly recommend it! references 3/3 glusker, ann (2017) book review: the data librarian’s handbook, iassist quarterly 41 (1-4), pp. 1-3. doi: https://doi.org/10.29173/iq14 goben, a. & raszewski, r. (2016). “data 101: learning and keeping current in data management skills”. in federer, l., ed. the medical library association guide to data management for librarians. lanham, md: rowman & littlefield. henderson, m. (2017). data management: a practical guide for librarians. lanham, md: rowman & littlefield. ann glusker phd mph mlis glusker@uw.edu librarian/research & data coordinator national network of libraries of medicine, pacific northwest region university of washington seattle, washington, usa ofltfl sirvuibe ii f1 liekfikli settifls bliss beckman simon data archivist/assistant professor library instruction services division baruch college, cuny historical background baruch college, originally the business school of the city college of new york (ccny), is, since 1968, one of the eight senior colleges in the city university of new york (cuny). although particularly strong in the field of business, most academic fields are represented in its curriculum with the departments organized into three schools: business, liberal arts and education. the college awards business and liberal arts undergraduate degrees as well as the mba, several other master's degrees and a ph.d. in business. increased interest in data resources on campus has paralleled a new emphasis on computerization. there had always been some use of secondary data on campus. several important machine-readable data files were already available, scattered through several different departments, unorganized and with little documentation or information available outside the particular departments which possessed the files. use was limited to those few who knew the files existed and who knew how to access them. under the leadership of professor thomas v. atkins, deputy chairman for library instruction services, the library was able to effectively convince the college administration that data files should be conceived of as basically an information resource and, as such, the college library was the natural place for a data service. in the spring of 1931, a data archives service (das) was established as part of the library's information services. the library administration added not only a data library but created an educational program whose function was to inform and instruct the baruch community about the use of secondary data as an information resource. because of its instructional orientation das was made part of the library instruction services division and was intended to complement services already provided by groups on campus such as the educational computer center, the 'statistics lab, etc. membership in icpsr and the roper center were begun immediately. data from sources other than icpsr or roper was purchased on an extremely selective basis due to budget restrictions. a reference collection of manuals and data catalogs was set up for use with the growing tape collection, for the first time centralizing to some extent the documentation for both mainframe software and secondary data. the training programs were begun, on a limited basis, almost immediately. almost at once, baruch college began to lobby for a university-wide icpsr membership. there was some precedent for this since there had been a previous city university icpsr membership which had lapsed due to administrative problems. baruch' s efforts joined those of a significant number of faculty and administrators at other cuny colleges who had long been interested in seeing a return of the university-wide membership. finally, in july of 1983, these combined efforts v^ere successful and the senior colleges of the city university of new york became a federated member of icpsr, one of the largest federations -continued 26 in the consortium. to begin with the federation included only the four-year institutions. it is assumed that the two-year colleges will join at a later time if there is sufficient interest. to encourage the success of the federation, baruch's administration, both college and library, willingly accepted the college's appointment as coordinator for the new membership. funding as is often the case with academic institutions, there was little additional funding available. when the service began it operated on the proverbial "shoestring" with support from the library budget and using library personnel. some additional financial assistance came from the title iii grant awarded to dr. atkins for the development of a graduate business resource and study center. sufficient money for computer use and tape storage were allocated from the general research funds of the baruch college educational computer center. when baruch became the coordinator of the city university icpsr membership, the university chancellor's office paid the federation's membership fee. halftime services of baruch's data archivist, additional student assistance, and some money for non-personnel expenses such as documentation, magnetic tapes, supplies, software and travel funds were funded by additional support from the chancellor's office and members of the icpsr federation. since baruch contributed its facilities and the services of its already established data library, it was not required to contribute further funds. university support is limited to expenditures associated with the icpsr federation while baruch uses its own funds for purchase of non-icpsr data, special equipment and its own data services. staffing baruch's data archivist is assigned part-time to the baruch service and part-time to the cuny center which is also staffed by a part-time graduate assistant and undergraduate student assistants. the archivist, a trained librarian, set up the data library, organized the tape collection, developed the documentation collection of appropriate codebooks and manuals, and established the research consultation service. once organized, basic maintenance of the tape and reference collection have been assumed by the graduate assistant, who also provides assistance with the development of educational programs. as the service expands, it is expected that additional graduate assistance will be needed. undergraduate students assist with clerical duties, as does the secretarial staff of the library instruction division. the staffing is based on an assumption that computer and statistical consultation is available from baruch college's educational computer center or, in the case of other cuny users, at their home campuses. equipment the city university has a large central computer installation (cuny/ucc) used by all the colleges as their main facility. the hardware at the ucc includes an ibm 3081, an ibm 3033, and an amdahl 470/v6-ii. in addition to the mainframes, there are high speed printers, graphics equipment and software installations of most of the major statistical packages. almost all the senior colleges, including baruch, have supplementary equipment including mainframes, minis and microcomputer labs. cda stores copies of its tapes at the cuny ucc on a permanent basis so that they are easily available to all campuses. for convenience baruch facilities are often used for small printing jobs since the cuny ucc is located some 50 blocks to the northwest. the center irself has had only a decwriter 300 band printing terminal for maintenance of its tape collection and "development of on-line demonstrations for seminars. recently a volcker craig crt and an ibm pc xt were received. this equipment will be used for present activities of the center in addition to development of instruction in and assistance with microcomputer data analysis, an area in which the cda intends to specialize. for training seminars which include on-line demonstrations of secondary data resources, the center has had access to the baruch college graduate business study and resource center seminar room which is equipped with several decwriters for workshop participants and an electrohome projector which projects an enlarged image from a video terminal. physical environment if the data archives service at baruch, or for that matter, the cuny data service had waited for proper housing, it would not exist today. space is at a premium almost everywhere in the university, but nowhere more than at baruch college. the data library was begun in one of the faculty offices of the library instruction division which at that time housed three other faculty members in addition to the part-time data archivist. at this writing, a separate room serves as office space for the archivist, as a workroom for maintenance for the data collection, a storage room for the documentation collection, and as consultation space. there is some additional storage space for the master tape copies. presently, plans are being made to acquire additional space which will serve as a workroom for the students in maintaining the tape collection and the data archives files. ' the original office will then be freed for use as office space and for data consultations. on the desiderata list is a terminal room for users which will encourage data use in the data library and allow for on-line data consultations. dissemination of information when the baruch data archives search service became the coorindator for the cuny icpsr membership it expanded upon its own primary emphasis, the education of faculty and graduate students, to bringing about a general awareness of the potential of secondary data analysis and icpsr files in particular. a monthly annotated list of icpsr data available to all cuny faculty is mailed to the campus liasons, the libraries and specific data users. at baruch, the cuny list is supplemented with a listing of datasets available to only baruch faculty and students. a baruch data directory is in preparation which will list and index by subject all data available on campus including nonbibliographic databases accessed through computer search services. these general listings are supplemented by data bibliographies on specific topics which are prepared for seminars and then mailed on request. the semi-annual newsletter published by the baruch graduate business study and resource center and mailed to all departments has carried a section on the baruch data archives since its inception. but timely information and publicity for the entire cuny community remains a difficult problem because staff is limited and the community to be served is large, disparate and separated by sizeable distances. for baruch and cuny the most important methods of reaching out to both new and sophisticated data users has been the training program and the seminar series. each seminar focuses on the information resources for the study and teaching of a particular inter-disciplinary topic. each seminar includes some mention of bibliographic databases available on the topic as well as a brief overview of information sources in print. the great est portion of the two hours, however, is devoted to machine-readable data files, particularly those from icpsr. where appropriate, government data files or other data archives which specialize in the topic are mentioned. an on-line, interactive demonstration of the contents of an important dataset in the field concludes each seminar. topics covered have been urban problems, women, youth, and consumer behavior. plans for the future include seminars in this format on the national elections, health care, education, marketing, etc. the seminars held at baruch have been supplemented by visits to the individual college campuses with a "what is icpsr" format also including on-line demonstrations. future develooments at this point it is expected that the cuny icpsr federation will continue as an integral part of the university's information resources and baruch will maintain its own data service as well. plans for the future for both services are inextricably linked. first on the agenda is the solution to the problem of reaching the huge cuny community. in some part this will be done by renewed emphasis on successful programs, but additional services will be offered as finances and personnel permit. we intend to continue and increase training in the availability of data resources and how they are used for research and teaching. these seminars, we believe, assist in motivating faculty to enlarge their uses of data and secondary analysis both in research and in instruction. both the interdisciplinary seminars held at baruch and the ones held at the individual colleges will be increased. in addition, "hands-on" workshops actually using data will be held in the specially equipped baruch on-line classrooms. not only our training will increase, but also our services. we intend to facilitate faculty and student use of microcomputers for data analysis by purchasing or preparing subsets on diskette and supporting microcomputer statistical packages. special workshops are planned, for example, on the use of abc, icpsr's instructional statistical package. as a part of this effort, special instructional packages will be developed similar to teaching packages prepared by the library instruction division for bibliographic instruction. these are intended for use in our course-related lectures series. our informational services will expand this year with a new data directory containing tape access information, a subject index and enlarged annotations, for baruch this will tie together all campus files, while the cuny edition will list icpsr data available. this directory is a preliminary step in our long-term goal of having an on-line dataset catalog. publication of our newsletter on a regular basis and special publications on major datasets are also in our plans for future development. data archive management -continued from page 23 statistical packages used for analysis. another useful manual would be an internal document describing the operational procedures of the archive for use in training new staff. program development archives can play a role in the development of new programs, particularly in the collection and creation of specialized data and by encouraging the creation of new computing power. the author wishes to thank david nasatir, whose article "operational considerations of archives" in howard white's reader in machine-readable social data served as the basis for both the workshop presentation and this article. edina iassist quarterly fall 2011 17 iassist quarterly abstract this paper will chart the development and delivery of a web 2.0 community engagement tool and application programming interface (api) developed at the edina in partnership with the national library of scotland, as part of the jisc digitisation and e-content programme. such a tool enables members of the community, both within and without academia (particularly local history groups and genealogists), to enhance and combine data from digitised historical scottish post office directories (pods) with contemporaneous large-scale historical maps. the paper discusses the background to post office directories and the corresponding geo-referenced historic maps for scotland, the technical platforms deployed including sustainable software components, and web applications and services. it also examines issues surrounding user generated content (ugc) created by the community such as mediation, validation and crosschecking, and the use of social media amplification for community engagement and future directions. to conclude, the paper argues that the success of online crowdsourcing tools such as the one developed for this project will ultimately be measured by continual and extended use within the wider community. introduction post office directories, precursors to modern day yellow pages, offer a fine-grained spatial and temporal view on important social, economic and demographic circumstances. they emerged during the late seventeenth century to meet the demand for accurate information about trade and industry due to the expansion of commerce during this period. they were published more frequently than the census and generally had information about local facilities, institutions and associations, listings for private residents, traders, trades and professions, sometimes details of important people, and advertisements. the ways in which publishers collected data varied considerably. some obtained information by personal canvassing and combined the results with existing trade listings. other publishers simply asked people to send in their names together with a small payment if they wanted to be included in the directory. by the early nineteenth century methods of compilation were more organised. in part, this reflected the growing links between directories and the post office. many postal officials turned their hand to directory publishing as a means of both aiding their work and augmenting their income. information was collected by letter carriers, who circulated forms during their postal rounds, and also delivered the finished directory on commission. for scotland there are at least 750 post office directories spanning the period 1770 – 1912. the nls are in the process of scanning using optical character recognition (ocr) techniques and publishing this historic collection in conjunction with the non-profit internet archive. during the 6 month project period the addressinghistory ‘crowdsourcing’ tool focussed on three volumes (1784-5; 1865; 1905-6) of the edinburgh digitised pods and maps addressinghistory: a web2.0 community engagement tool and api by stuart macdonald1 18 iassist quarterly fall 2011 iassist quarterly from the same periods. however the specifications were such as to accommodate the full scottish collection as and when they become available. the web 2.0 interface and back-end storage solutions were built to be both scalable and as far as was practicable, self-standing so that multiple independent instances can be supported and customised for different audiences. the edinburgh directories themselves are a unique and reliable collection of street, commercial, trades, law, court, parliamentary and postal information relating to the city of edinburgh. they also provide a wealth of detailed information regarding residential names, occupations and addresses and include maps of both edinburgh and leith indicating trade and residential origins and development. one significant deficiency of this collection, which the addressinghistory online tool aims to redress by ‘crowd sourcing’, is that the addresses are not geo-referenced. geo-referencing make possible explicit spatial search and discovery, whilst permitting a map based metaphor to be used in the exploration and visualisation of the resource. e.g. the historic distribution of shipwrights in edinburgh can be plotted on a base map or the map itself can be used to explore the spatial distribution of selected phenomena (and their variation over time). similarly, personalised maps illustrating family histories, maps tracking changes in local communities, and maps linking to other digitised materials such as census records and geo-referenced images, and historical addresses could all be explored through use of the application programming interface (api). the national library of scotland’s map library is one of the ten largest in the world; as the library of the faculty of advocates from 1689, maps of edinburgh were actively collected; as a copyright library, the collections are particularly strong in the printed mapping of scotland. since 1998, nls map library has scanned over 20,000 historical maps of scotland, including over 500 of edinburgh and its environs. it is the pre-existence of large scale geo-referenced and contemporaneous maps against which the historic post office directories were contextualised that allows manual (geo)referencing down to individual house address level to be accomplished. this is achieved by simply moving a pin on the map; i.e. the map is the mechanism through which the geo-reference is allocated by the user to a particular pod entry. to assist the geo-referencing exercise, addresses from each of the directories were parsed using google’s geocoding software2 in order to assign a geo-reference. there were issues with the legibility of the ocr’d text (especially for the pod for the earlier period) in addition to period addresses no longer being in existence or having suffered name changes. thus within the interface a ranking mechanism makes explicit the relative ‘accuracy’ of the geo-coded content. the user interface to the tool and associated api is intuitive and easyto-use to encourage researchers, local historians, genealogists and members of the wider community from across the age spectrum to discover, explore and contribute to rich records of social history and to create their own related maps and data sets for both academic and personal research. these were also developed to be sympathetic to tools developed by related projects including visualising urban geographies (vug), an online resource developing new insights into the spatial character and historical development of edinburgh (http:// geo.nls.uk/urbhist/). technologies overview the addressinghistory tool and api comprises several software components, each built with resilience and sustainability in mind. open source software was chosen in several instances, allowing for great flexibility and a feature-rich application, whilst containing costs. addressinghistory is built as a typical 3-tier web application. for the user-facing client presentation component, jisc recommended standards such as xhtml, css and other relevant w3c web standards (i.e. images etc) were employed. the web interface is supported by mainstream browsers and ogc standards including the web map service (wms) interface standard and openlayers were used for web mapping components. an api is available, allowing access to the raw data via multiple output formats. it is accessible via a restful web service development the project followed best practice for technical development, making extensive use of a number of common and well-established libraries including the java sdk, spring mcv framework and the jquery javascript libraries. unit testing was performed via the junit libraries. development initially began by scoping the application’s requirements, designing a database structure to store the information contained in the post office directories in conjunction with pre-processing and data-loading software. the structural interpretation and translation of the varied content from three eras of directories proved to be a time consuming exercise. the directory data was processed, and additional metadata such as the locations of addresses were added to the database. the api, following jisc recommendations for api good practice3 was designed to allow access to the raw data using a number of http get procedure queries including a parameter which allows web developers to specify the format (json, kml or txt) they want the result returned in. the client application was built upon the api, featuring web based mapping. to the openlayers mapping, we added a collection of historical maps from nls, contemporary to the three post office directories of interest. a user registration system, facilities to edit the stored data and suggest specific changes were added towards the end of the development, together with various enhancements – including a view to the original scanned directory pages. all components of the web accessible service and api are hosted via a solaris 10 virtual container, together with an established postgresql database, hosted at edina. throughout the project, the source code, tests and configuration files were stored in a subversion version controlled repository. software builds and releases were automated via the apache maven software project management tool. documentation is stored in a shared repository. user generated content the addressinghistory project raised a number of issues regarding user generated content (ugc) created by the community such as mediation, validation and cross-checking of ugc. at present the addressinghistory team retain the option to check ugc and will do so on a periodic basis. as part of a sustainability plan it is envisaged that once community participation reaches a certain level an ‘engaged iassist quarterly fall 2011 19 iassist quarterly user group’ comprising active members of the user community may volunteer to conduct validation and cross-checking of ugc through a devolved mediation process. a logging facility has been installed in order to identify inappropriate behaviour (e.g. spam) or inaccurate ugc. a registered user can be contacted in order to justify behaviour. potentially the user or more accurately the username can be prevented from editing further content. the addressinghistory backend database has been designed so that the original database and that containing the ugc are maintained as separate instances thus allowing inaccurate or inappropriate user generated content to be removed from the database. social media a key element in determining the success of the project was the establishment of a mechanism whereby the ‘crowd’ could contribute to the creation of a fully geo-coded version of the digitised directories. in part an avenue through which such community engagement could be realised was through working with edinburgh beltane – a national co-ordinating centre for public engagement and with the university of edinburgh college of humanities and social sciences knowledge transfer office. social media channels were also deployed to engage the public, to develop links within the local and family history communities, and to act as a vehicle to expose the tool and api to a wider audience. the following section describes both method and mechanism used to engender public collaboration and community engagement building and developing community connections at the outset of the project an information page was created on the edina website4 and later updated to connect to additional addressinghistory presences. a wordpress blog5 , was deployed as the key space for communicating and engaging with interested members of our target audiences. twitter was an unexpectedly useful space for the project (via the @ addresshistory account) and a facebook page6 was also created for addressinghistory for sharing short updates, useful links and to encourage viral sharing and recommendation. community engagement from the outset the project team encouraged blogging and discussion of the project and the project officer proactively sought out potential contacts, followers, and bloggers and responded to any comments on the project to ensure that mentions were complemented with links to the website and that questions were responded to. addressinghistory benefited from pre-existing genealogy and local history blogging and online communities, receiving regular mentions and links from a wide variety of sites and discussion boards7 . the social media and web presences helped reach out to many interested parties8. however outreach activities such as events, presentations, and print communication were also instrumental in exposing the project to a wider audience. ongoing activity as a longer term strategy we intend to maintain where practicable blog activity, facebook and twitter presences. a mailing list has been set up to ensure we can remain in contact with those interested in addressinghistory developments and a google group has been established aimed at users interested in using the addressinghistory api for their own websites, projects, or mashups. future directions features and functionality addressinghistory was an ambitious project from the outset, covering a range of technologies and featuring several disparate problems. the initial processing of data extracted from the historical directories through ocr, presents a unique challenge in terms of data errors and lack of structure. in spring 20119, as part of a second phase of development made possible through further funding, the addressinghistory team will investigate and further streamline data pre-processing and loading processes with a view to providing ‘cleaner’ output. in conjunction with this, it will add further content (for other areas of scotland) to broaden the user community and subsequent utility of the tool and api, as well as incorporating computer-generated metadata to pod entries such as categorising places and professions, or extracting multiple addresses. interesting further work may also involve investigating how best to capitalise on the social mechanics of addressinghistory. this unique application offers opportunities for game-mechanics, awarding users for their crowd sourced information and challenging users to contribute. another avenue of development under consideration is the inclusion of facilities to upload and attach geo-referenced content such as images, census records, videos and sound files to addressinghistory entries in the database. in this way, the directories would be extended to include photos of people, buildings, landmarks thus enriching the resource and broadening both utility and appeal. learning and teaching the addressinghistory project team have met with representatives from glow (a national intranet for education hosted by the scottish government10 ) with a view to using addressinghistory as a means to create learning and teaching materials by glow subject specialists for school pupils both within edinburgh and beyond. materials developed may have resonance with the recently launched digimap for schools11, an online mapping service for use by teachers and pupils in schools hosted by edina linked open data “great work! but (cries) this *really* should be done with linked data and rdf, endless scope and endless data!” comment received on the addressinghistory blog the pod data has been processed and structured in such a manner that every person entry in the directories has a unique identifier in the addressinghistory xml database. each entry is accessible via the api through a uri in much the same way that a unique place in geonames (an open geographical database) is referenced by a unique uri. as the rdf output from geonames gives you the data in an xml document by using a schema of tags defined by the geonames ontology/vocabulary, addressinghistory could theoretically provide output that included some place information using said geonames tags (i.e. you get a result for a person who lives in leith and the result also links to information describing the spatial footprint, population, economy, demography of leith). or, indeed, ‘celebrities’ present within the pods could be linked to dbpedia entries using the ‘sameas’ tag, 20 iassist quarterly fall 2011 iassist quarterly which declares a link between two resources that describe the same real-world thing. context versus content “a major feature of this project is the offer of maps, and maps which enable the user to explore and present historical information spatially. the outcome is visually attractive and exciting. there is a danger that the fun of producing the map acts as a barrier to thinking about what is happening.” quote from professor robert morris at the launch event professor robert morris, emeritus professor of social and economic history at the university of edinburgh who provided the introductory presentation at the addressinghistory launch12 had reservations regarding context versus content. he indicated that, where applicable, explanatory notes providing information about background, construct and content of the original directory listings should be made explicit. in addition underlying assumptions and rules about both the structure of the processed data and the translation of the structured data into a consumable and interactive format should be made clear. this has in part been addressed by inclusion within the interface of a range of help documents (including a post office directory guide and people, place & profession search guides), an api guide and frequently asked questions. as part of any future development, use cases and contextual essays (such those available from the statistical accounts for scotland13 ) should be considered. two videos which help to demonstrate context were created for the launch and shared via vimeo . one video discussed the background to the pods14 explaining their usefulness to researchers (amateur and professional), the second video, featured on our facebook page on the launch day, explains the pod digitisation process15. sustainability in accordance with the project plan the addressinghistory project partners are committed to supporting the resource for a minimum of one year whilst it gathers community traction. during this time consideration will be made to the processes necessary for ongoing dissemination, community take-up of the deliverables and their adoption by the community. addressinghistory aim to achieve this through those social media channels established as part of the project and an on-going relationship with edinburgh beltane16 and, in turn, to appropriate organizations engaged in local and family history projects. given the broad applicability of the resource it is envisaged that a range of communities may be interested in the longer term curation and continuance of the project tools e.g. the open street map community has an active user base interested in both contemporary and historical addresses. it is also anticipated that the active involvement of ‘engaged users’ throughout the project and beyond will provide direction on longer term sustainability issues. project partners will evaluate possible business models of sustainability based on levels of demand provided they remain consistent with the underlying open philosophy e.g. revenue generation through an online donations facility, subscription model (e.g. per annum, per month, per use), a ‘freemium model’ (e.g. free api download of a certain number of records with payment being required for further downloads), or academic advertising. conclusion addressinghistory was an ambitious project which combined a range of technologies from data processing and database design, to web 2.0 and web mapping services. much was achieved within the relatively short project in terms of public engagement and amplification through social media facilities and channels, and the delivery of a robust and scalable website and api capable of empowering the ‘crowd’ with the facility to search and edit geo-referenced content from the scottish post office directories and digitised historic maps from the same era. however, gauging the success of the project goes beyond the delivery of engaging and innovative online tools. it will ultimately be measured by continual and extended use within the wider community17. notes 1. this paper was shown as a poster presentation at the iassist conference 2010 at cornell university. stuart macdonald is the addressinghistory project manager at edina & data library, university of edinburgh. email: stuart.macdonald@ed.ac.uk 2. http://code.google.com/apis/maps/documentation/javascript/v2/ services.html#geocoding 3. http://ie-repository.jisc.ac.uk/344/ 4. http://edina.ac.uk/projects/addressinghistory_summary.html 5. http://addressinghistory.blogs.edina.ac.uk/ 6. http://www.facebook.com/addressinghistory/ 7. e.g. clan maclea/livingstone (http://clanlivingstone.info/forum/ viewtopic.php?f=5&t=1084&p=9800&hilit=addressinghistor y#p9800), rootschat (http://www.rootschat.com/forum/index. php?topic=496764) 8. indeed many of the social media monitoring techniques that were trialled on addressinghistory are now successfully being used to better monitor social media mentions of other edina projects and services. 9. at time of publication (february 2012) phase 2 of the project is nearing completion. work focused on streamlining the geo-parsing, and extending geographic and temporal coverage of post office directories within the online tool. an addressinghistory augmented reality application will also be available shortly. 10. http://www.ltscotland.org.uk/usingglowandict/glow/ 11. http://digimapforschools.edina.ac.uk 12. http://tinyurl.com/375czrb 13. http://edina.ac.uk/stat-acc-scot/reading/ 14. http://vimeo.com/16902845 15. http://vimeo.com/16906333 16. the partnership cultivated between addressinghistory and edinburgh research & innovation and edinburgh beltane has initiated ongoing communications between edina and both organisations with a view to enhancing community engagement from broader service level perspectives. 17. note: a free to access index for the glasgow post office directories from 1783-1911 is now available http://bizdirs.from-mt.com/ glasgow/ mag instructions for authors of the iassist quarterly 1/12 wiltshire, deborah (2024) developing canonical ‘safe researcher’ training materials for trusted research environments, iassist quarterly 48(1), pp 1-12. doi: https://doi.org/10.29173/iq1093 the creative commons-attribution-noncommercial license 4.0 international applies to all works published by iassist quarterly. authors will retain copyright of the work and full publishing rights. developing canonical ‘safe researcher’ training materials for trusted research environments deborah wiltshire 1 abstract social science and humanities research infrastructures allow the sharing and safe use of confidential, sensitive data for research via physical safe havens. in recent years there has been a shift towards virtual data enclaves or remote desktop systems that offer fewer physical controls. these controls need to be replaced with other safeguards, including mandatory ‘safe researcher’ training. this training aims to ensure that researchers are equipped with the knowledge required to use secure data safely. developing training is resource intensive so canonical training materials are an economical approach to providing standardized, high-quality training. the social sciences and humanities open cloud project deliverable ‘training materials of workshop for secure data facility professionals ́had two objectives. the first was the development of a set of canonical training materials that trusted research environments (tres) could use as a framework on which to build their own training course. the second objective was to hold a virtual workshop where the training materials could be demonstrated to a credible audience to gather feedback to inform the future development of the materials. we have now developed the canonical materials, building on the wealth of expertise and experience of uk-based tres. these training materials were then demonstrated at a virtual, two-hour stakeholder workshop that we organized in september 2021. following our demonstration of the materials, we facilitated small group discussions to gather vital feedback. the discussion groups formed a consensus that the materials were both comprehensive and clearly structured and would be a valuable resource to the tre community. keywords safe researcher training; canonical training materials, sensitive data, trusted research environments introduction social science and humanities research infrastructures provide a variety of resources and services that researchers can use and benefit from. they provide infrastructures that allow the sharing and safe use of confidential, sensitive data for research. these infrastructures are often in the form of trusted research environments (tres) that provide safe data access and use by creating highly secure digital https://doi.org/10.29173/iq1093 https://creativecommons.org/licenses/by-nc/4.0/ 2/12 wiltshire, deborah (2024) developing canonical ‘safe researcher’ training materials for trusted research environments, iassist quarterly 48(1), pp 1-12. doi: https://doi.org/10.29173/iq1093 environments. these digital environments are augmented by non-technical controls, and tres often choose to adopt the five safes framework as the basis of their security model. the five safes framework details five principles – safe projects, safe settings, safe data, safe people, and safe outputs – which can be successfully balanced to ensure the safe use of sensitive data (desai, ritchie, and welpton, 2016; woollard et al., 2021). originally access to these confidential data was only via safe havens or safe rooms physical, secure rooms often located within a data archive or secure data center2. these rooms are specially designed to ensure that strict physical controls are in place. such controls include lockable rooms with restricted access, the barring of personal items such as electronic devices in the safe room, and the prohibiting of taking handwritten notes whilst working in the room. in recent years there has been a shift towards virtual data enclaves or remote desktop systems via tres. these systems offer more flexibility as researchers can access the secure digital environments from their own institutional offices via a secure vpn connection, so they are highly popular. but they also introduce the potential for greater risk as they offer fewer physical controls. the controls lost need to be replaced with other safeguards in line with the five safes framework. these safeguards include legal measures in the form of licenses, data use agreements, and contracts – mandatory safe researcher training for researchers prior to data access being granted is often included as well. it is widely recognized within the tre community that the researcher is an integral part of any security model and should contribute to mitigating the risks associated with disclosive and sensitive data (lambert 1993; wiltshire 2022; bishop et al, 2022). therefore, the function of safe researcher training is to ensure that researchers have the knowledge and the appropriate attitude to avoid mistakes or poor practices that might otherwise lead to a data confidentiality breach (desai, ritchie and welpton, 2016). many tres in the uk have mandated safe researcher training as part of their security models for many years, and whilst some in the wider european tre community also offer training, it is not yet adopted as widely. this training can cover a range of different topics depending on the specific requirements of the tre but will typically include information on the legislative frameworks or constraints, service-specific information, and statistical disclosure. this ensures that researchers are equipped with the knowledge required to use sensitive data safely. the experience of tres across the uk and beyond shows that such training can have a significant positive impact on protecting the confidentiality of the data (bishop et al., 2022). one example is the uk safe researcher training which is a half-day training course complete with assessment which forms part of the accredited researcher scheme3 overseen by the office for national statistics. all researchers who wish to access sensitive data made available under the digital economy act must undergo this training. the training covers topics such as data security and the researchers’ responsibilities and statistical disclosure control, as well as service-specific information to help researchers use the service more effectively and efficiently. the training can be developed by any tre who make these data available, and this consortium-based approach enables researchers to train with one service but carry their ‘trained’ status to another service within the group. as tre community increasingly switches to remote access, rather than on-site access, discussions have turned to introducing training into our security models. with the additional driver of increasing remote https://doi.org/10.29173/iq1093 3/12 wiltshire, deborah (2024) developing canonical ‘safe researcher’ training materials for trusted research environments, iassist quarterly 48(1), pp 1-12. doi: https://doi.org/10.29173/iq1093 access connections that allow access to sensitive data across international borders, ensuring that we have consistent and comparative security models across the different tres is becoming a key priority. having some commonalities in the training that tres across the world offer can only be a positive addition to ongoing efforts to open up sensitive data access across international borders. the desire to implement training is not without challenges. for tres outside of this uk-based group, implementing such training would involve starting from scratch and this is potentially a prohibitive barrier for many. developing any training course is resource intensive, and for smaller tres with limited resources, designing and developing a safe researcher training course from scratch is a burden not easily overcome. looking for ways to overcome this barrier to implementing safe researcher training was the focus of one of the working groups of the social sciences and humanities open cloud (sshoc) project, a large project that aimed to expand access to sensitive data in europe. the social sciences and humanities open cloud project the sshoc project4 was an eu funded project that brought together 47 organisations from across europe and from across disciplinary boundaries to advance and to contribute to the work of the european open science cloud (eosc)5. the sshoc partners brought a breadth and depth of expertise and experience across the entire data cycle from data collection and curation to data re-use and training. the sshoc project ran between january 2019 and april 2022 and focused on transforming the heavily siloed social sciences & humanities data landscape into an integrated, cloud-based network of interconnected data infrastructures. it consisted of 9 work packages encompassing a range of different deliverables and milestones. work package 5 focused on innovations in data access and includes deliverables aimed at enhancing and extending the infrastructure for secure remote access to sensitive data. as part of this work package, myself and colleagues from tres across europe worked to deliver several key deliverables aimed at tackling some of the challenges and barriers to international access. in particular we were keen to find a way to facilitate tres, especially those with fewer resources, in implementing their own safe researcher training via the deliverable d5.20 ‘developing canonical training materials’. this deliverable had two objectives: the first was the development of a set of canonical safe researcher training materials that any tre looking to develop safe researcher training could use as a framework on which to build their own training course. the second was to hold a virtual workshop to debut the training materials in front of a credible audience of secure data access professionals, trainers, and researchers in order to gather feedback on the materials. the sshoc safe researcher canonical training materials and workshop as the lead of the secure data center team at gesis with many years’ experience of delivering safe researcher training in the uk, the task of overseeing the development of the sshoc canonical training materials fell to the author. this was opportune as while mandatory safe researcher training is not yet routinely in place at tres in germany, one of our priorities is to introduce it at the secure data center. but like many tres, we are a very small team and outside of the sshoc project, we would not have had the resources to develop our own training. when considering how to approach this task, our early decisions were heavily influenced by the https://doi.org/10.29173/iq1093 https://sshopencloud.eu/about-sshoc http://www.eosc-portal.eu/ https://www.sshopencloud.eu/partners 4/12 wiltshire, deborah (2024) developing canonical ‘safe researcher’ training materials for trusted research environments, iassist quarterly 48(1), pp 1-12. doi: https://doi.org/10.29173/iq1093 successful development and implementation of the safe researcher training (srt) scheme in the uk. under this scheme, a canonical set of training materials were developed by the team at the office for national statistics (ons) and then used by multiple uk-based tres to deliver their own training courses. i had worked with these training materials and delivered them to the research community over many years and knew that they both covered the topics required by tres and were well received by participants. therefore, as a project team we were confident that canonical training materials would be an economical approach to providing standardised, high quality safe researcher training programmes across multiple tres, and that the ons srt scheme provided a suitable model for us to adopt. safe researcher training is a very niche area, and to date there are very few practical examples to follow and not a large body of research available to us. therefore, as a starting point for developing our own canonical training materials, we approached felix ritchie, the lead author of the ons’ safe researcher training program to ask if we could use his materials as a basis for our own materials. we made the decision to base our own materials on these resources, because these are tried and tested over many years, and have evolved based on the experiences and feedback of both trainers and participants alike. whilst there is to date no formal research specifically into the efficacy of these training materials, anecdotally those who work in tres that implement such training, see a positive difference in key markers such as researcher attitudes towards data governance and the quality and safety of outputs. in july and august 2021, i developed a set of sshoc canonical set of materials consisting of a 94 slides powerpoint presentation. where appropriate, i added speaker notes that were designed to aid future trainers in delivering the key messages. following the completion of the sshoc project, we made the materials publicly and freely available via the sshoc projects’ zenodo account (wiltshire, 2021a)6. whilst the core content of the materials is very similar to the uk materials, there are a couple of key differences. firstly, the uk materials were developed by the team at the ons specifically for uk tres that make data available under the digital economy act. in contrast, with our materials we wanted to make them accessible to a more general audience so that they could be adapted to make a wide range of needs. secondly, the ons training materials are designed to be largely a finished course, with other tres adding a few additional slides to give service-specific information as the only edit. with our materials, our main aim was to produce something that tres could use as a starting point for developing their own training course. so, the materials provide information on core topics that could easily be adapted to the specific needs of tres across different countries and potentially different data types. in this sense, the materials are designed not to be a finished product but to provide a more flexible framework upon which tres can build their own training course. the structure of the training materials the training materials are arranged in six modules, each covering a distinct topic as follows: • module 1 introduction • module 2 understanding what impacts data access https://doi.org/10.29173/iq1093 5/12 wiltshire, deborah (2024) developing canonical ‘safe researcher’ training materials for trusted research environments, iassist quarterly 48(1), pp 1-12. doi: https://doi.org/10.29173/iq1093 • module 3 the role of legislation in data access • module 4 the five safes framework • module 5 – statistical disclosure control • module 6 – service-specific protocols i arranged the content into this modular format primarily to provide a clear structure and easy adaptability. i designed the module structure so that each module both builds on the previous module and provides the key concepts required to understand the next module. for example, module 2 discusses some of the factors which impact data access including the role of legislation which is discussed in more detail in module 3. having discussed these factors, the materials then move on to introducing the five safes framework as a way of thinking about and managing data access. this provides a clear, logical path through the materials that makes it easy for researchers to follow. the modular approach should also allow for easy adaptability, allowing tres to more easily identify and isolate content that they need to adapt. considering course delivery modes initially safe researcher training courses were run as in-person courses led by one or two trainers. during the recent covid-19 pandemic, training shifted to virtual delivery with great success and researchers were appreciative of not having to travel. although the pandemic is over, many tres have opted to stay with virtual delivery as this is a cheaper and more convenient option for both them and the participants. a consideration should be given to the course length, especially when thinking about virtual delivery. experience shows that virtual courses often need to be shorter to avoid ‘zoom fatigue’, therefore i developed these materials to function as a taught course that could work equally for in-person or virtual delivery. to aid tre teams in adapting the materials for virtual delivery, the slides include comments and recommendations on where changes may be appropriate. for example, suggesting content that could be moved from the presentation into a supplementary handout to reduce the overall course length (as seen in figure one). elsewhere, two different versions of a slide have been included and guidance given on which to choose for specific audiences. https://doi.org/10.29173/iq1093 6/12 wiltshire, deborah (2024) developing canonical ‘safe researcher’ training materials for trusted research environments, iassist quarterly 48(1), pp 1-12. doi: https://doi.org/10.29173/iq1093 figure 1 excerpt from the canonical safe researcher training materials another key feature of the training materials are suggestions for group exercises and where such exercises could be included. group exercises can play a key role in ensuring the participants are engaging with the materials and in helping embed the knowledge using practical exercises. these group exercises work well in an in-person setting, however, can be difficult to facilitate in a virtual setting. during the demonstration of the materials, discussed in more detail in the next section, i demonstrated how interactive whiteboards such as jamboard can be used during virtual courses to allow participants to collect their ideas and responses in a communal location (figure two)7. figure 2 excerpt from the canonical training materials showing an interactive group exercise https://doi.org/10.29173/iq1093 7/12 wiltshire, deborah (2024) developing canonical ‘safe researcher’ training materials for trusted research environments, iassist quarterly 48(1), pp 1-12. doi: https://doi.org/10.29173/iq1093 the sshoc workshop ‘developing canonical training materials’ once the development of the materials was complete, the sshoc project team organised a workshop entitled ‘developing canonical training materials’ which took place on 21 september 2021. we sent invitations to stakeholders within the sshoc and tre communities, and the sshoc team also advertised the workshop online via the project website8. as the key person involved in the development of the materials, i hosted the workshop, which due to the ongoing covid-19 pandemic, was held virtually via zoom. the workshop lasted around two hours and designed to be fully interactive with a targeted audience consisting of tre professionals, experienced trainers, and researchers. in total 22 people from across europe and the united states attended the event. the workshop was divided into two parts – during the first part of the workshop i presented the materials, discussing the motivation behind the development of the training materials and their proposed benefit to the tre community, followed a demonstration of how materials would be delivered. going through each of the six modules in turn, i talked through both the purpose of the module and the content, pointing out along the way where the content or the delivery could be adapted to suit individual tres. in the second part of the workshop, i divided the participants into small groups of 4-5 and sent them into breakout rooms for discussions focusing their thoughts about, and recommendations for the training materials. key themes from the discussions each group was presented with initial questions aimed to stimulate the discussion and steering it to the specific areas where we were particularly interested in getting feedback. we were particularly interested in gathering their thoughts around the following areas: 1. what do you think worked well? 2. do you think that the course structure is clear? 3. what do you think about the content? is it clear? is it comprehensive? 4. are there any topics that you feel are missing from the materials? what else would you need if you were delivering these materials? 5. any other comments? participants were given 30 minutes for the discussions to ensure that they had enough time to freely share their thoughts and ideas. all participants were very open to learning about the materials and this made for very lively and engaged conversations. several key themes emerged in the discussions which are summarised here. this summary includes also any recommendations made and our follow-up thoughts or actions. 1. the focus of the materials the first key theme centered around the focus of the materials. the primary focus of the materials is on quantitative analysis and therefore quantitative research which is more often the domain of the social sciences. some of the participants felt that these materials might not necessarily meet the needs of researchers in the humanities. with further discussion, the groups concluded that as both disciplinary fields adopt both quantitative and qualitative methodologies, it made more sense to think https://doi.org/10.29173/iq1093 8/12 wiltshire, deborah (2024) developing canonical ‘safe researcher’ training materials for trusted research environments, iassist quarterly 48(1), pp 1-12. doi: https://doi.org/10.29173/iq1093 of potential audiences as being either quantitative or qualitative researchers and consider how these materials will meet the needs of both. recommendations: 1. consider which groups will benefit from these materials and make this clearer in the materials. 2. adding some examples of outputs based on qualitative data i designed the materials originally with quantitative researchers in mind, and the participants agreed that the training would address the needs of quantitative researchers, as the content specifically addresses issues surrounding the sharing of numeric data sources. thus, any researcher carrying quantitative analyses, regardless of their academic discipline, could benefit from these materials. for qualitative researchers, the picture is less clear. some analysis software packages such as nvivo produce quantitative data, in which case these materials would retain some value. for those employing other qualitative analysis methodologies which do not produce quantitative data, the early modules would still be relevant, but the modules on disclosure risk will not be directly relevant. there is no immediate solution here. currently research is still ongoing on how disclosure control can be applied to qualitative research output, and a key outcome from this project is that the materials are to be further developed once this research bears fruit to include disclosure control examples for qualitative research outputs. 2. the inclusion of legalisation information the second theme focused on the role of legislation in data access which is covered in module 4. this divided opinions and prompted lively discussion among the participants with experience in the development and delivery of this kind of training. there were two schools of thought: one supporting the inclusion of information about legislation as a means of highlighting the importance of data protection and the role that tre procedures have in ensuring that researchers do not breach their legal responsibilities. the other perspective was that providing information about legislation is unnecessary as researchers should be encouraged to feel a sense of community and responsibility towards working with tre staff. this discussion mirrors previous discussions that i participated in as part of the srt expert group, set up in the uk to discuss and steer the safe researcher training courses. no firm recommendation emerged in this area within the small group discussions. within the scope of developing canonical materials that are adapted for a wide range of tres, i recognise that both perspectives are equally valid with the decision on how much information on legislation to include determined, at least in part, by the target audience. as a project, we decided to keep this content in the canonical materials to give the individual tres the option to decide how much legal information to include and whether to include this content in a presentation or in a supplementary handout. for those wishing to exclude the content on this topic, they could simply remove the complete module, without the need to carry out further edits to the rest of the materials. 3. the length of training courses based on these materials https://doi.org/10.29173/iq1093 9/12 wiltshire, deborah (2024) developing canonical ‘safe researcher’ training materials for trusted research environments, iassist quarterly 48(1), pp 1-12. doi: https://doi.org/10.29173/iq1093 the third theme was around how long courses based on these materials would take to deliver. among the participants in group 2 were several experienced ‘safe researcher training’ trainers and they along with myself were able to give guidance on estimated course lengths of between 2.5 and 5 hours depending on the mode of delivery. the conclusion of group 2 was that the length of the course is also determined in a large part by the level of prior experience of the participants, how communicative they are, and the delivery style of the trainers. recommendation: the key recommendation from the discussion in group 2 was that guidance on the length of the course should be included in the training materials. this will be added in future iterations of the materials. 4. delivery modality of the training materials another key theme to emerge in the discussion groups was how to deliver the training, i.e., which delivery modality would be the most effective method. i designed the materials to function as a taught course, that offered the possibility of developing either an in-person or virtual training. there are benefits to both delivery modes. delivering training in person allows trainers to build up a good relationship with the researchers who they will be supporting. this can be particularly useful when it is necessary to discuss potential problems with researchers. however, in-person training is time and cost intensive. for example, at the uk data service, the securelab team delivered safe researcher training with two trainers approximately every three weeks in london. the burden of traveling was also felt by researchers, who may have had to travel some distance to attend the course. virtual training offers a cheaper, less resource intensive option, but in group 1, participants felt that even virtual training might be too resource intensive for smaller tres. they felt that the modular design would allow the content to be adapted fairly easily to an online self-study format, should that be the preferred delivery mode. this may be an attractive option for services that do not have the resources to run regular taught training events, either in-person or virtually. recommendation: consider adding further adaptations or guidance to the materials to aid tres who wish to develop their training as an e-learning program. as part of the sshoc project we were not able to fully develop further iterations of the slides themselves, so this recommendation remains outstanding. however, we believe that the core content could be used and built upon to develop a self-study course material. for example, more information would need to be added to the slides, but this can easily be taken from the information i included in the delivery notes. depending on resources, tres could consider also producing short video recordings for some of the content to try to engage more active engagement with the content. 5. delivery notes for trainers the final theme to emerge was around how trainers could be supported in deliver training, especially those new to safe researcher training. some participants felt that although i included many delivery notes in the powerpoint slides, trainers new to the topic area would still need further guidance. they https://doi.org/10.29173/iq1093 10/12 wiltshire, deborah (2024) developing canonical ‘safe researcher’ training materials for trusted research environments, iassist quarterly 48(1), pp 1-12. doi: https://doi.org/10.29173/iq1093 also felt that some guidance for tres on deciding what content to include and what delivery mode to opt for would be helpful. to address these issues, the participants felt that a separate trainers guide would be a good addition to the materials. recommendation: a separate supplementary guide for trainers should be developed and included in the training materials. this guide should include not only delivery notes for trainers, but also some guidance on some of the issues discussed in this section. following the workshop, i produced a supplementary guide for trainers which starts with an overview of the purpose of the training materials, and then guidance on to tres on understanding their audience and how to decide on the delivery mode for their training courses. the remainder of the guide provides slide-by-slide guidance on the overall purpose of the slide and some ideas for key points to discuss. these notes are not intended to be a verbatim script but should guide trainers to the content and its purpose. this guide has now been added to the set of materials and made available online along with the powerpoint slides. as it stands, we have not had chance to gain feedback from tres and trainers, however, we hope to be able to assess the guides effectiveness in addressing these issues in the future, once we have some use cases for the materials. concluding comments and next steps the development of a new set of canonical safe researcher training materials is an important step in expanding sensitive data access. the materials were well received by those attending the workshop underlined the development of the materials as an important development in supporting tres to implement training for researchers applying to access sensitive data. the consensus in the small group discussions was that the comprehensive breadth of information included in the materials means that tres would be able to adapt them to their needs with relative ease. the materials are now freely accessible via the sshoc project zenodo account, and to date have been downloaded 128 times (wiltshire, 2021a). overall, it was widely agreed that these canonical training materials would be of great benefit to the international tre and research communities and could play a role in driving consistent standards of training across different organisations and countries. this is particularly important with the move towards opening up international data sharing where consistent data governance structures are vital. but this work has also highlighted the need to conduct further research into safe researcher training so that we can move beyond relying on anecdotal evidence. since the completion of the sshoc project, there have been two exciting developments. first, several tres in germany have now expressed interest in developing safe researcher training for their own services. as a result, i am utilizing and adapting the sshoc canonical training materials as part of a new collaborative project in germany, the ‘accrediting safe use of research environment (assured) project. we are currently working to develop a modular elearning training program for tres across germany, which initially will offer training pathways https://doi.org/10.29173/iq1093 11/12 wiltshire, deborah (2024) developing canonical ‘safe researcher’ training materials for trusted research environments, iassist quarterly 48(1), pp 1-12. doi: https://doi.org/10.29173/iq1093 for researchers wishing to access sensitive data via tres and for staff working in tres wishing to further develop their knowledge and improve their career development opportunities9. this will be the first use case of these materials and will offer over the coming months a chance to more formally assess their efficacy. the second development was a consensus among workshop participants that the opportunity to come together to discuss key aspects of secure data access and use, and to benefit from each other’s knowledge, was a valuable experience for those working within tres. this spirit of cooperation and engagement formed the basis for the new international secure data facility professionals network (isdanet) (lichtwardt, wiltshire and bishop, 2022). since the sshoc workshop, this new network has met biannually, bringing together tre professionals from across the globe to discuss different aspects of secure access to sensitive data. with this spirit of cooperation and collaboration, it is hoped that the endeavours to build consistent and comparative practices across the global tre community will further advance the move towards wider, more equitable access to sensitive data. references bishop, l., broeder, d., van den heuvel, h., kleiner, b., lichtwardt, b., wiltshire, d. and voronin, y. (2022). ’d5.10 white paper on remote access to sensitive data in the social sciences and humanities: 2021 and beyond (1.0)’, zenodo. available at: https://doi.org/10.5281/zenodo.6719121. desai t. and ritchie f. (2010) ’effective researcher management’, work session on statistical data confidentiality 2009, pp. 1-11. available at: https://unece.org/fileadmin/dam/stats/documents/ece/ces/ge.46/2009/wp.15.e.pdf. desai, t., ritchie, f. and welpton r. (2016) ‘five safes: designing data access for research’. economics working paper series 1601, pp. 1-27. available at: https://www2.uwe.ac.uk/faculties/bbs/documents/1601.pdf. lambert, d. (1993) ‘measures of disclosure and harm’, journal of official statistics, 9(2) 313-331. available at: https://www.researchgate.net/profile/dianelambert/publication/2337915_measures_of_disclosure_risks_and_harm/links/58486b7a08ae95e1 d1665c22/measures-of-disclosure-risks-and-harm.pdf. lichtwardt, b., wiltshire, d. & bishop. e.l (2022). ‘d5.12 international secure data facility professionals network (isdfpn’), zenodo. available at: https://doi.org/10.5281/zenodo.6583379. social science and humanities open cloud (2021) ‘workshop blog; workshop notes: developing canonical training materials for secure data access facility professionals’, sshoc, 30 september. available at: https://sshopencloud.eu/news/sshoc-workshophttps://doi.org/10.29173/iq1093 https://doi.org/10.5281/zenodo.6719121 https://unece.org/fileadmin/dam/stats/documents/ece/ces/ge.46/2009/wp.15.e.pdf https://www2.uwe.ac.uk/faculties/bbs/documents/1601.pdf https://www.researchgate.net/profile/diane-lambert/publication/2337915_measures_of_disclosure_risks_and_harm/links/58486b7a08ae95e1d1665c22/measures-of-disclosure-risks-and-harm.pdf https://www.researchgate.net/profile/diane-lambert/publication/2337915_measures_of_disclosure_risks_and_harm/links/58486b7a08ae95e1d1665c22/measures-of-disclosure-risks-and-harm.pdf https://www.researchgate.net/profile/diane-lambert/publication/2337915_measures_of_disclosure_risks_and_harm/links/58486b7a08ae95e1d1665c22/measures-of-disclosure-risks-and-harm.pdf https://doi.org/10.5281/zenodo.6583379 https://sshopencloud.eu/news/sshoc-workshop-notes-providing-canonical-training-materials-secure-data-facility-professionals 12/12 wiltshire, deborah (2024) developing canonical ‘safe researcher’ training materials for trusted research environments, iassist quarterly 48(1), pp 1-12. doi: https://doi.org/10.29173/iq1093 notes-providing-canonical-training-materials-secure-data-facility-professionals. wiltshire. d. (2021a). ‘d5.20 training materials of workshop for secure data facility professionals (v1.0)’, zenodo. available at: https://doi.org/10.5281/zenodo.5638596 wiltshire, d. (2021b). ‘sshoc workshop: providing canonical training materials for secure data facility professionals’, zenodo. available at: https://doi.org/10.5281/zenodo.5541587. wiltshire, d. and alvanides, s. (2022). ‘ensuring the ethical use of big data: lessons from secure data access’, heliyon, 8 (2), pp. 1-6. available at:https://doi.org/10.1016/j.heliyon.2022.e08981 woollard, m., lichtwardt, b., bishop, e.l. and müller, d. (2021). d5.9 framework and contract for international data use agreements on remote access to confidential data (v1.0)’. zenodo. available at: ‘https://doi.org/10.5281/zenodo.4534286 endnotes 1 deborah wiltshire, gesis-leibniz institute for the social sciences, unter sachsenhausen 6, 50667 cologne, germany, deborah.wiltshire@gesis.org; orcid 0000-0001-6533-2426 2 examples of a safe haven or safe room include the uk data service securelab (https://ukdataservice.ac.uk/help/secure-lab/what-is-securelab/) and the secure data center at gesis (https://www.gesis.org/en/services/processing-and-analyzing-data/analysis-of-sensitivedata/secure-data-center-sdc) [accessed 01/08/2023]. 3 become an accredited researcher, 2023 https://www.ons.gov.uk/aboutus/whatwedo/statistics/requestingstatistics/secureresearchservice/b ecomeanaccreditedresearcher#full-accredited-researcher-under-the-digital-economy-act-2017-dea [accessed 30/07/2023]. 4 https://sshopencloud.eu/about-sshoc [accessed 08/08/2023]. 5 https://eosc-portal.eu/ [accessed 29/07/2023]. 6 wiltshire. d. 2021. d5.20 training materials of workshop for secure data facility professionals (v1.0). zenodo. https://doi.org/10.5281/zenodo.5638596 7 jamboard is a free interactive whiteboard from google which allows participants in virtual classrooms to collect ideas and work collaboratively: https://jamboard.google.com/ [accessed 11/08/2023]. 8 see sshoc workshop: developing canonical training materials for secure data access facility professionals; https://www.sshopencloud.eu/events/sshoc-workshop-providing-canonical-trainingmaterials-secure-data-facility-professionals [accessed 11/08/2023]. 9 this project is a collaborative endeavour by the secure data center at gesis, the german human genome-phenome archive and berd @ nfdi. it does not yet have a dedicated website but a presentation outlining the project (using its original project name saras) can be found online: wiltshire, deborah. (2023, march 9). safe researcher accreditation system. zenodo. https://doi.org/10.5281/zenodo.8238606 https://doi.org/10.29173/iq1093 https://sshopencloud.eu/news/sshoc-workshop-notes-providing-canonical-training-materials-secure-data-facility-professionals https://doi.org/10.5281/zenodo.5638596 https://doi.org/10.5281/zenodo.5541587 https://doi.org/10.1016/j.heliyon.2022.e08981 https://doi.org/10.5281/zenodo.4534286 mailto:deborah.wiltshire@gesis.org https://ukdataservice.ac.uk/help/secure-lab/what-is-securelab/ https://www.gesis.org/en/services/processing-and-analyzing-data/analysis-of-sensitive-data/secure-data-center-sdc https://www.gesis.org/en/services/processing-and-analyzing-data/analysis-of-sensitive-data/secure-data-center-sdc https://www.ons.gov.uk/aboutus/whatwedo/statistics/requestingstatistics/secureresearchservice/becomeanaccreditedresearcher#full-accredited-researcher-under-the-digital-economy-act-2017-dea https://www.ons.gov.uk/aboutus/whatwedo/statistics/requestingstatistics/secureresearchservice/becomeanaccreditedresearcher#full-accredited-researcher-under-the-digital-economy-act-2017-dea https://sshopencloud.eu/about-sshoc https://eosc-portal.eu/ https://doi.org/10.5281/zenodo.5638596 https://jamboard.google.com/ https://www.sshopencloud.eu/events/sshoc-workshop-providing-canonical-training-materials-secure-data-facility-professionals https://www.sshopencloud.eu/events/sshoc-workshop-providing-canonical-training-materials-secure-data-facility-professionals https://doi.org/10.5281/zenodo.8238606 18 iassist quarterly educating the data user: the role of bibliographic instruction by kristin mcdonough' baruch college, city university of new york the program of credit courses in research methods and materials offered by the library instruction division at baruch college, cuny, has been in existence since the early 1970's. with the exponential growth of information and the advancement of technology the courses have shifted dramatically from a practical "how to use the library" approach to a more conceptually based one that integrates process and product, tools and techniques, traditional print and electronic information sources. what were pioneered as library research courses two decades ago have come to bear the titles " information research in business" and "information research in the social sciences and the humanities", testimony to the fact that the site of research is now just as often an online laboratory or a personal computer in the office or home as a library. 'presented at the international association for social science information service and technology (iassist) conference held in washington, d.c., may 26-29, 1988 of course, research involves much more than just a choice of site or form/media in which a body of disciplinary evidence or a synthesis of opinion, or string of raw numbers is stored. as research methods have evolved and changed, so has the content of these basic courses over the years. one of the most dramatic changes in the substantive content of the courses has been the increase in emphasis on access to, evaluation, and use of public data. our team of six bibliographic instruction librarians has a special commitment to alerting students to potential sources of data because of the nature of the institution in which we teach. baruch college is, arguably, the largest business school in the world. as such, it attracts students who are, or are quickly trying to become, quantitatively oriented, and offers business and social science courses that are, in the main, quantitatively based. by junior year a student can reasonably expect in one semester to be working on a demographic analysis for a market plan, an econometric projection for a finance course, and the comparison of a fictitious company's data with that of a national sample for an industrial management class. what is perhaps surprising is that emphasis more often is on the manipulation and application of numbers, using increasingly sophisticated statistical and spreadsheet software, than on identification and retrieval of the sources of these figures. by and large, students are provided with the numbers that they are expected to "crunch." our goal in the information research courses is to go a step further and tie the identification of authoritative sources of data to the secondary data analysis. we teach students how to identify and recognize potential sources of data from ever-growing core of government, institutional, corporate and private generators of data on both domestic and international levels because we are loathe to make them dependant on data derived from out-of-date textbooks and recycled classroom lectures. summer 1988 iassist quarterly 19 if out courses as now constituted succeed in the important task of convincing students of the relative availabilit\' of published data relevant to their particular needs, it is because we have made a conscious eftort to integrate this notion into the fabric of the courses. this was, unfortunately, not always so, at least not for those sections of the courses taught by a humanities-oriented, numbers-shy librarian such as myself. in fact, it is only within the past six or seven semesters that i have stopped treating statistics as a separate entity to be introduced toward the middle of the semester and confined to a fairly cursory treatment of standard sources, such as the specialized statistical indexes. there are several reasons for my former approach to statistics as a self contained unit broached halfway into a course. the first is simply that the material on statistical sources forms chapter 11 in each of the two in-house textbooks that we use in these courses: access information: research in the social sciences and the humanities, and access information: research in business . the position of this chapter, following a chapter on government documents, meant that an instructor following the chronology of the text waited to focus on statistical sources until the students had wrestled with government documents — that bibliographically unwieldy type of material daunting even to the most experienced of librarians. the reasorung behind this order of presentation seemed to be that since so many of the important statistical series are, in fact, government publications with complex corporate authorship and involved series added entries, it was best to deal with these later in the semester when the students would be more knowledgable. linking statistical sources with government documents not only reinforced the notion that statistics could be complicated to identify — as anyone who has searched under 'united states. bureau of the census' as author can anest — but, in a non-depository library like ours, difficult to actuallv locate. another reason for this artifical approach toward teaching public data sources was the use of the search strategy as a conceptual framework upon which to structure the presentation of instructional material and the completion of assigimients. a search strategy is a suggested sequence of steps to be followed in conducting research on almost any subject the order is, of course, approximate, and the object is to dispel the notion that relevant knowledge and information on a subject are acquired serendipitously rather than through a orderly process using standard bibliographic tools. using this approach, for example, librarians have students choose a topic of their choice and then introduce them, first, to the notion of background reading, sthen to the definition of terms using a thesarus, thirdly to the identification of a bibliography of previous research, fourthly to books using the catalog, fifthly to periodicals for current information, and finally, the icing on the cake, to recent statistics. one was lucky to have guided the students this far through a search strategy by the midterm point! below are selected examples of the approach adopted over the past few years in an attempt to underscore the centrality and virtual omnipresence of quantitative evidence in the sort of contemporary social science research in which students are expected to engage or with which, at minimum, they are expected to be familiar. though i continue to use a modified search strategy framework, my goal is to stress the fact that there are a number of ways to identify and access significant collections of published data. the focus of these examples is child day care, a timely and interesting topic for our largely working class students at baruch. very eariy in the semester students are taught to immerse themselves in a subject as they start their research. this initial immersion is referred to as backgroimd reading and yields a definition and condensed history of the subject an overview of the major issues involved, as well summer j 988 20 iassist quarterly as the identification of major associations and researchers who have contributed to the formation of the body knowledge in the field. a specialized encyclopedia, such as in this case the encyclopedia of social work, is often an ideal source of backgroimd reading. it features an expanded definition of the modalities of child care, with references to both individual and teams of researchers, as well as government agencies which have gathered data relating to "neighborhood care for several million families" (illus. 1). an additional point about the dme lag inherent in data collection and analyis can be made by noting how relatively dated are the references in, for example, the latest 1987 edition of an authoritative reference book, (illus. 1) in explaining the parenthetical citation form used in the encyclopedia, it is necessary to refer to the list of references appended at the end of each article. focusing on the organizations represented in the entries is an ideal way to underscore the number and variety of groups involved in data collecting, (illus. 2) profiles of these groups in the encyclopedia of asociations indicate those which have data gathering central to their mission, (illus.3) guides to the literature or research guides are a generic family of library tools that students are encouraged to use early in the semester. it seems relevant to introduce the latest edition of wasserman's statistics sources at the same time as webb's sources of information in the social sciences and friedes' literature and bibliography of the social sciences . as the illustration (illus. 4) suggests, students should be alerted to the fact that to maximize retrieval of information, fiexibility of approach is essential. important series of statistics on day care can be found by looking either under "child care arrangements" or under "children mothers working." that many of the publications identified in this guide are available on magnetic tapes as well as in paper is a point made again and again. this very question of research terminology is one best tackled right at the start of search strategy, with lc subject headings introduced as an example of a thesaurus. its fimction is to provide an authority list of terms to be used in searching the catalog for books on a topic. one of the key points is that once the researcher determines the conect heading or search term, quantitative data on that same subject can be found by employing the standard subdivision —statistics immediately following the heading, e.g. day care centers —united states —statistics. (illus. 5) another type of reference tool with which students should fairly quickly become familiar is the handbook. the statistical abstract of the united states is introduced as an example of the type of handbook that is a compilation of tables, as well as the first recourse a student has when confronted with the task of "finding statistics". but rather than emphasize only the technical features of this single volume wonder, with its tabular titles, headings and notes, contents tables and subject index, it is more effective to present this as a first step which offers a "snapshot" of the full range of statistical series available from various government agencies. in the illustration below (illus. 6), for example, the crucial part of the table is the source note which identifies a current population report by series number. that these periodic census updates are relatively easy to find and are available in machine-readable form are points dial can be made immediately and re-emphasized later in the course of reviewing the concept of series entries as one of the elements of the catalog, (illus. 7) by this very early point in the semester, then, students have been shown that a key publication such as cunent population reports can be located in a variety of ways, through references in a bibliographic guide ( statistics sources) , or summer 1988 iassist quarterly 21 those in a handbook (statistical abstract of the united states) , or through the subject or series approach to the hbrary's catalog. that there is more than one route in no way diminishes the key importance of american statistics index or the statistical reference index , which are now routinely introduced along with other periodical indexes. the power of these relatively sophisticated bibliographic tools and the level of detailed analysis they provide of statistical publications is impressed on the students. one effective way in which to start students thinking about the degree of complexity of a social issue such as day care is to have them simply scan the index volume of either asi or sri and note the various aspects of the subject on which data are generated and collected for subsequent analysis and policy implementation. this pair of bibliographic tools are no longer only viewed as access tools alone but also as a record or minor of the perspectives from which day care can be viewed: a service to working mothers, a fast growing service industry, a tax benefit to individuals and corporations, an employee benefit, a unit in the health and nutrition deliver)' system and so on. (illus. 9) one reason why it is important to familiarize students with a number of other indexes that can lead to statistical series is that the wealth of material identified by asi or sri , either online or in print, can be overwhelming and ultimately disappointing, especially to students using a non-depository library of moderate size which may not subscribe to all the publications indexed. odier indexes that are profitably introduced as adjuncts to, if not substitutes for, the above are monthly catalog of the united states (illus. 10) and the pais bulletin (illus. 11). the latter identifies quantitative studies in two ways : with the subdivision "statistics", or by means of a note in the citation indicating that the material contains graphs, tables, charts. each of these indexes generally employs subject terms identical to lc subject headings , with which, bv this time, the students feel familiar. references to public data can also be used eftectively when the class discusses the protocol of documentation, which is a concept that undergraduates often find difficult to grasp. "what kind of facts do i have to cite?" is one of the most frequently asked questions, to which for years i had been responding, "any opinion not your own, controversial ideas, facts that are not generally known." since adding to that not very helpful list "figures or data that are subject to change" i have begun to sense that at least a few of the students now comprehend. they are beginning to understand that a statement such as "albany is the capital of new york state" is both generally known and relatively stable but that a reference to the population of new york state should be docimiented since demographic figures change. in fact, as the students now realize, reference to statistics that are woefully in error is one of the surest signs that the sources on which a paper is based are either out-of-date or unreliable. then, too, even the most reliable and authoritative of sources is never entirely bias-free, and is certainly subject to misinterpretation, an observation that surfaces continually in class discussions on the importance of evaluating material. in response to an assignment to identify at least one publication or report the data in which have susequently been questioned, several students located accounts in the popular press or scholarly literatiire about surveys whose results had been either misrepresented or misinterpreted. a new york times article reponed an assertion by one researcher that the number of latchkey children in the u.s. is far greater than suspected, since the estimate of their numbers has largely been based on the self-reported responses of the parents. many working couples who are surveyed may not admit that their young children are left alone at home while they are at work. a union newspaper published by the aft contained an editorial repudiating the results of an nie report on school crime on the grounds that the summer 1988 22 iassist quarterly national survey had made virtually no distinction in the category "incidents of crime" between pranks, minor vandalism and armed assualts on teachers! as each student reported on the assignment orally to the class, it became clear that for the majority of students, this exercise really made the notion of data come alive. undoubtedly, the fact that bliss siman, the icpsr coordinator for all units of the city university of new york, is a dyanmic member of our teaching team has contributed to our determination to make awareness of potential sources and uses of survey, census, and lime-series data a vital pan of our credit courses. for several semesters she has been presenting sessions on the secondary analysis of data to all sections of the social sciences and business information research courses, using an approach that she describes elsewhere in this issue. the prime motive for the emphasis we place on the interdependence between data identification, retrieval and evaluation on the one hand and manipulation on the other, is to give the students the skills necessary to locate sources of authoritative data. there is yet another impetus behind our thrust toward familiarizing even our beginning, non-specialist students with sources of available data. we want, over the course of the undergraduate's career, to turn the student into a discerning and demanding consumer who will incorporate use of data into subsequent business and professional life. without a developed group of educated and expectant users coming out of our colleges, universities and professional schools, who will join with librarians and scholars to protest, for example, the bureau of the censusintent to make certain of their series available in electronic form only? in the future, when dollar values are put on information and access becomes a matter of economics and political will, we hope that our efforts in the classroom will have had some effect n bibliography american statistics index . washington, dc: congressional information service, 1974 -. encyclopedia of associations . gale research co., 1979 detroit, mich.: encyclopedia of social work . washington, dc: national association of social workers, 1974. friedes, thelma. literature and bibliography of the social sciences . lx)s angeles, ca.: melville pub. co., 1973. pais bulletin . new york, ny: public affairs information service, 1986 -. sources of information in the social sciences: a guide to the literature. 3rd, ed. chicago, ii: american library association, 1986. statistical reference index . bethesda, md: congressional information service, 1980 -. statistics sources . detroit mich.: gale research co., 1983. united states. bureau of the census cuneni population reports . washington, dc: u.s. government printing office, 1948 -. united states. bureau of the census statistical abstract of the united states . washington, dc: u.s. government printing office, "latest . united stales. library of congress. subject cataloging division library of congress subject headings . washington, dc: library of congress, 1986. summer 1988 iassist quarterly 23 illustration 1 encyclopedia of social work. 18th edition. association of social workers, 1987. edited by anne minahan. new york: national v older brothers and sisters, grandmothers, other kin, and householders to the extent possible, or even on the children themselves (werner, 1984). increasingly, however, families are turning to care outside the home. it is estimated that this was the case by 1980 for about half of all children under 6 (u.s. bureau of the census, 1982). smaller families and increased rates of maternal employment have gone hand in hand, and most families using out-of-home care are purchasing care for one child (emlen, 1974, 1982; hayghe, 1984). familydgy-cofe. the ldie ut^^tul4ren in a relatives home is less common than used to bv ( u.s. bureau of the census, 1982). day care is more likeiy to be nearoy wlih"a neighbor. family day care is provided by women who are not in the labor force, who have child care responsibilities of their own (usually involving larger families), and whose experience and motivations are suited to providing a child care service, typically involving three or four children—less than the limits imposed by regulation. family day care is the predominant resource used outside the home for infants and toddlers. it is also a major resource for school-age children. care in family homes affords flexibility in the ages of children accommodated and in the hours that care is provided. concern has been raised about the use of family day care in deteriorated neighborhoods, about the isolation of caregivers from social support and training, and about their inaccessibility to regulation or to i & r programs. family day care pers|st3 , hovyever, as a viable system of neighborhood care for~severahtnit hes (collins & watson, 1976; emler emlen & koren, 1984; fosburg ct werner, 1984). cenfer care. although nonprofit day care centers continue to provide a significant amount of subsidized care for lower-income 500 percent in 5 years (kinder-care, 1983) and has the largest market share of the center care business. the second-largest chain. la petite academy, has over 400 programs in 24 states, and children's world serves more than 20,000 children in 160 centers (friedman, 1985). these chains have been profitable, in part by achieving efficiencies from large numbers of children per center and minimum labor costs, as well as by marketing their discount programs to employers. treatment in day care settings. in any community, child care is recognized as occupying an important, though often neglected, position on a continuum of specialized services to families at risk of dissolution. whether for mental health or child welfare, child care is one of the least restrictive services that can be supportive of family functioning and of a child's treatment program. child care services play a part in the "reasonable effort" required as alternatives to placement in foster care or residential treatment facilities (adoption assistance and child welfare act of 1980, p.l. 96-272). employee assistance. of wider scope, however, are two kinds of services to families to help them cope with their child care responsibilities. one is the employee assistance program (eap), which began as a corporate approach to problems related to alcoholism and has been broadened to address the individualized child care needs of employees. employee assistance programs have expanded in scope as more attention has been paid to how employees manage child care, how it affects their work, and how company policies, in turn, facilitate or adversely affect the ability of employees to combine working with family responsibilities. the flexibility of policies concerning sick leave, maternity and paternity leave, flexible work hours, and absenteeism are being modified by companies. summer 1988 24 iassist quarterly niustration 2 encyclopedia of social work. 18th edition. r services. phase i results. cambridge, mass.: author. beer, e. (1957). working mothers and the day nursery. new york: whiteside. blank, h. (1985). fact sheet. washington, d.c.: children's defense fund. brookings institution. (1972). setting national priorities: the 1973 budget. washington, d.c.: author. bureau of national affairs. (1984). employers and child care: development ofa new employee benefit. washington, d.c.: author. burud, s., aschbacher, p., & mccroskey, j. (1984). employer-supported child care: investing in human resources. dover, mass.: auburn house. campbell, n. (1985). analysis ofinternal revenue service data. washington, d.c.: national women's law center. catalyst. (1983). child care information service: an option for employer-support of child care. new york: author. children's defense fund. (1982). employed parents and their children: a data book. washington, d.c.: author. city club of portland. (1985). survey ofemployeesponsored child care options. portland, oregon: author. class, n.. & english, j. (1985). "formulating valid standards for licensing." public welfare, 43 0), 31-35. clinton, l. (1985). "guess who stays home with a sick child?" working .mother, 5(10), 55-61. coelen, c, glantz, f., & calore, d. (1978). day on the wo portland, c emlen, a., et al porate fine service. p university fasciano, n. sneezles i> child care mother, 8 fosburg, s., el united sii^ no. [ohc u.s. govt friedman, d. ( care: ho\ expectatii. (10). 1-6. frieman, d. ( ance for c board. galinsky, e. ( policies." (eds.), 5boston: 1 grubb, w. n frontiers and parer lazerson americar. new yor hayghe, h. record 1 review , . hayghe, h.(l summer 1988 iassist quarterly 25 niustratiod 3 encyclopedia of associations. 22nd edition. detroit : gale, 1988. page 979 section 7 social welfare of 1^10212* children's defense fund (cdf) i22csl..n.w. phone:(202)628-8787 washington, dc 20001 marian wright edelman, pres. founded; 1973. staff: 60. budget: $4,000,000. provides systematic, longrange advocacy on behalf of the nation's children. engages in research, pubfic education, monitoring of federal agencies, litigation, legislative drafting and testimony, assistance to state and local groups, and community organizing in areas of child welfare, child health, adolescent pregnancy prevention, child care and development, family services, and child mental health. works with ixfividuals and groups to change policies and practices resulting in neglect or maltreatment of millions of children. advocates: access to existing programs and services; creation of new programs and services where necessary; enforcement of civo rights laws; program accountability; strong parent and community role in decision-making; ad^quatetortdiny tor-^^entiaj programs for children. maintains speakers' bureqjd^mpiles statisticsjjubfications: (1) cdf reports (newsletter), monthly; f2)~adolc3ecnt pregnancy prevention clearinghouse reports, bimonthly; also publishes series of books and handbooks on issues affecting children. formerfy: (1978) children's defense fund of the washington research project. convention/meeting: annual conference. •10213* children's rights group (crg) 693 mission st. phone:(415)495-7283 san francisco. ca 94105 vicki strang, deputy dir. founded: 1974. staff: 33. organization working primarily in the western and southwestern u.s. to help communities and parents utilize, upgrade, and expand available services for children. offers workshops, training seminars, and technical assistance to parents who seek to bring federally funded child nutrition programs into their community. makes a special effort to aid organizatk)ns that work with migrant farmworkers' families. sponsors project save, an energy conservation/youth employment program providing free home weatherization for tow-income househokjs in daly city and san francisco, ca. conducts analyses of issues and legislation that affect chikjren's services, particularly tax fimitation proposals such as califomia's proposition 13. lobbies for fair housing for children ordinances (making it illegal for landtords to refuse rental to families with children) in california. operates an employerrelated chiw care project to promote child care services for employees of major bay area emptoyers. compiles data on federal food program participatton. focuses research on children's services including health, nutrition, and chdd care. publications: community services bulletin, monthly; also pubfishes •102171 p.o. box washingt founded ports sta informati^ children's •10218^ 3955 crc p.o. box : colorado founded offers m vides for nancial si children's special c; sorship (1 children t (assists < and otha (assists n [ications: chures. f passion. ( •10219^ c/o coun 1920 ass reston, v founded: supportivi impaired ducts ser environmi scholarshi formerly: children; hospitafizi for excep summer 1988 26 iassist quarterly niustration 4 statistics sources. 10th edition. edited by paul wasserman. detroit: gale, 1986. statistics soi rces, eleventh edilion 19 chickens see poultry child abuse american humane association, 9725 east hampden, denv. colorado 80231; annual report, "national analysis of offici child neglect and abuse reportingchild care arrangements us. department of commerce, bureau of the census, suitlan maryland 20233; "current population reports." child support payments and alimony u.s. department of commerce, bu maryland 20233; "current populat 1 reports " children see also population and vital statistics children aid soclal welfare programs u.s. department of health and human services, social security administration, 6401 security boulevard, baltimore, maryland 21235; monthly report, "social security bulletin,"annual statistical supplement," monthly report, "public assistance statistics," and unpublished data. u.s. library of congress, 10 first street, se, washington, d c. 20540; "cash and non-cash benefits for persons with limited income: eligibility rules, recipient and expenditure data," september 1985. children aliens u.s. department of justice, immigration and naturaliiation service, 425 i street, nw, washington, dc 20536; "statistical yearbook," annual, and releases. children attending school u s department of commerce, bureau of the census, suitland, maryland 20233; "current population reports," and unpublished data. us department of education, 400 maryland avenue. sw. washington, dc. 20202; -biennial survey of education in the united states," chapter on statistical summary of education. annual reports, "digest of education statistics," and "statistics of public elementary and secondary schools system," "projections of education statistics," "estimate! of school statistics,-rankings of the states," and unpublished data children days lost from school nd human seru.s. departn service, 20( d c. 20201;" nt of health i independence aven ital and health statisti< vices. public health sw, washington, d unpublished data children fa.milies with children immunized against disease u s department of health and human services, center for disease control, 1600 clifton road, ne, atlanta, georgia 30333, annual report, "united states immunization sun-ey " children juvenile delinquency us department of justice, bureau of prisons, 320 first street, nw, washington, d c 20534, "statistical report. u.s. department of justice, law enforcement assistance administration. 633 indiana avenue. nw. washington. dc. 20531. -children in custody advance report on the 1982 census of public juvenile facilities,and "children in custody: advance report on the 1982 census of private juvenile facilities" children mothers working ' 4~~ us department of commerce, bureau of the census, suitland. maryland 20233, "current population reports," and unpublished u s department of labor, bureau of labor statistics, 200 constitution avenue, nw. washington, d c 20212. "special labor force reports," and unpublished data. children number per divorce decree us. department of health and human services, public health service, 200 independence avenue, sw, washington, d c. 20201. annual report, "vital statistics of the united states,"monthly vital statistics reports," and unpublished u s department of health and human se security administration, 6401 security bouleva maryland 21235, -annual statistical supplement security bulletin," and unpublished data. children orphans us. department of health and human se security administration, 6401 security bouleva maryland 20235, unpublished data us. department of agriculture, foo. fourteenth street and independence , dc. 20250, annual report, "agri unpublished data nd consumer services. :nue, sw, washington, tural statistics." and summer 1988 iassist quarterly 27 illustration 5 cuny librar*-. card catalogue. eef ha 203 a218 no. 14s day cake centefis united states statistics. brunof fiosalind b. after—school care of bchool— age children : decenber 1984 / by rosalind r. bruno. — vashlngtont d. c. : d.s. dept. of commercet bureau of the census : [o.s. g.p.o.. distributor] ; 1987. ivf 27 p. : i for« ; 28 cm. — (current population reports, special studies series p-23 ; no. 149) shipping list no.: 87-64-p. "issued january 1987." includes bibliographical references. c" nnbbc 29 sep 87 15238845 vvbrdc see next crd summer 1988 28 iassist quarterly illustration 6 u.s. department of commerce. statistical abstract of the u.s.. 108th ed. washington, d.c., 1988. no. 596. child care arrangements of children under 15 of employed mothers, by age of child and employment status of mother: 1984-1985 |ln thouund*, •xc«p( percent as ol wintef 1984-1985. oau were otilained lex \tw ttue^ youdqd^l chiloiun unckii is yoan okl linckickng any adoplad (x slepct\ikfettn n tr>u« caie) m me household. this repceminu appfoiimaluiy 90 peiceni ol eil chanen under is ytiatt oid ol vorlung women. bamd on the survey ol income and pi09iam pa/1w:ipdtjon. ism lexl secoon )4| usual weeiar cuajd camc amhamgement tow cd/a n cttkfi homeby lather ^ by grandparent . by olher raiatwa . by norveiativa ca/« n another home.. by grandpareni by olher (e<at)v« by norveialiv* oganued ched care lacilrbes.. oay/^oup care center nt^sery »cnuo</pre:>choqj _ kindofganen/giade ktvjol . ctm caroi lor vhi . paient car«3 lor child ' krcemt cnstkibution ciiie n cttkft home.. by tacrwr by grandpareni by o<hdr relauve by nonruiabve c<ire n arkilher homeby yarklpa/enl by ot-her reiatrv« by noraulativ« oodmrvd cttm care laalilies_ day/gtoi4) care centu hutiery »chool/prescftoo( _ kind^garumv/giade tchooi chad carei lor m)i pa/em cares lor chikj • chuxiren under 16 years 3e,455 . 4.699 2.496 712 604 667 3.601 t.l3a 467 2,196 2.411 1.440 671 13,815 488 u45 ^ n not afplkvalle • kuaom* job* c«>..i(«iiu 35 hourk or moie (ft wuth, » i; mother employed full ume ' 16.813 2.480 1.133 423 539 365 2.675 743 285 1,647 1.830 1.067 763 6,976 354 497 part ume t.643 i2l9 1,363 289 265 302 1.126 395 162 &49 561 373 208 4.639 134 74u 100q 230 14 1 30 2.7 3.1 117 8 1 39 22 502 1 4 78 8,168 2.534 1,282 467 306 479 3.020 633 366 1.819 1.668 1,142 746 663 100.0 310 15 7 6 7 37 59 37.0 102 4.5 223 23 i 140 8 1 children under i years mother employed fuu ume 6.060 1,235 542 259 163 251 2.135 633 212 1,390 1.415 635 560 252 100.0 24.4 107 6 1 36 5.0 422 105 42 275 28 165 115 part time s,108 1,300 740 209 123 228 664 300 155 429 473 307 166 412 418 238 6.7 4.0 7.3 28 4 c7 50 13 8 under 1 yea/ 13 («) 133 516 252 102 563 174 70 319 195 lie 79 100.0 373 18 2 74 32 65 40 6 126 5 i 23 152 9 9 6.4 6 3 5 7 1 and 2 years 3,267 1,066 528 208 147 165 1.368 361 130 677 563 401 162 26/ 327 162 57 41 9 it 40 266 172 123 50 n 950 502 157 114 115 1.069 2aa 167 624 1,131 625 506 265 270 143 45 33 50 31 65 47 17 7 322 178 14 4 17 m 8.1 diun 5-14 ye art. loial 10,217 118 66 ^52 27 32 icliidui momun wu<kir.g m hutm or io.iy summer 1988 iassisl quarterly — 29 illustration 7 cuny libran'. card catalogue. ref hb current population reports. series p— 820 70f household economic studies* — e36 no. 1(3rd quarter 19s3)-no. 6 (4th quarter 1984) ; no. 7• — washlngtont d.c. : u.s. dept. of commercet bureau of the census i for sale by the supt. of docs.f d.s. g.p.o., 1984v . \ 28 cm • quart er ly • "average monthly data from the survey of income and program participation." report nos. 1-6 have title: economic characteristics of households in the united states; report nos. 7have distinctive titles. report no. [2] has designation: p-70-83-4. i ' issn 0886-5^—^ 698 = current nnbbc 03 sep 87 1156j7017 vvbrdc see next crd summer 1988 30 iassisl quarterly illustration 8 american statistical index. 1987 ed. washington, d.c., congressional information services. tables: (table* 1-6 show dau for dec 1984. carctaken of children include parents, adult siblings, other adult relatives, unrelated adults, nonadults, and sclf-carc ' data by type of household arc shown for all. married couple; and female-headed households.] children [tables show number of children aged 5-13 years old enrolled in school] ^ 1-2. [by] afler-school caretaker of children,^!^ age of child, type of household, labor force status of mother, and race. (p. 7-12) 3. [by] after-school care [caretaier] of children whose mothers work full time by occupation and education of mother, family income, and race. (p. 13) 4. [by] hours of care for children who regularly spend time not under parents' supervision, type of caretaker, and period of day. (p. 15) households [tables show number of households with children aged 5-13 years old enrolled in school] 5. by whether fully cared for by parents after school and whether any child was regularly not in adult care, by type of household, labor force ^ status and education of female householder, / and family income, (p. 16) ( , „ , • ..,, . l ,1, , . „v estimafion. chapter 12. population and 6. by number of chjdren and whether any \ ,, .y ^ ,...,.' ... ^_ . „, .:, vjhqusing content items child was not m adult care after school, by —'"^ " inhot force status of female householder and described below. part a is described in asi 1986 annual (or 1986 monthly supplement 11) under thi*. number. the remaining 4 parts have not yet been issued. a similar report was issued for the 1970 census (see asi relrospcciivc edition and lsl-3rd annual ^uih^lemenls under 2557-1). 2555-2.'2: part b. chapter 4. census promotipn proflram. chapter 5. field enumeration [dec. 1986. 16+103 p. phc80-r2-b. price not given. asi/mf/4] contents: chapter 4. includes narrative discussion of census promotion program objectives and activities; supporting organizations and program participation; facsimile advertisements and other promotional materials; and program evaluation, (p. 4.1-4.16) chapter 5. includes narrative discussion of census field operations, organization structure, logistics, personnel and training, and mailing and interviewing procedures; lists of district offices and publicand field-use forms; facsimile reporting forms; staftmg calendars; and 6 methodo^-..^^logical tables, (p. 5.1-5.103) 2555-2.3: part c. chapter 7. sampling and race. (p. 17) trends 7. [number of children by] after-school child care arrangements [caretaker] for children 51 j years old [and] labor force status of mother oct 1974 and [dec.] 1984. (p. 17) [dec. 1986. 9-t-75 p. phc80-r2-c. price not [\ given. asl/mf/3] ^ \ contents; chapter 7. includes narrative discussion of sample dcsi|;n and features, estimation procedures, and sampling variability and errors; and list of references, (p. 7.1-7.9) chapur 12. includes narrative discussion of each population and housing questionnaire item, its purpose and history, user instructions, and computer editmg and processing specifications; facsimile survey forms; computer edit sequence; and lisu of instructional and classification codes. (p. 12.1-12.75) summer 1988 lassis! quarterly 31 eiustration 9 american statistical index. 1987 ed. washington, d.c., congressional information services. child day care afdc eligibuity and payment errors, by type and state, 2nd half fy84, semiannual rpt, 4692-1 afdc recipients demographic and financial characteristics, by slate, fy83, annual rpt, 4694-1 employer-sponsored child day care, finances and operations of federal program by agency, with data for selected private firms, 1985. gao rpt, 26119-104 employment in selected highand low-growth occupations, by sex and race, 1980 and projected to 1990, 9248-19 expenditures for child care by family composition, and working women's child care arrangements by type and payment source, 1981-82, article, 1702-1.610 food aid programs of usda, costs and participation by program, fy69-85, annual rpt, 1364-9 food aid programs of usda, participants and costs by program, region, and state, monthly rpt, 1362-14 food service estabushments and sales, by establishment type, 1977 and 1984, aimual rpt, 1544-22.4 handicapped children, by household composition and other characteristics, arrangement for care, and effects on family. 1981, 4948-5.2 health screening at child day care centers, costs and accuracy of diagnoses, local area study, 1986 article, 4042-3.614 hepatitis cases by infection source, age, sex, race, and state, and deaths, by strain, 1984 and trends from 1966. 4205-2 income tax returns of individuals, by filing status, tax item, and income level, 1985, annual article. 8302-2.618 income tax returns of individuals, detailed data, 1983, annual rpt, 8304-2 labor supply, demand, turnover, and training by source, by detailed occupation, 1984 and projected to 1995, biermial rpt, 6744-3 occupational outlook handbook. 1986-87. see also youth employment child support and alimony afdc eligibility and payment errors, by type and state, 2nd half fy84, semiannual rpt, 4692-1 afdc slate admin agencies performance measures, caseloads, payments, and costs, by stale, fy83-84. annual rpt, 4694-2 beneficiaries of noncash public and employer-based transfer programs, by income source and socioeconomic characteristics, 1984, annual current population rpt, 2546-6.46 child support enforcement program financial and operating data. fys 1-85, annual rpt, 4004-16 collection of child support. states using selected methods including wage garnishment, various dates 1986, gao rpt, 26121-119 fed govt spending in slates, by type. program, agency, and slate, fy85, annual rpt, 2464-2 hhs financial aid, by program, recipient. stale, and city, fy85, annual regional listings, 4004-3 income (household, family, and personal), by source, detailed characteristics and region, 1984. annual current population rpt. 2546-6.48 income lax reiums of high income individuals wilh and without tax liability, income and tax items, 1983, article, 8302-2.614 income tax returns of individuals, by filing status, tax item, and income level. 1985. annual article. 8302-2.618 income tax returns of individuals, detailed data. 1983. annual rpt, 8304-2 income tax returns of individuals, selected income and ux items by income. preliminary 1984. annual article. 8302-2.617 income tax returns of individuals, selected summer 1988 32 iassisl quarterly illustration 10 monthly catalog of the u.s. 1987 ed. washington, d.c., government printing office. n, v. : ill. ; 28 cm. annual $1.00 title from caption. previously classed: c 56.216:ma32 e shipping list no: 86-771-p. 1985. description based on: 1975. ©item 142-a s/n 003-024-06472-3 @ gpo 1. glassware— united states— statistics — periodicals. l united states. bureau of the census, sn-87042353 oclc 03060022 87-7041 c 3.158js1a 36 q (sshl current industrial reports. ma36q, semiconductors, printed circuit boards, and other electronic components /u.s. department of commerce, bureau of the censtis. washington, d.c. : the bureau : for sale by the supt of docs., u.s. g.p.o., 1986supl of docs., u.s. govt print off., washington, d.c. 20402 v. : ill. ; 28 cm. annual $1.00 1985title from caption. shipping list no.: 86-764-p. 1985. ©item 142-a s/n 003-o24-06215-1 @ gpo 1. semiconductors— statistics— periodicals. 2. sohd state electronics— statistics— periodicals. i. united states. bureau of the census. par 86-642605 oclc 14353189 87-7042 c 3.164:455/985/7.1 u.s. exports. world area and country by schedule e commodity groupings. washington, d.c. : u.s. dept of commerce, bureau of the census : for sale by the supt of docs., u.s. g.p.o., 1983supt of docs., u.s. govt print off., washington, dc 20402 v. ; 28 cm. annual $31.00 1982"ft 455." 1985 dec. and annual, v. 1. issued in 2 parts, 1982vols, for 1982distributed to depository ubraries in microfiche, •item 144-a-9 (microfiche) s/n 003-02406510-0 @ gpo issn 0741-8310 continues: u.s. exports. world area by commodity groupings issn 0360-2249 1. commercial products— united states— statistics — periodicals. 2. commercial products— united states — classification — periodicals. 3. united states— commerce — statistics — periodicals. i. united states. bureau of the census. hf105.b73d 83-647834 382/6/0973 /19 oclc 10092103 87-7043 c 3.186j>-23/149 bruno, rosalind r. after-school care of school-age children : december 1984 / by rosalind r. bruno. — washington, dc. : u.s. dept of commerce, bureau of the census : [u.s. g.p.o., distributor] ; 1987. page 28 annual $6.50 began with 1964/65. previously classed: c 3. ping list no.: 87-77-p. 1984-85. description basec earlier v. issued as part of: city finances. #item 003-024-06232-1 @ gpo issn 0082-9439 1. municipal finance— united states— statisti cals. 2. local finance — united states— statisti cals. i. united states. bureau of the census. ke government finances il series; government fina 4. hj9011.a4b 74-648912 //r81 336.73 ocu 87-7045 c 3j04/3584 county business patterns. united states. washington etept of commerce, bureau of the census : foi supt of docs., u.s. g.p.o., supt of docs., u.s. off., washington, d.c. 20402 v. : ill ; 28 cm. annual $5.50 began with 1973. previously classed: c 3.204: no: 86-778-p. "cbp-84-1." 1984. description 1978. ©item 133-a-52 s/n 003-024-06376-0 @ tinues in part: county business patterns, u.s. summ 1. united states — industries— periodicals. i. ' bureau of the census, sn-87042346 oclc 0753^ 87-7046 c 3.20s/3:wp-«5 world population profile. washington, d.c. : u.s. l merce. bureau of the census : for sale by the st u.s. g.p.o., 1986supt of docs., u.s. govt washington, d.c. 20402 v. : col. ill., col. maps ; 28 cm. $4.25 1985shipping list no.: 86-990-p. "wl •item 146-f s/n 003-024^218-6 @ gpc world population issn 0099-1139 i. population — statistics— periodicals. i. u bureau of the census, sn-87042039 oclc 15207 87-7047 c 3j15/16:982 current housing reports. series h-171, annual hoi supplementary reports. no. 1. summary of housing tics for selected metropolitan areas. washington, dept of commerce, bureau of the census : for sa services division, customer services (publications the census, data user services division, custot (publications), bureau of the census, washington, v. : ul., maps, form ; 28 cm. annua] summer j 988 iassist quarterly 33 illustration 11 pais bulletin. 1987. new york: public aftairs information service. tain prospects for isdn: a spanish '.ommunicatioas policy 10:313-24 d ccs digital network. senegal needs informatics skills, table :a and communicatioas kept 9:13-14 computers, telecommunications, and ;. advanced information and echnologies: challenges and ^reconomics 22:79-84 mr/ap '87 society, the economy, and ions. e. com. on public works and :asibility of allowing fiber optic cable ;e system: joint hearing, april 15, lubcommittcc on economic the subcomraittce on surface 7 iv+227p il ublc diags chans maps ess.) ([pubn. no.] 99-63) (sd cat. no. 3) pa — supt docs e. com. on the judiciary. subcom. on rtics, and the admin, of justice, anications privacy act; hearings, 55-march 5, 1986, on h.r. 3378. '86 99th cong., 2d scss.) (serial no. 50) lematioiial aspects last data transfer the german tmnsnadonaj data and iept 9:11-14 ag '86 k conceptual framework for the insborder dau flows, iiu'o society ;e flow of information, national lependent development. ide in data services: the international rts telccommiuxicatioas policy fmdings in his forthcoming book, sactions in services: the politics of lows.' li republic of germany, table ctc mal coqxs) reporter p 14-15+ regulation mcllors, colin and david poiliti. policing the communications revolution: a case-study of data protection legislation. west eur politics 9:195-211 o •86 role of the oecd and the council of europe; national legislation. social aspects case, donald and everett rogers. the adoption and social impacts of information technology in u.s. agriculture, bibl info society 5:57-66 ao 2 '87 case study of an experiment in which 200 kentucky farmers were given 'green thumb boxes,' videotex devices providing market, weather, and technological information through their television sets, 1980-81; conference paper. day care centers child care services in singapore, 1981-1986. table charts singapore stads news 9:1-4 no 1 '86 prepared by the child care branch, ministry of community development, t kahn, alfred j. and sheila b. kamcrman. child care: facing the hard choices. "87 xi-h273p tables index (lc 86-28710) (isbn 0-86569-164-9) $2&—auburn house major demographic and social developments; the reagan administration's emphasis on privatization; pros and cons of local and state initiatives; policy options. landers, roben k. new deal for the family, bibl il chan editonal research repts p 551-68 jl 25 '86 contents; juggling motherhood and jobs; feminists reconsider aims; pro-family politics. issue of maternity leave; day care policies; u.s. schwenk, frankig n.-child-caxc arrangements and cxpcnditurc^ibl tables ch^ns family ecoa r p 1-7 ao 4 '86 ' .— ^ united sutes. dead bodies (law) andrews, lori b. my body, my propcny. hastings center rept 16:28-38 o '86 whether patients (or their heirs) should be allowed to share in the profits derived from research conducted on the body pans and producu of the patient; the issue of informed consent. dealer relations summer 1988 ia5sist newsletter, vol. 2, no. 1 (winter 1978) action group reports the action group reports which follow are a result of the action group discussions at the lassist annual conference, itasca. illinois, in february, 1978. those not appearing in this issue of the newsletter will be published in subsequent issues. the reports are printed as tiled by the chairpersons of the action groups. if you would like a copy of the style manual, contact: barbara b. noble, center for advanced computaaction in the tion. university of illinois, urdocumeiitation action bana, il 61301, telephone (217) g^oup 333-3234 before april 1, 1978, or at bureau of social science rethe documentation action group search. suite 700, 1990 h st. n.h., is currently working on two prowashington, dc 20036 after april iects of interest to lassist mem1. bers; a study description form and . a style manual for documentation of sheldon laube, chair machine readable data. the study description form has been under development for some data organization and time, mostly in europe, under the ^1111511151 icttun auspices of the international fed5551? eration or data organizations (ifdo) . the purpose oi the study the dom ag met for two sessions description form is to present in a on thursday, february 9th at the standard outline format the essenmeeting held in itasca. during tial facts about a machine-readable these sessions the following activdata file. the main sections are: ities were carried out: identification and acknowledgements (study title, principal investiga1. ms. nancy morrison from tor, data distributor); analysis the university of illiconditions (purpose of study, numnois has been active in ber of units and variables, etc.) ; the formation of the new reanalysis conditions (data condispss user's group. nancy tion, accessibility, documentation, provided the ag with a etc.): references to publications; report on the progress of and background variables (demothe user's group. this graphic and socioeconomic) . batch report stressed four maand interactive programs in ibm jor points. os/360 assembler language and pl/1 are available for entering and a) an organizing commitprinting study descriptions. tee is presently preparing the by-laws for the style manual for documentathe group that will tion of machine-readable data is make it separate from the product of a research project spss, inc.. undertaken by the university of illinois for the us department of b) a membership drive is justice. the manual is intended to planned for april, be a general guide for the format 1978. individuals inand contents of user's guides terested in being (those things we sometimes call placed on the mailing •'codebooks") for hrdf. the basic list for membership sections of a user's guide dismaterials should write cussed are: title page, abstract, to 'as. nancy morrison, project history, processing hissocial sciences quantory, codebook, and appendices. titative laboratory, lincoln hall, univerthe documentation action group sity of illinois, urand the projects' sponsors are inbana. 111. 51801. terested in field testing both the study description form and the c) the group plans to act user's guide style manual, and in as a communication veany comments or suggestions. hide among spss users. for more information about the study description project contact: d) it will act as a codr. elliott mavendon, leisure herent group that can studies data bank, waterloo reprovide input to spss, search institute, university of inc. waterloo, waterloo, ontario n2l, 3g1, telephone (519) 885-1211. 23 ias5ist newsletter, vol. 2, no. 1 (winter 1978) 2. a workshop was conducted on the treatment of multiply-punched data. a round table approach was employed and individuals discussed procedures used by the organizations they represent. we were fortunate to have present during this session roald buhler of princeton university. dr. buhler discussed a program (stand alone) at princeton that he uses in conjunction with p-stat. in addition, mr. william gjertsen of the sa3 institute reported that the new version of sas will be capable of handling multi-punch data. terry stewart of the leisure studies data bank, university of waterloo and bill oammell and gary grandon, both from the university of connecticut, shared information on procedures used at their respective installations. 3. dr. richard roistacner reported on the progress of the data inter-cnange file concept which was first introduced to the dom ag at the first lassist north american working conference in floriaa last year. at the itasca conference, dick met with representatives of spss, sas, p-stat, osibis iii, osiris iv and tpl. tne outcome of this meeting was that everyone was in agreement that the time had come for an intercnange file prototype to be developed and tae actual coding should start in the near future. in addition, a national online conference will be set up this spring to keep the discussion of the interchange file active. future developments on this subject will be reported in lassist newsletters. h. there was an active discussion on the effectiveness of the dom ag m its present format. those of us who have been with this action group since its inception have become increasingly distressed with our inability to achieve a high level or activity on all the items in our ag mandate (see newsletter vol. 1, no. 1) we have experienced considerable success in providing a form for the transfer of information and informal training. unfortunately, the larger task of investigating and evaluating "existing procedures for data and documentation preparation and data management software and hardware capabilities" has only been marginally attended to. some of the reasons offered for this problem were lack of time together and resources, over ambitious expectations, and an overlapping ot interests with other ag's. whatever the reason, it aopears that tnis is a point of discussion that should involve all interested parties. you are encouraged to provide input to bill gammell, box u-164, university of connecticut, storrs, connecticut. 06268 on the brighter side of things, alice bobbin (dad a3) reported to us that our group supplied 5 of the 9 abstracts received for portions to providing ence data please refer ag report in this issue for aetails concerning this project. william gammell, chair action group on process?rdxuc1d~mti priority; #j we wound up discussions on our highest priority project, the directory of directories (d of d's). satisfied that we published a good start of the american list, (newsletter vol. 1, no. 4) , we turned our discussions to the canadian list. it was agreed the best wav to proceed would be to tap the resources of the data clearinghouse. we estimate that some information will be ready by march 31 and will be published m the forthcoming june issue of the newsletter™ this would be appended by addenda to the american list. this addenda will be additional listings of directoof a guide social sciservice. to the dad 24 lassist newsletter, vol. 2, no. 1 (winter 1978) ries, and another announcements and b are published separa rectories. what im to mind was announcem publicaly available statistical reporter, register, the data and others. when t ject is printed once it were) the action tempt to edit a comp which can be updated basis after that, elements for each ent vol. 1 no. 3 p. 13) w ered in that final ed vision this to be co the projected next n meeting in ottawa, ma is later than our pr cation for the upp however, absences tr conference and oth have retarded our pr assessment of priorit all along that it wa to be mastered and 1 be significant. on have been successful catagory of ulletins that tely from dimediately came ents of recent data in the the federal clearinghouse , he entire pro(by stages as group will atlete d of d's, on a yearly the essential ry (newsletter ill ee considition. we enmpleted before orth american y 1979. this ojected publisula meeting, om the itasca er exigencies ogress. our y #1 has been s small enough arge enough to that count, we priorities #2 and ±3 regarding our second priority. an inventory of existing guidelines, etc., and our third priority project. an inventory of procedures, etc. (newsletter vol. 1 nr. 2 pp. 14-15) we aecised to divide into subcommittees as follows: don harrison will begin to bring together into one package a discussion of the guidelines and procedures existing today in the united states federal government both in the national archives and in those federal agencies that create hrdf's. tony falsetto will bring together a similar document for the canadian federal government. charlotte bochan and harriet dhanak will form a subcommittee to explore a number of alternatives necessary to proceed with institutions in the private sector. the larger corporations such as hand and atst might provide a basis for a representative sampling of guidelines and procedures. smaller corporations might also be gueried but might not provide such a return. united states state governments and canadian provincial governments, while still on the threshold of creating and servicing mrdf's will also be significant. a projected telephone survey by the society of american archivists, when completed, might add to this and other projects by ppdag. look i subcom the w area w these freque bers i beth p newest subcom back t don harris nto this sur mittee compos ashington , ill attempt concepts nt local meet nclude don h owell, himi memfcer, de mittee repor o don by frid on vey. ed o dc to tog ings arri scha bbie ts ay m was asked to also, a f members in metropolitan tie some of ether with . these memson, elizade, and our pomerance. will be due arch 31. this project cannot be initiated before projects #2 and #3 are complete. therefore, no progress was reported. iliter^action grou£ cooperation ppd e do e a ript lope r of rev th n ovid on oble is s nth. ag w cume ppli ion d b pp lew on-s e f e ag ms i houl ill ntat cati form sev ag w the urve edba rega n it d be provide ion acti on of th which h eral memb ill work form for y data, ck to th rding an s applica complete feedb on gr e stu as b ers. with gene we wi e doc y. po lion d in ack to oup on dy deeen dea memthis ag ral use 11 also umentatential to ppd. about a deaccessioninq ppd duri asca me f led a ters d without researc agreed the abs tuted a ke will newslet problem critiqu sion at ng th( eting, ser ioi eacces cons; vi to wc ence ction prei ter tl and es in the ( works part pro ionin erati ue. k on f a roup re so t wi all f repar tawa hops icip blem g un on do this form for meth 11 i or at io meet of ants , da ique of po n h pro ally acqui ing denti comme n for ing. the itidentita cenhrdf's tential arrison ject in constisition. for t he the nts and a sesglossary 25 4 iassist quarterly 2015 iassist quarterly editor’s notes avoiding disclosure, using data, and improving digital preservation welcome! the three papers of this second issue of volume 39 of the iassist quarterly (iq 39:2, 2015) are about data. matters of privacy, disclosure and surveillance have been, and will continue to be, hot topics in society. within science we have obvious reasons to ensure that research cannot be accused of breaches, and therefore researchers need principles and guidance on how to avoid statistical disclosure. staff of research data centers and universities demonstrates in the first paper their investigation of this sensitive area in ways that we all can learn from. the second paper brings insights in the use of data collections and research design in business master’s theses at a canadian university. better insight of the use means better guidance for users. the third paper shows how modern data archives are continually improving their data procedures by applying standards and using measurement. all three papers will give you insight into these data areas and they are equally well furnished with references so you can explore the areas further. the first paper ‘principlesversus rules-based output statistical disclosure control in remote access environments’ is authored by felix ritchie from bristol business school, university of the west of england, and mark elliot, school of social sciences and data research institute, university of manchester. delivering data for research requires measures against disclosure risk. the authors demonstrate the differences between the two approaches where ‘rulesbased’ has been the traditional model. this is where a set of formulated rules are applied ahead of dissemination to researchers. a common example is ‘a table may only be released if there are at least three observations for each cell’. the set of rules can be comprehensive and each rule has a trade-off between confidentiality and efficiency. the authors argue that disclosure control based on principles is to be preferred, but demands more training which will in turn build a culture of confidentiality expertise. linda d. lowry is the liaison librarian at brock university in st. catharines, ontario. her paper ‘bridging the business data divide: insights into primary and secondary data use by business researchers’ is based on analysis of 32 master’s theses within management at brock university. the applied content analysis is demonstrated and the coding of the theses is explained in informative appendices. the analysis describes the research designs and data collection methods found as well as the distribution of theses in business subfields: accounting, finance, operations and information systems management (o & ism), marketing, and organization studies. it turns out that the distribution of master’s theses shows nearly half are within the finance subfield. furthermore, all of the finance theses are ‘archival quantitative’ while nearly all marketing theses are ‘questionnaire’-based. the conclusion is a recommendation of content analysis for examining data practices within social science disciplines. mari kleemola is information services manager at the finnish social science data archive (fsd), and she outlines fsd’s venture into digital preservation standards and assessments in the contribution ‘improving the quality of digital preservation using metrics’. the paper demonstrates the journey through standards, checklists and assessments, such as: open archival information system (oais), the audit and certification of trustworthy digital repositories (tdr) checklist, and finally the cessda trust process. the complete process is described with the key issues of measurements and requirements of each step in the process. what started as an internal handbook and a formation plan in 2003 has especially during the last four years been developed until the fsd received the data seal of approval certification in september 2014. articles for the iassist quarterly are always very welcome. they can be papers from iassist conferences or other conferences and workshops, from local presentations or papers especially written for the iq. when you are preparing a presentation, give a thought to turning your one-time presentation into a lasting contribution to continuing development. as an author you are permitted ‘deep links’ where you link directly to your paper published in the iq. chairing a conference session with the purpose of aggregating and integrating papers for a special issue iq is also much appreciated as the information reaches many more people than the session participants, and will be readily available on the iassist website at http://www. iassistdata.org. authors are very welcome to take a look at the instructions and layout: http://iassistdata.org/iq/instructions-authors authors can also contact me via e-mail: kbr@sam.sdu.dk. should you be interested in compiling a special issue for the iq as guest editor(s) i will also be delighted to hear from you. karsten boye rasmussen september 2015 editor http://www.iassistdata.org http://www.iassistdata.org http://iassistdata.org/iq/instructions mailto:kbr@sam.sdu.dk the organization and management of the norwegian social science data services (nsd) jarle brosveet, bj^rn henrichsen^ lars w. holm, and terje sande norwegian social science data services bergen, norway (this paper was deliveved at the 1981 ifdo/iassist conference, grenoble.) the early years the origins of the norwegian social science data services (nsd) go back to 1967 when the norwegian research council could no longer ignore various complaints about, long delays and frequent errors in the processing of social science data. ' as a result the council decided to set up a committee to study the reasons why bottlenecks occurred in computing services at the universities. the committee was also instructed to make proposals for improvement of social science computer facilities and services. the committee made a number of recommendations and took the first steps to implement these as part of its work. a series of intensive courses in computing for social scientists was organized, a variety of statistical packages and specialized programs were acquired and installed on the university computers, and steps were taken to build up a first set of data banks easily accessible to researchers at all universities. in 1970, the committee proposed to institutionalize these activities by establishing a nation-wide organization supplying various types of services to all social scientists. the research council approved the proposals and established the nsd for a trial period through 1974. later a new program of activities was approved for the next three-year period. as the activities proved a success, the nsd was given permanent status from january 1st, 1978. 1) for background information, see stein rokkan, "national primary socio-economic data structures iii: norway", international social science journal 30(3), 1978:621-652; also stein kuhnle and stein rokkan, "political research in norway 1960-1975", scandinavian political studies 12, 1978:127-156. 73 the organization the nsd differs from most other organizations of this type in three ways : it is a federally structured organization with offices at four universities and with close working arrangements with the regional colleges; it has built up a wide variety of data resources across all fields of the social sciences: the data holdings comprise surveys as well as large databanks for communes and census tracts, an archive of information about organizations, and a series of computerized files of data on the recruitment and careers of various elite groups; it has established close contacts with a wide range of users in the social sciences and in administrative bodies, partly through its regular newsletter (brukermelding) , partly through the annual meeting of representatives of a great majority of research institutes and teaching departments active in the social sciences throughout norway. the nsd has its headquarters at the university of bergen. in addition there are external data secretariats at each of the other universities: oslo, trondheim and troms0. the data the nsd has so far built up seven basic types of data holdings: i the first and largest is the database for communes . this database covers statistics for all local units of administration since 1800 and is linked up with a facility for computer cartography (polyvrt, calform, symap, figur). this is our most widely used facility and is constantly expanded and improved. a great amount of work has been expended on finding effective solutions to the problems posed by changes in boundaries and units . the database offers detailed documentation of these changes and provides the user with a variety of coefficients for recalculating data values whenever changes occur. ii the nsd has set up a cartographic service useful to a number of users. coordinate matrices for all commune boundaries in norway since 1800 have been computerized. the boundary segments are identified by a time key to allow the production of maps for every year since 1800, initially, maps were mostly conformant using cross-hatchings to represent the data values. later programs have been developed to display information at the centerpoint of every commune. this development makes it possible to display not only multivariate statistics, but also discrete data about the presence or absence of particular infrastructure elements such as airports, harbours, hospitals, schools, factories etc. iii to allow analyses at a lower level of aggregation the nsd has taken steps to organize a set of data at the lowest level of official 74 enumeration: the census tract . this data bank contains data from the censuses of 1960 and 1970 and has proved of great interest to city planners, geographers, anthropologists and sociologists. iv in dealing with survey data it has been our policy to link up response data for comparable questions over time , so we have deliberately declined to serve as a depository for individual studies. the largest file so far developed covers the surveys of the norwegian gallup institute (roughly monthly ones) from 1964 through 1976. currently, this archive is being updated with data from the morsk opinionsinstitutt a.s. (norwegian opinion institute). a similar linkage is being completed for the electoral surveys carried out by professor henry valen and his colleagues since 1957. some of the most thorough surveys carried out by the central bureau of statistics since 1967 are now at the disposal of academic users under an agreement with the nsd. the nsd has prepared detailed documentation for some of these surveys and reformatted them for use with spss to increase their availability. v a major effort has been made to develop a systematic databank for elite groups. the first file to be completed contains biographic information on members of parliament since 1814 , this archive gives information on father's occupation, education, early career, positions in legislative conmittees etc. the file has recently been linked with one for roll calls since the establishment of party fronts in the 1870's to allow analyses of interrelations between background factors and legislative behaviour. similar biographical files are being organized for members of the central administration as well as for graduates from the universities since 1811. vi another file taken over by the nsd is the archive of information about voluntary associations in norway 1971/1972 . vii the nsd has also got a file of test data for the recruits to the armed forces from 1951 to 1968. in organizing these as well as other data the nsd has emphasized the paramount importance of documentation and data linkage . the nsd does not wish to serve as a pure depository of data. consequently, data files are rarely incorporated into our holdings unless they can be fully documented and linked up into a system either topically or across time periods. to make sure that our facilities are used extensively, the nsd has given high priority to educational activities , courses, lectures and teaching packages. active support has been offered within the programme of the international social science council as one of the cross-national workbooks in the issc series has been produced in close cooperation with the nsd. steps have also been taken to compile a series of specifically norwegian teaching packages. one of these is based on fiscal -administrative data for communes , another focuses on time-series data for intermediary levels of regional aggregation (fylker ). further packages are based on data from electoral surveys and the file of biographical information on members of parliament. 75 matters of policy and economy up until 1977 nsd was almost exclusively financed by grants from the social science division of the research council. at that time it was realized that nsd holdings were used extensively by groups other than social scientists and for purposes other than research and teaching. even more important, some governmental agencies as well as the nsd realized that our data archives carried a huge potential for several kinds of governmental planning and research. this potential was likely to grow not least as a result of increased co-operation with the central bureau of statistics. consequently, our financing had to be altered in one way or another. several models were discussed, but in the end the nsd is still being financed on a project basis. a major advantage of this arrangement is that it allows the research council and the nsd to set the priorities. at the same time we can also choose the governmental projects that appear to give the greatest benefits to the social sciences in general. from 1978 to 1981 the basic grant to the nsd increased from about 1 mill, to 1.4 mill. nkr. , each year totalling about 8% of the research council allocation to social science projects. if this were to be our total budget, it would restrict our expansion considerably. however, we have been able to obtain supplementary funding elsewhere, including grants from other divisions of the research council, the nordic research councils and various governmental agencies. the current funding of tne nsd allows us to keep a reasonably large group of professionals permanently employed. as of today all full-time staff except the secretaries have a university degree. our recruitment policy has always been to emphasize social science background with experience in data analysis and programming. such experience is achieved through the training of recruits as assistants either at the nsd or on large research projects at the universities. the academic background of the staff covers almost the full range of the social sciences: economic:, sociclccy, political science, geography, public administration, history and information science. not only does this provide a good interdisciplinary working group, it also makes it easy to communicate with various research interests. the recruitment policy definitely reflects the tasks that have been assigned priority by the board of the nsd. to put it simply, data and documentation are given priority over software development and programming. however, the nsd is seldom involved in primary data collection. in effect, we place great store by establishing good working relations with data collecting agencies in administrative bodies, private firms and research institutes. obviously^ the most important of these is the central bureau of statistics. we now cooperate rather extensively with the bureau and also exchange services to some extent. this co-operation means that we are able to provide social scientists with virtually all data collected by the bureau at a very low cost that often amounts to no more than the copying of the magnetic tape. the nsd has also been assigned the responsibility for giving social scientists access to bureau data that are not completely anonymized, provided that the bureau receives feedback as to who is given access. furthermore, the nsd is represented on committees to review census questionnaires and bureau tables to be publicized using census iata. 76 to a certain extent we are developing similar relationships with other governmental agencies supporting data collection tasks. currently, we are cooperating with the ministry of consumer affairs on the computerization of a register of representatives in official boards and councils. we do part of the coding and punching as well as the running of simple tables to be used in the official report. later, these data will be established as an archive and included in the data holdings available to researchers free of charge. there are several reasons why such excellent and confident relationships can be developed vis-a-vis the governmental agencies. one of the most important reasons seems to be that there has never been any abuse of data by students or researchers. the nsd has also shown the ability to store and document data in a way that permits simple computer runs, map drawings etc. much faster, cheaper and often with better quality compared with other data distributors. since the holdings of the nsd contain data collected by means of public funds, there is also a strong argument for maximum public usage of these data. before leaving our relationship with governmental agencies, we must also mention the effect of privacy legislation on social research. as soon as the privacy law was proposed, the nsd took the initiative to help fulfill the intentions of the law as well as to counteract some of the negative effects experienced in other countries. after lengthy negotiations we are now in the process of implementing a strategy that we believe is advantageous for all parties concerned. first, the nsd has been given a general consession by the data inspectorate to archive and store relevant social science data. second, we have obtained a grant for acting as a liaison between the research community and the data inspectorate. it also enables us to set up a secretariat for the research council on matters concerning privacy and data collection. international cooperation on the international side, the nsd has always been cooperating actively with similar organizations in other countries. it was one of the founder members of the international federation of data organizations in 1977 and is active in a working group under its auspices for the coordination of local regional data bases and computer cartography . since 1971, information about european archiving efforts has been published regularly in the european political data newsletter set up by stein rokkan under the european consortium for political research and later co-sponsored by the norwegian social science data services. the epd newsletter has developed a classification scheme for the comparison of the contents of european time-series datasets and has repeatedly advocated actions to co-ordinate developments in this field. first, nsd and the norwegian research council succeeded in securing the sponsorship of the european science foundation for a meeting on "databases for regional analysis" in 1977. 77 second, nsd took steps to develop a proposal for the establishment of a joint nordic database for regional time series during 1977-1979. funding for a three-year period was granted by the four nordic social science research councils, i.e. the danish, finnish, norwegian and swedish. the project started in january 1979. also, at the ifdo meeting on regional data held in turin in march 1980, the task of coordinating an effort to summarize and structure information for a joint european database was entrusted to nsd. nsd was asked to complete the matrix or map of available data on the basis of the work started in turin, to disseminate a report and organize a workshop in which further action can be discussed. the effort implies the following: a. to identify organizations and scholars in each nation who are willing to cooperate, and to look for additional researchers and scholars in cases where regional or thematic gaps in the databases have been identified; b. to establish a cornnon standard for types of variables in these databases; c. to establish a common standard for documentation of these data; d. to carry out an effort as far as possible to establish the comparability of these data. this task does not necessarily imply the establishment of formal statistical comparability, but a more substantive kind of theoretical comparability, such as the possibility of using analogous measures for common research projects of analysis of social structures and processes. nsd concludes from this information-gathering effort that it will be feasible to start the building up of a first version of a joint european data base for regional time series. at this time we have established a network of contacts and compiled information about data matrices and variable lists. a report on the progress of the project has been submitted to the board of ifdo for further discussion. 78 iassist quarterly 2014 5 iassist quarterly editor’s notes getting the big picture and writing the thousand words welcome to the first issue of volume 38 of the iassist quarterly (iq (38):1, 2014). after the extended special issue of volume 37 we are offering you a regular issue with information on some of the areas presented and discussed at the latest iassist conference. at the 2014 iassist conference in toronto many papers and interesting ideas were presented. the number of participants in such gatherings will always be limited for a plenitude of practical reasons. i think i know what you are thinking now yes, often lack of funding is high on the list of obstacles to participation. but is participation necessary in our current era of blogging, tweeting, snapchatting, etcetera? well, the bandwidth when actually attending a conference is still tremendous compared with what we experience through our gadgets. iassist and thus also the iq have ‘technology’ in their names and we don’t perceive ourselves as being low-tech. an iq issue is produced through massive amounts of email correspondence, intense use of publishing software, and finally dissemination on the internet. however, old media like this journal continue to have great advantages. reading cannot be overrated and the use of text and the search for text on the internet should be sufficient proof of that. if you were absent from the great bandwidth of the iassist conference, this iq issue can at least partly bring you some of the information you missed. the chairs of the sessions at this year’s conference were given the opportunity to write short summaries of their sessions. san cannon – working at federal reserve board and currently the regional secretary of the usa – mailed the session chairs and asked for their contributions. as this was done ‘post festum’, as expected we received only two session summaries, and i hope that these are sufficient for our true purpose namely to receive your judgments about the usefulness of such summaries. if this is considered to be useful we will pursue this with an earlier notice to conference chairs at forthcoming conferences. from the 2014 conference, we present a summary of the session on ‘tools and services for supporting research data management’ produced by the chair carly strasser. a panel of four individuals from or affiliated with the uc curation center (uc3) at the california digital library presented tools for data management in libraries and for researchers sharing their datasets, including uploading and providing metadata. the second summary is from the session on ‘big picture metadata’, in which there were three presentations from different settings and parts of the world. the british library is involved in a collaborative project with participants around the world for better connection between researchers, authors and contributors of research data. from the university of michigan staff at isr presented a system for extracting ddi standard metadata from blaise databases; this software is freely available for all blaise users. from gesis in germany there was a presentation of the ddi handbook project; as the use of the ddi standards is steadily increasing, the idea is to produce a collection of best practices in a form between a book and a faq. they have also received suggestions on ‘what not to do’. although negative examples are normally not considered to be among the best pedagogic methods, there are tv shows that have become very popular even though signs warn the public: ‘don’t try this at home!’. in the pecha kucha session there was a huge number of great presentations in the dogmatic format of 20 slides for 20 seconds each. one of the presentations i especially enjoyed was ’data visualization and information literacy’ by ryan womack from rutgers university libraries; it had many pictures and some words but little text. ryan has been so kind as to produce a paper based on his talk. they say a picture is worth a thousand words. here you can experience an author who gave the time to write the thousand words, although not for each of his 20 slides. remember that presentations at the conference are available from the iassist website. if you are waiting in vain for a presentation to turn up as a paper in the iq you should take a look at the website. jungwon yang from the library at the university of michigan has experienced a growing interest among social science researchers in using geographic information systems (gis), and for more easily interpretable interdisciplinary dissemination of results. in the united states local government data as well as fine grained census data have become available, and are being supported by organizations. in the paper ‘discovering and accessing sub-national statistics and geospatial data of east asian countries: trends and obstacles’ she brings insights to these types of data in china, korea, and japan. as promised in the title, she describes barriers found in the use of the data. it is advisable to team up with librarians and specialists to overcome the issues. at the session on ‘research environments research data management’ san cannon from the federal reserve board gave the presentation ‘moving beyond research: building an enterprise data service from a research foundation’. san describes how the frb is now midway through a plan of centralized data management and governance, and is moving from the business-based silos designed to manage data sets that were only sparsely shared to an enterprise-focus with enhanced data governance, management and integration. the multitude of functions and departments of the frb has led to the focus of the project being on obtaining a precise inventory and metadata, where data are systemically and strategically catalogued within system-wide standards. the conclusion on data management is that there is no one-size-fits-all 6 iassist quarterly 2014 iassist quarterly technological solution. when we find smart solutions, it is good to consider if they are too smart, and to recall the saying attributed to h.l. mencken: ‘for every complex problem there is an answer that is clear, simple and wrong’. articles for the iassist quarterly are always very welcome. they can be papers from iassist conferences or other conferences and workshops, from local presentations or papers especially written for the iq. when you are preparing a presentation, give a thought to turning your one-time presentation into a lasting contribution to continuing development. as an author you are permitted “deep links” where you link directly to your paper published in the iq. chairing a conference session with the purpose of aggregating and integrating papers for a special issue iq is also much appreciated as the information reaches many more people than the session participants, and will be readily available on the iassist website at http://www.iassistdata.org. authors are very welcome to take a look at the instructions and layout: http://iassistdata.org/iq/instructions-authors authors can also contact me via e-mail: kbr@sam.sdu.dk. should you be interested in compiling a special issue for the iq as guest editor(s) i will also be delighted to hear from you. karsten boye rasmussen october 2014 editor http://www.iassistdata.org http://iassistdata.org/iq/instructions mailto:kbr@sam.sdu.dk iassist quarterly vol 24 no. 2 8 iassist quarterly summer 2000 web-based data enter classrooms in countries in transition: the unicef national program of education for development in slovakia by dr. dusan soltes * introduction the area of education has been a priority of the slovak committee for unicef’s national plan of activities since its beginning, dating back to the independence of the slovak republic as a sovereign nation on the 1st january 1993. in this respect the main goal has been to contribute as much as possible to the whole educational process in slovakia in such a way that education would become one of the driving forces of the country´s transition to the multiparty democracy, functioning market economy, economic prosperity, social justice and a modern society with a due respect for human rights as it has been in all developed countries of the world. one of the common denominators of these challenging tasks of education has been to prepare the young generation for the future integration of the country into the systems of regional and global systems and in particular to the european union. for the young people themselves it means to prepare them for their future role as full-fledged citizens of the future unified europe with the same rights, duties, but also opportuntites, as have their pals in the current european union. in this respect, the unicef world-wide „education for development“ (efd) programme, part of unesco´s „education for all“, another global education programme, has become a welcomed source of know how and an efficient vehicle for our national activities in the area of education. in the following parts of this paper we will be dealing in more details with some aspects but also problems of implementation of the efd world-wide program in the specific conditions of the slovak republic as one of the countries in the central and eastern europe (cee) in transition. some background information on the general context for the implementation of the efd in slovakia within other cee countries in transition in slovakia, as in all other cee countries, the process of implementation of the efd started in the early 90s, before 1993, as a part of the federal programme of the former czecho-slovak federal republic, and later as a part of its particular national program. in this respect, the main development objectives of the efd have been to contribute as much as possible to achieving the following main objectives of the whole transition process in relation to the young generation and its preparation for a proper understanding and active support and readiness for the following main goals: to establish a democratic, multiparty parliamentary system which would replace the former system of the one-party domination; to create a civic society with the due respect for human rights, rights of children, minorities, etc. and general equality of all citizens irrespective of their national background, gender, social status, etc.; to develop an efficient modern market economy securing a social and economic justice and thus giving the same chances for all in a fair competition on a market similarly as it has already been e.g. in the countries of the european union; to carry out modernization, restructuring, privatization and liberalization of all processes of the economic development internally and externally with the vital assistance of the know how and investments from the developed world and especially from the european union; to overcome a long-year isolation from the outside world and to establish and develop new dimensions of a global as well as regional and crossborder cooperation, communication and trade; and to overcome even more evident isolation and underdevelopment in the field of the freedom and free flow of information, media, ideas but also a free movement of persons, etc. in general, as it could be summed up, the main objective of the whole transition and the corresponding preparation of the young generation have been to prepare the whole country, and especially young people, for the future membership in the european union and in other european and transnational integration structures. d iassist quarterly summer 2000 9 this main objective has found its direct expression among others also in the so-called „copenhagen criteria“ as a criteria to be met by all candidates countries from the cee before they could become members of the european union (eu). the substance of these criteria are as follows: stability of institutions guaranteeing democracy, a rule of law, human rights and a respect for rights of minorities; functioning market economy and ability to withstand the pressure of competition and market forces; and ability to take over obligations of the membership (in the eu) as well as dedication to the objectives of the political, economic and monetary union. these main objectives, goals and criteria for the whole socio-economic development of the countries in transition have of course directly effected not only their whole further development but even more the whole system of education. it is then no surprise that also some articles of the particular association treaties between the eu and the countries in transition including slovakia have directly been dedicated to the development and harmonization of education, to the technical and financial assistance in transforming the whole system of education according to the standards of the eu. in this respect, the importance of education and, in particular, education for development has become even more important as it has to prepare the young generation for becoming citizens of the eu in equality with the same rights but also qualifications as needed for achieving such a challenging goal. just for illustration, one of the four basic freedoms i.e. the free movement of persons is unthinkable without a mutual recognition of qualification, education certificates, diplommas, without a good knowledge of foreign languages, etc. in general, the countries in transition and their education for development programs have to some respect, very similar or even identical features with the developed countries, such as: no illiteracy and the general level of education is relatively high and available for free to the whole population; no widespread poverty, famine, malnutriation and/ or starvation of children or young people as it is existing in some developing countries; a relatively still good standard of medical services, hygiene and thus overall living conditions; no child work or other forms of exploitation of children; no discrimination due to the gender, race, religion, etc.; and no armed conflicts, violence or other similar hardship circumstances which would be directly negatively effecting the life and education opportunities of young people. on the other hand it is fair to mention that there are also some specific conditions and/or features which to some extent differ the countries in transition from the developed countries of the eu and make them a relatively special group of countries regarding their education systems as e.g.: a relatively lower level of the overall socioeconomic development which in the terms of the gdp per capita represents only about 20-50% of the average of the eu. due to the above low level of the development, most of the countries in transition have not yet – even after the ten year period achieved their pre-transition level of the year 1989. it is quite evident that there has not been enough budget resources for any significant development of education in general not to mention its prodevelopment orientation, innovations, etc. the curricula at all levels of schools have not yet been fully corresponding to the needs of the challenges of the contemporary modern education methods, techniques, etc. there is an evident lack of modern didactic and computing technology. thus in the whole system of education has been prevailing an extensive system of memorizing instead of applying a modern principle of „learning by doing and doing while learning“ e.g. in relation to the utilization of modern information technologies, foreign languages labs, etc. there is a lack of qualified teachers for some subjects related e.g. to foreign laguages, information technologies, but also to human rights, civic society, etc.; and a lack of development issues in the national curriculum, a lack of alternative education, variability, etc. in the specific conditions of slovakia as a new independent nation, we have – in addition – to take into account that all these common development issues have been further effected by the necessity to solve some specific problems related to the new country i.e. to develop the necessary 10 iassist quarterly summer 2000 institutional framework, to introduce a new national curriculum better corresponding to its new statehood, independent identity, etc. preparation of the national efd program in slovakia in addition to the above general regional context of an economic and social transition as well as some specifics of slovakia as a new country, the whole process of preparation of the national program for the efd has been a rather complex and relatively long-term process . in order to contribute as much as possible to this process, the slovak national committee for unicef has been very active in supporting the rights for education as an integral part of the united nations convention on the rights of child and its implementation and monitoring in slovakia. among various other activities in support of their implementation, the national committee for unicef has – during the ten year period since its adoption – conducted two situation analyses viz. in years 1995 and in 1999. the first one in 1995 was the very first of that kind of analysis of children in the independent slovak republic. the second one was conducted in 1998-1999, and completed in the year of the tenth anniversary of the adoption of the convention. the second analysis in the area of education has shown that in spite of some progress achieved in comparison with the results of the similar analysis in 1995, there are still areas where the progress in the field of education has not been as expected. according to the individual conclusions and recommendations from the previous analysis, the achieved status in 1999 according to the particular situation analysis has been as follows: there still has not been prepared a long-term national strategy for education with the time horizon for next 10-20 years which would reflect the needs for preparation of the young generation for the future membership of slovakia in the eu with the horizon for accession in years 2004-5 or shortly afterwards; accordingly, also the process of transformation of individual levels of education has been proceeding relatively slowely and as a consequence there has been a growing number of unemployed graduates of individual types of schools. among others it is also one of the indicators that their qualification has not fully corresponded to the needs of the current labour market. for example the trend in unemployed graduates of different types of high schools has increased from 11.4%, 17.9%, 7.2% in 1994 to 15.02%, 22.63%, 8.23% in 1997 and has a tendency to grow further. in the case of university graduates it has increased even more dramatically from 3.0% in 1994 to 13.2% in 1997 and it is mostly due to the lack of structural adjustments of education to the needs of the labour market; there has not been achieved any significant progress in bringing the modern information and computing technology into the schools, classrooms, etc. the situation has to some extent even become worse now than before as the available computers are mostly older types unsuitable for the current networking opportunities, do not have parameters for modern software packages, etc. the whole this process has been negatively effected mainly by two reasons: the lack of funding for the procurement of the modern information technology; and -the lack of qualified teachers and instructors for the particular subjects as they find much better financial and career opportunities outside the education system and in particular in the newly arising private sector. the harmonization of the legislation in the area of education with the legislation of the eu has been in progress but to some extent it has been negatively effected by the fact that slovakia has not yet been selected for the direct accession negotiations with the eu. it is possible to expect that the helsinki summit of the eu in december 1999 and a subsequent start of the particular accession negotiations also with slovakia will bring the necessary acceleration also to this important problem area if we realize that pupils of current elementary schools could complete their highschool and/or university education already as citizens of the eu and to find their full utilization on the particular common labour market. in the area of the development of the alternative education only very little has been achieved in comparison with the results of the situation analysis of 1995. the main problem has again been a lack of funding. some progress has been achieved in the extension of the network of schools according to individual sectors and/or type of ownership (state, private, religious) or according to the main language of instruction, etc., but the further development has again been negatively effected by the worsening social and financial situation of the society as many families could hardly afford to pay for the high-school education at a private school, etc. a specific case of an alternative education regarding the minority education has not succeeded due to the objections of some minorities as they did not realized that improving their knowledge of the official language of the country could improve their chances for empolyment on the more and more competitive labour market. iassist quarterly summer 2000 11 in spite of various mainly budgetary problems so far it has been successfully secured that the whole system of education for the young generation has been free and thus the particular right as stipulated in the convention on the rights of child has fully been respected. but there has already been an existing trend to introduce some nominal fees for some kinds of schools, first of all at the university levels what to some extent could mean a kind of discrimination for students from social weaker families. the same situation as in the case of education in modern information technologies has been in the education of foreign languages i.e. the lack of funding for establishing modern foreign languages labs and the lack of qualified lecturers, teachers, instructors, etc. who again are lured by much better conditions in the private sector. thus, a big part of classes has to be carried out by external teachers what to some extent negatively effects the standard of the whole education. one solution could be to invite foreign instructors who are quite interested to come also to slovakia, but unfortunately the particular employment legislation regarding foreigners is so complex and unfavorable to foreign instructors that finally they usually start their assignments in the neighboring countries. in general, the situation in the foreign languages education has not improved but rather deteriorated in comparison with the past as now about one third of pupils of elementary schools have no foreign language classes at all. in the past at least russian has been a foreign language available for every pupil. now, one third of pupils of elementary schools has no such opportunity although otherwise there is formally much more opportunities to choose among six foreign languages (english, french, german, russian, spanish, italian), but again the problem is with the availability of teachers. as we have already mentioned, one of the main problems of education has been a lack of qualified teachers. this problem has been further deteriorating. not only that teachers have been underpaid and thus forced to seek better opportunities for employment in other sectors but even under such unfavorable conditions, there has been another threat to teachers. under the current plans of the government to reduce the state administration by 10%, the same reduction has to apply also to all types of schools. such a reduction would of course effect also teachers and especially those of the older generation i.e. those who in many cases are the only teaching staff as it is not at all attractive for young graduates to become teachers. one of the negative outcomes from the last situation analysis has been the fact that slovakia has still been one of the countries that has not yet introduced a post of an ombudsman for monitoring the rights of childern including those for education. in view of the above problems of the existing system of education in slovakia as revealed by the last situation analysis conducted by the slovak committee for unicef, the following main conclusions and recommendations for the further development of the efd have been formulated, which at the same time are also the main challenges for education in the forthcoming 21st century in general: to prepare and implement as soon as possible a comprehensive national strategy for the education on the principles of the education for all and for development and as a part of the system of the wholethe-life education process; to harmonize the education system according to the standards, rules and regulations in the countries of the eu and thus prepare the whole education system for its place in the unified europe. it concerns not only of the legislation, organization but also of all budget and financial implications, a direct support for research and development in the area of education; to participate actively in all education programs of the eu as e.g. tempus, socrates, leonardo, youth for europe, etc. and create all necessary conditions for the maximal mobility of teachers and students with the countries of the eu; to introduce and develop all kinds of european studies and thus to support the knowledge and true feeling of the common european identity, history, culture, cooperation, etc.; to develop and further promote all forms of education in the areas of human rights, child rights, minority rights, etc. as cornestones of the common european citizenship, free movement of persons, etc.; to maximize education, practical training and utilization of the modern information and communications technologies in all types of schools as there is already now existing an evident handicap in comparison with the situation in the eu. especially, it is necessary to enable young people an unrestricted access to internet and thus participate in various worldwide educational programs such as a voice of youth, etc.; to promote and further develop education in various global issues such as environmental protection, a healthy life style without drugs, smoking, alcoholism, protection against sexually transmitted diseases such as aids, etc. to the same category also 12 iassist quarterly summer 2000 belongs education in the field of international cooperation and understanding and against any forms of intolerance, xenophobia, rasism, stereotypes and prejudices in relation to other cultures, nations, etc.; to prepare the young generation in such a way that every young person will be able to master at least one of the official languages of the eu and in particular english as a language of the current globalization; and to modernize the whole educational system in the direction towards its further differentiation and pluralism, free and individual choice for an educational pattern, a whole-the-life education and an active participation of citizens and especially parents in the education of their children, etc. conclusion in order to actively contribute to the implementation of the above challenging tasks, the slovak committee for unicef has also launched in cooperation and with funding from the regional office for cee/cis and the baltics of the unicef geneva –its cmis – computerized monitoring information system in the area of education as one of the projects commemorating the 10th anniversary of the united nations convention on the rights of child. technically, the cmis project has been designed, developed and implemented at the department of information system of the faculty of management of the comenius university at bratislava. primarly, it has been formulated as a system for monitoring of the rights of minorities for education in the slovak republic, but, as a modern www based system (http://www.fm.uniba.sk/ projekty/cmis,) it is open and available for monitoring of the rights for education in general. in addition to this, its main monitoring oriented development and implementation strategy, it is also, at the same time, a system which is directly serving various efd-related goals: to enable young people a practical use of the modern computing and communication technology in the environment of the contemporary www; to use that technology for monitoring their own rights not only in education but also in general according to the particular convention on the rights of child and thus to contribute to their knowledge in that area; to learn directly about the situation regarding the same rights in other countries and thus to better understand the current globalized and ever more interrelated world and overcome some of existing stereotypes, misunderstandings, etc. to communicate directly with their partners in foreign countries and thus in many cases to acquire practical experiences in establishing and developing their own „foreign relations“, to increase their knowledge on globalization, on the world and foreign countries, foreign cultures, etc.; and to have an opportunity for practical utilization and improvement of their skills in foreign language modern on-line communications, etc. references 1) slovak committee for unicef, situation analysis, children – future of slovakia, unicef bratislava 1995 2) slovak committee for unicef, situation analysis, slovakia and children ´99, unicef slovakia 1999 3) soltes, d.: some specifics of implementation of the efd in slovakia as a country in transition, unicef efd workshop, kathmandu (nepal), april 1998 4) economic and social council: unicef efd, e/icep/1992/l.8, un new york, usa 1992 5) education for all, unicef response to the jomtien challenge, education section, programme division, unicef new york, usa 1992 6) commission of the european communities: preparation of associated countries of the cee for integration into the internal market of the union – white book, ces brussels 1995 7) soltes, d. at al: cmis – computerized monitoring information system (on the rights for education), bratislava 1999 * paper presented at the iassist conference, june 9, 2000, northwestern university, evanston, chigago. dr. dusan soltes, faculty of management, comenius university, bratislava, slovakia iassist quarterly summer 2000 13 la^biiit newsletter, vol. ^, no. i (ijummer lyyttj nttwokks anu hetwohking thomas win. madron western kentucky university absthact tne purpose ol this paper is to provide an overview or computer networks and tneir associated communications systems. in developing that overview we will first define the idea ot a "network", then describe communications systems including commentary on equipment and costs. having aealt with the role ol communications we will then turn to computer networks per se and review the contributions networks can make to computing. the discussion of computer networks will deal with shared hardware resources, shared software, types of systems, and network cont igurations. the general discussion ol networks will be followed by a description of a developing network: the kentucky tducationai computer network ikecntt;. the paper will conclude with a comment on some of the opportunities provided through the use of networks. in the following pages we will attempt to explore some of the characteristics, problems, and iwl kuuuctiun opportunities provided through the establishment and use of computer ihe title of this paper, "netnetworks. because computer networks and networking", may not proworks are dependent on communicavide the reader with the topic as tions networks, several attributes sell-evidently as it might seem. of communications systems--as those the terra "network", as used by systems apply to computers--wlll those of us who use computers, may first be described. ihe core ot refer to communications networks, the papei centers on a discussion computer networks, or both. the of computer networks and their topic of this paper is directed characteristics followed by a dlsprimarily toward computer networks, cussion of various configurations although it will become evident networks might take. ihe more genthat we will have need to speak of eral discussion of networks will be communications networks as well. made concrete by describing one operating network--the kentucky what are computer networks? one educational computer network (keccommoniy used definition oi a comnel j--wnich is still m the process puter network is the following: of growth. unally, we will one or more computers accessed by attempt to bring theory and pracusers via some communications nettise together m a summary discuswork. the computers may consist of sion ol why we should use networks main-site, or nost computers, and what applicatlons--especiaily and/or remote computing systems. in higher educatlon--are possible. generally it may be said that communications networks can exist without computers, but computer networks cannot exist without communications networks. bo iabiji5j new3ieti,er, vol. d. no. i ibumnier ly^tt) lummuhical lunla nhl wukk,b danowidtn. 10 transmit data over a telephone line, whether pudlic or a communications network conprivate, that data must be manipusists of a communications medium iated electronically so that it and computers or other devices used fits into some segment of the frein the control ol the communicaquencies into which a channel may tions process itself. the most be divided. this manipulation is pervasive communications system accomplished by a device called o available to most ot us lor voicemodem. ine speed of the line--tne grade exchange is the telephone rate at which data are transmitsystem. ihe telephone system has, ted--are measured in "bits per sec01 course, been witn us lor a conond", sometimes called baud rates, siderably longer period of time in the period from about ly^u to than have coiiiputers. while the about i^y^ there was over a 4uu telephone is the basis lor the most percent increase in he normal baud common communications network, rate at which data could be tranthere ,itq other such systems linked smitted over conventional telephone by land-lines, microwave or other lines (,cf. martin, lycu, p. y, for radio transmissions, and, most the earlier rates), recently, lasers. because most communications systems predated the need for data transmission, there has been a tendency to provide communications tquipment means oy wnicn digital data can coexist with voice messages. in while we cannot delve very far order to control the communications into tne establishment and control process it nas been necessary to ot communications networks in this provide instruments capable of propaper, it would oe well to explore viding the proper level of service a lew o^ the devices needed in the lor data handling. transmission of data. among those having importance for our purposes in a nor.iial telephone conversaare modems, and their extensions, tion, two or more people speak to multiplexors and concentrators, one another over lines connected by means oi public exchanges. jhese ihere is a relatively narrow lines are called "public" or range of frequencies which may "switched" lines and can be used travel without much distortion over for communicating data as well as telephone circuits. ihe human voice coinmunications. ihe alternavoice has a middle frequency ol tive to public lines--althougn about ibuu hertz. if we translate still provided by telephone compabinary data into a modulated frenie3--are "private" or "leased" quency ol about itiuo hertz, then lines which are connected permandata too can be transmitted without ently or semi-permanently in a data great distortion. a device called communications system. hegardless a modulator is used when sending of the kind of line used, the sigdata in order to achieve the necesnai carrying capacity of communicasary modulation ot the data. at tions lines is usually described in the other end, a demodulator is terms of frequencies those lines used to change the modulated carwiii carry. ihe range of frequenrier back co normal binary data a cies is called the bandwidth ot the computer can use and understand, channel and the quantity ot data normally these units--tne modulator that can be communicated over a and demod ulator--are combined into channel is proportional to tne ia5bii>t newsletter, vol. d, no. i lijummer ly/tt) a single unit with the abbreviated is consumed by a chunk of one ot name ot modem. modems are made by the signals to oe sent, independent vendors as well as by telephone companies. ma bell often while multiplexing equipment is calls modems "data sets." although relatively inexpensive, and while it is outside the scope ot this multiplexing may be a convenient paper, it should be noted that method ot using one high speed line there are at least three dilferent tor a number ot lower speed tasks, types of modulation. moreover, there are some occasions where muldata processing equipment may be tiplexing does not work well. muldirectly wired to the telephone tiplexing is limited by the total line or may be connected via an bandwidth available thus providing acoustic coupler when dial-up is a limit on the number ol channels the normal mode tor accessing a into which a given link may be subtarget computer. divided. une way ot getting around the limitations is to use a "hoidune ot the problems with teleand-torward concentrator" which is pnone lines is that they are relanothing more than a small computer tively expensive. une way ot cutwhich receives messages trom a varting costs, especially when using lety of terminals, stores the data relatively low speed communication in some memory (butters;, then equipment, is to employ multiplextires the signals up the line to ing . ihrough the use ol multiplexanother computer, ihe advantage of ing it is possible to send signals a concentrator is that many ditterirora more than one terminal over ent kinds of entry devices may be the same line at the same time. in attached to it, all running at dita multiplexed system, two or more tering baud rates, and the capacity signals are combined and transmitis limited only by the capabilities ted over a line with an appropriate ot the concentrator itself. unlike bandwidth. ihe terra "multiplexing" multiplexors, however, which are then, means the use of one facility transparent to the end-user, conto transmit in parallel several centrators may cause delays in turditterent channels oi data over a naround and, being more complex, single communications link. are more subject to tailures of various kinds. judgments as to une ot two multiplexing methods which kind of approach to use are normally used: f requency-divishould be governed largely by sion multiplexing or time-aivision cost/benefit considerations in tne multiplexing. t-requency-division overall design ot the network, multiplexing requires a modem or data set that provides several groups of channels over one line, dividing the total bandwidth of the cost considerations line into channels of some smaller frequency range. 5uch a device as martin has noted ciy^u, p. pertorms the normal t unctions of a i;, the "tacts about the transinismodem while carrying out the mulision links which most concern a plexing. by way ol contrast, sigsystems analyst are the cost of the nals may be sent in a round-robin links and the rate at which data fashion so that only one signal may be sent over them." it is occupies the channel or line at any becoming clear that the cost ot one time. ine time available is commuications is an increasingly divided into small slots and each important tactor in the cost ot iaijbibt newsletter, vol. d. no. i cbummer lyyo; computing and as large computer shared hardware kesources networks expand, and as we become more dependent on networks, commucertainly one ol the early motinications costs will expand at a vating tactors deiiind tne estaoisignilicant rate. it nas been ishment ol' computer networks was estimated tliat by lycit) ninety out the fact that tnrough networks many ol every one hundred computing doiusers could use expensive computer lars will be spent for data commusystems, thus spreading the cost ol nications rather than i'or the data such systems over a much larger processing itself ci-erreira and user base. very often the total niiles, lyybj. as a result of the number of users on a large network cost considerations and potential can support equipment on a scale no expenditures in the future, systems individual user or user institution analysts and organizational managcould manage alone. "why both neters will have to become more cogniworks and stand alone computers", zant of the demands for allocating one might ask, "when mini and micro resources for telecommunications. computers are getting to be so a byproduct of the projections inexpensive?" while it is true noted is that manufacturers av% that some micro computers cost litincreasingly turning to the productie more tnan expensive terminals, tion of communications hardware and and while the capabilities of such software (.xotaro, lyyo;. in the computers are great, they remain end, however, the objective is not inappropriate in situations where only to move the data, but to prothe same data base must be accessed cess it. consequently, we should by many different users, where consider the computer network which files are large and must be is supported by the communications accessed rapidly, and/or where user system just described. programs require very large amounts of memory to do the task needed. it may be helpful to look at each one of these points in turn. tumhutkh hetwuhks two of the three points just why should computers--and therenoted deal with capabilities lore end users--be organized into regarding files and file strucnetworks in the lirst place? with tures. although the cost of compucomputer hardware becoming less ters is falling rather rapidly, the expensive and at the same time more cost of large scale peripheral devsophisticated , the answer to the ices--devices which can be used to foregoing question is likely to be permanently store large quantities somewhat diiferent in the latter of data--are not decreasing in cost ly/u's than it was earlier in the m any substantial fasliion. moreuecdde. '.lo j iart^e extent the over, even snould large-scale perireiiidinder of the paper will be pheral devices fall in cost, the devoted to answering the question lower price for hardware does not "why networks?". before going necessarily justify using such devlurther, however, it will be helpices to keep multiple copies of the ful to explore some of the outcomes same data base. if, in a given and characteristics of networks, application, several small computhen turn to a description of the ters can legitimately be used to various configurations networks may replace one large computer, then take. the decision should be made on the basis of the comparative costs of by laiii>i3t newsletter, vol. d. no. i (.summer ly/a; tne two approaches when aii elements ol' the two are consiaered as complete systems. notwithstanding the comment concerning tne cost ol large-scale peripheral devices, we should note that devices such as large disk dr ives--wnich are very costly iteins--are likely to be replaced in the relatively near future py electronic memory systems (such as the bubble memory), thus eradicating the arguments just presented, similarly the question ol the amount ol main memory needed lor user programs will become less problematic as the cost ol' memory is reduced and as micro computers are designed to address ever larger chunks of memory. none ol the solutions or substitutions lor networks are nelptul, however, when one is dealing with a large database which must be used by various users located in diilerent places. shared software hesources while the ware capabi mini, and mic ing blurred , necessary for operation of any size is b portion of t ning such sys taries on t development o have suggeste several years more and mor languages mo natural langu programmer wi standing tha lact, however see signific personnel tim programs whic user to u differences lities among ro computers the personn the program computer sy ecoming a la he total cost terns. variou he direction f computer f d that over , as comput e to be prog re nearly r ages, the ro 11 change, t fact, or , we will co ant contribu e to the pro h will allow se computers in hardlarge , is becomel costs ming and stems of rger proof runs coramenof the acil ities the next ers come rammed in esembl ing le of the notwi thalleged ntinue to tions of vision 01 the end in a reason as a r sidera sidera cerns about the ha ciosel shared sectio provid contin works availa previo ing ex discus comput idea would notion lashio ably s esult tions , tions about the co rdware y rela data n. t es arg ued us can ble to usiy p ample sed in er co of com not b 01 ne n . traig of s the will hard st of l' ted noted ach urn en t e of also use resen ola some mere puter e tea twork htfo uch refo shi ware the his to t in of t s in net ma rs w t . n det ncin ized sibl s of rward perso re , ft f to peop same he p the hese fav works ke hich une dded ail b con wi som manner . nnel concost conrom conconcerns le to run issue is roblem of foregoing concerns or of the netresources were not interestresource , elow, is the very ferencing thout the e form or lypes of systems computer networks may be structured to support remote batch, or timesharing, or some combination of the two. alternative names for batch systems are "remote job entry ikje)" and/or "remote batch terminal ikbi;" systems; the most current buzz-word is "ktsj". liraesnaring systems are often called "interactive" systems. kbl systems are tnose which operate essentially with card input (or its equivalents in the form of floppy disks, online job entry, etc.;. jobs are placed in a job stream, given a priority for the order of execution, and a job in its proper turn is executed and returned to the user. unce the job is submitted it is out of the hands of the user and dependent on the characteristics of the operating system under which the job is being run. in contrast to a batch approach to networking is the use ol timesharing or interactive systems where the user sits at a iaii^ist newsletter^ vol. e' , no. i (summer lyvd) typewriter-styled terminal and the oversimplification ot networks, we computer and tne user interact with will review four characteristic one another in real time. in such conl igurations : point-to-point; systems the user always has control multipoint; centralized; and hierover what is happening to his/her archical, jod and may make immediate responses to any problem which arises. ut ten hyorid systems are developed such as ibm's conversational job pomt-to-hoint entry system ccj3; or utc ' s approach to tukjkan and tlubul on a point-to-point network is, tneir large systems. in doth without doudt, the simplest kind of approaches coding for programs is network for it consists ot a compuentered interactively through a ter, a telephone line, and one tertext-editor, but the program, when rainal at the other end ot the teletinished, is actually submitted as phone line. ihe terminal can be a batch job. both approaches to either an kbt or interactive. this networking have positive and negasimplest of systems is depicted m tive teatures, but what will be figure 1. many networks begin as said m subsequent sections ot this point-to-point systems and gradupaper apply equally to both types ally develop into more complex ot systems. entities. network uont igurations multi-point networks there are a variety of ways in multipoint networks constitute a which networks might be configured straightt orward extension ot ana many (perhaps most; networks point-to-point systems in that are in a constant state of change instead of a single remote station, and growth. if the computer netthere are multiple remote stations, work consists of only a main-site those remote stations may be either or host computer which does all kbt's or interactive or a combinadata processing from one or more tion ot both. ihe remote stations remotes, it is a centralized netmay be connected via independent work. it there are remote compucommunications lines to the computers processing jobs for end-users, ter or may be multiplexed over a as well as a main-site computer single line. such a system is (which is itselt optional;, then we illustrated in l-igure t^ . in either may have the beginnings of a disa point-to-point system or a multitributed network. a distributed point system, the characteristics network can be either centralized ot the remote work stations are a or hierarchical in lorm, but a netfunction of the work to be accomwork which does not involve distriplished at the remote site, buted processing can only be centralized since all data processing is done on a main-site computer. it is possible tor a single commucentralized networks nications system to provide communications services tor two or more as noted above, a centralized concurrently operating centralized network is one in which primary computer networks. although the computing is accomplished at a sincomments which follow constitute an iab:5i5l newsletter, vol. d, no. i tiiummer lyyb) hierarchical networks mgure i: point-to-point network gie site witn al feeding into that a system is thou network with e entering the cen single communicat m mgure f,. bot a star network ar terns in that they a single, cent lypically, howev network may not processing capab star network may ters out at the e cations lines. 1 ter which suppo network might , i into a star netwo 1 remote sta site. often ght of as a ach remote tral system ions line as h a multipoin e centralized are controll ralized comp er , a multi have distri ill ties whi nave other c nd of its com n tact, the c rts a multi tsell , de 1 rk . tions such star site via a noted t and sysed by uter . point buted le a ompum u n 1 ompupoint inked a hierarchical network represents a tully distributed network in which computers feed into computers which in turn feed into computers. the computers used for remote devices may have independent processing capabilities and draw upon the resources at higher or lower levels as intormatic i or other resources are required. such a network is pictured in kigure 4. a hierarchical network is not the only kind of distributed network but it is a completely distributed network. m figure ^: multipoint network iabbibi newsletter, vol. d. no. i ciiummer ly/b) hgure j centralized or 5tar network at se we have distribu two con opposite although can d i thought com put in gested b lusa ny coinputin part of ing of d tne pia nates an taming work." networks simple veral point spoken of ted proces cepts in n ends of such is no str ibuted of simply g. one d y data 1uu ybj states g places the prea ata, and a ces wnere d is used . central con in today' s which star star system in centr sing a etwork some t the proce as dec efinit as r that d "a s nd pos ccess the da . . w trol o comput t as oft this paper alized and s if the s were at continuum , case. nor ssing be entralized ion, sugeported by istr ibuted ubstantiai t-processto data at ta origihile mainf the neting world, relatively en become distributed systems without conscious design on the part of systems analysts. as often as not, a typical batch terminal in a network, rather than being a "dumb" terminal such as ibm's i^ou, will be a mini or micro processor based device providing local tile storage and processing capabilities as well as the capacity to read cards and print paper. he dec with 1 the de er tha ysts . ty of some ce , a ire a but the phone. ision nteli cisio n the if dial iown end micr progr iarg th entra tr ibu be sions 1 sit n whi e wi inati 1 com to acquire a igence may, in n of the end decision of n there is the ing into the speed teletyp user may dec ocoraputer lor am it to corarau er system vi us, the decis lized star n ted processin ven turtner r made by analy e. as an exam ch a network 11 turn next on ot the ke puter network termifact , user etwork possisystem e-iike ide to home nlcate a the ion to etwork g netemoved sts at pie of might to a ntucky ikecimt k.tnrulk.k huucatlunal cumputer k to p vice tion coram the bate capa were late base inte ecnt rovi s to s o onwe orig h a bill not ly d o ract t was establ de academic r the eight p 1 higher edu alth of k.entu inal design c nd interacti ties, intera actually est yt>. the ba n an ibm jy ive system i ished in computer ublic in cation 1 cky. ai ailed fo ve proc ctive se ablished tch syst u/lbt> an based lym serstitun the though r both essing rvices until em is d the on a bi labiiibt newsletter, vol. ^, no. j (summer 197«) on-line i ; on-line: storage \ || storage a higure 4: hierarcnicai or iree network utcsystem 10. the original batch system developed as a centralized star network and had its origins in tne use western and others had already had or the state's system at the bureau of computer services in t-ranktort, the state capital. the establishment of the lull network was conceived ol' as taking place in stages or phases, with the tirst phase connecting tne eight institutions into the computers at louisville and lexington over a common communications system. the original idea behind the first phase is illustrated in figure !? . the second stage was to expand the original communications system to that depicted m higure 0. ihe network currently approximates figure b, although variations in design have taken place. in particular, the node labeled bcs (bureau of computer services, hrankfort} is not now linked into the system at all, and may never be. ihat decision has been a political rather than a technical one. bigure b: kentucky educa computer network, fhas tional e 1 to illustrate the way in which the system has developed we might look more closely at the link between western kentucky university (wku) and the system. western constitutes a communications sublassist newsletter, vol, 2, no. 3 csumraer 197a) figure b: kentucky tducationai computer network, h'nase ii center in the system usin time-aivision multiplexor to ai a ytiuu paud line into lour baud channels. there are two centers on western's campus eac which, until recently, have one of the i?4uu baud channels w go to louisville and are switched over to lexington location of the ibm 3y0/lbb). ray otate university imu located further to the west linked to tne multiplexor at via a £?4uu baud line into a t channel. the fourth channel turttier time-division multipl into eight juu baud channels interactive computing at wester vide hbt h of used hich then (the murau; , is wku nird is exed for n . we cont i direc estab lexm two e upgra intel ibuu the u paduc are no guratio t link lished gton, a vents : de its ligent from an se ot ah comm w sh n as tat bet mov we hbt term ibm the unit ifting just d ybuu ba ween e made stern' s to a inai i. jyttu) vacated y colle slig escr ud; west nece dec high to a and en ge. htly the ibed. a is being ern and ssary by ision to er speed harris to allow annel by tor 1 puting , diverts system dependin sent a " sent tim in dire another , for the of the i computer puting sometime also be tern thus other th lexingto tucky un access t terminal central i into a 1 it wo in bowli microcom use , b automati two larg ing a lo thence state ' s situatio time of technolo enormous nteractive or a modem in the user to eit 1u or the 1 g upon whether 1" or a "2". e the two compu ct communicatio although that future. in ad nstitutions als s for administ and it is pro in the futur hooked into the providing com an those at lo n. already n iversity's ibm he system as a clearly wha zed star networ arge distribute on-line louis her the bm jyu a use at the ters ar n with is pi dition , o have rative jected e these kecntt puting uisvill orthern jyo/n high t was o k is tu d netwo coraville ubx/ibb, r has pree not one anned all local comthat may sysnodes e and kenb can speed nee a rning rk. uld ng u pute e ab call e sy cai bei syst n ha writ gica cos be possible reen to obt r system fo le to prog y access eit stems on k.ec call to the ng linked em. altho s not devel ing of this lly feasibl t to the end for someone ain a small r personal ram it to her of the ntt by makuniversity , into the ugh such a oped at the paper it is e without user . while a considerable amount of time and effort has entered the planning of ktxnb.t changes have been made in the original design. bti iai>6i5t newsletter, vol. d, no. i l^iummer lyyaj 5ome 01 the changes were made necessary as a tunction ot political aecisions inmuencing tne lunditig of the system, some changes were made on the basis of the availability ot hardware tcommunications) at unusually good prices, and some were made as a result of user institutions requiring better or ditferent kinds of services. unce operatio to restr opment . wn 1 c h m industry for pubi mile ot state go example , cost ot vate ent cost dif differen citic a depend in designm system . a networ do with a lar n it ic t t bu ight may ic ag lease vernm is 1 equi v erpr 1 feren t dec pproa g o g a in a kin ity ge is n he r rthe be not enci d t ent ess alen se . tial isio ches n w pub ny e oper network ot alway ange of rmore , good to be equ es. th elephone in kent than one t servic becaus s in pho ns conce might hether lie or vent, on ation, w comes into s possible its develdecisions r private ally good e cost per lines tor ucky, for third the es to prie of these ne service rning spebe made one was a private ce we have hat can we nktwukk ufpuhtunities there are a number of advantages land some disadvantages; in using networks. we will note three opportunities provided by such systems: tconomy ot scale; decentralization ot organization; and shared applications. kach of these will be considered in turn. equipme numbers the cos vide ea ger , mo sors an ind ivid resembl tems . systems same ar the lat tcon not onl if mul the sa necessa copy of for eac number reduce storage base, a in havi gle dat nt of ts ch u re p d pe ual ing alt is gume ter across users . it was ser wi ower fu r ipher cost that hough steadi nt can halt o reia by poss th 1 , ce als , s t on the ly de stil f the tively 1 so sprea ible to ccess to ntral pro yet to somet smaller cost ot s cl ining , 1 be mad lyyus. arge ding prolarceshold hing sysmall the e in omies y in tiple me d ry to the h pot 01 co the , ma nd ot ng mu abase of sea hardwar users atabase mainta data ra ential pies ca cost ot intenan her att itiple wi th m le are e , howev need ac , then in only ther tha user . be re expens ce ot t r ibutes users o ultiple obtained er , for cess to it is a single n copies when the duced we ive disk he datainherent 1 a sincopies. ihe key cons mg hardware be and communicat more important of communciatio cost of provid on the same s without incurri munications. a these cost dif become so grea end of networ while the argum works based on are still impor may come a time scale argument iderat comes ions is wh ns is ing lo cale ng the t leas ferent t as t ks . ent in an ec tant a when no ion ion as computless expensive costs become ether the cost exceeding the cai computing as a network cost of comt in the lyyus lais have not o require tne consequently , tavor of netonomy ot scale nd true , there the economy ot ger has merit. economy of scale as has already been noted, one 01 the earliest reasons tor networks was the ability to spread the costs of large scale computing decentralization of organizations if, indeed, economies of scale can be achieved through networking, then as large-scale organizations tend toward decentralization sucn labiiibt newsletter, vol. d, no. i csuraraer 197b) econoioies become ever more important. an apparent contemporary tact of lite in large-scale organizations today is a tendency toward at least geographical decentralization and possibly administrative decentralization. ket many large organizations continue to require access to common data, program libraries, and ttie like. as a result, witnin a large organization a computer network can provide tnose centralized services necessary, and at the same time help maintain coherence and standardization within a largely decentralized system. fcach individual component of the organization may be doing local computing, as well as remote computing on a network. particularly in higher education this argument nas merit. at least in kentucky not one of the eight state institutions of higher education would have, or could individually atford , computers of the size and performance of those m kecnet. applications bnared hesources including personnel resources, it should then become possible for remote sites to devote relatively more time and attention to user problems rather than to the problems of maintaining hardware and keeping it up and going. computerized conferencing and other opportunities unce a network has been established certain opportunities are provided which were not previously available. une such opportunity is computerized cont erencing . computerized conferencing is a process by which "groups may communicate about complex problems through members' terminals instead of telephones" cmckendree, 197tt) or in person. because of the computer's capacity for maintaining ongoing proceedings, the conference may extend for days or weeks or longer. there are several advantages wnich might accrue trom such conlerencing: 1 . participants are treed trom constraints of time. noted numerous times above is the tact that with networks come shared data, software, and hardware. within the context of a computer system there is another very major resource constituted by the skills ot the people associatied with the system. when working within the context of a network it is no longer necd*ssary for every local unit to provide staff with all possible skills, since each remote location nas the possibility of making use of the skills of those at other remote sites or at the central site if one exists. as a consequence ot the possibility ot making full use of all the resources available in the system. ^. inlormation related to the discussion is automatically processed and stored . i. text, statistics, and/or cases are printed in easily used anc' standardized formats and are available as needed. 4. periodic voting on proposals raised in the discussion may be accomplished with speed. b. conterees may be coupled to various modeling, simulation, or other comb/ iabbibi newsletter, vol. ^, no. i (summer lyya; putationai and planning tions tor communications proresources, such as uelphi cessing. uatamallun ^^ (june method studies. 1976), iiy-ipo. kelly, neil, et al . data communicomputerized conterencing would not cations, kart d. inkusiblt.mb de possible witnout networks. e^ (march lyro, i":>-^<^ . inrough the use ol' conferencing it kelly, neil, et al . information is possible to reduce travel time networks: a management update. and costs yet provide the basis tor iniuskslkmii <r'm (july ly/n, more thougnttul discussions among 4l-'4'4, lui. individuals. while networks are lusa, john m. et al . distributed costly they still provide services processing: alive and well. dillicuit to duplicate at remote infuslfstkms 2i (november 1970;, sites. jt)-^ 1 . lusa, john m. data communications, htferenchs part 1. infosystems 2^ (february 1977;, jy-4b. bump, kobert t. ijtarting up with martin, james. tflepkocfssin(^ netdatran, datamallun i?1 (april wukk uhlianlzatlon . (englewood 197b;, t)0. cliffs, n.j.: prentice-hall, dick, (jeorge m. the lowly modem. 197(j;. datamation di (march 197^;, mckendree, john d. computerized oy-y:s. conferencing. data manadement 'it\q communications channel: lb (january 197^;, 10b-11u. it's broken--now what? datamasanders, kay w. , and cerf, vinton tlun di (october "^^il), c. compatability or chaos in llj-l^b. communications. datamation dd digital. inthoduction to minicom(march 197b), bu-bb. puteh netwokks (maynard, mass.: tenkhoft , p. a. and collard, j. c. digital tquipment corporation, ihe common carriers' uncommon 197'*;. offerings. daiamation d\ ferreira, joseph and nilles, jack (april 197b;, ^o-'^l. m. five-year planning for data totaro, j. burt. communications communications. datamation dd processor survey. datamation (october 197b), bl-b7. ^^ (may 197b;, lbl-170. gray, james p., and blair, clark h. turoff, murray, and hiltz, starr ibm's systems network architeckoxanne. meeting through your ture. datamation d\ (april computer. specthum 14 (may 197b;, bl-bb, 1977;, ba-b4. hohri, william. mcl's microware woods, larry d. distributed proservice. uatamation d\ (april cessing in manul actur ing . i9^b;, 40-49. datamation di (october 1970, hirsch, phil. problems and predicbo-63. b« iassvol201 13winter 1996 comments on the data access and dissemination system by lisa j. neidert1, data archive, population studies center university of michigan the data access and dissemination system (dads) will be the vehicle for dissemination of census data in the year 2000. information distributed in published census volumes in 1990 will be accessed from the internet for future censuses. complete data files will no longer be written to media for redistribution to users. instead, users will access dads and pull off the tables they need. the advantages to the u.s. census bureau and its customers are quicker turnaround for release of files, cost-effectiveness, and increased access. several factors have probably motivated the census bureau to make this move. first, all federal agencies are responding to al gore’s call for internet access by january 1, 1996. second, changes in technology, such as the development of the internet, high speed computers, and low-cost storage have made this method of distribution feasible. in the past year, we have witnessed an explosion in the number of websites that distribute data (e.g. psid, hrs/ahead, national longitudinal surveys, ipums, milwaukee parental choice program, wisconsin longitudinal study, russian longitudinal monitoring survey, world fertility surveys, malaysia family life survey, survey of families and households.) finally, distributing information via dads is a cheaper alternative for the census bureau, particularly when compared with the cost of printing. as with any change, however, there are probably people who were better served by the old dissemination methods than they will be by dads. it is clear that the census bureau wants this system to serve all users. however, there are some shortcomings that should be solved in upcoming renditions of dads if the census bureau is to reach that goal. the data access and dissemination system doesn’t exist yet. it is still a concept. however, i will use the “data access” page on the census bureau’s web site provides a good working model of dads; and much of it is likely to be incorporated into the future operational dads. it is also likely that many features of this current system will be remodeled, so some of my comments may be “old news” to the inner circle of dads developers. the current configuration of dads needs three important improvements. first, dads should provide the same information that one could get using the old dissemination methods. the data may be in a different form than they were in the past, but the content of the data must not be compromised. when the psid changed from a family/individual file with a reocrd length approaching 32,767 to its new form of family records and individual records, the same information was still available. it takes new knowledge to work with the data, but users can still create the same sorts of tables they could create in the past. in contrast, the current configuration of dads does not allow users to create all the tables they could in the past. second, dads should accommodate the users who have access to high-speed computers and large amounts of disk space. some users would like to have the data on their own systems, rather than requesting tables and extracts from dads. the ftp access of raw files is weak in the data access site. if ftp access to raw files proves to be impossible because of the need to protect respondent confidentiality, then the extraction system should be improved. finally, and perhaps foremost in the minds of data librarians and archivists, there is the question of whether dads will meet the archival needs of future users of census data. expertise on and access to state and federal records tends to be fairly short-lived. thus, it is essential that the archival needs of future researcheers be considered in the development of dads. what sorts of records will be turned over to the national archives and in what form? loss of information with dads most users of summary tape files (stf) find the summary-level sequence charts confusing. however,the current configuration of the data access system does not make it clear that all the choices available in a typical summary tape file are available in the new system. users can get tabulations for states, counties, metropolitan statistical areas (msas), tracts, and blocks—the most typical choices. but can they get tabulations for central cities of msas (summary level 340) or for any of the american indian reservation categorizations (summary levels 210-221)? what about county-specific zip code statistics (summary level 820 versus summary level 810)? the way the data access system is currently configured some items in the geographic identification section are not accesible. occasionally users need the longitude and latitude or land area of census tracts or blocks for the computation of a summary 14 iassist quarterly measure such as a residential segregation index. however, users can’t select these items, or other items such as consolidated city population size code, place class code, or place description code from this section. dads works best when users are making a request for a small number of tables for a small number of geographic units. for example, this system works well for a user who wants to know the population size (a single cell) for all counties in michigan or the distribution of income in houston, texas. however, many analysts need perhaps 10 or 12 tables (which might mean 200 or even 1,000 cells) for all zip codes in the nation. to get these data from dads in its current configuration, one must list all the relevant zip code(s). typing in over 10,000 zip codes is not a very practical alternative! in a typical stf request based on data stored on-line, one would select the appropriate summary level for zip code (820) and would get all the tables for all zip codes with the execution of one job. the configuration for block groups and tracts is similar, but one only has to highlight the tracts or block groups rather than typing them out. dads can handle these requests for summary statistics for all zip codes or census tracks in the u.s., but it is a very labor-intensive task for the requester. the analyst who makes this sort of request is not just mining data that are never analyzed. the analyst is reducing 31,000 columns of information into 200-1,000 columns and then making the request for a unit of analysis that might range from around 3,000 for counties to more than 100,000 for block groups. i’m certain that the future dads will allow access to all summary tape file information. however, so far, only summary tape files 1 and 3 were released in cd-rom form. thus, none of the race-specific tabulations from stf2 and stf4 are available under the current data access system. one would hope that dads would allow the census bureau to eliminate the distinction between summary tape files and public use microdata files. it would be very useful for researchers to be able to define their own tables rather than being restricted to the limited number of tables supplied by the census bureau in its summary tape files. we had some researchers at michigan recently wanted to look at disability according to race and sex. however, our researchers needed an age breakdown other than the typical 18-64 and 65+ groupings. because the table they needed was not available in an stf file, they created one using pums files. using pums, however, gave them a smaller case base, and the census geography could not be perfectly duplicated. in general, if analysts make a table or summary statistic based on a small number of variables, they should be able to get the tabulation for any level of geography. however, if they want to use 15 variables to define a summary statist!ic, then the level of geography becomes much more restricted (state, msa, or puma). thus, another advantage of eliminating the distinction between summary tape files and public use microdata files would be the increased sample size for the public use microdata files. small populations such as male clerical workers, female pilots, 50 to 54 yearold women with own children under 5 years of age, or persons born in guam could be studied better with a 16% count rather than the 5% files available with public use microdata. (on the other hand, i shudder to think that some of our researchers would be tempted to try to swallow the 16% count of white prime-age males when even the 5% count proves to be fairly cumbersome.) ftp access and/or improving the extraction engine researchers who have excellent computing facilities may not want to get in the dads queue every time they want access to census data. if the demand for the system is great, the census bureau may want to allow for ftp access so that users who have the capacity can bypass dads except for quick exploration and for ftp access to the original raw files. if for reasons of confidentiality, the census bureau cannot provide access to the raw files—perhaps because confidentiality is built into the dads software rather than into the data (via sample size or census geography)—then the dads system needs to make improvements to the existing extraction engine. the systems developed by ciesin (ulysses) and public data queries (explore) are extremely fast. one of the reasons they are so fast is that they produce is tables instead of the cases and variables used to produce the tables (or summary statistics). any time one writes out individual records rather than tables or summary statistics, response time slows precipitously. if many users want micro-level extracts, as opposed to exploratory tables or even output from a summary tape file request (which is always a good example of data reduction), the response time will begin to discourage and irritate users. if a user has the capacity to handle the raw files, the census bureau should allow the user to do so, and thereby free up time for users who need the cpu. the creators of the integrated public use microdata samples (ipums) have found that their data cannot be used by all who might be interested in them, partly because of the sheer size of the files (125g) and partly because their primary audience (historians) traditionally has had limited access to powerful workstations. thus, the ipums creators developed an extraction system that allows a somewhat disenfranchised user to create a work file. (these users are not completely disenfranchised as they do have access to the internet.) however, response time will not be quick with the ipums data extraction system 15winter 1996 because, at least for the short run, all extracts will be executed on a single sparc20. clearly, a user with access to large disk storage and a powerful processor will be better off running the extract at his/her own desktop. of course, the calculus needed to figure out whether the extract should be executed by the ipums workstation or a local workstation is complicated by the fact that other products, such as an extract codebook and spss cards are created along with the ipums-based extract. researchers at my site, the population studies center at the university of michigan, have made countless extracts since the release of the 1990 pums files using an in-house program that rectangularizes the hierachical structure of these files. turnaround time is relatively quick (45 minutes 3 hours) depending on the sample being used (1%, 3%, 5%, 8%), number of states requested, the size of the file being written out, and the load on the system. we purchased most of the microdata from the census bureau for $4,800 (5% $4,000 and 1% $800). i don’t have a count of the number of extracts performed over the past few years but conservatively it has been 500 which works out to be $10 an extract and is more likely to be over a 1000. dads will not be able to provide this quick and cost-effective system for our users although many users will be ecstatic about the system that dads will provide. improving the extraction engine researchers often need access to more information than the tabular data provided through the stf data extraction system. the need for exploratory analysis can be met with tabular data and summary statistics; and more time spent exploring data before analysis often means less information actually ends up being extracted because the user has a much better idea of what is needed for the actual analysis. however, researchers often need access to microdata so that they can estimate equations. the systems developed by ciesin (ulysses) and public data queries (explore) allow researchers and policy analysts to get means and crosstabs from pums data in a matter of seconds; but not all statistical needs can be met with simple means and crosstabulations. the census bureau is aware of all of this and provides a data extraction engine for microdata. however, in the case of hierarchical files similar to census microdata (cps), the interface for extraction is quite awkward. the interface needs to be improved, particularly, if for reasons of confidentiality, access to microdata is limited to the census bureau extraction engine. currently, the extraction procedure requires users to extract the records separately by record type (household, family, person) even though almost all users want a rectangular product. although, the user certainly can merge the household, family, and person records to create a rectangular file there does not seem to be a rationale for adding this extra step to the procedure. in addtion, merging across record types increases the possibility that a novice user will end up with an erroneous file and not realize it. novice users would also benefit from features such as variables that provide counts across the household, (e.g. the number of children under 4 or the number of earners) and the option to rectangularize the record based on something other than record type (e.g., rectangularize by household relationship for husband/wife or a mother/child file). another common request is to select all person records if any person in a household meets a certain criteria, such as foreign birth, unemployment, age 60+, or interstate migration. of course, the more bells and whistles that are added to the extraction engine, the more likely it is that people will use it for data management rather than just data access. archival issues the final question that a system like dads invokes is how it can be archived. how will the census bureau unpack dads so that they can turn over raw data and a codebook to the national archives? will there even be a codebook if the census bureau intends not to disseminate raw files and technical documentation as it did in the past? if the confidentiality firewall is built into the software, how can this information become part of the raw data so that confidentiality requirements continue to be fulfilled? much of dads sounds dynamic, which suggests that the system will be updated to include more data and perhaps that variables will be recoded to meet the demands of users. at what point will dads be stabilized so that there is an archival record? table 1 has a list of questions that can help provoke our thinking and serve as guidelines in determining who should be responsible for making an archive out of dads. given the complexity and enormity of dads, we may be tempted to allow the census bureau to be the archive for the census of 2000 and for all future data products, particularly because the census bureau looks like the archival expert when compared to the national archives on many of these questions. however, it is important to remember that most state and federal data producers have poor long-term memories about old data (sometimes the definition of old is just a few years) and that there has been a lack of institutional memory within the census bureau about previous data losses due to poor archival policy. an article by dollar nicely summarizes the historical record of federal data producers and the national archives. in the decision on whether to archive summary statistics versus microdata from the 1940 census the reasoning was “if the government agency that created the records for statistical purposes did not fully exploit them, it is hardly likely that anyone else will.” (dollar, 198x: 79). thus, 1940 microdata were expendable. 16 iassist quarterly references dollar, charles. 1979. “machine-readable records of the federal government the national archives.”, archivists and machine-readable records: proceedings of the conference on archival management of machine-readable records. edited by carolyn g. geda, erik w. austin, and francis x. blouin, jr. hedstrom, margaret. 1991. “archives: to be or not to be: a commentary.” archives and museum informatics, technical report, number 13. 1.paper presented at the annual meetings of iassist, may 15, 1996, minneapolis, minnesota. table 1 who should be responsible for data (1)is there expertise in the creating agency that can explain the context, technicalities of the subject area, or the idiosyncrasies of the data which would not be available if the records were transferred to an archive? will that expertise remain available for all electronic records, or only for those in active systems. (2)what functionality of the system used to create the records is necessary to meet the needs of archival users? can the archives provide the necessary degree of functionality, or is the creating agency the only economically or technologically feasible place to preserve the data in a usable format? (3)will the creating agency guarantee equitable access within freedom of information and confidentiality policy guidelines? (4)do the records have continuing value to the creating agency so that it has an interest in and need to maintain the records beyond an external requirement? (5)will there be a duplication of effort if the archives acquire electronic records that have continuing value to the originating office? (6)where will the risk of loss or destruction be minimized? (7)can the creating agency guarantee that it will stabilize and not alter the archival record? (8)do regulations prohibit transfer of records from the custody of the original agency? (9)what is the total cost to the organization to maintain electronic records for accountability and research purposes? how can these costs be reduced for the institution as a whole, without eliminating services to users? source: hedstrom, margaret. 1991. “archives: to be or not to be: a commentary.” archives and museum informatics, technical report, number 13. vol25s.1 10 iassist quarterly spring 2001 the virtual training suite: internet skills for teaching and learning by heather dawson* introduction. this paper describes the development of the rdn virtual training suite which aims to support lecturers, students and researchers in finding and using resources on the internet. it will provide an overview of the aims of the virtual training suite and its content, focusing specifically on the social science related tutorials, which are available through sosig (the social science information gateway). it then gives some evidence of the way in which the tutorials are currently being used drawing upon the results of a recent evaluation study conducted by the university of bristol and practical examples of its incorporation into the teaching and learning experience at the london school of economics. what is the rdn virtual training suite? the virtual training suite1 is a series of 40 free web based internet tutorials which have been funded by the joint information systems committee (jisc)2, under their distributed national electronic resource (dner) programme, on behalf of higher education funding councils of england, scotland and wales3. the tutorials were created by the institute for learning and research technology, university of bristol4 with input from 30 universities, museums and research organisations across the uk. the first phase of the project was completed in july 2000 with the launch of the first 11 tutorials which included titles covering the physical sciences, social sciences and the humanities.5 a further 29 were launched in may 20016. again these encompassed subjects from a range of disciplines, including engineering, statistics, government and social welfare. the tutorials aim to cover a wide range of the academic subjects taught in uk universities and colleges. they also support the work of the resource discovery network (rdn)7. this is a national internet search service, which is being created, for academics and researchers based in british higher and further education institutions. it is composed of a co-operative network with a central organisation called the resource discovery network centre and a number of subject based independent service providers called ‘hubs’. there are currently 5 hubs. each has responsibility for selecting, cataloguing and classifying internet resources in a particular subject field. biome – health and life sciences; eevl – engineering, mathematics and computing; humbul – humanities; psigate – physical sciences and sosig – social sciences, business and law. what are the aims of the virtual training suite? the virtual training suite aims to offer a subject based introduction to locating and using internet resources. each tutorial addresses the needs of a particular subject community; taking them directly to the most important internet sites and offering subject specific guidance on what they need to know. it provides a flexible learning experience. all material is web based and can be accessed from anywhere at any time of the day or night! therefore users can directly control the pace at which they learn. the tutorials are designed for use by both academic staff and students. they include specialist sections for both these categories of user in which guidance is geared to their particular needs. the student sections include information on finding materials for essays and citation style guides. the lecturer resource section contains tips on tracing online course materials and syllabi; case studies of how teachers have used the internet and the facility to download the tutorial or print off posters for use in handouts. the tutorials aim to provide a structured learning environment with clearly defined learning objectives. after completion of the appropriate tutorial, users should be: • aware of the range of types of material that can be found on the internet and how these can be used to support their work • be able to identify the key internet sites and resources for their subject area; know how to use effectively the main tools and techniques for internet searching. • be able to critically evaluate the internet resources that they find. • the tutorials also seek to promote awareness and iassist quarterly spring 2001 11 effective use of the information gateways, which are being created by the rdn. how was the virtual training suite created? the virtual training suite was created by a consortium of authors from uk university libraries, academic departments, research bodies and organisations led by the institute of learning and research technology, university of bristol. contributors included: the national institute for social work (internet social worker); university of manchester (internet anthropologist ), edinburgh university data library (internet for social statistics) and the data archive, university of essex (internet for social research methods. this collaboration was able to draw upon the existing expertise of subject specialists. they were able to identify key issues of concern for their subject community and gear examples towards them. for example the internet medic identifies out of date material as a problem with many health related internet sites and provides guidance to users on how they might check the currency. close contact with the academic community also enabled testing and provision of qualitative feedback on the structure and content of individual tutorials. what do the tutorials contain? all the tutorials were constructed using calnet software8 developed in house by the institute for learning and research technology at the university of bristol. they share a common framework of 4 sections, which clearly structure the learning experience. users can choose to follow the whole tutorial sequentially or select the section most suited to their individual needs. • tour – the first section entitled tour provides a guided tour of key internet resources for the subject area. this highlights the range of materials available and directs users to the most important sites. for example, the internet for lawyers includes references to primary legal materials (legislation, treaties, law reports and judgements); secondary resources (journal articles, case commentaries and textbooks); finding tools (indexes to legislation, library catalogues and directories); organisational homepages (professional bodies, law forms, government departments and law departments in universities); statistical data and teaching materials (lecture notes and syllabi). provision has been made for updating these entries to take account of changes in urls and other alterations in content. • discover – the second discover section introduces the user to techniques for effective internet searching. it includes a comparison of the strengths and weaknesses of commercial search engines and information gateways with guidance on when it is most appropriate to use each. this section also has a particular aim in helping students to use the rdn hubs effectively. for instance the internet politician provides examples of how to use the advanced search form on sosig to truncate search terms and restrict searches to particular resource types. • review the third section is entitled review. this teaches skills for critically evaluating the quality of internet sites. this is an area of particular importance as the lack of quality control on the internet means that users must be careful to assess the value and authenticity of sources before they use them in their work. the tutorials use case study scenarios to highlight common pitfalls and offer tips on how to begin to assess quality. the process is broken down into a series of simple questions which students can ask relating to who? where? and why? the resource was placed on the internet. you might want to take a look at the internet sociologist 9 for an entertaining and effective example. this uses a site called ‘kill the television’ to discuss issues relating to bias on the internet. it provides guidance to users on how they might check the authenticity of an author and look for more information on his/her motives in placing a resource on the internet. • reflect – the final reflect section summarises the skills taught in the tutorial and provides case studies of how students, researchers and lecturers might incorporate usage of the internet into their working practices. a good example is provided in the internet aviator 10where the undergraduate scenario tells the story of ryan ayre who is looking for material about the military usages of unmanned aircraft for his essay. it takes him through the stages of research, showing how he can use the internet to find video footage, technical reports and relevant news items on the internet. the tone is light hearted but it does communicate useful lessons about the range of material that can be found on the internet and guidance on items which are currently available in paper only. • links basket another notable feature of the tutorials is the links basket. as they move through the tutorials users can collect useful urls in their ‘shopping basket’. at the end they can pick these up and save them as a series of useable bookmarks. • quizzes and practical exercises also lighten the learning experience and offer students the chance to test their understanding of the material. each section contains a selection of optional multiple choice, fill the gaps or more open ended questions. • additional features – these are geared towards student needs and include a glossary of commonly used internet terms and guides to citing internet resources in essays. • resources for teachers the second phase of the tutorials introduced a supporting “resources for 12 iassist quarterly spring 2001 teachers” section. this includes an introductory powerpoint presentation, student workbook and handout and lesson plans. teachers can use the “print/ download” option to print out or download the whole tutorial or specific chapters within it. these can then be used as slides or handouts. free posters may also be printed from the site. how is the virtual training suite being used? the first independent academic evaluation of the usage and value of the virtual training suite was completed by lin amber of the university of bristol information services in march 2001. 11 this provides both quantitative and qualitative data on initial usage. in terms of qualitative statistics, initial usage has been encouragingly high. from march-december 2000 there were over 43,000 log ins to the site, with an average of 204 sessions per day. interest in the future development of the project remained high as over 2,000 people signed up to receive notification of the launch of the second phase of tutorials. this means that usage is likely to rise further in the future. one of the most popular options was to download the tutorial (739 sessions recorded) and to print off posters (1473 requests). these trends were noted by the project organisers and as a result these features were incorporated into the second phase of the project and advertised more widely. analysis was also made from a total of 122 online feedback forms completed between december 2000 and march 2001. although the number of respondents only represents a small percentage of the total number of users, it does provide some interesting information on the type of users of the tutorials and their opinions of the content. 25% of the users were librarians seeking material for inclusion in user education sessions, 14% lecturers, 11% undergraduates, 8% researchers, 10% post graduates. the rest fell into other categories such as school students. 56% of users classified themselves as independent learners. this shows that the tutorials are reaching a wide audience, representing all the categories of user for which they were intended. the majority of the respondents felt that they had learnt something from using the materials. they were seen to be particularly useful starting point for novices. the most common problems experienced related to functionality. these were often local problems linked to the type of browser used and response time. opinions were also divided on tutorial length. as a result of this feedback the authors of the second stage tutorials were encouraged to adopt more concise styles of writing. technical features needed to support the quizzes were revisited and streamlined. feedback remains an ongoing process. users are encouraged to complete on-line forms with their comments and these have now been highlighted more prominently to encourage greater response. why is the virtual training suite being used at the london school of economics? the paper will now conclude by offering some qualitative examples of why the tutorials are currently being used by library staff at the london school of economics (lse). the lse is a renowned teaching and research institution for the social sciences with a student body of over 6,000, of whom 93% are studying for post-graduate degrees. there are currently over 800 part time students. usage of the tutorials was particularly attractive to us for a number of reasons. • the wide subject coverage offered. the 40 packages match a large number of the areas, which are taught and researched by school staff and students. subject based material was regarded by the users as more relevant than general courses as there was a feeling that it saved time by taking them directly to what they needed to know and avoided ‘irrelevant it terminology’. the virtual training suite also covers a number of subject areas, such as medicine, which have a marginal interest to lse researchers. the paper based library collections do not cover these areas comprehensively so information about the most important internet resources was welcome. • the flexibility of the packages. free internet access made it possible for part time students and students on fieldwork placements who could not attend timetabled library training sessions to be offered some training. it also enabled large numbers of students to receive training, avoiding the problems of restrictions on numbers in training rooms. • the downloadable materials and posters. these provided a readily-usable resource for preparing user education materials which was welcomed by library staff who were under pressure to provide more training for growing numbers of students. how is the virtual training suite being used by the london school of economics library? the range and flexibility of the package means that it has been possible to actively incorporate it into the learning programme in a number of ways: • large scale student inductions. at the start of each academic year subject liaison librarians give presentations for new lecturers and students in particular academic departments. these are large in scale and often take place in lecture theatres seating over 100 . they are intended as overviews to the services offered. slides have been taken from the tutorials and incorporated into powerpoint presentations to demonstrate the type of training courses the library offers. iassist quarterly spring 2001 13 • hands on workshops. liaison librarians regularly offer small workshops for research and masters students from their departments. these sessions provide hands on teaching for 10-15 individuals at a time. the tutorials have been incorporated into some of these as a useful resource for helping novice students to find out what is available on the internet for their subject area and improving the efficiency of their internet searching skills. • drop in sessions. in addition to targeted workshops. the library also offers regular drop in internet training sessions of around 90 minutes throughout term time. these are open to all staff and students without booking. the trainer usually demonstrates search techniques and important sites and then the students are able to engage in hands on exercises. as part of these sessions students are encouraged to work through tutorials as well as notes and materials prepared by more general internet courses such as the tonic netskills package. 12 trainers have found that while students are working independently at their own pace they have been able to offer a supportive atmosphere, which encourages learning. • training sessions for library staff. all new information desk staff at the lse receive training sessions on internet skills so that they can assist users in finding information. there is also a regular ongoing programme of refresher training sessions to remind established staff of new services. the tutorials have been used effectively in these. staff have been encouraged to explore the tutorials and to note key sites in the tour sections as these often contain specialist subject directories or gateways which are of value in providing starting points for research. they are particularly useful in directing staff who do not have a wide experience of the subject area. • inclusion in handouts. the library produces a number of guides for staff and students on electronic services. an example of these are quick reference guides for particular subjects which are intended to offer an overview of starting points for research in the subject area. they contain a list of the main classmarks for the subject area, lists of locations of important journals and information on key electronic resources such as cdroms, databases and internet sites. the virtual training suite tutorials have been listed in the appropriate guides as a good starting point for beginning internet research. • links on the library web pages. as the virtual training suite is intended as an independent learning package, links to it have been placed on the library training web pages. this has encouraged us to begin a process of redeveloping the information that we offer there in order to provide more full-text, self-contained training materials. the library is involved in a collaborative project with the teaching and learning technology project based in the academic section of the lse. this aims to develop online materials to cater more effectively for distance and part time learners who cannot attend formal training sessions. the first stage of the project has been the development of a package to train users how to search the library catalogue effectively.13 we also aim to develop more information on local fee based internet services. this will supplement the information contained within the virtual training suite. future developments and maintenance the second phase of the tutorials was launched on may 8th by michael wills, parliamentary under-secretary of state for education and technology via a live link up between simultaneous launch events held at the six universities of the rdn hubs. already users are asking if the suite will be expanded. they have pointed out gaps in subject coverage, such as art and design, music and archaeology and asked for courses in these areas. the project team has petitioned the funders to see if more money can be made available to expand coverage. they are currently awaiting a response from them. in the meantime any suggestions for new subjects to be covered are welcome, although no promises can be made at this stage. decisions are likely to be based on the response by users to the resource now that the full compliment of tutorials has gone live. the other main query about the virtual training suite is how it will be kept up to date. this is clearly an important issue where internet resources are concerned. as a result the resource has been firmly embedded into the organisational structure of its parent service. the tutorials are all hosted by the relevant rdn hub, so for example all the social science tutorials are hosted by sosig (the social science information gateway). the hubs will be responsible for the upkeep of their set of tutorials by maintaining regular link checking procedures. tutorial authors may also suggest edits to take account of new developments. ultimately we hope that word will spread about this new educational resource, and that it will enhance people’s experience of using the internet as a source of information, raising the skills level of academics and students nationwide. footnotes 1 the virtual training suite is located at: http:// www.vts.rdn.ac.uk/ 2 further information on jisc is located at: http:// www.jisc.ac.uk/ 3 hefce homepage is located at http://www.hefce.ac.uk/ http://www.vts.rdn.ac.uk/ http://www.vts.rdn.ac.uk/ http://www.jisc.ac.uk/ http://www.jisc.ac.uk/ http://www.hefce.ac.uk/ 14 iassist quarterly spring 2001 4 further information can be found at: http:// www.ilrt.bris.ac.uk/ 5 the first eleven tutorials were: internet medic, internet lawyer, internet politician, internet psychologist, internet social worker, internet sociologist, internet economist, internet manager, internet for history, internet for english, internet aviator. 6 social sciences internet anthropologist internet for development internet for education internet geographer internet for government internet for social policy internet for social research methods internet for social statistics internet for women’s studies reference internet instructor humanities internet for history and philosophy of science internet for modern languages internet philosopher internet for religious studies internet theologian health and life sciences internet for agriculture, food and forestry internet bioresearcher internet for nature internet vet physical sciences internet chemist internet earth scientist internet physicist engineering and mathematics internet civil engineer internet electrical engineer internet for health and safety internet materials engineer internet mathematician internet mechanical engineer internet offshore engineer 7 the rdn homepages are located at: http:// www.rdn.ac.uk/ 8 further information on this can be obtained from: http:// www.webecon.bris.ac.uk/calnet/ 9 see http://www.sosig.ac.uk/vts/sociologist/index.htm 10 see http://www.eevl.ac.uk/vts/aviator/index.htm 11 the full text can be viewed online from the virtual training suite pages http://www.vts.rdn.ac.uk/ 12 for further information on tonic see http:// www.netskills.ac.uk/tonicng/cgi/sesame?tng 13 see http://www.library.lse.ac.uk/infoskills/catalogue/ * paper presented at the iassist/ifdo conference 2001, amsterdam. heather dawson, british library of political and economic science, h.dawson@lse.ac.uk. with the assistance of debra hiom, institute for learning and research technology, university of bristol, d.hiom@bristol.ac.uk and emma place, institute for learning and research technology, university of bristol, emma.place@bristol.ac.uk http://www.ilrt.bris.ac.uk/ http://www.ilrt.bris.ac.uk/ http://www.rdn.ac.uk/ http://www.rdn.ac.uk/ http://www.webecon.bris.ac.uk/calnet/ http://www.webecon.bris.ac.uk/calnet/ http://www.sosig.ac.uk/vts/sociologist/index.htm http://www.eevl.ac.uk/vts/aviator/index.htm http://www.vts.rdn.ac.uk/ http://www.netskills.ac.uk/tonicng/cgi/sesame?tng http://www.netskills.ac.uk/tonicng/cgi/sesame?tng http://www.library.lse.ac.uk/infoskills/catalogue/ mailto:h.dawson@lse.ac.uk mailto:d.hiom@bristol.ac.uk mailto:emma.place@bristol.ac.uk spett 1/39 bonifacio, flavio (2018) differences in data-sharing attitudes and behaviours, iassist quarterly 42 (3), pp. 1-40. doi: https://doi.org/10.29173/iq912 differences in data-sharing attitudes and behaviours flavio bonifacio1 abstract this article reports the results of a survey conducted between 18th november and 18th december 2017 about different aspects of data sharing. after a short description of the data gathering task, the report describes the sample, the univariate distribution of the most important variables related to the work of data archiving and the attitudes concerning the data sharing activity: problems encountered, propensity to share the data, satisfaction obtained. part of the report illustrates models suitable for interpreting the results and finally gives some advice for promoting data services. some international comparisons of the results are proposed in the annex. background in october 2017 tuomas alatera began a mail discussion among members of iassist about data sharing, writing the email ‘i would share the data but…’. following this starting point many suggestions came from iassist members (see annex 2). in november 2017 at the university of turin a seminar was held on data archiving, dissemination and reuse. a backward sight to go ahead which addressed many similar questions. this cawi survey was conducted between 18th november and 18th december 2017 and reached 83 people working in the field of data curation and data analysis across the world (the sampling list used almost 500 email addresses of iassist members, 69 email addresses of eddi172 ( european data definition initiative, 17th congress, lausanne, 5-6 december 2017) participants and almost 500 participants at the turin seminar and from our mailing list. the questionnaire was built using selected questions from the email exchange among the members of iassist cited above and already used in the survey data sharing and data reuse practices and perceptions among scientists worldwide3. the questions regard different aspects of data sharing: tools used in building metadata, problems encountered in order to share the data, the propensity to share the data, and the satisfaction obtained from different working tasks. the problem we tackled was at first no more than a simple curiosity, ‘do there exist any differences in the attitudes about data sharing between italians and not italians’? which soon became a more exciting questions file. another idea came into my mind. given the decreasing amount of the traditional financial resources perhaps it is now necessary to sell these services to a wider and renewed market. https://doi.org/10.29173/iq912 2/39 bonifacio, flavio (2018) differences in data-sharing attitudes and behaviours, iassist quarterly 42 (3), pp. 1-40. doi: https://doi.org/10.29173/iq912 survey objectives the goal of this work is to analyse attitudes toward data sharing in order to find strategies to expand data sharing and to show the best paths to gain new ‘premium followers’, as we call the best performers in data sharing (below). this pragmatic point of view comes from empirical evidence that is compliant with other more detailed analysis coming from wider surveys. we confirm what other sources say: ‘results show that researchers in different regions have different perceptions about data and different data behaviours’4. although in literature there are many other well documented perspectives from which to examine data sharing, covering a range of aspects5 from the political to the technical, we think that one of the most pressing issues is how to expand best practices in sharing data. there are many purported motives getting in the way of data sharing: from modesty to ownership, from sabotage to fear.6 every sort of excuse is used to justify the dislike of data sharing. an overview of several collected opinions7 would be needed to get a more precise idea of the topic. this paper will focus on how to target the market to sell services related to data sharing: 1. is there a target for selling data-sharing services among data users? 2. are there any personal perspectives and attitudes that may influence data services diffusion? in order to answer these questions, we will: 1. examine the one-way variable distributions for background variables to describe the sample 2. examine the one-way variable distributions for attitudes variables 3. combine attitudes variables via pca analysis to build indexes (work satisfaction index, sharing problems index, sharing propensity index) 4. build a typology from the indexes 5. analyse bivariate models with country of origin as independent variables and attitudes variables, indexes and typology as dependent 6. analyse bivariate models with other background variables (gender, age, study title and work sector) as independent variables and indexes as dependent 7. analyse a three variables path with country of origin as exogenous variable, working sector as endogenous variable and propensity index as dependent variable 8. analyse bivariate models with background variables as independent variables and metadata use as dependent to reinforce the hypothesis of existence of differences among data users 9. definition of target variable for services promotion activity 10. build a logistic model to set the best predictors for the target. warnings: 1. the population of this survey are data users. the iassist member list, the list of participants at the meetings quoted above, are just the sources where the email addresses have been found, the sample lists. to be more precise our population is restricted to data users comprised in those lists. 2. in this paper we refer to a country of origin variable, italians and not italians. i did this not because i think italians are so important in the modern data world, but because i supposed they are less involved in data sharing practices and might be representative of other countries in the same condition. as the following report will show they may have different attitudes on data sharing. there are other reasons: there were also no funds and resources to expand the research, for example. 3. look at the annexes: there are research operations that may be clarified only with a methodological in-depth analysis. annex three about factorial analysis explains why i choose those factors i used. https://doi.org/10.29173/iq912 3 4. look at the dimension of the sample of this survey: we had 83 respondents. the confidence interval may become very large. every effort to build theories at this point would be premature and presumptuous and not sufficiently data based. instead, the results should work as a suggestion to continue the survey in a quantitative or qualitative manner. the survey results one-way distributions: background variables first, who are the respondents to the questionnaire? almost half of them work at a university, a quarter in private agencies, and one fifth in public administration or in a non-profit area (tab. 1). tab. 1 – work sector distribution n % university 40 48.2 public administration 10 12.0 private company 22 26.5 no profit 8 9.6 other 3 3.6 totale 83 100.0 although the sample drawn from the sample lists8 is not random, it is representative of data workers: all those interviewed are involved in the field of data analysis (at least as a user) or in data curation. the results must be viewed keeping this mind: we are dealing with a population of data field experts. among respondents there are slightly more men than women (tab. 2) and the most represented age cohort is 51-65 (tab. 3) who are the oldest ones. does this mean that the interest in data sharing is decreasing? tab. 2 – distribution by gender n % male 43 51.8 female 40 48.2 total 83 100.0 4 tab. 3 age distribution the age pyramid of respondents is reported in fig. 1. although there is not statistical significance between genders and age we note that females are more represented in the 18-30 and 51-65 age class and slightly more also in the 41-50 class. fig. 1 – age pyramid as we can expect, the majority have a post-graduate qualification (master or phd) and almost all are at least graduates (tab. 4). n % 18-30 8 9.6 31-40 12 14.5 41-50 24 28.9 51-65 31 37.3 over 65 8 9.6 total 83 100.0 males females over 65 51 – 65 41 – 50 31 – 40 18 – 30 0 5 10 15 20 20 15 10 5 0 5 tab. 4 – distribution by study title n % high school 2 2.4 graduation 34 41.0 master 31 37.3 phd 16 19.3 total 83 100.0 half of respondents come from social studies and one third from scientific studies (tab. 5). as expected social studies are preeminent. tab. 5 –study sector n % humanities 13 15.7 scientific 28 33.7 social (sociology, psychology, economics) 42 50.6 total 83 100.00 one-way distributions: data sharing attitudes as reported above, most of the questions have been extracted from imdsm2017 (see annex 2) and from the quoted survey changes in data sharing and data reuse practices and perceptions among scientists worldwide and grouped into three main conceptual frameworks: 1. problems emerging in data sharing (21 items) 2. items influencing work satisfaction (7 items) 3. items influencing data sharing propensity (6 items) to avoid bias due to oversampling of italian experts, frequencies are weighted in such a way that the italian subsample will lose weight in the figures. the solution is somewhat artificial and not optimal, due to the impossibility to calculate the effective weight of italian experts over the entire data expert population. if we count as a proxy the number of italian experts in the international organizations (i.e. iassist) this number tends to be zero. instead we have taken into account the conditional frequency of using metadata standards (the precise survey question is: what metadata standards do you currently use to describe your data?) given the subsample of not italians as a weight. the underlying hypothesis is that those who do not use metadata standards are less involved in the data sharing activity. the weight is given by the fraction .74286/.42169 for not italians and .25714/.57831 for italians where numerators are the proportions of not italians using (.74286) (or not using (.25714)) metadata standards and denominators the proportions of the two subsamples. as a result, italians will weigh almost half and not italians almost 1.76 times the real subsample weight. we will use this weight to present one-way frequencies of the attitudes shown by the people interviewed. this weight has no effect when we consider the conditional distributions given by the subsample type (not italian/italian subsamples). 6 problems emerging in data sharing among the problems mentioned, those recognized as creating more difficulties are, in order of frequency: confidentiality, lack of funding, lack of time, intellectual property, no clear definitions of ownership, my data are old and not sufficiently documented, i have not finished analysing the data yet, the data belong to my organization (agree strongly, agree somewhat over 30%, confidentiality over 50%, fig. 2) (the benchmark value is 21.78, the marginal distribution value over all items of the considered values). also mentioned are problems related to privacy and propriety, to resources (time and money), to data quality and documentation, and to data analysis. fig. 2 how much do you agree with these statements about the reasons that prevent data sharing? i would share the data but… items influencing work satisfaction data gathering, data analysing and data searching are the items that most influence work satisfaction (all over 70%). documentation (51%) and metadata preparation (49%, less than half) tools are the least satisfying (fig. 3, the benchmark value is 60.55, the marginal distribution value over all items). 7 fig. 3 the following statements relate to how you collect and use research data. tell us how much you agree with the following ways to complete this sentence: i am satisfied with the… items influencing data sharing propensity the most appealing items stimulating the propensity to share data seem to be sharing data among broad groups of researchers, creating new datasets from shared data and the ease of data access (agree strongly, agree somewhat over 80%, fig. 4). the least cited item is i would be willing to place all of my data into a central data repository with no restrictions which equals almost half of the expressed preferences (52%) (the benchmark value is 73.11, the marginal distribution value over all items). this order is interesting because it leaves “my data sharing”, the active part of data sharing, in the last position. fig. 4 the following statements relate to sharing scientific data. tell us how much you agree with each statement. 8 three attitude scales the items shown in the previous tables work very well together showing a very high reliability coefficient. we built three attitudes scales or indexes, one for every item battery: wsiwork satisfaction index with a standardized coefficient (cronbach’s alpha) of 0.90, spi-sharing problems importance index with a standardized coefficient (cronbach’s alpha) of 0.91, spn-sharing propensity index with a standardized coefficient (cronbach’s alpha) of 0.80)9. the weighted distributions10 show that satisfaction and propensity indexes are more concentrated over high scores: a high wsi score means that respondents find the tasks of their work more satisfying than the other respondents (39%, fig. 5), a high spn score means that respondents are more active in data sharing (41% fig. 7), a high spi score means that respondents think that there are problems in sharing data (24%, fig. 6). as we can see only spi has the lowest frequency for high score. fig. 5 – work satisfaction index fig. 6 – sharing problems importance index 9 fig. 7 – sharing propensity index if we look at the correlation among the indexes we find that they are almost uncorrelated, although they come from different pca analysis. that means that a high propensity to share data does not influence the attitude related to problems emerging in data sharing and vice versa, and neither of them influence the satisfaction index. the indexes measure truly different views of the respondents. moving the cut-off point of the index distributions to the mean11 and therefore reducing the numbers of the index categories from three to two, and combining the resulting values, the following figure shows a well-balanced distribution (fig. 8). the first12 category means low sharing propensity and few sharing problems (irreducible reluctant), the second category means low sharing propensity and high sharing problems (reducible reluctant), the third category high sharing propensity and many sharing problems (problematic follower13), the fourth category high sharing propensity and few sharing problems (premium follower). premium followers are the most represented category (almost one third) in the weighted distribution. fig. 8 – distribution by respondent’s typology 10 what those categories exactly mean comes from the pca analysis cited above and more precisely from the correlation matrix between factors and items used in building scales. those matrices are reported here with some comments in appendix 3. here just an explanation of the meaning given to the typology labels is provided: irreducible reluctants because they show low propensity in data sharing and do not recognise the related problems; reducible reluctants because they too show low propensity but have a feeling of the problems (perhaps they have a low propensity because of this feeling); problematic followers because they have a high level of propensity but also a high perception of problems (it seems that something is missing for them); premium followers because they have high propensity and do not perceive so many problems (presumably because they have solved them). if we think at the quadrant of customer satisfaction surveys, usually built combining importance and satisfaction of the items surveyed, each labelled action may be adapted to our typology: warning for irreducible reluctants, improvement for reducible reluctants, exploitation for problematic followers, maintenance for premium followers. in fig. 9 is shown our magic quadrant reporting the scores obtained by each respondent for both of the measures used by the typology. each point counts as one (it is not weighted). fig. 9 – data sharing magic quadrant 11/39 bonifacio, flavio (2018) differences in data-sharing attitudes and behaviours, iassist quarterly 42 (3), pp. 1-40. doi: https://doi.org/10.29173/iq912 models in this section some significant relations between the most important variables described above will be shown. the sample is small, so it is not possible to work on the model with more than one or two independent variables14. nevertheless, we have some interesting suggestions. first, there are the differences between the italian (as an example of countries with low interest in data sharing, where data sharing is not so widespread) and not italian experts interviewed. then some interesting models considering the work place and other variables are presented. first on the list, for each items battery, are the items where the p value of fischer f (homogeneity of variance test) is less than .1 indicating that the relation between the item and the subsample type is significant. tables 6, 7, 8 report the average of each item of the sharing problems importance index, coded from 1 (fewer problems recognized in data sharing) to 5 (more problems recognized in data sharing). we observe that for all the items where the statistic f is significant, and for the majority of the other items, not italians recognize fewer sharing problems (or assign less importance to the sharing problems) than italians. the biggest differences (the ones listed first) are for the opinions on data misuse, on economic convenience and on visibility of data sharing benefits (all significant at the 1% level). tab. 6 – items pertaining to the problem importance index ordered by ascending p value means difference between not italians and italians items label total p value not italians italians f statistic p value f people might misuse my data 2.039 1.771 2.813 13.57 0.00041 i spent a lot of money on this research, and it is not economically convenient to share it 1.565 1.400 2.042 9.522 0.00278 i can’t see the benefit 1.506 1.343 1.979 8.245 0.00521 i would lose control of the data 1.730 1.571 2.187 6.902 0.01029 my data change too quickly 1.879 1.714 2.354 6.414 0.01325 my data are old, they don’t answer the questions researchers ask today 1.586 1.429 2.042 6.061 0.01594 there is intellectual property in the data 2.645 2.457 3.188 4.456 0.03786 my data have been gathered under complete assurances of confidentiality 3.166 2.971 3.729 4.427 0.03847 it could negatively affect my career 1.522 1.429 1.792 2.942 0.09012 i would not know where and how to share data 1.868 1.743 2.229 2.694 0.10459 it’s my data. i don’t want to share it 1.474 1.400 1.687 1.743 0.19051 https://doi.org/10.29173/iq912 12 items label total p value not italians italians f statistic p value f i have not got time to prepare data for sharing 2.765 2.886 2.417 1.731 0.19197 data belong to my organization/university/company and it doesn’t give me the permission 2.464 2.343 2.812 1.641 0.20379 i don’t want other people taking credit for my work 1.947 1.857 2.208 1.363 0.24640 there are no clear definitions of ownership or usage rights 2.586 2.514 2.792 0.702 0.40465 lack of funding 2.840 2.914 2.625 0.574 0.45090 i have collected audio-visual data and cannot anonymise them 2.155 2.114 2.271 0.220 0.64012 we no longer have datasets 1.777 1.743 1.875 0.210 0.64819 i have not finished analyzing the data yet 2.559 2.543 2.604 0.032 0.85819 files are old or damaged, and there is not enough documentation (metadata) 2.479 2.486 2.458 0.006 0.93662 data cannot be anonymized 2.138 2.143 2.125 0.002 0.96063 not significant differences exist on the satisfaction index, where the average values are very similar between the two groups. there are no significant differences at the 10% level. the nearest values are on the cataloguing process. nevertheless, not italians show a little bit more satisfaction for every aspect. tab. 7 – items pertaining to the satisfaction index ordered by ascending p value means difference between not italians and italians items label total not italians italians f statistic p value f process for searching for my own data 3.611 3.686 3.396 1.155 0.28563 tools for preparing my documentation 3.414 3.486 3.208 1.148 0.28705 process for analysing my data 3.845 3.914 3.646 0.980 0.32524 tools for preparing metadata 3.441 3.486 3.313 0.406 0.52594 process for collecting my research data 3.675 3.714 3.563 0.338 0.56262 process for storing my data 3.649 3.686 3.542 0.299 0.58622 process for cataloguing/describing my data 3.478 3.514 3.375 0.264 0.60908 13 for every item of the sharing propensity item scale the propensity is higher for not italians, especially for the willingness to share data among a broad group of researchers and for the creation of new datasets moving from already shared data, both below 1% of significance. tab. 8 – items pertaining to the sharing propensity items ordered by ascending p value – means difference between not italians and italians items label total not italians italians f statistic p value f i would be willing to share data across a broad group of researchers 4.541 4.743 3.958 27.63 0.00000 it is appropriate to create new datasets from shared data 4.318 4.457 3.917 7.553 0.00738 i would use other researchers’ datasets if their datasets were easily accessible 4.313 4.429 3.979 3.886 0.05210 i would share my data if someone took care of the archiving process 3.893 4.000 3.583 2.144 0.14703 i would be willing to place all of my data into a central data repository with no restrictions 3.500 3.543 3.375 0.308 0.58038 i am satisfied with my ability to integrate data from disparate sources to address research questions 3.580 3.600 3.521 0.070 0.79240 the not surprising conclusion at this point is that the countries having an important data sharing tradition show fewer problems and a greater data sharing propensity. a synthetic overview what is more exciting is to analyse the differences with a more synthetic view directly over the indexes, instead of analysing each item, as we have done above. while the work satisfaction (0 not satisfied at all, 100 fully satisfied) index does not show significant differences by country, although not italians seem to be more satisfied, significant differences exist for sharing problems importance (0 not important at all, 100 highly important) and sharing propensity (0 not important at all, 100 highly important). not italians recognize fewer problems with data than italians and are more willing to share it (tab. 9). this fact is probably related to the lower data sharing propensity showed by italians. we noted above that data sharing propensity and sharing problem importance are uncorrelated. that is true. but if we run a model with sharing problem importance as dependent variable, sharing propensity as independent and country of origin as controlling variable, the regression shows significant estimates for the italian group with a positive correlation between sharing problem importance and sharing propensity. this relationship may be interpreted as an effect of problems encountered on the sharing propensity: the ones that have a higher sharing propensity recognize more problems. this may be counterintuitive, but may be plausible. only those who work in the field may see the difficulties. 14 tab. 9 – means difference over the three indexes: satisfaction, problems importance and sharing propensity by country of origin work satisfaction sharing problems importance sharing propensity f=0.97 p(f)=0.3265 f=5.66 p(f)=0.0197 f=8.99 p(f)=0.0036 n mean mean mean country 48 60.84 43.98 51.84 italians not italians 35 65.98 30.62 71.70 total 83 63.01 38.34 60.21 following our classification, not italians are more frequent among followers (premium and problematic), while italians are more present among reluctant (irreducible and reducible) (tab. 10). the table is significant at an alpha level of 5%. tab. 10 – means difference by typology and country of origin typology total irreducible reluctant premium follower problematic follower reducible reluctant % % % % n % country 25.00 12.50 22.92 39.58 48 100.0 italians not italians 17.14 40.00 25.71 17.14 35 100.0 other models given the dimension of the sample, it is hard to test a model with more than one or two independent variables. looking at models with one independent variable, we meet four significant relationships. gender is linked to the sharing problem importance index at a significant level: women seem to recognize more problems than men (tab. 11). it would be interesting to study this relation in a more detailed way. is it an indicator of a sort of digital (sharing) divide? or is it an indicator of a “natural” more attentive female attitude? the data do not permit the verification of such a hypothesis. further (qualitative?) studies would be needed on this topic. 15 tab. 11 – means by index type and gender work satisfaction sharing problems importance sharing propensity f=0.12 p(f)=0.7281 f=11.89 p(f)=0.0009 f=0.37 p(f)=0.5435 n mean mean mean 43 65.48 25.62 68.51 male female 40 63.89 41.97 64.79 total 83 64.66 34.05 66.59 although the older respondents seem to be more satisfied, to recognize fewer problems and share data more readily, age does not show a significant relationship with the measures tested (tab. 12). what is worth noting here is that sharing problems importance is decreasing as the age increase, while the sharing propensity moves in the opposite direction: the sharing propensity increases with age. tab. 12 – means by index type and age work satisfaction sharing problems importance sharing propensity f=0.89 p(f)=0.4761 f=1.12 p(f)=0.3523 f=1.34 p(f)=0.2641 n mean mean mean 8 62.49 43.14 58.16 18-30 31-40 12 65.21 27.61 61.95 41-50 24 59.24 31.95 60.71 51-65 31 66.49 38.33 71.52 over 65 8 74.51 26.16 81.35 total 83 64.66 34.05 66.59 the only significant effect of academic qualifications is on sharing propensity: the higher the qualifications, the higher the sharing propensity (tab. 13). 16 tab. 13 – means by index type and study title work satisfaction sharing problems importance sharing propensity f=1.22 p(f)=0.3019 f=1.03 p(f)=0.363 3 f=4.18 p(f)=0.0189 n mean mean mean 34 58.73 39.10 53.19 graduation master 31 66.71 30.62 67.94 phd 16 67.84 35.69 78.72 total 81 65.08 33.60 66.66 working in an academic or public environment fosters data sharing propensity too (tab. 14). tab. 14 – means by index type and work sector work satisfaction sharing problems importance sharing propensity f=0.26 p(f)=0.8570 f=1.84 p(f)=0.147 0 f=8.93 p(f)<0.0001 n mean mean mean 40 65.25 35.81 73.33 university public administration 10 59.83 27.52 77.76 private company 22 61.32 42.34 45.33 no profit 8 64.58 21.81 39.97 total 80 63.97 33.94 65.78 17/39 bonifacio, flavio (2018) differences in data-sharing attitudes and behaviours, iassist quarterly 42 (3), pp. 1-40. doi: https://doi.org/10.29173/iq912 a more complex model to analyse the effects of country of origin (v3) and work sector (v2) on the sharing propensity index (v1) we tested two structural models15. the first uses a reduced model not considering the effect of the country of origin on the work sector. this assumption does not reflect the reality because not italian respondents more often work at the university (tab. 15). tab. 15 – distribution by work sector by country of origin work sector total private sector university % % n % 54.17 45.83 48 100.00 italians not italians 20.00 80.00 35 100.00 coding italians as zero and not italians as one, working at university or in the public sector as one, otherwise coding 0, we obtain the model schema of the saturated model reported below (only direct effect, fig. 10), which seems to be the most appropriate and not reducible model as seen above. all the effects are significant at alpha=5%, also the indirect effect of the country of origin that influences the sharing propensity via the work sector, which is positive. the total effect of v3 (country) over v1 (data sharing propensity) is given by 0.1964+(0.3629*0.3298) that equals 0.3161. in other words, in this sample not italians more often work at university, so they add to their own higher sharing propensity also the fact that they work at university. we know that working at university has its own positive effect on data sharing propensity too, as the model shows. fig. 10 – path model with sharing propensity as dependent variable (v1) and work sector (v2) and country of origin (v3) https://doi.org/10.29173/iq912 18 looking for a target to sell data sharing services the data provides some other evidence that should be outlined before answering the question ‘what kind of target we are looking for?’ the previous paragraph describes a model where the dependent variable is the propensity to share data. now we observe the relations between the same independent variables with the use of metadata standard documentation, taken as a proxy of the real use of data sharing procedures. table 16 shows that not italians use metadata three times more than italians. tab. 16 – use of metadata standard by country of origin using metadata standard total no yes % % n % 70.83 29.17 48 100.00 italians not italians 25.71 74.29 35 100.00 table 17 shows that those working in the public sector use metadata more than the others, but the relation is not significant at the alpha level of .1. tab. 17 – use of metadata standard by work sector using metadata standard total no yes % % n % 48.25 51.75 33 100.00 private sector university 32.90 67.10 50 100.00 if we look at the relationship between the use of metadata and the typology, we discover that it is significant (p value less than 1%) and that 25 people (over 83) are in the previously defined condition of reducible reluctant (tab. 18), while 20 of them (80%) are not using metadata. (tab. 19) 19 tab. 18 – typology by use of metadata standard typology total irreducible reluctant reducible reluctant premium follower problematic follower % % % % n % 8.61 49.97 19.99 21.42 43 100.00 no yes 25.44 6.80 40.63 27.12 40 100.00 it seems reasonable to consider a target for a ‘marketing campaign’ aimed at promoting data sharing among those not familiar with meta documentation tools (tab. 19), having problems with data, working more often than others in a private environment (tab. 20) and more often in italy (tab. 21, remember that in this sample italy stands for any place where data sharing is not sufficiently widespread). that is the reducible reluctants, as we call them above. first, all tables show high level of significance. furthermore, we note that: 1. irreducible reluctant (low sharing data propensity-low sharing problems), premium followers (high sharing data propensity-low sharing problems) and problematic followers (high sharing data propensity-high sharing problems) all use metadata standard more than reducible reluctant (low sharing propensity-high sharing problems). worthy of note is the fact that, among irreducible reluctant, 9 respondents over 12 use metadata standardized within their organization, which is a limited standardization. 2. the followers more often come from university. among the premium followers academics are 10 times more than non-academics. 3. in each typology category, not italians are more present and in the special case of premium followers they are eleven times more than italians. 4. reducible reluctant are proportionally more among those not using metadata standard and among italians. they work at university more than irreducible reluctant. tab. 19 – use of metadata standard by typology no yes % % 16.77 83.23 irreducible reluctant reducible reluctant 81.38 18.62 premium follower 22.65 77.35 problematic follower 31.99 68.01 20 tab. 20 – work sector by typology private sector university % % 58.39 41.61 irreducible reluctant reducible reluctant 37.32 62.68 premium follower 8.07 91.93 problematic follower 25.56 74.44 tab. 21 – country of origin by typology italians not italians % % 33.55 66.45 irreducible reluctant reducible reluctant 44.42 55.58 premium follower 9.76 90.24 problematic follower 23.58 76.42 joining these variables with some personal characteristics, reducing to two the characters of typology (1 the reducible reluctant, 0 the others) and reversing the model taking as independent variables the work sector (0 private sector, 1 university and public sector), country of origin (0 italians, 1 not italians), the use of metadata standards (0 do not use, 1 use), gender (1 male, 2 female), age (1=18-30 … 5=over 65), academic qualifications (1 high school, 2 graduation, 3 phd, 4 master) we built a score using reducible reluctant as a target variable. the logistic model, which reaches a rescaled r square of 45% and is significant at an alpha level of 1%, retains as significant effects of independent variables at alpha level = .1 (tab. 22) the use of metadata standard (reducing the probability to be a target), the country of origin (reducing also the probability to be a target, i.e being italian increases the probability to be a target), gender (females increase the probability to be a target), age (being young increases the probability to be a target), qualifications (higher academic qualifications increases the probability to be target too). in other words, the model says that if we do not use metadata standard, we work in italy, we are female, young and have a higher academic qualification, the probability to be a target is higher. although the effect of the 21 work sector is not significant, we may argue that the probability to be a target is higher also if people work in the private sector (tab. 22). it seems that the model takes care of the warnings noted above. tab. 22 – logistic model: analysis of maximum likelihood estimates parameter df estimate standard error wald chi-square pr > chisq intercept 1 -2.6399 2.1437 1.5165 0.2182 s05 use of metadata standard 1 -2.4485 0.7556 10.5000 0.0012 ra01 – work sector 1 -0.1203 0.8031 0.0224 0.8809 rprov – country of origin 1 -2.1824 1.1335 3.7069 0.0542 f01 gender 1 1.3062 0.7345 3.1622 0.0754 f02 age 1 -0.5084 0.3108 2.6760 0.1019 rf03 studies 1 1.0625 0.6195 2.9415 0.0863 using 0.45 as a cut-off point for the estimated target, as suggested by the roc curve (fig. 11, .45 corresponds to .44 of sensitivity), there are 25 people in the estimated target (the same marginal distribution as the observed target). fig. 11 – roc curve used to set the cut-off point for logistic model comparing the estimated and the observed targets between the random model and the logistic model by means of the appropriate confusion matrices (tab. 23, 24), we observe that: while a random model guesses 28% of the target, the logistic model guesses 60% of the target, which is more than double. when this model predicts the target it is wrong in 40% of the cases given the prediction (false positive). when this model predicts not target it is wrong in 17% of the cases given the prediction. in other words, if we 22 promote data services when the model says ‘no’ we have a high risk of being unsuccessful (83%). this risk is more than halved if we promote services when the model says ‘yes’ (40%).16 tab. 23 confusion matrices for the random and logistic model: random model target total no yes % % n % target observed 68.97 31.03 58 100.00 no yes 72.00 28.00 25 100.00 tab. 24 confusion matrices for the random and logistic model: logistic model target total no yes % % n % target estimated 82.76 17.24 58 100.00 no yes 40.00 60.00 25 100.00 as an empirical view of model behaviour, once again we reverse the model to consider the conditional probability of the model’s independent variables given the estimated target prediction. so, by definition in our estimated target we have people that do not use metadata documentation tools (96%), work more often in the private sector (52%), work more often in italy (72%), are more often female (84%), more often young (80% less than fifty years old), more often have a higher academic qualification (52%). this is a clear indication for the market strategy direction (tab. 25). tab. 25 – characteristics of the estimated target independent model variables % p-value chisq do not use metadata standard 96% .0000 work in private sector 52% .1346 work in italy 72% .0861 are female 84% .0000 23 are young 80% .0104 have a higher academic qualification 52% .8638 24 some conclusions what kind of story does this survey tell us? first, there at least two different attitudes about data sharing, probably coming from different data culture (the way data are used in a research project). these different views make different attitudes: one, the not italians, more attentive to data sharing, recognising fewer problems doing it, being more satisfied. the other, the italians, are less proactive for every aspect considered. we might tell other stories too, regarding several aspects we encountered during the work: it may not sound so good that the youngest age class shows the lowest data sharing propensity while the oldest shows the highest. it is not surprising instead that we can find the highest sharing propensity among the public administration and at university or among those with the highest academic qualifications. also, not surprising is the fact that the highest evaluation of data sharing problems is among the private sector. a little bit more unexpected is that females evaluate sharing problems more than males, as we noted above. to conclude, in answer to the two questions reported at the beginning of this article: 1. is there a target for selling data-sharing services among data users? if we think of the listed objective reasons and attitudes the answer is: yes, if supported with a renewed promotion campaign. this target is composed by data users not having a great propensity for data sharing but recognising problems in data sharing (reducible reluctant). once the problems encountered are resolved, it is likely that also the data sharing propensity will increase. the models show who are the reducible reluctant, as we have seen above: italians (every country where data sharing is not so used), who does not use metadata standard, works in a private sector, is young and female with a higher academic qualification. this is most probably the most immediate target for data sharing promotion. that does not exclude that also the irreducible reluctants will be comprised in the target in the future. it will probably require more effort, given the characteristics of this data user group. 2. are there any personal perspectives and attitudes that may influence data services diffusion? yes, as stated by the indexes: propensity, problem, satisfaction indexes and by every item used to build them. among the first five items more evaluated as data sharing problems by the respondents, three concern the problem of data ownership and data access regulation. one question comes to mind: will be the gdpr17 be the right answer, although partial, to these troubles? but to answer this question we need further research… what to do? a hint comes first from the respondent distribution on the items regarding cooperation with a broad group of researchers and the opportunity to create new datasets from shared data, both of which scored highly (data sharing propensity scale). fostering cooperation among researchers and data reuse is the highway we have to take in order to gain more premium followers18. this statement is not a great novelty, especially for members of iassist that know the problem very well. nevertheless, it is worthy of note because the response to those items is significantly lower where the interviewed people are not very used to sharing data. furthermore, we have to broaden the scope of service promotion, moving from ‘developed countries’ to ‘developing countries’, those where data curation is less practised, to younger people, involving women in greater responsibility and more remunerative roles. how to do that is a matter that goes beyond the scope of this article, but one suggestion is to transform the self-referential meetings into open symposiums, for example moving from the usual locations (for iassiters usa, canada, north europe) to, perhaps, less easy locations, such as southern europe, africa or asia, and to less easy environments, outside the university, in an open public and private space. 25/39 bonifacio, flavio (2018) differences in data-sharing attitudes and behaviours, iassist quarterly 42 (3), pp. 1-40. doi: https://doi.org/10.29173/iq912 acknowledgements i would like to thank dr. veronica baldisserri for helping me in the questionnaire construction for the web and susan phillips for reviewing my english. the student silvia canavesio helped me in redaction of the final text and worked as a support in statistical analysis. a special thanks to the members of iassist that contributed to the question list and to everyone that participated in the survey https://doi.org/10.29173/iq912 26 annex 1: international survey results comparison as a final step we report some statistics on items comparable with the ones published in data sharing and data reuse practices and perceptions among scientists worldwide, quoted above (only the items present in both surveys). remember that the american surveys regard around 1000 people (tab. 26, 1329 baseline survey, 1015 follow-up survey), while this report has 83 people19). the following table, which orders the items on the absolute value of mean differences, show not so big differences, starting from 0.03 and reaching in two cases a difference greater than one. furthermore when comparing the item means of the actual survey with the item means of the american follow-up survey using the 99% confidence level, we cannot refute the hypotheses that both values are coming from the same population, at least for the first 7 smallest differences. means difference test between actual survey and follow-up survey ordered by absolute value of the difference items label follow-up mean americans actual survey mean differences between actual survey mean and follow up mean follow up mean is in the confidence interval at the level of 95% of the actual survey mean follow up mean is in the confidence interval at the level of 99% of the actual survey mean i would use other researchers’ datasets if their datasets were easily accessible 4.330 4.31 -.017 ** *** process for cataloguing / describing my data 3.520 3.48 -.042 ** *** it is appropriate to create new datasets from shared data 4.230 4.32 0.088 ** *** process for analyzing my data 3.940 3.85 -.095 ** *** i would be willing to share data across a broad group of researchers 4.390 4.54 0.151 *** process for searching for my own data 3.430 3.61 0.181 ** *** i would be willing to place all of my data into a central data repository with no restrictions 3.230 3.50 0.270 *** 27 items label follow-up mean americans actual survey mean differences between actual survey mean and follow up mean follow up mean is in the confidence interval at the level of 95% of the actual survey mean follow up mean is in the confidence interval at the level of 99% of the actual survey mean lack of access to data generated by other researchers or institutions is a major impediment to progress in science 3.990 4.27 0.275 tools for preparing my documentation 3.110 3.41 0.304 process for collecting my research data 4.050 3.68 -.375 i am satisfied with my ability to integrate data from disparate sources to address research questions 3.190 3.58 0.390 lack of access to data generated by other researchers or institutions has restricted my ability to answer scientific questions 3.360 3.78 0.416 others can access my data easily 3.150 2.60 -.553 tools for preparing metadata 2.870 3.44 0.571 process for storing my data 3.030 3.65 0.619 data may be misinterpreted 4.120 2.44 -1.68 data may be used in other ways than intended. 4.210 1.86 -2.35 28 annex 2: imdsm2017, iassit members data sharing mail, fall 2017 iassit members data sharing mail i would share my data but… members of iassist suggested text a my university holds ownership and won't let me my pi/collaborators won't share i'm not done with it yet (15 years later -haven't touched it in 14.5 years) b … it is difficult to compile various parts of data into a coherent dataset suitable for reuse … file formats are old or damaged, and there is not enough documentation (metadata) to be sure what to share … it is old, it doesn’t answer to the questions researchers ask today … there is no funding to produce a reusable copy of the data … there are no clear rules or recommendations on which datasets should be shared … data is classified or a nda was required by the funder or a partner (esp. a company) ... but i do not know how or why (nobody has taught me that). … but there isn't (enough) scientific merit in or professional incentive for sharing data. c why would anyone be interested in my data? my data are not of interest or use to anyone else. i have not got time or money to prepare data for sharing. data sharing makes it harder to recruit participants. if i ask my respondents for consent to share their data then they will not agree to participate in the study. i don’t mind making it open, but i worry someone else might object. people might misuse my data. we want people to come direct to us so we know why they want the data. i don’t want other people taking credit for my work. 29 i will if i can have an embargo…is 30 years ok? i want to publish my work before anyone else sees it. no way! my data on public attitudes towards the weather is incredibly sensitive and potentially disclosive. some of what you asked for is confidential. my data have been gathered under complete assurances of confidentiality. we’re worried about the data protection act. i have collected audio-visual data and cannot anonymise them; therefore, i cannot share these data. i am doing quantitative research and this combination of my variables discloses participants’ identity. my data collection contains data which i have purchased and it cannot be made public. there is intellectual property in the data. that data is already published via (external organisation x) d concerns about opening up data, and responses which have proved effective there’s no api to that system we’re worried about the data protection act (uk law) i don’t mind making it open, but i worry someone else might object it changes too quickly there’s already a project in progress which sounds similar some of what you asked for is confidential we don’t have that data that data is already published via (external organisation x) we can’t provide that dataset because one part is not possible what if something breaks and the open version becomes out of date? what if we want to sell access to this data? setting a dangerous precedent 30 fraudsters use data against us e i promised i would keep data locked in my office i don't want it to be used to reverse social policies my research has been used to support i want to be able to co-author all publications f ... but i have better things to do than fill out endless metadata fields that i already filled out elsewhere. ... but i don’t understand the fine print of your licence /agreement and i'm not a lawyer. ... but your clunky system (repository) makes me lose the will to live, never mind finish my deposit. ... but apparently even if i anonymise it there is still risk that individuals will be identified and armed, so i'll just think about it a while longer. ... but you haven't made the case to me that it will affect my academic career one iota. ... but nobody else in my department is doing it and really why should i be the first. g people will criticize my methods. my data is something special i can offer my own students. i spent a lot of money on this research, and it’s too valuable to just give away. what if another researcher misrepresents what i did? what if another researcher uses my research for commercial purposes? that’s not how my funder wanted this used. h someone may scoop me and find something interesting in it before i have a chance to publish it! i’m in a niche field. nobody else could possibly be interested in my data. documenting data so someone else can understand it is complicated. who has time? my institution doesn’t have a repository. i don’t have anywhere to share it. it’s so confusing! i don’t know where to start. someone may scoop me and find something interesting in it before i have a chance to publish it! b, d … legal reasons prevent it (for example because of personal data act) or research ethics code prevents sharing b, e … i’ve promised to my research subjects that i don’t share the data or i haven’t said anything about archiving (no informed consent) b, g … sharing might lead to need to give advice/guidance to those who reuse of the data (no time for it) 31 b, h … it would be a copyright infringement or i cannot make the copyright clearance a, b, e … no one else can understand the data, it is too personal b, c, e … it is customary to delete the raw data after the research has been carried out b, c, h … data security issues regarding identifiers prevent it; data cannot be anonymized or it is too expensive to do, or it becomes useless when anonymized b, c, d, i … there are no clear definitions of ownership or usage rights c, d we can’t see the benefit. if we publish this data, people might sue us. terrorists might use the data. we’ll get spam. it’s too big. c, h it’s my data. i don’t want to share it, and that’s all there is to it. other researchers would not understand my data at all – or may use them for the wrong purpose. c, d, i people will contact me to ask about stuff people will misinterpret the data my data is not very interesting i might want to use it in a research paper my data is too complicated. my data is embarrassingly bad it’s not a priority and i’m busy c, d, h i don’t own the data, so can’t give you permission. 32 annex 3: factor analysis for work satisfaction index, problems importance index and sharing propensity index for each factor analysis we report the factor variance (eigenvalues) scree plot and the factor pattern matrix. the scree plots report the extracted factors (components) ordered by the quantity of explained variance of the components extracted. in every case we observe that the first factor explains much more variance than the others, which will be the selected component for further analysis. the factor pattern matrix tells us in a synthetic way which of the original items counts more in the computation of the factor final value helping us to give a meaningful sense to the extracted factor. from these analyses we can say first that the factor extracted is for every analysis the most important factor (in terms of explained variance). second, we can get a more precise idea of the meaning of each factor. so, for the work satisfaction index we can say that all items used have a similar importance on the final factor. just the item “tools for preparing metadata” is slightly less important. the problems importance index summarizes several items. therefore, we see more items that seem to be less important in the computation of the factor. among them we find: “my data have been gathered under complete assurances of confidentiality”, “data belong to my organization/university/company and it doesn’t give me the permission” and “i have not got time to prepare data for sharing”. for the sharing propensity index the item “it is appropriate to create new datasets from shared data” shows a slightly smaller correlation coefficient with the respective factor values than the others. 33 factor analysis for work satisfaction index factor pattern wsi factor process for collecting my research data 0.81748 process for cataloging / describing my data 0.84129 process for storing my data 0.87918 process for searching for my own data 0.87144 process for analyzing my data 0.75725 tools for preparing metadata 0.65962 tools for preparing my documentation 0.83165 34 factor analysis for work problems importance index 35 factor pattern spifactor my data have been gathered under complete assurances of confidentiality 0.38103 there are no clear definitions of ownership or usage rights 0.65201 there is intellectual property in the data 0.66610 data cannot be anonymized 0.52725 i have collected audio-visual data and cannot anonymize them 0.44449 data belong to my organization/university/company and it doesn’t give me the permission 0.39950 my data change too quickly 0.66832 lack of funding 0.46256 i have not got time to prepare data for sharing 0.40837 files are old or damaged, and there is not enough documentation (metadata) 0.57591 we no longer have datasets 0.55692 i would not know where and how to share data 0.53569 my data are old, they don’t answer to the questions researchers ask today 0.67454 i can’t see the benefit 0.65204 people might misuse my data 0.67071 i spent a lot of money on this research, and it is not economically convenient to share it 0.66127 it could affect negatively my career 0.68575 it’s my data. i don’t want to share it 0.65698 i don’t want other people taking credit for my work 0.64718 i would lose control of the data 0.69237 i have not finished analyzing the data yet 0.60151 36 factor analysis for work data sharing propensity index * factor pattern spn factor i would use other researchers’ datasets if their datasets were easily accessible 0.74303 i would share my data if someone took care of the archiving process 0.74334 i would be willing to place all of my data into a central data repository with no restrictions 0.79848 i would be willing to share data across a broad group of researchers 0.77074 it is appropriate to create new datasets from shared data 0.66703 37/39 bonifacio, flavio (2018) differences in data-sharing attitudes and behaviours, iassist quarterly 42 (3), pp. 1-40. doi: https://doi.org/10.29173/iq912 references bonifacio, flavio (2017) ‘working across boundaries – public and private domains’, part 3 – a follow-up survey, (available at http://doi.org/10.5281/zenodo.1120237) doorn, peter and tjalsma, heiko (2007) ‘introduction: archiving research data’, springer science+business media b.v. hatcher, larry (1994) ‘sas system factor analysis and structural equation modelling’, sas institute horton, laurence (2016) ‘lse research data management data sharing objections faqs and naughts and crosses game’, (available at http://doi.org/10.5281/zenodo.61978) kim, youngseek and m. stanton, jeffrey (2012) ‘institutional and individual influences on scientists’ data sharing practices’, journal of computational science education, volume 3, issue 1 noble, susan; russel, celia and wiseman, richard (2012) ‘mind the gap: global data sharing’, iassist quarterly, vol. 35, n° 3 qualtrics (2010) ‘the 1936 election – a polling catastrophe’, (available at https://www.qualtrics.com/blog/the-1936-election-a-polling-catastrophe) rasmussen, karsten boye (2014) ‘social science metadata and the foundations of the ddi’, vol. 37, n° 1 ribeiro, cristina and matos fernandes, maria eugenia (2012) ‘data curation at u. porto: identifying current practices across disciplinary domains’, iassist quarterly, vol. 35, n° 4 tenopir, carol; d. dalton, elizabeth; allard, suzie; frame, mike; pjesivac, ivanka; birch, ben; pollock, danielle and dorsett, kristina (2015) ‘changes in data sharing and data reuse practices and perceptions among scientists worldwide’, (available at https://doi.org/10.1371/journal.pone.0134826) yang, meng-li (2013) ‘strategies of promoting the use of survey research data archive’, iassist quarterly, vol. 36, n° 1 notes 1flavio bonifacio, metis ricerche srl, via camerana 6, i-10128 torino, italy. contact e-mail: flavio.bonifacio@metis-ricerche.it. 2 european data definition initiative, 17th congress, lausanne, 5-6 december 2017. we included this file of emails to increase the number of respondents. there is no doubt that the participants at eddi17 are data users. https://doi.org/10.29173/iq912 http://doi.org/10.5281/zenodo.1120237 http://doi.org/10.5281/zenodo.61978 https://www.qualtrics.com/blog/the-1936-election-a-polling-catastrophe https://doi.org/10.1371/journal.pone.0134826 mailto:flavio.bonifacio@metis-ricerche.it 38 3 tenopir, carol and others, op. cit. 4tenopir, carol; d. dalton, elizabeth; allard, suzie; frame, mike; pjesivac, ivanka; birch, ben, et al. (2015) 5kim, youngseek and m. stanton, jeffrey, ( 2012), doorn, peter and tjalsma, heiko, (2007),noble, susan; russel, celia and wiseman, richard, (2012),ribeiro, cristina ad matos fernandes, maria eugenia, (2012) yang, meng-li, (2013),rasmussen, karsten boye, ( 2014) 6 see http://doi.org/10.5281/zenodo.61978 quoted by horton, laurence, iassist members data sharing mail 7 see iassist members data sharing mail, fall 2017 (imdsm2017 in the following) 8 the sample is not random because only those who want to respond to a questionnaire actually answer. the italian sample list was extracted from our mailing list and from participants at the turin seminar amounting to a total of about 500 people (525). the international sample list comes from the iassist member directory (431 people) and from the participants at the eddi17 congress (69 people), resulting in a total of 500 people (percentage of respondents: 9% and 7% respectively). 9 following the suggestions of the reliability analysis, for the last scale we excluded item 4, causing a decrement of cronbach’s alpha. 10 the indexes have been calculated using pca, rescaled with range 0-100 and recoded in three classes. the classes limits, expressed in standard units, are: 1, less than 0.431; 2, between -0.431 and +0.431; 3, greater than +0.431. in the case of normal distributions, the classes would be uniformly distributed. we have also recoded them into two categories, less or greater than the mean value 11 which is zero for the standardized indexes 12 first category includes the respondents with both indexes below the mean value, second category the ones with a propensity index below the mean value and sharing problems index over the mean (that means few sharing problems), etc. 13 definition: if data sharing were a subscriber of fb, then those that follow it are followers 14 only in the logistic model presented below do we use more variables as independent variables 15 we follow the notation proposed by l. hatcher, (1994) 16 we consider ‘at risk’ when we promote services to the wrong target. this happens every time the model predicts correctly ‘not target’ and predicts incorrectly ‘target’. this analysis is not complete because we do not have enough cases to test the model. 17 gdpr (general data protection regulation) regards the management of personal data. personal data here means: “any information relating to a data subject. sensitive personal data, which attracts a high degree of protection, is data which is in relation to race, political opinions, health, sexual life, religious and other similar belief, trade union membership and/or criminal records.” see the link https://www.audiencedatasharing.org/legal-information 18 bonifacio, flavio (2017) http://doi.org/10.5281/zenodo.61978 https://www.audiencedatasharing.org/legal-information 39 19 it is well known that the sample dimension is not the only thing that influences the outcomes of a survey. an example is the poll conducted in 1936 in usa about the presidential election, when the poll using 2 million surveyed persons predicted a. landon as winner against f.d. roosevelt https://www.qualtrics.com/blog/the-1936-election-a-polling-catastrophe/ https://www.qualtrics.com/blog/the-1936-election-a-polling-catastrophe/ iassisl quarterly 23 archives and dinosaurs by eric tanenbaum' introduction dinosaurs and social data archives have a lot in common. when botji began their existence they had their respective fields prett>' much to themselves. having almost exclusive control over their environment for a long period, dinosaurs and data archives both swept up material whenever possible and, in time. 'presented at lassist/ifdo international conference may 1985, amsterdam. this paper has been previouslv published in european political data newsletter no. 55:33-44, june 1985. appeared cumbersome and bottom-heavy. from this state both had to confront a changing environment however, for all the similarities between the two, dinosaurs differ from data archives in at least one important respect — they no longer exist thus while it is too late for dinosaurs to learn from data archives, archivists should consider the dinosaurs' progress if they wish to distance themselves from the dinosaurs' end. this note suggests how they might do so. palaeontologists may differ when assessing the relative weight of specific causes of the dinosaurs' demise, but there is common agreement that non-adaptation to changing climactic conditions is important in their undoing. in modem terms it could be said that dinosaurs were frozen out by a changing hardware environment archives also confront hardware changes, but their impact on archival work is confoimded by concurrent software developments. this paper describes major changes in several areas which affect computerized data archives. on the hardware side, the paper examines improvements in mass storage capacity and the ergonomics of computers (of all sizes). software developments, in parallel with these hardware changes, encotirage new orientations to social information. from among these the paper focuses on "new" database management techniques and electronic publishing — both 'have implications for archive growth. changes in hardware and software are combined by improved communication facilities; the catalyst producing the "alloy" lies in the imagination of information analysts (archive users) whose expectations are aroused by these more elemental)' developments. the paper describes aspects of the agents of change which are germane to the future operation of archives. an integrated systematic approach to the tasks reqiiired ensure that funire concludes this paper. spring 1986 24 iassist quarterly mass storage devices the history of computerized data archives for the social sciences illustrates the evolution of computerized mass storage devices. cardboard computer cards, or "ibm cards" as they were commonly known, were an early de facto standard medium for data storage. the "data archive movement" of the early 1960's was launched when it was recognized that these cards could be banked centrally for subsequent redistribution to other sites which supported this physical standard. although computer cards were reproduced and shipped "by the forest", the medium was not ideal. it is clumsy — cards get dropped, insecure — cards get torn, and expensive — bulk reproduction is a resource intensive activity. it also limited the analyst's access to large volumes of information. clearly faster forms of "data memory" would yield vast improvements in the kind of service that archives could provide resejirchers. the magnetic computer tape offered the medium of distribution that data repositories required. it is not as universal as computer cards, for each brand of computer uses a different mode of tape storage. however, almost all archives have computer softy/are that allows them to read and write tapes written in all formats used in their user constituency. thus, for example, the british esrc data archive maintains a suite of conversion routines that permits it to transform data from its own in-house standard to any form required by british users.' while magnetic computer tapes gave archives a cheap medium for transmitting subsets of their holdings to analysts working at remote sites, the medium constrains the kind of material that can be accessed. first, in almost all cases, it requires that information be stored as sequential files. this immediately limits the scope of data that can be transmitted to a few discrete chunks, if only because of the effort and skill required to reassemble anything more ambitious at the receiver's end. second, the medium itself has a small finite capacity. granted the volume of data that can be stored on magnetic tapes has increased dramatically from the 6.4 megabytes feasible with the earliest tapes to a current 210 megabytes,' but still requires six physical tapes to hold the results of the 1981 british population censuses after the data have been subjected to complex compression routine. operationally, this means that the analyst who wants to select census data from points across the nation is involved in considerable tape manipulation. third, and finally, tapes, which are volatile, offer poor archival security. ensuring the physical integrity of a tape-resident database is a labour and time intensive task which a central facility can perform because it can take advantage of economics of scale but which an individual would find restrictive. for these reasons archivists should welcome the recent emergence of new modes of mass data storage, two of which will be described here -as a prelude to a later discussion of how they should be incorporated into archival operations. several manufacturers have aimounced the development of disks that use laser techniques ^a side benefit of this mode of operation, initially designed to cope with the inelegancies of computer manufacturers' whims about tape standards, is that central archives have protected their, and thus their constituency's, data resources by creating a protective buffer between a single in-house standard to which all data are convened and changing external '(cont'd) reauirements. thus, when external technological changes occur the entire database can be transformed to the new requirement by a single routine operation which "maps" the old format to the new. 'the comparison is between a 2400' reel recorded at 200 bits per inch ("bpi") and one recorded at 6250 bpi. spring 1986 iassist quarterly 25 to input and output information at extremely high densities onto small robust platters. thus, for example, one firm's first release promises the storage of one gigabyte (i.e. 1,000,000,000 characters) on a single side of one physical disk. using the british population census again as an illustration, it ought to be possible to store the entire set of counts on a single disk. while at their initial release the disks, which cost about £200.00, are somewhat more expensive than the conventional computer tapes required to store a similar amoimt of information, the radical impact of these new devices will come both because they allow non-sequential access to data and because they are of archival quality, offering a minimum of ten years' secure storage. data analysts can now realistically contemplate linking large volumes of information from diverse sources in their pursuit of new cormections between and among social phenomena. in response to this facility, archives have to reconsider how they service their constituency. eventually, archives will have to meet the needs of analysts who have access to mass storage devices by supplying mass data packages. these will likely be based on diverse data sources which might in turn be linked by "discrete", but otherwise broad, "story lines".' this orientation to data, for which the italian and norwegian data archives' work constructing ecological databases is a precedent, will have to be extended to many areas of social inquiry and will conceivably require a more active intervention in the work of data archives by subject specialists acting in an editorial capacity. optical disks, because of their robustness and cheapness, are amenable to distribution in much the same way as traditional magnetic tapes are. their local use (by independent analysts) is feasible, as the manufacturers of optical disk drives generally use a standard "interface" between computer and drive. thus, imlike tape equipment, it is possible that this kind of mass storage will soon be available even for desktop "personal" computers. however, optical disks have value to the archives' own computer installations. for the british data archive, it is estimated that over 80% of its files are sufficiently stable' to make it sensible to transfer the bulk of its holdings to these devices. this would have the immediate advantage of simplifying internal operating procedures, even if the data archive continues to supply most of its users with copies of data files for access on their local machines. however, if one considers another development in technology, the "networking" of computers which permit individuals to address many computers directly from a single site, these storage devices assume a higher profile in the archives' future landscape, because, with a conceptually, if not technically, "simple" modification, they ofi'er an almost limitless volimie of fast access data retrieval. physically, optical disks resemble long playing gramophone records. thus, as with gramophone records, these disks can be stored in a machine similar to a "juke box" whereby a would be listener (analyst) can choose any song (data file) that is available within its confines. no human intervention, other than by the "listener", is required. the songs are permanently on-line. suggesting a machine that would keep the "top 40" data files readily accessible to analysts is 'in fact, this is analogous to the approach taken by the british broadcasting corporation's domesday project, which was described elsewhere during the conference and with which the british esrc data archive is collaborating. ^ the first optical disks on the market offer a "write once, read many times" facility. thus, for the moment at least, they are best considered devices for storing stable data. of course, from an archival perspective, the data security oftered by a non-erasable device is a bonus to the mass storage capacity. spring 1986 26 iassist quarterly noi fanciful. in fact, at least one optical disk developer (philips) supplies a "carousel" option for its "megadoc" system. although the system is initially directed to the storage of document images, there appears to be no reason it could not be adapted to numerical data bases. the juke-box approach to data storage is shared with another recently released mass volume device which is based on densely packed cassette-like tape cartridges. although these are not transportable in the way that optical disks are, they offer much more storage potential and have to be considered a likely enhancement to the hardware oftered by a data library service which wishes to support direct access to its holdings by analysts. as mentioned, the impact of improved inter-computer communication facilities on archiving is considered later in this paper. for now, it is sufticient to note that the potential for "on-line" access to masses of data which is made possible by the two devices just described will encourage social researchers to explore the use of the developing "network" capability, particularly as improved storage capacity is interacting with a radical change in the overall provision of computers themselves. a brief description of the "new ergonomics" of computer use is a useful prelude to a discussion of the effect of networks on archives. computers: a changing style traditionally, social science data archives could assume, reasonably, that their catchment area comprised all computer using social researchers. as computer use in social science was intricately unked to a quantitative orientation, computer users were numerate and usually shared a kit of tools that were applied to research tasks. moreover, the "conventional" computer oriented social investigator, who was most attracted to the "calculadve" power oftered by computers, was adequately served by the existing provision of computers in research environments. the computer, physically located in a central position in the institution, was fed numeric data, manipulated them, and then supplied the results of the manipulation. data archives, which were also centrally located, were well-suited to this mode of computer access and in most countries developed strong institutional ties with the providers of computer services used by the research community. in this way, archives could minimize the technical barriers which inhibited researchers' access to their holdings. the recent growth of desktop computers, cheap enough to be purchased by individuals, threatens the homogeneity of the computer-using community. a cursory glance at the "micro-computer" marketplace suffices to show that the main appeal of these machines is not that they are superior calculators but that they are remarkably sophisticated typewriters which manage to combine a keyboard, an electronic scissors and a truly non-spill gluepol more important, though, for archival development, these desktop machines are changing the prevailing view of what constitutes "machine-readable" data. it does not take long with a "word-processor" to recognize that semi-structured textual information often is more easily organized, manipulated and analysed with the help of a computer than it is manually. not surprisingly, facilities offered by desktop workstations are also changing the "traditional" computer analyst's orientation to computerized functions. granted, quantitative analyses of large data sources are still best done by large, central "mainframe" computers, but the post-processing of the results for research reports is now best accomplished with the software (and sometime hardware) facilities offered on micro-computers. thus, for example. spring 1986 iassist quarterly 27 the survey researcher will continue to manipulate the survey's data with a large computer to produce summary information. these results will be captured on the machine on which the report itself is composed. while this might be only to avoid re-typing tables or matrices, the analyst will likely also wish to apply micro-computer facilities like graph processors or spread-sheets to the reduced dataset for further refmemenl both of these instances of expanded computerization demonstrate an increasing integration of information processing which replaces the earlier compartmentalization of computers by specific tasks. clearly, computer use is no longer the exclusive prerogative of the specialist in qiiantitative techniques. this has profound implications for data archives. first, archives are bound to encounter a new community of users who regard them (archives) as just another source of information. these researchers will have been in contact with electronic publications of other types (e.g. bibliographic search services or reference publications) and will not immediately consider the traditional data archive as being in any way difterenl nor should they. it would be odd if the oldest purveyors of computerized information could not service the needs of the newest seekers of that kind of material. of course, data archives do not generally provide the summary (digested) style of information that most reference seekers want however, often that information is available to archives who, however, reject it because it is not their normal stock in trade. in the future, if archives are to serve this new market (which will include a significant proportion of their older market) they will have to make this kind of information accessible. second, the integration of information handling practised by the archives' traditional users will affect what these users expect archives to provide. they too will be less inclined to halt their work progress to permit conventional archive practices to work. they will demand more immediate access to these services, requiring that the archives' input to their information needs be much more transparent archives will (and likely should) become visible only when (infrequent) hitches develop which require intervention. these developing expectations can be traced to the influence of desktop workstations. however their fulfillment can only be realized by archives if the archives have access to the large scale storage capacity described earlier and if the archives can offer access to these facilities to remote users. commimication networks offer the linl computer networks when computers first became available to researchers employed in the british academic sector, they were provided by individual institutions. in time, an informally organized system of resource sharing developed, wherein some institutions assumed a responsibility to service some of the larger needs of institutions in a particular regioa by the end of the 1970's it was likely that a computer user in any given institution would have access to a larger regional centre as well as to a local center. indeed, communication with the larger machine was often via the smaller local computer. although this arrangement allowed researchers to use much more powerful installations than their own institutions could afford to provide had they stayed totally independent, they still offered a limited access route to the country's entire community of computers. for the british data archive this meant that there was little need to develop a facility that enabled direct spring 1986 28 iassist quarterly enquiries to its holdings by external users — most users continued to depend on a magnetic tape based service and the archive was best advised to devote its efforts to improving that mode of data dissemination. recent developments in data communication have changed this aspect of the computer user's working environment in many countries, most computers are now functionally no further away from the user than the nearest keyboard. in the british academic sector, for example, the computer board-sponsored joint academic network ("janet") offers researchers in that sector an appropriate inter-computer link which facilitates communciations among university computers in great britain (as well as with other "networks" in great britain and abroad). janet, as a communication path, meets the need for a facility that allows one computer to talk to many computers. its more important contribution, however, is to hide the intricacies of network use behind a facade which makes it simple for computer users to address multiple computers with little more knowledge than that required to use their local computers. it accomplishes this with "protocols" which standardize message transmission between sites. at the level of the network, messages may be commands to join a remote computer site as a "local" user or to retrieve information to the user's own site. it does not take too much imagination to see what the eftect of this new facility for communication with remote sites might be on the operating procedures of data archives. at the least, many analysts will want to interrogate a catalogue of archival holdings to determine which, if any, files contain information of interest to their projects. however, having located pertinent data sources many will wish to select only those parts that are relevant they may then be happy to "download" the data to their own installations for analysis but in theory, they could just as easily (or perhaps even with greater ease) analyse the data at the archive's site, particularly if the archive had implemented specialized software tools to facilitate secondary analysis. the last paragraph contains an implicit research agenda of projects that are necessary to build the interface for on-line access to archival holdings. an integrated approach to these essential tasks is mentioned in the conclusion to this paper and so the individual tasks need not be dwelt on here. however, at this point it is appropriate to note that the tasks that face an archive also confront any computerized information utility that wishes to encourage direct access to its wares. as these utilities increase in number and as demand for them grows pressure will develop for a coordinated approach to information management on a national (and possibly international) scale which will transcend particular subject orientations. data archives will have to join these integrated systems — they should be in the forefront of developments. in any event whatever their institutional inclinations two related features of the environment in which archives now work, the demand for multiple sources and the acceptance of the "relational model" of data management will push archives in this direction. multiple sources for years, advocates of secondary analysis as a research strategy for the social sciences, and thus supporters of social data archives, have argued that only this form of research allowed linkage of diverse data sources which was necessary to fully explore social phenomena. however, the records of data archives' use patterns suggest that these multiple linkages are rarely made — the majority of secondary analysis seem to be of single data sets. spring 1986 iassist quarterly 29 there are several reasons for this. the first might relate to the difficulties entailed in merging large masses of data supplied on magnetic computer tape. as described earlier in this paper, the user of archival material often had to request much more than was needed, largely because there were no facilities available for obtaining the subsets that were really required. at this level, it could be that the potential of direct user-archive access will be sufticient to encourage users to pre-process archival material before analysing il however the availability of many different kinds of information which are not only in computerized form but which, often, are in only machine readable form will force, or at least teach, researchers to address multiple sources of information. at the onset many of these will be "reference" works which are of interest because they yield independent "facts". however as experience is gained in locating "facts" from several (many?) places and retrieving them for assembly on a single computer, researchers' perspectives will become more ambitous. they will begin to want the same capability of addressing multiple numeric data files which they will tailor to a form which is adequate for reassembly into a purpose-built whole. besides the potential for network access and the increasing availability of multiple sources of computerized information, one more development on the computer landscape will have a great impact on archive users' expectations of the type of service that an information utility should provide. fortuitously, this development, the widespread acceptance of a relational model of data management, also provides users with the tool necessary to take advantage of multiple sources addressed on-line. the relational data mode further commentators on the penetration of computers into "everyday" life during the 1970's and 1980's will highlight the provision of "easy" database management techniques which permit researchers to take multiple logical perspectives of particular group of data. although several different modelling strategies are available, the "relational" approach to database management offers the most exciting and attractive prospect for social researchers because, among all the alternatives, the relational model most closely replicates the way analysts think about data. it thus offers a tool for analysts interested in analysing substantive problems rather than itself becoming tlie goal for which analysts strive. while it is not possible to delve into the details of the model here, it is worth noting that the model's strategy of simplifying the association between discrete sets of data supports the exploitation of multiple data sources when analysing a phenomenon. most importantly, from cm archive's viewpoint at least, it suggests that only those data that are required must be retained when assembling a file for analysis. as suggested earlier, this runs counter to the conventional archive practice of "user takes all", with its demand that the analyst cope with a massive body of unnecessary data. strangely, given the esoteric nature of database management, this is the change which could have the greatest single impact on the demands put to archives in the future. the elegance of the relational approach to data management has attracted many micro-computer program developers. consequently, social researchers who were first introduced to computers via these machines will have experienced "quasi"relational management systems and will have grown accustomed to applying their power. moreover, as micros have (until recently) offered only limited data storage capacity, these spring 1986 30 iassist quarterly new computer users will have learned to work within the confines of these machines. they will not appreciate that moving to larger machines, as they will do when accessing central information utilities, permits a more relaxed view of data storage. as these new computer users represent the "growth" potential for data archives, their influence on archival develoment cannot be ignored. thus it is appropriate that a description of the impact of "new technology" on archives conclude with a speculative note on the most powerful driving force for change, the new user community. between class and political participation is sufficiently special to warrant the cumbersome hurdles that now impede access to archival data. the archivist must be aware that barriers which reflect past contingencies will direct a major portion of their user community to other information services which ofter more flexible access to social information. having said this, it must be recognized that the transition from dinosaur to butterfly will not be an easy metamorphosis. one feasible route toward the changeover is described as a conclusion to this paper. changing people it will be evident from the remarks earlier in this essay about prevailing archival practice that users of social science archives almost invariably came from a small segment of the social science community. oriented toward "quantitative" social research, they grew up with archives and, like archives, learned to accept — and perhaps even like — the "user hostile" environment in which computer users were expected to work. the ethos of computer use has chsmged and new entrants will be unware of the need for a hairshirt archivists, who tend to be of the old school, will have to adjust their expectations of users to conespond to what their expanded catchment area expects of them. this new generation of computer users will treat computers with the same ease as they did typewriters a decade ago. for them, the computer is a general utilit)' for a wide range of tasks, among which is information gathering. people accustomed to interrogating a bank account on line or ordering furniture from a direct access shop, will not consider that assembling cross-national data on the association from dinosaur to butterfly: an easier metamorphosis there is a danger that the earlier discussion which related technological developments and current practice will leave the mistaken impression that archives are unresponsive to change. in practice, archives have worked to incorporate most technological advances into their operating procedures. in the area of networking, one could cite the eec-sponsored access project which is designed to produce a cross-national, integrated, bibliographic, on-line data base. the longstanding development of the cessda study description scheme fosters the bibliographic control crucial to the identifying sources of comparable data. there have been many examples of archival use of centralized mass storage facilities — for example, the british data archive's distributed anangements for the supply of the 1981 population census. however, for all these individual projects, the breakthrough to a comprehensive information service still seems a remote prospect although part of the problem is related to archival practices, a significant share of the difficulties are attributable to more general features of spring 1986 iassist quarterly 31 computer use. as these affect all information providers, the removal of these encumbrances on efficient infonnation access requires the development of an integrated system. it is to this joint eftort that archives should devote their resources. in most countries, the computer user can give an empathetic hearing to the tale of the tower of babel. while communication utilities like janet mask the intricacies of making connections between computers, they do little to improve users' access to different computer systems. in effect, the computer network gets the user to the computer's door but, in most cases, that door is locked against the user's entry unless the user possesses privileged knowledge, to say nothing of privileges. prevailing computer practices, which reflect a period when each institution offered its own computer power and each computer manufacturer devised its own operating system, throw up the greatest barrier to a "butterfly-like" access to information. until this artificial restriction on computer use is overcome, archives and users will be forced to work in an environment in which flexible approaches to information sources are blocked. however, the obstacle could be removed with a central computer-based facility, accessible to all by network commimications, which shields users from difterent computer environments and protects the environments from many difterent users. this facility would offer a classified catalogue of all available information sources in the united kingdom which contained information about the substance of each source and technical information about access arrangements. more importantly, the user would only use the database for subject searches — the technical information, which would be kept transparent to the user, would be used to "automatically" invoke the dialogue necessary to access the host information sites. social data archives should be promoting the development of a utihty like this. it requires more than a tmilateral ventixre from any single sector and demands more resources than archives themselves can expend. social data archives, nonetheless, have a privileged role among information providers for they were among the earliest to be computerized. thus they offer a rare perspective from which to view the changes described in this paper to those with whom they might cooperate. a central utility like this would benefit social researchers because the only "new" specialist skill required relates to the bibliographic search procedure, which would be common to all sources. it would reward social science because it would allow the exploitation of technological advances which would otherwise be barred to il it would be attractive to social information providers because they could work to one common standard. it should appeal to current data archives because it promises to provide the protection against the technological "chill" thai spelt the dinosaurs' demise." spring 1986 vol282-3.indd 6 iassist quarterly summer/fall 2004 by milo schield 1 information literacy, statistical literacy and data literacy introduction the evaluation of information is a key element in information literacy, statistical literacy and data literacy. as such, all three literacies are inter-related. it is difficult to promote information literacy or data literacy without promoting statistical literacy. while their relative importance varies with one’s perspective, these three literacies are united in dealing with similar problems that face students in college. more attention is needed on how these three literacies relate and how they may be taught synergistically. all librarians are interested in information literacy; archivists and data librarians are interested in data literacy. both should consider teaching statistical literacy as a service to students who need to critically evaluate information in arguments. information literacy the need for information literacy has been highlighted in the us by several organizations including the american library association (ala). in 1989 a call for information literacy was issued by the ala presidential committee on information literacy.2 in 1989, the national forum on information literacy3 was formed. and in 1998, the ala/acrl (association of college and research libraries) issued a progress report.4 each organization and each report had some differences in their approach to information literacy. but one element was common to all – the need for the critical evaluation of information. the ala and acrl issued a set of information literacy competency standards for higher education.5 information literacy is a set of abilities requiring individuals to “recognize when information is needed and have the ability to locate, evaluate, and use effectively the needed information.” an information literate individual is able to: (1) determine the extent of information needed, (2) access the needed information effectively and efficiently, (3) evaluate information and its sources critically, (4) incorporate selected information into one’s knowledge base, (5) use information effectively to accomplish a specific purpose, and (6) understand the economic, legal, and social issues surrounding the use of information, and access and use information ethically and legally. in their presentation of the standards for information literacy, the american association of school librarians association (aasl) for educational communications and technology presented evaluation as one of three standards6, and identified several related indicators. standard 2 the student who is information literate evaluates information critically and competently. the student who is information literate weighs information carefully and wisely to determine its quality. that student understands traditional and emerging principles for assessing the accuracy, validity, relevance, completeness, and impartiality of information. the student applies these principles insightfully across information sources and formats and uses logic and informed judgment to accept, reject, or replace information to meet a particular need. indicators: (1) determines accuracy, relevance, and comprehensiveness. (2) distinguishes among fact, point of view, and opinion. (3) identifies inaccurate and misleading information. (4) selects information appropriate to the problem or question at hand. evaluating information can be difficult. even government sources may have their agenda or ideology. but evaluating information can become much more difficult when the information involves statistics. there are numerous publications on information literacy.7 statistical literacy a great deal of information involves statistics. how would we talk about current social issues without using basic statistical ideas such as ratios, percents and rates? and with the advent of high speed computing and the internet, everyone is facing a flood of information in the form of statistics. it seems difficult to be considered information literate in the 21st century without being statistically literate. statistical literacy studies the use of statistics as evidence in arguments (schield, 1998, 1999). http://www.ala.org/ala/acrl/acrlpubs/whitepapers/presidential.htm http://www.ala.org/ala/acrl/acrlpubs/whitepapers/presidential.htm iassist quarterly summer/fall 2004 7 joel best (2001, 2004) identified the key to being statistical literate when he noted that all statistics are socially constructed. this isn’t some deep philosophical claim. it merely states that people choose what to count or measure, how to assemble those measurements into summary statistics, what comparisons to form from these statistics and how to communicate these statistics. a key element of statistical literacy is assembly: how the statistics are defined, selected and presented. • for example, in 1999 the difference in us mean household income before and after taxes was $14,000 whereas the difference in median household income was about $7,000.8 just changing the choice of the statistic cut the tax burden in half. • similarly, one person might note that in 1999 the mean us household income before taxes was almost $55,000 while another might note that the median household income after taxes was less than $34,000. an unwary reader might presume the entire difference of $21,000 was due entirely to taxes. • in the us consumer expenditure survey, the 1997-98 table for those under age 25, shows an average expenditure for alcohol of $330 per year. but in many families, this expenditure is zero so the average amount in families that spend money on alcohol can be much higher. the phrase “lies, damned lies and statistics” is well known (first proclaimed by benjamin disraeli and then popularized in the us by mark twain). in many ways statistics are just words in a different form. perhaps the years spent learning arithmetic have led us into thinking that since numbers don’t lie then neither do statistics. but statistics are more than numbers. statistics are numerical summaries about things in reality. the nature of the things being summarized can make a difference. see schield (20005a,b) for examples of statistical prevarication. consider this example. we all know that 6 plus 7 is 13 and that 60% plus 70% is 130%. so if a company has a 60% market share in the eastern us and has a 70% market share in the western us, what is their market share in the entire us? the math says 130%, but we all know that is wrong. market share has a particular meaning or nature. so for statistics, small changes in syntax can create large changes in semantics. as joel best put it, statistics are more like diamonds than like rocks. diamonds are carefully cut and displayed by people to produce a desired effect. statistics, like diamonds, are not 100% natural like sand or rocks. they are socially constructed. helping students see this is a major challenge. a second key element of statistical literacy is the importance of context and confounding. consider these examples involving simple rates or percentages. • in terms of people, the number of unemployed workers is much higher in the us than in canada. but does this mean unemployment is more prevalent in the us than in canada? no. the number who are unemployed is strongly influenced by the size of the population. to untangle the influence of population on the number who are unemployed we need to look at the unemployment rates: the percentage of those in the civilian labor force that are unemployed. taking into account the influence of a related factor can reverse this association between country and unemployment. the point is that the context is important. are we talking about counts or rates? what should we be talking about? but to be statistically literate one must go further than simply forming and comparing rates and percentages. • suppose someone claimed that mexico has a better medical system than the us. you might be skeptical. but consider this statistic: the death rate is lower in mexico than in the us. while there may be variation in how some statistics are assembled, there is little variation in what constitutes death. so what explains this lower death rate in mexico than in the us? it can’t be a difference in the size of the populations; rates take that into account. it could be the difference in medical care. but another relevant factor is the difference in ages. mexicans are much younger on average than those in the us. younger people are less likely to die in the next year than are older people. so perhaps taking into account the difference in the average ages in the two countries may decrease – if not reverse – the association between country and deaths. untangling the influence of confounding is a major element in being statistical literate. see schield (2004) and the statistical literacy website.9 see also gray (2003) and lackie (2004). data literacy while all students in majors that deal with information have a need for information literacy, those students in majors that also require a course in data analysis or statistics typically have a need for data literacy. such majors are often found in the social sciences and business. in these majors, students need training in how to obtain and manipulate data. data literacy is supported by the international association for social science information services and technology (iassist),10 by the association of public data users (apdu)11 and by related organizations, such as the inter8 iassist quarterly summer/fall 2004 ‘information literacy’ (1,498 citations), ‘quantitative literacy’, ‘statistical literacy’ and ‘data literacy’ are in their academic infancies with less than 65 citations each in the education resources information center (eric) database.13 with the internet, statistical summaries are more readily available. so if students are to be information literate, the must be statistically literate. analyzing, interpreting and evaluating statistics as evidence is a special skill. statistical literacy must be an essential component of information literacy. statistics summarize data. the numerical value of a statistic is heavily influenced by how the underlying data is selected, converted and manipulated. converting and manipulating data is a special skill that requires in-depth training. thus, data literacy must be an essential component of both information literacy and statistical literacy. how one organizes these literacies depends on one’s perspective. consider a perspective of those teaching in the social sciences. figure 2: discipline perspective from this disciplinary perspective, students need to be able to analyze, interpret and evaluate social science data. data literacy is needed to access, manipulate and summarize the data. but statistical literacy is needed to guide in that process while information literacy sets the overall context for evaluating the sources of data and the appropriate manipulations. regardless of the arrangement, students need to be able to access, analyze and evaluate information, and they must be able to communicate their findings, conclusions and recommendations accurately and effectively. the key point is that information literacy, statistical literacy and data literacy are tied together by a common critical thinking analysis, interpretation, evaluation information literacy statistical literacy data literacy university consortium for political and social research (icpsr).12 for more background, see papers by rice (2001), gray (2003), hunt (2004) and czarnocki and khouri (2004). data literacy may appear less technical than is either computer science or management information systems (mis). yet students need to understand a wide variety of tools for accessing, converting and manipulating data. these may need to understand structured query language (sql), relational databases (e.g. ms access), data manipulation techniques, statistical software (e.g., spss, stata, minitab and ms excel) and data presentation software (e.g., ms excel and ms powerpoint). inter-relation with the advent of the personal computer and the web, information literacy requires both statistical literacy and data literacy. students must be information literate: they must be able to think critically about concepts, claims and arguments: to read, interpret and evaluate information. statistical literacy is an essential component of information literacy. students must be statistically literate: they must be able to think critically about basic descriptive statistics. analyzing, interpreting and evaluating statistics as evidence is a special skill. and students must be data literate: they must be able to access, assess, manipulate, summarize, and present data. data literacy is an essential component of both information literacy and statistical literacy. see linden (2002). figure 1 illustrates one way of viewing the relation between the three literacies from a critical thinking perspective. figure 1: critical thinking perspective journalism schools work at integrating these three literacies from this information-literacy perspective. see ward and hansen (1997), and hansen and paul (2003). unlike ‘critical thinking’ (10,116 citations) and social science data analysis, interpretation, evaluation data literacy statistical literacy information literacy iassist quarterly summer/fall 2004 9 set of problems and a similar level of approach. all three are more general than specific, they each involve interdisciplinary study and they deal with fundamentals. they can be useful to students in any major. as such they should be core elements in a college education. organizational pressures and needs colleges and universities are under increasing financial constraints. at times, staff are sometimes viewed as overhead whereas faculty are viewed as a cost of production. this pressure may tend to reduce funds available for library staff. yet, librarians have a unique opportunity in view of their training. they are generalists, not specialists. their focus is not the focus of a particular discipline. as such they are eminently qualified to teach students how to think critically, how to become information literate, how to become statistically literate and how to become data literate. in a teaching capacity they may obtain a different status in relation to the ultimate mission of higher education than they would as library staff. in helping students think critically about information, statistics and data, their role might be considered mission critical given the importance of critical thinking as a strategic goal of higher education and the difficulties some students have in thinking critically about words – much less about numbers. teaching statistical literacy librarians should consider teaching statistical literacy as a component of information literacy. while they may not have a quantitative background or interest, statistical literacy is typically more about words than numbers, more about evidence than about formulas. librarians have always tried to be of service to the needs of those seeking to access and understand information. data librarians should definitely consider teaching statistical literacy as a component of data literacy. they typically have a quantitative background in the social sciences or they have acquired one on the job. they recognize that students need help in accessing data and in evaluating data. and they recognize that random assignment in the social sciences is often impossible or unethical, so statistical associations are being used as evidence for causal connections. helping students to think critically about such inferences would be a valuable service to students in the social sciences. training teachers in statistical literacy will be a major effort. see watkins (2004) for a promising approach. would the teaching of statistical literacy by trained staff undermine the role of the professional faculty teaching traditional statistics? my answer is “no!” most faculty teaching statistics are not looking to teach anything ‘below’ statistical inference and advanced data modeling. i believe the opposite is more likely. collaboration with such faculty would be welcomed and fruitful to the extent one could show that statistically-literate students are better prepared to understand and appreciate what they are learning in their advanced quantitative courses. would resources be available for this teaching? my answer is “yes!” while such funding is unlikely under existing library or staff budgets, funding is almost guaranteed from the academic budget as soon as the activity involves academic credit. having statistical literacy, data literacy and information literacy as an academic course that might be required of all students – or at least of those who fall below a certain level of proficiency – will open the door to a new source of funding. from the academic budget perspective, paying instructors on an adjunct or overload basis is much less costly than paying a tenured professor. from the staff perspective, money from the academic budget might go toward the general staff budget, it might go to the staff member teaching the course (provided the staff member is still working full-time on non-teaching activities) or it might involve an allocation between these so as to provide some incentive for staff to teach this kind of course. to obtain academic credit in higher education, a statistical literacy course must be viewed as non-remedial. for one approach, see schield (2004b). developing valid statistical literacy assessments will also be critical. to see how all this might work, review the structure and operation of quantitative literacy centers, programs and courses at selected us colleges. see http://www.statlit.org/ql2.htm. conclusion both information literacy and data literacy should be expanded to include critical thinking and statistical literacy. expanding information literacy to include statistical literacy will help students deal with information that involves statistics. expanding data literacy to include statistical literacy will help students in the social sciences deal with inferring causation from associations. as such, including statistical literacy with information literacy and with data literacy will provide more opportunities for librarians to be of service in helping students think critically. see schield (2005b). acknowledgments to iassist for promoting data literacy. to dianne harmon, joliet public library, and to boyd koehler, augsburg college library, for recommendations on information literacy. this work was conducted under a grant to augsburg college from the w. m. keck foundation “to develop statistical literacy as an interdisciplinary curriculum in the liberal arts.”14 references american library association (1998). a progress report 10 iassist quarterly summer/fall 2004 on information literacy: an update on the american library association presidential committee on information literacy: final report. best, joel (2001). damned lies and statistics: untangling numbers from the media, politicians, and activists. university of california press. best, joel (2004). more damned lies and statistics: how numbers confuse public issues. university of california press bruce, christine (1997). the seven faces of information literacy. auslib press, 1997. 203p. paperback corti, louise (2003) “exploiting uk survey data sources for teaching political science: experiences from the classroom”, presented at 2003 iassist conference, ottowa. available at: http://datalib.library.ualberta. ca/~humphrey/iassist2003/sessions/i03sessions.html#a20 czarnocki, susan and anastassia khouri (2004). filling the gap: doing stats in the library presented at iassist 2004 conference, madison. available at: http://dpls.dacc. wisc.edu/iassist2004/program.html eisenberg, michael b., carrie a. lowe, kathleen l. spitzer (2004). information literacy: essential skills for the information age second edition. libraries unlimited; 2nd edition goad, tom w. (2002). information literacy and workplace performance. green wood publishing group, 2002. 248p. gray, ann (2003). data and statistical literacy for librarians. presented at 2003 iassist conference, ottowa. abstract available at: http://datalib.library. ualberta.ca/~humphrey/iassist2003/sessions/i03sessions. html#a20 hansen, kathleen a. and nora paul (2003). behind the message : information strategies for communicators. allyn & bacon. hunt, karen (2004). the challenges of integrating data literacy into the curriculum in an undergraduate institution. presented at iassist 2004 conference, madison. available at: http://scholar.uwinnipeg.ca/khunt/ iassist2004/index.cfm lackie, paula (2004). understanding and using data: a discussion of the jargon and trends in “quantitative literacy.” presented at iassist 2004 conference, madison. available at: http://dpls.dacc.wisc.edu/ iassist2004/program.html linden, julie (2002). finding, evaluating and using numeric data . presented at iassist 2002 conference, storrs, connecticut. available at: http://ropercenter.uconn. edu/iassist2002/program.html loertscher, david v. and blanche woolls (2001). information literacy: a review of the research: a guide for practitioners and researchers, 2nd edition.. hi willow research & publishing, 2001. middle states commission on higher education (2003). developing research & communication skills: guidelines for information literacy in the curriculum. 112p. moore, penny (2002). information literacy: what’s it all about? new zealand council for educational research (nzcer). nims, julia, r. baier, e. owen, and r. bullard (2003). integrating information literacy into the college experience. library orientation series, pierian press. rice, robin (2001). data support for learning & teaching: the final frontier? . presented at iassist 2001 conference, amsterdam. available at: http://datalib.ed.ac. uk/projects/datateach/iassist2001/tsld001.htm rockman, ilene f. and associates (2004). integrating information literacy into the higher education curriculum: practical models for transformation. the jossey-bass higher and adult education series schield, milo (1998). statistical literacy and evidential statistics. asa proceedings of the section on statistical education, p. 137. available at: www.statlit.org/pdf/ 1998schieldasa.pdf schield, milo (1999). statistical literacy: thinking critically about statistics as evidence. of significance, volume 1, issue 1. association of public data users (apdu). available at: www.statlit.org/pdf/ 1999schieldapdu.pdf. schield, milo (2004a). statistical literacy and liberal education at augsburg college. peer review summer issue, assoc. of american colleges and universities. available at: www.statlit.org/pdf/2004schieldaacu.pdf and www.augsburg.edu/statlit. schield, milo (2004b). statistical literacy curriculum design. curriculum design roundtable sponsored by the international association of statistical educators (iase) in lund sweden. available at: www.statlit.org/pdf/ 2004schieldiase.pdf and www.augsburg.edu/statlit. schield, milo (2005a). statistical prevarication: telling half truth using statistics. communicating statistics http://www.amazon.com/exec/obidos/search-handle-url/index=books&field-author=michael%20b.%20eisenberg/103-3890030-2839804 http://www.amazon.com/exec/obidos/search-handle-url/index=books&field-author=carrie%20a.%20lowe/103-3890030-2839804 http://www.amazon.com/exec/obidos/search-handle-url/index=books&field-author=kathleen%20l.%20spitzer/103-3890030-2839804 http://ropercenter.uconn.edu/iassist2002/temp/julie_linden_wkshp.html http://ropercenter.uconn.edu/iassist2002/temp/julie_linden_wkshp.html http://www.amazon.com/exec/obidos/search-handle-url/index=books&field-author=julia%20k.%20nims/103-3890030-2839804 http://www.amazon.com/exec/obidos/search-handle-url/index=books&field-author=randal%20baier/103-3890030-2839804 http://www.amazon.com/exec/obidos/search-handle-url/index=books&field-author=eric%20owen/103-3890030-2839804 http://www.amazon.com/exec/obidos/search-handle-url/index=books&field-author=rita%20bullard/103-3890030-2839804 http://www.statlit.org/pdf/2004schieldiase.pdf http://www.statlit.org/pdf/2004schieldiase.pdf iassist quarterly summer/fall 2004 11 conference sponsored by the international association of statistical educators (iase). available at: www.statlit. org/pdf/2005schieldiase.pdf and www.augsburg.edu/ statlit. schield, milo (2005b). statistical literacy: an evangelical calling for statistical educators. invited paper at conference sponsored by the international statistical institute (isi) in sydney. available at: www.statlit.org/ pdf/2004schieldiase.pdf and www.augsburg.edu/statlit. timms-ferrara, lois, (2003). public opinion matters: a new roper center program designed to promote classroom use of public opinion data. presented at the 2003 iassist conference, in ottawa, canada. abstract available at: http://datalib.library.ualberta.ca/~humphrey/ iassist2003/sessions/i03sessions.html#a20 ward, jean and kathleen a. hansen (1997). search strategies in mass communications. longman. watkins, wendy (2004). do it yourselves: a peer-to-peer approach to professional training. presented at iassist 2004 conference, madison. abstract available at: http:// dpls.dacc.wisc.edu/iassist2004/program.html notes 1 contact: milo schield, professor of business administration and mis, director of the w. m. keck statistical literacy project, augsburg college, minneapolis, mn. information available at: www.augsburg.edu/statlit. email: schield@augsburg.edu 2 available at: www.ala.org/ala/acrl/acrlpubs/whitepapers/ presidential.htm and at www.infolit.org/documents/ 89report.htm 3 available at: www.infolit.org/ 4 available at: www.ala.org/ala/acrl/acrlpubs/whitepapers/ progressreport.htm 5available at: www.ala.org/ala/acrl/acrlstandards/informatio nliteracycompetency.htm 6 www.ala.org/ala/aasl/aaslproftools/informationpowerinfor mationliteracystandards_final.pdf 7 see bruce (1997), loertscher and woolls (2001), goad (2002), moore (2002), nims et al (2003), rockman et al (2004), and eisenberg et al (2004). 8 2001 us statistical abstract. table 664, p. 435. 9 statlit website at www.statlit.org 10 international association for social science information services and technology (iassist). available at: www. iassistdata.org and http://datalib.library.ualberta.ca/ 11 association of public data users (apdu). available at: www.apdu.org 12 icpsr. available at: www.icpsr.umich.edu/org/ and www.icpsr.umich.edu/access. 13 education resources information center (eric). available at: http://www.eric.ed.gov/ 14 w.m. keck foundation at: www.wmkeck.org/ http://www.statlit.org/pdf/2004schieldiase.pdf http://www.statlit.org/pdf/2004schieldiase.pdf http://www.statlit.org/pdf/2004schieldiase.pdf http://www.statlit.org/pdf/2004schieldiase.pdf on-line or on tape judith s. rowe princeton university computer center princeton, new jersey (this paper was delivered at the 1981 ifdo/iassist conference, grenoble.) few of us here remember the early days in which data were transported from one installation to another usinq metal tapes or punched cards. needless to say, not many data sets were transported. metal tapes are now historical artifacts, although most of our computers can still read punched cards. but for almost twenty years data archives and data libraries have used magnetic tapes as the medium of choice for data exchange. i'lith the advent of microcomputers some exchange uses floppy disks, but since no standard has been established for these and since their capacity is quite limited, they are not widely used for this purpose. to a large extent, access to data by local users is also via magnetic tapes, although under certain circumstances data can be made available permanently or temporarily on disks. the decision to provide local access and certainly to provide permanent storage by one device or another is essentially one of cost and is based largely on the local charging algorithms for the storage and use of the various media. the two key variables are the size of the data set and the frequency of its use. although two computer centers may both be attempting to recover costs, they may not necessarily be doing so in the same way. charges are normally set to encourage particular types of user behavior. most centers, for example, offer cheaper off-hour rates in order to even out the flow of work. however, depending on the size of the machine and/or storage area, and the number and the type of input-output devices, charging may encourage the use of tape storage or disk storage, high density or low density tapes, permanent or mountable disks. at princeton it clearly makes sense — and will increasingly in the future as we eliminate mountable disks -to store small, frequently accessed data sets on permanently mounted disks and large, infrequently accessed data sets on high density tapes. however, it should be noted that in order to keep the number of tapes in the machine room manageable, there is a small penalty for archival tapes which are seldom or ever mounted. as the computer center's largest tape owner, the data library owns its own tape rack and hence does not pay the "no use tax" except for rented tapes in the public use area. with the present array of storage options now available to center users, decisions concerning the middle-sized data set may be less clear. princeton provides charts to help users pick the most economical storage medium for a given data set. these have now been augmented with a small, publicly available program which allows the user to insert data set size and anticipated use variables in order to obtain a storage medium recommendation. the charts, however, are useful if one is not in a gray area. 79 until recently our concerns with these issues were purely local ones, but with the increasing amount of data now available from on-line services, a new dimension has been added. it is no longer enough to think in terms of appropriate media for local storage and use. many of us are already faced -and all of us will be increasingly in the future -with new decisions on the form in which data should be acquired. these decisions are not easy ones, and they will not get easier in the future. there is no single answer which applies to every data archive or data library, nor is there a single answer which applies to every data collection. we cannot say, e.g., that princeton should buy tapes and the university of grenoble should use on-line services, although depending on the local environment some institutions may lean in one direction or the other. nor can we say that everyone should acquire data set a on tape and data set b through an on-line service, although there, too, we may all veer in one direction or another. in making these decisions there are three major components for consideration: our local environments, which although essentially fixed are different for each of us; our user communities, which may vary both among us and for each of us over time; and the characteristics of the various data products available to us, essentially the same for each but differing from product to product. it's a little like running a restaurant. we know what our local resources are -the size of the ovens, the number of tables -and generally speaking, we can find out what products are on the market and the characteristics and costs associated with acquiring or using them -fresh, canned and frozen meat, fish and vegetables. the big question is how many people will turn up for dinner. if we have lots of group reservations made well in advance, we can plan more easily; but if we have only a walk-in clientele, then we have only our experience and the complaints of the customers to rely on. nonetheless, in spite of these local and market variabilities, it is likely that in the future we will each find ourselves using a mix of "on tape" and "on-line" resources the remainder of this presentation focuses on the elements which will determine your particular mix and mine. let us first look further at local computing environments. although even within a given institution some of these elements may vary for different classes of users, by and large there is a pattern which characterizes each environment. the first and perhaps the most critical element in this pattern is the type and the amount of money available. in most computing environments, we find two types of money. the first is real money, money which can be spent as easily outside the institution as inside, money to which there are no strings attached. the second type of money is internal -general funds or "funny money" — money which may be available only for local computing (or, in some cases, for other local services) but which cannot be spent outside. in each institution both the data archive or data library and the individual user nave some combination of these moneys. the total amount of each and the amount of computing and ancillary services which the internal component can purchase are key in the on-line or on tape decision. if there is little real money, but adequate 80 internal money and cheap computing, the inclination will be to keep it local, to buy tapes. assuming, however, that the money issue is not so elastic, what are the other elements in the local environment which will affect our decisions? certainly the quality of computing is one. although the reliability of the system, the cost and availability of storage devices, and turnaround or response time are important, the quality and variety of available software and programming support are key. if it is not possible for the user to do locally the kinds of analyses required for his work, and on-line services to do these things are available, the pressure to use them will be great. last but not least is the quality of the data archive or data library staff and the systems available for acquiring, storing, locating and accessing data. if these are all well-organized, ceteris paribus , users may be less interested in paying real money and learning new systems to use on-line services. other considerations peculiar to each local environment may also be at issue, but those i have mentioned are always relevant. now what of the individual data products? the most fundamental question is: is there a choice? some data are available only on-line and other data only on tape. one may characterize data files as consumable, durable, or permanent. the consumable files are like food. one must replace them daily, weekly or monthly. these are files, e.g., which provide input to economic models where the presence of timely data is essential. timeliness is generally far more essential in business than in academia. these data are generally available only on-line, which accounts in part for the fact that 90 percent of the use of on-line services for accessing numeric data is by business and only 10 percent by academia. for the rest of the explanation we must work back to differences in local environments, to the differences in how commercial and academic users measure costs. durable files are like cars, refrigerators, or even husbands. they may be inert files or dynamic files in which the currency of the models is valuable but not essential. the availability of these files on-line is usually a function of size and generality. large, small-area aggregate files or special-purpose microdata files are less likely to be available on-line than files of national, state, or provincial annual time series data. permanent files are like fine paintings or oriental rugs. they have long lives and small groups of devoted admirers, but seldom enough admirers to justify their availability on-line, whereas their long lives make them good investments for purchase on tape. assuming one has a choice between on-line or tape access, what other characteristics of the data product should be explored? size, cost, and subject matter are certainly factors, as are the relative ease and appropriateness of access and use -i.e., the source of the data, the form in which they are provided, and the availability of software. if the file is small and inexpensive, can be purchased and delivered within the appropriate time, can be analyzed locally, is of a permanent nature and relatively general interest, buy tapes. this decision, however, will also be affected by expected use. for exam81 pie, during the life of the file will there be many users or few? will they require access to the entire file or only to a small subset? if the latter, will they each require a different subset? and what will be the nature of their use? do they require only a one-time display of data or will they be subjecting the data to complex and repeated analyses? can one purchase only a subset of the data on tape or is it necessary to purchase the whole collection? is there a subscription fee for on-line access, or does one pay per use? if the cost of on-line access is moderate and the user needs to display only a few numbers, then purchasing the tape may not be justified unless there are clear indications of high future use. i have identified many variables in the decision equation which we may someday construct. however, our experience is as yet too limited to assign values to those variables. at this point we can only make subjective judgments, based on our awareness of the relevant factors. each of us behaving rationally may make different decisions, because our computing environments and our user communities are different. moreover, the advent of cheaper, larger on-line storage makes these decisions dynamic ones. some data we may all acquire on tape. other data we may all access online. but for an increasing amount of data our decisions will be individual ones. there is no single answer. a detailed look at a few specific examples of available data which some of us have acquired or might consider acquiring may serve to illustrate the interaction o1^ local environment, user community, and data product characteristics in determining data access choices. i have tried to choose an international group for these illustrations, but you will forgive me if i speak more about the data i know best. the united states bureau of labor statistics (bls) maintains a collection of 100,000 time series covering such subjects as national labor turnover; industry, producer, and consumer price indices; and national, state, and area employment. this collection is known as labstat, and in bls parlance labstat includes both the data and the software for its analysis. many of us over the years have purchased individual time series or groups of time series from the bureau for the cost of approximately $100 per reel, and have used locally installed time series packages such as troll, tsp or sas/ets for data analysis. a few of us have used one or more of the selected series which a number of the on-line services make available with accompanying analytic software. recently lockheed, which specializes in bibl iograohic data files, has made available to its users almost half of the labstat data collection. the primary omission is the series relating to unemployment. lockheed acquires each new update as soon as it is released and provides access to it, using the software which is familiar to its bibliographic data file users as dialog. there is no provision for any manipulation of the data, but selected displays are simple and inexpensive. as we discovered at princeton in providing access to the united states decennial census data, there is a large group of individuals who use machinereadable data products to find numbers which are not available in print or microform or which are more easily accessible in machine-readable form. 82 lockheed has made the same discovery, and the "labor statistics (labstat)" file is already being heavily used. a united states bureau of the census file called "u.s. exports" was added at the same time; and it is likely that lockheed and its competitors will soon add similar data files to their available holdings, and that they will concentrate on the display market rather than on the analysis market. for those of us who wish access to the full data base and to analytic capabilities, it is still necessary to purchase tapes. the international monetary fund produces several major data collections, including the world financial series (wfs) and the direction of trade. the former contains over 40,000 time series of financial statistics on over 160 countries, plus aggregate data for the world and over 50 regions. these data go back to 1948 and have been updated monthly since 1965. the direction of trade contains approximately 100,000 time series of import and export statistics for 230 nations and their trading partners. these data also go back to 1948 but have been updated monthly only since 1977. monthly tapes for ifs may be purchased directly from imf for an annual subscription of $400. icpsr members may obtain both collections on a less timely basis without cost. regardless of the source, when the time series package arrives it is necessary to use a special cobol program to unpack it and write it on disk, an assembly program to reformat it to tape, and a fortran program to retrieve it for analysis before using the data. data resources inc. (dri), possibly the largest of the commercial on-line sources of economic and other statistical data, adp network services inc., fri information services limited, rapidata inc., and telesystemes-eurodial , among others, provide subscribers with access to these data and to software for analyzina them. at princeton the international finance section of the economics department has adequate general funds money for computing and a graduate assistant to preprocess the tapes and to analyze each month's data. as a result, on-line offerings have not proved attractive. dri also provides united states current population survey annual time series data from 1968 to the present. these time series are based on the aagregate data published by the bureau of the census. for individuals interested in cps data in this form, there is no competing product. however, for individuals interested in cross-sectional microdata and specifically in the monthly supplement questions, the bureau sells each month's file for under $300. the quality of the files themselves and of their accompanying documentation has improved markedly in recent years, and normally the data are easily analyzed if software for processing hierarchical files is locally available. the popular march file, known as the annual demographic file, is available from icpsr. a somewhat unique case is the 1970 united states census of agriculture. a preliminary tape containing a subset of the preliminary published aggregate data is sold by the bureau at its customary $110 per reel for 5 reels. at the request of some members of the user community, final data were produced on tape as a special tabulation and sold for $1000. in the most recent directory of onl ine data bases , published by cuadra associates, an organization called on line research was reported to be making census of agriculture data available on-line. further investigation, however, uncovered the fact that this organization is no longer in business. in this instance it would appear that the 'display-only' customer will do well using the printed reports, and that 83 the researcher whose data needs are not excessive may find it more economical to reenter the necessary data himself -particularly if student or clerical time is freely available. this latter alternative is not one i would normally encourage, but it should not be totally overlooked. it may sometimes prove to be the most practical alternative. in discussing on-line access i have not distinguished between commercial vendors and academic sources. the issues are essentially the same. when should data be acquired and used locally, and when should it be accessed online from a remote site? we must each make our own decisions, and perhaps when we gather together next time, we will have a larger body of experiences to share. i realize that many of you had hoped for more specific recommendations on these issues. however, as with so many issues concerning the management of data archives and data libraries, we are each bound by our unique institutional settings and by the peculiar nature of our user communities. in the case of our local environments we must consider the amount and the type of money available for computing, the cost and the quality of local computing, the software available and supported, the amount of cheap labor, and the nature of our own operations. in terms of the user community we must consider the number of anticipated users and uses, the type of use, and the amount of data required. and finally, in looking at each data collection we must be aware of the options from which we may choose: the size of the file, its durability, its content, its completeness, its currency, and the software required and available for on-line or on tape use. 84 iassvol194 12 iassist quarterly abstract this paper documents progress on setting up a scottish household migration monitor and, in so doing, traces the path of a learning curve which operated from the inception of what came to be known as the migration and housing choice (scotland) survey through to the collection, processing, analysis and interpretation of the data it generated. hosted by edinburgh university data library, such a monitor will be available to the academic community, and to public and private sector housing and planning agencies. an innovative feature of the monitor is the inclusion of information on migration motivation as well as the numbers and patterns themselves. the inclusion of such qualitative as distinct from quantitative data posed questions about how best to handle both data types. it is an important issue to resolve since better informed decision-making is made possible as a result of documenting and understanding the decision to migrate and how it varies in space and time. for example, urban and regional planning authorities will be able to estimate housing demand and future housing land requirements more accurately. introduction the migration and housing choice (scotland) survey was conducted in the early nineties by researchers from two scottish universities; strathclyde and napier. the purpose of the survey was to discover the intentionality, namely, the motivation behind household migration patterns, to use the data for academic research and to inform decision making by urban and regional planning agencies in both the public and private sector. we regard the work done to date as a large pilot study for what may become an ongoing scottish household migration monitor. this paper contains descriptions of the following: * context of the survey * conduct of the survey * data handling issues * the role of edinburgh university data library * assessment of strategy and future plans context of the survey of the three components of population change, namely fertility, mortality and migration, it is migration which is the most significant, whether for facilities planners in the public sector or for market researchers in the private sector. with the decline of both fertility and mortality in recent decades, migration has become the primary process responsible for changes in population numbers and composition, with that importance magnified at the local scale (champion, 1993). however, migration researchers may experience greater problems than researchers into fertility and mortality when they attempt to use existing data sources relating to their respective fields of study. official data agencies release data at levels of aggregation which, on the one hand, protect confidentiality, but on the other hand, hide much migration activity. the following three diagrams illustrate how a move ‘on the ground’ becomes recognised as a migration only after it crosses a specified boundary, ie, is externalised. because the crossing of a boundary has become the defining feature of a migration, the importance of distance travelled is underplayed. for example, the moves depicted by the two horizontal arrows in figure 1(b) would both register as migrations, regardless of their markedly different lengths. yet the diagonal arrow in the same quadrant, although the same length as the longer of the two horizontal arrows, would not count as a migration at all, even although it is more than twice the length of the shorter horizontal arrow. only at the level of disaggregation shown in figure 1(c) does the long diagonal arrow register as a migration. however, even at this level the two short-distance moves in the bottom left quadrant do not count as migrations. thus, not only is some migration information concealed altogether, but also distance information is lost, which is developing a scottish migration monitor: a co-operative approach by alison mccleery* & emma forster department of economics, napier university heather ewington & peter burnhill edinburgh university data library, university of edinburgh 13winter 1995 significant because short-distance moves indicating residential mobility are more important numerically than long-distance migration and further, that the motivation behind each type of move is different. the solution suggested by forbes and mccleery (1991) is, in an ideal world, to attach a specific geographic reference to the data at the point of collection which would then permit analysis of all moves and provide the flexibility to aggregate the data to whatever level, with due regard to confidentiality. migration researchers in scotland are more fortunate than most in having access to the register of sasines1. the sasines register is a land register, unique to scotland, which records every private property transaction and therefore allows every moving household in the private sector to be identified. the information contained in each record is the name and address of the owner of the present property, a previous address (of questionable accuracy see below) and the price of the sale. used judiciously, the register of sasines provides a reasonably accurate record of the origin and destination of moving households by precise postal address. tracking of movements is theoretically possible at the lowest geographic level, namely the household. however, in practice, although the present address is accurate because it is recorded from legal documents relating to the sale of that property, the previous address information is collected mainly for administrative reasons. it may refer to the previous long-term address, but equally could refer to temporary accommodation occupied prior to the move or might even be the new address if there was a delay in completing the legal documentation. as mccleery (1980) emphasises, as a source of migration data, the register of sasines has to be used with care. furthermore, even if it is possible to produce accurate patterns of moves on the map, it should be recognised that these are merely the visible trace of an invisible, complex, increasingly segmented and largely unexplored household decision. it is the accumulated decisions of all the moving households which drive the migration streams. without some knowledge of these decision processes, and how they may vary from place to place, from time to time, and from household to household, it is impossible to estimate future volumes and directions of moves. thus is introduced an unknown factor in estimates of housing demand, and consequently housing land requirements are difficult to assess. although to date, some research has been carried out investigating movers in scotland (garner, 1980; forbes, 1989), there has yet to be a comprehensive survey of all tenures at the national level. forbes’ paper examined migration patterns in scotland’s largest city, glasgow, but was more innovative than either descriptive or analytical in that it used pre-existing data to explore the feasibility of an experimental migration monitor. now this present paper reports on the broadening of the initial work into this area. in so doing, it traces the path of a learning curve which operated from the inception of what came to be know as the migration and housing (scotland) survey through to the collection, processing, analysis and interpretation of the data it generated. conduct of the survey the survey’s two principal investigators were jean forbes of the centre for planning at strathclyde university in glasgow and alison mccleery of the department of social sciences at napier university in edinburgh. building on her previous work, the current project was initially designed by forbes to explore the various elements in the decision to move within the west of scotland. the subsequent involvement of mccleery allowed coverage to be extended to include the whole of mainland scotland. the data library, located at the neighbouring university of edinburgh, was brought in at a later stage to introduce a greater element of economy and efficiency to the questionnaire design and subsequent data processing. figure 1 (a) 1 registered migration (b) 1 + 3 migrations (c) 1 + 3 + 4 migrations 14 iassist quarterly the survey data was gathered by means of a postal questionnaire mailed to a 25% sample of households having moved between january and october 1990, identified from a computerised version of the sasines register held by the land value information unit at a fourth scottish university, paisley university. the unit is is a commercial concern and deals with requests for data from organisations such as chartered surveyors, builders, local authorities and researchers. computerised housing sales records for the whole of scotland are available from 1989 onwards and, for some parts of the country, from 1979. mailing of the questionnaire was made possible with the co-operation of a number of local government planning departments2 whom the investigators had succeeded in interesting in the project. the royal mail (post office) also helped because they had recognised that the survey would collect useful information about public awareness of the postcode3; the questionnaire asked respondents to provide both their previous and present postcodes. the design of the questionnaire was informed in two ways; firstly, it was influenced by the elements in an a priori model of the process of decision-making developed by forbes (1989) and secondly by elements related to housing supply and local environmental quality proposed by the collaborating planning departments. in essence, the questions were divided into 4 types: 1. question initiating open-ended answer as text where was your previous house? 2. question with pre-defined answer categories, only one answer per question what type of house? detached [ ] semi-detached [ ] terraced [ ] flat* [ ] *apartment 3. variation on the above, essentially one question, but with each possible factor presented as a separate question with only one answer per question, either yes or no what factors influenced your decision to move? job transfer [ ] 4. question inviting multiple, open-ended answers as text which other localities did you consider? during the course of the project, the questionnaire was enhanced as a result of further consultation with the planning authorities from whom we sought help with distribution of the questionnaires. variations between versions were minor, consisting mainly of additional questions, e.g. including a question about car ownership and a modification allowing respondents to specify other reasons for leaving the previous house and choosing the present house, other than the pre-defined reasons given on the questionnaire. the inclusion of the latter option allowed respondents to provide information in their own words about their motivation for moving. the challenge later was how best to translate these subjective snippets of information into objective data. apart from these additions, consistency was maintained to allow comparability of the core questions. as expected with a postal survey, the overall response rate fell short of 50%, although this varied decidedly from district to district. the resultant 10,006 cases represented about 1 in 10 of all private sector residential movers for the period. a proportion of strathclyde respondents volunteering their telephone numbers was also followed up with a more detailed survey of search patterns. data handling issues and the role of edinburgh university data library data collection the survey data were collected and processed at different times by different institutions. a pilot study was conducted first on one area of strathclyde region, the sprawling, industrial city of glasgow. the second, and larger survey covered six regions; the rest of strathclyde, dumfries and galloway, fife, grampian, tayside and central. forbes at strathclyde university organised data collection for both surveys and because of the large number of replies, considered it cost-effective to use an outside bureau agency to process the data. the two regions of fife and highland, took on responsibility for collecting and converting their own data and, with assistance from eudl, mccleery collected and processed data from the 15winter 1995 remaining two regions, lothian and borders. the smaller number of the east coast replies, together with financial constraints, prompted the investigation of an in-house solution to processing the data. the advantage of this approach was that we gained greater control over the data, but we still had to ensure that it was in a format consistent with the bureau-processed data which, because handled first, was regarded initially as the standard to follow. the west coast data had been converted into rectangular numeric data files and processed as spss system files. although strathclyde university operate a vax system and edinburgh and napier universities operate unix, we did not envisage any problems with data transfer between the two systems. in-house processing on advice from eudl, it was decided to use the database package filemaker pro to input and store the east coast results. eudl had used the software successfully on a number of other projects and thus was able to confirm its suitability for the purpose. in particular, it was a user-friendly package which a new keyboard operator could quickly learn how to use. the following positive attributes apply: * database definition is simple * the facility to construct different views of the database (layouts) allowed the construction of a data input screen which closely resembled the layout of the questionnaire we reckoned that the similarity between the questionnaire and input screen would reduce the number of input errors * field display options such as checkboxes, radio buttons, pop-up menus and pre-defined lists helped the operator to input the data efficiently * the lookup facility enabled automatic coding of some data items additional fields were created for each question requiring either a yes or no answer and the fields automatically filled with 0s and 1s depending on whether or not the checkbox was clicked * data validation mechanisms ensured that no invalid answers were input (e.g. text replies instead of numeric) and that no questionnaire could be entered into the database twice * the export facility provided output usable with spss. unlike the bureau processed data, we exported tabseparated output files and created free as opposed to fixed data format spss files, thus allowing us to store variable length data items, such as the postcode assessment of strategy overall, filemaker pro matched expectations. inputting data was relatively trouble-free and early results were obtained from the database prior to input and analysis within spss. it was useful to have these results to compare with the spss results to ensure that the data had transferred correctly between environments. also, we were able to input all data from the questionnaires into the filemaker pro database, both numeric and textual, whereas not all items from the west’s replies had been converted into machine-readable form. figure 2: the administrative map of mainland scotland key to regions 1 borders 2 central 3 dumfries and galloway 4 fife 5 grampian 6 highland 7 lothian 8 strathclyde 9 tayside 16 iassist quarterly specifically, postcode information had not been transcribed. although this information could still be retrieved from the questionnaires, it would be a time-consuming exercise and perhaps now, as the survey data ages, considered not worth the effort. eudl recommended that the postcode be recorded in our data because many official statistics are postcoded, including the population census, and thus the code would provide the link with other important datasets. the last census for the uk was carried out in 1991, shortly after our survey was conducted, and so we were keen to compare our results with those of the census. also, there is no machine-readable version of the exact text of the answers to the open-ended questions. a coding scheme for these answers had been devised and applied to the west’s data, but when it came to using it to code the east coast replies, we felt that a more detailed schema was required. and so we enhanced the original coding and applied it to our data. this has mean’t that some of the data has a more detailed level of classification for these variables and again, only by returning to the west’s questionnaires would be able to standardise the treatment of this data item. in sum, we think that we benefitted by building the filemaker pro database; we were able to stagger handling of the data, processing the less problematic data items first, producing some results and returning to the more difficult items (such as the open-ended questions) as and when time and resources permitted. the spatial element a fundamental, and what has proved to be an on-going operation has been processing the survey’s spatial data items; ie, the locality of previous and present houses, location of places of work and other areas searched. we wanted to attach a national grid4 reference to each of these items and be able to calculate, for example, the distance moved from previous to present house, how far respondents travel to work and the extent of their search area. also, with grid referenced data we would be able to use gis technology (e.g mapinfo ) and plot migration movements. the planners in particular had expressed interest in maps of our findings. however, for the purposes of this study it was decided not to record the address data from the sasines records. instead, having used the sasines successfully to identify the migrating households, we included questions in the questionnaire asking in which locality the respondent previously and presently lived. at the time, there were cogent reasons for this course of action. firstly, at the start of the pilot project, the team was not aware of an easy and quick way of grid referencing the precise postal address. secondly, there was the problem of potential inaccuracy of the previous address recorded in the sasines record. and thirdly, in this particular migration project, sensitive motivational migration information might not be revealed by respondents nervous in the knowledge that it was attached to their actual postal address. in retrospect, this decision was a mistake. if the survey were being repeated, we would pay more attention to the geographic detail and fully exploit the resources of the sasines, trying to overcome the confidentiality issue in a different way. despite the shortcomings of our spatial data it still had to be processed. although the first batch of replies was grid referenced using printed information sources, with the subsequent involvement of the data library, we learned that online grid referencing facilities were available which not only would automate a laborious task, but also provide sufficient geographic detail to produce meaningful maps. we used the following data library utilities to grid reference our data: * postzon file the postzon file provides 12 digit grid references to 10 metre resolution. the file is extracted from a central postcode directory maintained by the post office and contributed to by a number of official organisations such as the office of population, censuses and surveys, the general register office (scotland) and the welsh office. in addition to grid references, it contains information about postcodes (date of termination of a postcode) and area and country codes. * index of placenames (ipn) both datasets can be used interactively and in batch mode; the latter was the appropriate method for our application. a filemaker export file was produced consisting of, where it existed, the postcode and placename information for each reply, together with each reply’s unique reference number. the more specific match was on the postcode and so we used the postzon file first and only if unsuccessful, tried to match the placename using the ipn. both methods however were an improvement on a grid reference obtained from the printed sources in terms of speed of processing and the accuracy of the grid reference. 17winter 1995 data linkage methodological concerns the original decision not to record postcode information in the data from the first two surveys was taken partly because of an initial lack of awareness of its use to link with other datasets. it also related to doubts about its use in deriving a geographic reference; in particular, there was concern about the method by which a grid reference is allocated to a postcode. the practice of the general register office (scotland) has been to allocate the reference to the nearest 10 metres of the centre of the building judged by eye on the map to be the centroid of the area covered by the postcode. as forbes and mccleery (1991) have argued, postcode units vary widely in size and shape between different parts of the country. they are smallest in urban areas, but in sparsely populated highland scotland, they are large and straggling, and thus the process of allocating a centroid is a nonsense; subsequently to attach a 12 grid figure reference “lends spurious accuracy to a technique which is at the very least unsophisticated”. it could be that the effects are marginal. but a more pessimistic view cannot be dismissed: massive amounts of public money have been directed in recent years to areas of social and/or economic ‘need’, such areas being defined by mapped patterns of aggregated data. yet the need may not exist where the map says it does, and the map may fail entirely to reveal a need locality which does exist forbes and mccleery (1991) this draws attention to problems of spatial comparability. in addition there is also arguably a problem of comparability through time since the royal mail allocates and re-allocates postcodes as the built environment evolves, that is, as buildings appear and disappear. despite these acknowledged failings, the postcode is, as explained above, currently the best tool there is for achieving data linkage. data linkage practical applications as things turned out we had an earlier occasion to use the linkage facility; namely, to overcome the problem of the east’s unavoidably small sample. we had hoped to issue statistics for specific places, but in many cases the number of replies was too few to draw any meaningful conclusions and we were also wary of breaching confidentiality. although, all replies had been labelled by their district and regional identifier, we considered that statistics produced for these large and heterogeneous areas would be not very meaningful. fortunately, suitable sub-areas were identified by the planners using their local knowledge. they identified community council areas (ccas) which were largely homogeneous in terms of socio-economic composition and also constituted possible housing market areas. the ccas had been defined in terms of the geographic areas used to release the population census data, namely the census output areas, which in turn comprise one or more unit postcodes. our task therefore was to discover in which cca each reply was located. if we had had accurately grid referenced replies and if the ccas been available in digitised form, then the task could easily have been performed using the gis technique, ‘point-in-polygon’. however, we were not in this fortunate position and so we had to rely on another of the data library’s data services, ukborders. ukborders is a national, online service providing digitised boundary data for standard and user-defined areas based on the geography of the population census. by combining data from ukborders and information from the planners on the composition of the ccas, we were able to construct a lookup table detailing the cca and census output area for every postcode unit. the final step was to match the replies against this file and create two new variables of cca and census output area. we then re-released the earlier statistics produced for the newly-created areas. overview a project which started life as a rather modest example of legitimate, but by nature organic, academic enquiry in the best tradition of the ancient scottish universities, had slowly but surely changed into a monster. it was not so much nessie as messy! this situation arose mainly from the project’s funding arrangements. in short, there were no funds available at the start of the project and it was able to progress only because of personal investment by forbes who financed a small-scale, pilot study on glasgow, with the hope that the rest of the country could be surveyed in time as and when funds became available. it was fortunate that we were able to interest a number of organisations outwith the academic community in the project, who agreed to contribute either financial assistance or help in kind, e.g. posting the questionnaires. however, in return, we were obliged to consult with them about the content of the questionnaire and to provide them with basic results. the effect of having multiple contributors was both positive and negative: positive insofar as the survey would never have been carried out 18 iassist quarterly without their help, negative in that control of the project became dispersed, with the danger that the initial objectives of the survey would be lost as we tried to be all things to all men. it was unfortunate that eudl was not involved at the start of the project. their assistance with questionnaire design, data input and processing was important and much appreciated but, because of circumstances, consisted mainly of rectifying errors that might have been avoided if there had been better-informed planning at the start. ideally, their input should have been sought at the beginning, a project plan developed and adhered to throughout the project. however, we do regard the survey as a success and are currently collaborating to produce a national dataset to be deposited with eudl and thus be available for use by other researchers. dissemination through the internet eudl has also drawn attention to the potential use of the internet and the world wide web for both publicising the data and possibly providing access. we have created a few pages about the survey and have added them to the data library’s world wide web server, datalib. also, we have produced a wais (wide area information server) index of records about the survey questions, including details of variable and value labels, whether or not the variable exists for a particular area and how it was processed or derived. the organic growth of the dataset had given us the problem of where to record details of its changing structure, but fortunately, the web arrived during the course of the project and offered us a possible solution to the documentation problem. we thought it useful to let potential users view the questionnaire as an aid to helping them decide whether or not the data would be of value to them. we created created a gif image of the questionnaire using eudl’s scanning facilities. eudl is involved in a number of projects involving scanning. currently they are working with the university library to offer internet access to part of the library’s special collections by creating an index to scanned images of catalogue records. they are using a flatbed scanner controlled by the optical character reading (ocr) software formfile. we have created a wais index of information about the survey variables using the indexing software, freewais. users are prompted to search the index by keywords. additionally, they can view a question list in which each entry is a hypertext link to the appropriate index record. in addition to the above information, we are also investigating using the web to access the data itself directly and produce basic statistics for standard areas. we have noted other attempts to do this, including by the university of british columbia and statistics canada. however, we realise that direct access to our data may require to be limited because of the confidentiality issue. for experimental purposes we intend to hardwire into the system, a selection of counts for the larger geographical areas. conclusion although subject specialists provided the figure 3: home page the home page reproduced here lists the four main menu options: * about the survey * view the questionnaire * list questions * search index of questions 19winter 1995 initial stimulus for this investigation and are predominantly involved in the interpretation and exploitation of results, they did not have the information-handling skills or resources necessary to create the dataset from which the results would be derived. the involvement of data management specialists, edinburgh university data library, was therefore an important factor in the successful conduct of this project. despite the many flaws, results have been produced which have indicated potential further lines of enquiry. also, the nonacademic organisations who were initially involved have reacted well to our findings. furthermore, having been approached by a number of private sector organisations such as house builders and the government housing agency, scottish homes, we now feel confident in our ability to convince a consortium of these bodies to put together the funding for a comprehensive and efficiently organised migration monitor. during the course of the survey, vital lessons have been learned for the future. should the survey be repeated in a more favourable funding environment attention should be paid to the following issues: organisational * of paramount importance is the need to agree at the outset what the survey is about; then to ask questions which elicit the correct answers, that is, answers which provide useful information methodological * we have highlighted the issue of spatial referencing, that is, problems associated with the use of the postcode as the uk standard. frustrating though the postcode may be, nevertheless, by a process of historical accident, it has become firmly embedded in the data processing activities of uk official data collection agencies finally, two issues; one specific, the other more general arising from the foregoing discussion. firstly, the present investigators have restricted the use of the data generated to supporting the activities of the organisations which contributed funding to the project. however, with the deposit of the dataset with the data library, the data will become more widely available and used by researchers who may potentially adopt a more challenging attitude. for example, the widely accepted use of the housing market area (such as the cca supplied to us by the planners), could well be rejected. we have become aware that we might have been too much in the pocket of the planners. however, this is a dilemma faced by many in the uk with the shift during the 1980s from publicly-funded academic research to commercially-sponsored consultancy work. the question we may need to ask is this: what will happen if, with the wider release of the data, other users end up biting the hand that fed us! secondly, there may be a fundamental mismatch between the perspective of data professionals and the nature of academic enquiry. the former is rational, organised and has to adopt a pragmatic, problem-solving approach. the latter is organic and incremental, with a tendency to go off at tangents, some of which open up new and fruitful avenues of enquiry. for example, as we analysed the respondents own answers about why they moved house, the issue of health arose frequently; an option not given as one of the multiple-choice answers. figure 4: index of questions 20 iassist quarterly issues of this type point to one of the many benefits of an organisation such as iassist: it serves as a forum where data professionals meet to consider all aspects of the data provider/user interface and where subject specialists are welcome to learn from and contribute to the discussions. appendix migration and housing choice (scotland) survey: some key results these testify to the increasing complexity of motivation for migration. in other words people are citing more and different reasons for their move than previously. respondents gave a mean of 1.48 reasons for leaving their previous address, but 2.64 reasons for choosing their new one. while this confirms the operation of a trigger for leaving the old home, it also suggests that at the level of choosing, the decision is based on multiple factors. furthermore, it is apparent that a hard and fast distinction between pushes and pulls (i.e. from the origin or to the destination) is no longer entirely valid. for example, in response to a partner’s complaints that the present home is too cramped, a person may seek promotion at work and be moved to a branch manager’s position in a distant location. was the household pushed, pulled, or a bit of both? perhaps this is not a particularly well-chosen example, however, since another key finding relates to the declining significance of employment as a motivational factor relative to quality-of-life considerations. while short-distance moves have never traditionally been predominantly job-related (as might be expected), now, apparently, this is increasingly the case also with long-distance moves. this is only partly explained by the increase in retirement-related moves associated with the changing age structure in the uk and the western world universally. it also seems to reflect a change in the balance of importance between macro-structural and micro-behavioural determinants of or more properly influences upon migration. increasing segmentation of markets for goods and services mirrors precisely the growing profusion of life choices and chances. the consequent diversification in ‘lifestyle’ is itself associated with an almost complete disintegration of any remaining correlation with household income. so it is that classifications such as ‘the elderly’, ‘the average family’, ‘single-person households’ or even ‘multi-adult households’ are increasingly superficial and meaningless. elderly people may be active, semi-independent or dependent; they may be young elderly, elderly or very elderly; they may live alone whether as a result of being never-married or widowed or with a spouse, sibling, offspring or companion. their incomes, tastes and aspirations within each sub-group may vary in as many ways and more. a print-out of comprehensive cross-tabulations for the elderly group alone would make serious inroads into british columbia’s forests! yet for all that people attempt to carve out a very specific spatial and social niche for themselves according to their particular circumstances, there exist certain quality of life variables which, taken together, identify a highly desirable local environment which, other things being equal, most people would seek after for europeans perhaps to live in geneva, for canadians, vancouver? nor is it even necessary to compromise on lifestyle choices for the sake of a high quality-of-life, if the two are pursued at different scales. measurement of quality-of-life is, as suggested above, appropriate at the level of inter-urban comparison; satisfying individual lifestyle requirements is carried out at the intra-urban level. yet the two cannot be completely divorced from each other as the following example perhaps indicates. in the glasgow university quality-of-life ranking of the thirty-eight largest cities in britain, edinburgh came out top. in our survey fewer of the respondents moving either within or to glasgow cited liking the local environment as a reason for choosing their house than was the case in edinburgh (56% to 64%). evidently longstanding rivalry which has traditionally existed between glasgow and edinburgh is not yet dead, and the perception of glasgow as a declining industrial city with problems of urban decay and deprivation persists. edinburgh, by contrast, is associated with the successful financial services industry and with fine architecture and gracious parks. moreover, our survey found that it is precisely these types of quality of life attributes which seem to be accorded higher significance by many working people than proximity to their place of employment, the latter for the present being considered a stretchable link. some interesting relationships also emerged between house price and distance moved and between housing density and distance moved. a higher proportion of those moving into the most expensive properties moved either longer distances or intermediate distances. this could be interpreted as a sequential process of initially a coarse-grain employment-led interregional move and later a fine-grain environment-led intra-regional move. the latter is less constrained by the requirement to be close to work and therefore permits a wider search area and a longer move locally than in the case of the lower paid who cannot afford the cost of commuting. this speculation is supported by the finding that a lower proportion of households migrating intermediate distances mentioned convenience to work as an influence upon their choice of destination. 21winter 1995 finally, it was not surprising to find that in urban areas the typical move is very short, although this is offset by a small proportion of very long extra-area moves. less well appreciated is the opposite situation in rural areas, where the typical intra-regional move is longer, presumably because of the longer distances between localities offering a comparable level of housing choice. thus, in the very rural district of argyll, 62% of movers have travelled 30kms or more, as against 18% for glasgow, scotland’s urban heartland. the short distance move figures are the reverse, with 29% for argyll and 54% for glasgow. it would be possible to produce a phd on the interpretation of the results from the housing choice (scotland) survey. indeed this is precisely what one of the authors of this paper, emma forster, is currently undertaking. in so doing, she is transforming herself into that still rare breed of highly-skilled and highly-prized individual who is both a knowledgeable subject specialist and a competent data manager. * paper presented at iassist95 may 1995 quebec city, quebec, canada. bibliography champion, t. (1993). ‘introduction’ in t. champion (ed.). population matters: the local dimension. london : paul chapman publishing ltd, pp. 1-21. cansim data base: canadian socio-economic information management system [computer data]. ottawa, ont.: statistics canada [producer and distributor], [19—]. 1 data file and accompanying documentation. http://www.datalib.ubc.ca citibase, fame economic database 1946[computer file]. new york : fame, 1978-1994. http://www.datalib.ubc.ca datalib: edinburgh university data library’s www server [computer data]. edinburgh: data library, university of edinburgh, 1995http://datalib.ed.ac.uk filemaker pro [computer file]. filemaker pro 2.obv2, may 1993. computer program. santa clara, california : claris corporation. copyright 1988-1993 claris corporation. forbes, j. (1989). ‘migration monitoring and strategic planning’, in p. congdon & p. batey (eds.). advances in regional demography. london & new york : belhaven press, pp. 41-57. forbes, j. & mccleery, a. (1991). the 1991 census: spatial referencing considerations. esrc regional research laboratory for scotland working paper no. 21. edinburgh: esrc rrl scotland, 1991. formfile [computer file]. version 1.11. livingston, edinburgh : seel ltd, 1994. freewais-1.0-sf [computer file]. release 1.0 2/16/93. thinking machines, jim fullton, kevin gamiel, jane smith, tung huynh, ulrich pfeifer. garner, c.l. (1980). residential mobility in the local authority housing sector in edinburgh 1963-73, unpublished phd thesis, university of edinburgh. ipn [computer file]. computer program. edinburgh : edinburgh university data library, 1991. mapinfo [computer file]. mapinfo version 3.0.2. troy, ny : mapinfo corporation. copyright 1985-1994 mapinfo corporation. mccleery, a. (1980). the register of sasines as a source of migration data, british urban and regional information systems association newsletter 46: 16-17. postzon [computer file]. edinburgh : edinburgh university data library, 1991. rogerson, r. findlay, a, morris, a, paddison, r. august (1989). in cities. variations in quality of life in urban britain. pp. 227-233. the scottish national dictionary. edited by william grant and david d. murison. edinburgh : the scottish national 22 iassist quarterly dictionary association limited, 1952. ukborders: esrc national online service for the extraction of digital boundary data [computer data]. data library, university of edinburgh. http://borders.ed.ac.uk notes 1. sasines is an old scottish word meaning the act or procedure of giving possession of feudal property, until 1845 by the symbolical delivery of earth and stone on the property itself. symbolic delivery has now been abolished and all sasines are registered in the register of sasines which is now being converted into a computerised land registry. 2. at present, scotland’s local government administration is divided into two tiers; one level of 9 regions and a second level of 56 districts and 3 islands areas. planning responsibilities are shared between region and district, the former involved in strategic planning issues, the latter in development control matters. however, by the end of 1995, this two-tier arrangement will be replaced by one level consisting of 32 unitary authorities. 3. the primary purpose of the postcode is to assist the post office to deliver mail. the postcode is a combination of between five and seven letters and numbers which define four different levels of geographic unit; the postcode area (120 in the uk), the postcode district (2,700), the postcode sector (8,900) and the postcode unit (1.5 million). 4. the national grid is a reference system of squares overprinted on all ordnance survey maps since the 1940s. the system of breaking the country down into squares allows any place in the country to be given a unique reference code. 5. crofting is a system of land tenure. a croft is a smallholding worked by a tenant, comprising a plot of arable land attached to a house and a right of pastorage in common with others. the sale of crofts has traditionally been governed by crofting law which restricted free sale. the western isles and orkney and shetland were not included in the survey because of complications associated with the unique form of housing tenure called crofting which is peculiar to these areas and which distorts the housing market there. 1 steeves 1/14 steeves, vicky; rampin, rémi and chirigati, fernando (2018) using reprozip for reproducibility and library services, iassist quarterly 42 (1), pp. 1-14. doi https://doi.org/10.29173/iq18 using reprozip for reproducibility and library services vicky steeves1, rémi rampin2, and fernando chirigati3 4 abstract achieving research reproducibility is challenging in many ways: there are social and cultural obstacles as well as a constantly changing technical landscape that makes replicating and reproducing research difficult. users face challenges in reproducing research across different operating systems, in using different versions of software across long projects and among collaborations, and in using publicly available work. the dependencies required to reproduce research can be exceptionally hard to track – in many cases, these dependencies are hidden or nested too deeply to discover, and thus impossible to install on a new machine, which means the percentage of reproducible research remains low. in this paper, we present reprozip5, an open source tool to help overcome the technical difficulties involved in preserving and replicating research, applications, databases, software, and more. we will examine the current use cases of reprozip6, ranging from digital humanities to data science. we also explore potential library use cases for reprozip, particularly in digital libraries and archives, liaison librarianship, and other library services. we believe that libraries and archives can leverage reprozip to deliver more robust reproducibility services, repository services, as well as enhanced discoverability and preservation of research materials, applications, software, and computational environments. keywords reproducibility, data management, repository management, digital libraries, digital archiving, software preservation 1 introduction reproducibility is at the core of the research process: it is not only essential for verification and authentication of results, but also for driving a field forward. if a work is reproducible, newcomers to the field can easily learn methods and experienced researchers can build upon it. despite the widespread attention drawn to the subject following the reproducibility project: psychology, carried out by the center for open science (open science collaboration, 2015), reproducibility still remains an elusive target for many researchers (goodman, fanelli, and ioannidis 2016). while sharing research and materials should be easily achieved by using institutional or domain-specific repositories, this does not guarantee reproducibility. as research has become increasingly reliant on digital tools, the challenges in reproducibility have become more difficult. this is where the concept of computational reproducibility becomes important. before, researchers in the field captured their environments through observation, drawings, photographs, and videos; now, researchers and librarians who endeavor to preserve their work must begin to capture digital environments to achieve reproducibility (stodden et al. 2016). however, preserving digital environments is exceedingly difficult to do manually. for example, gronenschild et. al (2012) discussed how the results of data analyses in neuroscience performed with the same application differed based on the operating system: we investigated the effects of data processing variables such as freesurfer version (v4.3.1, v4.5.0, and v5.0.0), workstation (macintosh and hewlett-packard), and macintosh operating system version (os x 10.5 and os x 10.6). significant differences were revealed between freesurfer version v5.0.0 and the two earlier versions. [...] about a factor two https://doi.org/10.29173/iq18 https://reprozip.org/ https://examples.reprozip.org/ https://reprozip.org/ https://reprozip.org/ 2/14 steeves, vicky; rampin, rémi and chirigati, fernando (2018) using reprozip for reproducibility and library services, iassist quarterly 42 (1), pp. 1-14. doi https://doi.org/10.29173/iq18 smaller differences were detected between macintosh and hewlett-packard workstations and between mac os x 10.5 and mac os x 10.6. the observed differences are similar in magnitude as effect sizes reported in accuracy evaluations and neurodegenerative studies. in addition, there may be many unforeseen dependencies for each software or tool, of which different versions from the original configuration may give totally disparate results or not even run. to manually address these problems, collectively known as ‘dependency hell,’ researchers enter into an errorprone and resource-heavy process. they would have to create a file that encapsulates metadata about their computational environment, including the operating system, hardware architecture, and software library dependencies. tracking these dependencies is challenging – there are many layers of hardware and software which the average user has no skill or time to examine (marwick, 2015). there are a few classes of tools that can be used to capture digital environments. workflow systems, such as vistrails7 and kepler8, represent research as an executable diagram, depicting the flow of the different steps and processes of the research (davidson and freire, 2008). workflow systems help researchers conceptualize and manage the analysis process, support researchers by allowing the creation and reuse of analysis tasks, and (more recently) systematically record provenance information for later use. however, workflow systems have a steep learning curve and often represents too much of an ‘ask’ for users to adopt. in addition, software dependencies are rarely captured by these systems. virtual machines (vms) are an emulation of a computer system and provide the functionality of a physical computer. common implementations involve specialized hardware, software, or a combination of both (smith and ravi nair, 2005). the resulting files, called ‘images,’ are often large (gigabytes in size) since they encapsulate the entire operating system. while lightweight solutions exist for vms, such as vagrant and docker, the user is again tasked with building these images, which is timeconsuming. another potential solution are configuration management systems (feller, 2013), where users write ‘recipes’ that take the form of a requirements file that can be use to automate the installation of dependencies. again, these files must be created manually by the user and updated as their work takes new shape. porting existing research to these current solutions takes too much time and work to be of use in researchers’ day-to-day workflow. a class of tools that can automatically capture all the dependencies in the original environment and automatically set them up in another environment is needed to span this gap. there are a few tools that address parts of that gap, such as provenance-to-use (pham et al., 2013), which uses cde (code, data, and environment packaging for linux) (guo, 2012) to create a package that contains the dependencies necessary to reproduce the research. however, users can only reproduce these packages on linux, otherwise they would need to manually create a virtualization environment, which is less user-friendly and can be cumbersome for complex research. reprozip was created to fully bridge this gap in computational reproducibility. it is an open source, freely available program that allows users to make their work completely reproducible, down to the operating system level. reprozip can be used to reproduce a plethora of applications, including data analysis tools, scripts and software written in any language, graphical tools, interactive tools, clientserver applications (including databases), and jupyter notebooks9. reprozip works in two steps: automatically tracing the execution of work and then packaging all dependencies in a single, distributable package, a .rpz file (chirigati et al., 2016). the ease of use extends to the unpacking step as well. for others to reproduce research that was packed using reprozip, they use reprounzip, which works by providing different methods of unpacking from which the second user can choose. reprounzip then automatically sets up the environment by extracting metadata captured during the initial tracing process and building the new environment automatically. from there, the user only needs to rerun the work or application. packaging and reproducing can be https://doi.org/10.29173/iq18 https://www.vistrails.org/ https://kepler-project.org/ https://reprozip.org/ https://ipython.org/notebook.html 3/14 steeves, vicky; rampin, rémi and chirigati, fernando (2018) using reprozip for reproducibility and library services, iassist quarterly 42 (1), pp. 1-14. doi https://doi.org/10.29173/iq18 achieved in four simple steps — two steps to pack, and two to unpack — and makes reproducibility easier to achieve across research domains as compared to the options referenced above. users can heavily customize the way they use reprozip as well. the original user packing their work has a lot of control in what is included in a .rpz file. reprozip will automatically identify input files, dependencies, and output files, but users have the option of editing the configuration file to remove or add files, to label items more meaningfully, and more. reprozip packages are also highly portable, in that it automatically creates a virtual machine for the user – no extra work required beyond one click or command – allowing research to be reproduced across different operating systems. reprozip also uses an extensible architecture for unpacking, which means anyone can add an unpacker to reproduce research while maintaining compatibility with existing reprozip packages. for example, if docker ceased to exist, then another unpacker could be written and added to reprozip, and the packages would maintain forward compatibility. while reprozip has primarily been used in research, in this paper, we explore the many ways in which librarians can use reprozip, from helping user populations create well-managed, reproducible research, to preserving computational environments, and to building library infrastructure. 2 technical infrastructure reprozip has two stages, one for packing and another for unpacking. reprozip creates a small, selfcontained package (.rpz file) by automatically identifying, tracking, and capturing all required dependencies of research, applications, databases, and indeed, most digital work (chirigati et al., 2016). this package can easily be shared, as it is usually quite small, with reviewers, collaborators, or released to the general public (see figure 1). secondary users can unpack the .rpz using reprounzip, and reproduce the work on their machine, regardless of operating system (see figure 2). reprounzip then automatically sets up the environment for that user by extracting metadata captured during the initial tracing process and building the environment automatically. reprounzip's functionality is not limited to simple reproduction: it also allows users to modify the original work for reuse for new purposes, with very little effort. 2.1 packing currently, users can pack their work using reprozip on linux operating systems via the command line. future work includes adding in a gui for packing and support for non-linux environments (see future development work10). reprozip first needs to trace system calls, which identifies all the dependencies necessary to reproduce the environment and the research within. this is accomplished when the user prepends the command ‘reprozip trace’ to the execution of their research. for instance, if a user runs a python script to analyze data by typing ‘python analysis.py’ in the command line, to trace with reprozip, the user simply runs ‘reprozip trace python analysis.py’ instead. reprozip collects information including command-line arguments, environment variables, files read, and files written, and stores everything in a sqlite database. to detect detailed information about the dependencies, reprozip keeps strict provenance data. for instance, by observing which files were read and by using the software manager in the operating system, reprozip can identify the software libraries on which the research depends. reprozip uses rules to identify the role of files in the execution of the research, such as input files – those files that were only read and do not belong in a software package. all of this information is written to a humanreadable yaml configuration file. the user can then customize the package via the configuration file, which contains all of the required information for reproducing the work including commands, arguments, working directory, https://doi.org/10.29173/iq18 4/14 steeves, vicky; rampin, rémi and chirigati, fernando (2018) using reprozip for reproducibility and library services, iassist quarterly 42 (1), pp. 1-14. doi https://doi.org/10.29173/iq18 environment variables, operating system information, input files, output files, and software packages. this generated file is enough to create a reproducible package, however users can choose to modify it, e.g. to remove sensitive or proprietary information, give meaningful labels to input and output files, or include additional files to support more reproduction settings. lastly, users generate a package using the command ‘reprozip pack <package-name> ’ which compresses all the dependencies, files, and environmental information into a small .rpz file. this is usually small enough to send to reviewers of papers, deposit into an institutional repository, or simply be saved for archival purposes. figure 1. packing step on reprozip. 2.2 unpacking reprounzip allows users to unpack a .rpz package on any operating system, using the command line or the gui (figure 3). different unpackers are provided to reproduce the corresponding work. reprounzip was designed to allow anyone to create new plugins to unpack .rpz files, which will support every package created in the past. this model allows both the core maintainers of reprounzip and its users/contributors to plan for current unpackers (vagrant and docker) to become deprecated, or for future unpackers to be developed. this way, users are not locked into specific technologies, and so can use their reprozip packages to the fullest. figure 2. unpacking step on reprozip. unpacking works in two steps. the first is setting up the environment, which consists of extracting information from the .rpz, and can be accomplished by using the `reprounzip setup <package-name>` https://doi.org/10.29173/iq18 5/14 steeves, vicky; rampin, rémi and chirigati, fernando (2018) using reprozip for reproducibility and library services, iassist quarterly 42 (1), pp. 1-14. doi https://doi.org/10.29173/iq18 command, or by double-clicking the .rpz file to use with the gui. this step varies based on the unpacker chosen. for reproducing research across different operating systems, users can choose between vagrant and docker unpackers. for reproducing the research on linux, users can additionally use the `directory` or the `chroot` unpackers. for the latter, setting up the .rpz file means copying its contents into a single directory. for the former, a docker container or vagrant virtual machine is initialized. to re-execute the research, users can either use the `reprozip run <path-to-unpacked-work>` on the command line or click the ’run experiment’ button in the gui. if the user is unpacking using vagrant or docker, reprounzip will automatically create and start the virtual machine or docker container. the secondary user needs no knowledge of these programs – they simply need to use the setup and run commands, and reprounzip will take care of the rest. the `directory` and `chroot` options install all the required dependencies in the secondary users’ native system, and reruns the research inside a directory. these two unpackers are not recommended because they are not as reliable as vagrant or docker, which provide an isolated execution environment. figure 3. the graphical user interface for unpacking research. 2.2.1 data management reprozip identifies the role of files in the research (input and/or output) so users can take advantage of this in the unpacking step. after setting up and reproducing the research, users can download the resulting output files for their own inspection or reuse. users may also upload their own input files – this is useful when a researcher wants to use the same method on data they have collected independently of the original user. reprozip also allows users to visualize and edit their workflow through an integration with vistrails, a scientific workflow system (freire and silva, 2012). by working in vistrails, users can examine, modify, and play around with their own workflow, which is useful for evaluating the efficiency of methodologies used in the original work. this is made easier by the fact that users can re-execute unpacked work directly inside of vistrails. https://doi.org/10.29173/iq18 6/14 steeves, vicky; rampin, rémi and chirigati, fernando (2018) using reprozip for reproducibility and library services, iassist quarterly 42 (1), pp. 1-14. doi https://doi.org/10.29173/iq18 3 current use cases reprozip’s predominant use cases are for archiving scholarly works, publishing and sharing reproducible research, and reviewing results in publications for academic journals. within scholarly works, reprozip is used as a means to reproduce work, but also as a useful tool for reproducibility in many different fields, most recently data journalism, geoscience, and neuroscience, for example: boss and broussard, 2016; knoth and nüst, 2017; ghosh et al., 2017. since version 1.0 was released, reprozip has had 271 users, and a subset of those users have allowed us to publicize their work on github11 and through a website, reprozip-examples12. currently, there are four examples from digital humanities and social sciences, specifically from web publishing (including capturing client-server architecture and databases), history, and data journalism (one example has a playlist of demo videos for this example13 on youtube). additional examples include use cases in stem, including a neuroscience example, which also demonstrates the way reprozip can be used in tandem with the popular tool jupyter notebooks (there are also demo videos14 on youtube for this example); one interactive graphical application example, and many data science examples, specifically in machine learning, statistics (including simulations), and urban studies. these examples range not only in discipline, but also in technologies used. we have examples using java, r, python, the django web framework, postgresql databases, and more. the largest reprozip package in this gallery of examples is only 84 mb, and the smallest is 19 mb. figure 4. steps in using reprozip to package stackedup, a web application. https://doi.org/10.29173/iq18 https://github.com/vida-nyu/reprozip-examples https://examples.reprozip.org/ https://www.youtube.com/watch?v=soe2nejwylw&list=pljgz3v4gfxpxdprbafth42w3hrmmx2wfd https://www.youtube.com/watch?v=gs4s2jh8yw4&list=pljgz3v4gfxpw_lesrw1ze2ql1yxdsgzki 7/14 steeves, vicky; rampin, rémi and chirigati, fernando (2018) using reprozip for reproducibility and library services, iassist quarterly 42 (1), pp. 1-14. doi https://doi.org/10.29173/iq18 reprozip was also recently integrated into the national institute of standards and technology’s corr (cloud of reproducible records) web platform, designed as a gateway for simulation management tools to capture software executions. corr also allows users to store and view metadata associated with simulation records (faical yannick palingwende congo, 2015). the testing repository is on github at corr-reprozip15. reprozip has also been integrated into publication workflows for scholarly journals and conference proceedings. the journal information systems recommends using reprozip for its reproducibility section, which evaluates the reproducibility of submitted papers. reprozip is also recommended by the association of computing machinery (acm) artifact evaluation process guidelines, and in the acm special interest group on management of data (sigmod) reproducibility review for conference materials. there is continued work on expanding this to fields outside computing, given that reprozip has a demonstrated usefulness in digital humanities, physical sciences, and social sciences. 4 reprozip in librarianship librarians involved with research, whether their own or others’, can utilize reprozip to help make it reproducible. disciplines within librarianship can also use reprozip to great advantage, such as software preservation, repository management, and generally digital libraries and archives infrastructure. 4.1 digital libraries using reprozip, digital preservation work can be streamlined; users, data managers, and digital archivists/librarians can capture works, research, and most digital content at the environmental level to create a preservation-ready object. this file is a preservation-ready object for a number of reasons, but the most salient is the flexibility in unpacking technologies. one .rpz file can be rerun using many different unpackers in reprounzip. due to reprozip’s plugin model, unpackers can be added and removed as systems become more or less popular or usable. if vagrant no longer exists in ten years, then an unpacker for another virtual machine system can be written and the reprozip packages remain usable. however, if in the future there are no containers or virtual machines, then the archivist/librarian can still use the robust technical and administrative metadata in the configuration file that reprozip provides. this config.yml is machine-readable, and lists every single piece of information about the computational environment and dependencies necessary to recreate the environment. this can be extracted during the ‘setup’ phase of unpacking. as digital libraries and archives have begun collecting and preserving software alongside other materials they collect, this has presented many challenges in terms of long-term preservation, which reprozip could help alleviate. there is a lot of work being done in this area by not only institutional libraries, but also in consortial groups such as the software preservation network16, who seek to “save software together,” through community engagement, infrastructure support, and knowledge generation. however, software preservation has a unique problem in that the boundaries for the preservationready objects are not clearly defined. this is due in large part to the problem of dependency hell described earlier. some considerations include archiving simply the source code or executable, or the whole environment on which it runs (mcdonough and olendorf, 2011). in describing use cases for software preservation, rios et. al. (2017) call out first the “capture [of] the research transparency, reproducibility, and reuse scenarios of software preservation. the goals are to enable other researchers to examine and understand the preserved software and enable its re-execution and reuse.” https://doi.org/10.29173/iq18 https://github.com/usnistgov/corr-reprozip http://www.softwarepreservationnetwork.org/ 8/14 steeves, vicky; rampin, rémi and chirigati, fernando (2018) using reprozip for reproducibility and library services, iassist quarterly 42 (1), pp. 1-14. doi https://doi.org/10.29173/iq18 reprozip, in capturing computing environments in which research takes place, could be used to preserve software down to the operating system on which it runs. additionally, reprozip allows for users not only to reproduce and modify the packages, but also to reuse them for their own purposes. making collections accessible and usable are key components of libraries and archives, and reprozip facilitates this goal for research, applications, and software. users could leverage reprozip at library terminals, through obtaining a .rpz file via the catalog, and simply reproduce the research for their own learning and enrichment, or build on the contents for their own work. reprozip allows digital librarians and archivists to capture scholarly works, software, and more at the computing environmental level, and have a preservation-ready object at the end of it: the .rpz. 4.2 repository management open and accessible infrastructure is key to advancing openness in research. libraries have answered these challenges by either creating data repositories alongside their institutional repositories, or integrating deposit of data and other research output into their institutional repositories. as more repositories look to include research materials, building in tools to enable reproducibility is key. reprozip’s implementations in corr is one example of how it can be integrated into a repository environment. corr uses reprozip as an application that users registered to the platform could use to host and store their reprozip ’records,’ which are metadata records and .rpz files that are accessible, shareable, and published within corr. users can add to the ’reprozip trace’ command the configuration to automatically push .rpz files and metadata to the platform. this represents a way to integrate reprozip into the ingest process of repositories. while the ingest process in a library context would look different (a standard ingest process and descriptive metadata added), reprozip provides extensive metadata extraction options available from the configuration files that would help in automating this workflow. from the package, which contains extremely detailed technical and administrative metadata, can be exported as a json file, which allows for extensible models of metadata – such as crosswalking to the resource description framework (rdf) or dublin core (digital public library of america, 2017). aggregating metadata from diverse digital objects requires this level of extensibility, which reprozip automatically supplies within its own options for reading and extracting metadata. allowing reprozip packages to be a citable object within repositories would also go a long way in creating cultural shifts towards rewarding reproducible practices. as repositories begin to build curation workflows around reproducibility, reprozip can be built in to expand the creation of reproducible works. additionally, as these services grow and expand models of ingest workflows to include a variety of digital objects, repository managers can leverage the metadata created within the reprozip packages. 4.3 academic libraries with the rise of data services departments (also called data management, research data services, or statistical services) in academic libraries, their involvement with researchers’ data management has grown exponentially (crowe and crumpton, 2016). currently, most support in academic libraries for reproducibility comes out of data management services. however, data management is a means towards reproducibility, and libraries must begin to develop holistic support for reproducible research practices, extending beyond simply data management. 4.3.1 liaison librarians/subject specalists liaison librarians play a central role in identifying and reacting to user needs. liaison work includes assessment of services, collecting materials to satisfy patron needs, and locating useful resources to enhance scholarship. but perhaps most importantly, liaison work involves a focused and dedicated relationship with a subset of library users – the work is reliant on these relationships which requires two-way communication (silver, 2015). in this way, many library workers rely on liaison librarians to https://doi.org/10.29173/iq18 9/14 steeves, vicky; rampin, rémi and chirigati, fernando (2018) using reprozip for reproducibility and library services, iassist quarterly 42 (1), pp. 1-14. doi https://doi.org/10.29173/iq18 effectively communicate new initiatives, ideals, and services offered. liaison librarians are the bridge between students and faculty and the library. this offers an excellent opportunity to encourage and disseminate information on reproducible practices, particularly with reprozip. furthermore, liaison librarians also are research collaborators. a recent case study that exemplifies this idea is a paper co-authored by katherine boss, the liaison librarian at new york university for journalism and media studies, and meredith broussard, a faculty member in the arthur l. carter journalism institute at new york university. their paper, ’challenges facing the preservation of borndigital news applications,’ delves into how two different domain-specific sets of expertise work in tandem to solve a set of problems evident in both fields – digital preservation and archiving of news apps (such as dollars for docs17). on the library side, there’s the challenge of digital preservation and archiving software, but also conceptually of how to collect, disseminate, and make these news apps accessible (boss and broussard, 2016). on the journalism side, there is a serious concern that as funding falls away, so too will these important digital resources that organizations such as propublica provide. working together, these research questions are advanced in both fields, and accessibility to vetted news applications is advanced. boss and broussard are also exploring reprozip as a solution to news app preservation (see the stackedup18 example on reprozip-examples) as a direct result of collaboration with the reprozip development team (one member is a librarian) which resulted in boss communicating this tool to her department. generalizing from boss and broussard’s partnership, there is space and motivation for liaison specialists to work not only with other library partners, but also researchers. this can include using reprozip for mutual benefit -the library gains an excellent way to preserve and build collections of diverse, preservation-ready research outputs, and users easily integrate reproducible practices into their workflows. 4.3.2 data services as libraries add support for data services, which typically include data management training and software support for quantitative, geographic information systems (gis), and qualitative research, there should be support for reproducible research practices. as demonstrated, reprozip is an extremely user-friendly tool. in two steps, users can have reproducible research, and in another two steps, a secondary user can unpack the .rpz file and re-execute their work. this presents a low barrier not only to the research community, but also to those in data services departments involved with instruction and consultation. support for reproducibility using reprozip in research data services would draw in a lot of opportunities to cross-campus collaboration, as well as position data services departments and the library as a center for modern research practices. by adding reprozip to the curriculum of data services department workshops, and to the list of supported software, data services teams can become a center on campus for assisting the reproducibility of their user’s works. 5 future development work reprozip has been evolving since 2012 and has always been open source. development work on the tool has been initiated through assessing user needs, ascertained through one-on-one interactions, workshops, github issues, and/or the reprozip users mailing list. the hope is to continue this work as a core team of maintainers, but also solicit contributions from the community of users. 5.1 macos packing reprozip currently only allows for packing on linux os. across domains, this represents a minority of users, who tend to prefer macos and windows. https://doi.org/10.29173/iq18 https://projects.propublica.org/docdollars/ https://github.com/vida-nyu/reprozip-examples/tree/master/stacked-up 10/14 steeves, vicky; rampin, rémi and chirigati, fernando (2018) using reprozip for reproducibility and library services, iassist quarterly 42 (1), pp. 1-14. doi https://doi.org/10.29173/iq18 the team has chosen to begin with macos, as packing on this operating system is technically possible with a low barrier, based on our previous work. because macos and linux are both unix-based operating systems, reprozip could potentially use dtrace to capture dependencies on macos in place of the ptrace that is available for linux systems. the limitations of reproducing macos environments are significant in terms of licensing; as set out in the macos user agreement, users may only run macos on apple machines. therefore, unpacking would only be 'allowed' on another macos machine, which would be very limiting to the users. 5.2 reprozip-jupyter jupyter notebooks (also know as ipython notebooks) are an interactive computational environment where users can combine executable code, text, math, plots, and rich media (thomas et al., 2016). this tool is increasingly at the center of research initiatives across many domains of scholarship. however, jupyter notebooks are not inherently reproducible: they do not capture provenance information, dependencies, or the native environment. to add this missing element, the developers created the jupyter plugins for reprozip, which allow users to pack a notebook from the jupyter interface, and also to choose to start a notebook server and rerun when unpacking. users can see the source19 and watch a demo video20 demonstrating packing a notebook. 5.3 workflow visualizations and graphs currently, reprozip allows users to visualize the provenance and execution of the contents of a .rpz package using graphviz, an open source graph visualization software, used to represent structural information as diagrams of abstract graphs and networks (gansner and north, 2000). this allows users to graphically represent the flow of their work, both for internal exploratory purposes and for tracking provenance. there is room for improvement on this simple visualization by using the d3 javascript library21 to make the graph more appealing aesthetically, and to allow the user to interactively explore the provenance data which is often too much to be shown in full on a single graph. it will also be integrated into the gui of reprounzip, so users could interact with the graph and choose to save snapshots as image files for sharing and publication from there. this would help users in communicating the execution and provenance of their research or application. a prototype of this has been created but not yet integrated into reprozip. 5.4 mpi/hpc lastly, there is demand to add more support for researchers who use high performance computing (hpc) in their research. reprozip runs on linux and therefore on most hpc clusters. there has been success reproducing experiments spanning multiple machines with shared filesystems using the message passing interface (mpi) framework, merging the traces into a working .rpz file. in the future, the team aims to add functionality that is aware of the common scheduling tools and setups of hpc facilities to automatically run reprozip to trace all the processes without the user having to adapt his job submission routine. for unpacking, correctly rerunning an unpacked experiment through reprounzip for reproduction on a new cluster also requires extra functions. it would also be prudent to develop unpacker plugins specifically for hpc environment, where virtual machines are impractical and docker is generally not yet available. 6 conclusion the library community can leverage reprozip in instruction, consultation, repository services, digital archiving, and in their own research. it is extensible enough to be used for reproducibility across research domains as well as across library services. it is simple enough to use that it can be easily adopted to ensure reproducibility of scholarship, even if reproducibility is an afterthought. the user does not need to start their project with reproducibility in mind – but can, in the end, use reprozip to https://doi.org/10.29173/iq18 https://github.com/vida-nyu/reprozip/tree/ipython https://www.youtube.com/watch?v=y8ymgvyhhs8 11/14 steeves, vicky; rampin, rémi and chirigati, fernando (2018) using reprozip for reproducibility and library services, iassist quarterly 42 (1), pp. 1-14. doi https://doi.org/10.29173/iq18 create a compendium of their work that’s easily shareable, citable (if deposited in an institutional repository), and usable by themselves and the community at large. ultimately, this improves research by lowering the barriers to reproducibility, and also improves the way in which the lis community can preserve and make discoverable these research compendia. 7 acknowledgements we would like to acknowledge dr. juliana freire, the principal investigator of the reprozip project, for her support in continuing to build reprozip. we’d also like to thank dr. nicholas wolf, beth daniel lindsay, jonathan petters, and declan fleming for comments on this paper. lastly, we would like to acknowledge the support from the gordon and betty moore foundation as well as the alfred p. sloan foundation. the moore-sloan data science environment was vital to the development of reprozip. references boss, katherine, and meredith broussard. “challenges facing the preservation of born-digital news applications,” 2016. http://blogs.sub.uni-hamburg.de/ifla-newsmedia/wp-content/uploads/2016/04/boss-broussardchallenges-facing-the-preservation-of-born-digital-news-applications.pdf. chirigati, fernando, rémi rampin, dennis shasha, and juliana freire. “reprozip: computational reproducibility with ease,” 2085–88. acm press, 2016. doi:10.1145/2882903.2899401. crowe, kathryn, and michael crumpton. “defining the libraries’ role in research: a needs assessment; a case study,” 2016. https://libres.uncg.edu/ir/listing.aspx?id=19091. davidson, susan b., and juliana freire. “provenance and scientific workflows: challenges and opportunities.” in proceedings of the 2008 acm sigmod international conference on management of data, 1345–1350. sigmod ’08. new york, ny, usa: acm, 2008. doi:10.1145/1376616.1376772. digital public library of america. “an introduction to the dpla metadata model,” january 6, 2017. https://dp.la/info/wp-content/uploads/2015/03/intro_to_dpla_metadata_model.pdf. “docker.” docker, 2017. https://www.docker.com/. faical yannick palingwende congo. “building a cloud service for reproducible simulation management.” in proceedings of the 14th python in science conference, edited by kathryn huff and james bergstra, 195–201, 2015. https://conference.scipy.org/proceedings/scipy2015/pdfs/yannick_congo.pdf. feller, m. “configuration management.” ieee transactions on engineering management em-16, no. 2 (february 7, 2013): 64–66. doi:10.1109/tem.1969.6447051. fernando rios, nicole contaxis, bridget almas, paula jabloner, and heidi kelly. “exploring curationready software: use cases.” the software preservation network, april 14, 2017. http://www.softwarepreservationnetwork.org/exploring-curation-ready-software-use-cases/. freire, j., and c. t. silva. “making computations and publications reproducible with vistrails.” computing in science engineering 14, no. 4 (july 2012): 18–25. doi:10.1109/mcse.2012.76. https://doi.org/10.29173/iq18 12/14 steeves, vicky; rampin, rémi and chirigati, fernando (2018) using reprozip for reproducibility and library services, iassist quarterly 42 (1), pp. 1-14. doi https://doi.org/10.29173/iq18 gansner, emden r., and stephen c. north. “an open graph visualization system and its applications to software engineering.” software: practice and experience 30, no. 11 (september 2000): 1203–33. doi:10.1002/1097-024x(200009)30:11<1203::aid-spe338>3.0.co;2-n. ghosh, satrajit s., jean-baptiste poline, david b. keator, yaroslav o. halchenko, adam g. thomas, daniel a. kessler, and david n. kennedy. “a very simple, re-executable neuroimaging publication.” f1000research 6 (february 10, 2017): 124. doi:10.12688/f1000research.10783.1. goodman, steven n., daniele fanelli, and john p. a. ioannidis. “what does research reproducibility mean?” science translational medicine 8, no. 341 (june 1, 2016): 341ps12-341ps12. doi:10.1126/scitranslmed.aaf5027. gronenschild, ed h. b. m., petra habets, heidi i. l. jacobs, ron mengelers, nico rozendaal, jim van os, and machteld marcelis. “the effects of freesurfer version, workstation type, and macintosh operating system version on anatomical volume and cortical thickness measurements.” edited by satoru hayasaka. plos one 7, no. 6 (june 1, 2012): e38234. doi:10.1371/journal.pone.0038234. guo, philip. “cde: a tool for creating portable experimental software packages.” computing in science & engineering 14, no. 4 (july 2012): 32–35. doi:10.1109/mcse.2012.36. knoth, christian, and daniel nüst. “reproducibility and practical adoption of geobia with opensource software in docker containers.” remote sensing 9, no. 3 (march 18, 2017): 290. doi:10.3390/rs9030290. marwick, ben. “how computers broke science – and what we can do to fix it.” the conversation, november 9, 2015. http://theconversation.com/how-computers-broke-science-and-what-we-can-doto-fix-it-49938. mcdonough, jerome, and robert olendorf. “saving second life: issues in archiving a complex, multiuser virtual world.” international journal of digital curation 6, no. 2 (october 7, 2011): 89–108. doi:10.2218/ijdc.v6i2.192. open science collaboration. “estimating the reproducibility of psychological science.” science 349, no. 6251 (august 28, 2015): aac4716-aac4716. doi:10.1126/science.aac4716. pham, quan, tanu malik, and ian foster. “using provenance for repeatability.” in proceedings of the 5th usenix conference on theory and practice of provenance, 2–2. tapp’13. berkeley, ca, usa: usenix association, 2013. http://dl.acm.org/citation.cfm?id=2482613.2482615. silver, isabel d. “for your enrichment: outreach activities for librarian liaisons.” reference & user services quarterly 54, no. 2 (january 26, 2015): 8–14. smith, j.e., and ravi nair. “the architecture of virtual machines.” computer 38, no. 5 (may 2005): 32–38. doi:10.1109/mc.2005.173. stodden, victoria c. “the scientific method in practice: reproducibility in the computational sciences,” 2010. https://academiccommons.columbia.edu/catalog/ac:140117. stodden, victoria, marcia mcnutt, david h. bailey, ewa deelman, yolanda gil, brooks hanson, michael a. heroux, john p. a. ioannidis, and michela taufer. “enhancing reproducibility for https://doi.org/10.29173/iq18 13/14 steeves, vicky; rampin, rémi and chirigati, fernando (2018) using reprozip for reproducibility and library services, iassist quarterly 42 (1), pp. 1-14. doi https://doi.org/10.29173/iq18 computational methods.” science 354, no. 6317 (december 9, 2016): 1240–41. doi:10.1126/science.aah6168. thomas, kluyver, ragan-kelley benjamin, pérez fernando, granger brian, bussonnier matthias, frederic jonathan, kelley kyle, et al. “jupyter notebooks-a publishing format for reproducible computational workflows.” stand alone, 2016, 87–90. doi:10.3233/978-1-61499-649-1-87. endnotes 1vicky steeves is librarian for research data management and reproducibility, division of libraries & center for data science, new york university, vicky.steeves@nyu.edu 2rémi rampin is phd candidate, tandon school of engineering, new york university, remi.rampin@nyu.edu 3fernando chirigati is research engineer, center for data science, new york university, fchirigati@nyu.edu 4the authors presented this paper at the 2017 iassist conference. the presentation can be found here: https://vickysteeves.gitlab.io/2017-iassist-reprozip/#/ 5https://www.reprozip.org/ 6https://examples.reprozip.org/ 7https://www.vistrails.org/ 8https://kepler-project.org/ 9https://ipython.org/notebook.html 10#5 future development work 11https://github.com/vida-nyu/reprozip-examples 12https://examples.reprozip.org/ 13https://www.youtube.com/watch? v=soe2nejwylw&list=pljgz3v4gfxpxdprbafth42w3hrmmx2wfd 14https://www.youtube.com/watch?v=gs4s2jh8yw4&list=pljgz3v4gfxpw_lesrw1ze2ql1yxdsgzki 15https://github.com/usnistgov/corr-reprozip 16http://www.softwarepreservationnetwork.org/ 17https://projects.propublica.org/docdollars/ 18https://github.com/vida-nyu/reprozip-examples/tree/master/stacked-up https://doi.org/10.29173/iq18 mailto:vicky.steeves@nyu.edu mailto:remi.rampin@nyu.edu mailto:fchirigati@nyu.edu https://vickysteeves.gitlab.io/2017-iassist-reprozip/#/ https://www.reprozip.org/ https://examples.reprozip.org/ https://www.vistrails.org/ https://kepler-project.org/ https://ipython.org/notebook.html https://github.com/vida-nyu/reprozip-examples https://examples.reprozip.org/ https://www.youtube.com/watch?v=soe2nejwylw&list=pljgz3v4gfxpxdprbafth42w3hrmmx2wfd https://www.youtube.com/watch?v=soe2nejwylw&list=pljgz3v4gfxpxdprbafth42w3hrmmx2wfd https://www.youtube.com/watch?v=gs4s2jh8yw4&list=pljgz3v4gfxpw_lesrw1ze2ql1yxdsgzki https://github.com/usnistgov/corr-reprozip http://www.softwarepreservationnetwork.org/ https://projects.propublica.org/docdollars/ https://github.com/vida-nyu/reprozip-examples/tree/master/stacked-up 14/14 steeves, vicky; rampin, rémi and chirigati, fernando (2018) using reprozip for reproducibility and library services, iassist quarterly 42 (1), pp. 1-14. doi https://doi.org/10.29173/iq18 19https://github.com/vida-nyu/reprozip/tree/ipython 20https://www.youtube.com/watch?v=y8ymgvyhhs8 21https://d3js.org https://doi.org/10.29173/iq18 https://github.com/vida-nyu/reprozip/tree/ipython https://www.youtube.com/watch?v=y8ymgvyhhs8 by 10 iassist quarterly fall 2009 by flavio bonifacio1 numbers abstract the first section introduces the problem of numbers as a representation of reality. the second and third sections underline some difficulties that one may encounter in communicating numbers and introduce the problem of the institutionalisation of number processing production as a system of guarantee. the fourth and fifth sections see numbers as a part of a model. the sixth and eighth sections track a two-way path between numbers and society, while the seventh mentions the problem of the representation of numbers. the conclusion recognises as an urgent problem the need for an alert system on the uses and on the production processes of numbers. keywords: quantitative social research; communicating numbers; reality representation; use of numbers; 1. numbers and objectivity it is not easy to identify the border between what is certainly and objectively measurable and what is not. it would seem that everything that can be expressed with numbers or, more generally, that which can be formalized in mathematical models is objectively measurable. on the other hand, what it is not possible to formalize or express through numbers would seem not to be objectively measurable. in these brief notes i shall explain how the idea that numbers cannot be considered a criterion of demarcation (following what i have already discussed with regard to popper, [flavio bonifacio 1996, part iii]), how numbers are not more true than other representations of reality, that every statement about reality must be responsibly supported, that it is necessary to support this responsibility with an explicit agreement between the producers of the data, that this agreement must be institutionally granted, that only this agreement can make it possible not to surrender the objectivity of the measurements to the tastes of the moment. here below i report some difficulties in communication through numbers which i shall summarise in four questions: (1) what do we communicate with numbers? (2) who builds those numbers? (3) who communicates the numbers? and (4) to whom do they communicate the numbers? first of all i think i should warn the reader about the style in which this article is written. i have chosen to use direct language to stress the urgency of the problem posed, i.e. the misuse of numbers. what seem to be opinions are simply statements that concern my view of the real state of things, not their scientific representation. they concern what i feel to be real and what i want to communicate. it is not my intention to provide here an exhaustive description of the real meaning or interpretation of specific numbers. others have already done it better and more thoroughly and i shall direct the reader to these authors. i want to communicate instead the feeling coming from years of work in the field of data analysis, the feeling that i have contributed towards building a sand castle, the feeling of disillusion. nevertheless when possible i have used scientific methodology, see for example section 5, to illustrate my point of view which is, after all, optimistic: it is worth fighting for a better and more “true” representation of reality because objectivity is not given, objectivity is a conquest. despite the fact that it is rather nonsensical to say that reality is “what i feel to be real” i think that the sentiments here narrated are shared also by the target audience of this article who i think are mainly data workers and numeracy authors. 2. problems of communication with numbers typically, with numbers we communicate “quantities” the number of tourists in a place in a certain period of time, the numbers of those employed, the number of the victims in some natural disaster etc. while we are not surprised that a piece of news, a fact, can be described with words and in ways that are also very different and we have no difficulty in admitting that a description is constructed in different ways and that of a fact we can give different descriptions, we think that the “number” exists per se and that it is not therefore debatable. to show as a context may be differently represented with words, it is instructive to read the various descriptions that may be given of a banal event, such as the one which took place on what could be any bus s during the rush hour. this is narrated in 99 different ways by r. queneau in his book exercises in style, and some excerpts from the translation by barbara wright are provided briefly here below, just to iassist quarterly fall 2009 11 give a flavour of queneau’s book: notation (that is the master piece that is afterwards narrated in 99 different ways) in the s bus, in the rush hour. a chap of about 26, felt hat with a cord instead of a ribbon, neck too long, as if someone’s been having a tug-of-war with it…(continue) cockney (the cockney way) so a’m stand’n’ ahtsoider vis frog bus when a sees vis young froggy bloke, caw bloimey, a finks, ‘f’at ain’t ve most funniest look’n’ geezer wot ever a claps eyes on. bleed’n’ great neck, jus’ loike a tellyscope, ……(continue) sonnet (the sonnet way) glabrous was his dial and plaited was his bonnet,/ and he, a puny colt-(how sad the neck he bore,/and long)-was now intent on his quotidian chore-/the bus arriving full, of somehow getting on it……(continue) mathematical (the mathematical way) in a rectangular parallepiped moving along a line representing an integral solution of the second-order differential equation: y’’+pptb(x)y’+s=84 two homoids (of which only one, the homoid a, manifests a cylindrical element of length l>n ……(continue) [barbara wright, 1958] instead it would seem that reality translated into numbers is in some way more true than that translated into words. that regardless of evidence to the contrary: probably the “true” number of tourists, the “true” number of the employed, the “true” number of the victims of that natural disaster will never be known. as for other cases, we can think of the rate of inflation or the forecasts of the development of the gross domestic product (gdp), or the number of participants at a political or trade union demonstration. the construction process of such data is so complex and at times even biased 2, that it is difficult to trust in them blindly. 3. institutional truth of numbers additionally, a much simpler and immediate number, which would appear to be alien to any dispute as to its validity, is not truer than others. for example, the population of italy on the date of the population census is an “official” number, certainly, but it is not “true.” the number of residents (or of those present?) in italy is true because it is official and it is official because istat 3 says so. istat is the guarantor of “truth” for two reasons: 1) in accordance with the procedures and the methods of measurement that it uses and 2) because it is an organization of the italian state entrusted with that task. i would say that on account of this role istat is one of the most important institutions of the italian state. if it has been certified by istat the official number of residents becomes the “legal” number and we speak of the “legal” population. nevertheless there have been situations in which the “legal” number of the population of some italian municipalities was (and is) manifestly false: there have at times been, and at times still are, circumstances in which it is convenient for the municipalities to be above or below certain thresholds of population 4. usability of the number and models: the construction of the number is a craft as evidence of how numbers are taken seriously, at least apparently, we can think for example of how many administrations justify the numbers of their expenses with (the numbers) of other measures: the number of residents in a certain zone of the city to decide about the location of health services; the number of passengers on some railway lines to decide on whether they (the lines) should survive or be suppressed; the birth rate expected for the near future in order to organize a school system, the surface area of apartments or the number of family members to plan refuse collection. but where do all these numbers come from? leaving aside the problem of the production and preservation of the sources and in general of the collection of the data, let us focus our attention on the process of their manipulation and processing. the number is usually the result produced by a more or less complex calculation carried out by people who are experts in the techniques required, in particular in statistics. it ranges from the most elementary case, the enumeration of objects, to others that are less elementary, in which the processing of the number in question involves the use of complex computational procedures. in all cases, however, the procedures have a common point of departure the definition of the objects that will be subjected to the abstract procedures. for example, in the case of counting the population it is necessary to agree on whether to count the residents or those present. to count the employed it is necessary to agree on who should be included as employed: only those with an open-ended contract, workers with temporary but renewable contracts, those with temporary contracts for specific projects, or occasional workers. from the beginning there are therefore varying degrees of freedom. and each degree of freedom can be associated with a different interpretation. in other cases, which are less immediate, the production of the “number” in question involves not only counting or an estimate, but also a forecast or, better, a predictive model. the model is the crystal ball that the experts use to make their prophecies. for example, in estimating how much refuse will be collected in a certain area, it is necessary to think of the production of refuse per capita or by family and to attribute a quantity to each family, depending on the number in the family and on the surface area of their 12 iassist quarterly fall 2009 residence. the forecast will thus be “numerable” with an assessed number of people (p) and surface area (s). for example: q=a+b1p+b2s. the coefficients for the weight (b1 and b2 in the example) have to be observed and usually this is done by means of protocols that are realised in surveys. this procedure, whose description we have deliberately excluded from these brief remarks, is at times quite a complex one, and when all is said and done, it is on this procedure that all the results of the estimates depend. the simplest calculation or the most complicated equation are in any case a part of the same set of instruments with which we build estimates on aspects of real life, measured by protocols of observation. with these estimates we construct models that enable us to represent reality and establish by means of successive simulations a rule of behaviour in relation to some objective. in the case of the example, the objective is provided by the excellent distribution of a public service and in establishing for it a fair value for the contribution that families will be called upon to pay as their due. here it is important for us to stress that we are talking about at least two operations that are the task of experts: the definition of the model and the estimation of the same. more precisely, the expert will undertake to translate into a model the customer’s requests, whether the latter is public or private, and to conduct the necessary investigations to estimate its parameters. 5. more on numbers and models in this section we will see how models give sense to numbers. models may be viewed has a particular (formalized) point of view, or conceptual framework or context [paulos, 1998, p. 14 4 ] in which numbers find their true reference or correspondence to the reality that they describe. once the reality has been captured in a net of conceptual frameworks or models, numbers begin to be related to it in an unambiguous way and with an unambiguous meaning. let me show you an example. the example is drawn from a study about school achievement and socioeconomic status (ses) [bonifacio, f. 1987]. the study compares school achievement in two different types of high schools, one of which is of a general type, the other one of vocational type [oecd 1999]. the students of the first one show better achievement than the students in the other one. as students with lower ses are more likely to be enrolled in vocational schools there is a strong relationship between ses and achievement, although this may be not true inside a single school type (see also [raudenbush, s.w., bryk a.s. 2002, p.16-22]). imagine transferring the situation described for two schools into one school and referring the findings related to the observed differences in achievement in two school types to two classrooms. in the first classroom the teacher, aristogitone, is severe; in the other second classroom the teacher is valdo, who is less strict (for the sake of simplicity there is only one teacher per classroom). both of them are teachers who evaluate the achievement of their students in a strictly technical way and in their classrooms the probability of obtaining good achievements is exactly the same for each level of a student’s ses. the only difference is that with aristogitone the probability to fail is 30%, while with valdo it is 10%. both teachers are also perfectly impartial in their evaluations. but knowing that aristogitone’s students are more likely to come from a lower ses we would argue that they are, taken together, biased and unfair with respect to those students. both of these things are true: the teachers are impartial from the point of view of their own classroom; they are biased and unfair from the point of view of the school system. what is changing is the perspective of the observation point: we consider the general situation, the entire system (both the classrooms), and teachers as only one portion of it (just their classrooms). both the opinions are well supported, but what is more likely to happen is that teachers will think that their situation represents the true reality and therefore that there is no relationship between ses and achievement [bonifacio, 1987,p. 98]. this is one of those cases in which true observations give rise to false, unexpected consequences [effets pervers, see boudon 1977]. an analogous example is reported by paulos [paulos, j.h, 1998, p. 39]. the example is drawn from race relations in usa and shows how, supposing the same diffusion of racism among white and black people (10 percent of racists in each group), blacks will suffer disproportionately from racism due to the different marginal distribution of whites and blacks in the population 5. now i think that what i have stated above, that is that the true meaning of numbers comes from the model in which they are located, is clearer: by changing the perspective from which we consider the numbers, the numbers might change their meaning or support different interpretations. but there is a step further: the model may be formalised (in a more or less easy way). for instance the example related to achievement and ses may be formalised with the following expression [bonifacio,f. 1996, p. 45]: i believe that what numbers are telling us is neither the result of biased interpretations introduced by advocates or activists, nor a sum of technical mistakes in methodological issues such as guessing, defining, measuring, sampling [best, p.32-58] that might have been done more or less well. numbers are telling us the result of the application iassist quarterly fall 2009 13 of a model. therefore numbers (or more generally data) tell us things that may be seen as true only inside the model that generated them. in this sense numbers exist only in an interpreted form, shaping them in only one meaning. only in this sense are numbers true both for me and for you: it is just because we understand the underlying model that make them true. just for sake of completeness i shall report my view here on the relationship between natural language used in literature, for example (once again, see queneau), and scientific formalised models, or between storytelling and logic, mathematics or statistics as paulos would say [paulos 1998, p.104-105]. this view is reported in [bonifacio, 1996, p.9] and draws the mentioned relationship, referencing both literature and models to knowledge of the world. this view simply tells us that literature and mathematics are both “right” modes to know the world, but in a sort of specialized and complementary way. reality is not so simple to fit definitely in a model. or vice versa, a model could not be so complex to fit the reality at a delta level, where delta is less than a given ε taken as small as possible (and also even if it were possible it would be useless). all details of reality are not completely described by a model, which is in fact a simplification of the real world that gives us only an averaged rough idea of what is going on in real life. natural language embedded in literature (poems, novels, stories) gives us the possibility to go deeper into the real world, helping to depict also the most individualized and specific aspects of reality. in this way models and literature cooperate to build a more exhaustive world knowledge. literature starts where the model ends, so to speak. finally, this is the ontological possibility to lie that links words and numbers: both try to discover the real world, both do this in the midst of social constraints and relationships that conceal it and shape assertions about it. 6. the social construction of numbers 6 in order for there to be a model and an expert to assess it a problem arises from an objective to be reached which in the example reported in section 4 is the attribution of a fair price for families to pay for refuse collection. when the objectives are of this nature, in other words connected with public services7, the solution to the problem generated involves various social actors and various points of view. in this example the actors are the representatives of the administration that commissions the survey (the customer), the experts of the research firm, the families as they are part of the subject under investigation. the interaction among the protagonists (problem, objectives, assessment, model on the one hand; customer, experts, families on the other) constructs the number that represents the solution to the problem. or better, the number that constitutes the solution of the problem will emerge from this interaction. the expert is a subject of the interaction as he or she knows the process of the construction of the number and knows that the process of construction of the number is “methodologically” guaranteed. the methodology for the expert is not an esoteric mystery. it is composed of the discussion of problems in relation to objectives (for example, how much refuse disposal costs), of measurements, of statistical techniques for the modelling and forecasting, as we have seen in the previous part. the customer is the subject of interaction as he or she knows better than the others the objectives and problems that arise in the attempt to achieve them. above all he or she knows why it is necessary to propose certain objectives. he or she has, in other words, a political vision. the customer is the bearer of interests that in some way preexist the course of the construction of the number and that influence it. in turn the subject of the survey is the subject of interaction as he reacts to stimuli to which he is subjected. he or she may, for example, answer a questionnaire or refuse to do so. in this process of translation of the problem into a model, 14 iassist quarterly fall 2009 of estimate or measurement of the parameters of the same, of further translation of the model into procedures of calculation, of navigation in the midst of motivations, preconstituted interests and recalcitrant subjects of research, lies the objectivity of the “number.” the data are objective because there exists a controllable process that produces them and there exist those who produce them, who assume responsibility for them, a responsibility constructed among several subjects that interact in the attempt to solve the problems. finally, then, subjectivity, in the dual declination of control and responsibility, is the guarantee of objectivity. in other words the truth of the data does not lie, or does not only lie, in its congruity with reality, but also in the fact that we agree, for a series of reasons, to consider it real. this process of social validation of the data constructs the objectivity and gives authority to the data, in other words the data is institutionalised. this process of approval of the data is similar to what occurs during a trial with the jury [popper 1934]: despite the fact that the facts are facts, they become truly facts, in other words objectively facts, only when the members of the jury have confirmed them. qualifying objectivity as a two-way objectivity, one theoretical and empirical and the other one social, sets the research data on a “politically” marked path, thus rendering it permeable to the objectives of the actors involved, above all of the customer. the inclusion of the objectives in the process of production and evaluation of the data (the data is correct if it confirms the objective or at least it does not contradict it) risks flattening it to the existing situation and rendering it sensitive to the distribution of power. the data supports the opinions of those who produce it or, better, of the person who buys it: the process of institutionalisation that we have spoken about, instead, contrasts this outcome, recognising the two-way track from which the data originates, and its dual nature of guarantor of objective reality and subjective will. explicitly considering the point of view of the expert who produces the data, the procedure of institutionalization outlined above protects the data (and their interpretation) against incorrect practices that are aimed at fixing the data (and their interpretation) on preconstituted interests. paradoxically we have come to the point where, to defend the data from interference from politics, it is necessary to put them back inside politics. this can be done by constructing the political instruments for imposing respect for the procedures of construction, production and validation of data. in effect, with numbers we communicate solutions to problems, not objectivity tout court. the numbers are only apparently produced and communicated by an expert (consultant, person or company producing data) for a customer. in reality the numbers are produced by both, in a process that moreover exposes both to the intemperance of the subject under investigation 8. 7. the ultimate number transformation: graphics both parties (the expert and the customer) then communicate the numbers to a third party for whom those numbers become reality. the third party may be the board of directors of a company, the dean of a faculty, the readers of a newspaper, the audience of a tv programme, the participants at a convention. in this phase (but also before, in the meeting between the producer of the data and the customer) the numbers undergo a metamorphosis, which reduces them to static or animated little figures, figures with their own meanings that do not require any further interpretative effort -at last reality is within everyone’s grasp, easy to understand and, above all, objective. the number, stripped of its arabian consistency and dressed up in multi-coloured clothing, guarantees this. i am referring here to another source of misinterpretation of numbers pertaining to the communication world: numbers are not only presented as they are, but transformed in a simpler, more intuitive fashion. this simplification (nearly always graphication) is often an oversimplification in which something important gets lost: the way in which numbers have been built, that is the underlying logic or model 9. that is, as we have seen before, the only way to recognise the truth of numbers. one of the most common misinterpretations is induced by reporting quantities using quantities adverbs: very few, a few or many, a lot of, and so on. the right question, the question that must be modelled and that often gets lost here is: how much is it a few, how much is it much, or many, or a lot? the benchmark measure is often omitted. the same happens with graphics: quantities depend more on how the axes have been scaled than on real numbers (measures). for example, fig. 1 reports the scores achieved by the average customer in a hypothetical customer satisfaction survey on satisfaction for some items (x-axis) and their supposed importance (y-axis) scales. in fig. 2 the same results are reported, shifting the position of the axes. while in fig. 1 the threshold between bad and good is fixed to 5 (in a scale 0-10) in fig. 2 it is fixed at 7.5 instead and the related semantic has been changed accordingly. the model behind fig. 1 says: “values above the expected mean of a random scale ranging from 0 to 10 are to be considered good. therefore values above 5 are good ones in both scales”. this decision means that on average customers consider all items to be pretty good (see fig. 1). if we suspect that for some reasons customers over evaluate the product/services offered we can decide to increase the threshold value to the means of the observed scale, for instance. the model behind fig. 2 says: ”values above the observed mean of the measured scales lying in the observed range are to be iassist quarterly fall 2009 15 considered good. therefore values above 7,5 (for example) are good ones in both scales”. both figures represent true facts, but inside different models or interpretations. knowledge of the underlying models is essential to understand what the graphics mean and knowledge of pre-established goals (or interests) is essential to understand the models. if the interest is to reward workers or sellers for their good performance probably the first model will be chosen. if the interest is to underline aspects that are less appreciated than others by customers and therefore represent critical points the second model is better. in fig. 2 the items “response time” and “ability” are in a critical position, while they are not in fig. 1. in conclusion: graphics must not be taken as direct signal for the truth, but for the expression of model results; graphics, as stakeholders of numbers, must be interpreted inside a model, just as the numbers are. 8. the numerical construction of society despite this “intentionality” [paulos 1998, p. 92] of models, this plasticity of models, today nearly every opinion or policy maker uses numbers (data) to support his arguments without a model. or better they put data inside an empty model, that is a model that models nothing. despite the fact that they use more data than in the past and therefore they need even more than in the past empirical research tools such as opinion poll surveys, they have not developed what i would call an objective feeling, where objective feeling is the special aptitude to respect numbers importance 10 warning maintenance 5 0 improvement exploitation 10 ! ! ! ! ! ! reaction time customer care reliability ability politeness satisfaction fig. 1 – plot of mean values of satisfaction and importance for some customer satisfaction items. value of threshold is 5. importance 10 warning maintenance 5 0 improvement exploitation 10 ! ! ! ! ! ! reaction time customer care reliability ability politeness satisfaction importance 10 warning maintenance 7,5 5 improv 7,5 exploitation 10 ! ! ! ! ! ! reaction time customer care reliability ability politeness satisfaction ! ! ! ! ! ! fig. 2 plot of mean values of satisfaction and importance for some customer satisfaction items. value of threshold is 7,5. and their special ability to represent, in some way, reality. what happens instead is that everyone looks for the most convenient numbers that best apply to a particular, usually his own, point of view. this behaviour does not only have to do with the communication sphere. in the field of communication it would still be comprehensible and even safe. when a person knows what happens, he or she may decide whether it is safe to say it now or tomorrow, whether it is safe to say it in this or in another way. but what is wrong is that customers do not want to hear what numbers would help to discover. they fear novelties and truth, because novelties may be unexpected and truth may be dangerous in a particular and pre-established framework. this, in conjunction with the power that opinion and policy makers usually have for addressing goals by providing budgets for research work, makes it highly likely that there will be a representation of the world which has been obtained through rigged numbers, a biased mirror of the reality. real society does not appear in those numbers. what appear are societies that “i want to make existent and that are more convenient for my existence.” and so there exist as many descriptions of society as there are opinion and policy makers. 16 iassist quarterly fall 2009 conclusion as the picture i have described is rather a bleak and pessimistic one, it is natural to ask the following question: why continue to work with numbers (data)? there are at least two answers to this in my opinion. the first is a cynical one, but i think it is more common than generally believed. it goes more or less like this: “yes, i know that the numbers i am now working on will probably misused and have perhaps also been misworked. but this is what my customers are asking me for and i have to meet their needs. in the end the customer is always right, and he or she is even more right in the special case when he or she is wrong.” behind this way of thinking lies the belief that “all is for the best in the best of all possible worlds” and like voltaire’s candide we may only believe in this world, thinking in the exact way we are expected to think [voltaire, 1745]. the other answer is, to my mind, more challenging, although more expensive in terms of personal investments and less satisfying in terms of job returns (i.e., earnings and social appreciation): “yes, i know that the numbers i am now working on will probably misused. furthermore i am encouraged to work badly, in order to look for the numbers requested instead for the necessary data models. nevertheless, i think that customers have to learn how to know the real world, because only the most advanced knowledge of the real world will help them to succeed. and i will make every effort to help them to get the necessary (and perhaps) right knowledge. despite the pressures i accept the risk of being judged a bad supplier of numbers and i will continue to look for the (perhaps) right description of the world, i.e., for the (perhaps) right numbers instead for the requested numbers.” there are authors that recognise that there are problems in interpreting and presenting numbers 10 , recognise their social nature and the possibility that numbers may be shaped by different interests. their approach for solving the problem is different from the one presented here. i would indicate the former as an illuminist, educational and individualistic approach. illuminist because, in this interpretation, knowledge would suggest the right interpretation; educational because the right interpretation may be taught, or at least the methods to get the right interpretation may be taught; individualistic because the learning process to be used to get the right methods or interpretation is the result of an individual will (or several individual wills). in this view these conditions are necessary and sufficient to guarantee the right interpretation and presentation of numbers, or at least to recognise when the numbers published are suspicious. what i have suggested in this article is that all this is not sufficient, although it may be necessary, if there it is not an explicit agreement upon what has to be considered true or the right interpretation. the agreement must be stipulated among the supporters (advocates or activists) of different views or interpretations or disputed contexts. “supporters” here means every kind of person or institution that may influence the numbers’ production: governmental agencies, private firms, data producers, single researchers or university departments, journalists, political parties, trade unions, etc. as we have seen before, each one of them enters into interaction with the other to build what we call – in short a number. now we can add a new specificity: these supporters interact with their own point of view and their own defined context. these different views have to be collected in a sort of round table where differences will be recomposed (mediated) in the light of the procedures and methods used. while models and context may change according to a particular point of view, procedures and methods do not change across the context. for that reason procedures and methods establish the necessary common language among “supporters”. this capability of methods is science objectivity and the round table is the institution that guarantees its application. just like in democracy, where we have institutions that formally guarantee the equality of rights in a society where rights are unequal, in the number production process we need an institution that formally guarantees the equality of methods, by which models and numbers will be evaluated. what i have written merely signals a risk. the risk of being engaged in dirty, or at least not thoroughly honest, work against our own will. but a person’s will is not sufficient to counter rich and powerful forces. the principal way to counter these tendencies is the institutionalization of the relationships at work, just as they are, as i said before. in other words, the producer, recorders and archivists, customers, experts and so on must be put together to build an institutionalized warning system: a system that systematically surveys not only the physical or logical condition of the data, but first and foremost the use and abuse of the data and the entire process of the production of data and numbers. i know too that all this is not a novelty and that iassist and ifdo in their daily work have already been operating in this way for decades. perhaps one might think that all this would be useless because the custodian and guarantor of objectivity, the round table, is the university: i have some doubts that at present the university can do this alone. indeed there are some good reasons for stating that it is too late and that politics now governs procedures that do not belong to it (those of scientific research): for example, to determine careers. moreover, others will say that in the moment that the university enters the market, not so much through individual practices, consultancy by its professors which has always existed but as an institution that sells its own resources, it dirties its hands and can no longer act as disinterested guarantor (super partes). in these circumstances academia seems no longer able to act as the umpire. iassist quarterly fall 2009 17 i do not wish to be too drastic: it is certain, however, that the situation is not easy. anyone who bases his work on the expectation that the data are, in some way, immune to interested speculation, that there exists a hard core of knowledge, should be prepared to give the university a hand in operating as guarantor. this until the desire for honour and laurels rather than base profit returns to being the just aspiration of its followers and mentors 11. these short remarks are intended just for the record: it absolutely necessary for this alert system to survive to take advantage of the contribution that undoubtedly empirical social research may provide for a better and unbiased knowledge of the world. and when i say this i am thinking above all of my own country. notes 1 flavio bonifacio, metis ricerche srl, , via camerana 6, i-10128 torino, italy. contact e-mail: flavio.bonifacio@ metis-ricerche.it. 2 that is shaped by particular point of view:”…in short, even official statistics are social products, shaped by the people and organizations that create them…” [best, j. 2001, p. 26]; ”…advocates who conduct their own surveys can decide how to interpret the results..” [best, j. 2001, p. 48]. best’s book contains a lot of examples of uses and misuses of numbers to which i refer readers for reference. 3 istat is the italian national institute of statistics 4 “while some figures are almost self-explanatory, statistics without any context always run the risk of being arid, irrelevant, even meaningless.” 5 the probability that paulos and i both gave not only almost the same example, but also used actors with almost the same name, waldo [paulos, j.a. 1998, p.91] and valdo respectively, is very low. in some sense we always have to be careful with numbers, even with small numbers, in spite “of the stunning insignificance of the vast majority of coincidences” [paulos, 1998, p. 61] 6 the title of this section is not just aping the famous book of berger and luckmann [berger, luckmann 1966]. i am deeply convinced that numbers are fundamental in representing reality and therefore subject to the same constraints as reality is. in the same sense see [best, j. 2001, p. 27]: “all statistics are created through people’s actions: people have to decide what to count and how to count it, people have to do the counting and the other calculations, and people have to interpret the resulting statistics, to decide what the numbers mean. all statistics are social products, the results of people’s efforts”. see also [paulos, 1998, p. 84] 7 in reality the field of application is general and it regards the whole world of social research. 8 once again i shall turn to fundamental pieces of literature for help, and once again i shall refer the reader to queneau. two books are fundamental: the flight of icarus [queneau, r. 1973] and the blue flowers [queneau r., 1985]. in these books queneau uproots the protagonists from their context (novels in the case of the former, stories in the case of the latest) and makes them build their own “context” across the original ones from which they have been abruptly extracted. in the same sense the reader may be referred to the castle of crossed destinies by italo calvino [calvino, i. 1973]. in this book the context is randomly selected from and inspired by tarot cards. 9 here i am not thinking of the several ways in which liars may manipulate their graphic tools to induce false representations of reality. see for example [jones, g.e. 2007]. the “graphication” may lead to a false representation of reality because it is further away from reality than models and numbers are. 10 among them the quoted authors [belt, j. 2001] and [paulos, j.a. 1988, 1998]. 11 although the question is an ancient one: francesco petrarca, the great poet, already noted in the fourteenth century: “qual vaghezza di lauro, qual di mirto?/povera et nuda vai philosophia,/dice la turba al vil guadagno intesa./pochi compagni avrai per l’altra via/…” [petrarca, f. 1366]. the quotation is just paraphrased in the text bibliography best, j. 2001, damned lies and statistcs, university of california press, berkeley 2001 berger,luckmann, 1966 the social construction of reality. a treatise in the sociology of knowledge, 1966 bonifacio, f. 1987 – atteggiamento didattico, selezione nella scuola e differenze di fronte all’istruzione, angeli, milano, 1987 bonifacio, f. 1996 – pour un modèle scientifique du système scolaire, harmattan, paris, 1996 boudon, r. 1977, effets pervers et ordre social, puf, paris, 1977 calvino, i. 1979, the castle of crossed destinies, harvest pbk, 1979, italian edition, il castello dei destini incrociati, einaudi, torino 1973 jones,g. e. 2007,how to lie with charts, la puerta, s. monica, 2007 18 iassist quarterly fall 2009 paulos,j.a. 1988, innumeracy, hill and wang, new york, 1988 paulos,j.a. 1998, once upon a number, basic books, new york, 1998 queneau r. 1981, exercises of style, transl. barbara wright, new directions paperbook, 1981, first french edition, exercices de style, gallimard, paris, 1947 queneau r. 1973, the flight of icarus, transl. barbara wright, new directions paperbook, 1973, first french edition, le vol d’icare, gallimard, paris, 1968 queneau r. 1985, the blue flowers, transl. barbara wright, new directions paperbook, 1985, first french edition, les fleurs bleues, gallimard, paris, 1965 oecd 1999, classifiying educational programmes, 1999 edition petrarca, f. 1366, rerum vulgarium fragmenta, einaudi, canzoniere, torino, 1964 popper k.r. 1934, logik der forschung, engl. transl. the logic of scientific discovery, 1959 raudenbush s.w., bryk a.s. 2002, hierarchical linear models, applications and data analysis methods, second edition, sage publications inc., thousand oaks, california, 2002 voltaire, 1745 candide ou l’optimisme, librerie generale francaise, 1983 vol25s.1 iassist quarterly spring 2001 5 understanding barriers to the use of numeric data in learning and teaching by robin rice1* background uk higher education is rich in numeric datasets. in the socioeconomic field, for example, there are large-scale, representative sample surveys (e.g., general household survey), current and historical population censuses, international comparative datasets, longitudinal surveys, economic time series, and data about markets, companies, and commerce. in the uk a centrally funded system of national data services for higher education provides for the dissemination of much of this research data, which is free at the point-of-use and accessible over the internet (via janet, the uk academic network). however, these data resources are under-used in the learning and teaching environment. despite the potential gain in numeracy, critical use of evidence and empiricallybased knowledge by students conducting data analysis at both the postgraduate and undergraduate levels is infrequent, and obstacles exist that make integration of numeric data resources into coursework difficult. employing numeric data effectively in teaching requires specialised skills and more time for preparation than the use of printed materials or bibliographic databases, and both students and teachers require a high level of support. as expectations about the use of information technology in learning and teaching rise, the barriers that inhibit the use of this wealth of data in the classroom and in student projects need to be lowered. understanding statistical evidence is important not just for postgraduates learning to be researchers and entering the professions, but for undergraduates as well. milo schield has written widely about teaching statistical literacy in higher education. he explains it as a different and more fundamental skill than producing or ‘doing’ statistics: “statistical literacy focuses on making decisions using statistics as evidence just as reading literacy focuses on using words as evidence. statistical literacy is a competency just like reading, writing, or speaking.”2 the need for application of such a competency in many fields is readily apparent. the numeric data project this paper reports findings from a national collaborative project: “using numeric datasets in learning and teaching,” funded by the jisc (joint information systems committee, which itself is funded by the higher education funding councils). the lifetime of the project is february 2000 to september 2001. project partners are from three national data centres, edina, mimas, and the data archive, and two university data libraries, the university of edinburgh and the london school of economics. additionally, a task force of experienced academics from across the uk was recruited as volunteers to guide the enquiry and its outcomes. this partnership reflects the novel perspective taken by the project to examine use of the nationally-funded data services with particular reference to local support needs of teachers and learners within their universities. the project is one of several funded under the jisc’s learning and teaching development programme (see http:// www.jisc.ac.uk/dner/programmes/projects/ for a full list of projects). a major objective of the project was to generate knowledge on issues such as the extent of use and the practicalities of using data in teaching, and the experiences teachers have of data support from both national data services and support staff in local institutions. since user surveys tend to target those already registered for national services, there is no ready evidence about the larger population of uk university teaching staff on these issues. therefore, a nationally representative sample survey was needed to discover the current “state of play” before recommendations about how to lower barriers could be made. the survey was designed to ask teaching staff about their use of numeric data in teaching and supervising students, their experience of national data services, barriers to using data in teaching, and the extent of support available within their institutions. the teachers’ survey was enhanced by qualitative case studies of a diverse set of postgraduate and undergraduate classes using numerical data in teaching, which both inform the enquiry and also act as exemplars for other teachers. the full survey results and case studies are available on the project web site at http://datalib.ed.ac.uk/projects/ datateach.html. the final report with its recommendations, teaching resources, and other information is also available. http://www.jisc.ac.uk/dner/programmes/projects/ http://www.jisc.ac.uk/dner/programmes/projects/ http://datalib.ed.ac.uk/projects/datateach.html http://datalib.ed.ac.uk/projects/datateach.html 6 iassist quarterly spring 2001 survey methodology a sample postal survey was conducted of uk university teaching departments within the social sciences, plus other selected disciplines “outside” the social sciences, such as public health sciences. two hundred sixty-seven department heads were randomly selected from a universe of 1590 (1 in 6 sampling fraction). the sampling frame was purchased from the marketing company mardev, extracted from the worldwide academic & library file. department heads were asked to complete the four-page questionnaire themselves and to pass copies to relevant teaching colleagues to garner their participation. (a web version was also made available for on-line input.) there were 206 responses collected from 110 departments. fifteen records were removed as ineligible (e.g. non-teaching department). following telephone, e-mail, and postal follow-up requests to sample members, the final response rate (110 / 252) was 44 percent of departments sampled. survey results: use of data in teaching and learning due to the survey design and instructions to department heads, there was likely a skew toward data users among those in the sample who participated, as a result of selfselection. (non-data users tended not to respond to the survey, as it was not felt to be relevant to them.) seventynine percent of those survey respondents who taught or convened courses used data either “nearly always,” “often,” or “occasionally” (see chart 1). the sample also seemed to over-represent senior staff (perhaps because the request was sent to department heads), teachers of methods courses, and those committed to quantitative analysis. among those who used numeric data in teaching in some form, about two-thirds expected students to work with data chart 1: use of numeric data in this class by percent (n=181). on a computer, in “hands-on” fashion. as table 1 shows, a higher proportion of methods courses were hands-on than subject courses. [the categories of “methods-based” and”“subject-based” were coded during analysis, based on names of courses supplied by respondents.] surprisingly, neither course level nor class size appeared to affect whether the course was hands-on. table 1: whether course is “hands-on,” by course type although the survey was directed towards staff, not students, there was an attempt to understand the level of data use by students in their independent learning. ninetytwo percent of respondents who were either postor undergraduate supervisors recommended the use of numeric data for students’’ research at least occasionally (depending on the nature of the research project). below are “typical” responses for each category. • nearly always do (35 percent): “statements made need to be backed up with evidence – often of an empirical nature.” • often do (33 percent): “depends on topic, but statistical sources can contextualise a topic.” • only occasionally (21 percent): “many students are more inclined to qualitative research.” • never have and don’t plan to (6 percent): “not relevant to what i am teaching.” • haven’t yet but would like to (2 percent): “not always appropriate and [i am] insufficiently briefed on numeric data available.” burden of data preparation the survey instrument dealt directly with the issue of how burdened teachers felt regarding data preparation. as chart 2 shows, a slight majority felt that data preparation was a burden, but warranted. %loc sdohtem tcejbus lla no-sdnah 58 45 46 no-sdnahton 51 64 63 =n 64 001 641 iassist quarterly spring 2001 7 chart 2: burden of data preparation, percentage of respondents (n=140 respondents were also asked if they felt the need to update / refresh / revise the data used on a regular basis. of those responding (78 percent of those eligible, n=142), 57 percent said yes, and only 14 percent said no. however, 29 percent said yes, but there was insufficient time to do so. data sources and use of national data services the survey showed quite clearly that, although the use of numeric data among the survey respondents is high, the use of national data services that provide onor off-line access to secondary datasets is not. only one-quarter of the respondents who used data in teaching had “used or considered using” the national academic data services (namely the data archive, edina, and mimas) for teaching purposes. so what are the sources of numeric data used in higher education classes? most strikingly, half the teachers either required their students to collect their own data, or taught with data they collected themselves (see chart 3). nearly half, 44 percent, used print data sources, extracted from a monograph or serial. (print publications obviously do not provide the material needed for a “hands-on” component, which gives students practice at manipulating data on a computer, unless the data are hand-entered.) the rest of the sources, including from a colleague, freely available on the internet, or bundled with a textbook, were used by less than 20 percent of teachers who use data. twice as many respondents received data from a government agency or “directly from the data producer” as were registered with a national data service. these results indicate a need to further explore the nature of data sources needed by particular disciplines for teaching particular types of courses, and whether the national data services and local institutions are providing adequate collections. the findings also seem to undermine the notion that anything needed can be obtained freely on the internet. financial and company datasets, for example, are profitable information commodities, which require substantial academic discounts or subsidies to be affordable. would the national data services be more widely used if they were providing relevant collections to teaching departments? a closer look at the barriers to use of the national data services uncovers deeper issues than just ensuring that available sources exist. chart 3: source of data used in class (counts, n=181). 8 iassist quarterly spring 2001 barriers to using datasets in teaching those 46 respondents who were familiar with the national data services (one-quarter of those who teach with data) were asked to rank eight factors they thought might act as barriers in using national data services for learning and teaching purposes. table 2 shows the median score for each barrier, in descending order, and also the mean score. the two top-rated barriers were “lack of awareness of relevant materials,” and”“lack of sufficient time for preparation.” this issue was highlighted in a separate question, in which 57 percent agreed on the need to update /refresh /revise datasets used for teaching, but 29 percent had insufficient time to do so. the third greatest barrier was “registration procedures” [of the national data services]. however, the other barriers received high enough scores to also be considered seriously: namely, difficult data extraction interfaces, unsuitable file formats, inadequate dataset documentation, and lack of tailored teaching subsets. in an open-ended question, users were asked for positive changes the national services could make to support teachers and learners in the use of datasets. thirty-six out of 46 eligible respondents answered the question with a variety of useful suggestions. answers were grouped into the following four categories (with examples of actual responses): • easier access able to get data without learning special software.” • simple registration for students ––“make registration procedures simple and abolish restrictions on use (e.g. all students signing disclaimers).” • create relevant and interesting teaching datasets – –“rapid access to key summary economic data in form tailored for teaching.” • effective publicity ––“the initiative needs to come naem erocs naidem erocs slairetamfossenerawafokcal 5.6 7 noitaraperprofemitfokcal 4.6 7 serudecorpnoitartsiger 6.5 6 ecafretni 0.5 5 stesatadfotamrof 8.4 5 noitatnemucod 6.4 5 stesbusgnihcaetfokcal 4.4 5 sreirrabfogniknaregareva:2elbat .)tsewol=1,erocstsehgih=8( from the national services but better publicity would be a start.” support issues prior to the survey, only anecdotal evidence was available to determine how teachers obtained support for classroom use of datasets. members of the task force were familiar with the common reality of peer support for data use in both research and teaching via word-of-mouth. one member was aware that he was considered to be “the data guy” in the department, to whom others came for support. although two data librarians were involved in the project, specialised data libraries and data librarians are not common in uk universities. site representatives for the national data services can be based in the library, computing service, or elsewhere in an institution, but it was not known how much support they actually provide to users. to provide a baseline measure on this issue, the survey asked each respondent,”“from whom have you ever had support in obtaining or using data, whether for teaching or for research?” ” of those who responded, more than a third (37 percent) had received no support at all. more than one source could be ticked; the average number of sources of support received was two. peer support was the most common form, either from a project co-worker/assistant or another colleague (26 percent and 47 percent, respectively). the local computing service (26 percent) was roughly matched with the local library service (23 percent), which had helped about a quarter of respondents each. national service staff provided help to 10 percent of respondents, and their local site representatives only helped 7 percent of chart 4: level of local support provided, percentage of respondents (n=176). iassist quarterly spring 2001 9 them. as an indicator of the satisfaction level with this status quo, users were asked to characterise the level of data support provided in their institution. the results are shown in chart 4. notably, only 14 percent agreed that local support was “ very good across the board.” the majority, 62 percent, felt that support “tends to be ad-hoc.” to follow this up, the survey instrument anticipated a number of local support activities and asked respondents to tick all “forms of locally provided support needed by academic data users.” those who responded to this question (162 or 79 percent of total) reinforced the need for a number of forms of locally provided support, above all “data discovery / locating sources” (66 percent). all of the answers shown in chart 5 received “votes” from between one-third and two-thirds of those responding. the average number of needs ticked was three. an open-ended follow-up question tended to reinforce the forms of support suggested in the questionnaire, although a significant minority felt that no additional support was needed, or expressed concern about where the resources would come from. chart 5: forms of local support needed (counts, n=162). recommendations the task force and the project team provided the following recommendations to the jisc (project funder) at the close of the project. further elaboration may be found on the project web site. 1. a broad initiative is recommended to promote subject-based statistical literacy for students, coupled with tangible support for academic teaching staff who wish to incorporate empirical data into substantive courses. 2. the development of high-quality teaching materials for major uk datasets needs to be funded adequately, in order to provide salience to subject matter and demonstrate relevant methods for coursework. 3. the national data services need to improve the usability of their datasets for learning and teaching. 4. a more concerted and co-ordinated promotion of the national data services should then follow, which is responsive to user demand. 5. universities should develop it strategies that include data services and support for staff and students, and integration of empirical datasets into learning technologies. conclusion uk higher education is undergoing many changes. the renewed attention to “learning and teaching” is an impetus for change in university teaching practices. advances in information technology are creating new spaces for learning beyond the traditional classroom, and forms of teaching beyond the traditional lecture. yet the pressures on academic staff who are still rewarded primarily for research rather than innovative teaching are great. to ensure that statistical literacy is taught effectively, new products and resources must be developed and adequate levels of support and technology provided. 1 with acknowledgments to the project team: peter burnhill (project director), melanie wright, sean townsend; joan fairgrieve for statistical analysis; and the task force on the use of numeric data in learning and teaching. for membership see http:// datalib.ed.ac.uk/projects/datateach/ participants.html 2 m. schield (1999). “statistical literacy: thinking critically about statistics.””of significance (journal of the association of public data users):1. available as of 14 sep. 01: http://www.augsburg.edu/ppages/~schield/ milopapers/984statisticalliteracy6.pdf * paper presented at the iassist/ifdo conference 2001, amsterdamrobin rice, edinburgh university data library http://datalib.ed.ac.uk/projects/datateach/participants.html http://datalib.ed.ac.uk/projects/datateach/participants.html http://datalib.ed.ac.uk/projects/datateach/participants.html http://www.augsburg.edu/ppages/~schield/milopapers/984statisticalliteracy6.pdf http://www.augsburg.edu/ppages/~schield/milopapers/984statisticalliteracy6.pdf iassist quarterly 2014 5 iassist quarterly editor’s notes meaning and translations, thesauri and the potential of data archives welcome to the second issue of volume 38 of the iassist quarterly (iq (38):2, 2014). this issue brings you a potpourri of papers from the very theoretical or philosophical through the technical issues faced in wellestablished data archives to the supportive reasoning on the creation of new national data archives. this issue has two papers with the title ‘from data to the creation of meaning’. at the 2014 iassist conference in toronto, kristin partlo and justin joque presented their work in two parts in the session ‘developing meaningful data support roles and services’. a slide from the presentation reads: ‘our slides are rather sparse, but we are planning on publishing the papers in iassist quarterly ‘. i am happy to say that their planning worked very well! justin joque works as the visualization librarian at the university of michigan; his paper ‘from data to the creation of meaning part 1: unit of analysis as epistemological problem’ addresses how all data creation processes are based upon assumptions about the world and what existing unit can be analyzed. the assumptions present a hurdle when attempting to make data meaningful across disciplines. an example of the ideological framework that can generate incommensurabilities of datasets is the zip coding used in the united states. (probably similar examples can be found everywhere). zip codes are linked to addresses and not to areas, but they are often aggregated to use for areas. furthermore zip codes are dynamic and sometimes changed for greater efficiency of the postal system. this can present major problems, for example when research defines the unit of analysis. another example is the ‘family’ as a unit where political issues are involved. justin joque points to work of linnaeus and aristotle. when discussing categorization, i also recommend the book ‘women, fire, and dangerous things’ (lakoff, 1987). kristin partlo presents part 2 of this work. she works as reference & instruction librarian for social sciences and data at carleton college in minnesota. the paper ‘from data to the creation of meaning part 2: data librarian as translator’ highlights how data professionals use their role as translator for new students and researchers in a field, as well as in communication with people from other subject areas thus bridging and translating for meaning. kristin partlo also mentions how a shift away from her primary expertise combined with a similar shift to outside the expertise of the faculty made it possible to find common ground. kristin brings our awareness to the fact that: ‘infrastructure is something most people don’t see or think about until it breaks down’. data librarians show the challenges of aligning infrastructure, both technical and cultural. the 2014 iassist conference had a session on ‘harmonization, thesauri and indexing’. a paper by lorna bell and lucy bell describes a project that more literally concerns the issue of translation. the authors both work at the uk data archive at the university of essex. the data archive maintains a multilingual thesaurus to facilitate cross-national data retrieval. the work on this european language social science thesaurus (elsst) is based upon forty years of development of the monolingual thesaurus (hasset). the paper describes some of the developments and solutions. one of the important tasks of maintenance of a thesaurus is keeping the terms up-to-date, especially for the support of automatic indexing. the elsst project also focuses on moving from a term-based to a conceptbased thesaurus. elaborated thesauri are developed within the frame of standards; for example three types of mapping between thesauri equivalence, hierarchical, and associative defined in iso 25964. the controlled vocabularies in the thesauri support machine-to-machine communication and the development of a hub for social science data, and open up the world of linked data and the semantic web. at the 2013 iassist conference in cologne the session ‘(serscida) making new connections: developing data services in bosnia and herzegovina, croatia, and serbia’ included the presentation ‘research landscapes in social science data archiving in bosnia, croatia, and serbia’ by aleksandra bradic-martinovic. in this issue aleksandra has teamed up with aleksandar zdravkovic in producing the paper ‘researchers’ interest in data service in bosnia and herzegovina, croatia, and serbia’. new data archives are being brought into existence but even now there are no digital archives in social science in the western balkan region. this paper shows the potential for establishment of social sciences digital data archives in bosnia and herzegovina, croatia, and serbia. the potentials are presented through statistical results from a cross-country survey of social science researchers in these countries supported by an eu-project and cessda. the investigation found a positive attitude towards data sharing and the benefits of data archives. the paper describes the methodology, the characteristics of the researchers and the data they have produced. the researchers are very supportive of the idea of having better access to secondary data and there is a growing trend of producing datasets; however, the datasets are seldom provided with sufficient metadata. articles for the iassist quarterly are always very welcome. they can be papers from iassist conferences or other conferences and workshops, from local presentations or papers especially written for the iq. when you are preparingblished in the iq. chairing a conference session with the purpose of aggregating and integrating papers for a special issue iq is also much appreciated as the information reaches many more people than the session 6 iassist quarterly 2014 iassist quarterly participants, and will be readily available on the iassist website at http://www.iassistdata.org. authors are very welcome to take a look at the instructions and layout: http://iassistdata.org/iq/instructions-authors authors can also contact me via e-mail: kbr@sam.sdu.dk. should you be interested in compiling a special issue for the iq as guest editor(s) i will also be delighted to hear from you. karsten boye rasmussen febuary 2015 editor http://www.iassistdata.org http://iassistdata.org/iq/instructions mailto:kbr@sam.sdu.dk i^^sist newsletter vol. 1, no. 3 this meeting was also attended by observers from france, sweden and switzerland and a representative of the division for the international development of the social sciences of unesco. on sunday, 21st may, they decided on the establishment of an international federation of data organizations (if-do), designated mr. guido martinotti (adpss, milan) as its president, and entrusted the duties of secretary to mr. erwin scheuch (za, cologne). this federation is open to all organizations prepared to participate in the federation and cooperate in the continuing development of data archiving services in the social sciences. more information or a copy of the status may be available by contacting the secretary: dr. erwin scheuch zentralarchiv fur empirische sozialforschung universitat zu ktiln d 5 koln (deutschland) bachermerstrasse, 40 data organization and management the laboratory for computer graphics and spatial analysis within the graduate school of design at harvard university has just released a new edition of lab-log, its catalog of computer programs, data bases and publications. research at the laboratory is principally concerned with the analysis and graphic display of geographic data used in the planning process. lab-log describes various products which have resulted from this work and which are currently available for distribution to universities, government agencies and private organizations. lab-log includes a description of six different computer programs for use in the graphical display of spatial data via a line printer, line plotter and cathode ray tubes. a wide variety of cartographic (x-y coordinate) data bases are also described. publications are available on the subjects of automated cartography, theoretical cartography and theoretical geography. lab-log also contains a brief description of the laboratory's history, research directions and operating policies copies of lab-log are available at a cost of $1.00 each upon request to: the laboratory for computer graphics and spatial analysis 520 gund hall harvard university 48 quincy street cambridge, ma 02138 payment must accompany your order. 32. 4 iassist quarterly 2015 iassist quarterly editor’s notes data documentation initiative results, tools, and further initiatives welcome to the third issue of volume 39 of the iassist quarterly (iq 39:3, 2015). this special issue is guest edited by joachim wackerow of gesis – leibniz institute for the social sciences in germany and mary vardigan of icpsr at the university of michigan, usa. that sentence is a direct plagiarism from the editor’s notes of the recent double issue (iq 38:4 & 39:1). we are very grateful for all the work mary and achim have carried out and are developing further in the continuing story of the data documentation initiative (ddi), and for their efforts in presenting the work here in the iassist quarterly. as in the recent double issue on ddi this special issue also presents results, tools, and further initiatives. the ddi started 20 years ago and much has been accomplished. however, creative people are still refining and improving it, as well as developing new areas for the use of ddi. mary vardigan and joachim wackerow give on the next page an overview of the content of ddi papers in this issue. let me then applaud the two guest editors and also the many authors who made this possible: alerk amin, rand cooperation, www.rand.org, usa ingo barkow, associate professor for data management at the university for applied sciences eastern switzerland (htw chur), switzerland stefan kramer, american university, washington, dc, usa david schiller, research data centre (fdz) of the german federal employment agency (ba) at the institute for employment research (iab) jeremy williams, cornell institute for social and economic research, usa larry hoyle, senior scientist at the institute for policy & social research at the university of kansas, usa joachim wackerow, metadata expert at gesis leibniz institute for the social sciences, germany william poynter, ucl institute of education, london, uk jennifer spiegel, ucl institute of education, london, uk jay greenfield, health informatics architect working with data standards, usa sam hume, vice president of share technology and services at cdisc, usa sanda ionescu, user support for data and documentation, icpsr, usa jeremy iverson, co-founder and partner at colectica, usa john kunze, systems architect at the california digital library, usa barry radler, researcher at the university of wisconsin institute on aging, usa wendy thomas, director of the data access core in the minnesota population center (mpc) at the university of minnesota, usa mary vardigan, archivist at the inter-university consortium for political and social research (icpsr), usa stuart weibel, worked in oclc research, usa michael witt, associate professor of library science at purdue university, usa. i hope you will enjoy their work in this issue, and i am certain that the contact authors will enjoy hearing from you about new potential results, tools, and initiatives. articles for the iassist quarterly are always very welcome. they can be papers from iassist conferences or other conferences and workshops, from local presentations or papers especially written for the iq. when you are preparing a presentation, give a thought to turning your one-time presentation into a lasting contribution to continuing development. as an author you are permitted ‘deep links’ where you link directly to your paper published in the iq. chairing a conference session with the purpose of aggregating and integrating papers for a special issue iq is also much appreciated as the information reaches many more people than the session participants, and will be readily available on the iassist website at http://www. iassistdata.org. authors are very welcome to take a look at the instructions and layout: http://iassistdata.org/iq/instructions-authors authors can also contact me via e-mail: kbr@sam.sdu.dk. should you be interested in compiling a special issue for the iq as guest editor(s) i will also be delighted to hear from you. karsten boye rasmussen september 2015 editor iassist quarterly 2015 5 iassist quarterlyiassist quarterly this issue features four papers that look at leveraging the structured metadata provided by ddi in different ways. the first, “design considerations for ddi-based data systems,“ aims to help decisionmakers by highlighting the approach of using relational databases for data storage in contrast to representing ddi in its native xml format. the second paper, “ddi as a common format for export and import for statistical packages,” describes an experiment using the program stat/transfer to move datasets among five popular packages with ddi lifecycle as an intermediary format. the paper “protocol development for large-scale metadata archiving using ddi lifecycle” discusses the use of a ddi profile to document closer (cohorts and longitudinal studies enhancement resources, www.closer.ac.uk), which brings together nine of the uk’s longitudinal cohort studies by producing a metadata discovery platform (mdp). and finally, “ddi and enhanced data citation“ reports on efforts in extend data citation information in ddi to include a larger set of elements and a taxonomy for the role of research contributors. mary vardigan vardigan@umich.edu joachim wackerow joachim.wackerow@gesis.org new perspectives on ddi 1/3 rasmussen, karsten boye (2020) editor’s notes: sharing open data without risk, and with machine-actionable provenance metadata, iassist quarterly 44(4), pp. 1-2. doi https://doi.org/10.29173/iq990 editor's notes: sharing open data without risk, and with machine-actionable provenance metadata welcome to the fourth issue of 2020 and the last issue of volume 44 of the iassist quarterly (iq 44(4) 2020). at a future time, there might be a special issue of iassist quarterly on 'corona data'. right now, the numbers are rising as we are entering winter, but on the other hand vaccination is around the corner. i hope that only 2020 will be remembered as the year of the coronavirus, and that 2021 will bring us better times. open data and the sharing of data, including public data, is a foundation of democracy – by making data available for all and not only for the current political leaders. we are talking about free data like we talk about free speech. data archives around the world are contributing to making this possible, and this issue reveals the efforts made at the czech social science data archive towards involvement of the users in data sharing. we have to obtain a balance, and i shall return here to the phrase 'as open as possible, as closed as necessary'. open data must not compromise individuals and their right to privacy, and the need for protection of sensitive data. in this issue you will also find guidelines for assessing risks embedded in the data as well as remedies for de-identification and anonymization. such instructions for safe sharing and depositing of datasets are central for data producers, researchers, and students. we must tolerate speech we disagree with – and then we might decide to take the time and effort to refute the statements. likewise, we might experience data that call for scrutiny. for that purpose, the highest level of metadata is needed. i recall that many years ago, many hours were spent at the danish data archives trying to figure out how a central variable in an important survey was constructed. the structured data transformation language will solve such issues, by providing access to provenance metadata harvested from transformation scripts of common statistical languages. again, another accomplished contribution to the improvements in the sharing of quality data. great thanks are due to all the contributors at data archives, library schools, research institutions, and many more. michaela kudrnáčová and ilona trtíková are working at the czech social science data archive. michaela is a phd student with a focus on research and methodology, and ilona is data manager with expertise in retrieving and sharing research information. their work at the data archive includes a great deal of communication with students and researchers on issues of data management and data analysis. the archive views its role in open science as creating a trusted and sustainable environment for social science data. in order to better respond to users' demands and needs, data were collected on their users from several sources including user registrations, a survey, and interviews. the article on 'sustainability through the liaison with data archive users' thus includes some statistics on the users of the czech social science data archive (csda) since its availability for online data storing and sharing through nesstar. the presentation of the survey results includes distributions of purpose and frequencies of use of the csda. a very positive conclusion was the high degree to which users express their willingness to cooperate in the development of the csda functions and services. the investigations also revealed specific areas and functions that could be https://doi.org/10.29173/iq990 2/3 rasmussen, karsten boye (2020) editor’s notes: sharing open data without risk, and with machine-actionable provenance metadata, iassist quarterly 44(4), pp. 1-2. doi https://doi.org/10.29173/iq990 improved, and they are planning now to perform a short online survey of csda users every two years, supplemented with interviews. while data sharing is encouraged, at the same time sharing data from surveys often implies a risk for identification of individuals. that situation is addressed by the submission titled 'mathematics, risk, and messy survey data'. the authors provide researchers and data collectors with a better understanding of the theory and concepts of anonymization and risk assessment. the obvious first step in data anonymization is the removal of specific individual information, e.g., name, telephone numbers, and social media identifiers. however, demographic variables (quasi-identifiers) such as occupation and geography might also be sufficient to identify individuals, for example a small town's only doctor. many possible such quasi-identifiers might exist in a dataset. however, deleting quasiidentifiers may seriously decrease the value of the dataset. one widely used technique for ensuring anonymity is k-anonymity, which assesses how many records in the dataset have the same combination of quasi-identifiers. the authors illustrate the utility and challenges of k-anonymity by walking readers through the anonymization of two datasets. the authors kristi thompson and carolyn sullivan are at western university, canada, where kristi thompson is the research data management librarian and carolyn sullivan is a student of information and media studies. the article 'provenance metadata for statistical data: an introduction to structured data transformation language (sdtl)' presents a truly collective effort. the authors are george alter, darrell donakowski, jack gager, pascal heus, carson hunter, sanda ionescu, jeremy iverson, h v jagadish, carl lagoze, jared lyle, alexander mueller, sigbjorn revheim, matthew a. richardson, ornulf risnes, karunakara seelam, dan smith, tom smith, jie song, yashas jaydeep vaidya, and ole voldsater representing the institutions university of michigan, metadata technologies north america, algenta technologies, norwegian centre for research data, and norc. the continuous capture of metadata for statistical data project (c2metadata) created sdtl to capture provenance metadata from data transformation scripts found in statistical analysis software. the objective is to provide standardized, machine-actionable documentation to answer questions like 'which original variables were used to construct this derived variable?' sdtl works with metadata standards, like ddi, to add fileand variable-level provenance to data catalogs and codebooks. software created by the c2metadata project translates commands of five leading statistical packages (spss, stata, sas, r and python) into sdtl. sdtl covers basic data transformation commands for assigning and recoding values, metadata commands for setting labels and data types, and file-level processes like merging and appending rows. the article describes the sdtl approach and discusses similarities and differences among statistical analysis software. enjoy reading the three articles. submissions of papers for the iassist quarterly are always very welcome. we welcome input from iassist conferences or other conferences and workshops, from local presentations, or papers especially written for the iq. when you are preparing such a presentation, give a thought to turning your one-time presentation into a lasting contribution. doing that after the event also gives you the opportunity of improving your work after feedback. we encourage you to login or create an author profile at https://www.iassistquarterly.com (our open journal system application). we permit authors to have 'deep links' into the iq as well as deposition of the paper in your local repository. https://doi.org/10.29173/iq990 https://www.iassistquarterly.com/ 3/3 rasmussen, karsten boye (2020) editor’s notes: sharing open data without risk, and with machine-actionable provenance metadata, iassist quarterly 44(4), pp. 1-2. doi https://doi.org/10.29173/iq990 chairing a conference session or workshop with the purpose of aggregating and integrating papers for a special issue iq is also much appreciated as the information reaches many more people than the limited number of session participants and will be readily available on the iassist quarterly website at https://www.iassistquarterly.com. authors are very welcome to take a look at the instructions and layout: https://www.iassistquarterly.com/index.php/iassist/about/submissions. authors can also contact me directly via e-mail: kbr@sam.sdu.dk. should you be interested in compiling a special issue for the iq as guest editor(s) i will also be delighted to hear from you. karsten boye rasmussen december 2020 https://doi.org/10.29173/iq990 https://www.iassistquarterly.com/ https://www.iassistquarterly.com/index.php/iassist/about/submissions mailto:kbr@sam.sdu.dk microsoft word 49-3-brodsky.docx 1/23 brodsky, meryl & hannah chapman tripp (2025). data competencies for liaison librarians: a scoping review, iassist quarterly 49(3), pp. 1-23. doi: https://doi.org/10.29173/iq1154 the creative commons-attribution-noncommercial license 4.0 international applies to all works published by iassist quarterly. authors will retain copyright of the work and full publishing rights. data competencies for liaison librarians: a scoping review meryl brodsky1 and hannah chapman tripp2 abstract this retrospective scoping review explores the data-related competencies required by liaison and subject librarians to effectively support academic researchers. despite the growing demand for research data assistance, many librarians lack formal training (tenopir et al., 2014) or confidence (cox et al., 2012) in this area, often relying on self-taught skills. the objective of this review was to map data-related competencies over a ten-year period (2012-2022) with particular attention given to the skill sets that liaisons or non-data librarians may need to develop or hone. overall, the findings indicate a surprisingly stable list of skills over this period. this review finds that to support research data services on campus, librarians must rely on traditional skills including reference/consulting, teaching/training and collaboration/engagement as well as data-specific competencies, including metadata creation, data preparation for repositories, data preservation, data management plan (dmp) creation, and programming/data analysis. these competencies are essential for librarians to assist researchers with data queries. the study highlights the need for structured training and suggests which competencies to prioritize. the findings aim to guide the development of self-training resources and cross-training initiatives to better equip librarians in supporting data-rich research. keywords data, liaison librarians, skills, competencies, research data support, training introduction in 2012, tenopir, birch and allard published the acrl white paper ‘academic libraries and research data services: current practices and plans for the future.’ their proposition was that as science becomes more ‘collaborative, data intensive, and computational,’ (tenopir et al., 2012) researchers would have greater data management needs. the white paper cited the national science foundation’s ‘plan for open government’ (national science foundation, 2012) which required data management plans for all funded research projects. it stated, ‘it is critical that research data, regardless of the funding source, be properly managed, curated, standardized, citable, easily shared and made discoverable by others.’ the 2013 holdren memorandum on ‘increasing access to the results of federally funded scientific research’ required public access to research results and data (whitehouse office of science and technology policy (ostp), 2013). it directs federal agencies, such as the nih and nsf, to develop plans to support public access to the results of research, including data, funded by the federal government. given the changes in the way science was being conducted and analyzed, and the requirements for sharing of research results, college and university campuses were required to support data 2/23 brodsky, meryl & hannah chapman tripp (2025). data competencies for liaison librarians: a scoping review, iassist quarterly 49(3), pp. 1-23. doi: https://doi.org/10.29173/iq1154 management activities in ways not previously done. these changes combined with the acrl white paper, impelled academic libraries to help researchers fill these data management skill gaps. in 2011, librarians at purdue undertook a project to understand what data management skills researchers had, and what skills they needed via an imls grant to study ‘data information literacies’ (institute of museum and library services, 2011). purdue university libraries partnered with the university of minnesota libraries, the university of oregon libraries, and cornell university libraries to develop and implement data information literacy instruction for graduate students. an objective of this research was the development of instructional curricula as well as a community of trained librarians and disciplinary researchers who would share their skills with their institutions and local colleagues in a ‘train the trainer’ model (carlson, 2011). a product of this grant-funded research was the book ‘data information literacies: librarians, data, and the education of a new generation of researchers’ (carlson & johnston, 2015). this book guides librarians to teach data-related skills to researchers at their university. the focus is on researcher needs, but as a by-product of librarians leading these instructional sessions, librarians would also acquire and become proficient in these data-related skills. at about the same time, librarians in the uk formed responses to research data-related questions. rice et al. wrote an article on a pilot course on research data management (rdm) that data librarians led at the university of edinburgh in 2012-2013 for librarians (rice et al., 2013). this course was based on mantra, a course developed by the edina and data library, university of edinburgh, for early career researchers (university of edinburgh, n.d.). in 2011, in the netherlands, 3tu.datacentrum developed the course ‘data intelligence 4 librarians’ to provide online resources and training for digital preservation practitioners, specifically for library staff (de smaele et al., 2013). the site is now called ‘essentials 4 data support.’ it is an introductory course for those interested in supporting researchers in various data management activities including storing, managing, archiving, and sharing their research data (research data netherlands, n.d.). in 2016, rice and southall published ‘the data librarian’s handbook’ which outlines how librarians can work with researchers in a case study format (rice & southall, 2016). this book, together with ‘databrarianship: the academic data librarian in theory and practice’ edited by kellam and thompson, describes the breadth of data-related activities taking place in libraries, and how librarians and data professionals support research in an academic institution (kellam & thompson, 2016). the digital curation centre (dcc) was created in 2004 by a consortium comprising the universities of edinburgh and glasgow, ukoln at the university of bath, and stfc, which managed the rutherford appleton and daresbury laboratories. its mission was to solve challenges in digital curation that could not be tackled by any single institution or discipline. to that end, they developed models of research support (pryor, 2009), how-to-guides, and self-learning modules (digital curation centre, n.d.). in 2015, the data curation network (dcn) began in the u.s. as a grant-funded organization of institutional repositories and non-profit institutions whose vision was to advance open research by making data more ethical, reusable, and understandable. they also host primers and online modules (blake et al., 2022). the focus on data instruction and librarian learning led to the development of the association of college and research libraries research data management road show in the united states. the road show demonstrated how librarians might adapt their pre-existing information literacy skills to more specific research data management skills. for example, if a librarian is comfortable with reference interviewing, they could use that skill to have a consultation about the research process and data 3/23 brodsky, meryl & hannah chapman tripp (2025). data competencies for liaison librarians: a scoping review, iassist quarterly 49(3), pp. 1-23. doi: https://doi.org/10.29173/iq1154 needs (goben & sapp nelson, 2024). road show attendees were shown a chart with a scaffolded learning plan to develop rdm-specific knowledge (goben & sapp nelson, 2018). there has been a recent effort to create librarian positions focused solely on data, such as the data librarian or data specialist. some libraries even host entire departments that serve the data needs on campus. this practice has been concentrated at larger universities. researchers at smaller institutions still require help with data, but their needs are met by other librarian roles, such as liaisons or research and instruction librarians or scholarly communications librarians or other departments (tenopir et al., 2019). data and the sharing of data is a relatively new venture for librarians. there are courses on these topics in schools of information, but these classes are not required, so many aspiring librarians do not take them. many librarians who support research data are self-taught (thomas & urban, 2018). ‘working with research data’ does not usually top the list of skills or experiences in a subject or liaison librarian’s job description. increasingly though, researchers seek help in creating data management plans, identifying and preparing data to deposit in an appropriate repository, and creating metadata. academic librarians who may be unfamiliar with research data may find these requests overwhelming (fuhr, 2022). there are many places to find training, including self-paced online training, offered by mantra and the digital curation center. conferences with data-related sessions often offer workshops and instruction. however, it can be difficult to understand the breadth of the skills needed from the beginning to the end of a research project and beyond the project’s conclusion. in addition, the content of these discipline-agnostic introductory workshops may not meet the specific demands of scholars. objectives liaison librarians have helped researchers find and access data as part of traditional reference and consultation work for many years. tenopir et al. cite subject-based librarians as providing 61% of the ‘research data reference/consultation/instruction services to researchers’ (2015). this scoping review is intended as a comprehensive look at empirical literature to understand the data skills needed by liaison librarians in supporting data-rich research projects, from the basics of data literacy through data preservation. we wanted to look at what has happened in the past, how this has changed over time, and if there were points of consensus. we aimed to gather studies that reported data competencies for an audience of academic librarians collected through empirical methods. the primary research questions are: ● rq1 what does the literature indicate are the data-related competencies/skills that non-data liaison librarians (including reference, instruction, subject, and teaching librarians) should have to support researchers? ● rq2 what are the trends over this ten-year period in the literature related to data competencies? we use the word competencies, as defined by the national institutes of health, to mean ‘the knowledge, skills, abilities, and behaviors that contribute to individual and organizational performance’ (national institutes of health, 2024). by combining competencies with the word data, we mean using the knowledge, skills, abilities, and behaviors to support researchers in their work with data. 4/23 brodsky, meryl & hannah chapman tripp (2025). data competencies for liaison librarians: a scoping review, iassist quarterly 49(3), pp. 1-23. doi: https://doi.org/10.29173/iq1154 other studies that list data competencies data competency skill lists are cited in the library and information science literature with regards to librarians teaching researchers what they need to know. qin and d'ignazio discussed data-related skills as ‘science data literacy,’ differentiating it from information literacy and digital literacy. their efforts were around educating undergraduate and graduate stem students (calzada prado & marzal, 2013). carlson, et al. generated a list of ‘core competencies for data information literacy’ for students and faculty, including some of the same skills that qin and d’ignazio identified (2011). the carlson article suggested that librarians should ‘map the skill sets librarians currently have to the data information literacy objectives, either as stated here [in that article] or as they develop in practice.’ piorun linked learning objectives to the nsf’s data management plan requirements. piorun’s article described a training program geared towards medical, graduate, and undergraduate science students (piorun et al., 2012). the authors parsed learning objectives into modules that could be taught to researchers at different levels. prado and marzal described data literacy as having ties to information literacy and several other frameworks. in their article, they define and describe the competencies so that they could be separated into training modules (calzada prado & marzal, 2013). schneider combined the information literacy skills, and the digital preservation outreach & education (dpoe) curriculum from the library of congress for undergraduates, graduate students, lis (library and information science) students, data creators, data scientists, data librarians and data managers. the approach is useroriented and includes guidance on which competencies should be taught to each group (schneider, 2013). some of the competencies are not well defined, i.e., ‘extracting information from data models (and people).’ pothier and condon developed business data literacy competencies for business school students, in accordance with the needs of corporate institutions, having found the previously published data competencies more focused on the sciences (pothier & condon, 2020). risdale, et al., examined data competencies that were reported on in the research literature, and counted them to surface the best of the best (risdale et al., n.d.). they came up with 22 competencies in five different categories. however, the approach was not focused on librarian training or upskilling librarians. sapp nelson created a matrix of data management competencies that include learning goals (sapp nelson, 2017). she scaffolded it such that a researcher could start from an undergraduate level and learn to manage data through doctorate-level work and data stewardship. the skills build as one moves from personal information management through team data management, to research enterprise management. the competencies include observable activities that illustrate, through bloom’s taxonomy (cognitive, psychomotor and affective), whether the competency has been achieved. the thirty-six competencies identified in the matrix are similar to the ones we identified and used in our review. however, our purpose was to track what librarians need training on, so there are differences. we chose a scoping review methodology to assess these questions because we wanted to identify and track developments in the literature over the ten years following tenopir’s seminal white paper ‘academic libraries and research data services’ published by acrl in 2012 (tenopir et al., 2012). we wanted to identify if there were any trends related to the data competencies and to see if there was consensus in the skills librarians felt they needed to support data-related research. this white paper reported on a survey of arl libraries about the data-related services they were providing or planned to provide. tenopir et al., categorized the skills into ‘informational/consultative research data services (rds)’ and ‘technical assistance/hands-on research data services.’ this work provided us with an initial framework to examine the literature, but it needed to be enlarged and updated. we built on the skills/competencies through a careful reading of what we found in the 5/23 brodsky, meryl & hannah chapman tripp (2025). data competencies for liaison librarians: a scoping review, iassist quarterly 49(3), pp. 1-23. doi: https://doi.org/10.29173/iq1154 literature and devised four categories of competencies including data literacy, data management, data curation, and tools/technology/software. methods we used a scoping review methodology to identify articles. we developed the search strategy by brainstorming relevant keywords and term harvesting from known research in the library information science & technology abstracts (lista) database. the strategy was then tested against known relevant articles. several standardized vocabularies were evaluated, and we determined that adding specific vocabulary terms would not add relevant articles to the results. the search strategy used in lista (ebsco) with all fields being searched was: ( ("data literac*" or “data manag*” or “data curat*” or "data visualization" or “research data”) ) and ( (libraries or library or librarian or “information professional” or archivist) ) and ( (training or education or school or "continuing education" or course* or class* or skill* or competenc* or standard*) ) date limit: 2012-2022. no further limits were included in the strategy. the strategy can be further visualized in table 1. the search was run in the following databases: academic search complete, eric, information science & technology abstracts (ista), library & information science source, plus dissertations & theses global (proquest interface), and web of science: core collection. the first four database searches were conducted using the ebsco interface. the searches were all run on march 29, 2022. the choice of which databases to query was influenced by institutional availability, however the authors deemed the available databases to be of sufficient depth. the number of databases included achieved broad interdisciplinary coverage. the prisma diagram documenting the search numbers and exclusion decisions can be found in figure 1. supplemental searching included google scholar and handsearching the journal of escience librarianship. the authors chose to hand-search the journal of escience librarianship due to lack of indexing. only the first 200 google scholar results from a simplified search were examined. these results were included in the deduplication and screening and are represented as a database in the prisma chart. table 1. search terms developed and used in the search strategy (originally for lista) term categories data-related terms librarian-related terms competency-related terms term lists data literac* data manag* data curat* data visualization research data libraries library librarian information professional archivist training education school continuing education course* class* skill* competenc* standard* 6/23 brodsky, meryl & hannah chapman tripp (2025). data competencies for liaison librarians: a scoping review, iassist quarterly 49(3), pp. 1-23. doi: https://doi.org/10.29173/iq1154 figure 1. prisma flow diagram eligibility criteria & first round screening this review screened for articles that met the following inclusion criteria: 1. the article must present information concerning skill development, current knowledge or present an assessment/survey based on expanding librarian roles to include data. 2. the population must be librarians/information professionals working in a library. this group must work in a setting where they support the learning and/or research of faculty/students regardless of discipline. 3. the target setting must include higher education. 4. sources must report original research. these criteria were developed to suit our study objectives. the first criterion is the core of what we wanted to study. it’s broad enough to be inclusive of multiple study designs. the second is an important distinction because it centers the conversation on librarians. many studies were excluded because they discussed data competency assessment following a session aimed at student learning. the fourth criterion was designed to cast a broad net over types of research. we commonly see peer reviewed literature used as inclusion criteria; however, we decided instead to focus on original research. primary deduplication occurred using the deduplicator tool available through sr accelerator. a small number of additional duplicates were identified and removed during the subsequent title and abstract screening. all deduplication is accounted for in the prisma chart. title and abstract screening was performed using the web-based tool, rayyan. a pilot screening round of 100 randomly selected records was performed to ensure consistency in the application of predefined inclusion and exclusion criteria. table 2 outlines and defines our inclusion and exclusion criteria, including examples of when a study would be excluded. we utilized rayyan’s blind feature and resolved disagreements through 7/23 brodsky, meryl & hannah chapman tripp (2025). data competencies for liaison librarians: a scoping review, iassist quarterly 49(3), pp. 1-23. doi: https://doi.org/10.29173/iq1154 discussion and scope clarification. after we ran the pilot screening round, the blind feature was again used for the first round of screening. all results were independently assessed by both authors. we resolved discrepancies through discussion. articles that one or both of us deemed ‘maybe’ were included. the maybe category contained research that lacked an abstract or the title/abstract did not include enough context to exclude without reading the full text. ultimately, we excluded 2,007 articles, leaving 126 articles for the second-round screening. table 2. inclusion and exclusion criteria second round screening proceeding to the second round, we exported the 126 results from rayyan and added 4 articles found in hand-searching to begin our second round of screening. we followed the same discrepancy resolution as reported in the first round. criteria category definition code (rayyan label) inclusion the article must focus on skill development, current knowledge of librarians or present as assessment/survey based on expanding librarian roles to include data. right_focus the target population must be librarians/information professionals. this group must work in a setting where they support the learning and/or research of faculty/students regardless of discipline. right_audience the target setting must include higher education. right_setting sources must report on original research. right_publication type exclusion the article does not focus on skill development, knowledge of librarians or assessment/survey based on expanding librarian roles to include data. wrong_focus non-librarian learning focus. this might include presentation of a data literacy training program developed for students or for a department. wrong_audience non-academic library setting: primary & secondary school, special libraries, corporate libraries, public libraries. (exception: a health library that is serving a medical school.) wrong_setting review articles, overviews or opinion pieces that did not present original research wrong_publication_type 8/23 brodsky, meryl & hannah chapman tripp (2025). data competencies for liaison librarians: a scoping review, iassist quarterly 49(3), pp. 1-23. doi: https://doi.org/10.29173/iq1154 exclusion we decided not to translate articles that were not in english due to a lack of funds. this decision accounted for the exclusion of nine articles in portuguese, german, bosnian, chinese, and spanish. this brought us to 117 articles and research papers. from there, we excluded articles that we categorized as wrong focus (42). this included research that did not focus on data competencies or had an information school focus. we also eliminated those items of the wrong publication type (20). in general, these publications included review articles, overviews, or pieces that did not present original research. last, we excluded research for the wrong audience (7), which meant any audience not composed of academic, medical school librarians, or library staff. these decisions were established at the outset of the project to maintain a narrow focus with comparable competencies across articles. final studies once we added in 4 studies from hand searching we narrowed our results to 50 studies and coded these articles to understand which data competencies were identified as the ones that librarians should have. these studies are detailed in table 3. 9/23 brodsky, meryl & hannah chapman tripp (2025). data competencies for liaison librarians: a scoping review, iassist quarterly 49(3), pp. 1-23. doi: https://doi.org/10.29173/iq1154 author(s) year title journal study method purpose location tenopir et al. 2012 academic libraries and research data services: current practices and plans for the future. ala white papers & reports survey assess types of rdm support, planning, and staffing usa, can cox et al. 2012 upskilling liaison librarians for research data management. ariadne report connects rdm to traditional library roles and presents training uk reznikzellen et al. 2012 tiers of research data support services. journal of escience librarianship environ mental scan develops tiers of rdm support: education, consultation and infrastructure usa creamer et al. 2012 an assessment of needed competencies to promote the data curation and data management librarianship of health sciences and science and technology librarians in new england. journal of escience librarianship survey assess needed librarian competencies usa si et al. 2013 the cultivation of scientific data specialists: development of lis education oriented to e-science service requirements. library hi tech content analysis analyze data-related job ads and correlate ischool curricula multiple charbonn eau 2013 strategies for data management engagement. medical reference services quarterly report describes data-related opportunities for health sciences engagement usa corrall et al. 2013 bibliometrics and research data management services: emerging trends in library support for research. library trends survey investigate service developments including staff training needs, audiences and constraints aus, nz, irl, uk stewart & crossley 2013 library readiness for research data management. aliss quarterly report investigate what knowledge areas librarians need to support rdm uk antell et al. 2014 dealing with data: science librarians' participation in data management at association of research libraries institutions. college & research libraries survey assess awareness and involvement of science librarians in rdm-related activities usa xia & wang 2014 competencies and responsibilities of social science data librarians: an analysis of job descriptions. college & research libraries mixed methods assess position adds in order to clarify qualifications sought aus, can, ger, irl, ned, qa, sg, swe, uk, uae, usa tenopir et al. 2014 research data management services in academic research libraries and perceptions of librarians. library & information science research survey compare rdm activity frequency with library policy usa, can lee & stvilia 2014 data curation practices in institutional repositories: an exploratory study. proceedings of the association for information science & technology survey assess data curation practices in institutional repositories. usa cox et al. 2014 a spider, an octopus, or an animal just coming into existence? designing a curriculum for librarians to support research data management. journal of escience librarianship mixed methods introduces, explains and discusses assessment of rdm rose (training for librarians) uk brown et al. 2015 developing new skills for research support librarians. australian library journal case study offers perspective of academic librarians moving to a data librarian role aus davis & cross 2015 using a data management plan review service as a training ground for librarians. journal of librarianship & scholarly communication report evaluates a (dmp review) librarian training program based on predefined competencies usa johnson & bresnahan 2015 dataday! designing and assessing a research data workshop for subject librarians. journal of librarianship and scholarly communication report describes a liaison training day and post training assessment usa lockhart & leiß 2015 librarians' skills for e-research support–joint project at tu münchen and cput. proceedings of the iatul conferences. paper 2. case study librarians developed list of rdm services and determined required skills to offer services ger, sa 10/23 brodsky, meryl & hannah chapman tripp (2025). data competencies for liaison librarians: a scoping review, iassist quarterly 49(3), pp. 1-23. doi: https://doi.org/10.29173/iq1154 dér 2015 exploring the academic libraries' readiness for research data management: cases from hungary and estonia. master's thesis case study understand library staff member opinions on roles in rdm and institutional practice hu, ee cross et al. 2015 where do we go from here: choosing a framework for assessing research data services and training. charleston conference (2015) report evaluates existing frameworks for assessment of rdm usa tammaro et al. 2016 understanding roles and responsibilities of data curators: an international perspective. libellarium: journal for the research of writing, books, and cultural heritage institutions mixed methods chart key tasks and responsibilities of data curators it, ch, usa tenopir et al. 2017 research data services in european and north american libraries: current offerings and plans for the future. proceedings of the association for information science & technology survey establishes rdm services and goals by surveying directors, including staff skill development eu lee & stvilia. 2017 practices of research data curation in institutional repositories: a qualitative view from repository staff. plos one survey defines current ir practices through interviews at 13 research universities usa southall, & scutt 2017 training for research data management at the bodleian libraries: national contexts and local implementation for researchers and librarians. new review of academic librarianship report describes rdm training and development at a library uk johnston et al. 2017 results of the fall 2016 data curation pilot. university digital conservancy mixed methods develops a multi-institutional staffing model for data curation services usa kaushik 2017 perceptions of lis professionals about data curation. world digital libraries survey identifies views of lis professionals on data curation activities in federer 2018 defining data librarianship: a survey of competencies, skills, and training. journal of the medical library association survey describes data librarianship skills and tasks usa goben & sapp nelson 2018 the data engagement opportunities scaffold: development and implementation. journal of escience librarianship report presents a data engagement opportunities scaffold that merge lis skills with rdm and provide measurable outcomes usa read et al. 2019 a two-tiered curriculum to improve data management practices for researchers. plos one mixed methods librarians took online modules with embedded questions to track changes in understanding usa eclevia et al. 2019 what makes a data librarian: an analysis of job descriptions and specifications for data librarian. qualitative & quantitative methods in libraries job ad analysis presents analysis of data librarian job postings usa, can, uk, sg ohaji et al. 2019 the role of a data librarian in academic and research libraries. information research intervie w presents a “blueprint” mapping the development of the data librarian role based on interviews nz tammaro et al. 2019 data curator's roles and responsibilities: an international perspective. libri: international journal of libraries & information services mixed methods defines data curation through a multi-year investigation, including interviews and job posting analyses multiple li et al. 2019 research data management: what can librarians really help? the grey journal (tgj) case study describes developing rdm services at one institution usa cox et al. 2019 maturing research data services and the transformation of academic libraries. journal of documentation survey identifies changes from 2019 to 2014 questionnaires to analyze changes rdm in libraries aus, can, ger, irl, ned, nz, uk, usa tang & hu 2019 providing research data management (rdm) services in libraries: preparedness, roles, challenges, and training for rdm practice. data & information management survey presents results of a survey including questions on roles, readiness for rdm, and challenges aus, et, fi, fr, hu, jpn, kg, nz, nor, sa, rs, sg, sa, 11/23 brodsky, meryl & hannah chapman tripp (2025). data competencies for liaison librarians: a scoping review, iassist quarterly 49(3), pp. 1-23. doi: https://doi.org/10.29173/iq1154 es, ch, th, tt, tr, ug, zm, zw, can, uk, ned, in, jm, ger, uae, usa federer & qin 2019 beyond the data management plan: expanding roles for librarians in data science and open science. proceedings of the association for information science & technology mixed methods understand future development needs in rdm and open science usa tenopir et al. 2019 academic librarians and research data services: attitudes and practices. itlib: informacne technologie a kniznice survey presents results of a survey assessing perceived importance, confidence, and contribution of librarians to rds usa, can rice 2019 supporting research data management and open science in academic libraries: a data librarian's view. mitteilungen der vereinigung österreichischer bibliothekarinnen und bibliothekare case study discusses one institution's path to developing rdm sct joo et al. 2019 librarians' perceptions on skills/knowledge and resources needed for research data services: preliminary results. jcdl '19: proceedings of the 18th joint conference on digital libraries survey the survey assesses librarian opinions surrounding the importance of skill sets as they relate to rds usa chawinga & zinn 2020 research data management at an african medical university: implications for academic librarianship. journal of academic librarianship mixed methods assesses the rdm landscape and reveals an opportunity for libraries to step in sa ahmad et al. 2020 librarian's perspective for the implementation of big data analytics in libraries on the basis of lean-startup model. digital library perspectives survey questionnaire responses develop a path to big data analytics services in libraries pk chiware 2020 data librarianship in south african academic and research libraries: a survey. library management mixed methods define current competencies of practicing rds librarians in south africa sa federer et al. 2020 developing the librarian workforce for data science and open science. libellarium: journal for the research of writing, books, and cultural heritage institutions report report of a workshop to define skills needed for data and open science librarians usa ducas et al. 2020 reinventing ourselves: new and emerging roles of academic librarians in canadian research-intensive universities. college & research libraries survey discusses new roles in librarianship and the related skills required and confidence level can bishop et al. 2021 job analyses of earth science data librarians and data managers. bulletin of the american meteorological society intervie w present the skills of practicing librarians and data managers usa masinde et al. 2021 research librarians’ experiences of research data management activities at an academic library in a developing country. data and information management intervie w presents the skills related to rdm activities with aims to support growing curation needs ke ashiq et al. 2021 the perception of library and information science (lis) professionals about research data management services in university libraries of pakistan. libri: international journal of libraries & information services survey addresses lis professional opinions regarding rdm training needs and reasons for rdm support pk joo & schmidt 2021 research data services from the perspective of academic librarians. digital library perspectives survey presents survey results investigating librarian perceptions of rds in libraries usa 12/23 brodsky, meryl & hannah chapman tripp (2025). data competencies for liaison librarians: a scoping review, iassist quarterly 49(3), pp. 1-23. doi: https://doi.org/10.29173/iq1154 table 3. final studies coding for competencies during our second round of screening, we tested a preliminary collection of codes. shared coding was completed in excel with each competency represented by a column and each article represented by a row. the excel sheet produced a series of binary codes indicating whether the work included the identified competency. both authors examined each of the articles for inclusion/exclusion criteria and a trial round of coding was completed. the authors discussed discrepancies until an agreement was reached. the authors developed a formalized codebook and definitions based on the competencybased content identified in the articles. each author re-coded half the articles based on the updated competency-based coding sheet. the codebook is available via the open science framework repository here: https://osf.io/qa5vr/. the competency list is in table 4. during our work, we identified two larger themes among our categories: previously held liaison librarian competency areas and competency areas that represent data-specific skills. themes liaison librarian competencies data-specific competencies categories consultative/ informational data literacy data management data curation tools/tech/ software competencies reference & consulting & communication teaching & training collaboration & library engagement & outreach understand research methodologies disciplinary background find data access data cite data use data ethically data acquisition or deaccession research data lifecycle author & document identifiers bibliometrics promote open science data project planning dmp creation data documentation funder mandate familiarity copyright & license confidentiality repository selection repository use determine data quality data security data preservation data storage metadata prepare data for sharing data governance ir creation & management support multiple data types data cleaning data visualization programming data analysis statistical analysis tech infrastructure it competency table 4. themes, categories, and competencies used in coding the included articles kvale 2021 using personas to visualize the need for data stewardship. college & research libraries mixed methods explores possibilities of research data stewards to identify skills required of such a role nor borkakoti & singh 2021 research data management in central universities and institutes of national importance: a perspective from northeast india. library philosophy & practice survey discover opinions of current lis professionals regarding rdm in fuhr 2022 developing data services skills in academic libraries. college & research libraries survey assess current competencies and preferred mechanisms for upskilling can, usa, uk, aus 13/23 brodsky, meryl & hannah chapman tripp (2025). data competencies for liaison librarians: a scoping review, iassist quarterly 49(3), pp. 1-23. doi: https://doi.org/10.29173/iq1154 table 4 shows the classification of each of the competencies. ten of the skills fall into the data curation category. data curation is very important, as open data becomes mandatory as a condition of federal funding. it’s also an area that is well documented, due to organizations such as the data curation network and the digital curation centre who provide guidance and training. the next area of expertise falls into the data literacy category, which addresses nine more straightforward skills, such as finding data and data citation. data management contains nine skills, which are important as researchers begin working with data. they need to understand how to document their data processes and choose a repository for long-term storage in accordance with funder specifications. the tools/tech/software category contains seven skills. often expert researchers have familiarity with the tools they need to do their work, but obtaining access to those tools, or finding where to get help, can facilitate their work. results highlights of the results are the top ten data competencies, top categories, and categories by year. the top ten data competency areas and the affiliated categories respond to rq1. in response to rq2, the results revealed a stable list of competencies over the ten-year analysis period. top data competencies our review sought to identify which exact skills and competencies librarians need to acquire to work with researchers seeking help with data-related questions. figure 2 shows the top ten data competencies. figure 2. top ten data competencies the top skill or competency that librarians need to master is metadata or being able to help researchers understand what metadata is and how to create findable (useful metadata) that is disciplineor repository-appropriate. teaching researchers how to prepare data for use in a repository or for sharing is the second most cited competency. this is defined as organizing data for reuse, utilizing the fair principles, creating readme files, and ensuring files are in a software-agnostic file 14/23 brodsky, meryl & hannah chapman tripp (2025). data competencies for liaison librarians: a scoping review, iassist quarterly 49(3), pp. 1-23. doi: https://doi.org/10.29173/iq1154 format. data management plan (dmp) creation is next. in this competency, librarians must understand the parts of a data management plan and what information goes in each section. in addition, understanding the discipline-specific directorate-level requirements, can help with dmp creation. we also included the use of dmp generation tools, such as dmptool. data preservation is defined as selecting data for preservation and considering the ongoing roles and responsibilities for the data, including long-term storage. programming/data analysis was the fifth most cited competency. this includes purchasing and licensing tools for qualitative and quantitative analysis. it also encompasses understanding text mining, web apis, data modeling, and data science. teaching r and python via the software carpentries in the library is an example (pugachev, 2019). top competency categories while our focus was on skills and competencies, we also considered the categories of competencies. many of the articles focused on a specific category for training purposes. for instance, data curation was the focus of seven out of our 50 selected studies. table 5 shows the top data competencies with their categories. competency category metadata data curation prepare data for use in repository/data sharing data curation data preservation data curation dmp creation data management programming/data analysis tools/tech/software tech infrastructure/it competency tools/tech/software support multiple data types data curation data governance data curation repository use data curation data publication data literacy cite data data literacy table 5. top competency categories data competency categories by year we looked for trends in the competencies, but because there was so much variability in a competency from one year to the next, we couldn’t draw any conclusions. instead, we looked at categories to detect trends. first, we collected how many competencies/skills were noted in each year. most articles discussed more than just one category; hence the counts add up to more than 100%. we then normalized the impact of the number of studies per year by dividing the counts within each category by the number of studies in the year. what we show in figure 3 are the percent of skills addressed in each category. the lines show the trend over time. in the articles published in 2012, 63% of the categories in the data curation competency were discussed, 44% of the categories in the data 15/23 brodsky, meryl & hannah chapman tripp (2025). data competencies for liaison librarians: a scoping review, iassist quarterly 49(3), pp. 1-23. doi: https://doi.org/10.29173/iq1154 management competency were studied, 36% of the categories in the data literacy competency were mentioned, and 18% of the categories in the tools and technology competency were documented. figure 3. data competency categories by year the tools/tech/software category grew from 18% to over 60% in 2021, an overall growth rate of roughly 4% a year. data curation was relatively stable, in the 60-70% range. data management showed moderate growth, going from 44% in 2012 to 59% in 2021, growing about 3% per year. data literacy also remained relatively stable, rising from 36% in 2012 to 39% in 2021. discussion when we began the process of searching and collecting articles from the library literature related to data competencies, most of the articles had to do with researcher needs. the literature was filled with surveys of what researchers felt they needed to learn to competently work with data. some of the articles discussed how librarians addressed these needs. fewer articles focused on librarian needs. while the researcher perspective and the librarian perspective are sometimes linked, it was only in a handful of articles that authors suggested that librarians who taught researchers could or should also learn these skills. we started with the tenopir article and used other studies that mentioned data competencies to develop our competency list. after reviewing the literature, we created the following competency categories: data literacy, data management, data curation and tools/tech/software section. the skills in each category are unique and are often taught as separate units. the studies we identified are largely surveys of academic libraries or librarians. the volume of research in this area has been relatively consistent over time and includes articles from institutions across the globe. based on the literature, librarians all over the world do not feel confident in their ability to handle data-related queries. we were surprised to find how little these skills had changed over the ten-year period. there was a consensus in the top ten skills that librarians need to know for working with data. while the need for tools/tech/software category of competencies grew over time, more basic data handling skills are still required. the articles described research conducted on data competencies in 45 different countries. data is an issue of concern for librarians across the globe. the top six countries represented in our collection are the united states, canada, the uk, australia, germany, and south africa. research published from the 16/23 brodsky, meryl & hannah chapman tripp (2025). data competencies for liaison librarians: a scoping review, iassist quarterly 49(3), pp. 1-23. doi: https://doi.org/10.29173/iq1154 global south, such as india and south africa, has been published in more recent years. this suggests that data-rich research and thus data competencies for librarians are a growing need internationally. since most of the studies we found are surveys, this data comes right from the librarians themselves. the strength of this type of research is that we heard directly from the librarians about the specific data skills they are concerned about. a weakness of this type of study is that it is typically administered via online form and can “lead” or bias participants to respond based on the specific questions asked. for example, librarians might be apt to select skills they’d like training on simply because it’s on the survey. a weakness, and a gap in the literature, is that we don’t know why liaison librarians don’t have or feel confident about working with data. many information schools offer data-related courses. there are many online learning opportunities via self-paced modules from data-related organizations. professional conferences regularly offer workshops for librarians at all levels of data expertise. perhaps it has to do with the lack of time available for training, or the broad range or technical nature of data-related queries? or it may have to do with the hiring of data specialists? if this is the case, research which examines how data specialists work with librarians and their division of responsibilities could prove illuminating. one of the things that has shifted over time is federal funding requirements. though not evident in our review (due to the time period we examined), the national institutes of health (nih) has implemented stricter requirements for data management and sharing. in 2023, the nih instituted a new data management and sharing (dms) policy (national institutes of health, 2023). the policy emphasizes the completion of a data management plan that outlines how data and metadata will be managed and shared. it also encourages researchers to release the data as soon as an associated publication is ready, or at the completion of the grant, whichever comes first, without an embargo period. this indicates that managing and sharing data is becoming more important. another change the nih is encouraging is sharing software and code as it relates to funded research. as part of supporting the dissemination of ‘research products,’ this initiative is intended to allow for reproducibility, and contribute to the advancement of science (national institutes of health, 2023). thus, it’s likely that librarians will also need to develop software curation skills. limitations a scoping review presents some inherent limitations as a methodology. among them, are the sources used to search, and the algorithms/terms lists may not be entirely comprehensive despite the best efforts of the authors. the review type also does not prescribe an assessment of quality, so the included studies may range in their own methodological rigor (grant & booth, 2009). other limitations include the inability to translate studies into english. reviewing only studies published in english language journals limits the ability to have a global view of the competency landscape, which initially was one of our objectives. while we regret the inability to translate articles, we didn’t want our results to be limited to north america. here, we recognize the western bias in the publishing landscape and acknowledge that this bias is carried through to this study. another limitation was that our coding method generated a sheet of yes/no responses. using this method defined whether or not a competency was present in a research study. it did not drill down to the individual responses in a survey or quantify responses beyond the presence or absence of a competency. thus, this work does not measure changing data needs but tracks changes in the presence or absence of these competency areas in the literature. 17/23 brodsky, meryl & hannah chapman tripp (2025). data competencies for liaison librarians: a scoping review, iassist quarterly 49(3), pp. 1-23. doi: https://doi.org/10.29173/iq1154 initially, we planned to include liaison/subject librarians instead of all librarians and data-specific roles. however, this proved difficult as many of the surveys were not specifically sorted by the type of library role held by the respondent. thus, we jettisoned this exclusion criterion. though we speak broadly about developing data competencies, one of our objectives is to track what subject librarians/liaisons may need to support data. thus, we have listed this as a limitation of the study. many articles discussed the category level (i.e., data literacy) of need, whereas other articles spelled out the skills (i.e., data citation). we were most interested in research that captured the skills librarians need to know. one of the more significant contributions of this review may be the enumeration of each skill/competency. spelling out these skills and defining them makes them more transparent and can help librarians understand what type of training to take or offer. future directions research discovering how ingrained liaisons librarians are, specifically in data-rich research projects with departments, would benefit the broader data landscape. this work could help to better define a suspected liaison librarian task and potentially drive future hiring, particularly as data-rich research becomes the norm and open science practices are more fully embraced by researchers across higher education. additionally, an examination of the types of positions that currently list data-related tasks would help to identify who is being asked to work on data-related needs and address the question of how siloed or cohesive the library services and data-related areas are becoming. more research is needed to identify what kind of data learning is offered in the information school curricula as we continue to see non-mlis holders hired as data specialists. for librarians seeking to gain training, learning opportunities might be more useful if they specify what competencies will be gained or bolstered through the completion of the training and what prerequisites are recommended to grasp concepts. call to action: data literacy courses should be required in information schools, especially for academic librarian tracks. academic librarians need not be experts, but they should be familiar with data-rich research methods. conclusion many of the competencies identified in this retrospective scoping review are skills that liaison librarians are already well-versed in, such as finding data, metadata, and copyright. other competencies might require structured or self-paced learning such as data management plan specifics, repository selection, and some of the more technically focused competencies, such as programming and data visualization. we suggest that librarians start by learning the skills near the top of the list to master skills in frequently identified competency areas. for example, a librarian could focus on training in the areas of metadata, prepare data for use in repository/data sharing, data preservation, dmp creation, and programming/data analysis. alternatively, they may wish to fully invest in learning a category of competencies, and start with data curation, so they’ll be able to help researchers save their data which would enable reusability, potentially increasing researcher citation, and fulfill funder mandates. this category method of learning could prove particularly useful to libraries seeking to establish or broaden their data skills across multiple staff members. for example, two staff members could seek training in data curation, while others may seek training in data literacy. it’s clear upon completing this review that no one librarian is likely to be the expert in every competency. rather, a library could benefit from a staffing model in which librarians choose to upskill in specific competencies or categories, and other librarians or library staff take on other competencies. this shared expertise model could enhance the research support ecosystem by strengthening relationships among librarians and library staff in different units, particularly at larger institutions. 18/23 brodsky, meryl & hannah chapman tripp (2025). data competencies for liaison librarians: a scoping review, iassist quarterly 49(3), pp. 1-23. doi: https://doi.org/10.29173/iq1154 when institutions hire data librarians, they should be integrated into the collaborative ecosystem of research support that includes liaisons, public services experts, and technical services. this ecosystem would ideally operate seamlessly in sharing information and knowledge of each area of expertise. liaisons and other librarians have cultivated relationships with faculty and research groups on campus. building trust and a support structure between librarians and data librarian/specialist roles is essential to creating an ecosystem where data-specific roles are relied upon for their expertise. likewise, librarians who have a deep subject knowledge may be better equipped to work with a researcher on a project (federer, 2018) and may themselves be eager to develop data-related skills. as data-rich research continues to gain traction and be supported by libraries, the research support ecosystem is becoming more complex. as research services continue to develop, researchers and libraries alike will be better served by an integrated service model that allows librarians to explore and develop new competencies such as those described in this review. in conclusion, data is but another form of information. librarianship is the profession that helps users to find, evaluate and ethically use information. yet, we are grappling with adding new tasks to overfilled plates and trying to determine where data-intensive skills fit in our institutions. how these data-related demands will be met in the future may swing back toward requiring a credentialed librarian as more librarians seek to self-train and the credentialing process for academic librarians moves toward data competencies. funding the authors report that no funding was received to support this research. associated documents associated documents including protocol planning, definitions, and data charting can be consulted at the open science framework: https://osf.io/qa5vr/ references ahmad, k., jianming, z., & rafi, m. (2020). librarian’s perspective for the implementation of big data analytics in libraries on the bases of lean-startup model. digital library perspectives, 36(1), 21–37. https://doi.org/10.1108/dlp-04-2019-0016 american library association. association of college & research libraries. (2025). building your research data management toolkit: integrating rdm into your liaison work. roadshows. https://www.ala.org/acrl/conferences/roadshows/rdm antell, k., foote, j. b., turner, j., & shults, b. (2014). dealing with data: science librarians’ participation in data management at association of research libraries institutions. college & research libraries, 75(4), 557–574. https://doi.org/10.5860/crl.75.4.557 ashiq, m., saleem, q. u. a., & asim, m. (2021). the perception of library and information science (lis) professionals about research data management services in university libraries of pakistan. libri, 71(3), 239–249. https://doi.org/10.1515/libri-2020-0098 association for information science and technology. (n.d.). about asis&t. https://www.asist.org/about/ bishop, b. w., orehek, a. m., & collier, h. r. (2021). job analyses of earth science data librarians and data managers. bulletin of the american meteorological society, 102(7), e1384–e1393. https://doi.org/10.1175/bams-d-20-0163.1 blake, m., borda, s., carlson, j., darragh, j., fearon, d., hadley, h., herndon, j., johnston, l., kalt, m., kozlowski, w., hess, s. l., moore, j., narlock, m., scott, d., vitale, c. h., wham, b. e., & 19/23 brodsky, meryl & hannah chapman tripp (2025). data competencies for liaison librarians: a scoping review, iassist quarterly 49(3), pp. 1-23. doi: https://doi.org/10.29173/iq1154 wright, s. (2022). curated training. data curation network. https://datacurationnetwork.github.io/curated/ borkakoti, r., & singh, s. (2021). research data management in central universities and institutes of national importance: a perspective from north east india. library philosophy and practice (ejournal). https://digitalcommons.unl.edu/libphilprac/5848 brown, r. a., wolski, m., & richardson, j. (2015). developing new skills for research support librarians. the australian library journal, 64(3), 224–234. https://doi.org/10.1080/00049670.2015.1041215 calzada prado, j., & marzal, m. á. (2013). incorporating data literacy into information literacy programs: core competencies and contents. libri, 63(2). https://doi.org/10.1515/libri-20130010 carlson, j. (2011). data information literacy. data information literacy. https://www.datainfolit.org/ carlson, j., fossmire, m., miller, c., & sapp nelson, m. (2011). determining data information literacy needs: a study of students and research faculty (no. 23; libraries faculty and staff scholarship and research, pp. 1–30). purdue university. https://docs.lib.purdue.edu/lib_fsdocs/23/ carlson, j., & johnston, l. r. (eds.). (2015). data information literacy: librarians, data, and the education of a new generation of researchers. purdue university press. charbonneau, d. h. (2013). strategies for data management engagement. medical reference services quarterly, 32(3), 365–374. https://doi.org/10.1080/02763869.2013.807089 chawinga, w. d., & zinn, s. (2020). research data management at an african medical university: implications for academic librarianship. the journal of academic librarianship, 46(4), 102161. https://doi.org/10.1016/j.acalib.2020.102161 chiware, e. r. t. (2020). data librarianship in south african academic and research libraries: a survey. library management, 41(6/7), 401–416. https://doi.org/10.1108/lm-03-2020-0045 corrall, s., kennan, m. a., & afzal, w. (2013). bibliometrics and research data management services: emerging trends in library support for research. library trends, 61(3), 636–674. https://doi.org/10.1353/lib.2013.0005 cox, a. m., kennan, m. a., lyon, l., pinfield, s., & sbaffi, l. (2019). maturing research data services and the transformation of academic libraries. journal of documentation, 75(6), 1432–1462. https://doi.org/10.1108/jd-12-2018-0211 cox, a., verbaan, e., & sen, b. (2012). upskilling liaison librarians for research data management. ariadne, 70. http://www.ariadne.ac.uk/issue/70/cox-et-al/ cox, a., verbaan, e., & sen, b. (2014). a spider, an octopus, or an animal just coming into existence? designing a curriculum for librarians to support research data management. journal of escience librarianship, 3(1). https://doi.org/10.7191/jeslib.2014.1055 creamer, a., morales, m., crespo, j., kafel, d., & martin, e. (2012). an assessment of needed competencies to promote the data curation and management librarianship of health sciences and science and technology librarians in new england. journal of escience librarianship, 18–26. https://doi.org/10.7191/jeslib.2012.1006 cross, w. m., & davis, h. m. (2016). where do we go from here: choosing a framework for assessing research data services and training. where do we go from here? charleston conference proceedings 2015. charleston library conference. https://doi.org/10.5703/1288284316312 data curation network. (n.d.). https://datacurationnetwork.org/ davis, h. m., & cross, w. m. (2015). using a data management plan review service as a training ground for librarians. journal of librarianship and scholarly communication, 3(2), 1243. https://doi.org/10.7710/2162-3309.1243 de smaele, m., verbakel, e., potters, n., & noordegraaf, m. (2013). data intelligence training for library staff. international journal of digital curation, 8(1), 218–228. https://doi.org/10.2218/ijdc.v8i1.255 20/23 brodsky, meryl & hannah chapman tripp (2025). data competencies for liaison librarians: a scoping review, iassist quarterly 49(3), pp. 1-23. doi: https://doi.org/10.29173/iq1154 dér, á. (2015). exploring the academic libraries’ readiness for research data management: cases from hungary and estonia [master thesis, oslo and akershus university college of applied sciences]. https://oda.oslomet.no/oda-xmlui/handle/10642/3367 digital curation centre. (n.d.). https://www.dcc.ac.uk/ digital curation centre. (2004, 2024). how-to guides. dcc university of edinburgh. https://www.dcc.ac.uk/guidance/how-guides ducas, a., michaud-oystryk, n., & speare, m. (2020). reinventing ourselves: new and emerging roles of academic librarians in canadian research-intensive universities. college & research libraries, 81(1), 43–65. https://doi.org/10.5860/crl.81.1.43 eclevia, m. r., maestro, r. s., & jr, c. l. e. (2020). what makes a data librarian?: an analysis of job descriptions and specifications for data librarian. qualitative & quantitative methods in libraries, 8(3), 19. federer, l. (2018). defining data librarianship: a survey of competencies, skills, and training. journal of the medical library association, 106(3). https://doi.org/10.5195/jmla.2018.306 federer, l., clarke, s. c., & zaringhalam, m. (2020). developing the librarian workforce for data science and open science [preprint]. open science framework. https://osf.io/uycax federer, l. m., & qin, j. (2019). beyond the data management plan: expanding roles for librarians in data science and open science. proceedings of the association for information science and technology, 56(1), 529–531. https://doi.org/10.1002/pra2.82 fuhr, j. (2022). developing data services skills in academic libraries. college & research libraries, 83(3). https://doi.org/10.5860/crl.83.3.474 gamble, a. (2018, september 28). research data management librarian academy (rdmla) training. courses at ischools. https://github.com/rdmla/rdmla.github.io/blob/master/survey-documents/training.pdf goben, a., & raszewski, r. (2015). research data management self-education for librarians: a webliography. issues in science and technology librarianship, 82. https://doi.org/10.29173/istl1666 goben, a., & sapp nelson, m. (2018). the data engagement opportunities scaffold: development and implementation. journal of escience librarianship, 7(2), e1128. https://doi.org/10.7191/jeslib.2018.1128 goben, a., & sapp nelson, m. (2024, august 27). scholarly communication toolkit: acrl workshop: research data management. acrl libguides. https://acrl.libguides.com/scholcomm/toolkit/rdmworkshop grant, m. j., & booth, a. (2009). a typology of reviews: an analysis of 14 review types and associated methodologies. health information & libraries journal, 26(2), 91–108. https://doi.org/10.1111/j.1471-1842.2009.00848.x institute of museum and library services. (2011). national leadership grants—libraries award. imls awards. https://www.imls.gov/grants/awarded/lg-07-11-0232-11 johnson, a. m., & bresnahan, m. m. (2015). dataday!: designing and assessing a research data workshop for subject librarians. journal of librarianship and scholarly communication, 3(2), 1229. https://doi.org/10.7710/2162-3309.1229 johnston, l. r., carlson, j., hudson-vitale, c., imker, h., kozlowski, w., olendorf, r., & stewart, c. (2017). results of the fall 2016 data curation pilot [report]. http://conservancy.umn.edu/handle/11299/188640 joo, s., & peters, c. (2019). librarians’ perceptions on skills/knowledge and resources needed for research data services: preliminary results. 2019 acm/ieee joint conference on digital libraries (jcdl), 382–383. https://doi.org/10.1109/jcdl.2019.00081 joo, s., & schmidt, g. m. (2021). research data services from the perspective of academic librarians. digital library perspectives, 37(3), 242–256. https://doi.org/10.1108/dlp-10-2020-0106 kaushik, a. (2017). perceptions of lis professionals about data curation. world digital libraries – an international journal, 10(2). https://doi.org/10.18329/09757597/2017/10207 21/23 brodsky, meryl & hannah chapman tripp (2025). data competencies for liaison librarians: a scoping review, iassist quarterly 49(3), pp. 1-23. doi: https://doi.org/10.29173/iq1154 kellam, l. m., & thompson, k. (2016). databrarianship: the academic data librarian in theory and practice. association of college and research libraries, a division of the american library association. kvale, l. h. (2021). using personas to visualize the need for data stewardship. college & research libraries, 82(3). https://doi.org/10.5860/crl.82.3.332 lee, d. j., & stvilia, b. (2014). data curation practices in institutional repositories: an exploratory study. proceedings of the american society for information science and technology, 51(1), 1– 4. https://doi.org/10.1002/meet.2014.14505101085 lee, d. j., & stvilia, b. (2017). practices of research data curation in institutional repositories: a qualitative view from repository staff. plos one, 12(3), e0173987. https://doi.org/10.1371/journal.pone.0173987 li, y., dressel, w., & hersey, d. (2019). research data management: what can librarians really help? grey journal (tgj), 15(1), 9. lockhart, j., & leiß, c. (2015, july 7). librarians’ skills for e-research support – joint project at tu münchen and cput. proceedings of the iatul conferences. https://docs.lib.purdue.edu/iatul/2015/lil/2 masinde, j., chen, j., wambiri, d., & mumo, a. (2021). research librarians’ experiences of research data management activities at an academic library in a developing country. data and information management, 5(4), 412–424. https://doi.org/10.2478/dim-2021-0002 national institutes of health. (2024). what are competencies? office of human resources frequently asked questions. https://hr.nih.gov/about/faq/working-nih/competencies/what-arecompetencies national institutes of health. (2023). final nih policy for data management and sharing. frequently asked questions. https://grants.nih.gov/grants/guide/notice-files/not-od-21-013.html national institutes of health. (2023). best practices for sharing research software. national science foundation. (2012). open government plan 2.0. https://www.nsf.gov/pubs/2012/nsf12066/nsf12066.pdf ohaji, i. k., chawner, b., & yoong, p. (2019). the role of a data librarian in academic and research libraries. information research, 24(4), paper 884. https://informationr.net/ir/244/paper844.html piorun, m., kafel, d., leger-hornby, t., najafi, s., martin, e., colombo, p., & lapelle, n. (2012). teaching research data management: an undergraduate/graduate curriculum. journal of escience librarianship, 46–50. https://doi.org/10.7191/jeslib.2012.1003 pothier, w. g., & condon, p. b. (2020). towards data literacy competencies: business students, workforce needs, and the role of the librarian. journal of business & finance librarianship, 25(3–4), 123–146. https://doi.org/10.1080/08963568.2019.1680189 pryor, g. (2009, may). core skills for data management. rdmf3. dcc research data management forum 3: roles and responsibilities for effective data management, manchester, england. https://www.dcc.ac.uk/sites/default/files/documents/rdmf/rdmf3/08%20pryor.pdf pugachev, s. (2019). what are “the carpentries” and what are they doing in the library? portal: libraries and the academy, 19(2), 209–214. https://doi.org/10.1353/pla.2019.0011 read, k. b., larson, c., gillespie, c., oh, s. y., & surkis, a. (2019). a two-tiered curriculum to improve data management practices for researchers. plos one, 14(5), e0215509. https://doi.org/10.1371/journal.pone.0215509 research data netherlands. (n.d.). essentials 4 data support (english)—public. dans moodle. https://danstraining.moodlecloud.com/ reznik-zellen, r., adamick, j., & mcginty, s. (2012). tiers of research data support services. journal of escience librarianship, 27–35. https://doi.org/10.7191/jeslib.2012.1002 22/23 brodsky, meryl & hannah chapman tripp (2025). data competencies for liaison librarians: a scoping review, iassist quarterly 49(3), pp. 1-23. doi: https://doi.org/10.29173/iq1154 rice, r. (2019). supporting research data management and open science in academic libraries: a data librarian’s view. mitteilungen der vereinigung österreichischer bibliothekarinnen und bibliothekare, 72(2), 263–273. https://doi.org/10.31263/voebm.v72i2.3303 rice, r. c., & southall, j. (2016). the data librarian’s handbook. facet publishing. risdale, c., rothwell, j., smit, m., ali-hassan, h., bliemel, m., irvine, d., kelley, d., matwin, s., & wuetherick, b. (n.d.). strategies and best practices for data literacy education [knowledge synthesis report]. dalhousie university. https://dalspaceb.library.dal.ca/server/api/core/bitstreams/7c0db868-79ac-4c3c-9ad0f2a93cf2ccc6/content sapp nelson, m. (2017). a pilot competency matrix for data management skills: a step toward the development of systematic data information literacy programs. journal of escience librarianship, 6(1), e1096. https://doi.org/10.7191/jeslib.2017.1096 schneider, r. (2013). research data literacy. in s. kurbanoğlu, e. grassian, d. mizrachi, r. catts, & s. špiranec (eds.), worldwide commonalities and challenges in information literacy research and practice (vol. 397, pp. 134–140). springer international publishing. https://doi.org/10.1007/978-3-319-03919-0_16 si, l., zhuang, x., xing, w., & guo, w. (2013). the cultivation of scientific data specialists: development of lis education oriented to e-science service requirements. library hi tech, 31(4), 700–724. https://doi.org/10.1108/lht-06-2013-0070 southall, j., & scutt, c. (2017). training for research data management at the bodleian libraries: national contexts and local implementation for researchers and librarians. new review of academic librarianship, 23(2–3), 303–322. https://doi.org/10.1080/13614533.2017.1318766 stewart, j., & crossley, j. (2013). library readiness for research data management. aliss quarterly, 8(4). https://uwe-repository.worktribe.com/output/940457/library-readiness-for-researchdata-management tammaro, a. m., matusiak, k. k., sposito, f. a., & casarosa, v. (2019). data curator’s roles and responsibilities: an international perspective. libri, 69(2), 89–104. https://doi.org/10.1515/libri-2018-0090 tammaro, a. m., matusiak, k. k., sposito, f. a., pervan, a., & casarosa, v. (2017). understanding roles and responsibilities of data curators: an international perspective. libellarium: časopis za istraživanja u području informacijskih i srodnih znanosti, 9(2), 39–47. https://doi.org/10.15291/libellarium.v9i2.286 tang, r., & hu, z. (2019). providing research data management (rdm) services in libraries: preparedness, roles, challenges, and training for rdm practice. data and information management, 3(2), 84–101. https://doi.org/10.2478/dim-2019-0009 tenopir, c., allard, s., baird, l., sandusky, r., lundeen, a., hughes, d., & pollock, d. (2019). academic librarians and research data services: attitudes and practices. information technology and libraries journal, 1, 37. https://trace.tennessee.edu/utk_infosciepubs/99 tenopir, c., allard, s., university of tennessee, knoxville, birch, b., baird, l., sandusky, r., langseth, m., hughes, d., & lundeen, a. (2015). research data services in academic libraries: data intensive roles for the future? journal of escience librarianship, 4(2), e1085. https://doi.org/10.7191/jeslib.2015.1085 tenopir, c., birch, b., & allard, s. (2012). academic libraries and research data services: current practices and plans for the future (p. 56) [an acrl white paper]. association of college and research libraries. https://alair.ala.org/handle/11213/17190 tenopir, c., kaufman, j., sandusky, r., & pollock, d. (2019). research data services in academic libraries: where are we today? (acrl/choice) [white paper]. http://choice360.org/librarianship/whitepaper 23/23 brodsky, meryl & hannah chapman tripp (2025). data competencies for liaison librarians: a scoping review, iassist quarterly 49(3), pp. 1-23. doi: https://doi.org/10.29173/iq1154 tenopir, c., sandusky, r. j., allard, s., & birch, b. (2014). research data management services in academic research libraries and perceptions of librarians. library & information science research, 36(2), 84–90. https://doi.org/10.1016/j.lisr.2013.11.003 tenopir, c., talja, s., horstmann, w., late, e., hughes, d., pollock, d., schmidt, b., baird, l., sandusky, r. j., & allard, s. (2017). research data services in european academic research libraries. liber quarterly: the journal of european research libraries, 27(1), 23–44. https://doi.org/10.18352/lq.10180 thomas, c., & urban, r. (2018). what do data librarians think of the mlis? professionals’ perceptions of knowledge transfer, trends, and challenges. college & research libraries, 79(3), 401–423. https://doi.org/10.5860/crl.79.3.401 university of edinburgh. (n.d.). mantra research data management training. mantra training kit for information professionals. https://mantra.ed.ac.uk/ the whitehouse office of science and technology policy. (2013: february 22). increasing access to the results of federally funded scientific research. memorandum. https://obamawhitehouse.archives.gov/sites/default/files/microsites/ostp/ostp_public_acce ss_memo_2013.pdf xia, j., & wang, m. (2014). competencies and responsibilities of social science data librarians: an analysis of job descriptions. college & research libraries, 75(3), 362–388. https://doi.org/10.5860/crl13-435 1 meryl brodsky is the communication & information librarian at the university of texas at austin. she can be reached at meryl.brodsky@austin.utexas.edu 2 hannah chapman tripp is the instruction & faculty outreach librarian at st. mary’s university of minnesota. she can be reached at hannah.chapman.tripp@gmail.com microsoft word 49-3-sizer.docx 1/20 sizer, alison, andreas mastrosavvas, & oliver duke-williams (2025). the ons longitudinal study – opportunities for longitudinal research on the england and wales population, iassist quarterly 49(3), pp. 1-20. doi: https://doi.org/10.29173/iq1159 the creative commons-attribution-noncommercial license 4.0 international applies to all works published by iassist quarterly. authors will retain copyright of the work and full publishing rights. the ons longitudinal study – opportunities for longitudinal research on the england and wales population alison sizer1, andreas mastrosavvas2 and oliver duke-williams3 abstract comprising longitudinal data on around 1.1 million individuals, the office for national statistics longitudinal study (ons-ls) is the largest nationally representative longitudinal dataset in the united kingdom. it follows a 1% sample of the england & wales population drawn from the decennial census data (1971 – 2011), linked to some administrative data. currently comprising up to 46 years of data (1971 – 2017) on sample members, the forthcoming linkage of the 2021 england and wales census data to the ons-ls will extend this follow-up to 50 years. the centre for longitudinal study information and user support (celsius) provides assistance for researchers wishing to use the onsls in their research. based at university college london (ucl) and the office for national statistics (ons), it has been supporting academic and voluntary sector users of the ons-ls since 2001. its work includes helping researchers with their applications to use the ons-ls, supporting research projects and advising on research outputs. keywords census, longitudinal study, longitudinal research introduction this paper introduces the office for national statistics longitudinal study (ons-ls) and the opportunities that it offers for longitudinal research on the england and wales population. it details the information held in the ons-ls, and it introduces the centre for longitudinal study information and user support (celsius) and the work that it does to support current and prospective users of the ons-ls. it provides details of the ons-ls’s sister studies, the scottish longitudinal study and northern irish longitudinal study and the opportunities that the three studies offer for doing comparative research across the uk nations. it compares the ons-ls to other census-based longitudinal studies in the international context, and concludes by highlighting the challenges faced by these census-based studies, and the ons-ls specifically. what is the ons-ls? the ons-ls was set up in 1974 by the then office for population censuses and surveys (opcs), now office for national statistics (ons). initial sample members were drawn from the respondents to the 1971 census and were selected on the basis of four undisclosed birthdays, giving an initial sample of approximately 1% of the england and wales population (1.1 million individuals). the study is maintained through the annual addition of new births and immigrants with the same four birthdays. it includes the census forms from the 1971, 1981, 1991, 2001 and 2011 censuses, linked to some vital registrations data, including the births of new sample members, deaths of sample members, cancer registrations, live births to sample mothers, still births to sample mothers, infant deaths and widowerhoods. 2/20 sizer, alison, andreas mastrosavvas, & oliver duke-williams (2025). the ons longitudinal study – opportunities for longitudinal research on the england and wales population, iassist quarterly 49(3), pp. 1-20. doi: https://doi.org/10.29173/iq1159 the ons-ls is representative of the population for england and wales and in contrast to other longitudinal studies it includes individuals living in communal establishments, for example, hospitals, residential care homes and nursing homes, prisons, boarding schools, hostels and hotels. it also includes information from the census forms of other individuals living in the same household as a sample member at the time of the census. this means that in childhood other individuals in an ls member’s household might be their parents and siblings, but later in their life, it might be their partner or spouse and their children. this offers opportunities to investigate inter-generational social change. the large size of the ons-ls, currently around 1.1 million sample members, means that it is also possible to examine small population groups, for example specific groups of ethnic minorities, older age groups, and individuals in residential care homes, which may not be possible in other longitudinal studies, which have smaller sample sizes. the ons-ls is a dynamic sample of all individuals completing a census and usually resident in england and wales with one of the four confidential birthdays. the sample has been maintained since its start on census day in 1971 (26th april 1971) in the following ways:  children subsequently born on one of the four confidential birthdays since 26th april 1971 are added to the sample.  immigrants registering with the nhs central registry (nhscr) and being born on one of the four confidential birthdays are added to the sample.  any individuals completing a subsequent census form and giving their birth date as one of the four confidential birthdays is added to the sample if they are not already included.  former ons-ls members who left the sample through emigration are re-entered into the sample on re-registering with the nhscr.  deaths of ls members are recorded on their records but their information is retained.  ls members who emigrate have their date of embarkation recorded but their records are retained and they can therefore re-enter the study on re-registering with the nhs (see above).  ls members who enlist with the armed forces have their date of enlistment recorded. again, their records are retained and therefore similarly to emigrants, they can re-enter the study. figure 1 illustrates the dynamic nature of the ons-ls sample and how it is maintained. it shows the samples at each of the censuses included in the study, and the members of the sample that have been traced to the nhscr, which facilitates the linkage of their record to the vital registrations information that is included in the study. it also shows the approximate number of entry events (births, immigrations, re-entries) and exit events (deaths, emigrations and enlistments) over the course of the study to 2017. 3/20 sizer, alison, andreas mastrosavvas, & oliver duke-williams (2025). the ons longitudinal study – opportunities for longitudinal research on the england and wales population, iassist quarterly 49(3), pp. 1-20. doi: https://doi.org/10.29173/iq1159 figure 1: ons longitudinal study: samples at each census and the linked registry information figure 2 shows the approximate numbers of ls members who are enumerated in each individual census (the 1971, 1981, 1991, 2001 and 2011 censuses); two consecutive censuses (the 1971-1981, 1981-1991, 1991-2001, and 2001-2011 censuses); three consecutive censuses (the 1971-1991, 19812001, and 1991-2011 censuses); four consecutive censuses (the 1971-2001, and 1981-2011 censuses); and all five censuses (1971-2011 censuses). figure 2: ons longitudinal study: samples for single census years and multiple consecutive censuses 4/20 sizer, alison, andreas mastrosavvas, & oliver duke-williams (2025). the ons longitudinal study – opportunities for longitudinal research on the england and wales population, iassist quarterly 49(3), pp. 1-20. doi: https://doi.org/10.29173/iq1159 the information included in the ons-ls? household and sociodemographic information collected through the decennial census (1971 – 2021) is included in the ons-ls, along with variables that are derived from this information. the variables derived from the census information include social class, derived from the employment and occupation related information in the census, and 10-year migration indicators, derived by comparing the current address in successive censuses. the household information included in the ons-ls covers type of accommodation and whether it is self-contained, tenure, household amenities (e.g. type of central heating), number of rooms, and number of cars/ vans. the individual related topics included in the ons ls include age, sex, country of birth and marital/ civil partnership status; educational qualifications; economic position or activity (employed/ unemployed/ retired/ inactive); employment status (employed/ self-employed, full-time/ part-time), occupation, work address, number of hours worked and travel to work mode; and geographical location (current and historical). in addition to these variables, which are included across the censuses, topics were added to the censuses in 1981, 1991, 2001 and 2011. in 1991 (office of population censuses and surveys, 1991), ethnicity and longterm limiting illness4 were added to the census. in 2001 (office for national statistics, 2001), a voluntary question on religion5, and questions on informal caregiving6 and self-rated health7 were added. in 2011 (office for national statistics, 2011), questions were added related to national identity8, passports held9, time since arrival in the united kingdom10, and intended length of stay11. on 21st march 2021, the most recent census was held in england and wales12. further questions were added to the census, voluntary response questions related to gender identity13 and sexuality14, a revised question related to marital or civil partnership status15, and one related to being a veteran of the armed forces16. the linkage of the 2021 census to the ons-ls, and the testing of the linkage, is currently on-going and it is hoped that the linkage will be completed in the latter half of 2025. this linkage will extend the follow-up of the ons-ls to 2021 (50 years), enabling researchers to examine changes that have taken place in 2011 – 2021 period, a period which saw brexit and the covid 19 pandemic. table 1 lists the topics included in the census forms and the ons-ls, and the census years that they were addressed. table 1: topics included in the ons ls by census year (1971 – 2021) (data source: ons ls) topic census year household related 1971 1981 1991 2001 2011 2021 loca on (various geographical levels)       tenure       whether accommoda on is self-contained       type of accommoda on (caravan, flat, temporary structure, block of flats, whole house)      number of rooms       household ameni es (e.g. hea ng, hygiene)       number of cars/ vans       individual related 1971 1981 1991 2001 2011 2021 sex       date of birth       rela onship to head of household/ household reference person       usual resident or visitor at enumera on address  geographical loca on one-year ago       5/20 sizer, alison, andreas mastrosavvas, & oliver duke-williams (2025). the ons longitudinal study – opportunities for longitudinal research on the england and wales population, iassist quarterly 49(3), pp. 1-20. doi: https://doi.org/10.29173/iq1159 geographical loca on five-years ago  geographical loca on of second address   marital status (including civil partnership)       sex of marital or civil partner  marital history (women under-60 who are married/ widowed/ divorced)  live born children (women under-60 who are married/ widowed/ divorced)  country of birth       mother and father country of birth  na onal iden ty   passports held   year of arrival in uk (if born outside uk)   intended length of stay in uk   ethnic group     religion (voluntary answer)    welsh language (wales only) (speak, read, write)       main language, english language ability   sexuality (voluntary answer)  gender iden ty (voluntary answer)  educa on qualifica ons (nb: ques on varies over me)       working/ unemployed/ re red/ inac ve (previous week)       whether a student (previous week)       employment status (employed/ self-employed, full/ partme)       occupa on       loca on of main place of work  work address       occupa on one year ago  year last worked    hours worked       transport to work mode       long-term limi ng illness     self-rated health    informal caregiving    veteran status  life events data have also been linked to the ons-ls, which means that the ls also includes information that is not addressed through the decennial census, including births, deaths and cancer registrations. the life events data in the ons-ls are updated annually with a two-year delay. linkage happens through two processes:  the birth and death register and cancer registration files that are sent to ons annually are searched for individuals with one of the four ons-ls birthdates. these are then linked to actual sample members. 6/20 sizer, alison, andreas mastrosavvas, & oliver duke-williams (2025). the ons longitudinal study – opportunities for longitudinal research on the england and wales population, iassist quarterly 49(3), pp. 1-20. doi: https://doi.org/10.29173/iq1159  sample members who are marked as ons-ls sample members in the nhscr are sent to the longitudinal study team at ons when a birth or death is recorded in the register. these are then added to the ons-ls database. currently the following life events data from 1971 to 2017 have been linked to ons-ls study members records (centre for longitudinal study information and user support, 2018):  new births of sample members.  deaths of sample members.  live births and still births registered to sample mothers.  infant deaths of births registered to sample mothers.  cancer registrations.  widowerhoods (these are drawn from death registrations when an ls member’s legal spouse/ civil partner dies, and the ls members are identified through the date of birth of surviving spouse information in the death registration form). the information available in the ons-ls from the linkage comprises information that is recorded in the birth, death and cancer registration forms. table 2 shows the information available in the ons-ls according to the different life events that are linked to it. table 2: topics included in the life events information that is linked to ons-ls members (data source: ons-ls) life events data number of events (1971 – 2017) information available new births of sample members 340,542 day, month and year of birth* place and geographical location of birth sex birthweight whether multiple birth; type of multiple birth number of previous live or still births to mother parent’s age occupation of working parent deaths 293,931 day, month and year of death place and geographical location of death underlying and contributory causes of death (icd8, icd9, icd10)** whether the death was in a communal establishment and establishment type live births to sample mothers 331,304 date of birth sex place and geographical location of birth birthweight whether multiple birth; type of multiple birth number of previous live or still births to mother 7/20 sizer, alison, andreas mastrosavvas, & oliver duke-williams (2025). the ons longitudinal study – opportunities for longitudinal research on the england and wales population, iassist quarterly 49(3), pp. 1-20. doi: https://doi.org/10.29173/iq1159 age of sample mother at birth still births to sample mothers 1,864 date of birth sex place and geographical location of birth birthweight whether multiple birth; type of multiple birth number of previous live or still births to mother age of sample mother at birth whether the sample mother was married at the time of the birth (and year of current marriage) social class, employment status and occupation of sample mother and father infant deaths to sample mothers*** 2,418 date of birth of infant/ child date of death of infant/ child sex of child usual address of sample mother at time of infant/ child birth and death place of birth and place of death birthweight whether multiple birth; type of multiple birth cause of infant death (according to icd8, 9, 10) number of previous live and still births to mother age of sample mother at birth whether the sample mother was married at the time of the birth (and year of current marriage) social class, employment status and occupation of sample mother and father cancer registrations (to 2016) 160,727 date of cancer registration age at time of registration site and sub-site of cancer treatment type date of death (if dead) widow(er)hoods 99,892 age and date of death of ls member’s spouse cause of death of ls member’s (according to icd8, icd9, icd10) 8/20 sizer, alison, andreas mastrosavvas, & oliver duke-williams (2025). the ons longitudinal study – opportunities for longitudinal research on the england and wales population, iassist quarterly 49(3), pp. 1-20. doi: https://doi.org/10.29173/iq1159 notes to table 2: * access to the day and month of birth requires special permission from ons, because they will disclose the confidential birthdays used to select ons-ls sample members and hence disclose the identity of ls members. **three different versions of the international classification of diseases (icd) have been used over the course of the ons-ls: icd8 (from census day 1971 to 4th april 1981); icd9 (from 5th april 1981 to 31st december 2000); and icd10 (from 1st january 2001 onwards). *** until january 1993 this only included infants up to one-year old. following a change in the national database, from january 1993, the death of any child born in 1993 was included in the ls, which means that currently this includes deaths of children under 16-years old. the life courses of three hypothetical ons-ls members based on their enumeration through the census and the life events information that is linked to their ons-ls record through the nhscr are shown in figure 3. the first ons-ls member is included from the 1971 census on account of being born on one of the four ls birthdays. they have a child between the 1971 and 1981 censuses and then emigrate in 1985. they return to the uk and re-enter the sample in 1992 through registration with a gp. the ons-ls member then dies in 1997. the second ons-ls member migrates into england and wales in 1979 and enters the sample through registering with the nhscr and being born on one of the four confidential birthdays. they then complete the 1981 census. between the 1981 and 1991 censuses they have a live birth and subsequently that child dies (infant death). they complete the 1991 census and a further child is born between the 1991 and 2001 censuses. they then complete the 2001, but prior to the 2011 census their spouse dies and they are subsequently diagnosed with cancer in 2009. they go onto complete the 2011 census. the third ls member is born in 1978 and is then enumerated in the 1981, 1991 and 2001 censuses. after the 2001 census they are diagnosed with cancer and then they die in 2007. figure 3: the life courses of three hypothetical ons longitudinal study members geographical information included in the ons-ls the ons-ls includes the geographical location of ls members at the time of census based on the address on the front of the census form. the location is recorded at a range of different geographical levels including government office region, county and county district, local authority district and ward. location at lower geographical levels is also recorded including postcode area, output area 9/20 sizer, alison, andreas mastrosavvas, & oliver duke-williams (2025). the ons longitudinal study – opportunities for longitudinal research on the england and wales population, iassist quarterly 49(3), pp. 1-20. doi: https://doi.org/10.29173/iq1159 (oa)17, lower super output area (lsoa)18 and middle super output area (msoa)19 (office for national statistics, 2022). however, these lower-level geographical variables are subject to additional access permissions and restrictions owing to the potential for the identification (disclosure) of individual ls members. in addition to geographical location of the ls member at the time of the census, the census forms also ask for the address one year ago if it is different (the 1971 census also asked about the address five years ago). this information can be used to look at one-year moves for ls members, and comparing the geographical location of ls members in subsequent censuses enables 10-year moves to be examined. in addition to using the geographical location of ls members to examine migration between geographical areas in england and wales, the geographical location of ls members can be used to link ecological variables to the ons-ls, including neighbourhood deprivation (norman & boyle, 2014), population density (norman & boyle, 2014) and pollution data. travel to work areas (ttwas) and workplace zones of ls members are also included in ons-ls, based on their workplace address. ttwas are defined to approximate self-contained labour-market areas within which commuting to and from work occur. they are areas where the number of people living and working in the area should be at least 75% of the total number of workers living in the area and the total number of people working in the area (hattersley & creeser, 1995). in contrast to ttwas, workplace zones are areas where people work, and they are designed to have similar numbers of workers within them (200 625 workers). there were 60,709 workplace zones in 2011 (office for national statistics, 2012). previous research using the ons-ls the ons-ls currently has 46 years of follow-up data, 1971 – 2017, and it has been used by researchers to examine issues with social policy implications, including social inequalities in health, employment and education; social exclusion including the long-term health implications of non-employment and low educational status; housing and geographical mobility; and family policy. the research summarised below describes four recently completed and on-going research projects using the onsls. murray et al. (2020) and sacker et al (2024) used the ons-ls to examine the health of looked after children when they are grown up. their research found that adults who had been in care in any of the 1971 – 2001 censuses faced higher risk of mortality by 2013 than those ls members who had never been in care. they also found that the excess mortality was predominantly attributable to three causes, self-harm, accidents or mental and behavioural disorders (murray, lacey, maughan, & sacker, 2020; sacker, murray, maughan, & lacey, 2024). sturley et al. (2023) used the ons-ls to study variations in colorectal cancer incidence and survival over a 15-year period by individual characteristics and area type. the researchers found that compared to ls members without a degree, those with a degree had a lower incidence of colorectal cancer. it was also higher among individuals who were employed in manual occupations compared to ls members employed in non-manual occupations. furthermore, these disparities were even more pronounced in terms of colorectal cancer survival. although there was no clear relationship between area deprivation and colorectal cancer incidence among ls members, individuals living in the most deprived areas had higher probability of death than those living in the least deprived areas (sturley, norman, morris, & downing, 2023). researchers at the institute for fiscal studies (ifs) have used the ons-ls to examine intergenerational social mobility in england and wales (bourqin, joyce, & sturrock, 2021). they analysed data on ls members born in the 1960s and 1980s comparing their average inheritances to their lifetime projected income. the study found that for ls members born in the 1960s, average inheritances were worth 9% 10/20 sizer, alison, andreas mastrosavvas, & oliver duke-williams (2025). the ons longitudinal study – opportunities for longitudinal research on the england and wales population, iassist quarterly 49(3), pp. 1-20. doi: https://doi.org/10.29173/iq1159 of their projected household lifetime income (non-inheritance), whereas for those born in the 1980s this figure rose to 16%, suggesting that inheritances were set to increase inequalities in lifetime income between individuals with richer and poorer parents. finally, the ons-ls is being used to examine the labour market impact of trade shocks (irastorzafadrique, levell, & parey, 2024). this research is looking at the ls member’s own and their partner’s responses to the decline of manufacturing employment due to import substitution and has found that responses vary significantly by gender. men in households exposed to import competition respond to shocks by increasing their labour force participation into older ages, and by moving into solo selfemployment. furthermore, this response is observed both when men themselves lose employment and when their partners are affected. in contrast however, women did not increase their labour supply if their male partners were initially employed in exposed industries. applying to use the ons longitudinal study data although researchers are free to use the ons-ls data, access is strictly controlled. researchers wanting to access the ons-ls must be accredited researchers under the digital economy act (2017). this allows them to access secure ons data, such as the ons-ls, through the secure research service (srs) or another trusted research environment (tre). researchers apply to be an accredited researcher through the people and project services (pps) platform (see: https://integrateddataservice.gov.uk/apply-for-researcher-accreditation). accredited researchers must have an undergraduate degree or higher, with a significant proportion of maths or statistics, or be able to demonstrate at least three years of experience of quantitative research (office for national statistics, 2024b). once a researcher has applied to be an accredited researcher through the pps, they are invited to complete compulsory online safe researcher training. the training covers the “five safes”20 (ritchie, 2017), the safe handling of data, and the disclosure control rules for the research outputs. the disclosure control rules are applied to research outputs in order to minimise the risk of the disclosure of personal information about the data subjects. accredited researchers are required to renew their accreditation every five years, involving attending refresher training (uk statistics authority, 2020). accredited researchers wanting to use the ons-ls must apply for access to the ons-ls data. the application for the ons-ls data comprises two parts:  research accreditation application form.  ons longitudinal study supplementary form. both are available through the celsius website (centre for longitudinal study information and user support, 2024b). information required for the research accreditation application form focuses on the research aim and objectives, as well as the research questions and hypotheses. researchers are required to set out their research methodology and identify any research biases. in addition, all research using the ons-ls must demonstrate public benefit. the application form also asks if external data will be linked to the onsls, for example pollution data or neighbourhood deprivation levels. if such linkage is intended, researchers must explain the rationale for the linkage, provide the source and owner of the external data, confirm that permission to use the data has been granted, and describe the method of linkage to the ons-ls. in the application form all the members of the research team must be identified, and those requiring access to the data, rather than just reviewing outputs from the data, must be accredited researchers. (centre for longitudinal study information and user support, 2022) in the ons longitudinal study supplementary form the researcher describes their study sample and the variables that they will require for their research project. researchers will need to refer to the 11/20 sizer, alison, andreas mastrosavvas, & oliver duke-williams (2025). the ons longitudinal study – opportunities for longitudinal research on the england and wales population, iassist quarterly 49(3), pp. 1-20. doi: https://doi.org/10.29173/iq1159 ons-ls data dictionary for a list of the variables available for researchers to use in their ons-ls research project (see: https://iweb.dis.ucl.ac.uk/celsius/standalone/searchresults.php). the data dictionary groups the variables into different tables based on their source. for example, there are separate tables for each census and separate tables for ons-ls members (me71, me81, me91, me01 and me11) and non-members21 (nm71, nm81, nm91, nm01, nm11). there are also tables for each of the types of linked events data (for example the table called “deth” includes the variables drawn from the death registration, and the table called “nbir” includes the variables based on the birth registrations of new ons-ls study members). this grouping of the variables into a series of tables means that researchers can browse through the variables in each table. it is also possible to do a dictionary search of the data dictionary in order to search for specific variables or keywords. (centre for longitudinal study information and user support & office for national statistics, 2014) once a researcher has completed both parts of the research accreditation application (application form and ons ls supplementary form), they should send them to the centre for longitudinal study information and user support (celsius) to submit on their behalf. the application then undergoes a series of feasibility checks by ons, the ls data owners and ons-ls research board, and finally ons’s research accreditation panel. if none of these raise any serious concerns about the research project being applied for, the application is approved. once an application to use the ons-ls has been approved, all the accredited researchers named in the research team must sign an accredited researchers assurance registration (arar) form. the form sets out the commitments that accredited researchers must make before they will be able to access the secure research service (srs) environment through which they will access their data. it consists of sections which address, i) the responsibilities of the accredited researcher; ii) security of the data; iii) lawful access and data confidentiality; iv) data clearance; v) non-compliance and breach sanctions; vi) dispute procedures; and vii) incident management and reporting. accessing the ons ls data once an application to use the ons-ls has been approved and the data has been extracted, there are four routes through which researchers can access their project data. the first of these is through one of the ons safe researcher settings at their offices in titchfield (hampshire) or newport (wales). since august 2021 ons has allowed researchers to securely access their ons-ls data through one of a network of safepods22 around the uk. currently there are 23 safepods, 18 in england (including two in london), four in scotland and one in northern ireland (university of st. andrews, 2021). some organisations have an assured organisational connectivity (aoc) agreement with ons, which enables researchers to access their data via the srs from their office. subject to further ons approval, researchers may also be able to access their data from their home via their institution-owned computer (office for national statistics, 2024a). finally, for researchers who are unable to access their data in any other way, celsius user support officers will run researcher written syntax and then send them the output. the centre for longitudinal study information and user support (celsius) based at university college london (ucl), the centre for longitudinal study information and user support (celsius) was set up in 2001 to support academic and voluntary sector users of the ons-ls. this support includes helping researchers with their ons-ls research applications, guiding and assisting them with their analysis and advising on publications and other research outputs that they might want to consider using to share their research findings, e.g. podcasts and blogs. celsius maintains a website that provides access to a wealth of information on the ons-ls including the data dictionary; links to the census forms; a series of thematic guides to using the ons-ls; blogs and podcasts showcasing research using the ons-ls; and recordings of webinars delivered by the celsius team (centre for longitudinal study information and user support, 2024a). celsius user support 12/20 sizer, alison, andreas mastrosavvas, & oliver duke-williams (2025). the ons longitudinal study – opportunities for longitudinal research on the england and wales population, iassist quarterly 49(3), pp. 1-20. doi: https://doi.org/10.29173/iq1159 officers (usos) have also written syntax for deriving variables that many ls researchers use in their analysis, e.g. cause of death, and categorisations of country of birth. figure 4 shows the celsius website home page. figure 4: celsius website home page (centre for longitudinal study information and user support, 2024a) future directions of the ons-ls as stated above, access to ons-ls data is strictly controlled to protect identifying information, e.g., the confidential inclusion dates. although extensive user support from celsius is provided to guide researchers through application and data access processes, researchers can face challenges due to their inability to view data before approval. despite the existence of the ons-ls data dictionary, it lists thousands of variables, which can make identifying relevant variables difficult, particularly for new users, potentially resulting in delays to the application approval process. additionally, researchers cannot begin writing scripts or testing code until after gaining access to their data in the secure environment, further slowing the research process. to address these issues, celsius is developing the longitudinal impossible dataset (lids): an artificial, user-generated dataset that mimics the structure of ons-ls data without posing disclosure risks. lids supports users in project planning and script development by allowing them to select variables and download artificial data generated from publicly available metadata via an interactive interface. as lids is generated by taking random draws from publicly available code lists, it remains informative with respect to the coding of individual variables while eliminating real data disclosure risks. additional features, such as the inclusion of ‘impossible’ relationships between variables due to randomisation and adherence to the ons-ls longitudinal linkage convention of tracking participants across censuses via unique identifiers, further facilitate data discovery and script development while minimising the risk of perceived disclosure (mastrovasas & shelton, 2025). 13/20 sizer, alison, andreas mastrosavvas, & oliver duke-williams (2025). the ons longitudinal study – opportunities for longitudinal research on the england and wales population, iassist quarterly 49(3), pp. 1-20. doi: https://doi.org/10.29173/iq1159 the uk longitudinal studies and opportunities for research across the four nations of the united kingdom the ons-ls, which covers england and wales, is one of three sister studies that cover the four uk nations (england & wales, northern ireland, scotland). the other two cover northern ireland, the northern ireland longitudinal study (nils), and scotland, the scottish longitudinal study (sls). although they are all based on linking the decennial census forms and comprise dynamic samples, they differ in terms of the years that they cover, the sample selection and the sample size. as stated above, the ons-ls comprises a 1% sample of the population of england & wales, which is based on being born on one of four birthdays, and started with the 1971 census. nils is based on a 28% sample of the northern ireland population, starting with the 1981 census, selected using 104 birthdays. finally, sls is based on a 5.3% sample of the scottish population, starting with the 1991 census, selected using 20 birthdays. the three studies also differ in terms of the additional data that they are linked to. all the studies are linked to birth and death registrations, immigrations and embarkations, but there are additional linkages that differ between the studies. the ons-ls is also linked to cancer registrations. nils can also be linked to information on the type and value of accommodation from the land and property services, and with special approval can be linked to health data, including breast cancer screening, dental treatment and the prescription of medicines. the sls also includes marriages from the civil registration system, weather and pollution data, school level data from the scottish government directorate, and with additional permission, linkage to hospital episodes, maternity data and cancer data. the three uk longitudinal studies are supported by different agencies, and have separate websites with information about the study, research using it and how to apply for access to each study’s data. celsius supports current and prospective researchers using the ons-ls1; the scottish longitudinal study development and support unit (sls-dsu) supports researchers using the sls2; and the northern ireland longitudinal study research support unit (nils-rsu) supports researchers using nils3. the websites also include information on how researchers can access their data4. cross-nation research and analysis while it is possible to undertake comparative research using data from the three sister studies, owing to disclosure concerns, it is not possible to fully combine the three datasets to create a uk-wide sample. however, there are alternative possibilities, one of which is to undertake analysis on the three datasets separately and then compare the results or combine them on an ad hoc basis. however, this approach has several disadvantages. first, it is not possible to be sure that the datasets and variables which appear to be similar are comparable. second, an analysis that adjusts for covariates in each individual study will not be identical to one that would be obtained if the raw datasets were combined. finally, tests of region by covariate interactions cannot be carried out easily through the comparison or combination of the region-specific results (raab & diben, 2020). 1 website: h ps://www.ucl.ac.uk/popula on-health-sciences/epidemiology-health-care/research/ucl-researchdepartment-epidemiology-public-health/research/health-and-social-surveys-research-group/studies/celsius; email: celsius@ucl.ac.uk. 2 website: h ps://sls.lscs.ac.uk/; email: sls@lscs.ac.uk. 3 website: h ps://nils.ac.uk/; email: nils@qub.ac.uk. 4 ons-ls data can be accessed in four ways : safe researcher se ngs in the ons offices in titchfield and newport; safepods; through an aoc agreement with ons; or via a home access aoc agreement. nils and sls are only accessible through workspaces in a safe se ng room. for nils the safe se ng room is at the northern ireland sta s cs and research agency (nisra) research support unit in belfast. for sls the safe se ng room is at na onal records scotland in edinburgh. 14/20 sizer, alison, andreas mastrosavvas, & oliver duke-williams (2025). the ons longitudinal study – opportunities for longitudinal research on the england and wales population, iassist quarterly 49(3), pp. 1-20. doi: https://doi.org/10.29173/iq1159 in response to these disadvantages, the edatashield (e-mail data aggregation through anonymous summary statistics from harmonised individual level databases) system was developed by researchers at the scottish longitudinal study. this system is an adaptation of the datashield (data aggregation through anonymous summary statistics from harmonised individual level databases) system (see https://www.datashield.org/), whereby joint analysis of data from separate data centres is undertaken iteratively by linking a computer in each data centre to an analysis computer. the analysis computer, which does not hold any data, receives summary statistics from generalised linear models from each data centre and then combines them and passes them back to the individual data centres. the interface between the analysis computer and the individual data centres prevents any exchange of the raw data so addressing ethics-related data-sharing concerns related to privacy, confidentiality and the protection of study participants rights (budin-ljøsne et al., 2015; wolfson et al., 2010). however, it is not possible to link the three uk longitudinal studies through an analysis computer in this way owing to potential disclosure concerns. the edatashield protocol provides a work-around for this by allowing the exchange of the summary statistics by email, and so enabling generalised linear models or ordered logistic regression models to be fitted iteratively (raab, 2013). research that has been carried out using data from two or more of the studies, enabled by the edatashield protocol, includes examining levels of “excess” mortality in scotland and glasgow compared to elsewhere in the uk (ralston et al., 2017), and examining the effect of religion on mortality in scotland and northern ireland (wright et al., 2017). ralston et al. (2017) used the sls and ons-ls to examine the risk of all cause mortality between 2001 and 2010 and compared it for residents aged 35-74 between scotland and england and wales, and between residents of glasgow (scotland) and liverpool/ manchester (england). they found that after adjusting for age, gender and socio-economic status, all-cause mortality was 9% higher in scotland versus england and wales, and 27% higher in glasgow versus liverpool/ manchester (ralston et al., 2017). using the sls and nils, wright et al. (2017) compared socio-economic status and mortality among protestants and catholics. the researchers found that in fully-adjusted models catholic men and women in scotland experienced higher mortality risk compared to protestants, whereas no such difference was observed in northern ireland (wright et al., 2017). researchers who are interested in undertaking comparative research using data from the two or more of the three uk longitudinal studies will need to apply for accredited research projects for each of the studies that they want to include in their research. therefore, if a researcher wanted to undertake research using data from the ons longitudinal study and nils, they would need to submit separate research applications for the ons-ls and nils and submit them to their respective support units (celsius and nils-rsu). international context of the ons-ls the ons-ls and its sister studies, sls and nils, are among several other longitudinal census-based datasets. similar studies include ipums usa (integrated public use microdata series usa) and the australian census longitudinal study. ipums usa collects, preserves and harmonises us census microdata and provides access to it. it includes 50 high-precision samples of the american population drawn from fifteen federal censuses. they draw on every surviving census from 1850-2010, and each sample case represents between 20 and 1,000 individuals in the whole us population meaning that weights need to be used in analysis. the data includes harmonised income and occupation variables, but unlike the ons-ls and its sister studies, there is no guarantee that an individual in one sample is in the one for the next wave of data and there is limited geographic information (ipums usa, 2025). the australian census longitudinal study (acls) comprises four waves of census data (2006-2021: 2006-2011, 2011-2016, 2016-2021). three acls data panels, representing a 5% sample of individuals enumerated in australia on the 2006, 2011 and 2016 census nights. it began with a 5% panel sample 15/20 sizer, alison, andreas mastrosavvas, & oliver duke-williams (2025). the ons longitudinal study – opportunities for longitudinal research on the england and wales population, iassist quarterly 49(3), pp. 1-20. doi: https://doi.org/10.29173/iq1159 from the 2006 census, which was linked to the subsequent censuses, with samples from later censuses being formed using the same sample selection strategy. this was designed to maintain a linked sample size of 5% but also introduce new records to subsequent panels to account for new births, migrants and missed links in previous panels. similarly to the ons-ls, the acls includes people living in both private and communal establishments, but it doesnot include linkage to adminstrative data (e.g., births and deaths). currently it has been used to investigate how family structure has changed over time, the characterisitcs and changes resulting from single parenthood and whether australians who were unemployed in 2011 and moved regions by 2016 were more likely to be employed than those who remained in the same area (australian bureau of statistics, april 2024). a feasibility study for the establishing of a new zealand longitudinal census (nzlc) has also been done (didham, nissen, & dobson, 2014). it examined linking the 1981, 1986, 1991, 1996, 2001 and 2006 censuses, starting with the linkage of five census pairs (1981-1986, 1986-1991, 1991-1996, 19962001, 2001-2006) and then linking the five census pairs to each other, with linkage done using date of birth, sex and area unit of usual residence. the study also examined linking birth and death registrations and international migration data. despite demonstrating the feasibility of such a longitudinal census and making a series of recommendations including developing confidentiality rules to enable the secure remote access to the nzlc by approved researchers, finalising the linkage of birth and death registrations, and investigating linkage to other sources of administrative data, the nzlc remains as an experimental dataset. challenges to the ons-ls and other census-based longitudinal studies a generic threat to census-based longitudinal studies including the ons-ls is the possible cessation of traditional census approaches, in favour of methods that derive equivalent data from administrative sources. ons (2023) set out an ambitious intention to create a ”system [that] would primarily use administrative data like tax, benefit and border data, complemented by survey data and a wider range of data sources” (office for national statistics, 2023). this would most likely have removed the requirement for a further traditional census. the potential for use of administrative data has also been considered elsewhere, for example, australia bureau of statistics (2020) (australian bureau of statistics, october 2020) and stats new zealand (2022) (stats nz, 2022). whilst it is now recommended that full censuses are taken in the uk in 2031, a switch to administrative data would be a significant challenge. administrative records could continue to be attached to the ons-ls base as long as sufficient individual level data were available to carry out that linkage. however, upstream suppliers of data (government departments, etc.) may be unwilling to provide such data, or legally restricted from doing so. furthermore, characteristics collected in census results but not readily available through administrative data would not be included. for the purposes of crosssectional reporting, administrative data could be supplemented with data from regular specialised population surveys. however, any such surveys would have a different sampling framework to the ons-ls (and indeed could not replicate the ons-ls sample without disclosing the ons-ls confidential inclusion dates), and thus could not be appended to the the ons-ls database. the extent to which this challenge is faced by census-based longitudinal studies beyond the uk would depend on the sampling frameworks used, and the linking fields available. if there are common individual identifiers, such as id numbers, used in different data sources, then linkage would be, at least in principle, straightforward. however, where multiple non-comprehensive identifiers are used, linkage becomes harder. other challenges are economic. whilst such studies offer uniquely rich data, they are also complex and tend to be used only by data and subject specialists. they therefore run the risk of being seen as expensive to maintain, and thus liable to funding cuts. 16/20 sizer, alison, andreas mastrosavvas, & oliver duke-williams (2025). the ons longitudinal study – opportunities for longitudinal research on the england and wales population, iassist quarterly 49(3), pp. 1-20. doi: https://doi.org/10.29173/iq1159 conclusion the ons longitudinal study, follows a 1% sample of the england & wales population from the decennial census data (1971 – 2011), linked to births, deaths and cancer registration data. sample members are selected on the basis of four confidential birthdays, with new study members entering the study through birth on one of the four birthdays or immigration (and being born on one of the four birthdays) and leave through death or emigration. the main strength of the ls is its large sample size (>1 million), making it the largest nationally representative dataset in the uk, and allowing the analysis of small areas and specific population groups. currently, the ons-ls has 46 years of followup data 1971 – 2017, and it has been used to examine issues with social policy implications, for example social inequalities in health, employment and education; housing and geographical mobility; and family policy. the upcoming linkage of the 2021 census data to the ons-ls will extend the span of the study to 50 years, and will enable researchers to examine changes that have taken place in 2011 – 2021 period, a period which saw brexit and the covid 19 pandemic. the ons-ls is one of the three “sister” studies covering the united kingdom. while there are similarities between the three studies there are important differences related to their sample size and the additional information that they linked to. that being said all three of the studies provide opportunities for analyses with large sample sizes and they are supported by teams of researchers all of whom have undertaken research using one or more of the studies. research using data from two or more of the studies is possible enabled by the edatashield protocol. indicatively, past research using edatashield, has examined levels of “excess” mortality in scotland and glasgow compared to england and wales and liverpool/ manchester using data from the ons-ls and sls (ralston et al., 2017). it has also been used to examine the effect of religion on mortality in scotland and northern ireland (wright et al., 2017). the ons-ls and its sister studies are among several census-based longitudinal studies worlwide. however, such studies face the on-going threat of the possible cessation of traditional census approaches and the economic challenges posed by their perception as being expensive to run. that being said, census approaches do appear to be continuing, and funders continue to recognise the value such studies offer for research into social inequalities in health, education and employment, social exclusion and mobility, and family policy. acknowledgement and disclaimer the permission of the office for national statistics to use the longitudinal study is gratefully acknowledged, as is the help provided by staff of the centre for longitudinal study information & user support (celsius). celsius is funded by the esrc under project es/v003488/1. the authors alone are responsible for the interpretation of the data. the help provided by staff of the longitudinal studies centre—scotland (lscs) is acknowledged. the lscs is supported by the esrc/jisc, the scottish funding council, the chief scientist’s office, and the scottish government. the authors alone are responsible for the interpretation of the data. census output is crown copyright and is reproduced with the permission of the controller of hmso and the queen’s printer for scotland. the help provided by the staff of the northern ireland longitudinal study/northern ireland mortality study (nils/nims) and the nils research support unit is acknowledged. the nils/nims is funded by the health and social care research and development division of the public health agency (hsc r&d division) and nisra. the nils-rsu is funded by the esrc and the northern ireland government. the authors alone are responsible for the interpretation of the data and any views or opinions presented are solely those of the author and do not necessarily represent those of nisra/nils. 17/20 sizer, alison, andreas mastrosavvas, & oliver duke-williams (2025). the ons longitudinal study – opportunities for longitudinal research on the england and wales population, iassist quarterly 49(3), pp. 1-20. doi: https://doi.org/10.29173/iq1159 references australian bureau of statistics. (april 2024). microdata and table builder: australian census longitudinal dataset. retrieved from https://www.abs.gov.au/statistics/microdatatablebuilder/available-microdata-tablebuilder/australian-census-longitudinal-dataset#citewindow1 australian bureau of statistics. (october 2020). assessing administrative data quality to enhance the 2021 census. retrieved from https://www.abs.gov.au/statistics/research/assessingadministrative-data-quality-enhance-2021-census bourqin, p., joyce, r., & sturrock, d. (2021). inheritances and inequality over the life cycle: what will they mean for younger generations? retrieved from london, uk: https://ifs.org.uk/sites/default/files/output_url_files/r188-inheritances-and-inequalityover-the-lifecycle%252520%2525281%252529.pdf budin-ljøsne, i., burton, p., isaeva, j., gaye, a., turner, a., murtagh, m. j., . . . harris, j. r. (2015). datashield: an ethically robust solution to multiple-site individual-level data analysis. public health genomics, . public health genomics, 18(2), 87-96. doi:https://doi.org/10.1159/000368959 centre for longitudinal study information and user support. (2018). events. retrieved from london, united kingdom: https://www.ucl.ac.uk/epidemiology-health-care/research/epidemiologyand-public-health/research/health-and-social-surveys-research-group/studies-21 centre for longitudinal study information and user support. (2022). research project accreditation application guidance. london: university college london. centre for longitudinal study information and user support. (2024a, may 2024). celsius. retrieved from https://www.ucl.ac.uk/epidemiology-health-care/research/epidemiology-and-publichealth/research/health-and-social-surveys-research-group/studies-10 centre for longitudinal study information and user support. (2024b). want to use the ons ls for your research? retrieved from https://www.ucl.ac.uk/epidemiology-healthcare/research/epidemiology-and-public-health/research/health-and-social-surveysresearch-group/studies-14 centre for longitudinal study information and user support, & office for national statistics. (eds.). (2014). london, united kingdom: university college london. didham, r., nissen, k., & dobson, w. (2014). linking censuses: new zealand longitudinal census 1981-2006. retrieved from wellington, nz: https://www.stats.govt.nz/methods/linkingcensuses-new-zealand-longitudinal-census-19812006/ hattersley, l., & creeser, r. (1995). longitudinal study 1971-1991: history, organisation and quality of data. london: hmso retrieved from https://www.ucl.ac.uk/epidemiology-healthcare/sites/epidemiology-health-care/files/ls7.pdf ipums usa. (2025). us census data for social, economic and health research. retrieved from https://usa.ipums.org/usa/ irastorza-fadrique, a., levell, p., & parey, m. (2024). household responses to trade shocks. retrieved from london: https://ifs.org.uk/publications/household-responses-trade-shocks-0 mastrosavvas, a., & shelton, n. (2025). the longitudinal impossible dataset: helping users navigate the ons longitudinal study. paper presented at the iassist, 2025, bristol, uk. https://zenodo.org/records/15611228 murray, e. t., lacey, r., maughan, b., & sacker, a. (2020). association of childhood out-of-home care status with all-cause mortality up to 42-years later: office of national statistics longitudinal study. bmc public health, 20(1), 735. doi: https://doi.org/10.1186/s12889-020-08867-3 18/20 sizer, alison, andreas mastrosavvas, & oliver duke-williams (2025). the ons longitudinal study – opportunities for longitudinal research on the england and wales population, iassist quarterly 49(3), pp. 1-20. doi: https://doi.org/10.29173/iq1159 norman, p., & boyle, p. (2014). are health inequalities between differently deprived areas evident at different ages? a longitudinal study of census records in england and wales, 1991-2001. health place, 26, 88-93. doi: https://doi.org/10.1016/j.healthplace.2013.12.010 office for national statistics. (2001). census 2001: england household form. united kingdom: office for national statistics. office for national statistics. (2011). 2011 census: houshold questionnaire england. united kingdom: office for national statistics. office for national statistics. (2012). 2011 census geographies. retrieved from https://www.ons.gov.uk/methodology/geography/ukgeographies/censusgeographies/2011c ensusgeographies office for national statistics. (2022). census 2021 geographies. retrieved from https://www.ons.gov.uk/methodology/geography/ukgeographies/censusgeographies/censu s2021geographies office for national statistics. (2023). the future of population and migration statistics in england and wales: a consultation on ons proposals, 29 june 2023. retrieved from london: https://consultations.ons.gov.uk/ons/futureofpopulationandmigrationstatistics/supporting_ documents/future%20of%20population%20and%20migration%20statistics%20consultation %20document.pdf office for national statistics. (2024a). access the data securely: assured organisational connectivity (aoc). retrieved from https://www.ons.gov.uk/aboutus/whatwedo/statistics/requestingstatistics/secureresearchs ervice/accessthedatasecurely#assured-organisational-connectivity-aoc office for national statistics. (2024b). become an accredited researcher. retrieved from https://www.ons.gov.uk/aboutus/whatwedo/statistics/requestingstatistics/secureresearchs ervice/becomeanaccreditedresearcher office of population censuses and surveys. (1991). 1991 census england. united kingdom: office of population censuses and surveys. raab, g. (2013). e-datashield: e-mail data aggregation through anonymous summary statistics from harmonised individuals level databases. retrieved from edinburgh: raab, g., & diben, c. (2020). edatashield: running an analysis of combined data when the individual records cannot be combined. paper presented at the scottish centre for administrative data research, edinburgh. https://era.ed.ac.uk/handle/1842/36691 ralston, k., walsh, d., feng, z., dibben, c., mccartney, g., & o'reilly, d. (2017). do differences in religious affiliation explain high levels of excess mortality in the uk? journal of epidemiology and community health, 71(5), 493-498. doi: https://doi.org/10.1136/jech-2016-208176 ritchie, f. (2017). the "five safes": a framework for planning, designing and evaluating data access solutions. paper presented at the data for policy, london, uk. sacker, a., murray, e., maughan, b., & lacey, r. (2024). social care in childhood and adult outcomes. double whammy for minority children? longitudinal and life course studies, 15(2), 139-162. doi: https://doi.org/10.1332/17579597y2023d000000008 stats nz. (2022). experimental administrative population census: data sources, methods, and quality (second iteration). retrieved from wellington, new zealand: https://www.stats.govt.nz/research/experimental-administrative-population-census-datasources-methods-and-quality-seconditeration/#:~:text=this%20paper%20discusses%20the%20data%20and%20methods%20used ,provides%20information%20about%20the%20quality%20of%20the%20data. sturley, c., norman, p., morris, m., & downing, a. (2023). contrasting socio-economic influences on colorectal cancer incidence and survival in england and wales. social science and medicine, 333. doi: https://doi.org/10.1016/j.socscimed.2023.116138 19/20 sizer, alison, andreas mastrosavvas, & oliver duke-williams (2025). the ons longitudinal study – opportunities for longitudinal research on the england and wales population, iassist quarterly 49(3), pp. 1-20. doi: https://doi.org/10.29173/iq1159 uk statistics authority. (2020). research code of practice and accreditation criteria. section b: accreditation of researchers and peer reviewers. united kingdom: gov.uk retrieved from https://www.gov.uk/government/publications/digital-economy-act-2017-part-5-codes-ofpractice/research-code-of-practice-and-accreditation-criteria#section-b-accreditation-ofresearchers-and-peer-reviewers uk statistics authority. (2024). ethics self-assessment tool. london, united kingdom: uk statistics authority retrieved from https://uksa.statisticsauthority.gov.uk/the-authorityboard/committees/national-statisticians-advisory-committees-and-panels/nationalstatisticians-data-ethics-advisory-committee/ethics-self-assessment-tool/ university of st. andrews. (2021). safepod network. retrieved from https://safepodnetwork.ac.uk/ wolfson, m., wallace, s. e., masca, n., rowe, g., sheehan, n. a., ferretti, v., … burton, p. r. (2010). datashield: resolving a conflict in contemporary bioscience--performing a pooled analysis of individual-level data without sharing the data. international journal of epidemiology, 39(5), 1372-1382. doi: https://doi.org/10.1093/ije/dyq111 wright, d. m., rosato, m., raab, g., dibben, c., boyle, p., & o’reilly, d. (2017). does equality legislation reduce intergroup differences? religious affiliation, socio-economic status and mortality in scotland and northern ireland: a cohort study of 400,000 people. health & place, 45, 32-38. doi:https://doi.org/10.1016/j.healthplace.2017.02.009 endnotes 1 alison sizer is a celsius user support officer and associate researcher in the department of information studies, ucl. she can be reached by email: a.sizer.11@ucl.ac.uk. 2 andreas mastrosavvas, department of epidemiology and public health, university college london. he can be reached by email: a.mastrosavvas@ucl.ac.uk 3 oliver duke-williams is a senior advisor for celsius and a professor in the department of information studies, ucl. 4 ’do you have any long-term illness, health problem or handicap which limits the activities or work that you can do?’ 5 ’what is your religion?’ 6 ’do you look after, or give any help or support to family, friends, neighbours or others because of: long-term physical or mental ill-health or disability, or problems related to old age?’ 7 ’over the last twelve months would you say that your health has on the whole been: good; fairly good; not good?’ 8 ’how would you describe your national identity?’ 9 ‘what passports do you hold?’ 10 ‘if you were not born in the united kingdom, when did you most recently arrive to live here?’ 20/20 sizer, alison, andreas mastrosavvas, & oliver duke-williams (2025). the ons longitudinal study – opportunities for longitudinal research on the england and wales population, iassist quarterly 49(3), pp. 1-20. doi: https://doi.org/10.29173/iq1159 11 ‘including the time you have already spent here, how long do you intend to stay in the united kingdom?’ 12 the census was also held in northern ireland, but in scotland it was delayed by a year to 21st march 2022. 13 ’is the gender you identify with the same as your sex registered at birth?’ 14 ’which of the following best describes your sexual orientation?’, with tick boxes for straight/ heterosexual, gay or lesbian, bisexual, other sexual orientation. 15 ’who is (was) your legal marriage or registered civil partnership to?’, with tick boxes for someone of the opposite sex, or someone of the same sex. 16 ’have you previously served in the uk armed forces?’ 17 output areas are the lowest level of geographical area used for census statistics. they were created for the 2001 census and are made up of 40-250 households (100 – 625 individuals). 18 lower super output areas usually comprise 4 5 oas and 400 1,200 households (1,000 – 3,000 individuals). 19 middle super output areas usually comprise 4 – 5 lsoas and 2,000 – 6,000 households (5,000 – 15,000 individuals). 20 the “five safes” framework comprise a set of principles: safe data, safe projects, safe people, safe settings and safe outputs. 21 non-members are the household members who are living with an ls member at the time of the census. 22 a safepod is a standardised safe setting that provides the security and controls for data that requires secure access for research. a safepod includes a door control access system, cctv, a researcher area for dataset analysis, secure it cupboard and a height adjustable desk. vol252 iassist quarterly fall 2001 13 introduction i will be speaking today about an initiative currently underway in canada related to the preservation of data, specifically data generated by and useful for research in the social sciences and humanities. background i’d like to begin with some background information about my institution and its “relationship” with data over the years. the national archives of canada has existed since 1872 and has, compared to some other countries, an extremely wide collecting mandate which we have dubbed “total archives”. it means that we collect unpublished records from both public and private sector-sources in all media. it is the issue of publication which distinguishes our primary area of responsibility from that of the national library of canada, which collects published material. in the early 1970’s, an interest in “computer data” and its preservation developed within the institution and by 1973, a machine readable archives division had been created with a mandate to acquire research data from both public and private sector sources, in support of both the social and physical sciences. it became the de facto “national data archives” in canada. by the mid-1980’s, automation had begun to affect the creation of what archives considered their “traditional” records correspondence, memos, reports, case files, etc. in reacting to this situation, the national archives decided to”“integrate” the data archivists with the traditional archivists who would be most affected by the changes being wrought by automation, specifically those responsible for government paper, private paper, and cartographic records the intention was to cross-train everyone to do both paper and electronic records. for many reasons, including timing, the lack of human and financial resources, and other government priorities, the na’s role in, and commitment to data acquisition, preservation and access in canada slowly narrowed to focus on a very small number of government-generated databases, such as the census. immediate triggers two events have occurred in recent years to move the issue of data preservation and access back onto the government’s agenda. the first was a review of the national archives and national library of canada’s mandates, requested by heritage canada. dr. john english was appointed to investigate. among the interested groups who presented a submission during the hearings was the canadian association of public data users (capdu) represented by chuck humphrey, ernie boyko and wendy watkins, who are well-known to many members of iassist. dr. english’s most important recommendation, in this context at least, was his endorsement of capdu’s position that canada needed a national data management strategy: we endorse the canadian association of public data users proposal for a national data management strategy in which the national archives and the national library play a facilitative role. the two institutions should play a partnership role in such a data archive and coordinate the federal government’s relationship with such an archive. the second triggering event was a workshop co-sponsored by the social sciences and humanities research council (sshrc) and the organization for economic cooperation and development (oecd). in exploring canada’s social sciences infrastructure, participants highlighted the problems of access to canadian research data. this issue was of particular concern to sshrc, which funds a large proportion of academic research in canada. their grants enable a great deal of data collection and analysis and participants at the workshop emphasized the fact that the lack of a national data strategy has made the resulting data sets difficult to access, and has hindered canada’s ability to coordinate national developments and participate in international initiatives. joint investigation the impact of these two events led sshrc to contact the national archives of canada and propose a co-sponsored investigation. at this stage, it would be restricted to the a national research data management strategy for canada: the work of the national data archive consultation working group by yvette hackett* 14 iassist quarterly fall 2001 social sciences and humanities, with the hope that the work would attract the attention, and possibly, the participation of other “data” groups at a later date, including the natural sciences, health sciences, environmental sciences, etc. the national data archive consultation working group was formed over the summer of 2000. the 9-member group includes representation from a number of disciplines with experience in the creation and use of research data, such as political science, history and english. its membership also includes a representative from the archival community, luciana duranti, who some of you may know as the project director of interpares, an international research project investigating the authenticity of electronic records. the working group also includes one member from the data library community, chuck humphrey, and sue (gavrel) bryant, a former president of iassist, who had worked in the machine readable archives division at the national archives before moving on to a career in information management at treasury board, a central agency of the federal government. working group members john apsimon, chair special advisor to the president, carleton university gérard boismenu political science, université de montréal sue bryant treasury board, pki secretariat luciana duranti school of library, archival and information studies, university of british columbia chuck humphrey data library, university of alberta josé igartua département d’histoire, université du québec à montréal ian lancashire department of english, university of toronto michael murphy rogers communications centre, ryerson university matthew mendelsohn department of political studies, queen’s university in addition to the working group, a 14-member resource group was also invited to participate. the composition of the resource group is similar to that of the working group, adding subject expertise in sociology, geographic information systems, modern languages and new media. ernie boyko represents canada’s main statistical agency, statistics canada. wendy watkins provides additional representation from the data library community. chuck, ernie and wendy also bring extensive experience from both sides of canada’s data liberation initiative, an earlier and very successful project to improve research access to statistics canada’s data. my role, and that of my colleague from the national library of canada, focussed primarily on providing information relating to the mandates and current activities of our respective institutions. david moorman, a policy analyst at sshrc, coordinated the two groups’ activities. resource group members paul bernard dept. of sociology, université de montréal ernie boyko library and information centre, statistics canada martin brooks institute of information technology, national research council joseph desloges department of geography, university of toronto yvette hackett national archives of canada douglas hodges national library of canada terry kuny xist inc. timothy jackson ryerson university wanda noel barrister and solicitor frits pannekoek information resources, university of calgary michael ridley chief librarian, university of guelph geoffrey rockwell dept of modern languages, mcmaster university iassist quarterly fall 2001 15 fraser taylor department of geography, carleton university wendy watkins data centre, carleton university term of reference the terms of reference for the working group proposed a two-phase structure, with the second phase contingent on the results of the first phase. the focus of the investigation was distilled into a series of questions to be answered. the four phase one questions included: 1. to what extent is there a need for a unified and coordinated data archiving function? are modest changes to existing institutional policies and mechanisms adequate to meet current and future requirements? 2. what gaps exist in the mandates and structures of existing institutions in relation to management of research data? 3. who will benefit from the improved management of research data and to what degree? 4. how will effective research data management, preservation and access contribute to canadian research capacity? following a study of these“needs, gaps, benefits, and research capacity” questions, the working group would submit a preliminary report to sshrc and the national archives. the next steps would be determined both by the recommendation of the working group and the response of the sponsoring institutions. a working methodology rapidly evolved that depended equally on members of the working group and the resource group. the process began with a stakeholders’ meeting, held in ottawa in october 2000. fifty-five people attended, representing universities, federal government departments, research groups, academic associations, archives and libraries across the country. in the course of a day-long consultation, a wide range of problems with the current canadian situation were identified. the key ones included: difficulty in locating canadian data difficulty in gaining access to previously collected canadian data, due to costs and the lack of any kind of central resource directory, or depository service a weak tradition, within the canadian research community, of making data available for re-use or for replication studies a lack of “national data” leading to obstacles in canadian participation in multi-national studies a lack of a recognized national institution to facilitate canada’s participation in international research projects, in international associations and in standards development participants in the stakeholders’ meeting also heard of many initiatives currently underway, though most were addressing access issues only. they tended to be organized on an institutional, regional or disciplinary basis and were operating largely in isolation from each other. over 20 participants followed up with written submissions to the working group. the stakeholders meeting pointed out the need to clearly define the focus of the group’s investigations. the following “scope” statement was developed: a research data function would have the goal of preserving, managing and making publicly accessible digital information, structured through methodology and documentation, for the purpose of producing new knowledge. this function would address the gap that exists between the raw research materials and formally published results. acquisition would include digital information both produced by researchers and of interest to researchers. it emphasizes: the 3 facets of the required strategy access, management and preservation; the digital nature of the material, and the importance of a structured methodology; and the focus on the gap between raw data and published results. from a number of proposed research strategies, the working group focussed on three additional activities that they would undertake: the preparation of briefs outlining the “research data” situation in their particular areas of specialization; research to understand where “research data” fit into the mandates of existing institutions such as the national archives and the national library, as well as the role of the many university-based data libraries and archives. the organization of 4 surveys to elicit concrete data to support the opinions expressed at the stakeholders meeting; surveys were directed to sshrc-funded researchers; university data 16 iassist quarterly fall 2001 archivists; participating institutions in the data liberation initiative; and finally a list of general stakeholders who had attended the october meeting, or otherwise expressed their interest in this issue. the working group is scheduled to submit its phase one report in early june. the report confirms that a serious gap exists in canada’s research infrastructure and argues that the preservation of research data, and the facilitation of ongoing preservation, management and access to such data are important factors in building the research capacity so necessary to the growth of a “knowledge society”. as a result of these findings, the report will recommend that phase two be undertaken to study possible mechanisms to accomplish these goals. plans phase two the terms of reference for phase two have already been established, in anticipation of phase two proceeding. this time, the questions include: 1. is some form of national data archiving agency the right way to meet the needs of the research community? 2. are there alternative ways of meeting the needs of researchers? 3. if a new national facility is recommended, what functions should it perform and what institutional form should it take? 4. what is the most appropriate working relationship between a new facility and existing agencies such as the national archives of canada and the national library of canada? how should duplication of responsibilities and services be avoided? 5. how can a data preservation and access facility best take advantage of emerging information and communication technologies to increase its efficiency and effectiveness? a specific methodology for phase two has not yet been developed, but three obvious areas to pursue have emerged. the first would involve research into the various organizational models already in place in canada and abroad, and in specific disciplines. this process should include the identification of the strengths and weaknesses of each and an analysis of their applicability in the current canadian landscape. a second issue could address jurisdictional issues, as the needs of the private and public sectors are considered, including federal, provincial and municipal levels of government and the full spectrum of canada’s academic community, all of whom are potentially both creators and users of research data. a third issue would address the availability and appropriateness of various funding mechanisms. assuming a prompt and positive response from its sponsoring agencies, the working group hopes to complete work by december 2001. * paper presented at the iassist/ifdo conference 2001 in amsterdam. yvette hackett, national archives of canada, yhackett@archives.ca instructions for authors of the iassist quarterly 1/13 magnuson, diana l. (2024) stewarding our resources: building a sustainable ipums archival document access system, iassist quarterly 48(1), pp. 1-13. doi: https://doi.org/10.29173/iq1095 the creative commons-attribution-noncommercial license 4.0 international applies to all works published by iassist quarterly. authors will retain copyright of the work and full publishing rights. stewarding our resources: building a sustainable ipums archival document access system diana l. magnuson1 abstract ipums at the university of minnesota has created the world’s largest accessible database of census and survey microdata. the ipums suite of products contains nine harmonized data products. the largest of these products, ipums international (ipums-i), has supported the curation and preservation of ancillary materials received during data acquisition efforts. archival staff have preserved thousands of unique pieces of census and survey documentation, creating bibliographic records using an extended dublin core profile that supports the use of controlled vocabularies to enhance findability for the project staff and outside users. the goal of this curation work was to create a findable, searchable, and downloadable document access system for our internal use and to support ipums researchers. this paper describes our experience constructing a web interface that supports exploration and dissemination of these archived materials. during this development, we gained valuable insight about stewarding our resources that are applicable to research organizations responsible for curating, preserving, and disseminating archival materials. keywords archive, preservation, metadata, ipums introduction over the last thirty years, ipums at the university of minnesota has created the world’s largest accessible database of census and survey microdata (magnuson and ruggles 2022). the primary work of ipums is data harmonization—making census and survey data compatible across time and space. ipums integration and documentation makes it easy for researchers to study change, conduct comparative research, merge information across data types, and analyze individuals within family and community contexts. as of this writing, the ipums suite of products contains nine harmonized data collections. international data comes from over one hundred national and regional statistical organizations. all data are freely available to the global public. the context of this paper is specific to the ipums international (ipums-i) data project, the largest of the nine harmonized data collections. beginning in 1999, with a social science infrastructure grant from the national science foundation (nsf), ipums-i had a simple yet audaciously ambitious goal: preserve the world’s microdata resources and democratize access to those resources. twenty-five years later, the project goals continue to be: collecting and preserving census and survey data and documentation; harmonizing those data; and disseminating the harmonized data free of charge https://doi.org/10.29173/iq1095 https://creativecommons.org/licenses/by-nc/4.0/ 2/13 magnuson, diana l. (2024) stewarding our resources: building a sustainable ipums archival document access system, iassist quarterly 48(1), pp. 1-13. doi: https://doi.org/10.29173/iq1095 (ruggles and mccaa et al. 1999-2004; mccaa and ruggles 2000; ruggles et al. 2003). ipums-i data are coded and documented consistently across countries and over time to facilitate robust comparative research. ipums-i has amassed tens of thousands of ancillary materials in support of its data harmonization work. these materials came from united states census bureau (uscb), united nations statistical division (unsd), latin american and caribbean demographic center (celade), the east-west center, centre population et dévelopement (ceped), and over one hundred national statistical agencies. examples of this material include correspondence, maps, enumerator instructions, supervisor instructions, training materials, codebooks, publicity, reports, newspaper clippings, unpublished papers, census timetables, data processing materials, and technical manuals. the ancillary materials in our collection attest to the varied technical, business, social, and economic aspects of conducting censuses and surveys across time and space. acquisition, preservation, and dissemination of data products is central to the full ipums workflow (figure 1). for the purposes of this paper, archive-specific workflows are highlighted in blue. we used an open archival information system model (oais) to identify where our archive obtained submissions both external and internal (submission information package, sip), what actions we took once we obtained those submissions (archival information package, aip), and how we delivered the products to users (dissemination information package, dip) (magnuson and thomas 2023). figure 2 highlights the archive responsibility within the ipums workflow. this paper will focus on the development of the document access system within the archive workflow (figure 2, circled in red). figure 1. full ipums workflow (magnuson and thomas 2023) https://doi.org/10.29173/iq1095 3/13 magnuson, diana l. (2024) stewarding our resources: building a sustainable ipums archival document access system, iassist quarterly 48(1), pp. 1-13. doi: https://doi.org/10.29173/iq1095 figure 2. ipums archive workflow and document access system a portion of ipums-i grant money has funded the curation and preservation of the ancillary materials acquired by the project. for over two decades, archival staff have been preserving thousands of unique pieces of census and survey documentation, creating bibliographic records using an extended dublin core profile that supports the use of controlled vocabularies to enhance findability for the project staff and outside users. the goal of this work was the creation of a simple, findable, searchable, and downloadable document access system. an early iteration of the effort to make available basic international census documentation was the creation of a world population census forms webpage. the aim of the page was primarily to demonstrate to potential international data partners the scope and capacity of ipums-i. it provided a tabular list of census forms organized by country and year, with minimal functionality offered by hyperlinks to download the documents (ruggles et al. 2003).2 while useful for organizing, identifying, and accessing standard international census forms, this static utilitarian webpage challenged us to consider how we could provide access to the thousands of additional materials entrusted to us. we needed a simple but more sophisticated tool for stewarding our archival resources. building the ipums document access system building the ipums document access system required identifying shortand long-term goals of the archive, building a warehouse of flexible metadata, navigating the priorities and schedule of the institute for social research and data innovation (isrdi is ipums’ parent organization) and working https://doi.org/10.29173/iq1095 4/13 magnuson, diana l. (2024) stewarding our resources: building a sustainable ipums archival document access system, iassist quarterly 48(1), pp. 1-13. doi: https://doi.org/10.29173/iq1095 with the isrdi information technology (it) product team to develop a web application that was simple, sustainable, discoverable, and capable of disseminating basic metadata and pdfs of digitized archival materials. vision beginning in 1999, ancillary documents began streaming into the archive as a product of ipums-i acquisition efforts. it is likely that ipums-i uniquely holds some of these materials. a portion of the materials arrived with the express understanding that they would become available through some means after their lifecycle in the ipums-i microdata harmonization work had concluded. other acquired materials were broader than the ipums-i data collection and covered countries, censuses, and surveys for which ipums-i does not disseminate microdata. the fragility of some documents created a race against time to carefully preserve them for future use. for example, the high acid content of some paper documents meant they were brittle and at considerable risk of eventual disintegration. other materials were produced using thermofax technology (heat sensitive copy paper), which darkens over time, rendering the documents illegible. taken together, this collection reflects an exciting diversity of archival materials. we cannot predict the uses to which future researchers might put these materials, but it is our obligation to preserve them and make their innovative future use possible. data curator wendy thomas recognized the significance of all these materials and anticipated the day when an archival document access system would be a reality. thomas’ experience providing research support for social science data users and her technical expertise curating data and metadata positioned her well for envisioning an online ipums document access system. thomas immediately took steps to assemble the scaffolding necessary to build a flexible metadata warehouse. in this context, “flexible” refers to creating metadata in an open-source format that allowed for its mapping to a range of standards. from the beginning, thomas advocated for a simple online tool that would provide basic metadata and discoverable, accessible, searchable, and downloadable pdfs of digitized archival materials. the tool was envisioned to be sustainable long term by the archivist, with minimal intervention of it staff beyond the development of the web application (magnuson 2015). building a metadata warehouse building a metadata warehouse for the ipums archival document access system has been years in the making. first, thomas built up the scaffolding for logical intake and workflow for digital and manuscript materials (table 1). this was no small task. documents did not arrive on a fixed schedule or in uniform packaging, but rather, as agreements were made with supporting statistical entities (mccaa and ruggles 2000; ruggles et al. 2003). thousands of physical and digital documents needed to come under archival control before they were subject to the metadata creation stage. during this initial stage, thomas was balancing the creation and implementation of intake and workflow with a data harmonization project that was already underway. prior to thomas’ hire, ipumsi researchers had pulled a number of documents from their individual collections and scanned them for use in the ipums-i project. after thomas’ hire, archival staff were fielding requests from researchers to locate specific types of materials from within the newly acquired materials. thomas determined that the most expedient way to know the contents of the multilingual collection and to efficiently access the documents, was through scanning the cover or first page of all documents. https://doi.org/10.29173/iq1095 5/13 magnuson, diana l. (2024) stewarding our resources: building a sustainable ipums archival document access system, iassist quarterly 48(1), pp. 1-13. doi: https://doi.org/10.29173/iq1095 original boxes were sorted and labeled by region and country (using iso numeric codes) and then transferred into archival safe boxes. this approach facilitated an understanding of the scope of the collection by country, the percentage of documents needing special attention due to their physical condition and determining the extent of documents requiring language translation. thus, the shortterm needs of the ipums-i project followed a systematic, country-by-country approach to creating bibliographic records for this extensive collection. table 1. archival intake and workflow intake workflow acquisition acquire physical and digital documents from data partners accessioning e.g., unsd, uscb, celade, eastwest, ceped, other organization organize by source/region in file cabinets or boxes e.g., unsd, uscb, celade, eastwest, ceped, other label label box with standard [iso3166 region]-[iso3166-country]-[box ###] create country spreadsheet populate spreadsheet with item mpcacquisition number, folder, source, file name, region, country, notes scan follow scanning decision tree determine whether full, cover, or partial scan is warranted based on type, size, uniqueness, and/or fragility label scans follow mpc numbering taxonomy, e.g., 142-050-cb-09 or 142-048-001-09full bibliographic record creation extended dublin core profile utilize hand-tailored ipums-i controlled vocabulary next, thomas developed a process using a text editor for creating structured bibliographic records for each piece of archival material. using an extended dublin core profile, thomas hand-tailored a controlled vocabulary to the archival dimensions of ipums-i ancillary materials (table 2).3 the decision to use dublin core was a practical one. dublin core is recognized worldwide as a basic standard for creating bibliographic records. in addition, dublin core vocabulary is easily extended as needed to address new material types or issues of scale. cataloging systems that do not use dublin core invariably provide information on mapping to dublin core (thomas 2023).4 prior to the creation of the web application, archive staff used thomas’ controlled vocabulary to search for and organize materials, using text string searches employing simple programming scripts. table 2. ipums-international bibliographic record extended dublin core controlled vocabulary (excerpt) line no. text adds content 00 add document [includes lines: 01=title, 03=creator 11=mpcacq 12=fulluri 17=mpcloc 23=format[pdf] <mpcrecord> <dc:title xml:lang="en"></dc:title> <dc:creator></dc:creator> <dc:identifier xsi:type="mpcacq">mpc</dc:identifier> <dc:identifier xsi:type="fulluri"></dc:identifier> <dc:identifier xsi:type="mpcloc">box </dc:identifier> <dc:format>application/pdf</dc:format> https://doi.org/10.29173/iq1095 6/13 magnuson, diana l. (2024) stewarding our resources: building a sustainable ipums archival document access system, iassist quarterly 48(1), pp. 1-13. doi: https://doi.org/10.29173/iq1095 34=spatial 38=temporal 42=language 47=type 49=provenance] <dcterms:spatial xsi:type="iso3166_n"></dcterms:spatial> <dcterms:temporal xsi:type="census"></dcterms:temporal> <dc:language xsi:type="iso639-1">en</dc:language> <dc:type xsi:type="dcterms:dcmitype">physicalobject</dc:type> <dcterms:provenance>uscb</dcterms:provenance> </mpcrecord> 32 mpc subject <dc:subject>agricultural survey</dc:subject> <dc:subject>agricultural census</dc:subject> <dc:subject>economic census</dc:subject> <dc:subject>economic survey</dc:subject> <dc:subject>health survey</dc:subject> <dc:subject>housing census</dc:subject> <dc:subject>housing survey</dc:subject> <dc:subject>labor force survey</dc:subject> <dc:subject>population census</dc:subject> <dc:subject>population survey</dc:subject> <dc:subject>economic classifications</dc:subject> <dc:subject>educational classifications</dc:subject> <dc:subject>geographic classifications</dc:subject> <dc:subject>temporal classifications</dc:subject> 44 description <dc:description>data collection form</dc:description> <dc:description>non-census questionnaire</dc:description> <dc:description>process management form</dc:description> <dc:description>un questionnaire form</dc:description> <dc:description>data collection instructions</dc:description> <dc:description>supervisor instructions</dc:description> <dc:description>ipums-i inventory form</dc:description> <dc:description>geographic list</dc:description> <dc:description>from: ; to: </dc:description> <dc:description>[provider] documents inventory</dc:description> <dc:description>agenda</dc:description> <dc:description>budget</dc:description> <dc:description>census guide</dc:description> <dc:description>census methods</dc:description> <dc:description>census report</dc:description> <dc:description>census topics</dc:description> <dc:description>data processing</dc:description> <dc:description>evaluation</dc:description> <dc:description>hiring test</dc:description> <dc:description>information for schools</dc:description> <dc:description>meeting notes</dc:description> <dc:description>progress report</dc:description> <dc:description>un mission reports</dc:description> <dc:description></dc:description> 46 mpclimited types <dc:type xsi:type="mpclimited">code book</dc:type> <dc:type xsi:type="mpclimited">data</dc:type> <dc:type xsi:type="mpclimited">data collection form</dc:type> <dc:type xsi:type="mpclimited">data collection instructions</dc:type> <dc:type xsi:type="mpclimited">inventory</dc:type> <dc:type xsi:type="mpclimited">extract</dc:type> <dc:type xsi:type="mpclimited">form</dc:type> https://doi.org/10.29173/iq1095 7/13 magnuson, diana l. (2024) stewarding our resources: building a sustainable ipums archival document access system, iassist quarterly 48(1), pp. 1-13. doi: https://doi.org/10.29173/iq1095 <dc:type xsi:type="mpclimited">gov doc</dc:type> <dc:type xsi:type="mpclimited">image</dc:type> <dc:type xsi:type="mpclimited">instructions</dc:type> <dc:type xsi:type="mpclimited">correspondence</dc:type> <dc:type xsi:type="mpclimited">map</dc:type> <dc:type xsi:type="mpclimited">article</dc:type> <dc:type xsi:type="mpclimited">book</dc:type> <dc:type xsi:type="mpclimited">ephemera</dc:type> <dc:type xsi:type="mpclimited">publicity</dc:type> <dc:type xsi:type="mpclimited">regulations</dc:type> <dc:type xsi:type="mpclimited">serial</dc:type> <dc:type xsi:type="mpclimited">table</dc:type> <dc:type xsi:type="mpclimited">technical manual</dc:type> <dc:type xsi:type="mpclimited">training material</dc:type> <dc:type xsi:type="mpclimited">unpublished</dc:type> accessing archive metadata has consistently relied on use of controlled vocabularies to enhance findability and flexibility both for immediate use by project staff and for future external researchers. along the way, archivists developed and refined training and reference materials to support undergraduate student workers who created bibliographic records for each unique piece of archival material. an unintended but happy result of the development and refinement of these instructional and reference materials over time was the curation of institutional history documenting the work, growth, and development of the ipums archival processes. institutional priorities as noted above, census and survey data harmonization and dissemination are the central activities of ipums within isrdi at the university of minnesota. in this environment, it has been incumbent upon archival staff to nurture understanding and investment in expanding our institutional archival curation activities. ipums administrators, principal investigators, project managers, and research data scientists all recognize the importance and utility of an archival data access system, but quite naturally their priorities have historically focused on grant writing, data acquisition, data harmonization, and data dissemination. the key for archival planning and productivity, then, has been identifying ipums priorities that naturally connect with, and enhance, the shortand long-term goals of the archive. first, the bold mission of ipums is to democratize access to the world’s social and economic data for current and future generations.5 second, a “central goal” of ipums-i from the beginning was “to create an inventory of surviving census microdata and documentation,” including “enumerator instructions, census forms, codebooks, studies of data quality, and any other ancillary documentation we can locate for all countries that will allow us access to this documentation” (ruggles et al. 2003). third, the decision to begin assigning dois to ipums data products was triggered by the increasing requirements of external funding organizations to conform to “standard archival practice using the open archival information system (oais) model and digital object identifiers (doi)” (magnuson and thomas 2023). assigning dois forced broader internal discussions around access, curation, and preservation—all concerns of the archive. fourth, documenting these internal discussions became a steppingstone to building a successful core trust seal application, which in turn highlighted the preservation work of the archive (magnuson and thomas 2023).6 these priorities all have strong ties to ipums archive https://doi.org/10.29173/iq1095 8/13 magnuson, diana l. (2024) stewarding our resources: building a sustainable ipums archival document access system, iassist quarterly 48(1), pp. 1-13. doi: https://doi.org/10.29173/iq1095 concerns. intentionally communicating and nurturing those connections to internal ipums stakeholders has been essential to moving archival goals forward. the steady growth of the ipums-i data project, increasing expectations of external funding organizations regarding ipums preservation practices, and a transition in ipums archive leadership due to a retirement, eventually led to concrete steps toward building an online document access system. the ipums it product team, the group of developers responsible for the ipums web dissemination system and related components, added that task to their quarterly project calendar in late 2022 (ruggles et al. 2015; magnuson and thomas 2023). the persistent commitment to building a warehouse of flexible metadata positioned the archive well to take advantage of this opportunity to fulfill a long-standing grant deliverable. in addition, the it product team was looking for a discrete project to train a new hire and introduce them to project workflows. exploration of existing document access product solutions had occurred periodically, and the conclusion reached each time was that too much modification was necessary to leverage the use of our controlled vocabularies. now the evolutionary moment had finally arrived for ipums to undertake building an archival document access system. the it product team invited archival staff to propose a small defined project, and the timing was ideal. archive personnel had already submitted a presentation proposal to the 2023 iassist conference promising to document the effort to build a document access system. the acceptance in february 2023 for an iassist presentation in early june 2023 fit the ipums it product team timeline and meant the goal of the presentation would not merely be documentation but an actual demonstration of the new user interfaces. ipums it product team the process of working with the ipums it product team to develop a web interface was collaborative. the first formal step in the collaboration was completing an it project data sheet (pds) to record information from the archive covering four areas: 1) objective statement, project motivation, project context, primary stakeholder; 2) major milestones and deadlines, project scope; 3) minimum viable product (mvp), most important quality attributes, trade-off matrix; 4) known business issues and risks, and how success will be measured. working through the pds to its completion was a worthwhile intellectual exercise for archival stakeholders. it forced us to problematize the project in productive directions, identifying “must have,” “valuable to have,” and “aspirational” (future) developments to the system functionality and user interfaces. once the project was approved by the it product team and underway, there was regular written (email) and oral (zoom) dialogue clarifying project goals and developing a workflow. tasks were tracked via the web-based management tool trello. for five months leading up to the project “sprint” (final two weeks of the project), periodic check-ins to track progress and clarify action points were conducted over zoom every few weeks. archive stakeholders were first presented by the it product team software developers with proof of concept and shortly thereafter a prototype of the user interface. two weeks before the project deadline the it product team launched their sprint and communication increased to daily brief “standup” meetings. during the standup, it and archive contributors assessed and assigned tasks, asked questions, and held everyone accountable to the specifications of the pds. during the sprint period, the web developer created “wireframes” https://doi.org/10.29173/iq1095 9/13 magnuson, diana l. (2024) stewarding our resources: building a sustainable ipums archival document access system, iassist quarterly 48(1), pp. 1-13. doi: https://doi.org/10.29173/iq1095 (illustrations) of the proposed user interface content, functionalities, and intended behaviors for archival review and feedback (figure3). figure 3. example of ipums document collection wireframe the product team visualized a product workflow containing three dynamic data “warehouses” (figure 4). warehouse 1 contains pdfs of the scanned documents and the xml metadata working files created by undergraduate student workers. files in warehouse 1 are subject to archivist (“researcher”) validation against internal archive standards and cleaning protocols using perl scripts. after completing validation and cleaning the archivist moves the metadata to warehouse 2, where it is ready to be processed by the web application. the it product team then takes the pdf and xml metadata files from warehouse 2, subjecting these files to it computational validation and moves the files to warehouse 3, to be disseminated through the web interface. https://doi.org/10.29173/iq1095 10/13 magnuson, diana l. (2024) stewarding our resources: building a sustainable ipums archival document access system, iassist quarterly 48(1), pp. 1-13. doi: https://doi.org/10.29173/iq1095 figure 4. metadata workflow responsibilities for researcher and product team the collaboration between the it product team and archive personnel produced valuable learning in both directions. archive staff gained an appreciation and understanding of the elements necessary to build a scalable user interface: including consistency of content, reliable search filters, and clear expectations regarding the specific elements of the tool. in turn, the role and concerns of the archive are now in the it product team’s line of sight. this is especially important for the archive, as its work is often viewed as being conducted on the periphery of ipums product workflows. ipums document collection the ipums document collection web interface was the minimum viable product (mvp) resulting from the collaboration between archive staff and the ipums it product team. the launch of the ipums document collection (documents.ipums.org) from the development site to the live website in early june 2023 triggered a new phase in product development. pdf documents and metadata from the region of oceania served as a test dataset to represent the quality and content of the metadata records and to evaluate the ability of these records to support the functionality of the web interface. the oceania dataset is small (9 countries, 626 document records, 29 collection records, and 653 pdf files), but it is representative of the overall ipums-i collection. thus, the oceania dataset served as an essential diagnostic in the development of the web interface. first, the oceania dataset made clear for ipums it staff the nature of the files to be provided to them by the archive. second, archive staff divided responsibility for validation based on where in the workflow was most effective to assess the metadata. third, communication during the refinement of the mvp clarified record content in terms of immediate and future requirements that will support additional search options. going forward, the ipums document collection will be built out by archival staff as ipums-i ancillary materials from new regions are processed through warehouses 1, 2 and 3. https://doi.org/10.29173/iq1095 11/13 magnuson, diana l. (2024) stewarding our resources: building a sustainable ipums archival document access system, iassist quarterly 48(1), pp. 1-13. doi: https://doi.org/10.29173/iq1095 conclusion adopting a long-range strategy of acquisition, establishing archival control, and employing a process of flexible metadata creation, were essential to building and successfully launching the online ipums document collection. this effort took place over time, in expectation of an eventual funding and/or institutional opportunity. when the institutional opportunity presented itself, archival staff were ready. a strategy of anticipatory and flexible preparation can be adopted by archive staff who operate in an institutional setting in which their role supports the main product. as noted, the primary work of ipums is data harmonization, making census and survey data compatible across time and space. ipums integration and documentation makes it easy to study change, conduct comparative research, merge information across data types, and analyze individuals within family and community contexts. all ipums resources are free of charge. in this organizational context, the archive has a supporting but vital role. stewarding and expanding access to our archival resources in support of ipums involved: • articulating what areas of the organization our archival functions are related to. it was incumbent upon archive staff to clearly and consistently identify where the work of the archive touched the main product of the organization and how archival stewardship enhanced the main product. • knowing the priorities of our organization. framing archival goals within the context of ipums deliverables required archival staff to plan metadata creation in terms of how it would be efficiently conveyed to users of ipums data products. • investing in building flexible and scalable metadata processes and infrastructure. • leveraging work with the ipums it product team. through constructive and deliberate communication, archival staff developed shortand long-term goals, identified strengths and areas for refinement of archival data management, and educated ipums it staff on the vital role of the archive within the organization. • strategizing scalable deliverables. in the ipums-i context, we focused on developing the scaffolding for intake of digital and manuscript materials, tailoring a controlled vocabulary using an extended dublin core profile, and building a simple web interface. the ancillary materials in the ipums document collection attest to the technical, business, social, and economic aspects of conducting censuses and surveys across time and space. in addition, ancillary materials contextualize contemporary uses of, and responses to, historical censuses and surveys. expanding and deepening our search and delivery system will provide findability and accessibility to a rich set of supporting archival documentation that will illuminate census development and implementation processes across time and space. for example, access to materials documenting the development of enumeration forms and procedures over time supports researchers’ understanding of how statistical entities responded to the challenges of collecting demographic data on difficult to enumerate populations. creating a sustainable, discoverable, and searchable access system for a broad range of archival census and survey materials will support the ipums mission of democratizing access to the world’s social and economic data and support transformative scholarship. we believe curation and public availability of these materials enrich ipums products but also the innovative research of ipums data users. https://doi.org/10.29173/iq1095 12/13 magnuson, diana l. (2024) stewarding our resources: building a sustainable ipums archival document access system, iassist quarterly 48(1), pp. 1-13. doi: https://doi.org/10.29173/iq1095 while the ipums document collection and history is unique, it reflects a common situation for any archive that operates in a supporting role to a large research project. typically, research project managers focus on the dissemination of their primary product; the organization and management of ancillary materials is secondary to that effort. the role of the archivist is to ensure future researcher access to the rich documentary evidence amassed during any research project. this access may be a locally supported access system or deposited into an external system. at the beginning of the ipumsi project, we did not know who would be making the materials accessible, how they might be accessed, or what standards would be employed by this future system. for these reasons, we focused on organizing and creating a highly flexible structure for our metadata. the core of our system and that of any system operating in a similar environment is the need to accurately: • preserve the content in a manner that will support future research use. • track provenance of the materials, including distribution rights. • retain the logical relationship between the materials (based on archival needs). • select metadata and preservation standards that will ensure the greatest ease in depositing the content into one or more access systems. • provide as much descriptive information as possible in a manner that can be mapped to an existing or emerging access system and expanded as needed by the growth of the collection. developing shortand long-term strategies to organize and manage a collection within the constraints of current project funding is essential. in the ipums context, short-term efforts focused on: scanning cover pages so documents could be searched on a country-by-country basis; creating bibliographic records that formalized terms used by the project (mpclimited controlled vocabulary); and promoting the interests of funders by providing access to full documentation and provenance in some manner. the long-term ipums archival strategy included providing subsets of scans and bibliography records (in our context geographical regions) so that when the opportunity for developing an online access system arose, we were ready to populate it with a significant collection of materials that represented the range of materials in our collection. the payoff was an online collection with immediate usefulness to the global research community. these lessons are applicable across research organizations with unexploited or underexploited collections of valuable ancillary materials. https://doi.org/10.29173/iq1095 13/13 magnuson, diana l. (2024) stewarding our resources: building a sustainable ipums archival document access system, iassist quarterly 48(1), pp. 1-13. doi: https://doi.org/10.29173/iq1095 references magnuson, d.l. (2015) wendy thomas interview, university of minnesota, march 24, 2015. magnuson, d.l. and ruggles, s. (2022) “challenges of large-scale data processing in the 1990s: the ipums experience,” ieee annals of history and computing, pp. 71-83. https://ieeexplore.ieee.org/abstract/document/9972862 magnuson, d.l. and thomas, w.l. (2023) “expanding our perspective: building a sustainable metadata culture,” iassist quarterly, volume 47, no. 2. https://iassistquarterly.com/index.php/iassist/article/view/1046 mccaa, r. and ruggles, s. (2000) “ipums-international: a global project to preserve machinereadable census microdata and make them useable,” handbook of international historical microdata for population research, edited by patricia kelly hall, robert mccaa and gunnar thorvaldsen, pp. 335-346. https://international.ipums.org/international/microdata_handbook.shtml ruggles, s., king, m.l., levison, d., mccaa, r and sobek, m. (2003) “ipums international,” historical methods: a journal of quantitative and interdisciplinary history, volume 32, no. 2. https://www.tandfonline.com/doi/abs/10.1080/01615440309601215 ruggles, s., mccaa, r., sobek, m. and cleveland, l. (2015) “the ipums collaboration: integrating and disseminating the world’s population microdata,” journal of demographic economics, volume 81. https://doi.org/10.1017/dem.2014.6 ruggles, s., mccaa, r., levison, d., gardner, t. and sobek, m. (1999-2004) “international integrated microdata access system.” sbr9908380, methodology, measurement, and statistics program, nsf. thomas, w. (2023) “why dublin core,” general information, isrdi archive. endnotes 1 diana l. magnuson is curator and historian at the institute for social research and data innovation, university of minnesota. magnuson can be reached at magn0031@umn.edu. 2 https://international.ipums.org/international/census_forms_shtml 3 the use of “mpc” refers to the minnesota population center, which predated the institute for social research and data innovation (isrdi). 4 https://www.dublincore.org/specifications/dublin-core/profile-guidelines/ 5 https://www.ipums.org/mission-purpose 6 https://www.coretrustseal.org/ https://doi.org/10.29173/iq1095 https://ieeexplore.ieee.org/abstract/document/9972862 https://iassistquarterly.com/index.php/iassist/article/view/1046 https://international.ipums.org/international/microdata_handbook.shtml https://www.tandfonline.com/doi/abs/10.1080/01615440309601215 https://doi.org/10.1017/dem.2014.6 https://international.ipums.org/international/census_forms_shtml https://www.dublincore.org/specifications/dublin-core/profile-guidelines/ https://www.ipums.org/mission-purpose https://www.coretrustseal.org/ iassist quarterly iassist quarterly 2013 7 abstract aligning data and research infrastructure is, as our daily work often reminds us, a difficult process. while data professionals often focus on research lifecycles, incentives, storage and transmission technologies, metadata and data sharing we tend to overlook the epistemological incongruences of diverse research and data practices. all data creation processes, even if unknowingly, make assumptions about the world and what exists as a unique unit that can be analyzed. in attempting to make data meaningful to different audiences, especially across disciplines, we must pay attention to these epistemological assumptions. failure to do so will inevitably frustrate our attempts to develop meaningful infrastructure for research data and even potentially undermine effective research through misunderstandings of data. looking at census and zip code data as examples, this paper explores the issue of unit of analysis as an example of such disciplinary epistemological assumptions. the complexities that arise even in these simple examples suggest the importance of addressing the theoretical complexities of dealing with data collections, management and interpretation. keywords: data profession, categorization, harmonization, epistemology, data theory, critique. introduction there has been a growing interest in the humanities in adopting and developing digital methods from other disciplines, especially including techniques for collecting, describing and modeling large datasets. while there is much interesting work to be done in this regard, these interdisciplinary conversations will be most fruitful if the dialogue is carried out in both directions. by adopting an especially humanistic and critical perspective, this paper and kristin partlo’s, also published in this issue (see page 12: from data to the creation of meaning part ii: data librarian as translator), are thus an attempt to both historicize and theorize the work of describing and using data to generate new knowledge in the social sciences. in working with data it quickly becomes evident how difficult and problematic it can be to deal with multiple datasets or research questions that do not align with available data, but these difficulties are infrequently explored in a wider context. the difficult relationship between data and the world is often overlooked and not considered in a serious and abstract way. in day-today work with data the issues that arise tend to appear as accidental discrepancies rather than as endemic to the relationship between data and world. in this from data to the creation of meaning part 1: unit of analysis as epistemological problem by justin joque1 the difficult relationship between data and the world is often overlooked and not considered in a serious and abstract way. 8 iassist quarterly 2014 iassist quarterly light, this paper attempts to open the question of what exactly data measure and what relation that measurement has to the world, arguing that this relationship is necessarily problematic and difficult but still incredibly productive. ultimately, our attempts to describe the world through data are not processes of merely finding what is really there, but an active epistemological project of description and making the world knowable. all data creation processes, even if unknowingly, make assumptions about the world and what exists as a unique unit that can be analyzed and what ‘counts’. in attempting to make data meaningful to different audiences, especially across disciplines, data professionals must pay attention to these epistemological assumptions in their necessary diversity. with the possibilities and daily challenges of describing the world scientifically in mind, i would like, by tracing some of these problems historically as well as in relation to contemporary problems, to suggest that the difficulties we as data professionals face are a necessary difficulty of the relationship between data and the world that we will never be able to fully overcome. there have been many important contributions to confronting and describing these issues from epistemology to philosophy to science studies, including authors such as michel foucault and bruno latour.2 more closely related to our work with social science data, a number of authors working in information science have addressed some of these questions. ronald day’s (2014) recent work on the history of the documentary tradition, from early twentieth century information science to current work on big data, argues that the very act of describing of things—and the consequent use of this information as evidence—is a constructed, mediated and ideological process. his account of data and documentation in a historical context, suggests that all data and the description of information has always been mediated by technology and culture. furthermore, as geoffrey bowker (2014) has recently observed: big data appears to be able to efface categories in favor of temporary clusters of correlation (for example in the commercial world you no longer “need to know whether someone is male or female, queer or straight, you just need to know his or her patterns of purchases and find similar clusters” (bowker, 2014: 1796)). despite this, he argues social categories still produce real effects that need to be reckoned with and described. while algorithms may be able to bypass gender identity, the expression of gender in the real world creates undeniable effects. these categories cannot simply be ignored or replaced with uncategorized click-streams or dna sequences. with bowker and day’s arguments, it becomes clear that the categories through which data are identified, collected, aggregated and mediated are social and ideological constructs. despite this, these categories cannot simply be ignored. we must ultimately account for them and their effects, while also recognizing their constructed nature. all of the data work that social scientists, scientists, marketers and others do inevitability rests on a very complicated process of defining and recognizing categories of things in the world. the ideological work of harmonizing data the 2013 oecd global science forum report on new data for understanding the human condition, which provided the context for the most recent iassist meeting, makes explicit both the stakes and amount of work that are required to align infrastructure and data in such a way as to support data driven investigations: many of the research issues we face will require social scientists to work in close collaboration with scientific investigators in other disciplines, notably the biomedical and natural sciences, and across national boundaries. in part these changes arise from the need to adopt a multidisciplinary approach in our search for the causal mechanisms underlying and potential spread of communicable diseases, migration and human responses to climate change. but they also derive from a more ‘data driven’ approach to scientific investigation. the advances we have made in terms of our ability to generate, capture and re-use information on all aspects of human behaviour places us in a sea of data that has the potential to inform and inspire innovative approaches to scientific investigation. (oecd, 2013: preface) while the underlying idea of this project is desirable and of critical importance to the future of social science data, what i would like to attempt to articulate is precisely the ways in which this is an ideological project and what that means for such a project. i decidedly do not mean ideological in a negative way. slavoj žižek, a contemporary slovenian philosopher, has argued that to claim that one operates outsides of ideology is the ideological maneuver par excellence (žižek, 1989). for him there is no outside to ideology and in denying that one has an ideological agenda, one merely attempts to obfuscate and naturalize that ideology. thus, what i mean by ideology is merely how one relates to the world and what epistemological assumptions allow this relation to the world. so, the point is not the standard leftist ideology critique, but rather to attempt to open the question of what is ideologically and politically at stake in such projects. furthermore, what is at stake in such attempts to harmonize data across disciplines and countries is not merely the enactment of one universal political-ideology, but rather an attempt to deal with a whole spectrum of different and often times competing national, local and individual politicalideological frameworks. so, what ideological framework do our attempts to manage data, create repositories, create linked data, share data across disciplines, promote the reuse of data, etc. operate within? these questions can most clearly be answered by turning first to what is required to work with disparate datasets. as the oecd report states, to achieve this goal “data must be comparable across cultures, languages and environments. the concepts used to implement and communicate data at the international level should derive from universally recognized methods and standards” (oecd, 2013: 15). it is worth noting here the recommendation that one use universally recognized standards. while we will return to the notion of universally recognized standards, for the moment the important issue is that in order for data to be comparable and useful across datasets, it must of course be describing the same thing or at the very least a related phenomena. a dataset about zoning likely provides little additional information when combined with a dataset about astronomical objects. in our daily practice as data librarians, data producers and scientists, we are constantly confronted with the incommensurability of datasets and the difficulty of working with datasets that we wished were comparable. especially in working with spatial data, one often deals with things that do not coincide even if data users expect that they will. for instance, working with zip codes in the united states in a completely accurate way is a difficult, if not impossible, task. a large number of health and survey related datasets are anonymized and aggregated at the zip code level, since this address level data is often attached to the individual surveys. iassist quarterly 2014 9 iassist quarterly despite this common practice zip codes do not define areas. rather they are lists of addresses that often change between census years and for unpredictable reasons. while the census zip code tabulation areas (zctas) are often a workable stand-in for ‘zip code level demographics’, it is clear that computational comparison between zip code anonymized data and zctas is not necessarily a direct one-to-one relationship.3 what the available data describe is determined not by the pure desires of our knowledge nor by the requirements of our research questions, but rather by the accidental exigencies of the postal infrastructure. furthermore, similar discrepancies arise, especially when attempting to compare data internationally, not for infrastructural reasons but for more directly socio-political reasons. the canadian census definition of a family includes same sex couples whereas the united states definition does not include them, even in states where they are legally married (us census, 2012). while these discrepancies arise repeatedly in working with data, the most important point is that they are not merely contingent. despite un recommendations and other attempts to harmonize these types of data, a family is not at all a given unit. the canadian definition makes reference to the canberra report, which distinguishes the nuclear family from the economic family (statistics canada, 2013). this notion of an economic family suggests how politicized the naming of a unit of analysis potentially is. one could describe an entire marxist critique of the ‘economic family’, but for the time being that will have to be left aside. the point is that the family and questions of how it should function, what constitutes its boundaries, even how children should be socialized, etc. have been the site of political contestations for centuries if not millennia. so, one cannot simply state ‘what a family is’, without making at least some political-ideological assumptions about the world and how it functions or should function. thus, it is not only a problem of categorization and the delimitation of categories, but also the difficulty of delimiting the things that are counted themselves; it is a problem of defining the unit of analysis even before the problem of analysis or categorization. the question here is not what type of family is this, but what counts as part of a family. where, even before one may try to categorize a family, does the family itself begin and end? at least in the united states, the census’ primary aim is not a sociological one, but is purely political. it is designed to count the population for purposes of political representation and the demographic data that are produced are a secondary result. this political function has affected the way in which individuals are counted from the infamous 3/5th compromise to current debates about where prisoners should be counted. there is a concern now that with such large prison populations in the united states coming from inner-cities and being held in rural areas that there is now a noticeable process of exporting representation (and counts) from inner-cities to elsewhere. at first glance some of these issues may appear accidental or chosen for expediency, but the results they produce are contentious and thus carry political weight even when the initial decision may not have. moreover, as we have seen in the united states and to an even greater degree in canada the entire process of counting itself has become political (perhaps more accurately the always-political nature of counting has gained renewed political attention). replacing the census with a voluntary form has seriously called into question the validity and accuracy of the most recent national household survey. all of these political/ideological issues intervene to complicate any attempts at harmonization, crossnational studies and long-term comparisons. furthermore, the complications come not only from directly political questions but also as suggested earlier from a combination of political and infrastructural problems. i think it is worthwhile, in our work with social data, to think of these as two types of classification problems: on the one hand political decisions and on the other infrastructural problems, such as the need to change zip codes to make the postal system more efficient. of course these two categories overlap, but still they require different strategies to deal with them in terms of data harmonization. historical attempts at harmonizing the world ultimately, all of these complications are not merely technical data issues but rather directly political, ideological and also infrastructural problems. thus, the work required to overcome or at the very least deal with these issues is then not merely technical work but is political and epistemological work in its own right. it is here where it becomes apparent what ideological position underlies many of the attempts to internationally harmonize data. to suggest this in a larger context: in hoping to harmonize data across and between these political as well as infrastructural differences there exists a universalizing ideology that well predates our current work with data going back to linnaeus or likely even further to aristotle. attempts at international data harmonization can be seen as the most recent iteration of a long set of historical attempts to universally describe the world. while i lack both the space and the general expertise to trace these attempts at universal description historically in any sort of comprehensive manner, it is worthwhile to mention a couple examples from this history to at least begin to situate the type of work data professionals do with data in this larger tradition. charles godfray, chair of zoology at oxford university, summarized the relationship between linnaeus’ work and data well, saying “i like to think linnaeus faced the first bioinformatics crisis: the problem of organizing information about the increasing number of species that were being discovered in the eighteenth century, and he developed solutions using the best technologies available at the time” (paterlini, 2007: 814). linnaeus’ attempts to classify the living world provide an early example of these attempts to create a harmonized structure for defining things. while linnaeus’ work focused primarily on the biological, others have attempted even more comprehensive structures to describe the entire world that may more directly resemble modern attempts to describe in rigorous fashion all ‘social entities’. other philosophers, not to mention bibliographers, scientists, etc. at the time and since have attempted to make universal languages of description with the hope that all information could be knowable, retrievable and computable. more closely to our own time and work, the work of paul otlet, a belgian working on information science prior to world war ii, is indicative. as part of his major contributions to information science, he developed and advocated for what he called universal documentation. he claimed in a 1907 text, “through its collections and its various repertories [universal documentation] would truly become a ‘world memory.’ this would not be limited to recording facts, but would automatically and instantly permit their retrieval. it would be a vast intellectual mechanism designed to capture and condense scattered and diffuse information and then to distribute it everywhere it is needed” (otlet, 1990 [1907]: 110). as ronald day has begun to do, one could trace from paul otlet through 10 iassist quarterly 2014 iassist quarterly to current work on big data a fascinating variety of attempts to universally describe all knowledge in order to make it instantly retrievable (day, 2014). despite these lofty aims, these universal systems often fall flat. jorge luis borges, in a short essay on john wilkins, a natural philosopher who attempted, prior to linneaus, to create a universal language and classificatory scheme, offers a rather humorous and concise criticism of these attempts at the creation of universal systems of classification, citing: a certain chinese encyclopaedia entitled ‘celestial empire of benevolent knowledge’. in its remote pages it is written that the animals are divided into: (a) belonging to the emperor, (b) embalmed, (c) tame, (d) sucking pigs, (e) sirens, (f ) fabulous, (g) stray dogs, (h) included in the present classification, (i) frenzied, (j) innumerable, (k) drawn with a very fine camelhair brush, (l) et cetera, (m) having just broken the water pitcher, (n) that from a long way off look like flies. (borges, 1964[1952]: 103) borges’ taxonomy instantly suggests the difficulty and arbitrariness of any such classification system. as michel foucault says of borges’ taxonomy, “we apprehend in one great leap, the thing that, by means of the fable, is demonstrated as the exotic charm of another system of thought, is the limitation of our own, the stark impossibility of thinking that” (foucault, 1970: xv). in short what borges suggests, if one extrapolates slightly, is the impossibility of ever fully accounting for zip codes, families or other social categories in a comprehensive and totalizing manner. thus, what i hope to suggest is that our attempts to harmonize data and work internationally require coherent and agreed upon systems for naming and delimiting things; this is not a problem that is unique to the current epoch of ‘data.’ there is a long and fraught history of various universal attempts to name all things, often at least in the early years of such attempts under the belief that the abrahamic god created a well-organized and knowable world. while some of these attempts have had invaluable impact, especially those that have been limited to certain fields such as biological taxonomy, many in the light of history appear almost comical. furthermore, while none of these attempts have ever succeeded in creating a completely universal classification system or language, major breakthroughs were made in their pursuits. for instance, leibniz attempted to create a system for describing and calculating the answer to all questions, including philosophical inquiries, and in the process made a major breakthrough in binary calculation. likewise, while otlet’s more utopian dreams never came to fruition he made major contributions to information science and practice. the work of data returning to the question of ideology, it is now possible to point towards what is at stake in these attempts. all of these attempts to harmonize and create general descriptive languages are founded on a universalizing logic that in its most utopian dimensions believes that these political and infrastructural differences that stand in the way of classification are accidental and can ultimately be overcome. it is an ideology that believes that the world itself can successfully be homogenized and through erasing difference be made completely knowable. it is in short, in our time, a neo-liberal dream of flattening and connecting the entire globe. while most individuals working with social science data rarely, if ever, make such utopian claims, i think it is beneficial to consider the more totalizing historical antecedents to such work. those who work with data on a daily basis engage in a certain soft-utopianism that is entirely defensible and more often than not productive and beneficial, but placing that work in this larger historical context is helpful for thinking through the opportunities, challenges and risks of that work. thus, i do not at all mean to simply deride and criticize the very real and critical work that data professionals do to harmonize and integrate diverse datasets. rather, i hope to have suggested two things. first, that this sort of normalization of data across datasets, time and place is work. it is labor in a very real and measurable way. it is not simply a process of finding the ‘true name’ of things or the proper unit of analysis to delimit these things once and for all. it requires constant revision and integration of political change. second, i would wager, though i of course do not have the data to back up such claims, it is impossible to completely harmonize everything. additional types of data will be produced faster than anyone can ever deal with them, political differences and infrastructural exigencies will always intervene to guarantee that at the very least the texture and nuances of our data will be lost or at least smoothed away in combining and defining data that have been produced by diverse sources. we will never describe and know everything. instead we will always be engaged in a continuous process of discovery, rediscovery and translation between heterogeneous and at least partially incompatible places and times. the oecd report referenced above begins by commenting on how poorly anticipated the arab spring was and attributes this failure at least in part to the lack of data collected about new modes of communication. while of course it is a wholly worthwhile and valuable endeavor to know the world around us and the results of the arab spring are incredibly complicated, i must say that i am heartened by the fact that humanity and the world can still surprise us. acknowledgements i am grateful for the support and feedback from my colleagues at the university of michigan, especially nicole scholtz who is always willing to engage in conversations about the more theoretical implications of the work we do with data. most importantly, this paper never would have been possible without kristin partlo agreeing to explore these issues and present our initial findings together at iassist 2014. references borges l. (1964) other inquisitions 1937-1952. translated from the spanish by simms r. austin: university of texas press. (originally published in 1952) bowker, g. (2014) the theory/data thing. international journal of communication. [online] 8:1795–1799. available from: http://ijoc. org/index.php/ijoc/article/view/2190/1156 [accessed: 27 july 2014] day r. (2014). indexing it all: the modern documentary subsuming of the subject and its mediation of the real. in iconference 2014 proceedings 565–576. [online] doi:10.9776/14140. available from: http://hdl.handle.net/2142/47318 [accessed: 27 march 2014] foucault m. (1994 [1970]) the order of things: an archaeology of the human sciences. translated from the french. new york: random house. (originally published 1966) oecd (2013) new data for understanding the human condition: international perspectives. oecd global science forum report. [online] available from: http://www.oecd.org/sti/sci-tech/newdata-for-understanding-the-human-condition.htm [accessed: 2 june 2014] otlet p. (1990) “the systematic organization of documentation and the development of the international institute of bibliography” in: rayward w. (ed and trans) international organization and iassist quarterly 2014 11 iassist quarterly dissemination of knowledge: selected essays of paul otlet. amsterdam: elsevier. (originally published 1907) partlo k. (2014) from data to the creation of meaning part ii: data librarian as translator. iassist quarterly [online] 38(2). available from: http://iassistdata.org/iq/issue/38/2. [accessed 4 march 2015] paterlini m. (2007) there shall be order. embo reports. [online] 8(9): 814–816. available from: http://www.ncbi.nlm.nih.gov/pmc/articles/ pmc1973966/ [accessed: 3 june 2014] statistics canada (2013) economic family. [online] available from: http://www.statcan.gc.ca/concepts/definitions/fam-econ-eng.htm [accessed: 1 june 2014] united states census (2012) american community survey and puerto rico community survey 2012 subject definitions. [online] available from: http://www.census.gov/acs/www/downloads/data_ documentation/subjectdefinitions/2012_acssubjectdefinitions. pdf [accessed: 1 june 2014] žižek s. (1989) the sublime object of ideology. new york: verso. notes 1. justin joque is the visualization librarian at the university of michigan in ann arbor, michigan, usa. he can be reached by email at: joque@umich.edu. this paper was presented at the 2014 iassist conference in toronto, ontario, canada on 4 june, session 3j, along with its companion paper by kristin partlo, “from data to the creation of meaning part ii: data librarian as translator.” 2. for example: foucault m(1994 [1970]) the order of things. latour, b (1993). the pasteurization of france. harvard university press. 3. more information about zctas can be found here: https://www. census.gov/geo/reference/zctas.html elaborating a crosswalk between data documentation initiative (ddi) and encoded archival description (ead) for an emerging data archive service provider 1/24 peuch, benjamin (2018) elaborating a crosswalk between data documentation initiative (ddi) and encoded archival description (ead) for an emerging data archive service provider, iassist quarterly 42 (2), pp. 1-25. doi: https://doi.org/10.29173/iq924 elaborating a crosswalk between data documentation initiative (ddi) and encoded archival description (ead) for an emerging data archive service provider benjamin peuch1 abstract belgium has recently decided to integrate the consortium of european social science data archives (cessda). the social sciences data archive (soda) project aims at tackling the different challenges entailed by the setting up of a new research infrastructure in the form of a data archive. the soda project involves an archival institution, the state archives of belgium, which, like most other large archival repositories around the world, work with encoded archival description (ead) for managing their metadata. there exists at the state archives a large pipeline of programs and procedures that processes ead documents and channels their content through different applications, such as the online catalog of the institution. because there is a chance that the future belgian data archive will be part of the state archives and because ddi is the most widespread metadata standard in the social sciences as well as a requirement for joining cessda, the state archives have developed a ddi-to-ead crosswalk in order to re-use the state archives' infrastructure for the needs of the future belgian service provider. technical illustrations highlight the conceptual differences between ddi and ead and how these can be reconciled or escaped for the purpose of a data archive for the social sciences. keywords crosswalk, mapping, data documentation initiative (ddi), encoded archival description (ead), data archives, consortium of european social science data archives (cessda eric) introduction the consortium of european social science data archives (cessda) was created in 1976 (cessda eric, 2017). because it aims at bringing together researchers in social sciences in europe, it now represents one of the key international institutions in this field on the continent. as such, it also constitutes one of the main networks for ddi users since its first objective is to make exchange and re-use of social science research data possible (marker, 2013). fifteen european states ushered in this new organization in 1976, among which belgium (marker, 2013). even so, various factors of social and political nature led this country to leave the consortium prematurely and to fall behind in terms of research infrastructure development for the social sciences (schoups et al, 2008). today however, belgium has renewed its motivation to join cessda and become a full-fledged member in accordance with the organizational and technical requirements set out by the consortium. the setting up of a cessda-compliant data archive in belgium is the objective of the social sciences data archive (soda) project2. in this paper i would like to introduce a technical realization that was devised in the course of the ongoing soda project: a mapping between two xml-based metadata standards, ddi on the one hand https://doi.org/10.29173/iq924 2/24 peuch, benjamin (2018) elaborating a crosswalk between data documentation initiative (ddi) and encoded archival description (ead) for an emerging data archive service provider, iassist quarterly 42 (2), pp. 1-25. doi: https://doi.org/10.29173/iq924 and encoded archival description (ead) on the other hand. the idea of such a crosswalk stems from an institutional partnership between actors from the social science community in belgium and archivists working for the state. after i recount the context of the project, i will explain the rationale of the crosswalk and describe the materials and methods that were used to develop the mapping. i will then provide some technical illustrations for the obstacles that arose and the solutions that were proposed and conclude by outlining the next steps in the process. social scientists and archivists following in the footsteps of the other cessda service providers, the soda project seeks to bring together the belgian researchers in social sciences in order to promote sharing and re-use of research data. as such, this project involves representatives from the social science community, but also archivists. the latter often have academic backgrounds in history and historiography. as such, their discipline is more akin to the humanities than to the social sciences. still they are involved in the soda project for several reasons which will be explained hereafter. the soda project brings together three partners: 1) the vrije universiteit brussel3, a dutch-speaking university based in brussels; 2) the catholic university of louvain, the french-speaking university that used to manage the now-defunct belgian archives for social sciences (bass); and 3) the state archives of belgium, the institution responsible for preserving and giving access to administrative and historical records produced by belgian public offices, courts, notaries, as well as private organizations and influential families. the role of archives is to store documents that no longer serve a purpose for their original creators so that future historians may access them and use their contents to advance historical research. such documents come in many shapes: reports, minutes, letters, accounts, charters, drafts… and they now also come in both analog (paper) and digital form. the key organizing principle in archives is provenance: materials should be grouped according to their original source (e.g. a government department or a particular court of justice) in order to preserve, along with the materials themselves, the context of their creation (morris and rose, 2010). archivists document this context with finding aids, i.e. ‘description[s] of records that giv[e] the repository physical and intellectual control over the materials and that assis[t] users to gain access to and understand the materials’ (pearce-moses, 2005: 168). in other words, finding aids constitute the metadata of historiography. and just like the social science community procured an xml language for encoding its metadata in the form of ddi, historians devised the encoded archival description (ead) from that same ‘meta-markup language’ (van hooland and verborgh, 2014: 16, 28-43). the purpose of ead is to convert finding aids in digital form. like ddi, it was designed in the mid-1990s (pitti, 1997; 1999; dryden, 2010; rasmussen and blank, 2007). also like ddi, ead evolved over time and was further developed by its research community of origin so that it now comprises several versions (ddi alliance, 2017a; 2017b; stevenson, 2016). years ago, historians, genealogists and journalists had to go to the archives, sometimes at a heavy personal cost, before they could even be sure of the relevance of a repository’s collections to their work. nowadays, thanks to ead, archival institutions around the world can share their finding aids online, thus simplifying the first steps of historical research significantly. https://doi.org/10.29173/iq924 3/24 peuch, benjamin (2018) elaborating a crosswalk between data documentation initiative (ddi) and encoded archival description (ead) for an emerging data archive service provider, iassist quarterly 42 (2), pp. 1-25. doi: https://doi.org/10.29173/iq924 a crosswalk between two metadata standards there now exist, in large repositories like the state archives of belgium, whole infrastructures that were devised around ead so as to smoothen the process of writing finding aids, then converting them in machine-operable formats and finally making them accessible to users through open public access catalogs (opacs). because the current economic context deters the belgian stakeholders from investing much into scientific projects that do not involve the hard sciences, the actors of the soda project are working towards re-using the existing infrastructure and its underlying workflow at the state archives for the needs of the future belgian service provider. this workflow, or ‘pipeline’, comprises procedures, programs, tested and tried methods, and well-defined roles and tasks. if all or parts of it can be re-used for a social science data archive, then implementation will be quicker and costs will be reduced to a large extent. the gains from this joint venture would be twofold: first, it would support the setting up of the infrastructure for a brand new institution, thus advancing social science research in belgium; secondly, it would also benefit the field of historical research, as the state archives themselves could re-use social science data as well as make it accessible to its very own audience, thus promoting interdisciplinarity4. furthermore, the pilot study that is the soda project currently involves only two of the twelve belgian universities, but, as the future service provider will hopefully include all twelve in a cooperative endeavor along with the state archives, it will thus bring together the country’s two largest linguistic communities—the dutchand the french-speaking, upon whose governments the universities rely for funding—and the federal state, to which the state archives belong5. in this way, although the business plan of the future data archive is still in the making, the state archives will take on the tasks of archiving and disseminating the research data, while researchers and their universities will be the depositors and the end users of the subsequent new service. however, this entails bridging the gap between the aforementioned respective metadata standards: ddi, which is a cessda requirement (cmm group, 2016) on the one hand, and ead on the other. if ddi files can be converted into ead files in a fully or partially automated manner, then the tools and the skills of the state archives can be re-used for the needs of the future belgian service provider just like researchers will be able to re-use the data made available. the idea is not fully to replace ddi files with ead files, but to funnel the content of ddi files into the ‘pipeline’ of the state archives by means of ead even while retaining and making the ddi files accessible. precisely what actions and services will be performed with the ddi files on the one hand and the ead files on the other hand remain to be determined; yet the benefit of converting the former into the latter both for the future service provider and for the state archives is a given. hopefully this whole mapping enterprise comes off as heuristic, although it might also appear somewhat misguided. one of the reasons is the difficulty in envisioning the architecture of the future service provider, as much on the legal and institutional as on the technical and technological planes. tackling other questions such as the choice of a data management program or the crafting of a data transfer policy for future depositors will likely affect the crosswalk as well. however, in order to procure concrete deliverables that can be placed before the stakeholders, the decision was made to undertake the mapping. the conversion from ddi to ead can be done by means of a crosswalk, or ‘mapping’, from one element or ‘tag’ to another. a crosswalk can be defined as follows: ‘a crosswalk defines the semantic mappings of the fields of a source metadata schema to the fields of a target metadata schema, so as to https://doi.org/10.29173/iq924 4/24 peuch, benjamin (2018) elaborating a crosswalk between data documentation initiative (ddi) and encoded archival description (ead) for an emerging data archive service provider, iassist quarterly 42 (2), pp. 1-25. doi: https://doi.org/10.29173/iq924 semantically translate the description of sources encoded in different schemas. a crosswalk is expressed through a table that shows the equivalent metadata fields of the metadata schemas involved’ (gaitanou, bountouri and gergatsoulis, 2012: 264, emphasis in the original). the goal of crosswalks is often to allow ‘metadata created by one community to be used by another group that employs a different metadata standard’ (unesco, 2015: 53) as is one of the motivations of the soda project. preliminary searches aimed at uncovering existing crosswalks between ddi and ead proved unfruitful. the tag library of ead 2002 itself features crosswalks between ead and the dublin core as well as between ead and isad(g)6, the international standard for the description of archival materials (saa, 2002). however no crosswalks involving ead and ddi could be found, apart from a modest mapping between select elements from different schemas put out by the emory university libraries & information technology department (emory u, 2017). while thorough and well thought out, this crosswalk seems to proceed from the spirit of metadata standards such as the dublin core (dc), i.e. a more minimalist approach, one focused on the very essentials of metadata across fields and disciplines (doorn and tjalsma, 2007; méndez and van hooland, 2014). a metadata standard such as dc is meant to record the purely descriptive and administrative metadata, i.e. the information related to the identification, use, and management of the described objects (caplan, 2003). but archival description also heavily relies on technical, structural, and preservation metadata, which respectively specify the use requirements, the internal organization, and the storage conditions of the documents (oliver and harvey, 2016). the 2002 version of ead offers 146 elements for use (vs. 15 for dc) and though only a subset is mandatory, this number hints at the large and diverse range of information that may be encoded for the proper description of a group of archives. ddi and ead because ddi and ead are two complex, high-level markup languages applied to two different scientific disciplines, we at the state archives assumed at the outset of the project that a perfect one-to-one mapping between the two could not be achieved. however, as previously stated, the rationale of the crosswalk was not to do away with ddi, but to make the ddi-encoded metadata intelligible for the ead-minded workflow of the state archives. the future data archive will preserve the ddi files and use them whenever the already existing procedures cannot address the needs of the users7. little ought to be said about ddi here. however, it should be said that, for reasons that will be made clear further along, the version of ddi that was selected for the purpose of the mapping was ddicodebook because it was considered—with only a basic knowledge of the two ddi branches, codebook and lifecycle—the one most likely to be similar to ead. as for ead, as was previously stated, it now comes into three versions: ead 1.0, which was published in 1998 (pitti, 1999); ead 2002; and ead3, which was published in 2015 (stevenson, 2016). to this date the states archives of belgium as well as the archives portal europe (ape), which disseminates the metadata of archival repositories in europe, still use ead 2002. moreover, accounts of full-fledged implementations of ead3 or of transitions from ead 2002 to ead3 are still awaited. that is why the mapping was made only towards ead 2002 for now. it has been more than 15 years since the publication of ead 2002. contrary to ead3, the 2002 version came out at a time when yet few finding aids had been converted by archival repositories around the https://doi.org/10.29173/iq924 5/24 peuch, benjamin (2018) elaborating a crosswalk between data documentation initiative (ddi) and encoded archival description (ead) for an emerging data archive service provider, iassist quarterly 42 (2), pp. 1-25. doi: https://doi.org/10.29173/iq924 world: back then, archivists were still struggling to embrace the digital transition (allison-bunnell, 2016). today ead has truly become a standard, for it was adopted by most large archival institutions (francisco-revilla et al, 2014). that being said, it has drawn much criticism, even shortly after its publication. first, the complexity of the metadata language came under fire: as it turned out, the learning curve was very steep for such an expansive library of tags (yakel and kim, 2005). the fact that ead inherited the peculiar, much-criticized syntax of xml did not help8, nor did the fact that, while computer programmers themselves scoffed xml-based languages (pilgrim, 2009-2011), archivists and historians with varying, uneven levels of proficiency in computer science were bound to struggle with this new standard. it took some time for universities to include ead or at least xml training in their curricula destined to future archivists and librarians (fox, 2005), yet one can see now that ead is more and more frequently required by employers in the field (riggs, 2005)9. other critiques targeted the excessive flexibility of ead 2002: several authors pointed out that the lack of mandatory tags and of guidelines resulted in different archivists encoding the same kind of information in different places in ead files (shaw, 2001; francisco-revilla et al, 2014). for instance, luis francisco-revilla and his co-authors noted that such an essential element as ‘repository’ (<repository>), whose purpose is to record the name of the archival institution responsible for providing access to the described materials, can be left out of an ead 2002 document (franciscorevilla et al, 2014). however it is unfair, as many authors did, to heap on ead criticism that actually pertains to the deficiencies of the archival milieu at large. for instance, the subjective difficulty in appropriating ead was and still is largely due to the lack of training (or interest) from archivists into computer matters (shaw, 2001; yaco, 2008; dow, 2009). lack of funding for the purchase of computers, software, training sessions, and other such spendings for the implementation of ead cannot be seriously attributed to ead itself either: technology always comes at a cost. furthermore, as dennis meissner and his colleagues were brave enough to admit, attempts to convert finding aids into ead might incidentally reveal to some archivists the flaws in the finding aids themselves. meissner and his colleagues acted upon this by rewriting their finding aids, a task that proved strenuous and timeconsuming yet with great benefits (meissner, 1997). technological transitions, even the most successful ones, are always painful at any rate: as elizabeth dow reminds her readers: ‘[t]he archival community could take a lesson from the agonies librarians suffered as they converted card catalogs to online public access catalogs. by all accounts, the cleanliness of the data on the cards made a huge difference in the cost, efficiency, and sense of success of the project. the lesson: create clean data to start with’ (dow, 2009: 114). it is worth noting that much of the criticism leveled at ead echoes with the critiques of ddi: the latter was also said to be very complex and sometimes out of touch with the true needs of the social science researchers in terms of data documentation. similarly, problems that affect the community at large like the lack of proper tools or training, the need for organizational change, and a slow adoption rate hindered the initial widespread acceptance of ddi (wackerow and vardigan, 2013). nevertheless, just like ead within the archives, ddi has now made its way into the world of social sciences and is actively promoted through such vast networks as cessda or the standard's dedicated conferences, eddi and naddi. https://doi.org/10.29173/iq924 6/24 peuch, benjamin (2018) elaborating a crosswalk between data documentation initiative (ddi) and encoded archival description (ead) for an emerging data archive service provider, iassist quarterly 42 (2), pp. 1-25. doi: https://doi.org/10.29173/iq924 materials and methods ddi versions the choice of one of the two branches of ddi, codebook and lifecycle, was not an easy one, as it depended upon several factors of differing nature. first, there was the fact that the author does not have an extensive background in social sciences, so knowledge of ddi and of the current issues in metadata management in that same field had to be acquired through on-the-job training10. next, there was the fact that the soda partners knew little about the specific belgian context: how widespread are data documentation practices in the belgian social science academia? do researchers know about nesstar or dataverse? how much inclined would they be to share their datasets? a survey was designed to gather information on this topic, but the modest resources of the project and the tight schedule that they entailed forced the project members to move on and produce deliverables, sometimes in a counter-intuitive order. assuming, as we were eventually able to, that the data documentation practices are lacking in the belgian social science community, we saw in this situation an opportunity to instill good practices where there were few or none by offering to our target audience a consistent framework that had been thought out for their needs. however, such a framework had to be coherent with the state archives’ technical resources as well as its own interests in the matter. for one thing, like all archival repositories, the state archives have devised a particular ead template for their finding aids, selecting some of the tags from the library and leaving others out. knowing this, the question that arose was: which, of ddi-codebook and ddi-lifecycle, is more likely to parallel ead in general and the state archives’ own version more specifically? the idea arose that a codebook is more akin to a finding aid than a conceptual construct of the whole lifecycle of a dataset. as a discrete intellectual work that seeks to convey the meaning of codes who, without proper context, are unreadable, a codebook fulfills missions that are similar to that of finding aids, which provide researchers with the historical circumstances that determined the creation of records. just like a codebook will help researchers understand the role of variable ‘v29’, whose sole name is of no help in grasping its purpose, a finding aid will guide historians as they long to learn about, for example, the enigmatic usa board of tea appeals11. a valid critique of this stance is conveyed by the fact that, as finding aids serve to record the context of creation of documents, by documenting the custodial history and the conditions of acquisition of said documents by a repository before they were processed and arranged in a certain way by one or more archivists, one might say that finding aids in fact document the lifecycle of those records. this consideration regrettably occurred to the author only after the choice of ddi-codebook was made, and likely so for two reasons: first, to an outsider, ddi-lifecycle, both in its concept and in its peculiar xml rendition, can be very perplexing12, as it entails a good understanding of the actual lifecycle of social science research data nowadays as well as a very good command of xml; second, the idea of a codebook—both conceptually and, in more practical terms, as a literal ‘book of codes’ that one might hold in one’s hands—may prove more intuitive and relatable to other, more common forms of media. at first, only a high-level, almost superficial comparison of the general structure of ead 2002 on the one hand and of ddi-codebook and of -lifecycle on other hand could be made, working with peerreviewed publications as general guides and with the token ddi files and the ddi tag library (‘online https://doi.org/10.29173/iq924 7/24 peuch, benjamin (2018) elaborating a crosswalk between data documentation initiative (ddi) and encoded archival description (ead) for an emerging data archive service provider, iassist quarterly 42 (2), pp. 1-25. doi: https://doi.org/10.29173/iq924 field level documentation’) put out by the ddi alliance on its website (ddi alliance, 2017c). as will be shown when discussing the method behind the process, it seemed at this level of analysis that broad, general correspondences could be drawn between the key sections of ddi-codebook and ead, and while it arose along the mapping process that this claim ought to be mitigated, this further supported the choice of ddi-codebook. ddi 2 only builds upon the very first version of ddi (ddi alliance, 2017d), and because ddi 2.5, the latest codebook version, is backward compatible with ddi 2.1 (ddi alliance, 2017e), ddi 2.5 was chosen as the ‘source language’ for the crosswalk13. thus, a mapping that could process files following the rules of ddi 2.5 would also be able to handle files that draw on the ddi 1 specifications. a ddi corpus thanks to the hard work of the cessda service providers, many full-fledged data archives now exist and disseminate, especially via nesstar servers, datasets and metadata. the latter files are most often accessible to all, which is why actual instances of xml-ddi files could be gathered online and assembled into a corpus that enabled systematic search for specific occurrences of elements or attributes. in this way, it was possible to see them in context as well as to observe variation among the different uses and practices of metadata managers. for example, in cases where only one date was encoded in the <stdydscr> wrapper, suggesting that this was the date when the study as a whole had reached its conclusion, this date could be found either in proddate, in proddate att date, in distdate, in distdate att date, in depdate, in version, or in version att date. files produced in conformity with the syntax of ddi 2 were gathered from the online, open-access repositories of several data archives. the corpus was not put together in a thoroughly systematic fashion, by seeking files from all the european data archives for instance; rather, the governing criterion was the ease of retrieval of the metadata files. the ddi files were the following: title identifier data archive labour force survey, january 2017 [canada] lfs-71m0001-e-2017january abacus (canada) class and social structure of the population of czechoslovakia in 1984 module for individuals csda00111en csda (czechia) politbarometer short inquiry 2002 (accumulated dataset) za3851 gesis (germany) american national election study, 2004: preand postelection survey 4245 10.3886/icpsr04245.v2 icpsr (unites states of america) national travel survey, 2002-2015 5340 icpsr (unites states of america) euro-barometer 10 - october november, 1978 7728 icpsr (unites states of america) https://doi.org/10.29173/iq924 8/24 peuch, benjamin (2018) elaborating a crosswalk between data documentation initiative (ddi) and encoded archival description (ead) for an emerging data archive service provider, iassist quarterly 42 (2), pp. 1-25. doi: https://doi.org/10.29173/iq924 migrations between africa and europe – mafe senegal (2008) ie0216a ined (france) selected teagasc national farm survey data 2007 amended!teagasc! nfs!2007!portal!data issda (ireland) issp 2015: darbo orientacijos iv, 2015 m. spalis gruodis, 2 leidimas lida_issp_0296 _study_02 lida (lithuania) a smile is not enough, part i: developing an intervention for continuous positive feedback, 2013 nsd2045 nsd (norway) baromètre politique français 2006-2007 vague 1 fr.cdsp.ddi.bpf2007-r1 réseau quetelet (france) issp slovensko 2009-2010 sasd 2009001 sasd (slovaquia) institutional trust 2013 snd0963-001 snd (sweden) socioeconomic inequalities in thrace: living conditions, education and employment a1_en so.da.net (greece) general election study belgium 2003 electionsbelges2003b sohda (belgium)14 oecd main economic indicators databank, 1960-2017 4744 10.5257/oecd/mei/2013-04 ukda (united kingdom) multiscopo istat – time use 2002-2003 it.adpss-socio data.ddi.sn054 unidata (italy) table 1. list of ddi files that constitute the test corpus. several of these files only provide information for the description of the study and of the ddi document (<docdscr> and <stdydscr>); others feature variable-level description (<vardscr>). some tags or attributes could not be found, such as the sdatrefs, methrefs, and pubrefs attributes of such elements as <filedscr>, or the <controlledvocabused> element within <docdscr> and all of its subelements. yet the assumption was not made that those elements were rarely if ever used since the present corpus was created only as a tool for the purpose of the mapping and not as a representative sample of ddi files. software because a mapping consists in a correlation table, it seemed only natural to resort to microsoft excel so as to flesh out the correspondences in an orderly and systematic fashion. while the need for structure was high, the complexity of the ensuing model was low, which is why more versatile software like latex did not come under consideration. cross-files searches into the corpus of ddi files were performed with notepad++. https://doi.org/10.29173/iq924 9/24 peuch, benjamin (2018) elaborating a crosswalk between data documentation initiative (ddi) and encoded archival description (ead) for an emerging data archive service provider, iassist quarterly 42 (2), pp. 1-25. doi: https://doi.org/10.29173/iq924 method in order to present the method with which the mapping process was carried out, a word must be said about its two guiding principles. first, there was the ‘orientation’ of the mapping, i.e. determining which of ddi and ead would be the starting point of the mapping and which would be the receiving end. both configurations can be advocated for a number of reasons: on the one hand, ddi is the source material since this is all about the management of social science metadata; on the other hand, if the state archives are to fulfill several of the tasks of the future belgian service provider and if they are to do so by re-using their ead-based software infrastructure for the needs of the data archive, then ead must be given conceptual precedence. which standard ought to be adapted in order to transfer easily into the other proved to be a tricky question. eventually, the decision was made to begin with ddi because, since the author did not have an extensive knowledge of social sciences, it seemed more logical to embrace the whole of ddi and thus get better acquainted with this scientific field through its standard before looking for its conceptual equivalents in ead. right at the outset, the idea arose that, after the ddi-to-ead crosswalk would be finished, a ‘cross-mapping’ might have to be produced in turn, although more specifically from the state archives’ specific take on ead towards ddi. the second guiding principle consisted in the awareness of the opposition between syntax and semantics in computer science. depending on the context, precedence goes to one of these two concepts, which both originate from linguistics and which designate respectively ‘the way in which linguistic elements (such as words) are put together to form constituents (such as phrases or clauses)’ and ‘the historical and psychological study and the classification of changes in the signification of words or forms viewed as factors in linguistic development’ (merriam-webster, 2018). at the beginning of the project, the author had to acquire knowledge about the social sciences in general and about ddi in particular by himself, on the job. to this end, the insight of the university project partners, two belgian social scientists, proved most valuable. yet limitations came into light when the focus shifted to ddi: like most traditional historians and archivists do not know about ead, leaving such technical topics to computer scientists, it turned out that social scientists have often never heard about ddi, as they turn to academic librarians, the ict department in their institutions or graduate students for questions of documenting and archiving research materials. at best, researchers know about such programs as nesstar or dataverse, but the underlying technicalities such as the metadata standard in use remain in the shadows for them. and yet a metadata standard is certainly not the most ‘technical’ aspect of data management—i.e. neither the most complex, nor the most vital one—compared, for example, with the maintenance of the whole hardware infrastructure of a data archive. in most cases, the data encoded in ead or ddi are subject to only few checks and procedures, as opposed to other computer operations where the input of invalid data can lead to highly detrimental consequences such as the need to restart a lengthy and costly procedure or the corruption of data. let us imagine that a person in charge of documenting a dataset, instead of encoding the description of the study in stdyinfo/abstract, inserts the descriptive text in method/datacoll/timemeth. since <timemeth> allows for character data, this would not constitute a syntax error, although it would be a semantic one. while problematic, this mishap would probably not cripple the readability or usability of the dataset; at worst, the problem would be easily noticeable on the interface that displays the metadata of the dataset and could be quickly solved. compare this situation with that of a computer programmer who inadvertently removes a single https://doi.org/10.29173/iq924 10/24 peuch, benjamin (2018) elaborating a crosswalk between data documentation initiative (ddi) and encoded archival description (ead) for an emerging data archive service provider, iassist quarterly 42 (2), pp. 1-25. doi: https://doi.org/10.29173/iq924 comma from an algorithm upon which several automated processes rely. because of the unforgiving syntax rules of most programming languages, this would amount to a spanner thrown in the works, and it might be long before the cause of the ensuing problems is finally identified. apparently, the notion that the failure of the launch of the mariner 1 space shuttle was due to a misplaced comma in lieu of a period somewhere in the fortran computer code is inaccurate; still, part of the true cause of the problem was the (infinitesimal yet essential) lack of a ‘superscript bar’ in the transcript of a command for a smooth function (de florio, 2009: 32). that is to say, in effect, that contrary to operational computer functions where proper syntax is key, many elements in xml languages such as ead or ddi are but ‘meaningful titles’: they are primarily containers meant to store text, which will be parsed on a superficial level by various programs yet truly ‘processed’ only later on by human beings. like parsing programs verify the syntax of computer code, humans perform a similar operation when they process information, although on the semantic plane, where rules are far less strict and there is much more room for interpretation. one might say that the meaning attached to this or that element of an xml language is the ‘human factor’ of the standard: it is subject to interpretation, therefore to variation, and one must strive to apprehend its rationale so as not to adulterate its function. this was an important concern in the making of the mapping: while there was no doubt that many cases would prove to be problematic and require hard choices, much work was dedicated to grasping the original purpose of the ddi-codebook elements and attributes so as to reduce as much as possible the potential gap between them and ead’s own elements and attributes. results and discussion at the end of the mapping procedure, the starting hypothesis was successfully verified: because, in abstracto, ead is meant to document objects that require archiving, it could be bent for the needs of a data archive for the social sciences with appreciable pliability. like special characters that require particular encoding procedures in certain computer environments, some elements of ddi had to be ‘escaped’ in various ways in order to conform to the syntax of ead, yet the flexibility of ead allowed for such remedial reallocations. to illustrate, the following schemas show high-level templates for both metadata standards and the third figure presents a possible structural mapping: https://doi.org/10.29173/iq924 11/24 peuch, benjamin (2018) elaborating a crosswalk between data documentation initiative (ddi) and encoded archival description (ead) for an emerging data archive service provider, iassist quarterly 42 (2), pp. 1-25. doi: https://doi.org/10.29173/iq924 figure 1. high-level ddi-codebook template. figure 2. high-level ead template. <codebook> | <docdscr> | | information about the ddi file and the source codebook | </docdscr> | <stdydscr> | | information about the study and the dataset | </stdydscr> | <filedscr> | | information about individual files | </filedscr> | <datadscr> | | information about variables | </datadscr> | <othermat> | | other study-related materials | </othermat> </codebook> <ead> | <eadheader> | | bibliographic and descriptive information about the finding aid | </eadheader> | <archdesc> | | information about the content, context, and extent of the body of archival materials | | <dsc> | | | hierarchical groupings of the archival materials | | </dsc> | </archdesc> </ead> https://doi.org/10.29173/iq924 12/24 peuch, benjamin (2018) elaborating a crosswalk between data documentation initiative (ddi) and encoded archival description (ead) for an emerging data archive service provider, iassist quarterly 42 (2), pp. 1-25. doi: https://doi.org/10.29173/iq924 figure 3. possible structural mapping of ddi-codebook towards ead. figure 3 shows what could have been and what, in part, is in the crosswalk. mostly, the contents of the instances of <filedscr>, of <datadscr>, and the <othermat> section are conveniently transferred into ‘component’ (<c>) sections in ead with proper labeling, both to facilitate their identification should a technician need to look at the ‘source code’ in the archival institution (the ead file in this context) and to prepare the subsequent transfer of the formerly ddi-encoded information from the ead document to a more dynamic, user-friendly format—typically on a bibliographic record displayed through an opac. however the seemingly one-to-one relocation of the key <docdscr> and <stdydscr> sections into the ‘ead header’ (<eadheader>) and the ‘archival description’ (<archdesc>) sections respectively does not quite correspond to the eventual crosswalk. this is in spite of the fact that one could set down the following hypothetical analogy: ead ddi ‘ead header’ section <eadheader> ‘document description’ section <docdscr> concerns: the finding aid(s) covering the archival materials concerns: the ddi file and the source codebook ‘archival description’ section <archdesc> ‘study description’ section <stdydscr> <ead> | <eadheader> | | [docdscr] | </eadheader> | <archdesc> | | [stdydscr] | | <dsc> | | | <c> | | | | [filedscr] | | | </c> | | | <c> | | | | [datadscr] | | | </c> | | | <c> | | | | [othermat] | | | </c> | | </dsc> | </archdesc> </ead> https://doi.org/10.29173/iq924 13/24 peuch, benjamin (2018) elaborating a crosswalk between data documentation initiative (ddi) and encoded archival description (ead) for an emerging data archive service provider, iassist quarterly 42 (2), pp. 1-25. doi: https://doi.org/10.29173/iq924 concerns: the archival materials proper concerns: the study and the dataset table 2. a seemingly accurate parallel between the objects concerned by ead and ddi. at this level of analysis, these sections seem to mirror each other. the archival materials and the study results, on the one hand, constitute the ‘data’ of history and social sciences respectively, and, on the other hand, the finding aid and the ddi file (which explicitly corresponds to the codebook in this instance) are the ‘metadata’ proper. however, in depth-analysis soon reveals key structural differences between their respective encoding constraints. ‘document description’ (<docdscr>) and ‘study description’ (<stdydscr>) contain the most essential items of information concerning the dataset and its origins, specifically the descriptive and administrative metadata, such as the names of the people who partook in the study that resulted in the very dataset. such information could not simply be split as initially envisaged. a contributor who helped collect the data during the field work procedures should appear in the descriptive and administrative metadata along another contributor who, instead, encoded the information in the study-related codebook into a ddi file, even though their names and affiliations are not ‘stored’ in the same section in ddi. this reflects the choice of the developers of ddi, who decided to distribute the information into a potentially vast array of elements and subelements, e.g. the many possible instances of ‘bibliographic citation’. however, this was one of the instances where ead proved to work rather differently. an illustrative example to contrast these sections is the part played by, on the one hand, the large group of sub-elements contained in ‘bibliographic citation’ in ddi, and, on the other hand, the isolated ‘author’ (<author>) element in ead. authors and contributors ddi’s ‘bibliographic citation’ section allows for the use of a large group of elements that describe the various contributors and participants whether for the production of the dataset, the fulfillment of the study, the production of a work or the fulfillment of another study referenced in the metadata, etc. between the ‘producer’ (<producer>), the individuals or corporate bodies identified in the ‘version responsibility statement’ (<verresp>), those responsible for the various ‘notes and comments’ (<notes>, att resp), the ‘distributor’ (<distrbtr>), the ‘depositor’ (<depositr>), the ‘contact persons’ (<contact>), the ‘authoring entity’ or ‘primary investigator’ (<authenty>), and the ‘other identifications / acknowledgments’ (<othid>), the encoder is almost spoilt for choice when it comes to recording the names and roles of the various people who contributed in one way or another to the realization of the study and of its subsequent dataset. on the other hand, in the case of ead, there are quite fewer tags for encoding personor institutionrelated metadata, and they are fairly tightly confined to certain parts of the metadata document. essentially, the ‘author’ element records the ‘[n]ame(s) of institution(s) or individual(s) responsible for compiling the intellectual content of the finding aid’ (saa, 2002: 48) and, for other uses and requirements, the ‘name’ (<name>) element and its more specific variants, ‘family name’ (<famname>) and ‘corporate name’ (<corpname>), can be used. the latter three elements were designed for tagging significant names in the sections of the ead document that describe the archival materials per se, as opposed to ‘author’, whose purpose is to encode the name of the finding aid’s creator. for example, the names of certain public servants or historical figures may be tagged with https://doi.org/10.29173/iq924 14/24 peuch, benjamin (2018) elaborating a crosswalk between data documentation initiative (ddi) and encoded archival description (ead) for an emerging data archive service provider, iassist quarterly 42 (2), pp. 1-25. doi: https://doi.org/10.29173/iq924 <name> or <famname> so as to signal their presence in or relation to the archival materials and thus guide the researchers looking for archives that mention them. the following example shows a fictional instance of the ‘scope and content’ (<scopecontent>) section of an ead document, which is meant to ‘summarizing the range and topical coverage of the described materials’ (saa, 2002: 229), with illustrations of the use of the name-related tags: figure 4. a fictional example of a ‘scope and content’ section in a typical ead document. as illustrated, the various <name> elements in ead are tags meant to be used for labeling data and not metadata. this is why referencing the names of those who participated in fulfilling the study and producing the dataset—the people referenced in the ‘study description’ section—into the ‘archival description’ (<archdesc>) section in ead would inevitably break with the spirit of the archival standard. the problem might be circumvented by specifying the nature of the contribution of each participant through the att role of the ‘name’ (<name>) element. yet not only would an att type, which <name> is not endowed with, be preferable in this case; it would still be a hassle to separate, on the one hand, the names of those who partook in the creation of the source codebook and of the ensuing ddi file, and, on the other hand, the names of those who directly contributed to the realization of the study and the production of the related dataset15. while the question of authorship is an endless debate, especially in the case of complex creations such as motion pictures or social sciences datasets, the methods and precepts of archivists, which are fairly well translated by the structure of ead, advise for the concentration of all contributors in one large super-section, i.e. the ‘title statement’ (<titlestmt>) within ‘ead header’ (<eadheader>)16. the ‘author’ element in ead has few attributes that could help distinguish the various types of contributors. however, this can be easily remedied by automatically assigning meaningful identifiers to the instances of that element (all elements have a unique id attribute in ead). the resulting ‘title statement’ section in ead could then look as follows: <scopecontent> | | this collection contains records relating to the | <corpname>department of science and research | administration</corpname> when <name role=“secretary of | state”>ernest g. <famname>hunter</famname></name> was | secretary of cultural and scientific affairs between 1928 | and 1931. it includes a wide array of documents ranging | from minutes of internal and external meetings, reports on | various projects and matters, internal documentation | concerning the management of the department’s library as | well as administrative and particularly bookkeeping | correspondence. | </scopecontent> <titlestmt> | | <titleproper>employment survey</titleproper> | | <date>2017</date> | | <author id=doc_verresp_v1>jane doe</author> | <author id=doc_verresp_v2>jane doe</author> | <author id=doc_distrbtr>muhammad fayed</author> | <author id=stdy_authenty1>hwang seo-yun</author> | <author id=stdy_authenty2>oluwabusola adeyemi</author> | </titlestmt> https://doi.org/10.29173/iq924 15/24 peuch, benjamin (2018) elaborating a crosswalk between data documentation initiative (ddi) and encoded archival description (ead) for an emerging data archive service provider, iassist quarterly 42 (2), pp. 1-25. doi: https://doi.org/10.29173/iq924 figure 5. a fictional example of a ‘title statement’ section in a ddi-based ead document. to make sure that identifiers are, in conformity with the rules of ead, unique for each element in the file, an algorithm will have to generate identifiers with strict criteria: the identification of the provenance section; the retrieval of the role played by the contributor, whether as-is or according to a local controlled vocabulary; the inclusion of implicit information, i.e. information that is not originally contained within the mapped element itself, such as the version number as seen in figure 5; and, finally, when there are multiple individuals for one category, a disambiguation number, as in figure 5 in the case of the ‘authoring entities / primary investigators’. identifier overload another potentially problematic case is that of identifiers. here, ead shows practical flexibility, as each and every element can receive an id attribute, which must be unique. this allows for much granularity, which can be especially useful for complex archive collections. for example, a finding aid that describes a small group of documents might require only a few identifiers, such as a dedicated permanent identifier and a call number. on the other hand, more identifiers might be required for a finding aid that covers a very large and complex collection of archival materials. some of the most problematic cases include miscellaneous records from different archive producers, existing in various formats, some of which require specific preservation measures, which possibly entails scattering the materials across one or several repositories in the worst scenarios. in order to give both material and intellectual structure to such collections, archivists sometimes must compose very complex and hierarchical descriptions with many divisions and subdivisions, resulting in deep-reaching ‘trees’ of embedded ‘component’ (<c>) elements, as illustrated with figure 6: https://doi.org/10.29173/iq924 16/24 peuch, benjamin (2018) elaborating a crosswalk between data documentation initiative (ddi) and encoded archival description (ead) for an emerging data archive service provider, iassist quarterly 42 (2), pp. 1-25. doi: https://doi.org/10.29173/iq924 figure 6. a fictional example of a deep-reaching ead ‘tree’ of embedded ‘component’ (<c>) elements. in such cases, it can be especially helpful to allocate identifiers either to the deepest components, whether with the ‘id of the unit’ (<unitid>) element or with the id attribute of the ‘component’ elements, or to all echelons of the hierarchy, although the latter situation is rare since it requires much work of generation and management of identifiers. the problem that might arise when it comes to mapping ddi to ead would be the following: what is to be done with the identifiers already present in ddi? even if the future ead files are not meant fully to replace the ddi files, it would still be interesting to transfer the original ddi-encoded identifiers over to ead. yet the future data archive is likely to devise a policy for generating and managing identifiers of its own. would it be possible to retain the original identifiers on the one hand while enriching the ead documents with a set of new identifiers? it could be interesting to present the ‘structure’ of the dataset by listing the different files that make it up into ead’s ‘description of subordinate components’ (<dsc>) section. but where exactly should the identifying information transfer? into <unitid>, and/or in the id attribute of that same element, and/or in that of ‘title of the unit’ (<unittitle>), and/or in that of ‘descriptive identification’ (<did>), or perhaps in the id attribute of the highest element, <c>? further, there could be at least three identifiers to juggle with: the name of the file, the identifier of the file (allocated to it in the ddi-codebook file), and the data archive’s own identifier for each file. <archdesc> | <dsc> | | <head>hierarchical groupings of the materials</head> | | <c> | | | <did> | | | | <unittitle>i. irrigation and reclamation committee</unittitle> | | | </did> | | | <c> | | | | <did> | | | | | <unittitle>a. general information department</unittitle> | | | | </did> | | | | <c> | | | | | <did> | | | | | | <unittitle>1. correspondence</unittitle> | | | | | </did> | | | | | <c> | | | | | <did> | | | | | | <unittitle>a. public agencies and branches</unittitle> | | | | | </did> [...] | </dsc> </archdesc> https://doi.org/10.29173/iq924 17/24 peuch, benjamin (2018) elaborating a crosswalk between data documentation initiative (ddi) and encoded archival description (ead) for an emerging data archive service provider, iassist quarterly 42 (2), pp. 1-25. doi: https://doi.org/10.29173/iq924 this is one of those situations where ruling out technical questions shows that such questions bear larger organizational implications. the threat of ‘identifier overload’17 forces information systems engineers and decision-makers to determine what roles exactly the ead files will fulfill on the one hand and what other purposes the ddi files will serve on the other. in a way, it also broaches upon the question of the scale of the future institution: will it be only a branch of the state archives, which will simply draw upon the existing practices for identifier allocation, or will it be a larger entity with needs such as that of a specific policy for identifier generation and management? perhaps an exploratory study of the technical implications of the soda project was required before this question could come to light, but it must now be answered before a definitive technical solution can be implemented. distributing the information the fact that ddi-codebook and ead have their own inner logic transpires in how certain elements mapped well yet required that the various items of information sometimes take very different paths. for instance, ddi’s <holdings> (‘holdings information’), with its four specific attributes, ‘location’, ‘callno’, ‘uri’ and ‘media’, contains valuable information from the standpoint of ead. however, it is distributed as follows: ddi ead <holdings> ‘holdings information’ archdesc/custodhist or archdesc/repository/corpname (+ <address>) <holdings> att location <holdings> att callno archdesc/dsc/c/did/container or archdesc/dsc/c/did/unitid (att id) <holdings> att uri eadheader/eadid att urn or url <holdings> att media archdesc/phystech/genreform table 3. distribution of the information contained in <holdings> into different ead endpoints. as illustrated, the information that can be encoded within one single element along with its attributes in ddi has to be reallocated in possibly five different elements and/or attributes within all three main wrappers of ead, <eadheader>, <archdesc> and <dsc>. the element <holdings> itself or its location attribute seem to be where the name of the institution responsible for giving access to the dataset is most often encoded. if it gives the name of the ‘home institution’, i.e. the university from which the researchers behind the dataset hail, then its content should be transferred into the ead section ‘custodial history’ (<custodhist>); if, on the other hand, it already records the name of the data archive to which it is destined, then its content should be transferred to the key ead element <repository>, possibly followed by an <address> wrapper composed of <addressline> elements in case a postal address was also encoded. the call number (att callno) is transferred either to the ‘container’ (<container>) element in ead, or to the <unitid> element, or to that element’s id attribute. the ead 2002 tag library advises one of the latter two choices, acknowledging however that different practices https://doi.org/10.29173/iq924 18/24 peuch, benjamin (2018) elaborating a crosswalk between data documentation initiative (ddi) and encoded archival description (ead) for an emerging data archive service provider, iassist quarterly 42 (2), pp. 1-25. doi: https://doi.org/10.29173/iq924 coexist (saa, 2002: 80). next, the content of the uri attribute of <holdings> is copied over to either att urn or att url of <eadid>. finally, ead’s ‘genre / physical characteristics’ (<genreform>) element, which is meant to record ‘the types of material being described, by naming the style or technique of their intellectual content (genre); order of information or object function (form); and physical characteristics’ (saa, 2002: 150) is the perfect equivalent of <holdings>’ media attribute. we ought to point out that the abovementioned problem of ‘identifier overload’ might occur with the uri attribute of <holdings>. if the attribute contains a uri that is not a url hyperlink, then the question of identifier redundancy arises once more. but if, as was often observed in the ddi-codebook corpus, the attribute simply contains a url hyperlink that leads to the web page of the home institution of the researchers who authored the study, then we need only transfer said hyperlink to ‘custodial history’ and make it dynamic with the href attribute of the ‘external reference’ (<extref>) element in ead. while the various paths that each item of information must follow during the transfer might appear sprawling overall, the information nevertheless finds conveniently close equivalents in terms of xml elements according to the logic both of the source language, ddi, and the language of destination, ead. however, such scattering of the source information shows that the rationale of the mapping is a one-way transfer of information. it is unlikely that ddi files can be reconstructed from the ead files simply by reverting the ‘orientation’ of the mapping, hence the need to preserve the ddi files in spite of their conversion towards ead. conclusion and future work the potential inclusion of the state archives of belgium into the future data archive for the social sciences will require the harnessing of the existing infrastructures for the needs of said data archive. because the state archives are a cultural heritage institution, one of their missions is to preserve and communicate the context of the data stored in their collections, a task that they share with data archives. in order to perform this mission for the future belgian cessda service provider, a crosswalk between two metadata standards was undertaken and completed. at this stage the mapping of elements and attributes only exists, so to speak, in vitro: it has been laid out in a spreadsheet with much documentation as well as technical recommendations in a dedicated column of an xlsx file; but it has yet to be integrated into a program and then a series of processes— the state archives’ ‘pipeline’—that will ultimately produce the desired ead document. the following tasks will be undertaken in order to make the most of the crosswalk: • first, the mapping will need to be adapted to the template ead document laid out by the state archives, thus providing an example of adaptation to local needs and therefore a proper use case; • second, encoded archival context (eac) will have to be included if possible, so as to make the transfer of data about the scientific contributors even more effective; https://doi.org/10.29173/iq924 19/24 peuch, benjamin (2018) elaborating a crosswalk between data documentation initiative (ddi) and encoded archival description (ead) for an emerging data archive service provider, iassist quarterly 42 (2), pp. 1-25. doi: https://doi.org/10.29173/iq924 • thirdly, extensive testing will have to be performed with real instances of ddi files to make sure that the transfer is feasible and that the resulting ead documents conform to the rules of the standard as well as those of the state archives; • finally the mapping will have to fit into the pipeline of the state archives, along with secondary algorithms such as the one meant to identify (i.e. automatically generate identifiers for) the various elements that designate contributors to the original study. the spreadsheet that contains the mapping will likely be published, hopefully in open access, at some later point. additional research is required to determine what are the best formats (csv, tsv, pdf, xsd…) for making the crosswalk readable and readily re-usable by other users. this will benefit the publication of the subsequent adaptation to the state archives’ ead template, which will probably also take the form of a mapping. the rationale of such a crosswalk does not rest solely in the re-use of computer infrastructure, nor does it merely consist in reducing expenditures. it also demonstrates the validity of the archivist in the organization chart of a data archive. in several cessda service providers, the focus was visibly put on academic qualifications in social sciences for the constitution of the workforces. this is quite logical, considering the data archives federated by cessda permit the re-use of research data from the social sciences, by social scientists, and for social scientists. still, archivists seldom appear on the ‘who’s who’ webpages of cessda service providers. perhaps a ddi-ead crosswalk might inspire some curators and show the way towards fruitful collaborations. archivists are, after all, professional data and metadata managers, and they are used to handling highly heterogeneous archival objects. by working in close collaboration with the scientists from whom they receive data and for whom they will preserve and disseminate datasets, archivists can contribute to scientific research beyond their ‘organic’ attributions. although the soda project officially began several years ago, the involvement of the state archives is fairly recent, which is why many key questions still remain unanswered at this point: the exact business plan, cost model, and the roles that the state archives will take on in the new service provider; the legal form of the future entity; the types of contracts and/or agreements that will bind the data purveyors (the universities and their researchers) to the data archive. moreover, the transfer of (meta)data can only be done with actual (meta)data in the first place. the next grand step in terms of technical and organizational developments will entail gaining the trust of data providers and supplying the adequate tools and guidelines for documenting research data to them. acknowledgements the social sciences data archive (soda) project is funded by the belgian federal public planning service science policy office (belspo) under contract #fr/00/so3. i would like to express my thanks to my manager, dr rolande depoortere, whose careful and trustful supervision made it possible for me to move about and explore different possibilities before setting one course of action. i also wish to thank my colleague, samira hajji, whose hard work helped better envision the legal form of the future belgian service provider. my thanks go to the reviewers, who took the time to read my contribution and to formulate tactful suggestions for improvement. i am also grateful to mari kleemola and the cessda metadata management (cmm) working group, who https://doi.org/10.29173/iq924 20/24 peuch, benjamin (2018) elaborating a crosswalk between data documentation initiative (ddi) and encoded archival description (ead) for an emerging data archive service provider, iassist quarterly 42 (2), pp. 1-25. doi: https://doi.org/10.29173/iq924 showed kind interest into my work and who provided some valuable feedback, especially irena vipavc brvar and wolfgang zenk-möltgen. finally, i wish to thank everyone at the eddi 17 conference who attended and gave feedback after my presentation, especially mari kleemola, franck cotton, and joachim wackerow, as well as the organizing committee of the swiss centre of expertise in the social sciences (fors) at large. references akçeşme, banu, baktir, hasan and steele, eugene (eds.) (2016) interdisciplinarity, multidisciplinarity and transdisciplinarity in humanities, newcastle upon tyne, cambridge scholars. allison-bunnell, jodi (2016) ‘review of encoded archival description tag library – version ead3’, journal of western archives, vol. 7, no. 1, article 6. boydens, isabelle (2011) ‘strategic issues relating to data quality for e-government: learning from an approach adopted in belgium’. in: assar, saïd, boughzala, imed and boydens, isabelle (eds.) practical studies in e-government: best practices from around the world, new york, springer, pp. 113-130, https://doi.org/10.1007/978-1-4419-7533-1_7. caplan, priscilla (2003) metadata fundamentals for all librarians, chicago, american library association. cessda eric [consortium of european social science data archives european research infrastructure consortium] (2017) cessda eric – history, [online], available: https://www.cessda.eu/about/history [18 may 2018]. cmm group [cessda metadata management working group] (2016) cessda service providers’ metadata practices: standards, controlled vocabularies and requirements for the cessda portfolio. cessda metadata management project combined deliverable d1 & d2, [online], available: https://www.cessda.eu/content/download/834/7776/file/cmm_serviceprovidersmetadatapracti ces_2016.pdf [18 may 2018]. ddi alliance (2017a) history of the standard, [online], available: https://www.ddialliance.org/what/history.html [18 may 2018]. ddi alliance (2017b) ddi specification, [online], available: http://www.ddialliance.org/specification/ [18 may 2018]. ddi alliance (2017c) explore documentation, [online], available: https://www.ddialliance.org/explore-documentation [18 may 2018]. ddi alliance (2017d) ddi codebook 2.1, [online], available: http://www.ddialliance.org/specification/ddi-codebook/2.1/ [18 may 2018]. ddi alliance (2017e) ddi-codebook, [online], available: http://www.ddialliance.org/specification/ddi-codebook/ [18 may 2018]. de florio, vincenzo (2009) application-layer fault-tolerance protocols, hershey (pa), information science reference. deschouwer, kris (2005) ‘kingdom of belgium’. in: kincaid, john and tarr, g. alan (eds.) constitutional origins, structure, and change in federal countries: volume i. a global dialogue on federalism, montreal, mcgill-queen’s up, pp. 48-75. https://doi.org/10.29173/iq924 https://doi.org/10.1007/978-1-4419-7533-1_7 https://www.cessda.eu/about/history https://www.cessda.eu/content/download/834/7776/file/cmm_serviceprovidersmetadatapractices_2016.pdf https://www.cessda.eu/content/download/834/7776/file/cmm_serviceprovidersmetadatapractices_2016.pdf https://www.ddialliance.org/what/history.html http://www.ddialliance.org/specification/ https://www.ddialliance.org/explore-documentation http://www.ddialliance.org/specification/ddi-codebook/2.1/ http://www.ddialliance.org/specification/ddi-codebook/ 21/24 peuch, benjamin (2018) elaborating a crosswalk between data documentation initiative (ddi) and encoded archival description (ead) for an emerging data archive service provider, iassist quarterly 42 (2), pp. 1-25. doi: https://doi.org/10.29173/iq924 doorn, peter and tjalsma, heiko (2007) ‘introduction: archiving research data’, archival science, vol. 7, no. 1, pp. 1-20, https://doi.org/10.1007/s10502-007-9054-6. dow, elizabeth h. (2009) ‘encoded archival description as a halfway technology’, journal of archival organization, vol. 7, no. 3, pp. 108-115, https://doi.org/10.1080/15332740903117701. dryden, john. (2010) ‘a structure standard for archival context: eac-cpf is here’, journal of archival organization, vol. 8, no. 2, pp. 160-163, https://doi.org/10.1080/15332748.2010.513325. emory u [emory university, libraries & information technology] (2017) crosswalk of core metadata: emory core metadata mapping to selected standards and systems, [online], available: http://metadata.emory.edu/guidelines/descriptive/crosswalk.html [18 may 2018]. fox, michael (2005) ‘professional training for encoded archival description in europe’, journal of archival organization, vol. 3, nos. 2-3, pp. 71-82, https://doi.org/10.1300/j201v03n02_06. francisco-revilla, luis, trace, ciaran b., haoyang, li and buchanan, sarah a. (2014) ‘encoded archival description: data quality and analysis’, proceedings of the american society for information science and technology, vol. 51, no. 1, pp. 1-10, https://doi.org/10.1002/meet.2014.14505101043. gaitanou, panorea, bountouri, lina and gergatsoulis, manolis (2012) ‘automatic generation of crosswalks through cidoc crm’, metadata and semantics research: proceedings of the 6th research conference, mtsr 2012, berlin, springer, pp. 264-276, https://doi.org/10.1007/978-3642-35233-1_26. huddleston, rob (2008) xml: your visual blueprint™ for building expert web sites with xml, css, xhtml, and xslt, hoboken (nj), wiley. ica [international council on archives] (2000) z695.2.i83 2000 isad(g): general international standard archival description, 2nd edition, ottawa, international council on archives. lawson, gary (2004) federal administrative law, 3rd edition, st. paul (mn), west academic. marker, hans jørgen (2013) ‘data documentation, access, and dissemination systems’. in: kleiner, brian, renschler, isabelle, wernli, boris, farago, peter and joye, dominique (eds.) understanding research infrastructures in the social sciences, zurich, seismo, pp. 39-46. meissner, dennis (1997) ‘first things first: reengineering finding aids for implementation of ead’, the american archivist, vol. 60, no. 4, pp. 372-387, https://doi.org/10.17723/aarc.60.4.6405275227647220. méndez, eva and van hooland, seth (2014) ‘metadata typology and metadata uses’. in: sicilia, miguel-angel (ed.) handbook of metadata, semantics and ontologies, singapore, world scientific, pp. 9-40, https://doi.org/10.1142/9789812836304_0002. merriam-webster (2018) dictionary, [online], availaible: https://www.merriam-webster.com/ [18 may 2018]. morris, sammie l. and rose, shirley k. (2010) ‘invisible hands: recognizing archivists’ work to make records accessible’. in: ramsey, alexis e., sharer, wendy b., l’eplattenier, barbara and mastrangelo, lisa s. (eds.) working in the archives: practical research methods for rhetoric and composition, carbondale, southern illinois up, pp. 51-78. oliver, gillian and harvey, ross (2016) digital curation, 2nd edition, [no place], american library association. pearce-moses, richard (2005) a glossary of archival and records terminology, chicago (il), the society of american archivists. https://doi.org/10.29173/iq924 https://doi.org/10.1007/s10502-007-9054-6 https://doi.org/10.1080/15332740903117701 https://doi.org/10.1080/15332748.2010.513325 http://metadata.emory.edu/guidelines/descriptive/crosswalk.html https://doi.org/10.1300/j201v03n02_06 https://doi.org/10.1002/meet.2014.14505101043 https://doi.org/10.1007/978-3-642-35233-1_26 https://doi.org/10.1007/978-3-642-35233-1_26 https://doi.org/10.17723/aarc.60.4.6405275227647220 https://doi.org/10.1142/9789812836304_0002 https://www.merriam-webster.com/ 22/24 peuch, benjamin (2018) elaborating a crosswalk between data documentation initiative (ddi) and encoded archival description (ead) for an emerging data archive service provider, iassist quarterly 42 (2), pp. 1-25. doi: https://doi.org/10.29173/iq924 pilgrim, mark (2009-2011) everything you know about xhtml is wrong, [online], available: http://diveintohtml5.info/past.html [18 may 2018]. pitti, daniel v. (1997) ‘encoded archival description: the development of an encoding standard for archival finding aids’, the american archivist, vol. 60, no. 3, pp. 268-283, https://doi.org/10.17723/aarc.60.3.f5102tt644q123lx. pitti, daniel v. (1999) ‘encoded archival description: an introduction and overview’, d-lib, vol. 5, no. 11, https://doi.org/10.1080/13614579909516936. rasmussen, karsten boye and blank, grant (2007) ‘the data documentation initiative: a preservation standard for research’, archival science, vol. 7, no. 1, pp. 55-71, https://doi.org/10.1007/s10502006-9036-0. raţă, georgeta, arslan, hasan, runcan, patricia-luciana and akdemir, ali (eds.) (2014) interdisciplinary perspectives on social sciences, newcastle upon tyne, cambridge scholars. riggs, michelle (2005) ‘the correlation of archival education and job requirements since the advent of encoded archival description’, journal of archival organization, vol. 3, no. 1, pp. 61-79, https://doi.org/10.1300/j201v03n01_06. saa [society of american archivists] (2002) encoded archival description tag library, chicago (il), the society of american archivists. schoups, inge, lobet-maris, claire, laurent, véronique, poullet, yves and lefever, nathalie (2008) soda prospect study: feasibility study of a computerized archive service for the social sciences (soda) [studie soda: haalbaarheid van een data-archief voor de sociale wetenschappen / étude prospect soda : faisabilité d’un service d’archivage de données pour les sciences sociales], expertisecentrum david, cellule interdisciplinaire de technology assessment fundp and centre de recherche informatique et droit fundp, report without a number. shaw, elizabeth j. (2001) ‘rethinking ead: balancing flexibility and interoperability’, new review of information networking, vol. 7, no. 1, pp. 117-131, https://doi.org/10.1080/13614570109516972. stevenson, jane (2016) ‘encoded archival description tag library, version ead3’, archives and records, vol. 37, no. 2, pp. 257-260, https://doi.org/10.1080/23257962.2016.1220362. unesco (2015) interoperability and retrieval, paris, unesco, [online], available: http://unesdoc.unesco.org/images/0023/002321/232199e.pdf [18 may 2018]. van hooland, seth and verborgh, ruben (2014) linked data for libraries, archives and museums: how to clean, link and publish your metadata, london (uk), facet. wackerow, joachim (2017) current status of ddi 4. [presentation] 9th annual european ddi user conference (eddi17), swiss centre of expertise in the social sciences (fors), 6th december. wackerow, joachim and vardigan, mary (2013) ‘an established international metadata standard: the data documentation initiative (ddi)’. in: kleiner, brian, renschler, isabelle, wernli, boris, farago, peter and joye, dominique (eds.) understanding research infrastructures in the social sciences, zurich, seismo, pp. 158-167. yaco, sonia (2008) ‘it’s complicated: barriers to ead implementation’, the american archivist, vol. 71, no. 2, pp. 456-475, https://doi.org/10.17723/aarc.71.2.678t26623402p552. yakel, elizabeth and kim, jihyun (2005) ‘adoption and diffusion of encoded archival description’, journal of the american society for information science and technology, vol. 56, no. 13, pp. 14271437, https://doi.org/10.1002/asi.20236. https://doi.org/10.29173/iq924 http://diveintohtml5.info/past.html https://doi.org/10.17723/aarc.60.3.f5102tt644q123lx https://doi.org/10.1080/13614579909516936 https://doi.org/10.1007/s10502-006-9036-0 https://doi.org/10.1007/s10502-006-9036-0 https://doi.org/10.1300/j201v03n01_06 https://doi.org/10.1080/13614570109516972 https://doi.org/10.1080/23257962.2016.1220362 http://unesdoc.unesco.org/images/0023/002321/232199e.pdf https://doi.org/10.17723/aarc.71.2.678t26623402p552 https://doi.org/10.1002/asi.20236 23/24 peuch, benjamin (2018) elaborating a crosswalk between data documentation initiative (ddi) and encoded archival description (ead) for an emerging data archive service provider, iassist quarterly 42 (2), pp. 1-25. doi: https://doi.org/10.29173/iq924 end-notes 1 benjamin peuch is a contractual researcher at the state archives of belgium. a literature and information science graduate, he has been working on the social sciences data archive (soda) project since april 2017 and can be reached by email at benjamin.peuch@arch.be or benjamin.peuch@gmail.com. 2 the project was previously known as the ‘social sciences and humanities data archive (sohda) project’ and has been recently renamed. 3 the name of this university literally means ‘free university of brussels’. however, it is preferable not to translate it because there is another university in brussels whose name also literally translates like so: the french-speaking université libre de bruxelles. 4 on interdisciplinarity, see for example raţă et al, 2014 and akçeşme, baktır and steele, 2016. 5 for more information on the institutional landscape of belgium, see deschouwer, 2005. 6 ‘isad(g)’ stands for ‘general international standard archival description’. the standard was conceived by the international council on archives and published in 1988 (ica, 2000). 7 this is likely to be determined to a large extent by the kind of data management software, such as nesstar or dataverse, that will be chosen for the belgian data archive. a review has yet to be conducted before the choice is made. 8 ‘a common criticism of the w3c schema language is that it is too complex and tries to do too many things’ (huddleston, 2008: 52). 9 as a form of anecdotal evidence, in 2016, knowledge of ead was required for the current job position of the author at the state archives of belgium. 10 ‘metadata demands knowledge. obviously knowledge about the data and the research being carried out is needed, but the further cost is that it demands knowledge of metadata description, the structure and content of metadata and some technical insight regarding the practical arrangement of the metadata’ (rasmussen and blank, 2007: 60). 11 the board of tea appeals, abolished in 1996, was tasked with ‘adjudicat[ing] the claims of tea importers whose products were denied entry into the united states by federal tea-tasters’ (lawson, 2004: 7). https://doi.org/10.29173/iq924 mailto:benjamin.peuch@arch.be mailto:benjamin.peuch@gmail.com 24/24 peuch, benjamin (2018) elaborating a crosswalk between data documentation initiative (ddi) and encoded archival description (ead) for an emerging data archive service provider, iassist quarterly 42 (2), pp. 1-25. doi: https://doi.org/10.29173/iq924 12 a very bad pun to make about the core structure of ddi-lifecycle would be to say that it is exceptionally ‘fragmented’. 13 moreover, the ddi alliance announced at the 9th eddi conference that they would continue to support ddi 2 and ddi 3 even with the imminent release of ddi 4 (wackerow, 2017). 14 this file was produced as a small test case several years before the state archives became actively involved in the so(h)da project. 15 furthermore, experience shows that, oftentimes, in the few cases when information about the various contributors is encoded into the ‘document description’ section, it turns out that the researchers themselves encoded their own codebooks. 16 this is not entirely true when it comes to ead 2002 in particular: in this version, the identity and the roles played by the people who contributed to the creation of the ead file (as opposed to the usually paper-based finding aid) are to be encoded in the ‘creation’ (<creation>) section under ‘ead header’ (<eadheader>), and <author> may not be used in <creation>. this however was solved in ead 3, where <creation> was replaced by the ‘maintenance event’ (<maintenanceevent>) section, which allows for the use of a personal tag, ‘agent’ (<agent>). 17 on the problem of identifier overload, see boydens, 2011: 122-123. https://doi.org/10.29173/iq924 1/17 förster, andré; borschewski, kerrin; bolton, sharon; jääskeläinen, taina (2020) the matter of meta in research data management: introducing the cessda metadata office project, iassist quarterly 44(3), pp.1-17. doi: https://doi.org/10.29173/iq970 the matter of meta in research data management: introducing the cessda metadata office project andré förster1, kerrin borschewski2, sharon bolton3 and taina jääskeläinen4 abstract accompanying the growing importance of research data management, the provision and maintenance of metadata – understood as data about (research) data – have obtained a key role in contextualizing, understanding, and preserving research data. acknowledging the importance of metadata in the social sciences, the consortium of european social science data archives (cessda) started the metadata office project in 2019. this project report presents the various activities of the metadata office (mdo). metadata models, schemas, controlled vocabularies and thesauri are covered, including the mdo’s collaboration with the ddi alliance on multilingual translations of ddi vocabularies for cessda service providers. the report also summarizes the communication, training and advice provided by mdo, including ddi use across cessda, illustrates the impact of the project for the social sciences and research data management community, and offers an outline regarding future plans of the project. keywords research data management (rdm), metadata, controlled vocabularies, multilingualism, consortium of european social science data archives european research infrastructure consortium (cessdaeric) 1. introduction with the constant digitization of research and the ever-increasing relevance of open science, research data management (rdm) has become one of the most important indicators when assessing the quality of research. ‘good research data management is not a goal in itself, but rather the key conduit leading to knowledge discovery and innovation, and to subsequent data and knowledge integration and reuse’ (european commission 2016, p. 3). thus, it is a substantial part in the research (data) lifecycle, which is, for instance, reflected by the highly regarded fair principles for scientific data management and stewardship (wilkinson et al. 2016). according to the fair principles, guidelines concerning the documentation of research data and infrastructures for their long-term preservation should ensure that research data are findable, accessible, interoperable and reusable. these principles have become the guiding principles when handling research data, for research in general and for the social sciences in particular. accompanying these developments, research on rdm has become more important over the last years, although it is still considered to be at an early stage. for instance, some scholars investigate how information infrastructures at universities and research institutions have to be designed in order to optimally incorporate rdm processes (blask and förster 2019). other scholars focus on how to encourage researchers to perform rdm (van den eynden and bishop 2014), since for many researchers, the disadvantages of doing rdm still dominate. they believe that doing rdm and fostering open data means having additional work without personal gain. they fear the misuse of their https://doi.org/10.29173/iq970 2/17 förster, andré; borschewski, kerrin; bolton, sharon; jääskeläinen, taina (2020) the matter of meta in research data management: introducing the cessda metadata office project, iassist quarterly 44(3), pp.1-17. doi: https://doi.org/10.29173/iq970 data and that other researchers could beat them in publishing relevant findings their data provide, without integrating or citing the data producers. this means they would get less credit for their work, especially since publications are generally more appraised than the production of valuable data (wolf 2017). the most active field of research and services on rdm, however, is still the technical one. this is – among other things – due to the fact that many funding agencies have labelled rdm as means to fulfil specific mandatory funding requirements such as long-term preservation (deutsche forschungsgemeinschaft/german research foundation dfg 2013). since the provision and maintenance of metadata – understood as data about (research) data – is cardinal to understand research data, metadata also play an important part in ensuring technical rdm requirements. metadata describe publications and digital objects, make sure that research data can be contextualized, contribute to the implementation of the fair principles and thus help in following specific funding requirements (gregory et al. 2009). whereas in other sciences metadata are often not standardized according to reporting standards and information about the use of metadata schemas is not provided in the respective repositories, the quantitative social sciences have established the data documentation initiative (ddi)5 standard as the main standard of metadata documentation (vardigan 2014). the ddi standard is a free and international metadata standard for the description of research data from the social, behavioral, economic, and health sciences. it allows for a detailed and semantically rich documentation. ddi is the most commonly used metadata standard for social sciences survey data on an international scale. as the ddi documentation is captured in extensible markup language (xml), it is machine-actionable and fosters interoperability, metadata exchange and reuse of data. most of the social science data archives produce their metadata in the ddi format. currently, two different ddi specifications (each having different versions) exist – ddi codebook and ddi lifecycle. the codebook specification is focused on the after-the-factdocumentation. it includes information on document description, study description, variable description and file description. ddi lifecycle is the more elaborate of both specifications. it covers all aspects of the research data lifecycle (defined by the ddi alliance), starting with the planning and data collection right through to archiving (green and humphrey 2014; hoyle et al. 2011; rasmussen 2014; vardigan et al. 2008; zenk-möltgen 2012). over the last decades, research infrastructures in the social sciences have been established to manage the technical aspects of rdm – including metadata – within the whole research (data) lifecycle (kaase 2013; renschler et al. 2013). the consortium of european social science data archives (cessda)6 was established in 1976 and started out as an umbrella organization of seven social science data archives (mochmann 1998). as a leading research infrastructure for the social sciences cessda was awarded with the status of a european research infrastructure consortium (eric) by the european commission in 2017. the number of cessda member archives is increasing continually. the currently 20 cessda members work to improve data access for researchers. for this, cessda provides large-scale, integrated and sustainable data services to the social sciences and supports national and international research and cooperation (consortium of european social science data archives cessda 2020). acknowledging the importance of metadata in the social sciences and further backing this cause (van der eycken et al. 2019), cessda has started the metadata office project (mdo)7 in 2019. while cessda in general provides a full scale sustainable research infrastructure that supports social scientists in conducting high-quality research, mdo forms a core conceptual and strategic group consisting of six https://doi.org/10.29173/iq970 https://www.ddialliance.org/ https://www.cessda.eu/ https://www.cessda.eu/about/projects/work-plans/work-plan-2019#metaoff 3/17 förster, andré; borschewski, kerrin; bolton, sharon; jääskeläinen, taina (2020) the matter of meta in research data management: introducing the cessda metadata office project, iassist quarterly 44(3), pp.1-17. doi: https://doi.org/10.29173/iq970 partner institutions8 to maintain and manage cessda’s metadata-related material and monitors metadata developments and the metadata community. mdo not only oversees the strategic components and developments of all metadata related issues (including giving recommendations to cessda data archives, also named cessda service providers), but also manages and coordinates the content of the european language social science thesaurus (elsst)9 and related multilingual vocabulary services. this project report summarizes mdo’s various activities, illustrates its impact for the social science and rdm community, and finally offers an outline regarding mdo’s future plans. 2. development of metadata products mdo is not the first project dedicated to metadata within the cessda community. since its establishment, cessda has been active in metadata matters, e.g. by establishing a common catalogue, engaging in the ddi alliance, or in data exchange across europe. former projects include the eu funded projects nesstar, madiera, metadater, data without boundaries and cessda-ppp (jensen 2010; jensen and mochmann 2003; mauer 2012; silberman and tubaro 2008). the latest finished project on metadata issues within cessda and mdo’s forerunner is the cessda metadata management project. one of its main results was the development (zenk-möltgen et al. 2015) of a preliminary version of the cessda metadata model (cmm), which has been further developed and published as a first version within mdo (borschewski et al. 2019). we present the cmm in the following chapter (2.1). in general, it should be noted that projects on metadata accompany the developments of metadata standards such as ddi. while ddi is an elaborate metadata standard in the social sciences, its uptake and application vary between institutions, and even within cessda. thus, there is a need for a common semantic understanding of metadata issues that helps institutions to align their conceptual and technical metadata requirements with these standards. projects on metadata such as mdo help to establish this understanding within the cessda community. 2.1 the cessda metadata model (cmm) the cmm is built from the viewpoint of quantitative social science data. thus, it serves the purpose of helping cessda service providers to make their quantitative data more discoverable and comprehensible to users. the cmm is based on the ddi lifecycle metadata standard, because it is currently the most elaborate standard for the social sciences and due to its objective of interoperability. furthermore, most of the cessda archives use one of the ddi specifications for their metadata. for the cmm, the cessda metadata management project agreed on elements concerning the quantitative social sciences which were considered important for cessda tools, such as the cessda data catalogue (cdc)10, the cessda euro question bank (eqb)11, and for future tools of cessda. the aim of the cmm is to be an understandable, conceptual metadata schema. it is supposed to support cessda tools and also cessda service providers to have an overview of currently relevant metadata elements within cessda. facing this challenge of being as comprehensive as possible, but also as easy to handle as possible, the cmm currently does not include elements that do not adhere to the abovementioned characteristics. following these requirements, the cessda metadata management project decided to include elements based on corresponding ddi lifecycle 3.2 elements into the cmm. https://doi.org/10.29173/iq970 https://elsst.ukdataservice.ac.uk/ https://datacatalogue.cessda.eu/ https://datacatalogue.cessda.eu/ https://www.cessda.eu/about/projects/work-plans/work-plan-2019#eqb19 4/17 förster, andré; borschewski, kerrin; bolton, sharon; jääskeläinen, taina (2020) the matter of meta in research data management: introducing the cessda metadata office project, iassist quarterly 44(3), pp.1-17. doi: https://doi.org/10.29173/iq970 however, for the cmm elements, the project chose more conclusive element names. as all elements follow the ddi lifecycle 3.2 structure, the x-path examples can be found within the cmm. the structure of the cmm is based on the principle of reusing metadata elements. this means that wherever possible, information is referenced and reused according to ddi lifecycle 3.2, and the elements were stored within the ddi lifecycle resourcepackage. within resourcepackage information that is intended to be reused by multiple studies can be structured. the possibility to reference metadata information reduces the amount of work involved in the documentation process. the first version of the cmm contains more than 450 elements. however, the attributes concerning a certain element (such as information on language) are counted as separate elements. the latest version of the cmm was published in november 2019. it includes a mapping to the current version of the cdc metadata schema and corresponding ddi 2.5 x-paths, and it is also accompanied by a detailed documentation (storviken et al. 2019). the new version also includes a sheet where users can track all the changes that have been made to the cmm from version 0.1 to the current version 1.0. the cmm contains metadata elements, their definitions, examples of their use, and information on specific requirements, such as mandatories, repeatability and use of controlled vocabularies. table 1 presents the cmm overview, listing the information covered by the cmm elements on different levels. these levels are captured in the sections information on study, information on person(s), information on organization(s), information on dataset, information on instrument, information on questions and responses, information on concepts, information on further documents, information on publication, information on group of studies and information on document description. furthermore, the cmm overview offers explanations of the cmm characteristics. these are used to define the cmm elements in detail. specifically, the cmm includes information on the element number, on the name of an element, the definition on an element, and information on the status of an element, which is used to define whether the respective cmm element is mandatory, recommended or optional. moreover, the cmm offers information on standardized and controlled content for each element. this informs the user whether to employ a controlled vocabulary, an iso code or a thesaurus and also what type he or she should use. to give additional information to the status of an element and to make it easier for technicians to process the information, the cmm includes a column with the occurrences of an element. the last column of cmm contains ddi3.2 xpaths examples. the ddi3.2 x-path examples are exemplary ddi3.2 mappings to the element that are supposed to help the understanding of the respective element. to make this clearer, we present an example. figure 1 shows how the various characteristics are filled with information for each cmm element. for our example, we use the element ‘language of study https://doi.org/10.29173/iq970 5/17 förster, andré; borschewski, kerrin; bolton, sharon; jääskeläinen, taina (2020) the matter of meta in research data management: introducing the cessda metadata office project, iassist quarterly 44(3), pp.1-17. doi: https://doi.org/10.29173/iq970 title’. the number of the element is 1.1.3.1. this shows that it is a sub element of ‘bibliographic information’ (element number 1.1). the element ‘bibliographic information’ itself is a sub element of ‘study’ (element number 1). the definition of our example element with the element name ‘language of study title’ is ‘the language of the content of the element’. this means that the element is used to display the language in which the ‘study title’ was captured within the metadata. in the column for ‘status’ we find the information ‘m (for ddi3.2)’, meaning that if the top element of ‘language of study title’, namely the element ‘study title’ is used, the use of the ‘language of study title’ is mandatory for metadata captured in ddi3.2. the following column regarding standardization contains the text ‘use iso 639-1 (language code)’ for our example element 1.1.3.1. this means that the iso language code 639-1 has to be used to capture the information of the language. therefore, if the study title was in, say, finnish, the content of element 1.1.3.1 would be ‘fi’. the occurrence of our example element is ‘1’, meaning that for each time the element study title is used, the information ‘language of study title’ must be given. it also means that this information on the element can only be used once. the last column contains the ddi3.2 x-path example, which in our case is ‘ddi:ddiinstance/s:studyunit/r:citation/r:title/r:string/@xml:lang’. since the cmm has a conceptual nature and is independent of actual implementation, its further development by mdo supports cessda service providers and other research infrastructure institutions in the provision and maintenance of metadata while still offering them many possible ways to actually store, manage, organize and present metadata. thus, the final implementation of the cessda cmm remains in the service providers’ and institutions’ own responsibility. however, cessda has great interest in enabling its service providers to produce high-quality metadata, and one way to ensure this is fostering the standardization of metadata delivered by the service providers and received by cessda tools (such as cdc and eqb). standardization is achieved by giving clear instructions on how to use specific cmm elements, controlled vocabularies etc. more specifically, cessda’s long-term goal regarding standardization is that all service providers use ddi3.2. however, many providers lack resources to develop an editor being able to handle ddi3.2 references, leading to the majority of service providers still using ddi2.5. since cessda tools need to be able to harvest metadata even in the current situation, it would not make sense to define elements as mandatory that cannot be produced using the ddi2.5 specification. in general, fostering the process of standardization also means that metadata have to be checked continuously, errors have to be found and those errors have to be communicated and corrected. at the moment, mdo enters differences detected in harvested metadata into an issue tracker and assigns these issues to the respective service providers, asking them to amend their data (see also chapter 3). additionally, cessda is preparing a strategy on how it can help service providers to move towards ddi3.2. since different services require information in different languages, the cmm supports multilingualism by requesting users to produce language tags in their metadata and by providing controlled vocabularies in cessda languages (see also chapter 2.2). specifically, the cmm allows users to provide metadata in different languages, requesting the use of iso language tags in metadata. in doing so, different services are able to detect in which language the metadata and its elements actually are. this is particularly important when harvesting multilingual metadata files. in addition, cessda provides controlled vocabularies (e.g. the ddi and cessda vocabularies) included in the cmm in different cessda languages. the controlled vocabularies included in the cmm are currently provided https://doi.org/10.29173/iq970 6/17 förster, andré; borschewski, kerrin; bolton, sharon; jääskeläinen, taina (2020) the matter of meta in research data management: introducing the cessda metadata office project, iassist quarterly 44(3), pp.1-17. doi: https://doi.org/10.29173/iq970 in eleven different languages (see chapter 2.2). as mdo works closely together with the respective experts on controlled vocabularies within the ddi alliance, most vocabularies – with the exception of the cessda-specific topic classification – are ddi controlled vocabularies. both ddi and cessda vocabularies can be reused via the creative commons by 4.0 license12. multilingualism in the cmm is not only reflected in the controlled vocabularies, but also in certain text elements. language information is especially important for the cessda eqb. for elements of which eqb requires the language information, the use of the language attribute is mandatory. for instance, cmm includes the element ‘question item text’, which is a text element. since eqb also uses this element and requires information about language, the element has a mandatory attribute ‘language of question item text’. many language attribute elements are mandatory within cmm, apart from elements where this does not necessarily make sense. for instance, this accounts for the element ‘variable name’, since a variable name can also be an alphanumeric code not specific to a certain language and will not be translated, if a dataset is distributed in another language. apart from the maintenance and conceptual enhancement of the cmm and its documentation, mdo’s task is also to improve its technical representation. regarding a technically sophisticated form of the cmm, mdo is currently developing an xml schema definition (xsd) and a ddi profile derived from the model, accompanied by application profiles for other cessda tools and services, such as the eqb and the cessda vocabulary service (cvs)13. https://doi.org/10.29173/iq970 https://creativecommons.org/licenses/by/4.0/ https://vocabularies.cessda.eu/ 7/17 förster, andré; borschewski, kerrin; bolton, sharon; jääskeläinen, taina (2020) the matter of meta in research data management: introducing the cessda metadata office project, iassist quarterly 44(3), pp.1-17. doi: https://doi.org/10.29173/iq970 contained information and numbering: complete list of elements information on study: 1 information on person(s): 2 information on institution(s): 3 information on dataset: 4 information on instrument: 5 information on questions and responses: 6 information on concepts: 7 information on further documents: 8 information on publication: 9 information group of studies: 10 information on document description: 11 spreadsheet column headlines signification no. number of elements; represents the structure (1.1 means ”is child element of” 1) element name of element definition definition of element status (mandatory / recommended / optional) is this element mandatory (m), recommended (r) or optional (o). condition (if applicable) for m / r / o if applicable: under which condition is the element mandatory / recommended / optional? standardized/ controlled content for this element which cv / standard is used for the element (if ddi cv: http://ddialliance.org/controlled-vocabularies), or iso / thesaurus etc. or default values occurence occurence of element ddi 3.2 element x-path for ddi 3.2 elements. important remark: the x-paths are only preliminary and exemplary mappings. they are supposed to help with understanding the meaning of the elements. however, the final technical implementation of the elements will be up to the cessda sp. there will be no constraint to adopt them. mapping information: cdc element property name mapping cdc to cmm (version september 2018 https://docs.google.com/spreadsheets/d/1u9nsmvcwh1emcpgkrdomzv9uw0tbdtrdznr74vigky8/edit#gid=0) . cdc element property name mapping information: cdc element name for interface mapping cdc to cmm (version september 2018 https://docs.google.com/spreadsheets/d/1u9nsmvcwh1emcpgkrdomzv9uw0tbdtrdznr74vigky8/edit#gid=0) . cdc element name for interface mapping information: mandatoriness (mustshouldcould) mapping cdc to cmm (version september 2018 https://docs.google.com/spreadsheets/d/1u9nsmvcwh1emcpgkrdomzv9uw0tbdtrdznr74vigky8/edit#gid=0) . mapping information: mandatoriness (mustshouldcould) of cdc mapping information: schema element in ddi 2.5 mapping cdc to cmm (version september 2018 https://docs.google.com/spreadsheets/d/1u9nsmvcwh1emcpgkrdomzv9uw0tbdtrdznr74vigky8/edit#gid=0) . mapping information: schema element in ddi 2.5 by cdc mapping information: notes for ddi 2.5 schema mapping cdc to cmm (version september 2018 https://docs.google.com/spreadsheets/d/1u9nsmvcwh1emcpgkrdomzv9uw0tbdtrdznr74vigky8/edit#gid=0) . mapping information: notes for ddi 2.5 schema by cdc table 1: information covered by the cmm. https://doi.org/10.29173/iq970 8/17 förster, andré; borschewski, kerrin; bolton, sharon; jääskeläinen, taina (2020) the matter of meta in research data management: introducing the cessda metadata office project, iassist quarterly 44(3), pp.1-17. doi: https://doi.org/10.29173/iq970 figure 1: characteristics of the cmm elements. note that only a selection of elements is shown. https://doi.org/10.29173/iq970 9/17 förster, andré; borschewski, kerrin; bolton, sharon; jääskeläinen, taina (2020) the matter of meta in research data management: introducing the cessda metadata office project, iassist quarterly 44(3), pp.1-17. doi: https://doi.org/10.29173/iq970 2.2 the cessda vocabulary service (cvs) the cvs11 provides a user-friendly source of standardized controlled vocabularies. these controlled vocabularies can be defined as hierarchically organized lists of codes with descriptive terms and definitions in one or more languages. currently, there are 24 controlled vocabularies available in the cvs. the service contains both an editor for creating and updating vocabularies and a user interface where users can search and browse the published vocabularies and download them in different formats. the service also provides uniform resource names (urns) for both the controlled vocabulary and for each version of it, as well as an application programming interface (api). at concept level, code value is the identifier that stays the same across all language versions. figures 2 and 3 present the cvs search interface and a detailed view of a specific vocabulary. figure 2: cvs search interface. figure 3: detailed view of a specific vocabulary. https://doi.org/10.29173/iq970 https://vocabularies.cessda.eu/ 10/17 förster, andré; borschewski, kerrin; bolton, sharon; jääskeläinen, taina (2020) the matter of meta in research data management: introducing the cessda metadata office project, iassist quarterly 44(3), pp.1-17. doi: https://doi.org/10.29173/iq970 vocabulary management is done in the cvs editor. access to the editor is governed by agency and language specific user roles. source language administrators can create source vocabularies of their agency and translators translate them into their language. all administrators can publish and version their vocabularies, as well as browse all draft, unpublished agency vocabularies in any language. registered users receive training to ensure that they are following best practice. training includes a training webinar, power point slides and an extensive online user guide, thus making the use of the tool easy and straightforward. the cvs provides a respected and authoritative source for cessda controlled vocabularies which data producers can use to ensure systematic metadata and description for data assets. since ddi vocabularies form an important part in standardizing metadata, the ddi alliance uses the cvs editor for managing its vocabularies and their translations. their vocabularies are published on the ddi website, but the vocabularies are also published in the cvs user interface, where they can be browsed. the vocabulary service can also be used for managing other agency vocabularies, for instance those of other research infrastructures or data repositories. organizations interested in using the tool for creating their agency vocabularies or in producing and maintaining a new language version of ddi or cessda vocabularies can contact the mdo team.8 since ddi is an international standard, there is some incentive to add languages from outside of cessda. thus, the cvs provides cessda and other users with a robust multilingual service for the standardized description of social science data and associated materials. vocabulary information includes detailed documentation of any changes in published vocabularies, which enables users to update their legacy metadata. the cvs was created as an internal cessda project in 2017 to 2018, with participants from gesis, the finnish social science data archive (fsd), the united kingdom data service (ukds) and the swedish national data service (snd). when the project ended, fsd and ukds continued their work on the cvs within the mdo project. in 2019, they provided user feedback regarding the cvs system, tested some amendments and compiled an online user guide. fsd and ukds act as content managers of the cvs, handling user management, access and training. currently there are vocabularies available in eleven languages: danish, english, finnish, french, german, italian, norwegian, portuguese, serbian, slovenian and swedish. japanese and estonian are expected to be added next. 2.3 the european language social science thesaurus (elsst) the elsst9 is a broad-based, multilingual thesaurus for the social sciences. a thesaurus is also a controlled vocabulary. rather than a list, however, a thesaurus comprises a structure that consists of terms and the relationships between them. terms may be related hierarchically (broader/narrower) or non-hierarchically (related and synonymous). this is a more complex structure to the controlled vocabularies held within the cvs described above and means that elsst is not managed via the current cvs system but is held in a separate ontology management system that enables sophisticated hierarchical editing. also, their release schedule is currently different; elsst has an annual new version release including all languages. controlled vocabularies in cvs, on the other hand, are published as soon as they have been finalized, and each language is versioned and published separately with their own time schedule. controlled vocabularies and elsst keywords are utilized in separate ways within the metadata record. controlled vocabularies are used or are planned to be used as filters in search interfaces which places even stricter requirements for the harvested metadata to be standardized. https://doi.org/10.29173/iq970 https://elsst.ukdataservice.ac.uk/ 11/17 förster, andré; borschewski, kerrin; bolton, sharon; jääskeläinen, taina (2020) the matter of meta in research data management: introducing the cessda metadata office project, iassist quarterly 44(3), pp.1-17. doi: https://doi.org/10.29173/iq970 cessda service providers are required to update their legacy metadata after changes regarding the controlled vocabularies. they need human-readable, detailed version history to see what they need to change and for deciding which changes can be done by machine and which require manual updates. elsst keywords are different, as they describe in detail the actual subjects and concepts covered by the data, hence the more complex structural relationships between terms and editing facilities that elsst requires. in the mediumto long-term, it is planned that the management of cessda metadata ontologies including the cvs and elsst will be done within the same system, which will enable synergies between where possible although they are intended for different purposes. the elsst was originally based on the monolingual humanities and social science electronic thesaurus (hasset), developed by the uk data archive at the university of essex. in 2000, the eu-funded language independent metadata browsing of european resources (limber) project developed the first multilingual version of elsst, translating the english hasset into french, german and spanish. the thesaurus was further enhanced and extended through subsequent eu grants (such as madiera) and additional uk funding. in 2018, the cessda vocabulary services multilingual content management (voice) project took over development and management of elsst, moving forward with editorial work, software enhancements, and additional translations. the cessda multilinguality policy was also developed in voice with a view to managing elsst in the context of other multilingual controlled vocabularies, such as the cvs. from 2019, it was therefore logical to manage elsst within the mdo alongside cessda’s other key metadata assets, the cmm and cvs. the elsst is currently available in 14 languages and is managed by a dedicated team within the mdo project, in close cooperation with expert translators. in order to improve the functionality, currency, and utility of elsst, its multilingual content must be maintained and updated. therefore, regular collaboration takes place between mdo partners to review elsst content (with input from subject specialists) and updates are released annually. the 2019 release included a new language translation (dutch) and a new set of translated scope notes (slovenian), revised terms and hierarchy structures (groups of terms, held within a specific relationship to each other, that describe aspects of a social science concept), selected to strengthen subject coverage and reflect emerging topics in social science. elsst provides an indexing resource for data producers to aid in the curation and publishing of their data. in particular, updates to different language versions of elsst allow a common approach to indexing, which will benefit the cdc and the catalogues of cessda service providers, as elsst terms can be used for search and filtering purposes. guidance for users and translators is kept updated, ensuring that all the information they need to use the system is readily available. during 2020, elsst will move to a new technical software platform that will enable better linked open data capabilities and interoperability. the mdo’s elsst team is currently in consultation with the producers of other key international thesauri, such as eurovoc, the european union publications office official thesaurus, and the food and agriculture organization (fao) of the united nations’ agrovoc multilingual thesaurus, to exchange knowledge about multilingual thesaurus management and explore potential mappings between elsst and other key thesauri. due to elsst’s development history across various projects, database rights in the organisation and collection of the data, and the underlying application, are currently held by the university of essex. https://doi.org/10.29173/iq970 12/17 förster, andré; borschewski, kerrin; bolton, sharon; jääskeläinen, taina (2020) the matter of meta in research data management: introducing the cessda metadata office project, iassist quarterly 44(3), pp.1-17. doi: https://doi.org/10.29173/iq970 copyright in the natural language translations are held either by the individual translators or the relevant translating organisation. organisations can obtain a licence from the uk data service to use or adapt elsst. once elsst moves to its new technical platform, a more streamlined licencing model will be put in place using creative commons (similar to the cvs), administered by cessda. 3. communication, coordination, training and outreach an important result from previous research on rdm and from previous metadata related cessda projects is that communication, guidance and training are needed to achieve more involvement in the matter of research data services in general and specifically in metadata standards (tenopir et al. 2014). providing these resources also positively affects the implementation of standards within the cessda community and its institutions. since cessda has always been active in teaching and learning, its involvement established in the cessda training14 pillar, mdo has acted accordingly, for instance by integrating rather small and not yet well-established cessda data archives into the mdo project. in particular, the data centre serbia for social sciences (dcs) and the portuguese social information archive (apis) have worked on the translation of controlled vocabularies and on further developing the additional documentation materials of the cmm. furthermore, proper coordination as part of these efforts is also required when it comes to maintaining technical metadata issues (e.g. harvesting of metadata, discovery of studies etc.) and implementing cessda’s metadata standards in close cooperation between mdo and cessda service providers. therefore, mdo and the cessda main office have established a cessda internal metadata issue tracking system via a bitbucket15 issue tracker. since possibly sensitive technical questions might also be addressed there, access rights to the tracker are managed by the cessda main office, while mdo currently is responsible for managing the participation of the community. this way, we ensure that metadata experts at the cessda service providers are put in charge of solving the respective issues. in order to further increase the visibility of metadata issues within the cessda community, mdo has undertaken several additional outreach efforts. for instance, mdo has installed an mdo newsletter that informs the cessda community about the latest activities and future plans of the project. currently, it is published at irregular intervals via the cessda basecamp16 groups ‘cessda metadata office’ and ‘service providers’ forum’ and is also available on request from mdo. in general, the cessda community and its service providers take part in the further development of mdo’s activities in various ways. while a core group of six service providers manages the mdo project, all service providers can reach out to the mdo project team to give feedback regarding their metadata requirements and the use of the respective documents and tools (e.g. cmm, cvs). additionally, the mdo project has contacted each service provider and all the tools separately, asking for feedback concerning their needs on the mdo tasks. mdo is also exploring contacts within other cessda projects such as eqb and the cessda training group to raise awareness of mdo and ask what it can do for them. further activities include the participation in cessda events (such as the cessda expert seminar 2019) and the possibility to host additional new service providers within the project (such as dcs and apis). regarding public services and outreach beyond cessda, mdo has liaised with various metadata experts from other organizations. for instance, fsd has arranged a seminar on metadata, data catalogues and tools for findability, with participants from snd and the japanese national project on https://doi.org/10.29173/iq970 https://www.cessda.eu/training https://bitbucket.org/product/ https://basecamp.com/welcome-back 13/17 förster, andré; borschewski, kerrin; bolton, sharon; jääskeläinen, taina (2020) the matter of meta in research data management: introducing the cessda metadata office project, iassist quarterly 44(3), pp.1-17. doi: https://doi.org/10.29173/iq970 developing a joint data catalogue for the social sciences (laaksonen 2019). furthermore, mdo members of ukds attended the iassist 2019 conference and presented ‘sustainable european multilingual vocabularies: a model for cooperation in metadata management among european data archives’ (barbalet and bolton 2019). based on this presentation, ukds is working together with the australian data archive and the department of library and information science of punjabi university, patiala on further opportunities for cooperation. presentations on cvs, ddi vocabularies and their use in cessda were held at the eddi 2019 conference (bolton and jääskeläinen 2019; jääskeläinen and bolton 2019). furthermore, the cdc is also now included in the european open science cloud (eosc) marketplace17, which offers another source of (non-cessda) contacts and their feedback. additionally, mdo is participating in the research data alliance (rda)18 metadata interest group. all feedback is used to enhance the quality of the mdo metadata materials. more specifically, the cessda-external feedback enables mdo to expand its horizons. for instance, we can get an idea whether the cmm is used in more general contexts apart from cessda and which challenges other institutions encounter when using it. in general, we can learn about new questions and ideas in the metadata world. 4. future plans the future plans of mdo include a systematic review of international metadata developments in relevant consortia, institutions and projects. furthermore, since one of the general purposes of cessda’s metadata materials is the support of the fair principles, mdo will further develop these openly published materials, so that they clearly state in which way they contribute to aligning with the fair principles and how they support the obtainment of trust badges such as the core trust seal19. for instance, mdo will define a specific set of metadata profiles that are the standard for its service providers and monitor their application within the cdc. metadata services and research on rdm and metadata management serve the advancement of databased research, especially in the quantitative social sciences. by further developing cessda’s metadata products, mdo will continue its work within the cessda community and beyond, relying on the excellent working relationships built up in the first months of the project. with our work, we hope to contribute that data will always be more than just collections of numbers or codebooks. references barbalet, suzanne; bolton, sharon (2019): sustainable european multilingual vocabularies: a model for cooperation in metadata management among european data archives. available online at https://ukdataservice.ac.uk/media/622472/iassist19_sbarbalet_sbolton_02.pdf, checked on 7/23/2019. blask, katarina; förster, andré (2019): designing an information architecture for data management technologies: introducing the diamant model. in journal of librarianship and information science online first. doi: 10.1177/0961000619841419. bolton, sharon; jääskeläinen, taina (2019): ddi vocabularies and language versions. doi: 10.5281/zenodo.3596792. https://doi.org/10.29173/iq970 https://marketplace.eosc-portal.eu/services/cessda-data-catalogue https://www.rd-alliance.org/ https://www.coretrustseal.org/ https://ukdataservice.ac.uk/media/622472/iassist19_sbarbalet_sbolton_02.pdf https://dx.doi.org/10.1177/0961000619841419 http://doi.org/10.5281/zenodo.3596792 14/17 förster, andré; borschewski, kerrin; bolton, sharon; jääskeläinen, taina (2020) the matter of meta in research data management: introducing the cessda metadata office project, iassist quarterly 44(3), pp.1-17. doi: https://doi.org/10.29173/iq970 borschewski, kerrin; förster, andré; friedrich, tanja; zenk-möltgen, wolfgang; miranda, patrícia; moura ferreira, pedro et al. (2019): cmm cessda metadata model. doi: 10.5281/zenodo.3236171. consortium of european social science data archives (cessda) (2020): the cessda consortium. available online at https://www.cessda.eu/about/consortium, checked on 1/31/2020. deutsche forschungsgemeinschaft (dfg) (2013): vorschläge zur sicherung guter wissenschaftlicher praxis. proposals for safeguarding good scientific practice. available online at https://www.dfg.de/download/pdf/dfg_im_profil/reden_stellungnahmen/download/empfehlung_w iss_praxis_1310.pdf, checked on 7/18/2019. european commission (2016): guidelines on data management in horizon 2020. available online at https://ec.europa.eu/research/participants/data/ref/h2020/grants_manual/hi/oa_pilot/h2020-hioa-data-mgt_en.pdf, checked on 8/7/2019. green, ann e.; humphrey, chuck (2014): building the ddi. in iassist quarterly 37 (1-4), pp. 36–44. doi: 10.29173/iq500. gregory, arofan; heus, pascal; ryssevik, jostein (2009): metadata. in ratswd working paper series 57, pp. 1-22. doi: 10.2139/ssrn.1447866. hoyle, larry; castillo, fortunato; clark, benjamin; kashyap, neeraj; perpich, denise; wackerow, joachim; wenzig, knut (2011): metadata for the longitudinal data life cycle. doi: 10.3886/ddilongitudinal03. jääskeläinen, taina; bolton, sharon (2019): cessda vocabulary service – for managing vocabulary content. doi: 10.5281/zenodo.3600174. jensen, uwe (2010): data and metadata extensions of the cessda ri. enhancement of data and metadata infrastructures for the cessda ri. available online at https://ppp.cessda.eu/doc/d8.3_data_metdata_enhancement.pdf, checked on 8/7/2019. jensen, uwe; mochmann, ekkehard (2003): metadater: towards standards and tools for the description of comparative surveys. in za-information 52 (1), pp. 191–198. available online at https://www.gesis.org/fileadmin/upload/forschung/publikationen/zeitschriften/za_information/zainfo-52.pdf, checked on 8/7/2019. kaase, max (2013): research infrastructures in the social sciences: the long and winding road. in brian kleiner, isabelle renschler, boris wernli, peter farago, dominique joye (eds.): understanding research infrastructures in the social sciences. zürich: seismo press, pp. 19–27. karjalainen, merja; kleemola, mari; jensen, uwe (2013): metadata standards usage and needs in nsis and data archives. available online at http://www.dwbproject.org/export/sites/default/about/public_deliveraples/dwb_d7-1_metadatastandards-usage_report.pdf, checked on 8/16/2019. laaksonen, helena (2019): fsd's multilingual and qualitative data expertise brings in international visitors. available online at https://tietoarkistoblogi.blogspot.com/2019/04/jsps-seminar.html, checked on 7/23/2019. https://doi.org/10.29173/iq970 https://dx.doi.org/10.5281/zenodo.3236171 https://www.cessda.eu/about/consortium https://www.dfg.de/download/pdf/dfg_im_profil/reden_stellungnahmen/download/empfehlung_wiss_praxis_1310.pdf https://www.dfg.de/download/pdf/dfg_im_profil/reden_stellungnahmen/download/empfehlung_wiss_praxis_1310.pdf https://ec.europa.eu/research/participants/data/ref/h2020/grants_manual/hi/oa_pilot/h2020-hi-oa-data-mgt_en.pdf https://ec.europa.eu/research/participants/data/ref/h2020/grants_manual/hi/oa_pilot/h2020-hi-oa-data-mgt_en.pdf https://dx.doi.org/10.29173/iq500 https://dx.doi.org/10.2139/ssrn.1447866 https://dx.doi.org/10.3886/ddilongitudinal03 http://doi.org/10.5281/zenodo.3600174 https://ppp.cessda.eu/doc/d8.3_data_metdata_enhancement.pdf https://www.gesis.org/fileadmin/upload/forschung/publikationen/zeitschriften/za_information/za-info-52.pdf https://www.gesis.org/fileadmin/upload/forschung/publikationen/zeitschriften/za_information/za-info-52.pdf http://www.dwbproject.org/export/sites/default/about/public_deliveraples/dwb_d7-1_metadata-standards-usage_report.pdf http://www.dwbproject.org/export/sites/default/about/public_deliveraples/dwb_d7-1_metadata-standards-usage_report.pdf https://tietoarkistoblogi.blogspot.com/2019/04/jsps-seminar.html 15/17 förster, andré; borschewski, kerrin; bolton, sharon; jääskeläinen, taina (2020) the matter of meta in research data management: introducing the cessda metadata office project, iassist quarterly 44(3), pp.1-17. doi: https://doi.org/10.29173/iq970 mauer, reiner (2012): das gesis datenarchiv für sozialwissenschaften. in r. altenhöner, c. oellers (eds.): langzeitarchivierung von forschungsdaten. standards und disziplinspezifische lösungen. berlin: scivero, pp. 197–215. urn: nbn:de:0168-ssoar-46476-7. mochmann, ekkehard (1998): european co-operation in social science data dissemination. in rachel walker, marcia freed taylor (eds.): information dissemination and access in russia and eastern europe. problems and solutions in east and west. amsterdam: ios press, pp. 33–42. rasmussen, karsten boye (2014): social science metadata and the foundations of the ddi. in iassist quarterly 37 (1-4), pp. 28–35. doi: 10.29173/iq499. renschler, isabelle; kleiner, brian; wernli, boris (2013): concepts and key features for understanding social science research infrastructures. in brian kleiner, isabelle renschler, boris wernli, peter farago, dominique joye (eds.): understanding research infrastructures in the social sciences. zürich: seismo press, pp. 11–18. silberman, roxane; tubaro, paola (2008): data archives and access to government data for researchers. state of the art and future developments in europe in the eri perspective. available online at https://gala.gre.ac.uk/id/eprint/5437/1/6_1_cessda.pdf, checked on 8/7/2019. storviken, silje; hagen, sunniva; bockaj, brigita; bolko, irena; vipavc brvar, irena; fink kjeldgaard, anne sofie et al. (2019): user guide for the cessda metadata model. doi: 10.5281/zenodo.3236193. tenopir, carol; sandusky, robert j.; allard, suzie; birch, ben (2014): research data management services in academic research libraries and perceptions of librarians. in library & information science research 36, pp. 84–90. doi: 10.1016/j.lisr.2013.11.003. van den eynden, veerle; bishop, libby (2014): sowing the seed: incentives and motivations for sharing research data, a researchers' perspective. available online at http://repository.jisc.ac.uk/5662/1/ke_report-incentives-for-sharing-researchdata.pdf, checked on 7/18/2019. van der eycken, j.; styven, d.; gheldof, t.; depoortere, r. (2019): trust and understanding. the value of metadata in a digitally joined-up world: conclusion, a vision for the future. in r. depoortere, t. gheldof, d. styven, j. van der eycken (eds.): trust and understanding: the value of metadata in a digitally joined-up world. brussels: archives et bibliothèques de belgique, pp. 135–144. available online at https://hal.archives-ouvertes.fr/hal-02125062/document, checked on 8/7/2019. vardigan, mary (2014): the ddi matures: 1997 to the present. in iassist quarterly, pp. 45–50. available online at https://iassistquarterly.com/pdfs/iqvol371_4_vardigan.pdf, checked on 7/19/2019. vardigan, mary; heus, pascal; thomas, wendy (2008): data documentation initiative: toward a standard for the social sciences. in international journal of digital curation 3 (1), pp. 107–113. doi: 10.2218/ijdc.v3i1.45. wilkinson, mark d.; dumontier, michel; aalbersberg, ijsbrand jan; appleton, gabrielle; axton, myles; baak, arie et al. (2016): the fair guiding principles for scientific data management and stewardship. in sci. data 3: 160018. doi: 10.1038/sdata.2016.18. https://doi.org/10.29173/iq970 https://nbn-resolving.org/urn:nbn:de:0168-ssoar-46476-7 https://dx.doi.org/10.29173/iq499 https://gala.gre.ac.uk/id/eprint/5437/1/6_1_cessda.pdf https://dx.doi.org/10.5281/zenodo.3236193 https://dx.doi.org/10.1016/j.lisr.2013.11.003 http://repository.jisc.ac.uk/5662/1/ke_report-incentives-for-sharing-researchdata.pdf https://hal.archives-ouvertes.fr/hal-02125062/document https://iassistquarterly.com/pdfs/iqvol371_4_vardigan.pdf https://dx.doi.org/10.2218/ijdc.v3i1.45 https://dx.doi.org/10.1038/sdata.2016.18 16/17 förster, andré; borschewski, kerrin; bolton, sharon; jääskeläinen, taina (2020) the matter of meta in research data management: introducing the cessda metadata office project, iassist quarterly 44(3), pp.1-17. doi: https://doi.org/10.29173/iq970 wolf, christof (2017): implementing open science: the gesis perspective. talk given at institute day of gesis, 28 september 2017. in gesis papers (26), pp. 1–12. available online at https://www.ssoar.info/ssoar/bitstream/handle/document/54950/ssoar-2017-wolfimplementing_open_science_the_gesis.pdf?sequence=1&isallowed=y&lnkname=ssoar-2017-wolfimplementing_open_science_the_gesis.pdf, checked on 7/16/2019. zenk-möltgen, wolfgang (2012): metadaten und die data documentation initiative. in r. altenhöner, c. oellers (eds.): langzeitarchivierung von forschungsdaten. standards und disziplinspezifische lösungen. berlin: scivero, pp. 111–126. available online at https://www.ssoar.info/ssoar/bitstream/handle/document/46679/zenkmoeltgen_metadaten%20und%20die%20data%20documentation%20initiative.pdf?sequence=1, checked on 8/7/2019. zenk-möltgen, wolfgang; kleemola, mari; etheridge, anne; fink kjeldgaard, anne sofie (2015): introducing the cessda metadata management project. available online at http://www.eddiconferences.eu/ocs/index.php/eddi/eddi15/paper/view/190/174, checked on 8/7/2019. end-notes 1 andré förster was the head of the cessda metadata office project at the gesis – leibniz institute for the social sciences in 2019. he now works for the national coordination agency in education monitoring in germany and can be reached by email: andre.foerster@kommunalesbildungsmonitoring.de (version: march 2020) 2 kerrin borschewski works in the cessda metadata office project and in cessda training. she is a doctoral candidate at the gesis – leibniz institute for the social sciences and can be reached by email: kerrin.borschewski@gesis.org 3 sharon bolton is head of the cessda metadata office project at the uk data service. she can be reached by email: sharonb@essex.ac.uk 4 taina jääskeläinen works at the finnish social science data archive and participates in the cessda metadata office project. she can be reached by email: taina.jaaskelainen@tuni.fi 5 data documentation initiative; https://www.ddialliance.org/ 6 https://www.cessda.eu/ 7 https://www.cessda.eu/about/projects/work-plans/work-plan-2019#metaoff 8 the partners are the gesis – leibniz institute for the social sciences, the united kingdom data service (ukds), the finnish social science data archive (fsd), the norwegian centre for research data (nsd), the data centre serbia for social sciences (dcs) and the portuguese social information archive (apis). mdo can also be reached by email: metadata-office@cessda.eu 9 https://elsst.ukdataservice.ac.uk/ https://doi.org/10.29173/iq970 https://www.ssoar.info/ssoar/bitstream/handle/document/54950/ssoar-2017-wolf-implementing_open_science_the_gesis.pdf?sequence=1&isallowed=y&lnkname=ssoar-2017-wolf-implementing_open_science_the_gesis.pdf https://www.ssoar.info/ssoar/bitstream/handle/document/54950/ssoar-2017-wolf-implementing_open_science_the_gesis.pdf?sequence=1&isallowed=y&lnkname=ssoar-2017-wolf-implementing_open_science_the_gesis.pdf https://www.ssoar.info/ssoar/bitstream/handle/document/54950/ssoar-2017-wolf-implementing_open_science_the_gesis.pdf?sequence=1&isallowed=y&lnkname=ssoar-2017-wolf-implementing_open_science_the_gesis.pdf https://www.ssoar.info/ssoar/bitstream/handle/document/46679/zenk-moeltgen_metadaten%20und%20die%20data%20documentation%20initiative.pdf?sequence=1 https://www.ssoar.info/ssoar/bitstream/handle/document/46679/zenk-moeltgen_metadaten%20und%20die%20data%20documentation%20initiative.pdf?sequence=1 http://www.eddi-conferences.eu/ocs/index.php/eddi/eddi15/paper/view/190/174 http://www.eddi-conferences.eu/ocs/index.php/eddi/eddi15/paper/view/190/174 mailto:kerrin.borschewski@gesis.org mailto:sharonb@essex.ac.uk mailto:taina.jaaskelainen@tuni.fi https://www.ddialliance.org/ https://www.cessda.eu/ https://www.cessda.eu/about/projects/work-plans/work-plan-2019#metaoff mailto:metadata-office@cessda.eu https://elsst.ukdataservice.ac.uk/ 17/17 förster, andré; borschewski, kerrin; bolton, sharon; jääskeläinen, taina (2020) the matter of meta in research data management: introducing the cessda metadata office project, iassist quarterly 44(3), pp.1-17. doi: https://doi.org/10.29173/iq970 10 https://datacatalogue.cessda.eu 11 https://www.cessda.eu/about/projects/work-plans/work-plan-2019#eqb19 12 https://creativecommons.org/licenses/by/4.0/ 13 https://vocabularies.cessda.eu/ 14 https://www.cessda.eu/training 15 https://bitbucket.org/product/ 16 https://basecamp.com/welcome-back 17 https://marketplace.eosc-portal.eu/ 18 https://www.rd-alliance.org/ 19 https://www.coretrustseal.org/ https://doi.org/10.29173/iq970 https://datacatalogue.cessda.eu/ https://www.cessda.eu/about/projects/work-plans/work-plan-2019#eqb19 https://creativecommons.org/licenses/by/4.0/ https://vocabularies.cessda.eu/ https://www.cessda.eu/training https://bitbucket.org/product/ https://basecamp.com/welcome-back https://marketplace.eosc-portal.eu/ https://www.rd-alliance.org/ https://www.coretrustseal.org/ 62 iassist quarterly 2013 iassist quarterly adoption of data citation and in the promotion of data sharing and its benefits. the evolution of data citation: from principles to implementation by micah altman and mercè crosas 1 abstract data citation is rapidly emerging as a key practice in support of data access, sharing, reuse, and of sound and reproducible scholarship. in this article we review the evolution of data citation standards and practices – to which sue dodd was an early contributor – and the core principles of data citation that have emerged through a collaborative synthesis. we then discuss an example of the current state of the practice, and identify the remaining implementation challenges. keywords: data citation; bibliographic practices . background data is, as they say, the new black. scientific data are increasingly being made available online, and access to large collections of data is increasingly sought for education, science, policy, and commerce. lowering barriers to discovery and use of these data and increasing our ability to link data with publications have the potential to enable new forms of scholarly publishing, promote interdisciplinary research, strengthen the linkage between policy and science, and lower the costs of replicating and extending previous research. many problems arise when research findings become disconnected from the underlying data that forms the evidence for these findings. the most well-publicized of these problems is scientific fraud. access to data and the documentation of clear connections between the research results and the data facilitate detection of structural fraud both before and after publication. other problems arising from this disconnect include irreproducibility, lack of reuse and wasted effort collecting new data, a proliferation of unmanaged versions and subsets of the ‘same’ data, and weak incentives for data sharing. this is why the submission requirements for science, one of the most cited, read, and respected journals in the sciences, requires that “all data necessary to understand, assess, and extend the conclusions of the manuscript must be available to any reader of science” and that “citations to unpublished data and personal communications cannot be used to support claims in a published paper” (emphasis added). (science 2014) too often, this proscription, and others like it, have been honored only in the breach. the history of data sharing makes this clear – despite clear recognition of the benefits of data sharing (fienberg, et al. 1985) many research findings are based on data that is not made available -making this research surprisingly difficult to replicate and even more difficult to extend. furthermore, most research articles fail to provide clear citations to data, or the code necessary to reproduce, reuse, or extend results (codata 2013). within the social sciences, the vast majority of datasets produced by sponsored research is never deposited or shared (pienta 2006), and, as a result, reproducing published tables and figures, and directly extending prior results is often difficult or impossible (dewald, et al., 1986; altman, et al., 2003; hamermesh 2007). similar problems exist in other fields: a recent study by vines et al. (2014) of a sample of zoology articles found that less than 30% of even the most recent publications made data available, and that research data availability declined rapidly with article age, while loss of data increased. moreover, a study of articles published in high-impact journals during 2009 showed that only iassist quarterly 2013 63 iassist quarterly 41% minimally complied with the journal’s own data-sharing policies, and of these only 9% deposited the full primary raw data corresponding to the paper online (alsheikh-ali, et al 2011). the research community has begun to take wider notice of this. and in the past two years a number of efforts have been launched by publishers, funders, professional associations, and organized projects to improve reliability, reproducibility, and data availability across a variety of scientific fields. we are optimistic that these projects will succeed, and if they do a key part of their success is likely to be through better scholarly recognition of data authorship. there is increasing recognition that researchers are more inclined to share their data when they get credit (borgman, 2012, p. 1072). conversely, recent studies also suggest that researchers receive more credit when they share their data (piwowar & vision 2013). publications that shared data from earlier years yielded an increase in citations of up to 30%. data citation, which has existed for 40 years in principle, is finally emerging as a pivotal norm for promoting data accessibility and accountability. robust data citation practices and infrastructure will play a critical role in the widespread adoption of data citation and in the promotion of data sharing and its benefits. the emergence of data citation principles and practices within traditional print publishing, scholarly citation was widely formalized over a century ago. the first edition of the chicago manual of style, published in 1906 under the title manual of style: being a compilation of the typographical rules in force at the university of chicago press exemplified (and helped catalyze) the extent of standardization in scholarly citation. (pollack, 2006) within this tradition, a “bibliographic citation” referred to a formal, structured reference to another scholarly work that appeared in the text of a work. typically, citations were either marked off with parentheses or brackets, such as: “(altman 1992),” although in some fields footnotes were used. a standard reference entry included author(s), a title, a date, and a publisher (publishing house for books, journal name for articles) (van leunen 1992, pg. 186-208). in addition, citations could include “pinpointing” information that identified which part of the cited work was being referenced, typically in the form of a page range. citations to a single work could be repeated throughout the text. the reference list, typically appearing at the end of the main text, provided more detailed bibliographic information for each work cited in the text. many variations were used for references to archival sources, correspondence, government documents, and artworks. however, each of these reference formats provided as well as possible at least three elements: author/creator, dates of the work, and the publisher or distributor of the work. when the first scientific digital data archives were established in the late 1960s, their design focused on issues of access, storage, formatting, costs, and information retrieval (bisco 1965). bibliographic standards for cataloging data were developed over the next decade. in 1970 the american library association (ala) formed a subcommittee on rules for cataloging machinereadable data files (mrdf), and tasked it with, among other things, 1977$% 1998% icpr%% archive% marc% catalog%systems.% !"facilitate"descrip.on" "&"informa.on"retrieval" !"describe"data"in"archives"" !"describe"as"works"not"media" !"provide"author,".tle,"version." [avram"1975]" [dodd"1979]" [isbd"1990]" [iso"1997]" 1999$% 2003% nesstar% virtual%data%center%% !"facilitate"access" "&"persistence" !"cite"research"data"in"all" publica.ons"that"use"it." !"provide"ac.onable"uri’s" !"provide"persistent"iden.fiers" !"use"persistent"ins.tu.ons" [altman,"et"al."2001]" [ryssevik"&"musgrave"2001]" 2004$% 2009% tib%doi%service% dataverse%network" % !"facilitate"verifica.on" "&"reproducibility" !"provide"bit!"or"seman.c!"fixity" !"provide"granularity" [brase"2004]" [buneman"2006]" [altman"&"king"2007]" 2009$% dataverse%network% datacite" data%dryad" figshare" data%citaeon%index% !"facilitate"integra.on" !"include"data"cita.ons"in" standard"loca.ons"in"text" !"index"data"cita.ons"in"exis.ng" catalogs" !"integrate"data"cita.on"with"" [uhlir"(ed.)"2012]" [codata"2013]" [data"synthesis"group"2014]" exemplar%systems% core%principles% key%work% figure 1: a chronology of data citation principles and related systems 64 iassist quarterly 2013 iassist quarterly bringing bibliographic control to mrdf. it was a long time before academic citation practices started to catch up with archiving practices, as summarized in figure 1. (in the figure, “exemplar systems” indicate key software and or technical infrastructure supporting practices. “core principles” summarizes the principles identified for data citation – as described in the text below. “key work” indicates the work related to principles of data citation and bibliographic practice – not responsibility for exemplar systems. ) the american standard for bibliographic reference (ansi z29.291977, aka asbr) provided a minimal “data file” type to be used as part of the general material designator element in bibliographic metadata. dodd (1979) quickly noted the shortcomings of asbr in practice – notably the inconsistencies in describing the same dataset when presented in different physical formats, and the fact that the general approach conflated specific media with the “intellectual works” temporarily stored in those media. dodd proposed using existing asbr elements in a consistent and systematic way to bibliographically describe datasets as intellectual works. the key elements of dodd’s approach emphasized the use of consistent title, author, and edition (which included date). they were used along with a general media designator of “machine readable data file(s)” (mrdf), which was format and media agnostic. the recognition of data as a public good2 was, however, insufficient by itself to support or incentivize data sharing. in general, public goods in the absence of effective norms, regulation, or subsidies will be under-supplied. the state of the art in data citation, as well as in data sharing, did not progress quickly until catalyzed through advances in information technology, open source software development practices, and legal infrastructure. the growing recognition among scholars that data is a fundamental product of research, a trend identified in the national research council’s foundational report on data sharing (fienberg 1985), began to build slowly in momentum through the leadership of individual scholars such as sieber (1991) and king (1995). then, rapid advances in internet and web infrastructure greatly decreased the technical barriers to data sharing. more recently the rapid growth of the open software movement generally, together with development of the legal “technology” of robust standardized open licenses, have sparked initiatives in academia to build open tools in support of scholarly access, discovery, collaboration, and research sharing. building on these trends, and supported by the nsf digital library initiative (griffin 1998), altman, king & verba developed one of the first open source (and open access) data publishing systems, the virtual data center (altman et al. 2001). this system successfully fielded the largest federated catalog of social science datasets in the world (altman et al. 2009). the virtual data center was designed to support persistent access to research data through federated institutional curation. data citation was deeply integrated into the virtual data center – each dataset managed was assigned a persistent identifier, and a citation. moreover, the virtual data center was based on the principle that all data supporting published research should be cited, and that these citations and identifiers should be machine-actionable through the web (e.g. through machine-actionable uri’s). nesstar, a system developed in parallel by ryssevik & musgrave (2001), and later used by many european archives, also incorporated the concepts of actionable web links, and persistent federated curation – although it did not initially support or emphasize citation. incorporating work by altman, et al. (2003) and altman and king (2007), the virtual data center incorporated both support for “deep citations” (buneman 2006) that identify precise subsets of a larger dataset; and for semantic fixity information that enables verification of a dataset using the citation itself. these capabilities were further extended in the dataverse network (king 2007), which succeeded the virtual data center. the dataverse network has since been adopted by the harvard university as its data publication infrastructure and is used by hundreds of researchers in dozens of institutions to curate and publish data. (crosas 2011, 2013) in parallel work, brase (2004) lead an initiative to systematically archive datasets associated with research outputs, and to systematically associate these datasets with digital object identifiers (doi, 1997) – a robust form of persistent identifier used in the publication community. this was the first step toward integration of data citation and data publication into the larger publishing ecosystem. to summarize, from 1977 through 2009 there were three phases of development in the area of data citation. • the first phases of development focused on the role of citation to facilitate description and information retrieval. this phase introduced the principles that data in archives should be described as works rather than media, using author, title, and version. • the second phase extended citations to support data access and persistence. building upon the principle that research data used in publication should be cited, this phase introduced the principles that those citations should include persistent identifiers, and that the citations should be directly actionable on the web. • the third phase of development focused on using citations for verification and reproducibility. although verification and reproducibility had always been one of the motivations for data archiving – it had not been a focus of citation practice. this phase introduced the principles that citations should support verifiable linkage of data and published claims, and it started the trend towards wider integration with the publishing ecosystem. the importance and urgency of scientific data management and access is now starting to be recognized broadly. many publishers recognized this, in theory, in 2006, when the “brussel’s declaration” put forth the principle that data associated with publications should be openly available. this same year, the u.s. national science foundation introduced a policy requiring every grant proposal to be accompanied by a data management plan. also that same year, data management was the theme of the annual meeting of the society of scholarly publishers, the premier conference in that field. this continues a trend of funders and publishers adopting data publication and management policies. universities have likewise become involved and have started to develop their own policies requiring data management, while journals, archives, and research libraries are increasingly grappling, largely independently, with the issues of data management. even the media has taken note. this is reflected by numerous articles drawing attention to particular high-profile cases of scientific fraud, such as the stapel affair (e.g., carey 2011), to increased rates of retractions (e.g. ionaddis 2005, steen 2010, fang iassist quarterly 2013 65 iassist quarterly et al 2012), and to the practice of open science more generally (e.g, lin 2011). the culmination of this trend, thus far, is an increasingly widespread consensus by researchers and funders of research that data is a fundamental product of research and therefore a citable product. the fourth and current phase of data development work focuses on integration with the scholarly research and publishing ecosystem. this includes integration of data citation in standardized ways within publication, catalogs, tool chains, and larger systems of attribution. it is exemplified by systems such as data dryad (vision 2010) and figshare (hahnel 2013) which integrate data deposition into publisher workflows, and datacite and the thomson reuters data citation index, which integrate data citations into index and discovery of other published work; and by community standardizations efforts, such as those coordinated by the national academies (uhlir 2012), codata (2013), and the data citation synthesis group (2014). across these various groups there has been a developing agreement over the years that an essential part of connecting research publications or claims to data is formal data citation that includes a persistent link to guarantee long-term data accessibility. global persistent identifiers, such as dois and handles, offer a mechanism to provide a permanent link that can be configured to always resolve to a web page from which the data can be accessed, independent of whether the location of that page changes over time. an increasing number of data repositories generate dois which can be directly used in a publication to reference the data. however, until now, there has not been a single set of principles or guidelines for data citations which represents and is in agreement with all these initiatives.3 what has emerged in the bibliographic and research community is a substantial core of agreement over the need for citation to support attribution and verification; the recognition that citations must support both human and machine clients; the existence of robust persistent identifiers and the understanding of the core role; and the publication of key reference documents such as the national academies and codata reports. converging data citation principles given the rise of these parallel, variously implemented initiatives on data citation, as well as the lack of unified guidance for publishers, journal editors, and funding agencies, there was a need for a synthesis set of general recommendations and good practices for data citation. in the summer of 2013, a synthesis group was formed to unify the various recommendations. it came to be known as the data citation synthesis group. it met weekly from july to november of 2013 to thoroughly deconstruct previous data citation principles defined by codata, the amsterdam manifesto, and datacite, and to produce a synthesis set that included the input of more than 25 organizations. during that time, the group met as part of the rda (research data alliance) conference in washington, dc in september, in two half days of public workshop. as a result, in november 2013, the proposed joint declaration of data citation principles was released to the public for open comment, and finalized at the end of february 2014 (data citation synthesis group, 2014) the scope of the synthesis principles is solely to provide data citation recommendations, and does not intend to include detailed specifications for implementation or to focus on technologies or tools or research data repositories. the principles should extend to all disciplines and all types of data. some of the challenges for specific types of data will be discussed in the next sections. as will be seen below, the joint declaration of data citation principles reflect the various efforts described in the last section and a broad convergence on core principles: 1. importance. data should be considered legitimate, citable products of research. data citations should be accorded the same importance in the scholarly record as citations of other research objects, such as publications. 2. credit and attribution. data citations should facilitate giving scholarly credit and normative and legal attribution to all contributors to the data, recognizing that a single style or mechanism of attribution may not be applicable to all data. 3. evidence. in scholarly literature, whenever and wherever a claim relies upon data, the corresponding data should be cited. 4. unique identification. a data citation should include a persistent method for identification that is machine actionable, globally unique, and widely used by a community. 5. access. data citations should facilitate access to the data themselves and to such associated metadata, documentation, code, and other materials, as are necessary for both humans and machines to make informed use of the referenced data. 6. persistence. unique identifiers, and metadata describing the data, and its disposition, should persist -even beyond the lifespan of the data they describe. 7. specificity and verifiability. data citations should facilitate identification of, access to, and verification of the specific data that support a claim. citations or citation metadata should include information about provenance and fixity sufficient to facilitate verifying that the specific timeslice, version and/or granular portion of data retrieved subsequently is the same as was originally cited. 8. interoperability and flexibility. data citation methods should be sufficiently flexible to accommodate the variant practices among communities, but should not differ so much that they compromise interoperability of data citation practices across communities. at the time this article was completed, less than a month after the principles had been finalized, they had been officially endorsed by thirty organizations, including many major publishers and data archives. the synthesis group has also committed to a dissemination plan that includes reaching out to a large number of stakeholders from multiple organizations and disciplines for an endorsement of the principles. we anticipate that the impact of the unified, widely broadcasted joint declaration of data citation principles will be substantial and will: change current publication workflows, create new data citation technologies, define new metrics for scholarly impact and recognition, and, more importantly, provide persistent access to the data supporting scientific results to validate and extend previous scientific work. the principles will facilitate interoperability across existing and new implementations, and will help guide enhancements and new versions of the current implementations. several data repositories are already compliant, or close to compliant, with these principles (e.g., dataverse, datadryad). in section five, we describe, as an example, the dataverse network data citation implementation. a generic example a generic example for a data citation can be represented as: 66 iassist quarterly 2013 iassist quarterly author(s), year, dataset title, global persistent identifier, data repository or archive, version or subset the authors and the data repository or archive elements directly support principle two, providing credit and attributions to the creators of the data as well as to their publishers or distributors. as in citations of literature, in some cases the creators are not individual authors, but instead an entity or organization that produced the data. also as in citations of literature, authorship can be challenging and ill-defined in a simplified citation format when there is a large number of individuals who have contributed to the scholarly product in a wide range of ways (e.g., from designing the instrument and software to cleaning and analyzing the data.) we already find these authorship challenges in publications in highenergy physics, such as articles related to the observation of the higgs boson having nearly 3,000 authors (e.g., cms collaboration, 2012). this wide array of authors is more common for data products than for articles. the principles and this citation example do not address the authorship problem, but, as described below, the metadata associated with the dataset can allow annotation of various levels of contribution during the creation and processing of the data, and also allows reference to related datasets or other scholarly products. the year in which the dataset is first published and the title are not directly related to a principle. however, these elements are common in traditional literature citations, and such consistent and informative formats contribute toward giving data citation the same importance as citations of other scholarly records, as stated in principle one. the global persistent identifier is an essential piece of the citation of a digital object and directly supports principle four. the persistent identifier or url allows separation of the link given in the citation with the url to which it resolves, thus guaranteeing that even if the hosting or location of the dataset’s web page changes, the link in the citation will always go to the same dataset page. in a forthcoming article by pepe, et al (2014), based on a study of 7,641 astronomy publications from four main astronomy journals, we show that 44% of the links in publications from ten years ago are broken. these are regular links to web sites, and not global persistent identifiers. the persistent identifier or url solves a technical problem, but it is not sufficient without a publisher that supports and guarantees the validity of its persistent identifiers. in the case of data, the publisher is usually the data repository or archive. the more commonly used global persistent identifiers are handles (sun, et al. 2003) and dois (paskin, 2002). the persistent identifier in the data citation example also supports principles five and six. in support of principle five, the handle or doi should resolve to a dataset page, which contains sufficient information describing the data and facilitating their reuse. in the rare cases in which the data cannot be made accessible any longer or must be destroyed, the data citation should still be valid. that is, the persistent identifier should resolve to a page with information about the discontinuation of that dataset (principle six). the last element in the generic citation example is the version, subset, or timestamp, which supports principle seven. this element is particularly relevant when citing data. contrary to most literature publications, a dataset is often altered or expanded with time. the frequency with which a dataset might be changed can vary, from a static dataset that never changes once published, to a dataset that is updated once in a while with a new version, to datasets that are constantly changing, as is the case of dynamic data from meteorological sensors or streaming data twitter feeds that grow constantly over time. dynamic and streaming data offer a number of challenges for both citation and replication of published results contingent upon the reuse of a specific version of a dataset. those challenges are described in section six. the generic citation example might vary in style from community to community (principle eight), but across all cases it should be considered as important as other citations and should be part of either the standard reference section of a publication or a similar section for data citations, in accordance with principle one. the data citation synthesis group also recommends that when a published claim is made based on the data, enough information should be provided in the text to identify the data citation listed in the reference section, in the same fashion as other citations. when the published work makes a claim based on a subset of the data, specific information about the subset should be referenced by that claim. due to the possible complexity of such a citation, it is not always feasible to include in the reference section all the information needed to fulfill the core data citation principles. for this purpose, as stated in principle four, an important component of any data citation is machine-actionable metadata that is bound to the data citation and persists with it. for example, the datacite metadata schema and ontology (datacite 2013) describe a detailed set of fields that may be used to complete a data citation. typically, additional fixity and provenance information is required to support the verification requirements – such that future users of the citations can ensure that the data they use is identical to that cited. such information might include bit-level fixity information (such as a md5, sha-256 or other cryptographic hash), or preferably, where available, semantic fixity information (such as a unf or perceptual fingerprint). additional information on contributors will be required to fulfill the attribution requirements wherever the authors explicitly listed in the reference are ambiguous or incomplete. unstructured metadata such as a contributors list may fulfill the bare legal requirements for attribution; however, structured name authority or identifiers such as orcid’s (open research identifier) or isni’s (international standard name identifiers) are much preferred, because they facilitate scholarly attribution (credit). this information can be embedded in published documents in machine-accessible form, included in the metadata stored with the doi or other persistent identifier by its resolver service, or stored in an associated community index, such as crossref or datacite. such metadata should also be presented through the landing page provided to humans when the persistent identifier for the data is resolved implementing the state of the practice data repositories, or data publishers, are often responsible for implementing and generating data citations for the datasets hosted within them. as noted above, there are a number of repositories that are already generating data citations upon deposit of a dataset, and those citations are often compliant with the principles above (e.g., dryad, dataverse, figshare). the generic data citation example in section four is based on the citation format generated by the dataverse network software application. this application is a data repository platform that allows organizations to host dataverses, where each dataverse iassist quarterly 2013 67 iassist quarterly contains datasets, and where each dataset contains data files and metadata. a dataverse is, in essence, a virtual archive, which can be branded and administered individually, giving control to the data owner or distributor, while its data and metadata are stored by the repository in accordance to professional archival practices, metadata standards, and preservation formats (king 2007, crosas, 2011). the software is open-source and developed at the institute for quantitative social science at harvard university (king, 2014). the harvard dataverse is one of the dataverse network instances open to all researchers and to all data types. it supports a variety of types of dataverses, from journal dataverses, to dataverses for individual researchers, to dataverses for data associated with an institutional department (crosas, 2013). in this section, we describe the implementation of data citation as it is built in dataverse version 4.0. when a new dataset is added to a dataverse, the required metadata fields that must be entered by the depositor include the author(s) or producer organization and the dataset title. in addition, an extensive set of metadata fields are provided, some required and others optional. the citation metadata supported by dataverse maps closely to the datacite metadata, and can also be mapped to the format developed by the data documentation initiative (ddi, <http://www.ddialliance.org/specification/>) and dublin core metadata initiative terms (dcterms, <http://dublincore.org/ documents/dces/>). the dataset, when created, is in a draft form that is unpublished, and data files and additional metadata can be added at a later time. upon dataset creation, however, even if the dataset is not yet published, a draft data citation is instantly generated following these steps: 1. authors and title are obtained from the metadata fields entered by the data depositor. if instead of individual authors, a producer (organization or institution) is entered, the producer is used in place of the authors. 2. in the draft citation, the year is automatically populated by the year of deposit. at the time when the dataset is released, the final citation is updated with the year of the released or published date, which is often, but not always, the same year the dataset was deposited. 3. the dataverse network software supports both handles and dois as persistent identifiers. if a dataverse network is configured to use handles, each handle is registered to the handle system. the harvard dataverse is configured to use dois, which are registered to datacite through the ezid api (<http:// ezid.cdlib.org/home/documentation>). upon deposit, the dataset is registered with status “reserved”, an option provided by the ezid api. when the dataset is released, the status becomes “public”. this means that the doi at that point resolves to a public dataset page, which includes description information about the dataset, as well as information on how to access the data. even when data cannot be completely open, and one or more data files in the dataset are restricted due to data user agreements or confidential information, the doi resolves to a dataset page where access can be requested. 4. the publisher or data repository element in the citation is automatically populated as the repository name, in this case, the harvard dataverse. if additional distributors or archives are responsible for those data, they can be listed in the dataset page, as part of the additional metadata. 5. the dataverse network software supports versioning of datasets because, unlike traditional literature publications, data are often updated even after being published. the data citation generated by dataverse includes the version of the dataset. when the dataset is released, the version in the citation is set to 1. if the dataset metadata or files are updated in the future, a new version is created, and a new citation, with the same doi, but a new version number, is created. this allows reference to a specific previous version, and access to that version from the dataset page within a dataverse. it is important to note that a doi or other persistent identifier is not equal to a data citation. the data citation is the composition of all the elements that form it, and the doi is one of these elements. therefore, one can cite two versions of a dataset with the same doi, as long as the citation provides unambiguous information about the version. this is similar to citing a subset of the entire dataset, or in other type of citations, citing a set of pages in a book. the data citation generated by the dataverse network software also supports universal numerical fingerprints (unf) for tabular datasets (altman and king, 2007). the unf guarantees fixity; it’s a unique fingerprint on the semantics of a dataset. that is, even if a dataset changes format, if the data values remain the same, the unf remains the same. when a unf cannot be calculated, the dataverse calculates bit-level fixity information (the md5) of the data file(s) contained in the dataset. the dataverse network implementation is fully compliant with the data citation principles discussed throughout this article. however, it does not support, in its current form, dynamic or streaming data. this is discussed in more detail in the next section. remaining challenges at the broadest conceptual level, the substantial remaining challenges for implementing robust data citation systems fall into three categories:4 • challenges of provenance. provenance includes the chain of ownership of an object, and the history of transformations applied to it. models of provenance have strong implications for how data citation is integrated into the data curation workflow. • challenges of identity. these theories involve defining ‘data’ themselves, the identity of data and how to define equivalence and derivation relationships, and the granularity and structure of data. theories of data have strong implications for determining what should be cited. • challenges of attribution. attribution plays a key role in the incentives for citation. models of attribution have strong implications for determining the presentation of data citations. provenance is a particularly important concern because many data citations are used to document a direct evidentiary relationship between a published assertion and the underlying evidence that supports it. however, supporting this evidentiary relationship does not require recreating or establishing the entire provenance chain – and much of provenance can be considered as orthogonal to citation, as groth (2012) argues. notwithstanding, as smith (2012) points out, enabling readers to establish authenticity of the cited object is an important use for citation and requires that citation be connected to provenance information. the maintenance of this connection and of the associated provenance information is a major challenge for developing reliable citable scientific workflows. identity is close to the heart of creating a citation. to cite something requires it to be identified – the citation should enable 68 iassist quarterly 2013 iassist quarterly us to find the same thing that was used in the citing article. identity is relatively straightforward for immutable data in the original formats and used as a whole. however, when data that changes over time is manifested in different formats, or is used only in part, a number of practical questions emerge: • the equivalence question. how does one determine whether two data objects, not bitwise identical, are semantically equivalent (interchangeable for scientific computation and analysis)? • the versioning question. how does one unambiguously assign, at the time of citation, a ‘version’ to a data object, such that someone referencing the citation later can retrieve or recreate the data object in the same state that it was at the time of citation? • the granularity question. how does one unambiguously describe components and/or subsets of a data object for purposes of computations, provenance, and attribution? how does one incorporate this granularity with a bibliographic data citation to create a “deep” citation? although there are no complete solutions to these problems, a number of promising approaches are emerging. these approaches include: systematic identification of the “significant properties” of digital objects – those attributes that are used in later substantive/ semantic interpretation of the object (hedstrom and lee, 2002); creation of semantic fingerprints for data objects, such as unf’s (altman et al., 2003, 2008), which compute cryptographic hashes over canonicalized representations of an object; and perceptual fingerprints, which characterize uniquely the way that a data object is perceived (cano, et al. 2004). algorithms are being developed for generating persistent granular citations of specific forms of dynamic data objects, particularly of databases.5 moreover, open annotation frameworks and ontologies are being developed to allow interoperable annotation of digital objects that define spatial (logical) and temporal granularity which might be used generally to complement bibliographic data citations and support deep citation (van de sompel, 2012). natural corollaries to these questions involve considerations of scalability. for example, how does one track and recreate versions of large and dynamic databases? what data structures enable fine-grained access to data? how does one compute equivalence over the members of large collections for the purposes of de-duplication? a third challenge is that of attribution. citation should support unambiguous attribution of credit for all contributors. as the scale of the data increases, and more people contribute to its creation and maintenance, practical challenges with attribution arise. these include supporting attribution for contributors that may number in the hundreds of thousands in crowd-based citizen science (e.g. wiggins and crowston 2011), distinguishing among different contributor roles (iwcsa 2012), and capturing the nature of the relationship between the cited and citing objects (e.g. cronin 1984)6 summary scientific data are increasingly being made available online. lowering barriers to discovery and use of these data, and increasing our ability to link data with publications have the potential to enable new forms of scholarly publishing, promote interdisciplinary research, strengthen the linkage between policy and science, and lower the costs of replicating and extending previous research. robust data citation practices and infrastructure will play a critical role in achieving these outcomes. bibliographic standards for cataloging data developed gradually from the early days of data archives but it was a long time before academic citation practices started to catch up with archiving practices. over four decades ago, however, several core principles for data citation and bibliographic description were recognized – in part based on the pioneering work of sue dodd. for the next 25 years, data citations had little attention from or impact on either the scientific or library community – despite the fundamental soundness of many of the early principles and the implementation of citation practices by selected major data repositories. more recently data citation principles and practices have made a resurgence – fueled both by advances in web and network technologies and by a growing public and scientific recognition of the importance of scientific reproducibility, data sharing, and reuse. recently, a wide convergence on principles has emerged, and the deployment of production infrastructure to support data citation across the research lifecycle is rapidly advancing. key enablers of a successful synthesis process have included a substantial core of agreement concerning the need for citation to support attribution and verification; the recognition of the need for citation to support both human and machine clients; the existence of robust persistent identifiers and the understanding of their core role; and the publication of key reference documents such as the national academies and codata reports. a number of central challenges remain, particularly related to the frontiers of data – big data, complexly structured data, dynamic data, and data in changing formats. these are being addressed gradually through groups such as rda and through state-of-thepractice development of systems such as the dataverse network. acknowledgments we would like to thank the members of the data citation synthesis task group and of the co-data data citation working group for commentary on this paper and on the ideas leading into it: amy brand, amye kenall, andras rauber, anita dewaard, bonnie caroll, christinge borgman, dan cohen, david shotton, eefke smit, elizabeth arnaud, elizabeth iorns, fiona murphy, franciel linares, giri palanisami, hannelore vanhaverbeke heige sagen, hylke koers, ivan herman, jan brase, jianhui li, jo mcentyre, joan starr, joe hourcle, john helly maren morgenroth, kathleen cass, kerstin lehnert, koji zettsu, mark hahnel, mark parsons, martie van deventer, maryann martone, michael diepenbroek, michael wit, mustapha mokrane, natalia moanola, paul groth, paul uhlir, phil archer, puneet kishor, ruth duerr, sarah callaghan, simon hodson, stefan proell, stephanie hagstom, tim clark, tim smith, todd carpenter, vishwas chavan, yannis ionnadis, yasuhrio muryama references alsheikh-ali, a. a., w. qureshi, m.h. al-mallah, & j.p. ioannidis. (2011). “public availability of published research data in high-impact journals.” plos one, 6(9), e24357. < http://www. plosone.org/article/info%3adoi%2f10.1371%2fjournal. pone.0024357#pone-0024357-g001> altman, m. (2008) “a fingerprint method for scientific data verification.”advances in computer and information sciences and engineering. springer netherlands: 311-316. iassist quarterly 2013 69 iassist quarterly altman, m., l. andreev, m. diggory, g. king, a. sone, s. verba, and d.l l. kiskis. (2001) “a digital library for the dissemination and replication of quantitative social science research the virtual data center.” social science computer review 19(4): 458-470. altman, m., j. gill, and m.p. mcdonald. (2003). numerical issues in statistical computing for the social scientist. john wiley & sons. altman, m, and g. king. (2007) . “a proposed standard for the scholarly citation of quantitative data.” d-lib magazine 13.3/4. <http://www. dlib.org/dlib/march07/altman/03altman.html> altman, m., m.o. adams, j. crabtree, d. donakowski, m. maynard, a. pienta and c.h. young. (2009). “digital preservation through archival collaboration: the data preservation alliance for the social sciences.” american archivist 72, no. 1: 170-184. <http://archivists.metapress. com/content/eu7252lhnrp7h188/fulltext.pdf> avram, h. d. (1975). marc, its history and implications. washington: library of congress. bisco, r. l. (1965). “social science data archives: technical considerations.” social science information 4:3, 129-150. borgman, c. (2012) “why are the attribution and citation of scientific data important?” in p. f. uhlir, (ed.), for attribution: developing scientific data attribution and citation practices and standards: summary of an international workshop (pp. 1-10). washington, d.c.: national academies press. brase, j. (2004) “using digital library techniques–registration of scientific primary data.” in research and advanced technology for digital libraries, pp. 488-494. springer berlin heidelberg. buneman, p. (2006). “how to cite curated databases and how to make them citable.” proceedings of the 18th international conference on scientific and statistical database management (pp. 195-203). los alamitos, ca: ieee computer society. cano, e. batle, t. kalker, j. haistma, (2002) “a review of algorithms for audio fingerprinting”, ieee workshop on multimedia signal processing, ieee press:169173. carey, b. (2011), “fraud case seen as red flag for psychology research.” new york times, a3. november 3, 2011. cms collaboration, (2012) “observation of a new boson at a mass of 125 gev with the cms experiment at the lhc.” physics letters b, volume 716, issue 1, pages 30-61. codata/itsci task force on data citation, (2013). “out of cite, out of mind: the current state of practice, policy and technology for data citation.” data science journal 12: 1-75., <http://dx.doi.org/10.2481/ dsj.osom13-043> cronin, blaise. (1984)the citation process. the role and significance of citations in scientific communication. london: taylor graham. crosas, m. (2011). “the dataverse network: an open-source application for sharing, discovering and preserving data.” d-lib magazine 17 (1–2). <http://www.dlib.org/dlib/january11/ crosas/01crosas.html> crosas, m. (2013). “a data sharing story.” journal of escience librarianship 1 (3):173–79. <http://escholarship.umassmed.edu/ jeslib/vol1/iss3/7/> crosas, m., t. carpenter, c. borgman, d.m. shotton. (2013). “the amsterdam manifesto on data citation principles.” force11. data citation synthesis group, (2014). joint declaration of data citation principles, <http://www.force11.org/datacitation> datacite, (2013). “datacite metadata for the publication and citation of research data” doi:10.5438/0008 dewald, w.g., j.g. thursby, and r.g. anderson. (1986). “replication in empirical economics: the journal of money, credit and banking project.” american economic review, 76(4):587-603. dodd, s. a. (1979) “bibliographic reference for numeric social science data files: suggested guidelines.” american society for information science journal 30:2, 77-82. doi, (1997). doi handbook <http://www.doi.org/hb.html> fang, f. c., r.g. steen and a. casadevall. (2012). “misconduct accounts for the majority of retracted scientific publications.” proceedings of the national academy of sciences, 109(42), 17028-17033. fienberg, s. e., m.e. martin and m.l. straf. (1985). sharing research data. washington, d.c.: national academies press. griffin, s. (1998). “nsf/darpa/nasa digital libraries initiative.” d-lib mag, 4(7). <http://www.dlib.org/dlib/july98/07griffin.html> groth, p. (2012). “maintaining the scholarly value chain: authenticity, provenance, and trust.” in p. f. uhlir, (ed.), for attribution: developing scientific data attribution and citation practices and standards: summary of an international workshop, (pp. 31-42). washington, d.c.: national academies press. hahnel, m. (2013) “referencing: the reuse factor.” nature 502.7471: 298. hamermesh, d.s. (2007). “viewpoint: replication in economics,” canadian journal of economics. hedstrom, m. and c. lee (2002). “significant properties of digital objects: definitions, applications, implications.” proceedings of the dlm-forum: parallel session : 218-113. isbd (1990). international standard bibliographic description for computer files. recommended by the working group on the international standard bibliographic description for computer files set up by the ifla committee on cataloguing. isbn 0-903043-56-4 ioannidis, j. p.a. (2005). “why most published research findings are false.” plos medicine 2.8: e124. iso, (1997). information and documentation -bibliographic references -part 2: electronic documents or parts thereof. 690-2:1997. international standards organization. 70 iassist quarterly 2013 iassist quarterly iwcsa report. (2012). report on the international workshop on contributorship and scholarly attribution, may 16, 2012. harvard university and the wellcome trust. available at: <http://projects. iq.harvard.edu/attribution_workshop>. king, g. (1995). “replication, replication.” ps: political science and politics 28.3: 444-452. king, g. (2007). “an introduction to the dataverse network as an infrastructure for data sharing.” sociological methods and research 36 (2): 173–99. king, g. (2014). “restructuring the social sciences: reflections from harvard’s institute for quantitative social science.” ps: political science and politics 47, no. 1: 165-172. lin, t., (2011) “cracking open the scientific process.” new york times, d1. january 17, 2011. paskin, n. (2002). “digital object identifiers.” information services and use 22.2: 97-112. pepe, a., a. goodman, g. muench, m. crosas, c. erdmann (2014). “sharing, archiving and citing data in astronomy” plos one (forthcoming). pienta, a. (2006). “leads database identifies at-risk legacy studies.” icpsr bulletin 27(1). piwowar, h. and t. vision (2013). “data reuse and the open data citation advantage” peerj pollak, o. b. (2006). “the decline and fall of bottom notes, op. cit., loc. cit., and a century of the chicago manual of style.” journal of scholarly publishing 38.1: 14-30. proll, s. and a. rauber (2013). “scalable data citation in dynamic, large databases: model and reference implementation.” big data, 2013 ieee international conference on. ieee. ryssevik, j. & s. musgrave (2001). “the social science dream machine: resource discovery, analysis, and delivery on the web” social science computer review 19(2) 163-174. science (2014). general information for authors. retrieved from <http://www.sciencemag.org/site/feature/contribinfo/prep/gen_ info.xhtml> sieber, j. e. (1991). sharing social science data: advantages and challenges. sage publications, inc. smith, m. (2012), institutional perspectives on credit systems for research data. in p. f. uhlir, (ed.), for attribution: developing scientific data attribution and citation practices and standards: summary of an international workshop (pp. 77-80). washington, d.c.: national academies press steen, r.g. (2010). “retractions in the scientific literature: is the incidence of research fraud increasing?” journal of medical ethics 37: 1-5. sun, s., l. lannom, and b. boesch (2003). “handle system overview.” rfc 3650, november, 2003. uhlir, p. f., (ed.) (2012). for attribution: developing scientific data attribution and citation practices and standards: summary of an international workshop. washington, d.c.: national academies press. van de sompel, h. (2012), “data citation technical issues identification” in uhlir, p. f., (ed.) (2012). for attribution: developing scientific data attribution and citation practices and standards: summary of an international workshop. washington, d.c.: national academies press. van leunen, m. (1992). a handbook for scholars. new york, ny: oxford university press. vines, t. h.; a.y.k. albert, r.l. andrew, f. d barre, d.g. bock, m.t. franklin, k.j. gilbert, j-s moore, s. renaut, d.j. rennison (2014). “the availability of research data declines rapidly with article age” current biology 24 (1): 94 97. vision, t. j. (2010). “open data and the social contract of scientific publishing.”bioscience 60, (5): 330-331. wiggins, a., and k. crowston (2011). “from conservation to crowdsourcing: a typology of citizen science.” system sciences (hicss), 2011 44th hawaii international conference on. ieee. notes 1. authors are listed alphabetically; the authors have made equal contributions to this work. micah altman is director of research, mit libraries at the massachusetts institute of technology. he can be contacted at <escience@mit.edu>. mercè crosas is director of data science, institute for quantitative social science at harvard university. she can be contacted at <mcrosas@iq.harvard.edu>. 2. or more precisely, in some cases it is a “club good” – nonconsumptive and only partially excludable. 3. efforts in this area have been made by codata, as part of an extensive report on data citation (codata 2013), datacite principles, dcc as part of the core guidelines on data curation, harvard’s institute for quantitative social science through a data citation workshop hosted in 2012, and force11 in the form of the amsterdam manifesto for data citations principles, born at the beyond the pdf 2 conference in 2013 (crosas et al, 2013), among others, and by multiple research data repositories that offer to generate data citation upon deposit of a dataset (such as dataverse, datadryad, figshare, and the inter-university consortium for political and social research (icpsr)). 4. this section in part summarizes and updates section 7.2 in the codata report (2013), which was originally written by one of the authors of this article. 5. see buneman (2006) for fundamental work in this area; also proll and rauber (2013) for a more recent approach. 6. cronin (1984) reviews over 10 different proposed taxonomies of citation types and roles, some of which identify dozens of individual relationships. vol252 14 iassist quarterly summer 2001 a little bit of history the data archiving process, in romania, can be best described as an on-going process. we have great expectations and we are doing our best in order to see this happening. but our task is not an easy one, dealing in a more or less close environment. if some western european countries have 30 or 40 years of history in data archiving, the romanian saga is just beginning. before 1990, the concept of a public data archive was unconceivable. data were considered secret, with a limited access for only a few granted people. comparing to that period, after the 1989 revolution there was an ‘explosion’ in social research, with valuable data being collected. unfortunately, the old instincts are still in the system. research institutes, private and public, state institutions (like the national commission for statistics), even nongovernmental organizations still promote a close system policy, even if this means taking the risk of loosing the data (in many cases, they are not aware of that risk at all). a very good example that shows this kind of mentality is a story that happened to a fellow researcher. the national commission for statistics distributes a free booklet with information. the trouble is the booklet rapidly disappears because of the small number printed. when the researcher was told there were no more copies available, he asked ‘öbut why don’t you put them on the web?’. the answer came promptly: ‘no, no, this would mean to make them public!’. in 1990, the institute for quality of life research (iqlr), bucharest was founded. taking this institute for example, there is an increasing number of studies being pursued and a lot of data collected. naturally, problems of a comparative nature began to appear. feeling the potential of a data archive, following western examples, prof. catalin zamfir, ph.d. – the director of the institute (now the dean of the faculty of sociology and social work, as well) quickly realized that a romanian data archive was necessary. about three or four years ago, a team of researchers was formed, with the specific task of creating a data archive. they visited a couple of western countries with tradition, coming back full of ideas and enthusiasm. unfortunately, partly because of the poor romanian environment, a very sad but common event happened: many of them left the institute, for better paid jobs in the private sector or even leaving the country with the purpose of emigration. the remaining researchers lost their interest and the project got stuck, until recently. a new wave of young graduates from the faculty of sociology and social work were recruited, and new hopes are being raised. having the support of both iqlr and the faculty, these premises are very encouraging. a romanian social data archive (rsda) was set up, currently being under construction, and some visible results are expected in the next couple of months. institutional background the rsda will function as part of the iqlr, which is an academic research institute founded in 1990 under the aegis of the romanian academy for sciences and the national institute for economic research. the iqlr has relations of scientific cooperation with a large number of domestic and foreign institutions, promoting a sustained opening towards mass media. by the quality of its research activity and by the results it obtained, the institute earned its recognition both at national and international levels such as the presidency and government of romania, ministries, from other domestic and foreign research institutes and universities, as well as from several international organizations such as the council of europe, unicef, etc. starting with 1990, the institute publishes on an annual basis, the diagnosis of the quality of life in romania and evaluates the social policies adopted in our country during the period of transition. it also investigates and offers alternative solutions to the main social and economic problems of present day romania, by conducting empirical research on local national samples. the institute publishes books, studies, reviews, brochures, and research reports and provides consultancy in its field of expertise. the research activity of the institute is structured on five main directions as follows: • quality of life a data archive for social sciences in romania by adrian dusa 1 iassist quarterly summer 2001 15 • social policies • disadvantaged groups • human and communitarian development • interethnic relations the papers elaborated by the iqlr are oriented towards both fundamental and applied research in social sciences and quality of life, adhering to the values of the spirit of european integration and transition to a free-market oriented economy in romania. apart from this valuable support provided by the institute, the rsda receives financial and logistic support from the faculty of sociology and social work (the romanian acronym is sas), university of bucharest. the sas brought a great help in setting-up the romanian social data archive. apart from the financial support from the institute that covers the salaries, we received on behalf of the faculty the rights for using a network of 10 computers in one of the faculty’s establishments. to maintain this network, the faculty bought a server with a minimal configuration. so what do we do about it? setting up a data archive is a challenging thing to do. as already mentioned, the romanian data archive is under construction. practically, we only had a vague idea about what data archiving is, and a collection of data sets. in this stage, information is the key word, because we have to deal with some problems – some of them probably basic for long established data archives, but still problems for us. building a data archive from scratch means that we need answers for some elementary questions, like: – what exactly is a data archive? – how do we store data? – how do we build the metadata base? – how do we put it on the web? – what kind of procedures should we use? – what kind of software? of course, we didn’t have those answers; so in order to get a good idea about how to do data archiving we maintain a close contact with the other european data archives. their experience is very helpful, because they have already been here while ago. the good thing is that we do not have to ‘reinvent the wheel’ again, but to take advantage of the information they can provide and ‘jump’ to the construction phase. the first major step is to be recognized as an institution on the international level; and in order to do this, applications to cessda and ifdo will be pursued. we estimate this happening somewhere in the beginning of 2002, when a workshop for new data archives is planned, which will take place in berlin. later on we plan to apply for membership at icpsr as well. these international connections are very important indeed, as we will be able to have access to the enormous social science heritage in the world. we could not only provide access to the romanian data sets, but to the other international ones that are similar. although we are working on our archive’s development, the process is rather slow. there are several reasons; one of them is that our team is very small, with only three members, plus a part-time system administrator. all the tasks are distributed among these persons. we have to do everything on our own, and the lack of training is slowing the process even more. we will have to spend some time and money with staff training. defining procedures using international standards cannot be made until the staff fully understands what those standards are and how do they work. another reason is the lack of money. although the institute does pay us, there is no money for some professional help. for example, the team’s members are mainly sociologists, having no experience in web-design. we are currently working on the web page design, but i fell that we are loosing valuable time doing things that we are not trained for, instead of doing really important things. we will eventually do it, but with the cost of loosing time. one other reason (and i will continue no more after) is that the whole romanian society is inert. it is very hard to move fast forward in a static society (or even worse, moving backward). we have to adapt our speed to the state institutions’ speed, hoping that private research institutes will move faster. our plans after the system will be set and available data sets archived, we plan an extensive number of visits to the other research institutions, trying to bring more data sets in. things are easier now for long established data archives in the western europe, because the depositors are coming with their data to be archived. here, we will have a tough job persuading data sets’ owners that is worth depositing; until they will trust us, we will have to run and ‘chase’ data sets. the we will be able to offer an up to date catalogue with the romanian data sets archived, increasing high quality social research in the academic community. another future plan is to get some additional funds, from national and international sources. for a short term, international funds could be a solution, but in a longer perspective, funds from national institutions are necessary. we plan to attract funds from the romanian academy for sciences, from the education and research ministry, and from any other source possible. finally, international affiliations will bring the recognition we need in order to say: ‘yes we do have a romanian social data archive’. we would like to bring our thanks to the uk data archive and zentralarchiv f¸r empirische sozialforschung, for their valuable training and support they provided. without their help, nothing we have would have been possible. 1. contact: adrian dusa, institute for quality of life research 13 septembrie nr. 13, bucharest, romania. email: adi@iccv.ro vol271p.indd ����������������������������������������������������������������������������������������� ������������������������������������������������������������������������������������������������������������� �������������������������������� �������������� ��������������������� ���������������������������������� ��������������������������� ������������ ������������������������������������������� ��������������������������������������� ���������������������������������������� ������������������������������������� ����������������������������������� ����������������������������������������� �������������������������������������� ���������������������������������������� ��������������������������������������������������������� ���������������������������������������������������������� �������������������������������������������������������������� �������������� �������������������������������������������������������� ����������������������������������������������������� ����������������������������������������������������� ���������������������������������������������������������� ������������������������������������������������������������� ��������������������������������������������������� ��������������������������������������������� ������������������������������������������������������ ������������������������������������������������������������ �������������������������������������������������������� ������������������������������������������������������ ������������������������������������������������������ ������������������������������������������������������������� ���������������������������� ������������������������������ �������������������������������������������������� ������������������������������������������������������ ���������������������������������������������������������� �������������������������������������������������������������� ���������������������������������������������������������� ������������������������������������������������������� ���������������������������������������������������������� ��������������������������������������������������� ��������������������������������������������������������� ������������������������������������������������������� �������������������������������������������������������� ��������������������������������������������������������� ����������������������������������������������������������� ��������������������������������������������������������� ����������������������������������������������������������� �������������������������������������������������������� ���������������������������������������� ����������������������������������������� ��������������������������������������� ����������������������������������������� ��������������������������������������� ���������������������������������������� ������������������������������������������� ������������������������������������ ������������������ ������������������������������������������������������� �������������������������������������������������������� ��������������������������������������������������� ������������������������������������������������ ������������������������������������������������������ ������������������������������������������������������� ������������������������������������������������������ �������������������������������������������������������� ����������������������������������������������������������� ����������������������������������������������������������� ������������������������������������������������������� ����������������������������������������������������������� ������������������������������������������������������������ ������������������������������������������������������������ ������������ ������������������������������������������������������������� �������������������������������������������������������� �������������������������������������������������������� ��������������������������������������������������������� ����������������������������������������������������� ����������������������������������������������������� ��������������������������������������������������������������� ������������������������������������������������������� ����������������������������������������������������������� ��������������������������������������������������������� ����������� ��������������������������������������������������� ��������������������������������������������������� ���������������������������������������������� �������������������������������������������������������� ������������������������������������������������������� ��������������������������������������������������������� ��������������������������������������������������������� ���������������� � ���������������������������������� �������������������������� ������������������������������������������������������������� ������������� ��������������������������������� � ������������������������������������������������������ ��������������������������������������������������������� ������������������������������������������������������� ��������������������������������������������������������� ������������������������������������������������������������ ����������������������������������������������������������� ���������������������������������������������������������� ��������������������������������������������������������� ��������������������������������������������������������� ��������������������������������������������������������� ������������������������������������� �������������������������������������������������������������� ������������������������������������������������������������ ���������������������������������������������������������� �������������������������������������������������������������� ������������������������������������������������������� ������������������������������������������������������������ �������������������������������������������������������� �������������������������������������������������������� �������������������������������������������������������� ������������������������������������������������������������ ������������������������������������������������������������ ������������������������������������������������������������ ������������������������������������������������������������� ������������������������������������������������������������� ���������������������������������������������������������� �������������������������������������������������������� ����������������������������������������������������������� ��������������������������������������������������������������� ������������������������������� ���������������������������������������������������� �������������������������������������������������� ������������������������������������������������������� ���������������������������������������������������������� ������������������������������������������������������ �������������������������������������������� ����������������������������������� ����������������������������������������������������������� ����������������������������������������������������������� ���������������������������������������������������� ������������������������������������������������������ ������������������������������������������������������������ ����������������������������������������������������������� ������������������������������������������������������������ ������������������������������������������������������������� ���������������� ������������������������������������������������������ �������������������������������������������������������� ���������������������������������������������������������� ������������������������������������������������������ ����������������������������������������������������������� ����������������������������������������������������������������� ����������������������������������������������������������� ���������������������������������������������������������������� �������������������������������������������������� ������������������������������������������������������������ ������������������������������������������������������ �������������������������� ������������������������������������������������������ �������������������������������������������������������� ���������������������������������������������������������� ��������������������������������������������������������� ���������������������������������������������������������� �������������������������������������������������������� ��������������������������������������������������������������� ���������������������������������������������������������� ���������������������������������������������������������� ��������������������������������������������������������� ������������������������������������������������������������ ������������� �������������������������������������������������������� ������������������������������������������������ ������������������������������������������������ ������������������������������������������������������ ���������������������������������������������������� ��������������������������������������������������������� �������������������������������������������������������� ������������������������������������������������������ ��������������������������������������������������������� ������������������������������������������������������� ��������������������������������������������������� ��������������������������������������������������� ��������������������������������������������������������� ���������������������������������������������������������� ������������������������������������������������������������ ���������������������������������������������������������� ������� ��������������������������������������������������������� ����������������������������������������������������������� ���������������������������������������������������������� ���������������������������������������������������������� ����������������������������������������������������������� �������������������������������������������������������� ������������������������������������������������������� ����������������������������������������������������� ������������������������������������������� ������������ ���������������������������������������������������������� ������������������������������������������������������ ����������������������������������������������������� ������������������������������������������������������������ ���������������������������������������������������������� ���������������������������������������������������� ������������������������������������������������������ ������������������������������������������������������ ����������������������������������������������������� �������������������������������������������������������������������������������������� ����������������������������� ���������������������������������������������������������������� �������������� ���������������������������������� ����������������������������������������������������������� ��������������������������������������������������������� ���������������������������������������������������������� ������������������������������������������������������ ���������������������������������������������������������� ����������������������������������������������������� ���������������������������������������������������������� ��������������������������������������������������������� ��������������������������������������������������������� ���������������������������������������������������������� ���������������������������������������������������������� ����������������������������������������������������� ���������������������������������������������������������� �������������������������������������������������������� ��������������������������������������������������� ������������������������������������������������������� �������������������������������������������������������� ������������������������������������������������������� ��������������������������������������������������������� ������������������������������������������������������������� ���������������������������������������������������������� �������������������������������������������������������������� ���������������������������������������������������� ��������������������������������������������������������� ������������������������������������������������������� ��������������������������������������������������������� ��������������������������������������������������� ���������������������������������������������������������� ������������������������������������������������������������� ������������������������������������������������������� ����������������� ������������������������������������������������������������ ������������������������������������������������������ ������������������������������������������������������������� ��������������������������������������������������������������� ������������������������������������������������������������� ��������������������������������������������������������� �������������������������������������������������������������� ������������������������������������������������������������ �������������������������������������������������������� ����������������������� ���������� ����������������������������������������������������� �������������������������������������������������������� ������������������������������������������������������������� �������������������������������������������������������� ���������������������������������������������������������� �������������������������������������������������������������� �������������������������������������������������������� �������������������������������������������������������������� ������������������������������������������������������������ ������������������������������������������������������� ��������������������������������������������������������� ��������������������������������������������������������������� ����������������������������������������������������������� ����������������������������������������������������� ����������������������������������������������������������� ������������������������������������������������������������ ������������������������������������������������ �������������������������������������������������������� �������������������������������������������������������� ������������������������������������������������������������� ��������������������������������������������������������������� ������������������������������������������������������������ ���������������������������������������������������������� �������������������������������������������������������� ����������������������������������������������������������� ���������������������������������������������������������� ������������������������������������������������������ ������������������������������������������������� ���������������������������������������������������� ���������������������������������������������������� ���������������������� ���������� ������������������������������������������������������� �������������������������������������������������������� ������������������������ ������������������������������������������������������������ ������������������������� ������������������������������������������������������� �������������������������������������������������������� ������������������������������������������������������ �������������������������������������������������������� �������������������������������������������������������� ���������������������������������������������������� ������������������������������������������������� ����������������������������������������������������� ��������������������������������������������������� �������������������������������������������������� ����������������������������������� ������������������������������������������������� ���������������������������������������������������� ����������������������������������������������������� ���� ������������������������������������������������� ������������������������������������������������������ ������������������������������������������������������ ����������������������� ��������� � ���������������������������������� �������������������������� ������������������������������������������������������������� ������������� ��������������������������������� � ��������������������������������������� ���������������������������������������������������� ������������������������������������������������������� ���������������������������������������������������� �������������������������������������������������������� ��������������������������������������������������������� ��������������������������������� �������������������������������������������������������������������������������������� ����������������������������� ���������������������������������������������������������������� �������������� table 1: data type and source category data type data source demographics census australian bureau of statistics business victorian business registry australian bureau of statistics crime victorian crime statistics australian bureau of statistics education state schools – primary, secondary and technical victorian government: department of education & training website special developmental schools special schools language schools special education unites private school victorian independent schools directory victorian university federal government: department of education: science & technology website victorian tafes street directory index adult migration education yellow pages kindergartens yellow pages deaths coroners court data monash university cemeteries street directory index funeral directors yellow pages religious centres religious centres by denomination street directory index masonic centres street directory index recreation reserves/parks & leisure centres list of all athletics clubs bicycle tracks bmx tracks boat clubs boat launching sites bowling clubs indoor bowling clubs cinemas yellow pages croquet clubs street directory index golf clubs – public & private victorian golf association website golf driving ranges street directory index gun clubs victorian government: department of tourism, sport & the commonwealth games indoor cricket street directory index life saving clubs victorian government: department of tourism, sport & the commonwealth games motor cycle tracks street directory index night sports yellow pages racing tracks (horse and greyhound) street directory index riding schools and clubs (horse) scout parks public squash courts public skating rinks public swimming pools street directory index public tennis courts liquor licences victorian government: department of human services website gaming licences victorian government: department of human services website �� �������������������������������� �������������������������� ������������������������������������������������������������� ������������� ��������������������������������� � services ambulance services street directory index drug advice yellow pages law courts victorian government: department of human services website nursing homes street directory index police stations street directory index post offices australia post website rsl clubs returned services league website railway stations street directory index designated refuge areas street directory index state emergency services victorian government: department of infrastructure website animal hospitals street directory index veterinarians yellow pages cultural art galleries yellow pages theatres yellow pages libraries street directory index museums street directory index holiday caravan parks street directory index hotels and motels street directory index government services centrelink victorian government: department of employment website consulates street directory index state emergency services victorian government: department of infrastructure website community services community corrections centres street directory index community health centres victorian government: department of human services website halls street directory index health and community services victorian government: department of human services websitehospice and palliative care hospitals – public hospitals – private street directory index hostels – migrant street directory index hostels – youth maternal and child health centres street directory index municipal offices street directory index neighbourhood houses local government websites retirement villages yellow pages child care centers yellow pages lassist newsletter vol. 4 nos.3&4 private and public sector responsibility for the collection, distribution and analysis of statistical data joseph e. kasputys data resources inc lexington, mass. statistical information is a vital national resource quantitative data on who we are and what we do influence countless decisions in all parts of the hedera!. state and local governments, in commercial and investment banking, in manufacturing, in service and retail trade, and in various nonprofit enterprises such as hospitals, schools and foundations. such data also play a role in our personal lives, shaping education and career choices, personal finance and investment strategy, regional preferences and other lifetime decisions. the collection and distribution of statistics is an essential role of government like most resources, statistical data is indeed a limited resource. the american public, in general, and business, in particular, as evidenced in the mounting concern over paperwork burdens, have a finite capacity to respond to the ever-growing demands for information from the government it simply costs too much to respond to the total sum of information requirements arising from legislation, regulation, and other individually worthy program requirements. once the data are collected, this proliferation of information can have the effect of turning "more into less" by making it extremely difficult lo select relevant data from a myriad of conflicting sources, definitions, and time periods at the same time that both the capacity to produce and the capacity to utilize statistical data have approached the limits of reasonableness, we have had the accompanying orthogonal concerns over the right of the public to have access to data collected by the government and the right of the individual person or business to receive appropriate protection from invasions of privacy for these reasons, a more comprehensive effort does need to be made lo review and control federal statistical policy current programs conducted by the office of federal statistical policy and standards are certainly helpful to assure that maximum use is made of the existintg statistical system and that new requirements are appropriately designed to impose a minimum burden on the reporting public while meeting, to whatever extent possible, any unfulfilled needs for data that affect a broader user community. in my own view, these current programs are not sufficient. further, it is unlikely that any measures will be truly sufficient, including proposals lo establish an office of federal information policy or an independent office of statistical policy in the executive office of the president, until both the legislative and executive branches fully recognize statistical data as a scarce resource and begin to treat it accordingly. however, more central review and control should contribute lo this realization and. therefore. i would support some form of increased coordination and control by the executive office of the president in considering what types of control may be appropriate, it is helpful lo refiect on why the federal government collects statistical information i believe there are four major reasons: 1 regulation. regulation implies control, and control requires comparison of actual performance with a standard this comparison requires measurement, which converts into a need for statistical data. the explosion in federal regulatory activity has. in turn, generated enormous data requirements. legislators have also learned that statistical data options can affect regulatory outcomes as a result, legislation increasingly specifies the data lo be used, which limits the flexibility that should be present in the design of the federal statistical system. 2 program operation. many programs require data in order to operate at all major examples include revenue sharing, local public works, and similar grant programs lied lo specific formulas other programs, while not operated purely on a statistical basis, require data for evaluation of effectiveness and possible modification. as with regulation, those who design programs in the executive branch and enact them into law in the legislative branch have learned that the data used do have a material impact on program results. this is, of course, true with formula programs, where endless varieties of variables are tested in alternative formulas until one is found that provides the program designer with a distribution of funds that is intuitively acceptable. again, the needs of the federji s;.n,stical system are secondary in this process, which .nsicad encourages the development of increasing.' amounts of data collection to provide statistics that are uniquely appropriate for each program lassist newsletter vol. 4 nos.3&4 -''. policy analysis. slatislical daia musi be colleclcd lo provide information lo policymakers in ihe executive and legislative branches on economic and financial conditions that exist nationally and internationally data on the national income accounts, production, prices, population, trade, and similar items fall into this categoryit is probably the area that requires the most difficult decisions and one in which an office charged vnith statistical policy can have the greatest influence if the data needs of society could be correctly anticipated, and if proper discipline were exercised m program design, legislation, and regulation, practically all ihe statistical information needed for regulation and program operation would be regularly available through the thoughtful and coordinated development of ihe statistical system. while this may be an ideal that can never by reached, it is a worthwhile goal to keep working toward 4. information programs. certain government programs exist for the principal purpose of keeping the public informed the intent may vary from stimulating technological innovation tha>ugh the diffusion of knowledge to reducing market imperfections through improved information what statistical policy control and coordination will be effective given these sources of data requirements' for the first two sources, regulation and program support, it appears clear that statistical policy must be effectively incorporated into legislative proposals and regulatory actions this can be best done from the perspective of the executive office of the president hopefully, with growing public concern over the paperworl. and reporting burden, the legislative branch will become more sensitive to the need to make reference to statistical policy and precepts before enacting legislation and make increasing use of the capabilities of the proposed organization in the third area, statistical data for policy analysis, i believe the existing highly decentralized federal statistical system requires the improvement that can only be obtained through an organization thai has broad perspective over federal activities and sufficient clout to make a difference this again argues for a stronger central role from a high level. in the fourth area, government programs to provide information to the public are more controversial. prior to the extensive development of publishing, media, and information industries, the government indeed fulfilled an important need by supplying information facilitating commerce and industry that would not otherwise be available. however, in 1480, we find a highly developed information industry, utilizing the latest technology, that has a strong capability to identify information needs, collect data, deliver results frequently tailored to the needs of specific clients, and even provide special analysis on the meaning of information for business decisions in lieu of unilateral government determinations on the nature of private sector information requirements, market demand and the profit motive can now be used more extensively to govern statistical collection. equally important, the reporting public becomes free to chose whether lo respond to information requircmcnls generated by the private sector, which can be expccled lo be in propomon 10 the perceived value of the information the efforts of a strengthened central statistical policy office should focus on transfemng data collction and distribution of this nature to the private sector, while improving and rationalizing statistical operations in support of regulation, program operation and policy analysis. i would like to encourage any new statistical office to remain .separate from other aspects of information policy statistical policy is sufficiently unique and important to deserve separate attention. to be sure, statistical policy must be established with an awareness of telecommunications, adp management, pnvacy, and related concerns; but should not be merged with these other functions which are more involved with regulatory concerns. incidentally, with the reduced costs of computers and communications and rapidly rising costs of personnel, the emphasis on strong central control over the acquisition and use of adp and related services is probably misplaced. more emphasis should be given to ihe improvement of productivity through the use of these capabilities ihan to elaborate restrictions on procurement given that the federal govemmenl does, and must, collect statistics, the government bears a major responsibility to make this information available lo the public. divergent interpretations of the term "'making information available lo the public" are possible interpretations of this term can range from placing information in a public reading room somewh. re in washington. d.c.. such as the sec public reading room, lo ihe federal government's placing all available statistics in a shared computer and aggressively marketing these statistics to private sector organizations to the extent that controversy exists over the interpretation of ihe federal role in information dissemination, 1 believe thai much of it can be traced to two conflicting principles. on the one hand, we wish the federal government lo make available all information that is collected to the public, except where individual or business rights to proprietary and confidential information would be violated on the other hand, an equally important precept in our society is that the federal government should not encroach on activities that properly belong in the private sector it is fortunate for our economic health that the federal government has avoided enlering into legitimate private sector enterprises where not needed because, despite this restraint on private sector encroachment, we have only seen the role of the government grow bigger over time it is clear that an extensive information industry has developed in the united states to take full advantage of available communications and computer technology some testimony to the existence of this industry is an article in the may .s. 198u edition of fortune (which incidentally is the 25th anniversary edition of the fortune 500), entitled "everything you always wanted to know may soon beonline" this article explores ihe rapid development of ihe online information industry, which includes both bibliographic and statistical databases. the on-line database industry includes firms that specialize in organizing and delivering just one or a few databases, along with those that provide almost encyclopedic coverage of all available quantitative data lassist newsletter vol. 4 nos.3&4 many ol these firnii indeed carry dala that arc collecled and published by ihc federal govemmeni. since ihcy are involved in assessing market needs and matching these needs to dala availability, it is likely that most of the markets for statistical data, including data provicfcd by the federal government, that can be economically served either have been or will be identified products will be developed based upon available data by the private sector information industry to serve the needs of these markets it is important to emphasize that i know of no one in the industry who does not believe that the government should have access to all available technology to use in the dissemination of its data our principal concern is rather whether the government is competing with already existing private sector activities, which include computers, telecommunications and software that have been developed at considerable pnvate expense one consideration may be whether ihc existing industry efforts adequately serve all the markets for information that require such service there is little concern over whether the fortune 51hj companies, major banks, or large research organizations are being adequately served with data questions have been raised from lime to lime over whether the needs of the small organization or the individual researcher in the university are being adequately met, and whether the government should not step in. using the latest available technology, to meet these needs. this is clearly a matter for public policy to decide if such needs are not now' being met by the pnvate sector, it is likely that they cannot be met economically on a full cost recovery basis. at such time as the technology would permit such needs to be met economically. i am certain that the private sector would step in. since it is unlikely that the govemmeni could perform these services any more economically than the pnvate sector, we are then faced with a decision as to whether to subsidize information dissemination to these specialized markets. it is my own belief that in most cases, if the government makes information available through depositories and through statistical publications and reports, the small business or the individual researcher with an occasional need for information will be able to gain adequate access to the data and reports he or she seeks i would also encourage the government, as i mentioned earlier, to make full use of computer and communications technology in its own inlemal dissemination and utilization of statistical data. the government has been an advanced user in this regard all private citizens would hope this will continue. for it should lead to the more effective use of statistical data for public policy development it is axiomatic thai data always receive less analysis than they deserve. new technologies available through the computer and communications networks should permit higher levels of analysis to be conducted, leading to better decisions. the only caveat to be considered is the classical "make or buy" decision that is observed within government, which centers around whether to develop an inhouse capability for the storage, distnbulion and analysis of such data or whether to obtain it from pnvate contractors here again. i believe there is ample evidence to support the view that the pnvate sector is considerably advanced in the distribution and analysis of quantitative data and has much to offer the government in this regard . in order to gauge the adequacy of the on-line database information industry, it might be useful to provide a few additional words of claritication on industry operations dri serves as a good example of this industry and indeed was cited in the fortune article mentioned earlier as "probably the preeminent company in the on-line information field." dri has spent many millions of dollars over the past 1 3 years in collecting, organizing and documenting data only a small portion of these data originate in the federal government. while dri does maintain data on the national income accounts, prices, population, housing starts and the like, much of drl's data originates elsewhere, from such agencies as the oecd. imf. world bank, foreign governments; and from daily stock prices, commodity quantities and prices, market research surveys, and material published privately by trade associations a reprsentative listing of dri databases is attached in an appendix to this paper drl's current on-line storage capability for immediate access material is in the neighborhood of 36 billion characters and is scheduled to grow to 55 billion characters by the end of this year. we believe we have the largest collection of economic and financial information available anywhere. many of these databases were originally issued in cross-sectional format dri has taken many periods of cross-sectional data and converted them into consistent time series. this often involves dealing with definitional changes from one time period to the next, changes in the data collection approach, or changes in the basic entities measured once organized in a consistent time series formal, the dala are described both in on-line documentation and in reference manuals mnemonics are assigned for easy use in referencing through software. dri has developed software that permit any and all on-line data to be retrieved and manipulated with operations varying frotn seasonal adjustment to all forms of regression, to forming equations, to building models, to simulating the models, to producing reports and graphs for effective presentation of results more importantly, it has been drl's finding that on-line databases can only be efficiently and economically offered to markets with a good understanding of the applications to which the data will be put indeed, part of our product offerings include applications methodologies and consulting assistance that enable users to combine their own data on sales, production costs, and competition with the generally available data provided by dri. questions have also been raised on the appropnate role of the federal versus the pnvate sector in performing analysis of statistical data it is my belief that the federal statistical agencies should focus their analytical efforts on developing descriptive information from data that are collected. this involves organizing data to present it effectively, so that the users of the data can fully understand what the statistics mean this is a critical role which will often stimulate policy action, since proper presentation will of itself indicate that such action is needed. any analysis that goes beyond display or rearrangement of factual information should be limited to political appointees and policymakers, together with their supporting staffs, who should be kept separate from the statistical agencies and the information which they release [71 ) lassist newsletter vol. 4 nos.3&4 this latter type of analysis would include forecasts and interpretations of the underlying meaning of statistical information. it is important that analysis of this sort be released separately fropi the factual statistical information provided hy federal agencies so that the public can readily tell fact from opinion private sector firms, of course, should engage in any and all analysis appropriate for the markets being served indeed, this provides a plurality of information to the general public the information industry has an important role to play in providing alternative interpretations of statistical data to businesses and industry, which can then be used to form opinions on the validity of the analysis being provided through the political system. this plurality of information, together with multiple sources of statistical information delivery, is a healthy aspect of our society and is essential to maintaining an informed public which can reach independent judgments on major economic and financial issues. the development of our national statistical system continues to be a matter of great importance to all economic and financial activity the proper development of this system goes beyond narrow economic and financial concerns, and has direct infiuencc in maintaining the principles of our democracy the development and use of statistics should be a mutual undertaking, shared by both the federal and private sectors. 1 believe these sectors have worked together very cooperatively in the past to give the united states clearly the best statistical data available anywhere in the world we now face new challenges in how we can work together to fully utilize the advances that technology has given us i believe that continuing dialogues among all concerned will contribute materially to our ability to face these new challenges and meet the needs of the general public and our private clients in the most effective manner possible. appendix: a guide to dri data bases age-income model agriculture and weather automotive best executive data california canada canadian model data bank chemical data banks coal data bank coal model data bank commodities market data bank compuslat the conference board consumer expenditure survey flow of funds current population survey plan's data bank developing countries primary source data banks pro forma data bank drifacs site ii dri-sec standard & poor's industrial financial data drilling data banks state and area forecasting service data bank energy steel data banks forestry and wood service data banks ibrd's worid debt tables data bank imf's balance of payments imf's international financial statistics industry financial service insurance service data bank international energy data international trade information service data japan new york city model data bank oecd main economic indicators oecd national income acounts oecd trade series a cost forecasting service data banks paper and pulp data banks european model data bank target group index european national source data bank transportation fdic data u.s. central fiei us macro model data bank u.s. prices u.s. regional us weekly banking value line [72] applications of nformation science to social measurement by murray aborn national science foundation this paper was presented at the lassist annual conference^ may 19-22, 1983, in philadelphia, pennsylvania. in the 1960's, the growing influence of the computer caused dramatic changes to take place in the concept of scientific data and the character of data analysis. among these changes was the onset of a shift away from single-purpose data collections and analyses based on relatively small data sets, toward large-scale data collections and analyses based on data banks serving multiple applications and possessing widely accessible storage and retrieval systems. in social science this led to the establishment of data archives and an early attempt to regulate their functions. additional purposes were to keep these facilities abreast of a rapidly advancing technology, and enable them to remain au oourant with increasingly sophisticated management schemes for operating over larger and larger bodies of data (1). this paper briefly traces the role of the national science foundation (nsf) in these developments, discusses the current state of affairs with respect to social science data resources, and questions whether continued reliance on sheer data amassment is the true path to the further intellectual progress of the field. the building qf nsf data programs in the years immediately following the advent of the computer-based data archive, nsf involvement in the expansion and upgrading of the major sources of social data grew and intensified. to increase the research-return from the enormous investment society makes in the collection of social statistics, projects were supported to enhance the researcher access to them. to fill the gaps in the social science data base, projects were supported to maintain data series not covered by the federal statistical system but needed for the monitoring of social and economic trends and the modeling of longterm social change. direct support was afforded to archival facilities to help them expand their holdings and degray the costs of dissemination. on the user side, projects were supported to increase the research utilization of stored social data across disciplines, including projects to introduce bibliographic-type control over machine-readable data files in order to reduce duplicate data collection and help prevent incomplete data analyses. and alongside these data programs , projects were funded to improve existing methods and create new tools of broad utility in analyzing the growing stock of social and economic data becoming available to the research community. in 1976, a committee of the national academy of sciences surveying the social sciences at nsf acknowledged the role nsf programs were playing in the sphere of data resources. it declared: it is generally felt, and refleoted in the long-range plans of several of the social science programs in nsf, that deficiencies in the available base of social science data are seriously impeding the progress of research ( 2 ) . the committee did not refute this outlook; in fact, it ultimately recommended that such planning continue and include greater support for longitudinal studies over extended time periods and national facilities for survey research and large data bases. financial backing for the data programs described above was provided not only by nsf's division, of social and economic science, but by sections in other parts of the foundation, such as computer science and information science. over the next six years, however, computer science and information science turned inward, concentrating on their own disciplinary development and gradually eschewing applicational extensions to other fields of science. but in social science, the work went on. programs were maintained that to this day continue to build a data resource infrastructure capable of sustaining the large empirical researchtradition which characterizes contemporary social research. is it time for a change in orientation ? the national academy of sciences committee surveying nsf's social science programs in 1976 never made crystal clear precisely what was being referred to by "deficiencies in the available base of social science data," and which of these were more or less responsible for "impeding the progress of research." it is pretty apparent from the committee's report, however, that data resource planning in social science was largely oriented toward filling topical gaps and producing lower levels of aggregation, largerscale, longer-term data gathering efforts, and a more systematic approach. though the importance of methodological accompaniments to assure good data quality was neglected neither in nsf's programs nor the committee's report, the effect of stressing data gaps and data shortfalls inevitably leads to more and more data getting collected and more and more data being retained-and that is exactly what has happened. now, one question that arises at this point is whether an orientation toward data amassment has had negative as well as positive consequences. the answer is "yes." if negative consequences is too strong a term, then we can at least speak of limiting effects. and if there have been negative consequences or limiting effects, then it is time for a change in orientation. but before proceeding to describe what the required change appears to be, it is crucial to make clear that such change in no way gainsays the compelling arguments put forth in many recent publications regarding the value of secondary analysis. nor does it gainsay the need to have data available for reanalysis in order to test for bias in reported results, challenge data-driven theoretical assertions, and generally carry on the processes of scientific understanding in a field which is rarely able to conduct controlled experiments or reproduce the original conditions of an investigation. a change in orientation simply argues for diverting some amount of effort and devoting some portion of available resources to study the deeper aspects of the enterprise in which the field has become heavily engaged. one negative consequence of the data gathering enterprise has been the pejoration of the term "data." this is no doubt connected with the fact that the enterprise is largely concerned with quantitative data in computermanipulable form, but in any case the term data is now commonly used interchangeably with the terms observations, information and, worst of all, evidence. i daresay few really believe that data in and of themselves prove anything, but that's the way we have come to talk and, i fear, occasionally think. however, the more frequent tendency is to confuse data first with observations and then with information. i realize this gets pretty elementary, but contemporary social science data archives contain mostly recorded observations, not data. it sounds more imposing to speak of data archives, and it is certainly easier to raise money in the name of data than it is for just plain old observations, but the terminology is inaccurate. observations become data only after they are placed in some analytical framework. as it is obvious from the general-purpose nature of the data archives, the same observations are destined to be interpreted as more than one kind of data. a similar confusion prevails with respect to data and information. the two are not synonomous , though the exposition here is yery difficult inasmuch as the relation is inferential and dependent upon the application of external structures. simply state, it behooves us not to forget that any body of data is a mixture of information and noise, and that the proportions will vary according to the use to which the data are being put. in the main, the signal to noise ratio in social science is typically much lower than in the physical sciences, which is a way of saying that the information content of a data base can be very meager, particularly when the data are employed to test hypotheses far afield from the hypotheses which motivated the data collection originally. it is thus ironic that the very success of large-scale, integrated data bases and the attendant data-processing technology often leads to a confusion of the technology with the natural semantics of information. which is heavily context-dependent. thus the underlying assumptions appropriate to the context of one application may be totally inappropriate to the contexts of other applications. moreover, the difficulty is compounded by the fact that in their research, social scientists are heavily dependent upon data files which were not generated for scientific purposes, such as census data, voting records, police and court records, governmental budgets, and so forth, and whose informational value relative to the kinds of scientific questions social scientists ask may be completely uncertain. bringing information science into the picture in the previous section of this paper, mention was made of the ancillary role played by the computer and information sciences in the building of nsf data programs. it was noted that those roles, diminished after 1976, and that nsf's contributions to the data resource infrastructure of present-day social science has been carried on exclusively by the social science elements of nsf. this situation is changing. given recent advances in information science, it seems particularly important to begin to apply newly-formulated principles of knowledge management to social science data resources precisely because their holdings--observations of social and behavioral phenomena in digital form--tend to be incomplete, imprecise, and error-prone due to the fuzzy nature of the phenomena being observed and the looseness of the data gathering process. knowledge management facilitates the translation of user needs into expressions upon which a data base system can act. one example of possible applications to social data is the development of data base specification languages, that is, languages which would permit social science researchers to express their requirements in functional terms. these might then be translated into a database format, perhaps based on relational structures rather than representational ones, as is the present mode, which would help skirt the data dependence problem. other areas of potential application to social science data may come from information science's concern with descriptive classification, indexing, and the problems of relating variant terminology in a single retrieval system. the current work of dolby is an example (3). dolby argues that the correctness of data and data analysis involves correctness in meaning, and that correctness in meaning goes beyond matters of computer program correctness or the numerical accuracy of data. his approach concentrates on the use of classification structures to extend the formal treatment of meaning in computer-based data systems, and he has shown how such extensions can expose or reduce ambiguities and inconsistencies of meaning in such systems. there are some other, more practical reasons to believe that the time has come to test out achievements in information science as they may be applied to stored social data. urgencies created by current reductions in the quantity of social science-usable data generated by the federal statistical system is one reason; cutbacks in the funds available to support scientifically oriented social science data resources is another. it would help greatly if we could improve our ability to estimate the degree of redundancy (i.e., the amount of information overlap) among data collections, and if we could make progress in our ability to set data collection and maintenance priorities. considering the potential benefits of bringing information science into closer contact with social science data problems and opportunities, it has been decided to launch an initiative--still informal at this juncture--to make known our receptivity to proposals which combine or merge the subject matters normally covered by social science and information science independently. such proposals will be handled jointly by nsf's division of social and economic science and its division of information science and technology (4). the division of social and economic science supports the establishment, evaluation, and improvement of social science data resources, research on social data, and the development of methods for analyzing such data. the division of information science and technology supports research to increase understanding of the properties of information transfer. we believe the future will show that this initiative was well advised. notes and references (1) see, for example, glaser, william a. note on the work of the council of social science data archives. social science information , 1970 8(2) :159-176. it may be of some historical interest to point out that the council of social science data archives was the forerunner of lassist. (2) social and behavioral science programs in the national science foundation . final report of the committee on the social sciences in the national science foundation, national research council, national academy of sciences, washington, d.c. , 1976. (3) dolby, james l. meaning from data; implications for data analysis and database management systems. paper presented at the meeting of the american association for the advancement of science, detroit, michigan, 1983. (4) inquiries may be addressed to: program director for measurement methods and data resources, div. of social and economic science, nsf, washington, d.c. 1/13 kudrnáčová, michaela and ilona trtíková (2020) sustainability through the liaison with data archive users, iassist quarterly 44(4), pp. 1-13. doi: https://doi.org/10.29173/iq976 sustainability through the liaison with data archive users michaela kudrnáčová1 and ilona trtíková2 abstract as a social science data archive, we focus on collecting and archiving research data. however, there are more responsibilities that come with data archiving: cooperation on international social surveys (issp, ess), supporting secondary data analysis, and much more. a significant part of our work is to communicate with students and researchers, and to educate them about data management and data analysis. although the relationship we have is functional and seems sufficient, we tend to ask ourselves: who are the data archive users and what do they expect from us? we decided to employ user-centered design methods and tools to define a typical user of our services, find their motivations for using our data archive and the specific functions they do (or do not) use and appreciate, and thereby get a better understanding of their needs. moreover, we wondered about the role of open science and its impact on the users’ needs and future requirements arising from the open science environment. the information obtained is a starting point for redesigning archival services to satisfy new demands our users have regarding more data resources, new techniques for scientific work, and better interconnection between different platforms. keywords csda, data archive, open science, social science, user, survey, user-centered design https://doi.org/10.29173/iq976 2/13 kudrnáčová, michaela and ilona trtíková (2020) sustainability through the liaison with data archive users, iassist quarterly 44(4), pp. 1-13. doi: https://doi.org/10.29173/iq976 1. introduction ‘the czech social science data archive (csda) at the institute of sociology of the czech academy of sciences accesses, processes, documents and stores data files from social science research projects and promotes their dissemination to make them widely available for secondary use in academic research and for educational purposes.’ (csda, 2020) in order to support these primal activities, we occasionally come up with other activities and events such as seminars and workshops. due to the endeavor of constantly developing archive services, we decided to focus on understanding the needs of users: their perceived needs and how we, as a data archive, manage to meet these needs and what should be considered for their future needs. unfortunately, only a handful of studies focus on data consumers (borgman et al., 2015; late and kekäläinen, 2020). moreover, the cited articles only describe users’ backgrounds and the characteristics of the downloaded data. our ambition is to find out more about their interests and preferences, establish closer contacts, and involve our data consumers in the process of data archive development. we believe this topic is utterly important because the users are the key motivation that leads to building a sustainable data culture, data management and archiving, and making the overall concept work. this paper is organized as follows: firstly, we focus on the role of the data archive, stressing the general activities in which it is currently engaged and the need to promote and support open science; we then move on to the idea of applying user-centered design within social science data archives. the core of the paper is dedicated to csda users: what we already know about them and the description of a user survey and its methodology. thanks to the data we collected, we can describe ‘the typical user’ within a separate subchapter. finally, we will discuss the results of our survey, debate the employment of user-centered design in the context of the digital data archive, and determine if the chosen approach can sustainably ensure the effective functioning of the data archive. 2. the role of data archive the czech social science data archive, a national resource center for social science research, acquires, processes, documents, archives, and preserves digital datasets from czech and international social research and makes these data publicly available for two purposes: (1) secondary analysis in academic research and (2) training purposes at higher education institutions. at the same time, csda serves as a czech node within the pan-european distributed research infrastructure cessda eric (consortium of european social science data archives european research infrastructure consortium) and is the cessda eric service provider in the czech republic (csda, 2020). the data archive responds to changes in scientific work and supports open science. its data services are based on the fair data principles, emphasizing the re-use of data in academic research. it documents and processes data for purposes of secondary use, connects it with relevant research information and contextualizes it with other data and materials. at the same time, it is a source of research tools and procedures that have been verified in prior research, thus supporting the implementation of new surveys. moreover, it supports the use of secondary data analysis in research by: (1) providing training courses and taking part in educational programs at universities; (2) mapping https://doi.org/10.29173/iq976 3/13 kudrnáčová, michaela and ilona trtíková (2020) sustainability through the liaison with data archive users, iassist quarterly 44(4), pp. 1-13. doi: https://doi.org/10.29173/iq976 and analyzing available data sources, providing information services, and user support on data sources; and (3) connecting czech and international data sources and research in the field of data standardization and harmonization (csda, 2020). the data archive also promotes open science, which aims to open the entire research cycle, encouraging sharing and collaboration. it is based on collaboration and new ways of disseminating knowledge using digital technologies and collaboration tools (initiative, 2014). the data archive can be considered an essential open science service. the archive role in open science is to create a trusted and sustainable environment for social science data. the goal is, therefore, to remove any obstacles preventing the sharing of all types of outputs at any stage of the research process. it turns out that data sharing is highly dependent on the behavior of scientists (koltay, 2017). this is evidenced by various research, such as the data literacy multinational study questionnaire at charles university in the czech republic. the survey showed that social scientists are less willing to share their data and are more concerned about potential barriers (jarolimkova and drobikova, 2019). these findings correlate with other research around the world (tenopir et al., 2011; kim and adler, 2015) and also, to some extent, with our experience. as portrayed in the surveys mentioned in the previous paragraph, it is important to communicate with users to break down potential barriers. since the researchers are not only contributors but also data consumers, it is essential to ensure they feel support and concern from the digital data archive. therefore, it is desirable that archive services correspond to user needs. user-centered design can be the right way to design services that increase archive reliability and motivate users to collaborate through regular data sharing. 3. user-centered design user-centered design is the process by which end-users have an impact on design services. (abras, maloney-krichmar and preece, 2004) this process is comprised of a set of methods is used to develop applications and websites. the aim is to give users a quick orientation and thereby increase the use of services and websites. it is not limited to a specific area; the methods are universal for use in any development or redesign. several methods can be employed. designers and service designers choose appropriate methods, according to the particular situation, to meet the goal. the advantage of usercentered design is that a future user is considered from the beginning of the design (abras, maloneykrichmar and preece, 2004). first of all, we have to understand who our users are and what they desire (garrett, 2010b). the field of user research is devoted to collecting the data needed to develop that understanding. this design serves as a tool to define the users and to respond better to their demands and needs. it is necessary to collect data regarding what users want and need, mapping their experience in user-designed environments and their motivations to use specific applications or websites, and explore what makes it possible for them to use these features effectively versus which obstacles stand in the way of successful use. various methods, including analysis of web site visits, surveys, interviews, use cases, observations and more, can be used to collect data (still and crane, 2017). https://doi.org/10.29173/iq976 4/13 kudrnáčová, michaela and ilona trtíková (2020) sustainability through the liaison with data archive users, iassist quarterly 44(4), pp. 1-13. doi: https://doi.org/10.29173/iq976 an important step is to divide our users into smaller groups defined by key characteristics (garrett, 2010a). by creating these groups, we can then identify those groups that represent the user population and are best to work with when designing an application or a website. on the basis of information about user groups, we define typical users or personas with more detailed personal characteristics (garrett, 2010b). this helps to gain a better image of our users and will be useful for developers in tailoring the resulting design to the target users, in accordance with their needs, behaviors, and personal characteristics. 4. who are csda users? csda are at the beginning of the process of changing the system for managing and publishing research data. this is driven by several factors, including the currently used nesstar system’s lack of suitability for data security and, in our view, its limited user friendliness. however, beyond our personal perceptions, we believe it is necessary to involve the users in the process to get a better idea of their key issue and their perceptions of the current data archive’s functionality and services. we have accordingly implemented methods of user centered design. to define and understand our users and profile typical users, we have chosen to conduct a survey among existing archive users and interviews with selected users. moreover, we already have existing data at our disposal (see details in the chapter down below) and operational experience that have further helped us define the user. we have two different sources from which we can draw conclusions. the first is the registration form our users need to fill in while signing up, and the other is a short survey we conducted at the beginning of 2020. before addressing the survey, we will briefly reflect on some basic statistical measures that we are able to extract on existing data from registration forms from the year 2019. since the existence of csda dates from march 2007, the first figure (figure 1) shows a trend of growth. allowing for annual fluctuations, the number of registered users has increased over time. unfortunately, we can only guess the causes of the fluctuations. since the users are mostly students, reasons may include variations over time in the effectiveness with which they share information about the archive and varying numbers of students enrolled in data management courses. https://doi.org/10.29173/iq976 5/13 kudrnáčová, michaela and ilona trtíková (2020) sustainability through the liaison with data archive users, iassist quarterly 44(4), pp. 1-13. doi: https://doi.org/10.29173/iq976 figure 1: the number of users registered each year since its founding in 2007 until 2019 examining the figure below (figure 2), we can see the frequency of registrations is highest in march. this is most likely to be caused by students and their need to work with data for the purposes of assignments at the end of spring semester and due to the work on students’ theses. also, as mentioned earlier, this might be also caused by students being informed about the archive especially at the beginning of the courses in which they enroll. from june through august, there is a steady decrease in registrations, since most of the students have some time off, as is true for teachers and other academics on holidays. there is again an increase in october when the autumn semester begins. this is due to the same reasons as in the spring semester. figure 2: the average number of registrations since 2008 until 2019 note: the numbers are averages from each month from 2008 through 2019. 2007 was omitted because, as it is the year in which the archive was founded, some months would be missing and the whole year would be an outlier in comparison to all the other years. based on the information we get from new users when registering, we are able to roughly describe the users. for that purpose, we chose the most recent year at our disposal: 2019. during this year, a total of 346 people registered into our archive. about 55% stated studies as the main reason for registering, 19% academic research and almost 16% teaching. less common reasons included 2007 2008 2009 2010 2011 2012 2013 2014 2015 2016 2017 2018 2019 171 241 209 191 260 300 259 275 303 284 259 351 346 0 50 100 150 200 250 300 350 400 221 332 417 270 250 148 119 95 121 286 393 309 0 50 100 150 200 250 300 350 400 450 https://doi.org/10.29173/iq976 6/13 kudrnáčová, michaela and ilona trtíková (2020) sustainability through the liaison with data archive users, iassist quarterly 44(4), pp. 1-13. doi: https://doi.org/10.29173/iq976 postgraduate studies, individual interest, public sector and other. as for the affiliation, two thirds were affiliated with universities (charles university about 30%, masaryk university about 21%, palacky university olomouc about 11%), with students adding that they are mostly from faculties of arts or faculties of social sciences, leading us to believe it is mostly students/employees from the field of sociology, political studies, social studies and police academies. the last piece of information worth mentioning involves the country of origin: 95% of the users came from the czech republic, while others were from austria, china, france, germany, and other countries. 4.1. survey methodology to find out more about the csda users, it was decided to conduct a short online anonymous survey. overall, 3 398 people registered with our database were approached via their email address; 312 came back as “undeliverable”, 6 were manually excluded (due to death, terminated employment, or to their request to be excluded from the survey), and an additional 174 respondents deregistered themselves because they did not wish to be contacted in connection with the survey. over the period from the 6th of january until the 10th of february 2020, 564 individuals opened the survey, of whom 263 (46%) filled it in, representing 7.7% of the 3 398 registered users. the survey was conducted in czech language only, since most of our users is of czech nationality (about 90% in total); however, in the future, we aim to cover english speaking individuals as well. the survey consists of 10 brief questions developed to find out the usage frequency of the data archive, users’ affiliations, the types of activities they perform through the archive, the nesstar functions they use, citation practices, other data archives they use, and suggestions and complaints they have (this final question includes a space for expressing their wishes for improvements and other functionalities). within the survey, we focused solely on information relevant for csda functioning and therefore, we omitted questions regarding respondents’ age, sex/gender and other sociodemographic information. 4.2. survey results one question, included to help us understand respondents’ backgrounds, referred to the affiliation of respondents (figure 3). most of the users (82%) who responded to the survey come from some sort of public organization, while the rest are comprised of people working in commercial organizations (7%) and ngos (6.3%). interestingly, 13 users (4.8%) stated “self-interest”. https://doi.org/10.29173/iq976 7/13 kudrnáčová, michaela and ilona trtíková (2020) sustainability through the liaison with data archive users, iassist quarterly 44(4), pp. 1-13. doi: https://doi.org/10.29173/iq976 figure 3: 'under which affiliation do you use csda?' n=263 a question regarding the activities for which respondents needed data from csda (figure 4) allowed respondents to select more than one response: this revealed studying (28.6%) and academic research (28%) as the dominant activities. however, self-interest in the data (14.8%), teaching (14.3%) and applied research (12.5%) were also significant responses. the other two options were chosen by only a couple of respondents (“other” by 8 and “not using csda anymore” by 6). it is worth mentioning that among “other”, policy making, journalism and advocacy were mentioned. figure 4: 'for what purpose do you use csda?' n=447 the next question in the survey concerned the frequency of use of our data archive (figure 5). there seems to be quite a variation: almost half of the respondents stated they use the data archive “several times a year” (48.1%), slightly more than one third claimed to use it “less frequently” than once a week (38.5%), 10.7% use it “once a month” and only 7 respondents (2.7%) use the archive “at least once a week”. 223 19 17 13 public organization (school, academic institution, governmental institution, etc.) private commercial organization private ngo self-interest 128 125 64 62 54 8 6 0 20 40 60 80 100 120 140 studying academic research self-interest teaching applied research other not using csda (anymore) https://doi.org/10.29173/iq976 8/13 kudrnáčová, michaela and ilona trtíková (2020) sustainability through the liaison with data archive users, iassist quarterly 44(4), pp. 1-13. doi: https://doi.org/10.29173/iq976 figure 5: 'how often do you use csda?' n=262 it was also crucial to extract from the users what functions they need and use (figure 6). almost one third of the respondents (31.3%) were “downloading the data”, while 21.1% were “downloading data documentation” and 18.6% were “downloading the metadata”. the last two options (“using keywords to search nesstar” and “displaying tables in nesstar”) totaled around 15% of answers. figure 6: 'what nesstar functions do you use?' n=634 since the data archives have long been struggling with students and researchers not citing the data or not citing them properly, we decided to also address this in the survey (figure 7). about two thirds of respondents (66.4%) stated they “always” cite used research data and 16.2% admit “mostly” citing the data. as opposed to that, only 6 respondents said they do not cite the data in any form at all and, last but not least, 15.1% stated they do not cite the data simply because, since they do not use or share the data publicly or in class, they have no reason or platform for citing it. although there is no way for us to know if the users are indeed citing the data, we are at least able to say that this outcome shows they realize the norm (i.e., they should be citing the data regardless of whether they actually reference used data). 7 28 126 101 at least once a week once a month several times a year less frequently 197 134 118 94 91 0 50 100 150 200 250 downloading the data downloading data documentation (technical information, questionnaires, cards, etc.) downloading the metadata (basic research information) using keywords to search nesstar displaying tables in nesstar https://doi.org/10.29173/iq976 9/13 kudrnáčová, michaela and ilona trtíková (2020) sustainability through the liaison with data archive users, iassist quarterly 44(4), pp. 1-13. doi: https://doi.org/10.29173/iq976 figure 7: 'do you cite used data? either in written form (writing a paper) or oral form (for students during teaching classes) n=259 as seen in figure 8, slightly under half (44.7%) of respondents have no objections to the archive’s or nesstar’s functionality. weaknesses identified by the remaining respondents can be divided into three main groups, of which it seems the most burning issue is the “interface or nesstar functioning” (26.8%), followed by the “character of the data” (11.6%), and 10.2% wished for the amount of data and variety of the data to be broader in general. those citing the “character of the data” suggested that the users would like to have more quantitative data at their disposal, qualitative data and new types of data. among “other” desires (19 responses) for the archive, examples of specific complaints and wishes included the ordering of the data, conjoint datasets, nesstar failures, etc. figure 8: 'what do you find insufficient about the archive and nesstar functioning?' n=284 moreover, as for the nesstar functioning, a separate open-ended question gave respondents a chance to express precisely and more extensively their feelings about nesstar and its areas of insufficiency. we were stunned to find that 89 open-ended responses were submitted. most (41) referred to the interface and its chaotic arrangement, while 13 respondents commented on the functioning of the nesstar. interestingly, 23 responses involved wishes and interesting ideas for new features. the last 12 responses were generally positive, often mentioning similarity to other world-wide archives and stressing their “eventually finding what they need”. 172 42 6 39 always mostly yes no i have no reasong for citing the data since i do not share them in any form 127 76 33 29 19 without objection interface or functioning of nesstar character of the data (insufficient suply of quanti or quali data, big data, etc.) the amount or variety of the data other https://doi.org/10.29173/iq976 10/13 kudrnáčová, michaela and ilona trtíková (2020) sustainability through the liaison with data archive users, iassist quarterly 44(4), pp. 1-13. doi: https://doi.org/10.29173/iq976 the survey also asked about respondents’ use of other data archives ('do you use other data archives apart from csda for data viewing or downloading?'). based upon this, 53.8% of 251 respondents do not use other data archives; the rest do and mention data archives such as gesis, eurostat, čsú and many more. finally, but still importantly, we were interested in who brought the users to our archive (figure 9). close to half of responses (41 %) suggest the users encountered csda thanks to their school/university teachers, almost one fourth (23.3%) learned of csda through the internet, and about one fifth (21.1%) responded with the option “work colleagues”. only a small proportion of answers suggest classmates (7.1%) and friends (3.7%). among the “other” option (five responses) were presentations, seminars and scientific literature. seven respondents did not remember. figure 9: 'how did you find out about csda?' n = 257 the survey concluded with an open-ended question allowing the respondent to express anything regarding the data archive that was not addressed earlier in the survey. examples of 11 areas mentioned included the lack of localization in both czech and english, the arrangement of surveys, large-scale downloading, and design. moreover, we were very pleased and touched to get 20 positive comments, mostly simply thanking us for the data archive’s services and its existence. 4.3. discussion: defining typical user defining a typical user helps us to grasp our target group. the data we already had from our registration forms are more informative when combined with the user survey results. this is especially true since the survey results are not simply descriptive but were specifically designed to involve the data consumer in the data archive development and give them an opportunity to express their very own preferences in the matter. the csda seems to be mostly dealing with and serving people from academia: students, academics, and teachers. it makes sense that they tend to use the data from the archive for purposes of studying and academic research, as well as occasionally for teaching. this also explains why they usually 132 75 68 23 12 7 5 0 20 40 60 80 100 120 140 school teachers internet work colleagues classmates friends i do not remember/ idk other https://doi.org/10.29173/iq976 11/13 kudrnáčová, michaela and ilona trtíková (2020) sustainability through the liaison with data archive users, iassist quarterly 44(4), pp. 1-13. doi: https://doi.org/10.29173/iq976 download the data at least several times a year or even once a month: academics need to write articles and books, while students need to write theses or seminar papers. teachers might only be teaching, or their role may be broader since academics and students on a certain level of their studies have to teach as well1. due to this fact, we believe there are two specific types of users that are generally similar: academic and student. beyond their specific characteristics described at the very beginning of this subchapter, there is a bit more worth mentioning. both academics and students state that they always or almost always cite the data. a small part of the students, though, say that they have no reason to cite the data, meaning they could be downloading the data for inspiration while conducting their own research. as for usage of nesstar functions, it does not differ across the whole sample of respondents: academics and students mostly download the data and documentation, but generally use all offered functions. academics and students generally concur in their objections against nesstar, with students slightly more eager for greater quantities, variety, and types of data at their disposal. interestingly, with regard to the usage of other data archives, two-thirds of the students within our sample (about 80 people) use no other archives, while two-thirds of academics (about 75 people) do use other archives. these results are of a great importance to us. our users are genuinely interested in our data archive development and clearly know their wishes for data services. taken by themselves, the transaction log files, or registration forms would not have provided this precious information. figure 10: defining a typical user of csda student and researcher student academic studying at a university needs to write thesis or seminar papers one-time need due to fulfilling school duties uses nesstar functions beyond the documentation downloads data file is citing a dataset interested in other data does not use any other archives university employee needs to write articles and books repeated need already has ways of using nesstar focuses on search functions needs to search questions is citing a dataset does use other archives 5. conclusion within the article, we discussed the role of the data archive, specifically the role of csda, and our efforts to adjust to the needs of the users. we mentioned the importance of employing user-centered 1 it is important to know that these categories were constructed based on a question regarding user activity. a total of 42 people identified as both academic and teacher, while 45 identified as both student and academic, 80 only as academics, and 83 only as students. while talking about “specific types of users”, we compare academic respondents and student respondents, meaning there are duplications within the group. https://doi.org/10.29173/iq976 12/13 kudrnáčová, michaela and ilona trtíková (2020) sustainability through the liaison with data archive users, iassist quarterly 44(4), pp. 1-13. doi: https://doi.org/10.29173/iq976 design along with making use of the data we already possess about our users. a survey was conducted in order to get more user data. finally, thanks to the information we obtained about csda users, we were able to describe a typical user. contemporary science demands sharing and collaboration. of course, this affects those systems that provide these functions. to ensure maximum support for open cooperation and the sharing of scientific data, we try to set up our service so that it best suits our users’ needs. for this reason, we chose user-centered design methods that reflect the needs of our users. the survey that we conducted provided us with a lot of data. in particular, it made use of our users’ willingness to cooperate in the development of our services. we have clarified and confirmed a number of data that can help us both in prioritizing the services offered and in communicating with users. defining a typical user, the first of several phases of service design, is very important. we have acquired a great deal of information that we will use in the next stages of designing our services. we have also realized that we should put more emphasis on creating a detailed nesstar user’s manual for the students among our users, while also introducing them to other data archives around the world. to overcome existing restrictions on the use and sharing of research data, we will endeavor to remove possible obstacles in the system to publishing data and accessing services offered by our data archive. it turned out that we needed to help our users use data from other archives. the goal of our efforts is to simplify their involvement in the open science environment. lastly, we need to emphasize that more research is needed in this field. there is a knowledge gap regarding digital data archives and their purpose, along with knowledge of data consumers, that we intended to bridge. as is true for any business, data archives offer services, which is the reason why they should learn more about their target groups and their desires. we attempted to carry out this task differently and thereby managed to establish a greater level of contact with our users, through which we now know of their deep interest in taking part in the csda development. therefore, we intend to keep in touch with our users, involving them in the next stages of user-centered design to achieve well-targeted communication, website design, data cataloging and events. to verify the continuing validity of our typical user profiles, we will need to conduct a short online survey every two years, as well as a couple of in-depth interviews. this periodic updating of the user profile will promote the long-term sustainability of our user-centered design efforts. https://doi.org/10.29173/iq976 13/13 kudrnáčová, michaela and ilona trtíková (2020) sustainability through the liaison with data archive users, iassist quarterly 44(4), pp. 1-13. doi: https://doi.org/10.29173/iq976 references abras, c., maloney-krichmar, d. and preece, j. (2004) ‘user-centered design’, encyclopedia of human-computer interaction. sage publications. borgman, c. l. et al. (2015) ‘who uses the digital data archive? an exploratory study of dans’, proceedings of the association for information science and technology, 52(1), pp. 1–4. doi: https://doi.org/10.1002/pra2.2015.145052010096. csda (2020) about the czech social science data archive. available at: https://archiv.soc.cas.cz/en/about-czech-social-science-data-archive (accessed: 17 march 2020). garrett, j. j. (2010a) the elements of user experience: user-centered design for the web and beyond. 2 edition. berkeley: new riders. garrett, j. j. (2010b) the elements of user experience: user-centered design for the web and beyond (2nd edition) (voices that matter), elements. initiative, o. s. a. r. (2014) open science and research: the open science and research handbook, open science and research: the open science and research handbook. initiative, open science and research. available at: https://www.fosteropenscience.eu/sites/default/files/original/3986.pdf. jarolimkova, a. and drobikova, b. (2019) ‘data sharing in social sciences: case study on charles university’, in communications in computer and information science. doi: https://doi.org/10.1007/978-3-030-13472-3_52. kim, y. and adler, m. (2015) ‘social scientists’ data sharing behaviors: investigating the roles of individual motivations, institutional pressures, and data repositories’, international journal of information management. elsevier ltd, 35(4), pp. 408–418. doi: https://doi.org/10.1016/j.ijinfomgt.2015.04.007. koltay, t. (2017) ‘data literacy for researchers and data librarians’, journal of librarianship and information science. doi: https://doi.org/10.1177%2f0961000615616450. late, e. and kekäläinen, j. (2020) ‘use and users of a social science research data archive’, plos one, 15(8), p. e0233455. doi: https://doi.org/10.1371/journal.pone.0233455. still, b. and crane, k. (2017) fundamentals of user-centered design : a practical approach. crc press. tenopir, c. et al. (2011) ‘data sharing by scientists: practices and perceptions’, plos one. edited by c. neylon. public library of science, 6(6), p. e21101. doi: https://doi.org/10.1371/journal.pone.0021101. 1 michaela kudrnáčová is a phd student at the social science data archive focused on research and its methodology, and can be reached by email: michaela.kudrnacova@soc.cas.cz (2020). 2 ilona trtíková works as data manager at the social science data archive, her expertise involves sharing and retrieving research information. https://doi.org/10.29173/iq976 https://doi.org/10.1002/pra2.2015.145052010096 https://archiv.soc.cas.cz/en/about-czech-social-science-data-archive https://www.fosteropenscience.eu/sites/default/files/original/3986.pdf https://doi.org/10.1007/978-3-030-13472-3_52 https://doi.org/10.1016/j.ijinfomgt.2015.04.007 https://doi.org/10.1177%2f0961000615616450 https://doi.org/10.1371/journal.pone.0233455 https://doi.org/10.1371/journal.pone.0021101 mailto:michaela.kudrnacova@soc.cas.cz 1/22 ndhlovu, phillip and matingwina, thomas (2018) the state of preparedness for digital curation and preservation: a case study of a developing country academic library, iassist quarterly 42 (3), pp. 1-22. doi: https://doi.org/ 10.29173/iq929 the state of preparedness for digital curation and preservation: a case study of a developing country academic library phillip ndhlovu1 and thomas matingwina2 abstract digital technologies have allowed libraries to create, manipulate, store and make accessible vast amounts of digital content. however, they endanger the longevity of the very objects they produce and require very different management than the traditional paper-based world. despite the fact that the national university of science and technology (nust) library in zimbabwe has amassed a huge body of digital collections, there are no formal mechanisms to ensure accessibility and long-term preservation of digital content. the study assessed the state of preparedness of nust library for digital curation and preservation of its digital collections. the conceptual framework was based on sinclair et al. (2011) and boyle, eveleigh, and needham’s (2008) formulations. nust library preparedness for digital curation and preservation was assessed by examining awareness, competencies, technology infrastructure, digital disaster preparedness and challenges to digital curation and preservation. a mixed methods research design employing a case study research strategy was adopted for the study. the findings revealed a low level of awareness of digital curation and preservation. challenges to digital curation are mainly lack of policies, lack of expertise by library staff and lack of funding. it is recommended that the library should consider digital curation and preservation as one of the primary responsibilities and take staff members’ training in this area seriously in order to ensure current and future access to digital collections. keywords digital curation, digital preservation, digital disaster preparedness, data management 1. introduction and background as digital data and technologies have fast become an integral aspect of 21st century life, public organisations, particularly libraries, are facing increasing demands for digital services from users who “routinely and unthinkingly” use or depend upon digital information (pennock, 2007). many libraries have responded to digital technologies by setting up digital libraries and digital repositories, especially academic and research libraries (kim, warga and moen, 2012). this has opened up an opportunity for libraries to manage and make available different types of digital content. however, lee and tibbo (2007) are of the view that digital technologies pose a threat to the permanence of the very objects they produce and require very different management than what has been practiced in the paperbased world. addressing the preservation and long-term access issues for digital resources are some of the key challenges facing libraries and information centres today (alemneh, hastings and hartman, 2002). the challenges are enormous in developing countries due to lack of adequate resources and technologies for effective digital resources management and preservation (boamah, 2014). currently, libraries may manage digital content in three primary ways; providing access to metadata and electronic full-text for publisher or vendor content, managing digitized local collections, and managing institutional, scholarly digital assets (miller and blake, 2011). the researchers work at the national university of science and technology (nust) library in zimbabwe and have observed that the library has embraced the practice of e-collection and digitization by establishing a number of 2/22 ndhlovu, phillip and matingwina, thomas (2018) the state of preparedness for digital curation and preservation: a case study of a developing country academic library, iassist quarterly 42 (3), pp. 1-22. doi: https://doi.org/ 10.29173/iq929 digital collections in order to meet the demands of the 21st century clientele. as a result, the library has amassed a large body of valuable digital assets and information which include among other things institutional records, faculty and student research, theses and dissertations, university publications, past exam papers, multimedia collections and course materials. the digital collections which have been established over the years include; (1) the online public access catalogue (opac) database, (2) the nustone digital library, (3) research guides database, (4) the nust institutional repository (nuspace) as well as a collection of physical storage media such as cdrom, cds, dvds and floppy disks. these digital collections are complemented by providing access to e-books, e-journals, the library website as well as social networking platforms. the mission of one of its digital collections is to collect, preserve and disseminate the intellectual output of nust community and make academic work freely available to researchers and the general public (nuspace, 2016). however, there are no formal mechanisms and sustained strategies in place to ensure accessibility and long-term preservation of digital content at nust library. literature review there is a growing realisation that current and future access to digital resources is threatened by technology obsolescence, fragility of digital media as well digital disasters (mcgovern, 2009). given the dynamic nature of information technologies and the obsolescence issues associated with them, it is important to put in place digital preservation strategies to ensure that digital resources are preserved and remain accessible and useable over time (international records management trust, 2004). several studies have highlighted that it is impossible for digital information to survive or remain accessible by accident and that pro-active preservation strategies are vital. digital resources require well planned, well managed, and sustained strategies over their entire existence (yale university library, 2011). unlike print resources, benign neglect is not an option for digital information because of eventualities such as physical media decay, corruption of digital files and hardware and software obsolescence (miller and blake, 2011). the south african national research foundation (nrf) (2010) notes that sound digital curation and preservation requires inter alia; expertise, human and financial resources, technology infrastructure, adoption of standards, creation of guidelines and implementation of policies. however, most african libraries and information centres are poorly equipped for digital curation and preservation (kanyengo, 2006). a number of studies have highlighted problems that arise when libraries lack policy strategies for digital curation and preservation. these studies also give pointers to elements that should be looked at when assessing preparedness for digital curation and preservation. sinclair et al. (2011) reported on the findings of a planets survey which was done in 2009 of two hundred organisations, mainly european archives and libraries, to investigate their preparedness for digital curation and preservation. preparedness was assessed looking at awareness of digital preservation, digital preservation policies, and digital preservation’s inclusion in organisations’ general planning, budgets for digital preservation and timescales for investment (sinclair et al. 2011). results indicated that organisations without a digital preservation policy were four times more likely to have no experience or be unaware of the challenges presented by digital preservation, three times more likely to have no plans for the long-term management of digital information, and more than twice as likely to put off investing in digital preservation technological solutions. a survey of the preparedness for digital preservation of local authority archives in the united kingdom was conducted by boyle, eveleigh, and needham (2008) where over 80 percent of the respondents already held digital collections. preparedness for digital preservation was assessed in terms of digital preservation planning, general awareness of digital preservation, current practical digital preservation strategies, infrastructure and staffing requirements. the results indicated that awareness of essential issues of digital curation and preservation was particularly low in those organisations without a preservation policy. in the same study, barriers to digital curation and 3/22 ndhlovu, phillip and matingwina, thomas (2018) the state of preparedness for digital curation and preservation: a case study of a developing country academic library, iassist quarterly 42 (3), pp. 1-22. doi: https://doi.org/ 10.29173/iq929 preservation were identified as cultural (organisation, political, awareness, external partnerships/relations and motivation), resources (time, costs, funding and storage), and skills gap (training, competencies and information technology). other studies examined digital disaster preparedness as an essential component of preparedness for digital curation and preservation (zaveri, 2015; frank and yakel, 2013; jiazhen and daoling, 2007). frank and yakel (2013) note that our understanding of disaster planning for digital collections remains limited, both in terms of what constitutes disaster planning activities as well as whether any best practices have emerged in planning for different types of risk. a number of studies have showed that competencies or skills are vital for libraries and information centres to be fully prepared for digital curation and preservation (atkins et al, 2013; groenewald and breytenbach, 2011; hockx‐yu, 2006). however, most studies have concentrated on skills that are required for digital curation and preservation and have been carried out mainly in developed countries (raju, 2014). there are gaps in the literature for studies that assess possession of these skills by librarians working in a digital environment. uluocha (2014) argues that in many african countries, human resources with appropriate skills, competences and attitudes are not readily available to initiate, implement and sustain preservation projects. researchers have noticed gaps in academic libraries, particularly a need for appropriately trained information professionals to act on opportunities for supporting digital curation activities (soehner, steeves and ward, 2010). matsika (2014): , as cited in the nust vice chancellors annual report (2014:55): , notes that more students now have access to lap tops, i-pads and smart phones and prefer the library’s electronic information over print resources. hedstrom (2001) highlighting the importance of digital curation and preservation notes that once users become accustomed to accessing information online they do not want those resources to be removed or diminished. however, for libraries to ensure access to digital information to current and future users, there is need to assess their preparedness for digital curation and preservation in order to acquire the necessary information to allocate resources for their stewardship. 4/22 ndhlovu, phillip and matingwina, thomas (2018) the state of preparedness for digital curation and preservation: a case study of a developing country academic library, iassist quarterly 42 (3), pp. 1-22. doi: https://doi.org/ 10.29173/iq929 1.1 purpose of the study this study aimed to assess the state of preparedness of nust library for digital curation and preservation of its digital collections using the conceptual framework highlighted in figure 1. 5/22 ndhlovu, phillip and matingwina, thomas (2018) the state of preparedness for digital curation and preservation: a case study of a developing country academic library, iassist quarterly 42 (3), pp. 1-22. doi: https://doi.org/ 10.29173/iq929 the study’s conceptual framework was inspired by the criteria used by sinclair et al. (2011) in a survey of two hundred organisations, in mainly european archives and libraries, to investigate their preparedness for digital curation and preservation. it was also inspired by the criteria for preparedness used by boyle, eveleigh, and needham (2008) in a survey of the preparedness for digital preservation of local authority archives in the united kingdom. however, the researchers added some variables which they deemed necessary for assessing preparedness for digital curation and preservation based on the literature review. 1.1 research objectives the study addressed the following important research questions: i. what is the level of librarians’ awareness of digital curation and preservation? ii. to what extent are library staff competencies able to meet digital curation and preservation needs? iii. to what extent is the available information technology infrastructure able to support digital curation and preservation requirements? iv. what is the level of digital disaster preparedness at nust library? v. what challenges are faced by the library in digital curation and preservation? 2. methodology a mixed method research design was chosen for the study. johnson and onwuegbuzie (2004:18) define mixed methods research as the mixing or combination of quantitative and qualitative research techniques, methods, approaches, concepts or language into a single study. the following reasons cited by denscombe (2008) justify the adoption of a mixed research approach: 1) to improve the accuracy of research data 2) to produce a more complete picture by combining information from complementary kinds of data or sources 3) to avoid biases intrinsic to single-method approaches as a way of compensating specific strengths and weaknesses associated with particular methods. the case study approach was adopted as the research strategy for the study. yin (1984:23) defines the case study research method as “an empirical inquiry that investigates a contemporary phenomenon within its real-life context; when the boundaries between phenomenon and context are not clearly evident; and in which multiple sources of evidence are used”. mixed methods research design complements the case study research strategy as it allows the researcher to combine multiple sources of evidence and apply either quantitative and qualitative methods to the data (kitchenham, 2010). the library had a total population of 52 library staff members and purposive sampling was used to recruit 32 participants for the study. the rationale for this approach was that the researcher was able to select respondents who work with digital collections or were in library management and were in a position to provide relevant data. the following participants were included in the study based on the researcher's knowledge and research experience (welman, kruger, and mitchell, 2005: 69): 1) deputy librarian 6/22 ndhlovu, phillip and matingwina, thomas (2018) the state of preparedness for digital curation and preservation: a case study of a developing country academic library, iassist quarterly 42 (3), pp. 1-22. doi: https://doi.org/ 10.29173/iq929 2) library systems librarian/it manager 3) library it technician 4) assistant librarians and 5) chief/senior library assistants questionnaires, observation schedules, interview guides and focus group discussion guide were used as research instruments. the researcher started by individually notifying all respondents of the intention to do research at the nust library through email. details of the study were conveyed along with issues of confidentiality and or anonymity. the researcher made appointments and visited the offices of only two interview participants, the deputy librarian and the library information technology (it) manager to discuss appropriate date, venue and time for the interviews. the interviews were held in their offices at their convenience. the researcher decided to use the services of the student interns to distribute and collect questionnaires from the participants. this ensured strict anonymity of respondents. the researcher consulted the participants about the appropriate time and venue for the focus group discussion. a voice recorder was used with the permission of the participants. observation was done focusing on the items listed in the observation schedule. a digital camera was used to take visual images of some of the observations with the permission of the librarian and the university registrar. to ascertain the validity of the questionnaire, a pilot study was done with 5 respondents from among the library staff. content validity for the study was ensured by identifying all the major independent variables necessary in the existing literature. the pilot study revealed that the participants did not understand the phrase ‘digital curation’. the final questionnaire included a brief explanation in order to make it easier for participants to respond. other minor adjustments were made to the questionnaire based on the results of the pilot study. results of the study were shared with the participants and permission was sought to have their work role titles published in this research article. the study employed both qualitative and quantitative data analysis. descriptive analysis was used to report the study findings. the use of frequency distribution, percentages and averages was implemented to describe the results. miles and huberman’s (1994) qualitative data analysis techniques were used. these are; 1) organization/categorization of the data into concepts ; 2) connection of the data to show how one concept may influence another; 3) corroboration/legitimization by evaluating alternative explanations, disconfirming evidence, and searching for negative cases and 4) representing the account (reporting the findings). 2.1 limitations of the study the study was limited by its nature since ascertaining the preparedness may involve attitudinal investigation and attributes which may be regarded as sensitive and might hinder participants from responding objectively. another limitation was that the researcher is an employee of the nust library. there was, therefore, a possibility of bias on the part of the respondents. however, to mitigate this effect, the researcher carried out interviews and observation to support the stated facts. the field of digital curation is still emerging with many different contributions from a great number of scientists that make it even more difficult to define concepts and theories (palavitsinis, manouselis 7/22 ndhlovu, phillip and matingwina, thomas (2018) the state of preparedness for digital curation and preservation: a case study of a developing country academic library, iassist quarterly 42 (3), pp. 1-22. doi: https://doi.org/ 10.29173/iq929 and sanchez-alonso, 2010). however, the researcher narrowed the scope of digital curation and preservation activities by using the conceptual framework identified in literature. 3. results 3.1 description of participants thirty library staff took part in this study. questionnaires were distributed to 31 members of staff and 29 were returned representing a 94% response rate. however, interviews were conducted with the deputy and library it manager. figure 4.2 shows the number of staff members who participated in the study. the majority were chief/senior library assistants who numbered 19, representing 63% of library staff who took part in the study. figure 1: distribution of library staff by position 3.2 awareness of digital curation and preservation respondents were asked to indicate if they were familiar with the term digital curation or digital preservation. ten (34%) of library staff answered yes and 19 (66%) no. interview results with the deputy librarian revealed that she was aware of the term digital preservation and not digital curation. the library it manager also explained that the term digital curation is new to him. he however explained that he was aware of digital preservation “in the sense of using open file formats”. however, results from the focus group discussion revealed that few were aware of the term digital curation or digital preservation. one respondent said: “we were never taught that in library school”. 3.2.1 level of awareness of digital curation and preservation the researcher needed a way of ascertaining the level of awareness of the library staff concerning digital curation and preservation. therefore, questions that sought to measure the extent of awareness were included in the questionnaire. these included familiarities with different types of metadata and awareness of digital preservation strategies. 0 2 4 6 8 10 12 14 16 18 20 deputy librarian sub librarian assistant librarian chief / senior library assistants library it personnel n u m b er o f li b ra ry s ta ff library staff library staff 8/22 ndhlovu, phillip and matingwina, thomas (2018) the state of preparedness for digital curation and preservation: a case study of a developing country academic library, iassist quarterly 42 (3), pp. 1-22. doi: https://doi.org/ 10.29173/iq929 3.2.2 familiarity with metadata it was interesting to note that every staff member was aware of what metadata is. however, very few were familiar with the different metadata types. the exception was descriptive metadata where 86% of respondents were aware and 14% were not. interview results with the library it manager showed that he was aware of descriptive, technical and administrative metadata and not structural and preservation metadata. as indicated in table 2, most library staff were not completely familiar with most digital preservation strategies that may be applied in libraries. ninety percent were not familiar with emulation and encapsulation; 93% were not aware of technology preservation; 69% were not aware of migration. notably, however, only 31% were not aware of refreshing and 21% of metadata management. table 2: awareness of digital preservation strategies strategy n=29 familiar and have used it familiar and have not used it not familiar refreshing 17% 52% 31% technology preservation 7% 93% migration 7% 24% 69% emulation 10% 90% encapsulation 10% 90% metadata management 69% 10% 21% 3.5 digital curation and preservation competencies of library staff this study asked respondents to rate how they perceived their confidence in executing digital curation and preservation activities. the competencies and skills which were assessed were categorised into 4 categories namely; i. communication and interpersonal competency ii. curating and preserving content competency iii. curation technologies competency and iv. environmental scanning competency. 3.5.1 communication and interpersonal competency communication and interpersonal competencies measures included presentation skills, clear articulation of solutions to information technology problems, communication and advocacy, as well as writing skills in the context of writing persuasive grant proposals. it was interesting to note that the majority of library staff indicated that they had above ‘average’ to ‘excellent’ skills in all these areas. however, on the negative side, a significant number of respondents (62%) indicated that they were ‘poor’ in clear articulation of information technology problems. figure 3 captures the findings of library staff communication and interpersonal competencies. 9/22 ndhlovu, phillip and matingwina, thomas (2018) the state of preparedness for digital curation and preservation: a case study of a developing country academic library, iassist quarterly 42 (3), pp. 1-22. doi: https://doi.org/ 10.29173/iq929 figure 3: communication and interpersonal competency 3.5.2 curating and preserving content competency this category was operationalised into the following skills and competencies; electronic resources management, ability to use controlled vocabularies, ability to assign metadata to digital information, ability to use authority records, ability to use classification schemes and ability to identify file formats (figure 4). figure 4: curating and preserving content competency 0% 20% 40% 60% 80% 100% excellent good average poor not sure unaswered percentage of library staff n= 29 writing skills communication and advocacy clear articulation of solutions to it problems presentation skills 0% 20% 40% 60% 80% 100% electronic resource management ability to use controlled vocabularies ability to assign preservation metadata to digital information ability to use authority records including aacr2 and rda ability to use classification schema such as lcc, ddc, nlmc ability to identify file formats percentage of library staff n=29 not sure poor average good excellent 10/22 ndhlovu, phillip and matingwina, thomas (2018) the state of preparedness for digital curation and preservation: a case study of a developing country academic library, iassist quarterly 42 (3), pp. 1-22. doi: https://doi.org/ 10.29173/iq929 results show that most library staff are good at using classification schemes (90%), using controlled vocabularies (86%), using authority records (72%). however, most library staff rated themselves poor in assigning preservation metadata (86%) and identifying file formats (72%). most of the competencies in this category were traditional library skills such as ability to use classification schemes and using authority records. these were rated mostly above average. however, technical skills identifying file formats and assigning preservation metadata to digital resources were rated as poor. 3.5.3 curation technologies competency most skills in this category were rated as poor by most respondents (figure 5). eighty two percent said they were poor with database development and management, 69% with installing preservation systems, 61% with using digital curation workflow tools and 37% with using file conversion tools. on the positive side, 53% rated themselves as good in the use of scanners. figure 5: curation technologies competency 3.5.4 environmental scanning competency this competency was about the ability of library staff to identify and use online resources to stay current and on the leading edge regarding trends, technologies and practices that affect professional work and capabilities within the field of digital curation. most of the staff rated themselves above average (see figure 6). 0% 20% 40% 60% 80% 100% using digital curation workflow tools using a scanning machine to preserve document installing digital preservation systems e.g dspace using file conversion tools database development and management percentage of library staff n=29 not sure poor average good excellent 11/22 ndhlovu, phillip and matingwina, thomas (2018) the state of preparedness for digital curation and preservation: a case study of a developing country academic library, iassist quarterly 42 (3), pp. 1-22. doi: https://doi.org/ 10.29173/iq929 figure 6: environmental scanning competency the library it manager reported low level of skills to manage digital collections among library staff. he pointed out: “library staff are way behind in terms of skills necessary to deal with digital collections”. he lamented that most staff lacked the motivation to upgrade their skills. the focus group discussion revealed a low confidence among library staff in terms of digital curation and preservation skills. most respondents stated that there was no emphasis on information technology skills in their library school training. 3.6 information technology infrastructure library staff were asked to rate the extent to which they thought nust library had adequate technology infrastructure for digital curation and preservation. figure 7 highlights the results. two categories of infrastructure which were significantly rated as poor included, availability of desktop or work stations (63%), digital curation workflow tools (62%). the focus group discussion matched the ratings for availability of adequate computer workstations as they revealed that mostly senior library assistants were using slow and outdated computers. the deputy librarian also indicated that the technology infrastructure available, has not been replaced for many years. however, 66% of library staff rated as ‘excellent’ the availability of information discovery systems. interview results with the library it manager confirmed the positive ratings for information discovery systems as the library was using a state of the art proprietary discovery system called ebsco discovery tool. most of the technology infrastructure categories were rated as ‘not sure’ and these include storage media (52%), digital preservation management systems (55%) and backup and disaster recovery tools (48%). the focus group discussion confirmed these findings as most library staff revealed that they were not exposed to the library it department and therefore were not familiar with some of the technology infrastructure. the it manager revealed inadequacy of technology infrastructures in terms of servers built purposely for digital collection. he said: “we have converted desktop computers to serve as servers for the library website, nustone digital library and online research guides database and plans are underway to purchase a midrange server”. the it manager also indicated that there was need to revamp the library’s it infrastructure in general. 0% 10% 20% 30% 40% 50% 60% excellent good average poor not sure p e rc en ta ge o f li b ra ry s ta ff n=29 keeping abreast with latest trends on digital curation 12/22 ndhlovu, phillip and matingwina, thomas (2018) the state of preparedness for digital curation and preservation: a case study of a developing country academic library, iassist quarterly 42 (3), pp. 1-22. doi: https://doi.org/ 10.29173/iq929 figure 7: adequacy of information technology infrastructure 3.7 digital disaster preparedness preparedness for digital curation and preservation is also reflected in digital disaster preparedness. the library staff were asked their opinions on the likelihood of a number of digital disasters occurring at nust library as a way of ascertaining their level of awareness of the reality of digital disasters. table 3 summarises the results. adequate desktop/wor k stations storage media backup and disaster recovery tools digital preservation management systems digital curation software workflow tools information discovery systems excellent 3% 0% 0% 66% good 17% 14% 0% 17% 0% 17% average 17% 17% 17% 3% 3% 0% poor 63% 17% 35% 25% 62% 14% not sure 0% 52% 48% 55% 35% 3% 0% 10% 20% 30% 40% 50% 60% 70% p e rc en ta ge o f li b ra ry s ta ff n=29 13/22 ndhlovu, phillip and matingwina, thomas (2018) the state of preparedness for digital curation and preservation: a case study of a developing country academic library, iassist quarterly 42 (3), pp. 1-22. doi: https://doi.org/ 10.29173/iq929 table 3: likelihood of digital disasters digital disaster highly likely likely unlikely not sure fluctuation in power supply 42% 59% 0% 0% power outage 7% 90% 3% 0% software or hardware malfunctions 76% 24% 0% 0% computer viruses 69% 28% 3% 0% hacking of data 66% 31% 3% 0% human errors like spilling of liquids 24% 7% 69% 0% improper computer shutdown 69% 17% 14% 0% accidental deletion of data 48% 45% 0% 7% lightning/heavy rain 31% 14% 55% 0% fire 17% 21% 62% 0% collapse of whole or part of the building 79% 17% 3% 0% most digital disasters in the questionnaire were rated as ‘highly likely’ and the highest was collapse of whole or part of the building (79%) followed by software and hardware malfunctions with 76%. the it manager was also asked about the likelihood of digital disasters occurring. he explained that some digital collections were not being properly backed up such as the nustone digital collection and the nuspace digital collection and were at risk in the event of a digital disaster. he explained that workstation staff machines were being used as backups to save data. 3.7.1 digital disaster strategies at nust library in order to find out the extent to which nust library was prepared for digital disasters, respondents were asked an open ended question as to what strategies the library was using to prevent digital disasters. seven (24%) of the respondents highlighted that they were not aware of any strategy. eight questionnaires were not answered on this particular question. fourteen (48%) of the library staff who responded to the question highlighted the following strategies: i. use of security guards ii. use of surge protectors iii. use of uninterruptible power supply (ups) iv. daily tape backup of digital collections v. use of fire suppression equipment and extinguishers vi. use of antivirus software vii. use of passwords observation results concurred with the above findings and also reviewed the use of fire detection system, closed circuit television (cctv), fire extinguishers and air conditioning of the server room. 14/22 ndhlovu, phillip and matingwina, thomas (2018) the state of preparedness for digital curation and preservation: a case study of a developing country academic library, iassist quarterly 42 (3), pp. 1-22. doi: https://doi.org/ 10.29173/iq929 interview results with the library it manager revealed that not all it equipment was covered with the uninterruptible power supply mechanism. the it manager also indicated that backups of digital content were not regularly tested for readability and were not stored at a remote location in event of a digital disaster. 3.7.2 procedural and technical measures focus group discussion confirmed the above findings and some staff members indicated that the library is covered under the university wide insurance in case of disasters. some respondents said the following concerning insurance: “there is insurance cover whereby if we lost equipment we will get money to buy the equipment”. “i understand the library equipment is insured 100% ... if we lost everything we get it back”. interview results with the deputy librarian and the it manager also concurred with the above observations. however, when asked whether the library is fully prepared for digital preservation, the deputy librarian pointed out that lack of funding was hindering the library from purchasing state of the art digital disaster mechanisms. she also lamented: “fire canisters have not been serviced for the past two years”. she also highlighted: “the library does not have any policies to deal with digital disasters or any disaster for that matter”. focus group discussions revealed that library staff had been trained only on using fire canisters twice in the last 5 years. the library it manager highlighted some unique issues that were not mentioned by other participants. he pointed out that the library is using open standards for file formats and data encoding in its digital collections as a way of reducing risk to inaccessibility of digital materials. he also pointed out the use of firewalls to protect library’s digital collections and prevent unauthorised outsiders from accessing information. to this effect, the it manager stated: “the other measure we implemented to protect our systems and data from outside access, is putting in place a perimeter firewall which is cisco 5520 certified”. the it manager also highlighted the use of rights and privileges to ensure protection of digital collections. rights and privileges were used to control access to unauthorized data. this was normally done depending on the functions or duties of each different person using the data or databases and applications in the network. the it manager also highlighted the web access management system used to authenticate users who mainly want to access electronic journals and books off campus. 3.8 challenges to digital curation and preservation respondents were asked to indicate the challenges that they thought were making it hard for the library be fully engaged in digital curation and preservation activities. library staff members indicated most of the challenges indicated in the questionnaire were critical, with all respondents (100%) citing lack of funding and policy as major problem (see table 4). lack of technology infrastructure was rated a major challenge (93%, lack of skills (86%) and resistance to change (83%). table 4: challenges of digital curation and preservation challenge n= 29 major challenge minor challenge not a challenge lack of skills 86% 14% technological obsolescence 69% 21% 10% 15/22 ndhlovu, phillip and matingwina, thomas (2018) the state of preparedness for digital curation and preservation: a case study of a developing country academic library, iassist quarterly 42 (3), pp. 1-22. doi: https://doi.org/ 10.29173/iq929 lack of technology infrastructure 93% 7% lack of funding 100% lack of policy 100% resistance to change 83% 17% not in library’s mission 17% 52% 31% copyright clearance 66% 28% 7% focus group discussions confirmed findings of the questionnaire as lack of funding, policy and skills were highlighted by most respondents. the deputy librarian and the library it manager also echoed the same sentiments. the library it manager also highlighted that there is low appreciation of the value of preserving digital information among the library key decision makers. 4 discussion of findings 4.1 awareness of digital curation and preservation the results of this study showed that only 32% of the library staff members were aware of the term digital curation or digital preservation and the majority; 68% were not aware. the low level of awareness among library staff members is consistent with (ball, 2010) who highlights that the term ‘digital curation’ is relatively new, having been coined in 2001 as the title for a seminar on digital archives, libraries and e-science. the researcher probed the level of awareness further by asking library staff about their familiarity with different metadata types and awareness of digital preservation strategies. results showed that 83% were not aware of structural metadata, 76% with administrative metadata, 93% with preservation metadata and 79% with technical metadata. however, most library staff (86%) were aware of descriptive metadata. this might be because the digital collections the librarians dealt with require them to assign metadata. results are consistent with those of mutwiri (2014) who found high levels of awareness among library staff of descriptive metadata (86%) and high levels of unawareness of other metadata types; 71% structural metadata, 69% administrative metadata, 73% preservation metadata and 78% technical metadata. the study also revealed high levels of unawareness of digital preservation strategies; 90% were not familiar with emulation and encapsulation; 93% were not aware of technology preservation; 69% were not aware of migration. results are consistent with those of igberaese, sambo and saliu (2014) who found out low levels of awareness of digital preservation strategies by librarians in various libraries and institutions across nigeria. in this study, 14% of the participants had acquaintance with migration; 16% with emulation and 14% with technology preservation. the results were also consistent with those of mutwiri (2014) who found out low levels of awareness of between 18.8% and 43% in an investigation of library staff awareness of preservation strategies for digital materials. similarly, groenewald and breytenbach (2011) found low levels of awareness in south african libraries. however, this study also discovered that some librarians might be familiar with some of the digital preservation strategies but might have not been given the opportunity to use them in practice. 4.2 digital curation and preservation competencies excellent interpersonal, oral, written and online communication skills are highly desirable for personnel who work in the digital curation and preservation field (engelhardt, strathmann and 16/22 ndhlovu, phillip and matingwina, thomas (2018) the state of preparedness for digital curation and preservation: a case study of a developing country academic library, iassist quarterly 42 (3), pp. 1-22. doi: https://doi.org/ 10.29173/iq929 mccadden, 2012). it was interesting to note that the majority of the library staff members indicated that they had above ‘average’ to ‘excellent’ skills in the communication and interpersonal competency. this seems to suggest that communication and interpersonal competencies are generic skills which can be easily applied in the digital curation and preservation field. however, on the negative side, a significant number of respondents (62%) indicated that they were ‘poor’ in the clear articulation of information technology problems. this could be due to the fact that librarians are not exposed to the library it department as interview results with the deputy librarian showed. results of the curating and preserving content competency ratings showed that most library staff were good at using classification schemes (90%), using controlled vocabularies (86%) and using authority records (72%). these are traditional library skills and this might explain why library staff rated themselves highly. this also agrees with bahr, lindlar and vlaeminck (2011) who note that digital curation and preservation utilizes traditional library skills. however, most library staff rated themselves poor in assigning preservation metadata (86%) and identifying file formats (72%). the results were also similar to those obtained for curation technology competencies were rated poorly. eighty two percent said they were poor with database development and management, 69% with installing preservation systems, and 61% with using digital curation workflow tools. these skills are more on the technical side and might be explained by focus group discussion which reveals that the library school curriculum covers less practical information technology skills. 4.3 information technology infrastructure two categories of infrastructure which were rated as poor, availability of desktop or work stations (63%), digital curation workflow tools (62%). this could be explained by focus group discussions which indicated that mostly senior library assistants have not received any new machines for the past 15 years. they relied mostly on donated machines. this is consistent with the results obtained by zaveri (2015) who highlighted that libraries had not developed adequate it infrastructure to support digital resources in india. most of the technology infrastructure categories were rated as ‘not sure’ and include storage media (52%), digital preservation management systems (55%) and backup and disaster recovery tools (48%). focus group discussions revealed that these technologies are in the custody of the library it department which explains their limited knowledge. however, the library it manager also confirmed that these technology infrastructures are inadequate. according to (icpsr, 2009) technological infrastructure which are needed for digital curation and preservation include the requisite equipment, software, hardware, which include operating systems, computers and storage media. however, it is apparent that nust library lacks adequate infrastructure to support digital curation and preservation activities. 4.4 digital disaster preparedness the study sought to find out what the respondents thought about the level of preparedness for digital disasters. the study examined the likelihood of digital disasters at nust library, confidence in handling digital disasters and digital disaster prevention strategies being employed at the library. results showed that most digital disasters in the questionnaire were rated as ‘highly likely’ and ‘likely’ and the highest was the collapse of whole or part of the building (79%) followed by software and hardware malfunctions with 76%. this is in contrast to the findings of zaveri (2015) in an investigation of digital disasters in india where the results of the study indicated that just over 50% of the librarians perceived less than 20% chance of digital disasters, while only 7.61% perceived a probability of over 60%. the results of zaveri were, however, due to the fact that the proportion of digital resources were relatively low compared to print resources. observation results revealed a heavy reliance on digital resources at nust library while digital disaster preparedness was lacking. this might explain why respondents rated digital disasters as highly likely and likely. 17/22 ndhlovu, phillip and matingwina, thomas (2018) the state of preparedness for digital curation and preservation: a case study of a developing country academic library, iassist quarterly 42 (3), pp. 1-22. doi: https://doi.org/ 10.29173/iq929 the strategies that nust library was using to prevent digital disasters as indicated by the study results included use of security guards, use of surge protectors, use of interrupted power supply (ups), daily tape backup of digital collections, use of fire suppression equipment and use of antivirus software. this is in line with study findings by njoroge (2014) in an investigation of disaster preparedness and mitigation for computer based information systems in selected university libraries in kenya. results revealed a number of measures that libraries were using to prevent digital disasters. these included physical measures, procedural measures, technical measures and awareness of computer based information systems (cbis) measures. the it manager also indicated that backups to digital content were not regularly tested for readability and were not stored at a remote location in the event of a digital disaster. this is inconsistent with the recommendations by mutula and wamukoya (2007) who state that procedures for making backups should ensure that backup information should be stored at a location remote to the main site and backup data should regularly be tested for readability. 4.5 challenges to digital curation and preservation the results for challenges to digital curation and preservation showed that 100% of the library staff members regard lack of funding and lack of policy as major challenges. other major challenges include lack of skills (86%), lack of technology infrastructure (69%) and resistance to change. the results are similar to those of kenney and buckley (2005) in a survey carried out by the cornell university library in the united states. the results indicated that the main menace to digital materials were lack of policies and plans inside their institutions to carry out digital curation and preservation. the study results are also in line with duraspace (2013) which found out that 73% of selected academic libraries in the united states cited lack of funding. however, duraspace (2013) in the same study found out that only 23% cited lack of technical expertise and 23% cited lack of administrative support as barriers to digital curation and preservation. this is inconsistent with the study findings which cited lack of skills or technical expertise as a major challenge totaling 86% of the respondents. the library it manager highlighted lack of appreciation by key decision makers on the value of digital preservation at nust library. this is in agreement with voutssas, (2012) who highlights cultural factors which include a lack of awareness among planners and decision makers about the value of digital curation and preservation as challenges facing libraries. 5 conclusions and recommendations despite the fact that nust library has amassed a huge body of digital collections, there are no formal mechanisms in place to ensure accessibility and long-term preservation of digital content. the study sought to assess the preparedness of nust library for digital curation and preservation of its digital collections. the findings of the study revealed a low level awareness of digital curation and preservation. findings also indicated that the library lacks sufficient competencies or skills required for digital curation and preservation. however, the study established that librarians possess mostly traditional library skills which are also relevant in the field of digital curation and preservation. the study findings also showed that the library has inadequate infrastructure for digital curation and preservation and is lagging behind in comparisons libraries in the developed world. the preparedness of the library to handle digital disasters was of concern since no digital disaster plans were in place. the study revealed that most library staff members were not confident in handling digital disasters. however, the library has some mechanisms to prevent digital disasters, although these are not enough. challenges to digital curation were mainly the lack of policies, lack of expertise and lack of funding. the library is not in a position to introduce research data management services with the current state of affairs. the challenges presented in this study will have to be addressed first. 18/22 ndhlovu, phillip and matingwina, thomas (2018) the state of preparedness for digital curation and preservation: a case study of a developing country academic library, iassist quarterly 42 (3), pp. 1-22. doi: https://doi.org/ 10.29173/iq929 5.1 relevance of conceptual framework the conceptual framework was based on previous studies which assessed preparedness for digital curation and preservation of libraries and organisations in the field of information sector. the preceding discussion shows that the conceptual framework captured all the enabling and disabling factors for digital curation and preservation in libraries. 5.3 recommendations the following recommendations are made based on the findings of the study: i. the library should consider digital curation and preservation as one of the primary responsibilities in order to ensure current and future access to digital collections by coming up with policies for managing digital content. ii. digital curation and preservation is highly essential especially for libraries which deal with digital collections. the library should consider sending staff members for specialised training to institutions with established mechanisms for digital preservation to help them to be conversant with digital curation workflows and digital preservation tools. iii. the university should sponsor librarians to attend local and international workshops, conferences, trainings and seminars specific to digital curation and preservation in order to acquire the required skills. iv. the library should consider lobbying all library schools in zimbabwe to overhaul their curriculums in order to factor in digital curation and preservation as part of training required for information professionals. v. the library should consider exposing library staff members to the library it department as part of their duties in order to improve their information technology skills. vi. the library should come up with a disaster policy which incorporates digital disaster recovery plans. vii. the library should ensure training and awareness of current digital disasters strategies such as a fire detection system and use of fire canisters. viii. the library should conduct regular testing of backups to digital content as a way of ensuring that if anything happened to data or databases, then the backups could be used to restore the malfunctioning system. recommendations for further research the results of this study cannot be easily generalised to other university academic libraries in zimbabwe which are also on an accelerated drive to establish digital collections. research is also needed to assess how well they are prepared for digital curation and preservation. this study took a more generic approach in assessing the preparedness of nust library for digital curation and preservation. there is a need for studies to assess specific technology systems being used by nust library for digital preservation in order to provide additional insights on their capabilities to meet digital preservation requirements. 19/22 ndhlovu, phillip and matingwina, thomas (2018) the state of preparedness for digital curation and preservation: a case study of a developing country academic library, iassist quarterly 42 (3), pp. 1-22. doi: https://doi.org/ 10.29173/iq929 references atkins, w. et al. 2013. results of a survey of organizations preserving digital content. [online] available from: http://www.digitalpreservation.gov/documents/ndsa-staffing-survey-reportfinal122013.pdf [accessed 2017 june 6] ball, a. 2010. review of the state of the art of the digital curation of research data. [online] available from: http://opus.bath.ac.uk/19022/2/ [accessed 2017 june 6] bähr, t., lindlar, m. and vlaeminck, s. 2011. puzzling over digital preservation – identifying traditional and new skills needed for digital preservation. [online] available from: http://www.ifla.org/past-wlic/2011/217-bahr-en.pdf [accessed 2017 june 8] boamah, e. 2014. towards effective management and preservation of digital cultural heritage resources: an exploration of contextual factors in ghana. [online] available from: http://researcharchive.vuw.ac.nz/xmlui/bitstream/handle/10063/3270/thesis.pdf?sequence=2 [accessed 2017 june 10] boyle, f., eveleigh, a. and needham, h. 2008. report on the survey regarding digital preservation in local authority archive services. [online] available from: http://www.dpconline.org/index.php?option=com_docman&task=doc_download&gid=338 [accessed 2016 june 22] denscombe, m. 2008. communities of practice: a research paradigm for the mixed methods approach. journal of mixed methods research (2)3: 270-283. [online] available from: http://citeseerx.ist.psu.edu/viewdoc/download?doi=10.1.1.473.9571&rep=rep1&type=pdf [accessed 2017 oct 28] duraspace, 2013. managing digital collections survey results summary. [online] available from: http://www.duraspace.org/sites/duraspace.org/files/managing%20digital%20collections%20survey %20results%20summary.pdf [accessed 2018 feb 26] engelhardt, c., strathmann, s. and mccadden, k. 2012. digital curator vocational education europe: report on survey of sector training needs. [online] available from: http://www.adameurope.eu/prj/6880/prj/d3.1_trainingneedssurvey.pdf [accessed 2018 jan 21] frank, r.d. and yakel, e. 2013. disaster planning for digital repositories. [online] available from: https://www.asis.org/asist2013/proceedings/submissions/papers/59paper.pdf [accessed 2017 july 16] groenewald, r. and breytenbach, a. 2011. the use of metadata and preservation methods for continuous access to digital data. the electronic library, (29)2:236 – 248. [online] available from: http://dx.doi.org/10.1108/02640471111125195 [accessed 2017 june22] hedstrom, m. 2001. digital preservation: problems and prospects. [online] available from: http://www.dl.slis.tsukuba.ac.jp/dljournal/no_20/1-hedstrom/1-hedstrom.html [accessed 2017 june 18] hockx‐yu, h. 2006. digital preservation in the context of institutional repositories. program journal (40)3: 232 – 243. [online] available from: http://www.digitalpreservation.gov/documents/ndsa-staffing-survey-report-final122013.pdf http://www.digitalpreservation.gov/documents/ndsa-staffing-survey-report-final122013.pdf http://opus.bath.ac.uk/19022/2/ http://www.ifla.org/past-wlic/2011/217-bahr-en.pdf http://researcharchive.vuw.ac.nz/xmlui/bitstream/handle/10063/3270/thesis.pdf?sequence=2 http://www.dpconline.org/index.php?option=com_docman&task=doc_download&gid=338 http://citeseerx.ist.psu.edu/viewdoc/download?doi=10.1.1.473.9571&rep=rep1&type=pdf http://www.duraspace.org/sites/duraspace.org/files/managing%20digital%20collections%20survey%20results%20summary.pdf http://www.duraspace.org/sites/duraspace.org/files/managing%20digital%20collections%20survey%20results%20summary.pdf http://www.adam-europe.eu/prj/6880/prj/d3.1_trainingneedssurvey.pdf http://www.adam-europe.eu/prj/6880/prj/d3.1_trainingneedssurvey.pdf https://www.asis.org/asist2013/proceedings/submissions/papers/59paper.pdf http://dx.doi.org/10.1108/02640471111125195 http://www.dl.slis.tsukuba.ac.jp/dljournal/no_20/1-hedstrom/1-hedstrom.html 20/22 ndhlovu, phillip and matingwina, thomas (2018) the state of preparedness for digital curation and preservation: a case study of a developing country academic library, iassist quarterly 42 (3), pp. 1-22. doi: https://doi.org/ 10.29173/iq929 http://www.emeraldinsight.com/doi/full/10.1108/00330330610681312 [accessed 2017 aug 30] icpsr, interuniversity consortium for political and social research. 2009. principles and good practice for preserving data. [online] available from: http://www.ihsn.org/home/sites/default/files/resources/ihsn-wp003.pdf [accessed 2017 aug 25] international records management trust, 2004. international records management trust. [online] available from: http://www.irmt.org/ [accessed 2018 jan 5] jiazhen, l. and daoling, y. 2007. status of the preservation of digital resources in china: results of a survey. program journal, (41)1:35 – 46. [online] available from: http://www.emeraldinsight.com/doi/full/10.1108/00330330710724872 [accessed 2018 feb 6] johnson, r.b. and onwuegbuzie, a.j. 2004. a research paradigm whose time has come. educational researcher, (33)7: 14-26. [online] available from: http://edr.sagepub.com/content/33/7/14.full.pdf+html [accessed 2018 mar 5] kim, j., warga, e. and moen, w. 2012. digital curation in the academic library job market. [online] available from: https://www.asis.org/asist2012/proceedings/submissions/283.pdf [accessed 2017 aug 20] kitchenham, a.d.2010. mixed methods in case study research. [online] available from: https://srmo.sagepub.com/view/encyc-of-case-study-research/n208.xml [accessed 2017 oct 30]. lee, c.a., tibbo, h.r. and schaefer, j. 2007. defining what digital curators do and what they need to know: the digccurr project. paper presented at the 7th acm/ieee-cs joint conference on digital libraries (jcdl), vancouver, british columbia, canada. miles, m.b. and huberman, a.m. 1994. qualitative data analysis. thousand oaks, calif. sage. mutula, s. and wamukoya, j. 2007. web information management: a cross-disciplinary textbook. oxford: chandos publishing mutwiri, c.m. 2014. challenges facing academic staff in adopting open access outlets for disseminating research findings in selected university libraries in kenya. [online] available from: http://ir-library.ku.ac.ke/bitstream/handle/ [accessed 2017 july 2] national research foundation. 2010. managing digital collections: a collaborative initiative on the south african framework. pretoria: national research foundation. njoroge, r.w. 2014. an investigation on disaster preparedness and mitigation for computer based information systems in selected university libraries in kenya. [online] available from: http://irlibrary.ku.ac.ke/bitstream/handle/ [accessed 2017 july 4] nust, national university of science and technology vice chancellor’s annual report, 2014. nust library 2014 annual report. bulawayo: nust department of communication and marketing. pennock, m. 2006. managing digital cultural heritage resources: from digital creation to digital curation. information scotland, (1):1-3. [online] available from: http://www.ukoln.ac.uk/ukoln/staff/m.pennock/publications/docs/info-scotland_200610.pdf [accessed 2017 aug 25] http://www.emeraldinsight.com/doi/full/10.1108/00330330610681312 http://www.ihsn.org/home/sites/default/files/resources/ihsn-wp003.pdf http://www.irmt.org/ http://www.emeraldinsight.com/doi/full/10.1108/00330330710724872 http://edr.sagepub.com/content/33/7/14.full.pdf+html https://www.asis.org/asist2012/proceedings/submissions/283.pdf https://srmo.sagepub.com/view/encyc-of-case-study-research/n208.xml http://ir-library.ku.ac.ke/bitstream/handle/ http://ir-library.ku.ac.ke/bitstream/handle/ http://ir-library.ku.ac.ke/bitstream/handle/ http://www.ukoln.ac.uk/ukoln/staff/m.pennock/publications/docs/info-scotland_200610.pdf 21/22 ndhlovu, phillip and matingwina, thomas (2018) the state of preparedness for digital curation and preservation: a case study of a developing country academic library, iassist quarterly 42 (3), pp. 1-22. doi: https://doi.org/ 10.29173/iq929 raju, j. 2014. knowledge and skills for the digital era academic library. the journal of academic librarianship, (4002:163–170. [online] available from: http://www.sciencedirect.com/science/article/pii/s009913331400024x [accessed 2017 july 15]. sinclair, p. et al. 2011. are you ready? assessing whether organisations are prepared for digital preservation. the international journal of digital curation (6)1: 268-281. [online] available from: http://www.ijdc.net/index.php/ijdc/article/viewfile/178/247 [accessed 2017 sep 15] soehner, c., steeves, c and ward, j. 2010. e-science and data support services: a study of arl member institutions. [online] available from: http://www.arl.org/storage/documents/publications/escience-report-2010.pdf [accessed 2017 sep 15] uluocha, a. 2014. imperatives for storage and preservation of legal information resources in a digital era. information and knowledge management, (4)5:52-58. [online] available from: http://www.iiste.org/journals/index.php/ikm/article/download/12905/13246 [accessed 2017 oct 1] voutssas, j. 2012. long term digital information preservation: challenges in latin america. aslib proceedings, (64)1:83 – 96. [online] available from: http://www.emeraldinsight.com/doi/full/10.1108/00012531211196729 [accessed 2018 feb 6] welman, c., kruger, f., and mitchell, b. 2005. research methodology. oxford : oxford university press. yin, r. k. 1984. case study research: design and methods. newbury park, ca: sage. yale university library, 2011. best practice digital resource maintenance and preservation strategies. [online] available from: http://www.library.yale.edu/iac/dpc/maintenancepreservationstrategies09apr2007.pdf [accessed 2017 july 21] zaveri, p. 2015. digital disaster management in libraries in india. library hi tech (33)2: 230 – 244. [online] available from: http://www.emeraldinsight.com/doi/full/10.1108/lht-09-2014-0090 [accessed 2017 july 9] end-notes 1 phillip ndhlovu is the institutional repository librarian at the national university of science and technology. he is also the liaison librarian for the faculty of commerce and can be reached by email: ndhlovup@gmail.com (version: july 2018) 2 thomas matingwina is a lecturer in the department of library and information science at the national university of science and technology. (version: july 2018) http://www.sciencedirect.com/science/article/pii/s009913331400024x http://www.ijdc.net/index.php/ijdc/article/viewfile/178/247 http://www.arl.org/storage/documents/publications/escience-report-2010.pdf http://www.iiste.org/journals/index.php/ikm/article/download/12905/13246 http://www.emeraldinsight.com/doi/full/10.1108/00012531211196729 http://www.library.yale.edu/iac/dpc/maintenancepreservationstrategies09apr2007.pdf http://www.emeraldinsight.com/doi/full/10.1108/lht-09-2014-0090 mailto:ndhlovup@gmail.com 22/22 ndhlovu, phillip and matingwina, thomas (2018) the state of preparedness for digital curation and preservation: a case study of a developing country academic library, iassist quarterly 42 (3), pp. 1-22. doi: https://doi.org/ 10.29173/iq929 vol30-2neu.indd 18 iassist quarterly summer 2006 by micah altman and gary king* critical components of the scholarly and library community are use of a common language and universal standards for scholarly citations and credit attribution, to enable the location and retrieval of articles and books. we present a proposal for a similar universal standard for citing quantitative data that retains the advantages of print citations, adds other components made possible by, and needed due to, the digital form and systematic nature of quantitative datasets, and is consistent with most existing subfield-specific approaches. although the digital library field includes numerous creative ideas, we limit ourselves to only those elements that appear ready for easy practical use by scientists, journal editors, publishers, librarians, and archivists. we propose that citations to numerical data include, at a minimum, six required components. the first three components are traditional, directly paralleling print documents. they include the author(s) of the data set, the date the data set was published or otherwise made public, and the data set title. these are meant to be formatted in the style of the article or book in which the citation appears. the author, date, and title are useful for quickly understanding the nature of the data being cited, and when searching for the data. however, these attributes alone do not unambiguously identify a particular data set, nor can they be used for reliable location, retrieval, or verification of the study. thus, we add three components using modern technology, each of which is designed to persist even when the technology inevitably changes. they are also designed to take advantage of the digital form of quantitative data. the fourth component is a unique global identifier, which is a short name or character string guaranteed to be unique among all such names, that permanently identifies the data set independent of its location. we allow for any naming scheme to be chosen, so long as it (1) unambiguously identifies the data set object, (2) is globally unique, and (3) is associated with a naming resolution service that takes the name as input and shows how to find one or more copies of the identical data set. long-term persistence of the resolution service is meant to be guaranteed by the organization that operates it, although it is now becoming common to set up redundant multiple naming resolution services, so that archives can back each other up in case one goes out of business. unique global identifiers guarantee persistence of the link from the citation to the object, but we also need to guarantee and independently verify that the object does not change in any meaningful way, even when data storage formats change. to address this need, we add as the next component a universal numeric fingerprint, or unf. the unf is a short, fixed-length string of numbers and characters that summarize all the content in the data set, such that a change in any part of the data would produce a completely different unf. a unf works by first translating the data into a canonical form with fixed degrees of numerical precision, and then applies a cryptographic hash function to produce the short string. the advantage of canonicalization is that unfs (but not raw hash functions) are formatindependent: they keep the same value even if the data set is moved between software programs, file storage systems, compression schemes, operating systems, or hardware platforms. finally, since most web browsers do not currently recognize global unique identifiers directly (i.e., without typing them into a web form), we add as a final component of the citation standard a bridge service, which is designed to make this task easier in the medium term. given how web services are accessed presently, the bridge service should be a url, which can thus be recognized by any browser. we also offer a systematic way to add information to data citations that also retains complete flexibility in added content. for each added element, we recommend a two-part syntax composed of: the value of the content, a field name that describes the content being added, and an (optional) semicolon separator. for example: “value [fieldname];” or “ interuniversity consortium for political and social research [distributor];”. to encourage standardization, ion we recommend that field names be drawn from the ddi 2.1 specification elements for study and variable descriptions. if others are needed, additional items may be drawn from other metadata schemes and vocabularies by adding the identifier for that scheme in parentheses within the bracketed field name, such as “dataset [type (dc)]” or “current population survey supplements [series (iso 690-2)]”. in unusual cases, users overview of a proposed standard for the scholarly citation of quantitative data iassist quarterly summer 2006 19 could even easily add their own vocabulary if needed. this extended standard can be used to create citations similar to and compatible with some existing approaches, such as iso 690-2 (see iso, 1997) (although some aspects of these approaches may now be obsolete). together, the global unique identifier, unf, and bridge service ensure permanence, verifiability, and accessibility even in situations where the data are confidential, restricted, or proprietary; the sponsoring organization changes names, moves, or goes out of business; or new citation standards evolve. together with the author, title, and date, which are easier for humans and search engines to understand, all elements of the proposed full citation for quantitative data should achieve what print citations do and, in addition to being somewhat less redundant, take advantage of the special features of digital data to make the citation considerably more functional. * this extended abstract summarizes the proposed standard. this was presented at the iassist 2006 conference in ann arbor at the session “new standards in statistics and data citations” by micah altman, harvard university. micah altman is associate director, harvard-mit data center & senior research scientist, institute for quantitative social science; harvard university. gary king is david florence professor of government, institute for quantitative social science, harvard university our thanks to the national science foundation (ses0318275, iis-9874747) and the national institutes of aging (p01 ag17625-01) for research support. for full details see: micah altman & gary king, 2007. “a proposed standard for the scholarly citation of quantitative data”, d-lib magazine 13(3). http://dlib. org/dlib/march07/03contents.html iassist quarterly 41 australian health statistics by roger jones ' social science data archives. australian national university. canberra. australia. introduction the impetus for this paper was a workshop held earlier this year to identify current problems associated with australian health statistics and determine priorities for a national health data base. generally, academic researchers in australia have focussed on aetiology, that is the study of causes of diseases, and attention will be given to the data used to develop aetiological hypotheses and, ultimately, test these hypotheses. by identifying the major causes of avoidable death, disease or disability as well as possible with existing data, the gaps in the data will be identified and priorities for the collection of new data will be established. 'paper prepared for presentation at the ifdo/iassist conference. amsterdam, may 20 23, 1985. editors note: unfortunately the accompanying, tables were not submitted with this paper. we have therefore deleted all specific table references. with apologies to the author. data requirements increased attention is being given to the need for improved occupational health surveillance in western societies. in australia, concerns about potential health effects of exposure to the herbicide 2,4,5-t and its dioxin contaminant, vietnam war service in general, proximity to atomic testing, past employment in uranium mines, and exposure to asbestos and lead are prominent among the issues that have focussed attention on occupational risks. the risks of cancer resulting from occupational exposures have been discussed widely, and the patterns of accidents, injuries and illnesses associated with industry and occupation examined. personal and lifestyle factors such as diet, cigarette smoking, alcohol comsumption, stress and drug use are also now recognised as important 'risk factors' associated with the health status of the population. there is now evidence to support the view that the chronic, degenerative diseases such as heart disease and cancer which are the major killers of australians substantially result from these socio-behavioural factors rather than simply being the diseases of old age. the main causes of death and hospitalization in australia, for people below the age of 45, are motor vehicle accidents and other accidents, poisoning and violence, including suicide. for older people the chronic diseases of ischaemic heart disease, stroke and other diseases of the circulatory system and various cancers predominate. a number of factors are known to influence, or to be associated with these causes of death. the most important behavioural factors, in terms of the amount of resultant disease, are almost certainly smoking and alcohol consumption. the commonwealth department of health suggested that there were 16,200 summer 1986 42 iassist quarterly deaths in australia associated with tobacco use in 1980. equivalent to 110 deaths per 100.000 population. the diseases with which smoking is associated are clearly indicated in the warning ' smoking causes lung cancer, as well as heart and other lung disease ' which the nh & mrc (editor's note:) (national heath and medical research council) recommended to replace the less specific ' smoking is a health hazard' on cigarette packets sold in australia. almost half the number of deaths attributed to alcohol are associated with road traffic accidents, from which disability and injury also result diet is also impucated in several of the major causes of alcohol-related deaths. descriptive studies considerable difficulties arise when trying to establish causal links with diseases, such as the argument over smoking and cancer. individual studies are rarely definitive and considerable time may elapse before sufilcient evidence is available to assume a causal relationship. such studies can also be expensive. accordingly, less intensive descriptive studies are undertaken in the first instance, using readily available data sources from which crude measures may be derived in order to generate or provide an initial check of hypotheses. fully developed analytical studies would only follow when the results of exploratory investigations £ire sufficiently suggestive and important to wanant the expense of more rigorous confirmatory research. information for general descriptive studies may be available from routine collections of health event data or from prevalence surveys of population samples. routine health collections the routine health reporting systems in australia are poorly developed in comparison to other developed coimtries, particularly at the national level. under the federal/slate system of government that operates in australia, state governments are primarily responsible for the provision of hospital and health services within their own borders, and for the collection of health event data associated with these services. while all states administer similar collections, substantial variation exists in terms of coverage, scope and means of collection. mortality data play an important role as health status indicators. data are compiled in each state by the registrars of births, deaths and marriages and causes of death coding is added by the australian bureau of statistics (abs) from medical certificates. at the present time, there is no national compilation, although health ministers have agreed to establish a national death index, subject to appropriate confidentiality legislation being enacted. to obtain australia-wide mortality tapes at present, with personal identifiers removed, separate tapes for each state would have to be developed by the abs and sent to the state registrars for release, subject to their approval, and re-aggregated by the researcher. clearly this is a very cumbersome procedure. the main public source of data at present is the abs publication 'causes of death'. hospital morbidity collections based on data collected on patients leaving hospitals are compiled by the health authories in each state, providing limited information on the demographic background and illnesses of hospital in-patienls. in all states, data are collected on separations of public patients from public hospitals, but private hospitals, psychiatric hospitals and nursing homes may be included in some states and not in others, so that summer j 986 iassist quarterly 43 comparisons between slates and aggregation actoss states cannot be accomplished from published results, although a uniform minimal data set could be achieved with access to the computerised records. records of primary contact between the public and health professionals such as private medical practitioners, community health centres, hospital clinics and casualty services are totally lacking at present, at least in any unified form. however, the introduction of the universal health insurance scheme. medicare, in february 1984 has created an opportunity to obtain more information in this area. medicare gives automatic entitlement to a subsidy of 85% of the schedule fee for medical services provided by private practitioners, and approximately 100 miuion records per year relating to claims are stored and interfaced with a medicare enrolment file containing personal data. however, as yet, only the information necessary for payment of benefit is being collected: name, date of birth, sex, geographic area, and usual residence, and a provider's identifier. no data are collected about either diagnoses or procedures. with regard to specific diseases, each state now has a cancer registry based on notifications by hospitals and laboratories of cancer patients and by registrars of deaths attributed to cancer. a proposal to establish a national cancer statistics gearing house is currently before the nh & mrc and, if accepted, national figures should be available within a few years. instances of communicable diseases are reported to state health authorities and collated by the department of health but are of very hmited research value. information on work-related health problems can be obtained from compensation claims to the stale-based insurance schemes, but large numbers of workers are not covered and there would be under-reporting because neither medical practitioners nor workers are aware of the significance of work factors in the aetiology of many chronic diseases. police, legal authorities and traffic authorities collect information which yields statistics on road traffic accidents, and on drink-driving offences and drug offences. at present, these have no research value, but the department of transport is developing a national road traffic accidents data base based on police reports of accidents involving fatalities and casualties. how can these routine reporting systems be used to develop and check hypotheses about causes of disease? the traditional approach is to produce estimates of the risk of various health outcomes across various subgroups of the population under study. thus, for example, an indicator of the health risk associated with particular occupations is given by ratios of the number of deaths from various causes to the population at risk and making comparisons across occupational categories. the population at risk is usually obtained from census data. clearly there are severe limitations with this t>'pe of unlinked data analysis, particularly in the lack of ability to apply controls for what might be relevant covariables. this is limited to those factors coded in both the census and health event data, and there are usually very few in the latter age, sex, location and perhaps occupation, often in groups too broad to be useful. the value of this type of data is greatly enhanced when it can be linked directly at individual level to data on personal characteristics and lifestyle factors. in this case, files must include full names, previous surnames and dates of birth at least concern about confidentiality is, however, very high in australia and the opportunities for record linkage are very limited, although some useful work has been done by linking cancer records and death records to employment records. however, identifiable census returns have been destroyed in australia since the start of the century' and proposals that a sample of these be retained, as has been done recently in england and wales, seem unlikely to succeed. summer 1986 44 iassist quarterly population surveys given the poor state of routine health data collections, it is fortunate, though perhaps not unrelated, that australia shows up reasonably well in the area of health surveys in the international health data guide . this is probably due to the preference of the australian bureau of statistics for using its limited resources for population surveys which can cover a range of topics rather than for more specific health related collections. in additioa state health authorities and other agencies have undertaken a number of surveys relating to lifestyle factors, particularly cigarette smoking, alcohol consumption and drug use. prevalence surveys of population samples have the advantages that, for established risk factors, such as smoking, they can be used to describe cross-sectional and longitudinal variations which may relate to health outcomes, and they provide a basis for evaluating community behaviour intervention programmes. for postulated or possible risk factors, they can be used to examine relationships with health outcomes and may provide clues about causal factors in disease. the australian health surveys in particular provide valuable information on the health status of the population, particularly since recent changes to the abs act now permit the release of de-identified unit record data. a data file from the 1977-78 survey has been released, and the 1983 data file should be available later this yejir. a further survey is planned for 1986. the surveys give detailed information on a wide range of personal characteristics, although the categories of such important variables as occupation and birthplace were too broad in the unit record file for many research purposes. categories had been collapsed in order to ensure non-identifiability of respondents. particularly those in small subgroups of the population. for such subgroups, of course, sample surveys are of uttle value because of the small number, if any, of cases interviewed. nevertheless the practice of collapsing variables into standard classifications without giving sufticient thought to the potential research uses needs to be changed. interview data on self-reported health status may also be a suspect, and it would be useful to have some medical verification of a subsample of respondents or from pilot tests. another shortcoming of these data is the lack of information on health risk factors such as smoking and drinking, although 12 items from the general health questionnaire were included. several community health studies have been carried out in australia in recent years, and these may provide clues to the relationship between life-style characteristics and health status. one of the advantages of such studies is that they often include clinical checks on the self-reported health status of respondents. the main disadvantage is that the sample sizes are generally too small to allow tests of hypotheses. the prevalence of hean disease as a major cause of morbidity and mortality has generated a number of studies aimed at determining the associated life-style factors. the best of these are the two risk factor i*revalence studies conducted by the national heart foundation in 1980 and 1983. both studies used large national random samples and included clinical examinations to obtain height, weight, blood pressure and blood lipid levels in addition to interview data on tobacco, alcohol and medication consumption, diet, physical activity and psychological stress. while major programmes designed to infiuence the smoking behaviour of the community have been launched throughout australia, the abs has chosen to cease conducting surveys on smoking and to exclude questions on smoking summer 1986 iassist quarterly 45 from the health surveys. the last national survey on smoking conducted by the abs was in 1977. a series of much smaller national surveys, on about 6000 adult respondents, have been conducted by the anti-cancer council of victoria in 1974, 1976, 1980 and 1983. however, a large number of school-based surveys of alcohol, tobacco and drug use have been undertaken, although as a basis for national figures these have problems of comparability. two national surveys of school children aged 9-16 years were conducted by the nh & mrc in 1969 and 1973, and a similar survey has recently been carried out by the australian cancer society and the national heart foundation. efforts to reduce the number of road traffic accidents have centred on programmes designed to reduce alcohol use with random breath testing being introduced in most states of australia and substantial expenditure on television advertising campaigns. however, the availabihty of statistical data to evaluate the effectiveness of these approaches is limited largely to the mortality and morbidity statistics and some small studies on knowledge, attitudes and behaviour relating to drink-driving. concerns about invasion of privacy, the lack of legislation on the preservaton of confidentiality for some collections (morbidity), and over-rigorous interpretation of such legislation in others (census), have restricted the use that could be made of these data. population surveys have been adopted as an alternative, but these are generally too small for detailed analytic studies and are thus limited to monitoring the prevalence of established risk factors. the national health statistics workshop held in february expressed its concern over this lack of appropriate data and recommended that a national health statistics agency be established as part of the newly formed australian institute of health. this new agency should promote the development of national collections such the natural death index, cancer index and morbidity collections, and ensure that the necessary legislation is enacted to provide for preservation of confidentiality. p*riority should be given to assembling data already available in most cases at state level into unified national collections, to the development of record linkage procedures, and to risk factor surveys of diet, smoking, alcohol and illegal drugs. concluding comments as indicated in the above brief review of australian health status data, only limited attention has been given to the needs of researchers for aetiological analyses. routine health collections lack uniformity across state boundaries and do not include sufficient information on parents' background or possibly associated risk factors. record linkage could overcome some of these deficiencies but has generally been resisted by the appropriate authorties. this is obviously a large agendum which will require considerable resources and time to implement nevertheless, there is strong support behind the recommendations and a reasonable hope that a substantial improvement in australian health statistics will be achieved, n summer 1986 1/28 campbell, graeme, cuyler, katie & guindon, alex (2025). assessing the landscape for discovery and access to historical canadian census data, iassist quarterly 49(4), pp. 1-28. doi: https://doi.org/10.29713/iq1167 the creative commons-attribution-noncommercial license 4.0 international applies to all works published by iassist quarterly. authors will retain copyright of the work and full publishing rights. assessing the landscape for discovery and access to historical canadian census data graeme campbell1, katie cuyler2 and alex guindon3 abstract the canadian census is a primary source of information about canada and the people who live there, and this information is used by researchers, the private sector, public servants and residents. however, access to canadian census data is fragmented and inconsistent, with no single source of census data for all census years, or in all census data formats. this is a barrier to research, making systematic analysis, discovery, and reuse difficult. this article provides an overview of the current landscape of canadian census portals by data format. it includes an analysis of the coverage and usability of census portals and demonstrates the outstanding need for a single comprehensive access point for canadian census data. keywords canadian census, data discovery, data access, inventory introduction the census is a primary source of information about a country and the people who live there, and an important knowledge infrastructure used by researchers, the private sector, public servants and residents. however, access to canadian census data is fragmented and inconsistent, making systematic analysis, discovery, and reuse difficult. accessing historical census of canada data requires sleuthing and specialized knowledge about databases hosted by many institutions, projects and portals; each providing access to only some census years with incomplete content. there is no single source of census data for all census years or in all census data formats. this difficulty of census search and access presents a barrier to research exploration. this overview includes an analysis of the coverage and usability of census portals and demonstrates the outstanding need for a single comprehensive access point for canadian census data. background the first colonial population count, or census, in the territory now known as canada, was of colony inhabitants in new france in 1665-66 (statistics canada, 2015). this was followed by an inconsistent string of french or british colonial censuses until confederation in 1867 and the enactment of the census act, 1870 that mandated censuses be taken every ten years (statistics canada, 2015). since then, censuses have been conducted at regular intervals, with the exception of 2011, when the long form census was temporarily replaced with the voluntary national household survey (statistics https://doi.org/10.29713/iq1167 2/28 campbell, graeme, cuyler, katie & guindon, alex (2025). assessing the landscape for discovery and access to historical canadian census data, iassist quarterly 49(4), pp. 1-28. doi: https://doi.org/10.29713/iq1167 canada, 2024, november 26). census data is the most long standing comprehensive primary source of demographic information about the social, economic, and cultural aspects of the population that has lived in canada and the previous french and british colonies from the late 1600s to today. diverse products make up the corpus of available canadian census information, including statistical tables, microdata, maps, geospatial data and analytical reports. these products have been published in several formats such as print publications, cd-roms and electronic files. there is also a variety of accompanying and reference documentation useful for census research, such as survey questionnaires and enumerator instructions. the access points (which we also call “portals” in this article) to census products are varied in terms of the types of content and census years that they host. this is largely because of the technologies and digital file formats used to collect and disseminate the census over time, as well as the different objectives each portal had for collecting and analyzing census data. these varied access points make it difficult to conduct analyses with individual census data types longitudinally, or in a systematic way. the types of organizations creating, supporting, and hosting census portals tend to be government entities, universities, academic libraries, research groups, and community organizations. the current census access landscape is largely the result of how these hosting organizations and research projects have operated historically, with current challenges arising due to some portals having limited resources or lacking a mandate for long term content maintenance, or because they have been developed independently of the other projects. the authors of this paper were co-investigators on the canadian census data discovery partnership (ccddp), a social sciences and humanities research council (sshrc) partnership development grant funded initiative. the ccddp has been working to improve discoverability of census data and access to canadian census information by comprehensively inventorying census resources and making this inventory available as an open access database. as ccddp co-investigators, the authors are building upon knowledge gained through the work of the partnership to create the overview of the current landscape of canadian census portals detailed within this paper. scope all the portals included in this paper were initially identified in an early phase of the ccddp project where a comprehensive list of sources of canadian census data was created. since there are numerous access points and some degree of redundancy between them, there was a need to establish criteria to define the scope of this review. as such, instead of attempting an exhaustive inventory of portals to canadian census information, this investigation included portals matching one or more of the following four primary types: (1) portals created and hosted by institutions with legislative responsibility for creating and providing access to census content, such as statistics canada and library and archives canada; (2) portals of non-government projects that provide unique access to census data; (3) portals which provide improved organization to existing census content; and (4) portals most commonly relied upon within an academic, research setting. for instance, we excluded genealogical research sites. https://doi.org/10.29713/iq1167 https://cddp-pddr.ca/ 3/28 campbell, graeme, cuyler, katie & guindon, alex (2025). assessing the landscape for discovery and access to historical canadian census data, iassist quarterly 49(4), pp. 1-28. doi: https://doi.org/10.29713/iq1167 organization to facilitate information retrieval and better organize this review, census portals assessed herein are listed by content type. • census returns (household and individual information as collected through enumeration) • census publications (official government text-based publications issued after each census) • aggregate data (added up census data organized by theme and/or geography) • microdata (raw data observed or collected directly from a specific unit of observation – individual, family, household) • census maps and digital spatial data (digital spatial data are computer files meant to be used with a geographic information system (gis); for the census, they are mainly boundary files of geographic census units) each section starts with a brief definition of the type of materials covered. every section also includes a summary table that presents the main characteristics of the reviewed portals, such as census years covered, type of access, file formats, etc. this is followed by a short description of each portal. an alternative presentation of the section summary tables, as an alphabetical list of all portals inventoried, is included in appendix a. this paper concludes with an introduction to the ccddp. census returns the census of population returns comprise the information collected about individuals and households by the census program through its enumerators or self-enumeration. while early censuses often only enumerated households, from 1851 onwards, information was consistently collected for each individual as well (hillman, 1981, p. viii). there are strict prohibitions on viewing and disclosing census information after it is collected, but generally these rules cease to apply 92 years following the census year (see sections 17 and 18 of the statistics act, rsc 1985, c. s-19 for more information). as a result, the returns from 1931 are the most recently released to the public. in addition to the resources mentioned in this section, some other projects listed in the microdata section also provide census returns for specific periods. returns for some of the earliest preconfederation censuses may also be found via provincial archives, such as those of nova scotia or prince edward island, or on microfilm, as indexed by library and archives canada’s early census and related documents, 1640 to 1945 database or by hillman’s catalogue of census returns on microfilm 1666-1881. for the purposes of our assessment, we only included resources that were national in scope, provided digital online access to census returns, and were not primarily intended for genealogical researchers. https://doi.org/10.29713/iq1167 https://laws-lois.justice.gc.ca/eng/acts/s-19/index.html https://laws-lois.justice.gc.ca/eng/acts/s-19/index.html https://archives.novascotia.ca/census/ https://www.princeedwardisland.ca/en/information/education-and-early-years/genealogy-at-the-public-archives https://library-archives.canada.ca/eng/collection/research-help/genealogy-family-history/censuses/pages/other-census-related-documents.aspx https://library-archives.canada.ca/eng/collection/research-help/genealogy-family-history/censuses/pages/other-census-related-documents.aspx 4/28 campbell, graeme, cuyler, katie & guindon, alex (2025). assessing the landscape for discovery and access to historical canadian census data, iassist quarterly 49(4), pp. 1-28. doi: https://doi.org/10.29713/iq1167 website/portal creator access coverage formats host type héritage canadiana/crkn public 17th-19th centuries (selected) pdf, jpeg non-profit census search library and archives canada public 1825-1931 pdf, jpeg, csv*, and xml* government table 1: characteristics of census returns portals *search results can be downloaded in different formats. héritage the canadian research knowledge network (crkn) is a network of member institutions from the academic, research, and government and public library sectors, with a mandate that includes supporting “the digital infrastructure required to preserve and access critical canadian content” (crkn, 2024a). one way the crkn provides this support is through héritage, a collection of over 40 million digitized pages of historical archival documents, re-hosted and made publicly available by the crkn in a trustworthy digital repository (crkn, 2024b). a selection of census returns are found in héritage, most often as images digitized from microforms. while héritage does not have an advanced search interface, it provides keyword searching of document metadata and full text where available, and advanced search operators are available to refine search results. the census materials are not bundled together for easy browsing, but you can find many of them by searching for “census” or “recensement” in the title field. included are the census returns from lower canada/canada east 1825, manitoba and red river 1831-1870, upper canada/canada west 1841-1861, the city of victoria 1891, and several digitized reels of dépôt des papiers publics des colonies; état civil et recensements, série g 1 : recensements et documents divers, containing records from the general censuses of french colonies back to the 17th century. census records from the department of indian affairs and parish registers from manitoba, new brunswick, newfoundland and labrador, nova scotia, ontario, and quebec are also available. héritage’s content is freely available to the public, and users can download individual pages in jpeg or pdf format, and entire documents as pdfs. library and archives canada library and archives canada (lac)’s primary online historical census data resource is accessed through census search, a tool for searching and exploring census returns. coverage includes the returns from the decennial census for the dominion of canada between 1871 and 1931, the 1870 census of manitoba, the 1906 census of the northwest provinces, the 1916 and 1926 censuses of the prairie provinces, and various pre-confederation censuses of new brunswick, nova scotia, ontario (canada west), prince edward island, and quebec (lower canada, canada east) back to https://doi.org/10.29713/iq1167 https://heritage.canadiana.ca/ https://recherche-collection-search.bac-lac.gc.ca/eng/census/index 5/28 campbell, graeme, cuyler, katie & guindon, alex (2025). assessing the landscape for discovery and access to historical canadian census data, iassist quarterly 49(4), pp. 1-28. doi: https://doi.org/10.29713/iq1167 1825. digitized copies of the returns are hosted online by lac, and new census years are added as they become available as per section 18.1 of the statistics act. a variety of earlier census and related documents from 1640 to 1945 are also available, either digitally or on microform, though they are not currently indexed by personal name like those included in the census search. lac’s census search has basic and advanced search interfaces, allowing users to search by census year, geographic region, and various personal attributes of enumerated individuals (e.g. name, year of birth). matches are associated with digitized pages from enumeration records, which can then be freely downloaded in jpeg or pdf format. search results can also be exported in different formats, including csv and xml, allowing users to review and analyse custom datasets based on their own search criteria, including many (but not all) variables and up to 5000 rows. lac also identifies historical documents and publications related to each census year, like the published volumes of aggregate statistics and instructions to enumerators, in lac’s physical holdings or in digital format, with some often hosted by canadiana or the government of canada publications (gcp) directorate. lac also provides lists of districts and subdistricts for each historical census directly on their website as html content. census publications originally, the principal historical method of communicating census of canada information was in print publications. the practice of publishing the collated census results of multiple provinces in printed volumes stretches back before confederation to at least the 1851-52 census of the canadas and continued post confederation up until 1951. starting in 1956, the bound volumes issued previously were replaced by a series of reports that could be bound or consolidated by other means (canada, 1959, p. 79), and this practice continued until 2006, which was the last year for which the results of the census were published in print format (statistics canada, 2019, p. 5), though statistics canada continues to issue publications online in html and pdf formats that contain analyses of the results of the census and supporting documentation. the publications themselves comprise mostly aggregate data tables and contextual information based on the government’s own analysis of the census returns. as such, their contents vary from year to year based on differences in census questionnaires, geographical coverage, and the analytical aims of the government. while the scope and detail of the published analysis generally increased over time, some of the earlier census years also include publication volumes presenting aggregate data from previous census years (e.g. 1871), providing multi-year comparisons (e.g. 1871, 1931), or additional analysis of specialized topics and subjects (e.g. 1911, 1931). comprehensive holdings of printed census of canada publications can be found at various government, academic, and public libraries in large part due to historical distribution and release of these materials via the government of canada’s depository services program (dsp), though following the termination of the dsp's distribution program and depository library agreements in 2014, there is no guarantee that access to these physical materials will be maintained in the future. in more recent years, significant efforts have been made through various library and government https://doi.org/10.29713/iq1167 https://laws.justice.gc.ca/eng/acts/s-19/fulltext.html https://library-archives.canada.ca/eng/collection/research-help/genealogy-family-history/censuses/pages/other-census-related-documents.aspx https://library-archives.canada.ca/eng/collection/research-help/genealogy-family-history/censuses/pages/other-census-related-documents.aspx 6/28 campbell, graeme, cuyler, katie & guindon, alex (2025). assessing the landscape for discovery and access to historical canadian census data, iassist quarterly 49(4), pp. 1-28. doi: https://doi.org/10.29713/iq1167 initiatives to further digitize and expand access to these publications allowing them to be more easily findable and reused. website/portal creator access coverage formats host type canadiana canadiana / crkn public 18th-19th centuries (selected) pdf, jpeg non-profit government of canada publications government of canada publications directorate public 1851-2021 pdf government hathitrust digital library hathitrust / various academic institutions public / organizational membership 1861-1941* pdf, jpeg, tiff, epub**, and plain text non-profit / academic internet archive internet archive / various institutions public 1870-2006 pdf, daisy, epub, and plain text non-profit statistics canada website statistics canada public 1996-2021 pdf government table 2: characteristics of census publications portals *the census of canada publications available through the hathitrust digital library are only those they have verified to be in the public domain. **epub downloads are available for subscribers only. canadiana canadiana is another digital collection hosted and made publicly available at no charge by crkn in a trustworthy digital repository. it contains nearly 20 million digitized pages of historical publications (crkn, 2024b). canadiana includes some published census volumes from the 19th century, although they are not conveniently grouped together for easy discovery and perusal. of particular interest are digitized versions of a few publications reporting data aggregated from pre-1851 censuses, such as a census of the population of the province of new brunswick in the year 1840, statistics of the population of the british colonies in north america for the year 1833, and the recensement de la ville de québec pour 1716. there is no advanced search interface for canadiana, but it provides keyword searching of document metadata and full text, and a variety of search operators can be used to refine search results to https://doi.org/10.29713/iq1167 https://www.canadiana.ca/ 7/28 campbell, graeme, cuyler, katie & guindon, alex (2025). assessing the landscape for discovery and access to historical canadian census data, iassist quarterly 49(4), pp. 1-28. doi: https://doi.org/10.29713/iq1167 locate census publications (e.g. searching by title for “census”, “recensement” or “population”). canadiana’s content is freely available to the public, and users can download individual pages in jpeg or pdf format, and entire documents as pdfs. government of canada publications the government of canada publications (gcp) catalogue contains a comprehensive collection of the most recent born-digital census publications issued by statistics canada, digitized copies of historical census of canada publications from confederation to 2006, and digitized census publications from former colonies prior to joining confederation back to the 1851 census of the canadas. these publications are acquired and hosted by the gcp directorate of public services and procurement canada, and the catalogue and metadata are also maintained by them. publications are available in pdf format and are downloadable without charge. there is a basic and an advanced search interface, but at present only publication metadata are indexed, not full text. census publications are not easily browsed in the gcp catalogue in a way that would ensure users have a comprehensive view of what is available within each year, so it is better to use the advanced search interface. hathitrust digital library hathitrust is a consortium of academic and research libraries that collaborates on services and programs including the hathitrust digital library (hdl), which rehosts digital publications contributed by member libraries in a trustworthy digital repository (crl, 2011). access to the many digitized publications in the hdl that fall within the public domain are freely available to the public. items deemed to be in copyright cannot be read or downloaded, but their full text can be searched in a variety of ways, including through the hdl’s discovery interface. as a result, while there are hundreds of census of canada publications preserved in the hdl, only a fraction are available to be downloaded and viewed, namely those from the pre-confederation years to the middle of the 20th century. there are basic and advanced search interfaces that index publication metadata and full text, and notably, the database attempts to collocate census publications by year, which is a useful feature. while anyone can view the full text of individual publications in the public domain online, only individual pages can be downloaded in pdf, jpeg, tiff or plain text formats. users at hathitrust member institutions are able to download complete items in those formats and as epubs, and also have access to other services, including text and data mining. internet archive the internet archive (ia) is an american nonprofit organization that hosts a significant collection of freely downloadable digitized census of canada publications from the pre-confederation years to 2006 on its website. these publications have been contributed through the digitization efforts of various institutions at various times and levels of comprehensiveness. the ia has a very active expanding catalogue of publications, but acquisitions to a large extent are driven by what contributing organizations want to digitize. publications can be downloaded without charge in a variety of formats, including pdf, plain text, daisy, and epub. like the gcp catalogue, there are basic and advanced search interfaces, and it can be a challenge to browse for the publications of specific census years. since the metadata for ia items comes from heterogeneous sources, systematically searching using its metadata fields is also difficult. however, an important advantage over the gcp search interface is that users can also search the full text of items in the ia catalogue. https://doi.org/10.29713/iq1167 https://publications.gc.ca/site/eng/home.html https://www.hathitrust.org/ https://archive.org/ 8/28 campbell, graeme, cuyler, katie & guindon, alex (2025). assessing the landscape for discovery and access to historical canadian census data, iassist quarterly 49(4), pp. 1-28. doi: https://doi.org/10.29713/iq1167 statistics canada website the statistics canada website is a primary hosting and distribution channel for statistics canadaproduced publications, aggregate data, public use microdata, and survey and data collection information produced for public consumption. as such, it is updated as new information becomes available and is the principal source for the most recently published census of canada information. census of canada products form only part of the information available on the statistics canada website, and while census information can be accessed through the site’s primary search and browse interfaces, a separate subsite is also provided for each census year. the subsite for the most recent census can be accessed through a link in the statistics canada website’s main menu and is divided into three sections: census of population, census of agriculture, and census engagement. archived subsites for previous census years back to 1996 are linked to from the main landing pages for the censuses of population and agriculture respectively. while the statistics canada website is not an access point for the digitized historical census publications, hundreds of analytical publications related to the census of population and census of agriculture from about the year 2000 onwards can be found there and are freely downloadable. the 2006 census year was also the last year for which a set of printed census publications was issued, after which the results of the census were published exclusively online as one of the aforementioned subsites. these subsites effectively collocate information within each census year, and from 2011 onwards, they are what is issued in place of sets of printed census of canada publications. while individual census tables and documents for each year may be rehosted elsewhere, the 2011, 2016, and 2021 census of canada subsites are not replicated elsewhere except where they have been preserved by various initiatives such as the government of canada web archive. in addition to analytical publications, the statistics canada website also provides access to census reference materials back to the 1996 census of canada. these materials include information concerning how the census is carried out, changes from the last census, census terminology, sampling and weighting procedures, educational kits, and guides to assist users with interpretation of published census products. geographic reference materials are also provided, including boundary files and reference maps for various levels of census geography, road network files, correspondence files for previous versions of certain census geographies, and documentation for these materials. reference materials for pre-1996 censuses can be found on some of the other portals described in this paper, such as the gcp and ia sites, but the scope of what is available varies. non-confidential census of canada information products distributed via the statistics canada website fall under the statistics canada open licence (statistics canada, 2024, july 26), which currently allows for their royalty-free use, reproduction, and redistribution, subject to certain user responsibilities. aggregate data this section deals with portals that present census data in aggregate form, essentially data tables. these tables are either “born-digital” or, for older censuses, have been extracted from print publications via optical character recognition (ocr) or other techniques. although most of the tables were originally published by statistics canada as part of the census releases, some sites (like https://doi.org/10.29713/iq1167 https://www.statcan.gc.ca/ https://webarchiveweb.bac-lac.canada.ca/ https://www.statcan.gc.ca/en/reference/licence 9/28 campbell, graeme, cuyler, katie & guindon, alex (2025). assessing the landscape for discovery and access to historical canadian census data, iassist quarterly 49(4), pp. 1-28. doi: https://doi.org/10.29713/iq1167 uni·cen) have reformatted the tables to enhance their standardization. many of the sites listed in this section seek to supplement the data offering from statistics canada by providing a single access point to data collected from multiple censuses. others, like the canadian census analyser, aim to provide a very sophisticated filtering and downloading system that allows users to produce finely customized tables. finally, some sites, like the canadian historical geographic information system and the community data program, provide interactive maps that allow geographic exploration of the data. website/portal creator access coverage formats host type abacus university of british columbia public 1665-2016 ivt, csv, excel academic canadian census analyzer at chass data centre, faculty of arts & science, university of toronto subscription 1961-2021 text, html, csv, excel, sas, and spss academic canadian century research infrastructure consortium of research groups from 8 universities public 1911-1951 excel academic canadian historical geographic information system (chgis) the canadian peoples (tcp) project / historical gis lab, university of saskatchewan public 1851-1921* csv academic community data program community data program subscription 2006-2021* ivt non-profit odesi scholars portal public 1665-2021 ivt, csv academic statistics canada website statistics canada public 1981-2021* csv, xml, and ivt government uni·cen uni·cen public 1951-2021 dta, csv, dbf academic table 3: characteristics of aggregate data portals *coverage is not comprehensive for all years. https://doi.org/10.29713/iq1167 10/28 campbell, graeme, cuyler, katie & guindon, alex (2025). assessing the landscape for discovery and access to historical canadian census data, iassist quarterly 49(4), pp. 1-28. doi: https://doi.org/10.29713/iq1167 abacus abacus is a data repository shared by several universities from british columbia (simon fraser university, university of northern british columbia, university of victoria) and hosted by the university of british columbia. it is based on an instance of the dataverse platform and holds public and restricted data from various organizations such as statistics canada, the inter-university consortium for political and social research (icpsr), dmti, and the linguistic data consortium. of relevance to this article, it includes aggregate data from censuses ranging from 1665 to 2016. the census data tables are publicly downloadable and available in beyond 20/20 (ivt), csv and sometimes excel (xlsx) formats. the collection is extensive and appears to include most data tables published by statistics canada as well as census-related reference publications. although users cannot strictly limit their search to the collection of census tables, there is a subcollection (a “dataverse”) for all statistics canada materials with an open license, which include aggregate and geospatial data plus reference publications. within this sub-collection, one can either do a keyword search or browse the contents using the provided “keyword term” filters (“2001 census of canada”, for instance). an “advanced search” menu also allows for field-level keyword searches and is based on the metadata describing the various datasets. canadian census analyser at chass the canadian census analyser, developed by the computing in the humanities and social sciences (chass) data centre at the university of toronto, is a bilingual tool that provides access to census profiles, i.e. statistical portraits of different census geographic units going from the enumeration area (ea) level to the provincial and national levels. coverage spans 1961 to 2021, including the 2011 national household survey. this aggregate data (from the profiles) is available in multiple formats including text, html, csv, excel, sas and spss. content is hosted locally and updated as new census data becomes available, but access is only available to users at subscribing universities. one of the most important features of the census analyzer is the ability to filter data by profile variable and by specific geographic units (for instance, one can indicate exactly what census tracts are required). the census analyser does not provide a search tool but instead relies on a sophisticated set of browsing and filter features that allow users to precisely determine what years, geographic units, and variables they want from the census profiles. canadian century research infrastructure (ccri) the ccri project was developed by a multi-institutional research collaboration between eight canadian universities that used the census returns to create microdata for the five decennial censuses covering the first half of the 20th century (1911-1951) (see the microdata section below for more details). the research team also created spatial boundary files at the census division and census subdivision levels for all five censuses. the ccri project was active between 2003 to 2009 and its website and datasets are hosted by the university of alberta. copies of the datasets are also held at the université du québec à trois-rivières (uqtr). ccri scanned selected tables from the published census volumes (for the 1911 to 1951 period) and then used ocr software to extract the data and produce excel files. ccri also conducted a series of validation tests on the extracted data and added some annotation to the tables to indicate https://doi.org/10.29713/iq1167 https://abacus.library.ubc.ca/ https://dc1.chass.utoronto.ca/census/index.html https://ccri.library.ualberta.ca/ 11/28 campbell, graeme, cuyler, katie & guindon, alex (2025). assessing the landscape for discovery and access to historical canadian census data, iassist quarterly 49(4), pp. 1-28. doi: https://doi.org/10.29713/iq1167 correction to typographical errors for instance. in total 23 tables were selected, covering topics such as population, religion, dwellings and origin. the tables present data at the census division and census subdivision levels. ccri also created pdf files that provide the full name of the variables in french and english. the files are not directly available on the website of the project, but are publicly downloadable from the ccri dataverse hosted by borealis along with geospatial data files created by the project. the aggregate data files are also findable through odesi (see below) and downloadable from the associated census of population dataverse in borealis4. canadian historical geographic information system (chgis) the data available on the chgis website, hosted by the historical gis lab at the university of saskatchewan, was produced by the canadian peoples (tcp) project, itself based at the university of guelph and composed of researchers from several canadian universities and research teams. chgis offers selected aggregate census data tables from 1851 to 1921. it also provides geospatial data, which is described in the census maps and geospatial data section below. note that the data presented for the 1851 and 1861 pre-confederation censuses only covers lower and upper canada (currently the provinces of quebec and ontario). the data tables were obtained by performing ocr on digitized versions of the printed census volumes available from gcp and ia. the digitized tables are available in excel format and can be freely downloaded from the website. users can download a zipped file containing all tables for a given census year (the total number of tables varies from one to six depending on the census year). the website also allows users to view the data for a specific census subdivision5. this can be done in three different ways: 1) through a browsing feature that allows them to select the subdivision from a list that can be filtered by province and by census year; 2) through a simple search interface where they can enter the name of the subdivision and select the year and province; and 3) by using an interactive map that shows the boundaries of each census subdivision for every census. the available documentation includes a detailed description of the methodology used to digitize the files, explains what corrections to the original data were performed, and describes the various table formats that were created. a second document presents the list of all variables available for each table. community data program the community data program (cdp) is a member-funded, community non-profit organization that provides access to publications and data to facilitate community development initiatives by its members, which include community non-profits, educational institutions, and local governments. the data provided is intended to inform program and policy planning and implementation. census data is one of the types of data made available for this purpose. the cdp rehosts and makes available reference publications and aggregate data from the 2006-2021 censuses. access to publications and data from the censuses is incomplete and has been curated by cdp to serve the needs of their members. the data purchase and access working group (dpawg), which is made up of representatives from member organizations, meets regularly to identify data needs and to acquire data products for the program. data within the portal can be searched by topic, data source (census year and theme), data provider, title, geography, or year, but while the search interface is public, access to complete items and downloading of ivt files are restricted to members. https://doi.org/10.29713/iq1167 https://borealisdata.ca/dataverse/ccri https://borealisdata.ca/dataverse/census https://hgiscanada.usask.ca/ https://thecanadianpeoples.com/ https://communitydata.ca/ 12/28 campbell, graeme, cuyler, katie & guindon, alex (2025). assessing the landscape for discovery and access to historical canadian census data, iassist quarterly 49(4), pp. 1-28. doi: https://doi.org/10.29713/iq1167 odesi odesi is a search engine and discovery portal for statistical data developed and funded by the ontario council of university libraries (ocul) and maintained by scholars portal. content is hosted on the borealis (dataverse) platform, also maintained by scholars portal, and includes datasets from various organizations such as canadian opinion polling firms and icpsr. the database is continuously updated as new data tables and microdata files become available6. more importantly from our perspective, it includes an extensive number of aggregate data tables (and microdata files, as described below) from the canadian census of population, going back to the pre-confederation censuses, covering the period from 1665 to 2021. aggregate data is available in beyond 20/20 format (ivt) and occasionally in csv and tab-delimited formats. odesi provides a sophisticated search interface that allows searching files by title, author, variable, series, keyword and abstract. the search can be limited to specific collections (like the census) and filtered by year. the datasets are coded using the data description initiative (ddi) metadata schema. in addition to the search features, users can also choose to browse the different collections, including the census. browsing the census collection will present the data per census year and broken down by census topics (for the more recent censuses). both the search tool and the census data itself are open to the general public. statistics canada website as described above, statistics canada provides free public access to aggregate data products hosted and published on its website and is updated with each cycle of the census of canada. the years covered vary by access point. as already mentioned, there are census of population subsites available from 1996 onwards, with all but the most recent censuses marked as archived content. through the website’s main data search interface and the census datasets search interface, content is available back to the 1981 census, though most is from 1991 onwards. other census-related tools on the site have their own unique time period coverages. the newer census profile web data service and census program data viewer mediate access to content from the 2016 census forward. the main data interface can be searched by keyword and filtered by subject or level of geography. there is also a convenient filter specifically limiting all results just to the census of population, but one could also limit results similarly using the “survey or statistical program” filters. of the approximately 3000 census data products findable through this method, almost all are aggregated products, with most being data tables. there are also community or regional profiles, thematic maps, and data visualization products. the census datasets search interface is available through the main census of population subsite, which also includes a search interface for census profiles. it is possible to search or browse these two types of content back to 1996 as well. using the statistics canada website’s site search, the data search interface, or one of the censusspecific interfaces is mainly a matter of personal preference, as one should be able to find the same census of population aggregate data tables through a variety of methods. these tables may be somewhat fixed or may allow users to customize the view of the data by selecting the variables and/or variable values to display. there are often a variety of download options, including the extent to which the table and associated metadata are included, and formats like comma or tab-separated values, xml, or beyond 20/20. https://doi.org/10.29713/iq1167 https://odesi.ca/en https://scholarsportal.info/ https://borealisdata.ca/dataverse/census 13/28 campbell, graeme, cuyler, katie & guindon, alex (2025). assessing the landscape for discovery and access to historical canadian census data, iassist quarterly 49(4), pp. 1-28. doi: https://doi.org/10.29713/iq1167 unified infrastructure for canadian census research (uni·cen) uni·cen, a project from western university, includes publicly released univariate aggregate data tables at seven levels of geography: canada, provinces and territories, census metropolitan areas, agglomerations, divisions, subdivisions, and tracts7. currently, the aggregate data tables are available for the 1951–2021 censuses, but the project aims to add census data for the 1851 to 1951 period soon, adapting census division and subdivision data digitized and disseminated by the ccri and the canadian peoples project (tcp). the census boundary files are available for all censuses from 1851 to 2021. the original source files have been modified to standardize attribute tables and metadata, including the handling of geographic identifiers. uni·cen is currently working on producing additional resources such as a comprehensive database of census questions covering the 1851 to 2021 period, geographic concordance tables to link equivalent geographic units across censuses, and geographic and postal code lookup tables. access to uni·cen data is currently available in two places, with plans for additional access points. first, the aggregate data tables and boundary files can be freely downloaded from the uni·cen dataverse in the borealis repository. second, the uni·cen canadian neighbourhood change explorer 1951–2021, developed in partnership with esri canada, inc., and launched in march 2024, is a freelyavailable visualization tool allowing users to explore and compare, through a map interface, selected longitudinally available census data harmonized to 2021 census tract boundaries. future plans include access via the scholars geoportal platform and through a web portal that is still in development. microdata microdata is data observed or collected directly from a specific unit of observation. microdata from statistics canada usually comes in one of the two file types, public use microdata files (pumf) and master files (statistics canada, 2022). pumfs, as their name indicates, are meant to be accessible to the public. in those files, different techniques have been applied to maintain the respondents’ confidentiality. for instance, some of the variables are suppressed or categories are collapsed and, importantly, data is not generally available at fine geographic levels. master files, on the other hand, contain the complete data (all answers from all respondents) as collected from the census but are not publicly accessible. a network of research data centres (rdcs) has been created to allow vetted researchers to analyse these datasets in a secure environment. https://doi.org/10.29713/iq1167 https://observatory.uwo.ca/unicen/index.html https://borealisdata.ca/dataverse/unicen https://borealisdata.ca/dataverse/unicen https://edumaps.esri.ca/census/ https://edumaps.esri.ca/census/ 14/28 campbell, graeme, cuyler, katie & guindon, alex (2025). assessing the landscape for discovery and access to historical canadian census data, iassist quarterly 49(4), pp. 1-28. doi: https://doi.org/10.29713/iq1167 website/portal creator access coverage formats host type abacus university of british columbia public 1971-2021 spss, stata, sas, csv, tab academic canadian census analyzer at chass data centre, faculty of arts & science, university of toronto. subscription 1971-2016 csv and text academic ccri consortium of research groups from 8 universities. public 1911 & 1921* spss academic impq research teams from uqtr, uqac u de m. portal hosted by the cieq. public (account required) 1852-1911 csv academic odesi scholars’ portal public 1871 to 1911**, 1971-2021 r, spss, stata, sas, csv, tab academic prdh université de montréal public 1831 (qc) 1852 & 1881 spss & stata academic research data centres research data centres mediated / subscription 1911-2021 sas, spss, stata, csv and text government / academic statistics canada website statistics canada public 1991-2021 csv and text government table 4: characteristics of microdata portals *additional microdata created by ccri from 1921-1951 is available mediated through rdcs. ** these are population samples digitized by several research groups such as ccri. https://doi.org/10.29713/iq1167 15/28 campbell, graeme, cuyler, katie & guindon, alex (2025). assessing the landscape for discovery and access to historical canadian census data, iassist quarterly 49(4), pp. 1-28. doi: https://doi.org/10.29713/iq1167 abacus abacus (described in the aggregate data section) includes the various microdata files (pumfs) created by statistics canada for the censuses from 1971 to 2021. they can be downloaded as tab and csv files and, in most cases, syntax files for spss, sas and stata are available. the associated documentation (cobebooks, user guides, etc.) is also available. all files are freely downloadable. canadian census analyser at chass the canadian census analyser (mentioned above in the aggregate data section) also provides access to pumfs from 1971 to 2016. users can also download subsets of the standard pumfs via a sophisticated filtering system (based on the sda set of programs) allowing for variable and case selection via a browsing and filtering interface. data can be downloaded as text or csv files, with accompanying codebooks and syntax files for sas, spss, stata, ddi and sda. content is hosted locally by the service and updated as new census data becomes available, but access is only available to users at subscribing universities. canadian century research infrastructure (ccri) for a general description of the ccri project, see the aggregate data section. the microdata consists of samples of the population (5% for the 1911 census, 4% for the 1921 census and 3% for the 1931, 1941 and 1951 censuses). the 1911 and 1921 data can be freely downloaded directly from the ccri website in spss format and is also hosted in the ccri dataverse hosted in borealis and discoverable through odesi. data access for all five census years is available for researchers with approved projects via the statistics canada rdc network (see below). the spatial data files for all the census years covered by the project (boundary files for the various census geographic units) can be downloaded from scholars geoportal. the centre interuniversitaire d’études québécoises (cieq) at uqtr’s (one the ccri’s collaborating universities) échantillon de la population canadienne en 1911 also provides a tool with a sophisticated faceted search engine to explore the 1911 census microdata and the complete collection of geospatial data files created by the project. infrastructure intégrée des microdonnées historiques de la population du québec (impq) the impq project links data from the province of quebec civil records – specifically, records of marriage from 1621 to 1914; records of births and deaths between 1621 and 1849; and births and deaths recorded in the regions of saguenay and lac saint-jean before 1914 – with the complete (not sampled) microdata from seven canadian censuses held between 1852 and 1911. the portal features a search engine that allows faceted search based on the various census fields (last name, first name, age, profession, religion, etc.) and many criteria can be combined. geographic coverage is limited to québec city, trois‐rivières, and the saguenay, lac-saint-jean, côte‐nord, and gaspésie regions of the province of québec. there are two tiers of access to impq data, both of which require registration through the creation of an online account. registration for the public portal is available to anyone and allows users to search for and download data in excel format from the impq’s census microdata. users can also choose to view the census records individually or grouped by households. digitized versions (pdf files) of the census returns are also available. access to the impq’s quebec civil records data and to the links between them and the https://doi.org/10.29713/iq1167 https://dc1.chass.utoronto.ca/census/index.html https://sda.berkeley.edu/ https://borealisdata.ca/dataverse/ccri https://www.statcan.gc.ca/en/microdata/data-centres https://geo.scholarsportal.info/ https://ircs1911.cieq.ca/ https://ircs1911.cieq.ca/ https://impq.cieq.ca/ 16/28 campbell, graeme, cuyler, katie & guindon, alex (2025). assessing the landscape for discovery and access to historical canadian census data, iassist quarterly 49(4), pp. 1-28. doi: https://doi.org/10.29713/iq1167 census are only available to registered researchers, who must confirm their institutional affiliation and show that their use of the data will be for research purposes. impq is a partnership between research groups at the université du québec à chicoutimi, université de montréal and the uqtr. the research portal and data are hosted by uqtr. while the three collaborating research groups are still active, there is no mention of recent additions to the database or changes to the research portal. odesi odesi (mentioned in the aggregate data section above) also allows users to explore and download pumfs for the census of population from 1971 to 2021. the microdata files created by the ccri project for the 1911 census (see above) are also discoverable in odesi. also available are microdata datasets for samples of the population created by various research teams8 for the 1871, 1881, 1891 and 1901 censuses. the data files themselves are in the census dataverse hosted on borealis. users can download complete microdata files or extract specific variables. odesi has sophisticated search and browse functions described above, and microdata files can be downloaded in several formats, including r, spss, stata, sas, csv and tab-delimited. odesi is continuously updated as new content becomes available. programme de recherche en démographie historique (prdh) the prdh, based at université de montréal, created sets of microdata for the 1831 census of quebec, and the 1852 and 1881 censuses of canada based on the microfilmed copies of the census returns. the dataset for the 1831 census of quebec includes 100% of the population. for 1852, the dataset is a random sample representing 20% of the population in the surviving census returns (27% of the manuscripts returns were lost before they could be microfilmed). the 1881 dataset represents 100% of the population. in addition to the data files, some textual documentation is included. the content is entirely hosted on the prdh website, and while access is free and public, users must first register to create an online account. the microdata for all three (1831, 1852, and 1881) censuses can be downloaded from the prdh website in spss format, with syntax files available for the 1831 and 1881 censuses. the microdata for the 1831 census of quebec is additionally downloadable in stata format. the 1852 and 1881 databases can also be searched online via one of two data browsers that allow users to search for, view, and download pdf versions of the digitized microfilmed returns. each browser allows users to search based on characteristics of individuals, families, birthplace, religion, etc. the prdh is still active but appears to now be focused on the study of quebec demography. research data centres (rdcs) statistics canada rdcs are physically located within academic or government institutions and are jointly managed by statistics canada and host organizations. rdcs provide secure, non-public access to statistics canada microdata master files. data users are required to attain a security clearance, complete mandatory training, and swear or affirm the oath of office and secrecy to statistics https://doi.org/10.29713/iq1167 https://odesi.ca/en https://borealisdata.ca/dataverse/census https://www.prdh.umontreal.ca/census/en/main.aspx 17/28 campbell, graeme, cuyler, katie & guindon, alex (2025). assessing the landscape for discovery and access to historical canadian census data, iassist quarterly 49(4), pp. 1-28. doi: https://doi.org/10.29713/iq1167 canada. work with data must occur on site, and only vetted analysis, not raw data, can be taken out of the centre or published. holdings include a variety of survey and administrative data, including the census of population for every ten years from 1911-1951 and every five years from 1971-2021. note that only a 20% sample, and not the master file, is available for the 2011 census of population, but the master file is available for the 2011 national household survey, which replaced the long-form census that year. notably unique about rdc access is the availability of linked sets of microdata, including those linking censuses of population with a number of other statistics canada microdata sources. the number of datasets available to researchers grows as new products become available. a complete listing of data available in the rdcs can be found on the statistics canada website. statistics canada website pumfs for the census of population and the 2011 national household survey produced by statistics canada are hosted and made freely downloadable from the statistics canada website from the most recent census back to 1991. there is no search interface, but all available files are hyperlinked from a single index page. “individuals” files are available for 1991-present, “hierarchical” files from 2006 forward, and “families” and “households and housing” files from 1991-2001. each downloadable zip file includes the data itself in csv or text format, documentation, and license agreements. the use of statistics canada pumfs is governed by the statistics canada open licence (statistics canada, 2024, november 8). census maps and digital spatial data this section presents the main sites providing access to maps or geospatial data files related to the canadian census. in addition to thematic maps presenting visual representations of selected census variables or topics, statistics canada produces a series of reference maps that outline the boundaries of geographic units used to collect and organize the census data. the term “digital spatial data” refers to machine readable files that provide users with the ability to visualize and analyse data through the use of geographic information systems (gis) software. although these files were produced by statistics canada for recent censuses, the equivalent files for older censuses were created by various research groups or nonprofit organizations by scanning and georeferencing the original paper maps. https://doi.org/10.29713/iq1167 https://www.statcan.gc.ca/en/microdata/data-centres/data https://www.statcan.gc.ca/en/microdata/data-centres/data https://www150.statcan.gc.ca/n1/pub/98m0001x/index-eng.htm https://www.statcan.gc.ca/en/reference/licence 18/28 campbell, graeme, cuyler, katie & guindon, alex (2025). assessing the landscape for discovery and access to historical canadian census data, iassist quarterly 49(4), pp. 1-28. doi: https://doi.org/10.29713/iq1167 website/portal creator access coverage formats host type abacus university of british columbia public 1971-2016 shp, mapinfo, geojson, e00 academic canadian historical geographic information system (chgis) the canadian peoples (tcp) project / historical gis lab, university of saskatchewan public 1851-1921 gdb academic scholars geoportal scholars portal, ontario council of university libraries (ocul) public 1901-2021* shp academic statistics canada website statistics canada public 2001-2021 shp, gml, gdb, tab, pdf government uni·cen’s dataverse uni·cen public 1851-2021** gdb, geojson, shp academic table 5: characteristics of census maps and digital spatial content portals *1916, 1926 and 1946 (prairie provinces), 1945 (nl) and 1956 to 1976 (canada) are missing. **1956 to 1981 are missing abacus abacus (described in more details in the aggregate data section) has a collection of census geospatial files originally created by statistics canada that covers the period from 1971 to 2016. it includes street networks files as well as cartographic boundary files9 (cbfs) and digital boundary files10 (dbfs). the files are available in multiple formats including shp, mapinfo, e00, and geojson. associated documentation is also provided. all files are publicly downloadable. canadian historical geographic information system (chgis) this section only reviews the geospatial data made available by the chgis. for a more complete description of the project, see the entry for cghis in the aggregate data section of the article above. under the umbrella of tcp project, researchers from chgis at the university of saskatchewan along with the centre interuniversitaire d'études québécoises (cieq) at université laval (www.cieq.ca) created a series of boundary files at the census subdivision level for every census year covered by the project (from 1851 to 1921 inclusively). the data consists of one large geodatabase (gdb) that includes layers for each census year. the file is available to the public under a creative commons https://doi.org/10.29713/iq1167 http://www.cieq.ca/ 19/28 campbell, graeme, cuyler, katie & guindon, alex (2025). assessing the landscape for discovery and access to historical canadian census data, iassist quarterly 49(4), pp. 1-28. doi: https://doi.org/10.29713/iq1167 open data license and can be downloaded from the cghis website. unique identifiers were created for each census subdivision allowing the boundary files to be joined to the data tables created by tcp (see aggregate data section for more information on those tables). the available documentation describes the methodology used to create the boundary files and the unique identifiers. it also includes complete metadata and description of the attribute tables. finally, the document provides a detailed description and methodological notes for each series of boundary files (each census year). scholars geoportal scholars geoportal is a project supported the ocul, which provides a large collection of geospatial products (vector data, orthophotographs, aerial images, etc.) from various levels of government (municipal, provincial, federal) and from the collections of ontario university libraries. the portal is probably the most complete source of historical canadian census geospatial files as it includes an extensive collection of census boundary files (in addition to other spatial datasets such as road network files and geographic attribute files) covering a period from 1851 to 202111, although there are still a few gaps in the coverage and not all geographic levels are present for all census years. the files were originally created by several entities such as statistics canada, the ccri project (for the 1911 to 1951 censuses)12, the historical atlas of canada online learning project and the university of toronto map and data library (allen & leahey, 2018). in 2015, a project led by scholars portal and the university of toronto map and data library managed to gather the very dispersed and heterogeneous collection of boundary files, convert them to a current file format (shapefile) and create associated metadata. for recent years (1991 onward), the portal provides both the cbfs and the dbfs created by statistics canada. the scholars geoportal website provides a sophisticated search interface that allows users to look for datasets by keyword, place name or subject category (including “census and administrative boundaries”), or by drawing a rectangular area of interest on the interactive map. search criteria can also be combined. once a dataset has been selected, it can be visualized as a map or downloaded in shapefile format. all census datasets are publicly available. scholars geoportal is an ongoing project, and new datasets are being added as they become available. statistics canada website the census geography section of statistics canada’s website is the main resource for geographic reference products (illustrated glossary, reference guides, catalogue, working papers, etc.), thematic and reference maps, attribute information products (correspondence file and tools to explore the census geography) and geospatial products including boundary files and road network files. this collection is limited to recent censuses, from 2001 onward. the geosuite product can be used to navigate between the different geographic units and obtain lists of sub-units (along with basic data) corresponding to a larger geographic unit. the maps subsection of the website includes both reference maps that display some of the census geographic units and thematic maps that provide a cartographic view of certain census variables. the reference maps are mostly available at the census tract, federal electoral district and dissemination area levels in pdf format. also included in this section is an interactive mapping https://doi.org/10.29713/iq1167 https://hgiscanada.usask.ca/download https://geo1.scholarsportal.info/ https://www12.statcan.gc.ca/census-recensement/2021/geo/index-eng.cfm https://geosuite.statcan.gc.ca/geosuite/en/index 20/28 campbell, graeme, cuyler, katie & guindon, alex (2025). assessing the landscape for discovery and access to historical canadian census data, iassist quarterly 49(4), pp. 1-28. doi: https://doi.org/10.29713/iq1167 application called geosearch that allows users to locate reports and datasets corresponding to the various census geographic units. the boundary files include cbfs and dbfs as well as intercensal census subdivision boundary files (covering certain years between censuses) and are publicly available for download in various formats such as shapefile (shp), geography markup language (gml), file geodatabase (gdb), and mapinfo (tab). in 2021, statistics canada also added support for api-based mapping services (esri rest and web mapping service (wms)). there is no search engine, and access to the various documents is simply done by navigating the links and making selections (geographic unit, file format, year, etc.). all files are hosted locally and are publicly accessible. unified infrastructure for canadian census research (uni·cen) the uni-cen research group (described in more detail in the aggregate data section above) produced an extensive series of harmonized census boundary files to facilitate their use across time. according to uni-cen documentation, the team created a reformatted version of all publicly available census boundary files – originating mainly from statistics canada, tcp, ccri, and scholars geoportal. the harmonized files have standardised attribute table fields, consistent shorelines, are projected to the same coordinate system and are available in several modern file formats, namely shp, gdb and geojson. the files go back to very early censuses (1851) and are available at several geographic levels including, for the large majority of census years, census subdivision, census division and province. in addition to the census boundary files, uni-cen produced a comprehensive series of harmonized spatial files for the federal electoral districts (fed) going back to 1867. these can be used to map census data as statistics canada started to release profiles for feds in 1961. moreover, for the period covering the 1991 to 1951 censuses, uni-cen compiled fed data from donald blake’s profiles created for his phd thesis and profiles created by the ccri. the geospatial files are not directly available on uni-cen’s website, but are rather hosted on the project’s dataverse in borealis. from there, users can either do a generic keyword search or use the advanced search screen to interrogate the various metadata fields. from the main screen, they can also explore the data by using several types of filters such as time period, geographic unit, data type and keyword term. all files are publicly downloadable, and detailed documentation is also available in the dataverse. overall summary although we do not claim that this inventory of portals giving access to canadian census content is comprehensive – we have excluded genealogy sites for instance – it gives a good picture of what is available to researchers and the larger public. the two following tables present a breakdown by type of access and host type. note that some portals (abacus, canadian census analyser, etc.) appear in more than one content-type category and thus are counted more than once. https://doi.org/10.29713/iq1167 https://www.icpsr.umich.edu/web/icpsr/studies/39/versions/v2 https://www.icpsr.umich.edu/web/icpsr/studies/39/versions/v2 https://borealisdata.ca/dataverse/unicen_boundaries 21/28 campbell, graeme, cuyler, katie & guindon, alex (2025). assessing the landscape for discovery and access to historical canadian census data, iassist quarterly 49(4), pp. 1-28. doi: https://doi.org/10.29713/iq1167 portals by content type access type public subscription* mixed census returns 2 2 0 0 publications 5 4 0 1 aggregate data 8 6 2 0 microdata 8 6 2 0 maps and spatial data 5 5 0 0 total 28 23 4 1 percentage 100% 82% 14% 4% table 6: portals by access type *includes mediated access and organization membership of all the analysed websites, the vast majority (82%) provide public access to its content. the few subscription-based portals – the cdp for aggregate data, the rdcs for microdata and the census analyser that has both aggregate and microdata – offer significant added value in terms of data packaging or user assistance and thus must charge to cover the cost of those services. on the other hand, all the portals that give access to census returns and maps & digital spatial data (and all but one that offer census publications) make their content publicly and freely available. overall, we can say that, although discovering and using census data may be difficult given the fragmented landscape discussed in this paper, at least most of the products are freely available. https://doi.org/10.29713/iq1167 22/28 campbell, graeme, cuyler, katie & guindon, alex (2025). assessing the landscape for discovery and access to historical canadian census data, iassist quarterly 49(4), pp. 1-28. doi: https://doi.org/10.29713/iq1167 portals by content type host type academic government non-profit mixed census returns 2 0 1 1 0 publications 5 0 2 2 1 aggregate data 8 6 1 1 microdata 8 7 0 0 1 maps and spatial data 5 4 1 0 0 total 28 17 5 4 2 percentage 100% 61% 18% 14% 7% table 7: portals by host type this breakdown by host type shows the importance of academic initiatives, spearheaded by research groups or academic units, in providing and enhancing access to census data, and even creating derived products. of the 28 portals examined, 17 (or 61%) are the outcome of university-led initiatives. these projects offer a necessary complement to the government-based portals (18%), by providing researchers with a large selection of tools or products that facilitate the exploration of census data by date or geography. it is also interesting to note that non-profit organizations – namely internet archive, canadiana and hathitrust (which is categorized under mixed as it is also academic) – play an important role in giving access to census publications, especially those from historical censuses. conclusion in response to the fragmented context of access to census materials described in this paper, the authors joined other collaborators from canadian universities, statistics canada and lac, to form the canadian census data discovery partnership (ccddp). the ccddp’s main goal is to provide a https://doi.org/10.29713/iq1167 23/28 campbell, graeme, cuyler, katie & guindon, alex (2025). assessing the landscape for discovery and access to historical canadian census data, iassist quarterly 49(4), pp. 1-28. doi: https://doi.org/10.29713/iq1167 single point of access to all census materials at the item level (individual tables or documents) going back to the very first pre-confederation censuses. the first phase of the project – which is nearly completed at the time of writing – consisted of the creation of a comprehensive, bilingual (french and english), and item-level inventory of all census publications. the next objective was to provide a searchable database allowing researchers to identify tables or publications by subject or geography across time. a searchable prototype database, tentatively called the census discovery portal, is currently under development. the landscape of census information access is complex. while much of the current and historical information is available online in digital formats, it is held in numerous online locations, which vary in the type of information they host, the time period covered, the way they describe and provide access to the content, the type of organization supporting them, their attitude towards preservation, and their general user base. navigating this often-confusing network of resources was an important motivation for the ccddp to inventory the content hosted at multiple online locations and provide a single point of access using consistent metadata. looking to the future, one of the ccddp’s greatest challenges will be how to approach sustainability, that of the growth of the inventory as new canadian censuses are completed, that of the maintenance of the census discovery portal, and that of the metadata, which points to multiple online resources that each have their own site information architecture that may change over time. while identifying an expansive number of census data portals, which provide invaluable access to census data, this paper highlights gaps and areas for improvement. ideally researchers would have access to the entire corpus of census data in a variety of formats, consistent across census years. however, with this ideal not in existence, works like this paper, and the inventory created by the ccddp, seek to improve discoverability of the currently available data. specifically, we hope that this paper presents a well-organized map of the canadian census landscape, allowing librarians and researchers to locate the best sources of information for their specific needs whether those are defined by type of data, historic period, or format type. also, by clarifying what is available, and making it more discoverable, we think that gaps in content and data formats will be more readily apparent, and thus more data can be created and made available to fill these gaps. reference list allen, j., & leahey, a. (2016, spring/summer). improving access to digital historical census boundaries in canada. acmla bulletin, 153, 44-50. https://doi.org/10.31235/osf.io/eacs5 canada. dominion bureau of statistics. (1959). current publications: dominion bureau of statistics, 1959. https://perma.cc/nc7q-agpu cooper, a. & jadon, v. (2023, september). spotlight: odesi markit! program: a collaborative curation model for data. scholars portal newsletter. https://perma.cc/p9r8-4kxx crkn. (2024a). about crkn. https://perma.cc/n745-63sn https://doi.org/10.29713/iq1167 https://doi.org/10.31235/osf.io/eacs5 https://perma.cc/nc7q-agpu https://perma.cc/p9r8-4kxx https://perma.cc/n745-63sn 24/28 campbell, graeme, cuyler, katie & guindon, alex (2025). assessing the landscape for discovery and access to historical canadian census data, iassist quarterly 49(4), pp. 1-28. doi: https://doi.org/10.29713/iq1167 crkn. (2024b). canadiana collections. https://perma.cc/63bl-8x78 crl. (2011). hathitrust audit report 2011. https://perma.cc/3vpc-npz9 hillman, thomas a. (1981). catalogue of census returns on microfilm 1666-1881. public archives canada. statistics canada. (2015, december 30). history of the census of canada. https://perma.cc/jh4p6hq4 statistics canada. (2019, august 7). 100 years and counting: more than a century of censuses in canada. https://perma.cc/8x2d-ntnf statistics canada. (2022, november 30). dli survival guide. https://perma.cc/8s4a-fajx statistics canada. (2024, july 26). statistics canada open licence. https://perma.cc/5nf8-6kuh statistics canada. (2024, november 8). census of population: public use microdata files. https://perma.cc/wnx7-l6jt statistics canada. (2024, november 26). 100 years of the canadian census. https://perma.cc/vj749mxj https://doi.org/10.29713/iq1167 https://perma.cc/63bl-8x78 https://perma.cc/3vpc-npz9 https://perma.cc/jh4p-6hq4 https://perma.cc/jh4p-6hq4 https://perma.cc/8x2d-ntnf https://perma.cc/8s4a-fajx https://perma.cc/5nf8-6kuh https://perma.cc/wnx7-l6jt https://perma.cc/vj74-9mxj https://perma.cc/vj74-9mxj 25/28 campbell, graeme, cuyler, katie & guindon, alex (2025). assessing the landscape for discovery and access to historical canadian census data, iassist quarterly 49(4), pp. 1-28. doi: https://doi.org/10.29713/iq1167 appendix a: alphabetical list of census portals website / portal creator content type access coverage formats host type abacus university of british columbia a) aggregate data b) microdata c) maps and digital spatial public a) 16652016 b) 19712021 c) 19712016 a) ivt, csv, excel b) spss, stata, sas, csv, tab c) shp, mapinfo, geojson, e00 academic canadiana canadiana / crkn publications public 18th-19th centuries (selected) pdf, jpeg non-profit canadian census analyzer at chass data centre, faculty of arts & science, university of toronto. a) aggregate data b) microdata subscription a) 19612021 b) 19712016 a) text, html, csv, excel, sas, and spss b) csv and txt academic canadian century research infrastructure consortium of research groups from 8 universities. a) aggregate data b) microdata public a) 19111951 b) 1911 & 1921* a) excel b) spss academic canadian historical geographic information system (chgis) the canadian peoples (tcp) project / historical gis lab, university of saskatchewan a) aggregate data b) maps and digital spatial public a) 18511921 b) 18511921 a) csv b) gdb academic census search library and archives canada returns public 1825-1931 pdf, jpeg, csv**, and xml** government community data program community data program aggregate data subscription 20062021*** ivt non-profit government of canada publications government of canada publications directorate publications public 1851-2021 pdf government hathitrust digital library hathitrust / various academic institutions publications public / organizational membership 18611941**** pdf, jpeg, tiff, epub*****, and plain text non-profit / academic https://doi.org/10.29713/iq1167 26/28 campbell, graeme, cuyler, katie & guindon, alex (2025). assessing the landscape for discovery and access to historical canadian census data, iassist quarterly 49(4), pp. 1-28. doi: https://doi.org/10.29713/iq1167 website / portal creator content type access coverage formats host type héritage canadiana / crkn returns public 17th-19th centuries (selected) pdf, jpeg non-profit impq research teams from uqtr, uqac u de m. portal hosted by the cieq. microdata public (account required) 1852-1911 csv academic internet archive internet archive / various institutions publications public 1870-2006 pdf, daisy, epub, and plain text non-profit odesi scholars portal a) aggregate data b) microdata public a) 16652021 b)18711911******, 1971-2021 a) ivt, csv b) r, spss, stata, sas, csv, tab academic prdh université de montréal microdata public 1831 (qc) 1852 & 1881 spss & stata academic research data centres research data centres microdata mediated / subscription 1911-2021 sas, spss, stata, csv and text government / academic statistics canada website statistics canada a) publications b) aggregate data c) microdata d) maps and digital spatial public a) 19962021 b) 19812021*** c) 19912021 d) 20012021 a) pdf b) csv, xml, and ivt c) csv and txt d) shp, gml, gdb, tab, pdf government scholars geoportal scholars portal, ontario council of university libraries (ocul) maps and digital spatial public 1901-2021*** shp academic uni·cen uni·cen a) aggregate data b) maps and digital spatial public a) 19512021 b) 18512021*** a) dta, csv, dbf b) gdb, geojason,shp academic https://doi.org/10.29713/iq1167 27/28 campbell, graeme, cuyler, katie & guindon, alex (2025). assessing the landscape for discovery and access to historical canadian census data, iassist quarterly 49(4), pp. 1-28. doi: https://doi.org/10.29713/iq1167 *additional microdata created by ccri from 1921-1951 is available mediated through rdcs. **search results can be downloaded in different formats ***coverage is not comprehensive for all years ****the census of canada publications available through the hathitrust digital library are only those they have verified to be in the public domain. *****epub downloads are available for subscribers only. ******these are population samples digitized by several research groups such as ccri. endnotes 1 graeme campbell is the open government librarian at queen’s university library. he can be reached by email: graeme.campbell@queensu.ca. 2 katie cuyler is the open publishing & government information librarian at the university of alberta library. she can be reached by email: katie.cuyler@ualberta.ca. 3 alex guindon is the gis and data services librarian at concordia university library. he can be reached by email: alex.guindon@concordia.ca 4 borealis (https://borealisdata.ca/) is a research data repository based on the dataverse platform, supported by canadian universities and hosted by scholars portal and the university of toronto libraries. 5 note that the statistics are presented as text on a web page and cannot be downloaded in a spreadsheet format. 6 the update process is under the responsibility of the odesi markit program!, managed by a group of ontario-based universities members of ocul. for more information, see cooper and jadon (2023) 7 these are various geographic units used by statistics canada, for a detailed description see this illustrated glossary (https://www150.statcan.gc.ca/n1/pub/92-195-x/92-195-x2021001-eng.htm). 8 the research teams include the canadian historical mobility project from york university (1871 census), the 1891 canadian census project guelph university and the canadian families project from the university of victoria. 9 cartographic boundary files (cbfs) portray the boundaries of standard geographic areas, together with the shoreline around canada. selected inland lakes and rivers are available as supplementary layers. 10 digital boundary files (dbfs) depict the full extent of the boundaries of standard geographic areas established for the purpose of disseminating census data, including the coastal water area. 11 this information is based on the geoportal census boundaries inventory (https://tinyurl.com/ya2prxvt) a non-published working document produced by the managers of the geoportal, last updated on august 2, 2023. https://doi.org/10.29713/iq1167 mailto:graeme.campbell@queensu.ca mailto:katie.cuyler@ualberta.ca mailto:alex.guindon@concordia.ca https://borealisdata.ca/ https://www150.statcan.gc.ca/n1/pub/92-195-x/92-195-x2021001-eng.htm https://tinyurl.com/ya2prxvt 28/28 campbell, graeme, cuyler, katie & guindon, alex (2025). assessing the landscape for discovery and access to historical canadian census data, iassist quarterly 49(4), pp. 1-28. doi: https://doi.org/10.29713/iq1167 12 the ccri geospatial files are also available from the project’s dataverse (https://borealisdata.ca/dataverse/ccri) on borealis. https://doi.org/10.29713/iq1167 https://borealisdata.ca/dataverse/ccri abstract keywords introduction background scope organization census returns héritage library and archives canada census publications canadiana government of canada publications hathitrust digital library internet archive statistics canada website aggregate data abacus canadian census analyser at chass canadian century research infrastructure (ccri) canadian historical geographic information system (chgis) community data program odesi statistics canada website unified infrastructure for canadian census research (uni cen) microdata abacus canadian census analyser at chass canadian century research infrastructure (ccri) infrastructure intégrée des microdonnées historiques de la population du québec (impq) odesi programme de recherche en démographie historique (prdh) research data centres (rdcs) statistics canada website census maps and digital spatial data abacus canadian historical geographic information system (chgis) scholars geoportal statistics canada website unified infrastructure for canadian census research (uni cen) overall summary conclusion reference list appendix a: alphabetical list of census portals endnotes exchange of scanned documentation between social scientists and data archives: establishing an image file format and method of transfer by repke de vries and cor van der meer ' steinmetz data archivefor the social sciences amsterdam, holland introduction social science research uses as its raw material not only datasets but also the accompanying documentation: codebooks, questionnaires and so on. sometimes these "guidebooks" are machine readable and available as text files. bui older studies and questionnaires in their original form are all paper only documentation. other examples are handwritten comments on computer print out, sketches and black and while pictures as used in psychological research. needing this kind of documentation means repeated photocopying by archive or library and mail delivery whereas the actual data may travel by networks like the internet or be put on tape or any other computer medium. it's a situation disadvantageous to both the archiving world and the researcher in need of complementary documentation especially if both are geographically wide apart. wasn't the fax machine invented to do just that to get any sketch, image or piece of text instantaneously from a to b ? to an extent yes but the resolution is poor and it is still repeated "photocopying" sending and lots of paper again upon receiving. the image can't be pasted in a research paper, nor can it be stored in a database, viewed on screen or read by ocr packages. fax boards in a personal computer don't change that really: for one thing there can't be constant polling for incommg faxes or a direct telephone connection is not available to the researcher. and though a fax board gives you the image (whatever it is) for the first time as a file on the pc, the resolution is still not good enough. networking on the other hand is mature now: the integration of local area networks with interconnecting nets like the internet, often gives the desktop computer global networking facilities whereas the one fax machine for the department is down the corridor. obviously transferring codcbook pages etc. as images has to follow a different scenario, avoiding the repetition and manual labour in fax and taking advantage of network capabilities: the scanning of the document has to be separate from transfer. scanning should be a one time operation with adequate resolution. storage has to involve compression techniques. the collection of image files could be handled by a specialised database that also holds descriptive and administrative information. or the files might be the result of just scanning a few questionnaire pages with hand written comments. the advantage over fax is that once scanned and stored, sending out an image like any other computer file is easily repeated and initiated. and such scanning can be done at a much higher resolution. storage formats for scanned images can be the own pohcy of archive or library but an exchange format (and the necessary conversion ) should be accepted and adhered to by anyone offering documentation as image files. the transfer comes next and can be done in a number of ways, even as ordinary mail by reprinting the image on paper with a laser printer. network transfer though is easiest and fastest. the researcher needing the pages receives it as a series of small files on his or her own computer or personal file area in a local network. the last step involves a tool for the end user to decompress and actually use the images. ideally the images received can be handled as such by the usual word processing software available to social science researchers. but a free software program will otherwise translate back from the exchange format to a "common denominator" format, if need be. establishing an exchange standard for images. tiff as theformat of choicefor the exchange ofimages. the "lagged information file format" was launched by aldus corporation and microsoft in 1986 and revision 6.0 is now (april 1992) in draft 2 and finalizing. all this time careful attention has been paid to keep the skeleton of the tiff header and the mechanism of the format (a fx)inler structure) the same. older tiff readers or writers therefore can exit gracefully if confronted with a tiff file holding a state of the art colour image. another feature is the use of tags holding vital information about the kind of image, the compression type used for the image block inside the tiff file, but also texts of possibly any length describing the image. lassist quarterly if software can't read tiff' though it promises clearly to do so, it is just because of this versatility. often simpler compression types possible in tiff together with black and white images are handled but grey scale or a more complicated compression method are not reading appropriate tags in the tiff file these packages could have given you helpful hints why it was decided that your tiff variation can't be imported, but most of the time a misleading message on the screen mutters about "incompatible format". if one knows how to read the information, similarly a tiff header dumper program tells you straight away how the image in the tiff file is built up. for data archives and libraries starting the service of making documentation available as images, it is of paramount importance to choose a standard that: • has wide acceptance, is not in any way patented or licensed (with concern to the compression schemes), is not computer type or operating system dependent, has features to make it self-explaining (documentation tags) and is open to new developments in the imaging field but will never be changed in its basic format an indication of the acceptance of tiff as standard for an image file format is the publication last january of the memo "a file format for the exchange of images in the internet" by the network fax working group of the internet engineering task force. authors alan katz and danny cohen from usc information sciences institute, define "the standard file format for the exchange of bitmapped images within the internet" as a particular tiff variation. (tiff-b, preferably with compression type 4). tiff is the format read without any problem by the major ocr programs. format stability is an issue close to the heart of archives. for the storage of images the long term perspective is carefully planned for in the development of the tiff standard. on the other hand the tiff 6.0 revision draft also shows how flexible the standard really is: if libraries or archives take an interest in offering photographic information as images, the same tiff format can act as wrapper but this time with jpeg compression that is now accepted as one of the tiff compacting schemes. in choosing the right kind of tiff format for the exchange standard, the following is presumed: • foremost is the need for scanning and transferring of text together with some une drawings, as in questionnaires. these are called black and white or, bilevel images. • the scanning resolution should be 300 dpi. this gives adequate detail and matches best with the printing resolution of today's average laserprinter. mismatches complicate the software needed to either convert to the exchange standard or use the images afterwards. • the compression chosen should be optimized for bilevel images and pack as tightly as possible • each original page of information is kept as separate image and separate tiff file; tiff has a multi-page feature (one resulting file, holding a number of compressed images) but this option is for the moment not used. • there is a need for adding descriptive information to the image; tiff has tags that can be used for that purpose but this option is for the moment not used. this leads to the choice of tiff compression type 4. well described in the tiff 5.0 paper, still present in the tiff 6.0 revision draft (draft 1, february 1992) as one of the compression schemes. this compression type follows fax group 4. (the two numbers "four" are a coincidence). and fax group 4 is yet another standard and already fully described in the ccitt recommendation t.6. the compression and decompression techniques described in the recommendation are open to anybody for use in own programming. fax group 4 is optimized for bilevel images that hold a mix of text and lines: a lot of white with interspersed black dots. the tags used are (referring to tiff revision 6.0, february 1992): the architectural fields and the resolution fields, both baseline tiff fields. tiff 6.0 has paragraphs in "section 4" (another coincidence) that further define these fields given bi-level images. note that in the text mentioned, compression type 2 is used as a working example whereas the exchange standard employs type 4. in the future the tiff informational fields and document storage and retrieval tags could be exploited to make images self explaining. contrary to the (unused) multipage feature of tiff, these fields or tags can be handled and inspected by the user with any common file viewer: the information stands out as readable text among garbish (though this garbish is a sleeping beauty: it is the scanned and compressed image). either direct at the beginning of the tiff file or at the very end. the steinmeiz archive will help with all necessary spring/summer 1992 19 documentation and expertise if a data archive or library wishes to implement it's own tiff compression type 4 writer or reader. the archive will also provide a testbank service to judge if one starts using the exchange standard with indeed the right tiff 4 format further reading, commercial conversion packages, shareware tiff viewers/printers and the anonymous ftp availability of the excellent "sam leffler" toolkit to start programming for tiff are mentioned at the end of the paper. the katz and cohen proposal, also defines for bilevel images but is less strict in tiff compression type and resolution of the scanned images. the working group allows even uncompressed tiff for example, though tiff type 4 is to be preferred. multi-page files are supported. the perspective however of the proposal seems different from ours: katz and cohen have a strong emphasis on the actual transfer of images and leave it to the sending and receiving parties to negotiate a variation of their standard that both can handle. our emphasis is on establishing an exchange standard that ensures the researcher that he or she can always use the images received. hence one compression type and so on. producing images by document scanning. given the nature of printed text and line art, scanning black and white at 300 dpi or more is adequate. issues of preserving grays in the original or even colour are not involved. each separate page scans into one compressed file and these files are kept together by proper file names and subdirectories or folders to mimic the original chapters and separate volumes. especially in a closed system scanning station with proprietary software it is not always made clear by the suppher how the images are stored in terms of image format and compression type. the compression (decompression) more over is often done by additional, separate hardware. in order to exchange it is imperative that the system has exporting faciuties so that images can be converted to an established standard. next these converted images should be available in a general file area, open to networking and further handling. scanning and storing in a dos environment without specialized image bank software is more or less open by definition and can produce accessible images in a tiff format straight away, though often the less compressing tiff type 2 or even tiff packbits is used. both closed and open approach don't necessarily produce the tiff type 4 chosen as exchange format stfaight away. the following steps have to be taken; first case; a scanning station with own image database software and hardware compression and decompression. used for systematically scanning all paperwork of a number of studies. if there is a choice at all and if one only scans bilevel (black and white) printed source, tiff type 4 is a very good choice for an internal storage format as well. if the software is custom made, even the tiff documentation tags can be filled in to make the image files self explaining. second case: if the scanning station comes as is, caveat emperor: given your computing environment, the image files should still be open for access by other software if internal format "type x" is used, this format should be convertible by both software and hardware to the required tiff compression type 4, exchange format. "hardware" means that the scanning station software asks its compression/decompression board to do the conversion. but can it ? "software" means that a separate tool is available or can be written to convert to tiff type 4. if necessary both approaches can be split in successive steps: if only tiff 2 or 3 can be managed than an off the shelf graphics format conversion package, can do the tiff 2 or 3 to tiff 4 step. note that a scanning setup that uses a storage format that depends entirely on separate hardware, is a timebomb for a data archive or library. at some point in time the hardware board will fail and if a replacement is no longer available, the whole collection of scanned images is rendered useless. if tiff is used as internal format (and given bilevel images) it is very likely to be tiff type 4, because it is the most suitable compression scheme. if not than probably a standard package can do the conversion to the tiff 4 exchange format. third case; a simple scanner setup with some software for viewing and image manipulation, attached to a pc and used for per request scanning of documentary information. all involved packages in this case handle tiff but only aiming at import into desktop publishing software, of the wrong, simpler compression types. standard conversion packages can change into the required tiff 4. note: it is desirable to use this most compact tiff 4 format also for storage to accommodate future similar requests. transferring a group of images. assist quarterly one of the features of the tiff format is storing several different images (pages of text) in one file. scanning and storing however is done on a one page, one image basis. therefore extra processing would be needed to use this feature and it does not really improve the transfer. many smaller files travel easier over a net than one big chunk and dissecting and decompressing would be more complicated for the end user. consequently this feature is not yet part of the exchange standard. compressed images are binary files so in sending over a net, care should be taken to use the transfer protocols accordingly. if the faciuty is text oriented like send file in bitnet, extra steps are necessary, (uuencode and uudecode) to nevertheless preserve the binary nature of images. the ftp protocol used in internet, offers both binary transfer and multiple put to handle a stream of images with one command. name giving of the image files can be a problem: different operating systems have different conventions and too long a name gives trouble if the receiving end has a simpler scheme. the best seems the dos convention (8 characters, dot and a three character extension) just because it is the most restricted one and will fit into any other notation. unfortunately this means that if tiff type 4 is used for internal storage as well and the platform is unix with rich name giving possibilities, one still needs a conversion of the file names. this can be done while copying from the image storage area to the transfer area. the user side: transforming back the images to screen, paper or ocr file. implementing an exchange standard, the focus of attention should of course be the ease of operation for the user to wave the magic stick and have the requested documentation on screen or on laserjet printout. if the data archives or libraries offering the service take the trouble of converting to one and the same exchange format (tiff 4, 300 dpi, single page per file), the steps to be taken are well defined. if text processing software handles graphics, it can import tiff (but only the simpler compression types) and print it out drawing or imaging software does the same and offers viewing. disadvantage of this approach are the required expertise to import a received image into one's word processor and above all: get it printed. drawing software is specialized in importing, viewing and printing images but certainly not everybody masters that kind of software. for various platforms good shareware software is available that reads tiff (again: the simpler compression types) and lets you both view and print all software requires a setup of for the dos environment 286 or 386 pc with vga and a laserprinter available. this printer should be equipped to also print larger chunks of "graphics": an image holding a full page of text is a bit too much for laserprinters with limited graphics capabihties. a few pages of "how to do" information for some software common to most researchers, could ease this do it yourself approach to use tools ah-eady available. such a document will be made available by the steinmetz archive and will also point out the usefulness of some shareware already specialized in doing the job. remains the demand by popular software to be on a "simple compression tiff diet". clearly a conversion aid is needed to change tiff compression type 4 into one of the simpler schemes mentioned. the steinmetz archive has written a tool to do just that and will make it available to data archives to bundle it with requested images. ideally this tool will also have the option to print the decompressed image directly to a hp laserjet printer. printing is much easier accomplished than viewing because of the wide variety in display hardware and the limited resolution or viewing area of most pc screens. the hp printing option will most certainly be included in a future update by the steinmetz archive. (postscript printing would be desirable). with a growing user base for tiff type 4 bilevel images, software makers can be urged to implement importing and handling this tiff compression scheme as well. then the separate conversion step is no longer necessary. in the area of ocr programs this already is the case: tests showed that the market leading ocr programs for dos read the exchange standard format without need for conversion. (recognita, omnipage, prolector, liocr, k5200) summing up, these are the steps for the user to transfer back: 1 . use the free conversion tool to simplify the tiff compression type 2. with the help of the cookbook print or view the images by applying existing software: either commercial packages or shareware. the choice should be software commonly available to social science researchers. (the shareware can be redistributed together with the cookbook both as files through networking and electronic mail) 3. use the exchange format without further ado (for example with ocr software) further reading, availability of software and source spring/summer 1992 code. "a file format for the exchange of images in the internet" by the network fax working group of the internet engineering taskforce. authors: alan katz (katz at isi.edu) and danny cohen (cohen at isi.edu). phone: 310-822-1511. the sam leffler tiff toolkit: by anonymous ftp: sgi.com: /graphics/tiff/v3.0.tar.z (192.48.153.1) email : sam at sgi.com (this toolkit also includes the tiff 6.0 specification ) aldus can be reached at compuserve, again the tiff spec's and a much simpler tiffread toolkit commercial conversion packages to and from tiff type 4: (dos) huaak (dos and sun os): image alchemy (uucp: hsi at netcom.com or: apple! netcom!hsi) (dos) shareware graphics viewers and printers: among others: graphics workshop optiks (all three don't handle tiff 4 so need the free conversion tool first) pixfolio (runs in a dos windows environment) ' presented at the lassist 92 conference held in madison, wisconsin, u.s.a. may 26 29, 1992. this paper was also presented at css92, may 1992, ann arbor, michigan. lassist quarteriy 1/8 higgins, vanessa and carter, jackie (2022) developing data literacy: how data services and data fellowships are creating data skilled social researchers, iassist quarterly 46(2), pp. 1-8. doi: https://doi.org/10.29173/iq.1027 developing data literacy: how data services and data fellowships are creating data skilled social researchers vanessa higgins and jackie carter1 abstract this paper describes two successful approaches to data literacy training within the uk and the synergies and collaborations between these two programmes. the first is a data literacy training programme, being delivered by the uk data service, which focuses on training in basic data literacy skills. the second is a data fellows programme that has been developed to help undergraduate social science students gain real-world experience by applying their classroom skills in the workplace. the paper also discusses next steps in the global development of data literacy skills via the empoderadata project, which is trialling the data fellows programme in latin america. keywords data literacy, information literacy, data services, data skills, sustainable development goals introduction a lack of quantitative data skills among social scientists in the uk has been recognised for over twenty years by government, businesses and research funders (carter et al, 2021c; macinnes, 2009; uk dcms policy, 2021). the importance of this has become more apparent during the coronavirus pandemic when data literacy skills have been critical for research capabilities. as the world becomes increasingly data-rich, we need more data-literate social scientists to contribute to our social and economic understanding locally, nationally and globally. this is particularly important for countries to deliver against the sustainable development goal (sdgs), which rely heavily on quantitative data (united nations, 2015; higgins et al, 2019). there is a wealth of quantitative socio-economic data available for secondary reuse in national data archives/services. for instance, there are national and cross-national government survey data, census data and longitudinal data. these data are incredibly rich research resources, a lot of which is collected by government departments, and could be used much more widely for social research and policymaking if quantitative data literacy skills were more widespread among social researchers. this paper describes two successful approaches to data literacy training within the uk, the synergies and collaborations between these two programmes, and how one is being trialled in latin america. the first is a data literacy training programme being delivered by the uk data service, the second is a data fellows programme that has been developed to help undergraduate social science students gain realworld experience by applying their classroom skills in the workplace. throughout this paper we use the term ‘data literacy’ to describe the broad set of skills that are required to enable data that are hosted in data services and archives (such as the uk data service) to be discovered, critically evaluated and deployed in social research. we follow data-pop alliance’s approach in taking data literacy to mean “the desire and ability to constructively engage in society through or about data” (data-pop alliance, 2015). as such, data literacy is a broad and pragmatic term which intersects with and builds on other literacies, such as information and statistical literacies. data literacy as used here therefore requires conceptual and mechanical understanding in order to be able to do data analysis. moreover, we focus here on the use of quantitative data and analysis, and the data fellows programme is specifically aimed at increasing skills in this area to enable students to critically evaluate and use numerical data (usually but not always statistical data) through data-driven research projects. for further information about how we apply this to the sustainable development goals context see carter et al (2021c). https://doi.org/10.29173/iq.1027 https://ukdataservice.ac.uk/ 2/8 higgins, vanessa and carter, jackie (2022) developing data literacy: how data services and data fellowships are creating data skilled social researchers, iassist quarterly 46(2), pp. 1-8. doi: https://doi.org/10.29173/iq.1027 data literacy training via the uk data service many data archives or data services have data literacy training programmes to encourage and enable the use of the data that they make available for reuse, for example, the uk data service (ukds), gesis leibniz institute for the social sciences, the inter-university consortium for political and social research and consortium of european social science data archives. such programmes are vital to ensure the use of the data held by these archives/services because they provide training to use the specific data available to researchers, rather than generic statistical methodology courses. the ukds data literacy training programme provides a combination of training events and web-based on-demand training materials targeted at an introductory level of data literacy training. the programme includes online workshops with short practical sessions on topics such as getting started with secondary analysis and data in the spotlight workshops to introduce different data types, as well as other formats for training such as coding demos, where code is worked through line-by-line and drop-in sessions where attendees can drop in to meet experts from the ukds to get general advice on using data. on-demand training materials are also available from the ukds website and include short instructional videos and written guidance on topics such as weighting survey data and getting started with spss. the programme also includes a series of interactive data skills modules (figure 1) designed for anyone who wants to start using secondary data, with no prior knowledge required. learners can follow the modules in their own time, dipping in and out when needed, and they receive a certificate of completion at the end of each module. the uk data service training events and on-demand training materials are targeted at a basic level which encourages use of the data among those who have no/little formal data literacy training. this model encourages use among those who may want to dip their toes in the water but do not want to study on a more formal course. another attractive element for newcomers is that the training is free of charge and mostly online so there is no financial cost involved. this also encourages wider participation from those who would not normally be able to access such training, such as voluntary sector organisations, local government and students. figure 1: data skills module example the basic, introductory nature of the ukds training events and materials makes them very popular with over 7000 attendees at events per year and in excess of 7000 views, per month, of the materials https://doi.org/10.29173/iq.1027 https://ukdataservice.ac.uk/ https://www.gesis.org/institut https://www.gesis.org/institut https://www.icpsr.umich.edu/web/pages/ https://www.icpsr.umich.edu/web/pages/ https://www.cessda.eu/ https://trainingmodules.ukdataservice.ac.uk/surveys/#/ 3/8 higgins, vanessa and carter, jackie (2022) developing data literacy: how data services and data fellowships are creating data skilled social researchers, iassist quarterly 46(2), pp. 1-8. doi: https://doi.org/10.29173/iq.1027 on the ukds youtube channel. examples of the benefits of the training events are highlighted in the feedback we received from attendees: ‘it has given me a better idea of what data is already there and stimulated embryonic research ideas! thank you.’ (research fellow) ‘the presentation was extremely useful. it covered a lot of relevant content and contained lots of good links for further reading. the q & a session was helpful to understand how researchers could apply these rules in practice.’ (research technology specialist) ‘bravo! i am a trained computer scientist (hugely comfortable with assumptions, definition, classification, hierarchy) working in computational social science and striving hard to embrace and learn abstraction skills. have never seen anyone summarise this so well before! i feel seen and heard.’ (computer scientist) ‘this will mean i can make thematic maps myself, without having to outsource this to colleagues or contractors.’ (anon) ‘really helpful introduction and prompted me to sign up to the next one, which feels like a natural follow-on. thank you!’ (research associate) likewise, the utility of the data skills modules is highlighted in comments we received from learners on completion of the modules. ‘extremely useful, very thorough and engaging examples. layout is very easy to use and aesthetically pleasing.’ (undergraduate student) ‘i really enjoyed it and i have encouraged my students to complete the module as part of the learning content. a good, all-round introduction.’ (lecturer at higher education institution) ‘this module was very useful. further modules would be interesting on how to use different tests for survey data (anova, t-tests, etc).’ (not-for-profit sector user) data skills training via the university of manchester data fellows programme in 2013, the university of manchester, uk, established the q-step centre. ‘q-step’ was a government funded initiative to support the development of quantitatively skilled social science students across uk universities, to help ‘(i) create a step change in teaching undergraduate social science students quantitative research skills, and (ii) develop a talent pipeline for future careers in applied social research.’ (carter, 2021a). as part of the focus on application of data skills, the university of manchester q-step (umqstep) centre established work placements, which has come to be called the data fellowship programme. since the summer of 2014, three hundred undergraduates have been placed as data fellows into public, private and not-for-profit organisations to carry out data-driven research projects over a two-month period sandwiched between the end of the second and start of the third year of their degree. all students are paid the living wage, ensuring the placements are available to all and not the preserve of those who could afford to do these without payment. as a result, 25% of the placements have been taken up by those from less-advantaged backgrounds or under-represented groups, and 70% have been occupied by females. the programme is therefore addressing not only the need to help social science students acquire data skills practice in the workplace, but is also delivering a diverse talent pipeline into graduate social research careers. the learnings from the programme have been captured in a book on ‘work placements, internships and applied social research’ (carter, 2021b) which helps students find, prepare for, do and reflect on an https://doi.org/10.29173/iq.1027 4/8 higgins, vanessa and carter, jackie (2022) developing data literacy: how data services and data fellowships are creating data skilled social researchers, iassist quarterly 46(2), pp. 1-8. doi: https://doi.org/10.29173/iq.1027 internship. ten case studies and several vignettes are included in the book to illustrate the analytical and research skills, and professional skills, that can be developed through work experience focused on social research, together with examples of early career researchers working in social research organisations. it is important to note that the students who participate in the data fellowships programme are taught statistics and data analysis skills during the first two years of their undergraduate study. as social science majors they are on degree programmes through which they are studying subjects such as criminology, politics and international relations, sociology or social anthropology. the approach taken at the university has been to embed the teaching of data analysis into these substantive subjects (buckley et al, 2015). specific examples of course units teaching data skills can be found in carter (2021b). prior to starting placement students are given refresher courses to enable them to enter the workplace reasonably well-prepared for the data-driven project(s) they will work on. the combination of the classroom teaching and the embedding of data into their social science subject means that students will already have encountered real-world data through their education. for example, criminology students will have used nationally representative crime surveys, politics students will have analysed elections data from, for example, the british elections study, and sociology students will have worked with data from studies such as the british social attitudes survey. many will also have been exposed to international data sources through courses in their first year, for example the measuring inequalities unit a core first year course uses the world values survey and in the engaging with social research compulsory methods course all students on a ba in social science programme will have been introduced to qualitative and quantitative data sources used in their lecturers' research. in most cases they will have been introduced to this data in lectures and then handled the data in pc labs or virtually through practical sessions. consequently, manchester undergraduate social science students will have been exposed to realworld data through their first two years at university, and some will be starting to focus on these skills to specialise on a potential future career in applied social research, whilst others will be keeping their options open and following a broader curriculum. they will have used different software (ranging from excel to r) and have covered a range of methods (from descriptive statistics to bivariate analysis) by the time they take up a data fellowship. creating synergies between the two programmes the two programmes outlined above are both led by the university of manchester. with a proud history of supporting access to quantitative data for empirical research, and strengths in teaching data skills, the university was well-placed to create successful synergies between the two programmes. moreover, with its third strategic goal of ‘social responsibility’ and being a world leader in the impact of its research and teaching as measured against the un’s sustainable development goals (times higher, 2021), the university of manchester is a natural home of the development of data literacy and data skills in the social sciences. the two case studies below evidence how manchester undergraduates have, through the data fellowship programme, worked with the uk data service to enhance their professional and analytical skills and have used this learning in the subsequent steps into their careers. many others have also benefited from the connection between the two programmes, and the two included here are merely selected to illustrate the types of data-driven projects that enable students to springboard into careers that value the data literacy they acquire, coupled with the social science degree subject that they study. in the first case study the undergraduate was placed as a data fellow with a not-for-profit organisation. the experience opened her eyes to a career in supporting data uses and she now works with the uk https://doi.org/10.29173/iq.1027 5/8 higgins, vanessa and carter, jackie (2022) developing data literacy: how data services and data fellowships are creating data skilled social researchers, iassist quarterly 46(2), pp. 1-8. doi: https://doi.org/10.29173/iq.1027 data service. in the second example the student undertook her data fellowship with the uk data service, helping to update the index of multiple deprivation with the latest census data. she now works as a government analyst. case study 1: alle bloom as a data fellow alle completed a placement with respect helpline (a domestic abuse helpline service), then upon graduating completed a master's in social research methods and statistics at the university of manchester. she is now a research associate for the uk data service where she is teaching data literacy skills and helping others to extract value from data. alle says: ‘my data fellowship took me from looking at numbers on my university computer screen, to understanding the people and mechanisms behind them in the ‘real-world’. i’ve always learned better by ‘doing’ and after experiencing the whole process of taking data from collection through to reports that had a tangible impact, i realised this is what i wanted to do. i now work for the uk data service, helping others with their research and teaching, and still continue to learn more about our data and the research process everyday.’ case study 2: klara valentova klara was highly motivated by her firstand second-year course units to learn more about how data indices are created. she secured a data fellowship with the uk data service, helping them to update the carstairs index of deprivation that had been developed in the 1980s. as part of her fellowship she wrote a blog post (https://blog.ukdataservice.ac.uk/meet-our-interns-klara/) from which this extract is adapted: ‘carstairs … comprises four indicators from the census, which relate to material deprivation (overcrowding, male unemployment, low social class and lack of car ownership). some of these variables are a bit outdated, and so we include other indicators, which are more up to date. … [including] total unemployment (female and male combined) in our calculations as there are [many] more women in the labour force than there were nearly 40 years ago.’ she went on to use the findings (and outputs) from her placement in her final year dissertation, took that into her graduate study and is now working at the office for national statistics. the data fellowship opened up a career route for her that she had previously not considered, and she was able to evidence the skills and knowledge she had acquired during her undergraduate study. klara’s reflections on her time spent as a data fellow are captured here: ‘i realised just how powerful data can be. by examining a deprivation index created in the 1980s, it became clear that if outdated data or inappropriate methods are used, the outputs are likely to be biased and in turn lead to inaccurate policy decision making. on the other hand, if data is analysed correctly, it can be a powerful tool to help improve people’s lives. i am now working as a government analyst promoting best practice in data collection, analysis and dissemination so we can draw more value from statistics for public good.’ the combination of teaching with real-world data made available through the uk data service, and helping students acquire data skills through the workplace, is a powerful one. the success story is in the development of a pipeline of talented, curious, data literate social science graduates who can enter careers in the data professions. https://doi.org/10.29173/iq.1027 https://blog.ukdataservice.ac.uk/meet-our-interns-klara/ 6/8 higgins, vanessa and carter, jackie (2022) developing data literacy: how data services and data fellowships are creating data skilled social researchers, iassist quarterly 46(2), pp. 1-8. doi: https://doi.org/10.29173/iq.1027 developing data literacy globally – next steps the success of the data fellows programme has led to interest from education and civic society organisations interested in developing the model in their own countries. through a collaboration with datapop alliance, that resulted from a data fellow being placed with an organisation they work closely with (open data watch), we have been able to develop an international dimension to the initiative. we describe the origins and early stages of this research in carter et. al (2021c) where we discuss the empoderadata project which has explored the transferability of the data fellows scheme to colombia, mexico and brazil. the full results from the early stages of the empoderadata project are available in higgins et al (2019). the empoderadata research to date has uncovered a need for basic data skills training (similar to those covered in the uk data service data literacy programme) within the three case study countries. moreover, within brazil and colombia, the data fellowship model is perceived as a valuable tool for building basic data literacy capacity to help deliver the sdgs. the notion of hybrid professional teams has emerged which would bring together ‘fellows’ with complementary backgrounds to work collaboratively on sdg-related research projects at host organisations. examples of hybrid professional teams could be social scientists and data scientists working together or social scientists working alongside stem graduates (higgins et al, 2019; carter et al, 2021c). as a direct result of the empoderadata research project, two pilot data fellow programmes are being implemented in colombia and brazil (jones et al, 2021; carter et al 2021c). these early initiatives are a positive step forward and they illustrate that the university of manchester data fellows model can be adapted to different disciplines (business studies and mathematics) and has international relevance and appeal. conclusion there is a wealth of quantitative socio-economic data available for secondary reuse in national data archives/service but there is a lack of quantitative data literacy skills among social researchers, or others who may want to use these data; this is particularly important for the monitoring of the sdgs. different models of data literacy training such as the uk data service data literacy training programme and the university of manchester data fellows programme have created successful synergies to benefit undergraduate social science students’ data literacy training experience. the synergies outlined in this paper are a useful model for other data literacy programmes to apply to strengthen the pipeline of data skills training among undergraduates. the empoderadata research project has led to new initiatives to trial data fellow programmes in brazil and colombia. if successful, the programme has potential to be expanded to more countries and across different disciplines. this collaboration provides an exciting future research agenda which we will be evaluating carefully. acknowledgments the authors would like to thank julie ricard, valentina casasbeunas and emmanuel letouzé from data-pop alliance for their collaboration on empoderadata phase 1. a huge thanks to the uk data service staff members who deliver training events and materials and those who hosted the data fellows. thanks to klara valentova and alle bloom for their permission to include them as case studies within this paper. financial support for the research reported here has been thanks to various grants from the uk esrc (economic and research council), the nuffield foundation and the uk global challenges research fund (gcrf) as well as the university of manchester who co-fund the data fellowships. https://doi.org/10.29173/iq.1027 https://datapopalliance.org/empoderadata-project/ 7/8 higgins, vanessa and carter, jackie (2022) developing data literacy: how data services and data fellowships are creating data skilled social researchers, iassist quarterly 46(2), pp. 1-8. doi: https://doi.org/10.29173/iq.1027 references buckley, j, brown, m, thompson, s, olsen, w & carter, j 2015, 'embedding quantitative skills into the social science curriculum: case studies from manchester', international journal of social research methodology. https://doi.org/10.1080/13645579.2015.1062624. carter, j 2021a, 'developing a future pipeline of applied social researchers through experiential learning: the case of a data fellows programme', statistical journal of the iaos, pp. 1-16. https://doi.org/10.3233/sji-210844. carter, j (2021b) work placements, internships and applied social research, sage, london. carter, j, méndez-romero, ra, jones, p, higgins, v & samartini, als. (2021c). empoderadata: sharing a successful work-placement data skills training model within latin america, to develop capacity to deliver the sdg, statistical journal of the iaos. https://doi.org/10.3233/sji-210842. data-pop alliance. 2015 beyond data literacy: reinventing community engagement and empowerment in the age of data. white paper. http://datapopalliance.org/wpcontent/uploads/2015/10/beyonddataliteracy_datapopalliance_sept30.pdf. higgins, v, casasbuenas, v, ricard, j & carter, j 2019, empoderadata: data literacy assessment and sustainable development goals data gaps: brazil, colombia and mexico. data-pop alliance. https://datapopalliance.org/wpcontent/uploads/2019/11/empoderadatareport_final_oct2019.pdf. jones, p, carter, j, renken, j & arbeláez tobón, m 2021 'strengthening the skills pipeline for statistical capacity development to meet the demands of sustainable development: implementing a data fellowship model in colombia' digital development working paper series, 89 edn, university of manchester, global development institute, manchester. http://hummedia.manchester.ac.uk/institutes/gdi/publications/workingpapers/di/dd_wp89.p df. macinnes, j. (2009). final report: proposals to support and improve the teaching of quantitative research methods at undergraduate level in the uk. swindon: economic and social research council. times higher (2021) sdgs global impact rankings https://www.timeshighereducation.com/impactrankings#!/page/0/length/25/sort_by/rank/sort_or der/asc/cols/undefined [accessed 19 jan 2022]. uk dcms policy (2021) united kingdom government department for digital, culture, media and sport policy paper: national data strategy [internet] gov.uk 2020 [cited 25 may 2021]. available from: https://www.gov.uk/government/publications/uk-national-data-strategy/national-datastrategy. united nations. transforming our world: the 2030 agenda for sustainable development [internet]. sustainable development goals knowledge platform; 2015 [cited 2021 may 20] p. 41. report no.res/70/ https://sustainabledevelopment.un.org/content/documents/21252030%20agenda%20for%2 0sustainable%20development%20web.pdf. https://doi.org/10.29173/iq.1027 https://doi.org/10.1080/13645579.2015.1062624 https://doi.org/10.3233/sji-210844 https://doi.org/10.3233/sji-210842 http://datapopalliance.org/wp-content/uploads/2015/10/beyonddataliteracy_datapopalliance_sept30.pdf http://datapopalliance.org/wp-content/uploads/2015/10/beyonddataliteracy_datapopalliance_sept30.pdf https://www.research.manchester.ac.uk/portal/jackie.carter.html https://datapopalliance.org/wp-content/uploads/2019/11/empoderadatareport_final_oct2019.pdf https://datapopalliance.org/wp-content/uploads/2019/11/empoderadatareport_final_oct2019.pdf http://hummedia.manchester.ac.uk/institutes/gdi/publications/workingpapers/di/dd_wp89.pdf http://hummedia.manchester.ac.uk/institutes/gdi/publications/workingpapers/di/dd_wp89.pdf https://www.timeshighereducation.com/impactrankings#!/page/0/length/25/sort_by/rank/sort_order/asc/cols/undefined https://www.timeshighereducation.com/impactrankings#!/page/0/length/25/sort_by/rank/sort_order/asc/cols/undefined https://www.gov.uk/government/publications/uk-national-data-strategy/national-data-strategy https://www.gov.uk/government/publications/uk-national-data-strategy/national-data-strategy https://sustainabledevelopment.un.org/content/documents/21252030%20agenda%20for%20sustainable%20development%20web.pdf https://sustainabledevelopment.un.org/content/documents/21252030%20agenda%20for%20sustainable%20development%20web.pdf 8/8 higgins, vanessa and carter, jackie (2022) developing data literacy: how data services and data fellowships are creating data skilled social researchers, iassist quarterly 46(2), pp. 1-8. doi: https://doi.org/10.29173/iq.1027 endnotes 1 dr vanessa higgins, university of manchester, vanessa.higgins@manchester.ac.uk (corresponding author) 2 professor jackie carter, university of manchester, jackie.carter@manchester.ac.uk https://doi.org/10.29173/iq.1027 mailto:vanessa.higgins@manchester.ac.uk file:///c:/users/oschwart/stokes/iq/iq46_2/jackie.carter@manchester.ac.uk vol264 12 iassist quarterly winter 2002 iassist quarterly winter 2002 13 abstract: in the last five years the exponential growth and use of electronic resources has surpassed all former predictions. the rise in use of e-journals, bibliographic databases and research support facilities has resulted in the researcher relying far more heavily on web based resources than ever before. one of the acknowledged problems that now faces the researcher is how to filter, access and apply the information, using the wide variety of resource discovery tools now available to them. at mimas (manchester information and associated services, one of the three uk national data centres) we attempt to provide a one-stop-shop for our users ̓needs, be they undergraduate, postgraduate or post-doctorate. this paper seeks to show how mimas services can be used at every stage of the research development process, from the inception of an idea, through its development, to the final presentation of results. facilities like isiʼs web of science, copac (which merges online catalogues of 21 of the largest university research libraries in the uk and ireland), jstor (a retrospective digital archive of scholarly journals) together with many of the other services hosted at mimas now provide support for researchers at every level of the creative process. the world of scholarly writing has proven itself to be a highly dynamic environment. researchers are now expected to be literate in an ever increasing number of information seeking skills with the emergence of the internet as a new medium for resource location. the ability to search across huge databases, datasets and archives is now possible, in a way that would never have been conceivable to the first researchers. one thing that has not changed is the basic structure of a written piece of research. some sections may be slightly larger than others but in principle, there are six separate stages in the process of creating research, which can be identified. two authors clearly identify these stages. thomas and brubaker describe the structure of research papers in the following way. ‘the most popular six chapter structure for the dissertation in use today consists of (a) introduction, (b) review of the literature, (c) methodology, (d) results/ findings, (e) analysis and interpretation of findings, and (f) summary, conclusions, by jessica eustace* reaching your end-user with mimas applications and recommendations for further study.’1 edminster and moxley concur with this description of research structure. ʻalthough some modification of the six chapter format occurs in dissertations in the humanities, its influence is still clearly visible in most graduate work and the view that such work represents original thought by an individual author who merits recognition and reward for that originality continues to prevail.ʼ2 one question we must ask ourselves is, with the basic structure in place for the paper, why do some researchers struggle in creating their research paper? rosenfeldt and dowling refer to the tyranny of the blank page.3 research into the answer to this question can be found in literature on the creation of training programmes, studies into user searching skills, and many of the library and information journals have examined various facets of this question. i shall briefly look at some of the main causes for this problem. the most commonly recognised cause is computer anxiety. computer anxiety can occur in three distinct forms. these are recognised by torkzadeh and angulo in their paper in 1992. “a) psychological anxiety which includes fear of damaging the computer; b) sociological anxiety, which is characterized by the need for social contact and the fear of being replaced by a machine and c) operational anxiety, which includes an inability to type and/or operate the machine.”4 another reason for difficulty could be the sea of information now available to the researcher. the presentation and manipulation of information is now possible in such a huge variety of ways that it is no wonder many scholars are intimidated at the onset. many suffer with feelings of information overload. coupled with this there are so many new initiatives and projects providing new ways into the variety of data available to the scholar that where to start can be the greatest challenge. stewart highlights the dilemma researchers now face, ʻ… in this technological and techno phobic environment, novices are faced with a variety of systems with different interfaces, 14 iassist quarterly winter 2002 iassist quarterly winter 2002 15 search capabilities, commands, and screen layouts among other options.ʼ5 the amount of information available is not the only significant challenge which the change in scholarly writing has faced. the range of skills users need to seek the information they require has continued to diversify. researchers are no longer just required to be able to read and write but to be able to search across a variety of mediums from the book and journal to the internet, the world wide web and the huge range of searching tools which are now available. with growing numbers of mature and foreign students the diversity of their capabilities to manage in this environment is another challenge that service providers must consider. the ranges of abilities from novice to expert all require different levels of help. diane nahl discusses this in her paper, ʻsince the mid1980s, complex information retrieval systems have been available in academic libraries. as a result librarians have dealt with increasing numbers of novice end-users needing instruction in the use of database systems. at the turn of the century, the fast growing technological environment continues to introduce challenges for reference and instruction librarians who strive to help novices acquire effective search behavior.ʼ6 to allow this paper to have a focus i decided to select a topic that falls within all three major disciplines. the topic is the development of psychiatry. this topic can be taken from a scientific, a social scientific or a humanities angle depending on the researcherʼs emphasis. the reason why i have selected a cross disciplinary topic is due to the changing nature of research today. research no longer falls neatly within the old ideas of specific discipline but now calls upon a wide range of cross-disciplinary information. many of the tools i will discuss have acknowledged this change and allow cross-disciplinary searching. for our topic area the science researcher could view it as the development of psychiatry from a medical point of view. in the same way a social scientist may view this topic as the development of psychiatry as a social awareness topic. the humanities student could also look at this topic on the grounds of a historical project. lois buttlar in her discussion of the information sources for doctoral research discusses the importance of the understanding of cross-discipline searching. she refers to chubinʼs work in this area. ʻ… chubin (1976) discusses bradford notions of “core and scatter” noting that while a discipline is centred around an intellectual core, knowledge about communication outside the core (or scatter) indicates how the disciplines overlap. it is important for information professionals to understand the dynamics of knowledge overlap in research, especially since the hypertext technology has facilitated cross-disciplinary exchanges.ʼ7 the aim of this paper is to highlight how the crossdisciplinary services available at mimas can be used in the creation of a research paper. i shall illustrate how the researcher can draw on resources, which fall into a variety of disciplines, can be used throughout the research process and also make suggestions for their hypothetical use in a mock paper. in the following paper i shall illustrate how the resources of one data centre can be used to enhance, support and develop a researcherʼs paper. my aim is not to discuss how one data centre or one archive of information is better than another but simply to show how a wide variety of services covering cross-disciplinary data can be used in the creation of a single research paper. in this paper i will be looking at the mimas data centre. in the uk there are three data centres: mimas (manchester information and associated services) based at manchester university, edina (edinburgh data and information access) and bids (bath information data services) mimas is a jisc supported national data centre run by manchester computing, at the university of manchester, to serve the uk academic community. the range of services, hosted at mimas, provides flexible online access to socioeconomic, spatial and scientific data, and to bibliographic and electronic journal data services. the data service had previously been called midas (manchester information data sets) but due to copyright law the name was changed to mimas in july of 1999. i will briefly run through each of these services before looking at how they may be applied at the different stages of creating a piece of research. the services hosted at mimas can be broken down into six separate sections. the bibliographic reference services include four services, archives hub, copac, isi web of science and zetoc. the electronic journal services include jstor and nesli. scientific services include crossfire, csd and mossbaur but i will only be looking at casweb. the socio-economic data include census data that is accessed through casweb, surveys, ns databanks and international data. spatial data is made available in map data and satellite data. the archives hub service revolutionizes access to the archive collections held in uk universities and colleges by making descriptions of them available on the internet, many for the first time. there are currently over 40 institutions contributing to the collection. records listed on the service cover a wide range of sources, including papers of statesmen, scholars, soldiers, scientists and storytellers. the current descriptions are just “top-level”, giving an overview of each collection, with a brief biography of the creator of the documents, a note on the contents of the archive and idea of the quantity of material concerned. where possible this top-level description is linked to a full catalogue and each archive description is indexed to aid researchers in locating relevant material. 14 iassist quarterly winter 2002 iassist quarterly winter 2002 15 copac provides free access to the merged catalogues of major uk and irish university research libraries. within the database there are over 8 million records from the 22 contributing libraries, with more being loaded. records found in copac represent a variety of materials across most subject areas. some include links to full-text documents made available by other services. the materials range in date from c.1100 ad to the most recent documents with 300 languages being represented. further to the bibliographic details of the records, copac also has realtime local availability information that has been provided for some libraries. records can be downloaded via email and imported into reference management software. the web of science is the web interface providing access to isiʼs three central databases. these are: science citation index, social sciences citation index, arts & humanities citation index, and the interface was designed by the institute for scientific information (isi). the data in these databases covers articles dating back to 1981. isi proceedings, the sister service of web of science, has two databases, the science and technical edition (stp) and the social science and humanities edition (sshp). coverage in isi web of science and isi proceedings continues to grow. currently there are more than 8500 journals and records of nearly 2 million papers in proceedings have been indexed by isi, and over 20,000 new items are added each week. in addition to these two central services, the web of science service also provides access to additional isi products. these include current contents connect, journal citation reports, derwent innovations index, isi chemserv and biosis previews. zetoc provides z39.50-compliant access to the british libraryʼs electronic table of contents (etoc). the service was launched in september 2000. the database contains over 19 million article titles from over 20,000 of the most important research journals and 16,000 conference proceedings, covering every imaginable subject in science, technology, medicine, engineering, business, law, finance and the humanities. records start from 1993 and the database is updated daily with approximately 10,000 new records. the service also allows users to send themselves alerts when new issues of journals arrive. this is known as zetoc alert. macintyre and apps describe this service, ʻto supplement the basic search and retrieval functionality of the service, a current awareness alerting service based on the table of contents data was also developed. the aim of was to ʻfill the gap ̓left by the demise of the ʻautojournals ̓ service offered by bids, another of the ukʼs national academic datacentres, until july 2000.”8 jstor was established as an independent not-for-profit organization in august 1995; it is a retrospective collection of electronic full text articles. the mimas based jstor service is a mirror service of the american based service. jstor combines the advantages of page images with searchable full-text, and stores the data in both forms. jstor has a full-text electronic archive collection of noncurrent issues of over 100 core scholarly journals covering a wide range of subjects. coverage within the collection starts with the very first issues (many dating from 1800s). jstor continues to expand and in the last two years three new archive collections have been added to the service. the new collections are called ecology & botany collection, arts & sciences ii collection and the business collection. this year we will see the launch of a new language & literature collection. nesli delivers a national electronic journal service to the uk higher education and research community. the managing agent undertakes negotiations with journal publishers on behalf of uk higher education institutions for electronic journals at preferential terms. once a publisher has created a provisional offer it is then sent to the journals working group or jwg before the final offer is sent out. access to the subscribed journals is done in several ways. institutions have the option, depending on the deal, to access the information either directly from the publisherʼs site or through the nesli/swetsnetnavigator www gateway. this gateway provides a single point of access to the full text of electronic journals. the range of journals available through this access point is dependent on the publisherʼs agreement to have access through this point. site licenses are available for individual e-journals or packages of e-journals, depending on the current publisherʼs deal. for the period of 2002 there are currently 21 publisher deals in progress. material covered by these deals is cross-disciplinary in nature. crossfire is a comprehensive chemical information management system comprising the crossfire beilstein and crossfire gmelin databases and the beilstein commander and crossfire web software which are used to search and access the databases. crossfire beilstein contains over 8 million organic compounds, with over 8 million reactions and over 4 million citations. crossfire gmelin contains over 1.5 million inorganic and organometallic compounds, 1.3 million reactions and 1 million citations. casweb was originally developed by james harris at the esrc/jisc and funded by the census dissemination unit (cdu), which forms part of the mimas service. casweb is just one piece of software found in the cdu but is most appropriate to this paper. i will not be discussing the cdu as this topic will be covered by my colleague, jackie carter at a later stage of this conference. casweb is a web interface to statistics and related information from the 1991 united kingdom census of population. these large and complex datasets comprise aggregate counts of persons and households for various geographical units. the census sas/lbs datasets contain over 10,000 items of information, which are arranged for convenience into tables. the casweb system has preserved this arrangement 16 iassist quarterly winter 2002 iassist quarterly winter 2002 17 of the data. the landmap project uses data from two different satellites, landsat and spot. the landsat, an american satellite, consists of 32 scenes covering the uk. the landsat data maps allow users to view maps at 30m resolution. it is used in the study of marine environment, water resources, land use, ecological and engineering applications, agriculture and forestry monitoring. the spot, a french satellite, is composed of 152 scenes and covers the uk and republic of ireland. spot allows users to view data maps down to 10m resolution. it is used in much the same way as landsat but it is mainly concentrating on studies of land use mapping and digital terrain modeling. as i have illustrated the mimas services cover a wide range of disciplines and can be applied to research in a variety of ways. using the topic of the development of psychiatry i shall now show how these tools can be employed through the six stage process of creating a research paper. introduction rosenfeldt and dowling describe this stage in the following way, “… once you begin, you must face the tyranny of the blank page – a major cause of procrastination. the first step is to get something down. begin with a plan.”9 most of the work done before hand is centred upon reaching the core of the research project. “areas focused on here include drawing up a short list of topics, selecting a topic for investigation, formulating a general question and focusing a research question. the importance of spending as much time as necessary to get the question right is highlighted here as this often causes new researchers considerable difficulty.”10 it is the pre-cursor for the creation of your original idea, the core from where the research will start and finish. at this point the user reads a huge amount of literature to create this original idea. the introduction is the point where the net of research is the widest pulling in resources from cross-disciplinary sources. portals like those provided at ensure that the articles, journals and other resources found at mimas are of highest quality and suitable for inclusion in a piece of scholarly writing. mimas resources like copac are particularly useful at this stage. copac as a union catalogue of 22 libraries provides bibliographic details encompassing an enormous range of material ranging from maps to pictures to written texts. this allows the users to discover resources from a huge database of resources. the interface was designed to be as simple as possible. it is possible for users to search under author, keyword etc. in some cases articles in copac link straight through to the full text of articles. alternatively, users of copac can view their local holding details of libraries and can arrange for interlibrary loans. a search on psychiatry produces a large set of results, totalling 12,485, amongst these results we can see material including books, journal articles, medical textbooks, handbooks, illustrated texts and even lecture notes. another resource, which can be applied at this stage of the research process, is the archives hub. as i mentioned earlier the archives hub serves as a gateway for archives in the university sector. archives of relevant material can be located using this resource. information at the archive level includes details on the collection itself, but also details on the repository. details include important telephone numbers, opening hours, etc. in this way the researcher can identify archive collections which may be useful to their research topic. web of science is another useful tool that the researcher can utilize at this stage of the research process. for novice users this has proven to be a useful tool to help them through their research. two functions of searching exist within the service. full search that allows users to run advanced searches using boolean operators and more complex search queries. in full search it is possible to view 500 records at a time. the second search available is easy search; this allows users to search under person, place and topic. when the researcher searches the database for place, he can create a search to retrieve the most recent articles published by researchers working in a particular institution (college, university, company, etc.) or geographical area (country, city postal code, etc.). a search on person, will produce results on the person named as an author or a cited author or on an individual you want information on. topic searches allow the user to find articles that discuss the key topics in their research. this is a good starting place when the researcher is simply trying to come to grips with the research question. results from a topic search are displayed sorted by relevance. this means that the system will list first those articles that contain the most frequent occurrences of the words and/or phrases entered by the user to describe the topic. alternatively results can be displayed in reverse chronological order. this lists the most recent articles first based on the date on which the journal was processed at isi. zetoc is another mimas tool, which can be applied at this phase of the paper. zetoc can be used to search the tables of contents of journals hosted at the british library. zetoc can also be used to set up an alerting service that will notify the researcher when relevant journals he has recognized as useful to his research have been published. in using zetoc, researchers can ensure that the future work they do in their paper will be as up to date with theories in the field as possible prior to submission and or publication. literature review at this stage the focus of the research has been completed. the researcher now reads and gathers literature from the chief influences for the topic area. it is important that the 16 iassist quarterly winter 2002 iassist quarterly winter 2002 17 material included in this section is accurate and of the highest possible quality and relevance to the research topic, and the research question. the researcher must use this section of the paper to demonstrate the relationship between the proposed research and what has already been done in the area. web of science can be used in a slightly different way at this point. by using the full search as opposed to the easy search the researcher can track authors ̓works and also the authors that have cited the authorʼs work. we can also widen the number of articles by looking at related records. related records allow you to view articles that have shared one or more of the articles used in the original articleʼs reference list. web of science also allows users to run cited reference searches. this resource allows users to track a particular work by an author through its citations. in this way the web of science can be used as a tree so users can move from the start of a particular author and track backwards through the inspiration of the paper in question. they can move from branch to branch from the root of a particular topic. in the same way the user can then look beyond the original paper, and see how the paper has influenced later authors on the topic in question. users can use this unique method of moving along the branches of knowledge built up within the web of science. two services within the web of science allow users to move beyond the bibliographic records to the full text of an article. the first service is the holdings service. this service allows users to click on a button in the top right of the screen and link through to their host institutionʼs opac and see if the article in question is hosted at their own library. in this way they have access to the hard copy of the text. the second service is called the linkage service. isi makes it easy to link between the web of science and electronic full text. isiʼs goal is to provide users with bi-directional links for navigation back and forth between records in the web of science and corresponding full-text documents and references. isi links provides reliable access to full text. a link button will only appear when there is an active link to a document. this “no dead links” policy insures that web of science users will only get links that resolve to publishers ̓electronic full-text documents. isi actively seeks full-text linking partnerships with primary journal publishers and content hosts for all available electronic journals covered by the web of science. currently, over 5.4 million record links are available between the web of science and electronic journals and databases. currently isi has links set up with 14 core publishers. as i mentioned earlier in the paper, web of science has other allied tools from isi, one of which is the journal citation reports or jcr. this tool allows users to see which journals have been cited the most. this will help them determine which journals are the most pertinent and useful in the search for valued peer reviewed articles to support their work. this tool can be used in conjunction with zetoc during the creation of the journal alerting list. copac can again be used at this stage of the research process. now that the main authors on the topic have been identified, copac using the union catalogue can be used to track down less common articles and references. using the author/ title search allows the user to be far more precise than on earlier searches. it is at this point that the holdings function in copac becomes most useful. at the base of a full record, copac has holdings material that will allow researchers to either ask for an ill from a specific repository or travel to its location. in this way more obscure texts can be identified and located. jstor can be called upon at this stage to look for articles that have retrospective influence on the research topic. searches can be made as easy or as complex as the user wishes by building up searches on the search form. users have the option of selecting to search journal topic groups in clusters or to expand the list and search on individual journal titles. this later option tends to produce more accurate search results. once the list of results is compiled the researcher can click to view a pdf of the full text of the article. jstor uses high-resolution images to store, display and print faithful replications of the pages that make up the complete published record of the journals in its archive. users can view the article in an electronic copy of its original format. the ability to view the full text has proven itself to be very popular as researchers increasingly look for faster access to the full text of articles. nesli is another route for the researcher to reach the electronic journal. nesli uses the swetsnet navigator interface. users can log on to the service and search through the full text electronic journals of the nesli publisher offers which their institution has subscribed to. users can search on journal title, publisher or they can run a general or an advanced search. the journal title and publisher searches allow the researcher to type in one word from the journal title or publisher name and it will search for relevant records. the general search allows users to enter a date range between 1996 to the current year. they can then enter search terms, which will look for articles on article title, author, and abstract or on keywords. the advanced search allows users to select a date range again but they can create a more complex search using boolean operators. once a list of results is created the researcher is then able to click through to the abstract or to the full text of the article which they can view in a number of formats. most articles are viewable as a pdf. the article will include any images that authors have included in their original article. the researcher can then print these articles to read later. 18 iassist quarterly winter 2002 iassist quarterly winter 2002 19 methodology the methodology section of the paper is used to give a detailed description of how the research question or questions will be answered through the paper. this section will identify the steps the research will follow in order to answer the original research question. this stage is only possible through the work completed in the prior two stages. this section also normally discusses the application of either qualitative or quantitative research methods. qualitative research examines topics that cannot be examined through statistical analysis. this research area has been described as humanist, realist, subjectivist and observational. the researcher attempts to qualify reasons for certain behaviour and to explain their occurrence in a scientific way. quantitative research relies on empirical data. this is normally gathered through a structured questionnaire and basic statistical analysis. searching for qualitative and quantitative articles to support whichever method the researcher plans to utilise can be done using several of the bibliographic tools that were used in the introductory section. by simply using these terms in web of science the researcher can view articles on both areas of methodology and look at different applications of these theories within varied disciplines. copac can again be used to track down articles not cited in the web of science. jstor can be used to look at developments since the origins of the research methodology question in the journal archives. results/findings the results and findings section of the paper is the point where the researcher must provide the proof of the research undertaken. sources for his results and findings can come from an array of sources. casweb is the most obvious service to be used at this stage of the research process. casweb allows an intuitive access point to the census of 1991 using a familiar windows-based system. the researcher can click through the application first searching on specific tables of data, which have been retained for ease of use mirroring the original census format. this is a huge change from the earlier saspac (small area statistics package) system that required users to have knowledge of the structure and content of the census data. in this way it acts as an extraction process. while the user is able to extract specific information from the census they would then have to move the data from the remote machine before they could make use of it. casweb in direct contrast works as an acquisition process. it enables the user to extract the information but acquire the information directly onto their local machine. in this way the data becomes far more easily accessible. to access the data you follow a simple 3-stage process. 1. in step 1 you define the study area you are interested in and select the geographical units for which you wish to extract data. 2. step 2 is where you define what data you want to extract from the census database. the census sas/lbs datasets contain over 10,000 items of information 3. step 3 is the data extraction / download stage the result is the user has a route to a diverse array of data, which can be applied in many ways. for the purpose of our research topic use could be made of the data by looking at the number of psychiatric institutions that exist within a specific area and compare that data with figures found in past censuses for the same area. casweb also allows users to map the data they gather using arcview or mapinfo. in this way users can produce a visual representation of psychiatric institutions at ward level, county level or countrywide level allowing for an alternative view of the data away from the census tables. this can help the researcher to interpret these figures in a different way. an additional web of science product called derwent innovations index can also be utilised at this stage of the paper. derwent innovations index provides a unique access point into patent material over the last twenty years. users can search through patents based on patent topic area and discover innovations within their desired subject area. for the purpose of this paper typing in a topic word like psychiatry produces over 100 hits. within this material patents can be found on electrical equipment used for monitoring patients ̓responses to innovations in drugs for psychiatric patients. one of the results looks at the use of different coloured lenses within glasses to affect the mood of emotionally disturbed patients over long periods of time. jstor could also be used at this point of the paper. journal articles in jstor can be used to illustrate changes in technology and their application. since the archive carries papers from their first edition the development of theories, instruments and the success of certain drugs can be traced within journal articles. these articles can be used to support or refute current beliefs in the psychiatric field. the archive also holds images which were included in the original paper. these pictures can be used to illustrate changes in technology and their application for treatment, through the centuries, of psychiatric patients. crossfire is a chemists ̓tool but can also be made useful for this paper. the researcher can use crossfire to view the chemical structure of drugs that were prescribed to patients. the application provides highly detailed information on how to purify chemicals, their basic physical attributes and their reactions with other elements. if a paper concentrated on the development of chemical treatment of psychiatric patients this could prove to be a very useful tool. 18 iassist quarterly winter 2002 iassist quarterly winter 2002 19 landmap can be used at this stage to illustrate the layout of a specific geographical area. the satellite images allow a unique view of the land and geographical features. if we use this in conjunction with the data from casweb viewed in arcview it can allow the researcher to provide another view of an area of interest. hypothetically, if the research topic discussed a specific institution which was originally built at an earlier point in history, details recorded at that point, for example, a map of the institution and its surrounding areas could be compared with current satellite data. this could be used to illustrate the growth of an institution or to highlight changes to the surrounding areas influencing the growth of the institution. analysis of findings the analysis of findings is a section where the researcher compiles the knowledge, resources and understanding of the topic area, the discipline and his/her findings for the paper. it is not possible to suggest an appropriate mimas based resource tool, as the user will be applying the information gathered during the prior stages of the research process. to a greater or lesser extent this data will have been gathered from the resource tools i have already suggested, and it is in this way mimas resources have benefited the user at this section of the research topic. summary/conclusions the summary of the research paper is the final point where the researcher highlights how he has laid the case for the research question and provided answers for the query under discussion. it is at this point that the researcher culminates all of the work he has put into the paper and summarises how he has achieved the aims and goals he laid out in the introduction. this would not be possible without the influence of the tools he has applied throughout the project. as you can see the mimas tools are cross disciplinary, and yet all can be used to greater or lesser extent in various stages of the research process. it is important that researchers feel comfortable in using these resources and that they are capable of making use of these information seeking tools. i have provided numerous hypothetical applications of these tools to illustrate the cross disciplinary nature of research today and provide some suggestions for how researchers can make use of our resources. to conclude, mimas tools are multi-disciplinary. the researcher should not be intimidated by the variety available, but embrace them. as research continues to blur the old divisions of research between arts and social science and social science and science it is important that tools exist which encompass this move. the facilities at mimas have been designed to be as intuitive and user friendly as possible. each service has an individual helpdesk that provides weekday support when users become confused or lost. many of the applications utilise context sensitive help as a first port of call. all of the bibliographic software also has faqs (frequently asked questions) sections to try and answer some of the core common queries to help hesitant users from feeling intimidated in asking for help. zetoc and copac share the same interface allowing for familiarity of the resource tool for users. similarly the web of science and all of its additional products also use the same interface allowing for increased familiarity for the user on utilising the various products. casweb was designed specifically to be user-friendly. it heralded a move from the old saspak system to a familiar windows interface with online help. mimas services have been created to help the end user find the information they seek using intuitive, clear and user-friendly tools. in this way mimas can be viewed as an excellent foundation with reliable resource tools to aid researchers to reach their goal of creating a high quality, insightful research project within their field of expertise. paper presented at the iassist conference, june 2002, in storrs, ct, usa. * jessica eustace, university of manchester, uk. e-mail: jess.eustace@man.ac.uk bibliography: 1. barbara brien , beth hillman and victoria topp. effectiveness of hands-on instruction of electronic resources. research strategies 16 (1) 1998. pp41-51 2. lois buttlar. information sources in library and information science doctoral research. library and information science research 21(2) 1999. pp227-245 3. jude edminster and joe moxley. graduate education and the evolving genre of electronic theses and dissertations. computer and composition 116, 2002. pp116 4. emily fabiano, casting the net: reaching out to doctoral students in education. research strategies 14 (1996), 159-68 5. nancy fjallbrant. evaluation in a user education programme. journal of librarianship 9 (2) 1977. pp83-95 6. a.j jerebek and linda meter. “library anxiety” and “computer anxiety”: measures, validity and research implications. library and information science research, 23, 2001. pp277-289 7. ross macintyre and ann apps. working with the british library – the ʻzetoc ̓experience. 2002. http: //epub.mimas.ac.uk/papers/macappslwow4.pdf 8. c. a. mellon. library anxiety: a grounded theory and its development. college and research libraries 47, 1986. pp160-165 http://epub.mimas.ac.uk/papers/macappslwow4.pdf http://epub.mimas.ac.uk/papers/macappslwow4.pdf 20 iassist quarterly winter 2002 name: job title: organization: address: city: state/province: postal code: country: phone: fax: e-mail: url: iassist international association for social science information service and technology association internationale pour les services et techniques d'information en sciences sociales i would like to become a member of iassist. please see my choice below: options for payment in canadian dollars and by major credit card are available. see the following web site for details: http://datalib.library.ualberta.ca/membership/ membership.html $50 (us) regular member $25 student member $75 subscription (payment must be made in us$) list me in the membership directory add me to the iassist listserv membership form the international association for social science information services and technology (iassist) is an international association of individuals who are engaged in the acquistion, processing, maintenance, and distribution of machine readable text and/or numeric social science data. the membership includes information system specialists, data base librarians or administrators, archivists, researchers, programmers, and managers. their range of interests encompases hard copy as well as machine readable data paid-up members enjoy voting rights and receive the iassist quarterly. they also benefit from reduced fees for attendance at regional and international conferences sponsored by iassist. membership fees are: regular membership: $50.00 per calendar year. student membership: $25.00 per calendar year. institutional subcriptions to the quarterly are available, but do not confer voting rights or other membership benefits. institutional subcription: $75.00 per calendar year please make checks payable, in us funds, to iassist and mail to: iassist, assistant treasurer joann dionne 50360 warren road canton, mi 48187 usa 9. thomas r. murray and dale l. brubaker. theses and dissertations: a guide to planning, research and writing. westport beigin and gurvey. 2000. 10. diane nahl. creating user centred instructions for novice end-users. reference services review 27(3) 1999. pp280-286 11. jill newby. evolution of library research methods course for biology students. research strategies 17, 2000. pp57-62 12. barbara nelson. evaluating electronic resources. conference report/ library collections, acquisitions and technical service 24, 2000. pp 403-441 13. d. nunan, d. research methods in language learning. cambridge: cambridge university press. 1992. in brian paltridge. thesis and dissertation writing: preparing esl students for research. english for specific purposes 16 (1) 14. steve oʼconnor. economic and intellectual value in existing and new paradigms of electronic scholarly communication. library hi tech 18 (1) 2000. pp37-45 15. anthony onwuegbuzie. writing a research proposal: the role of library anxiety, statistics anxiety, and composition anxiety. library and information science research. 19 (1) 1997. pp5-33 16. ruth pagell. reaching for the bottle, not the glass; the enduser factor of electronic full text. database 16 (5) 1993. 17. sarah porter. into the future. scholarly needs, current provision, and future directions. 2000. http: //ahds.ac.uk/public/uneeds/un5.html 18. f.l.rosenfeldt, how to write a paper for publication. heart, lung and circulation. 9, 2000. pp 82-87 19. l. stewart. helping students during online searches: an evaluation. journal of academic librarianship. 18(6) 1993. pp 347-351 20. g torkzadeh and i.e. angulo the concept and correlates of computer anxiety. behavior and information technology. 11, 1992. pp99-108 21. bob travica. organizational aspects of the virtual library: a survey of academic libraries. library and information science research, 21 (2) 1999. pp173-203 22. denise troll. how and why the libraries are changing: what we know and what we need to know. libraries and the academy 2.1, 2002. pp99-123 footnotes 1 thomas r. murray and dale l. brubaker. theses and dissertations: a guide to planning, research and writing. westport beigin and gurvey. ̓2000. p. 8 2 jude edminster and joe moxley. graduate education and the evolving genre of electronic theses and dissertations. computer and composition 116, 2002. p.8 3 f.l.rosenfeldt, how to write a paper for publication. heart, lung and circulation. 9, 2000. p. 83 4 a.j jerebek and linda meter. “library anxiety” and “computer anxiety”: measures, validity and research implications. library and information science research, 23 2001. p.278 5 l. stewart. helping students during online searches: an evaluation. journal of academic librarianship. 18(6) p348 6 diane nahl. creating user centred instructions for novice end-users. reference services review 27(3) 1999. p280 7 lois buttlar. information sources in library and information science doctoral research. library and information science research, 21(2) p.229 8 ross macintyre and ann apps. working with the british library – the ʻzetoc ̓experience. 2002. http: //epub.mimas.ac.uk/papers/macappslwow4.pdf 9 f.l.rosenfeldt, how to write a paper for publication. heart, lung and circulation. 9, 2000. p. 83 10 d. nunan, d. research methods in language learning. cambridge: cambridge university press. 1992. in brian paltridge. thesis and dissertation writing: preparing esl students for research. english for specific purposes 16 (1) p.63 http://datalib.library.ualberta.ca/membership/membership.html http://datalib.library.ualberta.ca/membership/membership.html http://ahds.ac.uk/public/uneeds/un5.html http://ahds.ac.uk/public/uneeds/un5.html http://epub.mimas.ac.uk/papers/macappslwow4.pdf http://epub.mimas.ac.uk/papers/macappslwow4.pdf 20 iassist quarterly winter 2002 name: job title: organization: address: city: state/province: postal code: country: phone: fax: e-mail: url: iassist international association for social science information service and technology association internationale pour les services et techniques d'information en sciences sociales i would like to become a member of iassist. please see my choice below: options for payment in canadian dollars and by major credit card are available. see the following web site for details: http://datalib.library.ualberta.ca/membership/ membership.html $50 (us) regular member $25 student member $75 subscription (payment must be made in us$) list me in the membership directory add me to the iassist listserv membership form the international association for social science information services and technology (iassist) is an international association of individuals who are engaged in the acquistion, processing, maintenance, and distribution of machine readable text and/or numeric social science data. the membership includes information system specialists, data base librarians or administrators, archivists, researchers, programmers, and managers. their range of interests encompases hard copy as well as machine readable data paid-up members enjoy voting rights and receive the iassist quarterly. they also benefit from reduced fees for attendance at regional and international conferences sponsored by iassist. membership fees are: regular membership: $50.00 per calendar year. student membership: $25.00 per calendar year. institutional subcriptions to the quarterly are available, but do not confer voting rights or other membership benefits. institutional subcription: $75.00 per calendar year please make checks payable, in us funds, to iassist and mail to: iassist, assistant treasurer joann dionne 50360 warren road canton, mi 48187 usa http://datalib.library.ualberta.ca/membership/membership.html http://datalib.library.ualberta.ca/membership/membership.html 1/15 gillman, d. w. and waring, clayton (2023) modernizing data management at the us bureau of labor statistics, iassist quarterly 47(1), pp. 1-15. doi: https://doi.org/10.29173/iq1038 modernizing data management at the us bureau of labor statistics daniel w. gillman and clayton waring1 abstract the us bureau of labor statistics (bls) is undertaking initiatives to improve its data and metadata systems. planning for the replacement of the public facing labstat data query system and efforts within the office of productivity and technology to combine multiple production systems within a single cross-divisional database platform are examples. bls views time series data as a combination of three elemental components found in every time series. a measure element; a person, places, and things element; and a time element are the components. the authors turned this basic approach into a formal conceptual model represented in uml (unified modeling language). the uml model describes a flexible multi-dimensional data structure, of which time series are a kind, and supports any kind of query into the data. the office of productivity and technology adopted the model, and it is guiding their approach moving forward. the financial industry business ontology project under the object management group and the data documentation initiative cross-domain integration (ddicdi) development project also adopted the model. in this paper we describe the time series formulation and the uml conceptual model. then, the design of the opt system and its features are described. in doing so, we provide a thorough understanding of the structure of time series. keywords multi-dimensional data, time series, metadata, measures introduction the us bureau of labor statistics (bls) is undertaking initiatives to improve the way it manages its data and metadata systems. two examples include planning for the replacement of its public facing labstat data query system and efforts within its office of productivity and technology to combine multiple production systems within a single cross-divisional database platform. within these projects, bls views time series data as a combination of three mutually exclusive elemental components, which are found in every time series. they include a measure element; a person, place, and thing element; and a time element. the authors turned this basic approach into a more formal conceptual model represented in the unified modeling language under the object management group2 (omg) [omg, 2017a]. the uml model describes multi-dimensional data, of which time series are a kind, and is very flexible in that it supports any kind of query into the data. the office of productivity adopted the model, and it is guiding their approach moving forward. the omg financial industry business ontology – indices and indicators (fibo) [omg, 2017b] standard and the data documentation initiative (ddi) cross-domain integration (ddi-cdi) development project under the ddi alliance3 [ddi alliance, 2020b] also adopted the model. https://doi.org/do.be/doo 2/15 gillman, d. w. and waring, clayton (2023) modernizing data management at the us bureau of labor statistics, iassist quarterly 47(1), pp. 1-15. doi: https://doi.org/10.29173/iq1038 in this paper we describe the elemental components of time series and the uml conceptual model derived from them. then, the design of the opt system and its features are described. in doing so, we provide a thorough understanding of the structure of time series. history there have been many attempts to describe multi-dimensional data over the years. in all these efforts, the structure of the data is elemental. it defines the semantics (i.e., the meaning) of the data and the end points for beginning a query. in his phd dissertation in 1973, sundgren [sundgren, 1973] developed a mathematical treatment of multi-dimensional structures called boxes. a box has dimensions and a measure, and the cells defined by the elements of the dimensions store the values. additivity of cells, used to aggregate data and reduce the number of dimensions, is inherent in the approach. the main contribution is the idea that a table is just a presentation format for boxes. after the development of the relational model by codd [codd, 1970], data warehouses and the online analytical processing (olap) model followed, pioneered by e. codd, s. codd, and salley [codd, et al., 1993]. olap evolved as a counterpart to online transaction processing (oltp), the method for handling business transactions, to aid business analyses. as with sundgren, the focus of the olap approach was query, not semantics. the storm (statistical object representation model) model for describing multi-dimensional data in was introduced in 1990 [rafanelli and shoshani, 1990]. this approach recognized that a table presented an arbitrary ordering of the dimensions. the storm system provided the ability to manipulate the dimensions so a new table with a new ordering of the dimensions could be produced. it relied on additivity to facilitate the necessary transformations. in 2000, van bracht [van bracht, et al., 2000, and van bracht, 2004] resurrected sundgren’s box model and referred to the structure as an n-cube. these n-cubes include a measure associated with each of the cells. this paper cemented the distinction between n-cubes and tables in the statistical community. additivity of data and projections (i.e., aggregations and eliminating dimensions) were an important focus. the first, simple, standard in the ddi suite is ddi-codebook, first released in 2000. as of this writing, the current version is 2.5. [ddi alliance, 2012]. this standard contains a mechanical format description of a table. it includes the dimensions and measures in the description of each table. the statistical data and metadata exchange (sdmx) specification was released [sdmx, 2004] in 2004, and version 3.0 of this standard was released in 2021 [sdmx, 2021]. sdmx explicitly implements the idea of an n-cube. it is possible to treat time itself as a dimension, and sdmx incorporates the ability to include values from more than one measure in each cell. the focus of sdmx is the exchange of multi-dimensional data with metadata attached, mostly the dimensions. the world wide web consortium (w3c)4 issued their data cube vocabulary [w3c, 2010], a version of the sdmx data model rendered in rdf. the standard in the ddi suite based on the statistical lifecycle is ddi-lifecycle (ddi-l), first released in 2008. the latest is version 3.3, released in 2020 [ddi alliance, 2020a]. this includes the n-cube idea https://doi.org/do.be/doo 3/15 gillman, d. w. and waring, clayton (2023) modernizing data management at the us bureau of labor statistics, iassist quarterly 47(1), pp. 1-15. doi: https://doi.org/10.29173/iq1038 and has similar limitations and restrictions as in sdmx. the focus on the statistical lifecycle distinguishes ddi-l from sdmx. ddi-l also supports metadata reuse. in 2012 the unece released the generic statistical information model (gsim). the latest version, released in 2019, is version 1.2 [unece, 2019]. gsim is a conceptual model of the information needed to describe statistical data production, and it includes multi-dimensional data as n-cubes. it does not explicitly include time series. the treatment of n-cubes is like that found in sdmx and ddi-l. the new ddi cross-domain integration (ddi-cdi) standard, for release in early 2023, includes a full description of n-cubes [ddi alliance, 2020b]. multi-dimensional data, of which n-cubes and time series are kinds, are among the four major data structure types that ddi-cdi describes. the model presented in this paper is the basis for that descriptive capability. the omg fibo standard included the model presented here as well. background our problem is to find a simple and intuitive way to store and organize statistical data with the goal of making it easy to find and use the data. the challenge is to reduce the overall complexity of the problem by uncovering underlying principles from which to build. the semantic approach we propose contains these principles and supports storing, organizing, and querying data based on their meaning, not their structure. these principles are simple and lead to a natural design. we avoid complexity and rigidity by building from the principles. einstein is famously attributed as to having said, “everything should be made as simple as possible, but no simpler.” at bls, the labstat system is the main dissemination tool for bls time series, and, paradoxically, it is both too simple and complex. it is too simple because it does not support a wide variety of queries (functionality). for example, it is not possible to find all data bls has about nurses, hospitals, or north carolina through a simple query. but labstat is also too complex in that its underlying specialized design is based on diverse, siloed principles that do not easily scale or support flexibility. this complexity results in a rigid design that is difficult to improve. to advance simplicity and increase functionality, we begin with the notion that storing and organizing things (e.g., data) is a simple and intuitive concept. people in their routine lives store and organize dishes in kitchen cabinets, clothes in bedroom dressers, books on shelves, and cars in garages. fundamentally, storing and organizing statistical data should not be any different than these examples. one thing cabinets, dressers, bookshelves, and garages all have in common is that they are designed specifically to accommodate the common definitional attributes and characteristics that define dishes, clothes, books, and cars. in other words, the method of storage and organization depends on the items being stored and organized. so, a first step in storing and organizing data is to understand the common definitional attributes and characteristics of data. a time series is a set of numeric values that represents the change over time of measurable aspects of people, places, and things. under the approach presented here, statistical data is stored and organized around the elemental components: measures; people, places, and things; and time. the first elemental component is people, places, and things. this refers to groups that share a set of commonalities, such as nurses in north carolina, shirts sold in pittsburgh, u.s. manufacturing capital equipment, men in new england, heads of household over the age of 40, etc. the next component is https://doi.org/do.be/doo 4/15 gillman, d. w. and waring, clayton (2023) modernizing data management at the us bureau of labor statistics, iassist quarterly 47(1), pp. 1-15. doi: https://doi.org/10.29173/iq1038 measures, which generally refer to quantitative, or numeric, variables. these include things like consumer price index, unemployment rate, value of gdp, crop yields, etc. the time component is simply the time stamp that identifies the applicable time reference for the observation. these three common elemental components connect all statistical data through a shared understanding of the meaning of data. for example, the price of shirts in pittsburgh in the first quarter of 2016 and the unemployment rate of men in new england in 2020 both include measure; people, places, and things; and time components. so, one advantage of this approach is it allows seemingly diverse statistical data to be stored, organized, and queried within a single, unified framework. semantic approach rather than taking the usual structural approach to describing multi-dimensional cubes, time series, and their associated measures, we adopted a semantic one based upon the “measures / peopleplaces-things / time” model. by this we mean the focus is on the meaning of the data rather than on some predefined structure. the structure follows naturally from the semantics. people-places-things consider the description “nurses working in hospitals in north carolina”. this is an example of a ppt (people-places-things) description. technically, we call this a universe. it can be used to describe the interpretation of a cell in an n-cube or the focus of a time series without knowing anything about a particular measure. all statistical data have a geographic location to which data apply. in our example it is the state of north carolina. the physical location of hospitals as a place of employment necessitates the geographic component. this is a consideration in all cases. sometimes the area is just implied. the primary component of a universe is called a unit type. this addresses the basic units of analysis in the data. in other words, the unit type comprises the things in the world the data describe. unit types may be persons, households, business establishments, animals, commodities, economic constructs (output, labor, capital, and materials), or events (e.g., marriage, education, or hospitalization), among others. some unit types can be usefully partitioned. a business establishment can be partitioned into output, labor, capital, and materials. a household can be partitioned into liabilities and assets, and assets can be further partitioned into financial and material. these act as unit types themselves. unit types are specialized into universes by applying the other categories in a ppt description. however, starting with the description of a universe, the unit type is not necessarily implied. our example, “nurses working in hospitals in north carolina” is a case in point. there are at least two ways to unpack and understand this description, because different unit types might apply. starting with “nurses”, nursing by itself is an occupation, so “nurses” is a combination of people (a unit type) and an occupation. this is one way to interpret the example phrase. further, in this case, we only care about those nurses working in hospitals from the nurses’ perspective, as nurses can work in many other settings. on the other hand, we could start with hospitals, an industry, and recognize that hospitals are a kind of business establishment (another unit type). in this interpretation, we care about the nurses working https://doi.org/do.be/doo 5/15 gillman, d. w. and waring, clayton (2023) modernizing data management at the us bureau of labor statistics, iassist quarterly 47(1), pp. 1-15. doi: https://doi.org/10.29173/iq1038 in hospitals from the hospitals’ perspective. so, in this case, “nurses” refers to a category of employee in hospitals (along with others), and hospitals are business establishments specialized to an identified industry. in both perspectives we have a different unit type – people for the nurse perspective, and business establishments for the hospital perspective. therefore, we have 2 potential interpretations of the universe: 1. people working as nurses in hospitals in north carolina. 2. hospitals (which are business establishments) employing nurses in north carolina. the choice comes down to what the basic unit of analysis is designed to be, and this follows from the design and meaning of the underlying data. in each case, we have unpacked the description to identify the unit type and the categories used to specialize that unit type into a universe. even though the words used to name the universe in the two interpretations are similar, the underlying meaning is different, because the starting point, the unit type, is different. measures measures represent the basic notions of size and scale. therefore, measures can be thought of as pairing concepts of size or scale to numeric values. for example, the size of a house may be equal to 2000 ft2 in january 2022. this is subject to direct comparisons. for example, the number of square feet can used to compare with the size of a different house. alternatively, the square footage of a house can be compared over a time interval to see if the size is increasing or decreasing. the numeric values assigned to a measure are called measurements. so, the idea of the size of a house is a measure while the number of square feet of a particular house is a measurement. and, within this context, certain rules must be adhered to. for example, the size of a house cannot be quantified using measurements such as the miles per hour or degrees of temperature. size and scale may also be associated with economic concepts such as prices, employment, sales, wages, and others. for example, the price of a gallon of milk may be equal to $4.50 usd. separating the measure from the numeric measurement is useful for several reasons. first, it permits the description of the measure to be independent of the description of the measurement. for example, any number of descriptors can be added to a measure, such as price. the price can be described as a consumer price excluding sales promotions and prior to the addition of applicable sales taxes. similarly, other descriptors can be added to the numeric measurement. as above, $4.50 can be described in units of current u.s. dollars. another reason for separating the measure from the measurement is their relationship to time. the concept of a measure tends to be reliably steady; however, the numeric measurement may change over time. for example, the concept of a price remains the same, however, the actual price of an item, say a gallon of milk, changes quite frequently. consequently, a measurement is associated with a particular observation at a point or interval in time. in other words, the measure refers to a definition associated with something like price, whereas the measurement refers to a specific price observed at a known time. https://doi.org/do.be/doo 6/15 gillman, d. w. and waring, clayton (2023) modernizing data management at the us bureau of labor statistics, iassist quarterly 47(1), pp. 1-15. doi: https://doi.org/10.29173/iq1038 the last reason for separating the measure from the measurement is it reveals how multiple measurements can be assigned to the same measure for one reference time. for example, the price of an item on some date can be given in usd, pesos, or euros. within the model items are grouped by their similarities and separated by their differences. so, two measures of sales have similarities, but there may be differences that lead to different values for the observations. a measure of sales could be gross sales or value-added sales. the measures might also differ by the survey sample, the survey frequency, calculation methodology, or any adjustments that are made to the results. in these cases, the description of the measure is used to differentiate the understanding of outcomes for the purpose of comparisons. the proper description of a measure ensures the correct understanding of the measurements, reducing the risk of misinterpretation. time time, as might be expected, refers to our innate understanding of the sequence of events we observe in our daily lives. it is referred to in standardized units such as year, months, and days to enhance communication. further, time is referenced as a point in time or as an interval of time (with begin and end points). measures can refer to either, whereas time series refers to intervals. within our model time plays several important roles. first, time is used to compare measurements. measurements of wages, prices, and employment may fluctuate over time, and this notion of change over time is what characterizes measures. changes of values of numeric measurements over time are what economic analyses investigate. measures may have attribute qualifiers assigned to them, such as seasonally adjusted, but these attributes do not change over time. in other words, when seasonally adjusted employment increases, it is understood that employment is the concept that is increasing. in this way, ‘seasonally adjusted’ is the type or kind of employment that is increasing. as such, seasonal adjustment is not considered to be a measure in and by itself. so, one of the things that differentiates measures and the attributes assigned to them is their relationship to time. second, time is used to classify measures based how they relate to the passing of time. this is perhaps best illustrated by considering what would happen if time were to be held still as in a science fiction movie. with time held still certain measures would continue to exist, such as the number of employees at a firm or the value of the firm inventories. however, other measures such as the number of average weekly hours worked at the firm would cease to exist. it is impossible to work for an hour while time is held still. measures that exist at an instance of time are broadly referred to as stock measures while measures that require the passing of time are referred to as flow measures. however, measures of central tendency and dispersion, such as averages and standard deviation, apply to both stock and flow measures. these represent yet another class or type of measure. the third role time plays in the model is identifying the time stamp for the measure observations. the time stamp identifies the sequence of the observations. furthermore, the time stamps determine if the observations have regular intervals, with equal time intervals between the observations, or not. however, the meaning or interpretation of the timestamp is slightly different for stock and flow measures. a timestamp of 2021 represents an interval, typically from january 1st to december 31st. interval timestamps naturally fit with flow measure such as work hours and value of shipments. the measurement of value of shipments for 2021 would presumably be the sum of all shipments from the beginning to the end 2021. however, for stock measures, interval timestamps can be problematic. for example, consider the measurement of employment at a firm for 2021. employment cannot be https://doi.org/do.be/doo 7/15 gillman, d. w. and waring, clayton (2023) modernizing data management at the us bureau of labor statistics, iassist quarterly 47(1), pp. 1-15. doi: https://doi.org/10.29173/iq1038 measured as the sum of all employment from the beginning to the end of the year. more likely, the employment measurement may be some type of average over the year or the employment level on a specific day of the year. in each case additional information beyond a simple timestamp may be needed to provide clarity to the meaning of the measure and its relationship to time. uml model uml is a widely used modeling language. it is built for object-oriented analysis and design (ooad), a systems design paradigm. within that framework, several modeling techniques are defined and used, including class diagram, which show how entities (called classes in ooad) are described and related. class diagrams depict the entities, attributes, and relationships in a visual framework. for the uninitiated, there exist many online tutorials on how to construct and read uml class diagrams5. uml class diagrams are useful outside the specifics of ooad, and that is our approach in this paper. based on specificity, there are three levels of class diagram models for systems development:  conceptual  logical  physical these levels are explained in simple language in tutorials on the web6. in this paper, we are interested in the least specific of these, the conceptual model. a conceptual model is designed for human communication. it is used to communicate the essential requirements for some system, and these correspond to the entities, attributes, and relationships in the model. a system built to satisfy those essential requirements is said to conform to the conceptual model. logical models are used to establish and further refine the requirements for a system. the physical model contains all the requirements in a logical model and corresponds to the schema for an existing system. in this section, we describe the conceptual uml model that follows from the ppt, measures, and time components in our semantic approach. the model is presented in four class diagrams, figures 1-4. the colors in the boxes depicting classes indicate which of the components are addressed with each class: dark blue for ppt, light blue for measures, and brown for time and multi-dimensional structures. ppt model for people, places and things as described above, there are three main concerns in the construction of our conceptual model (for simplicity, “model”): ppt, measures, and time. in this section, we describe how ppts are modeled and why the model is structured as it is. the ppt construct translates into a universe in the bls model. a universe is all the specialized units that apply to some data. our previous example of nurses working in hospitals in north carolina is typical. the first consideration is the fundamental units we are measuring, e.g., people, households, establishments. for nurses working in hospitals in north carolina, placing nurses first in this formulation makes us think of people. if we had said hospitals employing nurses in north carolina, our https://doi.org/do.be/doo 8/15 gillman, d. w. and waring, clayton (2023) modernizing data management at the us bureau of labor statistics, iassist quarterly 47(1), pp. 1-15. doi: https://doi.org/10.29173/iq1038 focus shifts to hospitals. the main point is that word order really does not matter. let us assume we mean people as the fundamental units, which we refer to as unittype. nurse, hospital, and north carolina are each a category in our model. the set of categories from which nurse, hospital, or north carolina is drawn is a dimension. for example, the north american industry classification system7 (naics) is divided into levels, each expressing different levels of detail. hospital is the category labelled 622 in the subsector level of naics. so, the dimension in this case is a list of subsectors, with hospital listed as a category within. there are similar considerations for nurse in the standard occupational classification8 (soc) and north carolina in a list of us states such as the us postal state codes. the soc is a hierarchical listing of occupations, naics is a hierarchical listing of industries, and postal codes enumerate the us states. each of these general considerations – occupation, industry, and states of the us – is a characteristic. each category specializes the unittype people, for instance people residing in north carolina, people employed by hospitals, or people working as nurses. the combination of the three categories further specializes people in this case. the result is a universe. it is a specialization of a unittype. some unit types can be meaningfully broken into parts, and this is how the unittypepartition is used. each element of a partition can be a unittype itself, a facetedunittype. for example, establishments can be subdivided into output, labor, capital, and material. households can be subdivided into liabilities and assets. each could be used as a unit type. a combination of categories, each from one of the dimensions corresponding to each one of the characteristics, all applied to a unittype, forms a universe, which specifies a subset of the unittype. a scopedmeasure is the result of the association of a universe with a qualifiedmeasure, which are discussed in the measures section below. a universe describes the units (or objects) to which a scopedmeasure applies. all these ideas just discussed are illustrated in figure 1 below. measures model previously, we said measures and measurements are different, and among measures there are similarities and differences. so, we treat measures and measurements separately, and we specify several layers for measures. see figure 2 below to view the bls model for measures, which we describe in the following. in our formulation, a measure is a quantitative variable. it has a topic, or concept, associated with it, that groups broadly similar measures, but this is outside the scope of our model. one such concept might be wages, which can be specialized into various measures, such as weekly wages, hourly wages, average weekly wages, median weekly wages, etc. https://doi.org/do.be/doo 9/15 gillman, d. w. and waring, clayton (2023) modernizing data management at the us bureau of labor statistics, iassist quarterly 47(1), pp. 1-15. doi: https://doi.org/10.29173/iq1038 dimensioncategory 1..n 0..n restriction 2..n partitive 0..n 1..n element charactristic 1 1..n kind measure 1 0..n circumscribe refinerestrictscopedmeasure 1..n 0..n subject 1..n 1..n dimension unittypepartition 1..n 1 facets 1 0..n facetfacetedunittype 1 0..nsubject allowable unittype universe qualifiedmeasure a category that is a facet of a facetedunittype must be an element of the unittypepartition associated with the facetedunittype see measures and n-cubes pages for further information figure 1: ppt model a measure has a datatypefamily associated with it. this is a broad characterization of what one can do with the associated data. measures are quantitative variables, and the main quantitative types are interval and ratio. rates, percentages, and totals are typically ratio data. indexes are interval. the difference is whether a value of 0 means absence of the quantity: ratio is yes; interval is no. the unitofmeasure is the specific way quantities are named, such as dollars for expenditures or wages. there are many ways a measure can be made more specific. for instance, “average weekly wages” is produced in at least 2 ways at the bureau of labor statistics: 1) from the current employment statistics9 (ces) survey estimates of average hourly wages and weekly hours worked; and 2) from the quarterly census of employment and wages10 (qcew) estimate of average quarterly wages divided by 13. both measures fall under the same grouping: average weekly wages. we account for these differences by productionmethod. disclosureavoidance or adjustment, such as seasonal adjustment, are additional ways in which a measure can be differentiated. the resulting specialization of a measure is called a qualifiedmeasure. here, the dimensions associated with the generic measure are linked. for example, the set of sectors in naics is a specific instance of the characteristic industry. the combination of a universe and qualifiedmeasure is a scopedmeasure. this is the specialization of measure that we associate data with, and this done through measurement. each scopedmeasure may have many measurements. here, the universe tells us the relevant set of things (units or objects) in the world to which the data apply. a measurement has a datum and a timestamp (usually a reference date and a release date). sometimes, the data associated with a scopedmeasure for some reference date is revised (e.g., the total jobs added to the us economy in february 2021 was revised upward from 468,000 to 536,000 in june). for this reason, more than one datum may be associated with a single reference date. revision provides the reasons for why a number might be changed. the attribute vintage in the class datum is incremented by one for each new revision. this number allows researchers to put together series of data based on which revision is relevant for investigation. https://doi.org/do.be/doo 10/15 gillman, d. w. and waring, clayton (2023) modernizing data management at the us bureau of labor statistics, iassist quarterly 47(1), pp. 1-15. doi: https://doi.org/10.29173/iq1038 dimensioncategory 1..n 0..n restriction 0..n 1..n element characteristic 1 0..n kind measurerefinerestrict scoped measure 1..n 0..n subject 1..n 1..n dimension 0..1 0..n modify 0..n 0..n adjust 1 1..n subject area 0..1 0..n protect 0..1 0..n 0..n 1 intended unitof measure 0..n 1 characterization 0..1 0..n quantify revision 0..n 0..1 revisedatum timestamp 0..n 1 release 1 1..n produces 1 0..n reference 1 0..n generate measurement disclosure avoidance programarea production method qualified measure universe adjustment datatype family 1 0..n circumscribe figure 2: measures model multi-dimensional structures there are 2 ways to use a scopedmeasure: to express how measurements change over time in a timeseries; or as an expression of related measurements (with other scopedmeasures under a single qualifiedmeasure) in an n-cube for a single reference period. the models for each follow below, and entities with the same name in these models and the ones above are the same entity. figure 3 depicts the structure of time series, and figure 4 depicts the structure for an n-cube. these models describe the full range of multi-dimensional data produced by bls, with one caveat. some data produced by bls are in the form of time series, such as those from the quarterly census of employment and wages (qcew), but some of them are not strictly time series. qcew does not redefine historical data to align with changes to naics. so qcew aggregates (e.g., county employment) are time series, but the detailed industry data are only time series in a short-term sense. in figure 4, we distinguish between the structure of an n-cube (a structuraln-cube) from an n-cube with data applied (an n-cube). the structure is just all the dimensions and their underlying categories. figure 3 includes an entity called measurementgroup. it allows us to lump more than one scopedmeasure into a timeseries. this violates the definition of a time series, but bls uses it to describe inter-related series more easily. an example is the combination of the us unemployment rate and the percentage point change from the previous reference date. sdmx contains this feature as well. https://doi.org/do.be/doo 11/15 gillman, d. w. and waring, clayton (2023) modernizing data management at the us bureau of labor statistics, iassist quarterly 47(1), pp. 1-15. doi: https://doi.org/10.29173/iq1038 1 0..n circumscribe restrict scoped measure datum timestamp 0..n 1 release 1 1..n produces 1 0..n reference 1 0..n generate timeseries measurement group 1 0..n reference 0..n 1..n group elements 0..n 2..n series elements 1..n 0..n represents qualified measure universe measurement in a measurement group, the refernce date must be the same as that of all of its member measurements. figure 3: time series model dimensioncategory 1..n 0..n restriction 0..n 1..n element 1 0..n circumscribe restrict scoped measure 1..n 1..n dimension datum 1 1..n produces 1 0..n generate cellvalue 1 0..1 selection 0..1 1 empty cell 1..n 1..n filled cell 1 0..n represent 1 0..n common measure 0..n 1..n cell universe qualified measure measurement n-cube structuraln-cube cube datapoint datapoint unittype subject 1..n 0..n dimensionalize dimensions are the discrete axes of a structuraln-cube. figure 4: n-cube model applications the bls office of productivity and technology (opt), unlike other program offices in bls, does not process survey data directly but rather collects data from several sources. these data sources include other bls programs and outside sources such the bureau of economic analysis (bea), the census bureau, and others. these put opt into a situation where it is both a data user acquiring processed data from other sources and a data disseminator producing statistical data for public consumption. the problem facing opt was a common one: most data sources use their own idiosyncratic method to organize data. this leads to a siloed approach that inhibits interoperability within a statistical production system. the measures / people-places-things / time approach alleviated much of the https://doi.org/do.be/doo 12/15 gillman, d. w. and waring, clayton (2023) modernizing data management at the us bureau of labor statistics, iassist quarterly 47(1), pp. 1-15. doi: https://doi.org/10.29173/iq1038 siloed data issue and allowed opt to begin combining previously incongruent datasets into a uniform structure for the purpose of storing, managing, and querying data. all the statistical data in the system share a common definition of data that is expressed by the unifying semantic model presented here. currently, opt is combining two cross-division productivity production systems into a single system. in the example in table 1 below, the bls 2020 index value of labor productivity for workers in north carolina demonstrates how the model is used to organize and structure metadata elements within a statistical production system. each time series is identified by the unique combination of the field values associated with the people, places and things element and the measure element. the data value within the time series is then uniquely identified by the addition of the time stamp and vintage fields. the system also allows for additional metadata elements. the choice of the fields and their associated field values will differ from one dataset to another. model field field value description sc o p ed m ea su re ( ti m e se ri es ) p eo p le , p la ce s, a n d t h in gs ( u n iv er se ) unittype establishmen t sets the unittype as establishment with added specificity of private nonfarm business establishment. unittypepartion set as labor. establishmenttype private nonfarm business unittypepartition labor industry 113_81 sets the category of industry to the characteristics of naics code 113_81. sets the category of geography to the characteristic north carolina geography north carolina m ea su re measurelabel labor productivity sets the principal measure to labor productivity. labor productivity is output per unit where unit is set to the unittypepartition of labor. measure output per unit measureprogramare a bls office of productivity sets various measure attributes which qualifies the measure. measureadjustment not seasonally adjusted measuremethodolog y value added https://doi.org/do.be/doo 13/15 gillman, d. w. and waring, clayton (2023) modernizing data management at the us bureau of labor statistics, iassist quarterly 47(1), pp. 1-15. doi: https://doi.org/10.29173/iq1038 model field field value description measuretype index measureunit 2007=100 measuredisclosure public access measurementdatum 110.823 the assigned value of the labor productivity index. ti m e ti m e timestamp 2020 sets the timestamp to 2020 and sets the vintage to the publication date of may 27, 2021, for the purpose of tracking revisions over time. vintage 5/27/2021 table 1: opt example as mentioned in the introduction, the model described herein was adopted by two standardization efforts, fibo and ddi-cdi. as of this writing, each of these efforts is near to publishing the first version of its standard. the ddi-cdi effort is expected to produce its first version in 2022. fibo is in use, but the specification is still officially in draft. fibo is a development project under the omg to build a web ontology language (owl) ontology [owl, 2012] for describing and sharing financial business data in the form of time series. time series are the main way data are organized and presented. data from banks and stock exchanges are the primary sources under consideration. the dow-jones industrial average is a typical case. data from statistical agencies were also in scope, so bls time series are also typical cases. ddi-cdi is the latest standard in the suite of ddi statistical metadata standards. it is designed to address the needs of describing and integrating data from multiple sources, especially from the statistical point of view. in support of the goal to integrate data from multiple sources, a user needs to be able to describe data in varied structural arrangements. one of these is for multi-dimensional data. the ddi-cdi framework includes the model described here for that purpose. both fibo and ddi-cdi adopted the entire model because the authors of those standards want to be able to describe any multi-dimensional data. the semantic underpinning, flexibility, and completeness of our model were inducements for these adoptions. conclusion this paper describes the efforts at bls to build a scalable and flexible model for storing and organizing multi-dimensional data. a measures – people/places/things – time (m-ppt-t) semantic model was described, which accounts for all bls time series data. the formal uml representation of the model follows directly from the requirements imposed by the m-ppt-t approach. it incorporates some details that are missing from previous work. details about universes are given cursory descriptions in other models. they are not seen as having primary importance. our model shows they are the semantic factor that distinguishes the meaning of https://doi.org/do.be/doo 14/15 gillman, d. w. and waring, clayton (2023) modernizing data management at the us bureau of labor statistics, iassist quarterly 47(1), pp. 1-15. doi: https://doi.org/10.29173/iq1038 cells in a structural n-cube or distinguishes the meaning of one time series from another. the categories from each dimension are combined into the semantics for the cell, which follows our semantic approach. the combination of unit type and the categories from all the dimensions represents the ppt component. an additional feature is we do not use the combination of categories as an identifier for the cell. identifying cells this way depends on the name used for each category, not the combination of meanings of the categories. since names can change over time and do across programs, identifiers may change. further, in our model, structural n-cubes are reusable, and once more than one n-cube is constructed from the same structural n-cube these cell identifiers may no longer be unique. for complex n-cubes, assigning a new variable for each cell is a management and semantic disambiguation nightmare. it is not hard to construct a useful n-cube with thousands of cells, which would require describing and managing thousands of variables with only slight semantic differences among them. instead, our approach is to specialize the use of a measure so that the semantics of each cell arises from the “intersection” of a universe and a measure (the qualified measure in our model). the combined semantics from the components describes the cell rather than using the cell as the point from which descriptions start. the need to manage fewer objects follows. finally, as discussed, the bls the office of productivity and technology implemented the model, and it is the set of requirements for the modernization of the time series database. two independent standardization efforts adopted the model: the fibo effort under the omg and the ddi-cdi effort under the ddi-alliance. at this writing, the expected release date for ddi-cdi is in early 2023. references codd, e. f. (1970). a relational model of data for large shared data banks. commun. acm 13 (6): 377-387. codd, e., s. codd, and c. salley. (1993). providing olap to user-analysts: an it mandate. e. f. codd and associates. vol.32. ddi-alliance. (2012) ddi-c codebook. ddi-c. https://ddialliance.org/specification/ddicodebook/2.5/ ddi-alliance. (2020a). ddi-l lifecycle. (2022, nov). https://ddialliance.org/specification/ddilifecycle/3.3/. ddi-alliance. (2020b). ddi-cdi cross-domain integration. public review draft. april 16, 2020). https://ddialliance.org/specification/ddi-cdi. omg. (2017a). unified modeling language. https://www.omg.org/spec/uml/2.5.1/about-uml. omg. (2017b) financial industry business ontology – indices and indicator – v1.0 (fibo). (2022, nov). https://www.omg.org/spec/edmc-fibo/ind/rafanelli, m., and a. shoshani (1990) storm: a statistical object representation model. proceedings of international conference on scientific and statistical database management (ssdbm), pp14-29. sdmx. (2004). statistical data and metadata exchange, v1.0 technical specification. (2022, nov). sdmx. https://sdmx.org/?page_id=18#package https://doi.org/do.be/doo https://ddialliance.org/specification/ddi-codebook/2.5/ https://ddialliance.org/specification/ddi-codebook/2.5/ https://ddialliance.org/specification/ddi-lifecycle/3.3/ https://ddialliance.org/specification/ddi-lifecycle/3.3/ https://ddialliance.org/specification/ddi-cdi https://www.omg.org/spec/uml/2.5.1/about-uml https://sdmx.org/?page_id=18#package 15/15 gillman, d. w. and waring, clayton (2023) modernizing data management at the us bureau of labor statistics, iassist quarterly 47(1), pp. 1-15. doi: https://doi.org/10.29173/iq1038 sdmx. (2021). statistical data and metadata exchange, v3.0 technical specification. (2022, nov). sdmx. https://sdmx.org/?s=sdmx+3.0+technical+specification sundgren, b. (1973). an infological approach to data bases. stockholm university and statistics sweden, urval no 7. unece. (2019). generic statistical information model. (2022, jan). gsim. unece supporting standards group. https://statswiki.unece.org/display/gsim/generic+statistical+information+model. van bracht, e., de jonge, e. & kaper, e. (2000) cristal data objects. an object model for cubic, raw, or intermediate statistical data. netherlands: statistics netherlands. van bracht, e. (2004). cristal a model for data and metadata. working paper no. 29, unece work session on statistical metadata (metis), geneva. february 2004. w3c. (1998). resource description framework (rdf). (2022, nov). https://www.w3.org/rdf. w3c. (2010). rdf data cube vocabulary. (2022, nov). https://www.w3.org/tr/vocab-data-cube/. w3c. (2012). web ontology language (owl). (2012, dec). https://www.w3.org/owl/ endnotes 1 daniel w. gillman (gillman.daniel@bls.gov) and clayton waring (waring.clayton@bls.gov); us bureau of labor statistics; 2 massachusetts ave, ne; washington, dc; 20212; usa 2 object management group https://www.omg.org/ 3 ddi alliance – https://ddialliance.org 4 world wide web consortium https://www.w3.org/ 5 uml class diagram tutorial https://creately.com/blog/diagrams/class-diagram-tutorial 6 conceptual, logical, and physical models: https://www.guru99.com/data-modelling-conceptuallogical.html 7 north american industry classification system https://www.census.gov/naics/ 8 standard occupational classification https://www.bls.gov/soc/ 9 current employment statistics https://www.bls.gov/ces/ 10 quarterly census of employment and wages https://www.bls.gov/cew/ https://doi.org/do.be/doo https://sdmx.org/?s=sdmx+3.0+technical+specification https://statswiki.unece.org/display/gsim/generic+statistical+information+model https://www.w3.org/rdf https://www.w3.org/tr/vocab-data-cube/ https://www.w3.org/owl/ mailto:gillman.daniel@bls.gov mailto:waring.clayton@bls.gov https://www.omg.org/ https://ddialliance.org/ https://www.w3.org/ https://creately.com/blog/diagrams/class-diagram-tutorial https://www.guru99.com/data-modelling-conceptual-logical.html https://www.guru99.com/data-modelling-conceptual-logical.html https://www.census.gov/naics/ https://www.bls.gov/soc/ https://www.bls.gov/ces/ https://www.bls.gov/cew/ microsoft word 49-4-lafferty-hessfinal.docx 1/19 lafferty-hess, sophia, krenzer, william, ariansen, jenny & darragh, jennifer (2025) assessing data management and sharing plans: the “state of play” at duke and opportunities for cross-campus collaborations, iassist quarterly 49(4), pp. 1-19. doi: https://doi.org/10.29173/iq1168 the creative commons-attribution-noncommercial license 4.0 international applies to all works published by iassist quarterly. authors will retain copyright of the work and full publishing rights. assessing data management and sharing plans: the “state of play” at duke and opportunities for cross-campus collaborations sophia lafferty-hess,1 william krenzer,2 jenny ariansen,3 jennifer darragh4 abstract over the past few years, the united states has implemented a second round of data management policies, exemplified by the 2023 nih data management and sharing policy and 2022 “nelson memo.” effectively supporting public access to data and a data sharing culture at an academic research institution requires collaboration across various research support staff and central offices as well as knowledge of the current practices of researchers. two research support groups at duke university, the university libraries (dul) and the office of scientific integrity (dosi), have forged a strong working relationship for supporting data management and sharing practices, including an active teams channel for communication, developing tools collaboratively, delivering trainings, and providing co-consults for data management. to more effectively understand “the state of play” at our institution, dul and dosi analyzed data management and sharing plans (dmsps) submitted to the national science foundation (nsf) in 2021. the project team used a modified version of the dart rubric (https://osf.io/qh6ad/) to score dmsps against required elements in key areas, including types of data; standards for data and metadata; access, sharing, and preservation; limitations on access, distribution, and reuse; and roles and responsibilities. in this paper we will present the key findings from the dmsp assessment project and discuss how, as data management specialists, we can use this information to plan for ongoing education, training, and resource development using a cross-campus collaboration model. keywords data management plans, data sharing, institutional collaboration, data management services, repositories introduction the value of shared data has been widely discussed from its impact on the development of the rapid responses to global health crises (moorthy et al., 2020) to the use of shared data in creating models and training for artificial intelligence (bell & shimron, 2024). open data sharing is evolving from an aspirational goal to a standard requirement by funders (nih, 2020) and publishers alike (naughton & kernohan, 2016). with these growing requirements, academic institutions are developing and refining services to support researchers’ needs to effectively manage their data, write data management and 2/19 lafferty-hess, sophia, krenzer, william, ariansen, jenny & darragh, jennifer (2025) assessing data management and sharing plans: the “state of play” at duke and opportunities for cross-campus collaborations, iassist quarterly 49(4), pp. 1-19. doi: https://doi.org/10.29173/iq1168 sharing plans for grant proposals, package data for sharing through online repositories, and address often complex ethical and legal dimensions of data sharing. there has been an evolution of federal funding policies regarding data management and sharing in the united states over the past 15 years. the national science foundation (nsf) was one of the first funding agencies to implement the requirement for a “data management plan” (dmp) as part of all grant proposals in 2011, which was quickly followed by the office of science and technology policy memo “increasing access to the results of federally funded scientific research” released in 2013 (pasek, 2017). since 2020, there has been a new wave of data management and sharing policies within the us. these new policies began with the release of the national institutes of health (nih) data management and sharing policy in 2021 (nih, 2020), which went into effect january 25th, 2023, and was followed by the 2022 office of science and technology policy memo “ensuring free, immediate, and equitable access to federally funded research” (e.g., nelson memo) (ostp, 2022). a noteworthy shift in the nih policy was the more serious implications for non-compliance, as the data management and sharing plan (dmsp) is now considered a term and condition of the award, as well as clearer expectations that an established repository is used when making data publicly available (nih, 2020). in addition, the us government has defined a set of desirable characteristics for data repositories centering on making data more findable, accessible, interoperable, reuseable (wilkinson et al., 2016; national science and technology council, 2022). the shift towards asserting the importance of more formal sharing of data within repositories is also exemplified by other federal policies shifting to using the language “data management and sharing plans” versus “data management plans” (nsf, 2023, pp. 15). moving forward, in this paper when referring to data management plans, we will refer to them as dmsps. these types of policies have been one of the motivators for the rise of data management services being available within academic libraries (tenopir, sandusky, allard & birch, 2014). because of this, numerous data professionals and researchers over the years have begun examining these data management plans. bishoff and johnston (2015) used an “opt-in” approach for creating their sample of plans at the university of minnesota. others have gained access to a fuller corpus of dmps at their institution via a particular funding agency (van loon, 2017). a study at purdue focused their research on examining plans created within their institutional data management planning tool (green, cairns, & white, 2019). while many of these projects were based at individual institutions, parham et al. (2016) developed a standard rubric to assess nsf plans across numerous us-based institutions as part of the dart project. while examining dmps isn’t necessarily a “new area of study”; getting a glimpse into researcher’s practices can be particularly valuable for building a shared understanding about the current “state of play” at an institution regarding data management practices. while many data management services over the years have been centered within academic libraries, other units on campuses, particularly research offices, are increasingly important partners, especially regarding the importance of compliance with funder policies around data management and sharing (reinhart, 2016). these kinds of partnerships between groups helps provide various forms of data management support and are paramount for avoiding “siloing” individual groups and encouraging a collaborative environment for supporting an institution’s research community as a whole. for the past 3/19 lafferty-hess, sophia, krenzer, william, ariansen, jenny & darragh, jennifer (2025) assessing data management and sharing plans: the “state of play” at duke and opportunities for cross-campus collaborations, iassist quarterly 49(4), pp. 1-19. doi: https://doi.org/10.29173/iq1168 eight years, duke university libraries (dul) and the duke office of scientific integrity (dosi), have forged a strong working relationship around dmsps, including an active teams channel for communication, developing tools collaboratively, delivering trainings, and providing co-consults for data management. in 2020, dul and dosi began a collaboration on a new project aimed at assessing dmsps from nsf. the project team included two staff members from dul and two from dosi. the two staff from the library both had master’s degrees in library and/or information sciences and had worked as data librarians or data management professionals for numerous years. whereas the staff from dosi both had research experience, one had managed a large lab and had an ms in neurobiology, and the other had a phd in psychology. the team viewed this project as (1) a way to gain a better understanding of current practice, particularly as duke was working on implementing a new institutional research data policy,5 and (2) as a way to prepare for the new nih data management and sharing policy that had come out. an analysis of the plans would also be a foundation for the development of a strategy for enhancing data management education and areas for continuing collaboration between the two groups that would harness the varying expertise, values, and ideals that motivate the groups. methods in hopes of getting a clear and complete understanding of the planned data management practices of our university's researchers, we aimed to review an entire funding year’s worth of dmsps. to get this, the dosi team exported from our internal grant tracking system a list of all “new” proposal submissions (therefore excluding supplements, non-competing renewals, and competing renewals) to the nsf that occurred between january 1st, 2021 and december 31st, 2021. while there were 250 total submissions in this period, we limited the sample to those that received funding in order to have a more manageable and meaningful dataset. this query resulted in 81 proposals for review. of these 81, 10 were excluded because the proposal was mislabeled as new when it was actually a supplement for a previously funded proposal. an additional 11 proposals were removed from the sample because we were unable to locate the proposals, and therefore could not get the dmsps from them. lastly, an additional two proposals were excluded from the sample because the dmsps associated with the funded projects were nearly non-existent and their inclusion in the sample would have skewed the data set. this gave us a final sample size of 58 nsf funded proposals with dmsps to review from the following colleges: art & sciences, engineering, and environmental studies. for the review process, the authors used an adapted version of the dart rubric (whitmire et al., 2016) developed to assess nsf dmsps. the project team adapted the rubric to create generalized sections that also aligned conceptually with the new nih data management and sharing policy to allow for a better general understanding of data management practices (see the appendix for the full rubric). the primary modifications included having one section of the rubric focused more explicitly on access, sharing, and preservation and another section focused on elements that would limit the access, distribution and reuse of data, as well as adding an oversight, roles, and responsibilities section. each of the four authors individually ranked each element of the rubric on a 3-point performance level scale (complete/detailed (2); addressed issue, but incomplete (1); did not address (0)). the dmsps were also given a final score of exemplary, satisfactory, or needs improvement (using the same 3-point scale). 4/19 lafferty-hess, sophia, krenzer, william, ariansen, jenny & darragh, jennifer (2025) assessing data management and sharing plans: the “state of play” at duke and opportunities for cross-campus collaborations, iassist quarterly 49(4), pp. 1-19. doi: https://doi.org/10.29173/iq1168 the reviews of the dmsps were done individually via a redcap form that was generated for this project. from there, across a total of 15 sessions spanning 13-months (february 2022 to march 2023), the authors met to compare the set of ratings for each dmsp. the goal of these meetings was to come to a consensus on the rating of the dmsp for each individual question. while the group sought consensus, no rater was required to change their score (more details about scores and the rate of change can be found in table 1). the comparison process only used a subset of questions from the full rubric that either applied to all the nsf directories or that was identified as having particular importance to data management resource and service assessment at duke. a subset was also used for the in-depth comparison process as this was highly resource intensive, and the project team’s main goal was to gain a better conception of data management practices at a general level to determine areas for service and education enhancements. of particular interest to the project team was also whether a plan explicitly listed the name of a data repository within the plan and the mechanism for data sharing, as sharing data within a data repository is being increasingly encouraged and recommended in funder and journal policies (nih, 2020; plos, 2014). plans could include more than one repository as well as more than one mechanism for data sharing. one project team member then went through a process of coding repository and data sharing questions; these codes were then verified by another team member for correctness. when a repository was named, each repository was assigned a category of either institutional, domain-specific, general, or other. regarding the mechanism for data sharing, the categories included data repository, project website, journal supplement, article publication, by request, code repository, preprint server, other, or na.6 results over the course of 13 months, four raters reviewed 58 dmsps from funded nsf projects in 2021 on 10 adapted dart questions, for a total of 2,184 ratings given. during the comparison meetings, the group of raters in sum changed 13 of the scores to n/a, leading to a final count of 2,171 ratings given. with that information, a fleiss’ kappa test was run to test the level of agreement amongst the four raters in assigning the 3-point categorical level labels (fleiss, 1971). overall, the kappa statistics showed a substantial agreement amongst the raters when reviewing these nsf dmsps, κ = 0.71, p < .001 (landis & koch, 1977). while table 1 gives a breakdown of the kappa statistics for each question, we see that the raters had almost perfect to substantial agreement respectively with their ratings for questions “if the data are deemed to be of a "sensitive" nature, the dmsp describes what protections will be put into place to protect privacy or confidentiality” (κ = 0.86, p < .001), “provides details on when the data will be made publicly available” (κ = 0.81, p < .001), and “identifies metadata standards that will be used for the project” (κ = 0.80, p < .001). it is important to note that through the consensus meetings with the raters, approximately 12% of the ratings were changed after the meetings, which would increase our level of agreement. however, given that these meetings were used to discuss our scoring reasoning, and with the fact that no one was required to change their score, we do not see an issue with this. table 1: kappa score 5/19 lafferty-hess, sophia, krenzer, william, ariansen, jenny & darragh, jennifer (2025) assessing data management and sharing plans: the “state of play” at duke and opportunities for cross-campus collaborations, iassist quarterly 49(4), pp. 1-19. doi: https://doi.org/10.29173/iq1168 total ratings total changed % of change fleiss’ kappa p-value overall 2,171* 290 12.23 0.719 < .0001 describes what types of data will be captured, created, and collected 232 6 2.59 0.487 < .0001 identifies metadata standards that will be used for the project 232 35 15.09 0.800 < .0001 describes data formats created or used during the project 232 29 12.50 0.634 < .0001 provides details on when the data will be made publicly available 232 64 27.59 0.811 < .0001 describes how the data will be made publicly available 232 13 5.60 0.300 < .0001 describes plans for archiving and preserving digital data 232 26 11.21 0.534 < .0001 if the data are deemed to be of a "sensitive" nature, the dmp describes what protections will be put into place to protect privacy or confidentiality 232 26 11.21 0.865 < .0001 describes security measures that will be in place to protect the data from unauthorized access.** 94* 36 38.30 0.773 < .0001 if there are factors that limit the ability to share data (e.g., commercialization or proprietary nature of the data), the plan describes those conditions and describes to whom the data will be made available and under what conditions. 232 55 23.71 0.702 < .0001 final overall assessment of the dmp 232 0 0 0.366 < .0001 items are displayed by question order to keep thematic data management questions together and in the order they were answered. *the total ratings score does not include ratings that were switched in n/a while the change score does **question 8 on this list only appeared to reviewers that did not mark question 7 as n/a. 6/19 lafferty-hess, sophia, krenzer, william, ariansen, jenny & darragh, jennifer (2025) assessing data management and sharing plans: the “state of play” at duke and opportunities for cross-campus collaborations, iassist quarterly 49(4), pp. 1-19. doi: https://doi.org/10.29173/iq1168 review descriptions from there, we can look at the actual breakdown of how the questions were answered by the raters (for detailed breakdown, see table 2). starting with the overall assessment score, we see that the mode, the most common response, given by raters when reviewing the dmsps was ‘satisfactory’, making up 48% of the scores. the questions that mode response was ‘complete/detailed’ when reviewing dmsps were “describes what types of data will be captured, created, and collected” (71%), “describes how the data will be made publicly available” (53%), and “describes data formats created or used during the project” (50%). on the opposite end, the only question that the mode response was ‘did not address’ when reviewing dmsps was “identifies metadata standards that will be used for the project” (43%). taking all of this together, we can see that the raters found the dmsps to be very good at describing specific information about the data (i.e., types, formats, availability, etc.) but very poor at describing the metadata standards that would be used for the project. table 2: response counts complete / detailed (2) addressed issue but incomplete (1) did not address (0) n/a (3) mean describes what types of data will be captured, created, and collected 167 60 5 0 1.70 identifies metadata standards that will be used for the project 43 89 100 0 0.75 describes data formats created or used during the project 118 82 32 0 1.37 provides details on when the data will be made publicly available 41 109 82 0 0.82 describes how the data will be made publicly available 125 105 2 0 1.53 describes plans for archiving and preserving digital data 111 105 16 0 1.41 if the data are deemed to be of a "sensitive" nature, the dmp describes what protections will be put into place to protect privacy or confidentiality 16 24 9 183 1.14 7/19 lafferty-hess, sophia, krenzer, william, ariansen, jenny & darragh, jennifer (2025) assessing data management and sharing plans: the “state of play” at duke and opportunities for cross-campus collaborations, iassist quarterly 49(4), pp. 1-19. doi: https://doi.org/10.29173/iq1168 describes security measures that will be in place to protect the data from unauthorized access.* 15 24 8 36 1.15 if there are factors that limit the ability to share data (e.g., commercialization or proprietary nature of the data), the plan describes those conditions and describes to whom the data will be made available and under what conditions. 20 46 2 164 1.26 final overall assessment of the dmp 26 113 93 0 0.71 items are displayed by question order to keep thematic data management questions together and in the order they were answered. *question 8 on this regarding security measures list only appeared to reviewers that did not mark question 7 about the sensitive data as n/a. data sharing and categories of repositories lastly, we reviewed the ways in which dmsps discussed if, where, and how they would share their data. as a reminder, this data was created by two of the original raters (one did the initial coding of the repository information, while the second verified the information). additionally, any given dmsp could have named multiple mechanisms/repositories for how or where they would share their data, therefore there was not a 1-to-1 breakdown of dmsps reviewed and mechanisms/repositories mentioned. but, for our coding, a dmsp could not be coded for the same mechanism/repository multiple times. of the 58 dmsps that were reviewed, a total of 86 different mechanisms for how data would be shared were mentioned (see figure 1 for full breakdown), with sharing via a data repository as the most common mechanism, appearing in 22 of the dmsps. most surprising was that there were four dmsps that made no specific mention of how or where they would be sharing their data. from there we saw that of the 58 dmsps reviewed, 56% (33) of them mentioned a type of repository that they would use for sharing their data (figure 2). of these 33 dmsps, there were four types of repositories that were categorized by the coder: general, institutional, domain-specific, and other. as we can see from figure 3, domain-specific repositories were the most common type of repository, appearing in 57% of the dmsps, followed by other (39%), institutional (21%), and general (12%). lastly, we identified a total of 25 different repositories that dmsps mentioned for where they would share their data (figure 4). the most common repositories mentioned were: github (27%), duke research data repository (21%), sequence read archive (12%), genbank (12%), and arxiv.org (12%). the other remaining 20 repositories mentioned made up a total of 28 responses across the 33 dmsps that mentioned repositories (see the appendix for full list of repositories mentioned). figure 1. sharing mechanism identified in plans. plans could indicate more than one mechanism for sharing. 8/19 lafferty-hess, sophia, krenzer, william, ariansen, jenny & darragh, jennifer (2025) assessing data management and sharing plans: the “state of play” at duke and opportunities for cross-campus collaborations, iassist quarterly 49(4), pp. 1-19. doi: https://doi.org/10.29173/iq1168 9/19 lafferty-hess, sophia, krenzer, william, ariansen, jenny & darragh, jennifer (2025) assessing data management and sharing plans: the “state of play” at duke and opportunities for cross-campus collaborations, iassist quarterly 49(4), pp. 1-19. doi: https://doi.org/10.29173/iq1168 figure 2. count of plans that named a repository or repositories figure 3. categories of repositories listed in plans 10/19 lafferty-hess, sophia, krenzer, william, ariansen, jenny & darragh, jennifer (2025) assessing data management and sharing plans: the “state of play” at duke and opportunities for cross-campus collaborations, iassist quarterly 49(4), pp. 1-19. doi: https://doi.org/10.29173/iq1168 figure 4. count of repositories named in plans discussion assessing the content and quality of dm(s)ps is subjective while our kappa-level agreement overall showed significant agreement for plans, during the process of comparing coding across our key rubric elements, it became clear that while some metrics seem relatively straight forward, being able to parse the quality or presence of certain content within a dmsp is open to interpretation. the conversations the team had regarding why each of us scored an element the way we did illuminated both how plans require very close reading and re-reading to identify certain content and the variability in interpretation of certain terms. for instance, what do we mean by “metadata standards”? or what is the quality ranking for a plan that mentions that data will be shared as part of an article? during our re-coding process, the project team also added our own notes to the rubric to clarify some edge cases to aid in better consistency as we moved forward. while not necessarily surprising, this variability suggests that researchers may receive different advice/feedback when talking to different data management professionals. likewise, peer reviewers, in the case of nsf, and program officers, in the case of nih, that review the completeness of a plan may come to different conclusions regarding the quality of a plan. this variability also raises questions about how we might add more consistency into the dmsp review process. is additional training for plan reviewers needed to aid in better consistency and to build a shared understanding of what “makes a good plan”? could new ai tools assist in performing more machine-actionable assessments of dmsps? for instance, a recent project has been exploring the potential for developing machine11/19 lafferty-hess, sophia, krenzer, william, ariansen, jenny & darragh, jennifer (2025) assessing data management and sharing plans: the “state of play” at duke and opportunities for cross-campus collaborations, iassist quarterly 49(4), pp. 1-19. doi: https://doi.org/10.29173/iq1168 actionable dmsps more programmatically and this is an area that is worth watching closely as a community (praetzellis, gracy, & taylor, 2025). completeness in plans and areas for increasing education and training regarding the actual completeness of the answers to specific elements in a plan, the question where we found researchers had the most complete and detailed answers (mean 1.70) was related to the types of data they captured, created, and collected. whereas the question with the lowest completeness related to the plan identifying metadata standards (mean 0.75). this is not surprising that researchers generally understand the types of data they will be using in their project, whereas metadata standards are more obtuse or in some cases may not be available for a certain data type/discipline. however, this finding also suggests that more targeted education on what are metadata standards could aid more researchers in answering metadata related questions in dmsps. these forms of education regarding metadata standards may be included within general data management and sharing plan trainings or integrated into more concept specific trainings. for instance, dul has a responsible conduct of research training aimed at graduate students on the topic of data publishing that first demonstrates the use of metadata standards within several disciplinary repositories and then has participants find a repository based on a case study and rate the fairness of repositories asking them among other metrics to assess if and how a repository assigns standardized metadata. this approach goes beyond telling researchers what a metadata standard is and encourages them to engage with them in the real world. given the primary purpose of dmsps is increasingly being communicated in new regulatory documents (nih data management and sharing plan) as facilitating public access to research data, it was noteworthy that most plans did address in some way how data will be made publicly available and had the second highest mean for completeness (1.53). however, this is also an area where our team had the lowest percent of change in our scoring (5.6%) and the lowest kappa value (0.300). meaning this was a question where individual reviewers felt more strongly about their original score and were less likely to modify a score after our group conversation. this also suggests that team members interpreted different forms of making data publicly available (for instance using a data repository vs. making data available as a supplement within a journal article) with higher rates of variation. it is also noteworthy, although unsurprising given the scope of many nsf sponsored projects (versus what we might see for nih), that most plans did not have data that would be deemed sensitive or had factors that would limit the ability to share data. these questions were of particular interest to the project team as the risk factors to the institution when dealing with these forms of data increases substantially when sharing data. however, the lower percentage of plans that had complete answers to these questions, recommends this as an area for increasing outreach and resources to ensure that researchers are limiting exposure of sensitive data. finally the overall assessment scores for the dmsps (mean 0.71), suggests that more general education on writing quality dmsps will benefit the broader community. this is an area where the libraries and the office of scientific integrity have partnered extensively and invested significant time and resources for the implementation of the new nih data management and sharing policy. this has 12/19 lafferty-hess, sophia, krenzer, william, ariansen, jenny & darragh, jennifer (2025) assessing data management and sharing plans: the “state of play” at duke and opportunities for cross-campus collaborations, iassist quarterly 49(4), pp. 1-19. doi: https://doi.org/10.29173/iq1168 included offering a presentation to all nih funded departments prior to the implementation of the nih policy to raise awareness of dmsp review services, implementing new mechanisms to request support via our research portal myresearchhome, and providing a standard training every semester in partnership with our responsible conduct of research program for faculty and staff on “meeting data management and sharing requirements” as well as offering many sessions to individual departments, labs, and groups. within these trainings, while generally covering elements of dmsps and the new nih requirements, we have also integrated information on key terminology that we have seen more limited understanding from this research including metadata standards and what a repository is (and is not). dul and dosi have also been discussing how we might partner to reach nsf researchers when the new nsf public access policy goes into effect, albeit that timeline is more tenuous now with the current shifting political landscape within the united states. data sharing practices and use of repositories through the analysis of questions related to mechanisms for data sharing and the use of repositories, the project team found that over half of the plans (56%) did name the use of a “repository” for data sharing and the most common category of repository used were domain-specific repositories (60%). these findings suggest that some duke researchers are already well-versed in potential domainspecific repositories in their fields of study, which while not surprising is encouraging. however, it is also important to note that the project team construed “repository” broadly to also include code repositories (e.g., github) and preprints such as arxiv.org. these forms of “repositories” do not provide the necessary features to be considered “data repositories” that meet the desirable characteristics of data repositories or many fair principles. pointing to another area for increased outreach and education to help researchers understand the benefits and limitations of certain platforms and what a formal data repository can provide them. likewise, when considering the overall mechanism for data sharing, while “data repository” was the most commonly mentioned (22 plans), the next three highest mechanisms were “researcher website” (19 plans), “article publication” (11 plans), and “by request” (11 plans). these findings also point to the importance of continuing to educate on the value add of using a publicly accessible repository and that there is still much work to do to enhance a culture of public access to data within the united states. to further help duke researchers navigate the complexities of data sharing, dul and dosi developed a data sharing guidance page7 within myresearchpath that surfaces regulations, offices to consult, and duke supported repositories and resources. finally, the team members were also particularly interested in the prevalence of the use of the duke research data repository (rdr) within dmsps. while we were pleased to see that the rdr was named the second highest (7 plans); however, the fact that github was named the most (9 plans), suggests that additional outreach on the availability and features of our institutional data repository is an area to invest additional resources. dul and dosi have partnered over the last two years to extend this outreach including partnering with the associate vice president for research and innovation to provide presentations to departments and cross-promote the availability of the rdr in newsletters and training. 13/19 lafferty-hess, sophia, krenzer, william, ariansen, jenny & darragh, jennifer (2025) assessing data management and sharing plans: the “state of play” at duke and opportunities for cross-campus collaborations, iassist quarterly 49(4), pp. 1-19. doi: https://doi.org/10.29173/iq1168 the value of cross-unit collaborations this project not only helped the data management specialists at duke better understand researcher dmsp practices but also was a valuable mechanism for continuing to build a collaborative relationship between dul and dosi staff. working on a project with a shared goal allows units to learn from each other, build a shared understanding, and establish trust, thereby minimizing silos. this project also would not have been possible if undertaken solely by the library as the access to the dmsps was only possible because the dosi staff have permissions to access the duke grants system. the project team also benefited from the varying expertise and experiences that the cross-unit collaboration brought to bear with the library team members having deeper experience in general best practices in data management, while the dosi staff had more direct research experience. this allowed the team to have a theoretical/conceptual perspective as well as being grounded in real-world research experience. this can also provide more explanation to the different lenses different team members may have brought to their scoring and the value our conversations around scoring had on expanding our shared understanding of specific terms and practices. building on this project, this cross-functional team has engaged in ongoing collaborations including education initiatives, co-consultations, outreach for the duke research data repository (rdr), and developing a new workflow to assess the appropriateness of nih dmsps usage of the rdr as well as promote data sharing resources and guidance to funded projects. the relationships built through the nsf project and other initiatives will be maintained by continuing to share experiences and expertise and working together towards a shared goal to support duke researchers needing to comply with funder and journal data sharing mandates. conclusion as data management and sharing become increasingly a shared responsibility across various institutional research support offices, the importance of collaboration cannot be understated. at duke, the project to understand the “state of play” around data management by analyzing a set of nsf dmsps not only benefited our knowledge but also our relationships. while we did not approach this project with set hypotheses and viewed this research as more exploratory, the project has led us to identify areas for increased education and outreach, particularly in regard to metadata standards, key elements of dmsps and the interpretation of terminology, resources for the appropriate handling of sensitive data, the purpose and importance of data repositories, and the availability of our institutional data repository. the comparison process also highlighted the complexity inherent in coding qualitative dmsps in areas/topics that are open to interpretation. we have also found that the process itself helped the team come to a better shared understanding of various terms and practices. overall, this type of project can create a framework to conceptualize data management services within an institution as we work as a team to advance a data sharing culture at duke. references list bell, l. c., & shimron, e. (2024). sharing data is essential for the future of ai in medical imaging. radiology: artificial intelligence, 6(1), e230337. https://doi.org/10.1148/ryai.230337 14/19 lafferty-hess, sophia, krenzer, william, ariansen, jenny & darragh, jennifer (2025) assessing data management and sharing plans: the “state of play” at duke and opportunities for cross-campus collaborations, iassist quarterly 49(4), pp. 1-19. doi: https://doi.org/10.29173/iq1168 fleiss, j. l. (1971). measuring nominal scale agreement among many raters. psychological bulletin, 76(5), 378–382. https://doi.org/10.1037/h0031619 green, p., cairns, a., & white, h. (2019). evaluating data management plans—are they good and are they effective? proceedings of the iatul conferences. https://docs.lib.purdue.edu/iatul/2019/fair/3 lafferty-hess, s., krenzer, w., ariansen, j., darragh, j. (2025). data and code from: assessing data management and sharing plans: the "state of play" at duke and opportunities for cross-campus collaborations. duke research data repository. https://doi.org/10.7924/r4668q866 landis, j. r. & koch, g. g. (1977). the measurement of observer agreement for categorical data. biometrics, 33(1), 159–174. https://doi.org/10.2307/2529310 moorthy, v., henao restrepo, a. m., preziosi, m. p., & swaminathan, s. (2020). data sharing for novel coronavirus (covid-19). bulletin of the world health organization, 98(3), 150. https://doi.org/10.2471/blt.20.251561 national institutes of health (nih). (2020). final nih policy for data management and sharing. https://grants.nih.gov/grants/guide/notice-files/not-od-21-013.html national science and technology council. (2022). desirable characteristics of data repositories for federally funded research. https://bidenwhitehouse.archives.gov/wpcontent/uploads/2022/05/05-2022-desirable-characteristics-of-data-repositories.pdf national science foundation (2023). public access plan 2.0. https://nsf-govresources.nsf.gov/pubs/2023/nsf23104/nsf23104.pdf office of science and technology policy. (2022). ensuring free, immediate, and equitable access to federally funded research [memo]. https://bidenwhitehouse.archives.gov/wpcontent/uploads/2022/08/08-2022-ostp-public-access-memo.pdf naughton, l., & kernohan, d. (2016). making sense of journal research data policies. insights 29 (1): 84–89. http://doi.org/10.1629/uksg.284 parham, s. w., carlson, j., hswe, p., westra, b., & whitmire, a. (2016). using data management plans to explore variability in research data management practices across domains. international journal of digital curation, 11(1), 53–67. https://doi.org/10.2218/ijdc.v11i1.423 pasek, j. e. (2017). historical development and key issues of data management plan requirements for national science foundation grants: a review. issues in science and technology librarianship, 87. https://doi.org/10.29173/istl1709 plos one. (2014). data availability policy. https://journals.plos.org/plosone/s/data-availability 15/19 lafferty-hess, sophia, krenzer, william, ariansen, jenny & darragh, jennifer (2025) assessing data management and sharing plans: the “state of play” at duke and opportunities for cross-campus collaborations, iassist quarterly 49(4), pp. 1-19. doi: https://doi.org/10.29173/iq1168 praetzellis, m., grady, b., & taylor, s. (2025). piloting madmps for streamlined research data management workflows. international digital curation conference, the hague, netherlands. https://doi.org/10.5281/zenodo.14969693 rinehart, a. (2016). finding the connection: research data management and the office of research. bulletin of the association for information science and technology, 43(1), 28–30. https://doi.org/10.1002/bul2.2016.1720430107 tenopir, c., sandusky, r. j., allard, s., & birch, b. (2014). research data management services in academic research libraries and perceptions of librarians. library & information science research, 36(2), 84–90. https://doi.org/10.1016/j.lisr.2013.11.003 van loon, j. e., akers, k. g., hudson, c., & sarkozy, a. (2017). quality evaluation of data management plans at a research university. ifla journal, 43(1), 98–104. https://doi.org/10.1177/0340035216682041 whitmire, a. l., carlson, j., westra, b., hswe, p., & parham, s. (2016). the dart project: using data management plans as a research tool. open science framework. http://osf.io/kh2y6 wilkinson, m. d., dumontier, m., aalbersberg, ij. j., appleton, g., axton, m., baak, a., blomberg, n., boiten, j.-w., da silva santos, l. b., bourne, p. e., bouwman, j., brookes, a. j., clark, t., crosas, m., dillo, i., dumon, o., edmunds, s., evelo, c. t., finkers, r., … mons, b. (2016). the fair guiding principles for scientific data management and stewardship. scientific data, 3, 160018. https://doi.org/10.1038/sdata.2016.18 endnotes 1 sophia lafferty-hess is a senior research data management consultant at duke university libraries. she can be reached by email: sophia.lafferty.hess@duke.edu 2 william krenzer is a research project manager at duke office of scientific integrity. he can be reached by email: william.krenzer@duke.edu 3 jenny ariansen is the director of research integrity at duke university. they can be reached by email: ja267@duke.edu 4 jennifer darragh is a senior research data management consultant at duke university libraries. she can be reached by email: jennifer.darragh@duke.edu 5 https://policies.provost.duke.edu/docs/chapter-5-research-data 6 data and code are publicly available at: https://doi.org/10.7924/r4668q866 7 https://myresearchpath.duke.edu/topics/guidance-sharing-research-data 16/19 lafferty-hess, sophia, krenzer, william, ariansen, jenny & darragh, jennifer (2025) assessing data management and sharing plans: the “state of play” at duke and opportunities for cross-campus collaborations, iassist quarterly 49(4), pp. 1-19. doi: https://doi.org/10.29173/iq1168 appendix complete list of repositories coded 1: github 9 2: duke research data repository 7 3: genbank 4 4: ncbi sequence read archive 4 5: arxiv.org 4 6: bco-dmo 2 7: cuashi 2 8: dryad 2 9: lter data system 2 10: materialsmine 2 11: morphosource 2 12: ncbi geo 2 13: nsf pasta 2 14: biomodels 1 15: crawdad collection at ieee-dataport 1 16: cambridge structural database 1 17: dmref hybrid3 materials database 1 18: ecosis 1 19: figshare 1 20: lepbase 1 21: materials cloud 1 22: metabolights 1 23: ocb database 1 24: osf 1 25: streampulse 1 17/19 lafferty-hess, sophia, krenzer, william, ariansen, jenny & darragh, jennifer (2025) assessing data management and sharing plans: the “state of play” at duke and opportunities for cross-campus collaborations, iassist quarterly 49(4), pp. 1-19. doi: https://doi.org/10.29173/iq1168 scoresheet for assessment of duke data management plans the purpose of this scoresheet is to help duke research support staff to generally assess data management plans completeness and alignment with best practices for data sharing as a mechanism to more fully understand current research practices, compliance, areas for ongoing education, training, and resource development. this rubric was adapted from whitmire et al. (2017), which had an original purpose to assess nsf data management plans and has been generalized and harmonized with new nih requirements to allow for application across funding agencies. given that not all funders require the same elements for all plans, potential criteria that may be directorate specific within nsf or prioritized by new nih requirements are noted with an asterisk (*) or plus(+) respectively. only highlighted elements will be scored for the essential elements of dmps. any nas will not be counted when calculating the final score. section 1: types of data produced and tools/software used performance criteria complete/detailed (2) addressed issue, but incomplete (1) did not address (0) 1.1 describes what types of data will be captured, created, and collected 1.2* describes how data will be collected, captured, or created 1.3* identifies how much data (volume) will be produced 1.4+ describes what data will be shared from the project 1.5+ describes tools or software needed to access the data and how they will be made available section 2: standards for data and metadata performance criteria complete/detailed (2) addressed issue, but incomplete (1) did not address (0) 2.1 identifies metadata standards and/or metadata formats that will be used for the project 2.2 describes data formats created or used during project 2.3+ describes documentation that will be made available alongside the data section 3: access, sharing, and preservation 18/19 lafferty-hess, sophia, krenzer, william, ariansen, jenny & darragh, jennifer (2025) assessing data management and sharing plans: the “state of play” at duke and opportunities for cross-campus collaborations, iassist quarterly 49(4), pp. 1-19. doi: https://doi.org/10.29173/iq1168 performance criteria complete/detailed (2) addressed issue, but incomplete (1) did not address (0) 3.1 provides details on when the data will be made publicly available 3.2 describes how the data will be made publicly available 3.3+ provides the name of the repository where the data will be archived and made available (if 2 or 1 – drop-down to provide free-text for name of repository) name of repository listed in the dmp 3.4+ what is the mechanism indicated for sharing data? (free text) 3.5* describes how long the data will be retained and made available to people outside of the project removed from study instrument prior to analysis 3.6 describes plans for archiving and preserving digital data 3.7* describes plans for archiving and preserving physical data 3.8+ what is the mechanism indicated for archiving data? (free text) 3.9* identified timeframe for how long data will be archived? 3.10* describes the policies in place governing the use and reuse (i.e., redistribution, derivatives) of the data (i.e., relevant licenses, terms of use) 3.11* describes data types or formats that will be used for making data available section 4: limitations on access, distribution, and reuse performance criteria complete/detailed (2) addressed issue, but incomplete (1) did not address (0) 4.1 if the data are deemed to be of a "sensitive" nature, describes what protections will be put into place to protect privacy or confidentiality of research subjects 4.2 describes security measures that will be in place to protect the data from unauthorized access 4.3 if there are factors that limit the ability to share data, e.g. commercialization or proprietary nature of the data, plan describes those conditions and describes to whom the data will be made available and under what conditions 19/19 lafferty-hess, sophia, krenzer, william, ariansen, jenny & darragh, jennifer (2025) assessing data management and sharing plans: the “state of play” at duke and opportunities for cross-campus collaborations, iassist quarterly 49(4), pp. 1-19. doi: https://doi.org/10.29173/iq1168 4.4 describes what intellectual property rights to the data and supporting materials will be given to the public and which will be retained by project personnel (if any) 5. oversight, roles, and responsibilities performance criteria complete/detailed (2) addressed issue, but incomplete (1) did not address (0) 5.1 describes the roles and responsibilities of parties involved in the project 5.2+ describes how oversight and compliance of the plan with be monitored and managed final assessment total score: final assessment: exemplary ( ) satisfactory ( ) needs improvement ( ) notes: this scoresheet was adapted from: amanda l. whitmire, jake carlson, patricia m. hswe, susan wells parham, elizabeth rolando, & brian westra version: 1.0 | date: 23 march 2017 | contact: amanda whitmire at thalassa@stanford.edu this scoresheet is intended to be used with the nsf dmp rubric developed as part of this project. see whitmire, a. l., carlson, j., westra, b., hswe, p., & parham, s. (2017, march 24). rubric & related files. retrieved from http://osf.io/qh6ad *indicates that the performance criterion is nsf directorate or division-specific + indicates new element added to rubric to map to nih elements or other questions of interest to duke microsoft word iq49_2_editors_notes_final.docx 1/2 schwartz, ofira & hayslett, michele (2025) tides of transformation: enhancing research through data and community, iassist quarterly 49(2), pp. 1-2. doi: https://doi.org/10.29173/iq1170 the creative commons-attribution-noncommercial license 4.0 international applies to all works published by iassist quarterly. authors will retain copyright of the work and full publishing rights. editors’ notes: tides of transformation: enhancing research through data and community dear iassisters, welcome to iassist quarterly vol. 49 no. 2. it was wonderful to meet so many colleagues and to make new connections at the best iassist ever! in bristol, uk. it is always exciting to learn about the innovative work that is being done by members of this community. we would like to encourage those who presented at the conference to turn their conference presentation or poster into a paper and submit it to iq. this will allow you to share your research and expertise with a wider audience. we had several excellent submissions to this year's iassist conference paper competition. if you were not able to attend the conference, you may have missed the announcement about the 2025 competition’s winner. and the winner is: “assessing data management and sharing plans: the “state of play” at duke and opportunities for cross-campus collaborations." authored by sophia lafferty-hess, william krenzer, jenny ariansen, jennifer darragh. the paper describes a collaborative effort by the duke university libraries and the duke office of scientific integrity (dosi) to better understand the status of data management and sharing plans submitted from duke to the national science foundation (nsf) in 2021. it presents key findings from the data management and sharing plans (dmsp) assessment project and describes how this information can be used to plan ongoing education, training, and resource development using a cross-campus collaboration model. congratulations to this year's winning authors! the award for iassist conference paper competition winner includes free registration for the first author to the following year’s iassist conference, in addition to bragging rights. we are thankful to meryl brodsky, the competition coordinator, lauren phegley, last year’s paper competition winner, and rené duplain, who joined the journal’s editors in reviewing the competition submissions. 2/2 schwartz, ofira & hayslett, michele (2025) tides of transformation: enhancing research through data and community, iassist quarterly 49(2), pp. 1-2. doi: https://doi.org/10.29173/iq1170 we encourage all the paper competition authors to submit their manuscripts to iq and we look forward to publishing them in future iq issues. the four papers included in the current issue, iq 49(2), highlight innovative approaches to enhancing research practices and infrastructure across diverse domains within the social sciences. together, these works underscore the importance of data standardization, literacy, and community engagement in promoting collaboration and knowledge sharing. in their article “using common data elements to foster interoperability of research on health disparities,” megan chenoweth and john kubale discuss the use of common data elements (cdes) to support health disparities research in social, behavioral, and economic (sbe) research. the authors describe their process for identifying, validating, and building consensus on cdes related to covid public health policies. in “the role of fair principles in high-quality research data documentation: looking at national election studies”, author wolfgang zenk-möltgen uses national election studies as a case study to compare between manual and automatic assessments of fairness scores. the fair (findable, accessible, interoperable, and reusable) standards are among the core principles used for improving open science and research data management. technological developments may allow improvements of their application. the author compares the differences between manual and automated methods to assess fairness scores. margaret marchant and c. jeffrey belliston in their paper “data literacy in undergraduate research: a case study from student poster competitions” make use of information collected from undergraduate poster competitions, which is a common way for undergraduate students to share research, to examine data literacy skills of undergraduate students at brigham young university. the authors identify strengths and gaps in data literacy education and offer suggestions for supporting and encouraging undergraduate research and data literacy development beyond the traditional area of data analysis. the last article in this issue departs from our usual topics, but we felt it was important to include it. authors priya silverstein, julia g. bottesini, sebastian karcher, and colin elman, in their paper “introducing the journal editors discussion interface” (jedi), acquaint us with the online community they have helped found for journal editors in the social sciences, launched in 2021. jedi aims to increase uptake of open science at social science journals by providing journal editors with a space to learn and discuss. the paper explores jedi’s progress in its first two years, presenting data on membership, posts, and from a members’ survey. we, editors of the iq, found feedback from the jedi community to be not only informative and helpful but actually crucial at times since we became aware of it in march 2024. we publish it here not only as a window for readers into our world, but also as a potentially informative item for those of you who may be serving with other journals’ editorial staff or boards or have such people in your circles. we hope you enjoy reading and wish you a productive summer. michele hayslett and ofira schwartz, june 2025